Skip to content

ADR-0030: How we read Wikipedia prose

ADR-0030: How we read Wikipedia prose

  • Status: Accepted
  • Date: 2026-09-02
  • Supersedes: —
  • Superseded by: —

Context

Every game page on the site carries a title, a byline, three numbers and no prose. 0 of 333 live games have a description, and 0 of the 3973 harvested records do either — description is not a key in the corpus schema. WI-093 measured it; it is the thin-content problem, on the pages ADR-0009 stakes the whole web stack on ranking.

It also blocks enrichment. ADR-0017 derives mechanics and weight from prose, and promptFor() currently produces a 52-token prompt of title, year, designers and Wikidata genres. With no prose a model is not classifying supplied text, it is recalling the game from training data — which is ADR-0017’s own named failure mode, and it is worst exactly where the catalogue is longest.

ADR-0016 already names Wikipedia as the CC BY-SA supplement to Wikidata, and the attribution machinery is built: per-field provenance, and a CHECK constraint that refuses to store CC BY-SA material without a source URL. What was never decided is how to reach it.

Wikipedia’s robots.txt disallows both of its APIs. Measured:

User-agent: *
Allow: /w/api.php?action=mobileview&
Allow: /api/rest_v1/?doc
Disallow: /w/ <- the Action API
Disallow: /api/ <- REST v1

PoliteClient refuses both, correctly and by design (ADR-0016: honour robots.txt, never work around a block).

Decision

Read prose from api.wikimedia.org, the Wikimedia API gateway, and add no robots exemption. Decided by Simon on 2026-09-02, on the reasoning below.

GET https://api.wikimedia.org/core/v1/wikipedia/en/page/{title}/bare

Measured against it: HTTP 200, and the response states its own licence —

"license": {
"url": "https://creativecommons.org/licenses/by-sa/4.0/deed.en",
"title": "Creative Commons Attribution-Share Alike 4.0"
}

That last part matters more than the convenience. The obligation under CC BY-SA is to name the licence, and here the source tells us what it is per page rather than us assuming it — the same discipline WI-047 applied to Commons, where assuming one licence for all files was the trap.

Rationale

It needs no exemption, and an exemption is the expensive kind of decision. en.wikipedia.org’s robots.txt is addressed to crawlers walking article space; api.wikimedia.org is the gateway Wikimedia publishes for programmatic clients, with documented rate limits and an authentication story for higher ones. Using the front door beats reasoning our way through a side one.

One exemption is a precedent; two is a habit. ADR-0023 took the one that exists, for query.wikidata.org/sparql, and the reason it is defensible is that it is singular and argued. PoliteClient deliberately makes an exemption cost four required fields so it cannot be added casually. A second one, taken to save a hostname change, is how that becomes a formality.

What the gateway actually serves

Probed rather than assumed, because the obvious endpoint turned out to be the wrong one:

EndpointTransfer (gzip)Carries
/page/{title}/bare332 Bmetadata and licence — no prose
/page/{title}~20 KBwikitext and licence
/page/{title}/html~61 KBrendered HTML; a clean lead paragraph
/search/page?q=~400 Ba one-line description, and a match-highlighted excerpt

/bare is the one that looks right and is not. The lead paragraph comes from the rendered HTML, where it is a plain <p> — extracting it from wikitext would mean expanding templates, and the first 400 characters of any article’s source are an infobox.

Compression is not optional. 399 KB plain against 61 KB gzipped for one article; across the catalogue that is the difference between roughly 1.5 GB and 250 MB of somebody else’s bandwidth.

Titles must be resolved, never guessed

Measured on four games: Catan and Gloomhaven resolve, Azul lands on a disambiguation page and Brass: Birmingham 404s. A game’s name is not its article title, and the failures are silent — a 171-byte response is not an obvious error.

The SPARQL harvest already runs against query.wikidata.org under ADR-0023’s exemption, and one OPTIONAL clause returns the article:

OPTIONAL { ?article schema:about ?item ; schema:isPartOf <https://en.wikipedia.org/> . }

which gives Root_(board_game) and Summoner_Wars_(card_game) — precisely the disambiguation that guessing gets wrong. It costs no extra requests, because the harvest is already paging that query.

How much prose this actually buys

45.3%. Of the 4271 items carrying a BoardGameGeek id, 1935 have an English Wikipedia article and 2336 do not.

That is worth being plain about before the work starts. It takes the catalogue from no game having prose to most of the ones anybody has heard of having it, and it leaves more than half the records exactly as thin as they are now. The long tail of a BGG-id harvest is obscure by construction, and no source this project is willing to use has prose about it.

Consequences

  • A robots.txt that is not a robots.txt. api.wikimedia.org/robots.txt answers 301 to a documentation page. PoliteClient.rulesFor follows the redirect, gets a 200 of HTML, parses no rules from it, and allows everything. It reaches the right answer by accident — HTML is not permission. This ADR is not adopted until that is fixed: a response whose content-type is not text/plain should be treated as unknown, and unknown already fails closed. Recorded as the first item under it.
  • Title resolution is a second lookup. The corpus stores wikidataId and no sitelink, so the enwiki article title has to come from Wikidata’s wbgetentities (props=sitelinks, sitefilter=enwiki), which is already permitted and already used. Roughly 80 batched calls for 3973 games, then ~3973 page reads at the client’s one-per-second pacing — about 70 minutes, once, offline, committed as a diff like the rest of the corpus (ADR-0027).
  • Not every game has an article. A Wikidata item with no enwiki sitelink gets no prose, and that is a correct outcome rather than an error. The proportion is unknown until the run; the harvest must report it rather than quietly produce a shorter corpus, which is the mistake WI-076 made twice.
  • The corpus grows. 3973 extracts at a few hundred words each is single-digit megabytes on a 2.7 MB file. If that becomes uncomfortable the answer is a length cap per record, not a second store.
  • CC BY-SA attribution becomes load-bearing on every game page. Today the credits block renders for almost nobody. With prose it renders for most of the catalogue, and the licence must be the one the API reported for that page — not a constant.

Alternatives considered

A recorded robots exemption for en.wikipedia.org/w/api.php. Defensible on exactly ADR-0023’s logic: the directive addresses crawlers, and the Action API is documented for programmatic use with published etiquette. Rejected because the gateway exists and makes the argument unnecessary — and because the value of ADR-0023’s exemption is that it is the only one.

Wikimedia dumps (dumps.wikimedia.org). The sanctioned path for bulk access and unambiguously polite: no per-page requests at all. Rejected for this size — the enwiki abstracts dump is gigabytes to get 3973 extracts, and it needs a parse, a schedule and a store. Worth revisiting if the catalogue grows an order of magnitude, where it becomes the only right answer.

Do nothing and enrich from titles alone. Costed: under £6 for the whole catalogue at the most capable tier. Rejected on quality, not price. It would buy 3973 confident guesses, and ADR-0017 is explicit that a plausible wrong tag is worse than a missing one because it looks identical in the UI.