ADR-0030: How we read Wikipedia prose
ADR-0030: How we read Wikipedia prose
- Status: Accepted
- Date: 2026-09-02
- Supersedes: —
- Superseded by: —
Context
Every game page on the site carries a title, a byline, three numbers and no
prose. 0 of 333 live games have a description, and 0 of the 3973 harvested
records do either — description is not a key in the corpus schema.
WI-093 measured it; it is the thin-content
problem, on the pages ADR-0009 stakes the
whole web stack on ranking.
It also blocks enrichment. ADR-0017
derives mechanics and weight from prose, and promptFor() currently produces
a 52-token prompt of title, year, designers and Wikidata genres. With no prose
a model is not classifying supplied text, it is recalling the game from
training data — which is ADR-0017’s own named failure mode, and it is worst
exactly where the catalogue is longest.
ADR-0016 already names Wikipedia as the CC BY-SA supplement to Wikidata, and the attribution machinery is built: per-field provenance, and a CHECK constraint that refuses to store CC BY-SA material without a source URL. What was never decided is how to reach it.
Wikipedia’s robots.txt disallows both of its APIs. Measured:
User-agent: *Allow: /w/api.php?action=mobileview&Allow: /api/rest_v1/?docDisallow: /w/ <- the Action APIDisallow: /api/ <- REST v1PoliteClient refuses both, correctly and by design (ADR-0016: honour
robots.txt, never work around a block).
Decision
Read prose from api.wikimedia.org, the Wikimedia API gateway, and add no
robots exemption. Decided by Simon on 2026-09-02, on the reasoning below.
GET https://api.wikimedia.org/core/v1/wikipedia/en/page/{title}/bareMeasured against it: HTTP 200, and the response states its own licence —
"license": { "url": "https://creativecommons.org/licenses/by-sa/4.0/deed.en", "title": "Creative Commons Attribution-Share Alike 4.0"}That last part matters more than the convenience. The obligation under CC BY-SA is to name the licence, and here the source tells us what it is per page rather than us assuming it — the same discipline WI-047 applied to Commons, where assuming one licence for all files was the trap.
Rationale
It needs no exemption, and an exemption is the expensive kind of decision.
en.wikipedia.org’s robots.txt is addressed to crawlers walking article
space; api.wikimedia.org is the gateway Wikimedia publishes for programmatic
clients, with documented rate limits and an authentication story for higher
ones. Using the front door beats reasoning our way through a side one.
One exemption is a precedent; two is a habit. ADR-0023
took the one that exists, for query.wikidata.org/sparql, and the reason it is
defensible is that it is singular and argued. PoliteClient deliberately makes
an exemption cost four required fields so it cannot be added casually. A second
one, taken to save a hostname change, is how that becomes a formality.
What the gateway actually serves
Probed rather than assumed, because the obvious endpoint turned out to be the wrong one:
| Endpoint | Transfer (gzip) | Carries |
|---|---|---|
/page/{title}/bare | 332 B | metadata and licence — no prose |
/page/{title} | ~20 KB | wikitext and licence |
/page/{title}/html | ~61 KB | rendered HTML; a clean lead paragraph |
/search/page?q= | ~400 B | a one-line description, and a match-highlighted excerpt |
/bare is the one that looks right and is not. The lead paragraph comes from
the rendered HTML, where it is a plain <p> — extracting it from wikitext
would mean expanding templates, and the first 400 characters of any article’s
source are an infobox.
Compression is not optional. 399 KB plain against 61 KB gzipped for one article; across the catalogue that is the difference between roughly 1.5 GB and 250 MB of somebody else’s bandwidth.
Titles must be resolved, never guessed
Measured on four games: Catan and Gloomhaven resolve, Azul lands on a
disambiguation page and Brass: Birmingham 404s. A game’s name is not its
article title, and the failures are silent — a 171-byte response is not an
obvious error.
The SPARQL harvest already runs against query.wikidata.org under
ADR-0023’s exemption, and one OPTIONAL
clause returns the article:
OPTIONAL { ?article schema:about ?item ; schema:isPartOf <https://en.wikipedia.org/> . }which gives Root_(board_game) and Summoner_Wars_(card_game) — precisely the
disambiguation that guessing gets wrong. It costs no extra requests, because
the harvest is already paging that query.
How much prose this actually buys
45.3%. Of the 4271 items carrying a BoardGameGeek id, 1935 have an English Wikipedia article and 2336 do not.
That is worth being plain about before the work starts. It takes the catalogue from no game having prose to most of the ones anybody has heard of having it, and it leaves more than half the records exactly as thin as they are now. The long tail of a BGG-id harvest is obscure by construction, and no source this project is willing to use has prose about it.
Consequences
- A robots.txt that is not a robots.txt.
api.wikimedia.org/robots.txtanswers 301 to a documentation page.PoliteClient.rulesForfollows the redirect, gets a 200 of HTML, parses no rules from it, and allows everything. It reaches the right answer by accident — HTML is not permission. This ADR is not adopted until that is fixed: a response whose content-type is nottext/plainshould be treated as unknown, and unknown already fails closed. Recorded as the first item under it. - Title resolution is a second lookup. The corpus stores
wikidataIdand no sitelink, so the enwiki article title has to come from Wikidata’swbgetentities(props=sitelinks,sitefilter=enwiki), which is already permitted and already used. Roughly 80 batched calls for 3973 games, then ~3973 page reads at the client’s one-per-second pacing — about 70 minutes, once, offline, committed as a diff like the rest of the corpus (ADR-0027). - Not every game has an article. A Wikidata item with no enwiki sitelink gets no prose, and that is a correct outcome rather than an error. The proportion is unknown until the run; the harvest must report it rather than quietly produce a shorter corpus, which is the mistake WI-076 made twice.
- The corpus grows. 3973 extracts at a few hundred words each is single-digit megabytes on a 2.7 MB file. If that becomes uncomfortable the answer is a length cap per record, not a second store.
- CC BY-SA attribution becomes load-bearing on every game page. Today the credits block renders for almost nobody. With prose it renders for most of the catalogue, and the licence must be the one the API reported for that page — not a constant.
Alternatives considered
A recorded robots exemption for en.wikipedia.org/w/api.php. Defensible on
exactly ADR-0023’s logic: the directive addresses crawlers, and the Action API
is documented for programmatic use with published etiquette. Rejected because
the gateway exists and makes the argument unnecessary — and because the value
of ADR-0023’s exemption is that it is the only one.
Wikimedia dumps (dumps.wikimedia.org). The sanctioned path for bulk
access and unambiguously polite: no per-page requests at all. Rejected for
this size — the enwiki abstracts dump is gigabytes to get 3973 extracts, and
it needs a parse, a schedule and a store. Worth revisiting if the catalogue
grows an order of magnitude, where it becomes the only right answer.
Do nothing and enrich from titles alone. Costed: under £6 for the whole catalogue at the most capable tier. Rejected on quality, not price. It would buy 3973 confident guesses, and ADR-0017 is explicit that a plausible wrong tag is worse than a missing one because it looks identical in the UI.