Skip to content

ADR-0016: Open-data catalogue seeding with per-record provenance

ADR-0016: Open-data catalogue seeding with per-record provenance

  • Status: Accepted
  • Date: 2026-07-30
  • Supersedes: —
  • Superseded by: —

Context

ADR-0010 put the catalogue behind a CatalogueSource interface with a hand-seeded fixture, so that the unresolved BGG licensing question (WI-H01) blocks only the BGG sync worker.

But a hand-seeded fixture doesn’t scale past a few hundred games, and BGG may say no. The project needs a catalogue that can be built from sources whose terms are unambiguous.

Decision

Build a standalone tool, tools/catalogue-forge, that harvests from open and freely-licensed sources and emits seed data behind the existing CatalogueSource seam.

Source hierarchy:

SourceLicenceProvides
Wikidata (primary)CC0 — public domainStructured: title, designers, publishers, year, genre, BGG ID. SPARQL endpoint, no permission needed.
Wikipedia (secondary)CC BY-SA — attribution requiredProse descriptions, and the raw text the enrichment stage reads.
BGGPermission requiredOnly after WI-H01 returns in writing.

Every record carries provenance. Each field records where it came from, under what licence, and when it was fetched:

game_field_provenance: (game_id, field, source, source_url, licence, fetched_at, confidence)

Rationale

Wikidata is the answer to the BGG problem. It is CC0 — public domain, no attribution obligation, no commercial restriction, no permission to request — and it is genuinely structured, with board games modelled as first-class entities carrying designer, publisher, and publication-year properties, plus BGG IDs as external identifiers. It is the only source in this space whose terms cannot be withdrawn.

Wikipedia fills the prose gap but under CC BY-SA, which does carry an attribution obligation. That difference is exactly why provenance has to be per-field rather than per-record: a single game row can legitimately mix a CC0 designer list with a CC BY-SA description, and only one of those needs a credit line rendered on the page.

Provenance also makes the catalogue auditable. When a field is wrong, we can see which source produced it — and when a source’s terms change, we can find and purge exactly what came from it.

Amendment, 2026-08-04 (ADR-0023)

https://query.wikidata.org/robots.txt disallows /sparql to every user-agent. The conduct rule above and the mechanism named above therefore contradict each other, and ADR-0023 resolves it: the Query Service is used, because robots.txt is a crawler-control protocol while WDQS separately documents programmatic access, publishes rate limits, and asks for exactly the descriptive User-Agent this project sends.

“Honour robots.txt” still stands everywhere else, and the exception is not a general one. It is expressed in code as a RobotsExemption carrying an origin, a path prefix, a reason and the ADR that authorised it — all four required, or the client refuses to be constructed. query.wikidata.org/sparql is the only one, and it covers /sparql alone: /bigdata on the same host remains refused.

Second amendment, 2026-08-04 (WI-047): images

“Prefer bulk exports and structured endpoints over scraping rendered pages” inverts for Wikimedia Commons image licences, and the exception is recorded rather than taken quietly.

A Commons file is licensed file by file. The permitted structured endpoint, api.wikimedia.org/core/v1/commons/file/…, returns the file, its uploader and its image URLs and no licence at all. Everything that does carry one — /w/api.php, /w/rest.php, Commons’ Special:EntityData — is Disallowed, where Wikidata’s equivalent is explicitly allowed.

The file page is allowed, and Commons’ licence templates publish machine-readable fields on it precisely for reusers:

<span class="licensetpl_short">CC BY-SA 2.5</span>
<span class="licensetpl_link">https://creativecommons.org/licenses/by-sa/2.5</span>

So the rendered page is read, for the licence, the licence URL and the author — and no image is used whose licence cannot be read. That is the same rule as the CC BY-SA text: material that cannot be credited is not stored.

Consequences

  • Attribution is a rendering requirement, not a footnote. Any page showing a CC BY-SA-derived field must credit Wikipedia and link the source. The provenance table is what makes that automatic.
  • The fixture implementation from ADR-0010 becomes generated rather than hand-written. The interface is unchanged.
  • The tool is standalone and offline — tools/ per docs/standards/engineering-standards.md, never deployed, never in the request path. It cannot touch the 10ms Workers CPU ceiling because it never runs on a Worker.
  • A source whose terms are unclear is a needs:human decision, not an agent’s call — see docs/standards/agent-boundaries.md.
  • Scraping conduct is binding: honour robots.txt, prefer bulk exports and structured endpoints over scraping rendered pages, identify the client in the User-Agent, and never work around a block.
  • If BGG later says yes, it becomes a third source in the same pipeline with its own licence tag — not a rewrite.

Alternatives considered

BGG XML API as the primary source. Richest data by far. Rejected as primary: its terms are permissive for non-commercial use only and require permission for commercial use, which is the exact risk docs/architecture.md §4 calls “the one that could sink the project”. It remains a candidate secondary source pending WI-H01.

Hand-seeded catalogue only. Zero legal risk and zero cost. Rejected: it caps the catalogue at a few hundred games, which is too thin for discovery and search to be meaningfully exercised.

A third-party board game API. Rejected on the grounds already recorded in docs/architecture.md §4 — this category has a history of shutting down.