WI-073: A ceiling that cannot be talked round
WI-073: A ceiling that cannot be talked round
Wikidata models genre, not mechanics, and nothing free and open carries BGG-style complexity. Deriving those is the one job the structured sources genuinely cannot do, which is what ADR-0017 is for.
Nothing here has spent anything. The model is an interface with a stub behind it, deliberately: the ceiling and the confidence gate are the feature, and both are provably right before a single paid call.
Two ceilings, because they fail differently
A cost cap alone stops a run that has become expensive. A record cap alone stops a run that has become long. A pricing surprise breaks the first and a pagination bug breaks the second, and neither is a reason to keep going — so Ledger enforces both, in integer pence, and refuses to be constructed with a budget of nothing.
The check happens before the call. Noticing afterwards is not a ceiling; the money is already gone.
The gate is per field, not per record
ADR-0017’s failure mode is specific: a plausible, confidently-wrong mechanic tag is worse than a missing one. A missing tag makes a game harder to find; a wrong one actively misleads and looks identical to a right one on screen.
So anything at or below the threshold goes to review rather than the catalogue — per field, because a model can be certain about a game’s theme and guessing about its mechanics, and discarding the certain half wastes the call. A field the model did not rate at all is treated as unvouched-for: no score is not a high score.
An enumerated vocabulary, not free text
Asked for “the mechanics”, a model returns worker placement, Worker Placement, worker-placement and placing workers across four games, and discovery then filters on a field where nothing matches anything. The schema enumerates 29 mechanics and 23 themes; anything else fails validation and is reported.
The prompt says in terms that an empty list is a correct answer. Without that the model fills the field, which is the entire failure mode.
Definition of done
- A budget of nothing is refused.
- The cost ceiling stops a run.
- The record ceiling stops a run that still has money.
- Affordability is checked before the call, not after.
- Low-confidence values go to review, not the catalogue.
- Gating is per field.
- An unrated field is not published.
- A vocabulary violation is rejected.
- A paid-for unusable answer is reported, not dropped.
- Measured cost per record reported, per WI-H05.
- Each check proven able to fail.
- Gate green.
Verification
pnpm --filter @tabletop/catalogue-forge testpnpm gateSeeded failures
| Seed | Bit |
|---|---|
| Check the budget after the call instead of before | 1 |
| Cost ceiling only, no record ceiling | 1 |
| An unrated field counts as confident | 1 |
| Gate per record rather than per field | 1 |
| Publish everything regardless of confidence | 3 |
| Free-text mechanics instead of the vocabulary | 1 |
| Silently drop an answer that will not parse | 1 |
Not done here
No live model, and no spend. Wiring one needs an API key, which is WI-H04 and a person’s to provide. When it arrives the pilot is one command against a stated ceiling, and it reports measured cost per game — which is what WI-H05 asks for and what a full-catalogue decision should rest on.
Worth saying before that happens: the catalogue currently has themes from Wikidata genres and no mechanics at all, and no descriptions. That is thin input. Enrichment will do better with prose than without it, and prose is parked.