Skip to content

WI-061: Pages a crawler can read

WI-061: Pages a crawler can read

The last piece of the original architecture that did not exist at all. ADR-0009: a Letterboxd-alike lives or dies on its game pages ranking, and organic search is the only free acquisition channel this project has. A client-rendered shell forfeits that.

Server-rendered on Workers, and — for now — with no islands at all. Everything on a game page is public and static per game, and a crawler has no session. The interactive parts stay in the Expo app.

GET / 200 Tabletop Tracker — a record of the games you actually played
GET /games/catan 200 Catan (1995) — Tabletop Tracker
GET /games/azul 200 Azul (2017) — Tabletop Tracker
GET /games/not-a-game 404

The metadata is the feature

pageTitle, metaDescription, structuredData and isIndexable are pure functions with tests, because they are what a search result is made of rather than plumbing around it.

  • The snippet is built from facts, not prose. Most of the catalogue has no description, and a snippet present on a tenth of pages and missing from the rest is worse than one always there.
  • Game, not Product. This is not a shop, and claiming otherwise invites a rich result with a missing price.
  • Nothing is claimed that is not known. An empty author list is an assertion that nobody designed it.
  • Thin pages opt out. Fewer than three known facts and the page carries noindex, follow — a site full of thin pages ranks worse overall than a smaller site of good ones.

Attribution renders here too, per WI-072: CC BY-SA credits are a rendering obligation on every surface, not just the app.

Definition of done

  • Game pages render on the server, with title, description, canonical and JSON-LD.
  • A missing game is a 404, not an empty page.
  • Thin pages are not indexable.
  • Attribution renders where it is obliged.
  • Each check proven able to fail.
  • Gate green.

Verification

Terminal window
pnpm --filter @tabletop/web test
pnpm --filter @tabletop/web build && pnpm --filter @tabletop/web exec wrangler dev
pnpm gate

Seeded failures

SeedBit
A thin page is indexed anyway1
Claim an author for a game with no designers1
Describe it as a Product1
Truncate mid-word1
Drop the year from the title1
Playing time as a bare number, not a duration1

Six seeds that proved nothing, and one that lied

The first run of all six reported “no tests” and I nearly recorded six passes: the suite never collected, because seo.ts was written to src/lib/ while the test imported ../src/seo. A seed run against a suite that does not run is worse than no seed run, because it produces a green table.

Then the truncation seed still did not bite. The test used evenly-sized words, so the 155-character cut landed on a space by luck and passed whether or not the boundary logic existed. Rewritten as a property — the kept text is a whole-word prefix, checked against a rebuilt original — and it bites. My first replacement was also wrong, asserting the last word is not one or two characters, which fails on correct code the moment the text contains “an”.

Two versions I guessed at

@astrojs/cloudflare@12.6.10 peers on Astro ^5.7.0; this repo is on 7.1.5. The build succeeded with an IMPORT_IS_UNDEFINED warning about the adapter’s default export, and the worker then 500’d on every request. The correct version is 14.1.7, which in turn needs wrangler ^4.118.0 — pinned in apps/web alone, leaving the API’s 4.115.0 untouched rather than bumping its toolchain inside a web item.

Astro.locals.runtime.env was also wrong: removed in Astro v6. The API URL is public, so it is read from build-time PUBLIC_API_URL instead, which works in dev, test and production alike.

Both were caught by running the built worker, not by the build passing. A build that succeeds and a worker that serves are different claims.

Not done here

  • Deployment. The site builds and runs locally; putting it on a URL wants a Cloudflare project and is worth doing as its own step.
  • Search and browse pages. The client exists (searchGames); the pages do not.
  • A sitemap, which robots.txt already points at.