Skip to content

WI-018: Test the API in the Workers runtime

WI-018: Test the API in the Workers runtime

Problem

WI-012 shipped with a gap it recorded rather than hid. Its tests used Hono’s app.request() from Node — real routing, real SQL, but not the Workers runtime. So they could not see the bug that had broken that item’s first draft, and the regression test written afterwards passed with the bug reinstated.

../standards/testing.md asked for @cloudflare/vitest-pool-workers from the start. This does it, and it should land before WI-013 adds authentication — a second-request failure in a login flow is a worse thing to discover.

In scope

  • API tests run inside workerd via @cloudflare/vitest-pool-workers.
  • The Hyperdrive binding resolved to local Postgres, the way wrangler dev does.
  • SELF.fetch rather than app.request(), so requests go through the Worker as a client’s would.
  • Ambient types for cloudflare:test, with ProvidedEnv declared as our own Env.
  • Wire the real Hyperdrive configuration id into wrangler.toml.

It found two bugs on the first run

1. sql.end() produces an unhandled rejection on every request. postgres.js’s Workers polyfill runs a floating read loop over the socket; closing the socket makes that loop reject, and nothing owns the rejection. The end() promise itself was already caught — the rejection comes from inside the driver, so it cannot be caught at the call site.

Isolated by variant, two requests:

unhandled rejections
waitUntil(sql.end())2
awaited end()2
no end()0

2. The fix I reached for first was also wrong. waitUntil is the right mechanism for work that should outlive a response, and it made no difference here, because the rejection is not the close — it is the read loop noticing the socket went away.

Not closing is not a leak, and it was measured

Dropping end() looks like leaking a connection per request. It is not, because the socket belongs to the request and Hyperdrive — not this client — owns the pool.

400 requests at concurrency 16 against wrangler dev:

peak connectionsmedianafter settle
waitUntil(sql.end())1011
no end()1211

Identical. The first attempt at this measurement sampled pg_stat_activity after the load finished, reported 0 for both, and discriminated nothing — worth recording, because it looked like an answer.

This would be the wrong call against a bare Postgres with no pooler in front. It is right here, and the reason is the binding.

Definition of done

  • API tests run in workerd with the Hyperdrive binding available.
  • Reinstating the module-scope client cache fails the suite — the thing the Node tests could not do.
  • Reinstating sql.end() surfaces the unhandled rejection rather than passing quietly.
  • No unhandled errors in a clean run.
  • Repeated and concurrent requests are covered explicitly, since single-request tests are what hid both bugs.
  • The real Hyperdrive id is in wrangler.toml; no connection string is in the repo.
  • Gate green.

Verification

Terminal window
pnpm --filter @tabletop/db run db:up
pnpm --filter @tabletop/api test
pnpm gate

Notes

The pool’s API changed in 0.19: defineWorkersConfig from @cloudflare/vitest-pool-workers/config, which nearly every example still shows, no longer exists — that subpath was removed. It is now cloudflareTest(), a Vite plugin, from the package root.