Decision record · DR-004
Evaluate an LLM on the Card Table with the visitor's own key, from the browser, with a local audit log
- Status
- Accepted
- Date
- 2026-10
- Applies to
- /card-table/llm-eval, /ai-log, web/src/lib/ai, web/src/lib/go/llm-eval.ts
Decision in one line
The LLM evaluation is optional: it calls Anthropic or OpenAI directly from the visitor's browser with a key they paste in, scores every answer against the exact solver and two baselines with intervals, writes every call to an audit log on the visitor's device, and the site works fully without a key.
Amended by DR-005 and DR-006, which build on or change part of this decision; this record is kept as it was written.
Context
The site is a static Next.js app with no budget for AI calls. The best-partition question is a clean test of whether a language model can reason under hard combinatorial rules, because the exact answer is known for every hand. I wanted to show how I evaluate a model against a classical method, and how I keep a record of what an AI system was asked, what it answered and what a person decided about it. A shared key held by the site would cost money with every visit and would need abuse protection and secret management.
Decision
- The visitor picks Anthropic (default: Claude Haiku 4.5 at temperature 0, with Claude Sonnet 5.5 and a low, medium or high thinking effort as the option) or OpenAI (model id is free text) and pastes their own key. The key is kept in sessionStorage, moves to localStorage only if they turn on "remember on this device", and "Forget keys" clears both.
- Requests go from the browser straight to the provider. The site has no AI server code and never receives the key.
- The model sees the rules in my own words and one hand, and must reply with JSON matching a schema whose card list is that hand. zod checks the shape; the rules engine then judges the grouping. A reply that fails either check is recorded as such, never repaired.
- Every player (exact solver, greedy baseline, all singletons, LLM) is scored on the same seeded hands. Rates get Wilson 95% intervals; the mean gap to the best score gets a percentile bootstrap interval (10,000 resamples, seed 10001). An answer that is not a valid grouping is scored as the hand played as singletons, so failures count against the model instead of dropping out. The LLM is also compared with the greedy baseline hand by hand (paired bootstrap).
- A refused, truncated or malformed reply is a model outcome and is scored. A failed call (rate limit after retries, network) is not the model's answer: it is left out of the scores and reported separately. Errors that would repeat on every hand (bad key, unknown model) stop the run.
- No silent fallback to another model: if a provider declines, the hand is recorded as refused, so every score belongs to the model named next to it.
- Every call is appended to an audit log in IndexedDB (time, feature, provider, model, full
prompt, raw and parsed output, validation result, stop reason, latency, tokens, human
decision), scrubbed for the key before it is stored, viewable and exportable at
/ai-log. The visitor can accept or reject each answer, and that decision is stored with the call. - Every AI output carries an "AI-generated" label.
Options considered
- A server route with my key. Every visit would cost me money and need rate limits.
- A server route with the visitor's key. The key would pass through a server I run.
- Browser-direct calls with the visitor's key (chosen).
- No AI features. Simplest, but it leaves out the evaluation and audit work the upgrade is meant to show.
Why
Option 3 costs nothing to host, keeps the key between the visitor and their provider, and keeps the audit record on the visitor's device, where they can inspect and export it. Scoring against the exact solver means there is no judge model and no rubric to argue about: an answer is legal and optimal or it is not.
What happened
- Anthropic accepts browser calls only with the
anthropic-dangerous-direct-browser-accessheader. The name is a fair warning: any script on the page, including a browser extension, could read a key held there. The site loads no third-party scripts and keeps the key in sessionStorage by default, but it cannot protect a key from the visitor's own extensions. - The adapters, the validation, the retries, the judging and the audit log are covered by
unit tests with mocked
fetchresponses. No paid run is recorded in this repository, because a run needs a key and spends money. The page reports results only for runs made in the visitor's browser. - The baselines already make the point about uncertainty. On the default ten hands (seed 2019, ten cards each) the greedy baseline is optimal on 6 of 10, a Wilson interval of 31% to 83%, and its mean gap to the best score is 10.8 points (bootstrap interval 0.9 to 23.1). Ten hands cannot separate "usually optimal" from "optimal a third of the time", so the page offers up to 30 hands and says plainly that the intervals are wide.
What I'd change
- Run each hand several times, so run-to-run variation is measured as well as the mean.
- Try a second prompt (for example, one that lists the legal groups first) and report the difference as a paired comparison on the same hands.
- Test a repair loop (send the validation error back once) as a separate, labelled player, never mixed into the single-shot scores.
- Record the provider's request id in each audit entry, so an entry can be matched to the provider's usage logs.