Skip to content
COMP10001Playground
All decision records

Decision record · DR-005

Default the LLM evaluation to 30 hands and label its gap intervals as nominal

Status
Accepted
Date
2026-10
Applies to
/card-table/llm-eval, web/src/lib/go/llm-eval.ts, web/scripts/interval-coverage.sim.ts, docs/llm-evaluation.md

Decision in one line

The evaluation now deals 30 hands by default and offers 10 to 50, labels its bootstrap intervals for mean gaps "nominal 95%", says "no interval" when every hand has the same gap, and publishes a coverage simulation that shows how far short of 95% those intervals fall. DR-004's other decisions stand.

Amends DR-004.

Context

DR-004 set a default of 10 hands and reported the mean gap to the best score with a percentile bootstrap interval labelled 95%. A review of the upgrade asked whether that label was earned. The gaps here are the hard case for a bootstrap: most hands have a gap of 0 and a few have a large one, and with 10 hands the resamples cannot see the tail they have not drawn. If the interval is too narrow, the page overstates how much a run shows, which is the opposite of what the evaluation is for.

Decision

  • Default to 30 hands. Offer 10, 20, 30 or 50; drop 5.
  • Label the gap intervals "nominal 95%" on the page, in the export and in the evaluation card, and state the simulated coverage next to them. Keep Wilson intervals for rates unchanged.
  • When every hand in a sample has the same gap, print "same on all n hands; no interval" instead of a zero-width interval that reads as certainty.
  • Keep the simulation in the repository (cd web && pnpm sim:coverage) and quote its results in docs/llm-evaluation.md; a test checks the quoted numbers against its output.

Options considered

  1. Keep 10 hands and add a caveat. Cheapest for visitors, but the default run would still give an interval that misses about one time in five.
  2. Switch to a t interval or BCa. Possibly better, but I had not measured either on this data, and swapping methods without a check repeats the original mistake.
  3. 30 hands, a nominal label and a published simulation (chosen).
  4. 50 hands by default. Better coverage again, but longer runs and higher cost for every visitor, for a few points of coverage.

Why

Thirty calls to Claude Haiku 4.5 cost under 20 cents at list price even if every reply used its full output limit, and the page shows that bound before a run starts. That buys most of the available improvement. Labelling the interval as nominal and showing the simulated coverage is honest about the rest, and the simulation gives a way to test any better method before switching to it.

What happened

  • The simulation drew 4,000 evaluations per hand count from 3,000 dealt hands, with the greedy baseline standing in for a model. The nominal 95% gap interval held the true mean 69% of the time at 5 hands, 82% at 10, 88% at 20, 91% at 30 and 92% at 50. At 10 hands, 3% of samples had a gap of 0 on every hand, so the interval collapsed to a point.
  • For the all-singletons player, whose gaps are not zero-heavy, coverage was 90% at 10 hands and 94% at 30. The shape of the data matters more than the method's label.
  • The Wilson intervals for rates held their level: 93% to 97% exact coverage at the greedy baseline's optimality rate.
  • On the default set, the greedy optimality interval narrowed from 31% to 83% (6 of the first 10 hands) to 56% to 86% (22 of 30). Even at 30 hands the gap intervals are somewhat too narrow, and the page says so rather than claiming 95%.
  • I did not simulate the paired LLM-minus-greedy interval; it uses the same method and is labelled nominal for the same reason.

What I'd change

  • Measure BCa and bootstrap-t intervals in the same simulation, and switch if one holds its level better on this data.
  • Simulate the paired difference interval, not just the single-player mean.
  • Once real model runs exist, rerun the simulation with a model's own gaps as the population.
  • Report the share of hands with a gap of 0 next to the mean gap, since the mean hides how zero-heavy the data is.