Skip to content
COMP10001Playground

Methods

How the playground is checked

What is my original 2019 code and what was rebuilt, how each piece is tested, how the LLM evaluation is designed, and the decisions behind all of it, including the weak spots.

1 · Provenance

What is original and what was rebuilt

Only my Project 2 file survived from 2019. Everything else on the site is labelled with where it came from.

  • Rebuilt from the task (original lost)

    Ballot Box (Project 1, eVoting)

    • first_past_the_post
    • second_preference
    • multiple_preferences
    • is_valid_vote

    Checked by: The task's worked examples; 300 seeded elections and 200 ballots against the 2022 notebook code; four property-based tests.

  • Original 2019 code, ported line by line

    Falca's Cave (Project 2, Toy World), “As submitted” mode

    • build_cave
    • check_path
    • shortest_path
    • optimal_path

    Checked by: My unmodified 2019 file on 230 seeded caves (2 samples, 8 hand-made, 220 random), crashes included. CI checks the file still prints 23 and 14.

  • 2019 algorithms, bugs fixed in 2026

    Falca's Cave, “Spec-correct” mode

    • build_cave
    • check_path
    • shortest_path
    • optimal_path

    Checked by: The 2022 notebook code on the same 230 caves (one intended difference: at most three treasures); three property-based tests against brute force.

  • Rebuilt from the task (original lost)

    Card Table (Project 3, COMP10001-Go)

    • comp10001go_score_group
    • comp10001go_valid_groups
    • comp10001go_best_partitions

    Checked by: 500 groups, 200 groupings and 18 full best-partition searches against the 2022 notebook code; three property-based tests.

  • New in 2026, not coursework

    Evaluation and governance layer

    • greedy baseline
    • LLM harness
    • statistics helpers
    • AI audit log

    Checked by: Unit tests; statistics against statsmodels, SciPy, NumPy and R; AI calls with mocked network responses.

The decisions behind this split are DR-001 (rebuild lost work from the task) and DR-002 (keep the 2019 code as submitted). The original file is kept unchanged in the repository's coursework folder; the repository is private for now (DR-003).

2 · Verification

How the code is checked

  1. 1

    Worked examples

    Every example the tasks worked through is a unit test, including the two sample caves in my 2019 file (23 and 14 moves).

  2. 2

    Parity with Python

    scripts/generate_parity_fixtures.py runs the original Python on seeded inputs and records the answers. The TypeScript must reproduce every one. CI regenerates the 2019 fixture from my unmodified file and fails on any difference.

  3. 3

    Property-based tests

    fast-check generates 3,550 cases per CI run from fixed seeds and checks invariants against brute force, independent enumerations and recounts. Each suite asserts it generated exactly the number of cases listed below.

  4. 4

    Statistics against references

    Wilson intervals, normal quantiles, sample quantiles and percentile bootstrap intervals are tested against statsmodels, SciPy and NumPy (scripts/stats_reference.py), with the Wilson intervals also checked against R's prop.test. A coverage simulation checks how often each interval actually holds the truth at the sample sizes used.

3 · Property-based tests

3,550 generated cases, fixed seeds

Each property states an invariant, generates random inputs, and compares the code with an independent oracle. The counts and seeds below are read from the same file the tests use, and each test fails if fast-check generates a different number of cases.

Ballot Box

4 properties · 1,800 cases · src/lib/evoting/count.property.test.ts
  • First past the post names a unique top scorer or "tie" exactly when the top count is shared, and the answer does not depend on the order the ballots arrive in.

    Direct recount of the votes. 1 to 40 single-choice votes over 2 to 6 candidates.

    500 cases · seed 101

  • Second preference: an outright winner holds more than half the first choices; otherwise the eliminated candidate has the fewest first choices (ties go to the alphabetically first name), and no ballot is lost or double-counted in the transfer.

    Direct recount and conservation of ballots. 1 to 40 two-choice ballots over 2 to 6 candidates.

    500 cases · seed 102

  • Instant runoff: a declared winner holds a majority of the ballots still in play in the final round, each round eliminates the lowest candidate (ties to the alphabetically first name), ballots are conserved every round (counted plus exhausted equals cast), an eliminated candidate never returns, and ballot order does not matter.

    Round-by-round recount. 0 to 40 full or partial rankings over 2 to 6 candidates.

    500 cases · seed 103

  • is_valid_vote accepts every full ranking and rejects every malformed ballot (missing, repeated or unknown names, non-strings, not a list), so a count screened with it gives the same result as counting the valid ballots alone. This checks the validator: the Ballot Box itself counts partial rankings as written.

    Count of the valid ballots only. Valid rankings plus 1 to 10 deliberately broken ballots.

    300 cases · seed 104

Falca's Cave

3 properties · 800 cases · src/lib/cave/search.property.test.ts
  • Spec-correct shortest_path (breadth-first search) equals the shortest of all simple routes found by exhaustive depth-first enumeration, with and without the sword.

    Exhaustive enumeration of simple paths. Random valid caves of 3×3 to 5×5, random start and end squares.

    300 cases · seed 201

  • Spec-correct optimal_path (uniform-cost search) never exceeds the cost of any enumerated order of visiting the treasures, with or without fetching the sword; it equals the cheapest one, matches a full state-space search, and the route it draws passes check_path in exactly that many moves.

    Every waypoint order, plus breadth-first search over (square, sword, treasures). Random valid caves of 3×3 to 5×5 with 0 to 3 treasures.

    200 cases · seed 202

  • My 2019 shortest_path only guards the dragon's own square, so it is never longer than the spec-correct answer, finds a route whenever the spec does, and agrees exactly once Falca holds the sword.

    Spec-correct shortest_path. Random valid caves of 3×3 to 6×6, random distinct start and end squares.

    300 cases · seed 203

Card Table

3 properties · 950 cases · src/lib/go/partition.property.test.ts
  • comp10001go_best_partitions returns the true best score, every tied optimum and nothing else: each answer uses every card once, uses only legal groups and scores the best total.

    Independent enumeration of every set partition (restricted growth strings). Hands of 1 to 7 cards, half drawn from a narrow band of values so runs and sets occur.

    300 cases · seed 301

  • No legal grouping of a hand, whether random, all singletons or the greedy baseline's, ever scores more than the solver's best.

    Scores of random legal groupings and the greedy baseline. Hands of 2 to 10 cards with a random grouping of each.

    150 cases · seed 302

  • Group scores do not depend on card order; n of a kind scores value × n!; a constructed alternating-colour run (Aces filling inside gaps) scores the sum of its values; and two same-colour neighbours make it illegal.

    Scores computed from the rules by construction. Constructed sets and runs, with wild Aces substituted at random.

    500 cases · seed 303

To check the properties can fail, I broke the code on purpose: dropping the alphabetical tie-break, loosening the majority rule, letting Falca walk next to the dragon, skipping partitions and ignoring run colours each produced a shrunk counterexample within a few generated cases.

4 · Evaluation design

Can an LLM find the best partition?

The evaluation card for the Card Table evaluation. Rates use Wilson 95% intervals; gaps use percentile bootstrap intervals with 10,000 resamples and seed 10001, labelled nominal 95% because a coverage simulation showed them running narrow at these sample sizes.

This card describes the evaluation harness at /card-table/llm-eval. It evaluates a general language model on one narrow task with a known right answer. It is not a model card for any model, and it makes no claim about how a model behaves outside this task.

The question

COMP10001-Go's bonus question asked for every highest-scoring way to split a hand of up to ten cards into groups. The exact solver answers it by checking every set partition (115,975 for ten cards). The question here is how often a language model, given the rules and a hand, returns a grouping that is legal, and how often it is the best one.

Players

PlayerWhat it doesRole
Exact solvercomp10001go_best_partitions: checks every partitionThe reference; optimal by construction
Greedy baselinePuts down the best legal group left, and repeatsA cheap heuristic a person might use
All singletonsPlays every card on its ownThe floor
LLMThe visitor's chosen model, on their own keyThe system under test

Inputs

  • Hands are dealt by the same seeded dealer as the Card Table: hand 1 of a set is deal(seed), hand 2 is deal(seed + 1), and so on. The default set is 30 hands of 10 cards from seed 2019; the page offers 10 to 50 hands of 6 to 10 cards and any seed. The seed is shown next to every hand. The default was 10 hands until DR-005, which explains the change.
  • Every player sees exactly the same hands, so comparisons are paired.

What the model is given

  • A system prompt with the scoring rules in my own words (the task text is not reproduced), the card notation and the instruction to use every card exactly once.
  • One hand per request, and a JSON schema whose card list is that hand. The reply must be {"groups": [[...], ...], "claimed_score": n}.
  • Claude Haiku 4.5 (the default) runs at temperature 0 with up to 1,024 output tokens. Claude Sonnet 5.5 always thinks before answering; it runs at the effort the visitor picks (low by default) with up to 8,192 output tokens, because thinking counts against the limit. OpenAI models run at temperature 0 unless they are reasoning models, which do not accept it.
  • Hands are sent one at a time, and nothing from one hand is shown to the model for the next.

How an answer is scored

  • Valid: the reply matches the schema, uses every card in the hand exactly once, and every group of two or more cards is legal under the rules.
  • Optimal: valid, and it scores the exact solver's best total.
  • Score with fallback: the answer's score if valid; otherwise the hand played as singletons, which is what a careful player falls back to after rejecting an answer.
  • Gap: best score minus the score with fallback. It is never negative, and a failed answer counts against it instead of being dropped.
  • Claimed total: among valid answers, whether the model's own claimed_score equals the real score. A model's account of its own work is checked, not trusted.
  • Outcomes: answered (then judged), malformed reply, cut off at the token limit, refused, or failed call. A failed call (rate limit after retries, network) is not the model's answer, so it is left out of every rate and reported separately. When calls fail, every player on the scoreboard is scored on the hands the model was scored on, so the rows stay comparable; the export also has the baselines on the full set.
  • Latency: the response time of the final attempt. Retries and back-off waits are left out of the evaluation's latency figures; the audit log records both.

Uncertainty

  • Validity and optimality rates carry Wilson score 95% intervals, which behave well for small samples and for rates near 0% or 100%.
  • Mean gaps carry percentile bootstrap intervals, labelled nominal 95%: 10,000 resamples of hands, seed 10001. They run narrow at these sample sizes (see the coverage table below). When every hand in a sample has the same gap, the resamples cannot vary, so the page says "no interval" instead of printing a zero-width one.
  • The LLM and the greedy baseline are compared hand by hand: the mean of (LLM score minus greedy score) with a paired bootstrap interval, plus how many hands each was ahead on.
  • The statistics helpers are tested against statsmodels, SciPy and NumPy, and the Wilson intervals also against R's prop.test.

How often the intervals hold the truth

"95%" is a property of a method in the long run, and at small samples it can fall short. I checked both interval methods on the kind of data they are used for here, with cd web && pnpm sim:coverage (web/scripts/interval-coverage.sim.ts). The population is 3,000 hands of 10 cards dealt by the same dealer (seeds 100,000 to 102,999, apart from the evaluation's own seeds), solved exactly. The greedy baseline stands in for a model, because its gaps have the awkward shape a good model's would: 71% of hands have a gap of 0 and the rest spread out as far as 101 points. Each bootstrap cell is 4,000 simulated evaluations, so it carries about ±1.5 percentage points of Monte Carlo error. The Wilson column is exact.

HandsMean gap, greedy (nominal 95%)Mean gap, singletons (nominal 95%)Greedy optimality rate (Wilson 95%)
569%84%97%
1082%90%93%
2088%93%95%
3091%94%96%
5092%94%94%

The Wilson intervals hold their level. The bootstrap intervals for a mean gap do not: with 10 hands of zero-heavy gaps, a "95%" interval missed the true mean about one time in five, and in 3% of samples every gap was 0, so the interval collapsed to a point. That is why the default is now 30 hands, 5 hands is no longer offered, and the page labels the gap intervals "nominal 95%". Even at 30 hands they are somewhat too narrow, and the scoreboard says so.

Baselines on the default set

On the default 30 hands (seed 2019, 10 cards each), the greedy baseline is always valid and is optimal on 73% (56% to 86%) of hands. Its mean gap to the best score is 5.6 points (1.8 to 10.6). Playing every card alone is never optimal, 0% (0% to 11%), with a mean gap of 99.2 points (81.4 to 116.1).

On the first 10 of these hands, the old default, the same greedy baseline was optimal on 6, an interval of 31% to 83%. Thirty hands narrow that to 56% to 86%: enough to tell a strong player from a weak one, still not enough to rank two close ones.

What the judge can tell apart

  • Cards left out, used twice, or not in the hand at all.
  • Groups that break the rules: a run whose colours do not alternate, an Ace at the end of a run instead of inside it, a run with fewer than two non-Ace cards, or Aces as a set.
  • A legal grouping that is not the best one, and by how many points.
  • A claimed total that does not match the real score of the grouping.

Assumptions

  • The rules engine is right. It is checked against the 2022 reference notebooks (500 groups, 200 groupings, 18 best-partition searches) and by property-based tests against an independent exhaustive enumeration.
  • Hands dealt from a shuffled deck are a fair test. They are not weighted towards hard hands; many random hands have little to group, which favours the greedy baseline.
  • One request per hand, with no retries on a wrong answer, is the setting being measured.

Limitations

  • Each hand is asked once, so run-to-run variation is not measured.
  • The gap intervals are percentile bootstrap intervals on skewed data, and in the coverage simulation they held the true mean less often than their nominal 95% (91% at 30 hands). The greedy baseline's gaps stand in for a model's; a model whose gaps are shaped differently would get different coverage.
  • One prompt is used. A different wording could change the results; that is not measured.
  • Results are only as recent as the run: providers update models behind the same id.
  • Nothing here measures anything other than this card game; it says little about a model in general.

Audit and review

Every call is written to the AI audit log in the visitor's browser (/ai-log) with the full prompt, the JSON schema and generation settings (temperature, output-token limit, thinking effort), the raw and parsed reply, the validation result, the stop reason and any refusal category, latency (final attempt and total), token usage, and the visitor's accept or reject decision with an optional note. Results can be exported as CSV or JSON from the evaluation page, and the log from /ai-log; the evaluation's export reads decisions from the log, so a decision changed later is exported as it stands.

5 · Honesty

Assumptions, limits and what I'd change

Assumptions

  • The task paraphrases in coursework/README.md capture the rules the projects were marked on. The specifications themselves are not reproduced, so a reader cannot check the paraphrase against them here.
  • The 2022 notebook code is a useful oracle even though its provenance is unclear: it was written independently of the rebuilds, and agreement on hundreds of seeded inputs is evidence about the rebuilds, not about the notebooks.
  • Random inputs (elections from a left-right voter model, caves with random walls, hands from a shuffled deck) are representative enough to exercise the rules. They are not weighted towards the hardest cases.

Limitations

  • The rebuilt Projects 1 and 3 can match the task and still differ from what I submitted in 2019. Nothing here can recover the original marks or code.
  • Property-based tests check the invariants I thought to write down. A rule nobody encoded can still be wrong, and passing tests are evidence, not proof.
  • The brute-force cave oracle only reaches 5×5 caves and the exhaustive partition oracle only 7-card hands; larger inputs rely on the parity fixtures.
  • The LLM evaluation asks each hand once with one prompt, so it does not measure run-to-run variation or prompt sensitivity, and providers can change a model behind the same id.
  • No paid LLM run is stored in the repository. Every LLM result on the site comes from a run in the visitor's own browser.
  • The bootstrap intervals for mean gaps are nominal: on zero-heavy gaps they held the true mean 82% of the time at 10 hands and 91% at 30 in the coverage simulation (DR-005).

What I'd change

  • Use fast-check's shrinking to find the smallest caves where the 2019 and spec-correct modes disagree, and make those the presets.
  • Replace the notebook oracle with brute-force oracles of my own wherever a property can express the rule.
  • Run each LLM hand several times and with a second prompt, reported as paired comparisons on the same hands.
  • Test BCa and bootstrap-t intervals in the coverage simulation and switch if one holds its level better on this data.
  • Settle the history rewrite for the private notebooks before the repository is ever made public (DR-003).

6 · Transparency

AI use statement

Every AI call made from your browser is listed in the AI audit log.

This statement says where AI is used on the COMP10001 Playground, what it never does, what data leaves the visitor's browser, and where a person stays in charge. It is informed by the Australian Government's Policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act, and the NIST AI Risk Management Framework. This is a personal portfolio project: it is not assessed against any of them and makes no claim of compliance.

What AI does here

There is one AI feature, and it is optional: the evaluation "Can an LLM find the best partition?" on the Card Table (/card-table/llm-eval). When a visitor adds their own API key and starts a run, a language model is asked to group each seeded hand of cards, and its answers are scored against the exact solver.

What AI never does

  • It never runs without the visitor's own key and an explicit click on "Run".
  • It never changes the original 2019 code, the rebuilt functions or their results. The counting, the cave searches, the scoring and the exact solver are deterministic code.
  • It never scores itself or anyone else. Every answer is judged by the same rules engine as the rest of the site.
  • It never answers on behalf of another model. If the provider declines, the hand is recorded as refused; there is no silent fallback.
  • It never writes the text of this site at runtime.

Data sent to the provider

  • The scoring rules (a fixed system prompt), one hand of card codes such as 8S 8C 9D, and a JSON schema. No personal data is put in the request. Like any web request, the call reveals the visitor's IP address, browser user agent and this site's origin to the provider; cookies and the referring page are not sent.
  • The visitor's API key, in a request header, to the provider they chose: Anthropic (api.anthropic.com) or OpenAI (api.openai.com). Requests go directly from the browser. This site has no AI server code and never receives the key.
  • The provider handles requests under its own terms and retention policy.

Keys and storage

  • The key is kept in the browser's sessionStorage and is gone when the tab closes, unless the visitor turns on "remember on this device" (localStorage). "Forget keys" removes it.
  • Any script on the page could read a stored key, including browser extensions. The site loads no third-party scripts, and its Content-Security-Policy only lets the page connect to itself and the two provider APIs, so a script injected into the page could not send the key anywhere else with fetch (DR-006). Browser extensions are not bound by the policy, so visitors should still use a key with a spending limit.

Human in the loop

  • The visitor starts every run, can stop it at any time, and in-flight calls are cancelled if they leave the page.
  • Each answer is shown with an "AI-generated" label next to the exact answer and the greedy baseline's answer. The label marks content a model returned; failed calls and refusals are shown as such, without it. The visitor can accept or reject each answer, and add a note saying why; the decision and note are stored with the call.

Transparency and records

  • Every call is written to an audit log in the visitor's own browser (IndexedDB), with the prompt, the JSON schema and generation settings, the reply, the validation result, any refusal, the model, the latency, the token usage and the human decision. It is never uploaded. It can be viewed, filtered, exported as JSON or CSV, and cleared at /ai-log.
  • Scores are reported with intervals and the number of hands, and the evaluation design is documented on /methods, including how often the "95%" intervals actually held the truth in a simulation.
  • The recordings on the guided tour (/tour) and in the README show the evaluation running on mocked answers: the tour enters no key, and each provider call is answered inside the test browser by a fixed rule. Those frames are labelled "Mocked AI response for illustration", and their scores say nothing about any real model.

How this site was built

The 2019 Project 2 file was written by me, without AI tools. The 2026 revival, including this upgrade, was written with an AI coding assistant under my direction and review, and the commits record that.

7 · Decision records

Why it is built this way

Each record states the decision first, then the context, the options, why, what happened (weak numbers included) and what I would change. Past records are superseded, never edited.

  1. DR-001Accepted2026-10

    Rebuild the lost Project 1 and 3 submissions from the task, and label them as rebuilt

    Where my 2019 code is lost (Projects 1 and 3), the site runs a rebuild written from my own paraphrase of the task and its worked examples, every page says "Rebuilt from the task (original lost)", and code of unclear provenance is used only as a test oracle, never shipped.

    Read DR-001
  2. DR-002Accepted2026-10

    Keep my 2019 cave code exactly as submitted, with a separate spec-correct mode

    Falca's Cave runs a line-by-line port of my surviving 2019 file by default, bugs and crashes included, and a toggle switches to a "Spec-correct" mode that keeps the same algorithms with the bugs fixed, so nothing about the original is silently improved.

    Read DR-002
  3. DR-003Accepted2026-10

    Keep the 2022 reference notebooks private and publish only the answers their code gives

    The three 2022 notebooks stay on my machine, untracked and gitignored, because they quote the task text and may derive from staff sample solutions; the repository keeps only the inputs and answers their code produces, as test fixtures.

    Read DR-003
  4. DR-004Accepted2026-10· amended by DR-005, DR-006

    Evaluate an LLM on the Card Table with the visitor's own key, from the browser, with a local audit log

    The LLM evaluation is optional: it calls Anthropic or OpenAI directly from the visitor's browser with a key they paste in, scores every answer against the exact solver and two baselines with intervals, writes every call to an audit log on the visitor's device, and the site works fully without a key.

    Read DR-004
  5. DR-005Accepted2026-10

    Default the LLM evaluation to 30 hands and label its gap intervals as nominal

    The evaluation now deals 30 hands by default and offers 10 to 50, labels its bootstrap intervals for mean gaps "nominal 95%", says "no interval" when every hand has the same gap, and publishes a coverage simulation that shows how far short of 95% those intervals fall. DR-004's other decisions stand.

    Read DR-005
  6. DR-006Accepted2026-10

    Limit where the page can send a visitor's key with a Content-Security-Policy

    Every page is served with a Content-Security-Policy that lets it connect only to this site, api.anthropic.com and api.openai.com, load images only from this site, never be framed, and post forms only to itself; inline scripts stay allowed because the pages are prerendered without per-request nonces.

    Read DR-006