Skip to content
COMP10001Playground

Showcase

A guided tour

Three short walkthroughs of the playground, then a screenshot of each feature. Each video carries a numbered banner whose steps are written out beside it.

The videos are recorded by an automated browser tour that doubles as an end-to-end test of these journeys. Its inputs are fixed (an example electorate, an example cave, card deal seed 19 and the evaluation's default seed 2019), so every run records the same thing. No API key is used anywhere in it: the AI answers in workflow 3 are mocked and labelled as such.

Workflow 1

Count the votes

One electorate, three counting rules, three different winners, with every elimination round shown.

The numbered banner in the video matches the steps. Open the video file (MP4, 0.6 MB) or its captions (WebVTT).

  1. Step 1: Load the preset electorate: 13 ballots, 4 candidates

    Pick the “Same ballots, three winners” example: thirteen ranked ballots for alan, guido, ada and grace.

  2. Step 2: First past the post reads first choices only: alan wins with 4 of 13

    First past the post counts first choices only. alan has the most (4 of 13) and wins without a majority.

  3. Step 3: Instant runoff: no one has a majority, so the last-placed goes

    Switch to instant runoff and press Play. Nobody holds 7 of 13 first choices, so the lowest candidate is eliminated (ties go to the alphabetically first name).

  4. Step 4: Each round, the eliminated candidate's ballots move to their next choice

    ada goes in round 1 and alan in round 2. Striped bar segments are transferred ballots, coloured by their first choice.

  5. Step 5: guido reaches 8 of 13: same ballots, three different winners

    guido wins round 3 with an absolute majority. First past the post, a second-preference runoff and instant runoff crown alan, grace and guido.

Workflow 2

Falca's Cave

Edit a cave, watch breadth-first and uniform-cost search, then compare my 2019 code with the spec-correct version.

The numbered banner in the video matches the steps. Open the video file (MP4, 1.2 MB) or its captions (WebVTT).

  1. Step 1: Load “The long way round”: a 7×7 cave with three treasures

    Start from the 7×7 example cave: three treasures, a sword in the far corner and a dragon in the middle.

  2. Step 2: Edit the cave: paint a new wall along the top

    With the Wall tool, paint a new wall square along the top row.

  3. Step 3: Place the dragon right beside the top-left treasure

    With the Dragon tool, move the dragon next to the top-left treasure. The cave is still valid under both versions of the code.

  4. Step 4: shortest_path: breadth-first search floods out ring by ring, 12 moves

    Run shortest_path from the entrance to the exit. Each ring is one more move away; the exit is 12 moves out.

  5. Step 5: optimal_path as submitted in 2019: every treasure, then out, in 28 moves

    Run optimal_path in “As submitted (2019)” mode: my original code walks straight up to the treasure beside the dragon, 28 moves in all.

  6. Step 6: Spec-correct: the dragon guards 8 squares, so fetch the sword first: 40

    Flip to “Spec-correct”: the dragon guards its square and the eight around it until Falca has the sword, so the route detours to the sword and takes 40 moves.

Workflow 3

Card Table and LLM evaluation

Group a hand by hand, let the exact solver find the best grouping, then put a language model on the same task with your own key.

The numbered banner in the video matches the steps. Open the video file (MP4, 3.0 MB) or its captions (WebVTT).

Mocked AI response for illustration. The AI answers in steps 7 and 8 are mocked for illustration: no API key is entered, and the tour intercepts the provider call in the browser and answers it with a simple fixed rule, so no real model is called. The scores they show say nothing about any real model.

  1. Step 1: Deal a hand: seed 19, ten cards

    Press “Deal again”. The tour pins the seed to 19 so every run deals the same ten cards.

  2. Step 2: Group the three 3s: three of a kind scores +18

    Tap 3♠, 3♦ and 3♣, then “Make a group”: three of a kind scores 3 × 3! = 18. The score updates live.

  3. Step 3: Add the pair of 10s (+20): every loose card still counts against you

    Group 10♦ and 10♣ for +20. The five loose cards cost their face value (an Ace costs 20), so the table stands at −21.

  4. Step 4: The exact solver checks all 115,975 groupings: the best is 59

    “Find the best grouping” checks every partition of the ten cards in a Web Worker. The best uses the Ace as a wild card in the run J♠ A♥ K♠ and scores 59.

  5. Step 5: LLM evaluation: 30 seeded hands, baselines with 95% intervals

    Open the evaluation: the exact solver, a greedy baseline and all-singletons on 30 seeded hands, with Wilson intervals for rates and bootstrap intervals for gaps.

  6. Step 6: Bring your own key: AI settings. No key is entered in this tour

    AI settings take your own Anthropic or OpenAI key, kept in this browser only. The tour enters no key: it sets a placeholder model id and leaves the key field empty.

  7. Step 7: Mocked AI response for illustration: every answer is labelled and judged

    The run sends each hand to the provider one at a time. Here the call is intercepted and answered by a fixed rule; each answer is labelled AI-generated and judged by the same rules engine.

  8. Step 8: Accept or reject each answer; every call is in the AI audit log

    Review answers one by one. Every call, with the prompt, reply, latency, tokens and your decision, is kept in the browser's audit log and can be exported as JSON or CSV.

Every feature

Screenshots

Desktop at 1440 × 900 and phone at 390 × 844. Select one to enlarge it; the arrow keys step through them.

  • Landing pageThree 2019 projects, rebuilt as a playground that runs in the browser.
  • Landing page, dark themeThe same page in the dark theme.
  • Ballot BoxInstant runoff, round 2: transferred ballots are striped by first choice.
  • Falca's CaveSpec-correct mode: the dragon's reach is hatched and the route fetches the sword.
  • Breadth-first searchshortest_path floods out ring by ring; each number is a distance.
  • Card TableLive scoring as you group, and the exact solver's best grouping.
  • LLM evaluationExact solver, greedy and singletons on 30 seeded hands, with 95% intervals.
  • Bring your own keyAI settings: your own key, kept in this browser and sent only to the provider.
  • LLM run (mocked)Mocked AI response for illustration: answers labelled AI-generated, judged, and open to review.
  • AI audit logEvery AI call (here, the mocked ones) with prompt, reply and your decision.
  • MethodsProvenance, verification, evaluation design and decision records.
  • Mobile: landingThe landing page on a phone.
  • Mobile: Ballot BoxThe count on a phone.
  • Mobile: Falca's CaveThe cave painter on a phone.