ACI-045: Measuring with a Stochastic Instrument
Thesis
A model playing your product while a second model grades the run is a measuring instrument with sampling error, and its output means only what the apparatus can prove. Put the sample on every report and commit an experiment's design before any output exists. Prove the arms were two, and give every instrument, an ablation included, a red path you have seen fire.
Story
In June 2026 Lamplight, an Elixir platform for interactive fiction whose characters are voiced by a language model, added a test its suite could not otherwise run: whether a story plays well. A written brief sends a model through the product's real text interface, as a black box, until it finishes, fails or declares the product incoherent (Lamplight@d43a2e776, 2026-06-27). The product refusing a move is part of play; only a genuine crash or a lost session counts as incoherence (Lamplight@7ce377d0e, 2026-06-27), and the player can end a run with a typed abort and a reason (Lamplight@2d84cfcfb, 2026-06-27). A second model grades the transcript against a rubric written with the brief, seeing only the rubric, the player's view and the transcript. Its report "states it is one sampled run, never a pass/fail gate", and the grading seam is injected, so the harness's own tests need no live model (Lamplight@b1089a947, 2026-06-27).