205: Testing with Agents
Overview
104 taught one human to see a test go red before believing its green. This chapter tests at project scale, where agents write and run much of the suite and a language model may sit inside the product under test. In each failure here the suite reported what it was built to report, and what mattered happened where no assertion looked. A suite made a live, billed call for every memory or training record its tests created. A test ran the project's own health check and migrated the live store. Two arms of an experiment collapsed into one job and agreed perfectly. A fence against model version strings stayed green while it scanned nothing.