Agentic Coding

ACI-002: The Methodology Desert

Thesis

The agentic coding ecosystem has abundant tools and almost no reusable methodology frameworks: every major AI company has shipped a coding agent, and almost no one has shipped a framework for working with one over months of real development. The gap between "how to use the tool" and "how to build a practice" is largely unaddressed, so judge a tool by the practice around it as well as by what it can do.

Story

Search for "agentic coding" and you find tools: Claude Code, Cursor, Aider, Windsurf, Copilot, Devin, Cline, Continue, Codex CLI. Each has documentation explaining features, configuration and basic usage patterns. Anthropic tells you to use CLAUDE.md for project rules. Cursor tells you to use .cursorrules. Aider tells you to load a conventions file.

When this page was first written, in April 2026, a search for "agentic coding methodology" found almost nothing. No framework told you how to evolve your CLAUDE.md over time (ACI-010: The Guidance File). No guide explained what happens when your agent cross-contaminates between projects (ACI-012: Cross-Project Contamination). No standard existed for managing a session's lifecycle: how to start a session, orient the agent, verify its work, finish cleanly and hand context to the next session.

The closest things to methodology in the ecosystem were SWE-bench (a benchmark, not a practice), ChatDev and MetaGPT (multi-agent research frameworks, not practitioner methodology), and individual blog posts describing personal workflows that are useful but not reusable.

The multi-agent case is worth separating out, because it is easy to read "ChatDev and MetaGPT exist" as "multi-agent is covered". Those are autonomous agent swarms: agents coordinating with each other, with the human outside the loop. A human directing a fleet of agents is a different problem, and at the survey it had no practitioner framework either. The third year of this course takes it up, starting at 301: The Step Up.

Intent, the framework behind the projects this course draws its evidence from, is one attempt to fill the gap, and what it ships can be counted. As of 2026-09-14, at Intent@4c172d260, it ships a library of 68 rules, each addressed by an id; 23 skills, which are packaged instructions the agent loads for a kind of task; and 9 subagents, separate agents a session can hand work to, 6 of them critics that check code or prose against the rule library. Around those it organises work into steel threads and work packages, and every session into a start that loads the project's skills and a finish that records where the work stands. Whether that makes Intent more developed than other frameworks would take a survey this course has not made, so the course states what Intent ships and leaves the comparison to you.

The ecosystem optimises tool capabilities and assumes practitioners will work out the practice on their own. In the survey's reading, most developers therefore stayed at the autocomplete level and never progressed to agentic workflows, while enterprise teams deployed tools without a methodology and measured the wrong things.

Evidence

What Exists

Category Examples Methodology content
Coding agents Claude Code, Cursor, Aider, Windsurf, Copilot Tool documentation only
Benchmarks SWE-bench, Aider leaderboard Measures capability, not practice
Multi-agent research ChatDev, MetaGPT, OpenHands Agent-agent frameworks, not human-agent methodology
Blog posts Thorsten Ball, Simon Willison, swyx Individual workflows, not reusable frameworks
Enterprise guidance McKinsey, ThoughtWorks Radar Adoption patterns, not development methodology

What Did Not Exist Elsewhere

  • A standard for evolving CLAUDE.md or other project rules over time
  • A session lifecycle framework (orient, plan, execute, verify, finish, hand off)
  • Correction mining, or learning from an agent's mistakes
  • Verification methodology beyond "run the tests"
  • Debugging methodology specific to agent-generated code
  • A framework for multi-repo, long-duration agentic development
  • A standard for when to delegate and when to write the code yourself

Both lists are the April 2026 survey's and have not been re-checked.

What Intent Ships

As of 2026-09-14, at Intent@4c172d260:

Gap What Intent ships
Session lifecycle /in-session at every start and after every compaction, /in-finish to close, and hook templates for session start, each prompt and stop
Rules that evolve A library of 68 rules in nine packs, each rule addressed by an id such as IN-AG-HIGHLANDER-001
Enforcement 6 critic subagents that read the library and apply each rule's detection to code or prose
Correction learning The in-autopsy skill, which analyses Claude Code sessions against the project's memory rules
Code audit Five in-tca-* skills that run a Total Codebase Audit from set-up to finish
In all 23 skills and 9 subagents

The nine packs are one language-agnostic pack, five for programming languages and three for prose; five critics cover the programming languages and one covers prose. The lifecycle is taught in ACI-013: Session Lifecycle as Missing Primitive, and the library and its critics in ACI-027: Rules as a Library, Enforced by Critics.

Each figure is counted in Intent's tree at Intent@4c172d260 (as of 2026-09-14): rules as RULE.md files in the rule library, skills as SKILL.md files, subagents as agent definitions, and critics as the subagents named critic-*. These are counts of what ships, not evidence that any of it works.

Git References

  • Intent@4c172d260 (2026-09-14): the commit at which every count in the inventory is taken.

Application

  • When evaluating AI coding tools, evaluate the practice around the tool, not only the tool itself. A tool used inside a working practice (sessions that start and finish the same way, output that is verified, rules that evolve) can deliver more than a more capable tool used without one.
  • When a framework is described as the most developed of its kind, Intent included, ask what it ships at a stated version and date, and count it yourself. A framework's description of itself is a claim; its inventory at a commit can be checked.
  • When building an AI-assisted development practice, budget as much time for methodology as for tool selection. The tool choice matters less than you think, and the practice matters more.
  • Look for practitioners who share reusable frameworks, not just anecdotes. A blog post saying "I used AI to build X" is useful but not transferable. A framework saying "here is how to structure sessions, verify output and learn from mistakes" is transferable.
  • If you have no framework yet, build the smallest one a single developer can keep: a rules file you revise whenever the agent repeats a mistake, the same start and finish for every session, and a check on each report the agent gives you.

Anti-Pattern

"Tool Worship": An organisation spends six months evaluating whether to use Copilot, Cursor or Claude Code. They run benchmarks, compare completion quality, measure token costs, and read every blog post comparing features. They finally choose a tool and deploy it across the engineering team. Six months later most of those developers use it as autocomplete, and some have stopped using it altogether. The tool selection was meticulous. The methodology was non-existent. No one taught developers how to move from autocomplete to agentic workflows, how to verify agent output, or how to build durable project rules. The tool is capable; the practice is absent.