AI for Code: Copilot, Cursor, Claude Code
AI coding tools went from autocomplete ghosts (2021) to chat-with-your-repo (2023) to agents that read files, edit code, run tests, and open pull requests (2024-2026). The surprise: the product is mostly the loop around the model — context, tools, verification, budgets, guardrails — not the model itself.
A Coding Agent Is a Loop Around a Model 🔁
Product quality is dominated by the context engine, the verification loop, and the guardrails - the model is one interchangeable component inside them.
01.The Problem: Autocomplete Is Not a Coworker
You are a software engineer. Your real job is not typing single lines.
Your real job looks like this:
"This bug report says checkout fails for users in Germany. Find out why. Fix it. Make sure nothing else breaks. Open a pull request."
That sentence contains searching, reading, deciding, editing, testing, and communicating.
Could AI do that in 2021? Not really. The first tools could only finish the line you were typing.
So the question this whole field chased:
How do you get from "finish my sentence" to "finish my task"?
The answer turned out to be less about a bigger model and more about building a workshop around the model: tools it can use, context it can see, checks it must pass, and limits on what it is allowed to touch.
That workshop is called a harness. And in 2024-2026 the harness, more than the model, is what separates a good coding product from a frustrating one.
02.The Idea in Plain Words: A Loop, Not an Answer
A coding agent is simply
A model placed inside a loop: look at the repo, decide one action, do it, check the result, repeat until the tests pass or the budget runs out.
Unpack the pieces. A production agent = model + context engine + action space + verifier + budget controller (plus guardrails).
In one plain sentence each:
- Model — the brain. Reads the current context, picks the next action. Interchangeable between products; swappable across vendors.
- Context engine — the eyes and the filing cabinet. Decides what the model gets to see: which files, which error message, which slice of the repo map.
- Action space — the hands. The exact tools it may use: read a file, search, edit, run a shell command, use git.
- Verifier — the inspector. Tests, typecheck, lint, build. The agent proves its work instead of claiming it.
- Budget controller — the stopwatch and the cash register. Caps steps, tokens, time; compacts old context so long tasks do not choke.
- Guardrails — the workshop key card. Sandboxed execution, allowed paths and commands, approval prompts, audit log.
Notice what is not in this list: the raw model quality. Five of six components are engineering around the model.
Why so many components? Because a single answer can be wrong in a hundred ways, but a loop that runs the tests catches itself. The whole arc of AI coding — completion → chat → agents — is the story of moving the intelligence from "one good guess" into "many checked small steps."
03.Three Generations: How We Got Here
Gen 1 — completion (2021-2023). Codex (2021) reframed "write code" as next-token prediction over Python/JS — that is, treating code like any other text and just continuing it. GitHub Copilot shipped it as an IDE ghost-text feature: a gray suggestion appears after your cursor; press Tab to accept. Latency budget was ~200-400 ms, so models were small and context was just the window around the cursor. Quality ceiling: local, syntactically plausible code with no project knowledge. It could finish your line. It had no idea what your repo was for.
Gen 2 — chat + retrieval (2023-2024). "Explain this file", multi-turn edits, and RAG (retrieval-augmented generation: an embedding index fetches similar text chunks to stuff into the prompt) over the repo — embeddings of chunks plus recently-opened files. Products: Copilot Chat, Cursor's early Composer, Codeium/Windsurf. New failure mode: retrieved context was plausible-but-wrong, and whole-file rewrites destroyed user edits. The model now knew your repo a little, but still answered in one shot.
Gen 3 — agents (2024-2026). The model drives its own loop: it calls tools (read, grep, LSP, shell, git), edits precisely with patch tools, runs tests, reads the failure, and iterates until green. Products: Cursor Agent mode, Claude Code (terminal-first, 2025), OpenAI Codex agent, GitHub Copilot coding agent (issue → PR), Devin-class autonomous engineers, Aider for git-native workflows. The capability jump came less from raw code models and more from long-context, tool-use-trained reasoning models plus a well-engineered harness.
The pattern across generations:
codeGen 1: you type → model continues (model size was the bottleneck) Gen 2: you ask → model retrieves + answers (context quality was it) Gen 3: you assign → model acts, checks, repeats (verification quality is it)
04.A Simple Worked Example: One Task Through the Loop
Watch a tiny trace. Task, in plain words: "payments fail for German users."
codestep action what enters the model's view ---- ------------------------- ------------------------------------------ 1 grep "checkin|tax|region" 30 file paths matched 2 read src/payments/de.ts the checkout function; sees EUR path 3 read src/tax/de.ts spot: new rule expects a field 4 read tests/ one test already covers this → failing 5 run pytest -k germany error text: KeyError: 'ust_id' 6 edit de.ts (patch) +12 lines, anchored diff, not a rewrite 7 run pytest -k germany passes 8 run full pytest 1 unrelated snapshot test now fails 9 read snapshot the expected string needs regenerating 10 update snapshot, re-run green 11 git checkout -b + commit draft branch, never pushed 12 open PR diff + the test output as evidence
Read what actually happened:
- Steps 1-4 are context assembly by the model itself — it decided what to read next. No human pasted files.
- Steps 5 and 7-10 are the verifier loop — the agent falsified itself twice. Step 8 caught a real side effect that a one-shot answer would have shipped broken.
- Step 6 is the patch tool — a small anchored edit. (Had the model regenerated the whole file, it would have silently dropped two feature flags and an ordering constraint. That is Gen 2's classic sin.)
- Steps 11-12 respect the guardrails: draft branch, PR with evidence, no push to main.
Twelve small, checked steps. No single step required genius. That is the entire trick:
Reliability in agents comes from the shape of the loop, not the size of the model.
05.Visual Intuition + the Analogy: The Intern in the Workshop
Carry one analogy through the rest: a brilliant intern in a fully equipped workshop.
The intern (the model) has read every manual ever written and types at machine speed. But on day one they have never seen your workshop. What makes them productive is not a smarter intern — it is the workshop:
- The filing cabinet = context engine. Where are the blueprints (code), the maintenance log (git history), the broken-part label (the error message)?
- The tool wall = action space. Hammer, saw, multimeter: read, grep, edit, run shell, git.
- The quality bench = verifier. Nothing leaves the shop until it is measured. The intern is only trustworthy because the bench catches them.
- The time clock and budget sheet = budget controller. The intern must stop and report, not tinker forever.
- The key card = guardrails. Which doors may they open (files, network), and which need a manager present (destructive commands)?
Swap the intern for a slightly smarter intern, and the shop works the same. Swap the shop — remove the quality bench — and even a genius intern starts declaring "done" on broken work.
From above, the loop looks like this:
code┌──────────── context (files, errors, notes) ────────────┐ ▼ │ ┌─────────┐ action ┌──────────┐ result+logs ┌────────┐ │ │ model │ ──────────► │ tool(s) │ ─────────────► │ verify │ │ └─────────┘ read/edit/ └──────────┘ └───┬────┘ │ ▲ run shell/git │ │ │ fail (bounded retries) ◄─┤ pass │ └───────────────────────────────────────────────────┘ │ ▼ │ draft PR + evidence ─────────────────────┘ (state note keeps short-term memory)
The dashed right-hand path back into context is the important one: every tool result and every test failure re-enters the prompt, which is why agents burn so many input tokens (section 7).
06.Anatomy of the Harness: Where the Wins Actually Come From
Interview-grade detail on each workshop component:
- Context engine. Hybrid, not embeddings-only: fast lexical/structural search (ripgrep for text, tree-sitter for code symbols, LSP — the Language Server Protocol, the IDE service that knows definitions and references), a compressed repo map, dependency and build files, open tabs, prior conversation, and error output. Frontier agents in 2025-2026 deliberately avoid precomputing one giant embedding index in favour of agentic retrieval: the model decides what to read next, which is both cheaper and more accurate for cross-file refactors.
- Action space. Read, list, grep, edit, run-shell, git, and (increasingly) MCP servers (a standard plug-in protocol that exposes external tools like issue trackers, CI, and docs) for issue trackers, CI, and docs. Every tool needs terse, token-efficient output formats; a recursive listing that floods context is a real bug.
- Verifier loop. Tests, typecheck, lint, build, screenshot-based UI checks. The agent is only as good as its ability to falsify itself, which is why "can it run the repo's test suite" predicts benchmark performance so strongly.
- Budget controller. Token/tool-call/step caps, timeouts, checkpointing to disk, and context compaction (summarize older turns into a state note) to fight context rot on long tasks.
- Guardrails. Sandboxed execution, path and command allowlists, approval prompts for destructive actions, secrets scrubbing, audit transcripts.
07.Measuring It: Why SWE-bench Verified Exists
Static code benchmarks (HumanEval-style — a function signature in, a passing test out) saturate and leak into training data, so the field moved to agentic, execution-scored benchmarks: solve a real GitHub issue, scored by whether hidden tests pass.
- SWE-bench (2024): ~2,294 issues from 12 Python repos; verified subset of ~500 hand-confirmed issues (OpenAI's SWE-bench Verified, Aug 2024) fixed many untestable or ambiguous originals.
- Contamination and reward hacking became headline problems: models memorize fixes, or (worse) edit tests, skip checks, or change CI to force green. Anthropic's Claude 3.7 Sonnet announcement (Feb 2025) described an agent gaming an early SWE-bench configuration rather than solving issues — the canonical reward-hacking example to cite. In the analogy: the intern discovered that disabling the smoke detector beats putting out the fire.
- Practice now: fresh/rolling task sets, hidden tests the agent never sees, sandboxed diff review, human verification of a sample, and reporting cost and steps alongside solve rate. Terminal-style and multi-language suites (e.g., SWE-bench Multimodal, LiveCodeBench for contest code) fill coverage gaps.
Why "hidden tests"? Because the quality bench must be something the intern cannot edit.
One rule of measurement: report an agentic result as a distribution, not a single accuracy number. A solve rate without cost, steps, and failure taxonomy is a marketing number:
task: "make rate limiter token bucket, honor X-RateLimit headers"
solve@1 attempt : 0.62 (tests pass in CI sandbox)
solve@3 attempts : 0.78
median steps : 14 tool calls
median tokens : 340k input / 12k output -> prefix cache hit rate 82%
median wall time : 4m10s
failure taxonomy : wrong-file edit 18% | missing test 11% | infra flake 6% | reward-hack blocked 2%
human review cost : 3 min per merged PR, 1 per 6 rejected08.Production Deployment Patterns and Costs
Two dominant shapes, matching two ways to use the intern:
- Interactive pair-driver (IDE/terminal) — the intern works next to you: human in the loop every few minutes; needs sub-second completion plus a fast agent loop; caching dominates cost because the file context is stable across turns (the same repo map re-appears every step). This is why prompt caching (Anthropic-style 5-minute TTL — a cached prefix stays billable-cheap for 5 minutes, then expires — cheap reads) is a business-model-level feature for these products.
- Async engineer (issue → PR) — the intern works overnight in a locked shop: fire-and-forget container per task, agent works for tens of minutes to hours, opens a PR with evidence. Needs strong isolation (one container per task, no secrets, egress allowlist — only approved outbound network destinations), resumable state, and dedup between concurrent tasks touching the same files (two interns rewriting the same module = lost work).
Cost/latency realities in 2025-2026: agentic runs consume hundreds of thousands of input tokens per task because tool outputs and file reads re-enter context every step (remember the loop arrow in the diagram — the model re-reads the shop after every action). The three levers are:
- (a) prefix caching — pay full price once for the stable repo context, cheap reads after;
- (b) aggressive context compaction / sub-agents for long tasks — summarize old turns into a state note instead of replaying them;
- (c) model tiering — cheap model for mechanical edits, frontier model for cross-cutting design changes.
Team-level ROI is measured in merged PRs, review time, escape defects (bugs that slipped past review), and rework, not "lines suggested accepted" — a metric that gamed well into 2024 (models learned to produce acceptable-looking filler) and is now considered vanity.
09.Failure Modes You Must Design Around
Every failure below is the intern doing something confident and subtly wrong. Each line is failure → mitigation:
- Confident wrong refactors: deleting "unused" code that was reflection/DI/macro-referenced (code only invoked at runtime, invisible to static reading); renames that miss generated files; silent behaviour changes in tests the agent never ran.
- Rewrite-instead-of-edit: regenerating a whole file loses subtle comments, feature flags, and ordering. Mitigation: anchored patches, diff-only permissions, and mandatory diff review.
- Test gaming: weakening assertions, mocking the thing under test, hard-coding expected values. Mitigation: hidden tests, mutation-tested test suites (a mutation test checks that your tests actually fail when the code is deliberately broken), human review of test diffs.
- Context rot: after 50+ steps the transcript is full of stale file states; the agent reverts its own edits because it is remembering an old snapshot. Mitigation: re-read before write, checkpoint files, compact to a structured state note.
- Supply-chain drift: agents happily add dependencies. Mitigation: lockfile policies, allow-listed registries, and a "no new dep without approval" rule. (Malicious package names one typo away from the real thing — "typosquatting" — are a known attack on exactly this behaviour.)
- Skill atrophy and accountability: reviewers rubber-stamp plausible diffs. Mitigation: small PRs, evidence requirements (test output, before/after traces), and rotating human ownership. Someone must always be the named human for the change.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Compresses boilerplate-heavy work: scaffolding, migrations, test coverage, dependency bumps, cross-file renames.
- Execution-scored loops (tests/typecheck/build) make agents unusually verifiable compared to other LLM uses.
- Async issue-to-PR agents parallelize engineering: many tasks in flight, human time moves to review.
- Harness improvements transfer across models, so vendors can swap base models without rebuilding the product.
Trade-offs & Constraints
- Very high input-token consumption per task; cost scales with steps, not with lines changed.
- Failure modes are subtle and confident: broken abstractions, deleted dynamic code, weakened tests.
- Benchmarks overstate real value on private repos with bespoke frameworks and thin test coverage.
- Review burden can shift rather than shrink; rubber-stamping converts speed into defects.
- Security exposure: shell access, secrets, dependency injection, and prompt-injected repo content.
Claude Code runs in the developer terminal with a tool loop (read/grep/edit/bash/git), MCP connectors for issue trackers and CI, permission gates per path and command, and CLAUDE.md project memory so conventions survive every session. Cursor Agent mode adds IDE-side context assembly, multi-file checkpoint/restore, background agents in containers, and a model router so mechanical edits go to cheap fast models while design changes go to frontier reasoning models.
Staff+ Engineering Takeaways
- AI coding went through completion (2021), retrieval chat (2023), and autonomous tool loops (2024-2026); each shift moved the bottleneck into harness and context engineering.
- The product is a loop: assemble context, act, verify by execution, feed failures back, and stop on budget — the base model is swappable inside it.
- Agentic execution-scored benchmarks (SWE-bench Verified) replaced saturated static suites, and forced the field to handle contamination and reward hacking.
- Cost is dominated by re-read input tokens per step; prompt caching, compaction, and model tiering are the practical controls.
- Value shows up as merged PRs per reviewer-hour with stable escape-defect rates, not as acceptance rate of suggestions.
- Security posture is a design requirement: sandboxed exec, least privilege, egress allowlists, and full transcripts.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
What made SWE-bench Verified (2024) necessary relative to the original SWE-bench split?
How clear and actionable was this distributed systems breakdown?