Long-horizon Agentic Tasks
An agent that nails a five-minute task in 2024 can still derail on a five-hour one. The reason is math, not magic: success decays like p^N as dependent steps stack up, while context rots and the goal drifts. This topic covers the four failure modes and the architecture — durable state, exit criteria, sub-agents, verifier gates, checkpoints — that extends the reliable horizon.
Extending the Reliable Horizon 🧭
Long tasks succeed when state lives outside the context window, each step has an exit criterion, and verifier gates plus checkpoints bound the blast radius of inevitable failures.
01.The Problem: Five Minutes Works, Five Hours Does Not
Ask a 2026 agent to rename a variable across a file. Perfect.
Ask it to migrate an entire billing service to a new rate-limit API over an afternoon. It starts strong — reads the docs, edits a few files, runs tests — and then, around step 60, it forgets what it is doing. It re-adds code it already deleted. It "fixes" a bug it created itself. It announces success while the test suite is red.
Why does the same model that handles five minutes beautifully derail at five hours?
The answer is not that the model got dumb. The answer has three parts:
- Small error rates compound. A 2% mistake rate per step is excellent. Over 100 dependent steps it is fatal.
- The model's memory is its transcript. The context window (the rolling text file of everything the agent has seen this session) fills with stale information, and attention to the original goal degrades.
- Nobody is checking. Errors stay invisible until they surface far downstream — by which point the agent has built on top of them.
A task is long-horizon when success requires many dependent decisions over minutes to hours, where an early mistake is not locally observable and the environment mutates underneath you. Examples: migrate a service across frameworks, reproduce then fix a subtle bug, run a research-and-draft workflow, coordinate a multi-vendor procurement.
Everything in this topic is about one goal:
Make every mistake visible while it is still cheap.
02.The Idea in Plain Words: Success Decays Like p^N
The central math, stated simply:
If each step succeeds with probability p, and mistakes are never caught, the whole task succeeds with probability pᴺ over N dependent steps.
Unpack:
p= per-step success probability (say 0.98 — 98% of steps go right).N= number of dependent steps (step 30's work stands on step 29).pᴺ= probability all of them go right. No second chances.
And the two escape words:
- "dependent" — independent steps do not multiply like this; only chained ones do.
- "never caught" — the decay assumes errors are invisible. If a check catches and repairs 80% of mistakes, your effective
pclimbs toward 1, and the curve flattens dramatically.
That is the whole topic in one sentence: long-horizon engineering is the art of lowering N (smaller, independent chunks) and raising effective p (catching errors early).
The field's favourite measurement (popularized by METR from 2024 onward) turns this into a number you can track: the 50%-time-horizon — the human task duration at which the agent completes half of a task distribution at 80% reliability. Reported history: near-zero on software tasks until ~2022, then roughly doubling every ~7 months through 2024, with frontier reasoning agents reaching multi-hour ranges in 2025 and labs reporting further gains into 2026.
Two lessons from the shape of that curve:
- Reliability, not capability, is the constraint. Agents can attempt long tasks; they fail on compounding error and loss of situational awareness.
- Horizon scales super-linearly with scaffolding (state, verification, decomposition), which is why vendor-to-vendor spreads on the same benchmark are enormous.
03.A Simple Worked Example: 2% Wrong, 100 Steps
Crunch the numbers. Three agents get the same 50-step task. Each step succeeds with probability 0.98 — a great model, on paper.
- 10 steps:
0.98¹⁰ ≈ 0.82→ works 8 times out of 10. Fine. - 50 steps:
0.98⁵⁰ ≈ 0.36→ works barely 1 time in 3. Ugh. - 200 steps:
0.98²⁰⁰ ≈ 0.018→ essentially never. Dead.
Same model. Same per-step quality. The only difference is N.
Now add scaffolding: a verifier gate every 5 steps that catches and repairs 4 out of 5 mistakes. Effective per-step failure drops from 0.02 to about 0.004:
- 50 steps:
0.996⁵⁰ ≈ 0.82→ back to working 8 times out of 10. - 200 steps:
0.996²⁰⁰ ≈ 0.45→ now a coin flip instead of zero.
One more lever — decomposition. If the 200-step task becomes four independent 50-step work packages that only get merged at the end (Topic detail: sub-agents), the task succeeds when each package succeeds, and each package's repair loop is bounded to its own 50 steps. You have not changed the model at all. You changed the shape of the problem:
codeno gates: p^200 = 0.98^200 ≈ 0.02 ← doomed gates: p^200 = 0.996^200 ≈ 0.45 ← usable chunked: 4 × (0.996^50) fixes ≈ 0.82 ← shippable each, bounded repair
That is why "scaffolding extends the horizon" is an arithmetic statement, not a metaphor.
04.Visual Intuition: The Cliff and the Guardrails
Picture reliability over time as a walker on a mountain ridge. Without checkpoints and gates, every stumble carries you further downhill, unnoticed:
codereliability ▲ 1 │╲ │ ╲ no gates, no external state │ ╲ ╲ │ ╲ ╲ p^N cliff — errors invisible, │ ╲ ╲ the agent builds on its own mistakes │ ╲ ╲___ │ ╲ ╲______ ← hour 5: confidently wrong 0 └────────┬─────┬─────┬─────┴────► task time 1h 2h 3h ═══ gate ═══ gate ═══ gate ← checks every milestone 1 ─╲___╱──╲___╱──╲___╱────────── ← flat horizon: each stumble 0h 1h 2h 3h hour 5 caught while still cheap
And here is the second picture — where the agent's memory actually lives. The dangerous default is all in the transcript (context window):
codecontext window (shrinks in usefulness as it grows) ┌──────────────────────────────────────────┐ │ old file reads · stale plans · dead ends │ ← rot lives here └──────────────────────────────────────────┘ durable state (re-injected every step) D ┌───────────────┐ │ │ spec · notes │ ◄── the task file, git │ │ exit criteria │ branch, test status │ └───────────────┘
The design rule falls out of the picture: copy what matters out of the transcript and into durable state, and treat the transcript as disposable.
05.The Analogy: The Spacewalk
Carry one analogy through the rest: an astronaut on a six-hour spacewalk to repair a station module.
A spacewalk is the original long-horizon task: many dependent steps, an unforgiving environment, and a suit computer (the context window) that can only hold so much.
Every safety practice of the spacewalk maps to an agent pattern:
- The printed checklist, velcroed to the wrist = durable external state. The astronaut never trusts memory for where they are in the procedure. The task file is the checklist.
- "Confirm B, then proceed to C" callouts with Ground = verifier gates with exit criteria. Each step ends with a stated, checkable condition before the next one starts. No vibes.
- The tether = checkpointing/resumability. If the suit glitches (the process crashes), the astronaut does not float away: they reel back to the airlock — a journaled, resumable state — not to the beginning of the mission.
- Two-person rule on destructive switches = irreversible-action approval. Cutting a cable (force-push, delete, spend money) is never done solo by the spacewalker; it requires Mission Control sign-off.
- The suit is not the mission = the model is a flaky dependency. NASA does not assume the suit works forever; it builds around failure — backups, limits, abort paths. So should your agent.
The failure story this analogy warns against: an astronaut who works by feel, skips the callouts, and untethers "just for this one step." That is every derailment you have seen in an agent demo at hour four.
Everything in the next sections — planners, sub-agents, memory, gates — is only about building the checklist, the tether, and Mission Control around the model.
06.The Four Failure Modes of Long Tasks
Before the fixes, the diagnoses. In one bold line each, then the detail:
- Compounding error: a wrong assumption at step 3 silently invalidates steps 4-90. Because the agent has no reason to revisit settled premises, the mistake is laundered into confidence — every later step "sees" the wrong premise as established fact and builds on it. (The astronaut who miscounted turns at minute 10 never finds the bolt at minute 300.)
- Context rot and window pressure: the transcript fills with stale file contents, repeated tool output, and obsolete plans. Two documented consequences: needle-in-haystack recall (finding one fact buried in long context) degrades as context grows, and models anchor on early context, reverting later edits.
- Specification drift: the goal is implicitly renegotiated as the conversation accumulates. The agent optimizes the current paraphrase of the task rather than the original acceptance criteria. Each summarization of the goal is a tiny photocopy of a photocopy.
- Environment and state divergence: files change, dependencies update, branches move, sessions die mid-step. Without durable state, the agent cannot resume and starts a fresh half-plan — two half-plans are worse than none.
Add two operational ones every team hits:
- Cost/latency blow-ups from re-reading the world each step (the transcript grows, every step pays for all of it).
- Unsafe irreversible actions (force-push, delete, spend money) taken at 0.9 confidence. The p^N math says that 0.9 confidence is already the wrong number for anything you cannot undo.
07.Architecture Patterns That Actually Extend the Horizon
Seven patterns. Each one answers a failure mode from the last section. Think of them as the spacewalk safety system for your agent:
- Durable state outside context. (→ spec drift, state divergence) Write the spec, decisions, and current status to an artifact (task file, structured state note, DB row) that is re-injected each step — the wrist checklist. Git branches and worktrees (separate working copies of a repo sharing one history) are the accidental standard for code agents because they give cheap diff, rollback, and parallelism.
- Explicit plan with exit criteria. (→ compounding error) Decompose into milestones, each with a machine-checkable done condition ("test X passes", "endpoint returns 200 with schema Y"). Without exit criteria, "done" becomes a vibe.
- Sub-agents / work packages with isolated context. (→ context rot) Spawn a scoped agent per package with its own tool budget; the parent keeps only summaries. This is the primary defence against context rot and the main reason 2025-2026 harnesses (Claude Code sub-agents, Cursor background agents, orchestration frameworks) fan out.
- Verifier gates before merging state. (→ compounding error, the positive version) Deterministic checks first (tests, typecheck, schema validation, linters, dry-run APIs, policy checks), learned judges second (rubric LLM-as-judge with swapped positions — ask the judge twice with the two candidates reversed, to catch order bias). Gates are where reliability compounds positively instead of negatively.
- Checkpointing and resumability. (→ state divergence, crashes) Idempotent steps (re-running one is safe, not duplicative) journaled by step ID, artifacts snapshotted, and a resume path that reconstructs context from state rather than transcript. The tether.
- Stall detection and budget control. (→ cost blow-ups, loops) Detect loops (repeated identical tool calls, oscillating edits that undo each other), cap steps/tokens/cost per milestone, and escalate with a written summary of hypotheses tried. Mission Control abort authority.
- Reflection and consolidation passes. (→ rot, drift) Periodically summarize what is known, what changed, and what remains — a cheap "dreaming" step that fights drift and shrinks context.
A useful mental model: the agent is a state machine whose memory must survive its own context window.
Here is what a machine-checkable milestone spec looks like — the artifact the orchestrator reads, verifies, resumes, and escalates on:
task: migrate billing-service to the new rate-limit API
acceptance:
- id: rl-1
claim: "client sends X-RateLimit-Key and honors 429 Retry-After"
verify: "pytest tests/rate_limit.py -q" # deterministic gate
budget: { steps: 25, tokens: 400000, minutes: 20 }
blast_radius: reversible # branch + draft PR only
- id: rl-2
claim: "production config updated in us-east and eu-west"
verify: "terraform plan shows only 2 changed resources"
blast_radius: irreversible # -> requires human approval
stall_detection:
repeat_tool_call_threshold: 3
oscillating_edit_window: 5
on_exhaust: escalate_with(diff, hypotheses_tried, remaining_acceptance)08.Training and Evaluating for Long Horizons
Training. Post-training moved from single-turn instruction following to trajectory-level supervision: multi-turn tool-use SFT with synthetic trajectories (including error-recovery branches — deliberately broken tool results, bad assumptions revisited), and RL on verifiable environments where reward arrives only at the end (or at milestones). DeepSeek-R1-style RLVR (reinforcement learning with verifiable rewards) plus long-CoT training improved multi-step planning transferably, and 2025-2026 frontier models are explicitly tuned for 20-200 step agentic runs, tool-call format robustness, and "when to stop and ask".
Evaluation. Single-number task success is not enough for long horizons — a model that passes 40% of 3-hour tasks at 10× the cost of one that passes 35% may be the worse product. Report:
- Horizon curve: success versus step count / duration, and the 50%-time-horizon. (Not a leaderboard number — a reliability-versus-duration shape.)
- Process metrics: tool-call error rate, redundant-read ratio (how often the agent re-reads the same file pointlessly), edit oscillation, escalations per task.
- Blast-radius metrics: irreversible actions taken, wrong-file edits, tests weakened. These are the "near-misses" scoreboard.
- Cost per accepted outcome: dollars and minutes per merged PR or completed research artifact.
- Benchmarks to cite: SWE-bench Verified and its multimodal/live variants, WebArena/OSWorld for interfaces, Terminal-Bench-style environment tasks, GAIA for mixed-tool research, and long-horizon suites that reward sustained multi-hour work.
09.Design Checklist Before You Ship an Hour-Scale Agent
Before a long-horizon agent touches a real repo, real inbox, or real budget, run this list — it is the pre-launch go/no-go in one place:
- State is externalized and typed (spec, decisions, artifacts), not implicit in a transcript.
- Every milestone has a deterministic or near-deterministic verifier; the agent cannot edit the verifier. (No sawing your own tether.)
- Irreversible actions are classified and gated by policy with human approval; everything else defaults to sandbox and draft.
- Resumability is tested by killing the process mid-task, at scale, in production-like environments. (You do not trust the tether until someone has actually cut it in a drill.)
- Budgets exist per milestone and per task, with escalation that produces a useful handoff.
- Observability: full step journal, tool I/O sizes, cache hit rate, and cost per accepted outcome on a dashboard.
- A rollback story for damage already done (feature flags, git revert, API compensations).
- Concurrency controls when several agents write to the same repo, inbox, or account.
If any line makes you say "we would have to check", that is the real horizon limit of your system — not the model.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Verifier gates plus externalized state raise task success disproportionately to model capability gains.
- Sub-agent fan-out parallelizes work packages and keeps each context window clean.
- Checkpointing converts long runs from fragile gambles into resumable jobs with predictable cost.
- Explicit acceptance criteria make escalation and human takeover cheap and auditable.
Trade-offs & Constraints
- Scaffolding adds latency, tokens, and infrastructure: state store, sandboxes, gates, orchestration, dashboards.
- Naive scaffolding can lengthen attempts while reducing success, hiding regressions behind a single accuracy number.
- Milestone gates are brittle when the real objective is fuzzy; over-specified gates optimize the metric, not the intent.
- Parallel sub-agents create merge conflicts, duplicated effort, and coordination bugs that are hard to reproduce.
- Long runs amplify safety exposure: more tool calls means more chances for an irreversible action.
Teams run a planner that splits a migration into package specs with test-based exit criteria, spawns sub-agents per package inside isolated git worktrees, and gates each merge on build plus flake-free test runs. State lives in a task file plus branch metadata; a reconciler detects spec drift by diffing current behaviour against acceptance criteria; step journals make runs resumable after CI preemption. Humans review PRs at milestone boundaries instead of watching the loop.
Staff+ Engineering Takeaways
- Long-horizon difficulty is compounding error plus invisible state, not raw model intelligence: success decays roughly as p^N until errors become observable.
- The four failure modes are compounding error, context rot, specification drift, and environment/state divergence.
- Horizon-extending architecture: durable external state, milestone exit criteria, sub-agent context isolation, verifier gates, checkpoints, stall detection, escalation.
- The measured progress metric is a horizon curve (50%-time-horizon at a reliability threshold), reported as roughly doubling every ~7 months through 2024-2025.
- Scaffolding is not automatically good: it can lengthen attempts while lowering success, so evaluate process metrics and cost per accepted outcome.
- Classify irreversible actions and gate them; hour-scale agents are fault-tolerant batch jobs with a flaky dependency.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why does naive agent scaffolding sometimes reduce overall task success on long horizons?
How clear and actionable was this distributed systems breakdown?