Code Review Is Dead: Agents Write Faster Than Humans Read

Why line-by-line AI code review fails to scale, and how risk-based checks and context engineering unblock development.

  • Why line-by-line review stopped working
  • What companies do instead of reading line by line
  • Context determines whether there is anything worth checking
  • Simple looks easy, isn't

Why line-by-line review stopped working

  1. The autumn articles from The Pragmatic Engineer describe the same bottleneck at different companies: agents in Cursor, Claude Code, and Copilot produce code faster than teams can read it line by line. Asana completed a test framework migration in two weeks using AI—a task that had remained in the backlog for years as “too expensive.”

  2. The speed of writing code has increased, but the speed of reviewing it has not: the bottleneck has shifted, and review designed for human pace does not work at agent pace.

  3. Previously, a developer produced a limited amount of diff each day, and the reviewer could read every line before merging.

  4. An agent does in an hour what used to take a week, while the reviewer remains the same one person with the same limited attention.

  5. The PR queue grows faster than it can be processed, and at some point review either becomes a formality or blocks everything else.

  6. Both outcomes are equally bad for a company that promised its client a result by a deadline.

What companies do instead of reading line by line

The teams covered by The Pragmatic Engineer are adopting risk-based triage: critical paths (payments, authorization, and schema migrations) are reviewed manually, while everything else goes through an AI reviewer or is not read line by line at all. Instead of the diff, they check the outcome: whether tests pass, the data schema remains intact, and system behavior has not changed where it is measured. Control shifts from how code is written to the result—where it can actually be verified automatically.

Context determines whether there is anything worth checking

Dex Horthy of HumanLayer describes the root cause differently: an agent does not become ineffective on its own; it degrades when its context window is cluttered with irrelevant fragments—he calls this trajectory poisoning. In his assessment, managed context makes development agents 2–3 times more productive without sacrificing quality; unmanaged context produces code that no one can later review because it was written from a mix of outdated instructions and random files the agent picked up along the way.

Context engineering is not an option for enthusiasts; it is what makes risk-based review meaningful: when the agent's input is clean, its output is predictable and easier to validate with metrics rather than line by line.

Assess where AI can deliver impact in your process

Simple looks easy, isn't

The result business sees—“AI completed the migration in two weeks”—looks simple precisely because it rests on engineering no one sees: clean agent input context, test coverage that catches regressions automatically, and gates that prevent code from advancing until it passes validation.

At KT.Team, we build precisely this invisible layer: LLM & Security Gateway to control what the agent sees and where it sends data; MCP for structured access to internal systems instead of arbitrary chat context; and RAG to add up-to-date documentation instead of outdated prompt instructions. Each component reduces the chance of trajectory poisoning at the input—and therefore reduces the amount of output code that requires manual review.

Where manual review cannot be replaced—and why that is fine

  1. Hillel Wayne, a formal verification researcher, is candid: TLA+ and formal methods will remain niche tools for distributed systems—AI will broaden access to them but will not make them mainstream.

  2. This confirms a broader principle: code is not equally critical, and not all code can be validated with metrics.

  3. For a payment gateway or distributed-state consistency, risk-based triage keeps human review mandatory.

  4. The difference between a company drowning in a PR queue and one that completes a migration in two weeks is how precisely it drew this boundary in advance.

Conclusion

Line-by-line review of AI-generated code applies an industrial-era process to a pipeline that produces an order of magnitude more. Companies that continue either slow agents to human speed and lose the point of using them, or merge unchecked code and face incidents. The practical alternative is risk-based triage, outcome-based validation (tests, schemas, and metrics) instead of diff review, and investment in clean agent input context—not just faster responses.

Anyone who does not adapt the review process to the new speed of code production is postponing a bill that will arrive with the first serious incident.

Discuss the article: Code review is dead: agents write faster, …

Enter your email or phone number so we can get back to you.

Send via: