While tech review articles argue about what PR length is safe to let an agent produce, the review problem is set by a simpler number: how often the agent is wrong.
Specific Labs' Real-SWE benchmark, published September 2026, runs frontier agents on private, licensed enterprise codebases. The best model-and-harness combination resolves 38.8% of tasks, pass@1 averaged over eight runs per task. Everything else scores lower, down to 16.2%.
Read that the other way. The strongest agent setup, on realistic production changes across billing, tax, and multi-service migrations, is wrong on roughly six in ten tasks. Reference solutions touch a median of 11 files, against six in FrontierCode and DeepSWE.
The point for review teams is not the ranking. It is that the acceptance decision is the part that did not get faster. An agent can land a change in minutes. Deciding whether that change is correct, whether it preserves behavior across the other ten files it touched, still needs a human who understands the system. That is why review time climbs even when PRs land: every agent change carries a ~60% chance it needs real correction, and the correction is not free.
Two workflow implications.
First, treat agent-generated PRs as drafts by default, not as candidates. The benchmark says over half are wrong. Reviewing a draft and reviewing a "finished" change are different mindsets, and the second one skips the assumptions the agent should not have made.
Second, measure the acceptance rate per agent, per area of the codebase, and adjust the review depth from that, per path. If a codebase area trends low pass rates, that is where the attention budget should concentrate, not spread evenly across every PR. The codebase matters more than the raw volume, and the failure rate tells you which paths will burn the most reviewer time.
The expensive work moved downstream to the people who understand the system. Benchmarks like Real-SWE give you the number that says how much of it there is.
Top comments (1)
The number I keep coming back to is the median blast radius: 11 files touched per reference solution, against 6 on the public benchmarks. That is where our own agents fall over, not on single-function edits. We run a set of coding agents and the changes that pass review are the ones scoped to one module; once a task crosses service boundaries the diff is technically valid and still wrong. How are you normalising pass@1 against that? A task that touches two files and one that touches eleven are counted the same, so 38.8% on an 11-file median is a different animal than 38.8% weighted by changed lines. Per-path acceptance rates only really tell you where to spend reviewer time if the denominator is the size of the change, not the task count.