benchmarks
Claude's PRs merge at 84%, one point under humans
Claude News
benchmarksA study of the AIDev dataset, posted as a preprint on arXiv on July 23, 2026, puts the merge rate of pull requests opened by Claude at 84%, one percentage point below the 85% recorded for human-written pull requests in the same dataset.
At a glance
- The comparison rests on the AIDev dataset, which records pull requests opened by AI coding agents alongside human-generated ones, and the authors measure how the difference between the two moves over time.
- A second strand identifies the development tasks AI coding agents are predominantly applied to and tracks how those task distributions shift across development quarters rather than treating the mix as fixed.
- The third compares key characteristics of agentic and human pull requests with an eye to software quality and to the temporal dynamics of those characteristics across the development lifecycle.
Merge rate is a coarse proxy: it records what maintainers accept, not what merged code costs to review or maintain afterwards. Read that way, 84% against 85% for humans looks like parity on the acceptance step alone, and the longitudinal framing likely matters more for teams than the single headline number, since a gap that closes or widens quarter by quarter says something different about agent output than one snapshot does.
The analysis is drawn from the AIDev dataset. The authors set out to characterise agentic pull requests in comparison with human-generated ones and to examine how their properties change across different stages of the software development lifecycle. The paper states that the impact of coding agents on software quality remains insufficiently understood.
The framing given in the paper is that recent advances in large language models and their rapid adoption across software engineering tasks have made AI coding agents an integral part of modern development workflows, while their effect on quality has not been examined in the same detail.
The authors describe the result as an empirical and longitudinal perspective on the role of AI coding agents in software development, aimed at a more nuanced account of their benefits and limitations in real-world practice. How agentic contributions evolve across the lifecycle is the gap they name.
The preprint was submitted by Iren Mazloomzadeh and posted on July 23, 2026, carrying the identifier arXiv:2607.21832 in its first version, with a DOI issued through DataCite. The listing gives the submission size as 3,703 KB and files the work under software engineering, with machine learning as a secondary subject.
What the v1 listing leaves open
The version posted on July 23, 2026 is marked v1, and no revision or journal publication is listed alongside it. The abstract gives no per-quarter numbers for the merge gap, no sample size for the AIDev extract and no list of the quality characteristics compared, so those figures sit in the paper itself rather than in the record released with the listing.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
