Copilot wins GitHub's benchmark, ranks fourth on Martian's

GitHub built a new yardstick for AI code reviewers with Microsoft, and its own product came out on top: GitHub Copilot code review, in its Balanced configuration, leads with a 40.1% grounded F1 score. As The New Stack reports, Martian runs an independent leaderboard that ranks the same product fourth.
At a glance
- GitHub and Microsoft released ReviewBench on Monday. It is an open benchmark that asks AI code review agents to inspect 219 public pull requests from 187 repositories across 19 programming languages.
- Copilot code review in its Balanced configuration leads ReviewBench with 40.1% grounded F1. As of Oct. 6, Martian's online leaderboard has Cubic first at 64.9% F1 and Copilot fourth at 60.9%.
- GitHub's team ran every entry itself, and no vendor conducted or verified its own test. The runs were also months apart: Copilot was tested on Oct. 1, Cubic and Greptile back in June.
If you haven't been following, GitHub launched Copilot code review in October 2024 and made it available to all paid Copilot subscribers the following April. Since then the reviewer has moved to an agentic architecture that pulls in wider context from the repository. It now bills in GitHub Actions minutes on private repositories and can approve pull requests. On Sept. 28 a more thorough Balanced mode became the default.
ReviewBench draws 219 pull requests from a pool of 103.9 million
Every reviewer on ReviewBench inspects the same 219 public pull requests from 187 public repositories. GitHub says it picked them after analyzing 103.9 million pull requests, so the mix of languages and repository sizes closely mirrors GitHub overall. There is one deliberate skew: the set is weighted away from tiny, single-file changes.
The authors are Alejandro Carderera de Diego, a staff applied engineer at GitHub, and Michelle Zhou, an applied scientist at Microsoft. In a Monday blog post they explain why they thought the field needed another test:
Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.
Plenty of products could be measured. Qodo, Greptile, Cubic, Devin and Cursor all offer AI-assisted reviews, and CodeRabbit ranks high on other code-review leaderboards.
Claude Sonnet 5 sorts the findings, and a second LLM matches them
Scoring starts from an answer key. The benchmark builds a reference set of issues from human review comments, changes the authors made later, static-analysis tools and LLM reviewers. Claude Sonnet 5 classifies each finding as a true or false positive. A separate LLM matcher then decides whether a reviewer's comment points at the same underlying issue as an entry in the reference set.
Picture a teacher marking essays against a model answer that several colleagues wrote together. The hard part is deciding when a student's differently worded point is the same point, and that is the matcher's job. Grounded F1 then combines precision (how many comments flagged real issues) with recall (how many known issues got caught).
The methodology also reports how often the automated judge was wrong. People corrected 47 findings by hand that had first been classified as true positives. Human and classifier judgments agreed 96.6% of the time on whether a finding was a true or false positive.
GitHub ran every competitor's entry itself, months apart
GitHub states the caveats itself. Its team produced the first entries by running the publicly available version of each product, and the vendors neither conducted nor verified those tests. Copilot was tested on Oct. 1, Cubic and Greptile back in June. ReviewBench warns that the products may have changed since then and that results on its corpus may not carry over to a company's own code.
GitHub does publish the dataset, the methodology and the judging setup. Vendors can submit their own runs through a self-service system, and new results go on the leaderboard once they are scored. On LinkedIn on Monday, Carderera de Diego offered evidence from inside GitHub:
We've been using ReviewBench to improve GitHub Copilot Code Review, and its offline results have consistently anticipated the direction of later production experiments.
Martian's leaderboard puts Copilot fourth online and fifth offline
Martian, an AI research company, launched Code Review Bench in February. It pairs an offline test with an online tracker that follows how developers respond to review comments in real open source repositories. At launch, Martian said its offline results did not match what it saw in real-world use, so it made the online benchmark its headline metric.
As of Oct. 6, the online leaderboard has Cubic first with 64.9% F1, then Greptile and CodeRabbit, with Copilot fourth at 60.9%. Offline, Qodo Deep ranks first, then Cubic and Augment, and Copilot sits fifth with a 58% F2 score. F2 weights recall more heavily than precision, so it rewards a tool for catching a larger share of the relevant issues.
Martian's February launch post argued that credible benchmarks cost a lot to maintain, which has historically left the work to academics or vendors. Its benchmark and methodology are open source under an MIT license. The post described its own approach this way:
We're trying a third option: a well-funded research lab that doesn't train models or sell coding tools, and has no stake in which tool wins.
Three code review benchmarks, and none scores the same thing
The numbers above don't line up. ReviewBench also uses an F1-style measure for its default ranking, but it runs on a different dataset with a different scoring method. Martian's online and offline tests measure different kinds of evidence from each other.
There is also a naming clash. In July, LangChain published its own code review benchmark, also called ReviewBench. It is smaller: 59 tasks drawn from LangChain's LangSmith codebase, and no public vendor leaderboard. LangChain ran different models through the same Deep Agents setup, so it compares models under one agent configuration rather than ranking commercial products.
The obvious problem is that the vendor wrote the benchmark, its product came first, and it ran the tests for every rival. In our view the uneven dates are the bigger problem: Copilot was tested three days after Balanced became its default mode, while Cubic and Greptile were measured in June, which likely favors the freshest entry. The 96.6% agreement figure also covers only the true/false classification, not how accurately the LLM matcher pairs findings.
Whether rivals submit their own runs
The self-service submission system is the thing to watch. If Cubic, Greptile, Qodo and the others send in their own runs, GitHub's snapshot could turn into a leaderboard that vendors tested themselves. Each vendor will still choose which configuration to submit. GitHub has not given a timetable for new entries and says only that results are added as they are scored.
Related stories
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
