Back to Now
GitHub ReviewBench code review benchmark preview ecosystem Verified

GitHub opens ReviewBench for AI code review agents

The public preview provides a reproducible pull-request corpus and leaderboard, with limits on what its initial vendor comparisons show.

Why now GitHub published the ReviewBench announcement on October 5 at 15:59 UTC and made its site and repository available.

A shared test for review agents

GitHub introduced ReviewBench on October 5 as an open research preview for evaluating AI code review agents. The published corpus contains 219 pull requests from 187 public repositories across 19 languages, along with findings used to judge whether an agent catches useful issues. GitHub says the benchmark measures precision and recall and lets readers inspect results by issue severity and category. Its public repository includes the corpus, methodology and instructions for running another reviewer, while the website hosts a leaderboard and a submission path.

How to read the leaderboard

The ReviewBench site says its initial leaderboard entries were generated by the project’s own team using public code review products. The vendors did not run or verify those initial results, and the products may have changed since the recorded runs. The scores are therefore evidence about performance on this particular benchmark under its published method, not a general ranking of code review quality for every repository. GitHub says outside teams can submit their own agents for a fresh evaluation. That process and the public data make the comparison inspectable, while independent replication remains an important next step.