Evaluate coding agents on your own repository, not on a leaderboard

Most teams evaluate coding agents by reading a leaderboard, and a leaderboard answers a different question from the one being asked. It says how an agent did on tasks drawn from a fixed set of public repositories. The question a team has is how it will do on theirs, with their build, their test suite and their conventions, and the only way to answer that is to build an evaluation from the repository itself.

This post is the procedure: what public benchmarks measure, how to build a task set from your own closed issues, what to measure beyond pass rate, and why the evaluation has to run again every time the model underneath changes.

What a leaderboard measures and what it cannot

The dominant format comes from SWE-bench: a task is a real issue from a public repository plus a snapshot of the code before the fix, and the agent's patch is accepted if the tests that failed before the fix now pass and the tests that passed before still do. Earlier benchmarks such as HumanEval work at the level of a single function with a specification, which measures something narrower.

Repository-level benchmarks are a real improvement, and they still cannot tell you about your repository. The language mix, the build system, how long the suite takes, how the code is organised and what a reviewer there will accept are all specific. Public tasks can also have been seen during training, which inflates scores in a way nobody can measure from outside.

Use the leaderboard to shortlist. Use your repository to decide.

Build a task set from your own closed issues

The recipe mirrors how repository benchmarks are constructed, applied to code only you have.

  1. Select closed issues that were resolved by a merged change which added or modified tests. Issues fixed without a test are not usable as tasks, because there is nothing to check.
  2. Snapshot the repository at the parent of the fix. That snapshot plus the issue text is the task.
  3. Take the tests from the fix. In a clean container, confirm they fail on the snapshot and pass with the fix applied, and that the tests which pass on the snapshot still pass with the fix. Discard any task where that does not hold.
  4. Strip the fix. The agent gets the snapshot, the issue text and the ability to run the suite, not the solution and not the held-out tests.
  5. Label each task by area of the codebase and by the size of the human fix, so results can be read per area rather than as one number.
  6. Keep the set private. A task set that is published becomes training data.

This is the same verification that the code domain of Runix Data applies to every code task: run in a clean container, target tests pass with the patch and fail without it, and the rest of the suite still passes. If you also fine-tune on your own repository, the other rule from that layer matters too: evaluation data is split from training data by source, so that near-duplicates cannot sit on both sides.

How to evaluate coding agents beyond pass rate

Pass rate is necessary, and it is the metric most likely to mislead on its own. Agents are not deterministic, so run each task more than once and report a distribution rather than a single number. Then add the measures a team actually pays for.

MeasureWhat it catchesHow to get it
Diff size relative to the human fixOver-scoped changes that hide the real oneLines and files changed, divided by the merged fix's
Files touched outside the labelled areaChanges a reviewer did not expect to readCompare paths in the diff with the task's label
Review timeThe true cost of a passing diffReviewer minutes, or rounds of review before approval
FlakinessTasks the agent solves only sometimesVariance of outcome across repeated runs of the same task
Cost per merged changeExpensive successes and wasted failuresTokens priced per model plus review time, divided by changes merged

Two more are cheap and telling. Whether the agent modified tests, which should be read as a separate signal rather than folded into pass rate. And how often it stopped without a diff, which is an honest failure and better than a wrong one.

Cost per merged change needs cost attached at call time, per run and per task; cost attribution that survives audit explains why it cannot be reconstructed later. Review time needs the changes to arrive as diffs, which is the subject of why agent changes should land as reviewable diffs.

Run it again every time the model changes

An evaluation is a baseline, not a verdict. Providers retire and replace model versions on their own schedule, harnesses change their tools and prompts, and either can move the results without any error appearing. Your model id has an expiry date, and when it expires the evaluation is how you find out what the replacement did.

The same reasoning that applies to telling whether a model change made things worse applies here, with agent-shaped signals: whether the patch applies at all, the diff size distribution, the rate of runs that end without a change, and the share of diffs that reviewers merge. Compare populations of runs, not single runs, and keep the old numbers.

Refresh the task set as well. Issues closed since the last build are new tasks the model has not seen, and retiring the oldest keeps the set close to the code as it is now.

What to do with the result

The output is not a score but a map: per area of the codebase, how often the agent produces a mergeable change, how much it costs, and how much review it needs. That map decides where the agent may run on tasks unattended, where it should only assist, and where it should not be pointed yet. It also sets the spend ceiling with evidence rather than a guess, and gives the review gate a prior for how closely to read.

Runix Code is in development. It is designed as a coding agent grounded in your repository that lands every change as a reviewable diff, which is what makes diff size, review time and merge rate measurable in the first place. A coding agent that writes silently cannot be evaluated this way, and that is a reason to prefer one that does not.

Questions this raises

Why not just use SWE-bench scores to pick a coding agent?

Leaderboard tasks come from a fixed set of public repositories and may have been seen in training. They are useful for a shortlist; only tasks built from your own repository tell you how an agent will do on your code.

What makes a closed issue usable as an evaluation task?

It was fixed by a merged change that added or modified tests, and in a clean container those tests fail on the pre-fix snapshot and pass with the fix while the rest of the suite still passes. Issues fixed without a test cannot be checked.

How often should the evaluation be re-run?

Whenever the model version or the harness changes, and on a schedule in between, since providers replace models on their own calendar. Keep previous results so the comparison is between populations of runs rather than single runs.

Related to this post: Runix Code. Tell us what you are building and we reply within one business day.