A coding task for training or evaluating an agent is a bundle: a repository at a fixed commit, a statement of what to change, a reference patch, and tests. Fail-to-pass tests are the tests that fail on the repository before the patch and pass after it, and they are the part of the bundle that makes the task checkable by execution rather than by reading. This post explains that mechanism, the companion pass-to-pass tests, why the runs need clean containers, and which tasks have to be thrown away.
The reference point throughout is SWE-bench, whose paper introducing the benchmark described execution-based evaluation on 2,294 problems drawn from real GitHub issues and pull requests across 12 Python repositories, and whose Verified subset is the human-validated selection.
What execution-based verification means
Verification by reading asks a reviewer whether a patch looks right. Verification by execution runs the tests and records the outcome, and the outcome is a fact about the repository rather than an opinion about the diff. The SWE-bench Verified dataset card defines the two test lists plainly: FAIL_TO_PASS is the set of tests resolved by the pull request and tied to the issue, and PASS_TO_PASS is the set of tests that should pass both before and after the pull request is applied.
Three runs establish a task: the unpatched repository against the target tests, which must fail; the patched repository against the target tests, which must pass; and the patched repository against the rest of the suite, which must still pass. A task that meets all three is verified; a task that fails any of them is not a task yet, whatever the issue text says.
Fail-to-pass tests define the task
The target tests are the executable form of the task statement. They come from the pull request that resolved the original issue: the tests the author added or changed, run before and after the change. Where a statement says that the parser should accept trailing commas, the fail-to-pass test is the one that feeds a trailing comma and asserts the result, and it is the only part of the task that a grader can apply mechanically to a model's patch.
That gives the tests two jobs. They discriminate: a wrong patch must fail them. And they describe: a competent engineer, reading only the statement, should be able to write a patch that passes them. The first job is checked by running; the second is the subject of the Verified section below, and is where most invalid tasks hide.
Pass-to-pass tests guard the rest of the repository
A patch can make the target tests pass by breaking everything else: deleting a validation, special-casing the input the test uses, or changing a shared function so that the target test passes and unrelated tests fail. The pass-to-pass set is the rest of the suite that was green before the patch, and it must be green after. Without it, resolving the issue means satisfying one test, a weaker claim and an easier one to game.
The set is built from the unpatched run and confirmed on the patched one. Tests that fail before the patch for reasons unrelated to the issue are excluded, as are flaky tests that fail intermittently on the same commit, because a flaky test in the pass-to-pass set turns a valid patch into a coin toss. Finding them means running the unpatched suite more than once, a cost paid once per task rather than once per evaluation.
Clean containers, every run
The three runs are only evidence if they are reproducible, and a developer's machine is not reproducible: cached builds, globally installed packages, environment variables and network access all leak into results. The SWE-bench harness runs its evaluations in Docker containers for this reason, and the same discipline applies to building tasks, not only to scoring them.
Concretely: one image per repository and base commit, with dependencies pinned and the build recorded; a fresh container for every run, so that the unpatched run cannot leave state the patched run inherits; no network during the test run, so that a result cannot depend on an external service; and the image digest recorded in the task's provenance, so that a result can be rerun later on the same bits. Test output is captured in full, because a bare "passed" without the log is a claim again.
Why a task whose target tests pass without the patch is invalid
If the target tests pass on the unpatched repository, they do not distinguish a solution from no solution. Any patch, including an empty one, resolves the task. Such a task teaches a training set nothing, and in an evaluation set it inflates every model's score equally, which hides differences between models and rewards nobody.
It happens for mundane reasons: the pull request's test changes were unrelated to its code changes; the test already passed and was merely moved or renamed; the environment the task was built in differs from the one the original tests ran in, so the behaviour the issue describes is not reproduced; the patch is a refactor with no observable change.
None of these are visible by reading; all of them are visible in the unpatched run, which is why that run comes first and why a task that fails it is dropped, with the reason recorded. Evaluation tasks also have to come from repositories that contribute nothing to training, which is the split-by-source rule applied to code.
Underspecified statements, and what SWE-bench Verified screened for
The remaining failure is in the statement rather than the tests. An issue that says the program crashes on some inputs, with no example, or that asks for one behaviour while the tests assert another, cannot be solved from the statement alone. A model graded on such a task is graded on guessing which hidden test the author wrote. SWE-bench Verified is a subset of 500 tasks from the SWE-bench test set that human annotators validated for quality; the announcement the dataset card points to describes the screening as removing tasks whose issue descriptions were underspecified and tasks whose target tests checked behaviour the issue never asked for, along with other problems that made a task unfair.
The check that produced it is one a data pipeline can run on every task: a reviewer reads the statement, reads the target tests, and answers whether the statement describes what the tests check. Where it does not, the task is either rewritten so that it does, with the rewrite recorded, or dropped. The decision and the reason travel with the task, so that the quality report can say how many were cut and why.
Runix Data's code domain applies these rules: each task is run in a clean container, the target tests must pass with the patch and fail without it, the rest of the suite must still pass, a task whose target tests pass without the patch is dropped and reported, and statements are checked against what the tests actually test. Code data from Runix Data is in early access, by engagement, with dropped tasks counted in the quality report that accompanies every Runix Data delivery.
Questions this raises
What are fail-to-pass tests?
The tests tied to an issue that fail on the repository before the reference patch and pass after it. They are the executable form of the task statement and the only part of a task a grader can apply mechanically.
Why do pass-to-pass tests matter?
They are the rest of the suite that was green before the patch and must stay green after it. Without them a patch can satisfy the target tests by breaking unrelated behaviour.
Why is a task invalid if its target tests pass without the patch?
Because the tests then cannot tell a solution from no solution, so any patch, including an empty one, resolves the task. Such tasks teach nothing in training and inflate every score equally in evaluation.