What is a coding agent? The short answer: a program that takes a task described in words, produces a change to a repository, and gets there by running a loop of reading code, planning, editing files and running tests until a stopping condition is met. The model does the generating. A loop with tools does the rest, and an assistant stops at the generating.
That definition hides most of the useful detail, so the rest of this post takes it apart: what separates an agent from an assistant, what the loop looks like, which tools it needs, why grounding in the repository matters, where humans keep control, and what a coding agent is not.
What is a coding agent, as opposed to an assistant
An assistant produces text. Autocomplete proposes the next few tokens; chat answers a question or drafts a function. In both cases a person reads the output, decides whether it is right, pastes it, runs it and fixes it. The human is the loop.
An agent is the same model placed inside a loop that it drives itself. It reads files, decides what to change, applies the change, runs the tests, reads the failures and tries again. The output is not a suggestion but a change to the repository, plus a record of how it got there.
The difference is not the size of the model. The same model can be an assistant or an agent depending on who executes the next step, and that is a property of the harness around the model, not of the model.
The loop: read, plan, edit, run tests, revise
Most coding agents run some version of the interleaved reasoning-and-acting pattern described in the ReAct paper: think, act with a tool, observe the result, think again. For a repository the steps are concrete.
- Read. Locate the files the task touches, follow imports, read the existing tests, and look at how similar things were done nearby.
- Plan. Decompose the task into edits, decide the order, and note which tests should change and which must not.
- Edit. Apply the changes to files, usually as patches rather than whole-file rewrites, so the change stays small enough to review.
- Run tests. Execute the build, the linter and the relevant test subset, then the wider suite.
- Revise. Treat the failure output as the next input and go back to the read step.
The loop ends when the tests pass, when a budget of steps, tokens or time is exhausted, or when the agent detects that it is cycling. Which of those happened matters as much as the final diff, which is why the trace of the loop is part of the output.
The tools a coding agent needs
An agent is only as capable as the tools its harness exposes. Three are necessary for anything beyond trivial changes.
- Repository access. Reading files, searching across them, and seeing version control history. Without history the agent cannot tell how the code arrived at its current shape.
- A shell. To build, to run scripts, to inspect dependencies and to execute the project's own tooling. The shell is also the largest risk surface, which is why its permissions are a harness decision.
- A test runner. The clearest objective signal the agent gets, alongside the build and the linter. Most of what else it observes is text it has to interpret.
Useful but optional: a language server for symbol navigation and type errors, and a way to read documentation for the dependency versions actually installed. The SWE-agent paper makes the point that how tools are shaped for the model, not only which tools exist, changes what an agent can finish.
Grounding in the repository
Generic code is the common failure mode of assistants: the suggestion is correct for a tutorial and wrong for your codebase. Your repository has conventions for logging, errors and naming, a dependency set pinned to particular versions, and idioms that a reviewer will reject a change for ignoring.
Grounding means two things. First, retrieval: the agent reads the actual modules it is changing and the ones that call them, rather than reasoning from memory of similar projects. Second, verification: it runs the project's build and tests, so that a wrong assumption about a library version fails before a human sees it.
The second half is what an assistant cannot do and the reason grounding is stronger in an agent. A claim about how your code behaves can be checked by running it. Runix Code is being designed around exactly this: grounding in your repository, its modules, idioms and the versions you run, rather than in generic patterns.
Where humans gate
Autonomy inside the loop does not mean autonomy over the repository. The gates sit in three places.
- Before the run. Which repositories the agent may touch, which commands it may execute, which secrets it never sees, and how much it may spend.
- During the run. Tests as an objective check, and budgets that stop a cycling agent.
- After the run. The change arrives as a diff for a person to approve, request changes on, or reject. Nothing merges without that decision.
The third gate is the one that keeps an existing team's review habits intact, and it is the subject of why agent changes should land as reviewable diffs. The Runix Code rollout guide describes the same structure: assist mode in the editor and terminal, agent mode for tasks, and a review gate that is not optional.
What a coding agent is not
- Not a guarantee of correctness. Passing tests means the tests pass. It does not mean the change does what the issue asked, and an agent can make a failing test pass by changing the test.
- Not deterministic. The same task can produce different diffs on different runs. Evaluation has to account for that; see how to evaluate a coding agent on your own repository.
- Not a benchmark score. Public benchmarks such as SWE-bench measure resolution of tasks drawn from particular public repositories. Your repository is not one of them.
- Not a replacement for review. It removes the typing, not the judgement. A team that stops reading diffs because an agent wrote them has removed its last gate.
- Not an employee. It has no memory of last week's decision unless that decision is in the repository, and it does not know what the issue tracker left unsaid.
Runix Code is a coding agent in development, built for teams: grounded in the repository, designed to run as an independent agent for its domain with its own harness and to land every change as a reviewable diff rather than a silent write. It sits in the agents layer of the Runix stack; the waitlist is open, and this post describes what is being built, not a shipped product.
Questions this raises
Is a coding agent the same as an AI coding assistant?
No. An assistant produces text that a person applies and verifies; an agent runs the read, edit and test loop itself and produces a change to the repository. The same model can play either role depending on the harness around it.
Does a coding agent need access to a shell?
For anything beyond trivial edits, yes: it has to build the project and run tests to verify its own work. The shell is also the largest risk surface, so its permissions should be scoped by the harness rather than left open.
Can a coding agent merge its own changes?
It can be configured to, but it should not. The change should land as a diff that a person approves, so the team keeps its review habits and the blast radius of a wrong change stays bounded.