You can assign an agent a task, wait for changes to appear in a repository, and receive a draft pull request to review. But code also raises a question: does the result meet the task’s requirements, and have the tests missed an important problem?
On February 4, 2026, GitHub announced the public preview of Claude and Codex as coding agents. The company said that agents can be assigned tasks from issues and pull requests, after which developers can review the draft PR they prepare. This confirms that the workflow has appeared in a development interface, but it does not prove that it saves time or improves code quality.
The question, then, is not only whether an agent can write code. It is also what work people have to do around a delegated task: set constraints, monitor progress, and assess the result.
Oversight Is More Than a Final Review
In a preprint published on arXiv on September 21, 2026, researchers propose viewing oversight of a coding agent as a sequential process. The Work Behind Delegation draws on observations and workflow diagrams from 19 experienced developers. The authors describe seven stages of oversight and apply this framework to public developer discussions on Reddit.
Among the approaches the authors describe are paying more attention to planning, assigning some oversight tasks to other agents, and turning recurring instructions into reusable materials. This is an analytical framework, not a measure of time spent: the paper does not establish how long developers spend overseeing agents or whether this process is faster than doing the work manually. The arXiv page marks the preprint as under review, so its findings should be considered preliminary.
For practical purposes, it helps to distinguish three activities. Task definition sets the goal and constraints. Monitoring helps identify whether the agent has strayed from the intended course. Review makes it possible to assess whether the result can be accepted in light of the requirements, code, and tests. Success at one stage does not guarantee success at the next: a patch may look plausible but solve the wrong problem or require substantial rework.
Community Experience Is Not Statistics
In a discussion on r/LocalLLaMA, one participant wrote that their experience with local models and agents had been disappointing: they said they had to fix the output and remind agents of their instructions. Other participants in the same thread described a workflow they found more useful: limiting task scope, working in stages, and carefully reviewing the code. These are individual user observations, not a controlled comparison of models or a measure of team productivity.
These differing accounts show that experience may depend on the task, model, environment, and way of working. But a single discussion cannot establish which practices are more effective on average or how often problems arise. The thread focuses on local models, so its observations should not automatically be applied to all agent tools.
Human involvement does not, by itself, mean that delegation is useless. A developer can break a task into parts, review changes, and clarify requirements while remaining responsible for the outcome. But without accounting for time spent defining the task, making corrections, and reviewing the work, it is impossible to say whether the overall workload has decreased.
A Draft PR Is Not an Accepted PR
In the GitHub workflow described above, an agent can be assigned an issue and produce a draft pull request. Between its creation and acceptance, separate questions remain: does the change meet the requirements, are edge cases covered, are the tests sufficient, and does the solution fit the project’s architecture?
These criteria cannot be reduced to a single metric. The number of changes created is not the same as the number that are usable, and passing tests does not necessarily confirm every important property of a patch. Even a result that looks high-quality does not, by itself, show how much work went into preparing and reviewing it.
To assess the overall effect, the full cycle must be considered: defining the task, waiting for the result, reviewing it, making corrections, and maintaining the code afterward. The sources considered here do not provide such an overall comparison.
What We Can Conclude
The preprint offers a framework for discussing oversight, while GitHub has described a workflow in which an agent is assigned a task and its draft PR is reviewed. The Reddit discussion shows that individual users report both difficulties and useful ways of working with local models. Taken together, these materials raise the question of how to organize review of delegated code, but they do not answer whether it delivers a net productivity gain.
They do not show that developers as a whole have already shifted from writing code to overseeing it, that agents invariably create extra work, or that they increase productivity across the industry. Such conclusions would require comparable measurements of time and quality across different tasks and teams.
The practical question for a team is more specific: which changes can an agent prepare on its own, what needs to be checked before merging, and who is responsible for ensuring that the result solves the original problem? Delegation does not eliminate this engineering work—it changes where it begins and what it focuses on.