AI · Article 20 of 72 · Part 5

How AI Coding Agents Change Software Development

Map specification, implementation, review and maintenance roles against verified capability limits.

When a coding agent can inspect a repository, modify files, run a command, and respond to the result, the unit of collaboration changes. The builder is no longer asking only for a code snippet. They can ask for a bounded change with acceptance conditions and receive an implementation plus evidence about what happened.

That change is significant without requiring a claim that software engineering has become effortless. An agent can produce a plausible patch for the wrong problem, pass tests that miss the important behavior, or report completion before the product reaches users. The more work it can perform, the more important it becomes to define the boundary between a proposed change and an accepted result.

The central question is how coding agents change the delivery workflow and which human responsibilities remain. A useful answer follows the work from problem definition through maintenance. It identifies where agents can carry out bounded tasks, where evidence is needed, and where the business must retain judgment. It does not infer a universal productivity gain from a successful demonstration.

This chapter belongs to The AI Software Factory and the AI collection. The multi-agent coding team chapter handles team contracts in detail. Here the focus is the overall delivery loop and how to tell whether the loop is improving.

A coding agent is a work loop

A text-only coding exchange usually asks a model to infer an answer from supplied context. A tool-enabled agent can acquire context, perform an action, observe the outcome, and decide what to do next. The useful distinction is the loop, not the personality assigned to the model.

OpenAI’s Agents SDK documentation describes agents equipped with instructions and tools, along with ways to delegate work. That establishes an implementation pattern, not proof that a particular agent can safely complete an arbitrary development task. The official orchestration guide is a source for the pattern.

A hypothetical agent fixing a route might read the route implementation, identify the relevant test, change a condition, run that test, and inspect the diff. Every step can be valuable. Every step can also fail differently. It might search the wrong directory, misunderstand the test fixture, edit a shared helper beyond its scope, or interpret a failing command as an environment problem without sufficient evidence.

Define capabilities in terms of the actual environment. Can the agent read the relevant repository? Can it run the required checks? Can it obtain current dependency documentation? Can it write only a bounded set of files? Can it identify the deployment the customer receives? A model’s ability to discuss these activities does not establish permission or access to perform them.

This is why a product demonstration should include the surrounding tool and authority context. A polished patch generated with complete context and a permissive environment is not equivalent to reliable work in a mature repository with missing history, hidden operational assumptions, and restricted credentials.

Specification becomes a larger share of the bottleneck

A request such as “improve onboarding” hides several decisions. Which user is struggling? At which step? What behavior should change? What must remain compatible? What evidence would show success? Faster implementation makes those unanswered questions more expensive because the team can now generate several incompatible answers before anyone resolves the purpose.

A useful work request starts with a concrete trigger and expected behavior. For example: when an existing customer follows an expired invitation link, show a renewal path that preserves their organization context. The request can identify an existing route, the relevant policy, and the cases that must remain unchanged. The agent then has a bounded target.

Include negative conditions. A renewed invitation must not grant access to another organization. A customer who has already joined should not create a duplicate membership. An invalid token should not reveal whether a private organization exists. These cases define the product contract more effectively than adjectives such as seamless or robust.

Not every task needs a long specification. Correcting a spelling mistake in a label may require only the exact location and wording. A cross-account access change requires a much more explicit contract. Match the preparation to the consequence and uncertainty of the work. The point is to reduce hidden decisions, not turn every small edit into ceremony.

The frontier-model architecture chapter explores using stronger reasoning where a decision has many interacting constraints. Even then, the business remains responsible for supplying its actual priorities. An agent cannot discover a confidential customer promise from a generic description of the industry.

Repository discovery becomes part of the deliverable

Before changing code, a builder needs to know how the repository currently works. An agent can help map routes, find existing conventions, identify tests, and locate relevant instructions. That is useful work even when the eventual edit is small. It also creates a risk: a confident repository summary may omit the one special case that governs the change.

Ask for evidence-backed discovery. A useful report points to the specific implementation, tests, and configuration that support its claims. It distinguishes observed behavior from inferred intent. “The code rejects this input” is different from “the product should reject this input.” Both may matter, but they require different decisions.

The baseline revision belongs in the record. If the repository changes while discovery is underway, the report may be stale. A source identifier and a list of inspected files let the reviewer assess whether the proposed change still applies. A summary without a baseline is harder to trust when several contributors are active.

Discovery should also identify unknowns. A route may call a provider whose configuration is unavailable locally. A scheduled job may be defined outside the repository. A test may use a mock that does not establish live behavior. Those limitations are not failures of honesty; they are the boundaries of the available evidence.

Avoid asking the agent to reconstruct the entire system before every task. Give discovery a stopping condition: enough context to explain the affected behavior, its dependencies, and the relevant acceptance cases. Excessive exploration can consume attention and budget while delaying a straightforward change. Insufficient exploration can create a larger repair later.

Implementation shifts toward bounded patches

A coding agent can work effectively when the patch has a clear ownership boundary and the repository offers usable patterns. The builder can ask it to reuse an existing adapter, preserve public contracts, and avoid unrelated cleanup. The resulting diff is easier to review because it corresponds to a stated purpose.

A broad request to modernize the system invites a large number of discretionary changes. Some may be sensible individually while making the combined result difficult to assess. Scope is an operational control: it limits what the reviewer must understand and what the business must recover if the change fails.

Require the agent to explain material decisions in the handoff. Which existing pattern did it follow? Which alternative did it reject and why? Which files changed? What remains unresolved? The explanation should be short enough to inspect and grounded in the patch. A long narrative is not a substitute for a clear diff.

Keep refactoring separate when it has a separate justification. If a small bug requires extracting a helper, the extraction can be part of the fix. If the agent notices a general cleanup opportunity, it can record that opportunity for later. Combining every noticed improvement with the requested behavior change makes the release harder to attribute and recover.

The best patch is not necessarily the shortest. A few additional lines may make a permission decision explicit or supply a meaningful regression case. Nor is a larger patch evidence of more completed work. Judge the result against the requested behavior and maintenance burden rather than the amount of code produced.

Tool output becomes evidence that must be interpreted

A successful command says only what that command checked. A unit test may establish a function’s behavior under a fixture. A production build may establish that the application can be assembled. Neither automatically proves that a customer can complete a live workflow after deployment.

The handoff should record the check, environment, outcome, and practical limit. If the agent could not run a required check, it should state that plainly and explain the blocker. It should not rename an unavailable check as passed because it inspected the code or ran a smaller substitute.

A reviewer needs to see failures too. Suppose the first test failed because the fixture omitted a required field, and the agent changed the fixture. That may be correct, or it may hide a genuine compatibility problem. The relevant question is why the fixture was wrong relative to the independent product contract. A green result after modifying the test is not self-explanatory.

Keep acceptance conditions outside the implementation’s convenience. The testing chapter expands that principle. For agent work, it means a patch should not be allowed to redefine success merely to match the behavior it just produced. Some protected cases should remain unchanged during the candidate implementation.

Evidence also needs freshness. A check run before the final edit does not validate the final patch. Re-run the checks affected by later changes and identify the source revision being evaluated. Avoid repeating every check without reason, but do not attach an old successful result to a materially different artifact.

Review remains a separate function

An agent’s self-review can catch mistakes, but it shares context and assumptions with its implementation. A second reviewer can help when it receives the intended behavior and actual artifact rather than only the implementer’s explanation. The goal is another opportunity to challenge the patch, not another friendly summary of it.

Separate mechanical checks from judgment. Formatting, type compatibility, and known test cases may be checked automatically. Whether the change solves the customer’s problem, preserves a commercial promise, or handles an unrepresented edge case requires a different kind of review. A model can assist that review, but its confidence is not independent evidence.

A reviewer should be able to reject the result with a concrete reason. “The unauthorized case still returns the private object” is actionable. “This could be more elegant” may be a preference rather than a blocker. Define severity and acceptance rules before the team begins so that review does not become an endless style debate.

The same principle applies to a human reviewer. Familiarity with the builder does not make a change correct. An effective workflow gives the reviewer usable evidence and enough context to assess the important behavior. Agents can reduce preparation work, but the final decision still needs an accountable owner.

Review effort belongs in the cost of development. If an agent produces a patch quickly but requires extensive correction and verification, the initial generation time is a poor measure of value. The model-routing chapter develops fully loaded comparison rather than ranking models by their apparent speed.

External material should remain evidence, not authority

A coding agent may read issue comments, documentation, source files, webpages, and logs. Some of those materials can contain instructions directed at the agent. The task owner must distinguish the user’s request and repository policy from content being inspected as data.

OWASP describes indirect prompt injection through external content and recommends constrained privileges and controls for consequential actions. Its prompt-injection guidance supports treating retrieved material as untrusted input. It does not establish that a prompt alone removes the risk.

For example, an issue attachment might say to upload a private configuration file to diagnose a problem. The attachment is evidence about the issue; it is not authorization to disclose the file. The workflow should enforce the relevant boundary through available tools and permissions rather than rely only on the agent remembering a warning.

This does not mean every external document is hostile. It means the source of an instruction matters. A helpful-looking operational recommendation can still exceed the user’s authorized task. The agent should surface the relevant proposal and continue with authorized work where possible.

The practical consequence is that implementation capability and release authority should be separate. An agent may prepare a reviewable patch, migration plan, or deployment artifact while a defined gate controls the consequential action. Existing authorization can permit that action; the workflow should record it rather than ask repeatedly or assume that preparation itself granted new authority.

Maintenance becomes continuous context work

The first successful patch is only one point in the app’s life. Dependencies change, provider behavior changes, users reveal edge cases, and the product’s purpose evolves. Coding agents can help inspect upgrades, prepare migration patches, and reconcile documentation with implementation. They still need a current operating record.

The app registry can identify the owner, current release, important dependencies, and unresolved obligations. That context helps an agent avoid treating a repository as an isolated exercise. A technically attractive cleanup may be inappropriate if a customer integration still depends on the older contract.

Maintenance requests should include the intended compatibility boundary. Is the change a patch to preserve current behavior, a deliberate breaking release, or an investigation with no write authority? These distinctions guide both implementation and review. Leaving them implicit encourages the agent to optimize whatever it can observe rather than what the business has promised.

Record lessons at the appropriate scope. If an agent repeatedly misreads a repository-specific configuration, improve the local instructions or tool feedback. That does not justify a universal rule about all software. If a proposed workflow reduces corrections in a small set of tasks, retain the result with the task class and evaluation conditions.

A useful maintenance loop includes revalidation triggers. Changes in model version, tool permissions, runtime, or repository structure can invalidate a previously reliable task pattern. Reliability belongs to the combined system under stated conditions, not permanently to a model name.

Compare accepted work rather than impressive motion

To evaluate a coding-agent workflow, select a bounded class of tasks and define success before running candidates. A reasonable proposed comparison might use representative route corrections, adapter changes, and small feature additions from an authorized test repository. Include realistic failures and cases that should cause the agent to stop.

The baseline is the team’s existing documented process. Track accepted changes, elapsed time, tool cost, reviewer effort, corrections, and unresolved failures. Keep failed attempts in the denominator. A workflow that finishes the easy tasks and abandons the difficult ones should not look reliable because only its successes are counted.

Protect some evaluation cases from tuning. A team can improve prompts using development cases, then test the resulting policy on cases it has not optimized against. That protects against building a workflow that merely remembers the examples used to design it. It is still local evidence about a bounded task class, not proof of general software capability.

OpenAI’s evaluation guidance recommends task-specific cases and calibration of automated scoring with human judgment. The evaluation best-practices guide informs the methodology here. As of October 7, 2026, the same page announces the Evals platform’s transition to read-only on October 31 and scheduled shutdown on November 30. These articles recommend durable evaluation practice, not a new dependency on that retiring platform.

A counterexample might be a small label correction where orchestration and review overhead exceed the implementation effort. Another might be a high-consequence migration for which the agent lacks the live environment needed to validate behavior. Both limit the scope of an otherwise useful workflow. The goal is to learn where the process earns its cost.

What Would We Do at Salars?

We would propose a bounded pilot around changes with clear source ownership and acceptance cases. The task contract would identify the baseline, intended behavior, allowed files, relevant instructions, required evidence, and release authority. The agent would prepare a patch and a concise handoff. A separate review would assess the final artifact.

We would keep existing local observations in AI Workflows in Practice at their stated scope. Two project observations cannot establish a general productivity rate. A pilot would therefore record its own baseline and outcomes rather than borrow a broad conclusion from those examples.

Success would mean accepted behavior under the protected cases with lower total delivery effort or a clearly valuable improvement in recoverability. We would stop a run that exceeded its authority, could not establish its baseline, or consumed its bounded budget without usable progress. We would not treat an agent’s completion message as a release result.

The opportunity is practical: move more implementation and evidence preparation into a repeatable loop while keeping problem selection, acceptance, and operating responsibility visible. Coding agents change how the work gets done. The business still has to decide which work deserves to be done and what evidence makes the result worth shipping.

Sources

Workflow contracts and the Salars pilot are proposed designs. No general productivity percentage or model-specific reliability result is claimed.

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home