AI · Article 23 of 72 · Part 5

How to Build a Multi-Agent Coding Team

Define planner, implementer and independent reviewer contracts with ownership, handoff artifacts and acceptance checks.

A coordinator assigns bounded implementation and verification work, then reconciles evidence in a combined acceptance review.
Separate file ownership and acceptance evidence make parallel work reviewable.

Adding another coding agent is easy. Giving it a useful responsibility is harder. If two agents receive the same broad request, both may inspect the same files, make incompatible assumptions, and return patches that each appear reasonable. The builder then becomes a full-time reconciler of work that was supposed to reduce effort.

A multi-agent coding team needs the same basic clarity as any development team: a shared purpose, bounded responsibilities, explicit ownership, usable handoffs, and an independent definition of completion. Agents can operate quickly, but speed increases the cost of ambiguity. A mistaken interface can spread across several implementations before anyone reviews it.

The central question is how to divide work and converge on one reviewed change. Begin with a single accountable coordinating role. Give specialists bounded contracts. Preserve one source of truth for acceptance. Make each handoff an artifact rather than an optimistic message. Treat independent review as a challenge to the result, not as another member voting for it.

This cornerstone belongs to The AI Software Factory, within the AI collection. Deterministic orchestration handles scheduler state and terminal outcomes. Parallel repository work handles isolation. This chapter concentrates on roles, task contracts, and the evidence that lets a team converge.

Start with a reason to use more than one agent

A team is worthwhile when the work has meaningful separable responsibilities or benefits from an additional review perspective. It is not automatically worthwhile because a tool can spawn several agents. Each additional participant adds context preparation, handoff, coordination, and review cost.

A small wording change may be best handled by one agent and a quick inspection. A feature crossing a route, data adapter, and customer interface may justify separate implementation responsibilities if the contract between those pieces is clear. A consequential migration may justify a read-only reviewer even when a single implementer makes the patch.

Distinguish parallelism from specialization. Several agents can work sequentially in specialized roles. Several can also work concurrently on independent tasks. Specialization asks who is responsible for a kind of judgment. Parallelism asks which work can proceed without waiting. A good team design does not assume that every specialized role should run at the same time.

Write the expected benefit before building the team. It might be reduced elapsed time on independent modules, fewer missed acceptance cases, or a clearer division between implementation and review. Then include the coordination cost in the evaluation. If the benefit disappears after reconciliation, the team structure has not earned its place.

A useful counterexample is a tightly coupled change in one small function. Splitting its implementation between agents can create more negotiation than work. A second agent may still review the final behavior, but that is a different responsibility from dividing the edit. Team size should follow the dependency structure and consequence of the task.

One role owns convergence

The coordinating role is responsible for translating the authorized request into a bounded plan, assigning ownership, reconciling outputs, and delivering one coherent result. It does not need to write every file. It does need to know what has been accepted, what is unresolved, and which decisions remain with the user or release owner.

OpenAI’s Agents SDK distinguishes a manager calling specialist agents as tools from handoffs that transfer active responsibility. That is a useful implementation distinction for deciding who owns the user-facing result. The official orchestration documentation describes those patterns. Neither pattern removes the need for explicit team contracts.

For a coding task, the manager-style arrangement can preserve one place where evidence is reconciled. A specialist returns a bounded artifact. The coordinator checks whether it meets the contract, resolves cross-cutting conflicts, and prepares the combined result. A handoff can also be appropriate, but the receiving role must know that responsibility has actually moved.

Do not leave convergence to whichever agent finishes last. Completion order is not authority. An implementer may finish quickly while another has discovered a compatibility issue. The coordinator needs the state of all required work and the acceptance conditions before declaring the integrated change ready.

The coordinating role also controls scope. If a specialist discovers an adjacent cleanup, it records the finding. The coordinator decides whether it is necessary for the requested result or belongs in later work. Allowing every agent to expand the assignment independently turns a bounded task into an unplanned portfolio of changes.

Separate planning from unsupported certainty

The planner identifies the affected behavior, current baseline, dependencies, and possible work units. Its output should make uncertainty visible. A plan is not evidence that the implementation will succeed, and a confident decomposition is not proof that the pieces are independent.

Ask the planner to inspect the relevant repository and cite the files supporting its map. It should identify shared files, generated artifacts, public interfaces, migration requirements, and operational assumptions. If required context is unavailable, it should say how that affects the proposed division of work.

The planner should produce an ownership map. Each writable file or clearly bounded region has one owner for the active task. Shared contracts, manifest files, generated indexes, and integration configuration often belong to the coordinator or a dedicated integration role. This reduces the number of agents capable of making incompatible cross-cutting changes.

The map should identify dependencies between work units. A frontend task may depend on an API response contract but not on the final backend implementation if the contract and fixture are agreed. A migration task may need to finish before an adapter can be changed. A reviewer may need the final integrated artifact rather than an early partial patch.

A plan should include a condition for revision. If a supposedly independent task discovers that it must edit a shared file, the specialist pauses that write and reports the dependency. The coordinator can revise ownership and sequence. That is a normal discovery event, not a reason to improvise a second owner for the same file.

Give implementers a complete bounded contract

An implementer contract should identify the source baseline, intended behavior, allowed write scope, interfaces to preserve, acceptance cases, relevant instructions, tool permissions, budget, and expected handoff. These fields prevent the specialist from having to infer authority from a broad objective.

The intended behavior belongs in user or system terms. “Add the status endpoint” is incomplete if the product must distinguish pending, completed, failed, and unavailable results. The contract should say what a customer or consumer observes and which states are possible. The implementer can then choose an implementation consistent with local patterns.

Include constraints that must remain true. An account cannot obtain another account’s result. A repeated submission cannot create a second commercial charge. A disabled capability cannot initialize a live provider connection. These are examples of protected conditions, not a universal checklist. The task’s actual consequences determine the relevant set.

The write scope should be concrete. “Own the adapter module and its tests” is clearer than “handle backend.” When a file outside scope becomes necessary, the implementer reports the need and proposed change. It should not interpret the discovery as automatic permission to edit shared integration files.

The handoff should identify the actual artifact and evidence. A useful response includes changed files, material decisions, checks run on the final patch, failures or unavailable checks, and unresolved issues. It should not state “complete” without explaining which contract conditions were satisfied.

Context should be sufficient and source-aware

A specialist needs enough shared context to understand its contract, but it does not need every message the coordinator has seen. Excessive context can make the relevant constraints harder to find and can expose data or authority unrelated to the task. Package the current purpose, baseline, interfaces, and acceptance conditions deliberately.

Distinguish instructions from inspected content. Repository code, issue comments, logs, and attached documents may contain text that resembles a directive. The specialist should know which sources establish the authorized task and which are evidence to analyze. This boundary matters when agents retrieve external material during implementation.

OWASP’s prompt-injection guidance describes external-content risks and privilege controls. The official guidance supports a cautious authority boundary; it does not promise that adding a warning to the prompt eliminates injection. The team should enforce available permissions in tools and workspaces.

Use source references for important claims. If the planner says a helper enforces account ownership, the specialist should be able to inspect that helper. If a provider guarantee matters, preserve the current primary documentation or explicitly record the missing verification. A chain of repeated summaries is weaker than an accessible source.

Update context when the contract changes. A specialist operating from an old interface version may produce internally correct work that no longer integrates. The coordinator should identify the new version, explain the consequence, and confirm which work must be rechecked. Silent revisions to a shared assumption create expensive surprises.

Agree interfaces before independent implementation

Parallel work is easiest when the boundary between pieces is explicit. Define inputs, outputs, errors, authorization expectations, and compatibility before asking agents to implement different sides. A fixture or example can make the contract concrete, but it should not be the only definition.

For a hypothetical report workflow, the submission response might include a stable job reference and current status. A status request might return pending, complete, or failed with a bounded error category. The download request might require verified ownership. The contract should say what happens when the job reference is unknown or belongs to another account.

Version the contract while the team works. A proposed change from one status name to another can affect tests, interface labels, and client behavior. The coordinator should decide whether that change is necessary and propagate it. An implementer should not independently rename states because the new words appear cleaner.

Preserve ambiguity when it remains. If the product has not decided whether failed jobs can be retried under the same reference, the contract should identify that open question. The team may resolve it before implementation or constrain the first release. It should not let each side make a different guess.

The frontier-model architect can assist this preparation when several constraints interact. The contract still needs review against the actual business promise and current repository. Architectural assistance is not a replacement for an accountable acceptance decision.

Independence in review needs more than another model

A reviewer is independent in function when it can challenge the result against criteria not supplied solely by the implementer. It should receive the requested behavior, source baseline, final artifact, protected cases, and relevant evidence. The implementer’s narrative can help, but it should not define the truth of the patch.

A different model may still share the same mistaken assumption. A reviewer using the same incomplete specification may reproduce the implementation’s blind spot. Model diversity can be useful, but independence is primarily a property of evidence, role, and acceptance authority. It should not be claimed from a model name alone.

A read-only reviewer can inspect scope, public contracts, error paths, data boundaries, and verification claims. It can run authorized checks or request a concrete missing check. It should separate blocking failures from optional improvements. An endless list of preferences prevents convergence without necessarily improving the product.

Give the reviewer explicit questions. Does the artifact implement the promised trigger and outcome? Does it preserve the protected cases? Does its evidence correspond to the final revision? Does it cross an authority boundary? Are limitations reported accurately? These questions keep the review attached to the result.

The independent verification chapter develops this distinction further. In a coding team, the practical rule is that an implementer cannot become the sole authority deciding whether its own change satisfies the independent contract.

Review disagreements need a resolution process

Two agents can disagree for good reasons or because one has misunderstood the context. The coordinator should ask each for a specific claim, source, consequence, and proposed check. It should not resolve the disagreement by counting votes or choosing the more confident explanation.

Suppose the reviewer says a retry creates a duplicate record while the implementer says the database constraint prevents duplication. The coordinator can inspect the constraint, the transaction behavior, and a protected repeated-request case. The relevant evidence resolves the question more effectively than another round of broad discussion.

Some disagreements reveal a product decision. One agent may assume a failed report should be retried automatically; another may assume the customer must request a retry. Neither can establish the intended behavior from code alone if the business has not decided it. The coordinator should preserve that uncertainty and obtain the missing decision under the existing collaboration process.

Avoid letting review rewrite the assignment indefinitely. Once the blocking conditions are satisfied, optional improvements can be recorded separately. If new evidence exposes a consequential gap, the plan changes. If the reviewer merely prefers another style, the coordinator can retain the existing pattern and proceed.

Record the resolution near the artifact. A future maintainer should be able to see why an apparently odd condition exists. A concise decision note can preserve the counterexample and check that settled the issue without reproducing the entire conversation.

Handoffs should survive interruption

A handoff is useful when another role can continue from it after the original conversation is unavailable. It should identify the artifact, baseline, completed work, remaining work, assumptions, and evidence. “Everything looks good” cannot support that continuation.

For implementation, name the changed source and the final checks. For research, identify the primary passages read and the limits of inference. For architecture, include the decision record and unresolved constraints. For review, list blocking findings, optional observations, and the disposition of each required case.

Keep the handoff separate from the finished product where appropriate. An app’s interface should not display internal planning notes. A published article should not include its editorial ledger. The team needs those records for review and recovery, but the customer needs the product’s actual content and behavior.

A handoff should be current at the moment responsibility moves. If an agent changes the artifact after submitting it, the coordinator needs a new version and evidence. Otherwise the reviewer may approve a patch that no longer exists. Use a revision or artifact identifier to bind the review to the actual result.

Interrupted work should remain explicitly partial. If an agent stops after discovery, preserve what it learned and what it did not do. Another agent can take over the bounded task without treating the discovery summary as a completed implementation. Accurate partial state is more useful than a premature success message.

Converge through one integration path

Individually acceptable pieces can fail together. The combined system may have conflicting names, missing configuration, incompatible assumptions, or a shared dependency changed by one work unit. Integration deserves its own owner and checks rather than being treated as a mechanical final copy.

The coordinator should integrate only artifacts that meet their handoff conditions. It can then inspect the combined diff, resolve shared files, and run checks that exercise cross-boundary behavior. A successful local test in each specialist’s scope does not establish that the product works as a whole.

Keep the integration baseline explicit. If upstream changes during the work, reconcile them deliberately and rerun the checks affected by the new baseline. Avoid attaching old specialist evidence to a materially different integrated artifact without assessment.

The final result should have one coherent explanation. It identifies the problem, resulting behavior, material choices, validation, and limitations. It should not require the user to read several agent conversations to determine whether the task is complete. The coordinator’s job is to make the team output understandable.

Detailed workspace isolation belongs in parallel development. At the role level, the principle is simple: ownership prevents competing writes, and integration establishes the accepted combined state. Separate conversations do not by themselves provide either guarantee.

Authority should follow the task, not team size

Delegation cannot create authority the user never supplied. A coordinator can give a specialist a portion of the authorized work, but it cannot use the specialist to bypass a release boundary or access restriction. Each role should receive the minimum permissions necessary for its contract.

This principle also prevents needless approval loops. When the user has already authorized a reversible edit or a defined release action, the team should carry that work forward under the established conditions. It should not ask again merely because a different agent now owns part of the task. Record the relevant authority and make it available to the role that needs it.

Consequential actions should have a concrete reviewable artifact before any required final approval. A team can prepare the patch, evidence, migration plan, and deployment candidate while a release gate remains unresolved. Asking for abstract permission before preparing the result transfers planning work back to the user.

If a role encounters an action outside scope, it should explain the exact dependency and continue independent authorized work. It should not abandon the entire assignment when only one boundary is unresolved. Nor should it quietly expand access to keep the schedule moving.

The team contract should identify which role may prepare, review, execute, and verify a release. One process may hold several roles, but the evidence should still distinguish those events. A completion report is not proof that the release occurred, and a release operation is not proof that live behavior passed verification.

Budgets limit coordination as well as implementation

Agent teams can consume budget through repeated planning, duplicate research, long handoffs, and recursive delegation. Set limits on the number of active specialists, attempts, elapsed work, and model or tool spend. The limits should be visible to the coordinator and appropriate to the task.

A specialist should report diminishing progress. If it has repeated the same failed check without new evidence, another attempt may be wasteful. The coordinator can repair the environment, narrow the task, escalate to a different capability, or preserve a bounded failure. An endless retry loop is not persistence in a useful sense.

Do not measure progress by messages or tokens. Measure completed contract conditions and accepted artifacts. A quiet agent producing one correct bounded patch may be more valuable than a team producing a large amount of coordination text. The user should receive concise updates about findings, uncertainties, and meaningful milestones.

Include coordination effort in the evaluation. The manager’s time preparing tasks and reconciling work is part of the workflow. If a team improves elapsed implementation time but increases total review and correction effort, the business needs to see both outcomes.

The model-routing chapter handles candidate selection and fully loaded cost. Team design adds a second question: whether an additional role contributes enough independent value to justify its overhead. The cheapest specialist is still expensive if its output creates more reconciliation than it resolves.

A proposed report-feature team

Imagine an existing product needs a report workflow with submission, progress, and download. The current repository has a route layer, a processing module, and a customer interface. The request requires account isolation and recovery from interrupted requests. This is a hypothetical example, not a claim about a deployed Salars app.

The coordinator first inspects the baseline and defines the behavior contract. A planner maps existing modules and identifies shared files. The coordinator owns the public contract and integration configuration. One implementer owns the processing adapter and its scoped checks. Another owns the customer interface against the agreed fixture. A reviewer remains read-only.

Before work begins, the team agrees that submission returns a stable reference, status has defined outcomes, and download requires verified ownership. The unresolved retention promise is recorded as a business decision. The first implementation can be constrained until that promise is settled rather than making a permanent storage assumption.

The processing implementer returns its artifact, success and failure checks, and a note about a provider behavior that could not be verified locally. The interface implementer returns states and accessibility behavior against the contract. Neither edits the shared manifest or release settings. The coordinator resolves the unavailable provider check through the appropriate authorized environment.

The reviewer receives the combined artifact and protected cases: repeated submission, unknown reference, another account’s reference, processing failure, and interrupted download. It reports a concrete mismatch if the interface treats failure as pending forever. The coordinator resolves the mismatch and refreshes the affected evidence.

The final handoff identifies one source revision and one release candidate. It distinguishes local acceptance from live verification. If release execution is authorized and conditions are met, the designated role proceeds and records the actual result. If a required boundary remains unresolved, the team delivers the prepared artifact and the precise remaining condition.

Keep a contract change log

A team can lose alignment even when every member follows instructions faithfully. The coordinator may discover a new requirement after assigning the work. The planner may correct a mistaken repository map. The product owner may clarify an edge case. These changes need a visible record so that old and new assumptions do not coexist unnoticed.

A compact change entry can state the affected contract version, the new condition, the reason, and the work requiring recheck. If download authorization changes, both the route and interface tasks may be affected. If a wording preference changes, only the interface task may need adjustment. The record helps avoid restarting unrelated work merely because one assumption changed.

Ask receiving specialists to acknowledge material changes through their next handoff. The acknowledgement should identify the artifact or check updated, not merely say that the message was read. This creates evidence that the change reached the implementation rather than remaining in coordination text.

Close superseded findings deliberately. A reviewer observing an old mismatch should be told which new artifact resolves it and which check establishes the correction. Otherwise a team can spend several cycles debating a problem that no longer exists while failing to inspect the latest result.

Test the team against a simpler baseline

A proposed evaluation should compare the team with a documented single-agent or existing human-assisted process on a bounded task class. Define success independently before running either. Include correctness, acceptance completion, reviewer effort, scope compliance, elapsed time, and total cost.

Use development cases for tuning and protected cases for assessment. A parallel-friendly case can test whether the ownership map and interfaces reduce elapsed work. A tightly coupled small case can test whether the team recognizes that division is unnecessary. A disagreement case can test whether the coordinator resolves evidence rather than votes.

Also include an interruption case. Stop a specialist after a partial artifact and ask another to continue from the handoff. The goal is to test recoverability of the process, not to claim that interruption is beneficial. The record should make it clear what was executed and what remains.

A successful comparison would support only the tested scope. It would not prove that every project benefits from multiple agents or that a larger team is better. Retain the conditions under which the result holds and revalidate when models, tools, repository structure, permissions, or task classes change.

A severe boundary failure can stop the pilot regardless of its average speed. If a specialist exceeds its write scope or the coordinator reports completion without required evidence, the team policy needs repair before broader use. A fast incorrect process is not an acceptable substitute for a slower accountable one.

What Would We Do at Salars?

We would propose a small team with one coordinating role, one or two bounded implementers when the dependency structure justifies them, and a separate reviewer for consequential changes. The first pilot would use existing patterns and explicit acceptance cases. Shared integration files would have one owner, and every specialist would return an artifact-bound handoff.

We would compare that team against the current documented process using protected cases, including a small task where extra agents should not be used. Success would mean accepted behavior with useful improvements in elapsed work, missed-case detection, or recoverability at an acceptable total cost. This is a proposed evaluation, not a claim of measured Salars productivity.

We would stop expansion if coordination repeatedly obscured ownership, reviewers lacked independent criteria, or the team could not converge within its budget. A model, tool, permission, or repository change would trigger revalidation. Findings would remain bounded to the task classes actually tested.

A useful multi-agent team is a system for dividing responsibility and reconciling evidence. The agents may do much of the inspection, implementation, and review preparation. The coordinating role still has to deliver one understandable result whose behavior, authority, and limitations can be assessed.

Sources

Team roles, contracts, the report-feature example, and the Salars pilot are proposed designs. No agent-count productivity gain, universal reviewer independence, or deployed-team result is asserted.

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home