An agent can be uncertain about how to solve a problem. The system around it should be much less uncertain about whether the task is running, which dependencies are satisfied, who may write, how much budget remains, and what happens after failure. Those are orchestration decisions. Leaving all of them inside free-form conversation makes the work difficult to recover and the result difficult to inspect.
A swarm of agents amplifies that problem. Several tasks can finish out of order. A message can arrive after its recipient has changed state. A retry can repeat an action whose first outcome is unknown. A reviewer can inspect an artifact that has already been replaced. Each event may look ordinary in a chat transcript while the combined workflow becomes incoherent.
Deterministic orchestration means defined control rules for the workflow. Given the recorded state and an event, the system applies an explicit transition and records the outcome. It does not mean that model outputs are deterministic or that every external operation will succeed. It means uncertainty is handled through visible rules rather than hidden improvisation.
The central question is which control logic makes an agent swarm recoverable and inspectable. Begin with a dependency graph, state definitions, validated handoffs, bounded retries, budgets, permissions, and terminal outcomes. Preserve enough evidence to know what actually happened. Keep the control plane simple enough that an operator can explain it.
This cornerstone belongs to The AI Software Factory, within the AI collection. Multi-agent team design defines roles. Durable workflows develop persistence across long-running execution. Here the focus is scheduler logic: what may run, what may follow, and how the system decides to stop.
Separate reasoning from control
A model can propose a decomposition, classify a task, or suggest a recovery action. Those outputs are candidates. The orchestrator should validate them against the authorized workflow before granting work or accepting a transition. A natural-language proposal should not directly change authority, spending limits, or release state.
OpenAI’s Agents SDK describes orchestration through model decisions and through code, including staged and parallel patterns. The official orchestration guide supports that implementation distinction. The design inference here is to place consequential scheduling and permission rules in explicit control logic while using models where judgment is actually required.
A practical boundary is that the agent decides how to perform its bounded task while the orchestrator decides whether the task is eligible, which resources it receives, and whether its result meets the required handoff. A model may report that a dependency is complete; the orchestrator verifies the accepted artifact record before releasing the dependent task.
This boundary avoids circular authority. If an agent can declare itself successful, increase its budget, and authorize its own next action, the control system has become a set of suggestions. A trustworthy orchestrator distinguishes the agent’s report from the system’s accepted state.
Keep the model-facing contract clear. It should say which events the agent may submit, which evidence each event requires, and which decisions remain outside the agent’s authority. The agent can then make useful progress without having to guess how the scheduler interprets its messages.
Represent dependencies before launching work
A directed acyclic graph, or DAG, is a useful representation for a bounded development workflow. Nodes are work units. Edges state that one accepted result is required before another task can begin. The graph should represent actual dependencies, not merely the order in which the planner happened to list tasks.
For a hypothetical feature, repository discovery may precede the interface contract. Backend and interface implementation may proceed independently after that contract is accepted. Integration requires both artifacts. Final review requires the integrated revision. Release preparation requires review acceptance. Live verification follows any authorized release operation.
An edge should identify the artifact or condition it carries. “Research before implementation” is vague. “Implementation requires the accepted provider contract and the approved response schema” is more informative. It tells the scheduler what to validate and tells the operator why a task is waiting.
Validate the graph before execution. Required nodes should exist, references should resolve, and cycles should be rejected or redesigned as explicitly bounded iteration. A graph with a hidden cycle can leave every task waiting for another task that cannot finish. The operator needs a clear configuration error rather than a quiet stalled run.
Not every workflow is known completely in advance. The model can propose additional bounded nodes after discovery. The orchestrator can validate and version the graph change. Dynamic planning does not require unbounded authority: new nodes still need dependencies, ownership, budgets, and acceptance conditions.
Define node states with operational meaning
Use a small set of states that operators can understand. One proposed vocabulary is pending, ready, running, awaiting review, accepted, failed, blocked, canceled, and budget exhausted. These names are examples. Their transition rules matter more than the exact labels.
Pending means required dependencies are not yet accepted. Ready means the dependencies and scheduling conditions are satisfied. Running means a specific attempt holds the work assignment. Awaiting review means the attempt has submitted an artifact but the acceptance decision remains open. Accepted means the required conditions passed for an identified artifact.
Failed means the attempt or task did not meet the contract under a defined failure category. Blocked means a specified external condition prevents progress. Canceled means an authorized stop has ended the work. Budget exhausted means the system’s limit stopped further attempts. These outcomes should be visible rather than folded into a generic incomplete label.
Separate task state from attempt state. A task can remain active while its first attempt fails and a second attempt runs. The attempt record identifies the candidate, environment, input artifact, outputs, and evidence. The task record identifies the current accepted or terminal outcome. This distinction prevents a retry from erasing the history needed to understand cost and failure.
Define allowed transitions. Running can move to awaiting review when the required artifact is submitted. It cannot move directly to accepted merely because the agent wrote “done” unless the task’s contract genuinely permits an automatic acceptance check and that check passes. A canceled task should not return to running because a late progress message arrives.
Make events structured and versioned
Free-form messages are useful for explanation, but control events should have a defined shape. An event can include task identifier, attempt identifier, event type, timestamp, source, artifact reference, contract version, and evidence links. Validate required fields before changing state.
A completion submission should identify the exact artifact. A baseline revision, file digest, or build identifier binds the handoff to the result being reviewed. A message describing the artifact without identifying it allows later edits to silently invalidate the review.
Events also need an ordering policy. A late event from an older attempt should not replace the current result. The orchestrator can reject stale contract versions or retain the event as historical information without applying it to current state. The rejection should be inspectable so the agent or operator can correct a legitimate mismatch.
Do not infer control meaning from tone. “Looks good” may be a progress observation, not acceptance. “I cannot run this check” may be a blocked condition, a task limitation, or a request for environment repair. The structured event asks the role to state which outcome it intends and supply the relevant evidence.
A validation failure should produce useful feedback. Identify the missing artifact, invalid state transition, or stale version. A vague “event rejected” error forces the agent to guess and can create repeated submissions. Good control feedback reduces unnecessary reasoning work.
Schedule only eligible work
A task is eligible when its dependencies are accepted, its authority conditions are satisfied, its budget allows another attempt, and the required execution capacity is available. Each condition should be checked explicitly. A dependency graph alone does not address permission or resource constraints.
Concurrency limits should reflect the environment. Several independent agents can still compete for a shared build process, provider rate limit, or file. A scheduler can permit model reasoning in parallel while serializing access to a constrained resource. Independence in the task graph is not proof of independence in the operating environment.
Ownership needs a control record. A task assigned a file set should not overlap another active writer unless the workflow explicitly supports safe shared editing. Shared integration files usually deserve one owner. Detailed workspace isolation belongs in parallel development; the scheduler must still know which task holds the write responsibility.
A queue should expose why a ready task has not started. It may be waiting for capacity, a resource lease, or a spending window. Without that explanation, the coordinator may launch duplicate work because it mistakes waiting for failure. Observable scheduling state reduces improvised interventions.
Fairness also matters. A stream of small tasks should not permanently prevent a larger eligible task from running. The policy can be simple, such as a documented priority and waiting rule. Avoid building a sophisticated optimizer before the business has evidence that scheduling efficiency is its real bottleneck.
Give retries a reason and a ceiling
A retry should respond to a failure category that another attempt may resolve. A transient tool error, invalid output format, failed acceptance case, missing context, and exceeded authority are different categories. Repeating the same prompt is not an appropriate response to all of them.
For a transient failure, the policy may allow a bounded delay and another attempt. For invalid output, it may supply a specific validation error. For a failed protected case, it may request a focused correction. For missing access, it may block the task pending an authorized environment change. For an authority violation, it may stop the attempt and require policy review.
Set maximum attempts and a total task budget. Per-attempt limits alone allow repeated inexpensive failures to accumulate. A task should end visibly when the ceiling is reached. The operator can later authorize a revised bounded run with new evidence, but the original run should retain its terminal outcome.
Preserve the reason for each retry and the change in conditions. If the second attempt has no new context, capability, or correction, the system should be able to explain why it expects a different result. Otherwise the retry policy may be spending money on hope.
An unknown external outcome needs special handling. If a tool times out after a consequential operation, the orchestrator should not assume the operation failed and repeat it blindly. It should reconcile the actual state where possible. Idempotent agent work develops repeat-safe effects; scheduler logic must recognize when that protection is required.
Budgets belong in the control plane
A budget can limit elapsed time, model usage, tool spend, attempts, and operator effort. Not every limit can be measured precisely in real time. The workflow should state which values are measured, estimated, or unavailable, and use conservative stopping rules where necessary.
Allocate budget at both run and task levels. A task-level cap prevents one difficult node from consuming the entire run. A run-level cap prevents many individually bounded tasks from exceeding the business’s intended total. Include coordination and review work, not only implementation calls.
Budget changes require an authorized policy event. An agent may explain why more budget could help and what new evidence the next attempt would seek. It should not increase its own allowance by treating the original goal as unlimited spending authority.
A budget-exhausted result should preserve usable partial work. The record can identify completed discovery, a candidate artifact, failed checks, and the remaining condition. It should not discard the evidence or mislabel the task as accepted to make the dashboard look complete.
A useful control display distinguishes money spent from work accepted. A run can spend little and fail entirely, or spend more and produce a valuable verified result. The model-routing chapter explains fully loaded comparison. Orchestration supplies the accurate attempt history that such comparison requires.
Permissions must survive delegation and retries
The orchestrator should derive each task’s permissions from the authorized run and the task’s bounded needs. A specialist cannot receive authority the coordinator does not have. A retry cannot silently widen access because the first attempt found its constraints inconvenient.
Separate preparation, execution, and verification where the task requires it. A model can prepare a patch or release candidate. A defined role can execute an already authorized action after its conditions pass. Another check can establish the live outcome. The state machine should record those events distinctly.
This structure also preserves the user’s existing authorization. A new attempt or specialist should not require a repeated permission question when the authorized action and conditions are unchanged. The orchestrator should carry the relevant scope forward accurately. It should ask only when an actual required decision remains outside that scope.
Tool enforcement is stronger than narrative expectations. A task with read-only authority should use a read-only environment where available. A task owning one directory should not casually receive unrestricted write capability. The scheduler can record the intended boundary, but the execution environment should enforce what it can.
External content should not alter permissions. An issue comment or attached document can propose an action; it cannot grant access or release authority. The task’s source of authorization remains distinct from material being inspected. The agent may surface the proposal, while the orchestrator retains the control rule.
Cancellation is a real terminal path
A user can stop a run. A coordinator can cancel obsolete work under its authorized scope. The scheduler needs a defined cancellation behavior rather than assuming every agent will notice a conversational message at the same instant.
Cancellation should prevent new attempts and dependent scheduling. It should signal active workers, preserve their latest usable evidence, and record which effects may already have occurred. A task that was executing an external action at the moment of cancellation may require reconciliation before the final report can describe the result accurately.
Late outputs can be retained as historical artifacts without becoming accepted work. If a canceled implementer returns a patch, the coordinator may inspect it later under a new task. The old task should remain canceled. This preserves the truth about the run rather than rewriting history because useful output arrived afterward.
Cancellation also affects descendants. If a prerequisite is canceled, dependent tasks should become canceled or blocked according to a stated rule. They should not wait indefinitely for an accepted artifact that will never arrive. The run summary should identify the dependency consequence.
Do not equate cancellation with rollback. Stopping future work does not undo completed effects. A rollback requires its own authorized action and evidence. The orchestrator should make that distinction visible so the user understands whether the system stopped, reverted, or still needs operational repair.
Acceptance should bind evidence to artifacts
A reviewer or automated check accepts a specific artifact under a specific contract. The acceptance record should identify both. If the artifact changes, the relevant acceptance may need to be refreshed. If the contract changes, previously accepted work may need a new assessment.
Use acceptance conditions with practical limits. A local build can accept the buildability condition. It cannot accept live customer behavior. A unit test can accept a protected fixture behavior. It cannot establish provider account configuration. The orchestrator should not let one successful check satisfy unrelated conditions.
An integrated artifact often requires new checks even when every component has passed its own review. The scheduler can model integration and final review as separate nodes. That makes the combined verification a required dependency rather than an optional final habit.
The reviewer should return concrete findings and their disposition. Blocking findings prevent the relevant acceptance transition. Optional improvements can be recorded without holding the run open indefinitely. A finding resolved by a new artifact should link to the correction and refreshed evidence.
Acceptance rules should be independent of the candidate’s convenience. If a failed test is changed, the record needs the justification relative to the product contract. Otherwise the system can reach green status by weakening the conditions it was supposed to satisfy. Protected cases help preserve that independence.
Observe the run without drowning the operator
A useful run view shows task state, active attempt, waiting reason, budget, current artifact, latest meaningful event, and unresolved blockers. It should answer what is happening now and what condition determines the next step. A stream of every token is rarely the best operational view.
Preserve detailed events for investigation, but summarize them by outcome and dependency. A coordinator should be able to see that implementation is accepted, integration is running, and live verification has not begun. The distinction is more useful than a vague percentage complete.
Avoid inventing an overall progress percentage from task counts when tasks differ greatly in effort. Three completed discovery nodes and one unstarted migration do not necessarily mean the run is three-quarters complete. Report milestones and remaining uncertainty in plain language.
Logs need data boundaries. Record artifact references and necessary diagnostic information without indiscriminately copying customer content or secrets into the event store. Observability can itself create a handling obligation if it retains more sensitive data than the workflow needs.
A final run report should be self-contained. It identifies accepted outputs, executed actions, verification, failures, canceled work, budget outcome, and remaining conditions. The user should not need to reconstruct the result from several agents’ progress messages. The orchestration system exists partly to make that reconstruction unnecessary.
Recovery begins with recorded truth
If a process restarts, it must know which work was accepted, which attempts were active, and which effects have uncertain outcomes. Persistence and runtime durability are developed in durable workflows. At the scheduler level, the key is to resume from recorded evidence rather than assume every previously running task either succeeded or failed.
A recovered running attempt may require reconciliation. Its worker might still be active, have stopped, or have completed an external action without reporting it. The scheduler should identify the current attempt and check the relevant environment before launching a duplicate.
Accepted artifacts can be reused only when their baseline and contract remain valid. If an upstream dependency changed, the orchestrator should invalidate or recheck the affected descendants according to the graph. Unrelated accepted nodes need not be repeated merely because the run restarted.
A recovery operation should have its own event record. State what was observed, which assumption was made, and which check resolved uncertainty. This evidence distinguishes actual recovery from a narrative that the system “continued successfully.”
The control logic should be reproducible enough to replay state transitions for investigation. Replaying control events is different from repeating external effects. The former reconstructs what the scheduler decided. The latter can create new consequences and needs its own safeguards.
A proposed feature graph
Consider a hypothetical report feature. Discovery produces a repository map. Contract review produces accepted submit, status, and download behavior. Two implementation nodes depend on that contract. Integration depends on both accepted patches. Independent review depends on the integrated revision. Release preparation depends on that review, and live verification follows any authorized deployment.
At the start, discovery is ready and the other nodes are pending. The scheduler assigns one attempt with read authority and a bounded budget. The agent submits the map with source references. The acceptance check confirms the required evidence before the contract task becomes ready.
After contract acceptance, two implementers can run if capacity and ownership permit. One owns the processing adapter; the other owns the interface. The scheduler rejects a proposed shared-manifest write by either because the coordinator owns that integration file. The rejection identifies the boundary and allows the agent to report the needed change.
Suppose the processing test fails. The attempt submits the actual failure and a focused correction proposal. The retry policy allows one repair within the remaining budget. If the second attempt still fails, the node becomes failed or budget exhausted, and integration cannot start. The interface artifact may remain accepted while the run records the unresolved processing dependency.
Suppose instead the repair passes. Integration begins with both artifact identifiers and the accepted contract version. A reviewer finds that an interrupted request creates a duplicate job. The integrated artifact remains awaiting acceptance. The coordinator assigns the correction to the relevant owner and refreshes the affected checks.
Only after the final conditions pass does release preparation become accepted. If deployment is authorized and executed, the record identifies the actual artifact and target. Live verification has its own outcome. A successful deployment with failed live behavior is reported as that combination, not compressed into “done.”
Distinguish orchestration errors from task errors
An implementation can fail while the orchestrator behaves correctly. The system records the failed case, prevents dependent work from starting, and returns a bounded outcome. That is a task failure handled by the control policy. Conversely, every individual patch can appear correct while the scheduler accepts a stale artifact or loses a cancellation event. That is an orchestration failure even if the code itself is sound.
Keep these categories separate in evaluation. Otherwise a scheduler can be blamed for a model’s ordinary implementation mistake, or credited for a successful patch while its control rules remain untested. Each category has a different repair: improve the task context or candidate, revise the acceptance contract, correct the transition logic, or fix the environment.
The operator should also distinguish configuration errors from runtime failures. A missing dependency reference detected before launch is not the same as a worker disappearing during execution. The first requires correcting the graph. The second requires reconciling an attempt and possibly recovering it. Clear categories make recovery less dependent on improvised interpretation.
A final report can preserve both dimensions: the feature was not accepted because a protected case failed, and the orchestration correctly stopped integration. That truthful outcome is more useful than a single success score that conceals which part of the system needs attention.
Test control logic with protected failures
A proposed orchestration evaluation should begin with a simple serial baseline. Use the same bounded task contracts and acceptance conditions. Compare whether the orchestrator preserves correct state, avoids duplicate scheduling, respects authority, and produces an understandable terminal report. Speed is secondary to control correctness at this stage.
Protect failure cases from convenient simplification. Submit a stale completion event after a retry begins. Cancel a prerequisite while a dependent is queued. Exhaust a task budget. Return an artifact without its required identifier. Change a contract after one implementation is accepted. Simulate a tool timeout with an unknown external outcome.
Expected results should be defined before the run. The stale event must not overwrite current state. The canceled prerequisite must not release descendants. The budget limit must stop new attempts. The unidentified artifact must not be accepted. The changed contract must trigger the stated recheck. The unknown outcome must enter reconciliation rather than a blind retry.
These are proposed control tests, not a claim that a particular scheduler has passed them. Local simulation can establish behavior of the control implementation under the tested events. It cannot establish that every real provider or worker behaves as simulated. Preserve that limit and add real authorized integration checks where necessary.
A small workflow may not need a full orchestrator. A documented serial procedure can be a better starting point when the task has little branching or concurrency. The counterexample protects against turning every simple edit into infrastructure work. Build control machinery when recurring state ambiguity and recovery needs justify it.
What Would We Do at Salars?
We would propose an explicit small graph for one bounded coding workflow before expanding to a general swarm. Every node would have an owner, contract version, input and output artifacts, permitted tools, budget, acceptance conditions, and terminal states. Shared write ownership and release authority would be recorded separately from model reasoning.
We would compare it with the existing documented serial process using protected control failures. Success would require correct transitions, useful recovery evidence, no unauthorized expansion, and a self-contained result. We would not claim improved production reliability from simulation alone or infer a productivity gain from a successful demonstration.
We would stop expansion if the state machine became harder to explain than the underlying work, if retries repeated unknown effects, or if acceptance depended on unverified completion messages. A change in tools, runtime, permission model, event schema, or dependency structure would trigger revalidation of the affected rules.
Deterministic orchestration gives uncertain reasoning a stable operating frame. It lets agents explore and implement within bounded responsibilities while the system remains clear about what was authorized, what happened, what was accepted, and what must happen next.
Sources
- OpenAI Agents SDK: Agent orchestration — code-driven and model-driven orchestration patterns.
The state model, graph, event contracts, feature example, and Salars control evaluation are proposed designs. Deterministic control does not imply deterministic model output, repeat-safe external effects, or validated production reliability.
Loading comments…