The cheapest model call is not necessarily the cheapest accepted change. A low-cost response can require several retries, extensive review, or a correction after release. A more expensive response can also waste money if the task is simple and the extra reasoning adds no useful improvement. Model routing is the policy that decides which candidate should handle a particular class of work and when the process should escalate or stop.
The word reliably carries the real obligation in the title. It needs a defined task, acceptance bar, evaluation evidence, and operating boundary. Without those, routing becomes a collection of preferences about model names. A builder may save on visible tokens while spending more on hidden repair work.
The central question is how routing can reduce development cost while preserving a stated reliability bar. Begin with task classes and independent acceptance conditions. Compare complete workflows, including failed attempts and review. Use a fallback only when there is a reason to believe it addresses the failure. Preserve a terminal outcome for work that the system cannot safely finish.
This chapter belongs to The AI Software Factory, within the AI collection. Coding-agent workflows define the delivery loop. AI evaluation suites develop the larger evaluation discipline. Here the focus is the routing decision and its fully loaded economics.
Define task classes by behavior and consequence
A useful task class describes the work, available context, and consequences of error. “Coding” is too broad. A label correction, repository search, adapter migration, and cross-account authorization change require different evidence and may fail in different ways.
One proposed class might cover read-only repository discovery with source references. Another might cover small edits that follow an established pattern and preserve a public contract. A third might cover design analysis involving several interacting constraints. A fourth might cover implementation changes with money, access, or data-handling consequences.
These classes are routing inputs, not model rankings. The same model may be suitable for several classes under different tool and review conditions. A task can also move between classes when discovery reveals greater consequence. A request that appears to change a label might actually alter an entitlement decision if the label controls a hidden branch in the product.
Specify available evidence. A model handling a task with complete source context is not being tested under the same conditions as one that must search an unfamiliar repository. A model without required tool access cannot be made reliable simply by spending more reasoning effort. Route environment failures to a different resolution path from reasoning failures.
Keep classes small enough to evaluate and large enough to matter. One class for every individual task overfits the policy. One class for the entire engineering process hides important differences. Start with a few recurring workflows and revise the boundaries when failures show that the class contains incompatible demands.
Set the acceptance bar before selecting a candidate
The acceptance bar should express the desired behavior independently of the model’s output. For a repository-discovery task, it might require correct file references and explicit unknowns. For an implementation task, it might require the protected regression cases, a reviewable diff, and a truthful account of checks run.
Some conditions should be blocking. A patch that crosses an unauthorized file boundary, exposes a secret, or misses a required account-isolation case does not become acceptable because it is inexpensive. A single average score can conceal such failures. Separate correctness, authority, evidence quality, and operational readiness.
Reliability should include completion behavior. A workflow that confidently reports success for an incomplete task is different from one that accurately returns a bounded failure. The latter may be operationally safer, but it still has a completion cost. Record both rather than collapsing every non-success into the same label.
For high-consequence tasks, the bar may include human review or additional independent checks. That condition belongs in the route’s cost and latency. Do not compare a reviewed low-cost workflow with an unreviewed strong-model response and claim that the difference belongs only to the models.
A useful routing record states the class, allowed environment, acceptance requirements, candidate configurations, fallback rule, maximum budget, and terminal outcome. Someone should be able to inspect that record and understand what the policy promises. “Use the smart model if needed” is not a sufficiently defined route.
Compare configurations rather than names
A model configuration includes more than the model identifier. Context selection, reasoning settings, instructions, tools, output format, environment, retry policy, and review process all affect the result. Two uses of the same model can behave differently because those conditions differ.
Record versions and dates. Model availability, pricing, and tool behavior can change. A route qualified under one dated configuration should not be described as permanently reliable. Recheck the official provider information before implementation and preserve the actual configuration used in the comparison.
Do not turn public benchmarks into a local routing policy without evaluation. A benchmark may reveal useful capability signals, but it does not reproduce the repository, acceptance cases, permissions, and reviewer effort of the business’s workflow. Treat it as a reason to include a candidate, not as clearance to deploy the route.
The frontier-model architecture chapter suggests where stronger reasoning may be valuable. Routing should test that hypothesis in a bounded task class. It should not assume that every architectural question deserves the most expensive candidate or that every implementation task can use the cheapest one.
Include a simple non-agent baseline when appropriate. A deterministic search, formatter, or existing script may perform the task without model judgment. Routing between models is the wrong optimization if an ordinary tool already meets the contract more predictably and cheaply. The model’s role can be preparing the inputs or explaining the result rather than replacing the tool.
Build a representative development set and a protected set
A proposed routing comparison needs cases that resemble the work the team intends to route. Include common requests, important edges, and tasks where the correct behavior is to stop or ask for missing information. An evaluation made entirely of clean successful examples will underestimate operating difficulty.
Separate cases used to tune the policy from cases used to assess it. Development cases help refine prompts, context retrieval, and fallback categories. Protected cases are held apart until the candidate policy is ready for evaluation. Repeatedly changing the route after seeing every protected outcome weakens the independence of the comparison.
Use a fixed baseline environment for comparisons where possible. Preserve the repository revision, fixtures, task wording, tool permissions, and acceptance conditions. If a candidate receives additional context, record that difference and include the preparation cost. The comparison is then between workflows with stated inputs rather than a supposedly pure contest between models.
OpenAI’s evaluation guidance recommends task-specific cases, including edge and adversarial cases, and human calibration of automated assessments. The official evaluation guide supports that narrow methodology. This routing design does not depend on the retiring Evals platform; it can be implemented with a maintained, vendor-neutral evaluation harness.
A small evaluation supports only a small conclusion. If a candidate succeeds on a few examples, retain the result as provisional evidence about those conditions. Avoid presenting an observed percentage without the sample size, failure definitions, and uncertainty. The route can begin in a reviewed pilot while the business collects more representative outcomes.
Calculate the cost of accepted work
A proposed cost comparison can include model usage, tool charges, environment time, human review, correction work, and operational follow-up. The accounting method should be explicit. A business can choose its own internal hourly values, but it should apply them consistently and label estimates.
Use all attempted tasks in the comparison. If the cheaper route fails and a stronger route finishes the task, the cost includes both attempts plus the handoff and review. If no candidate finishes, preserve the unresolved task and its spent budget. Excluding failures makes the route look cheaper than the actual operating process.
A useful summary is total measured or estimated workflow cost divided by accepted tasks, accompanied by the acceptance rate and unresolved outcomes. Cost per accepted task alone can still hide a route that abandons many tasks. Report both so the business can see the tradeoff.
For a hypothetical example, suppose ten attempted tasks use $2 each in model and tool charges. Eight are accepted. Reviewer effort totals two hours, valued for this illustration at $40 per hour. The combined cost is $20 plus $80, or $100. Cost per accepted task is $12.50, with two unresolved tasks. These numbers illustrate an accounting method; they are not actual prices or Salars results.
Now suppose a second configuration costs $5 per attempt, accepts nine of ten, and takes one hour of reviewer effort at the same illustrative rate. Its combined cost is $50 plus $40, or $90, yielding $10 per accepted task. The higher visible call cost is compatible with lower fully loaded cost in this hypothetical comparison. The result would need actual measured inputs before informing a live policy.
Route on observable conditions
A routing rule needs inputs that the system can obtain and inspect. Task type, consequence category, required tools, context size, and presence of an established pattern may be useful. Some classifications can be deterministic. Others may require a model-assisted triage step with a bounded output and its own acceptance checks.
Do not let an unreviewed natural-language label grant privileged authority. A classifier might identify a task as routine, but the permission system should still enforce allowed files and tools. Routing changes the candidate doing the work; it should not silently expand the task’s authority boundary.
Record why a route was selected. The explanation can be compact: established adapter pattern, complete local context, bounded files, standard regression suite. A later failure can then reveal whether the classifier was wrong, the candidate was unsuitable, or the environment differed from the qualified conditions.
Avoid routing on confidence alone. A model’s self-reported certainty may not correspond to correctness on the task class. If confidence is used as an input, evaluate its relationship to actual outcomes before assigning it decision power. A confident wrong patch and an uncertain correct patch are both possible.
An escalation rule can respond to an observed failed check or unresolved constraint. It should specify what additional capability the next route provides: deeper reasoning, different context, a specialized tool, or human judgment. “Try a bigger model” may not solve missing credentials or a contradictory product requirement.
Make fallback bounded and informative
A fallback is a recovery path, not an unlimited loop. Define a maximum number of attempts, an elapsed-time or cost budget, and conditions that stop the run. A task should end in accepted, rejected, blocked by a stated condition, canceled, or budget exhausted. It should not disappear into repeated generation.
Pass the next candidate the useful evidence from the previous attempt: baseline, intended behavior, actual artifact, failed checks, and unresolved decisions. Do not pass only a confident narrative that the first model wrote. The fallback needs enough information to challenge the mistake rather than reproduce it.
A correction attempt should explain what changed. If a test failed because the implementation missed a case, another implementation may help. If the test environment is unavailable, a reasoning retry may add cost without new evidence. If the request exceeds authorization, the next route should preserve that boundary.
Some failures require abandoning the candidate patch and returning to the baseline. Others allow a focused repair. State the policy so a fallback does not accumulate partial changes from several incompatible approaches. The final reviewer should receive one coherent artifact and a history of relevant corrections.
A counterexample is a route that escalates every cheap-model result because the triage rule is too uncertain. Its total cost includes the first attempt and the fallback on nearly every task. That policy may be worse than starting with the stronger candidate. The comparison should test the actual escalation rate instead of assuming that cheap-first routing always saves money.
Separate routing success from release success
An accepted development artifact may still need deployment and live verification. A routing policy should not mark a product released merely because an implementation passed local checks. The route’s terminal condition must match the task it was authorized to perform.
If the task includes preparing a release, identify the source revision, built artifact, target environment, and required checks. If it includes executing the release under existing authorization, record the actual operation and live result. Model selection does not remove the need to know which artifact reached customers.
Operational failures can inform routing without being misattributed. A provider outage may prevent a sound patch from being verified. An incorrect assumption about the provider may reveal a design weakness. The failure record should distinguish those possibilities with evidence rather than automatically blame or credit the candidate model.
The AI software cost chapter addresses product-level economics. Development routing is a narrower cost center. It can reduce engineering effort while a product remains commercially unattractive, or increase development spend to avoid a costly operational failure. The registry and portfolio process should connect those costs without confusing them.
Keep the resulting policy readable. A few well-supported routes can be more useful than a sophisticated classifier whose behavior no one can explain. Complexity in the routing system creates its own testing, maintenance, and failure burden.
Monitor drift after the pilot
A route qualified under a bounded evaluation can deteriorate when inputs change. New repository patterns, larger files, different languages, altered provider tools, or model updates can move tasks outside the tested distribution. The policy needs a way to recognize and review those changes.
Record outcomes by task class and configuration. Watch correction categories, fallback rates, reviewer disagreement, unauthorized attempts, and unresolved tasks. A rising fallback rate may indicate class drift, a context-retrieval problem, or a candidate regression. It is a signal to investigate, not proof of a particular cause.
Do not feed every observed case directly into evaluation without examining it. Logs can include private data, unusual one-off conditions, or corrupted artifacts. Curate cases according to the team’s data-handling rules and preserve the distinction between development examples and protected evaluation.
Set revalidation triggers before launch: a model change, material tool change, authority change, significant failure, or sustained shift in the routed task class. The trigger should identify who reviews the policy and whether the route becomes more constrained while uncertainty is resolved.
A route can be retired. If it no longer offers a useful cost or reliability advantage, remove it from new work and preserve the evidence about why. Keeping an obsolete candidate because it once looked cheap creates another kind of software sprawl.
What Would We Do at Salars?
We would propose a small routing pilot for a few recurring development tasks, starting with read-only discovery and bounded changes that have independent acceptance cases. We would record the full configurations and compare them with the existing documented workflow. The experiment would use development cases and a protected set, with a human reviewer calibrating any automated assessment.
Success would require meeting the blocking conditions and a defined completion bar while reducing fully loaded accepted-work cost or providing a clearly justified reliability benefit. We would report failed and unresolved attempts, not only the patches that passed. The hypothetical arithmetic above would be replaced by dated actual charges and explicit reviewer-time estimates.
We would include the counterexample of a simple deterministic task and a case where missing environment access makes escalation ineffective. We would stop a route that repeatedly exceeded its scope, concealed failures, or exhausted the budget without useful evidence. A model, tool, repository, or authority change would trigger revalidation.
The useful policy is conditional: this configuration, for this task class, with these checks and boundaries, under these dated conditions. That is a stronger foundation than calling one model cheap and another intelligent. It lets the business choose the least costly workflow that actually meets its promise.
Sources
- OpenAI: Evaluation best practices — task-specific evaluation, representative cases, and human calibration.
Task classes, routing records, fallback rules, cost examples, and the Salars pilot are proposed methods. The article does not rank current models or assert measured savings.
Loading comments…