A powerful model can write a great deal of code before the builder notices that the design is wrong. The app may use an unnecessary service, mix customer data across a boundary, or assume a workflow that the business has never validated. Faster typing does not resolve those decisions. It can make their consequences arrive sooner.
The title’s recommendation is an allocation principle: consider using stronger reasoning for the decisions that constrain many later changes. It is not a universal claim that a frontier model is always the best architect or should never implement code. A strong model may be useful for implementation, and a familiar small change may need no model-assisted architecture work at all.
The central question is which software decisions justify spending more reasoning effort. Look for interacting constraints, costly reversals, missing evidence, and important counterexamples. Then ask the model to produce a reviewable decision artifact rather than an expansive design that merely sounds sophisticated. The artifact should expose choices so a builder can challenge them.
This chapter in The AI Software Factory, within the AI collection, narrows the role discussed in coding-agent workflows. Model routing addresses how to evaluate cost and reliability. Here the focus is the substance of an architectural decision and the conditions under which extra reasoning may earn its place.
Architecture begins with obligations
A design is a response to obligations: what the app must do, who it serves, what it must protect, what it can spend, and how it will be operated. Technology names are downstream choices. Starting with a fashionable stack encourages the builder to fit the product into whatever that stack makes convenient.
Give the model the concrete workflow. Who supplies the input? Which action creates value? Which parts can fail without breaking the promise? Which outputs affect money, access, or customer decisions? What current alternative does the customer use? These questions keep architecture attached to the purpose of the software.
Then state operating constraints. A one-person business maintaining a small app has different capacity from a team that can operate several independent services. A product requiring offline use has different data and synchronization needs from an internal dashboard. A customer contract limiting where data can be processed changes the set of acceptable providers.
Distinguish facts, preferences, and unresolved assumptions. “The customer requires this export format” is a fact only if the requirement has evidence. “We prefer one runtime” may be a reasonable maintenance preference. “Users will tolerate a delayed report” is an assumption to test. Mixing these categories makes the architecture look more certain than the product really is.
A useful first artifact is a constraint table with evidence, consequence, owner, and uncertainty. The model can help organize it and identify contradictions. It should not invent evidence to fill empty cells. An unresolved constraint is often the most valuable result of the exercise because it identifies the next conversation or test the business needs.
Spend reasoning on decisions with a large consequence surface
Some choices affect one implementation detail. Others spread through the app’s data, interfaces, operations, and commercial commitments. The latter deserve more deliberate analysis because mistakes become harder to isolate after customers and integrations depend on them.
Consider a hypothetical document-processing product. Choosing a button label is a small reversible decision. Choosing whether uploaded documents are retained, which account owns them, and whether processing can be repeated safely influences storage, access, support, billing, and deletion. A stronger model may help enumerate those consequences, but the retention promise still requires an accountable decision.
Look for boundaries that cross several disciplines: tenant identity, permission scope, data ownership, asynchronous work, pricing entitlement, external provider failure, and migration compatibility. These are places where local correctness can coexist with a broken overall workflow. A route handler can be correct while the surrounding retry policy creates duplicate customer work.
The model should map consequences, not just offer a preferred answer. If the design introduces a queue, what happens when the worker fails? If it separates services, who operates their interfaces? If it caches a result, how does the product determine freshness? If it creates a reusable package, who maintains compatibility?
Extra reasoning is less useful when the decision is already constrained by a reliable existing pattern and a small acceptance contract. A straightforward addition to an established module may be best handled by following the local convention and testing the behavior. Architecture effort can itself become overhead when it reopens settled decisions without new evidence.
Require genuine alternatives
Ask for at least two plausible designs and a simple baseline. The alternatives should differ in consequences rather than terminology. Three diagrams that all create the same service boundary are not three options. A useful comparison might consider a modular app with a background worker, a separate processing service, and a manual processing stage during early validation.
Each option should explain its strongest case. The simpler baseline may win because it has fewer operating obligations. The separated service may win when independent scaling, deployment, or isolation is actually required. The manual stage may win while the workflow is changing too quickly to automate responsibly.
Microsoft’s architecture guidance describes microservices as independently deployable and identifies added complexity around communication, consistency, testing, and versioning. Those tradeoffs help frame an alternative; they do not prove that any particular product needs microservices. The official microservices guide is the source for that limited comparison.
For a proposed design, include the condition that would make it lose. “Choose separate processing only if measured workloads or isolation obligations justify the interface and operations burden” is more useful than “microservices scale better.” The condition can later be tested against the product’s actual situation.
The modular-monolith chapter develops a simpler starting boundary. An architect model should be allowed to recommend that boundary even when asked about advanced orchestration. If every answer adds a new layer, the process is rewarding complexity rather than solving constraints.
Make the decision artifact small enough to review
A decision record can contain the question, current baseline, constraints, options, chosen direction, rejected alternatives, consequences, unknowns, and revalidation triggers. Its purpose is to preserve the reasoning necessary to assess the decision later. It is not a transcript of every generated thought.
Use concrete behavior examples. If the choice concerns a job queue, show one successful job, one failed job, and one repeated request. If it concerns tenant separation, show two accounts requesting the same object identifier. If it concerns a shared package, show how an older consumer behaves when the interface changes.
A diagram can clarify the flow, but it needs semantics. An arrow should indicate what crosses the boundary: request, event, data reference, or commercial entitlement. A box should identify ownership and state. A beautiful diagram without those details lets different readers imagine incompatible systems.
Specify what the implementation team receives. A reviewable interface contract might name inputs, outputs, failure categories, authorization expectations, and compatibility behavior. It should avoid unnecessary implementation constraints when several methods can meet the contract. The purpose is to align independently built pieces, not dictate every line of code.
Record uncertainty near the affected choice. If the product has not established its largest expected document size, do not hide that limitation in a generic appendix. State how the size assumption affects processing time, storage, and failure behavior. The eventual implementation can then preserve room for revision instead of treating the guessed value as permanent truth.
Use counterexamples to challenge the first design
The first coherent answer is a candidate. Ask which realistic case breaks it. A design for a subscription report tool might handle one user requesting one report. It may fail when a customer retries after a timeout, when two workers receive the same job, when an account loses entitlement during processing, or when a provider returns an incomplete result.
Counterexamples should follow the product’s actual consequences. Inventing exotic scenarios can consume attention without improving the decision. Start with known failures, edge cases in the workflow, and conditions where the selected option is weakest. The challenge is to reveal a mistaken assumption, not to prove that no system can ever be perfect.
Keep evaluation cases separate from the model’s proposed implementation. Otherwise the model can simplify the example until its architecture succeeds. A protected case might require that repeated requests produce one commercial charge while still allowing a customer to recover a missing result. The design must explain both aspects.
A useful reviewer asks whether the proposed control is enforceable. “The agent will be careful” is weaker than a tool that restricts access to the authorized account. “The worker avoids duplicates” is incomplete without a durable identity and conflict behavior. Architecture should identify where the product relies on code, storage guarantees, provider contracts, or operator judgment.
Counterexamples can also support a simpler design. If a manual review step handles the uncertain edge case safely during a small pilot, the business may not need a complex autonomous mechanism yet. The model should distinguish the early operating plan from the eventual automated design rather than pretend they must be identical.
Inspect the repository before prescribing a replacement
Architecture work in an existing product should begin with the current implementation and its constraints. A generic model recommendation can be technically reasonable and still inappropriate because it ignores established interfaces, migrations, operational experience, or customer integrations.
Give the model access to the relevant evidence where authorized, or supply a bounded repository map. Ask it to identify which claims come from inspected code and which are assumptions. A recommendation to replace a shared component should link to the affected consumers and explain why the existing contract cannot support the proposed behavior.
Preserve the baseline source revision in the artifact. A design based on last month’s branch may miss a correction already deployed. If the repository changes during the analysis, check whether the affected assumptions remain true. Architecture artifacts have freshness requirements just as test evidence does.
The output should include an incremental migration path when the product already operates. A new boundary may require old and new consumers to coexist. Explain the transition, verification conditions, and recovery options. A final-state diagram alone does not show how customers safely reach that state.
Sometimes the right result is no redesign. The model may discover that an existing module already supplies the needed capability or that a small adapter preserves the relevant contract. Count that result as useful analysis. Measuring the architect by the number of new components gives it an incentive to create unnecessary work.
Stronger reasoning does not grant stronger authority
An architectural role can propose interfaces, identify risks, and prepare a migration plan. It does not automatically authorize production writes, access to private data, changes to customer pricing, or removal of existing services. Those actions belong to the task’s established authority boundary.
Define the distinction in the work contract. An architect might have read access and permission to create a decision document. An implementer might have bounded write access in a branch. A release operator might execute a reviewed deployment under existing authorization. These roles can be held by one person or process, but the boundaries remain useful.
A model’s confidence should not change the boundary. If it concludes that a migration is safe, the conclusion is still a claim requiring evidence. The implementation must be checked, the artifact must be identified, and the relevant release conditions must be met. A persuasive explanation cannot substitute for the actual operation and verification.
This distinction is especially important when an architect delegates work. A subtask must receive no more authority than the parent has and no more access than the subtask needs. The multi-agent coding team chapter develops those role contracts. The architectural output should make the intended delegation clear enough to inspect.
An unresolved decision should remain unresolved. If the model cannot establish a required provider guarantee, it can state the missing evidence and propose a bounded check. It should not fill the gap with a plausible guarantee to keep the plan moving. Progress is useful only when it preserves the truth about the system being built.
Test whether the architect role earns its cost
A stronger model may cost more per run and create more material to review. The relevant question is whether that effort improves accepted decisions in a defined class of work. Avoid assuming that a higher model tier or a longer answer necessarily produces better architecture.
A proposed comparison can use the same constraint brief for a documented baseline process and a model-assisted process. Judge the resulting artifacts against independent criteria: correct treatment of constraints, useful alternatives, explicit failure paths, compatibility plan, unresolved assumptions, and implementation clarity. Include reviewer time and follow-up corrections.
Keep some cases protected. The team can tune prompts on familiar examples, then assess the resulting process on a different product scenario. A case with a stable existing pattern tests whether the model can refrain from needless redesign. A case with an important hidden coupling tests whether it notices the consequence rather than merely producing a diagram.
OpenAI’s evaluation guidance warns about model-judge biases such as preference for response order or verbosity. The official evaluation guide supports calibrating a model-assisted assessment against human judgment. For architecture, that means a longer document should not win unless its additional content changes the decision’s quality.
No single score should conceal a severe failure. An artifact that misses cross-account access requirements should not pass because it offers several elegant alternatives. Set independent blocking conditions and inspect disagreements. The outcome is a local finding about the selected task class, model, context, and review process.
A worked proposed decision
Imagine a small product that accepts a customer’s CSV, validates records, and produces a downloadable report. The business wants reliable retry behavior and a clear data-retention promise. It has one maintainer and a modest early audience. These are hypothetical constraints, not a description of a deployed Salars product.
The architect would first separate known requirements from guesses. The report format might be confirmed. The largest input size and acceptable waiting time might still be uncertain. The team would need to establish those before making a broad performance claim. The retention decision would identify who can approve the promise.
One option could process small files within the existing app. Another could create a background job with a bounded worker. A third could use a separate service. The comparison would ask how each handles interrupted requests, duplicate submission, large inputs, and operational recovery. It would also include the cost of maintaining interfaces.
A provisional choice might be a background job inside the product’s existing deployment structure, with explicit identity and status. The artifact would state the assumptions under which that choice holds. If measured processing demands or isolation obligations change, the team would revisit the service boundary. The model has helped expose a conditional decision; it has not established a universal architecture.
The implementation handoff would contain a contract for submit, status, and download; access expectations; failure categories; and protected acceptance cases. It would not ask an implementer to infer the retention policy from a storage adapter. That is the difference between a useful architectural artifact and a large amount of generated planning text.
What Would We Do at Salars?
We would propose using a stronger model for a bounded set of decisions involving several interacting constraints, especially data ownership, account boundaries, migration compatibility, and asynchronous work. The input would include the current repository evidence and the business’s actual operating capacity. The output would be a short decision record with alternatives and conditions.
We would compare that process with the existing documented decision method using protected cases and independent review. Success would require fewer missed consequential constraints or clearer implementation contracts at an acceptable total cost. We would not claim that an architect role had improved delivery until the comparison and follow-up evidence supported that conclusion.
We would stop a candidate analysis if it repeatedly invented requirements, ignored the existing baseline, or expanded the design beyond the authorized product question. A change in provider guarantees, customer promises, model version, or operating scale would trigger revalidation of the relevant decision.
The role earns its place by helping the builder see what a decision commits the business to. A frontier model can be an unusually capable partner in that work. The useful output is still a decision that can be questioned, implemented, and revisited by the people responsible for the product.
Sources
- Microsoft: Microservices architecture style — independent deployment and operational tradeoffs used in the architecture comparison.
- OpenAI: Evaluation best practices — human calibration and model-judge bias considerations.
The allocation principle, decision artifacts, worked example, and Salars comparison are proposed methods. No frontier model is declared universally superior, and no quantified architecture improvement is asserted.
Loading comments…