AI · Article 33 of 72 · Part 7

Never Let the Same AI Grade Its Own Work

Separate authored claims, deterministic checks, blinded evaluations and accountable human review.

The agent writes a tax calculation, writes tests using the same formula and reports that every test passes. Its arithmetic is consistent. Its assumption about the taxable amount is wrong. Asking the agent whether it is confident can produce a longer explanation of the same mistake.

An AI’s self-review can help improve a draft, but acceptance needs evidence that does not depend solely on the author’s assumptions. Independently specified outcomes, exact checks and accountable review reduce the chance that one error supplies both the work and its approval. A second model is useful only to the extent that its evidence and task create meaningful independence.

This chapter of The AI Software Factory examines that independence. The illustrative calculation is hypothetical and does not state a tax rule. Behavior tests explains test design; AI evaluation suites explains variable-output measurement. Here the question is who or what can legitimately establish that the author’s artifact deserves acceptance.

Self-review has a limited but useful role

An author can find omissions, inconsistent wording and obvious defects by rereading its work. A coding agent can run tools, inspect errors and repair a broken build. These steps often improve the artifact before another reviewer sees it.

The limitation is that the author may preserve the assumption that caused the defect. If it believes the wrong input field represents the price, it can write code, comments, tests and a summary that all agree. Internal consistency then becomes part of the problem.

The title’s rule concerns final grading authority. It does not require forbidding an agent from checking its own work. The useful operating design is self-check first, then acceptance evidence with an independent basis. An author can prepare the evidence while a responsible reviewer decides whether it answers the actual question.

That distinction keeps review proportional. A trivial formatting correction may need a focused deterministic check. A consequential calculation needs independently established expected results. Independence should match the failure mode rather than become a ritual involving extra agents for every small edit.

Independence begins before the artifact exists

The expected behavior should come from the customer’s obligation, authoritative rule or accepted specification. If the reviewer receives only the implementation and asks whether it looks correct, the implementation has already framed the answer.

For a hypothetical invoice calculation, the requirement might define which amounts are included, how currency is represented and which rounding rule applies. Representative expected results can be worked independently from that rule. The tests then compare the code with the specification rather than compare the code with itself.

When the rule is unclear, preserve that uncertainty. A reviewer should not invent an answer merely to finish the task. The responsible domain owner may need to decide the intended behavior, or the feature may need to exclude the unresolved case.

This is a practical reason to write acceptance cases at specification time. They expose ambiguous terms before code makes them expensive to change. “Correct total” is not a case until the included amounts, units and rounding boundary are understood.

Different models can share the same error

Two agents may use different model names and still rely on the same source, prompt framing or mistaken assumption. If both receive an incomplete policy, both can produce a confident answer that omits the same condition.

Independence has several dimensions: who specified the expected behavior, what evidence the reviewer receives, whether it sees the author’s conclusion, how it evaluates the artifact and whether it can challenge the governing assumption. A different provider can be one dimension, but it is not a general guarantee.

For an explanation feature, a reviewer given the source evidence and a claim-specific question can be more useful than a reviewer asked to “approve this excellent answer.” Blind candidate identity where practical. Avoid loading the review prompt with the author’s confidence or a summary of why the artifact should pass.

OpenAI’s evaluation guidance discusses model-grader position and verbosity biases and human calibration. Those are reminders that a judge is an instrument with behavior to evaluate, not an automatic authority.

Match the verifier to the claim

Exact arithmetic should be checked with exact arithmetic and known inputs. A route should be checked at the route boundary. A schema should be parsed. A user-facing interaction should be exercised. A factual claim should be compared with the supporting passage.

A model can help interpret a failure or identify missing cases, but it should not replace a precise check with a reassuring paragraph. If a deterministic tool can establish whether all seventy-two routes exist, another model’s opinion about the route list adds little evidence.

Judgment remains necessary for some questions. Does an explanation preserve the customer’s meaning? Does a product promise exceed its evidence? Is a design tradeoff appropriate? Those questions need a rubric, relevant expertise and inspectable reasons rather than an unsupported numerical score.

The verification plan can therefore be layered. Mechanical invariants receive executable checks. Source-supported claims receive passage review. Ambiguous quality receives calibrated judgment. Consequential decisions receive accountable approval at the existing authority boundary.

Keep the reviewer away from the author’s answer key

A reviewer should receive enough context to assess the work, but not an answer key created solely by its author. For a calculation, give the governing rule and input fixtures before showing the implementation’s result. For a source claim, give the passage and question rather than only the author’s summary.

This does not mean withholding all context. A reviewer needs the intended behavior, constraints and relevant environment. The goal is to prevent the author’s conclusion from becoming the easiest path through the review.

For a hypothetical merchant report, the reviewer can inspect source rows and the accepted definition of contribution. It can then ask whether the reported amount includes every named cost and whether unresolved values remain visible. If it sees only a polished report saying “verified profit,” it may approve the phrasing without noticing missing evidence.

Record disagreement precisely. Which claim or behavior failed? What input shows it? Which source or rule establishes the expected result? These details let the author repair a specific defect and let the release owner assess whether the repair addresses the actual problem.

Reviewer access should be broad enough to inspect and narrow enough to govern

A reviewer may need read access to source, tests and safe fixtures. It rarely needs authority to deploy, change billing or modify customer records merely to inspect a patch. Broad operational credentials can turn a review task into a new risk.

Separate review and repair where useful. A reviewer can return findings and suggested changes; an assigned implementer can edit the owned files. If the reviewer also edits, identify that transition and re-establish which artifact needs review afterward.

For concurrent work, use the ownership and isolation rules described in parallel agent development. A reviewer should not quietly rewrite the same file while the author continues editing it. Their collaboration can otherwise produce a patch that neither reviewed as a complete state.

The final approval should attach to the artifact, not merely the task label. A changed file after review may invalidate the evidence. The existing release process determines which changes require renewed checks or approval, and the team should retain that relationship explicitly.

Ask reviewers for concrete failures

“Review thoroughly” can produce general advice and an agreeable summary. A focused review asks for the first point where the promise fails, an unsupported claim, a missing boundary case, a contradictory state or a reproduction of incorrect behavior.

Give the reviewer the central question and acceptance conditions. For an endpoint, ask whether allowed and denied requests produce the intended state. For an article, ask which consequential claim lacks support and where the explanation stops delivering its title’s promise.

Findings should include evidence and a fix direction. A claim that “security could be improved” is less useful than showing that one customer identifier can retrieve another customer’s report. The latter identifies a concrete boundary and a case that can become a regression test.

Review should also acknowledge its limits. A source-only inspection cannot establish live behavior. A test against controlled fixtures cannot establish provider availability. A human reviewer unfamiliar with the domain cannot certify a specialized rule simply by reading fluent prose.

Make severity determine the next action

A typo, a weak explanation and an unauthorized data disclosure have different consequences. The review record should distinguish them so the release owner can act proportionately.

A critical failure in a customer obligation should block acceptance within that scope until repaired or the feature narrowed. A minor wording preference can be resolved by editorial judgment without another full test cycle. Unlimited iteration on valid stylistic alternatives can consume effort without improving the product.

For an AI feature, avoid an aggregate score that compensates for a critical error with good style. A clear answer that invents a refund policy still fails policy fidelity. The evaluation suite chapter shows how to report severity and case families separately.

The stopping rule should be explicit. Stop when material weaknesses are resolved and further revisions trade valid preferences, or when a real unresolved condition requires a scope decision. Do not stop merely because the author has become more confident or the review budget is nearly exhausted.

Verify the fix rather than the explanation of the fix

An author can respond to a finding with a plausible narrative while leaving the defect intact. The reviewer should inspect the changed artifact and rerun the reproducing case or other appropriate evidence.

If an unauthorized lookup was possible, confirm that the denied path creates no disclosure and that the allowed path still works. If a calculation omitted a cost, compare the repaired result with the independently established example. If an article overstated a source, inspect the revised claim and supporting passage.

A fix can create a new defect. Narrowing access too broadly can deny legitimate users. Rejecting every ambiguous input can make the app unusable when a safe clarification route exists. Include the relevant countercase so the repair preserves the intended capability.

Retain the finding and verification result together. The record should identify the final artifact and distinguish the original failure from the accepted repair. A comment marked resolved is a workflow state; it should correspond to actual evidence that the material issue was addressed.

Build a check that can fail meaningfully

A verifier that always agrees is not useful. Test the reviewer or check with known defects and known acceptable cases. It should reject the former without inventing problems in the latter.

For a model grader, a calibration set can include an unsupported policy exception, a correct concise answer and a verbose answer with the same defect. The reviewer should identify the consequential issue regardless of presentation. If it rewards verbosity, that bias can undermine acceptance.

For a deterministic validator, include a broken route, duplicate identifier or mismatched navigation pair in controlled fixtures. A passing validation run on the real artifact means more when the validator has demonstrated that it catches the relevant defect class.

These are local checks of the verification mechanism, not proof that it detects every error. Preserve scope and revalidation triggers. A validator for route existence says nothing about article truth; a claim reviewer says nothing about production deployment.

Preserve the distinction between inspection and experiment

A reviewer can inspect a design and identify a plausible failure. That finding is valuable, but it is different from executing a reproducing case. The report should say which happened. A proposed experiment should not become “tested” merely because its steps were described clearly.

For a hypothetical tenant-access defect, source inspection may show that the endpoint retrieves a report without checking ownership. An executed safe request can then establish whether the configured application actually discloses another fixture account’s report. The code finding and the observed result have different evidence and may reveal different conditions.

When execution is unavailable, retain the plausible finding and its limitations. The release owner can decide whether the source evidence is sufficient to block the change or whether a targeted test is needed. Inventing an executed result would make the record less useful and could direct the repair toward a condition that never occurred.

The same distinction applies to performance and economics. A reviewer may calculate that retries could multiply cost under stated assumptions. That is a scenario analysis. Measuring the actual attempt rate and bill under a real workload is another task. Both can inform a decision when the report labels them correctly.

A compact evidence record can include the artifact inspected, the claim, supporting location, test performed if any, observed result and unresolved condition. The format should reuse the team’s existing review convention. Its purpose is to keep proposed checks, executed checks and broader inference from collapsing into one confident summary.

That discipline also protects the reviewer. It can provide a clear, useful judgment within its actual access and competence instead of pretending to certify more than it examined. Independent verification gains credibility through accurate scope.

Independence has a cost worth measuring

Review consumes model calls, tools and human attention. A factory should measure accepted-change cost, including revision and integration, rather than only authoring speed. Some independent checks save effort by exposing defects early; others can become duplicated ceremony.

Choose the narrowest evidence that addresses the consequential uncertainty. A schema parse does not need an elaborate model debate. A complex business-policy question may need the domain owner rather than several generic agents. The appropriate verifier is the one that can inspect the actual claim.

A proposed pilot can compare defects found, false findings, repair time and escaped failures under a baseline and a new review path. Keep the tasks comparable and retain protected cases. One impressive review does not establish a universal productivity gain.

The economic result may support stronger review for high-consequence changes and lighter checks for reversible corrections. That is useful differentiation. The factory should avoid imposing the same heavyweight ritual on every artifact merely because agent capacity is available.

What Would We Do at Salars?

A proposed Forge team would distinguish author, verifier and release owner. The author could self-check and repair its patch. The verifier would receive the accepted specification and independent cases. The release owner would inspect the combined evidence under the repository’s existing workflow.

For Supplier Margin Guard, exact arithmetic and tenant access would use deterministic checks. Explanation claims would be compared with source rows and labeled assumptions. A model reviewer could flag unsupported language, but it would not certify actual merchant profit or authorize live price changes.

The pilot would include deliberately wrong calculations and unsupported explanations in safe fixtures to test the verification path. It would retain failures and disagreements rather than count approvals as evidence of quality. All of these checks are proposed here, not reported completed Salars results.

An accepted artifact should have support outside the story its author tells about it. That support can be compact, but it must reach the question the customer and operator actually need answered.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home