AI · Article 4 of 4

How to Test a Self-Checking System Without Believing Its Sales Pitch

A sorting example exposes a weak checker, then shows how to define a contract, reserve failure cases and keep the result within its actual scope.

A program receives the numbers 2, 2 and 1. It returns 1 and 2, then announces that its self-check passed. The output is certainly in order. One of the twos has disappeared.

This is an original teaching example, not a report from an operating AI service. Its smallness is useful. We can see precisely what the checker inspected, what it missed, and why the word “checked” did not settle the question.

A system becomes easier to trust when the checks have an explicit contract and face cases chosen to break it. That approach can improve an AI workflow without requiring a theory of consciousness or a universe that validates itself.

Define the answer before judging the checker

For a finite list of integers, an ascending sort has two requirements. Adjacent values must be in nondecreasing order, and the output must contain exactly the same values with the same multiplicities as the input. Multiplicity means how many times each value appears. Here, two copies of 2 must remain two copies.

The first requirement alone accepts an empty output for every input: there are no adjacent values out of order. It also accepts a beautifully ordered list of invented numbers. A system can meet that weak specification every time and never perform the intended task.

Adding the second requirement gives us a precise contract for this domain. It does not cover sorting customer records by several fields, preserving the original order of equal-key records, interpreting missing values or choosing locale-specific text order. Those tasks need further rules. Defining a small domain is how we make a result assessable.

The contract also separates proposal from acceptance. Any method may propose an output, but the checker should receive the original input through a path the proposer cannot silently replace. Otherwise the system could delete the second 2 from both records and congratulate itself on preserving everything it was shown.

Compare against a baseline

A comparison needs a starting point. For this demonstration, the baseline checks only whether the proposed output is ordered. The stronger checker inspects order and compares the input and output value counts. Two development examples established the intended behavior before the remaining cases were evaluated: an ordinary valid sort, and a sorted output with a value deleted.

Six separate evaluation cases were fixed before implementing the comparison. Their expected answers follow from the contract, not from the checker’s own verdict. The deterministic Python demonstration was then run for this series. These are the results:

Case Ordered-output check Order and value-count check Required verdict
Delete a repeated value Accept Reject Reject
Add a duplicate Accept Reject Reject
Replace one value with another Accept Reject Reject
Leave the output unsorted Reject Reject Reject
Empty input and output Accept Accept Accept
Correct sort with negatives and duplicates Accept Accept Accept

The baseline accepted three incorrect outputs. The stronger checker gave the required verdict on all six cases. This is a bounded demonstration of a specification gap, not a measured reliability rate for AI. There was no language-model trial, no physical experiment and no claim about how often failures occur in deployed systems.

The rule behind the stronger checker supports the conclusion for finite integer lists when implemented correctly: an ordered output that preserves the entire multiset satisfies the stated sorting contract. The six cases illustrate that rule and catch the deliberately weak baseline. They do not constitute a formal proof that the checker implementation has no bugs.

Protect the examples from the development process

A builder can accidentally teach a system to pass the test instead of doing the job. Once a failure case has guided a change, it becomes development material. Keep it as a regression case, then reserve new examples for the next evaluation.

This matters especially when a language model writes the proposed answer, the checking instructions and the explanation of the result. Agreement between those outputs may reflect a shared mistake. A separate expected answer, a calculation from independently specified inputs or a reviewer’s decision supplies a different basis for comparison.

For a first operational trial, use a small record containing the input, expected outcome, actual output, acceptance decision and failure reason. Record the model and tool versions when relevant. Keep missing information and service failures among the examples; they test whether a system can correctly decline to proceed.

The AI workflow testing guide provides concrete order-support cases. The answer-verification guide helps inspect evidence behind an individual response. Neither replaces a task-specific contract. A refund workflow, for example, must also establish authorization and policy conditions; arithmetic correctness alone does not authorize a refund.

Test the boundary, not just the happy path

The sorting demonstration has deliberately modest scope. Its checker assumes the two lists are the authentic input and proposed output, and that their elements are integers. A practical system must validate those conditions before relying on the comparison. Malformed data should produce an explicit failure, rather than an unexplained crash or an automatic acceptance.

Other boundaries require different tests. Can a stale record arrive with a current timestamp? Can the proposer change the checker’s configuration? Does a timeout get treated as success? Does a retry repeat an external action? Does a rejected answer still reach a user through another path?

These questions turn “self-checking” into operating behavior. A rejected result must stay rejected. An unavailable checker needs a stated fallback. When consequences warrant human review, the reviewer needs the evidence and the reason for the handoff. A green badge that does not control delivery is a decoration.

Do not aggregate failures in a way that hides the one you cannot accept. A system might be fluent on most examples and disclose protected data once. Whether it can proceed depends on the actual requirements, not the average pleasantness of its replies. Set the consequential rejection conditions before looking at the results.

Software success cannot validate a theory of mind

It is possible to build this kind of checking entirely through ordinary computation. That fact makes it useful while limiting what it says about Penrose’s larger questions. A successful sorting checker does not establish awareness. It also does not refute a proposal that human understanding has noncomputable ingredients. It performs a different task.

The distinction runs the other way as well. A new result in quantum foundations would need an engineering connection before it justified a claim about better AI verification. A physical theory’s vocabulary does not improve a checker unless it changes an operation whose result can be compared.

If someone proposes “self-checking reality” as a scientific hypothesis, ask which observable prediction it adds to specified alternatives, under what conditions, and what result would count against it. If the phrase refers to software, ask for the contract, artifact, checker and failure path. The two meanings need different evidence.

Keep the missing two in view

A useful next step for a team is to choose one narrow task, state its acceptance conditions, fix an ordinary baseline and reserve cases with known outcomes. Run the comparison, retain failures and stop at the conclusion the evidence supports. Extend the domain only when the new requirements have their own checks.

The small sorting failure is worth remembering because it is so easy to inspect. The list looked orderly. The check really did pass. The second 2 was still missing. Trust begins with asking what the check was designed to notice.

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home