AI · Article 31 of 72 · Part 7

Test the Behavior, Not the Code

Specify contract, integration and end-to-end tests from user obligations and failure cases.

A scheduling app’s tests pass after a refactor. The customer cancels an appointment, receives a confirmation and then receives a reminder for the cancelled visit. The unit tests proved that several functions returned expected values. They never asked whether cancellation stopped the promised reminder.

Software tests should establish the behavior the product owes its users, including forbidden effects and recovery paths. Implementation details can matter when they express a real constraint, but tests that merely repeat the current code can preserve a defect just as efficiently as a correct feature. AI-generated implementation makes an independent statement of expected behavior especially valuable.

This chapter of The AI Software Factory follows a hypothetical appointment app. It concerns conventional software behavior: state, permissions, persistence and visible outcomes. AI evaluation suites handles variable model outputs; independent verification examines who defines and checks the expectation.

Translate the promise into observable outcomes

“Customers can cancel appointments” contains several obligations. The account must identify the correct appointment. The cancellation must be authorized. Its state must persist. Reminders should stop. The customer should see a clear confirmation. Staff may need an updated schedule.

Write those outcomes before reading the proposed implementation. Otherwise the tests can inherit its omissions. A developer who focuses on changing a status field may test only that field, leaving the reminder worker and schedule view outside the acceptance boundary.

The specification should include conditions. Cancellation after a stated cutoff may follow a different policy. A customer should not cancel another account’s appointment. An unavailable service should produce an honest incomplete state rather than a false confirmation. These rules need an accountable owner, not a guess from the coding agent.

Observable behavior extends beyond the page. A database record, scheduled job and external notification can be part of the promise. The test design should identify which evidence establishes each outcome and which effects must remain absent.

Use contracts to expose meaning

An interface contract describes accepted inputs, outputs and failure behavior. It should name units, states and authority, not merely data types. A field containing a string called status does not explain whether “cancelled” means requested, effective or acknowledged by another system.

For the hypothetical app, a cancellation response might identify the appointment, effective state and confirmation time. It should distinguish a completed cancellation from a request waiting for staff approval if the product supports that distinction. The customer message must agree with the underlying state.

Contract tests can cover allowed and invalid combinations without depending on every internal helper. A valid owner request should produce the intended state. An unauthorized request should be denied with no appointment change. A repeated request should preserve the same cancellation rather than schedule another side effect.

These tests remain useful when an agent reorganizes code. The app may move from one function to several modules while keeping the same obligation. If the business rule changes intentionally, the contract and tests should change together with a retained explanation.

Choose test levels for different questions

A unit test can check a narrow rule quickly, such as whether a cancellation deadline includes its exact boundary. An integration test can check how the service updates persistence and the reminder queue. An end-to-end test can check the user’s complete cancellation route.

The levels complement each other. A browser test for every arithmetic branch can be slow and difficult to diagnose. A unit test alone cannot establish that the visible interface reaches the intended service. Match the mechanism to the evidence required.

For a deadline rule, a small table of times can cover before, at and after the boundary. For a reminder, an integration fixture can inspect whether the scheduled work is cancelled. For the user route, a browser test can sign in, select a safe appointment, cancel it and verify the resulting message and state.

The phrase “test behavior” does not prohibit inspecting internal state when that state determines an obligation. It discourages tying acceptance to accidental details, such as a private function name or a CSS class with no meaning for the user. The test should fail when the promise breaks, not whenever code moves.

Keep tests independent of each other

A test that creates an appointment and another that cancels it can become order-dependent. If the first fails, the second reports a misleading failure. If the suite runs concurrently, they can interfere. Each case should create or receive the state it needs.

Playwright’s best-practice guide recommends user-visible checks and isolated test state. It also recommends user-facing locators and assertions that wait for expected page conditions. These narrow practices support reproducible UI checks; they do not establish every aspect of product correctness.

For the appointment suite, each browser case should use its own fixture or isolated account state. The cancellation case should not depend on the booking case running first. Shared setup can be reused while preserving independent records.

Time is state too. A test that relies on the real current date can cross a deadline while it runs. Control the time boundary where the test framework permits it, or construct fixtures whose meaning remains stable during the test. The expected policy should be explicit about timezone and the instant at which the decision is evaluated.

Test absence as well as presence

Many important obligations describe something that must not happen. An unauthorized user must not see another customer’s details. A cancelled appointment must not trigger a reminder. A failed write must not display a success message.

A test that checks only the expected confirmation can miss these defects. The interface may show “cancelled” while the queue still contains the reminder. Inspect the relevant side-effect boundary or use a controlled service that records attempted effects.

Negative checks need a meaningful observation window. “No reminder appeared immediately” does not establish that a scheduled reminder will never be sent. The test should advance or simulate the relevant processing condition and inspect the result. A live observation over time answers a different question and needs its own evidence record.

Keep the forbidden behavior tied to the specification. Adding arbitrary negative assertions can make tests brittle without protecting a real obligation. The useful question is what a customer or operator would reasonably rely on remaining absent.

Control external services in reproducible checks

A live calendar or messaging service can be unavailable, rate-limited or changed by other users. A test depending on those conditions may fail despite correct application logic or pass without exercising the intended failure.

Controlled responses let the suite inspect known conditions: success, denial, timeout, duplicate event and malformed data. Playwright’s guidance discusses controlling third-party responses for reproducibility. A fake service can also count writes and simulate an ambiguous timeout after accepting a request.

The controlled test proves how the app behaves under that represented condition. It does not prove that the real provider currently behaves the same way. Keep a separate integration check against authorized test resources where needed, and describe its scope honestly.

The appointment app could test reminder cancellation against a local controlled queue while performing a safe provider check for the actual messaging integration. Both matter. Calling the controlled suite “live verified” would conceal the difference and leave the operator with a false picture of readiness.

Exercise failures where state crosses a boundary

A request can validate successfully and fail while saving. A database write can succeed while a notification fails. A worker can restart after creating a side effect but before recording its receipt. These are natural places for tests because the system’s state can become inconsistent.

For cancellation, test a failed persistence write and confirm the customer does not receive a completed state. Test a successful cancellation with failed staff notification and decide what remains completed and what needs retry. The policy should distinguish the appointment state from the notification state.

Repeated events require idempotent operations. A duplicate cancellation request should not create multiple refunds, notifications or audit entries that imply separate customer decisions. The test can count attempted effects and inspect the retained operation identity.

Recovery checks should verify the state after retry or repair. A test that merely expects an exception does not establish that the app remains usable. The rollback guide develops restoration and compensation, while behavior tests confirm the promised recovery route in the chosen configuration.

Avoid tests that copy the implementation’s conclusion

An agent may write a function and then create a test that calculates the expected value using the same function or same mistaken formula. The suite can pass because both sides share the defect.

Use independently specified examples. If the cancellation cutoff is twenty-four hours before the appointment in the customer’s stated timezone, work out representative boundary cases from that rule. The expected result should not be extracted from whatever the current code returns.

Generated tests can still be valuable. An agent can propose cases, locate missing branches and build fixtures. A reviewer should inspect whether the expected outcomes match the accepted policy and whether the cases challenge the implementation. The test author’s confidence is not an oracle.

A useful review question is: what plausible wrong implementation would this test reject? If the answer is unclear, the test may be checking activity rather than behavior. A cancellation test that never inspects reminder state would accept the defective implementation described in the opening.

Inspect coverage by obligation

Code coverage reports which statements or branches ran. They do not tell whether the asserted results represent the customer’s promise. A test can execute every line while checking almost nothing meaningful.

Build a compact obligation map for consequential behavior. Cancellation authorization, persisted state, deadline handling, reminder suppression and customer messaging each need a home in the evidence. The map can link to existing tests rather than require a duplicated checklist for every file.

This makes gaps visible without imposing a universal percentage target. A low-risk display change may need a focused check. A high-consequence access rule needs allowed and denied cases. The decision should reflect failure consequences and uncertainty, not the desire to maximize a dashboard number.

Review integration boundaries especially. Several modules can each have high coverage while their combined assumptions disagree. Units, timezone, identifier scope and error states often create those gaps. End-to-end examples should travel through the actual boundary that could fail.

Accessibility is behavior too

A cancellation route can work with a mouse while failing for a keyboard user. A success message may appear visually without being available to assistive technology. Those conditions affect whether the customer can complete the promised task.

Include relevant interaction checks and manual review where the product requires it. Automated browser tests can inspect labels, focus and keyboard paths, but they cannot establish complete accessibility by themselves. A passing locator test is not a conformance assessment.

The test should follow the actual task. Can a user find the appointment, understand the consequence and activate cancellation without a pointer? Does focus remain in a useful place after the state changes? Does the interface expose an actionable error if the request fails?

Product requirements may follow an applicable accessibility standard or contract. Verify the current authority and criteria for that use. The broader time-to-value chapter considers how barriers affect activation; here the test asks whether a named user route actually remains usable.

Keep the suite maintainable

A suite that takes too long or fails unpredictably becomes a temptation to bypass. Maintainability is part of its ability to protect the product. Consolidate duplicated setup, use data-driven cases where the invariant is shared and retain clear failure messages.

Do not broaden tests merely because automation makes it easy. Repeatedly verifying every unrelated page after a tiny isolated correction can consume resources without answering a new question. Run the required checks and the ones justified by the change’s boundary; broaden when new evidence or unresolved concerns warrant it.

When a test fails, preserve a reproduction and establish whether the problem is application behavior, fixture state or test design. Repair a flaky test rather than teach the team that red output is normal. Removing an assertion to make the dashboard green needs a legitimate change in the obligation, not frustration with the failure.

The suite should have an owner and a revision path. Customer incidents can become safe regression cases. Changes to policy should update expected behavior deliberately. Tests are maintained evidence about the product, not a museum of every implementation the app once had.

Establish the baseline before assigning the change

Run the relevant existing checks on the accepted starting state and retain material failures. Otherwise a worker can mistake a pre-existing defect for a regression or dismiss a new defect as something that was always broken. The baseline need not become a lengthy ceremony for every correction; it should answer the uncertainties that affect attribution.

For the appointment change, suppose reminder suppression already fails before the new cancellation interface is added. The new work may still need to repair it because the combined promise requires it. The baseline tells the team that the defect predates the patch; it does not make the defect irrelevant to acceptance.

Conversely, if the baseline succeeds and the proposed combined state fails, the new changes are a useful starting point for investigation. That association still needs a cause. A changed fixture, environment variable or dependency can produce the same timing as a code regression. Inspect the reproduction rather than rely on chronology alone.

Keep the baseline artifact and configuration identifiable. A test report from a branch that later moved cannot establish what was checked unless the report retains its revision. When another agent modifies shared files during the run, the result may describe a state nobody can reproduce. Isolation and serial integration help preserve that evidence.

The final report should state the useful difference: which obligations gained coverage, which failures were resolved and what remains outside the check. Counting added tests is less informative. One integration case that catches the cancellation/reminder contradiction can protect more customer value than many assertions about private helper names.

What Would We Do at Salars?

A proposed Forge template would start with behavior contracts for authentication, tenant access, operation state and the product’s first customer job. It would reuse the repository’s existing test system and add cases only where the new app creates a meaningful obligation.

For Supplier Margin Guard, tests would establish that ambiguous currency is rejected, source rows remain traceable and one merchant cannot view another merchant’s report. A repeated submission would preserve the intended operation. The suite would inspect customer-facing unresolved states as well as completed results.

An independent reviewer would define at least the consequential expected outcomes before inspecting generated code. Model explanations would receive a separate evaluation suite; deterministic arithmetic and access rules would receive exact tests. This separation keeps probabilistic output from weakening invariants that the application can enforce directly.

The pilot would report which tests ran, which integrations were controlled and which live checks remained. No passing local suite would be described as production verification. The acceptance evidence would follow the obligation all the way to the state the customer relies on.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home