AI · Article 27 of 72 · Part 6

GitHub as the Engineering Control Plane

Map issues, repositories, pull requests, Actions and branch controls to one governed development record.

A developer can find the code but cannot find the decision. Why does an overdue subscription still have access for three days? A comment says “temporary.” An agent’s old summary says “approved.” The customer-facing documentation says something else. The repository contains the implementation, yet the team has no reliable record of the rule it implements.

GitHub can serve as an engineering control plane when it connects requirements, proposed changes, checks, review and release evidence around an identifiable accepted state. The platform supplies useful surfaces. The team supplies the operating agreements that make those surfaces authoritative. A hundred issues and a green dashboard can still conceal a missing decision.

This chapter of The AI Software Factory concentrates on engineering coordination. GitHub issues as market research examines public complaints as demand signals; this chapter examines a team’s own issues as implementation records. The proposed subscription feature below is hypothetical, and no GitHub configuration is represented as already deployed at Salars.

Give the control plane a precise job

A control plane coordinates intended state and the operations that change it. In a software factory, that means making it possible to answer: what should be built, who owns the change, what evidence supports acceptance, what can merge and what was released?

GitHub does not need to hold every piece of customer data or every operational log. It can link to the relevant evidence without becoming a second production database. Its engineering record should retain enough context for contributors to understand the obligation and for reviewers to inspect the change.

The authoritative specification may be a repository document, an issue or another established system linked from the issue. Choose one home for the rule. Copying the same requirement into several task messages creates competing versions. An agent should be able to locate the accepted wording and distinguish it from discussion about future options.

The control plane should also identify its limits. A merged pull request does not prove that a release happened. A successful deployment does not prove that a customer workflow succeeded. Those states need their own evidence and should remain visible when a task reports completion.

Issues should describe the customer’s obligation

An implementation issue should begin with the observed problem and intended behavior. “Improve subscription logic” gives little guidance. “A customer who requests cancellation should retain already-paid access until the end of the period, with the effective date shown on the account page” provides a rule that can be tested.

Include safe reproduction details, constraints and acceptance cases. For the hypothetical cancellation change, identify the relevant subscription states, timing boundaries and expected customer message. Say what the feature must do when billing data is unavailable. Do not leave that decision for an agent to invent during error handling.

A useful issue also names exclusions. This change might not alter refund policy, plan prices or grace-period rules. Exclusions protect the current release from apparently helpful additions that create new obligations. A worker can propose an adjacent improvement without silently adding it to the accepted task.

Customer records should remain in authorized systems. Reproduction fixtures can remove names, account identifiers and unnecessary content while preserving the behavior. Public issues require particular care: a convenient debugging attachment can expose data far beyond the people investigating the problem.

Make task state describe evidence

A label such as “done” can mean code written, reviewed, merged, staged or verified live. A factory should use states that match its actual release process and avoid overloading one word with all of them.

For a small team, the states might be proposed, specified, in progress, ready for review, accepted and verified. Their exact names matter less than their transition rules. Ready for review requires an inspectable artifact and supporting checks. Verified requires evidence at the customer boundary named by the task.

State should not advance because an agent says it finished. The agent reports its work; the coordinator assesses whether that work meets the state requirements. Automated checks can enforce mechanical conditions, while unresolved policy decisions remain visible for an accountable owner.

Keep blocked work specific. “Blocked” should say which condition prevents progress and what evidence would resolve it. A failing baseline test, missing authorization and unclear requirement are different problems. Combining them into a general waiting state makes prioritization harder and encourages unnecessary retries.

Pull requests connect the artifact and its explanation

A pull request should explain the final behavior to a reviewer who did not participate in the chat. State the trigger, the change and the relevant validation. A small correction may need only a concise description. A migration or billing change requires enough context to assess compatibility and recovery.

GitHub’s pull request reference describes surfaces for conversation, commits, checks and changed files. Together they can connect explanation and evidence to a proposed merge. The team’s discipline determines whether those surfaces contain the information needed for a decision.

Link the pull request to the specification. If scope changes, revise the description around the final artifact. A long conversational history can obscure what is actually proposed. Abandoned approaches belong only when they explain a tradeoff the reviewer needs to understand.

An agent’s report should be attached to the exact change it describes. If a later commit alters behavior, check whether earlier evidence remains applicable. A review of one diff should not be treated as permission for any future diff that happens to occupy the same pull request.

Protect the accepted branch deliberately

GitHub branch protection can impose review and status-check requirements. Availability depends on repository visibility and account plan, and default rules can leave administrator bypasses. The current protected-branch documentation explains these settings and qualifications. Check the actual repository rather than assume the presence of a feature means it is enabled.

The engineering policy should name the intended accepted branch and the requirements for changes entering it. A green check is useful only if it corresponds to a meaningful test and is required at the right boundary. A check named “validation” that always succeeds gives the appearance of control without constraining anything.

Review authority should match the consequence. A UI wording correction differs from a data migration or credential change. The factory can use existing ownership conventions to route review while preserving a clear final release responsibility. Adding more approvers does not automatically improve judgment; the reviewers need a defined question and relevant evidence.

Document permitted exceptions through the existing process. Emergency recovery may require a different path, but that path should preserve accountability and a subsequent evidence record. Routine inconvenience is a poor reason to bypass a safeguard the team relies on.

Checks must test the combined state

A worker’s unit tests establish something about its working copy. The accepted branch includes other changes. Integration checks should evaluate the proposed change with the state it will join, especially when shared contracts or dependencies moved during development.

For the cancellation example, parser tests alone are insufficient. The combined behavior includes the billing event, entitlement update and account-page explanation. A test should establish that a cancellation request does not immediately remove paid access and that the effective date is displayed correctly. Recovery cases should cover a delayed or repeated event when those are realistic conditions.

Keep the checks proportional. A cosmetic wording change does not justify building an elaborate new test framework. A consequential billing rule does justify cases that distinguish the allowed and forbidden states. Behavior-based testing explains why the promise should determine the tests.

Checks also need maintainers. A permanently flaky test trains reviewers to dismiss failures. A check that depends on a live third-party service can fail for reasons unrelated to the patch. Distinguish controlled integration checks from live operational verification, and use each for the question it can actually answer.

Agent identities should be narrow and temporary

A coding agent may need to read source, propose a branch and inspect check results. It may not need repository administration, billing settings or production secrets. Its identity should reflect the task rather than the full authority of the owner who launched it.

Separate code preparation from merge and deployment where the existing release policy requires it. An agent able to edit a workflow file could propose a dangerous change to its own operating controls. Review that file under the appropriate ownership and security boundary rather than assume ordinary code review covers every consequence.

The same concern applies to incoming content. An issue, dependency README or failing-test log can contain instructions that attempt to redirect the agent. Those materials describe the task environment; they do not grant authority. Agent security treats this untrusted-input boundary in depth.

When a task ends, narrow or revoke its access without losing the artifact. The branch and report should remain inspectable even if the worker no longer has credentials. This makes task continuity less dependent on keeping broad permissions alive indefinitely.

Keep release state separate from merge state

A merge can trigger deployment, but the trigger does not guarantee deployment success. A release can fail after build, reach only some services or run with an unexpected configuration. The control plane should preserve the deployment result and verification state rather than close the issue at merge automatically.

Record the source revision and artifact identity where the tooling supports it. Link the release to the checks and approval applicable to that artifact. If a preview demonstrated a different revision, the discrepancy should be resolved before treating its evidence as production evidence.

For the hypothetical subscription feature, live verification might inspect a safe test account and confirm the effective date after an authorized cancellation. It should avoid changing a real customer’s plan merely to prove that the button works. The verification method should be designed before release so missing test access does not become a last-minute excuse for a broad action.

Cloud-native development follows this complete path across environments. GitHub’s role is to connect the engineering record; the runtime and customer system still supply their own observations.

Let failures create useful records

A failed check should preserve the reproduction, relevant output and the artifact tested. A bare red icon forces the next person to repeat the investigation. An enormous unfiltered log can conceal the failure and expose unnecessary secrets or customer data.

The issue should retain the smallest useful evidence. Which case failed? What was expected? What happened? Was the condition present before the change? A worker can provide this without claiming a cause it has not established. Association between a new patch and a failure is a lead, not proof that the patch caused it.

A release failure also needs a decision. Can the operator retry safely, restore a previous artifact or pause new work? An automated retry is appropriate only when it cannot create duplicate external effects or repeat an invalid change. The idempotency chapter explains why retry policy belongs beside operation identity.

After recovery, retain the accepted explanation and the new regression case where useful. The record should improve the next change. Repeating the same failure with more confident summaries is a sign that the control plane is recording activity rather than learning.

Avoid a second engineering bureaucracy

A factory can become burdened by its own tracking system. Every small patch gets an issue, a design document, a checklist, a dashboard entry and a separate agent transcript. The operator spends more time reconciling records than deciding whether the change is useful.

Reuse the existing conventions. If the repository already has an accepted issue template, release notes and deployment record, extend those surfaces only where a genuine gap exists. The software factory should make current responsibilities clearer rather than create a parallel queue that disagrees with the actual engineering process.

The amount of recordkeeping should scale with risk and uncertainty. A minor display correction may need a concise pull request and a focused check. A migration affecting customer entitlements needs the accepted rule, compatibility evidence, recovery plan and live verification. Proportionality keeps the record useful.

Automated summaries can help locate information, but they should link to the authoritative artifact. A summary saying “all checks passed” needs the check results it summarized. If source evidence changes, the summary becomes stale. The control plane should make that relationship visible rather than let polished prose replace the underlying record.

Test the record with an unfamiliar reviewer

An engineering record can look complete to the person who assembled it because that person remembers the missing context. Ask someone with appropriate access who did not work on the change to answer five questions from the retained artifacts: what behavior was promised, what changed, what evidence supports acceptance, what was deployed and what remains uncertain?

The exercise should not depend on guessing the team’s preferred answer. Give the reviewer the relevant artifact locations and allow them to identify contradictions. If a specification says the cancellation takes effect at period end while the pull request says access ends immediately, that disagreement is a material defect even if the code compiles.

A hypothetical pilot might sample one minor correction, one integration change and one consequential write. The cases would test whether the record scales with the decision. The small correction should not require an hour of reading; the consequential write should expose the exact approval and resulting state.

Retain the first point where the reviewer could not proceed. It may be a missing artifact identity, an unexplained check name or a policy decision hidden in a comment thread. Repair that gap directly. Adding another general checklist without locating the failure can make the record longer while leaving the same uncertainty intact.

This is an evaluation of the operating process rather than a claim about GitHub’s interface. The platform can display all the relevant information and the team can still fail to supply it. A successful exercise supports the record convention for those tested changes; it does not certify every future release.

Review the convention again when the product introduces a different consequence, such as customer billing, private data or multi-service deployment. New obligations can require evidence that the earlier examples never needed. The control plane should evolve with the actual decisions it coordinates.

What Would We Do at Salars?

A proposed Forge pilot would keep one issue for each accepted customer obligation and one inspectable change path. It would use the existing repository’s branch, check and release conventions. Agents would receive bounded tasks linked to the specification, exact ownership and the required handoff evidence.

For a proposed Supplier Margin Guard import correction, the issue would retain a sanitized missing-currency fixture. The pull request would explain the rejection behavior and show tests for missing, valid and conflicting currency labels. A preview would expose the intended merchant message. The release record would distinguish implementation, acceptance, deployment and live verification.

The pilot would inspect whether an unfamiliar authorized reviewer could reconstruct the decision from those records. If the reviewer needed the original chat to determine what was approved, the task record would be incomplete. That is a concrete failure the team could correct before expanding its agent count.

The control plane earns its role when the next contributor can see the obligation, the proposed artifact and the evidence connecting them. GitHub supplies a practical place for that record; responsible engineering gives the record its meaning.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home