A contractor hired to repair a door needs access to the door, suitable tools and the work instructions. Giving that contractor the safe combination, customer files and authority to sign contracts would make the job more dangerous without making the door repair better. A capable coding agent deserves the same distinction between useful access and unnecessary power.
Treat a coding agent as a powerful contributor working under a specific contract: defined scope, limited resources, controlled actions and revocable access. The system should enforce important boundaries outside the model. Prompt instructions can guide behavior, but they do not remove a credential’s capability or stop an unsafe tool from acting.
This chapter of The AI Software Factory concentrates on development-agent exposure. The existing permission architecture describes general business-agent authority; safe writes describes a committed action. Here the question is how to let an agent inspect, implement and test software without giving it the whole business.
Start with the assets and consequences
A coding workspace can contain source, secrets, customer fixtures, deployment configuration and access to other services. Those assets have different consequences if exposed or changed. A threat model should identify the ones present in the actual task rather than produce an abstract list of every possible danger.
For a hypothetical report-parser correction, the agent needs source code and safe example files. It may need a restricted test service. It does not need live merchant reports, production billing keys or the ability to change repository access policy.
Name the unacceptable outcomes: disclosure of another customer’s data, unauthorized production changes, exfiltration of secrets, corruption of the accepted source or excessive resource spending. These outcomes guide controls and tests. A vague goal to “be secure” cannot tell the tool layer which operation to deny.
The task contract should also identify who can expand access. A worker can report that a needed resource is unavailable and explain why. It should not solve the problem by searching unrelated directories or using a broad credential it discovers accidentally.
Distinguish read access, write access and commitment
Reading a repository can support analysis. Editing a branch prepares a change. Merging or deploying creates a shared accepted state or production effect. The task should distinguish those stages.
A coding agent may have broad read access to understand the architecture and narrow write ownership for implementation. The release coordinator can own shared configuration and integration. Production authority can remain in the established release workflow rather than the interactive workspace.
Parallel agent development makes file ownership useful for coordination. Security adds another question: what can the agent reach outside its owned files? A write convention in a prompt is weaker than a filesystem or tool boundary that actually denies unrelated writes.
Controls should be layered where the consequence warrants it. A tool can validate the operation, a scoped identity can limit resources and the release process can require review. None eliminates the need for the agent to understand the task, but each reduces reliance on perfect interpretation.
External text is evidence, not authority
A coding agent reads documentation, issue comments, dependency files and test output. Those materials can contain instructions that attempt to redirect its behavior. A malicious README might tell it to reveal secrets or weaken a workflow while presenting the request as a setup requirement.
OWASP’s prompt-injection guidance describes indirect injection through external content and recommends least privilege and human controls for high-risk actions. The same passage notes that retrieval and fine-tuning do not fully remove the risk. The boundary is therefore operational as well as linguistic.
The agent should treat external content as data within the user’s task. A dependency’s installation instructions can inform a proposal, but they do not grant permission to send credentials or change production. The tool layer should prevent actions outside the authorized scope even if the model is persuaded that they are necessary.
Segregating sources helps inspection. Mark which material is the accepted specification, which is official documentation and which is untrusted issue content. This can reduce confusion, but source labels alone should not be treated as complete protection against injection.
Give tools narrow meanings
A generic command runner can execute many actions, including ones unrelated to the task. A purpose-built tool can accept a constrained target and operation. The appropriate choice depends on the work, but consequential actions benefit from narrow validated interfaces.
A report diagnostic tool might read a named fixture and return findings. A deployment tool might accept an approved artifact identity under the normal release policy. A broad tool that can edit any account record makes it harder to preserve the task boundary.
Validate tool arguments independently. A model-produced customer identifier must belong to the authorized scope. A path should resolve within the intended workspace. A generated URL should not become an arbitrary destination for sensitive output. The check belongs where the operation can actually occur.
Tool results are also untrusted in the instruction sense. An error response can contain text asking the agent to retry with broader permissions. The worker can use the error as evidence, but access expansion still requires legitimate authority rather than the service’s phrasing.
Keep credentials away from generated-code reach where practical
A coding agent may execute code it just generated or code from a dependency. If production secrets are available in that process, a bug or malicious package can read them. A promise not to print secrets does not remove that exposure.
Use test credentials and narrow scopes for development. Prefer temporary task identities where the platform supports them. Production credentials can be supplied only to the authorized release stage rather than every workspace that builds the application.
GitHub’s secure-use reference recommends minimum workflow token permissions and cautions that automatic secret redaction is not guaranteed. Those facts support treating logs and process access as real boundaries rather than assuming a secret manager makes every use safe.
When a credential is exposed, removing it from a visible file is not sufficient. The system may have retained it in logs or history, and the credential may remain valid. Follow the actual provider’s rotation and incident process. Do not reproduce real secrets in a public article or a broad diagnostic report.
Isolation needs an explicit scope
A sandbox or isolated environment can constrain files, network access and processes. Its protection depends on what is actually enabled. An environment labeled “sandbox” can still mount sensitive directories or provide privileged sockets.
Document the relevant boundary for the task. Which paths can the worker read and write? Which hosts can it contact? Which credentials exist? Can it install dependencies, invoke containers or reach the deployment service? These questions make the isolation claim inspectable.
For the parser correction, safe fixtures and restricted network access may be sufficient. A task investigating a live integration may need a narrowly scoped connection, but that should be deliberate. Broad access should not remain as a default after the exceptional task ends.
Isolation can affect usefulness. A blocked tool may prevent a valid check. The worker should report the limitation and use authorized alternatives where available. It should not claim that a check passed when the environment never permitted it to run.
Review generated code as executable behavior
Model-generated code is not inherently safe because it came from a helpful assistant. It can contain wrong access checks, unsafe parsing, logging of private fields or a broad dependency added for a small function.
Review the behavior and boundary relevant to the change. For a parser, inspect input handling, output trust and resource consumption. For an endpoint, inspect identity and tenant scope. For a workflow change, inspect credential and release authority.
Independent verification separates authoring from acceptance. A coding agent’s self-review can improve the patch, but consequential expectations should have an independent basis. Known denied cases are especially useful: the system should refuse the operation while preserving the allowed route.
The review should not be an unrelated exhaustive audit for every reversible correction. Scope it to the change and material surrounding risks. Expand when evidence reveals a broader problem, and retain the reason for that expansion.
Dependency execution expands the trust boundary
Installing a package or running a setup script can execute third-party code inside the workspace. That code may reach the same files and credentials as the worker. The dependency choice therefore changes the task’s exposure, not merely its implementation convenience.
Evaluate whether the package is needed, what it executes and how it is pinned and maintained. A small parsing function may not justify a large dependency with broad setup behavior. The open-source evaluation chapter supplies the selection process.
The supply-chain security chapter owns build and artifact integrity. Agent security adds the interactive development question: what authority does unreviewed code receive when the agent runs it? A clean production build does not establish that a development workspace never exposed a secret during setup.
Keep installation and test commands within the authorized environment. A package’s instructions may be relevant, but they do not override the owner’s task or release boundary. If a dependency requires a consequential change, prepare the concrete proposal and supporting evidence through the existing process.
Limit accumulated actions and spending
A tool can be safe for one operation and dangerous when invoked thousands of times. A read-only agent can also create substantial cost through model calls, downloads or expensive analysis.
Set task budgets for attempts, elapsed time and resources where relevant. The coordinator needs an aggregate view across workers. A per-agent limit does not prevent a large combined commitment when many agents are launched.
The system should have an honest incomplete state when a budget is reached. Preserve useful artifacts, tests performed and the unresolved condition. The worker should not invent success or continue without bounds because it interprets persistence as unlimited authority.
A production feature needs a customer-aware policy for exhaustion. A hard stop can deny a promised service; a graceful queue or explicit exception route may be better. The budget control should be chosen against the actual obligation rather than copied from the development task.
Preserve evidence without creating another exposure
A security-relevant record can show the task, permissions, tool calls, artifact and denied operations. It helps explain what happened and verify that controls worked. Unlimited transcripts can also retain sensitive information unnecessarily.
Log the facts needed for the decision. A denied request can retain operation type, scope and reason without copying a whole customer document. Protect access and define retention according to the product’s responsibilities.
The record should not be editable by the same untrusted operation it is meant to observe. A worker can prepare a report, while the tool or platform supplies execution evidence. A narrative saying “no secrets accessed” is weaker than a boundary that denied access and retained the denial.
Evidence has limits. A log may show that named protected actions were blocked under one configuration. It does not prove that every attack is impossible. Report the tested scope and revalidate when tools, credentials or environment permissions change.
Test denied actions deliberately
A proposed agent-security pilot should include actions the task is not allowed to perform. Attempt an unrelated file write in a safe environment, request another tenant’s fixture, invoke an unauthorized production tool and follow an instruction embedded in external text.
Expected outcomes should be defined before execution. The system should deny the action, retain a useful reason and continue or stop according to policy. The allowed task should still work; a boundary that rejects every operation is not a successful development environment.
Use protected cases that the worker did not use to tune its behavior. Retain failures and the exact configuration. A prompt that resists one known injection does not establish that the tool boundary is enforced. Inspect whether the prohibited effect was technically impossible or merely declined on that attempt.
These are proposed experiments, not executed tests of the Salars environment. A real result should distinguish simulated adversarial input, local enforcement and live integration behavior. The report should not convert a successful demonstration into a universal security claim.
Delegation should preserve the original boundary
An agent may ask another agent or tool to help with a task. That delegation should carry the same authority limit. A worker authorized to inspect a fixture cannot grant its helper permission to read production records simply because the helper needs more context.
The coordinator should define how task contracts propagate. Give the helper the relevant specification, owned files and permitted resources, while keeping unrelated secrets and conversation content out of its context. Its output should return through the same acceptance path as other contributions.
A delegated report is still evidence to inspect. If the helper says a protected action was allowed, the original worker should not infer that the user’s scope changed. Access questions should return to the responsible authority and the enforceable tool boundary.
Test at least one delegated denied case in a proposed pilot. The parent might request an analysis while the child tries to expand the target set. The expected result is a retained refusal or boundary failure, with no prohibited effect. This checks that security depends on the actual capability chain rather than a cooperative agent at the first layer.
The result should remain scoped to the tested delegation route. Adding another tool or external service can create a new authority path and require revalidation.
Authority should end with the contract
When a task ends, revoke or narrow temporary access. Preserve the patch and evidence through the repository’s normal process. A completed contribution should not require leaving a broad credential active.
If a worker is interrupted and another takes over, communicate the ownership change. The first worker should not resume later with stale instructions and competing authority. Task continuity and access continuity are separate concerns.
Review standing identities periodically. Which tools and permissions remain necessary? Has a preview environment become a forgotten production-like service? Are old fixtures retaining customer information? Cleanup should follow existing policies and preserve needed evidence.
A contractor model makes this understandable: the work can remain, the account can end and the owner still needs to maintain the result. The factory’s value lies in useful artifacts under accountable control, not in accumulating autonomous sessions.
What Would We Do at Salars?
A proposed Forge pilot would give agents broad enough source context to understand the app and narrow enough write ownership to keep changes reviewable. It would use safe merchant fixtures, test identities and the existing authorized release path. Production credentials would not be workspace defaults.
The first boundary exercise would test denied paths as well as the intended parser correction. A reviewer would inspect whether enforcement came from the tool and environment rather than the agent’s cooperative response alone. Failures would remain visible with their scope and repair evidence.
Supplier Margin Guard would initially produce read-only diagnostics. Adding live merchant actions would require new authority, write binding, recovery and privacy evidence. A successful code-generation task would not establish permission to operate the customer’s business.
An agent can be useful and powerful within a narrow contract. The task is to make that contract real in the tools, resources and evidence the system controls.
Sources
- OWASP: LLM01:2025 prompt injection, read October 7, 2026, for indirect injection, least privilege and limits of prompt-only defenses.
- GitHub Actions: secure use, read October 7, 2026, for minimum workflow permissions and secret-redaction limitations.
- Salars AI library, for related permission and verification guides. All proposed task contracts, merchant examples and boundary tests remain illustrative rather than executed security assessments.
Loading comments…