AI · Article 15 of 54 · Part 4

The Permission Architecture for Safe AI Agents

How should an agent's permissions be constrained before it can act?

An assistant that drafts a purchase recommendation and an agent that submits the purchase may produce similar-looking text. Their consequences differ. One prepares a decision. The other commits money and creates obligations outside the conversation.

That boundary should be designed before the system acts. A message telling an agent to be careful does not remove a credential’s authority, prevent a tool from deleting a record, or make an external action reversible.

How should an agent’s permissions be constrained? Begin with the task’s actual needs. Separate observation, preparation, and commitment; limit the resources and operations available; enforce consequential boundaries outside the model; and preserve a way to stop and inspect actions.

“Safe” here means a design intended to reduce identified risks. It does not promise that every authorized action is correct or that a permission system eliminates all failure. Accuracy, security, and legitimate authority remain separate concerns.

Permission is narrower than capability

A model may be able to describe a command, infer a workflow, or propose a transaction. None of that establishes permission to carry it out.

Capability concerns what a system can do. Authority concerns what it may do under a particular user’s instructions and rights. Access concerns what connected tools and credentials technically permit. A sound design tries to align these boundaries rather than assume they are already identical.

For example, a supplier-analysis agent might need product specifications and prices. It need not receive the ability to change supplier bank details. A calendar assistant might need availability information without permission to expose confidential meeting descriptions to another service.

Those distinctions should survive changes in prompts and models. If a more capable model arrives, it should not acquire broader authority merely because it can now figure out how to use an existing tool creatively.

The worker-to-governor chapter concerns human authority and accountability. This chapter concerns the operational boundaries through which that authority becomes enforceable.

Divide the workflow into three stages

Observation collects information. Preparation transforms it into a proposed result. Commitment changes something outside the preparation environment.

This division provides a useful starting point for permissions. An observation stage might read an approved folder. A preparation stage might create a draft in a temporary location. A commitment stage might publish, send, purchase, or update an authoritative record.

The stages can share a system without sharing identical access. A research task does not need a publication credential just because its final output may later be published. A draft invoice does not need permission to change the bank account used to receive payment.

This also makes review more concrete. The reviewer sees the proposed commitment, its inputs, and the expected effect. Approval should apply to that proposal, not to an unspecified future sequence of actions that the agent may invent after the review.

A narrow routine can sometimes move through the stages automatically under standing authorization. The authorization should describe the routine, limits, and exceptions. A system should not reinterpret convenience as permission to expand its own role.

Write a permission contract

A permission contract is a plain statement of the allowed task. It can be implemented through existing access controls and tools; it does not require a new software platform.

Include the permitted inputs, outputs, resources, operations, and limits. State the actions that require escalation and what should happen when permission is absent or unclear.

For a hypothetical report assistant, the contract might allow reading one approved dataset, creating a draft report, and saving it in a review folder. It might prohibit sending the report externally, altering the source data, or including personal details outside the report’s purpose.

For a purchasing assistant, the contract would need different boundaries: approved products, suppliers, budget, delivery terms, and the person who can approve exceptions. Financial commitments require controls appropriate to their consequences.

The contract should be understandable to the person authorizing it. If the description is too vague to reveal what could happen, it is too vague to support meaningful delegation. “Handle the business” is an ambition, not an operational permission boundary.

Make tools fit the task

A narrowly designed operation is easier to constrain than an unrestricted interface. A tool that writes a draft to one folder differs from a tool that can run arbitrary commands anywhere its credential allows.

General tools are sometimes necessary, especially for development and investigation. Their use should make the larger access boundary visible. A system should not present a broad command runner as if it were equivalent to a single harmless function.

Before connecting a tool, inspect its operations. Can it read, create, update, delete, send, share, or change permissions? Which of those operations are needed? Can unused operations be removed or denied in the downstream service?

A tool’s friendly name is not enough. A “mail helper” might summarize messages and also send them. A “database assistant” might read records and also alter the schema. The actual operations and credential scope determine the exposure.

OWASP’s Excessive Agency guidance is relevant security context. The practical design below applies limited authority to concrete task boundaries rather than relying on a model’s promise of restraint.

Keep credentials out of generated-code reach where possible

Credentials are authority-bearing objects. If a generated program can read a powerful token, a prompt instruction about acceptable use is a weak boundary around that authority.

A preferable design uses an existing service or proxy to perform approved operations without exposing the underlying secret to the agent’s execution environment. The exact implementation depends on the platform and should be assessed by someone competent to configure it.

A current vendor example illustrates the architectural distinction. Anthropic’s April 2026 Managed Agents description separates session records, the agent harness, and execution sandboxes. It describes vault and proxy arrangements intended to keep external credentials away from generated-code environments. This is a documented product design, not proof that every use is secure. Engineering description.

Small operators do not need to imitate a vendor’s infrastructure. They should ask the simpler question: can the task be performed with narrower existing access, a draft-only integration, or an approved operation that does not reveal a general-purpose secret?

Treat incoming material as data

An agent may encounter documents, web pages, emails, or tool results that contain instructions. Those instructions are not automatically authorized by the owner.

A supplier page could include text telling the system to disregard a budget. A document could ask it to reveal private information. A customer message could request an action the sender has no right to authorize.

The system should distinguish the content being examined from the rules governing the examination. That distinction requires more than telling the model to ignore suspicious text. The permitted tools and downstream authorization should still prevent forbidden actions if the model misinterprets an input.

For example, a research assistant with no sending function cannot forward confidential documents through that function. A system limited to one review folder has less opportunity to alter authoritative records elsewhere. Those boundaries do not solve every possible attack, but they reduce dependence on perfect interpretation.

The inbox for reality concerns retaining observations honestly. Its content should inform decisions without acquiring authority to rewrite the permission contract.

Authorization should attach to a specific action

A review that approves an abstract plan may leave the actual commitment unclear. A stronger action packet identifies what will change, where, under whose authority, and with what expected consequence.

For a message, include the recipient and exact content. For a purchase, include the item, quantity, supplier, total commitment, and relevant terms. For a record update, include the prior value and proposed replacement where practical.

If the proposal changes after approval, the system should determine whether the authorization still applies. A formatting adjustment may be within scope. A new recipient, larger purchase, or changed payment detail may not be.

The aim is to prevent approval from becoming a reusable blank check. Authorization for one report does not imply authorization for every later report, especially if inputs or intended audience change.

Standing authorization can still be useful. It should describe a class of actions clearly enough that the system can check membership and route exceptions. A budget and recipient rule are more enforceable than “use good judgment.”

Put transaction checks at the commitment boundary

Important actions often have several conditions that must hold together. A purchase might need an approved supplier, sufficient budget, an acceptable item, and a valid delivery address.

Check those conditions close to commitment. A check performed earlier can become stale if records change before the action occurs. The system should avoid treating a valid proposal from yesterday as unquestionable authority today.

Retries require attention too. If a service times out after receiving a request, the agent may not know whether the action happened. Repeating the request could create a duplicate order or message.

Use existing transaction identifiers, status checks, and duplicate-prevention features where available. When uncertainty remains, route it for inspection instead of confidently assuming either success or failure. A tool error and an external outcome are not always the same event.

A permission design that governs only the first attempt leaves an important gap. It should also govern retries, partial completion, cancellations, and recovery.

Limit accumulation as well as individual actions

A per-action limit can conceal a larger commitment. Ten small purchases can exceed a budget even if each purchase is individually allowed. Several harmless-looking workflows can collectively expose sensitive information or consume too much review time.

Set aggregate boundaries where the consequence requires them. These can include total spending over a period, action frequency, simultaneous tasks, destinations, and the amount of information accessible to a workflow.

The limits should be based on the operation’s needs and the owner’s ability to absorb or correct mistakes. There is no universal safe dollar amount or number of actions.

A useful hypothetical purchasing design checks both the individual order and outstanding commitments. Money already promised should not disappear from the calculation merely because payment has not yet settled.

The kill engine handles stopping rules across the operation. Permission architecture supplies the boundaries that make those stops effective, including revoking access and preventing new commitments while existing obligations are addressed.

Preserve evidence without hoarding sensitive data

Records should make consequential actions inspectable. They can identify the initiating task, relevant input version, proposed action, authorization, tool attempt, and verified result.

The record should distinguish what the model said from what the external system did. A generated sentence claiming that a message was sent is weaker evidence than a verified send result associated with the approved recipient and content.

Logging should fit privacy and retention obligations. Storing every credential, private message, and raw document in a convenient transcript can create a new security problem. Record enough to support accountability without retaining unnecessary secrets.

The provenance chapter develops these distinctions. For permissions, the key question is whether a responsible person can inspect why an action was allowed and correct the result when the boundary failed.

A log is not prevention. Its value depends on someone being able to use it and on the system having a workable correction path.

Test boundaries with actions that should be denied

A permission system should be evaluated on forbidden and ambiguous requests, not only on successful authorized ones. These are proposed tests, not claims that a particular agent has passed them.

Ask whether a read-only task can alter a record. Try an out-of-scope destination, an expired approval, a duplicate request, a changed proposal, and a total commitment above the aggregate limit. Include an input document containing instructions that conflict with the task.

A meaningful test checks the external state, not just the agent’s verbal response. Saying “I cannot do that” while a downstream operation succeeds is a failure. Saying an operation succeeded when nothing changed is a different failure that should also be visible.

Preserve cases that previously exposed a gap. When tools or policies change, rerun relevant cases to make sure the boundary has not been weakened. The learning ledger records the supported scope and the trigger for revalidation.

Delegation should preserve the original boundary

One agent may ask another to perform a subtask. That handoff should not expand the authority of the initiating task. A research request remains a research request even if a second agent has access to publication tools.

Describe the delegated work with the same care as the original permission contract. Identify the allowed inputs, intended result, and prohibited commitments. Check that the receiving system can actually enforce the restriction rather than merely receive a polite instruction.

A failure can occur when authority is treated as belonging to the tool instead of the task. The second agent may possess broad credentials for other work, but those credentials do not establish that this particular request may use them. The external authorization should account for the initiating user and the permitted action.

Delegation also affects records. The owner should be able to follow which system proposed an action and which system performed it. A handoff should not erase responsibility or make an exception disappear inside a new conversation. If that continuity cannot be maintained, keep the subtask in preparation mode or use a simpler arrangement that the owner can inspect.

Expand authority only where the evidence applies

A preparation-only trial may show that a system is useful at gathering information. It does not automatically establish that autonomous commitment is justified.

An expansion should specify the new action class, its risk, the evidence supporting it, and the recovery path. Preserve the narrower mode as a fallback where possible.

The strongest objection to this approach is the review burden. Too many approvals can erase the productivity gain. That is a legitimate design concern. The response is to narrow and evaluate repeatable classes of action, improve the proposal packet, or use deterministic checks for mechanical conditions.

It should not be to remove consequential boundaries indiscriminately. A simpler task with narrow authority can be more useful than an ambitious agent whose output requires constant anxiety and inspection.

AI Leverage in Practice

What changed: agents can connect interpretation to external actions, making the distinction between preparation and commitment more consequential.

What you can do today: map one task’s inputs, tools, credentials, destinations, individual limits, aggregate limits, and exception route. Remove unnecessary operations. Test at least one action that should be denied and verify the external outcome.

What may come later: stronger models may use connected tools more effectively. Their improvement should not silently enlarge authority. Reassess the task and its boundaries whenever capabilities, integrations, or obligations change.

Authority that can be inspected

A useful permission architecture makes an agent’s authority narrower and more legible than its general capability. It gives the system enough access to do the authorized work and keeps consequential boundaries enforceable outside persuasive text.

Start with preparation, evaluate concrete actions, and expand deliberately. Keep stopping, recovery, and answerability part of the design.

Find the complete Age of AI Leverage series and the wider AI section.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home