Skip to content

Permissions, approvals and audit: the rules for a trustworthy AI agent

9/26/2026
Last updated: September 26, 2026. This article describes practices and sources available as of this date. The section on 2027 contains planning scenarios, not guaranteed predictions. Legal information is general and does not substitute for case-specific analysis.

An AI agent becomes truly useful when it goes beyond drafting a response and can actually consult applications, modify records, send messages, or kick off processes. That's also when a key question—often postponed in pilot projects—arises: on whose behalf does the agent act, what is it allowed to do, and how can we later prove what happened?

For an SME, the answer shouldn't be a heavy corporate infrastructure. What’s needed is a proportional set of controls, applied where they matter most. The language model can interpret a request and suggest the next step, but it shouldn't be the only component deciding whether that step is authorized. Permissions, approvals, and logging must be enforced by code, identity, and policies outside the model.

This guide explains a practical architecture for AI agents working with email, CRM, documents, invoices, websites, or infrastructure. It doesn’t promise zero risk and doesn’t treat human approval as a universal fix. The goal is more realistic: to limit the impact of an error, stop unauthorized actions before execution, and be able to reconstruct events without turning the log into another sensitive database.

Executive summary

A trustworthy agent needs six properties:

  1. Distinct identity and user assignment: the system must distinguish between the agent, the person who initiated the task, and the service executing the action.
  2. Minimum privileges: the agent gets access to only the tools, operations, data, and environments needed for the current task.
  3. Authorization at every significant effect: verification happens when the agent tries to use the tool, not just at login or the start of a conversation.
  4. Risk-proportionate approvals: routine reads can be automatic, while publishing, making payments, deleting, or sending data require explicit confirmation or dual control.
  5. End-to-end audit: every action must be linked to the initial request, identities, applied policy, approvals, and the actual outcome.
  6. Revocation and stop: access must expire or be quickly revoked, and there must be a way to halt tasks already in progress.

These principles aren’t tied to a specific model or provider. They adapt well-established access control and zero-trust architecture practices to the realities of AI agents: dynamic plans, tool calls, unsafe external content, delegation, and the ability to produce effects across systems.

Why an agent should not be treated as a simple chatbot

A chatbot typically generates text that the human reviews before using. An agent, on the other hand, may follow a longer loop:

  1. receives a goal;
  2. consults data or documents;
  3. decides which tool to use;
  4. executes an action;
  5. observes the result;
  6. continues until it considers the goal reached.

The risk doesn’t come just from a wrong answer. It may stem from a correctly executed, but inappropriate, action: an email sent to the wrong recipient, a discount given to an ineligible client, a record deleted, a confidential file uploaded to an external service, or a command run in production.

Research confirms that this is not just a theoretical issue. ToolEmu built scenarios for high-stakes tools and identified failures with possible financial or privacy consequences. AgentDojo evaluates agents using tools on unsafe external data and shows how prompt injection attacks can steer agent actions. These results don’t provide a universal risk rate for any implementation, but they do show why even a well-written prompt can’t replace technical controls.

The design rule is simple:

The model proposes. Policy authorizes. The tool executes. The log provides proof.

Five concepts that are often confused

Authentication

Authentication answers the question “who is this?”. For an agentic flow, there can be multiple identities involved at once:

  • the user who made the request;
  • the agent or its execution instance;
  • the application orchestrating the process;
  • the technical account used to access an API;
  • the person approving a sensitive action.

If all of these show up in the log under the same generic account, accountability becomes impossible to determine.

Authorization

Authorization answers the question “what can this identity do to this resource, in this context?”. It’s not enough for the agent to have access to the CRM. Policy must differentiate between reading a contact, modifying it, exporting a list, and deleting the record.

Delegation

Delegation means that a person or service grants the agent limited authority for a task. Delegated authority should never exceed that of the person and should not be automatically passed to a subagent without explicit limits.

Approval

Approval is a one-time decision regarding a specific action. It does not substitute for authorization. An interface asking the user to approve a payment must not allow the payment if the user never had the right to make it in the first place.

Audit

Audit is the ability to reconstruct who made the request, what the system decided, what policy was applied, what was executed, and with what result. Audit is not the same as real-time monitoring, nor is it equivalent to storing entire conversations.

The state of technology as of September 26, 2026

The fundamental principles are mature, but identity and delegation for agents are still a developing area. In February 2026, NIST published a conceptual paper on software and AI agent identity and authorization. Its status is conceptual, not a final standard. However, it demonstrates that the industry is trying to apply familiar identity and access management practices to agents, rather than inventing separate security for each platform.

NIST Zero Trust Architecture states a relevant principle: access is decided for every request, with least privilege, and no implicit trust based solely on network position. NIST SP 800-53 Rev. 5 provides established control families for access, least privilege, identification, audit, and log protection. These publications aren’t dedicated recipes just for agents, but form a solid foundation.

In the agent space, OWASP Top 10 for Agentic Applications 2026 and the guide Agentic AI Threats and Mitigations cover risks like objective hijacking, abusive tool use, identity and privileges, tainted memory, and cascading failures. Published on September 1, 2026, the OWASP Agent Control Standard suggests control points and declarative policies enforced during execution, but it’s a new open standard, not automatic proof that an implementation is secure. The current OpenAI documentation on guardrails and approvals recommends placing checks at the boundary of the tool producing the effect. Google Cloud describes dedicated identities, least privilege, deny policies, and human-in-the-loop modes for MCP tools.

The practical takeaway in 2026 is that there isn’t a single product that will automatically turn an agent into a trustworthy system. Control comes from combining identity, policy, isolation, approvals, logging, and testing.

The permissions model: more precise than just “read” and “write”

A useful permission for agents must be described along several axes:

Axis

Control question

Example

Tool

What integration can it use?

CRM, email, invoicing

Operation

What function can it call?

read, create, update, delete

Resource

On which objects?

only the leads of the Bucharest team

Fields

Which data can it view or change?

without personal ID numbers and without bank data

Environment

Where can it operate?

test, not production

Time

How long is access valid?

20 minutes for the current execution

Volume

How many operations can it perform?

maximum 25 messages per run

Value

Up to what limit?

offer under 5,000 lei, excluding payments

Destination

To whom can it send data?

only approved domains

Context

On whose behalf and for what purpose?

for user X, ticket Y

A generic role such as admin does not express these limits. Nor is a broad scope like crm.write sufficient for an agent who only needs to update the status of a lead.

Separate identity for each relevant agent

Avoid shared technical accounts among multiple agents. A distinct identity allows:

  • revoking a single agent without disrupting all automations;
  • different policies for sales, support, and finance;
  • the exact tracking of actions;
  • detecting an agent behaving unusually;
  • independent rotation and expiration of credentials.

The agent’s identity should not obscure the user. In a delegated flow, logs should retain both the agent’s identity and the user or process on whose behalf it acts.

Short-lived credentials, kept out of the prompt

API keys, access tokens, and passwords should not be placed in prompts, memory, or documents consulted by the model. A gateway, a vault, or even a simple internal service can attach the credential only after the policy has approved the request.

Technical preferences are:

  • short-lived tokens;
  • validated audience for the destination service;
  • scopes narrowed as much as possible;
  • automatic rotation;
  • central revocation;
  • no forwarding of received tokens to another service.

RFC 9700 brings together current OAuth 2.0 security practices, and RFC 9396 allows for more detailed authorization descriptions than a simple scope list. Not every SMB needs to implement these protocols themselves, but the chosen provider should be able to explain token audience, expiration, revocation, and privilege limitation.

Authorize at the tool, not via the model prompt

The phrase “do not delete anything without approval” in the prompt is useful as guidance, but it is not a security barrier. A prompt injection attack, a planning error, or ambiguous context may cause the model to ignore or misinterpret it.

The check must be repeated right before the operation:

  1. the agent proposes the tool call and its arguments;
  2. a policy enforcement point checks identity, resource, operation, and context;
  3. the policy allows, rejects, or requests approval;
  4. the necessary credential is attached only after the decision;
  5. the tool re-validates the right to the resource;
  6. the result is logged.

This double check prevents situations where the orchestrator accepts an action but the API executes it using an unlimited technical account.

Separation of read from write and planning from execution

An agent that analyzes invoices doesn’t need the implicit right to pay them. An agent drafting an article shouldn’t be able to publish it automatically. Separation can be accomplished with different tools, different identities, or distinct workflow steps.

For sensitive tasks, a two-step model is useful:

  • planning: the agent reads, analyzes, and produces a structured proposal;
  • execution: a deterministic service validates the proposal and applies only the approved operations.

The model does not get direct access to the database if a narrow function such as update_order_status(id, allowed_status), is sufficient.

Approvals: when they help and when they become security theater

Not every call requires human intervention. If a person has to confirm every read, the system slows down and people learn to hit “Approve” without checking. This leads to approval fatigue.

A practical classification can start from five questions:

  1. Is the action reversible?
  2. Does it impact money, production, reputation, or the rights of a person?
  3. Does it send data outside the organization?
  4. Is the volume or value unusual?
  5. Has the agent previously and successfully performed this combination of operation and context?

Indicative control matrix

Level

Examples

Recommended control

Low

searching in non-sensitive internal documents, reading product catalog

automatic execution, volume limits, logging

Moderate

creating drafts, updating a reversible field, internal reply

automatic policy, schema validation, human sampling

High

external sending, publishing, price modification, access to personal data

explicit approval before effect, current authentication

Critical

payment, mass deletion, privileged production access, changing bank account

dual control, additional authentication, strict limits or prohibition

This is a starting matrix, not a legal classification. Thresholds need to be adjusted to the company, process, and real-world consequences.

What the approver needs to see

A proper approval doesn’t require confirmation of a vague intent such as “the agent wants to continue.” The interface must display:

  • the exact action;
  • the system and target resource;
  • fields that are changing, before and after;
  • the recipient and data leaving the organization;
  • amount, currency, and beneficiary, if money is involved;
  • the operational reason and original request;
  • risks or rules that triggered the approval;
  • the period during which the approval is valid.

The approval must be cryptographically linked or uniquely identified to the specific action. If the agent changes the recipient, amount, or attachment after approval, the previous approval is invalidated.

When dual control is necessary

The four-eyes principle is suitable for high-impact actions: payments above threshold, changing a supplier’s account, massive data exports, privileged changes in production, or irreversible deletions. The second person must have an appropriate role and must not be the same identity that initiated the request.

Expiration, denial, and unavailability

An approval request must expire. If the policy or approval service is unavailable, sensitive actions should fail safely, not be executed by default. Denials must be communicated to the agent as a clear final state to prevent repeated re-tries of the same action until control is bypassed.

What the audit log must contain

A useful log answers questions, not just stores text. For each important run and call, it is recommended to have:

  • a unique identifier for the run and a correlation ID between systems;
  • the initiating user, the agent, the orchestrator, and the executing service;
  • the stated purpose and reference to the ticket, order, or process;
  • the agent configuration model and version;
  • the tool, operation, and target resource;
  • the relevant arguments, redacted or pseudonymized where necessary;
  • the source of external data used for the decision, identified by ID, version, or hash;
  • the policy version and the rule that allowed, denied, or escalated the call;
  • the approver, timestamp, decision, and the exact approved object;
  • the tool’s result, status code, and resulting effect;
  • retries, cancellations, and compensation operations;
  • duration, usage, and cost, if relevant for abuse detection.

For distributed flows, OpenTelemetry provides tracing conventions that can link the agent’s invocation, model calls, and tool executions. These technical traces must be supplemented with policy decisions and business approvals. A simple performance trace is not automatically an audit log.

It is not necessary to store the model’s “thoughts”

Auditability does not require retention of the full internal reasoning. This may be unavailable, unstable, or contain data we do not wish to keep. It is more useful to store:

  • only the strictly necessary operational inputs;
  • the plan or structured proposal presented to the policy system;
  • reason codes for the decision;
  • tool invocations and their results;
  • evidence and versions of relevant sources.

This way, we can explain the observable process without treating a model-generated text as legal justification or absolute truth.

The log should not become a data leak

A log that retains tokens, passwords, full prompts, attachments, or personal data without limitation may create a greater risk than the one it aims to reduce. EDPB guidance on the privacy risks of LLMs recommends access controls and logging, as well as minimizing data in logs.

Minimum practices:

  • do not log secrets and tokens;
  • mask sensitive fields;
  • restrict access to logs;
  • define retention periods by category;
  • protect log integrity and separate system administrators from audit administrators;
  • document who can export or delete records;
  • test whether an incident can be reconstructed before relying on these logs.

A suitable reference architecture for an SME

It is not necessary for each element to be a separate product. What matters is separation of responsibilities:

  1. The user interface authenticates the individual and captures the purpose of the request.
  2. The agent orchestrator plans steps and proposes tool invocations.
  3. The Policy Gateway verifies identity, action, arguments, environment, volume, and context.
  4. Approval Service interrupts execution and collects a human decision when required by policy.
  5. Credential Broker issues a limited token only after authorization.
  6. Tool Adapter validates the schema and calls the destination API.
  7. Audit and Monitoring System correlates proposal, decision, and result.

The policy must be enforced in a component that the model cannot alter. If the agent can edit the file that defines its boundaries, control is only superficial.

Example of a clearly expressed policy

The following rule is conceptual—not syntax for any specific product:

  • Agent: customer support.
  • Action: updating a contact in the CRM.
  • Allow if: the user has a support role, only the phone number or contact preference is being changed, and the client is part of their portfolio.
  • Limits: one record per call, in production environment.
  • Request approval if: a field classified as sensitive is being modified.
  • Deny if: the action requires deletion or external export.

The advantage of explicit policies is that they can be reviewed, tested, and versioned. Natural instructions in the prompt remain useful for behavior, but deterministic policy enforces authority.

Three practical examples

1. Agent for email and quoting

The agent reads a client request, consults the catalog, and prepares an offer. A balanced setup might automatically allow:

  • reading the current conversation;
  • consulting products and inventory;
  • creating an offer draft;
  • saving the draft in the CRM.

Approval would be required for:

  • sending the external email;
  • applying a discount above threshold;
  • attaching a document containing personal data;
  • adding a new recipient.

It would completely refuse exporting the client database and changing bank information. The log would link the sent offer to the prices and discount policy in effect at that moment.

2. Agent for invoices and payments

The agent can extract invoice data, verify the supplier, and prepare the payment. However, the right to read invoices does not imply the right to make payments.

A secure flow separates:

  • the analysis agent, with read-only access;
  • the deterministic validator for IBAN, duplicates, due date, and threshold;
  • the person who approves;
  • the payment service, with a short-lived token and amount limit;
  • subsequent reconciliation.

Changing the IBAN and the first payment to a beneficiary should trigger an additional check, not be learned automatically from a received email.

3. Agent for site administration

The agent researches and drafts an article in the CMS. Reading existing articles and saving a draft are reversible operations. Publishing, modifying an already published article, and deleting an image have external impact.

Permissions can be separated as follows:

  • read published content;
  • create and update only their own drafts;
  • no overwriting if the version has changed between reading and saving;
  • publish only after editorial approval;
  • deletion forbidden for the agent;
  • log source, version, approver, and resulting URL.

This model prevents both accidental publishing and the loss of changes made in the meantime by another person.

MCP and tool authorization

The Model Context Protocol standardizes how applications expose tools and resources to agents. The protocol does not remove the obligation to design permissions. An MCP server that exposes a dangerous function with an overly broad token remains dangerous.

Current MCP documentation and SDKs are based on OAuth and protected resource metadata. Relevant practices include:

  • validating that the token was issued for the specific server;
  • blocking token passthrough to downstream services;
  • using limited scopes and increasing them only as needed for each operation;
  • clear consent for clients and servers;
  • HTTPS and strict validation of redirects;
  • separate credentials for external services;
  • protection against the 'confused deputy' attack, where a service is tricked into using its authority on behalf of someone else.

The MCP specification and guides must be checked at implementation time, as the protocol evolves. For an SME, the essential criterion is not just MCP compatibility, but the provider's ability to limit each tool and generate useful audit trails.

Testing before autonomy

An agent should not be promoted to production just because it successfully completes a few demos. Testing must cover both success and refusal.

Functional and policy tests

Explicitly check:

  • operations permitted under normal conditions;
  • operations refused for the wrong role;
  • resources from another department or another client;
  • exceeding volume or value thresholds;
  • token expiration and access revocation;
  • modification of the action after approval;
  • service approval unavailability;
  • resuming a run after interruption;
  • accidentally doubling the same operation.

Adversarial tests

Insert into documents, emails, and test pages instructions that try to persuade the agent to ignore its objective, disclose data, or use another tool. The goal isn't just to check if the model says no. The crucial test is whether the authorization system blocks the effect even if the model suggests a wrongful action.

Run these tests after any change to the model, prompt, tools, policies, or connectors. An old evaluation doesn't guarantee the behavior of a new setup.

Shadow mode and gradual autonomy

Before going live, the agent can work in shadow mode: it proposes actions, and the system compares them to human decisions without any real impact. The next step allows low-risk operations to run automatically while requiring approval for others. Autonomy increases only based on observed data, not just because the demo “seems smart.”

Metrics that measure control, not just productivity

Besides time saved and completion rate, track:

  • percentage of calls denied by policy;
  • percentage of actions that require approval;
  • approval and denial rates by category;
  • approval wait times;
  • number of actions changed after being submitted to an approver;
  • attempts to access outside the allowed scope;
  • journal/log coverage for actions with impact;
  • number of overly broad or unused credentials;
  • time needed to revoke and stop the agent;
  • incidents, near-incidents, and compensated operations;
  • the gap between shadow mode proposals and human decisions.

A very high approval rate doesn't necessarily mean safety. It might signal unnecessary approvals or reviewer fatigue. A very high denial rate might point to a poorly configured agent or a policy that doesn't fit the process.

Legal and compliance aspects

Permissions and auditing are technical controls. They alone don't determine the legality of processing or classify an AI system.

If the agent processes personal data, GDPR principles still apply: defined purpose, data minimization, security, retention periods, and the operator’s ability to demonstrate compliance. The audit log must be designed in line with these principles. Keeping everything “for audit” indefinitely and without clear purpose isn’t a prudent solution.

European Union Artificial Intelligence Act includes requirements for automatic log retention and human oversight for some high-risk systems. These obligations do not make every internal agent a high-risk system. Classification depends on usage, the organization’s role, and context. In addition, the timeline for applying some rules to high-risk systems was modified and clarified in 2026. The European Commission maintains an updated page about AI Act application.

For decisions affecting employees, lending, access to essential services, or other rights, a specific legal assessment is required. This article provides a technical framework, not legal advice.

Common mistakes

“The agent uses the administrator account, but the prompt tells it to be careful”

The prompt doesn't reduce the account’s privileges. If the agent is compromised or makes a mistake, the full administrator impact still applies.

“The user accepted access during installation”

Initial consent for integration isn’t approval for all future actions. Sensitive operations must be assessed in context.

“All actions require approval, so we’re safe”

Too many approvals create routine. Furthermore, a person may be misled by an incomplete description. The interface must show the exact effect, and policy must block unauthorized actions regardless of superficial approval.

“We have the vendor’s logs”

Logs from a single vendor may show the model call, but not the end-to-end flow from user, orchestrator, policy, approval, to the final API. Check for correlation identifiers and export capabilities.

“We keep all prompts for investigations”

This can copy personal information, secrets, and entire documents into the log. Define required fields, mask content, and store references or hashes where full text isn’t justified.

“The sub-agent automatically inherits all rights”

Chain delegation can amplify privileges and break the link to the original user. Each sub-agent should be given a set of rights no greater than, and usually more restricted than, the previous agent, tailored to its subtask.

“If the operation fails, the agent may try again”

Retrying without an idempotency key and without checking state can result in duplicate payments, orders, or messages. Tools with side effects must distinguish a safe replay from a new operation.

Practical 30-day implementation plan

Week 1: inventory and classification

  • list agents, tools, technical accounts, and owners;
  • document accessible data and environments;
  • classify operations by impact and reversibility;
  • identify shared accounts and admin privileges;
  • define who can stop each agent.

Week 2: permissions and identity

  • create separate identities for important agents;
  • separate reading, writing, publishing, and deleting;
  • move secrets out of prompts and memory;
  • introduce short-lived tokens or rotation;
  • limit resources, environments, volumes, and destinations.

Week 3: approvals and audit

  • define risk matrix and thresholds;
  • implement interruption before effect, not after;
  • show concrete difference and approver's recipient;
  • add correlation IDs and policy version;
  • redact sensitive data from logs;
  • test export and run reconstruction.

Week 4: testing and gradual rollout

  • run permitted, denied, and adverse cases;
  • test revocation, expiration, and unavailability;
  • start in shadow mode or with reversible operations only;
  • analyze approvals for signs of fatigue;
  • establish monthly privilege and incident review;
  • document criteria for increased autonomy.

What could change in 2027

The following are plausible scenarios, not certainties.

Better-standardized agent identities

NIST initiatives and evolving cloud platforms point to the possibility of dedicated workload identities for agents, with better delegation and traceability. We are likely to see tighter integration between agent registries, IAM, gateways, and policy systems.

Finer and temporary authorization

Instead of static, broad scopes, platforms may increasingly use on-demand authorization, restricted by action, resource, amount, and time window. This would reduce permanent credentials but might increase complexity and reliance on central services.

Policy-assisted approvals and separate models

Moderate-risk operations could be automatically evaluated by a policy engine plus an independent control model. However, 2026 research shows that model-based monitors can themselves be bypassed. Therefore, such an evaluator should not replace deterministic barriers for critical operations.

Interoperable audit across multiple agents

As tasks move between agents and organizations, portable evidence of delegation, policies, and outcomes will be required. Observability conventions help, but a universal solution for non-repudiation and full provenance in agentic flows does not yet exist.

For an SME, the sound strategy is to use existing standards now, avoid lock-in with proprietary formats, and retain control over identities, policies, and essential logs.

Checklist before granting real access

  • Does the agent have a distinct identity?
  • Can we identify the person or process on whose behalf it is working?
  • Are permissions limited by operation, resource, environment, time, and volume?
  • Are secrets kept out of the prompt, memory, and logs?
  • Is policy enforced outside the model, even before the tool?
  • Do publishing, payment, deletion, and external transfer have clear thresholds?
  • Does the approver see the exact effect, not just a generic description?
  • Does approval expire and become invalid if the action changes?
  • Can we revoke access and stop active runs?
  • Are repeated operations idempotent?
  • Does the log link the request, policy, approval, and result?
  • Are logs minimized, protected, and do they have a retention period?
  • Have we tested for attacks via external data and for denial behavior?
  • Is there an owner, incident procedure, and periodic review?

Conclusion

Trust in an AI agent should not be measured by how convincingly it explains what it plans to do. It must be built through verifiable limits.

A trustworthy agent is not one to which we grant unlimited access simply because the model appears capable. It is an agent that can operate efficiently within a clear perimeter, uses temporary authority, requests appropriate approval before any effect, and leaves sufficient evidence for verification. If an error or attack occurs, the architecture limits the damage and allows intervention.

For most SMEs, the first step is not acquiring a complex platform. It is inventorying actions, separating reading from writing, removing administrative accounts from automated flows, and defining three categories: what the agent can do alone, what requires approval, and what it is never allowed to do.

Sources and recommended reading

Would you like us to assess a specific process, the necessary permissions, and the control architecture best suited to your company? We can turn a promising pilot into a measurable, auditable, and production-ready system.

Recommended for you

The 90-Day Adoption Plan: From First Pilot to a Measurable Agentic System

How much is an AI agent worth: total cost, KPIs, and investment returns

When the agent buys for us: budgets, payments, and autonomy limits

Cookies

We use cookies required for the site to work. With your consent we also enable additional features (videos, maps) or anonymous statistics. You can change your choice at any time from the site footer.

Cookie policyPrivacy noticeTerms and conditionsCookie preferences