Skip to content

When the agent buys on our behalf: budgets, payments, and limits of autonomy

9/27/2026

# When the agent buys on our behalf: budgets, payments, and limits of autonomy

An AI agent can compare offers, fill a cart, and, in some integrations, initiate steps in the payment process. For a small or medium business, the relevant question isn’t whether the agent “knows how to buy,” but what decisions we can delegate, within what limits, and who checks the result. A wrong order could mean a double payment, an unwanted subscription, delivery to the wrong address, or using an unapproved supplier.

This article looks at the company as the buyer. For a seller’s perspective on selling to agents, see the article on selling software as a service (SaaS) to AI agents. For access control and action logging, see the article on permissions and audit. The budget and threshold examples below are working scenarios, not universal recommendations or legal caps.

What an agent can do today and where company responsibility begins

As of September 2026, there are specifications and programs for agent-assisted transactions, but their availability depends on the provider, country, merchant, processor, and contract. For example, Google’s Agentic Payment Protocol (AP2) specification separates the mandate for the content of an order from the mandate for payment. Visa Intelligent Commerce documentation describes authenticated user instructions, limited payment credentials, and controls for merchant and amount. These are technical directions, not guarantees that any SME in Romania can activate all functions tomorrow.

Products also evolve. In its update from March 24, 2026, OpenAI explained that the initial Instant Checkout lacked needed flexibility, and is focusing on product discovery and merchant checkout experiences. By checkout we mean the final steps of purchasing—confirming cart, delivery, and payment method. So, before starting a project, check the actual flow available from your selected provider, not just the protocol’s name in presentations.

To decide what to automate, let’s separate four stages:

  1. Research and recommendation. The agent searches for products, compares prices, availability, terms, and conditions. It does not create commercial commitments yet.
  2. Order preparation. The agent fills out an internal request, cart, or purchase order. An internal request is not automatically an order accepted by the supplier.
  3. Placing the order and payment. The supplier may accept the order, and the bank or processor may authorize or process the payment. Order acceptance and payment confirmation are separate states.
  4. Receipt and reconciliation. The company checks the goods or services received, invoice, final amount, and possible returns. “Paid” doesn’t automatically mean “correctly delivered.”

Good autonomy at stage 1 doesn’t automatically justify autonomy at stage 3. The risk increases when an action incurs spending or creates obligations that are hard to reverse.

A practical scale of autonomy

The following scale is an editorial tool for design, not a legal standard:

  • Level 0, information: the agent provides options and arguments; the human does the rest.
  • Level 1, preparation: the agent builds the request and cart; the human reviews and submits.
  • Level 2, approved execution: the human approves the precise order content; the system submits it and then requests payment via the bank or processor’s accepted flow.
  • Level 3, limited restocking: the agent can repeatedly order only pre-approved items, from pre-approved suppliers, within automatic value and frequency limits. Exceptions are escalated to a human.
  • Level 4, delegated purchase and payment within limits: the agent can finalize a transaction through a payment integration that enforces verifiable instructions and limits. This level requires a compatible implementation, security checks, and agreement from participating institutions; it’s not achieved simply by installing an AI model.

For most SMEs, a pilot at levels 1 or 2 gives insight into recommendation quality and operations while keeping human approval for each purchase. The academic study τ-bench shows why rule compliance needs to be tested on concrete cases: evaluated agents had mixed results for tasks with domain policies and tools. The results relate to the benchmark scenarios—that is, the standardized test set—not a valid error rate for all commercial systems.

The budget must be enforced by the system, not just written in the prompt

The instruction sent to the model, called a prompt, may include “do not spend over 300 lei.” This can guide the agent, but real control needs to be in a separate component that checks the transaction before the order or payment is placed. The agent proposes; the procurement system decides if the proposal fits company policy. This is an architectural choice, supported by risks noted in research about agent policy compliance and in the OWASP guide for AI agent security.

Policies can, where relevant, include:

  • validated suppliers, their contractual identity, and official accounts or domains;
  • allowed categories, items, and variants, including bans on new or automatically renewed subscriptions;
  • order, user, department, and time-based caps, plus rules for repeated transactions and splitting one purchase into multiple orders;
  • acceptable total price, in the approved currency, including taxes, shipping, and known fees at checkout;
  • delivery address, maximum term, return conditions, and approval validity;
  • the approver, escalation thresholds, and who can change policy.

Policy should be checked on structured data from the cart or offer, not on the agent’s free-form summary. Additionally, the budget should be reserved while the order is in process, so two simultaneous sessions can’t each spend the same available amount. After a confirmed result, the reservation becomes a registered expense or is released, as appropriate. If the result is unknown, the reservation stays until clarified, to avoid spending the same funds again. These are design recommendations, not functions every platform offers out of the box.

Hypothetical scenario: an office allocates 2,000 lei per month for consumables, allows up to 300 lei per order, and only accepts products from a catalog and validated suppliers. The agent finds items for 280 lei, but shipping brings the total to 315 lei. The order doesn’t go through automatically, even if the product price is under the threshold. If, after approval, the quantity, supplier, address, or total changes, the system requires a new check. Splitting the same purchase into two orders shouldn’t bypass the approval threshold. Figures above are only to illustrate the mechanism.

Approval must be tied to the exact order

An “Approve purchase” button is ambiguous if the cart can change between approval and payment. The person approving should see at least the merchant, items and quantities, total cost and currency, delivery, possible recurring costs, and the approval’s validity period. The system saves the approved version and compares it to the submitted version. Any relevant change halts the transaction or sends the request back for approval.

This principle also appears in AP2, where the payment mandate is tied to a defined checkout. It's useful to distinguish between a commercial mandate, meaning the company's permission to buy under certain conditions, and payment authentication, performed by the payment service. Internal approval doesn’t replace the authentication required by the bank or processor.

In the EU, PSD2 Directive and Delegated Regulation (EU) 2018/389 set the framework for strong customer authentication in mandated cases. This uses at least two independent factors, such as something the user knows, possesses, and a biometric feature. Where strong authentication applies for remote electronic payment, Article 5 of the regulation requires the code to be tied to the agreed amount and payee—changing them makes the code invalid. There are exceptions under defined conditions, including for certain secure corporate processes, but you shouldn’t assume all business-to-business (B2B) payments are exempt. The actual flow needs validation with your payment service provider and, if needed, your legal advisers.

Card details and authentication should not be put in the agent’s conversation

An agent receiving a card number, security code, or authentication code over chat creates unnecessary risk. A safer route is a payment integration through an authorized provider, with limited credentials or tokens. A payment token is a representation used in the payment flow instead of raw card data; its value is tied to the purpose and restrictions set by the provider. Visa describes agent-specific tokens and instruction checks before issuing credentials, while OpenAI describes in its specification a delegated payment request with cap and expiry. These are technical models, not guarantees that any processor or merchant will accept them.

Before integration, ask your provider: who holds sensitive data, what permissions does the token have, which merchants and caps can it be used for, when it expires, how it’s revoked, how disputes are handled, and what happens if payment succeeds but order confirmation is delayed. Control separately the identity of the employee requesting the purchase and the identity of the app accessing the API—the interface through which software systems share data and execute actions.

What can still go wrong, even with a proper budget

Hidden instructions in external sources. A product page, offer PDF, or supplier email could have text attempting to alter the agent’s behavior. This is called an indirect prompt injection: information the agent reads is wrongly treated as instructions. AgentDojo studied such attacks in environments with tools. The practical solution is to treat offers as untrusted data and not let them modify policy, accounts, or approvals.

Offers that are no longer valid.Price, stock, shipping, or the rate may change between comparison and checkout. Check your final cart again. If the total or any key term differs from the approved version, request a new decision.

Duplicate orders. A dropped connection can cause the agent to resubmit the request without knowing the first attempt succeeded. Use a unique request identifier and idempotency, meaning that repeating the same request does not trigger a second purchase. Stripe’s documentation explains this principle for its API requests. The key must be reused for the same operation, and your integration should respect the provider’s retention period and rules. Safeguarding a payment request does not automatically eliminate duplicate orders in the system. Always check the order and payment status independently; a missing response is not proof of failure.

Costs that continue after the first payment. A SaaS subscription may renew automatically, bill per user, or have variable consumption charges. A limit on the first transaction does not necessarily cap future spending. Such an agreement requires a separate rule for subscription duration, amendments, termination, and the account owner.

Exposure of personal and business data. Employee names, delivery addresses, and invoices may reach external services. The European Data Protection Board explains the principle of data minimization for small businesses. Provide the agent and suppliers only the data needed for the purpose, set retention, and check contracts and access rights.

Who stops the agent and who resolves a problematic purchase

Before launching, assign a purchasing supervisor and a backup. They must be able to suspend new orders and request the revocation of delegated credentials. If budget or approval checks fail, the system halts execution and hands the case over to a human. This behavior is called fail closed: if a valid check is missing, the action is blocked. OWASP explicitly recommends this rule for high-impact operations.

Stopping the agent does not automatically cancel an already accepted order or a completed payment. For those, check with the supplier and payment provider, then follow the applicable cancellation, return, or refund flow. During the pilot, clearly assign who handles each exception and by when it must be resolved.

We recommend tracking the purchase process through to completion:

  • Order: what the supplier accepted, which identifier it has, and whether there are partial deliveries.
  • Payment: what amount is only authorized or reserved and what has actually been collected, according to the provider’s status.
  • Receipt: what was received and whether the items, quantities, and services match the order.
  • Invoice and adjustments: whether the documents and amounts correspond, and any refunds have been actually confirmed.

This matching is called reconciliation. An agent’s message like “I handled the return” is not enough to close the case. Keep the confirmations from the supplier and payment provider, linked to the same purchase. For instance, Stripe documentation about tracking payments describes verifying the result based on notifications sent to the server. For a purchasing firm, available data may also come from the procurement platform, bank portal, or supplier confirmations.

Three examples, three levels of control

Recurring consumables with stable specifications. The agent can compare offers in an internal catalog and prepare a restock. After a human-approved pilot, the company may consider automating orders to validated suppliers, with an aggregate cap and receipt verification. Price, stock, or delivery exceptions go to a human.

Software licenses. The agent can suggest plans and estimate costs for the projected number of users, but software usage rights, data access, contract length, and recurring costs must be checked. Approval usually involves the IT manager or budget owner, not just the requester.

Equipment or services with negotiated terms. The agent can gather specifications and compare offers, but differences in warranty, service, integration, and contractual terms may outweigh the displayed price. Here, it’s prudent to keep commercial decisions and commitments with people.

These are design scenarios, not claims that a given merchant or processor currently offers the automation described.

A measurable pilot for an SME

  1. Pick a single, narrow process. For example, consumables with a stable catalog and known suppliers. Record who requests, approves, pays, and confirms receipt in the current workflow.
  2. Pilot without purchases first. The agent prepares recommendations and draft carts. Compare total price, availability, specifications, and time spent vs. human procurement.
  3. Codify the policy in the system. Define suppliers, items, caps, exceptions, approvers, and error behavior. Test modified carts, prices over the cap, contention between two requests, and malicious messages from pages.
  4. Activate a stage with approval. Keep approval tied to the final cart and use the agreed payment flow with the provider. Log proposal, checks, approval, order, payment, and result without copying unnecessary sensitive data into logs.
  5. Measure and scale only based on evidence. Track actual review time saved, total price, wrong or duplicate orders, exceptions, returns, and policy compliance. Even if a pilot shows no unauthorized purchases, that doesn’t prove risk is zero. Review controls whenever the supplier, model, tools, or policy changes.

In the end, the company needs to answer simply: who requested the purchase, who approved which option, what rule allowed it, what was paid, and what was received? If reconstructing the agent’s conversation is needed to answer, the process is still not sufficiently controlled.

Conclusion

Agents can reduce repetitive procurement work, especially in search, comparison, and order preparation. Purchase autonomy is granted progressively, depending on risk and what systems, suppliers, and payment providers can check. Budgeting, approval, and reconciliation must remain enforceable rules and clear records, not merely intentions phrased in natural language.

Sources and further reading

If you want to identify a suitable procurement process for a pilot and set autonomy boundaries, talk to the i8.ro team.

Recommended for you

The 90-Day Adoption Plan: From First Pilot to a Measurable Agentic System

How Much Is an AI Agent Worth: Total Cost, KPIs, and Return on Investment

De unde începem: șapte procese potrivite pentru agenți AI într-un IMM

Cookies

We use cookies required for the site to work. With your consent we also enable additional features (videos, maps) or anonymous statistics. You can change your choice at any time from the site footer.

Cookie policyPrivacy noticeTerms and conditionsCookie preferences