
# When the agent buys on our behalf: budgets, payments, and limits of autonomy
An AI agent can compare offers, fill a cart, and, in some integrations, initiate steps in the payment process. For a small or medium business, the relevant question isn’t whether the agent “knows how to buy,” but what decisions we can delegate, within what limits, and who checks the result. A wrong order could mean a double payment, an unwanted subscription, delivery to the wrong address, or using an unapproved supplier.
This article looks at the company as the buyer. For a seller’s perspective on selling to agents, see the article on selling software as a service (SaaS) to AI agents. For access control and action logging, see the article on permissions and audit. The budget and threshold examples below are working scenarios, not universal recommendations or legal caps.
As of September 2026, there are specifications and programs for agent-assisted transactions, but their availability depends on the provider, country, merchant, processor, and contract. For example, Google’s Agentic Payment Protocol (AP2) specification separates the mandate for the content of an order from the mandate for payment. Visa Intelligent Commerce documentation describes authenticated user instructions, limited payment credentials, and controls for merchant and amount. These are technical directions, not guarantees that any SME in Romania can activate all functions tomorrow.
Products also evolve. In its update from March 24, 2026, OpenAI explained that the initial Instant Checkout lacked needed flexibility, and is focusing on product discovery and merchant checkout experiences. By checkout we mean the final steps of purchasing—confirming cart, delivery, and payment method. So, before starting a project, check the actual flow available from your selected provider, not just the protocol’s name in presentations.
To decide what to automate, let’s separate four stages:
Good autonomy at stage 1 doesn’t automatically justify autonomy at stage 3. The risk increases when an action incurs spending or creates obligations that are hard to reverse.
The following scale is an editorial tool for design, not a legal standard:
For most SMEs, a pilot at levels 1 or 2 gives insight into recommendation quality and operations while keeping human approval for each purchase. The academic study τ-bench shows why rule compliance needs to be tested on concrete cases: evaluated agents had mixed results for tasks with domain policies and tools. The results relate to the benchmark scenarios—that is, the standardized test set—not a valid error rate for all commercial systems.
The instruction sent to the model, called a prompt, may include “do not spend over 300 lei.” This can guide the agent, but real control needs to be in a separate component that checks the transaction before the order or payment is placed. The agent proposes; the procurement system decides if the proposal fits company policy. This is an architectural choice, supported by risks noted in research about agent policy compliance and in the OWASP guide for AI agent security.
Policies can, where relevant, include:
Policy should be checked on structured data from the cart or offer, not on the agent’s free-form summary. Additionally, the budget should be reserved while the order is in process, so two simultaneous sessions can’t each spend the same available amount. After a confirmed result, the reservation becomes a registered expense or is released, as appropriate. If the result is unknown, the reservation stays until clarified, to avoid spending the same funds again. These are design recommendations, not functions every platform offers out of the box.
Hypothetical scenario: an office allocates 2,000 lei per month for consumables, allows up to 300 lei per order, and only accepts products from a catalog and validated suppliers. The agent finds items for 280 lei, but shipping brings the total to 315 lei. The order doesn’t go through automatically, even if the product price is under the threshold. If, after approval, the quantity, supplier, address, or total changes, the system requires a new check. Splitting the same purchase into two orders shouldn’t bypass the approval threshold. Figures above are only to illustrate the mechanism.
An “Approve purchase” button is ambiguous if the cart can change between approval and payment. The person approving should see at least the merchant, items and quantities, total cost and currency, delivery, possible recurring costs, and the approval’s validity period. The system saves the approved version and compares it to the submitted version. Any relevant change halts the transaction or sends the request back for approval.
This principle also appears in AP2, where the payment mandate is tied to a defined checkout. It's useful to distinguish between a commercial mandate, meaning the company's permission to buy under certain conditions, and payment authentication, performed by the payment service. Internal approval doesn’t replace the authentication required by the bank or processor.
In the EU, PSD2 Directive and Delegated Regulation (EU) 2018/389 set the framework for strong customer authentication in mandated cases. This uses at least two independent factors, such as something the user knows, possesses, and a biometric feature. Where strong authentication applies for remote electronic payment, Article 5 of the regulation requires the code to be tied to the agreed amount and payee—changing them makes the code invalid. There are exceptions under defined conditions, including for certain secure corporate processes, but you shouldn’t assume all business-to-business (B2B) payments are exempt. The actual flow needs validation with your payment service provider and, if needed, your legal advisers.
An agent receiving a card number, security code, or authentication code over chat creates unnecessary risk. A safer route is a payment integration through an authorized provider, with limited credentials or tokens. A payment token is a representation used in the payment flow instead of raw card data; its value is tied to the purpose and restrictions set by the provider. Visa describes agent-specific tokens and instruction checks before issuing credentials, while OpenAI describes in its specification a delegated payment request with cap and expiry. These are technical models, not guarantees that any processor or merchant will accept them.
Before integration, ask your provider: who holds sensitive data, what permissions does the token have, which merchants and caps can it be used for, when it expires, how it’s revoked, how disputes are handled, and what happens if payment succeeds but order confirmation is delayed. Control separately the identity of the employee requesting the purchase and the identity of the app accessing the API—the interface through which software systems share data and execute actions.
Hidden instructions in external sources. A product page, offer PDF, or supplier email could have text attempting to alter the agent’s behavior. This is called an indirect prompt injection: information the agent reads is wrongly treated as instructions. AgentDojo studied such attacks in environments with tools. The practical solution is to treat offers as untrusted data and not let them modify policy, accounts, or approvals.
Offers that are no longer valid.Price, stock, shipping, or the rate may change between comparison and checkout. Check your final cart again. If the total or any key term differs from the approved version, request a new decision.
Duplicate orders. A dropped connection can cause the agent to resubmit the request without knowing the first attempt succeeded. Use a unique request identifier and idempotency, meaning that repeating the same request does not trigger a second purchase. Stripe’s documentation explains this principle for its API requests. The key must be reused for the same operation, and your integration should respect the provider’s retention period and rules. Safeguarding a payment request does not automatically eliminate duplicate orders in the system. Always check the order and payment status independently; a missing response is not proof of failure.
Costs that continue after the first payment. A SaaS subscription may renew automatically, bill per user, or have variable consumption charges. A limit on the first transaction does not necessarily cap future spending. Such an agreement requires a separate rule for subscription duration, amendments, termination, and the account owner.
Exposure of personal and business data. Employee names, delivery addresses, and invoices may reach external services. The European Data Protection Board explains the principle of data minimization for small businesses. Provide the agent and suppliers only the data needed for the purpose, set retention, and check contracts and access rights.
Before launching, assign a purchasing supervisor and a backup. They must be able to suspend new orders and request the revocation of delegated credentials. If budget or approval checks fail, the system halts execution and hands the case over to a human. This behavior is called fail closed: if a valid check is missing, the action is blocked. OWASP explicitly recommends this rule for high-impact operations.
Stopping the agent does not automatically cancel an already accepted order or a completed payment. For those, check with the supplier and payment provider, then follow the applicable cancellation, return, or refund flow. During the pilot, clearly assign who handles each exception and by when it must be resolved.
We recommend tracking the purchase process through to completion:
This matching is called reconciliation. An agent’s message like “I handled the return” is not enough to close the case. Keep the confirmations from the supplier and payment provider, linked to the same purchase. For instance, Stripe documentation about tracking payments describes verifying the result based on notifications sent to the server. For a purchasing firm, available data may also come from the procurement platform, bank portal, or supplier confirmations.
Recurring consumables with stable specifications. The agent can compare offers in an internal catalog and prepare a restock. After a human-approved pilot, the company may consider automating orders to validated suppliers, with an aggregate cap and receipt verification. Price, stock, or delivery exceptions go to a human.
Software licenses. The agent can suggest plans and estimate costs for the projected number of users, but software usage rights, data access, contract length, and recurring costs must be checked. Approval usually involves the IT manager or budget owner, not just the requester.
Equipment or services with negotiated terms. The agent can gather specifications and compare offers, but differences in warranty, service, integration, and contractual terms may outweigh the displayed price. Here, it’s prudent to keep commercial decisions and commitments with people.
These are design scenarios, not claims that a given merchant or processor currently offers the automation described.
In the end, the company needs to answer simply: who requested the purchase, who approved which option, what rule allowed it, what was paid, and what was received? If reconstructing the agent’s conversation is needed to answer, the process is still not sufficiently controlled.
Agents can reduce repetitive procurement work, especially in search, comparison, and order preparation. Purchase autonomy is granted progressively, depending on risk and what systems, suppliers, and payment providers can check. Budgeting, approval, and reconciliation must remain enforceable rules and clear records, not merely intentions phrased in natural language.
If you want to identify a suitable procurement process for a pilot and set autonomy boundaries, talk to the i8.ro team.