
An AI agent can compare offers, fill a cart, and, in certain integrations, initiate steps in the checkout process. For a small or medium business, the useful question isn’t whether the agent “knows how to buy,” but rather what decisions we can delegate to it, within what limits, and who checks the results. A wrong order could mean double payment, an unwanted subscription, delivery to the wrong address, or sourcing from an unapproved supplier.
This article looks at the company in the role of buyer. For the perspective of stores selling to agents, see the article on selling software as a service (SaaS) to AI agents. For access control and activity logging, see the article on permissions and audit. The budget and threshold examples below are working scenarios, not universal recommendations or legal ceilings.
As of September 2026, there are specifications and programs for agent-assisted transactions, but their availability depends on provider, country, merchant, processor, and contract. For example, Google’s Agentic Payment Protocol (AP2) specification separates the mandate for order content from the mandate for payment. Visa Intelligent Commerce documentation describes authenticated user instructions, limited payment credentials, and controls for the merchant and amount. These descriptions indicate a technical direction, not that any SME in Romania can enable all features tomorrow.
Products are changing, too. In its March 24, 2026 update, OpenAI explained that the initial version of Instant Checkout didn’t offer the desired flexibility and that the focus is now on product discovery and merchant checkout experiences. By checkout we mean the final steps of purchasing, where the cart, delivery, and payment method are confirmed. So, before starting a project, check the actual flow available at your chosen provider, not just the protocol’s name in presentations.
To decide what to automate, we separate four stages:
Good autonomy at stage 1 does not automatically justify autonomy for stage 3. The risk increases when actions incur costs or difficult-to-undo obligations.
The following scale is an editorial tool for design, not a legal standard:
For most SMEs, a pilot at levels 1 or 2 provides data on the quality of recommendations and operations while retaining human approval for each purchase. The academic study τ-bench shows why compliance with rules must be tested in real-world cases: assessed agents had uneven results on tasks involving tools and domain policies. The outcome pertains to the benchmark scenarios—that is, to the standardized test set—not to an error rate valid for all commercial systems.
The instruction sent to the model, called a prompt, can include “do not spend more than 300 lei.” This may guide the agent, but real control must exist in a separate component that checks the transaction before order or payment is placed. The agent makes a proposal; the procurement system decides if it fits the company’s policy. This is an architectural choice also supported by risks observed in research on agents’ policy compliance and in the OWASP guide to AI agent security..
Depending on the case, policy may include:
The policy should be checked against structured data from the cart or offer, not just the agent’s freeform summary. Additionally, budget is reserved while the order is pending, so that two simultaneous sessions can't both spend the same available funds. Once a result is confirmed, the reservation becomes a recorded expense, or is released, as appropriate. If the result is unknown, the reservation remains until clarified, to avoid double-spending. These rules are design recommendations, not features that all platforms provide by default.
Hypothetical scenario: an office allocates 2,000 lei per month for supplies, allows a maximum of 300 lei per order, and accepts only products from a catalog and validated suppliers. The agent finds items for 280 lei, but shipping brings the total to 315 lei. The order does not go through automatically, even if the product price is below the threshold. If, after approval, the quantity, supplier, address, or total changes, the system requires a new verification. Splitting the same purchase into two orders must not circumvent the approval threshold. The figures are solely to illustrate the mechanism.
An 'Approve purchase' button is ambiguous if the cart can change between approval and payment. The approver should at least see the merchant, items and quantities, total cost and currency, delivery details, any recurring charges, and the expiry of the approval. The system must retain the approved version and compare it to the transmitted one. Any relevant change stops the transaction or sends it back for approval.
This principle also appears in AP2, where the payment mandate is linked to a defined checkout. It is useful to distinguish between the commercial mandate, that is, company permission to purchase under specific conditions, and payment authentication, performed by the payment service. An internal approval does not replace authentication required by the bank or payment processor.
In the EU, the PSD2 Directive and Delegated Regulation (EU) 2018/389 set the framework for strong customer authentication in situations required by law. This uses at least two independent elements, from categories such as something the user knows, something they possess, and a biometric characteristic. When strong authentication is applied to a remote electronic payment, article 5 of the regulation requires linking the code to the approved amount and beneficiary, and changes to either invalidate it. There are exceptions under defined conditions, including for certain secured corporate processes, but it's not to be assumed that any business-to-business (B2B) payment is exempt. The concrete workflow should be validated with your payment service provider and, if needed, your legal counsel.
An agent who receives the card number, security code, or authentication codes via chat creates an unnecessary risk. The safer option is a payment integration through an authorized provider that only gives limited credentials or tokens. A payment token is a representation used in the payment flow instead of the raw card details; its value depends on its purpose and the restrictions set by the provider. Visa describes agent-specific tokens and instruction checks before issuing credentials, while OpenAI describes in its specification a delegated payment request with a ceiling and expiration. These are technical models, not guarantees that any processor or merchant will accept them.
Before integrating, ask your provider: who owns the sensitive data, what permissions the token has, which merchant and ceiling it can be used with, when it expires, how it can be revoked, how disputes are handled, and what happens if payment succeeds but order confirmation is delayed. Separately control the identity of the employee requesting the purchase and the application calling the API—the interface through which software systems exchange data and perform operations.
Instructions hidden in external sources. A product page, offer PDF, or supplier email can contain text trying to manipulate the agent's behavior. This is called an indirect prompt injection: data read by the agent is mistakenly treated as instructions. AgentDojo has studied this type of attack in tool-enabled environments. The practical solution is to treat offers as untrusted data and not allow them to change policy, accounts, or approvals.
Offers that are no longer valid. Price, stock, shipping, or exchange rate can change between comparison and checkout. Check the final cart again. If the total or an essential term differs from the approved version, ask for a new decision.
Duplicate orders. An interrupted connection may lead the agent to resubmit the request without knowing the first one succeeded. Use a unique request identifier and idempotency, meaning repeating the same request does not produce a second purchase. Stripe's documentation explains the principle for its API requests. The key should be reused for the same operation, and the integration must respect the provider’s retention period and rules. Protecting a payment request does not automatically remove duplicates from the order system. Then independently check the order and payment status; a missing response is not proof of failure.
Costs that continue after the first payment. A SaaS subscription might have automatic renewal, per-user pricing, or variable usage charges. A ceiling on the initial transaction does not necessarily limit future expenses. Such a contract requires a separate rule for subscription duration, changes, termination, and account holder.
Exposed personal and company data. Employee names, shipping addresses, and invoices can end up with external services. The European Data Protection Board explains the data minimization principle for small businesses. Only give agents and providers the data needed for the intended purpose, set retention period, and check contracts and access rights.
Before launch, designate a responsible person for purchases and a backup. They must be able to suspend new orders and request revocation of delegated credentials. If budget or approval checks fail, the system halts execution and escalates the case to a human. This behavior is called fail closed: if a valid check is missing, the action is blocked. OWASP explicitly recommends this rule for high-impact operations.
Stopping the agent doesn’t automatically cancel an accepted order or an executed payment. For these, check the supplier and payment provider status, then follow the applicable cancellation, return, or refund process. In the pilot, clearly define who handles each exception and by when it must be resolved.
We recommend the process tracks the purchase until closure:
This matching process is called reconciliation. An agent’s message such as “the return has been resolved” is not enough to close the case. Keep the supplier’s and payment provider’s confirmations, linked to the same purchase. For example, Stripe documentation on payment tracking describes checking the outcome based on notifications sent to the server. For a buying company, available data may also come from the procurement platform, bank portal, or supplier confirmations.
Recurring consumables with stable specifications. The agent can compare offers from an internal catalog and prepare a restock. After a human-approved pilot, the company can evaluate automated orders to validated suppliers, with an aggregate ceiling and verification upon receipt. Exceptions for price, stock, or delivery are escalated to a human.
Software licenses. The agent can suggest plans and simulate the cost for the estimated number of users, but the rights of use, data access, contract length, and recurring fees must be checked. Typically, approval also involves the IT lead or the person managing the budget, not just the requester.
Equipment or service with negotiated terms. The agent can gather specifications and compare offers, but differences in warranty, service, integration, and contract terms may matter more than the displayed price. Here it is wise to keep the decision and commercial commitment with people.
These are design scenarios, not statements that a particular merchant or processor currently offers the described automation.
In the end, the company should be able to answer simply: who requested the purchase, who approved which option, what rule allowed it, what amount was paid, and what was received? If the answer requires reconstructing the agent’s conversation, the process is still not sufficiently controlled.
Agents can reduce repetitive procurement work, especially for searching, comparison, and order preparation. Buying autonomy should be granted gradually, according to risk and what systems, suppliers, and payment providers can verify. Budget, approval, and reconciliation must remain executable rules and clear records—not just intentions phrased in natural language.
Do you want to identify a suitable procurement process for a pilot and define its autonomy boundaries?