Skip to content

Build, buy, or integrate? How to choose your AI agent platform

24.09.2026

A company wants to respond faster to customer requests. It already has email, a customer relationship management (CRM) system, and a product catalog. A demo with an AI agent can summarize a message and draft an offer in minutes. But in production, tough questions arise: can it see only the right clients? Who approves a discount? What happens if the sending fails? Where is the proof kept? And who fixes the integration when the CRM API changes?

That’s why choosing an AI agent platform doesn’t start with a list of models or the flashiest interface. It starts with the specific process, the systems it must interact with, and who is accountable for the final result. Based on the answers, we can either buy a ready-made feature, configure a visual platform, use a managed cloud service, or develop an agent within your own app. These options can be combined.

This article offers a decision framework for SMEs. It explains exactly what you’re buying in each scenario, how to compare real costs, which tests are worth requesting from vendors, and when a basic automation flow is enough.

Checked as of September 27, 2026. Platform examples reflect available documentation at this date; features, regions, and contracts may change. The numbers in the calculation below are pilot estimates, not market statistics.

The short answer

  • If the activity is clearly defined and the product you use already offers the needed feature, test that feature first. Check permissions, limits, and data export.
  • If the process involves just a few applications and the rules can be explicitly mapped out, an automation flow with limited AI steps is often the easiest to control.
  • If you need more tools, state management between steps, approvals, and integration with your existing cloud infrastructure, look at a managed platform or a development framework.
  • If your process logic is what differentiates your product, the rules are specific, and you have a team that can operate the software, in-house development may be worth the investment.
  • Whatever you choose, keep action authorization in verifiable systems, test with real company scenarios, and calculate the total cost—including oversight and error correction.

This is a guiding rule, not a universal product recommendation.

What do we actually mean by “agent platform”?

An AI agent combines a model that interprets a request with tools that read data or execute actions. A flow pre-defines the order of steps. In practice, many products mix the two: the model chooses among several tools, but a deterministic flow limits when and how they may be used.

There are five key layers worth separating:

  1. Model: the component that interprets and generates content. Its selection has its own criteria, discussed in the guide to selecting a large language model (LLM).
  2. Orchestration: rules that decide the steps, keep task state, resume executions, and escalate tricky cases to a human.
  3. Tools and data: connections to CRM, email, documents, databases, and APIs. Model Context Protocol (MCP) may standardize tool exposure, but does not replace authentication, authorization, or business logic.
  4. Controls: permissions, approvals, isolation, volume limits, logs, and incident response. These are detailed further in the article on permissions and audit.
  5. Operations: hosting, updates, monitoring, backup, support, versions, and costs.

A code library may solve orchestration without providing a full admin application. A hosting service may run code, but cannot define the company’s business policy. A software product delivered as SaaS may provide a great interface, but allow few changes to flow logic. When comparing offerings, make sure to explicitly ask which of these layers are included.

First filter: do you need an agent?

To extract a field from a stable form and save it to a system, a regular rule might be cheaper and more predictable. For classifying varied messages, summarizing documents, or suggesting a reply, a model can add value. For actually sending the response or changing a price, keep a deterministic validation step, and where needed, require human approval.

A simple test: if you can fully describe all steps and branches as fixed rules, start with a flow. If input data is highly variable and the next step depends on interpreting it, add AI only where the flexibility is worth the cost and risk.

Four implementation paths

| Path | What you get quickly | What remains company responsibility | When it makes sense |
|---|---|---|---|
| Agent feature in an existing product | interface, native integration, simplified admin | checking access, contract, quality & limits | process close to standard product features |
| Visual platforms (flows or agents) | connectors and fast process building | business rules, credentials, versions, tests & monitoring | a few systems, explicit steps, modest tech skills in team |
| Managed cloud service | infrastructure and operations managed by provider | agent design, data integration, controlling actions & cost | works with existing cloud, clear scale/governance reqs |
| Framework & in-house development | high flexibility over flows & interface | implementation, security, hosting, updates, support | unique logic, complex processes, team can operate the product |

Integration is not a fifth exclusive category. A company might buy a work interface, integrate flows into an existing system, and only develop custom tools for sensitive operations. Most of the time, the right decision focuses on what we keep under our own control, not on what we build end-to-end.

Separate from those four paths, you’ll need to decide where each component runs: orchestration, the model, document storage, and logs. A local solution can still call an external model, while a hybrid setup might combine an internal flow with cloud services. Map the data route and check separately the costs, latency, access, and contractual obligations for every component.

Buy features already present

Example: your team already uses a customer support tool that can suggest replies based on help articles. A key benefit is integration with existing users and tickets. However, check if the product enforces each operator’s permissions, shows answer sources, how it handles attachments, and whether you can export relevant history. Don’t pay for a separate platform before testing the feature with real scenarios.

Configure a visual flow

Tools like n8n and Dify let you compose AI flows and apps. n8n calls itself fair-code: the code is visible, but its license restricts certain commercial uses. Check the license terms for your case; don’t assume it’s a permissive open-source license. Dify’s docs include installation via Docker Compose, allowing a self-hosting option. But self-hosting means you must manage updates, dependencies, and backups. If the flow calls an external model, data sent to that model doesn’t automatically stay within your local infrastructure.

These products can speed up pilot development. For sensitive operations, check where credentials are stored, how action approvals work, what history can be exported, and what happens if an execution is resumed.

Use a managed service

As of this writing, Microsoft Foundry documents agent hosting, tools, observability and evaluations; the Gemini Enterprise Agent Platform brings together Google services for developing and running agents, including components previously branded as Vertex AI Agent Builder; Amazon Bedrock AgentCore documents components for the agent runtime, tools, identity, policy, and observability. These are service families, not identical offers. Check region, feature status, limits, extra required services, and your own contract’s costs.

One current note: AWS’s docs say that Bedrock Agents Classic is no longer open to new customers. A guide or tutorial on the Classic variant may remain technically useful, but should not be treated as a procurement path for new clients.

Develop on a framework

LangGraph offers mechanisms for stateful flows, interruptions, and resumptions; OpenAI Agents SDK documents orchestrating in your app, tools, and approvals. These are for developers, not automated security guarantees for your process. Storage, access policies, testing, UX, and operations may all remain the team’s responsibility.

A framework is a fit when you already have your own apps and want the agent to work through narrow, testable APIs. Flexibility comes with the obligation to maintain code, dependencies, and connector compatibility.

Three processes, three starting points

Support for an online store. The system receives questions around shipping and returns. If your current support platform can search the updated policy and propose a sourced reply, test that feature. The agent may draft, but a human double-checks and sends over the first weeks. Track unresolved issues and cases where the return policy doesn’t cover the scenario. A full in-house project is hard to justify until the existing product’s limitations are clear.

Offers to other companies (B2B) from e-mail, management system (ERP), and CRM. E-mails come in various formats, but prices, stock levels, and discounts are in dedicated systems. A visual workflow or a lightweight agent developed within the application can extract the data, prompt for missing information, consult APIs, and create a draft. Price calculation should remain in the ERP or in a service with explicit rules, and any external sending should require proper approval. The main cost may be proper integration, not the subscription to the model.

SaaS product that provides automation for its clients. Here, each client’s identity, data isolation, and in-product experience are part of the commercial offer. An integrated software framework within the app or a managed cloud runtime might be preferable to a separate visual workflow, since the team needs control over the interface, versioning, and organization-level access. Even in this case, the team may purchase components for hosting, evaluation, or observability, and only build the product-specific logic.

These are design scenarios, not guaranteed outcomes. Volume, data quality, and existing infrastructure can change the decision.

How to choose based on process, not just commercial promise

Write a one-page brief for the pilot process:

  • Objective: what measurable result should be delivered? For example, a correct offer draft, not simply “an intelligent assistant.”
  • Inputs: what types of e-mails and documents occur? In which languages? How often is data missing?
  • Systems: what are the official sources for price, stock, contracts, and client identity?
  • Actions: what can the agent read, propose, modify, publish, send, or delete?
  • Exceptions: what happens with conflicting data, response timeouts, duplicates, or accounts without rights?
  • Accountability: who approves, who is alerted, and who can stop the workflow?
  • Volume and result: how many cases per month, and how is quality measured?

If two teams interpret the same commercial rule differently, buying a platform won’t solve that ambiguity. First, stabilize the process and the source of truth.

Criterion 1: data and access rights

Ask whether the platform can access documents while respecting each user’s rights, separate clients or projects, and delete data from indexes, logs, and backups in line with applicable policies. A knowledge base or a retrieval-augmented generation (RAG) system must account for authorization at the time of reading, not just when documents are uploaded.

Check the data flow: browser to platform, platform to model provider, to monitoring system, and to any subcontractors. “EU region,” “on-premises installation,” and “data not used for training” are different statements. Request the documents and settings supporting each claim.

Criterion 2: tools and control over actions

A connector to CRM isn’t enough. The system must be able to “read client X” without being able to “export all clients,” validate prices before issuing offers, and require approval before sending. For a tool accessed via MCP or API, enforcement must happen before the operation and be confirmed by the destination system. OWASP Agentic Top 10 2026 offers a starting point for threat evaluation.

Concrete tests to request: is an unauthorized action refused? Does an expired approval become invalid? Can the agent change the recipient after approval? Does a faulty tool halt the operation instead of producing incomplete results?

Criterion 3: flow reliability

A ten-minute demo won’t show how tasks are resumed after errors. Ask for evidence of state retention, resumption, timeouts, queues, a unique operation identifier (idempotency key) and a rollback process. The identifier helps the system recognize repeated attempts of the same operation. An idempotent operation doesn’t produce, for example, two orders or two payments when the same request is retried. Without this safeguard, automatic retries can double the effect.

The τ-bench study tests agents in interactions with users, domain rules, and tools, and measures consistency under repeated attempts. The results of that benchmark do not predict your implementation’s performance. They justify repeatedly testing your own process, verifying the actual final state in applications, not just the model’s output.

Criterion 4: transparency and operations

Can you see for a request what data was accessed, which tools were called, what policy was applied, who approved, and what actually changed? Are there alerts for errors, abnormal costs, and access denials? Are test and production environments kept separate? Can logs be exported without unnecessarily copying personal data?

Technical observability isn’t the same as a business audit. A model call trace doesn’t prove the price was approved. Responsible individuals must be able to reconstruct the request path all the way to its effect in the destination system.

Criterion 5: team skills and time

A product can be configured in a day, but who will fix a connector in six months? An in-house solution may be flexible but relies on documentation, tests, and people who know the code. Record the hours available for design, operations, and support. A vendor with support and clear conditions might be cheaper than a poorly maintained in-house solution; for other cases, in-house control over the code might justify the investment.

Criterion 6: exit cost

Ask what can be exported: data, documents and metadata, flow definitions, prompts, rules, logs, configuration, and tests. Importing into another platform can require significant work even if JSON export is available. A common protocol for tools reduces some repeated integration but doesn’t guarantee portability of agent state, policies, or interface.

A good test is to ask the vendor to export the pilot and document the steps to rebuild it in a new environment.

What to include in the contract and operating plan

Document who owns each connector and who is notified if it fails. Set response times and incident procedures for provider-managed components, as well as who backs up data and who tests restores. Request a description of price changes, usage limits, and how incompatible API changes are communicated. If the vendor only provides the agent infrastructure, your team remains responsible for process logic and proprietary connectors.

An availability requirement expressed via a service-level agreement (SLA) must be considered along with exclusions and dependencies. An agent may run on a highly available service but won’t complete offers if the ERP or model is unresponsive. Define degraded behavior: delayed response, draft to operator, or fallback to manual process. Test this behavior before launch.

Privacy and compliance: contract questions

Before uploading real data, identify categories of personal data, processing purposes, roles of each party, sub-processors, retention periods, access control, and deletion. Check the data processor agreement, if applicable, and review cross-border transfer paths. The European Commission explains the mechanisms and safeguards for transfers outside the European Economic Area. A report drafted by external experts for the EDPB program proposes a methodology for privacy risk assessment in systems with language models; the report does not necessarily reflect the EDPB’s official position.

Avoid concluding that a platform is compliant just because it’s hosted in the EU. Execution logs, support services, model providers, and connected tools may follow different paths. Legal assessment depends on specific data and processes; this article does not determine the legal status of a product.

Calculate total cost, not just the subscription

A meaningful calculation is implementation cost + monthly operating cost + cost of interventions and errors + cost of platform switch. At minimum, include:

  • workflow and connector setup, data migration, and authentication integration;
  • license or subscription, hosting, storage and backups;
  • model calls, document preparation for search (indexing), search and re-ranking results, external tools;
  • monitoring, logging, regression testing, and security review;
  • approval time, result correction, and support for exceptions;
  • incidents, upgrades, and the cost of replacing a provider.

Not all of these elements are billed separately. A managed platform may include hosting but separate charges for the model or storage. A self-hosted solution may avoid a subscription yet carry substantial administrative costs. Compare on the same basis: total cost per correctly resolved request, at the same volume and control level.

An indicative calculation, with explicit assumptions

Suppose you have 1,000 requests per month, each requiring five minutes of manual work. That’s initially around 83 hours. In a hypothetical pilot, the agent prepares 700 cases for review; the other 300 are handled fully by humans. If all 700 are correct but still require two minutes of verification, the time saved is about 35 hours, before factoring in support, exceptions, and maintenance: 700 × (5 - 2) minutes / 60.

This calculation does not prove a positive return. Platform, call, and implementation costs are missing, and the quality of the 700 results must be measured. If review finds 20 errors each requiring 15 minutes to fix, another five hours are lost. If reviews take four minutes, the gross benefit drops below 12 hours even before errors. That’s why an impressive demonstration can’t replace proper measurement.

A selection matrix that doesn’t conceal risk

Start with eliminating conditions, then compare the remaining solutions. Possible conditions: acceptable data contract, sufficient permissions, ability to stop, export, maximum budget, mandatory system integration, and regional availability. If a solution fails a required condition, don’t rescue it because of a great interface score.

For the rest, you can use this sample matrix. The weights are working examples, totaling 100% and should be adapted for your process:

| Criterion | Example Weight | Evidence to Request in Pilot |
|---|---:|---|
| Result quality on internal cases | 25% | correct cases, errors, verifiable explanations |
| Permissions, data, and security | 20% | refusal, approval, deletion tests, data flow |
| Integration and Reliability | 15% | real-world system actions, replay without duplicates |
| Total Cost at Estimated Volume | 15% | cost simulation plus team hours |
| Operation and Support | 15% | logs, alerts, versions, incident owner |
| Portability | 10% | documented export and reconstruction |

Give each criterion a score from 0 to 5 and multiply it by the weight. For example, a score of 4/5 on a criterion with 20% weight contributes 0.8 points to the maximum total of 5.Don’t turn the score into an objective truth: two solutions might have identical totals but different risk profiles. Note evidence, assumptions, and unchecked issues. For an agent authorized to pay invoices, security and action control would likely carry more weight than for one drafting proposals.

Four-week Pilot

Week 1: lay the groundwork

Choose a single process and describe its current state: working time, error rate, cost, and service level. Prepare a set of anonymized or synthetic real cases, including rare scenarios: incomplete message, duplicate customer, expired price, wrong attachment, user without permission. Determine who judges whether a result is correct, acceptance thresholds, and situations that stop the pilot.

Week 2: build the minimum useful version

Connect only strictly necessary data and tools. Limit the agent to read/write or to a test environment. Place authorization rules in the tool or control service, not just in the prompt. Define the maximum budget and minimum required logging.

Week 3: test repeatedly and adversarially

Run the same set of cases multiple times. Measure the response and final state in the CRM or target system. Test documents and messages that contain hostile instructions for the agent, following the studied risk model in AgentDojo. Do not use actual sensitive data in uncontrolled tests.

Week 4: decide based on data

Compare two viable options on the same set, at the same volume. Check resolution rate, impactful errors, review time, latency, cost per resolved case, log quality, and recovery time after failure. Suspend testing if there’s unauthorized access, duplicate actions, or data leaks, and fix the root cause before resuming. The process owner and, as needed, IT, security, and data protection teams validate the launch. Decide whether to launch with approvals, remain in proposal mode, or revert to a non-autonomous flow.

An acceptance threshold must be set before testing begins, for the specific process. There is no universal success rate that makes an agent safe for any company.

Questions worth asking in a commercial demo

  1. Can you demonstrate with a user lacking permissions that the agent cannot access a customer belonging to another team?
  2. Where do the request content, attachments, conversations, and logs end up? What can be deleted?
  3. How do we approve a specific action, and how is approval invalidated if the data changes?
  4. What happens if the CRM API responds after a timeout, and the platform retries?
  5. Can we inspect the model, instruction, tool, and rule versions for an incident?
  6. Can we use another model or provider without rebuilding all tools? What are the limitations?
  7. How do we export flow definitions and test evidence, not just generated answers?
  8. What is the cost at our volume, with all auxiliary services included?
  9. Which features are currently available in our region and plan, which are in preview, and which require separate integration?
  10. Who is contractually responsible for support, incidents, and API changes?

Ask to see these answers demonstrated in the pilot. A feature list in the presentation shows what’s possible, but doesn’t prove behavior in your process.

Common mistakes

We choose based on the number of models or connectors. What matters is whether the connector respects permissions and can precisely perform the required operation, with errors handled correctly.

We build an agent for a process that’s still unclear. The model will try to resolve misunderstandings between teams, making result evaluation impossible.

We treat self-hosting as synonymous with exclusively local data. Models, telemetry, connectors, and support must each be inventoried.

We let the agent use an admin account. A system instruction does not reduce permissions for an overly privileged token.

We calculate only time saved on drafting. Revision, errors, integration, and support can completely change the economic result.

We confuse a protocol with an exit strategy. MCP enables some connections, but migrating state, policies, and approvals remains a separate project.

What is stable now, and what is still unsolved

By 2026 there are already documented options for visual flows, development frameworks, and managed cloud services. Principles such as least privilege, pre-approval for sensitive actions, observability, and evaluation on your own cases are well documented. NIST AI RMF and the profile for generative AI offer a voluntary framework for risk management, without prescribing any vendor.

Still variable, however, are feature availability across regions and plans, interoperability between platforms, agent quality for long processes, and the actual cost of oversight. Research on tool interactions and attacks shows important limitations but does not provide a universal ranking for your company. Data from your pilot is more valuable than a general promise of autonomy.

Scenarios for 2027, not predictions

It’s plausible that vendors will offer more standard components for agent identity, tool-applied policies, continuous evaluation, and export of traces. Some of the flows now built manually may be handled by products companies already use. Just as likely, differences in state, approval, and log formats may keep migration costs high. Plan contracts and architectures that accommodate change, without budgeting based on a feature promised for next year.

Main sources and useful reading

Documentation and standards: NIST AI 600-1, OWASP Top 10 for Agentic Applications 2026, European Commission on international data transfers, external expert report from the EDPB program, MCP, specification from July 28, 2026, and OpenTelemetry conventions for generative AI.

Research: τ-bench for evaluating tool and domain rule interaction, and AgentDojo for testing attacks via untrusted data.

Product documentation as of September 27, 2026: n8n, Dify, LangGraph, OpenAI Agents SDK, Microsoft Foundry, Gemini Enterprise Agent Platform and naming history, Amazon Bedrock AgentCore and AWS note on Bedrock Agents Classic.

Conclusion

The right platform is the one that solves a well-defined process at a justified cost, with controlled access and a capable team to operate it. For an SME, a sound initial decision might be to buy an existing function, integrate two apps via a clear flow, and custom-build only where specific rules demand special control. Another company may need an agent built into its own app. Testing with real data, permissions, and outcomes makes the difference.

If you want help choosing a pilot process, comparing platforms for your company’s infrastructure, and estimating total cost before implementation, reach out to the Imagine Infinity team. Together we can define what’s worth buying, integrating, and building, with clear criteria for release decisions.

Recommended for you

De unde începem: șapte procese potrivite pentru agenți AI într-un IMM

Permissions, Approvals, and Auditing: The Rules of a Trustworthy AI Agent

Prompt Injection and Personal Data: How to Use AI Agents Safely