
A 90-day program can turn a vague idea about AI agents into an evidence-based business decision. However, it cannot guarantee that every process will be automated, that integration will be simple, or that the system will be ready to scale. For an SME, the healthy outcome of this period might be a controlled launch, narrowing the use case, redesigning the solution, or even shutting it down.
This article proposes a practical framework for the first 90 days. The timeline is a working scenario, not an official standard or a guarantee of implementation. The actual duration depends on data, integrations, vendors, contractual requirements, and risk appetite. In regulated domains or for systems classified as “high risk”, legal review, assessment, and approvals may take longer.
An AI agent is not just a chatbot. Operationally, it is a system that receives an objective, interprets context, chooses steps, and can use tools such as the CRM, email, the knowledge base, or the invoicing app.
A measurable agentic system has at least seven components:
“Measurable” does not just mean counting API calls or conversations. It means being able to answer business questions: How many correct results were achieved? How much human time was saved? What was the cost per accepted result? What errors occurred? And is the residual risk acceptable?
Don’t start the clock just because you have an AI model subscription. Before the project begins, confirm five conditions.
“The team is wasting time” is an observation, not a use case. A better case would be: “employees manually check incoming requests, search for data in two applications, and prepare a response for approval.” The process must be described concretely enough that the same team can recognize a correct and an incorrect execution.
The process owner defines what a good outcome is, validates exceptions, and can halt the pilot. The technical provider cannot take on this business responsibility.
Measure the current state before automation: volume, processing time, cost, correction rate, delays, complaints, and incidents. If you don’t have historical data, the first phase must include manual observation of a representative sample.
Check who owns the data, what personal or confidential information appears, where it is processed, how long it’s kept, and whether vendors use it for training. Decide who can approve access to each tool.
A good first pilot is frequent, repeatable, valuable, and reversible. Avoid starting with decisions that are hard to dispute, large payments, layoffs, medical or legal evaluations, or actions that cannot be undone. For ideas about processes and selection criteria, also see 7 processes where an AI agent can deliver measurable results.
If the owner, data, or baseline are missing, the realistic goal for the 90 days is discovery and preparation—not production.
The first stage produces a one-page project brief and a map of the current process.
The brief should include:
Then chart the actual workflow, including exceptions. Observe a few cases from start to finish and talk to the people who handle them. Written procedures often omit shortcuts, informal checks, and rare situations that matter exactly when a system starts to act.
At this stage, establish the minimal alternative as well. Sometimes, it’s safer to solve the problem with a rule, a better form, a regular integration, or by having an assistant propose but not execute a response. The agent should be compared to these options—not just to manual work.
Decision Gate 1: continue only if the process is sufficiently clear, the outcome can be verified, data is accessible, and risk can be limited.
Now define the minimal architecture. For the first pilot, reduce the number of tools and privileges. If the agent only needs to check inventory and draft a response, don’t give it rights to change prices or issue refunds.
For each tool, note:
Apply the principle of least privilege and separate reading from writing. Dedicated accounts, permission lists, rate limits, and short-lived credentials all reduce the impact of mistakes. For a full control model, see Permissions, Identity, and Audit for AI Agents.
Build a set of cases before optimizing prompts. Include:
Attacks through hidden instructions in emails, documents, or web pages are a practical risk for agents that read unverified data and use tools. The AgentDojo research, published at NeurIPS 2024, evaluates this type of interaction via 97 tasks and 629 security tests. The practical takeaway for an SME is not that there’s a universal test, but that success on ordinary cases does not prove resistance to adversarial content.
Set thresholds before seeing results. For example:
The above are just examples. Actual thresholds depend on error cost, the baseline, and risk tolerance.
Decision Gate 2: do not build the pilot until permissions, test cases, thresholds, and ownership for every decision are explicit.
Implement the smallest workflow that can prove or disprove the hypothesis. Use a separate test environment, masked or synthetic data where possible, and simulated tools for impactful actions.
During this period, the team must be able to see both the final result and the relevant path: which data the agent accessed, what tool it called, what rules applied, where it asked for help, and why it failed. The trace should not unnecessarily expose personal data, nor should it be confused with the model’s internal explanation. Its purpose is operational audit.
Compare three variants:
If the first or the second variant delivers nearly the same value with less risk and cost, the simpler choice is usually better.
Don’t add new features with every demo. Temporarily lock the model version, prompts, and tool configuration, then log any change. Otherwise, you won’t know if an improvement or degradation is due to the product, data, or model.
Decision Gate 3: the prototype only advances if it meets minimum thresholds in the isolated environment and no critical failures are left without a compensating control.
Offline evaluation uses historical, synthetic, and specially constructed error cases. Don’t just measure the final response. Also check whether the agent used the right tool, in the allowed sequence, with correct arguments, and without forbidden steps.
This separation between result and trajectory is important. An apparently correct result can hide an unauthorized source, a double entry, or a shortcut unacceptable in production. The technical guide Demystifying evals for AI agents, published by Anthropic on January 9, 2026, recommends combining output evaluation, trajectory review, and human analysis. This is a vendor recommendation, not a legal requirement, but the distinction is useful regardless of the model.
After offline tests, use the parallel or “shadow mode.” The agent receives real cases and produces a recommendation, but it doesn't alter systems or communicate directly with the client. The operator resolves the case through the normal flow, and the evaluator compares the results.
Record:
A good average score can hide a category that consistently fails. Segment outcomes by case type and separately track rare, high-impact events.
Decision gate 4: the real-impact pilot starts only if parallel mode data confirms thresholds, the team can intervene, and the stop mechanism has been tested.
At this stage, the agent can act on a small sample with clear limits. For example, a single request type, an internal group, a product category, or a reduced value cap.
Use controls such as:
Idempotency means that repeating the same request does not produce the business effect twice. This is essential when a connection times out and the agent isn’t sure whether the first attempt succeeded.
Decide in advance who can stop the pilot and what happens after. The stop shouldn’t depend on the same model or service that failed.
Review incidents and rejected cases daily during the first week. Reduce frequency only once behavior stabilizes. Don’t change the model, instructions, data sources, and access rights simultaneously, as you’ll lose track of the root cause.
Decision gate 5: the pilot continues only if observed value remains positive after including review time and incidents are within agreed limits.
The final stage is not a celebratory demo, but a comparison to the starting point.
Include the cost of the model, infrastructure, integrations, licenses, evaluation, human review, incidents, and maintenance. For a complete methodology, refer to How much is an AI agent worth: total cost, KPIs, and return on investment.
A useful operational formula is:
Cost per accepted result = total period cost / number of accepted results
Compare it to the previous process, but keep quality and risk separate. A lower cost doesn’t justify an unacceptable error rate or increased legal exposure.
Stopping is a valid pilot outcome. It prevents scaling a solution that adds no value.
If you extend, document:
No large department required, but roles need to be visible. The same person can fill multiple roles, provided accountability isn’t lost.
In small firms, role conflicts deserve attention. The person building the flow shouldn’t be the only one declaring it safe and profitable.
NIST AI Risk Management Framework is voluntary and structures activities into four functions: Govern, Map, Measure, and Manage. For a pilot, the practical translation is simple: set responsibilities, understand context and risks, measure behavior, then manage risk throughout operation. NIST Profile for Generative Artificial Intelligence, published July 26, 2024, supplements the framework with risks specific to generative systems. These are guides, not certifications, and do not replace legal obligations.
In the European Union, the Artificial Intelligence Act, Regulation (EU) 2024/1689, entered into force on August 1, 2024 and generally becomes applicable on August 2, 2026, with exceptions and specific timelines. AI literacy obligations apply from February 2, 2025. As of this writing, a company must check its role in the chain, concrete purpose, risk category, and related law, including data protection, consumer rights, labor law, and sector rules. Classification shouldn’t be inferred just from the product’s commercial label.
For team training, combine theory with internal procedure: which tools are approved, what data can’t be entered, how an outcome is checked, who approves an action, and how an incident is reported. AI literacy isn’t just ticking off a generic course.
A compact dashboard should separate four dimensions.
“Zero observed incidents” does not mean “zero risk.” Also report test coverage, sample volume, and categories not yet encountered.
The agent amplifies ambiguity. Clarify the rules and decision ownership first.
The product should be chosen based on requirements, data, integrations, and risk. For criteria, see How to Choose an AI Platform for Your Company.
A demo is cherry-picked. A pilot should include routine, failed, and edge cases.
Setting thresholds after seeing scores leads to confirmation bias. Define them in advance and keep a holdout set for calibration.
Start with read, simulation, and human approval. Expand rights only with supporting evidence.
The number of messages, steps, or tokens is not a business outcome. Measure accepted results, time, cost, quality, and risk.
Models, prices, tools, and data change. Any material change must be recorded and may require reevaluation.
Regardless of the scaling decision, the project should leave behind:
This is the difference between an experiment and an operational capability. The experiment shows that the agent can produce a result. The capability shows when it produces it, at what cost, under what limits, and who is accountable when it doesn’t.
A good plan doesn’t start with the model name but with the process, thresholds, and accountability. If you want us to help select the right use case, define evaluation, and build a controlled pilot for your company, talk to the i8 team.