An AI agent can prepare responses in seconds, review documents, and execute steps in a process. For a small or medium business, however, the speed shown in a demo doesn’t reveal the true value of the investment. What matters are which outcomes are accepted, how much work remains for people, and what actual costs or revenues are affected.
Saving ten hours a week can be valuable even if it doesn’t immediately reduce any invoices. It might allow the team to respond faster or handle more requests. But for an investment decision, it’s important to distinguish freed-up capacity from cash savings and extra margin.
In the article on choosing the platform, we discussed implementation options and associated costs. Here, we build a framework for measuring value: the initial benchmark, total cost, performance indicators, and a complete calculation of profitability.
Sources verified as of September 27, 2026. Percentages cited from studies refer to the specific test environments. Figures used in the financial example are hypothetical and separate from research results.
TCO, Total Cost of Ownership, is the full cost of owning and operating the solution over a defined period. For an agent, this may include implementation, integration, subscriptions, usage, infrastructure, verification, maintenance, and supplier switching. The time frame is selected before comparing solutions: a high upfront cost looks very different over three months versus three years.
KPI, Key Performance Indicator, means a key performance metric. A useful KPI shows whether the process is delivering the desired result: for example, correctly solved requests, minutes of human labor, or cost per accepted case. The number of messages generated may describe activity but not measure its value.
ROI, Return on Investment, expresses the ratio between net benefit and the reference cost. Since various conventions are used in practice, it’s crucial to specify the time period, which benefits are included, and the denominator. In our example, we’ll calculate cash ROI for the first year in relation to all cash costs of the project in that year.
Payback period shows when cumulative net benefits cover the initial investment. On its own, it doesn’t reveal how good the investment will be after that point and doesn’t substitute for a quality or risk analysis.
“The agent is done” isn’t a sufficient definition of success. The outcome should be described in terms of the company’s process.
Take a distributor who receives product inquiries. The agent classifies the message, searches the catalog and approved commercial terms for information, drafts a response, and sends it to a colleague for validation. It doesn’t change prices or promise delivery terms the system hasn’t confirmed.
For this process, an accepted result might mean the product is correctly identified, information is up-to-date, the reply follows company policy, the request is handled on time, and required actions are logged. A fluent message with the wrong price does not pass the test; nor is a case automatically closed without actually being resolved a success.
We separate accepted outcomes that need no rework from those corrected by people. Both can have value but incur different costs. Routing an exception to the right person may be the agent’s correct behavior; even so, the dashboard must differentiate this from final case resolution.
This focus on outcomes aligns with the FinOps Foundation, which distinguishes unit technical costs, like cost per token, from business-related costs, like cost per case resolved. A token is a unit of text or other content processed by the model. It is useful for billing and optimization but does not directly represent a resolved request.
The initial benchmark, also called a baseline, describes how the process works without the new solution. For the chosen period, we record volume, case complexity, staff active time, rework, errors, and costs. Difficult or incomplete cases are included, not just examples that look good in a demo.
It’s helpful to separate two durations. Active time shows how much time a person is actually working on a case. Time to closure includes waiting on a colleague, supplier, or client. The agent might shorten one duration without affecting the other. For instance, drafting may speed up, but approval remains a bottleneck.
The comparison should use the same acceptance criteria and comparably difficult cases. Where feasible, random assignment of cases between options or use of a comparison group helps separate the effect of the solution from changes in volume and seasonality. If only a “before and after” comparison is possible, present the result as an estimate and explain what else changed between periods.
Also compare with a simple automation: a fixed rule, a better form, or a connection between two apps could solve the problem at lower cost. The seven processes suited for a first pilot offer starting points for choosing a narrow enough area.
For an initial estimate, we recommend four cost categories. This list is adapted to the process and does not assume every project will have all these costs.
Process mapping, data preparation, integration with apps, configuring access, building tests, and training the team all require time. Transition and a period where the new process runs in parallel with the old one may occur. If using in-house hours, decide how you value them and keep actual payments separate from accounting allocations.
Model subscriptions and usage are just part of the equation. There may be costs for search, extracting text from documents, databases, storage, monitoring, and hosting. API is the interface through which the agent communicates with another system; some services charge for calls or operations performed via API.
Claude pricing documentation, for example, separates input and output, cache operations, and certain tools. Cache allows reuse of already processed information, per the provider’s terms. A retry, extra check, or new agent step may change total usage. For budgeting, use measured executions and actual contract rates, not just a one-off call price.
In a local installation, costs include equipment, power, operations, and available capacity. A server already owned may reduce a project’s upfront payments, but the administrator’s time and consumed resources remain relevant.
Measure time spent reading and approving, documents reworked, manually handled cases, and error investigation. If the agent saves four minutes drafting but adds five for review, the net result is negative for that task.
Also include failed attempts: the cost of a process doesn’t vanish just because the outcome was rejected. Furthermore, a human check may be justified by risk, even if it reduces time savings. Permission and audit rules should be evaluated as part of the solution.
Data sources change, models are updated, and apps change their interfaces. Plan for reevaluation, integration fixes, and support. For a longer time frame, estimate data export or migration too. Avoid double-counting services already covered in a subscription, and document how shared costs are allocated across processes.
The following set is a proposal for a pilot, not a mandatory standard. The definitions and accountable owner for each indicator are set before testing.
A useful formula is:
Cost per accepted outcome = total cost attributed to the period / number of accepted outcomes in that period.
When comparing two options, use the same cost scope and outcome definition. Otherwise, the option excluding review work may appear artificially cheaper.
Anthropic’s agent assessment guide, published January 9, 2026, recommends result verification and combining evaluation methods. For an SME, the practical takeaway is simple: don’t let the agent self-confirm its success. Use evidence from operational systems and human checks appropriate to risk.
Freed-up capacity means people have available time. If payroll and other payments stay the same, the budget doesn’t decrease just because some tasks take less time. The benefit could instead be better service, higher volume, or less pressure on the team.
Cash savings occur when the company actually avoids payments: for example, overtime or outsourced services no longer needed. These must be compared with added costs from the solution. The distinction between cost reduction and available capacity is also present in the Government Efficiency Framework. Here we use the conceptual distinction, not the UK government’s reporting rules.
Additional revenue should be analyzed by the contribution left after costs of extra activity. If freed-up time allows for more sales, don’t assign all new turnover to the agent, and don’t ignore the costs of goods or delivery. It should also be shown there was actual demand and the process change contributed to the result.
Don’t add the estimated wage value of the same hours to savings or margin produced by using them. That would be double-counting the same benefit.
Let’s return to the distributor’s inquiry process. All figures below are hypothetical. They do not reflect i8.ro results, market prices, or promises of performance. We assume volume and quality are comparable and the process meets acceptance criteria.
The company receives 1,200 requests per month. The existing process averages 10 minutes of human work per request, which totals 200 hours per month. After introducing the agent, the average drops to four minutes, including reviews and exceptions, across all eligible requests. That gives 80 hours per month and a freed-up capacity of 120 hours.
At a reference labor cost of 60 lei per hour, the equivalent of 120 hours is 7,200 lei. This assessment describes the capacity, but does not prove that the company’s payments will decrease by 7,200 lei.
In the baseline scenario, we assume that 60 of the freed-up hours replace external services that were previously paid at 60 lei per hour. The company can actually stop these expenses, thus avoiding 3,600 lei per month. The other 60 hours free up internal capacity, without an immediate reduction in salaries.
Implementation costs 18,000 lei, paid up front. Additional operating costs are 2,400 lei per month and include all new recurring payments assumed in this example. Remaining internal work and ongoing salaries are included in the operational analysis and the TCO, but are not shown as cash savings.
The net monthly cash benefit is:
3,600 − 2,400 = 1,200 lei.
If benefits begin immediately and remain constant, the simple payback period for implementation is:
18,000 / 1,200 = 15 months.
In the first year, we have benefits of 3,600 × 12 = 43,200 lei and costs of 18,000 + 2,400 × 12 = 46,800 lei. The convention used here is:
Cash ROI in year one = (year one benefits − all year one project costs) / all year one project costs × 100.
The result is −7.7%. The project generates a positive monthly cash benefit after recurring costs, but does not recover the initial investment within the first 12 months. The two conclusions are compatible.
The calculation is simplified, without financial discounting or tax effects. It does not assume any other costs or benefits. If adoption is gradual, benefits will appear later, or if implementation goes over budget, payback takes longer. For large or multi-year investments, the timing of cash flows and the cost of financing must also be considered.
We keep the implementation at 18,000 lei and the additional monthly cost at 2,400 lei. We vary only how much of the freed-up capacity can avoid external payments, keeping all other assumptions the same:
The scenarios show how sensitive the calculation is to a business assumption: how much cost can actually be avoided. Improving the model is not enough if the company cannot make use of the freed time. A project can still have operational reasons to continue, such as reducing response time, but these need to be explicitly stated.
In Generative AI at Work, published in 2025 in the Quarterly Journal of Economics, the authors analyze the introduction of an assistant for 5,172 support employees. The number of problems resolved per hour increased on average by 15%, with different effects depending on experience. Humans controlled the conversations. The study does not measure ROI for an autonomous agent and does not allow converting the percentage into wage savings for any company.
Navigating the Jagged Technological Frontier, published in Organization Science in March 2026, uses a 2023 experiment with 758 consultants. For tasks suited to the model, participants with AI completed more tasks and worked faster. On a difficult problem outside tested capabilities, accuracy declined. The key lesson is that task selection influences the result.
In a METR experiment published in July 2025, 16 experienced developers worked on 246 real tasks. Tools from early 2025 increased time by 19% in the tested context. The February 2026 METR update explains why subsequent data does not reliably estimate the current effect, including due to participant selection and working with multiple agents simultaneously. We do not use the old result as a general statement about programmers in 2026.
These findings support the need for a local pilot, not a universal ROI percentage. We measure the company’s process with its own tools, people, and rules.
Before the pilot, the process owner and the person tracking the budget set the decision criteria. Quality and authorization of actions have their own conditions; we do not automatically offset them with an average time saving.
We continue or expand when results meet criteria, costs are understandable, and the benefit repeats across enough comparable cases. Expansion happens gradually, to observe whether exception rates and supervisory effort change.
We limit when value appears only in one category—for example, inquiries about standard products, but not quotes with negotiated terms. An agent useful in a small domain can be a better investment than one that tries to cover the whole process.
We stop or redesign when checks take more effort than they save, errors exceed acceptable limits, or there is no credible way to make use of the capacity. A large number of uses alone is not a reason to continue.
A good decision might also be to keep an assistant with human approval or to revert to simple automation. The goal is a more efficient and verifiable process, within the company’s economic limits.
If you want to evaluate an AI agent on your company’s data, talk to the Imagine Infinity team. We can define the accepted result, inventory the costs, and design a pilot with indicators to support your investment decision.