
An AI agent can prepare responses in seconds, check documents, and execute steps in a process. For a small or medium business, however, demo speed doesn’t reveal the value of the investment. What matters are the accepted results, how much work remains for people, and which expenses or revenues are actually changed.
Saving ten hours a week can be valuable even if it doesn’t immediately lower any invoice. It may enable the team to respond faster or handle more requests. But when making an investment decision, we must separate freed-up capacity from actual cash savings and any additional margin generated.
In the article on choosing a platform, we discussed implementation options and associated costs. Here, we build a framework for measuring value: the initial baseline, total cost, performance indicators, and a full return-on-investment calculation.
Sources verified as of September 27, 2026. Percentages from studies refer to the specific contexts tested. Figures in the financial example are hypothetical and separate from research results.
TCO, Total Cost of Ownership, is the total cost of owning and operating the solution over a defined period. For an agent, this may include implementation, integration, subscriptions, usage, infrastructure, verification work, maintenance, and supplier switching. The period is chosen before comparing solutions: a large upfront cost looks different over three months versus over three years.
KPI, Key Performance Indicator, is a main performance metric. A useful KPI shows whether the process gets the desired result; for example, correctly resolved requests, minutes of human labor, or cost per accepted case. The number of generated messages can reflect activity without actually measuring its value.
ROI, Return on Investment, expresses the ratio between net benefit and the base cost. Because different conventions are used in practice, the period, included benefits, and denominator must be specified. In our example, we’ll calculate the cash ROI from the first year, compared to all project cash costs for that year.
Payback period shows when cumulative net benefits cover the initial investment. On its own, it doesn’t show how good the investment is after that moment and doesn’t replace quality or risk analysis.
“Agent has finished” is not a sufficient definition of success. The outcome needs to be described in the firm’s process terms.
Let’s take a distributor receiving product inquiries. The agent classifies the message, looks up information in the catalog and approved commercial terms, drafts a reply, and sends it to a colleague for validation. The agent does not change prices or promise delivery dates unconfirmed by the system.
For this process, an accepted result might mean the product is correctly identified, the information is up to date, the response complies with company policy, the request is handled on time, and required actions are logged. A fluent message with an incorrect price does not pass the test. Nor does a case automatically closed without being resolved count as a success.
We distinguish accepted outcomes that require no rework from those corrected by people. Both have value, but entail different costs. Sending an exception to the right person may be the correct agent behavior; in the dashboard, however, it should be distinguished from final resolution of the request.
This results-focused approach aligns with FinOps Foundation, which differentiates technical costs per unit—like cost per token—from business-related costs such as cost per resolved case. A token is a unit of text or other content processed by the model. It’s useful for billing and optimization, but does not directly represent a resolved request.
The initial benchmark, also known as baseline, describes how the process works without the new solution. For the chosen period, record volume, case complexity, human active time, rework, errors, and costs. Include difficult and incomplete cases, not just showcase examples from a demo.
It’s useful to separate two durations. Active time indicates how long a person spends actually working on a case. Time to closure also includes waiting after a colleague, supplier, or customer. The agent may shorten one duration without changing the other. For instance, drafting may become quick, but approval remains the bottleneck.
The comparison must apply the same acceptance criteria and similar case difficulty. Where feasible, random distribution of cases between variants or a comparison group helps separate the solution’s effect from volume and seasonal changes. If all you have is a “before-and-after” comparison, present the result as an estimate and explain any other changes between the two periods.
Also compare to simple automation. A fixed rule, a better form, or connecting two apps may solve the problem at lower cost. The seven processes suitable for an initial pilot provide starting points for choosing a sufficiently narrow domain.
For an initial estimate, we suggest four categories. The list is adapted to the process and does not assume every project will have all costs.
Process mapping, data preparation, integration with apps, access setup, building tests, and training the team all take time. There may be a transition and overlap period when the new and old processes run in parallel. If you use internal hours, decide how to assign value and keep actual payments separate from internal cost allocations.
Model subscriptions and usage are just part of it. There may also be costs for search, extracting text from documents, databases, storage, monitoring, and hosting. API is the interface the agent uses to communicate with other systems; some services charge per call or per operation carried out through it.
Claude’s documentation on pricing, for instance, separates model input and output, cache operations, and certain tools. The cache enables reuse of previously processed data, subject to provider terms. A retry, extra validation, or a new agent step may all alter total usage. For budgeting, use measured executions and the contracted rates, not the price of a single isolated call.
For on-premises setups, costs include hardware, power, operation, and available capacity. An already purchased server may reduce the project’s initial outlay, but administrator time and used resources still matter.
Measure reading and approval time, documents redone, manually handled cases, and error investigations. If an agent saves four minutes drafting but adds five for checking, the net result is negative for that task.
Include failed attempts, too. The cost of a process doesn’t disappear because the outcome was rejected. Likewise, human checks may be justified by risk even if they reduce time saved overall. Permission and audit rules must be evaluated as part of the solution.
Data sources evolve, models update, and apps change their interfaces. Plan for re-evaluation, fixing integrations, and support. For a longer horizon, estimate data export or migration. Don’t double-count a service already included in your subscription, and document how you split shared costs between processes.
The following set is a proposal for a pilot, not a mandatory standard. Definitions and the person responsible for each KPI should be set before the test.
A useful formula is:
Cost per accepted result = total cost assigned to the period / number of accepted outcomes in that period.
When comparing two variants, use the same cost scope and the same outcome definition. Otherwise, omitting review work can make one option look artificially cheaper.
Anthropic’s guide to agent evaluation, published January 9, 2026, recommends reviewing results and combining evaluation methods. For SMEs, the practical lesson is simple: don’t let the agent self-certify its own success. Use working system evidence and human checks appropriate to risk level.
Freed-up capacity means people have more available time. If salaries and other outlays stay the same, the budget doesn’t drop just because some tasks take less time. The benefit may show up as better service, higher volume, or less pressure on the team.
Cash savings occur when a company actually avoids paying certain bills—for example, overtime or outsourcing no longer needed. These should be compared with additional solution costs. The distinction between lower spending and extra capacity also appears in the Government Efficiency Framework. We’re using the conceptual distinction here, not the UK government’s specific accounting rules.
Additional revenue needs to be analyzed in terms of the contribution after deducting the costs of the extra activity. If freed-up time enables more sales, don’t automatically credit the agent with the entire revenue, and don’t ignore product or delivery costs. You must show there was real demand and that the process change played a role.
Don’t add both the estimated wage value of the same hours and the margin or savings generated by using those hours—that would mean double-counting the same benefit.
Back to the distributor’s request process. All the following numbers are hypothetical. They do not represent i8.ro results, market prices, or a performance guarantee. We assume volume and quality remain comparable and that the process meets acceptance criteria.
The company receives 1,200 requests per month. The current process takes an average of 10 minutes of human work per request—that’s 200 hours per month. After introducing the agent, the average drops to four minutes, including checks and exceptions, averaged across all eligible requests. That’s 80 hours per month and a freed-up capacity of 120 hours.
At a reference labor cost of 60 lei per hour, the equivalent of 120 hours is 7,200 lei. This assessment describes capacity, but doesn't prove that the company's payments decrease by 7,200 lei.
In the central scenario, we assume that 60 of the freed hours replace external services that were previously paid at 60 lei per hour. The company can actually stop these expenses, so it avoids 3,600 lei per month. The other 60 hours free up internal capacity, without immediately reducing salaries.
Implementation costs 18,000 lei, paid upfront. Additional operating costs are 2,400 lei per month and include all new recurring payments assumed in this example. Remaining internal work and continuing salaries are included in the operational analysis and TCO, but are not presented as cash savings.
The net monthly cash benefit is:
3,600 − 2,400 = 1,200 lei.
If the benefits start immediately and remain constant, simple payback for the implementation is:
18,000 / 1,200 = 15 months.
In the first year we have benefits of 3,600 × 12 = 43,200 lei and costs of 18,000 + 2,400 × 12 = 46,800 lei. The convention used here is:
First-year cash ROI = (first-year benefits − all project costs in the first year) / all project costs in the first year × 100.
The result is −7.7%. The project produces a positive monthly benefit after recurring costs, but doesn't recover the initial investment within the first 12 months. These two conclusions are compatible.
The calculation is simplified, without financial discounting and without tax effects. It does not assume other costs or benefits. If adoption is gradual, benefits appear later, or implementation exceeds budget, payback takes longer. For large or multi-year investments, timing of cash flows and cost of financing must also be considered.
We keep the 18,000 lei implementation and the additional monthly cost of 2,400 lei. We vary only how much of the freed capacity can actually avoid external payments, keeping all other assumptions the same:
The scenarios demonstrate the calculation's sensitivity to a business assumption: how much spending can actually be avoided. It is not enough to improve the model if the company cannot capitalize on the freed-up time. A project may still have operational reasons to continue, such as reducing response times, but these must be explicitly reported.
In Generative AI at Work, published in 2025 in the Quarterly Journal of Economics, the authors analyze the introduction of an assistant for 5,172 support employees. The number of issues resolved per hour increased by an average of 15%, with different effects depending on experience. People controlled the conversations. The study does not measure the ROI of an autonomous agent and does not allow the percentage to be translated into wage savings for just any company.
Navigating the Jagged Technological Frontier, published in Organization Science in March 2026, uses a 2023 experiment with 758 consultants. For tasks suited to the model, participants with AI completed more tasks and worked more quickly. On a difficult problem outside the tested capabilities, accuracy dropped. The relevant lesson is that task selection influences the outcome.
In a METR experiment published in July 2025, 16 experienced developers worked on 246 real tasks. Tools available at the start of 2025 actually increased time spent by 19% in the tested context. The February 2026 METR update explains why later data do not reliably estimate the current effect, due in part to participant selection and working with multiple agents at once. We do not use the old result as a general claim about programmers in 2026.
These results support the need for a local pilot, not a universal profitability percentage. We measure the company's own process with its own tools, people, and rules.
Before the pilot, the process owner and budget manager set decision criteria. Action quality and authorization have their own requirements; we don't automatically trade them off for average time savings.
We continue or expand when results meet the criteria, costs are clear, and the benefit repeats across enough similar cases. Expansion is done gradually to observe if exception rate and supervision requirements change.
We restrict when value appears in just one category—for example, queries about standard products, but not custom offers. An agent useful in a small area may be a better investment than one trying to cover the whole process.
We stop or redesign when checks consume more than they save, errors exceed acceptable limits, or there is no credible way to capitalize on freed capacity. A high number of uses alone is not a reason to continue.
A good decision may also be keeping an assistant with human approval, or returning to simple automation. The goal is a more efficient, auditable process within the firm's economic limits.
If you want to evaluate an AI agent on your company’s data, contact the Imagine Infinity team. We can define the accepted result, inventory the costs, and design a pilot with indicators to support your investment decision.