Skip to content

How much is an AI agent worth: total cost, KPIs, and investment returns

9/27/2026

An AI agent can prepare answers in seconds, check documents, and execute steps within a process. However, for a small or medium-sized business, the speed shown in a demo doesn’t reveal the real value of the investment. What matters are the results that are actually accepted, how much work people still have to do, and what expenses or revenues truly change.

Saving ten hours per week can be valuable even if it doesn’t immediately lower any bills. It could allow the team to respond faster or handle more requests. But for an investment decision, we need to distinguish between capacity freed up, actual cash savings, and additional margin obtained.

In the article on choosing a platform we discussed implementation options and the costs involved. Here, we build a framework for measuring value: the initial benchmark, total cost, performance indicators, and a complete return analysis.

Sources verified as of September 27, 2026. Percentages from studies refer to the specific contexts tested. The figures in the financial example are hypothetical and are separate from research results.

Four terms for the same decision

TCO, Total Cost of Ownership, is the full cost of owning and operating the solution over a defined period. For an agent, this might include implementation, integration, subscriptions, usage, infrastructure, verification work, maintenance, and switching providers. We select the period before comparing solutions: a high upfront cost looks very different in a three-month calculation versus one over three years.

KPI, Key Performance Indicator, means a key performance indicator. A usable KPI shows whether the process delivers the desired outcome—for example, correctly resolved requests, minutes of human work, or cost per accepted case. The number of generated messages may describe activity without measuring its real value.

ROI, Return on Investment, expresses the ratio of net benefit to the cost used as a base. Since different conventions are used in practice, it’s important to specify the period, included benefits, and denominator. In our example, we'll calculate first-year cash ROI in relation to all project cash costs for that year.

Payback period shows when accumulated net benefits cover the initial investment. By itself, it doesn’t tell you how good the investment will be after that point, nor does it replace quality or risk analysis.

Start with a result that can be verified

“The agent has finished” isn’t a sufficient definition of success. The result needs to be described in the terms of the company’s process.

Let’s consider a distributor who receives product inquiries. The agent classifies the message, searches for information in the catalog and approved commercial terms, prepares a response, and sends it to a colleague for validation. It doesn’t alter prices or promise delivery times that aren’t confirmed by the system.

For this process, an accepted result might mean the product is correctly identified, information is current, the response aligns with company policy, the request is handled on time, and required interventions are logged. A fluent message with the wrong price doesn’t pass the test. Nor does an automatically closed case that wasn’t actually resolved.

We separate results accepted without rework from those corrected by people. Both can have value but incur different costs. Sending an exception to the right person can be the agent’s correct behavior; still, the dashboard should distinguish this from fully resolving a request.

This results-driven approach aligns with the FinOps Foundation, which differentiates technical unit costs, like cost per token, from business-related costs, such as cost per resolved case. A token is a unit of text or other content processed by the model. It’s useful for billing and optimization purposes, but doesn’t directly represent a resolved request.

Measure the current process before promising savings

The initial benchmark, also known as the baseline, describes how the process works without the new solution. For the selected period, we record volume, case complexity, people’s active time, rework, errors, and costs. Include tough and incomplete cases, not just examples that look good in a demo.

It’s useful to separate two types of duration. Active time shows how much time a person actually works on a case. Time to closure includes the waiting period after a colleague, supplier, or client. The agent may shorten one duration without changing the other. For example, drafting can become quick, but approvals might remain the bottleneck.

The comparison needs to use the same acceptance criteria and cases of comparable complexity. Where feasible, random assignment of cases between options or a comparison group helps isolate the solution’s impact from volume changes and seasonality effects. If we only have a 'before-and-after' comparison, present the outcome as an estimate and explain what else changed between the periods.

Also compare with a simple automation. A fixed rule, a better form, or a connection between two apps might solve the problem with far lower costs. The seven processes best suited for a first pilot offers starting points for choosing a sufficiently narrow field.

The total cost also includes the work around the agent.

For an initial estimate, we recommend four categories. The list adapts to the process; it doesn’t assume every project incurs all costs.

1. Initial Costs

Mapping the process, preparing the data, integrating with applications, configuring access, building tests, and training the team all take time. You may also encounter a transition period when the new process runs alongside the old one. If you use internal hours, decide how to value them and keep actual payments separate from economic allocations.

2. Recurring Technology Costs

Model subscriptions and usage are just part of the picture. There may also be costs for search, extracting text from documents, databases, storage, monitoring, and hosting. API is the interface through which the agent communicates with another system; some services charge for calls or operations made through it.

Claude’s pricing documentation, for example, separates model input and output, cache operations, and certain tools. Cache allows reuse of already processed information, based on the provider’s terms. A retry, an extra check, or a new step in the agent’s flow can change overall usage. For budgeting, use measured executions and your actual contract rates, not the cost of a single isolated call.

For on-premises setups, costs include hardware, electricity, maintenance, and available capacity. A server already purchased may reduce upfront payments for the project, but admin time and consumed resources still matter.

3. Validation, Correction, and Exceptions

Track reading and approval time, redone documents, cases handled manually, and error investigations. If an agent saves four minutes on drafting but adds five on checking, the net effect is unfavorable for that task.

Include failed attempts as well. The cost of a process doesn’t disappear just because the outcome was rejected. Also, human validation may be justified by risk even if it reduces time savings. Permission and audit rules should be evaluated as part of the solution.

4. Maintenance and Change

Data sources change, models are updated, and applications change their interfaces. Plan for review, fixing integrations, and support. Over a longer horizon, estimate data export or migration as well. Don’t double-count services already included in a subscription, and document how you split shared costs between processes.

Six indicators that can support the decision

The following set is a suggested pilot, not a mandatory standard. Definitions and the responsible party for each indicator are established before testing.

  1. Proportion of cases completed correctly and accepted. We count accepted results from all eligible cases entering the pilot. Failures, dropouts, and cases with no validation by deadline remain visible in records. At the end of the period, we distinguish cases still in progress from those whose deadline has passed.
  2. Share of results accepted with no corrections. This shows how many cases passed validation without human rework. We report these cases as a share of all eligible cases included in the pilot. A result checked by a colleague can be counted here; verification and correction are separate activities. Results salvaged by intervention are reported separately.
  3. Minutes of human work per case. We total the time spent on validation, corrections, and exceptions across all eligible cases and divide by the number of those cases. The time when the agent operates alone is not automatically employee work time.
  4. Total cost per accepted result. We divide the costs assigned to the process by the number of accepted outcomes in the same period. Costs include failed attempts. If we include part of the implementation, we explain the allocation. If no results are accepted, the unit cost cannot usefully be reported as zero.
  5. Time to closure. We track the median (the middle value in the distribution) and the 95th percentile (the duration by which about 95% of measured cases are closed). We keep the age of still-open cases separate, so the slowest ones don’t vanish from the report.
  6. Errors and their consequences. We separate an unclear phrasing from a wrong price quoted to a client or from an unauthorized action. A good average rate doesn’t negate an incident with high impact.

A useful formula is:

Cost per accepted result = total cost allocated to the period / number of results accepted in that period.

When comparing two variants, use the same cost perimeter and the same outcome definition. Otherwise, a variant that excludes revision work may appear artificially cheaper.

Anthropic’s agent evaluation guide, published on January 9, 2026, recommends checking outcomes and combining evaluation methods. For an SME, the practical application is simple: don’t let the agent confirm its own success. Use evidence from work systems and human checks tailored to the associated risk.

Released hours, cost savings, and revenues are different things

Released capacity means people have available time. If wages and other payments stay the same, the budget doesn't shrink just because some tasks take less time. The benefit might be better service, increased volume, or less pressure on the team.

Cash savings occur when the company actually avoids payments: for example, overtime or outsourcing that is no longer needed. These must be compared with the additional costs of the solution. The distinction between reducing expenses and available capacity also appears in the Government Efficiency Framework. Here we use the conceptual distinction, not the UK government's reporting rules.

Additional revenue should be analyzed by the remaining contribution after the costs of the extra activity. If freed time allows for more sales, we don't automatically credit the agent with the entire turnover and we don't ignore the cost of goods or delivery. It must also be shown that there was demand and that the process change contributed to the outcome.

We don't add the estimated wage value of the same hours to the savings or margin obtained from their use. That would mean counting the same benefit twice.

A full calculation, with explicit assumptions

Let's return to the distributor's request process. All the following figures are hypothetical. They do not represent an i8.ro result, a market price, or a performance guarantee. We assume that volume and quality remain comparable and that the process meets acceptance criteria.

From minutes to capacity

The company receives 1,200 requests per month. The current process consumes an average of 10 minutes of human work per request, that is, 200 hours per month. After introducing the agent, the average becomes four minutes, including checks and exceptions, calculated across all eligible requests. That results in 80 hours per month and a released capacity of 120 hours.

At a reference labor cost of 60 lei per hour, the equivalent of 120 hours is 7,200 lei. This estimate describes capacity, not proof that the company's payments will decrease by 7,200 lei.

From capacity to cash flow

In the main scenario, we assume that 60 of the released hours replace outsourced services previously paid at 60 lei per hour. The company can actually stop these expenses, thus avoiding 3,600 lei per month. The remaining 60 hours free up internal capacity, but don't immediately reduce salaries.

Implementation costs 18,000 lei, paid upfront. Additional operating costs are 2,400 lei per month and include all new recurring payments assumed in the example. The remaining internal work and the salaries that continue are included in the operational analysis and the TCO, but are not presented as cash savings.

The net monthly cash benefit is:

3,600 − 2,400 = 1,200 lei.

If the benefits start immediately and stay constant, the simple payback for implementation is:

18,000 / 1,200 = 15 months.

In the first year, benefits are 3,600 × 12 = 43,200 lei and costs are 18,000 + 2,400 × 12 = 46,800 lei. The convention used here is:

First-year cash ROI = (first-year benefits – all first-year project costs) / all first-year project costs × 100.

The result is −7,7%. The project produces a positive monthly benefit after recurring costs, but does not recover the initial investment in the first 12 months. The two conclusions are compatible.

The calculation is simplified, with no financial discounting and no tax effects. It assumes no other costs or benefits. If adoption is gradual and benefits appear later, or if implementation exceeds the budget, recovery takes longer. For large or multi-year investments, the timing of cash flows and financing costs must also be considered.

Three scenarios for the same decision

We keep the implementation cost of 18,000 lei and the additional monthly cost of 2,400 lei. The only variable is how much of the freed capacity can avoid external payments, while all other assumptions remain the same:

  • Prudent: 30 external hours avoided, or 1,800 lei per month. The difference compared to operation is negative, at –600 lei per month. The investment is not recovered through this steady flow; first-year ROI is approximately −53.8%.
  • Central: 60 external hours avoided, or 3,600 lei per month. The net benefit is 1,200 lei per month, payback takes 15 months, and first-year ROI is approximately −7.7%.
  • Favorable: 90 external hours avoided, or 5,400 lei per month. The net benefit is 3,000 lei per month, payback takes six months, and first-year ROI is approximately 38.5%.

These scenarios show the sensitivity of the calculation to a business assumption: how much spending can actually be avoided. It's not enough to improve the model if the company cannot make use of the freed-up time. A project may still have operational reasons to continue, such as reducing response times, but these must be stated explicitly.

What research tells us and what we cannot infer from it

In Generative AI at Work, published in 2025 in the Quarterly Journal of Economics, analyzes the introduction of an assistant for 5,172 support employees. The number of problems solved per hour increased by an average of 15%, with different effects depending on experience. People controlled the conversations. The study does not measure the ROI of an autonomous agent and does not allow converting the percentage into salary savings for any company.

Navigating the Jagged Technological Frontier, published in Organization Science in March 2026, uses a 2023 experiment with 758 consultants. For tasks suitable to the model, participants with AI completed more tasks and worked faster. On a difficult problem outside the model's tested capabilities, accuracy decreased. The relevant lesson is that task selection influences the outcome.

In a METR experiment published in July 2025, 16 experienced developers worked on 246 real tasks. Early 2025 tools increased completion time by 19% in the tested context. The METR update from February 2026 explains why later data do not reliably estimate the current effect, including participant selection and working with multiple agents at once. We do not use the old result as a general statement about programmers in 2026.

These results support the need for a local pilot, not a universal profitability percentage. We measure the company's process with its own tools, people, and rules.

When to continue, scale back, or stop

Before the pilot, the process owner and the person tracking the budget set the decision criteria. Action quality and authorization have their own requirements; we do not automatically offset them with an average time-saving.

We continue or expand when results meet the criteria, costs are transparent, and benefits recur across enough comparable cases. Expansion is done gradually to observe whether the rate of exceptions and supervision work changes.

We scale back when value appears only in one category: for example, queries about standard products, but not offers with negotiated terms. An agent that's useful in a small area may be a better investment than one that tries to cover the whole process.

We stop or redesign when reviews cost more than they save, errors exceed acceptable limits, or there is no credible mechanism for capturing the gains. A high number of uses alone is not a sufficient reason to continue.

A good decision can also mean keeping an assistant with human approval or returning to a simple automation. The aim is a more efficient and auditable process, within the company's economic constraints.

Main sources

Want to evaluate an AI agent on your company’s data? We can define the accepted outcome, inventory the costs, and design a pilot with indicators to support your investment decision.

Recommended for you

The 90-Day Adoption Plan: From First Pilot to a Measurable Agentic System

When the agent buys for us: budgets, payments, and autonomy limits

Where to Start: Seven Suitable Processes for AI Agents in an SME

Cookies

We use cookies required for the site to work. With your consent we also enable additional features (videos, maps) or anonymous statistics. You can change your choice at any time from the site footer.

Cookie policyPrivacy noticeTerms and conditionsCookie preferences