Securafy AI Lab

How to Measure AI Training ROI Without Inventing the Numbers

Written by Jillian O. | Aug 28, 2026, 12:00:00 PM

How Do You Actually Know Your AI Training Is Working?

Every CFO I have watched sit through an AI training ROI slide asks the same question within ninety seconds: where did that number come from? Not because they doubt AI is useful, but because "we saved 40,000 hours this year" sounds precise and falls apart the moment someone asks what it is made of.

Most AI training ROI claims are built the same way: survey employees on hours saved per week, multiply by a blended labor rate, multiply again by headcount, and annualize the result. That is four assumptions stacked on each other, and every one compounds error. Self-reported time estimates run generous. Blended rates ignore who actually did the work. Headcount extrapolation assumes uniform adoption that never exists. Annualizing bakes a full year out of a trend measured over a few weeks. By the time the number reaches a board deck, it has been laundered through enough multiplication that nobody can trace it back to anything real.

Why Is "Hours Saved" the Most Popular AI Metric and Also the Worst?

Hours saved is the most reported AI productivity metric because it is the easiest to collect: ask people how much faster they feel, and they give you a number. It is also the least defensible metric, because saved time is almost never converted into anything measurable. It gets reabsorbed. The work expands to fill it, and nothing downstream changes: output stays flat, headcount stays flat, revenue stays flat.

This is not a cynical guess. It matches what controlled research keeps finding about the gap between perceived and measured AI productivity. METR's randomized controlled study of experienced open-source developers found that when developers used AI coding tools on real issues in codebases they knew well, they took 19% longer, not less. Before the study, they predicted AI would speed them up by 24%. After finishing, they still believed they had been about 20% faster. Perception and measured reality moved in opposite directions, and the developers never noticed. If experienced engineers cannot judge their own speedup under controlled observation, a company-wide self-reported survey has no chance of producing a trustworthy figure.

Direct answer: the most defensible way to measure AI training ROI is to skip self-reported hours and track before-and-after changes in metrics your systems already record — cycle time on named processes, error and rework rates, ticket resolution, and sanctioned-tool adoption — because these numbers exist independent of how anyone feels about the training.

Is There Evidence That Company-Wide AI Gains Are Overstated?

Yes. MIT's NANDA initiative examined 300 publicly disclosed enterprise AI deployments and found that roughly 95% of generative AI pilots showed no measurable effect on profit and loss, despite an estimated $30 to $40 billion in enterprise spending. The report's own framing is blunt: the failure was not model quality, it was the absence of workflow integration and a way to tell whether anything had changed.

The pattern shows up at named companies too. In May 2026, Uber's chief operating officer, Andrew Macdonald, said the company could not draw a line from its AI coding tool spending to any measurable increase in shipped product features, even with heavy engineer adoption and a large share of code commits AI-assisted. Senior engineering leaders, pressed on which delayed projects had actually been unblocked, could not point to one, as Macdonald described on the record. That is heavy AI adoption with no defensible ROI story, from the company's own operations chief. The lesson matches MIT's data: adoption and impact are different measurements, and only one is worth reporting upward.

What Should You Measure Instead of Hours Saved?

Start with what your systems already record, before training happens. If you cannot state a baseline number from before the training rolled out, you cannot claim a before-and-after change afterward. This is the discipline missing from most AI training programs: nobody captured the starting point.

Pick two or three named processes rather than measuring the whole organization at once. A named process — client onboarding, invoice processing, first-tier support triage — has a defined start and end point, existing data, and people who will notice if the number moves. "Overall productivity" has none of those things, which is why it is impossible to falsify and impossible to trust.

  • Cycle time on a specific process: elapsed time from start to completion, measured the same way before and after training, not self-estimated.
  • Error and rework rates: how often AI-assisted output needs correction or redoing, which captures the hidden cost "hours saved" hides.
  • Ticket volume and first-contact resolution: whether trained support staff resolve more issues without escalation.
  • Sanctioned-tool adoption versus shadow usage: whether people use the tools your policy governs or route around them.
  • Time-to-competency for new hires: whether trained employees reach independent productivity faster, a number HR already tracks.

Separate leading indicators from lagging ones when you report these. Tool adoption and completed certifications are leading indicators — they tell you the program is running, not that it worked. Cycle time, error rates, and resolution rates are lagging indicators — they tell you whether it produced a result. Reporting a leading indicator as if it were a result is the second most common way AI training ROI claims collapse under scrutiny, right behind inflated hours-saved math.

How Should You Measure ROI Without Guessing at a Dollar Figure?

Two long-standing training evaluation frameworks were built for this problem before AI existed. The Kirkpatrick Model separates training evaluation into four levels: how learners reacted, what they learned, whether they changed on-the-job behavior, and whether that behavior produced a business result. Most organizations stop at the first two levels because they are easy to survey. The levels that matter for a CFO conversation are three and four, and they require operational data, not a feedback form.

The Phillips ROI Methodology extends that model with a fifth level that converts measured results into a monetary figure, but only after isolating training's specific contribution from other factors that could explain the same change. That isolation step is the part almost every AI training ROI claim skips. If cycle time improved the same quarter you also changed vendors or hired experienced staff, you cannot attribute the entire improvement to training, and a rigorous methodology says so rather than assuming it away.

Applying either framework means reporting Level 3 and Level 4 data, not Level 1 satisfaction scores. A defensible smaller number that survives isolation analysis beats an impressive unverifiable one, because the inflated figure is what gets the program cut at the next budget cycle when someone checks the math. This is where SMB AI initiatives most often fail — not at the training stage, but at the measurement stage that follows it.

Can You Put a Number on Prevented Security Incidents?

Not directly, and claiming you can is one of the fastest ways to lose credibility with finance. Prevented incidents are real value: training that changes how employees handle sensitive data in AI tools, or reduces near-misses reported through existing channels, is worth reporting. But you cannot count an incident that did not happen, and modeling a dollar figure for "breaches avoided" requires a base rate, a severity distribution, and a causal link no organization can actually verify.

IBM's 2025 Cost of a Data Breach Report found the average U.S. breach now costs $10.22 million, and that a high level of unsanctioned "shadow AI" use adds roughly $670,000 to the global average breach cost. That figure is a benchmark for why the behavior matters, not a number to claim you personally saved. The legitimate way to represent risk-avoidance value is to report measured behavior change directly: fewer policy violations logged, lower shadow-tool usage, faster reporting of suspicious AI outputs, more employees correctly routing sensitive data requests. Those are auditable. A projected breach-cost-avoided figure is not, and a CFO who has seen a few of these numbers before will ask exactly where it came from.

The NIST AI Risk Management Framework's Measure function makes the same distinction in more technical language: document which risks are tracked with real metrics, and note explicitly which cannot be measured directly, rather than substituting a modeled estimate for a number you do not have. That same honesty standard belongs in a training ROI report. Say what you measured. Say what you could not. Do not fill the gap with an assumption dressed up as a finding.

Which Metrics Survive a Second Question From Finance?

The table below separates the metrics that get repeated in decks but collapse on follow-up from the ones built to survive it.

Weak but popular metricDefensible alternative
Self-reported hours saved per employeeCycle time on a named process, measured before and after in system data
Hours saved × blended labor rate × headcountError and rework rate on AI-assisted output, isolated from other changes
Percentage of staff who "completed" trainingSanctioned-tool adoption rate versus shadow-tool usage after training
Modeled dollar value of breaches avoidedReduction in documented policy violations or near-miss reports
Annualized productivity gain extrapolated from a pilotTime-to-competency for new hires measured across actual cohorts

Notice what the right column has in common: every metric already exists somewhere in your systems before you introduce AI training. You are pulling a number your ticketing system, HR platform, or compliance log already produces, and comparing it against a baseline captured before training started.

How Does Securafy Help Clients Measure This Correctly?

When Securafy designs AI training for a client, the first working session is not about content, it is about instrumentation: which two or three processes we will track, what the baseline looks like now, and who owns pulling that number in ninety days. We push clients toward named processes with existing data instead of company-wide "AI adoption" targets, because that is the only way a result survives a second look from finance.

On the risk side, we build reporting around measured behavior — sanctioned-tool usage, policy violation trends, near-miss logs — rather than a modeled cost-avoided figure, because that number holds up in an audit. This is the same discipline behind helping clients secure Copilot and other AI agents before they become shadow IT, and setting clear boundaries for everyday tools like ChatGPT and Copilot.

That gap shows up whenever leadership treats AI adoption as a tool problem instead of a skills problem. A tool rollout produces adoption metrics; a training program, measured properly, produces a defensible business result. Securafy's approach treats AI governance as the foundation the measurement sits on, because you cannot isolate a training effect where nobody tracked usage beforehand. The broader case is laid out in what SMBs actually gain when AI is done right — a smaller, provable number, not a bigger, invented one.

Where To Go From Here

If your last AI training report leaned on hours saved or a headcount-extrapolated dollar figure, the fix is not a better spreadsheet. It is picking a smaller set of named processes, capturing a real baseline, and reporting the result your systems can actually back up.

If your team is moving faster with AI than your guardrails are, start with structured training rather than another tool. Securafy AI University gives your people role-based AI training with security built into the material, not bolted on afterward.

If you would rather talk through your specific environment first, book a strategy call with Securafy and we will walk your current AI usage, exposure, and the fastest path to safe adoption.