AI Readiness and Results

How Do You Measure Whether AI Is Actually Working in Your Business?

Most businesses using AI cannot say whether it is working, because nobody wrote down a before number. Here is a simple way to measure AI job by job: time, corrections, and where the saved hours go.

A hand-cut paper balance scale in a turquoise-to-violet gradient, one pan hanging lower than the other, standing for weighing a job before and after AI.

In brief

To measure whether AI is working, judge one named job, not the tool. Before AI goes in, record how long the job takes, how often it comes back for correction, and what the saved time should go to. Measure the same three after four weeks. If time fell, corrections did not rise and the hours went somewhere you can name, it is working.

Ask a room of business owners whether AI is saving their teams time and nearly every hand goes up. Ask how much time, on which job, compared with what, and the hands come down.

You know AI is working in your business when one named job takes less time, comes back for correction no more often than before, and the hours it saves go somewhere you can name. That means writing down a before number for the job before AI touches it, then measuring the same three things about four weeks later. Without the before number, you have an opinion, not a result.

This is the question small and medium-sized businesses (SMBs, or SMEs in the UK and Europe) are now asking, and it is the right one. The first wave of AI in small firms was about trying it. The second is about proving it.

Why Does AI Feel Like It Is Working When Nobody Can Prove It?

AI feels productive because it removes the blank page, and the relief of a quick first draft is easy to mistake for a finished job done faster.

One of the largest executive surveys on the subject shows the gap between feeling and evidence. In February 2026, a team of economists including Stanford's Nicholas Bloom published Firm Data on AI through the National Bureau of Economic Research. They surveyed nearly 6,000 senior executives in the US, UK, Germany and Australia. Sixty-nine percent of firms actively use AI. Yet "nine-in-ten" executives reported no impact on employment or productivity at their own firm over the past three years, while the same executives predicted AI would lift their productivity by an average of 1.4% over the next three.

British firms show the same pattern. The British Chambers of Commerce's Powering Productivity: AI and the Future of UK Work, published in March 2026, found 54% of UK firms actively using AI, up from 35% a year earlier. Firms already using it reported a net productivity expectation of +71 percentage points when asked what AI "will" do over the next twelve months. That is a measure of optimism. It is not a measure of what happened.

Then there is the problem that feelings can point the wrong way. In July 2025, the research group METR ran a controlled study with 16 experienced software developers working on open-source projects they knew well. Before the study they expected AI to speed them up by 24%. Afterwards they believed it had sped them up by 20%. Measured on the clock, they took 19% longer with AI. METR is careful to say the result may not generalize and that newer tools have since changed the picture. The lesson for an owner is not that AI slows people down. It is that the people doing the work could not tell, in either direction, without a clock.

My reading of all three: most businesses are not failing to get value from AI. They are failing to find out whether they are getting it.

What Is the Before Number Rule?

The Before Number Rule is how I suggest a small business measures whether AI is working: judge the job, not the tool, and never by how it feels. Before AI goes into a job, write down three before numbers: how long the job takes, how often it comes back for correction, and what the business would do with the time if it had it. Measure the same three after four weeks. If the time fell, the corrections did not rise and the hours went somewhere you can name, AI is working in that job. If you have no before number, you have an opinion, not a result.

Each of the three numbers catches a different way of fooling yourself.

Time catches the blank-page illusion. A first draft in thirty seconds means nothing if the person then spends forty minutes fixing it. Time the whole job, from the moment someone picks it up to the moment it leaves their desk finished.

Corrections catch the quality leak. If AI makes a job faster but a manager now sends one in three back, the time has not been saved. It has moved to someone more expensive. Count how often the finished work is returned, reworked or apologized for.

Destination catches the vanishing hour. Time saved that nobody can point to has not been saved in any sense a business can bank. Before you start, write one sentence on what the time is for: returning calls the same day, a second site visit, getting invoices out on Friday instead of Tuesday.

The rule is deliberately small. It asks for three numbers on one job, not a dashboard. It pairs with the One Job Rule, which tells an owner to choose one job, not one tool: the One Job Rule chooses the job, and the Before Number Rule tells you whether it worked.

How Do You Take a Before Number Without a Data Team?

You take a before number with a notepad, a phone timer and one honest week, because the aim is a fair comparison, not scientific precision.

Here is the method I suggest:

  1. Pick the job and write it down in one line. "Writing up a service visit report and emailing it to the customer." Not "admin".
  2. Time five to ten real instances the old way. Whoever does the job starts a timer when they begin and stops it when the work is finished and sent. Write each time down. Take the middle value, not the best one.
  3. Tally the corrections. For the same instances, note every time the work came back: a manager's edit, a customer's query, a fix after sending.
  4. Write the destination sentence. One line, agreed with the person doing the job, on what the saved time should go to.
  5. Bring AI in, then leave it alone for a fortnight. People are slower in the first days with any new way of working. Measuring the first week measures the learning, not the job.
  6. Time the same job again in weeks three and four. Same person, same definition of start and finish, same tally of corrections.

One condition matters more than the rest: the person doing the job must be the one keeping the numbers, and they must know the numbers will not be used against them. Measurement that feels like surveillance produces flattering numbers, and flattering numbers are worse than none.

What Does This Look Like in Practice?

Take a hypothetical 22-person heating and plumbing firm where engineers write a report after every service visit and the office emails it to the customer.

Before AI, the office manager times ten reports. The middle time is 25 minutes from notes to sent email. Two of the ten go back to an engineer because a detail is missing or unclear. The destination sentence reads: "Get every report out the same day, so customers stop calling to ask what we found."

The firm then has engineers dictate voice notes after each visit, and AI turns them into a draft report in the firm's standard layout, which the engineer checks and the office sends. After a fortnight of settling in, the office manager times ten more. The middle time is now 11 minutes. But three of the ten came back for correction, because the AI smoothed over details the engineer had left vague.

Under the Before Number Rule, that is not yet a success. Time fell, but corrections rose, so part of the saving is leaking into rework. The fix is not to drop AI. It is to change the brief, so the draft flags anything missing rather than filling the gap, and then to measure again. When the next ten show time holding at around 11 minutes and corrections back to two, the job passes. And because the firm wrote down its destination, it can check the thing it cared about: are reports going out the same day?

The numbers here are invented for illustration. The shape of the story is not. The second measurement almost always teaches you something the first could not.

Where Did the Saved Hours Go?

The saved hours go wherever the business has already decided they should, and if nobody decided, they go into more of the same work.

The US Chamber of Commerce Foundation and Ipsos asked this directly in their first Main Street AI Monitor, published in June 2026. They surveyed 1,070 people working at US small businesses. Half use AI at work, and only 10% reported formal AI training. Asked what they do with the time AI saves, 59% said they put it into "more work or higher-quality output", 43% into "learning, planning, or reviewing existing work", 28% into avoiding overtime and 23% into breaks or personal tasks.

None of those answers is wrong. A team that stops working overtime has gained something real. But notice who is making the choice: the individual, one day at a time, with no one keeping count. Across a 20-person business, that is how a genuine saving turns into a feeling that everyone is "a bit less stretched" and nothing an owner can point to at the end of the quarter.

The destination sentence is the cure. It turns saved time from a private benefit into a business decision. It also answers the question every owner eventually faces from a skeptical partner, lender or board: "What has AI actually done for us?"

The Mistake I See Most Often

The mistake I see most often is measuring the tool instead of the job: counting licenses, logins and how many people "use AI" as if adoption were the same thing as value.

Adoption numbers are easy to collect, which is why they fill slide decks. I have given more than 500 keynotes in 30 countries across five continents, and the AI report card leaders describe to me is remarkably consistent: seats bought, active users, a staff survey saying people like it. None of those tells you whether a single job got better. A team can use AI every day and save nothing, and a team of two can use it on one job and save an afternoon a week.

The second version of the mistake is averaging. An owner hears that AI "saves about an hour a day across the team" and treats that as a result. Averages bury the one job where AI transformed the work alongside the three where it quietly made things worse. Measure job by job, and keep the losers on the list, because they tell you where the brief, the training or the checking needs work.

This is why evidence is one of the Seven Dimensions of AI Readiness, and why that article asks a single question of it: could you show me one number that has changed because of AI? It is also why, when a business appoints an AI steward under the Steward Rule, one of the steward's three duties is to keep the evidence of what AI is saving. Evidence does not appear on its own. Someone has to keep it.

What If the Numbers Say AI Is Not Working?

If the numbers say AI is not working in a job, you have learned something valuable cheaply, and you have three sensible moves: fix the brief, fix the check, or drop the job.

Most failures I see trace to one of those. The brief was too thin, so the drafts needed heavy correction. The person checking the work did not know what to look for, so errors got through and came back later. Or the job was simply a poor fit: too varied, too dependent on judgment, too rare to be worth changing. The fourth question of the Freewheel Diagnostic exists for exactly this moment: "Does someone check whether the result is better?"

There is a fair counterargument. Some of AI's value does not show up in minutes: a better proposal that wins a contract, a manager who finally has time to think, a junior employee who learns faster. That is true, and the Before Number Rule does not pretend otherwise. But those gains still have a destination you can write down in advance and look for afterwards. "We will submit two more tenders a quarter" is measurable. "We will be more strategic" is not.

Dropping a job is not a failure. A business that tests three jobs, keeps two and drops one knows more than a business that rolled AI out everywhere and cannot say what it did.

What to Do This Week

You can have your first before number by Friday without buying anything or starting a project.

  • Monday: choose one recurring job where AI is already being used or is about to be. Write it in one line and name the person who does it.
  • Tuesday to Thursday: that person times every instance of the job and tallies anything that comes back for correction. If AI is already in the job, time it anyway: this becomes your baseline for improving the brief.
  • Thursday: agree the destination sentence together. What should the saved time go to?
  • Friday: put a date in the diary four weeks from now to measure the same three numbers again, and tell the team the results will be shared, whatever they show.

Ask your team one question at your next meeting: "Which job has AI made better, and how do you know?" The second half of that question is the one that matters.

Find Out Where Your Business Stands

If you are not sure whether AI is creating value across your business or only in a few people's inboxes, start with the free Business AI Readiness Scorecard. It takes about five minutes, scores six areas including value and improvement, and gives you a 90-day plan. If your leadership team wants help deciding which jobs to measure and what counts as success, that is the work I do with leadership teams through AI strategy for business. Organizations I have worked with include Google, Grammarly, Kahoot!, Welsh Water and Hyve Group.

Sources and further reading

Dan Fitzpatrick helps businesses use AI well through keynotes, practical AI training for teams, and AI strategy and governance for leadership teams. More about Dan.

Key takeaways

  • Most firms using AI cannot show it is working: in a February 2026 NBER survey of nearly 6,000 executives, nine in ten reported no productivity or employment impact at their own firm over three years.
  • How AI feels is a poor guide: in METR's 2025 study, experienced developers believed AI sped them up by 20% when the clock showed they took 19% longer.
  • The Before Number Rule measures a job, not a tool, using three numbers taken before AI goes in and again after four weeks: time, corrections and where the saved hours go.
  • Time a whole job from start to finished and sent, because a fast first draft that needs heavy fixing saves nothing.
  • Saved time that nobody directs goes into more of the same work, so write a destination sentence before you start.
  • Counting licenses, logins and active users measures adoption, not value, and averages across a team hide the jobs where AI made things worse.
  • If a job fails the measurement, fix the brief, fix the check or drop the job; dropping one is a cheap and useful result.

Frequently Asked Questions

How do I know if AI is actually saving my business time?

Measure one named job before and after AI. Time five to ten real instances the old way, count how often the work comes back for correction, then repeat both after four weeks of using AI. If the time fell and corrections did not rise, the saving is real rather than a feeling.

What should a small business measure to see if AI is working?

Measure three things on one job: how long it takes from start to finished, how often the work is returned for correction, and where the saved hours go. Licenses, logins and staff enthusiasm measure adoption, which is not the same as value, so leave them out of the verdict.

How long should I wait before measuring the impact of AI?

Wait about four weeks, and ignore the first fortnight. People are slower in the first days of any new way of working, so early numbers measure the learning rather than the job. Measure in weeks three and four with the same person and the same definition of finished.

Do I need a data team or software to measure AI results?

No. A notepad, a phone timer and one honest week are enough for a small business. The person doing the job keeps the numbers, knows they will not be used against them, and records the middle time of several instances rather than the best one.

Why do most businesses see no measurable productivity gain from AI yet?

Usually because they never tooka before number and never decided where saved time should go. A February 2026 NBER survey found most firms use AI, yet nine in ten executives reported no productivity or employment impact at their own firm so far.

What should I do if AI is not improving a job?

Fix the brief, fix the check, or drop the job. Thin instructions cause heavy correction, a checker who does not know what to look for lets errors through, and some jobs are simply a poor fit. Dropping one job after a four-week test is a useful, inexpensive result.

Dan Fitzpatrick helps businesses use AI well through keynotes, practical AI training for teams, and AI strategy and governance for leadership teams.

D
Dan Fitzpatrick

Delivered training to 150K+ educators | Founder of The AI Educator and AI Educator Tools | Forbes Contributor | International Keynote Speaker | 4 x #1 Bestselling Author