36nine

How to know if your AI is actually working

Most AI that gets called a success was never measured against anything that matters. It demoed well, it went live, and no one checked whether the number leadership cares about actually moved. This is a plain account of how to tell a working AI system from a busy one: the one number to judge it by, the metrics that flatter but do not count, and what honest measurement looks like when you run these systems in production.

A minimal brass measurement dial on a deep teal field with faint graphite tick marks and a single needle resting on one mark lit in pale teal, evoking an AI system judged against one real number.

Start from the number, not the tool

The question is not whether your AI is impressive. It is whether it moved a number a person in your business already watches. Those are different tests, and most deployments only ever pass the first. A system can answer fluently, sound intelligent, and light up a dashboard while the figure leadership actually reports to the board, response time, conversion, cost per case, sits exactly where it was before the project started.

So the measurement starts before the build, not after it. Pick the single business number the system is meant to move, write it down, and record where it stands today. Everything after that is judged against that baseline. Without it you cannot tell success from motion, and motion is what stalled AI produces in abundance.

"It's live" is not "it's working"

Going live feels like the finish line. It is not. A deployed system that no one measures against a real outcome is indistinguishable from a stalled one, and this is exactly where most AI quietly fails. MIT's State of AI in Business 2025 study found that about 95 percent of enterprise generative AI pilots never reach measurable impact on the P&L. The cause it identified was not weak models. It was an integration and measurement gap: systems that ran but were never wired to a number anyone acted on.

The tell is simple. Ask what business figure has changed since the system went live, and by how much. If the answer is a description of what the AI does rather than a number that moved, the system is live but not yet working, and a serious partner will not let you confuse the two.

The metrics that flatter but do not count

Vanity metrics are the ones that go up whether or not the business improves. Messages handled, questions answered, hours of conversation, tokens processed, uptime: all real, all easy to report, none of them proof that the system did its job. A chatbot can handle ten thousand messages and enrol no one. A dashboard full of activity is the most common disguise a failing deployment wears, because it looks like progress and asks nothing further of anyone.

The counter-question strips it back every time: which of these numbers would a head of admissions, a registrar, or a finance director have chosen to watch on their own, before AI entered the conversation? Those are the numbers that count. The rest are instrumentation, useful for debugging the system, useless for judging it.

Three layers worth watching

A working measurement setup usually has three layers, and only one of them is the verdict. Leading indicators are the operational signals that tell you the system is behaving: how fast it responds, how often it completes a task without a human, how many conversations it carries to a booked next step. These are early and fast, and they are how you tune.

The outcome number is the verdict: the single business figure the deployment exists to move. In an admissions operation that is usually one figure in the enrolment funnel: the response time that decides who enrols before the school down the road does, or the share of inquiries that convert to a booked tour. You watch the leading indicators daily to tune the system; you judge the whole engagement against the outcome number. A team that reports a fast response time while enrolments sit flat has measured the wrong layer.

What honest measurement looks like

Honest measurement means being willing to report the number when it is not flattering, and then rebuilding until it is. One voice agent 36nine built connected on fewer than one in ten calls when it first went live. That is not a number most vendors would put in front of a client. We measured it, said it plainly, and rebuilt the system until it cleared better than one in three. The value was made in that loop, not in the original build, and it only existed because the real number was being tracked from day one.

That is the shape of it: ship a narrow system into real production, measure it against the number you named up front, and rebuild what falls short. A partner who only ever shows you improving charts is either lucky or hiding the loop. The honest version has a number that starts low, is stated anyway, and climbs because someone kept measuring it.

What a leader should do next

Run one test on any AI you already have, or any a vendor is pitching. Name the single business number it is meant to move. Ask what that number was before the system, and what it is now. If no one can answer, the system is not being measured, and a system that is not measured is not yet working, whatever the dashboard says.

From there the path is narrow. Scope one real number, build the agent that moves it, run it in production, and judge it against the figure you started with. That is the whole test, and it is the same test we hold our own systems to. A strategy session with 36nine names that first number and the plan behind it.

Common questions

How do you measure if an AI system is actually working?
Judge it against a single business number that someone in your organisation already watches, response time to a first inquiry, conversion, cost per case, and compare where that number stood before the system to where it stands now. If the only evidence of success is a description of what the AI does, or a dashboard of activity, the system is live but not yet proven to work.
What is the difference between a vanity metric and a real AI metric?
A vanity metric goes up whether or not the business improves: messages handled, questions answered, hours of conversation, uptime. A real metric is the business outcome the deployment exists to move, the one a leader would have chosen to watch before AI entered the conversation. A chatbot can handle ten thousand messages and enrol no one; activity is not impact.
Why do most AI projects fail to show ROI?
MIT's State of AI in Business 2025 study found that about 95 percent of enterprise generative AI pilots never reach measurable P&L impact, and traced the cause to an integration and measurement gap rather than weak models. Systems went live but were never wired to a number anyone acted on, so no one could tell a working deployment from a busy one.
What metrics should we track for an AI deployment?
Track two layers for different jobs. Leading indicators, response speed, task completion without a human, conversations carried to a booked next step, tell you the system is behaving and are how you tune it daily. The outcome number, the single business figure the deployment exists to move, is the verdict you judge the whole engagement against. Do not mistake a good leading indicator for the outcome.
How soon should an AI deployment show a measurable result?
Within the first quarter, provided it was scoped to one number and shipped into real production early. The value is usually made in the rebuild loop: ship narrow, measure against the number you named up front, and rebuild what falls short. One voice agent 36nine built started below one in ten calls answered and was rebuilt until it cleared better than one in three, because the real number was tracked from day one.
What should we ask a vendor to prove their AI works?
Ask which single business number their system moved for a comparable client, what it was before, and what it became. Ask to see the number when it was not flattering and how they closed the gap. A partner who only ever shows improving charts is hiding the rebuild loop; an honest one can name a figure that started low, was reported anyway, and climbed because someone kept measuring it.

See where AI agents create measurable impact for your team.