Measure output that reaches a customer, not activity. Drafts produced, prompts run, and hours saved are unfalsifiable and therefore useless. Things published, replies received, and conversions attributed are checkable. And be careful with your visitor numbers — most tools report a sum of daily uniques, which is not a count of people.
The metrics that mean nothing
| Reported | Problem |
|---|---|
| “Saved 10 hours a week” | Unfalsifiable, and saved from what baseline? |
| Drafts generated | Production is not distribution |
| Prompts run | Measures usage of a tool, not value |
| Content volume | More pages is not more demand |
| Open rate | Increasingly inflated by privacy proxies |
The pattern: each measures the top of your own funnel rather than anything a customer did. They rise whenever you use the tool more, which makes them a measure of adoption masquerading as a measure of return.
“Hours saved” deserves special scepticism. It is not a business outcome until it becomes something downstream: more published work, faster response, or a task that previously did not happen at all. If the hours saved went into more meetings, the correct reported value is zero.
What to measure instead
- Things that reached a customer. Published, sent, answered. Production that never shipped is not output.
- Response time. Often the single biggest effect of automation, and easy to measure precisely. Inbound leads are lost to slowness far more than to bad conversations.
- Reply rate, not open rate. A reply is a human decision. An open is increasingly a proxy server.
- Tasks that used to be skipped. The most honest AI win is work that simply did not happen before — the fifth follow-up, the write-up after the call.
- Error rate. The one nobody reports. If AI doubled output and 44% of it was wrong, that is not a win, and you only know if you check.
That last point is not hypothetical. In an enrichment pipeline I measured, roughly 44% of records came back attached to the wrong employer, and every one looked correct. Volume metrics would have shown a triumph.
Attribution that survives cookie loss
Pixel-based measurement reports a shrinking sample and does not tell you it is shrinking. That is the important part: the failure is silent, so your dashboard looks healthy while covering less of reality each quarter.
What holds up:
- First-party measurement. Served from your own domain, so it is not blocked as a third party.
- A durable identifier you own, rather than one borrowed from an ad platform.
- Click identifiers captured on landing and carried through to conversion, so paid traffic can be tied to outcomes without depending on the ad platform’s own cookie.
- Server-side conversion recording where possible, since it does not depend on a browser cooperating.
- UTM discipline. Unglamorous, and the difference between a channel report and a guess.
One honest caveat about durable identifiers: if you create one, say so in your privacy policy. A measurement design that quietly persists an identity is a policy question, not just an engineering one.
The visitor-counting trap
Worth its own section because it silently inflates almost every report I see.
Many analytics tools report a sum of daily unique visitors. If identities rotate at midnight, one person visiting on three days counts as three. Over a week, a small loyal audience can look like a much larger one.
That is a perfectly reasonable traffic metric. It is not a count of humans, and the two get used interchangeably constantly. “We reached 500 people last month” from a tool reporting daily uniques is a claim you cannot support.
Check what your tool actually counts before any number goes in a deck. If it does not distinguish daily uniques from distinct people, do not describe the figure as people.
Set the baseline before you start
The most common measurement failure with AI is having nothing to compare against.
Before introducing a tool, record: how long a task takes now, how many get done per week, what the current response time is, what the current error rate is. Ten minutes of work, and without it every later claim is a story.
The same applies to publishing. Instrument before you publish, not after — data only accumulates once things are live, and the questions you will want to answer in three months depend on measurement that existed in month one.
Be honest about small numbers
A discipline worth adopting, and one I have had to apply to my own work.
I intended to write a study of traffic patterns across my whole portfolio of sites, and when I actually looked, the analytics held about two days of usable history across those properties. That cannot support a sentence like “a year of data across twenty sites,” so I did not write it.
The equivalent restraint in a marketing report: a 40% improvement on a base of five is noise, and reporting it as a percentage is a choice to obscure that. State the raw numbers alongside every rate. If the raw number is embarrassing, that is information too.
FAQ
How do you measure whether AI is helping?
By output that reached a customer, not activity. Drafts and hours saved are unfalsifiable; published work, replies, and attributed conversions are not.
Why is attribution harder now?
Third-party pixels are increasingly blocked, so they report a shrinking sample silently. First-party measurement with an identifier you own survives it.
Visitors versus people?
Most tools sum daily uniques, so one person over three days counts as three. Useful traffic figure, not a count of humans.
What if AI mostly saved time?
Measure what the time became. Saved hours are not an outcome until they turn into published work, faster response, or something new.
Related: an AI marketing stack for a small team and how to use AI for SEO.