Evaluation · a layer of Growth Engineering
Three times the output. The same growth.
More output is not growth, and output nobody checks is a liability. Evaluation checks every piece before it ships, and gives every decision a verdict afterwards against what it was supposed to do.
The trap
Time saved gets reported as return. Gartner's research on AI in marketing finds time savings rarely convert into business results, and that 45% of martech leaders say vendor AI agents fall short of the business performance they were promised. Meanwhile output volume keeps climbing, which is the easiest number to show and the least connected to growth.
Why more can mean less
In a market where machines assemble the shortlist, work that repeats what is already out there does not earn selection. It adds to the pile the engines average over. The unit of progress is not pieces shipped; it is whether the odds of being chosen moved.
The measurement problem underneath
Buyers are roughly 70% of the way to a decision before they contact a vendor, so the deciding moments happen where you cannot see them. We report honest signals, referrers from AI tools, branded search lift, how people say they heard about you, and never invent a revenue figure the data cannot support.
It has already been tested in public
A tribunal held Air Canada responsible for what its chatbot told a customer. Deloitte Australia refunded part of a fee after an AI-assisted report was found to contain fabricated citations. The rule that is forming is simple: it is your output, whatever produced it.
Rules are arriving with dates
The EU AI Act's transparency duties apply from 2 August 2026, with penalties up to 3% of global turnover, and the US Copyright Office has stated that prompts alone do not make you the author of what a model produces. Both change how marketing output must be handled and recorded.
What most teams have
A policy document, and no mechanism. Nothing sits between a confident draft and a published page, and nobody can reconstruct later who approved which claim on what evidence.
Before it ships
Every piece is checked for voice, for claims that cannot be proven, for whether it adds anything the already-cited sources do not, and for whether a machine can lift a clear answer out of it. Failing means it does not go out.
Refusal is the rare part
The market ships generation. Almost nothing on sale will stop a piece of work. A check that cannot fail is decoration, so ours ship with tests that prove they can.
After it ships
Each decision recorded what it expected and when to look again. When that date arrives it gets a verdict: right, wrong, or not enough evidence to say. Thin evidence closes nothing, which keeps the learning honest.
One number, agreed at the start
One number is agreed at the start, with a date to review it. Every piece of work exists to move it, each decision carries what it expects to achieve and what would make us stop, and the result is written down against the expectation. If the number has not moved by the review date, you can walk.
What you see
The number, the pace against it, and the evidence behind each claim. When a week teaches nothing, the report says so rather than dressing up activity as progress.
Everything we have written on this
- Agentic Marketing Platforms in 2026: Scored, Not Ranked
- Best AI Visibility Tools in 2026: An Honest List with Real Prices
- Profound vs Ahrefs Brand Radar: Why They Disagree
- Switch From Manual ChatGPT Checking to an Automated Panel
- Semrush Alternatives for AI Visibility: The Question Behind the Question
- Profound Alternatives (2026): Tools That Track AI Visibility, and What None of Them Measure
- AirOps Alternatives (2026): AEO Content Platforms Compared
- Jasper Alternatives (2026): AI Writing Tools Compared — and the Problem None of Them Solve
- The Build-vs-Buy Ledger: AEO Agency vs In-House Team vs AI Growth Operator, Priced
- Best AI Marketing Agencies for Startups in 2026: The Checkable List
- AI Detectors Measure the Wrong Thing
- The Trust Wall: Why We Measure Ourselves First
Where this layer is put to work
- Get recommended by AI: Reads how often each engine names you, as a trend across repeated runs, never a single screenshot.
- Demand creation: Reads what moved: replies, saves, branded search, and whether AI starts naming you.
- Outbound pipeline: Counts replies and meetings, not messages sent.
- Social audience: Measures after posting and keeps what worked for the next round.
Questions
What number should we pick?
Something a sceptic can check and your finance team already trusts: qualified enquiries, cost per customer, signups. We agree it before any work starts.
How do you handle attribution when nobody clicks?
We instrument what can be measured honestly and label the rest as unmeasured. A confident number nobody can verify costs more credibility than an honest gap.
What if the work does not move it?
You hear it from us first, with the evidence, and there is no lock-in past the review date.
Do we have to disclose AI use?
Increasingly yes, depending on where you operate, and the safer default is a clear record of what was produced how, kept before a regulator or a customer asks.
Can checks block real work?
They stop a piece three times at most, then the decision comes to you with the options: get the missing evidence, re-brief it, or drop it.
Who owns AI-generated content?
Authorship is not automatic from prompting, so the human contribution and the record of it matter. We keep that record as a matter of course.
What happens when a piece is refused three times?
It stops asking for another draft and comes to you with three options: get the missing evidence, re-brief it, or drop it.
Can we see the checks?
Yes. The standard is written down, and so is every exception to it.