You booked because talk is cheap.
This room exists so the call isn't spent explaining. Everything below is checkable — the studies are named, dated, and linked, and where a study limits itself, the limit travels with the number. Read what you want. Skip what you don't. Twenty-five minutes is enough time to decide something only if the deciding is the only thing left to do.
Buyers are asking machines about you. The machines are answering without you.
You know the behavior has shifted. What you may not know is how badly the answers hold up.
In March 2025, the Tow Center for Digital Journalism ran eight AI search engines through 1,600 queries — 200 per engine. Each was handed an excerpt from a real article and asked to identify it: headline, publisher, date, URL. Not a trick. The source material was in front of them.
More than 60% of the answers were wrong.
The spread is worth sitting with. Perplexity, the most accurate of the eight, missed 37% of the time. ChatGPT Search missed 67% — 134 of its 200 — and in that entire run it never once declined to answer. It signalled uncertainty fifteen times. DeepSeek misattributed the source in 115 of 200 responses. Grok-3 produced 154 citations that led to error pages.
These are systems performing a task where the correct answer was supplied to them.
What this study is, and is not. It measures source identification under one protocol, on February 2025 model versions, with one observed response per excerpt. It is not a general hallucination rate and not a current ranking. We say that plainly because the authors said it first, and because a firm that overstates its evidence has no business selling verification.
Now read the limitation the authors put on their own work, because it is the most important sentence in the paper: rerunning the same prompts would have a high chance of producing different outputs.
That is the finding. Not that the machines are wrong — that a single run cannot tell you whether they are. The most rigorous public study in this space says its own numbers would not reproduce.
Every dashboard selling you a visibility score ran fewer trials than that.
Why we're alone in this.
Everyone built the dashboard first.
The tools measuring your AI visibility were built to produce a number, and the number was built to be sold. Proof came later, or never, because proof was not the thing anyone was buying. To add it now, a vendor would have to rebuild what they are — a witness that does not grade, a record that cannot be altered, a signature that software cannot produce. None of that bolts on afterward.
We built the governance first and the product second. That order is not a virtue we're claiming. It's a constraint we accepted, and it is the only reason the next few paragraphs can be checked rather than believed.
The industry has already run this experiment once.
In 2016 the Association of National Advertisers — the advertisers, not the agencies — commissioned an independent investigation into media buying. K2 controlled the methodology, the sources, and the findings; source identities were never disclosed to ANA. Of 117 interviewees with direct experience of US media buying, 59 described non-transparent practices. Forty-one described rebate arrangements. In the principal transactions examined, markups ran roughly 30% to 90%.
The investigation did not calculate a total dollar loss, and we will not invent one.
What it produced instead was a diagnosis: the party placing the spend cannot be the party verifying it. The industry's answer was structural — contractual audit rights, disclosure, and clarity on whether your agency was acting as your agent or as a principal selling to you.
Seven years later the ANA measured what remained, using log-level data across 21 brands, 11 categories, and $123 million in spend. Of every dollar entering a demand-side platform, 36 cents reached the consumer. The modeled industry opportunity was $22 billion — an efficiency figure, not a proven loss, and we hold it as the modeled number it is.
Then the remedy was applied. In the following benchmark, among participating companies, made-for-advertising spend fell from 15% to 4%, and the average number of sites and apps fell from roughly 44,000 to 23,000.
The problem was real. The fix was structural. It worked.
It is happening again, on a surface where the fix does not exist yet.
And this is not only a marketing question.
The remedy that emerged in advertising — the Media Rating Council — does not accept a measurement provider's word about its own accuracy. Accredited services submit to full methodological disclosure, compliance with minimum standards, annual external audit by independent CPA teams, and committee review for continued accreditation.
That model is worth naming precisely, because it makes our claim narrower and more honest than the market's: independence there is constructed through disclosure, standards, and external audit — and the audited party pays the audit costs. Independence is a matter of structure and separation of function, not the absence of any commercial relationship. We hold ourselves to the same standard, and we tell you exactly which guarantees are enforced in machinery and which are certified by procedure.
As of 4 August 2026, we checked whether such a standard exists for AI-assistant brand visibility.
We checked the Media Rating Council, the IAB and IAB Tech Lab, ISO/IEC, NIST, the ANA, the WFA, the FTC, and the UK's ASA/CAP.
We found real and adjacent work. The MRC issued guidance on AI use in media measurement on 8 July 2026 and has phased standards work running into early 2027. The IAB's Project Eidos is releasing draft frameworks through 2026. IAB Tech Lab finalized a content-monetization protocol in April. ISO/IEC has published 42001, 42005 and 42006. NIST maintains its generative-AI risk profile.
What we did not find, in any of them, was a completed standard defining all of it together: the metric unit, the prompt sampling frame, model and version controls, repeat-run variance, and an accreditation procedure.
That is a scoped finding, dated, naming the bodies checked. It is not a claim that nothing exists anywhere, and the work in progress means it has a shelf life.
But today there is no accredited answer, and the market is selling numbers as though there were.
How the exam actually runs — three layers, no trust required.
Layer one — the executors. The agents that query surfaces and capture what comes back. They are collared: bounded in what they may do, what they may conclude, and what they may leave unsaid. They gather. They do not grade.
Layer two — the witness. A separate system that records what happened, hash-chained and signed. It observes the executors; it has no stake in what they find. The chain means an entry cannot be altered without breaking every entry after it. Nobody at JFD can quietly improve your result, including us.
Layer three — the hardware. A secure element — physical silicon holding a private key that never leaves it. Signatures are produced inside the chip. Software can request one; software cannot forge one. This is the floor under everything above it, and it is the layer that cannot be replicated by a competitor's next release.
The five instrument articles
For the reader who forwards this to their analyst.
MDE before probe. The minimum detectable effect is fixed before any surface is queried. We state in advance how large a difference the exam is capable of seeing. A result inside that floor is reported as inside the floor — never dressed as a finding.
Missingness is an outcome. When a surface returns nothing, that is data, and it is recorded as data. It is never dropped, never treated as a failed run, never quietly re-queried until it produces something. The most common way a visibility number gets inflated is by discarding the runs that didn't work.
The evidence packet. Every claim in your report carries the captured material it rests on — the query, the response, the timestamp, the model version. You are not asked to trust the summary. You are handed what the summary was made from.
The identity ladder. Every observation records how it was obtained and how strongly that method establishes it. A directly captured response and an inference from adjacent evidence are never rendered as the same thing, because they are not the same thing.
Consensus dedup. Repeated observations across runs are reconciled by an explicit compare step, not collapsed by counting. Two runs that agree and two runs that merely look alike produce different records.
The questions serious buyers ask.
Everyone else runs the query and reports the number. We lock the method before the query, record the run in a system that can't revise it, and sign the record in silicon. The difference is not the querying. It's that our answer arrives with the conditions of its own production attached, which means you can check it instead of believing it.
Because it isn't a feature. A witness that doesn't grade, an append-only signed record, and a hardware root of trust are three architectural commitments that have to precede the product. A vendor who built a dashboard first would have to rebuild what they are to add them. That's not a moat we dug. It's the shape of the order we happened to build in.
A signed verdict on the named decision you brought, the evidence packet every claim in it rests on, the pre-registration you countersigned before we looked, and the run record. Enough that your analyst can reconstruct our reasoning, and enough that your board can see what was fixed in advance.
Because every other structure pays us to reach a particular conclusion. A percentage of savings pays us to find waste. A retainer pays us to find ongoing work. A credit toward the next engagement pays us to find a next engagement. The fee is fixed and identical whichever way the verdict goes, and that is the only fee structure that leaves the finding alone.
Then the report says so, you pay the same amount, and you have something you did not have before: a signed, dated, independently produced baseline showing your position was sound. That is a real asset the next time someone in your organization proposes spending against a problem that isn't there. A firm that can't afford to return a clean verdict will never return one.
The oldest fix for a measurement nobody can trust.
Your pre-registration locks six things before a single surface is queried: the named decision, the asset perimeter, the competitor set, the prompt frame, the surface matrix, and the CERC anchors. You countersign it. Then we run.
This sequence is not our invention. It is the standard remedy in every field that has had to survive the accusation of finding what it went looking for.
Medicine. The International Committee of Medical Journal Editors requires a trial to be registered in a public registry at or before the first patient's consent, as a condition of being considered for publication. The World Health Organization treats prospective registration as a scientific and ethical responsibility — registered before the first subject is recruited. The stated reason in both cases is selective reporting: the temptation to decide what you were measuring after you've seen the results.
What changed when it was applied. A study of 55 large NHLBI-funded cardiovascular trials found that before prospective registration became standard, 17 of 30 trials — 57% — reported a significant benefit on the primary outcome. After, 2 of 25 — 8%. The authors are explicit that causality cannot be inferred; trial design and clinical practice changed over the same period, and that limitation travels with the number.
Psychology, cleaner. Registered Reports receive in-principle acceptance before results are known. Comparing 71 of them against a random sample of 152 conventional psychology articles: 96% of the conventional studies reported positive results. Among the Registered Reports, 44%.
Same field. Same rigor. Fifty-two points of difference, produced by nothing but the order of operations.
Assurance, where it becomes law. The SEC's auditor-independence doctrine names four structural threats: conflicting interest, auditing one's own work, acting as management, and acting as advocate for the client. The prohibited-services list exists to prevent the second. Internationally, IESBA bars non-assurance services to audit clients where a self-review threat arises. ISAE 3000 requires an assurance engagement to have a rational purpose, suitable criteria, and independence preconditions — established before the engagement, not asserted after.
Four disciplines — medicine, psychology, securities law, advertising — arrived independently at the same three controls: lock the method before you look, log what actually happened, and keep the interested party out of the grading.
We did not design a clever process. We implemented the one that already won everywhere else, in a market that has not yet adopted it.
The transfer is an institutional analogy and we mark it as one. Nobody has validated pre-registration inside AI-discovery measurement specifically, because nobody has run it there.
Your exam is where that starts.
The full technical grounding.
Not required reading. The call will be sharper if you've seen it.
Your 25 minutes.
No deck. No pitch.
We walk the specific decision your audit would inform, the six parameter classes we'd lock into your pre-registration, and whether the 3× test is met.
If it isn't met, we'll say so on the call.
You leave knowing exactly what the exam would cost, what it would measure, and what it would prove — whether or not you commission it.
The three sentences that cost us money.
The fee is fixed, and it is the same whichever way the verdict goes. Keep, cut, shift, do nothing — identical invoice. Not a percentage of what we find you can save, because that structure pays us to find waste. Not contingent on anything.
Nothing credits. The exam fee does not reduce the price of any subsequent engagement. A credit turns an audit into a sales funnel, and everyone downstream can tell.
Any remediation is a separate decision made after your verdict is signed — and we may decline it. We review what the exam found and elect whether to take the work on or refer it out. That election happens after a permanent signature exists, which means no finding can be shaped toward work nobody had on the table when it was signed.
If we do take remediation on, we tell you exactly what that changes: performance reporting on our own build is delivery accountability, and we will never sell it to you as independent verification. The independence claim attaches to the exam that sets your baseline. It does not launder anything that comes after.
Excellence in Governance — not our slogan. Our architecture, described.
Check it before we speak.
Everything above is checkable without taking our word for any of it. These are the artifacts, not a description of them.
The sealed question set and method — the questions committed cryptographically and externally timestamped before the first answer was collected. An error bar is a disclosure about the output; pre-registration is a constraint on the procedure. A margin of error can be authored after the fact. A timestamp that precedes the data cannot.
The receipts — every figure on the findings page rests on captured material: the query as issued, the response as returned, the timestamp, the engine. The index is public and each row verifies offline.
The findings, including our own — we ran the exam on ourselves and published the null. We executed a remedy on our own property, re-measured under the identical sealed method, and the number did not move: 0 of 144, before and after. We finished joint last in our own cohort of thirteen, sealed before we looked, published with no asterisk. Then we applied our own repetition-adequacy standard to our own published table and not one adjacent rank survived it — and we published that invalidation alongside the table.
The log — what changed, when, and why.
The boundary, said first rather than discovered: we can prove the diagnosis. We cannot yet prove the remedy moves anything — and neither can anyone selling you one. The difference is that we will tell you which way it went, in writing, after you have paid us. Every reading on this site is bounded to one engine: Perplexity.