Mission Log

Mission Log

Last modified 2026-08-17

The open letter to the IAB

Published, not sent.

Exhibit A is published below, inside Entry 001 §VII. It has not been sent to the IAB and no reply has been solicited. It publishes here first, in the open, because a letter that exists only in someone’s inbox can be received, thanked and buried at no cost to anyone. A published record stays checkable whether or not it is ever answered.

If the Project Eidos working group reads it and disagrees, we would rather hear that than not.

The position the letter will carry is already public and is set out in the first entry below. It does not depend on the letter being answered, or on it being sent at all.

The journey

Everything beyond the exam itself: what we are building, what we got wrong, and what changed because of it. Newest first. Each entry stands on its own and nothing here is edited after the fact — corrections arrive as new entries, dated, the way they do everywhere else on this property.

2026-08-16

Entry 001 — What We Thought We Had, and What Was Left

By K.L. Phillips · Jet Fyul Dynamics LLC · 2026-08-16


We publish this because a company that sells sealed measurement cannot keep a private list of the things it got wrong. The list is the product. So here is ours, for the month of August 2026, in the order the claims died.

The first claim was that nobody publishes their instrument’s instability. We carried it for three weeks as the differentiator. It was false when we started carrying it. Vendors publish variability today, and as of August 2026 the IAB’s own framework requires it — programs must document the variability observed across repeated identical queries and make it visible in reporting, and report a range rather than a single number. We did not discover this by reading the market. We discovered it by finally reading the document.

The second claim was a number: that AI answers account for two to four percent of product discovery. We used it to size the opportunity. When we went to source it, no methodology-consistent series existed behind it. Different studies, different denominators, welded into one figure by repetition. We killed it. Every value figure that survived in our own materials is marked as an estimate on its face, because that is what it is.

The third claim was that we could see a category forming. We reported it as a finding. It was a pattern we had produced, not one the engine had. We retracted it and withdrew the frame built on it from the offer.

The fourth claim was one word: “publicly-anchored.” It described our sealed readings as anchored outside our own systems. Nothing anchored them outside our own systems. The word was live in the generator and in four dossiers. We struck it, and then we went and built the thing the word had promised — the pre-registration hash is now timestamped to four independent calendars, and the Bitcoin commitment has since confirmed — block 962642.


The fifth claim is the one that produced this entry, and it is the most instructive, because our own gate caught it before it reached anyone.

We had written that the IAB’s framework requires variability be disclosed but sets no numerical threshold — the requirement with no number, the policy any vendor can satisfy while telling the buyer nothing. It was going to be the spine of a letter and of the exhibit that follows this entry.

Before publishing, we required a verbatim confirmation of every load-bearing sentence against the complete document, with a standing instruction: if any threshold, floor or recommended minimum exists anywhere in it, stop and publish nothing.

One does. Several do. The framework sets a fifty-query floor and calls it a floor. Its criteria matrix is explicitly a set of minimum thresholds. It fixes a seven-day window for reproducibility.

So we stopped, and nothing published.

The narrower claim survives, and it is stronger than what we wrote, because the framework states it itself: quantitative thresholds for acceptable variability remain, in its own words, an open working group question. And the sharpest version is one we had not seen at all — for decision-grade reproducibility the framework fixes the window and requires confidence levels, but leaves the acceptable variation range to the provider to define. A provider defining its own acceptable range, after seeing its own results, is the precise mechanism our whole method exists to remove.

We also checked whether the framework contains the concept of pre-registration. The term appears nowhere across all thirty-six pages. But an ex-ante requirement does exist — inclusion criteria are to be documented in advance rather than settled by ad hoc judgment. It is scoped to which platforms a program covers, and it is never extended to the measurement method itself. That is a more interesting fact than absence, and it is the one we will argue from.


What was left

Five claims died this month. Every one of them was something we said about the market or about ourselves. Not one was a thing the instrument measured.

That distinction is the entry.

What survived the month is the part nobody has to take our word for. The method and the exact question set were committed to an append-only ledger and hardware-signed before a single answer was collected — sealed at 05:19:36 UTC on 2026-08-13, nineteen minutes and nine seconds before the first capture landed. The question set is published. The pass count is fixed and each pass carries its own recorded window. Every reading names its engine, its frame, and its denominator. The instrument refuses a coverage verdict on a floor it cannot calibrate, and it refuses a prescription line that does not trace to a component that actually discriminates — in our own runs it refused between sixty-nine and one hundred percent of the lines we wanted.

And when we ran the remedy loop on ourselves and it produced nothing, we published the null.

We also scored zero. Asked thirty-six buying questions across three frames, four passes each, one hundred forty-four captures, in the category we sell measurement into, our own name was returned exactly zero times. We published that at joint-last with no asterisk. The honest answer to anyone who raises it is not a defence of the score. It is the seal and the method — which is the only answer we have ever been selling.

The boundary, stated first

We can prove the diagnosis. We cannot yet prove that any remedy moves it. Neither can anyone selling one; they do not say so. What we will do is measure again on the same sealed method after you engage us, and tell you in writing which way it went — up, down, or not at all.

VII. Exhibit

The argument this entry stumbled into is set out separately, in Exhibit A — For the Record: On a Requirement Without a Threshold. It is published, not sent. If the IAB’s Project Eidos working group reads it and disagrees, we would rather hear that than not.

Entry 001 · published 2026-08-16 · K.L. Phillips · Jet Fyul Dynamics LLC · nothing in this entry is new evidence; every measurement referenced is already on the receipts index

Open Letter and Commentary to the IAB · attached to Entry 001 §VII

EXHIBIT A — FOR THE RECORD

On a Requirement Without a Threshold

By K.L. Phillips · Jet Fyul Dynamics LLC · 2026-08-16 Published, not sent.


We are not members of the IAB, and we have no standing in Project Eidos. We are writing because the document is loud, because it will shape what buyers in this category are told for the next several years, and because we read all thirty-six pages of it before forming a view — which, from the evidence of how it is being characterised, is not universal.

Nothing below is a proposal. It is a record of what we found and what we do.


I. WHAT THE FRAMEWORK GOT RIGHT

The framework states that a single response to a single query is not measurement — that a brand’s visibility on a query is a distribution rather than a value, and a metric drawn from one response is a sample of one. It requires programs to document the variability observed across repeated identical queries within a defined window, expose that variability in reporting, and report results as a range rather than a single number.

That is correct, and it is more than this market was doing unprompted. Everything that follows takes it as the starting point rather than the target.


II. WHERE WE STAND RELATIVE TO IT, STATED FIRST

Our current instance runs thirty-six distinct questions, which is below the fifty-query floor the framework sets. That is a design choice, not an oversight: the floor guards against characterising a category on too few questions, and our design guards against characterising a question on too few responses. Both failures are real, and we built against the second because it is the one the framework leaves to the provider.

We also sell measurement. Everything in this document describes a control we already run and would benefit from seeing adopted, and that interest is stated here rather than left for a reader to find.


III. WHAT WE GOT WRONG FIRST

We had written that the framework requires variability be disclosed while setting no numerical threshold — the requirement with no number. Before publishing, we required verbatim confirmation of every load-bearing sentence against the complete document, with a standing instruction to halt if any threshold, floor or recommended minimum appeared anywhere in it.

Several do. The framework treats fewer than fifty queries per program as exploratory rather than directional, and calls it a floor. Its criteria matrix is presented as a set of minimum thresholds. It fixes a seven-day window for reproducibility.

The sentence was wrong. We halted, published nothing, and are stating the correction here rather than quietly narrowing it. Any argument premised on that silence should be disregarded, including the one we nearly published.


IV. THE GAP IS NOT A SILENCE. IT IS A DELEGATION.

The one number the framework does not set is the one governing whether a measurement can be trusted at all: how much same-query variability is too much.

It says so itself — quantitative thresholds for acceptable variability “remain an open working group question” — because appropriate ranges differ across metrics, platforms and query types. That reasoning is sound. A single cross-category number would be arbitrary.

But the parameter is not left unset. It is assigned. For decision-grade reproducibility, the provider must define acceptable variation ranges within the seven-day window and report confidence levels.

The provider defines the range against which the provider’s own results will be judged, and defines it after those results exist.

That is the finding. Not an omission — a handoff, to the one party with an interest in where the line falls.


V. WHY A DISCLOSURE OBLIGATION CANNOT CLOSE IT

Two providers measure the same brand in the same category.

The first returns an answer set that turns over substantially between passes taken minutes apart. The second is stable. Both document their variability. Both expose it in reporting. Both define an acceptable range — their own — and both report inside it.

Both are compliant. Nothing the buyer receives distinguishes them.

This is not a criticism of disclosure. It is an observation about what class of thing a disclosure is. An error bar is a statement about an output, and it can be composed after the output is known. Every control the framework places on variability is of that kind: retrospective, self-reported, and unverifiable by the reader.

A number would not repair it. A stricter retrospective figure is still a figure the provider applies to its own work after seeing its own results. The counterweight to a retrospective disclosure is not a tighter retrospective disclosure. It is a constraint that binds before the data exists.

The framework identifies the downstream consequence itself: buyers holding conflicting figures from two providers with no way to determine which is more accurate. A compliance tier every provider clears does not resolve that. It documents it.


VI. THE PRINCIPLE IS ALREADY IN THE DOCUMENT, ONCE

The term pre-registration appears nowhere across the thirty-six pages.

The principle does, in one place. Inclusion criteria — including the consumer share threshold a program applies — are to be documented in advance, so that platform additions are governed by stated standards rather than ad hoc judgment.

Advance documentation is treated as the appropriate remedy for a decision that would otherwise be made after the fact by an interested party. It is applied to which platforms a program covers. It is not applied to how the program measures.

We note the distinction and leave it there.


VII. WHAT WE DO INSTEAD

Described because it is our answer to the same problem, not because anyone asked for it. Any provider can run it without reference to us.

The method and the exact question set are committed to an append-only record, hardware-signed and externally timestamped, before the first response is collected. For the readings cited here: sealed at 05:19:36 UTC on 2026-08-13, nineteen minutes and nine seconds before the first capture landed, and timestamped to four independent calendars.

The question set is published in full and frozen with a hash, so anyone can re-run it.

The pass count is fixed in advance and every pass carries its own recorded timestamp, so the measurement window is a recorded fact rather than an assertion.

The artifact is verifiable offline — checkable without contacting us and without trusting us.

A timestamp that precedes the data cannot be authored after the fact. That is the whole of it. It is not a higher standard of candour and should not be read as a degree of the same virtue. It is a different mechanism: a constraint on the procedure rather than a description of the output.

Under it, the threshold question loses most of its urgency. A provider who fixed the method in advance and published the questions can report whatever variability it observes, because the disclosure can no longer be composed to fit the answer.


VIII. OUR OWN RECORD, WHICH IS THE ONLY WARRANT WE HAVE

We measured ourselves with our own instrument and were named in zero of one hundred forty-four captures, in the category we sell into. We published it at joint-last with no asterisk. We ran our own remedy loop, it produced nothing, and we published the null. We score identical sealed captures twice and publish our own instrument’s drift. We publish a repetition-adequacy method that reproduces Evertune’s published table at two of three points and fails at the third — marked unreconciled, with an invitation to correct us. We cite them because they published, and a provider who publishes is doing the thing this document argues for. Reproducing a published table is not an audit of the instrument behind it, and we have audited no one’s.

Five claims we had made about this market died in the month before this was written. Four were killed before anyone outside this company saw them. The fifth is the one in Section III, and it died at our own verification gate, which halted this document over a sentence its subject would have contested.

Our measurements are on one engine, in three categories, at one point in time, and we claim nothing about engines we did not measure.

None of that makes our reading of the framework correct. It is why this is published as a record rather than asserted as a conclusion. Anyone who can show it is wrong is welcome to, and everything referenced is already public and verifiable without our participation.

Exhibit A · published 2026-08-16 · K.L. Phillips · Jet Fyul Dynamics LLC · nothing here is new evidence; every measurement referenced is already on the public receipts index · the framework discussed throughout is Measuring Visibility in the AI Era, IAB / Project Eidos, August 2026 — linked, not reproduced; read it yourself rather than taking our account of it

2026-08-15

A requirement with no threshold

We have written to the IAB. The letter is in draft and the section above says so; it is not sent and its text is not published here yet.

The position it carries is not new and is not conditional on the letter. The IAB’s framework requires that same-query variability be documented and exposed, and it sets no numerical threshold. A requirement with no threshold can be satisfied by any vendor while telling the buyer nothing at all.

Every disclosure that framework requires is authored after the result is known. That is the structural problem, and no committee number fixes it. The counterweight is commitment before evidence, in a form the reader can verify without trusting the party who made it.

We hold ourselves to it in public: the findings carry a null we published on our own remedy, our own joint-last finish in our own cohort, and the repetition-adequacy test we wrote that invalidated our own headline table. The method was sealed before the evidence and the receipts verify offline.

2026-08-15

An uncounted capture and a capture that never happened were the same output

A capture harness can write its evidence to disk and then die before it writes the record that says the evidence exists. When that happens, nothing downstream can tell the difference between a run that produced nothing and a run whose record never landed. Both read as absence.

We found that state in our own instrument, and the honest description of it is that the defect was ours and it was live. The repair is an enumeration written before the evidence becomes durable, so that a crash leaves an entry marked incomplete rather than no entry at all, and a scorer that halts on a disagreement instead of skipping past it.

The rule underneath is the one this whole company runs on: two distinct states must never produce one identical output. An instrument that cannot tell absence from loss is not measuring; it is guessing politely.

2026-08-14

We published the cohort whole

Thirteen vendors, one sealed question set, four passes, one engine. The set was committed cryptographically and externally timestamped before the first answer was collected, and the rank order was published entire — including where we landed in it, which was joint last, with no asterisk.

Then we applied our own repetition-adequacy standard to the table and found that no adjacent rank in it was distinguishable. We published that invalidation next to the table rather than behind it.

Every reading described here is bounded to one engine: Perplexity. It is not a statement about AI systems generally, and we will not let it be read as one.

It starts with a reading.

Twenty-five minutes. No deck, no pitch. We walk the specific decision your audit would inform, the parameters we would lock into your pre-registration before anything is measured, and whether the test is met. If it is not met, we say so on the call.

Book the 25-minute scoping call →

Or check the work first, without speaking to anyone — the sealed question set and method · the receipts · the findings, including our own. Every reading on this site is bounded to one engine: Perplexity.