OA9 OppAlerts from Ben Wills
How AI Search Works: From Prompt to Response A guide for SEOs, AI search marketers, and marketing teams
Work with me
Work in progress Work in progress. Released about a week early, on purpose.

ChatGPT disclosed things in this session that I did not expect it to disclose, and there is no guarantee it stays available. I would rather people had time to use it than had a tidier version of it later.

So: every word of every ChatGPT response here is verbatim, and that part is checked automatically on every build. What is not finished is the presentation. The color coding on the code blocks is incomplete and some of it is imprecise, and there are notes to myself still sitting in the page.

Parts 1 and 2 are close to empty and are getting a lot of detail over the next week or so, along with a cleanup pass on everything else. Worth checking back.

Read the announcement →

Visibility measurement

How to Write Prompts Worth Tracking

What a prompt has to contain, how often to run it, and how to report AI search visibility honestly.

Draft

First pass. The structure is settled, the wording is not.

Parts 1 and 2 showed that one irrelevant word rewrites the answer, and that settings the person cannot see rewrite it again. That has a direct consequence for anyone trying to measure their visibility.

A short prompt is not a measurement. "Best CRM software" run once, tracked daily, produces a number that moves for reasons that have nothing to do with your marketing.

Put the person in the prompt

The hotel test works because the car acts as a persona. It is the smallest amount of information a model can use to decide what kind of buyer is asking, and it decides. If you leave the buyer out, the model invents one, and you have no idea which one it picked.

So state it. Every prompt you track should carry:

Part 5 is the reference for this. Each walkthrough opens by turning a sentence into a specification with hard constraints and preferences separated. Write your prompts so that step is already done.

Then keep them identical

Once a prompt is written, freeze it. Every word is a variable. Changing "cheap" to "affordable" between runs makes the two runs incomparable, and you will read the difference as movement in the market.

Run it many times

Two runs of an identical prompt, with every setting pinned, keep the same top result about 62% of the time. One run tells you almost nothing. A single answer is one sample from a distribution being reported as a fact.

Run each prompt enough times to see the spread, and report the spread.

Run it across the settings you cannot control

Reasoning level flips entire answers, and most people never touch it. Neither do the customers you are trying to reach, which is the point. You do not know which setting they are on, so cover the range and treat the variation as part of the result.

The same goes for search. Whether the model searches at all varies from 10% to 100% by model, so run each prompt with search on and off and expect two different pictures.

Slow down

Daily rank tracking makes sense when the underlying thing changes daily and the measurement is stable. Here neither is true. Monthly is enough. Run a large batch, look at where you land across all the variations, and let that set the direction for the next month.

Report it qualitatively

Across 1,363 answers to one question, no domain appeared in 90% of them. One cleared 75%. Five cleared 50%. Every one of the top ten domains occupied nearly the full range of positions from first to tenth.

A single rank number cannot describe that honestly. What can:

That last one is where Part 4 becomes practical. If you know what the machine opens and reads before it answers, you know what to go and build.