OA9 OppAlerts from Ben Wills
How AI Search Works: From Prompt to Response A guide for SEOs, AI search marketers, and marketing teams
Work with me
Work in progress Work in progress. Released about a week early, on purpose.

ChatGPT disclosed things in this session that I did not expect it to disclose, and there is no guarantee it stays available. I would rather people had time to use it than had a tidier version of it later.

So: every word of every ChatGPT response here is verbatim, and that part is checked automatically on every build. What is not finished is the presentation. The color coding on the code blocks is incomplete and some of it is imprecise, and there are notes to myself still sitting in the page.

Parts 1 and 2 are close to empty and are getting a lot of detail over the next week or so, along with a cleanup pass on everything else. Worth checking back.

Read the announcement →

AI Search Optimization

How AI Search Works, From Prompt to Response

A guide for SEOs, AI search marketers, and marketing teams

Placeholder

This section is not written yet. The research behind it is finished and the numbers below are verified, but the writing and the exhibits are still to come.

This part covers what happens when there is no search at all. The question goes to the model, the model answers from what it already knows, and nothing is fetched.

The test

One prompt asks for hotel recommendations in Los Angeles. It mentions, in passing, what car the traveler is flying in to pick up. The car has nothing to do with hotels.

Six wordings for the car, from "a used Honda Civic" through "a brand new Ferrari", plus one-word versions so that some pairs of prompts differ by a single word. Six models, three runs each, with temperature and reasoning level swept where the model supports them.

What it shows

Full write-up to come.

Look up any permutation yourself

Every answer below came from the same prompt, with only the car changing. Nothing here was fetched: web search was off for every one of these calls, so this is the model answering from what it already knows.

Pick a model and a dial setting. The three columns are the three cars, each one pooled across both of its wordings, and the tables show everything that exact slice recommended: how often each site made the top 10, and how often it was the #1 pick. Nothing is truncated. Every domain the slice ever returned is listed, so the whole tail moves as you flip one dial at a time.

Car against car, by website

Each car pooled across both of its wordings ("a Civic" with "a used Honda Civic", and so on), so the comparison is car against car with the phrasing averaged out. If the car itself did not matter, these three columns would hold the same list. In the grid there is one row per domain, in the order the columns introduce them: the first column's top 10, then whatever the second column adds, then whatever the third adds. Each cell is that domain's rank for that column, and a blank cell means it did not make that top 10 at all. Teal is a top-3 finish, blue is 4th to 7th, purple is 8th to 10th, and the shading runs dark at rank 1 to light at rank 10 straight through, so only the hue changes at a band edge. Read a row across for one brand's fate, or read the block shape: the solid corner that hollows out is the segment the car deletes.

The same thing again, by hotel name

We asked for websites, so that is what the models gave us, and they spelled the same hotel three different ways: a plain address like westin.com, a page inside the parent company like marriott.com/en-us/brands/westin, or a subdomain like westin.marriott.com. The table above treats the last two as Marriott, which quietly merges brands that compete with each other. Here is the same data with every address turned into the hotel it actually points at. The dropdowns above drive these tables too.

"In top-10" is the share of the selection's answers that contain the entry anywhere; "#1" is the share where it was the first recommendation. Rows are ordered by how often the entry appears, position-weighted on ties. Domains are rolled up to the registrable domain, and n under each car is the number of usable answers in the current selection. Each table scrolls; the full list continues past the first ten rows.