Jeremy Rivera had me on the Unscripted SEO Podcast yesterday. We spent most of the hour on one experiment and what it means for anyone trying to measure AI search.
The experiment was 1,374 runs of one prompt. I asked for a hotel in Los Angeles and said I was flying in to pick up a car. Then I changed the car. A Honda Civic got one set of hotels. A Ferrari got more expensive ones. I ran it across three ChatGPT models at different reasoning levels, and the differences got larger with more reasoning.
That's the part marketers need to sit with. Your customer gives the LLM details you never see, and those details change which businesses show up. If you only test the short, generic prompt, you miss it.
A few other things we got into:
- Why I test LLMs directly. You write the prompt, you get the result, and you learn something about the model. No waiting on an index to update.
- Correlation, not causality. Backlinks, rankings and LLM mentions move together overall, but the relationships change by industry and by persona. I use them to pick what to test next.
- Reasoning versus web retrieval. The answers change a lot once search is turned on, and an API result is a different measurement than a logged-out ChatGPT result.
- Search still matters. The more these systems retrieve, the more useful content, relevant publishers and backlinks stay part of the job.
I wrote up a full recap with the video, the audio, and the key quotes:
One Word Changed Every Answer: Ben Wills on Testing LLMs Directly
Written by Ben Wills, 26+ years across marketing and engineering.
Let’s discuss how we can work together.
An industry report, a custom dataset, a partnership; a couple of sentences is plenty. Every message comes straight to me, and I read all of them.
If it’s a fit, you’ll hear back quickly with next steps or a time to talk.