Live The AI Search Visibility & Ranking Factors reports are now live. Click here to view them →
AI Search Visibility

You Will Never, Ever, Ever Have Consistent AI Ranking Reports

You will never, ever, ever have consistent AI ranking reports. If you don’t thoroughly understand why, you don’t understand AI search. Imagine this prompt:

“I am flying into Los Angeles for a night to pick up my new car. It is a {{VEHICLE}}. Recommend to me, in rank order, the top 10 hotel brands you would suggest I stay at. I’m happy to stay anywhere in the greater LA area. Give me the websites only, no extra text.”

You’re just looking for a hotel. Do you expect to get different recommendations based on your vehicle?

If you don’t _confidently_ understand why buying a different vehicle will change the hotel recommendations you get, you do not understand AI search and LLMs.

Here are the next questions: What factors change the results you get from that prompt? And how many different permutations are there?

Most people don’t understand _deeply_ how the following things will affect your results:

  • The Model You Use: ChatGPT uses the chat-latest model for free accounts, but variations of that and gpt-5.3, 5.5, 5.6, and o3 with different reasoning levels in their chat interface….and that can change every couple of weeks.
  • The Reasoning Level: Higher levels of reasoning will introduce more variability in the results. For free accounts, reasoning is usually restricted to the lowest level(s).
  • Fanout Query Use: Higher reasoning is more likely to introduce more results from around the internet. This introduces exponentially more variance.
  • Temperature: This is something older model APIs allow you to configure, and influences how “plain vs creative” the result is. Algorithmically, it’s a randomization factor. There are 1,073,741,825 possible variations of this number. Internally, I _believe_ this is still used, except the provider decides the best temperature to use based on the prompt. When ChatGPT asks you to compare two results, assume you’re helping improve temperature and reasoning levels used (and more).

But this is just the start.

Do you understand that your IP address will almost certainly affect your results? It’s easy to tie IP addresses to specific locations, then turn that into a demographic profile that gets inserted into your prompt behind the scenes.

Do you understand how chat history/memory works? It’s short snippets that get sent with your prompt behind the scenes. If I’ve been talking about moving to Beverly Hills for 3 months, will I get different results than if I’m moving to Winchestertonfieldville Iowa?

This is why a _realistic_ representation of “rank tracking” requires using the APIs (eliminate chat history and IP skewing results, for a purer LLM representation), and measuring across different reasoning levels, fanout query abilities, etc.

I’ll show you the hard, concrete data tomorrow, from more than 1,000 prompts.

Ultimately, it’s our responsibility to help our clients understand this so that expectations can be properly held.

But it’s our responsibility to understand it, first.

Tomorrow, I will walk you through new research showing exactly why.

Written by Ben Wills, 26+ years across marketing and engineering.

Let's discuss how we can work together.

An industry report, a custom dataset, a partnership; a couple of sentences is plenty. Every message comes straight to me, and I read all of them.

If it's a fit, you'll hear back quickly with next steps or a time to talk.

Website Contact Form