Part II · Chapter 11 of 42

What Actually Correlates with AI Search Visibility

What the data says moves AI answers

This is the evidence chapter: the data behind the book's ordering of effort, so you can weigh the advice instead of taking it on faith. The data is OppAlerts' own study: AI recommendations across 100 industries and roughly 138,000 recommended brand domains, from 13 models answering from memory and 5 answering with live search, correlated against Google rankings, backlinks, Wikipedia, Wikidata, Reddit, news coverage, and video.Method, data sources, and full tables are in The Research Behind This Guide, part of the AI Search Visibility research. All figures in this chapter come from that study unless a different source is named. The findings reorder most instincts.

The leaderboard, and what "independent" means

Measured one at a time against how strongly a pure-memory model recommends a brand, the signals rank: backlink authority (Spearman rho 0.23), Reddit discussion volume (0.23), presence in category news (0.19), presence in Wikipedia (0.16), and a Wikidata knowledge-graph entity (0.14). The correlations are small in absolute terms because recommendation strength has many inputs; what matters is the ordering.

One-at-a-time is also the wrong test, because famous brands have all of these at once: the brand with the strong backlink profile is the same brand Reddit discusses and Wikipedia covers. The fix is partial correlation, which measures each signal's relationship to the AI recommendation while holding the others fixed, so only its independent contribution remains. That test collapses the field. Backlinks keep a real independent relationship (partial correlation 0.077), Reddit keeps a smaller one (0.046), and the encyclopedic signals drop to nothing: Wikidata 0.002, Wikipedia -0.011. About 98% of a Wikidata entity's apparent power, and effectively all of a Wikipedia page's, is shared with the other signals, mostly backlinks.

The concrete version is easier to hold. Brands in Wikipedia carry about 1.26x the AI recommendation strength of brands that are not; compare two brands with the same backlink authority and that lift falls to 1.00 to 1.03. The page was never the lever; the authority that earned the page was. Backlinks also win the industry-by-industry count, 75 of 100 industries to Reddit's 23, and rank as the strongest single signal inside all 13 memory models. Classic SEO visibility adds its own independent signal on top: Google organic visibility predicts the AI recommendation at a partial correlation of about 0.17 among brands measurable in both.

Thresholds versus gradients

Signals move the AI in two different shapes, and the shape decides how much of the work is worth doing. Wikipedia is a threshold. Whether a brand appears in Wikipedia at all predicts pure-LLM recommendation strength at rho 0.126; how many articles cite it predicts at 0.011, roughly eleven times weaker. A model that read Wikipedia in training seems to keep a yes/no "this is a notable brand" bit, set the moment the brand clears the bar, and piling on coverage after that does close to nothing. News works as both: being in the category's news cycle at all is worth about a 1.45x lift, and unlike Wikipedia, additional coverage volume keeps adding signal after the threshold.

So chase presence where the signal is a threshold and volume where it is a gradient, and stop when the shape says stop. The entity work this implies is in Entities: Becoming a Thing the Machine Knows.

What gets read and what gets credited

For the five search-grounded models, the study logged both what each model retrieved while answering and what it named as a reference. The two lists barely resemble each other. Reddit is the most retrieved third-party domain by a wide margin, and for every 100 Reddit retrievals the models credit it about once; Wikipedia runs at 3.5 credits per 100 retrievals. News, review, and analyst sites run the other way, credited at two to four times their retrieval share. The one near-universal third-party gatekeeper is Forbes: credited across 51 of 100 industries, by all five models. And the largest credited bucket is no third party at all: brands' own sites are at least 53.9% of named citations.

Video barely registers on either list: 0.5% of what the grounded models cite, against 9.2% for Reddit and 2.6% for Wikipedia, and most of that sliver is one model. Video presence correlates with AI recommendations anyway, because famous brands dominate the video shelf. That is a mirror of fame, and the mirror-versus-lever distinction runs through the practice chapters, starting with Reviews, Reddit, and Communities.

News predicts memory, not search

The expected result would be that search-grounded models, which can fetch this week's coverage, reward press more than memory models frozen at their cutoff. The data shows the opposite: on the same brands, news presence predicts memory recommendations (rho 0.169) more strongly than search recommendations (0.139), and 12 of the 13 memory models out-predict every search-grounded model on this signal, despite never having read the current news cycle. Press works on the AI as accumulated reputation, absorbed into training, and a fresh hit does little for next week's answers.

Correlation, causation, and how to read the rest of this book

Everything above is correlational. Famous brands have backlinks, Reddit threads, and Wikipedia pages because they are famous, and the models recommend them for the same reason, so no correlation here proves that adding a signal causes a recommendation. Two kinds of outside evidence narrow the gap. Controlled experiments exist on the content side: the GEO benchmark paper measured visibility changes from rewriting page content, adding citations, quotations, and statistics, and reports gains of up to 40% in generative-engine responses.Aggarwal et al., GEO: Generative Engine Optimization, which introduced the GEO-bench benchmark. A later critical survey reviews which of the field's claims have held up and which remain thin; the honest summary is that content-side experiments exist and authority-side experiments largely do not. And the measurement literature backs the sampling discipline this book keeps repeating: answers vary across runs, so visibility is a distribution, measured repeatedly, never a single observation.Don't Measure Once: Measuring Visibility in AI Search, which makes the empirical case for repeated measurement and distribution-level reporting.

So read the practice chapters this way: the correlations rank where the signal is, the mechanism chapters explain why it would be causal, and the honest position on most authority signals is "strongly consistent with, unproven as, a lever". Where a finding is a mirror rather than a lever, the chapter that covers it says so.

The finding-to-chapter map

FindingActs throughChapter
Backlinks: the strongest independent signal, in all 13 modelsBoth memory and retrievalLinks: The Signal That Refuses to Die
Reddit: the real number two, heavily read, rarely creditedBothReviews, Reddit, and Communities
Wikipedia and Wikidata: thresholds, redundant with authorityMemoryEntities: Becoming a Thing the Machine Knows
News: reputation signal, predicts memory over searchMemoryDigital PR, News, and the Gatekeeper Publishers
Own site: the largest credited citation bucketRetrievalWriting for Retrieval and Citation
Content changes move generative visibility (experimental)RetrievalContent Architecture: Building for the Chunk
Video: mirrors fame, moves nothing measuredNeitherMultimodal: Images, Video, and Audio

For any signal someone urges you to invest in, ask the three questions this chapter just asked: does it hold up independently, is it a threshold or a gradient, and is it a lever or a mirror. Most of the folklore in this space fails at least one.