Large-scale correlation research built from over 403,000 LLM prompts and 150,000 search engine queries, across 100 industries and 1,100 buyer personas. Every result is cross-referenced against nearly 15 billion web pages, the entire history of Reddit, all of Wikipedia, and billions of links.
Just the signals that move the needle in your market.
The top domains for the Airlines industry across four channels. Switch the model in the LLM and Fanout columns to see how each AI system's recommendations differ.
Compare how two AI models rank the same domains for the Cruises industry. Pick a model on each side; the rank slope shows where they agree and where they split.
Domains Google's AI Overview cites inline in the visible answer, next to domains it references as supplemental research, for Hotel Resorts. A wide gap flags a source it leans on but rarely shows, or the reverse.
Google vs. Bing, straight rank-weighted score, no per-engine weighting, for Consumer Banking.
Keywords LLMs associate with Fitness Clubs, ranked by a weighted score. Toggle models on or off and watch the ranking recalculate.
| Signal | Group | ρ (Spearman) | R² | n | Tier |
|---|---|---|---|---|---|
| SE Outbound Links | Search | +0.331 | 11% | 40,733 | Dominant |
| Homepage Keywords | Content | +0.204 | 4.2% | 41,494 | Strong |
| Domain PageRank | Backlinks | +0.192 | 3.7% | 97,628 | Confirmed |
| PageRank History | Backlinks | +0.183 | 3.3% | 98,546 | Confirmed |
| Search Engine Appearances | Search | +0.165 | 2.7% | 9,323 | Confirmed |
| Common Crawl | Web | +0.165 | 2.7% | 62,178 | Confirmed |
| Host Harmonic Centrality | Backlinks | +0.164 | 2.7% | 97,280 | Confirmed |
| Domain Backlinks | Backlinks | +0.160 | 2.6% | 74,885 | Confirmed |
| Host PageRank | Backlinks | +0.156 | 2.4% | 97,280 | Confirmed |
| Harmonic Centrality History | Backlinks | +0.153 | 2.4% | 98,546 | Confirmed |
| Domain Harmonic Centrality | Backlinks | +0.151 | 2.3% | 97,628 | Confirmed |
| Wikidata | Reference | +0.151 | 2.3% | 12,300 | Confirmed |
| Host Backlinks | Backlinks | +0.150 | 2.2% | 65,691 | Confirmed |
| Best Search Engine Rank | Search | +0.148 | 2.2% | 9,323 | Confirmed |
| Reddit Comments | Social | +0.148 | 2.2% | 48,298 | Confirmed |
| Reddit Posts | Social | +0.128 | 1.6% | 41,951 | Confirmed |
| Avg Search Engine Rank | Search | +0.077 | 0.6% | 9,323 | Emerging |
| Wikipedia Citations | Reference | +0.055 | 0.3% | 24,499 | Emerging |
1,100 buyer personas across 100 industries, run against 13 current models with over 403,000 prompts. Top recommended domains, search phrases, and on-page phrases captured per persona.
Seven web crawls analyzed, December 2025 through June 2026: nearly 15 billion pages scanned for general web presence and crawlability of every recommended domain.
The entire history of Reddit, 2005 through mid-2026: 3.7 billion submissions and 26.6 billion comments scanned for domain mentions.
More than 150,000 organic searches across Google and Bing, up to 100 results captured per query. Appearances and rank position recorded for every domain.
Wikidata entity associations plus English Wikipedia citations and outbound links, cross-referenced against 25 million+ articles and 120 million+ entities.
Every Common Crawl web graph release since 2017, 282 billion+ links analyzed. PageRank and Harmonic Centrality computed for every recommended domain.
Raw HTML downloaded for 3.67 million pages that appeared in the collected search results, and outbound links to every tracked domain extracted.
Homepage HTML downloaded for 176,000+ tracked domains, scored against industry and persona phrases weighted by placement as a content relevance signal.
SEO, analytics, and competitive intelligence teams who need someone fluent in both marketing and engineering, able to take a data project from concept to deliverable without hand-holding.
Your client needs data and research you can't build in-house. I build it, you deliver it: clean handoff, no drama, and you set your own margin.
On the marketing side, I've directed teams of 70–80 people responsible for over 1,400 SEO client accounts, led international SEO campaigns across 30–40 countries, and served as the weekly point of contact for Fortune 500 accounts.
On the engineering side, I've built large-scale web scraping and indexing systems processing billions of records, written a marketing SaaS platform from scratch in pure C, and compiled the complete Common Crawl web graph history into queryable SQLite databases.
We talk through what you're trying to accomplish, not a feature list. I define exactly what you'll receive.
No hourly billing surprises, no scope creep. You know what you're getting and what it costs before anything starts.
Delivered in your preferred format (CSV, JSON, Excel, SQLite): clean, documented, ready to plug into your workflows.
Once you see the data, you'll want it differently. That's expected; iteration is priced in from the start.
Parse and analyze pages across Common Crawl releases at scale: structured data from HTML, URLs, tags, and page elements across billions of records.
Hundreds of thousands of search queries, every ranked URL analyzed. Contact info, partnership opportunities, and competitive gaps, extracted and structured.
Track competitors across web properties, search results, and market signals. Build a database your team can search and act on immediately.
Monitor search results, news, RSS feeds, and web mentions for brand terms, product names, executives, and competitors, near real-time.
Searchable databases of sponsorship opportunities, link prospects, and outreach targets, nationwide or by specific location.
Phone numbers, emails, social accounts, and named entities: extracted, validated, and delivered in your preferred format.
Reports run on a quarterly subscription, but the underlying data itself refreshes roughly every two weeks, sometimes faster, sometimes a little slower, as new backlink, Common Crawl, Reddit, and other source data becomes available.
Yes, cancel anytime. You'll keep access through the end of the period you've already paid for.
Whatever works for your team: CSV, JSON, Excel, or SQLite, with field-level documentation and schema definitions. Web graph databases are delivered as SQLite.
Almost always fixed price. I scope the project, define clear deliverables, and quote a number before work begins. In-scope revisions are included.
It depends entirely on the project. Some are a few weeks, some are longer; you'll get a realistic timeline during the scoping conversation.
That's expected. Once people see their data, they almost always want adjustments. Iterations are baked into every project from the start.
Let me know what you're working on and we'll set up a time to talk: an industry report, a custom dataset, or both.
Reddit has influenced AI Search Visibility since day 1. ChatGPT now uses Bing for some web searches. Will Reddit influence AI Search as much…
Claude’s web search runs on Brave Search. Anthropic’s subprocessor registry says so, and the first part of this study confirms it independently, in the…