Part III · Chapter 17 of 42

Evidence, Expertise, and Original Research

Becoming a source worth citing

Citation selection favors sources that add verifiable, first-hand information, and that makes original data the content investment with the biggest payoff in this discipline. Everything else on your site competes with every competent competitor; a number only you can publish competes with nobody. This chapter covers why engines prefer evidence mechanically, the trust signals machines can parse, and how a small team produces research worth citing.

Why engines prefer evidence-dense sources, mechanically

An AI answer with citations has to justify its claims from the passages it retrieved. A passage stating "42% of the surveyed teams switched tools within a year, n=310, surveyed in March" gives the model an assertion it can make and attribute. A passage stating "many teams switch tools" gives it nothing to credit. That asymmetry is measurable: the GEO paper, the first systematic benchmark of content changes against generative engines, reports visibility lifts of up to 40% from content optimizations, with adding statistics, quotations, and cited sources among the tactics tested.GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024). Effectiveness varied by domain, so read the 40% as a ceiling observed in a benchmark, and expect drift as engines evolve.

The search side of the pipeline points the same way. Google's people-first content standard, E-E-A-T (experience, expertise, authoritativeness, trust), asks whether content demonstrates first-hand experience and whether a visitor can tell who wrote it and why they should be trusted.Google's guidance on helpful, reliable, people-first content, the document that defines E-E-A-T. Since AI search products retrieve through search indexes, the content that ranking systems trust is the content generative answers get built from.

The credibility stack: author, evidence, method, date

Four signals, all parseable by machines and checkable by people. Author: a named person with a findable identity and stated credentials, on the page, linked to a real author profile. Evidence: data and primary references rather than recycled secondary claims. Method: how the numbers were produced, on a methodology page you keep permanent. Date: when the claim was true, stated plainly, maintained honestly. Whether every engine weighs each signal today is not established, and I would not claim it is; the assessment is that these four are cheap, they compound, and every one of them also makes the content better for the human deciding whether to believe you.

Where credited citations actually point

OppAlerts' citation analysis across 100 industries looked at which domains AI answers actually credit. Two patterns matter for this chapter: credited citations skew heavily toward first-party sites, brands cited as the source about themselves, at roughly [number: verify] of named citations, and a short list of gatekeeper publishers recurs as credited sources across dozens of unrelated industries.The citation-mix and gatekeeper analyses in the AI Search Visibility research; method and tables in The Research Behind This Guide.

That split is why original research pays twice. Your own stat pages win live retrieval, where first-party sources get credited directly. And when gatekeeper publishers pick up your numbers, their coverage feeds the training data that builds model memory, the slow clock from Model Memory and Live Retrieval. One data asset, both mechanisms. Digital PR, News, and the Gatekeeper Publishers covers identifying your category's actual gatekeepers empirically.

Original research on a small team's budget

Scope it around data you already hold or can cheaply create: your own operational logs, aggregated and anonymized customer data, a survey of a few hundred practitioners, or a benchmark you can run yourself against products in your category. The bar is verifiable and first-hand, and a modest n honestly reported beats a grand claim nobody can check. Pick the one question in your category that gets argued about without data; that argument is your distribution plan already written.

Produce it with the credibility stack built in: named analyst, stated n, collection dates, method page. Publish the findings as extractable passages, one claim per sentence, numbers in text, following Writing for Retrieval and Citation. Then pitch the single most surprising number, not the report, to the publishers that gatekeep your category. Publishers cite numbers; nobody cites a PDF announcement.

Statistics pages: the assets that accumulate citations

A statistics page collects the key numbers on one topic, each stated in a self-contained, quotable sentence with its source and date, updated on a stated schedule. These pages accumulate citations for a simple reason: writers and machines both cite the page that made the number easy to lift. Build one per core topic, seed it with your original numbers alongside honestly attributed third-party ones, and keep the URL stable so its citations compound. This guide's own research part follows the same pattern, and it exists for the same reason: a brand that publishes the data stops being a subject other sources describe and starts being the source they cite.