Live The AI Search Visibility & Ranking Factors reports are now live. Click here to view them →

Content and Semantics Terms

Understand the terms behind semantic content analysis. Learn how embeddings, taxonomies and similarity relate to draft review, audience relevance and editorial decisions.

When you use semantic analysis in editorial work, you need to understand what the comparison says about your draft. Terms used across this site when discussing meaning-based content comparison.

Vector

An ordered list of numbers used as a position. On its own it means nothing; it is useful because positions can be compared.

Embedding

The vector produced from a piece of text by a model, and the process of producing it.

Text about similar subjects produces nearby positions, and text about unrelated subjects produces distant ones. In this product a piece of writing becomes 768 numbers.

The property that makes this useful is that the comparison is of meaning rather than of wording. Two passages describing the same thing in entirely different vocabulary have similar representations; two passages sharing vocabulary while discussing different subjects do not.

Two embeddings can only be compared when the same model produced both. Vectors from different models are numbers of the same shape describing different spaces, and comparing them returns a value that means nothing.

Similarity

How close two positions are. Displayed as a percentage in the content reports.

Four things similarity does not establish, each of which is regularly assumed:

  • It is not accuracy. A confident, well-written, factually wrong passage can closely match its subject.
  • It is not quality. A short focused piece can score higher for similarity than a thorough one that covers several subjects.
  • It is not a ranking signal. No engine is involved in producing it.
  • It is not keyword coverage. A high score does not mean particular terms are present, and adding terms does not reliably raise it.

Taxonomy

A named set of categories covering one dimension: subjects, audiences, industries, products, or locations.

A category belongs to one taxonomy. A content comparison therefore accepts one selected category per taxonomy. Each result reports that category's position; two selections within the same taxonomy would require separate positions.

Category

One entry in a taxonomy, with its own stored position.

Selecting a category states your intent before writing is measured. The comparison then reports whether the finished text reads as belonging to it, which is a different question from whether the text is good.

Nearest categories

The categories whose stored positions are closest to your text, returned in order.

Focus on the first few results. Later results may be the closest remaining categories without being relevant. Ten results do not necessarily mean ten suitable matches.

An unexpected category match can identify a mismatch with your intention. If your selected category is absent from the first few results, the writing is closer to other categories.

Topic and intent

Topic is what a piece is about. Intent is what the reader wants from it.

They come apart often, and the gap is a common cause of content that measures well and performs badly. An article about a product category, written for a reader who has already chosen and wants to compare two options, is on topic and wrong for the intent.

Comparing a draft against an audience taxonomy tests part of this, by asking who the writing reads as being addressed to.

Content brief

The instructions a writer works from: the audience, the questions to answer, the evidence available, the intended category, and the examples worth using.

A brief that names an intended category makes the later measurement meaningful, because there is a stated intent to compare the result against. Without one, a similarity report can only tell you what the writing became.

Phrase suggestions

Terms associated with a selected category that a draft does not contain.

These are evidence about a gap rather than a list to insert. Some will be irrelevant to your particular piece. Adding terms to move a number is the one use of the report that reliably produces worse writing, and the report cannot tell that it happened.

Input limit

The amount of text a comparison accepts in one request, here 8,192 bytes.

Most Latin characters take one byte; accented characters and other scripts take more, so text in those languages reaches the limit sooner than its character count suggests.

Longer writing is refused rather than trimmed. Scoring the opening of a long article and presenting the result as the whole article would describe a page that starts well and wanders as if it were fine.

Use these definitions when discussing a content report with a writer. They help you explain the finding and keep the revision tied to the brief.

See what vector similarity and keyword matching each measure or how category taxonomies help review a draft.

Let’s discuss how we can work together.

An industry report, a custom dataset, a partnership; a couple of sentences is plenty. Every message comes straight to me, and I read all of them.

If it’s a fit, you’ll hear back quickly with next steps or a time to talk.