Theme Customization

Customize your layout.
Layout Mode
Menu Mode
Menu Theme
Log in to Market Brew to compare using a URL already in your Ranking Sensor.

Drill into alignment

Start with a target query, then see whether the page's chunks land near the answer cluster or drift away from it.

drifted section
Free Market Brew SEO tool

How to use the BM25 Content Analyzer

Measure literal search-word relevance in visible content, with diminishing returns and page-length normalization.

Quick guide

  1. Use BM25 as a long-tail wording check: it finds the meaningful search words in visible content, including transparent word variations but not inferred synonyms.
  2. Review yellow matches for useful coverage. Rare terms matter more, repetition has diminishing returns, and page length affects the calculation.
  3. Compare match locations rather than copying another page's frequency; navigation and boilerplate are weaker evidence than one clear answer.
  4. Edit one section, click Measure This Change, and keep the revision only when it names the subject naturally without unnecessary repetition.

Use BM25 to examine literal body-language evidence

BM25 Content Analyzer measures how the actual words in a search phrase occur within visible page content. It is a lexical method: it recognizes normalized word forms but does not infer that a synonym, paraphrase, or concept has the same meaning. Enter the target and phrase, then use the yellow highlights to locate the exact occurrences included in the calculation. Unhighlighted related language may still help readers, but it is outside this factor’s literal evidence.

This narrow scope makes BM25 useful for checking whether a long page ever states the subject directly. A document can be semantically sophisticated yet avoid the vocabulary a searcher expects. It can also repeat the vocabulary many times without providing a good answer. The analyzer reveals the literal pattern so an editor can judge whether missing wording reflects a real communication gap or simply an acceptable choice of language.

Understand rarity, saturation, and length

BM25 gives more weight to informative query words that are less common across the modeled collection. Within one page, each additional occurrence contributes less than the previous occurrence. This saturation curve means the first useful mention can matter while the twentieth repeated mention adds very little. The calculation also normalizes for document length so a large page does not win solely because it has more opportunities to contain a word.

The displayed percentage is a diagnostic presentation of that formula, not a recommended keyword density. Inspect the underlying matched words and page context. If the score is low because one essential term never appears, a concise definition may help. If the page already uses all query words in useful sections, adding more copies is unlikely to be valuable. Length normalization also means deleting unrelated boilerplate can clarify the document without inserting any new term.

Know why BM25 is not TF-IDF or semantics

TF-IDF also combines term frequency with rarity, but it does not use BM25’s specific saturation curve and configurable document-length adjustment. Semantic similarity uses embeddings to compare meaning, allowing a paraphrase to align even when no search word appears. BM25 cannot make that inference. Its strength is transparent literal matching: every contribution begins with a query word that can be shown directly in the content preview.

Choose this tool when you want to audit explicit terminology in the body. Choose a semantic visualizer when the concern is topical meaning or conceptual alignment. A page may perform differently in the two analyses without either result being wrong. For example, an article can explain “large language models” thoroughly and receive strong semantic alignment for “LLM,” while BM25 reports that the abbreviation itself is absent.

Compare useful occurrences, not raw repetition

In comparison mode, inspect where each page earns its highlighted matches. One URL may use a query word in navigation, repeated cards, or footer modules, while the other uses it once in a decisive explanation. The bars summarize the formula but cannot replace editorial review of those locations. Look for missing answers, examples, qualifications, or labels that would naturally require the search vocabulary.

Do not copy the outperformer’s frequency. Its length, template, and collection statistics may differ, and its ranking can depend on signals outside BM25. Use the showdown to identify whether your page avoids a necessary term or carries too much diluting text. Then write the smallest original improvement that makes the answer more explicit. Preserve synonyms and natural variation that help readers even when BM25 does not reward them.

Measure a body edit and inspect the highlights

The editable preview for this analyzer focuses on visible content because metadata and path are not part of the BM25 body factor. Change one passage, press Measure This Change, and check which highlights appeared or disappeared. Also compare the edited percentage with the stored score. A useful added section may create a modest increase; repeated copies should show diminishing benefit rather than a linear jump.

If removing a large unrelated block improves or barely changes the result, that can confirm that the deleted material contributed little literal relevance. If a synonym-only edit leaves BM25 unchanged, the result is expected. Review the final copy for usefulness, factual support, and readability before publishing. BM25 is best used as a vocabulary diagnostic inside an editorial process, never as a recipe for keyword stuffing.

Investigate the document behind the BM25 number

Visible content can include more than the main article. Navigation, accordions, repeated cards, legal notices, forms, and embedded interfaces may all appear in the stored field. Scroll through the complete preview and locate every highlight. If most matches come from site chrome, the body may not state the subject as clearly as the percentage suggests. If a large unhighlighted section discusses a different topic, it may increase document length without contributing query evidence. These observations are often more actionable than the headline score.

Collection statistics matter too. The rarity value of a word depends on the pages against which the model was calibrated. A specialized term can carry more information than a common word, so two occurrences need not contribute equally. Avoid comparing raw BM25 behavior across unrelated Ranking Sensors as though they share one universal scale. Use the analyzer within the current modeled context, compare pages evaluated together, and focus on visible explanations that remain understandable regardless of statistical rarity.

For recurring audits, record which query words are absent, where the strongest occurrence appears, and whether boilerplate dominates the matched evidence. Pair that record with semantic and conversion review. A literal term may be necessary for recognition, while a paraphrase may communicate the idea better elsewhere. Keeping both forms can satisfy readers without excessive repetition. The appropriate endpoint is a body that names its subject plainly, develops it with original substance, and avoids unrelated bulk—not a prescribed occurrence count or density percentage.