Methodology and crawl policy

The Meta Tags Health Check inspects individual pages, but its main value comes from comparing those pages as a system: coverage, consistency, duplication, template behavior, canonical relationships, indexing instructions, and social-sharing presentation. It reports observable technical conditions — it does not, and cannot, promise search ranking or click-through-rate improvements.

Severity levels

Critical
Broad indexability or canonical failure.
High
A confirmed issue affecting discovery, canonicalization, or many pages.
Medium
A material inconsistency or missing metadata.
Low
An optimization or consistency opportunity.
Information
A useful observation without an implied defect.

Health dimensions

The report prioritizes findings over a single score, but a transparent health summary helps orientation. Confirmed contradictions always outweigh length heuristics, and broad template failures always outweigh isolated low-severity issues.

Indexability and directives
weight 30
Canonical integrity
weight 25
Identity coverage and validity
weight 20
Site-wide uniqueness and consistency
weight 15
Social-sharing metadata
weight 10

What gets checked

Document identity
Title, meta description, html lang, charset, author, and relevant article/author metadata.
Indexing and canonicalization
Canonical tags, robots meta (generic and bot-specific), X-Robots-Tag headers, sitemap inclusion, redirects, and hreflang.
Social sharing
Open Graph and Twitter/X card fields, with fallback behavior — a missing Twitter tag is not a failure when Open Graph provides an equivalent value.
Technical and presentation metadata
Viewport, favicon, theme color, manifest, and alternate feeds — weighted lower than indexing and identity problems.
Cross-tag relationships
Contradictions between tags on the same page, explained as one finding rather than several disconnected warnings.
Site-wide patterns
Coverage, exact and near-duplicate values, inferred templates, canonical topology, robots/sitemap consistency, social identity, and outliers.

Representative sampling

When a site's inventory exceeds the page limit, a stable sample is selected: the homepage, shallow top-level pages, a balanced selection from each path group, the oldest and newest entries by last-modified date, and a deterministic selection of the remainder. Results from a partial sample are described as "observed across the scanned pages," not as facts about every page on the domain.

Limits

  • Maximum 100 pages analyzed per report.
  • Maximum 5 sitemap documents fetched, including sitemap indexes.
  • Maximum 10,000 sitemap entries parsed before sampling.
  • Maximum 3 concurrent requests per target hostname.
  • Canonical, hreflang, and social-image targets are validated with a bounded request, not a full crawl.
  • One recent scan per normalized input reused from cache for 12 hours.

Privacy and crawl etiquette

Reports are private and non-indexable by default; publishing a report is an explicit, reversible choice. Running a check makes bounded requests to the submitted website from an identifiable scanner that respects robots.txt, rejects private/internal network addresses, and never forwards cookies, authorization headers, or visitor-controlled headers. Raw page HTML is never stored — only normalized tag values and bounded evidence.

Back to the Meta Tags Health Check

The Ace
Michal's assistant eye