Methodology and crawl policy
The Meta Tags Health Check inspects individual pages, but its main value comes from comparing those pages as a system: coverage, consistency, duplication, template behavior, canonical relationships, indexing instructions, and social-sharing presentation. It reports observable technical conditions — it does not, and cannot, promise search ranking or click-through-rate improvements.
Severity levels
- Critical
- Broad indexability or canonical failure.
- High
- A confirmed issue affecting discovery, canonicalization, or many pages.
- Medium
- A material inconsistency or missing metadata.
- Low
- An optimization or consistency opportunity.
- Information
- A useful observation without an implied defect.
Health dimensions
The report prioritizes findings over a single score, but a transparent health summary helps orientation. Confirmed contradictions always outweigh length heuristics, and broad template failures always outweigh isolated low-severity issues.
- Indexability and directives
- weight 30
- Canonical integrity
- weight 25
- Identity coverage and validity
- weight 20
- Site-wide uniqueness and consistency
- weight 15
- Social-sharing metadata
- weight 10
What gets checked
- Document identity
- Title, meta description, html lang, charset, author, and relevant article/author metadata.
- Indexing and canonicalization
- Canonical tags, robots meta (generic and bot-specific), X-Robots-Tag headers, sitemap inclusion, redirects, and hreflang.
- Social sharing
- Open Graph and Twitter/X card fields, with fallback behavior — a missing Twitter tag is not a failure when Open Graph provides an equivalent value.
- Technical and presentation metadata
- Viewport, favicon, theme color, manifest, and alternate feeds — weighted lower than indexing and identity problems.
- Cross-tag relationships
- Contradictions between tags on the same page, explained as one finding rather than several disconnected warnings.
- Site-wide patterns
- Coverage, exact and near-duplicate values, inferred templates, canonical topology, robots/sitemap consistency, social identity, and outliers.
Representative sampling
When a site's inventory exceeds the page limit, a stable sample is selected: the homepage, shallow top-level pages, a balanced selection from each path group, the oldest and newest entries by last-modified date, and a deterministic selection of the remainder. Results from a partial sample are described as "observed across the scanned pages," not as facts about every page on the domain.
Limits
- Maximum 100 pages analyzed per report.
- Maximum 5 sitemap documents fetched, including sitemap indexes.
- Maximum 10,000 sitemap entries parsed before sampling.
- Maximum 3 concurrent requests per target hostname.
- Canonical, hreflang, and social-image targets are validated with a bounded request, not a full crawl.
- One recent scan per normalized input reused from cache for 12 hours.
Privacy and crawl etiquette
Reports are private and non-indexable by default; publishing a report is an explicit, reversible choice. Running a check makes bounded requests to the submitted website from an identifiable scanner that respects robots.txt, rejects private/internal network addresses, and never forwards cookies, authorization headers, or visitor-controlled headers. Raw page HTML is never stored — only normalized tag values and bounded evidence.
