Blog

Indexability: canonical, robots, and HTTPS

Web Search

HTTPS redirects, self-referencing canonicals, and robots rules decide whether search can even consider your pages.

Before rankings and snippets come a simpler question: may this URL be indexed? VectoraPoint’s Indexability dimension looks at HTTPS (and HTTP→HTTPS redirects), meta robots / X-Robots-Tag, self-referencing canonicals, and whether robots.txt is blocking the crawler from the start.

Canonicals that tell the truth

A self-referencing canonical says “this URL is the preferred version of itself.” Cross-host or cross-path canonicals that point somewhere else can quietly drop the page you meant to rank. Staging hosts, trailing-slash mismatches, and parameter variants are classic sources of accidental noindex-by-canonical.

Robots: precise, not panicked

robots.txt is for crawl paths, not a privacy policy. Blocking entire sections you care about ranking for is a self-own. Conversely, leaving admin, cart, and filter-parameter chaos crawlable wastes budget. Pair robots rules with sitemap discovery so important URLs are both allowed and listed.

HTTPS is table stakes

Mixed HTTP/HTTPS sites, missing redirects, or certificate failures show up as indexability failures for a reason: crawlers and users expect a single secure entry point. Fix redirects first; then verify canonical and robots agree with that canonical host.

Use VectoraPoint’s Indexability insight to see which of these signals is costing you points — critical issues are listed before smaller hygiene items.

Audit my site More posts