Chapter Glossary
Canonical Tag An HTML element that signals to search engines which version of a duplicated or near-duplicate URL should be treated as the authoritative one
Chapter Glossary
- Canonical Tag
- An HTML element that signals to search engines which version of a duplicated or near-duplicate URL should be treated as the authoritative one for indexation and ranking purposes.
- Canonicalization
- The process by which search engines or site owners resolve duplicate content by designating one URL as the preferred version and consolidating signals toward it.
- Query Parameter
- A variable appended to a URL (e.g. ?sort=price) that modifies the content or presentation of a page without creating genuinely distinct content, often generating large numbers of near-duplicate URLs.
- Faceted Navigation
- A filtering system that allows users to narrow results by multiple attributes simultaneously, typically generating a combinatorial explosion of parameter-modified URLs.
- Pagination
- The practice of splitting a long list or article across multiple sequentially numbered URLs, each representing a portion of the whole.
- Lazy Loading
- A performance technique that defers the loading of offscreen images or content until a user scrolls to them, which can prevent crawlers from seeing content that is never triggered during a headless fetch.
- Cache Header
- An HTTP directive that instructs browsers and intermediary servers how long to store a cached version of a page before requesting a fresh copy, affecting what version a crawler may receive.
- Crawl Pipeline
- The sequential stages a search engine uses to process a URL: discovery, scheduling, fetching, rendering, and the separate decision of whether to index the rendered output.
- 301 Redirect
- A permanent redirect that signals to search engines the original URL has moved and that ranking signals should be consolidated at the destination.
- 302 Redirect
- A temporary redirect that signals the move is not permanent, instructing search engines to retain the original URL in their index rather than transferring signals to the destination.
- 307 Redirect
- A temporary redirect that preserves the HTTP method of the original request, functionally similar to a 302 but with stricter method-preservation behavior.
- Meta Refresh
- An HTML directive that instructs a browser to redirect to another URL after a specified delay, processed client-side after the original page has already loaded and rendered.
- Redirect Chain
- A sequence of redirects where the destination of one redirect is itself another redirect, compounding latency and signal loss with each additional hop.
- Redirect Loop
- A circular redirect sequence where following the chain eventually returns to the starting URL, making the destination unreachable for both crawlers and users.
- X-Robots-Tag
- An HTTP response header that applies noindex, nofollow, or other crawl directives at the server level, usable on any file type including PDFs and images that cannot contain HTML meta tags.
- JavaScript Redirect
- A redirect executed via client-side JavaScript after a page has loaded, meaning a crawler may index the original page without ever following the redirect to its intended destination.
- Site Architecture
- The hierarchical structure of a website's URLs and internal links, which determines how authority flows between pages and how many clicks separate any given page from the homepage.
- Orphaned Page
- A page with no internal links pointing to it, making it undiscoverable by any crawler that finds pages by following links rather than through a sitemap.
- Crawl Budget
- The finite amount of crawl activity a search engine allocates to a given site within a period, meaning that pages crawled for low-value URLs reduce the attention available for high-value ones.
- Log File Analysis
- The examination of server access logs to understand actual crawler behavior (what URLs were fetched, how often, and with what response codes) as distinct from tool-based estimates.
- Bingbot
- Microsoft Bing's web crawler, which operates on different scheduling, prioritisation, and indexation logic than Googlebot, and responds differently to certain technical signals.
- Privacy-First Search Engine
- A search engine that does not track individual user behavior across sessions, which removes personalisation signals from ranking and changes the relative weight of other ranking factors.
- Rendering
- The stage of the crawl pipeline where a search engine executes JavaScript and builds the full DOM of a page, which is a separate step from fetching the raw HTML and may happen with significant delay.
- CDN (Content Delivery Network)
- A distributed network of servers that delivers cached copies of a site's content from locations geographically closer to the user, reducing network latency as one component of overall page speed.