Chapter Glossary
Crawl budget The finite number of pages a search engine will crawl on a given site within a set period. Search engines allocate crawl attention based on si
Chapter Glossary
- Crawl budget
- The finite number of pages a search engine will crawl on a given site within a set period. Search engines allocate crawl attention based on site size, authority, and server responsiveness, meaning not every page on a large site will be crawled with equal frequency or priority.
- Core Web Vitals
- Three specific metrics Google uses to measure page experience: Largest Contentful Paint (how long the main content takes to load), Interaction to Next Paint (how quickly the page responds to user input), and Cumulative Layout Shift (how much the page layout moves unexpectedly during loading).
- Largest Contentful Paint (LCP)
- The point in a page's loading sequence when the largest visible content element (an image, a heading, or a block of text) has fully rendered. A slow LCP signals to search engines that the page feels slow to users.
- Interaction to Next Paint (INP)
- A Core Web Vitals metric that measures how quickly a page responds to user interactions such as clicks or taps. High INP values indicate a page that feels unresponsive.
- Cumulative Layout Shift (CLS)
- A Core Web Vitals metric measuring visual instability: how much page elements move around after initial rendering. High CLS creates a poor user experience and signals instability to search engines.
- Mobile-first indexing
- Google's practice of using the mobile version of a page as the primary version for crawling, indexing, and ranking. Content that exists only on the desktop version of a page is effectively invisible to Google's evaluation.
- JavaScript rendering
- The process by which a browser executes JavaScript to build the visible page. Search engine crawlers may not execute JavaScript during their first crawl pass, meaning content that only appears after JavaScript runs may not be seen or indexed immediately.
- HTTPS (HyperText Transfer Protocol Secure)
- An encrypted version of the standard web protocol that protects data exchanged between a user's browser and a web server. Search engines treat HTTPS as a baseline trust signal, and browsers actively warn users when a site uses unencrypted HTTP.
- Noindex
- A directive placed in a page's HTML or HTTP headers that instructs search engines not to include that page in their search results. The crawler can still visit and read the page; it simply will not show it in results.
- Disallow
- A directive placed in a site's robots.txt file that instructs crawlers not to visit specified URLs. Unlike noindex, disallow prevents the crawler from accessing the page at all, which means it cannot read any noindex tags present on the page.
- Robots.txt
- A plain-text file at the root of a domain that communicates crawling instructions to search engine bots. It controls access at the crawl stage, before indexing decisions are made.
- HTTP status code
- A three-digit number returned by a server in response to a page request, communicating the outcome of that request. Common codes include 200 (success), 301 (permanent redirect), 404 (not found), 410 (permanently gone), and 503 (temporarily unavailable).
- 301 redirect
- A permanent redirect that tells search engines a URL has moved to a new location and that ranking signals should be transferred to the destination URL.
- 404 status code
- A response indicating that a requested page cannot be found. Search engines interpret a 404 as a page that may return; they typically continue to check the URL periodically rather than immediately removing it from the index.
- 410 status code
- A response indicating that a page has been permanently removed and will not return. Search engines treat a 410 as a stronger signal to remove the URL from the index than a 404.
- Duplicate content
- Substantively identical or very similar content accessible at more than one URL. Search engines must choose one version to rank, which can dilute the ranking potential that would otherwise concentrate on a single canonical URL.
- Canonical tag
- An HTML element that signals to search engines which version of a URL should be treated as the primary, authoritative version. Canonical tags allow duplicate URLs to exist without splitting ranking signals across multiple versions.
- URL parameter
- A variable appended to a URL, often used for tracking, filtering, or session management. Parameters frequently generate duplicate or near-duplicate URLs that search engines must evaluate separately unless canonical tags or other controls are in place.
- Crawlability
- The degree to which a search engine can access, navigate, and read the pages on a site. Technical barriers such as blocked resources, server errors, or misconfigured robots.txt directives reduce crawlability and limit how much of a site enters the index.