Chapter Glossary
Site Structure The organisational framework that connects all pages on a website, including how they are linked to one another, how deep they sit within th
Chapter Glossary
- Site Structure
- The organisational framework that connects all pages on a website, including how they are linked to one another, how deep they sit within the hierarchy, and how consistently they are reachable from the homepage.
- Crawlability
- The degree to which a search engine crawler can access and follow links across a site. Pages that are not crawlable cannot be discovered through normal crawl processes.
- URL Structure
- The pattern and format of a page's web address. A clear URL communicates the topic and hierarchy of a page to both users and search engines before any content is read.
- Information Architecture
- The way content is organized and categorized across a site, including how topics relate to one another and how those relationships are expressed through navigation and linking.
- Flat Architecture
- A site structure where most pages are reachable within a small number of clicks from the homepage. Flat architectures are more efficient for crawlers and easier for users to navigate than deeply nested ones.
- Breadcrumbs
- A secondary navigation element that shows a user's position within the site hierarchy (for example, Home > Category > Article). Breadcrumbs serve user orientation, crawl efficiency, and eligibility for rich result display in search results.
- Internal Linking
- Links that connect one page on a site to another page on the same site. Internal links create crawl paths, distribute authority between pages, and signal to search engines which pages are related.
- Anchor Text
- The visible, clickable words inside a hyperlink. Anchor text carries a signal about the topic of the destination page that is independent of the destination page's own content.
- Orphaned Page
- A page that no other page on the site links to. Orphaned pages are difficult for crawlers to discover and receive no authority passed through internal links.
- XML Sitemap
- A file that lists a site's pages and submits them directly to search engines as a guide to what exists, rather than relying solely on link discovery. Accuracy is essential: a sitemap that includes broken or low-quality pages can reduce its own usefulness.
- Robots.txt
- A configuration file that instructs crawlers which parts of a site not to crawl. It controls crawl access only, not indexation. A page blocked by robots.txt can still be indexed if other sites link to it.
- Indexation
- The process by which a search engine adds a page to its index, making it eligible to appear in search results. A page must be crawled and evaluated before indexation can occur, but crawling does not guarantee indexation.
- Search Console
- A diagnostic tool provided by Google that shows how Google sees a site, including which pages are indexed, which are excluded, and what crawl errors have been recorded. It functions as a window into Google's view of a site's structure and health.
- Crawl Budget
- The amount of crawling activity a search engine allocates to a given site within a period of time. Sites with poor structure or many low-quality pages may find that crawlers spend their budget on unimportant pages rather than the ones that matter.
- Rich Results
- Enhanced search result formats that display additional information beyond a standard title and description, such as breadcrumb trails, star ratings, or FAQ entries. Eligibility for rich results depends partly on structured data markup.
- Schema Markup
- Structured data added to a page's code that helps search engines understand specific elements of the page, such as breadcrumb paths, product details, or article authorship, in a standardized format.
- Noindex
- A directive placed on a page that instructs search engines not to include it in their index. Unlike robots.txt, noindex allows crawling but prevents the page from appearing in search results.
- Crawl Path
- The sequence of links a crawler follows to move from one page to another across a site. Internal links, navigation menus, and sitemaps all contribute to the crawl paths available to a crawler.