Google crawls your site with a finite budget. Every second spent on a low-value page is a second not spent on content that matters. Most sites waste this budget by including every URL in their XML sitemap, treating it like a to-do list instead of a strategic tool.
The assumption is simple: more URLs in the sitemap means faster discovery. In practice, an oversized or poorly curated sitemap signals to Google that your site has more unique content than it actually does. This spreads crawl resources thin and delays indexing of your core pages.
This post covers which URLs belong in your sitemap, how to identify crawl-budget waste, and the audit steps to reclaim it.
What Crawl Budget Actually Is
Crawl budget is the number of URLs Google will crawl on your site in a given time period (usually per day). It has two components: crawl capacity (how many URLs Google's infrastructure can crawl) and crawl demand (how many URLs Google wants to crawl based on importance signals).
Google doesn't publish your crawl budget, but you can infer it from the Crawl Stats report in Google Search Console. If you see a plateau in URLs crawled per day, and your site has more unindexed pages, you've hit a budget ceiling.
The trap: Most sites assume bigger sitemap = higher crawl demand. The opposite is true. A sitemap stuffed with duplicate, thin, or redirect URLs tells Google your site is larger and less focused than it is. Google then spreads its crawl across more URLs, indexing fewer of your high-value pages per day.
URLs That Should Never Be in Your Sitemap
Start by removing these categories. They waste budget and add noise:
- Redirect chains. If a URL redirects to another URL, remove the redirect source from the sitemap. Google crawls the redirect, discovers the target, then crawls the target. Two crawls for one page. Include only the final destination.
- Canonicalized duplicates. If a page has a
rel="canonical"pointing elsewhere, remove it from the sitemap. Google will crawl it to find the canonical, then crawl the canonical. Keep only the canonical version in the sitemap. - Noindexed pages. A page with
noindexin the robots meta tag or HTTP header should not be in the sitemap. Google still crawls it (to read the noindex directive), but it doesn't index it. This is pure budget waste. Remove it. - Pagination without
rel="next/prev". If you have numbered pages (page 2, page 3, etc.) and you're not usingrel="next"andrel="prev"or self-referential canonicals, each page competes for budget independently. Either implement pagination markup or keep only page 1 in the sitemap. - Session IDs and tracking parameters. URLs with session tokens, UTM parameters, or other tracking variables should not be in the sitemap. Each unique parameter string becomes a separate URL in Google's index, fragmenting authority and crawl budget.
- Admin, login, and search pages. Remove
/admin,/login,/search,/cart, and other utility pages. They don't rank, don't need discovery, and waste budget. - Duplicate content with different parameters. Product filters (color, size, price) often generate hundreds of parameter combinations. If you're not using
rel="canonical"to consolidate them, each variant wastes budget. Use canonical tags or parameter handling in Search Console instead.
The Core/Supplemental Crawl Split
Google divides crawl budget into two pools: core crawl and supplemental crawl. Core crawl prioritizes pages Google deems important (high authority, frequent updates, user demand). Supplemental crawl handles lower-priority content.
Your sitemap influences which pool a page lands in. A page in your sitemap with no internal links, low engagement metrics, and infrequent updates will be crawled in supplemental, if at all. If you include it, you've signaled it's worth discovering, but Google's data says it's not. This creates friction.
The fix: Keep your sitemap small and high-signal. Only include pages that are either (1) core ranking targets, (2) updated frequently, or (3) difficult to discover through internal links alone. Everything else should be discoverable through your information architecture and internal linking.
Sitemap Size and Indexation Rate
A common metric to track: the ratio of indexed pages to sitemap URLs. If your sitemap has 10,000 URLs but Google indexes only 3,000, you have a problem.
This ratio signals one of three things:
- Your sitemap includes low-quality or duplicate content that Google doesn't consider worth indexing.
- Your crawl budget is insufficient to crawl all URLs in a reasonable timeframe, so many remain undiscovered.
- Your pages have
noindextags or other directives preventing indexation, yet they're still in the sitemap.
If your indexed-to-sitemap ratio is below 50%, audit your sitemap. Remove redirect sources, canonicalized duplicates, and noindexed pages. Then check whether the remaining URLs are truly ranking targets or just "nice to have" content.
How to Audit Your Sitemap
Start with Google Search Console. Pull the Crawl Stats report to see how many URLs Google crawls per day. Then export your sitemap and compare.
Use a spreadsheet or crawl tool to identify patterns in unindexed URLs. Common culprits:
- Pages with low word count (under 300 words)
- Pages with no internal links pointing to them
- Pages updated more than 6 months ago
- Pages with high bounce rate or low engagement metrics (if you have GA4 access)
- URLs with duplicate or near-duplicate content
If a page fits more than one of these patterns, it's a candidate for removal. Run a quick check: Does this page rank for any keywords? Does it receive organic traffic? If the answer is no to both, remove it from the sitemap.
For large sites (10,000+ pages), prioritize your sitemap by section. Create separate sitemaps for your core content (blog, product pages, guides), then a secondary sitemap for supporting content (FAQs, help articles, archive). Submit the core sitemap to Google Search Console and hold the secondary one back. Monitor crawl stats for 2–4 weeks. If crawl budget improves on core pages, you've found your threshold.
Frequency and Freshness Signals
The lastmod attribute in your sitemap tells Google when a page was last updated. This is a hint, not a directive. Google uses it to prioritize crawling pages that change often.
The mistake: Setting lastmod to today's date on every page, every time you deploy your site. This signals that all pages change constantly, so Google treats none as a priority. Keep lastmod accurate. Update it only when the page content actually changes, not when you deploy infrastructure updates.
The same logic applies to priority and changefreq` attributes. If you set priority to 1.0 (highest) on all pages, it's worthless. Google ignores these signals when they're uniform. Use them only if you're willing to differentiate (product pages 0.9, blog posts 0.7, archive 0.3). Otherwise, omit them.</p>
<h2>When to Use Multiple Sitemaps</h2>
<p>Google allows up to 50,000 URLs per sitemap file and 50 MB per file. If you hit either limit, use a sitemap index file to reference multiple sitemaps.</p>
<p>But you don't need multiple sitemaps just because you can. Use them strategically: one for high-priority content, one for supplemental content, one for media (if you have hundreds of images or videos). This lets you signal importance through structure and manage crawl budget more granularly.</p>
<p>A sitemap index is also useful if different parts of your site are maintained by different teams. Your blog team can manage the blog sitemap; your product team manages the product sitemap. Both feed into a single index that you submit to Google Search Console.</p>
<h2>Reality Check: Internal Links Matter More</h2>
<p>Your sitemap is a discovery tool, not a ranking signal. Internal linking architecture is far more important for crawl budget and ranking.</p>
<p>A page buried three clicks deep from your homepage, with no internal links pointing to it, will not rank well even if it's in your sitemap. Conversely, a page with strong internal linking (anchor text, placement, link authority) will be crawled and indexed even if it's not in your sitemap.</p>
<p>The best crawl budget strategy is not a perfect sitemap. It's a clear information architecture with strategic internal linking. The sitemap is the safety net, not the foundation.</p>
<h2>What to Do Next</h2>
<p>Pull your sitemap and your Search Console indexation report. Calculate your indexed-to-sitemap ratio. If it's below 70%, start removing redirect sources, canonicalized duplicates, and noindexed pages. Then run a crawl audit to identify pages that don't rank and receive no traffic. Remove those next.</p>
<p>Monitor your Crawl Stats report weekly for the next month. You should see either an increase in crawl volume (Google is more confident in your site structure) or a faster crawl of your core pages (Google is focused). Either is a win.</p>
<p>If you need help auditing your sitemap structure or identifying which pages waste budget, a <span class="iolink-hint">technical SEO audit</span> can pinpoint the issue and provide a removal roadmap.</p>
<hr>
<h2>FAQs</h2>
<p><strong>Q: Does my sitemap affect my rankings?</strong><br>A: No. Your sitemap is a discovery tool, not a ranking factor. Google uses it to find and crawl URLs, but the presence or absence of a page in the sitemap doesn't influence its ranking position.</p>
<p><strong>Q: Should I include images and videos in my sitemap?</strong><br>A: Only if you want Google to crawl and index them. Use image and video sitemaps if you have rich media that's not discoverable through page content or internal links. Otherwise, omit them.</p>
<p><strong>Q: How often should I update my sitemap?</strong><br>A: Update it when you add, remove, or significantly change pages. If your site is static (no new content), update it quarterly or annually. If you publish frequently, automate sitemap generation so it's always current.</p>
<p><strong>Q: What's a good crawl budget?</strong><br>A: There's no universal benchmark. Small sites (under 1,000 pages) typically see 100–500 URLs crawled per day. Large sites (10,000+ pages) might see 10,000+. The key is consistency. If your crawl volume is declining and you haven't changed your site, you may have a quality or structure issue.</p>
<hr>
<h2>People Also Ask</h2>
<p><strong>Q: Why is Google not crawling my new pages?</strong><br>A: New pages need discoverable paths to your homepage. Add them to your sitemap, but also link to them from existing high-authority pages. Google crawls links first; the sitemap is secondary.</p>
<p><strong>Q: Can I use robots.txt to block crawling instead of removing from the sitemap?</strong><br>A: No. If a URL is in your sitemap, Google will try to crawl it regardless of robots.txt. Use robots.txt to block crawling; remove the URL from the sitemap. Don't use both on the same URL.</p>
<p><strong>Q: How do I know if my crawl budget is the problem?</strong><br>A: Check Google Search Console's Crawl Stats report. If the number of URLs crawled per day is flat or declining, and you have unindexed pages, crawl budget may be the issue. Compare against your sitemap size. If you have 50,000 URLs in the sitemap but Google crawls only 500 per day, you'll need years to crawl everything.</p>
<p><strong>Q: Should I remove old blog posts from my sitemap?</strong><br>A: Only if they're noindexed, canonicalized, or receive no traffic. If an old post still ranks for a keyword or generates organic traffic, keep it in the sitemap. Google will crawl it less frequently (supplemental crawl), but it's still valuable.</p>
<p><strong>Q: What's the difference between a sitemap and a sitemap index?</strong><br>A: A sitemap is a file listing URLs on your site. A sitemap index is a file listing multiple sitemaps. Use an index when you have more than 50,000 URLs or want to organize sitemaps by content type or priority.</p>
<p><strong>Q: Can I submit multiple sitemaps to Google Search Console?</strong><br>A: Yes. You can submit individual sitemaps or a sitemap index. Google will crawl all of them. A sitemap index is cleaner for large sites.</p>
<p><strong>Q: Does the order of URLs in my sitemap matter?</strong><br>A: No. Google doesn't crawl URLs in the order they appear in the sitemap. It uses other signals (link authority, freshness, user demand) to prioritize.</p>
<p><strong>Q: What happens if I include a URL in my sitemap that's already canonicalized?</strong><br>A: Google crawls it to read the canonical tag, then crawls the canonical URL. This wastes crawl budget. Remove canonicalized URLs from the sitemap and keep only the canonical version.</p>
<p><strong>Q: How do I handle pagination in my sitemap?</strong><br>A: Use <code>rel="next" and rel="prev" tags on paginated pages, or self-referential canonicals on each page. Then include only page 1 in the sitemap. Google will discover subsequent pages through the pagination markup.
If this post is wrong, outdated, or you would take a different path
I write from work I have done on real sites. Search products change, and a step that was right when I published can go stale. I can also be wrong about the method.
If you disagree with the approach, the facts, or the outcome, I want the detail. Tell me what is off, what you would do instead, and where you saw it. I use that to correct the post so the next reader is not stuck.
This is not a comment thread. Use Contact me so the note is tied to this post and I can reply.
You are sending feedback for
Why Your XML Sitemap Strategy Might Be Hurting Your Crawl Budget
Technical SEO
https://hammadshk.com/blog/why-your-xml-sitemap-strategy-might-be-hurting-your-crawl-budget