Robots.txt and XML Sitemap: 6 Mistakes on B2B Sites
The robots.txt sitemap xml mistakes that quietly block indexing on B2B and catalog sites, plus a safe file template you can copy and adapt to your platform.
By Downway Team 3 min read
Robots.txt and the XML sitemap are two tiny files that can wipe out a site's traffic with one wrong line. A robots.txt sitemap xml setup is usually done once at launch and never reviewed again. These are the mistakes we see most on B2B and catalog sites.
Mistake 1: leaving the staging block live
During development it is common to block everything with Disallow: /. If that file is copied to the live site, Google stops crawling your pages. Consequence: rankings fall within weeks. How to avoid it: put a robots.txt review in the launch checklist and open the file's URL right after going live.
Mistake 2: blocking CSS, JavaScript or images
Rules that block script or style folders stop Google from rendering the page the way visitors see it. The result is a wrong reading of layout and mobile friendliness. Allow the resources needed for rendering and confirm with URL Inspection.
Mistake 3: using robots.txt to remove a page from Google
Robots.txt prevents crawling but does not guarantee removal from the index: the URL can still appear, just without a description. To remove a page, use a noindex tag and let the crawler reach the page to read it. Blocking and adding noindex together cancels the effect.
Mistake 4: junk URLs in the sitemap
The sitemap is a list of pages you want indexed. Keep these out of it:
- URLs that redirect or return a 404;
- pages with noindex or a canonical pointing elsewhere;
- filter, internal search and session-parameter URLs;
- http versions when the site runs on https.
A dirty sitemap triggers warnings in Search Console and weakens Google's trust in the list. If your catalog site generates the sitemap automatically, check that exclusion rules exist.
Mistake 5: a stale or unsubmitted sitemap
A new product missing from the sitemap takes longer to be discovered. Generate the file automatically with real last-modified dates and submit its address in Search Console. On large catalogs, split it into several sitemaps (products, categories, articles) to make diagnosis easier.
Mistake 6: not pointing to the sitemap in robots.txt
It is not mandatory, but it is one line that helps any crawler find the list. Another frequent slip is writing the sitemap path without the full address.
A safe template for a catalog site
Adapt it to your setup, since each platform has different folders. A conservative starting point:
- User-agent: * (applies to all crawlers);
- Disallow: /search/ and Disallow: /cart/ (pages with no search value);
- Disallow: /*?sort= (sorting parameters, if present);
- Sitemap: https://www.yourdomain.com/sitemap.xml (full address).
Do not block image folders or script and style files, and never ship Disallow: / to production.
A routine check
After any site change, open the file in a browser, view the pages report in Search Console and watch for increases in “blocked by robots.txt”. A quarterly check prevents nasty surprises. For a full technical review, see our industrial websites page.
Frequently asked questions
Does every site need a robots.txt file?
It is not mandatory. Without it Google crawls everything. But the file is useful for keeping crawlers out of low-value areas and for pointing to the sitemap.
How many URLs fit in one sitemap?
Each file accepts up to 50,000 URLs or 50 MB uncompressed. Beyond that, use a sitemap index with several files.
Does robots.txt protect confidential information?
No. The file is public and only guides well-behaved crawlers. Confidential content should sit behind a password or off the site.