Video transcript: WooCommerce Robots.txt For Better Site Speed & SEO
0:00Introduction to robots.txt
Hey, Brendan here from WP Speed Fix. In this video I want to talk about a robots.txt file that we use for WooCommerce sites to speed them up, reduce the load on the site and improve SEO. So you're probably familiar, if you're watching this video, that the robots.txt file is a file that all websites should have. Let me just show you ours, wpspeedfix.com/robots.txt. It should be there for every site.
Basically this tells crawlers like Google what they can crawl on the site and how they can crawl the site. If they're blocked, basically you deny them access to certain things. Not all crawlers will honor the robots.txt file, so keep that in mind, but big search engines and big tools do. It's best practice for crawlers to honor it.
We have a problem with WooCommerce sites. They have a lot of URLs, they have a lot of query strings, and they have a lot of dynamic things going on that, depending on how a theme is built or a site is built, can cause a lot of load on the site, and it can actually knock it over. By using exclusions or disallows in the robots file we can improve performance, stop that happening, and stop Google crawling things that it shouldn't. We can also block SEO tools that are aggressively crawling the site.
1:31Basic Crawl Optimization: Crawl Delay, Login and Search Pages
We have this example file, and this is kind of our base example. If you go to wpspeedfix.com/woocommerce-robots.txt you can copy and paste this and use it yourself. There are some comments here that explain what each section is, and we will update this over time, so it might look a bit different to this video if you're watching this in the future. Let me talk you through what these lines mean. We actually built this recently for a WooCommerce site that was having some load issues, and figured that we should make this a template and make it available.
So let's talk about it. We've added a crawl delay here. Because crawlers can be quite aggressive, they can load multiple URLs per second. We commonly see aggressive SEO crawlers loading 3 to 5 URLs per second, so that's mimicking five users on the site per second, which is quite a lot. If it's a small site, if you're on a small hosting plan, that can overload the hosting quite quickly. So if you add a crawl delay, crawl-delay 5 means that the crawler can only crawl the website every 5 seconds. Instead of multiple times per second, that's one crawl every 5 seconds, so that's quite reasonable. A lot of hosting companies do this by default. They'll have a crawl delay of 5 or 10 just to slow down crawlers and reduce load on the site.
Keep in mind that there is a scenario where this may have an impact on SEO. If you have a bigger site, if you've got thousands of URLs, this is going to slow down crawling, so a crawl delay of 5 might be too high, and there might not be enough time to crawl the site in a reasonable time period. Just keep that in mind. But for a mid-size WooCommerce site, let's say a 1,000 products, that's going to have maybe a couple of thousand pages, and this crawl delay is fine. The crawl delay slows down crawlers so they're not just chewing up all your CPU juice and resources. I think that makes sense. You can move this up and down. A crawl delay of 1 would even be good, considering we see crawlers hitting the site multiple times per second. I think that's self-explanatory.
Let's talk about login pages. We don't want crawlers crawling login pages, and this can happen in scenarios where you have hidden content or where crawlers are being redirected to login pages. What this does is just stop crawlers hitting login pages: the backend WP admin, the wp-login.php, and then the My Account page. There's no reason that a crawler should be hitting any of those, including the My Account page. This should be noindexed as well, but that just stops the crawler hitting those pages that it shouldn't.
Pretty straightforward, search pages. We don't want crawlers hitting search pages. WordPress can be searched in two different ways, with the search query string and also /search/ whatever. We don't want searches being indexed, and this is actually a vector for negative SEO. If your search pages are indexable, someone can hit your site with a negative SEO attack and it can cause problems. By disallowing those search pages we protect from negative SEO attacks as well, and these pages just shouldn't be crawled anyway.
3:29Advanced WooCommerce Blocks: Add to Cart, Filters, Feeds and Staging
The next one: on some sites or some themes, the add to cart and add to wishlist requests, the crawler will hit them. It'll hit the query string, add to cart, add to wishlist. What ends up happening is that multiple things are being added to the cart and added to the wishlist per second. If you have 3 to 5 crawlers per second hitting the site, that's 3 to 5 add to carts per second, or add to wishlist. In this example, for this customer, this was the problem. There were some aggressive crawlers hitting the add to cart link, and it was actually a targeted attack. We actually had this blocked in Cloudflare, but the rule wasn't quite right, so it was bypassing the rule and crawling those add to cart and add to wishlist links.
That again mimics having five people per second on the site adding things to the cart, which is quite a lot for a low to mid-size WooCommerce site. So this just stops that happening. You'll also see in Google Search Console some of these things, if the theme is pushing links to these query strings, and you'll see this as a problem in Google Search Console as well. So this is actually good for SEO and good for speed too.
This one's kind of an open one. A lot of WooCommerce sites have a search filter. Here's just a couple of common ones: order by and currency. Again, we typically don't want search filters crawled, but just be mindful that some sites have built their site structure around using search filters instead of proper pages. If you do block search filters, that may be an issue for your site structure, so just keep that in mind. Generally speaking, we don't want search filters being crawled. Yours might be filter by or something like that, so edit it as needed.
This one, we don't want to crawl the WordPress feed URL. That's a simple one, so we block that. If you're using RSS feeds for things, that may be a problem, so just keep that in mind. This is a common one: cPanel and Plesk's smart update system creates test sites, and in some edge case scenarios they may be crawled if you have a method or a plugin that is pushing indexation of pages. That's just a simple one if you're on Plesk or cPanel.
This one we add in because we block WP admin here and we want this admin-ajax call to be accessible, because it does impact particularly variations and some things like that. There are some edge cases where this is important. We want a sitemap URL. We always want to link to the sitemap in the robots file, so change that out to your sitemap link. You see it's yourdomain.com right now.
6:38Blocking Aggressive SEO Bots and Scrapers
Lastly, we want to block aggressive SEO tools. Ahrefs and Semrush are two very popular SEO tools, so in this case we've blocked them both here. That will stop them crawling the site. There's also a secondary benefit, not just the performance benefit: it will stop these tools from releasing SEO information, or sharing SEO information, about your site as well. So there's a little bit of SEO benefit there too.
I've also linked to this file that has a lot of entries. You can basically fatten this section out. These are basically bad bots or bad crawlers, so if you want to have the full list and give your site a little bit more protection, I've linked directly to those here so you can just basically copy and paste that. There's a lot in there, and you might just want to check that you're not using any of those tools, because some of them are genuine tools like Ahrefs and Semrush. Just go down, copy and paste it to the bottom of your robots file and you'll be good to go. These are common scrapers and crawlers and things that can be a bit malicious, so we just want to block those by default. We just link to that rather than fattening that page out a bit. Some of these ones, like this PetalBot, can particularly be aggressive, so that can be useful as well, and again, improving the performance of your site.
7:45Conclusion and Free Tools
So that's pretty much it. You can use this for other types of sites as well, so it's not just WooCommerce related, but this is our template for WooCommerce. If you want SEO help with your site or speed help, head to our website. We have some tools there. If you go to wpspeedfix.com, under more free stuff we've got four free tools: a free SEO audit, a free Core Web Vitals report, a free site speed test and a website valuation tool. The first three don't require an opt-in, so you can use those. They take 60 to 90 seconds to run, no opting required, so check those out.
If you want help with your site or with the SEO, submit a free audit on our homepage, and one of our team will have a look at your site and come back to you within a day or so. If you have any questions, just post in the comments below. Cheers.
WooCommerce sites are prone to aggressive crawling and unusual crawling of query strings, add to cart and add to wishlist links.
In this post and video we a share a robots.txt template you an use for your own site that will help reduce load on your site, improve site speed and improve your SEO through better crawling from Google.
Click this link to get the sample robots.txt file in the video: https://www.wpspeedfix.com/woocommerce-robots.txt
It’s also included inline below so you can copy and paste it.
Click play on the video to learn more – notes on the various recommendations in the video are below.
Head over to our FREE Site Audit page and provide some detail on where you’re stuff or what you’re looking to achieve and one of the team will review your site and tell you how we can help.
User-agent: *
#Added to slow down aggressive crawlers from causing a denial of service attack
Crawl-delay: 5
# Block crawling of logon pages
Disallow: /wp-admin/
Disallow: /*wp-login.php*
Disallow: /my-account/*
# Block search pages
Disallow: *s=*
Disallow: */search/*
# Block add to cart and wishlist links, some themes link directly to these and can cause high CPU usage
Disallow: *add-to-cart*
Disallow: *add_to_wishlist*
# Block common query strings, you may want to block other filter strings if you theme has a sidebar filter
Disallow: *currency=*
Disallow: /*?orderby*
# Block the WordPress feed URLs
Disallow: */feed/*
# Block Plesk and Cpanel smart update test site crawling
Disallow: *wp-toolkit*
# We need this allow link because we blocked wp-admin earlier
Allow: /wp-admin/admin-ajax.php
# Change this to your sitemap link
Sitemap: https://www.yourdomain.com/sitemap_index.xml
# Block aggressive SEO tools
# See more at https://github.com/mitchellkrogza/apache-ultimate-bad-bot-blocker/blob/master/robots.txt/robots.txt
User-agent: AhrefsBot
Disallow: /
User-agent: Semrush
Disallow:/
User-agent: SemrushBot
Disallow:/
Table of Contents
Implement a Crawl Delay
Aggressive crawlers can chew up your server’s CPU resources. Adding a Crawl-delay directive tells bots they must wait a certain amount of time between hits. For example, a Crawl-delay: 5 means bots can only crawl once every 5 seconds. Not all crawlers honour the crawl delay, e.g. Google ignores it but that doesn’t mean its completely useless.
Note: For very large sites (thousands of products), a high delay might prevent full indexing, so adjust accordingly
Block Login and Account Pages
There is no reason for search engines to crawl standard WordPress login areas or customer account pages. You should specifically disallow:
/wp-admin/
/wp-login.php
/my-account/
Disallow Internal Search Results
Allowing Google to index your internal search results is often a vector for negative SEO attacks. It is best practice to block crawlers from accessing these dynamic search pages to prevent them from being indexed.
Stop “Add to Cart” & “Add to Wishlist” Crawling
A major performance killer on WooCommerce sites occurs when bots crawl “Add to Cart” or “Add to Wishlist” links. This forces the server to process these actions as if real customers were performing them 3 to 5 times a second, causing massive load spikes. These query strings should be disallowed.
Filter Out Dynamic Query Strings
WooCommerce uses many filters, such as ?orderby= or currency switchers. Generally, you do not want these crawled as they create duplicate content issues and waste crawl budget.
Caution: Ensure your site structure doesn’t rely on these filters for standard navigation before blocking them.
Essential Inclusions: Admin-Ajax and Sitemaps
While you want to block most of /wp-admin/, you must allow /wp-admin/admin-ajax.php. WooCommerce relies heavily on Ajax for functionality like product variations, and blocking it can break your site’s features for both users and bots.
Always include a direct link to your sitemap_index.xml at the bottom of your robots file so crawlers can easily find your preferred URLs.
Block Aggressive SEO Tools
Tools like Ahrefs and Semrush can be very aggressive when crawling. If you don’t need their data, blocking them can save server resources and keep your site’s competitive data private.
Don’t Use Search Queries or Filters In Your Menu Structure
It’s not mentioned in the video but we regularly see sites that have built their menu system in WooCommerce partly using URLs that are a search string. In some cases, menus are linking to filters directly. This is generally a bad idea for speed AND SEO because these pages are not cached by default and they’re also not real pages so not counted from a SEO perspective.
Setting up caching for these query strings can help in some instances but you’re better off using real pages instead of query strings. How to achieve this really depends on the number of products, types of products and how you’re using the filters and searches. Broadly speaking, product categories and product tags are generally the best way to create various groupings of products.
Use Cloudflare Firewall Rules To Protect From Malicious Crawlers
While genuine crawlers like Google’s crawler bot honor the robots.txt file, many crawlers do not and will continue to aggressively crawl add-to-cart URLs and other URLs they shouldn’t be.
For this reason we also use Cloudflare Firewall rules to filter this traffic. If you check this post on Cloudflare Firewall Rules for WordPress it’ll walk you through some of the most common rules we use for WordPress and WooCommerce sites.