If you’re on this page you’re no doubt familiar with Google Search Console. In our SEO Audits we do a deep dive on Google Search Console and it can provide a wealth of insights and action items for improving a website’s SEO.
One common issue we see and more and more since mid 2024 is a large number of pages in the “Alternate page with proper canonical tag” in the indexation section of GSC. It’s not technically an error but it is a problem because these are pages Google is crawling that it probably shouldn’t be in most cases as they are alternate versions of other pages.
Fixing pages that appear in this section of GSC will improve how Google crawls the site and in many cases by fixing canonical errors we can improve rankings often quite dramatically.
In this post we’re going to walk through the process of fixing these pages in GSC – click play on the video below to learn more.
Video transcript: How To Fix "Alternate page with proper canonical tag" in Google Search Console
0:00What the Pages Aren't Indexed Report Tells You
Hey, it's Brendan here from WP Speed Fix. In this video we're going to talk about Google Search Console. One of the things we do in our SEO audits is a deep dive on Google Search Console, and we get a lot of questions about it from customers too. If you're not familiar, well, you're probably familiar if you're watching this video, it's Google's tool that tells you what's happening with your site, what it's doing in terms of indexing the site and how pages are performing. It will tell you a lot of things about errors and problems that are happening under the bonnet, and a lot of these things are easy to fix and can have a big effect on ranking.
I'm going to talk about this particular one. It's not an error necessarily, but it's a warning, in the indexing section, under "Why pages aren't indexed". I'm looking at our WP Speed Fix site here and we have 3155 pages that aren't indexed. Some of these aren't a problem at all, but we're going to dig into this particular area because it's one we see a lot, particularly in the last 6 months or so, since about mid 2024 when AI, particularly Google's AI, became more prominent. They seem to be crawling sites much more aggressively, and we're seeing a lot more warnings or errors pop up under this section of Google Search Console.
Let me dig into this a little bit and show you some examples. When you're in Search Console, ideally you want to be in the domain property, not in the individual website property. I'll just show you this site here, Didgeridoo Dojo. You'll see we have a domain level property, but you can also add all four variations of the URL: HTTPS, HTTP, and no www and with www. You want to be under the domain property because that covers everything, so it'll show you all the errors. Otherwise you'll need to check each individual one. If you use the domain property, that's the best way to go because it gives you the best picture of what's going on. I'll leave that alone because we already loaded it up.
2:00How Canonical Tags Work
As I said, it's under Pages here. Let's talk about canonical tags for a minute, because they're really misunderstood and a lot of people don't know what they are. One of our fundamental rules for SEO is that in order to rank for something you need a page for it, each page can only rank for one thing, and there should be only one page for that thing. A simple way to think about this is a dentist website. If a dentist wants to rank for teeth whitening, dental implants, Invisalign, all the different services they offer or procedures they do, they need a page of content for each one of those on their site. If they don't have a page for teeth whitening, they can't rank for teeth whitening. That's a very simple concept to grasp, and one of the first things we do with our customers is get them to list out all the products and services they sell. If there's not a page for it on the website, then they can't rank for it, so it's a simple case of adding that page.
You also don't want two pages, because Google gets confused. If a dentist website had two pages for teeth whitening, Google doesn't know which page to rank, since you've got two there. The ranking juice is split in half, and they both rank poorly. In SEO terms that's called duplicate content. But there's a middle ground here: each page should be unique, and we want each page to only load at one URL, but in reality pages can load at all sorts of URLs. They could be uppercase and lowercase, and those are counted as different URLs. There could be a forward slash on the end or no slash on the end. There could be all sorts of variations that a single page loads up as.
To solve this problem we have canonical tags. A canonical tag is a tag on a web page that tells Google that this is the primary page for this content, or this is the URL for this content that should rank. I'll give you an example, because I actually came across a problem with this URL. I'll load this up, and I've just got an SEO tool installed, but you can see the canonical tag for this matches the URL. If I add a query string on the end, say UTM, which is a tracking string for Google Analytics, you'll see that now we have a mismatch, so I have an error here in my tool. The URL that I've loaded up has a question mark and UTM on the end, but the canonical tag is still the same. So if Google comes across this URL, the canonical tag will tell it to ignore this page, because this is actually the real page to rank.
Basically, where Google comes across duplicate pages, the canonical tag sorts it out. This section is telling us where Google has come across duplicate pages when it's been crawling a website and has made a decision based on the canonical tag. Ideally we want to minimize this as much as possible because in theory we're wasting Google's crawler time, and any time you give Google mixed signals, that's a problem. It gets confused, so we want to avoid that. We really need to clean up this section with the canonical tag.
5:04Query Strings and Search Pages Being Crawled
Let's talk about the different types of errors you'll have here, and the most common one is query strings being crawled. Let me see if I have it here. Okay, this one is a WooCommerce one, and we see this all the time. You'll see there are nearly 2 and a half thousand pages in here, so this is a bit of a problem. This is a common thing we see with WooCommerce, especially, like I said, in the last six months as Google is getting more aggressive with crawling. You see here Google is actually crawling these add-to-cart URLs, and this is a problem. On the website, when someone wants to add to the cart, this is the URL that loads up: it adds the item to the cart and refreshes the page. We want a normal user to load that up, but we don't want Google to load that up.
This is a problem not just for SEO but for performance as well. Google is coming along to the site, crawling these URLs, and adding things to the cart. There's also a remove from cart query string, which it's not crawling in this case, but in some cases it'll be adding and removing to the cart multiple times per second, and that can chew up a ton of resources on the site. So these query strings are a problem. There are other query strings here where product attributes are being loaded up, and ideally we don't want that happening either. In some cases you'll have an add to wish list. I'm not sure if there is here, we just have attribute. Here we go, we have some other ones, "grid cookie equals list", which is probably review pages, I'm guessing. Let's have a look at what else there is. That's pretty much it, a lot of attribute stuff.
Okay, so those are the query strings. We'll commonly get other ones, not just query strings but search pages, particularly for WordPress. WordPress has two different ways you can search. The URL will have a query string, something like question mark s equals and then the search, or the other way that search pages will be indexed is slash search slash whatever. We don't want those crawled either. That's actually a negative SEO attack vector as well. We often see sites that have had negative SEO done on them and they have thousands of search pages indexed. What happens is a malicious crawler comes along, starts loading up these search pages, and then gets Google to go and index them. Google burns all its crawl budget, has tens of thousands of these pages that it's trying to access, and the whole website basically gets dinged because it looks like spam.
The WP login page is not a query string, but in some cases, if you have protected content, often the crawler will get redirected to a login page. We were having that on the WP Speed Fix website, and I'll show you in a second the robots.txt file that we've made to fix that. Like I said, WooCommerce add to cart pages, add to wish list being crawled, and product attributes. In some cases there are marketing strings, so I'll put these notes in the description. Marketing strings, things like UTM strings, are also a problem, and we don't really want those being crawled. You'll see for the WP Speed Fix website we've got a couple of affiliate strings here, and some of those are being crawled. We really don't want those crawled, and UTM strings as well. It's not a big deal, we just have one, but if we had hundreds that might be a problem. We're going to block those as well.
9:20Fixing Wrong Canonicals, Duplicate URLs and Crawled Query Strings
You also have genuine errors, where the canonical tag is set wrong. We actually had this page where the canonical tag had been manually set, and it was set wrong. So that's a problem too. In this case something went wrong when the content was published and it actually had the incorrect canonical tag. If I just load that up, when you're looking at these and you need more information, you can click here, oops, click here and it will bring up details. If we go to Inspect URL, it'll bring up the details of where it's finding those URLs. Let's just go back for one sec, that's not what we want. I'm in the wrong spot. We want to find out where Google is finding that page. There we go, page is indexed. You'll see here, oh, I've actually just fixed it, so it's not showing. We had the wrong canonical tag, it was manually set with a capital letter. The character case is important, because if you have capitals there, that's a different URL to all lowercase, so that can be a problem as well. This page is now fixed so that's not a good example, but the canonical was set incorrectly, so that was a problem there.
The other issue you commonly see here is a common one for older sites that have been around for a while. The website is loading over HTTP and HTTPS, so it's loading over both versions, there's not a redirect from HTTP to secure, and it's loading from www and no www, so both variations of the URL. There are a couple of fixes here. One is to load the website only over a secure channel, so redirect everything to HTTPS and enable HSTS, which we have a video about how to do in Cloudflare, so I'll link that up here. You also only want to load from either www or no www. Choose one of them, and the website should only load from that one, so there should be redirects from all these other versions to the primary URL. Those should be fixed, and websites typically don't have those errors anymore, but older sites might.
For these query strings, let's just go back there. Adding things to the robots.txt file and blocking these things is the way to fix those. This is the WP Speed Fix one, and I'll link that up. I'll put a note here so you can load that up directly. Here we had the crawler being redirected to the WP login page, so we blocked the crawler from that. We blocked the crawler from feed pages and search pages because we're getting a lot of search spam. This is patching some issues: the hosting is cloning the site and some of those URLs were getting accidentally indexed. This is another search string as well. We don't want searches indexed at all. Search is in there twice for some reason, but we really want this one here: disallow search followed by an asterisk. The asterisk is like a wildcard. We're also having some of the WordPress API crawled. So that was the robots.txt file we set up to fix the indexation issues.
If you have some of this add to cart stuff, like on this WooCommerce site, we want to have something like disallow, and we do add to cart equals and then an asterisk. We don't have the question mark in front of it, because if there are multiple query strings, you see in this one here where the attribute has an asterisk instead of a question mark, then if we had a question mark this variation of the URL would not have been picked up, and we want to block those ones as well. So I want an asterisk before the attribute. In this case I'll do attribute, underscore, asterisk, because you can see there are multiple attributes. One is attribute underscore pack size equals 35, and then this was attribute underscore color. There are multiple attributes being loaded up here and we want to block all of them, so we've just used this wildcard, which will stop all of those. We really don't want those URLs crawled because it's burning crawl budget and wasting Google's time. Again, we're chewing resources on the hosting itself, and we're also confusing Google. We're sending mixed signals.
12:34Crawl Stats and Validating Your Fixes
If you want to see what Google is doing with its crawling, I'll just move this out of the way. Over here, under Settings, you can also see what's happening with the crawl requests. This is another area to look at if you're trying to fix these canonical issues. Quite often, if you go in here and make this bigger, you can click here and see where errors are happening, things like 404 errors. That's all good, we don't necessarily want Google crawling this, so we might want to block this well-known page. There we go, well-known is coming up a couple of times. We have 820 crawl errors here, so we have a lot of issues there. We want to avoid Google crawling this well-known page for sure. Some of these other ones we need to add redirects for, but this one here, well-known, is the problem. Inside this crawl report you can get a better handle on what Google's doing, and you can block some of this stuff. We don't want Google crawling these JSON URLs as well.
We probably want to block Google from crawling this refresh fragments one, which has to do with the checkout in the header on WooCommerce. We want to avoid that being crawled, because again that's just chewing resources and hammering the hosting. There's a lot of stuff here that is trying to crawl that it shouldn't be, so we'd probably want to block those ones as well. The reason they're not showing up in the Pages section is that they're not full pages. So that's what's going on there.
That's pretty much it for this video. Once you've gone in to fix them, you want to click here and do Validate Fix. If your robots.txt file is working or the issues are fixed, they will go away. They won't go away permanently, they'll end up under here. I don't know if it has it on this website, I'll just switch across. Let me just go back here. You can see I actually started the validation on this one because I fixed this before the video. You'll see under "Blocked by robots.txt" that the errors and warnings will move over to here, and that's not an issue. These aren't necessarily problems, it's just telling you why pages aren't indexed.
You can see now, because we've put the WP login as a disallow in the robots file, it's being blocked. Google is trying to crawl it and it's being blocked. Eventually, over time, these will go away, but often it can take several months, even years, for this to go away. Those canonical errors or canonical warnings will move over here, so you won't ever get rid of them entirely. They might drop away over time, but they'll shift over here, and this isn't really a problem because we don't want any of this stuff crawled anyway.
15:15Free SEO Tools and SEO Audit Options
I think that's it for this video. That should be a good enough walkthrough. I'll add these notes and I'll link you up to this HSTS video in the description below. If you want more help with your site, head over to our website, wpspeedfix.com, and there are a few different ways we can help. For SEO, run our free SEO audit tool here. All these tools take 60 to 90 seconds and there's no opt-in required. Run an SEO audit and it will give you a breakdown of what's going on at the site, and a quick overview of any technical SEO issues that are immediately, obviously fixable. That's 90 seconds to run. This one here is the Core Web Vitals report. If you have enough traffic, you need 50 to 100 visits a day to have enough data to generate this report, it will give you a breakdown of your site speed and how real visitors are experiencing the site.
Then there's also our site speed test tool. This will give you detailed recommendations on how to improve your site speed. It also does some basic SEO checks, some different ones to the SEO audit, so it's worth running them both. We also have an SEO audit service, so if you want to deep dive in there, we have our technical and on-page SEO audit service. There are some example screenshots there of what that looks like, and we have two options: audit only, and then audit with implementation as well. If you're not sure, head over to the homepage and request a free audit here. We'll have a look at your site and one of the team will tell you what we can do, or how we can help, and come back to you. For any other questions, just post in the comments below. I hope you found this useful. Cheers.
Table of Contents
Understanding Canonical Tags
Before diving into the warning itself, let’s clarify what canonical tags are. In essence, a canonical tag is a way to tell search engines which version of a page is the preferred one to index when there are multiple versions with similar content. This is crucial because duplicate content can confuse search engines and dilute your site’s ranking potential.
You’re going to see this warning a lot with bigger sites and especially Woocommerce sites that by their very nature typically have multiple versions of a page due to things like product attributes or variations.
Common Causes of the Warning
The “Alternate page with proper canonical tag” warning typically arises when Google encounters duplicate pages on your site, but it can determine the preferred version based on the canonical tag. Here are some frequent culprits:
**-Query strings being crawled —Things like wp-login.php page being crawled —Search pages being crawled —In Woocommerce, add-to-cart, add-to-wishlist being crawled, product attributes —Marketing strings like UTM tags and Facebook FBCLID tags —Fix: update robots.txt to block these For Woocommerce: Disallow: add-to-cart= Disallow: attribute_ Disallow: /search/ Disallow: s= https://www.wpspeedfix.com/robots.txt **-The website loading on http AND https or www AND no-www —Fix here: load only over HTTPS —Enable HSTS, explainer video here https://www.wpspeedfix.com/hsts-reduc… —Only load from either WWW or no-WWW
- Query Strings: These are parameters added to URLs, often for tracking purposes or to filter content. Google might crawl these variations as separate pages, even though they essentially have the same content.
- Login pages being crawled: Wp-login.php or similar logon pages being crawled
- WooCommerce Issues: WooCommerce sites often generate dynamic URLs for actions like adding to cart or filtering products. These can trigger the warning if not properly managed.
- Search Pages: WordPress search results can create numerous URLs with duplicate content, especially if indexed by Google.
- Marketing Strings: UTM parameters used for tracking campaigns can also lead to duplicate content issues.
- Incorrect Canonical Tags: In some cases, the canonical tag itself might be set incorrectly, pointing to the wrong version of the page.
- Incorrect SSL Configuration: The website loading on both HTTP AND HTTPS
Addressing the Warning
Fortunately, there are several steps you can take to resolve this warning and improve your site’s SEO:
- Optimize Your robots.txt File: This file instructs search engines which pages to crawl. By adding specific rules, you can prevent Google from indexing unnecessary pages like those with query strings or search results.
- Fix Incorrect Canonical Tags: If the warning stems from incorrect canonical tags, ensure they are pointing to the correct, preferred version of the page.
- Address WooCommerce Issues: For WooCommerce sites, consider blocking unnecessary query strings and optimizing product filtering to minimize duplicate content. See the video below for more on this.
- Use 301 Redirects: If you have multiple versions of a page accessible through different URLs, implement 301 redirects to consolidate them to the preferred version.
- Fix SSL Redirect Issues & Enable HSTS: Ensuring all versions of URLs are loading via HTTPS and HTTP URLs redirect correctly will solve this issue and will also help improve site speed as http URLs cannot use the HTTP2 protocol. Getting HSTS enabled on your hosting or in Cloudflare will help even more!
- Validate Your Fix: Once you’ve implemented the necessary changes, use Google Search Console’s validation tool to confirm that the issue is resolved.
Robots File Specifically for Woocommerce
As above, WooCommerce sites are particularly prone to seeing this warning appear in Google Search Console. The video below walks you through a robots.txt file we’ve specifically developed for Woocommerce sites to resolve this error, boost SEO and performance.
/image
Read the full transcript of this video on Optimized robots.txt file for better WooCommerce SEO & Site Speed.
Need Help?
If you’re struggling with Technical SEO issues and wrangling with Google Search Console we can help. A good starting point is either our SEO Audit Service or alternatively our Google Search Console Audit service
If you head to our homepage and submit a FREE site audit along with some background about the problem you’re looking to solve or goals you’re looking to achieve, one of our team will review your site and come back to you with recommendations as to how we can help.