As an SEO I know that having duplicate content on a website hinders its chances at ranking high in Google’s search results. For those who are new to SEO I’ll explain why and for those who are experienced with SEO I’m also going to cover 2 different scenarios in duplicate content as well as explain how Google’s search algorithm deals with each of them and how that affects SEO performance.
To better understand how Google Search deals with each duplicate content scenario I had to find multiple sources and piece it all together since Google’s guidelines don’t have it all in one place and they don’t cover everything.
If you’re new to SEO then perhaps you’re aware that Google’s Search Console General Guidelines state; “Avoid creating duplicate content”.
SEO’s will know Google’s Webmaster Guidelines cover best practices and specific guidelines outlining things to avoid. In there Google says avoid “creating pages with little or no original content” and in another point says avoid “scraped content”.
To start, let’s just be clear about what duplicate content is. Since Google is the most popular search engine in the world, let’s go with their definition. Google’s Webmaster Guidelines state; “Duplicate content generally refers to substantive blocks of content within or across domains that either completely match other content or are appreciably similar.”
As per the definition, there are 2 scenarios when it comes to duplicate content. The first scenario is “within domains” and the second is “across domains”. The reason I’d like to discuss each scenario is that Google search deals with them differently.
In Google’s Webmaster Blog they identify these 2 scenarios as follows:
“Generally, we can differentiate between two major scenarios for issues related to duplicate content:
Basically, it’s copied content and therefore not original content. It’s taken from another website and republished with little or no changes.
Here’s how Google explains it:
“Some webmasters use content taken (“scraped”) from other, more reputable sites on the assumption that increasing the volume of pages on their site is a good long-term strategy regardless of the relevance or uniqueness of that content. Purely scraped content, even from high-quality sources, may not provide any added value to your users without additional useful services or content provided by your site; it may also constitute copyright infringement in some cases.
It’s worthwhile to take the time to create original content that sets your site apart. This will keep your visitors coming back and will provide more useful results for users searching on Google.
Some examples of scraping include:
In the first scenario, within-your-domain-duplicate-content MOZ provides a great explanation for why this is a problem for search engines and why it’s a problem for site owners and particularly SEO’s. Here’s how they put it:
“While not technically a penalty, duplicate content can still sometimes impact search engine rankings. When there are multiple pieces of, as Google calls it, “appreciably similar” content in more than one location on the Internet, it can be difficult for search engines to decide which version is more relevant to a given search query.
For search engines
Duplicate content can present three main issues for search engines:
For site owners
When duplicate content is present, site owners can suffer rankings and traffic losses. These losses often stem from two main problems:
This is the most common scenario that SEO’s deal with and Google Webmaster Guidelines outline how to deal with this. Here’s what Google recommends you do in this scenario.
If your site contains multiple pages with largely identical content, there are a number of ways you can indicate your preferred URL to Google. This is called “canonicalisation”.
There are some steps you can take to proactively address duplicate content issues, and ensure that visitors see the content you want them too.
Here’s what they’ve also added on the Google Webmaster Blog;
“You can take matters into your own hands to avoid Google indexing duplicate content on your site. Check out Adam Lasnik’s post Deftly dealing with duplicate content and Vanessa Fox’s Duplicate content summit at SMX Advanced, both of which give you some great tips on how to resolve duplicate content issues within your site.
Here’s one additional tip to help avoid content on your site being crawled as duplicate: include the preferred version of your URLs in your Sitemap file. When encountering different pages with the same content, this may help raise the likelihood of us serving the version you prefer. Some additional information on duplicate content can also be found in our comprehensive Help Center article discussing this topic.”
In the case of cross-domain-duplicate-content, this generally occurs from scraped content. That means content from your website has been taken and published on another website without your permission. Or, from the point of view of the scraper, you have taken content from another website and published it on yours as your own.
I would like to approach this scenario from both points of view and discuss how each website might perform in Google search.
If content has been scraped from your website then you don’t need to worry too much about your site suffering in search rankings or the copycat site outperforming yours. Google is pretty good at identifying the original source and it will only show this version of the content in search results as the algorithm won’t serve up 2 similar pages, it picks the best pages to show in its results.
Here’s how Google explains this scenario:
“You might have the case of someone scraping your content to put it on a different site, often to try to monetise it. It’s also common for many web proxies to index parts of sites which have been accessed through the proxy. When encountering such duplicate content on different sites, we look at various signals to determine which site is the original one, which usually works very well. This also means that you shouldn’t be very concerned about seeing negative effects on your site’s presence on Google if you notice someone scraping your content.”
If you have scraped content what this means for your website is that it has little to no chance of ranking in Google’s search results.
Google has a whole page discussing duplicate content issues. Most of the page covers duplicate content on your website with the last paragraph covering duplicate content found on another site. This would be a case where you see content you created for your website being used on another website without your permission.
Here’s how Google puts it:
“In rare situations, our algorithm may select a URL from an external site that is hosting your content without your permission. If you believe that another site is duplicating your content in violation of copyright law, you may contact the site’s host to request removal. In addition, you can request that Google remove the infringing page from our search results by filing a request under the Digital Millennium Copyright Act.”
If you have scraped content then your website will most likely not rank in Google’s search results. This means the website won’t get traffic from Google. If you’ve created a website with hopes to attract traffic from Google or search engines then scraping content is not going to help. You need to create original content.
Your website is also at risk of being removed from Google’s index should someone find out and report you to Google.
In conclusion, understanding and addressing internal duplicate content is crucial for maintaining a website’s SEO performance. Duplicate content, whether within your domain or across different domains, poses significant challenges for search engines in indexing and ranking web pages. Google’s guidelines, although not exhaustive, offer valuable insights into managing these issues.
Internal duplicate content can be mitigated through strategies like canonicalisation, consistent internal linking, and minimising repetitive content. On the other hand, external duplicate content, often arising from scraped content, requires vigilance and sometimes legal action to protect your site’s integrity and ranking.
It’s imperative for webmasters and SEO professionals to create and maintain original, high-quality content to ensure the best possible search performance and to provide a valuable, unique experience for users. This approach not only aligns with Google’s guidelines but also sets a standard for web content that benefits the entire online community.
Recent Comments