Product Page SEO at Scale: 10,000+ SKUs
BlogE-commerce SEO
+227%Monthly net sales · SMK

Product Page SEO at Scale: 10,000+ SKUs

11-minute read
August 27, 2026
M
Mubashar SharifVerified SEO Expert

Founder & SEO Expert · August 27, 2026

Follow on LinkedIn

Near-identical boilerplate is the most common reason large catalogues fail to index. Here is the brand-by-brand rewriting method we used on one 35,000-product store.

The duplicate content problem

Most e-commerce product pages are nearly identical — the same boilerplate description, the same spec table, the same FAQ block, with a model number swapped. Google does not owe you an index slot for each one, and on large catalogues it stops granting them.

You can see this in Search Console before you see it in revenue. Open the Pages report and look for "Crawled — currently not indexed". On a big catalogue that bucket often holds more URLs than the indexed one. It means Googlebot fetched the page, evaluated it, and decided storing it added nothing to what it already had.

That is a content judgement, not a technical fault. No sitemap, submission script or API changes it — a point worth making because the workaround industry around indexing is large. We wrote about that in The Google Indexing API Is Not a Shortcut.

Quick diagnostic: take twenty product URLs from the same category and paste their descriptions into a diff tool. If the only differences are the product name and a couple of numbers, you have found your indexing problem.

Brand-by-brand, not page-by-page

The instinct on a 35,000-SKU catalogue is to start at SKU 1 and work down. That fails on arithmetic — at ten minutes a page it is over a year of work — and it produces the wrong output anyway, because a writer working alphabetically has no context for what makes a product different from its siblings.

Work one brand at a time instead. Everything you learn researching a brand — its range, its materials, who buys it, how it differs from the brand next to it on the shelf — applies to every product you write under it. The first page in a brand takes an hour. The thirtieth takes ten minutes and is better, because by then you know what the buyer is actually choosing between.

It also gives you a natural release unit. Finish a brand, push it, watch indexing for that segment, and you learn whether the approach is working before you have committed the whole catalogue to it.

Sequence brands by commercial value, not alphabetically: revenue first, then search volume, then how badly the current pages are indexed.

The content template

A template that produces genuinely different pages constrains structure, not sentences. Ours has five parts:

  1. What it is, in two sentences. Written for someone who has landed from search and does not yet know if they are on the right page.
  2. Who it suits, and who it does not. The part almost nobody writes, and the part that is genuinely unique per product. Naming who should buy something else builds more trust than another paragraph of praise.
  3. How it differs from the nearest alternative — usually the next model up or down in the same range. This is impossible to write generically, which is exactly why it works.
  4. Specifications, as structured data and a table. Machine-readable, not prose.
  5. The two or three questions buyers actually ask, taken from support tickets and reviews rather than from a keyword tool.

Note what the template does not do: it does not specify sentence patterns or an opening formula. A template that dictates phrasing recreates the boilerplate you are trying to escape, just with fresher wording.

Length is not the target. Three hundred words that answer the buyer's actual question beat a thousand words of padding, and Google has been explicit that word count is not a ranking factor.

While you’re thinking about this

Is your own site making this mistake?

Send me your URL and I’ll check it myself against what you just read — competitor gaps, content gaps, and whether Google’s AI names you or them. Free, written by me, reply within 24 hours.

Where AI fits, and where it gets you hurt

Google's position is about value, not method: content is not spam because a model helped write it, and it is not acceptable because a human typed it. What the spam policies describe is scaled content abuse — generating pages at volume primarily to manipulate rankings rather than to help anyone.

The practical line we work to:

  • Reasonable: drafting from a real spec sheet you supply, restructuring existing copy, generating first passes a subject-matter reviewer then corrects, writing the spec table from structured data.
  • Not reasonable: generating thousands of descriptions from product names alone and publishing them unreviewed. That is the exact pattern that produced the boilerplate problem in the first place — it just produces it faster and in more fluent prose.

The test we apply before publishing: does this page contain at least one thing that could only have been written by someone who has handled the product or talked to a buyer? If not, it is a rewrite of the same page you already have.

Budget for review. On the projects where this worked, review time was roughly a third of total effort — and it is the third that produces the "who it does not suit" and "how it differs" sections that carry the whole approach.

What the results looked like

Two projects, reported honestly, including the one that went backwards before it went forwards.

SMK Store — over 35,000 product pages, barely indexed, with thin near-identical descriptions tripping duplicate-content filters and failing Core Web Vitals. We rewrote brand by brand, optimised crawl budget, implemented product schema and fixed Core Web Vitals, resubmitting in batches. Monthly net sales went from $5,832 in April 2026 to $19,100 in June 2026 (+227%) with no additional ad spend, as shown on the store's WooCommerce dashboard. Full detail in the SMK Store case study.

Michigan Outdoor Sports — brand pages never properly submitted, thin content causing mass non-indexing, crawl budget wasted. Organic clicks peaked at +476% in March 2026, then lost ground to a gradual de-indexing before we rebuilt to 11,549 indexed pages, up from roughly 3,000, and +83% US organic clicks by July. Written up in the Michigan Outdoor Sports case study.

That second trajectory is the more useful one to plan around. Indexing gains are held, not won — if the underlying content stays thin in places, pages drop back out.

Do this week

  1. Export the Pages report from Search Console and count how many URLs sit in "Crawled — currently not indexed".
  2. Diff twenty descriptions from one category. Confirm the cause before committing to a rewrite.
  3. Pick your highest-revenue brand and rewrite its full range against the five-part template.
  4. Push that brand alone and watch its indexing for three weeks. Do not start brand two until you know brand one worked.
#product pages#e-commerce#content#indexing
M

Mubashar Sharif

LinkedIn Profile
Founder & SEO Expert

Mubashar is an SEO analyst with 5+ years specializing in large-scale e-commerce SEO.

Share: LinkedIn

Send me your URL. I’ll tell you what’s wrong with it.

Two fields, a reply within 24 hours, written by me — the person you just read.

Or book a 30-min call