Competitive analysis falls apart at the collection step. Copying a competitor page by hand brings the navigation, the cookie notice and the related-articles rail along with the message you actually wanted, and by the twelfth page nobody is reading carefully any more.
The thinking part of this work is not usually the problem. Somebody who knows the market can read three competitor pages and tell you something useful about all three. What breaks is everything before that: opening the pages, getting the words out in a state you can compare, and still having the patience to look properly once you do.
- Decide what you are comparing before you collect anything, or you will collect everything.
- Four pages per competitor is usually enough. Twenty is a different project.
- Strip navigation, cookie notices and related-article rails before reading, not after.
- Keep the wording verbatim. Paraphrasing at the collection step destroys the comparison.
- Public pages only, and record where each extract came from.
Decide what you are comparing first
The most common way this goes wrong is trying to cover everything. A brief that tracks every page a competitor owns takes longer to produce and says less than one that tracks only the pages a buyer actually reads.
Pick a question before you open a browser tab. Something answerable, like how three competitors describe the same problem, or which objections each one handles on its pricing page. A question makes the collection finite, and finite is what makes it repeatable next quarter.
For most comparisons, four pages per competitor covers it: the homepage, the main product or feature page, the pricing page, and whichever page they send new visitors to first. If a fifth page keeps turning out to matter, add it deliberately rather than drifting into collecting the whole site.
What to pull from each page
You are collecting language, not design. The things worth having side by side are the sentences a company chose after arguing about them internally.
- The headline and subheading, which is the shortest version of their pitch.
- How they describe the problem, which tells you who they think has it.
- What they claim as the benefit, and whether they attach any evidence to it.
- Which objections they answer without being asked.
- How pricing is framed, separately from what it costs.
- What proof they lean on, whether that is named customers, numbers or nothing at all.
Keep the wording exactly as written. The temptation is to summarise while collecting, which feels efficient and quietly destroys the thing you came for, because the difference between two competitors is often one adjective.
Getting the text out cleanly
Selecting a page and copying it brings everything: the menu, the cookie banner, the footer sitemap, the newsletter box, the three related articles at the bottom. On one page that is mildly annoying. Across twelve it turns into an hour of deleting things, and the deleting is where attention runs out.
The tidying is not a small step before the work. On a twelve page comparison it is most of the work, which is why the reading gets rushed.
Whatever method you use, the goal is the same: the words a visitor came to read, without the furniture around them. Our Website Content Extractor does that part, and if a page blocks selection entirely there are several ways around it that need no extension.
One practical note: pages change. Put the date next to every extract. A comparison built from pages collected across three weeks is comparing a moving target, and you will not remember which was which.
Reading what you collected
With the text side by side, a few things tend to surface quickly.
Where everyone says the same thing. If four companies all lead with the same benefit, that phrase has stopped differentiating anyone. It is table stakes, and repeating it puts you fourth in a queue.
Where nobody says anything. Gaps are more useful than overlaps. A question every buyer asks that no competitor answers on the page is either an opportunity or a sign the answer is awkward. Both are worth knowing.
Who they are talking to. Vocabulary gives this away faster than any stated audience. A page written for a procurement committee reads differently from one written for the person who will use the thing on Monday.
What they defend. An objection answered unprompted is usually an objection they hear constantly. That is free research into their weak point, and often into yours.
Using a model to help, without letting it invent
A language model is useful here for the mechanical half: grouping similar claims, spotting repetition across a dozen pages, putting scattered wording into a table. It is not useful for telling you what any of it means, because it does not know your market.
Give it the extracted text and ask narrow questions. Three that work:
Gemini: "Here is text from four competitor homepages. List the distinct problems each one claims to solve, quoting the sentence for each. Do not summarise or infer."
ChatGPT: "Group these value propositions by the benefit they promise. Show which are near duplicates of each other and quote the wording."
Claude: "Which objections does each page answer without being asked? Quote the sentence that answers each one."
Notice that all three ask for quotes. That is the whole trick. A claim carrying the sentence it came from can be checked in a second, and a claim that was invented fails a text search instead of quietly making it into your brief. The same reasoning applies to any workflow with a model in it.
Staying on the right side of it
This is straightforward if you keep to a few limits.
Use publicly available pages only. Anything behind a login, a paywall or a signup form is not public, and agreeing to terms in order to collect against them is a different situation entirely. Respect what a site asks of automated access, and do not hammer a server.
Read competitor material to understand positioning, not to reuse it. Copying their sentences into your own pages is both a copyright problem and a strategic one, because it makes you the second company saying something.
Record where each extract came from and when. A source log turns an opinion into something a colleague can check, and it is the difference between analysis and assertion.
The collection step, handled
Our Website Content Extractor pulls the readable text off a page without the navigation, cookie notices and related-article rails. Free, no account, and the guides here cover what to do with the text once you have it.
Try the Website Content ExtractorKnowing whether it worked
The test is not how much you collected. It is whether the brief changed a decision. If you can point at a sentence you rewrote, a page you added, or an objection you started answering because of what you read, the exercise paid for itself. If nothing changed, either the question was too broad or the answer was that you are already well positioned, and both are worth writing down.
Do it again in a quarter with the same four pages per competitor. The second pass is where this becomes genuinely useful, because you are no longer looking at what they say, you are looking at what they changed.
How this was put together: this is the method we use for our own competitor reading, described plainly. It is not research, there is no dataset behind it, and we have quoted no figures about how long manual analysis takes or how much time extraction saves, because those numbers depend entirely on how many pages you collect and how tidy the sites are.