How do you research competitors with AI without copying them?

A woman writing in a notebook at a desk with a large monitor, a laptop and a tablet all showing different screens
Four screens and a notebook. The notebook is the part that matters, because the moment you stop writing things down is the moment observation and conclusion start blending together.

Use a model to organise public evidence and to keep track of who said what. Do not use it to decide what any of it means, and never let it write the thing you publish. That division sounds fussy until you notice the failure it prevents, which is not plagiarism. It is ending up with a position that only makes sense as a reply to somebody else's.

Everybody worries about the first risk, which is reproducing a competitor's words. That one is real, easy to understand and easy to avoid. The second risk is quieter and does more damage.

What this comes down to
  • Copying prose is the small risk. Absorbing their framing is the large one.
  • Write what you observed and what you concluded in two separate columns.
  • A model is good at sorting evidence and bad at judging a market it cannot see.
  • The output should be a list of unanswered questions, not a summary of their pages.
  • If a sentence you publish only exists because you read their site, cut it.

The risk that actually costs you something

Read five competitors closely and something happens without your permission. You start using their category name. You accept the axis they compete on. You answer the objection they invented. By the time you write anything, you are arguing inside a frame somebody else built, and inside that frame the company that built it always looks like the default.

Nobody notices derivative positioning, because it does not look copied. It looks like the market.

This is worth saying plainly because it changes what good research looks like. If your competitor document is a tidy summary of what four companies say, you have built a mirror. It will tell you what everyone already believes, which is precisely the thing that cannot differentiate you.

The countermeasure is structural rather than moral. Decide what you are trying to find out before you read anything, keep a hard line between what you saw and what you inferred, and judge the output by the questions it raises rather than the summary it produces.

Start from a decision, not from a competitor

The weakest version of this work begins with a name. Somebody says "have a look at what Acme are doing", you look, and you produce a document about Acme. It reads well and changes nothing, because no decision was waiting on it.

Begin with a decision that is actually pending. Whether to lead with speed or accuracy on the homepage. Whether the free tier is a mistake. Which objection to answer first in the sales conversation. Then ask what evidence would move that decision either way.

That reframing does two useful things. It bounds the collection, so you stop at four pages per company rather than drifting into an archive. And it makes the research falsifiable, because you can say afterwards whether the decision moved.

Separate what you saw from what you concluded

This is the single most useful habit here, and it takes a table.

ObservationInterpretationConfidenceStill unknown
Pricing page leads with an annual figure, monthly shown smaller underneath They are pushing annual commitment, possibly for cash flow or churn Medium Whether monthly converts worse, or is simply newer
Homepage names no competitor and no category Either they define the category or they are avoiding comparison Low Which, and nothing on the page distinguishes the two
Three of four pages answer a security question unprompted They hear this objection constantly High Whether it costs them deals or just meetings

The first column can be checked by anyone. The second is yours, and it is where the value is, but it is also where mistakes calcify. Left unseparated, an inference gets repeated in a meeting, loses its hedge, and is treated as fact by the third retelling.

The fourth column is the one people skip and the one worth most. Writing down what you still do not know is what stops a document from sounding more certain than the evidence supports.

What a model is genuinely good at here

Sorting. Given the extracted text from a dozen pages, a model will group similar claims, find every sentence mentioning a topic, and lay scattered wording into a table far faster than you will. That is real work and it is dull, which is exactly the kind of task worth handing over.

It is not good at telling you what a market means. It has no idea who your buyers are, what your sales team hears, or which of these companies is actually winning. Asked anyway, it will produce a confident and plausible answer, because fluency and correctness are separate properties.

Three requests that stay on the right side of that line:

Grouping. "Here is text from four competitor homepages. List the distinct problems each claims to solve, quoting the sentence for each. Do not infer motives."

Absence. "Across these pages, which of the following questions is never answered? Quote any partial answers."

Change. "These two versions of the same page are three months apart. List what was added, removed or reworded, quoting both."

Every one asks for quotes. That is not politeness, it is the verification step: a claim carrying its source sentence can be checked in seconds, and an invented one fails a text search instead of quietly reaching your summary.

An overhead view of handwritten pages spread across a wooden desk beside a laptop and a graphics tablet
Spread out like this, it is obvious which sheet holds what. That is the entire benefit of the four column split, and it disappears the moment everything gets merged into one summary.

Keep a source log

One line per extract: the page, the date you took it, and what you pulled. It takes seconds and it earns its keep twice.

First, pages change. A claim collected in March and quoted in June may no longer be on the site, and discovering that in front of other people is unpleasant. Second, somebody will eventually ask how you know. A source log turns "I think they are moving upmarket" into "their pricing page dropped the starter tier on the 4th, here is the line", which is a different kind of conversation.

It also keeps you honest about coverage. Four companies is a comparison. One company plus three you glanced at is an anecdote with decoration, and the log makes that visible before you present it.

Look for gaps, not summaries

The useful artefact at the end of this is short and mostly negative space.

  • Questions a buyer clearly asks that nobody answers on the page.
  • Segments everyone describes in the same words, which usually means nobody is serving them specifically.
  • Claims made repeatedly with no evidence attached anywhere.
  • Limitations that are obviously true and that nobody states.
  • Steps in the workflow every company skips past, because they are awkward.

Each of those is a position available to you that is not a reply to anyone. That is the difference between research that produces a strategy and research that produces a slightly different version of the market's existing sentence.

The test before you publish anything

Take whatever came out of this work, a page, a brief, a paragraph of copy, and go through it asking one question of each sentence: would this exist if I had not read their site?

If the honest answer is no, one of three things is true. It is a fact you now know, in which case keep it and attribute it. It is their framing wearing your vocabulary, in which case cut it. Or it is genuinely your idea that their page happened to trigger, in which case keep it and stop worrying.

Most people find more of the middle category than they expect. Finding it is the entire point of doing the check.

A woman resting her chin on her hand while reading an open book by a bright window
The useful part of competitor research is this pause, not the collecting. What you are looking for is the question none of them answered.

The collection half of this

Getting clean text off a page without the navigation and cookie notices is the boring first mile of any of this. Our Website Content Extractor does that part and stops there. Free, no account, and it runs in your browser.

Try the Website Content Extractor

Where the actual line is

Worth stating plainly, because vagueness here makes people either reckless or paralysed.

Public pages are fair to read, quote briefly with attribution, and analyse. That is ordinary competitive research and every company does it. Anything behind a login, a paywall or an agreement is not public, and accepting terms in order to collect against the company that set them is a different situation with different consequences.

Facts are not ownable. The observation that a competitor charges annually and displays it prominently is a fact. Their sentences, images and layout are theirs, so quote sparingly, attribute, and never reuse wording in your own material.

Respect what a site asks of automated access, and do not send enough traffic to matter. If you find yourself building something that needs rate limiting, you have left research and started scraping, and the rules for that are different.

None of this is legal advice. If a project sits near a line rather than well inside it, that is a question for somebody qualified rather than for a blog post.

Frequently asked questions

Is it legal to analyse a competitor's website?

Reading publicly available pages and analysing what they say is ordinary business practice. What is restricted is reproducing their copyrighted material, accessing anything behind a login or agreement, and automated collection that ignores what a site asks of it. This is general information rather than legal advice.

Can I use AI to rewrite a competitor's page for my own site?

No, and the reason is more practical than legal. A rewritten version of their page is still their argument, so it positions you as an alternative inside a frame they built. Even where it avoids a copyright problem, it guarantees a strategic one.

How many competitors should I look at?

Three or four, properly, rather than ten superficially. The value comes from reading the same four pages across several companies closely enough to notice what differs, and that attention does not survive being spread across a longer list.

How often should this be repeated?

Quarterly is enough for most markets, and the second pass is where it becomes genuinely useful, because you are reading the change rather than the snapshot. Keep the page set identical so the comparison holds.

What should I never put in a competitor document?

Anything you cannot source, anything obtained from a person under an obligation not to share it, and inference presented as observation. The first two are risks to the company, the third is a risk to every decision made from the document afterwards.

Does a model make this faster or just easier?

Faster at the sorting, which is most of the tedium. It does not make the judgement faster, because the judgement was never the bottleneck. Expect the collection and organisation to shrink and the thinking to take exactly as long as it did before.

Where to start this week

Pick one decision you are actually stuck on. Collect the same four pages from three competitors, using the collection method covered here. Then fill in the four column table for whatever you notice, and be strict about which column each line belongs in.

When you finish, look only at the last column. The list of things you still do not know is your next piece of research, and it is a much better brief than the one you started with.


How this was put together: this is the method we use for our own competitor reading, including the four column table and the source log. It is not research, there is no dataset behind it, and we have quoted no figures on how long this takes or how much it improves a decision, because both depend entirely on the decision. The section on where the legal line sits is general information and not legal advice.