What Is an Evidence Grade in Competitive Research?

A grading scheme for competitive claims, from an observed artifact down to inference, with the line between each grade drawn where a real claim gets sorted.

The slide says the competitor drops to thirty percent off at quarter end. A board member asks where that came from. The head of product says it came from the battlecard.

That is not an answer, and everyone in the room can hear it.

The claim might be perfectly true. The problem is that nobody can say how it got there. So nobody can say whether it is still true, and a claim you cannot trace is a claim you cannot spend on.

An evidence grade fixes that with one label. It ranks a claim by how you got it, from a document you can put on the table at the top down to something you worked out yourself at the bottom. The label rides along with the claim.

We sit competitor demos for the companies that hire us, which is a job mystery shopping for B2B SaaS describes end to end. From that seat you watch claims get their grade in real time.

The rep says a number. The screen shows a different one. Both are true, and they are not the same fact.

Ask What You Would Put on the Table

Here is the whole method in one question. If somebody challenged this claim right now, what would you put on the table?

The answer is not an opinion. It is a physical thing you either have or you do not.

Say the claim is that their enterprise floor is $140,000 a year. You might have a screenshot of their pricing page. You might have a rep saying the number on a recorded call.

Or you might have two people on your team who think they heard it somewhere. Or their smallest listed tier, times a plausible seat count, landing about there.

Same sentence, four different kinds of evidence. The grade belongs to what you are holding, not to the claim.

A true claim can sit at the bottom. A claim nobody disputes can sit at the top and still not be worth a slide.

The Ladder, From the Artifact Down to the Guess

Those four answers have names, and one more sits in the middle.

A, artifact. The thing itself: their pricing page, their contract, their documentation, a recording of their demo, the onboarding email with its timestamp. You produce it and the argument ends.

B, testimony. Someone who would know said it, on the record, in their own words: a rep on a call, a named customer. You have the words, not the thing the words describe.

C, corroboration. Two sources that got it separately say the same thing, and you have both.

D, report. One source said it. Nobody checked.

E, inference. Nobody said it. You worked it out.

Grades only help if the line between two of them is a question with one answer. Otherwise sorting a claim turns into an argument about taste.

BoundaryAsk thisHow it falls
A or BDoes it show the fact, or somebody saying the fact?A recording of a rep quoting a price is an artifact of the quoting and testimony about the price
B or CDo you have a named person’s own words?A paraphrase is a report, and so is somebody repeating what they heard
C or DDid the two sources get it separately?Two people who read the same review are one source
D or EDid anybody say it at all?If the sentence exists because you joined two facts, it is a read, however good

The line your board cares about sits between B and C. Above it, you can point at the thing: this document, this call, this person, this date. Below it, you are repeating what people say about calls you never sat in.

That gives you the rule for the room. Above the line, state the claim; below it, name its grade. One extra clause, and the room trusts everything else you say that afternoon.

What Twelve Comparison Pages Put on the Table

So where do real competitive claims fall? The easiest place to check is the page your own team probably lifted half its battlecard from.

On August 27, 2026 we opened twelve vendor comparison pages, one per publisher, across twelve categories of B2B software. Then we graded every claim each page made about its named rival, using one test: what does this page hand a reader who wants to check it? That came to 193 claims.

152 of the 193 came with nothing attached. No document, no named speaker, no outside source.

Most of them are probably accurate. A vendor usually does know its rival’s pricing tiers. But nothing on the page tells a reader which claims were checked, so all of them land on the bottom rung together.

Horizontal bar chart of 193 competitive claims sorted by evidence grade: Artifact 12, Testimony 8, Corroboration 1, Report 20, Inference 152. The Inference bar is far longer than the other four put together.
Every claim each page made about its named rival, graded by what the page put on the table beside it.

Twenty claims reached the top two grades. One claim in the whole set reached corroboration.

Where the good evidence sits is more interesting than the count. All twelve artifact-grade claims came from two pages, and eleven from one: a scheduling vendor that linked its competitor’s own pricing page next to every price it quoted.

Ten pages produced nothing at that grade. Two produced nothing above the bottom rung, and one of those is named after its competitor and says almost nothing about it.

The page I keep thinking about is the one that tried. It runs a review platform’s name and a data-as-of month across the top of a ten-row comparison table. Then no row says which figure supports it.

A source for the page is not a source for the claim. Row four is no easier to check than it would be with no banner at all.

None of this is dishonesty. It is just that nobody was grading.

What a Rep Says and What the Screen Shows

The boundary that costs the most money is the first one, and it is the one you cannot cross from a desk.

A recording of a rep saying their compression runs at six to one is an artifact. It is an artifact of the saying. It proves that on a given date, a named person at that company told a prospect that number.

What their software does when you point it at a real file is a separate claim, and it needs its own evidence.

Teams get this wrong in both directions. Treat the rep’s number as measured and you end up cutting your price against a figure nobody verified. Throw it out because it is only testimony and you lose the best signal you have about what that company thinks it can charge.

The page below is one dimension of our public example project, a fictional engagement we publish so the deliverable is not an abstraction. Every engagement gets pages like this, built from its own calls.

The first column is the ratio each vendor puts in its marketing. The second is what that vendor produced live in its demo. The third is how they produced it.

Comparison dimension from the Mystery Demo example board: five vendors with their headline compression ratio, the verified range measured in their own demo, the method each used, and the gap to the client's product.
Headline claim, verified range, and the method behind each: one dimension of the public example board.
Chief Mystery Officer
Mystery Demo
The number a rep says out loud and the number their product hits are two different facts, and the gap between them is where most deals get decided. We have sat calls where the headline figure in the deck was roughly double what the live run produced, and the rep was not lying, they were quoting their own marketing site. A recording tells you what a company claims. A recording plus the screen tells you what it can show, and only the second one survives a procurement review.

Watching those two grades come apart inside one meeting is most of why the meeting is worth attending. What a rep does under that pressure is its own subject, and it is where our competitor sales tactics research lives.

The point here is narrower. A recording proves what was said, not what is true, and a battlecard that does not mark the difference is one question away from falling over.

Two Sources That Read the Same Review Are One Source

Demos hand you the top two grades. The rest of the ladder you build at a desk, and the desk has a trapdoor.

Corroboration is not about how many sources you have. It is about where they got it.

Sit enough demos and you see this happen live. Two vendors quote you the same market statistic in the same week, and it sounds like the industry agreeing with itself. Both of them read it in the same analyst summary, so you have one source and two voices.

Your own team does it too, whenever two reps repeat a review neither of them mentions reading.

So the question to ask about a second source is simple. Where did you get that? If both answers point back to the same place, you have one source, and calling it corroboration just makes a rumor sound better dressed.

In our twelve pages, exactly one claim cleared this bar. Three customers had written the same thing about the same product, separately, and the page linked all three reviews so a reader could confirm they were three people and not one story repeated. That is the whole price of the grade, and eleven pages out of twelve never paid it.

A Claim Drops a Grade at Every Handoff

Paying it once is the easy part. Keeping the grade attached is harder, because evidence does not travel. The sentence does.

The recording stays in the folder where you collected it. The claim moves on: into the analysis, onto a battlecard, into a rep’s mouth on a call. Nothing along the way is stamped.

A grade-A observation and a grade-E guess reach the same meeting looking identical, and whoever has to defend one cannot tell which is which.

That is how a company with genuinely good research ends up with a battlecard describing a competitor that no longer exists, which is one of the biggest competitor analysis mistakes we see.

A date is a second axis, not part of the grade. That scheduling vendor with the linked pricing page is bylined October 1, 2024, and those claims are still grade A.

A reader in 2026 has no way to know whether the numbers held, because nothing on the page went back to look.

Grade tells you how you got a claim. Date tells you whether it survived. Run them together and you make both of the usual mistakes: trusting an old document because it is a document, and binning a sharp read because it has no document behind it.

Below is the top of one competitor’s page from that same example project. The company facts, the meeting that produced them with its date and length, and the recording sitting underneath. Nothing on that page is more than a click from where it came from.

Top of a per-vendor page on the Mystery Demo example board: company facts, an intro meeting callout with its date and length, and the recording embedded below it.
One competitor, one page: the facts, the meeting that produced them, and the recording sitting under both.

The fix is not more discipline further down the chain. It is keeping the artifact one click from the claim, so a reader can check the grade instead of taking somebody’s word for it.

The Grade Decides What You Do Next

Once a claim carries a grade, the grade tells you what you are allowed to do with it.

Grade A. State it flat, date it, and keep the artifact one click away.
Grade B. State it as what somebody said, and name who said it. The attribution is the strength, not a hedge.
Grade C. State it, and say why the two sources are separate, because that is the part somebody will push on.
Grade D. Use it to decide where to look next. Never use it to decide anything expensive.
Grade E. Show the reasoning next to the conclusion, so a reader can argue with the step instead of the verdict.

Inference sits at the bottom and gets treated as something to apologize for. It should not be. Some of the sharpest competitive reads you will ever produce are grade E, because nobody at the other company is going to explain what their roadmap means.

A good inference is not a confident one. It is one that shows its inputs.

Say what you saw. Say what you concluded. Never make a reader guess which is which.

That is what the analysis layer on each competitor’s page is doing in the header: how many meetings it draws on, how many minutes, which collateral was cross-checked. What follows is a read, and it says so by saying what it is made of.

Expert analysis section of a per-vendor page on the Mystery Demo example board, headed by a callout naming its inputs, two meetings and ninety-five minutes of them, above six bolded insights.
The inference layer with its inputs named, so a reader can weigh the conclusion against what produced it.

Which leaves one asymmetry worth pricing. Grades C, D, and E you can build yourself, with a browser and enough patience.

Grades A and B live inside your competitor’s funnel: in the pricing slide a rep shares, in the follow-up sequence, in what somebody says when a prospect pushes back on price. No amount of desk research gets you those. That gap is the work.

What we hand a partner is a board where every claim about a competitor sits next to the recording, transcript, or document it came from, so its grade is visible instead of remembered. If there is a line on your battlecard you would rather not defend in front of your own board, tell us which competitor it belongs to and we will go get the artifact.

Ready
To Connect?

Let's Connect