Every leadership team walks into a quarter carrying beliefs about its competitors that nobody has checked.
That your price sits at the bottom of the market, that the objection your reps hear most is the one the whole category hears, and that the feature beating you is the one on their homepage.
At least one of those is usually wrong. Finding out costs money and a fortnight, because somebody has to sit inside the competitor’s sales process as a real buyer, and that somebody is us.
So the real question is whether the information changes something you would otherwise get wrong. Interesting is cheap, and being wrong about a rival’s price for a quarter is not.
Measure the Artifact That Changed, Not a Modeled Lift
The temptation is to build a return calculation. Take a win rate, add a plausible lift, multiply by deal size, and produce a number with a decimal point in it.
That number is fiction, and everyone in the room knows it.

If neither Forrester nor Crayon can attribute revenue to a competitive intelligence program, the honest unit of return is something smaller: a named internal artifact that changed, with a person attached and a date it changed by.
A pricing rule, a battlecard field, the order of your demo, or the question that opens your next executive review.
You can walk into a room six weeks later and verify whether that happened. A modeled win-rate lift is not verifiable after the fact, which does not stop it arriving with two decimal places and a meeting.
Which Decisions Are Worth Testing This Way?
Three come up more than the rest. They are illustrations rather than a ranking, because no research establishes which of them pays best.
A pricing floor is only comparable once the conditions match, which is the whole of what a competitor pricing review captures: the metric, the included volume, the minimum commitment, the term, the implementation charge, the discount condition, and how long the quote stays valid.
A headline number without those attached is a rumor with a currency symbol on it.
A battlecard earns a rewrite when the line a competitor’s rep uses against you turns out to differ from the line your team assumed. Reps do not improvise objection handling. What comes back is one line, repeated, which is why a battlecard rewritten from a recording beats one rewritten from a guess.
The third is a product question. A capability described as important and never shown live is a different fact from one demonstrated on screen, and that difference should change what gets asked at your next executive review.
The Limits of a Trained Observer
That distinction is where this method is genuinely strong. A Reutlingen study of three retail service chains found mystery-shopping assessments were not related to sales performance, apart from a weak cross-selling effect, and were not effective at predicting customer satisfaction.
The same study was clear about what trained observers do well: they evaluate very specific aspects of an interaction that ordinary customers never recall. That is the part worth buying.
| The decision | Evidence to capture | What it does not prove |
|---|---|---|
| Your pricing floor | Metric, volume, minimum commitment, term, implementation charge, discount condition, quote validity | That their price predicts their revenue |
| A battlecard response | The exact objection line, the proof offered, the reframing question | That it is the line every rep uses |
| Executive review priorities | What was demonstrated live versus described, deferred, or promised | That an unshown feature does not exist |
One evaluation gets read by several people for different purposes, which is the other thing the research on this method is clear about. One project, several decisions.
Is 499 Euros a Lot of Money?
What you are buying is one Notion board per project: every demo recording with a timestamped transcript, the email trail and collateral each vendor sent, a page per competitor written by the person who sat the call, the comparison matrices, and an executive read on top.
It is a fixed published fee per competitor, which makes the comparison unusually simple to run.
One competitor is 499 euros, five are 2,495, and ten are 4,990. Extra meetings, follow-up, and the analysis all sit inside that fee rather than beside it, and a single competitor takes one to two weeks against four to eight for a full landscape.

We have run these for some of the best-known names in B2B SaaS and for companies nobody has heard of yet. The fee does not change because the competitor is famous.
Set the number against the decision, in that order:

A pricing rule that governs a large share of your outbound quotes is not a minor assumption to be carrying untested. Keep the fee beside that decision rather than inside an invented return model.
And if the decision at stake is genuinely smaller than the fee, you have your answer, and it is a perfectly good one.

A Finding With No Owner Changes Nothing
The failure mode here is not a bad piece of research. It is a good one that nobody acts on.
.png)

Any one of those can move a comparison further than the headline figure does. Which is why a number without its conditions attached is not a price yet, and why a pricing sheet built from published tiers is usually comparing the wrong column.
So the test at the end is one question, asked six weeks after delivery: did an artifact change, and can you name it?
If a pricing rule moved, a battlecard field got rewritten, or a demo path was reordered, the project earned its keep. If nothing moved, it did not, whatever the quality of the analysis.
The recordings are the evidence. The read on what they add up to is written by the people who sat the calls, which is the shift Forrester describes from delivering information to providing implications.
Most teams already know which belief they would least like to be wrong about. Bring that one to a call and we will tell you what it would take to settle it.
.png)