In the competitor demos we sit, the rep almost never tells you why they are better than the other name on the shortlist. Not because they are hiding it. Because nobody has asked yet, and the script does not get there on its own.
Ask directly and you learn something worse. A good number of them cannot answer.
Mystery Demo buys your competitors on your behalf. We go through their funnel as a genuine buyer, put the same questions to every vendor in the set, and hand back the recordings, the transcripts, the follow-up emails and our read on what all of it meant.
So we spend our weeks on the receiving end of other people’s differentiation. This is the report from that seat, and it is an opinion, held hard.
A difference that only surfaces when a buyer drags it out of you has never been tested under the one condition that matters, which is a buyer holding a live alternative, asking.
It has been written, approved, and put on a slide. Until somebody with a real choice has heard it and stayed, it is a claim.
Yours might be a good one. You do not know yet, and neither does anyone else in your building.
The Buyer Is Already Trying, and It Is Not Working
The comfortable version of this problem is that buyers do not read. It is a nice theory. The numbers do not support it.
Wynter surveyed a hundred B2B SaaS marketing leaders who had recently evaluated competing vendors, 73 percent of them VP level or above, in June 2026. When vendors seemed similar, 92 percent of them put real effort into finding distinctions and 56 percent dug deep. Half of the deep diggers still came away feeling everyone was the same.
Effort and clarity were uncorrelated. That is the entire indictment.
The buyer went looking. The buyer did not find it. Whatever you think you have, it was not on the surfaces they searched, which means the difference is not hidden, it is missing.
One of the CMOs in that survey said the quiet part out loud. “Their sellers were able to do a good job of explaining the differences that the website by itself was not able to do.”
That is the good outcome, by the way. That is a rep rescuing a claim that had already failed everywhere else it appeared.
Now put a clock on it. In 6sense’s 2025 study of nearly 4,000 B2B buyers, four out of five deals were won by the vendor the buyer already favored before speaking to a single seller, and 95 percent of the time the eventual winner was on the day-one shortlist. That producer sells software premised on exactly that finding, so weigh it accordingly, and it has held across its own earlier waves.
Read those two studies together and the seat you are sitting in gets uncomfortable. The buyer arrives having already leaned somewhere, and the call is close to the last room where that lean can move.
Most of the calls we sit use that room to confirm it.
You Cannot Differentiate From a Competitor You Have Never Watched Sell
That is not laziness, and it is not bad sales training either. Differentiation is a comparison, and most companies have only ever written one side of it.
You know your own product exhaustively. You know your rivals the way you know a restaurant from its menu.
Three things decide whether your claim survives contact, and public material gives you none of them with any confidence.
What the product is like when a person drives it live. A feature page tells you a capability exists. A demo tells you where it lives, how many clicks it takes, and whether the rep can do the thing or has to describe it.
The moment worth watching is the one where somebody says the words “we can get you a screenshot of that after the call,” because a capability that cannot be reached in the demo environment is a capability with an asterisk.
How they position when a real buyer is in the room. The homepage is written for everybody, so it says almost nothing. In a call, your name comes up, and the rep positions against you personally, in a sentence nobody outside that room has ever heard.
Their discovery questions are the other half of it: what a rep asks first tells you who they think they are for, and it is often not who the website says.
How they sell. This is the widest of the three and the one nobody researches, because it only exists in behavior.
None of that is on a comparison page, and none of it is guessable. Your differentiation sentence is a claim about a competitor, and you have been writing it from outside a building you have never entered.
The pattern of what reps hold back has its own page: the things competitors almost never do in a demo is the longer list. What follows here is the one question that does the most work.
.png)

What We Do About It, and What Comes Back
Somebody has to go and sit in the room, and that is the job. We are a team of operators who go and buy your competitors on your behalf, and nothing about the evaluation is staged except the assignment. If you want the category explained from the beginning rather than the method, mystery shopping for B2B SaaS covers what the work is and where it came from.
The whole method is one idea: hold the buyer constant, so that every difference you find afterwards belongs to the vendor.
It starts with the set, and the set is a decision, not a list. Which vendors, how many meetings each one takes, and why these and not the four others somebody suggested in the kickoff.
The screenshot below is from our public example project. It is an invented company evaluating invented vendors, built so we can show the work without showing anybody’s real project, and the names, numbers, and grades inside it are fiction. What is real is the shape.

Then one buying scenario, held constant across every call. A company profile, a use case, the constraints, the incumbent we are replacing, the objection we will raise. Same story, same questions, same order, every vendor.
That is the whole trick, and it is why a project like this beats a folder of notes from six sales calls that six different people happened to take. If the buyer changes between the calls, the differences you find afterwards are differences in the buyer.
Then we sit them. We take the follow-ups, we open the sequences, we ask for the reference, we let the security questionnaire run its course, and we record where the law allows it. Run across a whole field at once it becomes a competitive landscape review, which is a study rather than a shopping trip.
What comes back is a board. Not a deck, and not a slide with three bullets from each call.
There is a page per vendor with what was shown, what was said, and what was dodged. There is a comparison layer that puts them side by side on the dimensions the partner cares about. And there is a summary that commits to a verdict per competitor: here is what this one wins on, and here is where it loses.

That wins-on and loses-on line is the thing this whole subject is about. It is a differentiation claim written from the other side of the table, about a competitor, by people who sat in front of them for an hour.
The last layer is the boring one that makes the rest usable. Every claim on every page resolves to a named recording, transcript, or document, so nobody in your building has to take our word for anything.

Which is the whole point of the exercise. A verdict about a competitor is worth exactly as much as the artifact sitting under it, and most of the verdicts currently circulating inside your company have nothing under them at all.
499 euros a competitor, flat, everything included: the scenario, the booking, every meeting, the recordings, the transcripts and the analysis. One competitor takes a week or two; a whole field of ten to fifteen takes four to eight weeks, because the demos run in parallel.
Put a Rival’s Name in Your Sentence and See If It Still Reads True
You do not need any of that to start. There is one test you can run this afternoon without leaving the building.
Take the sentence your team agreed on. Replace your company’s name with your biggest competitor’s. Read it again.
That is the whole test. It costs a minute, and it settles most of the internal arguments that differentiation causes.
If the sentence still reads true, you have not written a differentiator. You have written a description of the category, and you have been paying to broadcast it.
The logic is not complicated. A difference is only a difference if it is false about the alternative, so a sentence that stays true when you swap the subject was never carrying a comparison in the first place.
Somebody has already run the buyer’s half of this experiment. Wynter took five real, unedited value propositions from the biggest names in CRM, stripped the company names off, and asked those hundred marketing leaders to match copy to company.

They scored 1.86 out of 5. Random guessing scores 1.0. The CMOs did worse than the room at 1.70, which is a hard number to look at if writing this copy is on your job description.
There is a worse one underneath it. Of those buyers, 41 percent said a genuinely different product capability was the single biggest reason they picked their vendor, and that group scored 1.63, below the average.
The people most certain that the product decided it were the worst in the room at recognizing a product from its own words.
So we ran the vendor’s half. We picked a B2B SaaS category we have never worked in, construction project management software, ran the search a buyer would run, and took every construction product that appeared on at least two of the five shortlists that came back. That is twenty-one products.
In one day in August 2026 we pulled each vendor’s own homepage and took its lead differentiation claim under a single rule: the hero headline, plus the subheadline where the subheadline completes the claim. Then we scored each claim three separate times, each pass blind to the others, and took the middle answer. It is one reader’s judgment run three ways, not a panel.
One question each time. Which of the other twenty products in the set could truthfully put this exact sentence on their own homepage? The figure for each claim is the middle answer of the three.

Thirteen of the twenty-one did not survive. Ten of them are true of at least five other products in the set. One is true of fifteen of the twenty, which means that vendor is not describing itself.
It is describing the aisle.
Six claims were built purely from category words, adjectives, and social proof, and five of the six failed badly. All-in-one, easy, affordable, powerful, leading, the new standard, results that matter, save time and money: swap the name and nothing moves, because nothing in the sentence was ever attached to the company saying it.
The obvious fix is to be more specific. The data will not cooperate.
Fifteen of the claims did name something concrete, a trade or a project type or a piece of work the software does, and only seven of those survived clean. The other eight ran up to eleven rivals apiece.
The concrete thing they named was job costing, or estimating, or payroll, or scheduling. In that category everybody does job costing.
Naming a real thing is not the same as naming your thing.
Six of the eight survivors named a buyer nobody else was chasing, a job nobody else was doing, or a mechanism nobody else had. One is built for demolition, abatement, and remediation contractors, which is three words that eliminate twenty rivals in a single line. Two more pick a project type and stop there: custom homes, and capital projects.
One is for construction workforce planning specifically rather than construction management generally. One is for roofing. One says it tracks crews by GPS, which sounds small until you notice it is the only sentence in the whole set that names a mechanism.
The last two survived on arithmetic. Nobody else could truthfully cite the same jobsite count, or the same customer count paired with the same decades of history, so those sentences are unrepeatable and they tell a buyer nothing about whether to buy. Unique and useful came apart right there in the data.

Three Gates a Difference Has to Get Through
The swap test tells you a claim is empty. It does not tell you which of the survivors to build a strategy on. Three gates do that, and a candidate has to pass all three.
Relevance. Does the buyer care. Not whether it is impressive, and not whether it was hard to build.
Name the decision it changes, out loud, in one sentence, or it is trivia with an engineering budget behind it.
Evidence. Does your proof match the kind of claim you are making. This is the gate teams fail without noticing, because every claim gets treated as though it needs the same kind of support.
The proof also has to be checkable quickly, because a buyer who has to take it on trust for two weeks does not experience it as a difference. They experience it as a thing you said. Separating what a competitor showed from what a competitor merely asserted is the whole job of an evidence grade, and it is the same discipline pointed at your own claims.
Resistance. Can the buyer get the same result somewhere else, with proof as good as yours. This is the gate almost nobody runs, because running it honestly requires knowing what the alternative shows a buyer, in a room you were not in.
That is the gate Mystery Demo exists for. We take your candidate difference into the rival’s own demo, ask them to show the thing your claim says they cannot do, and bring back whether they did it, dodged it, or offered to follow up next week.
Fail one gate and the claim comes out of the strategy. It does not go back to marketing for better wording, which is what happens instead, roughly quarterly, forever.
Differentiation Does Not Live Where You Are Looking For It
The gates tell you whether a candidate holds. They do not tell you where candidates come from, and most companies look in exactly one place.
Ask a leadership team where the company is different and the answer will be about the product, usually inside ten seconds. It is the surface that gets the roadmap, the budget, and the argument.
It is also the most copyable thing you own, and a buyer holding two feature grids reads them as noise. The surfaces that hold a difference for years are the ones nobody in the building has audited even once.
Here is the inventory, and what each surface gives up from each side.
| The surface | What the public page tells you | What only the call tells you |
|---|---|---|
| The product | A capability exists, in a screenshot their marketing chose | Where it lives, how many clicks it takes, and whether the rep can reach it live |
| Packaging | Three tiers and a Contact Us | What sits behind the enterprise wall, which line is a paid add-on, what the entry tier quietly excludes |
| Implementation | The words fast onboarding | Who does the work, how long until first value, and who is on the hook when it slips |
| Support | An SLA page | Response time versus resolution time, whose business hours, and the exclusions list |
| Security and compliance | A badge grid | What is certified, what is in progress, and whether a human answers the questionnaire |
| Integrations | A logo wall | Who built it, who maintains it, which objects sync, and in how many directions |
| References | Customer quotes and case studies | Whether anyone will take a call this week, and whether they will name a bad month |
| The demo | A Request a Demo button | How they sell, which is a product surface of its own |
Eight surfaces. In the evaluations we run, the vendors who can produce evidence on more than one or two of them are the exception, and the ones who have never been asked about the bottom half are the rule.
The right-hand column of that table is the board, one row at a time. Packaging becomes the pricing matrix with the quoted number beside the published one.
Support and security become the questionnaire thread with dates on it, and references become the name of the customer who did or did not take the call.
So here is what to do with the ones worth the most.
On the product, stop counting features and start counting clicks. Time the path to the thing you claim to be best at, then watch a rival demo the same job, and note which of you had to leave the software to explain it. The one who leaves is the one with the asterisk.
On packaging, price your own enterprise wall from the outside. Ask a colleague nobody at the vendor recognizes to request pricing at your size, then compare what came back to what your own sales team sends.
The gap between those two documents is a differentiator you own and have never described. Our SaaS enterprise pricing statistics cover how much of that wall is standard practice across the market.
On implementation, answer the question that decides renewals. Who does the work: their team, your team, or a partner the buyer has not met and will be paying separately. Say it in the demo, because whoever says it first sounds like the honest one.
On support, read your own SLA as an adversary would. Circle every clause that sounds like a promise and check whether it is one, and if the exclusions list is where the real policy lives, you have found either a weakness to fix or a strength nobody has ever put in a sentence.
On references, offer the one that costs you something. Everybody has a logo wall. Almost nobody offers a customer who will take a call this week, and the strongest trust signal we have ever watched a vendor produce was a reference who had churned and come back.
And treat the demo as a product surface, not a delivery mechanism. In Reprise’s survey of more than 300 sales and presales professionals, 73 percent of what makes a demo great came down to storytelling, discovery, and speed, with the technical factors sharing the remaining 27 percent. That is practitioner opinion rather than measured deal outcomes, and it points where everything we watch from the buyer’s chair also points.
The demo is not the place you show the product. It is the place you either prove the difference or spend forty minutes proving you belong in the category. What that gate costs you before anyone even reaches the call is its own subject: request a demo is the worst step in your funnel.
Every one of those eight is a place a competitor could be beating you today without anybody in your company knowing the score.
What to Do on Monday
Three things you can do on Monday without a budget, and between them they will show you exactly where the free half stops.
Run the swap on yourselves, from memory. Get the exec team in a room, no laptops, and have everyone write down the company’s differentiation in one sentence. If four people produce four different sentences, stop the meeting, because you have found the problem and it is not the market.
Then swap the names. Take the sentence you agreed on, drop your top competitor’s name into it, and read it aloud to the room. Then do the reverse: take the line off their homepage, put your name in it, and see whether anyone objects.
Then pick one surface you have never audited and go look. Choose the one you would bet you win, because that is the one you would hate to be wrong about, and being wrong about it quietly is what is currently happening.
That is a good afternoon. It does not get you into the conversation, and the conversation is where the last three of those eight surfaces live.
The Sentence Somebody Is Saying About You Right Now
There is a rep, this week, on a call with a buyer who is also talking to you. They are being asked how they compare, and they are answering.
You do not know what they said. You have never heard it. Your positioning was written against a website, and theirs is being delivered live against you, in a room built for exactly that purpose.
That gap is what we go and close. We buy your competitors on your behalf, sit their calls as genuine evaluators, and bring back a board with a page per vendor, the recordings and transcripts behind every line of it, and a verdict on what each of them wins on and where they lose.
Take the differentiation sentence you would least like to defend in front of your own board. Tell us which competitor it is aimed at, and we will go find out whether it survives them.
.png)