The Superhuman product-market fit engine is a method for measuring and improving product-market fit, described in 2018 by Superhuman founder Rahul Vohra. It asks active users how disappointed they would be if they could no longer use the product, and treats 40 percent answering 'very disappointed' as the benchmark. It builds on earlier work by Sean Ellis, adding a repeatable loop for raising the score.
Once you have at least 40 active users you can survey. The framework converts a vague concept (product-market fit) into a measurable score and a four-step loop you can run quarterly to push it upward.
What the Superhuman PMF engine is
The Superhuman product-market fit engine is a method for measuring product-market fit and then systematically improving it. It was described in 2018 by Rahul Vohra, founder of the email client Superhuman, in a First Round Review article.
It asks active users a single question: how would you feel if you could no longer use this product? The score is the share who answer “very disappointed”, and the benchmark is 40%. That benchmark comes from earlier work by Sean Ellis, who observed across a set of startups that products passing it tended to find sustainable growth and products below it tended to struggle.
Vohra’s addition is the engine. The score alone is a diagnosis, and a diagnosis you cannot act on is of limited use. The method turns the survey responses into a roadmap by segmenting them: study the fans to understand what you are already good at, and read the improvement requests from the winnable middle while ignoring the group that was never going to buy.
- Survey active users only Not signups
Someone who did the core action in the last 14 to 28 days. The benchmark assumes this population.
- Read the score against 40% A heuristic, not a law
The share answering "very disappointed". Direction over time matters more than the absolute number.
- Segment the responses Where the roadmap comes from
Fans define the segment worth serving. The middle is winnable. The third group is not your customer.
- Split the roadmap in half And re-measure quarterly
Half on what fans already love, half on what holds the middle back. Nothing for the rest.
Why it matters
“Product-market fit” is the most-cited concept in startup advice and the worst-defined. Most founders know they want it. Few can say whether they have it. Almost none can say whether they are getting closer month over month, which means the single most important thing about the company is being tracked by feel.
The engine’s contribution is that it makes fit a number you can move, and it does so without requiring scale. Superhuman ran this while still in a waitlisted beta.
That constraint matters more than it sounds. A method that needs ten thousand users before it produces a signal is useless exactly when you most need one.
- Very disappointed 46%
Your fans. The score is the share of this group and nothing else.
What to do with them: Study them. Their common attributes define the segment worth serving, and their reason for loving it is your positioning.
- Somewhat disappointed 31%
The winnable middle. They see something but not enough.
What to do with them: Read their improvement requests. This is the only group whose feature asks belong on the roadmap.
- Not disappointed 23%
Not your customer, whatever they say in the free-text box.
What to do with them: Nothing. Building for them dilutes what the fans love and lowers the score.
When to run it
- You have at least thirty genuinely active users and no idea whether you have fit.
- Growth is inconsistent and you cannot tell whether the product or the funnel is the problem.
- You are about to raise and need a defensible answer to "do you have product-market fit".
- The roadmap is driven by whoever complains loudest.
- You want a number that moves quarterly rather than a feeling that moves daily.
- You have very few users, so the proportion is unstable. Use The Mom Test →
- You need to know why people switch to you at all. Use Jobs to be Done →
- You need to sequence the work the survey generates. Use ICE scoring →
- The question is what to charge. Use Van Westendorp →
Running the survey properly
This is where most implementations go wrong, and the errors are quiet: the number still comes out, it just is not comparable to the benchmark it is being read against.
| Decision | The rule | Why it changes the number |
|---|---|---|
| Who to survey | Active users only | Someone who completed the core action in the last 14 to 28 days. Surveying signups, trialists or churned users mixes in people who never engaged and pushes the score down for reasons that are not about fit. |
| How many responses | At least 30, ideally 100+ | Below 30 the number swings on a handful of answers. Vohra ran Superhuman's on a few hundred. The score is a proportion, so its stability depends on the count, not on your user base size. |
| The main question | How would you feel if you could no longer use this product? | Three options: very disappointed, somewhat disappointed, not disappointed. Do not add a fourth, do not reword it, and do not soften it. The benchmark only means something against the standard wording. |
| The follow-up questions | Three, and the third is the one people skip | What type of person would benefit most? What is the main benefit you get? How can we improve it for you? The last one is read only from the somewhat-disappointed group. |
| Cadence | Quarterly | One measurement is a number. Two is a direction, and direction is what tells you whether the roadmap is working. Run it once and you have a fact you cannot act on. |
The roadmap the score produces
Built from: The main-benefit answers from the very-disappointed group
Protects the thing generating the score. Most teams neglect this half because it feels like standing still.
Built from: The improvement requests from the somewhat-disappointed group only
Converts the winnable group. Requests from the not-disappointed group are excluded on purpose.
Note who is absent. The 23% who answered “not disappointed” get no roadmap allocation at all. They are the loudest source of feature requests and the least likely to ever become fans, and building for them dilutes the product for the people already paying, which lowers the score you were trying to raise.
What to do below 40%
The benchmark is a heuristic and a low score is diagnostic rather than fatal. Three moves, in order:
- Re-segment before you rebuild. Vohra’s first move was not a product change. It was narrowing the population: filtering to a specific user type raised the score substantially without shipping anything. A mediocre score across everybody often conceals an excellent score inside a segment you have not named yet.
- Read the main-benefit answers from the fans. If the fans cannot articulate a consistent benefit, you do not have a positioning problem, you have a product one.
- Only then build. Half on the fans’ stated benefit, half on the middle’s blockers.
The failure mode at this stage is rebuilding for the not-disappointed group, on the reasoning that they are the ones who are unhappy. They are not unhappy. They are indifferent, and indifference is not a problem you can build your way out of.
The engine in practice: Superhuman
The case the method is named after, with the numbers it actually produced.
Superhuman · 2017 to 2019
Turned "do people love it" into a number, then moved the number deliberately.
Rahul Vohra was uncomfortable running a company on the vibe of product-market fit, so he adopted Sean Ellis's survey question: how would you feel if you could no longer use this product? The benchmark for fit is 40% answering "very disappointed".
Superhuman started at 22%. Rather than treating that as a verdict, Vohra segmented it: he separated the people who would be very disappointed, found what they had in common, and largely ignored the people who would not be, on the grounds that building for them would dilute what the first group valued.
The published figure rose to 58% over roughly a year of building specifically for that segment.
- Starting "very disappointed"
- 22%
- Benchmark for fit
- 40%
- After ~1 year
- 58%
What it shows: The number on its own is a thermometer. The method is the segmentation underneath it, which converts a score into a decision about who to disappoint.
The PMF engine vs the alternatives
How disappointed would active users be to lose this?
Gives you: One score against a 40% benchmark, plus a segmented roadmap
How likely are you to recommend us?
Gives you: A sentiment index. Easier to game with one good interaction
Do people keep coming back?
Gives you: Revealed behaviour. Stronger evidence, and much slower to read
Were you happy with this interaction?
Gives you: Service quality, which is not the same question as fit
When it won’t help you
- The 40% benchmark is not a constant
It comes from Sean Ellis observing a limited set of startups, mostly consumer-facing. It has never been established that the same threshold holds across B2B, enterprise, or categories with long replacement cycles, and it is quoted far more confidently than its evidence supports.
Instead: Track your own trajectory. A score moving from 22% to 31% is more informative than either number against 40.
- It measures stated feeling, not behaviour
Someone can sincerely say they would be very disappointed and still churn next month. High scores alongside high churn are not rare, and the survey cannot see the contradiction.
Instead: Read it beside cohort retention. If the two disagree, believe the behaviour.
- It rewards narrowing, which can hide a small market
Segmenting until the score passes 40% is the method working as designed, and it is also how you end up with excellent fit in a segment too small to build a company on.
Instead: Size the segment you narrowed to. A great score in a market of four hundred people is a lifestyle business at best.
- It needs enough active users to be stable
The score is a proportion. Below about thirty responses it swings on individual answers, and a quarterly change is indistinguishable from noise.
Instead: Below that threshold, run interviews instead. You will learn more from ten conversations than from a proportion of twelve.
ShipFit and the Superhuman PMF engine

The ICP definition locked at Stage 2 is what you segment the survey responses against, and refining it for the very-disappointed cohort is how positioning sharpens over time. ShipFit front-loads that definition so the segmentation step has something to segment by, rather than being invented after the first disappointing score arrives.
Where this sits in the sequence
The PMF engine runs after you have shipped something and have users engaged enough to have an opinion about losing it.
- Lean validation
Run the build-measure-learn loop on one risky assumption at a time.
- Superhuman PMF
Measure product-market fit and get a roadmap for raising it.
You are here
- ICE scoring
Rank what to try next as the evidence accumulates.
- CAC / LTV
Confirm the economics work before you pour money into growth.
Further reading
- Rahul Vohra, How Superhuman Built an Engine to Find Product/Market Fit, First Round Review (2018). The source.
- Sean Ellis’s earlier writing on the 40% benchmark, which is where the threshold originates.
- Lean Startup validation. The loop this measures the output of.
- ICE scoring. How to sequence the work the survey generates.
- Jobs to be Done. Why the fans love it, in a form you can act on.
- Churn rate. The behavioural number to read this against.
- ICP. What you segment the responses by.
How to apply Superhuman PMF Engine
- 1
Define an active user, then survey only them
Someone who completed the core action in the last 14 to 28 days. Surveying signups, trialists or churned users mixes in people who never engaged and depresses the score for reasons unrelated to fit. The 40% benchmark assumes an active-user population, so a score from any other population is not comparable to it.
- 2
Ask the question exactly as written
How would you feel if you could no longer use this product? Very disappointed, somewhat disappointed, not disappointed. Do not add a fourth option, soften the wording, or turn it into a rating scale. The benchmark only means something against the standard three-option form.
- 3
Add the three follow-up questions
What type of person would benefit most from this? What is the main benefit you get? How can we improve it for you? The first sharpens your segment, the second gives you positioning in customers' own words, and the third is read only from the somewhat-disappointed group.
- 4
Collect at least 30 responses, ideally over 100
The score is a proportion, so its stability depends on the response count rather than on how large your user base is. Below 30 the number swings on individual answers and a quarterly change becomes indistinguishable from noise.
- 5
Segment before you rebuild
If the score is below 40%, narrowing the population often raises it substantially without shipping anything. A mediocre score across everybody frequently conceals an excellent one inside a segment nobody has named yet. This was Vohra's first move at Superhuman, and it came before any product change.
- 6
Split the roadmap in half, then re-measure quarterly
Half on what the very-disappointed group already loves, to protect the thing generating the score. Half on what the somewhat-disappointed group says holds them back. Nothing for the not-disappointed group. One measurement is a number, two is a direction, and direction is what tells you whether the roadmap is working.
Common mistakes
- **Surveying the wrong audience.** Surveying signups, trialists, or churned users dilutes the score with people who never engaged. Survey active users only, defined as users who completed the core action in the last 14 to 28 days.
- **Treating 40% as a binary pass/fail.** The 40% threshold is a heuristic, not a law. Some excellent products plateau at 35%; some weak products spike to 45% in narrow segments. Use the score as a directional signal and pay attention to which way it's moving over time.
- **Acting on the 'not disappointed' segment.** The biggest temptation is to build features for the unhappy users who churned. Don't. Vohra is explicit: those users are not your customer. Building for them dilutes the experience for your fans and lowers your overall score.
- **Running the survey once.** The PMF Engine is a quarterly loop, not a one-time measurement. The score moves as you ship. If you don't measure quarterly, you cannot tell whether your roadmap is working.
- **Confusing PMF score with NPS.** NPS asks 'how likely to recommend.' PMF asks 'how disappointed if you lost it.' The second is a stronger signal because it implies the user has integrated the product into their workflow. NPS is easier to game with one good interaction.
How ShipFit operationalizes this
ShipFit treats the Superhuman PMF Engine as the canonical post-launch PMF measure, referenced in Stage 7 (Will They Pay?) and on subsequent product iterations. The 'very disappointed' percentage gives you a number to move; segmenting by 'who would be very disappointed' refines the [ICP](/glossary/icp) you defined at Stage 2. Sahil Lavingia's framing matters: PMF is felt by a specific buyer segment, not the whole user base.
ShipFit runs 55 frameworks across 9 decision stages
Superhuman PMF Engine is one tool in a bigger toolkit. The full library covers market sizing, buyer discovery, MVP scoping, pricing, and launch.
The Mom Test
Q3Rob Fitzpatrick
Validation question methodology, real interviews, not theater
Jobs-to-be-Done
Q2-Q4Clayton Christensen
Functional, social, and emotional jobs your product fulfills
7 Powers
Q4Hamilton Helmer
Strategic moats: Scale, Network, Counter-positioning, Switching, Brand, Cornered Resource, Process
Van Westendorp PSM
Q6Feature-weighted price sensitivity analysis without guessing
Blue Ocean Strategy
Q4Kim & Mauborgne
ERRC framework: Eliminate, Reduce, Raise, Create
Fake Door Testing
Q7Pre-build behavioral validation with landing pages and apology modals
+ 49 more: TAM/SAM/SOM Analysis, Porter's Five Forces, Market Timing Analysis, Unit Economics (LTV/CAC)...
Frequently asked questions
What is Superhuman's PMF Engine?
What is the 40% rule for PMF?
Who do I survey for the PMF score?
What's the difference between PMF score and NPS?
What if my PMF score is below 40%?
Can the PMF Engine work for B2B?
How does this differ from the Lean Startup's validation approach?
What is a good product-market fit score?
Who exactly should I survey for the PMF score?
What do I do if my PMF score is below 40%?
Why not use NPS instead?
Should I build what the not-disappointed users ask for?
Keep exploring
The 9-step playbook from market verdict to ship-ready spec.
The Mom Test is Rob Fitzpatrick's framework for customer interviews that generate real signal. Not praise. Three rules, applied step-by-step, with examples.
Four survey questions, four cumulative curves, four intersections. How to run the Van Westendorp price sensitivity meter, plot it, and read the price range.
Most founder market research is a TAM slide that nobody believes. The numbers that actually matter are smaller, harder to defend, and tell you whether the market exists for the ten-customer version of your business.
Most founders confuse idea validation with idea-receiving-encouragement. The two have nothing in common. Here's what real validation looks like, and the four methods that actually produce it.
What's your SaaS worth? Low, mid, high, in three inputs.
Narrow in four passes. (1) Start with the broad category your idea sits in. (2) Filter by buyer behavior: who currently has this problem and is doing something about it? (3) Filter by reach: who can you actually contact via the channels you have today? (4) Filter by willingness to pay: who has budget authority and a price point that clears your unit economics? The output is a specific buyer profile you could name 10 people who match. If you can't, you haven't narrowed enough.
Startup validation for first-time founders who don't know what they don't know. ShipFit forces 9 decisions and names the mistakes you can't see. Start free.
The Mom Test teaches you how to talk to customers without lying to yourself. ShipFit operationalizes that lesson alongside eight other decisions. Read the book; it's essential. Then use ShipFit to actually run the playbook.
Ready to make your next product a success?
9 decisions between your idea and a product worth building.