Method

Superhuman's PMF Engine: Rahul Vohra Made PMF a Number

Rahul Vohra's method for measuring product-market fit and improving it: the 40% benchmark, the survey mechanics that make it comparable, and the roadmap it produces.

Origin: Rahul Vohra, 'How Superhuman Built an Engine to Find Product/Market Fit', First Round Review, 2018. Builds on Sean Ellis's earlier work in 2009 on the 'very disappointed' survey question.
In short

The Superhuman product-market fit engine is a method for measuring and improving product-market fit, described in 2018 by Superhuman founder Rahul Vohra. It asks active users how disappointed they would be if they could no longer use the product, and treats 40 percent answering 'very disappointed' as the benchmark. It builds on earlier work by Sean Ellis, adding a repeatable loop for raising the score.

When to use

Once you have at least 40 active users you can survey. The framework converts a vague concept (product-market fit) into a measurable score and a four-step loop you can run quarterly to push it upward.

What the Superhuman PMF engine is

The Superhuman product-market fit engine is a method for measuring product-market fit and then systematically improving it. It was described in 2018 by Rahul Vohra, founder of the email client Superhuman, in a First Round Review article.

It asks active users a single question: how would you feel if you could no longer use this product? The score is the share who answer “very disappointed”, and the benchmark is 40%. That benchmark comes from earlier work by Sean Ellis, who observed across a set of startups that products passing it tended to find sustainable growth and products below it tended to struggle.

Vohra’s addition is the engine. The score alone is a diagnosis, and a diagnosis you cannot act on is of limited use. The method turns the survey responses into a roadmap by segmenting them: study the fans to understand what you are already good at, and read the improvement requests from the winnable middle while ignoring the group that was never going to buy.

  1. Survey active users only Not signups

    Someone who did the core action in the last 14 to 28 days. The benchmark assumes this population.

  2. Read the score against 40% A heuristic, not a law

    The share answering "very disappointed". Direction over time matters more than the absolute number.

  3. Segment the responses Where the roadmap comes from

    Fans define the segment worth serving. The middle is winnable. The third group is not your customer.

  4. Split the roadmap in half And re-measure quarterly

    Half on what fans already love, half on what holds the middle back. Nothing for the rest.

One quarter of the engine, start to finish. Survey active users, segment the answers, split the roadmap between protecting the fans and converting the middle, then re-measure.

Why it matters

“Product-market fit” is the most-cited concept in startup advice and the worst-defined. Most founders know they want it. Few can say whether they have it. Almost none can say whether they are getting closer month over month, which means the single most important thing about the company is being tracked by feel.

The engine’s contribution is that it makes fit a number you can move, and it does so without requiring scale. Superhuman ran this while still in a waitlisted beta.

That constraint matters more than it sounds. A method that needs ten thousand users before it produces a signal is useless exactly when you most need one.

  • Very disappointed 46%

    Your fans. The score is the share of this group and nothing else.

    What to do with them: Study them. Their common attributes define the segment worth serving, and their reason for loving it is your positioning.

  • Somewhat disappointed 31%

    The winnable middle. They see something but not enough.

    What to do with them: Read their improvement requests. This is the only group whose feature asks belong on the roadmap.

  • Not disappointed 23%

    Not your customer, whatever they say in the free-text box.

    What to do with them: Nothing. Building for them dilutes what the fans love and lowers the score.

The score and its composition. The number is one proportion, the share answering very disappointed, and the segments underneath it are where the roadmap comes from. Both are read from the same survey.

When to run it

Run it when
  • You have at least thirty genuinely active users and no idea whether you have fit.
  • Growth is inconsistent and you cannot tell whether the product or the funnel is the problem.
  • You are about to raise and need a defensible answer to "do you have product-market fit".
  • The roadmap is driven by whoever complains loudest.
  • You want a number that moves quarterly rather than a feeling that moves daily.
Do not run it when
When the number will mean something, and when it will not. The method needs an engaged user base to survey. Below roughly thirty active users the number swings on a handful of answers, and there are better questions to be asking at that stage.

Running the survey properly

This is where most implementations go wrong, and the errors are quiet: the number still comes out, it just is not comparable to the benchmark it is being read against.

Decision The rule Why it changes the number
Who to survey Active users only Someone who completed the core action in the last 14 to 28 days. Surveying signups, trialists or churned users mixes in people who never engaged and pushes the score down for reasons that are not about fit.
How many responses At least 30, ideally 100+ Below 30 the number swings on a handful of answers. Vohra ran Superhuman's on a few hundred. The score is a proportion, so its stability depends on the count, not on your user base size.
The main question How would you feel if you could no longer use this product? Three options: very disappointed, somewhat disappointed, not disappointed. Do not add a fourth, do not reword it, and do not soften it. The benchmark only means something against the standard wording.
The follow-up questions Three, and the third is the one people skip What type of person would benefit most? What is the main benefit you get? How can we improve it for you? The last one is read only from the somewhat-disappointed group.
Cadence Quarterly One measurement is a number. Two is a direction, and direction is what tells you whether the roadmap is working. Run it once and you have a fact you cannot act on.
Five decisions that change the number. The sampling rule is the one that matters most: a score computed from signups rather than active users is measuring something else and then being compared to a threshold that assumed active users.

The roadmap the score produces

Double down on what fans love 50%

Built from: The main-benefit answers from the very-disappointed group

Protects the thing generating the score. Most teams neglect this half because it feels like standing still.

Address what holds the middle back 50%

Built from: The improvement requests from the somewhat-disappointed group only

Converts the winnable group. Requests from the not-disappointed group are excluded on purpose.

Note who is absent. The 23% who answered “not disappointed” get no roadmap allocation at all. They are the loudest source of feature requests and the least likely to ever become fans, and building for them dilutes the product for the people already paying, which lowers the score you were trying to raise.

Half the roadmap protects what the fans already love, half converts the winnable middle. The instinct to build for the loudest complainers is exactly what the split exists to prevent.

What to do below 40%

The benchmark is a heuristic and a low score is diagnostic rather than fatal. Three moves, in order:

  1. Re-segment before you rebuild. Vohra’s first move was not a product change. It was narrowing the population: filtering to a specific user type raised the score substantially without shipping anything. A mediocre score across everybody often conceals an excellent score inside a segment you have not named yet.
  2. Read the main-benefit answers from the fans. If the fans cannot articulate a consistent benefit, you do not have a positioning problem, you have a product one.
  3. Only then build. Half on the fans’ stated benefit, half on the middle’s blockers.

The failure mode at this stage is rebuilding for the not-disappointed group, on the reasoning that they are the ones who are unhappy. They are not unhappy. They are indifferent, and indifference is not a problem you can build your way out of.

The engine in practice: Superhuman

The case the method is named after, with the numbers it actually produced.

Case study It worked

Superhuman · 2017 to 2019

Turned "do people love it" into a number, then moved the number deliberately.

Rahul Vohra was uncomfortable running a company on the vibe of product-market fit, so he adopted Sean Ellis's survey question: how would you feel if you could no longer use this product? The benchmark for fit is 40% answering "very disappointed".

Superhuman started at 22%. Rather than treating that as a verdict, Vohra segmented it: he separated the people who would be very disappointed, found what they had in common, and largely ignored the people who would not be, on the grounds that building for them would dilute what the first group valued.

The published figure rose to 58% over roughly a year of building specifically for that segment.

Starting "very disappointed"
22%
Benchmark for fit
40%
After ~1 year
58%

What it shows: The number on its own is a thermometer. The method is the segmentation underneath it, which converts a score into a decision about who to disappoint.

Source: Rahul Vohra, First Round Review, 2018; Sean Ellis, original survey methodology.

The PMF engine vs the alternatives

Superhuman PMF engine Fit

How disappointed would active users be to lose this?

Gives you: One score against a 40% benchmark, plus a segmented roadmap

Net Promoter Score Sentiment

How likely are you to recommend us?

Gives you: A sentiment index. Easier to game with one good interaction

Retention cohorts Behaviour

Do people keep coming back?

Gives you: Revealed behaviour. Stronger evidence, and much slower to read

Customer satisfaction Episode

Were you happy with this interaction?

Gives you: Service quality, which is not the same question as fit

What each measures. The Superhuman question asks about loss, which requires the respondent to have integrated the product into their life. That is a stronger signal than an intention to recommend.

When it won’t help you

  • The 40% benchmark is not a constant

    It comes from Sean Ellis observing a limited set of startups, mostly consumer-facing. It has never been established that the same threshold holds across B2B, enterprise, or categories with long replacement cycles, and it is quoted far more confidently than its evidence supports.

    Instead: Track your own trajectory. A score moving from 22% to 31% is more informative than either number against 40.

  • It measures stated feeling, not behaviour

    Someone can sincerely say they would be very disappointed and still churn next month. High scores alongside high churn are not rare, and the survey cannot see the contradiction.

    Instead: Read it beside cohort retention. If the two disagree, believe the behaviour.

  • It rewards narrowing, which can hide a small market

    Segmenting until the score passes 40% is the method working as designed, and it is also how you end up with excellent fit in a segment too small to build a company on.

    Instead: Size the segment you narrowed to. A great score in a market of four hundred people is a lifestyle business at best.

  • It needs enough active users to be stable

    The score is a proportion. Below about thirty responses it swings on individual answers, and a quarterly change is indistinguishable from noise.

    Instead: Below that threshold, run interviews instead. You will learn more from ten conversations than from a proportion of twelve.

Four honest limits. The 40% threshold is the one to hold loosely: it comes from a small sample of consumer-oriented startups and gets cited as though it were a constant.

ShipFit and the Superhuman PMF engine

ShipFit Stage 7, Will They Pay? Behavioral validation with demand signals and the evidence behind each.

The ICP definition locked at Stage 2 is what you segment the survey responses against, and refining it for the very-disappointed cohort is how positioning sharpens over time. ShipFit front-loads that definition so the segmentation step has something to segment by, rather than being invented after the first disappointing score arrives.

Where this sits in the sequence

The PMF engine runs after you have shipped something and have users engaged enough to have an opinion about losing it.

The PMF engine is the second step here. Ship an experiment, measure whether it produced fit, rank what to try next, then check the economics work before pouring money into growth.

Further reading

  • Rahul Vohra, How Superhuman Built an Engine to Find Product/Market Fit, First Round Review (2018). The source.
  • Sean Ellis’s earlier writing on the 40% benchmark, which is where the threshold originates.
  • Lean Startup validation. The loop this measures the output of.
  • ICE scoring. How to sequence the work the survey generates.
  • Jobs to be Done. Why the fans love it, in a form you can act on.
  • Churn rate. The behavioural number to read this against.
  • ICP. What you segment the responses by.

How to apply Superhuman PMF Engine

  1. 1

    Define an active user, then survey only them

    Someone who completed the core action in the last 14 to 28 days. Surveying signups, trialists or churned users mixes in people who never engaged and depresses the score for reasons unrelated to fit. The 40% benchmark assumes an active-user population, so a score from any other population is not comparable to it.

  2. 2

    Ask the question exactly as written

    How would you feel if you could no longer use this product? Very disappointed, somewhat disappointed, not disappointed. Do not add a fourth option, soften the wording, or turn it into a rating scale. The benchmark only means something against the standard three-option form.

  3. 3

    Add the three follow-up questions

    What type of person would benefit most from this? What is the main benefit you get? How can we improve it for you? The first sharpens your segment, the second gives you positioning in customers' own words, and the third is read only from the somewhat-disappointed group.

  4. 4

    Collect at least 30 responses, ideally over 100

    The score is a proportion, so its stability depends on the response count rather than on how large your user base is. Below 30 the number swings on individual answers and a quarterly change becomes indistinguishable from noise.

  5. 5

    Segment before you rebuild

    If the score is below 40%, narrowing the population often raises it substantially without shipping anything. A mediocre score across everybody frequently conceals an excellent one inside a segment nobody has named yet. This was Vohra's first move at Superhuman, and it came before any product change.

  6. 6

    Split the roadmap in half, then re-measure quarterly

    Half on what the very-disappointed group already loves, to protect the thing generating the score. Half on what the somewhat-disappointed group says holds them back. Nothing for the not-disappointed group. One measurement is a number, two is a direction, and direction is what tells you whether the roadmap is working.

Common mistakes

  • **Surveying the wrong audience.** Surveying signups, trialists, or churned users dilutes the score with people who never engaged. Survey active users only, defined as users who completed the core action in the last 14 to 28 days.
  • **Treating 40% as a binary pass/fail.** The 40% threshold is a heuristic, not a law. Some excellent products plateau at 35%; some weak products spike to 45% in narrow segments. Use the score as a directional signal and pay attention to which way it's moving over time.
  • **Acting on the 'not disappointed' segment.** The biggest temptation is to build features for the unhappy users who churned. Don't. Vohra is explicit: those users are not your customer. Building for them dilutes the experience for your fans and lowers your overall score.
  • **Running the survey once.** The PMF Engine is a quarterly loop, not a one-time measurement. The score moves as you ship. If you don't measure quarterly, you cannot tell whether your roadmap is working.
  • **Confusing PMF score with NPS.** NPS asks 'how likely to recommend.' PMF asks 'how disappointed if you lost it.' The second is a stronger signal because it implies the user has integrated the product into their workflow. NPS is easier to game with one good interaction.

How ShipFit operationalizes this

ShipFit treats the Superhuman PMF Engine as the canonical post-launch PMF measure, referenced in Stage 7 (Will They Pay?) and on subsequent product iterations. The 'very disappointed' percentage gives you a number to move; segmenting by 'who would be very disappointed' refines the [ICP](/glossary/icp) you defined at Stage 2. Sahil Lavingia's framing matters: PMF is felt by a specific buyer segment, not the whole user base.

Part of a larger playbook

ShipFit runs 55 frameworks across 9 decision stages

Superhuman PMF Engine is one tool in a bigger toolkit. The full library covers market sizing, buyer discovery, MVP scoping, pricing, and launch.

shipfit.ai/frameworks
Frameworks Library
55 frameworks, mapped to 9 stages

The Mom Test

Q3

Rob Fitzpatrick

Validation question methodology, real interviews, not theater

Jobs-to-be-Done

Q2-Q4

Clayton Christensen

Functional, social, and emotional jobs your product fulfills

7 Powers

Q4

Hamilton Helmer

Strategic moats: Scale, Network, Counter-positioning, Switching, Brand, Cornered Resource, Process

Van Westendorp PSM

Q6

Feature-weighted price sensitivity analysis without guessing

Blue Ocean Strategy

Q4

Kim & Mauborgne

ERRC framework: Eliminate, Reduce, Raise, Create

Fake Door Testing

Q7

Pre-build behavioral validation with landing pages and apology modals

+ 49 more: TAM/SAM/SOM Analysis, Porter's Five Forces, Market Timing Analysis, Unit Economics (LTV/CAC)...

Frequently asked questions

What is Superhuman's PMF Engine?
A framework documented by Rahul Vohra in 2018 for measuring and engineering product-market fit. It uses Sean Ellis's 'very disappointed' survey question as the core metric, segments respondents into fans, on-the-fence, and not-the-customer, then builds a roadmap that half doubles down on what fans love and half fixes what blocks the on-the-fence group. Run quarterly, the loop should push the PMF score upward.
What is the 40% rule for PMF?
Sean Ellis's heuristic, popularized by Vohra: if 40% or more of your active users say they would be 'very disappointed' if they could no longer use your product, you have likely achieved product-market fit. Below 40%, you don't yet. The threshold is a directional signal, not a hard binary. What matters more is whether the number is rising over time as you ship.
Who do I survey for the PMF score?
Active users only. Define active as 'completed the core product action in the last 14 to 28 days.' Do not survey signups, trialists, churned users, or one-time users. Including them dilutes the signal. You want a clean read on people who are actually using the product and could lose it.
What's the difference between PMF score and NPS?
NPS asks 'how likely are you to recommend this product?' on a 0-10 scale. PMF score asks 'how would you feel if you could no longer use this product?' with three options. PMF is a stronger signal because it implies workflow integration, not just a positive moment. A user can recommend a product they barely use; a user cannot honestly say they'd be very disappointed to lose a product they don't depend on.
What if my PMF score is below 40%?
Don't panic. Use the framework to find out why. Profile your fans (whoever said 'very disappointed') to identify your real ICP. Then read the open-text responses from the somewhat-disappointed group to see what they need fixed. Build a roadmap that doubles down on fan benefits and fixes the on-the-fence blockers. Re-run quarterly. The score should rise.
Can the PMF Engine work for B2B?
Yes. Vohra built it for Superhuman, which is B2B-ish (sold to individuals at companies). It works equally well for B2B SaaS aimed at end-users. For pure enterprise B2B where the buyer is not the user, you need to survey the user (for product fit) AND the buyer (for budget fit) separately. Same framework, two surveys.
How does this differ from the Lean Startup's validation approach?
The Lean Startup defines validated learning conceptually but doesn't give you a single number to track. The Superhuman PMF Engine gives you a number (the 'very disappointed' percentage) and a quarterly loop to move it. They're complementary: Lean Startup is the discipline; Superhuman PMF Engine is the operationalization for measuring product-market fit specifically.
What is a good product-market fit score?
40% of active users answering 'very disappointed' is the benchmark, drawn from Sean Ellis's observation across a set of startups that products above it tended to find sustainable growth. Treat it as a heuristic rather than a law: it comes from a limited and largely consumer-facing sample, and it has never been established that the same threshold holds in B2B or enterprise. Your own trajectory matters more than the absolute number, and a move from 22% to 31% is more informative than either figure against 40.
Who exactly should I survey for the PMF score?
Active users only, meaning people who completed your core action in the last 14 to 28 days. This is the single most consequential decision in the method. Surveying signups, trialists or churned users mixes in people who never engaged, which produces a lower number that is measuring something other than fit and then compares it to a threshold that assumed engaged users. Aim for at least 30 responses and preferably more than 100.
What do I do if my PMF score is below 40%?
Segment before you rebuild. Narrowing the population to a specific user type often raises the score substantially without shipping anything, because a mediocre score across everybody frequently conceals an excellent one inside a segment you have not named. That was Vohra's first move at Superhuman and it preceded any product change. Then read the main-benefit answers from your fans, and only then build: half on what fans love, half on what the somewhat-disappointed group says blocks them.
Why not use NPS instead?
They ask different things. NPS asks how likely you are to recommend, which is a social prediction and can be moved by a single good interaction. The Superhuman question asks how you would feel about losing the product, which requires you to have integrated it into your working life to answer at all. That makes it harder to answer positively out of politeness, and a better signal of whether the product has become load-bearing.
Should I build what the not-disappointed users ask for?
No, and this is the counterintuitive core of the method. That group is not unhappy, it is indifferent, and indifference is not a problem you can build your way out of. They are also the loudest source of feature requests, which is what makes the temptation real. Building for them dilutes what your fans value and lowers the score you were trying to raise, so their requests get no roadmap allocation at all.
Related on ShipFit

Keep exploring

Master guide
Validate your business idea

The 9-step playbook from market verdict to ship-ready spec.

Framework
The Mom Test

The Mom Test is Rob Fitzpatrick's framework for customer interviews that generate real signal. Not praise. Three rules, applied step-by-step, with examples.

Framework
Van Westendorp Price Sensitivity Meter

Four survey questions, four cumulative curves, four intersections. How to run the Van Westendorp price sensitivity meter, plot it, and read the price range.

Guide
Market Research

Most founder market research is a TAM slide that nobody believes. The numbers that actually matter are smaller, harder to defend, and tell you whether the market exists for the ten-customer version of your business.

Guide
Idea Validation

Most founders confuse idea validation with idea-receiving-encouragement. The two have nothing in common. Here's what real validation looks like, and the four methods that actually produce it.

Calculator
Startup valuation calculator

What's your SaaS worth? Low, mid, high, in three inputs.

Q&A
How do I find my target market?

Narrow in four passes. (1) Start with the broad category your idea sits in. (2) Filter by buyer behavior: who currently has this problem and is doing something about it? (3) Filter by reach: who can you actually contact via the channels you have today? (4) Filter by willingness to pay: who has budget authority and a price point that clears your unit economics? The output is a specific buyer profile you could name 10 people who match. If you can't, you haven't narrowed enough.

For founders
first-time founders

Startup validation for first-time founders who don't know what they don't know. ShipFit forces 9 decisions and names the mistakes you can't see. Start free.

Comparison
The Mom Test

The Mom Test teaches you how to talk to customers without lying to yourself. ShipFit operationalizes that lesson alongside eight other decisions. Read the book; it's essential. Then use ShipFit to actually run the playbook.

Ready to make your next product a success?

9 decisions between your idea and a product worth building.

No credit card required.

Try an example: