Model

Confidence Scoring: Rating How Much You Actually Know

A method for scoring the reliability of the evidence behind a decision, so a well-researched conclusion and a confident guess stop looking identical.

Origin: ShipFit's own model. It applies standard evidence-hierarchy thinking, familiar from research methodology, to the inputs behind a product decision. It is a working heuristic rather than a validated instrument.
In short

Confidence scoring rates how reliable the evidence behind a claim is, on a scale from an unsourced assumption to an observed outcome in your own product. Its purpose is to stop a well-researched conclusion and a confident guess from looking identical once they are both written down, which is the state most product decisions are made in.

When to use

Whenever several claims are feeding one decision and they came from very different places. It is a labelling discipline rather than an analysis method, and it takes minutes once the habit exists.

What confidence scoring is

Confidence scoring rates how reliable the evidence behind a claim is, on a scale from an unsourced assumption up to an outcome you observed in your own product.

It is ShipFit’s own model rather than an established framework. It applies evidence-hierarchy thinking, which is entirely standard in research methodology, to the inputs behind an ordinary product decision.

The problem it addresses is specific. Once a claim is written in a document, it loses its provenance. “Buyers will pay around £2,000 a year” reads identically whether it came from forty interviews or from one afternoon’s thinking, and six months later it will be quoted with the same authority either way.

  1. Five tiers of evidence Each worth more than the last

    Assumption, secondary source, stated preference, observed behaviour, your own result.

  2. Score the load-bearing claims Not the whole document

    Most documents contain a dozen assertions and rest on three.

  3. The weakest one governs Never average

    Averaging is how one assumption gets laundered by four good citations.

  4. Low confidence is not a blocker It sizes the bet

    It argues for a cheap reversible decision made quickly, not for more research.

Four things worth knowing before the tiers, and a map of this page.

Why provenance disappears

A claim travels and its source does not. Someone writes “our buyers approve up to £2,000 without escalation” in a strategy document, having heard it once. It gets quoted in a pricing discussion, then in a board deck, then in an onboarding document for a new hire, and by the fourth retelling it is a fact about the business.

Nothing dishonest happened at any step. Documents simply do not carry footnotes well, and the confident phrasing that makes a document readable is the same phrasing that strips out uncertainty.

Tagging costs seconds per claim and is the entire intervention.

It also feels faintly bureaucratic, which is why it does not happen. The moment it pays for itself arrives about nine months later, when somebody asks where the £2,000 figure came from and the honest answer turns out to be that a founder heard it at a conference.

The tiers

How much weight the claim can carry
  1. Assumption Tier 1

    Nobody checked. It sounds right, it came from experience elsewhere, or it arrived in a meeting and nobody objected.

    Sounds like: "Obviously buyers want..." with no source attached.

  2. Secondary source Tier 2

    A benchmark, a report, a competitor's public claim. Describes some population, and the question is whether you are in it.

    Sounds like: "Industry average conversion is 3%." For whom, measured how?

  3. Stated preference Tier 3

    What your own buyers said: interviews, surveys, expressions of interest. About your product, and about the future.

    Sounds like: "Nine of twelve said they would definitely pay for this."

  4. Observed behaviour Tier 4

    What people actually did somewhere. A workaround they built, a competitor they switched to, a process they maintain.

    Sounds like: "Seven of twelve maintain a spreadsheet that does part of this."

  5. Your own result Tier 5

    An outcome you produced and measured: a purchase, a cohort, a completed experiment with a pre-registered threshold.

    Sounds like: "Four of forty on the landing page left a deposit."

Five tiers, from an unsourced assumption to a result you produced yourself. The jump that matters most is between stated and revealed preference, because it is the one people most often treat as no jump at all.
Reported as high confidence

Four tier-4 claims and one tier-1 claim, averaged.

Average 3.4, described as "well evidenced". The tier-1 claim is the price, and the price is the decision.

Reported honestly

Same document, scored on the weakest load-bearing claim.

Tier 1, because the pricing assumption governs everything else. Cost to fix: ten conversations.

A decision is as reliable as its weakest load-bearing claim. Averaging the tiers is how one assumption gets laundered by four good citations, and it is what most confidence self-assessments actually do.

Two decision documents, each with five claims. Both look equally authoritative once written up, and one of them rests on a guess.

When to use it

Run it when
  • Several claims from very different sources are feeding one decision.
  • A number is circulating that nobody can attribute.
  • You are about to make an expensive or irreversible commitment.
  • A conclusion from months ago is still being quoted as current.
  • Two people disagree and cannot tell whether it is about evidence or interpretation.
Do not run it when
When the labelling is worth the seconds it costs, and when the problem is elsewhere.

When it won’t help you

  • It is a labelling discipline, not a measurement

    The tiers are a reasoned ordering with no calibration behind them. Nothing establishes that observed behaviour is worth twice a stated preference, only that it is worth more.

    Instead: Use it to make provenance visible. Do not present the number as though it quantified anything.

  • It can become a reason not to decide

    A low score is easily read as "we need more research", which is the comfortable conclusion. Plenty of good decisions are made deliberately on weak evidence because waiting costs more than being wrong.

    Instead: Pair every low score with the cost of the decision being wrong. Cheap and reversible beats slow and certain most of the time.

  • High-tier evidence about the wrong thing is still wrong

    A perfectly measured result from an unrepresentative sample scores highly and misleads. The tier describes how the evidence was gathered, not whether it was gathered about the right population.

    Instead: Score relevance separately from reliability. They are independent and both can fail.

  • It adds friction people will drop under pressure

    Any process requiring a tag on every claim will be skipped in the week it matters most, which is the week before a launch.

    Instead: Restrict it to the three or four load-bearing claims. A process that takes two minutes survives; one that takes an hour does not.

Four honest limits. The second is the way this model most often does damage, and it does it while looking like rigour.

Further reading

How to apply Confidence Scoring Framework

  1. 1

    List the claims the decision actually rests on

    Not the whole document. The three or four load-bearing statements, the ones where being wrong changes the answer. Most decision documents contain a dozen assertions and rest on very few of them.

  2. 2

    Tag each claim with where it came from

    Assumption, secondary source, stated preference, observed behaviour, or your own result. This takes seconds per claim and is the entire method: the labelling is what stops a guess and a finding reading identically on the page.

  3. 3

    Look at the tier of your weakest load-bearing claim

    A decision is only as reliable as the shakiest thing it rests on. One assumption underneath four well-evidenced claims makes the whole conclusion an assumption, and that is the number to report.

  4. 4

    Ask what it would cost to raise the weakest one

    Frequently very little. Ten interviews, a landing page, a query against your own data. Compare that cost to the cost of the decision being wrong, and the answer is usually obvious once both are written down.

  5. 5

    Record the tier alongside the conclusion

    Not in an appendix. A conclusion travels through an organisation and its provenance does not, so a claim sourced from one blog post is quoted six months later with exactly the authority of one sourced from your own experiment.

  6. 6

    Re-score when evidence arrives

    Tiers move upward as you learn. A claim that was an assumption in January and is an observed outcome by March should be relabelled, and decisions built on the old score revisited.

Common mistakes

  • **Scoring the document rather than the load-bearing claims.** Most documents contain a dozen assertions and rest on three. Scoring everything equally hides which ones matter.
  • **Averaging the tiers.** A decision is as reliable as its weakest load-bearing claim, not as the mean of its claims. Averaging is how one assumption gets laundered by four good citations.
  • **Treating a secondary source as evidence about you.** An industry benchmark describes a population you may not be in, and it is routinely quoted as though it described your product.
  • **Confusing stated with revealed preference.** People saying they would buy is a lower tier than people buying, and the gap is consistent and large.
  • **Hiding the score in an appendix.** Conclusions travel through organisations and provenance does not. The tier has to sit next to the claim or it is lost on the first retelling.
  • **Using it to block decisions.** Low confidence is a reason to make a cheap reversible decision quickly, not a reason to keep researching. The score informs how much to bet, not whether to move.

How ShipFit operationalizes this

ShipFit runs the Confidence Scoring Framework in Stage 6 (How to Charge?), where it rates the reliability of the data behind a pricing recommendation. Live market research and competitor pricing pulled at run time score differently from AI-generated hypotheses, and thin or contradictory inputs surface as Areas to Clarify rather than being presented with the same weight as evidence.

Part of a larger playbook

ShipFit runs 55 frameworks across 9 decision stages

Confidence Scoring Framework is one tool in a bigger toolkit. The full library covers market sizing, buyer discovery, MVP scoping, pricing, and launch.

shipfit.ai/frameworks
Frameworks Library
55 frameworks, mapped to 9 stages

The Mom Test

Q3

Rob Fitzpatrick

Validation question methodology, real interviews, not theater

Jobs-to-be-Done

Q2-Q4

Clayton Christensen

Functional, social, and emotional jobs your product fulfills

7 Powers

Q4

Hamilton Helmer

Strategic moats: Scale, Network, Counter-positioning, Switching, Brand, Cornered Resource, Process

Van Westendorp PSM

Q6

Feature-weighted price sensitivity analysis without guessing

Blue Ocean Strategy

Q4

Kim & Mauborgne

ERRC framework: Eliminate, Reduce, Raise, Create

Fake Door Testing

Q7

Pre-build behavioral validation with landing pages and apology modals

+ 49 more: TAM/SAM/SOM Analysis, Porter's Five Forces, Market Timing Analysis, Unit Economics (LTV/CAC)...

Frequently asked questions

What is confidence scoring?
Rating how reliable the evidence behind a claim is, so a well-researched conclusion and a confident guess stop looking identical once they are both written down. A typical scale runs from an unsourced assumption, through secondary sources and stated preferences, to observed behaviour and finally to a result you produced yourself. The purpose is not precision but visibility: most decisions are made from a mix of tiers with no indication of which is which.
What is the difference between stated and revealed preference?
Stated preference is what people say they will do: survey answers, interview claims, expressions of interest. Revealed preference is what they actually did: paid, switched, spent time. The gap between them is consistent and large, and it runs in one direction, with stated preference systematically more optimistic. Any confidence scale has to put them on different tiers, because treating them as equivalent is the single most common way a decision ends up resting on nothing.
How should I score a claim from an industry benchmark?
As a secondary source, which is a low tier when the claim is about your product. A benchmark describes a population, and the useful question is whether you are in it. Sample composition, definitions and time period all vary between published surveys, which is why two credible benchmarks routinely disagree. It is reasonable evidence about the shape of a market and weak evidence about what your own product will do.
Does low confidence mean I should not decide?
No, and using it that way is the main way this becomes harmful. Low confidence is a reason to make a cheap, reversible decision quickly rather than an expensive, irreversible one slowly. The score tells you how much to bet, not whether to move. Plenty of good decisions are made on weak evidence, deliberately, because the cost of being wrong is smaller than the cost of waiting.
Is this an established framework?
No. It is ShipFit's own model, applying evidence-hierarchy thinking that is standard in research methodology to the inputs behind a product decision. There is no published validation of this particular scale, and a different practitioner would reasonably draw the tiers differently. It is offered as a labelling discipline that costs seconds per claim, not as a measurement.
Related on ShipFit

Keep exploring

Master guide
Validate your business idea

The 9-step playbook from market verdict to ship-ready spec.

Framework
7 Powers

The 7 Powers, each with its benefit and its named barrier, plus the Power Progression that decides which of them a startup can realistically build and when.

Framework
The Lean Startup

Validated learning, the build-measure-learn loop, what an MVP actually is, the three engines of growth, and the ten pivots Ries names rather than one.

Guide
Market Research

Most founder market research is a TAM slide that nobody believes. The numbers that actually matter are smaller, harder to defend, and tell you whether the market exists for the ten-customer version of your business.

Guide
Idea Validation

Most founders confuse idea validation with idea-receiving-encouragement. The two have nothing in common. Here's what real validation looks like, and the four methods that actually produce it.

Calculator
Break-even calculator

How many units a month before the math stops bleeding?

Q&A
How do you write a startup problem statement?

Use the Jobs-to-be-Done switch format: 'When [situation], I want to [motivation], so I can [outcome].' Drawn from real buyer interviews, not your hypotheses. The statement should make a specific person feel a specific moment. Pass four tests: it names a real situation, the buyer can identify the moment, the outcome is measurable, and at least 3 unrelated buyers describe the same problem in similar words.

For founders
fintech founders

Fintech idea validation that tests demand, trust, and willingness to pay before you sink months into a regulated build. Forces 9 decisions. Start free.

Comparison
PRD generators

PRD generators turn your input into a tidy requirements doc fast. ShipFit forces the decisions that make the doc worth writing, then exports the spec. A polished PRD for an unvalidated idea is just well-formatted fiction.

Ready to make your next product a success?

9 decisions between your idea and a product worth building.

No credit card required.

Try an example: