Method

ICE Scoring: Prioritize Features Without Hand-Waving

ICE scoring ranks ideas by Impact times Confidence times Ease. How to score it honestly, where the arithmetic misleads, and how it really differs from RICE.

Origin: Sean Ellis, popularized via GrowthHackers (2014). The framework predates Ellis but he formalized and popularized it for growth experiment prioritization. Also widely used for product feature ranking.
In short

ICE scoring is a method for ranking ideas by multiplying three estimates: impact, confidence and ease. It was formalized and popularized by the growth marketer Sean Ellis around 2014 for prioritizing growth experiments. Its purpose is to make the reasoning behind a ranking explicit rather than to produce a precise number, and the confidence term is what forces evidence into the conversation.

When to use

When you have more candidate items than capacity to ship them. Most useful for feature ranking within an MVP scope, growth experiment selection, and channel prioritization. ICE is the quantitative ranking pass that comes after [MoSCoW](/frameworks/moscow) has done the qualitative bucket sort.

What ICE scoring is

ICE scoring is a method for ranking ideas by multiplying three estimates: Impact, Confidence and Ease. Each is rated 1 to 10, and the product is the score.

It was formalized and popularized by the growth marketer Sean Ellis around 2014, for prioritizing growth experiments at speed. The framework predates him in various forms; what Ellis contributed was the specific three-term version and the practice of scoring a backlog as a team.

Its purpose is not precision. Three subjective estimates multiplied together do not produce an accurate forecast, and treating the output as one is the most common misuse. The purpose is to make the reasoning behind a ranking explicit, and in particular to force the Confidence term into the conversation, where it functions as an audit of how much evidence anyone actually has.

  1. Name the metric first Or Impact is meaningless

    Impact on what? Different target metrics produce different rankings from the same backlog.

  2. Score Impact, Confidence and Ease 1 to 10 each

    All three carry equal weight because they multiply. A 9-9-1 scores below a 6-6-6.

  3. Multiply, then sort The easy part

    Impact × Confidence × Ease. No weighting, no division, which is what makes it usable in a meeting.

  4. Read it as an input, not a verdict And re-score

    It ranks; you still decide. Confidence moves as experiments run, so the ranking is live rather than fixed.

What the score is made of, and what it is worth. Three estimates multiply into one score, Confidence is the term that does the real work, and the ranking is an input to a decision rather than the decision.

Why it matters

Most founder feature lists are ranked by enthusiasm. What the founder is excited about floats to the top, what the team agrees is important sits in the middle, and what nobody loves slides off the bottom. This produces a defensible-looking order with no reasoning attached, which means it cannot be argued with and cannot be wrong.

ICE’s contribution is that it makes the disagreement specific. Two people who rank an item differently now have to disagree about a number and say which one, and that conversation is usually more valuable than the score it produces.

The Confidence term is where this bites. Asking “what is your confidence, and what evidence is it based on?” reliably exposes that the exciting idea is a hunch and the boring one has data behind it.

Expect that conversation to be mildly unpleasant the first two or three times. That is the framework working. A scoring session where everybody agrees immediately has not scored anything.

When to run it

Run it when
  • You have more validated candidates than capacity and need an order.
  • The roadmap is ranked by whoever argued hardest.
  • You are running growth experiments and need to pick this week's three.
  • Two people disagree about priority and cannot say why.
  • You want to make the evidence behind each estimate visible.
Do not run it when
When a ranking helps, and when it misleads. ICE ranks things you have already decided are in scope. It has no view on whether an item should be considered at all, and no view on what is required for a release.

A worked ranking

Candidate feature Impact Confidence Ease Score
Collect feedback from one channel 9 9 7 567
Email digest of new feedback 6 8 9 432
Slack integration 6 7 7 294
Summarize feedback automatically 9 6 4 216
Draft a spec from the summary 8 5 5 200
Public roadmap page 3 6 6 108
Custom fields and tagging 3 5 4 60
SSO and audit logs 4 7 2 56
The same eight candidate features that appear on the MoSCoW page, scored and sorted by ICE. Note what moves: the email digest outranks the spec drafter despite being far less important, because it is easy and everyone is sure it works.

That ordering is the framework’s characteristic distortion and it is worth sitting with. The email digest is a small feature nobody would call strategic. It ranks near the top because Ease and Confidence are both high, and they carry the same weight as Impact.

ICE will reliably surface cheap, certain, low-value work. That is a feature when you are running weekly growth experiments and a serious problem when you are building a product, which is why the ranking is an input to a decision rather than the decision.

Confidence is the audit

Impact and Ease are estimates. Confidence is a claim about evidence, and it is the only one of the three that can be checked.

  1. A hunch

    Somebody suggested it. No data, no research, no precedent in your own product.

  2. Indirect evidence

    A competitor does it, an article recommends it, or it worked at someone else's company. You cannot see whether it worked or why.

  3. Direct evidence

    User research, support tickets, or observed behaviour in your own product pointing the same way.

  4. Already proven here

    A prior experiment in your product, with the same buyer, produced this result. You are repeating something that worked.

Note where a competitor doing something lands. It feels like strong evidence and it is not, because you can see that they shipped it and not whether it worked, for whom, or whether they have since regretted it.

A published scale for what evidence justifies what score. Having one written down is what makes 'that is a 9?' a reasonable question in a scoring session rather than a challenge to somebody's judgement.

Scores drift upward over time if nobody enforces this, because a high Confidence gets your idea built. The term is worth defending precisely because it is the one with an incentive attached.

ICE vs RICE and the others

ICE
Impact × Confidence × Ease
Strength
Fast. Three estimates, no research required, usable in a meeting.
Weakness
Ignores how many people a change reaches, so a fix for 2% of users can outrank one for everybody.
Reach for it when
Growth experiments and early-stage backlogs where everything reaches roughly the same audience.
RICE
(Reach × Impact × Confidence) ÷ Effort
Strength
Adds reach, so scope of audience is priced in. Effort divides rather than multiplying, which penalises large work harder.
Weakness
Needs a reach estimate per item, which is the expensive input and often invented.
Reach for it when
Mature products with distinct user segments and features of wildly varying scope.
Weighted scoring
Σ (criterion × weight)
Strength
Any criteria you like, weighted to match strategy.
Weakness
The weights are an argument, and the framework hides the argument inside a number.
Reach for it when
Cross-functional decisions where the criteria genuinely differ from team to team.
MoSCoW
Four buckets, no arithmetic
Strength
Forces a binary on the only question that matters for a release: required or not.
Weakness
Does not sequence anything inside a bucket.
Reach for it when
The first cut. Run it before ICE, then score the Must list.
Four scoring methods and what each is for. The ICE-versus-RICE difference is narrower than it is usually presented: RICE adds reach and divides by effort, which matters when features reach very different audiences and does not when they all reach roughly the same people.

The practical rule: use ICE when everything on the list reaches a broadly similar audience, which is the normal early-stage case. Move to RICE when that stops being true, usually once you have distinct user segments and features that serve only one of them.

When it won’t help you

  • The output looks more precise than the inputs deserve

    Three subjective 1-to-10 estimates multiplied together produce a three-digit number, and three-digit numbers read as measurements. A score of 336 against 320 is noise, not a ranking.

    Instead: Treat the score as a band rather than a rank. Anything within roughly 15% is a tie, and ties are broken by judgement.

  • It systematically favours small, certain, low-value work

    Because Ease and Confidence carry the same weight as Impact, a cheap change everyone is sure about beats an ambitious one nobody has evidence for. Run a roadmap on ICE alone for a year and you get a very well-optimized version of what you already had.

    Instead: Ring-fence capacity for high-impact low-confidence work before you score. The framework cannot protect that work; only a policy can.

  • It ignores reach entirely

    A fix affecting 2% of users can outrank one affecting everybody, because nothing in the formula knows how many people are involved.

    Instead: Use RICE once your features stop reaching similar audiences.

  • It ranks, it does not decide

    The top five items by score may not add up to a coherent release. The framework has no concept of dependency, sequence or whether the result makes sense as a product.

    Instead: Sort with ICE, then read the top of the list and ask whether it is a thing. Frequently it is not.

Four honest limits. The first two are the ones that matter in practice: the score looks far more precise than its inputs justify, and the arithmetic quietly rewards small easy work.

ShipFit and ICE

ShipFit Stage 5, What's V1? Each feature carries an ICE-style score (0-10) and an S/M/L effort tag, sorted into Differentiator, Delight, and Operational buckets so the build sequence is ranked, not guessed.

ShipFit uses ICE-style scoring in Stage 5 (What’s V1?) to rank candidate features once the MoSCoW sort has decided what is in scope. Each feature carries a score and an effort tag, grouped into Differentiator, Delight and Operational buckets so the build sequence is ranked rather than guessed. The discipline ICE forces, naming the evidence behind each Confidence number instead of guessing, is the part founders skip, and the scored list makes that gap visible.

Where this sits in the sequence

ICE comes third: after the release goal is known and after the scope has been cut.

ICE is the third step. Establish what the release must accomplish, cut to what it cannot ship without, sequence what survived, then ship it and measure.

Further reading

  • Sean Ellis and Morgan Brown, Hacking Growth (2017). The ICE method in the context it was designed for.
  • Intercom’s writing on RICE, which introduced the reach term and the effort divisor.
  • MoSCoW. The first cut, run before this one.
  • Lean Startup validation. What the ranked experiments feed.
  • Superhuman PMF engine. Where the roadmap split comes from once you have users to survey.
  • Jobs to be Done. How to establish the metric that Impact is measured against.

How to apply ICE Scoring

  1. 1

    Name the metric before you score anything

    Impact on what? Signups, activation, revenue, retention? Different target metrics produce different rankings from the same backlog, so a score computed without a named metric is arbitrary. Write the metric at the top of the sheet.

  2. 2

    Score Impact against that metric

    1 to 10 against the named metric, not against how interesting the work is. A useful calibration: 10 is a change that doubles the metric, 5 is a meaningful but incremental improvement, 1 is a rounding error.

  3. 3

    Score Confidence, and demand evidence for anything above 7

    This is the audit. If you cannot name three pieces of supporting evidence, Confidence is 5 or below. Prior experiment results, user research and observed behaviour in your own product all count. A competitor doing it does not, because you cannot see whether it worked for them.

  4. 4

    Score Ease, and get the number from whoever will build it

    Effort and complexity, inverted: 10 is a copy change you ship today, 1 needs three engineers for a quarter. Founders without a build background systematically overestimate Ease, so this number should come from the person doing the work.

  5. 5

    Multiply, sort, and treat close scores as ties

    Impact times Confidence times Ease. Three subjective estimates multiplied together produce a number that looks far more precise than its inputs justify, so anything within roughly 15% is a tie to be broken by judgement rather than by the decimal.

  6. 6

    Read the top of the list and check it is a coherent release

    ICE has no concept of dependency or sequence, and the top five items by score frequently do not add up to something anyone would ship. The ranking is an input to the decision, not the decision. Re-score every four to eight weeks as Confidence moves.

Common mistakes

  • **Inflating Confidence for pet projects.** The most common failure. If you can't name three pieces of evidence supporting the Impact estimate, Confidence should be 5 or below. Honest Confidence scores are the framework's value.
  • **Treating Ease as a tiebreaker rather than a multiplier.** Ease has the same weight as Impact and Confidence in the multiplication. A 9-9-1 score (199) ranks below a 6-6-6 score (216). High-effort items need exceptional impact AND confidence to make the cut.
  • **Scoring without a defined metric.** 'Impact on what?' If the metric isn't named, scores are arbitrary. Define the target metric before scoring. Different metrics produce different rankings.
  • **Doing ICE once and never updating.** As experiments run, Confidence scores should update. As capacity changes, the cutoff line shifts. ICE is a live ranking, not a static one.
  • **Conflating ICE with deciding what to do.** ICE is a sorting tool. It tells you what's ranked higher than what. It doesn't tell you what to actually do. That requires the qualitative judgment about whether the highest-ranked items make a coherent product or campaign.

How ShipFit operationalizes this

ShipFit applies ICE-style scoring inside Stage 5 (What's V1?), where the feature matrix ranks every candidate by Impact, Confidence, and Ease (with an S/M/L effort tag). Stage 8 (How to Launch?) uses the same logic to rank channels. The discipline ICE forces, naming the evidence behind each Confidence number instead of guessing, is the part founders skip; the scored feature list makes that gap visible.

Part of a larger playbook

ShipFit runs 55 frameworks across 9 decision stages

ICE Scoring is one tool in a bigger toolkit. The full library covers market sizing, buyer discovery, MVP scoping, pricing, and launch.

shipfit.ai/frameworks
Frameworks Library
55 frameworks, mapped to 9 stages

The Mom Test

Q3

Rob Fitzpatrick

Validation question methodology, real interviews, not theater

Jobs-to-be-Done

Q2-Q4

Clayton Christensen

Functional, social, and emotional jobs your product fulfills

7 Powers

Q4

Hamilton Helmer

Strategic moats: Scale, Network, Counter-positioning, Switching, Brand, Cornered Resource, Process

Van Westendorp PSM

Q6

Feature-weighted price sensitivity analysis without guessing

Blue Ocean Strategy

Q4

Kim & Mauborgne

ERRC framework: Eliminate, Reduce, Raise, Create

Fake Door Testing

Q7

Pre-build behavioral validation with landing pages and apology modals

+ 49 more: TAM/SAM/SOM Analysis, Porter's Five Forces, Market Timing Analysis, Unit Economics (LTV/CAC)...

Frequently asked questions

What is ICE Scoring?
A prioritization framework that ranks candidate items by multiplying three 1-10 scores: Impact (how much it moves the metric), Confidence (how sure you are about Impact), and Ease (how easy to ship). ICE Score = I × C × E. Popularized by Sean Ellis at GrowthHackers around 2014 for growth experiment ranking, but applicable to any prioritization decision.
How is ICE different from RICE?
RICE adds Reach (how many users this affects) as a fourth multiplier: Reach × Impact × Confidence × Effort. RICE is more rigorous for B2C and large-user-base products where reach varies dramatically by feature. ICE is simpler and works well when reach is roughly equivalent across candidates (typical early-stage SaaS). Most ICE users default to ICE because it's faster; consider RICE when reach distinguishes meaningfully across your candidates.
Why does Confidence inflation matter?
Without honest Confidence scores, ICE becomes a tool for justifying preexisting preferences. Pet projects get Confidence 9; uninteresting projects get Confidence 5. The math then ranks the pet projects highest regardless of actual evidence. The framework's value depends on rigorous Confidence assessment, which requires forcing yourself (or a co-founder) to demand evidence for any Confidence above 7.
What's a good Impact score?
Impact is relative to your business stage. For a pre-PMF startup, Impact = 10 might mean 'doubles signup conversion' or 'launches a meaningful new revenue stream.' For a scaling company, Impact = 10 might mean '5% improvement in a $50M annual revenue line.' The scale is your business; the same point change means different things at different stages.
How often should I re-score?
Every sprint cycle (typically 2-4 weeks) for active growth experimentation. Quarterly for product roadmap items where Confidence shifts more slowly. Re-scoring is where Confidence updates as experiments produce signal. The framework is dynamic; static scores defeat the purpose.
Can ICE be used outside software?
Yes. ICE is a general-purpose prioritization tool. Marketing campaigns, hiring decisions, content topics, partnership negotiations. The 'Impact on metric' framing transposes naturally. The framework's origin in growth experimentation is incidental.
What's the cutoff for what to actually build vs deprioritize?
Depends on capacity, not score. Sort items by ICE score, then fill capacity from the top. If your team's monthly capacity is 200 engineer-hours and the top 5 items by ICE consume 220 hours, the 5th item gets cut. The cutoff is wherever capacity runs out. Items below the cutoff aren't bad; they're just lower priority for this cycle.
What is the difference between ICE and RICE scoring?
RICE adds Reach and divides by Effort: (Reach times Impact times Confidence) divided by Effort. ICE has three terms, all multiplied: Impact times Confidence times Ease. The practical difference is that RICE prices in how many people a change touches, which matters when your features serve very different audience sizes and does not when everything reaches roughly the same users. Use ICE for early-stage backlogs and growth experiments, and move to RICE once you have distinct user segments and features that serve only one of them.
Why does ICE favour small, easy work?
Because Ease and Confidence carry exactly the same weight as Impact in the multiplication. A cheap change everyone is sure about outranks an ambitious one nobody has evidence for, every time. Run a roadmap purely on ICE for a year and you get a very well-optimized version of what you already had. The fix is a policy rather than a formula: ring-fence capacity for high-impact low-confidence work before you score, because the framework cannot protect it.
How do I stop Confidence scores being inflated?
Require evidence for anything above 7 and say what counts: prior experiment results, user research, observed behaviour in your own product. A competitor shipping it does not count, because you cannot see whether it worked for them. Scores drift upward if nobody enforces this, since a high Confidence gets your idea built, which makes the term with the strongest incentive attached also the one most worth defending.
Is an ICE score of 336 better than 320?
Not meaningfully. Three subjective one-to-ten estimates multiplied together produce a three-digit number, and three-digit numbers read like measurements when they are not. Treat anything within roughly 15% as a tie and break it with judgement. The value of ICE is that it makes the reasoning explicit, not that it produces a precise ordering.
Who should do the scoring?
The people with the evidence, which is usually not one person. Impact is best estimated by whoever owns the metric, Confidence by whoever has seen the research, and Ease by whoever will build it. Scoring alone produces a defensible-looking list with one person's assumptions inside it, and the disagreements the exercise surfaces are usually more valuable than the ranking it produces.
Related on ShipFit

Keep exploring

Master guide
Validate your business idea

The 9-step playbook from market verdict to ship-ready spec.

Framework
Buyer Persona Canvas

Adele Revella's Five Rings of Buying Insight, the questions that produce each one, and why a persona built without buyer interviews informs no decision.

Framework
Superhuman PMF Engine

Rahul Vohra's method for measuring product-market fit and improving it: the 40% benchmark, the survey mechanics that make it comparable, and the roadmap it produces.

Guide
MVP Scope

Most founders ship an MVP that's actually V1.3 with bugs. Real MVP scoping cuts ruthlessly until you can name the one hypothesis V1 proves, and ships a product that tests it.

Guide
Competitive Analysis

Most early-stage competitive analysis is a 2x2 with your product in the top-right quadrant. The real version is harder, more boring, and tells you whether you can actually win.

Calculator
Pricing strategy calculator

Van Westendorp in 4 numbers. Skip the survey-platform fees.

Q&A
How do you write a startup problem statement?

Use the Jobs-to-be-Done switch format: 'When [situation], I want to [motivation], so I can [outcome].' Drawn from real buyer interviews, not your hypotheses. The statement should make a specific person feel a specific moment. Pass four tests: it names a real situation, the buyer can identify the moment, the outcome is measurable, and at least 3 unrelated buyers describe the same problem in similar words.

For founders
indie hackers

For indie hackers who've wasted months on dead ideas. ShipFit forces 9 decisions before you write a line of code. Proven frameworks, exports to Cursor.

Comparison
Product Hunt feedback

Posting on Product Hunt gets you real reactions from real people, after you've built something. ShipFit pressure-tests the idea privately before you write a line of code. Use ShipFit before, Product Hunt at launch.

Ready to make your next product a success?

9 decisions between your idea and a product worth building.

No credit card required.

Try an example: