In one sentence
ROI rubrics should suppress noise rather than display manufactured precision. Use coarse, well-defined Fermi estimates for impact and effort so rankings remain stable under normal estimation error and decisions are easier to explain.
Overview
Cohen begins with the familiar goal of maximizing value per unit of effort, then shows why ordinary ROI scoring often fails. Impact is difficult to define and predict; effort is systematically underestimated; combining uncertain inputs magnifies error. A spreadsheet may rank Feature A above Feature D by a small margin even though that distinction is not justified.
His remedy is to use Fermi-style estimates: values separated by large, meaningful intervals—typically powers of ten or similarly coarse categories. Instead of estimating $34,000 versus $38,000, use categories such as $1,000, $10,000, or $100,000. Instead of fine-grained effort estimates, use a few human-scaled choices such as two days, two weeks, or two months, rounding up.
For qualitative goals, convert vague concepts into concrete questions and define the possible scores behaviorally. If a presentation must reach many people, matter to them, offer genuine insight, and connect to the product, score each dimension with anchors such as “everyone,” “mission-critical,” “changes everything,” or “I’ll buy the product for that.” Multiply dimensions only when all are necessary for one goal; otherwise prioritize the main objective, or add scores when attributes are genuinely equally important.
Coarse scoring will produce ties. Cohen treats this as useful information: a tie means the evidence does not support a finer ranking. Resolve it with a runoff using another important dimension, intentional but carefully designed bias, or human factors such as team excitement and confidence.
Core ideas
Precision can be an illusion
ROI calculations inherit the uncertainty of their inputs. Impact is hard to attribute even after the fact, while effort commonly runs over schedule. More decimal places in the output do not make the decision more accurate.
Use Fermi-sized categories
Restrict estimates to widely separated choices—often powers of ten. Adjacent options should be clearly different, so minor disagreements do not dominate the ranking. The goal is not detailed prediction; it is robust comparison.
Make qualitative criteria observable
Replace broad labels such as “strategic value” or “customer delight” with specific questions. Then define score anchors in terms of recognizable reactions or outcomes, reducing disagreement about what a score means.
Choose the right aggregation rule
Multiply scores when every dimension is required to achieve one outcome. Do not combine unlike objectives such as revenue and delight as though they shared a unit. Prioritize the main objective, use another dimension to break ties, or add values only when the attributes are truly equivalent.
Coarse estimates expose strategic disagreement
If people disagree between $1M and $10M, that debate may reveal different assumptions about customers, markets, or execution. Disagreement between nearby values is usually not worth discussing because estimation error overwhelms it.
Ties are a feature, not a defect
A coarse rubric deliberately refuses to manufacture a winner when alternatives are effectively equivalent. Use a runoff, a justified bias in the spacing of values, or human considerations after the primary business comparison is tied.
Effort estimates should be few and rounded up
A small menu such as two days, two weeks, and two months makes scope disagreements visible and avoids hours of pseudo-analysis. The original Smart Bear process reportedly planned four months of work in a few hours and usually finished within about a week of the estimate.
Practical takeaways
- Replace fine-grained ROI inputs with a short, coarse scale. For impact, consider $1k/$10k/$100k or 1/10/100 customers; for effort, use a few concrete time buckets and round up.
- Write a definition beside every score. “Impact = 10” is weak; “customers are curious and ask to hear more” is usable.
- Before combining metrics, state the decision objective. If revenue is primary, rank by revenue first and use retention, delight, or brand as tie-breakers unless there is a defensible common goal.
- Use multiplication only when every factor is necessary. Otherwise, multiplication can pretend unrelated benefits belong to one coherent objective.
- Treat a narrow ranking as untrustworthy when normal estimation errors could reverse it. Prefer choices separated by an order of magnitude or by a clearly meaningful strategic difference.
- When items tie, run a second comparison focused on the most important unresolved dimension. Do not automatically add more numerical detail.
- Use team excitement or execution confidence only after the primary value comparison is tied, so these factors enrich the decision without obscuring the main strategy.
- Explain decisions in strategic language: “This option has dramatically more potential impact, so months of effort are justified,” rather than “its calculated score was 17 instead of 15.”
Caveats and counterpoints
- Fermi scoring does not make forecasts accurate; it makes the decision less sensitive to false precision. A wrong order-of-magnitude assumption can still produce a bad choice.
- The method depends on selecting meaningful dimensions and score anchors. Poorly chosen questions can make a coarse rubric confidently optimize the wrong thing.
- Multiplying qualitative scores can overweight a dimension or create extreme results. It is appropriate only when all dimensions are necessary to the stated outcome.
- The article’s examples are primarily product-planning and content-selection cases. Regulated, safety-critical, or financially material decisions may require finer evidence, formal uncertainty analysis, or independent review.
- Coarse scales can create many ties. That is informative, but organizations still need an explicit tie-breaking process and decision owner.
- Cohen presents his Smart Bear planning experience as evidence of usefulness, not as a controlled comparison proving that Fermi ROI universally outperforms other prioritization methods.
Questions worth revisiting
- Which objective is genuinely primary for the next planning period: revenue, retention, learning, reliability, strategic position, or something else?
- For each proposed item, what are the few order-of-magnitude impact choices that would be clearly distinguishable?
- What observable behavior should define each qualitative score?
- Which factors are necessary components of one goal, and which are separate objectives that should not be multiplied together?
- Would ordinary optimism or schedule slippage change the ranking? If so, is the ranking too close to trust?
- If several items tie, what runoff dimension would best express the real strategic preference?
- Are we using “confidence” to reflect execution knowledge, or merely disguising a risk estimate as a precise percentage?
Return to this when…
Return to this when a prioritization spreadsheet produces precise-looking rankings from speculative inputs, when stakeholders argue over small score differences, or when qualitative goals are being forced into arbitrary 1–5 ratings. The central diagnostic is: would ordinary estimation error reverse the decision? If yes, make the scale coarser and the definitions clearer.