Somewhere in your portfolio, your scarcest capacity is being spent right now on the strength of a number that nobody has ever checked, that nobody will ever check, and that most of the people in the room privately suspect is soft. It is not there because anyone lied. It got there by winning a comparison, and the comparison was rigged in its favor by the one property that should have counted against it: the fact that its value cannot be verified.
This is the half of portfolio economics that gets almost no attention. We have spent a great deal of effort on the denominator. Find your constraint, measure what work actually consumes of it, rank by value per constrained resource hour, stop pretending that a project's size tells you what it costs the system. All of that is right, and all of it is arithmetic performed on a ratio whose top line arrives from somewhere else entirely: a business case, written by the person who wants the project approved, at the moment they know least, and never compared with what happened.
If that top line were merely uncertain, you could live with it. Uncertainty averages out across a portfolio. The problem is that it is not uncertain, it is biased, and it is biased unevenly. Optimism in business cases is not spread flat across your portfolio. It runs along a slope, and the slope has a direction: the harder a claimed benefit is to check, the larger it tends to be, and the less likely it is to be delivered. I want to give that slope a name, because an unnamed distortion gets argued about case by case rather than corrected once: the optimism gradient.
Its consequence is precise and unpleasant. A ranking system fed by numbers that inflate with unverifiability will systematically hand your constraint to your least verifiable work. Not occasionally. Structurally, every cycle, in the same direction.
The half of the equation nobody audits
Start with the asymmetry that makes all of this possible.
Cost overruns self-report. If a project needs 40 per cent more money than it asked for, it has to come back and ask, and someone has to sign. If it needs more of the constraint than it budgeted, the constraint runs out and everything behind it goes late, loudly. The organization cannot avoid learning that its cost forecasts were wrong, because reality arrives with an invoice attached and forces the conversation.
Benefit shortfalls do not self-report. Nothing happens. There is no moment at which a project fails to deliver its benefits in a way that anyone must respond to, because the benefits were due 18 to 36 months after approval, by which point the project is closed, the team is dispersed, the sponsor has moved role, the baseline metric has been redefined twice in a reorganisation, and the system the benefits were supposed to come from has been superseded. The absence of a benefit makes no noise.
So the organization runs an error correction loop on one side of the ratio and no loop at all on the other. Twenty years of that produces exactly what you would predict: cost estimation that is bad but roughly calibrated, sitting on top of value estimation that has never once been marked against reality. Your finance function can tell you, to the dollar, what every project in the portfolio has spent. Ask it what proportion of forecast benefits the last thirty closed projects actually delivered and you will get a pause, then an offer to look into it, then nothing, because the data was never collected in a form that could answer the question.
That is not an oversight in the reporting pack. It means the input to every prioritization decision you make is an instrument that has never been calibrated.
Uniform optimism would be harmless
Here is the part that most people get wrong when they first meet this problem, and it matters, because it determines what the fix has to be.
Suppose every business case in your portfolio inflates its benefits by exactly the same factor. Every claim is double the truth. What happens to your ranking?
Nothing. If A claims $4m and B claims $2m, and both are twice reality, then A is still worth twice B, and the order is untouched. You are wrong about the absolute size of the prize, which matters for the go or no-go decision at the portfolio boundary, but every sequencing decision you make inside the portfolio is unaffected. Uniform optimism is a scaling error, and ranking is immune to scaling errors.
This is why "everyone exaggerates, we all know that, we take it with a pinch of salt" feels like an adequate response. If the exaggeration were uniform, it would be.
It is not uniform. It varies with a specific property of the benefit, and the direction is consistent. Rank your benefit types by how easy it is for a sceptical outsider to check the claim after the fact, and you get something like this. The percentages are illustrative, drawn from the pattern I see repeatedly rather than from your organization, and the whole point of the next section is that you should derive your own:
- Directly measurable cash. Revenue you will invoice, leakage you will stop, a contract line that falls, a licence you will not renew. Someone can go to the ledger in 18 months and settle the argument. These realize at something like 75 to 85 per cent of claim.
- Efficiency converted at an assumed rate. Hours saved, multiplied by a cost per hour. The hours saved are usually real. The money mostly is not, because the freed hours were absorbed rather than removed from the cost base, which makes the money a claim contingent on a management decision that nobody has actually committed to. Realisation of the cash runs nearer 40 to 50 per cent.
- Enablement and platform value. "This opens up $6m of downstream benefit." The value belongs to projects that are not approved, not scheduled, and in some cases not yet conceived. Realisation, honestly attributed, is often 15 to 25 per cent.
- Strategic, optionality and risk avoidance. Positioning, agility, brand, resilience, "the cost of not doing it". There is no measurement that could ever contradict the number, which is precisely what defines this tier. Attributable realisation is in the region of 10 per cent, and the honest answer is that it is unknowable, which for our purposes amounts to the same thing.
Now look at what that does. The tiers with the highest realisation rates are also, almost without exception, the tiers with the smallest claimed numbers. Nobody writes a $7m business case for stopping billing leakage, because billing leakage is a known quantity and $7m is not in it. Everybody can write a $7m business case for resilience.
So the two effects compound rather than cancel. Large claim, low realisation, at one end. Modest claim, high realisation, at the other. That is the gradient, and it is steep enough to invert a ranking.
Where the gradient comes from
Four mechanisms produce it, and none of them requires anyone to behave badly.
Scrutiny is applied in proportion to checkability. When a business case claims $2.4m of recovered billing revenue, finance can and does interrogate it: which customer segments, what leakage rate, over what base, compared with what benchmark. The number gets trimmed to $2.0m before it reaches the board, because it was contestable and someone contested it. When a case claims $5.6m of avoided strategic risk, there is nothing to interrogate. No analyst can mount a challenge to a number with no derivation. It arrives at the board intact. Scrutiny lands where evidence exists, which means it lands in inverse proportion to where it is needed, and the net effect of your governance process is to shrink your most reliable claims while waving through your least reliable ones.
Unfalsifiable claims have no ceiling. A cost saving is bounded by the cost line it comes out of. You cannot claim $3m of savings against a function that costs $1.8m to run, and if you try, someone will notice within a minute. Strategic value has no denominator anywhere. Nothing in the arithmetic caps it. So the number is not derived at all, it is chosen, and it is chosen to clear whatever bar the sponsor believes the investment committee is applying. Raise the bar and the unfalsifiable claims simply rise to meet it, while the measurable ones cannot, which means tightening your approval threshold actively steepens the gradient.
Enablement lets the same value be counted twice. The platform claims the benefits of the five things it will make possible. Two years later, each of those five arrives with its own business case claiming its own benefits, and nobody nets off the amount already booked, because the platform's case is closed and the person who wrote it has moved on. Meanwhile the platform's value was contingent on decisions that had not been taken, which is to say it was an option rather than a benefit. It was priced as a certainty because our case templates have no field for "this is worth $6m if we later approve three other things, and nothing at all if we do not".
The bias is self-reinforcing, because the winners are unauditable by construction. A portfolio learns from disappointment. The projects that win under the gradient are exactly the ones whose disappointment can never be established, so the feedback that would correct the bias is precisely the feedback that cannot be generated. The measurable work, when it underperforms, gets caught, discussed and remembered, which makes the next measurable case more conservative still. The gradient does not decay over time. It steepens.
A worked example
Four proposals compete for the same constrained resource in the coming year. The constraint is a small integration and data engineering gate, the one thing every substantial change in this organization has to pass through. It has about 1,700 hours of capacity available for new work this year once existing commitments are honoured. The numbers are illustrative and rounded, but the shape is one I have seen many times.
Each case has been through governance and arrives with a claimed three-year value and an estimate of the constraint-hours it will consume:
- Billing automation. Recovers invoicing leakage. Claimed $2.4m, needs 900 hours. Value per constraint-hour: $2,667.
- Warehouse process automation. Hours saved across two shifts, costed at an internal rate. Claimed $3.0m, needs 750 hours. Value per constraint-hour: $4,000.
- Unified data platform. Enables five named downstream initiatives. Claimed $7.5m, needs 1,500 hours. Value per constraint-hour: $5,000.
- Resilience program. Avoided outage and regulatory exposure. Claimed $5.6m, needs 800 hours. Value per constraint-hour: $7,000.
This portfolio is doing everything the good practice says. It is not ranking by project size, it is not ranking by whoever shouted loudest, it is ranking by return on the scarce resource, exactly as Stop Ranking Projects by How Big They Are argues it should. The order is unambiguous: resilience, then the platform, then the warehouse, then billing last.
Now apply the organization's own realisation history, using the tier factors above:
- Billing automation, cash tier at 80 per cent: $1.92m over 900 hours, $2,133 per hour.
- Warehouse automation, efficiency tier at 45 per cent: $1.35m over 750 hours, $1,800 per hour.
- Data platform, enablement tier at 20 per cent: $1.5m over 1,500 hours, $1,000 per hour.
- Resilience program, strategic tier at 10 per cent: $560k over 800 hours, $700 per hour.
The order is exactly reversed. The project ranked first is now ranked last, and the project ranked last is now first. I have made the inversion complete for clarity, and a real portfolio usually produces a partial reordering rather than a clean flip, which is less dramatic and costs very nearly as much, because the damage is done by whichever projects end up on the wrong side of the capacity line.
Follow the money through the year.
Under the claimed ranking, the constraint spends its 1,700 hours on the resilience program (800 hours, completed) and 900 of the 1,500 hours the platform needs, so the platform delivers nothing this year and carries into the next. Realistically attributable value delivered by year end: about $560k, plus 900 hours of work in progress sitting in a platform that returns nothing until it is finished, which is the batch size problem arriving on top of the ranking problem.
Under the calibrated ranking, the same 1,700 hours fund billing automation (900 hours) and warehouse automation (750 hours), both completed, with 50 hours spare. Realistically attributable value delivered: about $3.27m.
Same people. Same hours. Same year. Same governance process, run competently by serious people. The difference is $2.71m, and expressed as return on the scarce resource it is $329 per constraint-hour against $1,982 per constraint-hour, a factor of six.
Two things about that gap are worth sitting with. The first is that the resilience program and the platform were not canceled in the second plan, only sequenced later, so this is not an argument about which projects deserve to exist. It is entirely an argument about order, which is where most portfolio value is won and lost. The second is that nobody in the first plan will ever discover the $2.71m. There is no report on which it appears. The resilience program will be marked delivered, on time and to budget, and it will have been, which is the Project Illusion in its purest form: every project green, the portfolio running at a sixth of the value it could have returned.
Why single-project management is blind to this
The gradient survives because no single-project instrument can detect it, and a portfolio made of single-project instruments has no other eyes.
A business case is evaluated against a threshold, not against its peers. The question asked in the room is "does this clear our hurdle rate", which is a question about one project in isolation. The gradient is invisible at that altitude, because a distortion in the comparison between cases produces no anomaly in any individual case. Every one of the four proposals above is defensible on its own. The damage exists only in the ordering, and nothing in the governance process is looking at the ordering.
Benefits tracking, where it exists at all, is owned by the project that made the claim, which is a control failure so basic it would be unacceptable anywhere else in the organization. We would not let a supplier audit its own invoices. We do let a program grade its own benefits, after the point at which anyone still cares, using a baseline it is permitted to restate.
And the accountability horizon is shorter than the benefit horizon, always. The people who approve the case are assessed on delivery, not on realisation, because delivery is what happens during their tenure. So the entire incentive structure of the approval process points at getting things approved and finished, and stops precisely where the value was supposed to begin. That is the Execution Gap with the direction of causation made explicit: it is not only that strategy fails to reach delivery, it is that delivery never has to report back to strategy.
Underneath all of it sits the assumption that makes the ratio feel rigorous. You have worked hard on the denominator. You know your constraint, you know what each project draws from it, and you can express the answer in dollars per constraint-hour to a level of precision that feels like control. Precision in the denominator does not repair a bias in the numerator. It disguises it. An uncalibrated instrument that reads to three decimal places is more dangerous than one that admits it is a guess, because people act on it.
What to do instead
1. Measure realisation once, by benefit type, and you have most of the fix. Take the last twenty to thirty closed projects. For each, record the claimed benefit, the benefit type, and the best available evidence of what actually landed. The evidence will be poor. Use it anyway: a median ratio built from thirty imperfect observations is enormously better than the implicit factor of 1.0 you apply today, which is not a measurement at all but an assumption, and a demonstrably false one. This is a four to six week exercise for one analyst, it needs no new tooling, and it is very close to the argument in Your PMO Already Has the Data to Become a VMO. You are not building a benefits management system. You are calibrating an instrument you already rely on.
2. Make claims falsifiable as a condition of being ranked, not as a condition of being approved. Every case states the metric, its current level, the expected level, the date it will be checked, and the person who will be asked. Cases that cannot state those five things are not rejected. They are assigned the lowest realisation factor by default, because a claim nobody can check is, empirically, a claim that usually does not land. This converts an argument into a choice: a sponsor who wants a better factor makes their claim checkable, voluntarily, at the point where doing so is cheap. You will be surprised how much supposedly unmeasurable value acquires a metric in the week after this rule takes effect.
3. Apply the discount as a published standing policy, before the debate rather than inside it. The factors are set once, by benefit type, on the historical evidence, and applied mechanically to every case including the ones sponsored by the chief executive. This matters more than it sounds. If the haircut is applied case by case in the meeting, it becomes a negotiation about one person's credibility, it will be won by seniority, and it will not survive contact with a determined sponsor. If it is a published table applied to everyone before anyone walks in, the conversation is about the factor rather than the person, and the factor is a matter of evidence. Publish the factors. Republish them annually with the data behind them.
4. Refuse to let a case count value that depends on a project you have not approved. Enablement value is optionality, and it should be priced as such: at a fraction, or by bundling the enabled work into the same case so the whole chain is approved and sequenced together, or not at all. The test is a single question, and it is worth asking out loud in the room: if we approve this and nothing else, what does it return? For a large number of platform business cases the honest answer is nothing, and that answer changes where the work belongs in the queue.
5. Close the loop at the first measurable increment, not at the end of the benefits period. A 24 month feedback loop is not a feedback loop, it is an archive. Structure work so that something checkable arrives within a quarter, then check it, and let the result adjust both the project and your factors. This is the value-side reason for smaller projects, and it interacts directly with the absorption constraint: value that has been delivered but not yet absorbed by the business has not been realized either, and if you are measuring realisation you will find that out while there is still time to act on it.
6. Re-rank on the calibrated numbers, on a cadence. Calibration only pays if the ranking can move in response to it. If the portfolio is fixed for twelve months at a time, the best you can do is get one decision right a year, which is the problem described in The Frozen Portfolio. Quarterly re-ranking against calibrated value, with the constraint's capacity as the hard limit, is where this stops being an analysis exercise and starts being money. It is also, in practice, the thing that finally makes "everything is priority one" resolvable, because a comparison between two numbers derived the same way is a comparison people can actually have.
The objection you are already forming
"You are proposing an arbitrary haircut on exactly the work that matters most. Some of the most important things we do are genuinely hard to measure: regulatory compliance, resilience, security, the platform investments everything else depends on. Your scheme starves all of them and hands the portfolio to whoever is best at building a spreadsheet."
Three answers, and the third is the one that matters.
First, on arbitrariness. You are applying a factor today. It is 1.0, you apply it to every claim regardless of type, and the evidence says it is wrong by a wide margin and wrong differentially. A factor derived from thirty of your own closed projects is not arbitrary compared with that. It is the first non-arbitrary thing that will have happened to your numerator. If you dislike my illustrative percentages, good: they are placeholders for yours, and the entire point of the exercise is to replace them.
Second, most work described as unmeasurable is not. Regulatory compliance where non-delivery carries a defined penalty, a licence condition or a trading restriction is one of the most precisely measurable benefits in the portfolio, and it belongs in the cash tier, not the strategic one. Resilience is measurable the moment you are willing to state an outage probability and a cost per hour of outage, and if you are not willing to state those, notice that you have just declined to do the analysis that would have justified the investment. Much of what sits in the strategic tier is there through habit rather than necessity, and rule 2 above exists precisely to let it climb out.
Third, and this is the real answer: work that is genuinely important and genuinely unmeasurable should be funded, and it should be funded by a named allocation rather than by winning a comparison it cannot honestly enter. Ring-fence a share of constraint capacity, twenty per cent, whatever the number ought to be, for mandatory, regulatory and existential work. Decide that share as a deliberate policy question at portfolio level, in the open, and let everything else compete on calibrated value. This is better for the unmeasurable work, not worse. Today it is smuggled in via inflated benefit claims, which means its funding depends on nobody looking too closely and collapses the moment a rigorous chief financial officer arrives. Under an allocation it is protected explicitly, on the honest grounds that it is necessary, rather than the dishonest grounds that it returns $7,000 per constraint-hour. The alternative to a haircut is not more funding for important work. It is important work left permanently dependent on a number it cannot defend.
The uncomfortable part
Every prioritization method in this corpus, including the ones I have argued for hardest, is a ratio. Value per constrained resource hour is a ratio. Return on remaining investment is a ratio. Drag cost is a rate, and the shape it takes over time is a claim about future value. Every one of them has a forecast benefit somewhere inside it, and every one of them is only as good as that forecast.
Which means the highest-leverage improvement available to most portfolios is not a better method. It is a calibrated numerator, and it is available to any organization willing to spend six weeks looking backwards at what it promised and what it actually got.
Here is the test, and it takes a minute. Pick the project currently consuming the largest share of your constrained resource. Ask what benefit it claimed, what metric would show whether that benefit arrived, who is going to look at that metric, and when. If any of those four questions has no answer, that project did not win your scarcest capacity on merit. It won because the competition was checkable and it was not.

Leave a Reply