The hard half of a pre-priced proposal
Cost of action is observed. Cost of inaction is inferred, and an inferred number can be gamed as easily as it can be computed.

A defensible cost of inaction AI ROI figure comes from the same production baseline that generated the recommendation, not an imported industry benchmark. It names a real cohort, a real time window, and a documented trajectory a reviewer can check. Without that baseline, an inaction number is a guess wearing the format of a calculation.
A pre-priced proposal carries two numbers: what the action costs and what waiting costs. Only one of them is observed. The cost of action is the effort a workflow requires, the systems it touches, the ramp on the change itself, and every input for that number sits inside the workflow specification the system already produced. The cost of inaction is a different kind of number. Nobody observes what did not happen. Every inaction figure is an estimate of the trajectory the world would have followed without the action, and an estimate is the easiest kind of number to inflate, borrow from someone else's benchmark, or round toward whatever makes a proposal look more urgent than the evidence supports. The proposal that arrives pre-priced argued that both numbers belong on the proposal before a human reads it. This piece is about the one that is genuinely difficult to get right, and what separates a defensible version of it from a guess with a dollar sign in front.
Two numbers, two kinds of evidence
The cost of action answers a question with a knowable boundary: what does this workflow require. Systems touched, approvals needed, time to implement. A reviewer can check the estimate against the workflow specification itself, because the specification is the source of the number.
The cost of inaction answers a different question: what happens if we wait. That question has no boundary the same way. Waiting a day, a month, or a quarter produces a different answer each time, and the answer depends on an assumption about what the current trajectory looks like without intervention. That assumption is a modeling choice; nothing about it was observed directly. Two systems reasoning about the identical decision can produce two different inaction numbers, both internally consistent, because each one is built on a different guess about the baseline.
This is not a reason to drop the number. The status quo has a price is right that most organizations have never run the calculation and are worse off for it. It is a reason to be exact about what makes one version of the calculation trustworthy and another version theater. A proposal that states a cost of inaction without showing its baseline asks the reviewer to trust an assumption they cannot see. That is not a smaller ask than trusting the recommendation itself. It may be a larger one, because the recommendation at least names the action being taken.
What makes a baseline defensible
One entry condition and two real tests separate a defensible cost of inaction estimate from an invented one.
The first is source. The proposal that arrives pre-priced already established that the inaction figure should come from the same reasoning that produced the recommendation rather than a fresh calculation bolted on afterward. That is the entry condition, not the whole test. Passing it rules out the laziest failure, a number invented on the spot, without ruling out the more common one: a number pulled from the right system but built on a comparison that was never checked for legitimacy in the first place.
That is what the next two tests check.
The first test is a real comparison cohort with a stated time window, and honesty about what kind of comparison it is. Faster than average is not a baseline; it is a phrase. A defensible baseline names the group being compared, the period measured, the milestone used to mark completion, and whether the comparison is causal or observational. The anchor pilot at a Fortune 500 insurance carrier names a 47-day ramp gap between cohorts, labeled plainly as a historical-control comparison rather than a controlled experiment, multiplied by a per-person-per-day production value of $54.35 that did clear bootstrap, trimming, and further checks before it entered the approved evidence base (methodology: Decision Traces). The dollar figure earned its statistical validation. The cohort gap earned an honest label instead. Both belong in the estimate. Neither gets dressed up as the other.
The second test is that the estimate has to survive contact with new data. A number computed once and left alone drifts from whatever it originally measured. The same production system that generated the baseline keeps reading new outcomes, and a defensible inaction figure gets checked against them rather than frozen the day it was first calculated. If the gap between the current trajectory and the calibrated one changes, the number the proposal quotes has to change with it. A vendor benchmark imported once at contract signing and never revisited fails this test by construction. It was never connected to the buyer's own data in the first place.
None of this requires sophistication a reviewer cannot follow. It requires naming what would have to be true for the number to be wrong, and checking that nobody skipped that step, the same discipline Decision Traces documents for a claim before it enters an approved evidence base.
The guess that looks like a calculation
A cost of inaction figure without a visible baseline is indistinguishable, on the page, from one with a rigorous baseline behind it. Both arrive as a dollar amount attached to a proposal. Both look precise. That is the actual danger, and it is worse than presenting no number at all.
A reviewer facing a proposal with no financial context knows they are making a judgment call and can weigh it accordingly. A reviewer facing a proposal with a confident, specific inaction figure reasonably assumes the number was earned, because a specific figure reads as evidence rather than opinion. If the number was pulled from an unrelated benchmark or padded to clear an approval threshold, the reviewer has been handed false precision instead of information. The approval that follows is less accountable than a plain judgment call would have been, because it was made on a number that only looked like it had been checked.
This is the mechanism worth naming directly: precision is not evidence. A figure with two decimal places is not automatically more defensible than a rounded one. What makes a number defensible is whether a reviewer can trace it back to a cohort, a time window, and a data source they could ask about on a call, and reject the proposal if that trail is missing. A workflow proposal that shows its baseline invites that question. One that shows only the answer forecloses it.
Consider two versions of the same proposal side by side. Version one states a dollar figure for the cost of waiting and nothing else. Version two states the identical figure, then adds the cohort it came from, the ramp-days gap that produced it, and the date the underlying data was last refreshed. Either figure could be accurate. Only one of them gives the reviewer anything to check. A finance committee that has learned to ask for the second version stops approving proposals on faith and starts approving them on a record that survives an audit later, when someone asks why the pilot was funded and whether the number held up.
What this looked like in production
The $54.35 daily production constant and the 47-day ramp gap were not numbers a vendor supplied. Both came from the carrier's own system of record, over four years and 10,765 agents of production data. The dollar constant cleared bootstrap, trimming, and further statistical checks; the ramp gap is a labeled historical-control comparison across cohorts rather than a causal result, and the carrier states that plainly rather than dressing it up. That distinction, one figure statistically validated, the other honestly labeled instead of dressed up as something it is not, is what a defensible estimate looks like once it leaves the page and meets a reviewer's questions.
Buyers evaluating any AI workflow proposal, from any vendor, can ask for the same trail before accepting a cost of inaction figure: which cohort, which time window, which system of record, and whether the number gets rechecked against new outcomes. The hiring ROI calculator runs that math against a buyer's own inputs, which states the same requirement a different way. A number built from someone else's data is a benchmark. A number built from the buyer's own production data, with the baseline visible, is a calculation.
The cost of action was never the hard half of a pre-priced proposal. It was always visible in the workflow specification, available for anyone to check against what the system is asking to do. The cost of inaction is the half that requires an organization to show its work, and showing the work is a discipline, not a formula: pick the comparison honestly, write down where it came from, and keep checking it against what actually happens. A proposal that skips that step and still prints a confident number has not made the approval easier. It has made the approval harder to trust, dressed as if it were easier to make.
A pre-priced proposal is a claim about the system that produced it. If the cost of inaction figure cannot survive a reviewer asking where it came from, the recommendation attached to it deserves the same scrutiny. The number that holds is the one built from a baseline the buyer can see.
Sources
Saad Bin Shafiq is the founder of Nodes, serving data-sensitive enterprises. Methodology: Decision Traces.