Four years of outcomes exposed the cost of one untested filter.
One anonymized, NYSE-listed Fortune 500 insurance carrier tested its hiring assumptions against the production outcomes of 10,765 agents. The live Nodes deployment has scored 900,000+ candidates across 215+ locations.
Study period: 2022 to 2025 · Production boundary: one customer deployment · Results do not transfer without separate validation.
The filter had never been tested against production.
The carrier operated a high-volume hiring funnel across 215+ locations. Screening rules narrowed the population before anyone joined the hiring record to the production outcome.
The study connected applicant evidence, the decision, and the result that appeared after the hire. It then tested which familiar screening signals survived.
The familiar signals did not hold up.
Of 8,181 parsed skills, 3,597 had enough data to test. After Bonferroni correction, zero predicted the first production milestone and 30 were anti-predictive.
More approved evidence improved ranking.
Keyword screening alone produced AUC 0.558. Personality assessment alone produced AUC 0.647. Full ATS, assessment, and behavioral data fusion produced AUC 0.735 in the stated research sample.
The production milestone arrived 47 days earlier.
The reference cohorts moved from 109 days to 62 days to the first production milestone. The live requisition-to-hire loop moved from 127 days to 38 days.
One rule would have rejected 2,863 producing agents.
The historical replay identified producing agents the industry-experience filter would have removed. Their annual production represented $17.7M at risk in the retrospective counterfactual.
AUC describes ranking performance in the stated sample. The time and value findings come from one carrier. They do not establish universal accuracy, causality, or future lift.
The deployment moved from contract to production in 34 days.
Legal approval took 17 days after six AI hiring vendors had been rejected over 18 months on architecture. These are results from one deployment, not a timeline promise for another environment.
Evidence moved through a governed path before action.
The production program combined approved records, customer-specific calibration, named human authority, and a Decision Trace that preserved the recommendation and action.
Read approved records.
Production evidence stayed inside the approved customer VPC boundary.
Connect prior decisions to results.
The system evaluated historical evidence against the customer's defined production outcome.
Return the recommendation with evidence.
The named reviewer received the decision, limitations, and drafted next action.
Preserve the human decision and outcome.
The approved action and later result became part of the customer Decision Trace.
Observed, validated, and modeled claims stay distinct.
Each number retains its population, period, method, and limitations. The full evidence register remains available as a separate inspection surface.
| ID | Claim | Value | Population | Method | Status |
|---|---|---|---|---|---|
| C01 | Candidates scored | 900,000+ | One Fortune 500 carrier | Live production record since January 2025 | Observed |
| C05 | Research cohort | 10,765 agents | Agents at one carrier | Retrospective observational study, 2022 to 2025 | Observed |
| C20 | Q1 net savings | $1.58M | One live deployment | CFO-validated customer record | Observed |
| C21 | Annual production at risk | $17.7M | 2,863 producing agents | Retrospective counterfactual | Modeled |
| C25 | Time to first production milestone | 109 to 62 days | Reference carrier cohorts | Historical median comparison | Observed |
| C41 | Predictive keywords after correction | 0 of 3,597 | Testable keyword set | Bonferroni-corrected testing | Validated |
Methodology: Decision Traces, arXiv:2604.19819. Complete claim records and supporting artifacts are available through the evidence register and deployment data room.
“Absolutely love what you guys are doing.”
Talent acquisition leader · NYSE-listed Fortune 500 · 200+ locations · in production since January 2025
See what a different decision would have changed in your history.
A Decision Replay compares one named decision with outcomes already recorded in your systems. Past data only. No live decisions.