Case study · insurance talent · live since January 2025

Four years of outcomes exposed the cost of one untested filter.

One anonymized, NYSE-listed Fortune 500 insurance carrier tested its hiring assumptions against the production outcomes of 10,765 agents. The live Nodes deployment has scored 900,000+ candidates across 215+ locations.

Study period: 2022 to 2025 · Production boundary: one customer deployment · Results do not transfer without separate validation.

Study cohort10,765 agentsObserved
Live scale900,000+Observed
Deployment footprint215+ locationsObserved
Q1 net savings$1.58MObserved
The case

The filter had never been tested against production.

The carrier operated a high-volume hiring funnel across 215+ locations. Screening rules narrowed the population before anyone joined the hiring record to the production outcome.

The study connected applicant evidence, the decision, and the result that appeared after the hire. It then tested which familiar screening signals survived.

01 · Filter audit

The familiar signals did not hold up.

Of 8,181 parsed skills, 3,597 had enough data to test. After Bonferroni correction, zero predicted the first production milestone and 30 were anti-predictive.

02 · Ranking evidence

More approved evidence improved ranking.

Keyword screening alone produced AUC 0.558. Personality assessment alone produced AUC 0.647. Full ATS, assessment, and behavioral data fusion produced AUC 0.735 in the stated research sample.

03 · Time

The production milestone arrived 47 days earlier.

The reference cohorts moved from 109 days to 62 days to the first production milestone. The live requisition-to-hire loop moved from 127 days to 38 days.

04 · Value at risk

One rule would have rejected 2,863 producing agents.

The historical replay identified producing agents the industry-experience filter would have removed. Their annual production represented $17.7M at risk in the retrospective counterfactual.

AUC describes ranking performance in the stated sample. The time and value findings come from one carrier. They do not establish universal accuracy, causality, or future lift.

Observed production record

The deployment moved from contract to production in 34 days.

Legal approval took 17 days after six AI hiring vendors had been rejected over 18 months on architecture. These are results from one deployment, not a timeline promise for another environment.

17 days · legal approval · one customer record
34 days · contract to production · one deployment
Live since January 2025
Expanded two quarters ahead of schedule
Decision control

Evidence moved through a governed path before action.

The production program combined approved records, customer-specific calibration, named human authority, and a Decision Trace that preserved the recommendation and action.

01 · Ingest

Read approved records.

Production evidence stayed inside the approved customer VPC boundary.

02 · Test

Connect prior decisions to results.

The system evaluated historical evidence against the customer's defined production outcome.

03 · Propose

Return the recommendation with evidence.

The named reviewer received the decision, limitations, and drafted next action.

04 · Record

Preserve the human decision and outcome.

The approved action and later result became part of the customer Decision Trace.

Evidence record

Observed, validated, and modeled claims stay distinct.

Each number retains its population, period, method, and limitations. The full evidence register remains available as a separate inspection surface.

IDClaimValuePopulationMethodStatus
C01Candidates scored900,000+One Fortune 500 carrierLive production record since January 2025Observed
C05Research cohort10,765 agentsAgents at one carrierRetrospective observational study, 2022 to 2025Observed
C20Q1 net savings$1.58MOne live deploymentCFO-validated customer recordObserved
C21Annual production at risk$17.7M2,863 producing agentsRetrospective counterfactualModeled
C25Time to first production milestone109 to 62 daysReference carrier cohortsHistorical median comparisonObserved
C41Predictive keywords after correction0 of 3,597Testable keyword setBonferroni-corrected testingValidated

Methodology: Decision Traces, arXiv:2604.19819. Complete claim records and supporting artifacts are available through the evidence register and deployment data room.

Customer voice · on record

“Absolutely love what you guys are doing.”

Talent acquisition leader · NYSE-listed Fortune 500 · 200+ locations · in production since January 2025

One production reference
Insurance talent
Every additional decision type validated separately
Test one decision

See what a different decision would have changed in your history.

A Decision Replay compares one named decision with outcomes already recorded in your systems. Past data only. No live decisions.