From AUC 0.647 to 0.735: How Multi-System Data Fusion Improves Hiring Prediction
The broader historical study covered 10,765 agents at one insurance carrier. The AUC comparison came from the evaluable n=229 subset with personality data: personality assessment reached 0.647, fusion with behavioral and applicant tracking data reached 0.735, and keywords reached 0.558. These are ranked-separation results within that small research subset, not production accuracy or a rule for deciding about any person.
Source: "Decision Traces," Saad Bin Shafiq, Nodes, 2026. Model comparison on the evaluable n=229 subset with Predictive Index data at one carrier. Read it on arXiv.
What the model compared
The research compared regularized logistic regression models built from each system's features and their combinations. Cross-validation kept model fitting and evaluation separate.
| Model | Features | AUC |
|---|---|---|
| ATS keywords | keywords plus source channel | 0.558 |
| Personality type only | PI type | 0.647 |
| Full fusion | all features plus PI | 0.735 |
All three AUC values above come from the same evaluable n=229 subset at one carrier.
Personality was the strongest single signal
Predictive Index type reached an AUC of 0.647 in the evaluable n=229 subset, above the 0.558 result for ATS keywords. The comparison describes ranking performance in this small one-carrier dataset. Using personality type as a standalone hiring rule remains unsupported.
Read the 0.735 correctly
The 0.735 figure is cross-validated on the evaluable n=229 subset that had personality data at one carrier. It says nothing about expected performance at another company. AUC expresses ranking rather than the percentage of decisions the model gets right.
What the result supports
The comparison supports one narrow finding: joining signals from multiple systems improved ranking performance within the evaluable research sample. Causality, transfer to another role, and the value of a particular operating intervention remain separate questions that require analysis on observed outcomes. See the speed-to-production findings.
What this means
The signals were stored in different systems. Connecting them made the comparison possible and exposed whether each signal added ranking information within the study cohort. A decision trace keeps that result attached to its inputs, outcome, population, and limitations. See how.
Frequently asked questions
What is a good AUC for a hiring model? Context matters more than a single threshold. In the evaluable n=229 subset at one carrier, keyword screening reached 0.558, personality type 0.647, and full multi-system fusion 0.735.
Does AUC 0.735 mean the system is 73.5% accurate? No. AUC is a ranking measure rather than an accuracy percentage, and 0.735 was measured on the evaluable n=229 subset with personality data at one carrier.
Can the research AUC be used as a production benchmark? No. It was measured on an evaluable n=229 research subset at one carrier. Another organization must evaluate the model on its own population and defined outcome.
What predicted production best? Personality type was the strongest single predictor in this research sample. Fusing it with behavioral and ATS features ranked outcomes better than any single-system model in the same sample.
Related reading
- The economics of speed to production
- What 10,765 agents hired revealed about resume keywords
- Decision traces, explained
See what fusing your own systems would surface. Request a Decision Replay.