Prediction and Moderation: Two Different Tests for Hiring AI
Prediction asks who is most likely to reach a defined outcome. Moderation asks whether a condition changes that outcome differently across groups. A hiring model can be tested on both questions, but the tests require different evidence. A classification metric such as AUC cannot establish a moderation effect by itself.
Source: "Decision Traces," Saad Bin Shafiq, Nodes, 2026. Deployment at a Fortune 500 insurance carrier, N=10,765 agents. Read it on arXiv.
The standard test and its limit
The standard evaluation asks whether a model ranks people who later reach a defined outcome above those who do not. AUC measures that ranking. The conditional effect of onboarding, lead assignment, or another operating condition requires a separate test.
What a moderation test requires
A moderation analysis needs a defined condition, a defined outcome, and a prespecified interaction between the condition and the model signal. It also needs enough observations in each comparison group and a clear account of confounding factors. Without those pieces, a pattern should stay a research question rather than become a product claim. See the speed findings.
Why the distinction matters
Prediction and moderation support different decisions. Prediction can inform who is likely to reach an outcome. A validated moderation result could inform which operating condition to test for a defined group. Treating one measure as the other can turn a useful research question into an unsupported intervention.
The idea travels beyond hiring
The same distinction applies wherever an outcome may depend on both a person and a condition that can change. A decision trace can assemble the inputs, condition, and later outcome needed to test the interaction. An appropriate study design is still required. See decision traces.
What this does not say
The public evidence supports using moderation as an evaluation question. A settled public effect size for the behavioral-score interaction remains unsupported. Any such result would need a defined cohort, sufficient group sizes, uncertainty estimates, and validation beyond one carrier before it could guide an intervention.
Frequently asked questions
What is the difference between a predictor and a moderator in hiring? A predictor estimates who will reach an outcome. A moderator measures whether a defined condition changes that outcome differently across groups.
Is a low AUC a problem for a hiring model? AUC measures ranking for a defined outcome. Whether a result is useful depends on the study population, uncertainty, threshold policy, and the decision the model is meant to support.
How should hiring AI be evaluated? Ask both questions: does it predict outcomes, and does it identify who benefits most from favorable conditions. Evaluating only the first can undervalue a useful system.
Does this apply outside hiring? The reasoning extends to any decision system where individuals interact with conditions you control, including lending and clinical decisions.
Related reading
- From AUC 0.647 to 0.735: how data fusion improves prediction
- The economics of speed to production
- Decision traces, explained
See what your own data predicts, and who it says to invest in. Request a Decision Replay.