Incremental validity is evidence that a personality measure improves prediction or decision-making beyond information already available. A study usually compares a baseline model with a second model that adds the personality score, then examines the change in prediction, often reported as ΔR², a change in explained variance. The claim is worth making only when the baseline, outcome, sample, analysis, and intended use are clearly specified. A larger statistical fit is not automatically a meaningful improvement, and one positive result does not prove that a report is accurate for every person or purpose. A responsible report should name what was added, what it improved, for whom, under which conditions, and with what uncertainty.
The claim under review: “This report adds predictive value”
The phrase sounds stronger than it is. Incremental validity does not mean that a personality report is valid in the abstract, nor that its score reveals a hidden truth about a person. It addresses a comparative question: after a defined source of information is already used, does adding this measure improve a defined interpretation or prediction?
For example, a report might claim that a conscientiousness scale improves prediction of a particular work outcome beyond a structured interview. That is a narrower statement than “this personality test predicts job performance.” It also differs from a self-reflection report that helps a reader notice recurring preferences. The first claim needs evidence about an external outcome and a comparison procedure. The second may need evidence about score meaning, consistency, and useful interpretation, but it should not borrow the language of selection research without support.
This distinction matters because a report can quietly move from “the score added a small amount of prediction in one study” to “the score tells you how you will behave.” That second conclusion is not contained in the first. Personality measures describe tendencies under specified measurement conditions. They do not remove context, learning, incentives, relationships, or chance from behavior.
Why the idea sounds plausible
A new measure can contain information that overlaps with an existing measure and still add something distinct. Imagine a hiring study in which the baseline model uses a structured interview and prior experience. Researchers then add a personality scale intended to capture a different aspect of work behavior. If the expanded model predicts the prespecified outcome better, the scale may have incremental validity for that outcome and comparison.
The statistical picture is often expressed with two models. Model one uses the baseline predictors. Model two uses those same predictors plus the new personality score. The difference between their R² values is ΔR², pronounced “delta R squared.” R² is the proportion of variation in the outcome accounted for by the model in that sample. ΔR² therefore describes the additional in-sample variation associated with the added predictor, not a percentage of a person’s character that the test has discovered.
The comparison can be sensible when the added construct has a reason to matter. A meta-analysis by Dudley and colleagues found that narrow conscientiousness traits could add prediction beyond broad conscientiousness, but the amount depended on the performance criterion and occupation. That result illustrates both sides of the claim: narrower information may add detail, yet its value is conditional rather than universal.
What a proper incremental-validity study must specify
Start with the baseline. “Beyond existing information” is incomplete unless the report names what that information was. Age and education, a prior score, a structured interview, a job sample, another personality scale, and a clinician’s judgment are not interchangeable baselines. Changing the baseline changes the question and can change the result.
Next, identify the outcome. The added score might be evaluated against a later performance measure, a training result, a retention outcome, a coaching goal, or another assessment. Those outcomes require different interpretations. A result for one criterion should not be presented as evidence for every possible use. The Standards for Educational and Psychological Testing frame validity as support for interpretations of scores for proposed uses, not as a permanent property of a test detached from context.
The study should also describe the population and setting, the timing of measurement, how the criterion was obtained, the planned order of predictors, and how missing data and overlapping predictors were handled. Hunsley and Meyer’s review specifically identified predictor-entry order, the size of the validity increment, and artifacts in the criterion as issues that can affect conclusions. A report that offers only a bare phrase such as “demonstrated incremental validity” leaves the reader unable to audit the claim.
Finally, distinguish prediction from explanation. If adding a score improves a statistical prediction, that does not establish that the measured trait caused the outcome. Nor does it prove that the score will improve a real decision after administration costs, fairness concerns, and possible misuse are considered.

A worked example: what the result would and would not say
Consider an example. A researcher wants to know whether a short personality scale adds information beyond a structured interview when predicting a defined six-month work outcome. The baseline model uses the interview score. The expanded model adds the personality score. The study reports that the expanded model has a higher R² and that the change is statistically distinguishable from zero.
The defensible reading is: in this sample, for this outcome, the personality score contributed information beyond the interview score under the stated analysis. The reader should then ask how large the improvement was, how precisely it was estimated, and whether the result was tested in new data. A statistically detectable change can be too small to alter a decision. It can also be larger in the study sample than it will be in a new sample.
The indefensible readings are broader. The result does not show that the personality score is more accurate than the interview. It does not show that the score determines the outcome, that every facet is useful, or that a high or low score is good. It does not justify using the report for diagnosis. If the assessment is being considered for employment, the evidence also needs to fit that use, population, criterion, and the consequences of acting on the score.
A recent open-access analysis makes the practical point especially clearly: an out-of-sample approach can show what prediction improves on new data, while an in-sample ΔR² can make the gain look more impressive. The authors argue that the improvement should be weighed against administration, maintenance, and test-taker costs. That is the difference between a statistical increment and a useful assessment contribution.
Why a positive result can still be weak evidence
One problem is overfitting. A model can match quirks in the data used to build it and then perform less well on new cases. Out-of-sample evaluation, such as a properly separated test set or cross-validation, asks whether the gain survives data that did not determine the model. It does not solve every validity problem, but it is a useful check on whether the reported improvement travels beyond the original sample.
A second problem is measurement error. Personality scores are estimates, not perfectly observed quantities. Westfall and colleagues showed that common “after controlling for” arguments can produce misleading incremental-construct-validity conclusions when measurement error in predictors is ignored. A report should therefore avoid treating a regression coefficient as a clean measure of a person’s unique psychological contribution.
A third problem is redundancy. If the new score and the baseline measure largely capture the same information, the added contribution may be small or unstable. A small increment is not automatically worthless, especially when a decision affects many people or when the added information is inexpensive and fair. But it needs a practical justification. Conversely, a statistically larger increment can still be unsuitable if the criterion is weak, the sample is narrow, or the use creates unreasonable consequences.
There is also a reporting problem. Selective presentation can make one favorable result stand in for a mixed evidence base. Stronger documentation reports the prespecified outcome, comparison model, uncertainty, validation sample, and boundaries. It also says what the study did not test.

Should a personality report claim incremental validity?
Yes, but only as a bounded evidence claim. A report may use the term when it can identify the added measure, the baseline, the outcome, the relevant population, and the analysis that supports the comparison. It should state whether the evidence is from the same sample or new data, and it should describe the practical meaning of the improvement rather than relying on significance alone.
For a personal report used in self-reflection, the phrase may be unnecessary unless the report is explicitly comparing its score with another source of information. A reader usually needs clearer answers: What does this score represent? What norm group or reference frame is being used? How precise is the score? What should I do with the information? Adding a technical label without a decision it improves can make a report sound rigorous while making it less understandable.
For coaching or development, incremental evidence may support combining a personality measure with a stated goal, behavioral observations, or another relevant source. The score should prompt questions and experiments, not dictate a fixed identity. For employment selection, the burden is higher because the score can affect opportunity. Evidence should match the job-related interpretation and be reviewed for fairness and unintended consequences. For clinical decisions, a general personality report should not be presented as a diagnostic instrument unless the named instrument and intended clinical use have their own appropriate documentation.
The narrow conclusion is simple: claim incremental validity when the added value has been demonstrated for a specified comparison and use. Otherwise, describe the evidence more modestly and explain what remains unknown.
A reader’s checklist before trusting the claim
Use this short audit when a report says that its personality measure adds predictive value. First, underline the outcome. Is it a clearly defined future or external criterion, or is the report using “prediction” as a vague synonym for an appealing description? Second, identify the baseline. What information was already available, and why was that the comparison?
Then check the evidence. Does the report name the sample and setting? Does it show the size and uncertainty of the added contribution? Was the result evaluated in new data, or only in the data used to fit the model? Were the scores reliable enough for the claimed use, and were the predictors so similar that the increment is hard to interpret?
Finally, check the decision. What action is the report asking you to take? Is the added information worth the time, money, privacy cost, and possible harm? Would you reach the same conclusion if the score moved modestly because of measurement error? Does the report clearly state that the result describes a tendency rather than a diagnosis or destiny?
A practical next step is to write one sentence in your own words: “This measure adds [what kind of information] beyond [which baseline] for [which outcome] in [which population], with [what uncertainty].” If the report cannot support that sentence, treat incremental validity as an unsupported headline rather than a settled conclusion. For more guidance on score interpretation, norms, and responsible use, continue with the live topics library.
Sources and notes
- Reporting Scores and Other Results, Oxford Research Encyclopedia of Education
Supports the use-specific definition of validity and the connection between clear score reporting and valid interpretations.
- Standards for Educational and Psychological Testing: Essential Guidance and Key Developments
Supports the distinction between score interpretations and uses, the need for evidence matched to purpose, and the continuing nature of validation.
- The incremental validity of psychological testing and assessment: conceptual, methodological, and statistical issues
Supports the cautions about predictor-entry order, increment size, criterion artifacts, and design issues in incremental-validity research.
- An out-of-sample perspective on the assessment of incremental predictive validity
Supports evaluating incremental prediction on new data and weighing prediction gains against practical assessment costs.
- Statistically Controlling for Confounding Constructs Is Harder than You Think
Supports the warning that regression-based incremental claims can be distorted when measurement error in predictors is ignored.
- A meta-analytic investigation of conscientiousness in the prediction of job performance
Supports the example that narrow personality traits can add prediction beyond a broad trait depending on criterion and occupation.
- Principles for the Validation and Use of Personnel Selection Procedures
Supports matching selection validation evidence to the intended personnel use and considering the consequences of selection procedures.
Apply it to your work
Turn ‘that job was not for me’ into something more useful.
From this guide: Ask whether the report supplies enough detail to write a bounded sentence about what was added, beyond what, for whom, for which outcome, and with what uncertainty.
The Work Pattern Report can help you separate repeated preferences from one difficult environment by mapping ten work continuums and their intersections. Compare the pattern with the role’s pace, planning, feedback, conflict, ownership, and change demands without reducing the experience to personality alone.
