Criterion contamination happens when the outcome used to judge a personality assessment is influenced by factors that do not belong to the outcome being measured. In a validation study, the outcome is often called the criterion: for example, a defined set of work behaviors or a training result. If a supervisor’s rating is shaped by knowing an employee’s personality score, by unequal resources, or by a general impression that spills across rating categories, the rating may contain extra variance. A correlation with the personality score can then reflect the rating process as well as the person’s actual behavior. This does not mean that a report containing criterion-related evidence is automatically untrustworthy. It means you should ask what the criterion was, how it was collected, how reliable it was, and what irrelevant influences could have entered it. The practical comparison is between evidence tied to the intended outcome and evidence that quietly measures the outcome plus the evaluator, setting, or method. That distinction matters most when a report moves from describing a tendency to making a prediction about work, training, or another consequential result.
Start with the outcome, not the personality score
Imagine reading that an assessment score is related to performance. Before asking whether the relationship is large, ask what performance means in that sentence. A criterion is the outcome used as the reference point in a validation argument. It might be a defined work behavior, a training result, a turnover event, or another outcome that the proposed use actually cares about. It is not simply whatever information happens to be available.
Criterion contamination means that this outcome measure contains systematic influence from outside the intended outcome. The National Research Council describes contamination as including performance aspects that are not part of the job, or effects from factors unrelated to the criterion construct. The SIOP principles similarly say that a criterion is contaminated to the extent that it includes extraneous, systematic variance. Systematic matters here: this is not just random noise. It is a patterned influence that can push scores in a particular direction.
For a personality assessment, the predictor is usually a score or set of scores obtained from the instrument. The criterion is the later or related outcome used to examine the score’s meaning. If a report says that a trait score predicts teamwork, the reader needs to know whether teamwork was represented by observed coordination, a structured rating, a self-report, a manager’s global impression, or something else. Each choice answers a somewhat different question.
The two models that are easy to confuse
A useful way to read the evidence is to compare two models. In the first, the criterion is a reasonably clean approximation of the outcome: the assessment is scored in a standardized way, the outcome is defined in advance, and people measuring it have limited access to predictor scores. In the second, the criterion combines the intended outcome with an irrelevant influence: a rater knows the personality result, one group has better equipment, one shift has a different workload, or a pleasant general impression affects every rating. The second model can still produce an orderly table of correlations. Orderly numbers do not make the criterion clean.
Criterion deficiency is the neighboring problem. It occurs when the criterion leaves out relevant parts of the outcome. A narrow productivity count might omit safety, quality, or important collaborative work. Contamination adds something that should not be there; deficiency leaves something important out. A measure can have both problems at once. For example, a single supervisor rating might omit several important parts of performance and also be influenced by the supervisor’s knowledge of the employee’s assessment score.
The distinction changes the question you ask of a report. For contamination, ask, ‘What could have pushed the outcome measure for reasons unrelated to the outcome?’ For deficiency, ask, ‘Which important parts of the outcome were never measured?’ Neither question can be answered from a reliability coefficient alone. Reliability concerns consistency under specified conditions. It does not establish that the criterion represents the right construct or is free from irrelevant influence.
A worked comparison: one rating, two interpretations
Consider an unnamed employer testing whether a personality assessment is useful for a customer-support role. The proposed outcome is work performance. In a careful design, the employer first defines relevant behaviors, such as accurate handling of requests, appropriate escalation, and clear written communication. Ratings or records are collected using consistent instructions. People providing ratings do not receive the applicant’s assessment score, or the design otherwise limits its influence. The study still has uncertainty, but the evidence is aimed at the stated outcome.
Now change one feature. Supervisors receive the assessment reports before rating performance and are told that the assessment is intended to identify strong customer-support candidates. Their ratings may be influenced by what they expect to see. This is not proof that every rating is biased, and it does not prove that the assessment has no relationship to performance. It is a reason to treat the observed relationship as ambiguous: part of it may reflect work behavior, and part may reflect the rating process.
Change another feature instead. The workplace assigns some employees more difficult customer queues and gives others more experienced backup. Those conditions may affect the performance measure even if the raters never see assessment scores. The SIOP principles list differences in machinery quality, sales territories, job tenure, shifts, locations, and rater attitudes as possible contaminating factors. The exact factor will vary by setting, but the reading habit is general: inspect the conditions around the criterion, not only the instrument that produced the personality score.
This comparison also shows why a report should identify the unit of inference. Evidence that a score is associated with a supervisor rating in one organization does not by itself establish how a person will behave in every setting. The assessment may capture a tendency that appears differently under different demands, resources, roles, or observers.
Why contamination can change the apparent validity
Validity is claim-specific evidence supporting an interpretation and use of scores. A relationship with a criterion can support a narrower claim, such as an association with a particular rating under particular conditions. It does not automatically support a broad claim that the score reveals a fixed personal quality or determines future performance. The APA testing standards say that criterion variables should be described in terms of their reliability, their representation of the intended construct, and their possible exposure to extraneous sources of variance.
Contamination can make a relationship look stronger when the irrelevant influence is related to the predictor. Suppose a rater expects people with a certain assessment profile to be organized and therefore rates those people more favorably on several loosely defined dimensions. The association may be partly created by the rater’s expectation. Contamination can also hide a relationship when an irrelevant condition works against it. If one group receives poorer tools or harder assignments, its criterion scores may be lower for reasons not represented by personality.
The problem is especially important when several variables share the same method. A personality questionnaire and an outcome questionnaire completed by the same person at the same time can share response tendencies, mood, and wording effects. That does not make self-report evidence useless. It means the result should be interpreted as evidence from that method, not as a direct behavioral observation. In a peer-reviewed study of a five-factor inventory, evaluatively neutralized items showed similar criterion validity across three sets of self-rated behaviors while showing weaker relations to the social desirability of the criteria. The finding illustrates a design question: how were both the predictor and criterion worded and rated? It does not establish a universal correction for every assessment.
A careful report therefore avoids treating one coefficient as a verdict. It describes the criterion, the sample and setting, the direction and uncertainty of the relationship, and the limits on generalizing it. The Standards note that summary associations should be supplemented with information about the form and variability of the relationship, rather than being treated as enough to connect a particular score with a particular outcome.

What to look for in a report’s evidence section
A report does not need to use the phrase criterion contamination to let you evaluate the risk. Look for the ingredients of a transparent validation description. First, what interpretation is being supported? Is the report describing a score, estimating a related behavior, screening for a purpose, or recommending a decision? Evidence must be judged against the proposed use.
Second, what exactly was the criterion? ‘Performance’ or ‘success’ is too broad on its own. Look for the observable behavior, record, rating scale, or outcome definition. Ask which relevant parts were included and which may have been missed. Third, who supplied the information, when, and under what conditions? A manager, peer, customer, trainer, and the test taker may see different slices of behavior. Timing can also matter when the intended outcome is a stable pattern versus a current state.
Fourth, could the criterion provider have known the personality result? If so, the report should discuss how that knowledge was controlled or evaluated. Also check for unequal settings, different opportunities, range restriction, and common measurement methods. These are not automatic disqualifiers. They are design details that change how confidently the relationship can be interpreted.
Finally, separate evidence about the instrument from evidence about the use. A test manual or study may support score reliability in one population, while a separate validation study addresses a particular criterion and decision. The APA guidelines for psychological assessment emphasize integrating different sources of validity evidence and recognizing the limits of any one source. A polished narrative cannot substitute for evidence matched to the actual interpretation.
A responsible decision rule for readers
Criterion contamination is a reason to narrow an interpretation, not a reason to discard every personality assessment. For self-reflection, you can treat a result as a prompt for checking patterns against your own observations. For coaching, use it alongside concrete examples and goals rather than as a stand-alone explanation. For workplace selection or other high-stakes decisions, require evidence tied to the specific role, population, criterion, and procedure, and involve people who understand assessment standards and applicable rules. A general personality report is not a clinical diagnosis and cannot settle a person’s suitability by itself.
The simplest comparison is this: a clean criterion asks whether the score relates to the outcome you meant to study; a contaminated criterion asks the score to relate to that outcome plus hidden features of the measurement situation. The first can still be imperfect. The second makes the meaning of the observed relationship harder to identify.
Before accepting a prediction in a personality report, use this short checklist: identify the outcome; ask who measured it and under what conditions; check whether relevant parts were omitted; ask whether raters knew the assessment result; look for more than one source of evidence; and match the conclusion to the decision you are actually making. If those details are missing, the result may still prompt a useful question, but it should not carry more certainty than the evidence can support.
Questions readers ask
Does criterion contamination mean a personality assessment is invalid?
No. It means the outcome used to evaluate the assessment may include systematic influences unrelated to the intended outcome. The assessment may still have useful evidence for a narrower interpretation, but the reported relationship should be read with the criterion’s design, limits, and proposed use in view.
Sources and notes
- Principles for the Validation and Use of Personnel Selection Procedures, Fifth Edition
Defines criterion contamination and deficiency in personnel validation, and identifies rater and work-condition influences that can contaminate criteria.
- Performance Assessment for the Workplace: Evaluating the Quality of Performance Measures
Explains contamination and deficiency as threats to the meaning of performance criteria and shows why construct-irrelevant factors matter.
- Standards for Educational and Psychological Testing
Requires criterion-related evidence to describe criterion reliability, construct representation, and possible extraneous sources of variance.
- APA Guidelines for Psychological Assessment and Evaluation
Explains that validity evidence must be integrated for a particular purpose and that conclusions should be fair and minimize bias.
- Criterion Validity is Maintained When Items Are Evaluatively Neutralized
Reports a peer-reviewed comparison of evaluative and neutralized personality-inventory items, including their relations with self-rated criteria and social desirability.
Apply it to your work
Understand how you work before you choose what comes next.
From this guide: Carry this report-reading question into the work decision in front of you.
Build a private Work Pattern Report across ten workplace continuums, then compare the result with the demands of the role or environment in front of you.
