A personality report’s job-performance claim needs evidence that the score is measured and interpreted as stated, relates to behaviors important in the job, and connects to a defensible performance outcome in a relevant population and use. Well-matched earlier evidence may support the claim; a broad validation label or reliability figure alone cannot. Treat missing links as uncertainty, not an individual forecast.
The short answer: follow the score to the job outcome
A personality report’s job-performance claim needs evidence for three links: the score is measured and interpreted as stated; that interpretation concerns behavior important in the job; and the score relates to a defensible performance outcome in a relevant population. The Office of Personnel Management’s Assessment Strategy guidance explains that inconsistent scores cannot be expected to make useful predictions. Its Assessment Glossary defines job analysis as a systematic examination of work and required competencies. Score consistency alone does not establish a job link, and a plausible trait description does not show that a named outcome follows. A new study at the exact workplace is not the only possible support. Earlier evidence may apply if the assessment, behavior, outcome, population, and intended use match closely enough, with differences made clear. Evidence sufficient for self-reflection does not automatically justify screening, ranking, or an employment decision. Here, “supported,” “provisional,” and “unsupported” describe the documented fit of the exact claim; they are practical labels, not test ratings or legal findings. Read the report sentence one link at a time. Where evidence is missing, treat the conclusion as a hypothesis to examine, not an individual forecast or job verdict.
Sources: Designing an Assessment Strategy — U.S. Office of Personnel Management; Assessment Glossary — U.S. Office of Personnel Management
What exactly is the report claiming?
Rewrite the report’s sentence as a bounded claim. Identify the score, interpretation, target behavior, performance outcome, studied population, and proposed use. These are separate steps. “This score reflects orderly planning” interprets a score. “People with higher scores received stronger planning ratings in this job” claims a group relationship. “This person will perform well here” predicts an individual; “hire this person” adds an action. Each step requires its own support. Map the claim: score and scoring rule → construct interpretation → target behavior → observed criterion → population and context → intended use. A criterion is the outcome chosen to represent performance for a research question. The U.S. Office of Personnel Management’s Assessment Glossary defines criterion-related validity as the degree to which assessment results predict or relate to an important criterion, such as job performance, training success, or productivity. Evidence about one outcome does not establish a relationship with every meaning of success. At the first link, record the assessment name and version, scale or subscale, and scoring rule. Similar trait labels do not guarantee that two measures ask the same questions or combine responses alike. Evidence for one version cannot simply be assigned to another. Check the step from score meaning to workplace behavior: a scale may describe a general tendency, while the report claims a specific action. State and support that bridge. For example, imagine a report saying: “This scale reflects orderly planning”; “in the studied group, higher scores were associated with stronger deadline-coordination ratings”; “therefore, this applicant will coordinate deadlines well.” The middle statement describes a group association with an outcome. The last predicts one person, potentially in a different setting. Support for the association alone does not prove that forecast. The Assessment Glossary distinguishes assessment relationships with work outcomes and defines predictive evidence around testing applicants before later performance is evaluated. This reveals a possible jump in timing or population, not whether evidence transfers. Underline the report’s exact sentence and mark each link it asserts. Ask what evidence supports that link, not whether a nearby paragraph calls the instrument “validated.” A tendency statement is an interpretation; a job relationship needs a named behavior, outcome, and studied group. An individual forecast or action recommendation is stronger still. Match the evidence to the sentence actually printed.
Sources: Assessment Glossary — U.S. Office of Personnel Management; APA Guidelines for Psychological Assessment and Evaluation
What does the score itself need to show?
A report should let a reader follow the route from answers to score: which version was administered, how responses were combined, and what the number is meant to represent. Reliability means consistency under specified conditions. The U.S. Office of Personnel Management’s “Designing an Assessment Strategy” says documentation should report reliability and how it was computed, and explains that inconsistent scores limit useful prediction. Validity asks a different question: whether evidence supports a particular interpretation of scores for a particular purpose. Using a score for private reflection and using it to claim something about performance in a named job are different inferences. The “Assessment Glossary” from the U.S. Office of Personnel Management connects validation to the inferences drawn from scores and their intended use. Reliability supports prediction only as background; consistency alone cannot show that a score relates to the job outcome named in a report. The displayed score matters too. A raw total is produced by a scoring rule; by itself, it does not show how the person compares with others. A percentile gives a position relative to a specified norm group. It is not the percentage of answers correct or the probability of succeeding at work. A band such as “low,” “average,” or “high” groups scores according to stated boundaries. These formats answer different questions. The report should name its format and, for comparisons, the group, sample date, and basis. A percentile describes rank in that group, not an absolute amount of a trait. Some forced-choice assessments ask respondents to choose between statements rather than rate each independently. Interpret the output according to its documented method; do not treat it as an ordinary raw total or percentile unless the report explains how that interpretation was established. Changes in items, translation, administration, or scoring can weaken transfer of earlier evidence. The report should identify the version and conditions covered by its evidence. Measurement error means uncertainty in an observed score. A single value is not perfectly exact. Documentation may report a standard error of measurement or an interval around the score; check its population, conditions, and interpretation. Needed precision depends on the score and purpose; no coefficient is a universal pass mark. A transparent, reasonably consistent score establishes only the first link in a job-performance argument. If the report omits its scoring method, comparison group, or relevant uncertainty, there is less basis to interpret the number and accept the later job claim without qualification. The next question is whether the score’s interpretation connects to behaviors that matter in the specific role.
Sources: Designing an Assessment Strategy — U.S. Office of Personnel Management; Assessment Glossary — U.S. Office of Personnel Management; APA Guidelines for Psychological Assessment and Evaluation
Which work behavior is the score supposed to matter for?
A job claim should identify observable, important work behaviors and show how job analysis connects them to the role. The U.S. Office of Personnel Management’s Assessment Glossary defines job analysis as examining job tasks and required competencies. It describes tasks as activities stated as observable actions, and competencies as measurable patterns of knowledge, skills, abilities, behaviors, and other characteristics needed at work. Ask: what must a person do, and which actions does the report say its score relates to? A job title alone cannot answer that. Roles with the same title may differ, while different titles may share important tasks. The connection should be documented, not supplied by an appealing story about a trait. Suppose a report says a score reflects careful planning and claims this matters for handoffs. That is plausible: the work may involve checking dependencies, recording changes, and alerting colleagues before deadlines. But plausibility does not establish that the score captures the tendency as described, that it appears in those actions, or that those actions matter in this job. Evidence must test the links between score, behavior, and performance outcome. Job analysis can show that a behavior matters to the role; it does not show that a personality score predicts it. “Content validity” has a narrower meaning than a report may suggest. The Assessment Glossary describes content evidence as a match, based on job analysis and expert judgment, between assessment items or tasks and work tasks or competencies. That reasoning can fit a work sample resembling an actual task. A personality questionnaire generally measures an inferred tendency; its items do not directly sample a person coordinating handoffs at work. The EEOC’s Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures explains, within a U.S. employment-selection framework, that traits are intangible and not directly observable, so content-validity reasoning alone does not establish a trait measure’s job relevance. The guidance distinguishes these from procedures that operationally measure observable work behaviors. Similar wording in a questionnaire and job description is not proof of prediction. Related jobs can share core behaviors; evidence need not start from a unique title. The EEOC guidance says job analysis should describe important work behaviors, their relative importance and difficulty, and associated tasks. Applying evidence across roles rests on documented behavioral similarities, not simply a shared label such as “leadership” or “service.” Ask: Which observable behavior is this scale expected to relate to? How was it shown to matter in this job? Does the report name that behavior, or jump from a broad trait label straight to “performance”? Naming a relevant behavior gives the claim a target for testing. Until the score-to-behavior relationship is supported, it remains a hypothesis, not evidence that the score predicts performance.
Sources: Assessment Glossary — U.S. Office of Personnel Management; Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures — U.S. Equal Employment Opportunity Commission
What counts as performance in the evidence?
The study should define the outcome it calls performance and explain why that outcome represents an important part of the particular job. A personality score might relate to a supervisor’s assessment of teamwork but show little relation to task accuracy, customer service, productivity, or training success. Those findings answer different questions. The U.S. Office of Personnel Management (OPM) Assessment Glossary describes criterion-related evidence as the relationship between an assessment and an important criterion, such as job performance, training success, or productivity. A report that says only “predicts job success” leaves readers unable to tell what was measured. A useful account puts the work behavior and outcome side by side. If a report links orderly planning to performance, did the study measure timely completion, accurate records, a supervisor’s broad rating, or something else? Then ask whether that outcome reflects the behavior in the target role. A narrow measure can fit a narrow claim; a broad rating may cover several contributions while obscuring which one relates to the score. OPM’s Personality Tests guidance notes that relevant personality factors depend on the job and discusses distinct criteria such as overall performance, customer service, and teamwork. Require a clear match between claim and outcome. Each outcome has its own limits. A supervisor rating depends on whether the supervisor observed enough work, knew the standards, and used the scale consistently. A count of completed cases may seem objective, yet it can reflect case difficulty, tools, or workload as well as the worker’s contribution. An error count may miss careful work that prevents errors upstream. A training result informs a job-performance claim only if the study shows that the measure represents relevant work and is scored appropriately. Inspect these limits; they do not automatically invalidate a measure. Check whether the outcome is incomplete or partly entangled with the assessment. An incomplete criterion leaves out relevant performance: a productivity count may omit quality. A contaminated criterion includes influences that should not be mistaken for the target behavior; for example, a supervisor who knows assessment scores could let those expectations shape ratings. Ask who recorded the outcome, what standards they used, whether they knew the scores, and whether multiple criteria were examined. Mixed results may be meaningful: a score could relate to one documented behavior and not another. Read that as a bounded association, not proof of overall job success. The OPM Assessment Glossary and Personality Tests guidance support treating criteria as specified work outcomes; neither makes unlike outcomes interchangeable. Judge the evidence against the reported result, not the vague label “performance.”
Sources: Assessment Glossary — U.S. Office of Personnel Management; Personality Tests — U.S. Office of Personnel Management; Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures — U.S. Equal Employment Opportunity Commission
What can broad personality research establish—and what can’t it?
Meta-analyses can show whether personality measures have been associated with specified work outcomes across studies and whether results vary by measure and outcome. They cannot establish that an unnamed report, its current version, or one person’s score predicts success in one particular job. The useful role of broad research is to establish plausibility, expose variation, and guide the next question. It is background for a specific claim, not a substitute for evidence that connects that score to that job.
The meta-analysis Personality and Job Performance: The Big Five Revisited examined explicit Big Five measures and considered both job performance and contextual performance, meaning contributions beyond core task execution. Its abstract reports that job-performance results broadly paralleled two earlier reviews, while contextual-performance relations were more complex. This is evidence against saying personality measures never relate to work outcomes, but it also warns against treating different outcomes as one general measure of “success.” The study covers a family of Big Five measures across prior research; it does not validate a different instrument or establish what a score means for a named role.
Personality Measures as Predictors of Job Performance: A Meta-Analytic Review examined 494 studies and found usable results from 97 independent samples, totaling 13,521 people. In that historical evidence base, corrected mean validity was higher in studies using confirmatory rather than exploratory strategies, and higher where job analysis explicitly guided measure selection. This gives readers a practical reason to ask whether a score was chosen because it matched documented job requirements, rather than because a broad trait label seemed relevant after results were known. The authors also noted weaknesses in reporting study characteristics, limiting transfer of these aggregate estimates to a current report or job.
These findings support a measured middle position. Blanket skepticism misses evidence that some personality measures have related to work outcomes and that job-focused measure selection can matter. Blanket confidence treats an average across historical studies, measures, jobs, and criteria as a result for every instrument and person. Meta-analysis combines findings, but its conclusion depends on what counted as a measure, which jobs were represented, and how performance was recorded.
The Office of Personnel Management’s Personality Tests guidance also notes that relevant personality factors depend on the job and discusses distinct criteria such as overall performance, customer service, and teamwork. Broad evidence can therefore serve as a starting map: it may suggest which score–behavior relationship deserves examination and which outcomes should remain distinct. It cannot bridge the remaining distance by itself. When a provider cites a review, ask for its measures, populations, outcomes, and limits, then ask how those findings support this report’s score-to-behavior-to-performance claim.
Sources: Personality and Job Performance: The Big Five Revisited; Personality Measures as Predictors of Job Performance: A Meta-Analytic Review; Personality Tests — U.S. Office of Personnel Management
Does the study population and timing match the claim?
A predictive study measures applicants before they begin the job and compares those scores with performance measured later. A concurrent study measures people already doing the job and examines how their current assessment scores relate to their current performance. The distinction is about when and on whom the evidence was collected; neither design earns trust automatically. The U.S. Office of Personnel Management’s Assessment Glossary defines these designs in those terms and says whether concurrent evidence can stand in for predictive evidence depends on the measure and how closely the employee sample resembles the applicant population. For a claim about applicants, predictive evidence has a direct timing advantage: the score exists before the later outcome, as it would in a forecast. But time order alone does not make the forecast dependable. A study can follow applicants into a changed role, use a weak performance measure, or include too few or unrepresentative participants to support the report’s broad wording. Ask who was tested, when scores and performance were recorded, and whether the role and outcome match the report’s claim. If duties or standards changed, the later measure may answer a different question. Concurrent evidence may be useful when waiting for new hires to accumulate performance records is impractical. Current employees can provide relevant information about a real workplace, but they are not automatically interchangeable with applicants. Hiring and retention may have already selected who remains; employees may have acquired experience or training; and the sampled team may differ from the people to whom the report will be applied. These are reasons to examine transfer, not reasons to discard every concurrent study. The EEOC-hosted Questions and Answers on the Uniform Guidelines likewise says differences between incumbent and applicant groups should be considered when concurrent evidence is used. That document is U.S. employment-selection guidance, and its own page explains that it does not have the force and effect of law. The useful comparison is therefore between the target group and the study group, not simply between “predictive” and “concurrent.” Check whether the assessment version, language, role setting, experience level, selection history, and performance standards resemble the intended use. Fairness evidence matters in transfer too, alongside the score, job, outcome, and sample. Ask which differences were examined and how they affect the conclusion. If the report only says that current employees were studied, the supported statement is that a relationship was examined in that employee sample, not that the score has been shown to forecast every applicant’s performance. Readers can then distinguish “studied on current employees” from “shown to forecast applicants” and ask what evidence justifies the bridge.
Sources: Assessment Glossary — U.S. Office of Personnel Management; Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures — U.S. Equal Employment Opportunity Commission
When can evidence from another job or sample transfer?
Evidence from another job or sample can support a claim when the transfer is demonstrated across the score, important work behaviors, performance outcome, population, and proposed use. A new local study is not automatically stronger: well-matched external evidence may be more informative than a weak local study, while a familiar job title alone does not establish relevance. The U.S. Equal Employment Opportunity Commission’s Questions and Answers on the Uniform Guidelines says a study conducted elsewhere may provide sufficient evidence for employee selection when the procedure was valid in its original use; job analyses show a close match in major work behaviors; fairness evidence is considered; and differences in standards, work methods, sample representativeness, and currency are addressed. The document describes guidance, not law. This is a U.S. framework, not a universal rule or proof about a particular report. Compare both kinds of evidence on the same dimensions: | Check | Local prospective study | External or incumbent study | | --- | --- | --- | | Score | Exact version, scoring, and interpretation? | Same measure and interpretation? | | Job | Important behaviors identified? | Job analyses show shared behaviors? | | Outcome | Later performance represents the role? | Comparable outcomes and standards? | | Sample | People resemble the target group? | Effects of experience, prior selection, or sampling considered? | | Use | Proposed decision and fairness addressed? | Use, fairness, and study currency fit? | A prospective study measures the score before later performance, matching the timing of a forecast. But it cannot fix a poor outcome measure, unrepresentative sample, or changed job. An incumbent study measures current employees. The U.S. Office of Personnel Management’s Assessment Glossary says applying such findings to applicants depends on the measure and similarity between employee and applicant groups. The EEOC guidance also flags differences arising from work experience, prior selection, and sample choice. “Studied on employees” describes evidence; it does not prove applicant prediction. Transfer is stronger when the same score interpretation is studied, job analyses support shared important behaviors, outcomes align, and differences in population, fairness, method, and currency are addressed. It is weaker when links depend on labels or assumptions, and unresolved when documentation is missing. These are descriptions of fit, not numerical grades. Ask the provider or employer to identify the studies and compare the target job with each study setting on these dimensions. External evidence can suffice when the match is clear. Otherwise, the conclusion should be narrower than “this score predicts performance in this job.”
Sources: Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures — U.S. Equal Employment Opportunity Commission; Assessment Glossary — U.S. Office of Personnel Management
Does the score add useful information beyond what is already known?
A score may relate to job performance and still add little to a decision that already uses relevant information. Incremental validity means the additional prediction a measure provides beyond a stated baseline or existing method. The U.S. Office of Personnel Management’s “Designing an Assessment Strategy” defines the term this way and notes that a measure with limited standalone prediction could contribute when combined with another measure. That possibility is a reason to ask for a comparison, not proof that this report contributes useful information. The comparison needs a credible baseline: information or a method already used for the same outcome and purpose. If a report says its score improves a forecast, ask what it was compared with, whether that comparison reflects the actual decision context, and whether both measures were evaluated against the same performance criterion. A weak or unrelated baseline may make an addition look more useful than it would beside relevant evidence. The OPM “Assessment Glossary” defines incremental validity as the amount of predictive validity one tool adds relative to another; it does not set a universal baseline for every job or report. Then ask what “adds” means in practice. A detectable statistical increase may be too small to change a decision meaningfully, while a modest independent relationship could contribute distinct information. These are interpretations to investigate, not automatic conclusions from a correlation. Evidence should state the outcome, baseline, enough to judge whether a difference matters. An association between a score and performance alone does not show added value over information already available. Ask: “For this score and version, what evidence shows that it improves prediction of this named job outcome beyond information already used, and how much does that improvement change the intended conclusion?” A clear answer identifies the baseline and outcome and explains the practical implication. Without those details, the score may be a provisional clue, but its added contribution remains unestablished.
Sources: Designing an Assessment Strategy — U.S. Office of Personnel Management; Assessment Glossary — U.S. Office of Personnel Management
What should I ask before relying on the report?
Ask the provider four linked questions: Which instrument version and score produced this statement? Which observable behaviors in this job is the score meant to relate to, and how were they identified as important? What performance outcome did the supporting study measure, and who took part? What makes that evidence applicable to this job and proposed use? The U.S. Office of Personnel Management’s Assessment Strategy and Assessment Glossary distinguish reliability, job analysis, study design, and intended use. A general validation label leaves the link unresolved. Classify the statement as supported when the documented score, behavior, outcome, population, and use fit its wording; provisional when a plausible link exists but a material match remains unclear; unsupported when a necessary link is absent. For personal reflection, compare a reported tendency with repeated examples from your work, noting the conditions. Consider whether workload, authority, resources, or team expectations could also explain the pattern. Such observations guide reflection but do not validate the assessment. /assessment can organize questions about decisions, feedback, conflict, and change. It is a low-stakes self-report without norms, cutoffs, selection scores, or job predictions. Use it to choose a question to investigate, not as a verdict about job suitability.
Sources: Designing an Assessment Strategy — U.S. Office of Personnel Management; Assessment Glossary — U.S. Office of Personnel Management
Questions readers ask
Does a reliability figure show that a personality score predicts job performance?
No. Reliability concerns score consistency under specified conditions. It does not establish that the score relates to a particular job’s performance outcome.
Can evidence from another job support a report’s claim?
It can when the assessment, important work behaviors, outcome, population, and intended use match closely enough, and relevant differences are addressed.
Is a study of current employees evidence about applicants?
It may inform the question, but applying it to applicants requires examining how the employee and applicant groups, measures, and work context differ.
What should I ask about a report’s job-performance claim?
Ask which score and version were studied, which job behaviors and performance outcome were measured, who took part, and why that evidence applies to the job and proposed use.
Sources and notes
- Designing an Assessment Strategy — U.S. Office of Personnel Management
Supports the distinction between reliability, job analysis, predictive validity, and incremental validity, and explains why inconsistent scores limit useful prediction.
- Assessment Glossary — U.S. Office of Personnel Management
Defines job analysis, criterion, reliability, predictive and concurrent evidence, and incremental validity for interpreting assessment claims.
- Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures — U.S. Equal Employment Opportunity Commission
Supports the U.S.-specific discussion of transferring evidence across jobs and samples, fairness, job behaviors, criteria, and limits of content-validity reasoning for abstract traits.
- Personality Tests — U.S. Office of Personnel Management
Supports the point that relevant personality factors depend on the job and that overall performance, customer service, and teamwork are distinct criteria.
- Personality and Job Performance: The Big Five Revisited
The accessible abstract reports a meta-analysis of explicit Big Five measures and distinguishes job-performance findings from more complex contextual-performance relations.
- Personality Measures as Predictors of Job Performance: A Meta-Analytic Review
The publisher abstract reports findings from a historical meta-analysis that job-analysis-informed selection of personality measures was associated with higher corrected mean validity.
- APA Guidelines for Psychological Assessment and Evaluation
The accessible authoritative excerpt supports matching score inferences to specified purposes and populations and considering limits across evidence sources.
Apply it to your work
Turn a recurring work question into specific observations
From this guide: A general performance claim cannot tell you how your own tendencies show up in decisions, feedback, conflict, or change at work.
If the evidence leaves a personal question open, the Work Pattern Report can help you organize observations about how you approach decisions, feedback, conflict, and change. Use its results as prompts for reflection alongside your own work examples and context. It does not supply norms, selection scores, or job predictions.
