In brief

Sometimes. Evidence from one sample can support a prediction for another population or work setting when it supports both comparable score meaning and a relationship to the same relevant outcome under sufficiently similar conditions. A new local study is not always necessary, but a validation label or broad workplace finding cannot establish transfer alone. If the bridge is undocumented, treat the prediction as unestablished.

Can a personality report predict beyond its validation sample?

Sometimes, but only when evidence supports the inference being carried into the new setting. Transfer means applying evidence beyond the people or circumstances originally studied. A validation sample can contribute to a prediction about another population, but cannot establish by itself that the prediction travels. Evidence must support two links: that the score has comparable meaning for the target group, and that it relates to the same work outcome. Similarity may make transfer plausible; it does not prove either link. A broad finding that personality scores relate to work outcomes cannot establish what a particular report predicts for a different job. The National Academies’ chapter “5 Validity of the Achievement Levels” concerns educational testing, not personality reports. It explains that validity concerns evidence and theory supporting a proposed score interpretation and use, rather than a permanent property of a test. The chapter describes validation as an argument that can be strengthened or challenged. The question is not whether an instrument is called validated, but whether evidence supports this interpretation for this use. If transfer links are undocumented, treat the work prediction as unestablished. The report may prompt reflection, but cannot settle what a person will do at work.

Sources: 5 Validity of the Achievement Levels

What exactly is being transferred?

A report may transfer a description of a tendency, a comparison with a norm group, or a prediction about an outcome. These are different claims; evidence for one does not establish the others. “This score may describe how I usually approach planning” is not the same inference as “people with this score will perform better in this role.” The second predicts an outcome. First identify the precise conclusion being carried over. Validity is sometimes described as though it were a quality stamped onto an instrument: validated or not validated. The National Academies chapter “5 Validity of the Achievement Levels” instead frames validity around whether evidence and theory support interpretations of test scores for proposed uses. It explains that interpretations and uses, not a test in the abstract, are what require validation. Because the chapter concerns educational testing, it supplies a general principle, not empirical evidence about personality instruments. Ask: what evidence supports this score meaning and use? A score interpretation is the meaning assigned to a result. A report might say that a respondent tends to prefer a clear plan before beginning a task. Evidence should support that reading in its intended context. Even if supported for one group, it does not show how common the tendency is elsewhere, whether it appears similarly under another administration, or whether it predicts a work outcome. Each is an added claim. A norm comparison asks where a score falls relative to a specified reference group, called the norm group. A percentile expresses relative standing there, not a probability of success or employer preference. A comparison cannot simply transfer to a population with different composition or norms. The report should identify its group and comparison. Criterion-related validity concerns the relationship between scores and a specified external outcome, called a criterion. For a work prediction, the criterion might be a defined performance measure in a particular role. A relationship with one criterion does not establish one with another, and a group association does not determine an individual’s future. A tendency description and a score-to-criterion study answer separate questions, even for the same scale. Consider a hypothetical report based on people who completed a work-style questionnaire for personal reflection. Suppose its responses support a description of planning preferences in that group. A manager later claims applicants with a certain score will outperform others in a different role. The first finding does not establish this prediction: interpretation, target people, setting, and outcome have changed. This hypothetical is not an actual study; it shows how a conclusion can outrun evidence while the score stays the same. When reading a report, identify the sentence that matters and label it: tendency description, norm comparison, or outcome prediction. Then ask what evidence supports that statement, for which people, and for which use. This prevents treating evidence attached to a score as permission to carry every interpretation into a new setting. A validation sample can contribute to transfer, but the conclusion should not exceed evidence for the proposed interpretation and use.

Sources: 5 Validity of the Achievement Levels

Does comparable score meaning establish comparable prediction?

No. Evidence that a personality score has comparable meaning across groups can support comparing what the measure says about a tendency; it does not, by itself, show that the score predicts the same work outcome in each group or setting. A prediction joins a score and a named criterion. Evidence about the first link cannot stand in for the second. Measurement invariance is evidence that a measure relates to the underlying construct in a sufficiently similar way across specified groups. It addresses whether score differences can be interpreted comparably under the tested conditions. Researchers may ask whether the same broad pattern of items represents the construct in each group, then whether the scale supports the specific comparison being made. Similar labels or a familiar translation alone do not show that numerical differences carry the same meaning. Even when score meaning is sufficiently comparable, the score-to-outcome relationship is a separate question. Criterion-related validity means evidence about how assessment scores relate to a specified external outcome. The relationship applies to the people, conditions, and use represented in the evidence. A tendency can be measured similarly in two groups yet relate differently to a particular work criterion, because the criterion may capture different parts of performance or work demands may differ. A role centered on rapid decisions in one setting and careful review in another may not connect the same way to the measured tendency. This calls for examining the prediction; it does not establish that the tendency helps or hinders either group. The National Academies chapter “Validity of the Achievement Levels” explains the general measurement principle that validity concerns evidence and theory supporting a proposed interpretation and use of scores; the chapter concerns educational assessment, so it does not supply empirical evidence about personality reports or work prediction. This general principle clarifies why “the scale measures the same thing” is not a complete argument for a work prediction, which also needs a score-to-criterion link for the intended population and use. The Society for Industrial and Organizational Psychology’s “Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection” makes a related transfer point for selection: evidence that an assessment predicts in one context does not automatically show that its predictive ability applies in another. It says a transportability argument for an AI-based selection tool should start from a technically sound original study and compare job content, context, requirements, and applicant group. This is adjacent guidance, not direct validation evidence for ordinary personality questionnaires. The transferable lesson is to separate score meaning from prediction and specify which setting features support the inference. When a report predicts a work outcome, ask for both kinds of support: evidence that the score interpretation is comparable across the relevant groups, and evidence that this score relates to the named outcome in the target setting. If only the first is documented, the report may support cautious comparison of a measured tendency, but its work prediction remains unestablished. If the report does not name the group, administration, criterion, or intended use, those omissions limit what can be concluded; they do not prove the measure is useless. They do mean that a confident prediction goes beyond the evidence shown.

Sources: 5 Validity of the Achievement Levels; Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection

Which differences matter when the work setting changes?

A transfer check compares the source study with the proposed use along several dimensions, because ‘similar work’ can hide differences that matter. First, ask whether the assessment itself is the same: version, language, item format, scoring, and administration conditions. A translation, shortened form, revised scoring rule, or change from private reflection to an application requirement may alter what responses mean or how people answer. Second, compare people. Were validation participants current employees, applicants, volunteers, or another group? Do their characteristics and range resemble the population named in the report? The Virginia Tech dissertation, “Personality Test Validation Research: Present-employee and job applicant samples,” reviewed criterion-related evidence for seven personality constructs and found that sample type had a small moderating effect, although effects appeared for some constructs and their direction varied. That result argues against treating every employee sample as unusable for applicant questions; it does not establish that a test, language, applicant pool, or job has been validated. Third, compare the work. Similar titles do not guarantee similar duties, demands, or conditions. The Society for Industrial and Organizational Psychology’s 2023 “Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection” names job content, context, requirements, and applicant group as dimensions in a transportability argument. This guidance addresses AI-based employee selection, so it is an adjacent framework for organizing questions, not direct validation evidence for ordinary personality reports. It says accumulated evidence may sometimes support use in a new setting without a local study when validity generalizes across contexts and a compelling case connects the evidence to that setting. The argument starts from a technically sound study; resemblance is a reason to examine transfer, not proof of it. Fourth, identify the outcome precisely. “Success” might mean supervisor ratings, task completion, retention, or another criterion. These are not interchangeable. Ask how the criterion was defined and measured, whether it represents the work, and whether the report’s prediction concerns that same outcome. SIOP’s guidance emphasizes that outcome selection and measurement affect what the evidence can support, and that a personality score by itself is not a job outcome. Finally, match the intended decision. A tendency description used for private reflection makes a narrower claim than a score used to rank applicants. Applicant stakes may change response incentives; Bradley’s review examined test-taking status as a possible moderator. Ask whether the source evidence covers the administration and decision being proposed, rather than infer either that applicants always respond differently or that the difference never matters. These dimensions interact. One modest, well-documented difference may be addressed by evidence directly relevant to the target. Several unexamined changes at once—such as a translated version, a new applicant group, substantially different work, and a new performance criterion—leave more links in the inference unsupported. A larger sample does not, by itself, resolve those mismatches. When reading a report, ask its provider to identify the version and language studied, who took it under what conditions, which work and outcome were examined, and why those findings apply to the proposed use. If the answer supplies only a broad label such as “validated for workplace use,” the transfer rationale remains unclear.

Sources: Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection; Personality Test Validation Research: Present-employee and job applicant samples

What does pooled work evidence establish—and what does it not?

Meta-analyses can show whether score–performance relationships recur and whether study conditions help explain differences. Their averages apply most directly to the measures, populations, outcomes, and designs combined. A pooled result is evidence about a defined body of research, not a forecast for every report or workplace. Two reviews show why fit and coherence matter as much as size.

The 1991 review, “Personality Measures as Predictors of Job Performance: A Meta-Analytic Review,” began with 494 studies and identified usable findings from 97 independent samples, totaling 13,521 people. Across varied personality measures and performance studies, the authors reported corrected mean validity of .29 for confirmatory strategies, compared with .12 for exploratory strategies. Where job analysis guided measure choice, the mean was .38. These are historical averages, not probabilities of individual performance or coefficients for a current report. The review noted weaknesses in reporting study characteristics, limiting application to a present-day role or instrument.

The contrast among .12, .29, and .38 compares research approaches. Confirmatory work begins with an expectation about which measured tendency should matter for which job demand; job analysis makes that link explicit. The pattern suggests that prediction is more interpretable when measures are chosen for a defined work question. This is an interpretation, not proof that job analysis alone caused the larger estimate. Because the review combined varied measures and criteria, its figures do not establish transfer to an unnamed population or setting.

A narrower example is the 2005 study “Meta-analyses and Validity Generalization Studies of a Personality Test with Salaried Workers.” It examined one test used with Japanese salaried workers, with performance appraisals as the criterion. In the initial analysis, five of 17 scales had corrected validity coefficients above .10 in absolute value; the largest was .21, for Vitality. These findings describe that test, worker samples, and appraisal outcomes. They do not show that another instrument, a different population, or another definition of effective work will yield the same pattern.

The researchers then restricted a second analysis to studies from a specified period, using the same criterion and research purpose. They reported higher, more generalizable coefficients. Pooling is not improved simply by adding every result: differing measures, periods, purposes, and outcomes can blur which relationship is estimated. A narrower pool may be easier to interpret, but applies to a more bounded evidence base.

Heterogeneity means variation among studies in measures, samples, settings, criteria, or methods. It can reflect real differences in how a tendency relates to work or differences in measurement. An average may summarize an overall pattern while concealing outcomes or groups that fit poorly. A focused pool is not universal either: it applies most directly to targets resembling the studies retained. The Japanese analysis illustrates this trade-off, not transfer to every salaried workforce.

Together, the reviews show what pooled work evidence can establish. The broader 1991 review indicates that estimates varied with research strategy and job-analysis alignment. The 2005 single-test analysis shows that estimates changed when studies shared a period, criterion, and purpose. Neither supplies an estimate for an unspecified report, population, work demand, or outcome. To assess transfer, ask whether studies resemble the intended use and whether findings converge for its criterion. A larger pool helps only when its contents answer a sufficiently similar question.

Sources: Personality Measures as Predictors of Job Performance: A Meta-Analytic Review; Meta-analyses and Validity Generalization Studies of a Personality Test with Salaried Workers

Do employee samples ever support claims about applicants?

An employee sample can contribute evidence about applicants, but it does not settle the question by itself. The 2003 dissertation Personality Test Validation Research: Present-employee and job applicant samples is a counterexample to a blanket rule against transfer. Its quantitative review examined criterion-related validity for seven personality constructs and whether test-taking status moderated the relation between scores and job performance. The repository abstract reports that sample-type moderation was generally small. It also reports moderation for some constructs, with inconsistent direction: incumbent validity estimates were sometimes larger and sometimes smaller than applicant estimates. The dissertation’s simulations likewise supported incumbent samples as useful, subject to caveats. This is evidence against saying employee data can never inform applicant claims.

That result is bounded. It aggregates constructs and studies available to a dissertation completed in 2003; it does not validate a current report, occupation, or selection decision. A small average moderator also does not mean employees and applicants are interchangeable in every case. It means the reviewed evidence found limited average difference associated with sample type, while variation remained. Ask whether the evidence addresses the target score interpretation and outcome, and whether the change from employees to applicants could alter how the assessment works or relates to that outcome. The dissertation’s result is about a body of historical research, not a guarantee that the same relationship holds for each instrument or target group.

Study design clarifies the gap. A concurrent design measures the assessment and an existing workforce criterion around the same period; a predictive design measures the assessment first and collects the criterion later. A concurrent study with employees can show an association in that workforce, but does not reproduce the applicant decision and future-outcome sequence. Still, concurrent evidence is not worthless: it can contribute to a larger argument when other evidence supports the missing steps. Describe the design accurately, and do not expand the inference beyond it.

The 2023 meta-analysis The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check provides a narrower counterweight. It examined conscientiousness and supervisor-rated overall job performance in organizational field studies, combining 102 estimates from 23,305 participants. The overall correlation was .17. Estimates did not differ significantly across concurrent versus predictive designs or incumbent versus applicant samples. Yet only about 12% of studies used real applicants in predictive designs, a limit on evidence under realistic selection conditions. This study concerns conscientiousness, not every personality scale or report, and cannot establish whether another instrument predicts another job outcome.

Together, these findings support neither automatic rejection nor automatic transfer. The older review shows why employee data may inform a claim; the newer, construct-specific review shows why similar averages can coexist with thin evidence for exact applicant conditions. For a consequential prediction, ask for evidence connecting the score, administration, target group, and job criterion, and note whether support is concurrent, predictive, or cumulative. A fresh applicant-only study is not always required, but the transfer argument must address the changed sample and intended conditions. Without that account, the prediction remains unestablished for that use.

Sources: Personality Test Validation Research: Present-employee and job applicant samples; The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check

What changes when an assessment is used with applicants?

Applicant use changes what validation needs to establish. Evidence from current employees may show that scores relate to an outcome among those employees; it does not automatically show the same relationship when people answer while seeking a job. Applicants know responses may affect an opportunity, and later performance is measured after hiring. These differences narrow what employee evidence establishes. They do not prove every self-report prediction fails or employee samples are useless.

The 2023 meta-analysis “The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check” examined self-reported conscientiousness and supervisor-rated overall job performance in organizational field studies. Across 102 correlations and 23,305 participants, the overall correlation was .17. Analyses found no significant differences by design or sample type: concurrent studies reported .18 and predictive studies .15; incumbent samples .18 and applicant samples .14. This counters a blanket claim that incumbent findings cannot inform applicant use. Yet only about 12% of studies used real applicants in predictive designs, so evidence under realistic selection conditions was scarce. This does not validate a particular report, role, or population.

The scope matters. That review concerns conscientiousness, not every personality dimension or instrument, and uses supervisor ratings of overall performance, not every definition of work success. Its authors say whether self-report conscientiousness retains predictive validity when applicant responses may be shaped remains open. This is an evidence gap, not proof that applicants routinely distort answers or scores contain no useful information. The possibility of response shaping is a reason to study realistic conditions, not a basis for claims about an individual applicant.

An earlier counterpoint is the 2003 Virginia Tech dissertation “Personality Test Validation Research: Present-employee and job applicant samples.” Its review covered seven constructs, including conscientiousness, openness, and extraversion. The repository abstract reports generally small sample-type moderation of criterion-related validity, with inconsistent direction: incumbent estimates were sometimes higher and sometimes lower. It supports cautious use of employee samples, rather than automatic rejection. But it is historical and aggregates constructs, instruments, and occupations; it cannot validate a current report or hiring decision.

The findings fit when their questions remain distinct. The dissertation examines sample-type differences across constructs and finds small, inconsistent moderation. The later review focuses on conscientiousness and underscores how rarely real applicants were followed to later performance. An average showing no clear difference is not direct evidence for every target setting; scarce realistic studies do not prove transfer impossible. A report should identify its source group, assessment conditions, outcome, and intended decision, then explain their relevance to applicant use. Without that account, the prediction remains uncertain, not established.

Ask the provider: “Did applicants take this version under selection conditions, and what later work outcome was measured?” If evidence comes from current employees in a lower-stakes context, ask how it transfers. Comparable studies may support transfer; a new applicant sample is not the only possible evidence. Neither review settles a named instrument’s use. Applicant context is a specific uncertainty to investigate, not a verdict against personality assessment as a whole.

Sources: The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check; Personality Test Validation Research: Present-employee and job applicant samples

How strong is the bridge for this particular report?

A useful way to judge a report’s claim is to grade the bridge, rather than label the whole assessment either “validated” or “not validated.” The National Academies chapter on validity describes evidence as supporting particular interpretations and uses. Ask whether evidence reaches from this score and population to the proposed outcome and decision. Each link matters: the score must mean something comparable, and its relationship with the named outcome must hold under relevant conditions. At the strongest level, studies examine the relevant version and administration, a population close to the intended users, work demands that resemble the target setting, and an outcome that matches the report’s stated prediction. A second, still useful level is convergent evidence: several studies, not necessarily from the exact target site, point in a consistent direction across sufficiently comparable groups and settings. The 1991 meta-analysis “Personality Measures and Job Performance: A Meta-Analytic Review” reported different corrected mean validities by research strategy: .29 for confirmatory studies, .12 for exploratory studies, and .38 where job analysis guided measure selection, among 97 independent samples drawn from 494 reviewed studies. Those are historical averages across varied measures and criteria, not forecasts for a current report. This variation warns that study volume alone does not erase differences in how traits and work requirements were matched. The 2005 article “Meta-Analysis of Personality Test Validity for Personnel Selection in Japan” offers a different kind of illustration. It focused on one test among Japanese salaried workers and performance appraisal criteria. Five of 17 scales had corrected validity above .10 in absolute value, with a reported maximum of .21; the authors also found higher, more generalizable estimates in a restricted analysis using a common period, criterion, and research purpose. It illustrates why a coherent set of studies can be more informative than a larger but mixed collection. Together, they suggest asking whether accumulated evidence coheres around the intended claim. A broad analogy sits below that level. Suppose a report says, “This score predicts success in collaborative roles,” but gives no named outcome, tested group, or study conditions. If the only support offered is that the trait has been associated with performance somewhere, the relationship may be a reasonable hypothesis to investigate. It is not yet a supported prediction for the reader’s role. Sample size may improve precision within the sampled conditions, but cannot by itself show that a different population, criterion, or administration preserves the relationship. Likewise, one local study is not automatically better than several strong, relevant studies elsewhere. It may still use a weak outcome, changed questionnaire, or design that misses the proposed use. Conversely, a set of technically sound studies can support cautious transfer without a new local study if it covers the important differences and converges on the same interpretation. The Society for Industrial and Organizational Psychology’s guidance on transportability concerns AI-based employee selection, so it is adjacent rather than direct personality-test evidence. Its comparison of job content, context, requirements, and applicant group helps make “similar setting” more specific. Similarity must be evidenced along dimensions that could alter the claim. Consider this illustrative sentence in a report: “People with this pattern tend to perform well in fast-changing team roles.” Ask which people were studied, what counted as “perform well,” and whether the roles were fast-changing and team-based. If documentation reveals only an incumbent sample in another occupation and supervisor ratings, the claim should shrink to something like: “In the studied group, this score was associated with those ratings.” It separates a finding in a source group from a prediction about a different group whose population and criterion remain unshown. Finally, the acceptable bridge depends on what someone will do with the conclusion. A tentative pattern can be a prompt for development: compare it with observed situations, ask what counterexamples exist, and decide what to explore. The same indirect evidence should carry little weight in a consequential employment judgment, where an unsupported transfer could affect opportunity. State the boundary plainly: “This evidence supports a description,” “it supports a comparison with this reference group,” or “it supports a prediction for this specified outcome and setting.” If the report cannot identify which kind of claim it makes, the bridge remains too unclear to bear the prediction.

Sources: 5 Validity of the Achievement Levels; Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection; Personality Measures as Predictors of Job Performance: A Meta-Analytic Review; Meta-analyses and Validity Generalization Studies of a Personality Test with Salaried Workers

What should you do with a prediction whose transfer bridge is missing?

Ask the provider for the technical report behind the claim. It should identify the assessment version, people studied, administration, work outcome, and intended use. Compare these with the situation named in the prediction. The National Academies’ “Validity of the Achievement Levels” explains a general principle from educational testing: evidence supports an interpretation and proposed use. It does not establish personality-test findings, but it suggests the right question: what evidence connects this score to this conclusion? If a key link is missing, narrow the claim. A statement about a tendency can prompt reflection without establishing what one person will do in a different job. Compare it with repeated, observable situations: when did the pattern appear, when did it not, and what about the task or workplace may explain the difference? Keep counterexamples in view. This examines experience; it is not a substitute validation study or prediction score. For a private prompt about decision and collaboration tendencies, the live Work Pattern Report offers ten continuums and a local summary. Its result has no norms, cutoffs, type, or employment prediction. Use it to make a work-pattern question more specific, not to decide who should be hired or promoted. The verdict is conditional: a source sample can support transfer when evidence connects comparable score meaning to the same relevant outcome under sufficiently similar conditions. When that bridge is undocumented, treat the prediction as unestablished. Evidence for the target group, work setting, criterion, and use could change that conclusion.

Sources: Work Pattern Report; 5 Validity of the Achievement Levels

Sources and notes

  1. 5 Validity of the Achievement Levels

    Supports the general principle that validity evidence concerns score interpretations and proposed uses; the chapter covers educational testing, not personality tests.

  2. Considerations and Recommendations for the Validation and Use of AI-Based Assessments for Employee Selection

    Provides adjacent employee-selection guidance on comparing job content, context, requirements, and applicant groups when making a transportability argument.

  3. Personality Measures as Predictors of Job Performance: A Meta-Analytic Review

    Reports historical pooled personality-work validity estimates that varied by research strategy and job-analysis alignment; they are not forecasts for a particular current report or role.

  4. Meta-analyses and Validity Generalization Studies of a Personality Test with Salaried Workers

    Describes analyses of one test among Japanese salaried workers and illustrates how restricting the evidence to a common period, criterion, and purpose affected estimates.

  5. Personality Test Validation Research: Present-employee and job applicant samples

    Repository abstract reports a 2003 review across seven constructs in which sample-type moderation was generally small but inconsistent in direction.

  6. The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check

    Reports a 2023 conscientiousness meta-analysis and notes limited real-applicant predictive designs, leaving its findings bounded to that construct and selection context.

  7. Work Pattern Report

    Identifies the live 100-item self-report and its ten work-pattern continuums; it offers no norms, cutoffs, type, or employment prediction.

Apply it to your work

Turn a vague work-pattern question into something you can examine

From this guide: If the evidence does not establish what a report predicts in your setting, start with observable tendencies and the situations where they appear.

A general report cannot tell you how your own decision, planning, collaboration, or response to change shows up across specific work situations. The Work Pattern Report offers a private, low-stakes way to describe tendencies across ten continuums and make a reflection question more specific. It provides no norms, cutoffs, type, or employment prediction. Use it to guide reflection, then compare the pattern with your own observations.