In brief

Compare when the personality score and performance outcome were collected. A score recorded before a later outcome is prospective; a score compared with an outcome collected around the same time is concurrent. Prospective timing better fits a future-performance claim, but neither timing alone establishes a useful forecast. Check whether the instrument, sample, criterion, setting, and intended use match the claim.

What makes a study prospective rather than concurrent?

Classify the evidence by comparing when the personality score and outcome were collected. If the score came first and researchers measured the outcome later, the design is prospective. If the score is compared with an outcome already recorded at roughly the same time, the design is concurrent. The U.S. Office of Personnel Management’s ‘Designing an Assessment Strategy’ describes the prospective case: a personality test intended to forecast applicants’ job success should be related to subsequent job performance. The sequence matters more than a report’s use of ‘predicts’; publication date is not assessment date.

For a quick audit, make a timeline with the assessment date, the period represented by the performance outcome, and when that outcome was recorded. The predictor is the assessment score; the criterion is the outcome used to check the claim. Suppose an employee completes a questionnaire in June, and a supervisor’s rating covers July through December. The score came before the period being rated, so the design is prospective even though the participant was already an employee. If a June questionnaire is matched to a rating of work from January through May, it did not forecast that earlier performance. A rating entered later is not necessarily a later outcome: check the period it describes.

Abstracts may omit dates. Look in the methods section or the provider’s technical report for those dates. If you cannot reconstruct the sequence, mark the design unclear rather than guessing from promotional language. Prospective timing is a useful first classification, not a verdict on the report: it establishes that the score preceded the measured outcome, while leaving the outcome’s quality and relevance for separate examination.

Sources: Principles for the Validation and Use of Personnel Selection Procedures; Designing an Assessment Strategy

Why is the sample separate from the timeline?

‘Prospective’ describes when the score and outcome were measured; ‘applicant’ or ‘incumbent’ describes who supplied the data. These are separate features of a study. A study of current employees can be prospective if researchers record a personality score and then collect a later performance measure. Conversely, a study of applicants can be concurrent if it compares their assessment scores with an outcome already available at about the same time. The sample label alone cannot tell a reader whether a performance claim concerns the future. Who took part still matters because the people in a study may differ from the people to whom a report’s claim is applied. Incumbents have already entered a role. Their experience, training, and exposure to its routines may shape both their answers and the performance evidence available about them. Applicants answer in a setting where they may believe their responses affect selection. That incentive can matter for self-report measures, though it does not mean every applicant changes answers or that incumbent evidence has no value. The meta-analysis “The criterion-related validity of conscientiousness in personnel selection: A meta-analytic reality check” illustrates why the two labels should be inspected separately. Across 102 studies of self-reported conscientiousness and job performance, the authors reported estimates that did not differ significantly by design or sample type. The sample-type estimates were .18 for incumbents across 92 studies and .14 for applicants across 10; these did not differ significantly. Yet only about 12% of the reviewed studies used real applicants in predictive designs. The authors therefore described the evidence for realistic applicant prediction as unresolved, particularly where responses may be shaped by selection stakes. Sparse applicant studies leave that target population less directly represented in the comparison. For employment selection, the Uniform Guidelines on Employee Selection Procedures make sample fit an explicit question. They say that, for predictive and concurrent studies alike, participants should as far as feasible represent the candidates normally available for the relevant job or job group. They also call for attention to comparability when samples differ in the actual work performed or time on the job. A concurrent study of incumbents may be useful for a carefully bounded claim about those employees; it simply does not become applicant evidence because a report calls its result predictive. When reading a report, record the population description separately from the score and outcome dates: current employees or applicants, which job or group, and what relevant experience the study reports. This population check complements, rather than replaces, the timeline check: both are needed to state what the study actually supports.

Sources: The Criterion-Related Validity of Conscientiousness in Personnel Selection: A Meta-Analytic Reality Check; Uniform Guidelines on Employee Selection Procedures

What counts as the performance outcome?

A performance claim is only as specific as the outcome used to test it. A supervisor rating, training result, personnel record, output rate, or error count may capture different parts of work. Ask what was measured, who supplied or recorded it, and why it represents the performance named in the report. The outcome used to check an assessment score is called a criterion measure. The label can cover quite different things: a supervisor’s judgment, success during training, information in personnel records, or a count such as production or errors. They are not interchangeable simply because a report groups them under “performance.” They answer different questions. A supervisor may observe work over time, and may capture conduct a count misses. But a rating remains a judgment shaped by observation and rating standards. The Uniform Guidelines on Employee Selection Procedures call for careful development of supervisory rating methods and instructions, and attention to possible bias in choosing and applying criteria. A rating does not directly measure every aspect of performance. Records and numerical indicators can be more concrete, but may be narrow. An error count concerns recorded errors; it does not describe communication or prioritization without evidence connecting them. The Guidelines identify production rate, error rate, tardiness, absenteeism, and length of service as possible criteria, while requiring a criterion to represent important work behavior or outcomes in the employment context. Objectivity does not establish fit with a broader claim. Training results need a separate fit check. The Guidelines say training success should be measured properly and its relevance to the job shown, either by comparing training content with important work behaviors or by demonstrating a relationship between training and job performance. A training test may support a claim about training performance; it does not become evidence of later job performance merely because training is work-related. The distinction appears in research as well. “The Big Five Personality Dimensions and Job Performance: A Meta-Analysis” examined job proficiency, training proficiency, and personnel data across five occupational groups. Its abstract reports that conscientiousness had consistent relations across criteria and groups, while estimates for other dimensions varied by criterion and occupation. The criterion and occupation matter; these findings are not estimates for a particular report. When a report says a trait “relates to performance,” look for the criterion’s name and source. Was it an overall rating, training measure, personnel record, or specific work outcome? If it was an overall rating, does the study explain why it fits the job? If it was a narrow indicator, does the report keep its conclusion equally narrow? “Performance” is not one identical outcome across studies. Matching wording to the criterion keeps an association with one measured result from expanding into a claim about a person’s whole working ability.

Sources: The Big Five Personality Dimensions and Job Performance: A Meta-Analysis; Uniform Guidelines on Employee Selection Procedures

What does the broadest direct comparison show?

The clearest direct comparison in this evidence set comes from the meta-analysis The Criterion-Related Validity of Conscientiousness in Personnel Selection: A Meta-Analytic Reality Check. It asks whether the observed association between self-reported conscientiousness and overall job performance differed between concurrent and predictive validation studies. The authors pooled 102 studies, covering 23,305 participants, and reported an overall correlation of .17. A correlation summarizes how two measures varied together across people; it is an association, not a hit rate, a probability that one person will succeed, or a guarantee about an individual outcome. When the authors separated studies by design, the pooled correlation was .18 for concurrent studies and .15 for predictive studies. The concurrent estimate drew on 78 studies and 19,132 participants; the predictive estimate drew on 24 studies and 4,173 participants. The authors reported that the difference between designs was not statistically significant. Their finding complicates a simple rule that concurrent evidence must always produce a lower association than predictive evidence. It does not show that the designs are interchangeable, however. A non-significant difference does not demonstrate equivalence or show that the same result holds for every instrument, sample, trait, or outcome. The size and composition of the evidence base matter as much as the two point estimates. Although there were 24 predictive studies, only about 12 percent of all studies in the review used real applicants in predictive designs. The authors consequently described the question of whether self-report conscientiousness measures retain predictive validity under realistic applicant conditions as still open. Timing alone does not make a study representative of job applicants, and a predictive study of incumbents is not applicant evidence. Applicant-specific evidence remains limited, so the review cannot settle how responses under selection stakes relate to later performance across workplaces. The finding is also narrower than the phrase “personality predicts performance” suggests. The review focused on self-reported conscientiousness and supervisor-provided measures of overall job performance in organizational field studies. Its pooled results do not cover every personality trait, every kind of assessment, or every meaning of performance. An association for this construct does not validate a commercial report’s score interpretation or proposed use. Even within the review’s scope, studies differed; the authors reported meaningful variation among effects. A reader should therefore treat the estimate as evidence about this defined body of research, not as a transferable coefficient for an unnamed report. The practical conclusion has two parts. Concurrent evidence is not automatically worthless: this meta-analysis found a similar pooled association in its concurrent and predictive groups, so dismissing every same-time study would go beyond the result. Yet the comparison does not license a broad claim that same-time ratings forecast applicants’ future performance, or that a report predicts how a particular person will perform. A careful statement is that self-reported conscientiousness had modest pooled associations with supervisor-rated overall job performance, with no detected significant difference between designs. Sparse real-applicant predictive evidence and the study’s narrow trait and criterion keep that conclusion bounded. For a report reader, the comparison is a reason to inspect the study behind a claim, not a substitute for checking its dates, sample, outcome, and intended use.

Sources: The Criterion-Related Validity of Conscientiousness in Personnel Selection: A Meta-Analytic Reality Check

What does a local counterexample change—and what does it not?

Fine’s 2025 article, “Does concurrent validity really estimate predictive validity in psychological testing? Two local studies,” challenges the blanket claim that same-time evidence cannot approximate later evidence. It does not directly validate a personality report’s job-performance claim: the outcome was consumer loan default, not employee performance. The article compared the same assessment and loan-default criterion in incumbent and applicant consumer samples from two South American financial institutions. Across the two institutions, the abstract reports samples of 2,942 and 2,880 and no statistically significant differences in observed validity coefficients between incumbent and applicant groups. It also reports evidence of impression management in applicant samples. Holding the assessment and outcome consistent makes this a local comparison of concurrent and predictive estimates within a particular consumer-credit setting, rather than a comparison assembled from unrelated measures and outcomes. The instrument, Worthy Credit, was a 19-item self-report measure designed to predict loan defaults. The criterion was recorded defaults, defined as payments 30 to 60 days past due during the first six to twelve months of repayment. Thus, the study asks whether scores related similarly to this defined credit outcome in its incumbent and applicant samples. It does not test whether a general personality report forecasts effectiveness at work, predicts a supervisor’s later evaluation, or identifies an individual’s future performance. The full article adds qualifications. Applicant scores were higher than incumbent scores, which the author treats as evidence of impression management. Concurrent samples also had low response rates, and voluntary participation may have favored better-performing customers. These features do not erase the comparison, but they matter when deciding how far to carry it. Even with the same assessment and criterion, the ways groups were formed and outcomes gathered can affect interpretation. Range restriction means that scores or outcomes in a sample cover a narrower span than in the population of interest. Fine reports that correcting incumbent samples for apparent restriction in loan-default outcomes raised the concurrent coefficients; the corrected estimates then differed from the observed applicant estimates. The article cautions that these corrections would likely overestimate predictive validity in this setting and that the formula was not designed specifically for dichotomous outcomes. The uncorrected and corrected comparisons therefore tell different stories. An adjustment is not a neutral repair that automatically reveals the “true” relationship. This is a counterexample to an absolute rule, not proof of universal equivalence. “No significant difference” means the analysis did not establish a difference; it does not demonstrate that the estimates are equivalent or that concurrent designs work equally well elsewhere. Fine says the results depend on the particular test and consumer-loan products and that replication is needed before firmer conclusions can be drawn. For a reader auditing a work-performance claim, the study’s value is methodological: labels alone cannot settle the question. A close comparison may inform its own context when measure, outcome, and setting align. But consumer borrowers, a credit-risk questionnaire, and recorded loan defaults are not interchangeable with job applicants, a personality report, and workplace criteria. The finding narrows the verdict: concurrent evidence is not automatically worthless, yet this credit study cannot substitute for direct evidence that a report forecasts work performance.

Sources: Does concurrent validity really estimate predictive validity in psychological testing? Two local studies

Does a prospective association mean the report can forecast a person?

No. A prospective timeline shows that the score preceded the measured outcome; it does not prove that the association is strong enough for individual decisions, that the outcome captures the promised performance, or that the result applies to a different population or use. Timing answers whether the score came first; forecasting also depends on what was measured, in whom, and for what decision. The Standards for Educational and Psychological Testing frame validity around evidence and theory supporting a score interpretation for a proposed use. A study does not validate a score in the abstract. It supports a particular interpretation under conditions. If a study links a self-report score to a later supervisor rating among experienced employees, it concerns that measure, those employees, that rating, and that setting. It does not automatically show that the score forecasts applicants’ performance in another role or guides an individual’s career decision. A change in population or purpose creates a fit question. The outcome needs the same scrutiny. A criterion is the measure used as the performance result in a study. It may be a supervisor rating, training result, work record, or other outcome; its later date cannot show whether it represents the performance named in a report. A rating of one job duty is not necessarily overall effectiveness, and a recorded output can be concrete yet narrow. The criterion must match the claim. A correlation describes how scores and outcomes varied together across the studied group. It does not tell a reader with certainty what one person will do. Its practical value depends on the decision and information already available. The Office of Personnel Management’s “Designing an Assessment Strategy” distinguishes evidence that a measure relates to later performance from incremental validity: whether combining it with other measures adds useful predictive information. Do not treat a coefficient from other guidance or a different population as an estimate for this report. Prospective evidence is a closer design match to a future-oriented claim, but not a guarantee of a forecast. Stronger wording such as “forecasts applicant performance in this role” calls for evidence that fits the instrument, applicant population, criterion, work setting, and use. If a provider establishes a prospective association in a different sample, narrower wording should preserve those particulars. If the report cites only same-time ratings, it may describe an association in that sample; it should not silently turn it into a forecast of later performance. Reliability does not close this gap. Reliability concerns score consistency or precision; by itself, it cannot show that an interpretation about future work is supported. A polished narrative cannot supply missing evidence either. A useful report makes clear what its score represents and where the supporting studies apply. When those details are absent, conclude that the forecasting claim remains unestablished from the information presented, not that the measure has been proved useless. The practical test is fit: score, people, outcome, setting, and purpose must align with the wording. A prospective association is only part of that case.

Sources: The Standards for Educational and Psychological Testing; Designing an Assessment Strategy

How should a reader audit the wording in a report?

Translate each performance sentence into a bounded claim: which score, from which instrument, was associated with which later or same-time outcome, for which people and setting? Keep the verb no stronger than that evidence supports. The Standards for Educational and Psychological Testing frame validity around evidence supporting a particular interpretation of scores for a proposed use. So the practical question is not whether a test has ever been studied; it is whether the cited study supports this sentence, about this measure and outcome, for this purpose. Use a two-pass audit. First, reconstruct the study rather than paraphrasing its summary. Record the instrument and version if given, when participants completed the score measure, when the outcome was recorded, who took part, and what the criterion was. Note whether that criterion came from a supervisor rating, a training result, or a record,. The Uniform Guidelines on Employee Selection Procedures emphasize that employment criteria should reflect important work behavior and that subjective supervisory evaluations need scrutiny. This makes the rater and outcome source part of the claim, not minor study details. If the cited source does not disclose enough to establish the timeline or outcome, write “unclear” in your notes rather than filling the gap from the report’s wording. Second, compare those particulars with the claim’s scope. Suppose, purely as an illustration, a report says, “This scale predicts strong performance.” If the only evidence described is that current employees completed the scale and received supervisor ratings during the same period, the careful wording is: “In this studied employee group, scores were associated with supervisor ratings collected around the same time.” That describes a concurrent association. It does not establish that the score came before the work period or predicts future performance. The Office of Personnel Management’s Assessment Strategy distinguishes evidence based on subsequent job performance from other forms of evidence used to support assessment decisions; collection dates matter more than a marketing label. If documentation instead shows that scores were collected first and a later performance criterion followed, the wording can become: “In this sample and setting, earlier scores were associated with the later measured outcome.” That is a prospective association. It still does not automatically warrant “predicts strong performance” without qualification: the outcome may be narrower than overall performance, and a different group, job, or decision may require its own evidence. The Standards’ proposed-use principle is a reason to keep the instrument, population, criterion, and use visible in the sentence. Ask the provider for the technical report or full study citation when the summary leaves those details out. A focused request is: Which instrument version was studied? When were the scores and criterion collected? Who was in the sample? What outcome and rater or record were used? What use does the evidence support? A short promotional summary may not answer these questions, and a reader cannot certify a report’s validity from a citation alone. The audit has a more modest purpose: preserve the distinction between what the study shows and what the report asks the reader to infer. When dates, sample, outcome, or purpose remain unavailable, describe the fit as unclear and avoid repeating a broad forecast claim as established.

Sources: Designing an Assessment Strategy; The Standards for Educational and Psychological Testing; Uniform Guidelines on Employee Selection Procedures

What should you conclude when key study details are missing?

When a report cites a performance study but leaves out a key detail, treat the design or claim as unresolved. Missing dates do not show that a study was concurrent, prospective, or invalid. They show that you cannot yet classify its timing. The same restraint applies when the report does not identify who was studied, what counted as performance, or what interpretation the study was meant to support. Until those details are available, do not repeat the strongest version of the report’s claim as established. This is a limit on what can be concluded, not an accusation about the assessment. A summary may be brief while the underlying study or technical documentation contains the needed information. Conversely, a confident sentence in a report cannot fill gaps in the evidence it cites. The Standards for Educational and Psychological Testing frame validity around evidence and theory supporting a particular interpretation of scores for a proposed use. That principle makes missing context consequential: a study can be real and carefully conducted while still not answering the report’s broader question. Each missing detail blocks a different inference. Without the score and outcome collection periods, the reader cannot tell whether the evidence is prospective or concurrent. Without a sample description, the reader cannot judge whether results from current employees, applicants, or another group fit the people named in the claim. Without a defined criterion, the reader cannot know whether “performance” means a supervisor judgment, a training result, a personnel record, or something else. A general reference to “research” may also leave unclear whether the study used this instrument, a related measure, or a different interpretation altogether. These gaps should not be collapsed into one vague verdict that the report is either proven or useless. The next step is a focused request tied to the sentence you are assessing. Ask for the full study citation or technical report, the instrument and version, when the score and outcome were collected, who took part, how the criterion was defined and recorded, and the intended use. If the provider supplies only a short summary, ask where the methods and results can be reviewed. The Office of Personnel Management’s Designing an Assessment Strategy guidance connects evidence to the assessment purpose and outcome; it can orient this request, though its federal personnel-selection context does not validate a particular commercial report. You do not have to re-run a study to read a report responsibly. Your task is to keep the conclusion proportionate to what is visible. If the dates are absent, call the design unclear. If the sample or criterion is unspecified, say that fit cannot be judged from the information provided. If the purpose is unstated, do not assume evidence for one use carries over to another. Once documentation resolves those points, reconsider the claim against the actual study rather than the report’s summary alone. Until then, uncertainty is the accurate conclusion: neither endorsement nor dismissal, and no forecast stronger than the accessible evidence supports. A narrower statement protects the reader from overclaiming while leaving room to update the judgment when better documentation becomes available.

Sources: The Standards for Educational and Psychological Testing; Designing an Assessment Strategy

What is the next useful step?

Write down when the personality score and performance outcome were collected, who took part, what counted as performance, and the report’s intended use. Restate the claim at the narrowest level those details support. A score collected before a later outcome may support a prospective association for that sample and outcome; same-time ratings may support a contemporaneous association. Neither alone establishes an individual forecast or transfers the finding to a different population or use. The Standards for Educational and Psychological Testing frame validity around evidence supporting an interpretation for a proposed use. Ask the provider: Which study supports this sentence, when were score and outcome measured, who was studied, and what outcome did researchers count? If dates or other details are missing, leave the claim unresolved rather than filling gaps with ‘predicts.’ Reflection can still help you examine how a tendency appears in your work. Note a recent planning decision, the information you used, and what happened; treat this as a prompt for discussion, not proof of a stable trait or job fit. The live Work Pattern Report at /assessment offers private, low-stakes reflection across work continuums. It has no norms, cutoff, or selection score, so it does not validate performance claims or recommend a career. Use it to name questions about your patterns, then compare them with concrete work experiences.

Sources: The Standards for Educational and Psychological Testing

Questions readers ask

What makes a personality assessment study prospective?

It is prospective when the assessment score is collected before the later performance outcome being studied. Check the period represented by the outcome, not only when a rating was entered.

Can a study of current employees be prospective?

Yes. Prospective describes the order of measurement; applicant or incumbent describes the participants. A study of current employees can be prospective if their scores precede a later outcome.

Does a prospective study prove that a report predicts an individual’s job performance?

No. Timing establishes that the score preceded the outcome. A broader forecast also depends on whether the instrument, sample, criterion, setting, and intended use fit the claim.

What should I do if the report does not give study dates?

Treat the design as unclear and ask for the study or technical report, including score and outcome collection periods, participants, criterion, and intended use.

Sources and notes

  1. Principles for the Validation and Use of Personnel Selection Procedures

    Supports the distinction between prospective and concurrent designs based on whether there is a time lapse between predictor and criterion collection.

  2. Designing an Assessment Strategy

    Explains how subsequent job performance relates to applicant forecasting and distinguishes validity evidence from individual certainty.

  3. The Standards for Educational and Psychological Testing

    Defines validity in relation to evidence and theory supporting score interpretations for proposed uses.

  4. Uniform Guidelines on Employee Selection Procedures

    Supports checking job relevance, sample representativeness, and the development and possible bias of subjective supervisory criteria.

  5. The Criterion-Related Validity of Conscientiousness in Personnel Selection: A Meta-Analytic Reality Check

    Reports pooled concurrent and predictive associations for self-reported conscientiousness and job performance, and notes the limited real-applicant predictive evidence.

  6. Does concurrent validity really estimate predictive validity in psychological testing? Two local studies

    Reports local comparisons involving consumer loan-default outcomes; it complicates blanket assumptions about concurrent estimates but does not establish workplace performance validity.

  7. The Big Five Personality Dimensions and Job Performance: A Meta-Analysis

    Its abstract distinguishes job proficiency, training proficiency, and personnel data and reports that trait relations vary by criterion and occupation.

Apply it to your work

Turn a vague work-pattern question into observations you can examine

From this guide: After checking what a performance claim can support, compare a recurring work-friction question with concrete patterns in how you decide, plan, collaborate, handle conflict, adapt, and learn.

A study can help you judge a claim, but it cannot tell you how a tendency appears in your own work situations. The Work Pattern Report offers low-stakes self-reflection across ten work continuums. Use it to name patterns and questions to compare with specific experiences; it provides no norms, cutoff, or selection score, and does not recommend a job.