A higher score on a retaken employment personality test means the responses produced a higher score the second time. It does not, by itself, show that personality changed, that the first result was wrong, or that the person is now a better fit for a job. Research finds systematic response differences when stakes and feedback change, but group findings cannot explain one individual result. To interpret a difference, check whether it exceeds the instrument’s own expected measurement error, then compare the test version, interval, instructions, stakes, feedback, and life or work context. Even a difference larger than expected error establishes a change in measured score, not automatically a lasting change in a personality tendency.
What does employment retest evidence actually compare?
The clearest evidence on employment personality retesting does not simply compare a person’s personality at two points in time. It compares responses under different conditions. Jing Hu and Brian Connelly’s 2021 meta-analysis combined 20 within-person studies in which applicants completed a personality assessment once in a high-stakes applicant setting and again in a low-stakes setting. On average, applicant responses were moderately more favorable in socially desirable directions; score variability was slightly lower, and rank-order consistency was stronger. Order mattered: studies administering the high-stakes test first found a stronger effect than those administering the low-stakes test first. These are group patterns, not explanations for one person’s score. [Hu and Connelly](https://onlinelibrary.wiley.com/doi/10.1111/ijsa.12338)
Three different questions are often collapsed into “the score went up.” First, did the observed difference exceed the variation expected from measurement error? Measurement error is the score variation that can occur even if the measured tendency has not changed. Second, did response context change in a way that shifts scores across a group, such as the perceived stakes of applying? Third, did the person’s underlying tendency genuinely change? A high-stakes versus low-stakes comparison can inform the second question. It does not by itself answer the first or third.
The dates matter because this assignment asks what changed in recent evidence. The accessible evidence reviewed here does not establish a new 2026 consensus about individual personality retests. The direct applicant-versus-low-stakes synthesis was published in 2021. A broader 2017 review by Chad Van Iddekinge and John Arnold surveyed employment retesting research dating back nearly a century. Its abstract says retest scores tend to be higher, more varied, more reliable, and somewhat more related to criteria such as academic or job performance than initial scores. It also says research has not clearly identified the factors behind differences between initial and retest scores. This review covers employment tests broadly, not just personality measures, so it provides context rather than a personality-specific rule. [Van Iddekinge and Arnold](https://www.annualreviews.org/content/journals/10.1146/annurev-orgpsych-032516-113349)
The useful comparison is not simply an old number against a new one. Ask whether the same scale, version, format, instructions, and stakes were used. If an applicant first answered for a job and later completed the measure privately, the numbers may reflect different response frames. If conditions were similar, the difference still needs comparison with that scale’s measurement uncertainty. The design of the comparison determines what can be concluded.
Sources: Faking by actual applicants on personality tests: A meta-analysis of within-subjects studies; Retaking Employment Tests: What We Know and What We Still Need to Know
Can feedback or higher stakes explain a higher score?
They can explain some score movement, but the direction alone cannot establish that they did. In a 2013 study, Courtney Holladay, Emily David, and Stefanie Johnson examined actual applicants who completed an unproctored personality selection assessment a second time. Applicants told that a low assessment score was the reason they had not received a job were more likely to use response strategies aimed at presenting themselves favorably on the retest. They were predicted to show the greatest score change. Applicants without that feedback and a student comparison group showed less response change. The publisher abstract supports this description of the study; it does not show that every applicant with a higher score distorted responses. [Holladay, David, and Johnson](https://journals.sagepub.com/doi/10.2466/03.01.PR0.112.2.486-501)
Feedback offers a plausible mechanism: it can tell a person which part of an earlier result seemed consequential, making a desired impression easier to infer. A second exposure can also make questions feel more familiar. Response distortion means answering partly to create a favorable impression. It does not prove a specific person was dishonest, and a score cannot reveal motive. The Hu and Connelly synthesis also finds average differences between high- and low-stakes settings; its order effect shows that sequence can shape the comparison. Together, these findings make context one credible explanation. They do not rule out ordinary error, genuine change in behavior, or a change in the situations a person had in mind when answering.
Consider an unnamed applicant who receives feedback, retakes the same assessment, and sees a higher score on a scale related to planning. The result may reflect a learned sense of what the assessment seems to value, a different reading of the questions, changed planning habits at work, or ordinary variation within the measure’s uncertainty. The cited studies cannot choose among these explanations for that applicant. Compare recent examples: Have planning routines changed? Did the work itself change? Were the instructions or consequences different? This is an illustration of useful questions, not evidence about a test taker.
Record the interval, assessment version, format, instructions, stakes, feedback received, and relevant changes in work or routine. Compare like with like where possible. When conditions differed, say so. A report that records only “higher on retest” omits information needed to interpret why.
Sources: Faking by actual applicants on personality tests: A meta-analysis of within-subjects studies; Retesting Personality in Employee Selection: Implications of the Context, Sample, and Setting
What can a score increase establish about one person?
An instrument-specific uncertainty analysis can show whether a difference is larger than variation expected from measurement error under stated assumptions. It cannot by itself prove why the score changed. The Educational Testing Service’s guide to test reliability explains that reliability concerns consistency across occasions, test editions, or raters, and describes the standard error of measurement as one way to express score uncertainty. This is why the technical manual matters: uncertainty depends on the instrument and sometimes on the score range, not on a universal number that applies to every personality scale. [ETS](https://www.ets.org/research/policy_research_reports/publications/report/2018/jysw.html)
A reliable-change method uses instrument-specific information to judge whether a difference is larger than expected under a stated model of no change. It is not a threshold to borrow from another test. Check whether the assessment manual provides a standard error or a method for interpreting changes across administrations, and whether that method fits the scale and retest interval. If it does not, avoid calculating a cutoff from a generic online formula. An estimate based on the wrong reliability coefficient, score range, or conditions can look precise while answering the wrong question.
A careful interpretation follows four steps. First, verify that both results refer to the same construct, edition, and score metric. Second, find the manual’s measurement-error information or an appropriate reliable-change procedure. Third, check whether norms and administration conditions are comparable. Fourth, examine what else changed, including stakes, feedback, timing, instructions, and relevant work experience. If the report lacks information needed for a suitable calculation, the supported conclusion is limited: the observed response changed, but the evidence supplied does not show whether the difference exceeds expected error.
Even a change larger than expected measurement error establishes a change in measured score under the method used. A stronger claim about a lasting personality tendency needs evidence that addresses time, context, and intended interpretation. The broad employment review is a useful counterweight to the view that all movement is random: retest results can differ systematically and relate differently to criteria. Yet its authors say the factors behind those differences remain unclear. The evidence leaves room for meaningful change while withholding a verdict about why one individual’s score rose. A stable result also is not automatically valid for a particular employment decision; reliability and validity answer different questions.
For self-reflection or coaching, use the difference to start a grounded conversation: What behavior does the scale describe? Can you identify recent examples in more than one setting? What conditions make that behavior more or less likely? Compare the report with observable actions and context instead of treating a label as a fixed identity. In hiring or another employment decision, a retest increase is not a hiring score, job recommendation, diagnosis, or proof of improved competence. The evidence reviewed here does not validate score change for those conclusions.
The verdict is conditional. A higher retest score can mean more than measurement error when a suitable, instrument-specific analysis shows the difference exceeds expected uncertainty. Even then, it means the measured score changed, not necessarily that an enduring trait changed. The 2021 synthesis and 2013 applicant study make response context a credible alternative when stakes, feedback, or test conditions differ. The useful next action is to retrieve the scale’s own uncertainty guidance and note what changed between sessions before revising your view of a person’s work tendencies.
Sources: Retaking Employment Tests: What We Know and What We Still Need to Know; Test Reliability—Basic Concepts
Questions readers ask
If my employment personality score rose on retest, should I trust the newer result?
Do not choose between the scores by date alone. Check whether they use the same scale and score metric, review the assessment’s measurement-error information, and note changes in interval, stakes, feedback, or instructions. If the report provides no suitable method for interpreting change, treat the difference as an observation to discuss, not proof that either score is the true one.
Sources and notes
- Faking by actual applicants on personality tests: A meta-analysis of within-subjects studies
Supports average high-stakes versus low-stakes response differences across 20 within-person applicant studies, including order effects; it does not explain an individual score change.
- Retesting Personality in Employee Selection: Implications of the Context, Sample, and Setting
The accessible study abstract reports feedback-linked response strategies and greater predicted score change among actual applicants retesting an unproctored selection assessment.
- Retaking Employment Tests: What We Know and What We Still Need to Know
Summarizes broad employment retesting findings and states that causes of differences between initial and retest scores remain insufficiently delineated.
- Test Reliability—Basic Concepts
Defines reliability across occasions, editions, and raters and describes error of measurement and standard error of measurement for score interpretation.
Apply it to your work
Turn a changed score into work patterns you can observe
From this guide: If the retest leaves you unsure what changed in your own decisions or collaboration, compare specific examples across situations before drawing a career conclusion.
A personality score can point to a question, but it cannot settle what your work pattern looks like in practice. The Work Pattern Report offers a low-stakes self-reflection across decisions, planning, feedback, conflict, collaboration, change, and learning. Use it to name tendencies worth observing in your own examples, then bring those observations to a career conversation. It does not choose a role or evaluate employment suitability.
