Compare the assessments as measurement tools before comparing their results. Start with the decision you want the report to inform, then check whether both instruments measure the same construct, use comparable scoring and norm groups, and provide reliability and validity evidence for your intended use. A score of 70 on one assessment is not automatically higher, stronger, or more trustworthy than a score of 70 on another. If the instruments use different definitions, scales, or reference groups, treat the reports as two perspectives rather than placing them on one shared ranking. The useful question is not which test gives the better person-description. It is which report supports the specific conclusion you need, with the least uncertainty and the clearest limits.
Begin with the decision, not the scores
Suppose one report says you are high on a social trait and another places you near the middle. Before deciding that one test is wrong, write down what you are trying to decide. Are you choosing a tool for self-reflection, planning a coaching conversation, comparing a change over time, or making a consequential work decision? The same assessment may be suitable for one purpose and poorly supported for another.
This first step prevents a common mistake: treating a personality report as a general verdict about a person. The American Educational Research Association, American Psychological Association, and National Council on Measurement in Education describe validity as support for an intended interpretation and use of scores, not as a permanent property that makes every conclusion acceptable. The purpose therefore belongs at the top of your comparison sheet.
Write one sentence such as, ‘I want a useful prompt for reflecting on how I approach unfamiliar group work.’ That is narrower and safer than, ‘I want to know which test reveals my real personality.’ If the proposed use is hiring, selection, or a clinical decision, the evidence and safeguards required are much more demanding than for private reflection. A general personality report is not a diagnosis, and neither report should be used to make a medical conclusion.
Check whether the two assessments measure the same thing
Next, compare the construct: the attribute the assessment claims to measure. Do not rely on similar labels alone. Two reports may use words such as ‘communication,’ ‘confidence,’ or ‘openness’ while defining them through different item content, facets, scoring rules, or theoretical models. One may ask about typical behavior; another may ask about preferences, values, or reactions in specific situations. Their outputs can both be internally coherent without being interchangeable.
Read the technical description for each instrument and record four details: the construct name, its facets or subscales, the response format, and the interpretation promised by the publisher. Also note whether the instrument is self-report, observer-rated, performance-based, or a mixture. A self-report answer describes how a person reports themselves under the test instructions. It does not directly observe every behavior in every setting.
An especially useful comparison asks what each assessment leaves out. If one broad score combines several facets and the other reports only one narrow facet, disagreement may reflect different coverage rather than a contradiction. The ITC guidance recommends checking the scope and representativeness of test content before selecting an assessment. That is a more informative question than asking which report sounds more flattering.

Put the score formats on separate lines
Make a small table for each assessment, but do not combine the numbers yet. Record the raw score, the reported scale, any percentile, the band or category, and the norm group named in the report. A raw score is usually the count or sum produced before a transformation. A percentile describes the percentage of a reference group scoring at or below a position; it is not a percentage of the trait and it is not a pass mark. A band is a category created by a stated rule, which may hide small differences near a boundary.
Then ask whether either publisher provides a documented crosswalk between the instruments. A crosswalk is evidence-based information relating scores from two assessments. It is not created by noticing that both tests use a 0-to-100 display, by matching percentiles from unrelated norm groups, or by subtracting one result from another. Educational Testing Service explains that score linking requires a common scale and that equating, the strongest form of linking, is reserved for tests with the same construct and appropriate statistical conditions. The NCME glossary likewise treats comparability as a degree achieved through a linking procedure, not as an automatic feature of numerical labels.
For personality reports, a formal crosswalk may not exist, and that is normal. If it does not, compare patterns and interpretations at the construct level instead. You might say, ‘Both reports raise a question about how I handle unfamiliar social demands,’ while avoiding, ‘The second test proves I am 18 points more outgoing.’ The first statement is a cautious synthesis; the second invents a shared ruler.
Separate reliability from validity
A responsible comparison needs at least two different evidence questions. Reliability concerns consistency or precision. Would scores tend to be similar across repeated administrations, items, raters, or forms under relevant conditions? Measurement error is the expected imprecision around an observed score. A small difference between two results may be less meaningful than it looks if the reports do not provide enough precision for that use.
Validity asks a different question: does the available evidence support the interpretation and use you want to make? A report can describe a score consistently while still offering weak support for a particular conclusion. For example, evidence that a scale is stable does not by itself show that it predicts workplace performance, explains a relationship, or identifies a clinical condition. The APA’s guidance notes that unreliable information cannot support a valid inference, but that reliability alone is insufficient.
For each assessment, look for evidence tied to the actual claim. Check whether the documentation reports internal consistency, test-retest results, or another reliability estimate, and whether the relevant population resembles the people who will use the report. Then look for validity evidence about the construct and intended purpose. Prefer a technical manual, validation study, systematic review, or professional evaluation over testimonials, attractive sample reports, or a publisher’s broad promise. If one assessment publishes only a coefficient without explaining the sample, construct, or use, record that as incomplete evidence rather than treating the number as a quality badge.

Inspect the norm group and the uncertainty
A norm-referenced score tells you how a result compares with a stated reference group. That group may differ by age, language, country, education, occupation, or the date on which the norms were collected. The key question is not whether the group is large in the abstract. It is whether the comparison group is relevant to the interpretation being offered. The ITC guidelines caution against drawing conclusions from norms that are irrelevant or outdated.
Look for who was included, how the sample was recruited, which language version was used, and whether separate norms are supplied. If the report gives you a percentile but does not identify the reference group, the number has less meaning than its polished presentation suggests. If the two assessments use different norm groups, their percentiles should remain separate. A 60th percentile on one report and a 40th percentile on another do not establish a direct disagreement about you.
Also look for a confidence interval, standard error of measurement, score band, or other explanation of uncertainty. If none is provided, do not manufacture one. Instead, avoid over-reading a result near a cutoff or a small difference between retakes. Treat the report as a signal for reflection or further inquiry, with confidence matched to the evidence.
Check language, administration, and fairness
Different testing conditions can change what a score means. Compare the language version, instructions, completion time, accessibility arrangements, delivery mode, and whether the assessment was completed privately or under pressure. A person answering in a second language, using an adapted format, or responding after a major situational change may not be represented by the same evidence used to support the original interpretation.
The ITC translation and adaptation guidance treats language and cultural adaptation as an evidence question, not as a cosmetic rewrite. The ITC test-use guidance also asks users to consider linguistic, cultural, and situational differences, as well as the setting and recipient of the results. These checks do not prove that an assessment is unfair, but they can show where a confident comparison would be premature.
For personal use, write down any condition that could have shaped your answers: unfamiliar wording, fatigue, a recent conflict, uncertainty about whether someone else would see the result, or a change in the instructions. This is not an excuse to discard an unwanted score. It is context for deciding how much weight to give it. In work settings, ask who can access the report, what decision it informs, and whether the assessment has evidence for that population and purpose.

Make a modest conclusion and use the reports well
After the checks, classify the comparison rather than crowning a winner. If the instruments measure closely related constructs, use similar procedures, name relevant populations, and provide evidence for the same purpose, their patterns may offer converging information. If they measure related but distinct constructs, use the reports as complementary prompts. If their definitions, norms, or evidence are unclear, keep the conclusion tentative. If the proposed use is high stakes and the technical support is missing, do not use either report as the deciding evidence.
A practical report-reading checklist is short: What does each assessment measure? Who was the norm group? What does each score actually mean? What reliability evidence is relevant? What validity claim is supported? What uncertainty or context could change the interpretation? Can I describe the result without turning a tendency into an identity label? If you cannot answer these questions, the next step is to request the technical documentation or choose a more transparent instrument. The live Personality Report topics library can help you examine norms, percentiles, reliability, and validity one question at a time.
The most useful conversation is specific. You might ask a coach, provider, or colleague: ‘Both reports describe this area differently. What construct does each one measure, which norm group supports the comparison, and what decision is the evidence strong enough to inform?’ That question keeps the discussion open while placing the burden of interpretation where it belongs: on the report’s definitions, evidence, and limits, not on a label about who you are.
Questions readers ask
Can I compare the percentiles from two personality assessments?
Usually not as if they were measurements on one scale. Percentiles depend on each assessment’s norm group, scoring method, and construct. You can compare the underlying descriptions cautiously when the constructs and intended uses are similar, but do not treat a higher percentile on one report as proof of a higher trait level than a lower percentile on another.
What should I do if two personality reports disagree?
Check whether they measure the same construct, use comparable instructions and norm groups, and provide evidence for the same purpose. Then consider measurement uncertainty and the context in which you answered. The disagreement may reflect different questions rather than one report being false. Use the overlap as a reflection prompt and keep any conclusion modest.
Sources and notes
- Standards for Educational and Psychological Testing
The joint AERA, APA, and NCME standards provide the framework for evaluating score interpretation, validity, reliability, fairness, and test use.
- Professional Practice Guidelines for Occupationally Mandated Psychological Evaluations
APA guidance distinguishes reliability from validity and says assessment evidence should fit the population and intended inference.
- ITC Guidelines on Test Use
The International Test Commission guidance covers construct coverage, norm groups, reliability, validity, language, context, confidentiality, and responsible interpretation.
- The Practice of Comparing Scores on Different Tests
ETS explains why scores from different tests need a documented linking or concordance basis and why matching numerical displays is not enough.
- Glossary: Linking and Equating
The NCME glossary defines linking as relating scores so they have a common relative meaning and notes that comparability varies by linking method.
Apply it to your work
Understand how you work before you choose what comes next.
From this guide: Carry this report-reading question into the work decision in front of you.
Build a private Work Pattern Report across ten workplace continuums, then compare the result with the demands of the role or environment in front of you.
