In brief

Usually, no. Scores from two personality tests should not be placed on the same scale merely because both use numbers, percentiles, or labels such as low and high. A score is meaningful within the test that produced it: its items, scoring rules, construct definition, norm group, and intended use. A defensible comparison needs evidence that the tests measure the same or sufficiently similar construct, are used with comparable populations and purposes, and have been linked through an appropriate study. The result may be a prediction, an alignment, or a concordance table. Those are not automatically claims that the scores are equal or interchangeable. If no documented crosswalk exists, compare the reports at the level of constructs, evidence, and uncertainty rather than converting one number into the other.

The number is not the scale

Suppose one report gives a score of 72 and another gives a score of 6. The numbers do not tell you whether one person is higher, lower, or equal on a shared trait. One test may use a raw total, another a standardized score, and a third a percentile. A raw score is the total produced directly from scored responses. A standardized score is a transformed result whose distance from a reference mean is expressed in a chosen unit. A percentile rank describes the proportion of a specified comparison group scoring below a result. None of these formats, by itself, identifies the same underlying scale across tests.

The first question is not “Which score is bigger?” It is “What does each score mean in its own report?” Check the construct name, item content, scoring direction, possible range, norm group, date of the norms, and whether the result is norm-referenced or tied to a criterion. Two tests can both use the word confidence while asking about different behaviors or combining confidence with other content. A shared label is a starting point for investigation, not evidence of a shared ruler.

The Standards for Educational and Psychological Testing say that norms should refer to clearly described populations and support the intended interpretation. A percentile from one test cannot silently become a universal location on another test's scale.

First compare what the tests claim to measure

Before asking whether scores can be converted, compare the constructs. A construct is the psychological attribute or pattern a test is designed to represent. Look beyond broad names. One instrument may define assertiveness as speaking up and expressing disagreement. Another may treat it as a facet within a wider interpersonal style, or use items that also reflect social confidence. Those results may be related without being identical.

A useful comparison has four columns: construct definition, item content, score interpretation, and intended decision. Record what each test says the score represents, not what a familiar label suggests. A broad domain score should not be treated as interchangeable with a narrow facet simply because both are described as related to sociability. Nor should a self-reflection questionnaire be treated as a selection instrument because its report uses workplace language.

Score-linking guidance from the Educational Testing Service makes the same basic point in a different testing context: tests need to measure the same or highly similar constructs before their scores can be meaningfully linked. Conceptual resemblance can justify a comparison question, but it does not settle the measurement question.

Norms can make equal-looking results mean different things

A norm is information about how a score compares with a defined reference population. If two reports both show a result at the 70th percentile, that does not mean the tests measured the same amount of a trait. It means each result was above the scores of roughly 70 percent of its own stated reference group, subject to the report's method and rounding. If the groups, item sets, or construct definitions differ, the two percentiles answer different questions.

The same problem appears when reports use average bands. “High” may be a statistical description relative to a norm group, a threshold chosen for interpretation, or a broad editorial category. It is not automatically a common level across products. Check who was included in the norms, when the data were collected, whether the sample matches the person being assessed, and whether the report explains the precision of the norms. The testing standards call for norming documentation that identifies the sampled population, sampling procedures, testing dates, descriptive statistics, and precision.

Do not solve a norm mismatch by comparing the two averages. A mean of 50 on one standardized scale and a mean of 100 on another are conventions, not shared psychological quantities. A transformation can change the display without adding evidence that the transformed scores have the same meaning.

Two illustrated sheets with horizontal bars and geometric diagrams sit on opposite sides of a large question mark, above separate balance beams with rulers.
Two illustrated sheets with horizontal bars and geometric diagrams sit on opposite sides of a large question mark, above separate balance beams with rulers.

What a real crosswalk would require

A crosswalk is a documented relationship between scores on two instruments. Depending on the design, researchers may call the result linking, alignment, prediction, concordance, or equating. These terms should not be collapsed. A prediction model estimates one test's score from another and carries prediction error. An alignment may relate score levels or descriptors. A concordance table can show corresponding ranges for related tests. Equating makes the strongest claim: that scores can be treated as interchangeable under specified conditions.

For two personality tests, a credible study would need more than a convenient formula. It would normally need clear construct definitions, appropriate samples who take both measures under suitable conditions, transparent analysis, and evidence about uncertainty and relevant subgroups. The tests' purposes and score reliability also matter. If a score is unstable, a precise-looking conversion can give a false impression of precision.

The Standards require a clear rationale and supporting evidence for claims that scores may be used interchangeably, with technical information about the equating design, methods, examinees, and accuracy. ETS likewise warns that numerical operations alone do not create equating. A website table that maps one personality score to another without such documentation is an unsupported interpretation, not a common scale.

Correlation is useful, but it is not equivalence

Two tests can correlate and still produce non-interchangeable scores. Correlation asks whether people tend to be ordered similarly on two measures. It does not show that the numerical units match, that the tests have the same error, or that a particular result on one test has the same interpretation on the other. A high association can arise when tests share a broad theme while differing in emphasis.

When a report says that two instruments are “strongly related,” read the claim precisely. Does it mean the instruments correlate in one sample? Does it mean one result predicts the other within an interval? Does it support a decision threshold? Or does the provider claim interchangeability? Each statement requires different evidence. A correlation can support cautious evidence that measures overlap, but it does not authorize a score conversion by itself.

ETS's guidance states that even highly or statistically significant correlations do not make scores equivalent unless an equating approach supports that claim. Use a relationship between measures to understand overlap, not to manufacture a shared unit.

A brass balance scale stands between two illustrated profile sheets with colored bars; a pen, ruler, books, and leaves surround the scene.
A brass balance scale stands between two illustrated profile sheets with colored bars; a pen, ruler, books, and leaves surround the scene.

A worked comparison without a fake conversion

Consider an unnamed decision: a reader has completed two reports and wants to know whether a higher result on the second cancels out a lower result on the first. The reports use different score ranges and both mention “detail focus.” Instead of averaging or rescaling the numbers, make a comparison table with these questions: • What exact behavior or tendency does each report say the score represents? • Is the result raw, standardized, percentile-based, banded, or a facet score? • What norm group or reference rule gives the result its meaning? • Were the tests designed for the same purpose, such as reflection, coaching, development, or selection? • Is there a published crosswalk based on people who took both tests, and does it report uncertainty? • Does either report warn that the score should not be used for the decision you are considering?

If the answers show different constructs or missing documentation, the responsible conclusion is not that one test is wrong. The conclusion is narrower: the two numerical results cannot be interpreted as positions on one scale. You can still compare the descriptions. Both reports may point you toward reviewing how you handle detail-heavy tasks while disagreeing about the breadth or strength of that tendency. Use observable behavior to examine the question, not a converted number to settle it.

When a comparison is reasonable

A comparison becomes more defensible when the instruments are explicitly designed to assess the same construct for a closely related population and purpose, and when the provider supplies technical evidence for the claimed relationship. A published concordance may be useful as a limited correspondence if it names the sample, administration conditions, method, accuracy, and limits. Even then, read “corresponds to” according to the study's wording. It may mean an estimated relationship, not a shared psychological quantity.

One current example outside personality assessment shows the level of documentation involved. An ETS report on IELTS Academic and TOEFL iBT describes a concordance project in which test takers took both tests, the order was counterbalanced, score reports were verified, and equipercentile equating was conducted by an independent third party. That example does not validate a personality conversion. It illustrates why a serious crosswalk is a research project with a defined population and design, not a simple rescaling exercise.

For ordinary self-reflection, the safest useful comparison is often qualitative: identify overlapping constructs, note meaningful differences, and check whether both reports point to the same observable pattern. For coaching or development, use scores as prompts and verify them with behavior over time. For selection or other high-consequence decisions, require instrument-specific documentation and professional review rather than an informal conversion.

Two illustrated profile sheets with horizontal bars and small circular icons face each other across a vertical ruler; papers, leaves, books, and drafting tools surround them.
Two illustrated profile sheets with horizontal bars and small circular icons face each other across a vertical ruler; papers, leaves, books, and drafting tools surround them.

A practical report-reading checklist

Before placing two results side by side, write down this information from each report. If a field is absent, mark it as unknown rather than filling it with an assumption. 1. Construct: What does the instrument say it measures, and how narrowly or broadly is it defined? 2. Score type: Is the result raw, standardized, percentile-based, banded, or a facet? 3. Reference: Which norm group, criterion, or comparison rule gives the score meaning? 4. Purpose: Is the instrument intended for reflection, coaching, development, selection, or another use? 5. Evidence: Does the documentation support this interpretation and show reliability, validity, and relevant uncertainty? 6. Link: Is there a named, research-based crosswalk, and does it claim prediction, alignment, concordance, or interchangeability? 7. Decision: What would you do differently if the results agree, disagree, or remain incomparable?

If the last question leads to a consequential action, pause before treating either score as decisive. Return to the live topics library for further assessment-literacy guidance, and ask the publisher for the technical documentation behind any proposed conversion. The useful outcome is not always one combined number. It may be a clearer statement of what each report can and cannot tell you.

Questions readers ask

Can I convert two personality scores to percentiles and compare them?

Not safely by default. Percentiles are tied to each test's reference group and scoring method. Converting both results to percentiles does not show that the underlying constructs, norm groups, or units are the same. Use a documented crosswalk if one exists; otherwise compare the reports' constructs and limits.

Does the same trait name mean the same thing on two tests?

No. The same label can cover different item content, facets, scoring rules, and intended uses. Read each instrument's construct definition and technical documentation before treating the labels as equivalent.

Is a strong correlation enough to place scores on one scale?

No. Correlation can show that results are associated, but it does not establish equal units or interchangeable interpretations. A conversion needs evidence suited to the specific linking claim and should report its uncertainty.

What should I do when two personality reports disagree?

First check whether they measure the same construct, use comparable norms, and serve the same purpose. If not, disagreement may reflect different questions rather than a failed test. Treat both as limited sources of information and check the issue against observable behavior and the decision you need to make.

Sources and notes

  1. Standards for Educational and Psychological Testing

    Supports the requirements for clear norm populations, norm precision, and evidence before claiming score interchangeability.

  2. The Practice of Comparing Scores on Different Tests

    Explains distinctions among equating, concordance, norms, and comparisons across different test constructs and populations.

  3. Best Practices for Comparing TOEIC Speaking Test Scores to Other Assessments and Standards

    Supports the distinction between prediction, alignment, concordance, and equating, and warns against unsupported equivalence claims.

  4. Aligning Scores of Language Proficiency Tests: A Score Concordance Study Between IELTS Academic and TOEFL iBT

    Provides a documented example of cross-test concordance with matched test takers, counterbalanced order, verification, and independent equating.

Apply it to your work

Understand how you work before you choose what comes next.

From this guide: Carry this report-reading question into the work decision in front of you.

Build a private Work Pattern Report across ten workplace continuums, then compare the result with the demands of the role or environment in front of you.