In brief

Look for the comparison behind the interpretation. A norm-referenced personality report tells you where your score stands relative to a specified reference group. A criterion-referenced report relates your score to a defined standard, domain, competency, or decision rule. The number alone does not identify the approach. A raw score, a standardized score, or a band can be used in either kind of interpretation. Read the report's scoring notes, norm information, descriptors, and intended-use statement before deciding what the result means.

Start with the question the report is answering

The most useful first question is not whether the score looks high. It is: high or low compared with what? A personality report can describe a measured tendency, compare that tendency with other people, relate it to a stated standard, or combine those steps. The interpretation depends on the reference point.

In a norm-referenced interpretation, the reference point is a distribution of scores from a specified group. The report may say that a score is at a particular percentile, above or below the norm-group average, or within a relative band. The result answers a comparative question: where does this person's score sit among the people represented by the norms?

In a criterion-referenced interpretation, the reference point is a defined criterion. That criterion might describe a competency, a set of behaviors, a level of proficiency, or a rule for classifying results. The result answers a different question: does the score support a statement about performance against that defined standard?

These are interpretations of scores, not necessarily two permanently different species of questionnaire. The same raw responses could be converted into a percentile for one purpose and into a descriptor tied to a stated standard for another. This is why a report's documentation matters more than its visual design or the word ‘standardized.’

What norm-referenced language looks like

Look for a named norm group, a percentile rank, a standard score, an average, or wording such as ‘compared with adults in the reference sample.’ A norm is a statistical frame of reference for interpreting a score. It is not a target that every person should reach, and it does not by itself tell you whether a tendency is desirable.

A percentile rank tells you the percentage of scores in a specified distribution that fall below a given score. It does not mean that the person answered that percentage of items correctly, performed that percentage better, or has that percentage of a trait. The relevant distribution is essential. A percentile based on a broad adult sample answers a different question from one based on a local group or a job-specific sample.

A report using norms should tell you enough about the reference group to judge whether the comparison is relevant. Check the population, age range, language, country or cultural setting when provided, date of data collection, and whether the sample was representative or simply a group of previous users. The NCME distinguishes norms for a defined reference population from user norms based on a self-selected or otherwise limited set of test takers.

For example, a report template might contain the labels ‘percentile rank,’ ‘reference sample,’ and ‘relative standing.’ Those labels point toward norm-referenced interpretation. They do not prove that the norms are appropriate, current, or strong enough for an important decision. They tell you what comparison to investigate next.

What criterion-referenced language looks like

Criterion-referenced language usually names the standard rather than a comparison group. Look for phrases such as ‘meets the defined level,’ ‘demonstrates the listed competency,’ ‘proficient in the assessed domain,’ or ‘falls above the cut score.’ A cut score is a specified point on a scale at which results are reported, interpreted, or acted upon differently.

The criterion is not automatically the cutoff. The criterion is the domain or standard being assessed. A cutoff is a decision point used to classify results in relation to that standard. A responsible report should explain what the levels mean, how the descriptors were developed, and why the threshold is suitable for the stated use.

Personality measures require particular care here. A general inventory may describe typical reactions, preferences, or interpersonal tendencies. That does not turn a score into proof that someone possesses a fixed competency, will behave in one way, or is suitable for a role. If a report uses words such as ‘ready,’ ‘safe,’ ‘fit,’ or ‘effective,’ ask what observable behavior or external outcome those words are meant to represent and what evidence supports that interpretation.

A criterion-referenced statement can be useful when the report defines the target clearly. ‘The responses indicate a pattern consistent with the described level on this scale’ is narrower than ‘you are an excellent leader.’ The first points back to a defined score interpretation. The second makes a broad claim that needs separate evidence.

Use the report as a small evidence trail

To identify the approach, read the report in this order: score label, comparison statement, reference information, level descriptor, cutoff explanation, and intended use. You are looking for the report's chain from response to conclusion.

First, identify the score type. ‘Raw score’ usually means the direct total or combination of item responses. A standardized score is transformed onto a scale whose meaning depends on its reference system. A percentile is explicitly comparative, but it is still only as informative as the distribution behind it. A band such as low, typical, or high may be norm-based, criterion-based, or an unexplained editorial category.

Next, underline the verbs. ‘Compared with,’ ‘ranked among,’ and ‘relative to’ signal a norm comparison. ‘Meets,’ ‘demonstrates,’ ‘qualifies,’ and ‘below the required level’ signal a criterion or decision rule. Then look for the noun after the verb. Is the score compared with people, or with a defined domain and its descriptors?

Finally, find the technical manual, user guide, or scoring notes. The NCME describes these documents as sources for a test's purpose, appropriate uses, administration, scoring, normative data, and interpretation. If the report gives a label without explaining its reference point, record that as missing information rather than guessing. An attractive narrative is not a substitute for a scoring explanation.

Two cream-colored sheets display abstract charts: a bell-shaped curve with marked points on the left and four horizontal scales with dots on the right. A magnifying glass, books, pencils, ruler, and triangular ruler sit around them.
Two cream-colored sheets display abstract charts: a bell-shaped curve with marked points on the left and four horizontal scales with dots on the right. A magnifying glass, books, pencils, ruler, and triangular ruler sit around them.

A worked comparison without treating either result as a verdict

Consider this example: two report panels use the same broad scale name. Panel A says, ‘Your score is higher than the scores of many people in the adult reference sample.’ Panel B says, ‘Your score falls within the report's defined band for the described competency.’ Panel A is norm-referenced in that sentence. Panel B is criterion-referenced in that sentence.

The panels answer different questions. Panel A can help a reader understand relative standing. It cannot, on its own, establish that the person has a useful skill, will act consistently across settings, or should be selected for a job. Panel B can connect a score to a defined level, but its usefulness depends on the quality of the domain description, the way the threshold was set, and evidence that the interpretation fits the intended population and purpose.

Now imagine the report includes both panels. That is not necessarily a contradiction. A percentile can show relative position while a descriptor explains what the score range is intended to suggest. The reader should keep the claims separate. ‘Relatively higher than the reference group’ is not the same claim as ‘meets a behavioral standard.’

This comparison also shows why ‘high’ is incomplete language. High relative to whom? High enough for what? A high score can be interesting for self-reflection without being a pass mark, and a score near a criterion boundary may warrant careful interpretation rather than a sharp label. The report should make those limits visible.

Do not confuse norms with validity or reliability

A norm tells you how a score is being compared. Reliability concerns the consistency or precision of scores under specified conditions. Validity concerns whether accumulated evidence and theory support a particular interpretation for a particular use. These are related parts of assessment literacy, but they are not interchangeable.

A report can have carefully described norms and still lack evidence for a broad claim about work performance. It can also report a reliability coefficient without showing that the scale measures the intended construct or that a cutoff supports a consequential decision. The National Academies emphasizes that reliability, validity, and fairness all matter, with the required evidence depending on the intended use.

Self-report adds another interpretive layer. Responses describe what a person endorses about themselves under the conditions of that questionnaire. They may be affected by wording, context, self-knowledge, or a wish to present oneself favorably. That does not make every self-report useless. It does mean that a report should not silently turn a response pattern into an observed behavior or a diagnosis.

For a low-stakes reflection exercise, a norm comparison may help generate a question to examine in daily life. For coaching, the result may be one input among conversation and observation. For selection, licensing, or clinical work, the required evidence and professional safeguards are more demanding. The same score format does not make those uses equivalent.

The practical decision: what should you do with the report?

If the report is norm-referenced, use it to frame a relative comparison and inspect the norm group's fit. If it is criterion-referenced, inspect the standard, the level descriptors, the cutoff rationale, and the evidence for applying that standard to your purpose. If you cannot find either a clear reference group or a clear criterion, treat the interpretation as under-documented.

Then write one observable follow-up question. Instead of ‘Am I a poor communicator?’ ask, ‘In which situations do I hold back information, and what happens when I ask for clarification?’ Instead of ‘Does this prove I am suited to the role?’ ask, ‘What behavior does the report claim to represent, and what independent evidence would be relevant?’ These questions keep the report in its proper role: a structured source of information, not a final judgment.

Use this checklist before sharing the result or acting on it: identify the score type; locate the reference group or criterion; check whether the population fits; read the descriptor rather than only the label; note any measurement uncertainty; separate self-report from observed behavior; match the interpretation to the intended use; and avoid decisions that the documentation does not support.

The decision point is simple. If you can state what the score is compared with, what that comparison supports, and what it does not support, you are reading the report responsibly. If you cannot, pause and ask the publisher or qualified user for the scoring documentation. For more guides on reading scores, norms, and report sections, continue with the live [topics library](/topics).

Questions readers ask

Can a personality report be both norm-referenced and criterion-referenced?

Yes. A report may show a percentile relative to a reference group and also map a score to a defined descriptor or cutoff. Read each statement separately because the two interpretations answer different questions.

Does a percentile mean that I have that percentage of a personality trait?

No. A percentile rank describes the percentage of scores below a given score in a specified distribution. It is a relative position, not a percentage of a trait and not a percentage of correct answers.

Is a low or high personality score good or bad?

Not by itself. A norm-referenced score describes relative standing, while a criterion-referenced label depends on the defined standard and purpose. Meaning also depends on the construct, context, uncertainty, and intended use.

What if the report does not name its norm group or criterion?

Treat the interpretation as incomplete. Ask for the technical manual or scoring guide, including the reference population, criterion definitions, cutoff rationale, intended use, and relevant reliability and validity evidence before relying on the result.

Sources and notes

  1. NCME Assessment Glossary

    Supports the definitions of norm-referenced and criterion-referenced score interpretation, norms, percentiles, cut scores, technical manuals, reliability, validity, and user norms.

  2. Testing in the Public Service of Canada

    Supports the distinction between scores tied to defined behaviors and scores interpreted against relevant others, plus guidance on documenting norms and cut scores.

  3. The Standards for Educational and Psychological Testing

    Supports the professional standards framework for interpreting educational and psychological test scores and evaluating intended uses.

  4. Supporting Students' College Success: Assessment Methods for Intrapersonal and Interpersonal Competencies

    Supports matching assessment selection and interpretation to intended use, and considering reliability, validity, fairness, self-report limits, and context.

  5. Professional Practice Guidelines for Personality Assessment

    Supports using professional guidance for ethical personality assessment, appropriate applications, diversity considerations, and consumer understanding.

  6. Ten-Item Personality Inventory Technical Manual

    Provides a concrete published personality-inventory example showing raw scale averages converted to z scores and percentiles using stated norms.

Apply it to your work

Understand how you work before you choose what comes next.

From this guide: Proceed with reflection or a proportionate next question only when the report makes its reference point and intended use clear; otherwise request documentation before relying on it.

The Work Pattern Report maps how you decide, plan, collaborate, handle conflict, adapt, and learn across 100 workplace situations. Use the result to ask sharper questions about a role’s demands. It is a private reflection tool, not a job recommendation or hiring score.