In brief

A profile agreement statistic describes how closely two sets of scores line up across the same traits. For one person, it may summarize a self-rating against an observer rating, or one profile against a reference pattern. The number establishes correspondence under the report’s stated formula. It does not, by itself, show that either profile is accurate, stable over time, or useful for a particular decision. Before interpreting it, identify the two profiles, the traits compared, the coefficient, and any adjustment for the average pattern. A high raw correlation and a high norm-adjusted index answer related but different questions.

What is the statistic comparing?

A profile is a set of scores across several measured traits, considered together as a pattern. A profile agreement statistic compresses the similarity of two such patterns into one number. In the common self–observer case, one profile comes from the person’s answers and the other from someone who rates that same person. Both need to cover the same traits for the comparison to be interpretable.

This is different from asking whether people who score high on one trait also tend to be rated high on that trait. That trait-level question compares people across a sample. Profile agreement instead compares two raters across traits for a particular target. Allik and colleagues describe these as trait-centered and person-centered approaches, respectively, and analyzed ratings of 4,115 targets across four cultural-language samples. Their study shows why the unit matters: a group pattern across people and a particular pair’s profile are not the same result.

A simple illustration: imagine that both profiles place a person relatively higher on planning than on spontaneity, and relatively lower on social energy than on curiosity. A similarity coefficient may capture that ordering even if the two raters use different overall score levels. This is an illustration of the calculation idea, not an assessment result. Other coefficients also consider score levels or distances, so the report’s label alone is not enough.

The narrow conclusion is therefore conditional: two specified profiles corresponded to a stated degree under a stated method. Ask who supplied each profile, whether they rated the same person, which dimensions were included, and whether the value belongs to one pair or summarizes many pairs. A sample average cannot tell you the agreement statistic for your own report.

Sources: How are personality trait and profile agreement related?; Generalizability of self–other agreement from one personality trait to another

What kind of agreement does the number count?

The formula decides what the number rewards. A Pearson profile correlation focuses on whether scores rise and fall together across traits. If one rater tends to give higher scores across every trait, that overall elevation need not lower the correlation. A double-entry intraclass correlation, by contrast, is sensitive to both the shape of the pattern and its elevation. Other profile indices use other rules. These are not interchangeable merely because all are called agreement statistics.

A second choice is whether the comparison uses raw scores or adjusts for the average pattern in a reference group. Normativeness means resemblance to that average pattern. If the traits are standardized within a sample, common average differences between traits are removed; the comparison then focuses more on how each profile differs from the norm. In a four-sample study using NEO-family measures, Allik and colleagues found average raw profile correlations of .59 and average normalized correlations of .44. Those are study averages for those samples and methods, not personal benchmarks for an unspecified report.

This adjustment addresses a real interpretive issue. Two profiles can resemble each other partly because both reflect the common pattern in the population, even without much individual information. Yet the common pattern may also be meaningful for some questions. Removing it changes the question rather than automatically correcting the answer. The same study found differences among coefficients and transformations, while some coefficients were highly correlated within particular samples. That result supports checking the actual method, not assuming every method will agree in every context.

Use a raw index when the intended question includes overall resemblance across the measured dimensions. Use a norm-adjusted index when the question concerns matching in the person’s deviations from the reference pattern. Neither choice is universally better. Ask for the reference group and transformation, and check whether the report explains why that choice fits its interpretation. A larger value is not automatically stronger evidence if its recipe differs.

Sources: Generalizability of self–other agreement from one personality trait to another; The Counterpoint of Personality Assessment: Self Reports and Observer Ratings

Does agreement show that the report is accurate?

No. Agreement means that two ratings correspond under a defined comparison. Accuracy is a stronger claim: that the description captures relevant attributes of the person. Funder and West’s review treats consensus, self–other agreement, and accuracy as distinct ideas. McCrae’s discussion of self-reports and observer ratings makes a practical point: when ratings disagree, further information can help resolve the different explanations they suggest. Agreement can be informative, but it is not an independent check on truth by itself.

For example, two raters could share a mistaken assumption, use the same limited observations, or interpret an item in the same way. Their profiles could align while missing how the person behaves in another setting. Conversely, disagreement can reflect different vantage points: a colleague may see how someone handles deadlines, while the person has better access to private worry or effort. These are plausible mechanisms for interpretation, not claims about any particular pair.

Agreement is also not the same as reliability or validity. Reliability asks whether scores are consistent under specified repetitions or conditions. Validity concerns whether evidence supports a particular interpretation and use. A profile match does not provide test–retest evidence, a measurement-error estimate, or proof that the report predicts behavior. If a report makes a decision claim, the supporting evidence must address that claim and the intended population and context.

The counterpoint matters: agreement is not worthless. In research or careful feedback, multiple perspectives can reveal convergence and highlight where further observation is useful. McCrae’s paper describes using self and observer reports together and recommends seeking further information when they conflict. For personal reflection, treat agreement as a prompt to check examples, not a verdict. For consequential work decisions, ask for evidence designed for that use rather than relying on one profile statistic.

Sources: How are personality trait and profile agreement related?; Consensus, self-other agreement, and accuracy in personality judgment: an introduction

What should I ask before acting on my result?

Read the statistic as a method-bound comparison. First, establish its operands: whose reports are compared, whether they concern the same target, and whether they were completed at the same time or under comparable conditions. If the report compares a person’s profile with a model or norm instead of another rating of that individual, it answers a different question from self–observer agreement.

Second, request the coefficient’s name or formula and the score preparation. Does it compare rank-like shape, absolute level, or both? Were scores standardized, centered on a reference group, or otherwise transformed? The 2010 Allik study illustrates why these questions matter: Pearson correlation was insensitive to profile elevation and scatter, while the double-entry ICC also reflected elevation; normalization changed the study’s mean values. Its specific coefficients and NEO samples should not be transferred as cutoffs for another instrument.

Third, look for the comparison basis. If the report calls a result high, what reference distribution supports that label, and does it match the instrument, population, and profile length? McCrae’s profile-agreement work proposes interpretation guidance for a particular index and NEO-based data. That does not make its conventions universal. A label without its comparison basis is less informative than the number may appear.

Finally, match the conclusion to the evidence. For a low-stakes conversation, note where two views converge, then ask for concrete examples and context where they differ. Keep more than one explanation open. If someone claims the statistic proves accuracy, consistency, job fit, or future behavior, ask for separate evidence that directly tests that claim. The verdict is narrow but useful: profile agreement establishes correspondence between defined patterns. Its practical value lies in showing where to look next, not in settling what a person is like.

Sources: How are personality trait and profile agreement related?; Generalizability of self–other agreement from one personality trait to another; The Counterpoint of Personality Assessment: Self Reports and Observer Ratings; Consensus, self-other agreement, and accuracy in personality judgment: an introduction

Sources and notes

  1. How are personality trait and profile agreement related?

    Defines trait and profile agreement and clarifies that individual profile agreement compares two judges across traits for a target.

  2. Generalizability of self–other agreement from one personality trait to another

    Compares profile coefficients and raw versus normalized self-observer agreement in four samples using specific NEO-family measures.

  3. The Counterpoint of Personality Assessment: Self Reports and Observer Ratings

    Explains uses of combined self and observer reports and recommends further information when the ratings disagree.

  4. Consensus, self-other agreement, and accuracy in personality judgment: an introduction

    The accessible PubMed abstract distinguishes consensus, self-other agreement, and accuracy as separate concepts.

Apply it to your work

Turn a score pattern into a clearer work conversation

From this guide: If two descriptions of your work style differ, the useful next step is to identify which recurring situations each captures.

A profile statistic cannot resolve every disagreement about how you plan, decide, handle feedback, or respond to change. The Work Pattern Report offers a low-stakes self-reflection across these work continuums, so you can turn a broad question into specific observations to discuss. It is not a normed or validated employment score and does not recommend a job.