In brief

A self-report and an observer report can have similar average scores across a group while disagreeing for a particular person because group averages and individual agreement answer different questions. A mean compares typical score levels; it does not show whether each pair of ratings matches. For an individual, the difference may reflect what the person knows about their inner experience, what the observer has seen, the trait and setting being rated, or measurement uncertainty. Check whether the reports are comparable, then use the gap to ask about a specific behavior. Neither report automatically tells the whole story.

How can the averages match while one person's reports do not?

The claim under review is that similar self-report and observer-report averages mean that the two sources should agree for any one person. It sounds plausible because both reports use the same trait labels and may ask about the same target. But the conclusion does not follow. An average describes a group; an individual discrepancy describes one pair of scores.

A group mean is the arithmetic average of scores. If self-ratings average 50 and observer ratings also average 50, that does not mean every person received 50 from both sources. Some self-ratings can be above their observers' ratings and others below, with the differences offsetting one another. This simple illustration is hypothetical, not study data.

A second statistic answers a different question. A rank-order correlation asks whether people rated relatively higher by themselves also tend to be rated relatively higher by observers. Even a positive association allows score gaps: it describes ordering across people, not exact matches within each pair. Individual agreement asks how closely one person's two ratings align. These three ideas, mean level, rank order, and pair agreement, should not be treated as interchangeable.

[Kim, Di Domenico, and Connelly's meta-analysis](https://pubmed.ncbi.nlm.nih.gov/30481113/) compared Big Five self-report and informant-report means across 152 samples and 33,033 people. The authors found little average difference overall, with an average standardized mean difference of δ = −.038; moderate differences appeared when self-ratings were compared with stranger ratings. That is evidence against a broad claim that people generally rate themselves more positively than informants do. It is not evidence that every person's scores matched, or that all informants agreed with each target.

The comparison also has a defined scope: Big Five reports in the studies included in that meta-analysis. It does not establish the behavior of every instrument, trait scale, observer relationship, or report format. If a report pair differs, first check whether both forms cover the same trait, use matching scale direction, refer to the same time period, and invite judgments about the same context. Similar labels alone do not guarantee a like-for-like comparison.

So an individual gap does not contradict the group finding. The group result says that self and informant score levels were generally similar on average in the studied samples. It leaves room for a particular pair to differ, and it does not identify which score better describes that person's behavior.

Sources: Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means

What can make one person's ratings diverge?

Self and observer ratings draw on different access to information. A person can report private thoughts, intentions, or feelings that another person cannot directly see. An observer can describe visible actions in shared situations, including patterns the person may not notice. These are different viewpoints, not a built-in hierarchy in which one is always more accurate.

The distinction appears in research on individual trait and profile agreement. Trait agreement concerns whether a person and observer give similar ratings on one dimension. Profile agreement concerns whether the pattern across several dimensions resembles the other person's pattern. A pair could be close on one trait but not another; a profile could have a similar shape while the overall score levels differ. One kind of agreement cannot stand in for all the others.

[Allik and colleagues](https://www.ppw.kuleuven.be/okp/_pdf/Allik2010GOSAF.pdf) analyzed four samples from Estonia and Belgium using NEO-family measures. They found that self–other agreement did not generalize uniformly from one trait to another, and agreement contributions across traits were only moderately related to profile agreement. Their result supports a narrow conclusion: in those samples and measures, matching on one trait did not guarantee matching on the rest of a person's profile. It does not establish that the same pattern or size of differences applies to every assessment or population.

Who is doing the observing matters too. In a study of 184 targets rated by themselves, acquaintances, parents, and strangers, [Funder, Kolar, and Blackman](https://pubmed.ncbi.nlm.nih.gov/7473024/) reported stronger self–other agreement for acquaintances than for strangers. Knowing someone in a shared context enhanced agreement in one part of the study but was not necessary for agreement. Familiarity can give an observer more relevant examples, yet it cannot certify that a particular observer is right. A close relative may see home behavior; a colleague may see work behavior. Neither necessarily sees every setting.

A remaining explanation may lie in the measurement rather than either person's judgment. If an item is broad, such as describing whether someone speaks up, the self and observer may picture different meetings or different meanings of 'often.' Different time frames, opportunities to observe, and ordinary measurement error can also affect scores. Those are sensible questions to check, not causes diagnosed by the three cited studies for any one reader's discrepancy.

This is why the simple verdicts, 'the person lacks insight' or 'the observer sees the real person,' go beyond the evidence. A difference can be informative because it points to a place where perspectives or contexts diverge. On its own, it does not prove self-deception, observer bias, or a stable flaw.

Sources: Generalizability of self–other agreement from one personality trait to another; Agreement among judges of personality: interpersonal relations, similarity, and acquaintanceship

How should you use a difference in a real report?

Treat the mismatch as a prompt for a specific check, not a contest to decide who is correct. First compare the report details: the named trait or facet, instrument and version, score direction, reference period, and any context named in the instructions. If those differ, the apparent disagreement may partly be a comparison problem rather than a meaningful conflict.

Next, name the precise scale or statement that differs. Avoid turning one score into a global description of character. Ask what observable behavior the self-rating represents and what behavior the observer had in mind. For example, a person may feel hesitant before speaking in a meeting yet be seen by a colleague as decisive once they have stated a position. That is an illustration of how internal experience and visible action can diverge, not a documented case or a claim about any particular assessment result.

Then check the observation window. Does the observer see the person in the situation the report asks about? Are there repeated examples, or only one striking event? If the question matters for coaching or self-reflection, a concrete example and a second suitable perspective may clarify what each score captures. More ratings are not automatically better: their value depends on whether those people have relevant opportunities to observe the behavior and understand the question.

This sequence is practical guidance derived from the evidence, not a validated scoring rule. The studies show that mean similarity, individual agreement, trait agreement, profile agreement, and observer familiarity are distinct considerations. They do not provide a formula for resolving an individual disagreement. If accounts remain different, preserve that uncertainty and frame a follow-up question around context: 'What happens when I have to make a decision with limited discussion?' is more useful than 'Which report is right?'

The verdict is therefore qualified but clear: average agreement does not erase a person's discrepancy, and the discrepancy does not invalidate either report. It tells you that the two sources may be summarizing different information, or that the reports need a closer like-for-like comparison. General personality reports can support reflection or a coaching conversation; a score gap is not a diagnosis or a stand-alone hiring, promotion, or job-fit verdict.

If a recurring work disagreement is the concern, the next useful step is to choose one situation and observe what happens: who had information, when a decision was made, and how others responded. The live Work Pattern Report offers a low-stakes self-report across work tendencies as a way to organize reflection on decisions, feedback, conflict, and collaboration. It is not normed or validated for employment decisions, and it cannot tell which observer is right. Its role is to help turn a broad friction point into questions you can examine in real situations.

Sources: Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means; Generalizability of self–other agreement from one personality trait to another; Agreement among judges of personality: interpersonal relations, similarity, and acquaintanceship

Sources and notes

  1. Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means

    Supports the group-level comparison: 152 Big Five samples had very similar mean self and informant scores overall, with stranger-report differences; it does not establish individual pair agreement.

  2. Generalizability of self–other agreement from one personality trait to another

    Supports the distinction between trait agreement and profile agreement and reports that individual self–other agreement varied across traits in four Estonian and Belgian NEO-family samples.

  3. Agreement among judges of personality: interpersonal relations, similarity, and acquaintanceship

    Supports the bounded claim that acquaintances showed stronger self–other agreement than strangers among 184 targets in this study, not that a specific close observer is always accurate.

Apply it to your work

Turn a work-style disagreement into a question you can observe

From this guide: If different reports or colleagues describe the same work situation differently, identify the decision, feedback, or collaboration behavior each person is seeing.

A report can help name recurring tendencies, while a specific work exchange shows when they appear and what the setting asks of you. The Work Pattern Report gives you a low-stakes self-reflection across decisions, feedback, conflict, collaboration, and change. Use it to organize observations and questions for a coaching conversation; it does not verify another person's view or recommend a job.