Start by checking whether a self-rating and observer rating concern the same person, trait, period, and meaningfully similar questions and scales. These checks make a difference easier to discuss, but do not by themselves establish psychometric equivalence. Big Five research finds moderate, trait-varying convergence on average; a gap is a prompt to examine what each person could observe, not proof that either is wrong or more accurate in an individual case.
What does agreement between the two ratings mean?
Suppose a personality report says you are highly organized, while a colleague who has completed an observer form describes you as less consistent. Before deciding that one result is wrong, ask what the comparison can actually establish. A self-report is a person’s rating of their own tendencies; an informant report is another person’s rating of that target. Agreement, or convergence, means the two perspectives show some similarity. It does not mean that one score verifies the other as a fact.
A meta-analysis by Connolly and colleagues examined self and observer ratings of the Big Five, a widely researched set of broad personality dimensions. Its abstract reports corrected average correlations ranging from .46 for agreeableness (N=6,359 across 53 studies) to .62 for extraversion (N=7,725 across 50 studies); conscientiousness was .56 (N=6,754, 58 studies), emotional stability .51 (N=8,000, 55 studies), and openness .59 (N=5,333, 38 studies). The corrections account for scale alpha in self-ratings and inter-rater reliability in observer ratings. These are correlations across research samples, not percentages of people who agree or a measure of how close two scores should be for one reader. They show meaningful but incomplete average overlap in those studies, not a universal conversion rule for every personality report. [0]
A second distinction matters when reading your own pair of results: profile pattern versus score level. Pattern agreement asks whether the same traits appear relatively higher or lower across a set of traits. Level agreement asks whether both raters assign similar scores on a particular trait. For example, two raters might see a person as more planful than socially outgoing, yet differ in how high they rate both tendencies. That is an illustrative distinction, not a claim about a specific assessment result.
The two forms of agreement answer different questions. A similar profile can suggest that raters notice a shared pattern while disagreeing about its intensity. A level difference on one trait can coexist with broad pattern agreement. So inspect each trait separately and look at what the report actually displays. Do not treat two numbers as directly comparable merely because they carry the same trait label. A 2020 study of 3,253 people tested whether self and informant forms measured Five-Factor Model domains and facets comparably. Most domains and facets met the study’s threshold for comparing people’s relative standings, but Agreeableness and ten facets did not meet its recommended threshold for comparing mean scores. This study shows that evidence for one kind of comparison does not establish the other, and does not establish equivalence for every instrument. [3]
The fairest starting point is therefore modest: overlap is evidence that perspectives share some information, while disagreement signals that their observations or interpretations may differ. Neither result alone establishes a definitive trait level. The available meta-analytic evidence concerns Big Five measures across research samples; it does not validate a commercial report’s observer form or settle the meaning of an individual pair of scores.
Sources: The Convergent Validity between Self and Observer Ratings of Personality: A Meta-Analytic Review; Do Self-Reports and Informant-Ratings Measure the Same Personality Constructs?
What should I check before treating the gap as real?
Before interpreting a gap, check whether the ratings are practically comparable. Ask whether both concern the same person and trait definition; whether item wording, response options, scoring, reference period, and setting are close enough for the question you want to ask; and whether the scores have the same meaning. A raw score, standardized score, and percentile are different quantities, and a percentile refers to a specified norm group. Do not compare an observer raw total with a self percentile as if they shared a scale.
These are screening checks, not proof of psychometric equivalence. Matching labels or similar wording can make a conversation more focused, but a formal score comparison requires evidence that the forms measure the construct comparably for that purpose. The 3,253-person Five-Factor Model study found that its evidence differed by domain, facet, and comparison: relative-standing comparisons were generally supported, while mean-score comparisons were not supported for every domain and facet. Those results do not validate a different provider’s forms. If documentation is absent, use the pair to frame questions rather than treating the scores as interchangeable measurements. [3]
Access is another question, separate from score comparability. Connelly and Ones’ meta-analysis found that interaction frequency was associated with observer accuracy, and interpersonal intimacy was needed for substantial increases in other-rating accuracy. These are conditions that can matter; they do not make a familiar observer automatically accurate or objective, and the review cannot identify the better judge in one disagreement. [1]
Vazire’s study tested self, friend, and stranger ratings against behavioral criteria in a sample of 165 participants. Its trait-specific results differed by perspective: self-ratings were most informative for neuroticism-related traits, friends for intellect-related traits, and the perspectives were similarly informative for extraversion-related traits. That bounded result illustrates why access may vary with the trait; it does not establish a rule for every measure or person. [2]
For example, a person may report worry that coworkers never see, while a colleague may recall interruptions in meetings that the person did not notice. These are possible differences in access, not proof that either rating is correct. Keep the comparison exploratory if the forms or reference conditions are not established, and ask what specific observations each rating represents.
Sources: An Other Perspective on Personality: Meta-Analytic Integration of Observers’ Accuracy and Predictive Validity; Who Knows What About a Person? The Self–Other Knowledge Asymmetry Model; Do Self-Reports and Informant-Ratings Measure the Same Personality Constructs?
How can I use a disagreement without turning it into a verdict?
Use a remaining gap to find a behavior and setting worth discussing. Imagine an unnamed employee who rates their planning as strong, while a colleague’s observer rating is lower. Instead of concluding that the employee lacks insight or that the colleague is unfair, identify the evidence each may have used: Did one mean meeting deadlines, while the other meant sharing updates when plans changed? Did both think about the same projects and period? This is an illustration of how to ask, not a reported case or test result.
A practical sequence is: name the exact trait and score type; check form, scale, time frame, and setting; then ask for one recent example behind each rating. A concrete exchange might begin, “When you rated my follow-through lower, which recent situation were you thinking of? I was thinking about whether I completed the agreed work; were you thinking about how early I raised a delay?” The question helps distinguish different observations from different meanings of the same item.
Each viewpoint has limits. A self-rating can draw on private thoughts and experiences across settings, but it remains a person’s report rather than an unfiltered record. An observer can describe behavior they witnessed, but cannot directly report private experience and may have seen only a narrow slice. Evidence about group-level accuracy does not settle this individual pair.
For a work-related disagreement, keep the question specific: which behavior, in what setting and period, led each person to their rating? A difference is not by itself a hiring score, diagnosis, job recommendation, or verdict about character. If you also want a low-stakes way to reflect on your own tendencies, the live Work Pattern Report at /assessment offers a private self-report across ten work-pattern continuums. It has no norms or selection score and cannot adjudicate what another person observed.
The practical conclusion is to check the score meanings and conditions, then use one example to clarify what the ratings refer to. Big Five research supports partial average convergence, while evidence for formal comparison depends on the measure and the comparison being made. Ask: “Could we each name one recent example behind our rating, and check whether we were thinking about the same behavior and setting?”
Sources: The Convergent Validity between Self and Observer Ratings of Personality: A Meta-Analytic Review; An Other Perspective on Personality: Meta-Analytic Integration of Observers’ Accuracy and Predictive Validity
Questions readers ask
Does the observer’s rating count as more objective?
No. An observer can report visible behavior, but their view depends on what they have seen, their relationship with the person, and the setting. The self also has access to private experience that others cannot directly observe. Neither perspective is automatically the objective one.
Can I compare a self percentile with an observer raw score?
Not directly. A raw score is a total or item-based score; a percentile locates a score relative to a specified norm group. Compare quantities only when the report explains that both scores use compatible scales and interpretations.
What if the self and observer forms use different wording?
Check the provider’s documentation for evidence that the forms are equivalent in construct, scoring, and reference period. If that information is missing, treat the comparison as a conversation prompt rather than a psychometric comparison.
Does a rating gap prove I have a blind spot?
No. A gap may reflect different settings, examples, item meanings, or access to behavior and private experience. Ask what each rater had in mind before deciding what the difference means.
Sources and notes
- The Convergent Validity between Self and Observer Ratings of Personality: A Meta-Analytic Review
Supports the reported corrected average correlations between self and observer ratings across the Big Five dimensions, with trait-specific sample sizes (N) and numbers of studies (k) reported in the publisher abstract.
- An Other Perspective on Personality: Meta-Analytic Integration of Observers’ Accuracy and Predictive Validity
Supports the conditional role of interaction frequency and interpersonal intimacy in observer personality-rating accuracy.
- Who Knows What About a Person? The Self–Other Knowledge Asymmetry Model
Supports bounded sample findings that self, friend, and stranger ratings differed in accuracy by selected trait criteria.
- Do Self-Reports and Informant-Ratings Measure the Same Personality Constructs?
Reports a 3,253-person Five-Factor Model invariance analysis distinguishing relative-standing comparisons from mean-score comparisons.
Apply it to your work
Turn a broad work question into specific patterns
From this guide: A self-report can help you name your own decision and collaboration tendencies, while a conversation with an observer can clarify how particular behavior appeared in one setting.
If this comparison leaves you wondering how your own work tendencies combine across decisions, planning, feedback, and collaboration, the live Work Pattern Report can give you a structured starting point for reflection. It is a low-stakes self-report without norms or a hiring score, so use it to form questions about your experience rather than to settle another person's rating or choose a job.
