In brief

Don’t choose a winner from the disagreement alone. Check what the report’s wording or score means, then ask the observer for a specific example. Compare whether both accounts concern the same behavior, setting, time period, and evidence. A repeated, relevant observation may qualify the report’s description for that setting; neither account alone establishes your whole work style.

What should you do when the accounts differ?

When a personality report describes one work habit and someone you trust describes another, the disagreement alone does not tell you which account to accept. Inspect four things separately: the report’s exact sentence, what its score or narrative means, the observer’s concrete example, and the behavior claim both accounts seem to make. Then choose a bounded outcome for that interpretation: retain it if the same behavior is supported; qualify it to a particular setting or period if the accounts describe different conditions; or set it aside for now if the report’s meaning or the available example is too unclear. A specific repeated observation can expose a blind spot, but it needs to be examined rather than treated as an automatic verdict.

What does agreement between self and observer ratings show?

A self-report is a rating a person gives about their own personality; an observer rating is a judgment made by someone else who knows or has encountered that person. In “The Convergent Validity between Self and Observer Ratings of Personality: A meta-analytic review,” the researchers pooled Big Five self and observer ratings across studies. They reported corrected mean correlations of .46 for Agreeableness, .56 for Conscientiousness, .51 for Emotional Stability, .62 for Extraversion, and .59 for Openness. These are meaningful positive associations: across the reviewed samples, people’s self-ratings and observers’ ratings tended to move together, with the strength varying by domain.

A corrected correlation estimates an association after statistical corrections; it is not a percentage of behavior an observer got right. The five values show convergence, not a contest between sources. Even the highest value, for Extraversion, leaves room for disagreement; the lower Agreeableness value means less overlap, not that either source is useless. The review also reports substantial unique variance in both kinds of ratings: each perspective contains information not shared by the other in the pooled data.

That pattern makes sense because the two sources do not have identical vantage points. A person can report on private intentions, feelings, or behavior across situations an observer never sees. An observer can describe how the person came across in encounters the person remembers differently or did not notice. These are plausible interpretations of why ratings can overlap while retaining unique information; the meta-analysis establishes the statistical pattern, not which explanation applies to a particular pair of accounts.

The practical implication is to treat agreement as corroboration at the level the evidence supports. If a report and observer both point toward a similar broad tendency, that convergence can make the interpretation more worth exploring. It still does not establish that a particular sentence describes every setting, that the report’s score and the observer’s judgment use comparable scales, or that one person’s example is accurate. Conversely, a difference does not erase the shared pattern or prove that one source is mistaken; it identifies where the accounts may be drawing on different information.

These correlations summarize groups of ratings, not the probability that a specific coworker is right about a habit. The review covers Big Five ratings across studies; it cannot adjudicate an unnamed report against one colleague. Using pooled values as a personal tie-breaker would turn “How much do these ratings tend to converge?” into “Which account is true here?”—a conclusion the numbers do not provide. Neither source should be dismissed; examine the exact behavior on its own terms.

Sources: The Convergent Validity between Self and Observer Ratings of Personality: A meta-analytic review

Can the two scores be compared on the same scale?

A numerical gap supports score-level comparison only when both forms measure the same construct comparably and their units support the operation. Check whether both ratings use the same named assessment, version, items, response scale, and scoring method, or whether one is a report score and the other an informal judgment. Then consult technical documentation for scoring, norm population, and whether self and observer scores were designed and tested for direct comparison. A shared trait label such as conscientiousness is not enough: forms may differ in item content, reference period, scoring, or interpretation.

The key technical question is measurement invariance: whether a measure relates to the underlying construct in sufficiently similar ways across groups or forms. If the relationship changes, the same observed score can carry a different meaning. The study “Do self-reports and informant-ratings measure the same personality constructs?” tested this question for Five-Factor Model domains and 30 facets in 3,253 people. Its results distinguish two kinds of comparison. Metric invariance concerns whether score relationships, such as how people rank relative to one another, are comparable. Scalar invariance adds the condition needed to compare average score levels on a common footing. In plain language, metric results can support comparing who tends to be higher or lower; scalar results are needed before treating a numerical mean gap as a difference in level rather than a difference in how a form uses its scale.

In that study, all 30 facets and all Five-Factor Model domains except Agreeableness met the reported metric-invariance standard. This supports many relative-standing comparisons across the tested self and informant ratings. Scalar invariance was not supported for Agreeableness and ten facets, so mean-level comparisons for those measures were not justified in the same way. This is not a finding that self and informant ratings are inherently incomparable. It shows why the comparison must be specified: a form pair may support one question about relative standing while failing to support another question about the size of an average score difference.

For a reader looking at two numbers, the practical rule is to identify the intended operation before interpreting the gap. Comparing whether the same person tends to be rated higher on one trait than another is not the same as subtracting a self score from an observer score and calling the result a meaningful discrepancy. Averaging the two numbers makes an even stronger assumption: that they use a shared scale and that combining perspectives is warranted by the scoring model. Unless the manual or validation evidence supports that comparison, retain the scores as separate observations and examine the item or behavior content behind them. Do not calculate a difference simply because both reports show a number beside the same trait name.

The FFM study cannot answer whether an unnamed report and a particular colleague’s rating share a scale. Its sample, forms, constructs, and statistical tests belong to that research setup; it does not validate every provider’s paired forms or establish behavioral accuracy. If the provider offers no form-comparison evidence, the defensible conclusion is limited: the two scores may prompt a useful question, but their numerical distance has no established score-level meaning. You can still compare the underlying descriptions or episodes in ordinary language, provided you keep that separate from a psychometric claim that the scores are interchangeable.

Sources: Do self-reports and informant-ratings measure the same personality constructs?

What could each person actually observe?

To weigh the accounts, ask what information each person had a real chance to see. An observer’s access is relevant because a rating based on repeated encounters can draw on more examples than one based on a narrow slice of work. But contact is an opportunity to gather information, not proof that the resulting judgment is accurate. The reader may also know patterns across private decisions, preparation, or settings that the observer never encountered. The useful question is not simply who knows whom better; it is which source had access to the behavior named in the disputed description.

The meta-analysis “An Other Perspective on Personality: Meta-Analytic Integration of Observers' Accuracy and Predictive Validity” integrated results covering 44,178 targets in 263 independent samples. It examined three different kinds of evidence: consensus or reliability among observers, correspondence between self and observer ratings, and prediction of behavior. These are related but distinct criteria. Agreement among observers asks whether different observers give similar ratings; self-observer correlation asks whether two perspectives covary; behavioral prediction asks whether ratings anticipate measured behavior. None alone establishes that a trusted coworker is right about one specific habit, and a rating can perform differently depending on which criterion is used.

The article record reports that interaction frequency and intimacy were associated with observer accuracy, with intimacy contributing for less visible traits. That pattern gives a practical reason to ask about exposure: how often did this person see the relevant work, over what period, and in which situations? It also cautions against treating all contact as equivalent. A colleague may see meetings every day but not the decisions made afterward; a project partner may see follow-through closely but only during one unusually pressured month. These are questions about the observer’s sample of behavior, not a formula for awarding the more familiar person extra credibility.

Visibility changes what access can reveal. A behavior that leaves a public trace, such as whether a person shares a draft by an agreed deadline, may be available to several observers. A private intention, an unspoken concern, or a choice made outside shared meetings may be available mainly to the person making it. Even visible behavior can have competing explanations: an observer might see silence in a meeting but not know whether the person was listening, uncertain, excluded from the decision, or waiting for a promised turn. The meta-analysis concerns broad accuracy evidence across varied studies; it does not tell us which explanation fits a reader’s episode.

Map each account to its observation window. Ask the observer for the setting, approximate date, action they directly saw, and whether they saw the lead-up and outcome. Ask yourself whether your account includes observable follow-through as well as intention. Separate overlap from unseen portions: if both witnessed the same meeting, compare their descriptions of that episode; if one describes private preparation and the other only the meeting, they may draw on different evidence. A trusted relationship can ease candid conversation, but does not replace checking what was observed.

Frequent contact still allows selective attention, mistaken interpretation, and narrow sampling. Broad self-knowledge may include many events yet miss a repeated effect conspicuous to others. Treat access as a reason to inspect examples and their coverage, not as a ranking system. An observer’s report is more informative when it names a situation within their view; a self-account is more informative when it identifies behavior or context outside that view. Where the windows do not overlap, keep the broader trait unresolved and compare a future shared episode.

Sources: An Other Perspective on Personality: Meta-Analytic Integration of Observers' Accuracy and Predictive Validity

Why do trait labels invite more disagreement than visible actions?

A broad trait label is harder to compare because it compresses several possible actions and a judgment about what those actions mean. John and Robins’s (1993) paper, “Determinants of Interjudge Agreement on Personality Traits: The Big Five Domains, Observability, Evaluativeness, and the Unique Perspective of the Self,” analyzed personality judgments in two cross-validated samples. Its publisher abstract reports that agreement differed across Big Five domains, and was higher for more observable and less evaluative traits. The finding supports a narrow point: judges tend to agree more when the trait has visible behavioral anchors and carries less evaluative meaning. It does not show that an observer’s account of a particular work episode is necessarily correct.

The distinction matters because a word such as dependable sounds like one trait, yet people may use it to summarize different conduct. One person might mean that someone meets an agreed deadline; another may mean that the person gives early notice when a deadline is at risk, or reliably closes a task after a meeting. Those are related but nonidentical actions. Likewise, independent could refer to starting without prompting, making a decision without consultation, or completing work without help. If a report uses one adjective and an observer disputes it, the disagreement may concern which behavior counts, how often it must occur, or which standard applies—not necessarily a direct contradiction about the same event.

Observability helps only after the behavior is specified. A label about being organized can invite a global impression that blends planning, record keeping, prioritization, and deadline management. A bounded description—such as whether a particular handoff included the agreed document by the agreed time—gives the parties a more common object to discuss. Even then, context matters: the deadline may have changed, the document may have been blocked, or the observer may not have seen the handoff. Making the action concrete reduces ambiguity; it does not remove the need to check circumstances or establish that the action represents a lasting tendency.

Evaluativeness adds another source of variation. Words that imply virtue or fault can shift a judgment from “what did this person do?” toward “what kind of person are they?” The same choice may be called decisive by one judge and dismissive by another, depending on expectations about consultation and authority. John and Robins’s result is about agreement in trait judgments, not a test of these particular workplace interpretations. The practical inference is that a global adjective can bundle both behavior and a standard for approving it, so disagreement over the adjective alone cannot identify which part differs.

Translate the label into three questions: What happened that could be seen or recorded? In which setting, with what expectation? Over what period or how many occasions? Keep the report’s exact wording beside this translation; its documentation may define the term more narrowly or not explain it. Ask the observer for the action behind their description. Compare accounts only where behavior, setting, and time window overlap. If the observer describes missed updates during one project while the report summarizes a general tendency, they are not yet describing the same slice; the example should qualify the broad interpretation only if relevant to what the report claims.

A concrete action remains open to interpretation. A person can share a draft on time for reasons that do not establish dependability across other work, and a missed handoff does not by itself reveal motive or a stable trait. Visibility makes a judgment easier to compare, not objectively correct. The study’s cross-validated pattern gives a reason to move from evaluative adjectives toward observable descriptions; it does not supply a rule for deciding who is right in an individual disagreement. The useful result of translation is narrower: both people can identify whether they are discussing the same action under the same conditions, and what further example would be needed if they are not.

Sources: Determinants of Interjudge Agreement on Personality Traits: The Big Five Domains, Observability, Evaluativeness, and the Unique Perspective of the Self

How can two accounts each contain information the other misses?

Self and observer accounts can diverge because the two sources may have different access to behavior and different reasons for noticing or emphasizing it. The Self–Other Knowledge Asymmetry (SOKA) model proposes conditional advantages: people may have better access to less observable aspects of themselves, while other people may judge more observable aspects and socially evaluative traits differently. These are predictions about kinds of information and judgment, not a universal assignment of truth to self or observer. A reader should use the model to ask what each account could have sampled, not to decide in advance which perspective wins.

The study “Who Knows What About a Person? The Self–Other Knowledge Asymmetry (SOKA) Model” tested these ideas with 165 participants. Participants supplied self-ratings; four friends and up to four strangers rated them in a round-robin design, and researchers collected behavioral-test criterion measures. The design compared self and informant judgments with behavioral criteria in one sample. It was not a workplace study of colleagues disputing a report, so its predictions cannot decide a case about an employee. Its contribution is a conditional question: what could the self know that observers could not see, and what public behavior might observers notice that the self overlooks or describes differently?

A person may know their private preparation, hesitation, or reasoning, including work completed when no observer was present. That information can help explain why an action occurred, but it does not automatically establish how the action appeared to others or what effect it had. Observers may see a public consequence—a delayed handoff, a decision announced in a meeting, or an unresolved request—without seeing the context that preceded it. Their view can reveal impact or a repeated outward pattern, while remaining incomplete about intention, barriers, or other settings. Each account can therefore contain information unavailable to the other, but access alone does not make either account accurate.

The model makes motivation relevant without licensing a bias accusation. Ratings may reflect what each person wants to convey or considers important, depending on trait and situation. SOKA treats these as conditional predictions alongside observability; it does not assume self-description is defensive or a third party neutral. A gap alone reveals no hidden motive. Ask whether the statements concern private intent, public effect, or an unstated judgment standard. This may explain divergence without proving the accounts are complementary truths.

Hypothetical workplace illustration, not study evidence: during an early project stage, an employee may privately compare options, ask for input, and delay committing until a missing requirement is clear. A colleague who sees only the later meeting may remember a direct decision and describe the employee as independent. During a later stage, the same colleague might see the employee consult others before a handoff and describe them as dependent, while the employee thinks of the earlier analysis as independent work. The accounts refer to different stages and visible actions, so the adjectives need not contradict each other. This illustration shows how information windows can differ; it does not establish that either person’s recollection is accurate or that the pattern is stable.

A practical implication is to separate the episode each source can describe from the broader conclusion each draws. Ask yourself which parts of the work happened privately and which left observable effects. Ask the observer what they directly saw, what result they experienced, and which stages fell outside their view. If the report describes a broad tendency, check its own documentation before treating the word as a precise behavioral claim. Then compare overlapping episodes rather than combining unseen preparation with a public outcome as though they were identical observations. Differences in access can explain a discrepancy, but explanation is not adjudication.

Different vantage points do not guarantee that both accounts are partly right. Either person may misremember, infer too much from a limited event, or apply a different standard; both descriptions may also conflict about the same repeated behavior in the same setting. In that case, calling the accounts complementary would conceal the disagreement rather than resolve it. The SOKA study offers a bounded mechanism and a way to form better questions, not a method for scoring a coworker or deciding whose report to trust. Treat each perspective as a claim with a source and scope. Where the accounts cover distinct slices, keep those slices distinct; where they overlap, examine the shared behavior and the evidence for it.

Sources: Who Knows What About a Person? The Self–Other Knowledge Asymmetry (SOKA) Model

Does a disagreement prove that the self-report is self-protective?

No. A gap between your account and an observer’s does not, by itself, show that you are protecting your self-image. The meta-analysis “Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means” compared Big Five self-report and informant-report means across 152 samples and 33,033 targets. It found that the means generally did not differ, reporting an average delta of −.038. Moderate differences appeared when self-ratings were compared with ratings from strangers. This is counterevidence to a blanket claim that people systematically rate themselves more favorably than informants do.

That result is about averages across samples. It does not say that every person’s self-rating matches every informant, that either kind of rating is accurate, or that a particular report and colleague’s account measure the same thing. A small overall mean difference can coexist with substantial disagreements in individual cases: higher and lower differences may balance in the aggregate, and the informants in a study need not resemble the trusted observer in front of you. The stranger comparison also cautions against treating all informants as one interchangeable group. Relationship and access may matter, but this analysis does not adjudicate the source of a particular discrepancy.

A mean comparison asks whether scores averaged across study participants tended to differ. It does not measure each person’s gap, explain its cause, or test a single work behavior. A mean difference alone does not establish self-enhancement; results may depend on which informants were included and what their ratings captured. The stranger pattern reminds us that informant identity and relationship matter, but does not show that a trusted colleague is more accurate than a stranger in this case.

The inference is limited but useful: do not explain the gap by assigning a motive before examining the behavior. A self-protective account is possible in a particular situation; so are incomplete observation, different standards, recall error, or a report description that does not fit the episode. Conversely, a lack of general average self-enhancement cannot rule out a blind spot in one recurring work habit. The meta-analysis tests a group-level pattern, not a person’s honesty or insight, and the mean result cannot be converted into an individual diagnosis.

Treat the discrepancy as a reason to ask for specific evidence. What did the observer see you do, when, and against which expectation? What did you mean by the report’s wording, and which occasions support your interpretation? If concrete, repeated examples converge on a behavior you had not noticed, that evidence may qualify your view even though no study result proves you were defensive. If the examples concern different settings or standards, preserve that distinction. A story about bias should follow evidence about the behavior, not substitute for it.

Sources: Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means

How should you examine one disputed work habit?

Use one disputed habit as the unit of discussion. The aim is to make the descriptions comparable, not decide who has the better personality judgment. Keep the report sentence, observer example, behavior, and any score separate until their relationship is clear. A score summarizes responses under an instrument; it does not automatically verify a coworker’s interpretation of an episode.

First, copy the report’s exact sentence and locate its context. Does the report describe a broad tendency, a facet, or a result tied to a particular set of questions? Check the available explanation of what the scale or narrative represents, and whether it names a reference period or norm. If the report only supplies a broad adjective, do not silently translate it into a more precise claim. Write down the ordinary-language behavior you think the sentence refers to, while marking that as your interpretation rather than a definition supplied by the test.

Second, ask the observer for one recent episode that led them to their description. Invite them to identify what they directly saw, approximately when it happened, and the setting: a planning meeting, a handoff, an urgent decision, or another concrete moment. Ask what action or omission mattered to them. “You are not independent” is a conclusion; “you asked the group to approve the choice before sending it” points toward an event that can be discussed. The example may still be remembered imperfectly, but it gives both people something more specific than a character label.

Third, clarify the expected standard for that situation. Did the work call for deciding alone, consulting affected colleagues, making a recommendation, or following a decision already made? Those are different demands. An observer who expected a quick recommendation may call a person hesitant; the person may understand the task as requiring consultation before commitment. The disagreement can involve conduct, the standard applied to it, or both. State the standard explicitly rather than assuming that “independent” or “collaborative” has one shared workplace meaning.

Fourth, map what each person could observe. Ask whether the observer saw the preparation, discussion, decision, and follow-through, or only one stage. Note what you know from private work that was invisible to them, and what public result they experienced that you may have overlooked. Compare only the part both accounts cover. If your report interpretation concerns planning done alone and the colleague describes a decision meeting, those observations can inform different parts of the habit; they do not yet answer the same question. Do not fill an unseen stage with an assumption about motive.

Consider an invented illustration, not evidence: a person spends a morning comparing options independently, then brings two unresolved constraints to a team meeting and asks for input before choosing. A colleague who sees only the meeting may describe the person as reluctant to decide; the person may think of the morning’s analysis as independent work. To examine the difference, identify the behavior under discussion—starting analysis, consulting, making the final choice, or carrying it out—and ask which stages the colleague actually witnessed. “Independent” may bundle several behaviors that need separate descriptions.

Finally, if the accounts remain unresolved, name a future situation that could distinguish them. For example, agree to notice what happens when a decision has a clear owner and deadline: whether the person starts analysis, seeks input, makes a recommendation, decides, communicates the choice, and completes the next step. Record the context and the action rather than assigning a score. One vivid incident can clarify that episode but cannot establish a stable tendency; several observations may still represent only one role, team, or period. Keep any conclusion within that observed scope, and revisit it if the work conditions change.

This sequence is practical reasoning for a conversation, not a validated assessment, diagnostic method, or scoring rubric. It cannot establish a stable trait or decide whose account is objectively correct. It helps identify whether the disagreement concerns the same behavior, setting, period, and expectation. If evidence overlaps, discuss the action and its consequences. If it does not, keep each account within its observed scope and gather a more comparable example before drawing a broader conclusion.

What is the next step you can take?

Ask the observer for one recent example, then compare the same action in the same setting with the report sentence. If the example fits, retain the interpretation only as far as the report’s documentation supports it. If it shows a narrower pattern, qualify the interpretation for that setting. If the report does not explain its wording, suspend that interpretation rather than drawing a conclusion about your whole work style. When the example concerns another task or the observer saw only one stage, leave the conclusion provisional and notice another relevant instance. A specific example may change how you describe one recurring behavior without resolving every question about your personality or work contexts.

You do not need another assessment to settle a clear, specific question. If you want a private prompt for examining your own decision and collaboration tendencies, the Work Pattern Report at /assessment can help frame reflection; it has no norm, cutoff, type, or selection score, and cannot decide whose account is right. For broader help interpreting reports, visit /topics.

Questions readers ask

Does a disagreement between my personality report and a coworker mean the report is unreliable?

No. A disagreement alone does not establish that a report is unreliable. Check what the report measured and whether its wording and score are documented, then compare the behavior and setting each account describes.

Should I average my personality report score with an observer’s rating?

Only if the instrument provides a supported method for combining them. Similar trait labels do not establish that two scores use comparable scales or can be averaged.

Can a trusted observer be more accurate than my personality report?

Trust alone does not show whose account is more accurate. An observer may have evidence about visible behavior in a particular setting, while you may know private actions and other contexts. Compare specific examples before drawing a conclusion.

Sources and notes

  1. The Convergent Validity between Self and Observer Ratings of Personality: A meta-analytic review

    Across Big Five studies, self and observer ratings had moderate corrected correlations and each source retained unique variance; duration of acquaintance and observer type were examined as moderators.

  2. Do self-reports and informant-ratings measure the same personality constructs?

    For Five-Factor Model domains and 30 facets in N=3,253, metric invariance generally supported relative-standing comparisons; scalar invariance was not achieved for Agreeableness and ten facets, limiting mean-score comparisons.

  3. An Other Perspective on Personality: Meta-Analytic Integration of Observers' Accuracy and Predictive Validity

    The meta-analyses cover 44,178 targets in 263 independent samples and examine observer consensus, self-observer correlations, and behavioral prediction; the article record reports accuracy associations with interaction frequency and intimacy, with visibility relevant to less visible traits.

  4. Determinants of Interjudge Agreement on Personality Traits: The Big Five Domains, Observability, Evaluativeness, and the Unique Perspective of the Self

    John and Robins (1993) report cross-validated findings in two samples: agreement varied by Big Five domain, observability, evaluativeness, and whether the self was a judge; more observable and less evaluative traits elicited higher agreement.

  5. Who Knows What About a Person? The Self–Other Knowledge Asymmetry (SOKA) Model

    The study tests predictions about which personality aspects are better judged by self versus others; its design includes 165 participants, self-ratings, ratings from friends and strangers in a round-robin design, and behavioral criterion measures.

  6. Self-Other Agreement in Personality Reports: A Meta-Analytic Comparison of Self- and Informant-Report Means

    A meta-analysis of 152 samples and 33,033 targets found self-report and informant-report Big Five means generally did not differ (average delta=-.038); moderate differences appeared when self-reports were compared with stranger reports.

Apply it to your work

Turn a work-style disagreement into a specific question

From this guide: If the report and observer still seem to conflict, identify the decision or collaboration tendency you want to examine in your own work.

The Work Pattern Report can help you describe your own decision and collaboration tendencies across ten continuums, giving you a more specific starting point for reflection. It is a low-stakes self-report with no norms, cutoff, type, or selection score, so it cannot determine whether an observer is right. Use it when naming your own pattern would help you ask a clearer question about a particular work situation.