Berkeley Personality Lab’s current documentation says it has no official BFI-2 manual with published norms. The POMP figures it links come from an earlier BFI study and are sample averages, not BFI-2 percentiles. An online percentile may still describe a provider’s own comparison pool, but its meaning depends on the form, scoring method, and records behind it. Treat it as a population norm only if a suitable norm source is documented for that exact measure and interpretation.
Does “official norms” mean one specific thing here?
For the BFI-2, Berkeley’s current Personality Lab documentation says that there is no official manual with published norms. That is a precise statement about what the lab documents; it does not establish that no independent researcher or provider could publish a useful comparison. Here, “official” means that the norm source is explicitly issued or documented for the named measure, with an identifiable basis for interpreting scores. A percentile label on a webpage, or a calculation that produces a number from 0 to 100, does not by itself establish that provenance. The distinction matters because “Does the BFI-2 have official percentiles?” and “Can this service rank my result against some pool?” are different questions. The first concerns a documented norm source for the instrument; the second concerns the data and method behind one service’s output. Berkeley’s linked materials should therefore be read for what they actually report, while an online result should be assessed by identifying who produced it, what instrument and reference data it uses, and what comparison it claims to make. That leaves room for a separate, transparent resource to be informative on its own terms, without treating it as a Berkeley-issued BFI-2 norm. In the sections that follow, the useful first move is to trace the displayed result to its source and statistic, instead of inferring official standing from the number’s appearance.
What do Berkeley’s linked POMP figures actually describe?
A number on a 0–100 scale can describe position across the possible score range without describing rank among people. Berkeley’s linked paper, “Development of Personality in Early and Middle Adulthood: Set Like Plaster or Persistent Change?”, uses POMP, a linear transformation that maps the measure’s possible minimum to 0 and its possible maximum to 100. In the paper’s earlier BFI scoring, this transformation makes a raw scale value readable on a common range; it does not count how many members of a comparison group scored below an individual. For an abstract example, if a scale’s endpoints are transformed to 0 and 100, a converted value of 60 represents 60% of the possible interval under that transformation. It does not mean that the respondent scored above 60% of a reference group. That latter statement would require a defined distribution of other people’s scores and a rule for locating the respondent within it. The paper reports sample means and standard deviations after its conversion. A mean is an average of converted scores in the study sample; it is not an individual’s percentile, and the mean and standard deviation alone do not identify an individual rank without additional distributional assumptions and the individual’s position in the data. The metric and the comparison answer different questions: POMP asks where a score falls between the instrument’s possible endpoints, while a percentile asks where it falls relative to a specified group. Thus the linked POMP figures may describe the earlier BFI sample on a transformed scale, but they cannot be read as BFI-2 percentile norms. The distinction is useful because rescaling can make scores with different raw ranges easier to compare descriptively while leaving the reference population question unanswered. For instance, groups could share the same possible-score endpoints while having very different score distributions; the same transformed value would still occupy the same place on the scale, even though its rank within each group could differ. That is why a range conversion cannot substitute for information about who was compared and how their scores were distributed. The conversion changes the units in which a score is expressed; it does not add observations, select a comparison population, or estimate the share of that population below a given person. Nor does a sample average become a rank norm simply because it is printed on a 0–100 axis. Berkeley’s page identifies the cited study as the source of POMP means by age and gender, and the study’s reported statistic should retain that description when quoted. A reader looking at a POMP value should therefore ask whether the report gives a transformed score or a rank, and, for any claimed rank, where the comparison scores came from. No percentile for an individual can be recovered from the POMP label alone.
Sources: Development of Personality in Early and Middle Adulthood: Set Like Plaster or Persistent Change?; The Next Big Five Inventory (BFI-2): Developing and Assessing a Hierarchical Model With 15 Facets
Who took part in the Berkeley-linked study, and what can its averages travel to?
The study behind Berkeley’s linked age and gender figures describes a very large sample, but its size does not turn its averages into individual BFI-2 norms. In “Development of Personality in Early and Middle Adulthood: Set Like Plaster or Persistent Change?”, the researchers analyzed cross-sectional Internet responses from 132,515 adults aged 21 to 60 using the earlier Big Five Inventory (BFI), before the BFI-2. The paper asks how personality scores vary across adulthood and reports age-pattern analyses and sample summaries. Those features make it useful evidence for the developmental questions its design and measure address. They do not make it a table that locates a current BFI-2 respondent among a defined population. The distinction concerns the object of inference: findings about patterns in these collected responses are not the same output as a person’s rank against a documented norm group.
A mean can travel as far as its sampling and measurement support. Within this study, an average summarizes the converted scores of the participants included in the relevant analysis. Readers may use it to understand what that sample’s results looked like, and researchers may examine associations or patterns within the paper’s design. But an Internet sample is not automatically a probability sample of all adults. People who encounter and choose to complete an online measure can differ from those who do not; a large count reduces some kinds of numerical instability without removing that selection process. The ages also bound the observed range: adults outside 21–60 are not represented by that stated age band. These limits do not erase the study’s findings. They tell us which population and period those findings most directly describe, and why extending them requires an argument about how the study participants relate to the new target group.
Time and version add separate boundaries. The paper is a historical cross-sectional analysis, so it compares people of different ages at one study period rather than following each person through life. An age difference in that design cannot, by itself, tell a reader that an individual will change in the same way as they get older; cohort differences and other explanations remain possible. The measure matters too: the study used the earlier BFI, while BFI-2 is a later instrument. Similar construct labels do not establish that scores from different forms share an interchangeable scale or reference distribution. To use an older finding for a present-day BFI-2 score, one would need evidence connecting the forms and supporting the intended interpretation. The study’s large sample does not supply that connection simply by being large.
The practical reading is therefore neither “ignore the paper” nor “treat its averages as norms.” Read it as a substantial study of adult personality patterns in the specified Internet sample, measured with the earlier BFI and analyzed for developmental questions. Its means and standard deviations describe group-level results in that research context. A norm lookup would answer a different question: where a particular person stands relative to an identified comparison group under a stated scoring procedure. The paper’s purpose, design, instrument, and reported summaries do not by themselves provide that lookup for the BFI-2. A large sample is valuable when it strengthens the evidence for the question actually studied; it cannot change the question the statistic answers.
Sources: Development of Personality in Early and Middle Adulthood: Set Like Plaster or Persistent Change?
What does BFI-2 validation establish about an online result?
Validation evidence for the BFI-2 matters, but it does not automatically validate a percentile printed by an online service. The paper “The Next Big Five Inventory (BFI-2): Developing and Assessing a Hierarchical Model With 15 Facets” describes development and assessment of a 60-item instrument organized into five domains and 15 facets. Its evidence concerns the studied measure and the interpretations examined in the paper’s samples. That is relevant when deciding what the original BFI-2 was designed to measure and whether the paper supports those intended interpretations. It is a different evidentiary object from a particular website’s reference table: that table has its own data, scoring implementation, comparison population, and claims. The paper’s validation purpose cannot establish facts about a dataset its authors did not study.
The logical gap is easiest to see by separating two questions. First: does evidence support interpreting scores from a specified BFI-2 procedure as indicators of the intended personality domains and facets in the studied context? The development paper is relevant to that question. Second: does an online service’s number show a respondent’s rank within a defined group, and is that group suitable for the interpretation the service invites? Answering that requires evidence about the service’s score generation and comparison data. A measure can have evidence for its construct interpretation while an external percentile remains undocumented, poorly matched, or simply impossible to evaluate from the display. Conversely, a provider might disclose a clear local comparison table without establishing broad population standing. Instrument evidence and reference-data evidence bear on related but non-identical claims.
A percentile claim therefore needs a traceable chain beyond evidence that the BFI-2 has been studied. The service must specify what form and scoring procedure produced the score; identify the reference data against which it is ranked; describe how those data were sampled and processed; and state what population and use the comparison is intended to represent. The 2014 “Standards for Educational and Psychological Testing” treats score interpretation as dependent on evidence relevant to the intended use and describes norms in relation to a defined population. Applying that guidance here yields an inference, not a result reported by the BFI-2 development paper: a validated instrument alone cannot tell us whether an independent percentile table represents a suitable group. Each link answers a further question that the instrument’s construct evidence leaves open.
This distinction also prevents a common overreading of the word “validated.” Validation is not a permanent certificate attached to a name that transfers to every translation, web implementation, changed scoring key, or use. The claim must identify the score interpretation and context at issue. The BFI-2 paper remains relevant evidence for interpretations involving the studied BFI-2 procedure; it should not be discarded merely because a service adds a percentile. But the service needs evidence of its own for the added comparative claim. If the output does not name its reference group, show how the score was matched to that group, or explain the intended comparison, the paper’s validation findings cannot fill those missing details. That conclusion follows from the difference between evidence about a measure and evidence about a separate ranking procedure.
The counterpoint is important: an online percentile is not necessarily meaningless just because it is not part of the BFI-2 development study. A separately assembled table could support a narrow descriptive statement about ranks within its own documented records. The relevant questions then shift to the table’s provenance, eligibility rules, dates, representativeness, and the limits the provider places on interpretation. The instrument paper may help establish what the underlying domains mean, while the table documentation determines what “higher than this share of records” can mean. Neither source can do both jobs by implication. For a reader, “the BFI-2 has validation evidence” is a reason to examine what the result measures; it is not, by itself, a reason to accept the displayed percentile as a population norm.
Sources: The Next Big Five Inventory (BFI-2): Developing and Assessing a Hierarchical Model With 15 Facets; Standards for Educational and Psychological Testing (2014)
What changes to the form break the evidence link?
An adaptation is a change to a test’s content, format, language, scoring, or administration that may affect how responses produce and support an interpretation. Comparing an online result with the published BFI-2 therefore starts with form identity: the original BFI-2 is a 60-item measure, while a named short form has fewer items and a provider may present altered wording or procedures. Item count matters because removing questions changes which observations contribute to a score; wording and language matter because respondents may understand or answer prompts differently; response options and scoring keys matter because they change how answers become scale values. Administration can matter too, if instructions, context, or delivery change the response process. These are reasons to ask what was used and what evidence supports it, not grounds to declare a score invalid merely because the format differs. The “Standards for Educational and Psychological Testing (2014)” says subsets or rearrangements require evidence that the change has not distorted scores, cut scores, or norms. Its point is evidential: the comparison must be justified for the intended interpretation. A published measure’s name alone cannot establish that a changed procedure preserves the same score meaning. Even a response scale that looks familiar may use different anchors or a different number of choices, shifting how a response maps to a value. Likewise, reversing or reassigning items can alter the scoring key. Without the actual version and scoring description, visual similarity is not proof that two procedures yield interchangeable scores.
The short-form research shows why this distinction is practical rather than formal. “Short and extra-short forms of the Big Five Inventory–2: The BFI-2-S and BFI-2-XS” reports development and evaluation using an adult Internet selection sample of 1,000, an adult Internet validation sample of 2,000, and a university validation sample. It evaluates deliberately developed BFI-2 short forms, rather than an arbitrary online edit. The paper describes a real tradeoff: shorter forms save time, while reduced item coverage can cost reliability and validity, especially when interpreting facets from the extra-short version. This finding does not mean every short score is unusable or quantify the accuracy of a different test. It means that performance evidence belongs to a specified form and use; evidence for the 60-item measure does not automatically settle the precision of every shortened scale or its percentile comparison.
The “ITC Guidelines for Translating and Adapting Tests (Second Edition)” make the practical implication explicit. For an adaptation, they call for work such as pilot item analysis, reliability and validity analyses, samples relevant to the intended use, and documentation of the adaptation and score interpretation. That guidance applies to questions about changed wording, language, and procedure; it does not pronounce on any particular online BFI-2 result. A well-developed short form or translation can have its own supporting evidence. The reader’s comparison is therefore between identifiable procedures and their evidence: does the provider name the form, explain scoring and administration, and point to evidence for interpreting scores from that version? If those details are missing, equivalence remains unestablished. If they are documented, the next question is whether that evidence supports the particular claim being made, including any rank against a reference group. Form differences call for a closer evidence link, not an automatic verdict.
Sources: Standards for Educational and Psychological Testing (2014); Short and extra-short forms of the Big Five Inventory–2: The BFI-2-S and BFI-2-XS; The ITC Guidelines for Translating and Adapting Tests (Second Edition)
What population can a percentile refer to?
A reference group, sometimes called a norm group, is the set of people or records against which a score’s relative standing is calculated. A percentile answers “relative to whom?” A correctly computed percentile can be meaningful for one group and misleading if described as standing in another. It is not the percentage of questions answered correctly, a measure of personality quality, or a judgment that a score is good. The “Standards for Educational and Psychological Testing (2014)” frames norm-referenced interpretation in relation to a defined population and expects documentation that lets users understand the comparison. The population label is not a footnote: it is part of what the percentile means.
Consider a purely hypothetical illustration, not data: the same person receives the same underlying score, but one comparison pool consists mostly of people recruited from a particular training program, while another consists of a broader set of adults. Their score distributions could differ, so the person’s rank could differ too. Neither result is necessarily miscalculated. Each describes position in its own pool, and neither can be translated into the other without evidence linking the pools. A percentile from a local pool may be useful for a local descriptive question; it does not become a national estimate because it is displayed without qualification.
To judge how far a comparison can travel, the population description needs enough context to identify who could enter the pool and who actually did. Recruitment matters because volunteers, customers, students, or employees may differ from people who never joined. Dates matter because a table collected in one period describes those records, not automatically a present-day population. Language and exclusions matter because they shape who could complete the form and which results were retained. Those details are not a ritual checklist: each limits the claim’s reach. A well-documented service sample can support a carefully worded comparison among eligible records even if it does not represent a country. The boundary is to name that local comparison accurately. Without a stated group and its relevant composition, “the 70th percentile” leaves the reader unable to know what population standing it describes. For example, a result based on people who chose a particular service can describe those eligible records if the table’s rules and period are clear; it does not, by itself, estimate how all adults or workers would rank. A population name should therefore be read alongside the path into the sample, not as a guarantee of representativeness. Exclusions also define the comparison: removing incomplete forms may improve consistency of the table while narrowing the set of responses it describes. The appropriate interpretation follows the documented pool, while a broader claim needs evidence that the pool resembles the broader target on features relevant to the score. It is a comparison, not a population verdict.
Sources: Standards for Educational and Psychological Testing (2014)
Can a provider-built percentile still answer a narrow question?
A provider-built percentile can be informative when its claim is kept inside the boundaries of the provider’s own comparison table. “Team assessment methodology and evidence — SeeMyPersonality” is one documented example: the provider says its current English 60-item implementation adapts BFI-2 wording and uses a versioned table, identified as v1, with a recorded collection window. The page reports 58,401 distinct valid score codes after identical response profiles were deduplicated. Those details make the comparison more traceable than a bare percentile label: a reader can identify the form as the provider describes it, which table version is involved, and the period and records behind the ordering. This is an attributed description of that service, not a claim about online BFI-2 results generally. Versioning matters because a percentile is a relation to a particular distribution, not a permanent property of a person’s score: if the eligible pool or scoring implementation changes, the same score may occupy another position. The recorded window therefore helps define which comparison is being described and guards against silently treating successive tables as one unchanged norm.
The page also explains how it turns a score into a rank. It describes an empirical midrank calculation: the count of eligible codes below a score plus half the count tied at that score, divided by the eligible total, with the result expressed as a percentile. In plain terms, the method locates a score within the observed distribution and gives tied scores a midpoint position. Deduplication means the table counts each distinct score vector once, according to the provider’s stated rule. It does not establish that every code came from a different person, that every person answered honestly, or that the table has been independently audited. The number 58,401 therefore describes distinct valid score codes in the table, not 58,401 verified individuals.
The same methodology page expressly says the records are not a representative sample of adults, employees, or a general population. That qualification controls what the arithmetic can mean. If the provider’s description is accurate and the score being interpreted matches its specified implementation, one may describe where that score falls among eligible records in that particular versioned table. One cannot convert the result into “higher than this share of adults” or “higher than this share of workers.” The denominator is the provider’s eligible records, not a defined representative population. Nor does an orderly formula or a precise decimal change the identity of that denominator.
This is the strongest reasonable case for a non-official result: someone using the service might value a stable internal ordering, or use it as feedback about how their answer pattern compares with other eligible records collected by that service. The disclosure makes that narrow reading possible, though it remains the provider’s own account; independent verification and a stated rationale for a particular intended use are not supplied by the page reviewed. For broader interpretation, the missing bridge would include evidence about who contributed, how the records relate to the target population, whether the adapted form supports comparable scores, and whether the comparison works for the proposed use. Transparency can make a local comparison interpretable without making it official, representative, or suitable for decisions about people.
Sources: Team assessment methodology and evidence — SeeMyPersonality
What would make a provider’s percentile defensible for a broader claim?
A provider’s percentile becomes defensible for a broader claim only when the evidence supports each step between a response and that claim. The chain is: answers → score under an identified form and scoring rule → comparable records → people in a defined target population → interpretation for a stated use. A break at any link narrows what can be concluded. A technically reproducible rank may still describe only a selected service pool; a well-described population does not repair a score built from a materially different form; and a population comparison does not by itself show that the result is useful for a decision.
First, score generation must be clear enough to reproduce: what items, wording, response options, missing-answer rules, and scoring key produced the value? Then form comparability asks whether evidence supports interpreting the online procedure as comparable to the named BFI-2 or a particular short form. “Standards for Educational and Psychological Testing (2014)” treats changes to tests and the interpretations drawn from them as matters requiring supporting evidence. “The ITC Guidelines for Translating and Adapting Tests (Second Edition)” likewise call for evidence and documentation appropriate to an adaptation. These standards state criteria for a defensible claim; they do not report an empirical evaluation of SeeMyPersonality or any other provider in this article. The actual BFI-2 short-form research, “Short and extra-short forms of the Big Five Inventory–2: The BFI-2-S and the BFI-2-XS,” evaluates those specified forms and reports tradeoffs as length decreases. Its findings cannot be transferred to an arbitrary adapted implementation.
Next comes the bridge from records to people. A sampling frame defines who could enter the data; the target population defines the people the claim is about. Evidence must show how recruitment, eligibility, dates, language, and exclusions connect those two groups. A large n can reduce random fluctuation when estimating a quantity from an appropriate sampling process. It cannot make volunteers resemble nonparticipants, undo duplicated records, make a service sample representative, or align a mismatched version with the target population. In SeeMyPersonality’s disclosed table, deduplicating identical score codes may prevent repeated profiles from receiving repeated weight under its method, but it also means the reported count is codes rather than verified respondents. Neither that rule nor the count establishes representativeness. Data quality and uncertainty must be described in terms of what was checked and what remains unknown.
Finally, the evidence must match the intended use. Reliability concerns the consistency or precision of scores under specified conditions; it does not establish who the comparison records represent. Representativeness concerns how well the records support inference to a target population. Validity concerns whether evidence supports the particular interpretation and use proposed for scores. These are separate questions: a reliably computed rank can be unrepresentative, and a representative sample cannot by itself validate every consequential use. “Standards for Educational and Psychological Testing (2014)” ties interpretation to the intended use, while adaptation guidance asks for evidence relevant to that use. A source that supports one link should not be treated as proof of all the others.
For private curiosity, a cautious statement such as “my score ranked here among this service’s eligible records” may require less evidence than a claim about adults generally. The precision of the wording should track the evidence: local ordering can remain local, while generalization requires a relevant sampling frame, documented data quality, and support for the score’s meaning in the target group. For this provider, the methodology page itself disclaims representativeness, so it does not support population rank. The cited materials also do not justify using this percentile for hiring, promotion, compensation, diagnosis, or job-fit conclusions. Those applications would require separate evidence for the exact score interpretation, population, and decision, and no such support is established here.
Sources: Standards for Educational and Psychological Testing (2014); Short and extra-short forms of the Big Five Inventory–2: The BFI-2-S and BFI-2-XS; The ITC Guidelines for Translating and Adapting Tests (Second Edition); Team assessment methodology and evidence — SeeMyPersonality
What should I do with the result in front of me?
If the online BFI-2 percentile is still unclear, preserve the result but do not translate it into a claim about where you stand among adults. Save the provider name, the exact version or form it says it used, and the date you took it. Then ask what score the percentile was calculated from, which reference group supplied the comparison records, and where the provider documents those details. A clear answer lets you judge the stated comparison on its own terms. If the provider cannot identify the score source or comparison group, set the percentile aside; an unexplained rank does not become a population norm through repetition or a precise-looking number.
You can still treat a descriptive tendency as a prompt to notice recurring experiences, while keeping that reflection separate from norm interpretation or an employment conclusion. If your practical question is about how you approach decisions or collaborate, the live Work Pattern Report is a distinct, low-stakes self-reflection option. Its 100-item assessment describes patterns across ten work-pattern continuums; it supplies no BFI-2 norm, cutoff, or selection score, and it does not match people to careers. Use it to organize questions about your own habits, not to settle a job decision. [Explore the Work Pattern Report](/assessment).
Questions readers ask
Does the BFI-2 have official percentile norms?
Berkeley Personality Lab’s current documentation says there is no official BFI-2 manual with published norms. That describes the lab’s documentation; it does not rule out a separate resource supporting a defined comparison.
Are the BFI-2 POMP figures percentiles?
No. The linked figures come from a study of the earlier BFI and report sample averages on a 0–100 possible-score metric. They are not percentile ranks or BFI-2 norms.
What does an online BFI-2 percentile tell me?
It tells you only the comparison supported by the provider’s documented form, scoring method, and reference records. A provider-built percentile may describe a score’s position among its eligible records if its method and pool are documented, but that local comparison does not by itself establish standing among adults generally. Without those details, do not generalize the percentile.
Sources and notes
- Big Five Inventory — Berkeley Personality Lab
Current lab documentation says there is no official BFI-2 manual with published norms and identifies the cited age-20-to-60 American sample paper as containing POMP means by age and gender; the page also identifies the 60-item BFI-2 and shorter forms.
- Development of Personality in Early and Middle Adulthood: Set Like Plaster or Persistent Change?
The 2003 cross-sectional study used the earlier BFI with an Internet sample of 132,515 adults aged 21–60; its POMP conversion put raw BFI scale values onto a 0–100 possible-score metric, and it reports sample means and standard deviations, not percentile norms for BFI-2.
- The Next Big Five Inventory (BFI-2): Developing and Assessing a Hierarchical Model With 15 Facets
The original BFI-2 development/validation paper studies a 60-item instrument organized into five domains and 15 facets and supplies psychometric evidence for the studied measure and samples; it is not documentation that an unspecified website’s percentile reference is official or population-representative.
- Standards for Educational and Psychological Testing (2014)
The joint AERA/APA/NCME standards describe percentile-rank norms as standing within a defined population and require relevant evidence/documentation for intended interpretations; changed, shortened, or rearranged forms cannot simply assume undistorted scales or norms.
- Short and extra-short forms of the Big Five Inventory–2: The BFI-2-S and BFI-2-XS
The peer-reviewed form-development paper reports adult online selection and validation samples plus a university validation sample and evaluates shorter forms; the authors describe efficiency gains alongside losses in reliability/validity, demonstrating why similar naming or fewer items does not establish score interchangeability or a percentile norm.
- The ITC Guidelines for Translating and Adapting Tests (Second Edition)
The International Test Commission guidance calls for pilot item analysis, reliability and validity evidence for adapted tests, samples relevant to intended use, and documentation of adaptation and score interpretation; it explains why an adaptation’s evidence must be established rather than inherited by resemblance.
- Team assessment methodology and evidence — SeeMyPersonality
This provider says its current English 60-item implementation adapts BFI-2 wording, uses a dated versioned percentile table built from 58,401 distinct valid completed score codes, computes a midrank-style empirical percentile, and does not claim representativeness; it states codes are deduplicated score vectors rather than verified unique people.
- 100-item Work Pattern Report — Personality Report
The publication’s live 100-item assessment offers a separate result across ten work-pattern continuums and frames it as reflection rather than a hiring score; it is not a BFI-2 norm, cutoff, or validated career-match measure.
Apply it to your work
Turn a work-pattern question into specific observations
From this guide: If the BFI-2 result leaves a practical question about how you decide, plan, collaborate, handle conflict, adapt, or learn, use a separate reflection to examine those patterns.
A percentile cannot tell you which work pattern matters in a particular situation. The Work Pattern Report offers a separate low-stakes self-report across ten decision and collaboration continuums. Its result has no norms, cutoffs, type, or selection score, so use it to generate observations and questions for reflection, not as a job recommendation or employment judgment.
