No. The BFI-2 developers state that there is no official BFI-2 manual with published norms. That does not mean the instrument lacks research evidence: its development studies examined how its scales work, and researchers can describe scores in particular samples. But a percentile makes a different claim. It places a score within a defined reference population. A score converted to a precise-looking percentile is interpretable only when the report identifies a suitable population, the relevant BFI-2 form and scoring method, and the data and conversion used. The developers’ page mentions age- and gender-patterned descriptive data for the earlier BFI, not an official BFI-2 norm table. If a report cannot document its comparison group, read the scale score as a description of responses, not as an official ranking among people.
Does the BFI-2 provide an official percentile reference?
The developers say no: the Berkeley Personality Lab’s BFI page states that there is no official BFI-2 manual with published norms. The same page points readers to a paper with age 20 to 60 means from an American sample, converted to POMP, or percentage of maximum possible, and graphed by age and gender. It identifies that work as a study of the earlier BFI. The distinction matters. Those descriptive data are not an official BFI-2 percentile table.
A scale score summarizes a person’s answers under a scoring rule. A percentile is a comparison: it estimates the share of a specified reference population whose scores are at or below that score. The National Council on Measurement in Education defines norm-referenced interpretation in terms of a test taker’s performance relative to the score distribution in a specified reference population. A conversion can be mathematically neat and still have no defensible population behind it.
This is why ‘the site calculated my 74th percentile’ and ‘I rank at the 74th percentile in a defined population’ are not equivalent statements. The first describes an output. The second requires evidence about whose scores created the comparison and how the conversion was made. The developer statement answers the official-manual question; it does not prove that no researcher or provider has ever built a local comparison distribution. Any such distribution must be named and bounded on its own terms.
Sources: Big Five Inventory: Berkeley Personality Lab; NCME Assessment Glossary
What does BFI-2 research establish, and what does it leave open?
The PubMed abstract for Soto and John’s original BFI-2 paper describes three studies: the first specified a hierarchical model with 15 facets within five broad domains and developed a preliminary item pool; the second constructed domain and facet scales; and the third used two validation samples to evaluate measurement properties and relations with self- and peer-reported criteria. The authors report that the results indicate the BFI-2 is a reliable and valid measure. The abstract does not identify the validation samples’ populations, so it cannot establish from that record alone how far those findings generalize. It supports saying that the BFI-2 has published development and validation research, not that it is suitable for every population or use.
Validation and norming answer different questions. Validation asks whether evidence and theory support a proposed interpretation of scores for a use. Norming requires a score distribution from a defined reference population so that an individual’s relative position can be estimated. A measure may show a coherent structure or useful associations without supplying representative norms for every person who later takes it. The original paper’s validation work therefore cannot, by itself, establish an official percentile conversion.
A tempting counterargument is that studies contain means, standard deviations, or large samples, so a percentile can simply be calculated. Those statistics can support research comparisons within the sample and may inform a carefully described local reference. But a sample’s size alone does not establish that it represents a target population, and a study focused on scale development is not automatically a norming study. The developers’ older-BFI example makes the version issue especially clear: data for an earlier instrument should not be silently relabeled as BFI-2 norms.
So the BFI-2 can have useful evidence behind its scale scores while a percentile attached by a particular report remains unsupported, or applies only to a narrow sample. Neither result is a diagnosis, a measure of personal worth, or a conclusion about work suitability.
Sources: The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power — PubMed; Big Five Inventory: Berkeley Personality Lab; NCME Assessment Glossary
What information makes an individual percentile interpretable?
Start with the comparison group. NCME’s glossary distinguishes a reference population from user norms: percentiles calculated from a set of self-selected test takers or everyone tested during a period do not necessarily represent a well-defined population. Such values can describe that user group, but they should not be presented as if they locate the reader among all adults. The label ‘norm’ does not settle the issue; the report must say who is represented.
Then check the data and method. The GRoNC abstract describes a reporting guideline developed from a literature review and a two-round Delphi process involving theoretical experts and test developers. It says the guideline offers questions and explanations for test developers reporting standardized scores and reviewers evaluating them. Applied as a reader’s checklist, ask whether a BFI-2 report names its comparison group, explains how participants and scores were handled, and documents how the percentile was calculated. These questions help you inspect a report; the abstract does not evaluate the BFI-2 or prescribe a specific conversion.
A descriptive score and a percentile should be judged by the same purpose. A mean item response can tell a reader where answers fall on the questionnaire’s response scale; it does not say how common that score is among people. A sample-based percentile can say how the score compares with that sample, if the sample and method are disclosed. A population percentile makes a wider claim and needs a reference group and sampling evidence that justify it. More decimal places do not bridge the difference.
If the report is for self-reflection, a bounded comparison may be useful when it is labeled honestly. If the result will guide coaching or another consequential decision, ask whether the evidence supports that particular interpretation and use. Reliability, or score consistency, is not a substitute for evidence that the percentile means what the report claims.
Sources: NCME Assessment Glossary; The GRoNC: Guidelines for Reporting on Norm-Referenced and Criterion-Referenced Scores — PubMed
How should you read a BFI-2 percentile on a report?
Treat it as an official population rank only if the report identifies a suitable reference population and explains its BFI-2 scoring and percentile conversion. If it instead uses a clearly described group of website respondents, the number may be a user-group comparison, not a general population norm. If it gives no comparison source, the population meaning is unknown. This verdict is limited to the standard BFI-2 evidence and the official developer page reviewed here; it does not rule out local norms or later studies, and it does not make published research statistics useless.
A useful reading sequence is short. First, identify whether the report shows an item or scale score, a POMP-style position on the possible response scale, or a norm-referenced percentile. Second, locate the comparison population and check whether it fits the reader and context. Third, confirm the exact BFI-2 form and scoring procedure. Finally, match the conclusion to the evidence: a sample comparison is not a universal ranking, and neither is a verdict about a person.
When repeated work friction is the real question, a population percentile may not resolve what is happening in a particular collaboration. The live Work Pattern Report offers a separate, low-stakes self-report across decision and collaboration tendencies; it has no norms or selection score. It can give someone observations to discuss, not determine who is right or what job fits.
Ask the report provider for its norm-group description, sampling date and method, instrument version, and conversion procedure. If those details are absent, set the percentile aside and use the underlying scale score cautiously as a prompt for reflection. The decision point is simple: can the report show who the comparison represents? If not, it has not shown what the percentile claims to rank.
Sources: Big Five Inventory: Berkeley Personality Lab; NCME Assessment Glossary
Questions readers ask
Does a BFI-2 percentile mean the same thing as a high scale score?
No. A scale score summarizes responses according to the BFI-2 scoring procedure. A percentile locates that score relative to a specified comparison group. A score can be described without norms; a population percentile cannot be interpreted without knowing the reference distribution.
Can a research sample be used to calculate a BFI-2 percentile?
It can support a percentile relative to that sample if the data and calculation are adequate and the comparison is clearly labeled. That result does not automatically represent a wider population. The report should disclose how participants were sampled and what group the comparison can reasonably describe.
Sources and notes
- Big Five Inventory: Berkeley Personality Lab
The Berkeley Personality Lab states that there is no official BFI-2 manual with published norms. It points instead to age 20–60 means from an American sample, converted to POMP and graphed by age and gender, from a 2003 study of the earlier BFI.
- The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power — PubMed
The PubMed abstract reports three studies: model and item-pool development, construction of domain and facet scales, and evaluation of measurement properties and relations with self- and peer-reported criteria in two validation samples. It reports the authors’ conclusion that the BFI-2 is reliable and valid. The abstract does not identify the validation samples’ populations or establish official norms or a percentile conversion.
- NCME Assessment Glossary
NCME defines norm-referenced interpretation as comparison with the performance distribution in a specified reference population; it defines user norms as descriptive statistics, including percentile ranks, for a group that does not represent a well-defined reference population, such as self-selected test takers. It also defines percentiles in relation to a specified population.
- The GRoNC: Guidelines for Reporting on Norm-Referenced and Criterion-Referenced Scores — PubMed
The PubMed abstract describes GRoNC as a reporting guideline developed from a literature review and a two-round Delphi process involving theoretical experts and test developers. It says the guideline provides questions and explanations to support reporting and evaluation of standardized scores. It does not assess the BFI-2 or prescribe a particular percentile conversion.
Apply it to your work
Turn a broad work-style result into questions you can observe
From this guide: A population percentile cannot tell you how a recurring decision or collaboration tension appears in your own work.
If a BFI-2 report leaves you wondering why the same work friction keeps returning, compare the actual situations: how you decide, handle ambiguity, exchange feedback, and respond to change. The Work Pattern Report offers a private, low-stakes self-report across those tendencies so you can form clearer questions for reflection or conversation. It provides no norms or job recommendation.
