In brief

Adverse impact in workplace personality testing means that using a personality assessment or another selection procedure results in a substantially lower hiring, promotion, or other selection rate for a protected group than for the group with the highest rate. It is about the effect of a selection process on groups, not a verdict about an individual applicant’s character. In the United States, the Uniform Guidelines on Employee Selection Procedures describe a selection rate below four-fifths, or 80 percent, of the highest group’s rate as a general rule of thumb for identifying possible adverse impact. That ratio is a signal for investigation, not a complete legal finding. A responsible review also asks whether the assessment measures a job-related requirement, whether its scores predict an important work outcome for the stated use, whether the sample and administration are appropriate, and whether a substantially equally valid procedure would have less adverse impact.

Start with the decision, not the personality label

Imagine an employer uses a personality questionnaire to decide who moves from application to interview. The report may describe tendencies such as planning or sociability, but the immediate fairness question is different: who advances, who does not, and how does the assessment contribute to that decision? A report can be personally interesting while still being unsuitable for screening applicants.

This distinction matters because adverse impact is a group-level property of a selection procedure. It does not mean that every person in the lower-selection group received a poor score, that the assessment measured prejudice, or that the employer intentionally treated people differently. Nor does a difference in average personality scores by itself establish adverse impact. The relevant comparison is usually the rate at which people pass a stage or are selected, considered by the protected groups covered by the applicable law.

Two readings of the same report

There are two plausible ways to read a workplace personality report. The first is the individual-score reading: a score is compared with a norm group, placed in a band, or combined with other results to describe a tendency. That interpretation may be useful for self-reflection or development when the instrument, norms, and limits are clear.

The second is the selection-process reading: a score becomes one rule in a chain of decisions, such as an automatic cutoff, a ranking, or a recommendation to interview. Here the report cannot be judged only by whether its language sounds reasonable. The reviewer must examine the job analysis, the decision rule, selection rates, evidence for the intended inference, and possible effects on groups. The same scale can therefore be acceptable for a low-stakes conversation and inadequate as an automatic hiring gate.

What the four-fifths rule actually compares

The four-fifths rule compares selection rates. A selection rate is the number selected from a group divided by the number considered in that group. The impact ratio is the lower group’s selection rate divided by the highest group’s selection rate. A ratio below 0.80 generally prompts concern under the U.S. Uniform Guidelines.

For an illustrative calculation, suppose 30 of 100 applicants in Group A advance, while 50 of 100 applicants in Group B advance. The rates are 30 percent and 50 percent. Dividing 30 by 50 gives 0.60, or 60 percent. That result is a warning that the procedure deserves closer review. It does not tell us why the rates differ, whether the sample is large enough for a stable conclusion, or whether the assessment is the specific cause when several hurdles were used.

The rule is not a safe harbor. The Guidelines say smaller differences can still matter when statistically and practically significant, while larger differences may be unstable with small numbers. The correct comparison also depends on the stage being studied and on accurate records of who was considered and selected.

Open book showing a balance scale with groups of simplified people and report sheets, while more people stand in the foreground.
Open book showing a balance scale with groups of simplified people and report sheets, while more people stand in the foreground.

The first interpretation: a job-related measurement

The employer’s strongest interpretation may be that the questionnaire measures a work-related tendency and adds useful information to a broader assessment. That claim has to be specific. “Good personality” is not a job requirement. A defensible statement might identify a behavior relevant to a role, explain how the job analysis established its importance, and show why the assessment is related to an important work outcome.

Validity means evidence supporting a particular interpretation and use of scores. It is not a permanent label attached to a test. The U.S. Office of Personnel Management notes that a personality test intended to forecast later performance needs evidence connecting scores with subsequent job performance. Reliability is a separate question about score consistency. A consistent score can still measure the wrong thing or fail to support a hiring inference.

This interpretation becomes weaker when a report turns a broad trait into a rigid prediction, uses a cutoff without a job-based rationale, or relies on evidence from a different role, population, or decision. A report should not be treated as proof that one applicant will be a good manager, safe coworker, or cultural fit.

The second interpretation: a selection rule with unequal effects

The competing interpretation is that the assessment is functioning less like a useful supplement and more like a gate. A short questionnaire may be completed online, scored automatically, and used to reject applicants before anyone examines relevant experience. If one protected group advances at a much lower rate, the procedure has an adverse-impact question even if the report’s descriptions are polished and the same instructions were given to everyone.

This lens also examines construct-irrelevant demands: features of the test or its delivery that affect scores without representing the job requirement. Language complexity, inaccessible timing, unfamiliar response conventions, or a forced-choice format may matter, depending on the instrument and the applicants. These are hypotheses to investigate, not assumptions about any group.

A low impact ratio does not prove that the test is biased in its content. It does show that the selection process may be excluding people unevenly. The practical response is to examine the procedure, its evidence, and alternatives before defending the cutoff as inevitable.

Why a valid test can still require an alternative

Under the U.S. framework, showing that a procedure is job-related and consistent with business necessity does not end every question. The Equal Employment Opportunity Commission explains that, when a selection standard has a significant disparate impact, the employer must address whether a less discriminatory alternative could meet the business need. The Uniform Guidelines likewise direct users to consider suitable alternatives with less adverse impact when procedures are substantially equally valid for the purpose.

That comparison is not a contest between a personality test and a perfect replacement. It may involve changing a cutoff, adding a structured interview, using a work sample, reordering stages, or combining measures, but each proposed change needs evidence and consistent administration. A less discriminatory alternative that cannot measure the requirement or cannot be implemented reliably may not serve the same purpose.

The important point for a report reader is that “validated” does not mean “fair in every use.” Validity is claim-specific, and fairness includes how the instrument is administered, scored, interpreted, and used in an actual decision.

Brass balance scale on a light desk between stacked report papers, a dark pen, green folders, and a leafy plant.
Brass balance scale on a light desk between stacked report papers, a dark pen, green folders, and a leafy plant.

A worked review of one hiring funnel

Consider an unnamed employer using a conscientiousness-focused questionnaire after an application and before a structured interview. The assessment is not automatically improper, and no conclusion can be drawn from the trait name alone. A useful review would separate the stages: applicants entering the questionnaire, applicants meeting the cutoff, applicants invited to interview, and applicants hired.

The reviewer would calculate each group’s selection rate at each stage, identify where the largest difference appears, and check whether the denominator is the same population that was actually considered. They would then inspect the job analysis and ask what observable work behavior the questionnaire is meant to inform. Next comes the evidence: does the instrument’s documentation support this population, role, and use? Are scores sufficiently consistent for the decision? Was the cutoff set and reviewed using relevant evidence rather than convenience?

Finally, compare plausible alternatives. If a structured work sample measures the same essential behavior with similar or stronger evidence and a smaller disparity, it deserves serious attention. If the evidence is inconclusive, the prudent decision may be to stop using the questionnaire as an automatic rejection rule while collecting better data. This example illustrates a method, not an invented result.

What an applicant can ask for

An applicant usually cannot infer adverse impact from a personal report alone. A low percentile, an unexpected band, or a rejection after testing does not reveal the group selection rates or the employer’s decision rule. It is reasonable, however, to ask what the assessment measures, how it is used, whether it is one factor or an automatic screen, and what accommodations or alternative arrangements are available through the relevant process.

If the employer provides a report, read its purpose statement carefully. Look for the intended population, norm group, score meaning, uncertainty, evidence for the use, and limits on interpretation. A report that only gives a flattering label or a pass/fail verdict leaves out information needed for responsible use. For a workplace decision, also ask who receives the result, how long it is retained, and whether a human reviews the wider evidence.

The legal options and deadlines vary by country, state, protected characteristic, and employment context. A person who believes a selection process treated them unlawfully should preserve notices and records and seek advice from the relevant labor or equality authority or a qualified professional in their jurisdiction.

Report-style page showing a brass balance scale with three dark figures on one pan and three lighter figures on the other, surrounded by bars and circles.
Report-style page showing a brass balance scale with three dark figures on one pan and three lighter figures on the other, surrounded by bars and circles.

A responsible employer’s review checklist

Before relying on a personality report in selection, an employer should be able to answer a short chain of questions. What decision is being made? What job behavior or outcome matters? How did a current job analysis establish that requirement? What exactly does the instrument measure, and what evidence supports this interpretation for this role and population?

The review should then cover score consistency, missing responses, accommodations, administration conditions, cutoff or ranking rules, and the contribution of every other selection stage. Keep selection counts by relevant groups and inspect the total process as well as its components when the data indicate a problem. Treat the four-fifths ratio as an initial screen, then consider sample size, statistical and practical significance, and the possibility that several rules interact.

Most importantly, record the alternatives considered. The U.S. OPM recommends beginning selection design with critical competencies from job analysis, using strong validity evidence, and explaining to applicants what the device assesses and why it matters. That discipline keeps a personality report in its proper place: one source of evidence, never a verdict about a person.

The practical conclusion

When you hear that a workplace personality test may have adverse impact, ask two questions in order. First, did the selection rates differ substantially across relevant protected groups, and is the data strong enough to interpret? Second, can the employer show that the assessment and its exact use are job-related, supported by appropriate validity evidence, and not replaceable by a substantially equally valid procedure with less adverse impact?

For a reader, the next steps are concrete: identify the report’s purpose, separate your score from the employer’s selection rule, request an explanation of what was measured and how it was used, and record the dates and materials. For an employer, calculate the stage-by-stage rates, review the evidence and administration, and test reasonable alternatives. If those questions cannot be answered, the report may be too thin to support a high-stakes employment conclusion.

Questions readers ask

Does a score difference prove adverse impact?

No. Adverse impact concerns substantially different selection rates between groups in an actual selection process. A difference in average scores, percentiles, or trait levels may prompt study, but it is not the same calculation and does not by itself establish the cause or legal significance of an employment outcome.

Is the four-fifths rule a legal pass or fail test?

No. In the U.S. Uniform Guidelines it is a general rule of thumb: a group’s selection rate below 80 percent of the highest group’s rate generally signals adverse impact. Small samples, statistical and practical significance, applicant discouragement, and other evidence can change the interpretation.

Can an employer use a personality assessment if it has adverse impact?

Possibly, depending on the jurisdiction and evidence. In the U.S., the employer generally needs to show the procedure is job-related and consistent with business necessity, and must consider whether a substantially equally valid alternative would have less adverse impact. A general personality report without use-specific evidence is not enough.

Sources and notes

  1. 29 CFR § 1607.4, Information on impact

    Supports the four-fifths rule, selection-rate comparison, small-number caution, statistical significance, applicant discouragement, and careful review of the total selection process.

  2. 29 CFR § 1607.3, Discrimination defined

    Supports the relationship between adverse impact, validation, business justification, and the duty to consider suitable alternatives with less adverse impact.

  3. EEOC, Section 15: Race and Color Discrimination

    Supports the U.S. disparate-impact framework, professional validation of a personality test used for management selection, and consideration of less discriminatory alternatives.

  4. U.S. Office of Personnel Management, Designing an Assessment Strategy

    Supports distinctions among reliability, validity, job analysis, personality-test prediction, applicant reactions, selection stages, and practical design of an assessment process.

  5. U.S. Office of Personnel Management, Are we allowed to use personality tests to assess candidates?

    Supports the need for work-related use, the distinction from psychiatric testing, and compliance with the Uniform Guidelines when personality tests affect employment decisions.

Apply it to your work

Understand how you work before you choose what comes next.

From this guide: Carry this report-reading question into the work decision in front of you.

Build a private Work Pattern Report across ten workplace continuums, then compare the result with the demands of the role or environment in front of you.