Performance Review Bias Statistics: What the Research Shows
Research-backed performance review bias statistics covering recency, gender, race, ratings, language and calibration, with original study links.
Performance review bias is real, but the evidence is more specific—and more useful—than lists claiming that every manager displays every bias. Field studies show measurable differences in ratings, “potential” judgments and feedback language. They also show that process design can change outcomes.
This page distinguishes direct performance-review evidence from broader decision-making research. It reports sample context and avoids treating a result from one company as a universal workforce estimate.
Key performance review bias findings
- A 2026 American Economic Review study analyzed 29,809 management-track employees at a large North American retail chain. (American Economic Association)
- Women received higher performance ratings but substantially lower potential ratings than men in that company. (AEA)
- Differences in potential ratings explained approximately half of the gender promotion gap. (AEA)
- The lower potential scores were not accurate forecasts: women subsequently outperformed male colleagues on average and at the promotion margin. (AEA)
- A summary of the working-paper results reported women were 7.3% more likely than men to receive a high performance rating. (Yale Insights)
- The same summary reported women's potential ratings were 5.8% lower. (Yale Insights)
- Textio examined actual feedback delivered to more than 25,000 people at 250 organizations in its 2022 language-bias research. (Textio)
- Women received 22% more personality-related feedback than men. (Textio)
- Women received 30% more exaggerated feedback than men. (Textio)
- Workers under 40 reported being described as “ambitious” 2.5 times as often as workers aged 40 or older. (Textio)
- Asian employees received 25% more feedback than White employees by volume. (Textio)
- Black men received one-third less feedback than White women, measured by word count. (Textio)
- Black women received nearly nine times as much non-actionable feedback as White men under 40. (Textio)
- White employees were described as “geniuses” 2.5 times as often as Black employees in Textio's analysis. (Textio)
Newer research on language and self-ratings
- Textio's 2024 study found women were seven times more likely than men to internalize negative stereotypes such as “emotional.” (Textio)
- White and Asian employees were twice as likely as Hispanic/Latino and Black employees to be positively stereotyped as “intelligent.” (Textio)
- Men were four times as likely as people of other genders to be positively stereotyped as “likable.” (Textio)
- The same report found high performers received the lowest-quality feedback in its dataset, with high-performing women experiencing the greatest problem. (Textio)
- A 2025 multinational-company field study found women rated themselves lower than men across years. (Journal of Economic Behavior & Organization)
- Women of color gave themselves the lowest self-ratings in that dataset. (Study PDF)
- Managers rated employees of color lower than White employees in the study, while White women were rated higher than men. (Journal of Economic Behavior & Organization)
- When managers could not see current self-ratings, manager and employee ratings became less correlated. (Study PDF)
- Removing current self-ratings did not eliminate the general gender and race gaps because managers could still anchor on earlier ratings. (Study PDF)
- Among employees without a prior company rating, the study found suggestive—not causal—evidence that women of color benefited when managers could not see their self-ratings. (Study PDF)
Structure, comparison and calibration findings
- Stanford research found that equally assertive behavior could be described and valued differently depending on employee gender. (Stanford Graduate School of Business)
- Stanford's organizational work emphasizes that ambiguous evaluation criteria create more room for stereotypes to shape judgment. (Stanford Sociology)
- In a Harvard Business School experiment involving 654 participants, evaluating candidates jointly rather than separately eliminated the stereotype effect observed in separate evaluation. The study concerned candidate evaluation, so it is supporting—not direct performance-review—evidence. (HBS Working Knowledge)
- In a study of scientific review panels, reviewers changed their initial score 47% of the time after discussion, showing how group deliberation can materially alter ratings. (Harvard Business School Alumni)
- In that study, reviewers changed scores for “superstar” scientists 24% less often than for other candidates. (Harvard Business School Alumni)
- Women reviewers changed their initial scores 13% more often than men. (Harvard Business School Alumni)
- 84.7% of organizations in the Talent Strategy Group's 2026 survey used performance calibration meetings. (Talent Strategy Group)
- 63.6% provided assessors with an expected performance distribution. (Talent Strategy Group)
- Of organizations using distributions, 34.1% enforced a fixed distribution and 65.9% used guidance with discretion. (Talent Strategy Group)
- 56.9% of surveyed organizations with overall ratings used a five-point scale. (Talent Strategy Group)
- 20% used a four-point scale and 16.2% used a three-point scale. (Talent Strategy Group)
- Rating structure alone is not a bias control: the evidence above shows that bias can enter through criteria, written language, self-ratings, historical anchors and “potential” judgments even when a numeric scale is standardized.
What reduces bias in performance reviews?
No single technique makes a review bias-free. The strongest process combines several controls:
- Define criteria before seeing the person or outcome. Use role-specific behaviors and results, not vague traits such as “leadership presence.”
- Collect evidence throughout the period. Time-stamped examples reduce dependence on recent or memorable events.
- Separate performance from potential. If potential affects promotion, define and audit it independently; the AEA study shows why this subjective label is high risk.
- Delay or mask self-ratings where appropriate. Do not let an employee's confidence level become an anchor for the manager.
- Review language. Flag personality terms, absolutes and feedback that contains no observable behavior or next action.
- Compare against the rubric before comparing people. Calibration should test consistent use of criteria, not force a predetermined distribution.
- Audit outcomes. Compare ratings, written feedback, promotions and compensation by relevant demographic groups and reviewer, while protecting privacy and meeting local legal requirements.
EvalFlow supports structured evidence, configurable performance reviews and consistent review workflows. Continue with our performance review statistics, employee feedback statistics, recency bias guide and fair performance review guide.
Research note
Several findings come from single-company field studies. They demonstrate mechanisms and measurable disparities in those settings; they are not universal population rates. Textio's studies combine survey reports with analyzed feedback documents and should be interpreted using its published methodology. Calibration prevalence is a practice benchmark, not proof that calibration removes bias. Causal language is used only where the underlying design supports it.
Frequently asked questions
What is the most common bias in performance reviews?
There is no credible universal ranking. Recency, halo, leniency, similarity and confirmation biases are all plausible, but demographic and language-bias studies measure different constructs. Audit your own process instead of assuming one bias is always dominant.
Do self-evaluations create bias?
They can create an anchor. The 2025 multinational-company study found manager ratings were less correlated with self-ratings when current self-ratings were hidden, but historical ratings preserved broader gaps. Masking self-ratings is therefore one control, not a complete solution.
Do calibration meetings eliminate bias?
No. Calibration can expose inconsistent standards and require evidence, but social influence and status can also affect group decisions. Use a trained facilitator, a shared rubric, written evidence and an outcome audit.
Can AI remove performance review bias?
AI can flag inconsistent language or missing evidence, but it can also reproduce patterns in its training data or organizational history. Keep criteria transparent, test outputs across groups and require accountable human review.