Why Your Performance Ratings Are Wrong And It’s Not Your Managers’ Fault
Quick Answer
Why Your Performance Ratings Are Wrong And It’s Not Your Managers’ Fault
Quick Answer
Performance ratings become inaccurate not because managers are bad at their jobs but because without a calibration process, every manager uses the same rating scale differently. A ‘4 out of 5’ from one manager means something completely different to a ‘4’ from another. The result is performance decisions that cannot be fairly compared across the organisation.
Three years ago I sat in a room with twelve managers after a performance review cycle. We had just finished collecting ratings across a 400-person organisation. The data looked clean. The dashboard was complete. Every employee had a rating.
Then we put the rating distributions on the wall.
One manager had given 90 percent of their team a 4 or 5. Another had given 70 percent of their team a 3. A third had given out the first 1-rating in company history — to a high performer everyone else in the room knew was exceptional.
Same scale. Same definitions. Twelve completely different interpretations.
That is not a manager problem. That is a calibration problem.
What Performance Calibration Actually Is
Performance calibration is the structured process of reviewing and normalising employee ratings across the full manager cohort before any performance decisions are made. The goal is not to change every rating. The goal is to ensure that a 4 from Manager A and a 4 from Manager B actually represent the same level of performance.
Without calibration, performance decisions compensation, promotions, development opportunities, PIPs are made on data that cannot be fairly compared. The employee who happens to have a lenient manager gets a higher raise. The employee who happens to have a severe manager misses a promotion they earned.
This is what I mean when I say the ratings are wrong. Not inaccurate in isolation. Incomparable in combination.
The Three Bias Patterns That Appear in Every Uncalibrated Review Cycle
Leniency bias
The most common pattern. A manager rates the majority of their team in the top two performance buckets regardless of actual performance distribution. The underlying cause is almost never dishonesty. It is that most managers have a natural tendency to protect their team and to avoid difficult conversations that a lower rating might require.
In data, leniency bias shows up as a manager whose rating distribution sits significantly above the cohort average. A manager where 85 percent of their team is rated 4 or 5 in an organisation where the expected distribution is 60 percent is exhibiting leniency bias. Whether the ratings are wrong is a separate question. Whether they are comparable is not.
Severity bias
The opposite pattern. A manager consistently rates below the cohort average. Often driven by high personal standards, a belief that ratings should be earned not given, or a misunderstanding of what the rating scale is measuring. Severity bias is less common than leniency but carries a higher cost employees rated by severe managers receive lower compensation and fewer development opportunities regardless of their actual performance.
Recency bias
The tendency to weight the most recent months of the review period more heavily than the full year. An employee who had a strong Q1 and Q2 but a difficult Q3 will be underrated if the manager’s primary frame of reference is Q3. An employee who underperformed for three quarters but had a strong final month will be overrated.
Recency bias is the hardest to detect in uncalibrated data because it does not show up as a distribution anomaly. It shows up as a rating that does not match the evidence when you look closely — which is exactly why calibration sessions that review the actual evidence, not just the ratings, catch it.
What a Calibration Session Is Supposed to Do
A well-run calibration session does three things. First, it surfaces the distribution of ratings across the manager cohort so that outliers are visible before decisions are made. Second, it creates a shared conversation about the evidence behind specific ratings so that managers can recalibrate against each other’s standards. Third, it produces a documented audit trail of the decisions made and the reasoning behind them.
The failure mode is when calibration sessions become political rather than analytical. When managers defend their ratings instead of examining them. When the session is run without pre-prepared data and becomes a three-hour argument without resolution. When the outcome is ‘everyone agrees to disagree’ rather than a set of defensible, normalised ratings.
The reason calibration sessions go wrong is almost always the same: the HR team did not have the pre-session data ready, so the session starts blind.
Why Manual Calibration Preparation Takes Two to Three Days
Here is what HR teams typically do before a calibration session without automated tooling. They export all ratings from the performance management system. They build a spreadsheet with ratings by manager, team, and department. They calculate distribution percentages for each manager. They identify outlier managers visually. They flag employees whose ratings seem inconsistent with their documented performance. They prepare slides or a shared view for the session.
For a 200-person organisation with 15 managers, this takes two full working days. For a 500-person organisation with 30 managers, it takes three to four days. For a lean HR team of two or three people, this preparation is the largest single time cost in the entire review cycle.
And then they do it again six months later.
What Changes When Pre-Session Analysis Is Automated
At PerformSpark, I built TrAI specifically to solve this problem. Before every calibration session, TrAI runs automatically. It analyses rating distributions across the full manager cohort. It flags managers whose distributions deviate significantly from the cohort norm. It identifies leniency bias, severity bias, and recency patterns. It generates the shared pre-session view that HR needs without any manual preparation.
The calibration session still requires HR leadership. The decisions still require human judgment. The conversation between managers still requires skilled facilitation. What is eliminated is the two to three days of manual data preparation that happened before anyone walked into the room.
The HR teams using PerformSpark do not spend less time on calibration. They spend more time on the part that matters the conversation, the judgment, the documentation. Less time on the part that does not — building spreadsheets.
Internal link
For a full walkthrough of how TrAI generates pre-session calibration analysis, see the Performance Calibration guide at performspark.ai/calibration
The Three Things That Make a Calibration Session Work
• Pre-session data that is ready before anyone walks in: Rating distributions, outlier flags, and bias patterns prepared and shared before the session starts. Not assembled during the session.
• A shared definition of what each rating level means: Before the session, every manager must have reviewed the rating scale definitions. Calibration against different definitions is not calibration.
• A documentation system that records the decisions: Every rating change made in calibration, every flag discussed, every decision deferred — all logged with a timestamp. This is the audit trail that protects the organisation in a legal challenge and enables year-on-year calibration consistency.
What HR Leaders Should Ask in Every Calibration Session
The most useful questions to ask when reviewing individual ratings in a calibration session are not evaluative. They are evidential. ‘What specific examples support this rating?’ is more useful than ‘Do you agree with this rating?’ ‘How does this compare to the team member you rated a 3 in this same competency?’ is more useful than ‘Is this fair?’
The goal of calibration is not consensus. It is defensibility. A calibrated rating does not need every manager to agree with it. It needs the decision-making process behind it to be documented, evidence-based, and applied consistently across the cohort.
The Bottom Line
Your performance ratings are wrong not because your managers are wrong but because without a calibration process, ‘wrong’ is not a meaningful concept. Each manager’s ratings are internally consistent. They are not externally comparable. Calibration is what makes them comparable.
For HR leaders running performance cycles manually, the calibration preparation problem is real. The solution does not require changing your methodology. It requires changing how that methodology is prepared and facilitated.
If you are running calibration across 15 or more managers and spending two or more days building the pre-session data view, that time cost is worth examining. The preparation is not the valuable part. The conversation is.
Mahesh Kumar is the CEO of PerformSpark, a performance management platform for mid-market HR teams. PerformSpark’s TrAI engine automates pre-session calibration analysis so HR teams can focus on the decisions, not the data preparation. Learn more at performspark.ai/calibration
메타데이터
- post_id
- 2e1b9a29afbf
- slug
- why-your-performance-ratings-are-wrong-and-its-not-your-managers-fault-2e1b9a29afbf
- url
- https://medium.com/@performspark_maheshkumar/why-your-performance-ratings-are-wrong-and-its-not-your-managers-fault-2e1b9a29afbf
- canonical_url
- https://medium.com/@performspark_maheshkumar/why-your-performance-ratings-are-wrong-and-its-not-your-managers-fault-2e1b9a29afbf
- author_url
- https://medium.com/@performspark_maheshkumar
- status
- ok
- fetched_at
- 2026-06-09 15:37:30