← Back to list

Not All Calibration Failures Mean the Same Thing

Calibration drift is often treated as a technical problem.

Sauran Sarbassov · 2026-06-04 11:14 · 0 claps · 8.9 min read
#credit-risk #model-validation #risk-management #machine-learning #data-science
Open on Medium ↗
Wiki topics: ML · Machine Learning BIZ · Business Strategy EDU · Education & Learning 🔬 · Science · General

Not All Calibration Failures Mean the Same Thing

Calibration drift is often treated as a technical problem.

Predicted PDs no longer match observed default rates, calibration metrics deteriorate, and the model is sent for recalibration.

However, calibration drift is not a diagnosis.

It is a symptom.

The same calibration failure may emerge for very different reasons. Sometimes economic conditions change. Sometimes the portfolio itself evolves. And sometimes calibration errors reveal something more concerning: the model is gradually losing its understanding of risk.

Understanding these differences is critical because the appropriate response depends on the underlying cause.

In this article, we will explore the most common sources of calibration drift and discuss what calibration errors may be trying to tell us about the portfolio and the model itself.

The Anchor Beneath Calibration

Discrimination determines who is riskier and who is safer.

Calibration determines how much risk each borrower carries.

As long as discrimination remains broadly intact, calibration can be adjusted to reflect changing conditions.

For this reason, discrimination acts as the stable foundation beneath calibration.

Figure 1 — Recalibration is appropriate only when discrimination remains stable. When ranking deteriorates, the model itself must be investigated or rebuilt.

Figure 1 — Recalibration is appropriate only when discrimination remains stable. When ranking deteriorates, the model itself must be investigated or rebuilt.

Why Calibration Exists

Once a model has learned who is riskier and who is safer, another problem remains.

The model must determine how much risk each borrower actually carries.

This is the role of calibration.

Calibration translates relative risk rankings into actual probability estimates.

As long as the model continues to understand who is riskier and who is safer, calibration is routinely adjusted to reflect changing conditions and keep model outputs aligned with observed portfolio behavior.

For this reason, calibration often becomes one of the earliest indicators that something in the portfolio, the environment, or the model itself is beginning to change.

Type 1 — Level Shift

Calibration drift is often treated as a single problem.

In reality, calibration failures emerge for very different reasons. Understanding these differences is the first step toward interpreting what calibration errors are actually trying to tell us.

The first and arguably most common type of calibration failure occurs when the overall level of risk changes while the underlying ranking structure remains largely intact.

This type of calibration drift does not necessarily indicate a problem with the model. In fact, the model may still be functioning exactly as intended. What has changed is not the model, but the environment in which the model operates.

Suppose a model was developed during a relatively stable economic period. Several years later, economic conditions deteriorate.

  • Interest rates rise.
  • Unemployment increases.
  • Corporate earnings weaken.

As a result, default rates begin increasing across the portfolio.

Importantly, however, the economic deterioration primarily changes the overall level of risk rather than the ordering of risk. Most borrowers become riskier at the same time, even if the magnitude of the deterioration differs across segments. Some borrowers may be affected more strongly than others, but the overall direction remains the same — in this example, the portfolio as a whole becomes riskier.

This is the essence of a level shift. The entire portfolio moves in the same direction. Borrowers may become riskier during an economic downturn or safer during an economic boom. The increase in PD may differ from borrower to borrower, but the ranking of borrowers remains broadly intact. Borrowers previously identified as high risk continue to default more frequently than borrowers identified as low risk.

What has changed is not the ranking structure, but the overall level of risk across the portfolio.

Figure 2 — In a level shift scenario, borrowers remain correctly ranked by risk, but the probability scale becomes misaligned with reality. Calibration deteriorates while discrimination remains largely preserved.

Figure 2 — In a level shift scenario, borrowers remain correctly ranked by risk, but the probability scale becomes misaligned with reality. Calibration deteriorates while discrimination remains largely preserved.

In this situation, calibration drift is not necessarily a sign that the model has become structurally unreliable.

The model continues to rank borrowers correctly.

Reality has simply moved to a different risk level than the one reflected in the original calibration.

When Calibration Drift Is Expected

At first glance, a level shift may appear to be a model failure. Predicted PDs no longer match observed default rates, and calibration appears to have deteriorated.

However, this type of calibration drift is often a natural consequence of changing economic conditions rather than a sign of structural model deterioration.

In fact, level shifts occur regularly throughout the life of most credit risk models. Economic conditions constantly evolve. Interest rates rise and fall, unemployment fluctuates, and asset prices increase and decrease over time.

As a result, the overall level of risk within a portfolio also changes over time. The model may continue ranking borrowers correctly, yet the probability of default associated with each rating grade gradually changes as economic conditions evolve.

Viewed from this perspective, level shifts are not exceptional events. They are a natural consequence of a changing environment.

For this reason, level-shift calibration drift is often the least concerning type of calibration failure.

Under IFRS 9, changes in forward-looking economic forecasts frequently lead to adjustments in the overall level of portfolio risk. Although practitioners often refer to this process as incorporating forward-looking information (FLI), from a calibration perspective it is fundamentally a form of level-shift recalibration.

The model’s ranking structure remains unchanged. What changes is the overall level of risk assigned to borrowers in response to updated economic expectations.

Type 2 — Population Shift

Not all calibration drift originates from changing economic conditions.

Sometimes the environment remains broadly stable while the portfolio itself gradually changes.

  • New products are introduced.
  • Underwriting standards evolve.
  • New borrower segments enter the portfolio.
  • Geographic exposure shifts.
  • Customer acquisition strategies change.

Over time, the population to which the model is applied may become materially different from the population on which the model was originally developed.

This type of calibration drift is often referred to as population shift.

Unlike a level shift, where the overall level of risk changes across most borrower groups simultaneously, population shifts occur because the composition of the portfolio itself changes.

Figure 3— Population shift may break both ranking performance and probability calibration, limiting the effectiveness of recalibration.

Figure 3— Population shift may break both ranking performance and probability calibration, limiting the effectiveness of recalibration.

A useful way to think about population shifts is to imagine applying a model developed for one portfolio to a slightly different portfolio.

The model may continue to rank borrowers reasonably well. The overall level of risk may even remain similar. However, the relationship between borrower characteristics and default behavior may gradually begin to change. As a result, calibration starts drifting even though the economic environment appears relatively stable.

Population shifts are particularly important because they often develop slowly and may remain unnoticed for long periods of time. Unlike economic shocks, which are usually visible to everyone, portfolio changes frequently occur through a series of small business decisions made over many years.

For this reason, population shifts often sit somewhere between routine calibration drift and genuine model deterioration.

They are not necessarily a sign that the model is broken.

But they may represent the first indication that the portfolio is gradually becoming different from the one the model was originally designed to understand.

Population Shifts Deserve Attention

Unlike a level shift, a population shift often deserves closer investigation.

A level shift is usually driven by changing economic conditions. The model itself may remain perfectly healthy.

Population shifts are different.

The calibration drift may still be relatively small, but it now reflects gradual changes in the portfolio itself. New products, new customer segments, or evolving underwriting standards may slowly move the portfolio away from the population on which the model was originally developed.

This matters because statistical models are fundamentally built on historical experience. A model learns patterns from the population it has seen during development and assumes that these patterns will remain broadly relevant in the future. The further the portfolio moves away from that original population, the less certain this assumption becomes.

For this reason, population shifts often sit somewhere between routine recalibration and genuine model deterioration.

The model may still be functioning well. Moreover, it may continue functioning well for some time.

But the calibration drift itself is a signal. It may provide one of the earliest indications that the portfolio is gradually becoming different from the one the model was originally designed to understand.

Type 3 — Structural Calibration Drift

The third type of calibration failure is fundamentally different from the previous two.

Both level shifts and population shifts are primarily driven by external changes rather than serious problems with the model itself.

In a level shift, economic conditions change and the overall level of risk moves accordingly. The model may still rank borrowers correctly and require only a relatively simple calibration adjustment.

In a population shift, the portfolio gradually becomes different from the one on which the model was originally developed. The model may become somewhat less accurate, but the underlying relationships often remain broadly valid. The model still understands risk reasonably well, even if it now operates on a slightly different population.

Structural calibration drift goes one step further.

Structural drift occurs when the relationships learned during model development no longer describe borrower behavior as accurately as before. The borrowers themselves may remain broadly similar, but the predictive power of the underlying relationships on which the model was built begins to deteriorate.

As a result, calibration errors start appearing in specific regions of the score distribution rather than across the portfolio as a whole.

Figure 4 — Structural calibration drift reflects non-uniform calibration errors across the risk spectrum. Some borrower segments become systematically overestimated while others become underestimated.

Figure 4 — Structural calibration drift reflects non-uniform calibration errors across the risk spectrum. Some borrower segments become systematically overestimated while others become underestimated.

This is where calibration diagnostics become particularly informative.

In the previous two examples, calibration drift was largely driven by external factors. The economy changed. The portfolio changed. The calibration error itself was not necessarily telling us that something was wrong with the model.

Structural drift is different. Here, calibration errors begin carrying information about the model itself.

Rather than simply indicating that risk levels have changed, calibration patterns may start revealing weaknesses in the relationships on which the model was built.

In this sense, calibration becomes a diagnostic tool.

The objective is no longer simply to determine whether calibration is wrong. The objective is to understand what the calibration error is trying to tell us.

The shape of the calibration error may provide important clues about what exactly is starting to break inside the model.

Structural drift often manifests itself through distinctive calibration patterns. In more advanced stages, S-shaped or U-shaped calibration curves may begin to emerge.

Unlike level or population shifts, these patterns often indicate that the model’s portrait of a risky borrower is becoming outdated. Such patterns deserve careful investigation because they may represent the early stages of model deterioration.

For this reason, structural calibration drift deserves closer investigation than routine calibration deterioration.

The goal is no longer simply to restore calibration.

The goal is to understand why specific parts of the portfolio are no longer behaving as the model expects.

Not All Calibration Failures Mean the Same Thing

One of the most common mistakes in model monitoring is treating all calibration failures as if they were identical.

In practice, calibration drift can emerge for very different reasons.

Sometimes the economy changes and the overall level of risk shifts across the portfolio.

Sometimes the portfolio itself evolves and becomes different from the one on which the model was originally developed.

And sometimes calibration errors begin revealing something more concerning: the model’s understanding of risk itself is gradually becoming outdated.

These situations may produce similar calibration statistics while requiring very different responses.

  • A level shift may justify a relatively simple recalibration.
  • A population shift may require closer monitoring and investigation.
  • Structural drift may ultimately require redevelopment of the model itself.

For this reason, calibration errors should not be viewed merely as problems that need to be corrected.

They should be viewed as signals that need to be understood.

The first question is not how to fix the calibration error.

The first question is why the calibration error appeared in the first place.

Only after understanding the cause can we determine the appropriate response.

Conclusion

Calibration drift is often treated as a technical problem requiring technical fixes.

In reality, calibration errors frequently contain valuable information about both the portfolio and the model itself.

Some calibration failures are routine consequences of changing economic conditions.

Others may represent the early stages of model deterioration.

Understanding the difference is one of the most important skills in model monitoring and validation.

Before correcting a calibration error, we should first understand what the error is trying to tell us.

Calibration Series

This article is part of an ongoing series on calibration, discrimination, and model deterioration.

  1. Why AUC (Gini) Is Not Enough: Understanding Calibration in Credit Risk. Why good ranking does not guarantee accurate PD estimates.
  2. Why Calibration Breaks Even When Your Model Is Correct Why calibration drift may occur even when the model itself remains valid.
  3. When a Model Looks Right But Is Already Wrong How structural deterioration can remain hidden behind stable portfolio metrics.
  4. When Calibration Breaks: What To Do Next A practical framework for diagnosing and responding to calibration failures.
  5. Why Calibration Cannot Save a Broken Ranking Why recalibration stops working when the underlying ranking structure begins to deteriorate.
  6. Not All Calibration Failures Mean the Same Thing Why level shifts, population shifts, and structural drift should not be interpreted in the same way.

메타데이터
post_id
0d9bbec073a5
slug
not-all-calibration-failures-mean-the-same-thing-0d9bbec073a5
url
https://medium.com/@supersauran/not-all-calibration-failures-mean-the-same-thing-0d9bbec073a5
canonical_url
https://medium.com/@supersauran/not-all-calibration-failures-mean-the-same-thing-0d9bbec073a5
author_url
https://medium.com/@supersauran
status
ok
fetched_at
2026-06-25 16:53:31