Research — SEP 30, 2026
From Signal to Action: Selecting Thresholds for Credit Transition Signals
By Arun Kumar Singh and Michelle Cheong
Executive summary:
In today's complex global credit environment, market participants face an escalating challenge: navigating a rapidly shifting risk landscape. To address this challenge, S&P Global Market Intelligence developed Credit Transition Signals (follow the link here and select the analytics tab), a quantitative, model-driven surveillance tool designed to generate forward-looking probabilities of deterioration and improvement. By integrating ratings history, company fundamentals, and macroeconomic factors with real-time, market-implied signals, Credit Transitions Signal provides a continuous, high-frequency view of credit migration risk for rated entities; and transparency on key drivers of these measures.
Yet a fundamental question remains for practitioners: when is a credit transitions signal high enough to act on? This article provides some empirical evidence which clients can use to support operating ranges to act on this signal, grounded in empirical analysis. In this instalment, we will explore absolute thresholds. In the follow up article, we will look at credit rating level specific thresholds which can be useful for identifying potential fallen angels for example.
Key findings:
The data indicates that the candidate operating ranges for 12-month horizons are:
- For deterioration, 20%–30% can be a candidate operating range. It maximizes the signal-to-noise ratio while preserving enough alert volume for statistically reliable, high-conviction signals. For improvement, a range surrounding 20% can be considered.
- Calibrate operating ranges to your use case:
- Risk management teams, especially those with downstream processes or other metrics to further investigate alerts, could anchor toward 20% for broader number of alerts. This helps preserve capital and mitigate risks
- Use cases that require higher level of confidence (e.g. investment idea generation) can consider higher thresholds (e.g. > 40%), with smaller number of alerts.
Mapping the Distribution of Credit Transitions Signal Probabilities
Before testing the decision rules, we examine the distributional properties of Credit Transitions Signal across the rated universe. Figure 1 displays the density and frequency histograms for both the Probability of Deterioration (left) and the Probability of Improvement (right) over a 12-month horizon. Our sample consists of 16,930 probabilities of improvement and deterioration from 1,299 unique entities from June 2024 to July 2025, where we can map these to realized ratings changes over the subsequent 12-month periods.
Figure 1: Statistical Distribution of 12-Month Credit Transition Probabilities
Source: Credit Transition Signals on RatingsDirect® on Capital IQ Pro from S&P Global Market Intelligence. As of September 2026. For illustration only.
Three observations emerge from these distributions:
- Severe Right-Skewness (Positive Skew): Both distributions exhibit a right-skew, with most probabilities of improvements and deteriorations between 0% to 5%. This is consistent with the ratings migration patterns of investment-grade issuers that form most of our sample – where credit ratings are mostly unchanged over a 12-month period.
- Observed distribution differences between Deterioration vs. Improvement: While both charts are skewed, the deterioration distribution (left, brown) features a slightly thicker tail extending into the 30% and 40% ranges. Conversely, the improvement distribution (right, blue) tapers more quickly, with fewer entities registering elevated probabilities in our sample period.
Key takeaway: Given the right-skew of the CTS data, an absolute probability of 20% or 30% is at the upper percentiles of transition risk relative to the broader universe. Therefore, we can consider focusing on the 10% to 40% range.
Testing Hypothesis 1: which CTS thresholds have discriminatory power?
A straightforward approach is to define a probability threshold above which a signal is considered actionable, for example:
"If the Probability of Improvement exceeds X%, or the Probability of Deterioration exceeds X%, look closer at the entity’s fundamentals and other drivers."
We want to test whether entities with probabilities of improvement (deterioration) exceeding X% would see positive (negative) ratings trends vs. those below X%.
We will test using the following logistic regression for a particular threshold (X%)
- Indicator of credit ratings upgrade in 1-year1 (t+12 Months) = a0 + a1 x Indicator (Probability of Improvement in 12 months exceeds X%) + residual
- Indicator of credit ratings downgrade in 1-year (t+12 Months) = b0 + b1 x Indicator (Probability of Deterioration in 12 months exceeds X%) + residual
Statistical significance (Z-statistics) of the logistic regression coefficients a1 and b1 evaluated across absolute probability thresholds in 10% increments i.e. X = 10%, 20%, 30%, 40%, 50%,…,90%.)2.
Our findings are as follows:
1. All thresholds above 10% are effective in identifying early warnings of credit risk changes
In Figure 2, the horizontal dashed black line represents the 95% statistical significance threshold (z = 1.96, p < 0.05). Strikingly, both the Probability of Improvement and Probability of Deterioration Z-statistics remain substantially above this critical boundary for all thresholds ranging from 10% to 90%. From an application point of view, Z-statistics show the model's predictive validity is robust across the spectrum (i.e. any threshold above 10% should be useful), even though practical discriminatory power naturally varies depending on the chosen threshold.
Figure 2: Statistical Significance of Credit Transition Signals Across Absolute Probability Thresholds
Source: Research from Credit Solutions Thought Leadership. The underlying data is from Credit Transition Signals on RatingsDirect® on CaptialIQ Pro. Credit ratings history is sourced from RatingsXpress® and CreditPro® from S&P Global Market Intelligence. As of September 2026. For illustration only
2. AUC curves confirm strong discriminatory power
In Figure 3, we plotted Receiver Operating Characteristic (ROC) curves for Probability of Deterioration model (left graph) and Probability of Improvement model (right graph) across probability threshold. The dashed diagonal line represents a random baseline (AUC = 0.50) where the model has no estimating power. The more concave the ROC curve becomes, the better it is in estimating credit transitions.
Both the Probability of Deterioration model (red, AUC = 0.874) and the Probability of Improvement model (black, AUC = 0.817) demonstrate robust predictive capability, shifting notably toward the top-left quadrant. An AUC exceeding 0.80 suggests that the CTS algorithm reliably ranks transitioning entities higher than stable ones.
Figure 3: Receiver Operating Characteristic (ROC) Curves with Absolute Threshold Operating Points
Source: Research from Credit Solutions Thought Leadership. The underlying data is from Credit Transition Signals on RatingsDirect® on CaptialIQ Pro. Credit ratings history is sourced from RatingsXpress® and CreditPro® from S&P Global Market Intelligence. As of September 2026. For illustration only
2. Predictive Confidence is not uniform
Figure 2 also shows that although multiple thresholds from 10% to 90% have discriminatory power, their capability of estimating confidence varies, where Z-statistics exhibit a concave pattern that peaks at probability thresholds between 20% to 30%.
- Below 20%: At the 10% threshold, the statistical power is lower because the signal is overly sensitive, capturing too many credits with low likelihood of credit transitions (false positives).
- Above 30%: As the threshold becomes increasingly strict (30% to 90%), the Z-statistics steadily decay. This occurs because extreme probabilities are exceptionally rare. Whilst most of these lead to credit ratings in the same direction), the Volume of alerts (size of alerts to act on) shrinks dramatically, causing the statistical confidence (Z-statistic) of the regression model to wane.
Finding Potential Candidate Operating Ranges
Hence, there is a balancing act between capturing of an entity before ratings transitions and having too many signals in your portfolio (i.e. alert fatigue).
To find our candidate operating ranges for X, we first assess a threshold-independent discriminatory power of the Credit Transition Signal (CTS) model. In addition to ROC and AUC, Figure 3 illustrates the trade-off between the True Positive Rate (Recall) and the False Positive Rate (Noise) at each threshold.
Several additional insights emerge from the ROC charts:
- Confirming that 20% - 30% is a candidate operating range
By projecting the absolute thresholds directly onto the ROC curve, the inflection point, often referred to as the "elbow" is clearly illustrated. In both panels, the 20% and 30% thresholds cluster near this apex. This coordinate represents where the model maximizes the capture of true transitions without sliding horizontally into a disproportionate volume of false positives. - The Trade-off between Lenience and Strictness
The annotated points highlight the operational impact of deviating from the 20%–30% band.- Increased Lenience (10%): Lowering the threshold to 10% shifts the operating point horizontally to the right. The model achieves only a fractional gain in captured transitions while absorbing a notable increase in false positives (noise).
- Increased Strictness (40%+): Conversely, tightening the threshold from 30% to 40%, 50%, and beyond causes the operating points to move sharply down the Y-axis. While demanding this level of certainty effectively minimizes false positives, it meaningfully impairs the model's recall, limiting its ability to capture a broad set of valid market transitions.
- Furthermore, this visualization offers a quantitative rationale for not having overly restrictive (high) thresholds. For example, a threshold above 60% may miss many early-warning signals provided by the model.
Probability of Deterioration Model: Precision, Recall & F1
How can we measure this trade-off and find a candidate operating range? In statistics, Precision and Recall can be used to measure such tradeoffs:
- Precision measures the proportion of predictions that were correct. In this context, it answers: "Of all the times the model estimated a ratings downgrade in a 12-month horizon, how often was it right?"
- Recall measures the proportion of actual positives that were correctly identified by the model. In this context, it answers: "Of all the actual ratings downgrades in a 12-month horizon, how many did the model successfully estimate it ?"
In Figure 4, we empirically assess the tradeoff between signal accuracy (Precision/Hit Rate) and sensitivity (Recall/realized downgrades flagged by the model) across absolute probability thresholds ranging from 10% to 90%. This analysis will be repeated for upgrades in the next section (Figure 5).
We can see the typical tradeoff between precision and recall. Setting a high threshold (i.e. adopting a conservative approach) means the model only signals a downgrade when highly confident, resulting in high precision (few false positives) but low recall (missing more actual downgrades). Lowering the threshold allows the model to flag downgrades more freely, increasing recall (fewer missed downgrades) but decreasing precision (more false positives).
Figure 4: Probability of Deterioration Model: Precision, Recall & F1
Source: Research from Credit Solutions Thought Leadership. The underlying data is from Credit Transition Signals on RatingsDirect® on CaptialIQ Pro. Credit ratings history is sourced from RatingsXpress® and CreditPro® from S&P Global Market Intelligence. As of September 2026. For illustration only
Key Insights for Candidate Operating Ranges for Probability of Deterioration:
1. Balancing Signal Sensitivity and Accuracy
- At low thresholds (e.g. 10%), the model achieves high recall, capturing most realized downgrades, but precision drops. This results in many false positives among stable entities.
- Increasing the threshold to 40–50% significantly boosts precision, ensuring alerts are more likely to reflect actual deterioration, but recall declines, leaving more downgrades unflagged.
2. Optimizing Precision and Recall
There are two common approaches to balance these metrics:
- The intersection of precision and recall curves, where the risk of false alerts equals the risk of missing real downgrades. In Figure 4, this threshold is between 30–40%.
- If precision and recall are imbalanced (i.e. very different) at this intersection, the F1 statistic (harmonic mean of precision and recall) may be more informative. The candidate operating range based on F1 is 20–30% (grey area in Figure 4), which is useful for users with lower tolerance for alert fatigue. This approach maintains high recall while ensuring sufficient precision.
Probability of Improvement Model: Precision, Recall & F1
We apply the same precision-recall framework to probability of improvement. Figure 5 maps Precision (Hit Rate) against Recall (sensitivity) across absolute thresholds.
Figure 5: Probability of Improvement Model: Precision, Recall & F1
Source: Research from Credit Solutions Thought Leadership. The underlying data is from Credit Transition Signals on RatingsDirect® on CaptialIQ Pro. Credit ratings history is sourced from RatingsXpress® and CreditPro® from S&P Global Market Intelligence. As of September 2026. For illustration only
We find that:
- Balance between signal sensitivity and precision
The Precision and Recall curves intersect at an absolute probability threshold of ~20%, where both precision and recall are 50%, and F1 peak score is 49%. This represents a somewhat lower crossing point than observed for the Probability of Deterioration. - Gradual Precision Gains Across Higher Thresholds
While increasing the threshold steadily improves the signal's hit rate, the ascent of the Precision curve is relatively gradual, leveling off between 40% and 50% at higher thresholds. - Potential Operational Anchor for Investment use cases
For front office use cases, a 20% absolute threshold can be used as a cutoff. In our sample period, approximately 50% of entities breaching the threshold resulted in upgrades within a 12-month period (precision). Interestingly, recall is also 50%, which indicates that this threshold managed to capture 50% of the actual upgrades in a 12- month period.
Areas of future research
A more granular look at credit ratings transitions by credit ratings in our sample period (Figure 6) shows that speculative-grade ratings (BB+ and below) have much higher realized transition rates than that of investment-grade ratings (BBB- and above); more so as credit quality declines into high-yield and CCC/CC territory.
This points to a potential refinement of the approach set out in this paper. Specifically, to construct candidate operating ranges for each rating level: with lower thresholds for stable investment-grade ratings and higher thresholds for more volatile speculative-grade ratings. Such a rating-specific calibration could be especially valuable for flagging "fallen angels" (BBB- credits at risk of falling into high-yield status). We will explore this in our follow-up paper.
Naturally, as the sample data builds up to cover a full credit cycle, we can fine-tune these candidate operating ranges.
Figure 6: Relationship between historical 1-Year credit ratings Transition Rates by ratings level
Source: Research from Credit Solutions Thought Leadership. The underlying data is from Credit Transition Signals on RatingsDirect® on CaptialIQ Pro. Credit ratings history is sourced from CreditPro® from S&P Global Market Intelligence. As of September 2026. For illustration only
Conclusion and Applications:
In this article, we perform analysis to explore potential thresholds and decisions rules to decide what is potentially high enough level to act on credit transitions improvements and deteriorations. Based on a data sample from year 2024 – 2025 suggests that setting the candidate operating range at 20-30% for downgrades and 20% to trade-off between volume of alerts and missing out on potential changes in credit risk.
The thresholds can also differ for different institutional mandates and use cases. A conservative risk management team prioritizing capital preservation (aiming to minimize False Negatives/missed downgrades) could anchor their decision rule at a lower threshold to capture a higher volume of early warnings. On the other hand, investment use cases may wish to set higher thresholds to focus on entities which are highly likely to move in the direction of improvements or deteriorations; and filter on those which have not yet been priced in the prices.
Important Disclosures: Credit Transition Signals are not created or maintained by, associated with, or representative of any S&P Global Ratings credit rating or Outlook. The product is produced by S&P Global Market Intelligence, a division independent from S&P Global Ratings. S&P Global Ratings makes no representation regarding the product. The product is a statistical model that considers input based on publicly available information. The product is representative of probabilities and is not a prediction; it is not informed by knowledge of upcoming rating actions and is not a guarantee or indication of a future rating action. A credit rating action, if taken, may differ regarding timing, extent of movement, and/or direction compared to the product. The product may diverge from the Outlook published by S&P Global Ratings. For questions, contact S&P Global Market Intelligence at Capital IQ Pro Client Support. For more information, see the Credit Transition Signals Whitepaper.