Who Decides What Creditworthiness Means?
Executive Summary
Pedro Lange Machado’s 2026 paper argues that sovereign credit rating agencies embed a market-liberal theory of the state into the technical apparatus of creditworthiness, rewarding fiscal restraint and creditor protection while penalising the debt restructuring and public investment that climate adaptation requires. This essay accepts that diagnosis (for the purposes of asking a different question) and asks the question Machado leaves open: who, exactly, holds the authority to change it?
The answer requires separating three things commonly collapsed into one. Authorship of the concept, which characteristics count as evidence of a government's capacity to repay, is shared among the agencies, the Fund, the World Bank and a long professional tradition. Authorship of the operationalisation - the scorecards, weights, caps and overlays that turn those characteristics into a number - belongs substantially to S&P Global, Moody’s, and Fitch, and the essay documents that architecture in detail. Authorship through recognition, whether a rating is treated as meaningful by Basel weightings, index rules and portfolio mandates, belongs almost entirely to the wider financial system the agencies do not control.
The evidence bears this distinction out. The World Bank’s governance indicators sit inside the agencies’ scorecards as validated, weighted variables; climate risk does not, not because it is suppressed, but because it has not yet accumulated the decades of tested track record against realised outcomes that governance indicators have. Oligopoly, far from obviously sheltering the incumbents to experiment, is a structure built by regulators in 1936 and 1975 that the empirical record associates with eroded rather than sharpened rigour, though the commercial incentive this creates remains an inference rather than a demonstrated fact. Governments retain real agency to perform well against the given categories, as EU accession states did, but the evidence that any government has redefined those categories - AfCRA included - is considerably weaker.
The conclusion this essay reaches is that Machado’s market-liberal content is real, but its persistence is better explained as an institutional settlement distributed across raters, investors, regulators, and professional communities than as a conviction lodged inside three firms. That distinction matters practically. Persuading a rating committee to think differently would not, on its own, change what creditworthiness means, so long as the settlement around it continues to treat the old definition as the reasonable one.
Who Decides What Creditworthiness Means?
Machado’s recent paper for the Interdisciplinary Observatory on Climate Change at UERJ makes a claim that is easy to state and difficult to dismiss. The three dominant credit rating agencies - S&P Global, Moody’s and Fitch - evaluate the creditworthiness of sovereign states using frameworks built inside a particular understanding of how economies function, one in which fiscal restraint, private capital mobility, and the protection of creditor claims are treated as the ordinary preconditions of financial stability. Machado does not accuse the agencies of conspiring against climate policy: ‘promoting a just green transition is not part of the mandate of CRAs... they are private firms whose behaviour is shaped by institutional incentives, regulatory frameworks, and the need to maintain credibility with investors’ (Machado, 2026, p. 11). Sovereign rating methodologies incorporate climate risk on a short time horizon that privileges near-term fiscal and market effects over long-run physical and transition consequences. They treat fiscal consolidation as the default route to improved metrics, in a manner one of Machado’s sources, quoting Wolfgang Streeck, describes as an operationalisation of ‘bondholder value’, in which governments are expected to ‘persuade or compel their citizens to moderate their claims on the public purse’ in favour of the financial markets (Streeck, 2014, p. 92, quoted in Machado, 2026, p. 9). They treat debt restructuring and debt-for-nature swaps as departures from creditor rights warranting a negative response, so that Moody’s classified Ecuador as in default after it completed such a swap in 2023, while Colombia reportedly abandoned a comparable initiative over concern for its own rating (Machado, 2026, p. 10). In electoral contexts their commentary has treated market-friendly candidates as sources of confidence and their left-leaning opponents as sources of risk, as when Fitch’s analysts wrote in September 2018 that a Bolsonaro victory would see markets rally on his ‘investor-friendly advisors’, while a Haddad victory would send ‘bond yields’ spiking and ‘equities’ sharply lower (Fitch Ratings, 2018, quoted in Machado, 2026, p. 10). None of this required an analyst to hold a political opinion about the green transition. It required only that the apparatus used to answer a narrower question - will this government service its debt - contain assumptions about serviceability that happen to coincide with a programme commonly called market-liberal.
Machado’s paper belongs to a substantial literature that treats sovereign credit rating as a site of private authority within global finance, disseminating norms and institutional templates that governments must observe to retain access to capital on tolerable terms (Sinclair, 2005; Paudyn, 2014; both cited in Machado, 2026). This literature is right to insist that sovereign risk assessment is not a neutral technical exercise reading default probabilities off public accounts. Debt ratios, reserve levels, growth rates, and fiscal balances are observable. What they mean for a government’s capacity to service debt over a five- or ten-year horizon is not observable in the same way; it has to be inferred through a theory, stated or simply embedded in a model. Machado’s contribution is to show, carefully, that the theory embedded in sovereign rating practice at the Big Three tracks a recognisable programme of assumptions about fiscal discipline, private finance, and the proper size of the state.
The question this leaves open is one Machado does not set out to answer, and which this essay takes as its subject. If an analyst at S&P Global, or a rating committee at Moody’s, concluded that one of these embedded assumptions was incomplete, that a decade of austerity had produced not stability but eroded tax capacity and stranded human capital, could the agency simply change its framework to reflect that judgement? Of course a private company can revise its own methodology, and all three do so periodically. The harder question is whether a changed mind, on its own, would be sufficient to change what creditworthiness means in practice. Sovereign ratings are numbers wired into bond covenants, central bank collateral frameworks, pension mandates, Basel capital weightings, and index eligibility rules (International Monetary Fund, 2010, pp. 91 to 92), compared continuously against the ratings of two other agencies applying frameworks that are broadly similar because they were built inside the same professional tradition and are monitored by the same handful of regulators. To ask whether one agency could redefine creditworthiness is to ask a question about authorship that reaches well beyond any single rating committee.
Where the Framework Actually Lives
The published methodologies of the three agencies are unusually detailed for proprietary commercial documents – particularly since regulation has mandated that these documents be made wholly public - and reading them together clarifies how much interpretive apparatus sits between an observable fact and a rating. S&P Global’s sovereign methodology scores five components - institutional, economic, external, fiscal (split into performance and flexibility, and debt burden), and monetary assessment - each on a six point scale, combines two composite scores through an indicative rating matrix, and then permits supplemental adjustments subject to caps: an institutional assessment at the weakest end of the scale caps the sovereign rating at BB+ regardless of how the other four components score (S&P Global Ratings, 2026, para. 126). Moody’s is more explicit about its arithmetic. Economic strength and institutional and governance strength combine at equal weight into what it calls economic resiliency; resiliency then combines with fiscal strength, reweighted almost entirely toward debt burden for reserve-currency issuers, to produce government financial strength; a final adjustment for susceptibility to event risk is aggregated not by averaging but by taking the minimum across its four components, on the explicit ground that ‘the materialisation of even one of these risks can lead to a severe deterioration of a sovereign’s credit profile’ (Moody’s Ratings, 2026), risk behaving like a chain rather than a portfolio, itself a substantive causal claim about how sovereign crises unfold. Fitch discloses the most quantitatively explicit architecture of the three: an eighteen-variable regression in which structural features carry just under fifty four per cent of the weight, public finances just over nineteen, external finances just over seventeen, and macroeconomic performance under ten (Fitch Ratings, 2025, p. 6), followed by a Qualitative Overlay through which committee judgement can move the output by up to two notches within any pillar and three overall. Below the CCC+ threshold Fitch abandons the model altogether: ‘Fitch does not use the SRM and QO. Instead, ratings are directly based on Fitch’s Ratings Definitions’ (Fitch Ratings, 2025, p. 2), an admission that the entire quantitative apparatus simply stops applying once a sovereign crosses into serious distress.
What this comparison shows is that a sovereign rating emerges from a layered system rather than a single judgement: a scoring rule for each input, a combination rule across inputs, a discretionary overlay with its own bounds, and beneath all of it a rating committee whose deliberations are, by the agencies’ own account, only partially disclosed. The IMF’s 2010 review of the industry put the point plainly, describing ratings as combining quantitative inputs with a final committee judgement about a government’s ‘willingness to pay’, a variable that cannot by its nature be read off a balance sheet (International Monetary Fund, 2010, pp. 98 to 99). Kerwer observed that the industry has never developed anything resembling the agreed first principles that discipline financial reporting: ‘there exists nothing like a Generally Agreed Rating Principles’ (Kerwer, 2005, p. 472). An agency’s conception of creditworthiness can therefore shift at several points without any public act of reconceptualisation: a pillar reweighted, an overlay applied more generously in one direction within limits that are stated but not narrowly binding, a cap relaxed in a routine revision rather than announced as a change of philosophy. Much of what would have to change for an agency’s understanding of creditworthiness to change is therefore the pattern of discretion exercised inside a structure the criteria only partly specify, rather than the published criteria themselves.
Knowing Whether It Works
An agency that revises its framework needs some way of judging whether the revision is an improvement, and the evidence on how that judgement is made is narrower than either defenders or critics of the industry usually allow. S&P Global’s annual sovereign default and transition study measures discriminatory power through a Gini accuracy ratio computed from cohorts of issuers grouped by year-end rating and tracked forward, a ‘static pool’ method that produced a one-year ratio of 94.8 per cent for foreign currency sovereign ratings in 2024, above the long run average of 90.2 per cent (S&P Global Ratings, 2025). S&P bounds what this exercise establishes with some care: ‘the use of the term methodology in this article refers to data aggregation and calculation methods used in conducting the research. It does not relate to S&P Global Ratings’ methodologies, which are publicly available criteria used to determine credit ratings’ (S&P Global Ratings, 2025). A high Gini ratio shows that the ranking of sovereigns correlates with subsequent default. It says nothing about whether the weight placed on, say, fiscal consolidation relative to public investment capacity is the right one, since a differently weighted framework could in principle discriminate just as well, provided it were applied with equal consistency.
The regulatory response to the 2008 crisis pressed further into this gap, with mixed results. ESMA’s guidelines on methodology validation require agencies to demonstrate discriminatory power, predictive power, and historical robustness with ‘sufficient quantitative evidence’ (European Securities and Markets Authority, 2017, paras. 17 to 27), while leaving the threshold at which a poor result should trigger review to the agency’s own review function, and stating that a breach does not automatically compel a change (European Securities and Markets Authority, 2017, paras. 33 to 38). The feedback statement behind these guidelines records a genuine dispute about what validation is for. One agency objected that sovereign ratings ‘are forward-looking opinions about unlikely events, express creditworthiness as a relative rank order and are not predictive of a specific frequency of default or loss’, warning that predictive-power testing risked interfering with the ratings themselves. ESMA held the requirement while conceding the underlying point, replying that the guidelines ‘would not change the product that CRAs issuing ordinal credit ratings offer’ (European Securities and Markets Authority, 2016), a rare moment in which the limit of an external actor’s power to redefine what a rating claims to be becomes visible, and is then respected.
The performance record itself resists a single verdict. An IMF study covering 2005 to 2010 found sovereign accuracy ratios of 80 to 92 per cent, stronger than the equivalent corporate figures, with no investment grade sovereign defaulting over the period (Kiff, Nowak and Schumacher, 2012, p. 18), but also found pronounced stability failure during the 1997 to 1998 Asian crisis and the 2008 to 2010 global crisis, multi-notch downgrades concentrated among issuers previously rated in the more stable categories (Kiff, Nowak and Schumacher, 2012, pp. 21 to 24). Sovereign ratings are deliberately smoothed through the cycle, trading discrimination for stability in ordinary times at the cost of the cliff effects that follow once the smoothing breaks (International Monetary Fund, 2010, pp. 90 to 91).
A further possibility deserves honest treatment rather than dismissal or endorsement. If a downgrade tightens fiscal space, and tighter fiscal space then degrades the indicators the rating is meant to track, the rating and the reality it describes could move together in a self-confirming loop, close to what Barta means in describing credit rating agencies’ climate strategy as putting the ‘Tragedy of the Horizon... on steroids’ rather than helping to break it (Barta, 2026, p. 4, quoted in Machado, 2026, p. 6). This is a serious hypothesis, not an established finding about sovereign ratings specifically. The most careful empirical demonstration of a public evaluative measure reshaping the field it measures concerns United States law school rankings rather than sovereign finance: Espeland and Sauder traced cross-admit yields shifting from roughly a third choosing the higher-ranked school to between eighty and ninety per cent within a generation, driven by self-fulfilling prophecy and commensuration, and noted that the strength of this reactivity depended on facing a single dominant ranking rather than several competing ones (Espeland and Sauder, 2007, pp. 13 to 14, 32 to 33), a condition sovereign ratings do not meet. MacKenzie’s account of the Black Scholes formula, the most careful precedent for financial performativity, is equally careful about its limits, calling the effect ‘incomplete and historically specific’ and showing that fit with market prices deteriorated sharply after 1987 (MacKenzie, 2003, pp. 11, 55); what sustained that loop was continuous arbitrage, a mechanism sovereign ratings lack, though reserve managers surveyed after 2008 did name downgrades as their most common trigger for reallocation, a single A rating often an informal floor (Morahan and Mulder, 2013, pp. 15, 17 to 18). The mechanism is real and worth watching, but remains bounded and situational rather than the general reflexive dynamic a stronger reading of Machado’s argument might imply.
What Can Enter: Governance Indicators and the Limits of Climate Data
Peter Haas, writing about the international policy influence of expert networks rather than credit rating, drew a distinction that is useful here. An epistemic community is defined by shared causal beliefs and ‘internally defined notions of validity’, and its influence depends on gaining a foothold, in his phrase, by ‘occupying niches in advisory and regulatory bodies’ (Haas, 1992, pp. 3, 30). A body of knowledge can therefore be credible among specialists, epistemically admissible in that experts would accept it as sound, well before it becomes institutionally admissible, incorporated into the routines and comparability requirements of organisations that would need to use it. The World Bank’s Worldwide Governance Indicators illustrate a body of data that has cleared both hurdles. Moody’s scorecard names specific WGI thresholds directly: a government effectiveness or regulatory quality score above 1.5 maps to the strongest available subfactor grade (Moody’s Ratings, 2026). Fitch folds a composite of six WGI percentile ranks into its regression as a single variable carrying just over twenty-two per cent of the weight within its largest pillar (Fitch Ratings, 2025, pp. 8 to 9). Two IMF studies help explain how this indicator earned that position. Keita, Leon, and Lima found that a one standard deviation improvement in the WGI government effectiveness score raises the probability that a country has an internationally recognised rating at all by around thirty per cent, and improves ratings by roughly 1.3 notches for countries starting from weak governance scores (Keita, Leon, and Lima, 2019, p. 10); Arbatli and Escolano found comparable effects for a related IMF-constructed transparency index (Arbatli and Escolano, 2012, pp. 3 to 4). What these findings do show, beyond the correlation itself, is that the World Bank’s indicators have been tested repeatedly, across decades and a large panel of countries, against outcomes the agencies and the Fund both care about. Espeland and Stevens, writing about commensuration generally, note that a metric’s constitutive power over the field it measures depends on how far it has become institutionalised through repeated reliance across many independent users, rather than being a fixed property of the metric itself (Espeland and Stevens, 1998, pp. 316, 328 to 329). The WGI has, in effect, been through this kind of validation many times over, by many hands, long enough to accumulate the standing that both testifies to and constitutes its institutional position.
Set against this, the treatment of climate risk in the same three methodologies looks less like suppression and more like an unfinished admissions process. S&P Global’s criteria mention climate in three narrow places: as a possible influence on the institutional assessment through ‘policies to reduce dependence on sectors at risk from longer term energy transition’ or to ‘mitigate the adverse physical effects of climate change’, as a drag on the economic assessment through exposure to ‘natural disasters or adverse weather conditions’, and as a source of revenue volatility within the fiscal debt burden assessment (S&P Global Ratings, 2026, paras. 21, 41, 80), none a weighted, scored factor of the kind governance indicators have become. Moody’s routes environmental considerations through a separate Issuer Profile Score rather than the core scorecard, on the stated basis that they matter ‘primarily’ through downstream economic and fiscal effects (Moody’s Ratings, 2026). Fitch goes furthest, building a Climate Vulnerability Signal projected to 2050 in five-year steps, but is explicit that it functions as ‘a screener to identify credits with higher exposure to climate related risks’ which then ‘receives further analytical scrutiny’, with a stated combined threshold of fifty by 2035 that most sovereigns have not yet reached (Fitch Ratings, 2025, p. 33). Climate risk is, in Haas’s terms, epistemically admissible: all three agencies plainly believe it matters, enough to build dedicated language and, in Fitch’s case, a dedicated model, around it. It has not achieved institutional admissibility in the sense governance indicators have. Why not is a harder question than the comparison itself, and the evidence assembled here supports an inference rather than a settled answer. One reason may lie in validation. Climate risk lacks the multi-decade track record against realised default and market outcomes, at horizons short enough to be tested by the same static pool and accuracy ratio exercises described earlier, that governance indicators had the time to accumulate. This is not to say that climate as a concept should be a granular as governance issues are in terms of evaluation in order to be considered properly, but governance indicators are a benchmark to consider against. Barta’s complaint that rating methodology treats climate risk with a short time horizon is, on this evidence, correctly diagnosed, and the validation problem offers one plausible, partial explanation alongside whatever ideological content is also at work, since the entire apparatus that gives a methodology its claim to be more than opinion is built to operate at horizons where outcomes can be observed and counted, and a risk that manifests fully only by 2050 sits awkwardly inside infrastructure built to test itself against five and ten year records. That mismatch does not rule out other explanations. It offers one that does not require attributing the outcome to anyone’s intent.
The Trouble with ‘the Market’
Commentary on sovereign finance routinely invokes a singular actor - the market - said to require fiscal discipline, reward credibility, and dislike particular policies, as though a coherent subject stood behind the behaviour of commercial banks, sovereign wealth funds, index-tracking asset managers, hedge funds, central bank reserve managers, pension trustees, and the three credit rating agencies themselves. A study by economists at the Bank for International Settlements gives some sense of how little this abstraction survives contact with the data. Amstad, Remolona, and Shek extracted the principal common factor driving sovereign credit default swap spreads across eighteen emerging and ten advanced economies between 2004 and 2014, then tried to explain which countries loaded most heavily on that factor using the variables one would expect such as debt ratios, fiscal balances, growth, and sovereign ratings themselves. Almost none of them mattered; the only consistently significant variable was whether a country counted as an emerging market at all. ‘There seems to be no Fragile Five’, the authors concluded, ‘there are only emerging markets... a designation that lacks the kind of granularity that we would have expected for a fundamental on which investors’ risk assessments are based’ (Amstad, Remolona and Shek, 2016, p. 18), a pattern they attribute to benchmark-tracking behaviour among asset managers who buy and sell by index membership rather than country analysis. A related BIS study comparing agency ratings before and after 2008 found that once standard controls were included, any residual penalty on emerging sovereigns became statistically indistinguishable from zero, while agencies outside the traditional centres of the industry rate emerging sovereigns more generously than the Big Three, yet market spreads and independent analyst surveys track the majors far more closely. Its authors drew the plain conclusion: ‘if a bias does exist, it is one shared with financial markets and asset managers more generally’ (Amstad and Packer, 2015, p. 90). Whatever market-liberal content Machado is right to find embedded in sovereign rating criteria, it is not obviously proprietary to the three agencies.
This distribution of authorship across a wide field is what a body of organisational sociology, developed for entirely different purposes, would predict. Paul DiMaggio and Walter Powell argued that organisations within a shared field converge on similar structures through mechanisms that need not involve any actor’s conscious ideological commitment: coercive pressure from regulators, mimetic copying under uncertainty, and normative convergence produced by shared professional training, so that personnel become, in their words, ‘almost interchangeable’ across organisations in a field (DiMaggio and Powell, 1983, p. 152). They were explicit that the theory concerns ‘not the psychological states of actors but the structural determinants of the range of choices that actors perceive as rational or prudent’ (DiMaggio and Powell, 1983, p. 149, fn. 5). A mechanism of this kind could plausibly operate among rating analysts, Fund economists, sovereign debt lawyers and reserve managers; professions whose personnel move through overlapping training and sometimes rotate between the same institutions over a career, a channel sufficient for convergent assumptions without anyone defending those assumptions as doctrine. DiMaggio and Powell wrote about organisational fields in general rather than this profession, so the claim here extends their argument rather than restates a finding specific to sovereign rating. Theodore Porter’s account of quantification explains why the resulting judgement, once expressed as a letter grade, travels so easily and commands so much deference: numbers function as ‘a technology of distance’, coordinating action among parties who do not know or trust one another personally, because a quantified rule ‘excludes judgment’ at the point of transmission even though judgement was fully present at the point of construction (Porter, 1992, pp. 639 to 640). Barbara Levitt and James March’s work on organisational learning adds a dimension of time. Organisations, they write, learn by encoding the lessons of experience into routines, which then become ‘independent of the individual actors who execute them and are capable of surviving considerable turnover in individual actors’ (Levitt and March, 1988, p. 320), and can persist inside a rating framework simply because no one currently employed by the agency was present at their adoption.
None of this requires abandoning the observation that the resulting framework’s content tends to be market-liberal in a fairly specific sense. Timothy Sinclair’s description of the agencies as operators of what he calls embedded knowledge networks captures something true and important: a rating becomes, in his phrase, ‘just as commonplace, and just as unquestioned, an entity in these markets as a chair or table is in the domestic kitchen’ (Sinclair, 2001, p. 443), authority resting on having become infrastructure rather than opinion. Marion Fourcade and Kieran Healy, writing about consumer credit scoring, coin the term classification situation for the position a proprietary scoring technology assigns to the entity it scores, arguing that such classifications actively ‘recreate’ the differences they appear only to describe, becoming ‘the engine of modern class situations’ (Fourcade and Healy, 2013, pp. 561, 569), an analogy plausibly extending to the state. The question is whether the market-liberal content Machado identifies is best explained as something CRAs actively author, or as something that persists inside a much wider institutional field, disseminated by training pipelines, embedded in index rules, reproduced by routine, and codified and transmitted by the agencies with unusual visibility. The evidence assembled here favours the second reading more strongly than the first, without eliminating it, since a field can of course be populated by people who also hold the beliefs their routines encode. Kerwer’s framing of the agencies as private, non-majoritarian regulators holds both halves at once: bodies that ‘lack a formal element of coercion’ and yet are ‘often criticized for wielding illegitimate power’, their authority owed less to any mandate to hold market-liberal convictions than to having become, through regulatory embedding in instruments such as Basel capital rules, ‘access rules for financial markets’ a government cannot simply opt out of (Kerwer, 2005, pp. 453, 462).
Do the Evaluated Have a Voice?
It would be a mistake to let the preceding argument slide into a picture of governments as purely passive objects of judgement, though the case for their agency needs to be stated carefully, since agency within a settlement and authorship of its categories are not the same thing. The IMF’s 2010 review notes that sovereign ratings depend in part on ‘additional information supplied to them by the country authorities’ and on a continuing pattern of engagement, briefings, data provision, investor seminars, between issuing governments and analyst teams (International Monetary Fund, 2010, p. 99). Hauner, Jonas, and Kumar found that the Central and Eastern European states that joined the European Union in the 2000s enjoyed a rating and spread advantage over comparable emerging economies that fundamentals alone could not explain, roughly 1.7 notches and seventy to a hundred basis points on foreign and local currency yields, appearing in a structural break around 2001 to 2003, well before formal accession, what they call an ‘exogenous credibility infusion’ (Hauner, Jonas and Kumar, 2007, pp. 4, 14). Joining the Union did not change what evaluators mean by creditworthiness. It supplied a characteristic, binding oneself to a supranational rule structure, that evaluators had already learned to treat as informative, and the states that acquired it were rewarded within the existing categories rather than credited with having redefined them: a genuine form of agency, strategic positioning within an inherited settlement, but short of authorship over its underlying categories.
Direct contestation of those categories is rarer, and on the evidence available considerably less successful. In justifying its own new credit rating institution - the Africa Credit Rating Agency - the African Peer Review Mechanism noted that ‘the majority of governments did not participate’ in the G20’s Debt Service Suspension Initiative and Common Framework ‘for fear that it would lead to credit rating downgrades’, and that international agencies had in fact downgraded Cameroon, Cote d’Ivoire, Ethiopia, and Senegal specifically because they participated (African Peer Review Mechanism, 2025, p. 6, quoted in Machado, 2026, p. 10). Ecuador’s 2023 debt-for-nature swap was followed by a default classification from Moody’s, and Colombia is reported to have abandoned a comparable initiative over concern for its own rating (Machado, 2026, p. 10). In each case governments tried to alter the terms on which their conduct would be judged, and the incumbent agencies held the existing definition of creditor rights firm rather than treat the attempt as evidence the definition needed revising. AfCRA is the more interesting case precisely because of what it represents rather than what it has yet achieved: an attempt to move beyond performing creditworthiness under categories inherited from elsewhere, toward an institution that might, if it survives and is taken seriously by the audiences described throughout this essay, eventually participate in setting those categories itself. Reinhart, Rogoff, and Savastano’s account of debt intolerance helps explain why that move is difficult rather than merely slow: serial defaulters, a category including many of today’s developing economies, face elevated default risk at debt-to-output ratios as low as fifteen to thirty five per cent, thresholds unremarkable for an advanced economy, a pattern traced back through centuries in which Spain defaulted thirteen times between 1557 and 1882 and Venezuela nine times since 1824 (Reinhart, Rogoff and Savastano, 2003, pp. 3 to 5). The categories a government seeking to contest its treatment must reckon with are a longer accumulated record that predates any of them, not simply the current judgement of three firms. Governments possess real agency within the evaluative settlement, in the sense of performing well or badly against categories that are given to them. The evidence that they can alter what those categories are is considerably weaker, and where an attempt exists at all, as with AfCRA, it is a project still being tested rather than a result already achieved.
Could a Single Agency See Creditworthiness Differently?
Let us return to the question posed at the outset. Suppose an agency revised its theory of creditworthiness in some material respect, weighting public investment capacity more heavily relative to near-term fiscal balance, say, or treating a well-structured debt-for-nature swap as evidence of institutional capacity rather than a departure from creditor rights. Two different scenarios follow, and the evidence bears on each differently.
In the first, the revised theory leaves the resulting ratings largely where they were; the change is intellectual before it is numerical. There is reason to think such a change would attract very little market attention. Kiff, Nowak, and Schumacher found that sovereign credit default swap spreads respond significantly to changes in rating outlook and watch status, considerably more than to confirmed rating actions markets have already priced in, and most sharply when a rating crosses the investment-grade threshold that triggers mandate-based buying and selling (Kiff, Nowak, and Schumacher, 2012, pp. 11 to 12). What markets appear to price is the letter grade and its trajectory, not the reasoning behind it, much as Sinclair’s rating as unremarked financial furniture would suggest: an agency could publish an entirely new theoretical justification for an unchanged set of numbers, and the change might register only among specialists who read methodology documents closely.
The second scenario tests the limits of an agency’s authority: a change substantial enough that its ratings begin to diverge systematically from its two competitors. Divergence of this kind is not, on the evidence, a rare or destabilising event in itself. Vu et al, studying sixty four sovereigns from 1997 to 2011, found that ratings from the three major agencies disagreed on between fifty three and sixty seven per cent of daily observations, systematically rather than randomly: S&P was consistently the more conservative of the three, and the strongest predictor of disagreement was a measure of political risk rather than any economic fundamental - the World Bank’s rule of law indicator - whose explanatory power was considerably stronger outside Europe than within it (Vu, Alsakka, and Gwilym, 2017, pp. 4, 19 to 20). Disagreement is a routine feature of a practice that leaves real room for interpretation, and its ordinary occurrence has not visibly damaged the standing of any of the three agencies. What has not happened, and what this essay’s evidence cannot speak to, is a case in which a major agency abandoned the specific orthodoxy Machado identifies, treating fiscal expansion for climate investment as creditworthy or debt restructuring as institutional strength rather than weakness, at a scale large enough to produce sustained divergence from its peers. The nearest test of the market’s tolerance for redefinition on record concerns a change of epistemic self-description rather than a substantive change of view: the exchange recorded in ESMA’s feedback statement, in which an agency insisted its ratings were ordinal opinions rather than predictions of default frequency, and the regulator held its testing requirement while explicitly conceding this would not alter what the ordinal product itself claims to be (European Securities and Markets Authority, 2016). Even a regulator with statutory authority over methodology validation stopped short of compelling a change in what a rating is understood to mean. Whether a market of investors and reserve managers operating under fixed rating-linked mandates would extend the same latitude to an agency that changed what creditworthiness itself is taken to require, rather than merely how it is tested, is a question this essay cannot answer because no agency has yet run the experiment at scale. The infrastructure through which ratings operate - the Basel weightings, the index eligibility rules, the reserve manager mandates that name a single A rating as an informal floor - was built around the existing understanding and would not automatically follow an agency that moved away from it. An innovating agency would be asking institutions that currently defer to its judgement to defer to a different one instead, and that deference is itself the product of decades of routine, training, and regulatory embedding rather than a subscription that renews on its own.
Oligopoly as Shelter or Discipline
Two opposing intuitions are available for us to understand more about what the entrenched position of three agencies does to the prospect of change, and the evidence supports neither cleanly. The comforting intuition holds that competition between authoritative comparators disciplines any single agency against drifting too far from consensus, since divergence risks looking like error rather than insight. The less comfortable intuition holds that a settled oligopoly, sheltered from serious new entry, gives incumbents enough security to experiment without commercial consequence, precisely because clients and regulators have nowhere else credible to go.
The empirical record on competition points, if anything, in a direction that unsettles both intuitions. Bolton, Freixas, and Shapiro’s game-theoretic model of the ratings industry produces a genuinely counterintuitive result: under issuer-paid ratings with heterogeneous investors, some of whom trust ratings without independent verification (theoretically speaking), ‘a truth telling monopoly strictly dominates a truth telling duopoly’, because competition between two agencies gives issuers more scope to shop for a favourable opinion, a scope a monopoly cannot offer since there is nowhere else to shop (Bolton, Freixas and Shapiro, 2012, p. 101). Becker and Milbourn’s empirical study of Fitch’s expansion into corporate bond ratings between 1995 and 2006 found a version of the same result outside the model: as Fitch’s market share grew, incumbent ratings rose, the correlation between ratings and market-implied yields weakened by roughly a third, and the ratio of default rates between speculative-grade and investment-grade issuers collapsed from 7.7 to 2.2 under high competition compared with low (Becker and Milbourn, 2011, pp. 508, 510), which the authors attribute to an erosion of reputational incentives rather than to ratings shopping; here, competition makes accuracy less valuable to defend (Becker and Milbourn, 2011, pp. 497 to 498, 505). Neither study is direct evidence about sovereign methodology, but together they weigh against any easy assumption that adding competitors would push the industry toward either greater rigour or greater openness.
Lawrence White’s institutional history helps explain why the industry has not, in any case, faced serious new entry. The regulatory decisions that gave the three agencies their position, a 1936 bank regulation giving ‘the creditworthiness judgments of these third party raters... the force of law’, and the SEC’s 1975 creation of a recognised-rater designation that effectively grandfathered the existing firms, were deliberate acts of public policy rather than outcomes of market competition; one SEC commissioner later described the resulting structure as an oligopoly holding ‘more than 90 percent of the market share’ (White, 2010, pp. 213, 217). When the SEC designated seven additional recognised agencies between 2003 and 2008 to introduce competition, White reports the reform ‘had little substantial effect’ against the Big Three’s structural advantages (White, 2010, p. 221). The international soft-law layer above these national arrangements has focused less on methodological content than on conduct: supervisory colleges for the three major agencies, chaired variously by the SEC and by ESMA, review disclosure and governance rather than substantive judgement (International Organization of Securities Commissions, 2015). Kerwer adds that legal liability for a rating opinion is unusually weak, protected as speech in the United States and subject to a recklessness rather than a negligence standard, removing a mechanism that might otherwise force caution or defensible innovation (Kerwer, 2005, pp. 469, 472). Kerwer wrote this in 2005, years before the US and the EU both revolutionised their liability regimes for credit rating agencies; in reality, not much changed. This evidence supports a specific and cautious conclusion, not a sweeping one. The entrenched agencies have not, on this record, undertaken paradigmatic experimentation, and the clearest empirical test of increased competition found it eroding accuracy rather than disciplining toward it. What the record documents directly is a structure built by regulators rather than by the market its defenders sometimes invoke, closer to inertia than to either shelter for boldness or discipline toward correctness. What it does not document, and what can only be offered as an inference from that structure, is the commercial incentive presumed to follow from it: that this arrangement makes incremental adjustment within a shared professional consensus more attractive than either bold departure or externally enforced correction, because the consensus itself, rather than any single firm’s judgement, is what gives a rating its purchase on the institutions that rely on it. The structure is established. The incentive it is presumed to create remains a plausible reading of that structure rather than a separately demonstrated fact.
Drivers, Passengers, and the Return to Machado
It is now possible to give a more precise answer to this essay’s opening question than either ‘the agencies decide’ or ‘the market decides’ would allow. Credit rating agencies occupy a more complicated position than either formula suggests. The preceding sections have shown them exercising considerable authorship over the operational form of creditworthiness: the scorecards, matrices, overlays, and validation exercises through which observable facts become a judgement claiming to be more than opinion, while depending throughout on data, training, competitors, and infrastructure that they do not control: the World Bank’s governance indicators, professional pipelines shared with the Fund and investment banks, two competitors trained in the same tradition, and a regulatory and contractual apparatus that no single government or firm controls.
It helps, at this point, to separate three things that might be meant by authorship of creditworthiness, without turning the distinction into another named framework. One is authorship of the concept: which characteristics, fiscal balance, institutional quality, monetary credibility, ought to count as evidence of a government’s capacity to repay. A second is authorship of the operationalisation: how those characteristics become the variables, weights, caps, and overlays a rating committee can actually apply. A third is authorship through recognition: whether investors, regulators, index providers, and reserve managers agree to treat the resulting judgement as one worth acting upon. Credit rating agencies hold a great deal of the second kind - the architecture of scoring rules and adjustment thresholds is substantially theirs to set, as the earlier sections of this essay showed. They participate in the first, contributing to a conception of creditworthiness they did not originate alone and share with the Fund, the World Bank, and a long professional tradition. They depend, almost entirely, on the third: a rating means nothing until Basel weightings, index rules, and portfolio mandates agree to treat it as meaningful. Machado’s argument engages mainly the first kind, the content of the concept itself, and this essay does not dispute that the content is real. The second and third kinds help explain why that content persists so durably, and why moving it requires more than persuading a rating committee to think differently.
What Machado identifies as the market-liberal content of sovereign rating practice is therefore real, and the evidence gathered here does not dispute it. What that evidence suggests is that the mechanism by which this content persists looks less like a continuously renewed ideological commitment inside three firms and more like an institutional settlement, distributed across CRAs, investors, the Fund and World Bank, professional analyst communities, index providers, regulators, and governments themselves, in which each participant reproduces the settlement by doing its ordinary work rather than by defending a doctrine. Kerwer’s description of the agencies as non-majoritarian regulators, authoritative but without a formal mandate or a single principal to answer to, dependent for their standing on continued recognition by the rest of the system, is a reasonably accurate description of their position within that settlement, provided one adds that the settlement itself, and not any one node inside it, is where the market-liberal content actually resides.
This distinction matters for what follows from Machado’s argument more than for whether the argument is correct. If a single firm’s ideology were the mechanism, the remedy would be relatively direct: better regulation of that firm, more competition, perhaps public alternatives, until the offending judgement is corrected or displaced. Institutional reproduction is harder to correct, precisely because it is distributed across datasets that took decades to earn their standing, professional routines that outlive the people who first adopted them, and a regulatory infrastructure built around the existing definition of creditworthiness. New evidence on climate risk would need to travel the same route the World Bank’s governance indicators travelled (as a benchmark example), tested repeatedly against realised outcomes until it earns a place in a weighted scorecard rather than remaining a qualitative aside. Professional communities would need to treat a different weighting of public investment against fiscal restraint as defensible rather than as a departure from due diligence, investors and reserve managers would need to reward a rating that reflects it, and institutions such as the African Union’s new credit rating agency would need to survive long enough, and be taken seriously enough, to become a genuine second comparator rather than a marginal alternative the settlement can safely ignore.
Recognising this distributed mechanism does not weaken Machado’s diagnosis. Persuading S&P Global, Moody’s, or Fitch to think differently, even if such persuasion succeeded completely, would not by itself change what creditworthiness means in practice, so long as the wider settlement in which their judgements circulate continued to treat the old definition as the reasonable one. The harder question this essay has tried to locate, rather than resolve, is how an evaluative system that has spent the better part of a century learning to recognise fiscal restraint, creditor protection, and private capital mobility as the signatures of a government worth lending to might learn to recognise something else, sustained public investment against a warming climate among the more urgent candidates, as a signature of the same thing. No single actor in the settlement described here currently has the standing to make that recognition happen alone. This also raises the potential that, if the credit rating agencies were to diverge away from what appears to institutional logic, then would they still be useful to that same institutional network? Would they still be able to command some of the world’s highest profit margins? The answer to those questions, which are not attempted here, likely provide the reality of the situation more than any other answers could do.
References
African Peer Review Mechanism (2025) An Africa Credit Rating Agency (AfCRA). Addis Ababa: African Union. Available at: au.int/sites/default/files/documents/44466-doc-Brief_-_Africa_Credit_Rating_Agency_AfCRA.pdf (quoted in Machado, 2026).
Amstad, M. and Packer, F. (2015) ‘Sovereign Ratings of Advanced and Emerging Economies after the Crisis’, BIS Quarterly Review, December, pp. 77-91.
Amstad, M., Remolona, E. and Shek, J. (2016) How Do Global Investors Differentiate Between Sovereign Risks? The New Normal versus the Old, BIS Working Papers No. 541. Basel: Bank for International Settlements.
Arbatli, E. and Escolano, J. (2012) Fiscal Transparency, Fiscal Performance and Credit Ratings, IMF Working Paper WP/12/156. Washington, DC: International Monetary Fund.
Barta, Z. (2026) ‘Tragedy of the horizon on steroids: the green transition and credit ratings’, Review of International Political Economy. Available at: doi.org/10.1080/09692290.2026.2654619 (quoted in Machado, 2026).
Becker, B. and Milbourn, T. (2011) ‘How Did Increased Competition Affect Credit Ratings?’, Journal of Financial Economics, 101(3), pp. 493-514.
Bolton, P., Freixas, X. and Shapiro, J. (2012) ‘The Credit Ratings Game’, The Journal of Finance, 67(1), pp. 85-111.
DiMaggio, P. J. and Powell, W. W. (1983) ‘The Iron Cage Revisited: Institutional Isomorphism and Collective Rationality in Organizational Fields’, American Sociological Review, 48(2), pp. 147-160.
Espeland, W. N. and Sauder, M. (2007) ‘Rankings and Reactivity: How Public Measures Recreate Social Worlds’, American Journal of Sociology, 113(1), pp. 1-40.
Espeland, W. N. and Stevens, M. L. (1998) ‘Commensuration as a Social Process’, Annual Review of Sociology, 24, pp. 313-343.
European Securities and Markets Authority (2016) Final Report: Guidelines on the Validation and Review of Credit Rating Agencies’ Methodologies, ESMA/2016/1575. Paris: ESMA, 15 November 2016.
European Securities and Markets Authority (2017) Guidelines on the Validation and Review of Credit Rating Agencies’ Methodologies, ESMA/2016/1575. Paris: ESMA, 23 March 2017.
Fitch Ratings (2018) Fitch Solutions Election View: Bolsonaro Most Likely to Win in Close Election. New York: Fitch Ratings (quoted in Machado, 2026).
Fitch Ratings (2025) Global Sovereign Rating Criteria. New York: Fitch Ratings.
Fourcade, M. and Healy, K. (2013) ‘Classification Situations: Life-Chances in the Neoliberal Era’, Accounting, Organizations and Society, 38(8), pp. 559-572.
Haas, P. M. (1992) ‘Introduction: Epistemic Communities and International Policy Coordination’, International Organization, 46(1), pp. 1-35.
Hauner, D., Jonas, J. and Kumar, M. S. (2007) Policy Credibility and Sovereign Credit: The Case of the New EU Member States, IMF Working Paper WP/07/1. Washington, DC: International Monetary Fund.
International Monetary Fund (2010) ‘Chapter 3: The Uses and Abuses of Sovereign Credit Ratings’, in Global Financial Stability Report: Sovereigns, Funding, and Systemic Liquidity, October 2010. Washington, DC: International Monetary Fund, pp. 85-122.
International Organization of Securities Commissions (2015) Code of Conduct Fundamentals for Credit Rating Agencies, Final Report FR05/2015. Madrid: IOSCO.
Keita, K., Leon, G. and Lima, F. (2019) Do Financial Markets Value Quality of Fiscal Governance?, IMF Working Paper WP/19/218. Washington, DC: International Monetary Fund.
Kerwer, D. (2005) ‘Holding Global Regulators Accountable: The Case of Credit Rating Agencies’, Governance, 18(3), pp. 453-475.
Kiff, J., Nowak, S. and Schumacher, L. (2012) Are Rating Agencies Powerful? An Investigation into the Impact and Accuracy of Sovereign Ratings, IMF Working Paper WP/12/23. Washington, DC: International Monetary Fund.
Levitt, B. and March, J. G. (1988) ‘Organizational Learning’, Annual Review of Sociology, 14, pp. 319-340.
MacKenzie, D. (2003) ‘An Equation and its Worlds: Bricolage, Exemplars, Disunity and Performativity in Financial Economics’, Social Studies of Science, 33(6), pp. 831-868.
Machado, P. L. (2026) Climate Obstruction through Ratings: A Research Agenda, Cadernos do OIMC, no. 23/2026. Rio de Janeiro: Universidade do Estado do Rio de Janeiro, Interdisciplinary Observatory on Climate Change.
Moody’s Ratings (2026) Sovereign and Supranational Rating Methodology. New York: Moody’s Ratings.
Morahan, A. and Mulder, C. (2013) Survey of Reserve Managers: Lessons from the Crisis, IMF Working Paper WP/13/99. Washington, DC: International Monetary Fund.
Porter, T. M. (1992) ‘Quantification and the Accounting Ideal in Science’, Social Studies of Science, 22(4), pp. 633-651.
Reinhart, C. M., Rogoff, K. S. and Savastano, M. A. (2003) Debt Intolerance, NBER Working Paper No. 9908. Cambridge, MA: National Bureau of Economic Research.
S&P Global Ratings (2025) Default, Transition, and Recovery: 2024 Annual Global Sovereign Default and Rating Transition Study. New York: S&P Global Ratings.
S&P Global Ratings (2026) Criteria | Governments | Sovereigns: Sovereign Rating Methodology. New York: S&P Global Ratings.
Sinclair, T. J. (2001) ‘The Infrastructure of Global Governance: Quasi-Regulatory Mechanisms and the New Global Finance’, Global Governance, 7(4), pp. 441-451.
Vu, H., Alsakka, R. and Gwilym, O. (2017) ‘What Drives Differences of Opinion in Sovereign Ratings? The Roles of Information Disclosure and Political Risk’, International Journal of Finance & Economics, 22(3), pp. 216-233.
White, L. J. (2010) ‘Markets: The Credit Rating Agencies’, Journal of Economic Perspectives, 24(2), pp. 211-226.