Two findings that require no modeling
Neither of these is our statistical judgment. Both are facts about what the federal files contain.
Hospitals with a published patient-reported outcome
The CMS patient-reported outcomes facility file covers July 2023 to June 2024 and contains 4,626 rows. Every row reports no available result. No THA/TKA patient-reported outcome is published for any hospital in the United States — so the measure closest to what patients actually care about is not sparse, it is absent nationally.
Hospitals CMS cannot distinguish from average
On the THA/TKA complication measure, CMS classifies 1,692 of 1,706 hospitals as no different from the national rate. Eight are worse, six are better. Fourteen hospitals — 0.8% — are statistically distinguishable on the measure that dominates any composite built from these files.
This matters because it sets a ceiling. A composite that assigns a distinct rank to each of 1,596 hospitals derives that ordering almost entirely from differences the source data themselves declare indistinguishable from average. The honest output of such data is not a ranking — it is a statement that most institutions cannot be told apart, which is the opposite of what a consumer-facing index communicates.
What we built, and what it showed
Every US hospital in the CMS Provider Data Catalog with both a reported THA/TKA complication score and a 30-day readmission score, as of 14 September 2026. Domains converted to favorable within-cohort percentiles and combined on the prespecified weights, renormalized after the fourth domain proved uncomputable.
hospitals listed
fully scorable (29.5%)
composite reproduced by the complication percentile alone
association with an outside readmission construct
The domains do not measure one thing
A composite is defensible when its components express a shared underlying construct. These do not. Complications and readmissions correlate at 0.379; readmissions and patient experience at 0.103; complications and patient experience at 0.055. Combining near-independent measures into one number discards more information than it summarizes — and ranking on the complication percentile alone reproduces the whole composite at 0.910. The composite is, to a close approximation, the complication measure.
Rank stability depends on how much you were willing to disagree
Composite rankings are often defended by showing they survive alternative weights. Ours do — under narrow assumptions. Widen the disagreement and they do not.
| Weight uncertainty (20,000 draws) | Median rank correlation | Of the original top 10, retained |
|---|---|---|
| Very mild disagreement | 0.995 | 9.1 of 10 |
| Moderate disagreement | 0.963 | 7.5 of 10 |
| Genuine disagreement | 0.934 | 6.1 of 10 (as few as 2) |
Under genuine disagreement about how the domains should be weighted, an average of four of the top ten institutions change identity. The stability result that such indices usually report is a property of a narrow prior, not of the data.
There is nothing external to validate against
The only criterion available in public data is another CMS readmission construct. Measured over the same period it correlates at 0.998 with a domain already inside the composite — an association that looks like validation and is arithmetic. Measured over a partly different period it yields 0.329, falling to 0.246 once the composite's own readmission domain is removed. Public data permit no test at all against complications, revision, infection, function, or cost in a subsequent period.
Who can be scored is itself a filter
Only 29.5% of listed hospitals could be scored. Those that could not are disproportionately critical access, psychiatric, and children's hospitals; fewer than half carry a CMS overall star rating. Among hospitals reporting a THA/TKA denominator, the median was 76 in the scored group and 35.5 in the unscored group. Scorability is a volume filter, and any composite of this kind describes higher-volume institutions by construction.
Volume dominates the top of the distribution too: the twenty highest-scoring hospitals performed a median of 456 procedures against a cohort median of 76. Physician-owned hospitals were 1.8% of the cohort and 5.6% of the top decile — a threefold enrichment consistent with the case selection characteristic of focused elective facilities, which risk standardization on administrative codes does not fully adjust away. That pattern describes the ambulatory, elective setting this practice operates in, and we report it because it cuts against a flattering reading of our own results as much as anyone's.
Why better analysis cannot repair this
The reasonable objection is that CMS should simply measure better. The constraints follow from purpose rather than from effort.
Risk-standardized complication and readmission measures exist to administer payment programs — to distribute penalties across institutions with defensible statistical footing. Every design choice follows from that purpose, and each one is correct for a penalty program and disqualifying for a decision aid:
- A payment program must adjudicate within an administratively bounded period, which forecloses a one-year functional endpoint.
- It must attribute to the entities it contracts with, which forecloses episode or surgeon attribution — results produced by an episode get assigned to a building.
- It must report on a multi-year lag to achieve stable estimates, so the data describe care delivered years earlier.
- It must avoid accusing an institution wrongly — which is precisely why 99% of hospitals are classified as no different from average.
There is also the question of provenance. Risk standardization operates on administrative codes generated by the institutions being evaluated, so documentation and coding intensity influence measured performance independently of the care delivered. That can be adequate for administering a program. It carries an unavoidable limit on how much weight it should bear in a consequential individual decision.
And the blind spot widens. CMS removed hip and knee replacement from the inpatient-only list and later added them to the ambulatory surgery center covered-procedures list; utilization followed. The measurement apparatus remains an inpatient hospital apparatus. As the procedure moves toward the setting the payer itself enabled, the public system observes less of the care delivered, and improvement presents as missing data.
What a sufficient measurement layer would contain
None of this is exotic. All of it is collectable. None of it is publicly reported.
- Patient-reported function and pain before operation and at defined intervals afterward, scored against the minimal clinically important difference for the instrument used — measuring change, not level.
- Revision and reoperation at defined intervals.
- Periprosthetic joint infection reported separately rather than bundled into a composite complication measure.
- Return to work and to valued activity.
- Total episode cost, including the portion borne by the patient.
- Whether the patient, knowing the result, would choose the operation again.
Attribution would span the episode — indication, patient selection, optimization, technique, implant, rehabilitation, follow-up — rather than stopping at the facility, because the facility is one input among those and the others are where the result is actually produced.
Who should build it
Our conclusion is that arthroplasty surgeons should build and govern this layer themselves. That is an argument, not an empirical result, and we present it as such.
The precedent is established. Cardiac surgery answered the same question three decades ago by constructing a specialty-owned clinical database, developing and maintaining its own risk models, and setting the terms on which its results are reported — with the consequence that payers and regulators now defer to models the specialty built. Arthroplasty has demonstrated the same collective capacity through the American Joint Replacement Registry. What the specialty lacks is a layer that spans the episode and whose primary outcome data do not pass through the institutions being evaluated.
The obvious objection
A measurement system governed by surgeons invites the suspicion that surgeons will grade themselves favorably. For most clinical registries that suspicion is structurally justified, because the practice being measured also submits the data.
The answer to it
Collect patient-reported outcomes directly from the patient, under the patient's own authorization, by an entity that does not deliver the care. The specialty can govern the instrument, endpoints, risk models, and analysis plan without ever being the source of its own result.
Two commitments that have to come with it
No ordinal consumer ranking. A specialty that builds a league table reproduces the artifact it objects to, with more direct conflicts attached. The defensible output is a decomposed presentation of each domain, reported with its uncertainty and its missingness visible, and with the boundaries of the claim stated alongside it.
Open participation and a public method, or the layer becomes the private asset of whoever builds it first.
The risks of the alternative, stated honestly
Voluntary registries select for institutions confident in their results. Self-governance invites the suspicion it must answer. Collection is expensive in money and clinician time, and incomplete participation limits generalizability in the same way missingness limits the analysis above. A patient-facing output derived from a surgeon-governed layer can mislead as easily as one derived from CMS. These are reasons to design the layer carefully, publish its method, and submit it to audit. They are not reasons to keep deferring to a measurement system built for a different purpose by parties with their own.
Status and sources
All source files are public CMS Provider Data Catalog downloads of 14 September 2026.
| Source file | Measure | Period |
|---|---|---|
| Complications and Deaths – Hospital | COMP_HIP_KNEE, score and denominator | 04/2023–03/2025 |
| Unplanned Hospital Visits – Hospital | READM_30_HIP_KNEE | 07/2023–06/2025 |
| Patient survey (HCAHPS) – Hospital | H_STAR_RATING summary star | current file |
| Hospital Readmissions Reduction Program | READM-30-HIP-KNEE-HRRP excess readmission ratio | 07/2021–06/2024 |
| Patient-Reported Outcomes – Hospital | THA/TKA PRO-PM — no result published | 07/2023–06/2024 |
| Hospital General Information | type, ownership, overall star rating | current file |
Percentiles are computed within the analytic cohort, so they describe position among reporting hospitals rather than among all US hospitals, and are not comparable to percentiles from a differently defined cohort. Measure-level standard errors were not incorporated, so nothing here supports a claim that any two individual hospitals differ — the relevant statement is the CMS classification above, and it is that almost none do.
The manuscript reporting this analysis is in preparation. A full citation, the analysis script, and the derived dataset will be published here on acceptance. A machine-readable summary of the source manifest and findings is available at /avi/avi-spec.json.