Arthroplasty Quality Measurement

Who should measure arthroplasty?

A patient deciding on hip or knee replacement asks a few concrete questions. Will the joint work. How long will it last. What is the chance something goes wrong. What will it cost me. Would someone like me choose this again. When that patient looks for evidence, they are directed to publicly reported hospital measures — and none of those measures answers any of those five questions.

So we tested it directly: we built the most complete arthroplasty composite that public CMS data permits, and asked what it could establish. It proved internally coherent and substantively empty. The reasons are not defects of analysis. They follow from who owns and defines the measurement layer.

Two findings that require no modeling

Neither of these is our statistical judgment. Both are facts about what the federal files contain.

0

Hospitals with a published patient-reported outcome

The CMS patient-reported outcomes facility file covers July 2023 to June 2024 and contains 4,626 rows. Every row reports no available result. No THA/TKA patient-reported outcome is published for any hospital in the United States — so the measure closest to what patients actually care about is not sparse, it is absent nationally.

99.2%

Hospitals CMS cannot distinguish from average

On the THA/TKA complication measure, CMS classifies 1,692 of 1,706 hospitals as no different from the national rate. Eight are worse, six are better. Fourteen hospitals — 0.8% — are statistically distinguishable on the measure that dominates any composite built from these files.

1,692 NO DIFFERENT FROM THE NATIONAL RATE
1,692 no different8 worse6 betterTHA/TKA complication measure, 04/2023–03/2025

This matters because it sets a ceiling. A composite that assigns a distinct rank to each of 1,596 hospitals derives that ordering almost entirely from differences the source data themselves declare indistinguishable from average. The honest output of such data is not a ranking — it is a statement that most institutions cannot be told apart, which is the opposite of what a consumer-facing index communicates.

What we built, and what it showed

Every US hospital in the CMS Provider Data Catalog with both a reported THA/TKA complication score and a 30-day readmission score, as of 14 September 2026. Domains converted to favorable within-cohort percentiles and combined on the prespecified weights, renormalized after the fourth domain proved uncomputable.

5,419

hospitals listed

1,596

fully scorable (29.5%)

0.910

composite reproduced by the complication percentile alone

0.329

association with an outside readmission construct

The domains do not measure one thing

A composite is defensible when its components express a shared underlying construct. These do not. Complications and readmissions correlate at 0.379; readmissions and patient experience at 0.103; complications and patient experience at 0.055. Combining near-independent measures into one number discards more information than it summarizes — and ranking on the complication percentile alone reproduces the whole composite at 0.910. The composite is, to a close approximation, the complication measure.

Rank stability depends on how much you were willing to disagree

Composite rankings are often defended by showing they survive alternative weights. Ours do — under narrow assumptions. Widen the disagreement and they do not.

Weight uncertainty (20,000 draws)Median rank correlationOf the original top 10, retained
Very mild disagreement0.9959.1 of 10
Moderate disagreement0.9637.5 of 10
Genuine disagreement0.9346.1 of 10 (as few as 2)

Under genuine disagreement about how the domains should be weighted, an average of four of the top ten institutions change identity. The stability result that such indices usually report is a property of a narrow prior, not of the data.

There is nothing external to validate against

The only criterion available in public data is another CMS readmission construct. Measured over the same period it correlates at 0.998 with a domain already inside the composite — an association that looks like validation and is arithmetic. Measured over a partly different period it yields 0.329, falling to 0.246 once the composite's own readmission domain is removed. Public data permit no test at all against complications, revision, infection, function, or cost in a subsequent period.

Who can be scored is itself a filter

Only 29.5% of listed hospitals could be scored. Those that could not are disproportionately critical access, psychiatric, and children's hospitals; fewer than half carry a CMS overall star rating. Among hospitals reporting a THA/TKA denominator, the median was 76 in the scored group and 35.5 in the unscored group. Scorability is a volume filter, and any composite of this kind describes higher-volume institutions by construction.

Volume dominates the top of the distribution too: the twenty highest-scoring hospitals performed a median of 456 procedures against a cohort median of 76. Physician-owned hospitals were 1.8% of the cohort and 5.6% of the top decile — a threefold enrichment consistent with the case selection characteristic of focused elective facilities, which risk standardization on administrative codes does not fully adjust away. That pattern describes the ambulatory, elective setting this practice operates in, and we report it because it cuts against a flattering reading of our own results as much as anyone's.

Why better analysis cannot repair this

The reasonable objection is that CMS should simply measure better. The constraints follow from purpose rather than from effort.

Risk-standardized complication and readmission measures exist to administer payment programs — to distribute penalties across institutions with defensible statistical footing. Every design choice follows from that purpose, and each one is correct for a penalty program and disqualifying for a decision aid:

  • A payment program must adjudicate within an administratively bounded period, which forecloses a one-year functional endpoint.
  • It must attribute to the entities it contracts with, which forecloses episode or surgeon attribution — results produced by an episode get assigned to a building.
  • It must report on a multi-year lag to achieve stable estimates, so the data describe care delivered years earlier.
  • It must avoid accusing an institution wrongly — which is precisely why 99% of hospitals are classified as no different from average.

There is also the question of provenance. Risk standardization operates on administrative codes generated by the institutions being evaluated, so documentation and coding intensity influence measured performance independently of the care delivered. That can be adequate for administering a program. It carries an unavoidable limit on how much weight it should bear in a consequential individual decision.

And the blind spot widens. CMS removed hip and knee replacement from the inpatient-only list and later added them to the ambulatory surgery center covered-procedures list; utilization followed. The measurement apparatus remains an inpatient hospital apparatus. As the procedure moves toward the setting the payer itself enabled, the public system observes less of the care delivered, and improvement presents as missing data.

No measure of cost exists either. The CMS Provider Data Catalog publishes no arthroplasty-specific spending or episode payment measure. The only hospital-level spending constructs are hospital-wide and not procedure-specific. Nothing built from public data can incorporate cost — which is why the “Value Index” designation has been retired from this work. Nothing in the analysis measures value.

What a sufficient measurement layer would contain

None of this is exotic. All of it is collectable. None of it is publicly reported.

  • Patient-reported function and pain before operation and at defined intervals afterward, scored against the minimal clinically important difference for the instrument used — measuring change, not level.
  • Revision and reoperation at defined intervals.
  • Periprosthetic joint infection reported separately rather than bundled into a composite complication measure.
  • Return to work and to valued activity.
  • Total episode cost, including the portion borne by the patient.
  • Whether the patient, knowing the result, would choose the operation again.

Attribution would span the episode — indication, patient selection, optimization, technique, implant, rehabilitation, follow-up — rather than stopping at the facility, because the facility is one input among those and the others are where the result is actually produced.

Who should build it

Our conclusion is that arthroplasty surgeons should build and govern this layer themselves. That is an argument, not an empirical result, and we present it as such.

The precedent is established. Cardiac surgery answered the same question three decades ago by constructing a specialty-owned clinical database, developing and maintaining its own risk models, and setting the terms on which its results are reported — with the consequence that payers and regulators now defer to models the specialty built. Arthroplasty has demonstrated the same collective capacity through the American Joint Replacement Registry. What the specialty lacks is a layer that spans the episode and whose primary outcome data do not pass through the institutions being evaluated.

The obvious objection

A measurement system governed by surgeons invites the suspicion that surgeons will grade themselves favorably. For most clinical registries that suspicion is structurally justified, because the practice being measured also submits the data.

The answer to it

Collect patient-reported outcomes directly from the patient, under the patient's own authorization, by an entity that does not deliver the care. The specialty can govern the instrument, endpoints, risk models, and analysis plan without ever being the source of its own result.

Governance separated from self-report. That distinction is what the current arrangement cannot offer, and it is what makes the proposition credible rather than self-serving. With a published scoring method, an analysis plan prespecified before data collection, and independent audit, it is a stronger claim to credibility than either the current public system or a conventional registry can make.

Two commitments that have to come with it

No ordinal consumer ranking. A specialty that builds a league table reproduces the artifact it objects to, with more direct conflicts attached. The defensible output is a decomposed presentation of each domain, reported with its uncertainty and its missingness visible, and with the boundaries of the claim stated alongside it.

Open participation and a public method, or the layer becomes the private asset of whoever builds it first.

The risks of the alternative, stated honestly

Voluntary registries select for institutions confident in their results. Self-governance invites the suspicion it must answer. Collection is expensive in money and clinician time, and incomplete participation limits generalizability in the same way missingness limits the analysis above. A patient-facing output derived from a surgeon-governed layer can mislead as easily as one derived from CMS. These are reasons to design the layer carefully, publish its method, and submit it to audit. They are not reasons to keep deferring to a measurement system built for a different purpose by parties with their own.

Status and sources

The four-domain “Arthroplasty Value Index” previously published at this address is retired. It was built on the January 2026 CJR quality reporting production file, which is not publicly distributed and which a reader therefore cannot reconstruct. Three of its claims are superseded by the analysis above: the patient-reported outcome domain is not computable from public data for any hospital; the cohort was not publicly reproducible; and its reported convergent validity of 0.718 reflected a same-period comparison against a construct already inside the index. That index should not be computed, cited, or attributed to any program, including Total Joint Specialists. No score is published here for any hospital.

All source files are public CMS Provider Data Catalog downloads of 14 September 2026.

Source fileMeasurePeriod
Complications and Deaths – HospitalCOMP_HIP_KNEE, score and denominator04/2023–03/2025
Unplanned Hospital Visits – HospitalREADM_30_HIP_KNEE07/2023–06/2025
Patient survey (HCAHPS) – HospitalH_STAR_RATING summary starcurrent file
Hospital Readmissions Reduction ProgramREADM-30-HIP-KNEE-HRRP excess readmission ratio07/2021–06/2024
Patient-Reported Outcomes – HospitalTHA/TKA PRO-PM — no result published07/2023–06/2024
Hospital General Informationtype, ownership, overall star ratingcurrent file

Percentiles are computed within the analytic cohort, so they describe position among reporting hospitals rather than among all US hospitals, and are not comparable to percentiles from a differently defined cohort. Measure-level standard errors were not incorporated, so nothing here supports a claim that any two individual hospitals differ — the relevant statement is the CMS classification above, and it is that almost none do.

The manuscript reporting this analysis is in preparation. A full citation, the analysis script, and the derived dataset will be published here on acceptance. A machine-readable summary of the source manifest and findings is available at /avi/avi-spec.json.