How Reliable Is TKA Rotation Measurement?
CT measurement of knee component rotation has 95% limits of agreement wider than the thresholds it is used to judge. What that means for outcome studies and for AI planning.
Key takeaways
Component rotation after total knee arthroplasty is measured on CT, compared against thresholds of roughly 6 to 9 degrees, and used to explain painful knees. The problem is that the measurement's own 95% limits of agreement between observers span about 25 degrees, which is wider than the decision band it is meant to serve. The authors of the most rigorous reliability study put it plainly: these measurements are not sufficiently reliable for routine clinical use. This is not a fringe position and nobody has solved it. It also sets the bar for any software, ours included, that proposes to automate the measurement: automation is only an improvement if it is demonstrably more reproducible, and that has to be shown rather than assumed.
The reliability data
Toms AP, Rifai T, Whitehouse C, McNamara I. Eur Radiol 2022. PMID 35142899 · PMC9122870 full text
A prospective reliability study embedded in the CAPAbility randomised trial. n=80, pre- and postoperative CT, Berger protocol.
| Measurement | Inter-observer ICC | Inter-observer 95% limits of agreement (span) | Intra-observer span |
|---|---|---|---|
| Preoperative femoral version | 0.65-0.72 | 8.1° | 7.6° |
| Preoperative tibial version | 0.59-0.64 | 26.0° | 20.5° |
| Postoperative femoral rotation | 0.79-0.81 | 9.0° | 7.9° |
| Postoperative tibial rotation | 0.78-0.80 | 24.9° | 23.6° |
Two things about how to read this table, because both are commonly got wrong.
These are spans, not plus-or-minus values. A 24.9-degree span is roughly ±12.5 degrees around the mean difference. Quoting it as "±25 degrees" doubles the real figure.
The ICCs look respectable. Postoperative values sit at 0.78 to 0.81, which most papers would call good agreement. That is exactly why the ICC alone is misleading here: it is a ratio of between-subject to total variance, so a wide spread of true values can carry a decent ICC while absolute agreement stays poor. The limits of agreement are the number that matters when you are classifying an individual patient against a fixed threshold.
The thresholds it has to serve
| Source | Threshold for tibial malrotation |
|---|---|
| Bell SW, et al. 2014 (PMID 23140906) | 5.8° |
| Berger RA, et al. 1998 (PMID 9917679) | Severe category begins at 7° |
| Nicoll D, Rowley DI 2010 (PMID 20798441) | 9° |
A ±6 to ±9 degree threshold is a decision band 12 to 18 degrees wide. The inter-observer limits of agreement span 24.9 degrees. So the measurement uncertainty is roughly 1.4 to 2 times the width of the decision band.
That ratio is worth stating carefully, because the sloppier version of this argument divides 24.9 by 9 and announces "three times the threshold", which compares a two-sided span with a one-sided threshold. The corrected ratio is smaller and still decisive: you cannot reliably sort an individual knee into a category when your uncertainty is wider than the category.
The landmark is part of the problem
Heyse TJ, et al. Knee 2015 (PMID 26043879) compared reference axes in 55 knees (MRI-based, which is a limitation for direct transfer to CT protocols). The most reliable tibial references were the tibial epicondylar axis and the posterior tibial edge.
The least reliable was the tibial tubercle: inter-observer ICC 0.632, intra-observer ICC 0.526. That is the reference Berger's protocol uses, and with it most of the outcome literature that followed.
So a substantial part of the field's rotation evidence is built on the least reproducible landmark available, which is a structural problem rather than a criticism of any one paper.
What noise this size does to an outcome literature
When the measurement is this variable, the studies built on it behave strangely, and they do.
Young SW, Saffi M, Spangehl MJ, Clarke HD. Knee 2018 (PMID 29526396), described by its authors as the largest study of component rotation after TKA. 71 knees with unexplained pain versus 41 well-functioning controls:
| Painful | Control | p | |
|---|---|---|---|
| Femoral rotation | 0.6° external | 1.0° external | 0.4 |
| Tibial rotation | 11.2° internal | 9.5° internal | 0.3 |
| Combined | 10.5° internal | 8.5° internal | 0.25 |
All null. And the striking line: 59% of the painful knees and 49% of the well-functioning controls measured beyond 9 degrees of tibial internal rotation. Half the good knees are "malrotated" by the threshold that is supposed to explain the bad ones.
Two more studies make the same point by collision:
- Elkins JM, et al. J Arthroplasty 2023 (PMID 36963529): 93 well-functioning gap-balanced TKAs, minimum two years, mean Knee Society Score 185.7, mean range of motion 128.5°. Mean tibial rotation relative to the tubercle: 17.2 ± 7.9° internal.
- Bédard M, et al. 2011 (PMID 21533528): knees revised for stiffness had a mean of 13.7°.
The asymptomatic, excellent knees measure more malrotated than the stiff knees that required revision. That can only be reconciled by different tibial references, which is the whole point.
And a mechanical study to close it: Hutter EE, Granger JF, Beal MD, Siston RA. CORR 2013 (PMID 23392991) aligned the tibial component to four different axes in 10 cadaveric knees with custom navigation and found no change in knee stability or passive kinematics across the four techniques.
The foundational papers are smaller than their reputation
- Berger 1998, the source of the dose-response gradient that anchors this field: 30 patients, 20 controls, retrospective, single centre, 1998 implants, with heavily overlapping ranges (3-8° versus 7-17°).
- Barrack 2001 (PMID 11716424), the origin of the frequently repeated "five times more likely": CT obtained in 14 symptomatic and 14 matched controls. The ratio rests on n=14 versus n=14 and is quoted far beyond what it can carry.
- Valkering KP, et al. Acta Orthop 2015;86(4):432-9 (PMID 25708694), often cited as proof that rotation determines outcome: the correlations are computed between study-level means, not between patients. That is an ecological correlation, it inflates systematically relative to the individual-patient relationship, and it rests on eleven data points. The authors themselves could not derive a threshold.
None of this means rotation is irrelevant. It means the evidence tying a specific number to a specific patient's pain is far weaker than its citation frequency suggests, and measurement noise is a sufficient explanation for much of the inconsistency.
What this asks of planning software
This is the part that concerns us directly, and it cuts both ways.
The optimistic reading is that automated, volume-derived measurement should beat manual landmark picking, because the human variability above comes largely from where two observers put the same landmark. The honest qualification is that this has to be demonstrated on the same footing: reported as limits of agreement against a reference standard, on a defined population, not asserted because the pipeline is automated.
Three consequences we hold ourselves to:
- A single number without an uncertainty is a false claim of precision. If the underlying literature cannot separate 8 degrees from 14 degrees between two observers, a planner that reports "11.2°" and nothing else is presenting a confidence the field does not have.
- Classification near a boundary is where the honesty is tested. A phenotype or a malrotation category assigned to a knee sitting close to a cut-off should say so, rather than pick a side silently.
- Reproducibility is a claim that needs its own study. It is not a by-product of segmentation accuracy, and it is separate from any claim about clinical outcome.
Salnus software is Research Use Only. We treat measurement reproducibility as the thing to prove first, before any claim about alignment strategy or outcome, and the same standard applies to every tool in the 2026 comparison of orthopedic 3D planning software. It is also the reason we read the RACER-Knee trial narrowly rather than as vindication.
Bottom line
CT rotation measurement in TKA carries inter-observer limits of agreement spanning about 25 degrees against decision thresholds of 6 to 9 degrees, uses a tibial reference that is among the least reproducible available, and produces an outcome literature in which half the well-functioning knees fall on the pathological side of the line. The field states this openly and has not fixed it. For anyone building automated measurement, that is not a marketing opportunity, it is a specification: prove reproducibility first, report uncertainty always, and keep outcome claims out of it until there is evidence. For how alignment targets sit on top of these measurements, see TKA alignment philosophies compared and CPAK classification explained.
Reviewed by the Salnus biomedical engineering team.