8 min read

How Reliable Is TKA Rotation Measurement?

CT measurement of knee component rotation has 95% limits of agreement wider than the thresholds it is used to judge. What that means for outcome studies and for AI planning.

Burak Serteser
MeasurementReproducibilityTotal Knee ArthroplastyClinical EvidenceCTSurgical PlanningOrthopedic Surgery

Key takeaways

Component rotation after total knee arthroplasty is measured on CT, compared against thresholds of roughly 6 to 9 degrees, and used to explain painful knees. The problem is that the measurement's own 95% limits of agreement between observers span about 25 degrees, which is wider than the decision band it is meant to serve. The authors of the most rigorous reliability study put it plainly: these measurements are not sufficiently reliable for routine clinical use. This is not a fringe position and nobody has solved it. It also sets the bar for any software, ours included, that proposes to automate the measurement: automation is only an improvement if it is demonstrably more reproducible, and that has to be shown rather than assumed.

The reliability data

Toms AP, Rifai T, Whitehouse C, McNamara I. Eur Radiol 2022. PMID 35142899 · PMC9122870 full text

A prospective reliability study embedded in the CAPAbility randomised trial. n=80, pre- and postoperative CT, Berger protocol.

MeasurementInter-observer ICCInter-observer 95% limits of agreement (span)Intra-observer span
Preoperative femoral version0.65-0.728.1°7.6°
Preoperative tibial version0.59-0.6426.0°20.5°
Postoperative femoral rotation0.79-0.819.0°7.9°
Postoperative tibial rotation0.78-0.8024.9°23.6°

Two things about how to read this table, because both are commonly got wrong.

These are spans, not plus-or-minus values. A 24.9-degree span is roughly ±12.5 degrees around the mean difference. Quoting it as "±25 degrees" doubles the real figure.

The ICCs look respectable. Postoperative values sit at 0.78 to 0.81, which most papers would call good agreement. That is exactly why the ICC alone is misleading here: it is a ratio of between-subject to total variance, so a wide spread of true values can carry a decent ICC while absolute agreement stays poor. The limits of agreement are the number that matters when you are classifying an individual patient against a fixed threshold.

The thresholds it has to serve

SourceThreshold for tibial malrotation
Bell SW, et al. 2014 (PMID 23140906)5.8°
Berger RA, et al. 1998 (PMID 9917679)Severe category begins at 7°
Nicoll D, Rowley DI 2010 (PMID 20798441)

A ±6 to ±9 degree threshold is a decision band 12 to 18 degrees wide. The inter-observer limits of agreement span 24.9 degrees. So the measurement uncertainty is roughly 1.4 to 2 times the width of the decision band.

That ratio is worth stating carefully, because the sloppier version of this argument divides 24.9 by 9 and announces "three times the threshold", which compares a two-sided span with a one-sided threshold. The corrected ratio is smaller and still decisive: you cannot reliably sort an individual knee into a category when your uncertainty is wider than the category.

The landmark is part of the problem

Heyse TJ, et al. Knee 2015 (PMID 26043879) compared reference axes in 55 knees (MRI-based, which is a limitation for direct transfer to CT protocols). The most reliable tibial references were the tibial epicondylar axis and the posterior tibial edge.

The least reliable was the tibial tubercle: inter-observer ICC 0.632, intra-observer ICC 0.526. That is the reference Berger's protocol uses, and with it most of the outcome literature that followed.

So a substantial part of the field's rotation evidence is built on the least reproducible landmark available, which is a structural problem rather than a criticism of any one paper.

What noise this size does to an outcome literature

When the measurement is this variable, the studies built on it behave strangely, and they do.

Young SW, Saffi M, Spangehl MJ, Clarke HD. Knee 2018 (PMID 29526396), described by its authors as the largest study of component rotation after TKA. 71 knees with unexplained pain versus 41 well-functioning controls:

PainfulControlp
Femoral rotation0.6° external1.0° external0.4
Tibial rotation11.2° internal9.5° internal0.3
Combined10.5° internal8.5° internal0.25

All null. And the striking line: 59% of the painful knees and 49% of the well-functioning controls measured beyond 9 degrees of tibial internal rotation. Half the good knees are "malrotated" by the threshold that is supposed to explain the bad ones.

Two more studies make the same point by collision:

  • Elkins JM, et al. J Arthroplasty 2023 (PMID 36963529): 93 well-functioning gap-balanced TKAs, minimum two years, mean Knee Society Score 185.7, mean range of motion 128.5°. Mean tibial rotation relative to the tubercle: 17.2 ± 7.9° internal.
  • Bédard M, et al. 2011 (PMID 21533528): knees revised for stiffness had a mean of 13.7°.

The asymptomatic, excellent knees measure more malrotated than the stiff knees that required revision. That can only be reconciled by different tibial references, which is the whole point.

And a mechanical study to close it: Hutter EE, Granger JF, Beal MD, Siston RA. CORR 2013 (PMID 23392991) aligned the tibial component to four different axes in 10 cadaveric knees with custom navigation and found no change in knee stability or passive kinematics across the four techniques.

The foundational papers are smaller than their reputation

  • Berger 1998, the source of the dose-response gradient that anchors this field: 30 patients, 20 controls, retrospective, single centre, 1998 implants, with heavily overlapping ranges (3-8° versus 7-17°).
  • Barrack 2001 (PMID 11716424), the origin of the frequently repeated "five times more likely": CT obtained in 14 symptomatic and 14 matched controls. The ratio rests on n=14 versus n=14 and is quoted far beyond what it can carry.
  • Valkering KP, et al. Acta Orthop 2015;86(4):432-9 (PMID 25708694), often cited as proof that rotation determines outcome: the correlations are computed between study-level means, not between patients. That is an ecological correlation, it inflates systematically relative to the individual-patient relationship, and it rests on eleven data points. The authors themselves could not derive a threshold.

None of this means rotation is irrelevant. It means the evidence tying a specific number to a specific patient's pain is far weaker than its citation frequency suggests, and measurement noise is a sufficient explanation for much of the inconsistency.

What this asks of planning software

This is the part that concerns us directly, and it cuts both ways.

The optimistic reading is that automated, volume-derived measurement should beat manual landmark picking, because the human variability above comes largely from where two observers put the same landmark. The honest qualification is that this has to be demonstrated on the same footing: reported as limits of agreement against a reference standard, on a defined population, not asserted because the pipeline is automated.

Three consequences we hold ourselves to:

  • A single number without an uncertainty is a false claim of precision. If the underlying literature cannot separate 8 degrees from 14 degrees between two observers, a planner that reports "11.2°" and nothing else is presenting a confidence the field does not have.
  • Classification near a boundary is where the honesty is tested. A phenotype or a malrotation category assigned to a knee sitting close to a cut-off should say so, rather than pick a side silently.
  • Reproducibility is a claim that needs its own study. It is not a by-product of segmentation accuracy, and it is separate from any claim about clinical outcome.

Salnus software is Research Use Only. We treat measurement reproducibility as the thing to prove first, before any claim about alignment strategy or outcome, and the same standard applies to every tool in the 2026 comparison of orthopedic 3D planning software. It is also the reason we read the RACER-Knee trial narrowly rather than as vindication.

Bottom line

CT rotation measurement in TKA carries inter-observer limits of agreement spanning about 25 degrees against decision thresholds of 6 to 9 degrees, uses a tibial reference that is among the least reproducible available, and produces an outcome literature in which half the well-functioning knees fall on the pathological side of the line. The field states this openly and has not fixed it. For anyone building automated measurement, that is not a marketing opportunity, it is a specification: prove reproducibility first, report uncertainty always, and keep outcome claims out of it until there is evidence. For how alignment targets sit on top of these measurements, see TKA alignment philosophies compared and CPAK classification explained.

Ana Çıkarımlar

Total diz artroplastisi sonrası komponent rotasyonu BT'de ölçülür, yaklaşık 6 ile 9 derecelik eşiklerle karşılaştırılır ve ağrılı dizleri açıklamak için kullanılır. Sorun şu ki ölçümün kendi gözlemciler arası %95 uyum sınırları yaklaşık 25 derecelik bir aralığa yayılıyor ve bu, hizmet etmesi beklenen karar bandından geniş. En titiz güvenilirlik çalışmasının yazarları bunu açıkça söylüyor: bu ölçümler rutin klinik kullanım için yeterince güvenilir değil. Bu marjinal bir görüş değil ve kimse çözmedi. Aynı zamanda ölçümü otomatikleştirmeyi öneren her yazılım için, bizimki dahil, çıtayı belirliyor: otomasyon ancak gösterilebilir biçimde daha tekrarlanabilirse bir iyileştirmedir, ve bu varsayılmaz, gösterilir.

Güvenilirlik verisi

Toms AP, Rifai T, Whitehouse C, McNamara I. Eur Radiol 2022. PMID 35142899 · PMC9122870 tam metin

CAPAbility randomize çalışmasına yerleştirilmiş prospektif güvenilirlik çalışması. n=80, ameliyat öncesi ve sonrası BT, Berger protokolü.

ÖlçümGözlemciler arası ICCGözlemciler arası %95 uyum sınırı (aralık)Gözlemci içi aralık
Ameliyat öncesi femoral versiyon0,65-0,728,1°7,6°
Ameliyat öncesi tibial versiyon0,59-0,6426,0°20,5°
Ameliyat sonrası femoral rotasyon0,79-0,819,0°7,9°
Ameliyat sonrası tibial rotasyon0,78-0,8024,9°23,6°

Bu tabloyu okurken iki nokta önemli, çünkü ikisi de sık yanlış aktarılıyor.

Bunlar aralıktır, artı-eksi değildir. 24,9 derecelik aralık, ortalama farkın çevresinde kabaca ±12,5 derecedir. "±25 derece" diye aktarmak gerçek rakamı ikiye katlar.

ICC değerleri saygın görünüyor. Ameliyat sonrası değerler 0,78 ile 0,81 arasında, ki çoğu makale buna iyi uyum der. ICC'nin tek başına yanıltıcı olmasının nedeni tam olarak budur: denekler arası varyansın toplam varyansa oranıdır, dolayısıyla gerçek değerlerin geniş dağıldığı bir örneklemde mutlak uyum kötüyken ICC iyi kalabilir. Bireysel bir hastayı sabit bir eşiğe göre sınıflandırırken önemli olan sayı uyum sınırlarıdır.

Hizmet etmesi gereken eşikler

KaynakTibial malrotasyon eşiği
Bell SW, ve ark. 2014 (PMID 23140906)5,8°
Berger RA, ve ark. 1998 (PMID 9917679)Ağır kategori 7°'de başlıyor
Nicoll D, Rowley DI 2010 (PMID 20798441)

±6 ile ±9 derecelik bir eşik, 12 ile 18 derece genişliğinde bir karar bandıdır. Gözlemciler arası uyum sınırları ise 24,9 derecelik bir aralıktır. Yani ölçüm belirsizliği, karar bandının kabaca 1,4 ile 2 katı genişliğindedir.

Bu oranı dikkatle söylemek gerekiyor, çünkü bu argümanın özensiz sürümü 24,9'u 9'a bölüp "eşiğin üç katı" diyor; bu, iki yönlü bir aralığı tek yönlü bir eşiğe bölmektir. Düzeltilmiş oran daha küçük ve yine belirleyici: belirsizliğiniz kategoriden genişken bireysel bir dizi o kategoriye güvenilir biçimde yerleştiremezsiniz.

Landmark sorunun bir parçası

Heyse TJ, ve ark. Knee 2015 (PMID 26043879) 55 dizde referans eksenleri karşılaştırdı (MR tabanlı, ki BT protokollerine doğrudan aktarım açısından bir sınırlılıktır). En güvenilir tibial referanslar tibial epikondiler eksen ve posterior tibial kenar çıktı.

En güvenilmezi tibial tüberkül: gözlemciler arası ICC 0,632, gözlemci içi ICC 0,526. Berger protokolünün kullandığı referans budur ve ardından gelen sonuç literatürünün çoğu da onu kullanır.

Yani alanın rotasyon kanıtının önemli bir kısmı, eldeki en az tekrarlanabilir landmark üzerine kurulmuştur. Bu tek bir makaleye yönelik bir eleştiri değil, yapısal bir sorundur.

Bu büyüklükte gürültü sonuç literatürüne ne yapar

Ölçüm bu kadar değişkense üzerine kurulan çalışmalar tuhaf davranır, ve davranıyorlar.

Young SW, Saffi M, Spangehl MJ, Clarke HD. Knee 2018 (PMID 29526396), yazarlarının ifadesiyle TKA sonrası komponent rotasyonu üzerine bugüne kadarki en büyük çalışma. 71 açıklanamayan ağrılı diz ile 41 iyi çalışan kontrol:

AğrılıKontrolp
Femoral rotasyon0,6° dış1,0° dış0,4
Tibial rotasyon11,2° iç9,5° iç0,3
Kombine10,5° iç8,5° iç0,25

Hepsi null. Ve çarpıcı satır: ağrılı dizlerin %59'u, iyi çalışan kontrollerin %49'u 9 dereceden fazla tibial iç rotasyonda ölçüldü. Kötü dizleri açıklaması beklenen eşiğe göre iyi dizlerin yarısı "malrotasyonlu".

Aynı noktayı çarpışarak gösteren iki çalışma daha:

  • Elkins JM, ve ark. J Arthroplasty 2023 (PMID 36963529): 93 iyi çalışan gap-balanced TKA, en az iki yıl takip, ortalama Knee Society Score 185,7, ortalama hareket açıklığı 128,5°. Tüberküle göre ortalama tibial rotasyon: 17,2 ± 7,9° iç.
  • Bédard M, ve ark. 2011 (PMID 21533528): sertlik nedeniyle revize edilen dizlerde ortalama 13,7°.

Semptomsuz, mükemmel çalışan dizler, revizyon gerektiren sert dizlerden daha fazla malrotasyonlu ölçülüyor. Bu ancak farklı tibial referanslarla uzlaştırılabilir, ki mesele tam da budur.

Kapatmak için bir mekanik çalışma: Hutter EE, Granger JF, Beal MD, Siston RA. CORR 2013 (PMID 23392991) 10 kadavra dizinde tibial komponenti özel navigasyonla dört farklı eksene hizaladı ve dört teknik arasında diz stabilitesinde veya pasif kinematikte hiçbir değişiklik bulamadı.

Temel makaleler ünlerinden küçük

  • Berger 1998, bu alanı çıpalayan doz-yanıt gradyanının kaynağı: 30 hasta, 20 kontrol, retrospektif, tek merkez, 1998 implantları ve ağır biçimde örtüşen aralıklar (3-8° ile 7-17°).
  • Barrack 2001 (PMID 11716424), sık tekrarlanan "beş kat daha olası" ifadesinin kaynağı: BT yalnızca 14 semptomatik ve 14 eşleştirilmiş kontrolde çekilmiş. Oran n=14'e karşı n=14'e dayanıyor ve taşıyabileceğinin çok ötesinde alıntılanıyor.
  • Valkering KP, ve ark. Acta Orthop 2015;86(4):432-9 (PMID 25708694), rotasyonun sonucu belirlediğinin kanıtı olarak sık alıntılanıyor: korelasyonlar hastalar arasında değil çalışma düzeyi ortalamalar arasında hesaplanmış. Bu ekolojik korelasyondur, bireysel hasta ilişkisine göre sistematik olarak şişer ve on bir veri noktasına dayanır. Yazarlar da bir eşik türetememiştir.

Bunların hiçbiri rotasyonun önemsiz olduğu anlamına gelmez. Belirli bir sayıyı belirli bir hastanın ağrısına bağlayan kanıtın, alıntılanma sıklığının ima ettiğinden çok daha zayıf olduğu ve tutarsızlığın büyük kısmını ölçüm gürültüsünün açıklamaya yettiği anlamına gelir.

Bu, planlama yazılımından ne ister

Doğrudan bizi ilgilendiren kısım burası ve iki tarafı da kesiyor.

İyimser okuma şu: otomatik, hacimden türetilmiş ölçüm elle landmark seçmeyi geçmelidir, çünkü yukarıdaki insan değişkenliğinin büyük kısmı iki gözlemcinin aynı landmarkı nereye koyduğundan geliyor. Dürüst şerh ise şu: bunun aynı zeminde gösterilmesi gerekir, yani tanımlı bir popülasyonda, bir referans standarda karşı uyum sınırlarıyla raporlanarak; boru hattı otomatik diye varsayılarak değil.

Kendimizi bağladığımız üç sonuç:

  • Belirsizliği olmayan tek bir sayı, sahte bir kesinlik iddiasıdır. Altta yatan literatür iki gözlemci arasında 8 dereceyi 14 dereceden ayıramıyorsa, "11,2°" deyip susan bir planlayıcı alanın sahip olmadığı bir güveni sunuyordur.
  • Dürüstlük, sınıra yakın olguda sınanır. Bir eşiğe yakın duran dize atanan fenotip veya malrotasyon kategorisi bunu söylemeli, sessizce bir taraf seçmemelidir.
  • Tekrarlanabilirlik kendi çalışmasını gerektiren bir iddiadır. Segmentasyon doğruluğunun yan ürünü değildir ve klinik sonuca dair her iddiadan ayrıdır.

Salnus yazılımı Araştırma Amaçlıdır (RUO). Ölçüm tekrarlanabilirliğini, dizilim stratejisine veya sonuca dair her iddiadan önce kanıtlanması gereken şey olarak görüyoruz; aynı standart 2026 ortopedik 3B planlama yazılımı karşılaştırmasındaki her araç için geçerli. RACER-Knee çalışmasını bir zafer ilanı değil, dar bir sonuç olarak okumamızın nedeni de bu.

Sonuç

TKA'da BT rotasyon ölçümü, 6-9 derecelik karar eşiklerine karşı yaklaşık 25 derecelik gözlemciler arası uyum aralığı taşıyor, mevcut en az tekrarlanabilir tibial referanslardan birini kullanıyor ve iyi çalışan dizlerin yarısının patolojik tarafa düştüğü bir sonuç literatürü üretiyor. Alan bunu açıkça söylüyor ve düzeltmedi. Otomatik ölçüm geliştiren herkes için bu bir pazarlama fırsatı değil, bir şartname: önce tekrarlanabilirliği kanıtla, belirsizliği her zaman raporla, ve kanıt oluşana kadar sonuç iddiasını dışarıda tut. Dizilim hedeflerinin bu ölçümlerin üstüne nasıl oturduğu için TKA dizilim felsefeleri karşılaştırması ve CPAK sınıflandırması yazılarına bakabilirsiniz.

Reviewed by the Salnus biomedical engineering team.

Related Posts

Open Orthopaedic Imaging Datasets9 min readWhat the RACER-Knee Trial Actually Tested6 min readFemoral Rotation in TKA Planning4 min readLocal Planning Software for Domestic Makers7 min read
← All Posts

Orthopedic AI Research Updates

Monthly research digest, product updates, and clinical AI insights.

Unsubscribe anytime.

How Reliable Is TKA Rotation Measurement?, Salnus