Open Orthopaedic Imaging Datasets
What exists in open musculoskeletal imaging, what closed and why, and why knee CT segmentation is the gap. Based on 59 records verified against their sources.
Key takeaways
We built a verified index of open musculoskeletal imaging datasets and checked every record against its own source. Three things came out of it. Spine is well served and the knee is not: of 23 spine datasets, 12 carry CT with segmentation labels; of 25 knee datasets, two carry knee bone masks at all, and both are cadaveric. Datasets disappear, and nobody records why: three in this index were once public and can no longer be downloaded, including one withdrawn in 2019 when the implant manufacturer that supplied the data revoked permission. And licence terms are not a property of the dataset: they change by subset, so the part you need may be governed by different terms from the part you read about. The index is at salnus.com/datasets.
What is actually out there
Fifty-nine datasets, each verified against its own page, DOI record or publication. The distribution is not even:
| Anatomy | Datasets | With CT | CT + segmentation |
|---|---|---|---|
| Spine | 23 | 18 | 12 |
| Pelvis | 20 | 15 | 10 |
| Hip | 18 | 12 | 8 |
| Shoulder | 12 | 7 | 5 |
| Knee | 25 | 7 | 1 |
| Bone tumour | 7 | 5 | 3 |
Counted the way the index counts: a record enters the last column only if its label type is recorded as segmentation. Eighteen further records are labelled mixed, and some of those do carry masks alongside other annotation, so where that changes the answer we name the record.
The knee row is the striking one. It has the most datasets of any region and the fewest usable CT segmentations. The reason is that knee research has been MRI-led for two decades, driven by cartilage and the Osteoarthritis Initiative. That produced excellent MRI resources and left CT, which is what arthroplasty planning actually runs on, almost empty.
Going through the seven knee datasets that include CT, one by one, makes the gap concrete:
| Dataset | What limits it |
|---|---|
| 3D Visible Human | Femur, patella, tibia and fibula masks under CC BY 4.0. Two subjects, both cadaveric |
| VSDFullBodyBoneReconstruction | Complete knee-region bone masks, 30 postmortem forensic subjects. Licence is contradictory across sources: the article reads CC BY 4.0, the Zenodo record and repository licence read CC BY-NC-SA 4.0, so commercial use is unresolved |
| PlaTiF | Tibia masks and Schatzker grades, but the CT is one coronal slice per patient rather than a volume |
| Keast 2023 | Tibia and fibula meshes and a statistical shape model, no images: the source CT cannot be shared |
| NMDID | Over 15,000 postmortem whole-body CTs, no segmentation labels, institutional credentials and a data-use agreement required |
| MOST | Large in-vivo osteoarthritis cohort, CT only in later study cycles, classification labels, access by application |
| Fractured Limbs 5K+ | 16 knee cases, but the public repository exposes at most two cases per bone type and the rest require a password from the authors; no licence is stated |
Two of the seven carry volumetric knee bone masks, and both are cadaveric. A third, PlaTiF, does carry tibia masks from 186 living patients, but its CT is one coronal slice per patient rather than a volume. There is no open, volumetric, living-patient knee CT cohort with bone segmentation. That is the gap in one sentence.
Only seven datasets are both segmented and commercially usable
Across all anatomy, 21 of the 59 records have their label type recorded as segmentation. Of those, seven also record commercial use as permitted: 3D Visible Human, CT-ORG, BoneDat, SPIDER, CadAIver, Spinal-Multiple-Myeloma-SEG and Soft-Tissue-Sarcoma. Five of the seven are spine, pelvis or tumour. That number is the one to remember if you are building a product rather than writing a paper: the open-data route to a commercial segmentation model is narrower than the headline count suggests.
There is a second trap underneath it. Permission to use is not permission to redistribute. A dataset can allow commercial use and forbid re-hosting; 21 of the 59 records restrict redistribution. These are separate questions and they are answered in different clauses, which is why we record them as separate fields.
Datasets close, and the record of why they closed is disappearing
This is the part that surprised us, and the reason the index keeps closed datasets instead of deleting them.
SKI10 was the reference knee MRI segmentation benchmark for most of a decade. The challenge page now carries this sentence:
"In 2019, we were informed that we are no longer allowed to share the SKI10 dataset, which was kindly provided to us by Biomet, Inc."
An implant manufacturer supplied the data, and years later withdrew permission. The dataset had been cited in a generation of segmentation papers. Anyone trying to reproduce that work today cannot obtain the data, and there is no licence to appeal to: SKI10 never published one.
There is a second layer. The original ski10.org domain was lost and re-registered by someone else; it now redirects to an unrelated site, and 2025 archive captures of it show gambling pages. A researcher following an old citation lands somewhere that has nothing to do with the data. We record the challenge page instead, and we say plainly that the old domain must not be cited.
K2S, a 2022 challenge with 350 patients and six-class knee segmentation, states it directly:
"We are not sharing any data after the challenge has been concluded."
KneeMRI from Rijeka, 917 volumes annotated for ACL injury, simply went dark. The last working capture is May 2023; by August 2024 the page redirected, and today it returns 404. No announcement, no explanation.
Three datasets, three different ways to vanish: a supplier revoking permission, a challenge policy, and silence. None of the major dataset listings record any of this. A researcher searching for SKI10 finds a dead end and no explanation of what happened.
That is why closed records stay in this index, with the date and the reason. A resource that says "this existed, here is what it was, here is why you cannot have it" is more useful than one that quietly omits it.
Licences are read at subset level, not dataset level
The clearest example is the most widely used dataset in the list.
TotalSegmentator is CC BY 4.0 and genuinely permits commercial reuse. But of the knee bones, only the femur is in the main model. Tibia, patella and fibula live in a separate appendicular_bones task, that model carries its own licence requiring contact for commercial use, and its training data was never published.
So a team that reads "TotalSegmentator is CC BY 4.0", builds on the tibia labels and ships a product has made a licensing error while quoting the licence correctly. The dataset-level statement was true. The subset-level answer was different.
Two more of the same shape. TotalSegmentator MRI is a separate release from the same group under CC BY-NC-SA 2.0, so the CT dataset is commercially usable and the MRI one is not, despite sharing a name. And CADS aggregates 40 source datasets under mixed terms, with some sources marked "no training, no redistribution" in the paper's own source table.
Two things that look like data and are not
Two records in the index are phantom studies: imaging of physical models rather than people. We got this wrong at first and marked them as patient data, which is worth admitting because the error is easy to make and consequential. A phantom has no load-bearing posture and no physiological soft-tissue contrast, so alignment work validated on one has not been validated on anatomy.
Six further records are cadaveric, which carries a milder version of the same limitation: real anatomy, but unloaded. For a discipline where alignment is measured standing up, that distinction belongs in the metadata rather than in a footnote, so the index records what was imaged as its own field.
Why this matters beyond the index
If you are training or validating a knee CT segmentation model, there is no public dataset you can validate against. Your external validation has to come from a second clinical site, with the ethics and data-sharing work that implies. That is a real cost and it should be planned for, not discovered late.
It also sets a limit on what open-data claims can mean in this area. A model trained on whole-body CT where the knee appears incidentally has not been validated on the knee, and the published evidence to check that claim does not exist yet.
We have written separately about the layer beneath this, why measurement reproducibility is the harder problem, and about how planning software is classified under MDR.
The index
All 59 records are at salnus.com/datasets, filterable by modality, anatomy, label type and access route. Each record carries the licence, whether commercial use and redistribution are permitted, whether the subjects were patients, cadavers or phantoms, and the date it was last verified. Where something could not be confirmed, the field is empty and the reason is recorded rather than guessed.
Country, year, redistribution and subject status on some records are adapted from the BoneHub public dataset collection (CC BY 4.0, Alavi and Asseln, University of Twente), which covers datasets suitable for deriving 3D bone shape and goes deeper than we do on mesh, CAD and landmark detail.
Corrections are welcome. If a record is wrong, or a dataset has moved or closed since we checked, tell us and we will fix it and update the date.
Reviewed by the Salnus biomedical engineering team.