2026-09-02
On 1,285 pediatric hand films, readers disagree by 8.7 months SD under age 12; BoneXpert MAE 4.8 months and 0.7% errors >1.8 years vs 6.5 months and 3.4% for one reader.
Pediatric bone age is still mostly read with the 1959 Greulich-Pyle atlas: match the ossification centers on a left-hand radiograph to a standard plate, and write down a skeletal age. That number feeds decisions about growth hormone and precocious puberty. Different radiologists, same film, different ages. Automated systems are already in clinics. Until the size and shape of human disagreement are measured, there is no way to tell whether a model is actually better or merely sitting inside the noise.
Teams at Stanford, Yale, Boston Children's, Cincinnati Children's, and Emory, together with Visiana, the company behind BoneXpert, ran a secondary analysis on images from an earlier multicenter RCT (NCT03530098, Eng et al., 2021). The job is twofold: take human variability apart, then sit two automated methods next to a single reader.
The abstract describes a multi-reader, multi-case design. 1,285 left-hand and wrist films from five US academic centers; four radiologists independently scored each film with Greulich-Pyle, drawn from a pool of 22 readers at nine institutions. Mixed-effects models split variance across image, rater, and institution, then re-estimated inter-reader spread after adjusting for age and sex.
Two automated methods went through an interchangeability analysis against a three-rater consensus:
Reported metrics are MAE, RMSE, and the rate of substantial deviations, defined as missing the reference by more than 1.8 years. The mixed-model formula, the identity of the deep-learning system, and the numerical interchangeability threshold are not in the abstract, so they are not filled in here.
Human scatter depends on age. Under 12 years, the within-image standard deviation among radiologists is 8.7 months; above 12 it falls to 4.2 months (P<0.001). Between-institution spread looks large at 8.9 months, then collapses to 0.3 months after adjusting for patient age: centers see different age mixes, they are not using yardsticks that differ by almost a year. Residual inter-rater spread stays at 1.2 months after the same adjustment. Variability is higher in boys and in younger groups. Some readers carry systematic age- or sex-linked biases.
Against a single human reader, BoneXpert's MAE is 4.8 months versus 6.5 months for the radiologist. The unnamed deep-learning system lands at 6.3 months, essentially the single-reader band. Rates of errors larger than 1.8 years:
| Source | Rate >1.8 years |
| BoneXpert | 0.7% |
| Deep-learning algorithm | 2.4% |
| Single radiologist | 3.4% |
The conclusion is measured: automated methods sit inside, or past, the observed human range. That range is now a number later bone-age models can report against.
Imaging papers often average a few experts, call that ground truth, and announce a tie or a win. This study opens the ground truth. On a young child's film, experts already disagree by the better part of a year. BoneXpert's large-error rate is about one fifth of a single reader's (0.7% vs 3.4%), and MAE is 1.7 months lower. The unnamed deep net is in the same band as a single reader, so "AI beats the doctor" is not a property of every automated method.
For anyone shipping bone-age software, this reads as a calibration sheet. Quote MAE against a single reader or against consensus, and stratify by age and sex. Boys under 12 are the noisiest cell.
The Springer full text is paywalled. Stratified tables, Bland-Altman plots, which readers are biased, and the identity of the deep-learning algorithm are not available from the abstract. Crossref metadata records the conflict: Thodberg is CEO of Visiana, Thrane is an employee, and Visiana funded the study. The company supplied software and technical support, did not see the images or patient-level data, and academic investigators ran the statistics. The wall is there; the funding structure is still "the vendor of the winning tool paid." Discount BoneXpert's lead by a notch.
The data-availability note says no new datasets were generated, because this is a secondary analysis of an existing trial. The 22 readers are academic pediatric radiologists; community hospitals and non-US populations are a separate question. The abstract does not discuss how well Greulich-Pyle travels beyond children of European ancestry. The deep-learning system is anonymous, so it is unclear whether it is the 2017 RSNA challenge winner or the Larson/Eng product line.