Machine Learning Model Shows Promise for Improving Metanephrine Testing Accuracy
Posted on 30 Jul 2026
Pheochromocytomas and paragangliomas are rare tumors that form in or near the adrenal glands and cause overproduction of stress hormones. Plasma-free metanephrines are the recommended first-line test, but mild elevations in people without these tumors can generate false positives and lead to unnecessary follow-up. To improve interpretation, laboratories are investigating machine-learning approaches that contextualize results and improve rule-out accuracy. New findings demonstrate how such models can both enhance and complicate interpretation of metanephrine testing in routine care.
Researchers at Samsung Medical Center (Seoul, South Korea) tested several machine learning (ML) algorithms to augment plasma-free metanephrine testing for suspected pheochromocytoma and paragangliomas (PPGL). The approach combined metanephrine measurements with structured electronic health record (EHR) data to reduce false-positive results. Inputs described in the analysis included kidney and urine biomarkers, medication profiles, and the presence of other diseases.
In a real-world cohort, the team analyzed data from 20,516 adults who underwent metanephrine testing at Samsung Medical Center between 2011 and 2024. Of the 19,797 patients ultimately determined not to have PPGL, 25.2% showed metanephrine elevations that could prompt a false-positive result. Initial ML models suggested that integrating clinical context with metanephrine results improved real-world discrimination and appeared to boost the test’s performance to an excellent level.
Subsequent checks revealed important caveats. Additional analyses indicated that some apparent gains reflected “shortcut learning” from informative missingness, with models inferring clinician suspicion based on which follow-up tests were ordered rather than independent biochemical signal. The investigators emphasized that ML should be judged not only by headline performance metrics but also by whether it learns the intended clinical signal, highlighting the need for external validation and careful audits.
The findings were presented at the Association for Diagnostics & Laboratory Medicine (ADLM) 2026 meeting in Anaheim, California (Abstract B-091).
“After abstract submission, while testing the prototype app with the developed ML model, we realized that some of the apparent improvement might reflect patterns of clinical workup rather than independent biochemical information. Additional robustness analyses showed that much of the improvement was explained by shortcut learning from informative missingness,” said Se-eun Koo, one of the study’s co-authors and a clinical chemistry fellow in the department of laboratory medicine and genetics at Samsung Medical Center in Seoul, South Korea.
“This study shows that machine learning in laboratory medicine should be evaluated not only by performance metrics, but also by whether the model is learning the intended clinical signal. That is why external validation and careful audits for shortcut learning are so important,” Koo said.
Related Links
Samsung Medical Center