Could machine learning be used as a diagnostic aid?

A study analyzing data from over 20,000 patients demonstrates that integrating ML with clinical data enhances diagnostic precision for PPGL, but emphasizes the need to prevent shortcut learning and ensure models learn true clinical signals.

Research that was presented at ADLM 2026 shows how machine learning (ML) could play an important role in enhancing the accuracy of a standard blood test used to diagnose rare tumors that form in or near the adrenal glands. The findings are reported in an ADLM press release.

The study also highlights the complexity of using ML models in laboratory medicine. As healthcare strives to harness the power of ML, this research could guide labs in rigorously validating their ML algorithms by illustrating the kinds of pitfalls they need to look out for. 

Plasma-free metanephrines are the recommended first-line test for detecting masses known as pheochromocytomas and paragangliomas (PPGL), which cause the body to overproduce stress hormones. Left untreated, PPGL can cause heart problems, headaches, high blood pressure, and other issues.  

While the metanephrine test effectively detects people who have PPGL, it does less well at ruling out everyone who doesn’t have the condition. That’s because mild elevations in metanephrines — which are metabolites derived from stress hormones — are often seen in patients who do not have these tumors, leading to false-positive results. 

The researchers tested several ML algorithms to assess whether, and how much, they could bolster accuracy by reducing the likelihood of these false-positives. They analyzed data from 20,516 adults who underwent metanephrine testing at Samsung Medical Center between 2011 and 2024. Of the 19,797 patients tested who ultimately did not have PPGL, 25.2% demonstrated elevations in metanephrine that could trigger a false-positive result.  

“Our initial machine-learning models suggested that combining plasma metanephrine results with structured clinical information from the electronic health record could improve real-world discrimination,” said Se-eun Koo, one of the study’s co-authors and a clinical chemistry fellow in the department of laboratory medicine and genetics at Samsung Medical Center in Seoul, South Korea. 

Specifically, Koo and co-author Dr. Soo-Youn Lee found that, while metanephrine showed good reliability as a clinical marker, integrating ML to assess relevant clinical context — including kidney and urine biomarkers, patients’ medications, and the presence of other diseases — appeared to boost the test’s performance to make it excellent.  

After submitting their initial findings to the Association for Diagnostics & Laboratory Medicine (ADLM), Koo and Dr. Lee performed additional analyses to determine whether the ML algorithms were using any “shortcuts” to learning that could introduce errors or biases. (As an example: If ML consistently assesses images of computer hackers wearing hoodie sweatshirts, it might conclude that anyone wearing a hoodie is a hacker.)     

“After abstract submission, while testing the prototype app with the developed ML model, we realized that some of the apparent improvement might reflect patterns of clinical workup rather than independent biochemical information,” Koo said. “Additional robustness analyses showed that much of the improvement was explained by shortcut learning from informative missingness.”                

In other words, while the ML models seemed to improve diagnostic accuracy, they did so by learning which follow-up tests were ordered — something clinicians do when they already suspect PPGL. 

“This study shows that machine learning in laboratory medicine should be evaluated not only by performance metrics, but also by whether the model is learning the intended clinical signal,” Koo said. “That is why external validation and careful audits for shortcut learning are so important.” 

Visit ADLM for more information

About the Author

Sign up for our eNewsletters
Get the latest news and updates