Open-weight AI structures 2.18 million German radiology reports on hospital hardware

·

A German research team says it used an open-weight artificial intelligence model running on local hospital hardware to convert more than 2.18 million radiology reports into structured data, showing that archive-scale cleanup of medical records can be done on-premises rather than through outside AI services.

The work, described in a preprint posted on arXiv, focuses on radiology reports — the written summaries produced by clinicians after reviewing scans — not on interpreting medical images themselves. That distinction matters because hospitals sit on vast archives of free-text reports that are difficult to search, compare and aggregate for research, quality measurement and machine-learning dataset creation. Structured reporting has been encouraged for years, but adoption during routine dictation has been uneven. The new study instead tackles the backlog: converting existing archives into a more usable format while keeping patient data inside the hospital’s own systems.

The preprint, titled “Autonomous Structuring of Radiology Reports Across Modalities at Archive Scale Using an Open-Weight Large Language Model,” was led by first author Friedrich Puttkammer and senior author Keno K. Bressem, with researchers from Charité – Berlin and the Technical University of Munich, or TUM, among others. The team built a pipeline around 150 hierarchically organized templates covering five imaging modalities and 37 examination types. Using the open-weight model gpt-oss-120B with no fine-tuning, they developed templates on reports from Charité and then tested and deployed the system on German-language radiology reports from TUM dating from 2010 through 2025. In the full archive run, the system processed 2,186,982 eligible reports and produced structured output for 96.5% of them. Because some reports described more than one exam type, that yielded 2,401,544 structured reports. Reported throughput was 1,258 structured reports per hour on one NVIDIA H200 graphics processor.

The researchers also tested how well the system selected the right template and how faithfully it captured report content. In a set of 914 eligible reports labeled by experts, the system chose an optimal template set for 74.4%, or 680 reports, and an appropriate template set for 82.3%, or 752 reports. Performance was notably better for simpler cases: 87.7% appropriate template selection for single-region reports, compared with 54.1% for multi-region reports, which the paper identifies as the main source of errors.

For structuring quality, five radiology residents reviewed 920 radiography and CT reports field by field, with unclear cases adjudicated by two board-certified radiologists. Across 24,638 fields, reviewers left 88.7% unchanged. The paper reported macro semantic textual similarity — a measure of how closely the structured output matched corrected reference versions — of 0.95 for radiography and 0.97 for CT. Reviewers flagged unsupported content in 1.0% of radiography reports and 1.5% of CT reports in that sample.

The broader significance is less about a new model than about deployment. Earlier work, including a 2023 Radiology study using GPT-4, suggested large language models could structure radiology text on smaller or narrower datasets. What sets this paper apart is the claimed full-archive, multimodality run using an open-weight model hosted locally. For hospitals, that could reduce the need to send sensitive patient information to third-party application programming interfaces, while also giving institutions more control over cost and operations.

The findings come with important caveats. The study is a preprint, meaning it has not yet been peer-reviewed. And while the archive run covered multiple modalities, field-level structuring quality was assessed only for radiography and CT, not MRI, ultrasound or interventional radiology. Multi-region reports also remained a clear weak point. Even so, the paper argues that “an open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity.”

Tags: #ai, #radiology, #healthit, #medical-records