At the end of June, the IDEA4RC consortium gathered in Madrid for its final plenary meeting ahead of the project’s review this autumn. Hosted by the Universidad Politécnica de Madrid (UPM), the two-day meeting brought together the eleven clinical centres and technical partners that, over four years, have worked to build a federated platform for the re-use of clinical data on two types of rare cancers, sarcomas and head and neck cancers.
Opening the meeting, project coordinator Annalisa Trama, epidemiologist at the National Cancer Institute of Milan, offered a reflection that set the tone for the two days: a sound legal framework for sharing health data across borders is necessary, but on its own it is not enough. Re-using clinical data for research effectively and responsibly also requires technical capacity, careful data preparation and computing resources. That is a lesson the consortium believes is relevant well beyond IDEA4RC, as the European Health Data Space enters into force.



The platform in action
The meeting’s central event was the first opportunity for clinicians themselves to run the complete IDEA4RC workflow using real data contributed by multiple centres. Working through RAVEN, the project’s research interface, participants selected the tumour types and topographies they wanted to study, then built patient cohorts using a natural-language interface, typing a plain description such as “female patients over 65” and letting the tool translate it into the underlying database query. From there, they explored summary statistics, such as the distribution of tumour sites across the participating centres, and ran more advanced analyses through Vantage6, the federated learning platform that lets a statistical model travel to where each hospital’s data already sits, rather than requiring the data itself to move. Today, RAVEN gives authorised researchers access to the data of 7,869 sarcoma patients and 12,416 head and neck cancer patients contributed by six centres.



Building the infrastructure behind it
None of this, the cohort builder, the summary statistics, the federated analyses, would have been possible without first putting the underlying platform in place across the participating clinical centres. Eugenio Gaeta, a software engineer at UPM who led this strand of work, described deployment as a socio-technical process. Each clinical centre arrived with its own data sources, IT capacity, legal constraints and cybersecurity policies. Thus the methodology combined a common federated architecture with an implementation path tailored to each centre, coordinated through shared planning, recurring pilot meetings to track progress, and one-to-one technical follow-up.
Getting to this point required adapting the platform, step by step, to eleven centres with different clinical data systems, IT capacities and national requirements. The consortium describes this as a process that is structured but necessarily flexible, since no two hospitals arrived at the platform by quite the same route.
Among the components of the data space, Gaeta highlighted the data quality tool. According to Gaeta, passing the tool’s check should be regarded as a core readiness criterion for the centres. “In a federated environment, unlike in a centralised one where data are pull together in a single location, researchers cannot inspect data themselves and assess their consistency. So, the quality of the dataset should be checked for in advance to decide whether it is high enough to include it in the analysis”, Gaeta commented.
Extracting structured data from clinical notes
Beyond the platform demonstration, partners spent an afternoon on one of the project’s core research questions: how to extract structured information, the kind a researcher needs for analysis, from free-text clinical notes using large language models. As Annalisa Trama put it in introducing the topic, IDEA4RC was born the moment large language models entered the scene, a disruptive technology that changed the trajectory of the project.
The discussion was based on the knowledge gathered by considering two centres, the National Cancer Institute of Milan (Italy) and the Maria Sklodowska-Curie National Research Institute of Warsaw (Poland), which were able to share with IDEA4RC data scientists the clinicians’ notes and pathology or radiology reports annotated by their own experts.
IDEA4RC data scientists, coordinated by Unai Zulaika a data scientist at the University of Deusto, tested the performance of two models from the MedGEMMA family, one with 4 billion parameters and one with 27 billion parameters. “You can think of the 4-billion model as a junior trainee and the 27-billion one as an experienced senior,” Zulaika said. They asked each model to extract a series of items from the free-text, then compared its answers with the annotations given by the centres’ experts.
The bottom line was a tension between performance and sustainability. Larger models tend to produce more reliable output, but hospitals typically cannot run them: local computing resources are often limited, and commercial cloud-based systems are ruled out because patient data cannot leave the premises. That leaves smaller, open-weight models that can run inside a hospital’s own infrastructure, and that need to be guided carefully to perform well.
The consortium discussed several ways of doing this: pointing a model to the type of document most likely to contain a given piece of information, helping it reconstruct the chronological sequence of a patient’s clinical journey rather than treating each note as an isolated snapshot, and, perhaps most importantly, training it to recognise when information simply isn’t in the text, rather than inferring or fabricating an answer.
That last point emerged as a genuine open problem: telling a wrong answer apart from a fabricated one, and rewarding a model for saying “I don’t know”, is proving to be one of the harder challenges in applying language models to clinical text.
What comes next
The final part of the meeting turned to what comes next. With the project’s funding ending in August 2026, the consortium agreed on the remaining steps toward the technical report and the review meeting expected in the autumn. Just as importantly, partners discussed how the infrastructure could continue to operate afterwards. Several clinical partners expressed their commitment to maintaining and using the platform, making it available for further development and evaluation. Some partners are also looking at how individual components of the platform, besides the system as a whole, might be of use to other rare cancer registries and research initiatives.
As IDEA4RC approaches its official conclusion, the message from Madrid was consistent: the project’s formal end date is close, but the infrastructure, the working relationships across the consortium partners, and the research questions it has opened are only getting started.

