IDEA4RC

Intelligent ecosystem to improve
the governance, the sharing,

and the re-use of health data for rare cancers

Four years of IDEA4RC: lessons learned and what comes next

As IDEA4RC reaches the end of its four-year journey, we spoke with project coordinator Annalisa Trama, head of the Epidemiology Unit at the National Cancer Institute of Milan and responsible for the EURACAN registry, the registry of the European Reference Network for rare adult solid cancers. We asked her what the consortium has learned, how her own perspective has changed, and what comes next for rare cancer data in Europe.

Looking back after four years, what would you keep and what would you change in IDEA4RC’s approach to making rare cancer data reusable for research?

The objectives of IDEA4RC remain as relevant today as they were four years ago. Our experience over this period have reinforced my conviction that an infrastructure like the one we have developed is needed. Working with data, and with large volumes of data, remains an urgent need, even though there’s now a widespread narrative that we’re ready to ‘unlock the data’. What I would change is the amount of support provided to clinical centres and, as a consequence, my expectations about how quickly such an infrastructure could be built.

Clinical centres differ greatly in their technical capacity and in their culture around data, and the information they hold is extremely heterogeneous, particularly when you want to reuse it for research. IDEA4RC set out to use Natural Language Processing (NLP) to extract high quality data from clinical records, but clinical records themselves have intrinsic limitations.

They are written by doctors, often under considerable time pressure, and some details about the status of the disease may be omitted because they are not essential for deciding the best treatment plan, or because clinicians can retrieve them later from other documents if needed.

Collaborating with expert centers introduces an additional layer of complexity regarding data quality. Expert centres often see a patient only during one phase of the disease. A sarcoma patient, for example, may come to our institute for surgery and then continue treatment elsewhere, so we may know very little about what happens afterwards. Conversely, patients with head and neck cancer may arrive at an expert centre after a second or third recurrence, several years after their initial treatment elsewhere.

NLP can help us retrieve information that is contained in clinical notes, pathology reports and radiology reports, but it cannot recover information that was never recorded in the first place.

Does this mean that improving data reuse also requires changing how data are collected in the first place?

Absolutely. We need to work both downstream, by developing better tools to extract and harmonise existing information, and upstream, by improving the quality of data when they are first recorded.

But this does not mean asking clinicians to completely change the way they work. Technology should adapt to clinical practice and support it. Doctors cannot be expected to complete an electronic health record as if they were filling in a research database with hundreds of predefined fields. They need to be able to write clinical notes in natural language.

Within IDEA4RC, together with our partner CliniNote, we explored a different approach: supporting clinicians as they write, for example by suggesting more standardised terminology or prompting them to specify information that may not be essential for immediate patient care but could be important for future research. If a clinician writes that a tumour has increased or decreased in size, for instance, the system could prompt them to record the change according to a standard scale or to provide additional details regarding the magnitude of the change.

Could the IDEA4RC approach also serve as a blueprint beyond rare cancers?

I believe so. Because rare cancers are a subset of all cancers, the overall approach, architecture, and technological components developed in IDEA4RC can be adapted and extended to more common cancer types. A critical part of this adaptability is having a flexible yet standardised way to represent cancer data across different types. This need led us to develop the European Cancer Common Data Model (ECCDM), which provides a shared core framework for cancer information that supports both clinical care and research. We are already exploring how the ECCDM can be validated and expanded for use in common cancers such as breast cancer, making the approach scalable beyond rare cancers.

The work on ECCDM, led by HL7 within IDEA4RC, has attracted interest from major EU Cancer Mission projects like CANDLE and UNCAN-Connect, which see its potential as a foundation for wider cancer data initiatives. 

Starting with rare cancers made particular sense for a European project because individual centres rarely see enough patients on their own, making collaboration essential. We leveraged the expertise and networks of EURACAN, the European Reference Network for rare adult solid cancers, to build this infrastructure. The growing interest from other projects demonstrates that much of this work can be generalized across cancer types. 

The methodology and infrastructure we developed could even extend beyond oncology to other disease areas, though different data models would be required for those fields.

Interdisciplinarity has become a buzzword in research, but IDEA4RC had to put it into practice. How did the reality of interdisciplinary collaboration compare with your expectations?

Interdisciplinary collaboration is essential for progress, yet it remains complex and challenging to achieve. While individual experts need to be open to ongoing dialogue and learning across disciplines, no single person can change the culture alone. A broader cultural shift is necessary, one that influences organizational structures, education, and ways of working.

Technical experts and clinicians often work in separate spheres, with clinicians sometimes delegating data tasks due to time constraints or tradition. Changing this mindset towards shared responsibility is crucial for effective collaboration.

This cultural shift should begin early in education, supported by universities, scientific societies, and regulatory bodies, promoting autonomy and innovation. Rather than training all-round experts, the focus should be on developing professionals skilled at working across disciplines.

IDEA4RC succeeded in fostering genuine interdisciplinary collaboration because all partners came to appreciate its importance.

Patient organisations were actively involved in the early design of the data space. What did they bring to the project, and how did their perspective shape it?

At the beginning, it was important to manage expectations and explain that IDEA4RC was building a research infrastructure rather than something that would immediately change an individual patient’s treatment.

Once that was clear, patient representatives made a fundamental contribution, particularly by encouraging us to adopt a more open approach to data governance. Researchers and clinicians tended to be more cautious, partly because they operate under different incentives and obligations. But in rare cancers, openness and sharing are essential if we want to generate new ideas and identify new avenues for treatment.

Patients also showed considerable awareness of issues that may appear highly technical such as data quality.

Naturally, not every patient may have the knowledge or interest to engage in highly technical discussions. But some patient organisations and networks have representatives who are extremely knowledgeable and who contributed very constructively to the project.

Generative AI evolved extremely rapidly over IDEA4RC’s four years. How has the project changed your view of its potential and limitations in healthcare and health research?

In IDEA4RC, we employed open-weight large language models using few-shot prompting, where models are fine-tuned with limited annotated examples from clinical, pathology, and radiology reports. Larger models showed superior performance but require significant computational resources to run locally. Many clinical centres lack this necessary infrastructure.

This underscores the urgent need for coordinated EU and national investments to enhance local computing capacity, a key factor for both NLP advancement and European Health Data Space implementation.

Data access posed another major hurdle: only 3 of 11 centres could share annotated texts for model training and validation, mainly due to inconsistent privacy and security interpretations across institutions and countries. Harmonised data governance is crucial to overcome these barriers and enable scalable NLP applications.

What is the most important lesson from IDEA4RC for other European initiatives working on health-data reuse, and for the wider EU strategy in this area?

One of the most important lessons is that we need to invest to make clinical centers and all data providers technologically equipped to handle data sharing and interoperability. Capacity varies considerably between institutions and, in some cases, remains insufficient. Bringing centres to a comparable level cannot be achieved through individual research projects alone, it requires coordinated national and European investment.

Projects can identify promising approaches. IDEA4RC, for example, has explored federated rather than centralised data analysis. But whatever architecture is chosen, preparing high-quality data remains essential.

The European Health Data Space is moving in this direction because its implementation requires Member States to work together. National governments therefore have a fundamental role in creating the conditions that allow health data to be prepared, shared and reused effectively.

What future do you see for IDEA4RC after the project formally ends?

I think many of the individual components developed within IDEA4RC could continue to be used independently of the platform as a whole.

These include the cohort builder, which makes it easier for researchers to identify patient cohorts; the secure processing environments, or “capsules”, based on the Zero Trust approach and aligned with recommendations emerging from TEHDAS2 and the European Health Data Space; the data-quality tools and the user interface which has been designed to facilitate data exploration and analysis.

We are particularly interested in exploring how some of these tools could be deployed in centres willing to contribute to the EURACAN registry. One of IDEA4RC’s objectives from the outset was to find solutions that reduce the burden of data collection, making it easier for more centres to contribute to the registry.

Individual tools could also be useful to other cancer registries and data initiatives, including projects such as UNCAN-Connect and CANDLE. Throughout IDEA4RC, we have tried to keep the perspective of data holders firmly in mind, and these tools reflect that approach.