IDEA4RC played an instrumental role in launching the development of a European Cancer Common Data Model. The goal is to define a core set of clinical concepts shared across all types of cancer, providing a common foundation for cancer research in Europe. Since its launch in February 2025, the effort has reached a major milestone with the delivery of the conceptual and logical versions of the model, developed with contributions from European and national cancer registries. Major European cancer projects, such as CANDLE and UNCAN, have expressed interest in adopting the model and contributing to its extension. Now formally coordinated by HL7 Europe, the initiative plans to engage with the international cancer data model community to explore the development of a globally harmonized model. At the same time, discussions have begun with several international standardization communities to broaden the scope of the initiative beyond Europe. We asked Roberta Gazzarata of HL7 Europe, who leads the initiative, to share the latest developments and discuss the next steps, including how the work will continue beyond IDEA4RC.
Roberta, where do we stand now?
During two meetings with the HL7 community in February and May 2025, we agreed on a high-level conceptual data model that focused mainly on the first cancer diagnosis, treatment and response. Since then, we have refined the model, thanks to the contributions of clinicians, researchers and data scientists from across Europe. The model has been enriched to describe the evolution of the disease. We addressed questions such as: has the cancer gone into total or partial remission, remained stable, recurred or progressed? Has the cancer spread to other body sites beyond those affected at the first diagnosis?


We then moved to the logical version of the data model. While the conceptual model provides a high‑level, platform‑independent representation of the key cancer‑related concepts and their relationships, the logical model defines a more detailed, yet still platform‑independent, representation of concepts, attributes, data types, cardinalities and relationships.
Throughout the model development process, we also agreed on which would be the preferred sources for each information. For example, the tumor grade and histological behavior are defined by the results of a biopsy. The clinical stage is assessed through one or more imaging exams, while the pathological stage is defined through surgery results. Imaging is instead the primary source to define which body sites are affected by the disease.
Our goal is to identify the information that is essential for the secondary use of data, determine where it should be captured as part of routine patient care, and highlight the elements that are often missing, incomplete, or poorly structured in clinical reports.
Why does better data reuse start with better data collection?
There is a tendency to think that AI can solve everything, but data quality remains a fundamental challenge. If the quality of the data is poor at the start of the process, no algorithm can magically improve it. This is why we hope our initiative will have an impact not only on secondary data use, but also on primary data collection, which is the first step toward ensuring the quality of the final data.
Looking ahead, the model is intended to provide the foundation for the development of standards to ensure that information relevant for secondary use is made available more explicitly and consistently during primary use. On this basis, tools can then be developed to support clinicians in completing medical reports and electronic health records more rapidly and accurately, helping them capture information that is often implicit or not explicitly documented in routine practice.
Who has contributed to this journey so far?
We have worked to involve the major European cancer projects, as well as representatives from European and national cancer registries. The participation of these projects confirmed that this effort addresses a real need: many projects currently have to develop their own data models from scratch, resulting in duplicated work and inefficiencies. On the other hand, the experience of cancer registries was invaluable in defining the core concepts needed to study cancer more broadly. This has allowed us to move beyond the specific focus of IDEA4RC on sarcomas and head and neck cancers and develop a model that is suitable for describing all adult solid tumors. This is an intermediate step toward a model for all types of cancer.
What has been the impact of the initiative on other European projects?
Representatives from CANDLE, a major project within the EU Cancer Mission and Europe’s Beating Cancer Plan, have participated in the working group meetings since the end of 2025 and have now expressed their intention to adopt the model in their own work. They will likely need to extend it to include additional variables, such as specimen data, which we did not consider. These extensions would further enrich the model and make it more broadly applicable. CANDLE’s engagement has already sparked interest in the initiative among other projects, such as UNCAN. The model has been intentionally built to be open and extensible to include aspects that are not yet covered but could be useful for addressing different research questions and use cases.
Do you plan to involve additional stakeholders?
During the HL7 International Working Group Meeting held in Rotterdam at the end of May, we met the community behind mCODE, the US cancer common data standard. We hope to engage with them to learn from their experience and, ideally, converge towards a global cancer data model. While mCODE was conceived for primary use, meaning healthcare delivery, in the US setting, our model is intended to operate at a more abstract, standard-independent level, with the goal of facilitating data exchange. We believe these two perspectives can be complementary and progressively harmonized.
In addition, we have started a dialogue with the OHDSI Oncology Working Group. OHDSI is an interdisciplinary collaborative responsible for developing the OMOP health data standard. OMOP and FHIR—which is developed by HL7—are the two most widely used healthcare data standards in Europe. Over the past year, particularly through the work carried out within IDEA4RC, we have realized that many European centers use both standards. Implementing the logical model in both OMOP and FHIR will facilitate collaboration across centers, which is essential in rare cancers and is becoming increasingly important for common cancers as well.
More recently, we have begun discussions with the openEHR Foundation, which develops an open data standard for electronic health records.
What was the role of IDEA4RC in launching this initiative?
IDEA4RC and its coordinator, Annalisa Trama, were the driving forces behind the initiative. Annalisa recognized early on that it could have an impact beyond IDEA4RC and the field of rare cancers, and she supported its development every step of the way.
What are the next steps?
By the end of IDEA4RC, we plan to submit two papers detailing this activity, also to share the approach we adopted to involve different stakeholders and to incorporate their expertise and needs while at the same time remaining pragmatic. We deliberately selected a minimum set of variables to make the process feasible and complete this first phase, with the understanding that extensions will follow.
In the meantime, we are collecting feedback from the HL7 community. Once the model is finalized, we will map it to both FHIR and OMOP.
For this reason, the activities are expected to continue beyond IDEA4RC, also through the Phoenix initiative, which was created by HL7 Europe to provide continuity to this effort and to act as a connector across European cancer projects and related initiatives.
