The history of openEHR: before the revolution.

Sam Heard, Medical Practitioner, Chair of the openEHR Foundation

I am an Australian doctor who completed my medical degree and intern year in Adelaide, South Australia, before heading off to London on a motor bike through South East Asia and India with my girlfriend, who is now my wife. I was keen to be a good general practitioner (GP) and successfully applied for a three-year GP training position, based at St Bartholomew’s Hospital in Central London.

Early Days

My interest in health computing began in 1984. I had completed my GP training and became an Inner City Lecturer at the Department of General Practice, St Bartholomew’s Hospital Medical College, London University. The department had the first computerised general practice in the UK; a large machine in its own airconditioned room in the Wells Street practice in Hackney. My role as lecturer was both clinical and academic. I became the partner of an elderly solo practitioner in order to modernise the practice. It was in a tiny space in Shoreditch, providing comprehensive health care to 3500 patients  24/7 and, initially, had no receptionist. I saw 85 patients on the first day when my older partner became jaundiced and went home – never to return. The university wanted me to implement their second version computing system using concurrent CPM and I agreed. I wanted anything that might help the workload and organisation of care for this large group of patients.

The computer programmer from the Dept of General Practice sadly developed a brain tumour and had to stop working, leaving an incomprehensible program which did not work. This experience sowed a seed and I was really excited to help develop and use computers for primary care. He was using Q-Basic which I decided was too difficult for me to get my head around given my clinical responsibilities. In 1985 my father sent me his IBM Personal Computer (PC) from Australia as he had retired. I added a 10 MB hard drive! Borland had just published Paradox, and Tony Randall in Oxford encouraged me – he was using D-Base. These were the first PC-based databases with multiuser capability. No Windows, just MS-DOS  from Bill Gates. PC networks were just becoming stable. I used Paradox to create a patient management and follow up system; it was all a whole lot easier than using Q-Basic. Over 18 months I had things moving and surpassed the functionality of the original program.

I established a GP IT cooperative with the support of the Family Practice Committee (NHS payers) in 1987. Normalising addresses (autofill from a defined set) – now common – had not been done before and it had wonderful results, allowing us to achieve the highest immunisation and cervical screening rates in London in a very low socio-economic setting. About 80 GPs from East London joined the Coop and together we improved the care of the patients substantially.

More than Rows and Columns: GEHR

Very early in my days of building this simple system (usually done at home after our four children had gone to bed and as long as there were no house calls) I decided that files consisting of rows and columns were not suitable for health records (think Excel which did not exist at the time). MUMPS had been created by Octo Barnett in the 1960s and brought together the related rows and columns (a real speed advantage) whereas Paradox, the modern alternative, had all the data relating to one patient separated in many different tables. To add any new data, I had to change the database, often needing to add tables. This was hard to manage across many practices. There had to be a better way.

Alain Maskens from Belgium came to see me in 1989 promoting his software, HealthOne. This was a real breakthrough from my perspective. Gone were the rows and columns, and the data to be collected could be specified through the user interface! Alain was ambitious for his software and had built a consortium to make a bid for a large European health record project. He invited me to Brussels to see what could be put together. I had met Professor David Ingram, the first professor of health informatics in the UK, at my University and I was advised by Lesley Southgate, my then professor of General Practice, to go and see him. I invited David to come to Brussels – he agreed but was very clear he would not be getting involved. By the end of the meeting David was the leader of the project and put in a huge effort with Alain to successfully stand up what became known as the Good European Health Record (GEHR) project. In 1991 the GEHR project grant was the largest that St Bartholomew’s Medical College had ever received.

Joe Milan, a close associate and friend of David Ingram, joined the project. He was leading the UK health computing stakes with his MUMPS based system at the Royal Marsden Hospital and added weight to what was a colourful and diverse group of people from many European countries. I learned a great deal during the three years of this project despite having to relocate with my family to Australia in late 1992 and spending the last two years of the project flying back to London from Darwin.

Dipak Kalra (a GP using ParaDoc) and David Lloyd, prominent members of David Ingram’s team at the time, took major roles in the project. It was during this time that I met Thomas Beale, an Australian computer scientist and engineer, who was employed by Joe Milan and was seconded to the GEHR project. Thomas and I spent many evenings proposing ever more revolutionary solutions to what was now recognised as a ‘wicked problem’ but made little progress. There were very significant outcomes from the GEHR project, which contributed to what would later become openEHR. The first was to view the health record as a transactional artefact, enabling clear data provenance and roll-back to any point in time. The second was that it should work best for practicing clinicians. The final GEHR output was not suitable for implementation but was in the form of a technology independent specification, rather than software; a novel approach in the early 1990s.

The Birth of openEHR in Australia

Thomas Beale came back to Australia after the GEHR project finished in 1994 and dedicated one day per week to continue working with me on what we had come to see as a lifelong health record specification, with multilingual capabilities. I was working in the remote Northern Territory and leading the general practice educational unit of the new clinical school. This gave me time to collaborate with vendors and interested practitioners and I soon gained a reputation for innovation in this area.

I was asked to write a paper on a future view of health records for Australia by the new Government agency (which became NEHTA), giving me more time to dig into detail. Thomas was working in finance at the cutting edge of object-oriented programming with a language called Eiffel. Together we refined the structure and organisation of a lifelong health record, with some aspects to assist human navigation and organisation, others to enable safe automatic processing. But we could not solve the issue of the evolving clinical demand for new or extended data.

Working closely with Thomas, we learnt each other’s domains of knowledge to a considerable extent. Thomas was often assumed to be a clinician in meetings. After nearly 3 years of working together, including many, many hours on the phone and some high-level workshops in Darwin, Thomas and I decided in 1997 to spend two weeks dedicated to solving the problem of how to specify the clinical data. We agreed that if we failed we would give up. It had begun to feel a little futile working so far from other expert input. The fortnight was not very productive and at lunch on the last day – with just a few hours left – Thomas threw up his hands and said in frustration, “This would be easy if we just left out the clinical data.” We look at each other and thought about how we could do that. Within two hours we had the gist of openEHR down – with archetypes as the enabler. The EHR as a service was just a set of building blocks – datatypes, ways of aggregating, headings etc – some to assist safe processing and some to assist people using the record to understand the contents. The archetypes were the specification of how to represent clinical data using these building blocks. There were more features that quickly fell into place as we got into our stride over the next few months – safe display of any entries, roll back to a point in time, linking to multiple terminologies, a wealth of datatypes.

1998 Key ring from a Spanish Restaurant in Sydney the day Thomas and I created Ocean Informatics with David Rowed and Peter Schloefel after their IBM consultancy came across our work.

Thomas took on the enormous work of creating archetype definition language (ADL) – which was the first part of openEHR to be accepted by ISO about a decade lager. He ensured correctness of these archetype definitions and documented the reference model in great detail. The reference model is a logical model of what the computers process – in whatever language and using whatever data store is fashionable at the time. Refining this specification was slow, despite the increasing assistance from others, particularly David Ingram’s team now at University College London (UCL).

By 1999 I had completed programming the first archetype editor which enabled display of the archetypes – a giant leap to gaining acceptance and led to invitations to HL7 and CEN meetings. In 2006 Hugh Leslie and I built the template designer. I also built the first openEHR kernel to write and read data, with validation based on archetypes. I am proud to have contributed as a clinician to something that is designed to change health care fundamentally.

A CD-ROM posted by Thomas to me on the 27.04.2000 to ensure Ocean could prove to be the source of the (as yet unnamed) openEHR specifications – it remains unopened!

The Foundation and Beyond

In 2005 Sebastian Garde was living in Australia and doing a PhD with Evelyn Hovenga and he took on the web-based archetype management which became the Clinical Knowledge manager. He now lives back in Germany and his contribution has been extraordinary. Chunlan Ma completed her PhD on openEHR and by 2007 had created the OPT format as well as the Archetype Query Language. This revolutionised the power of openEHR as a solution for health information. Heather Leslie and Ian McNicoll (who had completed his Masters at UCL) began building more and more archetypes and have both made enormous contributions. In 2010 Heath Frankel and Chunlan Ma built the first openEHR system to manage the health records of remote, mobile Aboriginal people in  the Northern Territory. This openEHR project received recognition from the WHO. Heath Frankel went on to assist Better to get working with openEHR. Tomaz Gornik, the CEO of Better, has been an extraordinary commercial enabler and promoter of openEHR, and taken Better to be a recognised major international health system provider. He and his team have taken tooling to a new level.

David Ingram supported and encouraged our work and kept supporting Thomas and I during our very frequent trips to Europe and the US where HL7 was involving us in the CDA work but not interested in our approach. In 2005 we deemed it appropriate to publish the specification, David Ingram came up with the name and reserved the openEHR.org URL. University College London supported David in creating the openEHR Foundation which became the home of the specification. He organised the legal framework and other steps to enable the legal transfer from Ocean Informatics and ensure openEHR could move forward safely. That was all more than 15 years ago.

We have weathered incredulity over not specifying the storage format and about the maximal data models. These are now seen as key elements of a specification for a lifelong health record. We faced down Microsoft, Google and Apple’s efforts to create a personal health record. That was nerve wracking to say the least. And all the while David Ingram kept his hand on the helm and made the community something strong and resilient, long before it had any need to be. I am excited to see the new minds and new companies taking openEHR forward. There are still many familiar faces, individuals and teams, particularly in Europe, who have remained committed and passionate for many  years and have contributed considerably over that time.

Importantly, openEHR grew from a very deep collaboration between a computer scientist/engineer and a doctor – at a time when these professions were starting to have mutual concerns. Thomas still works as an engineer and computer scientist in the USA and Europe, and I have until recently been the medical director of a large Indigenous community controlled health service in Australia.

Finally my wife, Merridy Pitcher, has had a lot to do with the success of openEHR – from supporting all my activities, coaching on how to approach difficulties, co-writing documents, managing Ocean Informatics which was our early vehicle for supporting the work and never giving up on the vision of a lifelong health record that could make a difference.

Appendix

My CV Entries relating to openEHR

Appointments

1983-1992        Lecturer, Department of General Practice, Medical College of St. Bartholomew’s Hospital.

1984-1992        Principal in General Practice, The Lawson Practice, St. Leonards, Nuttall St., London N1 5LZ.

1989-1992        Adviser to the City & East London Family Health Authority on computerisation.

1989-1992        Director, Computer Development Unit, Department of General Practice, St. Bartholomew’s Medical College.

1992                 Clinical co-ordinator, European Health Record Project, AIM A2014, European Commission.

1993-1995        Research Fellow, St. Bartholomew’s Medical College, London.

1994-1997        Member of Wellcome Trust Tropical Medicine Resource, World Advisory Group

1996-                Director, Ocean Informatics

1997-1999        Program leader, Communications and Information, Cooperative Research Centre in Aboriginal and Tropical Health.

1998-1999        Member, Electronic Health Record Architecture Working Party, GPCG

1998-                Foundation member, openEHR Foundation

2002-2006        Chair, Standards Australia IT14.9.2 Working Group – EHR

  • Co-Chair, HL7 EHR Special Interest Group and Australian delegate to HL7 International
  • Australian delegate to Centre de European Normalisation (CEN)

2002-2005        Co-Chair HL7 EHR Technical Committee, Aus delegate to HL7

2002-                Director, openEHR Foundation (Ocean Informatics and University College London)

2002-2016        Honorary Senior Research Fellow, University College London

2004-2007        Adjunct Professor, Health Informatics, Central Queensland University, Queensland Australia.

2005-2008        Australian representative at CEN (European Standards)

2005-2007        Vice President, Foundation Fellow, Australian College of Health Informatics, now Australian Institute of Digital Health

2006-2010        Co-chair Standards Aus IT-14-9, Electronic Health Records

2007-2014        CEO, Ocean Informatics Pty Ltd

2011-2014        Chair, openEHR Foundation, London  UK

2019-2023        Director, Australian Digital Health Agency

2023-                Chair, openEHR Foundation

Research grants and competitive allocations

1990                 Computer Development, Family Health Authority £60,000, St Bartholomew’s Medical College

1991-1993        Medical Record structure, European Community £1.9 million, with Prof. David Ingram, Medical Informatics 1993 Healthy Heart Story, National Heart Foundation

2000                 The General Practice Computing Group (GPCG) Trial of the Good Electronic Health Record $230,000

2001                 GPCG Trial of conversion of legacy data with SA Health Commission, DSTC $100,000

2001                 GPCG Trial of Diabetic messaging between General Practice and Specialist Clinics with MCA, $70,000

2002                 GPCG Trail of conversion of GP data to the format of the Good Electronic Health Record $90,000

2002                 HealthConnect openEHR software trial design with the Distributed Systems Technology Centre $670,000 Phase 1

2004                 GPCG openEHR Archetype Tool Development Project developing software tools and documentation to support the uptake of archetypes. $70,000

2004                 ITOL grant to establish an EHR repository for the New South Wales Cancer Registry.  $30,000

2005                 GPCG Archetypes for Communication project. $105,000

Publications and Significant Products

1989                 Heard S.”ParaDoc, A health information system for inner city general practitioners”. This system was written in Paradox, a relational database. Innovations centred on the normalisation of practice data to the Post Office address database allowing strict knowledge of households. Tracing people for follow up or screening was greatly facilitated.

1991                 Heard S. “A standard medical record for Europe – a view from primary care” Abstract of presentation. 40th workshop in Clinical Decision systems, Royal College of Physicians, London.

1992                 Heard S. “The Good European Health Record” Abstract of Presentation. WONCA, Vancouver, Canada.

1992                 Heard S. “Clinical Requirements for a comprehensive Electronic Health Record.” Advanced Informatics in Medicine programme, Directorate General XIII, European Commission.

1993                 “Functional Specification for an Electronic Health Record”. Advanced Informatics in Medicine programme, Directorate General XIII, European Commission.

1993                 “Ethical and legal requirements for a European Electronic Health Record”. Advanced Informatics in Medicine programme, Directorate General XIII, European Commission. Heard S, Doyle L.

1995                 “The Good European Health Record” Final Report. Advanced Informatics in Medicine programme, Directorate General XIII, European Commission.

1996                 Heard S. “Support for remote health practitioners with Information Technology” Proceedings of the Joint Congress of the Australian College of Health Service Executives and the Royal Australian College of Medical Administrators, Darwin August 1996.

1997                 Heard S, Doyal L. “The importance of moral and legal regulation of the Electronic Patient Record.” British Journal of Healthcare Computing and Information Management. March 1997;14(2):26-28

1998                 Heard S. “The Good Electronic Health Record” Proceedings of the 1998 General Practice Evaluation Program Conference, Sydney, New South Wales. National Information Service.

1999                 Heard S. “The Good Electronic Health Record” Proceedings of the 1999 HISA Conference (HIC), Hobart, Tasmania. Health Informatics Society of Australia.

2000                 Heard S. “The Good Electronic Health Record.” British Journal of Health Care Computing and Information Management. February 2000.

2000                 Heard S, Grivel A, Schloeffel P, Doust J. “The benefits and difficulties of introducing a national approach to electronic health records in Australia” in “A Health Information Network for Australia” National Electronic Health Records Taskforce. Department of Health and Aged Care, Commonwealth of Australia. July 2000.

2000                 Heard S. “The GPCG Trial of the Good Electronic Health Record” Proceedings of the Health Informatics Conference, Health Informatics Society of Australia. August 2000.

2000                 Heard S. “Electronic Health Records”, The Australian Health Forum 2000;3:24-25.

2002                 Bird L, Goodchild A, Heard S. “Importing clinical data into electronic health records – lessons learnt from the first Australian GEHR trials.” Proceedings of the 2002 HISA Conference, Melbourne, Victoria. Health Informatics Society of Australia.

2003                 Heard S, Bird L, Warren J. “Editorial” and Guest Editors. Journal of Research and Practice in Information Technology Vol 35, 2. May 2003.

2003                 Barretto S, Warren J, Goodchild A, Bird L, Heard S, Stumptner M. “Linking Guidelines to Electronic Health Record Design for Improved Chronic Disease Management” Proceedings American Medical Informatics Association, November 2003.

2004                 Heard S, Fischetti L, Dickinson G. “HL7 EHR Systems Functional Model and Standard” Health Level Seven, http://www.hl7.org.

2005                 Heard S. openEHR Architype Editor 1.0 (Visual Basic), Ocean Informatics.

2005                 Hovenga E, Garde S, Heard S. “Nursing constraint models for electronic health records: a vision for domain knowledge governance”. Int J Med Inf 2005;74: 886-898.

2006                 Heard S. “Electronic Health Records” Chapter in Conrick M Ed. “Health Informatics: transforming healthcare with technology” Thompson Social Science Press, Australia 2006.

2006                 Heard S, Leslie Hugh. openEHR Template Designer (C#) Ocean Informatics.

2006                 Schuler T, Garde S, Heard S, Beale T. “Towards Automatically Generating Graphical User Interfaces from openEHR Archetypes” Proceedings MIE, Maastricht August 2006.

2007                 Garde S, Hovenga E, Heard S. “Towards Semantic Interoperability for Electronic Health Records: Domain Knowledge Governance for openEHR Archetypes”. Methods of Information in Medicine. 46(3): 332–343.

2007                 Garde S, Heard S, Leslie Heather, McNicol I. The openEHR Clinical Knowledge Manager, Ocean Informatics.

2011                 Frankel H, Ma C, Heard S. The Northern Territory of Australia “My eHealth Record” shared health record based on openEHR specifications, Ocean Informatics.

Presentations

1995                 Heard S. “Issues of ownership”- Data ownership and computerised health information. Central Australian Remote Practitioners Association Conference, Alice Springs, May 1995 .

1996                 Heard S. “The paperless medical record” IT Week, Queen Elizabeth Hospital, Adelaide, South Australia

1997                 Heard S. “The History of Computers in Medicine.” Fifth Biennial Conference of the Australian Society of the History of Medicine. Darwin, Northern Territory.

1998                 Workshop at “Working together for an Electronic Patient Record”, the National Meeting Health Informatics Society of Australia, Brisbane.

1998                 Heard S. “What sort of Electronic Health Record” Presentation at the National Scientific Meeting of the Royal Australian College of General Practitioners, Melbourne.

2000                 Heard S. “The Good Electronic Health Record” RACGP Computer Conference, Sydney, February 2000.

2001                 Heard S. “The Good Electronic Health Record” At Integrating Technology in Health Conference, Sheraton Towers Hotel, Melbourne, February 2001. Invited speaker.

2001                 Heard, S. “Electronic Health Records” Conference of the National Demonstrator Hospital Projects Phase 3, Sydney Convention Centre, October 2001. Invited Speaker

2001                 Heard, S. “Information Technology – Master or Slave”. Royal Australian College of General Practitioners 11th Computing Conference, Melbourne August 2001. Keynote Speaker.

2002                 Heard, S. “The EHR and the future: a clinician’s view.” EHR Clinical Aspects and Standards Workshop, Central Queensland University, Yepoon, August 2002. Invited Speaker.

2002                 Heard, S. “The EHR and HL7.” HL7 International Affiliates meeting, Melbourne August        Invited speaker.

2003                 Heard, S. “openEHR and International Developments” The Electronic Health Record, IIR Conferences, Sydney Australia, June 2003. Invited Speaker.

2003                 Heard S, Beale T, Schloeffel P. “openEHR Tutorial” Proceedings, of the 2003 HISA Conference and RACGP Computing conference (HIC) Sydney Australia, August 2003.

2004                 Heard S. “Archetypes – a user perspective” Keynote address at HL7 Australia National Meeting, Melbourne, Australia 2004.

2004                 Heard S, Elkin P. “The relationship between archetypes, templates and terminology”. HL7 Australia National Meeting, Melbourne, Australia 2004.

2005                 Heard S, Shabo A. “Lifelong health records: an investigation of the role of independent health record banks” HIC, Melbourne 2005

2005                 Heard S. “openEHR – recent developments.” Keynote address, HL7 Australia National Meeting, Sydney November 2005.

2006                 Heard S, Chen R “openEHR, a health computing platform for the future”. Workshop and address to the China Healthcare Information Network Conference (CHIMA, CHITA) Xi’an, China May 2006

2006                 Heard S “openEHR: the first formal specification for sharing and communicating electronic health records”. Key note address, HINZ Conference, Auckland, New Zealand, July 2006

2006                 Heard S, Beale T. “openEHR, a health computing platform for the future”. Workshop, Medical Informatics Europe Conference, Maastricht, Netherlands, August 2006.

2006                 Heard S. “Standards and Electronic Health Records” The World of Health IT Conference, Geneva, Switzerland, October 2006.

2006                 Heard S. “openEHR and semantic interoperability” Health Informatics Conference, Antalya, Turkey, November 2006

The Search for Logical Models and the place of the openEHR Reference Model

Sam Heard, April 2026 PDF download

PDF download?

Table of Contents

The search for a comprehensive set of logical models for health care has a long history. My experience began as a key investigator in the Good European Health Record project in the early 90s. It became clear during this project that 1) an all-inclusive single model of health data was not achievable, and 2) the data specifications had to meet the needs of the people who wanted to use the data.

I have always been interested in a single life-long health record that could grow over time and be available throughout a person’s life. Clearly it would not be appropriate to base this on a particular technology, terminology or other feature that is bound to evolve over time. This is considered to be the role of a logical model. Let us consider then what logical models are and why people are interested in them.

What is a logical model?

A logical model is a structured representation of data, showing entities, attributes, and relationships, independent of any specific database technology. It acts as a blueprint that defines how data is organized conceptually and ensures consistency, integrity, and clarity before moving to physical implementation.

Key Characteristics of a Logical Model

  • Technology-independent: Unlike physical models, logical models don’t depend on a particular database system.
  • Entities: Represent real-world objects or concepts (e.g., Customer, Product, Order).
  • Attributes: Describe properties of entities (e.g., Customer Name, Order Date).
  • Relationships: Define how entities connect (e.g., Customer places Order).
  • Normalization: Organizes data to reduce redundancy and improve integrity.
  • Semantic layer: Provides a business-friendly view of data, bridging raw data and meaningful insights.

Why Logical Models Matter

Standards bodies like the Centre for European Standards (CEN) have been creating logical models for more than 30 years in an attempt to standardise health information.

Logical models potentially offer:

  • Blueprint for databases: Guides the transition from conceptual ideas to physical database design.
  • Business alignment: Ensures data structures reflect business rules and requirements.
  • Data integrity: Prevents anomalies (like duplication or inconsistent updates).
  •         Scalability: Supports analytical applications by offering a consistent framework across systems.

Table 1: Comparison of data models

Model typeFocusExample use caseDependency on technology
ConceptualHigh-level business conceptsDefine what data exists (e.g., “Students enroll in Courses”)No
LogicalDetailed structure of entities, attributes, relationshipsDefine how data relates (e.g., “Student has StudentID, Course has CourseCode”)No
PhysicalActual implementation in a DBMSCreate tables, indexes, constraints in SQL Server, Oracle, etc.Yes

A Simple Example

Imagine designing a IT system for a university, using an SQL database. The different models developed might include:

  • Conceptual model: Students, Courses, Professors.
  • Logical model:
    • Entity: Student → Attributes: StudentID (PK), Name, Email
    • Entity: Course → Attributes: CourseCode (PK), Title, Credits
    • Relationship: Student enrolls in Course (many-to-many)
  • Physical model: Actual SQL tables with primary/foreign keys, indexes, and constraints.

 In short, a logical model is the bridge between abstract business concepts and concrete database implementation, ensuring clarity, consistency, and scalability.

 openEHR Reference Model and Logical Models

A design philosophy of openEHR is to provide a single logical model. The openEHR reference model can be frustrating for developers who want to see how to implement the specifications. What is the role of archetypes? The combination of a logical reference model (to ensure consistent representation of recurrent concepts) and archetypes (to define how to represent ever evolving clinical data using that reference model) is considered one of the most elegant aspects of the design.

Let’s consider how the reference model and archetypes relate to the idea of logical models.

First, the Reference Model (RM) defines the core entities, attributes, and relationships (e.g., Composition, Entry, AuditInfo).[1] This highly specified model captures data and relationships that:

  • Ensure data provenance, integrity, and consistency across implementations; and
  • Is technology-agnostic: it doesn’t care if the backend is relational, document-based, or graph-oriented.

This is essentially the logical model layer, specifying what data is and how it relates, independent of physical storage.

Second, the Archetypes which are built on top of the RM and constrain and specialize the generic structures into clinical concepts (e.g., “Blood Pressure Measurement,” “Medication Order”). Archetypes are:

  • Reusable, shareable definitions that map clinical semantics onto the logical backbone; and
  •  Acting as domain-specific logical models, expressing how to use the RM entities to represent real-world healthcare data, and requiring no change to the underlying database implementation.

How do these two models relate to Logical Models?

  • In database design, a logical model defines entities, attributes, and relationships without tying them to a physical schema.
  • In openEHR:
    • The Reference Model is the foundational logical model — the abstract schema of health record data. This provides all the entities and attributes required to ensure the integrity of a life-long health record.
    • Archetypes are applied logical models — they specify how to represent particular clinical concepts using the RM, and can be used to capture, query and validate data. These provide the changing, context specific, evolving and often critical data specifications for personal health care – at times fractal in nature.
  • Both layers remain agnostic to technology and DBMS, just like a logical model in traditional data modeling.

Table 2: Analogous equivalence in standard data modelling

Layer in openEHREquivalent in Data ModellingRole
Reference ModelLogical Data ModelDefines abstract entities, attributes, relationships necessary for data integrity and provenance of a life-long health record
ArchetypesDomain-specific constraints expressions independent of the RM implementationConstrains the logical RM model to provide the diverse and ever evolving clinical concepts necessary for health care
Physical SchemaPhysical data modelTables, indexes, Json docs of the reference model

There are a number of features of this two-level modelling that are a good fit for personal health data, which is almost infinitely complex and certainly evolving in nature. These include as key features:

  • Interoperability: Logical models (like openEHR RM) ensure that data can be exchanged across systems without loss of meaning.
  • Flexibility: Archetypes allow clinical communities to define semantics without changing the underlying logical backbone.
  • Future-proofing: Because both RM and archetypes are DB-agnostic, they can be implemented in SQL, NoSQL, or even distributed ledger systems without redesign.

In summary, openEHR’s Reference Model provides a generic, technology-independent information model for representing health record data. It defines the structural and lifecycle semantics of clinical information but does not represent specific clinical concepts.

Archetypes specialise the Reference Model by expressing domain-specific constraints that define clinical concepts such as observations, diagnoses, or procedures. Terminology bindings within archetypes further define the meaning of coded data.

In this openEHR architecture, the Reference Model enables structural interoperability between systems, while archetypes and terminology bindings enable semantic interoperability of clinical data.

A detailed description of the openEHR RM as a logical model

Introduction

The openEHR Reference Model (RM) is a foundational, formal information model that defines the structure, semantics, and behaviour of electronic health record (EHR) data. It is designed to enable semantic interoperability, robust data provenance, and clinical integrity across healthcare systems and over time. The RM achieves this by providing a stable, extensible set of information models and data types that serve as the backbone for all openEHR-based systems. Through its integration with archetypes and templates, the RM supports the consistent, high-fidelity representation of clinical data, facilitating safe data exchange, querying, and long-term record-keeping.

This section provides a comprehensive overview of the openEHR RM, detailing its key entities, attributes, and packages. It explains the role and utility of each major component—such as Composition, Entry, Section, Cluster, Element, and others—in the context of a longitudinal EHR. There is an emphasis on how the RM supports data provenance, clinical integrity, and semantic interoperability, and how it interacts with archetypes and templates to enable consistent clinical data representation across systems.

Overview of the openEHR Reference Model

Purpose and Scope

The openEHR RM defines a logical, interoperable architecture for EHRs. Its primary goals are to:

  • Enable consistent representation of clinical data across systems and over time.
  • Support semantic interoperability by providing a stable, well-defined set of structures and data types.
  • Ensure data provenance and clinical integrity through robust versioning, audit trails, and validation mechanisms.
  • Facilitate integration with archetypes and templates, allowing domain experts to define and constrain clinical content without altering the underlying technical model.

Package Structure

The openEHR RM documentation is organized into several interrelated packages, each responsible for a specific aspect of the EHR:

  • ehr: Defines the top-level EHR structure, including access control, status, compositions, folders, and contributions.
  • composition: Specifies the COMPOSITION class, the primary container for clinical content.
  • content: Contains navigation (SECTION) and entry (ENTRY and its subtypes) packages.
  • data_structures: Provides generic data structures such as ITEM_TREE, ITEM_LIST, ITEM_TABLE, and HISTORY.
  • data_types: Defines core data types (DV_TEXT, DV_CODED_TEXT, DV_QUANTITY, etc.).
  • common: Contains shared concepts like LOCATABLE, versioning, and participation.
  • support: Provides supporting types such as identifiers and URIs.
  • ehr_extract: Defines the EHR Extract model for data exchange.
  • integration: Supports integration with legacy and external systems.
  • (demographic: Models parties, roles, and demographic entities but is not included in this document as this package is not required for openEHR health record implementations).

Core RM Entities: Structure and Longitudinal Utility

The following Entities are the major components and provide the structure and organisation of health related data, which itself conforms to archetypes.

EHR

The EHR object is the root container for all health record content for a patient. It references all compositions, folders, status, access control, and contributions. The EHR object supports longitudinal record-keeping by maintaining a complete, versioned history of all clinical and administrative data for a subject. It is the central access point for querying, auditing, and managing the patient’s health information.

EHR-Level Structures: EHR_STATUS, EHR_ACCESS, Folder, Director

  • EHR_STATUS: Contains metadata about the EHR, such as subject, queryability, and modifiability.
  • EHR_ACCESS: Defines access control policies and settings.
  • FOLDER: Organizes compositions into a directory structure, supporting thematic or temporal grouping.
  • Directory: Optional hierarchical folder structure for advanced organization.

These structures support privacy, security, and efficient retrieval of clinical content.

Composition

A COMPOSITION represents a unit of information committed to the EHR, typically corresponding to a clinical document or event (e.g., a discharge summary, progress note, or lab report). Each composition is versioned, allowing for full audit trails and rollback. Compositions are categorized as:

  • Persistent: Represent long-term patient state (e.g., problem list, medication list).
  • Event: Capture time-specific clinical encounters.
  • Episodic: Support episode-based care (e.g., pregnancy episode).

Compositions encapsulate context (EVENT_CONTEXT), content (Sections and Entries), and metadata (language, territory, category, composer).

Event Context and Clinical Context

The EVENT_CONTEXT class captures the context of a clinical event, generally a composition, including:

•             start_time, end_time: Time interval of the event.

•             location: Physical location.

•             setting: Clinical setting (coded).

•             health_care_facility: Facility where the event occurred.

•             participations: List of parties involved.

EVENT_CONTEXT enables understanding of the circumstances surrounding clinical data, supporting provenance and contextualization.

Section

A SECTION is an organizational structure within a composition, used to group entries under logical headings (e.g., “History”, “Examination”, “Medications”). Sections can be nested, allowing for hierarchical organization of clinical content. While sections themselves do not carry clinical meaning, they provide structure for consistent documentation and facilitate navigation and querying over time.

Entry (and subtypes)

The ENTRY class is the abstract superclass for all clinical statements in the EHR. It is further specialized into several subtypes, each representing a distinct type of clinical information:

  • OBSERVATION: Records measurements, findings, and test results and data with specific timing. It utilises a time-series structure (see History below) to support repeated measurements or data streaming (e.g., vital signs, lab results).
  • EVALUATION: Captures clinical assessments, diagnoses, and opinions. Represents the outcome of clinical reasoning. The timing is drawn from the date in the archetyped data (e.g. date of onset) or from the composition.
  • INSTRUCTION: Specifies intended future actions, such as orders, prescriptions, or care plans. Supports workflow and care planning.
  • ACTION: Documents actions that have been performed, which may be in response to instructions (e.g., medication administration, procedures).
  • ADMIN_ENTRY: Records administrative or non-clinical data (e.g., admission details, social services information).

Each ENTRY subtype supports longitudinal tracking of clinical reasoning, interventions, and administrative events, enabling a comprehensive view of patient care over time.

Item Structures: Element, Cluster, Item_Tree, Item_List, Item_Table

Item structures provide the building blocks for organizing clinical data within entries:

  • ELEMENT: The atomic unit of data, holding a single value (e.g., blood pressure reading). Supports null_flavour for missing data semantics.
  • CLUSTER: Groups related elements or nested clusters, enabling modular and reusable data structures (e.g., a device cluster containing make and model elements).
  • ITEM_TREE: Represents hierarchical data, allowing for complex nested structures and is used in archetypes at the root of all data collections, although any structure is allowed in the reference model. The following ITEM_LIST and ITEM_TABLE are not used widely as they limit evolution of the data.
  • ITEM_LIST: Represents ordered lists of elements (e.g., problem lists).
  • ITEM_TABLE: Models tabular data with named columns and rows (e.g., lab result panels).

These structures enable flexible, archetypable organisation of clinical information, supporting both simple and complex data patterns.

History and Time-Series Modelling: History, Event, Point_Event, Interval_Event

The HISTORY structure models time-series data, essential for representing longitudinal observations such as vital signs or glucose tolerance tests. It contains:

  • origin: The starting point of the time series.
  • events: A list of EVENT instances, each representing a data point.

EVENT is the abstract superclass for time-stamped data points, with two main subtypes:

  • POINT_EVENT: Represents an instantaneous observation (e.g., a single blood pressure reading).
  • INTERVAL_EVENT: Represents data aggregated over a time interval (e.g., average blood pressure over 5 minutes), with attributes for width (duration), math_function (e.g., mean, max, min, change etc.), and sample_count.

The HISTORY structure supports both periodic and aperiodic sampling, efficient representation of fine-grained device data, and inclusion of summary data. It enables detailed temporal modeling of clinical phenomena, supporting longitudinal analysis and trend detection.

Data Types in the RM

The RM defines a comprehensive set of data types, all inheriting from the abstract DATA_VALUE class. Key data types include:

  • DV_TEXT: Free text with optional language and informal encoding.
  • DV_CODED_TEXT: Text with associated coded terminology, supporting terminology bindings and semantic interoperability. As this is a subclass of DV_TEXT, it enables any text to be formally coded when appropriate, linking the text and code from the terminology source.
  • DV_QUANTITY: Numeric values with units (e.g., 120 mmHg).
  • DV_DATE_TIME: Date and time values, supporting temporal precision.
  • DV_MULTIMEDIA: Encapsulated multimedia content (e.g., images, audio).
  • DV_BOOLEAN, DV_COUNT, DV_PROPORTION, DV_ORDINAL: Support for boolean, count, ratio, and ordinal data (name: value pair e.g. Apgar scoring).

These data types enable precise, semantically meaningful representation of clinical information, supporting both human readability and machine processability.

Versioning, Change-control, and Audit

The RM implements a robust versioning and change-control mechanism to ensure data provenance, integrity, and traceability:

  • VERSIONED_OBJECT<T>: A container for all versions of a top-level object (e.g., Composition).
  • VERSION<T>: Represents a single version, with audit details and digital signature.
  • CONTRIBUTION: Groups one or more versions committed together, representing a logical change-set.
  • AUDIT_DETAILS: Records metadata about each commit (who, when, what, why).
  • ATTESTATION: Supports legal attestation and digital signing of content.

This mechanism ensures that all changes are indelible (no physical deletion), fully auditable, and reconstructible, supporting medico-legal requirements and forensic examination of past states.

Feeder Audit and Provenance

The FEEDER_AUDIT class captures metadata about the origin of data imported from external systems. It includes:

  • originating_system_audit: Audit details from the source system.
  • feeder_system_audit: Audit details from any intermediate system.
  • original_content: Optional inclusion of the original content.

FEEDER_AUDIT enables traceability, trust, and integration of data from non-openEHR sources, supporting data provenance and duplicate detection.

Locatable, Identifiers, and Paths

The LOCATABLE class is the base for all archetypable objects in the RM. It provides:

  • uid: Unique identifier for the object.
  • archetype_node_id: Identifier linking the object to its archetype definition.
  • name: Human-readable name.
  • archetype_details: Metadata about the archetype and template used.
  • feeder_audit: Provenance information for imported data.

LOCATABLE supports path-based querying using Xpath-compatible syntax such as openEHR’s AQL, enabling precise referencing and navigation of any node within the EHR. This is essential for semantic querying, linking data, data extraction, and interoperability.

Interaction with Archetypes and Templates

openEHR employs a two-level modeling approach involving use of two models:

  1. Reference Model (RM): Provides a stable, technical foundation for data structures and types.
  2. Archetype Model (AM): Allows domain experts to define reusable, constraint-based models (archetypes) for clinical concepts (e.g., blood pressure, lab result).

The actual clinical data is specified as archetypes, which are statements of how to use the reference model to express specific data instances. The archetypes, often expressed as ADL, conform to the archetype model.

Templates combine multiple archetypes to define forms, documents, or messages for specific use cases (e.g., discharge summary, lab report).

Archetype and Template Binding

  • Archetypes constrain RM structures, specifying allowed data elements, value sets, and terminology bindings.
  • Templates further constrain archetypes for particular contexts, setting occurrence constraints, default values, and hiding irrelevant elements.

At runtime, archetype_node_id and archetype_details in LOCATABLE instances link data nodes to their generating archetypes and templates. This enables:

  • Validation: Ensuring data conforms to clinical and technical constraints.
  • Semantic querying: Using archetype paths and codes for precise data retrieval.
  • Consistent modification: Ensuring updates respect original constraints.

Archetypes and templates are stored separately from data, supporting reuse and independent evolution of clinical models and technical infrastructure.

Common features

  • Archetype-root points: ENTRY subtypes (OBSERVATION, EVALUATION, etc.) are typical root points for archetypes.
  • Templates: Used to define structures for common documents (e.g., discharge summaries, lab results, care plans).
    • Example structures: OGTT (oral glucose tolerance test), asthma management plan, multi-drug therapy.

Templates can be viewed and edited using tools such as the openEHR Template Designer, and numerous real-world examples are available in the openEHR Clinical Knowledge Manager (CKM).

Semantic Interoperability

Terminology Binding

The RM and archetypes support semantic interoperability through:

  • DV_CODED_TEXT and CODE_PHRASE: Allowing data values to be coded using standard terminologies (e.g., SNOMED CT, LOINC, ICD).
  • Archetype term definitions: Each archetype may contain an internal terminology for node definitions and value sets which can be directly linked to multiple terminologies through:
    • External bindings: Archetypes can bind internal codes to external terminologies, enabling mapping and querying across systems.

This approach ensures that clinical data retains its meaning across different systems, languages, and contexts, supporting consistent querying, data sharing, and decision support.

Path-Based Querying

Every data point can be accessed directly using path based queries.

  • LOCATABLE paths: providing Xpath-compatible syntax enables precise referencing of any node.
  • Archetype paths: Combine RM attribute names and archetype codes for semantic navigation.
  • AQL (Archetype Query Language): Supports expressive, semantically rich queries over archetyped data.

Semantic interoperability is further enhanced by the use of archetype constraints, which enforce consistent use of codes, value sets, and data structures across systems and allow tight data validation in forms or at the time of committing the data.

Clinical Integrity and Data Quality

The quality of data is increasingly important with more automatic processing of health data through AI and other mechanisms. It is worth summarising what openEHR offers in this domain.

Data Validation and Constraints

  • Archetype constraints: Define permissible values, structures, and terminologies for each data element.
  • Templates: Further constrain archetypes for specific contexts, ensuring only relevant data is captured.
  • Null_flavour: Indicates reasons for missing data (e.g., unknown, not applicable, masked), supporting accurate interpretation and data quality assessment.
  • Data validation: Implementations in all technologies are able to validate data against archetype and template constraints at data entry and modification time.

Attestation and Audit

Clinical accountability and medico-legal provenance is important for longitudinal health records. openEHR provides many features to support this, as well as versioning of all contributions to the health record.

  • ATTESTATION: Supports digital signing and legal attestation of content, ensuring accountability and trust.
  • AUDIT_DETAILS: Records who made changes, when, and why, supporting traceability and medico-legal compliance.

These mechanisms ensure that clinical data is accurate, complete, and trustworthy, supporting safe patient care and regulatory compliance.

Interoperability and Exchange

With the future of health care in mind, openEHR provides a means of extracting all or part of a health record and providing it to a new health record instance, while maintaining the provenance of the data, as well as the state of each health record instance at any point in time. This is considered to be critical for health care professionals.

EHR Extract Model

The EHR Extract model enables standardized exchange of health record content between systems, supporting:

  • Full, simplified, and synchronization extracts: For ad hoc queries, batch transfers, migrations, and system synchronization.
  • Compatibility with external standards: Such as ISO 13606, HL7 CDA, and FHIR.
  • Preservation of versioning and audit trails: Ensuring data integrity and provenance during exchange.

The extract model uses X_VERSIONED_OBJECT and related classes to serialize versioned content for lossless transmission, supporting a wide range of interoperability scenarios.

Integration with Other Standards

  • GENERIC_ENTRY: Supports integration with legacy and non-openEHR systems.
  • Terminology bindings: Enable mapping to external vocabularies and value sets.
  • Canonical and simplified serializations: Support XML, JSON, and other formats for data exchange and integration.

These features ensure that openEHR-based systems can interoperate with a diverse ecosystem of health IT solutions, supporting data liquidity and patient mobility.

Knowledge Engineering and Ontologies

There are many specialists who understand a great deal about the structure and organisation of knowledge, and the meaning of the words we use and the data we collect. Thomas Beale brought an ontological perspective to the design of the openEHR reference model – which represented the consistent features of health data, provenance, timing, language, and many other aspects.

Archetypes., although designed to work with any (and multiple) terminologies, were firmly based on clinical requirements and needed no semantic basis, at least at the outset, as clinicians already shared this very large data space that was common; straddling language, culture, time and space. Organising the archetypes to maximise interoperability and reuse was the challenge and it is a challenge that persists. Some data collections, such as laboratory results, were so well established that following the widespread pattern and using LOINC terms aided implementation.

The great benefits that openEHR offers are consistency of representation provided by the reference model and the community of practice to create the archetypes which together provide the context for the words and phrases in terminologies, and limit the set to a small number in many situations. All this makes meaning manageable; and automatic processing more resilient.

Conclusion

The openEHR Reference Model provides a robust, semantically rich, and longitudinally consistent framework for representing electronic health records. Its modular architecture, comprehensive set of entities and data types, and integration with archetypes and templates enable high-quality, interoperable clinical data across systems and over time. Through its support for versioning, audit, provenance, and semantic interoperability, the RM ensures clinical integrity, data quality, and safe data exchange in diverse healthcare environment.

The RM is explicitly designed to support longitudinal health records, enabling comprehensive tracking of patient data over time:

  • Persistent Compositions: Maintain single sources of truth for long-term patient state (e.g., problem list, medication list).
  • Event Compositions: Capture discrete clinical encounters or events.
  • Episodic Compositions: Support episode-based care (e.g., pregnancy, hospital admission).
  • HISTORY and EVENT structures: Model time-series data, supporting trend analysis and temporal reasoning.
  • Versioning and audit: Enable reconstruction of past informational states, supporting medico-legal requirements and forensic analysis.
  • Folders and directories: Organize data thematically or temporally for efficient retrieval and navigation.

This longitudinal utility is critical for supporting chronic disease management, population health analytics, and continuity of care across organizational boundaries.

By separating the technical foundation (RM) from clinical content modeling (archetypes and templates), openEHR empowers domain experts to define and evolve clinical models independently of technical infrastructure, fostering innovation, adaptability, and future-proof health information systems.

Appendix A: Major Reference Model classes, purpose and use

EntityMain AttributesPurposeLongitudinal Utility
The features of the entire longitudinal EHR and how it is organised
EHRsystem_id, ehr_id, contributions, ehr_status, ehr_access, compositions, foldersTop-level container for a patient’s health recordProvides the longitudinal anchor for all clinical data; ensures continuity across systems and time
EHR statusis_modifiable, is_queryable, subject, other_detailsDefines modifiability and queryability of the EHRControls lifecycle and accessibility of the record over time
EHR Accesssettings, policiesDefines access control policies for the EHRSupports longitudinal governance and security of patient data
Foldername, itemsLogical grouping of CompositionsSupports organization of longitudinal record across episodes or domains
The components of the EHR
Compositionlanguage, territory, category, context, composer, contentTop-level container for clinical content; represents a single clinical encounter or documentAnchors entries to time, place, and author; provides audit trail and context for longitudinal record
Sectionname, further sections or entriesOrganizational grouping within a CompositionSupports hierarchical structuring of content across encounters; aids navigation and retrieval and human readability
Entrylanguage, encoding, subject, provider, other_participations, data (ItemStructure)Abstract superclass for all clinical and administrative statementsProvides consistent framework for observations, evaluations, instructions, and actions across time
-Observationdata (history), state, protocolCaptures measured or observed dataSupports time-series data (e.g., vitals, labs) enabling longitudinal tracking and trend analysis
–Historyorigin (time of first event), events or serial measurements of fixed intervalTime-structured series of data pointsEnables single point in time or longitudinal representation of repeated measures (e.g., lab series). Longitudinal measures include maximum value, minimum value, average, range, or sum.
—Event (abstract)time, data and state informationSingle point in a History series. Events are either point in time or an interval of timeCaptures discrete longitudinal data points within a time series
—Point eventtime, data and state informationRepresents a single point in timeCaptures instantaneous observations
—Interval eventtime, width, math_function, dataRepresents an interval of timeCaptures aggregated or interval-based observations
-Evaluationdata, protocolRepresents clinical interpretation, diagnosis, or assessmentPreserves evolving clinical judgments and assessments over time
-Instructionnarrative, activities, workflow plan, expiry_time, wf_definitionPrescribes actions to be carried outCaptures intended care plans and orders; supports continuity of care across encounters
-Actiondescription, time, instruction details, ism_transition (state transition such as from prescribed to dispensed for a medication)Documents actions performed in response to instructionsProvides record of what was actually done; supports audit and outcome tracking
-Admin Entrydata (item_structure)Records administrative dataSupports non-clinical longitudinal data
Clusteritems, archetype_node_idStructure for nested or complex data – may be archetyped for reuse in other EntriesSupports detailed modeling of panels, imaging, or structured measurements across time
Elementname, valueLeaf node holding a single data valueBasic unit of clinical data; enables fine-grained longitudinal comparison (e.g., systolic BP values)
The features which allow the health record to be coherent, safe and useful
AuditInfocommitter, time_committed, change_typeMetadata for provenance and versioningEnsures traceability, accountability, and integrity across the evolving longitudinal record
Partyidentifiers, name, rolesRepresents people or organizations involvedSupports continuity by linking data to patients, clinicians, institutions across time
Locatablearchetype_node_id, nameBase class for all RM structuresProvides consistent identification and semantic anchoring across longitudinal data and enables a single data point to be queried and linked
Versionuid, lifecycle_state, contributionRepresents a version of an RM objectSupports version control and audit trail across longitudinal record, independent of the system it may be stored within
Contributionuid, versions, auditGroups versions committed togetherEnsures atomic commits and provenance across longitudinal updates
Archetypedarchetype_id, template_idLinks RM structures to archetypes/templatesProvides semantic binding for longitudinal consistency
Linkmeaning, type, targetRepresents relationships between RM objectsSupports semantic connections across longitudinal record
Attestationattester, proof, reason, itemsLegal attestation of contentSupports medico-legal integrity

[1] See table on page 16
ChatGPT was used to summarise the reference model specification and provided the many of the headings in this paper.