There's a statistic circulating in health tech circles: around 80% of the information in an electronic health record sits in unstructured free text: clinic letters, progress notes, radiology reports etc. It's a real problem and it's easy to see why it's become a rallying point. If you want research-grade data at hospital scale you have to be able to interpret what clinicians actually wrote, because that's where most of the clinical detail lives.

Increasingly, the answer being proposed is AI, large language models trained on millions of real patient records, deployed to mine years of unstructured narrative and pull structure out after the fact. There's genuine value in that: some of the most useful clinical detail really does only exist as free text. Unlocking that historical archive is worthwhile and AI is a reasonable tool for the job.

But I think the framing is backwards and solving for the symptom rather than the problem.

Why the Record Ends Up as Narrative

Clinicians don't default to free text because they prefer prose. In many cases they don't have a system that models the needs of the condition being treated. Whilst specialties like Cancer are well served by IT, there is a wide range of conditions that have no specialist system and the only option is recording unstructured clinic notes in the EPR or an Excel file. Every clinic involves time lost to pulling together the latest data for the patient and generating returns for registries is an even bigger pain point.

A Different Starting Point

At ProtoFlex, our view is that the more durable fix sits upstream: give clinical teams a platform flexible enough to model their actual processes and protocols, so that capturing coded, structured data accurately, at the right point in the pathway, is the easy option.

That means:

  • Modelling the pathway, not just the form. A protocol-driven structure that reflects how a condition is actually managed - the decision points, the branching logic, the fields that matter at each step - captures far more of the clinical picture as structured data than a generic template ever will, because it was built around the specific process rather than a lowest-common-denominator form.
  • Context that travels with the data. A value on its own is much less useful than a value with the protocol context attached - which pathway, which step, against which target, compared to what was expected. That context is exactly what turns a data point into something you can trend, audit, and act on, and it's exactly what gets lost when the same information is buried in a paragraph of narrative.
  • Capture at the point of care, not after it. Data entered as part of the clinical workflow, in the moment, is more accurate and more complete than anything reconstructed afterwards - by a person doing a retrospective coding pass, or by a model inferring intent from text written for a different purpose.

None of this eliminates the need for free text entirely, and it shouldn't try to. There will always be observations that don't fit a structured model cleanly, and clinicians should never be blocked from recording something because the form doesn't have a box for it. But if the platform is doing its job, free text becomes the exception that supplements a coded record - not the default that a model has to reconstruct structure from later.

The Case of Trial Matching

Trial matching is a good example of where the difference stops being theoretical. Recruiting patients to a trial means checking a cohort against a set of inclusion and exclusion criteria - diagnosis, staging, prior treatment lines, comorbidities, specific lab values within a specific window. Where that information is captured as structured, coded data, matching a cohort against a protocol is a query: instant, repeatable, and exhaustive across every eligible patient in the system, not just the ones a research nurse happened to think of.

Where the same information exists only as narrative, the alternative is scanning through pages of letters and notes per patient, trying to establish whether the criteria are met from context that may be worded a dozen different ways, may be incomplete, or may not be there at all. An AI tool can be pointed at that narrative and asked to make its best guess at each criterion in turn. That can genuinely help - it beats reading everything by hand - but it is still an interpretive step, run per patient, per criterion, over data that was never captured with that use in mind.

Coded data captured correctly at the point of care doesn't need that interpretive step at all. It's already answerable. That difference compounds every time the same cohort needs checking against a new trial, a new registry, or a new audit - the query gets faster and cheaper each time.

Structured Data That Actually Travels

There's a second half to this argument, and it's the one that matters most for anything spanning multiple organisations, which describes most chronic condition management. Structured data captured accurately at source has a property that AI-extracted data doesn't: it's immediately reusable and searchable, without a second interpretive step standing between the record and whoever needs to rely on it next.

That's where interoperability standards come in. Structured, protocol-aware, coded data - potentially as a FHIR resource or in an openEHR clinical data repository - is data that can move. It can feed the record shared with the next team in the care pathway. It can drive statutory returns without a manual reporting exercise bolted on at the end of the quarter. It can be pulled into a population-level view of how a pathway is performing, or matched against a trial, without anyone needing to re-derive what a given field actually meant.

Data pulled out of narrative by a model, however accurate the extraction, doesn't automatically have that property. It's a snapshot of an inference, generated after the fact, sitting alongside but not necessarily reconciled with the next system down the line. It can absolutely be useful for research and retrospective analysis. But it's not the same thing as a record that was born interoperable, built to be read and searched reliably by the next team, the next system, and the statutory return, without needing to be re-interpreted each time.

Both, But in the Right Order

There's a genuine and growing role for AI in unlocking the unstructured data that already exists - years of legacy notes aren't going to structure themselves retrospectively, and that's exactly the kind of problem AI is well suited to. That work has real value and should continue.

But for clinical teams choosing where to invest next, the order matters. Our approach doesn't burn compute trying to guess at meaning that was never captured cleanly to begin with - it puts a tool in front of clinical teams that lets them check and record accurate, coded data as part of their normal workflow, so it's easily searched and analysed the moment it's entered, not just once someone builds the tooling to go back and interpret it. A platform that sits over an EPR and helps clinicians capture data accurately, in structure, at the moment of care - feeding ongoing care decisions, statutory returns, trial and cohort matching, and reliable sharing with every other team who has a stake in that patient's outcome - does more for the long-term value of that data than a model trained to make sense of it after the fact ever will.

It's also, in plain financial terms, the cheaper path. An org-wide AI overlay has to earn its keep by continually re-evaluating narrative pulled from disparate, inconsistent sources - every letter, every note, every system it's pointed at - which means ongoing inference cost that scales with the volume of data and never really stops, because the underlying record is still unstructured and still has to be re-interpreted the next time someone needs an answer from it. A platform that captures structured, validated data at the point of entry pays that cost once, at the moment of capture, and the record is then simply queryable - no repeated interpretive pass, no compute bill that grows with every new question asked of the same patients. For a health economy already under budget pressure, that's not a minor operational detail; it's a materially different cost profile for getting to the same trustworthy answer.

Get in touch if you want to talk through where this fits for your service.


Marc Warburton is Managing Director at ProtoFlex Software.