← Back to list

From Bills of Mortality to Machine Learning: A History of Medical Coding

How a 17th-century statistician’s problem became the infrastructure of modern healthcare — and why we still haven’t solved it

Julie Ahlfeld RHIT, CCS · 2026-05-05 00:31 · 0 claps · 9.6 min read
#icd-10 #medical-coding #icd-10-cm #history #medical-history
Open on Medium ↗
Wiki topics: ML · Machine Learning HIS · History EDU · Education & Learning 💻 · Programming

From Bills of Mortality to Machine Learning: A History of Medical Coding

How a 17th century statistician’s problem became the infrastructure of modern healthcare — and why we still haven’t solved it

Here’s a puzzle that should bother you more than it does.

Every time you visit a hospital, someone translates the story of your illness into alphanumeric codes. Those codes determine how much your hospital gets paid, how your condition gets counted in public health data, and increasingly, how algorithms will be trained to classify the next patient who looks like you.

It began with a London haberdasher trying to count the dead.

In 1662, John Graunt published a small book analyzing London’s Bills of Mortality. He wanted to find patterns: which diseases killed the most people, which neighborhoods were deadliest. To do that, he needed categories. His categories were crude — deaths grouped under terms like “thrush, convulsions, rickets, teeth and worms, abortives, chrysomes, infants, livergrown, and overlaid.” The data was unreliable, the terminology inconsistent.

But Graunt’s estimate that about 36 percent of children died before age six turned out to be reasonably accurate. The imperfect system still produced usable knowledge.

That’s the foundational insight of medical coding history: a shared imperfect system consistently outperforms a collection of precise but incompatible local ones. This idea, born in 17th-century London, is still being argued about in 21st-century hospital billing departments.

The Man Who Convinced the World to Agree on Something

William Farr didn’t invent disease classification. He did something harder: he convinced people it mattered.

England’s first medical statistician, Farr inherited a classification system that hadn’t kept pace with medical science. His solution was radical in its simplicity: stop optimizing for precision in individual cases, start optimizing for comparability across all of them.

He drew two distinctions that still define the field. First: uniformity matters more than perfection. Second: a nomenclature (an exhaustive list of precise clinical terms) is a fundamentally different tool from a statistical classification (a manageable set of categories for counting and comparison). Confusing the two causes problems. We have been confusing the two ever since.

At the first International Statistical Congress in Brussels in 1853, Farr and Geneva physician Marc d’Espine were jointly tasked with creating a uniform international classification of causes of death. They submitted two incompatible proposals. The Congress adopted a compromise of 139 categories. Farr’s principle — organize by anatomical site — survived as the foundation for every subsequent revision through ICD-10.

Classification systems are not truth-seeking instruments. They are coordination devices. And coordination devices succeed when enough people agree to use them.

The Standard That Spread

Jacques Bertillon didn’t just create a better classification. He created the first one that stuck.

His 1893 classification drew on the Paris system — a synthesis of English, German, and Swiss approaches — and came in three versions: 44 titles, 99 titles, and 161 titles. Different users, different needs, same framework. By 1898, the American Public Health Association had recommended its adoption across Canada, Mexico, and the United States, and proposed something that would shape everything that followed: decennial revision. Update the system every ten years.

That seemingly mundane decision meant the classification would never be finished — always slightly behind medical reality, always requiring negotiation between what clinicians needed and what statisticians could count.

Bertillon drove the revisions of 1900, 1910, and 1920. His death in 1922 created a leadership gap that ultimately led the World Health Organization to assume responsibility after its founding in 1948. The Fifth Revision, under WHO stewardship, was the first called the International Classification of Diseases — ICD.

The Honest Confession Inside ICD-8

Most technical documents don’t admit their own limitations. The 1967 ICD-8 manual did.

Published by WHO based on a 1965 conference, ICD-8 acknowledged explicitly that the classification is built on necessary compromises — between systems organized by etiology, anatomical site, age, and clinical manifestation. No single logical axis could serve all users. It also acknowledged a documentation problem that would only worsen: physicians trained at different schools document conditions using mixed, unstandardized terminology. The classification must absorb all of it. There is no administrative solution — the problem is structural.

ICD-8’s seventeen chapters established the architectural logic ICD-10 still follows. But the revision exposed a problem that would haunt American healthcare for the next decade: the United States couldn’t agree on how to adapt it.

Two competing clinical modifications emerged simultaneously. The Commission on Professional and Hospital Activities published H-ICDA. The US government published ICDA-8. Each was used in roughly half the country’s hospitals — exactly the incompatible local systems Farr had warned against in 1838. The result was data that couldn’t be compared.

This chaotic experience became the defining lesson that shaped everything that followed.

The Committee That Refused to Let History Repeat Itself

Organizations rarely learn from their mistakes. The National Committee on Vital and Health Statistics did.

Established in 1949, the NCVHS had watched two incompatible ICD-8 adaptations divide the country’s hospitals for a decade. When ICD-9 arrived, the Committee actively opposed competing versions before they could take root, recommending that the Secretary of DHHS be personally responsible for a single integrated classification. No more splitting the country.

The result was ICD-9-CM, developed between 1977 and 1979 and implemented January 1, 1979. Congress mandated its use for Medicare billing in 1988. The US version also included something WHO hadn’t produced: Volume 3, a procedure coding system — a US-specific innovation that introduced a complication persisting today. Diagnoses and procedures are governed by different agencies, updated on different schedules, under different guidelines.

Then came the moment that changed everything.

When Medicare’s inpatient prospective payment system launched in 1983, ICD-9-CM became the basis for DRG assignment. Before prospective payment, coding was a documentation and statistical function. After it, coding determined hospital reimbursement. The stakes changed fundamentally — and with them the entire infrastructure of compliance, auditing, and clinical validation.

Systems designed for one purpose — statistical counting — get repurposed for another — financial payment — and the consequences cascade for decades. The classification hadn’t changed. The incentive structure had.

Twenty-Five Years to Implement a Known Improvement

The United States took more than twenty-five years after WHO endorsed ICD-10 to implement it. That fact deserves more astonishment than it typically gets.

By 1990, the NCVHS was already warning that ICD-9-CM had been “stressed to a point where the quality of the system may soon be compromised” (NCVHS Report, FY 1979 — 80). WHO endorsed ICD-10 in May 1990. The US began evaluating it in 1994. HHS published a final rule in 2009 requiring transition by October 2013. After two industry-driven delays, ICD-10-CM and ICD-10-PCS became effective October 1, 2015.

ICD-10-PCS was built on a new architecture, a seven-character alphanumeric structure where each character represents a specific axis of classification. Development under 3M’s CMS contract ran from 1995 through the system’s initial release in 1998, with annual updates since. Implementation in the United States faced sustained pushback from the AMA and over 100 physician specialty societies, which delayed the compliance date repeatedly. First from 2011 to 2013, then to 2014, and finally to October 2015. The delays weren’t driven by clinical logic. The industry objected, and the timeline bent accordingly.

This is how large-scale information systems actually evolve: not by optimizing for the best design, but by negotiating between the clinically ideal and what existing institutions can absorb.

The transition from ICD-9-CM to ICD-10-CM increased diagnostic codes more than fivefold. More specific codes capture clinical reality more accurately — but they also require more specific documentation, which existing clinical workflows were not designed to produce.

When the Tool That Was Supposed to Help Made Things Worse

Electronic medical records were supposed to solve documentation problems. They solved some. They created new ones.

ICD-10’s granularity arrived precisely when EMRs were transforming documentation in ways that undermined its quality. Legibility improved. Timeliness improved. Completeness, consistency, and reliability suffered in new and systematic ways.

Copy-paste, automated data import, templates, and macros let physicians document faster while increasing the risk of inaccurate, bloated, inconsistent records. CMS raised concerns about these features as early as 2016; the Office of Inspector General echoed them.

The picklist problem is the most consequential. Many EMRs let physicians select diagnoses from dropdown menus that automatically populate ICD-10-CM codes. A physician who selects “diabetes with complication” because it appears first has not documented a diagnosis — they’ve made an administrative selection. The resulting code may not reflect their actual clinical assessment, may conflict with documentation elsewhere in the record, and lacks the narrative support required for the code to be reportable on a claim.

We built a system that gets more specific with each revision, then built the tools to use it in a way that strips out exactly the specificity the system was designed to capture.

Computerized order entry added a parallel problem. Paper admission orders often included narrative context that helped coders identify the principal diagnosis. Computerized orders reduced that to a bed assignment. Less information flows to the people who need it.

Two Systems That Were Never Meant to Be Rivals

ICD was built as a statistical classification: finite categories, organized for counting and comparison. Electronic health records needed something different — a comprehensive clinical terminology capable of representing the full specificity of clinical observation in real time.

SNOMED CT filled part of that gap. Where ICD organizes diseases into standardized categories for statistical reporting, SNOMED CT organizes clinical concepts in multidimensional hierarchies for clinical use. In 2013, WHO and the International Health Terminology Standards Development Organisation formally agreed to link the two systems, mapping approximately 19,000 SNOMED CT concepts to ICD-10 codes. The goal: clinicians document in SNOMED CT’s clinical precision; administrative systems report in ICD’s statistical structure.

That’s exactly the distinction Farr drew in 1838 between nomenclature and classification. It took 175 years and a global standards agreement to formally institutionalize it.

The picklist problem is a symptom of failing to maintain that distinction. When an EMR asks a physician to select from ICD-10 code titles, it’s asking a clinical tool to perform a statistical classification function at the point of care, by someone whose cognitive priority is the patient in front of them.

The Automation Promise — And Its Limits

With ICD-10-CM exceeding 70,000 diagnostic codes and ICD-10-PCS exceeding 72,000 procedure codes, the case for automation was obvious. Coding errors cost an estimated $25 billion annually in the US. CMS reported an error payout rate of 6.8% in 2000.

Automated coding research has moved through three stages. Rule-based methods translated coding guidelines into logical programs — effective for small code sets (one study hit 90% accuracy on 45 codes) but unworkable as complexity grew. Traditional machine learning improved performance but treated each code as an independent problem, ignoring the hierarchical relationships central to how ICD actually works. Neural networks — the current stage — have achieved meaningful gains, particularly on the MIMIC ICU dataset, where the best models reach F1 scores around 0.72 on the top 50 codes. Performance drops sharply on full code sets.

Four problems define the current ceiling: the label space is enormous and grows with each revision; most codes appear rarely or never in training data (in MIMIC-III, more than 17,000 codes have zero instances); clinical documents are long, noisy, and full of copied text; and predicted codes must be traceable to specific supporting documentation for both clinical validity and audit purposes.

That last point matters most for any real-world deployment. An automated system that predicts a DRG without identifying the supporting documentation cannot substitute for human judgment about whether a code is defensible. Interpretability isn’t a nice-to-have — it’s the condition under which the output of an automated system can be trusted, reviewed, or contested.

Where Things Stand

Medical coding today sits at the intersection of four converging histories: the 170-year development of disease classification from Farr through WHO; US-specific clinical modifications driven by prospective payment; the EMR transition that expanded documentation capacity while degrading its quality; and the machine learning era that promises automation while confronting limits imposed by the first three.

ICD-11, available from WHO since 2018, is being implemented internationally. It incorporates lessons from SNOMED CT’s architecture and includes more than 55,000 codes. The United States has not adopted it for administrative reporting — continuing a pattern of delayed implementation driven by the complexity of transitioning a system underpinning billions of dollars in annual payments.

Every coding dispute, every documentation gap, every contested DRG traces back to the tension Farr identified in 1838: the gap between how physicians describe what they observe and what a statistical classification system needs to function. That gap has never closed. It has only become more expensive.

The River Is Still Flowing

The history of medical coding is the history of a problem that cannot be fully solved — only managed better or worse.

The classification will always be a compromise. The documentation feeding it will always reflect the mixed terminology of clinicians trained in different traditions. The payment system built on top of it will always create incentives that pull documentation in directions it wasn’t designed to go. The automated systems bridging the gaps will always encounter the limits of imperfect text.

When Graunt counted London’s dead in 1662, the stakes were statistical. When a coder assigns a DRG today, the stakes are financial, legal, and clinical simultaneously. The same fundamental act — classifying a human condition — now determines hospital reimbursement, measures healthcare quality, guides public health policy, and trains the algorithms that will attempt to do the classifying automatically.

The ICD-8 manual put it well, quoting a Victorian statistician: “The scientific purist, who will wait for medical statistics until they are nosologically exact, is no wiser than Horace’s rustic waiting for the river to flow away.” (ICD-8 Manual, WHO, 1967)

The river is still flowing. The people doing this work — coders, clinicians, informaticists, standards bodies — are not waiting for a perfect system. They’re making an imperfect one function.

That’s not a failure of ambition. That’s what serious work looks like.

Sources

International Classification of Diseases, Eighth Revision (WHO, 1967); ICD-10-CM and ICD-10-PCS Major Milestones (CMS/NCHS, 2019); ICD and SNOMED CT: A 21st Century Informatics Solution (WHO-FIC, 2013); “A Survey of Automated International Classification of Diseases Coding” (Yan et al., Intelligent Medicine, 2022); “Defining High-Quality Documentation” (Ericson, ICD10monitor, 2025); The National Committee on Vital and Health Statistics: Fiscal Years 1979 and 1980 (DHHS Publication No. (PHS) 82–1205, 1982); International Classification of Diseases, Adapted for Use in the United States, Eighth Revision (ICDA-8) (US Public Health Service, 1968).


메타데이터
post_id
3bd7d6a467d4
slug
from-bills-of-mortality-to-machine-learning-a-history-of-medical-coding-3bd7d6a467d4
url
https://medium.com/@Dharmabum84/from-bills-of-mortality-to-machine-learning-a-history-of-medical-coding-3bd7d6a467d4
canonical_url
https://medium.com/@Dharmabum84/from-bills-of-mortality-to-machine-learning-a-history-of-medical-coding-3bd7d6a467d4
author_url
https://medium.com/@Dharmabum84
status
ok
fetched_at
2026-06-09 15:37:30