Eight authoritative ways to say “diabetes”
Ask eight recognized stewards for the ICD-10-CM codes that identify diabetes and you get eight different answers, spanning 87 to 633 codes. Which one you grab quietly rewrites your cohort.
Gript Technologies · Health-data methods
Eight published diabetes value sets, compared on their ICD-10-CM members only.
Ask eight recognized stewards for the ICD-10-CM codes that identify diabetes and you get eight different answers, spanning 87 to 633 codes. Which one you grab quietly rewrites your cohort.
A data analyst gets a one-line request: pull the type 2 diabetes patients. They do the sensible, defensible thing, reaching for a published, steward-curated value set instead of hand-writing codes. The catch is that there are dozens of published diabetes value sets, they are not the same, and nothing about picking one over another gets written down. I pulled eight and measured exactly how much they disagree, using only their ICD-10-CM members, the vocabulary a claims or EHR analyst actually filters on.
To keep the comparison about definitions rather than organizations, each value set is labeled by the type of steward that publishes it: a professional society, a quality-measure body, a federal program, and so on. One extra line, the naive baseline of every billable code in ICD-10-CM chapter E11, stands in for what a rushed analyst types by hand.

Grab a diabetes set as a type-2 proxy, and four in five codes miss
The narrow, type-2-specific lists (rows B and C) stay close to chapter E11. The broad diabetes lists balloon by pulling in whole other conditions: the jump from about 280 to more than 400 codes (rows F onward) is almost entirely secondary diabetes (E08 and E09), 172 to 230 codes of it, plus type 1 (E10) and gestational (O24). Use one of those to answer a type 2 diabetes question, and four of every five codes you match describe something else.
Even the two curated lists that specifically say Type 2 in their title disagree with each other, 96 versus 206 codes, a 2.1x gap before you ever reach the broad sets.
Two families that barely overlap
Pairwise overlap, measured as Jaccard similarity, shared codes divided by combined codes, shows the eight are not scattered randomly. They fall into two camps that agree internally and disagree across the divide.

A small agreed core, a long idiosyncratic tail
Across all eight definitions there are 634 distinct ICD-10-CM codes. Only 86 appear in every one, the undisputed core of type-2 diabetes. At the other end, 120 codes appear in exactly one set: a single steward’s idiosyncratic inclusion that no one else counts.

What the broad sets drag in
Concretely, here are codes every definition agrees are in scope, next to a sample of what the broadest diabetes grouper adds that a type-2 cohort arguably should not include. All are public-domain ICD-10-CM.
| Code | Description | Status |
|---|---|---|
| E11.10 | Type 2 diabetes mellitus with ketoacidosis without coma | In all 8 (core) |
| E11.21 | Type 2 diabetes mellitus with diabetic nephropathy | In all 8 (core) |
| E11.311 | Type 2 diabetes mellitus with unspecified diabetic retinopathy with macular edema | In all 8 (core) |
| E10.10 | Type 1 diabetes mellitus with ketoacidosis without coma | Type 1, added by broad sets |
| E08.01 | Diabetes due to underlying condition, with hyperosmolarity with coma | Secondary, added by broad sets |
| O24.011 | Pre-existing type 1 diabetes mellitus, in pregnancy, first trimester | Gestational, added by broad sets |
Why this matters
None of these value sets is wrong. Each was built for a purpose, a quality measure, a comorbidity index, a regulatory program, and is defensible in that context. The problem is silent substitution: an analyst grabs whichever diabetes set is handy, the choice never enters the study record, and the cohort’s size, prevalence and downstream estimates shift by multiples depending on a decision no one documented.
Two codes changing hands is a rounding error. A 7.3x swing in what diabetes means is a study-design decision, and it deserves to be made deliberately and written down: which codes, on what authority, with what included and excluded, and why. The value sets compared here are anonymized by steward type because VSAC licensing restricts how individual value-set names and contents are displayed outside the licensed context; the provenance principle still applies to every set you use in your own work. That provenance is the difference between a cohort you can defend in review and one you merely hope was right.
Before you reuse a value set, look at what is actually in it, the type mix, the complications, the gestational and secondary codes, and record the choice as part of the protocol. A value set is a decision, not a default.
The tool behind this analysis
CodeSet Workbench drafts ICD-10-CM value sets from plain English, then shows the composition, comparator diffs and rationale behind every code, with FHIR export and a provenance record you can attach to a protocol. It is the workflow this note argues for, made routine. See it at codeset.gript.io, or browse the product page.
Method and data
Eight code lists for diabetes were compared using only their ICD-10-CM members. Seven are published value sets retrieved from the U.S. National Library of Medicine’s Value Set Authority Center (VSAC) via its FHIR terminology service ($expand); non-ICD-10-CM members (SNOMED CT, CPT, LOINC, RxNorm) were excluded. The eighth is a naive baseline: every billable code in ICD-10-CM chapter E11 (fiscal year 2026, public domain, from CMS and NCHS). Value sets are identified by steward type only. Metrics are computed on the code sets as retrieved; small cross-year vocabulary differences are not reconciled. This content includes material from VSAC and UMLS; UMLS Metathesaurus terms are used under the NLM License. For methodological illustration, not clinical advice. Value sets are identified by steward type because UMLS/VSAC license terms restrict redistribution of the lists themselves; the retrieval method above lets any licensed VSAC user reproduce the comparison.
Subscribe
New analysis, when it publishes.
Occasional notes on claims data, clinical informatics, and healthcare AI. No newsletter filler, no sequences, unsubscribe in one click.
Or follow the RSS feed