LAMBENCY
DeepCaffein · Public Data Series · 06 · Pharmaceuticals

The data is public.
Nobody's using it Pharmaceuticals · HIRA prescription data × drug registry · dose mix and regional demand · public-data analysis

Caff · August 2026
Data used — HIRA drug-utilization open API (top 100 classes × 3 time points; district and ingredient level; 3,695 queries), the reimbursed-drug ATC mapping list, resident population registry, HIRA research reports and ministry releases, peer-reviewed literature. Zero internal data.

I spent the first months of the pandemic on a hospital data team in the United States. There was no vaccine yet, and the doctors were testing drugs that might work. My job was to make those tests statistically real — starting with the math of how many patients you need before a difference stops being luck. What stayed with me from those months is this. The same drug looked different at different doses, and sometimes the difference dissolved when you measured it properly. But one thing was unambiguous: what dose reliably separated was not efficacy but safety. The high-dose arm got into trouble first. The conclusion those months left behind is simple — for the same molecule, the milligrams are a different event.

01The note-parsing years

This is not a story about making vaccines. I am a data person, and that work led me to a healthcare IT company, collecting oncology data from American hospitals. The data lived in three systems: Epic and Cerner, which run whole hospitals, and iKnowMed, used by community oncology clinics. They store the same items, but every hospital wrote them differently — some in codes, some in other units, some not at all. Every new hospital meant redoing the mapping from scratch.

The data had to flow in real time. If the pipeline died overnight, the next morning's clinic saw yesterday's numbers, so we rotated on-call shifts through weekends and holidays. And the hardest part was the prescriptions themselves. What drug a patient received, at how many milligrams, was in no column at all. It was inside the sentences of the physician's notes. Building the rules that pulled drug names and doses out of those sentences was half the project. The notes also carried patient identifiers, and in the end the project was stopped over exactly that. The prescription data I knew was that kind of thing — something you build until you have to stop.

02In Korea, it's a table

That is why I paused when I first opened Korea's public health-insurance data. Here, that data is a table. Ingredient, dose, and formulation live inside a single nine-digit code, and prescription quantities and amounts for each code are published by district, by month, nationwide, back to 2010. No notes to parse. No privacy problem — it is aggregated. It sits there, tidier than what we built on weekend shifts, for anyone to look at.

But being public and being usable turned out to be different things. The data lives behind a query screen, not a download. Pick one therapeutic class, one region, one month, and you get one table. It is a door built for a person with a mouse — assembling this article's national picture took 3,695 queries, over two days, at whatever pace the server would tolerate. The official file channels do not carry this combination at all. The data is open; the door is narrow. In America the wall was the physician's note. Here, the wall is this door.

How many people have actually walked through it can be counted. On Korea's open-data portal, applications for this prescription-data API total 2,352 in ten years (as of 2026-08-23). The hospital-location data the same agency opened the same year has 8,606. Apartment transaction prices collected 16,874 applications within a year and a half of registration. On GitHub, public repositories touching this API: zero. Repositories touching real-estate prices: 248. Sixteen thousand hands lined up for property numbers; this door saw two thousand in a decade, and I could not find a single public analysis they left behind. Why? Housing prices are everyone's own problem; prescription data belongs to an industry — an industry used to buying estimates, facing a narrow door. A guess, but probably all three.

Sixteen thousand hands gathered around real-estate prices in eighteen months. This door saw two thousand in ten years.

So this industry's gap has a different shape from the earlier essays. The gaps so far were data that didn't exist (the food registry records no deaths) or data in someone else's hands (the shelf belongs to the channel). Here it is the opposite: the data is this open, and nobody is using it. So let's use it.

This essay will follow one drug. The lipid-lowering combination — a statin and ezetimibe fused into a single pill, the kind you may know by names like Rosuzet. Why this one: over the past five years it became the most-prescribed drug class in Korea. You would expect the best-selling drug to be the one built with the best knowledge of demand. What the data shows is closer to the opposite. Here is that story, in order.

03A map of the market — its shape is changing

Data first. I took the 100 therapeutic classes with the most reimbursed products and collected nationwide claims for three months: June 2019, June 2024, and March 2026, the latest month available. The monthly reimbursement for these 100 classes went from ₩720bn to ₩1,190bn in five years. That is 64% growth — but what catches the eye before the growth is that the market is changing shape.

As of 2024, the biggest class is not statins. It is the lipid combination — a statin plus ezetimibe in one pill. ₩78bn a month, 4.4 times its level five years earlier, overtaking single statins (₩62.6bn, +50%). The same direction repeats elsewhere: blood-pressure combinations are up +94% and +279%, diabetes combinations +118%. Much of this is substitution rather than growth — lipid drugs as a whole grew +135% in five years, while the combination share within them went from 35% to 55%. A market of one ingredient per pill crossed the halfway mark into a market of several.

Other classes tell their own stories on the same map. In acid suppression the old generation (H2 blockers) fell -34% — the only decline in the top 100 — while the new generation (PPIs) filled the space at +174%. And the class containing choline alfoscerate, a "brain-function" drug under reimbursement review for disputed efficacy, still moves ₩36.4bn a month. But this essay doesn't linger there; it follows the class that became number one.

+0% +100% +200% +300% ₩20bn ₩40bn ₩60bn ₩80bn Lipid combos Statins PPIs Antiplatelets BP combos (ARB+CCB) Choline alfoscerate cls. Diabetes combos Contrast agents ARBs (single) Dementia drugs Ophthalmics B02BD C09DX C10BX H2 blockers Vertical = 5-yr growth · Horizontal = monthly reimbursement, Jun 2024 · combination drugs

Figure 1 · Size and 5-year growth of the top 100 classes (monthly reimbursement Jun 2024 × growth vs Jun 2019). Amber = combinations

(Note: the disease names attached to this data are the primary diagnosis on the prescription — the patient's representative condition, not the drug's indication. That is also why the most frequent diagnosis on statin prescriptions is not a lipid disorder but hypertension, at 25% of prescribed volume.)

You can also see the order in which this combination spread. Divide each region's prescriptions by its population — proportional is 1.0 — and in 2019, when the market was a third of its current size, Seoul was highest at 1.31. The new drug settled in Seoul first. The five years since were growth for everyone: Seoul's per-capita prescriptions rose 3.3-fold. The rest of the country simply rose faster — Gwangju and Daegu 4-fold, Jeju 4.3-fold, Sejong 5.2-fold. By 2024, per-capita prescriptions in Gwangju and Daegu (2.6 tablets per person-month) had reached Seoul's level. Seoul did not decline; the country caught up to Seoul. In the chart below, the multiples converging toward 1.0 are the trace of it. Who learned from whom — the path of transmission — is not visible in this data. What is visible is who settled first, who followed, and how fast the gap closed.

population parity 1.0 Sejong 0.51→0.72 Gwangju 1.05→1.16 Jeju 0.61→0.72 Daegu 1.08→1.17 Gyeongbuk 0.75→0.83 Ulsan 0.83→0.89 Chungbuk 0.87→0.93 Chungnam 0.77→0.82 Jeonnam 0.85→0.89 Gyeonggi 0.87→0.90 Gyeongnam 0.93→0.96 Daejeon 1.26→1.28 Busan 1.16→1.17 Jeonbuk 0.99→1.01 Gangwon 0.88→0.87 Incheon 1.08→1.00 Seoul 1.31→1.17 ○ 2019 · ● 2024

Figure 2 · The lipid combination's regional multiples, 2019 (○) → 2024 (●). 1.0 = proportional to population. Amber = rising

(Note: region here is the location of the prescribing institution, not the patient's address. Sejong's low multiple and Daejeon's high one are a pair — Sejong residents seen in Daejeon clinics — and should always be read together.)

04Same drug — which milligrams sell

This is the part the essay was written for. When combinations pass half the market, the number of products to make multiplies. Two ingredients make a pair, and every pair splits again by dose. What to make, and how much of it, becomes a much harder question than before.

Count the supply first. The 21,953 reimbursed drugs collapse to 2,268 ingredients by code. What matters is that the nine-digit code holds more than the ingredient: four digits for the molecule, two for the dose, three for the form. Which milligrams it is, is already split out in the code. The exact distinction we tried to parse out of physicians' notes in America is one line of code here. Two-thirds of ingredients (65.5%) come in a single dose — but there is another end. Donepezil, the dementia drug, comes in ten dose-forms made by 122 companies as 288 products. Rosuvastatin: seven, 124 companies, 330 products. For hospitals and pharmacies this is an inventory problem — stocking one drug means choosing among hundreds.

And because the dose is in the code, demand can be read by dose too. I split all dose combinations of the two number-one classes — lipid and diabetes combinations — by prescription volume. Doses are written the way product names write them: "ezetimibe 10mg + statin 5mg" becomes 10/5. The rosuvastatin combination sells in four doses: 48.5% of prescriptions go to 10/5 alone, 33.2% to 10/10. The atorvastatin combination is even more concentrated — 10/10 alone takes 67.6% of prescriptions. The metformin-dapagliflozin diabetes combination scatters across nine doses. Same "combination drug," very different shapes of demand.

Now set that demand beside the products companies have registered. The atorvastatin combination's products: 96 at 10/10, 97 at 10/20, 66 at 10/40 — spread almost evenly across the three doses. By product count alone, the three doses look equally important. Yet two-thirds of the prescriptions go to 10/10 alone. The rosuvastatin side shows the reverse mismatch too: the 10/20 dose takes 10% of prescriptions but holds 31% of the products, while the youngest dose, 10/2.5, where demand is growing, holds 6%. Prescriptions concentrate; products spread evenly — demand and supply have different shapes. In the chart below, the width mismatch between the upper and lower bars is that difference.

Prescriptions pile onto one dose; products sit evenly across three. And this, with demand published at national scale.

Same dose order. Top = where prescriptions go, bottom = where products sit (Jun 2024). The wider the mismatch, the bigger the gap Rosuvastatin + ezetimibe Rx 10/5 · 48% 10/10 · 33% 10% 9% Products 10/5 · 32% 10/10 · 32% 10/20 · 31% 6% Atorvastatin + ezetimibe Rx 10/10 · 68% 10/20 · 24% 9% Products 10/10 · 37% 10/20 · 37% 10/40 · 25%

Figure 3 · Prescription share (top) vs registered-product share (bottom) by dose, Jun 2024. Width mismatches are the gap

Reduced to one number: take the share of prescriptions going to each pair's most-prescribed dose, minus that dose's share of products — the atorvastatin combination is 67.6% versus 37.1%, a gap of 30.5 points. Weighted by volume across both classes, the average gap is 18 points. The more a dose is prescribed, the less of the product lineup it gets.

To be precise about what this does and does not say: product counts are registered products, not production volume. How much each company makes of each product is not in this data. So the number says one thing — the shape of the lineup differs from the shape of demand. What is oversupplied and what is short only comes out when each company joins its own production and shipment data. Public data shows the mismatch exists; that is where it stops.

Add time and one more thing appears. The rosuvastatin combination's lowest dose, 10/2.5, did not exist in 2019. It reached 8.5% in 2024 and 12% by March 2026. On the atorvastatin side, the low 10/5 dose only recently entered at 5%. The patients who needed lower doses existed all along. The products arrived late. And the signal had been there for years: a 2019 study found 15.6% of prescriptions involved tablet-splitting. That is what physicians and pharmacists do when the doses on sale don't match the doses needed.

I don't think the industry didn't know. But a public analysis that measures this mismatch at national scale, down to the individual dose, I could not find. What it took was not new data. It took joining two already-public tables — one from prescriptions, one from the registry — on a nine-digit code.

05The rulers this essay made

To make the observations reusable for any class, here they are as rulers. This essay measured two things.

One is the dose gap: within an ingredient pair, the prescription share of the most-prescribed dose minus that dose's share of registered products. Near zero when demand and lineup share a shape; the wider it opens, the emptier the shelf where demand concentrates. Across the two classes here it averaged +18 points; the atorvastatin combination was +30.5 points.

The other is the regional multiple: a region's share of prescriptions divided by its share of population. 1.0 is proportional; the further from 1.0, the more that region has its own story. With this ruler we measured the country catching up to Seoul, the multiples converging toward 1.0.

The decision rule: if the dose gap is wide, if corroborating signals overlap — tablet-splitting, doses arriving late — and if the regional multiples are moving, that class's lineup and inventory allocation likely deserve a redesign against demand data. But the conclusion is not drawn from public data alone. Product counts are not production volumes; the final call comes after joining each company's own production and shipment numbers.

RulerHow it's measuredReferenceWhat you do when it trips
Dose gapRx share of the most-prescribed dose − that dose's product share Near 0 = aligned. The two classes here average +18pts Starting with the widest pairs, check against your own production and shipment data — consider a lineup redesign
Regional multipleRegion's Rx share ÷ population share 1.0 = proportional to population Re-examine field and inventory allocation starting with regions far from 1.0 or moving year over year

06The money beside the data

Why it matters that nobody uses this data can be counted in money. By a HIRA research report's estimate (2018), duplicate prescriptions of the same ingredient leak ₩138.2bn a year — 1.16% of prescription value. Drugs prescribed but never taken are estimated at ₩218bn a year. Of leftover medicine, 55.2% went into the trash or the drain, and in 94.4% of cases the patient decided alone whether to take it.

It is not for lack of infrastructure. Hospital-grade EMR adoption was already 93.9% in the 2020 survey; certified institutions went from 41 in June 2020 to 4,052 by December 2024. The system that checks every outpatient prescription in real time for duplication and contraindications (DUR) screened about 1.49 billion prescriptions in 2025. This is a country where prescription data flows in real time. What's missing is the layer that reads that flow into demand and inventory — and the money leaks through that missing layer.

And there is one more hole. Patient-side waste at least has official estimates. For the inventory hospitals and pharmacies discard, I could not find a statistic at all. The eye that counts waste stops at the patient. Institutional inventory lives only in each institution's own ledger, and nobody adds it up.

The eye that counts waste stops at the patient. How much hospital and pharmacy stock gets discarded — there is no statistic.

07What this data cannot see

Its field of view ends at the reimbursement boundary. This is data that captures ₩1.2 trillion a month, and the Wegovy boom everyone knows about registers as zero inside it. Obesity drugs have no reimbursed products at all, and reimbursement claims for the GLP-1 class Wegovy belongs to measured zero at all three time points. A diabetes drug that did enter reimbursement — the SGLT2 inhibitors — shows up clearly, growing from ₩2.6bn to ₩9.1bn a month. The private-pay market is outside this lens.

These are prescriptions, not consumption. As the non-adherence numbers in section 06 show, the two differ — read prescription volume as demand with that margin in mind.

Diagnoses are primary diagnoses. The representative condition on the prescription, not the drug's indication. Use the distribution to see the patient profile, not as indication statistics.

Regions are institution locations. They can differ from where patients live, and regions dense with large hospitals look inflated. Hence dividing by population, and reading outflow-inflow pairs like Sejong-Daejeon together.

Masking and lag. Cells with three or fewer manufacturers are masked, and the aggregates run months behind — the latest month in this essay is March 2026.

08What this essay wants to say

In America we parsed physicians' notes to build prescription data, and had to stop. Korea publishes the same data by ingredient, dose, region, and month — and in ten years this door saw about two thousand applications and, as far as I could find, not one public analysis. Using it directly showed: while combinations crossed half the market, the best-selling drug's dose demand and its product lineup diverge by an average of 18 points; low doses arrived years after their demand; and the country's five-year catch-up to Seoul is stamped in the regional multiples. This was never a picture missing for lack of data. It was missing for lack of anyone looking. What the two rulers — dose gap and regional multiple — read for your own therapeutic class can be measured today, from public data.

Pharma · portfolio & production Start with your dose gap Does any of our combos put two-thirds of demand on a dose that holds a third of the products — and is a late-arriving low dose Pharma · sales & marketing Regional multiples, in motion Is our market one where the catch-up is done, or one still converging — the two need different field allocation Distribution · wholesale · pharmacy Re-rank inventory by demand share Combination drugs multiplied the SKUs — if half the prescriptions sit on one dose, the ordering priority is already Hospital · healthcare IT Does flowing data reach decisions Prescription data streams through EMR and DUR in real time — but no statistic even exists for institution-level

Figure 4 · What to look at first, wherever you sit

Method & limits

  • Sources — HIRA's drug-utilization open API (registered on Korea's open-data portal): nationwide claims for the top 100 therapeutic classes (ATC level 4) by reimbursed product count, at 2019-06, 2024-06 and 2026-03, plus district-level (298 official region codes) and ingredient-level sweeps for the lipid and diabetes combinations — 3,695 queries, dispensing basis, all payer types summed. The product and dose axis comes from the reimbursed-drug ATC mapping list (as of 2025-06-30), 21,953 products.
  • The dose-safety account in section 01 — grounded in public literature: a randomized trial halted early for significantly higher mortality in its high-dose arm (CloroCovid-19, 2020); efficacy differences were non-significant in most trials. Personal experience is used only as scene and timing.
  • Dose gap definition — within an ingredient pair, the volume share of the most-prescribed dose minus that dose's share of registered products. Product counts are registration counts, not production volumes — the gap speaks only to the shape of the lineup.
  • Regional multiple definition — a province's share of prescription volume ÷ its share of resident population (Ministry of the Interior, same-month basis). Gwangju and South Jeolla are computed on their pre-2026 administrative boundaries.
  • Reproducibility — the figures in this essay pass a verification script that recomputes them from the raw data; application counts and repository counts are as of 2026-08-23.

This is as far as public data can see. Repurchase, incrementality, attribution — the numbers that change decisions live inside your own data, and making them countable is what Lambency does.

caffrey.w.lee@gmail.com