LAMBENCY
DeepCaffein · Public Data Series · 11 · Pricing

The list price was honest.
That turned out to be the harder problem Pricing · 119 Olive Young shelves × 13,780 products × 7 weeks of price history · public-data analysis

Caff · September 2026
Data used — a weekly capture of the top 100 products on each category shelf of Olive Young's online store (2026-06-29 to 08-10, seven pulls · 119 shelves · 13,780 products · 2,044 brands · 75,149 rows), each row carrying that week's list price, selling price and discount rate. Zero internal data.

In the fifth piece in this series I pulled Olive Young's shelf rankings every week for seven weeks and counted which products sit next to which. Three columns in those files went untouched: list price, selling price, discount rate. Nothing new was scraped for this piece. The same seven weeks, opened with a different question.

My first thought was the obvious suspicion — that list prices are inflated so the discount looks bigger than it is. I checked. They are not. And once the list price checked out, something more awkward came into view: one product in four had not sold at that price once in seven weeks. Nobody is running a discount. That is simply the price.

This piece is about why the difference matters. In an order table the two look identical — "20% off" either way. The discount rate does not separate them either. And so a great deal of promotional measurement is running without knowing what it is measuring.

01One row in three had no list price at all

Seven weekly pulls, 29 June to 10 August 2026, top 100 of each of 119 shelves: 75,149 rows, 13,780 products, 2,044 brands. Every row carries that week's list price, selling price and discount rate.

The first thing to notice is an absence. The list-price column is empty in 24,775 rows — 33.0%. Empty means no discount is being shown. The other two thirds carry one.

WeekShare showing a discountMedian discount
2026-06-2963.4%20.0%
2026-07-1367.2%20.0%
2026-07-2765.8%20.0%
2026-08-1070.2%20.0%

Discounting spread from 63.4% of the shelf to 70.2% over seven weeks. The median discount was exactly 20.0% in every single one of them.

The reach moves; the depth does not. 31.6% of discounts land on a round 10, 20, 30, 40 or 50. These are not negotiated numbers. They are picked off a shelf of standard sizes.

02The suspicion was wrong — the list price is not inflated

Before going further I tried to kill the idea. If list prices were padded to manufacture a discount, they would drift, or nothing would ever change hands at them. Four tests.

TestResult
List price unchanged across all seven weeks5,675 of 5,898 (96.2%)
Selling price in undiscounted weeks == list price shown in discounted weeks4,042 of 4,159 (97.2%) exact
Rows where selling price exceeds list price0
Error when the discount rate is recomputed from the two pricesmedian 0.01pp, rows off by >1pp: 0

The second test is the one that decides it. When a product shows a discount in three weeks and none in the other four, you can hold the two against each other: what it actually sold for in the clean weeks, versus the list price it advertises in the discounted ones. They matched to the won in 97.2% of cases.

An empty list-price column is not a scraping failure. In that week the product really is selling at that price.

So the problem is not that the list price lies. The list price is true — and a large share of products have never sold at it. That is a much more awkward thing to deal with. A lie can be corrected. Here nobody has lied and the numbers still do not come out.

03One in four had never sold at full price

I narrowed to the 6,630 products that appear in all seven weeks (48% of 13,780) and asked two things of each: how many weeks did it show a discount, and did the selling price actually move.

6,630 products, all seven weeksPrice never movedPrice movedTotal
Never discounted61721638 (9.6%)
On and off (1–6 weeks)24,1574,159 (62.7%)
Discounted every week7451,0881,833 (27.6%)
6,630 products present in all seven weeks Never discounted 638 · 9.6% On and off 4,159 · 62.7% 745 Always on 1,833 · 27.6%

Figure 1 · The 6,630 seven-week residents, split three ways. The 745 in the lighter block held both price and rate for all seven weeks — that is a price, not a discount

The bottom row is the finding. 1,833 products — 27.6% — have zero observations of themselves at full price. And 745 of them go further still: same selling price and same discount rate for seven straight weeks. All 745 held the discount rate constant, at a median of 20.0%.

Seven weeks, one price, one discount rate. That is not a discount. That is the price.

I checked whether "present in all seven weeks" is doing the work. It is not. Loosen it to products seen in at least two weeks — 11,297 of them — and the share is 32.4%; at four weeks or more (9,019 products) it is 29.3%; the strict seven-week panel gives 27.6%. Widening the sample nudges the number up, not down.

04The order table cannot tell these three apart

The three rows above are different animals.

What it isWhat it should be in the books
Never discountedselling at pricerevenue
On and offcut for that week onlypromotional spend — measure the lift
Discounted every weekthis is the pricepricing — revise the list

But rows two and three are indistinguishable in an order table. Both land as "list 43,000 / sold 34,400 / 20% off". Looking at a single order, there is no way to tell whether that 20% is this week's decision or the last six months of them.

No schema draws this line. And yet marketing has to draw it, because one side is a budget you raise or cut and the other is a price you set. Left in the same column, the two decisions end up arguing over one number.

05"Deep means promotional" does not hold either

There is an obvious objection here. One order tells you nothing, fine — but the discount rate must. An always-on markdown will be shallow; a real event will be deep. It sounds right, so I measured it.

Discount depthAlways-on (baseline 0)Intermittent
Under 10%13.5%14.9%
10–20%29.6%33.2%
20–30%31.8%28.1%
30–50%21.6%20.1%
50% and over3.6%3.7%
Median20.0%20.0%
10% 20% 30% 13.5 14.9 under 10% 29.6 33.2 10~20% 31.8 28.1 20~30% 21.6 20.1 30~50% 3.6 3.7 50%+ Always-on (baseline 0) · n=12,831 Intermittent · n=18,844

Figure 2 · Discount depth. Both medians sit at 20.0%; the widest band gap is 3.6pp

That is 12,831 observations against 18,844. The medians agree to the decimal, and the widest gap in any band is 3.6pp. On means it is 21.7% against 20.9% — the always-on group is, if anything, marginally deeper.

As distributions the two groups are effectively the same. Depth will not separate them.

This matters because a lot of companies use exactly this rule as a stand-in — anything past some threshold counts as a promotion. It does not work here. What makes something a promotion is not how deep the cut is but whether the product ever sold without one, and that fact is not present in any single order. It only exists across time.

06So "promotional share" is not one number

Here is what the blurred line costs you. Across the seven-week panel there are 46,410 product-week observations.

ObservationsShare
Showing a discount31,67568.3%
Of those, belonging to baseline-zero products12,83140.5%
Reclassify baseline-zero as pricing18,84440.6%

From the same data, "promotional share" is 40.6% or it is 68.3%. The gap is 27.6 points (the difference before rounding — subtracting the two displayed figures gives 27.7).

Those 27.6 points are the number this piece produced. They appear in no source file; they only exist once you fold the presence or absence of a list price into product-by-week. And the width is not an error bar — it is resolution. It says the order table cannot tell you where inside that range you actually are.

The remaining 62.7% — the 4,159 products that genuinely go on and off — are not in good shape either. Their median count of undiscounted weeks is two. For 1,423 of them (34.2%) it is one; for 2,436 (58.6%) it is two or fewer. Measuring lift means comparing against the weeks you did nothing, and here "the weeks you did nothing" is one or two observations. A single week of weather, or one competitor's event, flips the answer.

07Only 4.1% were observed repeating

When comparison weeks are thin, the usual move is to pool repeats: if a product switched its discount on and off several times, each cycle is another trial. So I counted how many times each of the 4,159 intermittent products changed state across the seven weeks.

On↔off switches in 7 weeksProductsShare
12,14151.5%
2 (on then off — one complete promotion)1,84944.5%
3 or more (repeats)1694.1%

Half of them (51.5%) switched exactly once. That means the product turned its discount on and stayed on, or off and stayed off, inside the window — and within this window you cannot tell a promotion from a price change. The observation is cut off by the edge of the window.

A full cycle — on, then off — shows up for 44.5%. Two or more cycles: 169 products, 4.1%.

Buying precision through repetition is available for one product in twenty-five.

Pooling does not rescue this. It is the kind of gap that has to be closed at the moment of recording. Re-reading last quarter's orders will not recover a list price that was never written down.

08The gap is worst where the money is

The rate varies enormously by shelf. Ranking the 86 shelves that hold at least 30 seven-week products:

Highest baseline-zeroLowest
Hygiene — nappies79.2%Hobby — music0.0%
Sheet masks75.0%Fashion accessories2.9%
Home & baby68.5%Dermo — cleansing3.8%
Makeup — base65.4%Food — baby food7.5%
Health — fitness gear61.7%Hair tools & brushes8.1%

A spread of 79.2 points. Repeat-purchase commodities where brands are close substitutes sit at the top; categories where the list price is part of the identity sit at the bottom. A category-average diagnosis erases the entire spread.

Shelf position splits it too.

Shelf rankProductsBaseline zero
1–1055431.0%
11–2080232.3%
21–3080030.6%
31–501,56027.4%
51–1002,91425.1%

Around 30% in the top thirty, 25.1% down at 51–100. The better the slot, the less likely there is anything to compare against.

The unmeasurable stretch and the high-volume stretch are the same stretch. So "we track discount ROI" can be perfectly true and still undefined on the products that matter most.

At brand level the distribution is barbelled. Of the 435 brands with five or more seven-week products, 171 (39.3%) have no baseline-zero product at all, and 25 (5.7%) are baseline-zero throughout. This is not an industry habit. It is a choice, made differently by different companies.

09The metric I did not use, and why

A more obvious metric presented itself first: price amplitude, the spread between a product's highest and lowest weekly price over its mean. Median 14.8% across the panel, above 48% for the top decile. It writes itself — look how prices swing.

I dropped it. Opening the top of that ranking by hand turns up this:

(1 sheet / 30 sheets) · [1 / 3 / 4 / 10 pack] · single or gift set (50ml / 200ml / 500ml)

One product code, several sizes and pack counts underneath it, and the displayed price is one of them. When 900 won and 27,100 won live under the same code the amplitude reads 549% — but the price did not move; the visible option did. The data agrees: of the 41 products with amplitude above 100%, 43.9% carry an option marker in the name, against 11.9% across the panel.

Baseline availability is far less exposed to this, because it reads only the presence or absence of a list price. Products with option markers come in at 32.1% baseline-zero, those without at 27.1% — a five-point difference, not the multiples that amplitude produces.

The same data offered two metrics. I took the one the contamination could not reach. Choosing comes before measuring.

Seven weeks is a choice too. At four weeks an always-on price and a four-week promotion are the same object — both look discounted in every observed week. Stretch to six months and list-price revisions leak in, so the "baseline" becomes an old price. Seven weeks sits between those, which is why this piece only ever says "always-on" to mean "in all seven weeks".

Restricting to seven-week residents is a choice as well. The shorter the observation, the more baseline-zero appears by construction — a product seen in one week is baseline-zero as long as that single week was discounted. Measured: 57.5% across the 2,483 products seen in a single week, 35.2% at three weeks, 27.6% at seven. Mixing short observations inflates the headline. The most conservative cell is the one in the text.

Folding to product-by-week is the last one. The source is a per-shelf ranking, so a product can appear on up to four shelves in the same week (11.6% appear on two or more). Counted unfolded, products with wide shelf presence are counted repeatedly, and the result tilts toward always-on.

10What to actually do

That is as far as the outside view goes. Inside, there are three things.

(a) The metric. Per SKU, baseline availability = undiscounted weeks ÷ total observed weeks. It needs four fields: product code, week, the price it actually sold at, and the list price. Most companies already have all four. They cannot produce the metric because they do not keep them by week.

(b) The thresholds are 0 and 3. Neither is an arbitrary line.

Zero — lift is a difference between two states, and at a baseline of zero one of the states does not exist. The effect is not small; it is undefined. That boundary comes from logic, not from statistics.

Three — in a seven-week window, three undiscounted weeks means one odd week cannot move the median, because two remain. At two, one odd week is half your evidence. In a thirteen-week window the number becomes five. The rule is not the digit but "no single observation may be worth half the baseline."

(c) The decision rule. Put your own values in and the next move is fixed.

Baseline (weeks)What you do with that SKU
0It is not a promotion. Take it out of promotional spend and reclassify it as pricing. Reset the list price to what it actually sells for; the discount rate becomes zero, and the conversation moves from "should we discount less" to "is this the right price"
1–2Estimate, but never report a point value. Publish a range, and do not move budget on this SKU alone
3+Use it. Lift and elasticity get calculated here and nowhere else

What changes is not the budget but which line of the P&L the volume sits on. Leave baseline-zero volume in promotional spend and you will eventually be told that cutting discounts raises profit — when what is actually on the table is a price increase, and the volume leaves with it. Move the column and the pricing decision becomes visible as one.

You can run this on your own data today. Over the last thirteen weeks, count for each SKU how many weeks it sold at list, sort them into three buckets — 0, 1–2, 3+ — and take each bucket's share of revenue. The share sitting in the zero bucket is what has to come out of your promotional reporting. On the shelf data here that share was 40.5% of observations.

And if the list price for that week is not in your order table, the calculation is not available for the past at all. Then the job is not analysis, it is instrumentation: from today, write the list price and the realised price per SKU per week. Thirteen weeks later the three buckets exist for the first time.

11What this data cannot show

The boundaries, stated plainly.

1. There is no volume and no revenue. Every share here is a share of product-week observations, not of sales. Do not read "40.5%" as a revenue share. 2. Discounts below the sticker are invisible. Coupons, card promotions and membership pricing are not in this data. Real paid prices are lower than what I saw — which makes baselines scarcer, not more plentiful. 3. Only the top 100 of each shelf was collected, and baseline-zero is more common higher up (31.0% against 25.1%). Across a full shelf the 27.6% would fall. This number leans high. 4. Seven weeks is one season (29 Jun – 10 Aug). It does not contain a full promotional calendar. "Always-on" means always-on inside this window, and the 51.5% single-switch share in §07 is the size of that censoring. 5. No costs, no margins. No absolute ROI is computed here. The only verdict this piece issues is whether a baseline exists. 6. One channel. The same product may be priced differently on a brand's own store or elsewhere, in which case its baseline is not the one seen here.

12What this piece is saying

I opened the file suspecting the list price, found it honest, and then found the real problem behind it. One product in four (27.6%) carries a true list price it has never sold at. For those products "20% off" is not a discount, it is the price — and the order table records it exactly like a real promotion. The discount rate does not separate them either: both groups sit at a median of 20.0%.

The consequence is measurable. From identical data, "promotional share" is 40.6% or 68.3%, and nothing in the order table decides where inside those 27.6 points you are. Even the genuine promotions have a median of two comparison weeks, and only 4.1% were seen repeating. The shortage is worst exactly where the volume is.

Fixing it does not take new data. This piece did not scrape anything new either — it used three columns that were already sitting in a file we had. What it takes is keeping price by week, and dividing what accumulates by baseline availability. The work that comes before analysis is making the thing analysable, and in this category that work lives in the price history.

Where to start, depending on where you sit

If you areStart here
Brand marketingBaseline weeks per SKU for the last thirteen. Begin with the revenue share of the zero bucket — that is what leaves your promotional report
CRM / dataDoes your order table keep that week's list price? If not, baselines cannot be reconstructed after the fact. Recording from today is the only route
Pricing / merchandisingBaseline-zero SKUs are pricing decisions, not discount decisions. Reset the list to the realised price, then look again
LeadershipWhen someone reports that cutting discounts will raise profit, ask for the baseline availability of that volume. At zero, you are being handed a price increase

Method and limits

The source is a weekly capture of the top 100 products on each category shelf of Olive Young's online store, 2026-06-29 to 2026-08-10, seven pulls (119 shelves, 13,780 products, 2,044 brands, 75,149 rows). Each row carries that week's list price, selling price and discount rate. It is the same capture used for the fifth piece in this series, which did not touch the three price columns.

The unit of observation is product code × week. A product can appear on several shelves in one week (11.6% on two or more, up to four), so rows are folded before counting; unfolded, products with wide shelf presence are counted repeatedly. Within a product-week, selling prices disagreed in 2 cases and list prices in 1. An "undiscounted observation" is a week with an empty list-price field, and that definition is validated by the second test in §02 (97.2% exact agreement).

"Baseline zero" means every observed week showed a discount. Headline figures use the 6,630 products present in all seven weeks; loosening residency to two weeks or more (11,297 products) gives 32.4% and to four or more (9,019) gives 29.3%. Because short observations produce baseline-zero by construction (57.5% at one week), the most conservative cell is the one quoted. Shelf-level figures assign each product to its modal shelf and compare only the 86 shelves holding at least 30 resident products. Price amplitude was excluded as a metric because option bundling sits under the product code (§09). The reproduction script is kept with the research note.

No brand is named in this piece. The argument is about structure rather than any individual company, and the top of these shelves may include companies we are in live conversations with. Shelf names are Olive Young's category labels, not company names.

This is as far as public data can see. Repurchase, incrementality, attribution — the numbers that change decisions live inside your own data, and making them countable is what Lambency does.

caffrey.w.lee@gmail.com