The same customer is a different person
in every ledger
Customer data unification · 85 consumer/F&B/hospitality companies, public web records · public-data analysis
When I took over a retail data hub where the orders of 1.95 million customers and 256 brands came together, the first thing I met was not a fancy machine-learning model. It was the ledgers. The same product was scattered across half a dozen differently-spelled product names, and the same customer carried a different member ID on every channel. Nobody could answer the basic question "how many regulars do we have" — not because there was no number, but because there were several. Analysis could only begin after those ledgers were stitched into one person, one product.
01What happened
Was that one odd company? No — all 96 brands that passed through the hub were the same. Gentle Monster, RAWROW, Antoine Solais, Hunter. Big or small, thriving or not, the ledgers at the starting line were all in the same state. Half of the daily work at the hub was not modeling but two mappings — finding the same person, finding the same product.
So what about outside the hub? This time, we counted.
02It wasn't just me — 84 out of 85
You cannot see inside a company from the street. But you can read what the company wrote down itself — the privacy policy (which must legally list the systems and processors handling customer data), the structure of its own malls, official notices, corporate registrations and filings. The sample: 85 consumer, F&B and hospitality companies, measured August–September 2026.
The counting is simple. For each company, we counted the paths along which one customer's records can split into separate ledgers:
- two or more own malls?
- two or more customer-management systems?
- multiple sales/booking channels?
- separate legal entities per brand?
- M&A history?
Fragmentation rate (%) = (companies with ≥1 path) ÷ 85 × 100. The result: 84 of 85 companies — 99%. Inside or outside, almost everyone.
Figure 1 · Share of each fragmentation path — 85 consumer/F&B/hospitality companies, public web records
| Path | Of 85 | What splits |
|---|---|---|
| Multiple own malls | 49 (58%) | The same customer gets a different member ID per mall (up to 9 malls) |
| Multiple sales/booking channels | median 4 channels (max 11) | Revenue made by channels and revenue made by people blur together |
| Separate legal entities | 30 (35%) | Ledgers cannot legally merge across brand entities |
| M&A history | 9 (11%) | Systems and DBs never merged — only the closing numbers did |
By customer-system count, 45 companies (53%) run two or more, and the maximum is 17 — at one hospitality group, written in its own privacy policy's processor list. When bookings arrive through six channels and customer records sit in seventeen systems, even "is this guest a repeat visitor?" has no answer.
We counted only what each company had written down about itself. Even that was enough to find fragmentation at 84 of 85.
A caveat up front: these figures are a strict floor. We counted only what is publicly written, so the fragmentation that never reaches a policy page — a manager's spreadsheet, a dealer's paper roster, the legacy DB a departed employee built — is not in here. The real thing runs deeper.
03Why "just buy a CDP" is wrong
If one system fixed this, this piece would not exist. The measurement says the opposite — fragmentation is not one layer. Count the paths per company: only one company in 85 has none, the median is two layers, and two in five (36 companies, 42%) carry three or more at once.
Figure 2 · Layers per company — the same 85. Zero layers: one company; two in five carry three or more
One system cannot fix it because each layer needs a different remedy. Multiple malls are an ID-mapping problem; separate legal entities are a consent-and-contract problem — a legal one; acquired brands are an excavation problem, finding the customer DB left behind in someone else's system. A system purchase solves the first layer only; with the rest intact you get the familiar state of "we integrated, but the numbers don't match."
Buying a system settles the first layer. Every layer after that asks for a decision before it asks for budget.
For the same reason, "our own team can just pull it" is only half right. For the first layer it is true — with someone who writes good SQL, mapping IDs across malls takes days. But the legal-entity layer does not yield to SQL. That one is a question of whether you have grounds to bring another entity's customer data over, and no amount of query skill creates grounds that aren't there. Acquisitions are the same: when nobody knows which system still holds the data, adding people to pull it changes nothing. Some layers are solved by adding hands; others require a decision before any hands are added. Start without telling them apart and you get "we've been integrating for six months."
And one more distinction must be drawn. There are two kinds of fragmentation. The four paths above are data you *could* join but haven't. Platform and open-market sales are a different kind. As the supplements piece put it — sell on your own mall and the buyer's ID, repeat cycle and drop-offs all land on your own server; sell on an open market and what you receive is a settlement statement. Who that customer is doesn't come late — it never comes. The larger your platform share, the more your own channels are the only window onto your customers — and the measurement above says that even that one window is fragmented at 99% of companies. Data you cannot receive is beyond help — which makes joining the data you do receive all the more urgent.
04What a fragmented ledger hides
The numbers aren't wrong — the questions stop being answerable.
Repeat rate breaks. Someone who buys on your mall and again on a marketplace storefront is two new customers in a split ledger. Repeat rate reads too low, new-customer inflow too high — and the marketing budget follows those wrong numbers into "acquisition."
Segmentation breaks. RFM's F is the act of deciding "what counts as one purchase" — with purchases scattered across three ledgers, there is nothing to count. A loyal customer sits in the books as three drifters.
Incrementality breaks. Whether an ad and a promotion hit the same customer is visible only after orders are joined per person; per-channel ledgers will never show it.
Churn is a definition, not an event. Especially for memberships and subscriptions. Most teams pull their churn list by expiry date — but an expiry date is a calendar, not a signal. People don't leave because the term ended; the term ends after they have already left. The real signal sits earlier, in behavior: visits spacing out, same-day cancellations rising, the usual weekday shifting. And that behavior is scattered across the app, the on-site POS and the consultation notes. If a member who books in the app and pays at the counter exists as two people in two ledgers, "time between visits" cannot be computed at all. It is exactly RFM's F again — deciding what counts as churn comes before any model.
And people are not the only thing that splits. Products split too. When the same SKU is registered under channel-specific spellings, options and bundles, neither "this product's repeat purchases" nor "this product's margin by channel" can be counted.
Finally, a newer casualty — there is no history for AI to use. To make an AI persona act as your customer, you have to hand it a line like "over the past year you bought this brand three times, most recently on discount." Aggregate rates will not open its mouth — a later piece in this series measures exactly that. The problem is that the line can only be written from a unified history. Split the ledgers, and there is no sentence to give.
Split ledgers do not hand you wrong numbers. They take the questions away — repeat rate, segments, incrementality, and the one line a persona needs.
05What happens when you join them
We have done it. Together with Gentle Monster, we joined orders that had been arriving separately per channel into per-customer histories, and built RFM segments on top. The result was simple — click rates and revenue lined up in the exact order the segments were cut. Not because the model was clever, but because "what counts as one purchase" could finally be decided on a unified ledger. A CRM campaign run on that same foundation drove own-mall revenue to 4.3× in six weeks (ROAS 5,370%, z-test verified) — not because the campaign was brilliant, but because for the first time we could see who *not* to send to.
The order matters. Unification comes before analysis. A model built on split ledgers is precisely wrong in proportion to its sophistication.
06So what does your company actually do
Two things — one check today, one experiment on Monday.
Today: open your company's privacy policy. Exactly this piece's method. If the processor/system list runs past five lines, that list is the map of how many pieces your customer data is in. No meeting needed; five minutes.
Monday: re-measure your repeat rate, once. Pull the last twelve months of orders from your top two channels and join them on a single identifier — a phone number is enough. Not a development project; a one-day job. The gap between that number and the repeat rate on your dashboard is the size of what fragmentation has been hiding from you. If the gap is small, relax and move on. If it is large — that gap is the ROI estimate of a unification project. Either way, you get the answer in a day.
Then read your company's row — each layer has a different first move.
| Your company | First layer | What you do |
|---|---|---|
| Runs several brand malls | ID mapping | Ask whether member DBs are separate per mall — if so, start with a phone-number mapping table |
| Sells through 3+ channels | One identifier | Check whether orders carry a channel-common customer ID — if not, fix collection first |
| Has separate entities per brand | Law & consent | Before any system: build the legal basis (consent, contracts) for moving customer data between entities |
| Has acquired brands | DB excavation | Find which system still holds the pre-acquisition customer DB — usually nobody knows |
07What this piece is really saying
This work is usually filed under "boring cleanup before the analysis." It is the opposite. That 84 of 85 companies are fragmented means joining is the scarcest work there is. Anyone can build a dashboard; everyone uses the same models. But confirming that the customer in three ledgers is one person is a re-architecture of the data — and the moment it lands, repeat rates, segments, incrementality and AI personas all come alive at once. That is what we sell.
Method — 85 consumer/F&B/hospitality companies, public web records measured 2026-08-14 to 09-03 (privacy policies, own-mall structure, notices, registrations/filings). Fragmentation rate = companies with ≥1 path ÷ 85 × 100. A floor based on public records; internal ledger fragmentation is not included. The sample is our research universe, not a random sample. Sample company names are not disclosed. Gentle Monster collaboration figures are as published on this site.
Same method, different industry. The bottleneck differs every time — finding that difference is what this series does.
One industry, one bottleneck, public data only — straight to your inbox. Nothing else, ever.
This is as far as public data can see. Repurchase, incrementality, attribution — the numbers that change decisions live inside your own data, and making them countable is what Lambency does.
caffrey.w.lee@gmail.com