Numbers carry provenance.
숫자에는 출처가 있다.
∴ Every number ships with where it came from and when. If it can't be traced back, it doesn't ship — no matter how good it looks.
∴ 모든 숫자에는 출처와 날짜가 붙습니다. 어디서 나온 건지 되짚을 수 없는 숫자는, 아무리 좋아 보여도 내보내지 않습니다.
Data & AI Consulting · Seoul데이터·AI 컨설팅 · 서울
lam·ben·cyn.— light that grazes a surface without burning it, and in passing, reveals its shape.— 태우지 않고 스치는 빛. 닿는 자리마다 형태가 드러난다.
Analysis should work the same way.분석도 그래야 합니다.
Vires acquirit eundo
01 — De Me · About· 소개
A question changes a little every time it changes hands. The one who hears the problem, the one who writes the query, the one who builds the model, the one who explains it — by the fourth pair of hands, the answer often no longer fits the question. At Lambency, those four people are one person. The one who hears your problem writes the SQL, fits the model, ships the dashboard — and stands behind the number when it counts.
질문은 손을 탈 때마다 조금씩 다른 질문이 됩니다. 문제를 들은 사람, 쿼리를 짜는 사람, 모델을 만드는 사람, 그걸 설명하는 사람 — 네 번째 손에 닿을 즈음이면, 답은 처음의 질문과 어긋나 있기 일쑤입니다. Lambency에서는 그 네 사람이 한 사람입니다. 문제를 들은 사람이 직접 SQL을 짜고, 모델을 세우고, 대시보드를 넘깁니다. 그리고 그 숫자가 시험대에 오르는 순간까지 책임집니다.
Numbers carry different weights. Some, when wrong, make a dashboard look slightly off; others move a P&L, a city policy — sometimes a patient. The last ten years were spent with the second kind. Hospital operations data in Silicon Valley, COVID-era clinical studies and a cyclist-injury prediction model for the City Council in New York, 1.95M retail customers in Seoul, a Kyoto short-stay portfolio grown from 5 homes to 17. These days, LLM agents go into production pipelines, not demos. And two foundings left one conviction: an analysis that cannot answer so what do we do might as well not exist. So every engagement starts from the decision, and works backward.
숫자에도 무게가 있습니다. 틀려도 대시보드가 조금 어색해지고 마는 숫자가 있고, 틀리면 손익이, 도시의 정책이 — 때로는 환자가 흔들리는 숫자가 있습니다. 지난 10년은 후자의 자리에서 보냈습니다. 실리콘밸리의 병원 운영 데이터, 뉴욕의 COVID 임상연구와 시의회가 쓴 자전거 사고 예측 모델, 서울의 고객 195만 리테일, 그리고 5채에서 17채가 된 교토의 숙박 자산까지. 지금은 LLM 에이전트를 데모가 아니라 실제 파이프라인에 넣습니다. 그리고 두 번의 창업이 남긴 결론은 하나입니다. 그래서 뭘 해야 하나에 답하지 못하는 분석은, 없는 것과 같습니다. 그래서 모든 일을 내려야 할 결정에서 거꾸로 시작합니다.
Measured numbers only. Where the data can't answer, you'll hear I don't know — along with what it would take to find out. An analysis that can admit what it doesn't know is the only kind worth trusting when it says it does.
실측한 숫자만 씁니다. 데이터가 답을 못 하는 자리에서는 모른다고 말씀드립니다 — 무엇을 확보하면 알 수 있는지까지 함께. 모른다고 말할 수 있어야, 안다는 말도 믿을 수 있기 때문입니다.
02 — Opera · Practice· 분야
Scattered customer data, joined until you can see where revenue actually comes from. RFM, behavioral segments, cohort LTV. LTV never arrives as a single number — a point estimate looks precise and can't be planned against. You get a range, with the uncertainty attached.
흩어져 있는 고객 데이터를 하나로 붙여, 매출이 어디서 나오는지 보이게 만듭니다. RFM, 행동 세그먼트, 코호트 LTV까지. LTV는 “1인당 37만원”처럼 숫자 하나로 딱 잘라 드리지 않습니다. 정밀해 보이지만 그 숫자로는 계획을 못 세웁니다. 대신 오차를 붙인 범위로 드립니다.
"Did that actually work?" — answered with statistics, not enthusiasm. Power calculated before launch, empirical-Bayes shrinkage when segments thin out, holdouts kept clean, and quasi-experimental designs when randomizing isn't on the table.
“그래서 그게 효과가 있긴 했나요?” 이 질문에 느낌이 아니라 숫자로 답할 수 있게 만듭니다. 시작 전에 검정력부터 계산하고, 세그먼트가 얇아지면 empirical Bayes로 눌러주고, 홀드아웃이 오염되지 않게 지킵니다. 무작위 배정이 불가능한 상황이면 준실험으로 설계합니다.
Agents wired into the pipelines you already run, drafting the first pass of analysis so your team starts from a draft instead of a blank page. Shipped with an evaluation harness — "we tried it and it seemed fine" is not QA.
이미 돌아가고 있는 파이프라인에 에이전트를 붙여 분석 초안을 뽑게 합니다. 팀이 빈 화면이 아니라 초안에서 시작하게 하는 게 목적입니다. 넘길 때는 평가 하네스를 같이 드립니다. “돌려보니 잘 되던데요”는 검증이 아니니까요.
The unglamorous layer everything above sits on. One warehouse, versioned transformations, metrics defined once and reused everywhere — so two dashboards never show different revenue in the same meeting again.
눈에 잘 안 띄지만 위의 세 가지가 전부 여기에 얹힙니다. 데이터를 한 곳에 모으고, 변환 로직은 버전으로 관리하고, 지표는 한 번만 정의해 어디서든 같은 걸 쓰게 합니다. 회의에서 두 대시보드가 서로 다른 매출을 띄우는 일이 없어집니다.
03 — Modus · A Typical Engagement· 진행 방식
First, name the decision. Then audit the data that is supposed to support it. What can and cannot be answered at your current data quality is stated at the start, not at the end. The deliverable is a measurement plan, not a proposal deck.
먼저 무슨 결정을 내려야 하는지부터 정합니다. 그다음 그 결정을 뒷받침한다는 데이터를 열어봅니다. 지금 상태로 답할 수 있는 것과 없는 것은 끝이 아니라 시작할 때 말씀드립니다. 나오는 건 제안서가 아니라 측정 계획입니다.
Pipelines, models and experiments land in your repositories from the first week. Metrics are versioned, assumptions are written down, and estimates carry their intervals.
파이프라인·모델·실험은 첫 주부터 고객사 저장소에 커밋합니다. 지표는 버전으로 관리하고, 가정은 빠짐없이 적어두고, 추정치에는 신뢰구간을 답니다.
We operate it together, then I hand it over: documentation, runbooks, and a team that can run it without me. Becoming unnecessary is part of the engagement, by design.
한동안 같이 굴려본 다음 넘깁니다. 문서와 런북, 그리고 저 없이도 돌릴 수 있는 팀까지. 제가 없어도 돌아가게 만드는 것까지가 이 일의 결과물입니다.
04 — Principia · How I Work· 일하는 방식
∴ Every number ships with where it came from and when. If it can't be traced back, it doesn't ship — no matter how good it looks.
∴ 모든 숫자에는 출처와 날짜가 붙습니다. 어디서 나온 건지 되짚을 수 없는 숫자는, 아무리 좋아 보여도 내보내지 않습니다.
∴ You hear the limits before the findings. "I don't know" is a legitimate answer — and it comes with what it would cost to find out.
∴ 결론보다 한계를 먼저 말씀드립니다. “모른다”도 답입니다. 다만 거기서 끝내지 않고, 알아내려면 무엇이 얼마나 필요한지까지 붙여서 드립니다.
∴ Recommendations arrive as code that runs, not slides that don't. The engagement ends; the pipelines, models and documentation stay with your team.
∴ 조언은 문서가 아니라 돌아가는 코드로 넘깁니다. 프로젝트가 끝나도 파이프라인·모델·문서는 그대로 남습니다. 전부 그쪽 자산입니다.
Our team's data literacy has gone up.
우리 팀의 데이터 리터러시가 확 올라갔어요.
— a CRM client, a year into the engagement. Still the review that means the most. — 어느 CRM 클라이언트, 프로젝트 1년 차. 지금까지 들은 말 중에 제일 좋았습니다.
05 — Acta · Selected Work· 주요 사례
Experience includes함께한 곳
Lambency Engagement
Lambency 프로젝트
The metric this group's flagship brand actually needed was nowhere in its existing data. So we reframed the question together — data owners through decision-makers — and rebuilt the pipeline end to end, from instrumentation to modeling, across four markets: Korea, the U.S., Japan and Australia. It became a core CRM model and extended to the group's sister brands. The engagement was renewed every quarter for over a year.
이 그룹의 대표 브랜드가 정작 필요한 지표를 기존 데이터로는 뽑을 수 없었습니다. 그래서 데이터 담당자부터 의사결정권자까지 함께 앉아 질문을 다시 세웠습니다. 한국·미국·일본·호주 네 시장에 걸쳐 계측부터 모델링까지 파이프라인을 처음부터 다시 만들었습니다. 핵심 CRM 모델로 자리 잡았고 그룹 자매 브랜드로 확장됐습니다. 이 프로젝트는 1년 넘게 분기마다 갱신됐습니다.
Scope — KR · US · JP · AU · multi-market lifecycle & repurchase modeling
범위 — KR · US · JP · AU · 멀티마켓 라이프사이클 & 재구매 모델링
2024 — 2026 · Data Hub Lead
2024 — 2026 · 데이터 허브 리드
Ran the data hub inside a Series A customer-data platform — 1.95 million customers, 256 brands. RFM segments, collaborative filtering, Customer360, and a proprietary user–brand affinity index. LLM agents drafted the first pass of analysis. The most significant change was not technical: dashboards stopped being something reviewed before a meeting and became the basis for decisions made in it.
고객 195만 명, 브랜드 256개가 들어오는 시리즈 A 고객 데이터 플랫폼에서 데이터 허브를 맡았습니다. RFM 세그먼트, 협업 필터링, Customer360, 그리고 사용자와 브랜드가 얼마나 맞는지 재는 지수를 직접 만들었습니다. 분석 초안은 LLM 에이전트가 뽑게 했습니다. 가장 크게 달라진 건 기술이 아니었습니다. 대시보드가 회의 전에 열어보는 자료에서, 회의 중에 결정을 내리는 근거로 바뀌었습니다.
Stack — BigQuery · Python · LLM agents · reverse ETL
스택 — BigQuery · Python · LLM 에이전트 · reverse ETL
D2C F&B · Retail · Measurement Showcase
D2C 식음료 · 리테일 · 측정 쇼케이스
Two domestic D2C brands — one in food & beverage, one in lifestyle accessories. Audiences built with RFM and clustering, then tested against both broad and purchase-based targeting. Two-proportion z-tests, bootstrap resampling, Beta-Binomial conversion models. The F&B brand grew own-mall revenue 4.3× in six weeks; the accessories brand reached 16.5% CTR on its first campaign. Results of that magnitude invite the question of whether they were luck. The control group was held clean and the effect was tested against the statistics — and it survived.
국내 D2C 브랜드 두 곳, 식음료와 라이프스타일 잡화였습니다. RFM과 클러스터링으로 오디언스를 뽑고, 광범위 타겟팅과 구매 기반 타겟팅 양쪽에 붙여 비교했습니다. 2비율 z-검정, 부트스트랩 리샘플링, Beta-Binomial 전환 모델까지 돌렸습니다. 식음료 브랜드는 6주 만에 자사몰 매출이 4.3배가 됐고, 잡화 브랜드는 첫 캠페인에서 CTR 16.5%가 나왔습니다. 이 정도 배수가 나오면 보통 운이 좋았나 싶어집니다. 그래서 대조군을 오염 없이 유지한 채 통계로 눌러봤고, 그러고도 남았습니다.
Method — z-test · bootstrap · Beta-Binomial CVR · holdout
방법 — z-검정 · 부트스트랩 · Beta-Binomial CVR · 홀드아웃
Public Sector · Clinical Research
공공 · 임상 연구
A predictive model of cyclist killed-or-severely-injured intersections, built with Columbia University for the New York City Council and folded into Vision Zero road-safety policy. Power analysis and validation protocols for COVID-era clinical studies at Richmond University Medical Center. The domains differ; the discipline doesn't. When a number decides policy or patient care, the bar for evidence only goes up.
자전거 사망·중상 교차로를 예측하는 모델을 컬럼비아 대학교와 함께 뉴욕시의회를 위해 만들었고, Vision Zero 도로안전 정책에 반영됐습니다. 리치먼드 대학병원에서는 코로나 시기 임상 연구의 검정력 분석과 검증 프로토콜을 맡았습니다. 분야는 달라도 지켜야 하는 건 같습니다. 숫자가 정책이나 환자 치료를 좌우하는 자리에서는, 증거의 기준이 더 높아질 뿐입니다.
Scope — NYC Vision Zero · clinical statistics · trust & safety
범위 — NYC Vision Zero · 임상 통계 · 신뢰·안전(T&S)
Hospitality · Kyoto — Family Office
호스피탈리티 · 교토 — 패밀리 오피스
A family office running short-stay properties in Kyoto wanted dynamic pricing. The market data to price against did not exist — no one held it. So the first task was to build the dataset: competitors' nightly rates and booking patterns, assembled into something the market itself lacked. That priced the calendar. It also answered a question no one had asked — which property configurations were not worth building at all.
교토에서 단기 숙박을 운영하는 패밀리 오피스가 다이내믹 프라이싱을 하고 싶어 했습니다. 그런데 기준으로 삼을 시장 데이터가 없었습니다. 아무도 갖고 있지 않았습니다. 그래서 데이터를 만드는 일부터 했습니다. 경쟁 숙소의 1박 요금과 예약 상황을 모아, 시장에 존재하지 않던 데이터셋으로 엮었습니다. 그걸로 날짜별 단가를 매겼고, 아무도 묻지 않았던 답도 하나 나왔습니다. 어떤 구성의 집은 애초에 지을 가치가 없다는 것.
Scope — market construction · site selection · property design · dynamic pricing
범위 — 시장 데이터 구축 · 입지 발굴 · 숙소 설계 · 다이내믹 프라이싱
Also — two R packages on CRAN (getDTeval, a top-50 package; formulaic) · a comp-based real-estate pricing engine, in closed B2B beta.
그 외 — CRAN에 등록한 R 패키지 둘 (getDTeval, top-50 · formulaic) · 비교사례 기반 부동산 가격 산정 엔진, B2B 클로즈드 베타.
06 — Epistula · Contact· 문의
Tell me what you're trying to decide, and why you haven't been able to. You'll have a reply within two business days — from me, because there's no one else here.
무엇을 정해야 하는지, 그리고 왜 아직 못 정하고 있는지 적어주세요. 이틀 안에 답장 드립니다. 여기는 저 혼자라 제가 직접 읽고 제가 씁니다.
Write directly바로 메일 주세요
caffrey.w.lee@gmail.comThree things make the first reply far more useful: what you need to decide, what's blocking it, and by when.
이 세 가지가 있으면 첫 답장이 훨씬 쓸모 있어집니다. 정해야 할 것, 막고 있는 것, 그리고 언제까지인지.
One person reads this. It goes nowhere else.
저 말고는 아무도 안 봅니다. 다른 용도로도 쓰지 않습니다.