Jul 5, 2026A Pretty Accuracy Number Hid Dozens of Money-Moving Errors — How to Read the Eval to ShipOn a money-moving project I ran, the overall accuracy looked great; but pull the money-moving intents out on their own and the wrong-action rate was alarming — dozens of money-touching errors sat there the whole time, hidden by one blended number. In 5 minutes you'll see through "one accuracy figure to request launch"; in 10 you'll put a separate wrong-action gate on money-moving errors; in 20 you'll have a launch-decision flow: CI lower bound + per-scenario version cut + per-channel ramp.After Launch · 13 min read · EN · 中
Jul 3, 202650 Rows at 96%, Ship It? Size Labeling by the CI Lower Bound, Not the Pretty Number"We labeled 50, 96% correct — ship it?" No — the statistical lower bound is only 86%. When I designed the labeling gate for a customer-service agent, the first rule I pinned was: read the confidence-interval lower bound, not the point estimate. In 5 minutes you'll see through the small-sample 96% mirage; in 10 you'll have a "true accuracy → rows to label" table that explodes near the threshold; in 20, a CI stopping rule — label more where it's needed, don't waste it where it's done, and call a stop when the answer is fix-the-agent.After Launch · 12 min read · EN · 中