Series

Browse the writing by series, in order.

Agentic AI in Practice

16 posts

A reusable enterprise playbook for shipping agents — from grading projects to pre-launch gates.

  1. #1AI Agent Autonomy Levels: I Audited 28 'Agent' Projects — Only 5 Passed L0–L3May 12, 2026
  2. #2Five Architecture Decisions That Determine Whether Your Customer-Service Agent Can Ship | Agentic AI in Practice (II)May 16, 2026
  3. #3Why a 70% Critic Beats a 95% Critic — A Fail-Closed Design Deep Dive | Agentic AI in Practice (III)May 17, 2026
  4. #4Deploy and Abandon — The Costliest Misconception in AI Agent Projects | Agentic AI in Practice (IV)May 18, 2026
  5. #5Containment Rate vs Resolution Rate: The Only Customer-Service AI Metric That Matters (How "98% CSAT" Gets Faked)May 22, 2026
  6. #6Agent Skills vs Knowledge Base: Why Stuffing SOPs Into RAG Doesn't Make an Agent CapableMay 24, 2026
  7. #7Pytest-Green Doesn't Mean Ship-Ready: How to Actually Test an AI Agent (Dual-Track)May 25, 2026
  8. #8Intent Classification for Chatbots: Why Pure-Rule and Pure-LLM Both Fail (a 3-Tier Cascade)May 25, 2026
  9. #9Don't Let AI Agents Call APIs Directly — A 5-Layer Tool-Calling Stack + 25-API Contract ChecklistMay 27, 2026
  10. #10Corpus Drives Codebook — Why Your Intent Taxonomy Is Stuck at 60% and How It Evolves from 36 to 48 | Agentic AI in Practice (X)May 28, 2026
  11. #11LLM Fact-Checking with a Verifier Agent: 11 of 34 Facts in a 25-Page AI Plan Were FabricatedMay 29, 2026
  12. #12The Org Chart Is the Real Architecture Diagram — 90% of Stalled Agent Projects Aren't a Tech Problem | Agentic AI in Practice (XII)Jun 1, 2026
  13. #13Self-Serve Rate ≠ Correct Rate — The Gates a Customer-Service Agent Must Clear Before Launch | Agentic AI in Practice (XIII)Jun 2, 2026
  14. #14Your Code Is Fixed. Production Isn't.Jul 5, 2026
  15. #15Your Knowledge Center Is a Search BoxJul 6, 2026
  16. #16One Brain, Two Front-Ends: Refunded Twice on Sale DayJul 6, 2026

After Launch

6 posts

Launch isn't the finish line — it's where the machine starts to turn. From how to draw eval samples, how many to label, and whether the labels are reliable, to judging launch, guarding against silent drift, and spinning up the data flywheel — six posts to make an agent that gets up, stays up, and gets sharper with use.

  1. #1The 96% Turned Red the Moment We Split It by Channel — Two Fatal Failures of Eval SamplingJul 2, 2026
  2. #250 Rows at 96%, Ship It? Size Labeling by the CI Lower Bound, Not the Pretty NumberJul 3, 2026
  3. #3Your Labels Are Your Ceiling — One "Swap Half a Size Up" Gets Three Answers From Two AgentsJul 4, 2026
  4. #4A Pretty Accuracy Number Hid Dozens of Money-Moving Errors — How to Read the Eval to ShipJul 5, 2026
  5. #5Your Dashboards Are Green While the Agent Quietly Gets Dumber — Post-Launch Silent DriftJul 5, 2026
  6. #6Make the Agent Get Sharper With Use, Not Dumber — Spinning Up the Data FlywheelJul 5, 2026

Retail Agentic AI Handbook

3 posts

A three-part field guide to shipping agents in retail: pick scenarios, build the KB, run ops.

  1. #1Which 4 of Your 28 'Smart-X' AI Agent Projects to Start With — Retail Agentic AI Handbook (Part 1)Feb 28, 2026
  2. #2AI Agent Knowledge Base: 3-Layer Design + 4-Bucket Cost EstimateFeb 28, 2026
  3. #380% of Failed AI Agents Die in Ops, Not Tech — Post-Launch Loop, Safety Layer & 30-Day Monitoring PlanFeb 28, 2026

Research

2 posts

My academic papers, unpacked: from reward hacking to catastrophic forgetting — how the experiments were designed and what they mean for engineering.

  1. #1Reward Hacking in AI Agents: Trained 60,000 Steps, the Agent Learned to Delete Tickets (6 ITSM Patterns)Oct 10, 2025
  2. #2Catastrophic Forgetting in LLMs: 52 Domains Fine-Tuned, the Earlier 51 Regressed — A Dual-Replay Field ReportOct 13, 2025

Case Studies & Deep Dives

4 posts

Engineering deep dives paired with real production incidents: refund workflows, KB desync, event-loop blocking, Claude Code token attribution.

  1. #1What a Real, Money-Moving L2 Refund Workflow Actually Looks Like | Workflow Deep DiveJun 3, 2026
  2. #2Your KB Changed. The Search Index Didn't — Anatomy of a 9-Day Silent Desync | KB-Ops Deep DiveJun 3, 2026
  3. #3Your Observability Dashboard Is Throttling the Agent It Watches — an Async Latency PostmortemJun 10, 2026
  4. #4I Ran the Whole Token-Saving Playbook. The Savings Got Re-Spent.Jul 13, 2026

WeChat's AI Entry: Rebuilding the Merchant Agent

3 posts

WeChat is beta-testing an AI agent that can search, compare, order, and pay across 10M merchants. Three posts: how to get found, how to keep your customers, how to rebuild.

  1. #1AI Picked the First Store and the Other Four Vanished: WeChat's New Shelf for 10 Million MerchantsJun 29, 2026
  2. #2When the AI Becomes the Storefront, You Decay Into a Supplier: The Relationship Moat for 10 Million MerchantsJun 30, 2026
  3. #3The Day WeChat's AI Ordered for Users, 10 Million Merchants' Mini-Programs Expired — A Layered Rebuild BlueprintJun 24, 2026

Translations

15 posts

Curated translations of standout external writing — original source and translator always credited (not original work).

  1. #1创始人手册:打造 AI 原生初创公司May 18, 2026
  2. #2Fable 实战指南:找到你的未知Jul 5, 2026
  3. #3用 Fable 5 搭一个第二大脑Jul 5, 2026
  4. #4Loop Engineering:Karpathy 方法论,和那篇让它再快 5 倍的研究Jul 6, 2026
  5. #5用 Fable 5 搭一个自我改进的 agent 系统:14 步,loops、动态工作流、routinesJul 6, 2026
  6. #6四种 Loop 分别什么时候用:Claude Code 官方的 loop 入门指南Jul 8, 2026
  7. #7反向信息悖论:在 AI 时代,先泄密的是买方Jul 14, 2026
  8. #8用 12 步造一个会拒绝自己幻觉的研究 agent:搜索循环、来源打分、事实闸Jul 15, 2026
  9. #9企业财务的 AI 该怎么落地:别买平台,给「中间那个人」装个后台 agentJul 15, 2026
  10. #10让沙子学会思考之后:Demis Hassabis 的前沿 AI 治理框架Jul 15, 2026
  11. #11温度计与恒温器:让 eval 真正改变下一步的那一层Jul 29, 2026
  12. #12Harness / Loop / Graph:agent 系统坏了,先分清坏在哪一层Jul 29, 2026
  13. #13Loop Engineering 是一种模式,不是一个功能Jul 29, 2026
  14. #14让一个 Obsidian 研究系统替你读完这周的资料Jul 29, 2026
  15. #15记忆工程:让 agent 的记忆越用越聪明,而不是越用越臃肿Aug 3, 2026