一盏灯
首页文章视频生词本复习
首页文章视频生词本我的

When the Dashboard Lies: Making Decisions with Data

当仪表盘撒谎时:真正的数据驱动决策

科技互联网深度阅读高级约 5 分钟场景 · decision making# 数据与分析

人人都说自己数据驱动,但多数团队只是用图表装点早已做好的决定。真正的功夫在于:识破虚荣指标、区分相关与因果、警惕『指标变成目标就会失真』的古德哈特定律、尊重统计常识与新鲜感效应、让定量与定性互为表里、向高管汇报时先讲决定并如实交代不确定性——以及,知道什么时候应该推翻数据。

当前浏览器暂不支持语音朗读

Every company now describes itself as data-driven, in the same way every company describes itself as innovative — the phrase costs nothing and commits to less. Watch how decisions actually get made, though, and a familiar theatre appears: a leader forms a view, an analyst is dispatched to find supporting numbers, and a slide deck arrives decorated with charts that all point, miraculously, in the direction the leader already chose. That is not data-driven decision-making. It is decision-driven data-making, and the difference between the two determines whether your metrics are instruments or ornaments.

如今每家公司都自称数据驱动,就像每家公司都自称创新一样——这个说法不花一分钱,也不承诺任何事。但只要观察决策实际是如何做出的,一出熟悉的戏码就会浮现:领导先形成了看法,分析师被派去寻找支持性的数字,然后一份幻灯片如期而至,上面的图表全都奇迹般地指向领导早已选定的方向。这不是数据驱动的决策,这是决策驱动的数据,而这两者的区别,决定了你的指标究竟是仪器,还是装饰品。

The first discipline is telling vanity metrics from actionable ones. Cumulative registered users is the classic vanity metric: the chart can only ever go up and to the right, it flatters every all-hands meeting, and it cannot inform a single decision. Weekly active users who complete a core action, retention by signup month, revenue per customer — these are actionable, because they can go down, and when they do, they tell you where to look. A useful test for any number on your dashboard: if this metric dropped by a third tomorrow, would we do anything differently? If the honest answer is no, it is decoration.

第一门功夫,是分辨虚荣指标与可行动指标。累计注册用户数是经典的虚荣指标:这条曲线永远只会向右上方走,能给每一场全员大会贴金,却无法为任何一个决策提供信息。每周完成核心动作的活跃用户、按注册月份切分的留存率、单客收入——这些才是可行动的指标,因为它们会跌,而且跌的时候会告诉你该往哪里看。检验仪表盘上任何数字的实用标准是:如果这个指标明天跌去三分之一,我们会做出任何不同的动作吗?如果诚实的答案是『不会』,它就是装饰。

The second discipline is refusing to let correlation impersonate causation. Your analysis shows that users who enable the mobile app retain twice as well, so someone proposes forcing app installation at signup. But the app did not necessarily cause the retention; your most committed users may simply be the ones who bother to install apps. The only reliable way to separate the two is a controlled experiment — show the feature to a random half of new users and compare cohorts. Where experiments are impossible, at least say the honest sentence aloud: "these move together, and we do not yet know why." That sentence has saved companies millions.

第二门功夫,是拒绝让相关性冒充因果性。分析显示,启用了移动端 App 的用户留存率高出一倍,于是有人提议在注册时强制安装 App。但留存未必是 App 带来的:也许你最投入的那批用户,本来就是愿意费心装 App 的人。可靠区分两者的唯一办法是对照实验——把功能随机开放给一半新用户,然后比较两组人群。在无法做实验的场合,至少要把那句诚实的话说出口:『这两者一起变动,但我们还不知道为什么。』这句话为不少公司省下过数以百万计的学费。

Third, remember Goodhart's law: when a measure becomes a target, it ceases to be a good measure. Reward the support team for closing tickets within an hour, and tickets will close within an hour — resolved or not, because agents learn to close and reopen. Pay sales on signed contracts and revenue will arrive, trailed by churn twelve months later. None of this is dishonesty; it is people rationally optimising what you chose to count. The defence is to pair every target with a guardrail metric that catches the distortion: ticket closure time paired with customer satisfaction, contracts signed paired with second-year renewal.

第三,记住古德哈特定律:一旦某个度量成为目标,它就不再是好的度量。奖励客服团队一小时内关单,工单就会在一小时内被关掉——不管问题解决没解决,因为客服会学会先关再重开。按签约额给销售发提成,收入就会到账,而十二个月后跟着到来的是客户流失。这些都不是不诚实,而是人们在理性地优化你选择去计量的东西。防御之道,是给每个目标配一个护栏指标,专门捕捉变形:关单时长配客户满意度,签约数配次年续约率。

Fourth, respect the statistics you learned and then forgot. An A/B test with two hundred users per branch will "prove" almost anything if you stare at it long enough. Peeking at results daily and stopping the moment significance appears is the most common way teams manufacture false wins. So is ignoring the novelty effect: any visible change lifts engagement for a week, because users poke at whatever moved. Decide the sample size and the test duration before launch, write down the success threshold, and let the experiment finish. Discipline agreed in advance is the only known cure for wishful reading.

第四,尊重那些你学过又还给老师的统计学。每组只有两百名用户的 A/B 测试,只要你盯得够久,几乎可以『证明』任何结论。每天偷看结果、显著性一出现就立刻停止实验,是团队制造虚假胜利最常见的方式。忽视新鲜感效应也一样:任何看得见的改动都会让参与度升高一周,因为用户会去戳一戳任何变了样的东西。要在上线前定好样本量和实验时长,白纸黑字写下成功阈值,然后让实验跑完。事先约定的纪律,是治疗『许愿式读数』唯一已知的药方。

Numbers tell you what is happening; they are strangely silent on why. The retention chart shows users leaving in week two, but it took five customer interviews to learn the actual reason: the export feature they needed was hidden behind an unlabelled icon. This is why mature teams pair quantitative data with qualitative work — session recordings, support transcripts, open-ended interviews. An anecdote is not evidence, but it is an excellent hypothesis machine: the interview suggests the theory, the experiment tests it, the dashboard confirms the fix. Teams that use only one of these instruments are flying with one eye closed.

数字告诉你正在发生什么,却对『为什么』出奇地沉默。留存曲线显示用户在第二周流失,但真正的原因是靠五次客户访谈才问出来的:他们需要的导出功能,藏在一个没有文字标注的图标后面。这就是成熟团队让定量数据与定性工作配对的原因——录屏回放、客服对话记录、开放式访谈。轶事不是证据,但它是绝佳的假设生成器:访谈提出理论,实验检验理论,仪表盘确认修复效果。只用其中一种仪器的团队,等于闭着一只眼睛开飞机。

When you carry data into the boardroom, structure matters as much as substance. Lead with the decision you are asking for, then the two or three numbers that bear on it, then the confidence level and the caveats — in that order. Executives do not need your forty-slide methodology; they need to know what you recommend, how sure you are, and what would change your mind. Above all, resist the temptation to trim the caveats that weaken your case. Presenting the number that hurts your own argument is precisely what makes people trust the numbers that help it; credibility, once spent, does not refresh with the next quarter.

带着数据走进董事会时,结构与内容同样重要。先讲你请求拍板的决定,再讲与之相关的两三个数字,然后是置信程度和注意事项——按这个顺序。高管不需要你四十页的方法论,他们需要知道你的建议是什么、你有多确定、以及什么会让你改变主意。最重要的是,抵制住删掉不利于自己论点的那些保留条件的诱惑。敢于展示伤害自己论证的数字,恰恰是别人愿意相信那些支持你的数字的原因;信用一旦透支,不会随下个季度自动刷新。

Finally, know when to overrule the dashboard. Data describes the world that already exists; it is structurally conservative. No spreadsheet in 2007 argued for a phone without a keyboard, and no retention metric will justify the bet whose payoff sits three years out. When you enter a new market, ship a category-creating product, or make any call where the feedback loop is longer than your planning cycle, data thins out and judgment must carry the weight. The goal was never to remove human judgment from decisions — it was to stop dressing judgment up as certainty. Be data-informed, brutally honest about which one is speaking, and you will beat both the gut-only romantics and the spreadsheet-only bureaucrats.

最后,要知道什么时候该推翻仪表盘。数据描述的是已经存在的世界,它在结构上是保守的。2007 年没有任何一张电子表格会支持做一部没有键盘的手机,也没有任何留存指标能为一场三年后才见分晓的押注辩护。当你进入新市场、推出开创品类的产品,或做任何反馈周期长于规划周期的决策时,数据会变得稀薄,判断力必须挑起重担。数据驱动的目标从来不是把人的判断从决策中清除出去,而是不再把判断打扮成确定性。做一个数据启发的决策者,并对『此刻是谁在说话』保持残酷的诚实——你就能同时胜过只信直觉的浪漫派,和只信表格的官僚派。