Article

The Metric Improved. The Business Didn't.

Why optimizing one variable in a nonlinear system quietly breaks it — the dating site that A/B-tested itself into a slower funnel, and what that means for anyone tuning a system.

A dating site ran an A/B test. The hypothesis was reasonable, almost boring: longer user profiles give other users more to react to, so engagement should go up. They shipped the longer-profile variant to half the users, watched the dashboard, and the dashboard agreed. Clicks went up. Messages went up. Time on site went up. The team rolled it out to everyone and moved on to the next experiment.

Months later, somebody finally looked at the metric the dashboard wasn’t showing: the rate at which people who joined the site actually went on a date with someone. That number had quietly fallen off a cliff.

The story is from Bill Gurley — Benchmark partner, board member at the Santa Fe Institute, the institute that more or less invented complexity theory as an academic discipline. He uses it as a parable for what he calls “single-variable determinism” in systems that are anything but single-variable.

Here’s the mechanism, because I didn’t fully understand it the first time I heard it either. Short profiles are information-poor. You read a few lines, you don’t have enough to form a judgment, so you message the person to find out more. Longer profiles are information-rich. You read three paragraphs about hiking and a complicated relationship with cilantro, and now you do have enough to form a judgment — and the judgment is usually “no.” The variable the team optimized (engagement with a profile) and the variable that actually mattered (conversion to a real-world date) were not just different. They moved in opposite directions. More information upstream killed the downstream behavior the whole product existed to produce.

This is the shape of the trap, and it generalizes far beyond dating sites.

The second-order trap

In a linear system, you optimize a variable, the variable improves, the system improves. Cause and effect are local. You can A/B test your way to a better outcome.

In a nonlinear system — which is most systems involving humans, markets, ecologies, or any feedback loop — variables interact. Pushing harder on one input can push the system into a different regime where the rules quietly change. The first derivative (the metric you’re watching) goes the way you wanted. The second derivative (the rate of change of something else in the system, often the thing you actually care about) goes the other way. By the time you notice, you’ve been wrong for months.

A few examples to make this concrete:

  • Restaurant reviews. Sites that show more reviews per restaurant get more time-on-page. They also get fewer bookings, because past a certain density of opinions, the visitor’s uncertainty increases rather than decreases — there’s always one bad review, and now you’ve seen it.
  • Content moderation. Platforms that aggressively remove low-quality posts get cleaner feeds and lower engagement. Engagement was partially powered by people arguing about the low-quality posts. The visible metric (post quality) improved. The hidden metric (network effects) eroded.
  • Performance reviews. Companies that tie compensation to a single measurable output get more of that output and less of everything the output depended on but wasn’t measured. Sales teams hit their number by burning future pipeline. Engineering teams ship features by accumulating bugs. The dashboard is green. The system is rotting.

The pattern is always the same: a variable that correlates with the goal under normal conditions becomes the target. Once it’s the target, it stops correlating — because the system reorganizes itself around the new pressure. Goodhart’s Law is the famous one-line version: when a measure becomes a target, it ceases to be a good measure. But Gurley’s framing is more useful than Goodhart’s, because Goodhart sounds like a warning about gaming. The dating-site team wasn’t gaming anything. They were honestly trying to make the product better and used the best tool the industry has for that — controlled experimentation. The tool failed because the system underneath it was nonlinear, and the tool assumes linearity.

Why this is the hardest class of error to catch

A bug crashes the program. A bad deployment pages someone at 3 a.m. A regression shows up in the test suite. These errors are loud.

The second-order trap is silent. Every metric you instrumented says you won. You can’t catch it with more testing of the same kind — running the A/B test longer wouldn’t have helped the dating site, because the conversion damage was downstream of a behavior change that only emerged at scale, over months. You can’t catch it with cleaner data, because the data is clean. You catch it only by stepping back and asking: which variables in this system did I assume were independent, and what happens if they aren’t?

That question is uncomfortable, because the answer is usually “I have no idea, and finding out is expensive.” Which is why almost nobody asks it. It’s easier to celebrate the dashboard.

What complexity theory has been saying for forty years

The Santa Fe Institute exists because a group of physicists, biologists, and economists noticed in the 1980s that the systems they cared about — ecosystems, economies, immune systems, cities — kept producing behavior that their disciplines’ linear models couldn’t predict. Stuart Kauffman, Murray Gell-Mann, Brian Arthur, and others built a vocabulary for it: emergence, path dependence, phase transitions, power laws. The unifying claim is that systems with many interacting components don’t decompose cleanly into their parts. The whole has properties the parts don’t. Optimizing the parts can degrade the whole.

This is also, incidentally, why I find the physics training useful in markets. A physicist who’s studied phase transitions has already internalized the idea that a system can behave one way for a long stretch and then, at some threshold, flip into a regime where the old rules don’t apply. Water doesn’t get gradually more boiled. It’s liquid, liquid, liquid — then steam. The variable you were watching (temperature) doesn’t tell you what just happened to the system. You need a second variable (phase) and the knowledge that the relationship between them is discontinuous.

Most A/B testing implicitly assumes you’re in the “liquid” phase the whole time. The dating site found out, the hard way, that nudging one variable hard enough can push the whole product across a phase boundary.

The honest application

None of this means you should stop measuring things. (That would be the other failure mode — paralysis by complexity, where every decision becomes “well, who really knows.”) The point is narrower and more practical.

When you’re tuning a system you don’t fully understand — and you almost never fully understand the systems you care about — be suspicious of single-variable wins. Ask which other variables in the system might move in response, especially the ones you can’t easily measure. Hold the metric loosely. Treat the dashboard as a hypothesis, not a verdict. And when something works, ask why it works before you scale it, because the mechanism matters more than the result. A win you don’t understand is a win you can’t defend when the regime changes.

The dating site eventually figured this out and rolled the change back. But “we eventually figured it out” is the expensive version of the lesson. The cheap version is: the metric improved, and that should make you nervous, not relaxed. In a linear system, an improving metric is good news. In a nonlinear one, it might be the first signal of a regime shift you’d rather have seen coming.

Most of the systems worth caring about are nonlinear. The dashboard isn’t going to tell you which kind you’re in. You have to know.

Risk Disclaimer The contents are for information purposes only. Capital investment involves risks. No investment advice.

Live System

Social Trading Portfolio

Copy my trades in real-time on eToro – transparent and efficient.