Green Dashboards, Red Reality: When Your Engineering Metrics Are Just Cosplay
Photo by Photo by Mohammad Rahmani on Unsplash on Unsplash
There's a specific kind of meeting that happens in engineering orgs across the country, usually on a Thursday, usually with too many people on the call. Someone shares their screen. A dashboard appears. Numbers glow green. Deployment frequency is up 40% quarter-over-quarter. Test coverage just cracked 90%. Lead time for changes is trending in the right direction.
Everyone nods. The VP of Engineering looks satisfied. The meeting ends in 22 minutes instead of the scheduled 45, which itself gets logged somewhere as a productivity win.
And then, two days later, a hotfix goes out at 11 PM because something broke in production that nobody caught.
This is the metrics cosplay problem. And honestly, it's one of the more insidious things happening in software teams right now.
The Scoreboard Isn't the Game
Here's the uncomfortable truth about engineering metrics: most of the ones teams obsess over are outputs, not outcomes. There's a massive difference between those two things, and conflating them is where the trouble starts.
Deployment frequency, for example, is a genuinely useful signal — when it's measuring something real. The DORA metrics (developed by the DevOps Research and Assessment team at Google) weren't meant to be targets. They were meant to be diagnostic tools. But somewhere along the way, a lot of teams started optimizing for the metric rather than for the underlying health the metric was supposed to reflect.
So you get teams splitting large features into micro-deployments just to pump up the frequency number. You get test suites padded with trivially passing assertions to inch that coverage percentage upward. You get commit counts inflated by formatting commits, rename commits, and the classic "fix typo" trilogy that somehow spans three separate pushes.
The dashboard stays green. The software gets shakier.
What "Coverage" Actually Covers
Test coverage is probably the most widely misunderstood metric in the industry. An 85% coverage number sounds rigorous. It implies discipline. It implies that someone, somewhere, thought carefully about edge cases.
Except coverage tells you which lines were executed during tests — not whether those tests actually assert anything meaningful. A test that calls a function and checks that it doesn't throw an error will cover those lines. It will not tell you whether the function returned the right value, handled a null input gracefully, or interacted correctly with the database layer.
There's a term some folks use: "coverage theater." You can have a codebase sitting at 92% coverage that is, functionally, barely tested at all. The lines ran. The assertions were mostly vibes.
The more useful question isn't "what percentage of our code is covered?" It's "what percentage of our critical paths are covered by tests that would actually catch a regression?" That's harder to measure. It doesn't fit neatly into a CI dashboard widget. Which is probably why most teams don't measure it.
The Velocity Trap
Story points and sprint velocity deserve their own moment in the dock here.
Velocity was originally conceived as a planning tool — a way for teams to estimate how much work they could realistically take on in a given sprint based on historical data. It was never meant to be a performance benchmark. And yet.
When velocity becomes a KPI, teams get creative. Story point inflation is real and widespread. Tasks that would have been a 3 become a 5. Bugs that get fixed as part of a feature somehow acquire their own separate story point value. The velocity number climbs. The actual throughput of meaningful work does not.
Worse, velocity pressure actively degrades quality. When the implicit message is "ship more, faster," the things that get cut are the things that don't show up on the board: documentation, refactoring, code review depth, cross-team communication. The technical debt accumulates invisibly while the velocity chart trends upward.
So What Should You Actually Be Measuring?
This isn't an argument against metrics. Measurement matters — the right measurement, pointed at the right thing.
A few signals that tend to be more honest about actual engineering health:
Change failure rate. Not how often you deploy, but how often those deployments cause incidents, rollbacks, or emergency patches. A team deploying twice a week with a 2% change failure rate is in a fundamentally different position than a team deploying ten times a week with a 15% failure rate.
Mean time to recovery (MTTR). When something does break, how long does it take to get back to stable? This is a proxy for system observability, on-call culture, runbook quality, and architectural resilience all at once. It's hard to fake.
Escaped defect rate. How many bugs are your customers finding that your tests didn't? This is the brutal mirror that coverage percentages can't obscure. If users are filing bug reports for issues that your test suite should have caught, that's a direct indictment of test quality regardless of what the coverage number says.
Unplanned work percentage. What fraction of each sprint gets consumed by incidents, hotfixes, and "quick questions" that turn into two-day investigations? High unplanned work is a symptom of accumulated fragility. It's also something that velocity metrics actively hide, since unplanned work often gets retroactively pointed and added to the board.
The Cultural Layer Underneath
Here's the part that's genuinely tricky to fix: metrics cosplay isn't usually a technical problem. It's a cultural one.
Teams optimize for the metrics that leadership pays attention to. If your CTO celebrates high deployment frequency in all-hands meetings, engineers will find ways to produce high deployment frequency. If coverage percentage is the thing that gets mentioned in performance reviews, coverage percentage will go up — whether or not the underlying test quality does.
This means fixing the metrics problem requires fixing the incentive structure that surrounds them. That means leadership being honest about the difference between a metric that reflects health and a metric that looks like health. It means creating psychological safety around admitting that the current measurement framework might be pointing in the wrong direction.
It also means being willing to accept that some of the most important engineering work — the kind that makes everything else more stable, faster, and less painful — doesn't show up on any dashboard at all.
Asking Better Questions
The goal isn't to measure less. It's to measure honestly.
Before you add a new metric to the team dashboard, it's worth asking: could a team game this number while getting worse at engineering? If the answer is yes, it probably shouldn't be a target. It might still be a useful diagnostic signal, but the moment it becomes a goal, Goodhart's Law kicks in and it stops being a reliable measure of the thing you actually care about.
The best engineering teams tend to be a little suspicious of their own dashboards. They treat metrics as opening questions, not closing answers. They ask what the number doesn't capture. They look for the gap between what the data says and what the engineers in the room actually know to be true.
That gap is usually where the real story lives.