Kueri.me All articles
Developer Culture

Crying Wolf at Scale: How Dev Tools Trained Us to Ignore Everything

Kueri.me
Crying Wolf at Scale: How Dev Tools Trained Us to Ignore Everything

Photo: William Murphy from Dublin, Ireland, CC BY-SA 2.0, via Wikimedia Commons

There's a particular kind of numbness that sets in around week three of a new job. You've got Slack hooked up to your CI pipeline, PagerDuty pinging your phone, Datadog firing off emails, and your IDE underlining roughly forty percent of your codebase in red squiggles. By the end of month one, you've muted half of it. By month three, you've built an unconscious filter so aggressive that you almost miss the 3 a.m. alert that the payment service is down.

Almost.

This is where a lot of engineering teams quietly live — not in ignorance, but in a kind of learned helplessness. The tools are technically working. The notifications are technically accurate. And yet the signal has been so thoroughly buried in noise that the whole system is effectively broken.

The Notification Arms Race Nobody Asked For

Here's how it usually starts: someone on the team gets burned. A bug slips through, a deployment goes sideways, a database query starts timing out at 2 a.m. on a Friday. The postmortem conclusion, almost universally, is we need better visibility.

So you add monitoring. You add logging. You wire up alerts for CPU spikes, memory pressure, error rate thresholds, latency percentiles. You install a linter that flags every possible code smell. You set up a bot that posts to Slack whenever a PR sits open longer than 48 hours.

None of those decisions are wrong in isolation. Each one is a reasonable response to a real problem. But the cumulative effect is a development environment that treats every observation as equally urgent — which is functionally the same as treating nothing as urgent at all.

Human attention doesn't scale linearly with notification volume. It degrades. Psychologists call it habituation: repeated exposure to a stimulus reduces your response to it. Your brain is literally wired to stop caring about things that happen constantly. Tool designers know this, or should. And yet the default configuration for most observability platforms is closer to "alert on everything" than "alert on what matters."

What Gets Lost in the Static

The real cost of alert fatigue isn't the time spent dismissing notifications — though that's real and it adds up. The deeper cost is what happens to your team's relationship with the tools themselves.

Engineers are pattern matchers. We are very good at learning which signals to trust and which ones to tune out. The problem is that when a system cries wolf often enough, our brains don't distinguish between the false alarms and the real ones. We just stop processing the category.

This is how you end up in a postmortem saying something like, "Yeah, we were seeing elevated error rates in that service for a few days, but it's always a little noisy in there." That sentence is a sign that your observability layer has failed. Not because the data wasn't there — it was. But because the team had been conditioned to discount it.

There's also a subtler effect on junior engineers. When someone new joins a team and sees that experienced developers routinely ignore certain classes of warnings, they learn to do the same — before they've developed the judgment to know which ones actually matter. The institutional knowledge about what's signal and what's noise lives in people's heads, not in the tooling, which means it evaporates every time someone leaves the team.

The Configuration Nobody Has Time to Fix

Part of what makes this problem sticky is that fixing it requires time and attention — exactly the resources that alert fatigue has already depleted.

Tuning an alerting system well is genuinely hard work. You have to understand your baseline, define what "anomalous" actually means for your specific system, and resist the temptation to add alerts reactively every time something breaks. It requires saying no to things that feel reasonable in the moment. And in most engineering organizations, there's no dedicated sprint for "make our alerts less annoying." It's the kind of work that lives permanently on the backlog, right next to updating the README.

Vendors don't help here. Most monitoring and observability platforms are sold on the promise of comprehensive coverage. More dashboards, more integrations, more data. The sales pitch is almost never "we'll help you see less stuff, but make sure the stuff you see actually matters." Minimalism doesn't move enterprise software licenses.

Recalibrating What Deserves Your Attention

So what does a healthier approach actually look like? A few things that tend to work in practice:

Start with outcomes, not metrics. Before you wire up an alert, ask what user-facing or business-critical outcome it's protecting. If you can't answer that question in one sentence, you probably don't need the alert — at least not as an interrupt. Log it, sure. Track it on a dashboard. But don't let it page anyone.

Separate interrupts from observables. Not everything that's worth knowing is worth knowing right now. Build a clear distinction between alerts that require immediate human action and data that's useful for retrospective analysis. A weekly digest of slow queries is valuable. A Slack ping every time a query exceeds 200ms is a fast track to muted channels.

Audit your alert history. Pull the last 90 days of alerts and look at what actually got acted on. If a significant chunk of your alerts were acknowledged and dismissed without any follow-up action, those are candidates for downgrade or removal. This is a slightly uncomfortable exercise because it forces you to confront how much noise you've been generating, but it's clarifying.

Make noise reduction a first-class task. Put it in the sprint. Give it a ticket. Assign an owner. Treating alert hygiene as real engineering work — not cleanup, not housekeeping, but actual work that affects system reliability — is the only way it actually gets done.

Give engineers control over their own signal. Different roles have different tolerances and different needs. A platform engineer and a product engineer shouldn't necessarily be subscribed to the same alert channels. Letting people customize their notification surface, within guardrails, reduces resentment and improves the chance that the alerts they do receive actually get read.

Attention Is Infrastructure

We spend a lot of time in this industry talking about system reliability, uptime, and resilience. We're pretty good at reasoning about infrastructure in those terms. But we're less good at treating developer attention as infrastructure — as a finite, degradable resource that requires maintenance and protection.

Every false alarm is a small withdrawal from that resource. Enough withdrawals and you've got a team that's technically monitoring everything and effectively watching nothing. The irony is that the more you instrument, the blinder you can become — if you haven't done the work of deciding what actually deserves to interrupt a human being.

The tools aren't going to fix this on their own. The defaults will keep defaulting toward more. That's on us to push back against — deliberately, repeatedly, and probably more often than feels necessary.

Because the next real incident is coming. And when it does, you want someone to actually notice.

All Articles

Related Articles

Release Notes Written in Guilt: The Changelog Crisis Nobody Wants to Fix

Release Notes Written in Guilt: The Changelog Crisis Nobody Wants to Fix

Stack Overflow Made You a Faster Coder. It Also Made You a Shallower One.

Stack Overflow Made You a Faster Coder. It Also Made You a Shallower One.

You Don't Love the Framework. You Love Who You Were When You Chose It.

You Don't Love the Framework. You Love Who You Were When You Chose It.