Insights
SRE26 June 2026

I logged in this week. Datadog was down for two weeks. Nobody knew.

I joined a client this week and went to look at the application logs. There were none. The Datadog agent had been broken for weeks. Nobody had noticed. This is what monitoring theatre looks like in 2026.

Pattern·The monitoring stack is unmonitored

When the monitor is the thing that breaks

Dashboards still green. Logs stopped two weeks ago. Nobody knew.

Two weeks of silence

Day 0

Agent healthy

Day 1

Agent crashes

Day 3

Still broken

Week 2

Logs lost

Today

Someone looked

Dashboards remained green throughout. No alert fired.

Three failure modes stacked

1

Nobody monitors the monitor

No health check on the agent itself. No alert on log volume dropping to zero.

2

No log ingestion fallback

Single collector. No on-host buffer. No secondary path. Agent fails, logs are gone.

3

Stderr is noise, alerts ignore reality

Warnings and deprecations in stderr. Alerts fire on stderr volume. Team muted the channel.

Four questions to ask right now

  • 1. What alerts when your monitoring agent dies?
  • 2. What is your fallback when the primary collector fails?
  • 3. Are your error logs actually errors, or is stderr the dumping ground?
  • 4. Can you find a service's p99, error rate, traffic, and saturation in 30 seconds?

Treat your monitor as a critical service that needs monitoring.

I logged into a new client environment this week. First thing I do on any engagement is pull the last 24 hours of logs for the critical services. There were none. Not no errors. No logs at all.

I went looking. The Datadog agent had been broken. The agent status command returned nothing. To get it into a clean state we had to restart it just so we could stop it. While that was happening, none of the production services had been emitting logs to anyone. Nobody had been monitoring anything. The dashboards still looked green because nothing was being measured.

An observability stack that has stopped collecting data is worse than no observability stack. With no observability stack, you know you are flying blind. With a broken Datadog agent, the dashboards still report. The flight is blind anyway.

The fix was technical and took an hour. The reason this happened took longer to unpick. Three failure modes, all common, all underestimated.

One: nobody monitors the monitor. The Datadog agent itself was not in any health check. There was no alert that fires when log ingest drops to zero. The first signal that monitoring was broken was that I happened to look. That is not a strategy.

Two: there was no log ingestion strategy at all. The agent had been configured to ship application logs. It had silently stopped. There was no fallback. No secondary collector. No on-host buffer. When the agent failed, the logs were just lost.

Three: even when logs were flowing, they were unusable. The application was writing warnings, deprecation notices, and informational messages to standard error. The alerting rules were configured against standard error volume. Every routine warning was indistinguishable from a real error. The team had given up reading the alerts months ago.

  • Standard error was the dumping ground for everything that was not standard out, including non-errors
  • Alerts firing on stderr volume meant deprecation warnings paged people
  • After enough false positives, the team muted the channel
  • No golden signals dashboards (latency, traffic, errors, saturation) per service
  • During an incident, no one could drill into a service to find where the bottleneck was

If you cannot answer four questions about each of your critical services right now, your observability is in the same state as this client's was when I walked in.

  • Are your monitoring agents alive? What alerts when they are not?
  • Are logs being ingested? What is your fallback when the primary collector fails?
  • Are your error logs actually errors, or are they noise warning channels people have stopped reading?
  • If you opened a service dashboard during an incident right now, could you find the latency p99, error rate, traffic volume, and saturation in under 30 seconds?

Most teams discover their monitoring is broken during the incident the monitoring was supposed to warn them about. The fix is small. Add health checks for your observability infrastructure. Treat your monitor as a critical service that needs monitoring.

ShareLinkedIn

Get the next one in your inbox

One short, opinionated field note per fortnight on platform engineering, cloud, and making AI work in production. No spam. Unsubscribe anytime.

Senna Semakula

Senna Semakula

Founder, Atruvo

Bring your architecture diagram, cloud bill, or last incident summary.

I will tell you what is actually breaking.

30 minutes. No pitch. Ranked risks and a clear next step.