High cardinality is the bill nobody can read
The custom metrics overage on your Datadog invoice is almost always one or two high-cardinality tags turning a handful of metrics into millions. Here's how to find them.
Read post →Notes on observability, reliability, and the stuff that breaks at 3am.
The custom metrics overage on your Datadog invoice is almost always one or two high-cardinality tags turning a handful of metrics into millions. Here's how to find them.
Read post →In May 2026 a single AWS availability zone had a thermal event and 150+ services fell over. Most of them believed they were multi-AZ. Their failover only existed on paper.
Read post →Observability is 7 to 12 percent of cloud spend and climbing, yet in most orgs no single person is accountable for the number. Give the bill an owner.
Read post →Gartner published its first AI SRE Market Guide in January. The tech is real. The way most teams plan to deploy it will turn a blip into an outage.
Read post →An agent stuck in a retry loop can burn $40k overnight, and you'll find out from the invoice. Put the meter on the span.
Read post →OpenTelemetry graduated from the CNCF in May. The standards war is over. Operating the Collector is the new problem.
Read post →Beyla became an official OpenTelemetry project. Zero-touch instrumentation is real now, but it doesn't replace your SDKs and the vendor who built it says so.
Read post →The DORA 2025 data is blunt about it. AI raised your throughput and your change-failure risk at the same time.
Read post →Observability 2.0 isn't a vibe. ClickHouse runs a 100PB stack on it and published the CPU math to prove it.
Read post →Continuous profiling just became an official OpenTelemetry signal. The honest buy-or-wait call for a Series B-D team.
Read post →Ship an LLM agent and the request log stops being the unit of observability. The trace tree is. Most teams find out the hard way.
Read post →Cloudflare and AWS didn't fall over from too much traffic. A bad config and an empty DNS record did it.
Read post →Datadog meters logs twice: once to send them, again to make them searchable. The second one is where the bill quietly gets away from you.
Read post →Fourteen months, three engineers, a 400-line golden path two teams use. Platforms fail on politics, not tech, when you optimize for the 20% edge cases.
Read post →A 4-hour outage knocked $8M off a term sheet because nobody had the number. The hidden downtime costs that never show up in your MTTR report.
Read post →One team paid $8k to run a RAG chatbot and $89k to log it. AI observability breaks traditional APM. Here's how to instrument LLMs without going broke.
Read post →One team tried to kill Datadog and ended up with Datadog, Grafana Cloud, and New Relic, 40% over budget. Tool sprawl is a leadership problem, not a tech one.
Read post →On-call doesn't just burn SREs out, it pushes them into PM roles for a 40% bump and no pages. Each departure costs $150k+. Fixing on-call costs less.
Read post →Hiring a senior SRE who can read code takes 3-6 months and $200k+. Here's what to do about the 90-day gap while you search for that unicorn.
Read post →One user_id tag took a Datadog bill from $2k to $47k. High-cardinality tags multiply your metric series. Here's how to set an observability budget.
Read post →An OTel SDK bump from 1.28 to 2.1 caused a 3-hour outage. Why OpenTelemetry upgrades silently break prod and how to test the upgrade path first.
Read post →Finance flagged your Datadog spend? Tie telemetry to revenue, use tail-based sampling, and cut retention. How to defend the budget that actually matters.
Read post →Most SRE postings are 23-bullet wishlists nobody meets. The hire that matters can read your app code and instrument OTel, not just move YAML.
Read post →Most alerts page you about CPU and memory, not whether checkout works. Switch to SLO-based alerting and stop losing sleep over garbage.
Read post →A solid engineer spent three weeks adding tracing and still failed. OpenTelemetry nailed the hard tech but the usability is rough. Start with one path.
Read post →Your senior engineers burn 30% of their time on incidents because the system needs tribal knowledge to debug. That's $50k-$75k each, in salary alone.
Read post →Your 10x engineer is a single point of failure. When the one who holds the system in their head leaves, 20-minute fixes become 4-hour outages.
Read post →Multi-cloud observability means jumping CloudWatch, Azure Monitor, and GCP to chase one request. Here's how to get a real single pane of glass.
Read post →Eight senior engineers in a war room for four hours runs $4,500 a pop, and most teams do it weekly. Distributed tracing turns that into 30 seconds.
Read post →82% of orgs have MTTR over an hour and noise is the reason. Kill threshold alerts, switch to SLO burn-rate alerting, and let on-call sleep again.
Read post →You scaled to 23 microservices and nobody can draw how they connect. The fix for Series B observability isn't a rewrite, it's distributed tracing.
Read post →Two years of eBPF observability in prod: the ARM64 breakage, the BPFDoor security risk, and when traditional OTel instrumentation just wins.
Read post →Your board hears nines, they think dollars. Learn to report downtime as lost revenue so you actually get observability budget approved.
Read post →82% of companies now take over an hour to resolve incidents. More tools made you slower. The fix isn't adding a tool, it's removing three.
Read post →SOC2 auditors don't care about dashboards. They want proof you'd detect a data breach, an audit trail 6 months back, and an MTTR you can explain.
Read post →Trace context dies at Kafka, everyone names attributes differently, and the OTel collector silently drops spans. Hard lessons from 50+ microservices.
Read post →Your payment service can return 200 OK with a failed body. All green dashboards, 40 support tickets. Instrument business outcomes, not HTTP calls.
Read post →