Skip to main content
TACAVAR
•Build in Public

My Grafana Dashboard Looked Beautiful and Had Exactly One Operation in It

I probed my beautiful Grafana dashboards. 59 of 59 traces were one heartbeat operation.

The stack was up. Tempo ingesting, OpenTelemetry traces flowing through the collector, three dashboards rendering — Bailian Team Overview, Paperclip Swarm Observability, Tacavar Ops. Tight latency heatmaps. Populated bar charts. The kind of screen you screenshot for a board update. I was sure Tacavar's observability was working. Then I ran the queries behind the panels, and the whole thing collapsed in about ninety seconds.

The Beautiful Dashboard Illusion

A rendered dashboard is a claim, not evidence. Grafana will happily draw you a flame graph, a p95 line, a stacked bar chart of token spend — none of which require the underlying data to exist. Panels are opinionated about layout and silent about truth. A panel pointed at a metric that was never emitted renders as an empty styled timeseries: axes labeled, gridlines drawn, legend intact, zero points. Visually, that is indistinguishable from "low but real" traffic. If your baseline is a quiet cron service, an empty panel and a healthy panel look identical at a glance.

That is the trap. The dashboards were not lying in the sense of showing wrong numbers. They were lying in the sense of showing nothing with the same visual confidence as showing something. And I had been treating the render as the signal.

Probing the Queries: What 59 of 59 Traces Revealed

I went to the Tempo query behind the trace panel and counted rows in the last hour. 59 traces. Every single one was the same operation: paperclip_handle_heartbeat, a 60-second cron polling its own work queue. Median duration ~250ms. That is the sound of a system talking to itself.

Zero agent runs. Zero LLM calls. Zero tool calls. Zero task outcomes. The Paperclip Swarm Observability dashboard — the one with the heatmaps I was proud of — was rendering the pulse of a cron daemon and nothing else. The swarm, in trace terms, did not exist.

This is the core failure mode of opentelemetry traces in a young system: instrumentation gets added where it is easy, not where it matters. Cron wrappers are trivial to instrument. Agent run boundaries require you to define what a run is, thread context through the orchestrator, and decide what a span means when the agent retries. So you instrument the heartbeat, you see traces flowing, and you conclude you have observability. You have a liveness check with extra steps.

Missing Metrics That Render as Empty Timeseries

Bailian Team Overview was worse. It queried Prometheus for agent_calls_total, agent_tokens_total, and agent_cost_usd_total. None of those metrics exist in the Tacavar deployment. They were never emitted — the counters were named in a design doc and wired into the dashboard before anyone wrote the exporter.

Grafana's response to this is polite and devastating: it renders the panel anyway. No red error banner. No "no data" empty-state worth noticing. Just a styled timeseries with a flatline that reads as zero. Missing metrics monitoring is not a gap you notice; it is a gap that actively misleads you, because the absence is styled to look like a measurement.

The fix is not more panels. It is treating metric existence as an assertion you can test. Every dashboard query is a contract: this metric is emitted, at this cadence, by this service. If the contract is unverified, the panel is decoration.

Why Graceful Degradation Is an Anti-Feature

We treat graceful degradation as a virtue everywhere else. The service stays up, the request still returns, the UI still renders. In observability it is the opposite of a virtue. A broken dashboard makes you investigate. A beautifully empty one makes you think your system is working. One of those costs you an afternoon. The other costs you a quarter.

The failure mode compounds because empty is the default state of most infrastructure. Pre-launch, there is no traffic. Post-launch, low-traffic services look like pre-launch. A missing metric and a genuinely idle subsystem produce the same pixels. The dashboard cannot distinguish "nothing is happening" from "nothing is measured," and it will not try.

So the correct behavior for a missing metric, a query returning zero rows over a window where rows are expected, or a trace span that only ever contains heartbeats is to scream. Loud, ugly, unmissable. Panels should be allowed — encouraged — to fail hard when their data contract is broken. A dashboard that can silently show nothing is worse than no dashboard, because it consumes the attention you would otherwise spend finding the problem.

These are the observability anti-patterns worth naming: instrumenting what is easy instead of what is load-bearing, shipping panels before exporters, and trusting the render over the query.

Three Rules for Trustworthy Observability

1. Every panel has a query-and-count test. For any panel, you must be able to run the underlying query and get a non-zero row count within a known window, for a known reason. Not "someone looked at it and it seemed fine." A number. If the expected count is zero, the panel must be explicitly annotated as idle, not left to render empty.

2. Cardinality of operations, not volume of spans. Tacavar's trace pipeline should assert that the set of distinct operation names in the last hour includes agent_run, llm_call, and tool_call — not just paperclip_handle_heartbeat. Span volume is vanity. Operation diversity is the actual signal that the system you care about is being observed.

3. Alerts fire on absence, not just on thresholds. The dangerous case is not p99 latency spiking. It is agent_calls_total returning no series for six hours, which no threshold alert will ever catch. Absence alerts are the only ones that catch the empty-dashboard trap after the fact.

A grafana dashboard audit is not a design review. It is a query review. You are checking that each panel's data contract holds, and that the renderer is not doing the thinking for you.

How to Audit Your Own Dashboards in 10 Minutes

Open the query editor behind every panel. Set the window to the last hour. Run it. Count rows. For time series, check the series count, not the shape of the line. For traces, group by operation name and look at the distribution — if one operation dominates above 90%, your dashboard is describing a cron job, not your product.

Then do the reverse test: pick the three metrics your system's health actually depends on, and search Prometheus for them by name. If they are not there, you have found your missing metrics. Add the exporter before you add the panel. Anything else is theater.

That is the whole audit. It took me ten minutes to find that 59 of 59 traces were a heartbeat and three headline metrics did not exist. The dashboards had been beautiful for weeks.

Explore Tacavar's production observability stack for autonomous agents at tacavar.com