Part 4 · 9 chapters · ~45 min
Observability: being able to ask new questions
The test is simple: when something novel breaks, do you investigate, or do you add logging and wait for it to happen again? This part is about the second answer being unacceptable, and about the constraint that shapes every decision here, which is cardinality: metrics cannot carry an account id, logs and traces can, and the bill is decided by engineers adding a label without noticing the multiplication.
38
Monitoring answers known questions; observability answers new ones
worked numbers
monitoring: is the thing I expected to break, broken? dashboards and alerts for failure modes you predicted observability: can I answer a question I did not anticipate, without shipping new code? the test: when something novel breaks, do you investigate, or do you add logging and wait for it to happen again?
the practical difference
A system with good monitoring and poor observability produces incidents where the alert fires correctly and then the investigation requires a deploy. Part 0 made that measurable: "percentage of incidents diagnosed without shipping new instrumentation". If that number is low, you have monitoring, and the distinction is not academic: it is the difference between a twenty-minute incident and a two-day one.
39
The three signals, and the fourth for money
| Signal | Shape | Strength | Weakness |
|---|---|---|---|
| Metrics | Aggregated numbers over time | Cheap, fast, alertable | No cardinality. Cannot ask about one customer |
| Logs | Discrete events with fields | High cardinality, arbitrary detail | Expensive at volume, slow to aggregate |
| Traces | Causally linked spans | Shows where time went across services | Sampled, so the one you want may be missing |
| Correctness | Invariant assertions | Catches "wrong", which the others cannot | Entirely domain-specific. You build it |
the fourth signal, restated for this module
Part 13 of the CBA module called correctness the fifth golden signal. In the three-pillar framing it is the fourth, and the point is identical: metrics, logs and traces all describe what the system did, and none of them has an opinion about whether the result was right. A ledger that loses a million naira produces perfectly healthy telemetry. Invariant violations, suspense ageing and reconciliation breaks are the only signals that detect it.
40
Cardinality: the constraint that shapes everything
worked numbers
cardinality = the number of distinct values a field can take low: region (4), currency (4), rail (6), shard (8) high: account_id (20M), journal_id (unbounded), user agent a metric costs per unique combination of label values: region × currency × rail × shard = 4×4×6×8 = 768 fine add account_id = 15 billion no
the resolution
- Metrics carry low-cardinality dimensions only. Shard, region, rail, currency, status code, endpoint.
- Logs and traces carry high-cardinality fields. Account id, journal id, trace id, customer id.
- Structured events give you both: emit one wide event per operation with all the fields, derive low-cardinality metrics from it, and keep the event for querying. This is what EMF does on AWS and what wide-event tooling does generally.
- The correlation id bridges them. Part 13 used
journal_idas the one identifier present in every metric exemplar, every log line and every span.
the cost consequence, which is an engineering decision
The cloud module put observability at around 9% of the bill, above the warehouse. That figure is decided almost entirely by cardinality choices made by individual engineers adding a label without thinking about the multiplication. One careless dimension can double the bill, and nobody notices until a quarterly review.
41
Structured events over metrics, where it matters
code
// one wide event per operation, emitted once
{
"event": "transfer.completed",
// low cardinality → becomes metrics
"region": "ng", "rail": "nibss", "shard": 7, "status": "ok",
// high cardinality → stays in the event, queryable
"journal_id": "j-8f1a", "account_id": "acct-771",
"trace_id": "4bf92f35",
// the measurements
"duration_ms": 268, "fraud_ms": 142, "posting_ms": 91,
"amount_bucket": "10k_50k", "retry_count": 0
}why one wide event beats many narrow logs
- One write, not twelve. Cheaper to emit, cheaper to store, cheaper to ship.
- Every field is correlated by construction. You can ask "what was the fraud latency for transfers on shard 7 that retried" without joining anything.
- Metrics derive from it, so the metric and the event can never disagree, which is a surprisingly common bug with separate instrumentation.
- New questions need no new code. The field is already there. That is the definition of observability from chapter 38.
42
Sampling strategies that keep the interesting traces
| Strategy | How | Problem |
|---|---|---|
| Head-based, uniform | Decide at the start, 1% | Discards 99% of errors too |
| Head-based, weighted | Higher rate for some endpoints | Still blind to outcome |
| Tail-based | Decide after completion | Requires buffering. Worth it |
| Exemplar-linked | Metric buckets link to a trace | Best of both, needs support |
the tail-based policy for our system
- Keep 100% of errors. Rare and maximally informative.
- Keep 100% of traces slower than the SLO. These are the ones you will want.
- Keep 100% above a value threshold. A ₦5m transfer is worth the storage.
- Keep 100% where an invariant check or a reconciliation flag fired.
- Keep 0.1% of ordinary successes, for a baseline to compare against.
why head-based uniform sampling is the wrong default
It is the most common configuration and it optimises for exactly the wrong thing. At 1% head-based, an error occurring 50 times an hour is captured roughly once every two hours, so the trace you want during an incident is probably not there. Tail-based costs buffering and gives you every trace you actually need, which is one of the better trades in observability.
43
Dashboards people actually use
the three kinds, and most teams build only one
- The status dashboard. Is it working? Five to eight numbers, readable in ten seconds, on a wall. Should be boring.
- The diagnostic dashboard. One per service, showing its dependencies, saturation and errors. Used during an incident, linked from the alert.
- The exploration surface. Not a dashboard: a query interface over the structured events. This is where novel questions get answered, and its absence is what makes investigations require a deploy.
what makes a dashboard unused
- Too many panels. Forty graphs means nobody knows which matters, so nobody looks.
- No annotations. Without deploy markers, a step change in a graph has no explanation.
- Averages instead of percentiles. Per Part 2, it hides the thing you need to see.
- No link from the alert. If the page does not carry a link to the relevant view, the responder builds it from scratch at 3am.
- Never used in anger. A dashboard nobody consulted during the last three incidents should be deleted.
44
Alerting: the rules that stop alert fatigue
the rules
- Alert on symptoms, not causes. Part 13 said it: "transfer success below 99.5%" pages, "CPU above 80%" does not.
- Every page has a runbook, linked from the alert, with a do-not list.
- If it fires and nobody is affected, it should not have paged. Demote it to a ticket or delete it.
- Two tiers only. Page means wake a human now. Ticket means look at it in working hours. A third tier becomes a queue nobody reads.
- Burn-rate alerting over threshold alerting. Burning a month of error budget in an hour pages; burning it slowly raises a ticket.
- Review every page. Weekly: was it actionable, was it correct, did the runbook help? Alerts that fail this get changed or deleted.
worked numbers
the page budget, from Part 6:
target: fewer than 2 pages per on-call night
above that → responders stop reading carefully
→ the real one gets missed
alert fatigue is not a morale problem, it is a detection failure.45
Debugging a system you cannot reproduce
the method
- Establish the blast radius first. All shards or one? All regions or one? All customers or a segment? This single question eliminates most hypotheses.
- Establish the time boundary. When did it start, exactly? Correlate against deploys, config changes, traffic shifts and provider status.
- Find one concrete example. One journal id that exhibits it. Specific beats aggregate: one trace often tells you more than a day of graphs.
- Follow the correlation id across every service, per Part 13.
- Form a falsifiable hypothesis, then look for evidence that would disprove it. Confirmation bias is the main cost of long incidents.
- If you must add instrumentation, add it wide, not narrow. You will want the fields you did not think of.
the question that shortcuts most investigations
"What changed?" The overwhelming majority of incidents follow a change: a deploy, a flag, a config, a dependency version, a traffic pattern, or something on a provider's side. A timeline of changes next to a timeline of symptoms resolves more incidents than any amount of code reading, which is why deploy annotations on dashboards matter so much.
46
The cost of observability, and containing it
where the money goes, in order
- High-cardinality metrics. Per chapter 40, the dominant cost and the easiest to create accidentally.
- Log volume. Debug logging left on in production, or logging entire request bodies.
- Trace retention. Storing 100% of traces for 30 days when 0.1% for 7 days plus all errors for 30 would do.
- Log ingestion of things already in metrics. Paying twice for the same information.
the controls
- Cardinality limits enforced in the instrumentation library, so a dangerous label cannot be added without a deliberate override.
- Tiered retention. Debug 3 to 7 days, application 30, audit 7 years in cheap storage.
- Sample aggressively and keep what matters, per chapter 42.
- Attribute cost per team, exactly as Part 12 of the CBA module did with notification spend. Visible cost changes behaviour; an aggregate bill does not.
the balance to state explicitly
Under-instrumenting is more expensive than over-instrumenting, and both are expensive. A two-day incident that better instrumentation would have made twenty minutes costs more than a year of the extra telemetry. The goal is not minimum cost; it is the cheapest instrumentation that keeps investigations short, and that is a judgement that needs revisiting as the system changes.