Part 8 · 2 chapters · ~18 min
Observability and Cost
Metrics, traces and logs, what each answers and costs, joined by trace ids and exemplars through OpenTelemetry, with symptom-based alerts and question-driven dashboards; then the bill as a metric: tags, the billing export, unit cost, budgets and anomalies, the usual fixes and the review cadence.
22
Logs, metrics, traces and alerts
three signals, joined
- Metrics: RED per service and USE per resource. Watch cardinality.
- Traces: spans propagated by traceparent, with tail sampling that keeps errors and slow requests.
- Logs: structured and carrying the trace id. They are the most expensive signal, so log less and better.
- Join the three with trace ids and exemplars. Instrument once with OpenTelemetry.
- Page on symptoms users feel, with a runbook and an owner for each alert.
- Dashboards per service, per journey and for the platform, built from real incident questions.
code
// OpenTelemetry in a Node service: instrument once, export anywhere
import { NodeSDK } from '@opentelemetry/sdk-node';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { OTLPMetricExporter } from '@opentelemetry/exporter-metrics-otlp-http';
import { PeriodicExportingMetricReader } from '@opentelemetry/sdk-metrics';
new NodeSDK({
serviceName: 'payments-api',
traceExporter: new OTLPTraceExporter(), // to the Collector sidecar
metricReader: new PeriodicExportingMetricReader({ exporter: new OTLPMetricExporter() }),
instrumentations: [getNodeAutoInstrumentations()], // http, express, pg, redis, aws-sdk…
}).start();
// pino with the active trace id in every line
import pino from 'pino'; import { trace } from '@opentelemetry/api';
export const log = pino({ mixin: () => ({ trace_id: trace.getActiveSpan()?.spanContext().traceId }) });the client half
The Disciplines course part 5 starts the trace in the browser with traceparent. With this service continuing it, one trace runs from the tap on Pay to the payment provider's response, and the frontend and backend debug the same incident from the same evidence.
LOGS, METRICS, TRACES AND ALERTS
the three signals, what each answers, what each costs, and the alert that pages a human
swipe the figure sideways, or tap expand for full screen
1/6
metrics
Metrics: counters, gauges and histograms with a few labels, scraped or pushed every 10 to 60 seconds and stored as time series. The RED method per service (Rate, Errors, Duration) and USE per resource (Utilisation, Saturation, Errors). Cheap per data point; expensive in cardinality: a label with user ids creates a series per user and a bill to match.
23
The bill as a metric
measured, attributed, normalised, alerted, reviewed
- Tags on everything, enforced when resources are created. Untagged spend has no owner.
- Export billing data to a warehouse and query it.
- Unit cost per transaction, per user and per tenant, on the same dashboard as latency.
- Budgets and anomaly detection catch problems within a day.
- The usual fixes: endpoints, rightsizing, commitments, spot, lifecycle rules, log hygiene, schedules.
- Ownership and cadence: team dashboards, a monthly review, and cost estimates in design reviews.
code
# terraform: tags enforced by default on every AWS resource in the stack
provider "aws" {
region = "eu-west-1"
default_tags {
tags = { team = "payments", service = "payments-api", env = var.env, cost_centre = "cc-104" }
}
}
# a budget with forecast alerts, owned by the team
resource "aws_budgets_budget" "payments_prod" {
name = "payments-prod-monthly"
budget_type = "COST"
limit_amount = "6000"
limit_unit = "USD"
time_unit = "MONTHLY"
cost_filter { name = "TagKeyValue" values = ["user:team$payments"] }
notification {
comparison_operator = "GREATER_THAN" threshold = 80 threshold_type = "PERCENTAGE"
notification_type = "FORECASTED" subscriber_email_addresses = ["[email protected]"]
}
}the exercise
Pick your most important transaction and compute last month's cost per unit: the spend of the services in its path divided by the number completed. Write the number in the service README. When it moves, you will now notice.
THE BILL AS A METRIC
tagging, attribution, unit cost, anomalies and the fixes for the usual surprises
swipe the figure sideways, or tap expand for full screen
1/6
tags
Tags and labels: every resource carries team, service, environment and cost-centre tags, enforced at creation (tag policies, organisation policies, a Terraform module that requires them). Untagged spend is unowned spend; aim for 95% or more of the bill attributed. Shared costs (the NAT, the cluster, the observability stack) are split by a documented rule.