Part 2 · 1 chapters · ~8 min
Canary Releases and Automated Analysis
Choosing canary slices, comparing canary against baseline, enough traffic for significance, automated promotion and rollback with Argo Rollouts, Flagger and Kayenta, weighted routing through meshes and ingresses, and what canaries cannot catch.
3
Canaries that judge themselves
code
# Argo Rollouts: canary steps with automated analysis
apiVersion: argoproj.io/v1alpha1
kind: Rollout
spec:
strategy:
canary:
steps:
- setWeight: 5
- analysis: { templates: [ { templateName: error-rate-and-p99 } ] }
- setWeight: 25
- pause: { duration: 10m }
- setWeight: 50
- analysis: { templates: [ { templateName: error-rate-and-p99 } ] }
---
kind: AnalysisTemplate
spec:
metrics:
- name: error-rate
interval: 1m
count: 5
failureLimit: 1
successCondition: result[0] < 0.01
provider: { prometheus: { query: 'sum(rate(http_requests_total{app="ledger",version="{{args.version}}",code=~"5.."}[1m])) / sum(rate(http_requests_total{app="ledger",version="{{args.version}}"}[1m]))' } }What canaries miss: problems that only appear at full load (capacity, lock contention), problems in data written by the canary that the old version later reads, slow-burn leaks over hours, and features hidden behind flags (the canary never runs them). Combine canaries with load tests, compatibility rules and flagged releases.
A CANARY WITH AUTOMATED ANALYSIS
shift a little traffic, measure, and let the analysis decide
swipe the figure sideways, or tap expand for full screen
1/5
shift a little
Route 1-5% of traffic to the canary. Choose the slice deliberately: internal users first, then a country, then a percentage.
a small, chosen slice of trafficinternal users, then a region, then %