Part 7 · 1 chapters · ~8 min
Operating Both
Cassandra operations (nodetool, repair scheduling with Reaper, adding and removing nodes, bootstrapping, compaction and disk headroom, the metrics that mean an incident), MongoDB operations (rolling upgrades, index builds, currentOp, the profiler, Atlas), backups and point-in-time recovery for each, and capacity modelling.
11
Runbooks for both
code
# Cassandra
nodetool status # UN = up/normal; DN = down; load and token ownership per node
nodetool tpstats # thread pools: pending or dropped mutations = overload
nodetool tablestats bank.transactions_by_account # SSTable count, partition size max, tombstones per read
nodetool repair -pr # primary-range repair; schedule with Cassandra Reaper, within gc_grace_seconds
# keep ~50% disk free with STCS: compaction needs temporary space
# MongoDB
db.currentOp({ secs_running: { $gt: 5 } }) # long-running operations
db.setProfilingLevel(1, { slowms: 100 }) # profile slow operations
db.serverStatus().opcounters ; rs.status() # throughput and replica health
db.transfers.createIndex({ merchant_id: 1 }) # 4.2+: built without blocking reads and writes for long| Cassandra | MongoDB | |
|---|---|---|
| backup | per-node snapshots + incremental, Medusa | mongodump (small), snapshots or Atlas continuous backup with PITR |
| incident metrics | dropped mutations, pending compactions, read latency p99, tombstone warnings, disk | replication lag, cache eviction, queued operations, slow queries, oplog window |
| adding capacity | add nodes; data streams in; run cleanup | add shards; balancer migrates chunks |