Part 7 · 1 chapters · ~8 min

Operating Both

Cassandra operations (nodetool, repair scheduling with Reaper, adding and removing nodes, bootstrapping, compaction and disk headroom, the metrics that mean an incident), MongoDB operations (rolling upgrades, index builds, currentOp, the profiler, Atlas), backups and point-in-time recovery for each, and capacity modelling.

11

Runbooks for both

code
# Cassandra
nodetool status                    # UN = up/normal; DN = down; load and token ownership per node
nodetool tpstats                   # thread pools: pending or dropped mutations = overload
nodetool tablestats bank.transactions_by_account   # SSTable count, partition size max, tombstones per read
nodetool repair -pr                # primary-range repair; schedule with Cassandra Reaper, within gc_grace_seconds
# keep ~50% disk free with STCS: compaction needs temporary space

# MongoDB
db.currentOp({ secs_running: { $gt: 5 } })        # long-running operations
db.setProfilingLevel(1, { slowms: 100 })          # profile slow operations
db.serverStatus().opcounters ; rs.status()         # throughput and replica health
db.transfers.createIndex({ merchant_id: 1 })      # 4.2+: built without blocking reads and writes for long
CassandraMongoDB
backupper-node snapshots + incremental, Medusamongodump (small), snapshots or Atlas continuous backup with PITR
incident metricsdropped mutations, pending compactions, read latency p99, tombstone warnings, diskreplication lag, cache eviction, queued operations, slow queries, oplog window
adding capacityadd nodes; data streams in; run cleanupadd shards; balancer migrates chunks