Part 3 · 1 chapters · ~8 min

Cassandra Data Modelling

Query-first modelling, partition keys and clustering columns and the physical layout they produce, bounding wide partitions with time buckets, denormalisation and write fan-out, materialised views and why they are experimental, secondary indexes and SAI, collections, counters and their costs, time series modelling, anti-patterns, and CQL as not-quite-SQL.

6

One table per query

code
CREATE TABLE transactions_by_account (
  account_id uuid, month text, ts timestamp, txn_id timeuuid, amount_kobo bigint, type text,
  PRIMARY KEY ((account_id, month), ts, txn_id)
) WITH CLUSTERING ORDER BY (ts DESC, txn_id DESC)
  AND compaction = {'class': 'TimeWindowCompactionStrategy', 'compaction_window_unit': 'DAYS', 'compaction_window_size': 7};

SELECT * FROM transactions_by_account WHERE account_id = ? AND month = '2026-10' LIMIT 50;   -- one partition, sequential
-- SELECT * FROM transactions_by_account WHERE type = 'refund';   → needs ALLOW FILTERING: a full cluster scan

CQL looks like SQL and is not: no joins, no subqueries, WHERE only on partition keys and clustering prefixes (plus indexes), no foreign keys, and UPDATE is an upsert. Storage-attached indexes (SAI, Cassandra 5) make secondary queries far more practical than the old indexes, but the core rule stands: design tables for the queries.

QUERY-FIRST MODELLING
one table per query, partitioned by what you look up
query: transactions for an account, newest firstquery: transactions by merchant per daytransactions_by_accountPK ((account_id, month), ts DESC, txn_id)transactions_by_merchant_dayPK ((merchant_id, day), ts, txn_id)application writes bothbatch or async fan-out
swipe the figure sideways, or tap expand for full screen
1/5
queries first
In relational modelling you design entities, then write queries. In Cassandra you list the queries first, then design one table per query.
list queries before tablesthe inversion from relational