Part 3 · 1 chapters · ~8 min
Cassandra Data Modelling
Query-first modelling, partition keys and clustering columns and the physical layout they produce, bounding wide partitions with time buckets, denormalisation and write fan-out, materialised views and why they are experimental, secondary indexes and SAI, collections, counters and their costs, time series modelling, anti-patterns, and CQL as not-quite-SQL.
6
One table per query
code
CREATE TABLE transactions_by_account (
account_id uuid, month text, ts timestamp, txn_id timeuuid, amount_kobo bigint, type text,
PRIMARY KEY ((account_id, month), ts, txn_id)
) WITH CLUSTERING ORDER BY (ts DESC, txn_id DESC)
AND compaction = {'class': 'TimeWindowCompactionStrategy', 'compaction_window_unit': 'DAYS', 'compaction_window_size': 7};
SELECT * FROM transactions_by_account WHERE account_id = ? AND month = '2026-10' LIMIT 50; -- one partition, sequential
-- SELECT * FROM transactions_by_account WHERE type = 'refund'; → needs ALLOW FILTERING: a full cluster scanCQL looks like SQL and is not: no joins, no subqueries, WHERE only on partition keys and clustering prefixes (plus indexes), no foreign keys, and UPDATE is an upsert. Storage-attached indexes (SAI, Cassandra 5) make secondary queries far more practical than the old indexes, but the core rule stands: design tables for the queries.
QUERY-FIRST MODELLING
one table per query, partitioned by what you look up
swipe the figure sideways, or tap expand for full screen
1/5
queries first
In relational modelling you design entities, then write queries. In Cassandra you list the queries first, then design one table per query.
list queries before tablesthe inversion from relational