Cassandra Runbook

FOR OPERATORS

What to check when the cluster is already in trouble.

Short notes from production incidents. Each one is a decision you can make on a live ring: what the error actually means, which node to test, and what not to run.

NODE DOWN

Repair, or replace-address?

A dead node is not automatically a replace. The choice depends on how long it has been down, whether its data is still worth keeping, and what the rest of the ring is doing.

Read the decision tree

SAI INDEX

ReadFailure: INDEX_NOT_AVAILABLE

An indexed query failed because a replica could not serve its storage-attached index. Test each node at LOCAL_ONE, then rebuild the index only where it failed.

Read the check order

READ PATH

Tombstones overwhelming a read

The table is up. The query scanned deletion markers until Cassandra warned or failed. Bound the read, then compact only the table the log named.

Read the check order