Tombstones overwhelming a read
The table is up. The query is not. Cassandra had to walk a large number of deletion markers to return a small result, and it stopped or warned before the read finished. This shows up on every supported version. The log line changes slightly. The decision does not.
The log
Read 12 live rows and 104382 tombstone cells for query
SELECT ... FROM keyspace_name.table_name WHERE ...
(see tombstone_warn_threshold)
A harder failure is TombstoneOverwhelmingException, or a client ReadTimeout whose server log shows the same scan. The warn threshold is a count of tombstone cells scanned, not a count of deleted partitions. Default warn is often 1000. Default failure is often 100000. Raising either number does not make the read cheap.
What it means
A tombstone is a marker that a cell, row, or range was deleted or that a TTL expired. Cassandra must read those markers to avoid returning data that should be gone. They are dropped later, at compaction, and only after gc_grace_seconds if repair is how you protect a delete. Until then every read of that slice pays for them.
The usual writers are a table used as a queue, inserts followed by deletes of the same rows, inserts of nulls, collection overwrites, range deletes, and a query that reads a whole partition to find a few live rows. TTLs on a table that is still read after expiry look the same on the read path.
Confirm the query, then the table
- Take the statement from the log. Note the keyspace, table, partition key, and how wide the slice is. A read of one partition with no clustering bound is the usual offender.
- Find which replica logged it. The client timeout names a coordinator. The tombstone count is on the replica that scanned. Check
system.logon the replicas for that token, not only the coordinator. - Reproduce at
LOCAL_ONEfromcqlshon that replica, with the same bounds the application used. If a narrow slice returns quickly and the unbound partition does not, the partition is the problem. - Turn on tracing for one attempt if the log does not name the table. Do not leave tracing on. One traced read is enough.
CONSISTENCY LOCAL_ONE;
TRACING ON;
SELECT ... FROM keyspace_name.table_name WHERE pk = ...;
See how bad the table is
On the replica that logged the scan:
nodetool tablestats keyspace_name.table_name
Older versions call this cfstats. Look at maximum tombstones per slice, droppable tombstone ratio, and average live cells. A high droppable ratio means compaction can reclaim them. A low ratio means they are inside gc_grace_seconds and a compaction will not delete them yet.
SSTable metadata tells you whether a few files hold the scan. sstablemetadata on the largest files for that table shows estimated droppable tombstones and the deletion time. You want the files the read actually opens, not a cluster-wide average.
Fix, in this order
| What you found | What to do |
|---|---|
| The query reads a whole partition to return a few rows | Bound it. Add the clustering column, or page a narrow range. This is the fix that helps today, before any compaction. |
| Droppable tombstone ratio is high on that table | Compact the table on the replica that scanned. Prefer a compaction of that table over a major compaction of the node. Then re-run the read. |
| Most tombstones are not yet droppable | Compaction will not save the read. Stop the writer, or narrow the query, and wait out gc_grace_seconds unless you already understand the repair consequence of lowering it. |
| The writer inserts nulls, overwrites collections, or deletes by range | Fix the write. A null is a tombstone. Replacing a collection deletes the old cells. The read path cannot undo that. |
| A TTL table is read long after the rows expired | Stop reading expired slices. Time-window compaction only helps if the table is actually using it and old windows are dropping. |
nodetool garbagecollect on that table is the narrower reclaim when you need tombstones gone from SSTables that compaction is not visiting. It is still a local disk operation. Run it on the replica that failed the read, then confirm with the same LOCAL_ONE query.
Done when
This incident is done when the application statement returns without a tombstone warning on the replicas that previously logged one. tablestats moving is not the same thing. The read is.
The problem is done only if the writes stop producing markers faster than compaction can drop them. A table with steady insert-and-delete churn will fail the same query again after the next wave. Operator steps buy time. They do not replace a design that deletes rows a read still has to scan. Let them know they need to review their design. Otherwise AppDev will treat a successful compaction as a permanent fix.
Do not
- Raise
tombstone_failure_thresholdand call it fixed. The read still scans the markers. You only moved the point where Cassandra gives up. - Run a full repair because a read saw tombstones. Repair can spread the markers to replicas that did not have them.>
- Lower
gc_grace_secondson a table you repair, unless you accept that a missed repair can resurrect deleted data. - Major-compact every table on the node. Compact the table the log named.
- Tell AppDev the table is healthy because one compaction succeeded. If the application keeps deleting rows in the slice it reads, the warning comes back.
- Paste the raw query, key, or application name into a public note.
Thresholds and the stats command differ by version. Cassandra 3.x uses cfstats. Later versions use tablestats. Confirm the failure threshold in cassandra.yaml on the version you run. This is a decision aid, not a procedure for a specific cluster.