Cassandra Runbook

NOTE 02 · STORAGE-ATTACHED INDEX

ReadFailure: INDEX_NOT_AVAILABLE

A query that uses a storage-attached index failed because a replica could not serve that index. This is not a Solr error, and it is not proof the table is down. The coordinator asked one node for index results and that node refused.

The error

ReadFailure: received 0 responses and 1 failures
(INDEX_NOT_AVAILABLE=[/10.1.2.21])
error_code_map: {'10.1.2.21': '0x0002'}

0x0002 is INDEX_NOT_AVAILABLE. The address is the replica that failed the indexed read. Anonymize it in tickets and notes. The query was using an SAI index. A plain primary-key read against the same table does not take this path.

What a local check is for

The client error names one coordinator's failure. It does not tell you how many replicas have a bad index. Connect to each node and run the same query at LOCAL_ONE. That consistency level needs one replica in the local datacenter, and the node you connected to is the coordinator, so a failure stays on that node instead of being masked by a healthy replica.

One success does not clear the datacenter. In a six-node cluster, three nodes in each of two datacenters, the failure can sit in a single datacenter: two of the three nodes refuse the indexed read, and the third returns rows. The other datacenter can be clean. A client using LOCAL_QUORUM will keep hitting the bad pair whenever the coordinator is in the bad datacenter.

Reproduce it

  1. Take the application statement that failed. Keep the SAI predicate. Do not rewrite it as a primary-key read.
  2. From a shell on each node, open cqlsh against that node. Running it through another host defeats the check.
  3. Set the consistency, then run the query:
CONSISTENCY LOCAL_ONE;
SELECT ... FROM keyspace_name.table_name WHERE indexed_column = ...;
  1. Record pass or INDEX_NOT_AVAILABLE for every node. Do this in both datacenters before you rebuild anything.
  2. A node that returns rows has a usable SAI index for that query. A node that returns 0x0002 does not. Expect the bad nodes to cluster in one datacenter.

Rebuild only where it failed

SAI is built from the SSTables on that node. A cluster repair does not rebuild it. Replacing the node does not fix an index that is simply unreadable. On each node that failed the LOCAL_ONE check, rebuild that index:

nodetool rebuild_index keyspace_name table_name sai_index_name

The index name is the one from system_schema.indexes, not the column name, unless you created it with that name. Confirm the row before you run the command:

SELECT index_name, kind FROM system_schema.indexes
WHERE keyspace_name = 'keyspace_name' AND table_name = 'table_name';

Run it on the bad nodes in the affected datacenter. In the six-node case that was two nodes of three. Running it on the third node in that datacenter is reasonable if you want the datacenter uniform. Do not run it in the datacenter that already passed LOCAL_ONE.

One node at a time if the table is large. The rebuild reads local SSTables and writes a new index. Watch compaction and disk. When a node finishes, repeat the LOCAL_ONE query on that node before you start the next.

Done when

Every node that previously failed now returns rows at LOCAL_ONE. Then run the original client statement at the consistency the application uses. A retry before the rebuild finishes fails the same way. The base table was still readable by primary key the whole time.

Do not

Command names differ slightly by version. Confirm nodetool rebuild_index against the version you run. This is a decision aid, not a procedure for a specific cluster.

Back to notes