This article discusses a scenario where the control process is stuck in initializing state due to exceeding tombstone_failure_threshold in Cassandra Database and its possible remediation.
Contrail-status output:
Contrail control
control: initializing (Database:Cassandra connection down, IFMap Server End-Of-RIB not computed, No BGP configuration for self) nodemgr: active named: active dns: active Cassandra: Inactive
Cassandra error logs (system.log and debug.log):
ERROR [SharedPool-Worker-1] 2019-10-19 23:00:11,295 MessageDeliveryTask.java:77 - Scanned over 100001 tombstones in config_db_uuid.obj_fq_name_table; 100000 columns were requested; query aborted (see tombstone_failure_threshold;)
This issue occurs if Tombstone_failure_threshold has reached its default value of 100,000. If the number of tombstones scanned by a query exceeds this number, Cassandra will terminate the query. This is a mechanism to prevent one or more nodes from running out of memory and crashing.
The recommendation is to reduce the default value (10 days) for gc_grace_seconds to 3-5 days (depending on the information below).
The gc_grace_period controls how often major compaction is performed, which clears tombstone entries. If it is set to 3-5 days, the tombstone entries will be cleared fast (ie 10/3 , 3 or 2 times faster). Things to note in this approach is that if a DB is down, it must bring back the node within the gc_grace_period to avoid the possibility of reseeding the affected node (ie., clear the data from affected node, sync data from other nodes and not use the data from affected node).
Nodetool compact can be performed followed by changing the gc_grace_seconds value during a maintenance window:
docker exec -it config_database_cassandra_1 /bin/bash nodetool -p 7201 status nodetool -p 7201 compact nodetool -p 7201 status - Ensure everything loos correct
For Cassandra config database:
Enter into container (config_database_cassandra_1)
Note: Container name may vary based on different environments used for installation.
# docker exec -it config_database_cassandra_1 /bin/bash
Connect to contrail config database # cqlsh <controller ip> 9041
Connected to contrail_database at 10.85.216.9:9041.
See the current values:
cqlsh> SELECT table_name,gc_grace_seconds FROM system_schema.tables WHERE keyspace_name='config_db_uuid'; table_name | gc_grace_seconds
obj_fq_name_table | 864000 obj_shared_table | 864000 obj_uuid_table | 864000
Change the desired values to 259200 seconds (3 days):
Note: This examples shows value for 3 days, it can be set between 3-5 days.
cqlsh> ALTER TABLE config_db_uuid.obj_fq_name_table WITH gc_grace_seconds = 259200; cqlsh> ALTER TABLE config_db_uuid.obj_shared_table WITH gc_grace_seconds = 259200; cqlsh> ALTER TABLE config_db_uuid.obj_uuid_table WITH gc_grace_seconds = 259200;
Verification :
obj_fq_name_table | 259200 obj_shared_table | 259200 obj_uuid_table | 259200
2020-02-03: Corrected the port number 7201