Environment
- YugabyteDB YSQL
- YugabyteDB YCQL
- All supported versions. Behaviour changed in v2024.1, v2025.1.1.0 and v2025.2. Each change is marked below with the version it applies to.
Issue
YugabyteDB returns some errors that the client can safely retry. These errors do not mean the data is wrong or the cluster is unhealthy. They mean the operation met a conflict, a schema change, or a stale read snapshot, and the same operation can succeed on a fresh attempt.
This article lists the retryable errors, explains what the server already retries for you, and gives the options to reduce how often they reach the application.
Retryable errors
| Error text | API | SQLSTATE | Cause |
|---|---|---|---|
could not serialize access due to concurrent update |
YSQL | 40001 |
Two transactions modified the same rows. |
Conflicts with higher priority transaction |
YSQL | 40001 |
Fail-on-Conflict mode aborted the lower priority transaction. |
Transaction expired or aborted by a conflict |
YSQL | 40001 |
The transaction lost its conflict resolution or its heartbeat lapsed. |
Transaction was recently aborted |
YSQL | 40001 |
The transaction was already aborted by a conflict. |
Restart read required at |
YSQL | 40001 |
The read snapshot is no longer valid because of a concurrent commit. |
Catalog Version Mismatch: A DDL occurred while processing this query |
YSQL and YCQL | 40001 |
A DDL statement changed the catalog while this query was running. |
schema version mismatch for table |
YSQL and YCQL | 40001 |
The table schema changed while this operation was running. |
Value write after transaction start |
YSQL | 40001 |
A value was written after the transaction took its snapshot. |
deadlock detected |
YSQL | 40P01 |
Two or more transactions wait on each other in a cycle. Available when wait queues are enabled, which is the default from v2024.1. |
All transparent retries exhausted |
YSQL | 40001 |
The server retried internally up to yb_max_query_layer_retries and still could not make progress. |
Operation failed. Try again |
YCQL | Not applicable | Generic retryable condition on the YCQL API. |
Errors that are NOT retryable
Retrying these wastes work and can hide a real defect. Fix the cause instead.
| Error text | SQLSTATE | What to do |
|---|---|---|
duplicate key value violates unique constraint |
23505 |
The row already exists. Fix the application logic or the input data. |
current transaction is aborted, commands ignored until end of transaction block |
25P02 |
Issue ROLLBACK. Do not send more statements on that transaction. |
terminating connection due to idle-in-transaction timeout |
25P03 |
Reconnect. Replay the work only if it is safe to replay. |
| Syntax, permission, and undefined object errors | 42XXX |
Fix the SQL, the object reference, or the grants. |
| Internal errors | XX000 |
Do not retry on the SQLSTATE alone. See the note below. |
Some distributed conditions surface as XX000 because they do not map to a standard PostgreSQL SQLSTATE. Examples are Tablet leader changed during async write and Call waited in the queue past deadline. Retry these only when the message matches a condition you have already confirmed as transient, and only when the operation is safe to replay. Never write a rule of the form IF SQLSTATE = 'XX000' THEN retry, because that hides real internal failures.
Examples
ERROR: Operation failed. Try again.: [Operation failed. Try again. Conflicts with higher priority transaction: (transaction error 3)] (SQLSTATE 40001) ERROR: Catalog Version Mismatch: A DDL occurred while processing this query. Try Again ERROR: schema version mismatch for table 000033cf000030008000000000004000: expected 4, got 3 CONTEXT: Catalog Version Mismatch: A DDL occurred while processing this query. Try again. ERROR: All transparent retries exhausted. could not serialize access due to concurrent update ERROR: deadlock detected (...)
What the server already retries
Before you add retry logic, know which retries YugabyteDB performs for you. This decides how much the application still has to handle.
| Condition | Server behaviour | Version |
|---|---|---|
| Conflict on the first statement of a transaction | The server retries with exponential backoff, using a newer snapshot on each attempt. On failure it returns All transparent retries exhausted. |
All supported versions |
| Conflict on a later statement of a transaction | No server retry. The client must roll back and retry the whole transaction. | Repeatable Read and Serializable, all versions |
| Conflict in Read Committed isolation | The server retries the whole statement internally. The client does not need to handle 40001. |
Requires yb_enable_read_committed_isolation=true
|
Response larger than ysql_output_buffer_size
|
No retry is performed, because part of the result is already sent to the client. | All supported versions |
The retry count is controlled by the YSQL configuration parameter yb_max_query_layer_retries, default 60. From v2.21 and v2024.1 this single parameter replaced ysql_max_read_restart_attempts and ysql_max_write_restart_attempts. Set it per session with SET, or cluster wide with the YB-TServer flag ysql_pg_conf_csv.
Check the current value:
SHOW yb_max_query_layer_retries;
Resolution
1. Add retry logic in the application
This is the primary fix for 40001 and 40P01. Roll back, wait with backoff and jitter, then retry the full transaction.
max_attempts = 10 # max number of retries
sleep_time = 0.002 # 2 ms - base sleep time
backoff = 2 # exponential multiplier
attempt = 0
while attempt < max_attempts:
attempt += 1
try:
cursor = cxn.cursor()
cursor.execute("BEGIN")
# Execute transaction statements here
cursor.execute("COMMIT")
break
except psycopg2.errors.SerializationFailure as e:
cursor.execute("ROLLBACK")
if attempt < max_attempts:
time.sleep(sleep_time)
sleep_time *= backoffTwo cases where you must not retry:
- Unfamiliar errors, such as internal errors. A blind retry can duplicate work.
-
COMMITand auto-commit statements. You cannot tell whether the failure happened before or after the commit.
Match the retry unit to the application unit of work. A standalone statement is already an implicit transaction in YugabyteDB. Do not wrap it in BEGIN and COMMIT only to make the retry code look uniform, because that adds distributed transaction overhead.
2. Use Read Committed isolation
In Read Committed isolation the server retries the statement internally, so the application does not need to handle 40001. Only Repeatable Read and Serializable need client retry logic.
- Read Committed needs the YB-TServer flag
yb_enable_read_committed_isolation=true. - From v2025.2, new universes deployed with yugabyted, YugabyteDB Anywhere, or YugabyteDB Aeon have it enabled by default.
- Before v2025.2, and on manually deployed universes, the flag defaults to
false. In that caseREAD COMMITTEDfalls back to the stricter Snapshot isolation.
Check the effective isolation level:
SHOW default_transaction_isolation;
3. Confirm the concurrency control mode
Wait-on-Conflict is the default from v2024.1. Conflicting transactions wait for each other instead of aborting on priority, which lowers P99 latency and matches PostgreSQL semantics. It also enables deadlock detection, so 40P01 can be returned.
Before v2024.1, or wherever the flag was turned off, the cluster runs in priority based Fail-on-Conflict mode. That mode produces the Conflicts with higher priority transaction errors, and it does not detect deadlocks quickly.
Check the flag on a YB-TServer:
curl -s http://<tserver-ip:9000/varz?raw | grep -E "enable_wait_queues|yb_enable_read_committed_isolation"
A change to enable_wait_queues needs a rolling restart.
4. Use SAVEPOINT to avoid discarding the whole transaction
Savepoints are supported from v2.8. Set a savepoint before a statement that can fail, and roll back only to that point instead of aborting the whole transaction.
BEGIN; -- earlier statements SAVEPOINT before_insert; INSERT INTO txndemo VALUES (1, 30); -- on UniqueViolation: ROLLBACK TO SAVEPOINT before_insert; UPDATE txndemo SET v = 30 WHERE k = 1; COMMIT;
5. Use explicit locking to serialise access
SELECT ... FOR UPDATE takes a row lock and stops other transactions from modifying those rows until the transaction ends. This converts an abort into a wait, at the cost of reduced concurrency.
6. Reduce contention in the schema and the workload
Retries treat the symptom. Frequent retries point to a cause worth fixing.
- Spread hot keys. A single counter row or a monotonically increasing primary key concentrates all writes on one tablet.
- Keep transactions short. A long transaction holds provisional records and widens the conflict window.
- Batch related writes so that fewer transactions touch the same rows.
7. Handle Catalog Version Mismatch
This error means a DDL statement ran while the query was in flight.
- Run DDL statements one at a time, not in parallel.
- Retry the failed DML.
-
ANALYZEis a DDL-style operation in YugabyteDB, because it updates catalog metadata. Automatic analyze runs in the background and can collide with user DDL. - Event triggers can turn an ordinary user DDL into a breaking, cluster-wide catalog change with no sign in the statement text. Check
pg_event_triggerbefore you conclude that the application's own SQL is the whole story. See the MISMATCHED_SCHEMA article linked below.
Table-level locks (Early Access, v2025.1.1.0 and later). With object locking enabled, conflicting DDL waits instead of failing with a catalog version mismatch. The feature is disabled by default and needs two preview flags:
allowed_preview_flags_csv=enable_object_locking_for_table_locks,ysql_yb_ddl_transaction_block_enabled enable_object_locking_for_table_locks=true ysql_yb_ddl_transaction_block_enabled=true
Confirm both flags exist in your build before you plan a change, because preview flag names differ across release lines:
curl -s http://<tserver-ip:9000/varz?raw | grep -E "enable_object_locking_for_table_locks|ysql_yb_ddl_transaction_block_enabled"
Observe waiting object locks:
SELECT * FROM pg_locks WHERE NOT granted;
8. Confirm a unique constraint violation is genuine
duplicate key value violates unique constraint is not retryable. Confirm the duplicate with a direct query before changing anything:
db_name=# select item_id from item_mapping group by item_id having count(item_id) 1; item_id -------------- (0 rows)
9. Measure a slow DDL before you change timeouts
If a DDL statement fails on a timeout, run it with timing enabled and record how long it takes.
\timing \set ECHO all
When retries are a symptom, not a solution
Track the retry rate in the application. A rising rate points to a cause worth fixing.
| Signal | Likely cause |
|---|---|
| Retries concentrated on a few keys | Hot key or hot tablet |
| Retries rise with concurrency | Row level contention |
Catalog Version Mismatch in bursts |
Concurrent DDL, migrations, or automatic analyze |
Restart read required on long reads |
Long running reads against a write heavy table |
All transparent retries exhausted |
Contention that outlasts the backoff budget. Raising yb_max_query_layer_retries only masks it. |
Additional Information
Related support articles:
- Catalog version mismatch
- Database Transactions errors out with "Restart read required"
- How to troubleshoot MISMATCHED_SCHEMA Errors in YugabyteDB (Including Hidden DDLs like ALTER ROLE)
Documentation:
Comments
0 comments
Please sign in to leave a comment.