SQL Server, Swiss money-transfer platform
From 288 deadlocks a month to about four
Read Committed Snapshot Isolation and a redesign of the worst transactions, but the part that made it possible was capturing the deadlocks properly first.
The problem
Month-end was the bad time. The platform was logging 288 deadlocks a month, and 53 on the worst single day. Transfers failed with errors that meant nothing to the cashier, support retried them by hand, and nobody could say which transactions were colliding, because nothing was capturing them.
What I did
I set up an Extended Events session to record every deadlock graph in production. That gave me a ranked list of what was actually colliding, which is a better basis for a decision than anyone's theory about it. Two changes came out of the list. The database moved to Read Committed Snapshot Isolation so readers stopped blocking writers, and the transactions at the top of the list were redesigned.
Where it landed
About four deadlocks a month, down from 288. The capture still runs in production and pages someone when a new pattern appears, so the next regression shows up in days instead of at the next month-end close.
Evidence
Extended Events captures from before and after. The alerting is still live.