nak / counter-drift-race.md

Last active 1 week ago

Like 0
Unlisted

gpt-6-astra's assessment on why post counts drifted from reality

nak revised this gist 1 week ago ยท 0f63eee

1 file changed, 30 insertions

counter-drift-race.md (file created)
@@ -0,0 +1,30 @@
1 + # Reproduced cause of possible counter drift
2 +
3 + ## Finding
4 +
5 + A deterministic synthetic two-connection test reproduces understated profile post counts using the deployed `delete_post` SQL sequence under READ COMMITTED, the live database's confirmed default isolation level. Both deletions commit successfully; the second counts a row already removed by the first transaction.
6 +
7 + The reproduction uses a separate synthetic schema in the local restored database and removes that schema afterward. It does not operate on public Mitra tables or execute the Rust binary. Script: `reproduce-counter-race.py`; output: `counter-race.log`.
8 +
9 + ## Interleaving
10 +
11 + Start with actor 1 having three stored posts: a reply and two unrelated posts. Actor 2 has the reply's parent post.
12 +
13 + 1. Transaction A selects the reply for deletion and subtracts one from actor 1's counter (3 โ†’ 2), retaining the profile row lock.
14 + 2. Transaction B selects the parent and its reply. Its counter UPDATE starts with both posts visible but waits on A's profile row lock.
15 + 3. A deletes the reply, decrements the parent's reply count, and commits.
16 + 4. B resumes. Its UPDATE's earlier statement snapshot counted the reply, so it subtracts one again from actor 1's now-current counter (2 โ†’ 1). It also decrements actor 2's counter.
17 + 5. B deletes the parent and commits. The reply is already gone, so it is not deleted twice, but it was counted twice.
18 +
19 + Observed final result: actor 1 has stored count 1 and actual post count 2. Both transactions completed without constraint errors. Enough repetitions can later make an unrelated cleanup subtraction violate the nonnegative counter constraint.
20 +
21 + ## Source
22 +
23 + - Deployed Mitra 5.10.0 `mitra_models/src/posts/queries.rs:1907`: `delete_post` first enumerates descendants/reposts, then subtracts counts from profiles, then deletes the root. The selected posts are not locked before the accounting operation.
24 + - `mitra_models/src/profiles/queries.rs:550`: `delete_profile` uses the same enumerate/subtract/delete pattern. This is another path to audit; the executable reproduction specifically tests two overlapping post deletions.
25 + - `mitra_activitypub/src/handlers/delete.rs`: both federated Delete(Note) and Delete(Person) reach these operations. No surrounding serialization lock was found in that handler. The background retention worker also calls `delete_post`, providing another possible concurrent deletion source.
26 + - `delete_repost` also appears to omit decrementing its author's post counter despite creation incrementing it. That would produce overcounts, not the undercounts repaired here; it is a separate static-review observation, not the demonstrated race.
27 +
28 + ## Limits and next step
29 +
30 + This proves a counter-drift mechanism in the deployed SQL, not that particular users' deletion activity caused the seven historical inconsistencies. Busy deletion periods involving overlapping threads are consistent with the mechanism. A patch needs concurrency-safe accounting across deletion paths and a regression test reproducing this interleaving; simply skipping failed cleanup targets would not prevent new drift. No code patch has been applied.