Distributed Replicated Block Device (DRBD) development
 help / color / mirror / Atom feed
* [BUG] DRBD 9.3.4: acknowledgement counter rewind and stalled resync after failover
@ 2026-09-27  6:22 Amirhossein Rahimi
  2026-09-28 16:03 ` Philipp Reisner
  0 siblings, 1 reply; 2+ messages in thread
From: Amirhossein Rahimi @ 2026-09-27  6:22 UTC (permalink / raw)
  To: drbd-dev

Hi,

I'm reporting a recovery stall observed in a three-node DRBD 9.3.4 lab.
After a partition and failover, the returning former primary's
acknowledgement tracking moved backwards. A subsequent replica recovery
then remained stuck, although offline comparison showed identical data
on all three copies.

An experimental patch avoided the failure in one focused comparison.
Could you review the finding and advise whether there is an existing
fix or a preferred correction?

Environment
-----------
- DRBD 9.3.4, upstream revision:
  e59b287b199f7224468af85b68a10e83c68213f5
- Ubuntu, kernel 7.0.0-34-generic.
- Three diskful nodes, referred to here as A, B and C.
- Protocol A, majority quorum, suspended I/O on quorum loss, and native
  demotion of an outdated primary.
- Separate 32 MiB scratch LV on each node. No filesystem, applications
  or DRBD Reactor were involved in this reproducer.

Reproduction sequence
---------------------
1. C was Primary and completed synthetic synchronous 4 KiB writes.
2. A timed 75-second replication-only partition isolated C. A and B
   retained quorum; A promoted normally and accepted another write.
3. After connectivity returned, C demoted as obsolete. Its local write
   watermark remained 2704, but acknowledgement watermarks dropped
   from 2704 to 0.
4. B was taken offline. A accepted another 4 KiB write and demoted.
   B was then brought back.
5. B transferred the required data but remained Inconsistent/SyncTarget,
   with zero blocks outstanding against A. Its alternate path to C
   retained a 4 KiB out-of-sync indication.

At the stalled state, C was waiting for acknowledgement progress 2704
while recording 0; B was waiting for that flush acknowledgement before
completing resync. The peer connections were up and I/O queues were clear.

The stalled target and both sources were then stopped before recovery.
All 33,513,472 logical data bytes, excluding internal DRBD metadata,
matched by SHA256 across all three copies. In this reproduction, the
4 KiB indication was not a difference in stored data, but resync really
was unable to complete.

Suspected cause and experimental patch
-------------------------------------
A postponed write appears able to reach bitmap/peer-acknowledgement
cleanup before drbd_send_and_submit assigns its change tag. The request's
zero-initialized tag can then move acknowledgement tracking backwards,
leaving a wait for an earlier, higher watermark unsatisfied.

The attached experimental patch adds a dagtag_assigned flag, sets it at
actual tag assignment, and gates the relevant bitmap and peer-ack effects
on that flag. It does not simply reject numeric tag zero, which could be
valid after counter wrap. The patch retains the diagnostic and module
identification markers used during the lab comparison.

Only C temporarily ran the patched module; A and B retained their stock
builds. Those builds reported the same upstream revision, but B had a
different pre-existing srcversion, so these were not identical binaries.

In the patched run, the diagnostic caught an unsubmitted 4096-byte write
with tag=0 and state=110000. C's acknowledgement watermark stayed at 4080.
Repeating B's offline/write/return stage then completed recovery in about
seven seconds. All peers reached UpToDate/Established, bitmap counters
were zero, and offline full-data hashes matched again.

This was one focused stock-versus-patch comparison, with different write
counts but the same failure transition and follow-on recovery sequence.
It is not a full regression test. The patched module has been removed;
the lab was restored to stock and the patch has not been deployed in
production.

Could you confirm whether this diagnosis is consistent with the intended
request lifecycle, especially retry/reuse and acknowledgement ordering?
Is there already an upstream fix or supported workaround? Please also
advise what additional targeted diagnostics would be useful.

Thanks.

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-28 16:03 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27  6:22 [BUG] DRBD 9.3.4: acknowledgement counter rewind and stalled resync after failover Amirhossein Rahimi
2026-09-28 16:03 ` Philipp Reisner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox