From: Mohamed Khalfella <mkhalfella@purestorage.com>
To: linux-block@vger.kernel.org
Cc: shinichiro.kawasaki@wdc.com, Keith Busch <kbusch@kernel.org>,
Jens Axboe <axboe@kernel.dk>, Christoph Hellwig <hch@lst.de>,
Sagi Grimberg <sagi@grimberg.me>, Hannes Reinecke <hare@suse.de>,
John Meneghini <jmeneghi@redhat.com>,
Jesse Taube <jtaubepe@redhat.com>,
Randy Jennings <randyj@purestorage.com>,
Dhaval Giani <dgiani@purestorage.com>,
Mohamed Khalfella <mkhalfella@purestorage.com>
Subject: [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics
Date: Wed, 16 Sep 2026 20:06:20 -0600 [thread overview]
Message-ID: <20260917020752.1672578-1-mkhalfella@purestorage.com> (raw)
When a write times out on an NVMe target today, the initiator resets the
path and retries the write on another one right away. Nothing has told
it the original command is dead. The command was never acknowledged, and
the reset is a local action that does not reach into the fabric to
retire it. So the retry lands, a later write to the same LBA lands after
it, and the original command finally arrives at the controller and
overwrites the newer data. A read now returns a write the host gave up
on. This is the ABA ghost write: the block holds X, then Y, then X
again, and no layer above notices.
The host side of that window is visible in dmesg. The write times out,
error recovery starts, and the controller is torn down and scheduled for
reconnect, all while the command is still outstanding at the target. The
retry has already gone out on the other path by this point:
[ 240.598176] nvme nvme4: I/O tag 113 (0071) type 4 opcode 0x1 (Write) QID 5 timeout
[ 240.600233] nvme nvme4: starting error recovery
[ 240.619821] nvme nvme4: Reconnecting in 10 seconds...
[ 245.677026] nvme nvme5: Removing ctrl: NQN "blktests-subsystem-1"
[ 245.973568] nvme nvme4: Removing ctrl: NQN "blktests-subsystem-1"
[ 246.328666] nvme nvme4: Property Set error: 880, offset 0x14
Nothing here waits for tag 113 to be retired, because there is no
mechanism that would let it.
It is silent data corruption. Both writes were reported as successful,
the host has no error to act on, and the block now holds data that was
superseded. Anything layered above, a filesystem or a database, has
already committed on the strength of the acknowledgement it received for
Y and has no reason to read the block back.
NVMe defines two mechanisms to close this window. CQT (Command Quiesce
Time) tells the host how long after a timeout it must wait before a
command is guaranteed to be retired, and CCR (Cross Controller Reset)
lets a host reset a controller through another controller in the
subsystem, fencing off the commands still outstanding on the path that
went away. Linux implements neither today, on the host or the target
side, so there is no bound on the lifetime of a timed out command and
nothing prevents the retry from being overtaken by the original.
These patches therefore add a test that fails. nvme/070 fails on tcp,
rdma and fc as of today, and that failure is the point: it is the
missing CCR/CQT support made visible and reproducible. It passes on loop
only because nvme-loop defines no timeout callback, so the write is
never timed out and never retried. The test should start passing on the
remaining transports once CCR/CQT support lands.
This is what the failure looks like on tcp, tested on commit
fd9beb887073 ("nvme-tcp.h: drop kernel-doc comments, fix a few
descriptions"), tag nvme-7.3-2026-09-03. The detector reads the block
back and finds a pattern other than the one written last, so validation
fails on the first iteration and the corruption is caught directly, in
results/nodev_tr_tcp/nvme/070.out.bad:
Running nvme/070
Target ports: 2
starting nvme-ghost-write-detector test program
iteration number 0, writing data
validating written data
validation failed
finished nvme-ghost-write-detector test program
Test complete
Reproducing the race needs an IO that can be held at the target for
longer than the host's io_timeout without being failed. The first three
patches give miniublk that ability, since there was previously no way to
talk to a running ublk server at all:
1/5 adds a loopback UDP control channel to the daemon, served by a
detached thread, one reply datagram per request
2/5 adds an INJECT_DELAY command which delays a given number of reads
or writes as an io_uring timeout on the queue's own ring
3/5 exposes it as "miniublk inject", which returns only once the
daemon has armed the delay, so a test can start IO immediately
without racing it
4/5 adds nvme-ghost-write-detector, which writes distinct byte
patterns to a single LBA and reads the block back, so a
resurfaced write is identifiable by the pattern that comes back
5/5 adds nvme/070, which exports a ublk-backed nvmet namespace
through two ports, connects the host to both, drops io_timeout to
2 seconds, holds one write in the backstore for 4, and runs the
detector
Mohamed Khalfella (5):
src/miniublk: add a control channel to the daemon
src/miniublk: add IO delay injection
src/miniublk: add the inject command
src/nvme-ghost-write-detector: add an ABA ghost write detector
nvme/070: test for ABA ghost writes on a multipath fabrics namespace
src/.gitignore | 1 +
src/Makefile | 1 +
src/miniublk.c | 376 +++++++++++++++++++++++++++++++-
src/nvme-ghost-write-detector.c | 86 ++++++++
tests/nvme/070 | 98 +++++++++
tests/nvme/070.out | 35 +++
6 files changed, 594 insertions(+), 3 deletions(-)
create mode 100644 src/nvme-ghost-write-detector.c
create mode 100755 tests/nvme/070
create mode 100644 tests/nvme/070.out
--
2.55.0
next reply other threads:[~2026-09-17 2:08 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 2:06 Mohamed Khalfella [this message]
2026-09-17 2:06 ` [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 2/5] src/miniublk: add IO delay injection Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 3/5] src/miniublk: add the inject command Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector Mohamed Khalfella
2026-09-21 18:38 ` Jesse Taube
2026-09-23 17:00 ` Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace Mohamed Khalfella
2026-09-22 17:07 ` Jesse Taube
2026-09-23 16:56 ` Mohamed Khalfella
2026-09-23 8:18 ` [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Shin'ichiro Kawasaki
2026-09-23 17:02 ` Mohamed Khalfella
2026-10-02 3:03 ` Shin'ichiro Kawasaki
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917020752.1672578-1-mkhalfella@purestorage.com \
--to=mkhalfella@purestorage.com \
--cc=axboe@kernel.dk \
--cc=dgiani@purestorage.com \
--cc=hare@suse.de \
--cc=hch@lst.de \
--cc=jmeneghi@redhat.com \
--cc=jtaubepe@redhat.com \
--cc=kbusch@kernel.org \
--cc=linux-block@vger.kernel.org \
--cc=randyj@purestorage.com \
--cc=sagi@grimberg.me \
--cc=shinichiro.kawasaki@wdc.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox