* [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics
@ 2026-09-17 2:06 Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon Mohamed Khalfella
` (5 more replies)
0 siblings, 6 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
When a write times out on an NVMe target today, the initiator resets the
path and retries the write on another one right away. Nothing has told
it the original command is dead. The command was never acknowledged, and
the reset is a local action that does not reach into the fabric to
retire it. So the retry lands, a later write to the same LBA lands after
it, and the original command finally arrives at the controller and
overwrites the newer data. A read now returns a write the host gave up
on. This is the ABA ghost write: the block holds X, then Y, then X
again, and no layer above notices.
The host side of that window is visible in dmesg. The write times out,
error recovery starts, and the controller is torn down and scheduled for
reconnect, all while the command is still outstanding at the target. The
retry has already gone out on the other path by this point:
[ 240.598176] nvme nvme4: I/O tag 113 (0071) type 4 opcode 0x1 (Write) QID 5 timeout
[ 240.600233] nvme nvme4: starting error recovery
[ 240.619821] nvme nvme4: Reconnecting in 10 seconds...
[ 245.677026] nvme nvme5: Removing ctrl: NQN "blktests-subsystem-1"
[ 245.973568] nvme nvme4: Removing ctrl: NQN "blktests-subsystem-1"
[ 246.328666] nvme nvme4: Property Set error: 880, offset 0x14
Nothing here waits for tag 113 to be retired, because there is no
mechanism that would let it.
It is silent data corruption. Both writes were reported as successful,
the host has no error to act on, and the block now holds data that was
superseded. Anything layered above, a filesystem or a database, has
already committed on the strength of the acknowledgement it received for
Y and has no reason to read the block back.
NVMe defines two mechanisms to close this window. CQT (Command Quiesce
Time) tells the host how long after a timeout it must wait before a
command is guaranteed to be retired, and CCR (Cross Controller Reset)
lets a host reset a controller through another controller in the
subsystem, fencing off the commands still outstanding on the path that
went away. Linux implements neither today, on the host or the target
side, so there is no bound on the lifetime of a timed out command and
nothing prevents the retry from being overtaken by the original.
These patches therefore add a test that fails. nvme/070 fails on tcp,
rdma and fc as of today, and that failure is the point: it is the
missing CCR/CQT support made visible and reproducible. It passes on loop
only because nvme-loop defines no timeout callback, so the write is
never timed out and never retried. The test should start passing on the
remaining transports once CCR/CQT support lands.
This is what the failure looks like on tcp, tested on commit
fd9beb887073 ("nvme-tcp.h: drop kernel-doc comments, fix a few
descriptions"), tag nvme-7.3-2026-09-03. The detector reads the block
back and finds a pattern other than the one written last, so validation
fails on the first iteration and the corruption is caught directly, in
results/nodev_tr_tcp/nvme/070.out.bad:
Running nvme/070
Target ports: 2
starting nvme-ghost-write-detector test program
iteration number 0, writing data
validating written data
validation failed
finished nvme-ghost-write-detector test program
Test complete
Reproducing the race needs an IO that can be held at the target for
longer than the host's io_timeout without being failed. The first three
patches give miniublk that ability, since there was previously no way to
talk to a running ublk server at all:
1/5 adds a loopback UDP control channel to the daemon, served by a
detached thread, one reply datagram per request
2/5 adds an INJECT_DELAY command which delays a given number of reads
or writes as an io_uring timeout on the queue's own ring
3/5 exposes it as "miniublk inject", which returns only once the
daemon has armed the delay, so a test can start IO immediately
without racing it
4/5 adds nvme-ghost-write-detector, which writes distinct byte
patterns to a single LBA and reads the block back, so a
resurfaced write is identifiable by the pattern that comes back
5/5 adds nvme/070, which exports a ublk-backed nvmet namespace
through two ports, connects the host to both, drops io_timeout to
2 seconds, holds one write in the backstore for 4, and runs the
detector
Mohamed Khalfella (5):
src/miniublk: add a control channel to the daemon
src/miniublk: add IO delay injection
src/miniublk: add the inject command
src/nvme-ghost-write-detector: add an ABA ghost write detector
nvme/070: test for ABA ghost writes on a multipath fabrics namespace
src/.gitignore | 1 +
src/Makefile | 1 +
src/miniublk.c | 376 +++++++++++++++++++++++++++++++-
src/nvme-ghost-write-detector.c | 86 ++++++++
tests/nvme/070 | 98 +++++++++
tests/nvme/070.out | 35 +++
6 files changed, 594 insertions(+), 3 deletions(-)
create mode 100644 src/nvme-ghost-write-detector.c
create mode 100755 tests/nvme/070
create mode 100644 tests/nvme/070.out
--
2.55.0
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
@ 2026-09-17 2:06 ` Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 2/5] src/miniublk: add IO delay injection Mohamed Khalfella
` (4 subsequent siblings)
5 siblings, 0 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
There is currently no way to talk to a running ublk server. Every command
goes through /dev/ublk-control and is handled by the driver. Also, the
daemon only blocks on its per-queue io_uring.
Give it a loopback UDP socket served by a detached thread, one reply
datagram per request. The only command so far is PING, which returns the
device id and exists to exercise the transport. The port defaults to
61000 plus the device id. --control_port on "add" and "recover" overrides
it, and 0 disables it.
The listener is detached and never joined. It blocks in recvfrom() until
the process exits, and ublk_dev_unprep() closes the socket without
waking it, which is harmless because the daemon exits immediately
afterwards.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
src/miniublk.c | 153 ++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 151 insertions(+), 2 deletions(-)
diff --git a/src/miniublk.c b/src/miniublk.c
index b0c308b..30bd7f5 100644
--- a/src/miniublk.c
+++ b/src/miniublk.c
@@ -22,6 +22,8 @@
#include <sys/syscall.h>
#include <sys/mman.h>
#include <sys/ioctl.h>
+#include <sys/socket.h>
+#include <netinet/in.h>
#include <liburing.h>
#include <linux/ublk_cmd.h>
@@ -43,6 +45,22 @@
#define UBLK_DBG_CTRL_CMD (1U << 4)
#define UBLK_LOG (1U << 5)
+#define UBLK_CTRL_PORT_BASE 61000
+#define UBLK_CTRL_MSG_MAGIC 0x6b6c6275
+
+enum {
+ UBLK_CTRL_CMD_PING = 0,
+};
+
+struct ublk_ctrl_msg {
+ __u32 magic;
+ union {
+ __u32 cmd;
+ __s32 ret;
+ };
+ __u64 data[3];
+};
+
struct ublk_dev;
struct ublk_queue;
@@ -114,6 +132,10 @@ struct ublk_dev {
int ctrl_fd;
bool use_ioctl;
struct io_uring ring;
+
+ /* daemon control channel, 0 means it is disabled */
+ int ctrl_port;
+ int ctrl_sock;
};
#ifndef offsetof
@@ -486,6 +508,7 @@ static struct ublk_dev *ublk_ctrl_init()
int ret;
dev->use_ioctl = true; /* use ioctl opcodes by default */
+ dev->ctrl_sock = -1;
dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
if (dev->ctrl_fd < 0) {
@@ -607,6 +630,109 @@ static int ublk_queue_init(struct ublk_queue *q)
return -ENOMEM;
}
+static int ublk_ctrl_msg_handle(struct ublk_dev *dev,
+ const struct ublk_ctrl_msg *req, struct ublk_ctrl_msg *rsp)
+{
+ switch (req->cmd) {
+ case UBLK_CTRL_CMD_PING:
+ rsp->data[0] = dev->dev_info.dev_id;
+ return 0;
+ default:
+ ublk_dbg(UBLK_DBG_DEV, "%s: unknown control command %u\n",
+ __func__, req->cmd);
+ return -EINVAL;
+ }
+}
+
+static void *ublk_ctrl_sock_fn(void *data)
+{
+ struct ublk_dev *dev = data;
+ struct sockaddr_in peer;
+ struct ublk_ctrl_msg req;
+ struct ublk_ctrl_msg rsp;
+ socklen_t peer_len;
+ ssize_t len;
+
+ for (;;) {
+ peer_len = sizeof(peer);
+ len = recvfrom(dev->ctrl_sock, &req, sizeof(req), 0,
+ (struct sockaddr *)&peer, &peer_len);
+ if (len < 0) {
+ if (errno == EINTR)
+ continue;
+ ublk_err("dev %d: control socket recv failed: %s\n",
+ dev->dev_info.dev_id, strerror(errno));
+ break;
+ }
+
+ if (len != sizeof(req) || req.magic != UBLK_CTRL_MSG_MAGIC) {
+ ublk_dbg(UBLK_DBG_DEV,
+ "%s: dropped %zd byte datagram\n", __func__,
+ len);
+ continue;
+ }
+
+ memset(&rsp, 0, sizeof(rsp));
+ rsp.magic = UBLK_CTRL_MSG_MAGIC;
+ rsp.ret = ublk_ctrl_msg_handle(dev, &req, &rsp);
+
+ if (sendto(dev->ctrl_sock, &rsp, sizeof(rsp), 0,
+ (struct sockaddr *)&peer, peer_len) < 0)
+ ublk_dbg(UBLK_DBG_DEV, "%s: reply failed: %s\n",
+ __func__, strerror(errno));
+ }
+
+ return NULL;
+}
+
+static void ublk_ctrl_sock_init(struct ublk_dev *dev)
+{
+ struct sockaddr_in addr = {
+ .sin_family = AF_INET,
+ .sin_port = htons(dev->ctrl_port),
+ .sin_addr.s_addr = htonl(INADDR_LOOPBACK),
+ };
+ int dev_id = dev->dev_info.dev_id;
+ pthread_t thread;
+ int fd, ret;
+
+ if (!dev->ctrl_port)
+ return;
+
+ fd = socket(AF_INET, SOCK_DGRAM, 0);
+ if (fd < 0) {
+ ublk_err("dev %d: can't create control socket: %s\n",
+ dev_id, strerror(errno));
+ return;
+ }
+
+ if (bind(fd, (struct sockaddr *)&addr, sizeof(addr)) < 0) {
+ ublk_err("dev %d: can't bind control port %d: %s\n",
+ dev_id, dev->ctrl_port, strerror(errno));
+ close(fd);
+ return;
+ }
+
+ dev->ctrl_sock = fd;
+
+ /*
+ * The listener is detached and never joined. It blocks in recvfrom()
+ * until the daemon exits, and goes away with the process.
+ */
+ ret = pthread_create(&thread, NULL, ublk_ctrl_sock_fn, dev);
+ if (ret) {
+ ublk_err("dev %d: can't start control thread: %s\n",
+ dev_id, strerror(ret));
+ close(fd);
+ dev->ctrl_sock = -1;
+ return;
+ }
+ pthread_detach(thread);
+
+ ublk_log("dev %d: control socket listening on 127.0.0.1:%d\n",
+ dev_id, dev->ctrl_port);
+}
+
static int ublk_dev_prep(struct ublk_dev *dev)
{
int dev_id = dev->dev_info.dev_id;
@@ -621,6 +747,8 @@ static int ublk_dev_prep(struct ublk_dev *dev)
goto fail;
}
+ ublk_ctrl_sock_init(dev);
+
if (dev->dev_info.state != UBLK_S_DEV_QUIESCED && dev->tgt.ops->init_tgt)
ret = dev->tgt.ops->init_tgt(dev);
@@ -635,6 +763,11 @@ fail:
static void ublk_dev_unprep(struct ublk_dev *dev)
{
+ if (dev->ctrl_sock >= 0) {
+ close(dev->ctrl_sock);
+ dev->ctrl_sock = -1;
+ }
+
if (dev->tgt.ops->deinit_tgt)
dev->tgt.ops->deinit_tgt(dev);
close(dev->fds[0]);
@@ -968,6 +1101,7 @@ static int cmd_dev_add(int argc, char *argv[])
{ "queues", 1, NULL, 'q' },
{ "depth", 1, NULL, 'd' },
{ "recovery", 0, NULL, 'r' },
+ { "control_port", 1, NULL, 0},
{ "debug_mask", 1, NULL, 0},
{ "quiet", 0, NULL, 0},
{ NULL }
@@ -978,6 +1112,7 @@ static int cmd_dev_add(int argc, char *argv[])
int ret, option_idx, opt;
const char *tgt_type = NULL;
int dev_id = -1;
+ int ctrl_port = -1;
unsigned nr_queues = 2, depth = UBLK_QUEUE_DEPTH;
int user_recovery = 0;
@@ -1000,6 +1135,8 @@ static int cmd_dev_add(int argc, char *argv[])
user_recovery = 1;
break;
case 0:
+ if (!strcmp(longopts[option_idx].name, "control_port"))
+ ctrl_port = strtol(optarg, NULL, 10);
if (!strcmp(longopts[option_idx].name, "debug_mask"))
ublk_dbg_mask = strtol(optarg, NULL, 16);
if (!strcmp(longopts[option_idx].name, "quiet"))
@@ -1047,6 +1184,9 @@ static int cmd_dev_add(int argc, char *argv[])
goto fail;
}
+ if (ctrl_port < 0)
+ ctrl_port = UBLK_CTRL_PORT_BASE + dev->dev_info.dev_id;
+ dev->ctrl_port = ctrl_port;
ret = ublk_start_daemon(dev, false);
if (ret < 0) {
ublk_err("%s: can't start daemon id %d, type %s\n",
@@ -1066,6 +1206,7 @@ static int cmd_dev_recover(int argc, char *argv[])
static const struct option longopts[] = {
{ "type", 1, NULL, 't' },
{ "number", 1, NULL, 'n' },
+ { "control_port", 1, NULL, 0},
{ "debug_mask", 1, NULL, 0},
{ "quiet", 0, NULL, 0},
{ NULL }
@@ -1076,6 +1217,7 @@ static int cmd_dev_recover(int argc, char *argv[])
int ret, option_idx, opt;
const char *tgt_type = NULL;
int dev_id = -1;
+ int ctrl_port = -1;
while ((opt = getopt_long(argc, argv, "-:t:n:d:q:",
longopts, &option_idx)) != -1) {
@@ -1087,6 +1229,8 @@ static int cmd_dev_recover(int argc, char *argv[])
tgt_type = optarg;
break;
case 0:
+ if (!strcmp(longopts[option_idx].name, "control_port"))
+ ctrl_port = strtol(optarg, NULL, 10);
if (!strcmp(longopts[option_idx].name, "debug_mask"))
ublk_dbg_mask = strtol(optarg, NULL, 16);
if (!strcmp(longopts[option_idx].name, "quiet"))
@@ -1126,6 +1270,9 @@ static int cmd_dev_recover(int argc, char *argv[])
goto fail;
}
+ if (ctrl_port < 0)
+ ctrl_port = UBLK_CTRL_PORT_BASE + dev->dev_info.dev_id;
+ dev->ctrl_port = ctrl_port;
dev->tgt.ops = ops;
dev->tgt.argc = argc;
dev->tgt.argv = argv;
@@ -1309,16 +1456,18 @@ static int cmd_dev_list(int argc, char *argv[])
static int cmd_dev_help(int argc, char *argv[])
{
- printf("%s add -t {null|loop} [-q nr_queues] [-d depth] [-n dev_id] \n",
+ printf("%s add -t {null|loop} [-q nr_queues] [-d depth] [-n dev_id] [--control_port port] \n",
argv[0]);
printf("\t default: nr_queues=2(max 4), depth=128(max 128), dev_id=-1(auto allocation)\n");
printf("\t -t loop -f backing_file \n");
printf("\t -t null\n");
+ printf("\t --control_port default %d+dev_id, 0 disables the control socket\n",
+ UBLK_CTRL_PORT_BASE);
printf("%s del [-n dev_id] -a \n", argv[0]);
printf("\t -a delete all devices -n delete specified device\n");
printf("%s list [-n dev_id] -a \n", argv[0]);
printf("\t -a list all devices, -n list specified device, default -a \n");
- printf("%s recover -t {null|loop} [-n dev_id] \n", argv[0]);
+ printf("%s recover -t {null|loop} [-n dev_id] [--control_port port] \n", argv[0]);
printf("\t -t loop -f backing_file \n");
printf("\t -t null\n");
return 0;
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH blktests 2/5] src/miniublk: add IO delay injection
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon Mohamed Khalfella
@ 2026-09-17 2:06 ` Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 3/5] src/miniublk: add the inject command Mohamed Khalfella
` (3 subsequent siblings)
5 siblings, 0 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
Tests that exercise error handling in a block driver stacked on ublk need
a way to hold an IO back without failing it. Add an INJECT_DELAY control
command which delays a given number of reads or writes before they are
issued to the target.
The delay is an io_uring timeout on the same ring rather than a sleep in
the queue thread, marked with UBLK_DELAY_MARK in tgt_data to tell its
completion apart from a real IO. It is counted in io_inflight to be
waited for.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
src/miniublk.c | 111 ++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 109 insertions(+), 2 deletions(-)
diff --git a/src/miniublk.c b/src/miniublk.c
index 30bd7f5..8dc2a56 100644
--- a/src/miniublk.c
+++ b/src/miniublk.c
@@ -19,6 +19,7 @@
#include <limits.h>
#include <string.h>
#include <errno.h>
+#include <stdatomic.h>
#include <sys/syscall.h>
#include <sys/mman.h>
#include <sys/ioctl.h>
@@ -49,7 +50,8 @@
#define UBLK_CTRL_MSG_MAGIC 0x6b6c6275
enum {
- UBLK_CTRL_CMD_PING = 0,
+ UBLK_CTRL_CMD_PING = 0,
+ UBLK_CTRL_CMD_INJECT_DELAY = 1,
};
struct ublk_ctrl_msg {
@@ -84,6 +86,8 @@ struct ublk_io {
unsigned int flags;
unsigned int result;
+
+ struct __kernel_timespec delay_ts;
};
struct ublk_tgt_ops {
@@ -136,6 +140,10 @@ struct ublk_dev {
/* daemon control channel, 0 means it is disabled */
int ctrl_port;
int ctrl_sock;
+
+ atomic_int inject_op;
+ atomic_int inject_secs;
+ atomic_int inject_count;
};
#ifndef offsetof
@@ -190,6 +198,13 @@ static inline unsigned int user_data_to_op(__u64 user_data)
return (user_data >> 16) & 0xff;
}
+#define UBLK_DELAY_MARK 1
+
+static inline unsigned int user_data_to_tgt_data(__u64 user_data)
+{
+ return (user_data >> 24) & 0xff;
+}
+
static void ublk_err(const char *fmt, ...)
{
va_list ap;
@@ -509,6 +524,7 @@ static struct ublk_dev *ublk_ctrl_init()
dev->use_ioctl = true; /* use ioctl opcodes by default */
dev->ctrl_sock = -1;
+ dev->inject_op = -1;
dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
if (dev->ctrl_fd < 0) {
@@ -637,6 +653,26 @@ static int ublk_ctrl_msg_handle(struct ublk_dev *dev,
case UBLK_CTRL_CMD_PING:
rsp->data[0] = dev->dev_info.dev_id;
return 0;
+ case UBLK_CTRL_CMD_INJECT_DELAY: {
+ int op = req->data[0];
+ int secs = req->data[1];
+ int count = req->data[2];
+
+ if (op != UBLK_IO_OP_READ && op != UBLK_IO_OP_WRITE)
+ return -EINVAL;
+ if (secs <= 0 || count <= 0)
+ return -EINVAL;
+
+ dev->inject_op = op;
+ dev->inject_secs = secs;
+ /* arm the count last, the request becomes visible at once */
+ dev->inject_count = count;
+
+ ublk_log("dev %d: delay next %d %s io(s) by %d secs\n",
+ dev->dev_info.dev_id, count,
+ op == UBLK_IO_OP_READ ? "read" : "write", secs);
+ return 0;
+ }
default:
ublk_dbg(UBLK_DBG_DEV, "%s: unknown control command %u\n",
__func__, req->cmd);
@@ -773,6 +809,66 @@ static void ublk_dev_unprep(struct ublk_dev *dev)
close(dev->fds[0]);
}
+static int ublk_inject_claim(struct ublk_queue *q, unsigned int ublk_op)
+{
+ struct ublk_dev *dev = q->dev;
+ int secs;
+
+ if (dev->inject_count <= 0)
+ return 0;
+
+ if (dev->inject_op != (int)ublk_op)
+ return 0;
+
+ secs = dev->inject_secs;
+ if (secs <= 0)
+ return 0;
+
+ if (atomic_fetch_sub(&dev->inject_count, 1) <= 0) {
+ /* another queue took the last one */
+ atomic_fetch_add(&dev->inject_count, 1);
+ return 0;
+ }
+
+ return secs;
+}
+
+static int ublk_delay_io(struct ublk_queue *q, int tag)
+{
+ const struct ublksrv_io_desc *iod = ublk_get_iod(q, tag);
+ unsigned int ublk_op = ublksrv_get_op(iod);
+ struct ublk_io *io = &q->ios[tag];
+ struct io_uring_sqe *sqe;
+ int secs;
+
+ secs = ublk_inject_claim(q, ublk_op);
+ if (!secs)
+ return 0;
+
+ sqe = io_uring_get_sqe(&q->ring);
+ if (!sqe) {
+ ublk_err("%s: run out of sqe %d, tag %d\n", __func__,
+ q->q_id, tag);
+ atomic_fetch_add(&q->dev->inject_count, 1);
+ return 0;
+ }
+
+ io->delay_ts.tv_sec = secs;
+ io->delay_ts.tv_nsec = 0;
+
+ io_uring_prep_timeout(sqe, &io->delay_ts, 0, 0);
+ /* bit63 marks us as tgt io, tgt_data marks us as the delay timeout */
+ sqe->user_data = build_user_data(tag, ublk_op, UBLK_DELAY_MARK, 1);
+
+ q->io_inflight++;
+
+ ublk_dbg(UBLK_DBG_IO, "%s: dev %d q %d tag %d: delay %s io by %d secs\n",
+ __func__, q->dev->dev_info.dev_id, q->q_id, tag,
+ ublk_op == UBLK_IO_OP_READ ? "read" : "write", secs);
+
+ return 1;
+}
+
static int ublk_queue_io_cmd(struct ublk_queue *q,
struct ublk_io *io, unsigned tag)
{
@@ -896,6 +992,16 @@ static inline void ublksrv_handle_tgt_cqe(struct ublk_queue *q,
{
unsigned tag = user_data_to_tag(cqe->user_data);
+ /*
+ * A delay timeout completes with -ETIME. It is not a real io, so hand
+ * the tag back to queue_io() instead of completing it.
+ */
+ if (user_data_to_tgt_data(cqe->user_data) == UBLK_DELAY_MARK) {
+ q->io_inflight--;
+ q->tgt_ops->queue_io(q, tag);
+ return;
+ }
+
if (cqe->res < 0 && cqe->res != -EAGAIN)
ublk_err("%s: failed tgt io: res %d qid %u tag %u, cmd_op %u\n",
__func__, cqe->res, q->q_id,
@@ -937,7 +1043,8 @@ static void ublk_handle_cqe(struct io_uring *r,
if (cqe->res == UBLK_IO_RES_OK) {
ublk_assert(tag < q->q_depth);
- q->tgt_ops->queue_io(q, tag);
+ if (!ublk_delay_io(q, tag))
+ q->tgt_ops->queue_io(q, tag);
} else {
/*
* COMMIT_REQ will be completed immediately since no fetching
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH blktests 3/5] src/miniublk: add the inject command
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 2/5] src/miniublk: add IO delay injection Mohamed Khalfella
@ 2026-09-17 2:06 ` Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector Mohamed Khalfella
` (2 subsequent siblings)
5 siblings, 0 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
Add a CLI command which arms a delay on a running device:
miniublk inject -n dev_id -o {read|write} -d delay -c count
It sends one INJECT_DELAY datagram to the device's control port and
waits for the reply, so it returns only once the daemon has armed the
request. A test can therefore arm a delay and start IO immediately
without racing the daemon.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
src/miniublk.c | 114 +++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 114 insertions(+)
diff --git a/src/miniublk.c b/src/miniublk.c
index 8dc2a56..2dbd4aa 100644
--- a/src/miniublk.c
+++ b/src/miniublk.c
@@ -48,6 +48,7 @@
#define UBLK_CTRL_PORT_BASE 61000
#define UBLK_CTRL_MSG_MAGIC 0x6b6c6275
+#define UBLK_CTRL_RCV_TIMEO 3
enum {
UBLK_CTRL_CMD_PING = 0,
@@ -1561,6 +1562,114 @@ static int cmd_dev_list(int argc, char *argv[])
return 0;
}
+static int cmd_dev_inject(int argc, char *argv[])
+{
+ static const struct option longopts[] = {
+ { "number", 1, NULL, 'n' },
+ { "op", 1, NULL, 'o' },
+ { "delay", 1, NULL, 'd' },
+ { "count", 1, NULL, 'c' },
+ { "control_port", 1, NULL, 0},
+ { NULL }
+ };
+ struct timeval tv = { .tv_sec = UBLK_CTRL_RCV_TIMEO, };
+ struct ublk_ctrl_msg req = {};
+ struct ublk_ctrl_msg rsp;
+ struct sockaddr_in addr;
+ int dev_id = -1, op = -1, delay = 0, count = 1;
+ int ctrl_port = -1;
+ int fd, ret, option_idx, opt;
+
+ while ((opt = getopt_long(argc, argv, "n:o:d:c:",
+ longopts, &option_idx)) != -1) {
+ switch (opt) {
+ case 'n':
+ dev_id = strtol(optarg, NULL, 10);
+ break;
+ case 'o':
+ if (!strcmp(optarg, "read"))
+ op = UBLK_IO_OP_READ;
+ else if (!strcmp(optarg, "write"))
+ op = UBLK_IO_OP_WRITE;
+ break;
+ case 'd':
+ delay = strtol(optarg, NULL, 10);
+ break;
+ case 'c':
+ count = strtol(optarg, NULL, 10);
+ break;
+ case 0:
+ if (!strcmp(longopts[option_idx].name, "control_port"))
+ ctrl_port = strtol(optarg, NULL, 10);
+ break;
+ }
+ }
+
+ if (dev_id < 0 || op < 0 || delay <= 0 || count <= 0) {
+ ublk_err("%s: -n dev_id -o {read|write} -d delay -c count are required\n",
+ __func__);
+ return -EINVAL;
+ }
+
+ if (ctrl_port < 0)
+ ctrl_port = UBLK_CTRL_PORT_BASE + dev_id;
+
+ fd = socket(AF_INET, SOCK_DGRAM, 0);
+ if (fd < 0) {
+ ublk_err("%s: can't create socket: %s\n", __func__,
+ strerror(errno));
+ return -errno;
+ }
+ setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
+
+ memset(&addr, 0, sizeof(addr));
+ addr.sin_family = AF_INET;
+ addr.sin_port = htons(ctrl_port);
+ addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
+
+ /* connect so that a missing daemon shows up as ECONNREFUSED */
+ if (connect(fd, (struct sockaddr *)&addr, sizeof(addr)) < 0) {
+ ublk_err("%s: can't connect to port %d: %s\n", __func__,
+ ctrl_port, strerror(errno));
+ ret = -errno;
+ goto out;
+ }
+
+ req.magic = UBLK_CTRL_MSG_MAGIC;
+ req.cmd = UBLK_CTRL_CMD_INJECT_DELAY;
+ req.data[0] = op;
+ req.data[1] = delay;
+ req.data[2] = count;
+
+ if (send(fd, &req, sizeof(req), 0) != sizeof(req)) {
+ ublk_err("%s: can't send to port %d: %s\n", __func__,
+ ctrl_port, strerror(errno));
+ ret = -errno;
+ goto out;
+ }
+
+ if (recv(fd, &rsp, sizeof(rsp), 0) != sizeof(rsp)) {
+ ublk_err("%s: no reply from dev %d on port %d: %s\n", __func__,
+ dev_id, ctrl_port, strerror(errno));
+ ret = -errno;
+ goto out;
+ }
+
+ if (rsp.magic != UBLK_CTRL_MSG_MAGIC) {
+ ublk_err("%s: bad reply from port %d\n", __func__, ctrl_port);
+ ret = -EPROTO;
+ goto out;
+ }
+
+ ret = rsp.ret;
+ if (ret)
+ ublk_err("%s: dev %d rejected the request: %d\n", __func__,
+ dev_id, ret);
+out:
+ close(fd);
+ return ret;
+}
+
static int cmd_dev_help(int argc, char *argv[])
{
printf("%s add -t {null|loop} [-q nr_queues] [-d depth] [-n dev_id] [--control_port port] \n",
@@ -1577,6 +1686,9 @@ static int cmd_dev_help(int argc, char *argv[])
printf("%s recover -t {null|loop} [-n dev_id] [--control_port port] \n", argv[0]);
printf("\t -t loop -f backing_file \n");
printf("\t -t null\n");
+ printf("%s inject -n dev_id -o {read|write} -d delay -c count [--control_port port] \n",
+ argv[0]);
+ printf("\t delay the next <count> <read|write> ios by <delay> seconds\n");
return 0;
}
@@ -1915,6 +2027,8 @@ int main(int argc, char *argv[])
ret = cmd_dev_help(argc, argv);
else if (!strcmp(cmd, "recover"))
ret = cmd_dev_recover(argc, argv);
+ else if (!strcmp(cmd, "inject"))
+ ret = cmd_dev_inject(argc, argv);
out:
if (ret)
cmd_dev_help(argc, argv);
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
` (2 preceding siblings ...)
2026-09-17 2:06 ` [PATCH blktests 3/5] src/miniublk: add the inject command Mohamed Khalfella
@ 2026-09-17 2:06 ` Mohamed Khalfella
2026-09-21 18:38 ` Jesse Taube
2026-09-17 2:06 ` [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace Mohamed Khalfella
2026-09-23 8:18 ` [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Shin'ichiro Kawasaki
5 siblings, 1 reply; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
Consider a write X that is issued to an LBA and times out without ever
being acknowledged. The initiator retries it on another path, where it
lands. A write Y to the same LBA follows and lands as well. The original
X is still in flight somewhere in the fabric, and when it finally
reaches the device it overwrites Y. A read now returns X, a write the
initiator gave up on long ago.
Add a program to catch this. It opens the device O_DIRECT and, on each
iteration, issues a series of writes of distinct byte patterns to the
same LBA, the 4k block at offset 0, waits, then reads that block back
and checks every byte holds the pattern written last. Because the
patterns differ, the value that comes back identifies which write
reappeared.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
src/.gitignore | 1 +
src/Makefile | 1 +
src/nvme-ghost-write-detector.c | 86 +++++++++++++++++++++++++++++++++
3 files changed, 88 insertions(+)
create mode 100644 src/nvme-ghost-write-detector.c
diff --git a/src/.gitignore b/src/.gitignore
index e9869e1..9673be8 100644
--- a/src/.gitignore
+++ b/src/.gitignore
@@ -15,6 +15,7 @@
/zbdioctl
/miniublk
/nvme-passthrough-meta
+/nvme-ghost-write-detector
/ioctl-lbmd-query
/nvme-passthru-admin-uring
/nvme-delay-ioctl
diff --git a/src/Makefile b/src/Makefile
index dd64694..92b4d0c 100644
--- a/src/Makefile
+++ b/src/Makefile
@@ -24,6 +24,7 @@ C_TARGETS := \
mount_clear_sock \
nvme-delay-ioctl \
nvme-passthrough-meta \
+ nvme-ghost-write-detector \
ioctl-lbmd-query \
nbdsetsize \
openclose \
diff --git a/src/nvme-ghost-write-detector.c b/src/nvme-ghost-write-detector.c
new file mode 100644
index 0000000..bd42dde
--- /dev/null
+++ b/src/nvme-ghost-write-detector.c
@@ -0,0 +1,86 @@
+// SPDX-License-Identifier: GPL-3.0+
+// Copyright (C) 2026 Mohamed Khalfella
+
+#define _GNU_SOURCE
+#include <stdio.h>
+#include <stdlib.h>
+#include <fcntl.h>
+#include <unistd.h>
+#include <string.h>
+#include <malloc.h>
+#include <errno.h>
+#include <libgen.h>
+
+#define BUF_SIZE 4096
+#define ITERATIONS 10
+#define DELAY 3 /* seconds delay between iterations */
+#define WRITE_COUNT 10
+
+#define WRITE_OFFSET 0
+#define READ_OFFSET WRITE_OFFSET
+
+int main(int argc, char **argv)
+{
+ int fd, i, w, off, ret;
+ char *buff;
+
+ fprintf(stdout, "starting %s test program\n", basename(argv[0]));
+
+ if (argc < 2) {
+ fprintf(stderr, "usage: %s /dev/nvmeXnY", argv[0]);
+ return 1;
+ }
+
+ fd = open(argv[1], O_RDWR | O_DIRECT);
+ if (fd < 0) {
+ fprintf(stderr, "failed to open device, errno = %d\n", errno);
+ return 1;
+ }
+
+ ret = posix_memalign((void **)&buff, BUF_SIZE, BUF_SIZE);
+ if (ret) {
+ fprintf(stderr, "failed to allocate buffer, ret = %d\n", ret);
+ goto out;
+ }
+
+ for (i = 0; i < ITERATIONS; i++) {
+ fprintf(stdout, "iteration number %d, writing data\n", i);
+
+ for (w = 0; w < WRITE_COUNT; w++) {
+ memset(buff, w, BUF_SIZE);
+ ret = pwrite(fd, buff, BUF_SIZE, WRITE_OFFSET);
+ if (ret != BUF_SIZE) {
+ fprintf(stderr, "failed to write buff, "
+ "ret = %d, errno = %d\n",
+ ret, errno);
+ goto out;
+ }
+ }
+
+ sleep(5);
+ fprintf(stdout, "validating written data\n");
+
+ ret = pread(fd, buff, BUF_SIZE, READ_OFFSET);
+ if (ret != BUF_SIZE) {
+ fprintf(stderr, "failed to read buff, "
+ "ret = %d, errno = %d\n",
+ ret, errno);
+ goto out;
+ }
+
+ for (off = 0; off < BUF_SIZE; off++) {
+ if (buff[off] != WRITE_COUNT - 1) {
+ fprintf(stdout, "validation failed\n");
+ goto out;
+ }
+ }
+
+ fprintf(stdout, "successfully validated\n");
+ sleep(DELAY);
+ }
+
+out:
+ fprintf(stdout, "finished %s test program\n", basename(argv[0]));
+ close(fd);
+ return ret;
+}
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
` (3 preceding siblings ...)
2026-09-17 2:06 ` [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector Mohamed Khalfella
@ 2026-09-17 2:06 ` Mohamed Khalfella
2026-09-22 17:07 ` Jesse Taube
2026-09-23 8:18 ` [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Shin'ichiro Kawasaki
5 siblings, 1 reply; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-17 2:06 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
An unacknowledged write that is retried on another path can still be
alive in the fabric. If it reaches the target after a later write to the
same LBA has landed, it overwrites it, and a read returns stale data.
Nothing in the tree exercises that window.
Add a test that builds it deliberately. A ublk loop device backs an
nvmet namespace exported through two ports, and the host connects to
both, so nvme-multipath has a second path to fail over to. io_timeout on
the subsystem drops to 2 seconds, then miniublk's inject command holds
one write in the backstore for 4 seconds. The host times that write out,
retries it on the other path, and the held write completes at the target
afterwards.
nvme-ghost-write-detector then writes distinct patterns to a single LBA
and reads the block back, so a resurfaced write shows up as the wrong
pattern.
The test requires nvme_core.multipath=Y and a fabrics transport. As of
today it passes on loop, which defines no timeout callback and so never
times the write out and never retries it, and fails on tcp, rdma and fc.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
tests/nvme/070 | 98 ++++++++++++++++++++++++++++++++++++++++++++++
tests/nvme/070.out | 35 +++++++++++++++++
2 files changed, 133 insertions(+)
create mode 100755 tests/nvme/070
create mode 100644 tests/nvme/070.out
diff --git a/tests/nvme/070 b/tests/nvme/070
new file mode 100755
index 0000000..e9a4690
--- /dev/null
+++ b/tests/nvme/070
@@ -0,0 +1,98 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-3.0+
+# Copyright (C) 2026 Mohamed Khalfella
+
+. tests/nvme/rc
+. common/ublk
+
+DESCRIPTION="Test injecting delay on nvme-target backstore and expect no corruption"
+
+requires() {
+ _nvme_requires
+ _have_loop
+ _have_ublk
+ _have_module_param_value nvme_core multipath Y
+ _require_nvme_trtype_is_fabrics
+ _have_src_program nvme-ghost-write-detector
+}
+
+set_conditions() {
+ _set_nvme_trtype "$@"
+}
+
+count_paths_to_subsystem() {
+ local subsysnqn="$1"
+ local dev count
+
+ count=0
+ for dev in /sys/class/nvme/nvme*; do
+ [[ -e "${dev}/subsysnqn" ]] || continue
+ [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
+ count=$(( count + 1 ))
+ done
+ echo "${count}"
+}
+
+set_io_timeout_of_subsystem() {
+ local subsysnqn="$1"
+ local timeout="$2"
+ local dev
+
+ for dev in /sys/class/nvme/nvme*; do
+ [[ -e "${dev}/subsysnqn" ]] || continue
+ [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
+ if ! echo "${timeout}" > "${dev}/io_timeout" 2> /dev/null; then
+ echo "FAIL: can not set io_timeout on ${dev##*/}"
+ return 1
+ fi
+ done
+}
+
+test() {
+ echo "Running ${TEST_NAME}"
+
+ local ns port nr_paths
+ local -a ports
+
+ if ! _init_ublk; then
+ return 1
+ fi
+
+ truncate -s "${NVME_IMG_SIZE}" "${TMPDIR}/ublk-img"
+ if ! ${UBLK_PROG} add -t loop -f "${TMPDIR}/ublk-img" -n 0 > "$FULL" 2>&1; then
+ echo "fail to add ublk device"
+ _exit_ublk
+ return 1
+ fi
+ udevadm settle
+
+ _setup_nvmet
+ _nvmet_target_setup --ports 2 --blkdev none
+ _create_nvmet_ns --blkdev /dev/ublkb0 \
+ --uuid "${def_subsys_uuid}" > /dev/null
+
+ _get_nvmet_ports "${def_subsysnqn}" ports
+ echo "Target ports: ${#ports[@]}"
+ for port in "${ports[@]}"; do
+ _nvme_connect_subsys --port "${port}"
+ done
+
+ nr_paths=$(count_paths_to_subsystem "${def_subsysnqn}")
+ if (( nr_paths != 2 )); then
+ echo "FAIL: expected 2 paths, found ${nr_paths}"
+ fi
+
+ set_io_timeout_of_subsystem "${def_subsysnqn}" 2000
+
+ if ! ${UBLK_PROG} inject -n 0 -o write -d 4 -c 1 >> "$FULL" 2>&1; then
+ echo "FAIL: can not inject write delay"
+ fi
+
+ ns=$(_find_nvme_ns "${def_subsys_uuid}")
+ "$SRCDIR/nvme-ghost-write-detector" "/dev/${ns}"
+
+ _nvme_disconnect_subsys
+ _nvmet_target_cleanup
+ _exit_ublk
+ echo "Test complete"
+}
diff --git a/tests/nvme/070.out b/tests/nvme/070.out
new file mode 100644
index 0000000..b43fec7
--- /dev/null
+++ b/tests/nvme/070.out
@@ -0,0 +1,35 @@
+Running nvme/070
+Target ports: 2
+starting nvme-ghost-write-detector test program
+iteration number 0, writing data
+validating written data
+successfully validated
+iteration number 1, writing data
+validating written data
+successfully validated
+iteration number 2, writing data
+validating written data
+successfully validated
+iteration number 3, writing data
+validating written data
+successfully validated
+iteration number 4, writing data
+validating written data
+successfully validated
+iteration number 5, writing data
+validating written data
+successfully validated
+iteration number 6, writing data
+validating written data
+successfully validated
+iteration number 7, writing data
+validating written data
+successfully validated
+iteration number 8, writing data
+validating written data
+successfully validated
+iteration number 9, writing data
+validating written data
+successfully validated
+finished nvme-ghost-write-detector test program
+Test complete
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector
2026-09-17 2:06 ` [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector Mohamed Khalfella
@ 2026-09-21 18:38 ` Jesse Taube
2026-09-23 17:00 ` Mohamed Khalfella
0 siblings, 1 reply; 13+ messages in thread
From: Jesse Taube @ 2026-09-21 18:38 UTC (permalink / raw)
To: Mohamed Khalfella
Cc: linux-block, shinichiro.kawasaki, Keith Busch, Jens Axboe,
Christoph Hellwig, Sagi Grimberg, Hannes Reinecke, John Meneghini,
Randy Jennings, Dhaval Giani
On Wed, Sep 16, 2026 at 10:09 PM Mohamed Khalfella
<mkhalfella@purestorage.com> wrote:
>
> Consider a write X that is issued to an LBA and times out without ever
> being acknowledged. The initiator retries it on another path, where it
> lands. A write Y to the same LBA follows and lands as well. The original
> X is still in flight somewhere in the fabric, and when it finally
> reaches the device it overwrites Y. A read now returns X, a write the
> initiator gave up on long ago.
>
> Add a program to catch this. It opens the device O_DIRECT and, on each
> iteration, issues a series of writes of distinct byte patterns to the
> same LBA, the 4k block at offset 0, waits, then reads that block back
> and checks every byte holds the pattern written last. Because the
> patterns differ, the value that comes back identifies which write
> reappeared.
>
> Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
> ---
> src/.gitignore | 1 +
> src/Makefile | 1 +
> src/nvme-ghost-write-detector.c | 86 +++++++++++++++++++++++++++++++++
> 3 files changed, 88 insertions(+)
> create mode 100644 src/nvme-ghost-write-detector.c
>
> diff --git a/src/.gitignore b/src/.gitignore
> index e9869e1..9673be8 100644
> --- a/src/.gitignore
> +++ b/src/.gitignore
> @@ -15,6 +15,7 @@
> /zbdioctl
> /miniublk
> /nvme-passthrough-meta
> +/nvme-ghost-write-detector
> /ioctl-lbmd-query
> /nvme-passthru-admin-uring
> /nvme-delay-ioctl
> diff --git a/src/Makefile b/src/Makefile
> index dd64694..92b4d0c 100644
> --- a/src/Makefile
> +++ b/src/Makefile
> @@ -24,6 +24,7 @@ C_TARGETS := \
> mount_clear_sock \
> nvme-delay-ioctl \
> nvme-passthrough-meta \
> + nvme-ghost-write-detector \
> ioctl-lbmd-query \
> nbdsetsize \
> openclose \
> diff --git a/src/nvme-ghost-write-detector.c b/src/nvme-ghost-write-detector.c
> new file mode 100644
> index 0000000..bd42dde
> --- /dev/null
> +++ b/src/nvme-ghost-write-detector.c
> @@ -0,0 +1,86 @@
> +// SPDX-License-Identifier: GPL-3.0+
> +// Copyright (C) 2026 Mohamed Khalfella
> +
> +#define _GNU_SOURCE
> +#include <stdio.h>
> +#include <stdlib.h>
> +#include <fcntl.h>
> +#include <unistd.h>
> +#include <string.h>
> +#include <malloc.h>
> +#include <errno.h>
> +#include <libgen.h>
We don't need to include malloc.h
> +
> +#define BUF_SIZE 4096
> +#define ITERATIONS 10
> +#define DELAY 3 /* seconds delay between iterations */
> +#define WRITE_COUNT 10
> +
> +#define WRITE_OFFSET 0
> +#define READ_OFFSET WRITE_OFFSET
> +
> +int main(int argc, char **argv)
> +{
> + int fd, i, w, off, ret;
> + char *buff;
> +
This is a bit pedantic but:
if (argc < 1)
return 1;
> + fprintf(stdout, "starting %s test program\n", basename(argv[0]));
> +
> + if (argc < 2) {
> + fprintf(stderr, "usage: %s /dev/nvmeXnY", argv[0]);
> + return 1;
> + }
> +
> + fd = open(argv[1], O_RDWR | O_DIRECT);
> + if (fd < 0) {
> + fprintf(stderr, "failed to open device, errno = %d\n", errno);
> + return 1;
> + }
> +
> + ret = posix_memalign((void **)&buff, BUF_SIZE, BUF_SIZE);
> + if (ret) {
> + fprintf(stderr, "failed to allocate buffer, ret = %d\n", ret);
> + goto out;
> + }
> +
> + for (i = 0; i < ITERATIONS; i++) {
> + fprintf(stdout, "iteration number %d, writing data\n", i);
> +
> + for (w = 0; w < WRITE_COUNT; w++) {
> + memset(buff, w, BUF_SIZE);
> + ret = pwrite(fd, buff, BUF_SIZE, WRITE_OFFSET);
> + if (ret != BUF_SIZE) {
> + fprintf(stderr, "failed to write buff, "
> + "ret = %d, errno = %d\n",
> + ret, errno);
> + goto out;
> + }
> + }
> +
> + sleep(5);
There should be a macro for this.
Thanks,
Jesse Taube
> + fprintf(stdout, "validating written data\n");
> +
> + ret = pread(fd, buff, BUF_SIZE, READ_OFFSET);
> + if (ret != BUF_SIZE) {
> + fprintf(stderr, "failed to read buff, "
> + "ret = %d, errno = %d\n",
> + ret, errno);
> + goto out;
> + }
> +
> + for (off = 0; off < BUF_SIZE; off++) {
> + if (buff[off] != WRITE_COUNT - 1) {
> + fprintf(stdout, "validation failed\n");
> + goto out;
> + }
> + }
> +
> + fprintf(stdout, "successfully validated\n");
> + sleep(DELAY);
> + }
> +
> +out:
> + fprintf(stdout, "finished %s test program\n", basename(argv[0]));
> + close(fd);
> + return ret;
> +}
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace
2026-09-17 2:06 ` [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace Mohamed Khalfella
@ 2026-09-22 17:07 ` Jesse Taube
2026-09-23 16:56 ` Mohamed Khalfella
0 siblings, 1 reply; 13+ messages in thread
From: Jesse Taube @ 2026-09-22 17:07 UTC (permalink / raw)
To: Mohamed Khalfella
Cc: linux-block, shinichiro.kawasaki, Keith Busch, Jens Axboe,
Christoph Hellwig, Sagi Grimberg, Hannes Reinecke, John Meneghini,
Randy Jennings, Dhaval Giani
On Wed, Sep 16, 2026 at 10:09 PM Mohamed Khalfella
<mkhalfella@purestorage.com> wrote:
>
> An unacknowledged write that is retried on another path can still be
> alive in the fabric. If it reaches the target after a later write to the
> same LBA has landed, it overwrites it, and a read returns stale data.
> Nothing in the tree exercises that window.
>
> Add a test that builds it deliberately. A ublk loop device backs an
> nvmet namespace exported through two ports, and the host connects to
> both, so nvme-multipath has a second path to fail over to. io_timeout on
> the subsystem drops to 2 seconds, then miniublk's inject command holds
> one write in the backstore for 4 seconds. The host times that write out,
> retries it on the other path, and the held write completes at the target
> afterwards.
>
> nvme-ghost-write-detector then writes distinct patterns to a single LBA
> and reads the block back, so a resurfaced write shows up as the wrong
> pattern.
>
> The test requires nvme_core.multipath=Y and a fabrics transport. As of
> today it passes on loop, which defines no timeout callback and so never
> times the write out and never retries it, and fails on tcp, rdma and fc.
Can you share a tcp or fc variant of this test that demonstrates failure?
>
> Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Reviewed-by: Jesse Taube <jtaubepe@redhat.com>
Tested-by: Jesse Taube <jtaubepe@redhat.com>
> ---
> tests/nvme/070 | 98 ++++++++++++++++++++++++++++++++++++++++++++++
> tests/nvme/070.out | 35 +++++++++++++++++
> 2 files changed, 133 insertions(+)
> create mode 100755 tests/nvme/070
> create mode 100644 tests/nvme/070.out
btw this needs to be bumped to 071 as 070 got added recently.
Thanks,
Jesse Taube
>
> diff --git a/tests/nvme/070 b/tests/nvme/070
> new file mode 100755
> index 0000000..e9a4690
> --- /dev/null
> +++ b/tests/nvme/070
> @@ -0,0 +1,98 @@
> +#!/bin/bash
> +# SPDX-License-Identifier: GPL-3.0+
> +# Copyright (C) 2026 Mohamed Khalfella
> +
> +. tests/nvme/rc
> +. common/ublk
> +
> +DESCRIPTION="Test injecting delay on nvme-target backstore and expect no corruption"
> +
> +requires() {
> + _nvme_requires
> + _have_loop
> + _have_ublk
> + _have_module_param_value nvme_core multipath Y
> + _require_nvme_trtype_is_fabrics
> + _have_src_program nvme-ghost-write-detector
> +}
> +
> +set_conditions() {
> + _set_nvme_trtype "$@"
> +}
> +
> +count_paths_to_subsystem() {
> + local subsysnqn="$1"
> + local dev count
> +
> + count=0
> + for dev in /sys/class/nvme/nvme*; do
> + [[ -e "${dev}/subsysnqn" ]] || continue
> + [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
> + count=$(( count + 1 ))
> + done
> + echo "${count}"
> +}
> +
> +set_io_timeout_of_subsystem() {
> + local subsysnqn="$1"
> + local timeout="$2"
> + local dev
> +
> + for dev in /sys/class/nvme/nvme*; do
> + [[ -e "${dev}/subsysnqn" ]] || continue
> + [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
> + if ! echo "${timeout}" > "${dev}/io_timeout" 2> /dev/null; then
> + echo "FAIL: can not set io_timeout on ${dev##*/}"
> + return 1
> + fi
> + done
> +}
> +
> +test() {
> + echo "Running ${TEST_NAME}"
> +
> + local ns port nr_paths
> + local -a ports
> +
> + if ! _init_ublk; then
> + return 1
> + fi
> +
> + truncate -s "${NVME_IMG_SIZE}" "${TMPDIR}/ublk-img"
> + if ! ${UBLK_PROG} add -t loop -f "${TMPDIR}/ublk-img" -n 0 > "$FULL" 2>&1; then
> + echo "fail to add ublk device"
> + _exit_ublk
> + return 1
> + fi
> + udevadm settle
> +
> + _setup_nvmet
> + _nvmet_target_setup --ports 2 --blkdev none
> + _create_nvmet_ns --blkdev /dev/ublkb0 \
> + --uuid "${def_subsys_uuid}" > /dev/null
> +
> + _get_nvmet_ports "${def_subsysnqn}" ports
> + echo "Target ports: ${#ports[@]}"
> + for port in "${ports[@]}"; do
> + _nvme_connect_subsys --port "${port}"
> + done
> +
> + nr_paths=$(count_paths_to_subsystem "${def_subsysnqn}")
> + if (( nr_paths != 2 )); then
> + echo "FAIL: expected 2 paths, found ${nr_paths}"
> + fi
> +
> + set_io_timeout_of_subsystem "${def_subsysnqn}" 2000
> +
> + if ! ${UBLK_PROG} inject -n 0 -o write -d 4 -c 1 >> "$FULL" 2>&1; then
> + echo "FAIL: can not inject write delay"
> + fi
> +
> + ns=$(_find_nvme_ns "${def_subsys_uuid}")
> + "$SRCDIR/nvme-ghost-write-detector" "/dev/${ns}"
> +
> + _nvme_disconnect_subsys
> + _nvmet_target_cleanup
> + _exit_ublk
> + echo "Test complete"
> +}
> diff --git a/tests/nvme/070.out b/tests/nvme/070.out
> new file mode 100644
> index 0000000..b43fec7
> --- /dev/null
> +++ b/tests/nvme/070.out
> @@ -0,0 +1,35 @@
> +Running nvme/070
> +Target ports: 2
> +starting nvme-ghost-write-detector test program
> +iteration number 0, writing data
> +validating written data
> +successfully validated
> +iteration number 1, writing data
> +validating written data
> +successfully validated
> +iteration number 2, writing data
> +validating written data
> +successfully validated
> +iteration number 3, writing data
> +validating written data
> +successfully validated
> +iteration number 4, writing data
> +validating written data
> +successfully validated
> +iteration number 5, writing data
> +validating written data
> +successfully validated
> +iteration number 6, writing data
> +validating written data
> +successfully validated
> +iteration number 7, writing data
> +validating written data
> +successfully validated
> +iteration number 8, writing data
> +validating written data
> +successfully validated
> +iteration number 9, writing data
> +validating written data
> +successfully validated
> +finished nvme-ghost-write-detector test program
> +Test complete
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
` (4 preceding siblings ...)
2026-09-17 2:06 ` [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace Mohamed Khalfella
@ 2026-09-23 8:18 ` Shin'ichiro Kawasaki
2026-09-23 17:02 ` Mohamed Khalfella
5 siblings, 1 reply; 13+ messages in thread
From: Shin'ichiro Kawasaki @ 2026-09-23 8:18 UTC (permalink / raw)
To: Mohamed Khalfella
Cc: linux-block, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani
On Sep 16, 2026 / 20:06, Mohamed Khalfella wrote:
Mohamad, thanks for the series. Before I look in the details of the patches,
I would like make a comment below.
> Reproducing the race needs an IO that can be held at the target for
> longer than the host's io_timeout without being failed. The first three
> patches give miniublk that ability, since there was previously no way to
> talk to a running ublk server at all:
>
> 1/5 adds a loopback UDP control channel to the daemon, served by a
> detached thread, one reply datagram per request
> 2/5 adds an INJECT_DELAY command which delays a given number of reads
> or writes as an io_uring timeout on the queue's own ring
> 3/5 exposes it as "miniublk inject", which returns only once the
> daemon has armed the delay, so a test can start IO immediately
> without racing it
It is interesting to implement the delay feature to miniublk, and use it as the
testing tool. But I'm not sure if this is the best approach because of two
points. First, scsi_debug has jdelay and ndelay options. I wonder if this
existing feature can fulfill the requirements of the new test case. Second, when
rublk [*] is available in the PATH of the test system, blktests detect it and
use it instead of miniublk. I guess the added delay feature is not available in
rublk, and the test case won't work well. What do you think of these points?
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace
2026-09-22 17:07 ` Jesse Taube
@ 2026-09-23 16:56 ` Mohamed Khalfella
0 siblings, 0 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-23 16:56 UTC (permalink / raw)
To: Jesse Taube
Cc: linux-block, shinichiro.kawasaki, Keith Busch, Jens Axboe,
Christoph Hellwig, Sagi Grimberg, Hannes Reinecke, John Meneghini,
Randy Jennings, Dhaval Giani
On Tue 2026-09-22 13:07:04 -0400, Jesse Taube wrote:
> On Wed, Sep 16, 2026 at 10:09 PM Mohamed Khalfella
> <mkhalfella@purestorage.com> wrote:
> >
> > An unacknowledged write that is retried on another path can still be
> > alive in the fabric. If it reaches the target after a later write to the
> > same LBA has landed, it overwrites it, and a read returns stale data.
> > Nothing in the tree exercises that window.
> >
> > Add a test that builds it deliberately. A ublk loop device backs an
> > nvmet namespace exported through two ports, and the host connects to
> > both, so nvme-multipath has a second path to fail over to. io_timeout on
> > the subsystem drops to 2 seconds, then miniublk's inject command holds
> > one write in the backstore for 4 seconds. The host times that write out,
> > retries it on the other path, and the held write completes at the target
> > afterwards.
> >
> > nvme-ghost-write-detector then writes distinct patterns to a single LBA
> > and reads the block back, so a resurfaced write shows up as the wrong
> > pattern.
> >
> > The test requires nvme_core.multipath=Y and a fabrics transport. As of
> > today it passes on loop, which defines no timeout callback and so never
> > times the write out and never retries it, and fails on tcp, rdma and fc.
>
> Can you share a tcp or fc variant of this test that demonstrates failure?
Okay, will do that in next revision.
>
> >
> > Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
>
> Reviewed-by: Jesse Taube <jtaubepe@redhat.com>
> Tested-by: Jesse Taube <jtaubepe@redhat.com>
Thanks for testing and reviewing the change.
>
> > ---
> > tests/nvme/070 | 98 ++++++++++++++++++++++++++++++++++++++++++++++
> > tests/nvme/070.out | 35 +++++++++++++++++
> > 2 files changed, 133 insertions(+)
> > create mode 100755 tests/nvme/070
> > create mode 100644 tests/nvme/070.out
>
> btw this needs to be bumped to 071 as 070 got added recently.
Noted.
>
> Thanks,
> Jesse Taube
>
> >
> > diff --git a/tests/nvme/070 b/tests/nvme/070
> > new file mode 100755
> > index 0000000..e9a4690
> > --- /dev/null
> > +++ b/tests/nvme/070
> > @@ -0,0 +1,98 @@
> > +#!/bin/bash
> > +# SPDX-License-Identifier: GPL-3.0+
> > +# Copyright (C) 2026 Mohamed Khalfella
> > +
> > +. tests/nvme/rc
> > +. common/ublk
> > +
> > +DESCRIPTION="Test injecting delay on nvme-target backstore and expect no corruption"
> > +
> > +requires() {
> > + _nvme_requires
> > + _have_loop
> > + _have_ublk
> > + _have_module_param_value nvme_core multipath Y
> > + _require_nvme_trtype_is_fabrics
> > + _have_src_program nvme-ghost-write-detector
> > +}
> > +
> > +set_conditions() {
> > + _set_nvme_trtype "$@"
> > +}
> > +
> > +count_paths_to_subsystem() {
> > + local subsysnqn="$1"
> > + local dev count
> > +
> > + count=0
> > + for dev in /sys/class/nvme/nvme*; do
> > + [[ -e "${dev}/subsysnqn" ]] || continue
> > + [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
> > + count=$(( count + 1 ))
> > + done
> > + echo "${count}"
> > +}
> > +
> > +set_io_timeout_of_subsystem() {
> > + local subsysnqn="$1"
> > + local timeout="$2"
> > + local dev
> > +
> > + for dev in /sys/class/nvme/nvme*; do
> > + [[ -e "${dev}/subsysnqn" ]] || continue
> > + [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
> > + if ! echo "${timeout}" > "${dev}/io_timeout" 2> /dev/null; then
> > + echo "FAIL: can not set io_timeout on ${dev##*/}"
> > + return 1
> > + fi
> > + done
> > +}
> > +
> > +test() {
> > + echo "Running ${TEST_NAME}"
> > +
> > + local ns port nr_paths
> > + local -a ports
> > +
> > + if ! _init_ublk; then
> > + return 1
> > + fi
> > +
> > + truncate -s "${NVME_IMG_SIZE}" "${TMPDIR}/ublk-img"
> > + if ! ${UBLK_PROG} add -t loop -f "${TMPDIR}/ublk-img" -n 0 > "$FULL" 2>&1; then
> > + echo "fail to add ublk device"
> > + _exit_ublk
> > + return 1
> > + fi
> > + udevadm settle
> > +
> > + _setup_nvmet
> > + _nvmet_target_setup --ports 2 --blkdev none
> > + _create_nvmet_ns --blkdev /dev/ublkb0 \
> > + --uuid "${def_subsys_uuid}" > /dev/null
> > +
> > + _get_nvmet_ports "${def_subsysnqn}" ports
> > + echo "Target ports: ${#ports[@]}"
> > + for port in "${ports[@]}"; do
> > + _nvme_connect_subsys --port "${port}"
> > + done
> > +
> > + nr_paths=$(count_paths_to_subsystem "${def_subsysnqn}")
> > + if (( nr_paths != 2 )); then
> > + echo "FAIL: expected 2 paths, found ${nr_paths}"
> > + fi
> > +
> > + set_io_timeout_of_subsystem "${def_subsysnqn}" 2000
> > +
> > + if ! ${UBLK_PROG} inject -n 0 -o write -d 4 -c 1 >> "$FULL" 2>&1; then
> > + echo "FAIL: can not inject write delay"
> > + fi
> > +
> > + ns=$(_find_nvme_ns "${def_subsys_uuid}")
> > + "$SRCDIR/nvme-ghost-write-detector" "/dev/${ns}"
> > +
> > + _nvme_disconnect_subsys
> > + _nvmet_target_cleanup
> > + _exit_ublk
> > + echo "Test complete"
> > +}
> > diff --git a/tests/nvme/070.out b/tests/nvme/070.out
> > new file mode 100644
> > index 0000000..b43fec7
> > --- /dev/null
> > +++ b/tests/nvme/070.out
> > @@ -0,0 +1,35 @@
> > +Running nvme/070
> > +Target ports: 2
> > +starting nvme-ghost-write-detector test program
> > +iteration number 0, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 1, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 2, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 3, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 4, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 5, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 6, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 7, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 8, writing data
> > +validating written data
> > +successfully validated
> > +iteration number 9, writing data
> > +validating written data
> > +successfully validated
> > +finished nvme-ghost-write-detector test program
> > +Test complete
> > --
> > 2.55.0
> >
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector
2026-09-21 18:38 ` Jesse Taube
@ 2026-09-23 17:00 ` Mohamed Khalfella
0 siblings, 0 replies; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-23 17:00 UTC (permalink / raw)
To: Jesse Taube
Cc: linux-block, shinichiro.kawasaki, Keith Busch, Jens Axboe,
Christoph Hellwig, Sagi Grimberg, Hannes Reinecke, John Meneghini,
Randy Jennings, Dhaval Giani
On Mon 2026-09-21 14:38:12 -0400, Jesse Taube wrote:
> On Wed, Sep 16, 2026 at 10:09 PM Mohamed Khalfella
> <mkhalfella@purestorage.com> wrote:
> >
> > Consider a write X that is issued to an LBA and times out without ever
> > being acknowledged. The initiator retries it on another path, where it
> > lands. A write Y to the same LBA follows and lands as well. The original
> > X is still in flight somewhere in the fabric, and when it finally
> > reaches the device it overwrites Y. A read now returns X, a write the
> > initiator gave up on long ago.
> >
> > Add a program to catch this. It opens the device O_DIRECT and, on each
> > iteration, issues a series of writes of distinct byte patterns to the
> > same LBA, the 4k block at offset 0, waits, then reads that block back
> > and checks every byte holds the pattern written last. Because the
> > patterns differ, the value that comes back identifies which write
> > reappeared.
> >
> > Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
> > ---
> > src/.gitignore | 1 +
> > src/Makefile | 1 +
> > src/nvme-ghost-write-detector.c | 86 +++++++++++++++++++++++++++++++++
> > 3 files changed, 88 insertions(+)
> > create mode 100644 src/nvme-ghost-write-detector.c
> >
> > diff --git a/src/.gitignore b/src/.gitignore
> > index e9869e1..9673be8 100644
> > --- a/src/.gitignore
> > +++ b/src/.gitignore
> > @@ -15,6 +15,7 @@
> > /zbdioctl
> > /miniublk
> > /nvme-passthrough-meta
> > +/nvme-ghost-write-detector
> > /ioctl-lbmd-query
> > /nvme-passthru-admin-uring
> > /nvme-delay-ioctl
> > diff --git a/src/Makefile b/src/Makefile
> > index dd64694..92b4d0c 100644
> > --- a/src/Makefile
> > +++ b/src/Makefile
> > @@ -24,6 +24,7 @@ C_TARGETS := \
> > mount_clear_sock \
> > nvme-delay-ioctl \
> > nvme-passthrough-meta \
> > + nvme-ghost-write-detector \
> > ioctl-lbmd-query \
> > nbdsetsize \
> > openclose \
> > diff --git a/src/nvme-ghost-write-detector.c b/src/nvme-ghost-write-detector.c
> > new file mode 100644
> > index 0000000..bd42dde
> > --- /dev/null
> > +++ b/src/nvme-ghost-write-detector.c
> > @@ -0,0 +1,86 @@
> > +// SPDX-License-Identifier: GPL-3.0+
> > +// Copyright (C) 2026 Mohamed Khalfella
> > +
> > +#define _GNU_SOURCE
> > +#include <stdio.h>
> > +#include <stdlib.h>
> > +#include <fcntl.h>
> > +#include <unistd.h>
> > +#include <string.h>
> > +#include <malloc.h>
> > +#include <errno.h>
> > +#include <libgen.h>
>
> We don't need to include malloc.h
Okay, will drop it.
> > +
> > +#define BUF_SIZE 4096
> > +#define ITERATIONS 10
> > +#define DELAY 3 /* seconds delay between iterations */
> > +#define WRITE_COUNT 10
> > +
> > +#define WRITE_OFFSET 0
> > +#define READ_OFFSET WRITE_OFFSET
> > +
> > +int main(int argc, char **argv)
> > +{
> > + int fd, i, w, off, ret;
> > + char *buff;
> > +
>
> This is a bit pedantic but:
>
> if (argc < 1)
> return 1;
See below "if (argc < 2)". We need at least one argument the nvme device
to run IOs on. arg[0] is the program name, and arg[1] is the device name.
>
> > + fprintf(stdout, "starting %s test program\n", basename(argv[0]));
> > +
> > + if (argc < 2) {
> > + fprintf(stderr, "usage: %s /dev/nvmeXnY", argv[0]);
> > + return 1;
> > + }
> > +
> > + fd = open(argv[1], O_RDWR | O_DIRECT);
> > + if (fd < 0) {
> > + fprintf(stderr, "failed to open device, errno = %d\n", errno);
> > + return 1;
> > + }
> > +
> > + ret = posix_memalign((void **)&buff, BUF_SIZE, BUF_SIZE);
> > + if (ret) {
> > + fprintf(stderr, "failed to allocate buffer, ret = %d\n", ret);
> > + goto out;
> > + }
> > +
> > + for (i = 0; i < ITERATIONS; i++) {
> > + fprintf(stdout, "iteration number %d, writing data\n", i);
> > +
> > + for (w = 0; w < WRITE_COUNT; w++) {
> > + memset(buff, w, BUF_SIZE);
> > + ret = pwrite(fd, buff, BUF_SIZE, WRITE_OFFSET);
> > + if (ret != BUF_SIZE) {
> > + fprintf(stderr, "failed to write buff, "
> > + "ret = %d, errno = %d\n",
> > + ret, errno);
> > + goto out;
> > + }
> > + }
> > +
> > + sleep(5);
>
> There should be a macro for this.
Okay, will do that in next revision.
>
> Thanks,
> Jesse Taube
>
> > + fprintf(stdout, "validating written data\n");
> > +
> > + ret = pread(fd, buff, BUF_SIZE, READ_OFFSET);
> > + if (ret != BUF_SIZE) {
> > + fprintf(stderr, "failed to read buff, "
> > + "ret = %d, errno = %d\n",
> > + ret, errno);
> > + goto out;
> > + }
> > +
> > + for (off = 0; off < BUF_SIZE; off++) {
> > + if (buff[off] != WRITE_COUNT - 1) {
> > + fprintf(stdout, "validation failed\n");
> > + goto out;
> > + }
> > + }
> > +
> > + fprintf(stdout, "successfully validated\n");
> > + sleep(DELAY);
> > + }
> > +
> > +out:
> > + fprintf(stdout, "finished %s test program\n", basename(argv[0]));
> > + close(fd);
> > + return ret;
> > +}
> > --
> > 2.55.0
> >
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics
2026-09-23 8:18 ` [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Shin'ichiro Kawasaki
@ 2026-09-23 17:02 ` Mohamed Khalfella
2026-10-02 3:03 ` Shin'ichiro Kawasaki
0 siblings, 1 reply; 13+ messages in thread
From: Mohamed Khalfella @ 2026-09-23 17:02 UTC (permalink / raw)
To: Shin'ichiro Kawasaki
Cc: linux-block, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani
On Wed 2026-09-23 17:18:21 +0900, Shin'ichiro Kawasaki wrote:
> On Sep 16, 2026 / 20:06, Mohamed Khalfella wrote:
>
> Mohamad, thanks for the series. Before I look in the details of the patches,
> I would like make a comment below.
>
> > Reproducing the race needs an IO that can be held at the target for
> > longer than the host's io_timeout without being failed. The first three
> > patches give miniublk that ability, since there was previously no way to
> > talk to a running ublk server at all:
> >
> > 1/5 adds a loopback UDP control channel to the daemon, served by a
> > detached thread, one reply datagram per request
> > 2/5 adds an INJECT_DELAY command which delays a given number of reads
> > or writes as an io_uring timeout on the queue's own ring
> > 3/5 exposes it as "miniublk inject", which returns only once the
> > daemon has armed the delay, so a test can start IO immediately
> > without racing it
>
> It is interesting to implement the delay feature to miniublk, and use it as the
> testing tool. But I'm not sure if this is the best approach because of two
> points. First, scsi_debug has jdelay and ndelay options. I wonder if this
> existing feature can fulfill the requirements of the new test case. Second, when
> rublk [*] is available in the PATH of the test system, blktests detect it and
> use it instead of miniublk. I guess the added delay feature is not available in
> rublk, and the test case won't work well. What do you think of these points?
>
The two points make sense to me. Let me look at scsi_debug delay
injection mechanism first and see if it can be used here. I will also
look at rublk. Thanks for the pointers.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics
2026-09-23 17:02 ` Mohamed Khalfella
@ 2026-10-02 3:03 ` Shin'ichiro Kawasaki
0 siblings, 0 replies; 13+ messages in thread
From: Shin'ichiro Kawasaki @ 2026-10-02 3:03 UTC (permalink / raw)
To: Mohamed Khalfella
Cc: linux-block, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani
On Sep 23, 2026 / 10:02, Mohamed Khalfella wrote:
> On Wed 2026-09-23 17:18:21 +0900, Shin'ichiro Kawasaki wrote:
> > On Sep 16, 2026 / 20:06, Mohamed Khalfella wrote:
> >
> > Mohamad, thanks for the series. Before I look in the details of the patches,
> > I would like make a comment below.
> >
> > > Reproducing the race needs an IO that can be held at the target for
> > > longer than the host's io_timeout without being failed. The first three
> > > patches give miniublk that ability, since there was previously no way to
> > > talk to a running ublk server at all:
> > >
> > > 1/5 adds a loopback UDP control channel to the daemon, served by a
> > > detached thread, one reply datagram per request
> > > 2/5 adds an INJECT_DELAY command which delays a given number of reads
> > > or writes as an io_uring timeout on the queue's own ring
> > > 3/5 exposes it as "miniublk inject", which returns only once the
> > > daemon has armed the delay, so a test can start IO immediately
> > > without racing it
> >
> > It is interesting to implement the delay feature to miniublk, and use it as the
> > testing tool. But I'm not sure if this is the best approach because of two
> > points. First, scsi_debug has jdelay and ndelay options. I wonder if this
> > existing feature can fulfill the requirements of the new test case. Second, when
> > rublk [*] is available in the PATH of the test system, blktests detect it and
> > use it instead of miniublk. I guess the added delay feature is not available in
> > rublk, and the test case won't work well. What do you think of these points?
> >
>
> The two points make sense to me. Let me look at scsi_debug delay
> injection mechanism first and see if it can be used here. I will also
> look at rublk. Thanks for the pointers.
I've learned that now delay injection is being added to the block layer [*].
This looks more generic and could be the best way for this use case.
[*] https://lore.kernel.org/linux-block/20260928221634.43239-1-haris.iqbal@linux.dev/
^ permalink raw reply [flat|nested] 13+ messages in thread
end of thread, other threads:[~2026-10-02 3:03 UTC | newest]
Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 2:06 [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 1/5] src/miniublk: add a control channel to the daemon Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 2/5] src/miniublk: add IO delay injection Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 3/5] src/miniublk: add the inject command Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 4/5] src/nvme-ghost-write-detector: add an ABA ghost write detector Mohamed Khalfella
2026-09-21 18:38 ` Jesse Taube
2026-09-23 17:00 ` Mohamed Khalfella
2026-09-17 2:06 ` [PATCH blktests 5/5] nvme/070: test for ABA ghost writes on a multipath fabrics namespace Mohamed Khalfella
2026-09-22 17:07 ` Jesse Taube
2026-09-23 16:56 ` Mohamed Khalfella
2026-09-23 8:18 ` [PATCH blktests 0/5] nvme: detect ABA ghost writes on multipath fabrics Shin'ichiro Kawasaki
2026-09-23 17:02 ` Mohamed Khalfella
2026-10-02 3:03 ` Shin'ichiro Kawasaki
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox