* [PATCH blktests] nvme/071: test CCR and CQT recovery on a multipath fabrics namespace
@ 2026-09-21 20:59 Mohamed Khalfella
2026-09-25 3:08 ` Shin'ichiro Kawasaki
0 siblings, 1 reply; 2+ messages in thread
From: Mohamed Khalfella @ 2026-09-21 20:59 UTC (permalink / raw)
To: linux-block
Cc: shinichiro.kawasaki, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani, Mohamed Khalfella
nvme/070 demonstrated the ABA ghost write window. The kernel closes it
by fencing a controller before retrying timed-out writes on another
path, but nvme/070 does not functionally exercise CCR or CQT
themselves.
Add a functional test that drives both cross-controller reset (CCR)
and command quiesce time (CQT) recovery. The setup is the same as
nvme/070: a ublk loop device backs an nvmet namespace that is exported
through two ports, the host connects to both paths, and io_timeout on
the subsystem drops to 2 seconds. In addition, the target subsystem's
CQT is set to 30 seconds, since it is disabled by default.
The first scenario holds one write in the backstore for 4 seconds. The
host times the write out, starts fencing, and the CCR issued on the
other path completes within its budget. The second scenario holds a
write for 10 seconds, longer than CCR can wait, so the host gives up
on CCR and switches to time-based recovery, waiting out the remaining
CQT window before releasing the retry.
Each scenario runs nvme-ghost-write-detector while recovery is in flight
to check that no stale data surfaces, and confirms the recovery took the
expected route by watching dmesg after a marker written to /dev/kmsg.
The test requires nvme_core.multipath=Y and is limited to tcp, rdma
and fc, the transports that implement controller fencing.
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
---
tests/nvme/071 | 184 +++++++++++++++++++++++++++++++++++++++++++++
tests/nvme/071.out | 67 +++++++++++++++++
2 files changed, 251 insertions(+)
create mode 100755 tests/nvme/071
create mode 100644 tests/nvme/071.out
diff --git a/tests/nvme/071 b/tests/nvme/071
new file mode 100755
index 0000000..d700416
--- /dev/null
+++ b/tests/nvme/071
@@ -0,0 +1,184 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-3.0+
+# Copyright (C) 2026 Mohamed Khalfella
+
+. tests/nvme/rc
+. common/ublk
+
+DESCRIPTION="CCR/CQT functional test"
+
+requires() {
+ _nvme_requires
+ _have_loop
+ _have_ublk
+ _have_module_param_value nvme_core multipath Y
+ _require_nvme_trtype tcp rdma fc
+ _have_src_program nvme-ghost-write-detector
+}
+
+set_conditions() {
+ _set_nvme_trtype "$@"
+}
+
+count_paths_to_subsystem() {
+ local subsysnqn="$1"
+ local dev count
+
+ count=0
+ for dev in /sys/class/nvme/nvme*; do
+ [[ -e "${dev}/subsysnqn" ]] || continue
+ [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
+ count=$(( count + 1 ))
+ done
+ echo "${count}"
+}
+
+set_io_timeout_of_subsystem() {
+ local subsysnqn="$1"
+ local timeout="$2"
+ local dev
+
+ for dev in /sys/class/nvme/nvme*; do
+ [[ -e "${dev}/subsysnqn" ]] || continue
+ [[ "$(cat "${dev}/subsysnqn")" == "${subsysnqn}" ]] || continue
+ if ! echo "${timeout}" > "${dev}/io_timeout" 2> /dev/null; then
+ echo "FAIL: can not set io_timeout on ${dev##*/}"
+ return 1
+ fi
+ done
+}
+
+dmesg_mark() {
+ local marker="blktests ${TEST_NAME} $1"
+
+ echo "${marker}" >> /dev/kmsg
+ echo "${marker}"
+}
+
+dmesg_since_mark() {
+ dmesg | awk -v m="$1" 'index($0, m) { found = 1; next } found'
+}
+
+wait_for_dmesg_after_mark() {
+ local mark="$1"
+ local pattern="$2"
+ local timeout="$3"
+ local i
+
+ for ((i = 0; i < timeout * 2; i++)); do
+ if dmesg_since_mark "${mark}" | grep -q -e "${pattern}"; then
+ return 0
+ fi
+ sleep 0.5
+ done
+ return 1
+}
+
+test_injecting_delay_ccr_recovery() {
+ local mark ns
+
+ mark=$(dmesg_mark "ccr_recovery_started")
+
+ if ! ${UBLK_PROG} inject -n 0 -o write -d 4 -c 1 >> "$FULL" 2>&1; then
+ echo "FAIL: can not inject write delay"
+ fi
+
+ ns=$(_find_nvme_ns "${def_subsys_uuid}")
+ "$SRCDIR/nvme-ghost-write-detector" "/dev/${ns}"
+
+ # The injected delay is within the limits that CCR can handle.
+ # Expect fencing to start and CCR to succeed.
+ if ! wait_for_dmesg_after_mark "${mark}" "starting controller fencing" 30; then
+ echo "FAIL: controller fencing did not start"
+ fi
+
+ if ! wait_for_dmesg_after_mark "${mark}" "CCR succeeded using nvme" 30; then
+ echo "FAIL: cross controller reset did not succeed"
+ fi
+}
+
+test_injecting_delay_cqt_recovery() {
+ local mark ns
+
+ mark=$(dmesg_mark "cqt_recovery_started")
+
+ if ! ${UBLK_PROG} inject -n 0 -o write -d 10 -c 1 >> "$FULL" 2>&1; then
+ echo "FAIL: can not inject write delay"
+ fi
+
+ ns=$(_find_nvme_ns "${def_subsys_uuid}")
+ "$SRCDIR/nvme-ghost-write-detector" "/dev/${ns}"
+
+ # The injected delay is outside the limit that CCR can handle.
+ # Expect CCR to fail and CQT to save the day.
+ if ! wait_for_dmesg_after_mark "${mark}" "starting controller fencing" 30; then
+ echo "FAIL: controller fencing did not start"
+ fi
+
+ if ! wait_for_dmesg_after_mark "${mark}" "attempting CCR" 30; then
+ echo "FAIL: cross controller reset was not attempted"
+ fi
+
+ if ! wait_for_dmesg_after_mark "${mark}" "CCR failed, switch to time-based recovery" 30; then
+ echo "FAIL: cross controller did not fail as expected"
+ fi
+
+ if ! wait_for_dmesg_after_mark "${mark}" "Time-based recovery finished" 30; then
+ echo "FAIL: expected to see time-based recovery ends"
+ fi
+}
+
+test() {
+ echo "Running ${TEST_NAME}"
+
+ local port nr_paths
+ local attr_cqt
+ local -a ports
+
+ if ! _init_ublk; then
+ return 1
+ fi
+
+ truncate -s "${NVME_IMG_SIZE}" "${TMPDIR}/ublk-img"
+ if ! ${UBLK_PROG} add -t loop -f "${TMPDIR}/ublk-img" -n 0 > "$FULL" 2>&1; then
+ echo "fail to add ublk device"
+ _exit_ublk
+ return 1
+ fi
+ udevadm settle
+
+ _setup_nvmet
+ _nvmet_target_setup --ports 2 --blkdev none
+ _create_nvmet_ns --blkdev /dev/ublkb0 \
+ --uuid "${def_subsys_uuid}" > /dev/null
+
+ # CQT is disabled by default, enable it for 30s
+ attr_cqt="${NVMET_CFS}/subsystems/${def_subsysnqn}/attr_cqt"
+ if [[ -f "${attr_cqt}" ]]; then
+ echo 30000 > "${attr_cqt}"
+ else
+ echo "FAIL: failed to find attr_cqt: ${attr_cqt}"
+ fi
+
+ _get_nvmet_ports "${def_subsysnqn}" ports
+ echo "Target ports: ${#ports[@]}"
+ for port in "${ports[@]}"; do
+ _nvme_connect_subsys --port "${port}"
+ done
+
+ nr_paths=$(count_paths_to_subsystem "${def_subsysnqn}")
+ if (( nr_paths != 2 )); then
+ echo "FAIL: expected 2 paths, found ${nr_paths}"
+ fi
+
+ set_io_timeout_of_subsystem "${def_subsysnqn}" 2000
+
+ test_injecting_delay_ccr_recovery
+
+ test_injecting_delay_cqt_recovery
+
+ _nvme_disconnect_subsys
+ _nvmet_target_cleanup
+ _exit_ublk
+ echo "Test complete"
+}
diff --git a/tests/nvme/071.out b/tests/nvme/071.out
new file mode 100644
index 0000000..f55eefb
--- /dev/null
+++ b/tests/nvme/071.out
@@ -0,0 +1,67 @@
+Running nvme/071
+Target ports: 2
+starting nvme-ghost-write-detector test program
+iteration number 0, writing data
+validating written data
+successfully validated
+iteration number 1, writing data
+validating written data
+successfully validated
+iteration number 2, writing data
+validating written data
+successfully validated
+iteration number 3, writing data
+validating written data
+successfully validated
+iteration number 4, writing data
+validating written data
+successfully validated
+iteration number 5, writing data
+validating written data
+successfully validated
+iteration number 6, writing data
+validating written data
+successfully validated
+iteration number 7, writing data
+validating written data
+successfully validated
+iteration number 8, writing data
+validating written data
+successfully validated
+iteration number 9, writing data
+validating written data
+successfully validated
+finished nvme-ghost-write-detector test program
+starting nvme-ghost-write-detector test program
+iteration number 0, writing data
+validating written data
+successfully validated
+iteration number 1, writing data
+validating written data
+successfully validated
+iteration number 2, writing data
+validating written data
+successfully validated
+iteration number 3, writing data
+validating written data
+successfully validated
+iteration number 4, writing data
+validating written data
+successfully validated
+iteration number 5, writing data
+validating written data
+successfully validated
+iteration number 6, writing data
+validating written data
+successfully validated
+iteration number 7, writing data
+validating written data
+successfully validated
+iteration number 8, writing data
+validating written data
+successfully validated
+iteration number 9, writing data
+validating written data
+successfully validated
+finished nvme-ghost-write-detector test program
+Test complete
--
2.55.0
^ permalink raw reply related [flat|nested] 2+ messages in thread* Re: [PATCH blktests] nvme/071: test CCR and CQT recovery on a multipath fabrics namespace
2026-09-21 20:59 [PATCH blktests] nvme/071: test CCR and CQT recovery on a multipath fabrics namespace Mohamed Khalfella
@ 2026-09-25 3:08 ` Shin'ichiro Kawasaki
0 siblings, 0 replies; 2+ messages in thread
From: Shin'ichiro Kawasaki @ 2026-09-25 3:08 UTC (permalink / raw)
To: Mohamed Khalfella
Cc: linux-block, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, Hannes Reinecke, John Meneghini, Jesse Taube,
Randy Jennings, Dhaval Giani
On Sep 21, 2026 / 14:59, Mohamed Khalfella wrote:
> nvme/070 demonstrated the ABA ghost write window. The kernel closes it
> by fencing a controller before retrying timed-out writes on another
> path, but nvme/070 does not functionally exercise CCR or CQT
> themselves.
>
> Add a functional test that drives both cross-controller reset (CCR)
> and command quiesce time (CQT) recovery. The setup is the same as
> nvme/070: a ublk loop device backs an nvmet namespace that is exported
> through two ports, the host connects to both paths, and io_timeout on
> the subsystem drops to 2 seconds. In addition, the target subsystem's
> CQT is set to 30 seconds, since it is disabled by default.
>
> The first scenario holds one write in the backstore for 4 seconds. The
> host times the write out, starts fencing, and the CCR issued on the
> other path completes within its budget. The second scenario holds a
> write for 10 seconds, longer than CCR can wait, so the host gives up
> on CCR and switches to time-based recovery, waiting out the remaining
> CQT window before releasing the retry.
>
> Each scenario runs nvme-ghost-write-detector while recovery is in flight
> to check that no stale data surfaces, and confirms the recovery took the
> expected route by watching dmesg after a marker written to /dev/kmsg.
>
> The test requires nvme_core.multipath=Y and is limited to tcp, rdma
> and fc, the transports that implement controller fencing.
>
> Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Thanks. Overall, this patch looks valuable, and I did not see problems.
As I commented on the dependent series [*], this patch also depends on the
miniublk change. This approach looks useful for this kind of precise test
condition control, but the two points I noted will need consideration.
dmesg_mark() and dmesg_since_mark() can be useful for other future test
cases, and may worth adding to common/rc. Also, it can be refactored with
_dmesg_since_test_start(). Said that, I think current patch is fine and
we can do that refactoring work later.
[*] https://lore.kernel.org/linux-block/arOFG9-sM_GrrbQf@shinmob/
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-25 3:08 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 20:59 [PATCH blktests] nvme/071: test CCR and CQT recovery on a multipath fabrics namespace Mohamed Khalfella
2026-09-25 3:08 ` Shin'ichiro Kawasaki
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).