* [PATCH RFC 0/2] Support for Multiple Atomicity Mode
@ 2026-09-29 11:17 Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 1/2] block: add BLK_FEAT_ATOMIC_WRITE_MULTI Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 2/2] nvme: enable multiple atomicity mode Daniel Gomez
0 siblings, 2 replies; 3+ messages in thread
From: Daniel Gomez @ 2026-09-29 11:17 UTC (permalink / raw)
To: Jens Axboe, Keith Busch, Christoph Hellwig, Sagi Grimberg
Cc: linux-block, linux-kernel, linux-nvme, Andres Freund,
Pankaj Raghav, Daniel Gomez, GOST, Daniel Gomez
NVMe has 2 atomic modes: Single Atomicity Mode (SAM) and Multiple
Atomicity Mode (MAM). In SAM mode, as per spec, commands that cross
the atomic boundaries may or may not be guaranteed to be atomic. In MAM
mode, commands that cross the atomic boundaries will be guaranteed to
be atomic at the LBA subrange atomic boundaries, resulting in multiple
atomic operations.
This series adds a new block layer feature BLK_FEAT_ATOMIC_WRITE_MULTI
that gets enabled when an NVMe controller exposes MAM guarantees. In
this case, and when the block IO requests are SAM compliant, allow the
block layer to merge atomic command requests that are contiguous in the
LBA ranges, leveraging block plug-merge infra, and effectively
increasing the size of the command up to the maximum hardware
capabilities. The merged command is not atomic as a whole, so it becomes
a carrier of multiple atomic commands. This is possible because the
RWF_ATOMIC that the user has requested is honored at the LBA subrange
later at the controller side when the carrier command is split at the
atomic boundaries.
MAM QEMU NVMe series [1] covers emulation functionality for testing.
While I think this small feature is worth having (hence this RFC)
for users sending atomic write commands that are contiguous, the
question I'd also like to discuss really is how to also expose clear
MAM semantics to userspace for users who can send larger commands,
so they avoid syscall overhead [2] and splitting, and still get the
atomic guarantees in the LBA subranges. RWF_ATOMIC currently exposes
SAM semantics, but I think MAM may be confusing for users, as it's not
simply a larger atomic write but an atomic carrier. Thoughts? What can
work best to leverage MAM from userspace without conflicting with SAM?
Command line logs covering functionality with nvme-cli, fio and
blktrace:
nvme id-ns -H /dev/nvme0n1
NVME Identify Namespace 1:
{..}
nsfeat : 0x76
[7:7] : 0 NPRG, NPRA and NORS are Not Supported
[6:6] : 0x1 Multiple Atomicity Mode applies to write operations
[5:4] : 0x3 NPWG, NPWA, NPDG, NPDGL, NPDA, and NOWS are Supported
{..}
Sequential atomic write workflows can benefit from the block layer
plug-merge as the following fio atomic write workload:
fio --name=atomic-merge --filename=/dev/nvme0n1 \
--rw=write --bs=16k --size=64k \
--ioengine=io_uring \
--iodepth=4 --iodepth_batch_submit=4 \
--atomic=1 --direct=1
The output below from btrace confirms the merging results into one
single command:
btrace /dev/nvme0n1
259,2 3 1 0.000000000 2725 Q WS 0 + 32 [fio]
259,2 3 2 0.000003450 2725 G WS 0 + 32 [fio]
259,2 3 3 0.000004250 2725 P N [fio]
259,2 3 4 0.000005440 2725 Q WS 32 + 32 [fio]
259,2 3 5 0.000006400 2725 M WS 32 + 32 [fio]
259,2 3 6 0.000007360 2725 Q WS 64 + 32 [fio]
259,2 3 7 0.000007490 2725 M WS 64 + 32 [fio]
259,2 3 8 0.000008340 2725 Q WS 96 + 32 [fio]
259,2 3 9 0.000008460 2725 M WS 96 + 32 [fio]
259,2 3 10 0.000009270 2725 U N [fio] 1
259,2 3 11 0.000013460 2725 D WS 0 + 128 [fio]
259,2 3 12 0.000239642 0 C WS 0 + 128 [0]
Link: https://lore.kernel.org/all/20260825-nvme-mam-v1-0-afc38ac713ef@samsung.com/ [1]
Link: https://kernel-recipes.org/en/2026/postgres-on-vs-with-linux/ [2]
Signed-off-by: Daniel Gomez <da.gomez@samsung.com>
---
Daniel Gomez (2):
block: add BLK_FEAT_ATOMIC_WRITE_MULTI
nvme: enable multiple atomicity mode
block/blk-merge.c | 9 ++++++++-
block/blk-settings.c | 5 +++++
block/blk.h | 6 +++++-
drivers/nvme/host/core.c | 23 +++++++++++++++++++++++
include/linux/blkdev.h | 3 +++
include/linux/nvme.h | 1 +
6 files changed, 45 insertions(+), 2 deletions(-)
---
base-commit: e680312dd3990197297b485e27b53c33d371c279
change-id: 20260929-nvme-mam-073f2ed13bdc
Best regards,
--
Daniel Gomez <da.gomez@samsung.com>
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH RFC 1/2] block: add BLK_FEAT_ATOMIC_WRITE_MULTI
2026-09-29 11:17 [PATCH RFC 0/2] Support for Multiple Atomicity Mode Daniel Gomez
@ 2026-09-29 11:17 ` Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 2/2] nvme: enable multiple atomicity mode Daniel Gomez
1 sibling, 0 replies; 3+ messages in thread
From: Daniel Gomez @ 2026-09-29 11:17 UTC (permalink / raw)
To: Jens Axboe, Keith Busch, Christoph Hellwig, Sagi Grimberg
Cc: linux-block, linux-kernel, linux-nvme, Andres Freund,
Pankaj Raghav, Daniel Gomez, GOST, Daniel Gomez
From: Daniel Gomez <da.gomez@samsung.com>
Add a feature flag that allows the block layer to merge atomic write
commands into larger ones that are not atomic as a whole but are later
divided by the device at the atomic write boundaries, with each subrange
treated as atomic. NVMe calls this Multiple Atomicity Mode (MAM).
When the flag is set, stop limiting merged atomic writes at
the atomic write boundary and cap them at max(max_sectors,
atomic_write_max_sectors). Each individual atomic write still fits one
boundary window and every window is written atomically by the device, so
each merged write stays untorn.
The feature flag can only be enabled when an atomic write boundary
is set.
No consumer yet, so no behavior changes.
Assisted-by: LLM
Signed-off-by: Daniel Gomez <da.gomez@samsung.com>
---
block/blk-merge.c | 9 ++++++++-
block/blk-settings.c | 5 +++++
block/blk.h | 6 +++++-
include/linux/blkdev.h | 3 +++
4 files changed, 21 insertions(+), 2 deletions(-)
diff --git a/block/blk-merge.c b/block/blk-merge.c
index 258a726071d12..e3c6aae2b3125 100644
--- a/block/blk-merge.c
+++ b/block/blk-merge.c
@@ -523,7 +523,14 @@ static inline unsigned int blk_rq_get_max_sectors(struct request *rq,
struct request_queue *q = rq->q;
struct queue_limits *lim = &q->limits;
unsigned int max_sectors, boundary_sectors;
- bool is_atomic = rq->cmd_flags & REQ_ATOMIC;
+ /*
+ * The merged command is itself one atomic write and must not cross the
+ * atomic write boundary. But in the BLK_FEAT_ATOMIC_WRITE_MULTI case,
+ * the device writes each boundary window atomically, so its merged
+ * commands are not atomic as a whole and may cross the boundary.
+ */
+ bool is_atomic = (rq->cmd_flags & REQ_ATOMIC) &&
+ !(lim->features & BLK_FEAT_ATOMIC_WRITE_MULTI);
if (blk_rq_is_passthrough(rq))
return q->limits.max_hw_sectors;
diff --git a/block/blk-settings.c b/block/blk-settings.c
index 1f5ee2453269f..f0488dedb0b80 100644
--- a/block/blk-settings.c
+++ b/block/blk-settings.c
@@ -309,6 +309,10 @@ static void blk_validate_atomic_write_limits(struct queue_limits *lim)
boundary_sectors = lim->atomic_write_hw_boundary >> SECTOR_SHIFT;
+ if (WARN_ON_ONCE((lim->features & BLK_FEAT_ATOMIC_WRITE_MULTI) &&
+ !boundary_sectors))
+ lim->features &= ~BLK_FEAT_ATOMIC_WRITE_MULTI;
+
if (boundary_sectors) {
if (WARN_ON_ONCE(lim->atomic_write_hw_max >
lim->atomic_write_hw_boundary))
@@ -333,6 +337,7 @@ static void blk_validate_atomic_write_limits(struct queue_limits *lim)
return;
unsupported:
+ lim->features &= ~BLK_FEAT_ATOMIC_WRITE_MULTI;
lim->atomic_write_max_sectors = 0;
lim->atomic_write_boundary_sectors = 0;
lim->atomic_write_unit_min = 0;
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c532..8d902f41d93c0 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -241,8 +241,12 @@ static inline unsigned int blk_queue_get_max_sectors(struct request *rq)
if (unlikely(op == REQ_OP_WRITE_ZEROES))
return q->limits.max_write_zeroes_sectors;
- if (rq->cmd_flags & REQ_ATOMIC)
+ if (rq->cmd_flags & REQ_ATOMIC) {
+ if (q->limits.features & BLK_FEAT_ATOMIC_WRITE_MULTI)
+ return max(q->limits.max_sectors,
+ q->limits.atomic_write_max_sectors);
return q->limits.atomic_write_max_sectors;
+ }
return q->limits.max_sectors;
}
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6c..92efd4ff67f26 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -360,6 +360,9 @@ typedef unsigned int __bitwise blk_features_t;
#define BLK_FEAT_RAID_PARTIAL_STRIPES_EXPENSIVE \
((__force blk_features_t)(1u << 15))
+/* device writes each atomic write boundary window of a command atomically */
+#define BLK_FEAT_ATOMIC_WRITE_MULTI ((__force blk_features_t)(1u << 16))
+
/*
* Flags automatically inherited when stacking limits.
*/
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH RFC 2/2] nvme: enable multiple atomicity mode
2026-09-29 11:17 [PATCH RFC 0/2] Support for Multiple Atomicity Mode Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 1/2] block: add BLK_FEAT_ATOMIC_WRITE_MULTI Daniel Gomez
@ 2026-09-29 11:17 ` Daniel Gomez
1 sibling, 0 replies; 3+ messages in thread
From: Daniel Gomez @ 2026-09-29 11:17 UTC (permalink / raw)
To: Jens Axboe, Keith Busch, Christoph Hellwig, Sagi Grimberg
Cc: linux-block, linux-kernel, linux-nvme, Andres Freund,
Pankaj Raghav, Daniel Gomez, GOST, Daniel Gomez
From: Daniel Gomez <da.gomez@samsung.com>
Add support for Multiple Atomicity Mode (MAM), a superset of Single
Atomicity Mode (SAM) where the controller divides a write command that
crosses the atomic boundaries into per-window atomic writes. Honoring
the mode lets contiguous atomic writes merge past the atomic unit
limits.
Set BLK_FEAT_ATOMIC_WRITE_MULTI when the namespace reports MAM and its
atomic parameters are compliant; otherwise fall back to SAM. Re-derive
the flag on every rescan, and skip the single-write unit_max and
boundary checks because a merged atomic carrier can exceed both.
Assisted-by: LLM
Signed-off-by: Daniel Gomez <da.gomez@samsung.com>
---
drivers/nvme/host/core.c | 23 +++++++++++++++++++++++
include/linux/nvme.h | 1 +
2 files changed, 24 insertions(+)
diff --git a/drivers/nvme/host/core.c b/drivers/nvme/host/core.c
index 9bcab3dc4c118..69a565ac0da97 100644
--- a/drivers/nvme/host/core.c
+++ b/drivers/nvme/host/core.c
@@ -993,6 +993,9 @@ static bool nvme_valid_atomic_write(struct request *req)
struct request_queue *q = req->q;
u32 boundary_bytes = queue_atomic_write_boundary_bytes(q);
+ if (q->limits.features & BLK_FEAT_ATOMIC_WRITE_MULTI)
+ return true;
+
if (blk_rq_bytes(req) > queue_atomic_write_unit_max_bytes(q))
return false;
@@ -2038,12 +2041,24 @@ static void nvme_configure_metadata(struct nvme_ctrl *ctrl,
}
}
+static bool nvme_mam_compliant(struct nvme_id_ns *id)
+{
+ if (id->nabspf != id->nawupf)
+ return false;
+ if (id->nabsn && id->nabsn != id->nabspf)
+ return false;
+ if (id->nawun && id->nawun != id->nawupf)
+ return false;
+ return true;
+}
static u32 nvme_configure_atomic_write(struct nvme_ns *ns,
struct nvme_id_ns *id, struct queue_limits *lim, u32 bs)
{
u32 atomic_bs, boundary = 0;
+ lim->features &= ~BLK_FEAT_ATOMIC_WRITE_MULTI;
+
/*
* We do not support an offset for the atomic boundaries.
*/
@@ -2057,6 +2072,14 @@ static u32 nvme_configure_atomic_write(struct nvme_ns *ns,
atomic_bs = (1 + le16_to_cpu(id->nawupf)) * bs;
if (id->nabspf)
boundary = (le16_to_cpu(id->nabspf) + 1) * bs;
+
+ if (id->nsfeat & NVME_NS_FEAT_MAM) {
+ if (nvme_mam_compliant(id))
+ lim->features |= BLK_FEAT_ATOMIC_WRITE_MULTI;
+ else
+ dev_warn_once(ns->ctrl->device,
+ "Inconsistent MAM parameters, ignoring\n");
+ }
} else {
if (ns->ctrl->awupf)
dev_info_once(ns->ctrl->device,
diff --git a/include/linux/nvme.h b/include/linux/nvme.h
index 91ce434a7e8d9..8fdc4b91cf906 100644
--- a/include/linux/nvme.h
+++ b/include/linux/nvme.h
@@ -602,6 +602,7 @@ enum {
NVME_NS_FEAT_OPTPERF_MASK = 0x1,
/* Since version 2.1, OPTPERF is bits 4 and 5 of NSFEAT */
NVME_NS_FEAT_OPTPERF_MASK_2_1 = 0x3,
+ NVME_NS_FEAT_MAM = 1 << 6,
NVME_NS_ATTR_RO = 1 << 0,
NVME_NS_FLBAS_LBA_MASK = 0xf,
NVME_NS_FLBAS_LBA_UMASK = 0x60,
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-29 11:18 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 11:17 [PATCH RFC 0/2] Support for Multiple Atomicity Mode Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 1/2] block: add BLK_FEAT_ATOMIC_WRITE_MULTI Daniel Gomez
2026-09-29 11:17 ` [PATCH RFC 2/2] nvme: enable multiple atomicity mode Daniel Gomez
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox