* [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy
@ 2026-08-15 17:34 Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 01/10] block: expose blk_stat_{enable,disable}_accounting() to drivers Nilay Shroff
` (9 more replies)
0 siblings, 10 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
Hi,
This series introduces a new latency I/O policy for NVMe native
multipath. Existing policies such as numa, round-robin, and queue-depth
are static and do not adapt to real-time transport performance. The numa
selects the path closest to the NUMA node of the current CPU, optimizing
memory and path locality, but ignores actual path performance. The
round-robin distributes I/O evenly across all paths, providing fairness
but not performance awareness. The queue-depth reacts to instantaneous
queue occupancy, avoiding heavily loaded paths, but does not account for
actual latency, throughput, or link speed.
The new latency policy addresses these gaps selecting paths dynamically
based on measured I/O latency for both PCIe and fabrics. Latency is
derived by passively sampling I/O completions. Each path is assigned a
weight proportional to its latency score, and I/Os are then forwarded
accordingly. As condition changes (e.g. latency spikes, bandwidth
differences), path weights are updated, automatically steering traffic
toward better-performing paths.
Early results show reduced tail latency under mixed workloads and
improved throughput by exploiting higher-speed links more effectively.
For example, with NVMf/TCP using two paths (one throttled with ~30 ms
delay), fio results with random read/write/rw workloads (direct I/O)
showed:
numa round-robin queue-depth adaptive
----------- ----------- ----------- ---------
READ: 50.0 MiB/s 105 MiB/s 230 MiB/s 350 MiB/s
WRITE: 65.9 MiB/s 125 MiB/s 385 MiB/s 446 MiB/s
RW: R:30.6 MiB/s R:56.5 MiB/s R:122 MiB/s R:175 MiB/s
W:30.7 MiB/s W:56.5 MiB/s W:122 MiB/s W:175 MiB/s
This pathcset includes totla 8 patches:
[PATCH 1/10] block: expose blk_stat_{enable,disable}_accounting()
- Make blk_stat APIs available to block drivers.
- Needed for per-path latency measurement.
[PATCH 2/10] block: record I/O request start time for passthru request
- Record I/O start time for I/O passthru requests.
- This is prep patch which allows measuring I/O completion latency
for passthru requests.
[PATCH 3/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE
- Support nesting for QUEUE_FLAG_SAME_FORCE as multiple users
could toggle QUEUE_FLAG_SAME_FORCE.
[PATCH 4/10] nvme-multipath: pass I/O type to nvme_find_path()
- This is the prep patch which updates nvme_find_path() signature
[PATCH 5/10] nvme-multipath: add latency I/O policy
- Implement path scoring based on latency (EWMA).
- Distribute I/O proportionally to per-path weights.
[PATCH 6/10] nvme: add generic debugfs support
- Introduce generic debugfs support for NVMe module
[PATCH 7/10] nvme-multipath: add debugfs attribute latency_ewma_shift
- Adds a debugfs attribute to control ewma shift
[PATCH 8/10] nvme-multipath: add debugfs attribute latency_batch_timeout
- Adds a debugfs attribute to control latency batch window interval
[PATCH 9/10] nvme-multipath: add debugfs attribute latency_stat
- Add “latency_stat” under per-path and head debugfs directories to
expose latency policy state and statistics.
[PATCH 10/10] nvme-multipath: add documentation for latency I/O policy
- Includes documentation for latency I/O multipath policy.
LSFMM discussion:
=================
During lsfmm 2026, it was decided to rename this I/O policy from
"adaptive" to "latency". This series reflects that rename.
The discussion at lsfmm also focused extensively on the latency
measurement model, including whether latency should be tracked
per-CPU or per-NUMA, and whether separate I/O-size buckets should
be maintained for different request sizes.
After detailed discussion and evaluation of throughput results, the
consensus was to initially measure I/O completion latency on a
per-CPU basis. The available performance data showed that the
per-CPU implementation already provides sufficient averaging across
CPUs while keeping the design relatively simple.
The use of additional I/O-size buckets did not demonstrate meaningful
throughput improvement in the general case and would introduce extra
complexity into the fast path and accounting logic. As a result, the
consensus was to avoid I/O-size bucketing for now and keep the policy
focused on per-CPU latency measurement.
If future real-world workloads demonstrate a clear benefit from
I/O-size-aware latency accounting, the policy can be extended later
to support it.
As ususal, feedback and suggestions are most welcome!
Thanks!
Changes from v7:
- Rebased on nvme-7.3 (John Garry)
- Added Clang lock context annotations where appropriate (John Garry)
- Added support for nesting QUEUE_FLAG_SAME_FORCE in a new patch
3/10 (John Garry)
- Free head->latency_path in nvme_mpath_put_disk() instead of
nvme_free_ns_head() (John Garry)
- Replaced this_cpu_ptr() with per_cpu_ptr() (John Garry)
- Few miscellaneous and cosmetic changes such as adding comments
updating commit description etc.
Link to v7: https://lore.kernel.org/all/20260809100825.2014133-1-nilay@linux.ibm.com/
Changes from v6:
- Add patch 2/9, which records I/O start time for the passthru
requests (Guixin Liu)
- Correctly record the passthru I/O data direction (Guixin Liu)
- Clear NVME_NS_PATH_STAT before cancelling the latency weight work.
(Guixin Liu)
Link to v6: https://lore.kernel.org/all/20260520182112.863076-1-nilay@linux.ibm.com/
Changes from v5:
- Rename the policy from "adaptive" to "latency". The entire series
updates policy names, function names, and variable names accordingly,
without introducing any functional changes. (lsfmm discussion)
- The second patch is now splitted into two patches:
Patch #2: prep patch where we pass op_type to nvme_find_path()
Patch #3: core patch which introduces latency I/O policy
(Sagi)
- Rename ewma_update() to calc_ewma_update() (Sagi)
Link to v5: https://lore.kernel.org/all/20251105103347.86059-1-nilay@linux.ibm.com/
Changes from v4:
- Added patch #7 which includes the documentation for adaptive I/O
policy. (Guixin Liu)
Link to v4: https://lore.kernel.org/all/20251104104533.138481-1-nilay@linux.ibm.com/
Changes from v3:
- Update the adaptive APIs name (which actually enable/disable
adaptive policy) to reflect the actual work it does. Also removed
the misleading use of "current_path" from the adaptive policy code
(Hannes Reinecke)
- Move adaptive_ewma_shift and adaptive_weight_timeout attributes from
sysfs to debugfs (Hannes Reinecke)
Link to v3: https://lore.kernel.org/all/20251027092949.961287-1-nilay@linux.ibm.com/
Changes from v2:
- Addede a new patch to allow user to configure EWMA shift
through sysfs (Hannes Reinecke)
- Added a new patch to allow user to configure path weight
calculation timeout (Hannes Reinecke)
- Distinguish between read/write and other commands (e.g.
admin comamnd) and calculate path weight for other commands
which is separate from read/write weight. (Hannes Reinecke)
- Normalize per-path weight in the range from 0-128 instead
of 0-100 (Hannes Reinecke)
- Restructure and optimize adaptive I/O forwarding code to use
one loop instead of two (Hannes Reinecke)
Link to v2: https://lore.kernel.org/all/20251009100608.1699550-1-nilay@linux.ibm.com/
Changes from v1:
- Ensure that the completion of I/O occurs on the same CPU as the
submitting I/O CPU (Hannes Reinecke)
- Remove adapter link speed from the path weight calculation
(Hannes Reinecke)
- Add adaptive I/O stat under debugfs instead of current sysfs
(Hannes Reinecke)
- Move path weight calculation to a workqueue from IO completion
code path
Link to v1: https://lore.kernel.org/all/20250921111234.863853-1-nilay@linux.ibm.com/
Nilay Shroff (10):
block: expose blk_stat_{enable,disable}_accounting() to drivers
block: record I/O request start time for passthru request
block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE
nvme-multipath: pass I/O type to nvme_find_path()
nvme-multipath: add support for latency I/O policy
nvme: add generic debugfs support
nvme-multipath: add debugfs attribute latency_ewma_shift
nvme-multipath: add debugfs attribute latency_batch_timeout
nvme-multipath: add debugfs attribute latency_stat
nvme-multipath: add documentation for latency I/O policy
Documentation/admin-guide/nvme-multipath.rst | 19 +
block/blk-mq.c | 53 +-
block/blk-stat.h | 4 -
block/blk-sysfs.c | 6 +-
drivers/nvme/host/Makefile | 2 +-
drivers/nvme/host/core.c | 15 +-
drivers/nvme/host/debugfs.c | 347 ++++++++++++++
drivers/nvme/host/ioctl.c | 44 +-
drivers/nvme/host/multipath.c | 478 ++++++++++++++++++-
drivers/nvme/host/nvme.h | 110 ++++-
drivers/nvme/host/pr.c | 6 +-
drivers/nvme/host/sysfs.c | 2 +-
drivers/ufs/host/ufs-mediatek.c | 2 +-
include/linux/blk-mq.h | 6 +
include/linux/blkdev.h | 3 +
15 files changed, 1057 insertions(+), 40 deletions(-)
create mode 100644 drivers/nvme/host/debugfs.c
--
2.53.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v8 01/10] block: expose blk_stat_{enable,disable}_accounting() to drivers
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 02/10] block: record I/O request start time for passthru request Nilay Shroff
` (8 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
The functions blk_stat_enable_accounting() and
blk_stat_disable_accounting() are currently exported, but their
prototypes are only defined in a private header. Move these prototypes
into a common header so that block drivers can directly use these APIs.
Reviewed-by: John Garry <john.g.garry@oracle.com>
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
block/blk-stat.h | 4 ----
include/linux/blk-mq.h | 4 ++++
2 files changed, 4 insertions(+), 4 deletions(-)
diff --git a/block/blk-stat.h b/block/blk-stat.h
index cc5b66e7ee60..b1614a1ef4ad 100644
--- a/block/blk-stat.h
+++ b/block/blk-stat.h
@@ -70,10 +70,6 @@ void blk_free_queue_stats(struct blk_queue_stats *);
void blk_stat_add(struct request *rq, u64 now);
-/* record time/size info in request but not add a callback */
-void blk_stat_enable_accounting(struct request_queue *q);
-void blk_stat_disable_accounting(struct request_queue *q);
-
/**
* blk_stat_alloc_callback() - Allocate a block statistics callback.
* @timer_fn: Timer callback function.
diff --git a/include/linux/blk-mq.h b/include/linux/blk-mq.h
index af878597afb8..3956909764bf 100644
--- a/include/linux/blk-mq.h
+++ b/include/linux/blk-mq.h
@@ -753,6 +753,10 @@ int blk_rq_poll(struct request *rq, struct io_comp_batch *iob,
bool blk_mq_queue_inflight(struct request_queue *q);
+/* record time/size info in request but not add a callback */
+void blk_stat_enable_accounting(struct request_queue *q);
+void blk_stat_disable_accounting(struct request_queue *q);
+
enum {
/* return when out of requests */
BLK_MQ_REQ_NOWAIT = (__force blk_mq_req_flags_t)(1 << 0),
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 02/10] block: record I/O request start time for passthru request
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 01/10] block: expose blk_stat_{enable,disable}_accounting() to drivers Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 03/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE Nilay Shroff
` (7 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
While starting an I/O request, blk_mq_start_request() records the
request start timestamp only for non-passthrough requests when
QUEUE_FLAG_STATS is enabled.
However, the latency based multipath policy uses request completion
latency to evaluate path performance, and I/O is issued as passthrough
requests. Since passthru requests never initialize rq->io_start_time_ns,
their latency cannot be computed.
Record io_start_time_ns for all requests whenever QUEUE_FLAG_STATS is
enabled.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
block/blk-mq.c | 12 +++++++-----
1 file changed, 7 insertions(+), 5 deletions(-)
diff --git a/block/blk-mq.c b/block/blk-mq.c
index 2c850330a32b..38922209a24f 100644
--- a/block/blk-mq.c
+++ b/block/blk-mq.c
@@ -1340,12 +1340,14 @@ void blk_mq_start_request(struct request *rq)
trace_block_rq_issue(rq);
- if (test_bit(QUEUE_FLAG_STATS, &q->queue_flags) &&
- !blk_rq_is_passthrough(rq)) {
+ if (test_bit(QUEUE_FLAG_STATS, &q->queue_flags)) {
rq->io_start_time_ns = blk_time_get_ns();
- rq->stats_sectors = blk_rq_sectors(rq);
- rq->rq_flags |= RQF_STATS;
- rq_qos_issue(q, rq);
+
+ if (!blk_rq_is_passthrough(rq)) {
+ rq->stats_sectors = blk_rq_sectors(rq);
+ rq->rq_flags |= RQF_STATS;
+ rq_qos_issue(q, rq);
+ }
}
WARN_ON_ONCE(blk_mq_rq_state(rq) != MQ_RQ_IDLE);
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 03/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 01/10] block: expose blk_stat_{enable,disable}_accounting() to drivers Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 02/10] block: record I/O request start time for passthru request Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 04/10] nvme-multipath: pass I/O type to nvme_find_path() Nilay Shroff
` (6 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
QUEUE_FLAG_SAME_FORCE is currently used when setting rq_affinity
through sysfs as well as by UFS mediatek driver while configuring scsi
parameters. A subsequent patch adding a latency-based I/O policy
for NVMe multipath will also use this flag.
With multiple users able to set and clear QUEUE_FLAG_SAME_FORCE, the
flag needs to support nesting so that one user clearing the flag does
not inadvertently disable it for another user.
Add a nesting counter, q->same_force_depth, for QUEUE_FLAG_SAME_FORCE.
The flag is set when the first user acquires it and the nesting counter
is incremented for each subsequent user. Similarly, each user releases
its reference by decrementing the counter. The flag is cleared only
when the last user releases it and the counter reaches zero.
Preserve the existing sysfs rq_affinity semantics with a new
q->same_force_sysfs flag. When userspace enables QUEUE_FLAG_SAME_FORCE
by writing 2 to rq_affinity, mark q->same_force_sysfs as set and
increment q->same_force_depth by one. Subsequent writes of 2 to
rq_affinity while q->same_force_sysfs is already set are ignored, so
repeated writes of 2 from userspace do not increase q->same_force_depth.
Similarly, writing 0 or 1 decrements the q->same_force_depth and if
nesting counter reached to 0 then clears the QUEUE_FLAG_SAME_FORCE.
This ensures that multiple writes of 2 to rq_affinity do not require
multiple writes of 0 or 1.
This change ensures that sysfs interface retains its existing set/clear
semantics while also allowing other kernel users to hold or release
QUEUE_FLAG_SAME_FORCE.
Added two new APIs blk_mq_same_force_set() and blk_mq_same_force_clear()
to set and clear QUEUE_FLAG_SAME_FORCE respectively. Also, updated
existing call paths using these new APIs which toggles
QUEUE_FLAG_SAME_FORCE.
Cc: Peter Wang <peter.wang@mediatek.com>
Cc: Chaotian Jing <chaotian.jing@mediatek.com>
Cc: Stanley Jhu <chu.stanley@gmail.com>
Cc: linux-mediatek@lists.infradead.org
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
block/blk-mq.c | 41 +++++++++++++++++++++++++++++++++
block/blk-sysfs.c | 6 ++---
drivers/ufs/host/ufs-mediatek.c | 2 +-
include/linux/blk-mq.h | 2 ++
include/linux/blkdev.h | 3 +++
5 files changed, 50 insertions(+), 4 deletions(-)
diff --git a/block/blk-mq.c b/block/blk-mq.c
index 38922209a24f..17e8befb02bc 100644
--- a/block/blk-mq.c
+++ b/block/blk-mq.c
@@ -360,6 +360,47 @@ void blk_mq_unquiesce_tagset(struct blk_mq_tag_set *set)
}
EXPORT_SYMBOL_GPL(blk_mq_unquiesce_tagset);
+void blk_mq_same_force_set(struct request_queue *q, bool from_sysfs)
+{
+ unsigned long flags = 0;
+
+ spin_lock_irqsave(&q->queue_lock, flags);
+
+ if (from_sysfs) {
+ if (q->same_force_sysfs)
+ goto unlock;
+ q->same_force_sysfs = true;
+ }
+
+ if (!q->same_force_depth++)
+ blk_queue_flag_set(QUEUE_FLAG_SAME_FORCE, q);
+unlock:
+ spin_unlock_irqrestore(&q->queue_lock, flags);
+}
+EXPORT_SYMBOL_GPL(blk_mq_same_force_set);
+
+void blk_mq_same_force_clear(struct request_queue *q, bool from_sysfs)
+{
+ unsigned long flags = 0;
+
+ spin_lock_irqsave(&q->queue_lock, flags);
+
+ if (from_sysfs) {
+ if (!q->same_force_sysfs)
+ goto unlock;
+ q->same_force_sysfs = false;
+ }
+
+ if (WARN_ON_ONCE(q->same_force_depth <= 0))
+ goto unlock;
+
+ if (!--q->same_force_depth)
+ blk_queue_flag_clear(QUEUE_FLAG_SAME_FORCE, q);
+unlock:
+ spin_unlock_irqrestore(&q->queue_lock, flags);
+}
+EXPORT_SYMBOL_GPL(blk_mq_same_force_clear);
+
void blk_mq_wake_waiters(struct request_queue *q)
{
struct blk_mq_hw_ctx *hctx;
diff --git a/block/blk-sysfs.c b/block/blk-sysfs.c
index 520972676ab4..a3ec8ffab1ee 100644
--- a/block/blk-sysfs.c
+++ b/block/blk-sysfs.c
@@ -497,13 +497,13 @@ queue_rq_affinity_store(struct gendisk *disk, const char *page, size_t count)
*/
if (val == 2) {
blk_queue_flag_set(QUEUE_FLAG_SAME_COMP, q);
- blk_queue_flag_set(QUEUE_FLAG_SAME_FORCE, q);
+ blk_mq_same_force_set(q, true);
} else if (val == 1) {
blk_queue_flag_set(QUEUE_FLAG_SAME_COMP, q);
- blk_queue_flag_clear(QUEUE_FLAG_SAME_FORCE, q);
+ blk_mq_same_force_clear(q, true);
} else if (val == 0) {
blk_queue_flag_clear(QUEUE_FLAG_SAME_COMP, q);
- blk_queue_flag_clear(QUEUE_FLAG_SAME_FORCE, q);
+ blk_mq_same_force_clear(q, true);
}
#endif
return ret;
diff --git a/drivers/ufs/host/ufs-mediatek.c b/drivers/ufs/host/ufs-mediatek.c
index 3991a51263a6..8d80ac6fb5d2 100644
--- a/drivers/ufs/host/ufs-mediatek.c
+++ b/drivers/ufs/host/ufs-mediatek.c
@@ -2311,7 +2311,7 @@ static void ufs_mtk_config_scsi_dev(struct scsi_device *sdev)
dev_dbg(hba->dev, "lu %llu scsi device configured", sdev->lun);
if (sdev->lun == 2)
- blk_queue_flag_set(QUEUE_FLAG_SAME_FORCE, sdev->request_queue);
+ blk_mq_same_force_set(sdev->request_queue, false);
}
/*
diff --git a/include/linux/blk-mq.h b/include/linux/blk-mq.h
index 3956909764bf..30be3eb5a37b 100644
--- a/include/linux/blk-mq.h
+++ b/include/linux/blk-mq.h
@@ -943,6 +943,8 @@ void blk_mq_wait_quiesce_done(struct blk_mq_tag_set *set);
void blk_mq_quiesce_tagset(struct blk_mq_tag_set *set);
void blk_mq_unquiesce_tagset(struct blk_mq_tag_set *set);
void blk_mq_unquiesce_queue(struct request_queue *q);
+void blk_mq_same_force_set(struct request_queue *q, bool from_sysfs);
+void blk_mq_same_force_clear(struct request_queue *q, bool from_sysfs);
void blk_mq_delay_run_hw_queue(struct blk_mq_hw_ctx *hctx, unsigned long msecs);
void blk_mq_run_hw_queue(struct blk_mq_hw_ctx *hctx, bool async);
void blk_mq_run_hw_queues(struct request_queue *q, bool async);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 9213a5716f95..32d0fb47d73c 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -529,6 +529,9 @@ struct request_queue {
int quiesce_depth;
+ int same_force_depth;
+ bool same_force_sysfs;
+
struct gendisk *disk;
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 04/10] nvme-multipath: pass I/O type to nvme_find_path()
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (2 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 03/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 05/10] nvme-multipath: add support for latency I/O policy Nilay Shroff
` (5 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
Currently, nvme_find_path() only accepts an nvme_ns_head argument.
However, the upcoming latency-aware I/O policy also needs to know
the I/O type (read/write/other) associated with the request in order
to make path selection decisions.
Update nvme_find_path() to accept an additional argument describing
the I/O type. Classify requests into three categories: READ, WRITE,
and OTHER. Admin commands and I/O requests that are neither reads nor
writes are classified as OTHER.
This patch does not introduce any functional change and only prepares
the interface for subsequent latency-policy changes.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/ioctl.c | 44 +++++++++++++++++++++++++++---
drivers/nvme/host/multipath.c | 9 ++++---
drivers/nvme/host/nvme.h | 50 ++++++++++++++++++++++++++++++++++-
drivers/nvme/host/pr.c | 6 +++--
drivers/nvme/host/sysfs.c | 2 +-
5 files changed, 100 insertions(+), 11 deletions(-)
diff --git a/drivers/nvme/host/ioctl.c b/drivers/nvme/host/ioctl.c
index 6539d4750098..b0538ed2adbc 100644
--- a/drivers/nvme/host/ioctl.c
+++ b/drivers/nvme/host/ioctl.c
@@ -751,12 +751,25 @@ int nvme_ns_head_ioctl(struct block_device *bdev, blk_mode_t mode,
struct nvme_ns *ns;
int srcu_idx, ret = -EWOULDBLOCK;
unsigned int flags = 0;
+ unsigned int op_type = NVME_STAT_GROUP_OTHER;
if (bdev_is_partition(bdev))
flags |= NVME_IOCTL_PARTITION;
+ if (cmd == NVME_IOCTL_SUBMIT_IO) {
+ u8 opcode;
+
+ if (get_user(opcode, (u8 *)argp))
+ return -EFAULT;
+
+ if (opcode == nvme_cmd_write)
+ op_type = NVME_STAT_GROUP_WRITE;
+ else if (opcode == nvme_cmd_read)
+ op_type = NVME_STAT_GROUP_READ;
+ }
+
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, op_type);
if (!ns)
goto out_unlock;
@@ -785,9 +798,22 @@ long nvme_ns_head_chr_ioctl(struct file *file, unsigned int cmd,
void __user *argp = (void __user *)arg;
struct nvme_ns *ns;
int srcu_idx, ret = -EWOULDBLOCK;
+ unsigned int op_type = NVME_STAT_GROUP_OTHER;
+
+ if (cmd == NVME_IOCTL_SUBMIT_IO) {
+ u8 opcode;
+
+ if (get_user(opcode, (u8 *)argp))
+ return -EFAULT;
+
+ if (opcode == nvme_cmd_write)
+ op_type = NVME_STAT_GROUP_WRITE;
+ else if (opcode == nvme_cmd_read)
+ op_type = NVME_STAT_GROUP_READ;
+ }
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, op_type);
if (!ns)
goto out_unlock;
@@ -804,12 +830,24 @@ long nvme_ns_head_chr_ioctl(struct file *file, unsigned int cmd,
int nvme_ns_head_chr_uring_cmd(struct io_uring_cmd *ioucmd,
unsigned int issue_flags)
{
+ struct nvme_ns *ns;
+ unsigned int op_type;
struct cdev *cdev = file_inode(ioucmd->file)->i_cdev;
struct nvme_ns_head *head = container_of(cdev, struct nvme_ns_head, cdev);
int srcu_idx = srcu_read_lock(&head->srcu);
- struct nvme_ns *ns = nvme_find_path(head);
int ret = -EINVAL;
+ const struct nvme_uring_cmd *cmd = io_uring_sqe128_cmd(ioucmd->sqe,
+ struct nvme_uring_cmd);
+ __u8 opcode = READ_ONCE(cmd->opcode);
+
+ if (opcode == nvme_cmd_write)
+ op_type = NVME_STAT_GROUP_WRITE;
+ else if (opcode == nvme_cmd_read)
+ op_type = NVME_STAT_GROUP_READ;
+ else
+ op_type = NVME_STAT_GROUP_OTHER;
+ ns = nvme_find_path(head, op_type);
if (ns)
ret = nvme_ns_uring_cmd(ns, ioucmd, issue_flags);
srcu_read_unlock(&head->srcu, srcu_idx);
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 75dbb58286a3..95768eaef843 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -484,7 +484,8 @@ static struct nvme_ns *nvme_numa_path(struct nvme_ns_head *head)
return ns;
}
-inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head)
+inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head,
+ enum nvme_stat_group op_type)
{
switch (READ_ONCE(head->subsys->iopolicy)) {
case NVME_IOPOLICY_QD:
@@ -547,7 +548,7 @@ static void nvme_ns_head_submit_bio(struct bio *bio)
return;
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, __nvme_get_stat_group(bio_op(bio)));
if (likely(ns)) {
bio_set_dev(bio, ns->disk->part0);
/*
@@ -597,7 +598,7 @@ static int nvme_ns_head_get_unique_id(struct gendisk *disk, u8 id[16],
int srcu_idx, ret = -EWOULDBLOCK;
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, NVME_STAT_GROUP_OTHER);
if (ns)
ret = nvme_ns_get_unique_id(ns, id, type);
srcu_read_unlock(&head->srcu, srcu_idx);
@@ -613,7 +614,7 @@ static int nvme_ns_head_report_zones(struct gendisk *disk, sector_t sector,
int srcu_idx, ret = -EWOULDBLOCK;
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, NVME_STAT_GROUP_OTHER);
if (ns)
ret = nvme_ns_report_zones(ns, sector, nr_zones, args);
srcu_read_unlock(&head->srcu, srcu_idx);
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index 75e5d5a8a77c..5af272cf2f49 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -531,6 +531,20 @@ struct nvme_ns_ids {
u8 csi;
};
+/*
+ * Enum used to classify NVMe I/O type into a stat group. Read and write
+ * I/Os are classified as NVME_STAT_GROUP_READ and NVME_STAT_GROUP_WRITE
+ * respectively; all other I/Os or admin commands are classified as
+ * NVME_STAT_GROUP_OTHER.
+ */
+enum nvme_stat_group {
+ NVME_STAT_GROUP_READ,
+ NVME_STAT_GROUP_WRITE,
+ NVME_STAT_GROUP_OTHER,
+
+ NVME_NUM_STAT_GROUPS
+};
+
/*
* Anchor structure for namespaces. There is one for each namespace in a
* NVMe subsystem that any of our controllers can see, and the namespace
@@ -1046,8 +1060,42 @@ extern const struct attribute_group *nvme_dev_attr_groups[];
extern const struct block_device_operations nvme_bdev_ops;
void nvme_delete_ctrl_sync(struct nvme_ctrl *ctrl);
-struct nvme_ns *nvme_find_path(struct nvme_ns_head *head)
+struct nvme_ns *nvme_find_path(struct nvme_ns_head *head,
+ enum nvme_stat_group op_type)
__must_hold_shared(&head->srcu);
+
+static inline enum nvme_stat_group __nvme_get_stat_group(const enum req_op op)
+{
+ if (op == REQ_OP_READ)
+ return NVME_STAT_GROUP_READ;
+ else if (op == REQ_OP_WRITE)
+ return NVME_STAT_GROUP_WRITE;
+ else
+ return NVME_STAT_GROUP_OTHER;
+}
+
+static inline enum nvme_stat_group __nvme_get_passthru_stat_group(
+ enum nvme_opcode op)
+{
+ if (op == nvme_cmd_read)
+ return NVME_STAT_GROUP_READ;
+ else if (op == nvme_cmd_write)
+ return NVME_STAT_GROUP_WRITE;
+ else
+ return NVME_STAT_GROUP_OTHER;
+}
+
+static inline enum nvme_stat_group nvme_get_stat_group(struct request *req)
+{
+ if (blk_rq_is_passthrough(req)) {
+ struct nvme_request *nr = nvme_req(req);
+
+ return __nvme_get_passthru_stat_group(nr->cmd->common.opcode);
+ }
+
+ return __nvme_get_stat_group(req_op(req));
+}
+
#ifdef CONFIG_NVME_MULTIPATH
static inline bool nvme_ctrl_use_ana(struct nvme_ctrl *ctrl)
{
diff --git a/drivers/nvme/host/pr.c b/drivers/nvme/host/pr.c
index fe7dbe264815..715e6c242bd1 100644
--- a/drivers/nvme/host/pr.c
+++ b/drivers/nvme/host/pr.c
@@ -53,10 +53,12 @@ static int nvme_send_ns_head_pr_command(struct block_device *bdev,
struct nvme_command *c, void *data, unsigned int data_len)
{
struct nvme_ns_head *head = bdev->bd_disk->private_data;
- int srcu_idx = srcu_read_lock(&head->srcu);
- struct nvme_ns *ns = nvme_find_path(head);
+ int srcu_idx;
+ struct nvme_ns *ns;
int ret = -EWOULDBLOCK;
+ srcu_idx = srcu_read_lock(&head->srcu);
+ ns = nvme_find_path(head, NVME_STAT_GROUP_OTHER);
if (ns) {
c->common.nsid = cpu_to_le32(ns->head->ns_id);
ret = nvme_submit_sync_cmd(ns->queue, c, data, data_len);
diff --git a/drivers/nvme/host/sysfs.c b/drivers/nvme/host/sysfs.c
index abf8edaae371..e95543fecb2a 100644
--- a/drivers/nvme/host/sysfs.c
+++ b/drivers/nvme/host/sysfs.c
@@ -195,7 +195,7 @@ static int ns_head_update_nuse(struct nvme_ns_head *head)
return 0;
srcu_idx = srcu_read_lock(&head->srcu);
- ns = nvme_find_path(head);
+ ns = nvme_find_path(head, NVME_STAT_GROUP_OTHER);
if (!ns)
goto out_unlock;
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 05/10] nvme-multipath: add support for latency I/O policy
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (3 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 04/10] nvme-multipath: pass I/O type to nvme_find_path() Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 06/10] nvme: add generic debugfs support Nilay Shroff
` (4 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
This commit introduces a new I/O policy named "latency". Users can
configure it by writing "latency" to "/sys/class/nvme-subsystem/nvme-
subsystemX/iopolicy"
The "latency" policy dynamically distributes I/O based on measured I/O
completion latency. The main idea is to calculate latency for each path,
derive a weight, and then proportionally forward I/O according to those
weights.
To ensure scalability, path latency is measured per-CPU. Each CPU
maintains its own statistics, and I/O forwarding uses these per-CPU
values. Every ~15 seconds, a simple average latency of per-CPU batched
samples are computed and fed into an Exponentially Weighted Moving
Average (EWMA):
avg_latency = div_u64(batch, batch_count);
new_ewma_latency = (prev_ewma_latency * (WEIGHT-1) + avg_latency)/WEIGHT
With WEIGHT = 8, this assigns 7/8 (~87.5%) weight to the previous
latency value and 1/8 (~12.5%) to the most recent latency. This
smoothing reduces jitter, adapts quickly to changing conditions,
avoids storing historical samples, and works well for both low and
high I/O rates. Path weights are then derived from the smoothed (EWMA)
latency as follows (example with two paths A and B):
path_A_score = NSEC_PER_SEC / path_A_ewma_latency
path_B_score = NSEC_PER_SEC / path_B_ewma_latency
total_score = path_A_score + path_B_score
path_A_weight = (path_A_score * 64) / total_score
path_B_weight = (path_B_score * 64) / total_score
where:
- path_X_ewma_latency is the smoothed latency of a path in nanoseconds
- NSEC_PER_SEC is used as a scaling factor since valid latencies
are < 1 second
- weights are normalized to a 0–64 scale across all paths.
Path credits are refilled based on this weight, with one credit
consumed per I/O. When all credits are consumed, the credits are
refilled again based on the current weight. This ensures that I/O is
distributed across paths proportionally to their calculated weight.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/core.c | 12 +-
drivers/nvme/host/multipath.c | 462 +++++++++++++++++++++++++++++++++-
drivers/nvme/host/nvme.h | 48 +++-
3 files changed, 507 insertions(+), 15 deletions(-)
diff --git a/drivers/nvme/host/core.c b/drivers/nvme/host/core.c
index 1322c678f4eb..a480e33fd984 100644
--- a/drivers/nvme/host/core.c
+++ b/drivers/nvme/host/core.c
@@ -719,6 +719,7 @@ static void nvme_free_ns(struct kref *kref)
{
struct nvme_ns *ns = container_of(kref, struct nvme_ns, kref);
+ nvme_free_ns_stat(ns);
put_disk(ns->disk);
nvme_put_ns_head(ns->head);
nvme_put_ctrl(ns->ctrl);
@@ -4258,6 +4259,9 @@ static void nvme_alloc_ns(struct nvme_ctrl *ctrl, struct nvme_ns_info *info)
if (nvme_init_ns_head(ns, info))
goto out_cleanup_disk;
+ if (nvme_mpath_alloc_ns_stat(ns))
+ goto out_unlink_ns;
+
/*
* If multipathing is enabled, the device name for all disks and not
* just those that represent shared namespaces needs to be based on the
@@ -4282,7 +4286,7 @@ static void nvme_alloc_ns(struct nvme_ctrl *ctrl, struct nvme_ns_info *info)
}
if (nvme_update_ns_info(ns, info))
- goto out_unlink_ns;
+ goto out_free_ns_stat;
mutex_lock(&ctrl->namespaces_lock);
/*
@@ -4291,7 +4295,7 @@ static void nvme_alloc_ns(struct nvme_ctrl *ctrl, struct nvme_ns_info *info)
*/
if (test_bit(NVME_CTRL_FROZEN, &ctrl->flags)) {
mutex_unlock(&ctrl->namespaces_lock);
- goto out_unlink_ns;
+ goto out_free_ns_stat;
}
blk_queue_rq_timeout(ns->queue, ctrl->io_timeout);
nvme_ns_add_to_ctrl_list(ns);
@@ -4316,6 +4320,8 @@ static void nvme_alloc_ns(struct nvme_ctrl *ctrl, struct nvme_ns_info *info)
list_del_rcu(&ns->list);
mutex_unlock(&ctrl->namespaces_lock);
synchronize_srcu(&ctrl->srcu);
+out_free_ns_stat:
+ nvme_free_ns_stat(ns);
out_unlink_ns:
mutex_lock(&ctrl->subsys->lock);
list_del_rcu(&ns->siblings);
@@ -4355,7 +4361,7 @@ static void nvme_ns_remove(struct nvme_ns *ns)
/*
* Ensure that !NVME_NS_READY is seen by other threads to prevent
- * this ns going back into current_path.
+ * this ns going back into current_path/latency_path.
*/
synchronize_srcu(&ns->head->srcu);
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 95768eaef843..5501c22bd662 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -4,7 +4,10 @@
*/
#include <linux/backing-dev.h>
+#include <linux/blk-mq.h>
+#include <linux/math64.h>
#include <linux/moduleparam.h>
+#include <linux/rculist.h>
#include <linux/vmalloc.h>
#include <trace/events/block.h>
#include "nvme.h"
@@ -66,9 +69,10 @@ MODULE_PARM_DESC(multipath_always_on,
"create multipath node always except for private namespace with non-unique nsid; note that this also implicitly enables native multipath support");
static const char *nvme_iopolicy_names[] = {
- [NVME_IOPOLICY_NUMA] = "numa",
- [NVME_IOPOLICY_RR] = "round-robin",
- [NVME_IOPOLICY_QD] = "queue-depth",
+ [NVME_IOPOLICY_NUMA] = "numa",
+ [NVME_IOPOLICY_RR] = "round-robin",
+ [NVME_IOPOLICY_QD] = "queue-depth",
+ [NVME_IOPOLICY_LATENCY] = "latency",
};
static int iopolicy = NVME_IOPOLICY_NUMA;
@@ -107,7 +111,7 @@ static int nvme_get_iopolicy(char *buf, const struct kernel_param *kp)
module_param_call(iopolicy, nvme_set_iopolicy, nvme_get_iopolicy,
&iopolicy, 0644);
MODULE_PARM_DESC(iopolicy,
- "Default multipath I/O policy; 'numa' (default), 'round-robin' or 'queue-depth'");
+ "Default multipath I/O policy; 'numa' (default), 'round-robin' or 'queue-depth' or 'latency'");
void nvme_mpath_default_iopolicy(struct nvme_subsystem *subsys)
{
@@ -199,6 +203,204 @@ void nvme_mpath_start_request(struct request *rq)
}
EXPORT_SYMBOL_GPL(nvme_mpath_start_request);
+static void nvme_mpath_weight_work(struct work_struct *weight_work)
+{
+ int cpu, srcu_idx;
+ u32 weight;
+ struct nvme_ns *ns;
+ struct nvme_path_lat_stat *stat;
+ struct nvme_path_lat_work *work = container_of(weight_work,
+ struct nvme_path_lat_work, weight_work);
+ struct nvme_ns_head *head = work->ns->head;
+ enum nvme_stat_group op_type = work->op_type;
+ u64 total_score = 0;
+
+ cpu = get_cpu();
+
+ srcu_idx = srcu_read_lock(&head->srcu);
+ list_for_each_entry_srcu(ns, &head->list, siblings,
+ srcu_read_lock_held(&head->srcu)) {
+
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[op_type].stat;
+ if (!READ_ONCE(stat->slat_ns)) {
+ stat->score = 0;
+ continue;
+ }
+ /*
+ * Compute the path score as the inverse of smoothed
+ * latency, scaled by NSEC_PER_SEC. Floating point
+ * math is unavailable in the kernel, so fixed-point
+ * scaling is used instead. NSEC_PER_SEC is chosen
+ * because valid latencies are always < 1 second; longer
+ * latencies are ignored.
+ */
+ stat->score = div_u64(NSEC_PER_SEC, READ_ONCE(stat->slat_ns));
+
+ /* Compute total score. */
+ total_score += stat->score;
+ }
+
+ if (!total_score)
+ goto out;
+
+ /*
+ * After computing the total slatency, we derive per-path weight
+ * (normalized to the range 0–64). The weight represents the
+ * relative share of I/O the path should receive.
+ *
+ * - lower smoothed latency -> higher weight
+ * - higher smoothed slatency -> lower weight
+ *
+ * Next, while forwarding I/O, we assign "credits" to each path
+ * based on its weight (please also refer nvme_latency_path()):
+ * - Initially, credits = weight.
+ * - Each time an I/O is dispatched on a path, its credits are
+ * decremented proportionally.
+ * - When a path runs out of credits, it becomes temporarily
+ * ineligible until credit is refilled.
+ *
+ * I/O distribution is therefore governed by available credits,
+ * ensuring that over time the proportion of I/O sent to each
+ * path matches its weight (and thus its performance).
+ */
+ list_for_each_entry_srcu(ns, &head->list, siblings,
+ srcu_read_lock_held(&head->srcu)) {
+
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[op_type].stat;
+ weight = div_u64(stat->score * 64, total_score);
+
+ /*
+ * Ensure the path weight never drops below 1. A weight
+ * of 0 is used only for newly added paths. During
+ * bootstrap, a few I/Os are sent to such paths to
+ * establish an initial weight. Enforcing a minimum
+ * weight of 1 guarantees that no path is forgotten and
+ * that each path is probed at least occasionally.
+ */
+ if (!weight)
+ weight = 1;
+
+ WRITE_ONCE(stat->weight, weight);
+ }
+out:
+ srcu_read_unlock(&head->srcu, srcu_idx);
+ put_cpu();
+}
+
+/*
+ * Formula to calculate the EWMA (Exponentially Weighted Moving Average):
+ * ewma = (old_ewma * (EWMA_SHIFT - 1) + (EWMA_SHIFT)) / EWMA_SHIFT
+ * For instance, with EWMA_SHIFT = 3, this assigns 7/8 (~87.5 %) weight to
+ * the existing/old ewma and 1/8 (~12.5%) weight to the new sample.
+ */
+static inline u64 calc_ewma_update(u64 old, u64 new)
+{
+ return (old * ((1 << NVME_DEFAULT_LATENCY_EWMA_SHIFT) - 1)
+ + new) >> NVME_DEFAULT_LATENCY_EWMA_SHIFT;
+}
+
+static void nvme_mpath_add_sample(struct request *rq, struct nvme_ns *ns)
+ __must_hold_shared(&ns->head->srcu)
+{
+ int cpu;
+ enum nvme_stat_group op_type;
+ struct nvme_path_lat *path_lat;
+ struct nvme_path_lat_stat *stat;
+ u64 now, latency, slat_ns, avg_lat_ns;
+ struct nvme_ns_head *head = ns->head;
+
+ if (list_is_singular(&head->list))
+ return;
+
+ now = ktime_get_ns();
+ latency = now >= rq->io_start_time_ns ? now - rq->io_start_time_ns : 0;
+ if (!latency)
+ return;
+
+ /*
+ * As completion code path is serialized(i.e. no same completion queue
+ * update code could run simultaneously on multiple cpu) we can safely
+ * access per cpu nvme path stat here from another cpu (in case the
+ * completion cpu is different from submission cpu).
+ * The only field which could be accessed simultaneously here is the
+ * path ->weight which may be accessed by this function as well as I/O
+ * submission path during path selection logic and we protect ->weight
+ * using READ_ONCE/WRITE_ONCE. Yes this may not be 100% accurate but
+ * we also don't need to be so accurate here as the path credit would
+ * be anyways refilled, based on path weight, once path consumes all
+ * its credits. And we limit path weight/credit max up to 64. Please
+ * also refer nvme_latency_path().
+ */
+ cpu = blk_mq_rq_cpu(rq);
+ op_type = nvme_get_stat_group(rq);
+ path_lat = &per_cpu_ptr(ns->path_lat, cpu)[op_type];
+ stat = &path_lat->stat;
+
+ /*
+ * If latency > ~1s then ignore this sample to prevent EWMA from being
+ * skewed by pathological outliers (multi-second waits, controller
+ * timeouts etc.). This keeps path scores representative of normal
+ * performance and avoids instability from rare spikes. If such high
+ * latency is real, ANA state reporting or keep-alive error counters
+ * will mark the path unhealthy and remove it from the head node list,
+ * so we safely skip such sample here.
+ */
+ if (unlikely(latency > NSEC_PER_SEC)) {
+ stat->nr_ignored++;
+ dev_warn_once(ns->ctrl->device,
+ "ignoring sample with >1s latency (possible controller stall or timeout)\n");
+ return;
+ }
+
+ /*
+ * Accumulate latency samples and increment the batch count for each
+ * ~15 second interval. When the interval expires, compute the simple
+ * average latency over that window, then update the smoothed (EWMA)
+ * latency. The path weight is recalculated based on this smoothed
+ * latency.
+ */
+ stat->batch += latency;
+ stat->batch_count++;
+ stat->nr_samples++;
+
+ if (now > stat->last_batch_ts && ((now - stat->last_batch_ts) >=
+ NVME_DEFAULT_LATENCY_BATCH_TIMEOUT)) {
+
+ /*
+ * Find simple average latency for the last epoch (~15 sec
+ * interval).
+ */
+ avg_lat_ns = div_u64(stat->batch, stat->batch_count);
+ stat->last_batch_ts = now;
+
+ /*
+ * Calculate smooth/EWMA (Exponentially Weighted Moving Average)
+ * latency. EWMA is preferred over simple average latency
+ * because it smooths naturally, reduces jitter from sudden
+ * spikes, and adapts faster to changing conditions. It also
+ * avoids storing historical samples, and works well for both
+ * slow and fast I/O rates.
+ * Formula:
+ * slat_ns = (prev_slat_ns * (WEIGHT - 1) + (latency)) / WEIGHT
+ * With WEIGHT = 8, this assigns 7/8 (~87.5 %) weight to the
+ * existing latency and 1/8 (~12.5%) weight to the new latency.
+ */
+ if (unlikely(!stat->slat_ns))
+ WRITE_ONCE(stat->slat_ns, avg_lat_ns);
+ else {
+ slat_ns = calc_ewma_update(stat->slat_ns, avg_lat_ns);
+ WRITE_ONCE(stat->slat_ns, slat_ns);
+ }
+
+ stat->batch = stat->batch_count = 0;
+
+ /*
+ * Defer calculation of the path weight in per-cpu workqueue.
+ */
+ schedule_work_on(cpu, &path_lat->work.weight_work);
+ }
+}
+
void nvme_mpath_end_request(struct request *rq)
{
struct nvme_ns *ns = rq->q->queuedata;
@@ -206,6 +408,23 @@ void nvme_mpath_end_request(struct request *rq)
if (nvme_req(rq)->flags & NVME_MPATH_CNT_ACTIVE)
atomic_dec_if_positive(&ns->ctrl->nr_active);
+ if (test_bit(NVME_NS_PATH_STAT, &ns->flags)) {
+ int srcu_idx;
+
+ srcu_idx = srcu_read_lock(&ns->head->srcu);
+ /*
+ * We re-evaluate the NVME_NS_PATH_STAT here because the
+ * first test above is fast-path optimization to avoid taking
+ * srcu lock when latency policy is disabled. However the second
+ * test is needed to address the case when NVME_NS_PATH_STAT is
+ * cleared after the first test but before we acquire the srcu
+ * lock.
+ */
+ if (test_bit(NVME_NS_PATH_STAT, &ns->flags))
+ nvme_mpath_add_sample(rq, ns);
+ srcu_read_unlock(&ns->head->srcu, srcu_idx);
+ }
+
if (!(nvme_req(rq)->flags & NVME_MPATH_IO_STATS))
return;
bdev_end_io_acct(ns->head->disk->part0, req_op(rq),
@@ -239,6 +458,83 @@ static const char *nvme_ana_state_names[] = {
[NVME_ANA_CHANGE] = "change",
};
+static void nvme_reset_ns_latency_stat(struct nvme_ns *ns)
+{
+ struct nvme_path_lat_stat *stat;
+ int i, cpu;
+
+ for_each_possible_cpu(cpu) {
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[i].stat;
+ memset(stat, 0, sizeof(struct nvme_path_lat_stat));
+ }
+ }
+}
+
+static void nvme_cancel_ns_latency_weight_work(struct nvme_ns *ns)
+{
+ int i, cpu;
+ struct nvme_path_lat *path_lat;
+
+ for_each_possible_cpu(cpu) {
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ path_lat = &per_cpu_ptr(ns->path_lat, cpu)[i];
+ cancel_work_sync(&path_lat->work.weight_work);
+ }
+ }
+}
+
+static void nvme_enable_ns_latency_sampling(struct nvme_ns *ns)
+{
+ struct nvme_ns_head *head = ns->head;
+
+ if (!head->disk ||
+ READ_ONCE(head->subsys->iopolicy) != NVME_IOPOLICY_LATENCY)
+ return;
+
+ if (test_and_set_bit(NVME_NS_PATH_STAT, &ns->flags))
+ return;
+
+ /*
+ * Force I/O completion to run on the same CPU as I/O submission
+ * so that the per-CPU path statistics are updated on the CPU that
+ * submitted the I/O. This also helps avoid cross-CPU IPIs and
+ * associated cache bouncing.
+ */
+ blk_mq_same_force_set(ns->queue, false);
+ blk_stat_enable_accounting(ns->queue);
+}
+
+static bool nvme_disable_ns_latency_sampling(struct nvme_ns *ns)
+{
+ int cpu;
+ struct nvme_ns_head *head = ns->head;
+ bool changed = false;
+
+ if (!test_and_clear_bit(NVME_NS_PATH_STAT, &ns->flags))
+ return false;
+
+ for_each_possible_cpu(cpu) {
+ if (ns == READ_ONCE(*per_cpu_ptr(head->latency_path, cpu))) {
+ WRITE_ONCE(*per_cpu_ptr(head->latency_path, cpu), NULL);
+ changed = true;
+ }
+ }
+
+ blk_stat_disable_accounting(ns->queue);
+ blk_mq_same_force_clear(ns->queue, false);
+
+ /*
+ * Ensure that we wait until completion side samplings (if any sneaked
+ * in after we clear NVME_NS_PATH_STAT) are all scheduled before we
+ * start cancelling those.
+ */
+ synchronize_srcu(&head->srcu);
+ nvme_cancel_ns_latency_weight_work(ns);
+ nvme_reset_ns_latency_stat(ns);
+ return changed;
+}
+
bool nvme_mpath_clear_current_path(struct nvme_ns *ns)
{
struct nvme_ns_head *head = ns->head;
@@ -251,6 +547,10 @@ bool nvme_mpath_clear_current_path(struct nvme_ns *ns)
changed = true;
}
}
+
+ if (nvme_disable_ns_latency_sampling(ns))
+ changed = true;
+
return changed;
}
@@ -268,6 +568,45 @@ void nvme_mpath_clear_ctrl_paths(struct nvme_ctrl *ctrl)
srcu_read_unlock(&ctrl->srcu, srcu_idx);
}
+int nvme_mpath_alloc_ns_stat(struct nvme_ns *ns)
+{
+ int i, cpu;
+ struct nvme_path_lat_work *work;
+ gfp_t gfp = GFP_KERNEL | __GFP_ZERO;
+
+ if (!ns->head->disk)
+ return 0;
+
+ ns->path_lat = __alloc_percpu_gfp(NVME_NUM_STAT_GROUPS *
+ sizeof(struct nvme_path_lat),
+ __alignof__(struct nvme_path_lat), gfp);
+ if (!ns->path_lat)
+ return -ENOMEM;
+
+ for_each_possible_cpu(cpu) {
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ work = &per_cpu_ptr(ns->path_lat, cpu)[i].work;
+ work->ns = ns;
+ work->op_type = i;
+ INIT_WORK(&work->weight_work, nvme_mpath_weight_work);
+ }
+ }
+
+ return 0;
+}
+
+static void nvme_mpath_set_ctrl_paths(struct nvme_ctrl *ctrl)
+{
+ struct nvme_ns *ns;
+ int srcu_idx;
+
+ srcu_idx = srcu_read_lock(&ctrl->srcu);
+ list_for_each_entry_srcu(ns, &ctrl->namespaces, list,
+ srcu_read_lock_held(&ctrl->srcu))
+ nvme_enable_ns_latency_sampling(ns);
+ srcu_read_unlock(&ctrl->srcu, srcu_idx);
+}
+
void nvme_mpath_revalidate_paths(struct nvme_ns_head *head)
{
sector_t capacity = get_capacity(head->disk);
@@ -280,6 +619,8 @@ void nvme_mpath_revalidate_paths(struct nvme_ns_head *head)
srcu_read_lock_held(&head->srcu)) {
if (capacity != get_capacity(ns->disk))
clear_bit(NVME_NS_READY, &ns->flags);
+
+ nvme_reset_ns_latency_stat(ns);
}
srcu_read_unlock(&head->srcu, srcu_idx);
@@ -426,6 +767,91 @@ static struct nvme_ns *nvme_round_robin_path(struct nvme_ns_head *head)
return found;
}
+static inline bool nvme_state_is_live(enum nvme_ana_state state)
+{
+ return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
+}
+
+static struct nvme_ns *nvme_latency_path(struct nvme_ns_head *head,
+ enum nvme_stat_group op_type)
+{
+ struct nvme_ns *ns, *start, *found = NULL;
+ struct nvme_path_lat_stat *stat;
+ u32 weight;
+ int cpu;
+
+ cpu = get_cpu();
+ ns = READ_ONCE(*per_cpu_ptr(head->latency_path, cpu));
+ if (unlikely(!ns)) {
+ ns = list_first_or_null_rcu(&head->list,
+ struct nvme_ns, siblings);
+ if (unlikely(!ns))
+ goto out_put;
+ }
+found_ns:
+ start = ns;
+ while (nvme_path_is_disabled(ns) ||
+ !nvme_state_is_live(ns->ana_state)) {
+ ns = list_next_entry_circular(ns, &head->list, siblings);
+
+ /*
+ * If we iterate through all paths in the list but find each
+ * path in list is either disabled or dead then bail out.
+ */
+ if (ns == start)
+ goto out_put;
+ }
+
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[op_type].stat;
+
+ /*
+ * When the head path-list is singular we don't calculate the
+ * only path weight for optimization as we don't need to forward
+ * I/O to more than one path. The another possibility is when the
+ * path is newly added, we don't know its weight. So we go round
+ * -robin for each such path and forward I/O to it.Once we start
+ * getting response for such I/Os, the path weight calculation
+ * would kick in and then we start using path credit for
+ * forwarding I/O.
+ */
+ weight = READ_ONCE(stat->weight);
+ if (!weight) {
+ found = ns;
+ goto out;
+ }
+
+ /*
+ * To keep path selection logic simple, we don't distinguish
+ * between ANA optimized and non-optimized states. The non-
+ * optimized path is expected to have a lower weight, and
+ * therefore fewer credits. As a result, only a small number of
+ * I/Os will be forwarded to paths in the non-optimized state.
+ */
+ if (stat->credit > 0) {
+ --stat->credit;
+ found = ns;
+ } else {
+ /*
+ * Refill credit from path weight and move to next path. The
+ * refilled credit of the current path will be used next when
+ * all remainng paths exhaust its credits.
+ */
+ weight = READ_ONCE(stat->weight);
+ stat->credit = weight;
+ ns = list_next_entry_circular(ns, &head->list, siblings);
+ if (likely(ns))
+ goto found_ns;
+ }
+out:
+ if (found) {
+ stat->sel++;
+ WRITE_ONCE(*per_cpu_ptr(head->latency_path, cpu), found);
+ }
+out_put:
+ put_cpu();
+ return found;
+}
+
static struct nvme_ns *nvme_queue_depth_path(struct nvme_ns_head *head)
__must_hold_shared(&head->srcu)
{
@@ -488,6 +914,8 @@ inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head,
enum nvme_stat_group op_type)
{
switch (READ_ONCE(head->subsys->iopolicy)) {
+ case NVME_IOPOLICY_LATENCY:
+ return nvme_latency_path(head, op_type);
case NVME_IOPOLICY_QD:
return nvme_queue_depth_path(head);
case NVME_IOPOLICY_RR:
@@ -760,6 +1188,10 @@ int nvme_mpath_alloc_disk(struct nvme_ctrl *ctrl, struct nvme_ns_head *head)
if (!nvme_is_unique_nsid(ctrl, head))
return 0;
+ head->latency_path = alloc_percpu_gfp(struct nvme_ns*, GFP_KERNEL);
+ if (!head->latency_path)
+ return -ENOMEM;
+
blk_set_stacking_limits(&lim);
lim.dma_alignment = 3;
lim.features |= BLK_FEAT_IO_STAT | BLK_FEAT_NOWAIT |
@@ -768,8 +1200,10 @@ int nvme_mpath_alloc_disk(struct nvme_ctrl *ctrl, struct nvme_ns_head *head)
lim.features |= BLK_FEAT_ZONED;
head->disk = blk_alloc_disk(&lim, ctrl->numa_node);
- if (IS_ERR(head->disk))
+ if (IS_ERR(head->disk)) {
+ free_percpu(head->latency_path);
return PTR_ERR(head->disk);
+ }
head->disk->fops = &nvme_ns_head_ops;
head->disk->private_data = head;
@@ -825,6 +1259,14 @@ static void nvme_mpath_set_live(struct nvme_ns *ns)
}
mutex_unlock(&head->lock);
+ /*
+ * Serialize access to latency sampling with nvme_subsystems_lock
+ * to prevent nvme_subsys_iopolicy_update() from running concurrently.
+ */
+ mutex_lock(&nvme_subsystems_lock);
+ nvme_enable_ns_latency_sampling(ns);
+ mutex_unlock(&nvme_subsystems_lock);
+
synchronize_srcu(&head->srcu);
nvme_mpath_revalidate_zones(head);
kblockd_schedule_work(&head->requeue_work);
@@ -875,11 +1317,6 @@ static int nvme_parse_ana_log(struct nvme_ctrl *ctrl, void *data,
return 0;
}
-static inline bool nvme_state_is_live(enum nvme_ana_state state)
-{
- return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
-}
-
static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
struct nvme_ns *ns)
{
@@ -1057,10 +1494,12 @@ static void nvme_subsys_iopolicy_update(struct nvme_subsystem *subsys,
WRITE_ONCE(subsys->iopolicy, iopolicy);
- /* iopolicy changes clear the mpath by design */
+ /* iopolicy changes clear/reset the mpath by design */
mutex_lock(&nvme_subsystems_lock);
list_for_each_entry(ctrl, &subsys->ctrls, subsys_entry)
nvme_mpath_clear_ctrl_paths(ctrl);
+ list_for_each_entry(ctrl, &subsys->ctrls, subsys_entry)
+ nvme_mpath_set_ctrl_paths(ctrl);
mutex_unlock(&nvme_subsystems_lock);
pr_notice("subsysnqn %s iopolicy changed from %s to %s\n",
@@ -1431,6 +1870,7 @@ void nvme_mpath_put_disk(struct nvme_ns_head *head)
kblockd_schedule_work(&head->requeue_work);
flush_work(&head->requeue_work);
flush_work(&head->partition_scan_work);
+ free_percpu(head->latency_path);
put_disk(head->disk);
}
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index 5af272cf2f49..850a5f6afca5 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -28,7 +28,9 @@ extern unsigned int nvme_io_timeout;
extern unsigned int admin_timeout;
#define NVME_ADMIN_TIMEOUT (admin_timeout * HZ)
-#define NVME_DEFAULT_KATO 5
+#define NVME_DEFAULT_KATO 5
+#define NVME_DEFAULT_LATENCY_EWMA_SHIFT 3
+#define NVME_DEFAULT_LATENCY_BATCH_TIMEOUT (15 * NSEC_PER_SEC)
#ifdef CONFIG_ARCH_NO_SG_CHAIN
#define NVME_INLINE_SG_CNT 0
@@ -491,6 +493,7 @@ enum nvme_iopolicy {
NVME_IOPOLICY_NUMA,
NVME_IOPOLICY_RR,
NVME_IOPOLICY_QD,
+ NVME_IOPOLICY_LATENCY,
};
struct nvme_subsystem {
@@ -545,6 +548,30 @@ enum nvme_stat_group {
NVME_NUM_STAT_GROUPS
};
+struct nvme_path_lat_stat {
+ u64 nr_samples; /* total num of samples processed */
+ u64 nr_ignored; /* num. of samples ignored */
+ u64 slat_ns; /* smoothed (ewma) latency in nanoseconds */
+ u64 score; /* score used for weight calculation */
+ u64 last_batch_ts; /* timestamp when last time avg. latency is calculated */
+ u64 sel; /* num of times this path is selcted for I/O */
+ u64 batch; /* accumulated latency sum for current window */
+ u32 batch_count; /* num of samples accumulated in current window */
+ u32 weight; /* path weight */
+ u32 credit; /* path credit for I/O forwarding */
+};
+
+struct nvme_path_lat_work {
+ struct nvme_ns *ns; /* owning namespace */
+ struct work_struct weight_work; /* deferred work for weight calculation */
+ enum nvme_stat_group op_type; /* op type : READ/WRITE/OTHER */
+};
+
+struct nvme_path_lat {
+ struct nvme_path_lat_stat stat; /* path statistics */
+ struct nvme_path_lat_work work; /* background worker context */
+};
+
/*
* Anchor structure for namespaces. There is one for each namespace in a
* NVMe subsystem that any of our controllers can see, and the namespace
@@ -598,6 +625,8 @@ struct nvme_ns_head {
__guarded_by(&subsys->lock);
atomic_long_t io_requeue_no_usable_path_count;
atomic_long_t io_fail_no_available_path_count;
+ struct nvme_ns * __percpu *latency_path;
+
#define NVME_NSHEAD_DISK_LIVE 0
#define NVME_NSHEAD_QUEUE_IF_NO_PATH 1
#define NVME_NSHEAD_CDEV_LIVE 2
@@ -626,6 +655,7 @@ struct nvme_ns {
enum nvme_ana_state ana_state;
u32 ana_grpid;
atomic_long_t failover;
+ struct nvme_path_lat __percpu *path_lat;
#endif
atomic_long_t retries;
atomic_long_t errors;
@@ -640,6 +670,7 @@ struct nvme_ns {
#define NVME_NS_READY 4
#define NVME_NS_SYSFS_ATTR_LINK 5
#define NVME_NS_CDEV_LIVE 6
+#define NVME_NS_PATH_STAT 7
struct cdev cdev;
struct device cdev_device;
@@ -1127,6 +1158,7 @@ void nvme_mpath_clear_ctrl_paths(struct nvme_ctrl *ctrl);
void nvme_mpath_remove_disk(struct nvme_ns_head *head);
void nvme_mpath_start_request(struct request *rq);
void nvme_mpath_end_request(struct request *rq);
+int nvme_mpath_alloc_ns_stat(struct nvme_ns *ns);
static inline void nvme_trace_bio_complete(struct request *req)
{
@@ -1157,6 +1189,10 @@ static inline bool nvme_mpath_queue_if_no_path(struct nvme_ns_head *head)
return true;
return false;
}
+static inline void nvme_free_ns_stat(struct nvme_ns *ns)
+{
+ free_percpu(ns->path_lat);
+}
#else
#define multipath false
static inline bool nvme_ctrl_use_ana(struct nvme_ctrl *ctrl)
@@ -1248,6 +1284,16 @@ static inline bool nvme_mpath_queue_if_no_path(struct nvme_ns_head *head)
{
return false;
}
+static inline void nvme_cancel_ns_latency_weight_work(struct nvme_ns *ns)
+{
+}
+static inline int nvme_mpath_alloc_ns_stat(struct nvme_ns *ns)
+{
+ return 0;
+}
+static inline void nvme_free_ns_stat(struct nvme_ns *ns)
+{
+}
#endif /* CONFIG_NVME_MULTIPATH */
#if defined(CONFIG_NVME_MULTIPATH) && defined(CONFIG_BLK_DEV_ZONED)
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 06/10] nvme: add generic debugfs support
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (4 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 05/10] nvme-multipath: add support for latency I/O policy Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 07/10] nvme-multipath: add debugfs attribute latency_ewma_shift Nilay Shroff
` (3 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
Add generic infrastructure for creating and managing debugfs files in
the NVMe module. This introduces helper APIs that allow NVMe drivers to
register and unregister debugfs entries, along with a reusable attribute
structure for defining new debugfs files.
The implementation uses seq_file interfaces to safely expose per-NS and
per-NS-head statistics, while supporting both simple show callbacks and
full seq_operations.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/Makefile | 2 +-
drivers/nvme/host/core.c | 3 +
drivers/nvme/host/debugfs.c | 141 ++++++++++++++++++++++++++++++++++
drivers/nvme/host/multipath.c | 2 +
drivers/nvme/host/nvme.h | 10 +++
5 files changed, 157 insertions(+), 1 deletion(-)
create mode 100644 drivers/nvme/host/debugfs.c
diff --git a/drivers/nvme/host/Makefile b/drivers/nvme/host/Makefile
index 67563a69f7dc..314d3439406b 100644
--- a/drivers/nvme/host/Makefile
+++ b/drivers/nvme/host/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_NVME_FC) += nvme-fc.o
obj-$(CONFIG_NVME_TCP) += nvme-tcp.o
obj-$(CONFIG_NVME_APPLE) += nvme-apple.o
-nvme-core-y += core.o ioctl.o sysfs.o pr.o
+nvme-core-y += core.o ioctl.o sysfs.o pr.o debugfs.o
nvme-core-$(CONFIG_NVME_VERBOSE_ERRORS) += constants.o
nvme-core-$(CONFIG_TRACING) += trace.o
nvme-core-$(CONFIG_NVME_MULTIPATH) += multipath.o
diff --git a/drivers/nvme/host/core.c b/drivers/nvme/host/core.c
index a480e33fd984..0fde693dd9b8 100644
--- a/drivers/nvme/host/core.c
+++ b/drivers/nvme/host/core.c
@@ -4306,6 +4306,8 @@ static void nvme_alloc_ns(struct nvme_ctrl *ctrl, struct nvme_ns_info *info)
if (device_add_disk(ctrl->device, ns->disk, nvme_ns_attr_groups))
goto out_cleanup_ns_from_list;
+ nvme_debugfs_register(ns->disk);
+
if (!nvme_ns_head_multipath(ns->head))
nvme_add_ns_cdev(ns);
@@ -4388,6 +4390,7 @@ static void nvme_ns_remove(struct nvme_ns *ns)
nvme_mpath_remove_sysfs_link(ns);
+ nvme_debugfs_unregister(ns->disk);
del_gendisk(ns->disk);
mutex_lock(&ns->ctrl->namespaces_lock);
diff --git a/drivers/nvme/host/debugfs.c b/drivers/nvme/host/debugfs.c
new file mode 100644
index 000000000000..04f8250fe607
--- /dev/null
+++ b/drivers/nvme/host/debugfs.c
@@ -0,0 +1,141 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 IBM Corporation
+ * Nilay Shroff <nilay@linux.ibm.com>
+ */
+
+#include <linux/debugfs.h>
+#include <linux/seq_file.h>
+
+#include "nvme.h"
+
+struct nvme_debugfs_attr {
+ const char *name;
+ umode_t mode;
+ int (*show)(void *data, struct seq_file *m);
+ ssize_t (*write)(void *data, const char __user *buf, size_t count,
+ loff_t *ppos);
+ const struct seq_operations *seq_ops;
+};
+
+struct nvme_debugfs_ctx {
+ void *data;
+ struct nvme_debugfs_attr *attr;
+ int srcu_idx;
+};
+
+static int nvme_debugfs_show(struct seq_file *m, void *v)
+{
+ struct nvme_debugfs_ctx *ctx = m->private;
+ void *data = ctx->data;
+ struct nvme_debugfs_attr *attr = ctx->attr;
+
+ return attr->show(data, m);
+}
+
+static int nvme_debugfs_open(struct inode *inode, struct file *file)
+{
+ void *data = inode->i_private;
+ struct nvme_debugfs_attr *attr = debugfs_get_aux(file);
+ struct nvme_debugfs_ctx *ctx;
+ struct seq_file *m;
+ int ret;
+
+ ctx = kzalloc_obj(struct nvme_debugfs_ctx);
+ return -ENOMEM;
+
+ ctx->data = data;
+ ctx->attr = attr;
+
+ if (attr->seq_ops) {
+ ret = seq_open(file, attr->seq_ops);
+ if (ret) {
+ kfree(ctx);
+ return ret;
+ }
+ m = file->private_data;
+ m->private = ctx;
+ return ret;
+ }
+
+ if (WARN_ON_ONCE(!attr->show)) {
+ kfree(ctx);
+ return -EPERM;
+ }
+
+ ret = single_open(file, nvme_debugfs_show, ctx);
+ if (ret)
+ kfree(ctx);
+
+ return ret;
+}
+
+static ssize_t nvme_debugfs_write(struct file *file, const char __user *buf,
+ size_t count, loff_t *ppos)
+{
+ struct seq_file *m = file->private_data;
+ struct nvme_debugfs_ctx *ctx = m->private;
+ struct nvme_debugfs_attr *attr = ctx->attr;
+
+ if (!attr->write)
+ return -EPERM;
+
+ return attr->write(ctx->data, buf, count, ppos);
+}
+
+static int nvme_debugfs_release(struct inode *inode, struct file *file)
+{
+ struct seq_file *m = file->private_data;
+ struct nvme_debugfs_ctx *ctx = m->private;
+ struct nvme_debugfs_attr *attr = ctx->attr;
+ int ret;
+
+ if (attr->seq_ops)
+ ret = seq_release(inode, file);
+ else
+ ret = single_release(inode, file);
+
+ kfree(ctx);
+ return ret;
+}
+
+static const struct file_operations nvme_debugfs_fops = {
+ .owner = THIS_MODULE,
+ .open = nvme_debugfs_open,
+ .read = seq_read,
+ .write = nvme_debugfs_write,
+ .llseek = seq_lseek,
+ .release = nvme_debugfs_release,
+};
+
+
+static const struct nvme_debugfs_attr nvme_mpath_debugfs_attrs[] = {
+ {},
+};
+
+static const struct nvme_debugfs_attr nvme_ns_debugfs_attrs[] = {
+ {},
+};
+
+static void nvme_debugfs_create_files(struct request_queue *q,
+ const struct nvme_debugfs_attr *attr, void *data)
+{
+ if (WARN_ON_ONCE(!q->debugfs_dir))
+ return;
+
+ for (; attr->name; attr++)
+ debugfs_create_file_aux(attr->name, attr->mode, q->debugfs_dir,
+ data, (void *)attr, &nvme_debugfs_fops);
+}
+
+void nvme_debugfs_register(struct gendisk *disk)
+{
+ const struct nvme_debugfs_attr *attr;
+
+ if (nvme_disk_is_ns_head(disk))
+ attr = nvme_mpath_debugfs_attrs;
+ else
+ attr = nvme_ns_debugfs_attrs;
+
+ nvme_debugfs_create_files(disk->queue, attr, disk->private_data);
+}
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 5501c22bd662..8d5f1e42d10a 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -1137,6 +1137,7 @@ static void nvme_remove_head(struct nvme_ns_head *head)
if (test_and_clear_bit(NVME_NSHEAD_CDEV_LIVE, &head->flags))
nvme_cdev_del(&head->cdev, &head->cdev_device);
+ nvme_debugfs_unregister(head->disk);
del_gendisk(head->disk);
}
nvme_put_ns_head(head);
@@ -1244,6 +1245,7 @@ static void nvme_mpath_set_live(struct nvme_ns *ns)
}
nvme_add_ns_head_cdev(head);
queue_work(nvme_wq, &head->partition_scan_work);
+ nvme_debugfs_register(head->disk);
}
nvme_mpath_add_sysfs_link(ns->head);
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index 850a5f6afca5..040307c6534a 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -1127,6 +1127,16 @@ static inline enum nvme_stat_group nvme_get_stat_group(struct request *req)
return __nvme_get_stat_group(req_op(req));
}
+void nvme_debugfs_register(struct gendisk *disk);
+static inline void nvme_debugfs_unregister(struct gendisk *disk)
+{
+ /*
+ * Nothing to do for now. When the request queue is unregistered,
+ * all files under q->debugfs_dir are recursively deleted.
+ * This is just a placeholder; the compiler will optimize it out.
+ */
+}
+
#ifdef CONFIG_NVME_MULTIPATH
static inline bool nvme_ctrl_use_ana(struct nvme_ctrl *ctrl)
{
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 07/10] nvme-multipath: add debugfs attribute latency_ewma_shift
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (5 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 06/10] nvme: add generic debugfs support Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 08/10] nvme-multipath: add debugfs attribute latency_batch_timeout Nilay Shroff
` (2 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
By default, the EWMA (Exponentially Weighted Moving Average) shift
value, used for storing latency samples for latency iopolicy, is set
to 3. The EWMA is calculated using the following formula:
ewma = (old * ((1 << ewma_shift) - 1) + new) >> ewma_shift;
The default value of 3 assigns ~87.5% weight to the existing EWMA value
and ~12.5% weight to the new latency sample. This provides a stable
average that smooths out short-term variations.
However, different workloads may require faster or slower adaptation to
changing conditions. This commit introduces a new debugfs attribute,
latency_ewma_shift, allowing users to tune the weighting factor.
For example:
- latency_ewma_shift = 2 => 75% old, 25% new
- latency_ewma_shift = 1 => 50% old, 50% new
- latency_ewma_shift = 0 => 0% old, 100% new
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/debugfs.c | 46 +++++++++++++++++++++++++++++++++++
drivers/nvme/host/multipath.c | 9 ++++---
drivers/nvme/host/nvme.h | 1 +
3 files changed, 52 insertions(+), 4 deletions(-)
diff --git a/drivers/nvme/host/debugfs.c b/drivers/nvme/host/debugfs.c
index 04f8250fe607..ca1a670bea78 100644
--- a/drivers/nvme/host/debugfs.c
+++ b/drivers/nvme/host/debugfs.c
@@ -108,8 +108,54 @@ static const struct file_operations nvme_debugfs_fops = {
.release = nvme_debugfs_release,
};
+#ifdef CONFIG_NVME_MULTIPATH
+static int nvme_latency_ewma_shift_show(void *data, struct seq_file *m)
+{
+ struct nvme_ns_head *head = data;
+
+ seq_printf(m, "%u\n", READ_ONCE(head->latency_ewma_shift));
+ return 0;
+}
+
+static ssize_t nvme_latency_ewma_shift_store(void *data,
+ const char __user *ubuf, size_t count, loff_t *ppos)
+{
+ struct nvme_ns_head *head = data;
+ char kbuf[8];
+ u32 res;
+ int ret;
+ size_t len;
+ char *arg;
+
+ len = min(sizeof(kbuf) - 1, count);
+
+ if (copy_from_user(kbuf, ubuf, len))
+ return -EFAULT;
+
+ kbuf[len] = '\0';
+ arg = strstrip(kbuf);
+
+ ret = kstrtou32(arg, 0, &res);
+ if (ret)
+ return ret;
+
+ /*
+ * Values greater than 8 are nonsensical, as they effectively assign
+ * zero weight to new samples.
+ */
+ if (res > 8)
+ return -EINVAL;
+
+ WRITE_ONCE(head->latency_ewma_shift, res);
+ return count;
+}
+#endif
static const struct nvme_debugfs_attr nvme_mpath_debugfs_attrs[] = {
+#ifdef CONFIG_NVME_MULTIPATH
+ {"latency_ewma_shift", 0600, nvme_latency_ewma_shift_show,
+ nvme_latency_ewma_shift_store},
+#endif
{},
};
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 8d5f1e42d10a..82a4f2326b79 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -293,10 +293,9 @@ static void nvme_mpath_weight_work(struct work_struct *weight_work)
* For instance, with EWMA_SHIFT = 3, this assigns 7/8 (~87.5 %) weight to
* the existing/old ewma and 1/8 (~12.5%) weight to the new sample.
*/
-static inline u64 calc_ewma_update(u64 old, u64 new)
+static inline u64 calc_ewma_update(u64 old, u64 new, u32 ewma_shift)
{
- return (old * ((1 << NVME_DEFAULT_LATENCY_EWMA_SHIFT) - 1)
- + new) >> NVME_DEFAULT_LATENCY_EWMA_SHIFT;
+ return (old * ((1 << ewma_shift) - 1) + new) >> ewma_shift;
}
static void nvme_mpath_add_sample(struct request *rq, struct nvme_ns *ns)
@@ -388,7 +387,8 @@ static void nvme_mpath_add_sample(struct request *rq, struct nvme_ns *ns)
if (unlikely(!stat->slat_ns))
WRITE_ONCE(stat->slat_ns, avg_lat_ns);
else {
- slat_ns = calc_ewma_update(stat->slat_ns, avg_lat_ns);
+ slat_ns = calc_ewma_update(stat->slat_ns, avg_lat_ns,
+ READ_ONCE(head->latency_ewma_shift));
WRITE_ONCE(stat->slat_ns, slat_ns);
}
@@ -1170,6 +1170,7 @@ int nvme_mpath_alloc_disk(struct nvme_ctrl *ctrl, struct nvme_ns_head *head)
INIT_WORK(&head->requeue_work, nvme_requeue_work);
INIT_WORK(&head->partition_scan_work, nvme_partition_scan_work);
INIT_DELAYED_WORK(&head->remove_work, nvme_remove_head_work);
+ head->latency_ewma_shift = NVME_DEFAULT_LATENCY_EWMA_SHIFT;
/*
* If "multipath_always_on" is enabled, a multipath node is added
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index 040307c6534a..55e4d1344e5e 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -626,6 +626,7 @@ struct nvme_ns_head {
atomic_long_t io_requeue_no_usable_path_count;
atomic_long_t io_fail_no_available_path_count;
struct nvme_ns * __percpu *latency_path;
+ u32 latency_ewma_shift;
#define NVME_NSHEAD_DISK_LIVE 0
#define NVME_NSHEAD_QUEUE_IF_NO_PATH 1
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 08/10] nvme-multipath: add debugfs attribute latency_batch_timeout
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (6 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 07/10] nvme-multipath: add debugfs attribute latency_ewma_shift Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 09/10] nvme-multipath: add debugfs attribute latency_stat Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 10/10] nvme-multipath: add documentation for latency I/O policy Nilay Shroff
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
By default, the latency I/O policy accumulates latency samples over a
15-second window. When this window expires, the driver computes the
average latency and updates the smoothed (EWMA) latency value. The
path weight is then recalculated based on this data.
A 15-second window provides a good balance for most workloads, as it
helps smooth out transient latency spikes and produces a more stable
path weight profile. However, some workloads may benefit from faster
or slower adaptation to changing latency conditions.
This commit introduces a new debugfs attribute, latency_batch_timeout,
which allows users to configure the latency batch window and thus path
weight calculation interval based on their workload requirements.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/debugfs.c | 37 +++++++++++++++++++++++++++++++++++
drivers/nvme/host/multipath.c | 8 ++++++--
drivers/nvme/host/nvme.h | 1 +
3 files changed, 44 insertions(+), 2 deletions(-)
diff --git a/drivers/nvme/host/debugfs.c b/drivers/nvme/host/debugfs.c
index ca1a670bea78..b8850d37edce 100644
--- a/drivers/nvme/host/debugfs.c
+++ b/drivers/nvme/host/debugfs.c
@@ -149,12 +149,49 @@ static ssize_t nvme_latency_ewma_shift_store(void *data,
WRITE_ONCE(head->latency_ewma_shift, res);
return count;
}
+
+static int nvme_latency_batch_timeout_show(void *data, struct seq_file *m)
+{
+ struct nvme_ns_head *head = data;
+
+ seq_printf(m, "%llu\n",
+ div_u64(READ_ONCE(head->latency_batch_timeout), NSEC_PER_SEC));
+ return 0;
+}
+
+static ssize_t nvme_latency_batch_timeout_store(void *data,
+ const char __user *ubuf, size_t count, loff_t *ppos)
+{
+ struct nvme_ns_head *head = data;
+ char kbuf[8];
+ u32 res;
+ int ret;
+ size_t len;
+ char *arg;
+
+ len = min(sizeof(kbuf) - 1, count);
+
+ if (copy_from_user(kbuf, ubuf, len))
+ return -EFAULT;
+
+ kbuf[len] = '\0';
+ arg = strstrip(kbuf);
+
+ ret = kstrtou32(arg, 0, &res);
+ if (ret)
+ return ret;
+
+ WRITE_ONCE(head->latency_batch_timeout, res * NSEC_PER_SEC);
+ return count;
+}
#endif
static const struct nvme_debugfs_attr nvme_mpath_debugfs_attrs[] = {
#ifdef CONFIG_NVME_MULTIPATH
{"latency_ewma_shift", 0600, nvme_latency_ewma_shift_show,
nvme_latency_ewma_shift_store},
+ {"latency_batch_timeout", 0600, nvme_latency_batch_timeout_show,
+ nvme_latency_batch_timeout_store},
#endif
{},
};
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 82a4f2326b79..a4cca27e8eb5 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -362,8 +362,11 @@ static void nvme_mpath_add_sample(struct request *rq, struct nvme_ns *ns)
stat->batch_count++;
stat->nr_samples++;
- if (now > stat->last_batch_ts && ((now - stat->last_batch_ts) >=
- NVME_DEFAULT_LATENCY_BATCH_TIMEOUT)) {
+ if (now > stat->last_batch_ts) {
+ u64 timeout = READ_ONCE(head->latency_batch_timeout);
+
+ if ((now - stat->last_batch_ts) < timeout)
+ return;
/*
* Find simple average latency for the last epoch (~15 sec
@@ -1171,6 +1174,7 @@ int nvme_mpath_alloc_disk(struct nvme_ctrl *ctrl, struct nvme_ns_head *head)
INIT_WORK(&head->partition_scan_work, nvme_partition_scan_work);
INIT_DELAYED_WORK(&head->remove_work, nvme_remove_head_work);
head->latency_ewma_shift = NVME_DEFAULT_LATENCY_EWMA_SHIFT;
+ head->latency_batch_timeout = NVME_DEFAULT_LATENCY_BATCH_TIMEOUT;
/*
* If "multipath_always_on" is enabled, a multipath node is added
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index 55e4d1344e5e..133a22611597 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -627,6 +627,7 @@ struct nvme_ns_head {
atomic_long_t io_fail_no_available_path_count;
struct nvme_ns * __percpu *latency_path;
u32 latency_ewma_shift;
+ u64 latency_batch_timeout;
#define NVME_NSHEAD_DISK_LIVE 0
#define NVME_NSHEAD_QUEUE_IF_NO_PATH 1
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 09/10] nvme-multipath: add debugfs attribute latency_stat
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (7 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 08/10] nvme-multipath: add debugfs attribute latency_batch_timeout Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 10/10] nvme-multipath: add documentation for latency I/O policy Nilay Shroff
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
This commit introduces a new debugfs attribute, "latency_stat", under
both per-path and head debugfs directories (defined under /sys/kernel/
debug/block/). This attribute provides visibility into the internal
state of the latency I/O policy to aid in debugging and performance
analysis.
For per-path entries, "latency_stat" reports the corresponding path
statistics such as I/O weight, selection count, processed samples, and
ignored samples.
For head entries, it reports per-CPU statistics for each reachable path,
including I/O weight, path score, smoothed (EWMA) latency, selection
count, processed samples, and ignored samples.
These additions enhance observability of the I/O path selection behavior
and help diagnose imbalance or instability in multipath performance.
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
drivers/nvme/host/debugfs.c | 123 ++++++++++++++++++++++++++++++++++++
1 file changed, 123 insertions(+)
diff --git a/drivers/nvme/host/debugfs.c b/drivers/nvme/host/debugfs.c
index b8850d37edce..b3496951f0b1 100644
--- a/drivers/nvme/host/debugfs.c
+++ b/drivers/nvme/host/debugfs.c
@@ -184,6 +184,125 @@ static ssize_t nvme_latency_batch_timeout_store(void *data,
WRITE_ONCE(head->latency_batch_timeout, res * NSEC_PER_SEC);
return count;
}
+
+#define TO_CTX(m) ((struct nvme_debugfs_ctx *)m->private)
+#define TO_NS_HEAD(m) ((struct nvme_ns_head *)(TO_CTX(m)->data))
+
+static void *nvme_mpath_latency_stat_start(struct seq_file *m, loff_t *pos)
+ __acquires_shared(&TO_NS_HEAD(m)->srcu)
+{
+ struct nvme_ns *ns;
+ struct nvme_debugfs_ctx *ctx = m->private;
+ struct nvme_ns_head *head = ctx->data;
+ loff_t n = *pos;
+
+ /* Remember srcu index, so we can unlock later. */
+ ctx->srcu_idx = srcu_read_lock(&head->srcu);
+ ns = list_first_or_null_rcu(&head->list, struct nvme_ns, siblings);
+
+ while (n && ns) {
+ ns = list_next_or_null_rcu(&head->list, &ns->siblings,
+ struct nvme_ns, siblings);
+ n--;
+ }
+
+ return ns;
+}
+
+static void *nvme_mpath_latency_stat_next(struct seq_file *m, void *v,
+ loff_t *pos)
+{
+ struct nvme_ns *ns = v;
+ struct nvme_debugfs_ctx *ctx = m->private;
+ struct nvme_ns_head *head = ctx->data;
+
+ (*pos)++;
+
+ return list_next_or_null_rcu(&head->list, &ns->siblings,
+ struct nvme_ns, siblings);
+}
+
+static void nvme_mpath_latency_stat_stop(struct seq_file *m, void *v)
+ __releases_shared(&TO_NS_HEAD(m)->srcu)
+{
+ struct nvme_debugfs_ctx *ctx = m->private;
+ struct nvme_ns_head *head = ctx->data;
+
+ srcu_read_unlock(&head->srcu, ctx->srcu_idx);
+}
+
+static int nvme_mpath_latency_stat_show(struct seq_file *m, void *v)
+{
+ int i, cpu;
+ struct nvme_path_lat_stat *stat;
+ struct nvme_ns *ns = v;
+
+ seq_printf(m, "%s:\n", ns->disk->disk_name);
+ for_each_online_cpu(cpu) {
+ seq_printf(m, "cpu %d : ", cpu);
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[i].stat;
+ seq_printf(m, "%u %u %llu %llu %llu %llu %llu ",
+ stat->weight, stat->credit, stat->score,
+ stat->slat_ns, stat->sel,
+ stat->nr_samples, stat->nr_ignored);
+ }
+ seq_putc(m, '\n');
+ }
+ return 0;
+}
+
+static const struct seq_operations nvme_mpath_latency_stat_seq_ops = {
+ .start = nvme_mpath_latency_stat_start,
+ .next = nvme_mpath_latency_stat_next,
+ .stop = nvme_mpath_latency_stat_stop,
+ .show = nvme_mpath_latency_stat_show
+};
+
+static void nvme_latency_stat_read_all(struct nvme_ns *ns,
+ struct nvme_path_lat_stat *batch)
+{
+ int i, cpu;
+ u32 ncpu[NVME_NUM_STAT_GROUPS] = {0};
+ struct nvme_path_lat_stat *stat;
+
+ for_each_online_cpu(cpu) {
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ stat = &per_cpu_ptr(ns->path_lat, cpu)[i].stat;
+ batch[i].sel += stat->sel;
+ batch[i].nr_samples += stat->nr_samples;
+ batch[i].nr_ignored += stat->nr_ignored;
+ batch[i].weight += stat->weight;
+ if (stat->weight)
+ ncpu[i]++;
+ }
+ }
+
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ if (!ncpu[i])
+ continue;
+ batch[i].weight = DIV_U64_ROUND_CLOSEST(batch[i].weight,
+ ncpu[i]);
+ }
+}
+
+static int nvme_ns_latency_stat_show(void *data, struct seq_file *m)
+{
+ int i;
+ struct nvme_path_lat_stat stat[NVME_NUM_STAT_GROUPS] = {0};
+ struct nvme_ns *ns = (struct nvme_ns *)data;
+
+ if (!ns->head->disk)
+ return 0;
+
+ nvme_latency_stat_read_all(ns, stat);
+ for (i = 0; i < NVME_NUM_STAT_GROUPS; i++) {
+ seq_printf(m, "%u %llu %llu %llu ",
+ stat[i].weight, stat[i].sel,
+ stat[i].nr_samples, stat[i].nr_ignored);
+ }
+ return 0;
+}
#endif
static const struct nvme_debugfs_attr nvme_mpath_debugfs_attrs[] = {
@@ -192,11 +311,15 @@ static const struct nvme_debugfs_attr nvme_mpath_debugfs_attrs[] = {
nvme_latency_ewma_shift_store},
{"latency_batch_timeout", 0600, nvme_latency_batch_timeout_show,
nvme_latency_batch_timeout_store},
+ {"latency_stat", 0400, .seq_ops = &nvme_mpath_latency_stat_seq_ops},
#endif
{},
};
static const struct nvme_debugfs_attr nvme_ns_debugfs_attrs[] = {
+#ifdef CONFIG_NVME_MULTIPATH
+ {"latency_stat", 0400, nvme_ns_latency_stat_show},
+#endif
{},
};
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v8 10/10] nvme-multipath: add documentation for latency I/O policy
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
` (8 preceding siblings ...)
2026-08-15 17:34 ` [PATCH v8 09/10] nvme-multipath: add debugfs attribute latency_stat Nilay Shroff
@ 2026-08-15 17:34 ` Nilay Shroff
9 siblings, 0 replies; 11+ messages in thread
From: Nilay Shroff @ 2026-08-15 17:34 UTC (permalink / raw)
To: linux-nvme
Cc: hare, kbusch, hch, sagi, dwagner, kanie, jmeneghi, randyj,
martin.petersen, john.g.garry, gjoyce
Update the nvme-multipath documentation to describe the latency I/O
policy, its behavior, and when it is suitable for use.
Suggested-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Guixin Liu <kanie@linux.alibaba.com>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
Documentation/admin-guide/nvme-multipath.rst | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/Documentation/admin-guide/nvme-multipath.rst b/Documentation/admin-guide/nvme-multipath.rst
index 97ca1ccef459..41a8054638ff 100644
--- a/Documentation/admin-guide/nvme-multipath.rst
+++ b/Documentation/admin-guide/nvme-multipath.rst
@@ -70,3 +70,22 @@ When to use the queue-depth policy:
1. High load with small I/Os: Effectively balances load across paths when
the load is high, and I/O operations consist of small, relatively
fixed-sized requests.
+
+Latency
+--------
+
+The latency policy manages I/O requests based on path latency. It periodically
+calculates a weight for each path and distributes I/O accordingly. Paths with
+higher latency receive lower weights, resulting in fewer I/O requests being sent
+to them, while paths with lower latency handle a proportionally larger share of
+the I/O load.
+
+When to use the latency policy
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+1. Homogeneous Path Performance: Utilizes all available paths efficiently when
+ their performance characteristics (e.g., latency, bandwidth) are similar.
+
+2. Heterogeneous Path Performance: Dynamically distributes I/O based on per-path
+ performance characteristics. Paths with lower latency receive a higher share
+ of I/O compared to those with higher latency.
--
2.53.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-08-15 17:36 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-15 17:34 [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 01/10] block: expose blk_stat_{enable,disable}_accounting() to drivers Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 02/10] block: record I/O request start time for passthru request Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 03/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 04/10] nvme-multipath: pass I/O type to nvme_find_path() Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 05/10] nvme-multipath: add support for latency I/O policy Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 06/10] nvme: add generic debugfs support Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 07/10] nvme-multipath: add debugfs attribute latency_ewma_shift Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 08/10] nvme-multipath: add debugfs attribute latency_batch_timeout Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 09/10] nvme-multipath: add debugfs attribute latency_stat Nilay Shroff
2026-08-15 17:34 ` [PATCH v8 10/10] nvme-multipath: add documentation for latency I/O policy Nilay Shroff
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.