Linux block layer
 help / color / mirror / Atom feed
From: "Ionut Nechita (Wind River)" <ionut.nechita@windriver.com>
To: atomlin@atomlin.com
Cc: axboe@kernel.dk, bigeasy@linutronix.de, corbet@lwn.net,
	frederic@kernel.org, ionut.nechita@windriver.com,
	linux-block@vger.kernel.org, linux-doc@vger.kernel.org,
	linux-kernel@vger.kernel.org, linux-scsi@vger.kernel.org,
	marco.crivellari@suse.com, mingo@redhat.com, mkp@kernel.org,
	neelx@suse.com, peterz@infradead.org, tglx@kernel.org,
	vincent.guittot@linaro.org
Subject: Re: [PATCH v16 0/9] blk: honor isolcpus configuration
Date: Tue,  6 Oct 2026 16:57:44 +0300	[thread overview]
Message-ID: <20261006135749.348637-1-ionut.nechita@windriver.com> (raw)
In-Reply-To: <20261005152449.468696-1-ionut.nechita@windriver.com>

Hi Aaron,

Follow-up with the controlled data I promised, and a correction. I was
wrong to point at your series in my first mail. After a clean A/B the
control-plane stall is a shared blk-mq queue-contention problem on a
shared disk, reproducible on *plain* managed_irq. Your series does not
introduce it -- but it does remove the per-CPU-queue separation that
otherwise keeps it from happening, so I think there is still a real
I/O-QoS question for you. Details below, each claim backed by a debugfs
capture.

Setup
-----
  Kernel:   6.18.15-rt (PREEMPT_RT), v16 applied
  CPU:      2x Intel Xeon Gold 6338N, 128 CPUs, 2 NUMA nodes
            112 isolated, 16 housekeeping (8 carrying managed IRQs)
  Storage:  Dell PERC / MegaRAID (megaraid_sas) fronting a SATA SSD
            (Samsung PM893), passthrough. etcd and the fio target share
            the SAME physical disk (sda4 / cgts-vg).
  Workload: one fio pod, libaio iodepth=256 numjobs=4 direct=1,
            SCHED_OTHER (NOT RT -- RT is not required to reproduce),
            pinned to one isolated CPU (CPU 2) via the StarlingX
            isolcpus device-plugin label (verified with ps -eLo psr).
  Method:   per-hctx in-flight snapshot from
            /sys/kernel/debug/block/sda/hctx*/busy during the run,
            cross-referenced with each queue's mq/*/cpu_list and with
            etcd's CPU set (ps psr).

The decisive observable is not "etcd alive/dead" but which single
blk-mq queue the isolated pod's I/O lands on, and whether one of etcd's
CPUs is mapped to that same queue.

The clean A/B (identical kernel = plain managed_irq, identical fio 4k
randread, identical nr_requests=5089, fio always on isolated CPU 2;
only the device's blk-mq queue count differs)
----------------------------------------------------------------------

  (B) default MSI-X -> 120 blk-mq queues, ~1 CPU per queue
      All ~35 in-flight land on hctx61, a queue mapped to the isolated
      submitter and NOT used by any etcd CPU.
      -> no collision. etcd healthy. Device load ~82k IOPS 4k,
         sda aqu-sz ~1020, r_await ~12.3 ms.

  (C) megaraid_sas.msix_vectors=16 -> 15 blk-mq queues, ~8 CPUs each
      All ~33 in-flight land on hctx7, mapped to CPUs {0,4,64,68}.
      etcd runs on CPU0 and CPU68 -> same queue.
      -> collision. etcd apply 2-7 s, "request timed out",
         "context deadline exceeded", /registry/health failing.
         Control plane down -- on plain managed_irq, WITHOUT your series.

Same kernel, same disk, same fio, same depth; the only thing that
changed is whether there were enough hardware queues for the isolated
CPU to get one that etcd does not also use. That is the whole bug.

Supporting data point on the strict kernel
------------------------------------------
I also have one capture on the managed_irq_strict image (earlier run,
so not parameter-matched to B/C: it was a 64 KB read profile, and
/sys/block/sda/queue/nr_requests was already capped to 64 -- that cap
was my doing, the only way I could keep etcd / kube-apiserver alive
enough for the node to stay up and for me to collect the capture at
all; the default 5089 took the control plane down hard). It is still
informative directionally: with 0006 active, sda is confined to 16
blk-mq queues on the HK CPUs, fio on isolated CPU 2 funnels all ~30
in-flight onto hctx0 == CPU1, and CPU1 is an etcd CPU -> even with that
nr_requests=64 cap in place, etcd apply still blows out to multi-second
with context-deadline / health-probe failures. So strict reproduces the
same collision by construction, and the nr_requests knob only partially
masks it there.

Conclusion
----------
The failure is not strict-vs-non-strict. It is whether the isolated
pod's submission queue coincides with a queue an etcd CPU also uses:

  - On a device with enough hardware queues, plain managed_irq gives
    the isolated CPU its own blk-mq queue, disjoint from etcd's, and
    the heavy pod I/O does not head-of-line-block etcd (B).
  - Shrink the queue count until CPUs must share a queue -- via a low
    MSI-X count (C), or via your 0006 "use HK CPUs only" confinement
    (strict) -- and the isolated pod's deep queue lands on a queue etcd
    uses, and etcd starves.

So this is pre-existing shared-queue contention on a shared slow disk,
not a bug your series introduces. I retract the implication in my
previous mail. (It is also device-level, not CPU: during the stall 7 of
8 IRQ-HK cores are idle and the backlog is pure block-layer queueing,
aqu-sz ~1020 at ~70% util.)

Where I think the series is still relevant
------------------------------------------
0006 confines blk-mq to the HK CPUs unconditionally. On a
well-provisioned device (B, 120 queues) that is exactly the config that
*removes* the per-CPU-queue separation which was keeping isolated bulk
I/O off etcd's queues. In other words, strict turns a
"only on queue-starved devices" problem into "always, because every
isolated submission is forced onto the HK queue set the control plane
also lives on." That is a deliberate and correct part of the design --
isolated CPUs must not host the queues -- but it makes an I/O-QoS gap
unavoidable rather than incidental.

Mitigation
----------
Capping /sys/block/sda/queue/nr_requests on this device (megaraid_sas
with the SATA SSD exposed transparently / passthrough, e.g. 5089 -> a
small value in the 32-64 range -- lower for a slower disk) shortens the
head-of-line backlog and helps, but note it is weaker under strict/low-
queue configs: when all submissions pile on one etcd-shared queue, even
a depth of 32-64 on that single queue still hurts etcd. For production I'd
prefer cgroup v2 io.max / io.latency to protect control-plane I/O, or a
dedicated device for etcd.

Questions
---------
  1. When 0006 forces all isolated-pod block I/O onto the HK queue set,
     do you consider protecting co-located latency-critical services
     (etcd/kube-apiserver) from that I/O an operator concern
     (io.latency/io.weight, dedicated device), or is there a case for
     the series to leave a per-CPU submission path / reserved HK queue
     so bulk isolated I/O cannot head-of-line-block them?
  2. Is there an assumed minimum HK queue count relative to the
     aggregate block I/O of isolated pods? 8 IRQ-HK vs 112 isolated on
     one shared disk makes the HK queue set trivial to saturate on
     depth even while it is CPU-idle.

Raw debugfs hctx snapshots + iostat/etcd logs available on request.
Net: not a v16 regression; a shared-disk I/O-QoS gap that v16 makes
unavoidable by design, which may or may not be something you want to
address in the series rather than leave to operators.

Thanks, and sorry for the premature pointer in the first mail,
Ionut

      reply	other threads:[~2026-10-06 13:59 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 16:42 [PATCH v16 0/9] blk: honor isolcpus configuration Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 1/9] scsi: aacraid: use block layer helpers to calculate num of queues Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 2/9] lib/group_cpus: remove dead !SMP code Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 3/9] lib/group_cpus: Add group_mask_cpus_evenly() Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 4/9] sched/isolation: Prevent out-of-bounds read in isolcpus= boot parameter parser Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 5/9] isolation: Introduce managed_irq_strict isolcpus type Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 6/9] blk-mq: use hk cpus only when isolcpus=managed_irq_strict is enabled Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 7/9] blk-mq: prevent offlining hk CPUs with associated online isolated CPUs Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 8/9] genirq/affinity: Restrict managed IRQ affinity to housekeeping CPUs Aaron Tomlin
2026-09-10 16:42 ` [PATCH v16 9/9] docs: add managed_irq_strict flag to isolcpus Aaron Tomlin
2026-09-10 18:26 ` [PATCH v16 0/9] blk: honor isolcpus configuration Aaron Tomlin
2026-09-17 15:44 ` Ionut Nechita (Wind River)
2026-09-21 13:53   ` Ionut Nechita
2026-09-28 15:53     ` Aaron Tomlin
2026-10-05  9:34       ` [PATCH] genirq/affinity: confine reserved (pre/post) vectors to housekeeping CPUs Ionut Nechita (Wind River)
2026-10-05 15:24       ` [PATCH v16 0/9] blk: honor isolcpus configuration Ionut Nechita (Wind River)
2026-10-06 13:57         ` Ionut Nechita (Wind River) [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261006135749.348637-1-ionut.nechita@windriver.com \
    --to=ionut.nechita@windriver.com \
    --cc=atomlin@atomlin.com \
    --cc=axboe@kernel.dk \
    --cc=bigeasy@linutronix.de \
    --cc=corbet@lwn.net \
    --cc=frederic@kernel.org \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-scsi@vger.kernel.org \
    --cc=marco.crivellari@suse.com \
    --cc=mingo@redhat.com \
    --cc=mkp@kernel.org \
    --cc=neelx@suse.com \
    --cc=peterz@infradead.org \
    --cc=tglx@kernel.org \
    --cc=vincent.guittot@linaro.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox