From: Peter Xu <peterx@redhat.com>
To: Ming Lei <ming.lei@redhat.com>
Cc: Thomas Gleixner <tglx@linutronix.de>,
Juri Lelli <juri.lelli@redhat.com>, Ming Lei <minlei@redhat.com>,
Linux Kernel Mailing List <linux-kernel@vger.kernel.org>
Subject: Re: Kernel-managed IRQ affinity (cont)
Date: Thu, 19 Dec 2019 09:32:14 -0500 [thread overview]
Message-ID: <20191219143214.GA50561@xz-x1> (raw)
In-Reply-To: <20191219082819.GB15731@ming.t460p>
On Thu, Dec 19, 2019 at 04:28:19PM +0800, Ming Lei wrote:
> Hi Peter,
Hi, Ming,
>
> On Mon, Dec 16, 2019 at 02:57:12PM -0500, Peter Xu wrote:
> > Hi, Thomas,
> >
> > (Sorry I must have lost the discussion during an email migration, so
> > I'll start with a new one)
> >
> > This is a continued discussion of previous one on kernel managed IRQ
> > affinity [1]. I think at that time the conclusion is that we don't
> > have a usage scenario to change current policy [2]. However recently
> > I noticed that it is probably a very fundamental requirement for some
> > real-time scenarios, even when there's no multi-queue involved.
> >
> > In my test case, it was a very common realtime guest with 10 vcpus,
> > 0-1 are housekeeping vcpus, 2-9 are realtime vcpus. The guest has one
> > virtio-blk device as boot disk. With a distribution very close to
> > latest upstream, we can observe high spikes, probably due to the IRQs.
> >
> > To guarantee realtime responsiveness, we need to make sure the IRQs
> > will be managable, say, when I run a real-time workload on vcpu9, we
> > should be able to move all the IRQs from vcpu9 to the other vcpus
> > (most probably vcpu0 and vcpu1). However with the kernel managed IRQs
> > we can't echo to /proc/irq/N/smp_affinity. Here, vcpu9 gets IRQ 38
> > from the virtio-blk device:
> >
> > # cat /proc/interrupts | grep -w 38
> > 38: 0 0 0 0 0 0 0 0 0 15206 PCI-MSI 2621441-edge virtio2-req.0
> > # cat /proc/irq/38/smp_affinity
> > 3ff
> > # cat /proc/irq/38/effective_affinity
> > 200
> >
> > Meanwhile, I don't think there's anything special for VMs, so this
> > issue should exist even for hosts as long as the IRQ is managed in the
> > same way here as the virtio-blk device.
> >
> > As Ming has mentioned in previous discussions [3], I think it would be
> > at least good if the kernel IRQ system can respect "irqaffinity=" when
> > assigning IRQs to the cores. Currently it's not. What would you
> > suggest in this case? Do you think this is a valid user scenario?
> >
> > Thanks,
> >
> > [1] https://lkml.org/lkml/2019/3/18/15
> > [2] https://lkml.org/lkml/2019/3/25/562
> > [3] https://lkml.org/lkml/2019/3/25/308
>
> The following patch supposes to implementation the requirement for you,
> can you test it by passing 'isolcpus=managed_irq,X-Y'?
I really appreciate your patch! I'll keep this version, while before
I start to test it...
>
> With this kind of change, you can't run any IO from any isolated
> CPU core, otherwise, unpredictable error may be triggered, either oops or
> IO hang.
... I'm not sure whether this can be acceptable for a production
environment.
In our case, the IRQ should come from virtio-blk which is the root
disk, so I assume even the RT core could use it at least when loading
the executable into RAM. So...
>
> Another conservative approach is to only select effective CPU from
> non-isolated cpus, however, the assigned CPUs may not be balanced among
> interrupt vectors. But it is safer, since the system still works even if
> someone submits IO from any isolated cpu core.
... this one seems to be more appealing at least to me.
Thanks,
--
Peter Xu
next prev parent reply other threads:[~2019-12-19 14:32 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2019-12-16 19:57 Kernel-managed IRQ affinity (cont) Peter Xu
2019-12-19 8:28 ` Ming Lei
2019-12-19 14:32 ` Peter Xu [this message]
2019-12-19 16:11 ` Ming Lei
2019-12-19 18:09 ` Peter Xu
2019-12-23 19:18 ` Peter Xu
2020-01-09 20:02 ` Thomas Gleixner
2020-01-10 1:28 ` Ming Lei
2020-01-10 19:43 ` Thomas Gleixner
2020-01-11 2:48 ` Ming Lei
2020-01-14 13:45 ` Thomas Gleixner
2020-01-14 23:38 ` Ming Lei
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20191219143214.GA50561@xz-x1 \
--to=peterx@redhat.com \
--cc=juri.lelli@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=ming.lei@redhat.com \
--cc=minlei@redhat.com \
--cc=tglx@linutronix.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.