From: Alistair Francis <Alistair.Francis@wdc.com>
To: "mani@kernel.org" <mani@kernel.org>,
"cassel@kernel.org" <cassel@kernel.org>,
"dlemoal@kernel.org" <dlemoal@kernel.org>
Cc: "alistair23@gmail.com" <alistair23@gmail.com>,
"jasowangio@gmail.com" <jasowangio@gmail.com>,
"xuanzhuo@linux.alibaba.com" <xuanzhuo@linux.alibaba.com>,
"virtualization@lists.linux.dev" <virtualization@lists.linux.dev>,
"den@valinux.co.jp" <den@valinux.co.jp>,
"linux-scsi@vger.kernel.org" <linux-scsi@vger.kernel.org>,
"eperezma@redhat.com" <eperezma@redhat.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"Frank.Li@kernel.org" <Frank.Li@kernel.org>,
"mie@igel.co.jp" <mie@igel.co.jp>,
"mkp@kernel.org" <mkp@kernel.org>,
"mst@redhat.com" <mst@redhat.com>,
"James.Bottomley@hansenpartnership.com"
<James.Bottomley@hansenpartnership.com>,
"linux-pci@vger.kernel.org" <linux-pci@vger.kernel.org>,
"kishon@kernel.org" <kishon@kernel.org>
Subject: Re: [PATCH 0/2] scsi: Initial commit of VirtIO PCIe Endpoint Driver
Date: Fri, 4 Sep 2026 03:19:27 +0000 [thread overview]
Message-ID: <2daeb43d582f0d960370fcd415c3b5bd3d94c67c.camel@wdc.com> (raw)
In-Reply-To: <795ac93a-9377-4503-9dc4-58f21c815f71@kernel.org>
On Fri, 2026-09-04 at 09:12 +0900, Damien Le Moal wrote:
> On 9/3/26 20:51, Manivannan Sadhasivam wrote:
> > On Tue, Sep 01, 2026 at 10:35:49AM +0200, Niklas Cassel wrote:
> > > On Tue, Sep 01, 2026 at 11:46:48AM +1000,
> > > alistair23@gmail.com wrote:
> > > > From: Alistair Francis <alistair.francis@wdc.com>
> > > >
> > > > This series adds a VirtIO SCSI endpoint built on top of the
> > > > VirtIO PCIe
> > > > endpoint. This is a similar approach to the NVMe PCIe Endpoint
> > > > (drivers/nvme/target/pci-epf.c) but for SCSI.
> > > >
> > > > This does end up being somewhat similar to the pci-epf.c code,
> > > > but
> > > > re-written for SCSI.
> > > >
> > > > This approach allows a PCIe Endpoint device (tested on a
> > > > radxa-rock5b) to setup what appears to be a SCSI device, using
> > > > an
> > > > existing SCSI backend (tested using scsi_debug).
> > > >
> > > > At this point a host can connect over PCIe, ensure virtio_pci
> > > > and
> > > > virtio_scsi is loaded and on PCIe rescan will see a scsi
> > > > device.
> > > >
> > > > There are a few pain points with this approach though:
> > > > 1. We have to use the Legacy SCSI VirtIO driver. This is
> > > > because the
> > > > Raxda Rock5b (and AFAIK all PCIe Endpoint hardware) can't
> > > > add
> > > > capabilities. So we can't advertise the VirtIO Common
> > > > configuration
> > > > capability, which means we can't be a modern VirtIO SCSI
> > > > device.
> > > >
> > > > This is unfortunate, but there doesn't seem to be any way
> > > > around
> > > > this, at least with the current hardware.
> > > >
> > > > 1.2. Legacy virtio devices only have 32 feature bits and
> > > > therefore can't
> > > > set the VIRTIO_F_ACCESS_PLATFORM (bit 33) feature. This
> > > > means the
> > > > vring_use_map_api() function will return false.
> > > >
> > > > Currently Linux endpoint devices use the legacy virtio
> > > > interface as
> > > > they aren't able to advertise the Common configuration
> > > > capability.
> > > > As most PCI endpoint capable PCIe controllers do not allow
> > > > modifying the
> > > > capability list, and thus are unable to advertise the
> > > > Common configuration
> > > > capability. This means the device's inbound TLPs fault on
> > > > the host
> > > > SMMU because the vring descriptors carry raw physical
> > > > addresses.
> > > >
> > > > This series adds a quirk that forces a subset of legacy
> > > > virtio devices
> > > > to use the DMA Map API (vring_use_map_api() will return
> > > > true),
> > > > which fixes this issue.
> > > >
> > > > It's unideal that we have to hard code a quirk to basically
> > > > just
> > > > advertise the VIRTIO_F_ACCESS_PLATFORM feature, but (see 1)
> > > > as we
> > > > are stuck with legacy virtio devices there isn't much else
> > > > we can do.
> > > >
> > > > 2. We have to pin scsit_pci_epf_poll_cfg_thread() on a CPU in
> > > > order to
> > > > respond fast enough to the host. This means we effectivly
> > > > burn a CPU
> > > > to read and write some values. But as there are no
> > > > intterupts
> > > > generated on these events and we need to be very quick
> > > > there isn't
> > > > another option.
> > > >
> >
> > Most of these pain points will go away if you use virtio-msg [1]
> > transport
> > instead of the virtio-pci transport. Using the virtio-pci transport
> > on a real
> > PCIe device without a way to trap and emulate the config space
> > requests will
> > always be racy.
>
> There is nothing inherently racy about the config space. It is about
> the fact
> that most PCI endpoint controllers:
> 1) Do not raise an interrupt when PCI BARs or config space is written
> by the
> host RC, and
> 2) All PCI endpoint controllers that Linux supports do not allow
> drivers to
> create extended capabilities in the config space that can then be
> emulated in
> the endpoint driver (enabling that would require 1 to be supported,
> obviously).
>
> (2) can be delt with quirks. Not great, but simple enough. And in
> this case, we
> need it more because of the virtio-pci specs, which are not great to
> start with.
>
> And for (1), the only real problem that causes is that an endpoint
> driver needs
> to poll PCI BARs/submission queues to see if the host issued
> commands. Again not
> great, but that works just fine. Alistair's point about burning a CPU
> doing that
> is simply so that we can reduce command latency and get good enough
> performance.
It is actually racy. If we don't burn a CPU to check we end up racing,
with the host as we are too slow to update the config space.
>
> We went through all of that already with the NVMe PCI endpoint. Works
> well
> enough and does what is intended, which is the same here for the
> virtio-scsi
> endpoint driver: create a platform where one can emulate a SCSI
> device to
> experiment with new features etc. This is all intended as a
> development/test
> tool, not for production use.
>
> I do not know virtio-msg. First time I hear about it. And I am not
> sure if there
> is a standard way of exposing a SCSI host through that.
virtio-msg-amp does seem promising. I'll dig into it a bit more and
keep an eye on it.
As virtio-msg-amp is very new though, I'm not sure it solves the
problem right now.
Alistair
next prev parent reply other threads:[~2026-09-04 3:19 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <20260901014650.2728658-1-alistair.francis@wdc.com>
2026-09-01 8:35 ` [PATCH 0/2] scsi: Initial commit of VirtIO PCIe Endpoint Driver Niklas Cassel
2026-09-03 11:51 ` Manivannan Sadhasivam
2026-09-04 0:12 ` Damien Le Moal
2026-09-04 3:19 ` Alistair Francis [this message]
2026-09-04 4:41 ` Damien Le Moal
2026-09-04 5:26 ` Alistair Francis
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2daeb43d582f0d960370fcd415c3b5bd3d94c67c.camel@wdc.com \
--to=alistair.francis@wdc.com \
--cc=Frank.Li@kernel.org \
--cc=James.Bottomley@hansenpartnership.com \
--cc=alistair23@gmail.com \
--cc=cassel@kernel.org \
--cc=den@valinux.co.jp \
--cc=dlemoal@kernel.org \
--cc=eperezma@redhat.com \
--cc=jasowangio@gmail.com \
--cc=kishon@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linux-scsi@vger.kernel.org \
--cc=mani@kernel.org \
--cc=mie@igel.co.jp \
--cc=mkp@kernel.org \
--cc=mst@redhat.com \
--cc=virtualization@lists.linux.dev \
--cc=xuanzhuo@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox