Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Will Deacon <will@kernel.org>
To: Marc Zyngier <maz@kernel.org>
Cc: Fuad Tabba <fuad.tabba@linux.dev>,
	oupton@kernel.org, kvmarm@lists.linux.dev,
	linux-arm-kernel@lists.infradead.org,
	linux-kernel@vger.kernel.org, catalin.marinas@arm.com,
	joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com,
	yuzenghui@huawei.com, mark.rutland@arm.com, steven.price@arm.com,
	vdonnefort@google.com, qperret@google.com
Subject: Re: [PATCH v3 15/18] KVM: arm64: Reject host access to protected VM private state
Date: Fri, 18 Sep 2026 14:21:31 +0100	[thread overview]
Message-ID: <aq0628Ci4pfEj0KG@willie-the-truck> (raw)
In-Reply-To: <865x045mke.wl-maz@kernel.org>

Hi Marc, Fuad,

On Thu, Sep 17, 2026 at 09:06:25AM +0100, Marc Zyngier wrote:
> On Wed, 16 Sep 2026 20:07:28 +0100,
> Fuad Tabba <fuad.tabba@linux.dev> wrote:
> > On Wed, 16 Sept 2026 at 17:30, Marc Zyngier <maz@kernel.org> wrote:
> > [...]
> > > > diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
> > [...]
> > > > +/*
> > > > + * Once a protected vCPU has run, the host copy is not the guest's state,
> > > > + * and EL2 has read mp_state, which it does only at hyp vCPU creation.
> > > > + */
> > > > +static long pkvm_filter_vcpu_ioctl(struct kvm_vcpu *vcpu, unsigned int ioctl)
> > > > +{
> > > > +        switch (ioctl) {
> > > > +        case KVM_ARM_VCPU_INIT:
> > > > +        case KVM_SET_ONE_REG:
> > > > +        case KVM_GET_ONE_REG:
> > > > +                if (vcpu_is_protected(vcpu) && vcpu_has_run_once(vcpu))
> > > > +                        return -EPERM;
> > > > +        }
> > > > +
> > > > +        return 0;
> > > > +}
> > > > +
> > >
> > > I'm not keen on returning -EPERM for the ONE_REG stuff. For a start,
> > > X0 *is* valid on MMIO, and when you want to support LD64B and co,
> > > you'll need to show the actual data there.
> > >
> > > I'd rather return what is in the host vcpu structure, as normal.

I just wanted to chuck my thoughts in here, as some of this is my fault
and it would be really helpful to discuss it on the list. As you probably
know, on Android, pKVM has the semantics you describe above: the ONE_REG
calls succeed, but the register state isn't propagated to the guest once
it's started running. I think we largely did it this way because we were
solving a million problems at once during the initial development and
having to pipe-clean VMMs wasn't particularly appealing at the time.
It's also worth adding that, to my knowledge, we've never had any issues
in Android because of this decision.

So, based on that, it really sounds like it's the best option. If it ain't
broke, don't fix it!

*However*, I can't think of another upstream interface that behaves like
that and, when we came to document it, it was quite difficult to explain
it in a way that could be extended in the future. If we say that ONE_REG
doesn't propagate to the vCPU after first run and returns whatever was
last written, then userspace could (even accidentally) rely on that. If
we wanted to extend it, I think we'd not only need a method to advertise
the new behaviour (which I think you probably want even if ONE_REG
returned an error before) but also an opt-in for userspace to say that
it's ok with the new behaviour. In some ways, it feels a bit like the
"unchecked flags" problem for syscall arguments.

> > I think we should keep the error, for the reason s390 returns -EINVAL
> > from ONE_REG on a protected vCPU and x86 does from the register ioctls
> > for protected VM types:
> 
> News flash, this is not x86, nor s390. I don't feel constrained by
> other architecture, and we deviate *everywhere* already.

Right, this wasn't the rationale.

> > once the vCPU has run, the copy is the VMM's
> > boot state plus what the exit handlers copy out, and the VMM can't
> > distinguish them. x0 on MMIO is one of those fields, but the VMM reads
> > it from kvm_run->mmio. For LD64B, EL2 would copy the operands out the
> > same way and the error could be relaxed to those registers. Loosening
> > later breaks nobody. The comment and the message do overstate it by
> > calling the copy "not the guest's state", and I'll reword both in v4.
> > More on the kvmtool thread [1].
> 
> A protected-aware VMM already knows it cannot obtain the registers.

I think you can use that argument both ways: if the VMM knows it cannot
obtain the registers, then it's fine to return an error if the VMM does
something wrong.

In our recent investigation, it turned out that both kvmtool and crosvm
were perfectly happy with ONE_REG returning an error in normal operation
but we needed a handful of fixes for kvmtool [1] to handle things like
the 'debug' command.

If it's helpful, we can ask the crosvm developers here if they have
opinions for/against the two behaviours?

> I don't want to have to revisit the userspace interface once you have
> to relax it, because I know for sure that you will have to.

I think we'll have to do that either way, no? The VMM needs a way to
know that ONE_REG works and the situations in which it works. If the old
behaviour wasn't to return an error, it's also going to need a way to
enable the new behaviour.

> As far as LD64B is concerned, there is no place to copy anything in
> the run structure, and the relaxation would require to cover all the
> GPRs, ESR, and FAR. At this stage, returning whatever is there is the
> correct thing to do IMO.

The nice thing about that is that we can presumably tie it in with LD64B
support for protected VMs.

Will


  parent reply	other threads:[~2026-09-18 13:21 UTC|newest]

Thread overview: 49+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14 11:33 [PATCH v3 00/18] KVM: arm64: Confine protected VM vCPU state to EL2 Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 01/18] KVM: arm64: Sync HCR_EL2.VSE back to the host vCPU under pKVM Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 02/18] KVM: arm64: Validate the host vCPU's VM before reading it " Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 03/18] KVM: arm64: Pin the host vCPU before adjusting its PC " Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 04/18] KVM: arm64: Disable steal time for protected VMs Fuad Tabba
2026-09-22 14:34   ` Vincent Donnefort
2026-09-14 11:33 ` [PATCH v3 05/18] KVM: arm64: Introduce per-EC entry handlers for pKVM Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 06/18] KVM: arm64: Skip fixed-feature state flush for protected vCPUs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 07/18] KVM: arm64: Add {flush,sync}_hyp_timer_state() primitives Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 08/18] KVM: arm64: Add system register reset framework for protected VMs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 09/18] KVM: arm64: Implement HVC handling for protected guests at EL2 Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 10/18] KVM: arm64: Handle PSCI calls for protected VMs " Fuad Tabba
2026-09-22 16:35   ` Vincent Donnefort
2026-09-22 16:37     ` Vincent Donnefort
2026-09-22 17:07   ` Vincent Donnefort
2026-09-23  9:51     ` Fuad Tabba
2026-09-24  8:30       ` Will Deacon
2026-09-24 11:26         ` Fuad Tabba
2026-09-24 12:14           ` Will Deacon
2026-09-24 15:21             ` Fuad Tabba
2026-10-01 12:58               ` Will Deacon
2026-10-01 13:11                 ` Fuad Tabba
2026-10-01 12:59   ` Will Deacon
2026-10-01 13:11     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 11/18] KVM: arm64: Restrict KVM_ARM_VCPU_INIT and PSCI version for protected VMs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 12/18] KVM: arm64: Prevent host PC adjustments for protected vCPUs Fuad Tabba
2026-09-14 13:42   ` Marc Zyngier
2026-09-14 14:43     ` Fuad Tabba
2026-09-15 11:02       ` Marc Zyngier
2026-09-15 11:19         ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 13/18] KVM: arm64: Inject an UNDEF at EL2 for unhandled protected guest exits Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 14/18] KVM: arm64: Add per-EC entry/exit state marshalling for protected guests Fuad Tabba
2026-09-16 16:27   ` Marc Zyngier
2026-09-16 19:05     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 15/18] KVM: arm64: Reject host access to protected VM private state Fuad Tabba
2026-09-16 16:30   ` Marc Zyngier
2026-09-16 19:07     ` Fuad Tabba
2026-09-17  8:06       ` Marc Zyngier
2026-09-17 18:42         ` Fuad Tabba
2026-09-18 13:21         ` Will Deacon [this message]
2026-09-18 13:24           ` Will Deacon
2026-09-27  8:20           ` Marc Zyngier
2026-09-27 12:37             ` Fuad Tabba
2026-09-28 19:00             ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 16/18] KVM: arm64: Reject host power-on of a vCPU that EL2 holds powered off Fuad Tabba
2026-09-16 16:43   ` Marc Zyngier
2026-09-16 19:08     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 17/18] KVM: arm64: Advertise the capabilities that protected VMs support Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 18/18] KVM: arm64: Document the protected VM userspace API Fuad Tabba

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aq0628Ci4pfEj0KG@willie-the-truck \
    --to=will@kernel.org \
    --cc=catalin.marinas@arm.com \
    --cc=fuad.tabba@linux.dev \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=qperret@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=steven.price@arm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=vdonnefort@google.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox