From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 094F65427E for ; Thu, 18 Apr 2024 20:37:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1713472675; cv=none; b=C7Oc2KtIsRUDrI78fESzoZLeIkoOay1I3SyIOe2KLAqGIUu1bbvY2SyWtkFvQrmLwzsab3e9J0BHdljOdHM0RKpGXjj8G9cf5jOBe51oM0RCn5Q77v/vCSDrfn2Am7Xok0ENksvSfZ6c+yS39Ou2nHfC/72dMmc7P3AIqmaB4Yw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1713472675; c=relaxed/simple; bh=9rCCdJIOXGK5mMIJxWx6y9XaXwhCfvryMYkeNkglHYg=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=E4zhdjo40h6FAiumI2fydp4f6qGUyqFnxbfOXmxHlUwpKC/m8IkJWqQ5iTNJfgaSsu7eVPwtXfY1aDl+d5AUyyj2XmEXp4MI+Duko+9HDW/rUNdoJrBX59HDE76xAByvv2T/8V2XmXNyS4qVcZpUjRVn+Xril85BKyek9rmCCXo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=cBAbLwMf; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="cBAbLwMf" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1713472672; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Pn2/BXHQsJTtbsxPcLAeb1M670ku7QSTpIAaDmdaDfE=; b=cBAbLwMfNQG0lQbG9fmV5FHxM6dXz6p5JT2/b7wnj4xgl1s5g63HRRhIVGs/gfmeig0VoW qhUHJbUxvYul6sHawvn/EaYPj2vQ24SwRtPUKLcfmDsEOPNXXUhtPQKMEjSlVZsUb8OfVr mnfIgoDjldh5Dj6fPh4KVQQmQp/PVxA= Received: from mail-io1-f70.google.com (mail-io1-f70.google.com [209.85.166.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-452-YOU4VQmMNviioUpa1RF9mw-1; Thu, 18 Apr 2024 16:37:51 -0400 X-MC-Unique: YOU4VQmMNviioUpa1RF9mw-1 Received: by mail-io1-f70.google.com with SMTP id ca18e2360f4ac-7c8a960bd9eso164159939f.0 for ; Thu, 18 Apr 2024 13:37:51 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1713472670; x=1714077470; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=Pn2/BXHQsJTtbsxPcLAeb1M670ku7QSTpIAaDmdaDfE=; b=T677m1S2aNvGJr2TUk9S2CtW9gZgPMTOqmbqB/53jGM2YdpIzr0/1LHjImhBJluY8F epebdgEYEaPktkBSQerkyDNgG4vb+NQV3v8IrlkpUrd2hYgauegN4kzuAzz2LD+h3ccl gCzqQ53gMmUKi51iNjCuyErlsRc11gE4Y7ErefvCz2lRNxgS54k+4O1NPC6BNywWFmG/ ZA4YiIBbqjkbtf6skMD9BjFBfumM5b7AlTxqVaI/l82N6fvap/Xkvb57UwFeC2donzBj nSKEYKwWG9riQSi22ZKOZlwLGQG+JVUR4BoP7YZJeIo7O3COvdzW75PjkziUF7nCom9K dCJQ== X-Forwarded-Encrypted: i=1; AJvYcCV0K4HsM8FT+hvRnvdykbAslUd7aaVBiVODD2LQTx+V7TUwBwWa38Tar6vzyCJ7LPvFKIFcZvVbtr4Tx1Jen8pmmoWYh6o= X-Gm-Message-State: AOJu0YzX8lYSa8dIzOCpG/JCI1zAQiWTeI/HX9D63SyizqI+ixeClRLR A9vk6o38fNF1RpNO7iZe5rAfHTAZ1MRJazXd3fMtqnZ2MCSWq3J90QlAdCxtTIZY+PMBDQX63yD 013w9/nUtikNCksPtRGbFRDSrWypcYkCW3HYf8j4YBPWDjeptEz4o X-Received: by 2002:a05:6e02:1a28:b0:36b:2c7d:cd91 with SMTP id g8-20020a056e021a2800b0036b2c7dcd91mr306704ile.25.1713472670685; Thu, 18 Apr 2024 13:37:50 -0700 (PDT) X-Google-Smtp-Source: AGHT+IF1WTU9ZaZTA/+3VejLlx0uNmpHA2zKOaPH8HbGoZrE1W44hFuolUWSTZ3FJje7EeKvr4YJ3A== X-Received: by 2002:a05:6e02:1a28:b0:36b:2c7d:cd91 with SMTP id g8-20020a056e021a2800b0036b2c7dcd91mr306682ile.25.1713472670299; Thu, 18 Apr 2024 13:37:50 -0700 (PDT) Received: from redhat.com ([38.15.36.11]) by smtp.gmail.com with ESMTPSA id y26-20020a056638015a00b004829428d517sm633977jao.63.2024.04.18.13.37.49 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 18 Apr 2024 13:37:49 -0700 (PDT) Date: Thu, 18 Apr 2024 14:37:47 -0600 From: Alex Williamson To: Yi Liu Cc: "Tian, Kevin" , Jason Gunthorpe , "joro@8bytes.org" , "robin.murphy@arm.com" , "eric.auger@redhat.com" , "nicolinc@nvidia.com" , "kvm@vger.kernel.org" , "chao.p.peng@linux.intel.com" , "iommu@lists.linux.dev" , "baolu.lu@linux.intel.com" , "Duan, Zhenzhong" , "Pan, Jacob jun" Subject: Re: [PATCH v2 0/4] vfio-pci support pasid attach/detach Message-ID: <20240418143747.28b36750.alex.williamson@redhat.com> In-Reply-To: <4037d5f4-ae6b-4c17-97d8-e0f7812d5a6d@intel.com> References: <20240412082121.33382-1-yi.l.liu@intel.com> <20240416175018.GJ3637727@nvidia.com> <20240417122051.GN3637727@nvidia.com> <20240417170216.1db4334a.alex.williamson@redhat.com> <4037d5f4-ae6b-4c17-97d8-e0f7812d5a6d@intel.com> X-Mailer: Claws Mail 4.2.0 (GTK 3.24.41; x86_64-redhat-linux-gnu) Precedence: bulk X-Mailing-List: iommu@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Thu, 18 Apr 2024 17:03:15 +0800 Yi Liu wrote: > On 2024/4/18 08:06, Tian, Kevin wrote: > >> From: Alex Williamson > >> Sent: Thursday, April 18, 2024 7:02 AM > >> > >> On Wed, 17 Apr 2024 09:20:51 -0300 > >> Jason Gunthorpe wrote: > >> > >>> On Wed, Apr 17, 2024 at 07:16:05AM +0000, Tian, Kevin wrote: > >>>>> From: Jason Gunthorpe > >>>>> Sent: Wednesday, April 17, 2024 1:50 AM > >>>>> > >>>>> On Tue, Apr 16, 2024 at 08:38:50AM +0000, Tian, Kevin wrote: > >>>>>>> From: Liu, Yi L > >>>>>>> Sent: Friday, April 12, 2024 4:21 PM > >>>>>>> > >>>>>>> A userspace VMM is supposed to get the details of the device's > >> PASID > >>>>>>> capability > >>>>>>> and assemble a virtual PASID capability in a proper offset in the > >> virtual > >>>>> PCI > >>>>>>> configuration space. While it is still an open on how to get the > >> available > >>>>>>> offsets. Devices may have hidden bits that are not in the PCI cap > >> chain. > >>>>> For > >>>>>>> now, there are two options to get the available offsets.[2] > >>>>>>> > >>>>>>> - Report the available offsets via ioctl. This requires device-specific > >> logic > >>>>>>> to provide available offsets. e.g., vfio-pci variant driver. Or may the > >>>>> device > >>>>>>> provide the available offset by DVSEC. > >>>>>>> - Store the available offsets in a static table in userspace VMM. > >> VMM gets > >>>>> the > >>>>>>> empty offsets from this table. > >>>>>>> > >>>>>> > >>>>>> I'm not a fan of requesting a variant driver for every PASID-capable > >>>>>> VF just for the purpose of reporting a free range in the PCI config > >> space. > >>>>>> > >>>>>> It's easier to do that quirk in userspace. > >>>>>> > >>>>>> But I like Alex's original comment that at least for PF there is no > >> reason > >>>>>> to hide the offset. there could be a flag+field to communicate it. or > >>>>>> if there will be a new variant VF driver for other purposes e.g. > >> migration > >>>>>> it can certainly fill the field too. > >>>>> > >>>>> Yes, since this has been such a sticking point can we get a clean > >>>>> series that just enables it for PF and then come with a solution for > >>>>> VF? > >>>>> > >>>> > >>>> sure but we at least need to reach consensus on a minimal required > >>>> uapi covering both PF/VF to move forward so the user doesn't need > >>>> to touch different contracts for PF vs. VF. > >>> > >>> Do we? The situation where the VMM needs to wholly make a up a PASID > >>> capability seems completely new and seperate from just using an > >>> existing PASID capability as in the PF case. > >> > >> But we don't actually expose the PASID capability on the PF and as > >> argued in path 4/ we can't because it would break existing userspace. > > > Come back to this statement. > > > > Does 'break' means that legacy Qemu will crash due to a guest write > > to the read-only PASID capability, or just a conceptually functional > > break i.e. non-faithful emulation due to writes being dropped? I expect more the latter. > > If the latter it's probably not a bad idea to allow exposing the PASID > > capability on the PF as a sane guest shouldn't enable the PASID > > capability w/o seeing vIOMMU supporting PASID. And there is no > > status bit defined in the PASID capability to check back so even > > if an insane guest wants to blindly enable PASID it will naturally > > write and done. The only niche case is that the enable bits are > > defined as RW so ideally reading back those bits should get the > > latest written value. But probably this can be tolerated? Some degree of inconsistency is likely tolerated, the guest is unlikely to check that a RW bit was set or cleared. How would we virtualize the control registers for a VF and are they similarly virtualized for a PF or would we allow the guest to manipulate the physical PASID control registers? > > With that then should we consider exposing the PASID capability > > in PCI config space as the first option? For PF it's simple as how > > other caps are exposed. For VF a variant driver can also fake the > > PASID capability or emulate a DVSEC capability for unused space > > (to motivate the physical implementation so no variant driver is > > required in the future) > > If kernel exposes pasid cap for PF same as other caps, and in the meantime > the variant driver chooses to emulate a DVSEC cap, then userspace follows > the below steps to expose pasid cap to VM. If we have a variant driver, why wouldn't it expose an emulated PASID capability rather than a DVSEC if we're choosing to expose PASID for PFs? > > 1) Check if a pasid cap is already present in the virtual config space > read from kernel. If no, but user wants pasid, then goto step 2). > 2) Userspace invokes VFIO_DEVICE_FETURE to check if the device support > pasid cap. If yes, goto step 3). Why do we need the vfio feature interface if a physical or virtual PASID capability on the device exposes the same info? > 3) Userspace gets an available offset via reading the DVSEC cap. What's the scenario where we'd have a VF wanting to expose PASID support which doesn't also have a variant driver that could implement a virtual PASID? > 4) Userspace assembles a pasid cap and inserts it to the vconfig space. > > For PF, step 1) is enough. For VF, it needs to go through all the 4 steps. > This is a bit different from what we planned at the beginning. But sounds > doable if we want to pursue the staging direction. Seems like if we decide that we can just expose the PASID capability for a PF then we should just have any VF variant drivers also implement a virtual PASID capability. In this case DVSEC would only be used to provide information for a purely userspace emulation of PASID (in which case it also wouldn't necessarily need the vfio feature because it might implicitly know the PASID capabilities of the device). Thanks, Alex