From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD46720322 for ; Tue, 21 May 2024 18:19:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1716315594; cv=none; b=Dk7ig0++wl9Ngl3sg9ZEIh7X7ATXbMwpulWy/aDD1rbCpGC0sBYJf8/CvShJII8QJV+DCR1jPtUrcqCR7KOFZNC+BpflP52QGn1ggXM3gqEX7NDLCCPIJ/jHijZzLj5EuDFhVDsrXjdzrEgHvaiQoghjsQfLeab1d6W5/DMJKrQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1716315594; c=relaxed/simple; bh=0dqwcK/XPOcEv1mTsxYC7i1Y/75vW0HO1WRQSYOUEbI=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Q40RcZfqPRCNrIcRmVX4wJq3VAeZYZuHBfIdthtz/HfYqMC57I0LDCEonf9chU6t1HqTCBJbW3lI/qRFb1ujK6sr+93cibLb0qRNlT++SWxUYk2OGGy46bRVEF2AtGcmowbtB6s1+isIAQlOmBos23l8zWkiCoBLM4KnZcjcD7A= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=TDAOCVsp; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="TDAOCVsp" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1716315591; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=1Isde2dh7/647rfx8u6foFeQh7vHwio2r62p6n+pus0=; b=TDAOCVsp9vx7Reg95AqQGjVTB2PSj2XQpmMgIB6rjn8gBosQS4mpZeJqqTz06MtbkS8qXS bwkKXjRjKPyT6KtKnCTsJEmsKLbqisHTIgb3wyA2RmWiXOqYWY7u5RT5qm21FMld9Lv1dC e554nRi3fm+BbHY+7J3N753LzTS/wWk= Received: from mail-io1-f71.google.com (mail-io1-f71.google.com [209.85.166.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-677-rVvaoyMhNxmt9H7Yhzf-mg-1; Tue, 21 May 2024 14:19:50 -0400 X-MC-Unique: rVvaoyMhNxmt9H7Yhzf-mg-1 Received: by mail-io1-f71.google.com with SMTP id ca18e2360f4ac-7e1d56d36b9so1228993339f.1 for ; Tue, 21 May 2024 11:19:50 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1716315589; x=1716920389; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=1Isde2dh7/647rfx8u6foFeQh7vHwio2r62p6n+pus0=; b=mI5nTPoIdw+zHGeztV3vkSiQKspXQc+VCD++D9qLSitpDqLa0hM8/HEiDXShlWPjw/ qJHN3scVEYbsCT3lOLsXThOcu17xXt5tWwde+nYQSUBBFYhP9PdO+nwbt/AfnPL4Wr+w AFiENlv31cXa36FpV/mtVLvt/dgqYwNSEdRpdwTuB4d2RX5ccbOvMPqjxapB+9+n9oWd MvS3R9iK8LyBVwnRMHvFVknI1LORn5+3XSV4T3ie5RQhuL6ra9yl9qbvyoasOJVrRBeS HIMKm3MoI+7hsou+o6GwHfwS7PK1pW8T2crKBJs6p0u7Sb6QKkPWJAHMXFcCrN5rUuzE dTXA== X-Forwarded-Encrypted: i=1; AJvYcCXAS/dKRMrbyy+4TheiH5zQqVr6hGDZ9SIF3LspjuWazgOt1CnFpdwpoUwHl9CwL8CbEqGIGhB7eA6Da/NQhElSj17mJzI= X-Gm-Message-State: AOJu0Ywml9g+YltKNNeggfmdILs95cBqLlx3fsfetlhZAFboSrb9mngl kSxGhgSKteFP/L0xH4SknOdw7u80BnGRck53EcLjQGjNW8Am/KJ1awysiAzylml5OmnKXVi3CZT v0rGLYgzZvXGJ9Xp4ZZ6ipdEtDD3blibWgwOJIeV6JJa3Uoeok6yg X-Received: by 2002:a5d:9553:0:b0:7e1:973b:fbfc with SMTP id ca18e2360f4ac-7e1b520c7b7mr3444747639f.15.1716315589693; Tue, 21 May 2024 11:19:49 -0700 (PDT) X-Google-Smtp-Source: AGHT+IF4dEPJMmyq53uNgsrumO7N0Qr7jKz/Z8GDsO+x2zbKwDnCjIJWu7WA5azwt4gAJaAXiGpCZA== X-Received: by 2002:a5d:9553:0:b0:7e1:973b:fbfc with SMTP id ca18e2360f4ac-7e1b520c7b7mr3444744239f.15.1716315589310; Tue, 21 May 2024 11:19:49 -0700 (PDT) Received: from redhat.com ([38.15.36.11]) by smtp.gmail.com with ESMTPSA id 8926c6da1cb9f-489376de0e3sm7019591173.146.2024.05.21.11.19.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 May 2024 11:19:48 -0700 (PDT) Date: Tue, 21 May 2024 12:19:45 -0600 From: Alex Williamson To: Jason Gunthorpe Cc: "Tian, Kevin" , "Vetter, Daniel" , "Zhao, Yan Y" , "kvm@vger.kernel.org" , "linux-kernel@vger.kernel.org" , "x86@kernel.org" , "iommu@lists.linux.dev" , "pbonzini@redhat.com" , "seanjc@google.com" , "dave.hansen@linux.intel.com" , "luto@kernel.org" , "peterz@infradead.org" , "tglx@linutronix.de" , "mingo@redhat.com" , "bp@alien8.de" , "hpa@zytor.com" , "corbet@lwn.net" , "joro@8bytes.org" , "will@kernel.org" , "robin.murphy@arm.com" , "baolu.lu@linux.intel.com" , "Liu, Yi L" Subject: Re: [PATCH 4/5] vfio/type1: Flush CPU caches on DMA pages in non-coherent domains Message-ID: <20240521121945.7f144230.alex.williamson@redhat.com> In-Reply-To: <20240521163400.GK20229@nvidia.com> References: <20240509121049.58238a6f.alex.williamson@redhat.com> <20240510105728.76d97bbb.alex.williamson@redhat.com> <20240516143159.0416d6c7.alex.williamson@redhat.com> <20240517171117.GB20229@nvidia.com> <20240521160714.GJ20229@nvidia.com> <20240521102123.7baaf85a.alex.williamson@redhat.com> <20240521163400.GK20229@nvidia.com> X-Mailer: Claws Mail 4.2.0 (GTK 3.24.41; x86_64-redhat-linux-gnu) Precedence: bulk X-Mailing-List: iommu@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Tue, 21 May 2024 13:34:00 -0300 Jason Gunthorpe wrote: > On Tue, May 21, 2024 at 10:21:23AM -0600, Alex Williamson wrote: > > > > Intel GPU weirdness should not leak into making other devices > > > insecure/slow. If necessary Intel GPU only should get some variant > > > override to keep no snoop working. > > > > > > It would make alot of good sense if VFIO made the default to disable > > > no-snoop via the config space. > > > > We can certainly virtualize the config space no-snoop enable bit, but > > I'm not sure what it actually accomplishes. We'd then be relying on > > the device to honor the bit and not have any backdoors to twiddle the > > bit otherwise (where we know that GPUs often have multiple paths to get > > to config space). > > I'm OK with this. If devices are insecure then they need quirks in > vfio to disclose their problems, we shouldn't punish everyone who > followed the spec because of some bad actors. > > But more broadly in a security engineered environment we can trust the > no-snoop bit to work properly. The spec has an interesting requirement on devices sending no-snoop transactions anyway (regarding PCI_EXP_DEVCTL_NOSNOOP_EN): "Even when this bit is Set, a Function is only permitted to Set the No Snoop attribute on a transaction when it can guarantee that the address of the transaction is not stored in any cache in the system." I wouldn't think the function itself has such visibility and it would leave the problem of reestablishing coherency to the driver, but am I overlooking something that implicitly makes this safe? ie. if the function isn't permitted to perform no-snoop to an address stored in cache, there's nothing we need to do here. > > We also then have the question of does the device function > > correctly if we disable no-snoop. > > Other than the GPU BW issue the no-snoop is not a functional behavior. As with some other config space bits though, I think we're kind of hoping for sloppy driver behavior to virtualize this. The spec does allow the bit to be hardwired to zero: "This bit is permitted to be hardwired to 0b if a Function would never Set the No Snoop attribute in transactions it initiates." But there's no capability bit that allows us to report whether the device supports no-snoop, we're just hoping that a driver writing to the bit doesn't generate a fault if the bit doesn't stick. For example the no-snoop bit in the TLP itself may only be a bandwidth issue, but if the driver thinks no-snoop support is enabled it may request the device use the attribute for a specific transaction and the device could fault if it cannot comply. > > The more secure approach might be that we need to do these cache > > flushes for any IOMMU that doesn't maintain coherency, even for > > no-snoop transactions. Thanks, > > Did you mean 'even for snoop transactions'? I was referring to IOMMUs that maintain coherency regardless of no-snoop transactions, ie domain->enforce_cache_coherency (ex. snoop control/SNP on Intel), so I meant as typed, the IOMMU maintaining coherency even for no-snoop transactions. That's essentially the case we expect and we don't need to virtualize no-snoop enable on the device. > That is where this series is, it assumes a no-snoop transaction took > place even if that is impossible, because of config space, and then > does pessimistic flushes. So are you proposing that we can trust devices to honor the PCI_EXP_DEVCTL_NOSNOOP_EN bit and virtualize it to be hardwired to zero on IOMMUs that do not enforce coherency as the entire solution? Or maybe we trap on setting the bit to make the flushing less pessimistic? Intel folks might be able to comment on the performance hit relative to iGPU assignment of denying the device the ability to use no-snoop transactions (assuming the device control bit is actually honored). The latency of flushing caches on touching no-snoop enable might be prohibitive in the latter case. Thanks, Alex