From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C1BD719F48D for ; Fri, 17 Jan 2025 18:06:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737137211; cv=none; b=IF2amoUgCmpBU/TgentlUOAlZVWra0xAaRm8F6b8WktPJw/Q/SK3U6Cu6UREOPxNish3BKVfzcEX+45p+D8qalgiDLrdp4UC0G4EAHvipSNUhipFyTc8LojNS703VQyhSQuhSwOVVY5rF4brJ1a6reWhk6Ve008TfqnCRXPl/SQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737137211; c=relaxed/simple; bh=k8F+Py09ow5xNxpLo2lKsypfIEO5iGKvDE0OfOKSmpE=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=rpjXBRGuI6ossMZlU+SdLBCHFu3yOzbdeXGX9cIk/VigRJLxf0ife4DtOMf43YblwupQlwGs7kMn1JViariCqL2lifKA2JB0GKGGIIIW3QlMxvM/jZGfNhIidN2CdJcNblEW61aiXGcEYNVQuIz9csp+4XTz2usJCq8LHu6t3oE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=B8TLSD3v; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="B8TLSD3v" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1737137208; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=8mjFkUumJwJiAKwRbpAuQ4KVkrv1oWc374O0VRxVMSY=; b=B8TLSD3vIr6gSEMgTwwaW/q9Y2O/RdXYgkOOl36A1Zn+/bz8Bkzlx1qf1FPMEwIlg8GgG2 yq8SqSINK36o0Ohg6JScGT05LjspMFY51i/EtWVCJwHjNu2T/e/yVaocPz4i8d6YNm6vR4 v0vRhHr5rFoJV4rIAKUfL7+ff1OgV6Q= Received: from mail-il1-f200.google.com (mail-il1-f200.google.com [209.85.166.200]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-687-8m4KuHfIPSuiCLrIHpR62w-1; Fri, 17 Jan 2025 13:06:47 -0500 X-MC-Unique: 8m4KuHfIPSuiCLrIHpR62w-1 X-Mimecast-MFC-AGG-ID: 8m4KuHfIPSuiCLrIHpR62w Received: by mail-il1-f200.google.com with SMTP id e9e14a558f8ab-3cf788751a3so875455ab.3 for ; Fri, 17 Jan 2025 10:06:47 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737137207; x=1737742007; h=content-transfer-encoding:mime-version:organization:references :in-reply-to:message-id:subject:cc:to:from:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=8mjFkUumJwJiAKwRbpAuQ4KVkrv1oWc374O0VRxVMSY=; b=edzsr++dsbKAwjQZs8O169eRHo6I2Rea3GxjL/nDucfKbyMDUrLngfdc5zeBtEH6QR 5Jrz8O/vMlpnVjAxeNlFUVYaTVYDgSlosFLV20KxijFU0AE4739tF7HsSFJIm/fwM3oJ XNCnozLPgh5kURdQ0eFjQPEdNDP6jgRy3peFmpXOg14TQKrUcfaP/LTGCaascoMOc9Wo FhbLf9fWipJ3HpwZHlblup72QBXpvRCGeirAy139m3c+YAE2es45kCb+oho3RBE+dwCH X2BG0/KmjL0T2xK+QQcOFBwTLMizmS6z8sWAeX8eATM0sfSvCDZVWl5DKhWvNNtyZ+bu X2Pw== X-Forwarded-Encrypted: i=1; AJvYcCXEnmkNfwHUXyiDOo8dMVwFAl9de+21OQxnbJ+JzS3OqkMKvdP285mTPNpDpk8yAz9w1KiPjlQ53Ok8mKs=@vger.kernel.org X-Gm-Message-State: AOJu0Yz0XMHTzA4MS6PukqAaFe+cV7LZm77hkQp7esJ6nLNMO/jgBR8w TEJV7voyJUVO6k02MLMfYfROGFaYp63S/RmjxQSge+vLCRiwkf9bo43T8JMRrQAf/rBrdyIO1kw PvhF+ybTPVYFTap3qAuCIen+ggZG62Drtz1vLgCfcw0FX4zjusGfWOj0h6MxAoA== X-Gm-Gg: ASbGnctXL5O3riDiHIBMruF5BdWpNbapXr/xFMZkEfiCm56VEA+3fJ5vJZRJWBW5h42 6eNV+ViSr4Kkr77vutGwhTM2XbpJsv4JBf6whYQTVvWNH23OF6hsceTYqp8ekak9OotBKjonYZC UPWV823wm0e6yGnV5RKmnfvkBG4qGJ/V/m6HfSnqCQSog1+Kx31j+J786i9Yt0C5fc4/84kgxJF elShWYDdGVENgEQq8E3REH3IMMTodm3Np4cTiD4W+TzAI/IfwhXOKzfFxRc X-Received: by 2002:a05:6602:81c:b0:83a:abd1:6af2 with SMTP id ca18e2360f4ac-851b62a8c7dmr64612539f.3.1737137206560; Fri, 17 Jan 2025 10:06:46 -0800 (PST) X-Google-Smtp-Source: AGHT+IG0c9dKPF5thmIZVKOyk8ONBTOuSbhlRXF42EpWjIgTSp0o1bHg+IeKA9vmUbbBJczAakOb9g== X-Received: by 2002:a05:6602:81c:b0:83a:abd1:6af2 with SMTP id ca18e2360f4ac-851b62a8c7dmr64610239f.3.1737137206154; Fri, 17 Jan 2025 10:06:46 -0800 (PST) Received: from redhat.com ([38.15.36.11]) by smtp.gmail.com with ESMTPSA id ca18e2360f4ac-851b01f2e4asm72864339f.17.2025.01.17.10.06.43 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jan 2025 10:06:45 -0800 (PST) Date: Fri, 17 Jan 2025 13:06:29 -0500 From: Alex Williamson To: Cc: , , , , , , , , , , , , , , , , , Subject: Re: [PATCH v3 2/3] vfio/nvgrace-gpu: Expose the blackwell device PF BAR1 to the VM Message-ID: <20250117130629.03648fa9.alex.williamson@redhat.com> In-Reply-To: <20250117152334.2786-3-ankita@nvidia.com> References: <20250117152334.2786-1-ankita@nvidia.com> <20250117152334.2786-3-ankita@nvidia.com> Organization: Red Hat Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Fri, 17 Jan 2025 15:23:33 +0000 wrote: > From: Ankit Agrawal > > There is a HW defect on Grace Hopper (GH) to support the > Multi-Instance GPU (MIG) feature [1] that necessiated the presence > of a 1G region carved out from the device memory and mapped as > uncached. The 1G region is shown as a fake BAR (comprising region 2 and 3) > to workaround the issue. > > The Grace Blackwell systems (GB) differ from GH systems in the following > aspects: > 1. The aforementioned HW defect is fixed on GB systems. > 2. There is a usable BAR1 (region 2 and 3) on GB systems for the > GPUdirect RDMA feature [2]. > > This patch accommodate those GB changes by showing the 64b physical > device BAR1 (region2 and 3) to the VM instead of the fake one. This > takes care of both the differences. > > Moreover, the entire device memory is exposed on GB as cacheable to > the VM as there is no carveout required. > > Link: https://www.nvidia.com/en-in/technologies/multi-instance-gpu/ [1] > Link: https://docs.nvidia.com/cuda/gpudirect-rdma/ [2] > > Suggested-by: Alex Williamson > Signed-off-by: Ankit Agrawal > --- > drivers/vfio/pci/nvgrace-gpu/main.c | 65 ++++++++++++++++++----------- > 1 file changed, 41 insertions(+), 24 deletions(-) > > diff --git a/drivers/vfio/pci/nvgrace-gpu/main.c b/drivers/vfio/pci/nvgrace-gpu/main.c > index 85eacafaffdf..89d38e3c0261 100644 > --- a/drivers/vfio/pci/nvgrace-gpu/main.c > +++ b/drivers/vfio/pci/nvgrace-gpu/main.c > @@ -17,9 +17,6 @@ > #define RESMEM_REGION_INDEX VFIO_PCI_BAR2_REGION_INDEX > #define USEMEM_REGION_INDEX VFIO_PCI_BAR4_REGION_INDEX > > -/* Memory size expected as non cached and reserved by the VM driver */ > -#define RESMEM_SIZE SZ_1G > - > /* A hardwired and constant ABI value between the GPU FW and VFIO driver. */ > #define MEMBLK_SIZE SZ_512M > > @@ -72,7 +69,7 @@ nvgrace_gpu_memregion(int index, > if (index == USEMEM_REGION_INDEX) > return &nvdev->usemem; > > - if (index == RESMEM_REGION_INDEX) > + if (nvdev->resmem.memlength && index == RESMEM_REGION_INDEX) > return &nvdev->resmem; > > return NULL; > @@ -757,21 +754,31 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > u64 memphys, u64 memlength) > { > int ret = 0; > + u64 resmem_size = 0; > > /* > - * The VM GPU device driver needs a non-cacheable region to support > - * the MIG feature. Since the device memory is mapped as NORMAL cached, > - * carve out a region from the end with a different NORMAL_NC > - * property (called as reserved memory and represented as resmem). This > - * region then is exposed as a 64b BAR (region 2 and 3) to the VM, while > - * exposing the rest (termed as usable memory and represented using usemem) > - * as cacheable 64b BAR (region 4 and 5). > + * On Grace Hopper systems, the VM GPU device driver needs a non-cacheable > + * region to support the MIG feature owing to a hardware bug. Since the > + * device memory is mapped as NORMAL cached, carve out a region from the end > + * with a different NORMAL_NC property (called as reserved memory and > + * represented as resmem). This region then is exposed as a 64b BAR > + * (region 2 and 3) to the VM, while exposing the rest (termed as usable > + * memory and represented using usemem) as cacheable 64b BAR (region 4 and 5). > * > * devmem (memlength) > * |-------------------------------------------------| > * | | > * usemem.memphys resmem.memphys > + * > + * This hardware bug is fixed on the Grace Blackwell platforms and the > + * presence of fix can be determined through nvdev->has_mig_hw_bug_fix. > + * Thus on systems with the hardware fix, there is no need to partition > + * the GPU device memory and the entire memory is usable and mapped as > + * NORMAL cached. > */ > + if (!nvdev->has_mig_hw_bug_fix) > + resmem_size = SZ_1G; > + > nvdev->usemem.memphys = memphys; > > /* > @@ -780,23 +787,30 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > * memory (usemem) is added to the kernel for usage by the VM > * workloads. Make the usable memory size memblock aligned. > */ > - if (check_sub_overflow(memlength, RESMEM_SIZE, > + if (check_sub_overflow(memlength, resmem_size, > &nvdev->usemem.memlength)) { > ret = -EOVERFLOW; > goto done; > } > > - /* > - * The USEMEM part of the device memory has to be MEMBLK_SIZE > - * aligned. This is a hardwired ABI value between the GPU FW and > - * VFIO driver. The VM device driver is also aware of it and make > - * use of the value for its calculation to determine USEMEM size. > - */ > - nvdev->usemem.memlength = round_down(nvdev->usemem.memlength, > - MEMBLK_SIZE); > - if (nvdev->usemem.memlength == 0) { > - ret = -EINVAL; > - goto done; > + if (!nvdev->has_mig_hw_bug_fix) { > + /* > + * If the device memory is split to workaround the MIG bug, > + * the USEMEM part of the device memory has to be MEMBLK_SIZE > + * aligned. This is a hardwired ABI value between the GPU FW and > + * VFIO driver. The VM device driver is also aware of it and make > + * use of the value for its calculation to determine USEMEM size. > + * > + * If the hardware has the fix for MIG, there is no requirement > + * for splitting the device memory to create RESMEM. The entire > + * device memory is usable and will be USEMEM. > + */ > + nvdev->usemem.memlength = round_down(nvdev->usemem.memlength, > + MEMBLK_SIZE); > + if (nvdev->usemem.memlength == 0) { > + ret = -EINVAL; > + goto done; > + } Why does this operation need to be predicated on the buggy device? Does GB have memory that's not a multiple of 512MB? I was expecting this would be a no-op on GB and therefore wouldn't need to be conditional. Thanks, Alex > } > > if ((check_add_overflow(nvdev->usemem.memphys, > @@ -813,7 +827,10 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > * the BAR size for them. > */ > nvdev->usemem.bar_size = roundup_pow_of_two(nvdev->usemem.memlength); > - nvdev->resmem.bar_size = roundup_pow_of_two(nvdev->resmem.memlength); > + > + if (nvdev->resmem.memlength) > + nvdev->resmem.bar_size = > + roundup_pow_of_two(nvdev->resmem.memlength); > done: > return ret; > }