From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yx1-f43.google.com (mail-yx1-f43.google.com [74.125.224.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEC163DAAC7 for ; Sat, 22 Aug 2026 14:25:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.224.43 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787408756; cv=none; b=l6mqY/wtRkagm6veWxtgi3OtdcmqQ0/RIVK7fQlE8GLSGQmToxYdSl+t5jhsakmsSwwqoXg9P8OEl2bpnKUxqZZHj7paSTVgrPgT2GXN2xakmMgCTDmstRZfwtqKhfDv31nORnX6tdQ1sY7mq9M5utOFH/BlE7Q057BcKmcEv5I= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787408756; c=relaxed/simple; bh=ov6mNDQXtxxGgUhlfRQgI2X3AI5PJ8uI9LWkgqbYbb4=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=g0GhO0YbIK2p7pJSAy2O1jE6ilpikqtnZH6CbEzXnwxx5pxaPEzdmpIfmj4IXY1vGDkCMwjopScuAl5gwEsqFIlCXYB/uOvITzR8GBJh8r6OSRZ9qJWFKY0ZiqNmgX4KlGBcHyqIDpm4/RzHyQgpYMjZZB5wl1Pjl1lEKIR9pCY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=TBWAFwbi; arc=none smtp.client-ip=74.125.224.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="TBWAFwbi" Received: by mail-yx1-f43.google.com with SMTP id 956f58d0204a3-66c82b32121so2991923d50.1 for ; Sat, 22 Aug 2026 07:25:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787408742; x=1788013542; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=QxqAR5ocTjBBUS/ny625Id0I7Kf488fcpzoDzJMtITk=; b=TBWAFwbitAfm+YiTo+qdtitSLVqKwBhMMLq3+Bxy+9SaqxmjoDViuDjtuK41O21x6V htPBP/b/MNvgf7yy0xM/RtFIphJizHhyCIPw66OZaaHwmb4771HP3knHZ0gChO/3UXxp lXaawJidgROK28TT5ZnbbyKO0/2E2e4EinzvXkVEGuuOauN0jK4dQroJNsO1eQkY2KC8 rJp9nNsYCTmaPSaYUmgzi9Y/glfnK80FFQu541nV8mWFRrhcfDKTW1drVB1tN/t4AAPn jSa81XyjmW4QFOVmDFyKXeHx2Rr8r1g/Wl42qdLzsJ/OKlJYmBfxuZD6ZAV2yZkE0jZ5 fzfQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787408742; x=1788013542; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QxqAR5ocTjBBUS/ny625Id0I7Kf488fcpzoDzJMtITk=; b=qBV56EDkGgk8hud9cNQ+pRjE0WUTqKpmmxtJw3wv5N7Vbxt8upH9/UTVcxux0nH1tZ ZjF3u7pVq/mdFpxGOrlZxvimrHe+8tFCoytlXf0Zk9ou8XmZivsJVny5iSyMHw4LX3S9 pZ2acDyxOkOqxNqt2k8voTUQwckZCWZ/g9zaO82kA1PLQIg0uSTMgQy3R34hWpl4XV76 5l9SqhctOKHnr4kIE/coQHnftG8Hnio+E3OR63ftsxzytU9vDOJTJz7XhnTLl/DxEdZ/ EEmVZGSw0vpZZaxBt3oU8NN2Mni0emBajFuWYrDXlSSOkgcQ/Rue0RDAn2k0RYqLctWz p+rQ== X-Forwarded-Encrypted: i=1; AHgh+RoPDeRWrPi/+obCMj8leRmF6b5KOkXLjQOAG0RvA0JxymbLyAio/iwbqYoT9D2NVxzraBU47+TaziFaJUA=@vger.kernel.org X-Gm-Message-State: AFuF++k2XfLn29IZADpjewwkRvZjs2yDdepkag2n61/Ut+RQkKefzRuZ moHi2iLbbTRzFfRSIu3Tg5E3+i0BNXzCtC02vn87i41gEK7WiBvp6ts8 X-Gm-Gg: AR+sD11HYhOQDaq01nbPxRdwMKsVcocDwoeIvd/0VnB8VpcURMKtxx/mlzCddP6cHAd M+oFQVLXkyTtiYgq3n+E3f8nkITluwyFtHDDPDB9VMTZ0y9G90LdSThId4zQZ4jtlNdu1Uj/goI W7ST8Me14E6qSzDPhG1DDcSnG5RUmwHgL5lTFRAtQZEq0+FKOG+tdwCSNy9xYES3lsRXpp+G5kS tuWglWNAfcbvyOxdqp8f1Rp+pgQrYBx7eOyJW3vj/6G5bDH4oZ+E+Jty2V2FjQTOXXoGdWPB5fA 3XbzdlJcVTSCUIF8MHhlNzvy+bI5e93PfHmB0T+ZAAfHWAgsTRIik1XFQAuAXuZoYI1lWfQm754 KJsaf2sviPAOq4FwLP7K5z6CMxV8qFsRjxT3qDZUdXt3/2ZqEYAHF0vnLFCkp/Q0CXnW4tT5TwF Vye74tUdszl0dYhwm/IR0yKJNiB3fbXnJ1ufXJHyt2v+CY3cWdw2Ub2mjJAl/WvKVaqWSyfkDNr MRHeNzaBP1OXtGSeJ+jlGaTIlGZjzVyrYdQgHyYrWW0CKGmvBg2vju1OAd8ksllvNXY X-Received: by 2002:a05:690e:d8c:b0:66c:f689:8e60 with SMTP id 956f58d0204a3-66cf6899d6cmr910360d50.23.1787408742210; Sat, 22 Aug 2026 07:25:42 -0700 (PDT) Received: from DigitalOasis.tail29fcbc.ts.net (104-182-93-176.lightspeed.brhmal.sbcglobal.net. [104.182.93.176]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66cf4a2c3c2sm911142d50.17.2026.08.22.07.25.41 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 07:25:41 -0700 (PDT) From: Bocaj Gnuoy To: amd-gfx@lists.freedesktop.org Cc: Felix.Kuehling@amd.com, alexander.deucher@amd.com, christian.koenig@amd.com, airlied@gmail.com, simona@ffwll.ch, oded.gabbay@gmail.com, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, Bocaj Gnuoy Subject: [PATCH] drm/amdkfd: don't leak BOs when process teardown can't unmap them Date: Sat, 22 Aug 2026 09:25:21 -0500 Message-ID: <20260822142521.899098-1-bocajgnuoy@gmail.com> X-Mailer: git-send-email 2.55.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit kfd_process_device_free_bos() discards the return value of both amdgpu_amdkfd_gpuvm_unmap_memory_from_gpu() and amdgpu_amdkfd_gpuvm_free_memory_of_gpu(), then drops the idr handle unconditionally. That is fine as long as the free succeeds, but the unmap has to validate the page-table BOs first (vm_validate_pt_pd_bos()), and that allocates. When the process is dying *because* memory is exhausted - e.g. GTT pinned up to amdgpu.gttsize by a compute client that then aborted on ENOMEM - the unmap fails with "failed to validate PT BOs", mem->mapped_to_gpu_memory stays non-zero, and amdgpu_amdkfd_gpuvm_free_memory_of_gpu() bails out with -EBUSY before it removes the BO from process_info->kfd_bo_list. The handle is dropped anyway, so nothing will ever free that kgd_mem again. The consequences are worse than a leak. drm_release() -> amdgpu_vm_fini() -> amdgpu_amdkfd_gpuvm_destroy_cb() then trips WARN_ON(!list_empty(&process_info->kfd_bo_list)); WARN_ON(!list_empty(&process_info->userptr_inval_list)); and frees process_info regardless, while the leaked BO stays in TTM's eviction LRU with bo_vas pointing into the just-freed amdgpu_vm. A subsequent amdgpu_gem_create_ioctl() that forces TTM eviction can then walk into it and spin in amdgpu_vm_bo_move() -> _raw_spin_lock(), which produces repeating rcu_preempt self-detected stalls, blocks reclaim-related work (__lru_add_drain_all()) and eventually makes the machine unusable - within about ten minutes here. A recoverable ENOMEM becomes an unrecoverable hang. Note the free path itself does not allocate - it reserves the BO, detaches the attachments and drops the references - so it can complete even when the unmap could not. Give amdgpu_amdkfd_gpuvm_free_memory_of_gpu() a @force flag and set it on the teardown callers, which have no later chance to try again. The -EBUSY guard is kept for the ioctl paths, where userspace can still unmap and retry, and where kfd_ioctl_free_memory_of_gpu() already leaves the handle in place on failure. Reported on: RX 6900 XT (gfx1030), amdgpu 3.64.0, kernel 7.1.6/7.1.8, booted with amdgpu.gttsize=4096. The failing allocation was on the Navi 21 at 0000:06:00.0, which enumerates as ROCm0; a gfx1100 is also present in the same box but was not the device that ran out. Tooling disclosure, per Documentation/process/generated-content.rst: this bug was diagnosed and this patch written with the assistance of Claude Code (claude-opus-5) in a single interactive session. The input was a kernel trace captured on the reporting host - two WARNs out of amdgpu_amdkfd_gpuvm_destroy_cb() during an aborting client's teardown, followed by an unrecoverable deadlock in the TTM eviction path in the next process to allocate - together with the request to find the cause and, if possible, fix it. The assistant read the KFD BO lifecycle in the v7.1.8 and v7.2 sources, identified the -EBUSY early return as the point at which the BO is stranded on the process_info lists, and wrote both the diff and this changelog. The diff was then rebased onto amd-staging-drm-next and the call-site audit redone against that tree: one prototype, one definition and eight call sites, all converted. The rebase caught a real miss - the tree has gained kfd_process_free_gpuvm_map() alongside kfd_process_free_gpuvm(), and the original hunk's trailing context matched the wrong one of the two. Both are teardown-only (kfd_process_destroy_pdds() and the create_process() error unwind), so both take force=true. checkpatch.pl was the only additional analysis tool used. Tested on a RX 6900 XT + RX 7900 XTX box, kernel 7.2.0-1-cachyos, booted with amdgpu.gttsize=4096 so the GTT wall is reachable. A HIP-linked client (llama.cpp with both the HIP and Vulkan backends present) was made to pin host memory up to the 4 GiB cap and die on the resulting ENOMEM. Same kernel version, same cap, same workload, same binary either side; the patch is the only variable: unpatched patched peak GTT 4.00 GiB 4.00 GiB "failed to validate PT BOs" 34 36 exit 139/SEGV 139/SEGV teardown WARNs 2 0 "Force-freeing BO VA" 0 9 Both sides died the same way, on SIGSEGV, so the difference is in teardown and not in how the client failed. The force-free count of 9 matched the 9 KFD allocations an earlier kprobe run measured for this client, consistent with all of those allocations being reclaimed. End to end: on the unpatched kernel, relaunching the same workload after that abort is what hangs the machine - the new client's amdgpu_gem_create_ioctl walks the eviction LRU into the orphaned object. With the patch, a second run under heavier pressure (1904 "failed to validate PT BOs", 35 force-freed BOs) again left zero WARNs, and the relaunch then loaded in 12 seconds and served normally instead of deadlocking. GTT returned to its 0.05 GiB idle baseline afterwards, so the memory is genuinely reclaimed rather than merely not deadlocking. Across four runs the client died variously on SIGABRT and SIGSEGV; the leak tracked memory exhaustion at teardown, never the signal. Note the WARNs that fire vary between runs: the originally reported crash tripped :1608 and :1610 (kfd_bo_list, userptr_inval_list), while this reproduction tripped :1608 and :1609 (kfd_bo_list, userptr_valid_list). Both are the same skipped list_del_init(&mem->validate_list) - that node serves whichever list the BO currently sits on, so the leak is list-agnostic. The client also died on SIGSEGV rather than the abort seen originally; the leak depends on memory being exhausted at teardown, not on how the process died, which is consistent with SIGKILL alone never reproducing it. Limitations, per Documentation/process/coding-assistants.rst step 8: - The deadlock itself was reproduced on the unpatched kernel, but its stack trace was not captured: the jammed workqueue takes journald down with it, so nothing reaches disk. Capturing it needs netconsole or a serial console. The teardown WARNs, which are the direct measure of the leak, are captured in full. - The claim that the deadlocked spinlock lives inside the leaked object still rests on address arithmetic (lock at the WARNed structure + 0x30). Confirming the offset needs a debug build, which has not been done. - Verification is single-box: one gfx1100 + gfx1030 machine, one kernel version. The -EBUSY guard itself dates to commit a46a2cd103a8 ("drm/amdgpu: Add GPUVM memory management functions for KFD") and is correct for the ioctl path. The leak became reachable only once a caller started dropping the last handle regardless of the return value. Fixes: 52b29d73340d ("drm/amdkfd: Add per-process IDR for buffer handles") Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5672 Assisted-by: LLM checkpatch Tested-by: Bocaj Gnuoy Signed-off-by: Bocaj Gnuoy --- drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd.h | 2 +- .../gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c | 25 ++++++++++++++++--- drivers/gpu/drm/amd/amdkfd/kfd_chardev.c | 9 ++++--- drivers/gpu/drm/amd/amdkfd/kfd_process.c | 21 ++++++++++++---- 4 files changed, 43 insertions(+), 14 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd.h index 1b7dc0d3963b..4ad843105443 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd.h @@ -326,7 +326,7 @@ int amdgpu_amdkfd_gpuvm_alloc_memory_of_gpu( uint64_t *offset, uint32_t flags, bool criu_resume); int amdgpu_amdkfd_gpuvm_free_memory_of_gpu( struct amdgpu_device *adev, struct kgd_mem *mem, void *drm_priv, - uint64_t *size); + uint64_t *size, bool force); int amdgpu_amdkfd_gpuvm_map_memory_to_gpu(struct amdgpu_device *adev, struct kgd_mem *mem, void *drm_priv); int amdgpu_amdkfd_gpuvm_unmap_memory_from_gpu( diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c index 1e71829e0fc6..dc1fa664feca 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c @@ -1901,7 +1901,7 @@ int amdgpu_amdkfd_gpuvm_alloc_memory_of_gpu( int amdgpu_amdkfd_gpuvm_free_memory_of_gpu( struct amdgpu_device *adev, struct kgd_mem *mem, void *drm_priv, - uint64_t *size) + uint64_t *size, bool force) { struct amdkfd_process_info *process_info = mem->process_info; unsigned long bo_size = mem->bo->tbo.base.size; @@ -1922,9 +1922,26 @@ int amdgpu_amdkfd_gpuvm_free_memory_of_gpu( */ if (mapped_to_gpu_memory > 0) { - pr_debug("BO VA 0x%llx size 0x%lx is still mapped.\n", - mem->va, bo_size); - return -EBUSY; + /* + * Refusing to free a mapped BO is only meaningful while the + * process can still unmap it. On process teardown (@force) + * there is no such chance: the caller drops the last handle + * to @mem regardless, so bailing out here leaks the BO onto + * process_info->kfd_bo_list / userptr_inval_list. Those lists + * are then destroyed non-empty in + * amdgpu_amdkfd_gpuvm_destroy_cb(), leaving a BO in TTM's + * eviction LRU whose bo_vas point into the freed amdgpu_vm. + * The next client to trigger eviction deadlocks in + * amdgpu_vm_bo_move(). Tear the mappings down instead - the + * VM is going away right after us anyway. + */ + if (!force) { + pr_debug("BO VA 0x%llx size 0x%lx is still mapped.\n", + mem->va, bo_size); + return -EBUSY; + } + pr_warn("Force-freeing BO VA 0x%llx size 0x%lx still mapped %u time(s)\n", + mem->va, bo_size, mapped_to_gpu_memory); } /* At this point the BO is guaranteed to be freed, so unpin the diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c index 6fd18488d5cf..0121c6dc77b3 100644 --- a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c +++ b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c @@ -1219,7 +1219,7 @@ static int kfd_ioctl_alloc_memory_of_gpu(struct file *filep, err_free: amdgpu_amdkfd_gpuvm_free_memory_of_gpu(dev->adev, (struct kgd_mem *)mem, - pdd->drm_priv, NULL); + pdd->drm_priv, NULL, false); err_unlock: err_pdd: err_large_bar: @@ -1262,7 +1262,8 @@ static int kfd_ioctl_free_memory_of_gpu(struct file *filep, } ret = amdgpu_amdkfd_gpuvm_free_memory_of_gpu(pdd->dev->adev, - (struct kgd_mem *)mem, pdd->drm_priv, &size); + (struct kgd_mem *)mem, pdd->drm_priv, &size, + false); /* If freeing the buffer failed, leave the handle in place for * clean-up during process tear-down. @@ -1618,7 +1619,7 @@ static int kfd_ioctl_import_dmabuf(struct file *filep, err_free: amdgpu_amdkfd_gpuvm_free_memory_of_gpu(pdd->dev->adev, (struct kgd_mem *)mem, - pdd->drm_priv, NULL); + pdd->drm_priv, NULL, false); err_unlock: mutex_unlock(&p->mutex); return r; @@ -2483,7 +2484,7 @@ static int criu_restore_memory_of_gpu(struct kfd_process_device *pdd, if (idr_handle < 0) { pr_err("Could not allocate idr\n"); amdgpu_amdkfd_gpuvm_free_memory_of_gpu(pdd->dev->adev, *kgd_mem, pdd->drm_priv, - NULL); + NULL, false); return -ENOMEM; } diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_process.c b/drivers/gpu/drm/amd/amdkfd/kfd_process.c index 0a7c1900da95..6d5126aa6fe7 100644 --- a/drivers/gpu/drm/amd/amdkfd/kfd_process.c +++ b/drivers/gpu/drm/amd/amdkfd/kfd_process.c @@ -743,7 +743,7 @@ static void kfd_process_free_gpuvm(struct kgd_mem *mem, amdgpu_amdkfd_gpuvm_unmap_memory_from_gpu(dev->adev, mem, pdd->drm_priv); amdgpu_amdkfd_gpuvm_free_memory_of_gpu(dev->adev, mem, pdd->drm_priv, - NULL); + NULL, true); } static void kfd_process_free_gpuvm_map(struct kgd_mem *mem, @@ -759,7 +759,7 @@ static void kfd_process_free_gpuvm_map(struct kgd_mem *mem, amdgpu_amdkfd_gpuvm_unmap_memory_from_gpu(dev->adev, mem, pdd->drm_priv); amdgpu_amdkfd_gpuvm_free_memory_of_gpu(dev->adev, mem, pdd->drm_priv, - NULL); + NULL, true); } /* kfd_process_alloc_gpuvm - Allocate GPU VM for the KFD process @@ -814,7 +814,7 @@ static int kfd_process_alloc_gpuvm(struct kfd_process_device *pdd, err_map_mem: amdgpu_amdkfd_gpuvm_free_memory_of_gpu(kdev->adev, *mem, pdd->drm_priv, - NULL); + NULL, false); err_alloc_mem: *mem = NULL; *kptr = NULL; @@ -1119,18 +1119,29 @@ static void kfd_process_device_free_bos(struct kfd_process_device *pdd) * local memory object */ idr_for_each_entry(&pdd->alloc_idr, mem, id) { + int r; for (i = 0; i < p->n_pdds; i++) { struct kfd_process_device *peer_pdd = p->pdds[i]; if (!peer_pdd->drm_priv) continue; + /* + * This can fail under memory pressure: unmapping has + * to validate the page-table BOs first. Ignore it and + * force the free below - the handle is dropped either + * way, so a failed free would leak the BO onto the + * process_info lists and poison the eviction LRU. + */ amdgpu_amdkfd_gpuvm_unmap_memory_from_gpu( peer_pdd->dev->adev, mem, peer_pdd->drm_priv); } - amdgpu_amdkfd_gpuvm_free_memory_of_gpu(pdd->dev->adev, mem, - pdd->drm_priv, NULL); + r = amdgpu_amdkfd_gpuvm_free_memory_of_gpu(pdd->dev->adev, mem, + pdd->drm_priv, NULL, + true); + if (r) + pr_err("Failed to free BO on process teardown: %d\n", r); kfd_process_device_remove_obj_handle(pdd, id); } } -- 2.55.0