From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f179.google.com (mail-yw1-f179.google.com [209.85.128.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F95A25B0AB for ; Sat, 22 Aug 2026 15:11:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.179 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787411465; cv=none; b=keZfdApS2wowU9z9dKzUnH5rMZxYTQI0MCSbm9HpTkWHxgYwbC1vGQEEQWF/SxSf0SYNmRJpKXTRpf24c0O319wEMC+olBglEnIw1NLunFfYvKbUxLx653NzonGzG4s4zTPaqty8JV8jWvjwN1mpuC8Jd+OAy7W7v3GH1u170dU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787411465; c=relaxed/simple; bh=N3auuA2VSOZEVE+h5AXIcxVbUiADEtJyA9c739nlDn0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=axWxNTOshFIiYGF4qXUhMGS7RCOC1oODuaIuAWK+gOcoGEatS1jFQu8sTRBOD0gr3TnuTBbVuEx1EzcIxuFcsZSGgo+MNuxzmvDAI8wNButLviRs1jT9uFg9x14KRsSJ2IzuCfgduIeBNuP5oYJXlOgACK/BzWUxgRqsvoboq5Y= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=D/ZK1DYK; arc=none smtp.client-ip=209.85.128.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="D/ZK1DYK" Received: by mail-yw1-f179.google.com with SMTP id 00721157ae682-81f3b227a4aso32626027b3.1 for ; Sat, 22 Aug 2026 08:11:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787411461; x=1788016261; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tl+ydfXGpsApHyYdqtis4MbvgFxpMIFRVTomZ0enFKw=; b=D/ZK1DYKyjSt2UOR4M1ucvSnFC+n/1+YNNzAbzz5qd5aeCJfFYO+NpnucVNk0PHtPQ e1rihZ8TKTjjN6L0Go3opTSw/uF5CR/9Io6ucyd4KX8HKyk4kG1P1/OYKfeG9a1v/IMD LVnt19XlCB9OZSDGXTSzrtx2AbHOZgt2KEM5rg54HVxpxeThkIGjtt55KrvQRmfmh9iA 4EeOEepoDvEm0soLIIXI/yYtsZHXhEyQBN/qZzOZIDK+dHM6hYIv9lrAajoX9KF6Lo+P /2F5KtsX68+WJtvgZgdVtBJlJnK+IoLvtXGP86s5vAacLINcUKut6QQWlBr7i/SUicRU GwcQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787411461; x=1788016261; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=tl+ydfXGpsApHyYdqtis4MbvgFxpMIFRVTomZ0enFKw=; b=UEVu45CiBsw+VZbyUzCtLOSADdb28B8aHHUdTt7PBOruFon5FKNKIXo1WxzTq7J40e mhCsyavEAAeNnPiqhMquTHcLlNqFsdSiN2XlolyZ6+gSjPjT6wha7bRhw5Ay4Lt0SpTA 65TAiX26YjTthuEtHhmUMfLiU6vcOxeIAwtszJaTxVZtxN7KHl+lajgaKfnnueFtXrg7 h8sOL2X1jQd3S2+hkEc0I4zpDfp56Z/HwncPShzrPPJYQM4Xc1mOqZVtrI6b96JQJnZb PajCpdSRNjFsVfhL+BU/kE5lzXP9oYOmA/12Zmh9FQ68RZ5B4YZQcyIcIGFBtlGpzLPX H3fQ== X-Forwarded-Encrypted: i=1; AHgh+RogSFQz4/pEnN6FElwF3YilSjXiBe0XXXy4ie0jZVqY4WJiGLik+04EPIAExozk42DoJ2CdzjnHPXvADxY=@vger.kernel.org X-Gm-Message-State: AFuF++kghifDhloG8yMlV1Qizig9qrVybbV8kXCoDuKXwT6hhBn04SD4 2UQzofw5djR5lnFycW5brIrSva9BDhhtUTt9ByTyXh/TB0vt7m9Gfd/k3Y8rstNYSxE= X-Gm-Gg: AR+sD13mzfDwK9lWtavQNcSvUSPAttLk4ys0lChrmc2r6vHSJHC6KNJLCZmZyEuA55j kf3sd9hn3jvFtLt2rOz8ILCqacH7rtWdIhfa/eU7qw+52uFZ48TTF5Ex5bGap0Znc1Xhm+4IE1C 3+0Gp8rSTJjrFLGZHVlGOPvSjycFJfgvoWaogu/pbkxj0R8c06AlROCNPIWG4y7P7JpNnD8pRSp rysHNxDn7ExSKLGdRMObm2kSvZV+X1nA99pvzbJRF3guRJBrvHSLYdX3/qTzSEbB7v+2kh30UNs laOvEpTSCAtficJqam2af3F1YE2I2TCkwwz1pt80OpNwTGlsokKT3awFLyCKqfhv/anil0ulmP5 eZtY5Rjl8AczsmzTQMq4sA/FURJfhZUmWHDikGXUTFMHXqS1J+f6+DJFStafP42hck9ZJ/30y/q LavQqPU2/aCqSy3mwrbnv1Yw9u8lzRlHSbxodINcnbtFQ4+LBLjMftZRHYr20PRHr4aKQ7c4C/A HjWIiL7J9dNVdxsnECUzQ2vOkAalcbmvleS0kDsiMEj3Fztj184BA5XL3qjKqaKO++05XnHNIaS ebQ= X-Received: by 2002:a05:690c:7001:b0:82e:3aea:fa5d with SMTP id 00721157ae682-849f245444cmr61904547b3.15.1787411460944; Sat, 22 Aug 2026 08:11:00 -0700 (PDT) Received: from DigitalOasis.tail29fcbc.ts.net (104-182-93-176.lightspeed.brhmal.sbcglobal.net. [104.182.93.176]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84cac01026csm11076757b3.39.2026.08.22.08.11.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 08:11:00 -0700 (PDT) From: Bocaj Gnuoy To: amd-gfx@lists.freedesktop.org Cc: Felix.Kuehling@amd.com, alexander.deucher@amd.com, christian.koenig@amd.com, airlied@gmail.com, simona@ffwll.ch, oded.gabbay@gmail.com, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, Bocaj Gnuoy Subject: Re: [PATCH v2] drm/amdkfd: don't leak BOs when process teardown can't unmap them Date: Sat, 22 Aug 2026 10:10:46 -0500 Message-ID: <20260822151046.966976-1-bocajgnuoy@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260822144402.929677-1-bocajgnuoy@gmail.com> References: <20260822144402.929677-1-bocajgnuoy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Correcting a factual error in my own changelog before anyone spends review time on it. The changelog states: "Note the free path itself does not allocate - it reserves the BO, detaches the attachments and drops the references - so it can complete even when the unmap could not." The first clause is wrong. Immediately after the -EBUSY bypass, amdgpu_amdkfd_gpuvm_free_memory_of_gpu() calls: ret = reserve_bo_and_cond_vms(mem, NULL, BO_VM_ALL, &ctx); if (unlikely(ret)) return ret; reserve_bo_and_cond_vms() uses drm_exec_prepare_obj(), and drm_exec allocates its object array with kvmalloc_array()/kvrealloc(GFP_KERNEL) and returns -ENOMEM on failure; drm_exec_prepare_obj() also calls dma_resv_reserve_fences(). So the free path can allocate, and can fail with -ENOMEM. Bounded consequence ------------------- The forced free can therefore still fail, after which kfd_process_device_free_bos() drops the idr handle regardless - the same ownership violation the patch addresses, one level deeper. This is not detectable through the teardown WARNs. The BO is removed from process_info's lists before the reservation is attempted: /* Make sure restore workers don't access the BO any more */ mutex_lock(&process_info->lock); if (!list_empty(&mem->validate_list)) list_del_init(&mem->validate_list); mutex_unlock(&process_info->lock); ret = reserve_bo_and_cond_vms(mem, NULL, BO_VM_ALL, &ctx); if (unlikely(ret)) return ret; That ordering is deliberate, so reordering it is not a fix. The consequence is that if the reservation fails, the attachments are never detached and the free does not complete, but the lists that amdgpu_amdkfd_gpuvm_destroy_cb() checks are already empty. The stale lifetime condition implicated in the deadlock can therefore survive without the WARNs firing. I have not driven that particular residual failure through the eviction path and observed the deadlock, so I am describing a reachable state, not a reproduced one. Scope of the change ------------------- The underlying violation is not introduced by this patch. Upstream today can already reach it: mapped_to_gpu_memory == 0 -> delist -> reservation fails -> free returns error -> teardown discards the handle This patch additionally permits: mapped_to_gpu_memory > 0 after a failed unmap -> bypass -EBUSY -> delist -> reservation fails -> free returns error -> teardown discards the handle So it widens the set of states that can reach the violation rather than creating it. That is not offered as a justification, only as scope. What the testing does and does not show --------------------------------------- The pr_err() this patch adds to kfd_process_device_free_bos() is the only instrumentation in the patch that directly observes a post-delisting free failure. It did not fire in any run, including one with 1904 "failed to validate PT BOs" and 35 forced frees. That is consistent with the drm_exec allocations being far smaller than the page-table validation that failed, and therefore satisfiable under the same pressure. It is a probability argument, not an invariant. The accurate evidentiary statement is: the fix reliably completed in the tested pressure regime, but the implementation does not provide an invariant guaranteeing teardown completion under arbitrary allocation failure. What this does and does not change ---------------------------------- The reproduced failure is unaffected: the unmap fails in vm_validate_pt_pd_bos(), mapped_to_gpu_memory stays non-zero, the free returns -EBUSY before list_del_init(&mem->validate_list), and the BO is stranded. This patch removes that, and in every observed forced teardown the cleanup completed and the machine-killing relaunch stopped happening. What the error invalidates is the completeness claim, not the safety argument for forcing. Bypassing -EBUSY on irreversible teardown and detaching the bo_vas when the reservation succeeds is still supported by the evidence. What I can no longer claim is that the cleanup is guaranteed to succeed merely because the failing page-table operation was skipped. v3 will correct the changelog and state this as an explicit limitation. If the residual allocation failure should be repaired rather than documented, there appear to be several possible directions with different locking and lifetime implications - a retry, a reservation path that does not allocate, or deferring ownership so a later pass can free the BO, among others. I would rather have maintainer guidance on which is wanted than guess at a lifetime-ownership decision. Thanks, Bocaj