From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BABFE41A77C for ; Mon, 10 Aug 2026 16:28:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786379295; cv=none; b=QqilvfdES1y1DfOU1m3va7cliD+bG225Za4XrNlZ3aT9XHgEOqTrWTjnmovqd426ONzP9Or9TIz3SqqMd6BqeXat1qsZKy/XmWbnQGtZTxolavt5WjvD33ura5B28f3VIZXtShufrRa6hfhIg6phCwMhdPLaiApUzcw2Wxa1TPE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786379295; c=relaxed/simple; bh=31/7dsk0IdNCvznLHGQx3neWHmQouuUYea0saDN67NY=; h=Date:From:To:Cc:Subject:Message-ID:MIME-Version:Content-Type; b=nRxBRTkO+qx8W81F3QbuAqCis/XgD/Ot3yJ2QycfHbwFI3G78kQTCeHgKBwiyLIFHeie0GEPpXO38zuk6q48ldMmo1UatKMFFODVnPCZ/PpM6rwKmHtMgBDst8VuVhdxk7RmzisUq4LrGIEmnyD2bt9eQ17V0WD6zQFChj7UK1g= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=BuqO6uAK; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="BuqO6uAK" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:Content-Type:MIME-Version:Message-ID: Subject:Cc:To:From:Date:Sender:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: In-Reply-To:References:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=EANdORJlqyqj5ePn4ATBwyvNejleE/hNFHl3jIiiuyo=; b=BuqO6uAK48+UcGiJ3utN+DEA8n UsFiOfuSls7t4wj+LY291I5PwI7KM361FqAYM7LlSxvfna4Gh12i4UL8VkqpdBnzVpUPcSmeb8eME zBAKhzamnYXVqxbNdKoeYNgsvFka24ApTFXDzZb2ocWWEJdrnxYefCHDDFO2EZ8K/jg3KxDMssac8 wOk2Unsunhg1h6R70QtRAMZcUkWwKA05AIDoTdTliz+182f16U1c33g/SBwSxFL+0Y53S6kmQXbQC e3/lotKxcDf2tQB5YH6+rcJiDZSD/Bmihy6w9uNyQwf7oqpSJdiqE7HgOQm/A4xH1rnWhrNDa6Q78 K74bYrBw==; Received: from [2601:18c:8100:a0e0:5a47:caff:fe78:8708] (helo=fangorn) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtSqg-000000006jU-1WbH; Mon, 10 Aug 2026 12:27:38 -0400 Date: Mon, 10 Aug 2026 12:27:37 -0400 From: Rik van Riel To: Andrew Morton Cc: David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Mason , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [RFC v2 PATCH] mm/cma: don't release CMA pages still in use Message-ID: <20260810122737.030f8452@fangorn> X-Mailer: Claws Mail 4.4.0 (GTK 3.24.52; x86_64-redhat-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit When a driver calls dma_free_contiguous() before quiescing DMA, the page still has a reference from the device. put_page_testzero() there returns false, WARN fires, but the code proceeds to free_contig_frozen_range() putting a live page onto buddy and clearing the bitmap. Later cma_alloc() hands the same PFN to a new owner while the original holder still references it. A concurrent put_page() that drops the last reference between the testzero loop and free_contig_frozen_range() can double-queue the page via page->lru, corrupting buddy lists. Fix by freeing already-frozen pages in contiguous runs, while keeping still-referenced pages. When any live page exists, a future cma_alloc() on this address range will fail until whoever holds the extra references frees those pages. This change should be safe because nothing can get reallocated while it is still in use. Fixes: 9bda131c6093 ("mm: cma: add cma_alloc_frozen{_compound}()") Cc: stable@vger.kernel.org Assisted-by: Hermes:muse-spark-1.2 syzkaller Reported-by: Chris Mason Signed-off-by: Rik van Riel --- v2: - don't play clever games with the CMA bitmap, the page allocator alone will prevent re-use of not-free-yet pages - v1: https://lore.kernel.org/all/20260809210608.06b5ccb9@fangorn/ mm/cma.c | 39 +++++++++++++++++++++++++++++++++------ 1 file changed, 33 insertions(+), 6 deletions(-) diff --git a/mm/cma.c b/mm/cma.c index a13ce4999b39..cbf8dba8f077 100644 --- a/mm/cma.c +++ b/mm/cma.c @@ -998,7 +998,6 @@ static void __cma_release_frozen(struct cma *cma, struct cma_memrange *cmr, pr_debug("%s(page %p, count %lu)\n", __func__, (void *)pages, count); - free_contig_frozen_range(pfn, count); cma_clear_bitmap(cma, cmr, pfn, count); cma_sysfs_account_release_pages(cma, count); trace_cma_release(cma->name, pfn, pages, count); @@ -1018,18 +1017,45 @@ bool cma_release(struct cma *cma, const struct page *pages, unsigned long count) { struct cma_memrange *cmr; - unsigned long ret = 0; + unsigned long skipped = 0; unsigned long i, pfn; + unsigned long base_pfn; + unsigned long run_start = 0; + unsigned long run_len = 0; cmr = find_cma_memrange(cma, pages, count); if (!cmr) return false; - pfn = page_to_pfn(pages); - for (i = 0; i < count; i++, pfn++) - ret += !put_page_testzero(pfn_to_page(pfn)); + base_pfn = page_to_pfn(pages); + pfn = base_pfn; + for (i = 0; i < count; i++, pfn++) { + if (put_page_testzero(pfn_to_page(pfn))) { + /* Add it to the batch. */ + if (run_len == 0) + run_start = pfn; + run_len++; + } else { + /* + * This page is still in use! Free the freeable + * pages encountered so far, but skip this page. + */ + if (run_len) { + free_contig_frozen_range(run_start, run_len); + run_len = 0; + } + skipped++; + } + } + if (run_len) + free_contig_frozen_range(run_start, run_len); - WARN(ret, "%lu pages are still in use!\n", ret); + /* + * Some pages were still in use! This should not happen. + * Subsequent cma_alloc() calls to the same range will fail + * until whoever grabbed the extra refcounts frees the pages. + */ + WARN(skipped, "%lu pages are still in use!\n", skipped); __cma_release_frozen(cma, cmr, pages, count); @@ -1046,6 +1072,7 @@ bool cma_release_frozen(struct cma *cma, const struct page *pages, if (!cmr) return false; + free_contig_frozen_range(page_to_pfn(pages), count); __cma_release_frozen(cma, cmr, pages, count); return true; -- 2.55.0