From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-lj1-f180.google.com (mail-lj1-f180.google.com [209.85.208.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 485403EC824 for ; Thu, 30 Jul 2026 09:06:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.180 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785402415; cv=none; b=JMfe8Hd7GBDHXnoPogCPUcG2hEpXkRbU/eV//tb9K4ii301W3VOwXimBQQPCFFM80VwTb4kmP+VZqG/o0Xk8JdqJGTgjXqqvgwpoqRJmOOdDok0DoGlGrnPlGHWdzgMbxIujatUDGxrKuto2m0H1hgabb78z5RlfVtJUH/wk/J0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785402415; c=relaxed/simple; bh=ADFJtPWTPq536A8adbKhc7ZThBoqYBRrVdtbhXCP/4Y=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cxDk9PI0FihLhG1NQMWFy6v6af5i5vQjwpF4J3/fmB9+97GsZ8xf8uk5Ip3w/Y6IjGa01FWUJdIbiAy4qN3CLmQf4byJSHdNOf7CG3ta8OG35tQezW6EZE4bUunSCZdo0PfNOlM1MRN5FBlmlza51/L1K+tBHLJM27L2IZVwsbM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=V0kHIQ4B; arc=none smtp.client-ip=209.85.208.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="V0kHIQ4B" Received: by mail-lj1-f180.google.com with SMTP id 38308e7fff4ca-39c83acb86eso15734751fa.1 for ; Thu, 30 Jul 2026 02:06:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785402411; x=1786007211; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HdEg3D4TKN236ca6jYDlc0P5L4Wo24cZPSMmhVdO09s=; b=V0kHIQ4BvTaKOgsOSuGo5TrOpb3jQW0N75v5FNjaqbQJfwpBEN5ZSjTt4gIKsQ8/a+ DAe4YqZR0UrYki5BsB4eYtlsBV1ckEKvbCAj+GNqbUZUjqVIPCB20BWipy08g74JJfv8 AuHKnZpaoV8flGbt5aID9zYCdFgVJR6E9iAekT94st9gNUag9H0gCwd0lQLGbBxwbzFT VtPCi8LiuruHZDnnYQSaZ8ex9CPpsQ/Flhvu/Gv7TYhCOg2x/LeDNr8JT1OiI9Y2BGfY as2yF7wMm0fVSLyoq/4DZ+x4y3xc+kJI0SC9G1gyAqP3N7A9PnC6FEByyfRihmqVkW3W pegQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785402411; x=1786007211; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=HdEg3D4TKN236ca6jYDlc0P5L4Wo24cZPSMmhVdO09s=; b=tY89dvm7SrNmMucUJyTnd2jXJ6BaDv1+nsNFI/3NAVhGTZOMEK8M13VKn0sUQS2LoG tZgY7CaMvmJPb+cLKyLah4E8tqWAwFK04G/PlD8wXYjdn7QNowAVEKoGzGlozKbPUfpf /dRBfkBVXlVwuPqqzsgtvXg3dsF6JwqMhVX0Uxphi5I0a09eZ2mBUXY4qfBkoGrZlp/4 pA0bx90HNLmMrf/R/rUIUNxRkJPjJtC8Snx7kB3GHKFrEe9cBnPyk5mnpooJQ6lwICJ8 hQKunPzCHWSGDFo96v8NkOPU/EwwiGdaBRu+dpiBfHdN9ddu+93pfrP9skDjzl5yFcE5 34xg== X-Forwarded-Encrypted: i=1; AHgh+Rqz2KNBS/ty17Samz4o+Bs9znHMVy5KhZRIxHzSjFGKJngvzS10XX7jPt13h/XiUrculBxbFfbaJjnJOiQ=@vger.kernel.org X-Gm-Message-State: AOJu0YxwjGppYHA0+gPgNG1mBKz8yu43YGEmsdQtEEzgpqYYNpgjjp34 cuBGgwJGMtNjj2ftACPQBEK95AgRsZwOKCUHCBYA8dwPkIxjcIGhG2I9 X-Gm-Gg: AR+sD11teznPgA1Psiep+wFtxqW7+ENh4xbO/cpPOvJDczDXGmR0jYYUP9VpgkRkBiO V+Q5beHGs647GLkJDZQx8DCm5KFYgPkkMKicYvCz+JrUt3aDmmRkqT7slkAf1edeMocyxjlkuvM wUwF1Yo322kEpjYafOaB1R6nAk0J9VLDpB1HL4yVUsUdzvDZjd8rgc7krAFjcIse8Iks/HhnoEo XPNvQshX1ZuEno27bMYs57hyuwYlUoKZI39TKCLkGaR+jesq/BCXdDX7HmD+ft7/1ibH1sDeWEs G0xP8+/25FM6/3tfedYso96HVIXJlNrv1fiHgJe9wYee4hXYoBqDj8qbKYRdAfOLiFQZiFo0wEP CxyOWbrjq3gAyjlRdtGDmVt0/zqEhoSsvXiV1QiyRFHc3RCWLECAxknwXn6b/PtWwPI1c3aFvjW 7aEdTaS+SAgmNg6rA1sfR75jRSuNhQvCi5e0fuFxQmh29XTp07lOoPXb3OpZRZ2r+vvMuLAqrZd CeNzcGlmss3K9+F/OvwxYljsqtqv4vsIGfNe8FSDSCTHg0DXzQGF082gQ== X-Received: by 2002:a05:651c:18c6:b0:39c:a346:8a36 with SMTP id 38308e7fff4ca-39f6d0fbf2fmr3486971fa.8.1785402411014; Thu, 30 Jul 2026 02:06:51 -0700 (PDT) Received: from localhost.localdomain (46-138-176-102.dynamic.spd-mgts.ru. [46.138.176.102]) by smtp.gmail.com with ESMTPSA id 38308e7fff4ca-39f6ac80d6bsm1928641fa.29.2026.07.30.02.06.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 30 Jul 2026 02:06:50 -0700 (PDT) From: Artem Lytkin To: linux-mm@kvack.org Cc: akpm@linux-foundation.org, urezki@gmail.com, willy@infradead.org, shivamkalra98@zohomail.in, linux-kernel@vger.kernel.org Subject: [PATCH v2 1/3] mm/vmalloc: fix 32-bit truncation of the area size in vread_iter() Date: Thu, 30 Jul 2026 12:06:26 +0300 Message-ID: <20260730090628.65814-2-iprintercanon@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260729175708.7074-1-iprintercanon@gmail.com> References: <20260729175708.7074-1-iprintercanon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Commit 0bca23804632 ("mm/vmalloc: use physical page count in vread_iter() for VM_ALLOC areas") replaced get_vm_area_size(vm), which returns a size_t, with vm->nr_pages << PAGE_SHIFT. struct vm_struct::nr_pages is an unsigned int. The shift operator does not perform the usual arithmetic conversions: the integer promotions are applied to each operand and the type of the result is that of the promoted left operand. The expression is therefore evaluated in 32-bit arithmetic no matter how PAGE_SHIFT is typed, and no matter that the result is assigned to a size_t. Once an area reaches 4 GiB the byte count wraps, at 1 << 20 pages with 4 KiB pages, 1 << 18 with 16 KiB and 1 << 16 with 64 KiB. nr_pages counts base pages even for huge vmalloc mappings, since __vmalloc_area_node() requires area->nr_pages == nr_small_pages, so a huge page_order does not raise the threshold. The only limit on the size of a single area is the totalram_pages() check in __vmalloc_node_range_noprof(), so any machine with more than 4 GiB of memory can create an affected area. One example is the zram metadata table, which is a single vzalloc() of disksize / 256 with CONFIG_LOCKDEP=n: "echo 1T > /sys/block/zram0/disksize" allocates exactly 4 GiB and needs no 1 TiB of anything. Sufficiently large BPF array or prealloc-hash maps do it too, as does "modprobe test_vmalloc run_test_mask=1 nr_pages=1048576". The KVM dirty bitmap is one path that might be expected to reach it but cannot, being bounded by KVM_MEM_MAX_NR_PAGES to 512 MiB. alloc_large_system_hash() can reach it, but not on its own: it caps a table at a sixteenth of memory, which is 4 GiB once a machine has 64 GiB, and the automatic sizing stays orders of magnitude below that cap, so it takes an explicit table size on the command line. The effect is confined to /proc/kcore, the only caller. For an area whose size is an exact multiple of 4 GiB the computed size becomes 0, the if (addr >= vaddr + size) goto next_va; test succeeds and the whole area is skipped; for other sizes only the first nr_pages mod 2^20 pages are read and the rest is skipped. Either way the bytes are zero-filled at the finished_zero label, which returns the full requested length, so read_kcore_iter() sees a successful read and userspace gets neither an error nor a short read. Live inspection with drgn, crash or gdb silently observes zeros where the area is populated, and cannot distinguish that from genuinely zeroed memory. Before the offending commit the size came from vm->size in 64-bit arithmetic, so this is a v7.2 regression. The truncated value is always less than or equal to the true size, so vread_iter() can only under-read; there is no out-of-bounds access. Both the affected reader and every producer above are privileged: opening /proc/kcore requires CAP_SYS_RAWIO and is refused under lockdown. This is a correctness and debuggability problem, not a security one. Fix it by widening the shift, which also makes the expression consistent with the four (unsigned long)nr_pages << PAGE_SHIFT expressions in vrealloc_node_align_noprof(). On 32-bit a widening cast cannot help, size_t being 32 bits there as well, but a 4 GiB vmalloc area is not reachable on 32-bit either. On 64-bit the cast removes the truncation entirely, which is why replacing get_vm_area_size() introduced a regression rather than inheriting a pre-existing wart. Fixes: 0bca23804632 ("mm/vmalloc: use physical page count in vread_iter() for VM_ALLOC areas") Assisted-by: Claude:claude-fable-5 Signed-off-by: Artem Lytkin --- mm/vmalloc.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 26f32949c2f2e..34e10b825889a 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -4895,7 +4895,7 @@ long vread_iter(struct iov_iter *iter, const char *addr, size_t count) * mapping types (vmap, ioremap) don't set nr_pages. */ size = (vm->flags & VM_ALLOC && vm->nr_pages) ? - (vm->nr_pages << PAGE_SHIFT) : + ((unsigned long)vm->nr_pages << PAGE_SHIFT) : get_vm_area_size(vm); else size = va_size(va); -- 2.43.0