From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-lf1-f42.google.com (mail-lf1-f42.google.com [209.85.167.42]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 122003A48F6 for ; Wed, 29 Jul 2026 17:57:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.42 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785347863; cv=none; b=Of4EnJCouVMxxas+vUTMOQFW9MxIAWxWQj6eDzhk4fCF2/CjFNWi+MG5+RlBFUqx8WQ4vBp1VE2vzKZn8fZhwbvk314jVe4fpy1M+jF+ypKL6XqI/cTQ/9D3JzmxYQ8cvmSIDjaJj404q76SGF0kUdBmK7vzx8iN1IGDmxC+uLI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785347863; c=relaxed/simple; bh=fXBIqqjs77/m031nsB0oHCpk2vfHh+VgpeVzZSzbGCI=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=n2vht4a/w8Clwq0UzqNJ05jgF4L6Fg1J0trcdb2uX9sv7eil+Ze8SOIdw694iq/Vl2GG85QQLtGtOs8w9c+1tSBgZ9Ox0QJgz0s3ei89W9BoCqsJ6A74wEQBKlxFKRV6tnPzZ8D+8atmhzWiZiFrhgYO7PkN+GtvDkiEB9niAcc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=cfizVtGL; arc=none smtp.client-ip=209.85.167.42 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="cfizVtGL" Received: by mail-lf1-f42.google.com with SMTP id 2adb3069b0e04-5b015b2d792so1220498e87.3 for ; Wed, 29 Jul 2026 10:57:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785347850; x=1785952650; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=mhSA5XspJ7J1QMTT+GAfybOgMSS1170SbWNYCpn86sc=; b=cfizVtGL9HPP3ke1g2Vt78Xk9JGHpB53nhlfqdNDdvZ3UKsHtwVY6z5iDDhomB2dJZ gT21iQ4K/JarZO4dpiAremSPurkl/vFTGNSGSIyO84kt/0j2mNuwPqar0L2trP+hfqKn 0v3cEWrX5sqHFnfDkW2bAwpQTEEfGS6p0rXSXxPHJ3mD7nbYArqWStEFQ4F7zuql4SEs m06vLBCWu+gHyfcMQsMMbz68xPiBiIHS+dKX3/xrnUMDZQD09++L7NrHj8DE/GHyue3/ vExGec3CZplML+7gPxW3wJFolP5cnI2phwhjgr9IlxPdyJ81T4Z9jE8aB/TlZvHZnQvl u7Eg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785347850; x=1785952650; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=mhSA5XspJ7J1QMTT+GAfybOgMSS1170SbWNYCpn86sc=; b=rYJsijnyutO4M5T6/x59sScf2tC00gT+COj1Zixi+gJM3bJz9256o2KF1rjoMACznu CYtBlphgw2PyZEr4MpnWUfHRhyLid3vU5vG6vvr5A2JwyKb07/zY6bNRWGrrPgoYAo0c Xwo5r32kiW1FmRHKLPBIoN/s1VCeZS5bZDtH0rFCZ5XSE3RlnyFwqckpkREXszrfYgbu 0ze08jVrb+OHGziFdG1e3QAxGVZmyjISbcWdklVUHhM6SOCoTgoJfEtOPvU/KAlXO8CL QBC2NwG3bRUEGB023cPaLgj53RmvvQEbpB12nzw09JiQIoqxtWTVbs3LYoYFtwclnY9J XOSw== X-Forwarded-Encrypted: i=1; AHgh+Rr7LkjAqFRUf+yI2d43DFkmvJwK+7vnv2OEgLIKviOjmXmPpVUqZiRoCFoxhlLyvhPUF0CRWbxsJF5UOHs=@vger.kernel.org X-Gm-Message-State: AOJu0Yy8v3pyB3/n5i5IvA2waifJhpG27GzskiatRqWq3Fww6K934YB/ JnD2hA9sHaui7NFi95tVtjvnmg9GX5OYmctKJtL5/iVKdSw6hArF3kEG X-Gm-Gg: AR+sD10gAXE1hu/xnW0O+bZOBy67SgJsJOgm8VtZt5vbP8Xbm+lUkflJ6MCOQ0lUabc pgPJ3xLW2MS4v805JxOKngf/OzMTli32/lvUqOd2/pLW9WDbUvqxxmPb1P5X+z094rm2561S2Cl 9Zj9oYbooPxrAWnoceR89DbDjSLjFgYOka+EPF2hSlJNP5rpd6N5/53nq3Rjv0VGnUndJLcSOwh TqjeMtte7YsAYrSx3e55IIFE4wFZoyMvPLi7rVB+8SQdMw1Uu5nmd/Oroz2GVPiZovKF6iuctQE l1b2MnRYE3KvhwzFpL82wQ7/U/BQO31Hh2xWz64JdRLK+zRaK832KHJDx77DsyR7tDEcjXOFSQr ReYqmcnk2DJZwe45esEcRD8MKaWPkawIEZFmcjZ38Aq2cVEJgFCAXdVwgmD5z9yZO2GHzGbELKz zWY/C8vXnVqrtKULgcudssZuI2B0FVBaYEQk1vhs5ak2zyCuXU24TjInN+sr8DvSh/BNnp01DeQ nag37byW63qfVY2czTI6HJ3lIoJGOqGRIl61FnuEFwRE8s38n6/VqS4NA== X-Received: by 2002:a05:6512:3d21:b0:5b1:568d:971 with SMTP id 2adb3069b0e04-5b2d025dc53mr1561787e87.49.1785347849754; Wed, 29 Jul 2026 10:57:29 -0700 (PDT) Received: from localhost.localdomain (46-138-176-102.dynamic.spd-mgts.ru. [46.138.176.102]) by smtp.gmail.com with ESMTPSA id 2adb3069b0e04-5b2d78eaa08sm507610e87.54.2026.07.29.10.57.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 29 Jul 2026 10:57:29 -0700 (PDT) From: Artem Lytkin To: linux-mm@kvack.org Cc: akpm@linux-foundation.org, urezki@gmail.com, shivamkalra98@zohomail.in, linux-kernel@vger.kernel.org Subject: [PATCH] mm/vmalloc: make vm_struct.nr_pages an unsigned long Date: Wed, 29 Jul 2026 20:57:08 +0300 Message-ID: <20260729175708.7074-1-iprintercanon@gmail.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit A vmalloc area can hold more than 2^32 pages, but the page counts tracking it are 32 bit. That has produced the same truncation bug twice already, once in vread_iter() and once in the vrealloc() grow-in-place check, and it keeps the casts that hide it scattered around the file. Widen the field and everything that feeds or consumes it, so the byte counts derived from it are computed in 64 bit arithmetic without a cast at each site: - vm_area_alloc_pages() takes and returns a page count, and the result is stored into nr_pages, so its parameter, return type and the nr_allocated and nr_remaining accumulators widen too. The 100 page cap on a bulk request stays narrow, only its literal becomes 100UL. - nr_small_pages in __vmalloc_area_node() is derived from a 64 bit size and compared against nr_pages, so it widens as well. Before this the store truncated for a size of 16 TiB or more, which sized area->pages from the wrapped count while __vmap_pages_range() still walked the full size. The cast on array_size next to it was already redundant and goes away with it. - new_nr_pages and old_nr_pages in vrealloc_node_align_noprof() widen, which drops the four casts on them, and the casts added by the two preceding patches in vrealloc_node_align_noprof() and vread_iter() are no longer needed either. - vm_area_free_pages() takes a page index range. Three loop counters that index area->pages were plain int and would have overflowed at 2^31 pages: in set_area_direct_map(), in vm_reset_perms() and in the bulk allocation loop, where the index is seeded from nr_allocated. They become unsigned long. Printing needed fixing in two places. vmalloc_dump_obj() used %u, and vmalloc_info_show() printed the unsigned field with %d, which would have shown a negative page count for an area of 8 TiB or more. show_numa_info() used a single unsigned int as both the page index and the node id, so the page loop gets its own unsigned long counter. The page table walkers are not covered here. vmap_pages_pte_range() and the levels above it carry the page cursor as int *nr, so a mapping is still capped independently of the counters this patch widens. Widening that cursor touches the whole walker and is separate work. Nothing outside mm/vmalloc.c needs a change to build or to behave as before. Those users either index the pages array or multiply the count by PAGE_SIZE, which is unsigned long already. kexec_handover stores the count into a 32 bit field of its own ABI, which caps that interface exactly where it did before. On x86-64 this leaves sizeof(struct vm_struct) at 72 bytes when CONFIG_HAVE_ARCH_HUGE_VMALLOC=n, since the existing padding before phys_addr absorbs the change, and takes it to 80 when the config is enabled. Both sizes come from the same kmalloc-96 bucket that __get_vm_area_node() allocates from, so the footprint per area does not change either way. Suggested-by: Andrew Morton Assisted-by: Claude:claude-fable-5 Signed-off-by: Artem Lytkin --- include/linux/vmalloc.h | 2 +- mm/vmalloc.c | 62 ++++++++++++++++++++--------------------- 2 files changed, 31 insertions(+), 33 deletions(-) diff --git a/include/linux/vmalloc.h b/include/linux/vmalloc.h index d87dc7f77f4e8..a39c20c77efcd 100644 --- a/include/linux/vmalloc.h +++ b/include/linux/vmalloc.h @@ -62,7 +62,7 @@ struct vm_struct { #ifdef CONFIG_HAVE_ARCH_HUGE_VMALLOC unsigned int page_order; #endif - unsigned int nr_pages; + unsigned long nr_pages; phys_addr_t phys_addr; const void *caller; unsigned long requested_size; diff --git a/mm/vmalloc.c b/mm/vmalloc.c index baf7e396e3fc7..efd5e42c64e2e 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -3338,7 +3338,7 @@ struct vm_struct *remove_vm_area(const void *addr) static inline void set_area_direct_map(const struct vm_struct *area, int (*set_direct_map)(struct page *page)) { - int i; + unsigned long i; /* HUGE_VMALLOC passes small pages to set_direct_map */ for (i = 0; i < area->nr_pages; i++) @@ -3354,7 +3354,7 @@ static void vm_reset_perms(struct vm_struct *area) unsigned long start = ULONG_MAX, end = 0; unsigned int page_order = vm_area_page_order(area); int flush_dmap = 0; - int i; + unsigned long i; /* * Find the start and end range of the direct mappings to make sure that @@ -3427,10 +3427,10 @@ void vfree_atomic(const void *addr) * Caller is responsible for unmapping (vunmap_range) and KASAN * poisoning before calling this. */ -static void vm_area_free_pages(struct vm_struct *vm, unsigned int start_idx, - unsigned int end_idx) +static void vm_area_free_pages(struct vm_struct *vm, unsigned long start_idx, + unsigned long end_idx) { - unsigned int i; + unsigned long i; if (!(vm->flags & VM_MAP_PUT_PAGES)) { for (i = start_idx; i < end_idx; i++) @@ -3642,12 +3642,12 @@ static inline gfp_t vmalloc_gfp_adjust(gfp_t flags, const bool large) return flags; } -static inline unsigned int +static inline unsigned long vm_area_alloc_pages(gfp_t gfp, int nid, - unsigned int order, unsigned int nr_pages, struct page **pages) + unsigned int order, unsigned long nr_pages, struct page **pages) { - unsigned int nr_allocated = 0; - unsigned int nr_remaining = nr_pages; + unsigned long nr_allocated = 0; + unsigned long nr_remaining = nr_pages; unsigned int max_attempt_order = MAX_PAGE_ORDER; struct page *page; int i; @@ -3695,7 +3695,7 @@ vm_area_alloc_pages(gfp_t gfp, int nid, if (!order) { while (nr_allocated < nr_pages) { unsigned int nr, nr_pages_request; - int i; + unsigned long i; /* * A maximum allowed request is hard-coded and is 100 @@ -3703,7 +3703,7 @@ vm_area_alloc_pages(gfp_t gfp, int nid, * long preemption off scenario in the bulk-allocator * so the range is [1:100]. */ - nr_pages_request = min(100U, nr_pages - nr_allocated); + nr_pages_request = min(100UL, nr_pages - nr_allocated); /* memory allocation should consider mempolicy, we can't * wrongly use nearest node when nid == NUMA_NO_NODE, @@ -3849,12 +3849,12 @@ static void *__vmalloc_area_node(struct vm_struct *area, gfp_t gfp_mask, unsigned long addr = (unsigned long)area->addr; unsigned long size = get_vm_area_size(area); unsigned long array_size; - unsigned int nr_small_pages = size >> PAGE_SHIFT; + unsigned long nr_small_pages = size >> PAGE_SHIFT; unsigned int page_order; unsigned int flags; int ret; - array_size = (unsigned long)nr_small_pages * sizeof(struct page *); + array_size = nr_small_pages * sizeof(struct page *); /* __GFP_NOFAIL and "noblock" flags are mutually exclusive. */ if (!gfpflags_allow_blocking(gfp_mask)) @@ -4352,7 +4352,7 @@ void *vrealloc_node_align_noprof(const void *p, size_t size, unsigned long align } if (size <= old_size) { - unsigned int new_nr_pages = PAGE_ALIGN(size) >> PAGE_SHIFT; + unsigned long new_nr_pages = PAGE_ALIGN(size) >> PAGE_SHIFT; /* Zero out "freed" memory, potentially for future realloc. */ if (want_init_on_free() || want_init_on_alloc(flags)) @@ -4381,7 +4381,7 @@ void *vrealloc_node_align_noprof(const void *p, size_t size, unsigned long align !(vm->flags & (VM_FLUSH_RESET_PERMS | VM_USERMAP)) && gfp_has_io_fs(flags)) { unsigned long addr = (unsigned long)kasan_reset_tag(p); - unsigned int old_nr_pages = vm->nr_pages; + unsigned long old_nr_pages = vm->nr_pages; /* * Use the node lock to synchronize with concurrent @@ -4394,16 +4394,13 @@ void *vrealloc_node_align_noprof(const void *p, size_t size, unsigned long align spin_unlock(&vn->busy.lock); /* Notify kmemleak of the reduced allocation size before unmapping. */ - kmemleak_free_part( - (void *)addr + ((unsigned long)new_nr_pages - << PAGE_SHIFT), - (unsigned long)(old_nr_pages - new_nr_pages) - << PAGE_SHIFT); + kmemleak_free_part((void *)addr + + (new_nr_pages << PAGE_SHIFT), + (old_nr_pages - new_nr_pages) + << PAGE_SHIFT); - vunmap_range(addr + ((unsigned long)new_nr_pages - << PAGE_SHIFT), - addr + ((unsigned long)old_nr_pages - << PAGE_SHIFT)); + vunmap_range(addr + (new_nr_pages << PAGE_SHIFT), + addr + (old_nr_pages << PAGE_SHIFT)); vm_area_free_pages(vm, new_nr_pages, old_nr_pages); } @@ -4415,7 +4412,7 @@ void *vrealloc_node_align_noprof(const void *p, size_t size, unsigned long align /* * We already have the bytes available in the allocation; use them. */ - if (size <= (unsigned long)vm->nr_pages << PAGE_SHIFT) { + if (size <= vm->nr_pages << PAGE_SHIFT) { /* * No need to zero memory here, as unused memory will have * already been zeroed at initial allocation time or during @@ -4722,7 +4719,7 @@ long vread_iter(struct iov_iter *iter, const char *addr, size_t count) * mapping types (vmap, ioremap) don't set nr_pages. */ size = (vm->flags & VM_ALLOC && vm->nr_pages) ? - ((unsigned long)vm->nr_pages << PAGE_SHIFT) : + (vm->nr_pages << PAGE_SHIFT) : get_vm_area_size(vm); else size = va_size(va); @@ -5228,7 +5225,7 @@ bool vmalloc_dump_obj(void *object) struct vmap_area *va; struct vmap_node *vn; unsigned long addr; - unsigned int nr_pages; + unsigned long nr_pages; addr = PAGE_ALIGN((unsigned long) object); vn = addr_to_node(addr); @@ -5248,7 +5245,7 @@ bool vmalloc_dump_obj(void *object) nr_pages = vm->nr_pages; spin_unlock(&vn->busy.lock); - pr_cont(" %u-page vmalloc region starting at %#lx allocated at %pS\n", + pr_cont(" %lu-page vmalloc region starting at %#lx allocated at %pS\n", nr_pages, addr, caller); return true; @@ -5266,16 +5263,17 @@ bool vmalloc_dump_obj(void *object) static void show_numa_info(struct seq_file *m, struct vm_struct *v, unsigned int *counters) { - unsigned int nr; unsigned int step = 1U << vm_area_page_order(v); + unsigned long i; + unsigned int nr; if (!counters) return; memset(counters, 0, nr_node_ids * sizeof(unsigned int)); - for (nr = 0; nr < v->nr_pages; nr += step) - counters[page_to_nid(v->pages[nr])] += step; + for (i = 0; i < v->nr_pages; i += step) + counters[page_to_nid(v->pages[i])] += step; for_each_node_state(nr, N_HIGH_MEMORY) if (counters[nr]) seq_printf(m, " N%u=%u", nr, counters[nr]); @@ -5333,7 +5331,7 @@ static int vmalloc_info_show(struct seq_file *m, void *p) seq_printf(m, " %pS", v->caller); if (v->nr_pages) - seq_printf(m, " pages=%d", v->nr_pages); + seq_printf(m, " pages=%lu", v->nr_pages); if (v->phys_addr) seq_printf(m, " phys=%pa", &v->phys_addr); base-commit: 248951ddc14de84de3910f9b13f51491a8cd91df prerequisite-patch-id: bc462c518eb9ddcaccad5e1f5fed7cc24512b117 prerequisite-patch-id: 891688c2e93c570d1c9c7802b46aaba974002684 -- 2.43.0