Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v5 0/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
@ 2026-08-12  5:41 Hao Ge
  2026-08-12  5:41 ` [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags() Hao Ge
  2026-08-12  5:41 ` [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
  0 siblings, 2 replies; 11+ messages in thread
From: Hao Ge @ 2026-08-12  5:41 UTC (permalink / raw)
  To: Suren Baghdasaryan, Andrew Morton, Luis Chamberlain, Petr Pavlu,
	Daniel Gomez, Sami Tolvanen, Aaron Tomlin
  Cc: linux-modules, linux-kernel, linux-mm, Hao Ge

v3 was a single patch. After discussion with Suren and Andrew we went
for a more graceful approach: rather than failing the module load on
overflow, let it load without profiling. Once profiling is disabled,
codetag_needs_module_section() returns false, so on retry the codetag
section is placed as regular module data.

A new patch (1/2) is added to move release_module_tags() above
reserve_module_tags(), since the overflow path now has to call it and
the helper sits below it.

release_module_tags() is what module unload calls to drop a module's
reservation from the maple tree. By the time reserve_module_tags()
detects the overflow it has already stored that reservation, and the
-EAGAIN return skips vm_module_tags_populate(), so the backing pages
never get mapped. If reserve_module_tags() returns without calling
release_module_tags(), the stale entry keeps pointing at that unmapped
range; when the module is later unloaded, release_module_tags() walks
it and panics.

Tested on an x86_64 virtual machine:

# insmod overflow_tag.ko
# dmesg
  With module overflow_tag there are too many tags to fit in 13 page
  flag bits. Memory allocation profiling is disabled!
# rmmod overflow_tag
The module loads without profiling.

Changes in v5:
- add Fixes: and Cc: stable to patch 1/2 as well, since 2/2 does not
  compile without it (Andrew Morton)
- restore frob-adjusted mem[type].size on retry instead of zeroing,
  as s390 and parisc add GOT/PLT space there in
  module_frob_arch_sections() (Reported by Sashiko)
- drop the load_module() mem_profiling_support check; the percpu
  counter leak is pre-existing and orthogonal to this fix

Changes in v4:
- add a new patch (1/2) to move release_module_tags() above
  reserve_module_tags(); the overflow fix is 2/2
- release the reservation on the -EAGAIN path
- return -EAGAIN instead of -ENOMEM so the module can still load
  without profiling (Suren)
- reset sh_addr, mem[type].size and sym/str SHF_ALLOC before retry
- skip percpu counters in load_module() when profiling is off

Changes in v3:
- use pr_warn_once() instead of pr_warn()
- return -ENOMEM instead of -ENOSPC (Suren)
- expand the commit message to describe the /proc/allocinfo impact
  (Andrew)

Changes in v2:
- return an error after shutdown_mem_profiling() to skip
  vm_module_tags_populate()

v1: https://lore.kernel.org/all/20260804064408.105033-1-hao.ge@linux.dev/
v2: https://lore.kernel.org/all/20260804122038.190270-1-hao.ge@linux.dev/
v3: https://lore.kernel.org/all/20260805090633.141001-1-hao.ge@linux.dev/
v4: https://lore.kernel.org/all/20260810093955.153015-1-hao.ge@linux.dev/

Hao Ge (2):
  alloc_tag: move release_module_tags() above reserve_module_tags()
  alloc_tag: fix undetected compressed tag overflow when profiling is
    disabled

 kernel/module/main.c |  25 ++++++++++-
 mm/alloc_tag.c       | 100 ++++++++++++++++++++++---------------------
 2 files changed, 74 insertions(+), 51 deletions(-)

--
2.25.1


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags()
  2026-08-12  5:41 [PATCH v5 0/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
@ 2026-08-12  5:41 ` Hao Ge
  2026-08-15  3:40   ` Suren Baghdasaryan
  2026-08-12  5:41 ` [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
  1 sibling, 1 reply; 11+ messages in thread
From: Hao Ge @ 2026-08-12  5:41 UTC (permalink / raw)
  To: Suren Baghdasaryan, Andrew Morton, Luis Chamberlain, Petr Pavlu,
	Daniel Gomez, Sami Tolvanen, Aaron Tomlin
  Cc: linux-modules, linux-kernel, linux-mm, Hao Ge, stable

release_module_tags() is a cleanup helper. reserve_module_tags() can
also fail after storing the reservation in the maple tree, in which
case it should call release_module_tags() to undo it. Move the helper
above reserve_module_tags() so no forward declaration is needed.

No functional change.

Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
 mm/alloc_tag.c | 92 +++++++++++++++++++++++++-------------------------
 1 file changed, 46 insertions(+), 46 deletions(-)

diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
index 52aece27b00e..af44f90379f2 100644
--- a/mm/alloc_tag.c
+++ b/mm/alloc_tag.c
@@ -835,6 +835,52 @@ static int vm_module_tags_populate(void)
 	return 0;
 }
 
+static void release_module_tags(struct module *mod, bool used)
+{
+	MA_STATE(mas, &mod_area_mt, module_tags.size, module_tags.size);
+	struct alloc_tag *start_tag;
+	struct alloc_tag *end_tag;
+	struct module *val;
+
+	mas_lock(&mas);
+	mas_for_each_rev(&mas, val, 0)
+		if (val == mod)
+			break;
+
+	if (!val) /* module not found */
+		goto out;
+
+	if (!used)
+		goto release_area;
+
+	start_tag = (struct alloc_tag *)(module_tags.start_addr + mas.index);
+	end_tag = (struct alloc_tag *)(module_tags.start_addr + mas.last);
+	if (!clean_unused_counters(start_tag, end_tag)) {
+		struct alloc_tag *tag;
+
+		for (tag = start_tag; tag <= end_tag; tag++) {
+			struct alloc_tag_counters counter;
+
+			if (!tag->counters)
+				continue;
+
+			counter = alloc_tag_read(tag);
+			pr_info("%s:%u module %s func:%s has %llu allocated at module unload\n",
+				tag->ct.filename, tag->ct.lineno, tag->ct.modname,
+				tag->ct.function, counter.bytes);
+		}
+	} else {
+		used = false;
+	}
+release_area:
+	mas_store(&mas, used ? &unloaded_mod : NULL);
+	val = mas_prev_range(&mas, 0);
+	if (val == &prepend_mod)
+		mas_store(&mas, NULL);
+out:
+	mas_unlock(&mas);
+}
+
 static void *reserve_module_tags(struct module *mod, unsigned long size,
 				 unsigned int prepend, unsigned long align)
 {
@@ -922,52 +968,6 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
 	return (struct alloc_tag *)(module_tags.start_addr + offset);
 }
 
-static void release_module_tags(struct module *mod, bool used)
-{
-	MA_STATE(mas, &mod_area_mt, module_tags.size, module_tags.size);
-	struct alloc_tag *start_tag;
-	struct alloc_tag *end_tag;
-	struct module *val;
-
-	mas_lock(&mas);
-	mas_for_each_rev(&mas, val, 0)
-		if (val == mod)
-			break;
-
-	if (!val) /* module not found */
-		goto out;
-
-	if (!used)
-		goto release_area;
-
-	start_tag = (struct alloc_tag *)(module_tags.start_addr + mas.index);
-	end_tag = (struct alloc_tag *)(module_tags.start_addr + mas.last);
-	if (!clean_unused_counters(start_tag, end_tag)) {
-		struct alloc_tag *tag;
-
-		for (tag = start_tag; tag <= end_tag; tag++) {
-			struct alloc_tag_counters counter;
-
-			if (!tag->counters)
-				continue;
-
-			counter = alloc_tag_read(tag);
-			pr_info("%s:%u module %s func:%s has %llu allocated at module unload\n",
-				tag->ct.filename, tag->ct.lineno, tag->ct.modname,
-				tag->ct.function, counter.bytes);
-		}
-	} else {
-		used = false;
-	}
-release_area:
-	mas_store(&mas, used ? &unloaded_mod : NULL);
-	val = mas_prev_range(&mas, 0);
-	if (val == &prepend_mod)
-		mas_store(&mas, NULL);
-out:
-	mas_unlock(&mas);
-}
-
 static int load_module(struct module *mod, struct codetag *start, struct codetag *stop)
 {
 	/* Allocate module alloc_tag percpu counters */
-- 
2.25.1



^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-12  5:41 [PATCH v5 0/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
  2026-08-12  5:41 ` [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags() Hao Ge
@ 2026-08-12  5:41 ` Hao Ge
  2026-08-15  5:58   ` Suren Baghdasaryan
  1 sibling, 1 reply; 11+ messages in thread
From: Hao Ge @ 2026-08-12  5:41 UTC (permalink / raw)
  To: Suren Baghdasaryan, Andrew Morton, Luis Chamberlain, Petr Pavlu,
	Daniel Gomez, Sami Tolvanen, Aaron Tomlin
  Cc: linux-modules, linux-kernel, linux-mm, Hao Ge, stable

In reserve_module_tags(), the tag overflow check is gated on
mem_alloc_profiling_enabled():

    if (mem_alloc_profiling_enabled() && !tags_addressable())

If profiling is toggled off at runtime and a module is loaded whose
tags exceed the compressed-mode limit, shutdown_mem_profiling() is
skipped. vm_module_tags_populate() still maps memory for the tags and
the module loads successfully, but the total tag count now exceeds what
NR_UNUSED_PAGEFLAG_BITS can address.

Once profiling is re-enabled, ref_to_idx() computes each tag's index
as its position in the alloc_tag array. update_page_tag_ref() masks
it to alloc_tag_ref_mask before storing in page->flags. Indices
beyond the mask are truncated and idx_to_ref() resolves them to wrong
tags.

This silently corrupts /proc/allocinfo: allocated pages get attributed
to the wrong call sites, so the statistics it reports are wrong.

mem_alloc_profiling_enabled() and mem_profiling_compressed are
independent. Once compressed mode is established at boot, it stays
active regardless of runtime toggles of mem_profiling.

Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
profiling, release the reservation, and return -EAGAIN so that
layout_and_allocate() retries with profiling disabled: codetag sections
are then placed as regular module data and the module loads without
profiling rather than being rejected entirely.

Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
Cc: stable@vger.kernel.org
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
 kernel/module/main.c | 25 +++++++++++++++++++++++--
 mm/alloc_tag.c       |  8 +++++---
 2 files changed, 28 insertions(+), 5 deletions(-)

diff --git a/kernel/module/main.c b/kernel/module/main.c
index 46dd8d25a605..ed26f167be84 100644
--- a/kernel/module/main.c
+++ b/kernel/module/main.c
@@ -2944,6 +2944,7 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
 {
 	struct module *mod;
 	int err;
+	unsigned long frob_size[MOD_MEM_NUM_TYPES];
 
 	/* Allow arches to frob section contents and sizes.  */
 	err = module_frob_arch_sections(info->hdr, info->sechdrs,
@@ -2966,18 +2967,38 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
 	 */
 	module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
 
+	/*
+	 * Save the sizes reserved by module_frob_arch_sections() so they can
+	 * be restored if we retry below.
+	 */
+	for_each_mod_mem_type(type)
+		frob_size[type] = info->mod->mem[type].size;
+
 	/*
 	 * Determine total sizes, and put offsets in sh_entsize.  For now
 	 * this is done generically; there doesn't appear to be any
 	 * special cases for the architectures.
 	 */
+retry:
 	layout_sections(info->mod, info);
 	layout_symtab(info->mod, info);
 
 	/* Allocate and move to the final place */
 	err = move_module(info->mod, info);
-	if (err)
-		return ERR_PTR(err);
+	if (err) {
+		if (err != -EAGAIN)
+			return ERR_PTR(err);
+		/*
+		 * -EAGAIN means profiling was disabled but the module
+		 * can still load without it. Reset state and retry.
+		 */
+		rewrite_section_headers(info, flags);
+		for_each_mod_mem_type(type)
+			info->mod->mem[type].size = frob_size[type];
+		info->sechdrs[info->index.sym].sh_flags &= ~(unsigned long)SHF_ALLOC;
+		info->sechdrs[info->index.str].sh_flags &= ~(unsigned long)SHF_ALLOC;
+		goto retry;
+	}
 
 	/* Module has been copied to its final place now: return it. */
 	mod = (void *)info->sechdrs[info->index.mod].sh_addr;
diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
index af44f90379f2..0a7b657fe2de 100644
--- a/mm/alloc_tag.c
+++ b/mm/alloc_tag.c
@@ -950,10 +950,12 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
 		int grow_res;
 
 		module_tags.size = offset + size;
-		if (mem_alloc_profiling_enabled() && !tags_addressable()) {
+		if (!tags_addressable()) {
 			shutdown_mem_profiling(true);
-			pr_warn("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
-				mod->name, NR_UNUSED_PAGEFLAG_BITS);
+			pr_warn_once("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
+				     mod->name, NR_UNUSED_PAGEFLAG_BITS);
+			release_module_tags(mod, false);
+			return ERR_PTR(-EAGAIN);
 		}
 
 		grow_res = vm_module_tags_populate();
-- 
2.25.1



^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags()
  2026-08-12  5:41 ` [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags() Hao Ge
@ 2026-08-15  3:40   ` Suren Baghdasaryan
  0 siblings, 0 replies; 11+ messages in thread
From: Suren Baghdasaryan @ 2026-08-15  3:40 UTC (permalink / raw)
  To: Hao Ge
  Cc: Andrew Morton, Luis Chamberlain, Petr Pavlu, Daniel Gomez,
	Sami Tolvanen, Aaron Tomlin, linux-modules, linux-kernel,
	linux-mm, stable

On Tue, Aug 11, 2026 at 10:41 PM Hao Ge <hao.ge@linux.dev> wrote:
>
> release_module_tags() is a cleanup helper. reserve_module_tags() can
> also fail after storing the reservation in the maple tree, in which
> case it should call release_module_tags() to undo it. Move the helper
> above reserve_module_tags() so no forward declaration is needed.
>
> No functional change.
>
> Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
> Cc: stable@vger.kernel.org
> Signed-off-by: Hao Ge <hao.ge@linux.dev>

Acked-by: Suren Baghdasaryan <surenb@google.com>

> ---
>  mm/alloc_tag.c | 92 +++++++++++++++++++++++++-------------------------
>  1 file changed, 46 insertions(+), 46 deletions(-)
>
> diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
> index 52aece27b00e..af44f90379f2 100644
> --- a/mm/alloc_tag.c
> +++ b/mm/alloc_tag.c
> @@ -835,6 +835,52 @@ static int vm_module_tags_populate(void)
>         return 0;
>  }
>
> +static void release_module_tags(struct module *mod, bool used)
> +{
> +       MA_STATE(mas, &mod_area_mt, module_tags.size, module_tags.size);
> +       struct alloc_tag *start_tag;
> +       struct alloc_tag *end_tag;
> +       struct module *val;
> +
> +       mas_lock(&mas);
> +       mas_for_each_rev(&mas, val, 0)
> +               if (val == mod)
> +                       break;
> +
> +       if (!val) /* module not found */
> +               goto out;
> +
> +       if (!used)
> +               goto release_area;
> +
> +       start_tag = (struct alloc_tag *)(module_tags.start_addr + mas.index);
> +       end_tag = (struct alloc_tag *)(module_tags.start_addr + mas.last);
> +       if (!clean_unused_counters(start_tag, end_tag)) {
> +               struct alloc_tag *tag;
> +
> +               for (tag = start_tag; tag <= end_tag; tag++) {
> +                       struct alloc_tag_counters counter;
> +
> +                       if (!tag->counters)
> +                               continue;
> +
> +                       counter = alloc_tag_read(tag);
> +                       pr_info("%s:%u module %s func:%s has %llu allocated at module unload\n",
> +                               tag->ct.filename, tag->ct.lineno, tag->ct.modname,
> +                               tag->ct.function, counter.bytes);
> +               }
> +       } else {
> +               used = false;
> +       }
> +release_area:
> +       mas_store(&mas, used ? &unloaded_mod : NULL);
> +       val = mas_prev_range(&mas, 0);
> +       if (val == &prepend_mod)
> +               mas_store(&mas, NULL);
> +out:
> +       mas_unlock(&mas);
> +}
> +
>  static void *reserve_module_tags(struct module *mod, unsigned long size,
>                                  unsigned int prepend, unsigned long align)
>  {
> @@ -922,52 +968,6 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
>         return (struct alloc_tag *)(module_tags.start_addr + offset);
>  }
>
> -static void release_module_tags(struct module *mod, bool used)
> -{
> -       MA_STATE(mas, &mod_area_mt, module_tags.size, module_tags.size);
> -       struct alloc_tag *start_tag;
> -       struct alloc_tag *end_tag;
> -       struct module *val;
> -
> -       mas_lock(&mas);
> -       mas_for_each_rev(&mas, val, 0)
> -               if (val == mod)
> -                       break;
> -
> -       if (!val) /* module not found */
> -               goto out;
> -
> -       if (!used)
> -               goto release_area;
> -
> -       start_tag = (struct alloc_tag *)(module_tags.start_addr + mas.index);
> -       end_tag = (struct alloc_tag *)(module_tags.start_addr + mas.last);
> -       if (!clean_unused_counters(start_tag, end_tag)) {
> -               struct alloc_tag *tag;
> -
> -               for (tag = start_tag; tag <= end_tag; tag++) {
> -                       struct alloc_tag_counters counter;
> -
> -                       if (!tag->counters)
> -                               continue;
> -
> -                       counter = alloc_tag_read(tag);
> -                       pr_info("%s:%u module %s func:%s has %llu allocated at module unload\n",
> -                               tag->ct.filename, tag->ct.lineno, tag->ct.modname,
> -                               tag->ct.function, counter.bytes);
> -               }
> -       } else {
> -               used = false;
> -       }
> -release_area:
> -       mas_store(&mas, used ? &unloaded_mod : NULL);
> -       val = mas_prev_range(&mas, 0);
> -       if (val == &prepend_mod)
> -               mas_store(&mas, NULL);
> -out:
> -       mas_unlock(&mas);
> -}
> -
>  static int load_module(struct module *mod, struct codetag *start, struct codetag *stop)
>  {
>         /* Allocate module alloc_tag percpu counters */
> --
> 2.25.1
>


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-12  5:41 ` [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
@ 2026-08-15  5:58   ` Suren Baghdasaryan
  2026-08-15 10:45     ` Petr Pavlu
  0 siblings, 1 reply; 11+ messages in thread
From: Suren Baghdasaryan @ 2026-08-15  5:58 UTC (permalink / raw)
  To: Hao Ge
  Cc: Andrew Morton, Luis Chamberlain, Petr Pavlu, Daniel Gomez,
	Sami Tolvanen, Aaron Tomlin, linux-modules, linux-kernel,
	linux-mm, stable

On Tue, Aug 11, 2026 at 10:41 PM Hao Ge <hao.ge@linux.dev> wrote:
>
> In reserve_module_tags(), the tag overflow check is gated on
> mem_alloc_profiling_enabled():
>
>     if (mem_alloc_profiling_enabled() && !tags_addressable())
>
> If profiling is toggled off at runtime and a module is loaded whose
> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
> skipped. vm_module_tags_populate() still maps memory for the tags and
> the module loads successfully, but the total tag count now exceeds what
> NR_UNUSED_PAGEFLAG_BITS can address.
>
> Once profiling is re-enabled, ref_to_idx() computes each tag's index
> as its position in the alloc_tag array. update_page_tag_ref() masks
> it to alloc_tag_ref_mask before storing in page->flags. Indices
> beyond the mask are truncated and idx_to_ref() resolves them to wrong
> tags.
>
> This silently corrupts /proc/allocinfo: allocated pages get attributed
> to the wrong call sites, so the statistics it reports are wrong.
>
> mem_alloc_profiling_enabled() and mem_profiling_compressed are
> independent. Once compressed mode is established at boot, it stays
> active regardless of runtime toggles of mem_profiling.
>
> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
> profiling, release the reservation, and return -EAGAIN so that
> layout_and_allocate() retries with profiling disabled: codetag sections
> are then placed as regular module data and the module loads without
> profiling rather than being rejected entirely.
>
> Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
> Cc: stable@vger.kernel.org
> Suggested-by: Suren Baghdasaryan <surenb@google.com>
> Signed-off-by: Hao Ge <hao.ge@linux.dev>

I'm looking at your fixes for pre-existing issues reported by Sashiko
but this patch looks correct to me.

Acked-by: Suren Baghdasaryan <surenb@google.com>

> ---
>  kernel/module/main.c | 25 +++++++++++++++++++++++--
>  mm/alloc_tag.c       |  8 +++++---
>  2 files changed, 28 insertions(+), 5 deletions(-)
>
> diff --git a/kernel/module/main.c b/kernel/module/main.c
> index 46dd8d25a605..ed26f167be84 100644
> --- a/kernel/module/main.c
> +++ b/kernel/module/main.c
> @@ -2944,6 +2944,7 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>  {
>         struct module *mod;
>         int err;
> +       unsigned long frob_size[MOD_MEM_NUM_TYPES];
>
>         /* Allow arches to frob section contents and sizes.  */
>         err = module_frob_arch_sections(info->hdr, info->sechdrs,
> @@ -2966,18 +2967,38 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>          */
>         module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
>
> +       /*
> +        * Save the sizes reserved by module_frob_arch_sections() so they can
> +        * be restored if we retry below.
> +        */
> +       for_each_mod_mem_type(type)
> +               frob_size[type] = info->mod->mem[type].size;
> +
>         /*
>          * Determine total sizes, and put offsets in sh_entsize.  For now
>          * this is done generically; there doesn't appear to be any
>          * special cases for the architectures.
>          */
> +retry:
>         layout_sections(info->mod, info);
>         layout_symtab(info->mod, info);
>
>         /* Allocate and move to the final place */
>         err = move_module(info->mod, info);
> -       if (err)
> -               return ERR_PTR(err);
> +       if (err) {
> +               if (err != -EAGAIN)
> +                       return ERR_PTR(err);
> +               /*
> +                * -EAGAIN means profiling was disabled but the module
> +                * can still load without it. Reset state and retry.
> +                */
> +               rewrite_section_headers(info, flags);
> +               for_each_mod_mem_type(type)
> +                       info->mod->mem[type].size = frob_size[type];
> +               info->sechdrs[info->index.sym].sh_flags &= ~(unsigned long)SHF_ALLOC;
> +               info->sechdrs[info->index.str].sh_flags &= ~(unsigned long)SHF_ALLOC;
> +               goto retry;
> +       }
>
>         /* Module has been copied to its final place now: return it. */
>         mod = (void *)info->sechdrs[info->index.mod].sh_addr;
> diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
> index af44f90379f2..0a7b657fe2de 100644
> --- a/mm/alloc_tag.c
> +++ b/mm/alloc_tag.c
> @@ -950,10 +950,12 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
>                 int grow_res;
>
>                 module_tags.size = offset + size;
> -               if (mem_alloc_profiling_enabled() && !tags_addressable()) {
> +               if (!tags_addressable()) {
>                         shutdown_mem_profiling(true);
> -                       pr_warn("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
> -                               mod->name, NR_UNUSED_PAGEFLAG_BITS);
> +                       pr_warn_once("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
> +                                    mod->name, NR_UNUSED_PAGEFLAG_BITS);
> +                       release_module_tags(mod, false);
> +                       return ERR_PTR(-EAGAIN);
>                 }
>
>                 grow_res = vm_module_tags_populate();
> --
> 2.25.1
>


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-15  5:58   ` Suren Baghdasaryan
@ 2026-08-15 10:45     ` Petr Pavlu
  2026-08-16 15:16       ` Suren Baghdasaryan
  0 siblings, 1 reply; 11+ messages in thread
From: Petr Pavlu @ 2026-08-15 10:45 UTC (permalink / raw)
  To: Suren Baghdasaryan, Hao Ge
  Cc: Andrew Morton, Luis Chamberlain, Daniel Gomez, Sami Tolvanen,
	Aaron Tomlin, linux-modules, linux-kernel, linux-mm, stable

On 8/12/26 7:41 AM, Hao Ge wrote:
> In reserve_module_tags(), the tag overflow check is gated on
> mem_alloc_profiling_enabled():
> 
>     if (mem_alloc_profiling_enabled() && !tags_addressable())
> 
> If profiling is toggled off at runtime and a module is loaded whose
> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
> skipped. vm_module_tags_populate() still maps memory for the tags and
> the module loads successfully, but the total tag count now exceeds what
> NR_UNUSED_PAGEFLAG_BITS can address.
> 
> Once profiling is re-enabled, ref_to_idx() computes each tag's index
> as its position in the alloc_tag array. update_page_tag_ref() masks
> it to alloc_tag_ref_mask before storing in page->flags. Indices
> beyond the mask are truncated and idx_to_ref() resolves them to wrong
> tags.
> 
> This silently corrupts /proc/allocinfo: allocated pages get attributed
> to the wrong call sites, so the statistics it reports are wrong.
> 
> mem_alloc_profiling_enabled() and mem_profiling_compressed are
> independent. Once compressed mode is established at boot, it stays
> active regardless of runtime toggles of mem_profiling.
> 
> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
> profiling, release the reservation, and return -EAGAIN so that
> layout_and_allocate() retries with profiling disabled: codetag sections
> are then placed as regular module data and the module loads without
> profiling rather than being rejected entirely.

When the described overflow occurs, why should codetag sections be
placed as regular module data? Will the codetag support use them in any
way, or do they simply waste space? Is the issue that alloc_hooks()
creates relocations pointing into .codetag.alloc_tags?

> 
> Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
> Cc: stable@vger.kernel.org
> Suggested-by: Suren Baghdasaryan <surenb@google.com>
> Signed-off-by: Hao Ge <hao.ge@linux.dev>
> ---
>  kernel/module/main.c | 25 +++++++++++++++++++++++--
>  mm/alloc_tag.c       |  8 +++++---
>  2 files changed, 28 insertions(+), 5 deletions(-)
> 
> diff --git a/kernel/module/main.c b/kernel/module/main.c
> index 46dd8d25a605..ed26f167be84 100644
> --- a/kernel/module/main.c
> +++ b/kernel/module/main.c
> @@ -2944,6 +2944,7 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>  {
>  	struct module *mod;
>  	int err;
> +	unsigned long frob_size[MOD_MEM_NUM_TYPES];

frob_size is used to store values of module_memory::size, which has type
`unsigned int`. The types should match.

>  
>  	/* Allow arches to frob section contents and sizes.  */
>  	err = module_frob_arch_sections(info->hdr, info->sechdrs,
> @@ -2966,18 +2967,38 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>  	 */
>  	module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
>  
> +	/*
> +	 * Save the sizes reserved by module_frob_arch_sections() so they can
> +	 * be restored if we retry below.
> +	 */
> +	for_each_mod_mem_type(type)
> +		frob_size[type] = info->mod->mem[type].size;
> +
>  	/*
>  	 * Determine total sizes, and put offsets in sh_entsize.  For now
>  	 * this is done generically; there doesn't appear to be any
>  	 * special cases for the architectures.
>  	 */
> +retry:
>  	layout_sections(info->mod, info);
>  	layout_symtab(info->mod, info);
>  
>  	/* Allocate and move to the final place */
>  	err = move_module(info->mod, info);
> -	if (err)
> -		return ERR_PTR(err);
> +	if (err) {
> +		if (err != -EAGAIN)
> +			return ERR_PTR(err);

The move_module() logic is non-trivial. -EAGAIN could be returned by
other code, now or in the future.

> +		/*
> +		 * -EAGAIN means profiling was disabled but the module
> +		 * can still load without it. Reset state and retry.
> +		 */
> +		rewrite_section_headers(info, flags);
> +		for_each_mod_mem_type(type)
> +			info->mod->mem[type].size = frob_size[type];
> +		info->sechdrs[info->index.sym].sh_flags &= ~(unsigned long)SHF_ALLOC;
> +		info->sechdrs[info->index.str].sh_flags &= ~(unsigned long)SHF_ALLOC;

Why is it necessary to reset SHF_ALLOC for .symtab and .strtab here?

> +		goto retry;
> +	}
>  
>  	/* Module has been copied to its final place now: return it. */
>  	mod = (void *)info->sechdrs[info->index.mod].sh_addr;

I'm not sure this is the best approach. It's complex logic for what
appears to be an edge case related to a debugging facility. It will have
the usual problem of error paths not getting enough testing and breaking
subtly over time.

An alternative could be to reset SHF_ALLOC on the codetag section to
remove it from further processing and have relocations that point to
this section resolve to something else. It seems that alloc_hooks_tag()
could tolerate this, since it only needs to reference the associated
alloc_tag when mem_alloc_profiling_enabled() is true and that gets
disabled by reserve_module_tags() on the overflow.

It is also not an ideal approach, but I feel it could be less intrusive
to the module loader. I can put together a prototype if needed.

-- 
Thanks,
Petr


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-15 10:45     ` Petr Pavlu
@ 2026-08-16 15:16       ` Suren Baghdasaryan
  2026-08-17  2:24         ` Hao Ge
  2026-08-26  8:32         ` Petr Pavlu
  0 siblings, 2 replies; 11+ messages in thread
From: Suren Baghdasaryan @ 2026-08-16 15:16 UTC (permalink / raw)
  To: Petr Pavlu
  Cc: Hao Ge, Andrew Morton, Luis Chamberlain, Daniel Gomez,
	Sami Tolvanen, Aaron Tomlin, linux-modules, linux-kernel,
	linux-mm, stable

On Sat, Aug 15, 2026 at 3:45 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>
> On 8/12/26 7:41 AM, Hao Ge wrote:
> > In reserve_module_tags(), the tag overflow check is gated on
> > mem_alloc_profiling_enabled():
> >
> >     if (mem_alloc_profiling_enabled() && !tags_addressable())
> >
> > If profiling is toggled off at runtime and a module is loaded whose
> > tags exceed the compressed-mode limit, shutdown_mem_profiling() is
> > skipped. vm_module_tags_populate() still maps memory for the tags and
> > the module loads successfully, but the total tag count now exceeds what
> > NR_UNUSED_PAGEFLAG_BITS can address.
> >
> > Once profiling is re-enabled, ref_to_idx() computes each tag's index
> > as its position in the alloc_tag array. update_page_tag_ref() masks
> > it to alloc_tag_ref_mask before storing in page->flags. Indices
> > beyond the mask are truncated and idx_to_ref() resolves them to wrong
> > tags.
> >
> > This silently corrupts /proc/allocinfo: allocated pages get attributed
> > to the wrong call sites, so the statistics it reports are wrong.
> >
> > mem_alloc_profiling_enabled() and mem_profiling_compressed are
> > independent. Once compressed mode is established at boot, it stays
> > active regardless of runtime toggles of mem_profiling.
> >
> > Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
> > profiling, release the reservation, and return -EAGAIN so that
> > layout_and_allocate() retries with profiling disabled: codetag sections
> > are then placed as regular module data and the module loads without
> > profiling rather than being rejected entirely.
>
> When the described overflow occurs, why should codetag sections be
> placed as regular module data? Will the codetag support use them in any
> way, or do they simply waste space? Is the issue that alloc_hooks()
> creates relocations pointing into .codetag.alloc_tags?

Correct, alloc_hooks() will have references into .codetag.alloc_tags.
With mem_profiling_support=false they should technically never be used
but I don't think it's a good idea to skip .codetag.alloc_tags section
allocation and to leave dangling pointers. Also the case described
here is an outlier, so optimizing it would not yield much benefit.

>
> >
> > Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
> > Cc: stable@vger.kernel.org
> > Suggested-by: Suren Baghdasaryan <surenb@google.com>
> > Signed-off-by: Hao Ge <hao.ge@linux.dev>
> > ---
> >  kernel/module/main.c | 25 +++++++++++++++++++++++--
> >  mm/alloc_tag.c       |  8 +++++---
> >  2 files changed, 28 insertions(+), 5 deletions(-)
> >
> > diff --git a/kernel/module/main.c b/kernel/module/main.c
> > index 46dd8d25a605..ed26f167be84 100644
> > --- a/kernel/module/main.c
> > +++ b/kernel/module/main.c
> > @@ -2944,6 +2944,7 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
> >  {
> >       struct module *mod;
> >       int err;
> > +     unsigned long frob_size[MOD_MEM_NUM_TYPES];
>
> frob_size is used to store values of module_memory::size, which has type
> `unsigned int`. The types should match.
>
> >
> >       /* Allow arches to frob section contents and sizes.  */
> >       err = module_frob_arch_sections(info->hdr, info->sechdrs,
> > @@ -2966,18 +2967,38 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
> >        */
> >       module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
> >
> > +     /*
> > +      * Save the sizes reserved by module_frob_arch_sections() so they can
> > +      * be restored if we retry below.
> > +      */
> > +     for_each_mod_mem_type(type)
> > +             frob_size[type] = info->mod->mem[type].size;
> > +
> >       /*
> >        * Determine total sizes, and put offsets in sh_entsize.  For now
> >        * this is done generically; there doesn't appear to be any
> >        * special cases for the architectures.
> >        */
> > +retry:
> >       layout_sections(info->mod, info);
> >       layout_symtab(info->mod, info);
> >
> >       /* Allocate and move to the final place */
> >       err = move_module(info->mod, info);
> > -     if (err)
> > -             return ERR_PTR(err);
> > +     if (err) {
> > +             if (err != -EAGAIN)
> > +                     return ERR_PTR(err);
>
> The move_module() logic is non-trivial. -EAGAIN could be returned by
> other code, now or in the future.

That's a good point.

>
> > +             /*
> > +              * -EAGAIN means profiling was disabled but the module
> > +              * can still load without it. Reset state and retry.
> > +              */
> > +             rewrite_section_headers(info, flags);
> > +             for_each_mod_mem_type(type)
> > +                     info->mod->mem[type].size = frob_size[type];
> > +             info->sechdrs[info->index.sym].sh_flags &= ~(unsigned long)SHF_ALLOC;
> > +             info->sechdrs[info->index.str].sh_flags &= ~(unsigned long)SHF_ALLOC;
>
> Why is it necessary to reset SHF_ALLOC for .symtab and .strtab here?

I believe layout_symtab() sets that bit and to retry we need to reset
it. But I might be wrong here.

>
> > +             goto retry;
> > +     }
> >
> >       /* Module has been copied to its final place now: return it. */
> >       mod = (void *)info->sechdrs[info->index.mod].sh_addr;
>
> I'm not sure this is the best approach. It's complex logic for what
> appears to be an edge case related to a debugging facility. It will have
> the usual problem of error paths not getting enough testing and breaking
> subtly over time.
>
> An alternative could be to reset SHF_ALLOC on the codetag section to
> remove it from further processing and have relocations that point to
> this section resolve to something else. It seems that alloc_hooks_tag()
> could tolerate this, since it only needs to reference the associated
> alloc_tag when mem_alloc_profiling_enabled() is true and that gets
> disabled by reserve_module_tags() on the overflow.

Hmm, yeah if we redirect the references into .codetag.alloc_tags, that
would be much better.

>
> It is also not an ideal approach, but I feel it could be less intrusive
> to the module loader. I can put together a prototype if needed.

If your approach does not cause module loading to fail when we disable
profiling, then that sounds like a good idea. If it's not too much
trouble, could you please send an RFC?

>
> --
> Thanks,
> Petr


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-16 15:16       ` Suren Baghdasaryan
@ 2026-08-17  2:24         ` Hao Ge
  2026-08-26  8:32         ` Petr Pavlu
  1 sibling, 0 replies; 11+ messages in thread
From: Hao Ge @ 2026-08-17  2:24 UTC (permalink / raw)
  To: Suren Baghdasaryan, Petr Pavlu
  Cc: Andrew Morton, Luis Chamberlain, Daniel Gomez, Sami Tolvanen,
	Aaron Tomlin, linux-modules, linux-kernel, linux-mm, stable



On 2026/8/16 23:16, Suren Baghdasaryan wrote:
> On Sat, Aug 15, 2026 at 3:45 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>>
>> On 8/12/26 7:41 AM, Hao Ge wrote:
>>> In reserve_module_tags(), the tag overflow check is gated on
>>> mem_alloc_profiling_enabled():
>>>
>>>     if (mem_alloc_profiling_enabled() && !tags_addressable())
>>>
>>> If profiling is toggled off at runtime and a module is loaded whose
>>> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
>>> skipped. vm_module_tags_populate() still maps memory for the tags and
>>> the module loads successfully, but the total tag count now exceeds what
>>> NR_UNUSED_PAGEFLAG_BITS can address.
>>>
>>> Once profiling is re-enabled, ref_to_idx() computes each tag's index
>>> as its position in the alloc_tag array. update_page_tag_ref() masks
>>> it to alloc_tag_ref_mask before storing in page->flags. Indices
>>> beyond the mask are truncated and idx_to_ref() resolves them to wrong
>>> tags.
>>>
>>> This silently corrupts /proc/allocinfo: allocated pages get attributed
>>> to the wrong call sites, so the statistics it reports are wrong.
>>>
>>> mem_alloc_profiling_enabled() and mem_profiling_compressed are
>>> independent. Once compressed mode is established at boot, it stays
>>> active regardless of runtime toggles of mem_profiling.
>>>
>>> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
>>> profiling, release the reservation, and return -EAGAIN so that
>>> layout_and_allocate() retries with profiling disabled: codetag sections
>>> are then placed as regular module data and the module loads without
>>> profiling rather than being rejected entirely.
>>
>> When the described overflow occurs, why should codetag sections be
>> placed as regular module data? Will the codetag support use them in any
>> way, or do they simply waste space? Is the issue that alloc_hooks()
>> creates relocations pointing into .codetag.alloc_tags?
> 
> Correct, alloc_hooks() will have references into .codetag.alloc_tags.
> With mem_profiling_support=false they should technically never be used
> but I don't think it's a good idea to skip .codetag.alloc_tags section
> allocation and to leave dangling pointers. Also the case described
> here is an outlier, so optimizing it would not yield much benefit.
> 
>>
>>>
>>> Fixes: 4835f747d3ed ("alloc_tag: support for page allocation tag compression")
>>> Cc: stable@vger.kernel.org
>>> Suggested-by: Suren Baghdasaryan <surenb@google.com>
>>> Signed-off-by: Hao Ge <hao.ge@linux.dev>
>>> ---
>>>  kernel/module/main.c | 25 +++++++++++++++++++++++--
>>>  mm/alloc_tag.c       |  8 +++++---
>>>  2 files changed, 28 insertions(+), 5 deletions(-)
>>>
>>> diff --git a/kernel/module/main.c b/kernel/module/main.c
>>> index 46dd8d25a605..ed26f167be84 100644
>>> --- a/kernel/module/main.c
>>> +++ b/kernel/module/main.c
>>> @@ -2944,6 +2944,7 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>>>  {
>>>       struct module *mod;
>>>       int err;
>>> +     unsigned long frob_size[MOD_MEM_NUM_TYPES];
>>
>> frob_size is used to store values of module_memory::size, which has type
>> `unsigned int`. The types should match.
>>
>>>
>>>       /* Allow arches to frob section contents and sizes.  */
>>>       err = module_frob_arch_sections(info->hdr, info->sechdrs,
>>> @@ -2966,18 +2967,38 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>>>        */
>>>       module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
>>>
>>> +     /*
>>> +      * Save the sizes reserved by module_frob_arch_sections() so they can
>>> +      * be restored if we retry below.
>>> +      */
>>> +     for_each_mod_mem_type(type)
>>> +             frob_size[type] = info->mod->mem[type].size;
>>> +
>>>       /*
>>>        * Determine total sizes, and put offsets in sh_entsize.  For now
>>>        * this is done generically; there doesn't appear to be any
>>>        * special cases for the architectures.
>>>        */
>>> +retry:
>>>       layout_sections(info->mod, info);
>>>       layout_symtab(info->mod, info);
>>>
>>>       /* Allocate and move to the final place */
>>>       err = move_module(info->mod, info);
>>> -     if (err)
>>> -             return ERR_PTR(err);
>>> +     if (err) {
>>> +             if (err != -EAGAIN)
>>> +                     return ERR_PTR(err);
>>
>> The move_module() logic is non-trivial. -EAGAIN could be returned by
>> other code, now or in the future.
> 
> That's a good point.
> 
>>
>>> +             /*
>>> +              * -EAGAIN means profiling was disabled but the module
>>> +              * can still load without it. Reset state and retry.
>>> +              */
>>> +             rewrite_section_headers(info, flags);
>>> +             for_each_mod_mem_type(type)
>>> +                     info->mod->mem[type].size = frob_size[type];
>>> +             info->sechdrs[info->index.sym].sh_flags &= ~(unsigned long)SHF_ALLOC;
>>> +             info->sechdrs[info->index.str].sh_flags &= ~(unsigned long)SHF_ALLOC;
>>
>> Why is it necessary to reset SHF_ALLOC for .symtab and .strtab here?
> 
> I believe layout_symtab() sets that bit and to retry we need to reset
> it. But I might be wrong here.
> 

Thanks Suren.

And yes, IMHO layout_symtab() is the reason. In the module ELF, .symtab and.strtab
carry no flags at all. layout_symtab() sets SHF_ALLOC on them so that move_module()
will copy them, and it places them itself at the end of MOD_INIT_DATA.

Without the clearing, __layout_sections() on the second pass would pick the two
sections up again: SHF_ALLOC set, no SHF_WRITE, so they match its RODATA mask and get
some of MOD_RODATA reserved. Then layout_symtab() runs anyway, overwrites sh_entsize
and puts them into MOD_INIT_DATA, same as the first pass. The reserved MOD_RODATA is
never used by anything, so the module would just carry that dead space for no reason.
Hence the clearing.

Or am I missing any details?

>>
>>> +             goto retry;
>>> +     }
>>>
>>>       /* Module has been copied to its final place now: return it. */
>>>       mod = (void *)info->sechdrs[info->index.mod].sh_addr;
>>
>> I'm not sure this is the best approach. It's complex logic for what
>> appears to be an edge case related to a debugging facility. It will have
>> the usual problem of error paths not getting enough testing and breaking
>> subtly over time.
>>
>> An alternative could be to reset SHF_ALLOC on the codetag section to
>> remove it from further processing and have relocations that point to
>> this section resolve to something else. It seems that alloc_hooks_tag()
>> could tolerate this, since it only needs to reference the associated
>> alloc_tag when mem_alloc_profiling_enabled() is true and that gets
>> disabled by reserve_module_tags() on the overflow.
> 
> Hmm, yeah if we redirect the references into .codetag.alloc_tags, that
> would be much better.
> 
>>
>> It is also not an ideal approach, but I feel it could be less intrusive
>> to the module loader. I can put together a prototype if needed.
> 
> If your approach does not cause module loading to fail when we disable
> profiling, then that sounds like a good idea. If it's not too much
> trouble, could you please send an RFC?
> 
>>
>> --
>> Thanks,
>> Petr


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-16 15:16       ` Suren Baghdasaryan
  2026-08-17  2:24         ` Hao Ge
@ 2026-08-26  8:32         ` Petr Pavlu
  2026-08-27  0:47           ` Suren Baghdasaryan
  1 sibling, 1 reply; 11+ messages in thread
From: Petr Pavlu @ 2026-08-26  8:32 UTC (permalink / raw)
  To: Suren Baghdasaryan, Hao Ge
  Cc: Andrew Morton, Luis Chamberlain, Daniel Gomez, Sami Tolvanen,
	Aaron Tomlin, linux-modules, linux-kernel, linux-mm, stable

On 8/16/26 5:16 PM, Suren Baghdasaryan wrote:
> On Sat, Aug 15, 2026 at 3:45 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>> On 8/12/26 7:41 AM, Hao Ge wrote:
>>> In reserve_module_tags(), the tag overflow check is gated on
>>> mem_alloc_profiling_enabled():
>>>
>>>     if (mem_alloc_profiling_enabled() && !tags_addressable())
>>>
>>> If profiling is toggled off at runtime and a module is loaded whose
>>> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
>>> skipped. vm_module_tags_populate() still maps memory for the tags and
>>> the module loads successfully, but the total tag count now exceeds what
>>> NR_UNUSED_PAGEFLAG_BITS can address.
>>>
>>> Once profiling is re-enabled, ref_to_idx() computes each tag's index
>>> as its position in the alloc_tag array. update_page_tag_ref() masks
>>> it to alloc_tag_ref_mask before storing in page->flags. Indices
>>> beyond the mask are truncated and idx_to_ref() resolves them to wrong
>>> tags.
>>>
>>> This silently corrupts /proc/allocinfo: allocated pages get attributed
>>> to the wrong call sites, so the statistics it reports are wrong.
>>>
>>> mem_alloc_profiling_enabled() and mem_profiling_compressed are
>>> independent. Once compressed mode is established at boot, it stays
>>> active regardless of runtime toggles of mem_profiling.
>>>
>>> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
>>> profiling, release the reservation, and return -EAGAIN so that
>>> layout_and_allocate() retries with profiling disabled: codetag sections
>>> are then placed as regular module data and the module loads without
>>> profiling rather than being rejected entirely.
>>
>> When the described overflow occurs, why should codetag sections be
>> placed as regular module data? Will the codetag support use them in any
>> way, or do they simply waste space? Is the issue that alloc_hooks()
>> creates relocations pointing into .codetag.alloc_tags?
> 
> Correct, alloc_hooks() will have references into .codetag.alloc_tags.
> With mem_profiling_support=false they should technically never be used
> but I don't think it's a good idea to skip .codetag.alloc_tags section
> allocation and to leave dangling pointers. Also the case described
> here is an outlier, so optimizing it would not yield much benefit.
> 
[...]
>> I'm not sure this is the best approach. It's complex logic for what
>> appears to be an edge case related to a debugging facility. It will have
>> the usual problem of error paths not getting enough testing and breaking
>> subtly over time.
>>
>> An alternative could be to reset SHF_ALLOC on the codetag section to
>> remove it from further processing and have relocations that point to
>> this section resolve to something else. It seems that alloc_hooks_tag()
>> could tolerate this, since it only needs to reference the associated
>> alloc_tag when mem_alloc_profiling_enabled() is true and that gets
>> disabled by reserve_module_tags() on the overflow.
> 
> Hmm, yeah if we redirect the references into .codetag.alloc_tags, that
> would be much better.
> 
>>
>> It is also not an ideal approach, but I feel it could be less intrusive
>> to the module loader. I can put together a prototype if needed.
> 
> If your approach does not cause module loading to fail when we disable
> profiling, then that sounds like a good idea. If it's not too much
> trouble, could you please send an RFC?

The alternative approach I mentioned unfortunately doesn't work well,
since redirecting all relocations against .codetag.alloc_tag to
a different destination is nontrivial. It would require introducing
something like frob_relocation() that is called from each
architecture-specific apply_relocate()/apply_relocate_add() after the
addend has been decoded.

Another option I realized is to change the order in which module
sections are allocated. Rather than interleaving the allocation of
codetag and regular sections, the module loader could first try to
allocate codetag sections and then allocate regular sections. If
allocation of a codetag section fails, it can naturally fall back to
being treated as a regular section. This avoids retrying the allocation
process, which I would prefer to avoid.

A prototype is below.

-- 
Thanks,
Petr


diff --git a/include/linux/module.h b/include/linux/module.h
index 96cc98568eea..0c6f32ddcbf2 100644
--- a/include/linux/module.h
+++ b/include/linux/module.h
@@ -325,6 +325,8 @@ enum mod_mem_type {
 	MOD_INIT_RODATA,
 
 	MOD_MEM_NUM_TYPES,
+
+	MOD_STANDALONE = -2,
 	MOD_INVALID = -1,
 };
 
diff --git a/kernel/module/internal.h b/kernel/module/internal.h
index 061161cc79d9..217bb540e361 100644
--- a/kernel/module/internal.h
+++ b/kernel/module/internal.h
@@ -29,6 +29,10 @@
 #define SH_ENTSIZE_TYPE_MASK	((1UL << SH_ENTSIZE_TYPE_BITS) - 1)
 #define SH_ENTSIZE_OFFSET_MASK	((1UL << (BITS_PER_LONG - SH_ENTSIZE_TYPE_BITS)) - 1)
 
+#define SH_ENTSIZE_STANDALONE					\
+	(((unsigned long)MOD_STANDALONE & SH_ENTSIZE_TYPE_MASK)	\
+	 << SH_ENTSIZE_TYPE_SHIFT)
+
 /* Maximum number of characters written by module_flags() */
 #define MODULE_FLAGS_BUF_SIZE (TAINT_FLAGS_COUNT + 4)
 
diff --git a/kernel/module/main.c b/kernel/module/main.c
index d0e1e0bd2ad0..a86ae8774cd0 100644
--- a/kernel/module/main.c
+++ b/kernel/module/main.c
@@ -1624,7 +1624,7 @@ static int apply_relocations(struct module *mod, const struct load_info *info)
 		 * ELF template and subsequently copy it to the per-CPU destinations.
 		 */
 		if (!(info->sechdrs[infosec].sh_flags & SHF_ALLOC) &&
-		    (!infosec || infosec != info->index.pcpu))
+		    info->sechdrs[infosec].sh_entsize != SH_ENTSIZE_STANDALONE)
 			continue;
 
 		if (info->sechdrs[i].sh_flags & SHF_RELA_LIVEPATCH)
@@ -1722,20 +1722,6 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
 			if (WARN_ON_ONCE(type == MOD_INVALID))
 				continue;
 
-			/*
-			 * Do not allocate codetag memory as we load it into
-			 * preallocated contiguous memory.
-			 */
-			if (codetag_needs_module_section(mod, sname, s->sh_size)) {
-				/*
-				 * s->sh_entsize won't be used but populate the
-				 * type field to avoid confusion.
-				 */
-				s->sh_entsize = ((unsigned long)(type) & SH_ENTSIZE_TYPE_MASK)
-						<< SH_ENTSIZE_TYPE_SHIFT;
-				continue;
-			}
-
 			s->sh_entsize = module_get_offset_and_type(mod, type, s, i);
 			pr_debug("\t%s\n", sname);
 		}
@@ -1745,16 +1731,10 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
 /*
  * Lay out the SHF_ALLOC sections in a way not dissimilar to how ld
  * might -- code, read-only data, read-write data, small data.  Tally
- * sizes, and place the offsets into sh_entsize fields: high bit means it
- * belongs in init.
+ * sizes, and place the offsets into sh_entsize fields.
  */
 static void layout_sections(struct module *mod, struct load_info *info)
 {
-	unsigned int i;
-
-	for (i = 0; i < info->hdr->e_shnum; i++)
-		info->sechdrs[i].sh_entsize = ~0UL;
-
 	pr_debug("Core section allocation order for %s:\n", mod->name);
 	__layout_sections(mod, info, false);
 
@@ -2800,7 +2780,6 @@ static int move_module(struct module *mod, struct load_info *info)
 {
 	int i, ret;
 	enum mod_mem_type t = MOD_MEM_NUM_TYPES;
-	bool codetag_section_found = false;
 
 	for_each_mod_mem_type(type) {
 		if (!mod->mem[type].size) {
@@ -2818,36 +2797,14 @@ static int move_module(struct module *mod, struct load_info *info)
 	/* Transfer each section which specifies SHF_ALLOC */
 	pr_debug("Final section addresses for %s:\n", mod->name);
 	for (i = 0; i < info->hdr->e_shnum; i++) {
-		void *dest;
 		Elf_Shdr *shdr = &info->sechdrs[i];
-		const char *sname;
+		void *dest;
 
 		if (!(shdr->sh_flags & SHF_ALLOC))
 			continue;
 
-		sname = info->secstrings + shdr->sh_name;
-		/*
-		 * Load codetag sections separately as they might still be used
-		 * after module unload.
-		 */
-		if (codetag_needs_module_section(mod, sname, shdr->sh_size)) {
-			dest = codetag_alloc_module_section(mod, sname, shdr->sh_size,
-					arch_mod_section_prepend(mod, i), shdr->sh_addralign);
-			if (WARN_ON(!dest)) {
-				ret = -EINVAL;
-				goto out_err;
-			}
-			if (IS_ERR(dest)) {
-				ret = PTR_ERR(dest);
-				goto out_err;
-			}
-			codetag_section_found = true;
-		} else {
-			enum mod_mem_type type = shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT;
-			unsigned long offset = shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK;
-
-			dest = mod->mem[type].base + offset;
-		}
+		dest = mod->mem[shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT].base +
+		       (shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK);
 
 		if (shdr->sh_type != SHT_NOBITS) {
 			/*
@@ -2879,8 +2836,6 @@ static int move_module(struct module *mod, struct load_info *info)
 	module_memory_restore_rox(mod);
 	while (t--)
 		module_memory_free(mod, t);
-	if (codetag_section_found)
-		codetag_free_module_sections(mod);
 
 	return ret;
 }
@@ -2951,9 +2906,47 @@ static bool blacklisted(const char *module_name)
 }
 core_param(module_blacklist, module_blacklist, charp, 0400);
 
+/*
+ * Allocate codetag sections separately. They are loaded into preallocated
+ * contiguous memory because they may still be used after the module is
+ * unloaded.
+ *
+ * If the separate allocation overflows and fails, allocate the section normally
+ * so that the module can still be loaded.
+ */
+static void allocate_codetag_sections(struct load_info *info)
+{
+	for (unsigned int i = 1; i < info->hdr->e_shnum; i++) {
+		Elf_Shdr *shdr = &info->sechdrs[i];
+		const char *sname = info->secstrings + shdr->sh_name;
+		void *dest;
+
+		if (!(shdr->sh_flags & SHF_ALLOC) ||
+		    !codetag_needs_module_section(info->mod, sname,
+						  shdr->sh_size))
+			continue;
+
+		dest = codetag_alloc_module_section(
+			info->mod, sname, shdr->sh_size,
+			arch_mod_section_prepend(info->mod, i),
+			shdr->sh_addralign);
+		if (WARN_ON(!dest) || IS_ERR(dest)) {
+			/* Allocate the section as a regular section. */
+			continue;
+		}
+
+		if (shdr->sh_type != SHT_NOBITS)
+			memcpy(dest, (void *)shdr->sh_addr, shdr->sh_size);
+		shdr->sh_addr = (unsigned long)dest;
+		shdr->sh_flags &= ~(unsigned long)SHF_ALLOC;
+		shdr->sh_entsize = SH_ENTSIZE_STANDALONE;
+	}
+}
+
 static struct module *layout_and_allocate(struct load_info *info, int flags)
 {
 	struct module *mod;
+	unsigned int i;
 	int err;
 
 	/* Allow arches to frob section contents and sizes.  */
@@ -2967,9 +2960,6 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
 	if (err < 0)
 		return ERR_PTR(err);
 
-	/* We will do a special allocation for per-cpu sections later. */
-	info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
-
 	/*
 	 * Mark relevant sections as SHF_RO_AFTER_INIT so layout_sections() can
 	 * put them in the right place.
@@ -2977,18 +2967,27 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
 	 */
 	module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
 
-	/*
-	 * Determine total sizes, and put offsets in sh_entsize.  For now
-	 * this is done generically; there doesn't appear to be any
-	 * special cases for the architectures.
-	 */
+	/* Repurpose sh_entsize to track where each section is allocated. */
+	for (i = 0; i < info->hdr->e_shnum; i++)
+		info->sechdrs[i].sh_entsize = ~0UL;
+
+	/* We will do a special allocation for per-cpu sections later. */
+	info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
+	info->sechdrs[info->index.pcpu].sh_entsize = SH_ENTSIZE_STANDALONE;
+
+	/* Allow codetag sections to be allocated separately first. */
+	allocate_codetag_sections(info);
+
+	/* Determine total sizes and put offsets in sh_entsize. */
 	layout_sections(info->mod, info);
 	layout_symtab(info->mod, info);
 
 	/* Allocate and move to the final place */
 	err = move_module(info->mod, info);
-	if (err)
+	if (err) {
+		codetag_free_module_sections(mod);
 		return ERR_PTR(err);
+	}
 
 	/* Module has been copied to its final place now: return it. */
 	mod = (void *)info->sechdrs[info->index.mod].sh_addr;
diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
index 2070e682fe10..112a014d4b89 100644
--- a/mm/alloc_tag.c
+++ b/mm/alloc_tag.c
@@ -950,10 +950,12 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
 		int grow_res;
 
 		module_tags.size = offset + size;
-		if (mem_alloc_profiling_enabled() && !tags_addressable()) {
+		if (!tags_addressable()) {
 			shutdown_mem_profiling(true);
-			pr_warn("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
-				mod->name, NR_UNUSED_PAGEFLAG_BITS);
+			pr_warn_once("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
+				     mod->name, NR_UNUSED_PAGEFLAG_BITS);
+			release_module_tags(mod, false);
+			return ERR_PTR(-EAGAIN);
 		}
 
 		grow_res = vm_module_tags_populate();


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-26  8:32         ` Petr Pavlu
@ 2026-08-27  0:47           ` Suren Baghdasaryan
  2026-08-27  6:03             ` Hao Ge
  0 siblings, 1 reply; 11+ messages in thread
From: Suren Baghdasaryan @ 2026-08-27  0:47 UTC (permalink / raw)
  To: Petr Pavlu
  Cc: Hao Ge, Andrew Morton, Luis Chamberlain, Daniel Gomez,
	Sami Tolvanen, Aaron Tomlin, linux-modules, linux-kernel,
	linux-mm, stable

On Wed, Aug 26, 2026 at 1:32 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>
> On 8/16/26 5:16 PM, Suren Baghdasaryan wrote:
> > On Sat, Aug 15, 2026 at 3:45 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
> >> On 8/12/26 7:41 AM, Hao Ge wrote:
> >>> In reserve_module_tags(), the tag overflow check is gated on
> >>> mem_alloc_profiling_enabled():
> >>>
> >>>     if (mem_alloc_profiling_enabled() && !tags_addressable())
> >>>
> >>> If profiling is toggled off at runtime and a module is loaded whose
> >>> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
> >>> skipped. vm_module_tags_populate() still maps memory for the tags and
> >>> the module loads successfully, but the total tag count now exceeds what
> >>> NR_UNUSED_PAGEFLAG_BITS can address.
> >>>
> >>> Once profiling is re-enabled, ref_to_idx() computes each tag's index
> >>> as its position in the alloc_tag array. update_page_tag_ref() masks
> >>> it to alloc_tag_ref_mask before storing in page->flags. Indices
> >>> beyond the mask are truncated and idx_to_ref() resolves them to wrong
> >>> tags.
> >>>
> >>> This silently corrupts /proc/allocinfo: allocated pages get attributed
> >>> to the wrong call sites, so the statistics it reports are wrong.
> >>>
> >>> mem_alloc_profiling_enabled() and mem_profiling_compressed are
> >>> independent. Once compressed mode is established at boot, it stays
> >>> active regardless of runtime toggles of mem_profiling.
> >>>
> >>> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
> >>> profiling, release the reservation, and return -EAGAIN so that
> >>> layout_and_allocate() retries with profiling disabled: codetag sections
> >>> are then placed as regular module data and the module loads without
> >>> profiling rather than being rejected entirely.
> >>
> >> When the described overflow occurs, why should codetag sections be
> >> placed as regular module data? Will the codetag support use them in any
> >> way, or do they simply waste space? Is the issue that alloc_hooks()
> >> creates relocations pointing into .codetag.alloc_tags?
> >
> > Correct, alloc_hooks() will have references into .codetag.alloc_tags.
> > With mem_profiling_support=false they should technically never be used
> > but I don't think it's a good idea to skip .codetag.alloc_tags section
> > allocation and to leave dangling pointers. Also the case described
> > here is an outlier, so optimizing it would not yield much benefit.
> >
> [...]
> >> I'm not sure this is the best approach. It's complex logic for what
> >> appears to be an edge case related to a debugging facility. It will have
> >> the usual problem of error paths not getting enough testing and breaking
> >> subtly over time.
> >>
> >> An alternative could be to reset SHF_ALLOC on the codetag section to
> >> remove it from further processing and have relocations that point to
> >> this section resolve to something else. It seems that alloc_hooks_tag()
> >> could tolerate this, since it only needs to reference the associated
> >> alloc_tag when mem_alloc_profiling_enabled() is true and that gets
> >> disabled by reserve_module_tags() on the overflow.
> >
> > Hmm, yeah if we redirect the references into .codetag.alloc_tags, that
> > would be much better.
> >
> >>
> >> It is also not an ideal approach, but I feel it could be less intrusive
> >> to the module loader. I can put together a prototype if needed.
> >
> > If your approach does not cause module loading to fail when we disable
> > profiling, then that sounds like a good idea. If it's not too much
> > trouble, could you please send an RFC?
>
> The alternative approach I mentioned unfortunately doesn't work well,
> since redirecting all relocations against .codetag.alloc_tag to
> a different destination is nontrivial. It would require introducing
> something like frob_relocation() that is called from each
> architecture-specific apply_relocate()/apply_relocate_add() after the
> addend has been decoded.
>
> Another option I realized is to change the order in which module
> sections are allocated. Rather than interleaving the allocation of
> codetag and regular sections, the module loader could first try to
> allocate codetag sections and then allocate regular sections. If
> allocation of a codetag section fails, it can naturally fall back to
> being treated as a regular section. This avoids retrying the allocation
> process, which I would prefer to avoid.
>
> A prototype is below.

Thanks for following up on this, Petr!
At first glance, this seems like a much cleaner approach. But it's
also a sizable change, so it will need some testing. I'll try to run
some test scenarios over the weekend.

>
> --
> Thanks,
> Petr
>
>
> diff --git a/include/linux/module.h b/include/linux/module.h
> index 96cc98568eea..0c6f32ddcbf2 100644
> --- a/include/linux/module.h
> +++ b/include/linux/module.h
> @@ -325,6 +325,8 @@ enum mod_mem_type {
>         MOD_INIT_RODATA,
>
>         MOD_MEM_NUM_TYPES,
> +
> +       MOD_STANDALONE = -2,
>         MOD_INVALID = -1,
>  };
>
> diff --git a/kernel/module/internal.h b/kernel/module/internal.h
> index 061161cc79d9..217bb540e361 100644
> --- a/kernel/module/internal.h
> +++ b/kernel/module/internal.h
> @@ -29,6 +29,10 @@
>  #define SH_ENTSIZE_TYPE_MASK   ((1UL << SH_ENTSIZE_TYPE_BITS) - 1)
>  #define SH_ENTSIZE_OFFSET_MASK ((1UL << (BITS_PER_LONG - SH_ENTSIZE_TYPE_BITS)) - 1)
>
> +#define SH_ENTSIZE_STANDALONE                                  \
> +       (((unsigned long)MOD_STANDALONE & SH_ENTSIZE_TYPE_MASK) \
> +        << SH_ENTSIZE_TYPE_SHIFT)
> +
>  /* Maximum number of characters written by module_flags() */
>  #define MODULE_FLAGS_BUF_SIZE (TAINT_FLAGS_COUNT + 4)
>
> diff --git a/kernel/module/main.c b/kernel/module/main.c
> index d0e1e0bd2ad0..a86ae8774cd0 100644
> --- a/kernel/module/main.c
> +++ b/kernel/module/main.c
> @@ -1624,7 +1624,7 @@ static int apply_relocations(struct module *mod, const struct load_info *info)
>                  * ELF template and subsequently copy it to the per-CPU destinations.
>                  */
>                 if (!(info->sechdrs[infosec].sh_flags & SHF_ALLOC) &&
> -                   (!infosec || infosec != info->index.pcpu))
> +                   info->sechdrs[infosec].sh_entsize != SH_ENTSIZE_STANDALONE)
>                         continue;
>
>                 if (info->sechdrs[i].sh_flags & SHF_RELA_LIVEPATCH)
> @@ -1722,20 +1722,6 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
>                         if (WARN_ON_ONCE(type == MOD_INVALID))
>                                 continue;
>
> -                       /*
> -                        * Do not allocate codetag memory as we load it into
> -                        * preallocated contiguous memory.
> -                        */
> -                       if (codetag_needs_module_section(mod, sname, s->sh_size)) {
> -                               /*
> -                                * s->sh_entsize won't be used but populate the
> -                                * type field to avoid confusion.
> -                                */
> -                               s->sh_entsize = ((unsigned long)(type) & SH_ENTSIZE_TYPE_MASK)
> -                                               << SH_ENTSIZE_TYPE_SHIFT;
> -                               continue;
> -                       }
> -
>                         s->sh_entsize = module_get_offset_and_type(mod, type, s, i);
>                         pr_debug("\t%s\n", sname);
>                 }
> @@ -1745,16 +1731,10 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
>  /*
>   * Lay out the SHF_ALLOC sections in a way not dissimilar to how ld
>   * might -- code, read-only data, read-write data, small data.  Tally
> - * sizes, and place the offsets into sh_entsize fields: high bit means it
> - * belongs in init.
> + * sizes, and place the offsets into sh_entsize fields.
>   */
>  static void layout_sections(struct module *mod, struct load_info *info)
>  {
> -       unsigned int i;
> -
> -       for (i = 0; i < info->hdr->e_shnum; i++)
> -               info->sechdrs[i].sh_entsize = ~0UL;
> -
>         pr_debug("Core section allocation order for %s:\n", mod->name);
>         __layout_sections(mod, info, false);
>
> @@ -2800,7 +2780,6 @@ static int move_module(struct module *mod, struct load_info *info)
>  {
>         int i, ret;
>         enum mod_mem_type t = MOD_MEM_NUM_TYPES;
> -       bool codetag_section_found = false;
>
>         for_each_mod_mem_type(type) {
>                 if (!mod->mem[type].size) {
> @@ -2818,36 +2797,14 @@ static int move_module(struct module *mod, struct load_info *info)
>         /* Transfer each section which specifies SHF_ALLOC */
>         pr_debug("Final section addresses for %s:\n", mod->name);
>         for (i = 0; i < info->hdr->e_shnum; i++) {
> -               void *dest;
>                 Elf_Shdr *shdr = &info->sechdrs[i];
> -               const char *sname;
> +               void *dest;
>
>                 if (!(shdr->sh_flags & SHF_ALLOC))
>                         continue;
>
> -               sname = info->secstrings + shdr->sh_name;
> -               /*
> -                * Load codetag sections separately as they might still be used
> -                * after module unload.
> -                */
> -               if (codetag_needs_module_section(mod, sname, shdr->sh_size)) {
> -                       dest = codetag_alloc_module_section(mod, sname, shdr->sh_size,
> -                                       arch_mod_section_prepend(mod, i), shdr->sh_addralign);
> -                       if (WARN_ON(!dest)) {
> -                               ret = -EINVAL;
> -                               goto out_err;
> -                       }
> -                       if (IS_ERR(dest)) {
> -                               ret = PTR_ERR(dest);
> -                               goto out_err;
> -                       }
> -                       codetag_section_found = true;
> -               } else {
> -                       enum mod_mem_type type = shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT;
> -                       unsigned long offset = shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK;
> -
> -                       dest = mod->mem[type].base + offset;
> -               }
> +               dest = mod->mem[shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT].base +
> +                      (shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK);
>
>                 if (shdr->sh_type != SHT_NOBITS) {
>                         /*
> @@ -2879,8 +2836,6 @@ static int move_module(struct module *mod, struct load_info *info)
>         module_memory_restore_rox(mod);
>         while (t--)
>                 module_memory_free(mod, t);
> -       if (codetag_section_found)
> -               codetag_free_module_sections(mod);
>
>         return ret;
>  }
> @@ -2951,9 +2906,47 @@ static bool blacklisted(const char *module_name)
>  }
>  core_param(module_blacklist, module_blacklist, charp, 0400);
>
> +/*
> + * Allocate codetag sections separately. They are loaded into preallocated
> + * contiguous memory because they may still be used after the module is
> + * unloaded.
> + *
> + * If the separate allocation overflows and fails, allocate the section normally
> + * so that the module can still be loaded.
> + */
> +static void allocate_codetag_sections(struct load_info *info)
> +{
> +       for (unsigned int i = 1; i < info->hdr->e_shnum; i++) {
> +               Elf_Shdr *shdr = &info->sechdrs[i];
> +               const char *sname = info->secstrings + shdr->sh_name;
> +               void *dest;
> +
> +               if (!(shdr->sh_flags & SHF_ALLOC) ||
> +                   !codetag_needs_module_section(info->mod, sname,
> +                                                 shdr->sh_size))
> +                       continue;
> +
> +               dest = codetag_alloc_module_section(
> +                       info->mod, sname, shdr->sh_size,
> +                       arch_mod_section_prepend(info->mod, i),
> +                       shdr->sh_addralign);
> +               if (WARN_ON(!dest) || IS_ERR(dest)) {
> +                       /* Allocate the section as a regular section. */
> +                       continue;
> +               }
> +
> +               if (shdr->sh_type != SHT_NOBITS)
> +                       memcpy(dest, (void *)shdr->sh_addr, shdr->sh_size);
> +               shdr->sh_addr = (unsigned long)dest;
> +               shdr->sh_flags &= ~(unsigned long)SHF_ALLOC;
> +               shdr->sh_entsize = SH_ENTSIZE_STANDALONE;
> +       }
> +}
> +
>  static struct module *layout_and_allocate(struct load_info *info, int flags)
>  {
>         struct module *mod;
> +       unsigned int i;
>         int err;
>
>         /* Allow arches to frob section contents and sizes.  */
> @@ -2967,9 +2960,6 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>         if (err < 0)
>                 return ERR_PTR(err);
>
> -       /* We will do a special allocation for per-cpu sections later. */
> -       info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
> -
>         /*
>          * Mark relevant sections as SHF_RO_AFTER_INIT so layout_sections() can
>          * put them in the right place.
> @@ -2977,18 +2967,27 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>          */
>         module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
>
> -       /*
> -        * Determine total sizes, and put offsets in sh_entsize.  For now
> -        * this is done generically; there doesn't appear to be any
> -        * special cases for the architectures.
> -        */
> +       /* Repurpose sh_entsize to track where each section is allocated. */
> +       for (i = 0; i < info->hdr->e_shnum; i++)
> +               info->sechdrs[i].sh_entsize = ~0UL;
> +
> +       /* We will do a special allocation for per-cpu sections later. */
> +       info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
> +       info->sechdrs[info->index.pcpu].sh_entsize = SH_ENTSIZE_STANDALONE;
> +
> +       /* Allow codetag sections to be allocated separately first. */
> +       allocate_codetag_sections(info);
> +
> +       /* Determine total sizes and put offsets in sh_entsize. */
>         layout_sections(info->mod, info);
>         layout_symtab(info->mod, info);
>
>         /* Allocate and move to the final place */
>         err = move_module(info->mod, info);
> -       if (err)
> +       if (err) {
> +               codetag_free_module_sections(mod);
>                 return ERR_PTR(err);
> +       }
>
>         /* Module has been copied to its final place now: return it. */
>         mod = (void *)info->sechdrs[info->index.mod].sh_addr;
> diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
> index 2070e682fe10..112a014d4b89 100644
> --- a/mm/alloc_tag.c
> +++ b/mm/alloc_tag.c
> @@ -950,10 +950,12 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
>                 int grow_res;
>
>                 module_tags.size = offset + size;
> -               if (mem_alloc_profiling_enabled() && !tags_addressable()) {
> +               if (!tags_addressable()) {
>                         shutdown_mem_profiling(true);
> -                       pr_warn("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
> -                               mod->name, NR_UNUSED_PAGEFLAG_BITS);
> +                       pr_warn_once("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
> +                                    mod->name, NR_UNUSED_PAGEFLAG_BITS);
> +                       release_module_tags(mod, false);
> +                       return ERR_PTR(-EAGAIN);
>                 }
>
>                 grow_res = vm_module_tags_populate();


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled
  2026-08-27  0:47           ` Suren Baghdasaryan
@ 2026-08-27  6:03             ` Hao Ge
  0 siblings, 0 replies; 11+ messages in thread
From: Hao Ge @ 2026-08-27  6:03 UTC (permalink / raw)
  To: Suren Baghdasaryan, Petr Pavlu
  Cc: Andrew Morton, Luis Chamberlain, Daniel Gomez, Sami Tolvanen,
	Aaron Tomlin, linux-modules, linux-kernel, linux-mm, stable

Hi Suren and Petr


On 2026/8/27 08:47, Suren Baghdasaryan wrote:
> On Wed, Aug 26, 2026 at 1:32 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>>
>> On 8/16/26 5:16 PM, Suren Baghdasaryan wrote:
>>> On Sat, Aug 15, 2026 at 3:45 AM Petr Pavlu <petr.pavlu@suse.com> wrote:
>>>> On 8/12/26 7:41 AM, Hao Ge wrote:
>>>>> In reserve_module_tags(), the tag overflow check is gated on
>>>>> mem_alloc_profiling_enabled():
>>>>>
>>>>>     if (mem_alloc_profiling_enabled() && !tags_addressable())
>>>>>
>>>>> If profiling is toggled off at runtime and a module is loaded whose
>>>>> tags exceed the compressed-mode limit, shutdown_mem_profiling() is
>>>>> skipped. vm_module_tags_populate() still maps memory for the tags and
>>>>> the module loads successfully, but the total tag count now exceeds what
>>>>> NR_UNUSED_PAGEFLAG_BITS can address.
>>>>>
>>>>> Once profiling is re-enabled, ref_to_idx() computes each tag's index
>>>>> as its position in the alloc_tag array. update_page_tag_ref() masks
>>>>> it to alloc_tag_ref_mask before storing in page->flags. Indices
>>>>> beyond the mask are truncated and idx_to_ref() resolves them to wrong
>>>>> tags.
>>>>>
>>>>> This silently corrupts /proc/allocinfo: allocated pages get attributed
>>>>> to the wrong call sites, so the statistics it reports are wrong.
>>>>>
>>>>> mem_alloc_profiling_enabled() and mem_profiling_compressed are
>>>>> independent. Once compressed mode is established at boot, it stays
>>>>> active regardless of runtime toggles of mem_profiling.
>>>>>
>>>>> Remove the mem_alloc_profiling_enabled() guard. On overflow, shut down
>>>>> profiling, release the reservation, and return -EAGAIN so that
>>>>> layout_and_allocate() retries with profiling disabled: codetag sections
>>>>> are then placed as regular module data and the module loads without
>>>>> profiling rather than being rejected entirely.
>>>>
>>>> When the described overflow occurs, why should codetag sections be
>>>> placed as regular module data? Will the codetag support use them in any
>>>> way, or do they simply waste space? Is the issue that alloc_hooks()
>>>> creates relocations pointing into .codetag.alloc_tags?
>>>
>>> Correct, alloc_hooks() will have references into .codetag.alloc_tags.
>>> With mem_profiling_support=false they should technically never be used
>>> but I don't think it's a good idea to skip .codetag.alloc_tags section
>>> allocation and to leave dangling pointers. Also the case described
>>> here is an outlier, so optimizing it would not yield much benefit.
>>>
>> [...]
>>>> I'm not sure this is the best approach. It's complex logic for what
>>>> appears to be an edge case related to a debugging facility. It will have
>>>> the usual problem of error paths not getting enough testing and breaking
>>>> subtly over time.
>>>>
>>>> An alternative could be to reset SHF_ALLOC on the codetag section to
>>>> remove it from further processing and have relocations that point to
>>>> this section resolve to something else. It seems that alloc_hooks_tag()
>>>> could tolerate this, since it only needs to reference the associated
>>>> alloc_tag when mem_alloc_profiling_enabled() is true and that gets
>>>> disabled by reserve_module_tags() on the overflow.
>>>
>>> Hmm, yeah if we redirect the references into .codetag.alloc_tags, that
>>> would be much better.
>>>
>>>>
>>>> It is also not an ideal approach, but I feel it could be less intrusive
>>>> to the module loader. I can put together a prototype if needed.
>>>
>>> If your approach does not cause module loading to fail when we disable
>>> profiling, then that sounds like a good idea. If it's not too much
>>> trouble, could you please send an RFC?
>>
>> The alternative approach I mentioned unfortunately doesn't work well,
>> since redirecting all relocations against .codetag.alloc_tag to
>> a different destination is nontrivial. It would require introducing
>> something like frob_relocation() that is called from each
>> architecture-specific apply_relocate()/apply_relocate_add() after the
>> addend has been decoded.
>>
>> Another option I realized is to change the order in which module
>> sections are allocated. Rather than interleaving the allocation of
>> codetag and regular sections, the module loader could first try to
>> allocate codetag sections and then allocate regular sections. If
>> allocation of a codetag section fails, it can naturally fall back to
>> being treated as a regular section. This avoids retrying the allocation
>> process, which I would prefer to avoid.
>>
>> A prototype is below.
> 
> Thanks for following up on this, Petr!

+1

> At first glance, this seems like a much cleaner approach. But it's
> also a sizable change, so it will need some testing. I'll try to run
> some test scenarios over the weekend.

I believe Petr's approach can also address the race problem pointed out by this patch:
https://lore.kernel.org/all/20260813093421.135230-3-hao.ge@linux.dev/
I will also go through this patch and run some local tests as soon as possible.

Thanks
Best Regards
Hao

> 
>>
>> --
>> Thanks,
>> Petr
>>
>>
>> diff --git a/include/linux/module.h b/include/linux/module.h
>> index 96cc98568eea..0c6f32ddcbf2 100644
>> --- a/include/linux/module.h
>> +++ b/include/linux/module.h
>> @@ -325,6 +325,8 @@ enum mod_mem_type {
>>         MOD_INIT_RODATA,
>>
>>         MOD_MEM_NUM_TYPES,
>> +
>> +       MOD_STANDALONE = -2,
>>         MOD_INVALID = -1,
>>  };
>>
>> diff --git a/kernel/module/internal.h b/kernel/module/internal.h
>> index 061161cc79d9..217bb540e361 100644
>> --- a/kernel/module/internal.h
>> +++ b/kernel/module/internal.h
>> @@ -29,6 +29,10 @@
>>  #define SH_ENTSIZE_TYPE_MASK   ((1UL << SH_ENTSIZE_TYPE_BITS) - 1)
>>  #define SH_ENTSIZE_OFFSET_MASK ((1UL << (BITS_PER_LONG - SH_ENTSIZE_TYPE_BITS)) - 1)
>>
>> +#define SH_ENTSIZE_STANDALONE                                  \
>> +       (((unsigned long)MOD_STANDALONE & SH_ENTSIZE_TYPE_MASK) \
>> +        << SH_ENTSIZE_TYPE_SHIFT)
>> +
>>  /* Maximum number of characters written by module_flags() */
>>  #define MODULE_FLAGS_BUF_SIZE (TAINT_FLAGS_COUNT + 4)
>>
>> diff --git a/kernel/module/main.c b/kernel/module/main.c
>> index d0e1e0bd2ad0..a86ae8774cd0 100644
>> --- a/kernel/module/main.c
>> +++ b/kernel/module/main.c
>> @@ -1624,7 +1624,7 @@ static int apply_relocations(struct module *mod, const struct load_info *info)
>>                  * ELF template and subsequently copy it to the per-CPU destinations.
>>                  */
>>                 if (!(info->sechdrs[infosec].sh_flags & SHF_ALLOC) &&
>> -                   (!infosec || infosec != info->index.pcpu))
>> +                   info->sechdrs[infosec].sh_entsize != SH_ENTSIZE_STANDALONE)
>>                         continue;
>>
>>                 if (info->sechdrs[i].sh_flags & SHF_RELA_LIVEPATCH)
>> @@ -1722,20 +1722,6 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
>>                         if (WARN_ON_ONCE(type == MOD_INVALID))
>>                                 continue;
>>
>> -                       /*
>> -                        * Do not allocate codetag memory as we load it into
>> -                        * preallocated contiguous memory.
>> -                        */
>> -                       if (codetag_needs_module_section(mod, sname, s->sh_size)) {
>> -                               /*
>> -                                * s->sh_entsize won't be used but populate the
>> -                                * type field to avoid confusion.
>> -                                */
>> -                               s->sh_entsize = ((unsigned long)(type) & SH_ENTSIZE_TYPE_MASK)
>> -                                               << SH_ENTSIZE_TYPE_SHIFT;
>> -                               continue;
>> -                       }
>> -
>>                         s->sh_entsize = module_get_offset_and_type(mod, type, s, i);
>>                         pr_debug("\t%s\n", sname);
>>                 }
>> @@ -1745,16 +1731,10 @@ static void __layout_sections(struct module *mod, struct load_info *info, bool i
>>  /*
>>   * Lay out the SHF_ALLOC sections in a way not dissimilar to how ld
>>   * might -- code, read-only data, read-write data, small data.  Tally
>> - * sizes, and place the offsets into sh_entsize fields: high bit means it
>> - * belongs in init.
>> + * sizes, and place the offsets into sh_entsize fields.
>>   */
>>  static void layout_sections(struct module *mod, struct load_info *info)
>>  {
>> -       unsigned int i;
>> -
>> -       for (i = 0; i < info->hdr->e_shnum; i++)
>> -               info->sechdrs[i].sh_entsize = ~0UL;
>> -
>>         pr_debug("Core section allocation order for %s:\n", mod->name);
>>         __layout_sections(mod, info, false);
>>
>> @@ -2800,7 +2780,6 @@ static int move_module(struct module *mod, struct load_info *info)
>>  {
>>         int i, ret;
>>         enum mod_mem_type t = MOD_MEM_NUM_TYPES;
>> -       bool codetag_section_found = false;
>>
>>         for_each_mod_mem_type(type) {
>>                 if (!mod->mem[type].size) {
>> @@ -2818,36 +2797,14 @@ static int move_module(struct module *mod, struct load_info *info)
>>         /* Transfer each section which specifies SHF_ALLOC */
>>         pr_debug("Final section addresses for %s:\n", mod->name);
>>         for (i = 0; i < info->hdr->e_shnum; i++) {
>> -               void *dest;
>>                 Elf_Shdr *shdr = &info->sechdrs[i];
>> -               const char *sname;
>> +               void *dest;
>>
>>                 if (!(shdr->sh_flags & SHF_ALLOC))
>>                         continue;
>>
>> -               sname = info->secstrings + shdr->sh_name;
>> -               /*
>> -                * Load codetag sections separately as they might still be used
>> -                * after module unload.
>> -                */
>> -               if (codetag_needs_module_section(mod, sname, shdr->sh_size)) {
>> -                       dest = codetag_alloc_module_section(mod, sname, shdr->sh_size,
>> -                                       arch_mod_section_prepend(mod, i), shdr->sh_addralign);
>> -                       if (WARN_ON(!dest)) {
>> -                               ret = -EINVAL;
>> -                               goto out_err;
>> -                       }
>> -                       if (IS_ERR(dest)) {
>> -                               ret = PTR_ERR(dest);
>> -                               goto out_err;
>> -                       }
>> -                       codetag_section_found = true;
>> -               } else {
>> -                       enum mod_mem_type type = shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT;
>> -                       unsigned long offset = shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK;
>> -
>> -                       dest = mod->mem[type].base + offset;
>> -               }
>> +               dest = mod->mem[shdr->sh_entsize >> SH_ENTSIZE_TYPE_SHIFT].base +
>> +                      (shdr->sh_entsize & SH_ENTSIZE_OFFSET_MASK);
>>
>>                 if (shdr->sh_type != SHT_NOBITS) {
>>                         /*
>> @@ -2879,8 +2836,6 @@ static int move_module(struct module *mod, struct load_info *info)
>>         module_memory_restore_rox(mod);
>>         while (t--)
>>                 module_memory_free(mod, t);
>> -       if (codetag_section_found)
>> -               codetag_free_module_sections(mod);
>>
>>         return ret;
>>  }
>> @@ -2951,9 +2906,47 @@ static bool blacklisted(const char *module_name)
>>  }
>>  core_param(module_blacklist, module_blacklist, charp, 0400);
>>
>> +/*
>> + * Allocate codetag sections separately. They are loaded into preallocated
>> + * contiguous memory because they may still be used after the module is
>> + * unloaded.
>> + *
>> + * If the separate allocation overflows and fails, allocate the section normally
>> + * so that the module can still be loaded.
>> + */
>> +static void allocate_codetag_sections(struct load_info *info)
>> +{
>> +       for (unsigned int i = 1; i < info->hdr->e_shnum; i++) {
>> +               Elf_Shdr *shdr = &info->sechdrs[i];
>> +               const char *sname = info->secstrings + shdr->sh_name;
>> +               void *dest;
>> +
>> +               if (!(shdr->sh_flags & SHF_ALLOC) ||
>> +                   !codetag_needs_module_section(info->mod, sname,
>> +                                                 shdr->sh_size))
>> +                       continue;
>> +
>> +               dest = codetag_alloc_module_section(
>> +                       info->mod, sname, shdr->sh_size,
>> +                       arch_mod_section_prepend(info->mod, i),
>> +                       shdr->sh_addralign);
>> +               if (WARN_ON(!dest) || IS_ERR(dest)) {
>> +                       /* Allocate the section as a regular section. */
>> +                       continue;
>> +               }
>> +
>> +               if (shdr->sh_type != SHT_NOBITS)
>> +                       memcpy(dest, (void *)shdr->sh_addr, shdr->sh_size);
>> +               shdr->sh_addr = (unsigned long)dest;
>> +               shdr->sh_flags &= ~(unsigned long)SHF_ALLOC;
>> +               shdr->sh_entsize = SH_ENTSIZE_STANDALONE;
>> +       }
>> +}
>> +
>>  static struct module *layout_and_allocate(struct load_info *info, int flags)
>>  {
>>         struct module *mod;
>> +       unsigned int i;
>>         int err;
>>
>>         /* Allow arches to frob section contents and sizes.  */
>> @@ -2967,9 +2960,6 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>>         if (err < 0)
>>                 return ERR_PTR(err);
>>
>> -       /* We will do a special allocation for per-cpu sections later. */
>> -       info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
>> -
>>         /*
>>          * Mark relevant sections as SHF_RO_AFTER_INIT so layout_sections() can
>>          * put them in the right place.
>> @@ -2977,18 +2967,27 @@ static struct module *layout_and_allocate(struct load_info *info, int flags)
>>          */
>>         module_mark_ro_after_init(info->hdr, info->sechdrs, info->secstrings);
>>
>> -       /*
>> -        * Determine total sizes, and put offsets in sh_entsize.  For now
>> -        * this is done generically; there doesn't appear to be any
>> -        * special cases for the architectures.
>> -        */
>> +       /* Repurpose sh_entsize to track where each section is allocated. */
>> +       for (i = 0; i < info->hdr->e_shnum; i++)
>> +               info->sechdrs[i].sh_entsize = ~0UL;
>> +
>> +       /* We will do a special allocation for per-cpu sections later. */
>> +       info->sechdrs[info->index.pcpu].sh_flags &= ~(unsigned long)SHF_ALLOC;
>> +       info->sechdrs[info->index.pcpu].sh_entsize = SH_ENTSIZE_STANDALONE;
>> +
>> +       /* Allow codetag sections to be allocated separately first. */
>> +       allocate_codetag_sections(info);
>> +
>> +       /* Determine total sizes and put offsets in sh_entsize. */
>>         layout_sections(info->mod, info);
>>         layout_symtab(info->mod, info);
>>
>>         /* Allocate and move to the final place */
>>         err = move_module(info->mod, info);
>> -       if (err)
>> +       if (err) {
>> +               codetag_free_module_sections(mod);
>>                 return ERR_PTR(err);
>> +       }
>>
>>         /* Module has been copied to its final place now: return it. */
>>         mod = (void *)info->sechdrs[info->index.mod].sh_addr;
>> diff --git a/mm/alloc_tag.c b/mm/alloc_tag.c
>> index 2070e682fe10..112a014d4b89 100644
>> --- a/mm/alloc_tag.c
>> +++ b/mm/alloc_tag.c
>> @@ -950,10 +950,12 @@ static void *reserve_module_tags(struct module *mod, unsigned long size,
>>                 int grow_res;
>>
>>                 module_tags.size = offset + size;
>> -               if (mem_alloc_profiling_enabled() && !tags_addressable()) {
>> +               if (!tags_addressable()) {
>>                         shutdown_mem_profiling(true);
>> -                       pr_warn("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
>> -                               mod->name, NR_UNUSED_PAGEFLAG_BITS);
>> +                       pr_warn_once("With module %s there are too many tags to fit in %d page flag bits. Memory allocation profiling is disabled!\n",
>> +                                    mod->name, NR_UNUSED_PAGEFLAG_BITS);
>> +                       release_module_tags(mod, false);
>> +                       return ERR_PTR(-EAGAIN);
>>                 }
>>
>>                 grow_res = vm_module_tags_populate();


^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-08-27  6:02 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-12  5:41 [PATCH v5 0/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
2026-08-12  5:41 ` [PATCH v5 1/2] alloc_tag: move release_module_tags() above reserve_module_tags() Hao Ge
2026-08-15  3:40   ` Suren Baghdasaryan
2026-08-12  5:41 ` [PATCH v5 2/2] alloc_tag: fix undetected compressed tag overflow when profiling is disabled Hao Ge
2026-08-15  5:58   ` Suren Baghdasaryan
2026-08-15 10:45     ` Petr Pavlu
2026-08-16 15:16       ` Suren Baghdasaryan
2026-08-17  2:24         ` Hao Ge
2026-08-26  8:32         ` Petr Pavlu
2026-08-27  0:47           ` Suren Baghdasaryan
2026-08-27  6:03             ` Hao Ge

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox