* Re: [PATCH 6.12.y] mm/damon/core: skip aging from repeated aggressive merging
2026-09-09 4:38 ` [PATCH 6.12.y] mm/damon/core: skip aging from repeated aggressive merging SJ Park
@ 2026-09-09 4:46 ` sashiko-bot
2026-09-11 11:20 ` Sasha Levin
1 sibling, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-09-09 4:46 UTC (permalink / raw)
To: SJ Park; +Cc: damon
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
Pre-existing issues:
- [High] The do-while loop in `kdamond_merge_regions` can infinite-loop if `nr_regions` remains above `max_nr_regions`. It expects the loop to terminate when `threshold / 2 >= max_thres`, but `threshold` is a 32-bit `unsigned int` that overflows when doubled. If `max_thres` (derived from `aggr_interval / sample_interval`) is sufficiently large (e.g., > 1,073,741,824), the condition `threshold / 2 < max_thres` evaluates to true forever, locking up the kernel thread.
- [Medium] The C macro `abs()` is incorrectly used to compute the absolute difference between two `unsigned int` access counts (`r->nr_accesses`, `r->last_nr_accesses`, and `prev->nr_accesses`). For very large differences (exceeding 2 billion), unsigned wrap-around causes the result to evaluate as a negative 32-bit signed integer inside `abs()`, which negates it to a completely inaccurate value, causing DAMON to incorrectly reset region ages or improperly merge disjoint access patterns.
--
commit 2e588ff7cb5f9fc59f1692ecf781e53ba68d8c92
Author: SJ Park <sj@kernel.org>
mm/damon/core: skip aging from repeated aggressive merging
This commit fixes a corner case where kdamond_merge_regions() repeats
merging regions and updates region ages multiple times, which corrupts
monitoring results. It resolves the issue by skipping age counting during
the repeated merge operations.
Link: https://lore.kernel.org/20260712165432.87609-1-sj@kernel.org
Link: https://lore.kernel.org/20260621203548.10718-1-sj@kernel.org [1]
Link: https://lore.kernel.org/20260709145425.96247-1-sj@kernel.org [2]
> diff --git a/mm/damon/core.c b/mm/damon/core.c
> --- a/mm/damon/core.c
> +++ b/mm/damon/core.c
> @@ -1786,15 +1786,17 @@ static void damon_merge_two_regions(struct damon_target *t,
> * sz_limit size upper limit of each region
> */
> static void damon_merge_regions_of(struct damon_target *t, unsigned int thres,
> - unsigned long sz_limit)
> + unsigned long sz_limit, bool count_age)
> {
> struct damon_region *r, *prev = NULL, *next;
>
> damon_for_each_region_safe(r, next, t) {
> - if (abs(r->nr_accesses - r->last_nr_accesses) > thres)
> - r->age = 0;
> - else
> - r->age++;
> + if (count_age) {
> + if (abs(r->nr_accesses - r->last_nr_accesses) > thres)
[Severity: Medium]
This is a pre-existing issue, but does using the abs() macro here correctly
handle the absolute difference for large unsigned integer access counters?
For very large differences exceeding 2 billion, the unsigned wrap-around
could cause the result to evaluate as a negative 32-bit signed integer
inside abs(), which then negates it to an inaccurate positive value.
Could this cause DAMON to incorrectly reset region ages or improperly
merge disjoint access patterns when the ratio of aggr_interval to
sample_interval is very large? Would abs_diff() be a safer alternative here?
> + r->age = 0;
> + else
> + r->age++;
> + }
>
> if (prev && prev->ar.end == r->ar.start &&
> abs(prev->nr_accesses - r->nr_accesses) <= thres &&
[Severity: Medium]
This is a pre-existing issue, but similarly, could this abs() calculation
result in false positives during the threshold check, forcing aggressive
merges of disparate regions if the access frequencies are extremely large?
> damon_sz_region(prev) + damon_sz_region(r) <= sz_limit)
> damon_merge_two_regions(t, prev, r);
> else
> prev = r;
> }
> }
[ ... ]
> @@ -1828,15 +1830,18 @@ static void kdamond_merge_regions(struct damon_ctx *c, unsigned int threshold,
> struct damon_target *t;
> unsigned int nr_regions;
> unsigned int max_thres;
> + bool count_age = true;
>
> max_thres = c->attrs.aggr_interval /
> (c->attrs.sample_interval ? c->attrs.sample_interval : 1);
> do {
> nr_regions = 0;
> damon_for_each_target(t, c) {
> - damon_merge_regions_of(t, threshold, sz_limit);
> + damon_merge_regions_of(t, threshold, sz_limit,
> + count_age);
> nr_regions += damon_nr_regions(t);
> }
> + count_age = false;
> threshold = max(1, threshold * 2);
> } while (nr_regions > c->attrs.max_nr_regions &&
> threshold / 2 < max_thres);
[Severity: High]
This is a pre-existing issue, but can this do-while loop lock up the
kdamond kernel thread indefinitely?
The loop expects to terminate when threshold / 2 >= max_thres. However,
threshold is a 32-bit unsigned int that overflows when doubled.
If max_thres is derived to be a very large value during the assignment
at the start of kdamond_merge_regions() (e.g., if a user configures a
large aggregation and sample interval ratio), threshold could repeatedly
double until it exceeds 0x7FFFFFFF.
When it overflows, max(1, threshold * 2) resets the threshold to 1.
As a result, threshold / 2 would max out at 0x40000000, causing the
condition threshold / 2 < max_thres to remain permanently true.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260909043807.210326-1-sj@kernel.org?part=1
^ permalink raw reply [flat|nested] 3+ messages in thread