All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kunwu Chan <kunwu.chan@gmail.com>
To: "Barry Song (Xiaomi)" <baohua@kernel.org>
Cc: Kunwu Chan <kunwu.chan@linux.dev>,
	akpm@linux-foundation.org, lianux.mm@gmail.com,
	axelrasmussen@google.com, baolin.wang@linux.alibaba.com,
	baoquan.he@linux.dev, chenridong@xiaomi.com, david@kernel.org,
	hannes@cmpxchg.org, kasong@tencent.com,
	linux-kernel@vger.kernel.org, linux-mm@kvack.org, ljs@kernel.org,
	lyugaofei@xiaomi.com, mhocko@kernel.org, qi.zheng@linux.dev,
	shakeel.butt@linux.dev, stevensd@chromium.org,
	wangzicheng@honor.com, weixugc@google.com, yuanchu@google.com,
	zhangbo56@xiaomi.com, Xueyuan Chen <xueyuan.chen21@gmail.com>
Subject: Re: [PATCH v2 2/7] mm/mglru: batch update lrugen->nr_pages in inc_min_seq()
Date: Sun, 30 Aug 2026 11:58:38 +0800	[thread overview]
Message-ID: <20260830035843.712320-1-kunwu.chan@linux.dev> (raw)
In-Reply-To: <20260827234704.63163-3-baohua@kernel.org>

On Fri, 28 Aug 2026 07:46:59 +0800 "Barry Song (Xiaomi)" <baohua@kernel.org> wrote:

> Currently, folio_inc_gen() updates lrugen->nr_pages for every folio
> as it advances generations. Instead, accumulate the size changes
> and update lrugen->nr_pages in a batch after scanning the entire
> oldest generation, or when the scan stops because remaining reaches
> zero.
> 
> Since we only move folios from the oldest generation to the second
> oldest generation, the active/inactive state cannot change. We can
> therefore skip __lru_update_size().
> 
> Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
> Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
> ---
>  mm/vmscan.c | 23 ++++++++++++++++++-----
>  1 file changed, 18 insertions(+), 5 deletions(-)
> 
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 8f187d296b8e..07c22d51debd 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -3918,6 +3918,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
>  	struct lru_gen_folio *lrugen = &lruvec->lrugen;
>  	int hist = lru_hist_from_seq(lrugen->min_seq[type]);
>  	int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
> +	int target_gen = (old_gen + 1) % MAX_NR_GENS;
>  
>  	/* For file type, skip the check if swappiness is anon only */
>  	if (type && (swappiness == SWAPPINESS_ANON_ONLY))
> @@ -3927,35 +3928,47 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
>  	if (!type && !swappiness)
>  		goto done;
>  
> +	VM_WARN_ON_ONCE(get_nr_gens(lruvec, type) != MAX_NR_GENS);
> +	VM_WARN_ON_ONCE(lru_gen_is_active(lruvec, old_gen) !=
> +			lru_gen_is_active(lruvec, target_gen));
>  	/* prevent cold/hot inversion if the type is evictable */
>  	for (zone = 0; zone < MAX_NR_ZONES; zone++) {
>  		struct list_head *head = &lrugen->folios[old_gen][type][zone];
> +		long delta = 0;
>  
>  		while (!list_empty(head)) {
>  			struct folio *folio = lru_to_folio(head);
> +			long nr_pages = folio_nr_pages(folio);
>  			int refs = folio_lru_refs(folio);
>  			bool workingset = folio_test_workingset(folio);
> +			bool gen_increased;
>  
>  			VM_WARN_ON_ONCE_FOLIO(folio_test_unevictable(folio), folio);
>  			VM_WARN_ON_ONCE_FOLIO(folio_test_active(folio), folio);
>  			VM_WARN_ON_ONCE_FOLIO(folio_is_file_lru(folio) != type, folio);
>  			VM_WARN_ON_ONCE_FOLIO(folio_zonenum(folio) != zone, folio);
>  
> -			new_gen = folio_inc_gen(lruvec, folio);
> +			new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
>  			list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
> -
> +			if (gen_increased)
> +				delta += nr_pages;
>  			/* don't count the workingset being lazily promoted */
>  			if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
>  				int tier = lru_tier_from_refs(refs, workingset);
> -				int delta = folio_nr_pages(folio);
>  
>  				WRITE_ONCE(lrugen->protected[hist][type][tier],
> -					   lrugen->protected[hist][type][tier] + delta);
> +					   lrugen->protected[hist][type][tier] + nr_pages);
>  			}
>  
>  			if (!--remaining)
> -				return false;
> +				break;
>  		}
> +		WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone],
> +			   lrugen->nr_pages[old_gen][type][zone] - delta);
> +		WRITE_ONCE(lrugen->nr_pages[target_gen][type][zone],
> +			   lrugen->nr_pages[target_gen][type][zone] + delta);

Hi Barry,

One subtle point about the `remaining` handling:
when `remaining` reaches zero, we now `break` rather than return
so that the accumulated `delta` is applied before returning.

As I understand it, this is required because `__folio_inc_gen()`
has already changed the generation of the scanned folios,
while `lrugen->nr_pages[]` is now updated only in batch.

So the invariant is that every successful generation increment must
have its corresponding `delta` flushed before `inc_min_seq()` returns.

Is this the intended accounting invariant?

Thanks,
KunWu

> +		if (!remaining)
> +			return false;
>  	}
>  done:
>  	reset_ctrl_pos(lruvec, type, true);
> -- 
> 2.34.1
> 
> 



  reply	other threads:[~2026-08-30  3:59 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
2026-08-27 23:46 ` [PATCH v2 1/7] mm/mglru: separate folio generation update from LRU accounting Barry Song (Xiaomi)
2026-08-30  7:26   ` Lian Wang
2026-08-27 23:46 ` [PATCH v2 2/7] mm/mglru: batch update lrugen->nr_pages in inc_min_seq() Barry Song (Xiaomi)
2026-08-30  3:58   ` Kunwu Chan [this message]
2026-08-30  4:26     ` Barry Song
2026-08-30  4:47       ` KunWu Chan
2026-08-30  7:01   ` Lian Wang
2026-08-27 23:47 ` [PATCH v2 3/7] mm/mglru: enhance cold/hot inversion handling " Barry Song (Xiaomi)
2026-08-30  2:43   ` Ridong Chen
2026-08-30  4:28     ` Barry Song
2026-08-30  7:02   ` Lian Wang
2026-08-27 23:47 ` [PATCH v2 4/7] mm/mglru: exclude folios promoted by aging from protected " Barry Song (Xiaomi)
2026-08-30  6:13   ` Ridong Chen
2026-08-30  7:03   ` Lian Wang
2026-08-27 23:47 ` [PATCH v2 5/7] mm/mglru: make LRU folio prefetch helper an inline function Barry Song (Xiaomi)
2026-08-30  7:43   ` Lian Wang
2026-08-27 23:47 ` [PATCH v2 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail Barry Song (Xiaomi)
2026-08-30  7:44   ` Lian Wang
2026-08-27 23:47 ` [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU Barry Song (Xiaomi)
2026-08-28  3:29   ` Barry Song
2026-08-30  7:45   ` Lian Wang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260830035843.712320-1-kunwu.chan@linux.dev \
    --to=kunwu.chan@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=chenridong@xiaomi.com \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=kunwu.chan@linux.dev \
    --cc=lianux.mm@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=lyugaofei@xiaomi.com \
    --cc=mhocko@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=stevensd@chromium.org \
    --cc=wangzicheng@honor.com \
    --cc=weixugc@google.com \
    --cc=xueyuan.chen21@gmail.com \
    --cc=yuanchu@google.com \
    --cc=zhangbo56@xiaomi.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.