From: Ridong Chen <ridong.chen@linux.dev>
To: Andrew Morton <akpm@linux-foundation.org>,
Johannes Weiner <hannes@cmpxchg.org>
Cc: Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Baoquan He <baoquan.he@linux.dev>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"open list:MEMORY MANAGEMENT - MGLRU (MULTI-GEN LRU)"
<linux-mm@kvack.org>,
linux-kernel@vger.kernel.org, Ridong Chen <chenridong@xiaomi.com>,
xieym_ict@hotmail.com
Subject: Re: [PATCH v8] mm: vmscan: retry folios written back while isolated for traditional LRU
Date: Sun, 13 Sep 2026 18:16:40 +0800 [thread overview]
Message-ID: <46037a37-4cf6-448e-a94b-30a4d16e8814@linux.dev> (raw)
In-Reply-To: <20260913100013.3603815-1-ridong.chen@linux.dev>
On 9/13/2026 6:00 PM, Ridong Chen wrote:
> From: Ridong Chen <chenridong@xiaomi.com>
>
> As commit 359a5e1416ca ("mm: multi-gen LRU: retry folios written back
> while isolated") mentioned:
>
> The page reclaim isolates a batch of folios from the tail of one of the
> LRU lists and works on those folios one by one. For a suitable
> swap-backed folio, if the swap device is async, it queues that folio for
> writeback. After the page reclaim finishes an entire batch, it puts back
> the folios it queued for writeback to the head of the original LRU list.
>
> In the meantime, the page writeback flushes the queued folios also by
> batches. Its batching logic is independent from that of the page
> reclaim. For each of the folios it writes back, the page writeback calls
> folio_rotate_reclaimable() which tries to rotate a folio to the tail.
>
> folio_rotate_reclaimable() only works for a folio after the page reclaim
> has put it back. If an async swap device is fast enough, the page
> writeback can finish with that folio while the page reclaim is still
> working on the rest of the batch containing it. In this case, that folio
> will remain at the head and the page reclaim will not retry it before
> reaching there".
>
> The commit 359a5e1416ca ("mm: multi-gen LRU: retry folios written back
> while isolated") only fixed the issue for mglru. However, this issue
> also exists in the traditional active/inactive LRU and was found at [1].
>
> It can be reproduced with below steps:
>
> 1. Compile with CONFIG_TRANSPARENT_HUGEPAGE=y
> 2. Mount memcg v1, and create memcg named test_memcg and set
> limit_in_bytes=1G, memsw.limit_in_bytes=2G.
> 3. Create a 1G swap file, and allocate 1.35G anon memory in test_memcg.
>
Hi all,
I am raising this issue again. It has been a long time since the last version [1].
It was suspected that Kirill's "[PATCH 0/8] mm: Remove PG_reclaim" would solve
this issue, but the issue remains.
I am providing the reproducer(offered by Xuedong Zhao) in the hope that it will
help fix this issue.
memcg_malloc.c:
```
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#define ONE_GB (1024 * 1024 * 1024)
#define SIXTY_FOUR_MB (64 * 1024 * 1024)
/* non-zero fill: zero pages can be deduped/never written to swap, which hides
* the "written-back-while-isolated" leak. Use a real byte pattern. */
#define FILL_BYTE 0xAB
void allocate_memory(size_t size_in_bytes) {
size_t total_allocated = 0;
char *memory;
while (total_allocated + ONE_GB <= size_in_bytes) {
memory = (char *)malloc(ONE_GB);
if (memory == NULL) {
perror("malloc");
exit(EXIT_FAILURE);
}
memset(memory, FILL_BYTE, ONE_GB);
total_allocated += ONE_GB;
printf("Allocated %zu GB\n", total_allocated / ONE_GB);
sleep(1);
}
while (total_allocated + SIXTY_FOUR_MB <= size_in_bytes) {
memory = (char *)malloc(SIXTY_FOUR_MB);
if (memory == NULL) {
perror("malloc");
exit(EXIT_FAILURE);
}
memset(memory, FILL_BYTE, SIXTY_FOUR_MB);
total_allocated += SIXTY_FOUR_MB;
printf("Allocated %zu MB\n", total_allocated / (1024 * 1024));
sleep(1);
}
size_t remaining = size_in_bytes - total_allocated;
if (remaining > 0) {
memory = (char *)malloc(remaining);
if (memory == NULL) {
perror("malloc");
exit(EXIT_FAILURE);
}
memset(memory, FILL_BYTE, remaining);
total_allocated += remaining;
printf("Allocated remaining %zu bytes\n", remaining);
sleep(1);
}
printf("Total allocated: %zu bytes\n", total_allocated);
}
int main(int argc, char *argv[]) {
if (argc != 2) {
fprintf(stderr, "Usage: %s <size_in_gb>\n", argv[0]);
return EXIT_FAILURE;
}
double size_in_gb = atof(argv[1]);
if (size_in_gb <= 0) {
fprintf(stderr, "Invalid size: %s\n", argv[1]);
return EXIT_FAILURE;
}
size_t size_in_bytes = (size_t)(size_in_gb * ONE_GB);
allocate_memory(size_in_bytes);
sleep(3600);
return EXIT_SUCCESS;
}
```
test.sh:
```
#!/bin/bash
set -e
# Variables
MEMCG_NAME="test_memcg"
MEM_LIMIT="1G"
MEMSW_LIMIT="2G"
PROGRAM_PATH="./memcg_malloc"
PROGRAM_ARGS="1.35"
# ---- swapfile setup --------------------------------------------------------
SWAPFILE="/swapfile"
SWAP_SIZE="1G" # size of the swap device backing the test
setup_swap() {
# already have swap on? then nothing to do
if [ "$(swapon --show --noheadings | wc -l)" -gt 0 ]; then
echo "swap already active:"; swapon --show
return
fi
if [ ! -f "$SWAPFILE" ]; then
echo "Creating ${SWAP_SIZE} swapfile at ${SWAPFILE}"
# fallocate is fast; fall back to dd if the fs doesn't support it
fallocate -l "$SWAP_SIZE" "$SWAPFILE" 2>/dev/null || \
dd if=/dev/zero of="$SWAPFILE" bs=1M count=$((8*1024)) status=progress
chmod 600 "$SWAPFILE"
mkswap "$SWAPFILE"
fi
swapon "$SWAPFILE"
echo "swap enabled:"; swapon --show
}
setup_swap
# ---- reclaim preconditions -------------------------------------------------
# This bug is in the *traditional* active/inactive LRU; MGLRU already fixed it
# in commit 359a5e1416ca, so it must be disabled to reproduce.
[ -f /sys/kernel/mm/lru_gen/enabled ] && echo 0 > /sys/kernel/mm/lru_gen/enabled
# Large anon folios make the reclaim/writeback batching race easy to hit.
echo always > /sys/kernel/mm/transparent_hugepage/enabled
echo "lru_gen: $(cat /sys/kernel/mm/lru_gen/enabled 2>/dev/null) thp: $(cat
/sys/kernel/mm/transparent_hugepage/enabled)"
# Create the cgroup slice if it doesn't exist
if ! systemctl list-units --full -all | grep -q "${MEMCG_NAME}.slice"; then
echo "Creating cgroup slice ${MEMCG_NAME}.slice"
systemctl set-property --runtime -- ${MEMCG_NAME}.slice MemoryMax=${MEM_LIMIT}
systemctl set-property --runtime -- ${MEMCG_NAME}.slice
MemorySwapMax=${MEMSW_LIMIT}
fi
# Start the slice to apply the properties
systemctl start ${MEMCG_NAME}.slice
# Run the program in the cgroup
echo "Running programs in cgroup slice ${MEMCG_NAME}.slice"
systemd-run --unit=${MEMCG_NAME}_proc1 --slice=${MEMCG_NAME}.slice
${PROGRAM_PATH} ${PROGRAM_ARGS} &
# Pause to let reclaim settle under the 1G limit
sleep 60
# ---- measurement (cgroup v1) -----------------------------------------------
echo "########## measurement ##########"
CG=/sys/fs/cgroup/memory/${MEMCG_NAME}.slice/${MEMCG_NAME}_proc1.service
usage=$(cat "$CG/memory.usage_in_bytes" 2>/dev/null || echo 0)
memsw=$(cat "$CG/memory.memsw.usage_in_bytes" 2>/dev/null || echo 0)
# v1: swap charged to the memcg is memsw.usage - usage
swap_charged=$((memsw - usage))
# swap actually consumed on the device: /proc/swaps "Used" column is in KiB
dev_used_kb=$(awk 'NR>1 {sum += $4} END {print sum+0}' /proc/swaps)
dev_used=$((dev_used_kb * 1024))
# the bug wastes swap: slots written back while isolated are charged on the
# device but never reused, so device usage outruns what the cgroup accounts for
waste=$((dev_used - swap_charged))
to_mib() { awk -v b="$1" 'BEGIN { printf "%.0f MiB", b/1024/1024 }'; }
echo "memory.usage_in_bytes : ${usage} bytes ($(to_mib ${usage}))"
echo "memory.memsw.usage_in_bytes : ${memsw} bytes ($(to_mib ${memsw}))"
echo "swap charged (memsw - usage) : ${swap_charged} bytes ($(to_mib
${swap_charged}))"
echo "swap device used : ${dev_used} bytes ($(to_mib ${dev_used}))"
echo "wasted swap (device - cg) : ${waste} bytes ($(to_mib ${waste}))"
echo "-----"
free -h
echo "--- /proc/swaps ---"; cat /proc/swaps
# Wait for the processes to complete
wait
# Clean up
echo "Cleaning up"
# Reset failed state if the slice is still loaded
if systemctl list-units --full -all | grep -q "${MEMCG_NAME}.slice"; then
systemctl reset-failed ${MEMCG_NAME}.slice
fi
systemctl stop ${MEMCG_NAME}.slice
echo "Done"
```
Result shown as:
```
...
########## measurement ##########
memory.usage_in_bytes : 1070014464 bytes (1020 MiB)
memory.memsw.usage_in_bytes : 1413173248 bytes (1348 MiB)
swap charged (memsw - usage) : 343158784 bytes (327 MiB)
swap device used : 344248320 bytes (328 MiB)
wasted swap (device - cg) : 1089536 bytes (1 MiB)
-----
total used free shared buff/cache available
Mem: 1.6Gi 1.2Gi 316Mi 0.0Ki 85Mi 287Mi
Swap: 1.0Gi 328Mi 695Mi
...
```
[1] https://lore.kernel.org/linux-mm/20250113155206.GB829144@cmpxchg.org/#r
--
Best regards
Ridong
next prev parent reply other threads:[~2026-09-13 10:16 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-13 10:00 [PATCH v8] mm: vmscan: retry folios written back while isolated for traditional LRU Ridong Chen
2026-09-13 10:16 ` Ridong Chen [this message]
2026-09-13 11:33 ` Barry Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=46037a37-4cf6-448e-a94b-30a4d16e8814@linux.dev \
--to=ridong.chen@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=chenridong@xiaomi.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=weixugc@google.com \
--cc=xieym_ict@hotmail.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.