From: SJ Park <sj@kernel.org>
To: Anton Gavriliuk <antosha20xx@gmail.com>
Cc: SJ Park <sj@kernel.org>, damon@lists.linux.dev
Subject: Re: Memory tiering with DAMON/DAMOS auto-tuning
Date: Fri, 2 Oct 2026 04:02:15 -0700 [thread overview]
Message-ID: <20261002110216.41169-1-sj@kernel.org> (raw)
In-Reply-To: <CAAiJnjpBdRQnmsm8zB-Bh7taeBkOPadhQWOk6-a2dfTxYVMiZw@mail.gmail.com>
On Fri, 2 Oct 2026 13:02:13 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
> Peak sequential write to slower tier ~12 GB/s.
> I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded.
>
> For Memory Tiering with OLTP-like workloads needs to be
> promoted/demoted, 1.4 GB/s is OK.
> But for bandwidth intensive workloads such as analytics, AI, VDI,...
> 1.4 GB/s is too slow.
I think the factor deciding how fast migration should be is not bandwidth
intensiveness but how quickly access pattern changes. That is, migration speed
should be fast enough to get costs from migration itself be paid back, and get
additional benefits. That depends on the speed of the workloads' access
pattern change, and the amount of data of the changed pattern. For example,
let's suppose 10 GiB data of a workload becomes hot, but will be cold again
after 10 seconds. If we can migrate it to upper tier in one second, it will
get benefit from continued access for remaining 9 seconds. If it takes 10
seconds to migrate, the benefit from the migration will be much lower. It
might even lower than the migration work cost.
>
> On the server where I test Memory Tiering, Intel CPU 8280L installed.
> Despite being already 7 years old, any single CPU core has 13-14 GB/s
> bandwidth access to local memory.
That maese sense to me. Kernel level page migration requires not only the
content writes. It also need to do additional works. It should allocate pages
in destination node that the content will be copied to. It should update
mappings and related metadata. It should also handle possible races. Hence
page migration is much more expensive and slow than pure I/O.
>
> So from my point of view - even if we increase the number of cores
> keeping in mind 1.4 GB/s per core, we need so many cores to perform
> 10's GB/s demote/promote.
I agree.
> Firstly we need to improve demote/promote
> bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s.
I agree there could be workloads that could get benefit from faster migration.
And IIRC, there were a few people working on making migration faster and
lightweight. I'm not an expert in the domain and not involved to such works
for now, though. Nonetheless, having concrete data showing why 9-10 GB/s is
the right speed would be nice.
>
> > You could also confirm
> > this by measuring the migration speed using move_pages() like system call,
> > which also use the migration code.
>
> I'm not a developer and I need to spend more time thinking about how
> to do that, but I tried,
>
> I used migratepages command, which is probably use migrate_pages()
> instead of move_pages(), but I got the same poor bandwidth and stack
> as with DAMON/DAMOS Memory Tiering,
>
> top - 12:43:59 up 13 min, 2 users, load average: 0.65, 0.88, 0.76
> Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie
> %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
> MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used, 669.8 buff/cache
> MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6423508.+avail Mem
>
> PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
> 2935 root 20 0 2644 1764 1644 R 99.7 0.0 0:42.19
> migratepages
> 2941 root 20 0 10884 6192 3916 R 0.7 0.0 0:00.04 top
> 1238 systemd+ 20 0 15860 6996 5952 S 0.3 0.0 0:00.37
> systemd-oomd
> 1928 root 20 0 666268 24448 23440 S 0.3 0.0 0:00.34 rsyslogd
> 1 root 20 0 25840 15532 10492 S 0.0 0.0 0:02.15 systemd
> 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd
> 3 root 20 0 0 0 0 S 0.0 0.0 0:00.00
> pool_workqueue_release
> 4 root 0 -20 0 0 0 I 0.0 0.0 0:00.00
> kworker/R-rcu_gp
>
>
> |---------------------------------------||---------------------------------------|
> |-- System DRAM Read Throughput(MB/s): 759.41
> --|
> |-- System DRAM Write Throughput(MB/s): 751.36
> --|
> |-- System PMM Read Throughput(MB/s): 701.24
> --|
> |-- System PMM Write Throughput(MB/s): 1402.15
> --|
> |-- System Read Throughput(MB/s): 1460.65
> --|
> |-- System Write Throughput(MB/s): 2153.50
> --|
> |-- System Memory Throughput(MB/s): 3614.15
> --|
> |---------------------------------------||---------------------------------------|
>
>
> Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 48280162400 lost: 0/0 drop:
> Overhead Shared Object Symbol
> 49.82% [kernel] [k] migrate_folio_unmap
> 16.52% [kernel] [k] clear_highpages_kasan_tagged
> 13.56% [kernel] [k] copy_mc_fragile
> 1.80% [kernel] [k] smp_call_function_many_cond
> 1.06% [kernel] [k] page_vma_mapped_walk
> 0.95% [kernel] [k] try_to_migrate_one
> 0.94% [kernel] [k] folio_migrate_flags
> 0.71% [kernel] [k] rmqueue_bulk
> 0.61% [kernel] [k] remove_migration_pte
> 0.54% [kernel] [k] __free_one_page
> 0.50% [kernel] [k] migrate_folio_move
> 0.48% [kernel] [k] folio_batch_move_lru
> 0.47% [kernel] [k]
> __list_del_entry_valid_or_report
>
>
> Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 33608407521
> migrate_folio_unmap /proc/kcore [Percent: local period]
> Percent │ movq %r13,0x20(%rsp)
> 0.02 │ movq %rdx,%r13
> │ movq %r14,0x28(%rsp)
> │ movl %r9d,%r14d
> │ movq %r15,0x30(%rsp)
> 0.01 │ movq %r8,%r15
> │ → callq *%rax
> 0.04 │ testq %rax,%rax
> │ ↓ je 2e3
> 0.01 │ movq %rbp,0x10(%rsp)
> │ movq %rax,%rbp
> 0.03 │ movq %rax,(%r15)
> │ movq $0x0,0x28(%rax)
> 99.52 │ lock
> │ btsq $0x0,(%rbx)
> 0.01 │ ↓ jb 26b
> 0.13 │ 67: movq (%rbx),%rdx
> 0.01 │ testq $0x2,(%rbx)
> 0.01 │ ↓ je e2
> │ cmpl $0x2,%r14d
> │ ↓ je d2
> │ movq 0x40(%rsp),%r8
> │ movq %rbx,%rdi
> │ movl $0x1,%ecx
> │ xorl %edx,%edx
>
> Lastly I used migratepages between two fast tiers (numa nodes 0 & 1)
> and got 1.5 GB/s.
> All this is too slow for the coming CXL era.
Thank you for testing this and sharing the results. This confirms my previous
theory is not wrong.
>
> Anyway, I'm ready to proceed in assisting you in checking/testing if
> you are also interested.
I think my thoery (bottleneck is in migration, not DAMON) is already proven.
We also confirmed DAMON and DAMOS are working as expected so far. So I find no
more thing to test for now. If you have more ideas to test, I will be more
than happy to help.
Thanks,
SJ
[...]
next prev parent reply other threads:[~2026-10-02 11:02 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk
2026-09-28 8:15 ` SJ Park
2026-09-28 17:02 ` Anton Gavriliuk
2026-09-29 7:49 ` SJ Park
2026-09-29 8:41 ` Anton Gavriliuk
2026-09-29 9:25 ` SJ Park
2026-09-29 13:54 ` Anton Gavriliuk
2026-09-29 17:37 ` SJ Park
2026-09-30 4:03 ` Anton Gavriliuk
2026-09-30 8:35 ` SJ Park
2026-09-30 16:26 ` Anton Gavriliuk
2026-09-30 17:43 ` SJ Park
2026-10-01 9:58 ` Anton Gavriliuk
2026-10-01 10:26 ` SJ Park
2026-10-01 12:08 ` Anton Gavriliuk
2026-10-01 13:27 ` SJ Park
2026-10-01 16:54 ` Anton Gavriliuk
2026-10-02 8:34 ` SJ Park
2026-10-02 10:02 ` Anton Gavriliuk
2026-10-02 11:02 ` SJ Park [this message]
2026-10-02 13:06 ` Anton Gavriliuk
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002110216.41169-1-sj@kernel.org \
--to=sj@kernel.org \
--cc=antosha20xx@gmail.com \
--cc=damon@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox