From: SJ Park <sj@kernel.org>
To: Anton Gavriliuk <antosha20xx@gmail.com>
Cc: SJ Park <sj@kernel.org>, damon@lists.linux.dev
Subject: Re: Memory tiering with DAMON/DAMOS auto-tuning
Date: Fri, 2 Oct 2026 01:34:43 -0700 [thread overview]
Message-ID: <20261002083444.40553-1-sj@kernel.org> (raw)
In-Reply-To: <CAAiJnjpweWDACdt3HKtup=8Byp00ePakcHJWEY=WdGT7MT6yxA@mail.gmail.com>
On Thu, 1 Oct 2026 19:54:09 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
> > That said, I'm personally curious if such fast migration is really needed in
> > the real world workload. I assume the real production workload would have
> > stable access pattern but only occasionally get access pattern changes.
> > Sometimes temporal and rapid access pattern change could also be made. But for
> > such temporal pattern change, doing migration would only be costy, since the
> > data that suddenly hot could be soon be cold. And reliability is important in
> > production. Hence I was thinking slowly making the balance is better than too
> > quickly and reactively making migrations that will turn out to be no really
> > needed.
>
> > That said, if you want to test faster migration speed, the multiple kdamonds
> > usage could be one way to test. If it turns out it is really needed and
> > splitting address ranges has problems, we can consider adding DAMOS feature for
> > utilizing multiple CPUs for faster migrations.
>
>
> I agree with you!, but now it looks that demotion is too slow.
> The server is completely idle, only I play with memory tiering.
>
> As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices,
>
> [root@localhost anton]# ndctl list
> [
> {
> "dev":"namespace1.0",
> "mode":"devdax",
> "map":"dev",
> "size":3183575302144,
> "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f",
> "chardev":"dax1.0",
> "align":2097152
> },
> {
> "dev":"namespace0.0",
> "mode":"devdax",
> "map":"dev",
> "size":3183575302144,
> "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2",
> "chardev":"dax0.0",
> "align":2097152
> }
> ]
> [root@localhost anton]#
>
> and then configured as system ram by next command,
>
> daxctl reconfigure-device --mode=system-ram all
>
> The problem is during demotion (to keep 25% free in the fast tier),
> kdamond.0 utilizes single CPU core ~100%, but writing to the slower
> tier (System PMM Write) ~1.4 GB/s what is much slower than defined
> limit 8 GB/s.
The defined limit is only upper limit, so real speed could be slower than that.
FWIW, the example memory tiering script [1] uses 200 MiB/second as the
upperlimit. It was set by my gut feeling, not by some good data, though.
I personally feel like ~1.4 GB/s is still a good speed for long-running
production workloads, though I don't have a data to support that. Do you have
some data or reason to pursue 8 GB/s ?
>
>
> top - 16:26:06 up 12 min, 2 users, load average: 0.81, 0.75, 0.54
> Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie
> %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
> MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used, 465.2 buff/cache
> MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6424006.+avail Mem
>
> PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
> 3227 root 20 0 0 0 0 R 99.8 0.0 1:28.92 kdamond.0
> 2771 root 20 0 6932852 10260 5608 S 0.7 0.0 0:02.95
> pcm-memory
> 1593 root 20 0 17088 8352 7140 S 0.3 0.0 0:00.99
> systemd-logind
> 2263 root 20 0 300.4g 294.4g 6108 S 0.3 4.4 4:42.72
> valkey-server
> 3327 root 20 0 10916 6064 3764 R 0.3 0.0 0:00.16 top
> 1 root 20 0 25356 15232 10456 S 0.0 0.0 0:02.13 systemd
> 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd
>
>
>
> |---------------------------------------||---------------------------------------|
> |-- System DRAM Read Throughput(MB/s): 751.95
> --|
> |-- System DRAM Write Throughput(MB/s): 742.58
> --|
> |-- System PMM Read Throughput(MB/s): 693.12
> --|
> |-- System PMM Write Throughput(MB/s): 1387.47
> --|
> |-- System Read Throughput(MB/s): 1445.08
> --|
> |-- System Write Throughput(MB/s): 2130.05
> --|
> |-- System Memory Throughput(MB/s): 3575.12
> --|
> |---------------------------------------||---------------------------------------|
>
>
>
> Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 51312892466 lost: 0/0 drop
> Overhead Shared Object Symbol
> 47.50% [kernel] [k] migrate_folio_unmap
> 16.11% [kernel] [k] clear_highpages_kasan_tagged
> 14.93% [kernel] [k] copy_mc_fragile
> 1.63% [kernel] [k] smp_call_function_many_cond
> 1.44% [kernel] [k] page_vma_mapped_walk
> 0.97% [kernel] [k] folio_migrate_flags
> 0.87% [kernel] [k] try_to_migrate_one
> 0.66% [kernel] [k] rmqueue_bulk
> 0.61% [kernel] [k]
> __list_del_entry_valid_or_report
> 0.59% [kernel] [k] remove_migration_pte
> 0.50% [kernel] [k] rmap_walk_anon
> 0.45% [kernel] [k] __free_one_page
> 0.44% [kernel] [k] migrate_folio_move
> 0.44% [kernel] [k] folio_add_anon_rmap_ptes
> 0.40% [kernel] [k] lru_gen_add_folio
> 0.40% [kernel] [k] __mod_memcg_lruvec_state
> 0.39% [kernel] [k] mod_node_page_state
> 0.35% [kernel] [k] migrate_pages_batch
> 0.33% [kernel] [k] _raw_spin_lock
> 0.33% [kernel] [k] up_read
>
>
>
> Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 64449579226
> migrate_folio_unmap /proc/kcore [Percent: local period]
> Percent │ movq %r13,0x20(%rsp)
> 0.01 │ movq %rdx,%r13
> │ movq %r14,0x28(%rsp)
> │ movl %r9d,%r14d
> 0.00 │ movq %r15,0x30(%rsp)
> 0.01 │ movq %r8,%r15
> │ → callq *%rax
> 0.16 │ testq %rax,%rax
> │ ↓ je 2e3
> 0.01 │ movq %rbp,0x10(%rsp)
> 0.01 │ movq %rax,%rbp
> 0.04 │ movq %rax,(%r15)
> 0.00 │ movq $0x0,0x28(%rax)
> 99.27 │ lock
> │ btsq $0x0,(%rbx)
> 0.01 │ ↓ jb 26b
> 0.11 │ 67: movq (%rbx),%rdx
> 0.01 │ testq $0x2,(%rbx)
> 0.01 │ ↓ je e2
> │ cmpl $0x2,%r14d
> │ ↓ je d2
> │ movq 0x40(%rsp),%r8
> │ movq %rbx,%rdi
> │ movl $0x1,%ecx
> │ xorl %edx,%edx
>
> I also tried without "--damos_filter reject young" and "--damos_filter
> allow young", but got the same single CPU ~100% utilization and
> demotion poor performance.
Thank you for doing these great profiling and sharing the results, Anton.
Apparently the bottleneck is not in DAMON specific code including the filtering
part. Instead, the bottleneck is in the migration code, which is out of DAMON,
according to my understanding of the profiling result. You could also confirm
this by measuring the migration speed using move_pages() like system call,
which also use the migration code.
You may need to optimize the migration code, or make DAMON uses multiple CPUs
if faster speed is really what needed. I'd suggest using the address range
split approach that I suggested in the previous reply. If it becomes clear the
faster migration is really needed, we could start thinking about making DAMOS
use multiple CPUs without the address range split.
[1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
Thanks,
SJ
[...]
next prev parent reply other threads:[~2026-10-02 8:34 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk
2026-09-28 8:15 ` SJ Park
2026-09-28 17:02 ` Anton Gavriliuk
2026-09-29 7:49 ` SJ Park
2026-09-29 8:41 ` Anton Gavriliuk
2026-09-29 9:25 ` SJ Park
2026-09-29 13:54 ` Anton Gavriliuk
2026-09-29 17:37 ` SJ Park
2026-09-30 4:03 ` Anton Gavriliuk
2026-09-30 8:35 ` SJ Park
2026-09-30 16:26 ` Anton Gavriliuk
2026-09-30 17:43 ` SJ Park
2026-10-01 9:58 ` Anton Gavriliuk
2026-10-01 10:26 ` SJ Park
2026-10-01 12:08 ` Anton Gavriliuk
2026-10-01 13:27 ` SJ Park
2026-10-01 16:54 ` Anton Gavriliuk
2026-10-02 8:34 ` SJ Park [this message]
2026-10-02 10:02 ` Anton Gavriliuk
2026-10-02 11:02 ` SJ Park
2026-10-02 13:06 ` Anton Gavriliuk
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002083444.40553-1-sj@kernel.org \
--to=sj@kernel.org \
--cc=antosha20xx@gmail.com \
--cc=damon@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox