* Memory tiering with DAMON/DAMOS auto-tuning
@ 2026-09-27 17:26 Anton Gavriliuk
2026-09-28 8:15 ` SJ Park
0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-27 17:26 UTC (permalink / raw)
To: sj, damon
Hello
I would like to play with memory tiering with DAMON/DAMOS auto-tuning
(between numa nodes 0 & 2) based on valkey and memtier_benchmark.
The goal - keep numa node 0 50% free and promote cold pages from numa
node 2 to numa node 0 immediately when they become hot.
This is Fedora Server 44 up-to-date with the 7.2.8 kernel.
[root@localhost ~]# numactl -H
available: 4 nodes (0-3)
node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
22 23 24 25 26 27
node 0 size: 386590 MB
node 0 free: 337165 MB
node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
47 48 49 50 51 52 53 54 55
node 1 size: 387055 MB
node 1 free: 338204 MB
node 2 cpus:
node 2 size: 3033088 MB
node 2 free: 3032998 MB
node 3 cpus:
node 3 size: 3033088 MB
node 3 free: 3032998 MB
node distances:
node 0 1 2 3
0: 10 20 25 35
1: 20 10 35 25
2: 25 35 10 20
3: 35 25 20 10
[root@localhost ~]#
[root@localhost ~]# /home/anton/Linux/mlc
Intel(R) Memory Latency Checker - v3.13
Measuring idle latencies for sequential access (in ns)...
Numa node
Numa node 0 1 2 3
0 82.3 147.6 173.8 238.0
1 146.7 82.0 238.7 172.2
The Goal:
1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2
2. Keep fast tier 50% free
3. Demote valkey inactive >=5 min pages from fast to slow tier
4. Promote valkey hot pages from slow tier to fast tier immediately
when accessed
5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU
performance unlimited
6. Monitoring intervals for Demote/Promote 500ms
What I already done -
Launch Valkey Pinned to Node 0 DRAM
numactl --cpunodebind=0 --preferred=0 valkey-server \
--port 6379 \
--protected-mode no \
--save "" \
--appendonly no \
--maxmemory 300gb \
--maxmemory-policy noeviction &
Command to Load ~200+ GB
memtier_benchmark \
-s 127.0.0.1 -p 6379 \
-t 16 -c 16 \
-d 10240 \
--ratio=1:0 \
--key-pattern=P:P \
--distinct-client-seed \
--key-maximum=20000000 \
--pipeline=32 \
-n allkeys
[root@localhost anton]# numastat -p $(pgrep valkey-server)
Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
Node 0 Node 1 Node 2
--------------- --------------- ---------------
Huge 0.00 0.00 0.00
Heap 0.11 0.00 0.00
Stack 0.03 0.00 0.00
Private 238356.61 0.46 4.82
---------------- --------------- --------------- ---------------
Total 238356.75 0.46 4.82
Node 3 Total
--------------- ---------------
Huge 0.00 0.00
Heap 0.00 0.11
Stack 0.00 0.03
Private 0.00 238361.90
---------------- --------------- ---------------
Total 0.00 238362.04
damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s
\--damos_action migrate_cold 2 --damos_access_rate 0% 0%
\--damos_apply_interval 1s \--damos_quota_interval 1s
--damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
\--damos_filter reject young \--numa_node 2
--monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot
0 --damos_access_rate 5% max \--damos_apply_interval 1s
\--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
node_mem_used_bp 99.7% 0 \--damos_filter allow young
\--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
--nr_schemes 1 1 --nr_ctxs 1 1
And it is demoted much more, almost all pages than the goal of 50%
free numa node 0.
[root@localhost anton]# numastat -p $(pgrep valkey-server)
Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
Node 0 Node 1 Node 2
--------------- --------------- ---------------
Huge 0.00 0.00 0.00
Heap 0.02 0.00 0.09
Stack 0.02 0.00 0.01
Private 24389.88 0.46 213971.55
---------------- --------------- --------------- ---------------
Total 24389.92 0.46 213971.65
Node 3 Total
--------------- ---------------
Huge 0.00 0.00
Heap 0.00 0.11
Stack 0.00 0.03
Private 0.00 238361.90
---------------- --------------- ---------------
Total 0.00 238362.04
I just want that demotion process to stop when there is 50% free at numa node 0.
Where am I wrong, how to fix that ?
Anton
^ permalink raw reply [flat|nested] 21+ messages in thread* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk @ 2026-09-28 8:15 ` SJ Park 2026-09-28 17:02 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-28 8:15 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon Hello Anton, On Sun, 27 Sep 2026 20:26:21 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > Hello > > I would like to play with memory tiering with DAMON/DAMOS auto-tuning > (between numa nodes 0 & 2) based on valkey and memtier_benchmark. Thank you for sharing your use case and question! > The goal - keep numa node 0 50% free and promote cold pages from numa > node 2 to numa node 0 immediately when they become hot. > This is Fedora Server 44 up-to-date with the 7.2.8 kernel. > > [root@localhost ~]# numactl -H > available: 4 nodes (0-3) > node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 > 22 23 24 25 26 27 > node 0 size: 386590 MB > node 0 free: 337165 MB > node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 > 47 48 49 50 51 52 53 54 55 > node 1 size: 387055 MB > node 1 free: 338204 MB > node 2 cpus: > node 2 size: 3033088 MB > node 2 free: 3032998 MB > node 3 cpus: > node 3 size: 3033088 MB > node 3 free: 3032998 MB > node distances: > node 0 1 2 3 > 0: 10 20 25 35 > 1: 20 10 35 25 > 2: 25 35 10 20 > 3: 35 25 20 10 > [root@localhost ~]# > [root@localhost ~]# /home/anton/Linux/mlc > Intel(R) Memory Latency Checker - v3.13 > Measuring idle latencies for sequential access (in ns)... > Numa node > Numa node 0 1 2 3 > 0 82.3 147.6 173.8 238.0 > 1 146.7 82.0 238.7 172.2 > > > The Goal: > > 1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2 > 2. Keep fast tier 50% free > 3. Demote valkey inactive >=5 min pages from fast to slow tier > 4. Promote valkey hot pages from slow tier to fast tier immediately > when accessed > 5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU > performance unlimited > 6. Monitoring intervals for Demote/Promote 500ms > > What I already done - > > Launch Valkey Pinned to Node 0 DRAM > > numactl --cpunodebind=0 --preferred=0 valkey-server \ > --port 6379 \ > --protected-mode no \ > --save "" \ > --appendonly no \ > --maxmemory 300gb \ > --maxmemory-policy noeviction & > > > Command to Load ~200+ GB > > memtier_benchmark \ > -s 127.0.0.1 -p 6379 \ > -t 16 -c 16 \ > -d 10240 \ > --ratio=1:0 \ > --key-pattern=P:P \ > --distinct-client-seed \ > --key-maximum=20000000 \ > --pipeline=32 \ > -n allkeys > > > [root@localhost anton]# numastat -p $(pgrep valkey-server) > > Per-node process memory usage (in MBs) for PID 5631 (valkey-server) > Node 0 Node 1 Node 2 > --------------- --------------- --------------- > Huge 0.00 0.00 0.00 > Heap 0.11 0.00 0.00 > Stack 0.03 0.00 0.00 > Private 238356.61 0.46 4.82 > ---------------- --------------- --------------- --------------- > Total 238356.75 0.46 4.82 > > Node 3 Total > --------------- --------------- > Huge 0.00 0.00 > Heap 0.00 0.11 > Stack 0.00 0.03 > Private 0.00 238361.90 > ---------------- --------------- --------------- > Total 0.00 238362.04 > > > damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s The example memory tiering script [1] uses 4% as intervals goal. Is there a reason to use 97% as the goal instead? > \--damos_action migrate_cold 2 --damos_access_rate 0% 0% > \--damos_apply_interval 1s \--damos_quota_interval 1s > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0 > \--damos_filter reject young \ You mentioned you want to demote >=5 minutes inactive pages. But, the above command doesn't have '--damos_age' option. It means DAMON will demote node 0 pages as soon as it finds it was not accessed, even if it was not accessed for <5 minutes. You could let DAMON know you want to demote only >=5 minutes inactive pages by adding '--damos_age 5m max' option. > --numa_node 2 > --monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot Again, I'm curious why you use 97% goal. > 0 --damos_access_rate 5% max \--damos_apply_interval 1s > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal > node_mem_used_bp 99.7% 0 \--damos_filter allow young Is the above node_mum_used_bp what you really want? That means you want to promote hot pages from node 2 to node 0, until the node 0 memory utilization becomes 99.7%. That overlaps with the demotion goal (50% free memory of node 0) quite a lot. I'd suggest smaller overlap, say, 50.3%, to keep healthy circulation of hot/cold pages while not consuming too much resource under stabilized access pattern. > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1 > --nr_schemes 1 1 --nr_ctxs 1 1 > > And it is demoted much more, almost all pages than the goal of 50% > free numa node 0. > > [root@localhost anton]# numastat -p $(pgrep valkey-server) > > Per-node process memory usage (in MBs) for PID 5631 (valkey-server) > Node 0 Node 1 Node 2 > --------------- --------------- --------------- > Huge 0.00 0.00 0.00 > Heap 0.02 0.00 0.09 > Stack 0.02 0.00 0.01 > Private 24389.88 0.46 213971.55 > ---------------- --------------- --------------- --------------- > Total 24389.92 0.46 213971.65 > > Node 3 Total > --------------- --------------- > Huge 0.00 0.00 > Heap 0.00 0.11 > Stack 0.00 0.03 > Private 0.00 238361.90 > ---------------- --------------- --------------- > Total 0.00 238362.04 > > I just want that demotion process to stop when there is 50% free at numa node 0. > > Where am I wrong, how to fix that ? DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and temporal. As the documentation [2] explains, 'consistent' tuner assumes it should keep applying the action in a level to keep the goal achieved. In this case, for example, the demotion scheme assumes there will be continued promotion and therefore it keeps demoting. If there is no appropriate promotion, it could result in demoting more than expected amount. To me, this seems like the workload has no much hot data, so demotion is much more stronger. For path forward, I'd suggest trying 'temporal' tuner [2]. It is designed to immediately stop after achieving the goal. If it still doesn't work, I'd suggest tracing damon:damos_esz tracepoint with the numastat output. It will show if the autotune is working as expected. For long term production use case, I'd like to suggest running the example tiering config [1] and see if it also gives you unexpected results. [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh [2] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#aim-oriented-feedback-driven-auto-tuning Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-28 8:15 ` SJ Park @ 2026-09-28 17:02 ` Anton Gavriliuk 2026-09-29 7:49 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-09-28 17:02 UTC (permalink / raw) To: SJ Park; +Cc: damon > Thank you for sharing your use case and question! Thank you for your quick response and your answer/comments. We already had memory tiering implemented at hardware level 5-7 years ago with Intel's Optane PEM when configured in Memory Mode. Now we do the same thing at Linux level :-) I'm new at DAMON/DAMOS, there are lots of new tunables to me. > The example memory tiering script [1] uses 4% as intervals goal. Is there a > reason to use 97% as the goal instead? Oohhh... I moved to 5%. > Is the above node_mum_used_bp what you really want? That means you want to > promote hot pages from node 2 to node 0, until the node 0 memory utilization > becomes 99.7%. That overlaps with the demotion goal (50% free memory of node > 0) quite a lot. I'd suggest smaller overlap, say, 50.3%, to keep healthy > circulation of hot/cold pages while not consuming too much resource under > stabilized access pattern. > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and > temporal. As the documentation [2] explains, 'consistent' tuner assumes it > should keep applying the action in a level to keep the goal achieved. In this > case, for example, the demotion scheme assumes there will be continued > promotion and therefore it keeps demoting. If there is no appropriate > promotion, it could result in demoting more than expected amount. So let me firstly understand what I really want to test :-) Anton пн, 28 сент. 2026 г. в 11:15, SJ Park <sj@kernel.org>: > > Hello Anton, > > On Sun, 27 Sep 2026 20:26:21 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > Hello > > > > I would like to play with memory tiering with DAMON/DAMOS auto-tuning > > (between numa nodes 0 & 2) based on valkey and memtier_benchmark. > > Thank you for sharing your use case and question! > > > The goal - keep numa node 0 50% free and promote cold pages from numa > > node 2 to numa node 0 immediately when they become hot. > > This is Fedora Server 44 up-to-date with the 7.2.8 kernel. > > > > [root@localhost ~]# numactl -H > > available: 4 nodes (0-3) > > node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 > > 22 23 24 25 26 27 > > node 0 size: 386590 MB > > node 0 free: 337165 MB > > node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 > > 47 48 49 50 51 52 53 54 55 > > node 1 size: 387055 MB > > node 1 free: 338204 MB > > node 2 cpus: > > node 2 size: 3033088 MB > > node 2 free: 3032998 MB > > node 3 cpus: > > node 3 size: 3033088 MB > > node 3 free: 3032998 MB > > node distances: > > node 0 1 2 3 > > 0: 10 20 25 35 > > 1: 20 10 35 25 > > 2: 25 35 10 20 > > 3: 35 25 20 10 > > [root@localhost ~]# > > [root@localhost ~]# /home/anton/Linux/mlc > > Intel(R) Memory Latency Checker - v3.13 > > Measuring idle latencies for sequential access (in ns)... > > Numa node > > Numa node 0 1 2 3 > > 0 82.3 147.6 173.8 238.0 > > 1 146.7 82.0 238.7 172.2 > > > > > > The Goal: > > > > 1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2 > > 2. Keep fast tier 50% free > > 3. Demote valkey inactive >=5 min pages from fast to slow tier > > 4. Promote valkey hot pages from slow tier to fast tier immediately > > when accessed > > 5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU > > performance unlimited > > 6. Monitoring intervals for Demote/Promote 500ms > > > > What I already done - > > > > Launch Valkey Pinned to Node 0 DRAM > > > > numactl --cpunodebind=0 --preferred=0 valkey-server \ > > --port 6379 \ > > --protected-mode no \ > > --save "" \ > > --appendonly no \ > > --maxmemory 300gb \ > > --maxmemory-policy noeviction & > > > > > > Command to Load ~200+ GB > > > > memtier_benchmark \ > > -s 127.0.0.1 -p 6379 \ > > -t 16 -c 16 \ > > -d 10240 \ > > --ratio=1:0 \ > > --key-pattern=P:P \ > > --distinct-client-seed \ > > --key-maximum=20000000 \ > > --pipeline=32 \ > > -n allkeys > > > > > > [root@localhost anton]# numastat -p $(pgrep valkey-server) > > > > Per-node process memory usage (in MBs) for PID 5631 (valkey-server) > > Node 0 Node 1 Node 2 > > --------------- --------------- --------------- > > Huge 0.00 0.00 0.00 > > Heap 0.11 0.00 0.00 > > Stack 0.03 0.00 0.00 > > Private 238356.61 0.46 4.82 > > ---------------- --------------- --------------- --------------- > > Total 238356.75 0.46 4.82 > > > > Node 3 Total > > --------------- --------------- > > Huge 0.00 0.00 > > Heap 0.00 0.11 > > Stack 0.00 0.03 > > Private 0.00 238361.90 > > ---------------- --------------- --------------- > > Total 0.00 238362.04 > > > > > > damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s > > The example memory tiering script [1] uses 4% as intervals goal. Is there a > reason to use 97% as the goal instead? > > > \--damos_action migrate_cold 2 --damos_access_rate 0% 0% > > \--damos_apply_interval 1s \--damos_quota_interval 1s > > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0 > > \--damos_filter reject young \ > > You mentioned you want to demote >=5 minutes inactive pages. But, the above > command doesn't have '--damos_age' option. It means DAMON will demote node 0 > pages as soon as it finds it was not accessed, even if it was not accessed for > <5 minutes. You could let DAMON know you want to demote only >=5 minutes > inactive pages by adding '--damos_age 5m max' option. > > > --numa_node 2 > > --monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot > > Again, I'm curious why you use 97% goal. > > > 0 --damos_access_rate 5% max \--damos_apply_interval 1s > > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal > > node_mem_used_bp 99.7% 0 \--damos_filter allow young > > Is the above node_mum_used_bp what you really want? That means you want to > promote hot pages from node 2 to node 0, until the node 0 memory utilization > becomes 99.7%. That overlaps with the demotion goal (50% free memory of node > 0) quite a lot. I'd suggest smaller overlap, say, 50.3%, to keep healthy > circulation of hot/cold pages while not consuming too much resource under > stabilized access pattern. > > > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1 > > --nr_schemes 1 1 --nr_ctxs 1 1 > > > > And it is demoted much more, almost all pages than the goal of 50% > > free numa node 0. > > > > [root@localhost anton]# numastat -p $(pgrep valkey-server) > > > > Per-node process memory usage (in MBs) for PID 5631 (valkey-server) > > Node 0 Node 1 Node 2 > > --------------- --------------- --------------- > > Huge 0.00 0.00 0.00 > > Heap 0.02 0.00 0.09 > > Stack 0.02 0.00 0.01 > > Private 24389.88 0.46 213971.55 > > ---------------- --------------- --------------- --------------- > > Total 24389.92 0.46 213971.65 > > > > Node 3 Total > > --------------- --------------- > > Huge 0.00 0.00 > > Heap 0.00 0.11 > > Stack 0.00 0.03 > > Private 0.00 238361.90 > > ---------------- --------------- --------------- > > Total 0.00 238362.04 > > > > I just want that demotion process to stop when there is 50% free at numa node 0. > > > > Where am I wrong, how to fix that ? > > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and > temporal. As the documentation [2] explains, 'consistent' tuner assumes it > should keep applying the action in a level to keep the goal achieved. In this > case, for example, the demotion scheme assumes there will be continued > promotion and therefore it keeps demoting. If there is no appropriate > promotion, it could result in demoting more than expected amount. > > To me, this seems like the workload has no much hot data, so demotion is much > more stronger. > > For path forward, I'd suggest trying 'temporal' tuner [2]. It is designed to > immediately stop after achieving the goal. > > If it still doesn't work, I'd suggest tracing damon:damos_esz tracepoint with > the numastat output. It will show if the autotune is working as expected. > > For long term production use case, I'd like to suggest running the example > tiering config [1] and see if it also gives you unexpected results. > > [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh > [2] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#aim-oriented-feedback-driven-auto-tuning > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-28 17:02 ` Anton Gavriliuk @ 2026-09-29 7:49 ` SJ Park 2026-09-29 8:41 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-29 7:49 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Mon, 28 Sep 2026 20:02:43 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > Thank you for sharing your use case and question! > > Thank you for your quick response and your answer/comments. > > We already had memory tiering implemented at hardware level 5-7 years > ago with Intel's Optane PEM when configured in Memory Mode. > Now we do the same thing at Linux level :-) Interesting! > I'm new at DAMON/DAMOS, there are lots of new tunables to me. Indeed there are. I'd recommend starting with existing DAMON-based memory tiering solutions including SK Hynix HMSDK and my auto-tune based memory tiering script [1] and ask questions to the authors. > > > The example memory tiering script [1] uses 4% as intervals goal. Is there a > > reason to use 97% as the goal instead? > > Oohhh... I moved to 5%. I understand you mean 4%? > > > Is the above node_mum_used_bp what you really want? That means you want to > > promote hot pages from node 2 to node 0, until the node 0 memory utilization > > becomes 99.7%. That overlaps with the demotion goal (50% free memory of node > > 0) quite a lot. I'd suggest smaller overlap, say, 50.3%, to keep healthy > > circulation of hot/cold pages while not consuming too much resource under > > stabilized access pattern. > > > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and > > temporal. As the documentation [2] explains, 'consistent' tuner assumes it > > should keep applying the action in a level to keep the goal achieved. In this > > case, for example, the demotion scheme assumes there will be continued > > promotion and therefore it keeps demoting. If there is no appropriate > > promotion, it could result in demoting more than expected amount. > > So let me firstly understand what I really want to test :-) Sure, and please feel free to ask any question in any ways. If you prefer to, you could also ask questions privately to me or use DAMON Beer/Coffee/Tea Chat series [2]. [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh [2] https://docs.google.com/document/d/1v43Kcj3ly4CYqmAkMaZzLiM2GEnWfgdGbZAH3mi2vpM/edit?usp=sharing Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-29 7:49 ` SJ Park @ 2026-09-29 8:41 ` Anton Gavriliuk 2026-09-29 9:25 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-09-29 8:41 UTC (permalink / raw) To: SJ Park; +Cc: damon Hello I decided to move step-by-step, so firstly I would get on Linux/bare metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1 - https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency. "Threshold Activation: The VMkernel initiates memory tiering only when total host physical DRAM (Tier 0) consumption reaches approximately the 80% threshold. If consumption is below this trigger (e.g., at 70-75%), the system will not migrate pages to Tier 1. Page Classification: ESXi monitors page activity based on recency and frequency. Only memory pages strictly classified as "cold" (inactive) are migrated to the NVMe Tier 1 device. The active working set ("hot" pages) is deliberately retained in Tier 0 to prevent performance penalties. Scan Rate: The page scan and migration rate scale with memory pressure. At lower consumption levels (just crossing the 80% mark), the rate is highly conservative." Here is my corrected command, damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_cold 2 --damos_access_rate 0% 0% \--damos_apply_interval 1s \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0 \--damos_filter reject young \--numa_node 2 --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0 --damos_access_rate 5% max \--damos_apply_interval 1s \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal node_mem_used_bp 50.3% 0 \--damos_filter allow young \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1 --nr_schemes 1 1 --nr_ctxs 1 1 with consist -> temporal, [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner temporal [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner temporal But it again didn't stop ~50% [root@localhost ~]# numastat -p $(pgrep valkey-server) Per-node process memory usage (in MBs) for PID 18036 (valkey-server) Node 0 Node 1 Node 2 --------------- --------------- --------------- Huge 0.00 0.00 0.00 Heap 0.11 0.00 0.00 Stack 0.03 0.00 0.00 Private 228807.54 0.48 185.79 ---------------- --------------- --------------- --------------- Total 228807.67 0.48 185.79 Node 3 Total --------------- --------------- Huge 0.00 0.00 Heap 0.00 0.11 Stack 0.00 0.03 Private 0.00 228993.80 ---------------- --------------- --------------- Total 0.00 228993.94 [root@localhost ~]# numastat -p $(pgrep valkey-server) Per-node process memory usage (in MBs) for PID 18036 (valkey-server) Node 0 Node 1 Node 2 --------------- --------------- --------------- Huge 0.00 0.00 0.00 Heap 0.01 0.00 0.10 Stack 0.01 0.00 0.02 Private 13434.01 0.48 215559.31 ---------------- --------------- --------------- --------------- Total 13434.03 0.48 215559.43 Node 3 Total --------------- --------------- Huge 0.00 0.00 Heap 0.00 0.11 Stack 0.00 0.03 Private 0.00 228993.80 ---------------- --------------- --------------- Total 0.00 228993.94 > > Oohhh... I moved to 5%. > I understand you mean 4%? As far as I understood, this value represents "accuracy", so there shouldn't be a big difference between 4% vs 5%. Or please correct me if I'm wrong. Anton вт, 29 сент. 2026 г. в 10:49, SJ Park <sj@kernel.org>: > > On Mon, 28 Sep 2026 20:02:43 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > > Thank you for sharing your use case and question! > > > > Thank you for your quick response and your answer/comments. > > > > We already had memory tiering implemented at hardware level 5-7 years > > ago with Intel's Optane PEM when configured in Memory Mode. > > Now we do the same thing at Linux level :-) > > Interesting! > > > I'm new at DAMON/DAMOS, there are lots of new tunables to me. > > Indeed there are. I'd recommend starting with existing DAMON-based memory > tiering solutions including SK Hynix HMSDK and my auto-tune based memory > tiering script [1] and ask questions to the authors. > > > > > > The example memory tiering script [1] uses 4% as intervals goal. Is there a > > > reason to use 97% as the goal instead? > > > > Oohhh... I moved to 5%. > > I understand you mean 4%? > > > > > > Is the above node_mum_used_bp what you really want? That means you want to > > > promote hot pages from node 2 to node 0, until the node 0 memory utilization > > > becomes 99.7%. That overlaps with the demotion goal (50% free memory of node > > > 0) quite a lot. I'd suggest smaller overlap, say, 50.3%, to keep healthy > > > circulation of hot/cold pages while not consuming too much resource under > > > stabilized access pattern. > > > > > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and > > > temporal. As the documentation [2] explains, 'consistent' tuner assumes it > > > should keep applying the action in a level to keep the goal achieved. In this > > > case, for example, the demotion scheme assumes there will be continued > > > promotion and therefore it keeps demoting. If there is no appropriate > > > promotion, it could result in demoting more than expected amount. > > > > So let me firstly understand what I really want to test :-) > > Sure, and please feel free to ask any question in any ways. If you prefer to, > you could also ask questions privately to me or use DAMON Beer/Coffee/Tea Chat > series [2]. > > [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh > [2] https://docs.google.com/document/d/1v43Kcj3ly4CYqmAkMaZzLiM2GEnWfgdGbZAH3mi2vpM/edit?usp=sharing > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-29 8:41 ` Anton Gavriliuk @ 2026-09-29 9:25 ` SJ Park 2026-09-29 13:54 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-29 9:25 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Tue, 29 Sep 2026 11:41:15 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > Hello > > I decided to move step-by-step, so firstly I would get on Linux/bare > metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1 > - https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency. > > "Threshold Activation: The VMkernel initiates memory tiering only when > total host physical DRAM (Tier 0) consumption reaches approximately > the 80% threshold. If consumption is below this trigger (e.g., at > 70-75%), the system will not migrate pages to Tier 1. > Page Classification: ESXi monitors page activity based on recency and > frequency. Only memory pages strictly classified as "cold" (inactive) > are migrated to the NVMe Tier 1 device. The active working set ("hot" > pages) is deliberately retained in Tier 0 to prevent performance > penalties. > Scan Rate: The page scan and migration rate scale with memory > pressure. At lower consumption levels (just crossing the 80% mark), > the rate is highly conservative." Sounds good. > > Here is my corrected command, > > damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s Fyi, '--monitoring_intervals_autotune' is same to '--monitoring_intervals_goal 4% 3 5ms 10s'. > \--damos_action migrate_cold 2 --damos_access_rate 0% 0% > \--damos_apply_interval 1s \--damos_quota_interval 1s > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0 > \--damos_filter reject young \--numa_node 2 > --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0 > --damos_access_rate 5% max \--damos_apply_interval 1s > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal > node_mem_used_bp 50.3% 0 \--damos_filter allow young > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1 > --nr_schemes 1 1 --nr_ctxs 1 1 Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context, target and scheme parameters on the command line using --kdamond, --damon_ctx, --damon_target, and --damos_scheme. you coud use those instead of --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and --nr_ctxs. E.g., damo start \ ` # A kdamond to demote cold memory from node 0 to node 2 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 0 \ --damos_scheme \ --damos_action migrate_cold 2 \ --damos_access_rate 0% 0% --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50% node 0 free memory ` \ --damos_quota_goal node_mem_free_bp 50% 0 \ --damos_filter reject young \ ` # A kdamond to promote hot memory from node 2 to node 0 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 2 \ --damos_scheme \ --damos_action migrate_hot 0 \ --damos_access_rate 5% max --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50.3% node 0 memory utilization ` \ --damos_quota_goal node_mem_used_bp 50.3% 0 \ --damos_filter allow young \ > > with consist -> temporal, > > [root@localhost ~]# cat > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner > temporal > [root@localhost ~]# cat > /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner > temporal Seems you manually made this change. Has this executed after 'damo start'? Also, did you 'commit' the updated commit input? If any of your answer to the questions is not "yes", the tuner update may not applied. You can specify what tuner to use on 'damo' command together, using '--damos_quota_goal_tuner' option. I'd recommend using that. E.g., damo start \ ` # A kdamond to demote cold memory from node 0 to node 2 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 0 \ --damos_scheme \ --damos_action migrate_cold 2 \ --damos_access_rate 0% 0% --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50% node 0 free memory ` \ --damos_quota_goal node_mem_free_bp 50% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter reject young \ ` # A kdamond to promote hot memory from node 2 to node 0 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 2 \ --damos_scheme \ --damos_action migrate_hot 0 \ --damos_access_rate 5% max --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50.3% node 0 memory utilization ` \ --damos_quota_goal node_mem_used_bp 50.3% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter allow young \ > > But it again didn't stop ~50% > > [root@localhost ~]# numastat -p $(pgrep valkey-server) > > Per-node process memory usage (in MBs) for PID 18036 (valkey-server) > Node 0 Node 1 Node 2 > --------------- --------------- --------------- > Huge 0.00 0.00 0.00 > Heap 0.11 0.00 0.00 > Stack 0.03 0.00 0.00 > Private 228807.54 0.48 185.79 > ---------------- --------------- --------------- --------------- > Total 228807.67 0.48 185.79 > > Node 3 Total > --------------- --------------- > Huge 0.00 0.00 > Heap 0.00 0.11 > Stack 0.00 0.03 > Private 0.00 228993.80 > ---------------- --------------- --------------- > Total 0.00 228993.94 > > > [root@localhost ~]# numastat -p $(pgrep valkey-server) > > Per-node process memory usage (in MBs) for PID 18036 (valkey-server) > Node 0 Node 1 Node 2 > --------------- --------------- --------------- > Huge 0.00 0.00 0.00 > Heap 0.01 0.00 0.10 > Stack 0.01 0.00 0.02 > Private 13434.01 0.48 215559.31 > ---------------- --------------- --------------- --------------- > Total 13434.03 0.48 215559.43 > > Node 3 Total > --------------- --------------- > Huge 0.00 0.00 > Heap 0.00 0.11 > Stack 0.00 0.03 > Private 0.00 228993.80 > ---------------- --------------- --------------- > Total 0.00 228993.94 Interesting. I doubt if the tuner change is correctly made. 'damo report damon' can help us understand under what configuration DAMON is running. It could help us quickly see if the configuration is done as we intended. Could you share 'damo report damon' result on the final state? It would also be helpful if you could run 'damo report damon --damos_stats' periodically (say, once per 5-10 seconds) while the migration is ongoing and share the outputs with us. > > > > > > Oohhh... I moved to 5%. > > I understand you mean 4%? > > As far as I understood, this value represents "accuracy", so there > shouldn't be a big difference between 4% vs 5%. Or please correct me > if I'm wrong. Yes, 5% vs 4% should be an ignorable small difference. Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-29 9:25 ` SJ Park @ 2026-09-29 13:54 ` Anton Gavriliuk 2026-09-29 17:37 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-09-29 13:54 UTC (permalink / raw) To: SJ Park; +Cc: damon > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context, > target and scheme parameters on the command line using --kdamond, --damon_ctx, > --damon_target, and --damos_scheme. you coud use those instead of > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and > --nr_ctxs. E.g., Fedora shows old damo version, [root@localhost anton]# rpm -qa|grep -i damo damo-3.3.0-1.fc44.noarch [root@localhost anton]# so I configured latest available 3.4.1 [root@localhost damo]# damo version 3.4.1 [root@localhost damo]# which damo /home/anton/damo/damo [root@localhost damo]# > Seems you manually made this change. Has this executed after 'damo start'? > Also, did you 'commit' the updated commit input? If any of your answer to the > questions is not "yes", the tuner update may not applied. Yes, it was executed before 'damo start', but I didn't do 'commit'. > It would also be helpful if you could run 'damo report damon --damos_stats' > periodically (say, once per 5-10 seconds) while the migration is ongoing and > share the outputs with us. Ok, initially I have, [root@localhost ~]# numastat -z -p $(pgrep valkey-server) Per-node process memory usage (in MBs) for PID 25361 (valkey-server) Node 0 Node 1 Node 2 --------------- --------------- --------------- Heap 0.11 0.00 0.00 Stack 0.03 0.00 0.00 Private 238356.60 0.50 4.86 ---------------- --------------- --------------- --------------- Total 238356.73 0.50 4.86 Total --------------- Heap 0.11 Stack 0.03 Private 238361.95 ---------------- --------------- Total 238362.09 [root@localhost ~]# I run [root@localhost ~]# /home/anton/damo/damo start \ ` # A kdamond to demote cold memory from node 0 to node 2 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 0 \ --damos_scheme \ --damos_action migrate_cold 2 \ --damos_access_rate 0% 0% --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50% node 0 free memory ` \ --damos_quota_goal node_mem_free_bp 50% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter reject young \ ` # A kdamond to promote hot memory from node 2 to node 0 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 2 \ --damos_scheme \ --damos_action migrate_hot 0 \ --damos_access_rate 5% max --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50.3% node 0 memory utilization ` \ --damos_quota_goal node_mem_used_bp 50.3% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter allow young \ > sysinfo loading fail (info update fail (sysfs feature check fail (feature map making fail (staging damos goal feature check purpose kdamond failed)))) [root@localhost ~]# [root@localhost ~]# ps -ef|grep -i damo root 26051 25396 0 16:49 pts/2 00:00:00 grep --color=auto -i damo [root@localhost ~]# > It would also be helpful if you could run 'damo report damon --damos_stats' > periodically (say, once per 5-10 seconds) while the migration is ongoing and > share the outputs with us. Sure, I will start something like that, while true; do damo report damon --damos_stats >> /home/anton/damo_report_damon_damos_stats; sleep 10; done once damo will be started. > Interesting. I doubt if the tuner change is correctly made. 'damo report > damon' can help us understand under what configuration DAMON is running. It > could help us quickly see if the configuration is done as we intended. Could > you share 'damo report damon' result on the final state? There are errors above during damo start. Anton вт, 29 сент. 2026 г. в 12:25, SJ Park <sj@kernel.org>: > > On Tue, 29 Sep 2026 11:41:15 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > Hello > > > > I decided to move step-by-step, so firstly I would get on Linux/bare > > metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1 > > - https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency. > > > > "Threshold Activation: The VMkernel initiates memory tiering only when > > total host physical DRAM (Tier 0) consumption reaches approximately > > the 80% threshold. If consumption is below this trigger (e.g., at > > 70-75%), the system will not migrate pages to Tier 1. > > Page Classification: ESXi monitors page activity based on recency and > > frequency. Only memory pages strictly classified as "cold" (inactive) > > are migrated to the NVMe Tier 1 device. The active working set ("hot" > > pages) is deliberately retained in Tier 0 to prevent performance > > penalties. > > Scan Rate: The page scan and migration rate scale with memory > > pressure. At lower consumption levels (just crossing the 80% mark), > > the rate is highly conservative." > > Sounds good. > > > > > Here is my corrected command, > > > > damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s > > Fyi, '--monitoring_intervals_autotune' is same to > '--monitoring_intervals_goal 4% 3 5ms 10s'. > > > \--damos_action migrate_cold 2 --damos_access_rate 0% 0% > > \--damos_apply_interval 1s \--damos_quota_interval 1s > > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0 > > \--damos_filter reject young \--numa_node 2 > > --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0 > > --damos_access_rate 5% max \--damos_apply_interval 1s > > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal > > node_mem_used_bp 50.3% 0 \--damos_filter allow young > > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1 > > --nr_schemes 1 1 --nr_ctxs 1 1 > > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context, > target and scheme parameters on the command line using --kdamond, --damon_ctx, > --damon_target, and --damos_scheme. you coud use those instead of > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and > --nr_ctxs. E.g., > > damo start \ > ` # A kdamond to demote cold memory from node 0 to node 2 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 0 \ > --damos_scheme \ > --damos_action migrate_cold 2 \ > --damos_access_rate 0% 0% --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50% node 0 free memory ` \ > --damos_quota_goal node_mem_free_bp 50% 0 \ > --damos_filter reject young \ > ` # A kdamond to promote hot memory from node 2 to node 0 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 2 \ > --damos_scheme \ > --damos_action migrate_hot 0 \ > --damos_access_rate 5% max --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50.3% node 0 memory utilization ` \ > --damos_quota_goal node_mem_used_bp 50.3% 0 \ > --damos_filter allow young \ > > > > > with consist -> temporal, > > > > [root@localhost ~]# cat > > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner > > temporal > > [root@localhost ~]# cat > > /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner > > temporal > > Seems you manually made this change. Has this executed after 'damo start'? > Also, did you 'commit' the updated commit input? If any of your answer to the > questions is not "yes", the tuner update may not applied. > > You can specify what tuner to use on 'damo' command together, using > '--damos_quota_goal_tuner' option. I'd recommend using that. E.g., > > damo start \ > ` # A kdamond to demote cold memory from node 0 to node 2 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 0 \ > --damos_scheme \ > --damos_action migrate_cold 2 \ > --damos_access_rate 0% 0% --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50% node 0 free memory ` \ > --damos_quota_goal node_mem_free_bp 50% 0 \ > ` # using temporal tuner ` \ > --damos_quota_goal_tuner temporal \ > --damos_filter reject young \ > ` # A kdamond to promote hot memory from node 2 to node 0 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 2 \ > --damos_scheme \ > --damos_action migrate_hot 0 \ > --damos_access_rate 5% max --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50.3% node 0 memory utilization ` \ > --damos_quota_goal node_mem_used_bp 50.3% 0 \ > ` # using temporal tuner ` \ > --damos_quota_goal_tuner temporal \ > --damos_filter allow young \ > > > > > But it again didn't stop ~50% > > > > [root@localhost ~]# numastat -p $(pgrep valkey-server) > > > > Per-node process memory usage (in MBs) for PID 18036 (valkey-server) > > Node 0 Node 1 Node 2 > > --------------- --------------- --------------- > > Huge 0.00 0.00 0.00 > > Heap 0.11 0.00 0.00 > > Stack 0.03 0.00 0.00 > > Private 228807.54 0.48 185.79 > > ---------------- --------------- --------------- --------------- > > Total 228807.67 0.48 185.79 > > > > Node 3 Total > > --------------- --------------- > > Huge 0.00 0.00 > > Heap 0.00 0.11 > > Stack 0.00 0.03 > > Private 0.00 228993.80 > > ---------------- --------------- --------------- > > Total 0.00 228993.94 > > > > > > [root@localhost ~]# numastat -p $(pgrep valkey-server) > > > > Per-node process memory usage (in MBs) for PID 18036 (valkey-server) > > Node 0 Node 1 Node 2 > > --------------- --------------- --------------- > > Huge 0.00 0.00 0.00 > > Heap 0.01 0.00 0.10 > > Stack 0.01 0.00 0.02 > > Private 13434.01 0.48 215559.31 > > ---------------- --------------- --------------- --------------- > > Total 13434.03 0.48 215559.43 > > > > Node 3 Total > > --------------- --------------- > > Huge 0.00 0.00 > > Heap 0.00 0.11 > > Stack 0.00 0.03 > > Private 0.00 228993.80 > > ---------------- --------------- --------------- > > Total 0.00 228993.94 > > Interesting. I doubt if the tuner change is correctly made. 'damo report > damon' can help us understand under what configuration DAMON is running. It > could help us quickly see if the configuration is done as we intended. Could > you share 'damo report damon' result on the final state? > > It would also be helpful if you could run 'damo report damon --damos_stats' > periodically (say, once per 5-10 seconds) while the migration is ongoing and > share the outputs with us. > > > > > > > > > > > Oohhh... I moved to 5%. > > > I understand you mean 4%? > > > > As far as I understood, this value represents "accuracy", so there > > shouldn't be a big difference between 4% vs 5%. Or please correct me > > if I'm wrong. > > Yes, 5% vs 4% should be an ignorable small difference. > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-29 13:54 ` Anton Gavriliuk @ 2026-09-29 17:37 ` SJ Park 2026-09-30 4:03 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-29 17:37 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Tue, 29 Sep 2026 16:54:41 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context, > > target and scheme parameters on the command line using --kdamond, --damon_ctx, > > --damon_target, and --damos_scheme. you coud use those instead of > > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and > > --nr_ctxs. E.g., > > Fedora shows old damo version, > > [root@localhost anton]# rpm -qa|grep -i damo > damo-3.3.0-1.fc44.noarch > [root@localhost anton]# > > so I configured latest available 3.4.1 > > [root@localhost damo]# damo version > 3.4.1 > [root@localhost damo]# which damo > /home/anton/damo/damo > [root@localhost damo]# Thank you! Hopefully upgrading the version was not that difficult. You can simply git-clone the repo and use the 'damo' executable file under the local-cloned repo. > > > Seems you manually made this change. Has this executed after 'damo start'? > > Also, did you 'commit' the updated commit input? If any of your answer to the > > questions is not "yes", the tuner update may not applied. > > Yes, it was executed before 'damo start', but I didn't do 'commit'. To do this manually, you should write the files after 'damo start', and also do 'commit'. Anyway, this means your previous run was using 'consist' tuner. That explains why it didn't show any difference. > > > It would also be helpful if you could run 'damo report damon --damos_stats' > > periodically (say, once per 5-10 seconds) while the migration is ongoing and > > share the outputs with us. > > Ok, initially I have, > > [root@localhost ~]# numastat -z -p $(pgrep valkey-server) > > Per-node process memory usage (in MBs) for PID 25361 (valkey-server) > Node 0 Node 1 Node 2 > --------------- --------------- --------------- > Heap 0.11 0.00 0.00 > Stack 0.03 0.00 0.00 > Private 238356.60 0.50 4.86 > ---------------- --------------- --------------- --------------- > Total 238356.73 0.50 4.86 > > Total > --------------- > Heap 0.11 > Stack 0.03 > Private 238361.95 > ---------------- --------------- > Total 238362.09 > [root@localhost ~]# > > I run > > [root@localhost ~]# /home/anton/damo/damo start \ > ` # A kdamond to demote cold memory from node 0 to node 2 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 0 \ > --damos_scheme \ > --damos_action migrate_cold 2 \ > --damos_access_rate 0% 0% --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50% node 0 free memory ` \ > --damos_quota_goal node_mem_free_bp 50% 0 \ > ` # using temporal tuner ` \ > --damos_quota_goal_tuner temporal \ > --damos_filter reject young \ > ` # A kdamond to promote hot memory from node 2 to node 0 ` \ > --kdamond --damon_ctx --monitoring_intervals_autotune \ > --damon_target --numa_node 2 \ > --damos_scheme \ > --damos_action migrate_hot 0 \ > --damos_access_rate 5% max --damos_apply_interval 1s \ > ` # up to 8 GiB per second ` \ > --damos_quota_interval 1s --damos_quota_space 8G \ > ` # aiming at least 50.3% node 0 memory utilization ` \ > --damos_quota_goal node_mem_used_bp 50.3% 0 \ > ` # using temporal tuner ` \ > --damos_quota_goal_tuner temporal \ > --damos_filter allow young \ > > > sysinfo loading fail (info update fail (sysfs feature check fail > (feature map making fail (staging damos goal feature check purpose > kdamond failed)))) Oops... Seems you were running next branch of damo. There was a bug. I reproduced it on 7.2.8 kernel, and fixed it. The fix [1] is now pushed. Could you pull the 'next' branch and try again? > [root@localhost ~]# > [root@localhost ~]# ps -ef|grep -i damo > root 26051 25396 0 16:49 pts/2 00:00:00 grep --color=auto -i damo > [root@localhost ~]# > > > It would also be helpful if you could run 'damo report damon --damos_stats' > > periodically (say, once per 5-10 seconds) while the migration is ongoing and > > share the outputs with us. > > Sure, I will start something like that, > while true; do damo report damon --damos_stats >> > /home/anton/damo_report_damon_damos_stats; sleep 10; done > once damo will be started. Sounds good. > > > > Interesting. I doubt if the tuner change is correctly made. 'damo report > > damon' can help us understand under what configuration DAMON is running. It > > could help us quickly see if the configuration is done as we intended. Could > > you share 'damo report damon' result on the final state? > > There are errors above during damo start. Apparently it was a bug in damo's next branch. As I mentioned above, the fix is now pushed. Could you try again? [1] https://github.com/damonitor/damo/commit/fd2fb55b44ba03db55b8a025eef43dcf569831e5 Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-29 17:37 ` SJ Park @ 2026-09-30 4:03 ` Anton Gavriliuk 2026-09-30 8:35 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-09-30 4:03 UTC (permalink / raw) To: SJ Park; +Cc: damon [-- Attachment #1: Type: text/plain, Size: 15270 bytes --] > Oops... Seems you were running next branch of damo. There was a bug. I > reproduced it on 7.2.8 kernel, and fixed it. The fix [1] is now pushed. Could > you pull the 'next' branch and try again? Done. Ok, initially I have, [root@localhost ~]# numastat -z -p $(pgrep valkey-server) Per-node process memory usage (in MBs) for PID 25361 (valkey-server) Node 0 Node 1 Node 2 --------------- --------------- --------------- Heap 0.11 0.00 0.00 Stack 0.03 0.00 0.00 Private 238356.60 0.50 4.86 ---------------- --------------- --------------- --------------- Total 238356.73 0.50 4.86 Total --------------- Heap 0.11 Stack 0.03 Private 238361.95 ---------------- --------------- Total 238362.09 [root@localhost ~]# I run /home/anton/damo/damo start \ ` # A kdamond to demote cold memory from node 0 to node 2 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 0 \ --damos_scheme \ --damos_action migrate_cold 2 \ --damos_access_rate 0% 0% --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50% node 0 free memory ` \ --damos_quota_goal node_mem_free_bp 50% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter reject young \ ` # A kdamond to promote hot memory from node 2 to node 0 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 2 \ --damos_scheme \ --damos_action migrate_hot 0 \ --damos_access_rate 5% max --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 50.3% node 0 memory utilization ` \ --damos_quota_goal node_mem_used_bp 50.3% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter allow young \ > > [root@localhost ~]# ps -ef|grep -i dam root 29836 2 2 06:37 ? 00:00:00 [kdamond.0] root 29837 2 0 06:37 ? 00:00:00 [kdamond.1] root 29842 29644 0 06:37 pts/2 00:00:00 grep --color=auto -i dam [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner temporal [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner temporal [root@localhost ~]# > It would also be helpful if you could run 'damo report damon --damos_stats' > periodically (say, once per 5-10 seconds) while the migration is ongoing and > share the outputs with us. Sure, I will start something like that, while true; do /home/anton/damo/damo report damon --damos_stats >> /home/anton/damo_report_damon_damos_stats; sleep 10; done once damo will be started. Please check the attached file. It looks now it keeps ~50% free numa node 0, [root@localhost ~]# numastat -z -p $(pgrep valkey-server) Per-node process memory usage (in MBs) for PID 25361 (valkey-server) Node 0 Node 1 Node 2 --------------- --------------- --------------- Heap 0.05 0.00 0.06 Stack 0.01 0.00 0.02 Private 138725.66 0.50 99635.75 ---------------- --------------- --------------- --------------- Total 138725.71 0.50 99635.82 Total --------------- Heap 0.11 Stack 0.03 Private 238361.90 ---------------- --------------- Total 238362.04 [root@localhost ~]# > Interesting. I doubt if the tuner change is correctly made. 'damo report > damon' can help us understand under what configuration DAMON is running. It > could help us quickly see if the configuration is done as we intended. Could > you share 'damo report damon' result on the final state? [root@localhost ~]# /home/anton/damo/damo report damon kdamond 0 state: on, pid: 29836 context 0 ops: paddr target 0 pid: 0 region [4,096, 581,632) (564.000 KiB) region [589,824, 655,360) (64.000 KiB) region [1,048,576, 2,079,305,728) (1.936 GiB) region [2,079,363,072, 2,153,725,952) (70.918 MiB) region [2,153,926,656, 2,154,790,912) (844.000 KiB) region [2,154,811,392, 2,154,909,696) (96.000 KiB) region [2,154,930,176, 2,354,724,864) (190.539 MiB) region [2,355,859,456, 2,355,871,744) (12.000 KiB) region [2,372,751,360, 2,372,755,456) (4.000 KiB) region [2,372,759,552, 2,372,771,840) (12.000 KiB) region [2,391,846,912, 2,551,230,464) (152.000 MiB) region [2,604,707,840, 2,604,716,032) (8.000 KiB) region [2,605,244,416, 2,944,397,312) (323.441 MiB) region [2,944,397,372, 2,944,401,408) (3.941 KiB) region [4,294,967,296, 413,390,602,240) (381.000 GiB) intervals sample 5.120 s, aggr 1 m 42.400 s, update 1 s target 4 % accesses per 3 aggrs, [5 ms, 10 s] sampling interval nr_regions: [10, 1,000] scheme 0 action: migrate_cold to node 2 per 1 s target access pattern sz: [0 B, max] nr_accesses: [0 samples, 0 samples] age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals] quotas 0 ns / 8.000 GiB / 0 B per 1 s goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0 goal tuner: temporal priority: sz 0.1 %, nr_accesses 0.1 %, age 0.1 % watermarks metric none, interval 0 ns 0 %, 0 %, 0 % filter 0 reject young statistics tried 142 times (560.000 GiB) applied 55 times (97.342 GiB) 97.351 GiB passed filters quota exceeded 234 times 97.351 GiB tried 0 snapshots (max 0) tried regions (0 B) access sample control enabled primitives: page_table kdamond 1 state: on, pid: 29837 context 0 ops: paddr target 0 pid: 0 region [466,003,951,616, 3,646,427,234,304) (2.893 TiB) intervals sample 10 s, aggr 3 m 20 s, update 1 s target 4 % accesses per 3 aggrs, [5 ms, 10 s] sampling interval nr_regions: [10, 1,000] scheme 0 action: migrate_hot to node 0 per 1 s target access pattern sz: [0 B, max] nr_accesses: [1 samples, 3,689,348,814,741,910,528 samples] age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals] quotas 0 ns / 8.000 GiB / 8.000 GiB per 1 s goal 0: metric node_mem_used_bp (nid 0) target 5,030 current 0 goal tuner: temporal priority: sz 0.1 %, nr_accesses 0.1 %, age 0.1 % watermarks metric none, interval 0 ns 0 %, 0 %, 0 % filter 0 allow young statistics tried 0 times (0 B) applied 0 times (0 B) 0 B passed filters quota exceeded 150 times 0 B tried 95 snapshots (max 0) tried regions (0 B) access sample control enabled primitives: page_table damon_reclaim: off damon_stat: off [root@localhost ~]# DAMON & DAMOS are very flexible, so it requires a deep understanding of specific workload and how to configure memory tiering with DAMON & DAMOS. So firstly I decided to make it like VMware VCF 9.1 does :-)), next I will reduce the goal from 50% free to 25-25% for numa node 0. Based on the outputs above, please let me know if other improvements are required. Anton вт, 29 сент. 2026 г. в 20:37, SJ Park <sj@kernel.org>: > > On Tue, 29 Sep 2026 16:54:41 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context, > > > target and scheme parameters on the command line using --kdamond, --damon_ctx, > > > --damon_target, and --damos_scheme. you coud use those instead of > > > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and > > > --nr_ctxs. E.g., > > > > Fedora shows old damo version, > > > > [root@localhost anton]# rpm -qa|grep -i damo > > damo-3.3.0-1.fc44.noarch > > [root@localhost anton]# > > > > so I configured latest available 3.4.1 > > > > [root@localhost damo]# damo version > > 3.4.1 > > [root@localhost damo]# which damo > > /home/anton/damo/damo > > [root@localhost damo]# > > Thank you! Hopefully upgrading the version was not that difficult. You can > simply git-clone the repo and use the 'damo' executable file under the > local-cloned repo. > > > > > > Seems you manually made this change. Has this executed after 'damo start'? > > > Also, did you 'commit' the updated commit input? If any of your answer to the > > > questions is not "yes", the tuner update may not applied. > > > > Yes, it was executed before 'damo start', but I didn't do 'commit'. > > To do this manually, you should write the files after 'damo start', and also do > 'commit'. > > Anyway, this means your previous run was using 'consist' tuner. That explains > why it didn't show any difference. > > > > > > It would also be helpful if you could run 'damo report damon --damos_stats' > > > periodically (say, once per 5-10 seconds) while the migration is ongoing and > > > share the outputs with us. > > > > Ok, initially I have, > > > > [root@localhost ~]# numastat -z -p $(pgrep valkey-server) > > > > Per-node process memory usage (in MBs) for PID 25361 (valkey-server) > > Node 0 Node 1 Node 2 > > --------------- --------------- --------------- > > Heap 0.11 0.00 0.00 > > Stack 0.03 0.00 0.00 > > Private 238356.60 0.50 4.86 > > ---------------- --------------- --------------- --------------- > > Total 238356.73 0.50 4.86 > > > > Total > > --------------- > > Heap 0.11 > > Stack 0.03 > > Private 238361.95 > > ---------------- --------------- > > Total 238362.09 > > [root@localhost ~]# > > > > I run > > > > [root@localhost ~]# /home/anton/damo/damo start \ > > ` # A kdamond to demote cold memory from node 0 to node 2 ` \ > > --kdamond --damon_ctx --monitoring_intervals_autotune \ > > --damon_target --numa_node 0 \ > > --damos_scheme \ > > --damos_action migrate_cold 2 \ > > --damos_access_rate 0% 0% --damos_apply_interval 1s \ > > ` # up to 8 GiB per second ` \ > > --damos_quota_interval 1s --damos_quota_space 8G \ > > ` # aiming at least 50% node 0 free memory ` \ > > --damos_quota_goal node_mem_free_bp 50% 0 \ > > ` # using temporal tuner ` \ > > --damos_quota_goal_tuner temporal \ > > --damos_filter reject young \ > > ` # A kdamond to promote hot memory from node 2 to node 0 ` \ > > --kdamond --damon_ctx --monitoring_intervals_autotune \ > > --damon_target --numa_node 2 \ > > --damos_scheme \ > > --damos_action migrate_hot 0 \ > > --damos_access_rate 5% max --damos_apply_interval 1s \ > > ` # up to 8 GiB per second ` \ > > --damos_quota_interval 1s --damos_quota_space 8G \ > > ` # aiming at least 50.3% node 0 memory utilization ` \ > > --damos_quota_goal node_mem_used_bp 50.3% 0 \ > > ` # using temporal tuner ` \ > > --damos_quota_goal_tuner temporal \ > > --damos_filter allow young \ > > > > > sysinfo loading fail (info update fail (sysfs feature check fail > > (feature map making fail (staging damos goal feature check purpose > > kdamond failed)))) > > Oops... Seems you were running next branch of damo. There was a bug. I > reproduced it on 7.2.8 kernel, and fixed it. The fix [1] is now pushed. Could > you pull the 'next' branch and try again? > > > [root@localhost ~]# > > [root@localhost ~]# ps -ef|grep -i damo > > root 26051 25396 0 16:49 pts/2 00:00:00 grep --color=auto -i damo > > [root@localhost ~]# > > > > > It would also be helpful if you could run 'damo report damon --damos_stats' > > > periodically (say, once per 5-10 seconds) while the migration is ongoing and > > > share the outputs with us. > > > > Sure, I will start something like that, > > while true; do damo report damon --damos_stats >> > > /home/anton/damo_report_damon_damos_stats; sleep 10; done > > once damo will be started. > > Sounds good. > > > > > > > > Interesting. I doubt if the tuner change is correctly made. 'damo report > > > damon' can help us understand under what configuration DAMON is running. It > > > could help us quickly see if the configuration is done as we intended. Could > > > you share 'damo report damon' result on the final state? > > > > There are errors above during damo start. > > Apparently it was a bug in damo's next branch. As I mentioned above, the fix > is now pushed. Could you try again? > > [1] https://github.com/damonitor/damo/commit/fd2fb55b44ba03db55b8a025eef43dcf569831e5 > > Thanks, > SJ > > [...] [-- Attachment #2: damo_report_damon_damos_stats.gz --] [-- Type: application/x-gzip, Size: 833 bytes --] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-30 4:03 ` Anton Gavriliuk @ 2026-09-30 8:35 ` SJ Park 2026-09-30 16:26 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-30 8:35 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Wed, 30 Sep 2026 07:03:47 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > Oops... Seems you were running next branch of damo. There was a bug. I > > reproduced it on 7.2.8 kernel, and fixed it. The fix [1] is now pushed. Could > > you pull the 'next' branch and try again? > > Done. [...] > It looks now it keeps ~50% free numa node 0, Awesome :D [...] > [root@localhost ~]# /home/anton/damo/damo report damon > kdamond 0 > state: on, pid: 29836 > context 0 > ops: paddr [...] > scheme 0 > action: migrate_cold to node 2 per 1 s > target access pattern > sz: [0 B, max] > nr_accesses: [0 samples, 0 samples] > age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals] > quotas > 0 ns / 8.000 GiB / 0 B per 1 s > goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0 > goal tuner: temporal Ok, the above line confirms DAMON is running with 'temporal' tuner. [...] > DAMON & DAMOS are very flexible, so it requires a deep understanding > of specific workload and how to configure memory tiering with DAMON & > DAMOS. Indeed it is. Nonetheless, for the reason we provide DAMON modules [1] or damo scripts for commonly known DAMON usages. We have DAMON_RECLAIM module [2] for proactive memory reclamation, and mem_tier.sh [3] for memory tiering. My suggestion is to start from such examples, find what is not working and asking questions to me. So you are doing all great :D > > So firstly I decided to make it like VMware VCF 9.1 does :-)), next I > will reduce the goal from 50% free to 25-25% for numa node 0. Sounds like a good plan. > > Based on the outputs above, please let me know if other improvements > are required. This temporal tuner test has proved the quota system is not broken. It implies the previous test resulted in migrating nearly all data to the lower tier, because there was no hot data to promote back to the upper tier. In many common benchmarks using zipfian-like access distribution, hot data is always hot, and cold data is always cold. And usually the amount of cold data is much larger than hot data. I found [4] the pattern is making evaluation of tiering solutions difficult, and was using artificial access pattern mixing to work around. My imagined real world workload is, there will be hot and cold data. And the pattern will nearly always be stable. But, occasionally there will be changes. Say, a chunk of data will be hot for a few hours, but suddenly be cold and keep being cold for a few hours, then sudenly be warm for hours, and so on. The auto-tuning based DAMON tiering is designed with such workload in mind. For your testing, I think the setup is good. I'd suggest lower target free memory ratio, though. Assuming the theory (your workload has static access pattern of small hot data) is true, using consist tuner should also be fine. You will have more than expected data in lower tier, but if those are truly cold, why would we bother? If you still want strict upper tier utilization, you could make the access pattern and filter condition less strict. Ideally, the access pattern and filter conditions should all go away, assuming DAMON can find true hot and cold data. And I believe that would be the case for long-running real world workloads. Short-running test workload would show DAMON making wrong decisions in short term, though. If you want to make sure if my theory is true, you could also profile the access pattern of your test workload using DAMON. I actually did it for my auto-tuned tiering test [4] to find the fact. [1] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#special-purpose-access-aware-kernel-modules [2] https://origin.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html [3] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh [4] https://lkml.kernel.org/r/20250420194030.75838-1-sj@kernel.org Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-30 8:35 ` SJ Park @ 2026-09-30 16:26 ` Anton Gavriliuk 2026-09-30 17:43 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-09-30 16:26 UTC (permalink / raw) To: SJ Park; +Cc: damon I would go back to memory tiering VCF9.1-like behavior you helped me configure, for new few questions: 1. It works with the goal of 50% free memory for numa node 0, but when I tried to set 25% free memory, it didn't do anything... [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat Per-node process memory usage (in MBs) for PID 3808 (valkey-server) Node 0 Node 1 Node 2 Node 3 Total --------------- --------------- --------------- --------------- --------------- Huge 0.00 0.00 0.00 0.00 0.00 Heap 0.11 0.00 0.00 0.00 0.11 Stack 0.03 0.00 0.00 0.00 0.03 Private 238356.25 2.19 3.42 0.00 238361.86 ---------------- --------------- --------------- --------------- --------------- --------------- Total 238356.39 2.19 3.42 0.00 238362.00 [root@localhost ~]# /home/anton/damo/damo start \ ` # A kdamond to demote cold memory from node 0 to node 2 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 0 \ --damos_scheme \ --damos_action migrate_cold 2 \ --damos_access_rate 0% 0% --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 25% node 0 free memory ` \ --damos_quota_goal node_mem_free_bp 25% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter reject young \ ` # A kdamond to promote hot memory from node 2 to node 0 ` \ --kdamond --damon_ctx --monitoring_intervals_autotune \ --damon_target --numa_node 2 \ --damos_scheme \ --damos_action migrate_hot 0 \ --damos_access_rate 5% max --damos_apply_interval 1s \ ` # up to 8 GiB per second ` \ --damos_quota_interval 1s --damos_quota_space 8G \ ` # aiming at least 75.3% node 0 memory utilization ` \ --damos_quota_goal node_mem_used_bp 75.3% 0 \ ` # using temporal tuner ` \ --damos_quota_goal_tuner temporal \ --damos_filter allow young \ [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner temporal [root@localhost ~]# cat /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner temporal [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat Per-node process memory usage (in MBs) for PID 3808 (valkey-server) Node 0 Node 1 Node 2 Node 3 Total --------------- --------------- --------------- --------------- --------------- Huge 0.00 0.00 0.00 0.00 0.00 Heap 0.11 0.00 0.00 0.00 0.11 Stack 0.03 0.00 0.00 0.00 0.03 Private 238356.25 2.19 3.42 0.00 238361.86 ---------------- --------------- --------------- --------------- --------------- --------------- Total 238356.39 2.19 3.42 0.00 238362.00 [root@localhost ~]# 2. If there are no un-accessed pages to demote for achieving a defined goal, damon will demote pages with the most rare access ? Anton ср, 30 сент. 2026 г. в 11:35, SJ Park <sj@kernel.org>: > > On Wed, 30 Sep 2026 07:03:47 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > > Oops... Seems you were running next branch of damo. There was a bug. I > > > reproduced it on 7.2.8 kernel, and fixed it. The fix [1] is now pushed. Could > > > you pull the 'next' branch and try again? > > > > Done. > [...] > > It looks now it keeps ~50% free numa node 0, > > Awesome :D > > [...] > > [root@localhost ~]# /home/anton/damo/damo report damon > > kdamond 0 > > state: on, pid: 29836 > > context 0 > > ops: paddr > [...] > > scheme 0 > > action: migrate_cold to node 2 per 1 s > > target access pattern > > sz: [0 B, max] > > nr_accesses: [0 samples, 0 samples] > > age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals] > > quotas > > 0 ns / 8.000 GiB / 0 B per 1 s > > goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0 > > goal tuner: temporal > > Ok, the above line confirms DAMON is running with 'temporal' tuner. > > [...] > > DAMON & DAMOS are very flexible, so it requires a deep understanding > > of specific workload and how to configure memory tiering with DAMON & > > DAMOS. > > Indeed it is. Nonetheless, for the reason we provide DAMON modules [1] or damo > scripts for commonly known DAMON usages. We have DAMON_RECLAIM module [2] for > proactive memory reclamation, and mem_tier.sh [3] for memory tiering. > > My suggestion is to start from such examples, find what is not working and > asking questions to me. So you are doing all great :D > > > > > So firstly I decided to make it like VMware VCF 9.1 does :-)), next I > > will reduce the goal from 50% free to 25-25% for numa node 0. > > Sounds like a good plan. > > > > > Based on the outputs above, please let me know if other improvements > > are required. > > This temporal tuner test has proved the quota system is not broken. It implies > the previous test resulted in migrating nearly all data to the lower tier, > because there was no hot data to promote back to the upper tier. > > In many common benchmarks using zipfian-like access distribution, hot data is > always hot, and cold data is always cold. And usually the amount of cold data > is much larger than hot data. I found [4] the pattern is making evaluation of > tiering solutions difficult, and was using artificial access pattern mixing to > work around. > > My imagined real world workload is, there will be hot and cold data. And the > pattern will nearly always be stable. But, occasionally there will be changes. > Say, a chunk of data will be hot for a few hours, but suddenly be cold and keep > being cold for a few hours, then sudenly be warm for hours, and so on. The > auto-tuning based DAMON tiering is designed with such workload in mind. > > For your testing, I think the setup is good. I'd suggest lower target free > memory ratio, though. Assuming the theory (your workload has static access > pattern of small hot data) is true, using consist tuner should also be fine. > > You will have more than expected data in lower tier, but if those are truly > cold, why would we bother? If you still want strict upper tier utilization, > you could make the access pattern and filter condition less strict. Ideally, > the access pattern and filter conditions should all go away, assuming DAMON can > find true hot and cold data. And I believe that would be the case for > long-running real world workloads. Short-running test workload would show > DAMON making wrong decisions in short term, though. > > If you want to make sure if my theory is true, you could also profile the > access pattern of your test workload using DAMON. I actually did it for my > auto-tuned tiering test [4] to find the fact. > > [1] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#special-purpose-access-aware-kernel-modules > [2] https://origin.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html > [3] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh > [4] https://lkml.kernel.org/r/20250420194030.75838-1-sj@kernel.org > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-30 16:26 ` Anton Gavriliuk @ 2026-09-30 17:43 ` SJ Park 2026-10-01 9:58 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-09-30 17:43 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Wed, 30 Sep 2026 19:26:31 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > I would go back to memory tiering VCF9.1-like behavior you helped me > configure, for new few questions: > > 1. It works with the goal of 50% free memory for numa node 0, but > when I tried to set 25% free memory, it didn't do anything... > > [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat > > Per-node process memory usage (in MBs) for PID 3808 (valkey-server) > Node 0 Node 1 Node 2 > Node 3 Total > --------------- --------------- --------------- > --------------- --------------- > Huge 0.00 0.00 0.00 > 0.00 0.00 > Heap 0.11 0.00 0.00 > 0.00 0.11 > Stack 0.03 0.00 0.00 > 0.00 0.03 > Private 238356.25 2.19 3.42 > 0.00 238361.86 > ---------------- --------------- --------------- --------------- > --------------- --------------- > Total 238356.39 2.19 3.42 > 0.00 238362.00 > [root@localhost ~]# So, your workload is using ~232 GiB of node 0 memory and no demotion of it is occurred. According to your original mail [1], the node 0 has ~377 GiB memory. If we make 25% of it free, it means it should have 75% of it (~282 GiB) be utilized. If the node0 has only your workload, it means node 0 memory utilization is still only ~61%. In other words, 25% free memory goal is already achieved. Than DAMON wouldn't do any demotion. Does the theory makes sense? Maybe you can confirm by checking the free memory ratio of node 0. [...] > 2. If there are no un-accessed pages to demote for achieving a defined > goal, damon will demote pages with the most rare access ? No. Because you set '--access_rate 0% 0%' for the demotion scheme, it will do no demotion in the scenario. You could try '--access_rate 0% max' if you want. You may also need to remove '--damos_filter reject young' option from the command if that is really what you want to do. [1] https://lore.kernel.org/CAAiJnjp5F8ZuPG1gEsD_Wgs7Z+GToz67nzjr0490sz9y328=dg@mail.gmail.com Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-09-30 17:43 ` SJ Park @ 2026-10-01 9:58 ` Anton Gavriliuk 2026-10-01 10:26 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-10-01 9:58 UTC (permalink / raw) To: SJ Park; +Cc: damon > So, your workload is using ~232 GiB of node 0 memory and no demotion of it is > occurred. According to your original mail [1], the node 0 has ~377 GiB memory. > If we make 25% of it free, it means it should have 75% of it (~282 GiB) be > utilized. If the node0 has only your workload, it means node 0 memory > utilization is still only ~61%. In other words, 25% free memory goal is > already achieved. Than DAMON wouldn't do any demotion. OMG.... what was pretty easy. I feel ashamed. My apologies. I increased Valkey memory allocation to 300 GB and it works. > No. Because you set '--access_rate 0% 0%' for the demotion scheme, it will do > no demotion in the scenario. You could try '--access_rate 0% max' if you want. > You may also need to remove '--damos_filter reject young' option from the > command if that is really what you want to do. '--access_rate 0% max' means range 0% - 100% ? If yes, then I need memory management decisions based on criteria other than access frequency. Age-based filters or something else. The idea is for achieving a defined goal - keep 25% free for new hot allocations in the fastest tier, if there are no un-accessed pages, then start demoting the most-rare access or add aga-based filters. Complicated but interesting, I need to think more about that. Anton ср, 30 сент. 2026 г. в 20:43, SJ Park <sj@kernel.org>: > > On Wed, 30 Sep 2026 19:26:31 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > I would go back to memory tiering VCF9.1-like behavior you helped me > > configure, for new few questions: > > > > 1. It works with the goal of 50% free memory for numa node 0, but > > when I tried to set 25% free memory, it didn't do anything... > > > > [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat > > > > Per-node process memory usage (in MBs) for PID 3808 (valkey-server) > > Node 0 Node 1 Node 2 > > Node 3 Total > > --------------- --------------- --------------- > > --------------- --------------- > > Huge 0.00 0.00 0.00 > > 0.00 0.00 > > Heap 0.11 0.00 0.00 > > 0.00 0.11 > > Stack 0.03 0.00 0.00 > > 0.00 0.03 > > Private 238356.25 2.19 3.42 > > 0.00 238361.86 > > ---------------- --------------- --------------- --------------- > > --------------- --------------- > > Total 238356.39 2.19 3.42 > > 0.00 238362.00 > > [root@localhost ~]# > > So, your workload is using ~232 GiB of node 0 memory and no demotion of it is > occurred. According to your original mail [1], the node 0 has ~377 GiB memory. > If we make 25% of it free, it means it should have 75% of it (~282 GiB) be > utilized. If the node0 has only your workload, it means node 0 memory > utilization is still only ~61%. In other words, 25% free memory goal is > already achieved. Than DAMON wouldn't do any demotion. > > Does the theory makes sense? Maybe you can confirm by checking the free memory > ratio of node 0. > > [...] > > 2. If there are no un-accessed pages to demote for achieving a defined > > goal, damon will demote pages with the most rare access ? > > No. Because you set '--access_rate 0% 0%' for the demotion scheme, it will do > no demotion in the scenario. You could try '--access_rate 0% max' if you want. > You may also need to remove '--damos_filter reject young' option from the > command if that is really what you want to do. > > [1] https://lore.kernel.org/CAAiJnjp5F8ZuPG1gEsD_Wgs7Z+GToz67nzjr0490sz9y328=dg@mail.gmail.com > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-01 9:58 ` Anton Gavriliuk @ 2026-10-01 10:26 ` SJ Park 2026-10-01 12:08 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-10-01 10:26 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Thu, 1 Oct 2026 12:58:39 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > So, your workload is using ~232 GiB of node 0 memory and no demotion of it is > > occurred. According to your original mail [1], the node 0 has ~377 GiB memory. > > If we make 25% of it free, it means it should have 75% of it (~282 GiB) be > > utilized. If the node0 has only your workload, it means node 0 memory > > utilization is still only ~61%. In other words, 25% free memory goal is > > already achieved. Than DAMON wouldn't do any demotion. > > OMG.... what was pretty easy. > I feel ashamed. My apologies. No worry, I'm happy that I helped :) > > I increased Valkey memory allocation to 300 GB and it works. Awesome, thank you for confirming. > > > No. Because you set '--access_rate 0% 0%' for the demotion scheme, it will do > > no demotion in the scenario. You could try '--access_rate 0% max' if you want. > > You may also need to remove '--damos_filter reject young' option from the > > command if that is really what you want to do. > > '--access_rate 0% max' means range 0% - 100% ? That's correct. > If yes, then I need memory management decisions based on criteria > other than access frequency. > Age-based filters or something else. Maybe not. Under the (auto-tuned) quota, DAMOS applies the action to hottest or coldest quota amount of memory depend on the action. For example, in your current setup, let's suppose DAMON found 10 GiB of hot memory (have >=5% access rate) on node 2. And the quota is auto-tuned to 4 GiB. In the case, DAMOS will find hottest 5 GiB hot memory and migrate those to node 0. For finding the hottest memory, DAMOS calculates access temperature of each memory based on the access frequency and the age. If you set '--access_rate 0% max', DAMOS will show all memory in node 2 is eligible to migrate to node 0. But, because you have the auto-tuned quota that has 8 GiB/s upperlimit, it will migrate only up to 8 GiB hottest memory per second. Note that your goal also have '--damos_filter allow young' option for the promotion. That asks DAMOS to double check if each migration candidate page is marked as accessed since the last double check, by h/w. If it is not marked as accessed by h/w, DAMOS will not migrate it. So to allow promoting short-time unaccessed memory, you will need to turn off this double check, too. > > The idea is for achieving a defined goal - keep 25% free for new hot > allocations in the fastest tier, if there are no un-accessed pages, > then start demoting the most-rare access or add aga-based filters. > Complicated but interesting, I need to think more about that. That makes sense, and I believe removing '--access_rate' (absence of '--access_rate' is same to '--access_rate 0% max') and '--damos_filter' is one of the ways to achieve that. Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-01 10:26 ` SJ Park @ 2026-10-01 12:08 ` Anton Gavriliuk 2026-10-01 13:27 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-10-01 12:08 UTC (permalink / raw) To: SJ Park; +Cc: damon With your support, my progress tremendously faster than I expected :-) Moving forward. In my VCF9.1-like example Demote/Promote bandwidth is limited by admin up to 8GB/s. With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s. In this case, kdamond single CPU core limited or can use more CPU cores if/when required ? ~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5 DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s. Anton ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-01 12:08 ` Anton Gavriliuk @ 2026-10-01 13:27 ` SJ Park 2026-10-01 16:54 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-10-01 13:27 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Thu, 1 Oct 2026 15:08:38 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > With your support, my progress tremendously faster than I expected :-) Thank you. The pleasure is mine :) > > Moving forward. > > In my VCF9.1-like example Demote/Promote bandwidth is limited by admin > up to 8GB/s. > > With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s. > In this case, kdamond single CPU core limited or can use more CPU > cores if/when required ? > > ~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5 > DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s. It is basically limited to single CPU core per kdamond. You could split memory into multiple address ranges and assign kdamon per the range to utilize multi CPUs, if needed. SK Hynix [1] was using such an approach. That said, I'm personally curious if such fast migration is really needed in the real world workload. I assume the real production workload would have stable access pattern but only occasionally get access pattern changes. Sometimes temporal and rapid access pattern change could also be made. But for such temporal pattern change, doing migration would only be costy, since the data that suddenly hot could be soon be cold. And reliability is important in production. Hence I was thinking slowly making the balance is better than too quickly and reactively making migrations that will turn out to be no really needed. That said, if you want to test faster migration speed, the multiple kdamonds usage could be one way to test. If it turns out it is really needed and splitting address ranges has problems, we can consider adding DAMOS feature for utilizing multiple CPUs for faster migrations. [1] https://sched.co/2913n Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-01 13:27 ` SJ Park @ 2026-10-01 16:54 ` Anton Gavriliuk 2026-10-02 8:34 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-10-01 16:54 UTC (permalink / raw) To: SJ Park; +Cc: damon > That said, I'm personally curious if such fast migration is really needed in > the real world workload. I assume the real production workload would have > stable access pattern but only occasionally get access pattern changes. > Sometimes temporal and rapid access pattern change could also be made. But for > such temporal pattern change, doing migration would only be costy, since the > data that suddenly hot could be soon be cold. And reliability is important in > production. Hence I was thinking slowly making the balance is better than too > quickly and reactively making migrations that will turn out to be no really > needed. > That said, if you want to test faster migration speed, the multiple kdamonds > usage could be one way to test. If it turns out it is really needed and > splitting address ranges has problems, we can consider adding DAMOS feature for > utilizing multiple CPUs for faster migrations. I agree with you!, but now it looks that demotion is too slow. The server is completely idle, only I play with memory tiering. As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices, [root@localhost anton]# ndctl list [ { "dev":"namespace1.0", "mode":"devdax", "map":"dev", "size":3183575302144, "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f", "chardev":"dax1.0", "align":2097152 }, { "dev":"namespace0.0", "mode":"devdax", "map":"dev", "size":3183575302144, "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2", "chardev":"dax0.0", "align":2097152 } ] [root@localhost anton]# and then configured as system ram by next command, daxctl reconfigure-device --mode=system-ram all The problem is during demotion (to keep 25% free in the fast tier), kdamond.0 utilizes single CPU core ~100%, but writing to the slower tier (System PMM Write) ~1.4 GB/s what is much slower than defined limit 8 GB/s. top - 16:26:06 up 12 min, 2 users, load average: 0.81, 0.75, 0.54 Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used, 465.2 buff/cache MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6424006.+avail Mem PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 3227 root 20 0 0 0 0 R 99.8 0.0 1:28.92 kdamond.0 2771 root 20 0 6932852 10260 5608 S 0.7 0.0 0:02.95 pcm-memory 1593 root 20 0 17088 8352 7140 S 0.3 0.0 0:00.99 systemd-logind 2263 root 20 0 300.4g 294.4g 6108 S 0.3 4.4 4:42.72 valkey-server 3327 root 20 0 10916 6064 3764 R 0.3 0.0 0:00.16 top 1 root 20 0 25356 15232 10456 S 0.0 0.0 0:02.13 systemd 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd |---------------------------------------||---------------------------------------| |-- System DRAM Read Throughput(MB/s): 751.95 --| |-- System DRAM Write Throughput(MB/s): 742.58 --| |-- System PMM Read Throughput(MB/s): 693.12 --| |-- System PMM Write Throughput(MB/s): 1387.47 --| |-- System Read Throughput(MB/s): 1445.08 --| |-- System Write Throughput(MB/s): 2130.05 --| |-- System Memory Throughput(MB/s): 3575.12 --| |---------------------------------------||---------------------------------------| Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): 51312892466 lost: 0/0 drop Overhead Shared Object Symbol 47.50% [kernel] [k] migrate_folio_unmap 16.11% [kernel] [k] clear_highpages_kasan_tagged 14.93% [kernel] [k] copy_mc_fragile 1.63% [kernel] [k] smp_call_function_many_cond 1.44% [kernel] [k] page_vma_mapped_walk 0.97% [kernel] [k] folio_migrate_flags 0.87% [kernel] [k] try_to_migrate_one 0.66% [kernel] [k] rmqueue_bulk 0.61% [kernel] [k] __list_del_entry_valid_or_report 0.59% [kernel] [k] remove_migration_pte 0.50% [kernel] [k] rmap_walk_anon 0.45% [kernel] [k] __free_one_page 0.44% [kernel] [k] migrate_folio_move 0.44% [kernel] [k] folio_add_anon_rmap_ptes 0.40% [kernel] [k] lru_gen_add_folio 0.40% [kernel] [k] __mod_memcg_lruvec_state 0.39% [kernel] [k] mod_node_page_state 0.35% [kernel] [k] migrate_pages_batch 0.33% [kernel] [k] _raw_spin_lock 0.33% [kernel] [k] up_read Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): 64449579226 migrate_folio_unmap /proc/kcore [Percent: local period] Percent │ movq %r13,0x20(%rsp) 0.01 │ movq %rdx,%r13 │ movq %r14,0x28(%rsp) │ movl %r9d,%r14d 0.00 │ movq %r15,0x30(%rsp) 0.01 │ movq %r8,%r15 │ → callq *%rax 0.16 │ testq %rax,%rax │ ↓ je 2e3 0.01 │ movq %rbp,0x10(%rsp) 0.01 │ movq %rax,%rbp 0.04 │ movq %rax,(%r15) 0.00 │ movq $0x0,0x28(%rax) 99.27 │ lock │ btsq $0x0,(%rbx) 0.01 │ ↓ jb 26b 0.11 │ 67: movq (%rbx),%rdx 0.01 │ testq $0x2,(%rbx) 0.01 │ ↓ je e2 │ cmpl $0x2,%r14d │ ↓ je d2 │ movq 0x40(%rsp),%r8 │ movq %rbx,%rdi │ movl $0x1,%ecx │ xorl %edx,%edx I also tried without "--damos_filter reject young" and "--damos_filter allow young", but got the same single CPU ~100% utilization and demotion poor performance. Anton чт, 1 окт. 2026 г. в 16:27, SJ Park <sj@kernel.org>: > > On Thu, 1 Oct 2026 15:08:38 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > With your support, my progress tremendously faster than I expected :-) > > Thank you. The pleasure is mine :) > > > > > Moving forward. > > > > In my VCF9.1-like example Demote/Promote bandwidth is limited by admin > > up to 8GB/s. > > > > With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s. > > In this case, kdamond single CPU core limited or can use more CPU > > cores if/when required ? > > > > ~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5 > > DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s. > > It is basically limited to single CPU core per kdamond. You could split memory > into multiple address ranges and assign kdamon per the range to utilize multi > CPUs, if needed. SK Hynix [1] was using such an approach. > > That said, I'm personally curious if such fast migration is really needed in > the real world workload. I assume the real production workload would have > stable access pattern but only occasionally get access pattern changes. > Sometimes temporal and rapid access pattern change could also be made. But for > such temporal pattern change, doing migration would only be costy, since the > data that suddenly hot could be soon be cold. And reliability is important in > production. Hence I was thinking slowly making the balance is better than too > quickly and reactively making migrations that will turn out to be no really > needed. > > That said, if you want to test faster migration speed, the multiple kdamonds > usage could be one way to test. If it turns out it is really needed and > splitting address ranges has problems, we can consider adding DAMOS feature for > utilizing multiple CPUs for faster migrations. > > [1] https://sched.co/2913n > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-01 16:54 ` Anton Gavriliuk @ 2026-10-02 8:34 ` SJ Park 2026-10-02 10:02 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-10-02 8:34 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Thu, 1 Oct 2026 19:54:09 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > That said, I'm personally curious if such fast migration is really needed in > > the real world workload. I assume the real production workload would have > > stable access pattern but only occasionally get access pattern changes. > > Sometimes temporal and rapid access pattern change could also be made. But for > > such temporal pattern change, doing migration would only be costy, since the > > data that suddenly hot could be soon be cold. And reliability is important in > > production. Hence I was thinking slowly making the balance is better than too > > quickly and reactively making migrations that will turn out to be no really > > needed. > > > That said, if you want to test faster migration speed, the multiple kdamonds > > usage could be one way to test. If it turns out it is really needed and > > splitting address ranges has problems, we can consider adding DAMOS feature for > > utilizing multiple CPUs for faster migrations. > > > I agree with you!, but now it looks that demotion is too slow. > The server is completely idle, only I play with memory tiering. > > As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices, > > [root@localhost anton]# ndctl list > [ > { > "dev":"namespace1.0", > "mode":"devdax", > "map":"dev", > "size":3183575302144, > "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f", > "chardev":"dax1.0", > "align":2097152 > }, > { > "dev":"namespace0.0", > "mode":"devdax", > "map":"dev", > "size":3183575302144, > "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2", > "chardev":"dax0.0", > "align":2097152 > } > ] > [root@localhost anton]# > > and then configured as system ram by next command, > > daxctl reconfigure-device --mode=system-ram all > > The problem is during demotion (to keep 25% free in the fast tier), > kdamond.0 utilizes single CPU core ~100%, but writing to the slower > tier (System PMM Write) ~1.4 GB/s what is much slower than defined > limit 8 GB/s. The defined limit is only upper limit, so real speed could be slower than that. FWIW, the example memory tiering script [1] uses 200 MiB/second as the upperlimit. It was set by my gut feeling, not by some good data, though. I personally feel like ~1.4 GB/s is still a good speed for long-running production workloads, though I don't have a data to support that. Do you have some data or reason to pursue 8 GB/s ? > > > top - 16:26:06 up 12 min, 2 users, load average: 0.81, 0.75, 0.54 > Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie > %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st > MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used, 465.2 buff/cache > MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6424006.+avail Mem > > PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND > 3227 root 20 0 0 0 0 R 99.8 0.0 1:28.92 kdamond.0 > 2771 root 20 0 6932852 10260 5608 S 0.7 0.0 0:02.95 > pcm-memory > 1593 root 20 0 17088 8352 7140 S 0.3 0.0 0:00.99 > systemd-logind > 2263 root 20 0 300.4g 294.4g 6108 S 0.3 4.4 4:42.72 > valkey-server > 3327 root 20 0 10916 6064 3764 R 0.3 0.0 0:00.16 top > 1 root 20 0 25356 15232 10456 S 0.0 0.0 0:02.13 systemd > 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd > > > > |---------------------------------------||---------------------------------------| > |-- System DRAM Read Throughput(MB/s): 751.95 > --| > |-- System DRAM Write Throughput(MB/s): 742.58 > --| > |-- System PMM Read Throughput(MB/s): 693.12 > --| > |-- System PMM Write Throughput(MB/s): 1387.47 > --| > |-- System Read Throughput(MB/s): 1445.08 > --| > |-- System Write Throughput(MB/s): 2130.05 > --| > |-- System Memory Throughput(MB/s): 3575.12 > --| > |---------------------------------------||---------------------------------------| > > > > Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > 51312892466 lost: 0/0 drop > Overhead Shared Object Symbol > 47.50% [kernel] [k] migrate_folio_unmap > 16.11% [kernel] [k] clear_highpages_kasan_tagged > 14.93% [kernel] [k] copy_mc_fragile > 1.63% [kernel] [k] smp_call_function_many_cond > 1.44% [kernel] [k] page_vma_mapped_walk > 0.97% [kernel] [k] folio_migrate_flags > 0.87% [kernel] [k] try_to_migrate_one > 0.66% [kernel] [k] rmqueue_bulk > 0.61% [kernel] [k] > __list_del_entry_valid_or_report > 0.59% [kernel] [k] remove_migration_pte > 0.50% [kernel] [k] rmap_walk_anon > 0.45% [kernel] [k] __free_one_page > 0.44% [kernel] [k] migrate_folio_move > 0.44% [kernel] [k] folio_add_anon_rmap_ptes > 0.40% [kernel] [k] lru_gen_add_folio > 0.40% [kernel] [k] __mod_memcg_lruvec_state > 0.39% [kernel] [k] mod_node_page_state > 0.35% [kernel] [k] migrate_pages_batch > 0.33% [kernel] [k] _raw_spin_lock > 0.33% [kernel] [k] up_read > > > > Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > 64449579226 > migrate_folio_unmap /proc/kcore [Percent: local period] > Percent │ movq %r13,0x20(%rsp) > 0.01 │ movq %rdx,%r13 > │ movq %r14,0x28(%rsp) > │ movl %r9d,%r14d > 0.00 │ movq %r15,0x30(%rsp) > 0.01 │ movq %r8,%r15 > │ → callq *%rax > 0.16 │ testq %rax,%rax > │ ↓ je 2e3 > 0.01 │ movq %rbp,0x10(%rsp) > 0.01 │ movq %rax,%rbp > 0.04 │ movq %rax,(%r15) > 0.00 │ movq $0x0,0x28(%rax) > 99.27 │ lock > │ btsq $0x0,(%rbx) > 0.01 │ ↓ jb 26b > 0.11 │ 67: movq (%rbx),%rdx > 0.01 │ testq $0x2,(%rbx) > 0.01 │ ↓ je e2 > │ cmpl $0x2,%r14d > │ ↓ je d2 > │ movq 0x40(%rsp),%r8 > │ movq %rbx,%rdi > │ movl $0x1,%ecx > │ xorl %edx,%edx > > I also tried without "--damos_filter reject young" and "--damos_filter > allow young", but got the same single CPU ~100% utilization and > demotion poor performance. Thank you for doing these great profiling and sharing the results, Anton. Apparently the bottleneck is not in DAMON specific code including the filtering part. Instead, the bottleneck is in the migration code, which is out of DAMON, according to my understanding of the profiling result. You could also confirm this by measuring the migration speed using move_pages() like system call, which also use the migration code. You may need to optimize the migration code, or make DAMON uses multiple CPUs if faster speed is really what needed. I'd suggest using the address range split approach that I suggested in the previous reply. If it becomes clear the faster migration is really needed, we could start thinking about making DAMOS use multiple CPUs without the address range split. [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-02 8:34 ` SJ Park @ 2026-10-02 10:02 ` Anton Gavriliuk 2026-10-02 11:02 ` SJ Park 0 siblings, 1 reply; 21+ messages in thread From: Anton Gavriliuk @ 2026-10-02 10:02 UTC (permalink / raw) To: SJ Park; +Cc: damon Peak sequential write to slower tier ~12 GB/s. I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded. For Memory Tiering with OLTP-like workloads needs to be promoted/demoted, 1.4 GB/s is OK. But for bandwidth intensive workloads such as analytics, AI, VDI,... 1.4 GB/s is too slow. On the server where I test Memory Tiering, Intel CPU 8280L installed. Despite being already 7 years old, any single CPU core has 13-14 GB/s bandwidth access to local memory. So from my point of view - even if we increase the number of cores keeping in mind 1.4 GB/s per core, we need so many cores to perform 10's GB/s demote/promote. Firstly we need to improve demote/promote bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s. > You could also confirm > this by measuring the migration speed using move_pages() like system call, > which also use the migration code. I'm not a developer and I need to spend more time thinking about how to do that, but I tried, I used migratepages command, which is probably use migrate_pages() instead of move_pages(), but I got the same poor bandwidth and stack as with DAMON/DAMOS Memory Tiering, top - 12:43:59 up 13 min, 2 users, load average: 0.65, 0.88, 0.76 Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used, 669.8 buff/cache MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6423508.+avail Mem PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 2935 root 20 0 2644 1764 1644 R 99.7 0.0 0:42.19 migratepages 2941 root 20 0 10884 6192 3916 R 0.7 0.0 0:00.04 top 1238 systemd+ 20 0 15860 6996 5952 S 0.3 0.0 0:00.37 systemd-oomd 1928 root 20 0 666268 24448 23440 S 0.3 0.0 0:00.34 rsyslogd 1 root 20 0 25840 15532 10492 S 0.0 0.0 0:02.15 systemd 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd 3 root 20 0 0 0 0 S 0.0 0.0 0:00.00 pool_workqueue_release 4 root 0 -20 0 0 0 I 0.0 0.0 0:00.00 kworker/R-rcu_gp |---------------------------------------||---------------------------------------| |-- System DRAM Read Throughput(MB/s): 759.41 --| |-- System DRAM Write Throughput(MB/s): 751.36 --| |-- System PMM Read Throughput(MB/s): 701.24 --| |-- System PMM Write Throughput(MB/s): 1402.15 --| |-- System Read Throughput(MB/s): 1460.65 --| |-- System Write Throughput(MB/s): 2153.50 --| |-- System Memory Throughput(MB/s): 3614.15 --| |---------------------------------------||---------------------------------------| Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): 48280162400 lost: 0/0 drop: Overhead Shared Object Symbol 49.82% [kernel] [k] migrate_folio_unmap 16.52% [kernel] [k] clear_highpages_kasan_tagged 13.56% [kernel] [k] copy_mc_fragile 1.80% [kernel] [k] smp_call_function_many_cond 1.06% [kernel] [k] page_vma_mapped_walk 0.95% [kernel] [k] try_to_migrate_one 0.94% [kernel] [k] folio_migrate_flags 0.71% [kernel] [k] rmqueue_bulk 0.61% [kernel] [k] remove_migration_pte 0.54% [kernel] [k] __free_one_page 0.50% [kernel] [k] migrate_folio_move 0.48% [kernel] [k] folio_batch_move_lru 0.47% [kernel] [k] __list_del_entry_valid_or_report Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): 33608407521 migrate_folio_unmap /proc/kcore [Percent: local period] Percent │ movq %r13,0x20(%rsp) 0.02 │ movq %rdx,%r13 │ movq %r14,0x28(%rsp) │ movl %r9d,%r14d │ movq %r15,0x30(%rsp) 0.01 │ movq %r8,%r15 │ → callq *%rax 0.04 │ testq %rax,%rax │ ↓ je 2e3 0.01 │ movq %rbp,0x10(%rsp) │ movq %rax,%rbp 0.03 │ movq %rax,(%r15) │ movq $0x0,0x28(%rax) 99.52 │ lock │ btsq $0x0,(%rbx) 0.01 │ ↓ jb 26b 0.13 │ 67: movq (%rbx),%rdx 0.01 │ testq $0x2,(%rbx) 0.01 │ ↓ je e2 │ cmpl $0x2,%r14d │ ↓ je d2 │ movq 0x40(%rsp),%r8 │ movq %rbx,%rdi │ movl $0x1,%ecx │ xorl %edx,%edx Lastly I used migratepages between two fast tiers (numa nodes 0 & 1) and got 1.5 GB/s. All this is too slow for the coming CXL era. Anyway, I'm ready to proceed in assisting you in checking/testing if you are also interested. Anton пт, 2 окт. 2026 г. в 11:34, SJ Park <sj@kernel.org>: > > On Thu, 1 Oct 2026 19:54:09 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > > That said, I'm personally curious if such fast migration is really needed in > > > the real world workload. I assume the real production workload would have > > > stable access pattern but only occasionally get access pattern changes. > > > Sometimes temporal and rapid access pattern change could also be made. But for > > > such temporal pattern change, doing migration would only be costy, since the > > > data that suddenly hot could be soon be cold. And reliability is important in > > > production. Hence I was thinking slowly making the balance is better than too > > > quickly and reactively making migrations that will turn out to be no really > > > needed. > > > > > That said, if you want to test faster migration speed, the multiple kdamonds > > > usage could be one way to test. If it turns out it is really needed and > > > splitting address ranges has problems, we can consider adding DAMOS feature for > > > utilizing multiple CPUs for faster migrations. > > > > > > I agree with you!, but now it looks that demotion is too slow. > > The server is completely idle, only I play with memory tiering. > > > > As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices, > > > > [root@localhost anton]# ndctl list > > [ > > { > > "dev":"namespace1.0", > > "mode":"devdax", > > "map":"dev", > > "size":3183575302144, > > "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f", > > "chardev":"dax1.0", > > "align":2097152 > > }, > > { > > "dev":"namespace0.0", > > "mode":"devdax", > > "map":"dev", > > "size":3183575302144, > > "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2", > > "chardev":"dax0.0", > > "align":2097152 > > } > > ] > > [root@localhost anton]# > > > > and then configured as system ram by next command, > > > > daxctl reconfigure-device --mode=system-ram all > > > > The problem is during demotion (to keep 25% free in the fast tier), > > kdamond.0 utilizes single CPU core ~100%, but writing to the slower > > tier (System PMM Write) ~1.4 GB/s what is much slower than defined > > limit 8 GB/s. > > The defined limit is only upper limit, so real speed could be slower than that. > FWIW, the example memory tiering script [1] uses 200 MiB/second as the > upperlimit. It was set by my gut feeling, not by some good data, though. > > I personally feel like ~1.4 GB/s is still a good speed for long-running > production workloads, though I don't have a data to support that. Do you have > some data or reason to pursue 8 GB/s ? > > > > > > > top - 16:26:06 up 12 min, 2 users, load average: 0.81, 0.75, 0.54 > > Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie > > %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st > > MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used, 465.2 buff/cache > > MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6424006.+avail Mem > > > > PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND > > 3227 root 20 0 0 0 0 R 99.8 0.0 1:28.92 kdamond.0 > > 2771 root 20 0 6932852 10260 5608 S 0.7 0.0 0:02.95 > > pcm-memory > > 1593 root 20 0 17088 8352 7140 S 0.3 0.0 0:00.99 > > systemd-logind > > 2263 root 20 0 300.4g 294.4g 6108 S 0.3 4.4 4:42.72 > > valkey-server > > 3327 root 20 0 10916 6064 3764 R 0.3 0.0 0:00.16 top > > 1 root 20 0 25356 15232 10456 S 0.0 0.0 0:02.13 systemd > > 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd > > > > > > > > |---------------------------------------||---------------------------------------| > > |-- System DRAM Read Throughput(MB/s): 751.95 > > --| > > |-- System DRAM Write Throughput(MB/s): 742.58 > > --| > > |-- System PMM Read Throughput(MB/s): 693.12 > > --| > > |-- System PMM Write Throughput(MB/s): 1387.47 > > --| > > |-- System Read Throughput(MB/s): 1445.08 > > --| > > |-- System Write Throughput(MB/s): 2130.05 > > --| > > |-- System Memory Throughput(MB/s): 3575.12 > > --| > > |---------------------------------------||---------------------------------------| > > > > > > > > Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > > 51312892466 lost: 0/0 drop > > Overhead Shared Object Symbol > > 47.50% [kernel] [k] migrate_folio_unmap > > 16.11% [kernel] [k] clear_highpages_kasan_tagged > > 14.93% [kernel] [k] copy_mc_fragile > > 1.63% [kernel] [k] smp_call_function_many_cond > > 1.44% [kernel] [k] page_vma_mapped_walk > > 0.97% [kernel] [k] folio_migrate_flags > > 0.87% [kernel] [k] try_to_migrate_one > > 0.66% [kernel] [k] rmqueue_bulk > > 0.61% [kernel] [k] > > __list_del_entry_valid_or_report > > 0.59% [kernel] [k] remove_migration_pte > > 0.50% [kernel] [k] rmap_walk_anon > > 0.45% [kernel] [k] __free_one_page > > 0.44% [kernel] [k] migrate_folio_move > > 0.44% [kernel] [k] folio_add_anon_rmap_ptes > > 0.40% [kernel] [k] lru_gen_add_folio > > 0.40% [kernel] [k] __mod_memcg_lruvec_state > > 0.39% [kernel] [k] mod_node_page_state > > 0.35% [kernel] [k] migrate_pages_batch > > 0.33% [kernel] [k] _raw_spin_lock > > 0.33% [kernel] [k] up_read > > > > > > > > Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > > 64449579226 > > migrate_folio_unmap /proc/kcore [Percent: local period] > > Percent │ movq %r13,0x20(%rsp) > > 0.01 │ movq %rdx,%r13 > > │ movq %r14,0x28(%rsp) > > │ movl %r9d,%r14d > > 0.00 │ movq %r15,0x30(%rsp) > > 0.01 │ movq %r8,%r15 > > │ → callq *%rax > > 0.16 │ testq %rax,%rax > > │ ↓ je 2e3 > > 0.01 │ movq %rbp,0x10(%rsp) > > 0.01 │ movq %rax,%rbp > > 0.04 │ movq %rax,(%r15) > > 0.00 │ movq $0x0,0x28(%rax) > > 99.27 │ lock > > │ btsq $0x0,(%rbx) > > 0.01 │ ↓ jb 26b > > 0.11 │ 67: movq (%rbx),%rdx > > 0.01 │ testq $0x2,(%rbx) > > 0.01 │ ↓ je e2 > > │ cmpl $0x2,%r14d > > │ ↓ je d2 > > │ movq 0x40(%rsp),%r8 > > │ movq %rbx,%rdi > > │ movl $0x1,%ecx > > │ xorl %edx,%edx > > > > I also tried without "--damos_filter reject young" and "--damos_filter > > allow young", but got the same single CPU ~100% utilization and > > demotion poor performance. > > Thank you for doing these great profiling and sharing the results, Anton. > Apparently the bottleneck is not in DAMON specific code including the filtering > part. Instead, the bottleneck is in the migration code, which is out of DAMON, > according to my understanding of the profiling result. You could also confirm > this by measuring the migration speed using move_pages() like system call, > which also use the migration code. > > You may need to optimize the migration code, or make DAMON uses multiple CPUs > if faster speed is really what needed. I'd suggest using the address range > split approach that I suggested in the previous reply. If it becomes clear the > faster migration is really needed, we could start thinking about making DAMOS > use multiple CPUs without the address range split. > > [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-02 10:02 ` Anton Gavriliuk @ 2026-10-02 11:02 ` SJ Park 2026-10-02 13:06 ` Anton Gavriliuk 0 siblings, 1 reply; 21+ messages in thread From: SJ Park @ 2026-10-02 11:02 UTC (permalink / raw) To: Anton Gavriliuk; +Cc: SJ Park, damon On Fri, 2 Oct 2026 13:02:13 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > Peak sequential write to slower tier ~12 GB/s. > I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded. > > For Memory Tiering with OLTP-like workloads needs to be > promoted/demoted, 1.4 GB/s is OK. > But for bandwidth intensive workloads such as analytics, AI, VDI,... > 1.4 GB/s is too slow. I think the factor deciding how fast migration should be is not bandwidth intensiveness but how quickly access pattern changes. That is, migration speed should be fast enough to get costs from migration itself be paid back, and get additional benefits. That depends on the speed of the workloads' access pattern change, and the amount of data of the changed pattern. For example, let's suppose 10 GiB data of a workload becomes hot, but will be cold again after 10 seconds. If we can migrate it to upper tier in one second, it will get benefit from continued access for remaining 9 seconds. If it takes 10 seconds to migrate, the benefit from the migration will be much lower. It might even lower than the migration work cost. > > On the server where I test Memory Tiering, Intel CPU 8280L installed. > Despite being already 7 years old, any single CPU core has 13-14 GB/s > bandwidth access to local memory. That maese sense to me. Kernel level page migration requires not only the content writes. It also need to do additional works. It should allocate pages in destination node that the content will be copied to. It should update mappings and related metadata. It should also handle possible races. Hence page migration is much more expensive and slow than pure I/O. > > So from my point of view - even if we increase the number of cores > keeping in mind 1.4 GB/s per core, we need so many cores to perform > 10's GB/s demote/promote. I agree. > Firstly we need to improve demote/promote > bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s. I agree there could be workloads that could get benefit from faster migration. And IIRC, there were a few people working on making migration faster and lightweight. I'm not an expert in the domain and not involved to such works for now, though. Nonetheless, having concrete data showing why 9-10 GB/s is the right speed would be nice. > > > You could also confirm > > this by measuring the migration speed using move_pages() like system call, > > which also use the migration code. > > I'm not a developer and I need to spend more time thinking about how > to do that, but I tried, > > I used migratepages command, which is probably use migrate_pages() > instead of move_pages(), but I got the same poor bandwidth and stack > as with DAMON/DAMOS Memory Tiering, > > top - 12:43:59 up 13 min, 2 users, load average: 0.65, 0.88, 0.76 > Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie > %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st > MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used, 669.8 buff/cache > MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6423508.+avail Mem > > PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND > 2935 root 20 0 2644 1764 1644 R 99.7 0.0 0:42.19 > migratepages > 2941 root 20 0 10884 6192 3916 R 0.7 0.0 0:00.04 top > 1238 systemd+ 20 0 15860 6996 5952 S 0.3 0.0 0:00.37 > systemd-oomd > 1928 root 20 0 666268 24448 23440 S 0.3 0.0 0:00.34 rsyslogd > 1 root 20 0 25840 15532 10492 S 0.0 0.0 0:02.15 systemd > 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd > 3 root 20 0 0 0 0 S 0.0 0.0 0:00.00 > pool_workqueue_release > 4 root 0 -20 0 0 0 I 0.0 0.0 0:00.00 > kworker/R-rcu_gp > > > |---------------------------------------||---------------------------------------| > |-- System DRAM Read Throughput(MB/s): 759.41 > --| > |-- System DRAM Write Throughput(MB/s): 751.36 > --| > |-- System PMM Read Throughput(MB/s): 701.24 > --| > |-- System PMM Write Throughput(MB/s): 1402.15 > --| > |-- System Read Throughput(MB/s): 1460.65 > --| > |-- System Write Throughput(MB/s): 2153.50 > --| > |-- System Memory Throughput(MB/s): 3614.15 > --| > |---------------------------------------||---------------------------------------| > > > Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > 48280162400 lost: 0/0 drop: > Overhead Shared Object Symbol > 49.82% [kernel] [k] migrate_folio_unmap > 16.52% [kernel] [k] clear_highpages_kasan_tagged > 13.56% [kernel] [k] copy_mc_fragile > 1.80% [kernel] [k] smp_call_function_many_cond > 1.06% [kernel] [k] page_vma_mapped_walk > 0.95% [kernel] [k] try_to_migrate_one > 0.94% [kernel] [k] folio_migrate_flags > 0.71% [kernel] [k] rmqueue_bulk > 0.61% [kernel] [k] remove_migration_pte > 0.54% [kernel] [k] __free_one_page > 0.50% [kernel] [k] migrate_folio_move > 0.48% [kernel] [k] folio_batch_move_lru > 0.47% [kernel] [k] > __list_del_entry_valid_or_report > > > Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > 33608407521 > migrate_folio_unmap /proc/kcore [Percent: local period] > Percent │ movq %r13,0x20(%rsp) > 0.02 │ movq %rdx,%r13 > │ movq %r14,0x28(%rsp) > │ movl %r9d,%r14d > │ movq %r15,0x30(%rsp) > 0.01 │ movq %r8,%r15 > │ → callq *%rax > 0.04 │ testq %rax,%rax > │ ↓ je 2e3 > 0.01 │ movq %rbp,0x10(%rsp) > │ movq %rax,%rbp > 0.03 │ movq %rax,(%r15) > │ movq $0x0,0x28(%rax) > 99.52 │ lock > │ btsq $0x0,(%rbx) > 0.01 │ ↓ jb 26b > 0.13 │ 67: movq (%rbx),%rdx > 0.01 │ testq $0x2,(%rbx) > 0.01 │ ↓ je e2 > │ cmpl $0x2,%r14d > │ ↓ je d2 > │ movq 0x40(%rsp),%r8 > │ movq %rbx,%rdi > │ movl $0x1,%ecx > │ xorl %edx,%edx > > Lastly I used migratepages between two fast tiers (numa nodes 0 & 1) > and got 1.5 GB/s. > All this is too slow for the coming CXL era. Thank you for testing this and sharing the results. This confirms my previous theory is not wrong. > > Anyway, I'm ready to proceed in assisting you in checking/testing if > you are also interested. I think my thoery (bottleneck is in migration, not DAMON) is already proven. We also confirmed DAMON and DAMOS are working as expected so far. So I find no more thing to test for now. If you have more ideas to test, I will be more than happy to help. Thanks, SJ [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: Memory tiering with DAMON/DAMOS auto-tuning 2026-10-02 11:02 ` SJ Park @ 2026-10-02 13:06 ` Anton Gavriliuk 0 siblings, 0 replies; 21+ messages in thread From: Anton Gavriliuk @ 2026-10-02 13:06 UTC (permalink / raw) To: SJ Park; +Cc: damon > For example, > let's suppose 10 GiB data of a workload becomes hot, but will be cold again > after 10 seconds. If we can migrate it to upper tier in one second, it will > get benefit from continued access for remaining 9 seconds. If it takes 10 > seconds to migrate, the benefit from the migration will be much lower. It > might even lower than the migration work cost. Exactly right. It fully matches the requirement to improve the current 1.4 GB/s migration bandwidth. So based on our discussions, I will try to discuss that with Memory Management subsystem maintainers. Anton пт, 2 окт. 2026 г. в 14:02, SJ Park <sj@kernel.org>: > > On Fri, 2 Oct 2026 13:02:13 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote: > > > Peak sequential write to slower tier ~12 GB/s. > > I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded. > > > > For Memory Tiering with OLTP-like workloads needs to be > > promoted/demoted, 1.4 GB/s is OK. > > But for bandwidth intensive workloads such as analytics, AI, VDI,... > > 1.4 GB/s is too slow. > > I think the factor deciding how fast migration should be is not bandwidth > intensiveness but how quickly access pattern changes. That is, migration speed > should be fast enough to get costs from migration itself be paid back, and get > additional benefits. That depends on the speed of the workloads' access > pattern change, and the amount of data of the changed pattern. For example, > let's suppose 10 GiB data of a workload becomes hot, but will be cold again > after 10 seconds. If we can migrate it to upper tier in one second, it will > get benefit from continued access for remaining 9 seconds. If it takes 10 > seconds to migrate, the benefit from the migration will be much lower. It > might even lower than the migration work cost. > > > > > On the server where I test Memory Tiering, Intel CPU 8280L installed. > > Despite being already 7 years old, any single CPU core has 13-14 GB/s > > bandwidth access to local memory. > > That maese sense to me. Kernel level page migration requires not only the > content writes. It also need to do additional works. It should allocate pages > in destination node that the content will be copied to. It should update > mappings and related metadata. It should also handle possible races. Hence > page migration is much more expensive and slow than pure I/O. > > > > > So from my point of view - even if we increase the number of cores > > keeping in mind 1.4 GB/s per core, we need so many cores to perform > > 10's GB/s demote/promote. > > I agree. > > > Firstly we need to improve demote/promote > > bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s. > > I agree there could be workloads that could get benefit from faster migration. > And IIRC, there were a few people working on making migration faster and > lightweight. I'm not an expert in the domain and not involved to such works > for now, though. Nonetheless, having concrete data showing why 9-10 GB/s is > the right speed would be nice. > > > > > > You could also confirm > > > this by measuring the migration speed using move_pages() like system call, > > > which also use the migration code. > > > > I'm not a developer and I need to spend more time thinking about how > > to do that, but I tried, > > > > I used migratepages command, which is probably use migrate_pages() > > instead of move_pages(), but I got the same poor bandwidth and stack > > as with DAMON/DAMOS Memory Tiering, > > > > top - 12:43:59 up 13 min, 2 users, load average: 0.65, 0.88, 0.76 > > Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie > > %Cpu(s): 0.0 us, 1.8 sy, 0.0 ni, 98.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st > > MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used, 669.8 buff/cache > > MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 6423508.+avail Mem > > > > PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND > > 2935 root 20 0 2644 1764 1644 R 99.7 0.0 0:42.19 > > migratepages > > 2941 root 20 0 10884 6192 3916 R 0.7 0.0 0:00.04 top > > 1238 systemd+ 20 0 15860 6996 5952 S 0.3 0.0 0:00.37 > > systemd-oomd > > 1928 root 20 0 666268 24448 23440 S 0.3 0.0 0:00.34 rsyslogd > > 1 root 20 0 25840 15532 10492 S 0.0 0.0 0:02.15 systemd > > 2 root 20 0 0 0 0 S 0.0 0.0 0:00.02 kthreadd > > 3 root 20 0 0 0 0 S 0.0 0.0 0:00.00 > > pool_workqueue_release > > 4 root 0 -20 0 0 0 I 0.0 0.0 0:00.00 > > kworker/R-rcu_gp > > > > > > |---------------------------------------||---------------------------------------| > > |-- System DRAM Read Throughput(MB/s): 759.41 > > --| > > |-- System DRAM Write Throughput(MB/s): 751.36 > > --| > > |-- System PMM Read Throughput(MB/s): 701.24 > > --| > > |-- System PMM Write Throughput(MB/s): 1402.15 > > --| > > |-- System Read Throughput(MB/s): 1460.65 > > --| > > |-- System Write Throughput(MB/s): 2153.50 > > --| > > |-- System Memory Throughput(MB/s): 3614.15 > > --| > > |---------------------------------------||---------------------------------------| > > > > > > Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > > 48280162400 lost: 0/0 drop: > > Overhead Shared Object Symbol > > 49.82% [kernel] [k] migrate_folio_unmap > > 16.52% [kernel] [k] clear_highpages_kasan_tagged > > 13.56% [kernel] [k] copy_mc_fragile > > 1.80% [kernel] [k] smp_call_function_many_cond > > 1.06% [kernel] [k] page_vma_mapped_walk > > 0.95% [kernel] [k] try_to_migrate_one > > 0.94% [kernel] [k] folio_migrate_flags > > 0.71% [kernel] [k] rmqueue_bulk > > 0.61% [kernel] [k] remove_migration_pte > > 0.54% [kernel] [k] __free_one_page > > 0.50% [kernel] [k] migrate_folio_move > > 0.48% [kernel] [k] folio_batch_move_lru > > 0.47% [kernel] [k] > > __list_del_entry_valid_or_report > > > > > > Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.): > > 33608407521 > > migrate_folio_unmap /proc/kcore [Percent: local period] > > Percent │ movq %r13,0x20(%rsp) > > 0.02 │ movq %rdx,%r13 > > │ movq %r14,0x28(%rsp) > > │ movl %r9d,%r14d > > │ movq %r15,0x30(%rsp) > > 0.01 │ movq %r8,%r15 > > │ → callq *%rax > > 0.04 │ testq %rax,%rax > > │ ↓ je 2e3 > > 0.01 │ movq %rbp,0x10(%rsp) > > │ movq %rax,%rbp > > 0.03 │ movq %rax,(%r15) > > │ movq $0x0,0x28(%rax) > > 99.52 │ lock > > │ btsq $0x0,(%rbx) > > 0.01 │ ↓ jb 26b > > 0.13 │ 67: movq (%rbx),%rdx > > 0.01 │ testq $0x2,(%rbx) > > 0.01 │ ↓ je e2 > > │ cmpl $0x2,%r14d > > │ ↓ je d2 > > │ movq 0x40(%rsp),%r8 > > │ movq %rbx,%rdi > > │ movl $0x1,%ecx > > │ xorl %edx,%edx > > > > Lastly I used migratepages between two fast tiers (numa nodes 0 & 1) > > and got 1.5 GB/s. > > All this is too slow for the coming CXL era. > > Thank you for testing this and sharing the results. This confirms my previous > theory is not wrong. > > > > > Anyway, I'm ready to proceed in assisting you in checking/testing if > > you are also interested. > > I think my thoery (bottleneck is in migration, not DAMON) is already proven. > We also confirmed DAMON and DAMOS are working as expected so far. So I find no > more thing to test for now. If you have more ideas to test, I will be more > than happy to help. > > > Thanks, > SJ > > [...] ^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-10-02 13:06 UTC | newest] Thread overview: 21+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk 2026-09-28 8:15 ` SJ Park 2026-09-28 17:02 ` Anton Gavriliuk 2026-09-29 7:49 ` SJ Park 2026-09-29 8:41 ` Anton Gavriliuk 2026-09-29 9:25 ` SJ Park 2026-09-29 13:54 ` Anton Gavriliuk 2026-09-29 17:37 ` SJ Park 2026-09-30 4:03 ` Anton Gavriliuk 2026-09-30 8:35 ` SJ Park 2026-09-30 16:26 ` Anton Gavriliuk 2026-09-30 17:43 ` SJ Park 2026-10-01 9:58 ` Anton Gavriliuk 2026-10-01 10:26 ` SJ Park 2026-10-01 12:08 ` Anton Gavriliuk 2026-10-01 13:27 ` SJ Park 2026-10-01 16:54 ` Anton Gavriliuk 2026-10-02 8:34 ` SJ Park 2026-10-02 10:02 ` Anton Gavriliuk 2026-10-02 11:02 ` SJ Park 2026-10-02 13:06 ` Anton Gavriliuk
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox