DAMON development mailing list
 help / color / mirror / Atom feed
* Memory tiering with DAMON/DAMOS auto-tuning
@ 2026-09-27 17:26 Anton Gavriliuk
  2026-09-28  8:15 ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-27 17:26 UTC (permalink / raw)
  To: sj, damon

Hello

I would like to play with memory tiering with DAMON/DAMOS auto-tuning
(between numa nodes 0 & 2) based on valkey and memtier_benchmark.
The goal - keep numa node 0 50% free and promote cold pages from numa
node 2 to numa node 0 immediately when they become hot.
This is Fedora Server 44 up-to-date with the 7.2.8 kernel.

[root@localhost ~]# numactl -H
available: 4 nodes (0-3)
node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
22 23 24 25 26 27
node 0 size: 386590 MB
node 0 free: 337165 MB
node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
47 48 49 50 51 52 53 54 55
node 1 size: 387055 MB
node 1 free: 338204 MB
node 2 cpus:
node 2 size: 3033088 MB
node 2 free: 3032998 MB
node 3 cpus:
node 3 size: 3033088 MB
node 3 free: 3032998 MB
node distances:
node     0    1    2    3
   0:   10   20   25   35
   1:   20   10   35   25
   2:   25   35   10   20
   3:   35   25   20   10
[root@localhost ~]#
[root@localhost ~]# /home/anton/Linux/mlc
Intel(R) Memory Latency Checker - v3.13
Measuring idle latencies for sequential access (in ns)...
                Numa node
Numa node            0       1       2       3
       0          82.3   147.6   173.8   238.0
       1         146.7    82.0   238.7   172.2


The Goal:

1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2
2. Keep fast tier 50% free
3. Demote valkey inactive >=5 min pages from fast to slow tier
4. Promote valkey hot pages from slow tier to fast tier immediately
when accessed
5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU
performance unlimited
6. Monitoring intervals for Demote/Promote 500ms

What I already done -

Launch Valkey Pinned to Node 0 DRAM

numactl --cpunodebind=0 --preferred=0 valkey-server \
  --port 6379 \
  --protected-mode no \
  --save "" \
  --appendonly no \
  --maxmemory 300gb \
  --maxmemory-policy noeviction &


Command to Load ~200+ GB

memtier_benchmark \
  -s 127.0.0.1 -p 6379 \
  -t 16 -c 16 \
  -d 10240 \
  --ratio=1:0 \
  --key-pattern=P:P \
  --distinct-client-seed \
  --key-maximum=20000000 \
  --pipeline=32 \
  -n allkeys


[root@localhost anton]# numastat -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                         0.11            0.00            0.00
Stack                        0.03            0.00            0.00
Private                 238356.61            0.46            4.82
----------------  --------------- --------------- ---------------
Total                   238356.75            0.46            4.82

                           Node 3           Total
                  --------------- ---------------
Huge                         0.00            0.00
Heap                         0.00            0.11
Stack                        0.00            0.03
Private                      0.00       238361.90
----------------  --------------- ---------------
Total                        0.00       238362.04


damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s
\--damos_action migrate_cold 2 --damos_access_rate 0% 0%
\--damos_apply_interval 1s \--damos_quota_interval 1s
--damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
\--damos_filter reject young \--numa_node 2
--monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot
0 --damos_access_rate 5% max \--damos_apply_interval 1s
\--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
node_mem_used_bp 99.7% 0 \--damos_filter allow young
\--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
--nr_schemes 1 1 --nr_ctxs 1 1

And it is demoted much more, almost all pages than the goal of 50%
free numa node 0.

[root@localhost anton]# numastat -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                         0.02            0.00            0.09
Stack                        0.02            0.00            0.01
Private                  24389.88            0.46       213971.55
----------------  --------------- --------------- ---------------
Total                    24389.92            0.46       213971.65

                           Node 3           Total
                  --------------- ---------------
Huge                         0.00            0.00
Heap                         0.00            0.11
Stack                        0.00            0.03
Private                      0.00       238361.90
----------------  --------------- ---------------
Total                        0.00       238362.04

I just want that demotion process to stop when there is 50% free at numa node 0.

Where am I wrong, how to fix that ?

Anton

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk
@ 2026-09-28  8:15 ` SJ Park
  2026-09-28 17:02   ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-28  8:15 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

Hello Anton,

On Sun, 27 Sep 2026 20:26:21 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> Hello
> 
> I would like to play with memory tiering with DAMON/DAMOS auto-tuning
> (between numa nodes 0 & 2) based on valkey and memtier_benchmark.

Thank you for sharing your use case and question!

> The goal - keep numa node 0 50% free and promote cold pages from numa
> node 2 to numa node 0 immediately when they become hot.
> This is Fedora Server 44 up-to-date with the 7.2.8 kernel.
> 
> [root@localhost ~]# numactl -H
> available: 4 nodes (0-3)
> node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
> 22 23 24 25 26 27
> node 0 size: 386590 MB
> node 0 free: 337165 MB
> node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
> 47 48 49 50 51 52 53 54 55
> node 1 size: 387055 MB
> node 1 free: 338204 MB
> node 2 cpus:
> node 2 size: 3033088 MB
> node 2 free: 3032998 MB
> node 3 cpus:
> node 3 size: 3033088 MB
> node 3 free: 3032998 MB
> node distances:
> node     0    1    2    3
>    0:   10   20   25   35
>    1:   20   10   35   25
>    2:   25   35   10   20
>    3:   35   25   20   10
> [root@localhost ~]#
> [root@localhost ~]# /home/anton/Linux/mlc
> Intel(R) Memory Latency Checker - v3.13
> Measuring idle latencies for sequential access (in ns)...
>                 Numa node
> Numa node            0       1       2       3
>        0          82.3   147.6   173.8   238.0
>        1         146.7    82.0   238.7   172.2
> 
> 
> The Goal:
> 
> 1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2
> 2. Keep fast tier 50% free
> 3. Demote valkey inactive >=5 min pages from fast to slow tier
> 4. Promote valkey hot pages from slow tier to fast tier immediately
> when accessed
> 5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU
> performance unlimited
> 6. Monitoring intervals for Demote/Promote 500ms
> 
> What I already done -
> 
> Launch Valkey Pinned to Node 0 DRAM
> 
> numactl --cpunodebind=0 --preferred=0 valkey-server \
>   --port 6379 \
>   --protected-mode no \
>   --save "" \
>   --appendonly no \
>   --maxmemory 300gb \
>   --maxmemory-policy noeviction &
> 
> 
> Command to Load ~200+ GB
> 
> memtier_benchmark \
>   -s 127.0.0.1 -p 6379 \
>   -t 16 -c 16 \
>   -d 10240 \
>   --ratio=1:0 \
>   --key-pattern=P:P \
>   --distinct-client-seed \
>   --key-maximum=20000000 \
>   --pipeline=32 \
>   -n allkeys
> 
> 
> [root@localhost anton]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.11            0.00            0.00
> Stack                        0.03            0.00            0.00
> Private                 238356.61            0.46            4.82
> ----------------  --------------- --------------- ---------------
> Total                   238356.75            0.46            4.82
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       238361.90
> ----------------  --------------- ---------------
> Total                        0.00       238362.04
> 
> 
> damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s

The example memory tiering script [1] uses 4% as intervals goal.  Is there a
reason to use 97% as the goal instead?

> \--damos_action migrate_cold 2 --damos_access_rate 0% 0%
> \--damos_apply_interval 1s \--damos_quota_interval 1s
> --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
> \--damos_filter reject young \

You mentioned you want to demote >=5 minutes inactive pages.  But, the above
command doesn't have '--damos_age' option.  It means DAMON will demote node 0
pages as soon as it finds it was not accessed, even if it was not accessed for
<5 minutes.  You could let DAMON know you want to demote only >=5 minutes
inactive pages by adding '--damos_age 5m max' option.

> --numa_node 2
> --monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot

Again, I'm curious why you use 97% goal.

> 0 --damos_access_rate 5% max \--damos_apply_interval 1s
> \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
> node_mem_used_bp 99.7% 0 \--damos_filter allow young

Is the above node_mum_used_bp what you really want?  That means you want to
promote hot pages from node 2 to node 0, until the node 0 memory utilization
becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
circulation of hot/cold pages while not consuming too much resource under
stabilized access pattern.

> \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
> --nr_schemes 1 1 --nr_ctxs 1 1
> 
> And it is demoted much more, almost all pages than the goal of 50%
> free numa node 0.
> 
> [root@localhost anton]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.02            0.00            0.09
> Stack                        0.02            0.00            0.01
> Private                  24389.88            0.46       213971.55
> ----------------  --------------- --------------- ---------------
> Total                    24389.92            0.46       213971.65
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       238361.90
> ----------------  --------------- ---------------
> Total                        0.00       238362.04
> 
> I just want that demotion process to stop when there is 50% free at numa node 0.
> 
> Where am I wrong, how to fix that ?

DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
should keep applying the action in a level to keep the goal achieved.  In this
case, for example, the demotion scheme assumes there will be continued
promotion and therefore it keeps demoting.  If there is no appropriate
promotion, it could result in demoting more than expected amount.

To me, this seems like the workload has no much hot data, so demotion is much
more stronger.

For path forward, I'd suggest trying 'temporal' tuner [2].  It is designed to
immediately stop after achieving the goal.

If it still doesn't work, I'd suggest tracing damon:damos_esz tracepoint with
the numastat output.  It will show if the autotune is working as expected.

For long term production use case, I'd like to suggest running the example
tiering config [1] and see if it also gives you unexpected results.

[1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
[2] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#aim-oriented-feedback-driven-auto-tuning


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-28  8:15 ` SJ Park
@ 2026-09-28 17:02   ` Anton Gavriliuk
  2026-09-29  7:49     ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-28 17:02 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

> Thank you for sharing your use case and question!

Thank you for your quick response and your answer/comments.

We already had memory tiering implemented at hardware level 5-7 years
ago with Intel's Optane PEM when configured in Memory Mode.
Now we do the same thing at Linux level :-)
I'm new at DAMON/DAMOS, there are lots of new tunables to me.

> The example memory tiering script [1] uses 4% as intervals goal.  Is there a
> reason to use 97% as the goal instead?

Oohhh... I moved to 5%.

> Is the above node_mum_used_bp what you really want?  That means you want to
> promote hot pages from node 2 to node 0, until the node 0 memory utilization
> becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
> 0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
> circulation of hot/cold pages while not consuming too much resource under
> stabilized access pattern.

> DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
> temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
> should keep applying the action in a level to keep the goal achieved.  In this
> case, for example, the demotion scheme assumes there will be continued
> promotion and therefore it keeps demoting.  If there is no appropriate
> promotion, it could result in demoting more than expected amount.

So let me firstly understand what I really want to test :-)

Anton

пн, 28 сент. 2026 г. в 11:15, SJ Park <sj@kernel.org>:
>
> Hello Anton,
>
> On Sun, 27 Sep 2026 20:26:21 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > Hello
> >
> > I would like to play with memory tiering with DAMON/DAMOS auto-tuning
> > (between numa nodes 0 & 2) based on valkey and memtier_benchmark.
>
> Thank you for sharing your use case and question!
>
> > The goal - keep numa node 0 50% free and promote cold pages from numa
> > node 2 to numa node 0 immediately when they become hot.
> > This is Fedora Server 44 up-to-date with the 7.2.8 kernel.
> >
> > [root@localhost ~]# numactl -H
> > available: 4 nodes (0-3)
> > node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
> > 22 23 24 25 26 27
> > node 0 size: 386590 MB
> > node 0 free: 337165 MB
> > node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
> > 47 48 49 50 51 52 53 54 55
> > node 1 size: 387055 MB
> > node 1 free: 338204 MB
> > node 2 cpus:
> > node 2 size: 3033088 MB
> > node 2 free: 3032998 MB
> > node 3 cpus:
> > node 3 size: 3033088 MB
> > node 3 free: 3032998 MB
> > node distances:
> > node     0    1    2    3
> >    0:   10   20   25   35
> >    1:   20   10   35   25
> >    2:   25   35   10   20
> >    3:   35   25   20   10
> > [root@localhost ~]#
> > [root@localhost ~]# /home/anton/Linux/mlc
> > Intel(R) Memory Latency Checker - v3.13
> > Measuring idle latencies for sequential access (in ns)...
> >                 Numa node
> > Numa node            0       1       2       3
> >        0          82.3   147.6   173.8   238.0
> >        1         146.7    82.0   238.7   172.2
> >
> >
> > The Goal:
> >
> > 1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2
> > 2. Keep fast tier 50% free
> > 3. Demote valkey inactive >=5 min pages from fast to slow tier
> > 4. Promote valkey hot pages from slow tier to fast tier immediately
> > when accessed
> > 5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU
> > performance unlimited
> > 6. Monitoring intervals for Demote/Promote 500ms
> >
> > What I already done -
> >
> > Launch Valkey Pinned to Node 0 DRAM
> >
> > numactl --cpunodebind=0 --preferred=0 valkey-server \
> >   --port 6379 \
> >   --protected-mode no \
> >   --save "" \
> >   --appendonly no \
> >   --maxmemory 300gb \
> >   --maxmemory-policy noeviction &
> >
> >
> > Command to Load ~200+ GB
> >
> > memtier_benchmark \
> >   -s 127.0.0.1 -p 6379 \
> >   -t 16 -c 16 \
> >   -d 10240 \
> >   --ratio=1:0 \
> >   --key-pattern=P:P \
> >   --distinct-client-seed \
> >   --key-maximum=20000000 \
> >   --pipeline=32 \
> >   -n allkeys
> >
> >
> > [root@localhost anton]# numastat -p $(pgrep valkey-server)
> >
> > Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >                   --------------- --------------- ---------------
> > Huge                         0.00            0.00            0.00
> > Heap                         0.11            0.00            0.00
> > Stack                        0.03            0.00            0.00
> > Private                 238356.61            0.46            4.82
> > ----------------  --------------- --------------- ---------------
> > Total                   238356.75            0.46            4.82
> >
> >                            Node 3           Total
> >                   --------------- ---------------
> > Huge                         0.00            0.00
> > Heap                         0.00            0.11
> > Stack                        0.00            0.03
> > Private                      0.00       238361.90
> > ----------------  --------------- ---------------
> > Total                        0.00       238362.04
> >
> >
> > damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s
>
> The example memory tiering script [1] uses 4% as intervals goal.  Is there a
> reason to use 97% as the goal instead?
>
> > \--damos_action migrate_cold 2 --damos_access_rate 0% 0%
> > \--damos_apply_interval 1s \--damos_quota_interval 1s
> > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
> > \--damos_filter reject young \
>
> You mentioned you want to demote >=5 minutes inactive pages.  But, the above
> command doesn't have '--damos_age' option.  It means DAMON will demote node 0
> pages as soon as it finds it was not accessed, even if it was not accessed for
> <5 minutes.  You could let DAMON know you want to demote only >=5 minutes
> inactive pages by adding '--damos_age 5m max' option.
>
> > --numa_node 2
> > --monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot
>
> Again, I'm curious why you use 97% goal.
>
> > 0 --damos_access_rate 5% max \--damos_apply_interval 1s
> > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
> > node_mem_used_bp 99.7% 0 \--damos_filter allow young
>
> Is the above node_mum_used_bp what you really want?  That means you want to
> promote hot pages from node 2 to node 0, until the node 0 memory utilization
> becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
> 0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
> circulation of hot/cold pages while not consuming too much resource under
> stabilized access pattern.
>
> > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
> > --nr_schemes 1 1 --nr_ctxs 1 1
> >
> > And it is demoted much more, almost all pages than the goal of 50%
> > free numa node 0.
> >
> > [root@localhost anton]# numastat -p $(pgrep valkey-server)
> >
> > Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >                   --------------- --------------- ---------------
> > Huge                         0.00            0.00            0.00
> > Heap                         0.02            0.00            0.09
> > Stack                        0.02            0.00            0.01
> > Private                  24389.88            0.46       213971.55
> > ----------------  --------------- --------------- ---------------
> > Total                    24389.92            0.46       213971.65
> >
> >                            Node 3           Total
> >                   --------------- ---------------
> > Huge                         0.00            0.00
> > Heap                         0.00            0.11
> > Stack                        0.00            0.03
> > Private                      0.00       238361.90
> > ----------------  --------------- ---------------
> > Total                        0.00       238362.04
> >
> > I just want that demotion process to stop when there is 50% free at numa node 0.
> >
> > Where am I wrong, how to fix that ?
>
> DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
> temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
> should keep applying the action in a level to keep the goal achieved.  In this
> case, for example, the demotion scheme assumes there will be continued
> promotion and therefore it keeps demoting.  If there is no appropriate
> promotion, it could result in demoting more than expected amount.
>
> To me, this seems like the workload has no much hot data, so demotion is much
> more stronger.
>
> For path forward, I'd suggest trying 'temporal' tuner [2].  It is designed to
> immediately stop after achieving the goal.
>
> If it still doesn't work, I'd suggest tracing damon:damos_esz tracepoint with
> the numastat output.  It will show if the autotune is working as expected.
>
> For long term production use case, I'd like to suggest running the example
> tiering config [1] and see if it also gives you unexpected results.
>
> [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
> [2] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#aim-oriented-feedback-driven-auto-tuning
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-28 17:02   ` Anton Gavriliuk
@ 2026-09-29  7:49     ` SJ Park
  2026-09-29  8:41       ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-29  7:49 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Mon, 28 Sep 2026 20:02:43 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> > Thank you for sharing your use case and question!
> 
> Thank you for your quick response and your answer/comments.
> 
> We already had memory tiering implemented at hardware level 5-7 years
> ago with Intel's Optane PEM when configured in Memory Mode.
> Now we do the same thing at Linux level :-)

Interesting!

> I'm new at DAMON/DAMOS, there are lots of new tunables to me.

Indeed there are.  I'd recommend starting with existing DAMON-based memory
tiering solutions including SK Hynix HMSDK and my auto-tune based memory
tiering script [1] and ask questions to the authors.

> 
> > The example memory tiering script [1] uses 4% as intervals goal.  Is there a
> > reason to use 97% as the goal instead?
> 
> Oohhh... I moved to 5%.

I understand you mean 4%?

> 
> > Is the above node_mum_used_bp what you really want?  That means you want to
> > promote hot pages from node 2 to node 0, until the node 0 memory utilization
> > becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
> > 0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
> > circulation of hot/cold pages while not consuming too much resource under
> > stabilized access pattern.
> 
> > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
> > temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
> > should keep applying the action in a level to keep the goal achieved.  In this
> > case, for example, the demotion scheme assumes there will be continued
> > promotion and therefore it keeps demoting.  If there is no appropriate
> > promotion, it could result in demoting more than expected amount.
> 
> So let me firstly understand what I really want to test :-)

Sure, and please feel free to ask any question in any ways.  If you prefer to,
you could also ask questions privately to me or use DAMON Beer/Coffee/Tea Chat
series [2].

[1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
[2] https://docs.google.com/document/d/1v43Kcj3ly4CYqmAkMaZzLiM2GEnWfgdGbZAH3mi2vpM/edit?usp=sharing


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-29  7:49     ` SJ Park
@ 2026-09-29  8:41       ` Anton Gavriliuk
  2026-09-29  9:25         ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-29  8:41 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

Hello

I decided to move step-by-step, so firstly I would get on Linux/bare
metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1
- https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency.

"Threshold Activation: The VMkernel initiates memory tiering only when
total host physical DRAM (Tier 0) consumption reaches approximately
the 80% threshold. If consumption is below this trigger (e.g., at
70-75%), the system will not migrate pages to Tier 1.
Page Classification: ESXi monitors page activity based on recency and
frequency. Only memory pages strictly classified as "cold" (inactive)
are migrated to the NVMe Tier 1 device. The active working set ("hot"
pages) is deliberately retained in Tier 0 to prevent performance
penalties.
Scan Rate: The page scan and migration rate scale with memory
pressure. At lower consumption levels (just crossing the 80% mark),
the rate is highly conservative."

Here is my corrected command,

damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s
\--damos_action migrate_cold 2 --damos_access_rate 0% 0%
\--damos_apply_interval 1s \--damos_quota_interval 1s
--damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
\--damos_filter reject young \--numa_node 2
--monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0
--damos_access_rate 5% max \--damos_apply_interval 1s
\--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
node_mem_used_bp 50.3% 0 \--damos_filter allow young
\--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
--nr_schemes 1 1 --nr_ctxs 1 1

with consist -> temporal,

[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner
temporal
[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner
temporal

But it again didn't stop ~50%

[root@localhost ~]# numastat -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                         0.11            0.00            0.00
Stack                        0.03            0.00            0.00
Private                 228807.54            0.48          185.79
----------------  --------------- --------------- ---------------
Total                   228807.67            0.48          185.79

                           Node 3           Total
                  --------------- ---------------
Huge                         0.00            0.00
Heap                         0.00            0.11
Stack                        0.00            0.03
Private                      0.00       228993.80
----------------  --------------- ---------------
Total                        0.00       228993.94


[root@localhost ~]# numastat -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                         0.01            0.00            0.10
Stack                        0.01            0.00            0.02
Private                  13434.01            0.48       215559.31
----------------  --------------- --------------- ---------------
Total                    13434.03            0.48       215559.43

                           Node 3           Total
                  --------------- ---------------
Huge                         0.00            0.00
Heap                         0.00            0.11
Stack                        0.00            0.03
Private                      0.00       228993.80
----------------  --------------- ---------------
Total                        0.00       228993.94



> > Oohhh... I moved to 5%.
> I understand you mean 4%?

As far as I understood, this value represents "accuracy", so there
shouldn't be a big difference between 4% vs 5%.  Or please correct me
if I'm wrong.

Anton


вт, 29 сент. 2026 г. в 10:49, SJ Park <sj@kernel.org>:
>
> On Mon, 28 Sep 2026 20:02:43 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > > Thank you for sharing your use case and question!
> >
> > Thank you for your quick response and your answer/comments.
> >
> > We already had memory tiering implemented at hardware level 5-7 years
> > ago with Intel's Optane PEM when configured in Memory Mode.
> > Now we do the same thing at Linux level :-)
>
> Interesting!
>
> > I'm new at DAMON/DAMOS, there are lots of new tunables to me.
>
> Indeed there are.  I'd recommend starting with existing DAMON-based memory
> tiering solutions including SK Hynix HMSDK and my auto-tune based memory
> tiering script [1] and ask questions to the authors.
>
> >
> > > The example memory tiering script [1] uses 4% as intervals goal.  Is there a
> > > reason to use 97% as the goal instead?
> >
> > Oohhh... I moved to 5%.
>
> I understand you mean 4%?
>
> >
> > > Is the above node_mum_used_bp what you really want?  That means you want to
> > > promote hot pages from node 2 to node 0, until the node 0 memory utilization
> > > becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
> > > 0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
> > > circulation of hot/cold pages while not consuming too much resource under
> > > stabilized access pattern.
> >
> > > DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
> > > temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
> > > should keep applying the action in a level to keep the goal achieved.  In this
> > > case, for example, the demotion scheme assumes there will be continued
> > > promotion and therefore it keeps demoting.  If there is no appropriate
> > > promotion, it could result in demoting more than expected amount.
> >
> > So let me firstly understand what I really want to test :-)
>
> Sure, and please feel free to ask any question in any ways.  If you prefer to,
> you could also ask questions privately to me or use DAMON Beer/Coffee/Tea Chat
> series [2].
>
> [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
> [2] https://docs.google.com/document/d/1v43Kcj3ly4CYqmAkMaZzLiM2GEnWfgdGbZAH3mi2vpM/edit?usp=sharing
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-29  8:41       ` Anton Gavriliuk
@ 2026-09-29  9:25         ` SJ Park
  2026-09-29 13:54           ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-29  9:25 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Tue, 29 Sep 2026 11:41:15 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> Hello
> 
> I decided to move step-by-step, so firstly I would get on Linux/bare
> metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1
> - https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency.
> 
> "Threshold Activation: The VMkernel initiates memory tiering only when
> total host physical DRAM (Tier 0) consumption reaches approximately
> the 80% threshold. If consumption is below this trigger (e.g., at
> 70-75%), the system will not migrate pages to Tier 1.
> Page Classification: ESXi monitors page activity based on recency and
> frequency. Only memory pages strictly classified as "cold" (inactive)
> are migrated to the NVMe Tier 1 device. The active working set ("hot"
> pages) is deliberately retained in Tier 0 to prevent performance
> penalties.
> Scan Rate: The page scan and migration rate scale with memory
> pressure. At lower consumption levels (just crossing the 80% mark),
> the rate is highly conservative."

Sounds good.

> 
> Here is my corrected command,
> 
> damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s

Fyi, '--monitoring_intervals_autotune' is same to
'--monitoring_intervals_goal 4% 3 5ms 10s'.

> \--damos_action migrate_cold 2 --damos_access_rate 0% 0%
> \--damos_apply_interval 1s \--damos_quota_interval 1s
> --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
> \--damos_filter reject young \--numa_node 2
> --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0
> --damos_access_rate 5% max \--damos_apply_interval 1s
> \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
> node_mem_used_bp 50.3% 0 \--damos_filter allow young
> \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
> --nr_schemes 1 1 --nr_ctxs 1 1

Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context,
target and scheme parameters on the command line using --kdamond, --damon_ctx,
--damon_target, and --damos_scheme.  you coud use those instead of
--damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and
--nr_ctxs.  E.g.,

damo start \
        ` # A kdamond to demote cold memory from node 0 to node 2 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 0 \
                --damos_scheme \
                        --damos_action migrate_cold 2 \
                        --damos_access_rate 0% 0% --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50% node 0 free memory ` \
                        --damos_quota_goal node_mem_free_bp 50% 0 \
                        --damos_filter reject young \
        ` # A kdamond to promote hot memory from node 2 to node 0 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 2 \
                --damos_scheme \
                        --damos_action migrate_hot 0 \
                        --damos_access_rate 5% max --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50.3% node 0 memory utilization ` \
                        --damos_quota_goal node_mem_used_bp 50.3% 0 \
                        --damos_filter allow young \

> 
> with consist -> temporal,
> 
> [root@localhost ~]# cat
> /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner
> temporal
> [root@localhost ~]# cat
> /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner
> temporal

Seems you manually made this change.  Has this executed after 'damo start'?
Also, did you 'commit' the updated commit input?  If any of your answer to the
questions is not "yes", the tuner update may not applied.

You can specify what tuner to use on 'damo' command together, using
'--damos_quota_goal_tuner' option.  I'd recommend using that.  E.g.,

damo start \
        ` # A kdamond to demote cold memory from node 0 to node 2 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 0 \
                --damos_scheme \
                        --damos_action migrate_cold 2 \
                        --damos_access_rate 0% 0% --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50% node 0 free memory ` \
                        --damos_quota_goal node_mem_free_bp 50% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter reject young \
        ` # A kdamond to promote hot memory from node 2 to node 0 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 2 \
                --damos_scheme \
                        --damos_action migrate_hot 0 \
                        --damos_access_rate 5% max --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50.3% node 0 memory utilization ` \
                        --damos_quota_goal node_mem_used_bp 50.3% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter allow young \

> 
> But it again didn't stop ~50%
> 
> [root@localhost ~]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.11            0.00            0.00
> Stack                        0.03            0.00            0.00
> Private                 228807.54            0.48          185.79
> ----------------  --------------- --------------- ---------------
> Total                   228807.67            0.48          185.79
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       228993.80
> ----------------  --------------- ---------------
> Total                        0.00       228993.94
> 
> 
> [root@localhost ~]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.01            0.00            0.10
> Stack                        0.01            0.00            0.02
> Private                  13434.01            0.48       215559.31
> ----------------  --------------- --------------- ---------------
> Total                    13434.03            0.48       215559.43
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       228993.80
> ----------------  --------------- ---------------
> Total                        0.00       228993.94

Interesting.  I doubt if the tuner change is correctly made.  'damo report
damon' can help us understand under what configuration DAMON is running.  It
could help us quickly see if the configuration is done as we intended.  Could
you share 'damo report damon' result on the final state?

It would also be helpful if you could run 'damo report damon --damos_stats'
periodically (say, once per 5-10 seconds) while the migration is ongoing and
share the outputs with us.

> 
> 
> 
> > > Oohhh... I moved to 5%.
> > I understand you mean 4%?
> 
> As far as I understood, this value represents "accuracy", so there
> shouldn't be a big difference between 4% vs 5%.  Or please correct me
> if I'm wrong.

Yes, 5% vs 4% should be an ignorable small difference.


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-29  9:25         ` SJ Park
@ 2026-09-29 13:54           ` Anton Gavriliuk
  2026-09-29 17:37             ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-29 13:54 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

> Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context,
> target and scheme parameters on the command line using --kdamond, --damon_ctx,
> --damon_target, and --damos_scheme.  you coud use those instead of
> --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and
> --nr_ctxs.  E.g.,

Fedora shows old damo version,

[root@localhost anton]# rpm -qa|grep -i damo
damo-3.3.0-1.fc44.noarch
[root@localhost anton]#

so I configured latest available 3.4.1

[root@localhost damo]# damo version
3.4.1
[root@localhost damo]# which damo
/home/anton/damo/damo
[root@localhost damo]#

> Seems you manually made this change.  Has this executed after 'damo start'?
> Also, did you 'commit' the updated commit input?  If any of your answer to the
> questions is not "yes", the tuner update may not applied.

Yes, it was executed before 'damo start', but I didn't do 'commit'.

> It would also be helpful if you could run 'damo report damon --damos_stats'
> periodically (say, once per 5-10 seconds) while the migration is ongoing and
> share the outputs with us.

Ok, initially I have,

[root@localhost ~]# numastat -z -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 25361 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Heap                         0.11            0.00            0.00
Stack                        0.03            0.00            0.00
Private                 238356.60            0.50            4.86
----------------  --------------- --------------- ---------------
Total                   238356.73            0.50            4.86

                            Total
                  ---------------
Heap                         0.11
Stack                        0.03
Private                 238361.95
----------------  ---------------
Total                   238362.09
[root@localhost ~]#

I run

[root@localhost ~]# /home/anton/damo/damo start \
        ` # A kdamond to demote cold memory from node 0 to node 2 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 0 \
                --damos_scheme \
                        --damos_action migrate_cold 2 \
                        --damos_access_rate 0% 0% --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50% node 0 free memory ` \
                        --damos_quota_goal node_mem_free_bp 50% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter reject young \
        ` # A kdamond to promote hot memory from node 2 to node 0 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 2 \
                --damos_scheme \
                        --damos_action migrate_hot 0 \
                        --damos_access_rate 5% max --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50.3% node 0 memory utilization ` \
                        --damos_quota_goal node_mem_used_bp 50.3% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter allow young \
>
sysinfo loading fail (info update fail (sysfs feature check fail
(feature map making fail (staging damos goal feature check purpose
kdamond failed))))
[root@localhost ~]#
[root@localhost ~]# ps -ef|grep -i damo
root       26051   25396  0 16:49 pts/2    00:00:00 grep --color=auto -i damo
[root@localhost ~]#

> It would also be helpful if you could run 'damo report damon --damos_stats'
> periodically (say, once per 5-10 seconds) while the migration is ongoing and
> share the outputs with us.

Sure, I will start something like that,
while true; do damo report damon --damos_stats >>
/home/anton/damo_report_damon_damos_stats; sleep 10; done
once damo will be started.


> Interesting.  I doubt if the tuner change is correctly made.  'damo report
> damon' can help us understand under what configuration DAMON is running.  It
> could help us quickly see if the configuration is done as we intended.  Could
> you share 'damo report damon' result on the final state?

There are errors above during damo start.

Anton

вт, 29 сент. 2026 г. в 12:25, SJ Park <sj@kernel.org>:
>
> On Tue, 29 Sep 2026 11:41:15 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > Hello
> >
> > I decided to move step-by-step, so firstly I would get on Linux/bare
> > metal or non-VmWare setup memory tiering behaviour like on VmWare 9.1
> > - https://knowledge.broadcom.com/external/article/449016/understanding-nvme-memory-tiering-activa.html#:~:text=Cause.%20This%20behavior%20is%20by%20design.%20The,page%20activity%20based%20on%20recency%20and%20frequency.
> >
> > "Threshold Activation: The VMkernel initiates memory tiering only when
> > total host physical DRAM (Tier 0) consumption reaches approximately
> > the 80% threshold. If consumption is below this trigger (e.g., at
> > 70-75%), the system will not migrate pages to Tier 1.
> > Page Classification: ESXi monitors page activity based on recency and
> > frequency. Only memory pages strictly classified as "cold" (inactive)
> > are migrated to the NVMe Tier 1 device. The active working set ("hot"
> > pages) is deliberately retained in Tier 0 to prevent performance
> > penalties.
> > Scan Rate: The page scan and migration rate scale with memory
> > pressure. At lower consumption levels (just crossing the 80% mark),
> > the rate is highly conservative."
>
> Sounds good.
>
> >
> > Here is my corrected command,
> >
> > damo start \--numa_node 0 --monitoring_intervals_goal 5% 3 5ms 10s
>
> Fyi, '--monitoring_intervals_autotune' is same to
> '--monitoring_intervals_goal 4% 3 5ms 10s'.
>
> > \--damos_action migrate_cold 2 --damos_access_rate 0% 0%
> > \--damos_apply_interval 1s \--damos_quota_interval 1s
> > --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
> > \--damos_filter reject young \--numa_node 2
> > --monitoring_intervals_goal 5% 3 5ms 10s \--damos_action migrate_hot 0
> > --damos_access_rate 5% max \--damos_apply_interval 1s
> > \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
> > node_mem_used_bp 50.3% 0 \--damos_filter allow young
> > \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
> > --nr_schemes 1 1 --nr_ctxs 1 1
>
> Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context,
> target and scheme parameters on the command line using --kdamond, --damon_ctx,
> --damon_target, and --damos_scheme.  you coud use those instead of
> --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and
> --nr_ctxs.  E.g.,
>
> damo start \
>         ` # A kdamond to demote cold memory from node 0 to node 2 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 0 \
>                 --damos_scheme \
>                         --damos_action migrate_cold 2 \
>                         --damos_access_rate 0% 0% --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50% node 0 free memory ` \
>                         --damos_quota_goal node_mem_free_bp 50% 0 \
>                         --damos_filter reject young \
>         ` # A kdamond to promote hot memory from node 2 to node 0 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 2 \
>                 --damos_scheme \
>                         --damos_action migrate_hot 0 \
>                         --damos_access_rate 5% max --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50.3% node 0 memory utilization ` \
>                         --damos_quota_goal node_mem_used_bp 50.3% 0 \
>                         --damos_filter allow young \
>
> >
> > with consist -> temporal,
> >
> > [root@localhost ~]# cat
> > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner
> > temporal
> > [root@localhost ~]# cat
> > /sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner
> > temporal
>
> Seems you manually made this change.  Has this executed after 'damo start'?
> Also, did you 'commit' the updated commit input?  If any of your answer to the
> questions is not "yes", the tuner update may not applied.
>
> You can specify what tuner to use on 'damo' command together, using
> '--damos_quota_goal_tuner' option.  I'd recommend using that.  E.g.,
>
> damo start \
>         ` # A kdamond to demote cold memory from node 0 to node 2 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 0 \
>                 --damos_scheme \
>                         --damos_action migrate_cold 2 \
>                         --damos_access_rate 0% 0% --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50% node 0 free memory ` \
>                         --damos_quota_goal node_mem_free_bp 50% 0 \
>                         ` # using temporal tuner ` \
>                         --damos_quota_goal_tuner temporal \
>                         --damos_filter reject young \
>         ` # A kdamond to promote hot memory from node 2 to node 0 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 2 \
>                 --damos_scheme \
>                         --damos_action migrate_hot 0 \
>                         --damos_access_rate 5% max --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50.3% node 0 memory utilization ` \
>                         --damos_quota_goal node_mem_used_bp 50.3% 0 \
>                         ` # using temporal tuner ` \
>                         --damos_quota_goal_tuner temporal \
>                         --damos_filter allow young \
>
> >
> > But it again didn't stop ~50%
> >
> > [root@localhost ~]# numastat -p $(pgrep valkey-server)
> >
> > Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >                   --------------- --------------- ---------------
> > Huge                         0.00            0.00            0.00
> > Heap                         0.11            0.00            0.00
> > Stack                        0.03            0.00            0.00
> > Private                 228807.54            0.48          185.79
> > ----------------  --------------- --------------- ---------------
> > Total                   228807.67            0.48          185.79
> >
> >                            Node 3           Total
> >                   --------------- ---------------
> > Huge                         0.00            0.00
> > Heap                         0.00            0.11
> > Stack                        0.00            0.03
> > Private                      0.00       228993.80
> > ----------------  --------------- ---------------
> > Total                        0.00       228993.94
> >
> >
> > [root@localhost ~]# numastat -p $(pgrep valkey-server)
> >
> > Per-node process memory usage (in MBs) for PID 18036 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >                   --------------- --------------- ---------------
> > Huge                         0.00            0.00            0.00
> > Heap                         0.01            0.00            0.10
> > Stack                        0.01            0.00            0.02
> > Private                  13434.01            0.48       215559.31
> > ----------------  --------------- --------------- ---------------
> > Total                    13434.03            0.48       215559.43
> >
> >                            Node 3           Total
> >                   --------------- ---------------
> > Huge                         0.00            0.00
> > Heap                         0.00            0.11
> > Stack                        0.00            0.03
> > Private                      0.00       228993.80
> > ----------------  --------------- ---------------
> > Total                        0.00       228993.94
>
> Interesting.  I doubt if the tuner change is correctly made.  'damo report
> damon' can help us understand under what configuration DAMON is running.  It
> could help us quickly see if the configuration is done as we intended.  Could
> you share 'damo report damon' result on the final state?
>
> It would also be helpful if you could run 'damo report damon --damos_stats'
> periodically (say, once per 5-10 seconds) while the migration is ongoing and
> share the outputs with us.
>
> >
> >
> >
> > > > Oohhh... I moved to 5%.
> > > I understand you mean 4%?
> >
> > As far as I understood, this value represents "accuracy", so there
> > shouldn't be a big difference between 4% vs 5%.  Or please correct me
> > if I'm wrong.
>
> Yes, 5% vs 4% should be an ignorable small difference.
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-29 13:54           ` Anton Gavriliuk
@ 2026-09-29 17:37             ` SJ Park
  2026-09-30  4:03               ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-29 17:37 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Tue, 29 Sep 2026 16:54:41 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context,
> > target and scheme parameters on the command line using --kdamond, --damon_ctx,
> > --damon_target, and --damos_scheme.  you coud use those instead of
> > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and
> > --nr_ctxs.  E.g.,
> 
> Fedora shows old damo version,
> 
> [root@localhost anton]# rpm -qa|grep -i damo
> damo-3.3.0-1.fc44.noarch
> [root@localhost anton]#
> 
> so I configured latest available 3.4.1
> 
> [root@localhost damo]# damo version
> 3.4.1
> [root@localhost damo]# which damo
> /home/anton/damo/damo
> [root@localhost damo]#

Thank you!  Hopefully upgrading the version was not that difficult.  You can
simply git-clone the repo and use the 'damo' executable file under the
local-cloned repo.

> 
> > Seems you manually made this change.  Has this executed after 'damo start'?
> > Also, did you 'commit' the updated commit input?  If any of your answer to the
> > questions is not "yes", the tuner update may not applied.
> 
> Yes, it was executed before 'damo start', but I didn't do 'commit'.

To do this manually, you should write the files after 'damo start', and also do
'commit'.

Anyway, this means your previous run was using 'consist' tuner.  That explains
why it didn't show any difference.

> 
> > It would also be helpful if you could run 'damo report damon --damos_stats'
> > periodically (say, once per 5-10 seconds) while the migration is ongoing and
> > share the outputs with us.
> 
> Ok, initially I have,
> 
> [root@localhost ~]# numastat -z -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 25361 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Heap                         0.11            0.00            0.00
> Stack                        0.03            0.00            0.00
> Private                 238356.60            0.50            4.86
> ----------------  --------------- --------------- ---------------
> Total                   238356.73            0.50            4.86
> 
>                             Total
>                   ---------------
> Heap                         0.11
> Stack                        0.03
> Private                 238361.95
> ----------------  ---------------
> Total                   238362.09
> [root@localhost ~]#
> 
> I run
> 
> [root@localhost ~]# /home/anton/damo/damo start \
>         ` # A kdamond to demote cold memory from node 0 to node 2 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 0 \
>                 --damos_scheme \
>                         --damos_action migrate_cold 2 \
>                         --damos_access_rate 0% 0% --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50% node 0 free memory ` \
>                         --damos_quota_goal node_mem_free_bp 50% 0 \
>                         ` # using temporal tuner ` \
>                         --damos_quota_goal_tuner temporal \
>                         --damos_filter reject young \
>         ` # A kdamond to promote hot memory from node 2 to node 0 ` \
>         --kdamond --damon_ctx --monitoring_intervals_autotune \
>                 --damon_target --numa_node 2 \
>                 --damos_scheme \
>                         --damos_action migrate_hot 0 \
>                         --damos_access_rate 5% max --damos_apply_interval 1s \
>                         ` # up to 8 GiB per second ` \
>                         --damos_quota_interval 1s --damos_quota_space 8G \
>                         ` # aiming at least 50.3% node 0 memory utilization ` \
>                         --damos_quota_goal node_mem_used_bp 50.3% 0 \
>                         ` # using temporal tuner ` \
>                         --damos_quota_goal_tuner temporal \
>                         --damos_filter allow young \
> >
> sysinfo loading fail (info update fail (sysfs feature check fail
> (feature map making fail (staging damos goal feature check purpose
> kdamond failed))))

Oops...  Seems you were running next branch of damo.  There was a bug.  I
reproduced it on 7.2.8 kernel, and fixed it.  The fix [1] is now pushed.  Could
you pull the 'next' branch and try again?

> [root@localhost ~]#
> [root@localhost ~]# ps -ef|grep -i damo
> root       26051   25396  0 16:49 pts/2    00:00:00 grep --color=auto -i damo
> [root@localhost ~]#
> 
> > It would also be helpful if you could run 'damo report damon --damos_stats'
> > periodically (say, once per 5-10 seconds) while the migration is ongoing and
> > share the outputs with us.
> 
> Sure, I will start something like that,
> while true; do damo report damon --damos_stats >>
> /home/anton/damo_report_damon_damos_stats; sleep 10; done
> once damo will be started.

Sounds good.

> 
> 
> > Interesting.  I doubt if the tuner change is correctly made.  'damo report
> > damon' can help us understand under what configuration DAMON is running.  It
> > could help us quickly see if the configuration is done as we intended.  Could
> > you share 'damo report damon' result on the final state?
> 
> There are errors above during damo start.

Apparently it was a bug in damo's next branch.  As I mentioned above, the fix
is now pushed.  Could you try again?

[1] https://github.com/damonitor/damo/commit/fd2fb55b44ba03db55b8a025eef43dcf569831e5

Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-29 17:37             ` SJ Park
@ 2026-09-30  4:03               ` Anton Gavriliuk
  2026-09-30  8:35                 ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-30  4:03 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

[-- Attachment #1: Type: text/plain, Size: 15270 bytes --]

> Oops...  Seems you were running next branch of damo.  There was a bug.  I
> reproduced it on 7.2.8 kernel, and fixed it.  The fix [1] is now pushed.  Could
> you pull the 'next' branch and try again?

Done.

Ok, initially I have,

[root@localhost ~]# numastat -z -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 25361 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Heap                         0.11            0.00            0.00
Stack                        0.03            0.00            0.00
Private                 238356.60            0.50            4.86
----------------  --------------- --------------- ---------------
Total                   238356.73            0.50            4.86

                            Total
                  ---------------
Heap                         0.11
Stack                        0.03
Private                 238361.95
----------------  ---------------
Total                   238362.09
[root@localhost ~]#

I run

/home/anton/damo/damo start \
        ` # A kdamond to demote cold memory from node 0 to node 2 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 0 \
                --damos_scheme \
                        --damos_action migrate_cold 2 \
                        --damos_access_rate 0% 0% --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50% node 0 free memory ` \
                        --damos_quota_goal node_mem_free_bp 50% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter reject young \
        ` # A kdamond to promote hot memory from node 2 to node 0 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 2 \
                --damos_scheme \
                        --damos_action migrate_hot 0 \
                        --damos_access_rate 5% max --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 50.3% node 0 memory utilization ` \
                        --damos_quota_goal node_mem_used_bp 50.3% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter allow young \
>
>
[root@localhost ~]# ps -ef|grep -i dam
root       29836       2  2 06:37 ?        00:00:00 [kdamond.0]
root       29837       2  0 06:37 ?        00:00:00 [kdamond.1]
root       29842   29644  0 06:37 pts/2    00:00:00 grep --color=auto -i dam
[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner
temporal
[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner
temporal
[root@localhost ~]#


> It would also be helpful if you could run 'damo report damon --damos_stats'
> periodically (say, once per 5-10 seconds) while the migration is ongoing and
> share the outputs with us.

Sure, I will start something like that,
while true; do /home/anton/damo/damo report damon --damos_stats >>
/home/anton/damo_report_damon_damos_stats; sleep 10; done
once damo will be started.
Please check the attached file.
It looks now it keeps ~50% free numa node 0,

[root@localhost ~]# numastat -z -p $(pgrep valkey-server)

Per-node process memory usage (in MBs) for PID 25361 (valkey-server)
                           Node 0          Node 1          Node 2
                  --------------- --------------- ---------------
Heap                         0.05            0.00            0.06
Stack                        0.01            0.00            0.02
Private                 138725.66            0.50        99635.75
----------------  --------------- --------------- ---------------
Total                   138725.71            0.50        99635.82

                            Total
                  ---------------
Heap                         0.11
Stack                        0.03
Private                 238361.90
----------------  ---------------
Total                   238362.04
[root@localhost ~]#



> Interesting.  I doubt if the tuner change is correctly made.  'damo report
> damon' can help us understand under what configuration DAMON is running.  It
> could help us quickly see if the configuration is done as we intended.  Could
> you share 'damo report damon' result on the final state?

[root@localhost ~]# /home/anton/damo/damo report damon
kdamond 0
    state: on, pid: 29836
    context 0
        ops: paddr
        target 0
            pid: 0
            region [4,096, 581,632) (564.000 KiB)
            region [589,824, 655,360) (64.000 KiB)
            region [1,048,576, 2,079,305,728) (1.936 GiB)
            region [2,079,363,072, 2,153,725,952) (70.918 MiB)
            region [2,153,926,656, 2,154,790,912) (844.000 KiB)
            region [2,154,811,392, 2,154,909,696) (96.000 KiB)
            region [2,154,930,176, 2,354,724,864) (190.539 MiB)
            region [2,355,859,456, 2,355,871,744) (12.000 KiB)
            region [2,372,751,360, 2,372,755,456) (4.000 KiB)
            region [2,372,759,552, 2,372,771,840) (12.000 KiB)
            region [2,391,846,912, 2,551,230,464) (152.000 MiB)
            region [2,604,707,840, 2,604,716,032) (8.000 KiB)
            region [2,605,244,416, 2,944,397,312) (323.441 MiB)
            region [2,944,397,372, 2,944,401,408) (3.941 KiB)
            region [4,294,967,296, 413,390,602,240) (381.000 GiB)
        intervals
            sample 5.120 s, aggr 1 m 42.400 s, update 1 s
            target 4 % accesses per 3 aggrs, [5 ms, 10 s] sampling interval
        nr_regions: [10, 1,000]
        scheme 0
            action: migrate_cold to node 2 per 1 s
            target access pattern
                sz: [0 B, max]
                nr_accesses: [0 samples, 0 samples]
                age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals]
            quotas
                0 ns / 8.000 GiB / 0 B per 1 s
                goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0
                goal tuner: temporal

                priority: sz 0.1 %, nr_accesses 0.1 %, age 0.1 %
            watermarks
                metric none, interval 0 ns
                0 %, 0 %, 0 %
            filter 0
                reject young
            statistics
                tried 142 times (560.000 GiB)
                applied 55 times (97.342 GiB)
                97.351 GiB passed filters
                quota exceeded 234 times
                97.351 GiB
                tried 0 snapshots (max 0)
            tried regions (0 B)
        access sample control
        enabled primitives: page_table
kdamond 1
    state: on, pid: 29837
    context 0
        ops: paddr
        target 0
            pid: 0
            region [466,003,951,616, 3,646,427,234,304) (2.893 TiB)
        intervals
            sample 10 s, aggr 3 m 20 s, update 1 s
            target 4 % accesses per 3 aggrs, [5 ms, 10 s] sampling interval
        nr_regions: [10, 1,000]
        scheme 0
            action: migrate_hot to node 0 per 1 s
            target access pattern
                sz: [0 B, max]
                nr_accesses: [1 samples, 3,689,348,814,741,910,528 samples]
                age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals]
            quotas
                0 ns / 8.000 GiB / 8.000 GiB per 1 s
                goal 0: metric node_mem_used_bp (nid 0) target 5,030 current 0
                goal tuner: temporal

                priority: sz 0.1 %, nr_accesses 0.1 %, age 0.1 %
            watermarks
                metric none, interval 0 ns
                0 %, 0 %, 0 %
            filter 0
                allow young
            statistics
                tried 0 times (0 B)
                applied 0 times (0 B)
                0 B passed filters
                quota exceeded 150 times
                0 B
                tried 95 snapshots (max 0)
            tried regions (0 B)
        access sample control
        enabled primitives: page_table
damon_reclaim: off
damon_stat: off
[root@localhost ~]#


DAMON & DAMOS are very flexible, so it requires a deep understanding
of specific workload and how to configure memory tiering with DAMON &
DAMOS.

So firstly I decided to make it like VMware VCF 9.1 does :-)), next I
will reduce the goal from 50% free to 25-25% for numa node 0.

Based on the outputs above, please let me know if other improvements
are required.


Anton

вт, 29 сент. 2026 г. в 20:37, SJ Park <sj@kernel.org>:
>
> On Tue, 29 Sep 2026 16:54:41 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > > Also, fyi again, from damo v3.4.1, you can mark start of new kdamond, context,
> > > target and scheme parameters on the command line using --kdamond, --damon_ctx,
> > > --damon_target, and --damos_scheme.  you coud use those instead of
> > > --damos_nr_quota_goals, --damos_nr_filters, --nr_targets, --nr_schemes and
> > > --nr_ctxs.  E.g.,
> >
> > Fedora shows old damo version,
> >
> > [root@localhost anton]# rpm -qa|grep -i damo
> > damo-3.3.0-1.fc44.noarch
> > [root@localhost anton]#
> >
> > so I configured latest available 3.4.1
> >
> > [root@localhost damo]# damo version
> > 3.4.1
> > [root@localhost damo]# which damo
> > /home/anton/damo/damo
> > [root@localhost damo]#
>
> Thank you!  Hopefully upgrading the version was not that difficult.  You can
> simply git-clone the repo and use the 'damo' executable file under the
> local-cloned repo.
>
> >
> > > Seems you manually made this change.  Has this executed after 'damo start'?
> > > Also, did you 'commit' the updated commit input?  If any of your answer to the
> > > questions is not "yes", the tuner update may not applied.
> >
> > Yes, it was executed before 'damo start', but I didn't do 'commit'.
>
> To do this manually, you should write the files after 'damo start', and also do
> 'commit'.
>
> Anyway, this means your previous run was using 'consist' tuner.  That explains
> why it didn't show any difference.
>
> >
> > > It would also be helpful if you could run 'damo report damon --damos_stats'
> > > periodically (say, once per 5-10 seconds) while the migration is ongoing and
> > > share the outputs with us.
> >
> > Ok, initially I have,
> >
> > [root@localhost ~]# numastat -z -p $(pgrep valkey-server)
> >
> > Per-node process memory usage (in MBs) for PID 25361 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >                   --------------- --------------- ---------------
> > Heap                         0.11            0.00            0.00
> > Stack                        0.03            0.00            0.00
> > Private                 238356.60            0.50            4.86
> > ----------------  --------------- --------------- ---------------
> > Total                   238356.73            0.50            4.86
> >
> >                             Total
> >                   ---------------
> > Heap                         0.11
> > Stack                        0.03
> > Private                 238361.95
> > ----------------  ---------------
> > Total                   238362.09
> > [root@localhost ~]#
> >
> > I run
> >
> > [root@localhost ~]# /home/anton/damo/damo start \
> >         ` # A kdamond to demote cold memory from node 0 to node 2 ` \
> >         --kdamond --damon_ctx --monitoring_intervals_autotune \
> >                 --damon_target --numa_node 0 \
> >                 --damos_scheme \
> >                         --damos_action migrate_cold 2 \
> >                         --damos_access_rate 0% 0% --damos_apply_interval 1s \
> >                         ` # up to 8 GiB per second ` \
> >                         --damos_quota_interval 1s --damos_quota_space 8G \
> >                         ` # aiming at least 50% node 0 free memory ` \
> >                         --damos_quota_goal node_mem_free_bp 50% 0 \
> >                         ` # using temporal tuner ` \
> >                         --damos_quota_goal_tuner temporal \
> >                         --damos_filter reject young \
> >         ` # A kdamond to promote hot memory from node 2 to node 0 ` \
> >         --kdamond --damon_ctx --monitoring_intervals_autotune \
> >                 --damon_target --numa_node 2 \
> >                 --damos_scheme \
> >                         --damos_action migrate_hot 0 \
> >                         --damos_access_rate 5% max --damos_apply_interval 1s \
> >                         ` # up to 8 GiB per second ` \
> >                         --damos_quota_interval 1s --damos_quota_space 8G \
> >                         ` # aiming at least 50.3% node 0 memory utilization ` \
> >                         --damos_quota_goal node_mem_used_bp 50.3% 0 \
> >                         ` # using temporal tuner ` \
> >                         --damos_quota_goal_tuner temporal \
> >                         --damos_filter allow young \
> > >
> > sysinfo loading fail (info update fail (sysfs feature check fail
> > (feature map making fail (staging damos goal feature check purpose
> > kdamond failed))))
>
> Oops...  Seems you were running next branch of damo.  There was a bug.  I
> reproduced it on 7.2.8 kernel, and fixed it.  The fix [1] is now pushed.  Could
> you pull the 'next' branch and try again?
>
> > [root@localhost ~]#
> > [root@localhost ~]# ps -ef|grep -i damo
> > root       26051   25396  0 16:49 pts/2    00:00:00 grep --color=auto -i damo
> > [root@localhost ~]#
> >
> > > It would also be helpful if you could run 'damo report damon --damos_stats'
> > > periodically (say, once per 5-10 seconds) while the migration is ongoing and
> > > share the outputs with us.
> >
> > Sure, I will start something like that,
> > while true; do damo report damon --damos_stats >>
> > /home/anton/damo_report_damon_damos_stats; sleep 10; done
> > once damo will be started.
>
> Sounds good.
>
> >
> >
> > > Interesting.  I doubt if the tuner change is correctly made.  'damo report
> > > damon' can help us understand under what configuration DAMON is running.  It
> > > could help us quickly see if the configuration is done as we intended.  Could
> > > you share 'damo report damon' result on the final state?
> >
> > There are errors above during damo start.
>
> Apparently it was a bug in damo's next branch.  As I mentioned above, the fix
> is now pushed.  Could you try again?
>
> [1] https://github.com/damonitor/damo/commit/fd2fb55b44ba03db55b8a025eef43dcf569831e5
>
> Thanks,
> SJ
>
> [...]

[-- Attachment #2: damo_report_damon_damos_stats.gz --]
[-- Type: application/x-gzip, Size: 833 bytes --]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-30  4:03               ` Anton Gavriliuk
@ 2026-09-30  8:35                 ` SJ Park
  2026-09-30 16:26                   ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-30  8:35 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Wed, 30 Sep 2026 07:03:47 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> > Oops...  Seems you were running next branch of damo.  There was a bug.  I
> > reproduced it on 7.2.8 kernel, and fixed it.  The fix [1] is now pushed.  Could
> > you pull the 'next' branch and try again?
> 
> Done.
[...]
> It looks now it keeps ~50% free numa node 0,

Awesome :D

[...]
> [root@localhost ~]# /home/anton/damo/damo report damon
> kdamond 0
>     state: on, pid: 29836
>     context 0
>         ops: paddr
[...]
>         scheme 0
>             action: migrate_cold to node 2 per 1 s
>             target access pattern
>                 sz: [0 B, max]
>                 nr_accesses: [0 samples, 0 samples]
>                 age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals]
>             quotas
>                 0 ns / 8.000 GiB / 0 B per 1 s
>                 goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0
>                 goal tuner: temporal

Ok, the above line confirms DAMON is running with 'temporal' tuner.

[...]
> DAMON & DAMOS are very flexible, so it requires a deep understanding
> of specific workload and how to configure memory tiering with DAMON &
> DAMOS.

Indeed it is.  Nonetheless, for the reason we provide DAMON modules [1] or damo
scripts for commonly known DAMON usages.  We have DAMON_RECLAIM module [2] for
proactive memory reclamation, and mem_tier.sh [3] for memory tiering.

My suggestion is to start from such examples, find what is not working and
asking questions to me.  So you are doing all great :D

> 
> So firstly I decided to make it like VMware VCF 9.1 does :-)), next I
> will reduce the goal from 50% free to 25-25% for numa node 0.

Sounds like a good plan.

> 
> Based on the outputs above, please let me know if other improvements
> are required.

This temporal tuner test has proved the quota system is not broken.  It implies
the previous test resulted in migrating nearly all data to the lower tier,
because there was no hot data to promote back to the upper tier.

In many common benchmarks using zipfian-like access distribution, hot data is
always hot, and cold data is always cold.  And usually the amount of cold data
is much larger than hot data.  I found [4] the pattern is making evaluation of
tiering solutions difficult, and was using artificial access pattern mixing to
work around.

My imagined real world workload is, there will be hot and cold data.  And the
pattern will nearly always be stable.  But, occasionally there will be changes.
Say, a chunk of data will be hot for a few hours, but suddenly be cold and keep
being cold for a few hours, then sudenly be warm for hours, and so on.  The
auto-tuning based DAMON tiering is designed with such workload in mind.

For your testing, I think the setup is good.  I'd suggest lower target free
memory ratio, though.  Assuming the theory (your workload has static access
pattern of small hot data) is true, using consist tuner should also be fine.

You will have more than expected data in lower tier, but if those are truly
cold, why would we bother?  If you still want strict upper tier utilization,
you could make the access pattern and filter condition less strict.  Ideally,
the access pattern and filter conditions should all go away, assuming DAMON can
find true hot and cold data.  And I believe that would be the case for
long-running real world workloads.  Short-running test workload would show
DAMON making wrong decisions in short term, though.

If you want to make sure if my theory is true, you could also profile the
access pattern of your test workload using DAMON.  I actually did it for my
auto-tuned tiering test [4] to find the fact.

[1] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#special-purpose-access-aware-kernel-modules
[2] https://origin.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
[3] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
[4] https://lkml.kernel.org/r/20250420194030.75838-1-sj@kernel.org


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-30  8:35                 ` SJ Park
@ 2026-09-30 16:26                   ` Anton Gavriliuk
  2026-09-30 17:43                     ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-09-30 16:26 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

I would go back to memory tiering VCF9.1-like behavior you helped me
configure, for new few questions:

1.  It works with the goal of 50% free memory for numa node 0, but
when I tried to set 25% free memory, it didn't do anything...

[root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat

Per-node process memory usage (in MBs) for PID 3808 (valkey-server)
                           Node 0          Node 1          Node 2
    Node 3           Total
                  --------------- --------------- ---------------
--------------- ---------------
Huge                         0.00            0.00            0.00
      0.00            0.00
Heap                         0.11            0.00            0.00
      0.00            0.11
Stack                        0.03            0.00            0.00
      0.00            0.03
Private                 238356.25            2.19            3.42
      0.00       238361.86
----------------  --------------- --------------- ---------------
--------------- ---------------
Total                   238356.39            2.19            3.42
      0.00       238362.00
[root@localhost ~]#


/home/anton/damo/damo start \
        ` # A kdamond to demote cold memory from node 0 to node 2 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 0 \
                --damos_scheme \
                        --damos_action migrate_cold 2 \
                        --damos_access_rate 0% 0% --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 25% node 0 free memory ` \
                        --damos_quota_goal node_mem_free_bp 25% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter reject young \
        ` # A kdamond to promote hot memory from node 2 to node 0 ` \
        --kdamond --damon_ctx --monitoring_intervals_autotune \
                --damon_target --numa_node 2 \
                --damos_scheme \
                        --damos_action migrate_hot 0 \
                        --damos_access_rate 5% max --damos_apply_interval 1s \
                        ` # up to 8 GiB per second ` \
                        --damos_quota_interval 1s --damos_quota_space 8G \
                        ` # aiming at least 75.3% node 0 memory utilization ` \
                        --damos_quota_goal node_mem_used_bp 75.3% 0 \
                        ` # using temporal tuner ` \
                        --damos_quota_goal_tuner temporal \
                        --damos_filter allow young \


[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/schemes/0/quotas/goal_tuner
temporal
[root@localhost ~]# cat
/sys/kernel/mm/damon/admin/kdamonds/1/contexts/0/schemes/0/quotas/goal_tuner
temporal


[root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat

Per-node process memory usage (in MBs) for PID 3808 (valkey-server)
                           Node 0          Node 1          Node 2
    Node 3           Total
                  --------------- --------------- ---------------
--------------- ---------------
Huge                         0.00            0.00            0.00
      0.00            0.00
Heap                         0.11            0.00            0.00
      0.00            0.11
Stack                        0.03            0.00            0.00
      0.00            0.03
Private                 238356.25            2.19            3.42
      0.00       238361.86
----------------  --------------- --------------- ---------------
--------------- ---------------
Total                   238356.39            2.19            3.42
      0.00       238362.00
[root@localhost ~]#


2. If there are no un-accessed pages to demote for achieving a defined
goal, damon will demote pages with the most rare access ?

Anton

ср, 30 сент. 2026 г. в 11:35, SJ Park <sj@kernel.org>:
>
> On Wed, 30 Sep 2026 07:03:47 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > > Oops...  Seems you were running next branch of damo.  There was a bug.  I
> > > reproduced it on 7.2.8 kernel, and fixed it.  The fix [1] is now pushed.  Could
> > > you pull the 'next' branch and try again?
> >
> > Done.
> [...]
> > It looks now it keeps ~50% free numa node 0,
>
> Awesome :D
>
> [...]
> > [root@localhost ~]# /home/anton/damo/damo report damon
> > kdamond 0
> >     state: on, pid: 29836
> >     context 0
> >         ops: paddr
> [...]
> >         scheme 0
> >             action: migrate_cold to node 2 per 1 s
> >             target access pattern
> >                 sz: [0 B, max]
> >                 nr_accesses: [0 samples, 0 samples]
> >                 age: [0 aggr_intervals, 184,467,440,737,095 aggr_intervals]
> >             quotas
> >                 0 ns / 8.000 GiB / 0 B per 1 s
> >                 goal 0: metric node_mem_free_bp (nid 0) target 5,000 current 0
> >                 goal tuner: temporal
>
> Ok, the above line confirms DAMON is running with 'temporal' tuner.
>
> [...]
> > DAMON & DAMOS are very flexible, so it requires a deep understanding
> > of specific workload and how to configure memory tiering with DAMON &
> > DAMOS.
>
> Indeed it is.  Nonetheless, for the reason we provide DAMON modules [1] or damo
> scripts for commonly known DAMON usages.  We have DAMON_RECLAIM module [2] for
> proactive memory reclamation, and mem_tier.sh [3] for memory tiering.
>
> My suggestion is to start from such examples, find what is not working and
> asking questions to me.  So you are doing all great :D
>
> >
> > So firstly I decided to make it like VMware VCF 9.1 does :-)), next I
> > will reduce the goal from 50% free to 25-25% for numa node 0.
>
> Sounds like a good plan.
>
> >
> > Based on the outputs above, please let me know if other improvements
> > are required.
>
> This temporal tuner test has proved the quota system is not broken.  It implies
> the previous test resulted in migrating nearly all data to the lower tier,
> because there was no hot data to promote back to the upper tier.
>
> In many common benchmarks using zipfian-like access distribution, hot data is
> always hot, and cold data is always cold.  And usually the amount of cold data
> is much larger than hot data.  I found [4] the pattern is making evaluation of
> tiering solutions difficult, and was using artificial access pattern mixing to
> work around.
>
> My imagined real world workload is, there will be hot and cold data.  And the
> pattern will nearly always be stable.  But, occasionally there will be changes.
> Say, a chunk of data will be hot for a few hours, but suddenly be cold and keep
> being cold for a few hours, then sudenly be warm for hours, and so on.  The
> auto-tuning based DAMON tiering is designed with such workload in mind.
>
> For your testing, I think the setup is good.  I'd suggest lower target free
> memory ratio, though.  Assuming the theory (your workload has static access
> pattern of small hot data) is true, using consist tuner should also be fine.
>
> You will have more than expected data in lower tier, but if those are truly
> cold, why would we bother?  If you still want strict upper tier utilization,
> you could make the access pattern and filter condition less strict.  Ideally,
> the access pattern and filter conditions should all go away, assuming DAMON can
> find true hot and cold data.  And I believe that would be the case for
> long-running real world workloads.  Short-running test workload would show
> DAMON making wrong decisions in short term, though.
>
> If you want to make sure if my theory is true, you could also profile the
> access pattern of your test workload using DAMON.  I actually did it for my
> auto-tuned tiering test [4] to find the fact.
>
> [1] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#special-purpose-access-aware-kernel-modules
> [2] https://origin.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [3] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
> [4] https://lkml.kernel.org/r/20250420194030.75838-1-sj@kernel.org
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-30 16:26                   ` Anton Gavriliuk
@ 2026-09-30 17:43                     ` SJ Park
  2026-10-01  9:58                       ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-30 17:43 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Wed, 30 Sep 2026 19:26:31 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> I would go back to memory tiering VCF9.1-like behavior you helped me
> configure, for new few questions:
> 
> 1.  It works with the goal of 50% free memory for numa node 0, but
> when I tried to set 25% free memory, it didn't do anything...
> 
> [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat
> 
> Per-node process memory usage (in MBs) for PID 3808 (valkey-server)
>                            Node 0          Node 1          Node 2
>     Node 3           Total
>                   --------------- --------------- ---------------
> --------------- ---------------
> Huge                         0.00            0.00            0.00
>       0.00            0.00
> Heap                         0.11            0.00            0.00
>       0.00            0.11
> Stack                        0.03            0.00            0.00
>       0.00            0.03
> Private                 238356.25            2.19            3.42
>       0.00       238361.86
> ----------------  --------------- --------------- ---------------
> --------------- ---------------
> Total                   238356.39            2.19            3.42
>       0.00       238362.00
> [root@localhost ~]#

So, your workload is using ~232 GiB of node 0 memory and no demotion of it is
occurred.  According to your original mail [1], the node 0 has ~377 GiB memory.
If we make 25% of it free, it means it should have 75% of it (~282 GiB) be
utilized.  If the node0 has only your workload, it means node 0 memory
utilization is still only ~61%.  In other words, 25% free memory goal is
already achieved.  Than DAMON wouldn't do any demotion.

Does the theory makes sense?  Maybe you can confirm by checking the free memory
ratio of node 0.

[...]
> 2. If there are no un-accessed pages to demote for achieving a defined
> goal, damon will demote pages with the most rare access ?

No.  Because you set '--access_rate 0% 0%' for the demotion scheme, it will do
no demotion in the scenario.  You could try '--access_rate 0% max' if you want.
You may also need to remove '--damos_filter reject young' option from the
command if that is really what you want to do.

[1] https://lore.kernel.org/CAAiJnjp5F8ZuPG1gEsD_Wgs7Z+GToz67nzjr0490sz9y328=dg@mail.gmail.com


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-09-30 17:43                     ` SJ Park
@ 2026-10-01  9:58                       ` Anton Gavriliuk
  2026-10-01 10:26                         ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-10-01  9:58 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

> So, your workload is using ~232 GiB of node 0 memory and no demotion of it is
> occurred.  According to your original mail [1], the node 0 has ~377 GiB memory.
> If we make 25% of it free, it means it should have 75% of it (~282 GiB) be
> utilized.  If the node0 has only your workload, it means node 0 memory
> utilization is still only ~61%.  In other words, 25% free memory goal is
> already achieved.  Than DAMON wouldn't do any demotion.

OMG.... what was pretty easy.
I feel ashamed.  My apologies.

I increased Valkey memory allocation to 300 GB and it works.

> No.  Because you set '--access_rate 0% 0%' for the demotion scheme, it will do
> no demotion in the scenario.  You could try '--access_rate 0% max' if you want.
> You may also need to remove '--damos_filter reject young' option from the
> command if that is really what you want to do.

'--access_rate 0% max' means range 0% - 100% ?
If yes, then I need memory management decisions based on criteria
other than access frequency.
Age-based filters or something else.

The idea is for achieving a defined goal - keep 25% free for new hot
allocations in the fastest tier, if there are no un-accessed pages,
then start demoting the most-rare access or add aga-based filters.
Complicated but interesting, I need to think more about that.

Anton

ср, 30 сент. 2026 г. в 20:43, SJ Park <sj@kernel.org>:
>
> On Wed, 30 Sep 2026 19:26:31 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > I would go back to memory tiering VCF9.1-like behavior you helped me
> > configure, for new few questions:
> >
> > 1.  It works with the goal of 50% free memory for numa node 0, but
> > when I tried to set 25% free memory, it didn't do anything...
> >
> > [root@localhost ~]# COLUMNS=500 numastat -p $(pgrep valkey-server) | cat
> >
> > Per-node process memory usage (in MBs) for PID 3808 (valkey-server)
> >                            Node 0          Node 1          Node 2
> >     Node 3           Total
> >                   --------------- --------------- ---------------
> > --------------- ---------------
> > Huge                         0.00            0.00            0.00
> >       0.00            0.00
> > Heap                         0.11            0.00            0.00
> >       0.00            0.11
> > Stack                        0.03            0.00            0.00
> >       0.00            0.03
> > Private                 238356.25            2.19            3.42
> >       0.00       238361.86
> > ----------------  --------------- --------------- ---------------
> > --------------- ---------------
> > Total                   238356.39            2.19            3.42
> >       0.00       238362.00
> > [root@localhost ~]#
>
> So, your workload is using ~232 GiB of node 0 memory and no demotion of it is
> occurred.  According to your original mail [1], the node 0 has ~377 GiB memory.
> If we make 25% of it free, it means it should have 75% of it (~282 GiB) be
> utilized.  If the node0 has only your workload, it means node 0 memory
> utilization is still only ~61%.  In other words, 25% free memory goal is
> already achieved.  Than DAMON wouldn't do any demotion.
>
> Does the theory makes sense?  Maybe you can confirm by checking the free memory
> ratio of node 0.
>
> [...]
> > 2. If there are no un-accessed pages to demote for achieving a defined
> > goal, damon will demote pages with the most rare access ?
>
> No.  Because you set '--access_rate 0% 0%' for the demotion scheme, it will do
> no demotion in the scenario.  You could try '--access_rate 0% max' if you want.
> You may also need to remove '--damos_filter reject young' option from the
> command if that is really what you want to do.
>
> [1] https://lore.kernel.org/CAAiJnjp5F8ZuPG1gEsD_Wgs7Z+GToz67nzjr0490sz9y328=dg@mail.gmail.com
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-01  9:58                       ` Anton Gavriliuk
@ 2026-10-01 10:26                         ` SJ Park
  2026-10-01 12:08                           ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-10-01 10:26 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Thu, 1 Oct 2026 12:58:39 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> > So, your workload is using ~232 GiB of node 0 memory and no demotion of it is
> > occurred.  According to your original mail [1], the node 0 has ~377 GiB memory.
> > If we make 25% of it free, it means it should have 75% of it (~282 GiB) be
> > utilized.  If the node0 has only your workload, it means node 0 memory
> > utilization is still only ~61%.  In other words, 25% free memory goal is
> > already achieved.  Than DAMON wouldn't do any demotion.
> 
> OMG.... what was pretty easy.
> I feel ashamed.  My apologies.

No worry, I'm happy that I helped :)

> 
> I increased Valkey memory allocation to 300 GB and it works.

Awesome, thank you for confirming.

> 
> > No.  Because you set '--access_rate 0% 0%' for the demotion scheme, it will do
> > no demotion in the scenario.  You could try '--access_rate 0% max' if you want.
> > You may also need to remove '--damos_filter reject young' option from the
> > command if that is really what you want to do.
> 
> '--access_rate 0% max' means range 0% - 100% ?

That's correct.

> If yes, then I need memory management decisions based on criteria
> other than access frequency.
> Age-based filters or something else.

Maybe not.  Under the (auto-tuned) quota, DAMOS applies the action to hottest
or coldest quota amount of memory depend on the action.  For example, in your
current setup, let's suppose DAMON found 10 GiB of hot memory (have >=5% access
rate) on node 2.  And the quota is auto-tuned to 4 GiB.  In the case, DAMOS
will find hottest 5 GiB hot memory and migrate those to node 0.  For finding
the hottest memory, DAMOS calculates access temperature of each memory based on
the access frequency and the age.

If you set '--access_rate 0% max', DAMOS will show all memory in node 2 is
eligible to migrate to node 0.  But, because you have the auto-tuned quota that
has 8 GiB/s upperlimit, it will migrate only up to 8 GiB hottest memory per
second.

Note that your goal also have '--damos_filter allow young' option for the
promotion.  That asks DAMOS to double check if each migration candidate page is
marked as accessed since the last double check, by h/w.  If it is not marked as
accessed by h/w, DAMOS will not migrate it.  So to allow promoting short-time
unaccessed memory, you will need to turn off this double check, too.

> 
> The idea is for achieving a defined goal - keep 25% free for new hot
> allocations in the fastest tier, if there are no un-accessed pages,
> then start demoting the most-rare access or add aga-based filters.
> Complicated but interesting, I need to think more about that.

That makes sense, and I believe removing '--access_rate' (absence of
'--access_rate' is same to '--access_rate 0% max') and '--damos_filter' is one
of the ways to achieve that.


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-01 10:26                         ` SJ Park
@ 2026-10-01 12:08                           ` Anton Gavriliuk
  2026-10-01 13:27                             ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-10-01 12:08 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

With your support, my progress tremendously faster than I expected :-)

Moving forward.

In my VCF9.1-like example Demote/Promote bandwidth is limited by admin
up to 8GB/s.

With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s.
In this case, kdamond single CPU core limited or can use more CPU
cores if/when required ?

~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5
DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s.

Anton

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-01 12:08                           ` Anton Gavriliuk
@ 2026-10-01 13:27                             ` SJ Park
  2026-10-01 16:54                               ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-10-01 13:27 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Thu, 1 Oct 2026 15:08:38 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> With your support, my progress tremendously faster than I expected :-)

Thank you.  The pleasure is mine :)

> 
> Moving forward.
> 
> In my VCF9.1-like example Demote/Promote bandwidth is limited by admin
> up to 8GB/s.
> 
> With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s.
> In this case, kdamond single CPU core limited or can use more CPU
> cores if/when required ?
> 
> ~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5
> DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s.

It is basically limited to single CPU core per kdamond.  You could split memory
into multiple address ranges and assign kdamon per the range to utilize multi
CPUs, if needed.  SK Hynix [1] was using such an approach.

That said, I'm personally curious if such fast migration is really needed in
the real world workload.  I assume the real production workload would have
stable access pattern but only occasionally get access pattern changes.
Sometimes temporal and rapid access pattern change could also be made.  But for
such temporal pattern change, doing migration would only be costy, since the
data that suddenly hot could be soon be cold.  And reliability is important in
production.  Hence I was thinking slowly making the balance is better than too
quickly and reactively making migrations that will turn out to be no really
needed.

That said, if you want to test faster migration speed, the multiple kdamonds
usage could be one way to test.  If it turns out it is really needed and
splitting address ranges has problems, we can consider adding DAMOS feature for
utilizing multiple CPUs for faster migrations.

[1] https://sched.co/2913n


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-01 13:27                             ` SJ Park
@ 2026-10-01 16:54                               ` Anton Gavriliuk
  2026-10-02  8:34                                 ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-10-01 16:54 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

> That said, I'm personally curious if such fast migration is really needed in
> the real world workload.  I assume the real production workload would have
> stable access pattern but only occasionally get access pattern changes.
> Sometimes temporal and rapid access pattern change could also be made.  But for
> such temporal pattern change, doing migration would only be costy, since the
> data that suddenly hot could be soon be cold.  And reliability is important in
> production.  Hence I was thinking slowly making the balance is better than too
> quickly and reactively making migrations that will turn out to be no really
> needed.

> That said, if you want to test faster migration speed, the multiple kdamonds
> usage could be one way to test.  If it turns out it is really needed and
> splitting address ranges has problems, we can consider adding DAMOS feature for
> utilizing multiple CPUs for faster migrations.


I agree with you!, but now it looks that demotion is too slow.
The server is completely idle, only I play with memory tiering.

As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices,

[root@localhost anton]# ndctl list
[
  {
    "dev":"namespace1.0",
    "mode":"devdax",
    "map":"dev",
    "size":3183575302144,
    "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f",
    "chardev":"dax1.0",
    "align":2097152
  },
  {
    "dev":"namespace0.0",
    "mode":"devdax",
    "map":"dev",
    "size":3183575302144,
    "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2",
    "chardev":"dax0.0",
    "align":2097152
  }
]
[root@localhost anton]#

and then configured as system ram by next command,

daxctl reconfigure-device --mode=system-ram all

The problem is during demotion (to keep 25% free in the fast tier),
kdamond.0 utilizes single CPU core ~100%, but writing to the slower
tier (System PMM Write) ~1.4 GB/s what is much slower than defined
limit 8 GB/s.


top - 16:26:06 up 12 min,  2 users,  load average: 0.81, 0.75, 0.54
Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie
%Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used,    465.2 buff/cache
MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6424006.+avail Mem

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   3227 root      20   0       0      0      0 R  99.8   0.0   1:28.92 kdamond.0
   2771 root      20   0 6932852  10260   5608 S   0.7   0.0   0:02.95
pcm-memory
   1593 root      20   0   17088   8352   7140 S   0.3   0.0   0:00.99
systemd-logind
   2263 root      20   0  300.4g 294.4g   6108 S   0.3   4.4   4:42.72
valkey-server
   3327 root      20   0   10916   6064   3764 R   0.3   0.0   0:00.16 top
      1 root      20   0   25356  15232  10456 S   0.0   0.0   0:02.13 systemd
      2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd



|---------------------------------------||---------------------------------------|
|--            System DRAM Read Throughput(MB/s):        751.95
        --|
|--           System DRAM Write Throughput(MB/s):        742.58
        --|
|--             System PMM Read Throughput(MB/s):        693.12
        --|
|--            System PMM Write Throughput(MB/s):       1387.47
        --|
|--                 System Read Throughput(MB/s):       1445.08
        --|
|--                System Write Throughput(MB/s):       2130.05
        --|
|--               System Memory Throughput(MB/s):       3575.12
        --|
|---------------------------------------||---------------------------------------|



Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
51312892466 lost: 0/0 drop
Overhead  Shared Object                      Symbol
  47.50%  [kernel]                           [k] migrate_folio_unmap
  16.11%  [kernel]                           [k] clear_highpages_kasan_tagged
  14.93%  [kernel]                           [k] copy_mc_fragile
   1.63%  [kernel]                           [k] smp_call_function_many_cond
   1.44%  [kernel]                           [k] page_vma_mapped_walk
   0.97%  [kernel]                           [k] folio_migrate_flags
   0.87%  [kernel]                           [k] try_to_migrate_one
   0.66%  [kernel]                           [k] rmqueue_bulk
   0.61%  [kernel]                           [k]
__list_del_entry_valid_or_report
   0.59%  [kernel]                           [k] remove_migration_pte
   0.50%  [kernel]                           [k] rmap_walk_anon
   0.45%  [kernel]                           [k] __free_one_page
   0.44%  [kernel]                           [k] migrate_folio_move
   0.44%  [kernel]                           [k] folio_add_anon_rmap_ptes
   0.40%  [kernel]                           [k] lru_gen_add_folio
   0.40%  [kernel]                           [k] __mod_memcg_lruvec_state
   0.39%  [kernel]                           [k] mod_node_page_state
   0.35%  [kernel]                           [k] migrate_pages_batch
   0.33%  [kernel]                           [k] _raw_spin_lock
   0.33%  [kernel]                           [k] up_read



Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
64449579226
migrate_folio_unmap  /proc/kcore [Percent: local period]
Percent │       movq   %r13,0x20(%rsp)
   0.01 │       movq   %rdx,%r13
        │       movq   %r14,0x28(%rsp)
        │       movl   %r9d,%r14d
   0.00 │       movq   %r15,0x30(%rsp)
   0.01 │       movq   %r8,%r15
        │     → callq  *%rax
   0.16 │       testq  %rax,%rax
        │     ↓ je     2e3
   0.01 │       movq   %rbp,0x10(%rsp)
   0.01 │       movq   %rax,%rbp
   0.04 │       movq   %rax,(%r15)
   0.00 │       movq   $0x0,0x28(%rax)
  99.27 │       lock
        │       btsq   $0x0,(%rbx)
   0.01 │     ↓ jb     26b
   0.11 │ 67:   movq   (%rbx),%rdx
   0.01 │       testq  $0x2,(%rbx)
   0.01 │     ↓ je     e2
        │       cmpl   $0x2,%r14d
        │     ↓ je     d2
        │       movq   0x40(%rsp),%r8
        │       movq   %rbx,%rdi
        │       movl   $0x1,%ecx
        │       xorl   %edx,%edx

I also tried without "--damos_filter reject young" and "--damos_filter
allow young", but got the same single CPU ~100% utilization and
demotion poor performance.

Anton

чт, 1 окт. 2026 г. в 16:27, SJ Park <sj@kernel.org>:
>
> On Thu, 1 Oct 2026 15:08:38 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > With your support, my progress tremendously faster than I expected :-)
>
> Thank you.  The pleasure is mine :)
>
> >
> > Moving forward.
> >
> > In my VCF9.1-like example Demote/Promote bandwidth is limited by admin
> > up to 8GB/s.
> >
> > With CXL/PCIe 6.0 it may require moving pages between tiers 10's GB/s.
> > In this case, kdamond single CPU core limited or can use more CPU
> > cores if/when required ?
> >
> > ~1 year ago I played with AMD Zen5 9455 and 12 channels 6400 MT/s DDR5
> > DRAM, single thread core-to-local-memory bandwidth was ~49 GB/s.
>
> It is basically limited to single CPU core per kdamond.  You could split memory
> into multiple address ranges and assign kdamon per the range to utilize multi
> CPUs, if needed.  SK Hynix [1] was using such an approach.
>
> That said, I'm personally curious if such fast migration is really needed in
> the real world workload.  I assume the real production workload would have
> stable access pattern but only occasionally get access pattern changes.
> Sometimes temporal and rapid access pattern change could also be made.  But for
> such temporal pattern change, doing migration would only be costy, since the
> data that suddenly hot could be soon be cold.  And reliability is important in
> production.  Hence I was thinking slowly making the balance is better than too
> quickly and reactively making migrations that will turn out to be no really
> needed.
>
> That said, if you want to test faster migration speed, the multiple kdamonds
> usage could be one way to test.  If it turns out it is really needed and
> splitting address ranges has problems, we can consider adding DAMOS feature for
> utilizing multiple CPUs for faster migrations.
>
> [1] https://sched.co/2913n
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-01 16:54                               ` Anton Gavriliuk
@ 2026-10-02  8:34                                 ` SJ Park
  2026-10-02 10:02                                   ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-10-02  8:34 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Thu, 1 Oct 2026 19:54:09 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> > That said, I'm personally curious if such fast migration is really needed in
> > the real world workload.  I assume the real production workload would have
> > stable access pattern but only occasionally get access pattern changes.
> > Sometimes temporal and rapid access pattern change could also be made.  But for
> > such temporal pattern change, doing migration would only be costy, since the
> > data that suddenly hot could be soon be cold.  And reliability is important in
> > production.  Hence I was thinking slowly making the balance is better than too
> > quickly and reactively making migrations that will turn out to be no really
> > needed.
> 
> > That said, if you want to test faster migration speed, the multiple kdamonds
> > usage could be one way to test.  If it turns out it is really needed and
> > splitting address ranges has problems, we can consider adding DAMOS feature for
> > utilizing multiple CPUs for faster migrations.
> 
> 
> I agree with you!, but now it looks that demotion is too slow.
> The server is completely idle, only I play with memory tiering.
> 
> As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices,
> 
> [root@localhost anton]# ndctl list
> [
>   {
>     "dev":"namespace1.0",
>     "mode":"devdax",
>     "map":"dev",
>     "size":3183575302144,
>     "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f",
>     "chardev":"dax1.0",
>     "align":2097152
>   },
>   {
>     "dev":"namespace0.0",
>     "mode":"devdax",
>     "map":"dev",
>     "size":3183575302144,
>     "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2",
>     "chardev":"dax0.0",
>     "align":2097152
>   }
> ]
> [root@localhost anton]#
> 
> and then configured as system ram by next command,
> 
> daxctl reconfigure-device --mode=system-ram all
> 
> The problem is during demotion (to keep 25% free in the fast tier),
> kdamond.0 utilizes single CPU core ~100%, but writing to the slower
> tier (System PMM Write) ~1.4 GB/s what is much slower than defined
> limit 8 GB/s.

The defined limit is only upper limit, so real speed could be slower than that.
FWIW, the example memory tiering script [1] uses 200 MiB/second as the
upperlimit.  It was set by my gut feeling, not by some good data, though.

I personally feel like ~1.4 GB/s is still a good speed for long-running
production workloads, though I don't have a data to support that.  Do you have
some data or reason to pursue 8 GB/s ?

> 
> 
> top - 16:26:06 up 12 min,  2 users,  load average: 0.81, 0.75, 0.54
> Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie
> %Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
> MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used,    465.2 buff/cache
> MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6424006.+avail Mem
> 
>     PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
>    3227 root      20   0       0      0      0 R  99.8   0.0   1:28.92 kdamond.0
>    2771 root      20   0 6932852  10260   5608 S   0.7   0.0   0:02.95
> pcm-memory
>    1593 root      20   0   17088   8352   7140 S   0.3   0.0   0:00.99
> systemd-logind
>    2263 root      20   0  300.4g 294.4g   6108 S   0.3   4.4   4:42.72
> valkey-server
>    3327 root      20   0   10916   6064   3764 R   0.3   0.0   0:00.16 top
>       1 root      20   0   25356  15232  10456 S   0.0   0.0   0:02.13 systemd
>       2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd
> 
> 
> 
> |---------------------------------------||---------------------------------------|
> |--            System DRAM Read Throughput(MB/s):        751.95
>         --|
> |--           System DRAM Write Throughput(MB/s):        742.58
>         --|
> |--             System PMM Read Throughput(MB/s):        693.12
>         --|
> |--            System PMM Write Throughput(MB/s):       1387.47
>         --|
> |--                 System Read Throughput(MB/s):       1445.08
>         --|
> |--                System Write Throughput(MB/s):       2130.05
>         --|
> |--               System Memory Throughput(MB/s):       3575.12
>         --|
> |---------------------------------------||---------------------------------------|
> 
> 
> 
> Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 51312892466 lost: 0/0 drop
> Overhead  Shared Object                      Symbol
>   47.50%  [kernel]                           [k] migrate_folio_unmap
>   16.11%  [kernel]                           [k] clear_highpages_kasan_tagged
>   14.93%  [kernel]                           [k] copy_mc_fragile
>    1.63%  [kernel]                           [k] smp_call_function_many_cond
>    1.44%  [kernel]                           [k] page_vma_mapped_walk
>    0.97%  [kernel]                           [k] folio_migrate_flags
>    0.87%  [kernel]                           [k] try_to_migrate_one
>    0.66%  [kernel]                           [k] rmqueue_bulk
>    0.61%  [kernel]                           [k]
> __list_del_entry_valid_or_report
>    0.59%  [kernel]                           [k] remove_migration_pte
>    0.50%  [kernel]                           [k] rmap_walk_anon
>    0.45%  [kernel]                           [k] __free_one_page
>    0.44%  [kernel]                           [k] migrate_folio_move
>    0.44%  [kernel]                           [k] folio_add_anon_rmap_ptes
>    0.40%  [kernel]                           [k] lru_gen_add_folio
>    0.40%  [kernel]                           [k] __mod_memcg_lruvec_state
>    0.39%  [kernel]                           [k] mod_node_page_state
>    0.35%  [kernel]                           [k] migrate_pages_batch
>    0.33%  [kernel]                           [k] _raw_spin_lock
>    0.33%  [kernel]                           [k] up_read
> 
> 
> 
> Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 64449579226
> migrate_folio_unmap  /proc/kcore [Percent: local period]
> Percent │       movq   %r13,0x20(%rsp)
>    0.01 │       movq   %rdx,%r13
>         │       movq   %r14,0x28(%rsp)
>         │       movl   %r9d,%r14d
>    0.00 │       movq   %r15,0x30(%rsp)
>    0.01 │       movq   %r8,%r15
>         │     → callq  *%rax
>    0.16 │       testq  %rax,%rax
>         │     ↓ je     2e3
>    0.01 │       movq   %rbp,0x10(%rsp)
>    0.01 │       movq   %rax,%rbp
>    0.04 │       movq   %rax,(%r15)
>    0.00 │       movq   $0x0,0x28(%rax)
>   99.27 │       lock
>         │       btsq   $0x0,(%rbx)
>    0.01 │     ↓ jb     26b
>    0.11 │ 67:   movq   (%rbx),%rdx
>    0.01 │       testq  $0x2,(%rbx)
>    0.01 │     ↓ je     e2
>         │       cmpl   $0x2,%r14d
>         │     ↓ je     d2
>         │       movq   0x40(%rsp),%r8
>         │       movq   %rbx,%rdi
>         │       movl   $0x1,%ecx
>         │       xorl   %edx,%edx
> 
> I also tried without "--damos_filter reject young" and "--damos_filter
> allow young", but got the same single CPU ~100% utilization and
> demotion poor performance.

Thank you for doing these great profiling and sharing the results, Anton.
Apparently the bottleneck is not in DAMON specific code including the filtering
part.  Instead, the bottleneck is in the migration code, which is out of DAMON,
according to my understanding of the profiling result.  You could also confirm
this by measuring the migration speed using move_pages() like system call,
which also use the migration code.

You may need to optimize the migration code, or make DAMON uses multiple CPUs
if faster speed is really what needed.  I'd suggest using the address range
split approach that I suggested in the previous reply.  If it becomes clear the
faster migration is really needed, we could start thinking about making DAMOS
use multiple CPUs without the address range split.

[1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-02  8:34                                 ` SJ Park
@ 2026-10-02 10:02                                   ` Anton Gavriliuk
  2026-10-02 11:02                                     ` SJ Park
  0 siblings, 1 reply; 21+ messages in thread
From: Anton Gavriliuk @ 2026-10-02 10:02 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

Peak sequential write to slower tier ~12 GB/s.
I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded.

For Memory Tiering with OLTP-like workloads needs to be
promoted/demoted, 1.4 GB/s is OK.
But for bandwidth intensive workloads such as analytics, AI, VDI,...
1.4 GB/s is too slow.

On the server where I test Memory Tiering, Intel CPU 8280L installed.
Despite being already 7 years old, any single CPU core has 13-14 GB/s
bandwidth access to local memory.

So from my point of view - even if we increase the number of cores
keeping in mind 1.4 GB/s per core, we need so many cores to perform
10's GB/s demote/promote.  Firstly we need to improve demote/promote
bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s.

> You could also confirm
> this by measuring the migration speed using move_pages() like system call,
> which also use the migration code.

I'm not a developer and I need to spend more time thinking about how
to do that, but I tried,

I used migratepages command, which is probably use migrate_pages()
instead of move_pages(), but I got the same poor bandwidth and stack
as with DAMON/DAMOS Memory Tiering,

top - 12:43:59 up 13 min,  2 users,  load average: 0.65, 0.88, 0.76
Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie
%Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used,    669.8 buff/cache
MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6423508.+avail Mem

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   2935 root      20   0    2644   1764   1644 R  99.7   0.0   0:42.19
migratepages
   2941 root      20   0   10884   6192   3916 R   0.7   0.0   0:00.04 top
   1238 systemd+  20   0   15860   6996   5952 S   0.3   0.0   0:00.37
systemd-oomd
   1928 root      20   0  666268  24448  23440 S   0.3   0.0   0:00.34 rsyslogd
      1 root      20   0   25840  15532  10492 S   0.0   0.0   0:02.15 systemd
      2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd
      3 root      20   0       0      0      0 S   0.0   0.0   0:00.00
pool_workqueue_release
      4 root       0 -20       0      0      0 I   0.0   0.0   0:00.00
kworker/R-rcu_gp


|---------------------------------------||---------------------------------------|
|--            System DRAM Read Throughput(MB/s):        759.41
        --|
|--           System DRAM Write Throughput(MB/s):        751.36
        --|
|--             System PMM Read Throughput(MB/s):        701.24
        --|
|--            System PMM Write Throughput(MB/s):       1402.15
        --|
|--                 System Read Throughput(MB/s):       1460.65
        --|
|--                System Write Throughput(MB/s):       2153.50
        --|
|--               System Memory Throughput(MB/s):       3614.15
        --|
|---------------------------------------||---------------------------------------|


Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
48280162400 lost: 0/0 drop:
Overhead  Shared Object                      Symbol
  49.82%  [kernel]                           [k] migrate_folio_unmap
  16.52%  [kernel]                           [k] clear_highpages_kasan_tagged
  13.56%  [kernel]                           [k] copy_mc_fragile
   1.80%  [kernel]                           [k] smp_call_function_many_cond
   1.06%  [kernel]                           [k] page_vma_mapped_walk
   0.95%  [kernel]                           [k] try_to_migrate_one
   0.94%  [kernel]                           [k] folio_migrate_flags
   0.71%  [kernel]                           [k] rmqueue_bulk
   0.61%  [kernel]                           [k] remove_migration_pte
   0.54%  [kernel]                           [k] __free_one_page
   0.50%  [kernel]                           [k] migrate_folio_move
   0.48%  [kernel]                           [k] folio_batch_move_lru
   0.47%  [kernel]                           [k]
__list_del_entry_valid_or_report


Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
33608407521
migrate_folio_unmap  /proc/kcore [Percent: local period]
Percent │       movq   %r13,0x20(%rsp)
   0.02 │       movq   %rdx,%r13
        │       movq   %r14,0x28(%rsp)
        │       movl   %r9d,%r14d
        │       movq   %r15,0x30(%rsp)
   0.01 │       movq   %r8,%r15
        │     → callq  *%rax
   0.04 │       testq  %rax,%rax
        │     ↓ je     2e3
   0.01 │       movq   %rbp,0x10(%rsp)
        │       movq   %rax,%rbp
   0.03 │       movq   %rax,(%r15)
        │       movq   $0x0,0x28(%rax)
  99.52 │       lock
        │       btsq   $0x0,(%rbx)
   0.01 │     ↓ jb     26b
   0.13 │ 67:   movq   (%rbx),%rdx
   0.01 │       testq  $0x2,(%rbx)
   0.01 │     ↓ je     e2
        │       cmpl   $0x2,%r14d
        │     ↓ je     d2
        │       movq   0x40(%rsp),%r8
        │       movq   %rbx,%rdi
        │       movl   $0x1,%ecx
        │       xorl   %edx,%edx

Lastly I used migratepages between two fast tiers (numa nodes 0 & 1)
and got 1.5 GB/s.
All this is too slow for the coming CXL era.

Anyway, I'm ready to proceed in assisting you in checking/testing if
you are also interested.

Anton

пт, 2 окт. 2026 г. в 11:34, SJ Park <sj@kernel.org>:
>
> On Thu, 1 Oct 2026 19:54:09 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > > That said, I'm personally curious if such fast migration is really needed in
> > > the real world workload.  I assume the real production workload would have
> > > stable access pattern but only occasionally get access pattern changes.
> > > Sometimes temporal and rapid access pattern change could also be made.  But for
> > > such temporal pattern change, doing migration would only be costy, since the
> > > data that suddenly hot could be soon be cold.  And reliability is important in
> > > production.  Hence I was thinking slowly making the balance is better than too
> > > quickly and reactively making migrations that will turn out to be no really
> > > needed.
> >
> > > That said, if you want to test faster migration speed, the multiple kdamonds
> > > usage could be one way to test.  If it turns out it is really needed and
> > > splitting address ranges has problems, we can consider adding DAMOS feature for
> > > utilizing multiple CPUs for faster migrations.
> >
> >
> > I agree with you!, but now it looks that demotion is too slow.
> > The server is completely idle, only I play with memory tiering.
> >
> > As CPU-less NUMA nodes I use Intel Optane PMEM configured as DAX devices,
> >
> > [root@localhost anton]# ndctl list
> > [
> >   {
> >     "dev":"namespace1.0",
> >     "mode":"devdax",
> >     "map":"dev",
> >     "size":3183575302144,
> >     "uuid":"865e9ee4-6334-4094-ac56-8b452bf4f53f",
> >     "chardev":"dax1.0",
> >     "align":2097152
> >   },
> >   {
> >     "dev":"namespace0.0",
> >     "mode":"devdax",
> >     "map":"dev",
> >     "size":3183575302144,
> >     "uuid":"ce77e90e-31bd-4057-8076-1f5d4374e9a2",
> >     "chardev":"dax0.0",
> >     "align":2097152
> >   }
> > ]
> > [root@localhost anton]#
> >
> > and then configured as system ram by next command,
> >
> > daxctl reconfigure-device --mode=system-ram all
> >
> > The problem is during demotion (to keep 25% free in the fast tier),
> > kdamond.0 utilizes single CPU core ~100%, but writing to the slower
> > tier (System PMM Write) ~1.4 GB/s what is much slower than defined
> > limit 8 GB/s.
>
> The defined limit is only upper limit, so real speed could be slower than that.
> FWIW, the example memory tiering script [1] uses 200 MiB/second as the
> upperlimit.  It was set by my gut feeling, not by some good data, though.
>
> I personally feel like ~1.4 GB/s is still a good speed for long-running
> production workloads, though I don't have a data to support that.  Do you have
> some data or reason to pursue 8 GB/s ?
>
> >
> >
> > top - 16:26:06 up 12 min,  2 users,  load average: 0.81, 0.75, 0.54
> > Tasks: 765 total, 2 running, 763 sleep, 0 d-sleep, 0 stopped, 0 zombie
> > %Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
> > MiB Mem : 6839822.+total, 6439081.+free, 415816.0 used,    465.2 buff/cache
> > MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6424006.+avail Mem
> >
> >     PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
> >    3227 root      20   0       0      0      0 R  99.8   0.0   1:28.92 kdamond.0
> >    2771 root      20   0 6932852  10260   5608 S   0.7   0.0   0:02.95
> > pcm-memory
> >    1593 root      20   0   17088   8352   7140 S   0.3   0.0   0:00.99
> > systemd-logind
> >    2263 root      20   0  300.4g 294.4g   6108 S   0.3   4.4   4:42.72
> > valkey-server
> >    3327 root      20   0   10916   6064   3764 R   0.3   0.0   0:00.16 top
> >       1 root      20   0   25356  15232  10456 S   0.0   0.0   0:02.13 systemd
> >       2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd
> >
> >
> >
> > |---------------------------------------||---------------------------------------|
> > |--            System DRAM Read Throughput(MB/s):        751.95
> >         --|
> > |--           System DRAM Write Throughput(MB/s):        742.58
> >         --|
> > |--             System PMM Read Throughput(MB/s):        693.12
> >         --|
> > |--            System PMM Write Throughput(MB/s):       1387.47
> >         --|
> > |--                 System Read Throughput(MB/s):       1445.08
> >         --|
> > |--                System Write Throughput(MB/s):       2130.05
> >         --|
> > |--               System Memory Throughput(MB/s):       3575.12
> >         --|
> > |---------------------------------------||---------------------------------------|
> >
> >
> >
> > Samples: 111K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> > 51312892466 lost: 0/0 drop
> > Overhead  Shared Object                      Symbol
> >   47.50%  [kernel]                           [k] migrate_folio_unmap
> >   16.11%  [kernel]                           [k] clear_highpages_kasan_tagged
> >   14.93%  [kernel]                           [k] copy_mc_fragile
> >    1.63%  [kernel]                           [k] smp_call_function_many_cond
> >    1.44%  [kernel]                           [k] page_vma_mapped_walk
> >    0.97%  [kernel]                           [k] folio_migrate_flags
> >    0.87%  [kernel]                           [k] try_to_migrate_one
> >    0.66%  [kernel]                           [k] rmqueue_bulk
> >    0.61%  [kernel]                           [k]
> > __list_del_entry_valid_or_report
> >    0.59%  [kernel]                           [k] remove_migration_pte
> >    0.50%  [kernel]                           [k] rmap_walk_anon
> >    0.45%  [kernel]                           [k] __free_one_page
> >    0.44%  [kernel]                           [k] migrate_folio_move
> >    0.44%  [kernel]                           [k] folio_add_anon_rmap_ptes
> >    0.40%  [kernel]                           [k] lru_gen_add_folio
> >    0.40%  [kernel]                           [k] __mod_memcg_lruvec_state
> >    0.39%  [kernel]                           [k] mod_node_page_state
> >    0.35%  [kernel]                           [k] migrate_pages_batch
> >    0.33%  [kernel]                           [k] _raw_spin_lock
> >    0.33%  [kernel]                           [k] up_read
> >
> >
> >
> > Samples: 222K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> > 64449579226
> > migrate_folio_unmap  /proc/kcore [Percent: local period]
> > Percent │       movq   %r13,0x20(%rsp)
> >    0.01 │       movq   %rdx,%r13
> >         │       movq   %r14,0x28(%rsp)
> >         │       movl   %r9d,%r14d
> >    0.00 │       movq   %r15,0x30(%rsp)
> >    0.01 │       movq   %r8,%r15
> >         │     → callq  *%rax
> >    0.16 │       testq  %rax,%rax
> >         │     ↓ je     2e3
> >    0.01 │       movq   %rbp,0x10(%rsp)
> >    0.01 │       movq   %rax,%rbp
> >    0.04 │       movq   %rax,(%r15)
> >    0.00 │       movq   $0x0,0x28(%rax)
> >   99.27 │       lock
> >         │       btsq   $0x0,(%rbx)
> >    0.01 │     ↓ jb     26b
> >    0.11 │ 67:   movq   (%rbx),%rdx
> >    0.01 │       testq  $0x2,(%rbx)
> >    0.01 │     ↓ je     e2
> >         │       cmpl   $0x2,%r14d
> >         │     ↓ je     d2
> >         │       movq   0x40(%rsp),%r8
> >         │       movq   %rbx,%rdi
> >         │       movl   $0x1,%ecx
> >         │       xorl   %edx,%edx
> >
> > I also tried without "--damos_filter reject young" and "--damos_filter
> > allow young", but got the same single CPU ~100% utilization and
> > demotion poor performance.
>
> Thank you for doing these great profiling and sharing the results, Anton.
> Apparently the bottleneck is not in DAMON specific code including the filtering
> part.  Instead, the bottleneck is in the migration code, which is out of DAMON,
> according to my understanding of the profiling result.  You could also confirm
> this by measuring the migration speed using move_pages() like system call,
> which also use the migration code.
>
> You may need to optimize the migration code, or make DAMON uses multiple CPUs
> if faster speed is really what needed.  I'd suggest using the address range
> split approach that I suggested in the previous reply.  If it becomes clear the
> faster migration is really needed, we could start thinking about making DAMOS
> use multiple CPUs without the address range split.
>
> [1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-02 10:02                                   ` Anton Gavriliuk
@ 2026-10-02 11:02                                     ` SJ Park
  2026-10-02 13:06                                       ` Anton Gavriliuk
  0 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-10-02 11:02 UTC (permalink / raw)
  To: Anton Gavriliuk; +Cc: SJ Park, damon

On Fri, 2 Oct 2026 13:02:13 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> Peak sequential write to slower tier ~12 GB/s.
> I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded.
> 
> For Memory Tiering with OLTP-like workloads needs to be
> promoted/demoted, 1.4 GB/s is OK.
> But for bandwidth intensive workloads such as analytics, AI, VDI,...
> 1.4 GB/s is too slow.

I think the factor deciding how fast migration should be is not bandwidth
intensiveness but how quickly access pattern changes.  That is, migration speed
should be fast enough to get costs from migration itself be paid back, and get
additional benefits.  That depends on the speed of the workloads' access
pattern change, and the amount of data of the changed pattern.  For example,
let's suppose 10 GiB data of a workload becomes hot, but will be cold again
after 10 seconds.  If we can migrate it to upper tier in one second, it will
get benefit from continued access for remaining 9 seconds.  If it takes 10
seconds to migrate, the benefit from the migration will be much lower.  It
might even lower than the migration work cost.

> 
> On the server where I test Memory Tiering, Intel CPU 8280L installed.
> Despite being already 7 years old, any single CPU core has 13-14 GB/s
> bandwidth access to local memory.

That maese sense to me.  Kernel level page migration requires not only the
content writes.  It also need to do additional works.  It should allocate pages
in destination node that the content will be copied to.  It should update
mappings and related metadata.  It should also handle possible races.  Hence
page migration is much more expensive and slow than pure I/O.

> 
> So from my point of view - even if we increase the number of cores
> keeping in mind 1.4 GB/s per core, we need so many cores to perform
> 10's GB/s demote/promote.

I agree.

> Firstly we need to improve demote/promote
> bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s.

I agree there could be workloads that could get benefit from faster migration.
And IIRC, there were a few people working on making migration faster and
lightweight.  I'm not an expert in the domain and not involved to such works
for now, though.  Nonetheless, having concrete data showing why 9-10 GB/s is
the right speed would be nice.

> 
> > You could also confirm
> > this by measuring the migration speed using move_pages() like system call,
> > which also use the migration code.
> 
> I'm not a developer and I need to spend more time thinking about how
> to do that, but I tried,
> 
> I used migratepages command, which is probably use migrate_pages()
> instead of move_pages(), but I got the same poor bandwidth and stack
> as with DAMON/DAMOS Memory Tiering,
> 
> top - 12:43:59 up 13 min,  2 users,  load average: 0.65, 0.88, 0.76
> Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie
> %Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
> MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used,    669.8 buff/cache
> MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6423508.+avail Mem
> 
>     PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
>    2935 root      20   0    2644   1764   1644 R  99.7   0.0   0:42.19
> migratepages
>    2941 root      20   0   10884   6192   3916 R   0.7   0.0   0:00.04 top
>    1238 systemd+  20   0   15860   6996   5952 S   0.3   0.0   0:00.37
> systemd-oomd
>    1928 root      20   0  666268  24448  23440 S   0.3   0.0   0:00.34 rsyslogd
>       1 root      20   0   25840  15532  10492 S   0.0   0.0   0:02.15 systemd
>       2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd
>       3 root      20   0       0      0      0 S   0.0   0.0   0:00.00
> pool_workqueue_release
>       4 root       0 -20       0      0      0 I   0.0   0.0   0:00.00
> kworker/R-rcu_gp
> 
> 
> |---------------------------------------||---------------------------------------|
> |--            System DRAM Read Throughput(MB/s):        759.41
>         --|
> |--           System DRAM Write Throughput(MB/s):        751.36
>         --|
> |--             System PMM Read Throughput(MB/s):        701.24
>         --|
> |--            System PMM Write Throughput(MB/s):       1402.15
>         --|
> |--                 System Read Throughput(MB/s):       1460.65
>         --|
> |--                System Write Throughput(MB/s):       2153.50
>         --|
> |--               System Memory Throughput(MB/s):       3614.15
>         --|
> |---------------------------------------||---------------------------------------|
> 
> 
> Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 48280162400 lost: 0/0 drop:
> Overhead  Shared Object                      Symbol
>   49.82%  [kernel]                           [k] migrate_folio_unmap
>   16.52%  [kernel]                           [k] clear_highpages_kasan_tagged
>   13.56%  [kernel]                           [k] copy_mc_fragile
>    1.80%  [kernel]                           [k] smp_call_function_many_cond
>    1.06%  [kernel]                           [k] page_vma_mapped_walk
>    0.95%  [kernel]                           [k] try_to_migrate_one
>    0.94%  [kernel]                           [k] folio_migrate_flags
>    0.71%  [kernel]                           [k] rmqueue_bulk
>    0.61%  [kernel]                           [k] remove_migration_pte
>    0.54%  [kernel]                           [k] __free_one_page
>    0.50%  [kernel]                           [k] migrate_folio_move
>    0.48%  [kernel]                           [k] folio_batch_move_lru
>    0.47%  [kernel]                           [k]
> __list_del_entry_valid_or_report
> 
> 
> Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> 33608407521
> migrate_folio_unmap  /proc/kcore [Percent: local period]
> Percent │       movq   %r13,0x20(%rsp)
>    0.02 │       movq   %rdx,%r13
>         │       movq   %r14,0x28(%rsp)
>         │       movl   %r9d,%r14d
>         │       movq   %r15,0x30(%rsp)
>    0.01 │       movq   %r8,%r15
>         │     → callq  *%rax
>    0.04 │       testq  %rax,%rax
>         │     ↓ je     2e3
>    0.01 │       movq   %rbp,0x10(%rsp)
>         │       movq   %rax,%rbp
>    0.03 │       movq   %rax,(%r15)
>         │       movq   $0x0,0x28(%rax)
>   99.52 │       lock
>         │       btsq   $0x0,(%rbx)
>    0.01 │     ↓ jb     26b
>    0.13 │ 67:   movq   (%rbx),%rdx
>    0.01 │       testq  $0x2,(%rbx)
>    0.01 │     ↓ je     e2
>         │       cmpl   $0x2,%r14d
>         │     ↓ je     d2
>         │       movq   0x40(%rsp),%r8
>         │       movq   %rbx,%rdi
>         │       movl   $0x1,%ecx
>         │       xorl   %edx,%edx
> 
> Lastly I used migratepages between two fast tiers (numa nodes 0 & 1)
> and got 1.5 GB/s.
> All this is too slow for the coming CXL era.

Thank you for testing this and sharing the results.  This confirms my previous
theory is not wrong.

> 
> Anyway, I'm ready to proceed in assisting you in checking/testing if
> you are also interested.

I think my thoery (bottleneck is in migration, not DAMON) is already proven.
We also confirmed DAMON and DAMOS are working as expected so far.  So I find no
more thing to test for now.  If you have more ideas to test, I will be more
than happy to help.


Thanks,
SJ

[...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: Memory tiering with DAMON/DAMOS auto-tuning
  2026-10-02 11:02                                     ` SJ Park
@ 2026-10-02 13:06                                       ` Anton Gavriliuk
  0 siblings, 0 replies; 21+ messages in thread
From: Anton Gavriliuk @ 2026-10-02 13:06 UTC (permalink / raw)
  To: SJ Park; +Cc: damon

> For example,
> let's suppose 10 GiB data of a workload becomes hot, but will be cold again
> after 10 seconds.  If we can migrate it to upper tier in one second, it will
> get benefit from continued access for remaining 9 seconds.  If it takes 10
> seconds to migrate, the benefit from the migration will be much lower.  It
> might even lower than the migration work cost.

Exactly right.  It fully matches the requirement to improve the
current 1.4 GB/s migration bandwidth.

So based on our discussions, I will try to discuss that with Memory
Management subsystem maintainers.

Anton

пт, 2 окт. 2026 г. в 14:02, SJ Park <sj@kernel.org>:
>
> On Fri, 2 Oct 2026 13:02:13 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:
>
> > Peak sequential write to slower tier ~12 GB/s.
> > I set the upper limit 8 GB/s just to keep those PMEM modules not overloaded.
> >
> > For Memory Tiering with OLTP-like workloads needs to be
> > promoted/demoted, 1.4 GB/s is OK.
> > But for bandwidth intensive workloads such as analytics, AI, VDI,...
> > 1.4 GB/s is too slow.
>
> I think the factor deciding how fast migration should be is not bandwidth
> intensiveness but how quickly access pattern changes.  That is, migration speed
> should be fast enough to get costs from migration itself be paid back, and get
> additional benefits.  That depends on the speed of the workloads' access
> pattern change, and the amount of data of the changed pattern.  For example,
> let's suppose 10 GiB data of a workload becomes hot, but will be cold again
> after 10 seconds.  If we can migrate it to upper tier in one second, it will
> get benefit from continued access for remaining 9 seconds.  If it takes 10
> seconds to migrate, the benefit from the migration will be much lower.  It
> might even lower than the migration work cost.
>
> >
> > On the server where I test Memory Tiering, Intel CPU 8280L installed.
> > Despite being already 7 years old, any single CPU core has 13-14 GB/s
> > bandwidth access to local memory.
>
> That maese sense to me.  Kernel level page migration requires not only the
> content writes.  It also need to do additional works.  It should allocate pages
> in destination node that the content will be copied to.  It should update
> mappings and related metadata.  It should also handle possible races.  Hence
> page migration is much more expensive and slow than pure I/O.
>
> >
> > So from my point of view - even if we increase the number of cores
> > keeping in mind 1.4 GB/s per core, we need so many cores to perform
> > 10's GB/s demote/promote.
>
> I agree.
>
> > Firstly we need to improve demote/promote
> > bandwidth for single CPU core, from 1.4 GB/s to at least 9-10 GB/s.
>
> I agree there could be workloads that could get benefit from faster migration.
> And IIRC, there were a few people working on making migration faster and
> lightweight.  I'm not an expert in the domain and not involved to such works
> for now, though.  Nonetheless, having concrete data showing why 9-10 GB/s is
> the right speed would be nice.
>
> >
> > > You could also confirm
> > > this by measuring the migration speed using move_pages() like system call,
> > > which also use the migration code.
> >
> > I'm not a developer and I need to spend more time thinking about how
> > to do that, but I tried,
> >
> > I used migratepages command, which is probably use migrate_pages()
> > instead of move_pages(), but I got the same poor bandwidth and stack
> > as with DAMON/DAMOS Memory Tiering,
> >
> > top - 12:43:59 up 13 min,  2 users,  load average: 0.65, 0.88, 0.76
> > Tasks: 757 total, 2 running, 755 sleep, 0 d-sleep, 0 stopped, 0 zombie
> > %Cpu(s):  0.0 us,  1.8 sy,  0.0 ni, 98.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
> > MiB Mem : 6839822.+total, 6438483.+free, 416313.2 used,    669.8 buff/cache
> > MiB Swap:   8192.0 total,   8192.0 free,      0.0 used. 6423508.+avail Mem
> >
> >     PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
> >    2935 root      20   0    2644   1764   1644 R  99.7   0.0   0:42.19
> > migratepages
> >    2941 root      20   0   10884   6192   3916 R   0.7   0.0   0:00.04 top
> >    1238 systemd+  20   0   15860   6996   5952 S   0.3   0.0   0:00.37
> > systemd-oomd
> >    1928 root      20   0  666268  24448  23440 S   0.3   0.0   0:00.34 rsyslogd
> >       1 root      20   0   25840  15532  10492 S   0.0   0.0   0:02.15 systemd
> >       2 root      20   0       0      0      0 S   0.0   0.0   0:00.02 kthreadd
> >       3 root      20   0       0      0      0 S   0.0   0.0   0:00.00
> > pool_workqueue_release
> >       4 root       0 -20       0      0      0 I   0.0   0.0   0:00.00
> > kworker/R-rcu_gp
> >
> >
> > |---------------------------------------||---------------------------------------|
> > |--            System DRAM Read Throughput(MB/s):        759.41
> >         --|
> > |--           System DRAM Write Throughput(MB/s):        751.36
> >         --|
> > |--             System PMM Read Throughput(MB/s):        701.24
> >         --|
> > |--            System PMM Write Throughput(MB/s):       1402.15
> >         --|
> > |--                 System Read Throughput(MB/s):       1460.65
> >         --|
> > |--                System Write Throughput(MB/s):       2153.50
> >         --|
> > |--               System Memory Throughput(MB/s):       3614.15
> >         --|
> > |---------------------------------------||---------------------------------------|
> >
> >
> > Samples: 92K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> > 48280162400 lost: 0/0 drop:
> > Overhead  Shared Object                      Symbol
> >   49.82%  [kernel]                           [k] migrate_folio_unmap
> >   16.52%  [kernel]                           [k] clear_highpages_kasan_tagged
> >   13.56%  [kernel]                           [k] copy_mc_fragile
> >    1.80%  [kernel]                           [k] smp_call_function_many_cond
> >    1.06%  [kernel]                           [k] page_vma_mapped_walk
> >    0.95%  [kernel]                           [k] try_to_migrate_one
> >    0.94%  [kernel]                           [k] folio_migrate_flags
> >    0.71%  [kernel]                           [k] rmqueue_bulk
> >    0.61%  [kernel]                           [k] remove_migration_pte
> >    0.54%  [kernel]                           [k] __free_one_page
> >    0.50%  [kernel]                           [k] migrate_folio_move
> >    0.48%  [kernel]                           [k] folio_batch_move_lru
> >    0.47%  [kernel]                           [k]
> > __list_del_entry_valid_or_report
> >
> >
> > Samples: 37K of event 'cpu/cycles/P', 4000 Hz, Event count (approx.):
> > 33608407521
> > migrate_folio_unmap  /proc/kcore [Percent: local period]
> > Percent │       movq   %r13,0x20(%rsp)
> >    0.02 │       movq   %rdx,%r13
> >         │       movq   %r14,0x28(%rsp)
> >         │       movl   %r9d,%r14d
> >         │       movq   %r15,0x30(%rsp)
> >    0.01 │       movq   %r8,%r15
> >         │     → callq  *%rax
> >    0.04 │       testq  %rax,%rax
> >         │     ↓ je     2e3
> >    0.01 │       movq   %rbp,0x10(%rsp)
> >         │       movq   %rax,%rbp
> >    0.03 │       movq   %rax,(%r15)
> >         │       movq   $0x0,0x28(%rax)
> >   99.52 │       lock
> >         │       btsq   $0x0,(%rbx)
> >    0.01 │     ↓ jb     26b
> >    0.13 │ 67:   movq   (%rbx),%rdx
> >    0.01 │       testq  $0x2,(%rbx)
> >    0.01 │     ↓ je     e2
> >         │       cmpl   $0x2,%r14d
> >         │     ↓ je     d2
> >         │       movq   0x40(%rsp),%r8
> >         │       movq   %rbx,%rdi
> >         │       movl   $0x1,%ecx
> >         │       xorl   %edx,%edx
> >
> > Lastly I used migratepages between two fast tiers (numa nodes 0 & 1)
> > and got 1.5 GB/s.
> > All this is too slow for the coming CXL era.
>
> Thank you for testing this and sharing the results.  This confirms my previous
> theory is not wrong.
>
> >
> > Anyway, I'm ready to proceed in assisting you in checking/testing if
> > you are also interested.
>
> I think my thoery (bottleneck is in migration, not DAMON) is already proven.
> We also confirmed DAMON and DAMOS are working as expected so far.  So I find no
> more thing to test for now.  If you have more ideas to test, I will be more
> than happy to help.
>
>
> Thanks,
> SJ
>
> [...]

^ permalink raw reply	[flat|nested] 21+ messages in thread

end of thread, other threads:[~2026-10-02 13:06 UTC | newest]

Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk
2026-09-28  8:15 ` SJ Park
2026-09-28 17:02   ` Anton Gavriliuk
2026-09-29  7:49     ` SJ Park
2026-09-29  8:41       ` Anton Gavriliuk
2026-09-29  9:25         ` SJ Park
2026-09-29 13:54           ` Anton Gavriliuk
2026-09-29 17:37             ` SJ Park
2026-09-30  4:03               ` Anton Gavriliuk
2026-09-30  8:35                 ` SJ Park
2026-09-30 16:26                   ` Anton Gavriliuk
2026-09-30 17:43                     ` SJ Park
2026-10-01  9:58                       ` Anton Gavriliuk
2026-10-01 10:26                         ` SJ Park
2026-10-01 12:08                           ` Anton Gavriliuk
2026-10-01 13:27                             ` SJ Park
2026-10-01 16:54                               ` Anton Gavriliuk
2026-10-02  8:34                                 ` SJ Park
2026-10-02 10:02                                   ` Anton Gavriliuk
2026-10-02 11:02                                     ` SJ Park
2026-10-02 13:06                                       ` Anton Gavriliuk

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox