DAMON development mailing list
 help / color / mirror / Atom feed
From: SJ Park <sj@kernel.org>
To: Anton Gavriliuk <antosha20xx@gmail.com>
Cc: SJ Park <sj@kernel.org>, damon@lists.linux.dev
Subject: Re: Memory tiering with DAMON/DAMOS auto-tuning
Date: Mon, 28 Sep 2026 01:15:46 -0700	[thread overview]
Message-ID: <20260928081546.39695-1-sj@kernel.org> (raw)
In-Reply-To: <CAAiJnjp5F8ZuPG1gEsD_Wgs7Z+GToz67nzjr0490sz9y328=dg@mail.gmail.com>

Hello Anton,

On Sun, 27 Sep 2026 20:26:21 +0300 Anton Gavriliuk <antosha20xx@gmail.com> wrote:

> Hello
> 
> I would like to play with memory tiering with DAMON/DAMOS auto-tuning
> (between numa nodes 0 & 2) based on valkey and memtier_benchmark.

Thank you for sharing your use case and question!

> The goal - keep numa node 0 50% free and promote cold pages from numa
> node 2 to numa node 0 immediately when they become hot.
> This is Fedora Server 44 up-to-date with the 7.2.8 kernel.
> 
> [root@localhost ~]# numactl -H
> available: 4 nodes (0-3)
> node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
> 22 23 24 25 26 27
> node 0 size: 386590 MB
> node 0 free: 337165 MB
> node 1 cpus: 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
> 47 48 49 50 51 52 53 54 55
> node 1 size: 387055 MB
> node 1 free: 338204 MB
> node 2 cpus:
> node 2 size: 3033088 MB
> node 2 free: 3032998 MB
> node 3 cpus:
> node 3 size: 3033088 MB
> node 3 free: 3032998 MB
> node distances:
> node     0    1    2    3
>    0:   10   20   25   35
>    1:   20   10   35   25
>    2:   25   35   10   20
>    3:   35   25   20   10
> [root@localhost ~]#
> [root@localhost ~]# /home/anton/Linux/mlc
> Intel(R) Memory Latency Checker - v3.13
> Measuring idle latencies for sequential access (in ns)...
>                 Numa node
> Numa node            0       1       2       3
>        0          82.3   147.6   173.8   238.0
>        1         146.7    82.0   238.7   172.2
> 
> 
> The Goal:
> 
> 1. There are two tiers, fast tier - numa node 0; slow tier - numa node 2
> 2. Keep fast tier 50% free
> 3. Demote valkey inactive >=5 min pages from fast to slow tier
> 4. Promote valkey hot pages from slow tier to fast tier immediately
> when accessed
> 5. Demote and Promote bandwidth performance limit up to 8 GB/s; CPU
> performance unlimited
> 6. Monitoring intervals for Demote/Promote 500ms
> 
> What I already done -
> 
> Launch Valkey Pinned to Node 0 DRAM
> 
> numactl --cpunodebind=0 --preferred=0 valkey-server \
>   --port 6379 \
>   --protected-mode no \
>   --save "" \
>   --appendonly no \
>   --maxmemory 300gb \
>   --maxmemory-policy noeviction &
> 
> 
> Command to Load ~200+ GB
> 
> memtier_benchmark \
>   -s 127.0.0.1 -p 6379 \
>   -t 16 -c 16 \
>   -d 10240 \
>   --ratio=1:0 \
>   --key-pattern=P:P \
>   --distinct-client-seed \
>   --key-maximum=20000000 \
>   --pipeline=32 \
>   -n allkeys
> 
> 
> [root@localhost anton]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.11            0.00            0.00
> Stack                        0.03            0.00            0.00
> Private                 238356.61            0.46            4.82
> ----------------  --------------- --------------- ---------------
> Total                   238356.75            0.46            4.82
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       238361.90
> ----------------  --------------- ---------------
> Total                        0.00       238362.04
> 
> 
> damo start \--numa_node 0 --monitoring_intervals_goal 97% 3 5ms 10s

The example memory tiering script [1] uses 4% as intervals goal.  Is there a
reason to use 97% as the goal instead?

> \--damos_action migrate_cold 2 --damos_access_rate 0% 0%
> \--damos_apply_interval 1s \--damos_quota_interval 1s
> --damos_quota_space 8GB \--damos_quota_goal node_mem_free_bp 50% 0
> \--damos_filter reject young \

You mentioned you want to demote >=5 minutes inactive pages.  But, the above
command doesn't have '--damos_age' option.  It means DAMON will demote node 0
pages as soon as it finds it was not accessed, even if it was not accessed for
<5 minutes.  You could let DAMON know you want to demote only >=5 minutes
inactive pages by adding '--damos_age 5m max' option.

> --numa_node 2
> --monitoring_intervals_goal 97% 3 5ms 10s \--damos_action migrate_hot

Again, I'm curious why you use 97% goal.

> 0 --damos_access_rate 5% max \--damos_apply_interval 1s
> \--damos_quota_interval 1s --damos_quota_space 8GB \--damos_quota_goal
> node_mem_used_bp 99.7% 0 \--damos_filter allow young

Is the above node_mum_used_bp what you really want?  That means you want to
promote hot pages from node 2 to node 0, until the node 0 memory utilization
becomes 99.7%.  That overlaps with the demotion goal (50% free memory of node
0) quite a lot.  I'd suggest smaller overlap, say, 50.3%, to keep healthy
circulation of hot/cold pages while not consuming too much resource under
stabilized access pattern.

> \--damos_nr_quota_goals 1 1 --damos_nr_filters 1 1 \--nr_targets 1 1
> --nr_schemes 1 1 --nr_ctxs 1 1
> 
> And it is demoted much more, almost all pages than the goal of 50%
> free numa node 0.
> 
> [root@localhost anton]# numastat -p $(pgrep valkey-server)
> 
> Per-node process memory usage (in MBs) for PID 5631 (valkey-server)
>                            Node 0          Node 1          Node 2
>                   --------------- --------------- ---------------
> Huge                         0.00            0.00            0.00
> Heap                         0.02            0.00            0.09
> Stack                        0.02            0.00            0.01
> Private                  24389.88            0.46       213971.55
> ----------------  --------------- --------------- ---------------
> Total                    24389.92            0.46       213971.65
> 
>                            Node 3           Total
>                   --------------- ---------------
> Huge                         0.00            0.00
> Heap                         0.00            0.11
> Stack                        0.00            0.03
> Private                      0.00       238361.90
> ----------------  --------------- ---------------
> Total                        0.00       238362.04
> 
> I just want that demotion process to stop when there is 50% free at numa node 0.
> 
> Where am I wrong, how to fix that ?

DAMOS quota auto-tuning supports two tuner algorithms [2], consistent and
temporal.  As the documentation [2] explains, 'consistent' tuner assumes it
should keep applying the action in a level to keep the goal achieved.  In this
case, for example, the demotion scheme assumes there will be continued
promotion and therefore it keeps demoting.  If there is no appropriate
promotion, it could result in demoting more than expected amount.

To me, this seems like the workload has no much hot data, so demotion is much
more stronger.

For path forward, I'd suggest trying 'temporal' tuner [2].  It is designed to
immediately stop after achieving the goal.

If it still doesn't work, I'd suggest tracing damon:damos_esz tracepoint with
the numastat output.  It will show if the autotune is working as expected.

For long term production use case, I'd like to suggest running the example
tiering config [1] and see if it also gives you unexpected results.

[1] https://github.com/damonitor/damo/blob/next/scripts/mem_tier.sh
[2] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#aim-oriented-feedback-driven-auto-tuning


Thanks,
SJ

[...]

  reply	other threads:[~2026-09-28  8:15 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-27 17:26 Memory tiering with DAMON/DAMOS auto-tuning Anton Gavriliuk
2026-09-28  8:15 ` SJ Park [this message]
2026-09-28 17:02   ` Anton Gavriliuk
2026-09-29  7:49     ` SJ Park
2026-09-29  8:41       ` Anton Gavriliuk
2026-09-29  9:25         ` SJ Park
2026-09-29 13:54           ` Anton Gavriliuk
2026-09-29 17:37             ` SJ Park
2026-09-30  4:03               ` Anton Gavriliuk
2026-09-30  8:35                 ` SJ Park
2026-09-30 16:26                   ` Anton Gavriliuk
2026-09-30 17:43                     ` SJ Park
2026-10-01  9:58                       ` Anton Gavriliuk
2026-10-01 10:26                         ` SJ Park
2026-10-01 12:08                           ` Anton Gavriliuk
2026-10-01 13:27                             ` SJ Park
2026-10-01 16:54                               ` Anton Gavriliuk
2026-10-02  8:34                                 ` SJ Park
2026-10-02 10:02                                   ` Anton Gavriliuk
2026-10-02 11:02                                     ` SJ Park
2026-10-02 13:06                                       ` Anton Gavriliuk

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928081546.39695-1-sj@kernel.org \
    --to=sj@kernel.org \
    --cc=antosha20xx@gmail.com \
    --cc=damon@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox