Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Huang, Ying" <ying.huang@linux.alibaba.com>
To: Shivank Garg <shivankg@amd.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,  Zi Yan <ziy@nvidia.com>,
	 Matthew Brost <matthew.brost@intel.com>,
	 Joshua Hahn <joshua.hahnjy@gmail.com>,
	 Rakie Kim <rakie.kim@sk.com>,  Byungchul Park <byungchul@sk.com>,
	 Gregory Price <gourry@gourry.net>,
	 "Alistair Popple" <apopple@nvidia.com>,
	 Lorenzo Stoakes <ljs@kernel.org>,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 "Mike Rapoport" <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 "Michal Hocko" <mhocko@suse.com>,
	 Karim Manaouil <kmanaouil.dev@gmail.com>,
	 Frank van der Linden <fvdl@google.com>,
	 Teja Vojjala <tejavojjala@google.com>,
	Pravin Tamkhane <pravintamkhane@google.com>,
	 Kinsey Ho <kinseyho@google.com>,  Wei Xu <weixugc@google.com>,
	 Matthew Wilcox <willy@infradead.org>,
	 Davidlohr Bueso <dave@stgolabs.net>,
	 Vinod Koul <vkoul@kernel.org>,  Bharata B Rao <bharata@amd.com>,
	 SeongJae Park <sj@kernel.org>,
	 David Rientjes <rientjes@google.com>,
	 Xuezheng Chu <xuezhengchu@huawei.com>,
	 "Yiannis Nikolakopoulos" <yiannis@zptcorp.com>,
	Dave Hansen <dave.hansen@intel.com>,
	 Johannes Weiner <hannes@cmpxchg.org>,
	 John Hubbard <jhubbard@nvidia.com>,
	 Peter Xu <peterx@redhat.com>,  Rik van Riel <riel@surriel.com>,
	 Shakeel Butt <shakeel.butt@linux.dev>,
	 Tejun Heo <tj@kernel.org>,  Fan Ni <nifan.cxl@gmail.com>,
	 Jonathan Cameron <jic23@kernel.org>,
	 Aneesh Kumar K.V <aneesh.kumar@kernel.org>,
	 Nathan Lynch <nathan.lynch@amd.com>, Frank Li <Frank.li@nxp.com>,
	 Dan Williams <djbw@kernel.org>, <linux-mm@kvack.org>,
	 <linux-kernel@vger.kernel.org>,  Mike Day <michael.day@amd.com>
Subject: Re: [PATCH RFC v6 0/5] Accelerate page migration with batch copying and hardware offload
Date: Tue, 14 Jul 2026 19:29:40 +0800	[thread overview]
Message-ID: <87wlux6c7v.fsf@DESKTOP-5N7EMDA> (raw)
In-Reply-To: <20260630-shivank-batch-migrate-offload-v6-0-da95d7e8b8a2@amd.com> (Shivank Garg's message of "Tue, 30 Jun 2026 07:28:36 +0000")

Hi, Garg,

Thanks for updated patch.

Shivank Garg <shivankg@amd.com> writes:

[snip]

>
> PERFORMANCE RESULTS:
> --------------------
>
> AMD EPYC 7713 (Zen 3), 2 sockets, 32 cores, SMT on, 
> 1 NUMA node per socket, 256 GB/node, v7.2-rc1, DVFS=Performance, PTDMA
> (16 DMA channels).
>
> Benchmark: move_pages() syscall to move pages between two NUMA nodes.
>
> 1). Moving different sized folios such that total transfer size is constant
> (1GB), with different number of DMA channels. Throughput in GB/s.
>
> a. Baseline (vanilla kernel, single-threaded, serial folio_copy):
> ================================================================================
> 4K          | 16K        | 64K        | 256K       | 1M          | 2M          |
> ================================================================================
> 3.28±0.14   | 4.98±0.18  | 6.19±0.08  | 6.77±0.08  | 7.02±0.11   | 10.80±0.13  |
>
> b. DMA offload (Patched Kernel, dcbm driver, N DMA channels):
> ============================================================================================
> N channel| 4K        | 16K         | 64K         | 256K        | 1M          | 2M          |
> ============================================================================================
> 1      | 2.38±0.17   | 2.77±0.03   | 3.21±0.03   | 5.00±0.02   | 5.09±0.64   | 12.62±0.07  |
> 2      | 2.87±0.11   | 4.06±0.05   | 5.09±0.04   | 6.97±0.08   | 8.43±0.06   | 14.32±0.10  |
> 4      | 3.32±0.07   | 5.30±0.06   | 7.21±0.09   | 9.69±0.15   | 11.36±0.13  | 26.98±0.19  |
> 8      | 3.68±0.09   | 6.28±0.10   | 9.16±0.13   | 12.05±0.16  | 15.33±2.80  | 46.06±0.55  |
> 12     | 3.83±0.05   | 6.65±0.17   | 10.00±0.16  | 12.98±0.18  | 15.87±0.19  | 61.31±1.28  |
> 16     | 3.94±0.09   | 6.78±0.10   | 10.48±0.13  | 13.48±0.20  | 16.90±0.24  | 65.06±2.46  |
>
> 2). First-folio latency: custom tracepoints (in migrate_pages_batch enter/exit,
>     migrate_folio_done) measure latency per migrate_pages_batch() call.
>
>     Throughput (GB/s) and first-folio latency (us), median of 10 runs.
>
> a. Vanilla Kernel:
>
> NR_MAX_BATCHED_MIGRATION upstream default value is 512.
> --- Order 0 (4K folios) ---       --- Order 9 (2M folios) ---
>      n      vanilla/cpu                n      vanilla/cpu
> (folios)    GB/s | first(us)      (folios)    GB/s | first(us)
> --------------------------        --------------------------
>      1       0.03 |     24             1       6.86 |    204
>      4       0.13 |     30             4       8.68 |    191
>      8       0.27 |     27             8       7.92 |    207
>     16       0.43 |     34            16       6.77 |    234
>     64       1.12 |     51            64      10.44 |    179
>    256       1.67 |    166           256      10.43 |    181
>    512       1.98 |    255           512      10.55 |    179
>   2048       2.38 |    233
>   4096       2.42 |    168
>  16384       2.72 |    167
>  65536       3.00 |    156
> 262144       3.10 |    151
>
> b. Patched kernel:
> N = NR_MAX_BATCHED_MIGRATION (in pages), Total migrated data fixed at
> 1 GB. Change N with knob (just for testing) to measure impact of
> different max batched size.
>
> --- ORDER 0 (4K folios) ---
>
>      N         offload/dma1          offload/dma4          offload/dma16
>                GB/s | first(us)      GB/s | first(us)      GB/s | first(us)
> ------------------------------------------------------------------------
>    512         2.21 |    628         3.29 |    275         3.25 |    245
>   1024         2.06 |   1271         3.21 |    601         3.36 |    518
>   2048         2.02 |   2646         3.00 |   1388         3.20 |   1110
>   4096         2.08 |   4832         3.17 |   2514         3.41 |   2175
>   8192         2.16 |   9253         3.14 |   4839         3.62 |   3592
>  16384         2.24 |  17543         3.23 |   9680         3.58 |   7144
>  32768         2.22 |  36408         3.26 |  19301         3.67 |  14524
>  65536         2.12 |  82572         3.24 |  38091         3.62 |  29835
> 131072         2.08 | 153669         3.17 |  79744         3.48 |  62157
> 262144         2.05 | 332297         2.97 | 175315         3.33 | 134774
>
> --- ORDER 9 (2M folios) ---
>
>      N         offload/dma1          offload/dma4          offload/dma16
>                GB/s | first(us)      GB/s | first(us)      GB/s | first(us)
> ------------------------------------------------------------------------
>    512        11.74 |    160        11.71 |    160        11.75 |    159
>   1024        12.18 |    310        13.82 |    274        13.76 |    275
>   2048        12.39 |    612        25.55 |    290        25.69 |    289
>   4096        12.54 |   1211        26.25 |    564        42.36 |    334
>   8192        12.54 |   2421        26.82 |   1111        51.85 |    485
>  16384        12.61 |   4824        26.91 |   2209        54.26 |    925
>  32768        12.62 |   9652        27.04 |   4404        54.72 |   1942
>  65536        12.64 |  19287        26.95 |   8835        57.30 |   3535
> 131072        12.64 |  38824        26.95 |  17900        58.58 |   7747
> 262144        12.66 |  77610        26.95 |  35743        66.31 |  13801
>
> OPEN QUESTION:
> --------------
>
> The best batch size depends on the hardware, and bigger isn't always better.
> NR_MAX_BATCHED_MIGRATION decides how many pages we move at once.
> Higher batch size can help amortize the setup cost of migrator but
> increases the first-folio latency (the folio is inaccessible for this
> window).

Yes.  This is an important parameter for the batched migration.  We
really need more information from people who have workload information
to make a decision.  Can you try to reach them?

Also, can we find a sweet point where we can archive higher throughput
without increasing latency too much (e.g., < 1ms)?

> Goals could be workload dependent, e.g. higher throughput versus
> same throughput under a bounded latency.
>
> Should this be tunable to accommodate different hardware and goals?

[snip]

---
Best Regards,
Huang, Ying


      parent reply	other threads:[~2026-07-14 11:29 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-30  7:28 [PATCH RFC v6 0/5] Accelerate page migration with batch copying and hardware offload Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 1/5] mm/migrate: skip data copy for already-copied folios Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 2/5] mm/migrate: add batch-copy path in migrate_pages_batch Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 3/5] mm/migrate: add copy offload registration infrastructure Shivank Garg
2026-07-20 11:19   ` Huang, Ying
2026-07-20 14:44     ` Zi Yan
2026-07-20 15:32       ` Garg, Shivank
2026-07-20 15:45         ` Zi Yan
2026-07-21 11:32       ` Huang, Ying
2026-06-30  7:28 ` [PATCH RFC v6 4/5] drivers/migrate_offload: add DMA batch copy driver (dcbm) Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 5/5] mm/migrate: adjust NR_MAX_BATCHED_MIGRATION for testing Shivank Garg
2026-07-14 11:29 ` Huang, Ying [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=87wlux6c7v.fsf@DESKTOP-5N7EMDA \
    --to=ying.huang@linux.alibaba.com \
    --cc=Frank.li@nxp.com \
    --cc=akpm@linux-foundation.org \
    --cc=aneesh.kumar@kernel.org \
    --cc=apopple@nvidia.com \
    --cc=bharata@amd.com \
    --cc=byungchul@sk.com \
    --cc=dave.hansen@intel.com \
    --cc=dave@stgolabs.net \
    --cc=david@kernel.org \
    --cc=djbw@kernel.org \
    --cc=fvdl@google.com \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=jhubbard@nvidia.com \
    --cc=jic23@kernel.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kinseyho@google.com \
    --cc=kmanaouil.dev@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=michael.day@amd.com \
    --cc=nathan.lynch@amd.com \
    --cc=nifan.cxl@gmail.com \
    --cc=peterx@redhat.com \
    --cc=pravintamkhane@google.com \
    --cc=rakie.kim@sk.com \
    --cc=riel@surriel.com \
    --cc=rientjes@google.com \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=shivankg@amd.com \
    --cc=sj@kernel.org \
    --cc=surenb@google.com \
    --cc=tejavojjala@google.com \
    --cc=tj@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=vkoul@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=xuezhengchu@huawei.com \
    --cc=yiannis@zptcorp.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox