All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Huang, Ying" <ying.huang@linux.alibaba.com>
To: Shivank Garg <shivankg@amd.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,  Zi Yan <ziy@nvidia.com>,
	 Matthew Brost <matthew.brost@intel.com>,
	 Joshua Hahn <joshua.hahnjy@gmail.com>,
	 Rakie Kim <rakie.kim@sk.com>,  Byungchul Park <byungchul@sk.com>,
	 Gregory Price <gourry@gourry.net>,
	 "Alistair Popple" <apopple@nvidia.com>,
	 Lorenzo Stoakes <ljs@kernel.org>,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 "Mike Rapoport" <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 "Michal Hocko" <mhocko@suse.com>,
	 Karim Manaouil <kmanaouil.dev@gmail.com>,
	 Frank van der Linden <fvdl@google.com>,
	 Teja Vojjala <tejavojjala@google.com>,
	Pravin Tamkhane <pravintamkhane@google.com>,
	 Kinsey Ho <kinseyho@google.com>,  Wei Xu <weixugc@google.com>,
	 Matthew Wilcox <willy@infradead.org>,
	 Davidlohr Bueso <dave@stgolabs.net>,
	 Vinod Koul <vkoul@kernel.org>,  Bharata B Rao <bharata@amd.com>,
	 SeongJae Park <sj@kernel.org>,
	 David Rientjes <rientjes@google.com>,
	 Xuezheng Chu <xuezhengchu@huawei.com>,
	 "Yiannis Nikolakopoulos" <yiannis@zptcorp.com>,
	Dave Hansen <dave.hansen@intel.com>,
	 Johannes Weiner <hannes@cmpxchg.org>,
	 John Hubbard <jhubbard@nvidia.com>,
	 Peter Xu <peterx@redhat.com>,  Rik van Riel <riel@surriel.com>,
	 Shakeel Butt <shakeel.butt@linux.dev>,
	 Tejun Heo <tj@kernel.org>,  Fan Ni <nifan.cxl@gmail.com>,
	 Jonathan Cameron <jic23@kernel.org>,
	 Aneesh Kumar K.V <aneesh.kumar@kernel.org>,
	 Nathan Lynch <nathan.lynch@amd.com>, Frank Li <Frank.li@nxp.com>,
	 Dan Williams <djbw@kernel.org>, <linux-mm@kvack.org>,
	 <linux-kernel@vger.kernel.org>,  Mike Day <michael.day@amd.com>
Subject: Re: [PATCH RFC v6 0/5] Accelerate page migration with batch copying and hardware offload
Date: Tue, 14 Jul 2026 19:29:40 +0800	[thread overview]
Message-ID: <87wlux6c7v.fsf@DESKTOP-5N7EMDA> (raw)
In-Reply-To: <20260630-shivank-batch-migrate-offload-v6-0-da95d7e8b8a2@amd.com> (Shivank Garg's message of "Tue, 30 Jun 2026 07:28:36 +0000")

Hi, Garg,

Thanks for updated patch.

Shivank Garg <shivankg@amd.com> writes:

[snip]

>
> PERFORMANCE RESULTS:
> --------------------
>
> AMD EPYC 7713 (Zen 3), 2 sockets, 32 cores, SMT on, 
> 1 NUMA node per socket, 256 GB/node, v7.2-rc1, DVFS=Performance, PTDMA
> (16 DMA channels).
>
> Benchmark: move_pages() syscall to move pages between two NUMA nodes.
>
> 1). Moving different sized folios such that total transfer size is constant
> (1GB), with different number of DMA channels. Throughput in GB/s.
>
> a. Baseline (vanilla kernel, single-threaded, serial folio_copy):
> ================================================================================
> 4K          | 16K        | 64K        | 256K       | 1M          | 2M          |
> ================================================================================
> 3.28±0.14   | 4.98±0.18  | 6.19±0.08  | 6.77±0.08  | 7.02±0.11   | 10.80±0.13  |
>
> b. DMA offload (Patched Kernel, dcbm driver, N DMA channels):
> ============================================================================================
> N channel| 4K        | 16K         | 64K         | 256K        | 1M          | 2M          |
> ============================================================================================
> 1      | 2.38±0.17   | 2.77±0.03   | 3.21±0.03   | 5.00±0.02   | 5.09±0.64   | 12.62±0.07  |
> 2      | 2.87±0.11   | 4.06±0.05   | 5.09±0.04   | 6.97±0.08   | 8.43±0.06   | 14.32±0.10  |
> 4      | 3.32±0.07   | 5.30±0.06   | 7.21±0.09   | 9.69±0.15   | 11.36±0.13  | 26.98±0.19  |
> 8      | 3.68±0.09   | 6.28±0.10   | 9.16±0.13   | 12.05±0.16  | 15.33±2.80  | 46.06±0.55  |
> 12     | 3.83±0.05   | 6.65±0.17   | 10.00±0.16  | 12.98±0.18  | 15.87±0.19  | 61.31±1.28  |
> 16     | 3.94±0.09   | 6.78±0.10   | 10.48±0.13  | 13.48±0.20  | 16.90±0.24  | 65.06±2.46  |
>
> 2). First-folio latency: custom tracepoints (in migrate_pages_batch enter/exit,
>     migrate_folio_done) measure latency per migrate_pages_batch() call.
>
>     Throughput (GB/s) and first-folio latency (us), median of 10 runs.
>
> a. Vanilla Kernel:
>
> NR_MAX_BATCHED_MIGRATION upstream default value is 512.
> --- Order 0 (4K folios) ---       --- Order 9 (2M folios) ---
>      n      vanilla/cpu                n      vanilla/cpu
> (folios)    GB/s | first(us)      (folios)    GB/s | first(us)
> --------------------------        --------------------------
>      1       0.03 |     24             1       6.86 |    204
>      4       0.13 |     30             4       8.68 |    191
>      8       0.27 |     27             8       7.92 |    207
>     16       0.43 |     34            16       6.77 |    234
>     64       1.12 |     51            64      10.44 |    179
>    256       1.67 |    166           256      10.43 |    181
>    512       1.98 |    255           512      10.55 |    179
>   2048       2.38 |    233
>   4096       2.42 |    168
>  16384       2.72 |    167
>  65536       3.00 |    156
> 262144       3.10 |    151
>
> b. Patched kernel:
> N = NR_MAX_BATCHED_MIGRATION (in pages), Total migrated data fixed at
> 1 GB. Change N with knob (just for testing) to measure impact of
> different max batched size.
>
> --- ORDER 0 (4K folios) ---
>
>      N         offload/dma1          offload/dma4          offload/dma16
>                GB/s | first(us)      GB/s | first(us)      GB/s | first(us)
> ------------------------------------------------------------------------
>    512         2.21 |    628         3.29 |    275         3.25 |    245
>   1024         2.06 |   1271         3.21 |    601         3.36 |    518
>   2048         2.02 |   2646         3.00 |   1388         3.20 |   1110
>   4096         2.08 |   4832         3.17 |   2514         3.41 |   2175
>   8192         2.16 |   9253         3.14 |   4839         3.62 |   3592
>  16384         2.24 |  17543         3.23 |   9680         3.58 |   7144
>  32768         2.22 |  36408         3.26 |  19301         3.67 |  14524
>  65536         2.12 |  82572         3.24 |  38091         3.62 |  29835
> 131072         2.08 | 153669         3.17 |  79744         3.48 |  62157
> 262144         2.05 | 332297         2.97 | 175315         3.33 | 134774
>
> --- ORDER 9 (2M folios) ---
>
>      N         offload/dma1          offload/dma4          offload/dma16
>                GB/s | first(us)      GB/s | first(us)      GB/s | first(us)
> ------------------------------------------------------------------------
>    512        11.74 |    160        11.71 |    160        11.75 |    159
>   1024        12.18 |    310        13.82 |    274        13.76 |    275
>   2048        12.39 |    612        25.55 |    290        25.69 |    289
>   4096        12.54 |   1211        26.25 |    564        42.36 |    334
>   8192        12.54 |   2421        26.82 |   1111        51.85 |    485
>  16384        12.61 |   4824        26.91 |   2209        54.26 |    925
>  32768        12.62 |   9652        27.04 |   4404        54.72 |   1942
>  65536        12.64 |  19287        26.95 |   8835        57.30 |   3535
> 131072        12.64 |  38824        26.95 |  17900        58.58 |   7747
> 262144        12.66 |  77610        26.95 |  35743        66.31 |  13801
>
> OPEN QUESTION:
> --------------
>
> The best batch size depends on the hardware, and bigger isn't always better.
> NR_MAX_BATCHED_MIGRATION decides how many pages we move at once.
> Higher batch size can help amortize the setup cost of migrator but
> increases the first-folio latency (the folio is inaccessible for this
> window).

Yes.  This is an important parameter for the batched migration.  We
really need more information from people who have workload information
to make a decision.  Can you try to reach them?

Also, can we find a sweet point where we can archive higher throughput
without increasing latency too much (e.g., < 1ms)?

> Goals could be workload dependent, e.g. higher throughput versus
> same throughput under a bounded latency.
>
> Should this be tunable to accommodate different hardware and goals?

[snip]

---
Best Regards,
Huang, Ying


      parent reply	other threads:[~2026-07-14 11:29 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-30  7:28 [PATCH RFC v6 0/5] Accelerate page migration with batch copying and hardware offload Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 1/5] mm/migrate: skip data copy for already-copied folios Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 2/5] mm/migrate: add batch-copy path in migrate_pages_batch Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 3/5] mm/migrate: add copy offload registration infrastructure Shivank Garg
2026-07-20 11:19   ` Huang, Ying
2026-07-20 14:44     ` Zi Yan
2026-07-20 15:32       ` Garg, Shivank
2026-07-20 15:45         ` Zi Yan
2026-07-21 11:32       ` Huang, Ying
2026-06-30  7:28 ` [PATCH RFC v6 4/5] drivers/migrate_offload: add DMA batch copy driver (dcbm) Shivank Garg
2026-06-30  7:28 ` [PATCH RFC v6 5/5] mm/migrate: adjust NR_MAX_BATCHED_MIGRATION for testing Shivank Garg
2026-07-14 11:29 ` Huang, Ying [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=87wlux6c7v.fsf@DESKTOP-5N7EMDA \
    --to=ying.huang@linux.alibaba.com \
    --cc=Frank.li@nxp.com \
    --cc=akpm@linux-foundation.org \
    --cc=aneesh.kumar@kernel.org \
    --cc=apopple@nvidia.com \
    --cc=bharata@amd.com \
    --cc=byungchul@sk.com \
    --cc=dave.hansen@intel.com \
    --cc=dave@stgolabs.net \
    --cc=david@kernel.org \
    --cc=djbw@kernel.org \
    --cc=fvdl@google.com \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=jhubbard@nvidia.com \
    --cc=jic23@kernel.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kinseyho@google.com \
    --cc=kmanaouil.dev@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=michael.day@amd.com \
    --cc=nathan.lynch@amd.com \
    --cc=nifan.cxl@gmail.com \
    --cc=peterx@redhat.com \
    --cc=pravintamkhane@google.com \
    --cc=rakie.kim@sk.com \
    --cc=riel@surriel.com \
    --cc=rientjes@google.com \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=shivankg@amd.com \
    --cc=sj@kernel.org \
    --cc=surenb@google.com \
    --cc=tejavojjala@google.com \
    --cc=tj@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=vkoul@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=xuezhengchu@huawei.com \
    --cc=yiannis@zptcorp.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.