From: David Hildenbrand <david@redhat.com>
To: Andrew Morton <akpm@linux-foundation.org>, yangge1116@126.com
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
21cnbao@gmail.com, baolin.wang@linux.alibaba.com,
aisheng.dong@nxp.com, liuzixing@hygon.cn
Subject: Re: [PATCH V2] mm/cma: using per-CMA locks to improve concurrent allocation performance
Date: Tue, 18 Mar 2025 14:02:21 +0100 [thread overview]
Message-ID: <24ca760c-a861-4797-a434-d91a59513b12@redhat.com> (raw)
In-Reply-To: <20250317204325.99b45373023ad2f901c1152e@linux-foundation.org>
On 18.03.25 04:43, Andrew Morton wrote:
> On Mon, 10 Feb 2025 09:56:06 +0800 yangge1116@126.com wrote:
>
>> From: yangge <yangge1116@126.com>
>>
>> For different CMAs, concurrent allocation of CMA memory ideally should not
>> require synchronization using locks. Currently, a global cma_mutex lock is
>> employed to synchronize all CMA allocations, which can impact the
>> performance of concurrent allocations across different CMAs.
>>
>> To test the performance impact, follow these steps:
>> 1. Boot the kernel with the command line argument hugetlb_cma=30G to
>> allocate a 30GB CMA area specifically for huge page allocations. (note:
>> on my machine, which has 3 nodes, each node is initialized with 10G of
>> CMA)
>> 2. Use the dd command with parameters if=/dev/zero of=/dev/shm/file bs=1G
>> count=30 to fully utilize the CMA area by writing zeroes to a file in
>> /dev/shm.
>> 3. Open three terminals and execute the following commands simultaneously:
>> (Note: Each of these commands attempts to allocate 10GB [2621440 * 4KB
>> pages] of CMA memory.)
>> On Terminal 1: time echo 2621440 > /sys/kernel/debug/cma/hugetlb1/alloc
>> On Terminal 2: time echo 2621440 > /sys/kernel/debug/cma/hugetlb2/alloc
>> On Terminal 3: time echo 2621440 > /sys/kernel/debug/cma/hugetlb3/alloc
>>
>> We attempt to allocate pages through the CMA debug interface and use the
>> time command to measure the duration of each allocation.
>> Performance comparison:
>> Without this patch With this patch
>> Terminal1 ~7s ~7s
>> Terminal2 ~14s ~8s
>> Terminal3 ~21s ~7s
>>
>> To slove problem above, we could use per-CMA locks to improve concurrent
>> allocation performance. This would allow each CMA to be managed
>> independently, reducing the need for a global lock and thus improving
>> scalability and performance.
>
> This patch was in and out of mm-unstable for a while, as Frank's series
> "hugetlb/CMA improvements for large systems" was being added and
> dropped.
>
> Consequently it hasn't received any testing for a while.
>
> Below is the version which I've now re-added to mm-unstable. Can
> you please check this and retest it?
>
> Thanks.
>
> From: Ge Yang <yangge1116@126.com>
> Subject: mm/cma: using per-CMA locks to improve concurrent allocation performance
> Date: Mon, 10 Feb 2025 09:56:06 +0800
>
> For different CMAs, concurrent allocation of CMA memory ideally should not
> require synchronization using locks. Currently, a global cma_mutex lock
> is employed to synchronize all CMA allocations, which can impact the
> performance of concurrent allocations across different CMAs.
>
> To test the performance impact, follow these steps:
> 1. Boot the kernel with the command line argument hugetlb_cma=30G to
> allocate a 30GB CMA area specifically for huge page allocations. (note:
> on my machine, which has 3 nodes, each node is initialized with 10G of
> CMA)
> 2. Use the dd command with parameters if=/dev/zero of=/dev/shm/file bs=1G
> count=30 to fully utilize the CMA area by writing zeroes to a file in
> /dev/shm.
> 3. Open three terminals and execute the following commands simultaneously:
> (Note: Each of these commands attempts to allocate 10GB [2621440 * 4KB
> pages] of CMA memory.)
> On Terminal 1: time echo 2621440 > /sys/kernel/debug/cma/hugetlb1/alloc
> On Terminal 2: time echo 2621440 > /sys/kernel/debug/cma/hugetlb2/alloc
> On Terminal 3: time echo 2621440 > /sys/kernel/debug/cma/hugetlb3/alloc
>
> We attempt to allocate pages through the CMA debug interface and use the
> time command to measure the duration of each allocation.
> Performance comparison:
> Without this patch With this patch
> Terminal1 ~7s ~7s
> Terminal2 ~14s ~8s
> Terminal3 ~21s ~7s
>
> To solve problem above, we could use per-CMA locks to improve concurrent
> allocation performance. This would allow each CMA to be managed
> independently, reducing the need for a global lock and thus improving
> scalability and performance.
>
> Link: https://lkml.kernel.org/r/1739152566-744-1-git-send-email-yangge1116@126.com
> Signed-off-by: Ge Yang <yangge1116@126.com>
> Reviewed-by: Barry Song <baohua@kernel.org>
> Acked-by: David Hildenbrand <david@redhat.com>
> Reviewed-by: Oscar Salvador <osalvador@suse.de>
> Cc: Aisheng Dong <aisheng.dong@nxp.com>
> Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: David Hildenbrand <david@redhat.com>
--
Cheers,
David / dhildenb
prev parent reply other threads:[~2025-03-18 13:02 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-02-10 1:56 [PATCH V2] mm/cma: using per-CMA locks to improve concurrent allocation performance yangge1116
2025-02-10 3:44 ` Barry Song
2025-02-10 8:34 ` David Hildenbrand
2025-02-10 8:56 ` Ge Yang
2025-02-10 9:15 ` Oscar Salvador
2025-03-18 3:43 ` Andrew Morton
2025-03-18 7:21 ` Ge Yang
2025-03-18 13:02 ` David Hildenbrand [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=24ca760c-a861-4797-a434-d91a59513b12@redhat.com \
--to=david@redhat.com \
--cc=21cnbao@gmail.com \
--cc=aisheng.dong@nxp.com \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=liuzixing@hygon.cn \
--cc=yangge1116@126.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.