Linux Perf Users
 help / color / mirror / Atom feed
* [PATCH 0/2] perf/arm: Prefer large AUX mappings for CoreSight and SPE
@ 2026-08-10 14:44 Leo Yan
  2026-08-10 14:44 ` [PATCH 1/2] coresight: perf: Prefer large AUX mappings Leo Yan
  2026-08-10 14:44 ` [PATCH 2/2] perf: arm_spe: " Leo Yan
  0 siblings, 2 replies; 5+ messages in thread
From: Leo Yan @ 2026-08-10 14:44 UTC (permalink / raw)
  To: Suzuki K Poulose, Will Deacon, Peter Zijlstra, Mike Leach,
	James Clark, Anshuman Khandual, Mark Rutland, Tamas Petz,
	Tamas Zsoldos, Michiel van Tol, Dev Jain, David Hildenbrand,
	Yabin Cui
  Cc: coresight, linux-arm-kernel, linux-kernel, linux-perf-users,
	Leo Yan

Since the commit:

 18049c8cff9 ("perf/aux: Allocate non-contiguous AUX pages by default")

it changed the AUX buffer allocator to allocate AUX pages page-by-page
(order=0) unless a PMU explicitly asks for contiguous allocations via
the capability flag PERF_PMU_CAP_AUX_PREFER_LARGE. The goal was to make
AUX allocation more memory-friendly by default, because not all PMUs
require physically contiguous AUX pages and large contiguous allocations
can contribute to fragmentation on long-running systems.

However, Arm SPE and CoreSight/TRBE rely on page-table translation when
writing trace data to memory. With page-by-page AUX allocation, a large
AUX buffer is mapped with many small mappings. This increases TLB
pressure, in practice this can increase trace-buffer latency due to
table translation walks (TTW) and contribute to trace discontinuities.

This series restores large AUX allocation for Arm CoreSight and SPE by
setting PERF_PMU_CAP_AUX_PREFER_LARGE.

This is intended to work together with the mm large-mapping series [1].
That series allows vmap() to map physically contiguous pages with
larger granules. With this series, perf first tries larger-order AUX
allocations, and the vmap() code can then create larger mappings for the
contiguous chunks.

The fragmentation concern from commit 18049c8cff9c should not block this
opt-in. PERF_PMU_CAP_AUX_PREFER_LARGE is a preference, not a hard
requirement. The AUX allocator already falls back to smaller orders when
a high-order allocation fails. So this series gives Arm trace PMUs the
performance benefit when large chunks are available.

The comparison below uses a baseline that already includes the mm
large-mapping series [1]. "Baseline" means that the mm series is applied
but this Arm PMU series is not. "Large AUX" means the same kernel plus
this series. Some configurations to mitigate noise during test:

  1) The tests were run with CPU10 isolated with the kernel parameter
     "isolcpus=10".
  2) CPU10 was used as the traced CPU, the PMU counter CPU, and the
     workload CPU. The perf control tasks were pinned to CPU2 so that
     they did not add extra work on CPU10.
  3) Each test was run for 10 iterations, and the tables report the
     average counter values across those runs.

The results show that using larger AUX mappings reduces the TLB pressure.
This is mainly visible in the refill events: CoreSight/TRBE shows a
large drop in l2d_tlb_refill and a smaller reduction in l1d_tlb_refill,
while SPE also reduces l2d_tlb_refill. The dtlb_walk event also drops in
both tests, which shows fewer data TLB walks after the AUX buffer can be
mapped with larger granules.

ETM sparse branch delay (cs_etm, AUX 1GB)

  taskset -c 2 perf stat -C 10 -e cycles:u,instructions:u,dtlb_walk:u,l1d_tlb:u,l1d_tlb_refill:u,l2d_tlb_refill:u \
    -- taskset -c 2 perf record -C 10 -m ,1G -e cs_etm// \
    -- taskset -c 10 ./sparse_branch_delay.elf

 |                | Baseline  | Large map |            |         |
 | Metric         | Avg.      | Avg.      | Delta      | Change  |
 |----------------+-----------+-----------+------------+---------|
 | dtlb_walk      |      72.8 |      63.9 |       -8.9 | -12.23% |
 | l1d_tlb        |   7,434.4 |   1,982.2 |   -5,452.2 | -73.34% |
 | l1d_tlb_refill |     163.7 |     148.2 |      -15.5 |  -9.47% |
 | l2d_tlb_refill | 161,884.9 |     513.1 | -161,371.8 | -99.68% |

SPE dd memory copy (arm_spe, AUX 512MB)

  taskset -c 2 perf stat -C 10 -e cycles:u,instructions:u,dtlb_walk:u,l1d_tlb:u,l1d_tlb_refill:u,l2d_tlb_refill:u \
    -- taskset -c 2 perf record -C 10 -m ,512M -e arm_spe_0/ts_enable=1,pa_enable=1,period=64,min_latency=0/ \
    -- taskset -c 10 dd if=/dev/zero of=/dev/shm/dd_mem_test bs=1M count=1024 status=progress

 |                | Baseline  | Large map |            |         |
 | Metric         | Avg.      | Avg.      | Delta      | Change  |
 |----------------+-----------+-----------+------------+---------|
 | dtlb_walk      |   1,760.2 |   1,387.9 |     -372.3 | -21.15% |
 | l1d_tlb        | 257,312.4 | 251,460.9 |   -5,851.5 |  -2.27% |
 | l1d_tlb_refill |  15,921.9 |  15,933.6 |       11.7 |  +0.07% |
 | l2d_tlb_refill |   4,285.0 |   2,796.5 |   -1,488.5 | -34.74% |

Note that after setting PREFER_LARGE for CoreSight and SPE, the existing
AUX trace drivers either prefer large pages or, in the case of Intel
BTS/PT, use the stronger AUX_NO_SG constraint. We can refactor this
later by either dropping PREFER_LARGE entirely or reversing the flag if
a driver needs discrete pages. For now, keep PREFER_LARGE to preserve
flexibility in the allocation policy.

[1] https://lore.kernel.org/linux-mm/20260715120813.3609949-1-jiangwen6@xiaomi.com/

Signed-off-by: Leo Yan <leo.yan@arm.com>
---
Dev Jain (1):
      coresight: perf: Prefer large AUX mappings

Leo Yan (1):
      perf: arm_spe: Prefer large AUX mappings

 drivers/hwtracing/coresight/coresight-etm-perf.c | 3 ++-
 drivers/perf/arm_spe_pmu.c                       | 3 ++-
 2 files changed, 4 insertions(+), 2 deletions(-)
---
base-commit: db2ddb87143519e20a95aa36c60b36107b736a58
change-id: 20260717-perf_aux_trace_large_granule-d9b30cc14b5a

Best regards,
-- 
Leo Yan <leo.yan@arm.com>


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH 1/2] coresight: perf: Prefer large AUX mappings
  2026-08-10 14:44 [PATCH 0/2] perf/arm: Prefer large AUX mappings for CoreSight and SPE Leo Yan
@ 2026-08-10 14:44 ` Leo Yan
  2026-08-10 14:44 ` [PATCH 2/2] perf: arm_spe: " Leo Yan
  1 sibling, 0 replies; 5+ messages in thread
From: Leo Yan @ 2026-08-10 14:44 UTC (permalink / raw)
  To: Suzuki K Poulose, Will Deacon, Peter Zijlstra, Mike Leach,
	James Clark, Anshuman Khandual, Mark Rutland, Tamas Petz,
	Tamas Zsoldos, Michiel van Tol, Dev Jain, David Hildenbrand,
	Yabin Cui
  Cc: coresight, linux-arm-kernel, linux-kernel, linux-perf-users,
	Leo Yan

From: Dev Jain <dev.jain@arm.com>

Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
default") changed AUX allocation to use order-0 pages by default unless
a PMU explicitly asks for contiguous allocations. That reduces
unnecessary memory fragmentation for PMUs which do not require larger
AUX chunks.

TRBE relies on page-table translation for writing the AUX buffer. If a
large AUX buffer is built from order-0 pages, vmap() has to map it with
many small mappings. This adds TLB pressure from the trace unit itself,
and can increase trace-buffer latency and contribute to trace
discontinuities.

Set PERF_PMU_CAP_AUX_PREFER_LARGE for the CoreSight PMU. This asks the
generic AUX allocator to try larger-order allocations so that, with
the vmap() large-mapping support, contiguous chunks can be mapped with
larger granules.

Apply the same preference to traditional sinks such as ETR. ETR uses
double buffering (a bounce buffer and an AUX buffer) and does not use
CPU page table when accessing the bounce buffer, so this does not
benefit TTW latency there. It can still help when the driver or perf
tool accesses the AUX buffer.

With the mm large-mapping series already applied, a sparse branch test
using a 1GB AUX buffer with TRBE showed the following results over 10
iterations:

  l1d_tlb_refill: 163.7 -> 148.2  (-9.47%)
  l2d_tlb_refill: 161,884.9 -> 513.1  (-99.68%)
  dtlb_walk: 72.8 -> 63.9  (-12.23%)

This shows the intended reduction in TLB refill pressure and TLB walks
once the AUX buffer can use larger mappings.

Signed-off-by: Dev Jain <dev.jain@arm.com>
Signed-off-by: Leo Yan <leo.yan@arm.com>
---
 drivers/hwtracing/coresight/coresight-etm-perf.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/drivers/hwtracing/coresight/coresight-etm-perf.c b/drivers/hwtracing/coresight/coresight-etm-perf.c
index 09b21a711a8764ea429d712890265c84648e889e..9646a1aab65b5b0b75c622bd18667ba0b916674f 100644
--- a/drivers/hwtracing/coresight/coresight-etm-perf.c
+++ b/drivers/hwtracing/coresight/coresight-etm-perf.c
@@ -1036,7 +1036,8 @@ int __init etm_perf_init(void)
 
 	etm_pmu.capabilities		= (PERF_PMU_CAP_EXCLUSIVE |
 					   PERF_PMU_CAP_ITRACE |
-					   PERF_PMU_CAP_AUX_PAUSE);
+					   PERF_PMU_CAP_AUX_PAUSE |
+					   PERF_PMU_CAP_AUX_PREFER_LARGE);
 
 	etm_pmu.attr_groups		= etm_pmu_attr_groups;
 	etm_pmu.task_ctx_nr		= perf_sw_context;

-- 
2.34.1


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings
  2026-08-10 14:44 [PATCH 0/2] perf/arm: Prefer large AUX mappings for CoreSight and SPE Leo Yan
  2026-08-10 14:44 ` [PATCH 1/2] coresight: perf: Prefer large AUX mappings Leo Yan
@ 2026-08-10 14:44 ` Leo Yan
  2026-08-10 15:10   ` Will Deacon
  1 sibling, 1 reply; 5+ messages in thread
From: Leo Yan @ 2026-08-10 14:44 UTC (permalink / raw)
  To: Suzuki K Poulose, Will Deacon, Peter Zijlstra, Mike Leach,
	James Clark, Anshuman Khandual, Mark Rutland, Tamas Petz,
	Tamas Zsoldos, Michiel van Tol, Dev Jain, David Hildenbrand,
	Yabin Cui
  Cc: coresight, linux-arm-kernel, linux-kernel, linux-perf-users,
	Leo Yan

Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
default") made the AUX allocator use order-0 pages by default unless a
PMU explicitly asks for contiguous allocations.

SPE writes trace data to the AUX buffer via virtual addresses and relies
on page-table translation. When a large AUX buffer is allocated with
order-0, the buffer is mapped with many small mappings, increasing TLB
pressure from the trace unit itself. This can add translation latency
while collecting trace.

Set PERF_PMU_CAP_AUX_PREFER_LARGE for Arm SPE. This lets the generic AUX
allocator try larger-order chunks first, which can then be mapped by
vmap() with larger granules when the mm large-mapping support is present.

With the mm large-mapping series already applied, dd memory copy test
using a 512MB AUX buffer with SPE showed the following results over 10
iterations:

  l1d_tlb_refill: 15,921.9 -> 15,933.6  (+0.07%)
  l2d_tlb_refill: 4,285.0 -> 2,796.5  (-34.74%)
  dtlb_walk: 1,760.2 -> 1,387.9  (-21.15%)

The main improvement is the lower L2 data TLB refill count, with fewer data
TLB walks as well.

Signed-off-by: Leo Yan <leo.yan@arm.com>
---
 drivers/perf/arm_spe_pmu.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/drivers/perf/arm_spe_pmu.c b/drivers/perf/arm_spe_pmu.c
index dbd0da1116390f71edf47c93db2f6fa3b36739d1..02389d3842216d55cd06e9174d2d27ade9a2ce4b 100644
--- a/drivers/perf/arm_spe_pmu.c
+++ b/drivers/perf/arm_spe_pmu.c
@@ -1064,7 +1064,8 @@ static int arm_spe_pmu_perf_init(struct arm_spe_pmu *spe_pmu)
 	spe_pmu->pmu = (struct pmu) {
 		.module = THIS_MODULE,
 		.parent		= &spe_pmu->pdev->dev,
-		.capabilities	= PERF_PMU_CAP_EXCLUSIVE | PERF_PMU_CAP_ITRACE,
+		.capabilities	= PERF_PMU_CAP_EXCLUSIVE | PERF_PMU_CAP_ITRACE |
+				  PERF_PMU_CAP_AUX_PREFER_LARGE,
 		.attr_groups	= arm_spe_pmu_attr_groups,
 		/*
 		 * We hitch a ride on the software context here, so that

-- 
2.34.1


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings
  2026-08-10 14:44 ` [PATCH 2/2] perf: arm_spe: " Leo Yan
@ 2026-08-10 15:10   ` Will Deacon
  2026-08-10 17:41     ` Leo Yan
  0 siblings, 1 reply; 5+ messages in thread
From: Will Deacon @ 2026-08-10 15:10 UTC (permalink / raw)
  To: Leo Yan
  Cc: Suzuki K Poulose, Peter Zijlstra, Mike Leach, James Clark,
	Anshuman Khandual, Mark Rutland, Tamas Petz, Tamas Zsoldos,
	Michiel van Tol, Dev Jain, David Hildenbrand, Yabin Cui,
	coresight, linux-arm-kernel, linux-kernel, linux-perf-users

On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
> Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
> default") made the AUX allocator use order-0 pages by default unless a
> PMU explicitly asks for contiguous allocations.

But that commit specifically calls out SPE as benefitting from
non-contiguous pages:

  "For instance, ARM SPE and TRBE operate with virtual pages, and
   Coresight ETR allocates a separate buffer. For these PMUs,
   allocating contiguous AUX pages unnecessarily exacerbates memory
   fragmentation. This fragmentation can prevent their use on
   long-running devices."

so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the
problems that 18049c8cff9c was trying to solve?

Will

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings
  2026-08-10 15:10   ` Will Deacon
@ 2026-08-10 17:41     ` Leo Yan
  0 siblings, 0 replies; 5+ messages in thread
From: Leo Yan @ 2026-08-10 17:41 UTC (permalink / raw)
  To: Will Deacon
  Cc: Suzuki K Poulose, Peter Zijlstra, Mike Leach, James Clark,
	Anshuman Khandual, Mark Rutland, Tamas Petz, Tamas Zsoldos,
	Michiel van Tol, Dev Jain, David Hildenbrand, Yabin Cui,
	coresight, linux-arm-kernel, linux-kernel, linux-perf-users

On Mon, Aug 10, 2026 at 04:10:48PM +0100, Will Deacon wrote:
> On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
> > Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
> > default") made the AUX allocator use order-0 pages by default unless a
> > PMU explicitly asks for contiguous allocations.
> 
> But that commit specifically calls out SPE as benefitting from
> non-contiguous pages:
> 
>   "For instance, ARM SPE and TRBE operate with virtual pages, and
>    Coresight ETR allocates a separate buffer. For these PMUs,
>    allocating contiguous AUX pages unnecessarily exacerbates memory
>    fragmentation. This fragmentation can prevent their use on
>    long-running devices."
> 
> so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the
> problems that 18049c8cff9c was trying to solve?

The question is how "allocating contiguous AUX pages unnecessarily
exacerbates memory fragmentation." The relevant information I could find
is [1]:

 "On Android, we collect ETM data periodically on internal user devices
  for AutoFDO optimization (for both userspace libraries and the
  kernel). Allocating a large chunk of contiguous AUX pages (4M for each
  CPU) periodically is almost unbearable. The kernel may need to kill
  many processes to fulfill the request. It affects user experience even
  after using PMU."

We might have missed chance to clarify how the fragmentation issue
occurs in the first place. Let's say, a phone with 8 CPUs, allocating
4MB per CPU requires 32MB in total, which is a relatively small
portion of 4GiB or 8GiB of RAM commonly found in phones. Moreover, once
contiguous pages are freed, the buddy allocator can coalesce them
again into buddy list. It is not obvious to me that PREFER_LARGE
directly causes fragmentation.

One case where AUX allocation could exacerbate fragmentation is when the
system is already fragmented. If a high-order allocation fails and the
allocator falls back to smaller-order blocks, those allocations may
consume free blocks scattered across different buddy regions and make
subsequent high-order allocations more difficult.

If this is the main concern, I'd suggest using a smaller AUX buffer
(e.g. 1MB or even 512KB) for TRBE/SPE to reduce memory pressure.
Snapshot mode '-S' could also be considered, as it allows the buffer to
be allocated once and reused for subsequent recordings by signals.

OTOH, using only order-0 pages can significantly increase TTW overhead
on the trace path and lead to overflows, we observe this causes huge
trace discontinuity. In the end, we need to trace-off the fragmentation
concern against the trace discontinuity.

Thanks,
Leo

[1] https://lore.kernel.org/lkml/CALJ9ZPNLgEBxOmDim-vztUknEETwdL-Z2gJ8K9s44TiPgKZgHg@mail.gmail.com/

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-08-10 17:42 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10 14:44 [PATCH 0/2] perf/arm: Prefer large AUX mappings for CoreSight and SPE Leo Yan
2026-08-10 14:44 ` [PATCH 1/2] coresight: perf: Prefer large AUX mappings Leo Yan
2026-08-10 14:44 ` [PATCH 2/2] perf: arm_spe: " Leo Yan
2026-08-10 15:10   ` Will Deacon
2026-08-10 17:41     ` Leo Yan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox