Linux Perf Users
 help / color / mirror / Atom feed
* [PATCH] perf/x86/amd/lbr: Compress branch stack in a single pass
@ 2026-10-07 14:44 Sandipan Das
  2026-10-07 14:56 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Sandipan Das @ 2026-10-07 14:44 UTC (permalink / raw)
  To: linux-perf-users, linux-kernel
  Cc: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
	Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
	Ian Rogers, Adrian Hunter, James Clark, Thomas Gleixner,
	Borislav Petkov, Dave Hansen, x86, H . Peter Anvin, Ravi Bangoria,
	Ananth Narayan, Sandipan Das

Software filtering removes rejected branch records by shifting the
entire tail of the array down by one for each dropped entry. With k
invalid entries out of n, this costs O(k*n) copies. Switch to a two
pointer compaction that walks the array once and packs the remaining
entries towards the front, moving each entry at most once, for O(n)
total. No functional change intended.

Signed-off-by: Sandipan Das <sandipan.das@amd.com>
---
 arch/x86/events/amd/lbr.c | 17 +++++++----------
 1 file changed, 7 insertions(+), 10 deletions(-)

diff --git a/arch/x86/events/amd/lbr.c b/arch/x86/events/amd/lbr.c
index a55646fcb846..1e1b2f08cb51 100644
--- a/arch/x86/events/amd/lbr.c
+++ b/arch/x86/events/amd/lbr.c
@@ -141,17 +141,14 @@ static void amd_pmu_lbr_filter(void)
 		return;
 
 	/* Remove all invalid entries */
-	for (i = 0; i < cpuc->lbr_stack.nr; ) {
-		if (!cpuc->lbr_entries[i].from) {
-			j = i;
-			while (++j < cpuc->lbr_stack.nr)
-				cpuc->lbr_entries[j - 1] = cpuc->lbr_entries[j];
-			cpuc->lbr_stack.nr--;
-			if (!cpuc->lbr_entries[i].from)
-				continue;
-		}
-		i++;
+	for (i = 0, j = 0; i < cpuc->lbr_stack.nr; i++) {
+		if (!cpuc->lbr_entries[i].from)
+			continue;
+		if (j != i)
+			cpuc->lbr_entries[j] = cpuc->lbr_entries[i];
+		j++;
 	}
+	cpuc->lbr_stack.nr = j;
 }
 
 static const int lbr_spec_map[PERF_BR_SPEC_MAX] = {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH] perf/x86/amd/lbr: Compress branch stack in a single pass
  2026-10-07 14:44 [PATCH] perf/x86/amd/lbr: Compress branch stack in a single pass Sandipan Das
@ 2026-10-07 14:56 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-07 14:56 UTC (permalink / raw)
  To: Sandipan Das; +Cc: linux-perf-users

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: info leak of uncleared LBR tail entries via branch snapshot API
--

commit ed8227763c0f1dd19117eba0f7a4a321639e9390
Author: Sandipan Das <sandipan.das@amd.com>

perf/x86/amd/lbr: Compress branch stack in a single pass

This commit optimizes the software filtering of branch records by packing the
remaining valid entries towards the front of the array in a single pass,
rather than shifting the tail down for each dropped entry.

> diff --git a/arch/x86/events/amd/lbr.c b/arch/x86/events/amd/lbr.c
> index 9d9c961989d51..abb636d48df09 100644
> --- a/arch/x86/events/amd/lbr.c
> +++ b/arch/x86/events/amd/lbr.c
> @@ -141,17 +141,14 @@ static void amd_pmu_lbr_filter(void)
>  		return;
>  
>  	/* Remove all invalid entries */
> -	for (i = 0; i < cpuc->lbr_stack.nr; ) {
> -		if (!cpuc->lbr_entries[i].from) {
> -			j = i;
> -			while (++j < cpuc->lbr_stack.nr)
> -				cpuc->lbr_entries[j - 1] = cpuc->lbr_entries[j];
> -			cpuc->lbr_stack.nr--;
> -			if (!cpuc->lbr_entries[i].from)
> -				continue;
> -		}
> -		i++;
> +	for (i = 0, j = 0; i < cpuc->lbr_stack.nr; i++) {
> +		if (!cpuc->lbr_entries[i].from)
> +			continue;
> +		if (j != i)
> +			cpuc->lbr_entries[j] = cpuc->lbr_entries[i];
> +		j++;
>  	}
> +	cpuc->lbr_stack.nr = j;

[Severity: Medium]
Does this code leave stale entries at the tail of the cpuc->lbr_entries array?

When a privileged BPF program calls bpf_get_branch_snapshot(), it invokes
perf_snapshot_branch_stack(), which calls amd_pmu_v2_snapshot_branch_stack().

Looking at amd_pmu_v2_snapshot_branch_stack() in arch/x86/events/amd/core.c,
it unconditionally copies cnt (up to x86_pmu.lbr_nr) entries instead of
cpuc->lbr_stack.nr:

	memcpy(entries, cpuc->lbr_entries, sizeof(struct perf_branch_entry) * cnt);

Since amd_pmu_lbr_filter() drops invalid entries and shrinks
cpuc->lbr_stack.nr but doesn't clear the tail elements, will these uncleared
tail elements (which could contain filtered branches or branches from other
contexts) be leaked to BPF programs?

Additionally, if cpuc->lbr_users == 0, the LBRs are not read and the entire
stale array from a previous task could be leaked.

Should the remaining entries from j to the original cpuc->lbr_stack.nr be
zeroed out here, or should the snapshot APIs be updated to only copy
cpuc->lbr_stack.nr entries?

Note that Intel's BPF snapshot implementation
(__intel_pmu_snapshot_branch_stack in arch/x86/events/intel/core.c) and LBR
filter (intel_pmu_lbr_filter in arch/x86/events/intel/lbr.c) appear to have a
similar behavior where nr is shrunk without clearing tail elements.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261007144416.3446108-1-sandipan.das@amd.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-07 14:56 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 14:44 [PATCH] perf/x86/amd/lbr: Compress branch stack in a single pass Sandipan Das
2026-10-07 14:56 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox