Linux Perf Users
 help / color / mirror / Atom feed
* [PATCH] perf/core: Don't send SIGTRAP after exec removed the event
@ 2026-09-29 18:29 Danish Khateeb
  2026-09-29 18:46 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Danish Khateeb @ 2026-09-29 18:29 UTC (permalink / raw)
  To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
	Namhyung Kim
  Cc: Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
	Adrian Hunter, James Clark, Marco Elver, Frederic Weisbecker,
	linux-perf-users, linux-kernel, Danish Khateeb, stable

A sigtrap event must also set remove_on_exec, so that its SIGTRAP never
reaches a program after exec. But the signal is sent from task work,
which only runs on the way back to user space. If the event overflows
shortly before execve(), the task work can still be pending when the
task enters execve(), and then runs when execve() returns. By then
perf_event_exec() has removed the event and the new program has default
signal handlers, so the SIGTRAP kills it.

The exec_stress test in the remove_on_exec selftest catches this and
fails about half the time in a VM. A process that opens a sigtrap event
on itself and then calls execve() is killed by SIGTRAP in 15% of runs on
an AMD machine running v7.2, and in half of them in a VM.

perf_event_exit_event() sets PERF_EVENT_STATE_EXIT when it removes the
event, on exec and on exit. The exit case is already caught by the
PF_EXITING check in perf_sigtrap(), so an event in that state there was
removed by exec. Don't send the signal for it.

Fixes: 97ba62b27867 ("perf: Add support for SIGTRAP on perf events")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com>
---

Notes:
    Tested on v7.3-rc5 x86_64 under virtme-ng (KASAN, lockdep, 8 vCPUs on an
    AMD Zen 3 host):
    
    - selftests/perf_events/remove_on_exec: exec_stress failed in 10 of 20
      runs before, all 20 pass after.
    - The exec_stress pattern in a loop (30 inheriting children, 50 rounds):
      a child was killed by SIGTRAP in 20 rounds before, in none after.
    - Each child opens its own sigtrap + remove_on_exec event and execs:
      1078 of 2000 children were killed by SIGTRAP before, none after. The
      same program kills 297 of 2000 children on the bare-metal host running
      v7.2.6.
    - sigtrap_threads passes before and after. No new kernel warnings.
    
    v7.2 fails the same way in the VM. I did not test kernels older than
    v7.2. gcc W=1 and sparse show no new warnings.

 kernel/events/core.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/kernel/events/core.c b/kernel/events/core.c
index 634d2ccbab82..948583ffeb53 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -7631,6 +7631,14 @@ static void perf_sigtrap(struct perf_event *event)
 	if (current->flags & PF_EXITING)
 		return;
 
+	/*
+	 * exec() removed the event (remove_on_exec) after this signal was
+	 * queued. The new program has default signal handlers, so a SIGTRAP
+	 * would kill it.
+	 */
+	if (event->state == PERF_EVENT_STATE_EXIT)
+		return;
+
 	/*
 	 * We'd expect this to only occur if the irq_work is delayed and either
 	 * ctx->task or current has changed in the meantime. This can be the

base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH] perf/core: Don't send SIGTRAP after exec removed the event
  2026-09-29 18:29 [PATCH] perf/core: Don't send SIGTRAP after exec removed the event Danish Khateeb
@ 2026-09-29 18:46 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-29 18:46 UTC (permalink / raw)
  To: Danish Khateeb; +Cc: linux-perf-users

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] perf: strict event->state check allows fatal SIGTRAP on execve with O_CLOEXEC
--

commit 7195999db39d67f15f48d557a068f0dc7588f8ab
Author: Danish Khateeb <danishkhateeb03@gmail.com>

perf/core: Don't send SIGTRAP after exec removed the event

This commit prevents a SIGTRAP from reaching a newly executed program if an
event with remove_on_exec overflows right before execve(). It adds a check for
the PERF_EVENT_STATE_EXIT state in perf_sigtrap() to drop the signal if the
event was removed during exec.

> diff --git a/kernel/events/core.c b/kernel/events/core.c
> index 634d2ccbab82d..948583ffeb535 100644
> --- a/kernel/events/core.c
> +++ b/kernel/events/core.c
> @@ -7631,6 +7631,14 @@ static void perf_sigtrap(struct perf_event *event)
>  	if (current->flags & PF_EXITING)
>  		return;
>  
> +	/*
> +	 * exec() removed the event (remove_on_exec) after this signal was
> +	 * queued. The new program has default signal handlers, so a SIGTRAP
> +	 * would kill it.
> +	 */
> +	if (event->state == PERF_EVENT_STATE_EXIT)

[Severity: High]
Does this strict equality check account for the PERF_EVENT_STATE_DEAD state
caused by O_CLOEXEC processing?

If a perf event is configured with remove_on_exec=1, sigtrap=1, and the file
descriptor is opened with O_CLOEXEC, an event overflow queues
perf_pending_task() via task_work_add().

When the task begins execve(), do_close_on_exec() closes the O_CLOEXEC file
descriptor. fput() then queues ____fput via task_work_add(). Because the
task_works list executes in LIFO order (works are prepended to the head),
____fput is positioned to run before perf_pending_task().

After perf_event_exec() removes the event and sets 
event->state = PERF_EVENT_STATE_EXIT, execve() completes and task_work_run()
begins processing:

____fput()
   perf_release()
      perf_event_release_kernel()
         event->state = PERF_EVENT_STATE_DEAD;

When perf_pending_task() subsequently runs and calls perf_sigtrap(), the
event state is PERF_EVENT_STATE_DEAD (-5). Since -5 != PERF_EVENT_STATE_EXIT
(-3), the strict equality check bypasses the filter, erroneously sending
SIGTRAP and killing the new program.

Should this check for detached states via 
if (event->state <= PERF_EVENT_STATE_EXIT) instead?

> +		return;
> +
>  	/*
>  	 * We'd expect this to only occur if the irq_work is delayed and either
>  	 * ctx->task or current has changed in the meantime. This can be the

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260929182935.355892-1-danishkhateeb03@gmail.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-29 18:46 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 18:29 [PATCH] perf/core: Don't send SIGTRAP after exec removed the event Danish Khateeb
2026-09-29 18:46 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox