All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Mayank Rungta" <mrungta@google.com>
Cc: linux-watchdog@vger.kernel.org
Subject: Re: [PATCH v2 1/5] genirq: Synchronize in-flight handlers during NMI teardown
Date: Sat, 29 Aug 2026 01:14:49 +0000	[thread overview]
Message-ID: <20260829011449.87AF11F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260828-qcom-wdt-nmi-series-v2-1-363979fe6b6b@google.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

New issues:
- [High] Splitting the atomic NMI teardown in `__cleanup_nmi()` without holding `desc->request_mutex` introduces a critical race condition with concurrent `request_nmi()` calls.

Pre-existing issues:
- [Critical] Setting `desc->action = NULL` before synchronizing in-flight NMIs creates a fatal race window where concurrent NMIs dereference a NULL pointer.
--

commit c7c9e712a226c1915918d074d948b4cf1845990f
Author: Mayank Rungta <mrungta@google.com>

genirq: Synchronize in-flight handlers during NMI teardown

This patch modifies __cleanup_nmi() to wait for in-flight handlers using
__synchronize_hardirq() before tearing down the NMI configuration and
deactivating the domain. It also enforces that interrupt controllers claiming
NMI support must implement ->irq_get_irqchip_state().

> diff --git a/kernel/irq/manage.c b/kernel/irq/manage.c
> index 2fbff2618a1e2..61384925b921a 100644
> --- a/kernel/irq/manage.c
> +++ b/kernel/irq/manage.c

[ ... ]

> @@ -2047,11 +2051,20 @@ static const void *__cleanup_nmi(unsigned int irq, struct irq_desc *desc)
>  		desc->action = NULL;
>  
>  		irq_settings_clr_disable_unlazy(desc);
> -		irq_shutdown_and_deactivate(desc);
> +		irq_shutdown(desc);
>  	}
>  
>  	irq_proc_update_valid(desc);
>  
> +	/* Ensure all in-flight NMI handlers on other CPUs complete before freeing action */
> +	__synchronize_hardirq(desc, true);

[Severity: Critical]
This is a pre-existing issue, but does setting desc->action to NULL before
synchronizing in-flight NMIs expose a fatal race window?

If an NMI is delivered to a CPU after desc->action is set to NULL, but before
the hardware line is disabled (or if an NMI is already in flight), the NMI
handler could unconditionally dereference the nullified action pointer.

Unlike normal IRQ handlers, handle_fasteoi_nmi() assumes the action is
always valid:

kernel/irq/chip.c:handle_fasteoi_nmi() {
    ...
    /*
     * NMIs cannot be shared, there is only one action.
     */
    res = action->handler(irq, action->dev_id);
    ...
}

The attempt to fix the use-after-free via __synchronize_hardirq() appears
ineffective here because the pointer is nullified before the synchronization
barrier.

> +
> +	scoped_guard(raw_spinlock_irqsave, &desc->lock) {
> +		irq_nmi_teardown(desc);
> +		desc->istate &= ~IRQS_NMI;
> +		irq_domain_deactivate_irq(&desc->irq_data);
> +	}

[Severity: High]
Does splitting the atomic NMI teardown without holding desc->request_mutex
introduce a race condition with concurrent request_nmi() calls?

Since the lock is yielded to wait in __synchronize_hardirq(), a concurrent
request_nmi() could execute. Unlike free_irq(), __cleanup_nmi() does not hold
desc->request_mutex. This means request_nmi() could acquire the mutex,
configure the new NMI, and attach the action.

When __cleanup_nmi() resumes in this block, would it blindly execute
irq_nmi_teardown() and irq_domain_deactivate_irq(), inadvertently destroying
the newly requested hardware state and leaving the new NMI broken?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260828-qcom-wdt-nmi-series-v2-0-363979fe6b6b@google.com?part=1

  reply	other threads:[~2026-08-29  1:14 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-29  0:59 [PATCH v2 0/5] watchdog: qcom: Support NMI pretimeout warnings Mayank Rungta
2026-08-29  0:59 ` [PATCH v2 1/5] genirq: Synchronize in-flight handlers during NMI teardown Mayank Rungta
2026-08-29  1:14   ` sashiko-bot [this message]
2026-08-31 18:16   ` Doug Anderson
2026-08-29  0:59 ` [PATCH v2 2/5] genirq: Implement synchronous disable_nmi() Mayank Rungta
2026-08-29  1:18   ` sashiko-bot
2026-08-31 18:16   ` Doug Anderson
2026-08-29  0:59 ` [PATCH v2 3/5] genirq: Export NMI APIs Mayank Rungta
2026-08-29  1:15   ` sashiko-bot
2026-08-31 18:17   ` Doug Anderson
2026-08-29  0:59 ` [PATCH v2 4/5] watchdog: pretimeout: Protect governor access with RCU for NMI safety Mayank Rungta
2026-08-31 18:17   ` Doug Anderson
2026-08-29  0:59 ` [PATCH v2 5/5] watchdog: qcom: Register pretimeout interrupt as NMI Mayank Rungta
2026-08-29  1:15   ` sashiko-bot
2026-08-31 18:17   ` Doug Anderson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260829011449.87AF11F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-watchdog@vger.kernel.org \
    --cc=mrungta@google.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.