Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Jim Mattson <jmattson@google.com>
To: Manali Shukla <manali.shukla@amd.com>
Cc: Jim Mattson <jmattson@google.com>,
	seanjc@google.com, pbonzini@redhat.com,  mingo@redhat.com,
	bp@alien8.de, kvm@vger.kernel.org, x86@kernel.org,
	 santosh.shukla@amd.com, nikunj.dadhania@amd.com,
	Naveen.Rao@amd.com,  dapeng1.mi@linux.intel.com,
	ravi.bangoria@amd.com, peterz@infradead.org,
	 Sandipan.Das@amd.com, Yosry Ahmed <yosry@kernel.org>,
	linux-perf-users@vger.kernel.org
Subject: Re: [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS
Date: Thu,  8 Oct 2026 10:19:18 -0700	[thread overview]
Message-ID: <20261008171921.668587-1-jmattson@google.com> (raw)
In-Reply-To: <20260310060022.15120-2-manali.shukla@amd.com>

On Tue, Mar 10, 2026 at 06:00:13AM +0000, Manali Shukla wrote:
> Consider the following scenario,
>
> While scheduling out an IBS event from perf's core scheduling path,
> event_sched_out() disables the IBS event by clearing the IBS enable
> bit in perf_ibs_disable_event(). However, if a delayed IBS NMI is
> delivered after the IBS enable bit is cleared, the IBS NMI handler
> may still observe the valid bit set and incorrectly treat the sample
> as valid.

The sample is valid. It was collected while the event was active, and
it is correct to record it. The bug is only that the handler re-arms
the hardware after perf_ibs_stop() disables the event.

> As a result, it re-enables IBS by setting the enable bit,
> even though the event has already been scheduled out.
>
> This leads to a situation where IBS is re-enabled after being
> explicitly disabled, which is incorrect. Although this race does not
> have visible side effects, it violates the expected behavior of the
> perf subsystem.

This race does have visible side effects:

  1. When the delayed NMI arrives before perf_ibs_stop() clears
     IBS_STARTED, the handler takes the normal path (not the fail:
     path), leaves IBS_STOPPED set, and re-arms the hardware for one
     more period. When that extra period overflows after
     perf_ibs_stop() clears IBS_STARTED, a second NMI arrives. If an
     unrelated NMI arrives first, the IBS handler takes the fail:
     path, clears IBS_STOPPED, and claims that NMI. The second IBS
     NMI is then unhandled ("Uhhuh. NMI received for unknown
     reason").
  2. With VIBS enabled, on hardware without IBS_CAPS_DIS, if this
     race happens when perf schedules out a host IBS event before
     VMRUN, IbsFetchEn or IbsOpEn is 1 at VMRUN. APM vol. 2, section
     15.38, says that these bits must be 0 at VMRUN of an SEV-ES or
     SEV-SNP guest with IBS virtualization enabled.

> The race is particularly noticeable when userspace repeatedly disables
> and re-enables IBS using PERF_EVENT_IOC_DISABLE and
> PERF_EVENT_IOC_ENABLE ioctls in a loop.
>
> Fix this by checking the IBS_STOPPING bit in the IBS NMI handler before
> re-enabling the IBS event. If the IBS_STOPPING bit is set, it indicates
> that the event is either disabled or in the process of being disabled,
> and the NMI handler should not re-enable it.
>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>

I think this warrants a Fixes tag:

Fixes: 85dc600263c2 ("perf/x86/amd/ibs: Fix pmu::stop() nesting")

This fix does not depend on VIBS. It is probably better to send it
separately through tip/perf:core, so that it can go in before the
rest of this series.

> ---
>  arch/x86/events/amd/ibs.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
> index eeb607b84dda..09b56bab510a 100644
> --- a/arch/x86/events/amd/ibs.c
> +++ b/arch/x86/events/amd/ibs.c
> @@ -1582,7 +1582,8 @@ static int perf_ibs_handle_irq(struct perf_ibs *perf_ibs, struct pt_regs *iregs)
>  		}
>  		new_config |= period >> 4;
>
> -		perf_ibs_enable_event(perf_ibs, hwc, new_config);
> +		if (!test_bit(IBS_STOPPING, pcpu->state))
> +			perf_ibs_enable_event(perf_ibs, hwc, new_config);

This stops the late re-arm. perf_ibs_stop() sets IBS_STOPPING first,
with test_and_set_bit(). An NMI before that point can re-arm the
hardware, but perf_ibs_stop() then disables the hardware. An NMI
after that point does not re-arm. Both sides run on the same CPU, so
a plain test_bit() in NMI context is sufficient.

However, the skipped re-arm causes three problems when an NMI arrives
after perf_ibs_stop() sets IBS_STOPPING and before it clears
IBS_STARTED.

First, the event count can increase twice for the same sample:

  1. The handler calls perf_ibs_event_update() for the sample and
     adds a full period. perf_ibs_set_period() sets prev_count to 0.
  2. Because of this patch, the handler does not re-arm. CTL (and
     perf_ibs_stop()'s local config copy, if already read) still
     holds the old sample with Val=1, and the handler does not set
     PERF_HES_UPTODATE.
  3. perf_ibs_stop() clears Val in its config copy and calls
     perf_ibs_event_update() again, because PERF_HES_UPTODATE is
     clear. Because prev_count is now 0, the delta is the whole
     count field: CurCnt (op) or FetchCnt (fetch) from the old
     sample, as if it were progress in a new period.

The amount added in step 3 depends on what the hardware leaves in the
count fields after a sample. For op, the comment in
get_ibs_op_count() says that the lower 7 bits of CurCnt are
randomized after a rollover, so the amount is in general not zero.

Second, IBS_STOPPED can stay set in pcpu->state. perf_ibs_stop() sets
IBS_STOPPED so that a late NMI can clear it at the fail: label. When
the NMI instead arrives before perf_ibs_stop() clears IBS_STARTED,
the handler takes the normal path and does not clear IBS_STOPPED.
Before this patch, the re-armed period gave a second NMI that cleared
it. Now that the handler does not re-arm, IBS_STOPPED can stay set
and falsely claim a later unrelated NMI.

Third, the same sample can be recorded twice. After the handler skips
the re-arm, CTL still holds the old sample with Val=1, and
IBS_STARTED is still set. This is true at least until
perf_ibs_stop() calls perf_ibs_disable_event(). With IBS_CAPS_DIS,
that call writes only CTL2, so it stays true until perf_ibs_stop()
clears IBS_STARTED. If an unrelated NMI arrives in this window, the
IBS handler takes the normal path again, records the same sample a
second time, and adds another full period, because prev_count is 0.
Before this patch, the re-arm cleared Val, so this could not happen.

The throttle path (throttle != 0, so the handler does not re-arm) has
the first two problems already, when perf_event_overflow() calls
pmu::stop() from the NMI. They are not caused by this patch, but they
may be worth a look at the same time.

>  	}
>
>  	perf_event_update_userpage(event);

  reply	other threads:[~2026-10-08 17:19 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
2026-10-08 17:19   ` Jim Mattson [this message]
2026-03-10  6:00 ` [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests Manali Shukla
2026-10-08 17:28   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities Manali Shukla
2026-10-08 17:44   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf Manali Shukla
2026-10-08 17:59   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers Manali Shukla
2026-10-08 18:01   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization Manali Shukla
2026-10-08 19:19   ` Jim Mattson
2026-10-08 21:28     ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU Manali Shukla
2026-10-08 19:32   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init() Manali Shukla
2026-10-08 19:38   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs Manali Shukla
2026-10-08 21:03   ` Jim Mattson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008171921.668587-1-jmattson@google.com \
    --to=jmattson@google.com \
    --cc=Naveen.Rao@amd.com \
    --cc=Sandipan.Das@amd.com \
    --cc=bp@alien8.de \
    --cc=dapeng1.mi@linux.intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=manali.shukla@amd.com \
    --cc=mingo@redhat.com \
    --cc=nikunj.dadhania@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=ravi.bangoria@amd.com \
    --cc=santosh.shukla@amd.com \
    --cc=seanjc@google.com \
    --cc=x86@kernel.org \
    --cc=yosry@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox