Linux Perf Users
 help / color / mirror / Atom feed
From: "Mi, Dapeng" <dapeng1.mi@linux.intel.com>
To: Namhyung Kim <namhyung@kernel.org>,
	Ravi Bangoria <ravi.bangoria@amd.com>
Cc: Gennady Kupava <gennady.kupava@gmail.com>,
	linux-perf-users@vger.kernel.org
Subject: Re: [DISCUSSION] Three problems behind broken "perf --call-graph dwarf" on AMD: IP and stack dump mismatch, libdw fails on lld's layout, and the unwinder fallback never runs
Date: Thu, 3 Sep 2026 16:39:29 +0800	[thread overview]
Message-ID: <c38b254f-21e5-4a02-8934-7a4da61e4571@linux.intel.com> (raw)
In-Reply-To: <apbs8gywyAfFVzqm@google.com>


On 9/1/2026 11:19 PM, Namhyung Kim wrote:
> Hello,
>
> On Mon, Aug 31, 2026 at 10:35:06AM +0530, Ravi Bangoria wrote:
>>>>> 2. perf: precise events and DWARF call graphs are mutually exclusive on this
>>>>>    hardware, so arguably perf should simply not let the two be combined.  I
>>>>>    would suggest two rules rather than one, because a blanket refusal would
>>>>>    make the common case worse:
>>>>>
>>>>>    - for the default event, drop the P when --call-graph dwarf is requested,
>>>>>      silently and on PMUs where precision means IBS.  Otherwise a plain
>>>>>      "perf record -g --call-graph dwarf" starts failing outright on every AMD
>>>>>      box, which is worse than today.  The s390 case in evlist.c suggests this
>>>>>      kind of substitution is considered acceptable;
>>>>>
>>>>>    - if the user asked for a precise event explicitly and also asked for
>>>>>      DWARF call chains, refuse with a message that says why and what to do,
>>>>>      instead of quietly producing a useless result.
>>>>>
>>>>>    Note that this should be scoped to IBS.  On Intel, PEBS records the whole
>>>>>    register set, so precise events and DWARF unwinding work together there
>>>>>    and nothing needs restricting.  Frame-pointer call graphs are also fine
>>>>>    with precise events - only the leaf frame is off.
>>>> I think we can enable both precise IP and dwarf callchains by using
>>>> PERF_SAMPLE_REGS_INTR.  For unwinding, it should use PERF_REG_X86_IP
>>>> from the REGS_INTR instead of sample.ip
>>> Ok, it seems we already do this in the perf tools but it looks like the
>>> kernel already overwrote the PERF_REG_X86_IP with the precise IP.  I
>>> feel like we should fix the kernel.
>> This should resolve the DWARF unwinding issue with IBS PMUs:
>>
>> --- a/arch/x86/events/amd/ibs.c
>> +++ b/arch/x86/events/amd/ibs.c
>> @@ -1523,8 +1523,11 @@ static int perf_ibs_handle_irq(struct perf_ibs *perf_ibs, struct pt_regs *iregs)
>>  			goto out;
>>  		}
>>  
>> -		set_linear_ip(&regs, ibs_data.regs[1]);
>> -		regs.flags |= PERF_EFLAGS_EXACT;
>> +		if (event->attr.sample_type & PERF_SAMPLE_IP) {
>> +			data.ip = ibs_data.regs[1];
>> +			data.sample_flags |= PERF_SAMPLE_IP;

We may not set PERF_SAMPLE_IP here, otherwise the kernel address leakage
check in perf_instruction_pointer() could be bypassed.


>> +			regs.flags |= PERF_EFLAGS_EXACT;
>> +		}
>>  	}
>>  
>>  	if (((ibs_caps & IBS_CAPS_BIT63_FILTER) ||
>> ---
> Right, that's what I thought.  And I believe we should do similar on
> Intel and not update other registers.

For Intel PEBS, the call-chain would always use the interrupt regs instead
of the PEBS regs, so there would be no issues for Intel PEBS.

    /*
     * We must however always use iregs for the unwinder to stay sane; the
     * record BP,SP,IP can point into thin air when the record is from a
     * previous PMI context or an (I)RET happened between the record and
     * PMI.
     */
    perf_sample_save_callchain(data, event, iregs);

Thanks.


>
>> However, there's no way to address the stack being out of sync with the
>> IBS RIP, since the IBS HW does not capture GPRs alongside the sample.
> I think it's ok and we don't need to sync IP and stack.  The dwarf
> unwind should start from stack and we can see the skid between IP and
> the first entry of the callchain.
>
> Thanks,
> Namhyung

  reply	other threads:[~2026-09-03  8:39 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-21 21:49 [DISCUSSION] Three problems behind broken "perf --call-graph dwarf" on AMD: IP and stack dump mismatch, libdw fails on lld's layout, and the unwinder fallback never runs Gennady Kupava
2026-08-24 13:35 ` Ravi Bangoria
2026-08-25 18:05 ` Ian Rogers
2026-08-27 21:55 ` Namhyung Kim
2026-08-28 16:47   ` Namhyung Kim
2026-08-28 16:58   ` Namhyung Kim
2026-08-31  5:05     ` Ravi Bangoria
2026-09-01 15:19       ` Namhyung Kim
2026-09-03  8:39         ` Mi, Dapeng [this message]
2026-09-03 11:36           ` Ravi Bangoria
2026-09-03 11:59             ` Mi, Dapeng
2026-09-03 16:02               ` Ravi Bangoria
2026-09-04  0:22                 ` Mi, Dapeng
2026-09-05 10:14                   ` Gennady Kupava

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c38b254f-21e5-4a02-8934-7a4da61e4571@linux.intel.com \
    --to=dapeng1.mi@linux.intel.com \
    --cc=gennady.kupava@gmail.com \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=namhyung@kernel.org \
    --cc=ravi.bangoria@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox