* [RFC PATCH 0/1] perf scripts flamegraph: Add --asm option
@ 2026-08-21 8:50 Tudor-Stefan Magirescu
2026-08-21 8:50 ` [RFC PATCH 1/1] " Tudor-Stefan Magirescu
0 siblings, 1 reply; 3+ messages in thread
From: Tudor-Stefan Magirescu @ 2026-08-21 8:50 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, linux-perf-users, linux-kernel
Cc: Tudor-Stefan Magirescu
Flame graphs report samples at function granularity, so a wide frame shows
which function is hot but not which part of it. Obtaining that requires
perf annotate, which reports per-instruction counts without the calling
context of a flame graph. This adds an --asm option that appends the
sampled instruction as a leaf node, so both views are available at once.
While testing locally on x86_64 I noticed that when recording using
--call-graph fp and :p event modifiers, the top stack frame of the
callchain and the sample information don't agree. In my case, it appears
that the callchain's ip is exactly one instruction after the sample's ip.
perf annotate seems to ignore the callchain information and only uses the
sample ip to record the distributions, which I replicated in the script,
so that both tools attribute a sample to the same instruction.
This approach has 2 problems:
1) (Occurs only when recording with --call-graph fp and :p) The sample and
top of callchain might not refer to the same function, which means that
some instructions might be misattributed to a wrong call stack. Is a fixup
wanted here, and if so should it live in this script or where the
callchain is built?
2) A binary object might contain 2 or more symbols with the same name but
different code (e.g., when defining 2 static functions with the same name
in different translation units). In this case, the approach cannot
disambiguate between them, so instructions might get misattributed. Would
exporting the symbol start and end for the sample, as already present for
callchain entries, be acceptable? This would also remove the objdump -t
call entirely.
Tudor-Stefan Magirescu (1):
perf scripts flamegraph: Add --asm option
tools/perf/scripts/python/flamegraph.py | 120 ++++++++++++++++++++++++
1 file changed, 120 insertions(+)
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [RFC PATCH 1/1] perf scripts flamegraph: Add --asm option
2026-08-21 8:50 [RFC PATCH 0/1] perf scripts flamegraph: Add --asm option Tudor-Stefan Magirescu
@ 2026-08-21 8:50 ` Tudor-Stefan Magirescu
2026-08-21 9:05 ` sashiko-bot
0 siblings, 1 reply; 3+ messages in thread
From: Tudor-Stefan Magirescu @ 2026-08-21 8:50 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, linux-perf-users, linux-kernel
Cc: Tudor-Stefan Magirescu
The flamegraph script reports samples at function granularity, so a wide
frame shows which function is hot but not which part of it. Obtaining
that requires perf annotate, which reports per-instruction counts without
the calling context of a flame graph.
Add an --asm option to emit more fine-grained flame graphs by including
instruction-level information. By enabling this option, instructions
appear as the leaves in the flame graph call stacks:
main
a(int)
imul $0x3b9aca07,%rax,%rdx [0x13ad]
Signed-off-by: Tudor-Stefan Magirescu <tudor.magirescu@gmail.com>
---
tools/perf/scripts/python/flamegraph.py | 120 ++++++++++++++++++++++++
1 file changed, 120 insertions(+)
diff --git a/tools/perf/scripts/python/flamegraph.py b/tools/perf/scripts/python/flamegraph.py
index ad735990c5be..424919e24aed 100755
--- a/tools/perf/scripts/python/flamegraph.py
+++ b/tools/perf/scripts/python/flamegraph.py
@@ -19,14 +19,23 @@
# pylint: disable=missing-function-docstring
import argparse
+from dataclasses import dataclass
import hashlib
import io
import json
import os
import subprocess
import sys
+import re
from typing import Dict, Optional, Union
import urllib.request
+from perf_trace_context import perf_config_get
+
+
+def default_objdump():
+ config = perf_config_get("annotate.objdump")
+ return config if config else "objdump"
+
MINIMAL_HTML = """<head>
<link rel="stylesheet" type="text/css" href="https://cdn.jsdelivr.net/npm/d3-flame-graph@4.1.3/dist/d3-flamegraph.css">
@@ -67,10 +76,101 @@ class Node:
}
+@dataclass
+class Symbol:
+ start: int
+ size: int
+
+
+class Instructions:
+ def __init__(self):
+ self.symbols: dict[str, dict[str, Symbol]] = {}
+ self.instructions: dict[tuple[str, str], dict[int, str]] = {}
+ self.objdump_line: re.Pattern = re.compile(r"^\s+([0-9a-f]+):\s+(.*)")
+ self.symbol_line: re.Pattern = re.compile(
+ r"^([0-9a-f]+) (.{7})\s+\S+\s+([0-9a-f]+)\s+(.*)$")
+ self.objdump: str = default_objdump()
+
+ def load_symbols(self, dso: str) -> dict[str, Symbol]:
+ result = subprocess.run(
+ [self.objdump, "--demangle", "-t", dso],
+ capture_output=True, text=True, check=False
+ )
+
+ if result.returncode != 0:
+ return {}
+
+ symbols: dict[str, Symbol] = {}
+
+ for line in result.stdout.splitlines():
+ match = self.symbol_line.match(line)
+ if not match:
+ continue
+
+ value, flags, size_str, name = match.groups()
+ # the 7th flag column holds the symbol type, "F" for a function
+ if flags[6] != "F":
+ continue
+
+ size = int(size_str, 16)
+ if size == 0:
+ continue
+
+ # a name may be preceded by its visibility
+ parts = name.split(maxsplit=1)
+ if len(parts) == 2 and parts[0] in (".hidden", ".protected",
+ ".internal"):
+ name = parts[1]
+
+ symbols[name.strip()] = Symbol(start=int(value, 16), size=size)
+
+ return symbols
+
+ def load_instructions(self, dso: str, sym: Symbol) -> dict[int, str]:
+ result = subprocess.run(
+ [
+ self.objdump, "-d",
+ "--no-show-raw-insn",
+ f"--start-address=0x{sym.start:x}",
+ f"--stop-address=0x{sym.start + sym.size:x}",
+ dso,
+ ],
+ capture_output=True, text=True, check=False
+ )
+
+ instructions: dict[int, str] = {}
+
+ for line in result.stdout.splitlines():
+ match = self.objdump_line.match(line)
+ if not match:
+ continue
+ addr = int(match.group(1), 16)
+ insn = match.group(2).strip()
+ instructions[addr] = insn
+
+ return instructions
+
+ def lookup_instruction(self, dso: str, func: str,
+ off: int) -> Optional[tuple[int, Optional[str]]]:
+ if dso not in self.symbols:
+ self.symbols[dso] = self.load_symbols(dso)
+
+ sym = self.symbols[dso].get(func)
+ if sym is None:
+ return None
+
+ if (dso, func) not in self.instructions:
+ self.instructions[(dso, func)] = self.load_instructions(dso, sym)
+
+ addr = sym.start + off
+ return (addr, self.instructions[(dso, func)].get(addr))
+
+
class FlameGraphCLI:
def __init__(self, args):
self.args = args
self.stack = Node("all", "root")
+ self.instructions = Instructions() if args.asm else None
@staticmethod
def get_libtype_from_dso(dso: Optional[str]) -> str:
@@ -119,6 +219,22 @@ class FlameGraphCLI:
name = event.get("symbol", "[unknown]")
libtype = self.get_libtype_from_dso(event.get("dso"))
node = self.find_or_create_node(node, name, libtype)
+
+ if self.args.asm:
+ # use the sample IP directly rather than callchain[0], since with
+ # precise event recording (i.e. :pp) the top callchain entry may
+ # point to the next instruction rather than the sampled IP
+ sym_name = event.get("symbol")
+ sym_off = event.get("symoff")
+ dso = event.get("dso")
+
+ if sym_name and sym_off is not None and dso:
+ found = self.instructions.lookup_instruction(dso, sym_name, sym_off)
+ if found is not None:
+ addr, insn = found
+ name = f"{insn} [0x{addr:x}]" if insn else f"[0x{addr:x}]"
+ libtype = self.get_libtype_from_dso(dso)
+ node = self.find_or_create_node(node, name, libtype)
node.value += 1
def get_report_header(self) -> str:
@@ -259,6 +375,10 @@ if __name__ == "__main__":
dest="event_name",
type=str,
help="specify the event to generate flamegraph for")
+ parser.add_argument("--asm",
+ default=False,
+ action="store_true",
+ help="annotate leaf frames with instruction-level nodes")
cli_args = parser.parse_args()
cli = FlameGraphCLI(cli_args)
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [RFC PATCH 1/1] perf scripts flamegraph: Add --asm option
2026-08-21 8:50 ` [RFC PATCH 1/1] " Tudor-Stefan Magirescu
@ 2026-08-21 9:05 ` sashiko-bot
0 siblings, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-08-21 9:05 UTC (permalink / raw)
To: Tudor-Stefan Magirescu; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [Medium] The `--asm` option directly opens unverified `dso` paths on the host filesystem, bypassing perf's build-id cache.
- [Medium] The script invokes a new `objdump` process via `subprocess.run` for every unique sampled function, creating a severe performance bottleneck.
- [Medium] The script crashes with a raw `FileNotFoundError` if the configured or default `objdump` executable is missing.
--
--- Patch [1]: [RFC PATCH 1/1] perf scripts flamegraph: Add --asm option ---
commit e7196726e4cd528492356aed32496e94223294a5
Author: Tudor-Stefan Magirescu <tudor.magirescu@gmail.com>
perf scripts flamegraph: Add --asm option
The flamegraph script reports samples at function granularity, so a wide
frame shows which function is hot but not which part of it. This patch
adds an --asm option to emit more fine-grained flame graphs by including
instruction-level information.
> diff --git a/tools/perf/scripts/python/flamegraph.py b/tools/perf/scripts/python/flamegraph.py
> index ad735990c5bec..424919e24aed7 100755
> --- a/tools/perf/scripts/python/flamegraph.py
> +++ b/tools/perf/scripts/python/flamegraph.py
[ ... ]
> @@ -67,10 +76,101 @@ class Node:
> }
>
>
> +@dataclass
> +class Symbol:
> + start: int
> + size: int
> +
> +
> +class Instructions:
> + def __init__(self):
> + self.symbols: dict[str, dict[str, Symbol]] = {}
> + self.instructions: dict[tuple[str, str], dict[int, str]] = {}
> + self.objdump_line: re.Pattern = re.compile(r"^\s+([0-9a-f]+):\s+(.*)")
> + self.symbol_line: re.Pattern = re.compile(
> + r"^([0-9a-f]+) (.{7})\s+\S+\s+([0-9a-f]+)\s+(.*)$")
> + self.objdump: str = default_objdump()
> +
> + def load_symbols(self, dso: str) -> dict[str, Symbol]:
> + result = subprocess.run(
> + [self.objdump, "--demangle", "-t", dso],
> + capture_output=True, text=True, check=False
> + )
[Severity: Medium]
Does this code bypass perf's build-id cache by opening the dso path directly on
the host filesystem?
The objdump command attempts to open the dso string directly on the host
filesystem.
This can fail for kernel symbols like [kernel.kallsyms] or read incorrect host
binaries when analyzing cross-platform or container perf.data files.
[Severity: Medium]
Will this crash with a raw FileNotFoundError if the objdump executable is
missing?
Without an exception handler, if the configured or default objdump executable is
missing, the script will abruptly crash with a Python traceback.
It might be better to gracefully fall back to function-level nodes or display a
helpful error.
[ ... ]
> + def load_instructions(self, dso: str, sym: Symbol) -> dict[int, str]:
> + result = subprocess.run(
> + [
> + self.objdump, "-d",
> + "--no-show-raw-insn",
> + f"--start-address=0x{sym.start:x}",
> + f"--stop-address=0x{sym.start + sym.size:x}",
> + dso,
> + ],
> + capture_output=True, text=True, check=False
> + )
[Severity: Medium]
Will this also crash with a raw FileNotFoundError if the objdump executable is
missing?
Like load_symbols(), this subprocess.run() call lacks an exception handler for
when objdump cannot be found on the system path.
[ ... ]
> + def lookup_instruction(self, dso: str, func: str,
> + off: int) -> Optional[tuple[int, Optional[str]]]:
> + if dso not in self.symbols:
> + self.symbols[dso] = self.load_symbols(dso)
> +
> + sym = self.symbols[dso].get(func)
> + if sym is None:
> + return None
> +
> + if (dso, func) not in self.instructions:
> + self.instructions[(dso, func)] = self.load_instructions(dso, sym)
[Severity: Medium]
Does this create a severe performance bottleneck by invoking a new objdump
process for every unique sampled function?
Profiling an application with many unique functions causes thousands of objdump
processes to be spawned sequentially, resulting in a severe performance
bottleneck and drastically increasing execution time.
[ ... ]
> @@ -119,6 +219,22 @@ class FlameGraphCLI:
> name = event.get("symbol", "[unknown]")
> libtype = self.get_libtype_from_dso(event.get("dso"))
> node = self.find_or_create_node(node, name, libtype)
> +
> + if self.args.asm:
> + # use the sample IP directly rather than callchain[0], since with
> + # precise event recording (i.e. :pp) the top callchain entry may
> + # point to the next instruction rather than the sampled IP
> + sym_name = event.get("symbol")
> + sym_off = event.get("symoff")
> + dso = event.get("dso")
> +
> + if sym_name and sym_off is not None and dso:
> + found = self.instructions.lookup_instruction(dso, sym_name, sym_off)
[Severity: Medium]
Does this pass the target's unresolved dso path directly to the instruction
lookup logic?
When analyzing perf.data files from other architectures, containers, or kernel
space, using the dso string directly on the host filesystem will lead to missing
or incorrect instruction mappings because it bypasses the build-id cache.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260821085044.982835-2-tudor.magirescu@gmail.com?part=1
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-21 9:05 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-21 8:50 [RFC PATCH 0/1] perf scripts flamegraph: Add --asm option Tudor-Stefan Magirescu
2026-08-21 8:50 ` [RFC PATCH 1/1] " Tudor-Stefan Magirescu
2026-08-21 9:05 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox