All of lore.kernel.org
 help / color / mirror / Atom feed
From: Andi Kleen <ak@kernel.org>
To: linux-kernel@vger.kernel.org
Cc: mhiramat@kernel.org, oleg@redhat.com, peterz@infradead.org,
	tglx@kernel.org, x86@kernel.org, jolsa@kernel.org,
	linux-perf-users@vger.kernel.org, adrian.hunter@intel.com,
	Andi Kleen <ak@kernel.org>
Subject: [RFC v1 15/19] ptwrite uprobes: Add a tutorial and overview documentation
Date: Mon, 31 Aug 2026 08:04:51 -0700	[thread overview]
Message-ID: <20260831150651.1134594-16-ak@kernel.org> (raw)
In-Reply-To: <20260831150651.1134594-1-ak@kernel.org>

No code changes.

Assisted-by: omp:gpt-5.6-luna
Signed-off-by: Andi Kleen <ak@kernel.org>
---
 Documentation/trace/index.rst           |   1 +
 Documentation/trace/ptwrite-uprobes.rst | 390 ++++++++++++++++++++++++
 2 files changed, 391 insertions(+)
 create mode 100644 Documentation/trace/ptwrite-uprobes.rst

diff --git a/Documentation/trace/index.rst b/Documentation/trace/index.rst
index f4058e8e92e3..4ae7b158804b 100644
--- a/Documentation/trace/index.rst
+++ b/Documentation/trace/index.rst
@@ -90,6 +90,7 @@ interactions.
 .. toctree::
    :maxdepth: 1
 
+   ptwrite-uprobes
    user_events
    uprobetracer
 
diff --git a/Documentation/trace/ptwrite-uprobes.rst b/Documentation/trace/ptwrite-uprobes.rst
new file mode 100644
index 000000000000..79ab35824b93
--- /dev/null
+++ b/Documentation/trace/ptwrite-uprobes.rst
@@ -0,0 +1,390 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+===============
+ptwrite uprobes
+===============
+
+.. contents:: :local:
+
+Introduction
+============
+
+A classic uprobe enters the kernel for each probe. That causes overhead
+when the probe is executed frequently.
+
+ptwrite uprobes instead rely on hardware tracing that doesn't enter
+the kernel. It uses the ``PTWRITE`` instruction available on modern
+Intel CPUs to log data to the Processor Trace buffer. Processor
+Trace is configured and recorded by Linux perf.
+
+There are some limitations of the scheme (see below)
+but it is a lot faster than classic uprobes.
+
+Performance
+===========
+
+Measured on the development kernel in a KVM guest (2 vCPUs) running
+on a AlderLake laptop with a probe on a hot function called in a
+tight loop.
+
+    +---------------------------------------+----------------------+--------------------+------------------------+
+    | mode (2-arg probe)                    | PT off (% of classic)| full (% of classic)| snapshot (% of classic)|
+    +=======================================+======================+====================+========================+
+    | classic uprobe (tracefs)              | 100%                 | 100%               | 100%                   |
+    | classic uprobe (perf probe, trace r.) | 104%                 |                    |                        |
+    | classic uprobe (perf probe, perf ring)|                      | 156%               |                        |
+    | ptwrite %nopace                       | 2%                   | 2%                 | 8%                     |
+    | ptwrite default                       | 7%                   | 7%                 | 11%                    |
+    | perf probe ``--ptwrite``              | 7%                   | 7%                 | 12%                    |
+    +---------------------------------------+----------------------+--------------------+------------------------+
+
+The percentages are normalized to the classic tracefs uprobe in each
+recording mode; lower values represent lower cost per hit.
+
+A ``%nopace`` probe costs roughly **2%** as much as a classic uprobe
+with PT off (about 98% less). The default pacing (on unless ``%nopace``
+is given) costs roughly 7% as much (about 93% less) in the same mode.
+The default pacing slows down the probes to avoid data loss when they are
+too tightly spaced.
+
+``snapshot`` refers to ``perf record`` snapshot mode (``-S``) which doesn't
+save the PT ring buffer constantly.
+
+
+Requirements
+============
+
+- An Intel CPU with Intel PT and PTWRITE. When running as a guest Intel PT
+  needs to be exposed to the guest.
+  PT/PTWRITE are available when ``/sys/devices/intel_pt/format/ptw`` exists.
+- A kernel with ``CONFIG_UPROBE_EVENTS`` enabled.
+
+Quick start (tracefs)
+=====================
+
+Pick a probe site, register a probe at its file offset, enable it, run the
+program under PT, decode.
+
+Example 1: probe an existing instruction (punning)
+----------------------------------------------------
+
+Build a small program and probe the entry of ``main``::
+
+    $ cat > t.c <<'EOF'
+    #include <stdio.h>
+
+    __attribute__((noinline, noipa)) static unsigned long
+    target(unsigned long a, unsigned long b)
+    {
+        return a * 31 + b;
+    }
+
+    int main(void)
+    {
+        unsigned long i, acc = 0;
+        for (i = 0; i < 100; i++)
+            acc += target(i, i + 1);
+        printf("acc=%lu\n", acc);
+        return 0;
+    }
+    EOF
+    $ gcc -O2 -no-pie -fno-inline -o t t.c
+
+``objdump -F`` prints the file offset of every instruction::
+
+    $ objdump -d -F t | sed -n "/<main> (File Offset/,+1p"
+    0000000000401040 <main> (File Offset: 0x1040):
+      401040:	55			push   %rbp
+
+``1040`` is the file offset of ``main``'s first instruction, exactly
+what the probe line needs. Register the probe there, enable it, run
+the program under PT and decode::
+
+    # echo "ptw:e t:0x1040 %di %si" > /sys/kernel/tracing/uprobe_events
+    # echo 1 > /sys/kernel/tracing/events/uprobes/e/enable
+    # perf record -e intel_pt/ptw=1,fup_on_ptw=1/u -o perf.data ./t
+    # perf script --itrace=qwe -s uprobe-ptwrite-decode.py -i perf.data
+    record 1: event=uprobes/e id=0x6c2 args=[1, 140728356930600]
+    summary: records=1 dropped=0 stray=0 unknown=0 errors=0
+
+Since it isn't a nop the instruction is "punned": byte 0 is changed
+to a jump to a trampoline that logs the data and returns to the
+previous execution.
+
+Punning is a probabilistic method that depends on the existing
+instruction bytes and the placement of the executable in memory.
+It has a high chance of success on PIE/PIC binaries, but tends
+to work poorly on non PIE main executables.
+
+When punning is not possible the probe is rejected at install
+time. Options in this case:
+- Move the probe site to a different instruction which may work.
+- Rebuild with -fPIE if it's a main problem not using PIE.
+- Enable or disable /proc/sys/kernel/randomize_va_space. If the
+  randomization is enabled it may also just work on a rerun of
+  the program.
+- Fall back to a classic uprobes
+- Insert a 5 byte nop which is always supported (see below)
+
+Example 2: an explicit 5-byte NOP (inline assembly)
+---------------------------------------------------
+
+Add a 5-byte NOP at the probe point::
+
+    $ cat > t.c <<'EOF'
+    #include <stdio.h>
+
+    __attribute__((noinline)) static unsigned long
+    target(unsigned long a, unsigned long b)
+    {
+        asm volatile(".byte 0x0f, 0x1f, 0x44, 0x00, 0x00"); /* nopl */
+        return a * 31 + b;
+    }
+
+    int main(void)
+    {
+        unsigned long i, acc = 0;
+        for (i = 0; i < 100; i++)
+            acc += target(i, i + 1);
+        printf("acc=%lu\n", acc);
+        return 0;
+    }
+    EOF
+    $ gcc -O2 -no-pie -o t t.c
+
+    $ objdump -d -F t | sed -n "/<target> (File Offset/,+1p"
+    0000000000401170 <target> (File Offset: 0x1170):
+      401170:	0f 1f 44 00 00		nopl   0x0(%rax,%rax,1)
+
+Probe it exactly like example 1::
+
+    # echo "ptw:e t:0x1170 %di %si" > /sys/kernel/tracing/uprobe_events
+    # echo 1 > /sys/kernel/tracing/events/uprobes/e/enable
+    # perf record -e intel_pt/ptw=1,fup_on_ptw=1/u -o perf.data ./t
+    # perf script --itrace=qwe -s uprobe-ptwrite-decode.py -i perf.data
+    record 99: event=uprobes/e id=0x6c2 args=[98, 99]
+    record 100: event=uprobes/e id=0x6c2 args=[99, 100]
+    summary: records=100 dropped=0 stray=0 unknown=0 errors=0
+
+Configuring ptwrite uprobes
+===========================
+
+ptwrite uprobes is configured like normal uprobes by writing
+commands to ``/sys/kernel/tracing/uprobe_events``.
+
+  ptw[:[GRP/][EVENT]] PATH:OFFSET [FETCHARGS] : set a ptwrite probe
+  -:[GRP/][EVENT]                             : clear a probe
+
+  GRP      : group name. If omitted, "uprobes" is the default (the
+             event appears under events/uprobes/).
+  EVENT    : event name. If omitted, one is generated from PATH+OFFSET.
+  PATH     : path to an executable or a library.
+  OFFSET   : file offset of the probe site (0x-prefixed hex, see above).
+  FETCHARGS: probe arguments, up to 8 (see "Argument syntax" below).
+
+After creating the ptwrite uprobe it becomes available with its name
+in ``/sys/kernel/tracing/uprobe_events``. There it can be enabled
+by writing 1 to its enable field. However it only logs data
+when a Linux perf PT recording session with ptw=1 is active.
+
+perf probe
+----------
+
+``perf probe --ptwrite -x <file>`` creates ptwrite uprobes instead of
+the classic trap-based ones. The example below uses SDT probes.
+
+(this requires installing systemtap-devel or an equivalent package)
+
+    $ cat > t.c <<'EOF'
+    #include <stdio.h>
+    #include <sys/sdt.h>
+    __attribute__((noinline, noclone)) static unsigned long
+    target(unsigned long a, unsigned long b)
+    {
+        unsigned long local = a * 2;
+        STAP_PROBE1(test, rarg, a);
+        STAP_PROBE1(test, carg, 42);
+        STAP_PROBE2(test, marg, &local, b);
+        return a * 31 + b;
+    }
+    int main(void)
+    {
+        unsigned long i, acc = 0;
+        for (i = 0; i < 20; i++) {
+            acc += target(i, i + 1);
+            asm volatile("pause");
+        }
+        printf("acc=%lu\n", acc);
+        return 0;
+    }
+    EOF
+    $ gcc -O2 -no-pie -o t t.c
+
+The first probe point (``rarg``) is a nop 9 bytes into ``target``::
+
+    $ objdump -d t | sed -n "/<target>:/,+3p"
+    0000000000401180 <target>:
+      401180:	48 8d 04 3f		lea    (%rdi,%rdi,1),%rax
+      401184:	48 89 44 24 f8		mov    %rax,-0x8(%rsp)
+      401189:	90			nop
+
+Probe it with ``perf probe --ptwrite`` using the function+offset
+form, then enable, capture and delete it like any ptwrite probe::
+
+    $ perf probe --ptwrite -x ./t --add "target+9 %di %si"
+    Added new event:
+      probe_t:target      (on target+9 in ./t with %di %si)
+
+    # the tracefs line it wrote:
+    # ptw:probe_t/target ./t:0x1189 arg1=%di arg2=%si
+
+    # echo 1 > /sys/kernel/tracing/events/probe_t/target/enable
+    # perf record -e intel_pt/ptw=1,fup_on_ptw=1/u -o perf.data ./t
+    # perf script --itrace=qwe -s uprobe-ptwrite-decode.py -i perf.data
+    record 1: event=probe_t/target id=0x6a9 args=[0, 1]
+    record 20: event=probe_t/target id=0x6a9 args=[19, 20]
+    summary: records=20 dropped=0 stray=0 unknown=0 errors=0
+    # perf probe -d probe_t:target
+
+A ``nop`` instruction, as used by SDT probes, is not guaranteed to
+be ptwrite patchable. It needs a 5-byte NOP, but it can
+often be punned. If punning fails, the kernel reports
+``failed to install`` and the probe has to be moved to another site.
+
+GCC's ``-fpatchable-function-entry=5`` may emit five one-byte NOPs.
+To use that site, add ``%multinop`` to the tracefs probe offset::
+
+    # echo "ptw:e t:0x1170%multinop %di %si" > /sys/kernel/tracing/uprobe_events
+
+The five-byte run must start at an 8-byte-aligned address because it is
+patched with an atomic eight-byte store. An unaligned ``%multinop`` site is
+rejected; without ``%multinop``, the run is treated as a pun and may not
+always succeed.
+
+perf probe uses the standard argument syntax for the ptwrite subset
+(registers, ``$stack``/``$stackN``, ``+disp(%reg)`` memory reads, and
+``\0x2a``-style constants). Strings, arrays and typed suffixes are not
+supported by the ptwrite stub and are rejected by the kernel.
+``%return`` is refused (ptwrite probes are entry-only), and the mode
+requires ``-x``. The probes are enabled, captured and deleted like
+classic probes (``perf probe -l``, ``perf probe -d``).
+They carry the default (LFENCE) pacing. ``%nopace`` cannot be selected
+through perf probe. Write the tracefs line by hand for that.
+
+Argument syntax
+---------------
+
+ptwrite uprobes only support a limited number of argument types
+compared to classic uprobes.
+
+``ptw:<name> <path>:<offset> <arg> ... [options]`` where each ``<arg>`` is one
+of:
+
+- ``%di``, ``%si``, ``%ax`` ...: a live register.
+- ``\IMM``: a fixed constant (stored in the stub), e.g. ``\0x42``.
+- ``$stack``: the stack pointer value (never faults).
+- ``$stackN``: the Nth stack slot (``[%rsp + 8N]``). ``u64`` uses an
+  8-byte load on the fault-fixup path; ``u32``/``s32``/``x32`` use a 4-byte
+  load.
+- ``+<disp>(%reg)``: read memory at ``[reg + disp]``. ``u64`` uses an
+  8-byte load; ``u32``/``s32``/``x32`` use a 4-byte load (``ptwritel``).
+  A bad address writes ``0``.
+
+Options
+-------
+
+``%nopace`` disables artificial slowdown of the probes. This can cause
+data loss when they are tightly spaced or have many arguments, but
+speeds up the probes (see the benchmark section above)
+
+``%multinop`` lets users probe a 5-byte nop sequence that is not one
+instruction. A program could jump to a later nop, which would break when
+the probe rewrites the site.
+
+However there is a common case where gcc's -fpatchable-function-entry=5
+generates 5 nops for each function that are convenient points
+for patching, and nobody jumps into the middle of them.
+
+The 5-byte single nop sequence must be aligned to 8 bytes.
+
+
+The encoding format
+===================
+
+Each probe writes a header and the arguments to the PT stream.
+
+The header is a 64-bit word. Each argument is one PTWRITE payload exposed by
+perf as a ``u64`` value.
+
+    header word:      bits 63..48  event id (matches the tracefs id in sysfs)
+                      bits 47..40  number of argument words
+                      bits 39..0   fixed magic 0x5054525731 ("PTRW1")
+    arguments:        one PTWRITE payload per FETCHARG
+
+If the program itself also executes own ``PTWRITE``, those values mix with the
+uprobe output in the stream. The decoder uses the header magic to identify
+uprobe records. Other values are printed as ``manual ptwrite:`` lines (with
+their IP when ``fup_on_ptw`` is set) and counted in the summary's ``stray``
+field.
+
+To also print the decoded branch stream alongside the records, add
+``b`` to the itrace options and drop the ``q``
+
+    # ``perf script --itrace=web -s uprobe-ptwrite-decode.py -i perf.data``
+
+Each decoded branch prints as a ``branch:`` line (from => to, with
+symbols where resolvable), interleaved with the probe records and any
+manual ptwrites in delivery order.
+
+Other events in the recording, including classic uprobes, tracepoints,
+and sample events, are printed as ``event:`` lines unless disabled
+by the decoder.
+
+Unsupported instructions for probes
+===================================
+
+The following instructions are always refused for instrumentation::
+
+- Traps: ``int3``, ``int1``, ``int imm8``, ``into``, ``iret``
+  because they save the IP.
+- System instructions: ``syscall``, ``sysenter``, ``sysexit``
+  for similar reasons.
+- Far control flow: ``jmp far``, ``call far``, and the indirect far
+  forms (call-far, jmp-far).
+- Relative branches: ``jmp rel8/rel32``, ``jcc rel8/rel32``,
+  ``loop*``, ``jecxz/jrcxz`` (target-inside-window, see below), and
+  ``call rel32`` (its return address would point into the stub).
+- Indirect ``call``: the return-address problem applies
+  to the register/memory forms too.
+- Relative branches (``jmp``/``jcc``/``loop`` with a rel8/rel32
+  displacement, and ``call rel32``) are refused because the trampoline
+  may not be able to reach the target.
+- If the 5 byte area of the instruction crosses a page boundary it
+  currently cannot be probed (this applies to nop probes too).
+
+Other restrictions
+==================
+
+- Each probe needs 4K of process memory and roughly 1K extra in the kernel.
+- The probe pages are currently only freed on process exit.
+- Return probes (``%return``/``r:``): ptwrite probes are entry-only
+  and ``%return`` is refused.
+- The SDT reference counter (``(REF)``).
+- EBPF, perf actions, filters, event predicates, histograms, triggers,
+  profiling and similar advanced trace features are all not supported
+  since they would require a kernel entry. However some basic filtering
+  is possible at the perf recording level, for example limit the scope
+  to a CPU or to a process. PT also supports address filter ranges
+  that allow filtering by IP.
+- More than one probe at the same site
+- Only 4 and 8 byte memory references are supported.
+- Fetch argument variety: classic probes fetch strings
+  (``:string``/``:ustring``), arrays, bitfields, nested derefs,
+  ``$retval``, ``$comm`` and ``$argN``. Ptwrite probes only take live
+  registers, ``\IMM`` constants, ``$stack``/``$stackN`` and
+  ``+disp(%reg)`` memory reads. Memory reads are 8-byte words for ``u64``
+  and 4-byte words for ``u32``/``s32``/``x32``.
+  (some of this could be relaxed, but it would require a writable stack)
+- Like normal uprobes one byte of the instruction stream is overwritten
+  (or 5 bytes for the nop case). If the program reads its own code
+  it might see different values.
-- 
2.54.0


  parent reply	other threads:[~2026-08-31 15:07 UTC|newest]

Thread overview: 41+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31 15:04 [RFC] ptwrite uprobes Andi Kleen
2026-08-31 15:04 ` [RFC v1 01/19] uprobes: guard trace cleanup against error pointers Andi Kleen
2026-08-31 18:15   ` sashiko-bot
2026-09-01  0:49   ` Masami Hiramatsu
2026-08-31 15:04 ` [RFC v1 02/19] uprobes: Correctly reject anonymous VMAs for breakpoint installation Andi Kleen
2026-08-31 18:29   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 03/19] uprobes: Print warning for missing breakpoint install Andi Kleen
2026-08-31 18:42   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 04/19] ptwrite uprobes: Add infrastructure for ptwrite uprobes Andi Kleen
2026-08-31 18:55   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 05/19] ptwrite uprobes: Add minimal low level support for x86 Andi Kleen
2026-08-31 19:11   ` sashiko-bot
2026-09-02 16:35   ` Lorenzo Stoakes (ARM)
2026-08-31 15:04 ` [RFC v1 06/19] ptwrite uprobes: Add a sample module to exercise interface Andi Kleen
2026-08-31 19:19   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 07/19] ptwrite uprobes: Add support to tracing infrastructure Andi Kleen
2026-08-31 19:31   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 08/19] ptwrite uprobes / x86: Add a user fault notifier chain Andi Kleen
2026-08-31 19:38   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 09/19] ptwrite uprobes: Factor file-backed instruction reads Andi Kleen
2026-08-31 19:45   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 10/19] ptwrite uprobes: Minimal memory references and fault handling Andi Kleen
2026-08-31 19:59   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 11/19] ptwrite uprobes: Add multinop support Andi Kleen
2026-08-31 20:09   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 12/19] ptwrite uprobes: Add pacing to the probes Andi Kleen
2026-08-31 20:19   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 13/19] ptwrite uprobes: Support instruction puning Andi Kleen
2026-08-31 20:39   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 14/19] ptwrite uprobes: Use atomic patching for multinop sites Andi Kleen
2026-08-31 21:08   ` sashiko-bot
2026-08-31 15:04 ` Andi Kleen [this message]
2026-08-31 21:10   ` [RFC v1 15/19] ptwrite uprobes: Add a tutorial and overview documentation sashiko-bot
2026-08-31 15:04 ` [RFC v1 16/19] ptwrite uprobes / perf tools pt: Improve FUP error handling for ptwrite Andi Kleen
2026-08-31 21:19   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 17/19] ptwrite uprobes / perf tools probe: Add support of ptwrite probes Andi Kleen
2026-08-31 21:32   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 18/19] ptwrite uprobes / perf tools script: Add ptwrite uprobes decoder Andi Kleen
2026-08-31 21:39   ` sashiko-bot
2026-08-31 15:04 ` [RFC v1 19/19] ptwrite uprobes: Add self tests Andi Kleen
2026-08-31 21:47   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260831150651.1134594-16-ak@kernel.org \
    --to=ak@kernel.org \
    --cc=adrian.hunter@intel.com \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=mhiramat@kernel.org \
    --cc=oleg@redhat.com \
    --cc=peterz@infradead.org \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.