From: 袁晓峰 <yuanxiaofeng@eswincomputing.com>
To: "Nam Cao" <namcao@linutronix.de>
Cc: "Paul Walmsley" <pjw@kernel.org>,
"Palmer Dabbelt" <palmer@dabbelt.com>,
"Albert Ou" <aou@eecs.berkeley.edu>,
linux-riscv@lists.infradead.org
Subject: Re: Re: [PATCH v2 1/2] riscv: kprobes: simulate nop and c.nop instructions
Date: Wed, 2 Sep 2026 16:43:14 +0800 (GMT+08:00) [thread overview]
Message-ID: <3f4ecb79.9da1.1a061496120.Coremail.yuanxiaofeng@eswincomputing.com> (raw)
In-Reply-To: <87wltdt4md.fsf@yellow.woof>
From: Xiaofeng Yuan <yuanxiaofeng@eswincomputing.com>
To: Nam Cao <namcao@linutronix.de>
Cc: Paul Walmsley <pjw@kernel.org>, Palmer Dabbelt <palmer@dabbelt.com>,
Albert Ou <aou@eecs.berkeley.edu>, Nam Cao <namcaov@gmail.com>,
linux-riscv@lists.infradead.org
Subject: Re: [PATCH v2 1/2] riscv: kprobes: simulate nop and c.nop instructions
In-Reply-To: <87wltdt4md.fsf@yellow.woof>
Hi Nam,
Thanks for the detailed review. You were right on several points; let
me take them in order.
Xiaofeng Yuan <yuanxiaofeng@eswincomputing.com> writes:
> > nop and c.nop have no architectural effect, so allocating an
> [..]
> > pure overhead. [...] following the approach already
> > used on arm64.
>
> That is not the reason why arm64 simulate nop. And how is single
> stepping slower than simulating?
You are right about the first part: I should not have paraphrased the
arm64 motivation. ac4ad5c09b34 ("arm64: insn: Simulate nop
instruction for better uprobe performance") was motivated by Andrii's
uprobe benchmarks, where a probed nop was about 2x slower than an
already-emulated instruction.
On the second question, concretely what the extra cost is on riscv:
the "single-stepping" here is done by the XOL slot -- the breakpoint
trap runs the handler and redirects execution to a copy of the
probed instruction with another ebreak appended after it, which
traps back into the kernel a second time just to complete the step.
So each hit of a probed nop costs two exception round-trips (full
pt_regs + sret each) for the slot path versus one for simulation.
For uprobes it is worse still: the slot runs in user mode, so the
re-trap is a full kernel<->userspace round-trip. I measured
this on QEMU (RISC-V virt): a ftrace uprobe event on a
USDT-style nop site reduces the per-hit cost by roughly 80
percent compared to the XOL path; a kprobe on a nop by roughly
40 percent. (arm64 commit ac4ad5c09b34 measured ~2x on real
hardware for the same class of change.) v3 carries these
measurements instead of the arm64 rationale I mis-stated.
> > [...] This
> > is relevant for USDT probe sites in user-space binaries, which are
> [...]
>
> I can only find https://github.com/chrisa/libusdt, which does not have
> riscv support. Which USDT are you referring to?
Sorry for the vague shorthand -- by USDT I mean SystemTap-style SDT
probes (the .note.stapsdt format parsed by libbpf and used by tools
like bpftrace). On riscv a probe site is an SDT note whose recorded
Location is a plain nop by construction. In-tree evidence:
- tools/testing/selftests/bpf/sdt.h emits the location as
"990: _SDT_NOP"; _SDT_NOP is specialized only for ia64/s390
("nop"/"nop 0"), so riscv takes the generic "nop" path;
- libbpf documents the model: "USDT call is actually not a function
call, but is instead replaced by a single NOP instruction ... [the]
NOP instruction that kernel can replace with an interrupt
instruction" (tools/lib/bpf/usdt.c), and it already has riscv
specific note/argument parsing (the register map in
tools/lib/bpf/usdt.c);
- the bpf selftests USDT provider (urandom_read.c: STAP_PROBE1/3)
has riscv-specific build support
(tools/testing/selftests/bpf/Makefile);
- and the arm64 series ac4ad5c09b34 states it itself: "Typicall uprobe
is installed on 'nop' for USDT".
To make it concrete I cross-compiled a provider with the in-tree
header (riscv64-linux-musl-gcc -O2 -static):
$ readelf -n usdt_probe | grep -A4 stapsdt
[...]
Provider: "bench"
Name: "nop_site"
Location: 0x00000000000006ba, Base: 0x0000000000001052
$ objdump -d usdt_probe | grep -A2 '<probe_site>:'
00000000000006ba <probe_site>:
6ba: 0001 nop (rv64gc: compressed c.nop)
6bc: 8082 ret
and with -march=rv64imafd (no RVC):
6c8: 00000013 nop (4-byte nop)
so this is the ABI-defined instruction at every riscv USDT site --
exactly the two instructions this patch simulates.
To be honest about scope: I have not verified how widely USDT is
deployed on riscv workloads today. My claim is the narrower one:
libbpf parses riscv SDT notes, the bpf selftests' USDT provider has
riscv-specific build handling, and the sdt.h macro body emits the
note location as a plain nop/c.nop by construction -- so USDT is the
intended consumer of probed nops on riscv, the same model the arm64
series describes: "Typicall uprobe is installed on 'nop' for USDT".
> > In kernel text, nops are also found at ftrace
> > mcount call sites and disabled jump_label sites.
>
> Not sure about mcount, but isn't installing kprobe on jump labels
> forbidden?
You are right, and I dropped both claims in v3. register_kprobe()
rejects jump_label text sites via jump_label_text_reserved()
(kernel/kprobes.c), which covers the site regardless of whether it
currently holds a nop or a JAL (and static_call sites likewise), so
those nops are not natural probe candidates at all. The mcount claim
was misleading even where it is allowed: on riscv the function entry
instruction is the auipc of the mcount pair; the nop is only at +4.
> > Measured on QEMU [...]
> > roughly a 2x [...]
> Reading the arm64's commit, simulating nop should only improve uprobe,
> not kprobe; or am I confused somewhere?
Not confused about the motivation -- on arm64 the benchmark and the
title are uprobe-specific, and the common real-world producer of
probed nops is also uprobes (USDT sites), as above. The cost model
just doesn't stop at kprobes on our side: riscv_probe_decode_insn()
is shared by arch_prepare_kprobe() and arch_uprobe_analyze_insn(),
and the kprobe path is likewise a slot-replay-with-re-trap (second
breakpoint round-trip), so simulating the nop removes one exception
round-trip per hit of a probed nop for kprobes too. But your
framing is fair: the natural producer of this traffic is uprobes
on USDT sites, so v3 leads with the uprobe measurement (~50us ->
~9us per hit, QEMU) and keeps the kprobe reduction (~40 percent)
as the secondary number for the shared decode path.
v3 is out with these corrections; the diff itself is unchanged.
Best regards,
Xiaofeng Yuan
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
next prev parent reply other threads:[~2026-09-02 8:43 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 12:20 [PATCH v2 0/2] riscv: kprobes: simulate nop and c.nop instructions Xiaofeng Yuan
2026-08-25 12:20 ` [PATCH v2 1/2] " Xiaofeng Yuan
2026-08-26 13:04 ` Nam Cao
2026-09-02 8:43 ` 袁晓峰 [this message]
2026-08-25 12:20 ` [PATCH v2 2/2] riscv: kprobes: add nop and c.nop to the KUnit test Xiaofeng Yuan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3f4ecb79.9da1.1a061496120.Coremail.yuanxiaofeng@eswincomputing.com \
--to=yuanxiaofeng@eswincomputing.com \
--cc=aou@eecs.berkeley.edu \
--cc=linux-riscv@lists.infradead.org \
--cc=namcao@linutronix.de \
--cc=palmer@dabbelt.com \
--cc=pjw@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.