From: Kunwu Chan <kunwu.chan@gmail.com>
To: sj@kernel.org, akpm@linux-foundation.org
Cc: damon@lists.linux.dev, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org,
linux-kselftest@vger.kernel.org, rostedt@goodmis.org,
mhiramat@kernel.org, mathieu.desnoyers@efficios.com,
shuah@kernel.org, lianux.mm@gmail.com,
Kunwu Chan <kunwu.chan@gmail.com>
Subject: [RFC PATCH 6/7] Docs/mm/damon: document the perf observability framework
Date: Tue, 18 Aug 2026 14:10:30 +0800 [thread overview]
Message-ID: <20260818061031.827057-7-kunwu.chan@linux.dev> (raw)
In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev>
From: Lian Wang <lianux.mm@gmail.com>
Add documentation for the DAMON perf observability framework,
covering the CONFIG_DAMON_PERF_OBSERVE Kconfig option, the
debugfs perf_stats interface, the tracepoints, and the per-CPU
pipeline counter model. The debugfs format is explicitly marked
as unstable and must not be used by scripts.
Co-developed-by: Kunwu Chan <kunwu.chan@gmail.com>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
Signed-off-by: Lian Wang <lianux.mm@gmail.com>
---
Documentation/admin-guide/mm/damon/index.rst | 1 +
.../mm/damon/perf-observability.rst | 210 ++++++++++++++++++
2 files changed, 211 insertions(+)
create mode 100644 Documentation/admin-guide/mm/damon/perf-observability.rst
diff --git a/Documentation/admin-guide/mm/damon/index.rst b/Documentation/admin-guide/mm/damon/index.rst
index 3ce3164480c7..623a5c312b69 100644
--- a/Documentation/admin-guide/mm/damon/index.rst
+++ b/Documentation/admin-guide/mm/damon/index.rst
@@ -15,3 +15,4 @@ access monitoring and access-aware system operations.
reclaim
lru_sort
stat
+ perf-observability
diff --git a/Documentation/admin-guide/mm/damon/perf-observability.rst b/Documentation/admin-guide/mm/damon/perf-observability.rst
new file mode 100644
index 000000000000..3aa8185de314
--- /dev/null
+++ b/Documentation/admin-guide/mm/damon/perf-observability.rst
@@ -0,0 +1,210 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+=============================================
+DAMON Perf Event Observability Framework
+=============================================
+
+The DAMON perf event observability framework provides per-CPU counters and
+tracepoints for hardware-sampled access reports. When DAMON is configured to
+use a hardware PMU (e.g. AMD IBS, Intel PEBS, or ARM SPE) instead of page-table
+walks, this framework exposes raw pipeline diagnostics so that every stage of
+the PMU-to-DAMON pipeline can be inspected.
+
+Counters are best-effort: individual ``u64`` writes are atomic on 64-bit
+platforms, but no cross-field consistency is guaranteed. Do not build
+policy on snapshot reads. For stable, structured diagnostics, use the
+tracepoints under ``/sys/kernel/debug/tracing/events/damon/``.
+
+Pipeline Stages
+===============
+
+::
+
+ PMU hardware → overflow_handler / AUX drain
+ → damon_report_access() → per-CPU SPSC ring
+ → kdamond drain → target match → region update
+
+ Layer 1: Event Create perf_event_create_kernel_counter()
+ Layer 2: Event Bind per-CPU PMU attachment
+ Layer 3: Event Enable perf_event_enable()
+ Layer 4: Sampling callback / AUX record received
+ Layer 5: Ring SPSC enqueue / dequeue / overflow
+ Layer 6: Drain kdamond consumes entries from ring
+ Layer 7: Match & Update region access-rate update
+
+Each layer has a dedicated counter, and most layers have corresponding
+tracepoints. Per-CPU event state (UNINIT → CREATED → BOUND → ENABLED →
+RUNNING) is recorded unconditionally and exposed via the debugfs
+perf_stats file.
+
+Overhead Control
+================
+
+Two levels of overhead control are provided:
+
+1. **Compile-time** — ``CONFIG_DAMON_PERF_OBSERVE``
+ When set to ``n``, all observe functions are compiled to static-inline
+ no-ops. No code is generated and no runtime overhead exists.
+
+2. **Per-tracepoint on/off** — standard ftrace ``enable`` files
+ Individual tracepoints (``damon_perf_sample``, etc.) can be
+ enabled or disabled independently via
+ ``/sys/kernel/debug/tracing/events/damon/``. Counter increments are
+ unconditional (cheap per-CPU ``inc``); tracepoint decisions are
+ guarded by the ftrace static key and are zero-overhead when disabled.
+
+When ``CONFIG_DAMON_PERF_OBSERVE=y``, per-CPU counters always increment.
+There is no runtime toggle for counters; compile-time is the sole gate.
+
+Debugfs Interface
+=================
+
+Mount debugfs::
+
+ # mount -t debugfs none /sys/kernel/debug
+
+One file is created under ``/sys/kernel/debug/damon/``:
+
+perf_stats
+----------
+
+**DEBUG ONLY — format may change without notice.** Do not parse in
+scripts or tools. For stable diagnostics, use the tracepoints.
+
+Read-only. Aggregated counter table with all pipeline counters plus
+per-CPU breakdown::
+
+ # cat /sys/kernel/debug/damon/perf_stats
+ ┌──────────────┬──────────┐
+ │ Counter │ Value │
+ ├──────────────┼──────────┤
+ │ callback │ 233 │
+ │ valid │ 0 │
+ │ null │ 0 │
+ │ addr_zero │ 233 │
+ │ kernel │ 0 │
+ │ inv_phys │ 0 │
+ │ enqueue │ 0 │
+ │ dequeue │ 0 │
+ │ overflow │ 0 │
+ │ ring_peak │ 0 │
+ │ match │ 0 │
+ │ miss_tgid │ 0 │
+ │ miss_region │ 0 │
+ │ miss_bound │ 0 │
+ │ update │ 0 │
+ └──────────────┴──────────┘
+
+ Per-CPU (non-zero / non-UNINIT):
+ CPU00: st=BOUND cb=44 enq=0 deq=0 ovf=0 match=0 miss=0 upd=0
+ ...
+
+The ``st=<state>`` column shows the per-CPU event state machine position
+(UNINIT, CREATED, BOUND, ENABLED, RUNNING, ERROR), derived from the
+lifecycle observe calls. This allows verifying lifecycle progression
+without parsing dmesg.
+
+All counters are monotonic (cumulative since boot); userspace computes
+deltas between snapshots.
+
+Tracepoints
+===========
+
+Four tracepoints are defined::
+
+ damon_perf_sample
+ damon_perf_ring_overflow
+ damon_perf_report_missed
+ damon_perf_drain
+
+Enable via ftrace::
+
+ # echo 1 > /sys/kernel/debug/tracing/events/damon/damon_perf_sample/enable
+ # cat /sys/kernel/debug/tracing/trace_pipe
+
+Each ``damon_perf_sample`` record includes:
+
+ - ``addr``: the accessed virtual address (0 if the PMU did not populate)
+ - ``data_src``: PERF_MEM_* encoding (PMU-dependent)
+ - ``period``: sample period or frequency count
+ - ``cpu``: CPU that generated the sample
+ - ``reason``: 0=valid, 1=null-data, 2=addr-zero, 3=kernel-addr, 4=invalid-phys
+ - ``sample_flags``: what the PMU actually populated
+ - ``sample_type``: what DAMON requested
+ - ``context``: 0=process, 1=softirq, 2=hardirq, 3=NMI
+
+The ``context`` field is particularly useful for cross-PMU validation.
+For example, AMD IBS samples arrive in NMI context (context=3), while
+ARM SPE data from an AUX backend would arrive in process context (context=0).
+A mismatch between the expected and actual context is immediately visible.
+
+Selftest
+========
+
+A comprehensive automated test script is provided::
+
+ # cd tools/testing/selftests/damon
+ # sudo ./damon_perf_obs_test.sh --pmu arm_spe_0 --freq 0 --period 256
+
+The script performs a layer-by-layer validation:
+
+1. Checks kernel configuration (CONFIG_DAMON, CONFIG_DAMON_PERF_OBSERVE, etc.)
+2. Verifies PMU availability, tracepoints, and the debugfs perf_stats file
+3. Refuses to run if existing kdamonds are present (side-effect guard)
+4. Configures DAMON with the specified PMU via sysfs
+5. Runs a memory workload (stress-ng, stress, or dd fallback)
+6. Collects dmesg delta, trace output, and perf_stats
+7. Verifies per-CPU state progression and counter values
+
+Example output::
+
+ --- Layer 0: Environment ---
+ [PASS] CONFIG_DAMON_PERF_OBSERVE=y
+ [PASS] debugfs perf_stats exists
+
+ --- Layer 2-3: Enable & Run (via per-CPU state) ---
+ [PASS] Event Created (max per-CPU state >= CREATED)
+ [PASS] Event Bound (max per-CPU state >= BOUND)
+ [PASS] Event Enabled (max per-CPU state >= ENABLED)
+
+ --- Layer 4: Sampling (Callback) ---
+ [PASS] Sampling: 84532 callbacks received
+ Callback breakdown: valid=82103 addr_zero=0 kernel=2429
+
+ --- Layer 5: Ring ---
+ [PASS] Ring: enqueue > 0
+ [PASS] Ring: dequeue > 0
+ Ring: enqueue=82100 dequeue=81987 overflow=0
+
+ --- Layer 6: Drain & Match ---
+ [PASS] Drain & Match: 81987 matched
+ [PASS] Update: 81987 region updates
+
+Additional PMU examples::
+
+ # Software page-fault event (positive control):
+ sudo ./damon_perf_obs_test.sh --pmu software --freq 1 --sample-freq 100
+
+ # Any PMU by type number:
+ sudo ./damon_perf_obs_test.sh --pmu-type 38 --freq 0 --period 256
+
+Kernel Configuration
+====================
+
+Required for observability::
+
+ CONFIG_DAMON=y
+ CONFIG_DAMON_SYSFS=y
+ CONFIG_DAMON_VADDR=y
+ CONFIG_PERF_EVENTS=y
+ CONFIG_DEBUG_FS=y
+ CONFIG_TRACING=y
+ CONFIG_TRACEPOINTS=y
+
+Optional (enables observability framework)::
+
+ CONFIG_DAMON_PERF_OBSERVE=y
+
+When ``CONFIG_DAMON_PERF_OBSERVE=n``, ``/sys/kernel/debug/damon/perf_stats``
+is not created, tracepoints are not registered, and all observe functions
+are compiled to empty static inlines with zero overhead.
--
2.43.0
next prev parent reply other threads:[~2026-08-18 6:11 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 6:10 [RFC PATCH 0/7] mm/damon/perf: observability framework for hardware-sampled access reports Kunwu Chan
2026-08-18 6:10 ` [RFC PATCH 1/7] mm/damon/perf: add observability framework with tracepoints and CONFIG switch Kunwu Chan
2026-08-18 6:10 ` [RFC PATCH 2/7] mm/damon/perf: implement observe API and per-CPU statistics engine Kunwu Chan
2026-08-18 6:10 ` [RFC PATCH 3/7] mm/damon/perf: add debugfs statistics interface Kunwu Chan
2026-08-18 6:10 ` [RFC PATCH 4/7] mm/damon: integrate observe API into vaddr overflow handlers and core drain Kunwu Chan
2026-08-18 6:10 ` [RFC PATCH 5/7] selftests/damon: add automated layer-by-layer observability test Kunwu Chan
2026-08-18 6:10 ` Kunwu Chan [this message]
2026-08-18 6:10 ` [RFC PATCH 7/7] mm/damon/perf: add CONFIG_DAMON_PERF_DEBUG and pipeline health check Kunwu Chan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260818061031.827057-7-kunwu.chan@linux.dev \
--to=kunwu.chan@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=damon@lists.linux.dev \
--cc=lianux.mm@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=rostedt@goodmis.org \
--cc=shuah@kernel.org \
--cc=sj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.