* [PATCH v6 0/9] rv: Add task latency over budget RV monitor
@ 2026-08-20 16:45 wen.yang
2026-08-20 16:45 ` [PATCH v6 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
` (8 more replies)
0 siblings, 9 replies; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
This series introduces tlob (task latency over budget), a per-task
hybrid automaton RV monitor that measures elapsed wall-clock time across
a user-delimited code section and emits an error when the elapsed time
exceeds a configurable budget.
The monitor tracks three states (running, waiting, sleeping) driven by
sched_switch and sched_wakeup tracepoints. A single clock invariant,
clk_elapsed < BUDGET_NS(), is enforced by a per-task hrtimer
(HRTIMER_MODE_REL_HARD). On expiry, error_env_tlob is emitted, followed
by detail_env_tlob carrying a per-state time breakdown (running_ns,
waiting_ns, sleeping_ns).
Tasks are registered for monitoring by writing to the tracefs monitor
file:
echo "p /path/to/binary:OFFSET_START OFFSET_STOP threshold=NS" \
> /sys/kernel/tracing/rv/monitors/tlob/monitor
Two uprobes delimit the measured section without modifying the target
binary. Multiple uprobe pairs can be active simultaneously. A task may
repeat the START/STOP sequence any number of times: the entry parks in
the automaton's "stopped" state and later restarts in place, reusing the
same pool slot (see Patch 6).
Series structure
----------------
Patch 1: rv: Introduce DA_MON_ALLOCATION_STRATEGY
Compile-time allocation strategy selector for per-object DA storage:
DA_ALLOC_AUTO (default kmalloc), DA_ALLOC_POOL (pre-allocated mempool),
DA_ALLOC_MANUAL (caller-managed). DA_MON_POOL_SIZE alone selects pool
mode.
Patch 2: rv: Add generic uprobe infrastructure for RV monitors
Thin wrapper over uprobe_consumer for path resolution, registration,
and safe synchronous teardown. struct rv_uprobe embeds
uprobe_consumer directly and holds the probed binary's path for its
lifetime, avoiding a separate heap allocation per probe.
Patch 3: rv: Add tlob model DOT file
Graphviz DOT specification of the tlob hybrid automaton.
Patch 4: rv: Fix ha_invariant_passed_ns silent bypass of invariant check
Without this fix, ha_invariant_passed_ns() returns 0 on first call and
leaves env_store at U64_MAX. Every subsequent state transition resets
the hrtimer to the full budget instead of the remaining time, so the
budget never expires for tasks that transition between states. This
causes tlob's detail_sleeping and detail_waiting selftests to hang.
This patch is a minimal fix in the current dual-representation
framework. Nam's series [1] refactors ha_monitor.h to a
single-representation model; once it lands, tlob will need to be
updated to the new API and this patch will be superseded by Gabriele's
planned framework-level initialisation improvement. We carry it here
to keep the series self-contained and testable.
[1] https://lore.kernel.org/lkml/08188c28f274da63a3f8549add3086a92aef45e5.1780908661.git.namcao@linutronix.de
Patch 5: rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
#ifndef guards for da_monitor_reset_hook and EVENT_NONE_LBL.
Patch 6: rv: Add tlob hybrid automaton monitor
Main tlob implementation, including the parked-window redesign
described below.
Patches 7-9: Tests
KUnit tests for the uprobe-line parser; eight ftrace-style selftests
for tlob; the ftracetest walk-up change needed to run them from the
verification subdirectory.
Changes since v5
-----------------
Most of the fixes below address issues Sashiko (sashiko-bot@kernel.org)
flagged against this series; each is described on its own merits
below.
Patch 1: Reworded the DA_ALLOC_POOL comment; it could be misread as
allowing the allocating path to run from scheduling context.
Patch 3: Fixed a duplicate "running" node in the DOT file that made
rvgen mark it, instead of "stopped", as the model's final state.
Patch 6: Fixed automaton_tlob.final_states, generated before the DOT
fix above and inheriting the same wrong value (trace-only, no
functional impact).
Patch 7: Reworked per Gabriele's v4 feedback. Dropped
CONFIG_TLOB_KUNIT_TEST; tlob_parse_uprobe_line()/
tlob_parse_remove_line() stay static and are reached through a new
rv_tlob_kunit_ops struct, exported the same way nomiss.c/sco.c expose
rv_nomiss_ops/rv_sco_ops under the shared CONFIG_RV_MONITORS_KUNIT_TEST.
tlob_kunit.c now follows the monitors/*/*_kunit.c convention and is
included into rv_monitors_test.c instead of building standalone.
Patch 8: Fixed tracefs paths under ftracetest --rv mode; added a
VERIFICATIONTEST_BINDIR fallback for installed kselftest trees; made
uprobe_detail_waiting.tc skip on single-CPU machines instead of
pinning a SCHED_FIFO-99 busy loop; guarded tlob_target.c's empty
stop-probe bodies against being optimised away; added ELF
program-header bounds checking to tlob_sym.c.
Patch 9: Fixed an infinite loop and a local root privilege-escalation
window (sourcing test.d/functions from a world-writable directory)
in the walk-up loop.
Testing
-------
Environment: x86_64, CONFIG_PREEMPT_RT=y, kernel built with virtme-ng
(vng --build), selftests built on the host and run inside the VM:
- KUnit: rv_mon suite (CONFIG_RV_MONITORS_KUNIT_TEST), rv_test_tlob case
covering the uprobe-line parser
- 8 ftrace-style selftests: uprobe_bind, uprobe_violation,
uprobe_no_event, uprobe_multi, uprobe_detail_{running,sleeping,
waiting}, uprobe_restart
Both suites pass. The pre-existing rv_* tests in the same directory
keep passing as well.
v4: https://lore.kernel.org/all/cover.1783524627.git.wen.yang@linux.dev/
v3: https://lore.kernel.org/all/cover.1780847473.git.wen.yang@linux.dev/
v2: https://lore.kernel.org/all/cover.1778522945.git.wen.yang@linux.dev/
v1: https://lore.kernel.org/all/cover.1776020428.git.wen.yang@linux.dev/
Wen Yang (9):
rv: Introduce DA_MON_ALLOCATION_STRATEGY
rv: Add generic uprobe infrastructure for RV monitors
rv: Add tlob model DOT file
rv: Fix ha_invariant_passed_ns silent bypass of invariant check
rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
rv: Add tlob hybrid automaton monitor
rv: Add KUnit tests for the tlob monitor
selftests/verification: Add tlob selftests
selftests/ftrace: Walk up to find test.d/functions when a subdirectory
is passed
Documentation/trace/rv/index.rst | 1 +
Documentation/trace/rv/monitor_tlob.rst | 194 +++
include/rv/da_monitor.h | 152 ++-
include/rv/ha_monitor.h | 13 +-
include/rv/rv_uprobe.h | 90 ++
kernel/trace/rv/Kconfig | 5 +
kernel/trace/rv/Makefile | 2 +
kernel/trace/rv/monitors/tlob/Kconfig | 12 +
kernel/trace/rv/monitors/tlob/tlob.c | 1145 +++++++++++++++++
kernel/trace/rv/monitors/tlob/tlob.h | 149 +++
kernel/trace/rv/monitors/tlob/tlob_kunit.c | 87 ++
kernel/trace/rv/monitors/tlob/tlob_kunit.h | 17 +
kernel/trace/rv/monitors/tlob/tlob_trace.h | 48 +
kernel/trace/rv/rv_monitors_test.c | 2 +
kernel/trace/rv/rv_trace.h | 1 +
kernel/trace/rv/rv_uprobe.c | 91 ++
tools/testing/selftests/ftrace/ftracetest | 26 +-
.../testing/selftests/verification/.gitignore | 2 +
tools/testing/selftests/verification/Makefile | 4 +
.../test.d/tlob/run_tlob_tests.sh | 23 +
.../verification/test.d/tlob/uprobe_bind.tc | 39 +
.../test.d/tlob/uprobe_detail_running.tc | 53 +
.../test.d/tlob/uprobe_detail_sleeping.tc | 52 +
.../test.d/tlob/uprobe_detail_waiting.tc | 76 ++
.../verification/test.d/tlob/uprobe_multi.tc | 63 +
.../test.d/tlob/uprobe_no_event.tc | 17 +
.../test.d/tlob/uprobe_restart.tc | 79 ++
.../test.d/tlob/uprobe_violation.tc | 69 +
.../testing/selftests/verification/tlob_sym.c | 225 ++++
.../selftests/verification/tlob_target.c | 138 ++
tools/verification/models/tlob.dot | 24 +
31 files changed, 2884 insertions(+), 15 deletions(-)
create mode 100644 Documentation/trace/rv/monitor_tlob.rst
create mode 100644 include/rv/rv_uprobe.h
create mode 100644 kernel/trace/rv/monitors/tlob/Kconfig
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.h
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.h
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_trace.h
create mode 100644 kernel/trace/rv/rv_uprobe.c
create mode 100755 tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
create mode 100644 tools/testing/selftests/verification/tlob_sym.c
create mode 100644 tools/testing/selftests/verification/tlob_target.c
create mode 100644 tools/verification/models/tlob.dot
base-commit: 785095112f4198de49760552374f364043c8dbdf
--
2.25.1
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v6 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:45 ` [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
` (7 subsequent siblings)
8 siblings, 0 replies; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Per-object DA storage allocation is currently limited to kmalloc on
demand. Add a compile-time selector so monitors can choose among three
strategies:
DA_ALLOC_AUTO (default) - kmalloc per object on the monitor path
DA_ALLOC_POOL - pre-allocated fixed-size mempool;
selected by defining DA_MON_POOL_SIZE
DA_ALLOC_MANUAL - caller pre-inserts storage; framework
only links the target field
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/da_monitor.h | 152 +++++++++++++++++++++++++++++++++++++---
1 file changed, 142 insertions(+), 10 deletions(-)
diff --git a/include/rv/da_monitor.h b/include/rv/da_monitor.h
index e3cf85c9ce55..9e84f53b81cb 100644
--- a/include/rv/da_monitor.h
+++ b/include/rv/da_monitor.h
@@ -14,7 +14,54 @@
#ifndef _RV_DA_MONITOR_H
#define _RV_DA_MONITOR_H
+/*
+ * Allocation strategies for RV_MON_PER_OBJ monitors, selected by defining
+ * one of the following before including this header (never both):
+ *
+ * DA_MON_POOL_SIZE N - pool mode, N pre-allocated slots
+ * DA_MON_ALLOCATION_STRATEGY - explicit strategy (see below)
+ * (neither) - auto mode (default)
+ *
+ * DA_ALLOC_AUTO - kmalloc on demand; unbounded.
+ * DA_ALLOC_POOL - pre-allocated fixed-size pool; capped at DA_MON_POOL_SIZE
+ * entries. Uses mempool_alloc_preallocated() (spinlock_t);
+ * the start-event handler must run in task context.
+ * Other event handlers may run in any context as they use
+ * da_handle_event() and never allocate new storage.
+ * DA_ALLOC_MANUAL - caller pre-inserts storage; framework links the target.
+ */
+#define DA_ALLOC_AUTO 0
+#define DA_ALLOC_POOL 1
+#define DA_ALLOC_MANUAL 2
+
+#ifdef DA_MON_POOL_SIZE
+#ifdef DA_MON_ALLOCATION_STRATEGY
+#error "Define only one of DA_MON_POOL_SIZE or DA_MON_ALLOCATION_STRATEGY"
+#endif
+#if DA_MON_POOL_SIZE == 0
+#error "DA_MON_POOL_SIZE must be non-zero"
+#endif
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_POOL
+#endif /* DA_MON_POOL_SIZE */
+
+#ifndef DA_MON_ALLOCATION_STRATEGY
+#ifdef DA_SKIP_AUTO_ALLOC
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_MANUAL
+#else
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_AUTO
+#endif
+#endif /* DA_MON_ALLOCATION_STRATEGY */
+
+/* Zero default keeps pool-mode conditionals compile-time constant. */
+#ifndef DA_MON_POOL_SIZE
+#if DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL
+#error "DA_ALLOC_POOL requires DA_MON_POOL_SIZE to be defined and non-zero"
+#endif
+#define DA_MON_POOL_SIZE 0
+#endif /* DA_MON_POOL_SIZE */
+
#include <rv/automata.h>
+#include <linux/mempool.h>
#include <linux/rv.h>
#include <rv/kunit.h>
#include <linux/stringify.h>
@@ -67,6 +114,16 @@ static struct rv_monitor rv_this;
#define da_monitor_sync_hook()
#endif
+/*
+ * Per-object teardown hook, called after da_monitor_reset_all() +
+ * da_monitor_sync_hook() and before hash_del_rcu() for each entry.
+ * All HA timer callbacks have completed at this point.
+ * Define before including this header. Default: no-op.
+ */
+#ifndef da_extra_cleanup
+#define da_extra_cleanup(da_mon)
+#endif
+
/*
* Type for the target id, default to int but can be overridden.
* A long type can work as hash table key (PER_OBJ) but will be downgraded to
@@ -543,6 +600,59 @@ static inline monitor_target da_get_target_by_id(da_id_type id)
return mon_storage->target;
}
+/*
+ * Pre-allocated mempool for DA_ALLOC_POOL monitors: DA_MON_POOL_SIZE
+ * slots, eager-allocated at init. mempool_alloc_preallocated() pops a
+ * slot without touching the allocator (bounded start latency; NULL when
+ * exhausted). mempool_free() is safe from RCU-callback context.
+ * Non-pool monitors get a zero-initialised mempool_t; pool paths compile
+ * away.
+ */
+static mempool_t da_monitor_pool;
+
+static void da_pool_return_cb(struct rcu_head *head)
+{
+ struct da_monitor_storage *ms =
+ container_of(head, struct da_monitor_storage, rcu);
+
+ mempool_free(ms, &da_monitor_pool);
+}
+
+/*
+ * da_create_pool_storage - pop a free pool slot and insert it into the hash.
+ *
+ * Returns the new da_monitor, or NULL if the pool is exhausted. Finding
+ * an existing entry for the same id fires WARN_ON_ONCE (double-start bug).
+ *
+ * Caller must hold an RCU read-side CS and the monitor's serialisation lock.
+ */
+static inline struct da_monitor *
+da_create_pool_storage(da_id_type id, monitor_target target,
+ struct da_monitor *da_mon)
+{
+ struct da_monitor_storage *mon_storage, *existing;
+
+ if (da_mon)
+ return da_mon;
+
+ mon_storage = mempool_alloc_preallocated(&da_monitor_pool);
+ if (!mon_storage)
+ return NULL;
+ memset(mon_storage, 0, sizeof(*mon_storage));
+
+ mon_storage->id = id;
+ mon_storage->target = target;
+
+ /* Single consumer under the caller's lock; duplicate is a double-start bug. */
+ existing = __da_get_mon_storage(id);
+ if (WARN_ON_ONCE(existing)) {
+ mempool_free(mon_storage, &da_monitor_pool);
+ return NULL;
+ }
+ hash_add_rcu(da_monitor_ht, &mon_storage->node, id);
+ return &mon_storage->rv.da_mon;
+}
+
/*
* da_destroy_storage - destroy the per-object storage
*
@@ -564,7 +674,10 @@ static inline void da_destroy_storage(da_id_type id)
return;
da_monitor_reset_hook(&mon_storage->rv.da_mon);
hash_del_rcu(&mon_storage->node);
- kfree_rcu(mon_storage, rcu);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ call_rcu(&mon_storage->rcu, da_pool_return_cb);
+ else
+ kfree_rcu(mon_storage, rcu);
}
static void __da_monitor_reset_all(void (*reset)(struct da_monitor *))
@@ -590,6 +703,9 @@ static inline void da_monitor_reset_state_all(void)
static inline int da_monitor_init(void)
{
hash_init(da_monitor_ht);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ return mempool_init_kmalloc_pool(&da_monitor_pool, DA_MON_POOL_SIZE,
+ sizeof(struct da_monitor_storage));
return 0;
}
@@ -607,21 +723,37 @@ static inline void da_monitor_destroy(void)
* pending, we can safely assume no concurrent user.
*/
hash_for_each_safe(da_monitor_ht, bkt, tmp, mon_storage, node) {
+ da_extra_cleanup(&mon_storage->rv.da_mon);
hash_del_rcu(&mon_storage->node);
- kfree(mon_storage);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ mempool_free(mon_storage, &da_monitor_pool);
+ else
+ kfree(mon_storage);
+ }
+
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL) {
+ rcu_barrier();
+ mempool_exit(&da_monitor_pool);
}
}
/*
- * Allow the per-object monitors to run allocation manually, necessary if the
- * start condition is in a context problematic for allocation (e.g. scheduling).
- * In such case, if the storage was pre-allocated without a target, set it now.
+ * da_prepare_storage - allocate or link per-object monitor storage.
+ *
+ * Called only from da_handle_start_run_event(); must run in task context
+ * for DA_ALLOC_AUTO and DA_ALLOC_POOL (both take a spinlock_t internally).
+ * Subsequent event handlers use da_handle_event() and never allocate.
*/
-#ifdef DA_SKIP_AUTO_ALLOC
-#define da_prepare_storage da_fill_empty_storage
-#else
-#define da_prepare_storage da_create_storage
-#endif /* DA_SKIP_AUTO_ALLOC */
+static inline struct da_monitor *
+da_prepare_storage(da_id_type id, monitor_target target,
+ struct da_monitor *da_mon)
+{
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ return da_create_pool_storage(id, target, da_mon);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_MANUAL)
+ return da_fill_empty_storage(id, target, da_mon);
+ return da_create_storage(id, target, da_mon);
+}
#endif /* RV_MON_TYPE */
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-20 16:45 ` [PATCH v6 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:59 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 3/9] rv: Add tlob model DOT file wen.yang
` (6 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Monitors that instrument user-space function boundaries need to resolve
paths, register uprobes, and deregister them safely. Provide a thin
wrapper so monitors share a single implementation of this boilerplate.
struct rv_uprobe embeds struct uprobe_consumer directly, avoiding a
separate heap allocation per probe. The struct holds a struct path for
the probed binary so that the inode and its mount remain referenced for
the full uprobe lifetime; uprobe_register() does not take its own
reference to the inode. The path is released in
rv_uprobe_unregister_nosync() after the consumer has been removed.
rv_uprobe_sync() calls uprobe_unregister_sync() which performs
synchronize_rcu_tasks_trace(), waiting for all rcu_read_lock_trace()
readers (handler_chain()) to complete on all CPUs before returning;
the caller may then free the containing struct.
The API provides register, synchronous and nosync unregister, a global
handler barrier (rv_uprobe_sync), and an active-state predicate.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/rv_uprobe.h | 90 ++++++++++++++++++++++++++++++++++++
kernel/trace/rv/rv_uprobe.c | 91 +++++++++++++++++++++++++++++++++++++
2 files changed, 181 insertions(+)
create mode 100644 include/rv/rv_uprobe.h
create mode 100644 kernel/trace/rv/rv_uprobe.c
diff --git a/include/rv/rv_uprobe.h b/include/rv/rv_uprobe.h
new file mode 100644
index 000000000000..d0a9079ac5be
--- /dev/null
+++ b/include/rv/rv_uprobe.h
@@ -0,0 +1,90 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (C) 2026 Wen Yang <wen.yang@linux.dev> */
+/*
+ * Generic uprobe infrastructure for RV monitors.
+ *
+ */
+
+#ifndef _RV_UPROBE_H
+#define _RV_UPROBE_H
+
+#include <linux/path.h>
+#include <linux/types.h>
+#include <linux/uprobes.h>
+
+struct pt_regs;
+
+/**
+ * struct rv_uprobe - embeddable uprobe handle for RV monitors
+ *
+ * Embed via DECLARE_RV_UPROBE() and pass &name to rv_uprobe_register().
+ * The caller may free the containing struct after rv_uprobe_unregister()
+ * (or rv_uprobe_unregister_nosync() + rv_uprobe_sync()) returns.
+ *
+ * @uc: embedded uprobe_consumer; set handler/ret_handler before registering
+ * @uprobe: registered uprobe pointer (NULL when not registered)
+ * @path: path of the probed binary, held until unregistration
+ */
+struct rv_uprobe {
+ struct uprobe_consumer uc;
+ struct uprobe *uprobe;
+ struct path path;
+};
+
+/* Embed a named rv_uprobe inside a caller struct */
+#define DECLARE_RV_UPROBE(name) struct rv_uprobe name
+
+/**
+ * rv_uprobe_is_registered - test whether an uprobe is currently active
+ * @p: probe to test; may be NULL
+ */
+bool rv_uprobe_is_registered(const struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_register - initialise and register an uprobe
+ * @binpath: absolute path to the target binary
+ * @offset: byte offset within the binary
+ * @p: caller-provided rv_uprobe (embedded via DECLARE_RV_UPROBE);
+ * p->uc.handler and/or p->uc.ret_handler must be set before this call
+ *
+ * Resolves the path and registers p->uc with the uprobe subsystem.
+ * No heap allocation is performed.
+ *
+ * Returns 0 on success, negative errno on failure.
+ */
+int rv_uprobe_register(const char *binpath, loff_t offset, struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_unregister - synchronously unregister a uprobe
+ * @p: probe to unregister; may be NULL (no-op)
+ *
+ * Removes the consumer from the uprobe subsystem and waits for all in-flight
+ * handlers to complete (via synchronize_rcu_tasks_trace()). After this
+ * returns, the containing struct may be safely freed by the caller.
+ * Use rv_uprobe_unregister_nosync() + rv_uprobe_sync() to batch multiple
+ * deregistrations before a single synchronisation.
+ */
+void rv_uprobe_unregister(struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_unregister_nosync - dequeue an uprobe without waiting
+ * @p: probe to dequeue; may be NULL (no-op)
+ *
+ * Removes the consumer without waiting for in-flight handlers. The path
+ * (p->path) is NOT released here; the caller must call rv_uprobe_sync()
+ * followed by path_put(&p->path) before freeing the containing struct.
+ * Use rv_uprobe_unregister() to handle both in one step.
+ */
+void rv_uprobe_unregister_nosync(struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_sync - wait for all in-flight uprobe handlers to complete
+ *
+ * Global barrier: calls uprobe_unregister_sync(), which runs
+ * synchronize_rcu_tasks_trace() and synchronize_srcu(&uretprobes_srcu).
+ * After this returns, no handler_chain() iteration referencing any
+ * previously deregistered consumer is still in progress.
+ */
+void rv_uprobe_sync(void);
+
+#endif /* _RV_UPROBE_H */
diff --git a/kernel/trace/rv/rv_uprobe.c b/kernel/trace/rv/rv_uprobe.c
new file mode 100644
index 000000000000..b412a8e28a6e
--- /dev/null
+++ b/kernel/trace/rv/rv_uprobe.c
@@ -0,0 +1,91 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Generic uprobe infrastructure for RV monitors.
+ *
+ * rv_uprobe embeds struct uprobe_consumer; rv_uprobe_sync() drains in-flight
+ * handlers before the containing struct may be freed (see rv_uprobe.h).
+ */
+#include <linux/dcache.h>
+#include <linux/fs.h>
+#include <linux/namei.h>
+#include <linux/uprobes.h>
+#include <rv/rv_uprobe.h>
+
+/**
+ * rv_uprobe_register - initialise and register an uprobe
+ */
+int rv_uprobe_register(const char *binpath, loff_t offset, struct rv_uprobe *p)
+{
+ struct inode *inode;
+ int ret;
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &p->path);
+ if (ret)
+ return ret;
+
+ if (!d_is_reg(p->path.dentry)) {
+ path_put(&p->path);
+ return -EINVAL;
+ }
+
+ inode = d_real_inode(p->path.dentry);
+
+ /* uprobe_register() takes no inode reference; the path is held in p->path */
+ p->uprobe = uprobe_register(inode, offset, 0, &p->uc);
+ if (IS_ERR(p->uprobe)) {
+ ret = PTR_ERR(p->uprobe);
+ p->uprobe = NULL;
+ path_put(&p->path);
+ return ret;
+ }
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_register);
+
+/**
+ * rv_uprobe_is_registered - test whether an uprobe is currently active
+ */
+bool rv_uprobe_is_registered(const struct rv_uprobe *p)
+{
+ return p && p->uprobe;
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_is_registered);
+
+/**
+ * rv_uprobe_unregister - synchronously unregister a uprobe
+ */
+void rv_uprobe_unregister(struct rv_uprobe *p)
+{
+ if (!p || !p->uprobe)
+ return;
+
+ uprobe_unregister_nosync(p->uprobe, &p->uc);
+ p->uprobe = NULL;
+ rv_uprobe_sync();
+ path_put(&p->path);
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_unregister);
+
+/**
+ * rv_uprobe_unregister_nosync - dequeue an uprobe without waiting
+ */
+void rv_uprobe_unregister_nosync(struct rv_uprobe *p)
+{
+ if (!p || !p->uprobe)
+ return;
+
+ uprobe_unregister_nosync(p->uprobe, &p->uc);
+ p->uprobe = NULL;
+ /* path held; caller must call rv_uprobe_sync() then path_put(&p->path) */
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_unregister_nosync);
+
+/**
+ * rv_uprobe_sync - wait for all in-flight uprobe handlers to complete
+ */
+void rv_uprobe_sync(void)
+{
+ uprobe_unregister_sync();
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_sync);
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 3/9] rv: Add tlob model DOT file
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-20 16:45 ` [PATCH v6 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
2026-08-20 16:45 ` [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:53 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
` (5 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add the Graphviz DOT specification of the tlob hybrid automaton to
tools/verification/models/. The model has three states (running,
waiting, sleeping), five transitions (switch_in, preempt, wakeup,
sleep), and a single clock invariant clk_elapsed < BUDGET_NS() active
in all states.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
tools/verification/models/tlob.dot | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
create mode 100644 tools/verification/models/tlob.dot
diff --git a/tools/verification/models/tlob.dot b/tools/verification/models/tlob.dot
new file mode 100644
index 000000000000..515695599cf0
--- /dev/null
+++ b/tools/verification/models/tlob.dot
@@ -0,0 +1,24 @@
+digraph state_automaton {
+ center = true;
+ size = "7,11";
+ {node [shape = plaintext, style=invis, label=""] "__init_stopped"};
+ {node [shape = plaintext] "running"};
+ {node [shape = plaintext] "waiting"};
+ {node [shape = plaintext] "sleeping"};
+ {node [shape = plaintext] "stopped"};
+ "__init_stopped" -> "stopped";
+ "running" [label = "running\nclk_elapsed < BUDGET_NS()", color = green3];
+ "waiting" [label = "waiting\nclk_elapsed < BUDGET_NS()"];
+ "sleeping" [label = "sleeping\nclk_elapsed < BUDGET_NS()"];
+ "stopped" [label = "stopped"];
+ "running" -> "sleeping" [ label = "sleep" ];
+ "running" -> "waiting" [ label = "preempt" ];
+ "waiting" -> "running" [ label = "switch_in" ];
+ "sleeping" -> "waiting" [ label = "wakeup" ];
+ "running" -> "stopped" [ label = "stop" ];
+ "stopped" -> "running" [ label = "start;reset(clk_elapsed)" ];
+ { rank = min ;
+ "__init_stopped";
+ "stopped";
+ }
+}
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (2 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 3/9] rv: Add tlob model DOT file wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:58 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
` (4 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
When env_store is U64_MAX (its initial sentinel value),
ha_invariant_passed_ns() returns 0 immediately without initializing
env_store to the current clock. Subsequent calls to
ha_check_invariant_ns() then find env_store still at U64_MAX, causing
the elapsed comparison to wrap and always report the invariant as
satisfied, silently masking any violations.
Fix by calling ha_reset_clk_ns() to establish the guard on the first
invocation instead of returning early. Apply the same fix to
ha_invariant_passed_jiffy().
This is a stopgap: once the RV framework reworks the per-env clock
guard, this first-invocation reset should be subsumed.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/ha_monitor.h | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
index 6e1c7fe5449a..e1738d199b28 100644
--- a/include/rv/ha_monitor.h
+++ b/include/rv/ha_monitor.h
@@ -355,7 +355,7 @@ static inline u64 ha_invariant_passed_ns(struct ha_monitor *ha_mon, enum envs en
if (env < 0 || env >= ENV_MAX_STORED)
return 0;
if (ha_monitor_env_invalid(ha_mon, env))
- return 0;
+ ha_reset_clk_ns(ha_mon, env, time_ns);
return ha_get_env(ha_mon, env, time_ns);
}
@@ -375,6 +375,7 @@ static inline bool ha_check_invariant_jiffy(struct ha_monitor *ha_mon, enum envs
{
return time_after64(READ_ONCE(ha_mon->env_store[env]), get_jiffies_64() - expire_jiffy);
}
+
/*
* ha_invariant_passed_jiffy - prepare the invariant and return the time since reset
*/
@@ -383,7 +384,7 @@ static inline u64 ha_invariant_passed_jiffy(struct ha_monitor *ha_mon, enum envs
if (env < 0 || env >= ENV_MAX_STORED)
return 0;
if (ha_monitor_env_invalid(ha_mon, env))
- return 0;
+ ha_reset_clk_jiffy(ha_mon, env);
return ha_get_env(ha_mon, env, time_ns);
}
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (3 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:59 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor wen.yang
` (3 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Wrap both definitions with #ifndef guards so HA-based monitors can
substitute their own implementations before including this header.
tlob uses this to define a reset hook that cancels per-task hrtimers
on monitor teardown.
Overrides must still call ha_monitor_reset_env() or cancel outstanding
timers to avoid timer UAF.
No behaviour change for monitors that do not override either macro.
Reviewed-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/ha_monitor.h | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
index e1738d199b28..807b981eb548 100644
--- a/include/rv/ha_monitor.h
+++ b/include/rv/ha_monitor.h
@@ -36,8 +36,14 @@ static bool ha_monitor_handle_constraint(struct da_monitor *da_mon,
da_id_type id);
#define da_monitor_event_hook ha_monitor_handle_constraint
#define da_monitor_init_hook ha_monitor_init_env
+
+/* Overrides must still call ha_monitor_reset_env() or cancel the timer. */
+#ifndef da_monitor_reset_hook
#define da_monitor_reset_hook ha_monitor_reset_env
+#endif
+#ifndef da_monitor_sync_hook
#define da_monitor_sync_hook() synchronize_rcu()
+#endif
#if !defined(HA_SKIP_AUTO_CLEANUP) && RV_MON_TYPE == RV_MON_PER_TASK
/*
@@ -75,7 +81,9 @@ _Static_assert(offsetof(struct ha_monitor, da_mon) == 0,
#define ENV_INVALID_VALUE U64_MAX
/* Error with no event occurs only on timeouts */
#define EVENT_NONE EVENT_MAX
+#ifndef EVENT_NONE_LBL
#define EVENT_NONE_LBL "none"
+#endif
#define ENV_BUFFER_SIZE 64
#ifdef CONFIG_RV_REACTORS
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (4 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 17:03 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
` (2 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
tlob (task latency over budget) is a per-task hybrid automaton RV
monitor that tracks wall-clock time across a user-delimited code
section and emits an error when elapsed time exceeds a configurable
budget.
Four-state automaton (running, waiting, sleeping, stopped) driven by
sched_switch/sched_wakeup tracepoints and a user-visible start/stop
pair: "stop" only fires from running and parks the window in stopped,
where a later "start" restarts it in place (same pool slot, no
re-registration); both callers of stop run while the task is on CPU.
A single clk_elapsed < BUDGET_NS() invariant is enforced by a
per-task HRTIMER_MODE_REL_HARD timer; on expiry the monitor records a
per-state breakdown (running_ns, waiting_ns, sleeping_ns) before
emitting error_env_tlob.
Uprobe pairs are registered through a tracefs monitor file as
"p PATH:OFFSET_START OFFSET_STOP threshold=NS". A pre-allocated
mempool hard-caps concurrently monitored tasks at TLOB_MAX_MONITORED
(past it, fresh starts return -ENOSPC) with allocation-free start/stop
on the uprobe hot path.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
Documentation/trace/rv/index.rst | 1 +
Documentation/trace/rv/monitor_tlob.rst | 194 ++++
kernel/trace/rv/Kconfig | 5 +
kernel/trace/rv/Makefile | 2 +
kernel/trace/rv/monitors/tlob/Kconfig | 12 +
kernel/trace/rv/monitors/tlob/tlob.c | 1137 ++++++++++++++++++++
kernel/trace/rv/monitors/tlob/tlob.h | 149 +++
kernel/trace/rv/monitors/tlob/tlob_trace.h | 48 +
kernel/trace/rv/rv_trace.h | 1 +
9 files changed, 1549 insertions(+)
create mode 100644 Documentation/trace/rv/monitor_tlob.rst
create mode 100644 kernel/trace/rv/monitors/tlob/Kconfig
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.h
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_trace.h
diff --git a/Documentation/trace/rv/index.rst b/Documentation/trace/rv/index.rst
index 29769f06bb0f..1501545b5f08 100644
--- a/Documentation/trace/rv/index.rst
+++ b/Documentation/trace/rv/index.rst
@@ -16,5 +16,6 @@ Runtime Verification
monitor_wwnr.rst
monitor_sched.rst
monitor_rtapp.rst
+ monitor_tlob.rst
monitor_stall.rst
monitor_deadline.rst
diff --git a/Documentation/trace/rv/monitor_tlob.rst b/Documentation/trace/rv/monitor_tlob.rst
new file mode 100644
index 000000000000..2e606b0e67a4
--- /dev/null
+++ b/Documentation/trace/rv/monitor_tlob.rst
@@ -0,0 +1,194 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+Monitor tlob
+============
+
+- Name: tlob - task latency over budget
+- Type: per-object hybrid automaton (RV_MON_PER_OBJ)
+- Author: Wen Yang <wen.yang@linux.dev>
+
+Description
+-----------
+
+The tlob monitor tracks per-task elapsed wall-clock time (CLOCK_MONOTONIC,
+spanning running, waiting, and sleeping states) and reports a violation when
+the monitored task exceeds a configurable per-invocation budget threshold.
+
+The monitor implements a four-state hybrid automaton with a single clock
+environment variable ``clk_elapsed``. The clock invariant
+``clk_elapsed < BUDGET_NS()`` is active in the ``running``, ``waiting``, and
+``sleeping`` states (``stopped`` has no invariant, hence no timer); when it
+is violated the HA timer fires and the framework emits ``error_env_tlob``
+then calls ``da_monitor_reset()`` automatically::
+
+ | (initial)
+ v
+ +--------------+ +----------+
+ | running | --------> | stopped |
+ |->+--------------+ <-------- +----------+
+ switch_in preempt sleep
+ | | |
+ | | |
+ | v v
+ +---------+ +---------+
+ | waiting | | sleeping|
+ +---------+ +---------+
+ ^ v
+ | wakeup |
+ | |
+ +------------+
+
+ A fourth state, ``stopped``, has no clock invariant (hence no timer).
+ ``running`` reaches it on ``stop`` (``tlob_stop_task()``, window ended,
+ per-task state parked rather than freed) and returns to ``running`` on
+ ``start`` (``tlob_start_task()`` restarting the same task's parked
+ window).
+
+ Key transitions:
+ running --(sleep)------> sleeping (task blocks waiting for a resource)
+ running --(preempt)----> waiting (task preempted, back in runqueue)
+ sleeping --(wakeup)-----> waiting (resource available, enters runqueue)
+ waiting --(switch_in)--> running (scheduler picks task, back on CPU)
+ running --(stop)-------> stopped (tlob_stop_task(): window ended, parked)
+ stopped --(start)------> running (tlob_start_task(): window restarted)
+
+ ``tlob_start_task()`` calls ``da_handle_start_run_event(task->pid, ws, start_tlob)``.
+ The ``start_tlob`` edge goes ``stopped`` -> ``running`` for both a fresh
+ allocation (the initial state is ``stopped``) and a parked window's restart;
+ there is no ``start`` self-loop on ``running`` (a running task's START is
+ rejected with ``-EALREADY``). The transition triggers ``ha_setup_invariants()``,
+ which anchors ``clk_elapsed`` and arms the budget timer automatically.
+ ``tlob_stop_task()`` cancels the HA timer synchronously
+ via ``ha_cancel_timer_sync()``, then dispatches the ``stop_tlob`` event
+ (running -> stopped) instead of resetting the monitor: the per-task state
+ is parked, not freed, so a later ``tlob_start_task()`` call for the same
+ task can restart it without reallocating. Final teardown (task exit,
+ uprobe unbind, or monitor disable) is what actually calls
+ ``da_monitor_reset()`` and frees the state.
+
+The non-running condition (monitor not yet started, or reset after a budget
+violation or monitor disable) is handled implicitly by the RV framework
+(``da_mon->monitoring == 0``) - it is not an explicit DA state. A
+``tlob_stop_task()`` does not reset the monitor: the window parks in the
+explicit ``stopped`` state.
+
+Per-task state lives in ``struct tlob_task_state`` which is stored as
+``monitor_target`` in the framework's ``da_monitor_storage``, indexed by
+pid. The per-invocation ``threshold_ns`` is read via
+``ha_get_target(ha_mon)->threshold_ns`` inside the HA constraint functions,
+following the same pattern as the ``nomiss`` monitor.
+
+Usage
+-----
+
+tracefs interface (uprobe-based external monitoring)
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``monitor`` tracefs file instruments an unmodified binary via uprobes.
+The format follows the ftrace ``uprobe_events`` convention (``PATH:OFFSET``
+for the probe location, ``key=value`` for configuration parameters)::
+
+ p PATH:OFFSET_START OFFSET_STOP threshold=NS
+
+The uprobe at ``OFFSET_START`` fires ``tlob_start_task()``; the uprobe at
+``OFFSET_STOP`` fires ``tlob_stop_task()``. Both offsets are ELF file
+offsets of entry points in ``PATH``. ``PATH`` may contain ``:``; the last
+``:`` in the ``PATH:OFFSET_START`` token is the separator.
+
+A given task may hit the START/STOP pair any number of times: each pair
+of hits is one independent measurement window, and the underlying
+per-task state is reused across windows rather than reallocated each
+time. Removing a binding while one of its tasks is between windows
+(parked, having already hit STOP) frees that task's state immediately;
+a task still inside a window when its binding is removed keeps running
+unaffected and is cleaned up normally when it next exits.
+
+To remove a binding, use ``-PATH:OFFSET_START``::
+
+ echo 1 > /sys/kernel/tracing/rv/monitors/tlob/enable
+
+ echo "p /usr/bin/myapp:0x12a0 0x12f0 threshold=5000000" \
+ > /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # Remove a binding
+ echo "-/usr/bin/myapp:0x12a0" > /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # List registered bindings
+ cat /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # Read violations from the trace buffer
+ cat /sys/kernel/tracing/trace
+
+Violation tracepoints
+~~~~~~~~~~~~~~~~~~~~~
+
+Two tracepoints are emitted together on a budget violation:
+
+``error_env_tlob``
+ Standard HA clock-invariant tracepoint (emitted by the RV framework).
+ Fields: ``id`` (task pid), ``state``, ``event`` (``"budget_exceeded"``),
+ ``env`` (``"clk_elapsed"``).
+
+``detail_env_tlob``
+ Tlob-specific breakdown of elapsed time per DA state.
+ Fields: ``id`` (task pid), ``threshold_ns``, ``running_ns``,
+ ``waiting_ns``, ``sleeping_ns``.
+
+ Use ``detail_env_tlob`` to diagnose *which phase* consumed the budget:
+ high ``sleeping_ns`` indicates I/O latency; high ``waiting_ns`` indicates
+ scheduler pressure; high ``running_ns`` indicates a compute overrun.
+
+Example: correlate the two tracepoints to see the breakdown::
+
+ trace-cmd record -e error_env_tlob -e detail_env_tlob &
+ # ... run workload ...
+ trace-cmd report
+
+tracefs files
+~~~~~~~~~~~~~
+
+The following files are specific to tlob under
+``/sys/kernel/tracing/rv/monitors/tlob/``:
+
+``monitor`` (rw)
+ Write ``p PATH:OFFSET_START OFFSET_STOP threshold=NS``
+ to bind two entry uprobes. Write ``-PATH:OFFSET_START`` to remove a
+ binding. Read to list registered bindings in the same format.
+ See the `tracefs interface (uprobe-based external monitoring)`_ section above.
+
+Kernel API
+----------
+
+``tlob_start_task`` and ``tlob_stop_task`` are the implementation-level
+functions called by the uprobe entry/exit handlers; the interface is
+driven from userspace.
+
+.. kernel-doc:: kernel/trace/rv/monitors/tlob/tlob.c
+ :functions: tlob_start_task tlob_stop_task
+
+Design notes
+------------
+
+Limitations:
+
+- A fresh window dispatches ``start_tlob`` (initial ``stopped`` ->
+ ``running``) via ``da_handle_start_run_event(task->pid, ws, start_tlob)``,
+ so monitoring always begins in ``running``. Monitoring a non-current
+ task that is already in waiting or sleeping state at call time
+ misclassifies the first interval as ``running_ns``.
+- ``TASK_STOPPED`` and ``TASK_TRACED`` carry ``prev_state != 0`` and are
+ therefore counted as ``sleeping_ns``, indistinguishable from
+ I/O-blocked time.
+- ``sched_wakeup_new`` is not hooked. In practice this is not an issue
+ because ``tlob_start_task`` is always called from a running context.
+
+Specification
+-------------
+
+Graphviz DOT file in tools/verification/models/tlob.dot.
+
+KUnit tests under ``kernel/trace/rv/monitors/tlob/tlob_kunit.c``
+(CONFIG_TLOB_KUNIT_TEST).
+
+User-space integration tests under ``tools/testing/selftests/verification/``
+(requires CONFIG_RV_MON_TLOB=y and root).
diff --git a/kernel/trace/rv/Kconfig b/kernel/trace/rv/Kconfig
index efa930f94ea4..222b3bea8079 100644
--- a/kernel/trace/rv/Kconfig
+++ b/kernel/trace/rv/Kconfig
@@ -84,8 +84,13 @@ source "kernel/trace/rv/monitors/deadline/Kconfig"
source "kernel/trace/rv/monitors/nomiss/Kconfig"
# Add new deadline monitors here
+source "kernel/trace/rv/monitors/tlob/Kconfig"
# Add new monitors here
+config RV_UPROBE
+ bool
+ depends on RV && UPROBES
+
config RV_REACTORS
bool "Runtime verification reactors"
default y
diff --git a/kernel/trace/rv/Makefile b/kernel/trace/rv/Makefile
index cdbf68c84f5a..cd0ec11f0e05 100644
--- a/kernel/trace/rv/Makefile
+++ b/kernel/trace/rv/Makefile
@@ -21,7 +21,9 @@ obj-$(CONFIG_RV_MON_STALL) += monitors/stall/stall.o
obj-$(CONFIG_RV_MON_DEADLINE) += monitors/deadline/deadline.o
obj-$(CONFIG_RV_MON_NOMISS) += monitors/nomiss/nomiss.o
obj-$(CONFIG_RV_MON_WAKEUP) += monitors/wakeup/wakeup.o
+obj-$(CONFIG_RV_MON_TLOB) += monitors/tlob/tlob.o
# Add new monitors here
+obj-$(CONFIG_RV_UPROBE) += rv_uprobe.o
obj-$(CONFIG_RV_REACTORS) += rv_reactors.o
obj-$(CONFIG_RV_REACT_PRINTK) += reactor_printk.o
obj-$(CONFIG_RV_REACT_PANIC) += reactor_panic.o
diff --git a/kernel/trace/rv/monitors/tlob/Kconfig b/kernel/trace/rv/monitors/tlob/Kconfig
new file mode 100644
index 000000000000..aa43382073d2
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/Kconfig
@@ -0,0 +1,12 @@
+# SPDX-License-Identifier: GPL-2.0-only
+#
+config RV_MON_TLOB
+ bool "tlob monitor"
+ depends on RV && UPROBES && HIGH_RES_TIMERS
+ select HA_MON_EVENTS_ID
+ select RV_UPROBE
+ help
+ Enable the tlob (task latency over budget) hybrid-automaton RV
+ monitor. tlob tracks per-task elapsed wall-clock time across a
+ user-delimited code section and emits error_env_tlob when the
+ elapsed time exceeds a configurable per-invocation budget.
diff --git a/kernel/trace/rv/monitors/tlob/tlob.c b/kernel/trace/rv/monitors/tlob/tlob.c
new file mode 100644
index 000000000000..08b1bee884cc
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob.c
@@ -0,0 +1,1137 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob: task latency over budget monitor
+ *
+ * Tracks the elapsed wall-clock time (CLOCK_MONOTONIC) of a marked code
+ * path and flags per-task latency-budget overruns. The hrtimer callback
+ * emits error_env_tlob on violation plus detail_env_tlob, a per-state
+ * (running/waiting/sleeping) time breakdown.
+ *
+ * RV_MON_PER_OBJ: per-task state (struct tlob_task_state) lives as
+ * monitor_target in the framework's hash table. One HA clock invariant:
+ * clk_elapsed < BUDGET_NS() in running/waiting/sleeping (stopped parks).
+ *
+ * Copyright (C) 2026 Wen Yang <wen.yang@linux.dev>
+ */
+#include <linux/kernel.h>
+#include <linux/mempool.h>
+#include <linux/module.h>
+#include <linux/init.h>
+#include <linux/namei.h>
+#include <linux/rv.h>
+#include <linux/slab.h>
+#include <kunit/visibility.h>
+#include <rv/instrumentation.h>
+#include <rv/rv_uprobe.h>
+#include <rv.h>
+
+#define MODULE_NAME "tlob"
+
+#include <trace/events/sched.h>
+#include <rv_trace.h>
+
+/*
+ * Per-task latency monitoring state. One instance per monitoring window.
+ * Stored as monitor_target in da_monitor_storage; freed via call_rcu.
+ */
+enum tlob_acc_idx {
+ TLOB_ACC_RUNNING,
+ TLOB_ACC_WAITING,
+ TLOB_ACC_SLEEPING,
+ TLOB_ACC_MAX,
+};
+
+struct tlob_task_state {
+ struct task_struct *task; /* via get_task_struct */
+ u64 threshold_ns; /* budget in nanoseconds */
+
+ /*
+ * Per-window: 1 = this window ended (stop or timer expiry). Blocks
+ * timer re-arm in ha_setup_invariants(); cleared on restart.
+ */
+ atomic_t stopping;
+
+ /*
+ * Per-task, one-shot: final teardown has claimed this slot; never
+ * reset (a window can end and restart, the task cannot). atomic_t
+ * so cmpxchg is well-defined on every arch.
+ */
+ atomic_t destroying;
+
+ bool budget_exceeded;
+
+ /*
+ * Opaque owner: the binding that started this task. Set once on
+ * fresh allocation (NULL for callers with no binding), cleared by
+ * tlob_unbind_reap() for an active task whose binding is removed.
+ * Immutable elsewhere. Protected by tlob_ws_lock.
+ */
+ void *binding;
+ /*
+ * Linked into binding->started_list for the whole lifetime (not just
+ * while parked) so unbind reaping finds parked and active tasks.
+ * Protected by tlob_ws_lock.
+ */
+ struct list_head started_node;
+
+ /* Serialises accs_ns[]; held briefly (hardirq-safe). */
+ raw_spinlock_t entry_lock;
+ u64 accs_ns[TLOB_ACC_MAX]; /* per-state elapsed ns */
+ ktime_t last_ts;
+
+ struct rcu_head rcu;
+};
+
+#define RV_MON_TYPE RV_MON_PER_OBJ
+#define HA_TIMER_TYPE HA_TIMER_HRTIMER
+
+typedef struct tlob_task_state *monitor_target;
+
+static inline void tlob_reset_notify(struct da_monitor *da_mon);
+#define da_monitor_reset_hook tlob_reset_notify
+
+static inline void tlob_extra_cleanup(struct da_monitor *da_mon);
+#define da_extra_cleanup tlob_extra_cleanup
+
+#define EVENT_NONE_LBL "budget_exceeded"
+
+#include "tlob.h"
+
+#define DA_MON_POOL_SIZE TLOB_MAX_MONITORED
+
+#include <rv/ha_monitor.h>
+
+/*
+ * da_monitor_reset_hook: runs on hrtimer expiry, final teardown, and
+ * monitor disable. A normal stop never resets: tlob_stop_task() dispatches
+ * "stop" (running -> stopped, tlob.dot) instead. Only timer expiry is a
+ * genuine budget violation.
+ */
+static inline void tlob_reset_notify(struct da_monitor *da_mon)
+{
+ struct ha_monitor *ha_mon = to_ha_monitor(da_mon);
+ struct tlob_task_state *ws;
+
+ ha_monitor_reset_env(da_mon);
+
+ ws = ha_get_target(ha_mon);
+ if (!ws)
+ return;
+
+ /*
+ * stopping==1 means tlob_stop_task() ended this window already.
+ * acquire pairs with the _release clear in ha_setup_invariants().
+ */
+ if (atomic_read_acquire(&ws->stopping))
+ return;
+
+ /*
+ * Monitor disable (ha_mon_destroying set) is not a violation: the
+ * teardown paths free ws regardless. Couples to an HA-layer flag
+ * with no public contract; a framework-level equivalent would be
+ * cleaner.
+ */
+ if (unlikely(READ_ONCE(ha_mon_destroying)))
+ return;
+
+ /* Genuine expiry: end the window so a later start takes the restart path. */
+ atomic_set(&ws->stopping, 1);
+
+ /* Stamped regardless of the tracepoint; tlob_stop_task() reads it. */
+ WRITE_ONCE(ws->budget_exceeded, true);
+
+ if (!trace_detail_env_tlob_enabled())
+ return;
+
+ unsigned int curr_state = READ_ONCE(da_mon->curr_state);
+ u64 accs[TLOB_ACC_MAX], partial_ns;
+ unsigned long flags;
+
+ /* Snapshot accumulators; partial_ns covers curr_state time not yet folded in. */
+ raw_spin_lock_irqsave(&ws->entry_lock, flags);
+ partial_ns = ktime_get_ns() - ktime_to_ns(ws->last_ts);
+ accs[TLOB_ACC_RUNNING] = ws->accs_ns[TLOB_ACC_RUNNING] +
+ (curr_state == running_tlob ? partial_ns : 0);
+ accs[TLOB_ACC_WAITING] = ws->accs_ns[TLOB_ACC_WAITING] +
+ (curr_state == waiting_tlob ? partial_ns : 0);
+ accs[TLOB_ACC_SLEEPING] = ws->accs_ns[TLOB_ACC_SLEEPING] +
+ (curr_state == sleeping_tlob ? partial_ns : 0);
+ raw_spin_unlock_irqrestore(&ws->entry_lock, flags);
+
+ trace_detail_env_tlob(da_get_id(da_mon), ws->threshold_ns,
+ accs[TLOB_ACC_RUNNING],
+ accs[TLOB_ACC_WAITING],
+ accs[TLOB_ACC_SLEEPING]);
+}
+
+#define BUDGET_NS(ha_mon) (ha_get_target(ha_mon)->threshold_ns)
+
+/* HA constraint functions (called by ha_monitor_handle_constraint) */
+
+static u64 ha_get_env(struct ha_monitor *ha_mon, enum envs_tlob env,
+ u64 time_ns)
+{
+ if (env == clk_elapsed_tlob)
+ return ha_get_clk_ns(ha_mon, env, time_ns);
+ return ENV_INVALID_VALUE;
+}
+
+/*
+ * Invariant: clk_elapsed < BUDGET_NS in running/waiting/sleeping. "stopped"
+ * is exempt: the parked period must not be measured against the old window's
+ * clock anchor (restart from "stopped" would otherwise spuriously overrun).
+ */
+static inline bool ha_verify_invariants(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (curr_state == stopped_tlob)
+ return true;
+ return ha_check_invariant_ns(ha_mon, clk_elapsed_tlob, time_ns, BUDGET_NS(ha_mon));
+}
+
+/*
+ * The clock stays in guard (anchor) representation all window: env_store
+ * holds the window-start timestamp, re-anchored on start/restart.
+ * ha_invariant_passed_ns() never stores the deadline representation (the
+ * framework dropped ha_set_invariant_ns(), commit ab2900ae252b), so calling
+ * ha_inv_to_guard() here would subtract BUDGET_NS from the anchor and skew
+ * every check by one budget. nomiss likewise never converts.
+ */
+
+/* No per-event guard conditions for tlob; invariants suffice. */
+static inline bool ha_verify_guards(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ return true;
+}
+
+/*
+ * Guard on stopping: a sched_switch after ha_cancel_timer_sync() would
+ * re-arm the timer (ODEBUG splat). _acquire pairs with cmpxchg_release in
+ * tlob_stop_task.
+ *
+ * Entering stopped_tlob also resets env_store to the invalid sentinel, so a
+ * restart re-anchors the clock; a stale anchor would wrap the restart's
+ * timer delay to ~U64_MAX.
+ */
+static inline void ha_setup_invariants(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (next_state == stopped_tlob) {
+ /*
+ * Window ending: reset env_store to the invalid sentinel so
+ * the next window gets a fresh clock anchor. Keep stopping==1
+ * so __tlob_acc() continues to block sched events while parked.
+ */
+ ha_monitor_reset_all_stored(ha_mon);
+ return;
+ }
+
+ if (atomic_read_acquire(&ha_get_target(ha_mon)->stopping)) {
+ /*
+ * Restart (stopped -> running): arm the timer, then clear
+ * stopping so __tlob_acc() admits sched events only once the
+ * state is already running_tlob. _release pairs with the
+ * acquires in __tlob_acc/tlob_reset_notify.
+ */
+ if (next_state < state_max_tlob)
+ ha_start_timer_ns(ha_mon, clk_elapsed_tlob, BUDGET_NS(ha_mon), time_ns);
+ atomic_set_release(&ha_get_target(ha_mon)->stopping, 0);
+ return;
+ }
+
+ if (next_state < state_max_tlob)
+ ha_start_timer_ns(ha_mon, clk_elapsed_tlob, BUDGET_NS(ha_mon), time_ns);
+ else
+ ha_cancel_timer(ha_mon);
+}
+
+static bool ha_verify_constraint(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (!ha_verify_invariants(ha_mon, curr_state, event, next_state, time_ns))
+ return false;
+
+ if (!ha_verify_guards(ha_mon, curr_state, event, next_state, time_ns))
+ return false;
+
+ ha_setup_invariants(ha_mon, curr_state, event, next_state, time_ns);
+
+ return true;
+}
+
+/*
+ * Pre-allocated pool of TLOB_MAX_MONITORED slots. mempool_alloc_preallocated()
+ * pops a reserve slot without touching the allocator (bounded start latency;
+ * -ENOSPC past the cap). Slots return via destroy/cleanup; mempool_free() is
+ * safe from RCU-callback context.
+ */
+static mempool_t tlob_ws_pool;
+
+static void tlob_ws_return_cb(struct rcu_head *head)
+{
+ struct tlob_task_state *ws =
+ container_of(head, struct tlob_task_state, rcu);
+
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+/* Direct return without RCU delay (ws was never published to the hash). */
+static void tlob_ws_direct_return(struct tlob_task_state *ws)
+{
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+static struct tlob_task_state *tlob_ws_alloc(void)
+{
+ struct tlob_task_state *ws =
+ mempool_alloc_preallocated(&tlob_ws_pool);
+
+ if (!ws)
+ return NULL;
+
+ memset(ws, 0, sizeof(*ws));
+ INIT_LIST_HEAD(&ws->started_node);
+ return ws;
+}
+
+/*
+ * Uprobe binding list; protected by tlob_uprobe_mutex. When both are
+ * taken, tlob_uprobe_mutex is always acquired before tlob_ws_lock:
+ * inverting the order would be a silent lock-order inversion.
+ */
+static LIST_HEAD(tlob_uprobe_list);
+static DEFINE_MUTEX(tlob_uprobe_mutex);
+
+/* Serialises tlob_task_state ownership: restart, detach, unbind reap. */
+static DEFINE_SPINLOCK(tlob_ws_lock);
+
+/* Per-uprobe-binding state: a start + stop probe pair for one binary region. */
+struct tlob_uprobe_binding {
+ struct list_head list;
+ u64 threshold_ns;
+ char binpath[TLOB_MAX_PATH];
+ loff_t offset_start;
+ loff_t offset_stop;
+ /*
+ * All tlob_task_states this binding ever started, for each task's
+ * lifetime. Protected by tlob_ws_lock.
+ */
+ struct list_head started_list;
+ DECLARE_RV_UPROBE(start_probe);
+ DECLARE_RV_UPROBE(stop_probe);
+};
+
+/*
+ * Unlink ws from its binding's started_list before returning it to the pool.
+ * ws->binding is left stale: the next tlob_ws_alloc() memsets it, and the
+ * restart path checks destroying first. Idempotent (list_del_init no-op).
+ */
+static inline void tlob_detach_from_binding(struct tlob_task_state *ws)
+{
+ if (!ws->binding)
+ return;
+ guard(spinlock)(&tlob_ws_lock);
+ list_del_init(&ws->started_node);
+}
+
+/*
+ * Per-entry teardown during monitor disable. cmpxchg on destroying
+ * (0->1) claims ownership -- not stopping, which can be long-lived (a
+ * parked task).
+ *
+ * No timer cancel or locking needed: disable_tlob() already synced every
+ * uprobe/tracepoint, da_monitor_destroy() ran da_monitor_reset_all() +
+ * synchronize_rcu(), and ha_mon_destroying blocks new timer callbacks.
+ */
+static inline void tlob_extra_cleanup(struct da_monitor *da_mon)
+{
+ struct ha_monitor *ha_mon = to_ha_monitor(da_mon);
+ struct tlob_task_state *ws = ha_get_target(ha_mon);
+
+ if (!ws)
+ return;
+
+ if (atomic_cmpxchg_release(&ws->destroying, 0, 1) != 0)
+ return;
+
+ tlob_detach_from_binding(ws);
+ put_task_struct(ws->task);
+ /*
+ * da_monitor_destroy() has already called synchronize_rcu(); no
+ * reader holds ws. Return the slot directly without call_rcu.
+ */
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+/*
+ * Accumulate elapsed ns into accs_ns[idx] since last_ts and advance it.
+ * Returns true if monitored with an active window. The stopping gate is
+ * what keeps scheduler events from reaching a parked task (no "stopped"
+ * self-loops, see tlob.h) and keeps accs_ns[] from growing while parked.
+ */
+static inline bool __tlob_acc(struct task_struct *task, ktime_t now,
+ enum tlob_acc_idx idx)
+{
+ struct tlob_task_state *ws;
+ unsigned long flags;
+
+ guard(rcu)();
+ ws = da_get_target_by_id(task->pid);
+ /* acquire pairs with the _release clear in ha_setup_invariants(). */
+ if (!ws || atomic_read_acquire(&ws->stopping))
+ return false;
+ raw_spin_lock_irqsave(&ws->entry_lock, flags);
+ ws->accs_ns[idx] += ktime_to_ns(ktime_sub(now, ws->last_ts));
+ ws->last_ts = now;
+ raw_spin_unlock_irqrestore(&ws->entry_lock, flags);
+ return true;
+}
+
+static inline bool tlob_acc_running(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_RUNNING);
+}
+
+static inline bool tlob_acc_waiting(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_WAITING);
+}
+
+/*
+ * handle_sched_switch - advance the DA on every context switch.
+ *
+ * Emits sleep (running -> sleeping), preempt (running -> waiting) for prev,
+ * and switch_in (waiting -> running) for next. One ktime_get() shared by
+ * both acc calls keeps prev/next on the same context-switch timestamp.
+ *
+ * No waiting->sleeping edge: a task blocks (calls schedule()) only on CPU
+ * (running); waiting means TASK_RUNNING on the runqueue.
+ */
+static void handle_sched_switch(void *data, bool preempt_unused,
+ struct task_struct *prev,
+ struct task_struct *next,
+ unsigned int prev_state)
+{
+ ktime_t now = ktime_get();
+ bool prev_preempted = (prev_state == 0);
+
+ if (tlob_acc_running(prev, now))
+ da_handle_event(prev->pid, NULL,
+ prev_preempted ? preempt_tlob : sleep_tlob);
+ if (tlob_acc_waiting(next, now))
+ da_handle_event(next->pid, NULL, switch_in_tlob);
+}
+
+static inline bool tlob_acc_sleeping(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_SLEEPING);
+}
+
+/*
+ * handle_sched_wakeup - sleeping -> waiting transition. try_to_wake_up()
+ * skips TASK_RUNNING tasks, so this never fires for running/waiting.
+ */
+static void handle_sched_wakeup(void *data, struct task_struct *p)
+{
+ ktime_t now = ktime_get();
+
+ if (tlob_acc_sleeping(p, now))
+ da_handle_event(p->pid, NULL, wakeup_tlob);
+}
+
+/* Forward decl: used by handle_sched_process_exit() and tlob_unbind_reap(). */
+static int tlob_stop_task(struct task_struct *task, void *binding);
+static void tlob_destroy_task(struct task_struct *task);
+
+/*
+ * handle_sched_process_exit - clean up a task that exits without hitting its
+ * STOP uprobe (killed, unmapped mid-region, ...). The task is always in
+ * running_tlob here: do_exit() runs in the task's own context, which
+ * required passing through switch_in_tlob (running). tlob_stop_task() ends
+ * the window (or is a harmless -EAGAIN/-ESRCH); tlob_destroy_task() then
+ * frees the slot, as no restart can follow exit.
+ */
+static void handle_sched_process_exit(void *data, struct task_struct *p,
+ bool group_dead)
+{
+ tlob_stop_task(p, NULL);
+ tlob_destroy_task(p);
+}
+
+/**
+ * tlob_start_task - begin monitoring @task with budget @threshold_ns ns.
+ * @task: Task to monitor; may be current or another task.
+ * @threshold_ns: Budget in ns, in [1000, TLOB_MAX_THRESHOLD_NS].
+ * @binding: Opaque owner, recorded on fresh allocation and checked for
+ * an exact match on restart; NULL for callers that never
+ * restart a parked window.
+ *
+ * Allocates a fresh entry if @task has none, or restarts a parked entry in
+ * place when @binding matches (see tlob.dot: "start" fires from both
+ * running and stopped).
+ *
+ * Returns 0, -ENODEV, -ERANGE, -EALREADY, -ESRCH, or -ENOSPC (fresh start
+ * past pool capacity).
+ */
+static int tlob_start_task(struct task_struct *task, u64 threshold_ns, void *binding)
+{
+ struct tlob_task_state *ws;
+
+ if (!da_monitor_enabled())
+ return -ENODEV;
+
+ if (threshold_ns < TLOB_MIN_THRESHOLD_NS ||
+ threshold_ns > TLOB_MAX_THRESHOLD_NS)
+ return -ERANGE;
+
+ /* Serialise duplicate-check + pool-slot claim; see tlob_ws_lock. */
+ guard(spinlock)(&tlob_ws_lock);
+
+ /*
+ * da_get_target_by_id() uses hash_for_each_possible_rcu(), which
+ * requires an RCU read-side critical section.
+ */
+ scoped_guard(rcu) {
+ ws = da_get_target_by_id(task->pid);
+ if (ws) {
+ if (!atomic_read(&ws->stopping))
+ return -EALREADY;
+ if (atomic_read(&ws->destroying))
+ return -ESRCH;
+ /*
+ * Exact match only. An orphaned parked ws (binding
+ * cleared while active, then parked) is not adopted:
+ * that would need re-linking into the new binding's
+ * started_list. Accepted gap; the slot is reclaimed
+ * at task exit.
+ */
+ if (ws->binding != binding)
+ return -EALREADY;
+
+ /* Restart in place: same slot, hash entry, task ref, list node. */
+ ws->threshold_ns = threshold_ns;
+ WRITE_ONCE(ws->budget_exceeded, false);
+ memset(ws->accs_ns, 0, sizeof(ws->accs_ns));
+ ws->last_ts = ktime_get();
+
+ /*
+ * Keep stopping set: __tlob_acc() gates out sched
+ * events until ha_setup_invariants() clears it after
+ * the state is running_tlob. Clearing here would let
+ * events hit stopped_tlob (INVALID transitions).
+ */
+
+ /* Only failure here: monitor disabled since the check above. */
+ if (!da_handle_start_run_event(task->pid, ws, start_tlob))
+ return -ENODEV;
+ return 0;
+ }
+ }
+
+ ws = tlob_ws_alloc();
+ if (!ws)
+ return -ENOSPC;
+
+ ws->task = task;
+ get_task_struct(task);
+ ws->threshold_ns = threshold_ns;
+ ws->last_ts = ktime_get();
+ raw_spin_lock_init(&ws->entry_lock);
+ ws->binding = binding;
+ if (binding)
+ list_add_tail(&ws->started_node,
+ &((struct tlob_uprobe_binding *)binding)->started_list);
+
+ /* Dispatch failed (pool exhausted or monitor disabled): unwind the slot. */
+ if (!da_handle_start_run_event(task->pid, ws, start_tlob)) {
+ if (binding)
+ list_del_init(&ws->started_node);
+ /* stopping=1 short-circuits the reset hook; destroy before freeing ws. */
+ atomic_set(&ws->stopping, 1);
+ da_destroy_storage(task->pid);
+ put_task_struct(task);
+ tlob_ws_direct_return(ws);
+ return -ENOSPC;
+ }
+
+ return 0;
+}
+
+/**
+ * tlob_stop_task - end the current monitoring window for @task.
+ * @task: Task to stop.
+ * @binding: Opaque owner; must match ws->binding to end a normal (uprobe)
+ * window. NULL (task exit) skips the check.
+ *
+ * Ends the window (dispatches "stop") but does NOT free the entry: it stays
+ * parked so a later tlob_start_task() can restart it. Call
+ * tlob_destroy_task() once @task will never restart.
+ *
+ * cmpxchg on stopping (0->1) under RCU claims ownership; the winner cancels
+ * the timer synchronously.
+ *
+ * Returns 0, -EOVERFLOW (budget exceeded), -ESRCH (not monitored),
+ * -EAGAIN (window already ended), or -EALREADY (owned by another binding).
+ */
+static int tlob_stop_task(struct task_struct *task, void *binding)
+{
+ struct ha_monitor *ha_mon;
+ struct tlob_task_state *ws;
+ bool budget_exceeded;
+
+ scoped_guard(rcu) {
+ ha_mon = ha_get_monitor(task->pid, NULL);
+ if (!ha_mon)
+ return -ESRCH;
+
+ ws = ha_get_target(ha_mon);
+ if (WARN_ON_ONCE(!ws))
+ return -ESRCH;
+
+ /* Only the binding that opened the window may end it; NULL
+ * (task exit) skips the check. Symmetric with the restart
+ * check in tlob_start_task(). */
+ if (binding && ws->binding != binding)
+ return -EALREADY;
+
+ /* cmpxchg (0->1) claims the window under RCU; _release pairs
+ * with the acquire in ha_setup_invariants(). */
+ if (atomic_cmpxchg_release(&ws->stopping, 0, 1) != 0)
+ return -EAGAIN;
+
+ /*
+ * ws may be destroyed concurrently (unbind -> call_rcu), so
+ * keep its access under RCU; dispatch re-looks-up under RCU.
+ */
+ ha_cancel_timer_sync(ha_mon);
+ budget_exceeded = READ_ONCE(ws->budget_exceeded);
+ }
+
+ /* running -> stopped: no reset or destroy, the entry stays parked. */
+ da_handle_event(task->pid, NULL, stop_tlob);
+
+ return budget_exceeded ? -EOVERFLOW : 0;
+}
+
+/*
+ * tlob_destroy_task - final teardown for @task's entry: frees the pool slot,
+ * drops the task_struct ref, removes the hash entry, whether active or parked.
+ * Idempotent via the destroying cmpxchg (same pattern as tlob_extra_cleanup()).
+ * Callers must end the window first (see handle_sched_process_exit()).
+ */
+static void tlob_destroy_task(struct task_struct *task)
+{
+ struct ha_monitor *ha_mon;
+ struct tlob_task_state *ws;
+
+ scoped_guard(rcu) {
+ ha_mon = ha_get_monitor(task->pid, NULL);
+ if (!ha_mon)
+ return;
+ ws = ha_get_target(ha_mon);
+ if (WARN_ON_ONCE(!ws))
+ return;
+ if (atomic_cmpxchg_release(&ws->destroying, 0, 1) != 0)
+ return;
+ }
+
+ tlob_detach_from_binding(ws);
+
+ /* Force the window ended: @task may never have reached STOP or a timer. */
+ atomic_set(&ws->stopping, 1);
+ ha_cancel_timer_sync(ha_mon);
+
+ scoped_guard(rcu) {
+ da_monitor_reset(&ha_mon->da_mon);
+ }
+ da_destroy_storage(task->pid);
+
+ put_task_struct(ws->task);
+ call_rcu(&ws->rcu, tlob_ws_return_cb);
+}
+
+static int tlob_uprobe_entry_handler(struct uprobe_consumer *self,
+ struct pt_regs *regs, __u64 *data)
+{
+ struct tlob_uprobe_binding *b =
+ container_of(self, struct tlob_uprobe_binding, start_probe.uc);
+
+ tlob_start_task(current, b->threshold_ns, b);
+ return 0;
+}
+
+static int tlob_uprobe_stop_handler(struct uprobe_consumer *self,
+ struct pt_regs *regs, __u64 *data)
+{
+ struct tlob_uprobe_binding *b =
+ container_of(self, struct tlob_uprobe_binding, stop_probe.uc);
+
+ tlob_stop_task(current, b);
+ return 0;
+}
+
+/*
+ * Register start + stop entry uprobes for a binding.
+ * Called with tlob_uprobe_mutex held.
+ */
+static int tlob_add_uprobe(u64 threshold_ns, const char *binpath,
+ loff_t offset_start, loff_t offset_stop)
+{
+ struct tlob_uprobe_binding *tmp_b;
+ char pathbuf[TLOB_MAX_PATH];
+ struct inode *inode;
+ struct path path __free(path_put) = {};
+ char *canon;
+ int ret;
+
+ if (binpath[0] != '/')
+ return -EINVAL;
+
+ struct tlob_uprobe_binding *b __free(kfree) = kzalloc_obj(*b, GFP_KERNEL);
+ if (!b)
+ return -ENOMEM;
+
+ b->threshold_ns = threshold_ns;
+ b->offset_start = offset_start;
+ b->offset_stop = offset_stop;
+ INIT_LIST_HEAD(&b->started_list);
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &path);
+ if (ret)
+ return ret;
+
+ if (!d_is_reg(path.dentry))
+ return -EINVAL;
+
+ inode = d_real_inode(path.dentry);
+
+ /* Reject duplicate start offset for the same binary inode. */
+ list_for_each_entry(tmp_b, &tlob_uprobe_list, list) {
+ if (tmp_b->offset_start == offset_start &&
+ rv_uprobe_is_registered(&tmp_b->start_probe) &&
+ d_real_inode(tmp_b->start_probe.path.dentry) == inode)
+ return -EEXIST;
+ }
+
+ canon = d_path(&path, pathbuf, sizeof(pathbuf));
+ if (IS_ERR(canon))
+ return PTR_ERR(canon);
+ strscpy(b->binpath, canon, sizeof(b->binpath));
+
+ b->start_probe.uc.handler = tlob_uprobe_entry_handler;
+ ret = rv_uprobe_register(b->binpath, offset_start, &b->start_probe);
+ if (ret)
+ return ret;
+
+ b->stop_probe.uc.handler = tlob_uprobe_stop_handler;
+ ret = rv_uprobe_register(b->binpath, offset_stop, &b->stop_probe);
+ if (ret) {
+ rv_uprobe_unregister(&b->start_probe);
+ return ret;
+ }
+
+ /* NOT "b = no_free_ptr(b)": the re-assignment would free the live node. */
+ list_add_tail(&no_free_ptr(b)->list, &tlob_uprobe_list);
+ return 0;
+}
+
+/*
+ * tlob_unbind_reap - detach every task @b started, destroy the parked ones.
+ *
+ * Caller must have unregistered @b's uprobes and called rv_uprobe_sync():
+ * no start/stop can then be in flight for @b, so started_list is safe to
+ * walk. Active tasks are detached (binding cleared) and left running,
+ * matching unbind behaviour today; parked tasks are destroyed, or their
+ * pool slot leaks until the task next exits.
+ */
+static void tlob_unbind_reap(struct tlob_uprobe_binding *b)
+{
+ struct tlob_task_state *ws, *tmp;
+ LIST_HEAD(to_destroy);
+
+ scoped_guard(spinlock, &tlob_ws_lock) {
+ list_for_each_entry_safe(ws, tmp, &b->started_list, started_node) {
+ list_del_init(&ws->started_node);
+ ws->binding = NULL;
+ if (atomic_read(&ws->stopping))
+ list_add_tail(&ws->started_node, &to_destroy);
+ }
+ }
+
+ list_for_each_entry_safe(ws, tmp, &to_destroy, started_node) {
+ list_del_init(&ws->started_node);
+ tlob_destroy_task(ws->task);
+ }
+}
+
+static int tlob_remove_uprobe_by_key(loff_t offset_start, const char *binpath)
+{
+ struct tlob_uprobe_binding *b, *tmp;
+ struct path remove_path;
+ struct inode *inode;
+ int ret;
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &remove_path);
+ if (ret)
+ return ret;
+
+ inode = d_real_inode(remove_path.dentry);
+
+ ret = -ENOENT;
+ list_for_each_entry_safe(b, tmp, &tlob_uprobe_list, list) {
+ if (b->offset_start != offset_start)
+ continue;
+ if (d_real_inode(b->start_probe.path.dentry) != inode)
+ continue;
+ list_del(&b->list);
+ /*
+ * rv_uprobe_sync() may sleep; list_del() already made the
+ * binding invisible to new readers.
+ */
+ rv_uprobe_unregister_nosync(&b->start_probe);
+ rv_uprobe_unregister_nosync(&b->stop_probe);
+ rv_uprobe_sync();
+ tlob_unbind_reap(b);
+ path_put(&b->start_probe.path);
+ path_put(&b->stop_probe.path);
+ kfree(b);
+ ret = 0;
+ break;
+ }
+
+ path_put(&remove_path);
+ return ret;
+}
+
+static void tlob_remove_all_uprobes(void)
+{
+ struct tlob_uprobe_binding *b, *tmp;
+ LIST_HEAD(pending);
+
+ mutex_lock(&tlob_uprobe_mutex);
+ list_for_each_entry_safe(b, tmp, &tlob_uprobe_list, list) {
+ list_move(&b->list, &pending);
+ rv_uprobe_unregister_nosync(&b->start_probe);
+ rv_uprobe_unregister_nosync(&b->stop_probe);
+ }
+ mutex_unlock(&tlob_uprobe_mutex);
+
+ if (list_empty(&pending))
+ return;
+
+ /* One sync covers all dequeued probes: consumers are then safe to free. */
+ rv_uprobe_sync();
+
+ list_for_each_entry_safe(b, tmp, &pending, list) {
+ list_del(&b->list);
+ tlob_unbind_reap(b);
+ path_put(&b->start_probe.path);
+ path_put(&b->stop_probe.path);
+ kfree(b);
+ }
+}
+
+static ssize_t tlob_monitor_read(struct file *file,
+ char __user *ubuf,
+ size_t count, loff_t *ppos)
+{
+ const int line_sz = TLOB_MAX_PATH + 128;
+ struct tlob_uprobe_binding *b;
+ char *buf;
+ int n = 0, buf_sz, pos = 0;
+ ssize_t ret;
+
+ mutex_lock(&tlob_uprobe_mutex);
+ list_for_each_entry(b, &tlob_uprobe_list, list)
+ n++;
+
+ buf_sz = (n ? n : 1) * line_sz + 1;
+ buf = kmalloc(buf_sz, GFP_KERNEL);
+ if (!buf) {
+ mutex_unlock(&tlob_uprobe_mutex);
+ return -ENOMEM;
+ }
+
+ list_for_each_entry(b, &tlob_uprobe_list, list) {
+ pos += scnprintf(buf + pos, buf_sz - pos,
+ "p %s:0x%llx 0x%llx threshold=%llu\n",
+ b->binpath,
+ (unsigned long long)b->offset_start,
+ (unsigned long long)b->offset_stop,
+ b->threshold_ns);
+ }
+ mutex_unlock(&tlob_uprobe_mutex);
+
+ ret = simple_read_from_buffer(ubuf, count, ppos, buf, pos);
+ kfree(buf);
+ return ret;
+}
+
+/*
+ * Parse "p PATH:OFFSET_START OFFSET_STOP threshold=NS".
+ * PATH may contain ':'; the last ':' separates path from offset.
+ * Returns 0, -EINVAL, or -ERANGE.
+ */
+VISIBLE_IF_KUNIT int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
+ char **path_out,
+ loff_t *start_out, loff_t *stop_out)
+{
+ unsigned long long thr = 0, stop_val = 0;
+ long long start_val;
+ char *p, *path_token, *token, *colon;
+ bool got_stop = false, got_thr = false;
+ int n;
+
+ /* Must start with "p " */
+ if (buf[0] != 'p' || buf[1] != ' ')
+ return -EINVAL;
+
+ p = buf + 2;
+ while (*p == ' ')
+ p++;
+
+ /* First space-delimited token is PATH:OFFSET_START */
+ path_token = strsep(&p, " \t");
+ if (!path_token || !*path_token)
+ return -EINVAL;
+
+ /* Split at last ':' to handle paths that contain ':'. */
+ colon = strrchr(path_token, ':');
+ if (!colon || colon - path_token < 2)
+ return -EINVAL;
+ *colon = '\0';
+
+ if (path_token[0] != '/')
+ return -EINVAL;
+
+ n = 0;
+ if (sscanf(colon + 1, "%lli%n", &start_val, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (start_val < 0)
+ return -EINVAL;
+
+ /* Remaining tokens: OFFSET_STOP threshold=NS */
+ while (p && (token = strsep(&p, " \t")) != NULL) {
+ if (!*token)
+ continue;
+ if (strncmp(token, "threshold=", 10) == 0) {
+ if (kstrtoull(token + 10, 0, &thr))
+ return -EINVAL;
+ if (thr < TLOB_MIN_THRESHOLD_NS || thr > TLOB_MAX_THRESHOLD_NS)
+ return -ERANGE;
+ got_thr = true;
+ } else if (!got_stop) {
+ long long sv;
+
+ n = 0;
+ if (sscanf(token, "%lli%n", &sv, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (sv < 0)
+ return -EINVAL;
+ stop_val = (unsigned long long)sv;
+ got_stop = true;
+ } else {
+ return -EINVAL;
+ }
+ }
+
+ if (!got_stop || !got_thr)
+ return -EINVAL;
+ if (start_val == (long long)stop_val)
+ return -EINVAL;
+
+ *thr_out = thr;
+ *path_out = path_token;
+ *start_out = (loff_t)start_val;
+ *stop_out = (loff_t)stop_val;
+ return 0;
+}
+EXPORT_SYMBOL_IF_KUNIT(tlob_parse_uprobe_line);
+
+/*
+ * Parse "-PATH:OFFSET_START" (ftrace uprobe_events removal convention).
+ */
+VISIBLE_IF_KUNIT int tlob_parse_remove_line(char *buf, char **path_out,
+ loff_t *start_out)
+{
+ char *binpath, *colon;
+ long long off;
+ int n = 0;
+
+ if (buf[0] != '-')
+ return -EINVAL;
+ binpath = buf + 1;
+ if (binpath[0] != '/')
+ return -EINVAL;
+ colon = strrchr(binpath, ':');
+ if (!colon || colon - binpath < 2)
+ return -EINVAL;
+ *colon = '\0';
+ if (sscanf(colon + 1, "%lli%n", &off, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (off < 0)
+ return -EINVAL;
+ *path_out = binpath;
+ *start_out = (loff_t)off;
+ return 0;
+}
+EXPORT_SYMBOL_IF_KUNIT(tlob_parse_remove_line);
+
+static int tlob_create_or_delete_uprobe(char *buf)
+{
+ loff_t offset_start, offset_stop;
+ u64 threshold_ns;
+ char *binpath;
+ int ret;
+
+ if (buf[0] == '-') {
+ ret = tlob_parse_remove_line(buf, &binpath, &offset_start);
+ if (ret)
+ return ret;
+ mutex_lock(&tlob_uprobe_mutex);
+ ret = tlob_remove_uprobe_by_key(offset_start, binpath);
+ mutex_unlock(&tlob_uprobe_mutex);
+ return ret;
+ }
+ ret = tlob_parse_uprobe_line(buf, &threshold_ns, &binpath,
+ &offset_start, &offset_stop);
+ if (ret)
+ return ret;
+ mutex_lock(&tlob_uprobe_mutex);
+ ret = tlob_add_uprobe(threshold_ns, binpath, offset_start, offset_stop);
+ mutex_unlock(&tlob_uprobe_mutex);
+ return ret;
+}
+
+static ssize_t tlob_monitor_write(struct file *file,
+ const char __user *ubuf,
+ size_t count, loff_t *ppos)
+{
+ char buf[TLOB_MAX_PATH + 128];
+
+ if (count >= sizeof(buf))
+ return -EINVAL;
+ if (copy_from_user(buf, ubuf, count))
+ return -EFAULT;
+ buf[count] = '\0';
+ if (count > 0 && buf[count - 1] == '\n')
+ buf[count - 1] = '\0';
+ return tlob_create_or_delete_uprobe(buf) ?: (ssize_t)count;
+}
+
+static const struct file_operations tlob_monitor_fops = {
+ .open = simple_open,
+ .read = tlob_monitor_read,
+ .write = tlob_monitor_write,
+ .llseek = noop_llseek,
+};
+
+static int __tlob_init_monitor(void)
+{
+ int retval;
+
+ retval = mempool_init_kmalloc_pool(&tlob_ws_pool, TLOB_MAX_MONITORED,
+ sizeof(struct tlob_task_state));
+ if (retval)
+ return retval;
+
+ retval = ha_monitor_init();
+ if (retval) {
+ mempool_exit(&tlob_ws_pool);
+ return retval;
+ }
+
+ rv_this.enabled = 1;
+ return 0;
+}
+
+static void __tlob_destroy_monitor(void)
+{
+ rv_this.enabled = 0;
+ tlob_remove_all_uprobes();
+ /*
+ * A grace period only makes the call_rcu()'d tlob_ws_return_cb()
+ * callbacks eligible to run; rcu_barrier() waits until they have all
+ * returned their slots before the pool is destroyed.
+ */
+ ha_monitor_destroy();
+ rcu_barrier();
+ mempool_exit(&tlob_ws_pool);
+}
+
+static int tlob_enable_hooks(void)
+{
+ rv_attach_trace_probe("tlob", sched_switch, handle_sched_switch);
+ rv_attach_trace_probe("tlob", sched_wakeup, handle_sched_wakeup);
+ rv_attach_trace_probe("tlob", sched_process_exit, handle_sched_process_exit);
+ return 0;
+}
+
+static void tlob_disable_hooks(void)
+{
+ rv_detach_trace_probe("tlob", sched_switch, handle_sched_switch);
+ rv_detach_trace_probe("tlob", sched_wakeup, handle_sched_wakeup);
+ rv_detach_trace_probe("tlob", sched_process_exit, handle_sched_process_exit);
+}
+
+static int enable_tlob(void)
+{
+ int retval;
+
+ retval = __tlob_init_monitor();
+ if (retval)
+ return retval;
+
+ return tlob_enable_hooks();
+}
+
+static void disable_tlob(void)
+{
+ tlob_disable_hooks();
+ __tlob_destroy_monitor();
+}
+
+static struct rv_monitor rv_this = {
+ .name = "tlob",
+ .description = "Per-task latency-over-budget monitor.",
+ .enable = enable_tlob,
+ .disable = disable_tlob,
+ .reset = da_monitor_reset_all,
+ .enabled = 0,
+};
+
+static int __init register_tlob(void)
+{
+ int ret;
+
+ ret = rv_register_monitor(&rv_this, NULL);
+ if (ret)
+ return ret;
+
+ if (rv_this.root_d) {
+ if (!rv_create_file("monitor", RV_MODE_WRITE, rv_this.root_d, NULL,
+ &tlob_monitor_fops)) {
+ rv_unregister_monitor(&rv_this);
+ return -ENOMEM;
+ }
+ }
+
+ return 0;
+}
+
+static void __exit unregister_tlob(void)
+{
+ rv_unregister_monitor(&rv_this);
+}
+
+module_init(register_tlob);
+module_exit(unregister_tlob);
+
+MODULE_LICENSE("GPL");
+MODULE_AUTHOR("Wen Yang <wen.yang@linux.dev>");
+MODULE_DESCRIPTION("tlob: task latency over budget per-task monitor.");
diff --git a/kernel/trace/rv/monitors/tlob/tlob.h b/kernel/trace/rv/monitors/tlob/tlob.h
new file mode 100644
index 000000000000..0a80a66d94bc
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob.h
@@ -0,0 +1,149 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _RV_TLOB_H
+#define _RV_TLOB_H
+
+/*
+ * C representation of the tlob hybrid automaton (see tlob.dot).
+ *
+ * States: stopped (initial; parked), running (on CPU), waiting (runqueue),
+ * sleeping (blocked). Events: start/stop (tlob_start_task/tlob_stop_task),
+ * sleep/preempt/wakeup/switch_in (sched tracepoints).
+ *
+ * "stop" fires only from running (both callers run on CPU); "stopped"
+ * leaves only via "start" (fresh start or in-place restart). running[start]
+ * is INVALID: a stray re-start must not silently reset the budget clock.
+ *
+ * Invariant: clk_elapsed < BUDGET_NS() in running/waiting/sleeping; stopped
+ * parks the window, no clock while parked. start re-inits the monitor
+ * (da_handle_start_run_event()); stop dispatches after ha_cancel_timer_sync();
+ * final teardown uses ha_cancel_timer_sync() + da_monitor_reset() +
+ * da_destroy_storage().
+ *
+ * Format: Documentation/trace/rv/deterministic_automata.rst
+ */
+
+#include <linux/rv.h>
+#include <linux/sched.h>
+
+#define MONITOR_NAME tlob
+
+enum states_tlob {
+ stopped_tlob,
+ running_tlob,
+ sleeping_tlob,
+ waiting_tlob,
+ state_max_tlob,
+};
+
+#define INVALID_STATE state_max_tlob
+
+enum events_tlob {
+ preempt_tlob,
+ sleep_tlob,
+ start_tlob,
+ stop_tlob,
+ switch_in_tlob,
+ wakeup_tlob,
+ event_max_tlob,
+};
+
+/*
+ * HA clock env: clk_elapsed, wall-clock since the window start; anchored in
+ * running/waiting/sleeping, cleared on stop.
+ */
+enum envs_tlob {
+ clk_elapsed_tlob,
+ env_max_tlob,
+ env_max_stored_tlob = env_max_tlob,
+};
+
+_Static_assert(env_max_stored_tlob <= MAX_HA_ENV_LEN, "Not enough slots");
+#define HA_CLK_NS
+
+struct automaton_tlob {
+ char *state_names[state_max_tlob];
+ char *event_names[event_max_tlob];
+ char *env_names[env_max_tlob];
+ unsigned char function[state_max_tlob][event_max_tlob];
+ unsigned char initial_state;
+ bool final_states[state_max_tlob];
+};
+
+static const struct automaton_tlob automaton_tlob = {
+ .state_names = {
+ "stopped",
+ "running",
+ "sleeping",
+ "waiting",
+ },
+ .event_names = {
+ "preempt",
+ "sleep",
+ "start",
+ "stop",
+ "switch_in",
+ "wakeup",
+ },
+ .env_names = {
+ "clk_elapsed",
+ },
+ .function = {
+ /* stopped (initial; window parked, sched events not routed) */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (not on CPU) */
+ running_tlob, /* start (tlob_start_task, fresh or restart) */
+ INVALID_STATE, /* stop (already stopped) */
+ INVALID_STATE, /* switch_in (not on CPU) */
+ INVALID_STATE, /* wakeup (not on CPU) */
+ },
+ /* running */
+ {
+ waiting_tlob, /* preempt (sched_switch, prev_state == 0) */
+ sleeping_tlob, /* sleep (sched_switch, prev_state != 0) */
+ INVALID_STATE, /* start (running task's START is -EALREADY) */
+ stopped_tlob, /* stop (tlob_stop_task) */
+ INVALID_STATE, /* switch_in (already on CPU) */
+ INVALID_STATE, /* wakeup (TASK_RUNNING can't be woken) */
+ },
+ /* sleeping */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (already sleeping) */
+ INVALID_STATE, /* start (not in running state) */
+ INVALID_STATE, /* stop (not in running state) */
+ INVALID_STATE, /* switch_in (must go through waiting first) */
+ waiting_tlob, /* wakeup */
+ },
+ /* waiting */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (not on CPU) */
+ INVALID_STATE, /* start (not in running state) */
+ INVALID_STATE, /* stop (not in running state) */
+ running_tlob, /* switch_in */
+ INVALID_STATE, /* wakeup (already TASK_RUNNING) */
+ },
+ },
+ .initial_state = stopped_tlob,
+ .final_states = { 1, 0, 0, 0 },
+};
+
+/*
+ * Hard cap on concurrently monitored tasks. tlob_ws_pool pre-allocates
+ * this many slots; a fresh start past the cap returns -ENOSPC with bounded
+ * latency (mempool_alloc_preallocated() never touches the allocator).
+ * Restarts reuse the same slot.
+ */
+#define TLOB_MAX_MONITORED 64U
+
+/* Maximum binary path length for uprobe binding. */
+#define TLOB_MAX_PATH 256
+
+/* Minimum monitoring budget (1 us). */
+#define TLOB_MIN_THRESHOLD_NS 1000ULL
+
+/* Upper budget bound (1 hour): keeps the u64 ns accumulators far from overflow. */
+#define TLOB_MAX_THRESHOLD_NS 3600000000000ULL
+
+#endif /* _RV_TLOB_H */
diff --git a/kernel/trace/rv/monitors/tlob/tlob_trace.h b/kernel/trace/rv/monitors/tlob/tlob_trace.h
new file mode 100644
index 000000000000..b3a7cf4ad3ea
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob_trace.h
@@ -0,0 +1,48 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+
+/*
+ * Snippet to be included in rv_trace.h
+ */
+
+#ifdef CONFIG_RV_MON_TLOB
+DEFINE_EVENT(event_da_monitor_id, event_tlob,
+ TP_PROTO(int id, char *state, char *event,
+ char *next_state, bool final_state),
+ TP_ARGS(id, state, event, next_state, final_state));
+
+DEFINE_EVENT(error_da_monitor_id, error_tlob,
+ TP_PROTO(int id, char *state, char *event),
+ TP_ARGS(id, state, event));
+
+DEFINE_EVENT(error_env_da_monitor_id, error_env_tlob,
+ TP_PROTO(int id, char *state, char *event, char *env),
+ TP_ARGS(id, state, event, env));
+
+/*
+ * detail_env_tlob - per-state latency breakdown on budget violation.
+ * Emitted right after error_env_tlob from the hrtimer callback.
+ */
+TRACE_EVENT(detail_env_tlob,
+ TP_PROTO(int id, u64 threshold_ns,
+ u64 running_ns, u64 waiting_ns, u64 sleeping_ns),
+ TP_ARGS(id, threshold_ns, running_ns, waiting_ns, sleeping_ns),
+ TP_STRUCT__entry(
+ __field(int, id)
+ __field(u64, threshold_ns)
+ __field(u64, running_ns)
+ __field(u64, waiting_ns)
+ __field(u64, sleeping_ns)
+ ),
+ TP_fast_assign(
+ __entry->id = id;
+ __entry->threshold_ns = threshold_ns;
+ __entry->running_ns = running_ns;
+ __entry->waiting_ns = waiting_ns;
+ __entry->sleeping_ns = sleeping_ns;
+ ),
+ TP_printk("pid=%d threshold_ns=%llu"
+ " running_ns=%llu waiting_ns=%llu sleeping_ns=%llu",
+ __entry->id, __entry->threshold_ns,
+ __entry->running_ns, __entry->waiting_ns, __entry->sleeping_ns)
+);
+#endif /* CONFIG_RV_MON_TLOB */
diff --git a/kernel/trace/rv/rv_trace.h b/kernel/trace/rv/rv_trace.h
index 2f8a932432c9..4bfa39717cef 100644
--- a/kernel/trace/rv/rv_trace.h
+++ b/kernel/trace/rv/rv_trace.h
@@ -189,6 +189,7 @@ DECLARE_EVENT_CLASS(error_env_da_monitor_id,
#include <monitors/stall/stall_trace.h>
#include <monitors/nomiss/nomiss_trace.h>
+#include <monitors/tlob/tlob_trace.h>
// Add new monitors based on CONFIG_HA_MON_EVENTS_ID here
#endif
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 7/9] rv: Add KUnit tests for the tlob monitor
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (5 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:45 ` [PATCH v6 8/9] selftests/verification: Add tlob selftests wen.yang
2026-08-20 16:45 ` [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
8 siblings, 0 replies; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add a test case to the shared rv_monitors_test.c suite (gated by the
existing CONFIG_RV_MONITORS_KUNIT_TEST, same as nomiss/sco/etc.)
covering the uprobe-line parser.
tlob_parse_uprobe_line() and tlob_parse_remove_line() stay static and
unconditional: production code (tlob_create_or_delete_uprobe()) always
needs them, so they can't be compiled out, and un-hiding them via
VISIBLE_IF_KUNIT would tie their linkage to CONFIG_KUNIT while nothing
else in tlob.c depends on it. Instead, expose them to the test only
through a const rv_tlob_kunit_ops struct of function pointers, built
and exported the same way nomiss.c/sco.c expose their rv_<mon>_ops:
the struct itself, its declaration in tlob_kunit.h, and its use in
tlob_kunit.c are all gated on the single CONFIG_RV_MONITORS_KUNIT_TEST
symbol, so there is no separate prototype to go out of sync with a
visibility macro.
tlob_kunit.c follows the monitors/*/*_kunit.c convention: one
rv_test_tlob() case textually included into rv_monitors_test.c,
guarded by IS_REACHABLE(CONFIG_RV_MON_TLOB) with an rv_test_stub()
fallback so the shared suite still builds when RV_MON_TLOB=n (that
symbol is independent of CONFIG_RV_MONITORS_KUNIT_TEST). Drop the
per-monitor TLOB_KUNIT_TEST Kconfig entry, .kunitconfig, and Makefile
line that a standalone test module would have needed.
Cases cover valid inputs, malformed paths and offsets (including
negative values), out-of-range thresholds, and valid and invalid
remove lines.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
kernel/trace/rv/monitors/tlob/tlob.c | 24 ++++--
kernel/trace/rv/monitors/tlob/tlob_kunit.c | 87 ++++++++++++++++++++++
kernel/trace/rv/monitors/tlob/tlob_kunit.h | 17 +++++
kernel/trace/rv/rv_monitors_test.c | 2 +
4 files changed, 122 insertions(+), 8 deletions(-)
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.h
diff --git a/kernel/trace/rv/monitors/tlob/tlob.c b/kernel/trace/rv/monitors/tlob/tlob.c
index 08b1bee884cc..18150cbf57a5 100644
--- a/kernel/trace/rv/monitors/tlob/tlob.c
+++ b/kernel/trace/rv/monitors/tlob/tlob.c
@@ -20,7 +20,6 @@
#include <linux/namei.h>
#include <linux/rv.h>
#include <linux/slab.h>
-#include <kunit/visibility.h>
#include <rv/instrumentation.h>
#include <rv/rv_uprobe.h>
#include <rv.h>
@@ -877,9 +876,9 @@ static ssize_t tlob_monitor_read(struct file *file,
* PATH may contain ':'; the last ':' separates path from offset.
* Returns 0, -EINVAL, or -ERANGE.
*/
-VISIBLE_IF_KUNIT int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
- char **path_out,
- loff_t *start_out, loff_t *stop_out)
+static int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
+ char **path_out,
+ loff_t *start_out, loff_t *stop_out)
{
unsigned long long thr = 0, stop_val = 0;
long long start_val;
@@ -951,13 +950,12 @@ VISIBLE_IF_KUNIT int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
*stop_out = (loff_t)stop_val;
return 0;
}
-EXPORT_SYMBOL_IF_KUNIT(tlob_parse_uprobe_line);
/*
* Parse "-PATH:OFFSET_START" (ftrace uprobe_events removal convention).
*/
-VISIBLE_IF_KUNIT int tlob_parse_remove_line(char *buf, char **path_out,
- loff_t *start_out)
+static int tlob_parse_remove_line(char *buf, char **path_out,
+ loff_t *start_out)
{
char *binpath, *colon;
long long off;
@@ -980,7 +978,17 @@ VISIBLE_IF_KUNIT int tlob_parse_remove_line(char *buf, char **path_out,
*start_out = (loff_t)off;
return 0;
}
-EXPORT_SYMBOL_IF_KUNIT(tlob_parse_remove_line);
+
+#if IS_ENABLED(CONFIG_RV_MONITORS_KUNIT_TEST)
+#include <kunit/visibility.h>
+#include "tlob_kunit.h"
+
+const struct rv_tlob_kunit_ops rv_tlob_kunit_ops = {
+ .parse_uprobe_line = tlob_parse_uprobe_line,
+ .parse_remove_line = tlob_parse_remove_line,
+};
+EXPORT_SYMBOL_IF_KUNIT(rv_tlob_kunit_ops);
+#endif
static int tlob_create_or_delete_uprobe(char *buf)
{
diff --git a/kernel/trace/rv/monitors/tlob/tlob_kunit.c b/kernel/trace/rv/monitors/tlob/tlob_kunit.c
new file mode 100644
index 000000000000..a8987fba2533
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob_kunit.c
@@ -0,0 +1,87 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <linux/kernel.h>
+#include <linux/string.h>
+#include "tlob_kunit.h"
+
+#if IS_REACHABLE(CONFIG_RV_MON_TLOB)
+
+/* Valid "p PATH:START STOP threshold=NS" lines. */
+static const char * const tlob_parse_valid[] = {
+ "p /usr/bin/myapp:4768 4848 threshold=5000000",
+ "p /usr/bin/myapp:0x12a0 0x12f0 threshold=10000000",
+ "p /opt/my:app/bin:0x100 0x200 threshold=1000000",
+};
+
+/* Malformed "p ..." lines that must be rejected with -EINVAL. */
+static const char * const tlob_parse_invalid[] = {
+ "p :0x100 0x200 threshold=5000",
+ "p /usr/bin/myapp:0x100 threshold=5000",
+ "p /usr/bin/myapp:-1 0x200 threshold=5000",
+ "p /usr/bin/myapp:0x100 -1 threshold=5000000", /* negative stop offset */
+ "p /usr/bin/myapp:0x100 0x200",
+ "p /usr/bin/myapp:0x100 0x100 threshold=5000",
+};
+
+/* threshold_ns out of valid range => -ERANGE. */
+static const char * const tlob_parse_out_of_range[] = {
+ "p /usr/bin/myapp:0x100 0x200 threshold=0",
+ "p /usr/bin/myapp:0x100 0x200 threshold=999",
+ "p /usr/bin/myapp:0x100 0x200 threshold=3600000000001",
+};
+
+/* Valid "-PATH:OFFSET_START" remove lines. */
+static const char * const tlob_remove_valid[] = {
+ "-/usr/bin/myapp:0x100",
+ "-/opt/my:app/bin:0x200",
+};
+
+/* Malformed remove lines that must be rejected with -EINVAL. */
+static const char * const tlob_remove_invalid[] = {
+ "-usr/bin/myapp:0x100",
+ "-/usr/bin/myapp",
+ "-/:0x100",
+ "-/usr/bin/myapp:-1", /* negative offset */
+ "-/usr/bin/myapp:abc",
+};
+
+static void rv_test_tlob(struct kunit *test)
+{
+ u64 thr;
+ char *path;
+ loff_t start, stop;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_valid); i++) {
+ strscpy(buf, tlob_parse_valid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, rv_tlob_kunit_ops.parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), 0);
+ }
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_invalid); i++) {
+ strscpy(buf, tlob_parse_invalid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, rv_tlob_kunit_ops.parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), -EINVAL);
+ }
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_out_of_range); i++) {
+ strscpy(buf, tlob_parse_out_of_range[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, rv_tlob_kunit_ops.parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), -ERANGE);
+ }
+
+ for (i = 0; i < ARRAY_SIZE(tlob_remove_valid); i++) {
+ strscpy(buf, tlob_remove_valid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, rv_tlob_kunit_ops.parse_remove_line(buf, &path, &start), 0);
+ }
+
+ for (i = 0; i < ARRAY_SIZE(tlob_remove_invalid); i++) {
+ strscpy(buf, tlob_remove_invalid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, rv_tlob_kunit_ops.parse_remove_line(buf, &path, &start),
+ -EINVAL);
+ }
+}
+
+#else
+#define rv_test_tlob rv_test_stub
+#endif
diff --git a/kernel/trace/rv/monitors/tlob/tlob_kunit.h b/kernel/trace/rv/monitors/tlob/tlob_kunit.h
new file mode 100644
index 000000000000..4c1081871ea3
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob_kunit.h
@@ -0,0 +1,17 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef __TLOB_KUNIT_H
+#define __TLOB_KUNIT_H
+
+#if IS_ENABLED(CONFIG_RV_MONITORS_KUNIT_TEST)
+
+#include <linux/types.h>
+
+extern const struct rv_tlob_kunit_ops {
+ int (*parse_uprobe_line)(char *buf, u64 *thr_out, char **path_out,
+ loff_t *start_out, loff_t *stop_out);
+ int (*parse_remove_line)(char *buf, char **path_out, loff_t *start_out);
+} rv_tlob_kunit_ops;
+
+#endif
+
+#endif /* __TLOB_KUNIT_H */
diff --git a/kernel/trace/rv/rv_monitors_test.c b/kernel/trace/rv/rv_monitors_test.c
index 3ad11195e664..791df0fe03e3 100644
--- a/kernel/trace/rv/rv_monitors_test.c
+++ b/kernel/trace/rv/rv_monitors_test.c
@@ -153,6 +153,7 @@ static void rv_test_dummy(struct kunit *test)
#include "monitors/nomiss/nomiss_kunit.c"
#include "monitors/pagefault/pagefault_kunit.c"
#include "monitors/sleep/sleep_kunit.c"
+#include "monitors/tlob/tlob_kunit.c"
static struct kunit_case rv_mon_test_cases[] = {
KUNIT_CASE(rv_test_dummy),
@@ -163,6 +164,7 @@ static struct kunit_case rv_mon_test_cases[] = {
KUNIT_CASE(rv_test_nomiss),
KUNIT_CASE(rv_test_pagefault),
KUNIT_CASE(rv_test_sleep),
+ KUNIT_CASE(rv_test_tlob),
{}
};
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 8/9] selftests/verification: Add tlob selftests
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (6 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:56 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add seven ftrace-style test scripts for the tlob RV monitor under
tools/testing/selftests/verification/test.d/tlob/. The tests cover
uprobe binding management, budget violation detection, and per-state
time accounting.
run_tlob_tests.sh is a thin wrapper: it builds the helpers and delegates
all argument handling to ftracetest via exec "$FTRACETEST" -K --rv "$@".
All .tc files that start background processes set up a trap teardown EXIT
immediately after launch so background tasks are cleaned up even when a
subsequent assertion fails under set -e.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
.../testing/selftests/verification/.gitignore | 2 +
tools/testing/selftests/verification/Makefile | 4 +
.../test.d/tlob/run_tlob_tests.sh | 23 ++
.../verification/test.d/tlob/uprobe_bind.tc | 39 +++
.../test.d/tlob/uprobe_detail_running.tc | 53 +++++
.../test.d/tlob/uprobe_detail_sleeping.tc | 52 ++++
.../test.d/tlob/uprobe_detail_waiting.tc | 76 ++++++
.../verification/test.d/tlob/uprobe_multi.tc | 63 +++++
.../test.d/tlob/uprobe_no_event.tc | 17 ++
.../test.d/tlob/uprobe_restart.tc | 79 ++++++
.../test.d/tlob/uprobe_violation.tc | 69 ++++++
.../testing/selftests/verification/tlob_sym.c | 225 ++++++++++++++++++
.../selftests/verification/tlob_target.c | 138 +++++++++++
13 files changed, 840 insertions(+)
create mode 100755 tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
create mode 100644 tools/testing/selftests/verification/tlob_sym.c
create mode 100644 tools/testing/selftests/verification/tlob_target.c
diff --git a/tools/testing/selftests/verification/.gitignore b/tools/testing/selftests/verification/.gitignore
index 2659417cb2c7..d2f231f1bacb 100644
--- a/tools/testing/selftests/verification/.gitignore
+++ b/tools/testing/selftests/verification/.gitignore
@@ -1,2 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
logs
+tlob_sym
+tlob_target
diff --git a/tools/testing/selftests/verification/Makefile b/tools/testing/selftests/verification/Makefile
index aa8790c22a71..41445d15b86a 100644
--- a/tools/testing/selftests/verification/Makefile
+++ b/tools/testing/selftests/verification/Makefile
@@ -5,4 +5,8 @@ TEST_PROGS := verificationtest-ktap
TEST_FILES := test.d settings
EXTRA_CLEAN := $(OUTPUT)/logs/*
+TEST_GEN_FILES := tlob_sym tlob_target
+
include ../lib.mk
+
+export VERIFICATIONTEST_BINDIR := $(OUTPUT)
diff --git a/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh b/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
new file mode 100755
index 000000000000..d8f96c4ecc2c
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
@@ -0,0 +1,23 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+#
+# Standalone runner for tlob selftests
+
+SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
+BINDIR="$(cd "$SCRIPT_DIR/../.." && pwd)"
+FTRACETEST="$(cd "$SCRIPT_DIR/../../../ftrace" && pwd)/ftracetest"
+
+# Rebuild helpers when the Makefile is present (source-tree run).
+# Skip silently in installed kselftest environments where source is absent.
+if [ -f "$BINDIR/Makefile" ]; then
+ make -C "$BINDIR" tlob_target tlob_sym
+fi
+
+if [ ! -x "$BINDIR/tlob_target" ] || [ ! -x "$BINDIR/tlob_sym" ]; then
+ echo "ERROR: tlob_target or tlob_sym not found in $BINDIR" >&2
+ exit 1
+fi
+
+export VERIFICATIONTEST_BINDIR="$BINDIR"
+
+exec "$FTRACETEST" -K --rv "$SCRIPT_DIR" "$@"
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
new file mode 100644
index 000000000000..950ef0497542
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
@@ -0,0 +1,39 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor uprobe binding (visible in monitor file, removable, duplicate rejected)
+# requires: tlob:monitor
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > monitors/tlob/enable
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=5000000000" > "$TLOB_MONITOR"
+
+# Binding must appear in monitor file with canonical hex-offset format.
+grep -qE "^p ${UPROBE_TARGET}:0x[0-9a-f]+ 0x[0-9a-f]+ threshold=[0-9]+$" "$TLOB_MONITOR"
+grep -q "threshold=5000000000" "$TLOB_MONITOR"
+
+# Duplicate offset_start must be rejected.
+! echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=9999000" > "$TLOB_MONITOR" 2>/dev/null || false
+
+# Remove the binding; it must no longer appear.
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR"
+! grep -q "^p .*:0x${busy_offset#0x} " "$TLOB_MONITOR" || false
+
+echo 0 > monitors/tlob/enable
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
new file mode 100644
index 000000000000..e1e857144237
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
@@ -0,0 +1,53 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail running (running_ns dominates when task busy-spins between probes)
+# requires: tlob:monitor
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 5000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# 10 us budget; task busy-spins 200 ms per iteration -> running_ns dominates.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=10000" > "$TLOB_MONITOR"
+
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" ../trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > ../events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" ../trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+# Busy-spin keeps the task on-CPU: running_ns must exceed sleeping_ns.
+[ "$running" -gt "$sleeping" ]
+
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
new file mode 100644
index 000000000000..7165190cef7b
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
@@ -0,0 +1,52 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail sleeping (sleeping_ns dominates when task blocks between probes)
+# requires: tlob:monitor
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 5000 sleep &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# 50 ms budget; task sleeps 200 ms per iteration -> sleeping_ns dominates.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=50000000" > "$TLOB_MONITOR"
+
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" ../trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > ../events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" ../trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+[ "$sleeping" -gt "$((running + waiting))" ]
+
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
new file mode 100644
index 000000000000..798a04012b6e
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
@@ -0,0 +1,76 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail waiting (waiting_ns dominates when task is preempted between probes)
+# requires: tlob:monitor chrt:program taskset:program
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+# This test pins a SCHED_FIFO-99 hog on a dedicated CPU; at least 2 CPUs
+# are required so the test runner itself is never starved.
+[ "$(nproc)" -ge 2 ] || exit_unresolved
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work_done 2>/dev/null)
+
+# Pick the last CPU to avoid cpu0 which is used by vng infrastructure.
+cpu=$(($(nproc) - 1))
+
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# tlob_target loops calling tlob_preempt_work(200) / tlob_preempt_work_done()
+# in 200 ms wall-clock iterations. The stop probe fires when
+# tlob_preempt_work_done() is called, which cancels the budget timer.
+# Budget must be less than 200 ms so the HA timer fires while the target is
+# still inside tlob_preempt_work() and before the stop probe fires.
+# 150 ms gives a comfortable margin: waiting_ns ≈ 140 ms >> running_ns < 10 ms.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=150000000" > "$TLOB_MONITOR"
+
+# Start the RT hog BEFORE the target so the target is immediately preempted
+# when it calls tlob_preempt_work() (start probe fires), minimising running_ns.
+chrt -f 99 taskset -c "$cpu" sh -c 'while true; do :; done' 2>/dev/null &
+hog_pid=$!
+teardown() {
+ kill "$hog_pid" 2>/dev/null || true; wait "$hog_pid" 2>/dev/null || true
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.02
+
+taskset -c "$cpu" "$UPROBE_TARGET" 5000 preempt &
+busy_pid=$!
+
+# Poll up to 3 s (budget 150 ms + generous margin).
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" ../trace && { found=1; break; }
+ i=$((i+1))
+done
+
+# Kill the RT hog first so tlob_target can release any in-flight SRCU read
+# section from uprobe_notify_resume; otherwise probe removal blocks in
+# synchronize_srcu with the hog monopolising the CPU at FIFO-99.
+kill "$hog_pid" 2>/dev/null || true; wait "$hog_pid" 2>/dev/null || true
+kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > ../events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" ../trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+[ "$waiting" -gt "$((running + sleeping))" ]
+
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
new file mode 100644
index 000000000000..b24ec9473f3c
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
@@ -0,0 +1,63 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor multiple uprobe bindings (different offsets fire independently)
+# requires: tlob:monitor
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+busy_stop=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+sleep_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work 2>/dev/null)
+sleep_stop=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 & # busy mode: tlob_busy_work fires every 200 ms
+busy_pid=$!
+"$UPROBE_TARGET" 30000 sleep & # sleep mode: tlob_sleep_work fires every 200 ms
+sleep_pid=$!
+teardown() {
+ kill "$sleep_pid" 2>/dev/null || true; wait "$sleep_pid" 2>/dev/null || true
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > ../events/rv/error_env_tlob/enable
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# Binding A: 5 s budget on the busy probe - must not fire in 200 ms loops.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${busy_stop} threshold=5000000000" > "$TLOB_MONITOR"
+# Binding B: 10 us budget on the sleep probe - fires on first invocation.
+echo "p ${UPROBE_TARGET}:${sleep_offset} ${sleep_stop} threshold=10000" > "$TLOB_MONITOR"
+
+# Wait up to 2 s for error_env_tlob from binding B.
+found=0; i=0
+while [ "$i" -lt 20 ]; do
+ sleep 0.1
+ grep -q "error_env_tlob" ../trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo "-${UPROBE_TARGET}:${sleep_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > monitors/tlob/enable
+echo 0 > ../events/rv/error_env_tlob/enable
+echo 0 > ../events/rv/detail_env_tlob/enable
+
+[ "$found" = "1" ]
+# error_env_tlob payload: clock variable must be present.
+# The event field can be "budget_exceeded" (hrtimer path) or the DA event
+# name ("sleep", "preempt") depending on which fires first; don't constrain it.
+grep "error_env_tlob" ../trace | head -n 1 | grep -q "clk_elapsed="
+# detail_env_tlob must appear alongside the error.
+grep -q "detail_env_tlob" ../trace
+
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
new file mode 100644
index 000000000000..23f8abeca090
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
@@ -0,0 +1,17 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor no spurious events without active uprobe binding
+# requires: tlob:monitor
+
+echo 1 > ../events/rv/error_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+sleep 0.5
+
+! grep -q "error_env_tlob" ../trace || false
+
+echo 0 > monitors/tlob/enable
+echo 0 > ../events/rv/error_env_tlob/enable
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
new file mode 100644
index 000000000000..fc3788c5fd0a
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
@@ -0,0 +1,79 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor restarting a parked window on the same task (no cross-window accumulator leak, unbind of a repeatedly-started task works)
+# requires: tlob:monitor
+
+set -x
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+# Diagnostic dump for a budget-violation regression: prints the exact
+# error_env_tlob entry (state + clk_elapsed at expiry), the per-state
+# accumulator breakdown, and the full start/stop transition sequence for
+# the monitored pid, so a false-positive violation can be told apart from
+# a genuine per-window overrun (and, if it is the latter, whether the
+# accumulators leaked across a restart).
+dump_tlob_diag() {
+ TRACE=../trace
+ echo "===== tlob restart diagnostic (pid ${busy_pid}) =====" >&2
+ echo "--- error_env_tlob (violations) ---" >&2
+ grep "error_env_tlob" "$TRACE" >&2 || true
+ echo "--- detail_env_tlob (accumulator breakdown) ---" >&2
+ grep "detail_env_tlob" "$TRACE" >&2 || true
+ echo "--- event_tlob (state transitions for target pid) ---" >&2
+ grep "event_tlob" "$TRACE" | grep ":${busy_pid}:" >&2 || true
+ echo "--- last 50 trace lines (full context) ---" >&2
+ tail -50 "$TRACE" >&2 || true
+}
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+# tlob_target loops calling tlob_busy_work(200)/tlob_busy_work_done() every
+# ~200ms for the whole run: each call is one start/stop window on the SAME
+# task, driving tlob_start_task()'s restart path (same pid, parked ws) on
+# every iteration after the first. 900ms gives ~4 such windows.
+"$UPROBE_TARGET" 900 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > ../events/rv/event_tlob/enable
+echo 1 > ../events/rv/error_env_tlob/enable
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# 300ms budget: comfortably covers one ~200ms window. If a restart failed
+# to reset ws->accs_ns[]/budget_exceeded (a regression this test exists to
+# catch), running_ns would keep growing across windows and the SECOND
+# window would already exceed budget (~400ms cumulative > 300ms). With a
+# correct reset, no window ever exceeds it.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=300000000" > "$TLOB_MONITOR"
+
+wait "$busy_pid" || true
+
+if grep -q "error_env_tlob" ../trace; then
+ dump_tlob_diag
+ false
+fi
+
+# The task has exited (tlob_destroy_task() ran via handle_sched_process_exit);
+# removing its now-stale binding must still succeed cleanly.
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR"
+! grep -q "^p .*:0x${busy_offset#0x} " "$TLOB_MONITOR" || false
+
+echo 0 > monitors/tlob/enable
+echo 0 > ../events/rv/event_tlob/enable
+echo 0 > ../events/rv/error_env_tlob/enable
+echo 0 > ../events/rv/detail_env_tlob/enable
+echo > ../trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
new file mode 100644
index 000000000000..cf8d38b3a2bc
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
@@ -0,0 +1,69 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor budget violation (error_env_tlob and detail_env_tlob fire with correct fields)
+# requires: tlob:monitor
+
+# Fall back to the verification directory relative to FTRACETEST_ROOT when not
+# set by make (e.g. in installed kselftest environments).
+: "${VERIFICATIONTEST_BINDIR:="$FTRACETEST_ROOT/../verification"}"
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > ../events/rv/error_env_tlob/enable
+echo 1 > ../events/rv/detail_env_tlob/enable
+echo 1 > ../tracing_on
+echo 1 > monitors/tlob/enable
+echo > ../trace
+
+# 10 us budget - fires almost immediately; task is busy-spinning on-CPU.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=10000" > "$TLOB_MONITOR"
+
+# wait up to 2 s for detail_env_tlob
+found=0; i=0
+while [ "$i" -lt 20 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" ../trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > ../events/rv/error_env_tlob/enable
+echo 0 > ../events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+# error_env_tlob must carry the clk_elapsed environment field.
+# The event label is "budget_exceeded" when detected by the hrtimer callback,
+# or the triggering sched event name when detected by the constraint path on a
+# preemption that races with the timer (common on PREEMPT_RT / VM). Both are
+# valid detections; check the env field instead of the label.
+grep "error_env_tlob" ../trace | head -n 1 | grep -q "clk_elapsed="
+
+# detail_env_tlob must have all five fields with the correct threshold
+line=$(grep "detail_env_tlob" ../trace | head -n 1)
+echo "$line" | grep -q "pid="
+echo "$line" | grep -q "threshold_ns=10000"
+echo "$line" | grep -q "running_ns="
+echo "$line" | grep -q "waiting_ns="
+echo "$line" | grep -q "sleeping_ns="
+
+# Busy-spin keeps the task on-CPU: running_ns must exceed sleeping_ns.
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+[ "$running" -gt "$sleeping" ]
+
+echo > ../trace
diff --git a/tools/testing/selftests/verification/tlob_sym.c b/tools/testing/selftests/verification/tlob_sym.c
new file mode 100644
index 000000000000..2d9561331d2f
--- /dev/null
+++ b/tools/testing/selftests/verification/tlob_sym.c
@@ -0,0 +1,225 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob_sym.c - ELF symbol-to-file-offset utility for tlob selftests
+ *
+ * Usage: tlob_sym sym_offset <binary> <symbol>
+ *
+ * Prints the ELF file offset of <symbol> in <binary> to stdout.
+ *
+ * Exit: 0 = found, 1 = error / not found.
+ */
+#include <elf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+static int sym_offset(const char *binary, const char *symname)
+{
+ int fd;
+ struct stat st;
+ void *map;
+ Elf64_Ehdr *ehdr;
+ Elf32_Ehdr *ehdr32;
+ int is64;
+ uint64_t sym_vaddr = 0;
+ int found = 0;
+ uint64_t file_offset = 0;
+
+ fd = open(binary, O_RDONLY);
+ if (fd < 0) {
+ fprintf(stderr, "open %s: %s\n", binary, strerror(errno));
+ return 1;
+ }
+ if (fstat(fd, &st) < 0) {
+ close(fd);
+ return 1;
+ }
+ map = mmap(NULL, (size_t)st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);
+ close(fd);
+ if (map == MAP_FAILED) {
+ fprintf(stderr, "mmap: %s\n", strerror(errno));
+ return 1;
+ }
+
+ ehdr = (Elf64_Ehdr *)map;
+ ehdr32 = (Elf32_Ehdr *)map;
+ if (st.st_size < 4 ||
+ ehdr->e_ident[EI_MAG0] != ELFMAG0 ||
+ ehdr->e_ident[EI_MAG1] != ELFMAG1 ||
+ ehdr->e_ident[EI_MAG2] != ELFMAG2 ||
+ ehdr->e_ident[EI_MAG3] != ELFMAG3) {
+ fprintf(stderr, "%s: not an ELF file\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ is64 = (ehdr->e_ident[EI_CLASS] == ELFCLASS64);
+
+ if (is64) {
+ Elf64_Shdr *shdrs;
+ Elf64_Shdr *shstrtab_hdr;
+
+ if (ehdr->e_shnum == 0 || ehdr->e_shstrndx >= ehdr->e_shnum ||
+ (uint64_t)ehdr->e_shoff +
+ (uint64_t)ehdr->e_shnum * sizeof(Elf64_Shdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF section table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ shdrs = (Elf64_Shdr *)((char *)map + ehdr->e_shoff);
+ shstrtab_hdr = &shdrs[ehdr->e_shstrndx];
+ const char *shstrtab = (char *)map + shstrtab_hdr->sh_offset;
+ int si;
+
+ for (int pass = 0; pass < 2 && !found; pass++) {
+ const char *target = pass ? ".dynsym" : ".symtab";
+
+ for (si = 0; si < ehdr->e_shnum && !found; si++) {
+ Elf64_Shdr *sh = &shdrs[si];
+ const char *name = shstrtab + sh->sh_name;
+
+ if (strcmp(name, target) != 0)
+ continue;
+
+ Elf64_Shdr *strtab_sh = &shdrs[sh->sh_link];
+ const char *strtab = (char *)map + strtab_sh->sh_offset;
+ Elf64_Sym *syms = (Elf64_Sym *)((char *)map + sh->sh_offset);
+ uint64_t nsyms = sh->sh_size / sizeof(Elf64_Sym);
+ uint64_t j;
+
+ for (j = 0; j < nsyms; j++) {
+ if (strcmp(strtab + syms[j].st_name, symname) == 0) {
+ sym_vaddr = syms[j].st_value;
+ found = 1;
+ break;
+ }
+ }
+ }
+ }
+
+ if (!found) {
+ fprintf(stderr, "symbol '%s' not found in %s\n", symname, binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ if (ehdr->e_phnum == 0 ||
+ (uint64_t)ehdr->e_phoff +
+ (uint64_t)ehdr->e_phnum * sizeof(Elf64_Phdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF program table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ Elf64_Phdr *phdrs = (Elf64_Phdr *)((char *)map + ehdr->e_phoff);
+ int pi;
+
+ for (pi = 0; pi < ehdr->e_phnum; pi++) {
+ Elf64_Phdr *ph = &phdrs[pi];
+
+ if (ph->p_type != PT_LOAD)
+ continue;
+ if (sym_vaddr >= ph->p_vaddr &&
+ sym_vaddr < ph->p_vaddr + ph->p_filesz) {
+ file_offset = sym_vaddr - ph->p_vaddr + ph->p_offset;
+ break;
+ }
+ }
+ } else {
+ Elf32_Shdr *shdrs;
+ Elf32_Shdr *shstrtab_hdr;
+
+ if (ehdr32->e_shnum == 0 || ehdr32->e_shstrndx >= ehdr32->e_shnum ||
+ (uint64_t)ehdr32->e_shoff +
+ (uint64_t)ehdr32->e_shnum * sizeof(Elf32_Shdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF section table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ shdrs = (Elf32_Shdr *)((char *)map + ehdr32->e_shoff);
+ shstrtab_hdr = &shdrs[ehdr32->e_shstrndx];
+ const char *shstrtab = (char *)map + shstrtab_hdr->sh_offset;
+ int si;
+ uint32_t sym_vaddr32 = 0;
+
+ for (int pass = 0; pass < 2 && !found; pass++) {
+ const char *target = pass ? ".dynsym" : ".symtab";
+
+ for (si = 0; si < ehdr32->e_shnum && !found; si++) {
+ Elf32_Shdr *sh = &shdrs[si];
+ const char *name = shstrtab + sh->sh_name;
+
+ if (strcmp(name, target) != 0)
+ continue;
+
+ Elf32_Shdr *strtab_sh = &shdrs[sh->sh_link];
+ const char *strtab = (char *)map + strtab_sh->sh_offset;
+ Elf32_Sym *syms = (Elf32_Sym *)((char *)map + sh->sh_offset);
+ uint32_t nsyms = sh->sh_size / sizeof(Elf32_Sym);
+ uint32_t j;
+
+ for (j = 0; j < nsyms; j++) {
+ if (strcmp(strtab + syms[j].st_name, symname) == 0) {
+ sym_vaddr32 = syms[j].st_value;
+ found = 1;
+ break;
+ }
+ }
+ }
+ }
+
+ if (!found) {
+ fprintf(stderr, "symbol '%s' not found in %s\n", symname, binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ if (ehdr32->e_phnum == 0 ||
+ (uint64_t)ehdr32->e_phoff +
+ (uint64_t)ehdr32->e_phnum * sizeof(Elf32_Phdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF program table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ Elf32_Phdr *phdrs = (Elf32_Phdr *)((char *)map + ehdr32->e_phoff);
+ int pi;
+
+ for (pi = 0; pi < ehdr32->e_phnum; pi++) {
+ Elf32_Phdr *ph = &phdrs[pi];
+
+ if (ph->p_type != PT_LOAD)
+ continue;
+ if (sym_vaddr32 >= ph->p_vaddr &&
+ sym_vaddr32 < ph->p_vaddr + ph->p_filesz) {
+ file_offset = sym_vaddr32 - ph->p_vaddr + ph->p_offset;
+ break;
+ }
+ }
+ sym_vaddr = sym_vaddr32;
+ }
+
+ munmap(map, (size_t)st.st_size);
+
+ if (!file_offset && sym_vaddr) {
+ fprintf(stderr, "could not map vaddr 0x%lx to file offset\n",
+ (unsigned long)sym_vaddr);
+ return 1;
+ }
+
+ printf("0x%lx\n", (unsigned long)file_offset);
+ return 0;
+}
+
+int main(int argc, char *argv[])
+{
+ if (argc != 4 || strcmp(argv[1], "sym_offset") != 0) {
+ fprintf(stderr, "Usage: %s sym_offset <binary> <symbol>\n", argv[0]);
+ return 1;
+ }
+ return sym_offset(argv[2], argv[3]);
+}
diff --git a/tools/testing/selftests/verification/tlob_target.c b/tools/testing/selftests/verification/tlob_target.c
new file mode 100644
index 000000000000..b32ee6243c00
--- /dev/null
+++ b/tools/testing/selftests/verification/tlob_target.c
@@ -0,0 +1,138 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob_target.c - uprobe target binary for tlob selftests.
+ *
+ * Provides three start/stop probe pairs, each designed to exercise a
+ * different dominant component of the detail_env_tlob ns breakdown:
+ *
+ * tlob_busy_work / tlob_busy_work_done - busy-spin: running_ns dominates
+ * tlob_sleep_work / tlob_sleep_work_done - nanosleep: sleeping_ns dominates
+ * tlob_preempt_work / tlob_preempt_work_done - busy-spin + RT competitor:
+ * waiting_ns dominates
+ *
+ * Usage: tlob_target <duration_ms> [mode]
+ *
+ * mode is one of: busy (default), sleep, preempt.
+ * Loops in 200 ms iterations until <duration_ms> has elapsed
+ * (0 = run for ~24 hours).
+ */
+#define _GNU_SOURCE
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <time.h>
+
+#ifndef noinline
+#define noinline __attribute__((noinline))
+#endif
+
+static inline int timespec_before(const struct timespec *a,
+ const struct timespec *b)
+{
+ return a->tv_sec < b->tv_sec ||
+ (a->tv_sec == b->tv_sec && a->tv_nsec < b->tv_nsec);
+}
+
+static void timespec_add_ms(struct timespec *ts, unsigned long ms)
+{
+ ts->tv_sec += ms / 1000;
+ ts->tv_nsec += (long)(ms % 1000) * 1000000L;
+ if (ts->tv_nsec >= 1000000000L) {
+ ts->tv_sec++;
+ ts->tv_nsec -= 1000000000L;
+ }
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_busy_work_done(void)
+{
+ asm volatile("" ::: "memory");
+}
+
+/* start probe; busy-spin so running_ns dominates */
+noinline void tlob_busy_work(unsigned long duration_ms)
+{
+ struct timespec start, now;
+ unsigned long elapsed;
+
+ clock_gettime(CLOCK_MONOTONIC, &start);
+ do {
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ elapsed = (unsigned long)(now.tv_sec - start.tv_sec)
+ * 1000000000UL
+ + (unsigned long)(now.tv_nsec - start.tv_nsec);
+ } while (elapsed < duration_ms * 1000000UL);
+
+ tlob_busy_work_done();
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_sleep_work_done(void)
+{
+ asm volatile("" ::: "memory");
+}
+
+/* start probe; nanosleep so sleeping_ns dominates */
+noinline void tlob_sleep_work(unsigned long duration_ms)
+{
+ struct timespec ts = {
+ .tv_sec = duration_ms / 1000,
+ .tv_nsec = (long)(duration_ms % 1000) * 1000000L,
+ };
+ nanosleep(&ts, NULL);
+ tlob_sleep_work_done();
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_preempt_work_done(void)
+{
+ asm volatile("" ::: "memory");
+}
+
+/*
+ * start probe; busy-spin so an RT competitor on the same CPU drives
+ * waiting_ns (prev_state==0 -> preempt event, task stays runnable off-CPU).
+ */
+noinline void tlob_preempt_work(unsigned long duration_ms)
+{
+ struct timespec start, now;
+ unsigned long elapsed;
+
+ clock_gettime(CLOCK_MONOTONIC, &start);
+ do {
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ elapsed = (unsigned long)(now.tv_sec - start.tv_sec)
+ * 1000000000UL
+ + (unsigned long)(now.tv_nsec - start.tv_nsec);
+ } while (elapsed < duration_ms * 1000000UL);
+
+ tlob_preempt_work_done();
+}
+
+int main(int argc, char *argv[])
+{
+ unsigned long duration_ms = 0;
+ const char *mode = "busy";
+ struct timespec deadline, now;
+
+ if (argc >= 2)
+ duration_ms = strtoul(argv[1], NULL, 10);
+ if (argc >= 3)
+ mode = argv[2];
+
+ clock_gettime(CLOCK_MONOTONIC, &deadline);
+ timespec_add_ms(&deadline, duration_ms ? duration_ms : 86400000UL);
+
+ do {
+ if (strcmp(mode, "sleep") == 0)
+ tlob_sleep_work(200);
+ else if (strcmp(mode, "preempt") == 0)
+ tlob_preempt_work(200);
+ else
+ tlob_busy_work(200);
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ } while (timespec_before(&now, &deadline));
+
+ return 0;
+}
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
` (7 preceding siblings ...)
2026-08-20 16:45 ` [PATCH v6 8/9] selftests/verification: Add tlob selftests wen.yang
@ 2026-08-20 16:45 ` wen.yang
2026-08-20 16:58 ` sashiko-bot
8 siblings, 1 reply; 17+ messages in thread
From: wen.yang @ 2026-08-20 16:45 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
When a test directory that does not itself contain test.d/functions is
passed to ftracetest (e.g. verification/test.d/tlob/), ftracetest fell
back to its own functions file and lost the rv-specific check_requires
handling for ':monitor' and ':reactor' requirements.
Walk up the directory tree from OPT_TEST_DIR until a directory containing
test.d/functions is found. This allows monitor subdirectories to be passed
directly as the test root without placing a functions shim in each one.
The RV verification suite uses this so that
tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
can pass test.d/tlob/ to ftracetest and have it source
verification/test.d/functions (which understands ':monitor'/':reactor').
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
tools/testing/selftests/ftrace/ftracetest | 26 ++++++++++++++++++++---
1 file changed, 23 insertions(+), 3 deletions(-)
diff --git a/tools/testing/selftests/ftrace/ftracetest b/tools/testing/selftests/ftrace/ftracetest
index 0a56bf209f6c..8f9d9291bf4c 100755
--- a/tools/testing/selftests/ftrace/ftracetest
+++ b/tools/testing/selftests/ftrace/ftracetest
@@ -159,9 +159,29 @@ parse_opts() { # opts
if [ -n "$OPT_TEST_CASES" ]; then
TEST_CASES=$OPT_TEST_CASES
fi
- if [ -n "$OPT_TEST_DIR" -a -f "$OPT_TEST_DIR"/test.d/functions ]; then
- TOP_DIR=$OPT_TEST_DIR
- TEST_DIR=$TOP_DIR/test.d
+ if [ -n "$OPT_TEST_DIR" ]; then
+ # Walk up from OPT_TEST_DIR to find the nearest ancestor that contains
+ # test.d/functions. This allows a monitor subdirectory (e.g.
+ # verification/test.d/tlob/) to be passed directly without placing a
+ # dummy functions shim in each new subdirectory.
+ # Security: only walk within directories owned by root or the current user
+ # to prevent sourcing a functions file planted in a world-writable ancestor.
+ dir=$(realpath "$OPT_TEST_DIR" 2>/dev/null || echo "$OPT_TEST_DIR")
+ while [ "$dir" != "/" ] && [ "$dir" != "//" ]; do
+ if [ -f "$dir/test.d/functions" ]; then
+ # Verify the directory is not world-writable to prevent LPE.
+ dirperms=$(stat -c "%a" "$dir" 2>/dev/null || stat -f "%Lp" "$dir" 2>/dev/null)
+ case "$dirperms" in
+ *2|*3|*6|*7) ;; # world-writable: skip this candidate
+ *)
+ TOP_DIR=$dir
+ TEST_DIR=$TOP_DIR/test.d
+ break
+ ;;
+ esac
+ fi
+ dir=$(dirname "$dir")
+ done
fi
}
--
2.25.1
^ permalink raw reply related [flat|nested] 17+ messages in thread
* Re: [PATCH v6 3/9] rv: Add tlob model DOT file
2026-08-20 16:45 ` [PATCH v6 3/9] rv: Add tlob model DOT file wen.yang
@ 2026-08-20 16:53 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:53 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] The implemented DOT model contradicts explicit claims made in the commit message regarding states, transitions, and invariants.
--
--- Patch [3]: [PATCH v6 3/9] rv: Add tlob model DOT file ---
commit 8a3ea61817d2834335f8a9230c5ffaf2a08353b7
Author: Wen Yang <wen.yang@linux.dev>
rv: Add tlob model DOT file
This commit adds the Graphviz DOT specification of the tlob hybrid
automaton to the kernel's verification models.
> Add the Graphviz DOT specification of the tlob hybrid automaton to
> tools/verification/models/. The model has three states (running,
> waiting, sleeping), five transitions (switch_in, preempt, wakeup,
> sleep), and a single clock invariant clk_elapsed < BUDGET_NS() active
> in all states.
[Severity: Medium]
Does this description accurately reflect the implemented model?
The text mentions three states, five transitions, and an invariant active
in all states. However, the included DOT file defines a fourth "stopped"
state, two additional transitions ("stop" and "start"), and omits the
invariant from the "stopped" state:
tools/verification/models/tlob.dot:
"running" [label = "running\nclk_elapsed < BUDGET_NS()", color = green3];
"waiting" [label = "waiting\nclk_elapsed < BUDGET_NS()"];
"sleeping" [label = "sleeping\nclk_elapsed < BUDGET_NS()"];
"stopped" [label = "stopped"];
"running" -> "sleeping" [ label = "sleep" ];
"running" -> "waiting" [ label = "preempt" ];
"waiting" -> "running" [ label = "switch_in" ];
"sleeping" -> "waiting" [ label = "wakeup" ];
"running" -> "stopped" [ label = "stop" ];
"stopped" -> "running" [ label = "start;reset(clk_elapsed)" ];
Could the commit message be updated to match the actual automaton
implemented in the code?
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=3
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 8/9] selftests/verification: Add tlob selftests
2026-08-20 16:45 ` [PATCH v6 8/9] selftests/verification: Add tlob selftests wen.yang
@ 2026-08-20 16:56 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:56 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] Using `$(nproc) - 1` to determine a target CPU ID assumes that CPU IDs are contiguous and start from 0, which breaks on sparse affinity masks or offline CPUs.
- [Medium] The ELF parser in `tlob_sym.c` lacks sufficient bounds checks when validating headers and iterating over sections, leading to out-of-bounds memory accesses.
--
--- Patch [8]: [PATCH v6 8/9] selftests/verification: Add tlob selftests ---
commit cb49807af49abc34cb32dee7d7e1d06575cff80a
Author: Wen Yang <wen.yang@linux.dev>
selftests/verification: Add tlob selftests
This commit adds seven ftrace-style test scripts for the tlob RV monitor,
covering uprobe binding management, budget violation detection, and per-state
time accounting. It also includes a standalone runner and C helper programs
for target execution and ELF symbol parsing.
> diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
> new file mode 100644
> index 0000000000000..798a04012b6ea
> --- /dev/null
> +++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
[ ... ]
> +start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work 2>/dev/null)
> +stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work_done 2>/dev/null)
> +
> +# Pick the last CPU to avoid cpu0 which is used by vng infrastructure.
> +cpu=$(($(nproc) - 1))
> +
[Severity: Medium]
Does this CPU ID calculation safely handle systems with sparse CPU affinity
masks or offline CPUs?
If a system has non-contiguous CPU IDs (for example, if CPU 1 is offline but
CPUs 0 and 2 are online), nproc will evaluate to 2. This calculation would
yield CPU 1, which might cause the subsequent taskset commands to fail with
an invalid argument error since CPU 1 is unavailable.
> +echo 1 > ../events/rv/detail_env_tlob/enable
> +echo 1 > ../tracing_on
> +echo 1 > monitors/tlob/enable
> +echo > ../trace
[ ... ]
> diff --git a/tools/testing/selftests/verification/tlob_sym.c b/tools/testing/selftests/verification/tlob_sym.c
> new file mode 100644
> index 0000000000000..2d9561331d2f9
> --- /dev/null
> +++ b/tools/testing/selftests/verification/tlob_sym.c
[ ... ]
> + for (int pass = 0; pass < 2 && !found; pass++) {
> + const char *target = pass ? ".dynsym" : ".symtab";
> +
> + for (si = 0; si < ehdr->e_shnum && !found; si++) {
> + Elf64_Shdr *sh = &shdrs[si];
> + const char *name = shstrtab + sh->sh_name;
[Severity: Medium]
Is there a risk of an out-of-bounds read when accessing the string table
here?
It appears sh->sh_name is used as an offset into shstrtab and passed to
strcmp() without bounds checking against the actual string table size. Could
a malformed binary cause a segmentation fault when the tool attempts to read
beyond the mapped memory?
> +
> + if (strcmp(name, target) != 0)
> + continue;
> +
> + Elf64_Shdr *strtab_sh = &shdrs[sh->sh_link];
[Severity: Medium]
Could this array indexing access memory outside the section header table?
The sh->sh_link field is used directly as an index into shdrs. Since there
is no verification that sh->sh_link is less than ehdr->e_shnum, a malformed
ELF file could cause the tool to read arbitrary memory when mapping headers.
> + const char *strtab = (char *)map + strtab_sh->sh_offset;
> + Elf64_Sym *syms = (Elf64_Sym *)((char *)map + sh->sh_offset);
> + uint64_t nsyms = sh->sh_size / sizeof(Elf64_Sym);
> + uint64_t j;
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=8
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed
2026-08-20 16:45 ` [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
@ 2026-08-20 16:58 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:58 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] The directory traversal logic fails to check directory ownership and fails open on `stat` error, allowing arbitrary code execution as root via an attacker-controlled ancestor directory.
--
--- Patch [9]: [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed ---
commit 71305ea3fc4b8a7294d098d4bcdd8624fd21ca70
Author: Wen Yang <wen.yang@linux.dev>
selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed
This commit modifies ftracetest to walk up the directory tree from a passed
test directory until it finds a test.d/functions file. This allows monitor
subdirectories to be used directly as the test root without needing dummy
functions shims.
> diff --git a/tools/testing/selftests/ftrace/ftracetest b/tools/testing/selftests/ftrace/ftracetest
> index 0a56bf209f6c0..8f9d9291bf4ce 100755
> --- a/tools/testing/selftests/ftrace/ftracetest
> +++ b/tools/testing/selftests/ftrace/ftracetest
[ ... ]
> @@ -159,9 +159,29 @@ parse_opts() { # opts
> if [ -n "$OPT_TEST_CASES" ]; then
> TEST_CASES=$OPT_TEST_CASES
> fi
> - if [ -n "$OPT_TEST_DIR" -a -f "$OPT_TEST_DIR"/test.d/functions ]; then
> - TOP_DIR=$OPT_TEST_DIR
> - TEST_DIR=$TOP_DIR/test.d
> + if [ -n "$OPT_TEST_DIR" ]; then
> + # Walk up from OPT_TEST_DIR to find the nearest ancestor that contains
> + # test.d/functions. This allows a monitor subdirectory (e.g.
> + # verification/test.d/tlob/) to be passed directly without placing a
> + # dummy functions shim in each new subdirectory.
> + # Security: only walk within directories owned by root or the current user
> + # to prevent sourcing a functions file planted in a world-writable ancestor.
> + dir=$(realpath "$OPT_TEST_DIR" 2>/dev/null || echo "$OPT_TEST_DIR")
> + while [ "$dir" != "/" ] && [ "$dir" != "//" ]; do
> + if [ -f "$dir/test.d/functions" ]; then
> + # Verify the directory is not world-writable to prevent LPE.
> + dirperms=$(stat -c "%a" "$dir" 2>/dev/null || stat -f "%Lp" "$dir" 2>/dev/null)
> + case "$dirperms" in
> + *2|*3|*6|*7) ;; # world-writable: skip this candidate
> + *)
> + TOP_DIR=$dir
[Severity: High]
Does this logic actually check directory ownership as mentioned in the comments?
It appears the code only checks if the directory is world-writable, but omits
validating that it is owned by root or the current user. An attacker could
potentially create a directory with standard 0755 permissions, plant a
malicious test.d/functions script, and if a privileged user or automated test
harness runs ftracetest against a subdirectory, the script would traverse up
and source the payload.
Additionally, if the stat command fails for any reason, dirperms will be empty.
Would an empty string match the *) fallback case and cause the script to silently
fail open, accepting even a world-writable directory?
> + TEST_DIR=$TOP_DIR/test.d
> + break
> + ;;
> + esac
> + fi
> + dir=$(dirname "$dir")
> + done
> fi
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=9
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check
2026-08-20 16:45 ` [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
@ 2026-08-20 16:58 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:58 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
New issues:
- [High] Check-then-act (TOCTOU) race condition in first-invocation guard reset allows concurrent thread execution to overwrite an initialized invariant timestamp.
Pre-existing issues:
- [High] `ha_check_invariant_ns` calculates `time_ns - expire_ns`, which can underflow and cause valid invariants to falsely fail when the system uptime is less than `expire_ns`.
- [High] `ha_start_timer_ns` and `ha_start_timer_jiffy` suffer from integer underflow when elapsed time (`passed`) exceeds the timeout (`expire`), causing monitor timers to be scheduled far into the future (or at a random time) instead of expiring immediately.
--
--- Patch [4]: [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check ---
commit 1e15107a6865e6fddc264988e2a5c771a535486a
Author: Wen Yang <wen.yang@linux.dev>
rv: Fix ha_invariant_passed_ns silent bypass of invariant check
This patch fixes a silent bypass in ha_invariant_passed_ns() and
ha_invariant_passed_jiffy() by initializing env_store to the current clock
on the first invocation. It replaces the early return with a call to
ha_reset_clk_ns() to establish the invariant guard.
> diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
> index 6e1c7fe5449a3..e1738d199b289 100644
> --- a/include/rv/ha_monitor.h
> +++ b/include/rv/ha_monitor.h
> @@ -355,7 +355,7 @@ static inline u64 ha_invariant_passed_ns(struct ha_monitor *ha_mon, enum envs en
> if (env < 0 || env >= ENV_MAX_STORED)
> return 0;
> if (ha_monitor_env_invalid(ha_mon, env))
> - return 0;
> + ha_reset_clk_ns(ha_mon, env, time_ns);
[Severity: High]
Does this lockless state check and modification allow a check-then-act race?
Looking at ha_invariant_passed_ns() in include/rv/ha_monitor.h, if concurrent
threads execute this path for the same multi-state monitor, could Thread A and
Thread B both evaluate ha_monitor_env_invalid() as true?
This could lead to one thread blindly overwriting the invariant start time
initialized by the other, inappropriately pushing the epoch forward and masking
violations.
> return ha_get_env(ha_mon, env, time_ns);
> }
>
> @@ -375,6 +375,7 @@ static inline bool ha_check_invariant_jiffy(struct ha_monitor *ha_mon, enum envs
> {
> return time_after64(READ_ONCE(ha_mon->env_store[env]), get_jiffies_64() - expire_jiffy);
> }
> +
[Severity: High]
This isn't a bug introduced by this patch, but is there a risk of integer
underflow in the related ns check function?
include/rv/ha_monitor.h:ha_check_invariant_ns() {
return READ_ONCE(ha_mon->env_store[env]) >= time_ns - expire_ns;
}
If the system uptime (time_ns) is smaller than expire_ns (e.g. for a monitor
started early during boot with a large timeout budget), could time_ns - expire_ns
underflow?
This would immediately trigger a false positive invariant violation that could
cause a kernel panic if the panic reactor is configured.
[Severity: High]
This is a pre-existing issue, but could integer underflow occur in the timer
setup functions when the elapsed time exceeds the timeout?
include/rv/ha_monitor.h:ha_start_timer_ns() {
u64 passed = ha_invariant_passed_ns(ha_mon, env, time_ns);
ha_start_timer_jiffy(ha_mon, ENV_MAX_STORED,
nsecs_to_jiffies(expire - passed + TICK_NSEC - 1), time_ns);
}
If passed is greater than expire, expire - passed underflows to a massive unsigned
64-bit value.
When passed to nsecs_to_jiffies() (or jiffies_to_nsecs() in ha_start_timer_jiffy),
this can result in timers being scheduled far into the future (or at a random
time) instead of expiring immediately, nullifying the monitor's ability to
enforce latency budgets.
> /*
> * ha_invariant_passed_jiffy - prepare the invariant and return the time since reset
> */
> @@ -383,7 +384,7 @@ static inline u64 ha_invariant_passed_jiffy(struct ha_monitor *ha_mon, enum envs
> if (env < 0 || env >= ENV_MAX_STORED)
> return 0;
> if (ha_monitor_env_invalid(ha_mon, env))
> - return 0;
> + ha_reset_clk_jiffy(ha_mon, env);
[Severity: High]
Can concurrent threads race when resetting the jiffy clock here as well?
Similarly to ha_invariant_passed_ns(), could multiple CPUs evaluating
ha_monitor_env_invalid() at the same time in ha_invariant_passed_jiffy() cause
an initialized invariant timestamp to be silently overwritten?
> return ha_get_env(ha_mon, env, time_ns);
> }
>
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=4
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors
2026-08-20 16:45 ` [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
@ 2026-08-20 16:59 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:59 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] API state tracking design in `rv_uprobe` leaves stale pointers on error and breaks state validation for batch cleanup, trapping callers into a double-free or memory leak.
- [Medium] The commit message explicitly states a factually incorrect resource lifecycle that contradicts the code and creates an RCU use-after-free hazard if implemented as described.
--
--- Patch [2]: [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors ---
commit be3e5812586424192ba2d5028d63b1a9b959099b
Author: Wen Yang <wen.yang@linux.dev>
rv: Add generic uprobe infrastructure for RV monitors
Monitors that instrument user-space function boundaries need to resolve
paths, register uprobes, and deregister them safely. Provide a thin
wrapper so monitors share a single implementation of this boilerplate.
> struct rv_uprobe embeds struct uprobe_consumer directly, avoiding a
> separate heap allocation per probe. The struct holds a struct path for
> the probed binary so that the inode and its mount remain referenced for
> the full uprobe lifetime; uprobe_register() does not take its own
> reference to the inode. The path is released in
> rv_uprobe_unregister_nosync() after the consumer has been removed.
[Severity: Medium]
This isn't a bug in the code, but this statement contradicts the actual
implementation. Looking at the code, rv_uprobe_unregister_nosync() explicitly
leaves the path held to prevent RCU use-after-free hazards, relying on the
caller to release it after rv_uprobe_sync(). Could this description be
updated to match the code?
> diff --git a/kernel/trace/rv/rv_uprobe.c b/kernel/trace/rv/rv_uprobe.c
> new file mode 100644
> index 0000000000000..b412a8e28a6ec
> --- /dev/null
> +++ b/kernel/trace/rv/rv_uprobe.c
> @@ -0,0 +1,91 @@
[ ... ]
> +int rv_uprobe_register(const char *binpath, loff_t offset, struct rv_uprobe *p)
> +{
> + struct inode *inode;
> + int ret;
> +
> + ret = kern_path(binpath, LOOKUP_FOLLOW, &p->path);
> + if (ret)
> + return ret;
> +
> + if (!d_is_reg(p->path.dentry)) {
> + path_put(&p->path);
> + return -EINVAL;
> + }
[Severity: Medium]
Do these error paths leave a stale pointer in the caller's struct?
When path_put() is called, p->path.dentry is not cleared. If a caller is
using an array of struct rv_uprobe and tries to batch clean up failed
registrations, could they accidentally use this stale pointer and cause a
double-free?
> +
> + inode = d_real_inode(p->path.dentry);
> +
> + /* uprobe_register() takes no inode reference; the path is held in p->path */
> + p->uprobe = uprobe_register(inode, offset, 0, &p->uc);
> + if (IS_ERR(p->uprobe)) {
> + ret = PTR_ERR(p->uprobe);
> + p->uprobe = NULL;
> + path_put(&p->path);
> + return ret;
> + }
> +
> + return 0;
> +}
[ ... ]
> +void rv_uprobe_unregister(struct rv_uprobe *p)
> +{
> + if (!p || !p->uprobe)
> + return;
[Severity: Medium]
Does this early return cause a memory leak if the caller previously used
rv_uprobe_unregister_nosync()?
Since rv_uprobe_unregister_nosync() clears p->uprobe to NULL, a caller
attempting to use rv_uprobe_unregister() on it later will hit this early
return. This skips path_put(&p->path), permanently leaking the path
reference.
> +
> + uprobe_unregister_nosync(p->uprobe, &p->uc);
> + p->uprobe = NULL;
> + rv_uprobe_sync();
> + path_put(&p->path);
> +}
> +EXPORT_SYMBOL_GPL(rv_uprobe_unregister);
> +
> +/**
> + * rv_uprobe_unregister_nosync - dequeue an uprobe without waiting
> + */
> +void rv_uprobe_unregister_nosync(struct rv_uprobe *p)
> +{
> + if (!p || !p->uprobe)
> + return;
> +
> + uprobe_unregister_nosync(p->uprobe, &p->uc);
> + p->uprobe = NULL;
[Severity: Medium]
If p->uprobe is cleared here, how does the caller know which paths need to be
released later?
The documentation suggests batching deregistrations. However, if p->uprobe is
set to NULL here, any subsequent iteration over an array to conditionally call
path_put() has no reliable state indicator left to determine if the path was
actually held, potentially leading to incorrect cleanup.
> + /* path held; caller must call rv_uprobe_sync() then path_put(&p->path) */
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=2
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
2026-08-20 16:45 ` [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
@ 2026-08-20 16:59 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 16:59 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [High] Non-synchronous timer cancellation combined with `synchronize_rcu()` or `kfree_rcu()` leads to a Use-After-Free during monitor teardown.
--
--- Patch [5]: [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable ---
commit 693c5f5465506e090a4b8b44d5e3c72769ec0685
Author: Wen Yang <wen.yang@linux.dev>
rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
This commit wraps da_monitor_reset_hook and EVENT_NONE_LBL with ifndef guards
to allow HA-based monitors to provide their own implementations. It enables
monitors like tlob to define custom reset hooks for tasks such as cancelling
per-task hrtimers.
> diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
> index e1738d199b289..807b981eb548a 100644
> --- a/include/rv/ha_monitor.h
> +++ b/include/rv/ha_monitor.h
> @@ -36,8 +36,14 @@ static bool ha_monitor_handle_constraint(struct da_monitor *da_mon,
> da_id_type id);
> #define da_monitor_event_hook ha_monitor_handle_constraint
> #define da_monitor_init_hook ha_monitor_init_env
> +
> +/* Overrides must still call ha_monitor_reset_env() or cancel the timer. */
> +#ifndef da_monitor_reset_hook
> #define da_monitor_reset_hook ha_monitor_reset_env
> +#endif
> +#ifndef da_monitor_sync_hook
> #define da_monitor_sync_hook() synchronize_rcu()
> +#endif
[Severity: High]
This isn't a bug introduced by this patch, but is it safe to rely on
synchronize_rcu() to wait for cancelled timers during monitor teardown?
Looking at the global teardown path:
include/rv/da_monitor.h:da_monitor_destroy() {
...
da_monitor_reset_all();
da_monitor_sync_hook();
...
}
And per-object teardown:
include/rv/da_monitor.h:da_destroy_storage() {
...
da_monitor_reset_hook(&mon_storage->rv.da_mon);
hash_del_rcu(&mon_storage->node);
if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
call_rcu(&mon_storage->rcu, da_pool_return_cb);
else
kfree_rcu(mon_storage, rcu);
...
}
Both paths cancel timers non-synchronously in ha_monitor_reset_env():
include/rv/ha_monitor.h:ha_monitor_reset_env() {
...
ha_cancel_timer(ha_mon);
...
}
They then rely on RCU grace periods to wait for executing callbacks,
assuming the callbacks execute within an RCU read-side critical section:
include/rv/ha_monitor.h:__ha_monitor_timer_callback() {
...
guard(rcu)();
...
}
If the timer callback is dispatched but preempted or delayed before
entering its RCU read-side critical section, synchronize_rcu() will not
wait for it. When the callback resumes, couldn't it access the freed
monitor storage and cause a use-after-free?
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=5
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor
2026-08-20 16:45 ` [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor wen.yang
@ 2026-08-20 17:03 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-08-20 17:03 UTC (permalink / raw)
To: wen.yang; +Cc: linux-trace-kernel
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
New issues:
- [Critical] Concurrent task exit causes Use-After-Free and list corruption in `tlob_unbind_reap`.
- [High] Race between `tlob_stop_task` claiming the window and the HA timer callback suppresses the detail trace and hides budget violations.
- [High] `tlob_ws_lock` is a `spinlock_t` acquired in tracepoint and uprobe exception contexts, causing sleeping in atomic context on PREEMPT_RT.
Pre-existing issues:
- [Critical] Concurrent lockless calls to `hash_add_rcu` and `hash_del_rcu` on `da_monitor_ht` corrupt the hash bucket lists.
--
--- Patch [6]: [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor ---
commit ecef387f80f071ac66f0456fa83acfd88043286c
Author: Wen Yang <wen.yang@linux.dev>
Subject: rv: Add tlob hybrid automaton monitor
This patch introduces tlob (task latency over budget), a per-task hybrid
automaton RV monitor. It tracks wall-clock time across a user-delimited code
section using uprobes and emits warnings when the elapsed time exceeds a
configured threshold.
> diff --git a/kernel/trace/rv/monitors/tlob/tlob.c b/kernel/trace/rv/monitors/tlob/tlob.c
> new file mode 100644
> index 0000000000000..08b1bee884cc4
> --- /dev/null
> +++ b/kernel/trace/rv/monitors/tlob/tlob.c
[ ... ]
> +/* Serialises tlob_task_state ownership: restart, detach, unbind reap. */
> +static DEFINE_SPINLOCK(tlob_ws_lock);
[Severity: High]
Could this spinlock cause sleeping in atomic context warnings on PREEMPT_RT
kernels?
On PREEMPT_RT, spinlock_t maps to a sleepable rt_mutex. However, this lock
is acquired in atomic contexts. For example, in the tracepoint callback path:
handle_sched_process_exit()
tlob_destroy_task()
tlob_detach_from_binding()
And in the uprobe exception path:
tlob_uprobe_entry_handler()
tlob_start_task()
Should this use a raw_spinlock_t instead to avoid violating PREEMPT_RT
rules?
[ ... ]
> +static int tlob_stop_task(struct task_struct *task, void *binding)
> +{
> + struct ha_monitor *ha_mon;
> + struct tlob_task_state *ws;
> + bool budget_exceeded;
> +
> + scoped_guard(rcu) {
> + ha_mon = ha_get_monitor(task->pid, NULL);
> + if (!ha_mon)
> + return -ESRCH;
> +
> + ws = ha_get_target(ha_mon);
> + if (WARN_ON_ONCE(!ws))
> + return -ESRCH;
> +
> + /* Only the binding that opened the window may end it; NULL
> + * (task exit) skips the check. Symmetric with the restart
> + * check in tlob_start_task(). */
> + if (binding && ws->binding != binding)
> + return -EALREADY;
> +
> + /* cmpxchg (0->1) claims the window under RCU; _release pairs
> + * with the acquire in ha_setup_invariants(). */
> + if (atomic_cmpxchg_release(&ws->stopping, 0, 1) != 0)
> + return -EAGAIN;
[Severity: High]
Is there a race condition here if the latency budget expires at the exact
same time the task hits the stop uprobe?
If tlob_stop_task() wins the cmpxchg, it claims the window. The hrtimer
callback could then run concurrently and call tlob_reset_notify(), which
checks the stopping flag and returns early. This would skip emitting the
detail tracepoint and skip setting budget_exceeded to true.
Since tlob_stop_task() then reads budget_exceeded as false and returns 0,
could this silently hide genuine budget violations?
[ ... ]
> +static void tlob_destroy_task(struct task_struct *task)
> +{
[ ... ]
> + tlob_detach_from_binding(ws);
> +
> + /* Force the window ended: @task may never have reached STOP or a timer. */
> + atomic_set(&ws->stopping, 1);
> + ha_cancel_timer_sync(ha_mon);
> +
> + scoped_guard(rcu) {
> + da_monitor_reset(&ha_mon->da_mon);
> + }
> + da_destroy_storage(task->pid);
[Severity: Critical]
This is a pre-existing issue, but is it safe to call da_destroy_storage()
here without holding a lock?
Looking at da_destroy_storage() in include/rv/da_monitor.h, it calls
hash_del_rcu(&mon_storage->node). If multiple monitored tasks whose PIDs
map to the same hash bucket exit concurrently, or one exits while another
starts, calling hash_del_rcu() and hash_add_rcu() concurrently without
mutual exclusion on the same bucket can corrupt the hlist pointers.
[ ... ]
> +static void tlob_unbind_reap(struct tlob_uprobe_binding *b)
> +{
> + struct tlob_task_state *ws, *tmp;
> + LIST_HEAD(to_destroy);
> +
> + scoped_guard(spinlock, &tlob_ws_lock) {
> + list_for_each_entry_safe(ws, tmp, &b->started_list, started_node) {
> + list_del_init(&ws->started_node);
> + ws->binding = NULL;
> + if (atomic_read(&ws->stopping))
> + list_add_tail(&ws->started_node, &to_destroy);
> + }
> + }
> +
> + list_for_each_entry_safe(ws, tmp, &to_destroy, started_node) {
[Severity: Critical]
Can this iteration trigger a use-after-free if a parked task exits
concurrently?
When tlob_unbind_reap() moves parked windows to the local to_destroy list
and drops the spinlock, it iterates the list without holding rcu_read_lock().
If handle_sched_process_exit() invokes tlob_destroy_task() concurrently,
it skips detaching from the binding list (since ws->binding is already NULL)
and calls call_rcu() to free the task state. Without RCU read protection
here, the grace period could expire and free the memory while it is still
linked on the local to_destroy list.
> + list_del_init(&ws->started_node);
> + tlob_destroy_task(ws->task);
> + }
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1787243842.git.wen.yang@linux.dev?part=6
^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-08-20 17:03 UTC | newest]
Thread overview: 17+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-20 16:45 [PATCH v6 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-20 16:45 ` [PATCH v6 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
2026-08-20 16:45 ` [PATCH v6 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
2026-08-20 16:59 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 3/9] rv: Add tlob model DOT file wen.yang
2026-08-20 16:53 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
2026-08-20 16:58 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
2026-08-20 16:59 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 6/9] rv: Add tlob hybrid automaton monitor wen.yang
2026-08-20 17:03 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
2026-08-20 16:45 ` [PATCH v6 8/9] selftests/verification: Add tlob selftests wen.yang
2026-08-20 16:56 ` sashiko-bot
2026-08-20 16:45 ` [PATCH v6 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
2026-08-20 16:58 ` sashiko-bot
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.