* [PATCH v5 0/9] rv: Add task latency over budget RV monitor
@ 2026-08-19 18:15 wen.yang
2026-08-19 18:15 ` [PATCH v5 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
` (8 more replies)
0 siblings, 9 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
This series introduces tlob (task latency over budget), a per-task
hybrid automaton RV monitor that measures elapsed wall-clock time across
a user-delimited code section and emits an error when the elapsed time
exceeds a configurable budget.
The monitor tracks three states (running, waiting, sleeping) driven by
sched_switch and sched_wakeup tracepoints. A single clock invariant,
clk_elapsed < BUDGET_NS(), is enforced by a per-task hrtimer
(HRTIMER_MODE_REL_HARD). On expiry, error_env_tlob is emitted, followed
by detail_env_tlob carrying a per-state time breakdown (running_ns,
waiting_ns, sleeping_ns).
Tasks are registered for monitoring by writing to the tracefs monitor
file:
echo "p /path/to/binary:OFFSET_START OFFSET_STOP threshold=NS" \
> /sys/kernel/tracing/rv/monitors/tlob/monitor
Two uprobes delimit the measured section without modifying the target
binary. Multiple uprobe pairs can be active simultaneously. A task may
repeat the START/STOP sequence any number of times: the entry parks in
the automaton's "stopped" state and later restarts in place, reusing the
same pool slot (see Patch 6).
Series structure
----------------
Patch 1: rv: Introduce DA_MON_ALLOCATION_STRATEGY
Compile-time allocation strategy selector for per-object DA storage:
DA_ALLOC_AUTO (default kmalloc), DA_ALLOC_POOL (pre-allocated mempool),
DA_ALLOC_MANUAL (caller-managed). DA_MON_POOL_SIZE alone selects pool
mode.
Patch 2: rv: Add generic uprobe infrastructure for RV monitors
Thin wrapper over uprobe_consumer for path resolution, registration,
and safe synchronous teardown. struct rv_uprobe embeds
uprobe_consumer directly and holds the probed binary's path for its
lifetime, avoiding a separate heap allocation per probe.
Patch 3: rv: Add tlob model DOT file
Graphviz DOT specification of the tlob hybrid automaton.
Patch 4: rv: Fix ha_invariant_passed_ns silent bypass of invariant check
Without this fix, ha_invariant_passed_ns() returns 0 on first call and
leaves env_store at U64_MAX. Every subsequent state transition resets
the hrtimer to the full budget instead of the remaining time, so the
budget never expires for tasks that transition between states. This
causes tlob's detail_sleeping and detail_waiting selftests to hang.
This patch is a minimal fix in the current dual-representation
framework. Nam's series [1] refactors ha_monitor.h to a
single-representation model; once it lands, tlob will need to be
updated to the new API and this patch will be superseded by Gabriele's
planned framework-level initialisation improvement. We carry it here
to keep the series self-contained and testable.
[1] https://lore.kernel.org/lkml/08188c28f274da63a3f8549add3086a92aef45e5.1780908661.git.namcao@linutronix.de
Patch 5: rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
#ifndef guards for da_monitor_reset_hook and EVENT_NONE_LBL.
Patch 6: rv: Add tlob hybrid automaton monitor
Main tlob implementation, including the parked-window redesign
described below.
Patches 7-9: Tests
KUnit tests for the uprobe-line parser; eight ftrace-style selftests
for tlob; the ftracetest walk-up change needed to run them from the
verification subdirectory.
Changes since v4
----------------
Series-wide:
- Subject tags normalized to a single "rv:" prefix (was rv/da:, rv/ha:,
rv/tlob:) with summary words capitalized, matching the trace
subsystem style
- Comments tightened throughout all patches: narrative prose and
duplicated explanations removed, concise technical comments kept
- The ftracetest walk-up change (previously folded into the selftests
patch) is split out as its own final patch, per Gabriele's suggestion;
the series now has 9 patches
Patch 1 (DA storage):
- Pool storage redesigned from the lock-free llist stack to a
pre-allocated mempool: mempool_init_kmalloc_pool() pre-allocates all
slots at init and mempool_alloc_preallocated() pops one without
touching the allocator, bounding start latency and returning -ENOSPC
when exhausted; mempool_free() is safe from RCU-callback context
- da_extra_cleanup() hook added so per-object monitors can release
pool-backed storage; the pool is torn down only after rcu_barrier()
Patch 2 (uprobe):
- struct rv_uprobe keeps a struct path (not an inode pointer) to the
probed binary; the path is held from registration until after the sync
barrier and released with path_put(), since uprobe_register() takes no
inode reference of its own
Patch 6 (tlob): parked-window redesign
- The automaton gains a "stopped" state (see tlob.dot): ending a window
(STOP uprobe, budget-timer expiry) no longer frees the entry; it stays
parked in the hash table, and repeated START/STOP of the same region
restarts in place, reusing the pool slot, hash entry and task_struct
reference (no allocation, no get_task_struct, no hash churn)
- ws->stopping is now per-window (cleared on restart) while final
teardown is claimed per-task through ws->destroying; the monitor
disable path switches its claim to destroying as well
- Per-binding started_list plus tlob_unbind_reap(): removing a binding
destroys tasks currently parked under it and detaches active ones,
closing the leak where parked tasks lingered until exit
- Pool slots fenced at TLOB_MAX_MONITORED and backed by the mempool;
fresh starts past the cap return -ENOSPC
- __tlob_acc() gates on stopping so scheduler events never reach a
parked task (hence no "stopped" self-loops in the model)
- tlob_stop_task() now takes the firing binding and rejects (-EALREADY)
a STOP that does not own the active window, so nested monitored
regions cannot truncate a window mid-flight; task-exit cleanup passes
NULL to skip the check, symmetric with the restart check
Patch 8 (selftests):
- uprobe_restart.tc added: drives the target through START/STOP twice in
a loop, checking the second start restarts the parked window, plus an
unbind-while-parked scenario leaving no stale state
- Test helpers moved out of test.d/tlob/ into the verification test
root as tlob_sym.c / tlob_target.c; the per-directory Makefile is no
longer needed
- The verification Makefile now follows the standard lib.mk selftest
layout (TEST_GEN_FILES := tlob_sym tlob_target, export
VERIFICATIONTEST_BINDIR := $(OUTPUT)); the original empty all: target
is kept, so the diff over the base Makefile is exactly those two
lines. tlob_sym resolves ELF symbols in C, so the tests run without
any binutils dependency.
Testing
-------
Environment: x86_64, CONFIG_PREEMPT_RT=y, kernel built with virtme-ng
(vng --build), selftests built on the host and run inside the VM:
- KUnit test suite: tlob uprobe-line parser tests
- 8 ftrace-style selftests: uprobe_bind, uprobe_violation,
uprobe_no_event, uprobe_multi, uprobe_detail_{running,sleeping,
waiting}, uprobe_restart
Both suites pass. The pre-existing rv_* tests in the same directory
keep passing as well.
v4: https://lore.kernel.org/all/cover.1783524627.git.wen.yang@linux.dev/
v3: https://lore.kernel.org/all/cover.1780847473.git.wen.yang@linux.dev/
v2: https://lore.kernel.org/all/cover.1778522945.git.wen.yang@linux.dev/
v1: https://lore.kernel.org/all/cover.1776020428.git.wen.yang@linux.dev/
Wen Yang (9):
rv: Introduce DA_MON_ALLOCATION_STRATEGY
rv: Add generic uprobe infrastructure for RV monitors
rv: Add tlob model DOT file
rv: Fix ha_invariant_passed_ns silent bypass of invariant check
rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
rv: Add tlob hybrid automaton monitor
rv: Add KUnit tests for the tlob monitor
selftests/verification: Add tlob selftests
selftests/ftrace: Walk up to find test.d/functions when a subdirectory
is passed
Documentation/trace/rv/index.rst | 1 +
Documentation/trace/rv/monitor_tlob.rst | 194 +++
include/rv/da_monitor.h | 140 +-
include/rv/ha_monitor.h | 13 +-
include/rv/rv_uprobe.h | 90 ++
kernel/trace/rv/Kconfig | 5 +
kernel/trace/rv/Makefile | 3 +
kernel/trace/rv/monitors/tlob/.kunitconfig | 5 +
kernel/trace/rv/monitors/tlob/Kconfig | 22 +
kernel/trace/rv/monitors/tlob/tlob.c | 1134 +++++++++++++++++
kernel/trace/rv/monitors/tlob/tlob.h | 155 +++
kernel/trace/rv/monitors/tlob/tlob_kunit.c | 139 ++
kernel/trace/rv/monitors/tlob/tlob_trace.h | 48 +
kernel/trace/rv/rv_trace.h | 1 +
kernel/trace/rv/rv_uprobe.c | 91 ++
tools/testing/selftests/ftrace/ftracetest | 17 +-
.../testing/selftests/verification/.gitignore | 2 +
tools/testing/selftests/verification/Makefile | 4 +
.../test.d/tlob/run_tlob_tests.sh | 20 +
.../verification/test.d/tlob/uprobe_bind.tc | 35 +
.../test.d/tlob/uprobe_detail_running.tc | 49 +
.../test.d/tlob/uprobe_detail_sleeping.tc | 48 +
.../test.d/tlob/uprobe_detail_waiting.tc | 68 +
.../verification/test.d/tlob/uprobe_multi.tc | 59 +
.../test.d/tlob/uprobe_no_event.tc | 17 +
.../test.d/tlob/uprobe_restart.tc | 75 ++
.../test.d/tlob/uprobe_violation.tc | 65 +
.../testing/selftests/verification/tlob_sym.c | 209 +++
.../selftests/verification/tlob_target.c | 138 ++
tools/verification/models/tlob.dot | 25 +
30 files changed, 2860 insertions(+), 12 deletions(-)
create mode 100644 Documentation/trace/rv/monitor_tlob.rst
create mode 100644 include/rv/rv_uprobe.h
create mode 100644 kernel/trace/rv/monitors/tlob/.kunitconfig
create mode 100644 kernel/trace/rv/monitors/tlob/Kconfig
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.h
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_trace.h
create mode 100644 kernel/trace/rv/rv_uprobe.c
create mode 100755 tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
create mode 100644 tools/testing/selftests/verification/tlob_sym.c
create mode 100644 tools/testing/selftests/verification/tlob_target.c
create mode 100644 tools/verification/models/tlob.dot
base-commit: 785095112f4198de49760552374f364043c8dbdf
--
2.25.1
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH v5 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
` (7 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Per-object DA storage allocation is currently limited to kmalloc on
demand. Add a compile-time selector so monitors can choose among three
strategies:
DA_ALLOC_AUTO (default) - kmalloc per object on the monitor path
DA_ALLOC_POOL - pre-allocated fixed-size mempool;
selected by defining DA_MON_POOL_SIZE
DA_ALLOC_MANUAL - caller pre-inserts storage; framework
only links the target field
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/da_monitor.h | 140 ++++++++++++++++++++++++++++++++++++++--
1 file changed, 133 insertions(+), 7 deletions(-)
diff --git a/include/rv/da_monitor.h b/include/rv/da_monitor.h
index e3cf85c9ce55..48c534324cbb 100644
--- a/include/rv/da_monitor.h
+++ b/include/rv/da_monitor.h
@@ -14,7 +14,50 @@
#ifndef _RV_DA_MONITOR_H
#define _RV_DA_MONITOR_H
+/*
+ * Allocation strategies for RV_MON_PER_OBJ monitors, selected by defining
+ * one of the following before including this header (never both):
+ *
+ * DA_MON_POOL_SIZE N - pool mode, N pre-allocated slots
+ * DA_MON_ALLOCATION_STRATEGY - explicit strategy (see below)
+ * (neither) - auto mode (default)
+ *
+ * DA_ALLOC_AUTO - kmalloc on demand; unbounded.
+ * DA_ALLOC_POOL - pre-allocated fixed-size pool.
+ * DA_ALLOC_MANUAL - caller pre-inserts storage; framework links the target.
+ */
+#define DA_ALLOC_AUTO 0
+#define DA_ALLOC_POOL 1
+#define DA_ALLOC_MANUAL 2
+
+#ifdef DA_MON_POOL_SIZE
+#ifdef DA_MON_ALLOCATION_STRATEGY
+#error "Define only one of DA_MON_POOL_SIZE or DA_MON_ALLOCATION_STRATEGY"
+#endif
+#if DA_MON_POOL_SIZE == 0
+#error "DA_MON_POOL_SIZE must be non-zero"
+#endif
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_POOL
+#endif /* DA_MON_POOL_SIZE */
+
+#ifndef DA_MON_ALLOCATION_STRATEGY
+#ifdef DA_SKIP_AUTO_ALLOC
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_MANUAL
+#else
+#define DA_MON_ALLOCATION_STRATEGY DA_ALLOC_AUTO
+#endif
+#endif /* DA_MON_ALLOCATION_STRATEGY */
+
+/* Zero default keeps pool-mode conditionals compile-time constant. */
+#ifndef DA_MON_POOL_SIZE
+#if DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL
+#error "DA_ALLOC_POOL requires DA_MON_POOL_SIZE to be defined and non-zero"
+#endif
+#define DA_MON_POOL_SIZE 0
+#endif /* DA_MON_POOL_SIZE */
+
#include <rv/automata.h>
+#include <linux/mempool.h>
#include <linux/rv.h>
#include <rv/kunit.h>
#include <linux/stringify.h>
@@ -67,6 +110,16 @@ static struct rv_monitor rv_this;
#define da_monitor_sync_hook()
#endif
+/*
+ * Per-object teardown hook, called after da_monitor_reset_all() +
+ * da_monitor_sync_hook() and before hash_del_rcu() for each entry.
+ * All HA timer callbacks have completed at this point.
+ * Define before including this header. Default: no-op.
+ */
+#ifndef da_extra_cleanup
+#define da_extra_cleanup(da_mon)
+#endif
+
/*
* Type for the target id, default to int but can be overridden.
* A long type can work as hash table key (PER_OBJ) but will be downgraded to
@@ -543,6 +596,59 @@ static inline monitor_target da_get_target_by_id(da_id_type id)
return mon_storage->target;
}
+/*
+ * Pre-allocated mempool for DA_ALLOC_POOL monitors: DA_MON_POOL_SIZE
+ * slots, eager-allocated at init. mempool_alloc_preallocated() pops a
+ * slot without touching the allocator (bounded start latency; NULL when
+ * exhausted). mempool_free() is safe from RCU-callback context.
+ * Non-pool monitors get a zero-initialised mempool_t; pool paths compile
+ * away.
+ */
+static mempool_t da_monitor_pool;
+
+static void da_pool_return_cb(struct rcu_head *head)
+{
+ struct da_monitor_storage *ms =
+ container_of(head, struct da_monitor_storage, rcu);
+
+ mempool_free(ms, &da_monitor_pool);
+}
+
+/*
+ * da_create_pool_storage - pop a free pool slot and insert it into the hash.
+ *
+ * Returns the new da_monitor, or NULL if the pool is exhausted. Finding
+ * an existing entry for the same id fires WARN_ON_ONCE (double-start bug).
+ *
+ * Caller must hold an RCU read-side CS and the monitor's serialisation lock.
+ */
+static inline struct da_monitor *
+da_create_pool_storage(da_id_type id, monitor_target target,
+ struct da_monitor *da_mon)
+{
+ struct da_monitor_storage *mon_storage, *existing;
+
+ if (da_mon)
+ return da_mon;
+
+ mon_storage = mempool_alloc_preallocated(&da_monitor_pool);
+ if (!mon_storage)
+ return NULL;
+ memset(mon_storage, 0, sizeof(*mon_storage));
+
+ mon_storage->id = id;
+ mon_storage->target = target;
+
+ /* Single consumer under the caller's lock; duplicate is a double-start bug. */
+ existing = __da_get_mon_storage(id);
+ if (WARN_ON_ONCE(existing)) {
+ mempool_free(mon_storage, &da_monitor_pool);
+ return NULL;
+ }
+ hash_add_rcu(da_monitor_ht, &mon_storage->node, id);
+ return &mon_storage->rv.da_mon;
+}
+
/*
* da_destroy_storage - destroy the per-object storage
*
@@ -564,7 +670,10 @@ static inline void da_destroy_storage(da_id_type id)
return;
da_monitor_reset_hook(&mon_storage->rv.da_mon);
hash_del_rcu(&mon_storage->node);
- kfree_rcu(mon_storage, rcu);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ call_rcu(&mon_storage->rcu, da_pool_return_cb);
+ else
+ kfree_rcu(mon_storage, rcu);
}
static void __da_monitor_reset_all(void (*reset)(struct da_monitor *))
@@ -590,6 +699,9 @@ static inline void da_monitor_reset_state_all(void)
static inline int da_monitor_init(void)
{
hash_init(da_monitor_ht);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ return mempool_init_kmalloc_pool(&da_monitor_pool, DA_MON_POOL_SIZE,
+ sizeof(struct da_monitor_storage));
return 0;
}
@@ -607,8 +719,17 @@ static inline void da_monitor_destroy(void)
* pending, we can safely assume no concurrent user.
*/
hash_for_each_safe(da_monitor_ht, bkt, tmp, mon_storage, node) {
+ da_extra_cleanup(&mon_storage->rv.da_mon);
hash_del_rcu(&mon_storage->node);
- kfree(mon_storage);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ mempool_free(mon_storage, &da_monitor_pool);
+ else
+ kfree(mon_storage);
+ }
+
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL) {
+ rcu_barrier();
+ mempool_exit(&da_monitor_pool);
}
}
@@ -617,11 +738,16 @@ static inline void da_monitor_destroy(void)
* start condition is in a context problematic for allocation (e.g. scheduling).
* In such case, if the storage was pre-allocated without a target, set it now.
*/
-#ifdef DA_SKIP_AUTO_ALLOC
-#define da_prepare_storage da_fill_empty_storage
-#else
-#define da_prepare_storage da_create_storage
-#endif /* DA_SKIP_AUTO_ALLOC */
+static inline struct da_monitor *
+da_prepare_storage(da_id_type id, monitor_target target,
+ struct da_monitor *da_mon)
+{
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_POOL)
+ return da_create_pool_storage(id, target, da_mon);
+ if (DA_MON_ALLOCATION_STRATEGY == DA_ALLOC_MANUAL)
+ return da_fill_empty_storage(id, target, da_mon);
+ return da_create_storage(id, target, da_mon);
+}
#endif /* RV_MON_TYPE */
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 2/9] rv: Add generic uprobe infrastructure for RV monitors
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-19 18:15 ` [PATCH v5 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 3/9] rv: Add tlob model DOT file wen.yang
` (6 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Monitors that instrument user-space function boundaries need to resolve
paths, register uprobes, and deregister them safely. Provide a thin
wrapper so monitors share a single implementation of this boilerplate.
struct rv_uprobe embeds struct uprobe_consumer directly, avoiding a
separate heap allocation per probe. The struct holds a struct path for
the probed binary so that the inode and its mount remain referenced for
the full uprobe lifetime; uprobe_register() does not take its own
reference to the inode. The path is released in
rv_uprobe_unregister_nosync() after the consumer has been removed.
rv_uprobe_sync() calls uprobe_unregister_sync() which performs
synchronize_rcu_tasks_trace(), waiting for all rcu_read_lock_trace()
readers (handler_chain()) to complete on all CPUs before returning;
the caller may then free the containing struct.
The API provides register, synchronous and nosync unregister, a global
handler barrier (rv_uprobe_sync), and an active-state predicate.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/rv_uprobe.h | 90 ++++++++++++++++++++++++++++++++++++
kernel/trace/rv/rv_uprobe.c | 91 +++++++++++++++++++++++++++++++++++++
2 files changed, 181 insertions(+)
create mode 100644 include/rv/rv_uprobe.h
create mode 100644 kernel/trace/rv/rv_uprobe.c
diff --git a/include/rv/rv_uprobe.h b/include/rv/rv_uprobe.h
new file mode 100644
index 000000000000..d0a9079ac5be
--- /dev/null
+++ b/include/rv/rv_uprobe.h
@@ -0,0 +1,90 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (C) 2026 Wen Yang <wen.yang@linux.dev> */
+/*
+ * Generic uprobe infrastructure for RV monitors.
+ *
+ */
+
+#ifndef _RV_UPROBE_H
+#define _RV_UPROBE_H
+
+#include <linux/path.h>
+#include <linux/types.h>
+#include <linux/uprobes.h>
+
+struct pt_regs;
+
+/**
+ * struct rv_uprobe - embeddable uprobe handle for RV monitors
+ *
+ * Embed via DECLARE_RV_UPROBE() and pass &name to rv_uprobe_register().
+ * The caller may free the containing struct after rv_uprobe_unregister()
+ * (or rv_uprobe_unregister_nosync() + rv_uprobe_sync()) returns.
+ *
+ * @uc: embedded uprobe_consumer; set handler/ret_handler before registering
+ * @uprobe: registered uprobe pointer (NULL when not registered)
+ * @path: path of the probed binary, held until unregistration
+ */
+struct rv_uprobe {
+ struct uprobe_consumer uc;
+ struct uprobe *uprobe;
+ struct path path;
+};
+
+/* Embed a named rv_uprobe inside a caller struct */
+#define DECLARE_RV_UPROBE(name) struct rv_uprobe name
+
+/**
+ * rv_uprobe_is_registered - test whether an uprobe is currently active
+ * @p: probe to test; may be NULL
+ */
+bool rv_uprobe_is_registered(const struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_register - initialise and register an uprobe
+ * @binpath: absolute path to the target binary
+ * @offset: byte offset within the binary
+ * @p: caller-provided rv_uprobe (embedded via DECLARE_RV_UPROBE);
+ * p->uc.handler and/or p->uc.ret_handler must be set before this call
+ *
+ * Resolves the path and registers p->uc with the uprobe subsystem.
+ * No heap allocation is performed.
+ *
+ * Returns 0 on success, negative errno on failure.
+ */
+int rv_uprobe_register(const char *binpath, loff_t offset, struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_unregister - synchronously unregister a uprobe
+ * @p: probe to unregister; may be NULL (no-op)
+ *
+ * Removes the consumer from the uprobe subsystem and waits for all in-flight
+ * handlers to complete (via synchronize_rcu_tasks_trace()). After this
+ * returns, the containing struct may be safely freed by the caller.
+ * Use rv_uprobe_unregister_nosync() + rv_uprobe_sync() to batch multiple
+ * deregistrations before a single synchronisation.
+ */
+void rv_uprobe_unregister(struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_unregister_nosync - dequeue an uprobe without waiting
+ * @p: probe to dequeue; may be NULL (no-op)
+ *
+ * Removes the consumer without waiting for in-flight handlers. The path
+ * (p->path) is NOT released here; the caller must call rv_uprobe_sync()
+ * followed by path_put(&p->path) before freeing the containing struct.
+ * Use rv_uprobe_unregister() to handle both in one step.
+ */
+void rv_uprobe_unregister_nosync(struct rv_uprobe *p);
+
+/**
+ * rv_uprobe_sync - wait for all in-flight uprobe handlers to complete
+ *
+ * Global barrier: calls uprobe_unregister_sync(), which runs
+ * synchronize_rcu_tasks_trace() and synchronize_srcu(&uretprobes_srcu).
+ * After this returns, no handler_chain() iteration referencing any
+ * previously deregistered consumer is still in progress.
+ */
+void rv_uprobe_sync(void);
+
+#endif /* _RV_UPROBE_H */
diff --git a/kernel/trace/rv/rv_uprobe.c b/kernel/trace/rv/rv_uprobe.c
new file mode 100644
index 000000000000..b412a8e28a6e
--- /dev/null
+++ b/kernel/trace/rv/rv_uprobe.c
@@ -0,0 +1,91 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Generic uprobe infrastructure for RV monitors.
+ *
+ * rv_uprobe embeds struct uprobe_consumer; rv_uprobe_sync() drains in-flight
+ * handlers before the containing struct may be freed (see rv_uprobe.h).
+ */
+#include <linux/dcache.h>
+#include <linux/fs.h>
+#include <linux/namei.h>
+#include <linux/uprobes.h>
+#include <rv/rv_uprobe.h>
+
+/**
+ * rv_uprobe_register - initialise and register an uprobe
+ */
+int rv_uprobe_register(const char *binpath, loff_t offset, struct rv_uprobe *p)
+{
+ struct inode *inode;
+ int ret;
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &p->path);
+ if (ret)
+ return ret;
+
+ if (!d_is_reg(p->path.dentry)) {
+ path_put(&p->path);
+ return -EINVAL;
+ }
+
+ inode = d_real_inode(p->path.dentry);
+
+ /* uprobe_register() takes no inode reference; the path is held in p->path */
+ p->uprobe = uprobe_register(inode, offset, 0, &p->uc);
+ if (IS_ERR(p->uprobe)) {
+ ret = PTR_ERR(p->uprobe);
+ p->uprobe = NULL;
+ path_put(&p->path);
+ return ret;
+ }
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_register);
+
+/**
+ * rv_uprobe_is_registered - test whether an uprobe is currently active
+ */
+bool rv_uprobe_is_registered(const struct rv_uprobe *p)
+{
+ return p && p->uprobe;
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_is_registered);
+
+/**
+ * rv_uprobe_unregister - synchronously unregister a uprobe
+ */
+void rv_uprobe_unregister(struct rv_uprobe *p)
+{
+ if (!p || !p->uprobe)
+ return;
+
+ uprobe_unregister_nosync(p->uprobe, &p->uc);
+ p->uprobe = NULL;
+ rv_uprobe_sync();
+ path_put(&p->path);
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_unregister);
+
+/**
+ * rv_uprobe_unregister_nosync - dequeue an uprobe without waiting
+ */
+void rv_uprobe_unregister_nosync(struct rv_uprobe *p)
+{
+ if (!p || !p->uprobe)
+ return;
+
+ uprobe_unregister_nosync(p->uprobe, &p->uc);
+ p->uprobe = NULL;
+ /* path held; caller must call rv_uprobe_sync() then path_put(&p->path) */
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_unregister_nosync);
+
+/**
+ * rv_uprobe_sync - wait for all in-flight uprobe handlers to complete
+ */
+void rv_uprobe_sync(void)
+{
+ uprobe_unregister_sync();
+}
+EXPORT_SYMBOL_GPL(rv_uprobe_sync);
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 3/9] rv: Add tlob model DOT file
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-19 18:15 ` [PATCH v5 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
2026-08-19 18:15 ` [PATCH v5 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
` (5 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add the Graphviz DOT specification of the tlob hybrid automaton to
tools/verification/models/. The model has three states (running,
waiting, sleeping), five transitions (switch_in, preempt, wakeup,
sleep), and a single clock invariant clk_elapsed < BUDGET_NS() active
in all states.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
tools/verification/models/tlob.dot | 25 +++++++++++++++++++++++++
1 file changed, 25 insertions(+)
create mode 100644 tools/verification/models/tlob.dot
diff --git a/tools/verification/models/tlob.dot b/tools/verification/models/tlob.dot
new file mode 100644
index 000000000000..7f6f09c3d7cf
--- /dev/null
+++ b/tools/verification/models/tlob.dot
@@ -0,0 +1,25 @@
+digraph state_automaton {
+ center = true;
+ size = "7,11";
+ {node [shape = plaintext, style=invis, label=""] "__init_stopped"};
+ {node [shape = ellipse] "running"};
+ {node [shape = plaintext] "running"};
+ {node [shape = plaintext] "waiting"};
+ {node [shape = plaintext] "sleeping"};
+ {node [shape = plaintext] "stopped"};
+ "__init_stopped" -> "stopped";
+ "running" [label = "running\nclk_elapsed < BUDGET_NS()", color = green3];
+ "waiting" [label = "waiting\nclk_elapsed < BUDGET_NS()"];
+ "sleeping" [label = "sleeping\nclk_elapsed < BUDGET_NS()"];
+ "stopped" [label = "stopped"];
+ "running" -> "sleeping" [ label = "sleep" ];
+ "running" -> "waiting" [ label = "preempt" ];
+ "waiting" -> "running" [ label = "switch_in" ];
+ "sleeping" -> "waiting" [ label = "wakeup" ];
+ "running" -> "stopped" [ label = "stop" ];
+ "stopped" -> "running" [ label = "start;reset(clk_elapsed)" ];
+ { rank = min ;
+ "__init_stopped";
+ "stopped";
+ }
+}
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (2 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 3/9] rv: Add tlob model DOT file wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
` (4 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
When env_store is U64_MAX (its initial sentinel value),
ha_invariant_passed_ns() returns 0 immediately without initializing
env_store to the current clock. Subsequent calls to
ha_check_invariant_ns() then find env_store still at U64_MAX, causing
the elapsed comparison to wrap and always report the invariant as
satisfied, silently masking any violations.
Fix by calling ha_reset_clk_ns() to establish the guard on the first
invocation instead of returning early. Apply the same fix to
ha_invariant_passed_jiffy().
This is a stopgap: once the RV framework reworks the per-env clock
guard, this first-invocation reset should be subsumed.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/ha_monitor.h | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
index 6e1c7fe5449a..e1738d199b28 100644
--- a/include/rv/ha_monitor.h
+++ b/include/rv/ha_monitor.h
@@ -355,7 +355,7 @@ static inline u64 ha_invariant_passed_ns(struct ha_monitor *ha_mon, enum envs en
if (env < 0 || env >= ENV_MAX_STORED)
return 0;
if (ha_monitor_env_invalid(ha_mon, env))
- return 0;
+ ha_reset_clk_ns(ha_mon, env, time_ns);
return ha_get_env(ha_mon, env, time_ns);
}
@@ -375,6 +375,7 @@ static inline bool ha_check_invariant_jiffy(struct ha_monitor *ha_mon, enum envs
{
return time_after64(READ_ONCE(ha_mon->env_store[env]), get_jiffies_64() - expire_jiffy);
}
+
/*
* ha_invariant_passed_jiffy - prepare the invariant and return the time since reset
*/
@@ -383,7 +384,7 @@ static inline u64 ha_invariant_passed_jiffy(struct ha_monitor *ha_mon, enum envs
if (env < 0 || env >= ENV_MAX_STORED)
return 0;
if (ha_monitor_env_invalid(ha_mon, env))
- return 0;
+ ha_reset_clk_jiffy(ha_mon, env);
return ha_get_env(ha_mon, env, time_ns);
}
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (3 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 6/9] rv: Add tlob hybrid automaton monitor wen.yang
` (3 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Wrap both definitions with #ifndef guards so HA-based monitors can
substitute their own implementations before including this header.
tlob uses this to define a reset hook that cancels per-task hrtimers
on monitor teardown.
Overrides must still call ha_monitor_reset_env() or cancel outstanding
timers to avoid timer UAF.
No behaviour change for monitors that do not override either macro.
Reviewed-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
include/rv/ha_monitor.h | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/include/rv/ha_monitor.h b/include/rv/ha_monitor.h
index e1738d199b28..807b981eb548 100644
--- a/include/rv/ha_monitor.h
+++ b/include/rv/ha_monitor.h
@@ -36,8 +36,14 @@ static bool ha_monitor_handle_constraint(struct da_monitor *da_mon,
da_id_type id);
#define da_monitor_event_hook ha_monitor_handle_constraint
#define da_monitor_init_hook ha_monitor_init_env
+
+/* Overrides must still call ha_monitor_reset_env() or cancel the timer. */
+#ifndef da_monitor_reset_hook
#define da_monitor_reset_hook ha_monitor_reset_env
+#endif
+#ifndef da_monitor_sync_hook
#define da_monitor_sync_hook() synchronize_rcu()
+#endif
#if !defined(HA_SKIP_AUTO_CLEANUP) && RV_MON_TYPE == RV_MON_PER_TASK
/*
@@ -75,7 +81,9 @@ _Static_assert(offsetof(struct ha_monitor, da_mon) == 0,
#define ENV_INVALID_VALUE U64_MAX
/* Error with no event occurs only on timeouts */
#define EVENT_NONE EVENT_MAX
+#ifndef EVENT_NONE_LBL
#define EVENT_NONE_LBL "none"
+#endif
#define ENV_BUFFER_SIZE 64
#ifdef CONFIG_RV_REACTORS
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 6/9] rv: Add tlob hybrid automaton monitor
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (4 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
` (2 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
tlob (task latency over budget) is a per-task hybrid automaton RV
monitor that tracks wall-clock time across a user-delimited code
section and emits an error when elapsed time exceeds a configurable
budget.
Four-state automaton (running, waiting, sleeping, stopped) driven by
sched_switch/sched_wakeup tracepoints and a user-visible start/stop
pair: "stop" only fires from running and parks the window in stopped,
where a later "start" restarts it in place (same pool slot, no
re-registration); both callers of stop run while the task is on CPU.
A single clk_elapsed < BUDGET_NS() invariant is enforced by a
per-task HRTIMER_MODE_REL_HARD timer; on expiry the monitor records a
per-state breakdown (running_ns, waiting_ns, sleeping_ns) before
emitting error_env_tlob.
Uprobe pairs are registered through a tracefs monitor file as
"p PATH:OFFSET_START OFFSET_STOP threshold=NS". A pre-allocated
mempool hard-caps concurrently monitored tasks at TLOB_MAX_MONITORED
(past it, fresh starts return -ENOSPC) with allocation-free start/stop
on the uprobe hot path.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
Documentation/trace/rv/index.rst | 1 +
Documentation/trace/rv/monitor_tlob.rst | 194 ++++
kernel/trace/rv/Kconfig | 5 +
kernel/trace/rv/Makefile | 2 +
kernel/trace/rv/monitors/tlob/Kconfig | 12 +
kernel/trace/rv/monitors/tlob/tlob.c | 1132 ++++++++++++++++++++
kernel/trace/rv/monitors/tlob/tlob.h | 149 +++
kernel/trace/rv/monitors/tlob/tlob_trace.h | 48 +
kernel/trace/rv/rv_trace.h | 1 +
9 files changed, 1544 insertions(+)
create mode 100644 Documentation/trace/rv/monitor_tlob.rst
create mode 100644 kernel/trace/rv/monitors/tlob/Kconfig
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.c
create mode 100644 kernel/trace/rv/monitors/tlob/tlob.h
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_trace.h
diff --git a/Documentation/trace/rv/index.rst b/Documentation/trace/rv/index.rst
index 29769f06bb0f..1501545b5f08 100644
--- a/Documentation/trace/rv/index.rst
+++ b/Documentation/trace/rv/index.rst
@@ -16,5 +16,6 @@ Runtime Verification
monitor_wwnr.rst
monitor_sched.rst
monitor_rtapp.rst
+ monitor_tlob.rst
monitor_stall.rst
monitor_deadline.rst
diff --git a/Documentation/trace/rv/monitor_tlob.rst b/Documentation/trace/rv/monitor_tlob.rst
new file mode 100644
index 000000000000..2e606b0e67a4
--- /dev/null
+++ b/Documentation/trace/rv/monitor_tlob.rst
@@ -0,0 +1,194 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+Monitor tlob
+============
+
+- Name: tlob - task latency over budget
+- Type: per-object hybrid automaton (RV_MON_PER_OBJ)
+- Author: Wen Yang <wen.yang@linux.dev>
+
+Description
+-----------
+
+The tlob monitor tracks per-task elapsed wall-clock time (CLOCK_MONOTONIC,
+spanning running, waiting, and sleeping states) and reports a violation when
+the monitored task exceeds a configurable per-invocation budget threshold.
+
+The monitor implements a four-state hybrid automaton with a single clock
+environment variable ``clk_elapsed``. The clock invariant
+``clk_elapsed < BUDGET_NS()`` is active in the ``running``, ``waiting``, and
+``sleeping`` states (``stopped`` has no invariant, hence no timer); when it
+is violated the HA timer fires and the framework emits ``error_env_tlob``
+then calls ``da_monitor_reset()`` automatically::
+
+ | (initial)
+ v
+ +--------------+ +----------+
+ | running | --------> | stopped |
+ |->+--------------+ <-------- +----------+
+ switch_in preempt sleep
+ | | |
+ | | |
+ | v v
+ +---------+ +---------+
+ | waiting | | sleeping|
+ +---------+ +---------+
+ ^ v
+ | wakeup |
+ | |
+ +------------+
+
+ A fourth state, ``stopped``, has no clock invariant (hence no timer).
+ ``running`` reaches it on ``stop`` (``tlob_stop_task()``, window ended,
+ per-task state parked rather than freed) and returns to ``running`` on
+ ``start`` (``tlob_start_task()`` restarting the same task's parked
+ window).
+
+ Key transitions:
+ running --(sleep)------> sleeping (task blocks waiting for a resource)
+ running --(preempt)----> waiting (task preempted, back in runqueue)
+ sleeping --(wakeup)-----> waiting (resource available, enters runqueue)
+ waiting --(switch_in)--> running (scheduler picks task, back on CPU)
+ running --(stop)-------> stopped (tlob_stop_task(): window ended, parked)
+ stopped --(start)------> running (tlob_start_task(): window restarted)
+
+ ``tlob_start_task()`` calls ``da_handle_start_run_event(task->pid, ws, start_tlob)``.
+ The ``start_tlob`` edge goes ``stopped`` -> ``running`` for both a fresh
+ allocation (the initial state is ``stopped``) and a parked window's restart;
+ there is no ``start`` self-loop on ``running`` (a running task's START is
+ rejected with ``-EALREADY``). The transition triggers ``ha_setup_invariants()``,
+ which anchors ``clk_elapsed`` and arms the budget timer automatically.
+ ``tlob_stop_task()`` cancels the HA timer synchronously
+ via ``ha_cancel_timer_sync()``, then dispatches the ``stop_tlob`` event
+ (running -> stopped) instead of resetting the monitor: the per-task state
+ is parked, not freed, so a later ``tlob_start_task()`` call for the same
+ task can restart it without reallocating. Final teardown (task exit,
+ uprobe unbind, or monitor disable) is what actually calls
+ ``da_monitor_reset()`` and frees the state.
+
+The non-running condition (monitor not yet started, or reset after a budget
+violation or monitor disable) is handled implicitly by the RV framework
+(``da_mon->monitoring == 0``) - it is not an explicit DA state. A
+``tlob_stop_task()`` does not reset the monitor: the window parks in the
+explicit ``stopped`` state.
+
+Per-task state lives in ``struct tlob_task_state`` which is stored as
+``monitor_target`` in the framework's ``da_monitor_storage``, indexed by
+pid. The per-invocation ``threshold_ns`` is read via
+``ha_get_target(ha_mon)->threshold_ns`` inside the HA constraint functions,
+following the same pattern as the ``nomiss`` monitor.
+
+Usage
+-----
+
+tracefs interface (uprobe-based external monitoring)
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``monitor`` tracefs file instruments an unmodified binary via uprobes.
+The format follows the ftrace ``uprobe_events`` convention (``PATH:OFFSET``
+for the probe location, ``key=value`` for configuration parameters)::
+
+ p PATH:OFFSET_START OFFSET_STOP threshold=NS
+
+The uprobe at ``OFFSET_START`` fires ``tlob_start_task()``; the uprobe at
+``OFFSET_STOP`` fires ``tlob_stop_task()``. Both offsets are ELF file
+offsets of entry points in ``PATH``. ``PATH`` may contain ``:``; the last
+``:`` in the ``PATH:OFFSET_START`` token is the separator.
+
+A given task may hit the START/STOP pair any number of times: each pair
+of hits is one independent measurement window, and the underlying
+per-task state is reused across windows rather than reallocated each
+time. Removing a binding while one of its tasks is between windows
+(parked, having already hit STOP) frees that task's state immediately;
+a task still inside a window when its binding is removed keeps running
+unaffected and is cleaned up normally when it next exits.
+
+To remove a binding, use ``-PATH:OFFSET_START``::
+
+ echo 1 > /sys/kernel/tracing/rv/monitors/tlob/enable
+
+ echo "p /usr/bin/myapp:0x12a0 0x12f0 threshold=5000000" \
+ > /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # Remove a binding
+ echo "-/usr/bin/myapp:0x12a0" > /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # List registered bindings
+ cat /sys/kernel/tracing/rv/monitors/tlob/monitor
+
+ # Read violations from the trace buffer
+ cat /sys/kernel/tracing/trace
+
+Violation tracepoints
+~~~~~~~~~~~~~~~~~~~~~
+
+Two tracepoints are emitted together on a budget violation:
+
+``error_env_tlob``
+ Standard HA clock-invariant tracepoint (emitted by the RV framework).
+ Fields: ``id`` (task pid), ``state``, ``event`` (``"budget_exceeded"``),
+ ``env`` (``"clk_elapsed"``).
+
+``detail_env_tlob``
+ Tlob-specific breakdown of elapsed time per DA state.
+ Fields: ``id`` (task pid), ``threshold_ns``, ``running_ns``,
+ ``waiting_ns``, ``sleeping_ns``.
+
+ Use ``detail_env_tlob`` to diagnose *which phase* consumed the budget:
+ high ``sleeping_ns`` indicates I/O latency; high ``waiting_ns`` indicates
+ scheduler pressure; high ``running_ns`` indicates a compute overrun.
+
+Example: correlate the two tracepoints to see the breakdown::
+
+ trace-cmd record -e error_env_tlob -e detail_env_tlob &
+ # ... run workload ...
+ trace-cmd report
+
+tracefs files
+~~~~~~~~~~~~~
+
+The following files are specific to tlob under
+``/sys/kernel/tracing/rv/monitors/tlob/``:
+
+``monitor`` (rw)
+ Write ``p PATH:OFFSET_START OFFSET_STOP threshold=NS``
+ to bind two entry uprobes. Write ``-PATH:OFFSET_START`` to remove a
+ binding. Read to list registered bindings in the same format.
+ See the `tracefs interface (uprobe-based external monitoring)`_ section above.
+
+Kernel API
+----------
+
+``tlob_start_task`` and ``tlob_stop_task`` are the implementation-level
+functions called by the uprobe entry/exit handlers; the interface is
+driven from userspace.
+
+.. kernel-doc:: kernel/trace/rv/monitors/tlob/tlob.c
+ :functions: tlob_start_task tlob_stop_task
+
+Design notes
+------------
+
+Limitations:
+
+- A fresh window dispatches ``start_tlob`` (initial ``stopped`` ->
+ ``running``) via ``da_handle_start_run_event(task->pid, ws, start_tlob)``,
+ so monitoring always begins in ``running``. Monitoring a non-current
+ task that is already in waiting or sleeping state at call time
+ misclassifies the first interval as ``running_ns``.
+- ``TASK_STOPPED`` and ``TASK_TRACED`` carry ``prev_state != 0`` and are
+ therefore counted as ``sleeping_ns``, indistinguishable from
+ I/O-blocked time.
+- ``sched_wakeup_new`` is not hooked. In practice this is not an issue
+ because ``tlob_start_task`` is always called from a running context.
+
+Specification
+-------------
+
+Graphviz DOT file in tools/verification/models/tlob.dot.
+
+KUnit tests under ``kernel/trace/rv/monitors/tlob/tlob_kunit.c``
+(CONFIG_TLOB_KUNIT_TEST).
+
+User-space integration tests under ``tools/testing/selftests/verification/``
+(requires CONFIG_RV_MON_TLOB=y and root).
diff --git a/kernel/trace/rv/Kconfig b/kernel/trace/rv/Kconfig
index efa930f94ea4..222b3bea8079 100644
--- a/kernel/trace/rv/Kconfig
+++ b/kernel/trace/rv/Kconfig
@@ -84,8 +84,13 @@ source "kernel/trace/rv/monitors/deadline/Kconfig"
source "kernel/trace/rv/monitors/nomiss/Kconfig"
# Add new deadline monitors here
+source "kernel/trace/rv/monitors/tlob/Kconfig"
# Add new monitors here
+config RV_UPROBE
+ bool
+ depends on RV && UPROBES
+
config RV_REACTORS
bool "Runtime verification reactors"
default y
diff --git a/kernel/trace/rv/Makefile b/kernel/trace/rv/Makefile
index cdbf68c84f5a..cd0ec11f0e05 100644
--- a/kernel/trace/rv/Makefile
+++ b/kernel/trace/rv/Makefile
@@ -21,7 +21,9 @@ obj-$(CONFIG_RV_MON_STALL) += monitors/stall/stall.o
obj-$(CONFIG_RV_MON_DEADLINE) += monitors/deadline/deadline.o
obj-$(CONFIG_RV_MON_NOMISS) += monitors/nomiss/nomiss.o
obj-$(CONFIG_RV_MON_WAKEUP) += monitors/wakeup/wakeup.o
+obj-$(CONFIG_RV_MON_TLOB) += monitors/tlob/tlob.o
# Add new monitors here
+obj-$(CONFIG_RV_UPROBE) += rv_uprobe.o
obj-$(CONFIG_RV_REACTORS) += rv_reactors.o
obj-$(CONFIG_RV_REACT_PRINTK) += reactor_printk.o
obj-$(CONFIG_RV_REACT_PANIC) += reactor_panic.o
diff --git a/kernel/trace/rv/monitors/tlob/Kconfig b/kernel/trace/rv/monitors/tlob/Kconfig
new file mode 100644
index 000000000000..aa43382073d2
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/Kconfig
@@ -0,0 +1,12 @@
+# SPDX-License-Identifier: GPL-2.0-only
+#
+config RV_MON_TLOB
+ bool "tlob monitor"
+ depends on RV && UPROBES && HIGH_RES_TIMERS
+ select HA_MON_EVENTS_ID
+ select RV_UPROBE
+ help
+ Enable the tlob (task latency over budget) hybrid-automaton RV
+ monitor. tlob tracks per-task elapsed wall-clock time across a
+ user-delimited code section and emits error_env_tlob when the
+ elapsed time exceeds a configurable per-invocation budget.
diff --git a/kernel/trace/rv/monitors/tlob/tlob.c b/kernel/trace/rv/monitors/tlob/tlob.c
new file mode 100644
index 000000000000..99acd34726f1
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob.c
@@ -0,0 +1,1132 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob: task latency over budget monitor
+ *
+ * Tracks the elapsed wall-clock time (CLOCK_MONOTONIC) of a marked code
+ * path and flags per-task latency-budget overruns. The hrtimer callback
+ * emits error_env_tlob on violation plus detail_env_tlob, a per-state
+ * (running/waiting/sleeping) time breakdown.
+ *
+ * RV_MON_PER_OBJ: per-task state (struct tlob_task_state) lives as
+ * monitor_target in the framework's hash table. One HA clock invariant:
+ * clk_elapsed < BUDGET_NS() in running/waiting/sleeping (stopped parks).
+ *
+ * Copyright (C) 2026 Wen Yang <wen.yang@linux.dev>
+ */
+#include <linux/kernel.h>
+#include <linux/mempool.h>
+#include <linux/module.h>
+#include <linux/init.h>
+#include <linux/namei.h>
+#include <linux/rv.h>
+#include <linux/slab.h>
+#include <kunit/visibility.h>
+#include <rv/instrumentation.h>
+#include <rv/rv_uprobe.h>
+#include <rv.h>
+
+#define MODULE_NAME "tlob"
+
+#include <trace/events/sched.h>
+#include <rv_trace.h>
+
+/*
+ * Per-task latency monitoring state. One instance per monitoring window.
+ * Stored as monitor_target in da_monitor_storage; freed via call_rcu.
+ */
+enum tlob_acc_idx {
+ TLOB_ACC_RUNNING,
+ TLOB_ACC_WAITING,
+ TLOB_ACC_SLEEPING,
+ TLOB_ACC_MAX,
+};
+
+struct tlob_task_state {
+ struct task_struct *task; /* via get_task_struct */
+ u64 threshold_ns; /* budget in nanoseconds */
+
+ /*
+ * Per-window: 1 = this window ended (stop or timer expiry). Blocks
+ * timer re-arm in ha_setup_invariants(); cleared on restart.
+ */
+ atomic_t stopping;
+
+ /*
+ * Per-task, one-shot: final teardown has claimed this slot; never
+ * reset (a window can end and restart, the task cannot). atomic_t
+ * so cmpxchg is well-defined on every arch.
+ */
+ atomic_t destroying;
+
+ bool budget_exceeded;
+
+ /*
+ * Opaque owner: the binding that started this task. Set once on
+ * fresh allocation (NULL for callers with no binding), cleared by
+ * tlob_unbind_reap() for an active task whose binding is removed.
+ * Immutable elsewhere. Protected by tlob_ws_lock.
+ */
+ void *binding;
+ /*
+ * Linked into binding->started_list for the whole lifetime (not just
+ * while parked) so unbind reaping finds parked and active tasks.
+ * Protected by tlob_ws_lock.
+ */
+ struct list_head started_node;
+
+ /* Serialises accs_ns[]; held briefly (hardirq-safe). */
+ raw_spinlock_t entry_lock;
+ u64 accs_ns[TLOB_ACC_MAX]; /* per-state elapsed ns */
+ ktime_t last_ts;
+
+ struct rcu_head rcu;
+};
+
+#define RV_MON_TYPE RV_MON_PER_OBJ
+#define HA_TIMER_TYPE HA_TIMER_HRTIMER
+
+typedef struct tlob_task_state *monitor_target;
+
+static inline void tlob_reset_notify(struct da_monitor *da_mon);
+#define da_monitor_reset_hook tlob_reset_notify
+
+static inline void tlob_extra_cleanup(struct da_monitor *da_mon);
+#define da_extra_cleanup tlob_extra_cleanup
+
+#define EVENT_NONE_LBL "budget_exceeded"
+
+#include "tlob.h"
+
+#define DA_MON_POOL_SIZE TLOB_MAX_MONITORED
+
+#include <rv/ha_monitor.h>
+
+/*
+ * da_monitor_reset_hook: runs on hrtimer expiry, final teardown, and
+ * monitor disable. A normal stop never resets: tlob_stop_task() dispatches
+ * "stop" (running -> stopped, tlob.dot) instead. Only timer expiry is a
+ * genuine budget violation.
+ */
+static inline void tlob_reset_notify(struct da_monitor *da_mon)
+{
+ struct ha_monitor *ha_mon = to_ha_monitor(da_mon);
+ struct tlob_task_state *ws;
+
+ ha_monitor_reset_env(da_mon);
+
+ ws = ha_get_target(ha_mon);
+ if (!ws)
+ return;
+
+ /*
+ * stopping==1 means tlob_stop_task() ended this window already.
+ * acquire pairs with the _release clear in ha_setup_invariants().
+ */
+ if (atomic_read_acquire(&ws->stopping))
+ return;
+
+ /*
+ * Monitor disable (ha_mon_destroying set) is not a violation: the
+ * teardown paths free ws regardless. Couples to an HA-layer flag
+ * with no public contract; a framework-level equivalent would be
+ * cleaner.
+ */
+ if (unlikely(READ_ONCE(ha_mon_destroying)))
+ return;
+
+ /* Genuine expiry: end the window so a later start takes the restart path. */
+ atomic_set(&ws->stopping, 1);
+
+ /* Stamped regardless of the tracepoint; tlob_stop_task() reads it. */
+ WRITE_ONCE(ws->budget_exceeded, true);
+
+ if (!trace_detail_env_tlob_enabled())
+ return;
+
+ unsigned int curr_state = READ_ONCE(da_mon->curr_state);
+ u64 accs[TLOB_ACC_MAX], partial_ns;
+ unsigned long flags;
+
+ /* Snapshot accumulators; partial_ns covers curr_state time not yet folded in. */
+ raw_spin_lock_irqsave(&ws->entry_lock, flags);
+ partial_ns = ktime_get_ns() - ktime_to_ns(ws->last_ts);
+ accs[TLOB_ACC_RUNNING] = ws->accs_ns[TLOB_ACC_RUNNING] +
+ (curr_state == running_tlob ? partial_ns : 0);
+ accs[TLOB_ACC_WAITING] = ws->accs_ns[TLOB_ACC_WAITING] +
+ (curr_state == waiting_tlob ? partial_ns : 0);
+ accs[TLOB_ACC_SLEEPING] = ws->accs_ns[TLOB_ACC_SLEEPING] +
+ (curr_state == sleeping_tlob ? partial_ns : 0);
+ raw_spin_unlock_irqrestore(&ws->entry_lock, flags);
+
+ trace_detail_env_tlob(da_get_id(da_mon), ws->threshold_ns,
+ accs[TLOB_ACC_RUNNING],
+ accs[TLOB_ACC_WAITING],
+ accs[TLOB_ACC_SLEEPING]);
+}
+
+#define BUDGET_NS(ha_mon) (ha_get_target(ha_mon)->threshold_ns)
+
+/* HA constraint functions (called by ha_monitor_handle_constraint) */
+
+static u64 ha_get_env(struct ha_monitor *ha_mon, enum envs_tlob env,
+ u64 time_ns)
+{
+ if (env == clk_elapsed_tlob)
+ return ha_get_clk_ns(ha_mon, env, time_ns);
+ return ENV_INVALID_VALUE;
+}
+
+/*
+ * Invariant: clk_elapsed < BUDGET_NS in running/waiting/sleeping. "stopped"
+ * is exempt: the parked period must not be measured against the old window's
+ * clock anchor (restart from "stopped" would otherwise spuriously overrun).
+ */
+static inline bool ha_verify_invariants(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (curr_state == stopped_tlob)
+ return true;
+ return ha_check_invariant_ns(ha_mon, clk_elapsed_tlob, time_ns, BUDGET_NS(ha_mon));
+}
+
+/*
+ * The clock stays in guard (anchor) representation all window: env_store
+ * holds the window-start timestamp, re-anchored on start/restart.
+ * ha_invariant_passed_ns() never stores the deadline representation (the
+ * framework dropped ha_set_invariant_ns(), commit ab2900ae252b), so calling
+ * ha_inv_to_guard() here would subtract BUDGET_NS from the anchor and skew
+ * every check by one budget. nomiss likewise never converts.
+ */
+
+/* No per-event guard conditions for tlob; invariants suffice. */
+static inline bool ha_verify_guards(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ return true;
+}
+
+/*
+ * Guard on stopping: a sched_switch after ha_cancel_timer_sync() would
+ * re-arm the timer (ODEBUG splat). _acquire pairs with cmpxchg_release in
+ * tlob_stop_task.
+ *
+ * Entering stopped_tlob also resets env_store to the invalid sentinel, so a
+ * restart re-anchors the clock; a stale anchor would wrap the restart's
+ * timer delay to ~U64_MAX.
+ */
+static inline void ha_setup_invariants(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (next_state == stopped_tlob) {
+ /*
+ * Window ending: reset env_store to the invalid sentinel so
+ * the next window gets a fresh clock anchor. Keep stopping==1
+ * so __tlob_acc() continues to block sched events while parked.
+ */
+ ha_monitor_reset_all_stored(ha_mon);
+ return;
+ }
+
+ if (atomic_read_acquire(&ha_get_target(ha_mon)->stopping)) {
+ /*
+ * Restart (stopped -> running): arm the timer, then clear
+ * stopping so __tlob_acc() admits sched events only once the
+ * state is already running_tlob. _release pairs with the
+ * acquires in __tlob_acc/tlob_reset_notify.
+ */
+ if (next_state < state_max_tlob)
+ ha_start_timer_ns(ha_mon, clk_elapsed_tlob, BUDGET_NS(ha_mon), time_ns);
+ atomic_set_release(&ha_get_target(ha_mon)->stopping, 0);
+ return;
+ }
+
+ if (next_state < state_max_tlob)
+ ha_start_timer_ns(ha_mon, clk_elapsed_tlob, BUDGET_NS(ha_mon), time_ns);
+ else
+ ha_cancel_timer(ha_mon);
+}
+
+static bool ha_verify_constraint(struct ha_monitor *ha_mon,
+ enum states curr_state, enum events event,
+ enum states next_state, u64 time_ns)
+{
+ if (!ha_verify_invariants(ha_mon, curr_state, event, next_state, time_ns))
+ return false;
+
+ if (!ha_verify_guards(ha_mon, curr_state, event, next_state, time_ns))
+ return false;
+
+ ha_setup_invariants(ha_mon, curr_state, event, next_state, time_ns);
+
+ return true;
+}
+
+/*
+ * Pre-allocated pool of TLOB_MAX_MONITORED slots. mempool_alloc_preallocated()
+ * pops a reserve slot without touching the allocator (bounded start latency;
+ * -ENOSPC past the cap). Slots return via destroy/cleanup; mempool_free() is
+ * safe from RCU-callback context.
+ */
+static mempool_t tlob_ws_pool;
+
+static void tlob_ws_return_cb(struct rcu_head *head)
+{
+ struct tlob_task_state *ws =
+ container_of(head, struct tlob_task_state, rcu);
+
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+/* Direct return without RCU delay (ws was never published to the hash). */
+static void tlob_ws_direct_return(struct tlob_task_state *ws)
+{
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+static struct tlob_task_state *tlob_ws_alloc(void)
+{
+ struct tlob_task_state *ws =
+ mempool_alloc_preallocated(&tlob_ws_pool);
+
+ if (!ws)
+ return NULL;
+
+ memset(ws, 0, sizeof(*ws));
+ INIT_LIST_HEAD(&ws->started_node);
+ return ws;
+}
+
+/*
+ * Uprobe binding list; protected by tlob_uprobe_mutex. When both are
+ * taken, tlob_uprobe_mutex is always acquired before tlob_ws_lock:
+ * inverting the order would be a silent lock-order inversion.
+ */
+static LIST_HEAD(tlob_uprobe_list);
+static DEFINE_MUTEX(tlob_uprobe_mutex);
+
+/* Serialises tlob_task_state ownership: restart, detach, unbind reap. */
+static DEFINE_SPINLOCK(tlob_ws_lock);
+
+/* Per-uprobe-binding state: a start + stop probe pair for one binary region. */
+struct tlob_uprobe_binding {
+ struct list_head list;
+ u64 threshold_ns;
+ char binpath[TLOB_MAX_PATH];
+ loff_t offset_start;
+ loff_t offset_stop;
+ /*
+ * All tlob_task_states this binding ever started, for each task's
+ * lifetime. Protected by tlob_ws_lock.
+ */
+ struct list_head started_list;
+ DECLARE_RV_UPROBE(start_probe);
+ DECLARE_RV_UPROBE(stop_probe);
+};
+
+/*
+ * Unlink ws from its binding's started_list before returning it to the pool.
+ * ws->binding is left stale: the next tlob_ws_alloc() memsets it, and the
+ * restart path checks destroying first. Idempotent (list_del_init no-op).
+ */
+static inline void tlob_detach_from_binding(struct tlob_task_state *ws)
+{
+ if (!ws->binding)
+ return;
+ guard(spinlock)(&tlob_ws_lock);
+ list_del_init(&ws->started_node);
+}
+
+/*
+ * Per-entry teardown during monitor disable. cmpxchg on destroying
+ * (0->1) claims ownership -- not stopping, which can be long-lived (a
+ * parked task).
+ *
+ * No timer cancel or locking needed: disable_tlob() already synced every
+ * uprobe/tracepoint, da_monitor_destroy() ran da_monitor_reset_all() +
+ * synchronize_rcu(), and ha_mon_destroying blocks new timer callbacks.
+ */
+static inline void tlob_extra_cleanup(struct da_monitor *da_mon)
+{
+ struct ha_monitor *ha_mon = to_ha_monitor(da_mon);
+ struct tlob_task_state *ws = ha_get_target(ha_mon);
+
+ if (!ws)
+ return;
+
+ if (atomic_cmpxchg_release(&ws->destroying, 0, 1) != 0)
+ return;
+
+ tlob_detach_from_binding(ws);
+ put_task_struct(ws->task);
+ /*
+ * da_monitor_destroy() has already called synchronize_rcu(); no
+ * reader holds ws. Return the slot directly without call_rcu.
+ */
+ mempool_free(ws, &tlob_ws_pool);
+}
+
+/*
+ * Accumulate elapsed ns into accs_ns[idx] since last_ts and advance it.
+ * Returns true if monitored with an active window. The stopping gate is
+ * what keeps scheduler events from reaching a parked task (no "stopped"
+ * self-loops, see tlob.h) and keeps accs_ns[] from growing while parked.
+ */
+static inline bool __tlob_acc(struct task_struct *task, ktime_t now,
+ enum tlob_acc_idx idx)
+{
+ struct tlob_task_state *ws;
+ unsigned long flags;
+
+ guard(rcu)();
+ ws = da_get_target_by_id(task->pid);
+ /* acquire pairs with the _release clear in ha_setup_invariants(). */
+ if (!ws || atomic_read_acquire(&ws->stopping))
+ return false;
+ raw_spin_lock_irqsave(&ws->entry_lock, flags);
+ ws->accs_ns[idx] += ktime_to_ns(ktime_sub(now, ws->last_ts));
+ ws->last_ts = now;
+ raw_spin_unlock_irqrestore(&ws->entry_lock, flags);
+ return true;
+}
+
+static inline bool tlob_acc_running(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_RUNNING);
+}
+
+static inline bool tlob_acc_waiting(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_WAITING);
+}
+
+/*
+ * handle_sched_switch - advance the DA on every context switch.
+ *
+ * Emits sleep (running -> sleeping), preempt (running -> waiting) for prev,
+ * and switch_in (waiting -> running) for next. One ktime_get() shared by
+ * both acc calls keeps prev/next on the same context-switch timestamp.
+ *
+ * No waiting->sleeping edge: a task blocks (calls schedule()) only on CPU
+ * (running); waiting means TASK_RUNNING on the runqueue.
+ */
+static void handle_sched_switch(void *data, bool preempt_unused,
+ struct task_struct *prev,
+ struct task_struct *next,
+ unsigned int prev_state)
+{
+ ktime_t now = ktime_get();
+ bool prev_preempted = (prev_state == 0);
+
+ if (tlob_acc_running(prev, now))
+ da_handle_event(prev->pid, NULL,
+ prev_preempted ? preempt_tlob : sleep_tlob);
+ if (tlob_acc_waiting(next, now))
+ da_handle_event(next->pid, NULL, switch_in_tlob);
+}
+
+static inline bool tlob_acc_sleeping(struct task_struct *task, ktime_t now)
+{
+ return __tlob_acc(task, now, TLOB_ACC_SLEEPING);
+}
+
+/*
+ * handle_sched_wakeup - sleeping -> waiting transition. try_to_wake_up()
+ * skips TASK_RUNNING tasks, so this never fires for running/waiting.
+ */
+static void handle_sched_wakeup(void *data, struct task_struct *p)
+{
+ ktime_t now = ktime_get();
+
+ if (tlob_acc_sleeping(p, now))
+ da_handle_event(p->pid, NULL, wakeup_tlob);
+}
+
+/* Forward decl: used by handle_sched_process_exit() and tlob_unbind_reap(). */
+static int tlob_stop_task(struct task_struct *task, void *binding);
+static void tlob_destroy_task(struct task_struct *task);
+
+/*
+ * handle_sched_process_exit - clean up a task that exits without hitting its
+ * STOP uprobe (killed, unmapped mid-region, ...). The task is always in
+ * running_tlob here: do_exit() runs in the task's own context, which
+ * required passing through switch_in_tlob (running). tlob_stop_task() ends
+ * the window (or is a harmless -EAGAIN/-ESRCH); tlob_destroy_task() then
+ * frees the slot, as no restart can follow exit.
+ */
+static void handle_sched_process_exit(void *data, struct task_struct *p,
+ bool group_dead)
+{
+ tlob_stop_task(p, NULL);
+ tlob_destroy_task(p);
+}
+
+/**
+ * tlob_start_task - begin monitoring @task with budget @threshold_ns ns.
+ * @task: Task to monitor; may be current or another task.
+ * @threshold_ns: Budget in ns, in [1000, TLOB_MAX_THRESHOLD_NS].
+ * @binding: Opaque owner, recorded on fresh allocation and checked for
+ * an exact match on restart; NULL for callers that never
+ * restart a parked window.
+ *
+ * Allocates a fresh entry if @task has none, or restarts a parked entry in
+ * place when @binding matches (see tlob.dot: "start" fires from both
+ * running and stopped).
+ *
+ * Returns 0, -ENODEV, -ERANGE, -EALREADY, -ESRCH, or -ENOSPC (fresh start
+ * past pool capacity).
+ */
+static int tlob_start_task(struct task_struct *task, u64 threshold_ns, void *binding)
+{
+ struct tlob_task_state *ws;
+
+ if (!da_monitor_enabled())
+ return -ENODEV;
+
+ if (threshold_ns < TLOB_MIN_THRESHOLD_NS ||
+ threshold_ns > TLOB_MAX_THRESHOLD_NS)
+ return -ERANGE;
+
+ /* Serialise duplicate-check + pool-slot claim; see tlob_ws_lock. */
+ guard(spinlock)(&tlob_ws_lock);
+
+ /*
+ * da_get_target_by_id() uses hash_for_each_possible_rcu(), which
+ * requires an RCU read-side critical section.
+ */
+ scoped_guard(rcu) {
+ ws = da_get_target_by_id(task->pid);
+ if (ws) {
+ if (!atomic_read(&ws->stopping))
+ return -EALREADY;
+ if (atomic_read(&ws->destroying))
+ return -ESRCH;
+ /*
+ * Exact match only. An orphaned parked ws (binding
+ * cleared while active, then parked) is not adopted:
+ * that would need re-linking into the new binding's
+ * started_list. Accepted gap; the slot is reclaimed
+ * at task exit.
+ */
+ if (ws->binding != binding)
+ return -EALREADY;
+
+ /* Restart in place: same slot, hash entry, task ref, list node. */
+ ws->threshold_ns = threshold_ns;
+ WRITE_ONCE(ws->budget_exceeded, false);
+ memset(ws->accs_ns, 0, sizeof(ws->accs_ns));
+ ws->last_ts = ktime_get();
+
+ /*
+ * Keep stopping set: __tlob_acc() gates out sched
+ * events until ha_setup_invariants() clears it after
+ * the state is running_tlob. Clearing here would let
+ * events hit stopped_tlob (INVALID transitions).
+ */
+
+ /* Only failure here: monitor disabled since the check above. */
+ if (!da_handle_start_run_event(task->pid, ws, start_tlob))
+ return -ENODEV;
+ return 0;
+ }
+ }
+
+ ws = tlob_ws_alloc();
+ if (!ws)
+ return -ENOSPC;
+
+ ws->task = task;
+ get_task_struct(task);
+ ws->threshold_ns = threshold_ns;
+ ws->last_ts = ktime_get();
+ raw_spin_lock_init(&ws->entry_lock);
+ ws->binding = binding;
+ if (binding)
+ list_add_tail(&ws->started_node,
+ &((struct tlob_uprobe_binding *)binding)->started_list);
+
+ /* Dispatch failed (pool exhausted or monitor disabled): unwind the slot. */
+ if (!da_handle_start_run_event(task->pid, ws, start_tlob)) {
+ if (binding)
+ list_del_init(&ws->started_node);
+ put_task_struct(task);
+ tlob_ws_direct_return(ws);
+ return -ENOSPC;
+ }
+
+ return 0;
+}
+
+/**
+ * tlob_stop_task - end the current monitoring window for @task.
+ * @task: Task to stop.
+ * @binding: Opaque owner; must match ws->binding to end a normal (uprobe)
+ * window. NULL (task exit) skips the check.
+ *
+ * Ends the window (dispatches "stop") but does NOT free the entry: it stays
+ * parked so a later tlob_start_task() can restart it. Call
+ * tlob_destroy_task() once @task will never restart.
+ *
+ * cmpxchg on stopping (0->1) under RCU claims ownership; the winner cancels
+ * the timer synchronously.
+ *
+ * Returns 0, -EOVERFLOW (budget exceeded), -ESRCH (not monitored),
+ * -EAGAIN (window already ended), or -EALREADY (owned by another binding).
+ */
+static int tlob_stop_task(struct task_struct *task, void *binding)
+{
+ struct ha_monitor *ha_mon;
+ struct tlob_task_state *ws;
+ bool budget_exceeded;
+
+ scoped_guard(rcu) {
+ ha_mon = ha_get_monitor(task->pid, NULL);
+ if (!ha_mon)
+ return -ESRCH;
+
+ ws = ha_get_target(ha_mon);
+ if (WARN_ON_ONCE(!ws))
+ return -ESRCH;
+
+ /* Only the binding that opened the window may end it; NULL
+ * (task exit) skips the check. Symmetric with the restart
+ * check in tlob_start_task(). */
+ if (binding && ws->binding != binding)
+ return -EALREADY;
+
+ /* cmpxchg (0->1) claims the window under RCU; _release pairs
+ * with the acquire in ha_setup_invariants(). */
+ if (atomic_cmpxchg_release(&ws->stopping, 0, 1) != 0)
+ return -EAGAIN;
+
+ /*
+ * ws may be destroyed concurrently (unbind -> call_rcu), so
+ * keep its access under RCU; dispatch re-looks-up under RCU.
+ */
+ ha_cancel_timer_sync(ha_mon);
+ budget_exceeded = READ_ONCE(ws->budget_exceeded);
+ }
+
+ /* running -> stopped: no reset or destroy, the entry stays parked. */
+ da_handle_event(task->pid, NULL, stop_tlob);
+
+ return budget_exceeded ? -EOVERFLOW : 0;
+}
+
+/*
+ * tlob_destroy_task - final teardown for @task's entry: frees the pool slot,
+ * drops the task_struct ref, removes the hash entry, whether active or parked.
+ * Idempotent via the destroying cmpxchg (same pattern as tlob_extra_cleanup()).
+ * Callers must end the window first (see handle_sched_process_exit()).
+ */
+static void tlob_destroy_task(struct task_struct *task)
+{
+ struct ha_monitor *ha_mon;
+ struct tlob_task_state *ws;
+
+ scoped_guard(rcu) {
+ ha_mon = ha_get_monitor(task->pid, NULL);
+ if (!ha_mon)
+ return;
+ ws = ha_get_target(ha_mon);
+ if (WARN_ON_ONCE(!ws))
+ return;
+ if (atomic_cmpxchg_release(&ws->destroying, 0, 1) != 0)
+ return;
+ }
+
+ tlob_detach_from_binding(ws);
+
+ /* Force the window ended: @task may never have reached STOP or a timer. */
+ atomic_set(&ws->stopping, 1);
+ ha_cancel_timer_sync(ha_mon);
+
+ scoped_guard(rcu) {
+ da_monitor_reset(&ha_mon->da_mon);
+ }
+ da_destroy_storage(task->pid);
+
+ put_task_struct(ws->task);
+ call_rcu(&ws->rcu, tlob_ws_return_cb);
+}
+
+static int tlob_uprobe_entry_handler(struct uprobe_consumer *self,
+ struct pt_regs *regs, __u64 *data)
+{
+ struct tlob_uprobe_binding *b =
+ container_of(self, struct tlob_uprobe_binding, start_probe.uc);
+
+ tlob_start_task(current, b->threshold_ns, b);
+ return 0;
+}
+
+static int tlob_uprobe_stop_handler(struct uprobe_consumer *self,
+ struct pt_regs *regs, __u64 *data)
+{
+ struct tlob_uprobe_binding *b =
+ container_of(self, struct tlob_uprobe_binding, stop_probe.uc);
+
+ tlob_stop_task(current, b);
+ return 0;
+}
+
+/*
+ * Register start + stop entry uprobes for a binding.
+ * Called with tlob_uprobe_mutex held.
+ */
+static int tlob_add_uprobe(u64 threshold_ns, const char *binpath,
+ loff_t offset_start, loff_t offset_stop)
+{
+ struct tlob_uprobe_binding *tmp_b;
+ char pathbuf[TLOB_MAX_PATH];
+ struct inode *inode;
+ struct path path __free(path_put) = {};
+ char *canon;
+ int ret;
+
+ if (binpath[0] != '/')
+ return -EINVAL;
+
+ struct tlob_uprobe_binding *b __free(kfree) = kzalloc_obj(*b, GFP_KERNEL);
+ if (!b)
+ return -ENOMEM;
+
+ b->threshold_ns = threshold_ns;
+ b->offset_start = offset_start;
+ b->offset_stop = offset_stop;
+ INIT_LIST_HEAD(&b->started_list);
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &path);
+ if (ret)
+ return ret;
+
+ if (!d_is_reg(path.dentry))
+ return -EINVAL;
+
+ inode = d_real_inode(path.dentry);
+
+ /* Reject duplicate start offset for the same binary inode. */
+ list_for_each_entry(tmp_b, &tlob_uprobe_list, list) {
+ if (tmp_b->offset_start == offset_start &&
+ rv_uprobe_is_registered(&tmp_b->start_probe) &&
+ d_real_inode(tmp_b->start_probe.path.dentry) == inode)
+ return -EEXIST;
+ }
+
+ canon = d_path(&path, pathbuf, sizeof(pathbuf));
+ if (IS_ERR(canon))
+ return PTR_ERR(canon);
+ strscpy(b->binpath, canon, sizeof(b->binpath));
+
+ b->start_probe.uc.handler = tlob_uprobe_entry_handler;
+ ret = rv_uprobe_register(b->binpath, offset_start, &b->start_probe);
+ if (ret)
+ return ret;
+
+ b->stop_probe.uc.handler = tlob_uprobe_stop_handler;
+ ret = rv_uprobe_register(b->binpath, offset_stop, &b->stop_probe);
+ if (ret) {
+ rv_uprobe_unregister(&b->start_probe);
+ return ret;
+ }
+
+ /* NOT "b = no_free_ptr(b)": the re-assignment would free the live node. */
+ list_add_tail(&no_free_ptr(b)->list, &tlob_uprobe_list);
+ return 0;
+}
+
+/*
+ * tlob_unbind_reap - detach every task @b started, destroy the parked ones.
+ *
+ * Caller must have unregistered @b's uprobes and called rv_uprobe_sync():
+ * no start/stop can then be in flight for @b, so started_list is safe to
+ * walk. Active tasks are detached (binding cleared) and left running,
+ * matching unbind behaviour today; parked tasks are destroyed, or their
+ * pool slot leaks until the task next exits.
+ */
+static void tlob_unbind_reap(struct tlob_uprobe_binding *b)
+{
+ struct tlob_task_state *ws, *tmp;
+ LIST_HEAD(to_destroy);
+
+ scoped_guard(spinlock, &tlob_ws_lock) {
+ list_for_each_entry_safe(ws, tmp, &b->started_list, started_node) {
+ list_del_init(&ws->started_node);
+ ws->binding = NULL;
+ if (atomic_read(&ws->stopping))
+ list_add_tail(&ws->started_node, &to_destroy);
+ }
+ }
+
+ list_for_each_entry_safe(ws, tmp, &to_destroy, started_node) {
+ list_del_init(&ws->started_node);
+ tlob_destroy_task(ws->task);
+ }
+}
+
+static int tlob_remove_uprobe_by_key(loff_t offset_start, const char *binpath)
+{
+ struct tlob_uprobe_binding *b, *tmp;
+ struct path remove_path;
+ struct inode *inode;
+ int ret;
+
+ ret = kern_path(binpath, LOOKUP_FOLLOW, &remove_path);
+ if (ret)
+ return ret;
+
+ inode = d_real_inode(remove_path.dentry);
+
+ ret = -ENOENT;
+ list_for_each_entry_safe(b, tmp, &tlob_uprobe_list, list) {
+ if (b->offset_start != offset_start)
+ continue;
+ if (d_real_inode(b->start_probe.path.dentry) != inode)
+ continue;
+ list_del(&b->list);
+ /*
+ * rv_uprobe_sync() may sleep; list_del() already made the
+ * binding invisible to new readers.
+ */
+ rv_uprobe_unregister_nosync(&b->start_probe);
+ rv_uprobe_unregister_nosync(&b->stop_probe);
+ rv_uprobe_sync();
+ tlob_unbind_reap(b);
+ path_put(&b->start_probe.path);
+ path_put(&b->stop_probe.path);
+ kfree(b);
+ ret = 0;
+ break;
+ }
+
+ path_put(&remove_path);
+ return ret;
+}
+
+static void tlob_remove_all_uprobes(void)
+{
+ struct tlob_uprobe_binding *b, *tmp;
+ LIST_HEAD(pending);
+
+ mutex_lock(&tlob_uprobe_mutex);
+ list_for_each_entry_safe(b, tmp, &tlob_uprobe_list, list) {
+ list_move(&b->list, &pending);
+ rv_uprobe_unregister_nosync(&b->start_probe);
+ rv_uprobe_unregister_nosync(&b->stop_probe);
+ }
+ mutex_unlock(&tlob_uprobe_mutex);
+
+ if (list_empty(&pending))
+ return;
+
+ /* One sync covers all dequeued probes: consumers are then safe to free. */
+ rv_uprobe_sync();
+
+ list_for_each_entry_safe(b, tmp, &pending, list) {
+ list_del(&b->list);
+ tlob_unbind_reap(b);
+ path_put(&b->start_probe.path);
+ path_put(&b->stop_probe.path);
+ kfree(b);
+ }
+}
+
+static ssize_t tlob_monitor_read(struct file *file,
+ char __user *ubuf,
+ size_t count, loff_t *ppos)
+{
+ const int line_sz = TLOB_MAX_PATH + 128;
+ struct tlob_uprobe_binding *b;
+ char *buf;
+ int n = 0, buf_sz, pos = 0;
+ ssize_t ret;
+
+ mutex_lock(&tlob_uprobe_mutex);
+ list_for_each_entry(b, &tlob_uprobe_list, list)
+ n++;
+
+ buf_sz = (n ? n : 1) * line_sz + 1;
+ buf = kmalloc(buf_sz, GFP_KERNEL);
+ if (!buf) {
+ mutex_unlock(&tlob_uprobe_mutex);
+ return -ENOMEM;
+ }
+
+ list_for_each_entry(b, &tlob_uprobe_list, list) {
+ pos += scnprintf(buf + pos, buf_sz - pos,
+ "p %s:0x%llx 0x%llx threshold=%llu\n",
+ b->binpath,
+ (unsigned long long)b->offset_start,
+ (unsigned long long)b->offset_stop,
+ b->threshold_ns);
+ }
+ mutex_unlock(&tlob_uprobe_mutex);
+
+ ret = simple_read_from_buffer(ubuf, count, ppos, buf, pos);
+ kfree(buf);
+ return ret;
+}
+
+/*
+ * Parse "p PATH:OFFSET_START OFFSET_STOP threshold=NS".
+ * PATH may contain ':'; the last ':' separates path from offset.
+ * Returns 0, -EINVAL, or -ERANGE.
+ */
+static int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
+ char **path_out,
+ loff_t *start_out, loff_t *stop_out)
+{
+ unsigned long long thr = 0, stop_val = 0;
+ long long start_val;
+ char *p, *path_token, *token, *colon;
+ bool got_stop = false, got_thr = false;
+ int n;
+
+ /* Must start with "p " */
+ if (buf[0] != 'p' || buf[1] != ' ')
+ return -EINVAL;
+
+ p = buf + 2;
+ while (*p == ' ')
+ p++;
+
+ /* First space-delimited token is PATH:OFFSET_START */
+ path_token = strsep(&p, " \t");
+ if (!path_token || !*path_token)
+ return -EINVAL;
+
+ /* Split at last ':' to handle paths that contain ':'. */
+ colon = strrchr(path_token, ':');
+ if (!colon || colon - path_token < 2)
+ return -EINVAL;
+ *colon = '\0';
+
+ if (path_token[0] != '/')
+ return -EINVAL;
+
+ n = 0;
+ if (sscanf(colon + 1, "%lli%n", &start_val, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (start_val < 0)
+ return -EINVAL;
+
+ /* Remaining tokens: OFFSET_STOP threshold=NS */
+ while (p && (token = strsep(&p, " \t")) != NULL) {
+ if (!*token)
+ continue;
+ if (strncmp(token, "threshold=", 10) == 0) {
+ if (kstrtoull(token + 10, 0, &thr))
+ return -EINVAL;
+ if (thr < TLOB_MIN_THRESHOLD_NS || thr > TLOB_MAX_THRESHOLD_NS)
+ return -ERANGE;
+ got_thr = true;
+ } else if (!got_stop) {
+ long long sv;
+
+ n = 0;
+ if (sscanf(token, "%lli%n", &sv, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (sv < 0)
+ return -EINVAL;
+ stop_val = (unsigned long long)sv;
+ got_stop = true;
+ } else {
+ return -EINVAL;
+ }
+ }
+
+ if (!got_stop || !got_thr)
+ return -EINVAL;
+ if (start_val == (long long)stop_val)
+ return -EINVAL;
+
+ *thr_out = thr;
+ *path_out = path_token;
+ *start_out = (loff_t)start_val;
+ *stop_out = (loff_t)stop_val;
+ return 0;
+}
+
+/*
+ * Parse "-PATH:OFFSET_START" (ftrace uprobe_events removal convention).
+ */
+static int tlob_parse_remove_line(char *buf, char **path_out,
+ loff_t *start_out)
+{
+ char *binpath, *colon;
+ long long off;
+ int n = 0;
+
+ if (buf[0] != '-')
+ return -EINVAL;
+ binpath = buf + 1;
+ if (binpath[0] != '/')
+ return -EINVAL;
+ colon = strrchr(binpath, ':');
+ if (!colon || colon - binpath < 2)
+ return -EINVAL;
+ *colon = '\0';
+ if (sscanf(colon + 1, "%lli%n", &off, &n) != 1 || n == 0)
+ return -EINVAL;
+ if (off < 0)
+ return -EINVAL;
+ *path_out = binpath;
+ *start_out = (loff_t)off;
+ return 0;
+}
+
+static int tlob_create_or_delete_uprobe(char *buf)
+{
+ loff_t offset_start, offset_stop;
+ u64 threshold_ns;
+ char *binpath;
+ int ret;
+
+ if (buf[0] == '-') {
+ ret = tlob_parse_remove_line(buf, &binpath, &offset_start);
+ if (ret)
+ return ret;
+ mutex_lock(&tlob_uprobe_mutex);
+ ret = tlob_remove_uprobe_by_key(offset_start, binpath);
+ mutex_unlock(&tlob_uprobe_mutex);
+ return ret;
+ }
+ ret = tlob_parse_uprobe_line(buf, &threshold_ns, &binpath,
+ &offset_start, &offset_stop);
+ if (ret)
+ return ret;
+ mutex_lock(&tlob_uprobe_mutex);
+ ret = tlob_add_uprobe(threshold_ns, binpath, offset_start, offset_stop);
+ mutex_unlock(&tlob_uprobe_mutex);
+ return ret;
+}
+
+static ssize_t tlob_monitor_write(struct file *file,
+ const char __user *ubuf,
+ size_t count, loff_t *ppos)
+{
+ char buf[TLOB_MAX_PATH + 128];
+
+ if (count >= sizeof(buf))
+ return -EINVAL;
+ if (copy_from_user(buf, ubuf, count))
+ return -EFAULT;
+ buf[count] = '\0';
+ if (count > 0 && buf[count - 1] == '\n')
+ buf[count - 1] = '\0';
+ return tlob_create_or_delete_uprobe(buf) ?: (ssize_t)count;
+}
+
+static const struct file_operations tlob_monitor_fops = {
+ .open = simple_open,
+ .read = tlob_monitor_read,
+ .write = tlob_monitor_write,
+ .llseek = noop_llseek,
+};
+
+static int __tlob_init_monitor(void)
+{
+ int retval;
+
+ retval = mempool_init_kmalloc_pool(&tlob_ws_pool, TLOB_MAX_MONITORED,
+ sizeof(struct tlob_task_state));
+ if (retval)
+ return retval;
+
+ retval = ha_monitor_init();
+ if (retval) {
+ mempool_exit(&tlob_ws_pool);
+ return retval;
+ }
+
+ rv_this.enabled = 1;
+ return 0;
+}
+
+static void __tlob_destroy_monitor(void)
+{
+ rv_this.enabled = 0;
+ tlob_remove_all_uprobes();
+ /*
+ * A grace period only makes the call_rcu()'d tlob_ws_return_cb()
+ * callbacks eligible to run; rcu_barrier() waits until they have all
+ * returned their slots before the pool is destroyed.
+ */
+ ha_monitor_destroy();
+ rcu_barrier();
+ mempool_exit(&tlob_ws_pool);
+}
+
+static int tlob_enable_hooks(void)
+{
+ rv_attach_trace_probe("tlob", sched_switch, handle_sched_switch);
+ rv_attach_trace_probe("tlob", sched_wakeup, handle_sched_wakeup);
+ rv_attach_trace_probe("tlob", sched_process_exit, handle_sched_process_exit);
+ return 0;
+}
+
+static void tlob_disable_hooks(void)
+{
+ rv_detach_trace_probe("tlob", sched_switch, handle_sched_switch);
+ rv_detach_trace_probe("tlob", sched_wakeup, handle_sched_wakeup);
+ rv_detach_trace_probe("tlob", sched_process_exit, handle_sched_process_exit);
+}
+
+static int enable_tlob(void)
+{
+ int retval;
+
+ retval = __tlob_init_monitor();
+ if (retval)
+ return retval;
+
+ return tlob_enable_hooks();
+}
+
+static void disable_tlob(void)
+{
+ tlob_disable_hooks();
+ __tlob_destroy_monitor();
+}
+
+static struct rv_monitor rv_this = {
+ .name = "tlob",
+ .description = "Per-task latency-over-budget monitor.",
+ .enable = enable_tlob,
+ .disable = disable_tlob,
+ .reset = da_monitor_reset_all,
+ .enabled = 0,
+};
+
+static int __init register_tlob(void)
+{
+ int ret;
+
+ ret = rv_register_monitor(&rv_this, NULL);
+ if (ret)
+ return ret;
+
+ if (rv_this.root_d) {
+ if (!rv_create_file("monitor", RV_MODE_WRITE, rv_this.root_d, NULL,
+ &tlob_monitor_fops)) {
+ rv_unregister_monitor(&rv_this);
+ return -ENOMEM;
+ }
+ }
+
+ return 0;
+}
+
+static void __exit unregister_tlob(void)
+{
+ rv_unregister_monitor(&rv_this);
+}
+
+module_init(register_tlob);
+module_exit(unregister_tlob);
+
+MODULE_LICENSE("GPL");
+MODULE_AUTHOR("Wen Yang <wen.yang@linux.dev>");
+MODULE_DESCRIPTION("tlob: task latency over budget per-task monitor.");
diff --git a/kernel/trace/rv/monitors/tlob/tlob.h b/kernel/trace/rv/monitors/tlob/tlob.h
new file mode 100644
index 000000000000..94e7382c2130
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob.h
@@ -0,0 +1,149 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _RV_TLOB_H
+#define _RV_TLOB_H
+
+/*
+ * C representation of the tlob hybrid automaton (see tlob.dot).
+ *
+ * States: stopped (initial; parked), running (on CPU), waiting (runqueue),
+ * sleeping (blocked). Events: start/stop (tlob_start_task/tlob_stop_task),
+ * sleep/preempt/wakeup/switch_in (sched tracepoints).
+ *
+ * "stop" fires only from running (both callers run on CPU); "stopped"
+ * leaves only via "start" (fresh start or in-place restart). running[start]
+ * is INVALID: a stray re-start must not silently reset the budget clock.
+ *
+ * Invariant: clk_elapsed < BUDGET_NS() in running/waiting/sleeping; stopped
+ * parks the window, no clock while parked. start re-inits the monitor
+ * (da_handle_start_run_event()); stop dispatches after ha_cancel_timer_sync();
+ * final teardown uses ha_cancel_timer_sync() + da_monitor_reset() +
+ * da_destroy_storage().
+ *
+ * Format: Documentation/trace/rv/deterministic_automata.rst
+ */
+
+#include <linux/rv.h>
+#include <linux/sched.h>
+
+#define MONITOR_NAME tlob
+
+enum states_tlob {
+ stopped_tlob,
+ running_tlob,
+ sleeping_tlob,
+ waiting_tlob,
+ state_max_tlob,
+};
+
+#define INVALID_STATE state_max_tlob
+
+enum events_tlob {
+ preempt_tlob,
+ sleep_tlob,
+ start_tlob,
+ stop_tlob,
+ switch_in_tlob,
+ wakeup_tlob,
+ event_max_tlob,
+};
+
+/*
+ * HA clock env: clk_elapsed, wall-clock since the window start; anchored in
+ * running/waiting/sleeping, cleared on stop.
+ */
+enum envs_tlob {
+ clk_elapsed_tlob,
+ env_max_tlob,
+ env_max_stored_tlob = env_max_tlob,
+};
+
+_Static_assert(env_max_stored_tlob <= MAX_HA_ENV_LEN, "Not enough slots");
+#define HA_CLK_NS
+
+struct automaton_tlob {
+ char *state_names[state_max_tlob];
+ char *event_names[event_max_tlob];
+ char *env_names[env_max_tlob];
+ unsigned char function[state_max_tlob][event_max_tlob];
+ unsigned char initial_state;
+ bool final_states[state_max_tlob];
+};
+
+static const struct automaton_tlob automaton_tlob = {
+ .state_names = {
+ "stopped",
+ "running",
+ "sleeping",
+ "waiting",
+ },
+ .event_names = {
+ "preempt",
+ "sleep",
+ "start",
+ "stop",
+ "switch_in",
+ "wakeup",
+ },
+ .env_names = {
+ "clk_elapsed",
+ },
+ .function = {
+ /* stopped (initial; window parked, sched events not routed) */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (not on CPU) */
+ running_tlob, /* start (tlob_start_task, fresh or restart) */
+ INVALID_STATE, /* stop (already stopped) */
+ INVALID_STATE, /* switch_in (not on CPU) */
+ INVALID_STATE, /* wakeup (not on CPU) */
+ },
+ /* running */
+ {
+ waiting_tlob, /* preempt (sched_switch, prev_state == 0) */
+ sleeping_tlob, /* sleep (sched_switch, prev_state != 0) */
+ INVALID_STATE, /* start (running task's START is -EALREADY) */
+ stopped_tlob, /* stop (tlob_stop_task) */
+ INVALID_STATE, /* switch_in (already on CPU) */
+ INVALID_STATE, /* wakeup (TASK_RUNNING can't be woken) */
+ },
+ /* sleeping */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (already sleeping) */
+ INVALID_STATE, /* start (not in running state) */
+ INVALID_STATE, /* stop (not in running state) */
+ INVALID_STATE, /* switch_in (must go through waiting first) */
+ waiting_tlob, /* wakeup */
+ },
+ /* waiting */
+ {
+ INVALID_STATE, /* preempt (not on CPU) */
+ INVALID_STATE, /* sleep (not on CPU) */
+ INVALID_STATE, /* start (not in running state) */
+ INVALID_STATE, /* stop (not in running state) */
+ running_tlob, /* switch_in */
+ INVALID_STATE, /* wakeup (already TASK_RUNNING) */
+ },
+ },
+ .initial_state = stopped_tlob,
+ .final_states = { 0, 1, 0, 0 },
+};
+
+/*
+ * Hard cap on concurrently monitored tasks. tlob_ws_pool pre-allocates
+ * this many slots; a fresh start past the cap returns -ENOSPC with bounded
+ * latency (mempool_alloc_preallocated() never touches the allocator).
+ * Restarts reuse the same slot.
+ */
+#define TLOB_MAX_MONITORED 64U
+
+/* Maximum binary path length for uprobe binding. */
+#define TLOB_MAX_PATH 256
+
+/* Minimum monitoring budget (1 us). */
+#define TLOB_MIN_THRESHOLD_NS 1000ULL
+
+/* Upper budget bound (1 hour): keeps the u64 ns accumulators far from overflow. */
+#define TLOB_MAX_THRESHOLD_NS 3600000000000ULL
+
+#endif /* _RV_TLOB_H */
diff --git a/kernel/trace/rv/monitors/tlob/tlob_trace.h b/kernel/trace/rv/monitors/tlob/tlob_trace.h
new file mode 100644
index 000000000000..b3a7cf4ad3ea
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob_trace.h
@@ -0,0 +1,48 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+
+/*
+ * Snippet to be included in rv_trace.h
+ */
+
+#ifdef CONFIG_RV_MON_TLOB
+DEFINE_EVENT(event_da_monitor_id, event_tlob,
+ TP_PROTO(int id, char *state, char *event,
+ char *next_state, bool final_state),
+ TP_ARGS(id, state, event, next_state, final_state));
+
+DEFINE_EVENT(error_da_monitor_id, error_tlob,
+ TP_PROTO(int id, char *state, char *event),
+ TP_ARGS(id, state, event));
+
+DEFINE_EVENT(error_env_da_monitor_id, error_env_tlob,
+ TP_PROTO(int id, char *state, char *event, char *env),
+ TP_ARGS(id, state, event, env));
+
+/*
+ * detail_env_tlob - per-state latency breakdown on budget violation.
+ * Emitted right after error_env_tlob from the hrtimer callback.
+ */
+TRACE_EVENT(detail_env_tlob,
+ TP_PROTO(int id, u64 threshold_ns,
+ u64 running_ns, u64 waiting_ns, u64 sleeping_ns),
+ TP_ARGS(id, threshold_ns, running_ns, waiting_ns, sleeping_ns),
+ TP_STRUCT__entry(
+ __field(int, id)
+ __field(u64, threshold_ns)
+ __field(u64, running_ns)
+ __field(u64, waiting_ns)
+ __field(u64, sleeping_ns)
+ ),
+ TP_fast_assign(
+ __entry->id = id;
+ __entry->threshold_ns = threshold_ns;
+ __entry->running_ns = running_ns;
+ __entry->waiting_ns = waiting_ns;
+ __entry->sleeping_ns = sleeping_ns;
+ ),
+ TP_printk("pid=%d threshold_ns=%llu"
+ " running_ns=%llu waiting_ns=%llu sleeping_ns=%llu",
+ __entry->id, __entry->threshold_ns,
+ __entry->running_ns, __entry->waiting_ns, __entry->sleeping_ns)
+);
+#endif /* CONFIG_RV_MON_TLOB */
diff --git a/kernel/trace/rv/rv_trace.h b/kernel/trace/rv/rv_trace.h
index 2f8a932432c9..4bfa39717cef 100644
--- a/kernel/trace/rv/rv_trace.h
+++ b/kernel/trace/rv/rv_trace.h
@@ -189,6 +189,7 @@ DECLARE_EVENT_CLASS(error_env_da_monitor_id,
#include <monitors/stall/stall_trace.h>
#include <monitors/nomiss/nomiss_trace.h>
+#include <monitors/tlob/tlob_trace.h>
// Add new monitors based on CONFIG_HA_MON_EVENTS_ID here
#endif
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 7/9] rv: Add KUnit tests for the tlob monitor
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (5 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 6/9] rv: Add tlob hybrid automaton monitor wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 8/9] selftests/verification: Add tlob selftests wen.yang
2026-08-19 18:15 ` [PATCH v5 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add CONFIG_TLOB_KUNIT_TEST (tristate, depends on RV_MON_TLOB && KUNIT,
default KUNIT_ALL_TESTS) with a test suite covering the uprobe-line
parser.
Tests call tlob_parse_uprobe_line() and tlob_parse_remove_line()
directly rather than going through the top-level write handler.
Mark both functions VISIBLE_IF_KUNIT and export with
EXPORT_SYMBOL_IF_KUNIT. Update the IS_ENABLED guard in tlob.h to use
CONFIG_TLOB_KUNIT_TEST instead of CONFIG_KUNIT.
Reviewed-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
kernel/trace/rv/Makefile | 1 +
kernel/trace/rv/monitors/tlob/.kunitconfig | 5 +
kernel/trace/rv/monitors/tlob/Kconfig | 10 ++
kernel/trace/rv/monitors/tlob/tlob.c | 6 +-
kernel/trace/rv/monitors/tlob/tlob.h | 6 +
kernel/trace/rv/monitors/tlob/tlob_kunit.c | 139 +++++++++++++++++++++
6 files changed, 165 insertions(+), 2 deletions(-)
create mode 100644 kernel/trace/rv/monitors/tlob/.kunitconfig
create mode 100644 kernel/trace/rv/monitors/tlob/tlob_kunit.c
diff --git a/kernel/trace/rv/Makefile b/kernel/trace/rv/Makefile
index cd0ec11f0e05..92a04367ce93 100644
--- a/kernel/trace/rv/Makefile
+++ b/kernel/trace/rv/Makefile
@@ -22,6 +22,7 @@ obj-$(CONFIG_RV_MON_DEADLINE) += monitors/deadline/deadline.o
obj-$(CONFIG_RV_MON_NOMISS) += monitors/nomiss/nomiss.o
obj-$(CONFIG_RV_MON_WAKEUP) += monitors/wakeup/wakeup.o
obj-$(CONFIG_RV_MON_TLOB) += monitors/tlob/tlob.o
+obj-$(CONFIG_TLOB_KUNIT_TEST) += monitors/tlob/tlob_kunit.o
# Add new monitors here
obj-$(CONFIG_RV_UPROBE) += rv_uprobe.o
obj-$(CONFIG_RV_REACTORS) += rv_reactors.o
diff --git a/kernel/trace/rv/monitors/tlob/.kunitconfig b/kernel/trace/rv/monitors/tlob/.kunitconfig
new file mode 100644
index 000000000000..82a6f121016e
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/.kunitconfig
@@ -0,0 +1,5 @@
+CONFIG_KUNIT=y
+CONFIG_UPROBES=y
+CONFIG_RV=y
+CONFIG_RV_MON_TLOB=y
+CONFIG_TLOB_KUNIT_TEST=y
diff --git a/kernel/trace/rv/monitors/tlob/Kconfig b/kernel/trace/rv/monitors/tlob/Kconfig
index aa43382073d2..f4fce35d38fb 100644
--- a/kernel/trace/rv/monitors/tlob/Kconfig
+++ b/kernel/trace/rv/monitors/tlob/Kconfig
@@ -10,3 +10,13 @@ config RV_MON_TLOB
monitor. tlob tracks per-task elapsed wall-clock time across a
user-delimited code section and emits error_env_tlob when the
elapsed time exceeds a configurable per-invocation budget.
+
+config TLOB_KUNIT_TEST
+ tristate "KUnit tests for tlob monitor" if !KUNIT_ALL_TESTS
+ depends on RV_MON_TLOB && KUNIT
+ default KUNIT_ALL_TESTS
+ help
+ Enable KUnit unit tests for the tlob RV monitor. The tests
+ cover the uprobe-line parser (tlob_parse_uprobe_line) and the
+ remove-line parser (tlob_parse_remove_line), verifying valid,
+ invalid, and out-of-range inputs without requiring a running kernel.
diff --git a/kernel/trace/rv/monitors/tlob/tlob.c b/kernel/trace/rv/monitors/tlob/tlob.c
index 99acd34726f1..e109390ba3ad 100644
--- a/kernel/trace/rv/monitors/tlob/tlob.c
+++ b/kernel/trace/rv/monitors/tlob/tlob.c
@@ -874,7 +874,7 @@ static ssize_t tlob_monitor_read(struct file *file,
* PATH may contain ':'; the last ':' separates path from offset.
* Returns 0, -EINVAL, or -ERANGE.
*/
-static int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
+VISIBLE_IF_KUNIT int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
char **path_out,
loff_t *start_out, loff_t *stop_out)
{
@@ -948,11 +948,12 @@ static int tlob_parse_uprobe_line(char *buf, u64 *thr_out,
*stop_out = (loff_t)stop_val;
return 0;
}
+EXPORT_SYMBOL_IF_KUNIT(tlob_parse_uprobe_line);
/*
* Parse "-PATH:OFFSET_START" (ftrace uprobe_events removal convention).
*/
-static int tlob_parse_remove_line(char *buf, char **path_out,
+VISIBLE_IF_KUNIT int tlob_parse_remove_line(char *buf, char **path_out,
loff_t *start_out)
{
char *binpath, *colon;
@@ -976,6 +977,7 @@ static int tlob_parse_remove_line(char *buf, char **path_out,
*start_out = (loff_t)off;
return 0;
}
+EXPORT_SYMBOL_IF_KUNIT(tlob_parse_remove_line);
static int tlob_create_or_delete_uprobe(char *buf)
{
diff --git a/kernel/trace/rv/monitors/tlob/tlob.h b/kernel/trace/rv/monitors/tlob/tlob.h
index 94e7382c2130..6ad9d5179ab6 100644
--- a/kernel/trace/rv/monitors/tlob/tlob.h
+++ b/kernel/trace/rv/monitors/tlob/tlob.h
@@ -146,4 +146,10 @@ static const struct automaton_tlob automaton_tlob = {
/* Upper budget bound (1 hour): keeps the u64 ns accumulators far from overflow. */
#define TLOB_MAX_THRESHOLD_NS 3600000000000ULL
+#if IS_ENABLED(CONFIG_TLOB_KUNIT_TEST)
+int tlob_parse_uprobe_line(char *buf, u64 *thr_out, char **path_out,
+ loff_t *start_out, loff_t *stop_out);
+int tlob_parse_remove_line(char *buf, char **path_out, loff_t *start_out);
+#endif /* CONFIG_TLOB_KUNIT_TEST */
+
#endif /* _RV_TLOB_H */
diff --git a/kernel/trace/rv/monitors/tlob/tlob_kunit.c b/kernel/trace/rv/monitors/tlob/tlob_kunit.c
new file mode 100644
index 000000000000..6a6fb57d0678
--- /dev/null
+++ b/kernel/trace/rv/monitors/tlob/tlob_kunit.c
@@ -0,0 +1,139 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * KUnit tests for the tlob RV monitor.
+ *
+ */
+#include <kunit/test.h>
+
+#include "tlob.h"
+
+MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING");
+
+/* Valid "p PATH:START STOP threshold=NS" lines. */
+static const char * const tlob_parse_valid[] = {
+ "p /usr/bin/myapp:4768 4848 threshold=5000000",
+ "p /usr/bin/myapp:0x12a0 0x12f0 threshold=10000000",
+ "p /opt/my:app/bin:0x100 0x200 threshold=1000000",
+};
+
+/* Malformed "p ..." lines that must be rejected with -EINVAL. */
+static const char * const tlob_parse_invalid[] = {
+ "p :0x100 0x200 threshold=5000",
+ "p /usr/bin/myapp:0x100 threshold=5000",
+ "p /usr/bin/myapp:-1 0x200 threshold=5000",
+ "p /usr/bin/myapp:0x100 -1 threshold=5000000", /* negative stop offset */
+ "p /usr/bin/myapp:0x100 0x200",
+ "p /usr/bin/myapp:0x100 0x100 threshold=5000",
+};
+
+/* threshold_ns < TLOB_MIN_THRESHOLD_NS or > TLOB_MAX_THRESHOLD_NS => -ERANGE. */
+static const char * const tlob_parse_out_of_range[] = {
+ "p /usr/bin/myapp:0x100 0x200 threshold=0",
+ "p /usr/bin/myapp:0x100 0x200 threshold=999",
+ "p /usr/bin/myapp:0x100 0x200 threshold=3600000000001",
+};
+
+/* Valid "-PATH:OFFSET_START" remove lines. */
+static const char * const tlob_remove_valid[] = {
+ "-/usr/bin/myapp:0x100",
+ "-/opt/my:app/bin:0x200",
+};
+
+/* Malformed remove lines that must be rejected with -EINVAL. */
+static const char * const tlob_remove_invalid[] = {
+ "-usr/bin/myapp:0x100",
+ "-/usr/bin/myapp",
+ "-/:0x100",
+ "-/usr/bin/myapp:-1", /* negative offset */
+ "-/usr/bin/myapp:abc",
+};
+
+static void tlob_parse_valid_accepted(struct kunit *test)
+{
+ u64 thr;
+ char *path;
+ loff_t start, stop;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_valid); i++) {
+ strscpy(buf, tlob_parse_valid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, tlob_parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), 0);
+ }
+}
+
+static void tlob_parse_invalid_rejected(struct kunit *test)
+{
+ u64 thr;
+ char *path;
+ loff_t start, stop;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_invalid); i++) {
+ strscpy(buf, tlob_parse_invalid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, tlob_parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), -EINVAL);
+ }
+}
+
+static void tlob_parse_out_of_range_rejected(struct kunit *test)
+{
+ u64 thr;
+ char *path;
+ loff_t start, stop;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_parse_out_of_range); i++) {
+ strscpy(buf, tlob_parse_out_of_range[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, tlob_parse_uprobe_line(buf, &thr, &path,
+ &start, &stop), -ERANGE);
+ }
+}
+
+static void tlob_remove_valid_accepted(struct kunit *test)
+{
+ char *path;
+ loff_t start;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_remove_valid); i++) {
+ strscpy(buf, tlob_remove_valid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, tlob_parse_remove_line(buf, &path, &start), 0);
+ }
+}
+
+static void tlob_remove_invalid_rejected(struct kunit *test)
+{
+ char *path;
+ loff_t start;
+ char buf[128];
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(tlob_remove_invalid); i++) {
+ strscpy(buf, tlob_remove_invalid[i], sizeof(buf));
+ KUNIT_EXPECT_EQ(test, tlob_parse_remove_line(buf, &path, &start), -EINVAL);
+ }
+}
+
+static struct kunit_case tlob_parse_cases[] = {
+ KUNIT_CASE(tlob_parse_valid_accepted),
+ KUNIT_CASE(tlob_parse_invalid_rejected),
+ KUNIT_CASE(tlob_parse_out_of_range_rejected),
+ KUNIT_CASE(tlob_remove_valid_accepted),
+ KUNIT_CASE(tlob_remove_invalid_rejected),
+ {}
+};
+
+static struct kunit_suite tlob_parse_suite = {
+ .name = "tlob_parse",
+ .test_cases = tlob_parse_cases,
+};
+
+kunit_test_suite(tlob_parse_suite);
+
+MODULE_DESCRIPTION("KUnit tests for the tlob RV monitor");
+MODULE_LICENSE("GPL");
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 8/9] selftests/verification: Add tlob selftests
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (6 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
@ 2026-08-19 18:15 ` wen.yang
2026-08-19 18:15 ` [PATCH v5 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
Add seven ftrace-style test scripts for the tlob RV monitor under
tools/testing/selftests/verification/test.d/tlob/. The tests cover
uprobe binding management, budget violation detection, and per-state
time accounting.
Helper binaries tlob_sym and tlob_target are placed in
tools/testing/selftests/verification/ and built via TEST_GEN_FILES in
the top-level Makefile, following the standard kselftests convention.
VERIFICATIONTEST_BINDIR is exported after include ../lib.mk so OUTPUT
is already populated by lib.mk when the variable is assigned.
run_tlob_tests.sh is a thin wrapper: it builds the helpers and delegates
all argument handling to ftracetest via exec "$FTRACETEST" -K --rv "$@".
All .tc files that start background processes set up a trap teardown EXIT
immediately after launch so background tasks are cleaned up even when a
subsequent assertion fails under set -e.
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
.../testing/selftests/verification/.gitignore | 2 +
tools/testing/selftests/verification/Makefile | 4 +
.../test.d/tlob/run_tlob_tests.sh | 20 ++
.../verification/test.d/tlob/uprobe_bind.tc | 35 +++
.../test.d/tlob/uprobe_detail_running.tc | 49 ++++
.../test.d/tlob/uprobe_detail_sleeping.tc | 48 ++++
.../test.d/tlob/uprobe_detail_waiting.tc | 68 ++++++
.../verification/test.d/tlob/uprobe_multi.tc | 59 +++++
.../test.d/tlob/uprobe_no_event.tc | 17 ++
.../test.d/tlob/uprobe_restart.tc | 75 +++++++
.../test.d/tlob/uprobe_violation.tc | 65 ++++++
.../testing/selftests/verification/tlob_sym.c | 209 ++++++++++++++++++
.../selftests/verification/tlob_target.c | 138 ++++++++++++
13 files changed, 789 insertions(+)
create mode 100755 tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
create mode 100644 tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
create mode 100644 tools/testing/selftests/verification/tlob_sym.c
create mode 100644 tools/testing/selftests/verification/tlob_target.c
diff --git a/tools/testing/selftests/verification/.gitignore b/tools/testing/selftests/verification/.gitignore
index 2659417cb2c7..d2f231f1bacb 100644
--- a/tools/testing/selftests/verification/.gitignore
+++ b/tools/testing/selftests/verification/.gitignore
@@ -1,2 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
logs
+tlob_sym
+tlob_target
diff --git a/tools/testing/selftests/verification/Makefile b/tools/testing/selftests/verification/Makefile
index aa8790c22a71..41445d15b86a 100644
--- a/tools/testing/selftests/verification/Makefile
+++ b/tools/testing/selftests/verification/Makefile
@@ -5,4 +5,8 @@ TEST_PROGS := verificationtest-ktap
TEST_FILES := test.d settings
EXTRA_CLEAN := $(OUTPUT)/logs/*
+TEST_GEN_FILES := tlob_sym tlob_target
+
include ../lib.mk
+
+export VERIFICATIONTEST_BINDIR := $(OUTPUT)
diff --git a/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh b/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
new file mode 100755
index 000000000000..13adf1eec9a2
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
@@ -0,0 +1,20 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# Standalone runner for tlob selftests
+
+set -e
+
+SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
+FTRACETEST="$SCRIPT_DIR/../../../ftrace/ftracetest"
+
+# Build test helpers
+echo "Building tlob test helpers..."
+make -C "$SCRIPT_DIR/../.." all
+
+# Export VERIFICATIONTEST_BINDIR so test scripts can find tlob_target and
+# tlob_sym (built in the verification directory)
+export VERIFICATIONTEST_BINDIR="$(realpath "$SCRIPT_DIR/../..")"
+
+# Forward all options to ftracetest
+exec "$FTRACETEST" -K --rv "$SCRIPT_DIR" "$@"
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
new file mode 100644
index 000000000000..baa6c5fa0ff2
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_bind.tc
@@ -0,0 +1,35 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor uprobe binding (visible in monitor file, removable, duplicate rejected)
+# requires: tlob:monitor
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > monitors/tlob/enable
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=5000000000" > "$TLOB_MONITOR"
+
+# Binding must appear in monitor file with canonical hex-offset format.
+grep -qE "^p ${UPROBE_TARGET}:0x[0-9a-f]+ 0x[0-9a-f]+ threshold=[0-9]+$" "$TLOB_MONITOR"
+grep -q "threshold=5000000000" "$TLOB_MONITOR"
+
+# Duplicate offset_start must be rejected.
+! echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=9999000" > "$TLOB_MONITOR" 2>/dev/null || false
+
+# Remove the binding; it must no longer appear.
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR"
+! grep -q "^p .*:0x${busy_offset#0x} " "$TLOB_MONITOR" || false
+
+echo 0 > monitors/tlob/enable
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
new file mode 100644
index 000000000000..f9a412930cda
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_running.tc
@@ -0,0 +1,49 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail running (running_ns dominates when task busy-spins between probes)
+# requires: tlob:monitor
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 5000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# 10 us budget; task busy-spins 200 ms per iteration -> running_ns dominates.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=10000" > "$TLOB_MONITOR"
+
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" /sys/kernel/tracing/trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" /sys/kernel/tracing/trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+# Busy-spin keeps the task on-CPU: running_ns must exceed sleeping_ns.
+[ "$running" -gt "$sleeping" ]
+
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
new file mode 100644
index 000000000000..9b9e1c98d7e1
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_sleeping.tc
@@ -0,0 +1,48 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail sleeping (sleeping_ns dominates when task blocks between probes)
+# requires: tlob:monitor
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 5000 sleep &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# 50 ms budget; task sleeps 200 ms per iteration -> sleeping_ns dominates.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=50000000" > "$TLOB_MONITOR"
+
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" /sys/kernel/tracing/trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" /sys/kernel/tracing/trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+[ "$sleeping" -gt "$((running + waiting))" ]
+
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
new file mode 100644
index 000000000000..68aeb1fa2172
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_detail_waiting.tc
@@ -0,0 +1,68 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor detail waiting (waiting_ns dominates when task is preempted between probes)
+# requires: tlob:monitor chrt:program taskset:program
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+start_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_preempt_work_done 2>/dev/null)
+
+# Pick the last CPU to avoid cpu0 which is used by vng infrastructure.
+cpu=$(($(nproc) - 1))
+
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# tlob_target loops calling tlob_preempt_work(200) / tlob_preempt_work_done()
+# in 200 ms wall-clock iterations. The stop probe fires when
+# tlob_preempt_work_done() is called, which cancels the budget timer.
+# Budget must be less than 200 ms so the HA timer fires while the target is
+# still inside tlob_preempt_work() and before the stop probe fires.
+# 150 ms gives a comfortable margin: waiting_ns ≈ 140 ms >> running_ns < 10 ms.
+echo "p ${UPROBE_TARGET}:${start_offset} ${stop_offset} threshold=150000000" > "$TLOB_MONITOR"
+
+# Start the RT hog BEFORE the target so the target is immediately preempted
+# when it calls tlob_preempt_work() (start probe fires), minimising running_ns.
+chrt -f 99 taskset -c "$cpu" sh -c 'while true; do :; done' 2>/dev/null &
+hog_pid=$!
+teardown() {
+ kill "$hog_pid" 2>/dev/null || true; wait "$hog_pid" 2>/dev/null || true
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.02
+
+taskset -c "$cpu" "$UPROBE_TARGET" 5000 preempt &
+busy_pid=$!
+
+# Poll up to 3 s (budget 150 ms + generous margin).
+found=0; i=0
+while [ "$i" -lt 30 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" /sys/kernel/tracing/trace && { found=1; break; }
+ i=$((i+1))
+done
+
+# Kill the RT hog first so tlob_target can release any in-flight SRCU read
+# section from uprobe_notify_resume; otherwise probe removal blocks in
+# synchronize_srcu with the hog monopolising the CPU at FIFO-99.
+kill "$hog_pid" 2>/dev/null || true; wait "$hog_pid" 2>/dev/null || true
+kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+echo "-${UPROBE_TARGET}:${start_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+line=$(grep "detail_env_tlob" /sys/kernel/tracing/trace | head -n 1)
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+waiting=$(echo "$line" | sed 's/.*waiting_ns=\([0-9]*\).*/\1/')
+[ "$waiting" -gt "$((running + sleeping))" ]
+
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
new file mode 100644
index 000000000000..d5bfa4e3f695
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_multi.tc
@@ -0,0 +1,59 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor multiple uprobe bindings (different offsets fire independently)
+# requires: tlob:monitor
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+busy_stop=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+sleep_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work 2>/dev/null)
+sleep_stop=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_sleep_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 & # busy mode: tlob_busy_work fires every 200 ms
+busy_pid=$!
+"$UPROBE_TARGET" 30000 sleep & # sleep mode: tlob_sleep_work fires every 200 ms
+sleep_pid=$!
+teardown() {
+ kill "$sleep_pid" 2>/dev/null || true; wait "$sleep_pid" 2>/dev/null || true
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# Binding A: 5 s budget on the busy probe - must not fire in 200 ms loops.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${busy_stop} threshold=5000000000" > "$TLOB_MONITOR"
+# Binding B: 10 us budget on the sleep probe - fires on first invocation.
+echo "p ${UPROBE_TARGET}:${sleep_offset} ${sleep_stop} threshold=10000" > "$TLOB_MONITOR"
+
+# Wait up to 2 s for error_env_tlob from binding B.
+found=0; i=0
+while [ "$i" -lt 20 ]; do
+ sleep 0.1
+ grep -q "error_env_tlob" /sys/kernel/tracing/trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo "-${UPROBE_TARGET}:${sleep_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > monitors/tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+
+[ "$found" = "1" ]
+# error_env_tlob payload: clock variable must be present.
+# The event field can be "budget_exceeded" (hrtimer path) or the DA event
+# name ("sleep", "preempt") depending on which fires first; don't constrain it.
+grep "error_env_tlob" /sys/kernel/tracing/trace | head -n 1 | grep -q "clk_elapsed="
+# detail_env_tlob must appear alongside the error.
+grep -q "detail_env_tlob" /sys/kernel/tracing/trace
+
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
new file mode 100644
index 000000000000..30e78f2b1475
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_no_event.tc
@@ -0,0 +1,17 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor no spurious events without active uprobe binding
+# requires: tlob:monitor
+
+echo 1 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+sleep 0.5
+
+! grep -q "error_env_tlob" /sys/kernel/tracing/trace || false
+
+echo 0 > monitors/tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
new file mode 100644
index 000000000000..de7229744969
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_restart.tc
@@ -0,0 +1,75 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor restarting a parked window on the same task (no cross-window accumulator leak, unbind of a repeatedly-started task works)
+# requires: tlob:monitor
+
+set -x
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+# Diagnostic dump for a budget-violation regression: prints the exact
+# error_env_tlob entry (state + clk_elapsed at expiry), the per-state
+# accumulator breakdown, and the full start/stop transition sequence for
+# the monitored pid, so a false-positive violation can be told apart from
+# a genuine per-window overrun (and, if it is the latter, whether the
+# accumulators leaked across a restart).
+dump_tlob_diag() {
+ TRACE=/sys/kernel/tracing/trace
+ echo "===== tlob restart diagnostic (pid ${busy_pid}) =====" >&2
+ echo "--- error_env_tlob (violations) ---" >&2
+ grep "error_env_tlob" "$TRACE" >&2 || true
+ echo "--- detail_env_tlob (accumulator breakdown) ---" >&2
+ grep "detail_env_tlob" "$TRACE" >&2 || true
+ echo "--- event_tlob (state transitions for target pid) ---" >&2
+ grep "event_tlob" "$TRACE" | grep ":${busy_pid}:" >&2 || true
+ echo "--- last 50 trace lines (full context) ---" >&2
+ tail -50 "$TRACE" >&2 || true
+}
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+# tlob_target loops calling tlob_busy_work(200)/tlob_busy_work_done() every
+# ~200ms for the whole run: each call is one start/stop window on the SAME
+# task, driving tlob_start_task()'s restart path (same pid, parked ws) on
+# every iteration after the first. 900ms gives ~4 such windows.
+"$UPROBE_TARGET" 900 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > /sys/kernel/tracing/events/rv/event_tlob/enable
+echo 1 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# 300ms budget: comfortably covers one ~200ms window. If a restart failed
+# to reset ws->accs_ns[]/budget_exceeded (a regression this test exists to
+# catch), running_ns would keep growing across windows and the SECOND
+# window would already exceed budget (~400ms cumulative > 300ms). With a
+# correct reset, no window ever exceeds it.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=300000000" > "$TLOB_MONITOR"
+
+wait "$busy_pid" || true
+
+if grep -q "error_env_tlob" /sys/kernel/tracing/trace; then
+ dump_tlob_diag
+ false
+fi
+
+# The task has exited (tlob_destroy_task() ran via handle_sched_process_exit);
+# removing its now-stale binding must still succeed cleanly.
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR"
+! grep -q "^p .*:0x${busy_offset#0x} " "$TLOB_MONITOR" || false
+
+echo 0 > monitors/tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/event_tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc b/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
new file mode 100644
index 000000000000..772f79b0d609
--- /dev/null
+++ b/tools/testing/selftests/verification/test.d/tlob/uprobe_violation.tc
@@ -0,0 +1,65 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0-or-later
+# description: Test tlob monitor budget violation (error_env_tlob and detail_env_tlob fire with correct fields)
+# requires: tlob:monitor
+
+UPROBE_TARGET="${VERIFICATIONTEST_BINDIR}/tlob_target"
+TLOB_SYM="${VERIFICATIONTEST_BINDIR}/tlob_sym"
+TLOB_MONITOR=monitors/tlob/monitor
+
+busy_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work 2>/dev/null)
+stop_offset=$("$TLOB_SYM" sym_offset "$UPROBE_TARGET" tlob_busy_work_done 2>/dev/null)
+
+"$UPROBE_TARGET" 30000 &
+busy_pid=$!
+teardown() {
+ kill "$busy_pid" 2>/dev/null || true; wait "$busy_pid" 2>/dev/null || true
+}
+trap teardown EXIT
+sleep 0.05
+
+echo 1 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 1 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 1 > /sys/kernel/tracing/tracing_on
+echo 1 > monitors/tlob/enable
+echo > /sys/kernel/tracing/trace
+
+# 10 us budget - fires almost immediately; task is busy-spinning on-CPU.
+echo "p ${UPROBE_TARGET}:${busy_offset} ${stop_offset} threshold=10000" > "$TLOB_MONITOR"
+
+# wait up to 2 s for detail_env_tlob
+found=0; i=0
+while [ "$i" -lt 20 ]; do
+ sleep 0.1
+ grep -q "detail_env_tlob" /sys/kernel/tracing/trace && { found=1; break; }
+ i=$((i+1))
+done
+
+echo "-${UPROBE_TARGET}:${busy_offset}" > "$TLOB_MONITOR" 2>/dev/null
+echo 0 > /sys/kernel/tracing/events/rv/error_env_tlob/enable
+echo 0 > /sys/kernel/tracing/events/rv/detail_env_tlob/enable
+echo 0 > monitors/tlob/enable
+
+[ "$found" = "1" ]
+
+# error_env_tlob must carry the clk_elapsed environment field.
+# The event label is "budget_exceeded" when detected by the hrtimer callback,
+# or the triggering sched event name when detected by the constraint path on a
+# preemption that races with the timer (common on PREEMPT_RT / VM). Both are
+# valid detections; check the env field instead of the label.
+grep "error_env_tlob" /sys/kernel/tracing/trace | head -n 1 | grep -q "clk_elapsed="
+
+# detail_env_tlob must have all five fields with the correct threshold
+line=$(grep "detail_env_tlob" /sys/kernel/tracing/trace | head -n 1)
+echo "$line" | grep -q "pid="
+echo "$line" | grep -q "threshold_ns=10000"
+echo "$line" | grep -q "running_ns="
+echo "$line" | grep -q "waiting_ns="
+echo "$line" | grep -q "sleeping_ns="
+
+# Busy-spin keeps the task on-CPU: running_ns must exceed sleeping_ns.
+running=$(echo "$line" | sed 's/.*running_ns=\([0-9]*\).*/\1/')
+sleeping=$(echo "$line" | sed 's/.*sleeping_ns=\([0-9]*\).*/\1/')
+[ "$running" -gt "$sleeping" ]
+
+echo > /sys/kernel/tracing/trace
diff --git a/tools/testing/selftests/verification/tlob_sym.c b/tools/testing/selftests/verification/tlob_sym.c
new file mode 100644
index 000000000000..a92fc49d1304
--- /dev/null
+++ b/tools/testing/selftests/verification/tlob_sym.c
@@ -0,0 +1,209 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob_sym.c - ELF symbol-to-file-offset utility for tlob selftests
+ *
+ * Usage: tlob_sym sym_offset <binary> <symbol>
+ *
+ * Prints the ELF file offset of <symbol> in <binary> to stdout.
+ *
+ * Exit: 0 = found, 1 = error / not found.
+ */
+#include <elf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+static int sym_offset(const char *binary, const char *symname)
+{
+ int fd;
+ struct stat st;
+ void *map;
+ Elf64_Ehdr *ehdr;
+ Elf32_Ehdr *ehdr32;
+ int is64;
+ uint64_t sym_vaddr = 0;
+ int found = 0;
+ uint64_t file_offset = 0;
+
+ fd = open(binary, O_RDONLY);
+ if (fd < 0) {
+ fprintf(stderr, "open %s: %s\n", binary, strerror(errno));
+ return 1;
+ }
+ if (fstat(fd, &st) < 0) {
+ close(fd);
+ return 1;
+ }
+ map = mmap(NULL, (size_t)st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);
+ close(fd);
+ if (map == MAP_FAILED) {
+ fprintf(stderr, "mmap: %s\n", strerror(errno));
+ return 1;
+ }
+
+ ehdr = (Elf64_Ehdr *)map;
+ ehdr32 = (Elf32_Ehdr *)map;
+ if (st.st_size < 4 ||
+ ehdr->e_ident[EI_MAG0] != ELFMAG0 ||
+ ehdr->e_ident[EI_MAG1] != ELFMAG1 ||
+ ehdr->e_ident[EI_MAG2] != ELFMAG2 ||
+ ehdr->e_ident[EI_MAG3] != ELFMAG3) {
+ fprintf(stderr, "%s: not an ELF file\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ is64 = (ehdr->e_ident[EI_CLASS] == ELFCLASS64);
+
+ if (is64) {
+ Elf64_Shdr *shdrs;
+ Elf64_Shdr *shstrtab_hdr;
+
+ if (ehdr->e_shnum == 0 || ehdr->e_shstrndx >= ehdr->e_shnum ||
+ (uint64_t)ehdr->e_shoff +
+ (uint64_t)ehdr->e_shnum * sizeof(Elf64_Shdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF section table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ shdrs = (Elf64_Shdr *)((char *)map + ehdr->e_shoff);
+ shstrtab_hdr = &shdrs[ehdr->e_shstrndx];
+ const char *shstrtab = (char *)map + shstrtab_hdr->sh_offset;
+ int si;
+
+ for (int pass = 0; pass < 2 && !found; pass++) {
+ const char *target = pass ? ".dynsym" : ".symtab";
+
+ for (si = 0; si < ehdr->e_shnum && !found; si++) {
+ Elf64_Shdr *sh = &shdrs[si];
+ const char *name = shstrtab + sh->sh_name;
+
+ if (strcmp(name, target) != 0)
+ continue;
+
+ Elf64_Shdr *strtab_sh = &shdrs[sh->sh_link];
+ const char *strtab = (char *)map + strtab_sh->sh_offset;
+ Elf64_Sym *syms = (Elf64_Sym *)((char *)map + sh->sh_offset);
+ uint64_t nsyms = sh->sh_size / sizeof(Elf64_Sym);
+ uint64_t j;
+
+ for (j = 0; j < nsyms; j++) {
+ if (strcmp(strtab + syms[j].st_name, symname) == 0) {
+ sym_vaddr = syms[j].st_value;
+ found = 1;
+ break;
+ }
+ }
+ }
+ }
+
+ if (!found) {
+ fprintf(stderr, "symbol '%s' not found in %s\n", symname, binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ Elf64_Phdr *phdrs = (Elf64_Phdr *)((char *)map + ehdr->e_phoff);
+ int pi;
+
+ for (pi = 0; pi < ehdr->e_phnum; pi++) {
+ Elf64_Phdr *ph = &phdrs[pi];
+
+ if (ph->p_type != PT_LOAD)
+ continue;
+ if (sym_vaddr >= ph->p_vaddr &&
+ sym_vaddr < ph->p_vaddr + ph->p_filesz) {
+ file_offset = sym_vaddr - ph->p_vaddr + ph->p_offset;
+ break;
+ }
+ }
+ } else {
+ Elf32_Shdr *shdrs;
+ Elf32_Shdr *shstrtab_hdr;
+
+ if (ehdr32->e_shnum == 0 || ehdr32->e_shstrndx >= ehdr32->e_shnum ||
+ (uint64_t)ehdr32->e_shoff +
+ (uint64_t)ehdr32->e_shnum * sizeof(Elf32_Shdr) > (uint64_t)st.st_size) {
+ fprintf(stderr, "%s: malformed ELF section table\n", binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+ shdrs = (Elf32_Shdr *)((char *)map + ehdr32->e_shoff);
+ shstrtab_hdr = &shdrs[ehdr32->e_shstrndx];
+ const char *shstrtab = (char *)map + shstrtab_hdr->sh_offset;
+ int si;
+ uint32_t sym_vaddr32 = 0;
+
+ for (int pass = 0; pass < 2 && !found; pass++) {
+ const char *target = pass ? ".dynsym" : ".symtab";
+
+ for (si = 0; si < ehdr32->e_shnum && !found; si++) {
+ Elf32_Shdr *sh = &shdrs[si];
+ const char *name = shstrtab + sh->sh_name;
+
+ if (strcmp(name, target) != 0)
+ continue;
+
+ Elf32_Shdr *strtab_sh = &shdrs[sh->sh_link];
+ const char *strtab = (char *)map + strtab_sh->sh_offset;
+ Elf32_Sym *syms = (Elf32_Sym *)((char *)map + sh->sh_offset);
+ uint32_t nsyms = sh->sh_size / sizeof(Elf32_Sym);
+ uint32_t j;
+
+ for (j = 0; j < nsyms; j++) {
+ if (strcmp(strtab + syms[j].st_name, symname) == 0) {
+ sym_vaddr32 = syms[j].st_value;
+ found = 1;
+ break;
+ }
+ }
+ }
+ }
+
+ if (!found) {
+ fprintf(stderr, "symbol '%s' not found in %s\n", symname, binary);
+ munmap(map, (size_t)st.st_size);
+ return 1;
+ }
+
+ Elf32_Phdr *phdrs = (Elf32_Phdr *)((char *)map + ehdr32->e_phoff);
+ int pi;
+
+ for (pi = 0; pi < ehdr32->e_phnum; pi++) {
+ Elf32_Phdr *ph = &phdrs[pi];
+
+ if (ph->p_type != PT_LOAD)
+ continue;
+ if (sym_vaddr32 >= ph->p_vaddr &&
+ sym_vaddr32 < ph->p_vaddr + ph->p_filesz) {
+ file_offset = sym_vaddr32 - ph->p_vaddr + ph->p_offset;
+ break;
+ }
+ }
+ sym_vaddr = sym_vaddr32;
+ }
+
+ munmap(map, (size_t)st.st_size);
+
+ if (!file_offset && sym_vaddr) {
+ fprintf(stderr, "could not map vaddr 0x%lx to file offset\n",
+ (unsigned long)sym_vaddr);
+ return 1;
+ }
+
+ printf("0x%lx\n", (unsigned long)file_offset);
+ return 0;
+}
+
+int main(int argc, char *argv[])
+{
+ if (argc != 4 || strcmp(argv[1], "sym_offset") != 0) {
+ fprintf(stderr, "Usage: %s sym_offset <binary> <symbol>\n", argv[0]);
+ return 1;
+ }
+ return sym_offset(argv[2], argv[3]);
+}
diff --git a/tools/testing/selftests/verification/tlob_target.c b/tools/testing/selftests/verification/tlob_target.c
new file mode 100644
index 000000000000..adf4c2397fb3
--- /dev/null
+++ b/tools/testing/selftests/verification/tlob_target.c
@@ -0,0 +1,138 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * tlob_target.c - uprobe target binary for tlob selftests.
+ *
+ * Provides three start/stop probe pairs, each designed to exercise a
+ * different dominant component of the detail_env_tlob ns breakdown:
+ *
+ * tlob_busy_work / tlob_busy_work_done - busy-spin: running_ns dominates
+ * tlob_sleep_work / tlob_sleep_work_done - nanosleep: sleeping_ns dominates
+ * tlob_preempt_work / tlob_preempt_work_done - busy-spin + RT competitor:
+ * waiting_ns dominates
+ *
+ * Usage: tlob_target <duration_ms> [mode]
+ *
+ * mode is one of: busy (default), sleep, preempt.
+ * Loops in 200 ms iterations until <duration_ms> has elapsed
+ * (0 = run for ~24 hours).
+ */
+#define _GNU_SOURCE
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <time.h>
+
+#ifndef noinline
+#define noinline __attribute__((noinline))
+#endif
+
+static inline int timespec_before(const struct timespec *a,
+ const struct timespec *b)
+{
+ return a->tv_sec < b->tv_sec ||
+ (a->tv_sec == b->tv_sec && a->tv_nsec < b->tv_nsec);
+}
+
+static void timespec_add_ms(struct timespec *ts, unsigned long ms)
+{
+ ts->tv_sec += ms / 1000;
+ ts->tv_nsec += (long)(ms % 1000) * 1000000L;
+ if (ts->tv_nsec >= 1000000000L) {
+ ts->tv_sec++;
+ ts->tv_nsec -= 1000000000L;
+ }
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_busy_work_done(void)
+{
+ /* empty: uprobe fires on entry */
+}
+
+/* start probe; busy-spin so running_ns dominates */
+noinline void tlob_busy_work(unsigned long duration_ms)
+{
+ struct timespec start, now;
+ unsigned long elapsed;
+
+ clock_gettime(CLOCK_MONOTONIC, &start);
+ do {
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ elapsed = (unsigned long)(now.tv_sec - start.tv_sec)
+ * 1000000000UL
+ + (unsigned long)(now.tv_nsec - start.tv_nsec);
+ } while (elapsed < duration_ms * 1000000UL);
+
+ tlob_busy_work_done();
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_sleep_work_done(void)
+{
+ /* empty: uprobe fires on entry */
+}
+
+/* start probe; nanosleep so sleeping_ns dominates */
+noinline void tlob_sleep_work(unsigned long duration_ms)
+{
+ struct timespec ts = {
+ .tv_sec = duration_ms / 1000,
+ .tv_nsec = (long)(duration_ms % 1000) * 1000000L,
+ };
+ nanosleep(&ts, NULL);
+ tlob_sleep_work_done();
+}
+
+/* stop probe; noinline keeps the entry point visible to uprobes */
+noinline void tlob_preempt_work_done(void)
+{
+ /* empty: uprobe fires on entry */
+}
+
+/*
+ * start probe; busy-spin so an RT competitor on the same CPU drives
+ * waiting_ns (prev_state==0 -> preempt event, task stays runnable off-CPU).
+ */
+noinline void tlob_preempt_work(unsigned long duration_ms)
+{
+ struct timespec start, now;
+ unsigned long elapsed;
+
+ clock_gettime(CLOCK_MONOTONIC, &start);
+ do {
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ elapsed = (unsigned long)(now.tv_sec - start.tv_sec)
+ * 1000000000UL
+ + (unsigned long)(now.tv_nsec - start.tv_nsec);
+ } while (elapsed < duration_ms * 1000000UL);
+
+ tlob_preempt_work_done();
+}
+
+int main(int argc, char *argv[])
+{
+ unsigned long duration_ms = 0;
+ const char *mode = "busy";
+ struct timespec deadline, now;
+
+ if (argc >= 2)
+ duration_ms = strtoul(argv[1], NULL, 10);
+ if (argc >= 3)
+ mode = argv[2];
+
+ clock_gettime(CLOCK_MONOTONIC, &deadline);
+ timespec_add_ms(&deadline, duration_ms ? duration_ms : 86400000UL);
+
+ do {
+ if (strcmp(mode, "sleep") == 0)
+ tlob_sleep_work(200);
+ else if (strcmp(mode, "preempt") == 0)
+ tlob_preempt_work(200);
+ else
+ tlob_busy_work(200);
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ } while (timespec_before(&now, &deadline));
+
+ return 0;
+}
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH v5 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
` (7 preceding siblings ...)
2026-08-19 18:15 ` [PATCH v5 8/9] selftests/verification: Add tlob selftests wen.yang
@ 2026-08-19 18:15 ` wen.yang
8 siblings, 0 replies; 10+ messages in thread
From: wen.yang @ 2026-08-19 18:15 UTC (permalink / raw)
To: Gabriele Monaco; +Cc: Nam Cao, linux-trace-kernel, linux-kernel, Wen Yang
From: Wen Yang <wen.yang@linux.dev>
When a test directory that does not itself contain test.d/functions is
passed to ftracetest (e.g. verification/test.d/tlob/), ftracetest fell
back to its own functions file and lost the rv-specific check_requires
handling for ':monitor' and ':reactor' requirements.
Walk up the directory tree from OPT_TEST_DIR until a directory containing
test.d/functions is found. This allows monitor subdirectories to be passed
directly as the test root without placing a functions shim in each one.
The RV verification suite uses this so that
tools/testing/selftests/verification/test.d/tlob/run_tlob_tests.sh
can pass test.d/tlob/ to ftracetest and have it source
verification/test.d/functions (which understands ':monitor'/':reactor').
Suggested-by: Gabriele Monaco <gmonaco@redhat.com>
Signed-off-by: Wen Yang <wen.yang@linux.dev>
---
tools/testing/selftests/ftrace/ftracetest | 17 ++++++++++++++---
1 file changed, 14 insertions(+), 3 deletions(-)
diff --git a/tools/testing/selftests/ftrace/ftracetest b/tools/testing/selftests/ftrace/ftracetest
index 0a56bf209f6c..3ba929820dd2 100755
--- a/tools/testing/selftests/ftrace/ftracetest
+++ b/tools/testing/selftests/ftrace/ftracetest
@@ -159,9 +159,20 @@ parse_opts() { # opts
if [ -n "$OPT_TEST_CASES" ]; then
TEST_CASES=$OPT_TEST_CASES
fi
- if [ -n "$OPT_TEST_DIR" -a -f "$OPT_TEST_DIR"/test.d/functions ]; then
- TOP_DIR=$OPT_TEST_DIR
- TEST_DIR=$TOP_DIR/test.d
+ if [ -n "$OPT_TEST_DIR" ]; then
+ # Walk up from OPT_TEST_DIR to find the nearest ancestor that contains
+ # test.d/functions. This allows a monitor subdirectory (e.g.
+ # verification/test.d/tlob/) to be passed directly without placing a
+ # dummy functions shim in each new subdirectory.
+ dir=$OPT_TEST_DIR
+ while [ "$dir" != "/" ]; do
+ if [ -f "$dir/test.d/functions" ]; then
+ TOP_DIR=$dir
+ TEST_DIR=$TOP_DIR/test.d
+ break
+ fi
+ dir=$(dirname "$dir")
+ done
fi
}
--
2.25.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
end of thread, other threads:[~2026-08-19 18:16 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-19 18:15 [PATCH v5 0/9] rv: Add task latency over budget RV monitor wen.yang
2026-08-19 18:15 ` [PATCH v5 1/9] rv: Introduce DA_MON_ALLOCATION_STRATEGY wen.yang
2026-08-19 18:15 ` [PATCH v5 2/9] rv: Add generic uprobe infrastructure for RV monitors wen.yang
2026-08-19 18:15 ` [PATCH v5 3/9] rv: Add tlob model DOT file wen.yang
2026-08-19 18:15 ` [PATCH v5 4/9] rv: Fix ha_invariant_passed_ns silent bypass of invariant check wen.yang
2026-08-19 18:15 ` [PATCH v5 5/9] rv: Make da_monitor_reset_hook and EVENT_NONE_LBL overridable wen.yang
2026-08-19 18:15 ` [PATCH v5 6/9] rv: Add tlob hybrid automaton monitor wen.yang
2026-08-19 18:15 ` [PATCH v5 7/9] rv: Add KUnit tests for the tlob monitor wen.yang
2026-08-19 18:15 ` [PATCH v5 8/9] selftests/verification: Add tlob selftests wen.yang
2026-08-19 18:15 ` [PATCH v5 9/9] selftests/ftrace: Walk up to find test.d/functions when a subdirectory is passed wen.yang
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox