All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/3] workqueue: Add telemetry tracepoints for CPU hogs, distress, and BH budget yields
@ 2026-08-29 23:05 Aaron Tomlin
  2026-08-29 23:05 ` [PATCH 1/3] workqueue: Add workqueue_cpu_intensive tracepoint Aaron Tomlin
                   ` (3 more replies)
  0 siblings, 4 replies; 7+ messages in thread
From: Aaron Tomlin @ 2026-08-29 23:05 UTC (permalink / raw)
  To: tj
  Cc: jianshanlai, rostedt, mhiramat, osandov, atomlin, neelx, sean,
	linux-kernel, linux-trace-kernel

Hi Tejun, Lai,

While the workqueue subsystem maintains rich internal telemetry via
pwq->stats[], wq_cpu_intensive_thresh_us, and distress mechanisms, several
critical state transitions and execution anomalies currently lack real-time
event notifications.

Across production fleets and low-latency networking workloads, polling
pwq->stats[] or running drgn scripts is impractical for detecting
intermittent stalls. Tail-latency spikes and packet drops often stem from
softirq overruns or latency-critical work items queueing behind CPU-bound
tasks. These event-driven tracepoints allow zero-overhead eBPF tools and
latency profilers to capture stack traces and kernel context at the exact
moment a starvation event or softirq budget exhaustion occurs.

This patch series introduces lightweight tracepoints for these key
operational boundaries:

Patch 1 adds workqueue_cpu_intensive tracepoint. When a concurrency-managed
worker runs for longer than wq_cpu_intensive_thresh_us without sleeping,
wq_worker_tick() marks it as WORKER_CPU_INTENSIVE and kicks the pool to
prevent queue starvation. The tracepoint will capture the offending work
function, workqueue name, CPU, and elapsed duration.

Patch 2 adds workqueue_mayday and workqueue_rescued tracepoints. One when
worker allocation stalls trigger mayday distress, and another when pending
work items are handed off to the rescuer thread to ensure forward progress.

Patch 3 adds workqueue_bh_budget_yield tracepoint. Bottom-Half (BH)
workqueues enforce execution limits in softirq context (BH_WORKER_JIFFIES
and BH_WORKER_RESTARTS). When a BH worker hits these limits while pending
work remains, it yields and re-raises the softirq. The tracepoint can be
used to identify softirq budget saturation and track whether yielding
occurred due to time slice expiration or restart counts.

Aaron Tomlin (3):
  workqueue: Add workqueue_cpu_intensive tracepoint
  workqueue: Add workqueue_mayday and workqueue_rescued tracepoints
  workqueue: Add workqueue_bh_budget_yield tracepoint

 include/trace/events/workqueue.h | 146 +++++++++++++++++++++++++++++++
 kernel/workqueue.c               |  24 ++++-
 2 files changed, 168 insertions(+), 2 deletions(-)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-31 21:08 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-29 23:05 [PATCH 0/3] workqueue: Add telemetry tracepoints for CPU hogs, distress, and BH budget yields Aaron Tomlin
2026-08-29 23:05 ` [PATCH 1/3] workqueue: Add workqueue_cpu_intensive tracepoint Aaron Tomlin
2026-08-29 23:14   ` sashiko-bot
2026-08-29 23:05 ` [PATCH 2/3] workqueue: Add workqueue_mayday and workqueue_rescued tracepoints Aaron Tomlin
2026-08-29 23:05 ` [PATCH 3/3] workqueue: Add workqueue_bh_budget_yield tracepoint Aaron Tomlin
2026-08-31 17:41   ` kernel test robot
2026-08-31 21:08 ` [PATCH 0/3] workqueue: Add telemetry tracepoints for CPU hogs, distress, and BH budget yields Tejun Heo

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.