Linux CXL
 help / color / mirror / Atom feed
* [PATCH v2 0/2] cxl/memdev: Fix poison debugfs vs unbind deadlock
@ 2026-08-31 12:48 Guixin Liu
  2026-08-31 12:48 ` [PATCH v2 1/2] driver core: Add conditional guard support for device_trylock() Guixin Liu
  2026-08-31 12:48 ` [PATCH v2 2/2] cxl/memdev: Fix deadlock between poison debugfs and cxl_mem unbind Guixin Liu
  0 siblings, 2 replies; 3+ messages in thread
From: Guixin Liu @ 2026-08-31 12:48 UTC (permalink / raw)
  To: Davidlohr Bueso, Jonathan Cameron, Dave Jiang, Alison Schofield,
	Vishal Verma, Dan Williams, Ira Weiny, Li Ming,
	Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
	Shaikh Kamaluddin
  Cc: linux-cxl, driver-core

Writing a memdev poison debugfs file while cxl_mem is being unbound
deadlocks. Patch 2 fixes it by not waiting for the device lock in those
handlers. Patch 1 adds the trylock guard it uses.

Patch 2 does not build without patch 1, so the two need to travel
together.

Testing:

Reproduced on a QEMU CXL topology whose type3 devices advertise poison
inject support:

  while :; do echo 0 > /sys/kernel/debug/cxl/mem0/inject_poison; done &
  while :; do
      echo mem0 > /sys/bus/cxl/drivers/cxl_mem/unbind
      echo mem0 > /sys/bus/cxl/drivers/cxl_mem/bind
  done

Without the fix the unbind wedges within seconds. The three tasks
involved, from /proc/<pid>/stack:

  writer, state S, holds the debugfs reference and waits for the lock
    cxl_debugfs_poison_inject+0x25/0xa0 [cxl_mem]
    debugfs_attr_write+0x61/0xb0
    full_proxy_write+0xfc/0x1c0
    vfs_write+0x1d4/0xe60

  unbind, state D, holds the lock and waits for the reference to drain
    remove_one+0x27f/0x3d0
    debugfs_remove+0x44/0x60
    release_nodes+0xfa/0x2c0
    devres_release_all+0x113/0x1a0
    device_unbind_cleanup+0x76/0x260
    device_release_driver_internal+0x3eb/0x540
    unbind_store+0xde/0x100

  cxl_port workqueue, state D, blocked on the same lock
    device_release_driver_internal+0x96/0x540
    detach_memdev+0x79/0xb0 [cxl_core]
    process_one_work+0x6b0/0xfb0

The writer is in interruptible sleep and can be killed; the unbind
cannot, and because cxl_bus_wq is an ordered workqueue the wedged
detach_memdev() blocks every other CXL bus work item behind it.

With both patches applied, 46 unbind/bind cycles against the same writer
loop all completed, no task was left in D state, and the writer collected
10920 EBUSY returns from the contended trylock. Note that lockdep stays
quiet either way: one leg of the cycle is the debugfs active_users
completion rather than a lock it tracks.

v1 -> v2:
- add the device_trylock() guard and use ACQUIRE(device_try, ...) instead
  of open-coding device_trylock()/device_unlock(), keeping the style the
  Fixes: commit established (Shaikh Kamaluddin)
- cut the changelog down to the failing condition, the consequence and
  the fix; the call graph and the reproducer live here instead
- say how the issue was found and how it was tested

v1:
https://lore.kernel.org/linux-cxl/\
20260826125248.4003792-1-kanie@linux.alibaba.com/

Guixin Liu (2):
  driver core: Add conditional guard support for device_trylock()
  cxl/memdev: Fix deadlock between poison debugfs and cxl_mem unbind

 drivers/cxl/mem.c      | 14 ++++++++++----
 include/linux/device.h |  1 +
 2 files changed, 11 insertions(+), 4 deletions(-)


base-commit: 7098e9cd98a05c0c5de2fae0c2465f9d966fdd07
-- 
2.43.7

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-08-31 12:48 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-31 12:48 [PATCH v2 0/2] cxl/memdev: Fix poison debugfs vs unbind deadlock Guixin Liu
2026-08-31 12:48 ` [PATCH v2 1/2] driver core: Add conditional guard support for device_trylock() Guixin Liu
2026-08-31 12:48 ` [PATCH v2 2/2] cxl/memdev: Fix deadlock between poison debugfs and cxl_mem unbind Guixin Liu

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox