* [PATCH 0/3] thunderbolt: Reset affected AMD host interfaces before DMA HopID reuse
@ 2026-10-05 13:43 Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail Basavaraj Natikar
` (2 more replies)
0 siblings, 3 replies; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-05 13:43 UTC (permalink / raw)
To: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni
Cc: linux-usb, linux-doc, Mario Limonciello, Sanath S,
Basavaraj Natikar
Some AMD USB4 host routers hang a TX ring when a DMA HopID is reused after
a DMA path is torn down without a host interface reset in between. Resetting
on every teardown clears the state but also disrupts other active DMA
tunnels on the same host interface.
Track the TX and RX DMA HopIDs programmed since the last reset and prefer
unused HopIDs when allocating rings. Check reuse at tb_ring_start() as well,
since networking retains its rings across reconnect, and return -EAGAIN
instead of reprogramming a HopID that still needs a reset. Run the reset
from a per-NHI work item once all DMA rings are idle, serialized with the
connection manager and with the control channel stopped only over the reset.
Enable it only for the affected AMD routers via QUIRK_RESET_DMA_ON_REUSE,
and only on v1 host interfaces.
On -EAGAIN networking retries login asynchronously; stream and DMA-test
users retry the operation. An active tunnel is left running even if it
blocks another tunnel from starting until it stops.
Basavaraj Natikar (3):
thunderbolt: Allow tb_ring_start() to fail
thunderbolt: Reset the host interface before reusing a DMA HopID
thunderbolt: Add quirk to reset host interface for AMD USB4 routers
Documentation/admin-guide/thunderbolt.rst | 6 +
drivers/net/thunderbolt/main.c | 120 +++++++----
drivers/thunderbolt/ctl.c | 36 +++-
drivers/thunderbolt/ctl.h | 1 +
drivers/thunderbolt/dma_test.c | 32 ++-
drivers/thunderbolt/domain.c | 66 +++++-
drivers/thunderbolt/nhi.c | 249 ++++++++++++++++++++--
drivers/thunderbolt/nhi.h | 13 ++
drivers/thunderbolt/nhi_regs.h | 4 +
drivers/thunderbolt/pci.c | 21 ++
drivers/thunderbolt/stream.c | 175 +++++++++------
include/linux/thunderbolt.h | 22 +-
12 files changed, 612 insertions(+), 133 deletions(-)
--
2.34.1
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail
2026-10-05 13:43 [PATCH 0/3] thunderbolt: Reset affected AMD host interfaces before DMA HopID reuse Basavaraj Natikar
@ 2026-10-05 13:43 ` Basavaraj Natikar
2026-10-06 4:27 ` Mika Westerberg
2026-10-05 13:43 ` [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 3/3] thunderbolt: Add quirk to reset host interface for AMD USB4 routers Basavaraj Natikar
2 siblings, 1 reply; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-05 13:43 UTC (permalink / raw)
To: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni
Cc: linux-usb, linux-doc, Mario Limonciello, Sanath S,
Basavaraj Natikar
tb_ring_start() returns void, so its callers cannot tell when a ring fails
to start and keep building an unusable tunnel. On some host interfaces a
DMA HopID also cannot be reprogrammed until the host interface has been
reset.
Hence, let tb_ring_start() return an error and unwind the callers on
failure: stop an already started TX ring when its RX peer fails to start,
and disable the DMA paths enabled before the rings were started.
A stream can also stay open after a failed resume. Therefore, free the
partial allocations, clear the ring pointers, and let the subsequent I/O
and close return without touching the freed rings. Check readiness under
the device mutex and use a wake token so a wakeup is not lost across the
unlocked sleep.
Co-developed-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
---
drivers/net/thunderbolt/main.c | 14 ++-
drivers/thunderbolt/ctl.c | 23 ++++-
drivers/thunderbolt/dma_test.c | 32 +++++-
drivers/thunderbolt/nhi.c | 12 ++-
drivers/thunderbolt/stream.c | 175 +++++++++++++++++++++------------
include/linux/thunderbolt.h | 2 +-
6 files changed, 181 insertions(+), 77 deletions(-)
diff --git a/drivers/net/thunderbolt/main.c b/drivers/net/thunderbolt/main.c
index cf51b9c39f4e..93ccccc5cf8b 100644
--- a/drivers/net/thunderbolt/main.c
+++ b/drivers/net/thunderbolt/main.c
@@ -669,8 +669,16 @@ static void tbnet_connected_work(struct work_struct *work)
* the Rx ring before any incoming packets are allowed to
* arrive.
*/
- tb_ring_start(net->tx_ring.ring);
- tb_ring_start(net->rx_ring.ring);
+ ret = tb_ring_start(net->tx_ring.ring);
+ if (ret) {
+ netdev_dbg(net->dev, "failed to start Tx ring, ret=%d\n", ret);
+ goto err_release_hopid;
+ }
+ ret = tb_ring_start(net->rx_ring.ring);
+ if (ret) {
+ netdev_dbg(net->dev, "failed to start Rx ring, ret=%d\n", ret);
+ goto err_stop_tx;
+ }
ret = tbnet_alloc_rx_buffers(net, TBNET_RING_SIZE);
if (ret)
@@ -701,7 +709,9 @@ static void tbnet_connected_work(struct work_struct *work)
tbnet_free_buffers(&net->rx_ring);
err_stop_rings:
tb_ring_stop(net->rx_ring.ring);
+err_stop_tx:
tb_ring_stop(net->tx_ring.ring);
+err_release_hopid:
tb_xdomain_release_in_hopid(net->xd, net->remote_transmit_path);
tbnet_connect_failed(net);
}
diff --git a/drivers/thunderbolt/ctl.c b/drivers/thunderbolt/ctl.c
index 965988b18608..94b29430ffe4 100644
--- a/drivers/thunderbolt/ctl.c
+++ b/drivers/thunderbolt/ctl.c
@@ -728,10 +728,27 @@ void tb_ctl_free(struct tb_ctl *ctl)
*/
void tb_ctl_start(struct tb_ctl *ctl)
{
- int i;
+ int i, ret;
tb_ctl_dbg(ctl, "control channel starting...\n");
- tb_ring_start(ctl->tx); /* is used to ack hotplug packets, start first */
- tb_ring_start(ctl->rx);
+
+ /*
+ * TX is used to ack hotplug packets so start it first. -ENODEV
+ * means the host controller itself is already gone (expected on
+ * an unplug-during-suspend resume), so do not warn about that.
+ */
+ ret = tb_ring_start(ctl->tx);
+ if (ret) {
+ if (ret != -ENODEV)
+ tb_ctl_WARN(ctl, "failed to start TX ring\n");
+ return;
+ }
+ ret = tb_ring_start(ctl->rx);
+ if (ret) {
+ if (ret != -ENODEV)
+ tb_ctl_WARN(ctl, "failed to start RX ring\n");
+ tb_ring_stop(ctl->tx);
+ return;
+ }
for (i = 0; i < TB_CTL_RX_PKG_COUNT; i++)
tb_ctl_rx_submit(ctl->rx_packets[i]);
diff --git a/drivers/thunderbolt/dma_test.c b/drivers/thunderbolt/dma_test.c
index bcecb0edcb81..e9c01bcfedf9 100644
--- a/drivers/thunderbolt/dma_test.c
+++ b/drivers/thunderbolt/dma_test.c
@@ -203,12 +203,36 @@ static int dma_test_start_rings(struct dma_test *dt)
return ret;
}
- if (dt->tx_ring)
- tb_ring_start(dt->tx_ring);
- if (dt->rx_ring)
- tb_ring_start(dt->rx_ring);
+ if (dt->tx_ring) {
+ ret = tb_ring_start(dt->tx_ring);
+ if (ret)
+ goto err_disable_paths;
+ }
+ if (dt->rx_ring) {
+ ret = tb_ring_start(dt->rx_ring);
+ if (ret)
+ goto err_stop_tx;
+ }
return 0;
+
+err_stop_tx:
+ tb_xdomain_disable_paths(dt->xd, dt->tx_hopid,
+ dt->tx_ring ? dt->tx_ring->hop : -1,
+ dt->rx_hopid,
+ dt->rx_ring ? dt->rx_ring->hop : -1);
+ if (dt->tx_ring)
+ tb_ring_stop(dt->tx_ring);
+ goto err_free;
+err_disable_paths:
+ tb_xdomain_disable_paths(dt->xd, dt->tx_hopid,
+ dt->tx_ring ? dt->tx_ring->hop : -1,
+ dt->rx_hopid,
+ dt->rx_ring ? dt->rx_ring->hop : -1);
+err_free:
+ dma_test_free_rings(dt);
+
+ return ret;
}
static void dma_test_stop_rings(struct dma_test *dt)
diff --git a/drivers/thunderbolt/nhi.c b/drivers/thunderbolt/nhi.c
index e99a3fcc4a29..1960fa30e13a 100644
--- a/drivers/thunderbolt/nhi.c
+++ b/drivers/thunderbolt/nhi.c
@@ -741,18 +741,24 @@ EXPORT_SYMBOL_GPL(tb_ring_alloc_rx);
* @ring: Ring to start
*
* Must not be invoked in parallel with tb_ring_stop().
+ *
+ * Returns %0 on success and negative errno in case of failure.
*/
-void tb_ring_start(struct tb_ring *ring)
+int tb_ring_start(struct tb_ring *ring)
{
u16 frame_size;
+ int ret = 0;
u32 flags;
spin_lock_irq(&ring->nhi->lock);
spin_lock(&ring->lock);
- if (ring->nhi->going_away)
+ if (ring->nhi->going_away) {
+ ret = -ENODEV;
goto err;
+ }
if (ring->running) {
dev_WARN(ring->nhi->dev, "ring already started\n");
+ ret = -EBUSY;
goto err;
}
dev_dbg(ring->nhi->dev, "starting %s %d\n",
@@ -810,6 +816,8 @@ void tb_ring_start(struct tb_ring *ring)
err:
spin_unlock(&ring->lock);
spin_unlock_irq(&ring->nhi->lock);
+
+ return ret;
}
EXPORT_SYMBOL_GPL(tb_ring_start);
diff --git a/drivers/thunderbolt/stream.c b/drivers/thunderbolt/stream.c
index 4f9a57b77bfa..48b7e3f01312 100644
--- a/drivers/thunderbolt/stream.c
+++ b/drivers/thunderbolt/stream.c
@@ -259,6 +259,9 @@ static void tbstream_ring_free(struct tbstream_ring *ring)
enum dma_data_direction dir;
int i;
+ if (!ring->frames)
+ return;
+
if (ring->ring->is_tx)
dir = DMA_TO_DEVICE;
else
@@ -279,6 +282,7 @@ static void tbstream_ring_free(struct tbstream_ring *ring)
ring->prod = 0;
ring->cons = 0;
kfree(ring->frames);
+ ring->frames = NULL;
}
static inline bool tbstream_ring_available(const struct tbstream_ring *ring)
@@ -575,6 +579,9 @@ static int tbstream_dev_send_close(struct tbstream_dev *sdev)
struct tbstream_frame *sf;
ktime_t timeout;
+ if (!sdev->tx_ring.ring)
+ return -ESHUTDOWN;
+
/*
* Wait for the ring to have available slots before we send the
* CLOSE packet.
@@ -645,7 +652,7 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
ret = tbstream_dev_alloc_tx_buffers(sdev);
if (ret)
- goto err_free_tx;
+ goto err_free_tx_buffers;
e2e_tx_hop = ring->hop;
sof_mask = BIT(TBSTREAM_FRAME_START);
@@ -672,8 +679,12 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
sdev->rx_pending = false;
- tb_ring_start(sdev->tx_ring.ring);
- tb_ring_start(sdev->rx_ring.ring);
+ ret = tb_ring_start(sdev->tx_ring.ring);
+ if (ret)
+ goto err_disable_paths;
+ ret = tb_ring_start(sdev->rx_ring.ring);
+ if (ret)
+ goto err_stop_tx;
ret = tbstream_dev_alloc_rx_buffers(sdev);
if (ret)
@@ -682,13 +693,20 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
err_stop:
tb_ring_stop(sdev->rx_ring.ring);
+ tbstream_ring_free(&sdev->rx_ring);
+err_stop_tx:
tb_ring_stop(sdev->tx_ring.ring);
+err_disable_paths:
+ tb_xdomain_disable_paths(xd, sdev->out_hopid, sdev->tx_ring.ring->hop,
+ sdev->in_hopid, sdev->rx_ring.ring->hop);
err_free_rx:
tb_ring_free(sdev->rx_ring.ring);
+ sdev->rx_ring.ring = NULL;
err_free_tx_buffers:
tbstream_ring_free(&sdev->tx_ring);
-err_free_tx:
tb_ring_free(sdev->tx_ring.ring);
+ sdev->tx_ring.ring = NULL;
+ wake_up_interruptible(&sdev->wait);
return ret;
}
@@ -708,6 +726,10 @@ static void tbstream_dev_stop(struct tbstream_dev *sdev)
{
struct tb_xdomain *xd;
+ /* Starting may have failed and freed the rings already */
+ if (!sdev->tx_ring.ring)
+ return;
+
if (sdev->busy_poll) {
/*
* When busy polling we must advance the ring ourselves
@@ -742,6 +764,7 @@ static void tbstream_dev_stop(struct tbstream_dev *sdev)
tbstream_ring_free(&sdev->tx_ring);
tb_ring_free(sdev->tx_ring.ring);
sdev->tx_ring.ring = NULL;
+ wake_up_interruptible(&sdev->wait);
}
/* Use only with read_iter/write_iter() to handle nowait */
@@ -757,29 +780,6 @@ static int tbstream_dev_lock(struct tbstream_dev *sdev, bool nowait)
return 0;
}
-/* Must not be called with @sdev->lock held */
-static int tbstream_dev_busy_poll_wait(struct tbstream_dev *sdev,
- struct tbstream_ring *ring)
-{
- for (;;) {
- if (signal_pending(current))
- return -ERESTARTSYS;
- if (tb_ring_poll_pending(ring->ring))
- return 0;
- /*
- * For TX ring we need to check the RX side too because
- * it might have received CLOSE packet.
- */
- if (ring == &sdev->tx_ring &&
- tb_ring_poll_pending(sdev->rx_ring.ring))
- return 0;
- if (tbstream_dev_valid(sdev) != 0 ||
- tbstream_dev_closed(sdev) || tbstream_dev_removed(sdev))
- return 0;
- cond_resched();
- }
-}
-
static bool
tbstream_dev_has_event(struct tbstream_dev *sdev, struct tbstream_ring *ring)
{
@@ -796,6 +796,52 @@ tbstream_dev_has_event(struct tbstream_dev *sdev, struct tbstream_ring *ring)
return tb_ring_poll_pending(sdev->rx_ring.ring);
}
+static bool tbstream_dev_ready(struct tbstream_dev *sdev,
+ struct tbstream_ring *ring)
+{
+ lockdep_assert_held(&sdev->lock);
+
+ /* Starting may have failed and freed the rings already */
+ if (!sdev->tx_ring.ring)
+ return true;
+
+ return tbstream_dev_has_event(sdev, ring) ||
+ tbstream_dev_close_received(sdev) ||
+ tb_ring_poll_pending(ring->ring);
+}
+
+/* Must not be called with @sdev->lock held. */
+static int tbstream_dev_wait(struct tbstream_dev *sdev,
+ struct tbstream_ring *ring)
+{
+ DEFINE_WAIT_FUNC(wait, woken_wake_function);
+ int ret = 0;
+
+ add_wait_queue(&sdev->wait, &wait);
+ for (;;) {
+ bool ready;
+
+ ret = mutex_lock_interruptible(&sdev->lock);
+ if (ret)
+ break;
+ ready = tbstream_dev_ready(sdev, ring);
+ mutex_unlock(&sdev->lock);
+ if (ready)
+ break;
+ if (signal_pending(current)) {
+ ret = -ERESTARTSYS;
+ break;
+ }
+ if (sdev->busy_poll)
+ cond_resched();
+ else
+ /* The wake token bridges the unlocked check-to-sleep gap. */
+ wait_woken(&wait, TASK_INTERRUPTIBLE, MAX_SCHEDULE_TIMEOUT);
+ }
+ remove_wait_queue(&sdev->wait, &wait);
+ return ret;
+}
+
static ssize_t
tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
{
@@ -814,6 +860,11 @@ tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
return ret;
for (;;) {
+ if (!sdev->tx_ring.ring) {
+ mutex_unlock(&sdev->lock);
+ return -ESHUTDOWN;
+ }
+
/* Advance RX completions */
tbstream_dev_advance_rx(sdev);
@@ -838,16 +889,9 @@ tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
if (nowait)
return -EAGAIN;
- if (sdev->busy_poll) {
- ret = tbstream_dev_busy_poll_wait(sdev, &sdev->rx_ring);
- if (ret)
- return ret;
- } else {
- ret = wait_event_interruptible(sdev->wait,
- tbstream_dev_has_event(sdev, &sdev->rx_ring));
- if (ret)
- return ret;
- }
+ ret = tbstream_dev_wait(sdev, &sdev->rx_ring);
+ if (ret)
+ return ret;
ret = tbstream_dev_lock(sdev, nowait);
if (ret)
@@ -957,6 +1001,11 @@ tbstream_dev_fops_write_iter(struct kiocb *kiocb, struct iov_iter *from)
return ret;
for (;;) {
+ if (!sdev->tx_ring.ring) {
+ mutex_unlock(&sdev->lock);
+ return -ESHUTDOWN;
+ }
+
/* Advance TX (and RX) completions */
tbstream_dev_advance_both(sdev);
@@ -984,17 +1033,9 @@ tbstream_dev_fops_write_iter(struct kiocb *kiocb, struct iov_iter *from)
if (nowait)
return -EAGAIN;
- if (sdev->busy_poll) {
- ret = tbstream_dev_busy_poll_wait(sdev, &sdev->tx_ring);
- if (ret)
- return ret;
- } else {
- ret = wait_event_interruptible(sdev->wait,
- tbstream_dev_has_event(sdev, &sdev->tx_ring) ||
- tbstream_dev_close_received(sdev));
- if (ret)
- return ret;
- }
+ ret = tbstream_dev_wait(sdev, &sdev->tx_ring);
+ if (ret)
+ return ret;
ret = tbstream_dev_lock(sdev, nowait);
if (ret)
@@ -1041,7 +1082,7 @@ tbstream_dev_fops_poll(struct file *file, struct poll_table_struct *wait)
poll_wait(file, &sdev->wait, wait);
guard(mutex)(&sdev->lock);
- if (tbstream_dev_valid(sdev) != 0)
+ if (tbstream_dev_valid(sdev) != 0 || !sdev->tx_ring.ring)
return EPOLLHUP | EPOLLERR;
/*
@@ -1105,6 +1146,11 @@ static int tbstream_dev_fops_open(struct inode *inode, struct file *file)
}
}
+ if (sdev->users && !sdev->tx_ring.ring) {
+ ret = -ESHUTDOWN;
+ goto err_unlock;
+ }
+
/* Only on first open we allocate rings and enable paths */
if (!sdev->users++) {
ret = tbstream_dev_start(sdev);
@@ -1136,7 +1182,7 @@ static int tbstream_dev_fops_release(struct inode *inode, struct file *file)
struct tbstream_dev *sdev = to_tbstream_dev(file->private_data);
mutex_lock(&sdev->lock);
- if (--sdev->users == 0) {
+ if (--sdev->users == 0 && sdev->tx_ring.ring) {
/*
* Advance now in case there is CLOSE waiting in the RX
* ring.
@@ -1884,14 +1930,14 @@ static int __maybe_unused tbstream_suspend(struct device *dev)
if (!sg)
return 0;
+ mutex_lock(&sg->lock);
list_for_each_entry_reverse(sdev, &sg->dev_list, list) {
- tbstream_dev_get(sdev);
- /* Stop the stream (if it was open) */
+ mutex_lock(&sdev->lock);
if (sdev->users)
tbstream_dev_stop(sdev);
- tbstream_dev_put(sdev);
+ mutex_unlock(&sdev->lock);
}
-
+ mutex_unlock(&sg->lock);
config_group_put(&sg->group);
return 0;
}
@@ -1902,28 +1948,27 @@ static int __maybe_unused tbstream_resume(struct device *dev)
struct tbstream *stream = tb_service_get_drvdata(svc);
struct tbstream_group *sg;
struct tbstream_dev *sdev;
+ int ret = 0;
sg = tbstream_group_find(stream);
if (!sg)
return 0;
+ mutex_lock(&sg->lock);
list_for_each_entry(sdev, &sg->dev_list, list) {
- tbstream_dev_get(sdev);
+ mutex_lock(&sdev->lock);
if (sdev->users) {
- int ret;
+ int err = tbstream_dev_start(sdev);
- ret = tbstream_dev_start(sdev);
- if (ret) {
- tbstream_dev_put(sdev);
- config_group_put(&sg->group);
- return ret;
- }
+ if (err && !ret)
+ ret = err;
}
- tbstream_dev_put(sdev);
+ mutex_unlock(&sdev->lock);
+ wake_up_interruptible(&sdev->wait);
}
-
+ mutex_unlock(&sg->lock);
config_group_put(&sg->group);
- return 0;
+ return ret;
}
static const struct dev_pm_ops tbstream_pm_ops = {
diff --git a/include/linux/thunderbolt.h b/include/linux/thunderbolt.h
index 57502da29080..5fce1b67c376 100644
--- a/include/linux/thunderbolt.h
+++ b/include/linux/thunderbolt.h
@@ -672,7 +672,7 @@ struct tb_ring *tb_ring_alloc_rx(struct tb_nhi *nhi, int hop, int size,
unsigned int flags, int e2e_tx_hop,
u16 sof_mask, u16 eof_mask,
void (*start_poll)(void *), void *poll_data);
-void tb_ring_start(struct tb_ring *ring);
+int tb_ring_start(struct tb_ring *ring);
bool tb_ring_flush(struct tb_ring *ring, unsigned int timeout_msec);
void tb_ring_stop(struct tb_ring *ring);
void tb_ring_free(struct tb_ring *ring);
--
2.34.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-05 13:43 [PATCH 0/3] thunderbolt: Reset affected AMD host interfaces before DMA HopID reuse Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail Basavaraj Natikar
@ 2026-10-05 13:43 ` Basavaraj Natikar
2026-10-05 14:35 ` Mika Westerberg
2026-10-06 4:33 ` Mika Westerberg
2026-10-05 13:43 ` [PATCH 3/3] thunderbolt: Add quirk to reset host interface for AMD USB4 routers Basavaraj Natikar
2 siblings, 2 replies; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-05 13:43 UTC (permalink / raw)
To: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni
Cc: linux-usb, linux-doc, Mario Limonciello, Sanath S,
Basavaraj Natikar, Mika Westerberg
Reusing a DMA HopID without an intervening host interface reset can hang
the TX ring on some host routers. Resetting on every DMA tunnel teardown
clears the state but, as the reset affects all rings, also disturbs
unrelated active tunnels.
Hence, track the DMA HopIDs programmed since the last reset, prefer unused
HopIDs when allocating rings, and check for reuse at tb_ring_start() too,
since networking retains its rings across reconnect. Return -EAGAIN instead
of programming a HopID that still needs a reset.
Run the reset from a work item once all DMA rings are idle: serialize it
with the connection manager, stop the control channel around it, and block
DMA rings from starting during the reset. Fence the work across domain
removal and PM transitions, preserve live DMA rings across freeze/thaw, and
restore the interrupt-mask shadow under the NHI lock after the reset.
On -EAGAIN, retry the networking login asynchronously and block the work
producers before teardown cancels the workers. Keep peer disconnection
separate from administrative shutdown so it cannot reopen the login gate,
while stream and DMA-test callers unwind immediately and return the error
to userspace.
Enable this only for reset-capable host interfaces marked with
QUIRK_RESET_DMA_ON_REUSE. A competing DMA tunnel must stop before its dirty
HopID can be reused, and retrying does not interrupt that tunnel.
Suggested-by: Mika Westerberg <mika.westerberg@linux.intel.com>
Link: https://lore.kernel.org/linux-usb/20260831130638.GK124825@black.igk.intel.com/T/#m668c2efcfe2f298632537721d72b17445b39211a
Co-developed-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
---
Documentation/admin-guide/thunderbolt.rst | 6 +
drivers/net/thunderbolt/main.c | 106 ++++++----
drivers/thunderbolt/ctl.c | 13 ++
drivers/thunderbolt/ctl.h | 1 +
drivers/thunderbolt/domain.c | 66 +++++-
drivers/thunderbolt/nhi.c | 239 ++++++++++++++++++++--
drivers/thunderbolt/nhi.h | 1 +
drivers/thunderbolt/nhi_regs.h | 4 +
include/linux/thunderbolt.h | 20 ++
9 files changed, 399 insertions(+), 57 deletions(-)
diff --git a/Documentation/admin-guide/thunderbolt.rst b/Documentation/admin-guide/thunderbolt.rst
index ff25fe853706..5595d307ac86 100644
--- a/Documentation/admin-guide/thunderbolt.rst
+++ b/Documentation/admin-guide/thunderbolt.rst
@@ -410,6 +410,12 @@ transfer data::
host2 # cat /dev/tbstream0
host1 # dmesg > /dev/tbstream0
+On affected host interfaces a DMA HopID cannot be reused until the host
+interface has been reset. Opening a stream or starting a DMA test can return
+``EAGAIN`` until all other DMA tunnels stop and the reset completes. Retry
+the complete operation; it does not interrupt an active networking tunnel.
+If stream resume fails, close and reopen the stream to retry setup.
+
Once you are done with the stream you can remove them::
host2 # cd /sys/kernel/config/thunderbolt/stream
diff --git a/drivers/net/thunderbolt/main.c b/drivers/net/thunderbolt/main.c
index 93ccccc5cf8b..f2dbb6fb30c6 100644
--- a/drivers/net/thunderbolt/main.c
+++ b/drivers/net/thunderbolt/main.c
@@ -160,6 +160,9 @@ struct tbnet_ring {
* @login_sent: ThunderboltIP login message successfully sent
* @login_received: ThunderboltIP login message received from the remote
* host
+ * @stopping: Administrative teardown blocks all connection work
+ * @disconnecting: Peer logout blocks login until teardown completes
+ * @connected: DMA rings and paths were successfully enabled
* @local_transmit_path: HopID we are using to send out packets
* @remote_transmit_path: HopID the other end is using to send packets to us
* @connection_lock: Lock serializing access to @login_sent,
@@ -190,6 +193,9 @@ struct tbnet {
atomic_t command_id;
bool login_sent;
bool login_received;
+ bool stopping;
+ bool disconnecting;
+ bool connected;
int local_transmit_path;
int remote_transmit_path;
struct mutex connection_lock;
@@ -311,17 +317,21 @@ static void start_login(struct tbnet *net)
{
netdev_dbg(net->dev, "login started\n");
- mutex_lock(&net->connection_lock);
+ guard(mutex)(&net->connection_lock);
net->login_sent = false;
net->login_received = false;
- mutex_unlock(&net->connection_lock);
-
+ net->stopping = false;
+ net->disconnecting = false;
queue_delayed_work(system_long_wq, &net->login_work,
msecs_to_jiffies(1000));
}
static void stop_login(struct tbnet *net)
{
+ scoped_guard(mutex, &net->connection_lock)
+ net->stopping = true;
+
+ cancel_work_sync(&net->disconnect_work);
cancel_delayed_work_sync(&net->login_work);
cancel_work_sync(&net->connected_work);
@@ -371,11 +381,7 @@ static void tbnet_tear_down(struct tbnet *net, bool send_logout)
netif_carrier_off(net->dev);
netif_stop_queue(net->dev);
- stop_login(net);
-
- mutex_lock(&net->connection_lock);
-
- if (net->login_sent && net->login_received) {
+ if (net->connected) {
int ret, retries = TBNET_LOGOUT_RETRIES;
while (send_logout && retries-- > 0) {
@@ -413,13 +419,13 @@ static void tbnet_tear_down(struct tbnet *net, bool send_logout)
net->remote_transmit_path = 0;
}
+ guard(mutex)(&net->connection_lock);
+ net->connected = false;
net->login_retries = 0;
net->login_sent = false;
net->login_received = false;
netdev_dbg(net->dev, "network traffic stopped\n");
-
- mutex_unlock(&net->connection_lock);
}
static int tbnet_handle_packet(const void *buf, size_t size, void *data)
@@ -459,7 +465,11 @@ static int tbnet_handle_packet(const void *buf, size_t size, void *data)
if (!ret) {
netdev_dbg(net->dev, "remote login response sent\n");
- mutex_lock(&net->connection_lock);
+ guard(mutex)(&net->connection_lock);
+ if (net->stopping || net->disconnecting ||
+ (net->login_received &&
+ net->remote_transmit_path != pkg->transmit_path))
+ break;
net->login_received = true;
net->remote_transmit_path = pkg->transmit_path;
@@ -473,8 +483,6 @@ static int tbnet_handle_packet(const void *buf, size_t size, void *data)
queue_delayed_work(system_long_wq,
&net->login_work, 0);
}
- mutex_unlock(&net->connection_lock);
-
queue_work(system_long_wq, &net->connected_work);
}
break;
@@ -484,7 +492,12 @@ static int tbnet_handle_packet(const void *buf, size_t size, void *data)
ret = tbnet_logout_response(net, route, sequence, command_id);
if (!ret) {
netdev_dbg(net->dev, "remote logout response sent\n");
- queue_work(system_long_wq, &net->disconnect_work);
+ guard(mutex)(&net->connection_lock);
+ if (netif_running(net->dev) && !net->stopping &&
+ !net->disconnecting) {
+ net->disconnecting = true;
+ queue_work(system_long_wq, &net->disconnect_work);
+ }
}
break;
@@ -644,9 +657,9 @@ static void tbnet_connected_work(struct work_struct *work)
if (netif_carrier_ok(net->dev))
return;
- mutex_lock(&net->connection_lock);
- connected = net->login_sent && net->login_received;
- mutex_unlock(&net->connection_lock);
+ scoped_guard(mutex, &net->connection_lock)
+ connected = !net->stopping && !net->disconnecting &&
+ net->login_sent && net->login_received;
if (!connected)
return;
@@ -697,8 +710,13 @@ static void tbnet_connected_work(struct work_struct *work)
goto err_free_tx_buffers;
}
- netif_carrier_on(net->dev);
- netif_start_queue(net->dev);
+ scoped_guard(mutex, &net->connection_lock) {
+ net->connected = true;
+ if (!net->stopping && !net->disconnecting) {
+ netif_carrier_on(net->dev);
+ netif_start_queue(net->dev);
+ }
+ }
netdev_dbg(net->dev, "network traffic started\n");
return;
@@ -714,24 +732,34 @@ static void tbnet_connected_work(struct work_struct *work)
err_release_hopid:
tb_xdomain_release_in_hopid(net->xd, net->remote_transmit_path);
tbnet_connect_failed(net);
+ if (ret == -EAGAIN) {
+ guard(mutex)(&net->connection_lock);
+ if (!net->stopping && !net->disconnecting)
+ queue_delayed_work(system_long_wq, &net->login_work,
+ msecs_to_jiffies(TBNET_LOGIN_DELAY));
+ }
}
static void tbnet_login_work(struct work_struct *work)
{
struct tbnet *net = container_of(work, typeof(*net), login_work.work);
unsigned long delay = msecs_to_jiffies(TBNET_LOGIN_DELAY);
- int ret;
-
- if (netif_carrier_ok(net->dev))
- return;
+ int ret, retries;
- netdev_dbg(net->dev, "sending login request, retries=%u\n",
- net->login_retries);
+ scoped_guard(mutex, &net->connection_lock) {
+ if (net->stopping || net->disconnecting || net->connected)
+ return;
+ retries = net->login_retries;
+ }
- ret = tbnet_login_request(net, net->login_retries % 4);
+ netdev_dbg(net->dev, "sending login request, retries=%u\n", retries);
+ ret = tbnet_login_request(net, retries % 4);
if (ret) {
netdev_dbg(net->dev, "sending login request failed, ret=%d\n",
ret);
+ guard(mutex)(&net->connection_lock);
+ if (net->stopping || net->disconnecting)
+ return;
if (net->login_retries++ < TBNET_LOGIN_RETRIES) {
queue_delayed_work(system_long_wq, &net->login_work,
delay);
@@ -741,13 +769,12 @@ static void tbnet_login_work(struct work_struct *work)
} else {
netdev_dbg(net->dev, "received login reply\n");
- net->login_retries = 0;
-
- mutex_lock(&net->connection_lock);
- net->login_sent = true;
- mutex_unlock(&net->connection_lock);
-
- queue_work(system_long_wq, &net->connected_work);
+ guard(mutex)(&net->connection_lock);
+ if (!net->stopping && !net->disconnecting) {
+ net->login_retries = 0;
+ net->login_sent = true;
+ queue_work(system_long_wq, &net->connected_work);
+ }
}
}
@@ -755,7 +782,11 @@ static void tbnet_disconnect_work(struct work_struct *work)
{
struct tbnet *net = container_of(work, typeof(*net), disconnect_work);
+ cancel_delayed_work_sync(&net->login_work);
+ cancel_work_sync(&net->connected_work);
tbnet_tear_down(net, false);
+ guard(mutex)(&net->connection_lock);
+ net->disconnecting = false;
}
static bool tbnet_check_frame(struct tbnet *net, const struct tbnet_frame *tf,
@@ -1008,9 +1039,8 @@ static int tbnet_stop(struct net_device *dev)
{
struct tbnet *net = netdev_priv(dev);
+ stop_login(net);
napi_disable(&net->napi);
-
- cancel_work_sync(&net->disconnect_work);
tbnet_tear_down(net, true);
tb_ring_free(net->rx_ring.ring);
@@ -1397,6 +1427,7 @@ static int tbnet_probe(struct tb_service *svc)
net->svc = svc;
net->dev = dev;
net->xd = xd;
+ net->stopping = true;
tbnet_generate_mac(dev);
@@ -1456,7 +1487,10 @@ static void tbnet_remove(struct tb_service *svc)
static void tbnet_shutdown(struct tb_service *svc)
{
- tbnet_tear_down(tb_service_get_drvdata(svc), true);
+ struct tbnet *net = tb_service_get_drvdata(svc);
+
+ stop_login(net);
+ tbnet_tear_down(net, true);
}
static int tbnet_suspend(struct device *dev)
diff --git a/drivers/thunderbolt/ctl.c b/drivers/thunderbolt/ctl.c
index 94b29430ffe4..368b50b1ef2a 100644
--- a/drivers/thunderbolt/ctl.c
+++ b/drivers/thunderbolt/ctl.c
@@ -779,6 +779,19 @@ void tb_ctl_stop(struct tb_ctl *ctl)
tb_ctl_dbg(ctl, "control channel stopped\n");
}
+/* Close the enqueue gate only after all existing requests have drained. */
+bool tb_ctl_stop_if_idle(struct tb_ctl *ctl)
+{
+ scoped_guard(mutex, &ctl->request_queue_lock) {
+ if (!ctl->running || !list_empty(&ctl->request_queue))
+ return false;
+ ctl->running = false;
+ }
+
+ tb_ctl_stop(ctl);
+ return true;
+}
+
/* public interface, commands */
/**
diff --git a/drivers/thunderbolt/ctl.h b/drivers/thunderbolt/ctl.h
index db1646eb4fd0..a88c61fd628a 100644
--- a/drivers/thunderbolt/ctl.h
+++ b/drivers/thunderbolt/ctl.h
@@ -25,6 +25,7 @@ struct tb_ctl *tb_ctl_alloc(struct tb_nhi *nhi, int index, int timeout_msec,
event_cb cb, void *cb_data);
void tb_ctl_start(struct tb_ctl *ctl);
void tb_ctl_stop(struct tb_ctl *ctl);
+bool tb_ctl_stop_if_idle(struct tb_ctl *ctl);
void tb_ctl_free(struct tb_ctl *ctl);
/* configuration commands */
diff --git a/drivers/thunderbolt/domain.c b/drivers/thunderbolt/domain.c
index 24611f05b3cd..9b814fe1b68c 100644
--- a/drivers/thunderbolt/domain.c
+++ b/drivers/thunderbolt/domain.c
@@ -314,11 +314,20 @@ const struct bus_type tb_bus_type = {
.shutdown = tb_service_shutdown,
};
+/* Reset work needs tb->lock to finish. */
+static void tb_domain_cancel_nhi_reset(struct tb *tb)
+{
+ lockdep_assert_not_held(&tb->lock);
+ cancel_delayed_work_sync(&tb->nhi->reset_work);
+}
+
static void tb_domain_release(struct device *dev)
{
struct tb *tb = container_of(dev, struct tb, dev);
struct tb_nhi *nhi = tb->nhi;
+ /* The host interface reset runs against this domain */
+ tb_domain_cancel_nhi_reset(tb);
tb_ctl_free(tb->ctl);
destroy_workqueue(tb->wq);
ida_free(&tb_domain_ida, tb->index);
@@ -505,8 +514,14 @@ void tb_domain_remove(struct tb *tb)
tb->cm_ops->stop(tb);
/* Stop the domain control traffic */
tb_ctl_stop(tb->ctl);
+ /* Keep reset_work from restarting it again below */
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->removing = true;
+ }
mutex_unlock(&tb->lock);
+ tb_domain_cancel_nhi_reset(tb);
flush_workqueue(tb->wq);
if (tb->cm_ops->deinit)
@@ -535,10 +550,24 @@ int tb_domain_suspend_noirq(struct tb *tb)
mutex_lock(&tb->lock);
if (tb->cm_ops->suspend_noirq)
ret = tb->cm_ops->suspend_noirq(tb);
- if (!ret)
+ if (!ret) {
+ /* Keep reset_work from restarting the control channel below */
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = true;
+ }
tb_ctl_stop(tb->ctl);
+ }
mutex_unlock(&tb->lock);
+ /*
+ * Make sure a host interface reset queued by cm_ops->suspend_noirq()
+ * tearing down a DMA tunnel above either finishes or is cancelled
+ * outright before the NHI is powered down below - either outcome
+ * is fine since a dirty bitmap is handled again on resume.
+ */
+ tb_domain_cancel_nhi_reset(tb);
+
return ret;
}
@@ -556,6 +585,10 @@ int tb_domain_resume_noirq(struct tb *tb)
int ret = 0;
mutex_lock(&tb->lock);
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = false;
+ }
tb_ctl_start(tb->ctl);
if (tb->cm_ops->resume_noirq)
ret = tb->cm_ops->resume_noirq(tb);
@@ -576,10 +609,18 @@ int tb_domain_freeze_noirq(struct tb *tb)
mutex_lock(&tb->lock);
if (tb->cm_ops->freeze_noirq)
ret = tb->cm_ops->freeze_noirq(tb);
- if (!ret)
+ if (!ret) {
+ /* Keep reset_work from restarting the control channel below */
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = true;
+ }
tb_ctl_stop(tb->ctl);
+ }
mutex_unlock(&tb->lock);
+ tb_domain_cancel_nhi_reset(tb);
+
return ret;
}
@@ -588,6 +629,10 @@ int tb_domain_thaw_noirq(struct tb *tb)
int ret = 0;
mutex_lock(&tb->lock);
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = false;
+ }
tb_ctl_start(tb->ctl);
if (tb->cm_ops->thaw_noirq)
ret = tb->cm_ops->thaw_noirq(tb);
@@ -609,13 +654,30 @@ int tb_domain_runtime_suspend(struct tb *tb)
if (ret)
return ret;
}
+
+ /* Keep reset_work from restarting the control channel below */
+ mutex_lock(&tb->lock);
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = true;
+ }
tb_ctl_stop(tb->ctl);
+ mutex_unlock(&tb->lock);
+
+ tb_domain_cancel_nhi_reset(tb);
+
return 0;
}
int tb_domain_runtime_resume(struct tb *tb)
{
+ mutex_lock(&tb->lock);
+ scoped_guard(spinlock_irq, &tb->nhi->lock) {
+ if (tb->nhi->dma_hops_used)
+ tb->nhi->suspended = false;
+ }
tb_ctl_start(tb->ctl);
+ mutex_unlock(&tb->lock);
if (tb->cm_ops->runtime_resume) {
int ret = tb->cm_ops->runtime_resume(tb);
if (ret)
diff --git a/drivers/thunderbolt/nhi.c b/drivers/thunderbolt/nhi.c
index 1960fa30e13a..a68ecabadcb5 100644
--- a/drivers/thunderbolt/nhi.c
+++ b/drivers/thunderbolt/nhi.c
@@ -17,6 +17,7 @@
#include <linux/iommu.h>
#include <linux/lockdep.h>
#include <linux/module.h>
+#include <linux/pci.h>
#include <linux/delay.h>
#include <linux/property.h>
#include <linux/string_choices.h>
@@ -175,7 +176,8 @@ static void ring_interrupt_active(struct tb_ring *ring, bool active)
/*
* nhi_disable_interrupts() - disable interrupts for all rings
*
- * Use only during init and shutdown.
+ * During runtime reset the caller must hold nhi->lock.
+ * Init and shutdown callers must exclude concurrent ring operations.
*/
void nhi_disable_interrupts(struct tb_nhi *nhi)
{
@@ -513,7 +515,8 @@ void tb_ring_poll_complete(struct tb_ring *ring)
spin_lock_irqsave(&ring->nhi->lock, flags);
spin_lock(&ring->lock);
- if (ring->start_poll)
+ if (ring->start_poll && ring->running && !ring->nhi->resetting &&
+ !ring->nhi->going_away)
__ring_interrupt_mask(ring, false);
spin_unlock(&ring->lock);
spin_unlock_irqrestore(&ring->nhi->lock, flags);
@@ -549,6 +552,84 @@ irqreturn_t ring_msix(int irq, void *data)
return IRQ_HANDLED;
}
+static bool ring_is_dma(const struct tb_ring *ring)
+{
+ return ring->hop >= RING_FIRST_USABLE_HOPID;
+}
+
+/* TX and RX HopIDs are tracked separately */
+static unsigned int nhi_dma_hops_bits(const struct tb_nhi *nhi)
+{
+ return 2 * nhi->hop_count;
+}
+
+static unsigned int nhi_hop_bit(const struct tb_nhi *nhi, bool is_tx,
+ unsigned int hop)
+{
+ return is_tx ? hop : nhi->hop_count + hop;
+}
+
+/* Returns %true if a DMA HopID has been programmed since the last reset */
+static bool nhi_dma_hops_dirty(const struct tb_nhi *nhi)
+{
+ return nhi->dma_hops_used &&
+ !bitmap_empty(nhi->dma_hops_used, nhi_dma_hops_bits(nhi));
+}
+
+/* Returns %true if the HopID of @ring cannot be programmed again yet */
+static bool ring_needs_reset(const struct tb_ring *ring)
+{
+ const struct tb_nhi *nhi = ring->nhi;
+
+ lockdep_assert_held(&nhi->lock);
+
+ if (!nhi->dma_hops_used || !ring_is_dma(ring))
+ return false;
+
+ return test_bit(nhi_hop_bit(nhi, ring->is_tx, ring->hop),
+ nhi->dma_hops_used);
+}
+
+static bool nhi_dma_rings_running(const struct tb_nhi *nhi)
+{
+ unsigned int i;
+
+ lockdep_assert_held(&nhi->lock);
+
+ /* ring->running is only updated under nhi->lock */
+ for (i = RING_FIRST_USABLE_HOPID; i < nhi->hop_count; i++) {
+ if (nhi->tx_rings[i] && nhi->tx_rings[i]->running)
+ return true;
+ if (nhi->rx_rings[i] && nhi->rx_rings[i]->running)
+ return true;
+ }
+
+ return false;
+}
+
+static int nhi_find_hop(const struct tb_nhi *nhi, const struct tb_ring *ring,
+ unsigned int start_hop, bool skip_used)
+{
+ unsigned int i;
+
+ lockdep_assert_held(&nhi->lock);
+
+ for (i = start_hop; i < nhi->hop_count; i++) {
+ if (skip_used && test_bit(nhi_hop_bit(nhi, ring->is_tx, i),
+ nhi->dma_hops_used))
+ continue;
+ if (ring->is_tx) {
+ if (!nhi->tx_rings[i])
+ return i;
+ } else {
+ if (!nhi->rx_rings[i])
+ return i;
+ }
+ }
+
+ return -EBUSY;
+}
+
static int nhi_alloc_hop(struct tb_nhi *nhi, struct tb_ring *ring)
{
unsigned int start_hop = RING_FIRST_USABLE_HOPID;
@@ -566,25 +647,20 @@ static int nhi_alloc_hop(struct tb_nhi *nhi, struct tb_ring *ring)
spin_lock_irq(&nhi->lock);
if (ring->hop < 0) {
- unsigned int i;
+ int hop;
/*
* Automatically allocate HopID from the non-reserved
- * range 1 .. hop_count - 1.
+ * range 1 .. hop_count - 1, preferring the ones that do
+ * not need a host interface reset first.
*/
- for (i = start_hop; i < nhi->hop_count; i++) {
- if (ring->is_tx) {
- if (!nhi->tx_rings[i]) {
- ring->hop = i;
- break;
- }
- } else {
- if (!nhi->rx_rings[i]) {
- ring->hop = i;
- break;
- }
- }
- }
+ hop = -EBUSY;
+ if (nhi->dma_hops_used)
+ hop = nhi_find_hop(nhi, ring, start_hop, true);
+ if (hop < 0)
+ hop = nhi_find_hop(nhi, ring, start_hop, false);
+ if (hop >= 0)
+ ring->hop = hop;
}
if (ring->hop > 0 && ring->hop < start_hop) {
@@ -736,6 +812,73 @@ struct tb_ring *tb_ring_alloc_rx(struct tb_nhi *nhi, int hop, int size,
}
EXPORT_SYMBOL_GPL(tb_ring_alloc_rx);
+/*
+ * Brings the registers in the memory BAR back to their default state and
+ * clears the End-to-End Flow Control state. The caller must stop the
+ * control channel over the reset.
+ */
+static void nhi_reset_interface(struct tb_nhi *nhi)
+{
+ u32 val;
+
+ val = ioread32(nhi->iobase + REG_CAPS);
+ /* Only v1 host interfaces implement the reset */
+ if (FIELD_GET(REG_CAPS_VERSION_MASK, val) >= REG_CAPS_VERSION_2)
+ return;
+
+ dev_dbg(nhi->dev, "issuing host interface reset\n");
+
+ iowrite32(REG_HOST_INTERFACE_RESET_RST,
+ nhi->iobase + REG_HOST_INTERFACE_RESET);
+ /* Flush the posted write without accessing the resetting BAR. */
+ pci_read_config_dword(to_pci_dev(nhi->dev), PCI_VENDOR_ID, &val);
+ /* Wait for tHIReset (10 ms) to complete */
+ usleep_range(10000, 20000);
+
+ /* Serialize the interrupt shadow with late polling completions. */
+ scoped_guard(spinlock_irqsave, &nhi->lock)
+ nhi_disable_interrupts(nhi);
+}
+
+/* Cancel only through tb_domain_cancel_nhi_reset(), outside tb->lock. */
+static void nhi_reset_work(struct work_struct *work)
+{
+ struct tb_nhi *nhi = container_of(to_delayed_work(work), struct tb_nhi,
+ reset_work);
+ struct tb *tb = dev_get_drvdata(nhi->dev);
+
+ /* The connection manager must be blocked over the reset */
+ guard(mutex)(&tb->lock);
+
+ scoped_guard(spinlock_irq, &nhi->lock) {
+ if (nhi->going_away || nhi->removing || nhi->suspended)
+ return;
+ if (bitmap_empty(nhi->dma_hops_used, nhi_dma_hops_bits(nhi)))
+ return;
+ if (nhi_dma_rings_running(nhi))
+ return;
+
+ /* Keep the DMA rings from starting over the reset */
+ nhi->resetting = true;
+ }
+
+ if (!tb_ctl_stop_if_idle(tb->ctl)) {
+ scoped_guard(spinlock_irq, &nhi->lock) {
+ nhi->resetting = false;
+ queue_delayed_work(system_long_wq, &nhi->reset_work,
+ msecs_to_jiffies(100));
+ }
+ return;
+ }
+ nhi_reset_interface(nhi);
+ tb_ctl_start(tb->ctl);
+
+ scoped_guard(spinlock_irq, &nhi->lock) {
+ bitmap_zero(nhi->dma_hops_used, nhi_dma_hops_bits(nhi));
+ nhi->resetting = false;
+ }
+}
+
/**
* tb_ring_start() - enable a ring
* @ring: Ring to start
@@ -752,7 +895,7 @@ int tb_ring_start(struct tb_ring *ring)
spin_lock_irq(&ring->nhi->lock);
spin_lock(&ring->lock);
- if (ring->nhi->going_away) {
+ if (ring->nhi->going_away || ring->nhi->removing) {
ret = -ENODEV;
goto err;
}
@@ -761,6 +904,16 @@ int tb_ring_start(struct tb_ring *ring)
ret = -EBUSY;
goto err;
}
+ if (ring->nhi->dma_hops_used && ring_is_dma(ring) &&
+ (ring->nhi->suspended || ring->nhi->resetting ||
+ ring_needs_reset(ring))) {
+ ret = -EAGAIN;
+ if (!ring->nhi->suspended && !ring->nhi->resetting &&
+ !nhi_dma_rings_running(ring->nhi))
+ queue_delayed_work(system_long_wq,
+ &ring->nhi->reset_work, 0);
+ goto err;
+ }
dev_dbg(ring->nhi->dev, "starting %s %d\n",
RING_TYPE(ring), ring->hop);
@@ -813,6 +966,9 @@ int tb_ring_start(struct tb_ring *ring)
if (!(ring->flags & RING_FLAG_NO_INTERRUPT))
ring_interrupt_active(ring, true);
ring->running = true;
+ if (ring->nhi->dma_hops_used && ring_is_dma(ring))
+ __set_bit(nhi_hop_bit(ring->nhi, ring->is_tx, ring->hop),
+ ring->nhi->dma_hops_used);
err:
spin_unlock(&ring->lock);
spin_unlock_irq(&ring->nhi->lock);
@@ -885,6 +1041,11 @@ void tb_ring_stop(struct tb_ring *ring)
ring->notify_pending = false;
ring->running = false;
+ if (ring->nhi->dma_hops_used && ring_is_dma(ring) &&
+ !ring->nhi->removing && !ring->nhi->suspended &&
+ !nhi_dma_rings_running(ring->nhi))
+ queue_delayed_work(system_long_wq, &ring->nhi->reset_work, 0);
+
err:
spin_unlock(&ring->lock);
spin_unlock_irq(&ring->nhi->lock);
@@ -1118,10 +1279,28 @@ static int nhi_freeze_noirq(struct device *dev)
return tb_domain_freeze_noirq(tb);
}
+static void nhi_reset_if_idle(struct tb_nhi *nhi)
+{
+ scoped_guard(spinlock_irq, &nhi->lock) {
+ if (nhi->going_away || nhi->removing ||
+ !nhi_dma_hops_dirty(nhi) || nhi_dma_rings_running(nhi))
+ return;
+ nhi->resetting = true;
+ }
+
+ nhi_reset_interface(nhi);
+ scoped_guard(spinlock_irq, &nhi->lock) {
+ bitmap_zero(nhi->dma_hops_used, nhi_dma_hops_bits(nhi));
+ nhi->resetting = false;
+ }
+}
+
static int nhi_thaw_noirq(struct device *dev)
{
struct tb *tb = dev_get_drvdata(dev);
+ /* Freeze can preserve live DMA rings. */
+ nhi_reset_if_idle(tb->nhi);
return tb_domain_thaw_noirq(tb);
}
@@ -1166,6 +1345,8 @@ static int nhi_resume_noirq(struct device *dev)
return ret;
}
+ nhi_reset_if_idle(nhi);
+
return tb_domain_resume_noirq(tb);
}
@@ -1221,6 +1402,8 @@ static int nhi_runtime_resume(struct device *dev)
return ret;
}
+ nhi_reset_if_idle(nhi);
+
return tb_domain_runtime_resume(tb);
}
@@ -1318,6 +1501,7 @@ int nhi_probe(struct tb_nhi *nhi)
{
struct device *dev = nhi->dev;
struct tb *tb;
+ u32 caps;
int res;
if (!nhi->ops)
@@ -1326,7 +1510,8 @@ int nhi_probe(struct tb_nhi *nhi)
if (!nhi->ops->init_interrupts)
return dev_err_probe(dev, -EINVAL, "missing required NHI ops\n");
- nhi->hop_count = ioread32(nhi->iobase + REG_CAPS) & 0x3ff;
+ caps = ioread32(nhi->iobase + REG_CAPS);
+ nhi->hop_count = caps & 0x3ff;
dev_dbg(dev, "total paths: %d\n", nhi->hop_count);
nhi->tx_rings = devm_kcalloc(dev, nhi->hop_count,
@@ -1339,6 +1524,22 @@ int nhi_probe(struct tb_nhi *nhi)
if (!nhi->tx_rings || !nhi->rx_rings || !nhi->interrupt_mask)
return -ENOMEM;
+ INIT_DELAYED_WORK(&nhi->reset_work, nhi_reset_work);
+
+ if (nhi->quirks & QUIRK_RESET_DMA_ON_REUSE) {
+ /* Only v1 host interfaces implement the reset */
+ if (FIELD_GET(REG_CAPS_VERSION_MASK, caps) < REG_CAPS_VERSION_2) {
+ nhi->dma_hops_used = devm_bitmap_zalloc(dev,
+ nhi_dma_hops_bits(nhi),
+ GFP_KERNEL);
+ if (!nhi->dma_hops_used)
+ return -ENOMEM;
+ } else {
+ dev_warn(dev,
+ "reset-on-reuse quirk requires a v1 host interface, disabling\n");
+ }
+ }
+
nhi_reset(nhi);
/* In case someone left them on. */
diff --git a/drivers/thunderbolt/nhi.h b/drivers/thunderbolt/nhi.h
index 88af9ab9fde7..13cb94cae658 100644
--- a/drivers/thunderbolt/nhi.h
+++ b/drivers/thunderbolt/nhi.h
@@ -123,6 +123,7 @@ struct tb_nhi_ops {
/* Host interface quirks */
#define QUIRK_AUTO_CLEAR_INT BIT(0)
#define QUIRK_E2E BIT(1)
+#define QUIRK_RESET_DMA_ON_REUSE BIT(2)
/*
* Minimal number of vectors when we use MSI-X. Two for control channel
diff --git a/drivers/thunderbolt/nhi_regs.h b/drivers/thunderbolt/nhi_regs.h
index d6a197fabc74..99df60b6db36 100644
--- a/drivers/thunderbolt/nhi_regs.h
+++ b/drivers/thunderbolt/nhi_regs.h
@@ -115,6 +115,10 @@ struct ring_desc {
#define REG_CAPS_VERSION_MASK GENMASK(23, 16)
#define REG_CAPS_VERSION_2 0x40
+/* Host Interface Reset - resets TX/RX rings and E2E flow control counters */
+#define REG_HOST_INTERFACE_RESET 0x39858
+#define REG_HOST_INTERFACE_RESET_RST BIT(0)
+
#define REG_DMA_MISC 0x39864
#define REG_DMA_MISC_INT_AUTO_CLEAR BIT(2)
#define REG_DMA_MISC_DISABLE_AUTO_CLEAR BIT(17)
diff --git a/include/linux/thunderbolt.h b/include/linux/thunderbolt.h
index 5fce1b67c376..2d5efc54c0b6 100644
--- a/include/linux/thunderbolt.h
+++ b/include/linux/thunderbolt.h
@@ -515,6 +515,21 @@ void tb_service_properties_changed(struct tb_service *svc);
* MSI-X is used.
* @hop_count: Number of rings (end point hops) supported by NHI.
* @quirks: NHI specific quirks if any
+ * @dma_hops_used: Bitmap of the TX and RX DMA HopIDs programmed after the
+ * last host interface reset. Only with
+ * %QUIRK_RESET_DMA_ON_REUSE.
+ * @resetting: The host interface reset is in progress. Only with
+ * %QUIRK_RESET_DMA_ON_REUSE.
+ * @removing: The domain this host interface belongs to is being removed.
+ * Set before releasing the domain lock so
+ * reset_work does not restart the control channel while the
+ * domain is tearing down. Only with %QUIRK_RESET_DMA_ON_REUSE.
+ * @suspended: The domain's control channel is stopped for system or runtime PM.
+ * Set before the control channel is stopped so reset_work
+ * does not restart it while the NHI is about to be powered
+ * down. Only with %QUIRK_RESET_DMA_ON_REUSE.
+ * @reset_work: Work that runs the host interface reset once all the DMA
+ * rings are idle. Only with %QUIRK_RESET_DMA_ON_REUSE.
* @domain_released: Completed when domain has been fully released
* @host_reset: Host router was reset on driver load, or forced on system
* shutdown/reboot. When set, tb_stop() asserts DPR on connected
@@ -535,6 +550,11 @@ struct tb_nhi {
struct work_struct interrupt_work;
u32 hop_count;
unsigned long quirks;
+ unsigned long *dma_hops_used;
+ bool resetting;
+ bool removing;
+ bool suspended;
+ struct delayed_work reset_work;
struct completion domain_released;
bool host_reset;
};
--
2.34.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* [PATCH 3/3] thunderbolt: Add quirk to reset host interface for AMD USB4 routers
2026-10-05 13:43 [PATCH 0/3] thunderbolt: Reset affected AMD host interfaces before DMA HopID reuse Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID Basavaraj Natikar
@ 2026-10-05 13:43 ` Basavaraj Natikar
2 siblings, 0 replies; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-05 13:43 UTC (permalink / raw)
To: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni
Cc: linux-usb, linux-doc, Mario Limonciello, Sanath S,
Basavaraj Natikar, Mika Westerberg
Affected AMD USB4 host routers hang a TX ring when a DMA HopID is
programmed again after the DMA paths are torn down without an intervening
host interface reset.
Hence, set QUIRK_RESET_DMA_ON_REUSE for these routers so a used DMA HopID
is not reprogrammed until the host interface has been reset. The reset
waits for all DMA rings to stop, leaving unrelated active tunnels
undisturbed.
Suggested-by: Mika Westerberg <mika.westerberg@linux.intel.com>
Co-developed-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Sanath S <Sanath.S@amd.com>
Signed-off-by: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
---
drivers/thunderbolt/nhi.h | 12 ++++++++++++
drivers/thunderbolt/pci.c | 21 +++++++++++++++++++++
2 files changed, 33 insertions(+)
diff --git a/drivers/thunderbolt/nhi.h b/drivers/thunderbolt/nhi.h
index 13cb94cae658..70d6ee0e960a 100644
--- a/drivers/thunderbolt/nhi.h
+++ b/drivers/thunderbolt/nhi.h
@@ -118,6 +118,18 @@ struct tb_nhi_ops {
#define PCI_DEVICE_ID_INTEL_PTL_P_NHI0 0xe433
#define PCI_DEVICE_ID_INTEL_PTL_P_NHI1 0xe434
+#define PCI_DEVICE_ID_AMD_1AH_M60H_NHI0 0x1120
+#define PCI_DEVICE_ID_AMD_1AH_M60H_NHI1 0x1121
+#define PCI_DEVICE_ID_AMD_1AH_M68H_NHI0 0x113b
+#define PCI_DEVICE_ID_AMD_1AH_M68H_NHI1 0x113c
+#define PCI_DEVICE_ID_AMD_1AH_M80H_NHI0 0x1155
+#define PCI_DEVICE_ID_AMD_1AH_M80H_NHI1 0x1158
+#define PCI_DEVICE_ID_AMD_1AH_M80H_NHI2 0x1159
+#define PCI_DEVICE_ID_AMD_1AH_M24H_NHI0 0x151c
+#define PCI_DEVICE_ID_AMD_1AH_M24H_NHI1 0x151d
+#define PCI_DEVICE_ID_AMD_1AH_M70H_NHI0 0x158d
+#define PCI_DEVICE_ID_AMD_1AH_M70H_NHI1 0x158e
+
#define PCI_CLASS_SERIAL_USB_USB4 0x0c0340
/* Host interface quirks */
diff --git a/drivers/thunderbolt/pci.c b/drivers/thunderbolt/pci.c
index 4408229d8ac7..b26f5b39e923 100644
--- a/drivers/thunderbolt/pci.c
+++ b/drivers/thunderbolt/pci.c
@@ -63,6 +63,27 @@ static void nhi_pci_check_quirks(struct tb_nhi_pci *nhi_pci)
nhi->quirks |= QUIRK_E2E;
break;
}
+ } else if (pdev->vendor == PCI_VENDOR_ID_AMD) {
+ switch (pdev->device) {
+ case PCI_DEVICE_ID_AMD_1AH_M60H_NHI0:
+ case PCI_DEVICE_ID_AMD_1AH_M60H_NHI1:
+ case PCI_DEVICE_ID_AMD_1AH_M68H_NHI0:
+ case PCI_DEVICE_ID_AMD_1AH_M68H_NHI1:
+ case PCI_DEVICE_ID_AMD_1AH_M80H_NHI0:
+ case PCI_DEVICE_ID_AMD_1AH_M80H_NHI1:
+ case PCI_DEVICE_ID_AMD_1AH_M80H_NHI2:
+ case PCI_DEVICE_ID_AMD_1AH_M24H_NHI0:
+ case PCI_DEVICE_ID_AMD_1AH_M24H_NHI1:
+ case PCI_DEVICE_ID_AMD_1AH_M70H_NHI0:
+ case PCI_DEVICE_ID_AMD_1AH_M70H_NHI1:
+ /*
+ * These hosts may hang the Tx ring if its HopID is
+ * programmed again without a host interface reset
+ * in between.
+ */
+ nhi->quirks |= QUIRK_RESET_DMA_ON_REUSE;
+ break;
+ }
}
}
--
2.34.1
^ permalink raw reply related [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-05 13:43 ` [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID Basavaraj Natikar
@ 2026-10-05 14:35 ` Mika Westerberg
2026-10-05 15:12 ` Mario Limonciello
2026-10-05 16:50 ` Basavaraj Natikar
2026-10-06 4:33 ` Mika Westerberg
1 sibling, 2 replies; 10+ messages in thread
From: Mika Westerberg @ 2026-10-05 14:35 UTC (permalink / raw)
To: Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Mario Limonciello, Sanath S
Hi,
On Mon, Oct 05, 2026 at 07:13:40PM +0530, Basavaraj Natikar wrote:
> +++ b/drivers/thunderbolt/nhi.c
> @@ -17,6 +17,7 @@
> #include <linux/iommu.h>
> #include <linux/lockdep.h>
> #include <linux/module.h>
> +#include <linux/pci.h>
I haven't looked at the code at all yet but this one is no-go. We are
trying to make the NHI code generic that can work with any interconnect,
like the controllers on ARM based SoCs (aka platform bus) and that's why we
have pci.c for the PCIe based NHIs.
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-05 14:35 ` Mika Westerberg
@ 2026-10-05 15:12 ` Mario Limonciello
2026-10-05 16:50 ` Basavaraj Natikar
1 sibling, 0 replies; 10+ messages in thread
From: Mario Limonciello @ 2026-10-05 15:12 UTC (permalink / raw)
To: Mika Westerberg, Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Sanath S
On 10/5/26 09:35, Mika Westerberg wrote:
> Hi,
>
> On Mon, Oct 05, 2026 at 07:13:40PM +0530, Basavaraj Natikar wrote:
>> +++ b/drivers/thunderbolt/nhi.c
>> @@ -17,6 +17,7 @@
>> #include <linux/iommu.h>
>> #include <linux/lockdep.h>
>> #include <linux/module.h>
>> +#include <linux/pci.h>
>
> I haven't looked at the code at all yet but this one is no-go. We are
> trying to make the NHI code generic that can work with any interconnect,
> like the controllers on ARM based SoCs (aka platform bus) and that's why we
> have pci.c for the PCIe based NHIs.
To add to that - what tree did you base on? I tried on thunderbolt/next
(a93a8e3200002f0c345fec4c340390e9a60aabc7) and it doesn't apply.
╰─❯ b4 shazam
https://lore.kernel.org/linux-usb/cover.1790854235.git.Basavaraj.Natikar@amd.com/T/#m2e70efa2f627df231d9f975424096587b819a354
Looking up
https://lore.kernel.org/all/cover.1790854235.git.Basavaraj.Natikar@amd.com/
Grabbing thread from
lore.kernel.org/all/cover.1790854235.git.Basavaraj.Natikar@amd.com/t.mbox.gz
Checking for newer revisions
Grabbing search results from lore.kernel.org
Analyzing 5 messages in the thread
Analyzing 0 code-review messages
Checking attestation on all messages, may take a moment...
---
✓ [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail
✓ [PATCH 2/3] thunderbolt: Reset the host interface before reusing a
DMA HopID
✓ [PATCH 3/3] thunderbolt: Add quirk to reset host interface for AMD
USB4 routers
---
✓ Signed: DKIM/amd.com
---
Total patches: 3
---
Applying: thunderbolt: Allow tb_ring_start() to fail
Applying: thunderbolt: Reset the host interface before reusing a DMA HopID
Patch failed at 0002 thunderbolt: Reset the host interface before
reusing a DMA HopID
When you have resolved this problem, run "git am --continue".
If you prefer to skip this patch, run "git am --skip" instead.
To restore the original branch and stop patching, run "git am --abort".
error: patch failed: drivers/thunderbolt/nhi.c:17
error: drivers/thunderbolt/nhi.c: patch does not apply
error: patch failed: drivers/thunderbolt/nhi.h:123
error: drivers/thunderbolt/nhi.h: patch does not apply
error: patch failed: drivers/thunderbolt/nhi_regs.h:115
error: drivers/thunderbolt/nhi_regs.h: patch does not apply
hint: Use 'git am --show-current-patch=diff' to see the failed patch
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-05 14:35 ` Mika Westerberg
2026-10-05 15:12 ` Mario Limonciello
@ 2026-10-05 16:50 ` Basavaraj Natikar
1 sibling, 0 replies; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-05 16:50 UTC (permalink / raw)
To: Mika Westerberg, Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Mario Limonciello, Sanath S
Hi Mika,
On 10/5/2026 8:05 PM, Mika Westerberg wrote:
> Hi,
>
> On Mon, Oct 05, 2026 at 07:13:40PM +0530, Basavaraj Natikar wrote:
>> +++ b/drivers/thunderbolt/nhi.c
>> @@ -17,6 +17,7 @@
>> #include <linux/iommu.h>
>> #include <linux/lockdep.h>
>> #include <linux/module.h>
>> +#include <linux/pci.h>
> I haven't looked at the code at all yet but this one is no-go. We are
> trying to make the NHI code generic that can work with any interconnect,
> like the controllers on ARM based SoCs (aka platform bus) and that's why we
> have pci.c for the PCIe based NHIs.
I rebased the series onto thunderbolt/next
a93a8e3200002f0c345fec4c340390e9a60aabc7. The reworked v2 removes
the PCI-specific dependency from generic nhi.c and uses the existing
reset_interface operation instead.
Thanks,
--
Basavaraj
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail
2026-10-05 13:43 ` [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail Basavaraj Natikar
@ 2026-10-06 4:27 ` Mika Westerberg
0 siblings, 0 replies; 10+ messages in thread
From: Mika Westerberg @ 2026-10-06 4:27 UTC (permalink / raw)
To: Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Mario Limonciello, Sanath S
Hi,
On Mon, Oct 05, 2026 at 07:13:39PM +0530, Basavaraj Natikar wrote:
> tb_ring_start() returns void, so its callers cannot tell when a ring fails
> to start and keep building an unusable tunnel. On some host interfaces a
> DMA HopID also cannot be reprogrammed until the host interface has been
> reset.
>
> Hence, let tb_ring_start() return an error and unwind the callers on
> failure: stop an already started TX ring when its RX peer fails to start,
> and disable the DMA paths enabled before the rings were started.
>
> A stream can also stay open after a failed resume. Therefore, free the
> partial allocations, clear the ring pointers, and let the subsequent I/O
> and close return without touching the freed rings. Check readiness under
> the device mutex and use a wake token so a wakeup is not lost across the
> unlocked sleep.
>
> Co-developed-by: Sanath S <Sanath.S@amd.com>
> Signed-off-by: Sanath S <Sanath.S@amd.com>
> Signed-off-by: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
> ---
>
> drivers/net/thunderbolt/main.c | 14 ++-
> drivers/thunderbolt/ctl.c | 23 ++++-
> drivers/thunderbolt/dma_test.c | 32 +++++-
> drivers/thunderbolt/nhi.c | 12 ++-
> drivers/thunderbolt/stream.c | 175 +++++++++++++++++++++------------
> include/linux/thunderbolt.h | 2 +-
> 6 files changed, 181 insertions(+), 77 deletions(-)
This is pretty invasive change. How well did you test this?
I would have been hoping we can avoid touching the service drivers
completely.
> diff --git a/drivers/net/thunderbolt/main.c b/drivers/net/thunderbolt/main.c
> index cf51b9c39f4e..93ccccc5cf8b 100644
> --- a/drivers/net/thunderbolt/main.c
> +++ b/drivers/net/thunderbolt/main.c
> @@ -669,8 +669,16 @@ static void tbnet_connected_work(struct work_struct *work)
> * the Rx ring before any incoming packets are allowed to
> * arrive.
> */
> - tb_ring_start(net->tx_ring.ring);
> - tb_ring_start(net->rx_ring.ring);
> + ret = tb_ring_start(net->tx_ring.ring);
> + if (ret) {
> + netdev_dbg(net->dev, "failed to start Tx ring, ret=%d\n", ret);
the ret= is not consistent wit the rest of the driver and should this be
_warn() instead?
> + goto err_release_hopid;
> + }
> + ret = tb_ring_start(net->rx_ring.ring);
> + if (ret) {
> + netdev_dbg(net->dev, "failed to start Rx ring, ret=%d\n", ret);
> + goto err_stop_tx;
> + }
>
> ret = tbnet_alloc_rx_buffers(net, TBNET_RING_SIZE);
> if (ret)
> @@ -701,7 +709,9 @@ static void tbnet_connected_work(struct work_struct *work)
> tbnet_free_buffers(&net->rx_ring);
> err_stop_rings:
> tb_ring_stop(net->rx_ring.ring);
> +err_stop_tx:
> tb_ring_stop(net->tx_ring.ring);
> +err_release_hopid:
> tb_xdomain_release_in_hopid(net->xd, net->remote_transmit_path);
> tbnet_connect_failed(net);
> }
> diff --git a/drivers/thunderbolt/ctl.c b/drivers/thunderbolt/ctl.c
> index 965988b18608..94b29430ffe4 100644
> --- a/drivers/thunderbolt/ctl.c
> +++ b/drivers/thunderbolt/ctl.c
> @@ -728,10 +728,27 @@ void tb_ctl_free(struct tb_ctl *ctl)
> */
> void tb_ctl_start(struct tb_ctl *ctl)
> {
> - int i;
> + int i, ret;
> tb_ctl_dbg(ctl, "control channel starting...\n");
> - tb_ring_start(ctl->tx); /* is used to ack hotplug packets, start first */
> - tb_ring_start(ctl->rx);
> +
> + /*
> + * TX is used to ack hotplug packets so start it first. -ENODEV
> + * means the host controller itself is already gone (expected on
> + * an unplug-during-suspend resume), so do not warn about that.
What?
> + */
> + ret = tb_ring_start(ctl->tx);
> + if (ret) {
> + if (ret != -ENODEV)
> + tb_ctl_WARN(ctl, "failed to start TX ring\n");
> + return;
> + }
> + ret = tb_ring_start(ctl->rx);
> + if (ret) {
> + if (ret != -ENODEV)
> + tb_ctl_WARN(ctl, "failed to start RX ring\n");
> + tb_ring_stop(ctl->tx);
> + return;
> + }
> for (i = 0; i < TB_CTL_RX_PKG_COUNT; i++)
> tb_ctl_rx_submit(ctl->rx_packets[i]);
>
> diff --git a/drivers/thunderbolt/dma_test.c b/drivers/thunderbolt/dma_test.c
> index bcecb0edcb81..e9c01bcfedf9 100644
> --- a/drivers/thunderbolt/dma_test.c
> +++ b/drivers/thunderbolt/dma_test.c
> @@ -203,12 +203,36 @@ static int dma_test_start_rings(struct dma_test *dt)
> return ret;
> }
>
> - if (dt->tx_ring)
> - tb_ring_start(dt->tx_ring);
> - if (dt->rx_ring)
> - tb_ring_start(dt->rx_ring);
> + if (dt->tx_ring) {
> + ret = tb_ring_start(dt->tx_ring);
> + if (ret)
> + goto err_disable_paths;
> + }
> + if (dt->rx_ring) {
> + ret = tb_ring_start(dt->rx_ring);
> + if (ret)
> + goto err_stop_tx;
> + }
>
> return 0;
> +
> +err_stop_tx:
> + tb_xdomain_disable_paths(dt->xd, dt->tx_hopid,
> + dt->tx_ring ? dt->tx_ring->hop : -1,
> + dt->rx_hopid,
> + dt->rx_ring ? dt->rx_ring->hop : -1);
> + if (dt->tx_ring)
> + tb_ring_stop(dt->tx_ring);
> + goto err_free;
This is way too fragile :-(
> +err_disable_paths:
> + tb_xdomain_disable_paths(dt->xd, dt->tx_hopid,
> + dt->tx_ring ? dt->tx_ring->hop : -1,
> + dt->rx_hopid,
> + dt->rx_ring ? dt->rx_ring->hop : -1);
> +err_free:
> + dma_test_free_rings(dt);
> +
> + return ret;
> }
>
> static void dma_test_stop_rings(struct dma_test *dt)
> diff --git a/drivers/thunderbolt/nhi.c b/drivers/thunderbolt/nhi.c
> index e99a3fcc4a29..1960fa30e13a 100644
> --- a/drivers/thunderbolt/nhi.c
> +++ b/drivers/thunderbolt/nhi.c
> @@ -741,18 +741,24 @@ EXPORT_SYMBOL_GPL(tb_ring_alloc_rx);
> * @ring: Ring to start
> *
> * Must not be invoked in parallel with tb_ring_stop().
> + *
> + * Returns %0 on success and negative errno in case of failure.
> */
> -void tb_ring_start(struct tb_ring *ring)
> +int tb_ring_start(struct tb_ring *ring)
> {
> u16 frame_size;
> + int ret = 0;
> u32 flags;
>
> spin_lock_irq(&ring->nhi->lock);
> spin_lock(&ring->lock);
> - if (ring->nhi->going_away)
> + if (ring->nhi->going_away) {
> + ret = -ENODEV;
> goto err;
> + }
> if (ring->running) {
> dev_WARN(ring->nhi->dev, "ring already started\n");
> + ret = -EBUSY;
> goto err;
> }
> dev_dbg(ring->nhi->dev, "starting %s %d\n",
> @@ -810,6 +816,8 @@ void tb_ring_start(struct tb_ring *ring)
> err:
> spin_unlock(&ring->lock);
> spin_unlock_irq(&ring->nhi->lock);
> +
> + return ret;
> }
> EXPORT_SYMBOL_GPL(tb_ring_start);
>
> diff --git a/drivers/thunderbolt/stream.c b/drivers/thunderbolt/stream.c
> index 4f9a57b77bfa..48b7e3f01312 100644
> --- a/drivers/thunderbolt/stream.c
> +++ b/drivers/thunderbolt/stream.c
> @@ -259,6 +259,9 @@ static void tbstream_ring_free(struct tbstream_ring *ring)
> enum dma_data_direction dir;
> int i;
>
> + if (!ring->frames)
> + return;
> +
> if (ring->ring->is_tx)
> dir = DMA_TO_DEVICE;
> else
> @@ -279,6 +282,7 @@ static void tbstream_ring_free(struct tbstream_ring *ring)
> ring->prod = 0;
> ring->cons = 0;
> kfree(ring->frames);
> + ring->frames = NULL;
> }
>
> static inline bool tbstream_ring_available(const struct tbstream_ring *ring)
> @@ -575,6 +579,9 @@ static int tbstream_dev_send_close(struct tbstream_dev *sdev)
> struct tbstream_frame *sf;
> ktime_t timeout;
>
> + if (!sdev->tx_ring.ring)
> + return -ESHUTDOWN;
> +
> /*
> * Wait for the ring to have available slots before we send the
> * CLOSE packet.
> @@ -645,7 +652,7 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
>
> ret = tbstream_dev_alloc_tx_buffers(sdev);
> if (ret)
> - goto err_free_tx;
> + goto err_free_tx_buffers;
>
> e2e_tx_hop = ring->hop;
> sof_mask = BIT(TBSTREAM_FRAME_START);
> @@ -672,8 +679,12 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
>
> sdev->rx_pending = false;
>
> - tb_ring_start(sdev->tx_ring.ring);
> - tb_ring_start(sdev->rx_ring.ring);
> + ret = tb_ring_start(sdev->tx_ring.ring);
> + if (ret)
> + goto err_disable_paths;
> + ret = tb_ring_start(sdev->rx_ring.ring);
> + if (ret)
> + goto err_stop_tx;
>
> ret = tbstream_dev_alloc_rx_buffers(sdev);
> if (ret)
> @@ -682,13 +693,20 @@ static int tbstream_dev_start(struct tbstream_dev *sdev)
>
> err_stop:
> tb_ring_stop(sdev->rx_ring.ring);
> + tbstream_ring_free(&sdev->rx_ring);
> +err_stop_tx:
> tb_ring_stop(sdev->tx_ring.ring);
> +err_disable_paths:
> + tb_xdomain_disable_paths(xd, sdev->out_hopid, sdev->tx_ring.ring->hop,
> + sdev->in_hopid, sdev->rx_ring.ring->hop);
> err_free_rx:
> tb_ring_free(sdev->rx_ring.ring);
> + sdev->rx_ring.ring = NULL;
> err_free_tx_buffers:
> tbstream_ring_free(&sdev->tx_ring);
> -err_free_tx:
> tb_ring_free(sdev->tx_ring.ring);
> + sdev->tx_ring.ring = NULL;
> + wake_up_interruptible(&sdev->wait);
>
> return ret;
> }
> @@ -708,6 +726,10 @@ static void tbstream_dev_stop(struct tbstream_dev *sdev)
> {
> struct tb_xdomain *xd;
>
> + /* Starting may have failed and freed the rings already */
> + if (!sdev->tx_ring.ring)
> + return;
> +
> if (sdev->busy_poll) {
> /*
> * When busy polling we must advance the ring ourselves
> @@ -742,6 +764,7 @@ static void tbstream_dev_stop(struct tbstream_dev *sdev)
> tbstream_ring_free(&sdev->tx_ring);
> tb_ring_free(sdev->tx_ring.ring);
> sdev->tx_ring.ring = NULL;
> + wake_up_interruptible(&sdev->wait);
> }
>
> /* Use only with read_iter/write_iter() to handle nowait */
> @@ -757,29 +780,6 @@ static int tbstream_dev_lock(struct tbstream_dev *sdev, bool nowait)
> return 0;
> }
>
> -/* Must not be called with @sdev->lock held */
> -static int tbstream_dev_busy_poll_wait(struct tbstream_dev *sdev,
> - struct tbstream_ring *ring)
> -{
> - for (;;) {
> - if (signal_pending(current))
> - return -ERESTARTSYS;
> - if (tb_ring_poll_pending(ring->ring))
> - return 0;
> - /*
> - * For TX ring we need to check the RX side too because
> - * it might have received CLOSE packet.
> - */
> - if (ring == &sdev->tx_ring &&
> - tb_ring_poll_pending(sdev->rx_ring.ring))
> - return 0;
> - if (tbstream_dev_valid(sdev) != 0 ||
> - tbstream_dev_closed(sdev) || tbstream_dev_removed(sdev))
> - return 0;
> - cond_resched();
> - }
> -}
> -
> static bool
> tbstream_dev_has_event(struct tbstream_dev *sdev, struct tbstream_ring *ring)
> {
> @@ -796,6 +796,52 @@ tbstream_dev_has_event(struct tbstream_dev *sdev, struct tbstream_ring *ring)
> return tb_ring_poll_pending(sdev->rx_ring.ring);
> }
>
> +static bool tbstream_dev_ready(struct tbstream_dev *sdev,
> + struct tbstream_ring *ring)
> +{
> + lockdep_assert_held(&sdev->lock);
> +
> + /* Starting may have failed and freed the rings already */
> + if (!sdev->tx_ring.ring)
> + return true;
> +
> + return tbstream_dev_has_event(sdev, ring) ||
> + tbstream_dev_close_received(sdev) ||
> + tb_ring_poll_pending(ring->ring);
> +}
> +
> +/* Must not be called with @sdev->lock held. */
> +static int tbstream_dev_wait(struct tbstream_dev *sdev,
> + struct tbstream_ring *ring)
> +{
> + DEFINE_WAIT_FUNC(wait, woken_wake_function);
woken_wake_function?
> + int ret = 0;
> +
> + add_wait_queue(&sdev->wait, &wait);
> + for (;;) {
> + bool ready;
> +
> + ret = mutex_lock_interruptible(&sdev->lock);
> + if (ret)
> + break;
> + ready = tbstream_dev_ready(sdev, ring);
> + mutex_unlock(&sdev->lock);
> + if (ready)
> + break;
> + if (signal_pending(current)) {
> + ret = -ERESTARTSYS;
> + break;
> + }
> + if (sdev->busy_poll)
> + cond_resched();
> + else
> + /* The wake token bridges the unlocked check-to-sleep gap. */
> + wait_woken(&wait, TASK_INTERRUPTIBLE, MAX_SCHEDULE_TIMEOUT);
> + }
> + remove_wait_queue(&sdev->wait, &wait);
This is also really complex to just handle the error in tb_ring_start(). I
wonder what scenarios were actually tested?
> + return ret;
> +}
> +
> static ssize_t
> tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
> {
> @@ -814,6 +860,11 @@ tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
> return ret;
>
> for (;;) {
> + if (!sdev->tx_ring.ring) {
> + mutex_unlock(&sdev->lock);
> + return -ESHUTDOWN;
Now the ring went away behind the reader?
I don't think this is a good approach to be honest.
> + }
> +
> /* Advance RX completions */
> tbstream_dev_advance_rx(sdev);
>
> @@ -838,16 +889,9 @@ tbstream_dev_fops_read_iter(struct kiocb *kiocb, struct iov_iter *to)
> if (nowait)
> return -EAGAIN;
>
> - if (sdev->busy_poll) {
> - ret = tbstream_dev_busy_poll_wait(sdev, &sdev->rx_ring);
> - if (ret)
> - return ret;
> - } else {
> - ret = wait_event_interruptible(sdev->wait,
> - tbstream_dev_has_event(sdev, &sdev->rx_ring));
> - if (ret)
> - return ret;
> - }
> + ret = tbstream_dev_wait(sdev, &sdev->rx_ring);
> + if (ret)
> + return ret;
>
> ret = tbstream_dev_lock(sdev, nowait);
> if (ret)
> @@ -957,6 +1001,11 @@ tbstream_dev_fops_write_iter(struct kiocb *kiocb, struct iov_iter *from)
> return ret;
>
> for (;;) {
> + if (!sdev->tx_ring.ring) {
> + mutex_unlock(&sdev->lock);
> + return -ESHUTDOWN;
> + }
> +
> /* Advance TX (and RX) completions */
> tbstream_dev_advance_both(sdev);
>
> @@ -984,17 +1033,9 @@ tbstream_dev_fops_write_iter(struct kiocb *kiocb, struct iov_iter *from)
> if (nowait)
> return -EAGAIN;
>
> - if (sdev->busy_poll) {
> - ret = tbstream_dev_busy_poll_wait(sdev, &sdev->tx_ring);
> - if (ret)
> - return ret;
> - } else {
> - ret = wait_event_interruptible(sdev->wait,
> - tbstream_dev_has_event(sdev, &sdev->tx_ring) ||
> - tbstream_dev_close_received(sdev));
> - if (ret)
> - return ret;
> - }
> + ret = tbstream_dev_wait(sdev, &sdev->tx_ring);
> + if (ret)
> + return ret;
>
> ret = tbstream_dev_lock(sdev, nowait);
> if (ret)
> @@ -1041,7 +1082,7 @@ tbstream_dev_fops_poll(struct file *file, struct poll_table_struct *wait)
>
> poll_wait(file, &sdev->wait, wait);
> guard(mutex)(&sdev->lock);
> - if (tbstream_dev_valid(sdev) != 0)
> + if (tbstream_dev_valid(sdev) != 0 || !sdev->tx_ring.ring)
> return EPOLLHUP | EPOLLERR;
>
> /*
> @@ -1105,6 +1146,11 @@ static int tbstream_dev_fops_open(struct inode *inode, struct file *file)
> }
> }
>
> + if (sdev->users && !sdev->tx_ring.ring) {
> + ret = -ESHUTDOWN;
> + goto err_unlock;
> + }
> +
> /* Only on first open we allocate rings and enable paths */
> if (!sdev->users++) {
> ret = tbstream_dev_start(sdev);
> @@ -1136,7 +1182,7 @@ static int tbstream_dev_fops_release(struct inode *inode, struct file *file)
> struct tbstream_dev *sdev = to_tbstream_dev(file->private_data);
>
> mutex_lock(&sdev->lock);
> - if (--sdev->users == 0) {
> + if (--sdev->users == 0 && sdev->tx_ring.ring) {
> /*
> * Advance now in case there is CLOSE waiting in the RX
> * ring.
> @@ -1884,14 +1930,14 @@ static int __maybe_unused tbstream_suspend(struct device *dev)
> if (!sg)
> return 0;
>
> + mutex_lock(&sg->lock);
> list_for_each_entry_reverse(sdev, &sg->dev_list, list) {
> - tbstream_dev_get(sdev);
> - /* Stop the stream (if it was open) */
> + mutex_lock(&sdev->lock);
> if (sdev->users)
> tbstream_dev_stop(sdev);
> - tbstream_dev_put(sdev);
> + mutex_unlock(&sdev->lock);
What's this?
> }
> -
> + mutex_unlock(&sg->lock);
> config_group_put(&sg->group);
> return 0;
> }
> @@ -1902,28 +1948,27 @@ static int __maybe_unused tbstream_resume(struct device *dev)
> struct tbstream *stream = tb_service_get_drvdata(svc);
> struct tbstream_group *sg;
> struct tbstream_dev *sdev;
> + int ret = 0;
>
> sg = tbstream_group_find(stream);
> if (!sg)
> return 0;
>
> + mutex_lock(&sg->lock);
> list_for_each_entry(sdev, &sg->dev_list, list) {
> - tbstream_dev_get(sdev);
> + mutex_lock(&sdev->lock);
> if (sdev->users) {
> - int ret;
> + int err = tbstream_dev_start(sdev);
>
> - ret = tbstream_dev_start(sdev);
> - if (ret) {
> - tbstream_dev_put(sdev);
> - config_group_put(&sg->group);
> - return ret;
> - }
> + if (err && !ret)
> + ret = err;
> }
> - tbstream_dev_put(sdev);
> + mutex_unlock(&sdev->lock);
> + wake_up_interruptible(&sdev->wait);
> }
> -
> + mutex_unlock(&sg->lock);
> config_group_put(&sg->group);
> - return 0;
> + return ret;
> }
>
> static const struct dev_pm_ops tbstream_pm_ops = {
> diff --git a/include/linux/thunderbolt.h b/include/linux/thunderbolt.h
> index 57502da29080..5fce1b67c376 100644
> --- a/include/linux/thunderbolt.h
> +++ b/include/linux/thunderbolt.h
> @@ -672,7 +672,7 @@ struct tb_ring *tb_ring_alloc_rx(struct tb_nhi *nhi, int hop, int size,
> unsigned int flags, int e2e_tx_hop,
> u16 sof_mask, u16 eof_mask,
> void (*start_poll)(void *), void *poll_data);
> -void tb_ring_start(struct tb_ring *ring);
> +int tb_ring_start(struct tb_ring *ring);
> bool tb_ring_flush(struct tb_ring *ring, unsigned int timeout_msec);
> void tb_ring_stop(struct tb_ring *ring);
> void tb_ring_free(struct tb_ring *ring);
> --
> 2.34.1
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-05 13:43 ` [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID Basavaraj Natikar
2026-10-05 14:35 ` Mika Westerberg
@ 2026-10-06 4:33 ` Mika Westerberg
2026-10-06 14:47 ` Basavaraj Natikar
1 sibling, 1 reply; 10+ messages in thread
From: Mika Westerberg @ 2026-10-06 4:33 UTC (permalink / raw)
To: Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Mario Limonciello, Sanath S
Hi,
On Mon, Oct 05, 2026 at 07:13:40PM +0530, Basavaraj Natikar wrote:
> Reusing a DMA HopID without an intervening host interface reset can hang
> the TX ring on some host routers. Resetting on every DMA tunnel teardown
> clears the state but, as the reset affects all rings, also disturbs
> unrelated active tunnels.
>
> Hence, track the DMA HopIDs programmed since the last reset, prefer unused
> HopIDs when allocating rings, and check for reuse at tb_ring_start() too,
> since networking retains its rings across reconnect. Return -EAGAIN instead
> of programming a HopID that still needs a reset.
>
> Run the reset from a work item once all DMA rings are idle: serialize it
> with the connection manager, stop the control channel around it, and block
> DMA rings from starting during the reset. Fence the work across domain
> removal and PM transitions, preserve live DMA rings across freeze/thaw, and
> restore the interrupt-mask shadow under the NHI lock after the reset.
>
> On -EAGAIN, retry the networking login asynchronously and block the work
> producers before teardown cancels the workers. Keep peer disconnection
> separate from administrative shutdown so it cannot reopen the login gate,
> while stream and DMA-test callers unwind immediately and return the error
> to userspace.
>
> Enable this only for reset-capable host interfaces marked with
> QUIRK_RESET_DMA_ON_REUSE. A competing DMA tunnel must stop before its dirty
> HopID can be reused, and retrying does not interrupt that tunnel.
>
> Suggested-by: Mika Westerberg <mika.westerberg@linux.intel.com>
Yeah, I'm not really sure I suggested this :(
My idea was to done it so that it is nicely contained inside nhi.c without
distracting the service drivers. What you are doing is complete opposite of
that.
Given the complexity I would then rather just take the previous quirk with
the deadlock fixed.
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID
2026-10-06 4:33 ` Mika Westerberg
@ 2026-10-06 14:47 ` Basavaraj Natikar
0 siblings, 0 replies; 10+ messages in thread
From: Basavaraj Natikar @ 2026-10-06 14:47 UTC (permalink / raw)
To: Mika Westerberg, Basavaraj Natikar
Cc: Andreas Noever, Mika Westerberg, Yehezkel Bernat, Jonathan Corbet,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, linux-usb, linux-doc, Mario Limonciello, Sanath S
Hi Mika,
On 10/6/2026 10:03 AM, Mika Westerberg wrote:
> Hi,
>
> On Mon, Oct 05, 2026 at 07:13:40PM +0530, Basavaraj Natikar wrote:
>> Reusing a DMA HopID without an intervening host interface reset can hang
>> the TX ring on some host routers. Resetting on every DMA tunnel teardown
>> clears the state but, as the reset affects all rings, also disturbs
>> unrelated active tunnels.
>>
>> Hence, track the DMA HopIDs programmed since the last reset, prefer unused
>> HopIDs when allocating rings, and check for reuse at tb_ring_start() too,
>> since networking retains its rings across reconnect. Return -EAGAIN instead
>> of programming a HopID that still needs a reset.
>>
>> Run the reset from a work item once all DMA rings are idle: serialize it
>> with the connection manager, stop the control channel around it, and block
>> DMA rings from starting during the reset. Fence the work across domain
>> removal and PM transitions, preserve live DMA rings across freeze/thaw, and
>> restore the interrupt-mask shadow under the NHI lock after the reset.
>>
>> On -EAGAIN, retry the networking login asynchronously and block the work
>> producers before teardown cancels the workers. Keep peer disconnection
>> separate from administrative shutdown so it cannot reopen the login gate,
>> while stream and DMA-test callers unwind immediately and return the error
>> to userspace.
>>
>> Enable this only for reset-capable host interfaces marked with
>> QUIRK_RESET_DMA_ON_REUSE. A competing DMA tunnel must stop before its dirty
>> HopID can be reused, and retrying does not interrupt that tunnel.
>>
>> Suggested-by: Mika Westerberg <mika.westerberg@linux.intel.com>
> Yeah, I'm not really sure I suggested this :(
>
> My idea was to done it so that it is nicely contained inside nhi.c without
> distracting the service drivers. What you are doing is complete opposite of
> that.
>
> Given the complexity I would then rather just take the previous quirk with
> the deadlock fixed.
Agreed. The current approach spreads the handling into the service drivers
and is more complex than intended.
Thanks,
--
Basavaraj
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2026-10-06 14:47 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 13:43 [PATCH 0/3] thunderbolt: Reset affected AMD host interfaces before DMA HopID reuse Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 1/3] thunderbolt: Allow tb_ring_start() to fail Basavaraj Natikar
2026-10-06 4:27 ` Mika Westerberg
2026-10-05 13:43 ` [PATCH 2/3] thunderbolt: Reset the host interface before reusing a DMA HopID Basavaraj Natikar
2026-10-05 14:35 ` Mika Westerberg
2026-10-05 15:12 ` Mario Limonciello
2026-10-05 16:50 ` Basavaraj Natikar
2026-10-06 4:33 ` Mika Westerberg
2026-10-06 14:47 ` Basavaraj Natikar
2026-10-05 13:43 ` [PATCH 3/3] thunderbolt: Add quirk to reset host interface for AMD USB4 routers Basavaraj Natikar
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox