* [PATCH net-next 0/2] net/smc: fix v2 slot clearing and reduce TX slot contention @ 2026-09-10 10:44 D. Wythe 2026-09-10 10:44 ` [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() D. Wythe 2026-09-10 10:44 ` [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait D. Wythe 0 siblings, 2 replies; 6+ messages in thread From: D. Wythe @ 2026-09-10 10:44 UTC (permalink / raw) To: mjambigi, wenjia, wintera, dust.li, tonylu, guwen Cc: kuba, davem, netdev, linux-s390, linux-rdma, leonro, pabeni, edumazet, sidraya, jaka, oliver.yang This series contains the two reviewed patches from the previously posted "net/smc: transition to RDMA core CQ pooling" series (v4), reposted as a standalone series so they can land independently. The remaining CQ pooling patch will be reposted separately after these two are merged. Patch 1 fixes smc_wr_tx_put_slot() to clear the v2 pending slot and buffer structures instead of the pointer variables, the memset targets were the 8-byte pointers themselves so the structures were never actually cleared. Patch 2 reduces TX slot contention by switching TX slot allocation from non-exclusive wait_event() to prepare_to_wait_exclusive(), avoiding thundering-herd wakes when slots are scarce. For patch 2, uperf numbers are now included in the commit message, as requested by Mahanta during v1 review. The short version is that the gain tracks how often the TX slot wait path is actually taken: with the default sysctl settings, where a link group multiplexes many connections over a small send queue, throughput improves by 134% on the 200x1000 request/response workload and by 458% on the 1-byte ping-pong workload; with a tuned configuration where slots are rarely exhausted, the change is neutral (+3.4% and +1.0%, i.e. within noise). See patch 2 for the full table. Link: https://lore.kernel.org/netdev/20260721175309.321b6503@kernel.org/ Link: https://lore.kernel.org/netdev/20260716113745.65234-1-alibuda@linux.alibaba.com/ Link: https://lore.kernel.org/netdev/20260821091702.21458-2-alibuda@linux.alibaba.com/ D. Wythe (2): net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() net/smc: reduce TX slot contention with exclusive wait net/smc/smc_wr.c | 44 ++++++++++++++++++++++++++++++-------------- 1 file changed, 30 insertions(+), 14 deletions(-) -- 2.45.0 ^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() 2026-09-10 10:44 [PATCH net-next 0/2] net/smc: fix v2 slot clearing and reduce TX slot contention D. Wythe @ 2026-09-10 10:44 ` D. Wythe 2026-09-11 10:45 ` sashiko-bot 2026-09-10 10:44 ` [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait D. Wythe 1 sibling, 1 reply; 6+ messages in thread From: D. Wythe @ 2026-09-10 10:44 UTC (permalink / raw) To: mjambigi, wenjia, wintera, dust.li, tonylu, guwen Cc: kuba, davem, netdev, linux-s390, linux-rdma, leonro, pabeni, edumazet, sidraya, jaka, oliver.yang smc_wr_tx_put_slot() tries to reset the v2 pending slot and buffer with memset(&link->wr_tx_v2_pend, 0, sizeof(link->wr_tx_v2_pend)) and the equivalent for wr_tx_buf_v2. Both are pointers, so this zeroes the 8-byte pointer variable instead of the structure it points to. The pending slot and buffer are therefore never actually cleared, and the pointers get overwritten with NULL. Pass the pointers directly and use sizeof(*pointer) so the intended structures are cleared. Fixes: 8799e310fb3f ("net/smc: add v2 support to the work request layer") Signed-off-by: D. Wythe <alibuda@linux.alibaba.com> Reviewed-by: Dust Li <dust.li@linux.alibaba.com> Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com> --- net/smc/smc_wr.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/net/smc/smc_wr.c b/net/smc/smc_wr.c index 97ba46893b17..def2ab84b0c7 100644 --- a/net/smc/smc_wr.c +++ b/net/smc/smc_wr.c @@ -288,10 +288,10 @@ int smc_wr_tx_put_slot(struct smc_link *link, } else if (link->lgr->smc_version == SMC_V2 && pend->idx == link->wr_tx_cnt) { /* Large v2 buffer */ - memset(&link->wr_tx_v2_pend, 0, - sizeof(link->wr_tx_v2_pend)); - memset(&link->lgr->wr_tx_buf_v2, 0, - sizeof(link->lgr->wr_tx_buf_v2)); + memset(link->wr_tx_v2_pend, 0, + sizeof(*link->wr_tx_v2_pend)); + memset(link->lgr->wr_tx_buf_v2, 0, + sizeof(*link->lgr->wr_tx_buf_v2)); return 1; } -- 2.45.0 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() 2026-09-10 10:44 ` [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() D. Wythe @ 2026-09-11 10:45 ` sashiko-bot 0 siblings, 0 replies; 6+ messages in thread From: sashiko-bot @ 2026-09-11 10:45 UTC (permalink / raw) To: D. Wythe Cc: Christian Borntraeger, Vasily Gorbik, linux-s390, Heiko Carstens, linux-rdma, Alexander Gordeev Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: Pre-existing issues: - [High] DMA use-after-free/unmap during link teardown due to ignored v2 TX pending state. - [High] TOCTOU data race condition in V2 slot management causes data corruption of outgoing LLC messages due to lockless memset. -- commit 4dd2538a0566b239179ada14640f3f763de057d5 Author: D. Wythe <alibuda@linux.alibaba.com> net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() This commit fixes a bug where pointers were being zeroed instead of the structures they point to when clearing the v2 pending slot and buffer. By passing the pointers directly and using sizeof(*pointer), it ensures the intended structures are actually zeroed. > diff --git a/net/smc/smc_wr.c b/net/smc/smc_wr.c > index 97ba46893b17f..def2ab84b0c71 100644 > --- a/net/smc/smc_wr.c > +++ b/net/smc/smc_wr.c > @@ -288,10 +288,10 @@ int smc_wr_tx_put_slot(struct smc_link *link, > } else if (link->lgr->smc_version == SMC_V2 && > pend->idx == link->wr_tx_cnt) { > /* Large v2 buffer */ > - memset(&link->wr_tx_v2_pend, 0, > - sizeof(link->wr_tx_v2_pend)); > - memset(&link->lgr->wr_tx_buf_v2, 0, > - sizeof(link->lgr->wr_tx_buf_v2)); > + memset(link->wr_tx_v2_pend, 0, > + sizeof(*link->wr_tx_v2_pend)); > + memset(link->lgr->wr_tx_buf_v2, 0, > + sizeof(*link->lgr->wr_tx_buf_v2)); [Severity: High] This isn't a bug introduced by this patch, but since the memset is now correctly clearing the struct, could this create a data race with smc_wr_tx_get_v2_slot()? The first memset clears the entire smc_wr_tx_pend structure, which includes the idx field. Once idx becomes 0, a concurrent thread calling smc_wr_tx_get_v2_slot() might locklessly see that the slot is available since it checks: if (link->wr_tx_v2_pend->idx == link->wr_tx_cnt) return -EBUSY; If the concurrent thread grabs the slot and starts writing data into the payload buffer, the second memset here will then overwrite that newly claimed buffer with zeroes. Could this cause the hardware to transmit a corrupted or empty message? [Severity: High] This is also a pre-existing issue, but while reviewing the V2 TX pending state logic, could there be a DMA use-after-free during link teardown? During teardown, smc_wr_free_link() waits for pending sends to finish by calling smc_wr_tx_wait_no_pending_sends(), which relies on: static inline bool smc_wr_is_tx_pend(struct smc_link *link) { ... return !bitmap_empty(link->wr_tx_mask, link->wr_tx_cnt); } This only checks the wr_tx_mask bitmap. However, the V2 TX slot does not use the mask; its pending state is tracked by link->wr_tx_v2_pend->idx. If only a V2 TX request is in flight, would smc_wr_is_tx_pend() return false and allow smc_wr_tx_wait_no_pending_sends() to return early? This could allow subsequent DMA unmap operations to free memory while the hardware is still transmitting the V2 message. > return 1; > } -- Sashiko AI review · https://sashiko.dev/#/patchset/20260910104446.41423-1-alibuda@linux.alibaba.com?part=1 ^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait 2026-09-10 10:44 [PATCH net-next 0/2] net/smc: fix v2 slot clearing and reduce TX slot contention D. Wythe 2026-09-10 10:44 ` [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() D. Wythe @ 2026-09-10 10:44 ` D. Wythe 2026-09-11 3:54 ` Mahanta Jambigi 2026-09-11 10:45 ` sashiko-bot 1 sibling, 2 replies; 6+ messages in thread From: D. Wythe @ 2026-09-10 10:44 UTC (permalink / raw) To: mjambigi, wenjia, wintera, dust.li, tonylu, guwen Cc: kuba, davem, netdev, linux-s390, linux-rdma, leonro, pabeni, edumazet, sidraya, jaka, oliver.yang smc_wr_tx_get_free_slot() waits for a free TX slot with wait_event_interruptible_timeout(). Since the wait_event family enqueues waiters as non-exclusive, wake_up() may wake multiple waiters even though only one can use the slot, causing thundering-herd contention when slots are scarce. Use an exclusive wait loop with prepare_to_wait_exclusive() so wake_up() wakes only one waiter per freed slot. smc_wr_wakeup_tx_wait() still uses wake_up_all() during link teardown, so teardown behavior is unchanged. This also corrects the return value on a pending signal: the previous wait_event_interruptible_timeout() path fell through to the "no free slot" case and returned -EPIPE, masking the signal as a connection error. The open-coded loop now returns -ERESTARTSYS, matching the standard interruptible-wait semantics and letting the syscall restart machinery handle it. Performance =========== Measured with uperf between two peers over SMC-R. The benefit depends on how often the TX slot wait path is actually taken. With the default settings, where many connections share a link group and the send queue is small, slots are scarce and the wait path is hot: net.smc.smcr_max_conns_per_lgr = 255 net.smc.smcr_max_send_wr = 16 net.smc.smcr_max_recv_wr = 48 workload baseline patched delta --------------------------------------------------------- rr1c-200x1000-50.xml 655.06 Mb/s 1.53 Gb/s +134% rr1c-1x1-250.xml 371.03 Kb/s 2.07 Mb/s +458% With a tuned configuration, where slots are mostly available and the wait path is rarely entered, the change is neutral to slightly positive: net.smc.smcr_max_conns_per_lgr = 32 net.smc.smcr_max_send_wr = 64 net.smc.smcr_max_recv_wr = 64 workload baseline patched delta --------------------------------------------------------- rr1c-200x1000-50.xml 1.74 Gb/s 1.80 Gb/s +3.4% rr1c-1x1-250.xml 3.11 Mb/s 3.14 Mb/s +1.0% So the change does not regress the uncontended case, and recovers most of the throughput lost to thundering-herd wakeups once slots become scarce. Signed-off-by: D. Wythe <alibuda@linux.alibaba.com> Reviewed-by: Wen Gu <guwen@linux.alibaba.com> Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com> --- net/smc/smc_wr.c | 36 ++++++++++++++++++++++++++---------- 1 file changed, 26 insertions(+), 10 deletions(-) diff --git a/net/smc/smc_wr.c b/net/smc/smc_wr.c index def2ab84b0c7..64413008f9c1 100644 --- a/net/smc/smc_wr.c +++ b/net/smc/smc_wr.c @@ -198,11 +198,13 @@ int smc_wr_tx_get_free_slot(struct smc_link *link, struct smc_rdma_wr **wr_rdma_buf, struct smc_wr_tx_pend_priv **wr_pend_priv) { + unsigned long timeout = SMC_WR_TX_WAIT_FREE_SLOT_TIME; struct smc_link_group *lgr = smc_get_lgr(link); struct smc_wr_tx_pend *wr_pend; u32 idx = link->wr_tx_cnt; struct ib_send_wr *wr_ib; u64 wr_id; + DEFINE_WAIT(wait); int rc; *wr_buf = NULL; @@ -212,17 +214,31 @@ int smc_wr_tx_get_free_slot(struct smc_link *link, if (rc) return rc; } else { - rc = wait_event_interruptible_timeout( - link->wr_tx_wait, - !smc_link_sendable(link) || - lgr->terminating || - (smc_wr_tx_get_free_slot_index(link, &idx) != -EBUSY), - SMC_WR_TX_WAIT_FREE_SLOT_TIME); - if (!rc) { - /* timeout - terminate link */ - smcr_link_down_cond_sched(link); - return -EPIPE; + rc = 0; + for (;;) { + prepare_to_wait_exclusive(&link->wr_tx_wait, &wait, + TASK_INTERRUPTIBLE); + if (!smc_link_sendable(link) || lgr->terminating || + smc_wr_tx_get_free_slot_index(link, &idx) != -EBUSY) + break; + timeout = schedule_timeout(timeout); + /* re-check */ + if (!smc_link_sendable(link) || lgr->terminating || + smc_wr_tx_get_free_slot_index(link, &idx) != -EBUSY) + break; + if (!timeout) { + /* timeout - terminate link */ + smcr_link_down_cond_sched(link); + break; + } + if (signal_pending(current)) { + rc = -ERESTARTSYS; + break; + } } + finish_wait(&link->wr_tx_wait, &wait); + if (rc) + return rc; if (idx == link->wr_tx_cnt) return -EPIPE; } -- 2.45.0 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait 2026-09-10 10:44 ` [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait D. Wythe @ 2026-09-11 3:54 ` Mahanta Jambigi 2026-09-11 10:45 ` sashiko-bot 1 sibling, 0 replies; 6+ messages in thread From: Mahanta Jambigi @ 2026-09-11 3:54 UTC (permalink / raw) To: D. Wythe, wenjia, wintera, dust.li, tonylu, guwen Cc: kuba, davem, netdev, linux-s390, linux-rdma, leonro, pabeni, edumazet, sidraya, jaka, oliver.yang On 10/09/26 4:14 pm, D. Wythe wrote: > smc_wr_tx_get_free_slot() waits for a free TX slot with > wait_event_interruptible_timeout(). Since the wait_event family > enqueues waiters as non-exclusive, wake_up() may wake multiple > waiters even though only one can use the slot, causing > thundering-herd contention when slots are scarce. > > Use an exclusive wait loop with prepare_to_wait_exclusive() so > wake_up() wakes only one waiter per freed slot. > smc_wr_wakeup_tx_wait() still uses wake_up_all() during link > teardown, so teardown behavior is unchanged. > > This also corrects the return value on a pending signal: the previous > wait_event_interruptible_timeout() path fell through to the "no free > slot" case and returned -EPIPE, masking the signal as a connection > error. The open-coded loop now returns -ERESTARTSYS, matching the > standard interruptible-wait semantics and letting the syscall restart > machinery handle it. > > Performance > =========== > > Measured with uperf between two peers over SMC-R. The benefit depends > on how often the TX slot wait path is actually taken. > > With the default settings, where many connections share a link group > and the send queue is small, slots are scarce and the wait path is hot: > > net.smc.smcr_max_conns_per_lgr = 255 > net.smc.smcr_max_send_wr = 16 > net.smc.smcr_max_recv_wr = 48 > > workload baseline patched delta > --------------------------------------------------------- > rr1c-200x1000-50.xml 655.06 Mb/s 1.53 Gb/s +134% > rr1c-1x1-250.xml 371.03 Kb/s 2.07 Mb/s +458% > > With a tuned configuration, where slots are mostly available and the > wait path is rarely entered, the change is neutral to slightly positive: > > net.smc.smcr_max_conns_per_lgr = 32 > net.smc.smcr_max_send_wr = 64 > net.smc.smcr_max_recv_wr = 64 > > workload baseline patched delta > --------------------------------------------------------- > rr1c-200x1000-50.xml 1.74 Gb/s 1.80 Gb/s +3.4% > rr1c-1x1-250.xml 3.11 Mb/s 3.14 Mb/s +1.0% > > So the change does not regress the uncontended case, and recovers most > of the throughput lost to thundering-herd wakeups once slots become > scarce. > > Signed-off-by: D. Wythe <alibuda@linux.alibaba.com> > Reviewed-by: Wen Gu <guwen@linux.alibaba.com> > Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com> I don't see a changelog describing what has changed since the previous version. Any Sashiko comments were addressed? Could you explain why this is being reposted? ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait 2026-09-10 10:44 ` [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait D. Wythe 2026-09-11 3:54 ` Mahanta Jambigi @ 2026-09-11 10:45 ` sashiko-bot 1 sibling, 0 replies; 6+ messages in thread From: sashiko-bot @ 2026-09-11 10:45 UTC (permalink / raw) To: D. Wythe Cc: Vasily Gorbik, Alexander Gordeev, Christian Borntraeger, linux-rdma, linux-s390, Heiko Carstens > smc_wr_tx_get_free_slot() waits for a free TX slot with > wait_event_interruptible_timeout(). Since the wait_event family > enqueues waiters as non-exclusive, wake_up() may wake multiple > waiters even though only one can use the slot, causing > thundering-herd contention when slots are scarce. > > Use an exclusive wait loop with prepare_to_wait_exclusive() so > wake_up() wakes only one waiter per freed slot. > smc_wr_wakeup_tx_wait() still uses wake_up_all() during link > teardown, so teardown behavior is unchanged. > > This also corrects the return value on a pending signal: the previous > wait_event_interruptible_timeout() path fell through to the "no free > slot" case and returned -EPIPE, masking the signal as a connection > error. The open-coded loop now returns -ERESTARTSYS, matching the > [ ... ] > So the change does not regress the uncontended case, and recovers most > of the throughput lost to thundering-herd wakeups once slots become > scarce. > > Signed-off-by: D. Wythe <alibuda@linux.alibaba.com> Sashiko has reviewed this patch and found no issues. It looks great! -- Sashiko AI review · https://sashiko.dev/#/patchset/20260910104446.41423-1-alibuda@linux.alibaba.com?part=2 ^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-11 10:45 UTC | newest] Thread overview: 6+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-10 10:44 [PATCH net-next 0/2] net/smc: fix v2 slot clearing and reduce TX slot contention D. Wythe 2026-09-10 10:44 ` [PATCH net-next 1/2] net/smc: clear the correct v2 slot and buffer in smc_wr_tx_put_slot() D. Wythe 2026-09-11 10:45 ` sashiko-bot 2026-09-10 10:44 ` [PATCH net-next 2/2] net/smc: reduce TX slot contention with exclusive wait D. Wythe 2026-09-11 3:54 ` Mahanta Jambigi 2026-09-11 10:45 ` sashiko-bot
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox