Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: Long Li <longli@microsoft.com>
To: Long Li <longli@microsoft.com>,
	Konstantin Taranov <kotaranov@microsoft.com>,
	Jakub Kicinski <kuba@kernel.org>,
	"David S . Miller" <davem@davemloft.net>,
	Paolo Abeni <pabeni@redhat.com>,
	Eric Dumazet <edumazet@google.com>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	Jason Gunthorpe <jgg@ziepe.ca>, Leon Romanovsky <leon@kernel.org>,
	Haiyang Zhang <haiyangz@microsoft.com>,
	"K . Y . Srinivasan" <kys@microsoft.com>,
	Wei Liu <wei.liu@kernel.org>, Dexuan Cui <decui@microsoft.com>,
	shradhagupta@linux.microsoft.com, Simon Horman <horms@kernel.org>,
	ernis@linux.microsoft.com, stephen@networkplumber.org
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [PATCH net v6 0/7] net: mana: HW channel reliability and hardening fixes
Date: Mon, 10 Aug 2026 19:38:14 -0700	[thread overview]
Message-ID: <20260811023823.2391255-1-longli@microsoft.com> (raw)

This series fixes a set of latent bugs and robustness gaps in the MANA
Hardware Channel (HWC), the control path the driver uses to talk to the
device.  The issues range from a use-after-free of completion queues
during teardown to buffer mis-sizing, unsafe teardown ordering, missing
validation of device-supplied RX metadata, and stale-response handling
after a command timeout.

Patch overview:

  1  RCU-protect gc->cq_table lookups against concurrent CQ destroy
     The EQ interrupt handler dereferences CQ pointers from gc->cq_table
     while teardown can free them.  Put the table under RCU and wait a
     grace period before freeing, closing the use-after-free.

  2  fix HWC RQ/SQ buffer size swap
     init_queues() sized the RQ with max_req_msg_size and the SQ with
     max_resp_msg_size -- backwards.  Correct the swap; both sizes are
     equal in practice, so this is a latent-correctness fix.

  3  free HWC comp_buf after destroying the EQ
     Reorder teardown so the EQ is destroyed (readers quiesced) before
     comp_buf and the CQ are freed, preventing a late EQ-handler access
     to freed memory.

  4  validate hardware-supplied values in the HWC RX path
     Bounds-check the SGE, verify the recovered slot index and SGE
     address, and validate response length and msg_id before use, so
     malformed or hostile DMA metadata cannot cause wrong-slot completion
     or out-of-bounds access.

  5  fix HWC teardown safety with setup_active flag and destroy ordering
     Track setup activation explicitly, tear the EQ/CQ down before the
     TXQ/RXQ, and on an unrecoverable teardown failure leak the HWC
     resources rather than free memory the device may still DMA into.

  6  fix stale HWC response after command timeout
     Replace the inflight-slot semaphore with a bitmap + waitqueue and
     per-slot refcount/lock; latch the channel on timeout so no new slots
     are handed out, drop duplicate/late responses, and ignore a zero
     firmware-supplied timeout.

  7  keep max_num_cqs immutable once cq_table is allocated
     gc->max_num_cqs is set once when cq_table is allocated and never
     reset, so a spoofed post-init HWC event cannot inflate the bound
     past the allocation and drive an out-of-bounds cq_table access.

Follow-up feature work (net-next, sent separately):

The original series also contained two patches that are improvements, not
fixes:

  net: mana: support concurrent HWC requests
  net: mana: add dynamic HWC queue depth with reinit path

Per the netdev tree rules, fixes go to 'net' and features/improvements go
to 'net-next', and the two must not be combined in a single submission.
Those two patches build on the locking and teardown groundwork in this
series, so they will be posted as a separate net-next series only after
these fixes have propagated from net into net-next through the usual
periodic merge.

Changes since v5:

Addressed the automated reviews of v5 and a net-next apply conflict
reported by the netdev CI.  No behavioural change to the patch intent;
the deltas below harden the v5 implementation.

 - patch 1: mana_gd_unpublish_cq() and mana_ib_remove_cq_cb() now clear a
   cq_table slot only when it still points at the CQ being torn down, so
   the two-pass netdev teardown cannot wipe an entry a concurrent RDMA CQ
   create legitimately recycled during the grace period.
   mana_gd_process_eqe() drops an already-unpublished (NULL) slot quietly
   instead of tripping WARN_ON_ONCE() during a normal ifdown/MTU change,
   and reads gc->cq_table before gc->max_num_cqs with an smp_rmb()
   between them so a shrinking re-establish cannot pair a stale bound
   with a newly published, smaller table.
 - patch 5: set setup_active before calling mana_smc_setup_hwc(), which
   activates the device MST entries before it can report a late failure;
   on an unrecoverable teardown failure keep the HWC context reachable
   and retry the teardown on the next bring-up instead of orphaning it.
 - patch 6: initialise the caller_ctx refcount/state before publishing
   the inflight bitmap bit so a racing or forged response cannot observe
   an uninitialised slot; honour a response only while the sender still
   owns the slot (output_buf published and not yet reclaimed), so a
   premature response cannot free the slot while its command is still in
   flight; do not latch hwc_timed_out for the deliberate no-wait teardown
   (hwc_timeout == 0), applied on both the admission gate and the
   post-wait check in mana_hwc_get_msg_index().
 - Moved the max_num_cqs/cq_table lifetime documentation onto the
   cq_table field so the series no longer rewords the comment above
   max_num_cqs, which a pending net-next change (cqe8_coalescing_sup)
   also touches; this resolves the netdev CI net-next apply conflict.
   Comment-only, no functional change.
 - Rebased onto current net/main.
 - v5: https://lore.kernel.org/netdev/20260809190047.2086368-1-longli@microsoft.com/

Changes since v4:

 - No code changes.  Resending as a standalone thread; v4 was
   accidentally sent in-reply-to the v3 thread.
 - v4: https://lore.kernel.org/netdev/20260808023417.1746886-1-longli@microsoft.com/

Changes since v3:

Addressed the netdev-ai and sashiko.dev automated reviews of v3.

 - New patch 7 ("keep max_num_cqs immutable once cq_table is
   allocated"): gc->max_num_cqs is set once and never reset, so a
   spoofed post-init HWC event cannot inflate the bound past the
   allocation and cause an out-of-bounds cq_table access.
 - patch 1: replaced the per-CQ synchronize_rcu() in the netdev teardown
   paths with a two-pass quiesce/free that takes a single grace period
   per teardown; snapshot cq->id and max_num_cqs with READ_ONCE() so the
   same value sizes, bounds and indexes cq_table.
 - patch 4: removed the short-response early return so a malformed
   response reaches verify_resp_msg() -> -EPROTO and completes the
   sender instead of hanging it; account leaked RX WQEs and trip
   hwc_timeout on RQ exhaustion; read the device-supplied
   inline_oob_size_div4 and sge->address with READ_ONCE() and reject any
   value other than the one the driver programs.
 - patch 5: arm setup_active after mana_smc_setup_hwc() succeeds; destroy
   the EQ (IRQ deregister + drain) before the CQ; drop the redundant
   teardown in mana_hwc_establish_channel() that caused a double hardware
   timeout and masked the original error.
 - patch 6: take both the sender and response-side references up front in
   mana_hwc_get_msg_index() so an early/stale/forged response cannot free
   the slot before the sender posts; changed caller_ctx::error from u32
   to int; reject a zero firmware-supplied HWC timeout in the query path
   as well as the reconfig path.
 - v3: https://lore.kernel.org/netdev/20260803234355.636038-1-longli@microsoft.com/

Changes since v2:

 - Per maintainer feedback, split the original combined series: the
   fixes here target 'net'; the two feature patches now go to 'net-next'
   and are sent separately (see above).  Rebased the fixes onto net.
 - Dropped the pcie_flr()-based reset fallback from the teardown path;
   on an unrecoverable teardown failure the driver now leaks the HWC
   resources instead of touching memory the device may still DMA into.

Long Li (7):
  net: mana: RCU-protect gc->cq_table lookups against concurrent CQ
    destroy
  net: mana: fix HWC RQ/SQ buffer size swap
  net: mana: free HWC comp_buf after destroying the EQ
  net: mana: validate hardware-supplied values in the HWC RX path
  net: mana: fix HWC teardown safety with setup_active flag and destroy
    ordering
  net: mana: fix stale HWC response after command timeout
  net: mana: keep max_num_cqs immutable once cq_table is allocated

 drivers/infiniband/hw/mana/cq.c               |  51 +-
 .../net/ethernet/microsoft/mana/gdma_main.c   |  72 ++-
 .../net/ethernet/microsoft/mana/hw_channel.c  | 521 +++++++++++++++---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 136 ++++-
 include/net/mana/gdma.h                       |  37 +-
 include/net/mana/hw_channel.h                 |  44 +-
 6 files changed, 743 insertions(+), 118 deletions(-)


base-commit: f1b3416ceaf7ca4cb5cbd986ee8fe3ffaeda2d48
-- 
2.43.0


             reply	other threads:[~2026-08-11  2:38 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-11  2:38 Long Li [this message]
2026-08-11  2:38 ` [PATCH net v6 1/7] net: mana: RCU-protect gc->cq_table lookups against concurrent CQ destroy Long Li
2026-08-11  8:18   ` Leon Romanovsky
2026-08-11  2:38 ` [PATCH net v6 2/7] net: mana: fix HWC RQ/SQ buffer size swap Long Li
2026-08-11  2:38 ` [PATCH net v6 3/7] net: mana: free HWC comp_buf after destroying the EQ Long Li
2026-08-11  2:38 ` [PATCH net v6 4/7] net: mana: validate hardware-supplied values in the HWC RX path Long Li
2026-08-11  2:38 ` [PATCH net v6 5/7] net: mana: fix HWC teardown safety with setup_active flag and destroy ordering Long Li
2026-08-11  2:38 ` [PATCH net v6 6/7] net: mana: fix stale HWC response after command timeout Long Li
2026-08-11  2:38 ` [PATCH net v6 7/7] net: mana: keep max_num_cqs immutable once cq_table is allocated Long Li

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260811023823.2391255-1-longli@microsoft.com \
    --to=longli@microsoft.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=decui@microsoft.com \
    --cc=edumazet@google.com \
    --cc=ernis@linux.microsoft.com \
    --cc=haiyangz@microsoft.com \
    --cc=horms@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=kotaranov@microsoft.com \
    --cc=kuba@kernel.org \
    --cc=kys@microsoft.com \
    --cc=leon@kernel.org \
    --cc=linux-hyperv@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=shradhagupta@linux.microsoft.com \
    --cc=stephen@networkplumber.org \
    --cc=wei.liu@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox