Linux NFS development
 help / color / mirror / Atom feed
From: Benjamin Coddington <ben.coddington@hammerspace.com>
To: Trond Myklebust <trondmy@kernel.org>, Anna Schumaker <anna@kernel.org>
Cc: linux-nfs@vger.kernel.org,
	Jonathan Curley <jcurley@purestorage.com>,
	Mike Snitzer <snitzer@kernel.org>,
	Jeff Layton <jlayton@kernel.org>,
	Junrui Luo <moonafterrain@outlook.com>
Subject: [PATCH v3 15/24] NFSv4/flexfiles: Honor ndc_immediate on CB_NOTIFY_DEVICEID CHANGE
Date: Fri,  4 Sep 2026 12:53:14 -0400	[thread overview]
Message-ID: <47cdf067fb0582ff5cbcd229f3781787095d9701.1788530385.git.bcodding@hammerspace.com> (raw)
In-Reply-To: <cover.1788530385.git.bcodding@hammerspace.com>

When a CHANGE notification carries ndc_immediate, RFC 8881 Section
20.12 says the change is enforced immediately and the client might not
be able to complete pending I/O.  In addition to un-pinning the stripe's
device node, mark the old node unavailable.

Marking does not recall the references already handed out.  A write
whose DS connection is already up keeps using the old node until that
I/O errors: nfs4_ff_layout_prepare_ds() returns early on a live
ds_clp, and the unavailable flag is only consulted when a connection
is being established.  What the mark does change is that a read skips
the node while another mirror is usable, and that an IOMODE_RW segment
still pinning it stops counting as fully available -- so I/O the
server rejects falls back to the MDS rather than being retried against
a mapping the server has already withdrawn.

The walk can also exchange out a node that already carries the new
mapping: a re-resolve that completed between the unhash and the walk
reaching the stripe, raced by this walk or by the walk of a later
CHANGE for the same deviceid.  Ripping such a node out is harmless (it
is still hashed, so the next I/O re-pins it from the cache), but it
must not be marked unavailable.  A superseded node is distinguishable
by hashed-ness: it was unhashed before its walk began and is never
re-inserted, while a fresh node is inserted before it is installed.
Only mark nodes that are no longer hashed.

That test is hlist_unhashed_lockless(): the hook runs under the layout
inode's i_lock and rcu_read_lock(), but not under nfs4_deviceid_lock,
which is what serializes the writers of node.pprev -- and __hlist_del()
stores a neighbour's pprev with WRITE_ONCE(), so removing any other
entry in the same bucket can write the field this test reads.

Without ndc_immediate, pending I/O drains on the old mapping and only
new I/O re-resolves, as before.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
 fs/nfs/flexfilelayout/flexfilelayout.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index a7fda4422c5f..9caf4b3b45ae 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -2560,6 +2560,13 @@ static void ff_layout_reresolve_deviceid(struct pnfs_layout_hdr *lo,
 				kfree(put);
 				continue;
 			}
+			/* A node still hashed was fetched after the unhash
+			 * and carries the new mapping; mark only the
+			 * superseded ones.
+			 */
+			if (immediate &&
+			    hlist_unhashed_lockless(&old->id_node.node))
+				nfs4_mark_deviceid_unavailable(&old->id_node);
 			put->dev = &old->id_node;
 			list_add(&put->node, head);
 		}
-- 
2.53.0


  parent reply	other threads:[~2026-09-04 16:53 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 16:52 [PATCH v3 00/24] NFS: flexfiles device notifications and caching for wide striped layouts Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 01/24] NFSv4/pnfs: Free the netid when draining a data-server address list Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 02/24] NFSv4/flexfiles: Use the full 64-bit stripe_unit Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 03/24] NFSv4/pnfs: bound the CB_NOTIFY_DEVICEID array count before allocating Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 04/24] pNFS: Fix CB_NOTIFY_DEVICEID CHANGE to consume ndc_immediate Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 05/24] NFSv4/flexfiles: Use the full 64-bit offset for read DS selection Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 06/24] NFSv4/flexfiles: Bound page coalescing on the absolute stripe offset Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 07/24] NFSv4/filelayout: Anchor page coalescing on pattern_offset Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 08/24] NFSv4/flexfiles: Reference the device node across DS setup Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 09/24] NFSv4/flexfiles: Carry the device node reference across each I/O Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 10/24] NFSv4/flexfiles: Hold a device node reference for layoutstats encoding Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 11/24] NFSv4/flexfiles: Make the pinned device node pointer RCU-managed Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 12/24] pNFS: Add a reresolve_deviceid layout driver hook Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 13/24] NFSv4/flexfiles: Implement in-place device re-resolve on CHANGE Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 14/24] NFSv4: Dispatch CB_NOTIFY_DEVICEID CHANGE to an in-place refresh Benjamin Coddington
2026-09-04 16:53 ` Benjamin Coddington [this message]
2026-09-04 16:53 ` [PATCH v3 16/24] pNFS: Discard a GETDEVICEINFO reply that raced a CHANGE notification Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 17/24] pNFS: Add deviceid reference query and collection walkers Benjamin Coddington
2026-09-10 16:58   ` Anna Schumaker
2026-09-10 17:24     ` Benjamin Coddington
2026-09-10 18:05       ` Anna Schumaker
2026-09-04 16:53 ` [PATCH v3 18/24] NFSv4/pnfs: Recover revoked layouts on a deleted deviceID Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 19/24] NFSv4/pnfs: Confirm a deviceID delete via GETDEVICEINFO Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 20/24] NFSv4/pnfs: Dispatch CB_NOTIFY_DEVICEID DELETE to race recovery Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 21/24] NFSv4/pnfs: Grow the deviceid cache hash table Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 22/24] NFSv4/pnfs: Re-home the data-server cache onto hash buckets Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 23/24] NFSv4/pnfs: Key the data-server cache by its address set and version Benjamin Coddington
2026-09-04 16:53 ` [PATCH v3 24/24] NFSv4/flexfiles: Add a dataserver_nconnect cap Benjamin Coddington

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=47cdf067fb0582ff5cbcd229f3781787095d9701.1788530385.git.bcodding@hammerspace.com \
    --to=ben.coddington@hammerspace.com \
    --cc=anna@kernel.org \
    --cc=jcurley@purestorage.com \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=moonafterrain@outlook.com \
    --cc=snitzer@kernel.org \
    --cc=trondmy@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox