From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f177.google.com (mail-pf1-f177.google.com [209.85.210.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEDC33081A2 for ; Wed, 26 Nov 2025 02:13:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.177 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764123204; cv=none; b=mkhE2/NtmmIBlL87l1gVb94EGSeBAha4oEzdc/k4iTHMfnsLUlLimuanmemFi/BsVd7iB0xTpEUidvN2yeceJ2tfuoDMkwAeHkhNNTHjqffa/LHNoL/745zx/ESPpAVyGvJBJNsr9Dg4oquYyUOF30eWfpl2wVO72IX6A9Q65RA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764123204; c=relaxed/simple; bh=vVtLE64Vh6DaD2m3HWNfJ4gz/HGfyITLeNeWKxfeGqk=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=F+AbbFqmcb2QIEXAKq6j7KM7e+bM4ado1g2FFwI8FvhmAyU9m/OUWnQZWoW3nDTrVtXSzm1ebG8dTCpJFddijAT1dAsRGmnMu3+QsGOuhOmZgo0m6Eg+0FSKiswQLYEKGe4+pTL1Vp+XAdI3l4unptClBro6QiIY71hENj736Lg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=purestorage.com; spf=fail smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=ao8PERGd; arc=none smtp.client-ip=209.85.210.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="ao8PERGd" Received: by mail-pf1-f177.google.com with SMTP id d2e1a72fcca58-7aad4823079so5433781b3a.0 for ; Tue, 25 Nov 2025 18:13:22 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1764123202; x=1764728002; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=0DYMwugwCCMRbCMYT3zoF/0wVBmZFEJOOVa8ErucyPQ=; b=ao8PERGdYAOJ+bqsbKUw8kIROyteewFgMY7/9KSIO+Qcs9iasIwTcMmJgR411jxQ4R /DT4fxQ7NHaJLok3/yhxZsQVQHsCp4pNMF5kvCJJoR2PBdUnHfy6x66fh1vmaccaqkal dxSayQjgt5xJsV9Dp2rQRDrh0fGtlsfmwyPLgZDuzYZRZojl6bInRmDGCeSHxw6MtiTw Q8eSwlCCMqYfQwHlOAze+MM50R1GQ3L1yrlUybdZ83pVZcanqvqIGEVvvb3xyESMpPao SAWV3n04U56vpclKMYZGIkEPLMfJeNVlFhbwS7123Q/44ABXnhT6fAJ6loRnSRfXhrNU DVyQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1764123202; x=1764728002; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=0DYMwugwCCMRbCMYT3zoF/0wVBmZFEJOOVa8ErucyPQ=; b=uWmnq9Ao52dv6FXUAFb15Aip11fwYCAn23Ann97v8//gcnZuhtCyFfF+wHH3jNzx7q YPxvpUfRj3JNvQOJbv/N50kzLRCE25M9sFWBbM9ipb1Xutvvv8vwcUx1Ndnw0871ZPjG i/CzyMD9i7tDok3g90Hr3bQt38hXkE9P/ShFBQJi7g20snugVn/JSiI+kOsd46cvL5S+ ojvpjMPiRYnyFWkuIllc0IANPSfrCtGn861FhQc/m0K0xe+HYnnZaQ/B4IZKKCaypGT0 BJnfjYuHUUSTt7O/E66Upej78gzIkOE7XzozwbH0iVecF/7s1Av3s+C7sYhmXC9DVIrZ K/iA== X-Forwarded-Encrypted: i=1; AJvYcCUFmObc6ncjshlX61F1ACvf/1uK6SBHhJvgnLNUzQrCZU0XtcoaKfUqfNVls0Xyl3XkVa4DAssWuXfShRI=@vger.kernel.org X-Gm-Message-State: AOJu0YwFoW9ZVvLTXlWn3O9mv537SfxNodHJkgDsqVHyZCEuL1Cu19j8 iZ0GXlrcWuqn+pN7jRYKEdXutSz1sigsHmAdT4t1tQkuC8SBeLPjF9EgIhJhAfHx+vI= X-Gm-Gg: ASbGncuz2h9VgD/8hcYfsd87dfh0dAZ2RuAr92ZOK2obC99rejB/jjIGCvGQH+ScwDV opw8lldrvgCpcafjkLFa4K7O8dx8t1iI3umkhLGg/d0h9+yWip/AXI74tpaOTfSljZ291Y/hBqY xb83axZoGTjCxnXmSpcIFPBIT7V7wkLIZQsjQOFrg93bWgn3accePMmJkfa1NUBGdXTzXonNmEo ZYXpwYiLjqsoS/ReoLbndyx9u3Gxr4R+iK4p799mdxDZC94YFH10Y/QOw76DcMc3ZKFnImoOCqU VbyH2leGM8J49AHbS0H8I3DwnSzzYqZwsbZuxvtO3aaQYfVLJYTa+aw3HCM5uaDpUbosMvvkA+/ KKk19qpFrRBZ5QxSLHi11W9+EsX0+GnjWBk3dTRzBjItL8O+rOuBvvoC//EOLTp8Reye9k/6cLI Zw0nG4WXc/gzCuZXXyn8U5vGP8qgFe/otuqw== X-Google-Smtp-Source: AGHT+IHX3yT5pXdf7LYjSNE6WgFqzyyar12yID69DwGd9ndEqYxMAf61RI2rL1JknD6xXgHVKt+nNg== X-Received: by 2002:a05:7022:ec16:b0:119:e56b:91da with SMTP id a92af1059eb24-11c9d811990mr11265772c88.11.1764123201516; Tue, 25 Nov 2025 18:13:21 -0800 (PST) Received: from apollo.purestorage.com ([208.88.152.253]) by smtp.googlemail.com with ESMTPSA id a92af1059eb24-11cc631c236sm17922979c88.7.2025.11.25.18.13.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 25 Nov 2025 18:13:21 -0800 (PST) From: Mohamed Khalfella To: Chaitanya Kulkarni , Christoph Hellwig , Jens Axboe , Keith Busch , Sagi Grimberg Cc: Aaron Dailey , Randy Jennings , John Meneghini , Hannes Reinecke , linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, Mohamed Khalfella Subject: [RFC PATCH 00/14] TP8028 Rapid Path Failure Recovery Date: Tue, 25 Nov 2025 18:11:47 -0800 Message-ID: <20251126021250.2583630-1-mkhalfella@purestorage.com> X-Mailer: git-send-email 2.51.2 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit This patchset adds support for TP8028 Rapid Path Failure Recovery for both nvme target and initiator. Rapid Path Failure Recovery brings Cross-Controller Reset (CCR) functionality to nvme. This allows nvme host to send an nvme command to source nvme controller to reset impacted nvme controller. Provided that both source and impacted controllers are in the same nvme subsystem. The main use of CCR is when one path to nvme subsystem fails. Inflight IOs on impacted nvme controller need to be terminated first before they can be retried on another path. Otherwise data corruption may happen. CCR provides a quick way to terminate these IOs on the unreachable nvme controller allowing recovery to move quickly and avoiding unnecessary delays. In case of CCR is not possible, then inflight requests are held for duration defined by TP4129 KATO Corrections and Clarifications before they are allowed to be retried. On the target side: - New struct members have been added to support CCR. struct nvme_id_ctrl has been updated with CIU (Controller Instance Uniquifier), CIRN (Controller Instance Random Number), and CQT (Command Quiesce Time). The combination of CIU, CNTLID, and CIRN is used to identify impacted controller in CCR command. - CCR nvme command implemented on the target causes impacted controller to fail and drop connections to host. - CCR logpage contains the status of pending CCR requests. An entry is added to the logpage after CCR request is validated. Completed CCR requests are removed from the logpage when controller becomes ready or when requested in get logpage command. - An AEN is sent when CCR completes to let the host know that it is safe to retry inflight requests. On the host side: - CIU, CIRN, and CQT have been added to struct nvme_ctrl. CIU and CIRN have been added to sysfs to make the values visible to user. CIU and CIRN can be used to construct and manually send admin-passthru CCR commands. - New controller state NVME_CTRL_RECOVERING has been added to prevent cancelling timed out inflight requests while CCR is in progress. Controller flag NVME_CTRL_RECOVERED was also added to signal end of time-based recovery. - Controller recovery in nvme_recover_ctrl() is invoked when LIVE controller hits an error or when a request times out. CCR is attempted to reset impacted controller. - Updated nvme fabric transports nvme-tcp, nvme-rdma, and nvme-fc to use CCR recovery. Ideally all inflight requests should be held during controller recovery and only retried after recovery is done. However, there are known situations that is not the case in this implementation. These gaps will be addressed in future patches: - Manual controller reset from sysfs will result in controller going to RESETTING state and all inflight requests to be canceled immediately and maybe retried on another path. - Manual controller delete from sysfs will also result in all inflight requests to be canceled immediately and maybe retried on another path. - In nvme-fc nvme controller will be deleted if remote port disappears with no timeout specified. This results in immediate cancellation of requests that maybe retried on another path. - In nvme-rdma if HCA is removed all nvme controllers will be deleted. This results in canceling inflight IOs and maybe they will be retred on another path. - In nvme-fc if controller is LIVE and an IO ends with an error from LLDD, only this IO will be completed immediately. However, the rest of inflight IOs will be held correctly because the controller will have transitioned to RECOVERING state. Mohamed Khalfella (14): nvmet: Rapid Path Failure Recovery set controller identify fields nvmet/debugfs: Add ctrl uniquifier and random values nvmet: Implement CCR nvme command nvmet: Implement CCR logpage nvmet: Send an AEN on CCR completion nvme: Rapid Path Failure Recovery read controller identify fields nvme: Add RECOVERING nvme controller state nvme: Implement cross-controller reset recovery nvme: Implement cross-controller reset completion nvme-tcp: Use CCR to recover controller that hits an error nvme-rdma: Use CCR to recover controller that hits an error nvme-fc: Decouple error recovery from controller reset nvme-fc: Use CCR to recover controller that hits an error nvme-fc: Hold inflight requests while in RECOVERING state drivers/nvme/host/constants.c | 1 + drivers/nvme/host/core.c | 197 +++++++++++++++++++++++++++++++- drivers/nvme/host/fc.c | 194 ++++++++++++++++++++----------- drivers/nvme/host/nvme.h | 24 ++++ drivers/nvme/host/rdma.c | 51 +++++++-- drivers/nvme/host/sysfs.c | 24 ++++ drivers/nvme/host/tcp.c | 52 +++++++-- drivers/nvme/target/admin-cmd.c | 127 ++++++++++++++++++++ drivers/nvme/target/core.c | 103 ++++++++++++++++- drivers/nvme/target/debugfs.c | 21 ++++ drivers/nvme/target/nvmet.h | 18 ++- include/linux/nvme.h | 57 ++++++++- 12 files changed, 778 insertions(+), 91 deletions(-) base-commit: fd95357fd8c6778ac7dea6c57a19b8b182b6e91f -- 2.51.2