All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v5 0/3] Error recovery for zPCI passthrough devices
@ 2026-09-14 17:44 Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3 Farhan Ali
                   ` (2 more replies)
  0 siblings, 3 replies; 9+ messages in thread
From: Farhan Ali @ 2026-09-14 17:44 UTC (permalink / raw)
  To: qemu-devel, qemu-s390x; +Cc: alifm, mjrosato, farman, cohuck, alex, clg, armbru

Hi,

This patch series introduces support for error recovery for passthrough PCI
devices on System Z (s390x). This is the user space component for the Linux
kernel patches [1]. The kernel patches were merged for 7.3 release.

The current design for QEMU updates the callback for vfio error notifier to
an s390x specific error handler if the kernel supports the new device
feature VFIO_DEVICE_FEATURE_ZPCI_ERROR, for s390 vfio-pci devices. So now
on an eventfd notification for error notifier it will invoke the s390x
specific error handler.  The s390x error handler will retrieve the
architecture specific PCI error information and inject the information into
the guest. Once the guest receives the error information, the guest drivers
will drive the error recovery.  Typically recovery involves a device reset
which translate to CLP disable/enable cycle for the device.

I would appreciate some feedback on this patch series.

Thanks Farhan

[1] https://lore.kernel.org/all/20260818164326.387bb27b@shazbot.org/

ChangeLog
---------
v4 https://lore.kernel.org/all/20260831183151.12626-1-alifm@linux.ibm.com/
v4 -> v5
    - Remove err_handler() callback in vfio pci core.
    - Update the vfio error notifier callback to an s390x specific callback
    for s390x devices
    - Include linux headers for 7.3-rc3.
v3 https://lore.kernel.org/qemu-devel/20250925174852.1302-1-alifm@linux.ibm.com/
v3 -> v4
    - Include linux headers for 7.3-rc1.
    - Rework VFIO API changes based on the kernel API (patch 3).
    - Address Markus's comments from v3 (patch 2).

v2 https://lore.kernel.org/qemu-devel/20250825212434.2255-1-alifm@linux.ibm.com/
v2 -> v3
    - Update arch_err_handler to err_handler and include Error ** in
    function definition. (patch 2)

    - Introduce helper function to hide the internal indirection of device_feature()
    (patch 3)

    - Update function definitions to include Error ** (patch 4)
    


v1 https://lore.kernel.org/qemu-devel/20250813174152.1238-1-alifm@linux.ibm.com/
v1 -> v2
   - Use VFIO_DEVICE_FEATURE ioctl to get device error information.
   (Based on Alex's feedback on kernel series)


Farhan Ali (3):
  linux-headers: Update Linux header to 7.3-rc3
  s390x/pci: Add PCI error handling for vfio pci devices
  s390x/pci: Reset a device in error state

 hw/s390x/s390-pci-bus.c                   |  13 ++
 hw/s390x/s390-pci-vfio-stubs.c            |  10 ++
 hw/s390x/s390-pci-vfio.c                  | 124 +++++++++++++
 include/hw/s390x/s390-pci-bus.h           |   1 +
 include/hw/s390x/s390-pci-vfio.h          |   2 +
 include/standard-headers/drm/drm_fourcc.h | 209 ++++++++++++++++++++--
 include/standard-headers/linux/fuse.h     |  63 ++++++-
 linux-headers/asm-arm64/kvm.h             |   1 +
 linux-headers/asm-riscv/kvm.h             |  13 ++
 linux-headers/linux/kvm.h                 |   1 +
 10 files changed, 417 insertions(+), 20 deletions(-)

-- 
2.43.0



^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3
  2026-09-14 17:44 [PATCH v5 0/3] Error recovery for zPCI passthrough devices Farhan Ali
@ 2026-09-14 17:44 ` Farhan Ali
  2026-09-15 22:32   ` Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 3/3] s390x/pci: Reset a device in error state Farhan Ali
  2 siblings, 1 reply; 9+ messages in thread
From: Farhan Ali @ 2026-09-14 17:44 UTC (permalink / raw)
  To: qemu-devel, qemu-s390x; +Cc: alifm, mjrosato, farman, cohuck, alex, clg, armbru

Update the headers to include VFIO and zPCI VFIO changes.

Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
---
 include/standard-headers/drm/drm_fourcc.h | 209 ++++++++++++++++++++--
 include/standard-headers/linux/fuse.h     |  63 ++++++-
 linux-headers/asm-arm64/kvm.h             |   1 +
 linux-headers/asm-riscv/kvm.h             |  13 ++
 linux-headers/linux/kvm.h                 |   1 +
 5 files changed, 267 insertions(+), 20 deletions(-)

diff --git a/include/standard-headers/drm/drm_fourcc.h b/include/standard-headers/drm/drm_fourcc.h
index 2ba0b5c925..ca4f4e3d8c 100644
--- a/include/standard-headers/drm/drm_fourcc.h
+++ b/include/standard-headers/drm/drm_fourcc.h
@@ -1645,36 +1645,69 @@ drm_fourcc_canonicalize_nvidia_format_mod(uint64_t modifier)
  * For multi-plane formats the above surfaces get merged into one plane for
  * each format plane, based on the required alignment only.
  *
- * Bits  Parameter                Notes
- * ----- ------------------------ ---------------------------------------------
+ * Bits    Parameter                Notes
+ * ------- ------------------------ ---------------------------------------------
+ *
+ * DRM format modifier fields on AMD GPUs:
+ *     7:0 TILE_VERSION             Values are AMD_FMT_MOD_TILE_VER_*
+ *    12:8 TILE                     Values are AMD_FMT_MOD_TILE_<version>_*
+ *      13 DCC                      Delta Color Compression, supported on GFX8 and newer
+ *   55:14 (chip specific)          See below for details, depends on GFX block version
+ *   63:56 Vendor                   Value is DRM_FORMAT_MOD_VENDOR_AMD
+ *
+ * Chip specific fields on Gfx9 and newer:
+ *      14 DCC_RETILE
+ *      15 DCC_PIPE_ALIGN
+ *      16 DCC_INDEPENDENT_64B
+ *      17 DCC_INDEPENDENT_128B
+ *   19:18 DCC_MAX_COMPRESSED_BLOCK Values are AMD_FMT_MOD_DCC_BLOCK_*
+ *      20 DCC_CONSTANT_ENCODE
+ *   23:21 PIPE_XOR_BITS            Only for some chips
+ *   26:24 BANK_XOR_BITS            Only for some chips
+ *   29:27 PACKERS                  Only for some chips
+ *   32:30 RB                       Only for some chips
+ *   35:33 PIPE                     Only for some chips
+ *   55:36 -                        Reserved for future use, must be zero
+ *
+ * Chip specific fields on Gfx6-8:
+ *   16:14 MICROTILE                Micro tile format
+ *   21:17 PIPE_CONFIG              Number of pipes and how pipes are interleaved
+ *   24:22 TILE_SPLIT               Tile split size
+ *   26:25 BANK_WIDTH               Number of tiles in the X direction in the same bank
+ *   28:27 BANK_HEIGHT              Number of tiles in the Y direction in the same bank
+ *   30:29 MACRO_TILE_ASPECT        Macro tile aspect ratio
+ *   32:31 NUM_BANKS                Number of banks
+ *   55:33 -                        Reserved for future use, must be zero
  *
- *   7:0 TILE_VERSION             Values are AMD_FMT_MOD_TILE_VER_*
- *  12:8 TILE                     Values are AMD_FMT_MOD_TILE_<version>_*
- *    13 DCC
- *    14 DCC_RETILE
- *    15 DCC_PIPE_ALIGN
- *    16 DCC_INDEPENDENT_64B
- *    17 DCC_INDEPENDENT_128B
- * 19:18 DCC_MAX_COMPRESSED_BLOCK Values are AMD_FMT_MOD_DCC_BLOCK_*
- *    20 DCC_CONSTANT_ENCODE
- * 23:21 PIPE_XOR_BITS            Only for some chips
- * 26:24 BANK_XOR_BITS            Only for some chips
- * 29:27 PACKERS                  Only for some chips
- * 32:30 RB                       Only for some chips
- * 35:33 PIPE                     Only for some chips
- * 55:36 -                        Reserved for future use, must be zero
  */
 #define AMD_FMT_MOD fourcc_mod_code(AMD, 0)
 
 #define IS_AMD_FMT_MOD(val) (((val) >> 56) == DRM_FORMAT_MOD_VENDOR_AMD)
 
-/* Reserve 0 for GFX8 and older */
+#define AMD_FMT_MOD_TILE_VER_GFX6 0
 #define AMD_FMT_MOD_TILE_VER_GFX9 1
 #define AMD_FMT_MOD_TILE_VER_GFX10 2
 #define AMD_FMT_MOD_TILE_VER_GFX10_RBPLUS 3
 #define AMD_FMT_MOD_TILE_VER_GFX11 4
 #define AMD_FMT_MOD_TILE_VER_GFX12 5
 
+/*
+ * Gfx6-8 tiling modes.
+ * A complete reference implementation is found in addrlib in the Mesa code base.
+ *
+ * - Microtiled modes (1D):
+ *   Pixel data is organized into micro tiles of 8x8 pixels.
+ *
+ * - Macrotiled modes (2D):
+ *   Micro tiles are further organized into macro tiles.
+ *   These are optimized for even load distribution among memory channels.
+ *
+ * Note that only THIN1 modes are exposed here.
+ * THICK and XTHICK are for 3D images and not relevant to DRM format modifiers.
+ */
+#define AMD_FMT_MOD_TILE_GFX6_1D_TILED_THIN1 0x2
+#define AMD_FMT_MOD_TILE_GFX6_2D_TILED_THIN1 0x4
+
 /*
  * 64K_S is the same for GFX9/GFX10/GFX10_RBPLUS and hence has GFX9 as canonical
  * version.
@@ -1773,6 +1806,146 @@ drm_fourcc_canonicalize_nvidia_format_mod(uint64_t modifier)
 #define AMD_FMT_MOD_PIPE_SHIFT 33
 #define AMD_FMT_MOD_PIPE_MASK 0x7
 
+/*
+ * MICRO_TILE_MODE, 3 bits. Determines the micro tile format.
+ * Only relevant to Gfx6-8.
+ *
+ * DISPLAY - Displayable tiling
+ * THIN - Non-displayable tiling, a.k.a thin micro tiling
+ * DEPTH, THICK - not exposed, not relevant to DRM format modifier use cases
+ * ROTATED - not exposed, not implemented in Linux or Mesa
+ */
+#define AMD_FMT_MOD_MICROTILE_SHIFT 14ULL
+#define AMD_FMT_MOD_MICROTILE_MASK 0x7
+
+#define AMD_FMT_MOD_MICROTILE_DISPLAY 0x0
+#define AMD_FMT_MOD_MICROTILE_THIN 0x1
+
+/*
+ * PIPE_CONFIG, 5 bits. Number of pipes and how pipes are interleaved on the surface,
+ * which means the shader engine tile size and packer tile size.
+ * Typically matches the number of memory channels, or number of RBs.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ *
+ * P<n>_<a>x<b>_<c>x<d>
+ * where:
+ * <n> - number of pipes
+ * <a>x<b> - shader engine tile size
+ * <c>x<d> - packer tile size
+ */
+#define AMD_FMT_MOD_PIPE_CONFIG_SHIFT 17ULL
+#define AMD_FMT_MOD_PIPE_CONFIG_MASK 0x1f
+
+#define AMD_FMT_MOD_PIPE_CONFIG_P2 0x0
+#define AMD_FMT_MOD_PIPE_CONFIG_P4_8x16 0x4
+#define AMD_FMT_MOD_PIPE_CONFIG_P4_16x16 0x5
+#define AMD_FMT_MOD_PIPE_CONFIG_P4_16x32 0x6
+#define AMD_FMT_MOD_PIPE_CONFIG_P4_32x32 0x7
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x16_8x16 0x8
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x32_8x16 0x9
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_8x16 0xa
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x32_16x16 0xb
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_16x16 0xc
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_16x32 0xd
+#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x64_32x32 0xe
+#define AMD_FMT_MOD_PIPE_CONFIG_P16_32x32_8x16 0x10
+#define AMD_FMT_MOD_PIPE_CONFIG_P16_32x32_16x16 0x11
+
+/*
+ * TILE_SPLIT, 3 bits.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ *
+ * On GFX6 (or with depth tiling modes on GFX7 and newer),
+ * the GFX block uses the GB_TILE_MODE.TILE_SPLIT field directly.
+ *
+ * On GFX7 and newer with non-depth tiling modes, the GFX block uses a
+ * split factor which is stored in the GB_TILE_MODE.SAMPLE_SPLIT field.
+ * SAMPLE_SPLIT may be: 0 - 1 byte; 1 - 2 bytes; 2 - 4 bytes; 3 - 8 bytes.
+ * The actual tile size and tile split bytes are calculated as follows:
+ *
+ *    bpp = ... <- bits per pixel in the current image
+ *    thickness = ... <- depends on array mode; may be: 1, 4, 8
+ *    num_samples = ... <- number of samples in the current image
+ *    tile_size_pixels = 8 * 8
+ *    tile_bytes_1x = thickness * tile_size_pixels * bpp / 8
+ *    sample_split_factor = 1 << SAMPLE_SPLIT
+ *    tile_split_bytes = clamp(tile_bytes_1x * sample_split_factor, 256, dram_row_size_bytes)
+ *    tile_bytes = clamp(tile_bytes_1x * num_samples, 64, tile_split_bytes)
+ *
+ * In both cases, the display block (DCE) has no SAMPLE_SPLIT
+ * and just needs the tile split bytes in the GRPH_CONTROL.GRPH_TILE_SPLIT field.
+ * To maximize compatibility between GFX6-7, we don't include the SAMPLE_SPLIT
+ * in the format modifiers.
+ *
+ * The actual tile split in bytes is: 64 << field value
+ * Possible values of this field:
+ *
+ * 0 - Tile split is 64 bytes
+ * 1 - Tile split is 128 bytes
+ * 2 - Tile split is 256 bytes
+ * 3 - Tile split is 512 bytes
+ * 4 - Tile split is 1 KiB
+ * 5 - Tile split is 2 KiB
+ * 6 - Tile split is 4 KiB
+ */
+#define AMD_FMT_MOD_TILE_SPLIT_SHIFT 22ULL
+#define AMD_FMT_MOD_TILE_SPLIT_MASK 0x7
+
+/*
+ * BANK_WIDTH, 2 bits. Number of tiles in the X direction in the same bank.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ * The actual bank width is: 1 << field value
+ * Possible values:
+ *
+ * 0 - bank width is 1
+ * 1 - bank width is 2
+ * 2 - bank width is 4
+ * 3 - bank width is 8
+ */
+#define AMD_FMT_MOD_BANK_WIDTH_SHIFT 25ULL
+#define AMD_FMT_MOD_BANK_WIDTH_MASK 0x3
+
+/*
+ * BANK_HEIGHT, 2 bits. Number of tiles in the Y direction in the same bank.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ * The actual bank height is: 1 << field value
+ * Possible values:
+ *
+ * 0 - bank height is 1
+ * 1 - bank height is 2
+ * 2 - bank height is 4
+ * 3 - bank height is 8
+ */
+#define AMD_FMT_MOD_BANK_HEIGHT_SHIFT 27ULL
+#define AMD_FMT_MOD_BANK_HEIGHT_MASK 0x3
+
+/*
+ * MACRO_TILE_ASPECT, 2 bits. Macro tile aspect ratio.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ * Possible values:
+ *
+ * 0 - aspect ratio is 1:1
+ * 1 - aspect ratio is 4:1
+ * 2 - aspect ratio is 16:1
+ * 3 - aspect ratio is 64:1
+ */
+#define AMD_FMT_MOD_MACRO_TILE_ASPECT_SHIFT 29ULL
+#define AMD_FMT_MOD_MACRO_TILE_ASPECT_MASK 0x3
+
+/*
+ * NUM_BANKS, 2 bits. Number of banks.
+ * Only relevant to Gfx6-8 macro tiled modes.
+ * The actual number of banks is: 2 << field value
+ * Possible values:
+ *
+ * 0 - number of banks is 2
+ * 1 - number of banks is 4
+ * 2 - number of banks is 8
+ * 3 - number of banks is 16
+ */
+#define AMD_FMT_MOD_NUM_BANKS_SHIFT 31ULL
+#define AMD_FMT_MOD_NUM_BANKS_MASK 0x3
+
 #define AMD_FMT_MOD_SET(field, value) \
 	((uint64_t)(value) << AMD_FMT_MOD_##field##_SHIFT)
 #define AMD_FMT_MOD_GET(field, value) \
diff --git a/include/standard-headers/linux/fuse.h b/include/standard-headers/linux/fuse.h
index abf3a78858..cee8342d4e 100644
--- a/include/standard-headers/linux/fuse.h
+++ b/include/standard-headers/linux/fuse.h
@@ -240,6 +240,14 @@
  *  - add FUSE_COPY_FILE_RANGE_64
  *  - add struct fuse_copy_file_range_out
  *  - add FUSE_NOTIFY_PRUNE
+ *
+ *  7.46
+ *  - add FUSE_IO_URING_CMD_ADD_QUEUE
+ *  - add FUSE_HAS_IO_URING_BUFPOOL
+ *  - add fuse_uring_cmd_req bufpool struct
+ *  - add bufpool offset field to fuse_uring_ent_in_out struct
+ *  - add FUSE_URING_ZERO_COPY, FUSE_URING_ENT_ZERO_COPY, and
+ *    FOPEN_IO_URING_ZERO_COPY flag
  */
 
 #ifndef _LINUX_FUSE_H
@@ -271,7 +279,7 @@
 #define FUSE_KERNEL_VERSION 7
 
 /** Minor version number of this interface */
-#define FUSE_KERNEL_MINOR_VERSION 45
+#define FUSE_KERNEL_MINOR_VERSION 46
 
 /** The node ID of the root inode */
 #define FUSE_ROOT_ID 1
@@ -379,6 +387,12 @@ struct fuse_file_lock {
  * FOPEN_NOFLUSH: don't flush data cache on close (unless FUSE_WRITEBACK_CACHE)
  * FOPEN_PARALLEL_DIRECT_WRITES: Allow concurrent direct writes on the same inode
  * FOPEN_PASSTHROUGH: passthrough read/write io for this open file
+ * FOPEN_IO_URING_ZERO_COPY: use io-uring zero-copy for reads/writes on this
+ *                           open file. Honored only when the serving io-uring
+ *                           queue was set up for zero-copy
+ *                           (FUSE_URING_ZERO_COPY) and the request carries page
+ *                           payload. Otherwise reads/writes fall back to
+ *                           copying.
  */
 #define FOPEN_DIRECT_IO		(1 << 0)
 #define FOPEN_KEEP_CACHE	(1 << 1)
@@ -388,6 +402,7 @@ struct fuse_file_lock {
 #define FOPEN_NOFLUSH		(1 << 5)
 #define FOPEN_PARALLEL_DIRECT_WRITES	(1 << 6)
 #define FOPEN_PASSTHROUGH	(1 << 7)
+#define FOPEN_IO_URING_ZERO_COPY (1 << 8)
 
 /**
  * INIT request/reply flags
@@ -444,6 +459,7 @@ struct fuse_file_lock {
  * FUSE_OVER_IO_URING: Indicate that client supports io-uring
  * FUSE_REQUEST_TIMEOUT: kernel supports timing out requests.
  *			 init_out.request_timeout contains the timeout (in secs)
+ * FUSE_HAS_IO_URING_BUFPOOL: kernel supports io-uring buffer pools
  */
 #define FUSE_ASYNC_READ		(1 << 0)
 #define FUSE_POSIX_LOCKS	(1 << 1)
@@ -491,6 +507,7 @@ struct fuse_file_lock {
 #define FUSE_ALLOW_IDMAP	(1ULL << 40)
 #define FUSE_OVER_IO_URING	(1ULL << 41)
 #define FUSE_REQUEST_TIMEOUT	(1ULL << 42)
+#define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43)
 
 /**
  * CUSE INIT request/reply flags
@@ -1247,6 +1264,13 @@ struct fuse_supp_groups {
 #define FUSE_URING_IN_OUT_HEADER_SZ 128
 #define FUSE_URING_OP_IN_OUT_SZ 128
 
+/**
+ * fuse_uring_ent_in_out flags
+ *
+ * FUSE_URING_ENT_ZERO_COPY: Set if the ent's payload is zero-copied
+ */
+#define FUSE_URING_ENT_ZERO_COPY	(1 << 0)
+
 /* Used as part of the fuse_uring_req_header */
 struct fuse_uring_ent_in_out {
 	uint64_t flags;
@@ -1259,7 +1283,9 @@ struct fuse_uring_ent_in_out {
 
 	/* size of user payload buffer */
 	uint32_t payload_sz;
-	uint32_t padding;
+
+	/* Offset into the bufpool, if bufpools are used */
+	uint32_t offset;
 
 	uint64_t reserved;
 };
@@ -1288,8 +1314,22 @@ enum fuse_uring_cmd {
 
 	/* commit fuse request result and fetch next request */
 	FUSE_IO_URING_CMD_COMMIT_AND_FETCH = 2,
+
+	/* add a queue */
+	FUSE_IO_URING_CMD_ADD_QUEUE = 3,
+
+	/* add a bufpool to a queue */
+	FUSE_IO_URING_CMD_ADD_BUFPOOL = 4,
 };
 
+/*
+ * fuse_uring_cmd_req flags for FUSE_IO_URING_CMD_ADD_QUEUE
+ *
+ * FUSE_URING_ZERO_COPY is only supported for queues with bufpools on privileged
+ * servers
+ */
+#define FUSE_URING_ZERO_COPY		(1 << 0)
+
 /**
  * In the 80B command area of the SQE.
  */
@@ -1302,6 +1342,25 @@ struct fuse_uring_cmd_req {
 	/* queue the command is for (queue index) */
 	uint16_t qid;
 	uint8_t padding[6];
+
+	union {
+		struct {
+			/* base address of bufpool */
+			uint64_t uaddr;
+			uint32_t len;
+			uint32_t reserved;
+		} bufpool;
+
+		/*
+		 * Index of this entry's slot in the server's io_uring
+		 * registered buffer table, where the kernel registers the
+		 * request's pages for zero-copy. Set for
+		 * FUSE_IO_URING_CMD_REGISTER cmds only, and only on queues
+		 * created with FUSE_URING_ZERO_COPY. On a non-zero-copy queue
+		 * this must be 0
+		 */
+		uint16_t ent_zero_copy_buf_index;
+	};
 };
 
 #endif /* _LINUX_FUSE_H */
diff --git a/linux-headers/asm-arm64/kvm.h b/linux-headers/asm-arm64/kvm.h
index 6aefe79738..780cc7f729 100644
--- a/linux-headers/asm-arm64/kvm.h
+++ b/linux-headers/asm-arm64/kvm.h
@@ -106,6 +106,7 @@ struct kvm_regs {
 #define KVM_ARM_VCPU_PTRAUTH_GENERIC	6 /* VCPU uses generic authentication */
 #define KVM_ARM_VCPU_HAS_EL2		7 /* Support nested virtualization */
 #define KVM_ARM_VCPU_HAS_EL2_E2H0	8 /* Limit NV support to E2H RES0 */
+#define KVM_ARM_VCPU_PMU_V3_STRICT	9 /* No default PMU creation */
 
 struct kvm_vcpu_init {
 	__u32 target;
diff --git a/linux-headers/asm-riscv/kvm.h b/linux-headers/asm-riscv/kvm.h
index 504e733053..20d9959ca4 100644
--- a/linux-headers/asm-riscv/kvm.h
+++ b/linux-headers/asm-riscv/kvm.h
@@ -102,6 +102,11 @@ struct kvm_riscv_smstateen_csr {
 	unsigned long sstateen0;
 };
 
+/* Zicfiss CSR for KVM_GET_ONE_REG and KVM_SET_ONE_REG */
+struct kvm_riscv_zicfiss_csr {
+	unsigned long ssp;
+};
+
 /* TIMER registers for KVM_GET_ONE_REG and KVM_SET_ONE_REG */
 struct kvm_riscv_timer {
 	__u64 frequency;
@@ -199,6 +204,8 @@ enum KVM_RISCV_ISA_EXT_ID {
 	KVM_RISCV_ISA_EXT_ZCLSD,
 	KVM_RISCV_ISA_EXT_ZILSD,
 	KVM_RISCV_ISA_EXT_ZALASR,
+	KVM_RISCV_ISA_EXT_ZICFILP,
+	KVM_RISCV_ISA_EXT_ZICFISS,
 	KVM_RISCV_ISA_EXT_MAX,
 };
 
@@ -240,6 +247,9 @@ struct kvm_riscv_sbi_fwft_feature {
 struct kvm_riscv_sbi_fwft {
 	struct kvm_riscv_sbi_fwft_feature misaligned_deleg;
 	struct kvm_riscv_sbi_fwft_feature pointer_masking;
+	struct kvm_riscv_sbi_fwft_feature pte_ad_hw_updating;
+	struct kvm_riscv_sbi_fwft_feature landing_pad;
+	struct kvm_riscv_sbi_fwft_feature shadow_stack;
 };
 
 /* If you need to interpret the index values, here is the key: */
@@ -263,12 +273,15 @@ struct kvm_riscv_sbi_fwft {
 #define KVM_REG_RISCV_CSR_GENERAL	(0x0 << KVM_REG_RISCV_SUBTYPE_SHIFT)
 #define KVM_REG_RISCV_CSR_AIA		(0x1 << KVM_REG_RISCV_SUBTYPE_SHIFT)
 #define KVM_REG_RISCV_CSR_SMSTATEEN	(0x2 << KVM_REG_RISCV_SUBTYPE_SHIFT)
+#define KVM_REG_RISCV_CSR_ZICFISS	(0x3 << KVM_REG_RISCV_SUBTYPE_SHIFT)
 #define KVM_REG_RISCV_CSR_REG(name)	\
 		(offsetof(struct kvm_riscv_csr, name) / sizeof(unsigned long))
 #define KVM_REG_RISCV_CSR_AIA_REG(name)	\
 	(offsetof(struct kvm_riscv_aia_csr, name) / sizeof(unsigned long))
 #define KVM_REG_RISCV_CSR_SMSTATEEN_REG(name)  \
 	(offsetof(struct kvm_riscv_smstateen_csr, name) / sizeof(unsigned long))
+#define KVM_REG_RISCV_CSR_ZICFISS_REG(name)  \
+	(offsetof(struct kvm_riscv_zicfiss_csr, name) / sizeof(unsigned long))
 
 /* Timer registers are mapped as type 4 */
 #define KVM_REG_RISCV_TIMER		(0x04 << KVM_REG_RISCV_TYPE_SHIFT)
diff --git a/linux-headers/linux/kvm.h b/linux-headers/linux/kvm.h
index aea4ab5c69..f21c47bf8d 100644
--- a/linux-headers/linux/kvm.h
+++ b/linux-headers/linux/kvm.h
@@ -987,6 +987,7 @@ struct kvm_enable_cap {
 #define KVM_CAP_S390_VSIE_ESAMODE 248
 #define KVM_CAP_S390_HPAGE_2G 249
 #define KVM_CAP_PPC_COMPAT_CAPS 250
+#define KVM_CAP_ARM_PMU_V3_STRICT 251
 
 struct kvm_irq_routing_irqchip {
 	__u32 irqchip;
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices
  2026-09-14 17:44 [PATCH v5 0/3] Error recovery for zPCI passthrough devices Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3 Farhan Ali
@ 2026-09-14 17:44 ` Farhan Ali
  2026-09-16  8:13   ` Cédric Le Goater
  2026-09-14 17:44 ` [PATCH v5 3/3] s390x/pci: Reset a device in error state Farhan Ali
  2 siblings, 1 reply; 9+ messages in thread
From: Farhan Ali @ 2026-09-14 17:44 UTC (permalink / raw)
  To: qemu-devel, qemu-s390x; +Cc: alifm, mjrosato, farman, cohuck, alex, clg, armbru

Add an s390x specific handler for vfio error notifier. For s390x pci devices,
we have platform specific error information. We need to retrieve this error
information for passthrough devices. This is done via a VFIO_DEVICE_FEATURE
ioctl which exposes that information.

Once this error information is retrieved we can then inject an error into
the guest, and let the guest drive the recovery.

Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
---
 hw/s390x/s390-pci-bus.c          |   6 ++
 hw/s390x/s390-pci-vfio-stubs.c   |   6 ++
 hw/s390x/s390-pci-vfio.c         | 115 +++++++++++++++++++++++++++++++
 include/hw/s390x/s390-pci-bus.h  |   1 +
 include/hw/s390x/s390-pci-vfio.h |   1 +
 5 files changed, 129 insertions(+)

diff --git a/hw/s390x/s390-pci-bus.c b/hw/s390x/s390-pci-bus.c
index 2eb4e8cec4..b2967dacba 100644
--- a/hw/s390x/s390-pci-bus.c
+++ b/hw/s390x/s390-pci-bus.c
@@ -1085,6 +1085,7 @@ static void s390_pcihost_plug(const HotplugHandler *hotplug_dev, DeviceState *de
     S390pciState *s = S390_PCI_HOST_BRIDGE(hotplug_dev);
     PCIDevice *pdev = NULL;
     S390PCIBusDevice *pbdev = NULL;
+    Error *local_err = NULL;
     int rc;
 
     if (object_dynamic_cast(OBJECT(dev), TYPE_PCI_BRIDGE)) {
@@ -1175,6 +1176,11 @@ static void s390_pcihost_plug(const HotplugHandler *hotplug_dev, DeviceState *de
             pbdev->iommu->dma_limit = s390_pci_start_dma_count(s, pbdev);
             /* Fill in CLP information passed via the vfio region */
             s390_pci_get_clp_info(pbdev);
+            /* Setup error handler for error recovery */
+            if (!s390_pci_setup_err_handler(pbdev, &local_err)) {
+                warn_report_err(local_err);
+            }
+
             if (!pbdev->interp) {
                 /* Do vfio passthrough but intercept for I/O */
                 pbdev->fh |= FH_SHM_VFIO;
diff --git a/hw/s390x/s390-pci-vfio-stubs.c b/hw/s390x/s390-pci-vfio-stubs.c
index d9882b7aad..9fc84ca135 100644
--- a/hw/s390x/s390-pci-vfio-stubs.c
+++ b/hw/s390x/s390-pci-vfio-stubs.c
@@ -30,3 +30,9 @@ bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh)
 void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
 {
 }
+
+bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
+{
+    error_setg(errp, "VFIO not available, cannot setup error handler");
+    return false;
+}
diff --git a/hw/s390x/s390-pci-vfio.c b/hw/s390x/s390-pci-vfio.c
index db6de00bd2..6c072005fd 100644
--- a/hw/s390x/s390-pci-vfio.c
+++ b/hw/s390x/s390-pci-vfio.c
@@ -10,6 +10,7 @@
  */
 
 #include "qemu/osdep.h"
+#include "qemu/error-report.h"
 
 #include <sys/ioctl.h>
 #include <linux/vfio.h>
@@ -105,6 +106,85 @@ void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt)
     }
 }
 
+static bool s390_pci_get_feature_err(VFIOPCIDevice *vfio_pci,
+                                    PciCcdfErr *ccdf,
+                                    uint32_t ccdf_err_length,
+                                    Error **errp)
+{
+    int ret;
+    size_t total_size;
+    struct vfio_device_feature_zpci_err *err;
+    g_autofree void *buf = NULL;
+    g_autofree struct vfio_device_feature *feature = NULL;
+
+    total_size = sizeof(*feature) + sizeof(*err);
+    feature = g_malloc(total_size);
+    feature->argsz = total_size;
+    feature->flags = VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_ZPCI_ERROR;
+
+    buf = g_malloc(ccdf_err_length);
+    err = (void *)feature->data;
+    err->data = (uint64_t)buf;
+    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
+
+    if (ret) {
+        if (ret != -ENOMSG) {
+            error_setg(errp, "Failed feature get VFIO_DEVICE_FEATURE_ZPCI_ERROR"
+                              " (rc=%d)", ret);
+        }
+        return false;
+    }
+
+    memcpy(ccdf, (PciCcdfErr *) err->data, ccdf_err_length);
+
+    return true;
+}
+
+static void s390_pci_err_handler(void *opaque)
+{
+    VFIOPCIDevice *vfio_pci;
+    S390PCIBusDevice *pbdev;
+    Error *errp = NULL;
+    PciCcdfErr ccdf;
+    bool ret = true;
+
+    vfio_pci = opaque;
+    if (!event_notifier_test_and_clear(&vfio_pci->err_notifier)) {
+        return;
+    }
+
+    pbdev = s390_pci_find_dev_by_target(s390_get_phb(),
+                                        DEVICE(&vfio_pci->parent_obj)->id);
+
+    if (!pbdev) {
+        error_report("No matching zpci device found");
+        return;
+    }
+    pbdev->state = ZPCI_FS_ERROR;
+
+    if (sizeof(ccdf) != pbdev->ccdf_err_length) {
+        error_report(
+                   "CCDF size mismatch expected size=%zu, provided size=%d",
+                   sizeof(ccdf), pbdev->ccdf_err_length);
+        return;
+    }
+
+    while (ret) {
+        ret = s390_pci_get_feature_err(vfio_pci, &ccdf,
+                                       pbdev->ccdf_err_length, &errp);
+        if (!ret) {
+            if (errp) {
+                error_report_err(errp);
+            }
+            break;
+        }
+        s390_pci_generate_error_event(ccdf.pec, pbdev->fh, pbdev->fid,
+                                      ccdf.faddr, ccdf.e);
+    }
+
+    return;
+}
+
 static void s390_pci_read_base(S390PCIBusDevice *pbdev,
                                struct vfio_device_info *info)
 {
@@ -134,6 +214,10 @@ static void s390_pci_read_base(S390PCIBusDevice *pbdev,
     /* Store function type separately for type-specific behavior */
     pbdev->pft = cap->pft;
 
+    if (hdr->version >= 3) {
+        pbdev->ccdf_err_length = cap->ccdf_err_length;
+    }
+
     /*
      * If the device is a passthrough ISM device, disallow relaxed
      * translation.
@@ -371,3 +455,34 @@ void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
     s390_pci_read_util(pbdev, info);
     s390_pci_read_pfip(pbdev, info);
 }
+
+bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
+{
+    int ret;
+    int32_t fd;
+    VFIOPCIDevice *vfio_pci = VFIO_PCI_DEVICE(pbdev->pdev);
+    uint64_t buf[DIV_ROUND_UP(sizeof(struct vfio_device_feature),
+                              sizeof(uint64_t))] = {};
+    struct vfio_device_feature *feature = (struct vfio_device_feature *)buf;
+
+    feature->argsz = sizeof(buf);
+    feature->flags = VFIO_DEVICE_FEATURE_PROBE | VFIO_DEVICE_FEATURE_ZPCI_ERROR;
+
+    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
+
+    if (ret != 0) {
+        if (ret == -ENOTTY) {
+            error_setg(errp, "Automated error recovery unavailable for device");
+        } else {
+            error_setg(errp,
+                       "Failed to probe for VFIO_DEVICE_FEATURE_ZPCI_ERROR (ret=%d)",
+                       ret);
+        }
+        return false;
+    }
+
+    fd = event_notifier_get_fd(&vfio_pci->err_notifier);
+    qemu_set_fd_handler(fd, s390_pci_err_handler, NULL, vfio_pci);
+
+    return true;
+}
diff --git a/include/hw/s390x/s390-pci-bus.h b/include/hw/s390x/s390-pci-bus.h
index 9228523ce8..c2348ede86 100644
--- a/include/hw/s390x/s390-pci-bus.h
+++ b/include/hw/s390x/s390-pci-bus.h
@@ -364,6 +364,7 @@ struct S390PCIBusDevice {
     bool forwarding_assist;
     bool aif;
     bool rtr_avail;
+    uint32_t ccdf_err_length;
     QTAILQ_ENTRY(S390PCIBusDevice) link;
 };
 
diff --git a/include/hw/s390x/s390-pci-vfio.h b/include/hw/s390x/s390-pci-vfio.h
index f7d6149daf..c7886b63ea 100644
--- a/include/hw/s390x/s390-pci-vfio.h
+++ b/include/hw/s390x/s390-pci-vfio.h
@@ -20,5 +20,6 @@ S390PCIDMACount *s390_pci_start_dma_count(S390pciState *s,
 void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt);
 bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh);
 void s390_pci_get_clp_info(S390PCIBusDevice *pbdev);
+bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp);
 
 #endif
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH v5 3/3] s390x/pci: Reset a device in error state
  2026-09-14 17:44 [PATCH v5 0/3] Error recovery for zPCI passthrough devices Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3 Farhan Ali
  2026-09-14 17:44 ` [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices Farhan Ali
@ 2026-09-14 17:44 ` Farhan Ali
  2 siblings, 0 replies; 9+ messages in thread
From: Farhan Ali @ 2026-09-14 17:44 UTC (permalink / raw)
  To: qemu-devel, qemu-s390x; +Cc: alifm, mjrosato, farman, cohuck, alex, clg, armbru

For passthrough devices in error state, for a guest driven reset of the
device we can attempt a reset to recover the device. A reset of the device
will trigger a CLP disable/enable cycle on the host to bring the device
into a recovered state.

Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
---
 hw/s390x/s390-pci-bus.c          | 7 +++++++
 hw/s390x/s390-pci-vfio-stubs.c   | 4 ++++
 hw/s390x/s390-pci-vfio.c         | 9 +++++++++
 include/hw/s390x/s390-pci-vfio.h | 1 +
 4 files changed, 21 insertions(+)

diff --git a/hw/s390x/s390-pci-bus.c b/hw/s390x/s390-pci-bus.c
index b2967dacba..8418da9372 100644
--- a/hw/s390x/s390-pci-bus.c
+++ b/hw/s390x/s390-pci-bus.c
@@ -1505,6 +1505,8 @@ static void s390_pci_device_reset(DeviceState *dev)
         return;
     case ZPCI_FS_STANDBY:
         break;
+    case ZPCI_FS_ERROR:
+        break;
     default:
         pbdev->fh &= ~FH_MASK_ENABLE;
         pbdev->state = ZPCI_FS_DISABLED;
@@ -1517,6 +1519,11 @@ static void s390_pci_device_reset(DeviceState *dev)
     } else if (pbdev->summary_ind) {
         pci_dereg_irqs(pbdev);
     }
+
+    if (pbdev->state == ZPCI_FS_ERROR) {
+        s390_pci_reset(pbdev);
+    }
+
     if (pbdev->iommu->enabled) {
         pci_dereg_ioat(pbdev->iommu);
     }
diff --git a/hw/s390x/s390-pci-vfio-stubs.c b/hw/s390x/s390-pci-vfio-stubs.c
index 9fc84ca135..c68c612038 100644
--- a/hw/s390x/s390-pci-vfio-stubs.c
+++ b/hw/s390x/s390-pci-vfio-stubs.c
@@ -36,3 +36,7 @@ bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
     error_setg(errp, "VFIO not available, cannot setup error handler");
     return false;
 }
+
+void s390_pci_reset(S390PCIBusDevice *pbdev)
+{
+}
diff --git a/hw/s390x/s390-pci-vfio.c b/hw/s390x/s390-pci-vfio.c
index 6c072005fd..37d2f49658 100644
--- a/hw/s390x/s390-pci-vfio.c
+++ b/hw/s390x/s390-pci-vfio.c
@@ -185,6 +185,15 @@ static void s390_pci_err_handler(void *opaque)
     return;
 }
 
+void s390_pci_reset(S390PCIBusDevice *pbdev)
+{
+    VFIOPCIDevice *vfio_pci = VFIO_PCI_DEVICE(pbdev->pdev);
+    if (ioctl(vfio_pci->vbasedev.fd, VFIO_DEVICE_RESET)) {
+        error_report("Failed to reset PCI device %s : %s ",
+                      vfio_pci->vbasedev.name, strerror(errno));
+    }
+}
+
 static void s390_pci_read_base(S390PCIBusDevice *pbdev,
                                struct vfio_device_info *info)
 {
diff --git a/include/hw/s390x/s390-pci-vfio.h b/include/hw/s390x/s390-pci-vfio.h
index c7886b63ea..38ccf445ea 100644
--- a/include/hw/s390x/s390-pci-vfio.h
+++ b/include/hw/s390x/s390-pci-vfio.h
@@ -21,5 +21,6 @@ void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt);
 bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh);
 void s390_pci_get_clp_info(S390PCIBusDevice *pbdev);
 bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp);
+void s390_pci_reset(S390PCIBusDevice *pbdev);
 
 #endif
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3
  2026-09-14 17:44 ` [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3 Farhan Ali
@ 2026-09-15 22:32   ` Farhan Ali
  0 siblings, 0 replies; 9+ messages in thread
From: Farhan Ali @ 2026-09-15 22:32 UTC (permalink / raw)
  To: qemu-devel, qemu-s390x; +Cc: mjrosato, farman, cohuck, alex, clg, armbru

It looks like I didn't correctly update the headers in this version and 
so it didn't pick the VFIO header changes correctly. Will wait for some 
feedback before spinning a new version.

Thanks

Farhan

On 9/14/2026 10:44 AM, Farhan Ali wrote:
> Update the headers to include VFIO and zPCI VFIO changes.
>
> Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
> ---
>   include/standard-headers/drm/drm_fourcc.h | 209 ++++++++++++++++++++--
>   include/standard-headers/linux/fuse.h     |  63 ++++++-
>   linux-headers/asm-arm64/kvm.h             |   1 +
>   linux-headers/asm-riscv/kvm.h             |  13 ++
>   linux-headers/linux/kvm.h                 |   1 +
>   5 files changed, 267 insertions(+), 20 deletions(-)
>
> diff --git a/include/standard-headers/drm/drm_fourcc.h b/include/standard-headers/drm/drm_fourcc.h
> index 2ba0b5c925..ca4f4e3d8c 100644
> --- a/include/standard-headers/drm/drm_fourcc.h
> +++ b/include/standard-headers/drm/drm_fourcc.h
> @@ -1645,36 +1645,69 @@ drm_fourcc_canonicalize_nvidia_format_mod(uint64_t modifier)
>    * For multi-plane formats the above surfaces get merged into one plane for
>    * each format plane, based on the required alignment only.
>    *
> - * Bits  Parameter                Notes
> - * ----- ------------------------ ---------------------------------------------
> + * Bits    Parameter                Notes
> + * ------- ------------------------ ---------------------------------------------
> + *
> + * DRM format modifier fields on AMD GPUs:
> + *     7:0 TILE_VERSION             Values are AMD_FMT_MOD_TILE_VER_*
> + *    12:8 TILE                     Values are AMD_FMT_MOD_TILE_<version>_*
> + *      13 DCC                      Delta Color Compression, supported on GFX8 and newer
> + *   55:14 (chip specific)          See below for details, depends on GFX block version
> + *   63:56 Vendor                   Value is DRM_FORMAT_MOD_VENDOR_AMD
> + *
> + * Chip specific fields on Gfx9 and newer:
> + *      14 DCC_RETILE
> + *      15 DCC_PIPE_ALIGN
> + *      16 DCC_INDEPENDENT_64B
> + *      17 DCC_INDEPENDENT_128B
> + *   19:18 DCC_MAX_COMPRESSED_BLOCK Values are AMD_FMT_MOD_DCC_BLOCK_*
> + *      20 DCC_CONSTANT_ENCODE
> + *   23:21 PIPE_XOR_BITS            Only for some chips
> + *   26:24 BANK_XOR_BITS            Only for some chips
> + *   29:27 PACKERS                  Only for some chips
> + *   32:30 RB                       Only for some chips
> + *   35:33 PIPE                     Only for some chips
> + *   55:36 -                        Reserved for future use, must be zero
> + *
> + * Chip specific fields on Gfx6-8:
> + *   16:14 MICROTILE                Micro tile format
> + *   21:17 PIPE_CONFIG              Number of pipes and how pipes are interleaved
> + *   24:22 TILE_SPLIT               Tile split size
> + *   26:25 BANK_WIDTH               Number of tiles in the X direction in the same bank
> + *   28:27 BANK_HEIGHT              Number of tiles in the Y direction in the same bank
> + *   30:29 MACRO_TILE_ASPECT        Macro tile aspect ratio
> + *   32:31 NUM_BANKS                Number of banks
> + *   55:33 -                        Reserved for future use, must be zero
>    *
> - *   7:0 TILE_VERSION             Values are AMD_FMT_MOD_TILE_VER_*
> - *  12:8 TILE                     Values are AMD_FMT_MOD_TILE_<version>_*
> - *    13 DCC
> - *    14 DCC_RETILE
> - *    15 DCC_PIPE_ALIGN
> - *    16 DCC_INDEPENDENT_64B
> - *    17 DCC_INDEPENDENT_128B
> - * 19:18 DCC_MAX_COMPRESSED_BLOCK Values are AMD_FMT_MOD_DCC_BLOCK_*
> - *    20 DCC_CONSTANT_ENCODE
> - * 23:21 PIPE_XOR_BITS            Only for some chips
> - * 26:24 BANK_XOR_BITS            Only for some chips
> - * 29:27 PACKERS                  Only for some chips
> - * 32:30 RB                       Only for some chips
> - * 35:33 PIPE                     Only for some chips
> - * 55:36 -                        Reserved for future use, must be zero
>    */
>   #define AMD_FMT_MOD fourcc_mod_code(AMD, 0)
>   
>   #define IS_AMD_FMT_MOD(val) (((val) >> 56) == DRM_FORMAT_MOD_VENDOR_AMD)
>   
> -/* Reserve 0 for GFX8 and older */
> +#define AMD_FMT_MOD_TILE_VER_GFX6 0
>   #define AMD_FMT_MOD_TILE_VER_GFX9 1
>   #define AMD_FMT_MOD_TILE_VER_GFX10 2
>   #define AMD_FMT_MOD_TILE_VER_GFX10_RBPLUS 3
>   #define AMD_FMT_MOD_TILE_VER_GFX11 4
>   #define AMD_FMT_MOD_TILE_VER_GFX12 5
>   
> +/*
> + * Gfx6-8 tiling modes.
> + * A complete reference implementation is found in addrlib in the Mesa code base.
> + *
> + * - Microtiled modes (1D):
> + *   Pixel data is organized into micro tiles of 8x8 pixels.
> + *
> + * - Macrotiled modes (2D):
> + *   Micro tiles are further organized into macro tiles.
> + *   These are optimized for even load distribution among memory channels.
> + *
> + * Note that only THIN1 modes are exposed here.
> + * THICK and XTHICK are for 3D images and not relevant to DRM format modifiers.
> + */
> +#define AMD_FMT_MOD_TILE_GFX6_1D_TILED_THIN1 0x2
> +#define AMD_FMT_MOD_TILE_GFX6_2D_TILED_THIN1 0x4
> +
>   /*
>    * 64K_S is the same for GFX9/GFX10/GFX10_RBPLUS and hence has GFX9 as canonical
>    * version.
> @@ -1773,6 +1806,146 @@ drm_fourcc_canonicalize_nvidia_format_mod(uint64_t modifier)
>   #define AMD_FMT_MOD_PIPE_SHIFT 33
>   #define AMD_FMT_MOD_PIPE_MASK 0x7
>   
> +/*
> + * MICRO_TILE_MODE, 3 bits. Determines the micro tile format.
> + * Only relevant to Gfx6-8.
> + *
> + * DISPLAY - Displayable tiling
> + * THIN - Non-displayable tiling, a.k.a thin micro tiling
> + * DEPTH, THICK - not exposed, not relevant to DRM format modifier use cases
> + * ROTATED - not exposed, not implemented in Linux or Mesa
> + */
> +#define AMD_FMT_MOD_MICROTILE_SHIFT 14ULL
> +#define AMD_FMT_MOD_MICROTILE_MASK 0x7
> +
> +#define AMD_FMT_MOD_MICROTILE_DISPLAY 0x0
> +#define AMD_FMT_MOD_MICROTILE_THIN 0x1
> +
> +/*
> + * PIPE_CONFIG, 5 bits. Number of pipes and how pipes are interleaved on the surface,
> + * which means the shader engine tile size and packer tile size.
> + * Typically matches the number of memory channels, or number of RBs.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + *
> + * P<n>_<a>x<b>_<c>x<d>
> + * where:
> + * <n> - number of pipes
> + * <a>x<b> - shader engine tile size
> + * <c>x<d> - packer tile size
> + */
> +#define AMD_FMT_MOD_PIPE_CONFIG_SHIFT 17ULL
> +#define AMD_FMT_MOD_PIPE_CONFIG_MASK 0x1f
> +
> +#define AMD_FMT_MOD_PIPE_CONFIG_P2 0x0
> +#define AMD_FMT_MOD_PIPE_CONFIG_P4_8x16 0x4
> +#define AMD_FMT_MOD_PIPE_CONFIG_P4_16x16 0x5
> +#define AMD_FMT_MOD_PIPE_CONFIG_P4_16x32 0x6
> +#define AMD_FMT_MOD_PIPE_CONFIG_P4_32x32 0x7
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x16_8x16 0x8
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x32_8x16 0x9
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_8x16 0xa
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_16x32_16x16 0xb
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_16x16 0xc
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x32_16x32 0xd
> +#define AMD_FMT_MOD_PIPE_CONFIG_P8_32x64_32x32 0xe
> +#define AMD_FMT_MOD_PIPE_CONFIG_P16_32x32_8x16 0x10
> +#define AMD_FMT_MOD_PIPE_CONFIG_P16_32x32_16x16 0x11
> +
> +/*
> + * TILE_SPLIT, 3 bits.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + *
> + * On GFX6 (or with depth tiling modes on GFX7 and newer),
> + * the GFX block uses the GB_TILE_MODE.TILE_SPLIT field directly.
> + *
> + * On GFX7 and newer with non-depth tiling modes, the GFX block uses a
> + * split factor which is stored in the GB_TILE_MODE.SAMPLE_SPLIT field.
> + * SAMPLE_SPLIT may be: 0 - 1 byte; 1 - 2 bytes; 2 - 4 bytes; 3 - 8 bytes.
> + * The actual tile size and tile split bytes are calculated as follows:
> + *
> + *    bpp = ... <- bits per pixel in the current image
> + *    thickness = ... <- depends on array mode; may be: 1, 4, 8
> + *    num_samples = ... <- number of samples in the current image
> + *    tile_size_pixels = 8 * 8
> + *    tile_bytes_1x = thickness * tile_size_pixels * bpp / 8
> + *    sample_split_factor = 1 << SAMPLE_SPLIT
> + *    tile_split_bytes = clamp(tile_bytes_1x * sample_split_factor, 256, dram_row_size_bytes)
> + *    tile_bytes = clamp(tile_bytes_1x * num_samples, 64, tile_split_bytes)
> + *
> + * In both cases, the display block (DCE) has no SAMPLE_SPLIT
> + * and just needs the tile split bytes in the GRPH_CONTROL.GRPH_TILE_SPLIT field.
> + * To maximize compatibility between GFX6-7, we don't include the SAMPLE_SPLIT
> + * in the format modifiers.
> + *
> + * The actual tile split in bytes is: 64 << field value
> + * Possible values of this field:
> + *
> + * 0 - Tile split is 64 bytes
> + * 1 - Tile split is 128 bytes
> + * 2 - Tile split is 256 bytes
> + * 3 - Tile split is 512 bytes
> + * 4 - Tile split is 1 KiB
> + * 5 - Tile split is 2 KiB
> + * 6 - Tile split is 4 KiB
> + */
> +#define AMD_FMT_MOD_TILE_SPLIT_SHIFT 22ULL
> +#define AMD_FMT_MOD_TILE_SPLIT_MASK 0x7
> +
> +/*
> + * BANK_WIDTH, 2 bits. Number of tiles in the X direction in the same bank.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + * The actual bank width is: 1 << field value
> + * Possible values:
> + *
> + * 0 - bank width is 1
> + * 1 - bank width is 2
> + * 2 - bank width is 4
> + * 3 - bank width is 8
> + */
> +#define AMD_FMT_MOD_BANK_WIDTH_SHIFT 25ULL
> +#define AMD_FMT_MOD_BANK_WIDTH_MASK 0x3
> +
> +/*
> + * BANK_HEIGHT, 2 bits. Number of tiles in the Y direction in the same bank.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + * The actual bank height is: 1 << field value
> + * Possible values:
> + *
> + * 0 - bank height is 1
> + * 1 - bank height is 2
> + * 2 - bank height is 4
> + * 3 - bank height is 8
> + */
> +#define AMD_FMT_MOD_BANK_HEIGHT_SHIFT 27ULL
> +#define AMD_FMT_MOD_BANK_HEIGHT_MASK 0x3
> +
> +/*
> + * MACRO_TILE_ASPECT, 2 bits. Macro tile aspect ratio.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + * Possible values:
> + *
> + * 0 - aspect ratio is 1:1
> + * 1 - aspect ratio is 4:1
> + * 2 - aspect ratio is 16:1
> + * 3 - aspect ratio is 64:1
> + */
> +#define AMD_FMT_MOD_MACRO_TILE_ASPECT_SHIFT 29ULL
> +#define AMD_FMT_MOD_MACRO_TILE_ASPECT_MASK 0x3
> +
> +/*
> + * NUM_BANKS, 2 bits. Number of banks.
> + * Only relevant to Gfx6-8 macro tiled modes.
> + * The actual number of banks is: 2 << field value
> + * Possible values:
> + *
> + * 0 - number of banks is 2
> + * 1 - number of banks is 4
> + * 2 - number of banks is 8
> + * 3 - number of banks is 16
> + */
> +#define AMD_FMT_MOD_NUM_BANKS_SHIFT 31ULL
> +#define AMD_FMT_MOD_NUM_BANKS_MASK 0x3
> +
>   #define AMD_FMT_MOD_SET(field, value) \
>   	((uint64_t)(value) << AMD_FMT_MOD_##field##_SHIFT)
>   #define AMD_FMT_MOD_GET(field, value) \
> diff --git a/include/standard-headers/linux/fuse.h b/include/standard-headers/linux/fuse.h
> index abf3a78858..cee8342d4e 100644
> --- a/include/standard-headers/linux/fuse.h
> +++ b/include/standard-headers/linux/fuse.h
> @@ -240,6 +240,14 @@
>    *  - add FUSE_COPY_FILE_RANGE_64
>    *  - add struct fuse_copy_file_range_out
>    *  - add FUSE_NOTIFY_PRUNE
> + *
> + *  7.46
> + *  - add FUSE_IO_URING_CMD_ADD_QUEUE
> + *  - add FUSE_HAS_IO_URING_BUFPOOL
> + *  - add fuse_uring_cmd_req bufpool struct
> + *  - add bufpool offset field to fuse_uring_ent_in_out struct
> + *  - add FUSE_URING_ZERO_COPY, FUSE_URING_ENT_ZERO_COPY, and
> + *    FOPEN_IO_URING_ZERO_COPY flag
>    */
>   
>   #ifndef _LINUX_FUSE_H
> @@ -271,7 +279,7 @@
>   #define FUSE_KERNEL_VERSION 7
>   
>   /** Minor version number of this interface */
> -#define FUSE_KERNEL_MINOR_VERSION 45
> +#define FUSE_KERNEL_MINOR_VERSION 46
>   
>   /** The node ID of the root inode */
>   #define FUSE_ROOT_ID 1
> @@ -379,6 +387,12 @@ struct fuse_file_lock {
>    * FOPEN_NOFLUSH: don't flush data cache on close (unless FUSE_WRITEBACK_CACHE)
>    * FOPEN_PARALLEL_DIRECT_WRITES: Allow concurrent direct writes on the same inode
>    * FOPEN_PASSTHROUGH: passthrough read/write io for this open file
> + * FOPEN_IO_URING_ZERO_COPY: use io-uring zero-copy for reads/writes on this
> + *                           open file. Honored only when the serving io-uring
> + *                           queue was set up for zero-copy
> + *                           (FUSE_URING_ZERO_COPY) and the request carries page
> + *                           payload. Otherwise reads/writes fall back to
> + *                           copying.
>    */
>   #define FOPEN_DIRECT_IO		(1 << 0)
>   #define FOPEN_KEEP_CACHE	(1 << 1)
> @@ -388,6 +402,7 @@ struct fuse_file_lock {
>   #define FOPEN_NOFLUSH		(1 << 5)
>   #define FOPEN_PARALLEL_DIRECT_WRITES	(1 << 6)
>   #define FOPEN_PASSTHROUGH	(1 << 7)
> +#define FOPEN_IO_URING_ZERO_COPY (1 << 8)
>   
>   /**
>    * INIT request/reply flags
> @@ -444,6 +459,7 @@ struct fuse_file_lock {
>    * FUSE_OVER_IO_URING: Indicate that client supports io-uring
>    * FUSE_REQUEST_TIMEOUT: kernel supports timing out requests.
>    *			 init_out.request_timeout contains the timeout (in secs)
> + * FUSE_HAS_IO_URING_BUFPOOL: kernel supports io-uring buffer pools
>    */
>   #define FUSE_ASYNC_READ		(1 << 0)
>   #define FUSE_POSIX_LOCKS	(1 << 1)
> @@ -491,6 +507,7 @@ struct fuse_file_lock {
>   #define FUSE_ALLOW_IDMAP	(1ULL << 40)
>   #define FUSE_OVER_IO_URING	(1ULL << 41)
>   #define FUSE_REQUEST_TIMEOUT	(1ULL << 42)
> +#define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43)
>   
>   /**
>    * CUSE INIT request/reply flags
> @@ -1247,6 +1264,13 @@ struct fuse_supp_groups {
>   #define FUSE_URING_IN_OUT_HEADER_SZ 128
>   #define FUSE_URING_OP_IN_OUT_SZ 128
>   
> +/**
> + * fuse_uring_ent_in_out flags
> + *
> + * FUSE_URING_ENT_ZERO_COPY: Set if the ent's payload is zero-copied
> + */
> +#define FUSE_URING_ENT_ZERO_COPY	(1 << 0)
> +
>   /* Used as part of the fuse_uring_req_header */
>   struct fuse_uring_ent_in_out {
>   	uint64_t flags;
> @@ -1259,7 +1283,9 @@ struct fuse_uring_ent_in_out {
>   
>   	/* size of user payload buffer */
>   	uint32_t payload_sz;
> -	uint32_t padding;
> +
> +	/* Offset into the bufpool, if bufpools are used */
> +	uint32_t offset;
>   
>   	uint64_t reserved;
>   };
> @@ -1288,8 +1314,22 @@ enum fuse_uring_cmd {
>   
>   	/* commit fuse request result and fetch next request */
>   	FUSE_IO_URING_CMD_COMMIT_AND_FETCH = 2,
> +
> +	/* add a queue */
> +	FUSE_IO_URING_CMD_ADD_QUEUE = 3,
> +
> +	/* add a bufpool to a queue */
> +	FUSE_IO_URING_CMD_ADD_BUFPOOL = 4,
>   };
>   
> +/*
> + * fuse_uring_cmd_req flags for FUSE_IO_URING_CMD_ADD_QUEUE
> + *
> + * FUSE_URING_ZERO_COPY is only supported for queues with bufpools on privileged
> + * servers
> + */
> +#define FUSE_URING_ZERO_COPY		(1 << 0)
> +
>   /**
>    * In the 80B command area of the SQE.
>    */
> @@ -1302,6 +1342,25 @@ struct fuse_uring_cmd_req {
>   	/* queue the command is for (queue index) */
>   	uint16_t qid;
>   	uint8_t padding[6];
> +
> +	union {
> +		struct {
> +			/* base address of bufpool */
> +			uint64_t uaddr;
> +			uint32_t len;
> +			uint32_t reserved;
> +		} bufpool;
> +
> +		/*
> +		 * Index of this entry's slot in the server's io_uring
> +		 * registered buffer table, where the kernel registers the
> +		 * request's pages for zero-copy. Set for
> +		 * FUSE_IO_URING_CMD_REGISTER cmds only, and only on queues
> +		 * created with FUSE_URING_ZERO_COPY. On a non-zero-copy queue
> +		 * this must be 0
> +		 */
> +		uint16_t ent_zero_copy_buf_index;
> +	};
>   };
>   
>   #endif /* _LINUX_FUSE_H */
> diff --git a/linux-headers/asm-arm64/kvm.h b/linux-headers/asm-arm64/kvm.h
> index 6aefe79738..780cc7f729 100644
> --- a/linux-headers/asm-arm64/kvm.h
> +++ b/linux-headers/asm-arm64/kvm.h
> @@ -106,6 +106,7 @@ struct kvm_regs {
>   #define KVM_ARM_VCPU_PTRAUTH_GENERIC	6 /* VCPU uses generic authentication */
>   #define KVM_ARM_VCPU_HAS_EL2		7 /* Support nested virtualization */
>   #define KVM_ARM_VCPU_HAS_EL2_E2H0	8 /* Limit NV support to E2H RES0 */
> +#define KVM_ARM_VCPU_PMU_V3_STRICT	9 /* No default PMU creation */
>   
>   struct kvm_vcpu_init {
>   	__u32 target;
> diff --git a/linux-headers/asm-riscv/kvm.h b/linux-headers/asm-riscv/kvm.h
> index 504e733053..20d9959ca4 100644
> --- a/linux-headers/asm-riscv/kvm.h
> +++ b/linux-headers/asm-riscv/kvm.h
> @@ -102,6 +102,11 @@ struct kvm_riscv_smstateen_csr {
>   	unsigned long sstateen0;
>   };
>   
> +/* Zicfiss CSR for KVM_GET_ONE_REG and KVM_SET_ONE_REG */
> +struct kvm_riscv_zicfiss_csr {
> +	unsigned long ssp;
> +};
> +
>   /* TIMER registers for KVM_GET_ONE_REG and KVM_SET_ONE_REG */
>   struct kvm_riscv_timer {
>   	__u64 frequency;
> @@ -199,6 +204,8 @@ enum KVM_RISCV_ISA_EXT_ID {
>   	KVM_RISCV_ISA_EXT_ZCLSD,
>   	KVM_RISCV_ISA_EXT_ZILSD,
>   	KVM_RISCV_ISA_EXT_ZALASR,
> +	KVM_RISCV_ISA_EXT_ZICFILP,
> +	KVM_RISCV_ISA_EXT_ZICFISS,
>   	KVM_RISCV_ISA_EXT_MAX,
>   };
>   
> @@ -240,6 +247,9 @@ struct kvm_riscv_sbi_fwft_feature {
>   struct kvm_riscv_sbi_fwft {
>   	struct kvm_riscv_sbi_fwft_feature misaligned_deleg;
>   	struct kvm_riscv_sbi_fwft_feature pointer_masking;
> +	struct kvm_riscv_sbi_fwft_feature pte_ad_hw_updating;
> +	struct kvm_riscv_sbi_fwft_feature landing_pad;
> +	struct kvm_riscv_sbi_fwft_feature shadow_stack;
>   };
>   
>   /* If you need to interpret the index values, here is the key: */
> @@ -263,12 +273,15 @@ struct kvm_riscv_sbi_fwft {
>   #define KVM_REG_RISCV_CSR_GENERAL	(0x0 << KVM_REG_RISCV_SUBTYPE_SHIFT)
>   #define KVM_REG_RISCV_CSR_AIA		(0x1 << KVM_REG_RISCV_SUBTYPE_SHIFT)
>   #define KVM_REG_RISCV_CSR_SMSTATEEN	(0x2 << KVM_REG_RISCV_SUBTYPE_SHIFT)
> +#define KVM_REG_RISCV_CSR_ZICFISS	(0x3 << KVM_REG_RISCV_SUBTYPE_SHIFT)
>   #define KVM_REG_RISCV_CSR_REG(name)	\
>   		(offsetof(struct kvm_riscv_csr, name) / sizeof(unsigned long))
>   #define KVM_REG_RISCV_CSR_AIA_REG(name)	\
>   	(offsetof(struct kvm_riscv_aia_csr, name) / sizeof(unsigned long))
>   #define KVM_REG_RISCV_CSR_SMSTATEEN_REG(name)  \
>   	(offsetof(struct kvm_riscv_smstateen_csr, name) / sizeof(unsigned long))
> +#define KVM_REG_RISCV_CSR_ZICFISS_REG(name)  \
> +	(offsetof(struct kvm_riscv_zicfiss_csr, name) / sizeof(unsigned long))
>   
>   /* Timer registers are mapped as type 4 */
>   #define KVM_REG_RISCV_TIMER		(0x04 << KVM_REG_RISCV_TYPE_SHIFT)
> diff --git a/linux-headers/linux/kvm.h b/linux-headers/linux/kvm.h
> index aea4ab5c69..f21c47bf8d 100644
> --- a/linux-headers/linux/kvm.h
> +++ b/linux-headers/linux/kvm.h
> @@ -987,6 +987,7 @@ struct kvm_enable_cap {
>   #define KVM_CAP_S390_VSIE_ESAMODE 248
>   #define KVM_CAP_S390_HPAGE_2G 249
>   #define KVM_CAP_PPC_COMPAT_CAPS 250
> +#define KVM_CAP_ARM_PMU_V3_STRICT 251
>   
>   struct kvm_irq_routing_irqchip {
>   	__u32 irqchip;


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices
  2026-09-14 17:44 ` [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices Farhan Ali
@ 2026-09-16  8:13   ` Cédric Le Goater
  2026-09-16 17:52     ` Farhan Ali
  2026-09-17  1:21     ` Matthew Rosato
  0 siblings, 2 replies; 9+ messages in thread
From: Cédric Le Goater @ 2026-09-16  8:13 UTC (permalink / raw)
  To: Farhan Ali, qemu-devel, qemu-s390x; +Cc: mjrosato, farman, cohuck, alex, armbru

On 9/14/26 19:44, Farhan Ali wrote:
> Add an s390x specific handler for vfio error notifier. For s390x pci devices,
> we have platform specific error information. We need to retrieve this error
> information for passthrough devices. This is done via a VFIO_DEVICE_FEATURE
> ioctl which exposes that information.
> 
> Once this error information is retrieved we can then inject an error into
> the guest, and let the guest drive the recovery.
> 
> Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
> ---
>   hw/s390x/s390-pci-bus.c          |   6 ++
>   hw/s390x/s390-pci-vfio-stubs.c   |   6 ++
>   hw/s390x/s390-pci-vfio.c         | 115 +++++++++++++++++++++++++++++++
>   include/hw/s390x/s390-pci-bus.h  |   1 +
>   include/hw/s390x/s390-pci-vfio.h |   1 +
>   5 files changed, 129 insertions(+)
> 
> diff --git a/hw/s390x/s390-pci-bus.c b/hw/s390x/s390-pci-bus.c
> index 2eb4e8cec4..b2967dacba 100644
> --- a/hw/s390x/s390-pci-bus.c
> +++ b/hw/s390x/s390-pci-bus.c
> @@ -1085,6 +1085,7 @@ static void s390_pcihost_plug(const HotplugHandler *hotplug_dev, DeviceState *de
>       S390pciState *s = S390_PCI_HOST_BRIDGE(hotplug_dev);
>       PCIDevice *pdev = NULL;
>       S390PCIBusDevice *pbdev = NULL;
> +    Error *local_err = NULL;
>       int rc;
>   
>       if (object_dynamic_cast(OBJECT(dev), TYPE_PCI_BRIDGE)) {
> @@ -1175,6 +1176,11 @@ static void s390_pcihost_plug(const HotplugHandler *hotplug_dev, DeviceState *de
>               pbdev->iommu->dma_limit = s390_pci_start_dma_count(s, pbdev);
>               /* Fill in CLP information passed via the vfio region */
>               s390_pci_get_clp_info(pbdev);
> +            /* Setup error handler for error recovery */
> +            if (!s390_pci_setup_err_handler(pbdev, &local_err)) {
> +                warn_report_err(local_err);
> +            }
> +
>               if (!pbdev->interp) {
>                   /* Do vfio passthrough but intercept for I/O */
>                   pbdev->fh |= FH_SHM_VFIO;
> diff --git a/hw/s390x/s390-pci-vfio-stubs.c b/hw/s390x/s390-pci-vfio-stubs.c
> index d9882b7aad..9fc84ca135 100644
> --- a/hw/s390x/s390-pci-vfio-stubs.c
> +++ b/hw/s390x/s390-pci-vfio-stubs.c
> @@ -30,3 +30,9 @@ bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh)
>   void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
>   {
>   }
> +
> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
> +{
> +    error_setg(errp, "VFIO not available, cannot setup error handler");
> +    return false;
> +}
> diff --git a/hw/s390x/s390-pci-vfio.c b/hw/s390x/s390-pci-vfio.c
> index db6de00bd2..6c072005fd 100644
> --- a/hw/s390x/s390-pci-vfio.c
> +++ b/hw/s390x/s390-pci-vfio.c
> @@ -10,6 +10,7 @@
>    */
>   
>   #include "qemu/osdep.h"
> +#include "qemu/error-report.h"
>   
>   #include <sys/ioctl.h>
>   #include <linux/vfio.h>
> @@ -105,6 +106,85 @@ void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt)
>       }
>   }
>   
> +static bool s390_pci_get_feature_err(VFIOPCIDevice *vfio_pci,
> +                                    PciCcdfErr *ccdf,
> +                                    uint32_t ccdf_err_length,
> +                                    Error **errp)
> +{
> +    int ret;
> +    size_t total_size;
> +    struct vfio_device_feature_zpci_err *err;
> +    g_autofree void *buf = NULL;

could be  buf = g_malloc(ccdf_err_length);


> +    g_autofree struct vfio_device_feature *feature = NULL;
> +
> +    total_size = sizeof(*feature) + sizeof(*err);
> +    feature = g_malloc(total_size);
> +    feature->argsz = total_size;
> +    feature->flags = VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_ZPCI_ERROR;
> +
> +    buf = g_malloc(ccdf_err_length);
> +    err = (void *)feature->data;
> +    err->data = (uint64_t)buf;
> +    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
> +
> +    if (ret) {
> +        if (ret != -ENOMSG) {
> +            error_setg(errp, "Failed feature get VFIO_DEVICE_FEATURE_ZPCI_ERROR"
> +                              " (rc=%d)", ret);
> +        }
> +        return false;

returning false without setting errp :/

> +    }
> +
> +    memcpy(ccdf, (PciCcdfErr *) err->data, ccdf_err_length);
> +
> +    return true;
> +}
> +
> +static void s390_pci_err_handler(void *opaque)
> +{
> +    VFIOPCIDevice *vfio_pci;
> +    S390PCIBusDevice *pbdev;
> +    Error *errp = NULL;

a 'local_err' name would be preferred.
> +    PciCcdfErr ccdf;
> +    bool ret = true;
> +
> +    vfio_pci = opaque;
> +    if (!event_notifier_test_and_clear(&vfio_pci->err_notifier)) {

This means that the vfio_pci->err_notifier eventfd was initialized.
IOW, pci_aer is true.

Is that the case for Z ? If not, this needs its own eventfd setup.

> +        return;
> +    }
> +
> +    pbdev = s390_pci_find_dev_by_target(s390_get_phb(),
> +                                        DEVICE(&vfio_pci->parent_obj)->id);
> +
> +    if (!pbdev) {
> +        error_report("No matching zpci device found");
> +        return;
> +    }
> +    pbdev->state = ZPCI_FS_ERROR;
> +
> +    if (sizeof(ccdf) != pbdev->ccdf_err_length) {
> +        error_report(
> +                   "CCDF size mismatch expected size=%zu, provided size=%d",
> +                   sizeof(ccdf), pbdev->ccdf_err_length);
> +        return;
> +    }
> +
> +    while (ret) {
> +        ret = s390_pci_get_feature_err(vfio_pci, &ccdf,
> +                                       pbdev->ccdf_err_length, &errp);
> +        if (!ret) {
> +            if (errp) {
> +                error_report_err(errp);
> +            }
> +            break;

The 'local_err' not being set with a 'false' returned value is not
following the qapi/error.h guidelines. This is unexpected.
It is also wrong. ERRP_GUARD() is needed. please read qapi/error.h.



> +        }
> +        s390_pci_generate_error_event(ccdf.pec, pbdev->fh, pbdev->fid,
> +                                      ccdf.faddr, ccdf.e);
> +    }
> +
> +    return;
> +}
> +
>   static void s390_pci_read_base(S390PCIBusDevice *pbdev,
>                                  struct vfio_device_info *info)
>   {
> @@ -134,6 +214,10 @@ static void s390_pci_read_base(S390PCIBusDevice *pbdev,
>       /* Store function type separately for type-specific behavior */
>       pbdev->pft = cap->pft;
>   
> +    if (hdr->version >= 3) {
> +        pbdev->ccdf_err_length = cap->ccdf_err_length;

So ccdf_err_length can be 0. Is that expected ?

> +    }
> +
>       /*
>        * If the device is a passthrough ISM device, disallow relaxed
>        * translation.
> @@ -371,3 +455,34 @@ void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
>       s390_pci_read_util(pbdev, info);
>       s390_pci_read_pfip(pbdev, info);
>   }
> +
> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
> +{
> +    int ret;
> +    int32_t fd;
> +    VFIOPCIDevice *vfio_pci = VFIO_PCI_DEVICE(pbdev->pdev);
> +    uint64_t buf[DIV_ROUND_UP(sizeof(struct vfio_device_feature),
> +                              sizeof(uint64_t))] = {};
> +    struct vfio_device_feature *feature = (struct vfio_device_feature *)buf;
> +
> +    feature->argsz = sizeof(buf);
> +    feature->flags = VFIO_DEVICE_FEATURE_PROBE | VFIO_DEVICE_FEATURE_ZPCI_ERROR;
> +
> +    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
> +
> +    if (ret != 0) {
> +        if (ret == -ENOTTY) {
> +            error_setg(errp, "Automated error recovery unavailable for device");
> +        } else {
> +            error_setg(errp,
> +                       "Failed to probe for VFIO_DEVICE_FEATURE_ZPCI_ERROR (ret=%d)",
> +                       ret);
> +        }
> +        return false;
> +    }
> +
> +    fd = event_notifier_get_fd(&vfio_pci->err_notifier);
> +    qemu_set_fd_handler(fd, s390_pci_err_handler, NULL, vfio_pci);

Shouldn't we check ccdf_err_length before installing the handler ? because
it won't run cleanly anyhow.

Thanks,

C.


> +
> +    return true;
> +}
> diff --git a/include/hw/s390x/s390-pci-bus.h b/include/hw/s390x/s390-pci-bus.h
> index 9228523ce8..c2348ede86 100644
> --- a/include/hw/s390x/s390-pci-bus.h
> +++ b/include/hw/s390x/s390-pci-bus.h
> @@ -364,6 +364,7 @@ struct S390PCIBusDevice {
>       bool forwarding_assist;
>       bool aif;
>       bool rtr_avail;
> +    uint32_t ccdf_err_length;
>       QTAILQ_ENTRY(S390PCIBusDevice) link;
>   };
>   
> diff --git a/include/hw/s390x/s390-pci-vfio.h b/include/hw/s390x/s390-pci-vfio.h
> index f7d6149daf..c7886b63ea 100644
> --- a/include/hw/s390x/s390-pci-vfio.h
> +++ b/include/hw/s390x/s390-pci-vfio.h
> @@ -20,5 +20,6 @@ S390PCIDMACount *s390_pci_start_dma_count(S390pciState *s,
>   void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt);
>   bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh);
>   void s390_pci_get_clp_info(S390PCIBusDevice *pbdev);
> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp);
>   
>   #endif



^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices
  2026-09-16  8:13   ` Cédric Le Goater
@ 2026-09-16 17:52     ` Farhan Ali
  2026-09-17  1:21     ` Matthew Rosato
  1 sibling, 0 replies; 9+ messages in thread
From: Farhan Ali @ 2026-09-16 17:52 UTC (permalink / raw)
  To: Cédric Le Goater, qemu-devel, qemu-s390x
  Cc: mjrosato, farman, cohuck, alex, armbru

Hi Cedric,

On 9/16/2026 1:13 AM, Cédric Le Goater wrote:
> On 9/14/26 19:44, Farhan Ali wrote:
>> Add an s390x specific handler for vfio error notifier. For s390x pci 
>> devices,
>> we have platform specific error information. We need to retrieve this 
>> error
>> information for passthrough devices. This is done via a 
>> VFIO_DEVICE_FEATURE
>> ioctl which exposes that information.
>>
>> Once this error information is retrieved we can then inject an error 
>> into
>> the guest, and let the guest drive the recovery.
>>
>> Signed-off-by: Farhan Ali <alifm@linux.ibm.com>
>> ---
>>   hw/s390x/s390-pci-bus.c          |   6 ++
>>   hw/s390x/s390-pci-vfio-stubs.c   |   6 ++
>>   hw/s390x/s390-pci-vfio.c         | 115 +++++++++++++++++++++++++++++++
>>   include/hw/s390x/s390-pci-bus.h  |   1 +
>>   include/hw/s390x/s390-pci-vfio.h |   1 +
>>   5 files changed, 129 insertions(+)
>>
>> diff --git a/hw/s390x/s390-pci-bus.c b/hw/s390x/s390-pci-bus.c
>> index 2eb4e8cec4..b2967dacba 100644
>> --- a/hw/s390x/s390-pci-bus.c
>> +++ b/hw/s390x/s390-pci-bus.c
>> @@ -1085,6 +1085,7 @@ static void s390_pcihost_plug(const 
>> HotplugHandler *hotplug_dev, DeviceState *de
>>       S390pciState *s = S390_PCI_HOST_BRIDGE(hotplug_dev);
>>       PCIDevice *pdev = NULL;
>>       S390PCIBusDevice *pbdev = NULL;
>> +    Error *local_err = NULL;
>>       int rc;
>>         if (object_dynamic_cast(OBJECT(dev), TYPE_PCI_BRIDGE)) {
>> @@ -1175,6 +1176,11 @@ static void s390_pcihost_plug(const 
>> HotplugHandler *hotplug_dev, DeviceState *de
>>               pbdev->iommu->dma_limit = s390_pci_start_dma_count(s, 
>> pbdev);
>>               /* Fill in CLP information passed via the vfio region */
>>               s390_pci_get_clp_info(pbdev);
>> +            /* Setup error handler for error recovery */
>> +            if (!s390_pci_setup_err_handler(pbdev, &local_err)) {
>> +                warn_report_err(local_err);
>> +            }
>> +
>>               if (!pbdev->interp) {
>>                   /* Do vfio passthrough but intercept for I/O */
>>                   pbdev->fh |= FH_SHM_VFIO;
>> diff --git a/hw/s390x/s390-pci-vfio-stubs.c 
>> b/hw/s390x/s390-pci-vfio-stubs.c
>> index d9882b7aad..9fc84ca135 100644
>> --- a/hw/s390x/s390-pci-vfio-stubs.c
>> +++ b/hw/s390x/s390-pci-vfio-stubs.c
>> @@ -30,3 +30,9 @@ bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, 
>> uint32_t *fh)
>>   void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
>>   {
>>   }
>> +
>> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
>> +{
>> +    error_setg(errp, "VFIO not available, cannot setup error handler");
>> +    return false;
>> +}
>> diff --git a/hw/s390x/s390-pci-vfio.c b/hw/s390x/s390-pci-vfio.c
>> index db6de00bd2..6c072005fd 100644
>> --- a/hw/s390x/s390-pci-vfio.c
>> +++ b/hw/s390x/s390-pci-vfio.c
>> @@ -10,6 +10,7 @@
>>    */
>>     #include "qemu/osdep.h"
>> +#include "qemu/error-report.h"
>>     #include <sys/ioctl.h>
>>   #include <linux/vfio.h>
>> @@ -105,6 +106,85 @@ void s390_pci_end_dma_count(S390pciState *s, 
>> S390PCIDMACount *cnt)
>>       }
>>   }
>>   +static bool s390_pci_get_feature_err(VFIOPCIDevice *vfio_pci,
>> +                                    PciCcdfErr *ccdf,
>> +                                    uint32_t ccdf_err_length,
>> +                                    Error **errp)
>> +{
>> +    int ret;
>> +    size_t total_size;
>> +    struct vfio_device_feature_zpci_err *err;
>> +    g_autofree void *buf = NULL;
>
> could be  buf = g_malloc(ccdf_err_length);

okay, will change.


>
>
>> +    g_autofree struct vfio_device_feature *feature = NULL;
>> +
>> +    total_size = sizeof(*feature) + sizeof(*err);
>> +    feature = g_malloc(total_size);
>> +    feature->argsz = total_size;
>> +    feature->flags = VFIO_DEVICE_FEATURE_GET | 
>> VFIO_DEVICE_FEATURE_ZPCI_ERROR;
>> +
>> +    buf = g_malloc(ccdf_err_length);
>> +    err = (void *)feature->data;
>> +    err->data = (uint64_t)buf;
>> +    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
>> +
>> +    if (ret) {
>> +        if (ret != -ENOMSG) {
>> +            error_setg(errp, "Failed feature get 
>> VFIO_DEVICE_FEATURE_ZPCI_ERROR"
>> +                              " (rc=%d)", ret);
>> +        }
>> +        return false;
>
> returning false without setting errp :/

An ENOMSG would indicate there are no more pending errors to be handled 
for the device. It doesn't indicate a critical error. So I don't think 
we want to set the errp? I am open to suggestion on this, should errp be 
set to warn in this case?


>
>> +    }
>> +
>> +    memcpy(ccdf, (PciCcdfErr *) err->data, ccdf_err_length);
>> +
>> +    return true;
>> +}
>> +
>> +static void s390_pci_err_handler(void *opaque)
>> +{
>> +    VFIOPCIDevice *vfio_pci;
>> +    S390PCIBusDevice *pbdev;
>> +    Error *errp = NULL;
>
> a 'local_err' name would be preferred.

okay, will change.

>
>> +    PciCcdfErr ccdf;
>> +    bool ret = true;
>> +
>> +    vfio_pci = opaque;
>> +    if (!event_notifier_test_and_clear(&vfio_pci->err_notifier)) {
>
> This means that the vfio_pci->err_notifier eventfd was initialized.
> IOW, pci_aer is true.
>
> Is that the case for Z ? If not, this needs its own eventfd setup.

Yes, it is. AFAIU the pci_aer is set if VFIO_PCI_ERR_IRQ_INDEX is 
supported, which it is on Z. The recent kernel change [1] enables it on 
all supported devices on Z.

[1] 
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=4e3c1fc8abcb8eff062150b4340fa4569696d645


>
>
>> +        return;
>> +    }
>> +
>> +    pbdev = s390_pci_find_dev_by_target(s390_get_phb(),
>> + DEVICE(&vfio_pci->parent_obj)->id);
>> +
>> +    if (!pbdev) {
>> +        error_report("No matching zpci device found");
>> +        return;
>> +    }
>> +    pbdev->state = ZPCI_FS_ERROR;
>> +
>> +    if (sizeof(ccdf) != pbdev->ccdf_err_length) {
>> +        error_report(
>> +                   "CCDF size mismatch expected size=%zu, provided 
>> size=%d",
>> +                   sizeof(ccdf), pbdev->ccdf_err_length);
>> +        return;
>> +    }
>> +
>> +    while (ret) {
>> +        ret = s390_pci_get_feature_err(vfio_pci, &ccdf,
>> + pbdev->ccdf_err_length, &errp);
>> +        if (!ret) {
>> +            if (errp) {
>> +                error_report_err(errp);
>> +            }
>> +            break;
>
> The 'local_err' not being set with a 'false' returned value is not
> following the qapi/error.h guidelines. This is unexpected.
> It is also wrong. ERRP_GUARD() is needed. please read qapi/error.h.

Should the ERRP_GUARD() be added here for local_err or in 
s390_pci_get_feature_err()?
I am open to suggestions on how we can handle the case of ENOMSG which 
is not a critical error.

>
>
>
>> +        }
>> +        s390_pci_generate_error_event(ccdf.pec, pbdev->fh, pbdev->fid,
>> +                                      ccdf.faddr, ccdf.e);
>> +    }
>> +
>> +    return;
>> +}
>> +
>>   static void s390_pci_read_base(S390PCIBusDevice *pbdev,
>>                                  struct vfio_device_info *info)
>>   {
>> @@ -134,6 +214,10 @@ static void s390_pci_read_base(S390PCIBusDevice 
>> *pbdev,
>>       /* Store function type separately for type-specific behavior */
>>       pbdev->pft = cap->pft;
>>   +    if (hdr->version >= 3) {
>> +        pbdev->ccdf_err_length = cap->ccdf_err_length;
>
> So ccdf_err_length can be 0. Is that expected ?

On kernels that don't support the vfio device feature, the kernel 
doesn't provide ccdf_err_length. So in that case pbdev->ccdf_err_length 
can be 0.


>
>> +    }
>> +
>>       /*
>>        * If the device is a passthrough ISM device, disallow relaxed
>>        * translation.
>> @@ -371,3 +455,34 @@ void s390_pci_get_clp_info(S390PCIBusDevice *pbdev)
>>       s390_pci_read_util(pbdev, info);
>>       s390_pci_read_pfip(pbdev, info);
>>   }
>> +
>> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp)
>> +{
>> +    int ret;
>> +    int32_t fd;
>> +    VFIOPCIDevice *vfio_pci = VFIO_PCI_DEVICE(pbdev->pdev);
>> +    uint64_t buf[DIV_ROUND_UP(sizeof(struct vfio_device_feature),
>> +                              sizeof(uint64_t))] = {};
>> +    struct vfio_device_feature *feature = (struct 
>> vfio_device_feature *)buf;
>> +
>> +    feature->argsz = sizeof(buf);
>> +    feature->flags = VFIO_DEVICE_FEATURE_PROBE | 
>> VFIO_DEVICE_FEATURE_ZPCI_ERROR;
>> +
>> +    ret = vfio_device_get_feature(&vfio_pci->vbasedev, feature);
>> +
>> +    if (ret != 0) {
>> +        if (ret == -ENOTTY) {
>> +            error_setg(errp, "Automated error recovery unavailable 
>> for device");
>> +        } else {
>> +            error_setg(errp,
>> +                       "Failed to probe for 
>> VFIO_DEVICE_FEATURE_ZPCI_ERROR (ret=%d)",
>> +                       ret);
>> +        }
>> +        return false;
>> +    }
>> +
>> +    fd = event_notifier_get_fd(&vfio_pci->err_notifier);
>> +    qemu_set_fd_handler(fd, s390_pci_err_handler, NULL, vfio_pci);
>
> Shouldn't we check ccdf_err_length before installing the handler ? 
> because
> it won't run cleanly anyhow.
>
> Thanks,
>
> C.
>
I can move the ccdf check before installing the handler.

Thanks

Farhan

>
>> +
>> +    return true;
>> +}
>> diff --git a/include/hw/s390x/s390-pci-bus.h 
>> b/include/hw/s390x/s390-pci-bus.h
>> index 9228523ce8..c2348ede86 100644
>> --- a/include/hw/s390x/s390-pci-bus.h
>> +++ b/include/hw/s390x/s390-pci-bus.h
>> @@ -364,6 +364,7 @@ struct S390PCIBusDevice {
>>       bool forwarding_assist;
>>       bool aif;
>>       bool rtr_avail;
>> +    uint32_t ccdf_err_length;
>>       QTAILQ_ENTRY(S390PCIBusDevice) link;
>>   };
>>   diff --git a/include/hw/s390x/s390-pci-vfio.h 
>> b/include/hw/s390x/s390-pci-vfio.h
>> index f7d6149daf..c7886b63ea 100644
>> --- a/include/hw/s390x/s390-pci-vfio.h
>> +++ b/include/hw/s390x/s390-pci-vfio.h
>> @@ -20,5 +20,6 @@ S390PCIDMACount 
>> *s390_pci_start_dma_count(S390pciState *s,
>>   void s390_pci_end_dma_count(S390pciState *s, S390PCIDMACount *cnt);
>>   bool s390_pci_get_host_fh(S390PCIBusDevice *pbdev, uint32_t *fh);
>>   void s390_pci_get_clp_info(S390PCIBusDevice *pbdev);
>> +bool s390_pci_setup_err_handler(S390PCIBusDevice *pbdev, Error **errp);
>>     #endif
>


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices
  2026-09-16  8:13   ` Cédric Le Goater
  2026-09-16 17:52     ` Farhan Ali
@ 2026-09-17  1:21     ` Matthew Rosato
  2026-09-17  6:51       ` Cédric Le Goater
  1 sibling, 1 reply; 9+ messages in thread
From: Matthew Rosato @ 2026-09-17  1:21 UTC (permalink / raw)
  To: Cédric Le Goater, Farhan Ali, qemu-devel, qemu-s390x, farman
  Cc: cohuck, alex, armbru


>> +    }
>> +
>> +    fd = event_notifier_get_fd(&vfio_pci->err_notifier);
>> +    qemu_set_fd_handler(fd, s390_pci_err_handler, NULL, vfio_pci);

@Eric - Since this version of the series is now all living in s390x
code, I expect it will eventually go via an s390x PR -- but I'd like to
make sure we at minimum get an ACK from Alex and/or Cédric for this
patch due to the change above.

I already see Cédric giving review comments (thank you!) but wanted to
make the requirement explicit.
Thanks,
Matt


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices
  2026-09-17  1:21     ` Matthew Rosato
@ 2026-09-17  6:51       ` Cédric Le Goater
  0 siblings, 0 replies; 9+ messages in thread
From: Cédric Le Goater @ 2026-09-17  6:51 UTC (permalink / raw)
  To: Matthew Rosato, Farhan Ali, qemu-devel, qemu-s390x, farman
  Cc: cohuck, alex, armbru

On 9/17/26 03:21, Matthew Rosato wrote:
> 
>>> +    }
>>> +
>>> +    fd = event_notifier_get_fd(&vfio_pci->err_notifier);
>>> +    qemu_set_fd_handler(fd, s390_pci_err_handler, NULL, vfio_pci);
> 
> @Eric - Since this version of the series is now all living in s390x
> code, I expect it will eventually go via an s390x PR -- but I'd like to
> make sure we at minimum get an ACK from Alex and/or Cédric for this
> patch due to the change above.
> 
> I already see Cédric giving review comments (thank you!) but wanted to
> make the requirement explicit.

Sure.

Nothing major on my side. The 'pci_aer=true' is implicit and the Error
handling a bit weird.

Thanks,

C.




^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-17  6:51 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-14 17:44 [PATCH v5 0/3] Error recovery for zPCI passthrough devices Farhan Ali
2026-09-14 17:44 ` [PATCH v5 1/3] linux-headers: Update Linux header to 7.3-rc3 Farhan Ali
2026-09-15 22:32   ` Farhan Ali
2026-09-14 17:44 ` [PATCH v5 2/3] s390x/pci: Add PCI error handling for vfio pci devices Farhan Ali
2026-09-16  8:13   ` Cédric Le Goater
2026-09-16 17:52     ` Farhan Ali
2026-09-17  1:21     ` Matthew Rosato
2026-09-17  6:51       ` Cédric Le Goater
2026-09-14 17:44 ` [PATCH v5 3/3] s390x/pci: Reset a device in error state Farhan Ali

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.