* [PATCH v3 0/6] Add support to handle memory double-bit ecc errors
@ 2026-09-28 6:18 Riana Tauro
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
` (8 more replies)
0 siblings, 9 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
This series introduces the firmware commands required to handle
uncorrectable memory double-bit ecc errors.
When a error is detected, the driver determines whether the affected
page can be offlined and sends an offline/remove request to firmware
via sysctrl interface. Pages belonging to critical Bos or stolen memory
require a Secondary Bus Reset(SBR) to recover.
Pages configured for log-only handling are removed from queue.
For the rest of the valid pages, the first error is a poison error and a
remove is sent to firmware to avoid marking the page as bad and to
remove from queue. On the second occurence at the same address, we
request the firmware to offline the page as this would indicate a
Double-bit ECC error. Any subsequent errors at the same address should be
because of memory scrubber and a remove is sent to firmware to remove the
page from queue. This is done by tracking the firmware offlined pages by Xe KMD.
At module load, the driver queries firmware for both the offline list
and the queue and processes them accordingly.
Rev2: Rebase
fix sigid logs
fix sashiko review comments
Rev3: add a common function for queue and list
fix other review comments
Riana Tauro (6):
drm/xe/xe_ras: Handle page offline requests for device memory ecc
errors
drm/xe/xe_ras: Add support to query page offline queue and list
drm/xe: Separate drm-ras netlink data from device and firmware RAS
state
drm/xe/xe_ras: Add function to get maximum pages firmware can store
drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace
drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
drivers/gpu/drm/xe/xe_device.c | 4 +-
drivers/gpu/drm/xe/xe_device_types.h | 11 +-
drivers/gpu/drm/xe/xe_drm_ras.c | 16 +-
drivers/gpu/drm/xe/xe_drm_ras_types.h | 3 -
drivers/gpu/drm/xe/xe_hw_error.c | 6 +-
drivers/gpu/drm/xe/xe_ras.c | 280 +++++++++++++++++-
drivers/gpu/drm/xe/xe_ras.h | 5 +-
drivers/gpu/drm/xe/xe_ras_types.h | 85 ++++++
drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 6 +
drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 9 +-
drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h | 2 -
11 files changed, 387 insertions(+), 40 deletions(-)
--
2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-09-28 6:34 ` sashiko-bot
` (2 more replies)
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
` (7 subsequent siblings)
8 siblings, 3 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
Add basic support for sending page offline/remove requests to system
controller and use it for device memory ECC error handling.
Pages that belong to critical BOs cannot be handled by offlining and
require a SBR (Secondary Bus Reset).
Pages that are configured for log-only handling are not marked as bad by
firmware.
For all other valid page addresses, the first occurrence of error
indicates a poison error and the page is offlined only by software.
Firmware avoids permanently marking the page as bad. The second occurrence
of an error indicates a Double-bit ECC error and the firmware
permanently marks the page as bad.
Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
---
v2: use ret in sigid logging (Mallesh)
remove additional log
use xe_assert (Michal)
v3: align address to PAGE_SIZE (sashiko, Himal)
rename decline to remove
add more descriptive logs (Himal)
---
drivers/gpu/drm/xe/xe_ras.c | 127 +++++++++++++++++-
drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++
drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 2 +
3 files changed, 159 insertions(+), 5 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index 7a85735c57d5..1225c561a872 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -3,6 +3,8 @@
* Copyright © 2026 Intel Corporation
*/
+#include "xe_assert.h"
+#include "xe_bo.h"
#include "xe_configfs.h"
#include "xe_debugfs.h"
#include "xe_device.h"
@@ -16,6 +18,7 @@
#include "xe_sysctrl_event_types.h"
#include "xe_sysctrl_mailbox.h"
#include "xe_sysctrl_mailbox_types.h"
+#include "xe_ttm_vram_mgr.h"
#define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
/*
@@ -201,6 +204,119 @@ static inline const char *comp_to_str(u8 component)
return xe_ras_components[component];
}
+static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
+ enum xe_ras_page_action action)
+{
+ struct xe_sysctrl_mailbox_command command = {0};
+ struct xe_ras_page_offline_request request = {0};
+ struct xe_ras_page_offline_response response = {0};
+ size_t rlen;
+ int ret;
+
+ if (!xe->info.has_sysctrl)
+ return 0;
+
+ xe_assert(xe, action < XE_RAS_PAGE_ACTION_MAX);
+
+ request.page_address = page_address;
+ request.action = action;
+
+ if (action == XE_RAS_PAGE_ACTION_OFFLINE)
+ xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to offline page 0x%llx\n",
+ page_address);
+ else
+ xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to remove page 0x%llx from queue\n",
+ page_address);
+
+ xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP, XE_SYSCTRL_CMD_PAGE_OFFLINE,
+ &request, sizeof(request), &response, sizeof(response));
+
+ ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
+ if (ret) {
+ xe_log_err(xe, SYSCTRL, ret, "failed to send page offline command\n");
+ return ret;
+ }
+
+ if (rlen != sizeof(response)) {
+ xe_log_err(xe, SYSCTRL, -EINVAL,
+ "unexpected page offline response length %zu (expected %zu)\n",
+ rlen, sizeof(response));
+ return -EINVAL;
+ }
+
+ ret = ras_status_to_errno(response.status);
+ if (ret)
+ xe_log_err(xe, SYSCTRL, ret, "page offline command failed with status %u\n",
+ response.status);
+
+ return ret;
+}
+
+static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send_cmd)
+{
+ enum xe_ras_page_action action;
+ u64 addr;
+ int ret = 0;
+
+ if (!IS_ALIGNED(page_address, XE_PAGE_SIZE)) {
+ xe_log_err(xe, SYSCTRL, -EINVAL, "Unaligned physical page address: 0x%llx\n",
+ page_address);
+ return -EINVAL;
+ }
+
+ addr = ALIGN_DOWN(page_address, PAGE_SIZE);
+
+ ret = xe_ttm_vram_handle_addr_fault(xe, addr);
+
+ /*
+ * Handle return code from address fault handling function:
+ * 0: Page softofflined, remove from firmware queue
+ * -EIO: Address belongs to a critical BO/stolen area that cannot be offlined
+ * -EOPNOTSUPP: Address is valid and can be offlined but user policy is not to offline
+ * -EEXIST: Address is soft offlined but yet to be offlined by firmware for second
+ * occurrence
+ */
+
+ switch (ret) {
+ case 0:
+ action = XE_RAS_PAGE_ACTION_REMOVE;
+ xe_log_err(xe, DEVICE_MEMORY, ret,
+ "Poison detected at physical address 0x%llx, page soft-offlined\n",
+ page_address);
+ break;
+ /* User policy set to decline page offlining */
+ case -EOPNOTSUPP:
+ action = XE_RAS_PAGE_ACTION_REMOVE;
+ xe_log_err(xe, DEVICE_MEMORY, ret,
+ "Poison detected at physical address 0x%llx, user policy set to decline soft-offlining\n",
+ page_address);
+ break;
+ case -EIO:
+ xe_log_err(xe, DEVICE_MEMORY, ret,
+ "Poison detected at physical address 0x%llx, page belongs to critical BO and cannot be soft-offlined\n",
+ page_address);
+ return ret;
+ case -EEXIST:
+ action = XE_RAS_PAGE_ACTION_OFFLINE;
+ xe_log_err(xe, DEVICE_MEMORY, ret,
+ "Double-bit ECC error detected at physical address 0x%llx, page soft-offlined\n",
+ page_address);
+ break;
+ default:
+ xe_log_err(xe, DEVICE_MEMORY, ret, "Failed to handle address fault at physical address 0x%llx\n",
+ page_address);
+ return 0;
+ }
+
+ if (send_cmd) {
+ ret = send_page_offline_cmd(xe, page_address, action);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class *counter)
{
u8 severity = counter->common.severity;
@@ -368,11 +484,12 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a
static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_array *arr)
{
struct xe_ras_memory_error *info = (void *)arr->details;
+ int ret;
/*
* For memory errors, the recovery action depends on the error category
*
- * TODO: Double-bit ECC errors: Page offlining
+ * Double-bit ECC errors: Page offlining
* Poison and data parity errors: Log only
* For any other memory errors, request a reset as recovery mechanism
*/
@@ -384,10 +501,10 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_
xe_info(xe, "[RAS]: Data parity error detected\n");
break;
case XE_RAS_MEMORY_DB_ECC:
- xe_info(xe, "[RAS]: Double-bit ECC error detected at sw address 0x%llx\n",
- info->sw_address);
- /* TODO: Add page offlining for Double-bit ECC error */
- fallthrough;
+ ret = handle_page_offline(xe, info->sw_address, true);
+ if (ret)
+ return XE_RAS_RECOVERY_ACTION_RESET;
+ break;
default:
return XE_RAS_RECOVERY_ACTION_RESET;
}
diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
index fe6f3658a2a4..f119489bcdf2 100644
--- a/drivers/gpu/drm/xe/xe_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_ras_types.h
@@ -17,6 +17,19 @@
#define XE_RAS_MEMORY_POISON BIT(2)
#define XE_RAS_MEMORY_DATA_PARITY BIT(5)
+/**
+ * enum xe_ras_page_action - Page offline actions for page offline request
+ *
+ * @XE_RAS_PAGE_ACTION_OFFLINE: Instruct firmware to offline the page
+ * @XE_RAS_PAGE_ACTION_REMOVE: Instruct firmware to remove the page from queue
+ * @XE_RAS_PAGE_ACTION_MAX: Max value
+ */
+enum xe_ras_page_action {
+ XE_RAS_PAGE_ACTION_OFFLINE,
+ XE_RAS_PAGE_ACTION_REMOVE,
+ XE_RAS_PAGE_ACTION_MAX
+};
+
/**
* enum xe_ras_recovery_action - RAS recovery actions
*
@@ -295,6 +308,28 @@ struct xe_ras_memory_error {
u32 reserved2[10];
} __packed;
+/**
+ * struct xe_ras_page_offline_request - Request for page offline command
+ */
+struct xe_ras_page_offline_request {
+ /** @page_address: Page address (4KB aligned) */
+ u64 page_address;
+ /** @action: Action to be performed, see &enum xe_ras_page_action */
+ u32 action;
+ /** @reserved: Reserved for future use */
+ u32 reserved;
+} __packed;
+
+/**
+ * struct xe_ras_page_offline_response - Response from page offline command
+ */
+struct xe_ras_page_offline_response {
+ /** @status: Status of the page offline request */
+ u32 status;
+ /** @reserved: Reserved for future use */
+ u32 reserved;
+} __packed;
+
/**
* struct xe_ras_get_health_request - Request structure for obtaining gpu health
*/
diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
index c236e5377f30..a01576bf2e73 100644
--- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
+++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
@@ -30,6 +30,7 @@ enum xe_sysctrl_group {
* @XE_SYSCTRL_CMD_GET_THRESHOLD: Retrieve error threshold
* @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
* @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
+ * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a page
* @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
* @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
*/
@@ -40,6 +41,7 @@ enum xe_sysctrl_gfsp_cmd {
XE_SYSCTRL_CMD_GET_THRESHOLD = 0x05,
XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
+ XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
};
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-09-28 6:35 ` sashiko-bot
` (2 more replies)
2026-09-28 6:18 ` [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state Riana Tauro
` (6 subsequent siblings)
8 siblings, 3 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
Add support to query page offline list and queue from firmware
during module load. The page offline list command retrieves pages that
are already offlined by the firmware. The page offline queue command
retrieves the pages pending to be offlined by the firmware.
Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
---
v2: rebase
store total pages once per response (Sashiko)
v3: common function for offline and queue (Himal)
---
drivers/gpu/drm/xe/xe_ras.c | 74 +++++++++++++++++++
drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++++++
drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 4 +
3 files changed, 113 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index 1225c561a872..752754f09314 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -335,6 +335,77 @@ static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class
return true;
}
+static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t req_size,
+ void *resp, size_t resp_size,
+ struct xe_ras_offline_common *common, bool offline)
+{
+ struct xe_sysctrl_mailbox_command command = {0};
+ struct xe_ras_offline_list_request *list_req;
+ u32 total_pages = 0, count = 0;
+ ssize_t rlen;
+ int ret, i;
+
+ list_req = req ? req : NULL;
+
+ xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP, cmd, req, req_size, resp,
+ resp_size);
+
+ do {
+ memset(resp, 0, resp_size);
+
+ if (list_req)
+ list_req->index = count;
+
+ ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
+ if (ret) {
+ xe_log_err(xe, SYSCTRL, ret, "failed to get page offline data, cmd=%#x\n",
+ cmd);
+ return;
+ }
+
+ if (rlen != resp_size) {
+ xe_log_err(xe, SYSCTRL, -EINVAL,
+ "unexpected page offline response length %zu (expected %zu), cmd=%#x\n",
+ rlen, resp_size, cmd);
+ return;
+ }
+
+ for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
+ handle_page_offline(xe, common->page_addresses[i], offline);
+
+ count += common->pages_returned;
+ if (!common->pages_returned)
+ break;
+
+ if (!total_pages)
+ total_pages = common->total_pages;
+
+ if (count > total_pages) {
+ xe_log_err(xe, SYSCTRL, -EINVAL,
+ "Pages returned exceed total pages %u, returned %u, cmd=%#x\n",
+ total_pages, count, cmd);
+ return;
+ }
+ } while (common->additional_data);
+}
+
+static void get_queued_pages(struct xe_device *xe)
+{
+ struct xe_ras_offline_common response = {0};
+
+ get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL, 0, &response,
+ sizeof(response), &response, true);
+}
+
+static void get_offlined_list(struct xe_device *xe)
+{
+ struct xe_ras_offline_list_response response = {0};
+ struct xe_ras_offline_list_request request = {0};
+
+ get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request, sizeof(request),
+ &response, sizeof(response), &response.common, false);
+}
+
static struct pci_dev *find_usp_dev(struct pci_dev *pdev)
{
struct pci_dev *vsp;
@@ -1049,6 +1120,9 @@ void xe_ras_init(struct xe_device *xe)
if (IS_ENABLED(CONFIG_PCIEAER))
ras_usp_aer_init(xe);
+ get_queued_pages(xe);
+ get_offlined_list(xe);
+
ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
if (ret)
xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
index f119489bcdf2..021ffbd6d4e2 100644
--- a/drivers/gpu/drm/xe/xe_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_ras_types.h
@@ -10,6 +10,7 @@
#define XE_RAS_NUM_COUNTERS 16
#define XE_RAS_NUM_ERROR_ARR 3
+#define XE_RAS_NUM_PAGES 25
/* Error bits in IEH global error status register */
#define XE_RAS_SOC_IEH_PUNIT BIT(1)
/* Device memory error categories */
@@ -330,6 +331,40 @@ struct xe_ras_page_offline_response {
u32 reserved;
} __packed;
+/**
+ * struct xe_ras_offline_common - Common structure for offline list and queue
+ */
+struct xe_ras_offline_common {
+ /** @total_pages: Total number of queued pages */
+ u32 total_pages;
+ /** @pages_returned: Number of pages returned in this response */
+ u32 pages_returned;
+ /** @page_addresses: Array of page addresses (4KB aligned) */
+ u64 page_addresses[XE_RAS_NUM_PAGES];
+ /** @additional_data: Indicates if more data is available */
+ u8 additional_data;
+ /** @reserved: Reserved for future use */
+ u8 reserved[3];
+} __packed;
+
+/**
+ * struct xe_ras_offline_list_request - Request for get offline list command
+ */
+struct xe_ras_offline_list_request {
+ /** @index: Zero-based index into the offline page list */
+ u32 index;
+} __packed;
+
+/**
+ * struct xe_ras_offline_list_response - Response from get offline list command
+ */
+struct xe_ras_offline_list_response {
+ /** @max_entries: Total no of pages that can be stored in flash */
+ u32 max_entries;
+ /** @common: Common offline page information */
+ struct xe_ras_offline_common common;
+} __packed;
+
/**
* struct xe_ras_get_health_request - Request structure for obtaining gpu health
*/
diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
index a01576bf2e73..3a71ed446949 100644
--- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
+++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
@@ -31,6 +31,8 @@ enum xe_sysctrl_group {
* @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
* @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
* @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a page
+ * @XE_SYSCTRL_CMD_GET_OFFLINE_LIST: Retrieve list of all offlined pages from flash
+ * @XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE: Retrieve list of offlined queued pages from firmware
* @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
* @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
*/
@@ -42,6 +44,8 @@ enum xe_sysctrl_gfsp_cmd {
XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
+ XE_SYSCTRL_CMD_GET_OFFLINE_LIST = 0x09,
+ XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE = 0x0A,
XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
};
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-09-28 6:51 ` sashiko-bot
2026-09-28 8:59 ` Ghimiray, Himal Prasad
2026-09-28 6:18 ` [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store Riana Tauro
` (5 subsequent siblings)
8 siblings, 2 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
Keep the DRM RAS data used for netlink reporting separate from the
device and firmware RAS state in xe_device.
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
---
v2: remove inline function (Himal)
---
drivers/gpu/drm/xe/xe_device_types.h | 11 +++++++++--
drivers/gpu/drm/xe/xe_drm_ras.c | 16 ++++++++--------
drivers/gpu/drm/xe/xe_drm_ras_types.h | 3 ---
drivers/gpu/drm/xe/xe_hw_error.c | 6 +++---
drivers/gpu/drm/xe/xe_ras.c | 26 ++++++++++++++++++--------
drivers/gpu/drm/xe/xe_ras.h | 2 ++
drivers/gpu/drm/xe/xe_ras_types.h | 10 ++++++++++
drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 3 ++-
8 files changed, 52 insertions(+), 25 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h
index bb4234f00453..c0b53af9bdd6 100644
--- a/drivers/gpu/drm/xe/xe_device_types.h
+++ b/drivers/gpu/drm/xe/xe_device_types.h
@@ -21,6 +21,7 @@
#include "xe_platform_types.h"
#include "xe_pmu_types.h"
#include "xe_pt_types.h"
+#include "xe_ras_types.h"
#include "xe_sriov_pf_types.h"
#include "xe_sriov_types.h"
#include "xe_sriov_vf_types.h"
@@ -570,8 +571,14 @@ struct xe_device {
/** @pmu: performance monitoring unit */
struct xe_pmu pmu;
- /** @ras: RAS structure for device */
- struct xe_drm_ras ras;
+ /** @ras: RAS (Reliability, Availability, Serviceability) structures */
+ struct {
+ /** @ras.nl_data: drm-ras netlink data */
+ struct xe_drm_ras nl_data;
+
+ /** @ras.state: RAS device and firmware state */
+ struct xe_ras_state state;
+ } ras;
/** @i2c: I2C host controller */
struct xe_i2c *i2c;
diff --git a/drivers/gpu/drm/xe/xe_drm_ras.c b/drivers/gpu/drm/xe/xe_drm_ras.c
index 7f3695707611..38d77561facf 100644
--- a/drivers/gpu/drm/xe/xe_drm_ras.c
+++ b/drivers/gpu/drm/xe/xe_drm_ras.c
@@ -20,7 +20,7 @@ static int query_error_counter(struct xe_device *xe,
enum drm_xe_ras_error_severity severity,
u32 error_id, const char **name, u32 *val)
{
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
if (!info || !info[error_id].name)
@@ -41,7 +41,7 @@ static int clear_error_counter(struct xe_device *xe,
enum drm_xe_ras_error_severity severity,
u32 error_id)
{
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
if (!info || !info[error_id].name)
@@ -90,7 +90,7 @@ static int query_correctable_error_threshold(struct drm_ras_node *ep, u32 error_
const char **name, u32 *threshold)
{
struct xe_device *xe = ep->priv;
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[DRM_XE_RAS_ERR_SEV_CORRECTABLE];
if (!info || !info[error_id].name)
@@ -106,7 +106,7 @@ static int query_correctable_error_threshold(struct drm_ras_node *ep, u32 error_
static int set_correctable_error_threshold(struct drm_ras_node *ep, u32 error_id, u32 threshold)
{
struct xe_device *xe = ep->priv;
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[DRM_XE_RAS_ERR_SEV_CORRECTABLE];
if (!info || !info[error_id].name)
@@ -142,7 +142,7 @@ static int assign_node_params(struct xe_device *xe, struct drm_ras_node *node,
const enum drm_xe_ras_error_severity severity)
{
struct pci_dev *pdev = to_pci_dev(xe->drm.dev);
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
const char *device_name;
device_name = kasprintf(GFP_KERNEL, "%04x:%02x:%02x.%d",
@@ -190,7 +190,7 @@ static void cleanup_node(struct drm_device *drm, void *node)
static int register_nodes(struct xe_device *xe)
{
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct drm_ras_node *node;
int i, ret;
@@ -230,7 +230,7 @@ static int register_nodes(struct xe_device *xe)
*/
void xe_drm_ras_event(struct xe_device *xe, u32 component, u32 severity, u32 value)
{
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
struct drm_ras_node *node;
int ret;
@@ -260,7 +260,7 @@ void xe_drm_ras_event(struct xe_device *xe, u32 component, u32 severity, u32 val
*/
int xe_drm_ras_init(struct xe_device *xe)
{
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct drm_ras_node *node;
int err;
diff --git a/drivers/gpu/drm/xe/xe_drm_ras_types.h b/drivers/gpu/drm/xe/xe_drm_ras_types.h
index 0be218ba2db7..8d729ad6a264 100644
--- a/drivers/gpu/drm/xe/xe_drm_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_drm_ras_types.h
@@ -43,9 +43,6 @@ struct xe_drm_ras {
/** @info: info array for all types of errors */
struct xe_drm_ras_counter *info[DRM_XE_RAS_ERR_SEV_MAX];
-
- /** @disable_vram_page_offline: cached configfs policy, immutable after init */
- bool disable_vram_page_offline;
};
#endif
diff --git a/drivers/gpu/drm/xe/xe_hw_error.c b/drivers/gpu/drm/xe/xe_hw_error.c
index 5f2abc9485ff..f53a6b3055de 100644
--- a/drivers/gpu/drm/xe/xe_hw_error.c
+++ b/drivers/gpu/drm/xe/xe_hw_error.c
@@ -240,7 +240,7 @@ static void log_soc_error(struct xe_tile *tile, const char * const *reg_info,
{
const char *severity_str = error_severity[severity];
struct xe_device *xe = tile_to_xe(tile);
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
const char *name;
@@ -260,7 +260,7 @@ static void gt_hw_error_handler(struct xe_tile *tile, const enum hardware_error
{
const enum drm_xe_ras_error_severity severity = hw_err_to_severity(hw_err);
struct xe_device *xe = tile_to_xe(tile);
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
struct xe_mmio *mmio = &tile->mmio;
unsigned long err_stat = 0;
@@ -422,7 +422,7 @@ static void hw_error_source_handler(struct xe_tile *tile, const enum hardware_er
const enum drm_xe_ras_error_severity severity = hw_err_to_severity(hw_err);
const char *severity_str = error_severity[severity];
struct xe_device *xe = tile_to_xe(tile);
- struct xe_drm_ras *ras = &xe->ras;
+ struct xe_drm_ras *ras = &xe->ras.nl_data;
struct xe_drm_ras_counter *info = ras->info[severity];
unsigned long flags, err_src;
u32 err_bit;
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index 752754f09314..9a1a8190578b 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -1094,6 +1094,17 @@ static const struct attribute_group gpu_health_group = {
.attrs = gpu_health_attrs,
};
+/**
+ * xe_ras_get_disable_page_offline - Get the page offline user policy
+ * @xe: xe device instance
+ *
+ * Return: true if the page offline policy is disabled, false otherwise.
+ */
+bool xe_ras_get_disable_page_offline(struct xe_device *xe)
+{
+ return xe->ras.state.disable_page_offline;
+}
+
/**
* xe_ras_init - Initialize Xe RAS
* @xe: xe device instance
@@ -1104,19 +1115,18 @@ void xe_ras_init(struct xe_device *xe)
{
int ret;
- /*
- * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
- * once the feature flag is plumbed through device info.
- */
- if (xe->info.platform == XE_CRESCENTISLAND)
- xe->ras.disable_vram_page_offline =
- xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
-
xe_drm_ras_init(xe);
if (!xe->info.has_sysctrl)
return;
+ /*
+ * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
+ * once the feature flag is plumbed through device info.
+ */
+ xe->ras.state.disable_page_offline =
+ xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
+
if (IS_ENABLED(CONFIG_PCIEAER))
ras_usp_aer_init(xe);
diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
index 0b8669f28d56..ac1253d6ead9 100644
--- a/drivers/gpu/drm/xe/xe_ras.h
+++ b/drivers/gpu/drm/xe/xe_ras.h
@@ -7,6 +7,7 @@
#define _XE_RAS_H_
#include <linux/types.h>
+
#include "xe_ras_types.h"
struct xe_device;
@@ -20,5 +21,6 @@ int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8 component, u32 *t
int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8 component, u32 threshold);
void xe_ras_init(struct xe_device *xe);
enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe);
+bool xe_ras_get_disable_page_offline(struct xe_device *xe);
#endif
diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
index 021ffbd6d4e2..40224db38906 100644
--- a/drivers/gpu/drm/xe/xe_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_ras_types.h
@@ -406,4 +406,14 @@ struct xe_ras_set_health_response {
/** @reserved1: Reserved for future use */
u32 reserved1[2];
} __packed;
+
+/* Device structures */
+
+/**
+ * struct xe_ras_state - RAS device and firmware state
+ */
+struct xe_ras_state {
+ /** @disable_page_offline: cached configfs policy, immutable after init */
+ bool disable_page_offline;
+};
#endif
diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
index 9a514d983e90..2c4722a956a0 100644
--- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
+++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
@@ -24,6 +24,7 @@
#include "xe_mmio.h"
#include "xe_pm.h"
#include "xe_printk.h"
+#include "xe_ras.h"
#include "xe_res_cursor.h"
#include "xe_ttm_stolen_mgr.h"
#include "xe_ttm_vram_mgr.h"
@@ -933,7 +934,7 @@ int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr)
vram_mgr = &vr->ttm;
mm = &vram_mgr->mm;
- if (xe->ras.disable_vram_page_offline) {
+ if (xe_ras_get_disable_page_offline(xe)) {
xe_err(xe, "0x%llx is reported as corrupted address by HW\n",
addr);
return -EOPNOTSUPP;
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (2 preceding siblings ...)
2026-09-28 6:18 ` [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-09-28 6:43 ` sashiko-bot
2026-10-01 11:36 ` Upadhyay, Tejas
2026-09-28 6:18 ` [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace Riana Tauro
` (4 subsequent siblings)
8 siblings, 2 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
Add function to get maximum number of pages that firmware can store for
offline tracking. This will be used to report max pages to userspace.
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
v2: fix commit message (Himal)
---
drivers/gpu/drm/xe/xe_ras.c | 15 +++++++++++++++
drivers/gpu/drm/xe/xe_ras.h | 1 +
drivers/gpu/drm/xe/xe_ras_types.h | 2 ++
3 files changed, 18 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index 9a1a8190578b..cfc7f07e9ced 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -401,9 +401,13 @@ static void get_offlined_list(struct xe_device *xe)
{
struct xe_ras_offline_list_response response = {0};
struct xe_ras_offline_list_request request = {0};
+ struct xe_ras_state *state = &xe->ras.state;
get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request, sizeof(request),
&response, sizeof(response), &response.common, false);
+
+ if (response.max_entries)
+ state->max_pages = response.max_entries;
}
static struct pci_dev *find_usp_dev(struct pci_dev *pdev)
@@ -1105,6 +1109,17 @@ bool xe_ras_get_disable_page_offline(struct xe_device *xe)
return xe->ras.state.disable_page_offline;
}
+/**
+ * xe_ras_get_max_pages - Get the maximum number of pages
+ * @xe: xe device instance
+ *
+ * Return: Maximum number of pages that can be stored in flash
+ */
+u32 xe_ras_get_max_pages(struct xe_device *xe)
+{
+ return xe->ras.state.max_pages;
+}
+
/**
* xe_ras_init - Initialize Xe RAS
* @xe: xe device instance
diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
index ac1253d6ead9..561c656e2ad3 100644
--- a/drivers/gpu/drm/xe/xe_ras.h
+++ b/drivers/gpu/drm/xe/xe_ras.h
@@ -22,5 +22,6 @@ int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8 component, u32 th
void xe_ras_init(struct xe_device *xe);
enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe);
bool xe_ras_get_disable_page_offline(struct xe_device *xe);
+u32 xe_ras_get_max_pages(struct xe_device *xe);
#endif
diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
index 40224db38906..e09b50a6f77b 100644
--- a/drivers/gpu/drm/xe/xe_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_ras_types.h
@@ -415,5 +415,7 @@ struct xe_ras_set_health_response {
struct xe_ras_state {
/** @disable_page_offline: cached configfs policy, immutable after init */
bool disable_page_offline;
+ /** @max_pages: Total number of pages that can be stored by firmware */
+ u32 max_pages;
};
#endif
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (3 preceding siblings ...)
2026-09-28 6:18 ` [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-10-01 11:37 ` Upadhyay, Tejas
2026-09-28 6:18 ` [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates Riana Tauro
` (3 subsequent siblings)
8 siblings, 1 reply; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
Report the maximum number of pages returned by firmware at probe,
resolving the existing TODO.
Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
v2: fix commit message (Himal)
---
drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 6 +-----
drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h | 2 --
2 files changed, 1 insertion(+), 7 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
index 2c4722a956a0..cc33ba7e23c2 100644
--- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
+++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
@@ -1012,11 +1012,7 @@ static int vram_bad_pages_show(struct seq_file *m, void *unused)
struct xe_tile *tile;
u8 id;
- man = ttm_manager_type(&xe->ttm, XE_PL_VRAM0);
- if (man)
- /* TODO Hook with RAS to show max_pages fetched from FW */
- seq_printf(m, "max_pages: %d\n",
- to_xe_ttm_vram_mgr(man)->max_pages);
+ seq_printf(m, "max_pages: %u\n", xe_ras_get_max_pages(xe));
for_each_tile(tile, xe, id) {
struct xe_vram_region *vr = tile->mem.vram;
diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h b/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
index efcf3e1d4e80..dc97b0ad0e51 100644
--- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
+++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
@@ -37,8 +37,6 @@ struct xe_ttm_vram_mgr {
struct mutex lock;
/** @mem_type: The TTM memory type */
u32 mem_type;
- /** @max_pages: max pages that can be in offline queue retrieved from FW */
- u16 max_pages;
};
/**
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (4 preceding siblings ...)
2026-09-28 6:18 ` [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace Riana Tauro
@ 2026-09-28 6:18 ` Riana Tauro
2026-09-28 7:07 ` sashiko-bot
2026-09-28 9:00 ` Ghimiray, Himal Prasad
2026-09-28 14:42 ` ✓ CI.KUnit: success for Add support to handle memory double-bit ecc errors (rev3) Patchwork
` (2 subsequent siblings)
8 siblings, 2 replies; 28+ messages in thread
From: Riana Tauro @ 2026-09-28 6:18 UTC (permalink / raw)
To: intel-xe
Cc: riana.tauro, anshuman.gupta, rodrigo.vivi, aravind.iddamsetty,
badal.nilawar, raag.jadav, ravi.kishore.koppuravuri,
mallesh.koujalagi, tejas.upadhyay, himal.prasad.ghimiray
A memory scrubber can report multiple errors at the same address. Track
pages already successfully offlined by the firmware so that subsequent
reports for the same address are removed from the firmware queue instead
of issuing redundant page offline requests.
Signed-off-by: Riana Tauro <riana.tauro@intel.com>
---
v2: use xe page shift (Sashiko)
---
drivers/gpu/drm/xe/xe_device.c | 4 ++-
drivers/gpu/drm/xe/xe_ras.c | 44 ++++++++++++++++++++++++++++---
drivers/gpu/drm/xe/xe_ras.h | 2 +-
drivers/gpu/drm/xe/xe_ras_types.h | 3 +++
4 files changed, 47 insertions(+), 6 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
index 205cb4e7f9e8..1ec7c74cbd58 100644
--- a/drivers/gpu/drm/xe/xe_device.c
+++ b/drivers/gpu/drm/xe/xe_device.c
@@ -957,7 +957,9 @@ int xe_device_probe(struct xe_device *xe)
if (err)
return err;
- xe_ras_init(xe);
+ err = xe_ras_init(xe);
+ if (err)
+ return err;
/*
* Now that GT is initialized (TTM in particular),
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index cfc7f07e9ced..7dee1c804eeb 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -20,6 +20,9 @@
#include "xe_sysctrl_mailbox_types.h"
#include "xe_ttm_vram_mgr.h"
+/* Any non-null entry marks the page as offlined by firmware */
+#define XE_RAS_PAGE_OFFLINED xa_mk_value(1)
+
#define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
/*
* Uncorrectable error type for core compute errors.
@@ -210,6 +213,7 @@ static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
struct xe_sysctrl_mailbox_command command = {0};
struct xe_ras_page_offline_request request = {0};
struct xe_ras_page_offline_response response = {0};
+ struct xe_ras_state *state = &xe->ras.state;
size_t rlen;
int ret;
@@ -249,11 +253,16 @@ static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
xe_log_err(xe, SYSCTRL, ret, "page offline command failed with status %u\n",
response.status);
+ if (action == XE_RAS_PAGE_ACTION_OFFLINE)
+ xa_store(&state->offlined_pages, page_address >> XE_PTE_SHIFT,
+ XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
+
return ret;
}
static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send_cmd)
{
+ struct xe_ras_state *state = &xe->ras.state;
enum xe_ras_page_action action;
u64 addr;
int ret = 0;
@@ -297,7 +306,11 @@ static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send
page_address);
return ret;
case -EEXIST:
- action = XE_RAS_PAGE_ACTION_OFFLINE;
+ if (xa_load(&state->offlined_pages, page_address >> XE_PTE_SHIFT))
+ action = XE_RAS_PAGE_ACTION_REMOVE;
+ else
+ action = XE_RAS_PAGE_ACTION_OFFLINE;
+
xe_log_err(xe, DEVICE_MEMORY, ret,
"Double-bit ECC error detected at physical address 0x%llx, page soft-offlined\n",
page_address);
@@ -341,6 +354,8 @@ static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t r
{
struct xe_sysctrl_mailbox_command command = {0};
struct xe_ras_offline_list_request *list_req;
+ struct xe_ras_state *state = &xe->ras.state;
+ unsigned long index;
u32 total_pages = 0, count = 0;
ssize_t rlen;
int ret, i;
@@ -370,8 +385,15 @@ static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t r
return;
}
- for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
+ for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++) {
handle_page_offline(xe, common->page_addresses[i], offline);
+ /* The pages are already offlined by firmware */
+ if (!offline) {
+ index = common->page_addresses[i] >> XE_PTE_SHIFT;
+ xa_store(&state->offlined_pages, index,
+ XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
+ }
+ }
count += common->pages_returned;
if (!common->pages_returned)
@@ -1098,6 +1120,13 @@ static const struct attribute_group gpu_health_group = {
.attrs = gpu_health_attrs,
};
+static void ras_fini(void *arg)
+{
+ struct xe_device *xe = arg;
+
+ xa_destroy(&xe->ras.state.offlined_pages);
+}
+
/**
* xe_ras_get_disable_page_offline - Get the page offline user policy
* @xe: xe device instance
@@ -1125,15 +1154,18 @@ u32 xe_ras_get_max_pages(struct xe_device *xe)
* @xe: xe device instance
*
* Initialize Xe RAS
+ *
+ * Return: 0 on success, negative error code otherwise
*/
-void xe_ras_init(struct xe_device *xe)
+int xe_ras_init(struct xe_device *xe)
{
+ struct xe_ras_state *state = &xe->ras.state;
int ret;
xe_drm_ras_init(xe);
if (!xe->info.has_sysctrl)
- return;
+ return 0;
/*
* TODO: Replace platform check with xe->info.has_disable_vram_page_offline
@@ -1145,10 +1177,14 @@ void xe_ras_init(struct xe_device *xe)
if (IS_ENABLED(CONFIG_PCIEAER))
ras_usp_aer_init(xe);
+ xa_init(&state->offlined_pages);
+
get_queued_pages(xe);
get_offlined_list(xe);
ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
if (ret)
xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
+
+ return devm_add_action_or_reset(xe->drm.dev, ras_fini, xe);
}
diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
index 561c656e2ad3..e41a52869085 100644
--- a/drivers/gpu/drm/xe/xe_ras.h
+++ b/drivers/gpu/drm/xe/xe_ras.h
@@ -19,7 +19,7 @@ int xe_ras_get_counter(struct xe_device *xe, u8 severity, u8 component, u32 *val
int xe_ras_clear_counter(struct xe_device *xe, u8 severity, u8 component);
int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8 component, u32 *threshold);
int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8 component, u32 threshold);
-void xe_ras_init(struct xe_device *xe);
+int xe_ras_init(struct xe_device *xe);
enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe);
bool xe_ras_get_disable_page_offline(struct xe_device *xe);
u32 xe_ras_get_max_pages(struct xe_device *xe);
diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
index e09b50a6f77b..50d70e2e2fd3 100644
--- a/drivers/gpu/drm/xe/xe_ras_types.h
+++ b/drivers/gpu/drm/xe/xe_ras_types.h
@@ -7,6 +7,7 @@
#define _XE_RAS_TYPES_H_
#include <linux/types.h>
+#include <linux/xarray.h>
#define XE_RAS_NUM_COUNTERS 16
#define XE_RAS_NUM_ERROR_ARR 3
@@ -417,5 +418,7 @@ struct xe_ras_state {
bool disable_page_offline;
/** @max_pages: Total number of pages that can be stored by firmware */
u32 max_pages;
+ /** @offlined_pages: XArray of pages offlined by firmware */
+ struct xarray offlined_pages;
};
#endif
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* Re: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
@ 2026-09-28 6:34 ` sashiko-bot
2026-09-28 8:55 ` Ghimiray, Himal Prasad
2026-10-01 10:50 ` Upadhyay, Tejas
2 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-28 6:34 UTC (permalink / raw)
To: Riana Tauro; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] xe_ras: unhandled page offlining failures bypass fallback GPU reset
- [Medium] xe_ras: out-of-bounds memory errors mischaracterized as user policy disable
--
commit 68428995a0f62e84ab14e112b823a469bad65984
Author: Riana Tauro <riana.tauro@intel.com>
drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
Add basic support for sending page offline/remove requests to system
controller and use it for device memory ECC error handling.
Pages that belong to critical BOs cannot be handled by offlining and
require a SBR (Secondary Bus Reset).
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 7a85735c57d58..96ad473388a95 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
[ ... ]
> @@ -201,6 +204,119 @@ static inline const char *comp_to_str(u8 component)
[ ... ]
> +static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send_cmd)
> +{
[ ... ]
> + switch (ret) {
> + case 0:
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page soft-offlined\n",
> + page_address);
> + break;
> + /* User policy set to decline page offlining */
> + case -EOPNOTSUPP:
[Severity: Medium]
When xe_ttm_vram_handle_addr_fault() encounters an out-of-bounds address, does
it also return -EOPNOTSUPP?
If firmware reports an error at an invalid physical address outside VRAM/GSM,
it looks like this will mischaracterize the error as a user policy choice and
still send a XE_RAS_PAGE_ACTION_REMOVE command.
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, user policy set to decline soft-offlining\n",
> + page_address);
> + break;
> + case -EIO:
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page belongs to critical BO and cannot be soft-offlined\n",
> + page_address);
> + return ret;
> + case -EEXIST:
> + action = XE_RAS_PAGE_ACTION_OFFLINE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Double-bit ECC error detected at physical address 0x%llx, page soft-offlined\n",
> + page_address);
> + break;
> + default:
> + xe_log_err(xe, DEVICE_MEMORY, ret, "Failed to handle address fault at physical address 0x%llx\n",
> + page_address);
> + return 0;
[Severity: High]
Does returning 0 here mask the error code from xe_ttm_vram_handle_addr_fault()?
If an internal error occurs (like -ENOMEM), returning 0 instead of the error
code causes the caller, handle_device_memory_errors(), to evaluate if (ret)
as false. This masks the offlining failure and bypasses the fallback hardware
reset, potentially leaving corrupted memory in use.
> + }
> +
> + if (send_cmd) {
> + ret = send_page_offline_cmd(xe, page_address, action);
> + if (ret)
> + return ret;
> + }
> +
> + return 0;
> +}
[ ... ]
> @@ -384,10 +501,10 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_
> xe_info(xe, "[RAS]: Data parity error detected\n");
> break;
> case XE_RAS_MEMORY_DB_ECC:
> - xe_info(xe, "[RAS]: Double-bit ECC error detected at sw address 0x%llx\n",
> - info->sw_address);
> - /* TODO: Add page offlining for Double-bit ECC error */
> - fallthrough;
> + ret = handle_page_offline(xe, info->sw_address, true);
> + if (ret)
> + return XE_RAS_RECOVERY_ACTION_RESET;
> + break;
> default:
> return XE_RAS_RECOVERY_ACTION_RESET;
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928061830.417395-8-riana.tauro@intel.com?part=1
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
@ 2026-09-28 6:35 ` sashiko-bot
2026-10-01 11:24 ` Upadhyay, Tejas
2026-10-01 12:11 ` Ghimiray, Himal Prasad
2 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-28 6:35 UTC (permalink / raw)
To: Riana Tauro; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [Medium] xe_ras: missing bounds check on pages_returned causes silent page skipping
- [Medium] xe_ras: unhandled errors in handle_page_offline abort queue fetching
- [Medium] xe_ras: integer overflow in loop termination check causes hang
- [Medium] xe_ras: missing endianness conversions in sysctrl payload structures
- [Low] build: pointer sign mismatch in xe_sysctrl_send_command call
--
commit c79936a34d5a6eb6a5f10a97e90fa4f355928529
Author: Riana Tauro <riana.tauro@intel.com>
drm/xe/xe_ras: Add support to query page offline queue and list
This patch adds support to query the list of pages offlined by firmware
and the queue of pages pending to be offlined. It implements mailbox
commands to fetch the data from the firmware during module load.
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 96ad473388a95..f1da7f983637a 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -335,6 +335,77 @@ static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class
> return true;
> }
>
> +static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t req_size,
> + void *resp, size_t resp_size,
> + struct xe_ras_offline_common *common, bool offline)
> +{
> + struct xe_sysctrl_mailbox_command command = {0};
> + struct xe_ras_offline_list_request *list_req;
> + u32 total_pages = 0, count = 0;
> + ssize_t rlen;
> + int ret, i;
> +
> + list_req = req ? req : NULL;
> +
> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP, cmd, req, req_size, resp,
> + resp_size);
> +
> + do {
> + memset(resp, 0, resp_size);
> +
> + if (list_req)
> + list_req->index = count;
[Severity: Medium]
Since the sysctrl mailbox header explicitly uses little-endian
data conversions, does the payload data also need to be little-endian?
Defining the index as a native u32 and assigning it directly without
cpu_to_le32() might send corrupted values to the firmware on big-endian
architectures.
> +
> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
[Severity: Low]
Does passing a pointer to the signed ssize_t rlen variable here cause a
pointer sign mismatch compiler warning? It looks like
xe_sysctrl_send_command() expects a size_t * for its third argument.
> + if (ret) {
> + xe_log_err(xe, SYSCTRL, ret, "failed to get page offline data, cmd=%#x\n",
> + cmd);
> + return;
> + }
> +
> + if (rlen != resp_size) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "unexpected page offline response length %zu (expected %zu), cmd=%#x\n",
> + rlen, resp_size, cmd);
> + return;
> + }
> +
> + for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
> + handle_page_offline(xe, common->page_addresses[i], offline);
[Severity: Medium]
Can handle_page_offline() fail (for instance with -EIO for a critical BO) and
return early without sending the offline command to the firmware?
Because the return value is ignored here, the page might be left in the
firmware's pending queue. Would the firmware then repeatedly return the same
unhandled pages on the next fetch, eventually causing the loop below to
abruptly abort the entire queue processing once count exceeds total_pages?
> +
> + count += common->pages_returned;
[Severity: Medium]
If the firmware returns a pages_returned value greater than XE_RAS_NUM_PAGES
(25), the loop above only processes up to 25 pages. Does incrementing count by
the full unvalidated pages_returned value incorrectly shift the list_req->index
for the next request, causing the driver to silently skip the remaining
poisoned pages?
> + if (!common->pages_returned)
> + break;
> +
> + if (!total_pages)
> + total_pages = common->total_pages;
> +
> + if (count > total_pages) {
[Severity: Medium]
Both count and total_pages are 32-bit integers. If the firmware provides
total_pages as the maximum u32 value and additional_data remains set, could
count wrap around to 0? If so, this termination check would never be true,
resulting in an infinite loop.
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "Pages returned exceed total pages %u, returned %u, cmd=%#x\n",
> + total_pages, count, cmd);
> + return;
> + }
> + } while (common->additional_data);
> +}
[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
> index f119489bcdf21..021ffbd6d4e2c 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
[ ... ]
> @@ -330,6 +331,40 @@ struct xe_ras_page_offline_response {
> u32 reserved;
> } __packed;
>
> +/**
> + * struct xe_ras_offline_common - Common structure for offline list and queue
> + */
> +struct xe_ras_offline_common {
> + /** @total_pages: Total number of queued pages */
> + u32 total_pages;
> + /** @pages_returned: Number of pages returned in this response */
> + u32 pages_returned;
> + /** @page_addresses: Array of page addresses (4KB aligned) */
> + u64 page_addresses[XE_RAS_NUM_PAGES];
> + /** @additional_data: Indicates if more data is available */
> + u8 additional_data;
> + /** @reserved: Reserved for future use */
> + u8 reserved[3];
> +} __packed;
> +
> +/**
> + * struct xe_ras_offline_list_request - Request for get offline list command
> + */
> +struct xe_ras_offline_list_request {
> + /** @index: Zero-based index into the offline page list */
> + u32 index;
> +} __packed;
[Severity: Medium]
Should the index field (and fields in the common structure above) be defined
using __le32 or __le64 types instead of native types to ensure correct
endianness conversion when communicating with the sysctrl firmware?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928061830.417395-8-riana.tauro@intel.com?part=2
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store
2026-09-28 6:18 ` [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store Riana Tauro
@ 2026-09-28 6:43 ` sashiko-bot
2026-10-01 11:36 ` Upadhyay, Tejas
1 sibling, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-28 6:43 UTC (permalink / raw)
To: Riana Tauro; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] xe: max_entries from firmware response lost during paginated fetch
--
commit 7e32ecac1950627e3285eb91a4b340622535f4c7
Author: Riana Tauro <riana.tauro@intel.com>
drm/xe/xe_ras: Add function to get maximum pages firmware can store
Add function to get maximum number of pages that firmware can store for
offline tracking. This will be used to report max pages to userspace.
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index f7232429080a8..749cf25e9148a 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -401,9 +401,13 @@ static void get_offlined_list(struct xe_device *xe)
> {
> struct xe_ras_offline_list_response response = {0};
> struct xe_ras_offline_list_request request = {0};
> + struct xe_ras_state *state = &xe->ras.state;
>
> get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request, sizeof(request),
> &response, sizeof(response), &response.common, false);
> +
> + if (response.max_entries)
> + state->max_pages = response.max_entries;
[Severity: Medium]
Is it possible for the max_entries value to be lost here if the offline
list requires multiple chunks to fetch?
Looking at get_offline_pages() in the same file, the response buffer is
cleared at the beginning of every iteration:
static void get_offline_pages(...)
{
...
do {
memset(resp, 0, resp_size);
...
ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
...
} while (common->additional_data);
}
If the firmware only populates max_entries in the first chunk of the
paginated fetch, wouldn't the memset() overwrite it with zeros in
subsequent iterations?
By the time get_offlined_list() checks response.max_entries, it only
contains data from the final chunk. Could this result in max_pages
remaining 0?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928061830.417395-8-riana.tauro@intel.com?part=4
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state
2026-09-28 6:18 ` [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state Riana Tauro
@ 2026-09-28 6:51 ` sashiko-bot
2026-09-28 8:59 ` Ghimiray, Himal Prasad
1 sibling, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-28 6:51 UTC (permalink / raw)
To: Riana Tauro; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] drm/xe/ras: accidental deletion of platform check in xe_ras_init
--
commit 680ae4844f08c353c254033fac23eb66e90d2147
Author: Riana Tauro <riana.tauro@intel.com>
drm/xe: Separate drm-ras netlink data from device and firmware RAS state
This commit separates the DRM RAS data used for netlink reporting from the
device and firmware RAS state in xe_device. It introduces a new xe_ras_state
structure to encapsulate configuration such as the disable_page_offline flag.
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index f1da7f983637a..f7232429080a8 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
[ ... ]
> @@ -1104,19 +1115,18 @@ void xe_ras_init(struct xe_device *xe)
> {
> int ret;
>
> - /*
> - * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
> - * once the feature flag is plumbed through device info.
> - */
> - if (xe->info.platform == XE_CRESCENTISLAND)
> - xe->ras.disable_vram_page_offline =
> - xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
> -
> xe_drm_ras_init(xe);
>
> if (!xe->info.has_sysctrl)
> return;
>
> + /*
> + * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
> + * once the feature flag is plumbed through device info.
> + */
> + xe->ras.state.disable_page_offline =
> + xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
[Severity: Low]
Was the removal of the platform check intentional here?
The TODO comment mentions replacing the platform check, but it looks like the
if (xe->info.platform == XE_CRESCENTISLAND) condition was simply deleted
rather than replaced.
This changes the evaluation scope of the configfs policy in xe_ras_init(). It
will now be evaluated for all platforms that have has_sysctrl set to true,
rather than being restricted to XE_CRESCENTISLAND.
Additionally, because VFs always have has_sysctrl set to 0:
drivers/gpu/drm/xe/xe_pci.c:xe_info_init() {
...
if (IS_SRIOV_VF(xe)) {
xe->info.has_sysctrl = 0;
...
}
The policy will no longer be initialized on VFs due to the early return.
While this shouldn't have a functional impact on VFs since they don't handle
physical VRAM faults, was it intended to evaluate this policy on all other
sysctrl-capable platforms without plumbing the feature flag first?
> +
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928061830.417395-8-riana.tauro@intel.com?part=3
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
2026-09-28 6:18 ` [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates Riana Tauro
@ 2026-09-28 7:07 ` sashiko-bot
2026-09-28 9:00 ` Ghimiray, Himal Prasad
1 sibling, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-28 7:07 UTC (permalink / raw)
To: Riana Tauro; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] xe: failed page offline requests are erroneously cached as successful
- [Medium] xe: page offline deduplication fails during initialization due to wrong order
--
commit 9f25f9a9144c678c3a1d0b655504ad65d4f2849b
Author: Riana Tauro <riana.tauro@intel.com>
drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
A memory scrubber can report multiple errors at the same address. Track
pages already successfully offlined by the firmware so that subsequent
reports for the same address are removed from the firmware queue instead
of issuing redundant page offline requests.
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 749cf25e9148a..d2bcbecd1ba66 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -249,11 +253,16 @@ static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
> xe_log_err(xe, SYSCTRL, ret, "page offline command failed with status %u\n",
> response.status);
>
> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
> + xa_store(&state->offlined_pages, page_address >> XE_PTE_SHIFT,
> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
> +
> return ret;
> }
[Severity: High]
If ras_status_to_errno(response.status) returns an error, the code logs the
failure but execution continues. Does this cause the failed page offline
request to be erroneously cached as successful in state->offlined_pages?
If so, future reports of the same error will trigger a REMOVE action instead
of attempting to OFFLINE again, permanently dropping hardware memory errors
that failed to be offlined.
[ ... ]
> @@ -1145,10 +1177,14 @@ void xe_ras_init(struct xe_device *xe)
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
> + xa_init(&state->offlined_pages);
> +
> get_queued_pages(xe);
> get_offlined_list(xe);
[Severity: Medium]
Will page offline deduplication fail for queued pages because of the
initialization order?
During get_queued_pages(), pending errors are processed via
handle_page_offline(), which checks the xarray to detect duplicates:
handle_page_offline()
if (xa_load(&state->offlined_pages, ...))
action = XE_RAS_PAGE_ACTION_REMOVE;
Since get_offlined_list() has not yet populated the xarray with the existing
offlined pages, won't the xarray be empty at this point?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928061830.417395-8-riana.tauro@intel.com?part=6
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
2026-09-28 6:34 ` sashiko-bot
@ 2026-09-28 8:55 ` Ghimiray, Himal Prasad
2026-10-01 10:50 ` Upadhyay, Tejas
2 siblings, 0 replies; 28+ messages in thread
From: Ghimiray, Himal Prasad @ 2026-09-28 8:55 UTC (permalink / raw)
To: Riana Tauro, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 11:48, Riana Tauro wrote:
> Add basic support for sending page offline/remove requests to system
> controller and use it for device memory ECC error handling.
> Pages that belong to critical BOs cannot be handled by offlining and
> require a SBR (Secondary Bus Reset).
> Pages that are configured for log-only handling are not marked as bad by
> firmware.
>
> For all other valid page addresses, the first occurrence of error
> indicates a poison error and the page is offlined only by software.
> Firmware avoids permanently marking the page as bad. The second occurrence
> of an error indicates a Double-bit ECC error and the firmware
> permanently marks the page as bad.
>
> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: use ret in sigid logging (Mallesh)
> remove additional log
> use xe_assert (Michal)
>
> v3: align address to PAGE_SIZE (sashiko, Himal)
> rename decline to remove
> add more descriptive logs (Himal)
> ---
> drivers/gpu/drm/xe/xe_ras.c | 127 +++++++++++++++++-
> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++
> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 2 +
> 3 files changed, 159 insertions(+), 5 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 7a85735c57d5..1225c561a872 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -3,6 +3,8 @@
> * Copyright © 2026 Intel Corporation
> */
>
> +#include "xe_assert.h"
> +#include "xe_bo.h"
> #include "xe_configfs.h"
> #include "xe_debugfs.h"
> #include "xe_device.h"
> @@ -16,6 +18,7 @@
> #include "xe_sysctrl_event_types.h"
> #include "xe_sysctrl_mailbox.h"
> #include "xe_sysctrl_mailbox_types.h"
> +#include "xe_ttm_vram_mgr.h"
>
> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
> /*
> @@ -201,6 +204,119 @@ static inline const char *comp_to_str(u8 component)
> return xe_ras_components[component];
> }
>
> +static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
> + enum xe_ras_page_action action)
> +{
> + struct xe_sysctrl_mailbox_command command = {0};
> + struct xe_ras_page_offline_request request = {0};
> + struct xe_ras_page_offline_response response = {0};
> + size_t rlen;
> + int ret;
> +
> + if (!xe->info.has_sysctrl)
> + return 0;
> +
> + xe_assert(xe, action < XE_RAS_PAGE_ACTION_MAX);
> +
> + request.page_address = page_address;
> + request.action = action;
> +
> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to offline page 0x%llx\n",
> + page_address);
> + else
> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to remove page 0x%llx from queue\n",
> + page_address);
> +
> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP, XE_SYSCTRL_CMD_PAGE_OFFLINE,
> + &request, sizeof(request), &response, sizeof(response));
> +
> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
> + if (ret) {
> + xe_log_err(xe, SYSCTRL, ret, "failed to send page offline command\n");
> + return ret;
> + }
> +
> + if (rlen != sizeof(response)) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "unexpected page offline response length %zu (expected %zu)\n",
> + rlen, sizeof(response));
> + return -EINVAL;
> + }
> +
> + ret = ras_status_to_errno(response.status);
> + if (ret)
> + xe_log_err(xe, SYSCTRL, ret, "page offline command failed with status %u\n",
> + response.status);
> +
> + return ret;
> +}
> +
> +static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send_cmd)
> +{
> + enum xe_ras_page_action action;
> + u64 addr;
> + int ret = 0;
> +
> + if (!IS_ALIGNED(page_address, XE_PAGE_SIZE)) {
> + xe_log_err(xe, SYSCTRL, -EINVAL, "Unaligned physical page address: 0x%llx\n",
> + page_address);
> + return -EINVAL;
> + }
> +
> + addr = ALIGN_DOWN(page_address, PAGE_SIZE);
> +
> + ret = xe_ttm_vram_handle_addr_fault(xe, addr);
> +
> + /*
> + * Handle return code from address fault handling function:
> + * 0: Page softofflined, remove from firmware queue
> + * -EIO: Address belongs to a critical BO/stolen area that cannot be offlined
> + * -EOPNOTSUPP: Address is valid and can be offlined but user policy is not to offline
> + * -EEXIST: Address is soft offlined but yet to be offlined by firmware for second
> + * occurrence
> + */
> +
> + switch (ret) {
> + case 0:
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page soft-offlined\n",
> + page_address);
> + break;
> + /* User policy set to decline page offlining */
> + case -EOPNOTSUPP:
xe_ttm_vram_handle_addr_fault() can also return -EOPNOTSUPP, since
xe_ttm_vram_addr_to_region() may return that error as well. Either add a
policy check alongside the existing check -EOPNOTSUPP, or change
xe_ttm_vram_addr_to_region() to return 0 instead.
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, user policy set to decline soft-offlining\n",
> + page_address);
> + break;
> + case -EIO:
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page belongs to critical BO and cannot be soft-offlined\n",
> + page_address);
> + return ret;
> + case -EEXIST:
> + action = XE_RAS_PAGE_ACTION_OFFLINE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Double-bit ECC error detected at physical address 0x%llx, page soft-offlined\n",
> + page_address);
> + break;
> + default:
> + xe_log_err(xe, DEVICE_MEMORY, ret, "Failed to handle address fault at physical address 0x%llx\n",
> + page_address);
> + return 0;
> + }
> +
> + if (send_cmd) {
> + ret = send_page_offline_cmd(xe, page_address, action);
> + if (ret)
> + return ret;
> + }
> +
> + return 0;
> +}
> +
> static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class *counter)
> {
> u8 severity = counter->common.severity;
> @@ -368,11 +484,12 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a
> static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_array *arr)
> {
> struct xe_ras_memory_error *info = (void *)arr->details;
> + int ret;
>
> /*
> * For memory errors, the recovery action depends on the error category
> *
> - * TODO: Double-bit ECC errors: Page offlining
> + * Double-bit ECC errors: Page offlining
> * Poison and data parity errors: Log only
> * For any other memory errors, request a reset as recovery mechanism
> */
> @@ -384,10 +501,10 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_
> xe_info(xe, "[RAS]: Data parity error detected\n");
> break;
> case XE_RAS_MEMORY_DB_ECC:
> - xe_info(xe, "[RAS]: Double-bit ECC error detected at sw address 0x%llx\n",
> - info->sw_address);
> - /* TODO: Add page offlining for Double-bit ECC error */
> - fallthrough;
> + ret = handle_page_offline(xe, info->sw_address, true);
> + if (ret)
> + return XE_RAS_RECOVERY_ACTION_RESET;
> + break;
> default:
> return XE_RAS_RECOVERY_ACTION_RESET;
> }
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
> index fe6f3658a2a4..f119489bcdf2 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -17,6 +17,19 @@
> #define XE_RAS_MEMORY_POISON BIT(2)
> #define XE_RAS_MEMORY_DATA_PARITY BIT(5)
>
> +/**
> + * enum xe_ras_page_action - Page offline actions for page offline request
> + *
> + * @XE_RAS_PAGE_ACTION_OFFLINE: Instruct firmware to offline the page
> + * @XE_RAS_PAGE_ACTION_REMOVE: Instruct firmware to remove the page from queue
> + * @XE_RAS_PAGE_ACTION_MAX: Max value
> + */
> +enum xe_ras_page_action {
> + XE_RAS_PAGE_ACTION_OFFLINE,
> + XE_RAS_PAGE_ACTION_REMOVE,
> + XE_RAS_PAGE_ACTION_MAX
> +};
> +
> /**
> * enum xe_ras_recovery_action - RAS recovery actions
> *
> @@ -295,6 +308,28 @@ struct xe_ras_memory_error {
> u32 reserved2[10];
> } __packed;
>
> +/**
> + * struct xe_ras_page_offline_request - Request for page offline command
> + */
> +struct xe_ras_page_offline_request {
> + /** @page_address: Page address (4KB aligned) */
> + u64 page_address;
> + /** @action: Action to be performed, see &enum xe_ras_page_action */
> + u32 action;
> + /** @reserved: Reserved for future use */
> + u32 reserved;
> +} __packed;
> +
> +/**
> + * struct xe_ras_page_offline_response - Response from page offline command
> + */
> +struct xe_ras_page_offline_response {
> + /** @status: Status of the page offline request */
> + u32 status;
> + /** @reserved: Reserved for future use */
> + u32 reserved;
> +} __packed;
> +
> /**
> * struct xe_ras_get_health_request - Request structure for obtaining gpu health
> */
> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> index c236e5377f30..a01576bf2e73 100644
> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> @@ -30,6 +30,7 @@ enum xe_sysctrl_group {
> * @XE_SYSCTRL_CMD_GET_THRESHOLD: Retrieve error threshold
> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
> + * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a page
> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
> */
> @@ -40,6 +41,7 @@ enum xe_sysctrl_gfsp_cmd {
> XE_SYSCTRL_CMD_GET_THRESHOLD = 0x05,
> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
> + XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
> };
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state
2026-09-28 6:18 ` [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state Riana Tauro
2026-09-28 6:51 ` sashiko-bot
@ 2026-09-28 8:59 ` Ghimiray, Himal Prasad
1 sibling, 0 replies; 28+ messages in thread
From: Ghimiray, Himal Prasad @ 2026-09-28 8:59 UTC (permalink / raw)
To: Riana Tauro, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 11:48, Riana Tauro wrote:
> Keep the DRM RAS data used for netlink reporting separate from the
> device and firmware RAS state in xe_device.
>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: remove inline function (Himal)
> ---
> drivers/gpu/drm/xe/xe_device_types.h | 11 +++++++++--
> drivers/gpu/drm/xe/xe_drm_ras.c | 16 ++++++++--------
> drivers/gpu/drm/xe/xe_drm_ras_types.h | 3 ---
> drivers/gpu/drm/xe/xe_hw_error.c | 6 +++---
> drivers/gpu/drm/xe/xe_ras.c | 26 ++++++++++++++++++--------
> drivers/gpu/drm/xe/xe_ras.h | 2 ++
> drivers/gpu/drm/xe/xe_ras_types.h | 10 ++++++++++
> drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 3 ++-
> 8 files changed, 52 insertions(+), 25 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h
> index bb4234f00453..c0b53af9bdd6 100644
> --- a/drivers/gpu/drm/xe/xe_device_types.h
> +++ b/drivers/gpu/drm/xe/xe_device_types.h
> @@ -21,6 +21,7 @@
> #include "xe_platform_types.h"
> #include "xe_pmu_types.h"
> #include "xe_pt_types.h"
> +#include "xe_ras_types.h"
> #include "xe_sriov_pf_types.h"
> #include "xe_sriov_types.h"
> #include "xe_sriov_vf_types.h"
> @@ -570,8 +571,14 @@ struct xe_device {
> /** @pmu: performance monitoring unit */
> struct xe_pmu pmu;
>
> - /** @ras: RAS structure for device */
> - struct xe_drm_ras ras;
> + /** @ras: RAS (Reliability, Availability, Serviceability) structures */
> + struct {
> + /** @ras.nl_data: drm-ras netlink data */
> + struct xe_drm_ras nl_data;
> +
> + /** @ras.state: RAS device and firmware state */
> + struct xe_ras_state state;
> + } ras;
>
> /** @i2c: I2C host controller */
> struct xe_i2c *i2c;
> diff --git a/drivers/gpu/drm/xe/xe_drm_ras.c b/drivers/gpu/drm/xe/xe_drm_ras.c
> index 7f3695707611..38d77561facf 100644
> --- a/drivers/gpu/drm/xe/xe_drm_ras.c
> +++ b/drivers/gpu/drm/xe/xe_drm_ras.c
> @@ -20,7 +20,7 @@ static int query_error_counter(struct xe_device *xe,
> enum drm_xe_ras_error_severity severity,
> u32 error_id, const char **name, u32 *val)
> {
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
>
> if (!info || !info[error_id].name)
> @@ -41,7 +41,7 @@ static int clear_error_counter(struct xe_device *xe,
> enum drm_xe_ras_error_severity severity,
> u32 error_id)
> {
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
>
> if (!info || !info[error_id].name)
> @@ -90,7 +90,7 @@ static int query_correctable_error_threshold(struct drm_ras_node *ep, u32 error_
> const char **name, u32 *threshold)
> {
> struct xe_device *xe = ep->priv;
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[DRM_XE_RAS_ERR_SEV_CORRECTABLE];
>
> if (!info || !info[error_id].name)
> @@ -106,7 +106,7 @@ static int query_correctable_error_threshold(struct drm_ras_node *ep, u32 error_
> static int set_correctable_error_threshold(struct drm_ras_node *ep, u32 error_id, u32 threshold)
> {
> struct xe_device *xe = ep->priv;
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[DRM_XE_RAS_ERR_SEV_CORRECTABLE];
>
> if (!info || !info[error_id].name)
> @@ -142,7 +142,7 @@ static int assign_node_params(struct xe_device *xe, struct drm_ras_node *node,
> const enum drm_xe_ras_error_severity severity)
> {
> struct pci_dev *pdev = to_pci_dev(xe->drm.dev);
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> const char *device_name;
>
> device_name = kasprintf(GFP_KERNEL, "%04x:%02x:%02x.%d",
> @@ -190,7 +190,7 @@ static void cleanup_node(struct drm_device *drm, void *node)
>
> static int register_nodes(struct xe_device *xe)
> {
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct drm_ras_node *node;
> int i, ret;
>
> @@ -230,7 +230,7 @@ static int register_nodes(struct xe_device *xe)
> */
> void xe_drm_ras_event(struct xe_device *xe, u32 component, u32 severity, u32 value)
> {
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
> struct drm_ras_node *node;
> int ret;
> @@ -260,7 +260,7 @@ void xe_drm_ras_event(struct xe_device *xe, u32 component, u32 severity, u32 val
> */
> int xe_drm_ras_init(struct xe_device *xe)
> {
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct drm_ras_node *node;
> int err;
>
> diff --git a/drivers/gpu/drm/xe/xe_drm_ras_types.h b/drivers/gpu/drm/xe/xe_drm_ras_types.h
> index 0be218ba2db7..8d729ad6a264 100644
> --- a/drivers/gpu/drm/xe/xe_drm_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_drm_ras_types.h
> @@ -43,9 +43,6 @@ struct xe_drm_ras {
>
> /** @info: info array for all types of errors */
> struct xe_drm_ras_counter *info[DRM_XE_RAS_ERR_SEV_MAX];
> -
> - /** @disable_vram_page_offline: cached configfs policy, immutable after init */
> - bool disable_vram_page_offline;
> };
>
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_hw_error.c b/drivers/gpu/drm/xe/xe_hw_error.c
> index 5f2abc9485ff..f53a6b3055de 100644
> --- a/drivers/gpu/drm/xe/xe_hw_error.c
> +++ b/drivers/gpu/drm/xe/xe_hw_error.c
> @@ -240,7 +240,7 @@ static void log_soc_error(struct xe_tile *tile, const char * const *reg_info,
> {
> const char *severity_str = error_severity[severity];
> struct xe_device *xe = tile_to_xe(tile);
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
> const char *name;
>
> @@ -260,7 +260,7 @@ static void gt_hw_error_handler(struct xe_tile *tile, const enum hardware_error
> {
> const enum drm_xe_ras_error_severity severity = hw_err_to_severity(hw_err);
> struct xe_device *xe = tile_to_xe(tile);
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
> struct xe_mmio *mmio = &tile->mmio;
> unsigned long err_stat = 0;
> @@ -422,7 +422,7 @@ static void hw_error_source_handler(struct xe_tile *tile, const enum hardware_er
> const enum drm_xe_ras_error_severity severity = hw_err_to_severity(hw_err);
> const char *severity_str = error_severity[severity];
> struct xe_device *xe = tile_to_xe(tile);
> - struct xe_drm_ras *ras = &xe->ras;
> + struct xe_drm_ras *ras = &xe->ras.nl_data;
> struct xe_drm_ras_counter *info = ras->info[severity];
> unsigned long flags, err_src;
> u32 err_bit;
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 752754f09314..9a1a8190578b 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -1094,6 +1094,17 @@ static const struct attribute_group gpu_health_group = {
> .attrs = gpu_health_attrs,
> };
>
> +/**
> + * xe_ras_get_disable_page_offline - Get the page offline user policy
> + * @xe: xe device instance
> + *
> + * Return: true if the page offline policy is disabled, false otherwise.
> + */
> +bool xe_ras_get_disable_page_offline(struct xe_device *xe)
> +{
> + return xe->ras.state.disable_page_offline;
> +}
> +
> /**
> * xe_ras_init - Initialize Xe RAS
> * @xe: xe device instance
> @@ -1104,19 +1115,18 @@ void xe_ras_init(struct xe_device *xe)
> {
> int ret;
>
> - /*
> - * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
> - * once the feature flag is plumbed through device info.
> - */
> - if (xe->info.platform == XE_CRESCENTISLAND)
> - xe->ras.disable_vram_page_offline =
> - xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
> -
> xe_drm_ras_init(xe);
>
> if (!xe->info.has_sysctrl)
> return;
>
> + /*
> + * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
> + * once the feature flag is plumbed through device info.
> + */
> + xe->ras.state.disable_page_offline =
> + xe_configfs_get_disable_vram_page_offline(to_pci_dev(xe->drm.dev));
> +
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
> index 0b8669f28d56..ac1253d6ead9 100644
> --- a/drivers/gpu/drm/xe/xe_ras.h
> +++ b/drivers/gpu/drm/xe/xe_ras.h
> @@ -7,6 +7,7 @@
> #define _XE_RAS_H_
>
> #include <linux/types.h>
> +
> #include "xe_ras_types.h"
>
> struct xe_device;
> @@ -20,5 +21,6 @@ int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8 component, u32 *t
> int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8 component, u32 threshold);
> void xe_ras_init(struct xe_device *xe);
> enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe);
> +bool xe_ras_get_disable_page_offline(struct xe_device *xe);
>
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
> index 021ffbd6d4e2..40224db38906 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -406,4 +406,14 @@ struct xe_ras_set_health_response {
> /** @reserved1: Reserved for future use */
> u32 reserved1[2];
> } __packed;
> +
> +/* Device structures */
> +
> +/**
> + * struct xe_ras_state - RAS device and firmware state
> + */
> +struct xe_ras_state {
> + /** @disable_page_offline: cached configfs policy, immutable after init */
> + bool disable_page_offline;
> +};
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> index 9a514d983e90..2c4722a956a0 100644
> --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> @@ -24,6 +24,7 @@
> #include "xe_mmio.h"
> #include "xe_pm.h"
> #include "xe_printk.h"
> +#include "xe_ras.h"
> #include "xe_res_cursor.h"
> #include "xe_ttm_stolen_mgr.h"
> #include "xe_ttm_vram_mgr.h"
> @@ -933,7 +934,7 @@ int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr)
> vram_mgr = &vr->ttm;
> mm = &vram_mgr->mm;
>
> - if (xe->ras.disable_vram_page_offline) {
> + if (xe_ras_get_disable_page_offline(xe)) {
> xe_err(xe, "0x%llx is reported as corrupted address by HW\n",
> addr);
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> return -EOPNOTSUPP;
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
2026-09-28 6:18 ` [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates Riana Tauro
2026-09-28 7:07 ` sashiko-bot
@ 2026-09-28 9:00 ` Ghimiray, Himal Prasad
2026-09-28 9:17 ` Ghimiray, Himal Prasad
1 sibling, 1 reply; 28+ messages in thread
From: Ghimiray, Himal Prasad @ 2026-09-28 9:00 UTC (permalink / raw)
To: Riana Tauro, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 11:48, Riana Tauro wrote:
> A memory scrubber can report multiple errors at the same address. Track
> pages already successfully offlined by the firmware so that subsequent
> reports for the same address are removed from the firmware queue instead
> of issuing redundant page offline requests.
>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: use xe page shift (Sashiko)
> ---
> drivers/gpu/drm/xe/xe_device.c | 4 ++-
> drivers/gpu/drm/xe/xe_ras.c | 44 ++++++++++++++++++++++++++++---
> drivers/gpu/drm/xe/xe_ras.h | 2 +-
> drivers/gpu/drm/xe/xe_ras_types.h | 3 +++
> 4 files changed, 47 insertions(+), 6 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index 205cb4e7f9e8..1ec7c74cbd58 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
> @@ -957,7 +957,9 @@ int xe_device_probe(struct xe_device *xe)
> if (err)
> return err;
>
> - xe_ras_init(xe);
> + err = xe_ras_init(xe);
> + if (err)
> + return err;
>
> /*
> * Now that GT is initialized (TTM in particular),
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index cfc7f07e9ced..7dee1c804eeb 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -20,6 +20,9 @@
> #include "xe_sysctrl_mailbox_types.h"
> #include "xe_ttm_vram_mgr.h"
>
> +/* Any non-null entry marks the page as offlined by firmware */
> +#define XE_RAS_PAGE_OFFLINED xa_mk_value(1)
> +
> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
> /*
> * Uncorrectable error type for core compute errors.
> @@ -210,6 +213,7 @@ static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
> struct xe_sysctrl_mailbox_command command = {0};
> struct xe_ras_page_offline_request request = {0};
> struct xe_ras_page_offline_response response = {0};
> + struct xe_ras_state *state = &xe->ras.state;
> size_t rlen;
> int ret;
>
> @@ -249,11 +253,16 @@ static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
> xe_log_err(xe, SYSCTRL, ret, "page offline command failed with status %u\n",
> response.status);
>
> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
> + xa_store(&state->offlined_pages, page_address >> XE_PTE_SHIFT,
> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
> +
> return ret;
> }
>
> static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send_cmd)
> {
> + struct xe_ras_state *state = &xe->ras.state;
> enum xe_ras_page_action action;
> u64 addr;
> int ret = 0;
> @@ -297,7 +306,11 @@ static int handle_page_offline(struct xe_device *xe, u64 page_address, bool send
> page_address);
> return ret;
> case -EEXIST:
> - action = XE_RAS_PAGE_ACTION_OFFLINE;
> + if (xa_load(&state->offlined_pages, page_address >> XE_PTE_SHIFT))
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + else
> + action = XE_RAS_PAGE_ACTION_OFFLINE;
> +
> xe_log_err(xe, DEVICE_MEMORY, ret,
> "Double-bit ECC error detected at physical address 0x%llx, page soft-offlined\n",
> page_address);
> @@ -341,6 +354,8 @@ static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t r
> {
> struct xe_sysctrl_mailbox_command command = {0};
> struct xe_ras_offline_list_request *list_req;
> + struct xe_ras_state *state = &xe->ras.state;
> + unsigned long index;
> u32 total_pages = 0, count = 0;
> ssize_t rlen;
> int ret, i;
> @@ -370,8 +385,15 @@ static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t r
> return;
> }
>
> - for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
> + for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++) {
> handle_page_offline(xe, common->page_addresses[i], offline);
> + /* The pages are already offlined by firmware */
> + if (!offline) {
> + index = common->page_addresses[i] >> XE_PTE_SHIFT;
> + xa_store(&state->offlined_pages, index,
> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
> + }
> + }
>
> count += common->pages_returned;
> if (!common->pages_returned)
> @@ -1098,6 +1120,13 @@ static const struct attribute_group gpu_health_group = {
> .attrs = gpu_health_attrs,
> };
>
> +static void ras_fini(void *arg)
> +{
> + struct xe_device *xe = arg;
> +
> + xa_destroy(&xe->ras.state.offlined_pages);
> +}
> +
> /**
> * xe_ras_get_disable_page_offline - Get the page offline user policy
> * @xe: xe device instance
> @@ -1125,15 +1154,18 @@ u32 xe_ras_get_max_pages(struct xe_device *xe)
> * @xe: xe device instance
> *
> * Initialize Xe RAS
> + *
> + * Return: 0 on success, negative error code otherwise
> */
> -void xe_ras_init(struct xe_device *xe)
> +int xe_ras_init(struct xe_device *xe)
> {
> + struct xe_ras_state *state = &xe->ras.state;
> int ret;
>
> xe_drm_ras_init(xe);
>
> if (!xe->info.has_sysctrl)
> - return;
> + return 0;
>
> /*
> * TODO: Replace platform check with xe->info.has_disable_vram_page_offline
> @@ -1145,10 +1177,14 @@ void xe_ras_init(struct xe_device *xe)
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
> + xa_init(&state->offlined_pages);
> +
> get_queued_pages(xe);
> get_offlined_list(xe);
>
> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
> if (ret)
> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
> +
> + return devm_add_action_or_reset(xe->drm.dev, ras_fini, xe);
> }
> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
> index 561c656e2ad3..e41a52869085 100644
> --- a/drivers/gpu/drm/xe/xe_ras.h
> +++ b/drivers/gpu/drm/xe/xe_ras.h
> @@ -19,7 +19,7 @@ int xe_ras_get_counter(struct xe_device *xe, u8 severity, u8 component, u32 *val
> int xe_ras_clear_counter(struct xe_device *xe, u8 severity, u8 component);
> int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8 component, u32 *threshold);
> int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8 component, u32 threshold);
> -void xe_ras_init(struct xe_device *xe);
> +int xe_ras_init(struct xe_device *xe);
> enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe);
> bool xe_ras_get_disable_page_offline(struct xe_device *xe);
> u32 xe_ras_get_max_pages(struct xe_device *xe);
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
> index e09b50a6f77b..50d70e2e2fd3 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -7,6 +7,7 @@
> #define _XE_RAS_TYPES_H_
>
> #include <linux/types.h>
> +#include <linux/xarray.h>
>
> #define XE_RAS_NUM_COUNTERS 16
> #define XE_RAS_NUM_ERROR_ARR 3
> @@ -417,5 +418,7 @@ struct xe_ras_state {
> bool disable_page_offline;
> /** @max_pages: Total number of pages that can be stored by firmware */
> u32 max_pages;
> + /** @offlined_pages: XArray of pages offlined by firmware */
> + struct xarray offlined_pages;
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> };
> #endif
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
2026-09-28 9:00 ` Ghimiray, Himal Prasad
@ 2026-09-28 9:17 ` Ghimiray, Himal Prasad
2026-09-28 9:22 ` Tauro, Riana
0 siblings, 1 reply; 28+ messages in thread
From: Ghimiray, Himal Prasad @ 2026-09-28 9:17 UTC (permalink / raw)
To: Riana Tauro, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 14:30, Ghimiray, Himal Prasad wrote:
>
>
> On 28-09-2026 11:48, Riana Tauro wrote:
>> A memory scrubber can report multiple errors at the same address. Track
>> pages already successfully offlined by the firmware so that subsequent
>> reports for the same address are removed from the firmware queue instead
>> of issuing redundant page offline requests.
>>
>> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
>> ---
>> v2: use xe page shift (Sashiko)
>> ---
>> drivers/gpu/drm/xe/xe_device.c | 4 ++-
>> drivers/gpu/drm/xe/xe_ras.c | 44 ++++++++++++++++++++++++++++---
>> drivers/gpu/drm/xe/xe_ras.h | 2 +-
>> drivers/gpu/drm/xe/xe_ras_types.h | 3 +++
>> 4 files changed, 47 insertions(+), 6 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/
>> xe_device.c
>> index 205cb4e7f9e8..1ec7c74cbd58 100644
>> --- a/drivers/gpu/drm/xe/xe_device.c
>> +++ b/drivers/gpu/drm/xe/xe_device.c
>> @@ -957,7 +957,9 @@ int xe_device_probe(struct xe_device *xe)
>> if (err)
>> return err;
>> - xe_ras_init(xe);
>> + err = xe_ras_init(xe);
>> + if (err)
>> + return err;
>> /*
>> * Now that GT is initialized (TTM in particular),
>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
>> index cfc7f07e9ced..7dee1c804eeb 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.c
>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>> @@ -20,6 +20,9 @@
>> #include "xe_sysctrl_mailbox_types.h"
>> #include "xe_ttm_vram_mgr.h"
>> +/* Any non-null entry marks the page as offlined by firmware */
>> +#define XE_RAS_PAGE_OFFLINED xa_mk_value(1)
>> +
>> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
>> /*
>> * Uncorrectable error type for core compute errors.
>> @@ -210,6 +213,7 @@ static int send_page_offline_cmd(struct xe_device
>> *xe, u64 page_address,
>> struct xe_sysctrl_mailbox_command command = {0};
>> struct xe_ras_page_offline_request request = {0};
>> struct xe_ras_page_offline_response response = {0};
>> + struct xe_ras_state *state = &xe->ras.state;
>> size_t rlen;
>> int ret;
>> @@ -249,11 +253,16 @@ static int send_page_offline_cmd(struct
>> xe_device *xe, u64 page_address,
>> xe_log_err(xe, SYSCTRL, ret, "page offline command failed
>> with status %u\n",
>> response.status);
if (!ret && action == XE_RAS_PAGE_ACTION_OFFLINE)
xa_store(...);
>> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
>> + xa_store(&state->offlined_pages, page_address >> XE_PTE_SHIFT,
>> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
>> +
>> return ret;
>> }
>> static int handle_page_offline(struct xe_device *xe, u64
>> page_address, bool send_cmd)
>> {
>> + struct xe_ras_state *state = &xe->ras.state;
>> enum xe_ras_page_action action;
>> u64 addr;
>> int ret = 0;
>> @@ -297,7 +306,11 @@ static int handle_page_offline(struct xe_device
>> *xe, u64 page_address, bool send
>> page_address);
>> return ret;
>> case -EEXIST:
>> - action = XE_RAS_PAGE_ACTION_OFFLINE;
>> + if (xa_load(&state->offlined_pages, page_address >>
>> XE_PTE_SHIFT))
>> + action = XE_RAS_PAGE_ACTION_REMOVE;
>> + else
>> + action = XE_RAS_PAGE_ACTION_OFFLINE;
>> +
>> xe_log_err(xe, DEVICE_MEMORY, ret,
>> "Double-bit ECC error detected at physical address
>> 0x%llx, page soft-offlined\n",
>> page_address);
>> @@ -341,6 +354,8 @@ static void get_offline_pages(struct xe_device
>> *xe, u32 cmd, void *req, size_t r
>> {
>> struct xe_sysctrl_mailbox_command command = {0};
>> struct xe_ras_offline_list_request *list_req;
>> + struct xe_ras_state *state = &xe->ras.state;
>> + unsigned long index;
>> u32 total_pages = 0, count = 0;
>> ssize_t rlen;
>> int ret, i;
>> @@ -370,8 +385,15 @@ static void get_offline_pages(struct xe_device
>> *xe, u32 cmd, void *req, size_t r
>> return;
>> }
>> - for (i = 0; i < common->pages_returned && i <
>> XE_RAS_NUM_PAGES; i++)
>> + for (i = 0; i < common->pages_returned && i <
>> XE_RAS_NUM_PAGES; i++) {
>> handle_page_offline(xe, common->page_addresses[i],
>> offline);
>> + /* The pages are already offlined by firmware */
>> + if (!offline) {
>> + index = common->page_addresses[i] >> XE_PTE_SHIFT;
>> + xa_store(&state->offlined_pages, index,
>> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
>> + }
>> + }
Do we need change order for
get_offlined_list and get_queued_pages ?
Missed them earlier.>> count += common->pages_returned;
>> if (!common->pages_returned)
>> @@ -1098,6 +1120,13 @@ static const struct attribute_group
>> gpu_health_group = {
>> .attrs = gpu_health_attrs,
>> };
>> +static void ras_fini(void *arg)
>> +{
>> + struct xe_device *xe = arg;
>> +
>> + xa_destroy(&xe->ras.state.offlined_pages);
>> +}
>> +
>> /**
>> * xe_ras_get_disable_page_offline - Get the page offline user policy
>> * @xe: xe device instance
>> @@ -1125,15 +1154,18 @@ u32 xe_ras_get_max_pages(struct xe_device *xe)
>> * @xe: xe device instance
>> *
>> * Initialize Xe RAS
>> + *
>> + * Return: 0 on success, negative error code otherwise
>> */
>> -void xe_ras_init(struct xe_device *xe)
>> +int xe_ras_init(struct xe_device *xe)
>> {
>> + struct xe_ras_state *state = &xe->ras.state;
>> int ret;
>> xe_drm_ras_init(xe);
>> if (!xe->info.has_sysctrl)
>> - return;
>> + return 0;
>> /*
>> * TODO: Replace platform check with xe-
>> >info.has_disable_vram_page_offline
>> @@ -1145,10 +1177,14 @@ void xe_ras_init(struct xe_device *xe)
>> if (IS_ENABLED(CONFIG_PCIEAER))
>> ras_usp_aer_init(xe);
>> + xa_init(&state->offlined_pages);
>> +
>> get_queued_pages(xe);
>> get_offlined_list(xe);
>> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
>> if (ret)
>> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
>> +
>> + return devm_add_action_or_reset(xe->drm.dev, ras_fini, xe);
>> }
>> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
>> index 561c656e2ad3..e41a52869085 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.h
>> +++ b/drivers/gpu/drm/xe/xe_ras.h
>> @@ -19,7 +19,7 @@ int xe_ras_get_counter(struct xe_device *xe, u8
>> severity, u8 component, u32 *val
>> int xe_ras_clear_counter(struct xe_device *xe, u8 severity, u8
>> component);
>> int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8
>> component, u32 *threshold);
>> int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8
>> component, u32 threshold);
>> -void xe_ras_init(struct xe_device *xe);
>> +int xe_ras_init(struct xe_device *xe);
>> enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device
>> *xe);
>> bool xe_ras_get_disable_page_offline(struct xe_device *xe);
>> u32 xe_ras_get_max_pages(struct xe_device *xe);
>> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/
>> xe_ras_types.h
>> index e09b50a6f77b..50d70e2e2fd3 100644
>> --- a/drivers/gpu/drm/xe/xe_ras_types.h
>> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
>> @@ -7,6 +7,7 @@
>> #define _XE_RAS_TYPES_H_
>> #include <linux/types.h>
>> +#include <linux/xarray.h>
>> #define XE_RAS_NUM_COUNTERS 16
>> #define XE_RAS_NUM_ERROR_ARR 3
>> @@ -417,5 +418,7 @@ struct xe_ras_state {
>> bool disable_page_offline;
>> /** @max_pages: Total number of pages that can be stored by
>> firmware */
>> u32 max_pages;
>> + /** @offlined_pages: XArray of pages offlined by firmware */
>> + struct xarray offlined_pages;
>
> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Sashiko comments look valid and Need addressing.
>
>> };
>> #endif
>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates
2026-09-28 9:17 ` Ghimiray, Himal Prasad
@ 2026-09-28 9:22 ` Tauro, Riana
0 siblings, 0 replies; 28+ messages in thread
From: Tauro, Riana @ 2026-09-28 9:22 UTC (permalink / raw)
To: Ghimiray, Himal Prasad, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 14:47, Ghimiray, Himal Prasad wrote:
>
>
> On 28-09-2026 14:30, Ghimiray, Himal Prasad wrote:
>>
>>
>> On 28-09-2026 11:48, Riana Tauro wrote:
>>> A memory scrubber can report multiple errors at the same address. Track
>>> pages already successfully offlined by the firmware so that subsequent
>>> reports for the same address are removed from the firmware queue
>>> instead
>>> of issuing redundant page offline requests.
>>>
>>> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
>>> ---
>>> v2: use xe page shift (Sashiko)
>>> ---
>>> drivers/gpu/drm/xe/xe_device.c | 4 ++-
>>> drivers/gpu/drm/xe/xe_ras.c | 44
>>> ++++++++++++++++++++++++++++---
>>> drivers/gpu/drm/xe/xe_ras.h | 2 +-
>>> drivers/gpu/drm/xe/xe_ras_types.h | 3 +++
>>> 4 files changed, 47 insertions(+), 6 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/
>>> xe_device.c
>>> index 205cb4e7f9e8..1ec7c74cbd58 100644
>>> --- a/drivers/gpu/drm/xe/xe_device.c
>>> +++ b/drivers/gpu/drm/xe/xe_device.c
>>> @@ -957,7 +957,9 @@ int xe_device_probe(struct xe_device *xe)
>>> if (err)
>>> return err;
>>> - xe_ras_init(xe);
>>> + err = xe_ras_init(xe);
>>> + if (err)
>>> + return err;
>>> /*
>>> * Now that GT is initialized (TTM in particular),
>>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
>>> index cfc7f07e9ced..7dee1c804eeb 100644
>>> --- a/drivers/gpu/drm/xe/xe_ras.c
>>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>>> @@ -20,6 +20,9 @@
>>> #include "xe_sysctrl_mailbox_types.h"
>>> #include "xe_ttm_vram_mgr.h"
>>> +/* Any non-null entry marks the page as offlined by firmware */
>>> +#define XE_RAS_PAGE_OFFLINED xa_mk_value(1)
>>> +
>>> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
>>> /*
>>> * Uncorrectable error type for core compute errors.
>>> @@ -210,6 +213,7 @@ static int send_page_offline_cmd(struct
>>> xe_device *xe, u64 page_address,
>>> struct xe_sysctrl_mailbox_command command = {0};
>>> struct xe_ras_page_offline_request request = {0};
>>> struct xe_ras_page_offline_response response = {0};
>>> + struct xe_ras_state *state = &xe->ras.state;
>>> size_t rlen;
>>> int ret;
>>> @@ -249,11 +253,16 @@ static int send_page_offline_cmd(struct
>>> xe_device *xe, u64 page_address,
>>> xe_log_err(xe, SYSCTRL, ret, "page offline command failed
>>> with status %u\n",
>>> response.status);
>
> if (!ret && action == XE_RAS_PAGE_ACTION_OFFLINE)
> xa_store(...);
Removed ret and missed while rebasing.
Thanks for catching this
>
>>> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
>>> + xa_store(&state->offlined_pages, page_address >> XE_PTE_SHIFT,
>>> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
>>> +
>>> return ret;
>>> }
>>> static int handle_page_offline(struct xe_device *xe, u64
>>> page_address, bool send_cmd)
>>> {
>>> + struct xe_ras_state *state = &xe->ras.state;
>>> enum xe_ras_page_action action;
>>> u64 addr;
>>> int ret = 0;
>>> @@ -297,7 +306,11 @@ static int handle_page_offline(struct xe_device
>>> *xe, u64 page_address, bool send
>>> page_address);
>>> return ret;
>>> case -EEXIST:
>>> - action = XE_RAS_PAGE_ACTION_OFFLINE;
>>> + if (xa_load(&state->offlined_pages, page_address >>
>>> XE_PTE_SHIFT))
>>> + action = XE_RAS_PAGE_ACTION_REMOVE;
>>> + else
>>> + action = XE_RAS_PAGE_ACTION_OFFLINE;
>>> +
>>> xe_log_err(xe, DEVICE_MEMORY, ret,
>>> "Double-bit ECC error detected at physical address
>>> 0x%llx, page soft-offlined\n",
>>> page_address);
>>> @@ -341,6 +354,8 @@ static void get_offline_pages(struct xe_device
>>> *xe, u32 cmd, void *req, size_t r
>>> {
>>> struct xe_sysctrl_mailbox_command command = {0};
>>> struct xe_ras_offline_list_request *list_req;
>>> + struct xe_ras_state *state = &xe->ras.state;
>>> + unsigned long index;
>>> u32 total_pages = 0, count = 0;
>>> ssize_t rlen;
>>> int ret, i;
>>> @@ -370,8 +385,15 @@ static void get_offline_pages(struct xe_device
>>> *xe, u32 cmd, void *req, size_t r
>>> return;
>>> }
>>> - for (i = 0; i < common->pages_returned && i <
>>> XE_RAS_NUM_PAGES; i++)
>>> + for (i = 0; i < common->pages_returned && i <
>>> XE_RAS_NUM_PAGES; i++) {
>>> handle_page_offline(xe, common->page_addresses[i],
>>> offline);
>>> + /* The pages are already offlined by firmware */
>>> + if (!offline) {
>>> + index = common->page_addresses[i] >> XE_PTE_SHIFT;
>>> + xa_store(&state->offlined_pages, index,
>>> + XE_RAS_PAGE_OFFLINED, GFP_KERNEL);
>>> + }
>>> + }
>
> Do we need change order for
> get_offlined_list and get_queued_pages ?
Yeah sashiko comments seem valid. Will re-order it.
Thanks
Riana
>
> Missed them earlier.>> count += common->pages_returned;
>>> if (!common->pages_returned)
>>> @@ -1098,6 +1120,13 @@ static const struct attribute_group
>>> gpu_health_group = {
>>> .attrs = gpu_health_attrs,
>>> };
>>> +static void ras_fini(void *arg)
>>> +{
>>> + struct xe_device *xe = arg;
>>> +
>>> + xa_destroy(&xe->ras.state.offlined_pages);
>>> +}
>>> +
>>> /**
>>> * xe_ras_get_disable_page_offline - Get the page offline user policy
>>> * @xe: xe device instance
>>> @@ -1125,15 +1154,18 @@ u32 xe_ras_get_max_pages(struct xe_device *xe)
>>> * @xe: xe device instance
>>> *
>>> * Initialize Xe RAS
>>> + *
>>> + * Return: 0 on success, negative error code otherwise
>>> */
>>> -void xe_ras_init(struct xe_device *xe)
>>> +int xe_ras_init(struct xe_device *xe)
>>> {
>>> + struct xe_ras_state *state = &xe->ras.state;
>>> int ret;
>>> xe_drm_ras_init(xe);
>>> if (!xe->info.has_sysctrl)
>>> - return;
>>> + return 0;
>>> /*
>>> * TODO: Replace platform check with xe-
>>> >info.has_disable_vram_page_offline
>>> @@ -1145,10 +1177,14 @@ void xe_ras_init(struct xe_device *xe)
>>> if (IS_ENABLED(CONFIG_PCIEAER))
>>> ras_usp_aer_init(xe);
>>> + xa_init(&state->offlined_pages);
>>> +
>>> get_queued_pages(xe);
>>> get_offlined_list(xe);
>>> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
>>> if (ret)
>>> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n",
>>> ret);
>>> +
>>> + return devm_add_action_or_reset(xe->drm.dev, ras_fini, xe);
>>> }
>>> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h
>>> index 561c656e2ad3..e41a52869085 100644
>>> --- a/drivers/gpu/drm/xe/xe_ras.h
>>> +++ b/drivers/gpu/drm/xe/xe_ras.h
>>> @@ -19,7 +19,7 @@ int xe_ras_get_counter(struct xe_device *xe, u8
>>> severity, u8 component, u32 *val
>>> int xe_ras_clear_counter(struct xe_device *xe, u8 severity, u8
>>> component);
>>> int xe_ras_get_threshold(struct xe_device *xe, u8 severity, u8
>>> component, u32 *threshold);
>>> int xe_ras_set_threshold(struct xe_device *xe, u8 severity, u8
>>> component, u32 threshold);
>>> -void xe_ras_init(struct xe_device *xe);
>>> +int xe_ras_init(struct xe_device *xe);
>>> enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device
>>> *xe);
>>> bool xe_ras_get_disable_page_offline(struct xe_device *xe);
>>> u32 xe_ras_get_max_pages(struct xe_device *xe);
>>> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/
>>> xe_ras_types.h
>>> index e09b50a6f77b..50d70e2e2fd3 100644
>>> --- a/drivers/gpu/drm/xe/xe_ras_types.h
>>> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
>>> @@ -7,6 +7,7 @@
>>> #define _XE_RAS_TYPES_H_
>>> #include <linux/types.h>
>>> +#include <linux/xarray.h>
>>> #define XE_RAS_NUM_COUNTERS 16
>>> #define XE_RAS_NUM_ERROR_ARR 3
>>> @@ -417,5 +418,7 @@ struct xe_ras_state {
>>> bool disable_page_offline;
>>> /** @max_pages: Total number of pages that can be stored by
>>> firmware */
>>> u32 max_pages;
>>> + /** @offlined_pages: XArray of pages offlined by firmware */
>>> + struct xarray offlined_pages;
>>
>> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>
> Sashiko comments look valid and Need addressing.
>>
>>> };
>>> #endif
>>
>
^ permalink raw reply [flat|nested] 28+ messages in thread
* ✓ CI.KUnit: success for Add support to handle memory double-bit ecc errors (rev3)
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (5 preceding siblings ...)
2026-09-28 6:18 ` [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates Riana Tauro
@ 2026-09-28 14:42 ` Patchwork
2026-09-28 15:27 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-28 17:40 ` ✓ Xe.CI.FULL: " Patchwork
8 siblings, 0 replies; 28+ messages in thread
From: Patchwork @ 2026-09-28 14:42 UTC (permalink / raw)
To: Tauro, Riana; +Cc: intel-xe
== Series Details ==
Series: Add support to handle memory double-bit ecc errors (rev3)
URL : https://patchwork.freedesktop.org/series/172710/
State : success
== Summary ==
+ trap cleanup EXIT
+ kunitconfigs=('/kernel/drivers/gpu/tests/.kunitconfig' '/kernel/drivers/gpu/drm/xe/.kunitconfig' '/kernel/drivers/gpu/drm/tests/.kunitconfig' '/kernel/drivers/gpu/drm/ttm/tests/.kunitconfig' '/kernel/drivers/dma-buf/.kunitconfig')
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/tests/.kunitconfig
[14:40:42] Configuring KUnit Kernel ...
Generating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[14:40:46] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[14:41:07] Starting KUnit Kernel (1/1)...
[14:41:07] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[14:41:07] ============= refcount_interrupt (4 subtests) ==============
[14:41:07] [PASSED] test_single_irq_change
[14:41:07] [PASSED] test_nested_irq_change
[14:41:07] [PASSED] test_multiple_irq_change
[14:41:07] [PASSED] test_irq_save
[14:41:07] =============== [PASSED] refcount_interrupt ================
[14:41:07] ================= gpu_buddy (14 subtests) ==================
[14:41:07] [PASSED] gpu_test_buddy_alloc_limit
[14:41:07] [PASSED] gpu_test_buddy_alloc_optimistic
[14:41:07] [PASSED] gpu_test_buddy_alloc_pessimistic
[14:41:07] [PASSED] gpu_test_buddy_alloc_pathological
[14:41:07] [PASSED] gpu_test_buddy_alloc_contiguous
[14:41:07] [PASSED] gpu_test_buddy_alloc_clear
[14:41:07] [PASSED] gpu_test_buddy_alloc_range
[14:41:07] [PASSED] gpu_test_buddy_alloc_range_bias
[14:41:08] [PASSED] gpu_test_buddy_fragmentation_performance
[14:41:09] [PASSED] gpu_test_buddy_dirty_tracker_performance
[14:41:09] [PASSED] gpu_test_buddy_alloc_exceeds_max_order
[14:41:09] [PASSED] gpu_test_buddy_offset_aligned_allocation
[14:41:09] [PASSED] gpu_test_buddy_subtree_offset_alignment_stress
[14:41:09] [PASSED] gpu_test_buddy_addr_to_block
[14:41:09] ==================== [PASSED] gpu_buddy ====================
[14:41:09] ============================================================
[14:41:09] Testing complete. Ran 18 tests: passed: 18
[14:41:09] Elapsed time: 26.713s total, 4.385s configuring, 20.410s building, 1.853s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/xe/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/xe/.kunitconfig
[14:41:09] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[14:41:11] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[14:41:44] Starting KUnit Kernel (1/1)...
[14:41:44] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[14:41:44] ============= refcount_interrupt (4 subtests) ==============
[14:41:44] [PASSED] test_single_irq_change
[14:41:44] [PASSED] test_nested_irq_change
[14:41:44] [PASSED] test_multiple_irq_change
[14:41:44] [PASSED] test_irq_save
[14:41:44] =============== [PASSED] refcount_interrupt ================
[14:41:44] ================== guc_buf (11 subtests) ===================
[14:41:44] [PASSED] test_smallest
[14:41:44] [PASSED] test_largest
[14:41:44] [PASSED] test_granular
[14:41:44] [PASSED] test_unique
[14:41:44] [PASSED] test_overlap
[14:41:44] [PASSED] test_reusable
[14:41:44] [PASSED] test_too_big
[14:41:44] [PASSED] test_flush
[14:41:44] [PASSED] test_lookup
[14:41:44] [PASSED] test_data
[14:41:44] [PASSED] test_class
[14:41:44] ===================== [PASSED] guc_buf =====================
[14:41:44] =================== guc_dbm (7 subtests) ===================
[14:41:44] [PASSED] test_empty
[14:41:44] [PASSED] test_default
[14:41:44] ======================== test_size ========================
[14:41:44] [PASSED] 4
[14:41:44] [PASSED] 8
[14:41:44] [PASSED] 32
[14:41:44] [PASSED] 256
[14:41:44] ==================== [PASSED] test_size ====================
[14:41:44] ======================= test_reuse ========================
[14:41:44] [PASSED] 4
[14:41:44] [PASSED] 8
[14:41:44] [PASSED] 32
[14:41:44] [PASSED] 256
[14:41:44] =================== [PASSED] test_reuse ====================
[14:41:44] =================== test_range_overlap ====================
[14:41:44] [PASSED] 4
[14:41:44] [PASSED] 8
[14:41:44] [PASSED] 32
[14:41:44] [PASSED] 256
[14:41:44] =============== [PASSED] test_range_overlap ================
[14:41:44] =================== test_range_compact ====================
[14:41:44] [PASSED] 4
[14:41:44] [PASSED] 8
[14:41:44] [PASSED] 32
[14:41:44] [PASSED] 256
[14:41:44] =============== [PASSED] test_range_compact ================
[14:41:44] ==================== test_range_spare =====================
[14:41:44] [PASSED] 4
[14:41:44] [PASSED] 8
[14:41:44] [PASSED] 32
[14:41:44] [PASSED] 256
[14:41:44] ================ [PASSED] test_range_spare =================
[14:41:44] ===================== [PASSED] guc_dbm =====================
[14:41:44] =================== guc_idm (6 subtests) ===================
[14:41:44] [PASSED] bad_init
[14:41:44] [PASSED] no_init
[14:41:44] [PASSED] init_fini
[14:41:44] [PASSED] check_used
[14:41:45] [PASSED] check_quota
[14:41:45] [PASSED] check_all
[14:41:45] ===================== [PASSED] guc_idm =====================
[14:41:45] ============== guc_klv_helpers (13 subtests) ===============
[14:41:45] [PASSED] test_count
[14:41:45] [PASSED] test_encode_u32
[14:41:45] [PASSED] test_encode_u64
[14:41:45] [PASSED] test_encode_string
[14:41:45] [PASSED] test_encode_object_raw
[14:41:45] [PASSED] test_encode_object_klv
[14:41:45] [PASSED] test_encode_object_nested
[14:41:45] [PASSED] test_encode_object_basic
[14:41:45] ===================== test_decode_u16 =====================
[14:41:45] [PASSED] empty
[14:41:45] [PASSED] valid
[14:41:45] [PASSED] max
[14:41:45] [PASSED] big
[14:41:45] [PASSED] long
[14:41:45] ================= [PASSED] test_decode_u16 =================
[14:41:45] ===================== test_decode_u32 =====================
[14:41:45] [PASSED] empty
[14:41:45] [PASSED] valid
[14:41:45] [PASSED] max
[14:41:45] [PASSED] long
[14:41:45] ================= [PASSED] test_decode_u32 =================
[14:41:45] ===================== test_decode_u64 =====================
[14:41:45] [PASSED] empty
[14:41:45] [PASSED] short
[14:41:45] [PASSED] valid
[14:41:45] [PASSED] max
[14:41:45] [PASSED] long
[14:41:45] ================= [PASSED] test_decode_u64 =================
[14:41:45] ===================== test_decode_str =====================
[14:41:45] [PASSED] empty
[14:41:45] [PASSED] one
[14:41:45] [PASSED] one_trash
[14:41:45] [PASSED] two
[14:41:45] [PASSED] two_trash
[14:41:45] [PASSED] three
[14:41:45] [PASSED] dword
[14:41:45] [PASSED] four
[14:41:45] [PASSED] five
[14:41:45] [PASSED] six
[14:41:45] [PASSED] seven
[14:41:45] [PASSED] qword
[14:41:45] ================= [PASSED] test_decode_str =================
[14:41:45] [PASSED] test_print
[14:41:45] ================= [PASSED] guc_klv_helpers =================
[14:41:45] =================== xe_log (4 subtests) ====================
[14:41:45] [PASSED] demo_cper
[14:41:45] [PASSED] demo_dmesg
[14:41:45] ======================= test_dmesg ========================
[14:41:45] [PASSED] test_fatal
[14:41:45] [PASSED] test_fatal_tile
[14:41:45] [PASSED] test_fatal_gt
[14:41:45] [PASSED] test_fatal_comp
[14:41:45] [PASSED] test_fatal_comp_tile
[14:41:45] [PASSED] test_fatal_comp_gt
[14:41:45] [PASSED] test_fatal_all
[14:41:45] [PASSED] test_recoverable
[14:41:45] [PASSED] test_recoverable_tile
[14:41:45] [PASSED] test_recoverable_gt
[14:41:45] [PASSED] test_recoverable_comp
[14:41:45] [PASSED] test_recoverable_comp_tile
[14:41:45] [PASSED] test_recoverable_comp_gt
[14:41:45] [PASSED] test_recoverable_all
[14:41:45] [PASSED] test_info
[14:41:45] [PASSED] test_info_tile
[14:41:45] [PASSED] test_info_gt
[14:41:45] [PASSED] test_info_err
[14:41:45] [PASSED] test_info_comp
[14:41:45] [PASSED] test_info_comp_tile
[14:41:45] [PASSED] test_info_comp_gt
[14:41:45] [PASSED] test_info_all
[14:41:45] [PASSED] test_hw_fatal
[14:41:45] [PASSED] test_hw_recoverable
[14:41:45] [PASSED] test_hw_corrected
[14:41:45] [PASSED] test_hw_informational
[14:41:45] =================== [PASSED] test_dmesg ====================
[14:41:45] ====================== test_invalid =======================
[14:41:45] [SKIPPED] no-component no-location no-warn (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] reserved location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] unknown location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] nonzero-device-id location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] invalid-tile-id location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] invalid-gt-id location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] unknown component class (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] unknown system component (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] unknown hardware component (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] [SKIPPED] unknown component and location (requires CONFIG_DRM_XE_DEBUG)
[14:41:45] ================== [SKIPPED] test_invalid ==================
[14:41:45] ===================== [PASSED] xe_log ======================
[14:41:45] ================== no_relay (3 subtests) ===================
[14:41:45] [PASSED] xe_drops_guc2pf_if_not_ready
[14:41:45] [PASSED] xe_drops_guc2vf_if_not_ready
[14:41:45] [PASSED] xe_rejects_send_if_not_ready
[14:41:45] ==================== [PASSED] no_relay =====================
[14:41:45] ================== pf_relay (14 subtests) ==================
[14:41:45] [PASSED] pf_rejects_guc2pf_too_short
[14:41:45] [PASSED] pf_rejects_guc2pf_too_long
[14:41:45] [PASSED] pf_rejects_guc2pf_no_payload
[14:41:45] [PASSED] pf_fails_no_payload
[14:41:45] [PASSED] pf_fails_bad_origin
[14:41:45] [PASSED] pf_fails_bad_type
[14:41:45] [PASSED] pf_txn_reports_error
[14:41:45] [PASSED] pf_txn_sends_pf2guc
[14:41:45] [PASSED] pf_sends_pf2guc
[14:41:45] [SKIPPED] pf_loopback_nop (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[14:41:45] [SKIPPED] pf_loopback_echo (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[14:41:45] [SKIPPED] pf_loopback_fail (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[14:41:45] [SKIPPED] pf_loopback_busy (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[14:41:45] [SKIPPED] pf_loopback_retry (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[14:41:45] ==================== [PASSED] pf_relay =====================
[14:41:45] ================== vf_relay (3 subtests) ===================
[14:41:45] [PASSED] vf_rejects_guc2vf_too_short
[14:41:45] [PASSED] vf_rejects_guc2vf_too_long
[14:41:45] [PASSED] vf_rejects_guc2vf_no_payload
[14:41:45] ==================== [PASSED] vf_relay =====================
[14:41:45] ================ pf_gt_config (9 subtests) =================
[14:41:45] [PASSED] fair_contexts_1vf
[14:41:45] [PASSED] fair_doorbells_1vf
[14:41:45] [PASSED] fair_ggtt_1vf
[14:41:45] ====================== fair_vram_1vf ======================
[14:41:45] [PASSED] 3.50 GiB
[14:41:45] [PASSED] 11.5 GiB
[14:41:45] [PASSED] 15.5 GiB
[14:41:45] [PASSED] 31.5 GiB
[14:41:45] [PASSED] 63.5 GiB
[14:41:45] [PASSED] 1.91 GiB
[14:41:45] ================== [PASSED] fair_vram_1vf ==================
[14:41:45] ================ fair_vram_1vf_admin_only =================
[14:41:45] [PASSED] 3.50 GiB
[14:41:45] [PASSED] 11.5 GiB
[14:41:45] [PASSED] 15.5 GiB
[14:41:45] [PASSED] 31.5 GiB
[14:41:45] [PASSED] 63.5 GiB
[14:41:45] [PASSED] 1.91 GiB
[14:41:45] ============ [PASSED] fair_vram_1vf_admin_only =============
[14:41:45] ====================== fair_contexts ======================
[14:41:45] [PASSED] 1 VF
[14:41:45] [PASSED] 2 VFs
[14:41:45] [PASSED] 3 VFs
[14:41:45] [PASSED] 4 VFs
[14:41:45] [PASSED] 5 VFs
[14:41:45] [PASSED] 6 VFs
[14:41:45] [PASSED] 7 VFs
[14:41:45] [PASSED] 8 VFs
[14:41:45] [PASSED] 9 VFs
[14:41:45] [PASSED] 10 VFs
[14:41:45] [PASSED] 11 VFs
[14:41:45] [PASSED] 12 VFs
[14:41:45] [PASSED] 13 VFs
[14:41:45] [PASSED] 14 VFs
[14:41:45] [PASSED] 15 VFs
[14:41:45] [PASSED] 16 VFs
[14:41:45] [PASSED] 17 VFs
[14:41:45] [PASSED] 18 VFs
[14:41:45] [PASSED] 19 VFs
[14:41:45] [PASSED] 20 VFs
[14:41:45] [PASSED] 21 VFs
[14:41:45] [PASSED] 22 VFs
[14:41:45] [PASSED] 23 VFs
[14:41:45] [PASSED] 24 VFs
[14:41:45] [PASSED] 25 VFs
[14:41:45] [PASSED] 26 VFs
[14:41:45] [PASSED] 27 VFs
[14:41:45] [PASSED] 28 VFs
[14:41:45] [PASSED] 29 VFs
[14:41:45] [PASSED] 30 VFs
[14:41:45] [PASSED] 31 VFs
[14:41:45] [PASSED] 32 VFs
[14:41:45] [PASSED] 33 VFs
[14:41:45] [PASSED] 34 VFs
[14:41:45] [PASSED] 35 VFs
[14:41:45] [PASSED] 36 VFs
[14:41:45] [PASSED] 37 VFs
[14:41:45] [PASSED] 38 VFs
[14:41:45] [PASSED] 39 VFs
[14:41:45] [PASSED] 40 VFs
[14:41:45] [PASSED] 41 VFs
[14:41:45] [PASSED] 42 VFs
[14:41:45] [PASSED] 43 VFs
[14:41:45] [PASSED] 44 VFs
[14:41:45] [PASSED] 45 VFs
[14:41:45] [PASSED] 46 VFs
[14:41:45] [PASSED] 47 VFs
[14:41:45] [PASSED] 48 VFs
[14:41:45] [PASSED] 49 VFs
[14:41:45] [PASSED] 50 VFs
[14:41:45] [PASSED] 51 VFs
[14:41:45] [PASSED] 52 VFs
[14:41:45] [PASSED] 53 VFs
[14:41:45] [PASSED] 54 VFs
[14:41:45] [PASSED] 55 VFs
[14:41:45] [PASSED] 56 VFs
[14:41:45] [PASSED] 57 VFs
[14:41:45] [PASSED] 58 VFs
[14:41:45] [PASSED] 59 VFs
[14:41:45] [PASSED] 60 VFs
[14:41:45] [PASSED] 61 VFs
[14:41:45] [PASSED] 62 VFs
[14:41:45] [PASSED] 63 VFs
[14:41:45] ================== [PASSED] fair_contexts ==================
[14:41:45] ===================== fair_doorbells ======================
[14:41:45] [PASSED] 1 VF
[14:41:45] [PASSED] 2 VFs
[14:41:45] [PASSED] 3 VFs
[14:41:45] [PASSED] 4 VFs
[14:41:45] [PASSED] 5 VFs
[14:41:45] [PASSED] 6 VFs
[14:41:45] [PASSED] 7 VFs
[14:41:45] [PASSED] 8 VFs
[14:41:45] [PASSED] 9 VFs
[14:41:45] [PASSED] 10 VFs
[14:41:45] [PASSED] 11 VFs
[14:41:45] [PASSED] 12 VFs
[14:41:45] [PASSED] 13 VFs
[14:41:45] [PASSED] 14 VFs
[14:41:45] [PASSED] 15 VFs
[14:41:45] [PASSED] 16 VFs
[14:41:45] [PASSED] 17 VFs
[14:41:45] [PASSED] 18 VFs
[14:41:45] [PASSED] 19 VFs
[14:41:45] [PASSED] 20 VFs
[14:41:45] [PASSED] 21 VFs
[14:41:45] [PASSED] 22 VFs
[14:41:45] [PASSED] 23 VFs
[14:41:45] [PASSED] 24 VFs
[14:41:45] [PASSED] 25 VFs
[14:41:45] [PASSED] 26 VFs
[14:41:45] [PASSED] 27 VFs
[14:41:45] [PASSED] 28 VFs
[14:41:45] [PASSED] 29 VFs
[14:41:45] [PASSED] 30 VFs
[14:41:45] [PASSED] 31 VFs
[14:41:45] [PASSED] 32 VFs
[14:41:45] [PASSED] 33 VFs
[14:41:45] [PASSED] 34 VFs
[14:41:45] [PASSED] 35 VFs
[14:41:45] [PASSED] 36 VFs
[14:41:45] [PASSED] 37 VFs
[14:41:45] [PASSED] 38 VFs
[14:41:45] [PASSED] 39 VFs
[14:41:45] [PASSED] 40 VFs
[14:41:45] [PASSED] 41 VFs
[14:41:45] [PASSED] 42 VFs
[14:41:45] [PASSED] 43 VFs
[14:41:45] [PASSED] 44 VFs
[14:41:45] [PASSED] 45 VFs
[14:41:45] [PASSED] 46 VFs
[14:41:45] [PASSED] 47 VFs
[14:41:45] [PASSED] 48 VFs
[14:41:45] [PASSED] 49 VFs
[14:41:45] [PASSED] 50 VFs
[14:41:45] [PASSED] 51 VFs
[14:41:45] [PASSED] 52 VFs
[14:41:45] [PASSED] 53 VFs
[14:41:45] [PASSED] 54 VFs
[14:41:45] [PASSED] 55 VFs
[14:41:45] [PASSED] 56 VFs
[14:41:45] [PASSED] 57 VFs
[14:41:45] [PASSED] 58 VFs
[14:41:45] [PASSED] 59 VFs
[14:41:45] [PASSED] 60 VFs
[14:41:45] [PASSED] 61 VFs
[14:41:45] [PASSED] 62 VFs
[14:41:45] [PASSED] 63 VFs
[14:41:45] ================= [PASSED] fair_doorbells ==================
[14:41:45] ======================== fair_ggtt ========================
[14:41:45] [PASSED] 1 VF
[14:41:45] [PASSED] 2 VFs
[14:41:45] [PASSED] 3 VFs
[14:41:45] [PASSED] 4 VFs
[14:41:45] [PASSED] 5 VFs
[14:41:45] [PASSED] 6 VFs
[14:41:45] [PASSED] 7 VFs
[14:41:45] [PASSED] 8 VFs
[14:41:45] [PASSED] 9 VFs
[14:41:45] [PASSED] 10 VFs
[14:41:45] [PASSED] 11 VFs
[14:41:45] [PASSED] 12 VFs
[14:41:45] [PASSED] 13 VFs
[14:41:45] [PASSED] 14 VFs
[14:41:45] [PASSED] 15 VFs
[14:41:45] [PASSED] 16 VFs
[14:41:45] [PASSED] 17 VFs
[14:41:45] [PASSED] 18 VFs
[14:41:45] [PASSED] 19 VFs
[14:41:45] [PASSED] 20 VFs
[14:41:45] [PASSED] 21 VFs
[14:41:45] [PASSED] 22 VFs
[14:41:45] [PASSED] 23 VFs
[14:41:45] [PASSED] 24 VFs
[14:41:45] [PASSED] 25 VFs
[14:41:45] [PASSED] 26 VFs
[14:41:45] [PASSED] 27 VFs
[14:41:45] [PASSED] 28 VFs
[14:41:45] [PASSED] 29 VFs
[14:41:45] [PASSED] 30 VFs
[14:41:45] [PASSED] 31 VFs
[14:41:45] [PASSED] 32 VFs
[14:41:45] [PASSED] 33 VFs
[14:41:45] [PASSED] 34 VFs
[14:41:45] [PASSED] 35 VFs
[14:41:45] [PASSED] 36 VFs
[14:41:45] [PASSED] 37 VFs
[14:41:45] [PASSED] 38 VFs
[14:41:45] [PASSED] 39 VFs
[14:41:45] [PASSED] 40 VFs
[14:41:45] [PASSED] 41 VFs
[14:41:45] [PASSED] 42 VFs
[14:41:45] [PASSED] 43 VFs
[14:41:45] [PASSED] 44 VFs
[14:41:45] [PASSED] 45 VFs
[14:41:45] [PASSED] 46 VFs
[14:41:45] [PASSED] 47 VFs
[14:41:45] [PASSED] 48 VFs
[14:41:45] [PASSED] 49 VFs
[14:41:45] [PASSED] 50 VFs
[14:41:45] [PASSED] 51 VFs
[14:41:45] [PASSED] 52 VFs
[14:41:45] [PASSED] 53 VFs
[14:41:45] [PASSED] 54 VFs
[14:41:45] [PASSED] 55 VFs
[14:41:45] [PASSED] 56 VFs
[14:41:45] [PASSED] 57 VFs
[14:41:45] [PASSED] 58 VFs
[14:41:45] [PASSED] 59 VFs
[14:41:45] [PASSED] 60 VFs
[14:41:45] [PASSED] 61 VFs
[14:41:45] [PASSED] 62 VFs
[14:41:45] [PASSED] 63 VFs
[14:41:45] ==================== [PASSED] fair_ggtt ====================
[14:41:45] ======================== fair_vram ========================
[14:41:45] [PASSED] 1 VF
[14:41:45] [PASSED] 2 VFs
[14:41:45] [PASSED] 3 VFs
[14:41:45] [PASSED] 4 VFs
[14:41:45] [PASSED] 5 VFs
[14:41:45] [PASSED] 6 VFs
[14:41:45] [PASSED] 7 VFs
[14:41:45] [PASSED] 8 VFs
[14:41:45] [PASSED] 9 VFs
[14:41:45] [PASSED] 10 VFs
[14:41:45] [PASSED] 11 VFs
[14:41:45] [PASSED] 12 VFs
[14:41:45] [PASSED] 13 VFs
[14:41:45] [PASSED] 14 VFs
[14:41:45] [PASSED] 15 VFs
[14:41:45] [PASSED] 16 VFs
[14:41:45] [PASSED] 17 VFs
[14:41:45] [PASSED] 18 VFs
[14:41:45] [PASSED] 19 VFs
[14:41:45] [PASSED] 20 VFs
[14:41:45] [PASSED] 21 VFs
[14:41:45] [PASSED] 22 VFs
[14:41:45] [PASSED] 23 VFs
[14:41:45] [PASSED] 24 VFs
[14:41:45] [PASSED] 25 VFs
[14:41:45] [PASSED] 26 VFs
[14:41:45] [PASSED] 27 VFs
[14:41:45] [PASSED] 28 VFs
[14:41:45] [PASSED] 29 VFs
[14:41:45] [PASSED] 30 VFs
[14:41:45] [PASSED] 31 VFs
[14:41:45] [PASSED] 32 VFs
[14:41:45] [PASSED] 33 VFs
[14:41:45] [PASSED] 34 VFs
[14:41:45] [PASSED] 35 VFs
[14:41:45] [PASSED] 36 VFs
[14:41:45] [PASSED] 37 VFs
[14:41:45] [PASSED] 38 VFs
[14:41:45] [PASSED] 39 VFs
[14:41:45] [PASSED] 40 VFs
[14:41:45] [PASSED] 41 VFs
[14:41:45] [PASSED] 42 VFs
[14:41:45] [PASSED] 43 VFs
[14:41:45] [PASSED] 44 VFs
[14:41:45] [PASSED] 45 VFs
[14:41:45] [PASSED] 46 VFs
[14:41:45] [PASSED] 47 VFs
[14:41:45] [PASSED] 48 VFs
[14:41:45] [PASSED] 49 VFs
[14:41:45] [PASSED] 50 VFs
[14:41:45] [PASSED] 51 VFs
[14:41:45] [PASSED] 52 VFs
[14:41:45] [PASSED] 53 VFs
[14:41:45] [PASSED] 54 VFs
[14:41:45] [PASSED] 55 VFs
[14:41:45] [PASSED] 56 VFs
[14:41:45] [PASSED] 57 VFs
[14:41:45] [PASSED] 58 VFs
[14:41:45] [PASSED] 59 VFs
[14:41:45] [PASSED] 60 VFs
[14:41:45] [PASSED] 61 VFs
[14:41:45] [PASSED] 62 VFs
[14:41:45] [PASSED] 63 VFs
[14:41:45] ==================== [PASSED] fair_vram ====================
[14:41:45] ================== [PASSED] pf_gt_config ===================
[14:41:45] ===================== lmtt (1 subtest) =====================
[14:41:45] ======================== test_ops =========================
[14:41:45] [PASSED] 2-level
[14:41:45] [PASSED] multi-level
[14:41:45] ==================== [PASSED] test_ops =====================
[14:41:45] ====================== [PASSED] lmtt =======================
[14:41:45] ================= sriov_packet (1 subtest) =================
[14:41:45] [PASSED] test_descriptor_init
[14:41:45] ================== [PASSED] sriov_packet ===================
[14:41:45] ================= pf_service (11 subtests) =================
[14:41:45] [PASSED] pf_negotiate_any
[14:41:45] [PASSED] pf_negotiate_base_match
[14:41:45] [PASSED] pf_negotiate_base_newer
[14:41:45] [PASSED] pf_negotiate_base_next
[14:41:45] [SKIPPED] pf_negotiate_base_older (no older minor)
[14:41:45] [PASSED] pf_negotiate_base_prev
[14:41:45] [PASSED] pf_negotiate_latest_match
[14:41:45] [PASSED] pf_negotiate_latest_newer
[14:41:45] [PASSED] pf_negotiate_latest_next
[14:41:45] [SKIPPED] pf_negotiate_latest_older (no older minor)
[14:41:45] [SKIPPED] pf_negotiate_latest_prev (no prev major)
[14:41:45] =================== [PASSED] pf_service ====================
[14:41:45] ================= xe_guc_g2g (2 subtests) ==================
[14:41:45] ============== xe_live_guc_g2g_kunit_default ==============
[14:41:45] ========= [SKIPPED] xe_live_guc_g2g_kunit_default ==========
[14:41:45] ============== xe_live_guc_g2g_kunit_allmem ===============
[14:41:45] ========== [SKIPPED] xe_live_guc_g2g_kunit_allmem ==========
[14:41:45] =================== [SKIPPED] xe_guc_g2g ===================
[14:41:45] =================== xe_mocs (2 subtests) ===================
[14:41:45] ================ xe_live_mocs_kernel_kunit ================
[14:41:45] =========== [SKIPPED] xe_live_mocs_kernel_kunit ============
[14:41:45] ================ xe_live_mocs_reset_kunit =================
[14:41:45] ============ [SKIPPED] xe_live_mocs_reset_kunit ============
[14:41:45] ==================== [SKIPPED] xe_mocs =====================
[14:41:45] ================= xe_migrate (2 subtests) ==================
[14:41:45] ================= xe_migrate_sanity_kunit =================
[14:41:45] ============ [SKIPPED] xe_migrate_sanity_kunit =============
[14:41:45] ================== xe_validate_ccs_kunit ==================
[14:41:45] ============= [SKIPPED] xe_validate_ccs_kunit ==============
[14:41:45] =================== [SKIPPED] xe_migrate ===================
[14:41:45] ================== xe_dma_buf (1 subtest) ==================
[14:41:45] ==================== xe_dma_buf_kunit =====================
[14:41:45] ================ [SKIPPED] xe_dma_buf_kunit ================
[14:41:45] =================== [SKIPPED] xe_dma_buf ===================
[14:41:45] ================= xe_bo_shrink (1 subtest) =================
[14:41:45] =================== xe_bo_shrink_kunit ====================
[14:41:45] =============== [SKIPPED] xe_bo_shrink_kunit ===============
[14:41:45] ================== [SKIPPED] xe_bo_shrink ==================
[14:41:45] ==================== xe_bo (2 subtests) ====================
[14:41:45] ================== xe_ccs_migrate_kunit ===================
[14:41:45] ============== [SKIPPED] xe_ccs_migrate_kunit ==============
[14:41:45] ==================== xe_bo_evict_kunit ====================
[14:41:45] =============== [SKIPPED] xe_bo_evict_kunit ================
[14:41:45] ===================== [SKIPPED] xe_bo ======================
[14:41:45] =================== xe_any (9 subtests) ====================
[14:41:45] [PASSED] test_to_xe
[14:41:45] [PASSED] test_to_dev
[14:41:45] [PASSED] test_to_pdev
[14:41:45] [PASSED] test_to_drm
[14:41:45] [PASSED] test_if_pdev
[14:41:45] [PASSED] test_if_xe
[14:41:45] [PASSED] test_if_tile
[14:41:45] [PASSED] test_if_gt
[14:41:45] [PASSED] test_to_id
[14:41:45] ===================== [PASSED] xe_any ======================
[14:41:45] ==================== args (13 subtests) ====================
[14:41:45] [PASSED] count_args_test
[14:41:45] [PASSED] call_args_example
[14:41:45] [PASSED] call_args_test
[14:41:45] [PASSED] drop_first_arg_example
[14:41:45] [PASSED] drop_first_arg_test
[14:41:45] [PASSED] first_arg_example
[14:41:45] [PASSED] first_arg_test
[14:41:45] [PASSED] last_arg_example
[14:41:45] [PASSED] last_arg_test
[14:41:45] [PASSED] pick_arg_example
[14:41:45] [PASSED] if_args_example
[14:41:45] [PASSED] if_args_test
[14:41:45] [PASSED] sep_comma_example
[14:41:45] ====================== [PASSED] args =======================
[14:41:45] =================== xe_pci (3 subtests) ====================
[14:41:45] ==================== check_graphics_ip ====================
[14:41:45] [PASSED] 12.00 Xe_LP
[14:41:45] [PASSED] 12.10 Xe_LP+
[14:41:45] [PASSED] 12.55 Xe_HPG
[14:41:45] [PASSED] 12.60 Xe_HPC
[14:41:45] [PASSED] 12.70 Xe_LPG
[14:41:45] [PASSED] 12.71 Xe_LPG
[14:41:45] [PASSED] 12.74 Xe_LPG+
[14:41:45] [PASSED] 20.01 Xe2_HPG
[14:41:45] [PASSED] 20.02 Xe2_HPG
[14:41:45] [PASSED] 20.04 Xe2_LPG
[14:41:45] [PASSED] 30.00 Xe3_LPG
[14:41:45] [PASSED] 30.01 Xe3_LPG
[14:41:45] [PASSED] 30.03 Xe3_LPG
[14:41:45] [PASSED] 30.04 Xe3_LPG
[14:41:45] [PASSED] 30.05 Xe3_LPG
[14:41:45] [PASSED] 35.10 Xe3p_LPG
[14:41:45] [PASSED] 35.11 Xe3p_XPC
[14:41:45] ================ [PASSED] check_graphics_ip ================
[14:41:45] ===================== check_media_ip ======================
[14:41:45] [PASSED] 12.00 Xe_M
[14:41:45] [PASSED] 12.55 Xe_HPM
[14:41:45] [PASSED] 13.00 Xe_LPM+
[14:41:45] [PASSED] 13.01 Xe2_HPM
[14:41:45] [PASSED] 20.00 Xe2_LPM
[14:41:45] [PASSED] 30.00 Xe3_LPM
[14:41:45] [PASSED] 30.02 Xe3_LPM
[14:41:45] [PASSED] 35.00 Xe3p_LPM
[14:41:45] [PASSED] 35.03 Xe3p_HPM
[14:41:45] ================= [PASSED] check_media_ip ==================
[14:41:45] =================== check_platform_desc ===================
[14:41:45] [PASSED] 0x9A60 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A68 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A70 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A40 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A49 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A59 (TIGERLAKE)
[14:41:45] [PASSED] 0x9A78 (TIGERLAKE)
[14:41:45] [PASSED] 0x9AC0 (TIGERLAKE)
[14:41:45] [PASSED] 0x9AC9 (TIGERLAKE)
[14:41:45] [PASSED] 0x9AD9 (TIGERLAKE)
[14:41:45] [PASSED] 0x9AF8 (TIGERLAKE)
[14:41:45] [PASSED] 0x4C80 (ROCKETLAKE)
[14:41:45] [PASSED] 0x4C8A (ROCKETLAKE)
[14:41:45] [PASSED] 0x4C8B (ROCKETLAKE)
[14:41:45] [PASSED] 0x4C8C (ROCKETLAKE)
[14:41:45] [PASSED] 0x4C90 (ROCKETLAKE)
[14:41:45] [PASSED] 0x4C9A (ROCKETLAKE)
[14:41:45] [PASSED] 0x4680 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4682 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4688 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x468A (ALDERLAKE_S)
[14:41:45] [PASSED] 0x468B (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4690 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4692 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4693 (ALDERLAKE_S)
[14:41:45] [PASSED] 0x46A0 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46A1 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46A2 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46A3 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46A6 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46A8 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46AA (ALDERLAKE_P)
[14:41:45] [PASSED] 0x462A (ALDERLAKE_P)
[14:41:45] [PASSED] 0x4626 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x4628 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46B0 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46B1 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46B2 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46B3 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46C0 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46C1 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46C2 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46C3 (ALDERLAKE_P)
[14:41:45] [PASSED] 0x46D0 (ALDERLAKE_N)
[14:41:45] [PASSED] 0x46D1 (ALDERLAKE_N)
[14:41:45] [PASSED] 0x46D2 (ALDERLAKE_N)
[14:41:45] [PASSED] 0x46D3 (ALDERLAKE_N)
[14:41:45] [PASSED] 0x46D4 (ALDERLAKE_N)
[14:41:45] [PASSED] 0xA721 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7A1 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7A9 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7AC (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7AD (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA720 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7A0 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7A8 (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7AA (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA7AB (ALDERLAKE_P)
[14:41:45] [PASSED] 0xA780 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA781 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA782 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA783 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA788 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA789 (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA78A (ALDERLAKE_S)
[14:41:45] [PASSED] 0xA78B (ALDERLAKE_S)
[14:41:45] [PASSED] 0x4905 (DG1)
[14:41:45] [PASSED] 0x4906 (DG1)
[14:41:45] [PASSED] 0x4907 (DG1)
[14:41:45] [PASSED] 0x4908 (DG1)
[14:41:45] [PASSED] 0x4909 (DG1)
[14:41:45] [PASSED] 0x56C0 (DG2)
[14:41:45] [PASSED] 0x56C2 (DG2)
[14:41:45] [PASSED] 0x56C1 (DG2)
[14:41:45] [PASSED] 0x7D51 (METEORLAKE)
[14:41:45] [PASSED] 0x7DD1 (METEORLAKE)
[14:41:45] [PASSED] 0x7D41 (METEORLAKE)
[14:41:45] [PASSED] 0x7D67 (METEORLAKE)
[14:41:45] [PASSED] 0xB640 (METEORLAKE)
[14:41:45] [PASSED] 0x56A0 (DG2)
[14:41:45] [PASSED] 0x56A1 (DG2)
[14:41:45] [PASSED] 0x56A2 (DG2)
[14:41:45] [PASSED] 0x56BE (DG2)
[14:41:45] [PASSED] 0x56BF (DG2)
[14:41:45] [PASSED] 0x5690 (DG2)
[14:41:45] [PASSED] 0x5691 (DG2)
[14:41:45] [PASSED] 0x5692 (DG2)
[14:41:45] [PASSED] 0x56A5 (DG2)
[14:41:45] [PASSED] 0x56A6 (DG2)
[14:41:45] [PASSED] 0x56B0 (DG2)
[14:41:45] [PASSED] 0x56B1 (DG2)
[14:41:45] [PASSED] 0x56BA (DG2)
[14:41:45] [PASSED] 0x56BB (DG2)
[14:41:45] [PASSED] 0x56BC (DG2)
[14:41:45] [PASSED] 0x56BD (DG2)
[14:41:45] [PASSED] 0x5693 (DG2)
[14:41:45] [PASSED] 0x5694 (DG2)
[14:41:45] [PASSED] 0x5695 (DG2)
[14:41:45] [PASSED] 0x56A3 (DG2)
[14:41:45] [PASSED] 0x56A4 (DG2)
[14:41:45] [PASSED] 0x56B2 (DG2)
[14:41:45] [PASSED] 0x56B3 (DG2)
[14:41:45] [PASSED] 0x5696 (DG2)
[14:41:45] [PASSED] 0x5697 (DG2)
[14:41:45] [PASSED] 0xB69 (PVC)
[14:41:45] [PASSED] 0xB6E (PVC)
[14:41:45] [PASSED] 0xBD4 (PVC)
[14:41:45] [PASSED] 0xBD5 (PVC)
[14:41:45] [PASSED] 0xBD6 (PVC)
[14:41:45] [PASSED] 0xBD7 (PVC)
[14:41:45] [PASSED] 0xBD8 (PVC)
[14:41:45] [PASSED] 0xBD9 (PVC)
[14:41:45] [PASSED] 0xBDA (PVC)
[14:41:45] [PASSED] 0xBDB (PVC)
[14:41:45] [PASSED] 0xBE0 (PVC)
[14:41:45] [PASSED] 0xBE1 (PVC)
[14:41:45] [PASSED] 0xBE5 (PVC)
[14:41:45] [PASSED] 0x7D40 (METEORLAKE)
[14:41:45] [PASSED] 0x7D45 (METEORLAKE)
[14:41:45] [PASSED] 0x7D55 (METEORLAKE)
[14:41:45] [PASSED] 0x7D60 (METEORLAKE)
[14:41:45] [PASSED] 0x7DD5 (METEORLAKE)
[14:41:45] [PASSED] 0x6420 (LUNARLAKE)
[14:41:45] [PASSED] 0x64A0 (LUNARLAKE)
[14:41:45] [PASSED] 0x64B0 (LUNARLAKE)
[14:41:45] [PASSED] 0xE202 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE209 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE20B (BATTLEMAGE)
[14:41:45] [PASSED] 0xE20C (BATTLEMAGE)
[14:41:45] [PASSED] 0xE20D (BATTLEMAGE)
[14:41:45] [PASSED] 0xE210 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE211 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE212 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE216 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE220 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE221 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE222 (BATTLEMAGE)
[14:41:45] [PASSED] 0xE223 (BATTLEMAGE)
[14:41:45] [PASSED] 0xB080 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB081 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB082 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB083 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB084 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB085 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB086 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB087 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB08F (PANTHERLAKE)
[14:41:45] [PASSED] 0xB090 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB0A0 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB0A1 (PANTHERLAKE)
[14:41:45] [PASSED] 0xB0B0 (PANTHERLAKE)
[14:41:45] [PASSED] 0xFD80 (PANTHERLAKE)
[14:41:45] [PASSED] 0xFD81 (PANTHERLAKE)
[14:41:45] [PASSED] 0xD740 (NOVALAKE_S)
[14:41:45] [PASSED] 0xD741 (NOVALAKE_S)
[14:41:45] [PASSED] 0xD742 (NOVALAKE_S)
[14:41:45] [PASSED] 0xD743 (NOVALAKE_S)
[14:41:45] [PASSED] 0xD745 (NOVALAKE_S)
[14:41:45] [PASSED] 0xD74A (NOVALAKE_S)
[14:41:45] [PASSED] 0xD74B (NOVALAKE_S)
[14:41:45] [PASSED] 0x674C (CRESCENTISLAND)
[14:41:45] [PASSED] 0x674D (CRESCENTISLAND)
[14:41:45] [PASSED] 0x674E (CRESCENTISLAND)
[14:41:45] [PASSED] 0x674F (CRESCENTISLAND)
[14:41:45] [PASSED] 0x6750 (CRESCENTISLAND)
[14:41:45] [PASSED] 0xD750 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD751 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD752 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD753 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD754 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD755 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD756 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD757 (NOVALAKE_P)
[14:41:45] [PASSED] 0xD75F (NOVALAKE_P)
[14:41:45] =============== [PASSED] check_platform_desc ===============
[14:41:45] ===================== [PASSED] xe_pci ======================
[14:41:45] ============= xe_rtp_tables_test (5 subtests) ==============
[14:41:45] ================== xe_rtp_table_gt_test ===================
[14:41:45] [PASSED] gt_was/14011060649
[14:41:45] [PASSED] gt_was/14011059788
[14:41:45] [PASSED] gt_was/14015795083
[14:41:45] [PASSED] gt_was/16021867713
[14:41:45] [PASSED] gt_was/14019449301
[14:41:45] [PASSED] gt_was/16028005424
[14:41:45] [PASSED] gt_was/14026578760
[14:41:45] [PASSED] gt_was/1409420604
[14:41:45] [PASSED] gt_was/1408615072
[14:41:45] [PASSED] gt_was/22010523718
[14:41:45] [PASSED] gt_was/14011006942
[14:41:45] [PASSED] gt_was/14014830051
[14:41:45] [PASSED] gt_was/18018781329
[14:41:45] [PASSED] gt_was/1509235366
[14:41:45] [PASSED] gt_was/18018781329
[14:41:45] [PASSED] gt_was/16016694945
[14:41:45] [PASSED] gt_was/14018575942
[14:41:45] [PASSED] gt_was/22016670082
[14:41:45] [PASSED] gt_was/22016670082
[14:41:45] [PASSED] gt_was/14017421178
[14:41:45] [PASSED] gt_was/16025250150
[14:41:45] [PASSED] gt_was/14021871409
[14:41:45] [PASSED] gt_was/16021865536
[14:41:45] [PASSED] gt_was/14021486841
[14:41:45] [PASSED] gt_was/14025160223
[14:41:45] [PASSED] gt_was/14026144927, 16029437861, 14026127056
[14:41:45] [PASSED] gt_was/14025635424
[14:41:45] [PASSED] gt_was/16028005424
[14:41:45] ============== [PASSED] xe_rtp_table_gt_test ===============
[14:41:45] ================== xe_rtp_table_gt_test ===================
[14:41:45] [PASSED] gt_tunings/Tuning: Blend Fill Caching Optimization Disable
[14:41:45] [PASSED] gt_tunings/Tuning: 32B Access Enable
[14:41:45] [PASSED] gt_tunings/Tuning: L3 cache
[14:41:45] [PASSED] gt_tunings/Tuning: L3 cache - media
[14:41:45] [PASSED] gt_tunings/Tuning: Compression Overfetch
[14:41:45] [PASSED] gt_tunings/Tuning: Compression Overfetch - media
[14:41:45] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3
[14:41:45] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3 - media
[14:41:45] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only
[14:41:45] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only - media
[14:41:45] [PASSED] gt_tunings/Tuning: Stateless compression control
[14:41:45] [PASSED] gt_tunings/Tuning: Stateless compression control - media
[14:41:45] [PASSED] gt_tunings/Tuning: L3 RW flush all Cache
[14:41:45] [PASSED] gt_tunings/Tuning: L3 RW flush all cache - media
[14:41:45] [PASSED] gt_tunings/Tuning: Set STLB Bank Hash Mode to 4KB
[14:41:45] ============== [PASSED] xe_rtp_table_gt_test ===============
[14:41:45] ================== xe_rtp_table_oob_test ==================
[14:41:45] [PASSED] oob_was/1607983814
[14:41:45] [PASSED] oob_was/16010904313
[14:41:45] [PASSED] oob_was/18022495364
[14:41:45] [PASSED] oob_was/22012773006
[14:41:45] [PASSED] oob_was/14014475959
[14:41:45] [PASSED] oob_was/22011391025
[14:41:45] [PASSED] oob_was/22012727170
[14:41:45] [PASSED] oob_was/22012727685
[14:41:45] [PASSED] oob_was/22016596838
[14:41:45] [PASSED] oob_was/18020744125
[14:41:45] [PASSED] oob_was/1409600907
[14:41:45] [PASSED] oob_was/22014953428
[14:41:45] [PASSED] oob_was/16017236439
[14:41:45] [PASSED] oob_was/14019821291
[14:41:45] [PASSED] oob_was/14015076503
[14:41:45] [PASSED] oob_was/22016122933
[14:41:45] [PASSED] oob_was/14018913170
[14:41:45] [PASSED] oob_was/14018094691
[14:41:45] [PASSED] oob_was/18024947630
[14:41:45] [PASSED] oob_was/16022287689
[14:41:45] [PASSED] oob_was/13011645652
[14:41:45] [PASSED] oob_was/14022293748
[14:41:45] [PASSED] oob_was/22019794406
[14:41:45] [PASSED] oob_was/22019338487
[14:41:45] [PASSED] oob_was/16023588340
[14:41:45] [PASSED] oob_was/14019789679
[14:41:45] [PASSED] oob_was/14022866841
[14:41:45] [PASSED] oob_was/16021333562
[14:41:45] [PASSED] oob_was/14016712196
[14:41:45] [PASSED] oob_was/14015568240
[14:41:45] [PASSED] oob_was/18013179988
[14:41:45] [PASSED] oob_was/1508761755
[14:41:45] [PASSED] oob_was/16023105232
[14:41:45] [PASSED] oob_was/16026508708
[14:41:45] [PASSED] oob_was/14020001231
[14:41:45] [PASSED] oob_was/16023683509
[14:41:45] [PASSED] oob_was/14025515070
[14:41:45] [PASSED] oob_was/15015404425_disable
[14:41:45] [PASSED] oob_was/16026007364
[14:41:45] [PASSED] oob_was/14020316580
[14:41:45] [PASSED] oob_was/14025883347
[14:41:45] [PASSED] oob_was/16029380221
[14:41:45] [PASSED] oob_was/22022079272
[14:41:45] [PASSED] oob_was/16029897822
[14:41:45] [PASSED] oob_was/14027054324
[14:41:45] [PASSED] oob_was/14025941587
[14:41:45] ============== [PASSED] xe_rtp_table_oob_test ==============
[14:41:45] ================ xe_rtp_table_dev_oob_test ================
[14:41:45] [PASSED] device_oob_was/22010954014
[14:41:45] [PASSED] device_oob_was/15015404425
[14:41:45] [PASSED] device_oob_was/22019338487_display
[14:41:45] [PASSED] device_oob_was/14022085890
[14:41:45] [PASSED] device_oob_was/14026539277
[14:41:45] [PASSED] device_oob_was/14026633728
[14:41:45] [PASSED] device_oob_was/14026746987
[14:41:45] [PASSED] device_oob_was/14026779378
[14:41:45] ============ [PASSED] xe_rtp_table_dev_oob_test ============
[14:41:45] ========== xe_rtp_table_missing_upper_bound_test ==========
[14:41:45] [PASSED] register_whitelist/WaAllowPMDepthAndInvocationCountAccessFromUMD, 1408556865
[14:41:45] [PASSED] register_whitelist/1508744258, 14012131227, 1808121037
[14:41:45] [PASSED] register_whitelist/1806527549
[14:41:45] [PASSED] register_whitelist/allow_read_ctx_timestamp
[14:41:45] [PASSED] register_whitelist/allow_read_queue_timestamp
[14:41:45] [PASSED] register_whitelist/16014440446
[14:41:45] [PASSED] register_whitelist/16017236439
[14:41:45] [PASSED] register_whitelist/16020183090
[14:41:45] [PASSED] register_whitelist/14024997852
[14:41:45] [PASSED] register_whitelist/14024997852
[14:41:45] ====== [PASSED] xe_rtp_table_missing_upper_bound_test ======
[14:41:45] =============== [PASSED] xe_rtp_tables_test ================
[14:41:45] =================== xe_rtp (3 subtests) ====================
[14:41:45] =================== xe_rtp_rules_tests ====================
[14:41:45] [PASSED] no
[14:41:45] [PASSED] yes
[14:41:45] [PASSED] no-and-no
[14:41:45] [PASSED] no-and-yes
[14:41:45] [PASSED] yes-and-no
[14:41:45] [PASSED] yes-and-yes
[14:41:45] [PASSED] no-or-no
[14:41:45] [PASSED] no-or-yes
[14:41:45] [PASSED] yes-or-no
[14:41:45] [PASSED] yes-or-yes
[14:41:45] [PASSED] no-yes-or-yes-no
[14:41:45] [PASSED] no-yes-or-yes-yes
[14:41:45] [PASSED] yes-yes-or-no-yes
[14:41:45] [PASSED] yes-yes-or-yes-yes
[14:41:45] [PASSED] no-no-or-yes-or-no
[14:41:45] [PASSED] or
[14:41:45] [PASSED] or-yes
[14:41:45] [PASSED] or-no
[14:41:45] [PASSED] yes-or
[14:41:45] [PASSED] no-or
[14:41:45] [PASSED] no-or-or-yes
[14:41:45] [PASSED] yes-or-or-no
[14:41:45] [PASSED] no-or-or-no
[14:41:45] [PASSED] missing-context-engine-class
[14:41:45] [PASSED] missing-context-engine-class-or-yes
[14:41:45] [PASSED] missing-context-engine-class-or-or-yes
[14:41:45] =============== [PASSED] xe_rtp_rules_tests ================
[14:41:45] =============== xe_rtp_process_to_sr_tests ================
[14:41:45] [PASSED] coalesce-same-reg
[14:41:45] [PASSED] coalesce-same-reg-literal-and-func
[14:41:45] [PASSED] no-match-no-add
[14:41:45] [PASSED] two-regs-two-entries
[14:41:45] [PASSED] clr-one-set-other
[14:41:45] [PASSED] set-field
[14:41:45] [PASSED] conflict-duplicate
[14:41:45] [PASSED] conflict-not-disjoint
[14:41:45] [PASSED] conflict-not-disjoint-literal-and-func
[14:41:45] [PASSED] conflict-reg-type
[14:41:45] [PASSED] bad-mcr-reg-forced-to-regular
[14:41:45] [PASSED] bad-regular-reg-forced-to-mcr
[14:41:45] =========== [PASSED] xe_rtp_process_to_sr_tests ============
[14:41:45] ================== xe_rtp_process_tests ===================
[14:41:45] [PASSED] active1
[14:41:45] [PASSED] active2
[14:41:45] [PASSED] active-inactive
[14:41:45] [PASSED] inactive-active
[14:41:45] [PASSED] inactive-active-inactive
[14:41:45] [PASSED] inactive-inactive-inactive
[14:41:45] ============== [PASSED] xe_rtp_process_tests ===============
[14:41:45] ===================== [PASSED] xe_rtp ======================
[14:41:45] ==================== xe_wa (1 subtest) =====================
[14:41:45] ======================== xe_wa_gt =========================
[14:41:45] [PASSED] TIGERLAKE B0
[14:41:45] [PASSED] DG1 A0
[14:41:45] [PASSED] DG1 B0
[14:41:45] [PASSED] ALDERLAKE_S A0
[14:41:45] [PASSED] ALDERLAKE_S B0
[14:41:45] [PASSED] ALDERLAKE_S C0
[14:41:45] [PASSED] ALDERLAKE_S D0
[14:41:45] [PASSED] ALDERLAKE_P A0
[14:41:45] [PASSED] ALDERLAKE_P B0
[14:41:45] [PASSED] ALDERLAKE_P C0
[14:41:45] [PASSED] ALDERLAKE_S RPLS D0
[14:41:45] [PASSED] ALDERLAKE_P RPLU E0
[14:41:45] [PASSED] DG2 G10 C0
[14:41:45] [PASSED] DG2 G11 B1
[14:41:45] [PASSED] DG2 G12 A1
[14:41:45] [PASSED] METEORLAKE 12.70(Xe_LPG) A0 13.00(Xe_LPM+) A0
[14:41:45] [PASSED] METEORLAKE 12.71(Xe_LPG) A0 13.00(Xe_LPM+) A0
[14:41:45] [PASSED] METEORLAKE 12.74(Xe_LPG+) A0 13.00(Xe_LPM+) A0
[14:41:45] [PASSED] LUNARLAKE 20.04(Xe2_LPG) A0 20.00(Xe2_LPM) A0
[14:41:45] [PASSED] LUNARLAKE 20.04(Xe2_LPG) B0 20.00(Xe2_LPM) A0
[14:41:45] [PASSED] BATTLEMAGE 20.01(Xe2_HPG) A0 13.01(Xe2_HPM) A1
[14:41:45] [PASSED] PANTHERLAKE 30.00(Xe3_LPG) A0 30.00(Xe3_LPM) A0
[14:41:45] ==================== [PASSED] xe_wa_gt =====================
[14:41:45] ====================== [PASSED] xe_wa ======================
[14:41:45] ============================================================
[14:41:45] Testing complete. Ran 822 tests: passed: 794, skipped: 28
[14:41:45] Elapsed time: 36.039s total, 1.814s configuring, 33.510s building, 0.707s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/tests/.kunitconfig
[14:41:45] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[14:41:47] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[14:42:12] Starting KUnit Kernel (1/1)...
[14:42:12] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[14:42:12] ============= refcount_interrupt (4 subtests) ==============
[14:42:12] [PASSED] test_single_irq_change
[14:42:12] [PASSED] test_nested_irq_change
[14:42:12] [PASSED] test_multiple_irq_change
[14:42:12] [PASSED] test_irq_save
[14:42:12] =============== [PASSED] refcount_interrupt ================
[14:42:12] ============ drm_test_pick_cmdline (2 subtests) ============
[14:42:12] [PASSED] drm_test_pick_cmdline_res_1920_1080_60
[14:42:12] =============== drm_test_pick_cmdline_named ===============
[14:42:12] [PASSED] NTSC
[14:42:12] [PASSED] NTSC-J
[14:42:12] [PASSED] PAL
[14:42:12] [PASSED] PAL-M
[14:42:12] =========== [PASSED] drm_test_pick_cmdline_named ===========
[14:42:12] ============== [PASSED] drm_test_pick_cmdline ==============
[14:42:12] == drm_test_atomic_get_connector_for_encoder (1 subtest) ===
[14:42:12] [PASSED] drm_test_drm_atomic_get_connector_for_encoder
[14:42:12] ==== [PASSED] drm_test_atomic_get_connector_for_encoder ====
[14:42:12] =========== drm_validate_clone_mode (2 subtests) ===========
[14:42:12] ============== drm_test_check_in_clone_mode ===============
[14:42:12] [PASSED] in_clone_mode
[14:42:12] [PASSED] not_in_clone_mode
[14:42:12] ========== [PASSED] drm_test_check_in_clone_mode ===========
[14:42:12] =============== drm_test_check_valid_clones ===============
[14:42:12] [PASSED] not_in_clone_mode
[14:42:12] [PASSED] valid_clone
[14:42:12] [PASSED] invalid_clone
[14:42:12] =========== [PASSED] drm_test_check_valid_clones ===========
[14:42:12] ============= [PASSED] drm_validate_clone_mode =============
[14:42:12] ============= drm_validate_modeset (1 subtest) =============
[14:42:12] [PASSED] drm_test_check_connector_changed_modeset
[14:42:12] ============== [PASSED] drm_validate_modeset ===============
[14:42:12] ====== drm_test_bridge_get_current_state (1 subtest) =======
[14:42:12] [PASSED] drm_test_drm_bridge_get_current_state_atomic
[14:42:12] ======== [PASSED] drm_test_bridge_get_current_state ========
[14:42:12] ====== drm_test_bridge_helper_reset_crtc (3 subtests) ======
[14:42:12] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic
[14:42:12] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic_disabled
[14:42:12] [PASSED] drm_test_drm_bridge_helper_hdmi_output_bus_fmts
[14:42:12] ======== [PASSED] drm_test_bridge_helper_reset_crtc ========
[14:42:12] ============== drm_bridge_alloc (2 subtests) ===============
[14:42:12] [PASSED] drm_test_drm_bridge_alloc_basic
[14:42:12] [PASSED] drm_test_drm_bridge_alloc_get_put
[14:42:12] ================ [PASSED] drm_bridge_alloc =================
[14:42:12] ============= drm_bridge_bus_fmt (5 subtests) ==============
[14:42:12] [PASSED] drm_test_bridge_rgb_yuv_rgb
[14:42:12] [PASSED] drm_test_bridge_must_convert_to_yuv444
[14:42:12] [PASSED] drm_test_bridge_hdmi_auto_rgb
[14:42:12] [PASSED] drm_test_bridge_auto_first
[14:42:12] [PASSED] drm_test_bridge_rgb_yuv_no_path
[14:42:12] =============== [PASSED] drm_bridge_bus_fmt ================
[14:42:12] ============= drm_cmdline_parser (40 subtests) =============
[14:42:12] [PASSED] drm_test_cmdline_force_d_only
[14:42:12] [PASSED] drm_test_cmdline_force_D_only_dvi
[14:42:12] [PASSED] drm_test_cmdline_force_D_only_hdmi
[14:42:12] [PASSED] drm_test_cmdline_force_D_only_not_digital
[14:42:12] [PASSED] drm_test_cmdline_force_e_only
[14:42:12] [PASSED] drm_test_cmdline_res
[14:42:12] [PASSED] drm_test_cmdline_res_vesa
[14:42:12] [PASSED] drm_test_cmdline_res_vesa_rblank
[14:42:12] [PASSED] drm_test_cmdline_res_rblank
[14:42:12] [PASSED] drm_test_cmdline_res_bpp
[14:42:12] [PASSED] drm_test_cmdline_res_refresh
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_margins
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_force_off
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_analog
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_digital
[14:42:12] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced_margins_force_on
[14:42:12] [PASSED] drm_test_cmdline_res_margins_force_on
[14:42:12] [PASSED] drm_test_cmdline_res_vesa_margins
[14:42:12] [PASSED] drm_test_cmdline_name
[14:42:12] [PASSED] drm_test_cmdline_name_bpp
[14:42:12] [PASSED] drm_test_cmdline_name_option
[14:42:12] [PASSED] drm_test_cmdline_name_bpp_option
[14:42:12] [PASSED] drm_test_cmdline_rotate_0
[14:42:12] [PASSED] drm_test_cmdline_rotate_90
[14:42:12] [PASSED] drm_test_cmdline_rotate_180
[14:42:12] [PASSED] drm_test_cmdline_rotate_270
[14:42:12] [PASSED] drm_test_cmdline_hmirror
[14:42:12] [PASSED] drm_test_cmdline_vmirror
[14:42:12] [PASSED] drm_test_cmdline_margin_options
[14:42:12] [PASSED] drm_test_cmdline_multiple_options
[14:42:12] [PASSED] drm_test_cmdline_bpp_extra_and_option
[14:42:12] [PASSED] drm_test_cmdline_extra_and_option
[14:42:12] [PASSED] drm_test_cmdline_freestanding_options
[14:42:12] [PASSED] drm_test_cmdline_freestanding_force_e_and_options
[14:42:12] [PASSED] drm_test_cmdline_panel_orientation
[14:42:12] ================ drm_test_cmdline_invalid =================
[14:42:12] [PASSED] margin_only
[14:42:12] [PASSED] interlace_only
[14:42:12] [PASSED] res_missing_x
[14:42:12] [PASSED] res_missing_y
[14:42:12] [PASSED] res_bad_y
[14:42:12] [PASSED] res_missing_y_bpp
[14:42:12] [PASSED] res_bad_bpp
[14:42:12] [PASSED] res_bad_refresh
[14:42:12] [PASSED] res_bpp_refresh_force_on_off
[14:42:12] [PASSED] res_invalid_mode
[14:42:12] [PASSED] res_bpp_wrong_place_mode
[14:42:12] [PASSED] name_bpp_refresh
[14:42:12] [PASSED] name_refresh
[14:42:12] [PASSED] name_refresh_wrong_mode
[14:42:12] [PASSED] name_refresh_invalid_mode
[14:42:12] [PASSED] rotate_multiple
[14:42:12] [PASSED] rotate_invalid_val
[14:42:12] [PASSED] rotate_truncated
[14:42:12] [PASSED] invalid_option
[14:42:12] [PASSED] invalid_tv_option
[14:42:12] [PASSED] truncated_tv_option
[14:42:12] ============ [PASSED] drm_test_cmdline_invalid =============
[14:42:12] =============== drm_test_cmdline_tv_options ===============
[14:42:12] [PASSED] NTSC
[14:42:12] [PASSED] NTSC_443
[14:42:12] [PASSED] NTSC_J
[14:42:12] [PASSED] PAL
[14:42:12] [PASSED] PAL_M
[14:42:12] [PASSED] PAL_N
[14:42:12] [PASSED] SECAM
[14:42:12] [PASSED] MONO_525
[14:42:12] [PASSED] MONO_625
[14:42:12] =========== [PASSED] drm_test_cmdline_tv_options ===========
[14:42:12] =============== [PASSED] drm_cmdline_parser ================
[14:42:12] ========== drmm_connector_hdmi_init (20 subtests) ==========
[14:42:12] [PASSED] drm_test_connector_hdmi_init_valid
[14:42:12] [PASSED] drm_test_connector_hdmi_init_bpc_8
[14:42:12] [PASSED] drm_test_connector_hdmi_init_bpc_10
[14:42:12] [PASSED] drm_test_connector_hdmi_init_bpc_12
[14:42:12] [PASSED] drm_test_connector_hdmi_init_bpc_invalid
[14:42:12] [PASSED] drm_test_connector_hdmi_init_bpc_null
[14:42:12] [PASSED] drm_test_connector_hdmi_init_formats_empty
[14:42:12] [PASSED] drm_test_connector_hdmi_init_formats_no_rgb
[14:42:12] === drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[14:42:12] [PASSED] supported_formats=0x9 yuv420_allowed=1
[14:42:12] [PASSED] supported_formats=0x9 yuv420_allowed=0
[14:42:12] [PASSED] supported_formats=0x5 yuv420_allowed=1
[14:42:12] [PASSED] supported_formats=0x5 yuv420_allowed=0
[14:42:12] === [PASSED] drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[14:42:12] [PASSED] drm_test_connector_hdmi_init_null_ddc
[14:42:12] [PASSED] drm_test_connector_hdmi_init_null_product
[14:42:12] [PASSED] drm_test_connector_hdmi_init_null_vendor
[14:42:12] [PASSED] drm_test_connector_hdmi_init_product_length_exact
[14:42:12] [PASSED] drm_test_connector_hdmi_init_product_length_too_long
[14:42:12] [PASSED] drm_test_connector_hdmi_init_product_valid
[14:42:12] [PASSED] drm_test_connector_hdmi_init_vendor_length_exact
[14:42:12] [PASSED] drm_test_connector_hdmi_init_vendor_length_too_long
[14:42:12] [PASSED] drm_test_connector_hdmi_init_vendor_valid
[14:42:12] ========= drm_test_connector_hdmi_init_type_valid =========
[14:42:12] [PASSED] HDMI-A
[14:42:12] [PASSED] HDMI-B
[14:42:12] ===== [PASSED] drm_test_connector_hdmi_init_type_valid =====
[14:42:12] ======== drm_test_connector_hdmi_init_type_invalid ========
[14:42:12] [PASSED] Unknown
[14:42:12] [PASSED] VGA
[14:42:12] [PASSED] DVI-I
[14:42:12] [PASSED] DVI-D
[14:42:12] [PASSED] DVI-A
[14:42:12] [PASSED] Composite
[14:42:12] [PASSED] SVIDEO
[14:42:12] [PASSED] LVDS
[14:42:12] [PASSED] Component
[14:42:12] [PASSED] DIN
[14:42:12] [PASSED] DP
[14:42:12] [PASSED] TV
[14:42:12] [PASSED] eDP
[14:42:12] [PASSED] Virtual
[14:42:12] [PASSED] DSI
[14:42:12] [PASSED] DPI
[14:42:12] [PASSED] Writeback
[14:42:12] [PASSED] SPI
[14:42:12] [PASSED] USB
[14:42:12] ==== [PASSED] drm_test_connector_hdmi_init_type_invalid ====
[14:42:12] ============ [PASSED] drmm_connector_hdmi_init =============
[14:42:12] ============= drmm_connector_init (3 subtests) =============
[14:42:12] [PASSED] drm_test_drmm_connector_init
[14:42:12] [PASSED] drm_test_drmm_connector_init_null_ddc
[14:42:12] ========= drm_test_drmm_connector_init_type_valid =========
[14:42:12] [PASSED] Unknown
[14:42:12] [PASSED] VGA
[14:42:12] [PASSED] DVI-I
[14:42:12] [PASSED] DVI-D
[14:42:12] [PASSED] DVI-A
[14:42:12] [PASSED] Composite
[14:42:12] [PASSED] SVIDEO
[14:42:12] [PASSED] LVDS
[14:42:12] [PASSED] Component
[14:42:12] [PASSED] DIN
[14:42:12] [PASSED] DP
[14:42:12] [PASSED] HDMI-A
[14:42:12] [PASSED] HDMI-B
[14:42:12] [PASSED] TV
[14:42:12] [PASSED] eDP
[14:42:12] [PASSED] Virtual
[14:42:12] [PASSED] DSI
[14:42:12] [PASSED] DPI
[14:42:12] [PASSED] Writeback
[14:42:12] [PASSED] SPI
[14:42:12] [PASSED] USB
[14:42:12] ===== [PASSED] drm_test_drmm_connector_init_type_valid =====
[14:42:12] =============== [PASSED] drmm_connector_init ===============
[14:42:12] ========= drm_connector_dynamic_init (6 subtests) ==========
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_init
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_init_null_ddc
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_init_not_added
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_init_properties
[14:42:12] ===== drm_test_drm_connector_dynamic_init_type_valid ======
[14:42:12] [PASSED] Unknown
[14:42:12] [PASSED] VGA
[14:42:12] [PASSED] DVI-I
[14:42:12] [PASSED] DVI-D
[14:42:12] [PASSED] DVI-A
[14:42:12] [PASSED] Composite
[14:42:12] [PASSED] SVIDEO
[14:42:12] [PASSED] LVDS
[14:42:12] [PASSED] Component
[14:42:12] [PASSED] DIN
[14:42:12] [PASSED] DP
[14:42:12] [PASSED] HDMI-A
[14:42:12] [PASSED] HDMI-B
[14:42:12] [PASSED] TV
[14:42:12] [PASSED] eDP
[14:42:12] [PASSED] Virtual
[14:42:12] [PASSED] DSI
[14:42:12] [PASSED] DPI
[14:42:12] [PASSED] Writeback
[14:42:12] [PASSED] SPI
[14:42:12] [PASSED] USB
[14:42:12] = [PASSED] drm_test_drm_connector_dynamic_init_type_valid ==
[14:42:12] ======== drm_test_drm_connector_dynamic_init_name =========
[14:42:12] [PASSED] Unknown
[14:42:12] [PASSED] VGA
[14:42:12] [PASSED] DVI-I
[14:42:12] [PASSED] DVI-D
[14:42:12] [PASSED] DVI-A
[14:42:12] [PASSED] Composite
[14:42:12] [PASSED] SVIDEO
[14:42:12] [PASSED] LVDS
[14:42:12] [PASSED] Component
[14:42:12] [PASSED] DIN
[14:42:12] [PASSED] DP
[14:42:12] [PASSED] HDMI-A
[14:42:12] [PASSED] HDMI-B
[14:42:12] [PASSED] TV
[14:42:12] [PASSED] eDP
[14:42:12] [PASSED] Virtual
[14:42:12] [PASSED] DSI
[14:42:12] [PASSED] DPI
[14:42:12] [PASSED] Writeback
[14:42:12] [PASSED] SPI
[14:42:12] [PASSED] USB
[14:42:12] ==== [PASSED] drm_test_drm_connector_dynamic_init_name =====
[14:42:12] =========== [PASSED] drm_connector_dynamic_init ============
[14:42:12] ==== drm_connector_dynamic_register_early (4 subtests) =====
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_early_on_list
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_early_defer
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_early_no_init
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_early_no_mode_object
[14:42:12] ====== [PASSED] drm_connector_dynamic_register_early =======
[14:42:12] ======= drm_connector_dynamic_register (7 subtests) ========
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_on_list
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_no_defer
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_no_init
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_mode_object
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_sysfs
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_sysfs_name
[14:42:12] [PASSED] drm_test_drm_connector_dynamic_register_debugfs
[14:42:12] ========= [PASSED] drm_connector_dynamic_register ==========
[14:42:12] = drm_connector_attach_broadcast_rgb_property (2 subtests) =
[14:42:12] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property
[14:42:12] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property_hdmi_connector
[14:42:12] === [PASSED] drm_connector_attach_broadcast_rgb_property ===
[14:42:12] ========== drm_get_tv_mode_from_name (2 subtests) ==========
[14:42:12] ========== drm_test_get_tv_mode_from_name_valid ===========
[14:42:12] [PASSED] NTSC
[14:42:12] [PASSED] NTSC-443
[14:42:12] [PASSED] NTSC-J
[14:42:12] [PASSED] PAL
[14:42:12] [PASSED] PAL-M
[14:42:12] [PASSED] PAL-N
[14:42:12] [PASSED] SECAM
[14:42:12] [PASSED] Mono
[14:42:12] ====== [PASSED] drm_test_get_tv_mode_from_name_valid =======
[14:42:12] [PASSED] drm_test_get_tv_mode_from_name_truncated
[14:42:12] ============ [PASSED] drm_get_tv_mode_from_name ============
[14:42:12] = drm_test_connector_hdmi_compute_mode_clock (12 subtests) =
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc_vic_1
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc_vic_1
[14:42:12] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_double
[14:42:12] = drm_test_connector_hdmi_compute_mode_clock_yuv420_valid =
[14:42:12] [PASSED] VIC 96
[14:42:12] [PASSED] VIC 97
[14:42:12] [PASSED] VIC 101
[14:42:12] [PASSED] VIC 102
[14:42:12] [PASSED] VIC 106
[14:42:12] [PASSED] VIC 107
[14:42:12] === [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_valid ===
[14:42:12] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_10_bpc
[14:42:12] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_12_bpc
[14:42:12] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_8_bpc
[14:42:12] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_10_bpc
[14:42:12] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_12_bpc
[14:42:12] === [PASSED] drm_test_connector_hdmi_compute_mode_clock ====
[14:42:12] == drm_hdmi_connector_get_broadcast_rgb_name (2 subtests) ==
[14:42:12] === drm_test_drm_hdmi_connector_get_broadcast_rgb_name ====
[14:42:12] [PASSED] Automatic
[14:42:12] [PASSED] Full
[14:42:12] [PASSED] Limited 16:235
[14:42:12] === [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name ===
[14:42:12] [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name_invalid
[14:42:12] ==== [PASSED] drm_hdmi_connector_get_broadcast_rgb_name ====
[14:42:12] == drm_hdmi_connector_get_output_format_name (2 subtests) ==
[14:42:12] === drm_test_drm_hdmi_connector_get_output_format_name ====
[14:42:12] [PASSED] RGB
[14:42:12] [PASSED] YUV 4:2:0
[14:42:12] [PASSED] YUV 4:2:2
[14:42:12] [PASSED] YUV 4:4:4
[14:42:12] === [PASSED] drm_test_drm_hdmi_connector_get_output_format_name ===
[14:42:12] [PASSED] drm_test_drm_hdmi_connector_get_output_format_name_invalid
[14:42:12] ==== [PASSED] drm_hdmi_connector_get_output_format_name ====
[14:42:12] ============= drm_damage_helper (21 subtests) ==============
[14:42:12] [PASSED] drm_test_damage_iter_no_damage
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_fractional_src
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_src_moved
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_fractional_src_moved
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_not_visible
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_no_crtc
[14:42:12] [PASSED] drm_test_damage_iter_no_damage_no_fb
[14:42:12] [PASSED] drm_test_damage_iter_simple_damage
[14:42:12] [PASSED] drm_test_damage_iter_single_damage
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_intersect_src
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_outside_src
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_fractional_src
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_intersect_fractional_src
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_outside_fractional_src
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_src_moved
[14:42:12] [PASSED] drm_test_damage_iter_single_damage_fractional_src_moved
[14:42:12] [PASSED] drm_test_damage_iter_damage
[14:42:12] [PASSED] drm_test_damage_iter_damage_one_intersect
[14:42:12] [PASSED] drm_test_damage_iter_damage_one_outside
[14:42:12] [PASSED] drm_test_damage_iter_damage_src_moved
[14:42:12] [PASSED] drm_test_damage_iter_damage_not_visible
[14:42:12] ================ [PASSED] drm_damage_helper ================
[14:42:12] ============== drm_dp_mst_helper (3 subtests) ==============
[14:42:12] ============== drm_test_dp_mst_calc_pbn_mode ==============
[14:42:12] [PASSED] Clock 154000 BPP 30 DSC disabled
[14:42:12] [PASSED] Clock 234000 BPP 30 DSC disabled
[14:42:12] [PASSED] Clock 297000 BPP 24 DSC disabled
[14:42:12] [PASSED] Clock 332880 BPP 24 DSC enabled
[14:42:12] [PASSED] Clock 324540 BPP 24 DSC enabled
[14:42:12] ========== [PASSED] drm_test_dp_mst_calc_pbn_mode ==========
[14:42:12] ============== drm_test_dp_mst_calc_pbn_div ===============
[14:42:12] [PASSED] Link rate 2000000 lane count 4
[14:42:12] [PASSED] Link rate 2000000 lane count 2
[14:42:12] [PASSED] Link rate 2000000 lane count 1
[14:42:12] [PASSED] Link rate 1350000 lane count 4
[14:42:12] [PASSED] Link rate 1350000 lane count 2
[14:42:12] [PASSED] Link rate 1350000 lane count 1
[14:42:12] [PASSED] Link rate 1000000 lane count 4
[14:42:12] [PASSED] Link rate 1000000 lane count 2
[14:42:12] [PASSED] Link rate 1000000 lane count 1
[14:42:12] [PASSED] Link rate 810000 lane count 4
[14:42:12] [PASSED] Link rate 810000 lane count 2
[14:42:12] [PASSED] Link rate 810000 lane count 1
[14:42:12] [PASSED] Link rate 540000 lane count 4
[14:42:12] [PASSED] Link rate 540000 lane count 2
[14:42:12] [PASSED] Link rate 540000 lane count 1
[14:42:12] [PASSED] Link rate 270000 lane count 4
[14:42:12] [PASSED] Link rate 270000 lane count 2
[14:42:12] [PASSED] Link rate 270000 lane count 1
[14:42:12] [PASSED] Link rate 162000 lane count 4
[14:42:12] [PASSED] Link rate 162000 lane count 2
[14:42:12] [PASSED] Link rate 162000 lane count 1
[14:42:12] ========== [PASSED] drm_test_dp_mst_calc_pbn_div ===========
[14:42:12] ========= drm_test_dp_mst_sideband_msg_req_decode =========
[14:42:12] [PASSED] DP_ENUM_PATH_RESOURCES with port number
[14:42:12] [PASSED] DP_POWER_UP_PHY with port number
[14:42:12] [PASSED] DP_POWER_DOWN_PHY with port number
[14:42:12] [PASSED] DP_ALLOCATE_PAYLOAD with SDP stream sinks
[14:42:12] [PASSED] DP_ALLOCATE_PAYLOAD with port number
[14:42:12] [PASSED] DP_ALLOCATE_PAYLOAD with VCPI
[14:42:12] [PASSED] DP_ALLOCATE_PAYLOAD with PBN
[14:42:12] [PASSED] DP_QUERY_PAYLOAD with port number
[14:42:12] [PASSED] DP_QUERY_PAYLOAD with VCPI
[14:42:12] [PASSED] DP_REMOTE_DPCD_READ with port number
[14:42:12] [PASSED] DP_REMOTE_DPCD_READ with DPCD address
[14:42:12] [PASSED] DP_REMOTE_DPCD_READ with max number of bytes
[14:42:12] [PASSED] DP_REMOTE_DPCD_WRITE with port number
[14:42:12] [PASSED] DP_REMOTE_DPCD_WRITE with DPCD address
[14:42:12] [PASSED] DP_REMOTE_DPCD_WRITE with data array
[14:42:12] [PASSED] DP_REMOTE_I2C_READ with port number
[14:42:12] [PASSED] DP_REMOTE_I2C_READ with I2C device ID
[14:42:12] [PASSED] DP_REMOTE_I2C_READ with transactions array
[14:42:12] [PASSED] DP_REMOTE_I2C_WRITE with port number
[14:42:12] [PASSED] DP_REMOTE_I2C_WRITE with I2C device ID
[14:42:12] [PASSED] DP_REMOTE_I2C_WRITE with data array
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream ID
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with client ID
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream event
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with valid stream event
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream behavior
[14:42:12] [PASSED] DP_QUERY_STREAM_ENC_STATUS with a valid stream behavior
[14:42:12] ===== [PASSED] drm_test_dp_mst_sideband_msg_req_decode =====
[14:42:12] ================ [PASSED] drm_dp_mst_helper ================
[14:42:12] ================== drm_exec (7 subtests) ===================
[14:42:12] [PASSED] sanitycheck
[14:42:12] [PASSED] test_lock
[14:42:12] [PASSED] test_lock_unlock
[14:42:12] [PASSED] test_duplicates
[14:42:12] [PASSED] test_prepare
[14:42:12] [PASSED] test_prepare_array
[14:42:12] [PASSED] test_multiple_loops
[14:42:12] ==================== [PASSED] drm_exec =====================
[14:42:12] =========== drm_format_helper_test (17 subtests) ===========
[14:42:12] ============== drm_test_fb_xrgb8888_to_gray8 ==============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========== [PASSED] drm_test_fb_xrgb8888_to_gray8 ==========
[14:42:12] ============= drm_test_fb_xrgb8888_to_rgb332 ==============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb332 ==========
[14:42:12] ============= drm_test_fb_xrgb8888_to_rgb565 ==============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb565 ==========
[14:42:12] ============ drm_test_fb_xrgb8888_to_xrgb1555 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_xrgb1555 =========
[14:42:12] ============ drm_test_fb_xrgb8888_to_argb1555 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_argb1555 =========
[14:42:12] ============ drm_test_fb_xrgb8888_to_rgba5551 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_rgba5551 =========
[14:42:12] ============= drm_test_fb_xrgb8888_to_rgb888 ==============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb888 ==========
[14:42:12] ============= drm_test_fb_xrgb8888_to_bgr888 ==============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========= [PASSED] drm_test_fb_xrgb8888_to_bgr888 ==========
[14:42:12] ============ drm_test_fb_xrgb8888_to_argb8888 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_argb8888 =========
[14:42:12] =========== drm_test_fb_xrgb8888_to_xrgb2101010 ===========
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======= [PASSED] drm_test_fb_xrgb8888_to_xrgb2101010 =======
[14:42:12] =========== drm_test_fb_xrgb8888_to_argb2101010 ===========
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======= [PASSED] drm_test_fb_xrgb8888_to_argb2101010 =======
[14:42:12] ============== drm_test_fb_xrgb8888_to_mono ===============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ========== [PASSED] drm_test_fb_xrgb8888_to_mono ===========
[14:42:12] ==================== drm_test_fb_swab =====================
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ================ [PASSED] drm_test_fb_swab =================
[14:42:12] ============ drm_test_fb_xrgb8888_to_xbgr8888 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_xbgr8888 =========
[14:42:12] ============ drm_test_fb_xrgb8888_to_abgr8888 =============
[14:42:12] [PASSED] single_pixel_source_buffer
[14:42:12] [PASSED] single_pixel_clip_rectangle
[14:42:12] [PASSED] well_known_colors
[14:42:12] [PASSED] destination_pitch
[14:42:12] ======== [PASSED] drm_test_fb_xrgb8888_to_abgr8888 =========
[14:42:12] ================= drm_test_fb_clip_offset =================
[14:42:12] [PASSED] pass through
[14:42:12] [PASSED] horizontal offset
[14:42:12] [PASSED] vertical offset
[14:42:12] [PASSED] horizontal and vertical offset
[14:42:12] [PASSED] horizontal offset (custom pitch)
[14:42:12] [PASSED] vertical offset (custom pitch)
[14:42:12] [PASSED] horizontal and vertical offset (custom pitch)
[14:42:12] ============= [PASSED] drm_test_fb_clip_offset =============
[14:42:12] =================== drm_test_fb_memcpy ====================
[14:42:12] [PASSED] single_pixel_source_buffer: XR24 little-endian (0x34325258)
[14:42:12] [PASSED] single_pixel_source_buffer: XRA8 little-endian (0x38415258)
[14:42:12] [PASSED] single_pixel_source_buffer: YU24 little-endian (0x34325559)
[14:42:12] [PASSED] single_pixel_clip_rectangle: XB24 little-endian (0x34324258)
[14:42:12] [PASSED] single_pixel_clip_rectangle: XRA8 little-endian (0x38415258)
[14:42:12] [PASSED] single_pixel_clip_rectangle: YU24 little-endian (0x34325559)
[14:42:12] [PASSED] well_known_colors: XB24 little-endian (0x34324258)
[14:42:12] [PASSED] well_known_colors: XRA8 little-endian (0x38415258)
[14:42:12] [PASSED] well_known_colors: YU24 little-endian (0x34325559)
[14:42:12] [PASSED] destination_pitch: XB24 little-endian (0x34324258)
[14:42:12] [PASSED] destination_pitch: XRA8 little-endian (0x38415258)
[14:42:12] [PASSED] destination_pitch: YU24 little-endian (0x34325559)
[14:42:12] =============== [PASSED] drm_test_fb_memcpy ================
[14:42:12] ============= [PASSED] drm_format_helper_test ==============
[14:42:12] ================= drm_format (18 subtests) =================
[14:42:12] [PASSED] drm_test_format_block_width_invalid
[14:42:12] [PASSED] drm_test_format_block_width_one_plane
[14:42:12] [PASSED] drm_test_format_block_width_two_plane
[14:42:12] [PASSED] drm_test_format_block_width_three_plane
[14:42:12] [PASSED] drm_test_format_block_width_tiled
[14:42:12] [PASSED] drm_test_format_block_height_invalid
[14:42:12] [PASSED] drm_test_format_block_height_one_plane
[14:42:12] [PASSED] drm_test_format_block_height_two_plane
[14:42:12] [PASSED] drm_test_format_block_height_three_plane
[14:42:12] [PASSED] drm_test_format_block_height_tiled
[14:42:12] [PASSED] drm_test_format_min_pitch_invalid
[14:42:12] [PASSED] drm_test_format_min_pitch_one_plane_8bpp
[14:42:12] [PASSED] drm_test_format_min_pitch_one_plane_16bpp
[14:42:12] [PASSED] drm_test_format_min_pitch_one_plane_24bpp
[14:42:12] [PASSED] drm_test_format_min_pitch_one_plane_32bpp
[14:42:12] [PASSED] drm_test_format_min_pitch_two_plane
[14:42:12] [PASSED] drm_test_format_min_pitch_three_plane_8bpp
[14:42:12] [PASSED] drm_test_format_min_pitch_tiled
[14:42:12] =================== [PASSED] drm_format ====================
[14:42:12] ============== drm_framebuffer (10 subtests) ===============
[14:42:12] ========== drm_test_framebuffer_check_src_coords ==========
[14:42:12] [PASSED] Success: source fits into fb
[14:42:12] [PASSED] Fail: overflowing fb with x-axis coordinate
[14:42:12] [PASSED] Fail: overflowing fb with y-axis coordinate
[14:42:12] [PASSED] Fail: overflowing fb with source width
[14:42:12] [PASSED] Fail: overflowing fb with source height
[14:42:12] ====== [PASSED] drm_test_framebuffer_check_src_coords ======
[14:42:12] [PASSED] drm_test_framebuffer_cleanup
[14:42:12] =============== drm_test_framebuffer_create ===============
[14:42:12] [PASSED] ABGR8888 normal sizes
[14:42:12] [PASSED] ABGR8888 max sizes
[14:42:12] [PASSED] ABGR8888 pitch greater than min required
[14:42:12] [PASSED] ABGR8888 pitch less than min required
[14:42:12] [PASSED] ABGR8888 Invalid width
[14:42:12] [PASSED] ABGR8888 Invalid buffer handle
[14:42:12] [PASSED] No pixel format
[14:42:12] [PASSED] ABGR8888 Width 0
[14:42:12] [PASSED] ABGR8888 Height 0
[14:42:12] [PASSED] ABGR8888 Out of bound height * pitch combination
[14:42:12] [PASSED] ABGR8888 Large buffer offset
[14:42:12] [PASSED] ABGR8888 Buffer offset for inexistent plane
[14:42:12] [PASSED] ABGR8888 Invalid flag
[14:42:12] [PASSED] ABGR8888 Set DRM_MODE_FB_MODIFIERS without modifiers
[14:42:12] [PASSED] ABGR8888 Valid buffer modifier
[14:42:12] [PASSED] ABGR8888 Invalid buffer modifier(DRM_FORMAT_MOD_SAMSUNG_64_32_TILE)
[14:42:12] [PASSED] ABGR8888 Extra pitches without DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] ABGR8888 Extra pitches with DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] NV12 Normal sizes
[14:42:12] [PASSED] NV12 Max sizes
[14:42:12] [PASSED] NV12 Invalid pitch
[14:42:12] [PASSED] NV12 Invalid modifier/missing DRM_MODE_FB_MODIFIERS flag
[14:42:12] [PASSED] NV12 different modifier per-plane
[14:42:12] [PASSED] NV12 with DRM_FORMAT_MOD_SAMSUNG_64_32_TILE
[14:42:12] [PASSED] NV12 Valid modifiers without DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] NV12 Modifier for inexistent plane
[14:42:12] [PASSED] NV12 Handle for inexistent plane
[14:42:12] [PASSED] NV12 Handle for inexistent plane without DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] YVU420 DRM_MODE_FB_MODIFIERS set without modifier
[14:42:12] [PASSED] YVU420 Normal sizes
[14:42:12] [PASSED] YVU420 Max sizes
[14:42:12] [PASSED] YVU420 Invalid pitch
[14:42:12] [PASSED] YVU420 Different pitches
[14:42:12] [PASSED] YVU420 Different buffer offsets/pitches
[14:42:12] [PASSED] YVU420 Modifier set just for plane 0, without DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] YVU420 Modifier set just for planes 0, 1, without DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] YVU420 Modifier set just for plane 0, 1, with DRM_MODE_FB_MODIFIERS
[14:42:12] [PASSED] YVU420 Valid modifier
[14:42:12] [PASSED] YVU420 Different modifiers per plane
[14:42:12] [PASSED] YVU420 Modifier for inexistent plane
[14:42:12] [PASSED] YUV420_10BIT Invalid modifier(DRM_FORMAT_MOD_LINEAR)
[14:42:12] [PASSED] X0L2 Normal sizes
[14:42:12] [PASSED] X0L2 Max sizes
[14:42:12] [PASSED] X0L2 Invalid pitch
[14:42:12] [PASSED] X0L2 Pitch greater than minimum required
[14:42:12] [PASSED] X0L2 Handle for inexistent plane
[14:42:12] [PASSED] X0L2 Offset for inexistent plane, without DRM_MODE_FB_MODIFIERS set
[14:42:12] [PASSED] X0L2 Modifier without DRM_MODE_FB_MODIFIERS set
[14:42:12] [PASSED] X0L2 Valid modifier
[14:42:12] [PASSED] X0L2 Modifier for inexistent plane
[14:42:12] =========== [PASSED] drm_test_framebuffer_create ===========
[14:42:12] [PASSED] drm_test_framebuffer_free
[14:42:12] [PASSED] drm_test_framebuffer_init
[14:42:12] [PASSED] drm_test_framebuffer_init_bad_format
[14:42:12] [PASSED] drm_test_framebuffer_init_dev_mismatch
[14:42:12] [PASSED] drm_test_framebuffer_lookup
[14:42:12] [PASSED] drm_test_framebuffer_lookup_inexistent
[14:42:12] [PASSED] drm_test_framebuffer_modifiers_not_supported
[14:42:12] ================= [PASSED] drm_framebuffer =================
[14:42:12] ================ drm_gem_shmem (8 subtests) ================
[14:42:12] [PASSED] drm_gem_shmem_test_obj_create
[14:42:12] [PASSED] drm_gem_shmem_test_obj_create_private
[14:42:12] [PASSED] drm_gem_shmem_test_pin_pages
[14:42:12] [PASSED] drm_gem_shmem_test_vmap
[14:42:12] [PASSED] drm_gem_shmem_test_get_sg_table
[14:42:12] [PASSED] drm_gem_shmem_test_get_pages_sgt
[14:42:12] [PASSED] drm_gem_shmem_test_madvise
[14:42:12] [PASSED] drm_gem_shmem_test_purge
[14:42:12] ================== [PASSED] drm_gem_shmem ==================
[14:42:12] === drm_atomic_helper_connector_hdmi_check (29 subtests) ===
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode_vic_1
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode_vic_1
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode_vic_1
[14:42:12] ====== drm_test_check_broadcast_rgb_cea_mode_yuv420 =======
[14:42:12] [PASSED] Automatic
[14:42:12] [PASSED] Full
[14:42:12] [PASSED] Limited 16:235
[14:42:12] == [PASSED] drm_test_check_broadcast_rgb_cea_mode_yuv420 ===
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_changed
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_not_changed
[14:42:12] [PASSED] drm_test_check_disable_connector
[14:42:12] [PASSED] drm_test_check_hdmi_funcs_reject_rate
[14:42:12] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_rgb
[14:42:12] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_yuv420
[14:42:12] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv422
[14:42:12] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv420
[14:42:12] [PASSED] drm_test_check_driver_unsupported_fallback_yuv420
[14:42:12] [PASSED] drm_test_check_output_bpc_crtc_mode_changed
[14:42:12] [PASSED] drm_test_check_output_bpc_crtc_mode_not_changed
[14:42:12] [PASSED] drm_test_check_output_bpc_dvi
[14:42:12] [PASSED] drm_test_check_output_bpc_format_vic_1
[14:42:12] [PASSED] drm_test_check_output_bpc_format_display_8bpc_only
[14:42:12] [PASSED] drm_test_check_output_bpc_format_display_rgb_only
[14:42:12] [PASSED] drm_test_check_output_bpc_format_driver_8bpc_only
[14:42:12] [PASSED] drm_test_check_output_bpc_format_driver_rgb_only
[14:42:12] [PASSED] drm_test_check_tmds_char_rate_rgb_8bpc
[14:42:12] [PASSED] drm_test_check_tmds_char_rate_rgb_10bpc
[14:42:12] [PASSED] drm_test_check_tmds_char_rate_rgb_12bpc
[14:42:12] ============ drm_test_check_hdmi_color_format =============
[14:42:12] [PASSED] AUTO -> RGB
[14:42:12] [PASSED] YCBCR422 -> YUV422
[14:42:12] [PASSED] YCBCR420 -> YUV420
[14:42:12] [PASSED] YCBCR444 -> YUV444
[14:42:12] [PASSED] RGB -> RGB
[14:42:12] ======== [PASSED] drm_test_check_hdmi_color_format =========
[14:42:12] ======== drm_test_check_hdmi_color_format_420_only ========
[14:42:12] [PASSED] RGB should fail
[14:42:12] [PASSED] YUV444 should fail
[14:42:12] [PASSED] YUV422 should fail
[14:42:12] [PASSED] YUV420 should work
[14:42:12] ==== [PASSED] drm_test_check_hdmi_color_format_420_only ====
[14:42:12] ===== [PASSED] drm_atomic_helper_connector_hdmi_check ======
[14:42:12] === drm_atomic_helper_connector_hdmi_reset (6 subtests) ====
[14:42:12] [PASSED] drm_test_check_broadcast_rgb_value
[14:42:12] [PASSED] drm_test_check_bpc_8_value
[14:42:12] [PASSED] drm_test_check_bpc_10_value
[14:42:12] [PASSED] drm_test_check_bpc_12_value
[14:42:12] [PASSED] drm_test_check_format_value
[14:42:12] [PASSED] drm_test_check_tmds_char_value
[14:42:12] ===== [PASSED] drm_atomic_helper_connector_hdmi_reset ======
[14:42:12] = drm_atomic_helper_connector_hdmi_mode_valid (7 subtests) =
[14:42:12] [PASSED] drm_test_check_mode_valid
[14:42:12] [PASSED] drm_test_check_mode_valid_reject
[14:42:12] [PASSED] drm_test_check_mode_valid_reject_rate
[14:42:12] [PASSED] drm_test_check_mode_valid_reject_max_clock
[14:42:12] [PASSED] drm_test_check_mode_valid_yuv420_only_max_clock
[14:42:12] [PASSED] drm_test_check_mode_valid_reject_yuv420_only_connector
[14:42:12] [PASSED] drm_test_check_mode_valid_accept_yuv420_also_connector_rgb
[14:42:12] === [PASSED] drm_atomic_helper_connector_hdmi_mode_valid ===
[14:42:12] = drm_atomic_helper_connector_hdmi_infoframes (5 subtests) =
[14:42:12] [PASSED] drm_test_check_infoframes
[14:42:12] [PASSED] drm_test_check_reject_avi_infoframe
[14:42:12] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_8
[14:42:12] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_10
[14:42:12] [PASSED] drm_test_check_reject_audio_infoframe
[14:42:12] === [PASSED] drm_atomic_helper_connector_hdmi_infoframes ===
[14:42:12] ================= drm_managed (2 subtests) =================
[14:42:12] [PASSED] drm_test_managed_release_action
[14:42:12] [PASSED] drm_test_managed_run_action
[14:42:12] =================== [PASSED] drm_managed ===================
[14:42:12] =================== drm_mm (6 subtests) ====================
[14:42:12] [PASSED] drm_test_mm_init
[14:42:12] [PASSED] drm_test_mm_debug
[14:42:12] [PASSED] drm_test_mm_align32
[14:42:12] [PASSED] drm_test_mm_align64
[14:42:12] [PASSED] drm_test_mm_lowest
[14:42:12] [PASSED] drm_test_mm_highest
[14:42:12] ===================== [PASSED] drm_mm ======================
[14:42:12] ============= drm_modes_analog_tv (5 subtests) =============
[14:42:12] [PASSED] drm_test_modes_analog_tv_mono_576i
[14:42:12] [PASSED] drm_test_modes_analog_tv_ntsc_480i
[14:42:12] [PASSED] drm_test_modes_analog_tv_ntsc_480i_inlined
[14:42:12] [PASSED] drm_test_modes_analog_tv_pal_576i
[14:42:12] [PASSED] drm_test_modes_analog_tv_pal_576i_inlined
[14:42:12] =============== [PASSED] drm_modes_analog_tv ===============
[14:42:12] ============== drm_panic_helper (3 subtests) ===============
[14:42:12] ============= drm_test_panic_screen_user_map ==============
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] ========= [PASSED] drm_test_panic_screen_user_map ==========
[14:42:12] ============= drm_test_panic_screen_user_page =============
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] ========= [PASSED] drm_test_panic_screen_user_page =========
[14:42:12] ========== drm_test_panic_screen_user_set_pixel ===========
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[14:42:12] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[14:42:12] ====== [PASSED] drm_test_panic_screen_user_set_pixel =======
[14:42:12] ================ [PASSED] drm_panic_helper =================
[14:42:12] ============== drm_plane_helper (2 subtests) ===============
[14:42:12] =============== drm_test_check_plane_state ================
[14:42:12] [PASSED] clipping_simple
[14:42:12] [PASSED] clipping_rotate_reflect
[14:42:12] [PASSED] positioning_simple
[14:42:12] [PASSED] upscaling
[14:42:12] [PASSED] downscaling
[14:42:12] [PASSED] rounding1
[14:42:12] [PASSED] rounding2
[14:42:12] [PASSED] rounding3
[14:42:12] [PASSED] rounding4
[14:42:12] =========== [PASSED] drm_test_check_plane_state ============
[14:42:12] =========== drm_test_check_invalid_plane_state ============
[14:42:12] [PASSED] positioning_invalid
[14:42:12] [PASSED] upscaling_invalid
[14:42:12] [PASSED] downscaling_invalid
[14:42:12] ======= [PASSED] drm_test_check_invalid_plane_state ========
[14:42:12] ================ [PASSED] drm_plane_helper =================
[14:42:12] ====== drm_connector_helper_tv_get_modes (1 subtest) =======
[14:42:12] ====== drm_test_connector_helper_tv_get_modes_check =======
[14:42:12] [PASSED] None
[14:42:12] [PASSED] PAL
[14:42:12] [PASSED] NTSC
[14:42:12] [PASSED] Both, NTSC Default
[14:42:12] [PASSED] Both, PAL Default
[14:42:12] [PASSED] Both, NTSC Default, with PAL on command-line
[14:42:12] [PASSED] Both, PAL Default, with NTSC on command-line
[14:42:12] == [PASSED] drm_test_connector_helper_tv_get_modes_check ===
[14:42:12] ======== [PASSED] drm_connector_helper_tv_get_modes ========
[14:42:12] ================== drm_rect (9 subtests) ===================
[14:42:12] [PASSED] drm_test_rect_clip_scaled_div_by_zero
[14:42:12] [PASSED] drm_test_rect_clip_scaled_not_clipped
[14:42:12] [PASSED] drm_test_rect_clip_scaled_clipped
[14:42:12] [PASSED] drm_test_rect_clip_scaled_signed_vs_unsigned
[14:42:12] ================= drm_test_rect_intersect =================
[14:42:12] [PASSED] top-left x bottom-right: 2x2+1+1 x 2x2+0+0
[14:42:12] [PASSED] top-right x bottom-left: 2x2+0+0 x 2x2+1-1
[14:42:12] [PASSED] bottom-left x top-right: 2x2+1-1 x 2x2+0+0
[14:42:12] [PASSED] bottom-right x top-left: 2x2+0+0 x 2x2+1+1
[14:42:12] [PASSED] right x left: 2x1+0+0 x 3x1+1+0
[14:42:12] [PASSED] left x right: 3x1+1+0 x 2x1+0+0
[14:42:12] [PASSED] up x bottom: 1x2+0+0 x 1x3+0-1
[14:42:12] [PASSED] bottom x up: 1x3+0-1 x 1x2+0+0
[14:42:12] [PASSED] touching corner: 1x1+0+0 x 2x2+1+1
[14:42:12] [PASSED] touching side: 1x1+0+0 x 1x1+1+0
[14:42:12] [PASSED] equal rects: 2x2+0+0 x 2x2+0+0
[14:42:12] [PASSED] inside another: 2x2+0+0 x 1x1+1+1
[14:42:12] [PASSED] far away: 1x1+0+0 x 1x1+3+6
[14:42:12] [PASSED] points intersecting: 0x0+5+10 x 0x0+5+10
[14:42:12] [PASSED] points not intersecting: 0x0+0+0 x 0x0+5+10
[14:42:12] ============= [PASSED] drm_test_rect_intersect =============
[14:42:12] ================ drm_test_rect_calc_hscale ================
[14:42:12] [PASSED] normal use
[14:42:12] [PASSED] out of max range
[14:42:12] [PASSED] out of min range
[14:42:12] [PASSED] zero dst
[14:42:12] [PASSED] negative src
[14:42:12] [PASSED] negative dst
[14:42:12] ============ [PASSED] drm_test_rect_calc_hscale ============
[14:42:12] ================ drm_test_rect_calc_vscale ================
[14:42:12] [PASSED] normal use
[14:42:12] [PASSED] out of max range
[14:42:12] [PASSED] out of min range
[14:42:12] [PASSED] zero dst
[14:42:12] [PASSED] negative src
[14:42:12] [PASSED] negative dst
[14:42:12] ============ [PASSED] drm_test_rect_calc_vscale ============
[14:42:12] ================== drm_test_rect_rotate ===================
[14:42:12] [PASSED] reflect-x
[14:42:12] [PASSED] reflect-y
[14:42:12] [PASSED] rotate-0
[14:42:12] [PASSED] rotate-90
[14:42:12] [PASSED] rotate-180
[14:42:12] [PASSED] rotate-270
[14:42:12] ============== [PASSED] drm_test_rect_rotate ===============
[14:42:12] ================ drm_test_rect_rotate_inv =================
[14:42:12] [PASSED] reflect-x
[14:42:12] [PASSED] reflect-y
[14:42:12] [PASSED] rotate-0
[14:42:12] [PASSED] rotate-90
[14:42:12] [PASSED] rotate-180
[14:42:12] [PASSED] rotate-270
[14:42:12] ============ [PASSED] drm_test_rect_rotate_inv =============
[14:42:12] ==================== [PASSED] drm_rect =====================
[14:42:12] ============ drm_sysfb_modeset_test (1 subtest) ============
[14:42:12] ============ drm_test_sysfb_build_fourcc_list =============
[14:42:12] [PASSED] no native formats
[14:42:12] [PASSED] XRGB8888 as native format
[14:42:12] [PASSED] remove duplicates
[14:42:12] [PASSED] convert alpha formats
[14:42:12] [PASSED] random formats
[14:42:12] ======== [PASSED] drm_test_sysfb_build_fourcc_list =========
[14:42:12] ============= [PASSED] drm_sysfb_modeset_test ==============
[14:42:12] ================== drm_fixp (2 subtests) ===================
[14:42:12] [PASSED] drm_test_int2fixp
[14:42:12] [PASSED] drm_test_sm2fixp
[14:42:12] ==================== [PASSED] drm_fixp =====================
[14:42:12] ============================================================
[14:42:12] Testing complete. Ran 671 tests: passed: 671
[14:42:12] Elapsed time: 26.882s total, 1.796s configuring, 24.620s building, 0.414s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/ttm/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/ttm/tests/.kunitconfig
[14:42:12] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[14:42:14] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[14:42:24] Starting KUnit Kernel (1/1)...
[14:42:24] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[14:42:24] ============= refcount_interrupt (4 subtests) ==============
[14:42:24] [PASSED] test_single_irq_change
[14:42:24] [PASSED] test_nested_irq_change
[14:42:24] [PASSED] test_multiple_irq_change
[14:42:24] [PASSED] test_irq_save
[14:42:24] =============== [PASSED] refcount_interrupt ================
[14:42:24] ================= ttm_device (5 subtests) ==================
[14:42:24] [PASSED] ttm_device_init_basic
[14:42:24] [PASSED] ttm_device_init_multiple
[14:42:24] [PASSED] ttm_device_fini_basic
[14:42:24] [PASSED] ttm_device_init_no_vma_man
[14:42:24] ================== ttm_device_init_pools ==================
[14:42:24] [PASSED] No DMA allocations, no DMA32 required
[14:42:24] [PASSED] DMA allocations, DMA32 required
[14:42:24] [PASSED] No DMA allocations, DMA32 required
[14:42:24] [PASSED] DMA allocations, no DMA32 required
[14:42:24] ============== [PASSED] ttm_device_init_pools ==============
[14:42:24] =================== [PASSED] ttm_device ====================
[14:42:24] ================== ttm_pool (8 subtests) ===================
[14:42:24] ================== ttm_pool_alloc_basic ===================
[14:42:24] [PASSED] One page
[14:42:24] [PASSED] More than one page
[14:42:24] [PASSED] Above the allocation limit
[14:42:24] [PASSED] One page, with coherent DMA mappings enabled
[14:42:24] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[14:42:24] ============== [PASSED] ttm_pool_alloc_basic ===============
[14:42:24] ============== ttm_pool_alloc_basic_dma_addr ==============
[14:42:24] [PASSED] One page
[14:42:24] [PASSED] More than one page
[14:42:24] [PASSED] Above the allocation limit
[14:42:24] [PASSED] One page, with coherent DMA mappings enabled
[14:42:24] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[14:42:24] ========== [PASSED] ttm_pool_alloc_basic_dma_addr ==========
[14:42:24] [PASSED] ttm_pool_alloc_order_caching_match
[14:42:24] [PASSED] ttm_pool_alloc_caching_mismatch
[14:42:24] [PASSED] ttm_pool_alloc_order_mismatch
[14:42:24] [PASSED] ttm_pool_free_dma_alloc
[14:42:24] [PASSED] ttm_pool_free_no_dma_alloc
[14:42:24] [PASSED] ttm_pool_fini_basic
[14:42:24] ==================== [PASSED] ttm_pool =====================
[14:42:24] ================ ttm_resource (8 subtests) =================
[14:42:24] ================= ttm_resource_init_basic =================
[14:42:24] [PASSED] Init resource in TTM_PL_SYSTEM
[14:42:24] [PASSED] Init resource in TTM_PL_VRAM
[14:42:24] [PASSED] Init resource in a private placement
[14:42:24] [PASSED] Init resource in TTM_PL_SYSTEM, set placement flags
[14:42:24] ============= [PASSED] ttm_resource_init_basic =============
[14:42:24] [PASSED] ttm_resource_init_pinned
[14:42:24] [PASSED] ttm_resource_fini_basic
[14:42:24] [PASSED] ttm_resource_manager_init_basic
[14:42:24] [PASSED] ttm_resource_manager_usage_basic
[14:42:24] [PASSED] ttm_resource_manager_set_used_basic
[14:42:24] [PASSED] ttm_sys_man_alloc_basic
[14:42:24] [PASSED] ttm_sys_man_free_basic
[14:42:24] ================== [PASSED] ttm_resource ===================
[14:42:24] =================== ttm_tt (15 subtests) ===================
[14:42:24] ==================== ttm_tt_init_basic ====================
[14:42:24] [PASSED] Page-aligned size
[14:42:24] [PASSED] Extra pages requested
[14:42:24] ================ [PASSED] ttm_tt_init_basic ================
[14:42:24] [PASSED] ttm_tt_init_misaligned
[14:42:24] [PASSED] ttm_tt_fini_basic
[14:42:24] [PASSED] ttm_tt_fini_sg
[14:42:24] [PASSED] ttm_tt_fini_shmem
[14:42:24] [PASSED] ttm_tt_create_basic
[14:42:24] [PASSED] ttm_tt_create_invalid_bo_type
[14:42:24] [PASSED] ttm_tt_create_ttm_exists
[14:42:24] [PASSED] ttm_tt_create_failed
[14:42:24] [PASSED] ttm_tt_destroy_basic
[14:42:24] [PASSED] ttm_tt_populate_null_ttm
[14:42:24] [PASSED] ttm_tt_populate_populated_ttm
[14:42:24] [PASSED] ttm_tt_unpopulate_basic
[14:42:24] [PASSED] ttm_tt_unpopulate_empty_ttm
[14:42:24] [PASSED] ttm_tt_swapin_basic
[14:42:24] ===================== [PASSED] ttm_tt ======================
[14:42:24] =================== ttm_bo (14 subtests) ===================
[14:42:24] =========== ttm_bo_reserve_optimistic_no_ticket ===========
[14:42:24] [PASSED] Cannot be interrupted and sleeps
[14:42:24] [PASSED] Cannot be interrupted, locks straight away
[14:42:24] [PASSED] Can be interrupted, sleeps
[14:42:24] ======= [PASSED] ttm_bo_reserve_optimistic_no_ticket =======
[14:42:24] [PASSED] ttm_bo_reserve_locked_no_sleep
[14:42:24] [PASSED] ttm_bo_reserve_no_wait_ticket
[14:42:24] [PASSED] ttm_bo_reserve_double_resv
[14:42:24] [PASSED] ttm_bo_reserve_interrupted
[14:42:24] [PASSED] ttm_bo_reserve_deadlock
[14:42:24] [PASSED] ttm_bo_unreserve_basic
[14:42:24] [PASSED] ttm_bo_unreserve_pinned
[14:42:24] [PASSED] ttm_bo_unreserve_bulk
[14:42:24] [PASSED] ttm_bo_fini_basic
[14:42:24] [PASSED] ttm_bo_fini_shared_resv
[14:42:24] [PASSED] ttm_bo_pin_basic
[14:42:24] [PASSED] ttm_bo_pin_unpin_resource
[14:42:24] [PASSED] ttm_bo_multiple_pin_one_unpin
[14:42:24] ===================== [PASSED] ttm_bo ======================
[14:42:24] ============== ttm_bo_validate (22 subtests) ===============
[14:42:24] ============== ttm_bo_init_reserved_sys_man ===============
[14:42:24] [PASSED] Buffer object for userspace
[14:42:24] [PASSED] Kernel buffer object
[14:42:24] [PASSED] Shared buffer object
[14:42:24] ========== [PASSED] ttm_bo_init_reserved_sys_man ===========
[14:42:24] ============== ttm_bo_init_reserved_mock_man ==============
[14:42:24] [PASSED] Buffer object for userspace
[14:42:24] [PASSED] Kernel buffer object
[14:42:24] [PASSED] Shared buffer object
[14:42:24] ========== [PASSED] ttm_bo_init_reserved_mock_man ==========
[14:42:24] [PASSED] ttm_bo_init_reserved_resv
[14:42:24] ================== ttm_bo_validate_basic ==================
[14:42:24] [PASSED] Buffer object for userspace
[14:42:24] [PASSED] Kernel buffer object
[14:42:24] [PASSED] Shared buffer object
[14:42:24] ============== [PASSED] ttm_bo_validate_basic ==============
[14:42:24] [PASSED] ttm_bo_validate_invalid_placement
[14:42:24] ============= ttm_bo_validate_same_placement ==============
[14:42:24] [PASSED] System manager
[14:42:24] [PASSED] VRAM manager
[14:42:24] ========= [PASSED] ttm_bo_validate_same_placement ==========
[14:42:24] [PASSED] ttm_bo_validate_failed_alloc
[14:42:24] [PASSED] ttm_bo_validate_pinned
[14:42:24] [PASSED] ttm_bo_validate_busy_placement
[14:42:24] ================ ttm_bo_validate_multihop =================
[14:42:24] [PASSED] Buffer object for userspace
[14:42:24] [PASSED] Kernel buffer object
[14:42:24] [PASSED] Shared buffer object
[14:42:24] ============ [PASSED] ttm_bo_validate_multihop =============
[14:42:24] ========== ttm_bo_validate_no_placement_signaled ==========
[14:42:24] [PASSED] Buffer object in system domain, no page vector
[14:42:24] [PASSED] Buffer object in system domain with an existing page vector
[14:42:24] ====== [PASSED] ttm_bo_validate_no_placement_signaled ======
[14:42:24] ======== ttm_bo_validate_no_placement_not_signaled ========
[14:42:24] [PASSED] Buffer object for userspace
[14:42:24] [PASSED] Kernel buffer object
[14:42:24] [PASSED] Shared buffer object
[14:42:24] ==== [PASSED] ttm_bo_validate_no_placement_not_signaled ====
[14:42:24] [PASSED] ttm_bo_validate_move_fence_signaled
[14:42:24] ========= ttm_bo_validate_move_fence_not_signaled =========
[14:42:24] [PASSED] Waits for GPU
[14:42:24] [PASSED] Tries to lock straight away
[14:42:24] ===== [PASSED] ttm_bo_validate_move_fence_not_signaled =====
[14:42:24] [PASSED] ttm_bo_validate_swapout
[14:42:24] [PASSED] ttm_bo_validate_happy_evict
[14:42:24] [PASSED] ttm_bo_validate_all_pinned_evict
[14:42:24] [PASSED] ttm_bo_validate_allowed_only_evict
[14:42:24] [PASSED] ttm_bo_validate_deleted_evict
[14:42:24] [PASSED] ttm_bo_validate_busy_domain_evict
[14:42:24] [PASSED] ttm_bo_validate_evict_gutting
[14:42:24] [PASSED] ttm_bo_validate_recrusive_evict
[14:42:24] ================= [PASSED] ttm_bo_validate =================
[14:42:24] ============================================================
[14:42:24] Testing complete. Ran 106 tests: passed: 106
[14:42:24] Elapsed time: 12.030s total, 1.818s configuring, 9.997s building, 0.182s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/dma-buf/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/dma-buf/.kunitconfig
[14:42:24] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[14:42:26] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[14:42:35] Starting KUnit Kernel (1/1)...
[14:42:35] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[14:42:35] ============= refcount_interrupt (4 subtests) ==============
[14:42:35] [PASSED] test_single_irq_change
[14:42:35] [PASSED] test_nested_irq_change
[14:42:35] [PASSED] test_multiple_irq_change
[14:42:35] [PASSED] test_irq_save
[14:42:35] =============== [PASSED] refcount_interrupt ================
[14:42:35] =============== dma-buf-fence (12 subtests) ================
[14:42:35] [PASSED] test_sanitycheck
[14:42:35] [PASSED] test_signaling
[14:42:35] [PASSED] test_add_callback
[14:42:35] [PASSED] test_late_add_callback
[14:42:35] [PASSED] test_rm_callback
[14:42:35] [PASSED] test_late_rm_callback
[14:42:35] [PASSED] test_status
[14:42:35] [PASSED] test_error
[14:42:35] [PASSED] test_wait
[14:42:35] [PASSED] test_wait_timeout
[14:42:35] [PASSED] test_stub
[14:42:35] [SKIPPED] test_race_signal_callback (requires at least 2 CPUs)
[14:42:35] ================== [PASSED] dma-buf-fence ==================
[14:42:35] ============ dma-buf-fence-chain (11 subtests) =============
[14:42:35] [PASSED] test_sanitycheck
[14:42:35] [PASSED] test_find_seqno
[14:42:35] [PASSED] test_find_signaled
[14:42:35] [PASSED] test_find_out_of_order
[14:42:40] [PASSED] test_find_gap
[14:42:40] [PASSED] test_find_race
[14:42:40] [PASSED] test_signal_forward
[14:42:40] [PASSED] test_signal_backward
[14:42:40] [PASSED] test_wait_forward
[14:42:40] [PASSED] test_wait_backward
[14:42:40] [PASSED] test_wait_random
[14:42:40] =============== [PASSED] dma-buf-fence-chain ===============
[14:42:40] ============ dma-buf-fence-unwrap (10 subtests) ============
[14:42:40] [PASSED] test_sanitycheck
[14:42:40] [PASSED] test_unwrap_array
[14:42:40] [PASSED] test_unwrap_chain
[14:42:40] [PASSED] test_unwrap_chain_array
[14:42:40] [PASSED] test_unwrap_merge
[14:42:40] [PASSED] test_unwrap_merge_duplicate
[14:42:40] [PASSED] test_unwrap_merge_seqno
[14:42:40] [PASSED] test_unwrap_merge_order
[14:42:40] [PASSED] test_unwrap_merge_complex
[14:42:40] [PASSED] test_unwrap_merge_complex_seqno
[14:42:40] ============== [PASSED] dma-buf-fence-unwrap ===============
[14:42:40] ================ dma-buf-resv (5 subtests) =================
[14:42:40] [PASSED] test_sanitycheck
[14:42:40] ===================== test_signaling ======================
[14:42:40] [PASSED] kernel
[14:42:40] [PASSED] write
[14:42:40] [PASSED] read
[14:42:40] [PASSED] bookkeep
[14:42:40] ================= [PASSED] test_signaling ==================
[14:42:40] ====================== test_for_each ======================
[14:42:40] [PASSED] kernel
[14:42:40] [PASSED] write
[14:42:40] [PASSED] read
[14:42:40] [PASSED] bookkeep
[14:42:40] ================== [PASSED] test_for_each ==================
[14:42:40] ================= test_for_each_unlocked ==================
[14:42:40] [PASSED] kernel
[14:42:40] [PASSED] write
[14:42:40] [PASSED] read
[14:42:40] [PASSED] bookkeep
[14:42:40] ============= [PASSED] test_for_each_unlocked ==============
[14:42:40] ===================== test_get_fences =====================
[14:42:40] [PASSED] kernel
[14:42:40] [PASSED] write
[14:42:40] [PASSED] read
[14:42:40] [PASSED] bookkeep
[14:42:40] ================= [PASSED] test_get_fences =================
[14:42:40] ================== [PASSED] dma-buf-resv ===================
[14:42:40] ============================================================
[14:42:40] Testing complete. Ran 54 tests: passed: 53, skipped: 1
[14:42:40] Elapsed time: 15.744s total, 1.775s configuring, 8.646s building, 5.284s running
+ cleanup
++ stat -c %u:%g /kernel
+ chown -R 1003:1003 /kernel
^ permalink raw reply [flat|nested] 28+ messages in thread
* ✓ Xe.CI.BAT: success for Add support to handle memory double-bit ecc errors (rev3)
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (6 preceding siblings ...)
2026-09-28 14:42 ` ✓ CI.KUnit: success for Add support to handle memory double-bit ecc errors (rev3) Patchwork
@ 2026-09-28 15:27 ` Patchwork
2026-09-28 17:40 ` ✓ Xe.CI.FULL: " Patchwork
8 siblings, 0 replies; 28+ messages in thread
From: Patchwork @ 2026-09-28 15:27 UTC (permalink / raw)
To: Tauro, Riana; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 1536 bytes --]
== Series Details ==
Series: Add support to handle memory double-bit ecc errors (rev3)
URL : https://patchwork.freedesktop.org/series/172710/
State : success
== Summary ==
CI Bug Log - changes from xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5_BAT -> xe-pw-172710v3_BAT
====================================================
Summary
-------
**SUCCESS**
No regressions found.
Participating hosts (15 -> 13)
------------------------------
Missing (2): bat-bmg-vm bat-ptl-vm
Known issues
------------
Here are the changes found in xe-pw-172710v3_BAT that come from known issues:
### IGT changes ###
#### Issues hit ####
* igt@core_hotunplug@unbind-rebind:
- bat-bmg-2: [PASS][1] -> [ABORT][2] ([Intel XE#8007])
[1]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/bat-bmg-2/igt@core_hotunplug@unbind-rebind.html
[2]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/bat-bmg-2/igt@core_hotunplug@unbind-rebind.html
[Intel XE#8007]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8007
Build changes
-------------
* Linux: xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5 -> xe-pw-172710v3
IGT_9113: 691cc6fd50d5a3f6f2b5805c9698af1df28f6830 @ https://gitlab.freedesktop.org/drm/igt-gpu-tools.git
xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5: 183265fddd55ef951adb7bcd5703927abaa76bb5
xe-pw-172710v3: 172710v3
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/index.html
[-- Attachment #2: Type: text/html, Size: 2101 bytes --]
^ permalink raw reply [flat|nested] 28+ messages in thread
* ✓ Xe.CI.FULL: success for Add support to handle memory double-bit ecc errors (rev3)
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
` (7 preceding siblings ...)
2026-09-28 15:27 ` ✓ Xe.CI.BAT: " Patchwork
@ 2026-09-28 17:40 ` Patchwork
8 siblings, 0 replies; 28+ messages in thread
From: Patchwork @ 2026-09-28 17:40 UTC (permalink / raw)
To: Tauro, Riana; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 42659 bytes --]
== Series Details ==
Series: Add support to handle memory double-bit ecc errors (rev3)
URL : https://patchwork.freedesktop.org/series/172710/
State : success
== Summary ==
CI Bug Log - changes from xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5_FULL -> xe-pw-172710v3_FULL
====================================================
Summary
-------
**SUCCESS**
No regressions found.
Participating hosts (3 -> 3)
------------------------------
No changes in participating hosts
Known issues
------------
Here are the changes found in xe-pw-172710v3_FULL that come from known issues:
### IGT changes ###
#### Issues hit ####
* igt@kms_big_fb@4-tiled-max-hw-stride-64bpp-rotate-0-hflip-async-flip:
- shard-ptl: NOTRUN -> [SKIP][1] ([Intel XE#6925])
[1]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_big_fb@4-tiled-max-hw-stride-64bpp-rotate-0-hflip-async-flip.html
* igt@kms_big_fb@x-tiled-32bpp-rotate-90:
- shard-ptl: NOTRUN -> [SKIP][2] ([Intel XE#5808]) +6 other tests skip
[2]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_big_fb@x-tiled-32bpp-rotate-90.html
* igt@kms_big_fb@x-tiled-addfb-size-overflow:
- shard-ptl: NOTRUN -> [SKIP][3] ([Intel XE#5855])
[3]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_big_fb@x-tiled-addfb-size-overflow.html
* igt@kms_big_fb@y-tiled-64bpp-rotate-180:
- shard-bmg: NOTRUN -> [SKIP][4] ([Intel XE#1124])
[4]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_big_fb@y-tiled-64bpp-rotate-180.html
* igt@kms_big_fb@yf-tiled-addfb:
- shard-ptl: NOTRUN -> [SKIP][5] ([Intel XE#5850])
[5]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_big_fb@yf-tiled-addfb.html
* igt@kms_bw@connected-linear-tiling-4-displays-target-2560x1440p:
- shard-ptl: NOTRUN -> [SKIP][6] ([Intel XE#7679])
[6]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_bw@connected-linear-tiling-4-displays-target-2560x1440p.html
* igt@kms_bw@linear-tiling-2-displays-target-3840x2160p:
- shard-ptl: NOTRUN -> [SKIP][7] ([Intel XE#367])
[7]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@kms_bw@linear-tiling-2-displays-target-3840x2160p.html
* igt@kms_ccs@missing-ccs-buffer-y-tiled-gen12-rc-ccs-cc:
- shard-ptl: NOTRUN -> [SKIP][8] ([Intel XE#5822]) +9 other tests skip
[8]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_ccs@missing-ccs-buffer-y-tiled-gen12-rc-ccs-cc.html
* igt@kms_ccs@random-ccs-data-4-tiled-dg2-rc-ccs-cc:
- shard-bmg: NOTRUN -> [SKIP][9] ([Intel XE#2887]) +1 other test skip
[9]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@kms_ccs@random-ccs-data-4-tiled-dg2-rc-ccs-cc.html
* igt@kms_chamelium_audio@dp-audio-after-suspend:
- shard-ptl: NOTRUN -> [SKIP][10] ([Intel XE#2252]) +3 other tests skip
[10]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_chamelium_audio@dp-audio-after-suspend.html
* igt@kms_chamelium_color@ctm-negative:
- shard-ptl: NOTRUN -> [SKIP][11] ([Intel XE#5824]) +1 other test skip
[11]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_chamelium_color@ctm-negative.html
* igt@kms_chamelium_color@ctm-red-to-blue:
- shard-bmg: NOTRUN -> [SKIP][12] ([Intel XE#2325] / [Intel XE#7358])
[12]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_chamelium_color@ctm-red-to-blue.html
* igt@kms_chamelium_color_pipeline@plane-lut1d-ctm3x4:
- shard-ptl: NOTRUN -> [SKIP][13] ([Intel XE#8391]) +1 other test skip
[13]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_chamelium_color_pipeline@plane-lut1d-ctm3x4.html
* igt@kms_color_pipeline@plane-lut1d-post-ctm3x4@pipe-d-plane-4:
- shard-ptl: NOTRUN -> [SKIP][14] ([Intel XE#6969]) +7 other tests skip
[14]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_color_pipeline@plane-lut1d-post-ctm3x4@pipe-d-plane-4.html
* igt@kms_content_protection@dp-mst-type-1-suspend-resume:
- shard-ptl: NOTRUN -> [SKIP][15] ([Intel XE#6974])
[15]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_content_protection@dp-mst-type-1-suspend-resume.html
* igt@kms_content_protection@lic-type-0:
- shard-ptl: NOTRUN -> [SKIP][16] ([Intel XE#7642]) +2 other tests skip
[16]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_content_protection@lic-type-0.html
* igt@kms_cursor_crc@cursor-offscreen-128x42:
- shard-bmg: NOTRUN -> [SKIP][17] ([Intel XE#2320])
[17]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@kms_cursor_crc@cursor-offscreen-128x42.html
* igt@kms_cursor_crc@cursor-sliding-128x42:
- shard-ptl: NOTRUN -> [SKIP][18] ([Intel XE#5900]) +1 other test skip
[18]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_cursor_crc@cursor-sliding-128x42.html
* igt@kms_cursor_legacy@cursorb-vs-flipb-toggle:
- shard-ptl: NOTRUN -> [SKIP][19] ([Intel XE#5962] / [Intel XE#7343]) +1 other test skip
[19]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_cursor_legacy@cursorb-vs-flipb-toggle.html
* igt@kms_dp_link_training@non-uhbr-mst:
- shard-ptl: NOTRUN -> [SKIP][20] ([Intel XE#5882])
[20]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_dp_link_training@non-uhbr-mst.html
* igt@kms_dsc@dsc-fractional-bpp-ultrajoiner:
- shard-ptl: NOTRUN -> [SKIP][21] ([Intel XE#5907]) +2 other tests skip
[21]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@kms_dsc@dsc-fractional-bpp-ultrajoiner.html
* igt@kms_feature_discovery@dp-mst:
- shard-ptl: NOTRUN -> [SKIP][22] ([Intel XE#2375])
[22]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_feature_discovery@dp-mst.html
* igt@kms_flip@2x-flip-vs-suspend-interruptible:
- shard-bmg: [PASS][23] -> [ABORT][24] ([Intel XE#9362]) +1 other test abort
[23]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@kms_flip@2x-flip-vs-suspend-interruptible.html
[24]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@kms_flip@2x-flip-vs-suspend-interruptible.html
* igt@kms_flip@2x-modeset-vs-vblank-race:
- shard-ptl: NOTRUN -> [SKIP][25] ([Intel XE#2316]) +2 other tests skip
[25]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_flip@2x-modeset-vs-vblank-race.html
* igt@kms_flip@2x-wf_vblank-ts-check:
- shard-bmg: [PASS][26] -> [FAIL][27] ([Intel XE#5408] / [Intel XE#6266]) +1 other test fail
[26]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-6/igt@kms_flip@2x-wf_vblank-ts-check.html
[27]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@kms_flip@2x-wf_vblank-ts-check.html
* igt@kms_flip@flip-vs-expired-vblank@b-edp1:
- shard-lnl: [PASS][28] -> [FAIL][29] ([Intel XE#301])
[28]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-lnl-1/igt@kms_flip@flip-vs-expired-vblank@b-edp1.html
[29]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-lnl-3/igt@kms_flip@flip-vs-expired-vblank@b-edp1.html
* igt@kms_flip_scaled_crc@flip-32bpp-linear-to-32bpp-linear-reflect-x:
- shard-ptl: NOTRUN -> [SKIP][30] ([Intel XE#7179])
[30]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_flip_scaled_crc@flip-32bpp-linear-to-32bpp-linear-reflect-x.html
* igt@kms_flip_scaled_crc@flip-64bpp-ytile-to-32bpp-ytilercccs-downscaling:
- shard-ptl: NOTRUN -> [SKIP][31] ([Intel XE#7178]) +2 other tests skip
[31]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_flip_scaled_crc@flip-64bpp-ytile-to-32bpp-ytilercccs-downscaling.html
* igt@kms_frontbuffer_tracking@fbcdrrs-1p-offscreen-pri-indfb-draw-render:
- shard-bmg: NOTRUN -> [SKIP][32] ([Intel XE#2311]) +3 other tests skip
[32]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_frontbuffer_tracking@fbcdrrs-1p-offscreen-pri-indfb-draw-render.html
* igt@kms_frontbuffer_tracking@fbcdrrs-1p-primscrn-cur-indfb-move:
- shard-ptl: NOTRUN -> [SKIP][33] ([Intel XE#5800] / [Intel XE#6312]) +5 other tests skip
[33]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_frontbuffer_tracking@fbcdrrs-1p-primscrn-cur-indfb-move.html
* igt@kms_frontbuffer_tracking@fbcdrrs-abgr161616f-draw-blt:
- shard-ptl: NOTRUN -> [SKIP][34] ([Intel XE#7061] / [Intel XE#7356]) +1 other test skip
[34]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcdrrs-abgr161616f-draw-blt.html
* igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-indfb-plflip-blt:
- shard-ptl: NOTRUN -> [SKIP][35] ([Intel XE#6312]) +7 other tests skip
[35]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-indfb-plflip-blt.html
* igt@kms_frontbuffer_tracking@fbcdrrshdr-abgr161616f-draw-blt:
- shard-ptl: NOTRUN -> [SKIP][36] ([Intel XE#7061])
[36]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcdrrshdr-abgr161616f-draw-blt.html
* igt@kms_frontbuffer_tracking@fbchdr-rgb565-draw-mmap-wc:
- shard-ptl: NOTRUN -> [SKIP][37] ([Intel XE#7865]) +18 other tests skip
[37]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_frontbuffer_tracking@fbchdr-rgb565-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@fbcpsr-argb161616f-draw-blt:
- shard-bmg: NOTRUN -> [SKIP][38] ([Intel XE#7061] / [Intel XE#7356])
[38]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_frontbuffer_tracking@fbcpsr-argb161616f-draw-blt.html
* igt@kms_frontbuffer_tracking@fbcpsr-tiling-linear:
- shard-bmg: NOTRUN -> [SKIP][39] ([Intel XE#2313])
[39]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_frontbuffer_tracking@fbcpsr-tiling-linear.html
* igt@kms_frontbuffer_tracking@fbcpsrhdr-2p-scndscrn-pri-shrfb-draw-mmap-wc:
- shard-ptl: NOTRUN -> [SKIP][40] ([Intel XE#5812]) +49 other tests skip
[40]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcpsrhdr-2p-scndscrn-pri-shrfb-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@fbcpsrhdr-tiling-y:
- shard-ptl: NOTRUN -> [SKIP][41] ([Intel XE#7399])
[41]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcpsrhdr-tiling-y.html
* igt@kms_hdr@invalid-hdr:
- shard-bmg: [PASS][42] -> [SKIP][43] ([Intel XE#1503])
[42]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-8/igt@kms_hdr@invalid-hdr.html
[43]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_hdr@invalid-hdr.html
* igt@kms_hdr@static-toggle-suspend:
- shard-bmg: [PASS][44] -> [DMESG-WARN][45] ([Intel XE#9365]) +1 other test dmesg-warn
[44]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-2/igt@kms_hdr@static-toggle-suspend.html
[45]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@kms_hdr@static-toggle-suspend.html
* igt@kms_joiner@basic-ultra-joiner:
- shard-ptl: NOTRUN -> [SKIP][46] ([Intel XE#6900])
[46]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_joiner@basic-ultra-joiner.html
* igt@kms_plane@pixel-format-yf-tiled-ccs-modifier-source-clamping:
- shard-ptl: NOTRUN -> [SKIP][47] ([Intel XE#7283]) +3 other tests skip
[47]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_plane@pixel-format-yf-tiled-ccs-modifier-source-clamping.html
* igt@kms_plane@plane-panning-bottom-right-suspend:
- shard-ptl: [PASS][48] -> [ABORT][49] ([Intel XE#9362]) +6 other tests abort
[48]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-6/igt@kms_plane@plane-panning-bottom-right-suspend.html
[49]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-7/igt@kms_plane@plane-panning-bottom-right-suspend.html
* igt@kms_plane_lowres@tiling-4:
- shard-ptl: NOTRUN -> [SKIP][50] ([Intel XE#5844]) +4 other tests skip
[50]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_plane_lowres@tiling-4.html
* igt@kms_pm_rpm@modeset-non-lpsp-stress-no-wait:
- shard-ptl: NOTRUN -> [SKIP][51] ([Intel XE#5849])
[51]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@kms_pm_rpm@modeset-non-lpsp-stress-no-wait.html
* igt@kms_psr2_sf@fbc-pr-overlay-plane-move-continuous-exceed-sf:
- shard-ptl: NOTRUN -> [SKIP][52] ([Intel XE#7304]) +3 other tests skip
[52]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@kms_psr2_sf@fbc-pr-overlay-plane-move-continuous-exceed-sf.html
* igt@kms_psr2_sf@fbc-psr2-cursor-plane-update-sf@pipe-b-edp-1:
- shard-ptl: NOTRUN -> [SKIP][53] ([Intel XE#5858] / [Intel XE#7304]) +3 other tests skip
[53]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_psr2_sf@fbc-psr2-cursor-plane-update-sf@pipe-b-edp-1.html
* igt@kms_psr2_sf@fbc-psr2-overlay-plane-move-continuous-exceed-fully-sf@pipe-a-edp-1:
- shard-ptl: NOTRUN -> [SKIP][54] ([Intel XE#5858]) +1 other test skip
[54]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_psr2_sf@fbc-psr2-overlay-plane-move-continuous-exceed-fully-sf@pipe-a-edp-1.html
* igt@kms_psr2_sf@fbc-psr2-overlay-plane-update-sf-dmg-area:
- shard-bmg: NOTRUN -> [SKIP][55] ([Intel XE#1489])
[55]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_psr2_sf@fbc-psr2-overlay-plane-update-sf-dmg-area.html
* igt@kms_psr2_su@page_flip-nv12:
- shard-ptl: NOTRUN -> [SKIP][56] ([Intel XE#7413])
[56]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_psr2_su@page_flip-nv12.html
* igt@kms_psr@fbc-pr-basic:
- shard-bmg: NOTRUN -> [SKIP][57] ([Intel XE#2234] / [Intel XE#2850])
[57]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_psr@fbc-pr-basic.html
* igt@kms_psr@fbc-pr-no-drrs:
- shard-ptl: NOTRUN -> [SKIP][58] ([Intel XE#1406]) +1 other test skip
[58]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@kms_psr@fbc-pr-no-drrs.html
* igt@kms_psr@fbc-psr2-primary-blt@edp-1:
- shard-ptl: NOTRUN -> [SKIP][59] ([Intel XE#1406] / [Intel XE#5835]) +1 other test skip
[59]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_psr@fbc-psr2-primary-blt@edp-1.html
* igt@kms_psr@psr-suspend@edp-1:
- shard-ptl: NOTRUN -> [ABORT][60] ([Intel XE#9362]) +1 other test abort
[60]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@kms_psr@psr-suspend@edp-1.html
* igt@kms_rotation_crc@primary-4-tiled-reflect-x-0:
- shard-ptl: NOTRUN -> [SKIP][61] ([Intel XE#5821]) +2 other tests skip
[61]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@kms_rotation_crc@primary-4-tiled-reflect-x-0.html
* igt@kms_rotation_crc@sprite-rotation-90-pos-100-0:
- shard-bmg: NOTRUN -> [SKIP][62] ([Intel XE#3904] / [Intel XE#7342])
[62]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_rotation_crc@sprite-rotation-90-pos-100-0.html
* igt@kms_vrr@flip-dpms@pipe-a-edp-1:
- shard-ptl: [PASS][63] -> [FAIL][64] ([Intel XE#4227] / [Intel XE#7397]) +1 other test fail
[63]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-5/igt@kms_vrr@flip-dpms@pipe-a-edp-1.html
[64]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-2/igt@kms_vrr@flip-dpms@pipe-a-edp-1.html
* igt@kms_vrr@flip-suspend:
- shard-bmg: NOTRUN -> [SKIP][65] ([Intel XE#1499] / [Intel XE#9271])
[65]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_vrr@flip-suspend.html
* igt@xe_ccs@vm-bind-fault-mode-decompress:
- shard-ptl: NOTRUN -> [SKIP][66] ([Intel XE#7644])
[66]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@xe_ccs@vm-bind-fault-mode-decompress.html
* igt@xe_configfs@ctx-restore-mid-bb-invalid:
- shard-bmg: [PASS][67] -> [ABORT][68] ([Intel XE#8007])
[67]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-2/igt@xe_configfs@ctx-restore-mid-bb-invalid.html
[68]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@xe_configfs@ctx-restore-mid-bb-invalid.html
* igt@xe_create@create-big-vram:
- shard-ptl: NOTRUN -> [SKIP][69] ([Intel XE#5853] / [Intel XE#7457])
[69]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_create@create-big-vram.html
* igt@xe_evict@evict-cm-threads-small-multi-vm:
- shard-ptl: NOTRUN -> [SKIP][70] ([Intel XE#5764]) +3 other tests skip
[70]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_evict@evict-cm-threads-small-multi-vm.html
* igt@xe_evict_ccs@evict-overcommit-standalone-instantfree-reopen:
- shard-ptl: NOTRUN -> [SKIP][71] ([Intel XE#6144])
[71]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_evict_ccs@evict-overcommit-standalone-instantfree-reopen.html
* igt@xe_exec_balancer@twice-virtual-basic:
- shard-ptl: NOTRUN -> [SKIP][72] ([Intel XE#7482]) +12 other tests skip
[72]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_exec_balancer@twice-virtual-basic.html
* igt@xe_exec_basic@multigpu-no-exec-rebind:
- shard-bmg: NOTRUN -> [SKIP][73] ([Intel XE#2322] / [Intel XE#7372])
[73]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@xe_exec_basic@multigpu-no-exec-rebind.html
* igt@xe_exec_basic@multigpu-once-bindexecqueue-userptr-invalidate-race:
- shard-ptl: NOTRUN -> [SKIP][74] ([Intel XE#5575] / [Intel XE#7484]) +5 other tests skip
[74]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_exec_basic@multigpu-once-bindexecqueue-userptr-invalidate-race.html
* igt@xe_exec_fault_mode@many-execqueues-multi-queue-prefetch:
- shard-ptl: NOTRUN -> [SKIP][75] ([Intel XE#8374]) +7 other tests skip
[75]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_exec_fault_mode@many-execqueues-multi-queue-prefetch.html
* igt@xe_exec_multi_queue@many-queues-preempt-mode-fault-priority:
- shard-ptl: NOTRUN -> [SKIP][76] ([Intel XE#8364]) +19 other tests skip
[76]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_exec_multi_queue@many-queues-preempt-mode-fault-priority.html
* igt@xe_exec_multi_queue@one-queue-preempt-mode-fault-dyn-priority-smem:
- shard-bmg: NOTRUN -> [SKIP][77] ([Intel XE#8364]) +2 other tests skip
[77]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@xe_exec_multi_queue@one-queue-preempt-mode-fault-dyn-priority-smem.html
* igt@xe_exec_reset@cm-multi-queue-gt-reset:
- shard-ptl: NOTRUN -> [SKIP][78] ([Intel XE#8369]) +2 other tests skip
[78]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_exec_reset@cm-multi-queue-gt-reset.html
* igt@xe_exec_threads@threads-hang-userptr-invalidate-race:
- shard-bmg: [PASS][79] -> [FAIL][80] ([Intel XE#9096])
[79]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
[80]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
- shard-lnl: [PASS][81] -> [FAIL][82] ([Intel XE#9096])
[81]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-lnl-1/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
[82]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-lnl-3/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
* igt@xe_exec_threads@threads-multi-queue-mixed-fd-userptr-invalidate:
- shard-ptl: NOTRUN -> [SKIP][83] ([Intel XE#8378]) +2 other tests skip
[83]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_exec_threads@threads-multi-queue-mixed-fd-userptr-invalidate.html
* igt@xe_multigpu_svm@mgpu-atomic-op-conflict:
- shard-ptl: NOTRUN -> [SKIP][84] ([Intel XE#6964]) +2 other tests skip
[84]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_multigpu_svm@mgpu-atomic-op-conflict.html
* igt@xe_multigpu_svm@mgpu-xgpu-access-basic:
- shard-bmg: NOTRUN -> [SKIP][85] ([Intel XE#6964])
[85]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@xe_multigpu_svm@mgpu-xgpu-access-basic.html
* igt@xe_peer2peer@write:
- shard-ptl: NOTRUN -> [SKIP][86] ([Intel XE#6588])
[86]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_peer2peer@write.html
* igt@xe_pm@s2idle-d3cold-basic-exec:
- shard-ptl: NOTRUN -> [SKIP][87] ([Intel XE#5741])
[87]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_pm@s2idle-d3cold-basic-exec.html
* igt@xe_pm@s4-basic:
- shard-lnl: [PASS][88] -> [ABORT][89] ([Intel XE#9362]) +2 other tests abort
[88]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-lnl-5/igt@xe_pm@s4-basic.html
[89]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-lnl-1/igt@xe_pm@s4-basic.html
* igt@xe_query@multigpu-query-engines:
- shard-ptl: NOTRUN -> [SKIP][90] ([Intel XE#6019]) +1 other test skip
[90]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_query@multigpu-query-engines.html
* igt@xe_sriov_vram@vf-access-after-resize-up:
- shard-ptl: NOTRUN -> [SKIP][91] ([Intel XE#6376] / [Intel XE#7422]) +1 other test skip
[91]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_sriov_vram@vf-access-after-resize-up.html
* igt@xe_vm@overcommit-nonfault-vram-lr-external-nodefer:
- shard-ptl: NOTRUN -> [SKIP][92] ([Intel XE#7892])
[92]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-5/igt@xe_vm@overcommit-nonfault-vram-lr-external-nodefer.html
#### Possible fixes ####
* igt@kms_async_flips@alternate-sync-async-flip:
- shard-bmg: [FAIL][93] ([Intel XE#3718] / [Intel XE#6078]) -> [PASS][94]
[93]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-9/igt@kms_async_flips@alternate-sync-async-flip.html
[94]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@kms_async_flips@alternate-sync-async-flip.html
* igt@kms_async_flips@alternate-sync-async-flip@pipe-a-dp-2:
- shard-bmg: [FAIL][95] ([Intel XE#6078]) -> [PASS][96]
[95]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-9/igt@kms_async_flips@alternate-sync-async-flip@pipe-a-dp-2.html
[96]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@kms_async_flips@alternate-sync-async-flip@pipe-a-dp-2.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs:
- shard-bmg: [INCOMPLETE][97] ([Intel XE#7084] / [Intel XE#8150]) -> [PASS][98] +1 other test pass
[97]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-10/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs.html
[98]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs.html
* igt@kms_flip@plain-flip-ts-check-interruptible@a-dp2:
- shard-bmg: [FAIL][99] ([Intel XE#3098]) -> [PASS][100] +1 other test pass
[99]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@kms_flip@plain-flip-ts-check-interruptible@a-dp2.html
[100]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@kms_flip@plain-flip-ts-check-interruptible@a-dp2.html
* igt@kms_vrr@flip-basic-fastset:
- shard-ptl: [FAIL][101] ([Intel XE#4227] / [Intel XE#7397]) -> [PASS][102]
[101]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-8/igt@kms_vrr@flip-basic-fastset.html
[102]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-3/igt@kms_vrr@flip-basic-fastset.html
* igt@kms_vrr@flip-basic-fastset@pipe-a-edp-1:
- shard-ptl: [FAIL][103] ([Intel XE#6293] / [Intel XE#7397]) -> [PASS][104]
[103]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-8/igt@kms_vrr@flip-basic-fastset@pipe-a-edp-1.html
[104]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-3/igt@kms_vrr@flip-basic-fastset@pipe-a-edp-1.html
* igt@xe_module_load@load:
- shard-bmg: ([PASS][105], [PASS][106], [PASS][107], [PASS][108], [PASS][109], [PASS][110], [PASS][111], [PASS][112], [PASS][113], [PASS][114], [PASS][115], [SKIP][116], [PASS][117], [PASS][118], [PASS][119], [PASS][120], [PASS][121], [PASS][122], [PASS][123], [PASS][124], [PASS][125], [PASS][126], [PASS][127], [PASS][128], [PASS][129], [PASS][130]) ([Intel XE#2457] / [Intel XE#7405]) -> ([PASS][131], [PASS][132], [PASS][133], [PASS][134], [PASS][135], [PASS][136], [PASS][137], [PASS][138], [PASS][139], [PASS][140], [PASS][141], [PASS][142], [PASS][143], [PASS][144], [PASS][145], [PASS][146], [PASS][147], [PASS][148], [PASS][149], [PASS][150], [PASS][151], [PASS][152], [PASS][153], [PASS][154], [PASS][155])
[105]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-2/igt@xe_module_load@load.html
[106]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-8/igt@xe_module_load@load.html
[107]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-8/igt@xe_module_load@load.html
[108]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-2/igt@xe_module_load@load.html
[109]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-4/igt@xe_module_load@load.html
[110]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-4/igt@xe_module_load@load.html
[111]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-10/igt@xe_module_load@load.html
[112]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@xe_module_load@load.html
[113]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-3/igt@xe_module_load@load.html
[114]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-3/igt@xe_module_load@load.html
[115]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-7/igt@xe_module_load@load.html
[116]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@xe_module_load@load.html
[117]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-2/igt@xe_module_load@load.html
[118]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-9/igt@xe_module_load@load.html
[119]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-5/igt@xe_module_load@load.html
[120]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-9/igt@xe_module_load@load.html
[121]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-6/igt@xe_module_load@load.html
[122]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-3/igt@xe_module_load@load.html
[123]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-6/igt@xe_module_load@load.html
[124]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-10/igt@xe_module_load@load.html
[125]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-9/igt@xe_module_load@load.html
[126]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-5/igt@xe_module_load@load.html
[127]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@xe_module_load@load.html
[128]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@xe_module_load@load.html
[129]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-5/igt@xe_module_load@load.html
[130]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-7/igt@xe_module_load@load.html
[131]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-7/igt@xe_module_load@load.html
[132]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@xe_module_load@load.html
[133]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@xe_module_load@load.html
[134]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-4/igt@xe_module_load@load.html
[135]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@xe_module_load@load.html
[136]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-1/igt@xe_module_load@load.html
[137]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-4/igt@xe_module_load@load.html
[138]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-7/igt@xe_module_load@load.html
[139]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-7/igt@xe_module_load@load.html
[140]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-8/igt@xe_module_load@load.html
[141]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-1/igt@xe_module_load@load.html
[142]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-2/igt@xe_module_load@load.html
[143]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-9/igt@xe_module_load@load.html
[144]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-9/igt@xe_module_load@load.html
[145]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-2/igt@xe_module_load@load.html
[146]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@xe_module_load@load.html
[147]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-3/igt@xe_module_load@load.html
[148]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-10/igt@xe_module_load@load.html
[149]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@xe_module_load@load.html
[150]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@xe_module_load@load.html
[151]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-6/igt@xe_module_load@load.html
[152]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-10/igt@xe_module_load@load.html
[153]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-2/igt@xe_module_load@load.html
[154]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-5/igt@xe_module_load@load.html
[155]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-1/igt@xe_module_load@load.html
* igt@xe_pm@s4-basic:
- shard-ptl: [ABORT][156] ([Intel XE#9362]) -> [PASS][157] +1 other test pass
[156]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-4/igt@xe_pm@s4-basic.html
[157]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-8/igt@xe_pm@s4-basic.html
#### Warnings ####
* igt@kms_tiled_display@basic-test-pattern-with-chamelium:
- shard-bmg: [SKIP][158] ([Intel XE#2509] / [Intel XE#7437]) -> [SKIP][159] ([Intel XE#2426] / [Intel XE#5848])
[158]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-bmg-1/igt@kms_tiled_display@basic-test-pattern-with-chamelium.html
[159]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-bmg-7/igt@kms_tiled_display@basic-test-pattern-with-chamelium.html
* igt@xe_pat@pat-sw-hw-suspend:
- shard-ptl: [ABORT][160] ([Intel XE#9362]) -> [FAIL][161] ([Intel XE#7695])
[160]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5/shard-ptl-6/igt@xe_pat@pat-sw-hw-suspend.html
[161]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/shard-ptl-4/igt@xe_pat@pat-sw-hw-suspend.html
[Intel XE#1124]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1124
[Intel XE#1406]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1406
[Intel XE#1489]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1489
[Intel XE#1499]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1499
[Intel XE#1503]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1503
[Intel XE#2234]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2234
[Intel XE#2252]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2252
[Intel XE#2311]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2311
[Intel XE#2313]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2313
[Intel XE#2316]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2316
[Intel XE#2320]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2320
[Intel XE#2322]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2322
[Intel XE#2325]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2325
[Intel XE#2375]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2375
[Intel XE#2426]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2426
[Intel XE#2457]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2457
[Intel XE#2509]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2509
[Intel XE#2850]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2850
[Intel XE#2887]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2887
[Intel XE#301]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/301
[Intel XE#3098]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3098
[Intel XE#367]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/367
[Intel XE#3718]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3718
[Intel XE#3904]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3904
[Intel XE#4227]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4227
[Intel XE#5408]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5408
[Intel XE#5575]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5575
[Intel XE#5741]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5741
[Intel XE#5764]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5764
[Intel XE#5800]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5800
[Intel XE#5808]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5808
[Intel XE#5812]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5812
[Intel XE#5821]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5821
[Intel XE#5822]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5822
[Intel XE#5824]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5824
[Intel XE#5835]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5835
[Intel XE#5844]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5844
[Intel XE#5848]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5848
[Intel XE#5849]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5849
[Intel XE#5850]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5850
[Intel XE#5853]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5853
[Intel XE#5855]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5855
[Intel XE#5858]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5858
[Intel XE#5882]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5882
[Intel XE#5900]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5900
[Intel XE#5907]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5907
[Intel XE#5962]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5962
[Intel XE#6019]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6019
[Intel XE#6078]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6078
[Intel XE#6144]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6144
[Intel XE#6266]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6266
[Intel XE#6293]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6293
[Intel XE#6312]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6312
[Intel XE#6376]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6376
[Intel XE#6588]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6588
[Intel XE#6900]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6900
[Intel XE#6925]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6925
[Intel XE#6964]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6964
[Intel XE#6969]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6969
[Intel XE#6974]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6974
[Intel XE#7061]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7061
[Intel XE#7084]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7084
[Intel XE#7178]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7178
[Intel XE#7179]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7179
[Intel XE#7283]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7283
[Intel XE#7304]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7304
[Intel XE#7342]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7342
[Intel XE#7343]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7343
[Intel XE#7356]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7356
[Intel XE#7358]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7358
[Intel XE#7372]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7372
[Intel XE#7397]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7397
[Intel XE#7399]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7399
[Intel XE#7405]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7405
[Intel XE#7413]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7413
[Intel XE#7422]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7422
[Intel XE#7437]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7437
[Intel XE#7457]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7457
[Intel XE#7482]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7482
[Intel XE#7484]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7484
[Intel XE#7642]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7642
[Intel XE#7644]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7644
[Intel XE#7679]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7679
[Intel XE#7695]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7695
[Intel XE#7865]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7865
[Intel XE#7892]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7892
[Intel XE#8007]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8007
[Intel XE#8150]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8150
[Intel XE#8364]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8364
[Intel XE#8369]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8369
[Intel XE#8374]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8374
[Intel XE#8378]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8378
[Intel XE#8391]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8391
[Intel XE#9096]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9096
[Intel XE#9271]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9271
[Intel XE#9362]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9362
[Intel XE#9365]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9365
Build changes
-------------
* Linux: xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5 -> xe-pw-172710v3
IGT_9113: 691cc6fd50d5a3f6f2b5805c9698af1df28f6830 @ https://gitlab.freedesktop.org/drm/igt-gpu-tools.git
xe-5834-183265fddd55ef951adb7bcd5703927abaa76bb5: 183265fddd55ef951adb7bcd5703927abaa76bb5
xe-pw-172710v3: 172710v3
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-172710v3/index.html
[-- Attachment #2: Type: text/html, Size: 46886 bytes --]
^ permalink raw reply [flat|nested] 28+ messages in thread
* RE: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
2026-09-28 6:34 ` sashiko-bot
2026-09-28 8:55 ` Ghimiray, Himal Prasad
@ 2026-10-01 10:50 ` Upadhyay, Tejas
2026-10-01 11:27 ` Tauro, Riana
2 siblings, 1 reply; 28+ messages in thread
From: Upadhyay, Tejas @ 2026-10-01 10:50 UTC (permalink / raw)
To: Tauro, Riana, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
> -----Original Message-----
> From: Tauro, Riana <riana.tauro@intel.com>
> Sent: 28 September 2026 11:49
> To: intel-xe@lists.freedesktop.org
> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
> <himal.prasad.ghimiray@intel.com>
> Subject: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for
> device memory ecc errors
>
> Add basic support for sending page offline/remove requests to system
> controller and use it for device memory ECC error handling.
> Pages that belong to critical BOs cannot be handled by offlining and require a
> SBR (Secondary Bus Reset).
> Pages that are configured for log-only handling are not marked as bad by
> firmware.
>
> For all other valid page addresses, the first occurrence of error indicates a
> poison error and the page is offlined only by software.
> Firmware avoids permanently marking the page as bad. The second
> occurrence of an error indicates a Double-bit ECC error and the firmware
> permanently marks the page as bad.
>
> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: use ret in sigid logging (Mallesh)
> remove additional log
> use xe_assert (Michal)
>
> v3: align address to PAGE_SIZE (sashiko, Himal)
> rename decline to remove
> add more descriptive logs (Himal)
> ---
> drivers/gpu/drm/xe/xe_ras.c | 127 +++++++++++++++++-
> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++
> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 2 +
> 3 files changed, 159 insertions(+), 5 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
> 7a85735c57d5..1225c561a872 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -3,6 +3,8 @@
> * Copyright © 2026 Intel Corporation
> */
>
> +#include "xe_assert.h"
> +#include "xe_bo.h"
> #include "xe_configfs.h"
> #include "xe_debugfs.h"
> #include "xe_device.h"
> @@ -16,6 +18,7 @@
> #include "xe_sysctrl_event_types.h"
> #include "xe_sysctrl_mailbox.h"
> #include "xe_sysctrl_mailbox_types.h"
> +#include "xe_ttm_vram_mgr.h"
>
> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
> /*
> @@ -201,6 +204,119 @@ static inline const char *comp_to_str(u8
> component)
> return xe_ras_components[component];
> }
>
> +static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
> + enum xe_ras_page_action action)
> +{
> + struct xe_sysctrl_mailbox_command command = {0};
> + struct xe_ras_page_offline_request request = {0};
> + struct xe_ras_page_offline_response response = {0};
> + size_t rlen;
> + int ret;
> +
> + if (!xe->info.has_sysctrl)
> + return 0;
> +
> + xe_assert(xe, action < XE_RAS_PAGE_ACTION_MAX);
> +
> + request.page_address = page_address;
> + request.action = action;
> +
> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to
> offline page 0x%llx\n",
> + page_address);
> + else
> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to
> remove page 0x%llx from queue\n",
> + page_address);
> +
> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP,
> XE_SYSCTRL_CMD_PAGE_OFFLINE,
> + &request, sizeof(request), &response,
> sizeof(response));
> +
> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
> + if (ret) {
> + xe_log_err(xe, SYSCTRL, ret, "failed to send page offline
> command\n");
> + return ret;
> + }
> +
> + if (rlen != sizeof(response)) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "unexpected page offline response length %zu
> (expected %zu)\n",
> + rlen, sizeof(response));
> + return -EINVAL;
> + }
> +
> + ret = ras_status_to_errno(response.status);
> + if (ret)
> + xe_log_err(xe, SYSCTRL, ret, "page offline command failed with
> status %u\n",
> + response.status);
> +
> + return ret;
> +}
> +
> +static int handle_page_offline(struct xe_device *xe, u64 page_address,
> +bool send_cmd) {
> + enum xe_ras_page_action action;
> + u64 addr;
> + int ret = 0;
> +
> + if (!IS_ALIGNED(page_address, XE_PAGE_SIZE)) {
xe_ttm_vram_handle_addr_fault() itself asserts against PAGE_SIZE, so it'd be more consistent to validate against PAGE_SIZE here too.
Also, please have a look at Sashiko comments if applicable.
Tejas
> + xe_log_err(xe, SYSCTRL, -EINVAL, "Unaligned physical page
> address: 0x%llx\n",
> + page_address);
> + return -EINVAL;
> + }
> +
> + addr = ALIGN_DOWN(page_address, PAGE_SIZE);
> +
> + ret = xe_ttm_vram_handle_addr_fault(xe, addr);
> +
> + /*
> + * Handle return code from address fault handling function:
> + * 0: Page softofflined, remove from firmware queue
> + * -EIO: Address belongs to a critical BO/stolen area that cannot be
> offlined
> + * -EOPNOTSUPP: Address is valid and can be offlined but user policy is
> not to offline
> + * -EEXIST: Address is soft offlined but yet to be offlined by firmware
> for second
> + * occurrence
> + */
-ENOMEM - allocation failure; next action is reset. --> we should reset on this error.
> +
> + switch (ret) {
> + case 0:
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page
> soft-offlined\n",
> + page_address);
> + break;
> + /* User policy set to decline page offlining */
> + case -EOPNOTSUPP:
> + action = XE_RAS_PAGE_ACTION_REMOVE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, user
> policy set to decline soft-offlining\n",
> + page_address);
> + break;
> + case -EIO:
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Poison detected at physical address 0x%llx, page
> belongs to critical BO and cannot be soft-offlined\n",
> + page_address);
> + return ret;
> + case -EEXIST:
> + action = XE_RAS_PAGE_ACTION_OFFLINE;
> + xe_log_err(xe, DEVICE_MEMORY, ret,
> + "Double-bit ECC error detected at physical address
> 0x%llx, page soft-offlined\n",
> + page_address);
> + break;
> + default:
> + xe_log_err(xe, DEVICE_MEMORY, ret, "Failed to handle
> address fault at physical address 0x%llx\n",
> + page_address);
> + return 0;
> + }
> +
> + if (send_cmd) {
> + ret = send_page_offline_cmd(xe, page_address, action);
> + if (ret)
> + return ret;
> + }
> +
> + return 0;
> +}
> +
> static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class
> *counter) {
> u8 severity = counter->common.severity; @@ -368,11 +484,12 @@
> static u8 handle_soc_internal_errors(struct xe_device *xe, struct
> xe_ras_error_a static u8 handle_device_memory_errors(struct xe_device *xe,
> struct xe_ras_error_array *arr) {
> struct xe_ras_memory_error *info = (void *)arr->details;
> + int ret;
>
> /*
> * For memory errors, the recovery action depends on the error
> category
> *
> - * TODO: Double-bit ECC errors: Page offlining
> + * Double-bit ECC errors: Page offlining
> * Poison and data parity errors: Log only
> * For any other memory errors, request a reset as recovery
> mechanism
> */
> @@ -384,10 +501,10 @@ static u8 handle_device_memory_errors(struct
> xe_device *xe, struct xe_ras_error_
> xe_info(xe, "[RAS]: Data parity error detected\n");
> break;
> case XE_RAS_MEMORY_DB_ECC:
> - xe_info(xe, "[RAS]: Double-bit ECC error detected at sw
> address 0x%llx\n",
> - info->sw_address);
> - /* TODO: Add page offlining for Double-bit ECC error */
> - fallthrough;
> + ret = handle_page_offline(xe, info->sw_address, true);
> + if (ret)
> + return XE_RAS_RECOVERY_ACTION_RESET;
> + break;
> default:
> return XE_RAS_RECOVERY_ACTION_RESET;
> }
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
> b/drivers/gpu/drm/xe/xe_ras_types.h
> index fe6f3658a2a4..f119489bcdf2 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -17,6 +17,19 @@
> #define XE_RAS_MEMORY_POISON BIT(2)
> #define XE_RAS_MEMORY_DATA_PARITY BIT(5)
>
> +/**
> + * enum xe_ras_page_action - Page offline actions for page offline
> +request
> + *
> + * @XE_RAS_PAGE_ACTION_OFFLINE: Instruct firmware to offline the page
> + * @XE_RAS_PAGE_ACTION_REMOVE: Instruct firmware to remove the page
> +from queue
> + * @XE_RAS_PAGE_ACTION_MAX: Max value
> + */
> +enum xe_ras_page_action {
> + XE_RAS_PAGE_ACTION_OFFLINE,
> + XE_RAS_PAGE_ACTION_REMOVE,
> + XE_RAS_PAGE_ACTION_MAX
> +};
> +
> /**
> * enum xe_ras_recovery_action - RAS recovery actions
> *
> @@ -295,6 +308,28 @@ struct xe_ras_memory_error {
> u32 reserved2[10];
> } __packed;
>
> +/**
> + * struct xe_ras_page_offline_request - Request for page offline
> +command */ struct xe_ras_page_offline_request {
> + /** @page_address: Page address (4KB aligned) */
> + u64 page_address;
> + /** @action: Action to be performed, see &enum xe_ras_page_action
> */
> + u32 action;
> + /** @reserved: Reserved for future use */
> + u32 reserved;
> +} __packed;
> +
> +/**
> + * struct xe_ras_page_offline_response - Response from page offline
> +command */ struct xe_ras_page_offline_response {
> + /** @status: Status of the page offline request */
> + u32 status;
> + /** @reserved: Reserved for future use */
> + u32 reserved;
> +} __packed;
> +
> /**
> * struct xe_ras_get_health_request - Request structure for obtaining gpu
> health
> */
> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> index c236e5377f30..a01576bf2e73 100644
> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> @@ -30,6 +30,7 @@ enum xe_sysctrl_group {
> * @XE_SYSCTRL_CMD_GET_THRESHOLD: Retrieve error threshold
> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
> + * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a
> + page
> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
> */
> @@ -40,6 +41,7 @@ enum xe_sysctrl_gfsp_cmd {
> XE_SYSCTRL_CMD_GET_THRESHOLD = 0x05,
> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
> + XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
> };
> --
> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* RE: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
2026-09-28 6:35 ` sashiko-bot
@ 2026-10-01 11:24 ` Upadhyay, Tejas
2026-10-01 11:37 ` Tauro, Riana
2026-10-01 12:11 ` Ghimiray, Himal Prasad
2 siblings, 1 reply; 28+ messages in thread
From: Upadhyay, Tejas @ 2026-10-01 11:24 UTC (permalink / raw)
To: Tauro, Riana, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
> -----Original Message-----
> From: Tauro, Riana <riana.tauro@intel.com>
> Sent: 28 September 2026 11:49
> To: intel-xe@lists.freedesktop.org
> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
> <himal.prasad.ghimiray@intel.com>
> Subject: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline
> queue and list
>
> Add support to query page offline list and queue from firmware during
> module load. The page offline list command retrieves pages that are already
> offlined by the firmware. The page offline queue command retrieves the pages
> pending to be offlined by the firmware.
>
> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: rebase
> store total pages once per response (Sashiko)
>
> v3: common function for offline and queue (Himal)
> ---
> drivers/gpu/drm/xe/xe_ras.c | 74 +++++++++++++++++++
> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++++++
> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 4 +
> 3 files changed, 113 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
> 1225c561a872..752754f09314 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -335,6 +335,77 @@ static bool ras_counter_is_valid(struct xe_device
> *xe, struct xe_ras_error_class
> return true;
> }
>
> +static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t
> req_size,
> + void *resp, size_t resp_size,
> + struct xe_ras_offline_common *common, bool
> offline) {
> + struct xe_sysctrl_mailbox_command command = {0};
> + struct xe_ras_offline_list_request *list_req;
> + u32 total_pages = 0, count = 0;
> + ssize_t rlen;
size_t rlen;
> + int ret, i;
> +
> + list_req = req ? req : NULL;
equivalent to list_req = req
> +
> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP,
> cmd, req, req_size, resp,
> + resp_size);
> +
> + do {
> + memset(resp, 0, resp_size);
> +
> + if (list_req)
> + list_req->index = count;
> +
> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
> + if (ret) {
> + xe_log_err(xe, SYSCTRL, ret, "failed to get page offline
> data, cmd=%#x\n",
> + cmd);
> + return;
> + }
> +
> + if (rlen != resp_size) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "unexpected page offline response length
> %zu (expected %zu), cmd=%#x\n",
> + rlen, resp_size, cmd);
> + return;
> + }
> +
> + for (i = 0; i < common->pages_returned && i <
> XE_RAS_NUM_PAGES; i++)
> + handle_page_offline(xe, common->page_addresses[i],
> offline);
> +
> + count += common->pages_returned;
> + if (!common->pages_returned)
> + break;
> +
> + if (!total_pages)
> + total_pages = common->total_pages;
> +
> + if (count > total_pages) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "Pages returned exceed total pages %u,
> returned %u, cmd=%#x\n",
> + total_pages, count, cmd);
> + return;
> + }
> + } while (common->additional_data);
> +}
> +
> +static void get_queued_pages(struct xe_device *xe) {
> + struct xe_ras_offline_common response = {0};
> +
> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL,
> 0, &response,
> + sizeof(response), &response, true); }
> +
> +static void get_offlined_list(struct xe_device *xe) {
> + struct xe_ras_offline_list_response response = {0};
> + struct xe_ras_offline_list_request request = {0};
> +
> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST,
> &request, sizeof(request),
> + &response, sizeof(response), &response.common,
> false); }
Please consider following if it looks ok. To make it less confusing and naming it what it actually does,
/*
* Fetch the next batch of page addresses for @cmd from firmware and process
* each one locally via handle_page_offline(). @notify_fw controls whether
* firmware is told back (XE_SYSCTRL_CMD_PAGE_OFFLINE) once a page has been
* handled - see the two callers below for why that differs per source.
*/
static void xe_ras_process_offline_pages(struct xe_device *xe, u32 cmd, void *req,
size_t req_size, void *resp, size_t resp_size,
struct xe_ras_offline_common *common, bool notify_fw)
{
...
for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
handle_page_offline(xe, common->page_addresses[i], notify_fw);
...
}
/*
* Firmware's pending queue: addresses it hasn't finished offlining yet.
* Drain it and ack each page back so firmware can dequeue it.
*/
static void xe_ras_drain_offline_queue(struct xe_device *xe)
{
struct xe_ras_offline_common response = {0};
xe_ras_process_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL, 0,
&response, sizeof(response), &response, true);
}
/*
* Firmware's persisted (flash) list of already-offlined pages. Just replay
* them into local VRAM tracking on driver load; firmware already has them.
*/
static void xe_ras_restore_offlined_pages(struct xe_device *xe)
{
struct xe_ras_offline_list_response response = {0};
struct xe_ras_offline_list_request request = {0};
xe_ras_process_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request,
sizeof(request), &response, sizeof(response),
&response.common, false);
}
Tejas
> +
> static struct pci_dev *find_usp_dev(struct pci_dev *pdev) {
> struct pci_dev *vsp;
> @@ -1049,6 +1120,9 @@ void xe_ras_init(struct xe_device *xe)
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
> + get_queued_pages(xe);
> + get_offlined_list(xe);
> +
> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
> if (ret)
> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
> b/drivers/gpu/drm/xe/xe_ras_types.h
> index f119489bcdf2..021ffbd6d4e2 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -10,6 +10,7 @@
>
> #define XE_RAS_NUM_COUNTERS 16
> #define XE_RAS_NUM_ERROR_ARR 3
> +#define XE_RAS_NUM_PAGES 25
> /* Error bits in IEH global error status register */
> #define XE_RAS_SOC_IEH_PUNIT BIT(1)
> /* Device memory error categories */
> @@ -330,6 +331,40 @@ struct xe_ras_page_offline_response {
> u32 reserved;
> } __packed;
>
> +/**
> + * struct xe_ras_offline_common - Common structure for offline list and
> +queue */ struct xe_ras_offline_common {
> + /** @total_pages: Total number of queued pages */
> + u32 total_pages;
> + /** @pages_returned: Number of pages returned in this response */
> + u32 pages_returned;
> + /** @page_addresses: Array of page addresses (4KB aligned) */
> + u64 page_addresses[XE_RAS_NUM_PAGES];
> + /** @additional_data: Indicates if more data is available */
> + u8 additional_data;
> + /** @reserved: Reserved for future use */
> + u8 reserved[3];
> +} __packed;
> +
> +/**
> + * struct xe_ras_offline_list_request - Request for get offline list
> +command */ struct xe_ras_offline_list_request {
> + /** @index: Zero-based index into the offline page list */
> + u32 index;
> +} __packed;
> +
> +/**
> + * struct xe_ras_offline_list_response - Response from get offline list
> +command */ struct xe_ras_offline_list_response {
> + /** @max_entries: Total no of pages that can be stored in flash */
> + u32 max_entries;
> + /** @common: Common offline page information */
> + struct xe_ras_offline_common common;
> +} __packed;
> +
> /**
> * struct xe_ras_get_health_request - Request structure for obtaining gpu
> health
> */
> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> index a01576bf2e73..3a71ed446949 100644
> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> @@ -31,6 +31,8 @@ enum xe_sysctrl_group {
> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
> * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a
> page
> + * @XE_SYSCTRL_CMD_GET_OFFLINE_LIST: Retrieve list of all offlined
> + pages from flash
> + * @XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE: Retrieve list of offlined
> queued
> + pages from firmware
> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
> */
> @@ -42,6 +44,8 @@ enum xe_sysctrl_gfsp_cmd {
> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
> XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
> + XE_SYSCTRL_CMD_GET_OFFLINE_LIST = 0x09,
> + XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE = 0x0A,
> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
> };
> --
> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory ecc errors
2026-10-01 10:50 ` Upadhyay, Tejas
@ 2026-10-01 11:27 ` Tauro, Riana
0 siblings, 0 replies; 28+ messages in thread
From: Tauro, Riana @ 2026-10-01 11:27 UTC (permalink / raw)
To: Upadhyay, Tejas, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
On 01-10-2026 16:20, Upadhyay, Tejas wrote:
>
>> -----Original Message-----
>> From: Tauro, Riana <riana.tauro@intel.com>
>> Sent: 28 September 2026 11:49
>> To: intel-xe@lists.freedesktop.org
>> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
>> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
>> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
>> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
>> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
>> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
>> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
>> <himal.prasad.ghimiray@intel.com>
>> Subject: [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for
>> device memory ecc errors
>>
>> Add basic support for sending page offline/remove requests to system
>> controller and use it for device memory ECC error handling.
>> Pages that belong to critical BOs cannot be handled by offlining and require a
>> SBR (Secondary Bus Reset).
>> Pages that are configured for log-only handling are not marked as bad by
>> firmware.
>>
>> For all other valid page addresses, the first occurrence of error indicates a
>> poison error and the page is offlined only by software.
>> Firmware avoids permanently marking the page as bad. The second
>> occurrence of an error indicates a Double-bit ECC error and the firmware
>> permanently marks the page as bad.
>>
>> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
>> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
>> ---
>> v2: use ret in sigid logging (Mallesh)
>> remove additional log
>> use xe_assert (Michal)
>>
>> v3: align address to PAGE_SIZE (sashiko, Himal)
>> rename decline to remove
>> add more descriptive logs (Himal)
>> ---
>> drivers/gpu/drm/xe/xe_ras.c | 127 +++++++++++++++++-
>> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++
>> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 2 +
>> 3 files changed, 159 insertions(+), 5 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
>> 7a85735c57d5..1225c561a872 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.c
>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>> @@ -3,6 +3,8 @@
>> * Copyright © 2026 Intel Corporation
>> */
>>
>> +#include "xe_assert.h"
>> +#include "xe_bo.h"
>> #include "xe_configfs.h"
>> #include "xe_debugfs.h"
>> #include "xe_device.h"
>> @@ -16,6 +18,7 @@
>> #include "xe_sysctrl_event_types.h"
>> #include "xe_sysctrl_mailbox.h"
>> #include "xe_sysctrl_mailbox_types.h"
>> +#include "xe_ttm_vram_mgr.h"
>>
>> #define CORE_COMPUTE_UNCORR_TYPE GENMASK(26, 25)
>> /*
>> @@ -201,6 +204,119 @@ static inline const char *comp_to_str(u8
>> component)
>> return xe_ras_components[component];
>> }
>>
>> +static int send_page_offline_cmd(struct xe_device *xe, u64 page_address,
>> + enum xe_ras_page_action action)
>> +{
>> + struct xe_sysctrl_mailbox_command command = {0};
>> + struct xe_ras_page_offline_request request = {0};
>> + struct xe_ras_page_offline_response response = {0};
>> + size_t rlen;
>> + int ret;
>> +
>> + if (!xe->info.has_sysctrl)
>> + return 0;
>> +
>> + xe_assert(xe, action < XE_RAS_PAGE_ACTION_MAX);
>> +
>> + request.page_address = page_address;
>> + request.action = action;
>> +
>> + if (action == XE_RAS_PAGE_ACTION_OFFLINE)
>> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to
>> offline page 0x%llx\n",
>> + page_address);
>> + else
>> + xe_log_err(xe, DEVICE_MEMORY, 0, "Requesting firmware to
>> remove page 0x%llx from queue\n",
>> + page_address);
>> +
>> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP,
>> XE_SYSCTRL_CMD_PAGE_OFFLINE,
>> + &request, sizeof(request), &response,
>> sizeof(response));
>> +
>> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
>> + if (ret) {
>> + xe_log_err(xe, SYSCTRL, ret, "failed to send page offline
>> command\n");
>> + return ret;
>> + }
>> +
>> + if (rlen != sizeof(response)) {
>> + xe_log_err(xe, SYSCTRL, -EINVAL,
>> + "unexpected page offline response length %zu
>> (expected %zu)\n",
>> + rlen, sizeof(response));
>> + return -EINVAL;
>> + }
>> +
>> + ret = ras_status_to_errno(response.status);
>> + if (ret)
>> + xe_log_err(xe, SYSCTRL, ret, "page offline command failed with
>> status %u\n",
>> + response.status);
>> +
>> + return ret;
>> +}
>> +
>> +static int handle_page_offline(struct xe_device *xe, u64 page_address,
>> +bool send_cmd) {
>> + enum xe_ras_page_action action;
>> + u64 addr;
>> + int ret = 0;
>> +
>> + if (!IS_ALIGNED(page_address, XE_PAGE_SIZE)) {
> xe_ttm_vram_handle_addr_fault() itself asserts against PAGE_SIZE, so it'd be more consistent to validate against PAGE_SIZE here too.
>
> Also, please have a look at Sashiko comments if applicable.
This is firmware response check. The response from firmware should be 4k
aligned.
As suggested by Himal i have aligned the address to the nearest
page_size below to be consistent with xe_ttm_vram_handle_addr_fault
For sashiko comments
[Severity: Medium]When xe_ttm_vram_handle_addr_fault() encounters an
out-of-bounds address, does it also return -EOPNOTSUPP?
Will add a configfs check before calling xe_ttm_vram_handle_addr_fault.
[Severity: High]Does returning 0 here mask the error code from
xe_ttm_vram_handle_addr_fault()? For -ENOMEM, i can add a ret. But for
the rest of the errors, triggering sbr or wedging is not required,
Thanks Riana
>
> Tejas
>
>> + xe_log_err(xe, SYSCTRL, -EINVAL, "Unaligned physical page
>> address: 0x%llx\n",
>> + page_address);
>> + return -EINVAL;
>> + }
>> +
>> + addr = ALIGN_DOWN(page_address, PAGE_SIZE);
>> +
>> + ret = xe_ttm_vram_handle_addr_fault(xe, addr);
>> +
>> + /*
>> + * Handle return code from address fault handling function:
>> + * 0: Page softofflined, remove from firmware queue
>> + * -EIO: Address belongs to a critical BO/stolen area that cannot be
>> offlined
>> + * -EOPNOTSUPP: Address is valid and can be offlined but user policy is
>> not to offline
>> + * -EEXIST: Address is soft offlined but yet to be offlined by firmware
>> for second
>> + * occurrence
>> + */
> -ENOMEM - allocation failure; next action is reset. --> we should reset on this error.
>
>> +
>> + switch (ret) {
>> + case 0:
>> + action = XE_RAS_PAGE_ACTION_REMOVE;
>> + xe_log_err(xe, DEVICE_MEMORY, ret,
>> + "Poison detected at physical address 0x%llx, page
>> soft-offlined\n",
>> + page_address);
>> + break;
>> + /* User policy set to decline page offlining */
>> + case -EOPNOTSUPP:
>> + action = XE_RAS_PAGE_ACTION_REMOVE;
>> + xe_log_err(xe, DEVICE_MEMORY, ret,
>> + "Poison detected at physical address 0x%llx, user
>> policy set to decline soft-offlining\n",
>> + page_address);
>> + break;
>> + case -EIO:
>> + xe_log_err(xe, DEVICE_MEMORY, ret,
>> + "Poison detected at physical address 0x%llx, page
>> belongs to critical BO and cannot be soft-offlined\n",
>> + page_address);
>> + return ret;
>> + case -EEXIST:
>> + action = XE_RAS_PAGE_ACTION_OFFLINE;
>> + xe_log_err(xe, DEVICE_MEMORY, ret,
>> + "Double-bit ECC error detected at physical address
>> 0x%llx, page soft-offlined\n",
>> + page_address);
>> + break;
>> + default:
>> + xe_log_err(xe, DEVICE_MEMORY, ret, "Failed to handle
>> address fault at physical address 0x%llx\n",
>> + page_address);
>> + return 0;
>> + }
>> +
>> + if (send_cmd) {
>> + ret = send_page_offline_cmd(xe, page_address, action);
>> + if (ret)
>> + return ret;
>> + }
>> +
>> + return 0;
>> +}
>> +
>> static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class
>> *counter) {
>> u8 severity = counter->common.severity; @@ -368,11 +484,12 @@
>> static u8 handle_soc_internal_errors(struct xe_device *xe, struct
>> xe_ras_error_a static u8 handle_device_memory_errors(struct xe_device *xe,
>> struct xe_ras_error_array *arr) {
>> struct xe_ras_memory_error *info = (void *)arr->details;
>> + int ret;
>>
>> /*
>> * For memory errors, the recovery action depends on the error
>> category
>> *
>> - * TODO: Double-bit ECC errors: Page offlining
>> + * Double-bit ECC errors: Page offlining
>> * Poison and data parity errors: Log only
>> * For any other memory errors, request a reset as recovery
>> mechanism
>> */
>> @@ -384,10 +501,10 @@ static u8 handle_device_memory_errors(struct
>> xe_device *xe, struct xe_ras_error_
>> xe_info(xe, "[RAS]: Data parity error detected\n");
>> break;
>> case XE_RAS_MEMORY_DB_ECC:
>> - xe_info(xe, "[RAS]: Double-bit ECC error detected at sw
>> address 0x%llx\n",
>> - info->sw_address);
>> - /* TODO: Add page offlining for Double-bit ECC error */
>> - fallthrough;
>> + ret = handle_page_offline(xe, info->sw_address, true);
>> + if (ret)
>> + return XE_RAS_RECOVERY_ACTION_RESET;
>> + break;
>> default:
>> return XE_RAS_RECOVERY_ACTION_RESET;
>> }
>> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
>> b/drivers/gpu/drm/xe/xe_ras_types.h
>> index fe6f3658a2a4..f119489bcdf2 100644
>> --- a/drivers/gpu/drm/xe/xe_ras_types.h
>> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
>> @@ -17,6 +17,19 @@
>> #define XE_RAS_MEMORY_POISON BIT(2)
>> #define XE_RAS_MEMORY_DATA_PARITY BIT(5)
>>
>> +/**
>> + * enum xe_ras_page_action - Page offline actions for page offline
>> +request
>> + *
>> + * @XE_RAS_PAGE_ACTION_OFFLINE: Instruct firmware to offline the page
>> + * @XE_RAS_PAGE_ACTION_REMOVE: Instruct firmware to remove the page
>> +from queue
>> + * @XE_RAS_PAGE_ACTION_MAX: Max value
>> + */
>> +enum xe_ras_page_action {
>> + XE_RAS_PAGE_ACTION_OFFLINE,
>> + XE_RAS_PAGE_ACTION_REMOVE,
>> + XE_RAS_PAGE_ACTION_MAX
>> +};
>> +
>> /**
>> * enum xe_ras_recovery_action - RAS recovery actions
>> *
>> @@ -295,6 +308,28 @@ struct xe_ras_memory_error {
>> u32 reserved2[10];
>> } __packed;
>>
>> +/**
>> + * struct xe_ras_page_offline_request - Request for page offline
>> +command */ struct xe_ras_page_offline_request {
>> + /** @page_address: Page address (4KB aligned) */
>> + u64 page_address;
>> + /** @action: Action to be performed, see &enum xe_ras_page_action
>> */
>> + u32 action;
>> + /** @reserved: Reserved for future use */
>> + u32 reserved;
>> +} __packed;
>> +
>> +/**
>> + * struct xe_ras_page_offline_response - Response from page offline
>> +command */ struct xe_ras_page_offline_response {
>> + /** @status: Status of the page offline request */
>> + u32 status;
>> + /** @reserved: Reserved for future use */
>> + u32 reserved;
>> +} __packed;
>> +
>> /**
>> * struct xe_ras_get_health_request - Request structure for obtaining gpu
>> health
>> */
>> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> index c236e5377f30..a01576bf2e73 100644
>> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> @@ -30,6 +30,7 @@ enum xe_sysctrl_group {
>> * @XE_SYSCTRL_CMD_GET_THRESHOLD: Retrieve error threshold
>> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
>> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
>> + * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a
>> + page
>> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
>> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
>> */
>> @@ -40,6 +41,7 @@ enum xe_sysctrl_gfsp_cmd {
>> XE_SYSCTRL_CMD_GET_THRESHOLD = 0x05,
>> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
>> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
>> + XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
>> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
>> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
>> };
>> --
>> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* RE: [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store
2026-09-28 6:18 ` [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store Riana Tauro
2026-09-28 6:43 ` sashiko-bot
@ 2026-10-01 11:36 ` Upadhyay, Tejas
2026-10-01 11:40 ` Tauro, Riana
1 sibling, 1 reply; 28+ messages in thread
From: Upadhyay, Tejas @ 2026-10-01 11:36 UTC (permalink / raw)
To: Tauro, Riana, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
> -----Original Message-----
> From: Tauro, Riana <riana.tauro@intel.com>
> Sent: 28 September 2026 11:49
> To: intel-xe@lists.freedesktop.org
> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
> <himal.prasad.ghimiray@intel.com>
> Subject: [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages
> firmware can store
>
> Add function to get maximum number of pages that firmware can store for
> offline tracking. This will be used to report max pages to userspace.
>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> ---
> v2: fix commit message (Himal)
> ---
> drivers/gpu/drm/xe/xe_ras.c | 15 +++++++++++++++
> drivers/gpu/drm/xe/xe_ras.h | 1 +
> drivers/gpu/drm/xe/xe_ras_types.h | 2 ++
> 3 files changed, 18 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
> 9a1a8190578b..cfc7f07e9ced 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -401,9 +401,13 @@ static void get_offlined_list(struct xe_device *xe) {
> struct xe_ras_offline_list_response response = {0};
> struct xe_ras_offline_list_request request = {0};
> + struct xe_ras_state *state = &xe->ras.state;
>
> get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST,
> &request, sizeof(request),
> &response, sizeof(response), &response.common,
> false);
Isnt it better to snapshot max_entries into a local in get_offlined_list() as soon as the first response comes back, not after the whole paginated fetch completes. I mean better to not depend on "whatever's left in resp when the function returns."
Tejas
> +
> + if (response.max_entries)
> + state->max_pages = response.max_entries;
> }
>
> static struct pci_dev *find_usp_dev(struct pci_dev *pdev) @@ -1105,6
> +1109,17 @@ bool xe_ras_get_disable_page_offline(struct xe_device *xe)
> return xe->ras.state.disable_page_offline;
> }
>
> +/**
> + * xe_ras_get_max_pages - Get the maximum number of pages
> + * @xe: xe device instance
> + *
> + * Return: Maximum number of pages that can be stored in flash */
> +u32 xe_ras_get_max_pages(struct xe_device *xe) {
> + return xe->ras.state.max_pages;
> +}
> +
> /**
> * xe_ras_init - Initialize Xe RAS
> * @xe: xe device instance
> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h index
> ac1253d6ead9..561c656e2ad3 100644
> --- a/drivers/gpu/drm/xe/xe_ras.h
> +++ b/drivers/gpu/drm/xe/xe_ras.h
> @@ -22,5 +22,6 @@ int xe_ras_set_threshold(struct xe_device *xe, u8
> severity, u8 component, u32 th void xe_ras_init(struct xe_device *xe); enum
> xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe); bool
> xe_ras_get_disable_page_offline(struct xe_device *xe);
> +u32 xe_ras_get_max_pages(struct xe_device *xe);
>
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
> b/drivers/gpu/drm/xe/xe_ras_types.h
> index 40224db38906..e09b50a6f77b 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -415,5 +415,7 @@ struct xe_ras_set_health_response { struct
> xe_ras_state {
> /** @disable_page_offline: cached configfs policy, immutable after init
> */
> bool disable_page_offline;
> + /** @max_pages: Total number of pages that can be stored by
> firmware */
> + u32 max_pages;
> };
> #endif
> --
> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* RE: [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace
2026-09-28 6:18 ` [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace Riana Tauro
@ 2026-10-01 11:37 ` Upadhyay, Tejas
0 siblings, 0 replies; 28+ messages in thread
From: Upadhyay, Tejas @ 2026-10-01 11:37 UTC (permalink / raw)
To: Tauro, Riana, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
> -----Original Message-----
> From: Tauro, Riana <riana.tauro@intel.com>
> Sent: 28 September 2026 11:49
> To: intel-xe@lists.freedesktop.org
> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
> <himal.prasad.ghimiray@intel.com>
> Subject: [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by
> firmware to userspace
>
> Report the maximum number of pages returned by firmware at probe,
> resolving the existing TODO.
>
> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> ---
> v2: fix commit message (Himal)
> ---
> drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 6 +-----
> drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h | 2 --
> 2 files changed, 1 insertion(+), 7 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> index 2c4722a956a0..cc33ba7e23c2 100644
> --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> @@ -1012,11 +1012,7 @@ static int vram_bad_pages_show(struct seq_file
> *m, void *unused)
> struct xe_tile *tile;
> u8 id;
>
> - man = ttm_manager_type(&xe->ttm, XE_PL_VRAM0);
> - if (man)
> - /* TODO Hook with RAS to show max_pages fetched from FW
> */
> - seq_printf(m, "max_pages: %d\n",
> - to_xe_ttm_vram_mgr(man)->max_pages);
> + seq_printf(m, "max_pages: %u\n", xe_ras_get_max_pages(xe));
Reviewed-by: Tejas Upadhyay <tejas.upadhyay@intel.com>
Tejas
>
> for_each_tile(tile, xe, id) {
> struct xe_vram_region *vr = tile->mem.vram; diff --git
> a/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
> b/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
> index efcf3e1d4e80..dc97b0ad0e51 100644
> --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
> +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr_types.h
> @@ -37,8 +37,6 @@ struct xe_ttm_vram_mgr {
> struct mutex lock;
> /** @mem_type: The TTM memory type */
> u32 mem_type;
> - /** @max_pages: max pages that can be in offline queue retrieved
> from FW */
> - u16 max_pages;
> };
>
> /**
> --
> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list
2026-10-01 11:24 ` Upadhyay, Tejas
@ 2026-10-01 11:37 ` Tauro, Riana
0 siblings, 0 replies; 28+ messages in thread
From: Tauro, Riana @ 2026-10-01 11:37 UTC (permalink / raw)
To: Upadhyay, Tejas, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
On 01-10-2026 16:54, Upadhyay, Tejas wrote:
>
>> -----Original Message-----
>> From: Tauro, Riana <riana.tauro@intel.com>
>> Sent: 28 September 2026 11:49
>> To: intel-xe@lists.freedesktop.org
>> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
>> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
>> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
>> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
>> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
>> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
>> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
>> <himal.prasad.ghimiray@intel.com>
>> Subject: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline
>> queue and list
>>
>> Add support to query page offline list and queue from firmware during
>> module load. The page offline list command retrieves pages that are already
>> offlined by the firmware. The page offline queue command retrieves the pages
>> pending to be offlined by the firmware.
>>
>> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
>> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
>> ---
>> v2: rebase
>> store total pages once per response (Sashiko)
>>
>> v3: common function for offline and queue (Himal)
>> ---
>> drivers/gpu/drm/xe/xe_ras.c | 74 +++++++++++++++++++
>> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++++++
>> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 4 +
>> 3 files changed, 113 insertions(+)
>>
>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
>> 1225c561a872..752754f09314 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.c
>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>> @@ -335,6 +335,77 @@ static bool ras_counter_is_valid(struct xe_device
>> *xe, struct xe_ras_error_class
>> return true;
>> }
>>
>> +static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t
>> req_size,
>> + void *resp, size_t resp_size,
>> + struct xe_ras_offline_common *common, bool
>> offline) {
>> + struct xe_sysctrl_mailbox_command command = {0};
>> + struct xe_ras_offline_list_request *list_req;
>> + u32 total_pages = 0, count = 0;
>> + ssize_t rlen;
> size_t rlen;
Sure will fix
>
>> + int ret, i;
>> +
>> + list_req = req ? req : NULL;
> equivalent to list_req = req
I had initially added all conversions. missed this while removing . yeah
will fix it
>
>> +
>> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP,
>> cmd, req, req_size, resp,
>> + resp_size);
>> +
>> + do {
>> + memset(resp, 0, resp_size);
>> +
>> + if (list_req)
>> + list_req->index = count;
>> +
>> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
>> + if (ret) {
>> + xe_log_err(xe, SYSCTRL, ret, "failed to get page offline
>> data, cmd=%#x\n",
>> + cmd);
>> + return;
>> + }
>> +
>> + if (rlen != resp_size) {
>> + xe_log_err(xe, SYSCTRL, -EINVAL,
>> + "unexpected page offline response length
>> %zu (expected %zu), cmd=%#x\n",
>> + rlen, resp_size, cmd);
>> + return;
>> + }
>> +
>> + for (i = 0; i < common->pages_returned && i <
>> XE_RAS_NUM_PAGES; i++)
>> + handle_page_offline(xe, common->page_addresses[i],
>> offline);
>> +
>> + count += common->pages_returned;
>> + if (!common->pages_returned)
>> + break;
>> +
>> + if (!total_pages)
>> + total_pages = common->total_pages;
>> +
>> + if (count > total_pages) {
>> + xe_log_err(xe, SYSCTRL, -EINVAL,
>> + "Pages returned exceed total pages %u,
>> returned %u, cmd=%#x\n",
>> + total_pages, count, cmd);
>> + return;
>> + }
>> + } while (common->additional_data);
>> +}
>> +
>> +static void get_queued_pages(struct xe_device *xe) {
>> + struct xe_ras_offline_common response = {0};
>> +
>> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL,
>> 0, &response,
>> + sizeof(response), &response, true); }
>> +
>> +static void get_offlined_list(struct xe_device *xe) {
>> + struct xe_ras_offline_list_response response = {0};
>> + struct xe_ras_offline_list_request request = {0};
>> +
>> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST,
>> &request, sizeof(request),
>> + &response, sizeof(response), &response.common,
>> false); }
> Please consider following if it looks ok. To make it less confusing and naming it what it actually does,
>
> /*
> * Fetch the next batch of page addresses for @cmd from firmware and process
> * each one locally via handle_page_offline(). @notify_fw controls whether
> * firmware is told back (XE_SYSCTRL_CMD_PAGE_OFFLINE) once a page has been
> * handled - see the two callers below for why that differs per source.
> */
> static void xe_ras_process_offline_pages(struct xe_device *xe, u32 cmd, void *req,
> size_t req_size, void *resp, size_t resp_size,
> struct xe_ras_offline_common *common, bool notify_fw)
process_offline_pages is indeed better than get_offline_pages. Will
avoid confusion
Will change it. File prefix is used in xe driver for non-static functions.
> {
> ...
> for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
> handle_page_offline(xe, common->page_addresses[i], notify_fw);
> ...
> }
>
> /*
> * Firmware's pending queue: addresses it hasn't finished offlining yet.
> * Drain it and ack each page back so firmware can dequeue it.
> */
> static void xe_ras_drain_offline_queue(struct xe_device *xe)
> {
> struct xe_ras_offline_common response = {0};
>
> xe_ras_process_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL, 0,
> &response, sizeof(response), &response, true);
> }
>
> /*
> * Firmware's persisted (flash) list of already-offlined pages. Just replay
> * them into local VRAM tracking on driver load; firmware already has them.
> */
> static void xe_ras_restore_offlined_pages(struct xe_device *xe)
Drain does make sense for queue but restore doesn't for list. We are not
restoring the pages, they are still offlined.
Wouldn't it be better to just retain the command names. Adding
description is better . will add that in new rev.
Thanks
Riana
> {
> struct xe_ras_offline_list_response response = {0};
> struct xe_ras_offline_list_request request = {0};
>
> xe_ras_process_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request,
> sizeof(request), &response, sizeof(response),
> &response.common, false);
> }
>
> Tejas
>> +
>> static struct pci_dev *find_usp_dev(struct pci_dev *pdev) {
>> struct pci_dev *vsp;
>> @@ -1049,6 +1120,9 @@ void xe_ras_init(struct xe_device *xe)
>> if (IS_ENABLED(CONFIG_PCIEAER))
>> ras_usp_aer_init(xe);
>>
>> + get_queued_pages(xe);
>> + get_offlined_list(xe);
>> +
>> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
>> if (ret)
>> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
>> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
>> b/drivers/gpu/drm/xe/xe_ras_types.h
>> index f119489bcdf2..021ffbd6d4e2 100644
>> --- a/drivers/gpu/drm/xe/xe_ras_types.h
>> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
>> @@ -10,6 +10,7 @@
>>
>> #define XE_RAS_NUM_COUNTERS 16
>> #define XE_RAS_NUM_ERROR_ARR 3
>> +#define XE_RAS_NUM_PAGES 25
>> /* Error bits in IEH global error status register */
>> #define XE_RAS_SOC_IEH_PUNIT BIT(1)
>> /* Device memory error categories */
>> @@ -330,6 +331,40 @@ struct xe_ras_page_offline_response {
>> u32 reserved;
>> } __packed;
>>
>> +/**
>> + * struct xe_ras_offline_common - Common structure for offline list and
>> +queue */ struct xe_ras_offline_common {
>> + /** @total_pages: Total number of queued pages */
>> + u32 total_pages;
>> + /** @pages_returned: Number of pages returned in this response */
>> + u32 pages_returned;
>> + /** @page_addresses: Array of page addresses (4KB aligned) */
>> + u64 page_addresses[XE_RAS_NUM_PAGES];
>> + /** @additional_data: Indicates if more data is available */
>> + u8 additional_data;
>> + /** @reserved: Reserved for future use */
>> + u8 reserved[3];
>> +} __packed;
>> +
>> +/**
>> + * struct xe_ras_offline_list_request - Request for get offline list
>> +command */ struct xe_ras_offline_list_request {
>> + /** @index: Zero-based index into the offline page list */
>> + u32 index;
>> +} __packed;
>> +
>> +/**
>> + * struct xe_ras_offline_list_response - Response from get offline list
>> +command */ struct xe_ras_offline_list_response {
>> + /** @max_entries: Total no of pages that can be stored in flash */
>> + u32 max_entries;
>> + /** @common: Common offline page information */
>> + struct xe_ras_offline_common common;
>> +} __packed;
>> +
>> /**
>> * struct xe_ras_get_health_request - Request structure for obtaining gpu
>> health
>> */
>> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> index a01576bf2e73..3a71ed446949 100644
>> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
>> @@ -31,6 +31,8 @@ enum xe_sysctrl_group {
>> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
>> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
>> * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a
>> page
>> + * @XE_SYSCTRL_CMD_GET_OFFLINE_LIST: Retrieve list of all offlined
>> + pages from flash
>> + * @XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE: Retrieve list of offlined
>> queued
>> + pages from firmware
>> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
>> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
>> */
>> @@ -42,6 +44,8 @@ enum xe_sysctrl_gfsp_cmd {
>> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
>> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
>> XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
>> + XE_SYSCTRL_CMD_GET_OFFLINE_LIST = 0x09,
>> + XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE = 0x0A,
>> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
>> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
>> };
>> --
>> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store
2026-10-01 11:36 ` Upadhyay, Tejas
@ 2026-10-01 11:40 ` Tauro, Riana
0 siblings, 0 replies; 28+ messages in thread
From: Tauro, Riana @ 2026-10-01 11:40 UTC (permalink / raw)
To: Upadhyay, Tejas, intel-xe@lists.freedesktop.org
Cc: Gupta, Anshuman, Vivi, Rodrigo,
aravind.iddamsetty@linux.intel.com, Nilawar, Badal, Jadav, Raag,
Koppuravuri, Ravi Kishore, Koujalagi, Mallesh,
Ghimiray, Himal Prasad
On 01-10-2026 17:06, Upadhyay, Tejas wrote:
>
>> -----Original Message-----
>> From: Tauro, Riana <riana.tauro@intel.com>
>> Sent: 28 September 2026 11:49
>> To: intel-xe@lists.freedesktop.org
>> Cc: Tauro, Riana <riana.tauro@intel.com>; Gupta, Anshuman
>> <anshuman.gupta@intel.com>; Vivi, Rodrigo <rodrigo.vivi@intel.com>;
>> aravind.iddamsetty@linux.intel.com; Nilawar, Badal
>> <badal.nilawar@intel.com>; Jadav, Raag <raag.jadav@intel.com>;
>> Koppuravuri, Ravi Kishore <ravi.kishore.koppuravuri@intel.com>; Koujalagi,
>> Mallesh <mallesh.koujalagi@intel.com>; Upadhyay, Tejas
>> <tejas.upadhyay@intel.com>; Ghimiray, Himal Prasad
>> <himal.prasad.ghimiray@intel.com>
>> Subject: [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages
>> firmware can store
>>
>> Add function to get maximum number of pages that firmware can store for
>> offline tracking. This will be used to report max pages to userspace.
>>
>> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
>> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>> ---
>> v2: fix commit message (Himal)
>> ---
>> drivers/gpu/drm/xe/xe_ras.c | 15 +++++++++++++++
>> drivers/gpu/drm/xe/xe_ras.h | 1 +
>> drivers/gpu/drm/xe/xe_ras_types.h | 2 ++
>> 3 files changed, 18 insertions(+)
>>
>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index
>> 9a1a8190578b..cfc7f07e9ced 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.c
>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>> @@ -401,9 +401,13 @@ static void get_offlined_list(struct xe_device *xe) {
>> struct xe_ras_offline_list_response response = {0};
>> struct xe_ras_offline_list_request request = {0};
>> + struct xe_ras_state *state = &xe->ras.state;
>>
>> get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST,
>> &request, sizeof(request),
>> &response, sizeof(response), &response.common,
>> false);
> Isnt it better to snapshot max_entries into a local in get_offlined_list() as soon as the first response comes back, not after the whole paginated fetch completes. I mean better to not depend on "whatever's left in resp when the function returns."
yeah i am adding this for next rev. The last response can be 0s cannot
be guarantee to have max entries
Thanks
Riana
>
> Tejas
>> +
>> + if (response.max_entries)
>> + state->max_pages = response.max_entries;
>> }
>>
>> static struct pci_dev *find_usp_dev(struct pci_dev *pdev) @@ -1105,6
>> +1109,17 @@ bool xe_ras_get_disable_page_offline(struct xe_device *xe)
>> return xe->ras.state.disable_page_offline;
>> }
>>
>> +/**
>> + * xe_ras_get_max_pages - Get the maximum number of pages
>> + * @xe: xe device instance
>> + *
>> + * Return: Maximum number of pages that can be stored in flash */
>> +u32 xe_ras_get_max_pages(struct xe_device *xe) {
>> + return xe->ras.state.max_pages;
>> +}
>> +
>> /**
>> * xe_ras_init - Initialize Xe RAS
>> * @xe: xe device instance
>> diff --git a/drivers/gpu/drm/xe/xe_ras.h b/drivers/gpu/drm/xe/xe_ras.h index
>> ac1253d6ead9..561c656e2ad3 100644
>> --- a/drivers/gpu/drm/xe/xe_ras.h
>> +++ b/drivers/gpu/drm/xe/xe_ras.h
>> @@ -22,5 +22,6 @@ int xe_ras_set_threshold(struct xe_device *xe, u8
>> severity, u8 component, u32 th void xe_ras_init(struct xe_device *xe); enum
>> xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe); bool
>> xe_ras_get_disable_page_offline(struct xe_device *xe);
>> +u32 xe_ras_get_max_pages(struct xe_device *xe);
>>
>> #endif
>> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h
>> b/drivers/gpu/drm/xe/xe_ras_types.h
>> index 40224db38906..e09b50a6f77b 100644
>> --- a/drivers/gpu/drm/xe/xe_ras_types.h
>> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
>> @@ -415,5 +415,7 @@ struct xe_ras_set_health_response { struct
>> xe_ras_state {
>> /** @disable_page_offline: cached configfs policy, immutable after init
>> */
>> bool disable_page_offline;
>> + /** @max_pages: Total number of pages that can be stored by
>> firmware */
>> + u32 max_pages;
>> };
>> #endif
>> --
>> 2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
2026-09-28 6:35 ` sashiko-bot
2026-10-01 11:24 ` Upadhyay, Tejas
@ 2026-10-01 12:11 ` Ghimiray, Himal Prasad
2 siblings, 0 replies; 28+ messages in thread
From: Ghimiray, Himal Prasad @ 2026-10-01 12:11 UTC (permalink / raw)
To: Riana Tauro, intel-xe
Cc: anshuman.gupta, rodrigo.vivi, aravind.iddamsetty, badal.nilawar,
raag.jadav, ravi.kishore.koppuravuri, mallesh.koujalagi,
tejas.upadhyay
On 28-09-2026 11:48, Riana Tauro wrote:
> Add support to query page offline list and queue from firmware
> during module load. The page offline list command retrieves pages that
> are already offlined by the firmware. The page offline queue command
> retrieves the pages pending to be offlined by the firmware.
>
> Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
> Signed-off-by: Riana Tauro <riana.tauro@intel.com>
> ---
> v2: rebase
> store total pages once per response (Sashiko)
>
> v3: common function for offline and queue (Himal)
> ---
> drivers/gpu/drm/xe/xe_ras.c | 74 +++++++++++++++++++
> drivers/gpu/drm/xe/xe_ras_types.h | 35 +++++++++
> drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h | 4 +
> 3 files changed, 113 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 1225c561a872..752754f09314 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -335,6 +335,77 @@ static bool ras_counter_is_valid(struct xe_device *xe, struct xe_ras_error_class
> return true;
> }
>
> +static void get_offline_pages(struct xe_device *xe, u32 cmd, void *req, size_t req_size,
> + void *resp, size_t resp_size,
> + struct xe_ras_offline_common *common, bool offline)
> +{
> + struct xe_sysctrl_mailbox_command command = {0};
> + struct xe_ras_offline_list_request *list_req;
> + u32 total_pages = 0, count = 0;
> + ssize_t rlen;
> + int ret, i;
> +
> + list_req = req ? req : NULL;
> +
> + xe_sysctrl_create_command(&command, XE_SYSCTRL_GROUP_GFSP, cmd, req, req_size, resp,
> + resp_size);
> +
> + do {
> + memset(resp, 0, resp_size);
> +
> + if (list_req)
> + list_req->index = count;
> +
> + ret = xe_sysctrl_send_command(&xe->sc, &command, &rlen);
> + if (ret) {
> + xe_log_err(xe, SYSCTRL, ret, "failed to get page offline data, cmd=%#x\n",
> + cmd);
> + return;
> + }
> +
> + if (rlen != resp_size) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "unexpected page offline response length %zu (expected %zu), cmd=%#x\n",
> + rlen, resp_size, cmd);
> + return;
> + }
> +
> + for (i = 0; i < common->pages_returned && i < XE_RAS_NUM_PAGES; i++)
> + handle_page_offline(xe, common->page_addresses[i], offline);
> +
> + count += common->pages_returned;
> + if (!common->pages_returned)
> + break;
> +
> + if (!total_pages)
> + total_pages = common->total_pages;
> +
> + if (count > total_pages) {
> + xe_log_err(xe, SYSCTRL, -EINVAL,
> + "Pages returned exceed total pages %u, returned %u, cmd=%#x\n",
> + total_pages, count, cmd);
> + return;
> + }
> + } while (common->additional_data);
> +}
> +
> +static void get_queued_pages(struct xe_device *xe)
> +{
> + struct xe_ras_offline_common response = {0};
> +
> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE, NULL, 0, &response,
> + sizeof(response), &response, true);
> +}
> +
> +static void get_offlined_list(struct xe_device *xe)
> +{
> + struct xe_ras_offline_list_response response = {0};
> + struct xe_ras_offline_list_request request = {0};
> +
> + get_offline_pages(xe, XE_SYSCTRL_CMD_GET_OFFLINE_LIST, &request, sizeof(request),
> + &response, sizeof(response), &response.common, false);
> +}
> +
> static struct pci_dev *find_usp_dev(struct pci_dev *pdev)
> {
> struct pci_dev *vsp;
> @@ -1049,6 +1120,9 @@ void xe_ras_init(struct xe_device *xe)
> if (IS_ENABLED(CONFIG_PCIEAER))
> ras_usp_aer_init(xe);
>
> + get_queued_pages(xe);
> + get_offlined_list(xe);
I believe these also needs to be called during resume in case of
suspend-resume flow.
> +
> ret = devm_device_add_group(xe->drm.dev, &gpu_health_group);
> if (ret)
> xe_err(xe, "Failed to create GPU health sysfs, err=%d\n", ret);
> diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h
> index f119489bcdf2..021ffbd6d4e2 100644
> --- a/drivers/gpu/drm/xe/xe_ras_types.h
> +++ b/drivers/gpu/drm/xe/xe_ras_types.h
> @@ -10,6 +10,7 @@
>
> #define XE_RAS_NUM_COUNTERS 16
> #define XE_RAS_NUM_ERROR_ARR 3
> +#define XE_RAS_NUM_PAGES 25
> /* Error bits in IEH global error status register */
> #define XE_RAS_SOC_IEH_PUNIT BIT(1)
> /* Device memory error categories */
> @@ -330,6 +331,40 @@ struct xe_ras_page_offline_response {
> u32 reserved;
> } __packed;
>
> +/**
> + * struct xe_ras_offline_common - Common structure for offline list and queue
> + */
> +struct xe_ras_offline_common {
> + /** @total_pages: Total number of queued pages */
> + u32 total_pages;
> + /** @pages_returned: Number of pages returned in this response */
> + u32 pages_returned;
> + /** @page_addresses: Array of page addresses (4KB aligned) */
> + u64 page_addresses[XE_RAS_NUM_PAGES];
> + /** @additional_data: Indicates if more data is available */
> + u8 additional_data;
> + /** @reserved: Reserved for future use */
> + u8 reserved[3];
> +} __packed;
> +
> +/**
> + * struct xe_ras_offline_list_request - Request for get offline list command
> + */
> +struct xe_ras_offline_list_request {
> + /** @index: Zero-based index into the offline page list */
> + u32 index;
> +} __packed;
> +
> +/**
> + * struct xe_ras_offline_list_response - Response from get offline list command
> + */
> +struct xe_ras_offline_list_response {
> + /** @max_entries: Total no of pages that can be stored in flash */
> + u32 max_entries;
> + /** @common: Common offline page information */
> + struct xe_ras_offline_common common;
> +} __packed;
> +
> /**
> * struct xe_ras_get_health_request - Request structure for obtaining gpu health
> */
> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> index a01576bf2e73..3a71ed446949 100644
> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox_types.h
> @@ -31,6 +31,8 @@ enum xe_sysctrl_group {
> * @XE_SYSCTRL_CMD_SET_THRESHOLD: Set error threshold
> * @XE_SYSCTRL_CMD_GET_PENDING_EVENT: Retrieve pending event
> * @XE_SYSCTRL_CMD_PAGE_OFFLINE: Instruct firmware to offline/remove a page
> + * @XE_SYSCTRL_CMD_GET_OFFLINE_LIST: Retrieve list of all offlined pages from flash
> + * @XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE: Retrieve list of offlined queued pages from firmware
> * @XE_SYSCTRL_CMD_GET_HEALTH: Retrieve gpu health
> * @XE_SYSCTRL_CMD_SET_HEALTH: Set gpu health
> */
> @@ -42,6 +44,8 @@ enum xe_sysctrl_gfsp_cmd {
> XE_SYSCTRL_CMD_SET_THRESHOLD = 0x06,
> XE_SYSCTRL_CMD_GET_PENDING_EVENT = 0x07,
> XE_SYSCTRL_CMD_PAGE_OFFLINE = 0x08,
> + XE_SYSCTRL_CMD_GET_OFFLINE_LIST = 0x09,
> + XE_SYSCTRL_CMD_GET_OFFLINE_QUEUE = 0x0A,
> XE_SYSCTRL_CMD_GET_HEALTH = 0x0B,
> XE_SYSCTRL_CMD_SET_HEALTH = 0x0C,
> };
^ permalink raw reply [flat|nested] 28+ messages in thread
end of thread, other threads:[~2026-10-01 12:11 UTC | newest]
Thread overview: 28+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 6:18 [PATCH v3 0/6] Add support to handle memory double-bit ecc errors Riana Tauro
2026-09-28 6:18 ` [PATCH v3 1/6] drm/xe/xe_ras: Handle page offline requests for device memory " Riana Tauro
2026-09-28 6:34 ` sashiko-bot
2026-09-28 8:55 ` Ghimiray, Himal Prasad
2026-10-01 10:50 ` Upadhyay, Tejas
2026-10-01 11:27 ` Tauro, Riana
2026-09-28 6:18 ` [PATCH v3 2/6] drm/xe/xe_ras: Add support to query page offline queue and list Riana Tauro
2026-09-28 6:35 ` sashiko-bot
2026-10-01 11:24 ` Upadhyay, Tejas
2026-10-01 11:37 ` Tauro, Riana
2026-10-01 12:11 ` Ghimiray, Himal Prasad
2026-09-28 6:18 ` [PATCH v3 3/6] drm/xe: Separate drm-ras netlink data from device and firmware RAS state Riana Tauro
2026-09-28 6:51 ` sashiko-bot
2026-09-28 8:59 ` Ghimiray, Himal Prasad
2026-09-28 6:18 ` [PATCH v3 4/6] drm/xe/xe_ras: Add function to get maximum pages firmware can store Riana Tauro
2026-09-28 6:43 ` sashiko-bot
2026-10-01 11:36 ` Upadhyay, Tejas
2026-10-01 11:40 ` Tauro, Riana
2026-09-28 6:18 ` [PATCH v3 5/6] drm/xe/xe_ttm_vram: Report max_pages reported by firmware to userspace Riana Tauro
2026-10-01 11:37 ` Upadhyay, Tejas
2026-09-28 6:18 ` [PATCH v3 6/6] drm/xe/xe_ras: Track offlined pages by firmware to avoid duplicates Riana Tauro
2026-09-28 7:07 ` sashiko-bot
2026-09-28 9:00 ` Ghimiray, Himal Prasad
2026-09-28 9:17 ` Ghimiray, Himal Prasad
2026-09-28 9:22 ` Tauro, Riana
2026-09-28 14:42 ` ✓ CI.KUnit: success for Add support to handle memory double-bit ecc errors (rev3) Patchwork
2026-09-28 15:27 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-28 17:40 ` ✓ Xe.CI.FULL: " Patchwork
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox