From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from gabe.freedesktop.org (gabe.freedesktop.org [131.252.210.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 7096DC88E7F for ; Wed, 16 Sep 2026 10:21:26 +0000 (UTC) Received: from gabe.freedesktop.org (localhost [127.0.0.1]) by gabe.freedesktop.org (Postfix) with ESMTP id 22DCB10E495; Wed, 16 Sep 2026 10:21:26 +0000 (UTC) Authentication-Results: gabe.freedesktop.org; dkim=pass (2048-bit key; unprotected) header.d=intel.com header.i=@intel.com header.b="KwxXRRD6"; dkim-atps=neutral Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.6]) by gabe.freedesktop.org (Postfix) with ESMTPS id 4747F10E495 for ; Wed, 16 Sep 2026 10:21:24 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1789554084; x=1821090084; h=date:from:to:cc:subject:message-id:references: mime-version:in-reply-to; bh=oAQYql8VscjKiE2KddR9LFMArPHDruovTvqMziya3JQ=; b=KwxXRRD6UifzMpxT6yvYthz/LanZON8mc6WkhlxcbPY2Nwr0tiCBKJen izJOJr4SXnqCauzOHEuTCIXVmW38919uwIAQsc9akRzgOB6hB61jX+hEA z+aj+bzc6fsdLETx4VJ83rDOY8H/qOUzVBYwz8qdC4g7y0jwN3YPD57Mr YMqafACevs6I/uwRqHnwCpvwHvdtn2RJ996oSKbIHSaHI0Ffjs0oaqFN7 6SMCYtmb6Za8vnUAYQMGzd6yOTh1dSbFsjfOxfRq6hYB9G4xefVXe6ByF A7+hQZ2gHZBzaY3sADw3yAjYXoBk18fIRrgxp4CVO3idDUaByC13B0/Ns A==; X-CSE-ConnectionGUID: oU++HW5vTfWaMr5wZLH3Qg== X-CSE-MsgGUID: sPtpKyCwSLmPej4zCKZuCA== X-IronPort-AV: E=McAfee;i="6800,10657,11905"; a="431958" X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="431958" Received: from orviesa002.jf.intel.com ([10.64.159.142]) by fmvoesa116.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 16 Sep 2026 03:21:24 -0700 X-CSE-ConnectionGUID: 4671Ith4SDueOaawHOFoVA== X-CSE-MsgGUID: 40FXllj0RimUtsNEUT5fYA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="303182060" Received: from black.igk.intel.com ([10.91.253.5]) by orviesa002.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 16 Sep 2026 03:21:22 -0700 Date: Wed, 16 Sep 2026 12:21:19 +0200 From: Raag Jadav To: Riana Tauro Cc: intel-xe@lists.freedesktop.org, anshuman.gupta@intel.com, rodrigo.vivi@intel.com, aravind.iddamsetty@linux.intel.com, badal.nilawar@intel.com, ravi.kishore.koppuravuri@intel.com, mallesh.koujalagi@intel.com Subject: Re: [PATCH v2 1/2] drm/xe/xe_ras: Add support for handling Fabric errors Message-ID: References: <20260916052836.2953509-4-riana.tauro@intel.com> <20260916052836.2953509-5-riana.tauro@intel.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260916052836.2953509-5-riana.tauro@intel.com> X-BeenThere: intel-xe@lists.freedesktop.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Intel Xe graphics driver List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: intel-xe-bounces@lists.freedesktop.org Sender: "Intel-xe" On Wed, Sep 16, 2026 at 10:58:38AM +0530, Riana Tauro wrote: > Add support for handling fabric errors. For System Agent Fabric errors > caused by data payload parity, log the error and return. For all other > errors or causes, request a Seconday Bus Reset(SBR). > > Signed-off-by: Riana Tauro > --- > v2: add mapping to cper severity (Badal, Rodrigo) > add abbreviations > add new line in header (Raag) > --- > drivers/gpu/drm/xe/xe_ras.c | 35 +++++++++++++++++++++++++++++++ > drivers/gpu/drm/xe/xe_ras_types.h | 6 ++++++ > 2 files changed, 41 insertions(+) > > diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c > index 7a85735c57d5..c1939fd71f2d 100644 > --- a/drivers/gpu/drm/xe/xe_ras.c > +++ b/drivers/gpu/drm/xe/xe_ras.c > @@ -185,6 +185,20 @@ static int ras_status_to_errno(u32 status) > } > } > > +static u8 ras_sev_to_cper_sev(u8 sev) Group these with existing *_to_severity() switcheroos, also use consistent naming (both helper and parameter). > +{ > + switch (sev) { > + case XE_RAS_SEV_CORRECTABLE: > + return CPER_SEV_CORRECTED; > + case XE_RAS_SEV_UNCORRECTABLE: > + return CPER_SEV_RECOVERABLE; > + case XE_RAS_SEV_INFORMATIONAL: > + return CPER_SEV_INFORMATIONAL; > + default: > + return CPER_SEV_RECOVERABLE; > + } > +} > + > static inline const char *sev_to_str(u8 severity) > { > if (severity >= XE_RAS_SEV_MAX) > @@ -395,6 +409,24 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_ > return XE_RAS_RECOVERY_ACTION_RECOVERED; > } > > +static u8 handle_fabric_errors(struct xe_device *xe, struct xe_ras_error_array *arr) > +{ > + struct xe_ras_error_product *product = &arr->counter.product; > + struct xe_ras_ieh_error *info = (void *)arr->details; > + u8 severity = arr->counter.common.severity; > + > + if ((info->global_error_status & XE_RAS_FAB_IEH_SAF_MHB) && > + product->cause.cause == XE_RAS_FAB_CAUSE_PAYLOAD) { > + xe_log_comp(xe, ras_sev_to_cper_sev(severity), FABRIC, &arr->counter, > + sizeof(arr->counter), "SAF MHB error detected\n"); > + return XE_RAS_RECOVERY_ACTION_RECOVERED; > + } > + > + xe_log_comp(xe, ras_sev_to_cper_sev(severity), FABRIC, &arr->counter, > + sizeof(arr->counter), "Errors detected\n"); > + return XE_RAS_RECOVERY_ACTION_RESET; > +} > + > void xe_ras_counter_threshold_crossed(struct xe_device *xe, > struct xe_sysctrl_event_response *response) > { > @@ -556,6 +588,9 @@ enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe) > case XE_RAS_COMP_DEVICE_MEMORY: > action = handle_device_memory_errors(xe, arr); > break; > + case XE_RAS_COMP_FABRIC: > + action = handle_fabric_errors(xe, arr); > + break; > default: > /* For any other component, reset */ > action = XE_RAS_RECOVERY_ACTION_RESET; > diff --git a/drivers/gpu/drm/xe/xe_ras_types.h b/drivers/gpu/drm/xe/xe_ras_types.h > index fe6f3658a2a4..4a8cd6fc5ba2 100644 > --- a/drivers/gpu/drm/xe/xe_ras_types.h > +++ b/drivers/gpu/drm/xe/xe_ras_types.h > @@ -12,6 +12,12 @@ > #define XE_RAS_NUM_ERROR_ARR 3 Nit: Worth a blank line here. Raag > /* Error bits in IEH global error status register */ > #define XE_RAS_SOC_IEH_PUNIT BIT(1) > +/* Bits 16-31 represent individual SAF MHB unit */ > +#define XE_RAS_FAB_IEH_SAF_MHB GENMASK(31, 16) > + > +/* Fabric Data payload parity errors */ > +#define XE_RAS_FAB_CAUSE_PAYLOAD BIT(2) > + > /* Device memory error categories */ > #define XE_RAS_MEMORY_DB_ECC BIT(1) > #define XE_RAS_MEMORY_POISON BIT(2) > -- > 2.47.1 >