From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from gabe.freedesktop.org (gabe.freedesktop.org [131.252.210.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 1F7FDCA5FFE for ; Mon, 5 Oct 2026 19:06:18 +0000 (UTC) Received: from gabe.freedesktop.org (localhost [127.0.0.1]) by gabe.freedesktop.org (Postfix) with ESMTP id A01F110EE1A; Mon, 5 Oct 2026 19:06:17 +0000 (UTC) Authentication-Results: gabe.freedesktop.org; dkim=pass (2048-bit key; unprotected) header.d=intel.com header.i=@intel.com header.b="WeqqavLP"; dkim-atps=neutral Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) by gabe.freedesktop.org (Postfix) with ESMTPS id 732AD10E5F5 for ; Mon, 5 Oct 2026 19:06:15 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1791227175; x=1822763175; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=r0Eeg07DIFE2WcudOKrlD7eUOk8yhHlZu4uEAGqR0c8=; b=WeqqavLPvSpmu7sZAn+5f10bkE6Eyma3mBsAR0DA0MRS/jNdPrRmxxbP 4nUdJ9y7zfNB7WHlulrcustWg2Bj+ikZmDw/gTZSXEhKt006HA2o+X54j ZleYZNXPmfbK0aa+Sil9ld2ARsUynv3Yt61r21pG3bEF7ikj7gtZ1+nul 3HVbIMKmk4IewkmX6pRq1jrGrBMqMjAouT9afMif6tGIZqyXPJ/929PL5 2lXrxNgcffsHDpHwpGgCdG5f9sI4ckX5RRpRM6SQ5u7KAL1h69B6we8TZ RUh61I+hSxbIgMGaZc5PnVZdlqE1+uAI1MPCAyVfiZWNRm/Lk7U+SV/df Q==; X-CSE-ConnectionGUID: RqK1rLWETn+yFdudjvNxhw== X-CSE-MsgGUID: aw+1wZe3Q6uacfAiyN0v5Q== X-IronPort-AV: E=McAfee;i="6800,10657,11926"; a="102578418" X-IronPort-AV: E=Sophos;i="6.27,142,1787036400"; d="scan'208";a="102578418" Received: from orviesa006.jf.intel.com ([10.64.159.146]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Oct 2026 12:06:15 -0700 X-CSE-ConnectionGUID: IP49fPsRQi2zyTSY5G0QWA== X-CSE-MsgGUID: DbizEg7rTTuy+WLsWZIdZw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,142,1787036400"; d="scan'208";a="274690861" Received: from dut4435arlh.fm.intel.com ([10.105.8.126]) by orviesa006-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Oct 2026 12:06:15 -0700 From: Stuart Summers To: Cc: intel-xe@lists.freedesktop.org, rodrigo.vivi@intel.com, matthew.brost@intel.com, umesh.nerlige.ramappa@intel.com, gustavo.sousa@intel.com, matthew.d.roper@intel.com, daniele.ceraolospurio@intel.com, shuicheng.lin@intel.com, Stuart Summers Subject: [PATCH 03/15] drm/xe/configfs: Copy wa_bb out under the configfs lock Date: Mon, 5 Oct 2026 19:06:13 +0000 Message-ID: <20261005190611.332940-20-stuart.summers@intel.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20261005190611.332940-17-stuart.summers@intel.com> References: <20261005190611.332940-17-stuart.summers@intel.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-BeenThere: intel-xe@lists.freedesktop.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Intel Xe graphics driver List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: intel-xe-bounces@lists.freedesktop.org Sender: "Intel-xe" The ctx_restore_{mid,post}_bb getters handed the caller a pointer into storage owned by the configfs group and then dropped both the lock and the group reference. A concurrent store can krealloc() that buffer and an rmdir can free it outright, leaving the caller in xe_lrc.c to memcpy() from freed memory. Copy the batch into a caller supplied buffer while the lock is still held instead, which also folds the length check the callers were doing into the getters. Fixes: 6c6988c5e03d ("drm/xe/lrc: Allow to add user commands on context switch") Fixes: 39ac06f70062 ("drm/xe/configfs: Add post context restore bb") Signed-off-by: Stuart Summers Assisted-by: LLM --- drivers/gpu/drm/xe/xe_configfs.c | 52 ++++++++++++++++++++++---------- drivers/gpu/drm/xe/xe_configfs.h | 24 +++++++-------- drivers/gpu/drm/xe/xe_lrc.c | 38 +++++++---------------- 3 files changed, 59 insertions(+), 55 deletions(-) diff --git a/drivers/gpu/drm/xe/xe_configfs.c b/drivers/gpu/drm/xe/xe_configfs.c index 1c7a096aa046..6e62fdccb4c4 100644 --- a/drivers/gpu/drm/xe/xe_configfs.c +++ b/drivers/gpu/drm/xe/xe_configfs.c @@ -1441,25 +1441,34 @@ bool xe_configfs_get_disable_vram_page_offline(struct pci_dev *pdev) * xe_configfs_get_ctx_restore_mid_bb - get configfs ctx_restore_mid_bb setting * @pdev: pci device * @class: hw engine class - * @cs: pointer to the bb to use - only valid during probe + * @cs: destination buffer for the batch, or NULL to only query the length + * @max_len: capacity of @cs, in dwords * - * Return: Number of dwords used in the mid_ctx_restore setting in configfs + * Copy the configured batch into @cs while holding the configfs lock, so the + * caller never gets a pointer to storage owned by the configfs group. + * + * Return: Number of dwords used in the mid_ctx_restore setting in configfs, or + * -ENOSPC if it does not fit in @max_len */ -u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs) +ssize_t xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len) { struct xe_config_group_device *dev = find_xe_config_group_device(pdev); - u32 len; + ssize_t len; if (!dev) return 0; scoped_guard(mutex, &dev->lock) { - if (cs) - *cs = dev->config.ctx_restore_mid_bb[class].cs; - len = dev->config.ctx_restore_mid_bb[class].len; + if (cs && len) { + if (len > max_len) + len = -ENOSPC; + else + memcpy(cs, dev->config.ctx_restore_mid_bb[class].cs, + len * sizeof(u32)); + } } config_group_put(&dev->group); @@ -1470,23 +1479,34 @@ u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, * xe_configfs_get_ctx_restore_post_bb - get configfs ctx_restore_post_bb setting * @pdev: pci device * @class: hw engine class - * @cs: pointer to the bb to use - only valid during probe + * @cs: destination buffer for the batch, or NULL to only query the length + * @max_len: capacity of @cs, in dwords * - * Return: Number of dwords used in the post_ctx_restore setting in configfs + * Copy the configured batch into @cs while holding the configfs lock, so the + * caller never gets a pointer to storage owned by the configfs group. + * + * Return: Number of dwords used in the post_ctx_restore setting in configfs, or + * -ENOSPC if it does not fit in @max_len */ -u32 xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs) +ssize_t xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len) { struct xe_config_group_device *dev = find_xe_config_group_device(pdev); - u32 len; + ssize_t len; if (!dev) return 0; scoped_guard(mutex, &dev->lock) { - *cs = dev->config.ctx_restore_post_bb[class].cs; len = dev->config.ctx_restore_post_bb[class].len; + if (cs && len) { + if (len > max_len) + len = -ENOSPC; + else + memcpy(cs, dev->config.ctx_restore_post_bb[class].cs, + len * sizeof(u32)); + } } config_group_put(&dev->group); diff --git a/drivers/gpu/drm/xe/xe_configfs.h b/drivers/gpu/drm/xe/xe_configfs.h index 10a00d6bbf8f..14e49f23306f 100644 --- a/drivers/gpu/drm/xe/xe_configfs.h +++ b/drivers/gpu/drm/xe/xe_configfs.h @@ -26,12 +26,12 @@ bool xe_configfs_get_psmi_enabled(struct pci_dev *pdev); bool xe_configfs_get_enable_multi_queue(struct pci_dev *pdev); u32 xe_configfs_get_migrate_ulls_period_ms(struct pci_dev *pdev); bool xe_configfs_get_disable_vram_page_offline(struct pci_dev *pdev); -u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs); -u32 xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs); +ssize_t xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len); +ssize_t xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len); #ifdef CONFIG_PCI_IOV unsigned int xe_configfs_get_max_vfs(struct pci_dev *pdev); bool xe_configfs_admin_only_pf(struct pci_dev *pdev); @@ -48,12 +48,12 @@ static inline bool xe_configfs_get_psmi_enabled(struct pci_dev *pdev) { return f static inline bool xe_configfs_get_enable_multi_queue(struct pci_dev *pdev) { return true; } static inline u32 xe_configfs_get_migrate_ulls_period_ms(struct pci_dev *pdev) { return 5; } static inline bool xe_configfs_get_disable_vram_page_offline(struct pci_dev *pdev) { return false; } -static inline u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs) { return 0; } -static inline u32 xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, - enum xe_engine_class class, - const u32 **cs) { return 0; } +static inline ssize_t xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len) { return 0; } +static inline ssize_t xe_configfs_get_ctx_restore_post_bb(struct pci_dev *pdev, + enum xe_engine_class class, + u32 *cs, size_t max_len) { return 0; } #ifdef CONFIG_PCI_IOV static inline unsigned int xe_configfs_get_max_vfs(struct pci_dev *pdev) { diff --git a/drivers/gpu/drm/xe/xe_lrc.c b/drivers/gpu/drm/xe/xe_lrc.c index de9a9560ecf3..f751687dc568 100644 --- a/drivers/gpu/drm/xe/xe_lrc.c +++ b/drivers/gpu/drm/xe/xe_lrc.c @@ -94,7 +94,7 @@ gt_engine_needs_indirect_ctx(struct xe_gt *gt, enum xe_engine_class class) return true; if (xe_configfs_get_ctx_restore_mid_bb(to_pci_dev(xe->drm.dev), - class, NULL)) + class, NULL, 0) > 0) return true; if (gt->ring_ops[class]->emit_aux_table_inv) @@ -1186,17 +1186,12 @@ static ssize_t setup_configfs_post_ctx_restore_bb(struct xe_lrc *lrc, bool indirect) { struct xe_device *xe = gt_to_xe(lrc->gt); - const u32 *user_batch; - u32 *cmd = batch; - u32 count; + ssize_t count; count = xe_configfs_get_ctx_restore_post_bb(to_pci_dev(xe->drm.dev), - hwe->class, &user_batch); - if (!count) - return 0; - - if (count > max_len) - return -ENOSPC; + hwe->class, batch, max_len); + if (count <= 0) + return count; /* * This should be used only for tests and validation. Taint the kernel @@ -1204,10 +1199,7 @@ static ssize_t setup_configfs_post_ctx_restore_bb(struct xe_lrc *lrc, */ add_taint(TAINT_TEST, LOCKDEP_STILL_OK); - memcpy(cmd, user_batch, count * sizeof(u32)); - cmd += count; - - return cmd - batch; + return count; } static ssize_t setup_configfs_mid_ctx_restore_bb(struct xe_lrc *lrc, @@ -1216,17 +1208,12 @@ static ssize_t setup_configfs_mid_ctx_restore_bb(struct xe_lrc *lrc, bool indirect) { struct xe_device *xe = gt_to_xe(lrc->gt); - const u32 *user_batch; - u32 *cmd = batch; - u32 count; + ssize_t count; count = xe_configfs_get_ctx_restore_mid_bb(to_pci_dev(xe->drm.dev), - hwe->class, &user_batch); - if (!count) - return 0; - - if (count > max_len) - return -ENOSPC; + hwe->class, batch, max_len); + if (count <= 0) + return count; /* * This should be used only for tests and validation. Taint the kernel @@ -1234,10 +1221,7 @@ static ssize_t setup_configfs_mid_ctx_restore_bb(struct xe_lrc *lrc, */ add_taint(TAINT_TEST, LOCKDEP_STILL_OK); - memcpy(cmd, user_batch, count * sizeof(u32)); - cmd += count; - - return cmd - batch; + return count; } static ssize_t setup_invalidate_state_cache_wa(struct xe_lrc *lrc, -- 2.43.0