From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ej1-f44.google.com (mail-ej1-f44.google.com [209.85.218.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAE644C9564 for ; Fri, 21 Aug 2026 15:23:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.44 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325794; cv=none; b=N1ZuIyFBSTF+/vw9Lyva4Grs8D9iezYFQmZvnP4YX4maLuecRoCuqA3DdjFBpb/X9ce2caaR6nsuH9vddbvbi5Uf7KVOo5B7zVG+nu38tNHlYatXEPXhLZhgh/kVrkAAVJRkoUU+CboboBgv6mQtISM65M9nbvwxhVudBzh1xcI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325794; c=relaxed/simple; bh=mDUYiCIFTwYuylrBDcfPAZgmy/Xj0GyFwnBLs10N0Wg=; h=From:Date:Subject:To:Cc:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=JV2mH4ObTmbO7CkKl8pxm/cIl2QZacvBluvwZgc5Ixd6+tgNRmIoHT5T3aYRYd0DegqptbhKe4+BjX5BcD7wYsFdz7M64dz/r5bzX75E+x28IyTJFem95txuvfigwkMCaR71CdtuWJhjLkFrN1bT4IkbdBgejlC3Cb2v2gbU22Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=a+R9ixyI; arc=none smtp.client-ip=209.85.218.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="a+R9ixyI" Received: by mail-ej1-f44.google.com with SMTP id a640c23a62f3a-c1676497000so12243366b.3 for ; Fri, 21 Aug 2026 08:23:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787325791; x=1787930591; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=vYeEFSz/we+8LtDsD3eJWBaG2LH+3C5IKZLmMvYVchk=; b=a+R9ixyIjLFQA7TMBYom+YBcNCouijIWyGVROeCBpBkJF59Qhz5A3b+LxV+G0+G0lK TMIUsF0dt4ui4NTvuScXanFqwMu625t6iR7I0gvTKwTWZ8NwdMhhV1w+QH2L67Af2W5h C02F8Zs6fDHemQb59nCbPewjNqjWP8Yu32TCTICTQ982RCaIAJeo7FjUp7YfmIVYXmN/ tz2Jg+8kh/XWg2kGoskhWEsNVRR8bwIjzNR8o+rAFlWJg4vAaX268Q+ZwjGJ7imfVZzC cYjYguwsH4w2kj32bcRDIBKwaB5ETN3gHkkvlwAxdOE+6LXFOEdZLyylfPE/PlgSl2nN D5xg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787325791; x=1787930591; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=vYeEFSz/we+8LtDsD3eJWBaG2LH+3C5IKZLmMvYVchk=; b=Ol1MaP7y6taefXXrM5ctBPuJYhdF7ILtwdv/qUnW3K+T4qfewN8EPIT8t1/9fxcxIf uSglLsnuDJaWNAE6roEGl5P07cL0otykENR95LA5ZMmhn49EjiGOeHHUIdQC6XOk/svR x6q35UiyHr1PQ5cukcex7bOoBe/mPtYKpcoFfidHebreP5uS9v2ICeieK1QHYx2/yHe5 qhrUW687f9/rxGi2SBY9pTZ03+OVFMYAZIpysTT9av6C41mc8dtfsqvpz+LzdFdDsjai 5CuyDTiONa6MZBk/e8xLSUHsvnjJTYxausSA4WPZ3EGrZnFP4HQ+pxF0AZZauBhTxgrR 6O4Q== X-Gm-Message-State: AFuF++mEOp9NxYeqWq+7wkIeaE5etRESCWZoaABk24y64lz4xc1RrxW0 +6OxHXs82GUii/UICb27FufhBr/CZ48MYWBtSTg24QDtB0gx3oX17riE X-Gm-Gg: AR+sD11xhkN076aK57R7bkU0JidWRmSr95jfFOwc4IIZdZUmSwVKF98JCVznPMaKg1I Sir3AQTxAXkOCreMj8bwj1hf5qKET6u0tgFDyU8vEIDWEnwdnQ79QywYT64yi5db3atbV0OP0JF fojKPCvgOgAhijGVX3oVSxAeAXo+8AlwxlHA7v2Qc+vFmw622QTOdYFGBkMa1N76HIf/tjc2Zs3 /dkUyIbTeE6e0pX8kPzfuULkYZD1E9V5sJcMp6+eg0SZIGrEXcEDFsq3pfFzNKaGUocPrRfmj/2 Yl/y0IoTm+po1CYNK0fGvO6UO3biqrMAMZX3/JDXgVZc4DmXn6UeHjmnk/FO1q6x4WMEsNqkJ0e 8zBIasMZUdfyOb51hVw05saCS7b1G2l3Qe94mLcwtcqmu8xm7Qkxt7cIysVAcJGWAgPC9/MOM5H TTX5LcKFx8n2Q/VN5kQ3/8acplsqWgtY8d4j7CLue0P8P9KJUSPK4qoQs0wpmqS6fPlbXM5ZndD FVVniHkLMiaDDt25DXXzUX9Mgr//zUynczg0ve8fw== X-Received: by 2002:a17:907:84e:b0:c20:7897:bc67 with SMTP id a640c23a62f3a-c246a6a8e0bmr384314166b.3.1787325791134; Fri, 21 Aug 2026 08:23:11 -0700 (PDT) Received: from [127.0.0.1] (ip-109-193-028-127.um39.pools.vodafone-ip.de. [109.193.28.127]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24591dc662sm506198566b.42.2026.08.21.08.23.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 21 Aug 2026 08:23:10 -0700 (PDT) From: Marek Czernohous X-Google-Original-From: Marek Czernohous Date: Fri, 21 Aug 2026 17:23:01 +0200 Subject: [PATCH v4 2/3] drm/nouveau: don't kill a fence context that is not ready yet To: nouveau@lists.freedesktop.org, dri-devel@lists.freedesktop.org Cc: linux-kernel@vger.kernel.org, Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Ben Skeggs Message-ID: <178732578167.167481.4590474568281802870@gmail.com> In-Reply-To: <178732578167.167481.5619512544301226563@gmail.com> References: <178732578167.167481.5619512544301226563@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit nouveau_channel_init() arms the channel-kill subscription early, right after mapping userd, and only creates the fence context at the very end of the same function. The handler it installs, nouveau_channel_killed(), reaches nouveau_fence_context_kill(chan->fence). The NULL check in nouveau_channel_kill() does not cover the window in between. Every backend that can reach it publishes the pointer before the context is usable: fctx = chan->fence = kzalloc_obj(*fctx); if (!fctx) return -ENOMEM; nouveau_fence_context_new(chan, &fctx->base); and nouveau_fence_context_new() is what runs spin_lock_init(&fctx->lock) and INIT_LIST_HEAD(&fctx->pending). An event arriving after the assignment but before that call finds chan->fence non-NULL and unusable: nouveau_fence_context_kill() takes a lock that was never initialised and walks a list head whose next pointer is still the NULL left by kzalloc(). Give the fence context a ->ready flag and hand the kill over through it. nouveau_fence_context_arm() sets the flag once nouveau_channel_init() has finished building the context, and nouveau_channel_kill() leaves the context alone until it is set. A kill arriving while the context is still being built is no longer lost either: it is recorded in chan->killed, and nouveau_fence_context_arm() acts on it as soon as there is a context to kill. The two sides hand over rather than exclude each other, because the kill side must not touch fctx->lock at all before the context is built, which is the very bug being fixed. Each stores its own flag before it loads the other's, so at least one of them observes the other. Both observing it is harmless: nouveau_fence_context_kill() then walks a list the first caller has already emptied. This does not close the other window. A kill delivered before nouveau_channel_init() subscribes is still not observed at all, and nvkm_uchan_init() makes the channel schedulable before that point. Closing that one means subscribing before the channel becomes schedulable, which is a larger change than this fix. The approach is Lyude Paul's suggestion. It is implemented with two differences from the sketch, both following from the same detail. The sketch checks chan->killed before setting ->ready. Both sides have to store their own flag before loading the other's, or the interleaving loses the kill: arm() reads killed == 0, kill() sets killed and reads ready == false, arm() then sets ready, and neither calls nouveau_fence_context_kill(). That outcome is reachable under sequential consistency, so no barrier can forbid it and the two accesses have to be the other way round in program order. Swapped, and with the smp_mb() on each side, this is the store-buffering pattern of tools/memory-model/litmus-tests/SB+fencembonceonces.litmus. The sketch also holds fctx->lock across the handover. The kill side cannot join it, because reaching fctx->lock is exactly what has to be avoided until the context is built: on those backends chan->fence is published by the allocation, before nouveau_fence_context_new() calls spin_lock_init(). So ->ready is read outside the lock. That answers the open question in the sketch as well: it does not have to be atomic_t, but it does have to be published with release and read with acquire, so that a caller that sees it set also sees the initialised lock and list. Fixes: ea13e5abf807 ("drm/nouveau: signal pending fences when channel has been killed") Cc: stable@vger.kernel.org Suggested-by: Lyude Paul Assisted-by: Claude:claude-opus-5 Signed-off-by: Marek Czernohous --- drivers/gpu/drm/nouveau/nouveau_chan.c | 19 ++++++++++++++++--- drivers/gpu/drm/nouveau/nouveau_fence.c | 19 +++++++++++++++++++ drivers/gpu/drm/nouveau/nouveau_fence.h | 8 ++++++++ 3 files changed, 43 insertions(+), 3 deletions(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_chan.c b/drivers/gpu/drm/nouveau/nouveau_chan.c index f142f6310596..605ce74c0d15 100644 --- a/drivers/gpu/drm/nouveau/nouveau_chan.c +++ b/drivers/gpu/drm/nouveau/nouveau_chan.c @@ -43,9 +43,17 @@ module_param_named(vram_pushbuf, nouveau_vram_pushbuf, int, 0400); void nouveau_channel_kill(struct nouveau_channel *chan) { + struct nouveau_fence_chan *fctx; + atomic_set(&chan->killed, 1); - if (chan->fence) - nouveau_fence_context_kill(chan->fence, -ENODEV); + + /* Pairs with the smp_mb() in nouveau_fence_context_arm(). */ + smp_mb(); + + fctx = READ_ONCE(chan->fence); + /* Pairs with the smp_store_release() there. */ + if (fctx && smp_load_acquire(&fctx->ready)) + nouveau_fence_context_kill(fctx, -ENODEV); } static int @@ -494,7 +502,12 @@ nouveau_channel_init(struct nouveau_channel *chan, u32 vram, u32 gart) } /* initialise synchronisation */ - return nouveau_fence(drm)->context_new(chan); + ret = nouveau_fence(drm)->context_new(chan); + if (ret) + return ret; + + nouveau_fence_context_arm(chan); + return 0; } int diff --git a/drivers/gpu/drm/nouveau/nouveau_fence.c b/drivers/gpu/drm/nouveau/nouveau_fence.c index edbe9e08ba0f..2fed631d44ba 100644 --- a/drivers/gpu/drm/nouveau/nouveau_fence.c +++ b/drivers/gpu/drm/nouveau/nouveau_fence.c @@ -93,6 +93,25 @@ nouveau_fence_context_kill(struct nouveau_fence_chan *fctx, int error) spin_unlock_irqrestore(&fctx->lock, flags); } +/* + * Declare a finished fence context killable. A kill can arrive while the + * caller is still building the context, so this and nouveau_channel_kill() + * hand over through fctx->ready and chan->killed. + */ +void +nouveau_fence_context_arm(struct nouveau_channel *chan) +{ + struct nouveau_fence_chan *fctx = chan->fence; + + /* Pairs with the smp_load_acquire() in nouveau_channel_kill(). */ + smp_store_release(&fctx->ready, true); + /* Pairs with the smp_mb() there: store-buffering, one side always sees the other. */ + smp_mb(); + + if (atomic_read(&chan->killed)) + nouveau_fence_context_kill(fctx, -ENODEV); +} + void nouveau_fence_context_del(struct nouveau_fence_chan *fctx) { diff --git a/drivers/gpu/drm/nouveau/nouveau_fence.h b/drivers/gpu/drm/nouveau/nouveau_fence.h index 183dd43ecfff..d9fede5dcba6 100644 --- a/drivers/gpu/drm/nouveau/nouveau_fence.h +++ b/drivers/gpu/drm/nouveau/nouveau_fence.h @@ -53,6 +53,13 @@ struct nouveau_fence_chan { struct work_struct uevent_work; struct nvif_event event; int notify_ref, dead, killed; + + /* + * Set by nouveau_fence_context_arm() once the context is complete. + * Read without fctx->lock, which nouveau_channel_kill() may not + * touch until it is set. + */ + bool ready; }; struct nouveau_fence_priv { @@ -71,6 +78,7 @@ void nouveau_fence_context_new(struct nouveau_channel *, struct nouveau_fence_ch void nouveau_fence_context_del(struct nouveau_fence_chan *); void nouveau_fence_context_free(struct nouveau_fence_chan *); void nouveau_fence_context_kill(struct nouveau_fence_chan *, int error); +void nouveau_fence_context_arm(struct nouveau_channel *chan); int nv04_fence_create(struct nouveau_drm *); int nv04_fence_mthd(struct nouveau_channel *, u32, u32, u32); -- 2.54.0