From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6BF6034750F for ; Tue, 25 Aug 2026 14:43:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787669029; cv=none; b=BHYgkKZPRMgpkH6oImrUQn07lAbVZabt0yHTFr74hBDtsSQAnVVEhysK1xgef8/aULkcTkb1n9AGcfAxoKFFUo61YqWhRJoxQWOFwlV4EYwitMaTZxCDvBChi+W02yC0kppI+zxdvs3xbROWjD+A6RlH8832HhgUZFoHincWibs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787669029; c=relaxed/simple; bh=+mzjBvtJiXWlRYg5vALkaQ/lMTmfc5pn1N/aoubfE2Y=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=QGTlqMOpPlZF9yAdQK5PqwpvENXoQ6A0o0B0eq5aiKykieYl3oiVliSvA63kZFf5l031Hj8lUe1DBKas7T0S3r1tfTmqDBzXNSdSktdDwlJUOBP7H5gVfgK9YUf/raVaDv0pwP77neIn7vjjPSmYLESjjtQ9JAlW3t0/iTvqKAY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=dnqLgrcg; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=EBC6rTnL; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="dnqLgrcg"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="EBC6rTnL" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1787669026; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:autocrypt:autocrypt; bh=+mzjBvtJiXWlRYg5vALkaQ/lMTmfc5pn1N/aoubfE2Y=; b=dnqLgrcgaDqckoQ1yckdu4wladHeLr5byqoSwaP6W9ogRV4+Wv+79Zs9ujuGN9E+PL7qxZ eZZwHcl0DilqxFa4l0gr3deDWBWhBRpa5ixkpQfWF89TrIWa0IieybqNcqqRBjaUKzu5vj yopEVmRkGebYrdaKpXW48ip/RrBCzWY= Received: from mail-ej1-f72.google.com (mail-ej1-f72.google.com [209.85.218.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-128-WpIl-ev6Miy-HgDK3yGeUg-1; Tue, 25 Aug 2026 10:43:40 -0400 X-MC-Unique: WpIl-ev6Miy-HgDK3yGeUg-1 X-Mimecast-MFC-AGG-ID: WpIl-ev6Miy-HgDK3yGeUg_1787669018 Received: by mail-ej1-f72.google.com with SMTP id a640c23a62f3a-c15d4224f9dso162074766b.3 for ; Tue, 25 Aug 2026 07:43:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1787669018; x=1788273818; darn=vger.kernel.org; h=mime-version:user-agent:content-transfer-encoding:content-type :autocrypt:references:in-reply-to:date:cc:to:from:subject:message-id :from:to:cc:subject:date:message-id:reply-to:content-type; bh=+mzjBvtJiXWlRYg5vALkaQ/lMTmfc5pn1N/aoubfE2Y=; b=EBC6rTnL4kZ4N/A+LgeVoLM/k3+SLpZoTg2TOhVImNjHtU1PWqTXV0jMq/75pjB314 FcaLzmiLnpzpxj2L7Bj/FHSYedIP4nVSSDfczCEH+gWRPkMjvQOeNK4NU3sYy1y8zGjG F9SBSkGXLqAzSw3S7LV5UT21dUtzORU6OFFeB65N0XKWnrujebWULOnDhSO/QfUuOCJv DljKrqprYhEB1PQ0WS/ObGj46XUynGj2Jo7uPfNtwk4XViGaOrnkpZTTPfkC0EVLTXnp kfHzRez06rLAhy1w0ysqAVwUQtWyUR8b5QHPHY30OkxpBD4cXvrSF96NNSNBTn6MXGMV SYaQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787669018; x=1788273818; h=mime-version:user-agent:content-transfer-encoding:content-type :autocrypt:references:in-reply-to:date:cc:to:from:subject:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=+mzjBvtJiXWlRYg5vALkaQ/lMTmfc5pn1N/aoubfE2Y=; b=F3mVkYe8sOjfZ1Pgqu03LvupuTG6WHTxacsrTjScmYkMfxBW5CzuAcuw3+y1iQfyw5 yuRJt7cAauOxcJu6DeBYGxZvXH5xCGnsTO7w+VlroTbT7Yk4cF1jnV7/HJJ1qCBD09Ho 779zpfesKuRiKexcwQoROEE8MTyJLt6T13fti1YWGk/4+AVLovJJItfCgpj5OhM5xMtn srZwEMPrZrV1Gk+JmvBpOw5uKYFr16uNKx7tVSh8TJ70flZPOZSlMAutAFDpaIYJrxLp LREio/vDMNcpvdAIANVtX6Wu9RhzEeoz2ASxVlJdwCxc2cIaC+M/yQ0ZqvQaQP4AX1UG PoGg== X-Forwarded-Encrypted: i=1; AHgh+RrxONWk0qSfUpBBNfc3D/Azi3/87lfCnL3tPdwMeruIHKmVXAZN0aMtxmMnKmwH5FGsOdrnpe1rVzgSnn4=@vger.kernel.org X-Gm-Message-State: AFuF++lfLKXRt9J6IV3+rvUw6j91XKNhTzK47/E0E81Uct024GcDF3RJ jiTq3Zpe3JZloEfHmtEqJViYTT4ypjb7jTMb2swG4Xi2RdXwW0NYySirtRZh+OjEjtlifBj9Z83 Kuukkmxo8wAA4VPjYMgA8NVc1kZfh+ht/UCwqzpwatvERiE5sZ/oERu/6BxenkOVR8cEuNkwXaD xr X-Gm-Gg: AR+sD133fTu3oYiwAJW7DPFP4HD9acoSdPl7BZU5Al7430vGWhMj8LGwm0mntxKdShG SqJvOlXnSuK3ZW8s73TCoxsjcHAkR01a0aimgxSSeXsyC0n7V2SkuJ+/BVzheiXDVFD//uL9iqz b28reg84X+21rSsVKB+qyOFwwYl6aZm7aWeM3AM4hiQO5fXvj0qXD3AO9D1dOH4kdWWUiq5JcAd ZHLs+PxLKoBaseZp5mYFyVIBc1lMoH/fS5Sja8ZPHKcJEnRm7QnWFCMQ/iu7vQsAo4RvuIwVqDr /ZJkud5OS7o61QsGkgWMvdi1z4TeKdcA/e8kfJ5LwTejXs29/MKVuq4Wf9BfCRhKiWezzFsX9+c JzGXxRqsoTljOKtbsHHEI+/+SCEYDW8J0qsD40sBBi568KVq9qi4cgO7r9lWm+vGWml0qBw== X-Received: by 2002:a17:907:72cc:b0:c1f:8145:820c with SMTP id a640c23a62f3a-c2492579d7cmr2900603066b.2.1787669017991; Tue, 25 Aug 2026 07:43:37 -0700 (PDT) X-Received: by 2002:a17:907:72cc:b0:c1f:8145:820c with SMTP id a640c23a62f3a-c2492579d7cmr2900596466b.2.1787669017406; Tue, 25 Aug 2026 07:43:37 -0700 (PDT) Received: from gmonaco-thinkpadt14gen3.rmtit.csb (212-8-243-115.hosted-by-worldstream.net. [212.8.243.115]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24966f8f08sm1819198666b.35.2026.08.25.07.43.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 25 Aug 2026 07:43:36 -0700 (PDT) Message-ID: <1b3bb36a8acfa64dd250482e92a23dadcc8a6329.camel@redhat.com> Subject: Re: [PATCH v3 1/4] rv/reactors: use context-sensitive lockdep wait type in rv_react() From: Gabriele Monaco To: Wen Yang , Thomas =?ISO-8859-1?Q?Wei=DFschuh?= Cc: Nam Cao , linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org Date: Tue, 25 Aug 2026 16:43:34 +0200 In-Reply-To: References: <26526e555baa5118325b2383e9f7f0f8f9b6a199.1786294920.git.wen.yang@linux.dev> <38decc4f7ef5b6f03b37705c7e97f22a7e40d7e2.camel@redhat.com> <87zeylnos3.fsf@yellow.woof> <20260819112038-e7b033f8-3711-4acd-ba05-dbef015c6fc8@linutronix.de> Autocrypt: addr=gmonaco@redhat.com; prefer-encrypt=mutual; keydata=mDMEZuK5YxYJKwYBBAHaRw8BAQdAmJ3dM9Sz6/Hodu33Qrf8QH2bNeNbOikqYtxWFLVm0 1a0JEdhYnJpZWxlIE1vbmFjbyA8Z21vbmFjb0BrZXJuZWwub3JnPoiZBBMWCgBBFiEEysoR+AuB3R Zwp6j270psSVh4TfIFAmjKX2MCGwMFCQWjmoAFCwkIBwICIgIGFQoJCAsCBBYCAwECHgcCF4AACgk Q70psSVh4TfIQuAD+JulczTN6l7oJjyroySU55Fbjdvo52xiYYlMjPG7dCTsBAMFI7dSL5zg98I+8 cXY1J7kyNsY6/dcipqBM4RMaxXsOtCRHYWJyaWVsZSBNb25hY28gPGdtb25hY29AcmVkaGF0LmNvb T6InAQTFgoARAIbAwUJBaOagAULCQgHAgIiAgYVCgkICwIEFgIDAQIeBwIXgBYhBMrKEfgLgd0WcK eo9u9KbElYeE3yBQJoymCyAhkBAAoJEO9KbElYeE3yjX4BAJ/ETNnlHn8OjZPT77xGmal9kbT1bC1 7DfrYVISWV2Y1AP9HdAMhWNAvtCtN2S1beYjNybuK6IzWYcFfeOV+OBWRDQ== Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.60.2 (3.60.2-1.fc44) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Thu, 2026-08-20 at 01:28 +0800, Wen Yang wrote: > So I don't think LD_WAIT_SPIN for everything is safe, it's not just=20 > "less strict on paper". Yes it won't be safe, we could have wrong reactors not reported by lockdep,= but so would your conditional check. We'd be catching everything in interrupts,= but if a model never runs in interrupt context, we'd be demoted to LD_WAIT_SPIN nonetheless. Your condition is kind of a best-effort to do a bit better than just LD_WAIT_SPIN. Again, I'm not completely against it, but it isn't a final solution (not that I can think of a better one) and if most people find it confusing, we're probably better off without it. Or am I missing something? > If a reactor ever does raw_spin_lock() by mistake while running from=20 > nrp's path, check_wait_context() takes the wait type straight from the= =20 > override map: >=20 > =C2=A0=C2=A0=C2=A0=C2=A0 if (unlikely(class->lock_type =3D=3D LD_LOCK_WAI= T_OVERRIDE)) > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = curr_inner =3D prev_inner;=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 /* SPI= N */ > =C2=A0=C2=A0=C2=A0=C2=A0 if (next_outer > curr_inner) > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = return print_lock_invalid_wait_context(...); >=20 > SPIN nested in SPIN is 2 > 2, which is false, so it just passes. > We'd be silently giving up the one case (hardirq/NMI) where the "no=20 > locks" rule for reactors actually matters, which is the same thing that= =20 > bit you with the signal reactor. True, that is one, but it isn't the only case where using locks can be an i= ssue, namely sched_switch is particularly hard to please (interrupts are disabled= for a reason). > And for what it's worth, checking context to decide the wait type isn't= =20 > something we'd be introducing -- lockdep does the same thing to get the= =20 > baseline before any override is applied, in the exact function that=20 > produced this splat: >=20 > =C2=A0=C2=A0=C2=A0=C2=A0 static inline short task_wait_context(struct tas= k_struct *curr) > =C2=A0=C2=A0=C2=A0=C2=A0 { > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = if (lockdep_hardirq_context()) { > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 if (curr->hardirq_threaded= || curr->irq_config) > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2= =A0=C2=A0=C2=A0=C2=A0 return LD_WAIT_CONFIG; > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 return LD_WAIT_SPIN; > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = } else if (curr->softirq_context) { > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 return LD_WAIT_CONFIG; > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = } > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = return LD_WAIT_MAX; > =C2=A0=C2=A0=C2=A0=C2=A0 } >=20 > That's four cases. Our in_nmi() || in_hardirq() is a simplification of > what's already there, not a new habit. I'm not familiar with the lockdep code, so I may be missing something here,= but isn't this function getting the current context? It isn't used to tune the lockdep logic for some specific function, it /is/ the lockdep logic itself. Is this really comparable? > Thomas, on disabling preemption instead: it does fix the pagefault > case, and it's a no-op for nrp since preempt_count is already elevated= =20 > there. But it changes what a reactor is allowed to do, for every=20 > reactor, not just the two above.cspin_lock() on RT checks=20 > might_resched() before it even looks at thevlock: >=20 > =C2=A0=C2=A0=C2=A0=C2=A0 static __always_inline void __rt_spin_lock(spinl= ock_t *lock) > =C2=A0=C2=A0=C2=A0=C2=A0 { > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = rtlock_might_resched();=C2=A0=C2=A0 /* unconditional */ > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = rtlock_lock(&lock->lock); > =C2=A0=C2=A0=C2=A0=C2=A0 } >=20 > so any future reactor using a plain spinlock would hit that every single= =20 > time, lock contended or not, on top of whatever it's already called for.= =20 > That seems worse than the thing we're trying to fix. I'm also against disabling preemption, we may even be able to avoid schedul= ing by also disabling interrupts in a convenient order, but in principle I'd av= oid that like plague. RV was born as a tool for real-time and disabling preempt= ion only to get better lockdep reporting doesn't look right to me. Sure we could do it only with lockdep and assume lockdep's overhead is alre= ady enough, but we're again complicating things. > Since neither struct rv_reactor nor rv_react() actually documents what > context a callback may run in, maybe that's worth spelling out > separately regardless of what we do here -- something close to what > printk already does for the same reason (reactor_printk's > vprintk_deferred() leans on this internally: is_printk_legacy_deferred() > checks in_nmi() to decide whether it's safe to take console_lock or > whether it has to go through the lock-free irq_work path instead). You raise a fair point though, we should document all this somewhere. I bel= ieve this is the best thing we can do at this point. The peculiarity of reactors and RV in general is that we run from tracepoin= ts, so we could run virtually from anywhere, that's why there's no one size fit= s all.. Overall I'm now leaning towards using only LD_WAIT_SPIN and carefully docum= ent why we had to do this. What do you all think? Thanks, Gabriele