From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f49.google.com (mail-pj1-f49.google.com [209.85.216.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 30ABB64 for ; Sun, 26 Apr 2026 01:47:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777168079; cv=none; b=r+oFadg+iUyA4LK9Pc3sGV6q/tX+EnQZhYreOa8AxEgte/nyP1JXnkh/TQf2pyJaSVwSDhVKona7onv285ICEHWiWWBJ0ucbXtjJJ62Gmg1JiJ2mE2c7Waq622s8rLKbeWCOzhGlsfxi63ThsV7EJsTkPMZX6CuxGALjrCXm2Tw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777168079; c=relaxed/simple; bh=QrvqSDc041TNuXNt3BBlwMz8vZoL2QGWRmws13hXG+Q=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=RSNx++daRDgcvPADuUnN/hbQ8WtT/Io3Y+xsc9o8bmOpHBXhj7pPNOmZ63/ioGYsIFOxKGYfRmzcUY7CHdxyuIaVIi2OFaAuThDc3u+m/g0bkwzdPuGw1qY4toAfyCCUMZl2Ry0DKa69Hrvv/Wij8WUlwECKAtevULeGIvKDQPE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=MlCeL93i; arc=none smtp.client-ip=209.85.216.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="MlCeL93i" Received: by mail-pj1-f49.google.com with SMTP id 98e67ed59e1d1-35d94f4ee36so5542100a91.3 for ; Sat, 25 Apr 2026 18:47:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1777168077; x=1777772877; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=RL7xdU0D39uNUreJOg8ut3Rg0sXkiiyngpi4Y/rXepg=; b=MlCeL93i+k5BnXpoadVv3h20tmwbQfh1RoGBECjnNylbPg5k5eAWULQYDzmSVLY/t/ 0oNej1XGJH3ZGIqUxndUJ5uM/1Q+kmu2MRQ9zXpuwZYyg6HWzR8hWDmc7YEsL9OiFOKa XPkL5pxAlbadb7UdytZAZJaJKxiz9hUfMaAIzFAXgNlj8R87Jf4lfPIhNlvTpdO1SR7B 3FQgFed8mULWJ9wzOK/O026i6sN0ChHiukh6X4cA79I4xm480V03fjLBh18c1nOlCiXn LN4Q8SOlRnuqpFV9mJqUii494pIyl7Qo59rZC+8ZoLCkzzqpeP+3b3dPkb7OLC5lMmWe yMHw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1777168077; x=1777772877; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=RL7xdU0D39uNUreJOg8ut3Rg0sXkiiyngpi4Y/rXepg=; b=sz+fKkvexCyVY3EqvH2e+Rtk5Z0H8AnLz5zn5nSv1BSfSilOyvSz/FhdrQ3ux+dmgq rms9VIg0W4q9pn3pi70eJICnfHkkmf4YzLFWkZd2sTlqXRAzQWuWdOVkcXvEHZq/27tU gonuKAffiIjcvV8TXL3Jr10LzV6IZcWQX8g2pvtYDPG2/wgq3PMCGIPSabSjMYZ8ZL/1 3Oy0JCX+IlFAqLWzgE8GN8QYjlpodVQ2LkDOfRuDcJ0dLdRFzQvQN+OuwPK9n4+0MZMv OObUfqSTiXn6zOG0VRmMAsLeC8UnUnH1IUd0BYaqJo6nu5KsV0rTQOOKvVflzCJW0b2N ES9A== X-Forwarded-Encrypted: i=1; AFNElJ++GIPjYiiyOciYiYk3uLhij13Rp2tByC+5Ngb/I5Bimgp+v54xZd6QDsiw0xZh78VAnhoMaK8OBzCIRIA=@vger.kernel.org X-Gm-Message-State: AOJu0YwRQCIh7JSjYlxEcWgqws/qXfFFI5vJcLnJ8Zo55V75gXhC8Ohf XYaqmkpWHhKXbTZ5C40tEXjjqCPGfK5rs+e5mj+M/I7KkHHQVW8Kxfj6uKH8DPBD X-Gm-Gg: AeBDiev1Hco6VzaPXGdDZEcfTJtrXpapuk8OYN5MF7aRs+0DG7ZzPzKun3aVrkvAgZq 2+sKeT5fYxXzYdqVQmK1DW3PqT7neiGwPaKg07nsKG0aZrlaDi6Ure5Et0oAV8cAPC/psW1yyLy aer2wyz6I+iQClu5PmBJvuCpnbwwTzWDbnHflnVSt1BwiylM7oAaicdT70uXbxUC+Ay64NuCxP5 GxEN1bvz7uTn6FFkZMckwQcMf7chJkrR8jNKveA6TyPsPNlVFasVmu8MzkCb9ucN16qS5ZmlWW2 BSFQkUZBOa09bvnnW08irfvR4mo7tVAMBfISGRaxdmDgYNGAAfNDcTuDxXeTXB2xHxuqW8zbX/2 NH+offr1eTeHGQMIr2ETz+x4udiDEwDsWiGy19vTTOrqKNFgsUhT9IZPzG/Zvo14rWlM1t5wfoO X9sArVbSqEZZxmOFtBmaLSYZvTPFMQiaRnnqxRp12zFGJEULsWCi1Ywi4M8BP+3IkChNZ9FNzVD i5XsKtSn7srfECf X-Received: by 2002:a17:90b:3504:b0:35f:bb33:d727 with SMTP id 98e67ed59e1d1-36140478be0mr37352571a91.20.1777168077417; Sat, 25 Apr 2026 18:47:57 -0700 (PDT) Received: from cchengyang.duckdns.org (36-225-83-234.dynamic-ip.hinet.net. [36.225.83.234]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-36141973c57sm27058415a91.14.2026.04.25.18.47.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 25 Apr 2026 18:47:56 -0700 (PDT) Date: Sun, 26 Apr 2026 09:47:53 +0800 From: Cheng-Yang Chou To: Kuba Piecuch Cc: Tejun Heo , Andrea Righi , David Vernet , Changwoo Min , Emil Tsalapatis , Christian Loehle , Daniel Hodges , sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org, Ching-Chun Huang , Chia-Ping Tsai Subject: Re: [PATCH v2 sched_ext/for-7.1] sched_ext: Invalidate dispatch decisions on CPU affinity changes Message-ID: <20260426093756.Gd781@cchengyang.duckdns.org> References: <20260319083518.94673-1-arighi@nvidia.com> <20260422142633.G7180@cchengyang.duckdns.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: Hi Kuba, On Thu, Apr 23, 2026 at 01:32:20PM +0000, Kuba Piecuch wrote: > > On Mon, Mar 23, 2026 at 01:13:20PM -1000, Tejun Heo wrote: > >> > The simple way to do this is to do scx_bpf_dsq_insert() at the very beginning, > >> > once we know which task we would like to dispatch, and cancel the pending > >> > dispatch via scx_bpf_dispatch_cancel() if any of the pre-dispatch checks fail > >> > on the BPF side. This way, the "critical section" includes BPF-side checks, and > >> > SCX will ignore the dispatch if there was a dequeue/enqueue racing with the > >> > critical section. > >> > > >> > With this solution, we can throw an error if task_can_run_on_remote_rq() is > >> > false, because we know that there was no racing cpumask change (if there was, > >> > it would have been caught earlier, in finish_dispatch()). > >> > >> Yeah, I think this makes more sense. qseq is already there to provide > >> protection against these events. It's just that the capturing of qseq is too > >> late. If insert/cancel is too ugly, we can introduce another kfunc to > >> capture the qseq - scx_bpf_dsq_insert_begin() or something like that - and > >> stash it in a per-cpu variable. That way, qseq would be cover the "current" > >> queued instance and the existing qseq mechanism would be able to reliably > >> ignore the ones that lost race to dequeue. > > > > Since this has been stale for a while, I prepared a patch to implement > > scx_bpf_dsq_insert_begin() as suggested. > > Thanks for creating the patch. A couple of thoughts: > > 1. Do we have a use case that requires dsq_insert_begin() that isn't > satisfied using the "insert and then cancel if needed" approach? IIUC, yes. scx_bpf_dispatch_cancel() is only registered in scx_kfunc_ids_dispatch, so it is only callable from ops.dispatch(). dsq_insert_begin(), on the other hand, is available from both ops.enqueue() and ops.dispatch() (SCX_KF_ENQUEUE | SCX_KF_DISPATCH). Since there is nothing to cancel in ops.enqueue(), the insert-and-cancel approach simply doesn't work there. > > 2. Do we want to restrict ourselves through the one qseq slot provided by > dsq_insert_begin()? The most flexible approach IMO would be to simply > allow BPF to read the qseq directly via a kfunc and then supply it to > dsq_insert() later. With this, we can have multiple qseqs saved at the > same time, and we can even pass them between CPUs, e.g. if one CPU > dequeues a task for a sibling CPU, but we want the checks to be made inside > the sibling's ops.dispatch() (I just made this use case it up, it may not > be practical.) > That said, exposing an internal thing like qseq to BPF may be a step too far. In Tejun's reply back in [1], he suggested dsq_insert_begin() precisely to avoid promoting qseq into the BPF ABI — which matches your own concern. The single per-CPU slot is sufficient for the one-task-per-iteration dispatch loops used by existing schedulers (e.g., scx_central). If a concrete cross-CPU use case materializes later, we can always extend dsq_insert() to accept an explicit qseq without breaking the current, simpler path. [1]: https://lore.kernel.org/all/acHJED4iAeytdC2l@slm.duckdns.org/ > Let me know what you think. > Please correct me if I'm missing something, thanks! ^0^ -- Cheers, Cheng-Yang