From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from va-1-113.ptr.blmpb.com (va-1-113.ptr.blmpb.com [209.127.230.113]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2E9702E7372 for ; Fri, 26 Jun 2026 15:47:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.230.113 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782488876; cv=none; b=boWNMMdV6qlG3Xt7N39CG3aE7RdztikT2GXMN6B+eWReQBgfdcWhh0BQ+fdd25GZ1AoXB8jvTgyM588cpVKjUbgD0N3Kz8cdxSBBrLCH0zNv/NJnr2Z3fY7OgJxiS4S4Q2PMxT+TG02MJjoF6ztFBVzN1TC/aZJuSOBdddIsm0E= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782488876; c=relaxed/simple; bh=/tITg8s8T0btWnhBxfBqFrFoB7H9ez338jyx26AGnDg=; h=Mime-Version:References:Cc:Message-Id:To:Date:Content-Type: In-Reply-To:From:Subject; b=smhNqg5VqCYand29MpWyqvhISo82lLD2aN9jn60jQeGHUyHaJTBETbfeuJlj2wMzjXTZhOLngXYIs+pfjvyeysVDg65Ex7MFoIn/N5uXjB2mPffPFy+eHSPu0A1CICMBHZmhSp8JGdsZCUiO8VV1lv9/B5Z5Q0ZjCc+n3GzbY6A= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=e8OSi7fB; arc=none smtp.client-ip=209.127.230.113 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="e8OSi7fB" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=2212171451; d=bytedance.com; t=1782488864; h=from:subject: mime-version:from:date:message-id:subject:to:cc:reply-to:content-type: mime-version:in-reply-to:message-id; bh=MmBm10iT129m3PZ/tatNSgger/doUXn6PYN6R4ihZxI=; b=e8OSi7fBhpZfM0ztk7IPL2KiTynULpX0+UewtfcEIL6ZpVdXJRljykUNDVGRJ/2Sm6MBEY /t1QCGPuSSOYct5Yw8mjvFRpwPB8zM+Ir0aSTqlY60ppIIeGj/STErCvzn/kFur8Ns+A51 g2W9gjZ/mxnbTC4OZzxn5MQq2YbYmoIoveJXnkKUcNXdBdj4vUdexfflxEDbgIvcoYnkEY AEEtCQuyv3Ddzq1vmA/CFbmRPgh+wlRLSXpTY7+TTB9RHI8Xp4r7Bi5/nArPnGrkp1t62B +Bycx1Ulm5E0EB8CKGD+0R61rZN5fvg2azaerDmk38lS/JwyY17hGxB+3dmLNQ== Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Transfer-Encoding: 7bit References: <20260616111127.966468-1-zhouchuyi@bytedance.com> <20260616111127.966468-5-zhouchuyi@bytedance.com> <871pdtjryo.ffs@fw13> User-Agent: Mozilla Thunderbird Cc: Message-Id: <8d3587e6-e3a1-40f5-ba0d-65583a2f1ecb@bytedance.com> X-Lms-Return-Path: To: "Thomas Gleixner" , , , , , , , , , , , , , Date: Fri, 26 Jun 2026 23:47:21 +0800 Content-Type: text/plain; charset=UTF-8 X-Original-From: Chuyi Zhou In-Reply-To: <871pdtjryo.ffs@fw13> From: "Chuyi Zhou" Subject: Re: [PATCH v8 04/14] smp: Use task-local IPI cpumask in smp_call_function_many_cond() On 2026-06-26 10:29 p.m., Thomas Gleixner wrote: > On Tue, Jun 16 2026 at 19:11, Chuyi Zhou wrote: >> This patch prepares the task-local IPI cpumask during thread creation, and >> uses the local cpumask to replace the percpu cfd cpumask in >> smp_call_function_many_cond(). We will enable preemption during >> csd_lock_wait() later, and this can prevent concurrent access to the >> cfd->cpumask from other tasks on the current CPU. For cases where >> cpumask_size() is smaller than or equal to the pointer size, it tries to >> stash the cpumask in the pointer itself to avoid extra memory allocations. > > This one fails the comprehensible test and also does not match the rules of > how change logs should be written. > >> +#if defined(CONFIG_SMP) && defined(CONFIG_PREEMPTION) >> + union { >> + cpumask_t *ipi_mask_ptr; >> + unsigned long ipi_mask_val; > > Indentation of the variable name wants TABs not spaces > >> @@ -933,10 +934,14 @@ static struct task_struct *dup_task_struct(struct task_struct *orig, int node) >> #endif >> account_kernel_stack(tsk, 1); >> >> - err = scs_prepare(tsk, node); >> + err = smp_task_ipi_mask_alloc(tsk); > > Hrm. So we unconditionally allocate another per task CPU mask. How many > task actually utilize it? > > We keep making task_struct and the related things larger every other > release without actually looking at the resulting overall memory > consumption. > Thanks, this is a fair concern. The task-local cpumask approach came from the earlier discussion with Sebastian and Nadav. The problem we tried to solve there was the lifetime of the wait mask once the later patch re-enables preemption before csd_lock_wait(). At that point the wait mask can no longer be the per-CPU cfd->cpumask: the task may be preempted or migrate while it is still iterating the mask, and another task running on the original CPU could enter smp_call_function_many_cond() and reuse that per-CPU mask. I agree that the memory cost needs to be called out explicitly. The current implementation trades one task-local cpumask for a stable mask lifetime and avoids adding allocation/failure handling to the generic IPI path. I considered avoiding the fork-time allocation, but the alternatives do not look straightforward: - stack storage is not suitable for large NR_CPUS/CPUMASK_OFFSTACK configurations; - per-CPU storage is exactly what becomes unsafe once the wait is made preemptible; - allocating the mask in smp_call_function_many_cond() would put an allocation in the generic IPI path. It also cannot rely on a sleeping allocation because this function is entered from contexts which have historically only required preemption to be disabled. Using GFP_ATOMIC would need a failure/fallback path, in which case the latency improvement becomes opportunistic rather than guaranteed. For the motivating x86 TLB flush paths, the users are also not a small static set of tasks. Ordinary tasks can hit this through exit, unmap, reclaim, etc., so I do not see a clean way to allocate this only for a pre-identifiable subset of tasks.