From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A8E8619E83C for ; Fri, 17 Jan 2025 16:53:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737132824; cv=none; b=pUdY+xUYx5lbvaElrUN5CBqOQJrHpp+XP2U+CM6iHzXaIeNyQQOgO5fUdl3a36cD15D6gb8tWUUSJp+PRkq7j3S2bGm2G1j8fxFWjTiYxilps3H0bFQ8S6IH+uxoeaIo3jtyKxQIMXOszGdOKu5p2zOJLbWovo7r7HsgBfZltjA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737132824; c=relaxed/simple; bh=N/IFo0/H4hlssGmiFiddM6Zrvy7F7bLkhW3DK1UInsM=; h=From:To:Cc:Subject:In-Reply-To:References:Date:Message-ID: MIME-Version:Content-Type; b=aefjG2D5D6dohcm2nIsXVuVW1MmI/fBx7CtsZ5ukc7d32O04BF8pRVwjuS/j+Ltm2u88ZBqbbT9M63JmQyhTo8ZNe1uFqKN6ULMcQ4GkTQVDExmAkC18AzwoCoUh+bKwrpx8qLrwrHqh7/Q1PkZhYXXCH/m+Bkle9qrvYV9/6JY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=QEeLCMX/; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="QEeLCMX/" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1737132820; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=N/IFo0/H4hlssGmiFiddM6Zrvy7F7bLkhW3DK1UInsM=; b=QEeLCMX/oNZYqE/f8zniWr+O3tTGE7QzyOo76ZLV1xaKZ7Sswa1q64nQ7NEl2L4882SJnI Y5gAa+uHyhYbtg4WaUyE97RLvPoddPSC2ve+KEFdnYybpCb/Yqfx9SMZ0SZa+C0F+6jRGV 8j3l5ytcmSbal3aAmTJSYg7sF1xNzSg= Received: from mail-wr1-f71.google.com (mail-wr1-f71.google.com [209.85.221.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-371-CtF6wLLiPhObmQBgYGXplw-1; Fri, 17 Jan 2025 11:53:39 -0500 X-MC-Unique: CtF6wLLiPhObmQBgYGXplw-1 X-Mimecast-MFC-AGG-ID: CtF6wLLiPhObmQBgYGXplw Received: by mail-wr1-f71.google.com with SMTP id ffacd0b85a97d-38a2140a400so1882831f8f.0 for ; Fri, 17 Jan 2025 08:53:38 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737132817; x=1737737617; h=content-transfer-encoding:mime-version:message-id:date:references :in-reply-to:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=zS2KlTiZnGu6aQKTYj9yMifWHTI073WZ1gonGb0ZMuU=; b=jCoy97TkheNFo6gfUDd/ZlRfZMLWe32x/KNwPu5B9Z6v5+/6c9YJjoY9IOTAuRPVN6 P09rZaXYDokKvgbiB8PsUt3vqFheESpujf59n/OQ4Pe77E+sayfJNX6rfzxZObbB+5X3 gD/owuQ2AGrwjGy7D+ds1i3RQggSmtLdhQQpIY9sMkH0kOEGn7WKwwcUh7smBe6dt1WQ P84ENfQUdqaaWqh8x/is5v9XEL4RljHj/ITbJ7L5tyuJaEEJd6IqC4koGFqDT9IWEGY+ ySkU22WXvgJlz8fjEas5TtTDBoD/H2LR8Lv+ddC9Ysr8sUp31cuJLjauB/UF2QmT42+h 1HcQ== X-Forwarded-Encrypted: i=1; AJvYcCW7JQC/waXPGOCZ0cp1qIXPyJZyTkNnfDoblbPw030WBj1O+hIv5XlysXxVYmBk00k/uarAXsrZW6jp7Rsklg==@lists.linux.dev X-Gm-Message-State: AOJu0YyIkjjLkUvJpP5DVuNm8SR1dN3qToEzz65ZdxOEVpyvZ4oRdQPR LVA/Yrz6Sv7r6EnoJCrerYEXn8iiQ5fRnAyvGJcjTSboqIPfWyyEe4+v81GhT47kHTJQE8zrQVv 75ahCbslJV+XFxQNIZlJIxvsniAFWcgCX17poalEv72FJ3vR4DSHF3rOypKmHHWct X-Gm-Gg: ASbGncvGo5+waFfs4fkPtgLR+hEdnHmktSRJapDp9bZoo/eAv6FI4KO7QnvvUYf9TTo 2Ph7kNm3GlgZhHxhbDE0pzMSY2mHopA4BaTgGNvcWV0NhsrxLfrgiA005GU7pSM9rma1ziPuFim A2F+EiPpUtVyTXf7R2CIbqq+cBd6o/uh8WmGXtCjazoHX7Ruc9WJghkVB2jgohOh+cJH4CL9ykK UKzamDmMYIbPhP/ofGwbmLCWUrPIIgpAl+dt/1pRJ6JOJd9UzjSjmX177ngA3tcDSj5xwTY+a6I 4TcmKAxztKmWliglC/GpNcT2V85B8t6+Hgsy6bE1sQ== X-Received: by 2002:a5d:6da4:0:b0:38b:e32a:10a6 with SMTP id ffacd0b85a97d-38bf57a9932mr3394936f8f.41.1737132817365; Fri, 17 Jan 2025 08:53:37 -0800 (PST) X-Google-Smtp-Source: AGHT+IGgg/tn25OOVaMMDMlg+FUQ1oxKEYFZYFbhCFaXXhI5w4vXLbPHJsNwkMnyD4aTZ+GRHXAAGg== X-Received: by 2002:a5d:6da4:0:b0:38b:e32a:10a6 with SMTP id ffacd0b85a97d-38bf57a9932mr3394815f8f.41.1737132816637; Fri, 17 Jan 2025 08:53:36 -0800 (PST) Received: from vschneid-thinkpadt14sgen2i.remote.csb (213-44-141-166.abo.bbox.fr. [213.44.141.166]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-38bf3221db2sm2893201f8f.29.2025.01.17.08.53.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jan 2025 08:53:36 -0800 (PST) From: Valentin Schneider To: Jann Horn Cc: linux-kernel@vger.kernel.org, x86@kernel.org, virtualization@lists.linux.dev, linux-arm-kernel@lists.infradead.org, loongarch@lists.linux.dev, linux-riscv@lists.infradead.org, linux-perf-users@vger.kernel.org, xen-devel@lists.xenproject.org, kvm@vger.kernel.org, linux-arch@vger.kernel.org, rcu@vger.kernel.org, linux-hardening@vger.kernel.org, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, bpf@vger.kernel.org, bcm-kernel-feedback-list@broadcom.com, Juergen Gross , Ajay Kaher , Alexey Makhalov , Russell King , Catalin Marinas , Will Deacon , Huacai Chen , WANG Xuerui , Paul Walmsley , Palmer Dabbelt , Albert Ou , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , "H. Peter Anvin" , Peter Zijlstra , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , "Liang, Kan" , Boris Ostrovsky , Josh Poimboeuf , Pawan Gupta , Sean Christopherson , Paolo Bonzini , Andy Lutomirski , Arnd Bergmann , Frederic Weisbecker , "Paul E. McKenney" , Jason Baron , Steven Rostedt , Ard Biesheuvel , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juri Lelli , Clark Williams , Yair Podemsky , Tomas Glozar , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Kees Cook , Andrew Morton , Christoph Hellwig , Shuah Khan , Sami Tolvanen , Miguel Ojeda , Alice Ryhl , "Mike Rapoport (Microsoft)" , Samuel Holland , Rong Xu , Nicolas Saenz Julienne , Geert Uytterhoeven , Yosry Ahmed , "Kirill A. Shutemov" , "Masami Hiramatsu (Google)" , Jinghao Jia , Luis Chamberlain , Randy Dunlap , Tiezhu Yang Subject: Re: [PATCH v4 29/30] x86/mm, mm/vmalloc: Defer flush_tlb_kernel_range() targeting NOHZ_FULL CPUs In-Reply-To: References: <20250114175143.81438-1-vschneid@redhat.com> <20250114175143.81438-30-vschneid@redhat.com> Date: Fri, 17 Jan 2025 17:53:33 +0100 Message-ID: Precedence: bulk X-Mailing-List: virtualization@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-Mimecast-Spam-Score: 0 X-Mimecast-MFC-PROC-ID: D6iodIH0CnwsTLb85gLfoKxsKOrVWpVmEiPurppzMug_1737132817 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable On 17/01/25 16:52, Jann Horn wrote: > On Fri, Jan 17, 2025 at 4:25=E2=80=AFPM Valentin Schneider wrote: >> On 14/01/25 19:16, Jann Horn wrote: >> > On Tue, Jan 14, 2025 at 6:51=E2=80=AFPM Valentin Schneider wrote: >> >> vunmap()'s issued from housekeeping CPUs are a relatively common sour= ce of >> >> interference for isolated NOHZ_FULL CPUs, as they are hit by the >> >> flush_tlb_kernel_range() IPIs. >> >> >> >> Given that CPUs executing in userspace do not access data in the vmal= loc >> >> range, these IPIs could be deferred until their next kernel entry. >> >> >> >> Deferral vs early entry danger zone >> >> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D >> >> >> >> This requires a guarantee that nothing in the vmalloc range can be vu= nmap'd >> >> and then accessed in early entry code. >> > >> > In other words, it needs a guarantee that no vmalloc allocations that >> > have been created in the vmalloc region while the CPU was idle can >> > then be accessed during early entry, right? >> >> I'm not sure if that would be a problem (not an mm expert, please do >> correct me) - looking at vmap_pages_range(), flush_cache_vmap() isn't >> deferred anyway. > > flush_cache_vmap() is about stuff like flushing data caches on > architectures with virtually indexed caches; that doesn't do TLB > maintenance. When you look for its definition on x86 or arm64, you'll > see that they use the generic implementation which is simply an empty > inline function. > >> So after vmapping something, I wouldn't expect isolated CPUs to have >> invalid TLB entries for the newly vmapped page. >> >> However, upon vunmap'ing something, the TLB flush is deferred, and thus >> stale TLB entries can and will remain on isolated CPUs, up until they >> execute the deferred flush themselves (IOW for the entire duration of th= e >> "danger zone"). >> >> Does that make sense? > > The design idea wrt TLB flushes in the vmap code is that you don't do > TLB flushes when you unmap stuff or when you map stuff, because doing > TLB flushes across the entire system on every vmap/vunmap would be a > bit costly; instead you just do batched TLB flushes in between, in > __purge_vmap_area_lazy(). > > In other words, the basic idea is that you can keep calling vmap() and > vunmap() a bunch of times without ever doing TLB flushes until you run > out of virtual memory in the vmap region; then you do one big TLB > flush, and afterwards you can reuse the free virtual address space for > new allocations again. > > So if you "defer" that batched TLB flush for CPUs that are not > currently running in the kernel, I think the consequence is that those > CPUs may end up with incoherent TLB state after a reallocation of the > virtual address space. > Ah, gotcha, thank you for laying this out! In which case yes, any vmalloc that occurred while an isolated CPU was NOHZ-FULL can be an issue if said CPU accesses it during early entry; > Actually, I think this would mean that your optimization is disallowed > at least on arm64 - I'm not sure about the exact wording, but arm64 > has a "break before make" rule that forbids conflicting writable > address translations or something like that. > On the bright side of things, arm64 is not as bad as x86 when it comes to IPI'ing isolated CPUs :-) I'll add that to my notes, thanks! > (I said "until you run out of virtual memory in the vmap region", but > that's not actually true - see the comment above lazy_max_pages() for > an explanation of the actual heuristic. You might be able to tune that > a bit if you'd be significantly happier with less frequent > interruptions, or something along those lines.)