From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EEBC31E8855 for ; Mon, 20 Jan 2025 16:09:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737389383; cv=none; b=HeA/8hbMyo7GZFejGg7VUzfmIh2ZxIpkUGAIJDPme1jVbmutpm7x7X8uegZHT8OwuKyUF/fBhniTopUbARqzNWAExHddOcRo8I7p6Ie0n+xzhGcxYIe9lWSMOeSeE22mh+luju7H9MmmeZDjz74GZKi4Z88Wxz301XqR2t7nJyc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737389383; c=relaxed/simple; bh=SNeDTVMDd6LFz2Ibm+JZz5IT863oQDJPOaC5/U8lcJk=; h=From:To:Cc:Subject:In-Reply-To:References:Date:Message-ID: MIME-Version:Content-Type; b=IfRqFExovW53fYAL9h01f4O3gFDiNP5XVfrUS/q6ZCjVq64MydsRtkOX/xLzwR7YnmVOV+ihuuvfyK+Mk9jwJs+1SsPo4HTPOgA6oVhZU9EJ/HRyDu3s3VPkCGOVqhVIfPuYTkl3bKnERi08nbZJtiqB7a/1c8g+j+eJR3yp1jE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=YCuIi2gq; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="YCuIi2gq" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1737389381; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=SNeDTVMDd6LFz2Ibm+JZz5IT863oQDJPOaC5/U8lcJk=; b=YCuIi2gqY27/RenYo+i67O1JK7ekVse6cs1fAn7xO6n6ukpSgzz3wokTwcvwoSIDPChQFO yupdrm0+05XyuGVby5P4zH4FFDrt382SGWD+3X6j1lkD4aqbEDrQBw4NKoLCHKliqBmMj5 WCxV8GAHEnJFtnnMaRdblDg5U18r8xM= Received: from mail-wr1-f72.google.com (mail-wr1-f72.google.com [209.85.221.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-207-jkZGZXSNMGy4w5nm-ssG2Q-1; Mon, 20 Jan 2025 11:09:39 -0500 X-MC-Unique: jkZGZXSNMGy4w5nm-ssG2Q-1 X-Mimecast-MFC-AGG-ID: jkZGZXSNMGy4w5nm-ssG2Q Received: by mail-wr1-f72.google.com with SMTP id ffacd0b85a97d-3862e986d17so1960707f8f.3 for ; Mon, 20 Jan 2025 08:09:39 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737389378; x=1737994178; h=content-transfer-encoding:mime-version:message-id:date:references :in-reply-to:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=uPYF3/ytKmNcBW6xY8UN6OXUapKiaxgyBXzq3dB0L9c=; b=SV+BM0bmfvdVdffvOxnFsEktb26PnYmyZujvUtJiPE5DpJ0IcpwW5AErjf+T5Ts5Oe eE17tQp+8xYfOI+A+115maGGOjJdEz3tCNFk32IFs796hGqxmWBr2O+WDjc53VjWfYHo ewHq54gYA1xix0VHn3vogj/erutDf/Rw5Eodh14KUANsVZD56VyoosXqE3dhCILfurFq b7WB/2uZDrbYV/mfW5xfbuCcqyFGO7vvT6lss+NUqre6y2pxGSwRQUTYBJgHiOnbOcBj +Vgxn3w0zRAg2RXew1NxicY5tUoCsRJq8vgkFQSlAv0g0vmoCit0kmocqpec13RJMVbn tpRQ== X-Forwarded-Encrypted: i=1; AJvYcCVm8eztY3mdRSignpHdgrI15xh2a+hjVDgOvXMFcgFhyCjJUAS7Mz/Ep1UIbFDiuPmEqo4CTADePmLZiKYMPg==@lists.linux.dev X-Gm-Message-State: AOJu0Yy5AAupunADklx7842+NL/monhVQhaLvJiyHazeq1aNUew6dUFx SVk2/XDDI7CIWOcYe+PHAeTqWtT7dOONV78iVZLNQmyRgQG94cHB1VCCMDOobLgxcmNgSryoxe9 iERcEx3hB2OPuId79p4VAhaRJaFqgnflQJIJDh3KRkVmYku3IY0QTvjJD9O5LrJs9 X-Gm-Gg: ASbGncvHA6DM0smw23vi0Q9sqXzxdsp5LiXdpH/XR/BSfGRNjqLTMyygCSQEmz9zzXK zWkGoOj/7Ug9qjwcK+84TOSdYvN9gBq/s98DZj7vsdyEB5zqn/AAoy0KOqsodCbe4FD0Emzb19Z aKCA2VHCWHwUSTVN1ymbESuRMrZfK6eyACMJG3hf2KGK6Qdb11ItPXxN9V8XWw/49CE3jF5arv/ SwzWC+PHWxWFJYfuqrYeoKdNI9ogODqIVsg6RrRvgxI5aHparpPuTnao+BcqW3VGOnNO1RVHwDU JJtGz3FJGRNttyAual0Apa20aP2uldZsoN1c+PIoM8L9KvxsbX4GFeo= X-Received: by 2002:adf:f682:0:b0:38b:e26d:ea0b with SMTP id ffacd0b85a97d-38bf566c314mr10592135f8f.25.1737389378237; Mon, 20 Jan 2025 08:09:38 -0800 (PST) X-Google-Smtp-Source: AGHT+IE7jMmKmijGDhYJpj3v2AzDeWwh7lxfHaye7x+JEw9SklvyOVZq+Wb4Niy5wxwqRIeZEeA2rg== X-Received: by 2002:adf:f682:0:b0:38b:e26d:ea0b with SMTP id ffacd0b85a97d-38bf566c314mr10592030f8f.25.1737389377661; Mon, 20 Jan 2025 08:09:37 -0800 (PST) Received: from vschneid-thinkpadt14sgen2i.remote.csb (213-44-141-166.abo.bbox.fr. [213.44.141.166]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-38bf3221b70sm10695813f8f.26.2025.01.20.08.09.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 20 Jan 2025 08:09:37 -0800 (PST) From: Valentin Schneider To: Uladzislau Rezki Cc: Uladzislau Rezki , Jann Horn , linux-kernel@vger.kernel.org, x86@kernel.org, virtualization@lists.linux.dev, linux-arm-kernel@lists.infradead.org, loongarch@lists.linux.dev, linux-riscv@lists.infradead.org, linux-perf-users@vger.kernel.org, xen-devel@lists.xenproject.org, kvm@vger.kernel.org, linux-arch@vger.kernel.org, rcu@vger.kernel.org, linux-hardening@vger.kernel.org, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, bpf@vger.kernel.org, bcm-kernel-feedback-list@broadcom.com, Juergen Gross , Ajay Kaher , Alexey Makhalov , Russell King , Catalin Marinas , Will Deacon , Huacai Chen , WANG Xuerui , Paul Walmsley , Palmer Dabbelt , Albert Ou , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , "H. Peter Anvin" , Peter Zijlstra , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , "Liang, Kan" , Boris Ostrovsky , Josh Poimboeuf , Pawan Gupta , Sean Christopherson , Paolo Bonzini , Andy Lutomirski , Arnd Bergmann , Frederic Weisbecker , "Paul E. McKenney" , Jason Baron , Steven Rostedt , Ard Biesheuvel , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juri Lelli , Clark Williams , Yair Podemsky , Tomas Glozar , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Kees Cook , Andrew Morton , Christoph Hellwig , Shuah Khan , Sami Tolvanen , Miguel Ojeda , Alice Ryhl , "Mike Rapoport (Microsoft)" , Samuel Holland , Rong Xu , Nicolas Saenz Julienne , Geert Uytterhoeven , Yosry Ahmed , "Kirill A. Shutemov" , "Masami Hiramatsu (Google)" , Jinghao Jia , Luis Chamberlain , Randy Dunlap , Tiezhu Yang Subject: Re: [PATCH v4 29/30] x86/mm, mm/vmalloc: Defer flush_tlb_kernel_range() targeting NOHZ_FULL CPUs In-Reply-To: References: <20250114175143.81438-1-vschneid@redhat.com> <20250114175143.81438-30-vschneid@redhat.com> Date: Mon, 20 Jan 2025 17:09:34 +0100 Message-ID: Precedence: bulk X-Mailing-List: virtualization@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-Mimecast-Spam-Score: 0 X-Mimecast-MFC-PROC-ID: uvPbWUtOeai9GhILuWfp0P5dBaRKL-A7un60B-Mz86k_1737389378 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable On 20/01/25 12:15, Uladzislau Rezki wrote: > On Fri, Jan 17, 2025 at 06:00:30PM +0100, Valentin Schneider wrote: >> On 17/01/25 17:11, Uladzislau Rezki wrote: >> > On Fri, Jan 17, 2025 at 04:25:45PM +0100, Valentin Schneider wrote: >> >> On 14/01/25 19:16, Jann Horn wrote: >> >> > On Tue, Jan 14, 2025 at 6:51=E2=80=AFPM Valentin Schneider wrote: >> >> >> vunmap()'s issued from housekeeping CPUs are a relatively common s= ource of >> >> >> interference for isolated NOHZ_FULL CPUs, as they are hit by the >> >> >> flush_tlb_kernel_range() IPIs. >> >> >> >> >> >> Given that CPUs executing in userspace do not access data in the v= malloc >> >> >> range, these IPIs could be deferred until their next kernel entry. >> >> >> >> >> >> Deferral vs early entry danger zone >> >> >> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D >> >> >> >> >> >> This requires a guarantee that nothing in the vmalloc range can be= vunmap'd >> >> >> and then accessed in early entry code. >> >> > >> >> > In other words, it needs a guarantee that no vmalloc allocations th= at >> >> > have been created in the vmalloc region while the CPU was idle can >> >> > then be accessed during early entry, right? >> >> >> >> I'm not sure if that would be a problem (not an mm expert, please do >> >> correct me) - looking at vmap_pages_range(), flush_cache_vmap() isn't >> >> deferred anyway. >> >> >> >> So after vmapping something, I wouldn't expect isolated CPUs to have >> >> invalid TLB entries for the newly vmapped page. >> >> >> >> However, upon vunmap'ing something, the TLB flush is deferred, and th= us >> >> stale TLB entries can and will remain on isolated CPUs, up until they >> >> execute the deferred flush themselves (IOW for the entire duration of= the >> >> "danger zone"). >> >> >> >> Does that make sense? >> >> >> > Probably i am missing something and need to have a look at your patche= s, >> > but how do you guarantee that no-one map same are that you defer for T= LB >> > flushing? >> > >> >> That's the cool part: I don't :') >> > Indeed, sounds unsafe :) Then we just do not need to free areas. > >> For deferring instruction patching IPIs, I (well Josh really) managed to >> get instrumentation to back me up and catch any problematic area. >> >> I looked into getting something similar for vmalloc region access in >> .noinstr code, but I didn't get anywhere. I even tried using emulated >> watchpoints on QEMU to watch the whole vmalloc range, but that went abou= t >> as well as you could expect. >> >> That left me with staring at code. AFAICT the only vmap'd thing that is >> accessed during early entry is the task stack (CONFIG_VMAP_STACK), which >> itself cannot be freed until the task exits - thus can't be subject to >> invalidation when a task is entering kernelspace. >> >> If you have any tracing/instrumentation suggestions, I'm all ears (eyes?= ). >> > As noted before, we defer flushing for vmalloc. We have a lazy-threshold > which can be exposed(if you need it) over sysfs for tuning. So, we can ad= d it. > In a CPU isolation / NOHZ_FULL context, isolated CPUs will be running a single userspace application that will never enter the kernel, unless forced to by some interference (e.g. IPI sent from a housekeeping CPU). Increasing the lazy threshold would unfortunately only delay the interference - housekeeping CPUs are free to run whatever, and so they will eventually cause the lazy threshold to be hit and IPI all the CPUs, including the isolated/NOHZ_FULL ones. I was thinking maybe we could subdivide the vmap space into two regions with their own thresholds, but a task may allocate/vmap stuff while on a HK CPU and be moved to an isolated CPU afterwards, and also I still don't have any strong guarantee about what accesses an isolated CPU can do in its early entry code :( > -- > Uladzislau Rezki