From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f180.google.com (mail-pf1-f180.google.com [209.85.210.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E40F32E73A for ; Thu, 18 Dec 2025 22:24:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.180 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766096694; cv=none; b=WN9sJ9raBfPVzVQDJQM+PFb+U++KpTspIpQlx9m9YVqiKYuxcBpAslZBqO912dmIN214oigOFZ80qxIlCUohCKVOepjfwq9Ssf7VofT6gxYDu2dqOyiqOiLW7l14VLDmHDVNYIYYqmHiOl10qFftUX1ODgCUwPwZ87OdwfKotgE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766096694; c=relaxed/simple; bh=/eLfiFUuH9yvahCPVcjPkezg5kYE43CsIOzLJ1Q0GsQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=WfSU7VWdcU4DzTD8cuTP5roC9decoa1zU4sH2p9OekluP0CJhCN9wP3Cm99958HRGByLS+5E3hexQg10pHwOh1MuhzK2FUrujMeYMSwJ4jflj5QV9JpzXTOloJr2gBYAI4tgFfzYMg+81eZ+BKRLfBv2NbCjC5OYyq8xFBdKVPA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=l3GqovP2; arc=none smtp.client-ip=209.85.210.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="l3GqovP2" Received: by mail-pf1-f180.google.com with SMTP id d2e1a72fcca58-7b89c1ce9easo1357776b3a.2 for ; Thu, 18 Dec 2025 14:24:52 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1766096692; x=1766701492; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:feedback-id:from:to:cc:subject:date :message-id:reply-to; bh=N8rgecbz9Q7Jzp6eajm21hLSWx7IeyY0MrZP+QvLVqg=; b=l3GqovP2ynSkiEaLpL43wxGQ5RcOjg13ZIfXrv+pepS2OygStoRAHo/l4uUYCVYbdH hpsYskwC4lfGMqrtIyQFGLWZe+2iQW4xaITFI1KDK4SwX2B/S7IkHxQg+/BeAaqo7/UQ 1R4HQWJGewveX7VP6Y1Mv/9FknAlv7lrDzjWxxatkUUzgqOwXAGfNCgJZ3mr7AejyLpk pIEngxw+l7Qn9WLONHIbrIVOz96WrxCWDj1fm/zOnp2GL9INTVOomUW7YkUp0D2CRzl/ 1Kczt6+t1UQJt1I2QKavYA+LEv1HRQ6f2nggLttXX0YE6Esn3vhEKAHgjTIQdaUuj4Ey LzQg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766096692; x=1766701492; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:feedback-id:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=N8rgecbz9Q7Jzp6eajm21hLSWx7IeyY0MrZP+QvLVqg=; b=f2f/GyIMdGG5HCbl+uIv0lE1BUAFO6ApvEN18o0lh2I86fpSn0Q7UPUVhmrATEe7NN nS/KD3YUHWNozaqegUpocASMXv8c2/XDy9e+qLnJPLJwdrkTENY7aBMBl5H8tjKux6Gx rRVNMAbWj75ZOW19DTe/wsLLAU1VAAR/+1nWliYwkz0KQ+wVLntFpvdrC64JRXnEXBJS Luxw5guQi9IkETZ4Vjmt17H3a5utGIalKJxqiZgZmg1t7fNKJpw4kcyCkPqARIfgfktk PSX1kYposCqLzMw7ikZXC6ojV5JmomHvOdwYz790Z/YDCvepNEC/QVE3UjppSf7nQsVJ NPkg== X-Forwarded-Encrypted: i=1; AJvYcCV1kSjxuRvRZ/nTxiKNI+y5Wz30oexWM6ZpcrX9oJDqcJt2xy0AXi6ijO8qZjeDSnGY/jYLa3Lv8kmBTZ4=@vger.kernel.org X-Gm-Message-State: AOJu0Yw/Kv+mGKIOSP6KNcqdWsoYI7iN71JzCqb9nMVRDqvVWMLrJ4+2 dND3eF0PpBBFs8/sxGFzz+EeeOj9KVwkpwO/xlk4D5s9B4o9yzZldW0vF624orcsC1E= X-Gm-Gg: AY/fxX43utiYME0zMOhm4OgOC4B1J1sBa3CsSzej2RfQ+GEa8AR+MzDJbgz8jZ4tvg9 FqQt2D+GXMK2sCj9swiBCwSNqQw7lRPemlbwkFtqBxA79uVC4P535glvOtO+jfXAN+ZhR9ZnB9i pplVrLhctnEsdvD/CiM1yksjXIbeLH0jhk/Qt4f9eC6Za1J3gGHlj4opjYHDH4aRnMtqIuLyQBf UqpJWNMcx0Ka4SU12tl5KUevrIkvM7mkjQmUtLeXOFrjJzKAykJNU+FiKy11nHZ9VnTndIYq5B9 IEtw+HhP1BuNV6FeLadjxLKCEav0mY+oauBigZ2Zt3F0THjvZKzhNol+WaBtCF3ROTZogSUcdDj ZPWhDWl8sheflp6FenbbbO/fXlibcsNm4JE1MSi9Aw2gNyRlIBzwsizguxKkwgMWvbeDsfXxnkE lIvDJeSMca6//YyYYwALH24TYseAyVxY8H2A7T9LS7v76WX2JX2idYDPideO//jKUisgfX+ePxn 7oLZfgYPyvKeDH0Kix0ioH6Rg== X-Google-Smtp-Source: AGHT+IEspr+mA8yBfzyu92XEiMHrHEfySL5XUdcBH3qNwt/XL6ga6kYNOb/BEKUmPIRMr7Vo9NJQvA== X-Received: by 2002:a05:622a:1801:b0:4f0:2b7d:5e05 with SMTP id d75a77b69052e-4f4abdc5894mr6756991cf.73.1766089353934; Thu, 18 Dec 2025 12:22:33 -0800 (PST) Received: from fauth-a2-smtp.messagingengine.com (fauth-a2-smtp.messagingengine.com. [103.168.172.201]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-88d973a7f17sm3470116d6.22.2025.12.18.12.22.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 18 Dec 2025 12:22:33 -0800 (PST) Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfauth.phl.internal (Postfix) with ESMTP id 6CAEFF40068; Thu, 18 Dec 2025 15:22:32 -0500 (EST) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Thu, 18 Dec 2025 15:22:32 -0500 X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: gggruggvucftvghtrhhoucdtuddrgeefgedrtddtgdegieefjecutefuodetggdotefrod ftvfcurfhrohhfihhlvgemucfhrghsthforghilhdpuffrtefokffrpgfnqfghnecuuegr ihhlohhuthemuceftddtnecusecvtfgvtghiphhivghnthhsucdlqddutddtmdenucfjug hrpeffhffvvefukfhfgggtuggjsehttdertddttddvnecuhfhrohhmpeeuohhquhhnucfh vghnghcuoegsohhquhhnrdhfvghnghesghhmrghilhdrtghomheqnecuggftrfgrthhtvg hrnheptddtudegueevgefhgfeuffetffeuheekgedtffefhefhjeffhffgfeeggeetgefh necuffhomhgrihhnpegvfhhfihgtihhoshdrtghomhenucevlhhushhtvghrufhiiigvpe dtnecurfgrrhgrmhepmhgrihhlfhhrohhmpegsohhquhhnodhmvghsmhhtphgruhhthhhp vghrshhonhgrlhhithihqdeiledvgeehtdeigedqudejjeekheehhedvqdgsohhquhhnrd hfvghngheppehgmhgrihhlrdgtohhmsehfihigmhgvrdhnrghmvgdpnhgspghrtghpthht ohepfeefpdhmohguvgepshhmthhpohhuthdprhgtphhtthhopehmrghthhhivghurdguvg hsnhhohigvrhhssegvfhhfihgtihhoshdrtghomhdprhgtphhtthhopehjohgvlhesjhho vghlfhgvrhhnrghnuggvshdrohhrghdprhgtphhtthhopehprghulhhmtghksehkvghrnh gvlhdrohhrghdprhgtphhtthhopehlihhnuhigqdhkvghrnhgvlhesvhhgvghrrdhkvghr nhgvlhdrohhrghdprhgtphhtthhopehnphhighhgihhnsehgmhgrihhlrdgtohhmpdhrtg hpthhtohepmhhpvgesvghllhgvrhhmrghnrdhiugdrrghupdhrtghpthhtohepghhrvghg khhhsehlihhnuhigfhhouhhnuggrthhiohhnrdhorhhgpdhrtghpthhtohepsghighgvrg hshieslhhinhhuthhrohhnihigrdguvgdprhgtphhtthhopeifihhllheskhgvrhhnvghl rdhorhhg X-ME-Proxy: Feedback-ID: iad51458e:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 18 Dec 2025 15:22:31 -0500 (EST) Date: Fri, 19 Dec 2025 05:22:29 +0900 From: Boqun Feng To: Mathieu Desnoyers Cc: Joel Fernandes , "Paul E. McKenney" , linux-kernel@vger.kernel.org, Nicholas Piggin , Michael Ellerman , Greg Kroah-Hartman , Sebastian Andrzej Siewior , Will Deacon , Peter Zijlstra , Alan Stern , John Stultz , Neeraj Upadhyay , Linus Torvalds , Andrew Morton , Frederic Weisbecker , Josh Triplett , Uladzislau Rezki , Steven Rostedt , Lai Jiangshan , Zqiang , Ingo Molnar , Waiman Long , Mark Rutland , Thomas Gleixner , Vlastimil Babka , maged.michael@gmail.com, Mateusz Guzik , Jonas Oberhauser , rcu@vger.kernel.org, linux-mm@kvack.org, lkmm@lists.linux.dev Subject: Re: [RFC PATCH v4 3/4] hazptr: Implement Hazard Pointers Message-ID: References: <20251218014531.3793471-1-mathieu.desnoyers@efficios.com> <20251218014531.3793471-4-mathieu.desnoyers@efficios.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Thu, Dec 18, 2025 at 12:35:18PM -0500, Mathieu Desnoyers wrote: > On 2025-12-18 03:36, Boqun Feng wrote: > > On Wed, Dec 17, 2025 at 08:45:30PM -0500, Mathieu Desnoyers wrote: > [...] > > > +static inline > > > +void *hazptr_acquire(struct hazptr_ctx *ctx, void * const * addr_p) > > > +{ > > > + struct hazptr_slot *slot = NULL; > > > + void *addr, *addr2; > > > + > > > + /* > > > + * Load @addr_p to know which address should be protected. > > > + */ > > > + addr = READ_ONCE(*addr_p); > > > + for (;;) { > > > + if (!addr) > > > + return NULL; > > > + guard(preempt)(); > > > + if (likely(!hazptr_slot_is_backup(ctx, slot))) { > > > + slot = hazptr_get_free_percpu_slot(); > > > > I need to continue share my concerns about this "allocating slot while > > protecting" pattern. Here realistically, we will go over a few of the > > per-CPU hazard pointer slots *every time* instead of directly using a > > pre-allocated hazard pointer slot. > > No, that's not the expected fast-path with CONFIG_PREEMPT_HAZPTR=y > (introduced in patch 4/4). > I see, I was missing the patch #4, will take a look and reply accordingly. > With PREEMPT_HAZPTR, using more than one hazard pointer per CPU will > only happen if there are nested hazard pointer users, which can happen > due to: > > - Holding a hazard pointer across function calls, where another hazard > pointer is used. > - Using hazard pointers from interrupt handlers (note: my current code > only does preempt disable, not irq disable, this is something I'd need > to change if we wish to acquire/release hazard pointers from interrupt > handlers). But even that should be a rare event. > > So the fast-path has an initial state where there are no hazard pointers > in use on the CPU, which means hazptr_acquire() finds its empty slot at > index 0. > > > Could you utilize this[1] to see a > > comparison of the reader-side performance against RCU/SRCU? > > Good point ! Let's see. > > On a AMD 2x EPYC 9654 96-Core Processor with 192 cores, > hyperthreading disabled, > CONFIG_PREEMPT=y, > CONFIG_PREEMPT_RCU=y, > CONFIG_PREEMPT_HAZPTR=y. > > scale_type ns > ----------------------- > hazptr-smp-mb 13.1 <- this implementation > hazptr-barrier 11.5 <- replace smp_mb() on acquire with barrier(), requires IPIs on synchronize. > hazptr-smp-mb-hlist 12.7 <- replace per-task hp context and per-cpu overflow lists by hlist. > rcu 17.0 > srcu 20.0 > srcu-fast 1.5 > rcu-tasks 0.0 > rcu-trace 1.7 > refcnt 1148.0 > rwlock 1190.0 > rwsem 4199.3 > lock 41070.6 > lock-irq 46176.3 > acqrel 1.1 > > So only srcu-fast, rcu-tasks, rcu-trace and a plain acqrel > appear to beat hazptr read-side performance. > Could you also see the reader-side performance impact when the percpu hazard pointer slots are used up? I.e. the worst case. > [...] > > > > +/* > > > + * Perform piecewise iteration on overflow list waiting until "addr" is > > > + * not present. Raw spinlock is released and taken between each list > > > + * item and busy loop iteration. The overflow list generation is checked > > > + * each time the lock is taken to validate that the list has not changed > > > + * before resuming iteration or busy wait. If the generation has > > > + * changed, retry the entire list traversal. > > > + */ > > > +static > > > +void hazptr_synchronize_overflow_list(struct overflow_list *overflow_list, void *addr) > > > +{ > > > + struct hazptr_backup_slot *backup_slot; > > > + uint64_t snapshot_gen; > > > + > > > + raw_spin_lock(&overflow_list->lock); > > > +retry: > > > + snapshot_gen = overflow_list->gen; > > > + list_for_each_entry(backup_slot, &overflow_list->head, node) { > > > + /* Busy-wait if node is found. */ > > > + while (smp_load_acquire(&backup_slot->slot.addr) == addr) { /* Load B */ > > > + raw_spin_unlock(&overflow_list->lock); > > > + cpu_relax(); > > > > I think we should prioritize the scan thread solution [2] instead of > > busy waiting hazrd pointer updaters, because when we have multiple > > hazard pointer usages we would want to consolidate the scans from > > updater side. > > I agree that batching scans with a worker thread is a logical next step. > > > If so, the whole ->gen can be avoided. > > How would it allow removing the generation trick without causing long > raw spinlock latencies ? > Because we won't need to busy-wait for the readers to go away, we can check whether they are still there in the next scan. so: list_for_each_entry(backup_slot, &overflow_list->head, node) { /* Busy-wait if node is found. */ if (smp_load_acquire(&backup_slot->slot.addr) == addr) { /* Load B */ > > > > However this ->gen idea does seem ot resolve another issue for me, I'm > > trying to make shazptr critical section preemptive by using a per-task > > backup slot (if you recall, this is your idea from the hallway > > discussions we had during LPC 2024), > > I honestly did not remember. It's been a whole year! ;-) > > > and currently I could not make it > > work because the following sequeue: > > > > 1. CPU 0 already has one pointer protected. > > > > 2. CPU 1 begins the updater scan, and it scans the list of preempted > > hazard pointer readers, no reader. > > > > 3. CPU 0 does a context switch, it stores the current hazard pointer > > value to the current task's ->hazard_slot (let's say the task is task > > A), and add it to the list of preempted hazard pointer readers. > > > > 4. CPU 0 clears its percpu hazptr_slots for the next task (B). > > > > 5. CPU 1 continues the updater scan, and it scans the percpu slot of > > CPU 0, and finds no reader. > > > > in this situation, updater will miss a reader. But if we add a > > generation snapshotting at step 2 and generation increment at step 3, I > > think it'll work. > > > > IMO, if we make this work, it's better than the current backup slot > > mechanism IMO, because we only need to acquire the lock if context > > switch happens. > > With PREEMPT_HAZPTR we also only need to acquire the per-cpu overflow > list raw spinlock on context switch (preemption or blocking). The only Indeed, pre-allocating the slot on the stack to save the percpu slot when context switch seems easier and quite smart ;-) Let me take a look. Regards, Boqun > other case requiring it is hazptr nested usage (more than 8 active > hazptr) on a thread context + nested irqs. > > Thanks, > > Mathieu > > -- > Mathieu Desnoyers > EfficiOS Inc. > https://www.efficios.com