* [PATCH net v2 1/1] net/sched: cls_flower: avoid stale mask references after delete
From: Ren Wei @ 2026-04-21 16:03 UTC (permalink / raw)
To: netdev
Cc: jhs, jiri, davem, edumazet, kuba, pabeni, horms, sbrivio, vladbu,
yuantan098, yifanwucs, tomapufckgml, bird, kanolyc, z1652074432,
n05ec
From: Yuhang Zheng <z1652074432@gmail.com>
cls_flower keeps filter and mask state separately. After a filter is
removed or replaced, some paths can still need the mask data associated
with that filter.
Cache the mask key and dissector in struct cls_fl_filter when the mask
is assigned, and use the cached copies in dump and offload paths. This
avoids depending on the external mask object's lifetime after delete or
replace.
Fixes: 92149190067d ("net: sched: flower: set unlocked flag for flower proto ops")
Cc: stable@kernel.org
Reported-by: Yuan Tan <yuantan098@gmail.com>
Reported-by: Yifan Wu <yifanwucs@gmail.com>
Reported-by: Juefei Pu <tomapufckgml@gmail.com>
Reported-by: Xin Liu <bird@lzu.edu.cn>
Tested-by: Yucheng Lu <kanolyc@gmail.com>
Signed-off-by: Yuhang Zheng <z1652074432@gmail.com>
Signed-off-by: Ren Wei <n05ec@lzu.edu.cn>
---
changes in v2:
- target the net tree instead of nf
- fix the Fixes tag to the first triggerable introduction
- correct the Yuan Tan Reported-by address and add the forwarder sign-off
- v1 Link: https://lore.kernel.org/all/0fdcae6ac3e07afbbd43958f6b42e2ed6281e3d2.1773559972.git.z1652074432@gmail.com
net/sched/cls_flower.c | 19 ++++++++++++++-----
1 file changed, 14 insertions(+), 5 deletions(-)
diff --git a/net/sched/cls_flower.c b/net/sched/cls_flower.c
index 099ff6a3e1f5..c1f10b4ec748 100644
--- a/net/sched/cls_flower.c
+++ b/net/sched/cls_flower.c
@@ -124,8 +124,10 @@ struct cls_fl_head {
struct cls_fl_filter {
struct fl_flow_mask *mask;
+ struct flow_dissector mask_dissector;
struct rhash_head ht_node;
struct fl_flow_key mkey;
+ struct fl_flow_key mask_key;
struct tcf_exts exts;
struct tcf_result res;
struct fl_flow_key key;
@@ -445,6 +447,12 @@ static void fl_destroy_filter_work(struct work_struct *work)
__fl_destroy_filter(f);
}
+static void fl_filter_copy_mask(struct cls_fl_filter *f)
+{
+ f->mask_key = f->mask->key;
+ f->mask_dissector = f->mask->dissector;
+}
+
static void fl_hw_destroy_filter(struct tcf_proto *tp, struct cls_fl_filter *f,
bool rtnl_held, struct netlink_ext_ack *extack)
{
@@ -476,8 +484,8 @@ static int fl_hw_replace_filter(struct tcf_proto *tp,
tc_cls_common_offload_init(&cls_flower.common, tp, f->flags, extack);
cls_flower.command = FLOW_CLS_REPLACE;
cls_flower.cookie = (unsigned long) f;
- cls_flower.rule->match.dissector = &f->mask->dissector;
- cls_flower.rule->match.mask = &f->mask->key;
+ cls_flower.rule->match.dissector = &f->mask_dissector;
+ cls_flower.rule->match.mask = &f->mask_key;
cls_flower.rule->match.key = &f->mkey;
cls_flower.classid = f->res.classid;
@@ -2489,6 +2497,7 @@ static int fl_change(struct net *net, struct sk_buff *in_skb,
err = fl_check_assign_mask(head, fnew, fold, mask);
if (err)
goto unbind_filter;
+ fl_filter_copy_mask(fnew);
err = fl_ht_insert_unique(fnew, fold, &in_ht);
if (err)
@@ -2705,8 +2714,8 @@ static int fl_reoffload(struct tcf_proto *tp, bool add, flow_setup_cb_t *cb,
cls_flower.command = add ?
FLOW_CLS_REPLACE : FLOW_CLS_DESTROY;
cls_flower.cookie = (unsigned long)f;
- cls_flower.rule->match.dissector = &f->mask->dissector;
- cls_flower.rule->match.mask = &f->mask->key;
+ cls_flower.rule->match.dissector = &f->mask_dissector;
+ cls_flower.rule->match.mask = &f->mask_key;
cls_flower.rule->match.key = &f->mkey;
err = tc_setup_offload_action(&cls_flower.rule->action, &f->exts,
@@ -3709,7 +3718,7 @@ static int fl_dump(struct net *net, struct tcf_proto *tp, void *fh,
goto nla_put_failure_locked;
key = &f->key;
- mask = &f->mask->key;
+ mask = &f->mask_key;
skip_hw = tc_skip_hw(f->flags);
if (fl_dump_key(skb, net, key, mask))
--
2.43.0
^ permalink raw reply related
* Re: [PATCH 18/23] cpu/hotplug: Add a new cpuhp_offline_cb() API
From: Thomas Gleixner @ 2026-04-21 16:17 UTC (permalink / raw)
To: Waiman Long, Tejun Heo, Johannes Weiner, Michal Koutný,
Jonathan Corbet, Shuah Khan, Catalin Marinas, Will Deacon,
K. Y. Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
Guenter Roeck, Frederic Weisbecker, Paul E. McKenney,
Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
Uladzislau Rezki, Steven Rostedt, Mathieu Desnoyers,
Lai Jiangshan, Zqiang, Anna-Maria Behnsen, Ingo Molnar,
Chen Ridong, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman
Cc: cgroups, linux-doc, linux-kernel, linux-arm-kernel, linux-hyperv,
linux-hwmon, rcu, netdev, linux-kselftest, Costa Shulyupin,
Qiliang Yuan, Waiman Long
In-Reply-To: <20260421030351.281436-19-longman@redhat.com>
On Mon, Apr 20 2026 at 23:03, Waiman Long wrote:
> Add a new cpuhp_offline_cb() API that allows us to offline a set of
> CPUs one-by-one, run the given callback function and then bring those
> CPUs back online again while inhibiting any concurrent CPU hotplug
> operations from happening.
Please provide a properly structured change log which explains the
context, the problem and the solution in separate paragraphs and this
order. This is not new. It's documented...
> This new API can be used to enable runtime adjustment of nohz_full and
> isolcpus boot command line options. A new cpuhp_offline_cb_mode flag
> is also added to signal that the system is in this offline callback
> transient state so that some hotplug operations can be optimized out
> if we choose to.
We chose nothing.
> +#include <linux/cpumask_types.h>
What for? This header only needs a 'struct cpumask' forward declaration
so that the compiler can handle the pointer argument, no?
> +typedef int (*cpuhp_cb_t)(void *arg);
You couldn't come up with a more generic name for this, right?
> struct device;
>
> extern int lockdep_is_cpus_held(void);
> @@ -29,6 +31,8 @@ void clear_tasks_mm_cpumask(int cpu);
> int remove_cpu(unsigned int cpu);
> int cpu_device_down(struct device *dev);
> void smp_shutdown_nonboot_cpus(unsigned int primary_cpu);
> +int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg);
Ditto.
> +extern bool cpuhp_offline_cb_mode;
Groan. The only users are in the cpusets code which invokes this muck
and should therefore know what's going on, no?
> #else /* CONFIG_HOTPLUG_CPU */
>
> @@ -43,6 +47,11 @@ static inline void cpu_hotplug_disable(void) { }
> static inline void cpu_hotplug_enable(void) { }
> static inline int remove_cpu(unsigned int cpu) { return -EPERM; }
> static inline void smp_shutdown_nonboot_cpus(unsigned int primary_cpu) { }
> +static inline int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg)
> +{
> + return -EPERM;
-EPERM?
> +/**
> + * cpuhp_offline_cb - offline CPUs, invoke callback function & online CPUs afterward
> + * @mask: A mask of CPUs to be taken offline and then online
> + * @func: A callback function to be invoked while the given CPUs are offline
> + * @arg: Argument to be passed back to the callback function
> + *
> + * Return: 0 if successful, an error code otherwise
> + */
> +int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg)
> +{
> + int off_cpu, on_cpu, ret, ret2 = 0;
> +
> + if (WARN_ON_ONCE(cpumask_empty(mask) ||
> + !cpumask_subset(mask, cpu_online_mask)))
> + return -EINVAL;
No line break required. You have 100 characters.
But what's worse is that the access to cpu_online_mask is not protected
against a concurrent CPU hotplug operation.
> +
> + pr_debug("%s: begin (CPU list = %*pbl)\n", __func__, cpumask_pr_args(mask));
Tracing?
> + lock_device_hotplug();
> + cpuhp_offline_cb_mode = true;
> + /*
> + * If all offline operations succeed, off_cpu should become nr_cpu_ids.
> + */
> + for_each_cpu(off_cpu, mask) {
> + ret = device_offline(get_cpu_device(off_cpu));
> + if (unlikely(ret))
> + break;
> + }
> + if (!ret)
> + ret = func(arg);
> +
> + /* Bring previously offline CPUs back online */
> + for_each_cpu(on_cpu, mask) {
> + int retries = 0;
> +
> + if (on_cpu == off_cpu)
> + break;
> +
> +retry:
> + ret2 = device_online(get_cpu_device(on_cpu));
> +
> + /*
> + * With the unlikely event that CPU hotplug is disabled while
> + * this operation is in progress, we will need to wait a bit
> + * for hotplug to hopefully be re-enabled again. If not, print
> + * a warning and return the error.
> + *
> + * cpu_hotplug_disabled is supposed to be accessed while
> + * holding the cpu_add_remove_lock mutex. So we need to
> + * use the data_race() macro to access it here.
> + */
> + while ((ret2 == -EBUSY) && data_race(cpu_hotplug_disabled) &&
> + (++retries <= 5)) {
> + msleep(20);
> + if (!data_race(cpu_hotplug_disabled))
> + goto retry;
> + }
> + if (ret2) {
> + pr_warn("%s: Failed to bring CPU %d back online!\n",
> + __func__, on_cpu);
Provide a proper text and not this silly __func__ thing.
> + break;
> + }
> + }
TBH. This is unreviewable gunk and the whole 'unlikely event that CPU
hotplug is disabled' is just a lazy hack.
All of this can be avoided including this made up callback function.
It's not rocket science to provide:
1) A function which serializes against any other CPU hotplug
related action.
2) A function which brings the CPUs in a given CPU mask down
3) A function which brings the CPUs in a given CPU mask up
4) A function which undoes #1
Yeah I know, it's more work and not convoluted enough. But see below.
That brings me to that other hack namely cpuhp_offline_cb_mode, which
you self described as such in patch 21/23:
> + /*
> + * Hack: In cpuhp_offline_cb_mode, pretend all partitions are empty
> + * to prevent unnecessary partition invalidation.
> + */
> + if (cpuhp_offline_cb_mode)
> + return false;
> +
We are not merging hacks. End of story. But you knew that already, no?
Let's take a step back and see what you really need to achieve:
1) Update tick_nohz_full_mask
2) Update the managed interrupt mask
3) Update CPU sets
Independent of the direction of this update you need to ensure that the
affected functionality keeps working correctly.
You achieve that by bulk offlining the affected CPUs, invoking a magic
callback and then bulk onlining the affected CPUs again, which requires
that ill defined cpuhp_offline_cb_mode hackery and probably some more
hacks all over the place.
You can achieve the same by doing CPU by CPU operations in the right
order without this mode hack, when you establish proper limitations for
this:
At no point in time it's allowed to empty a CPU set or a affected CPU
mask, except when you completely undo the isolation of CPUs.
That can be computed upfront w/o changing anything at all. Once the
validity is established, the update can proceed. Or you can leave it
to user space which can keep the pieces if it gets it wrong.
That's a reasonable limitation as there is absolutely zero justification
to support something like:
housekeeping_cpus = [CPU 0], isolated_cpus = [CPU 1]
---> housekeeping_cpus = [CPU 1], isolated_cpus = [CPU 0]
just because we can with enough horrible hacks.
If you get that out of the way, then a CPU by CPU update becomes the
obvious and simplest solution. The ordering constraints can be computed
in user space upfront and there is no reason to do any of this in the
kernel itself except for an eventual validation step. It might be a tad
slower, but this is all but a hotpath operation.
Just for the record. I suggested exactly this more than a year ago and
it's still the right thing to do.
And of course neither your cover letter nor any of the patches give a
proper rationale why you think that your bulk hackery is better. For the
very simple reason that there is no rationale at all.
This bulk muck is doomed when your ultimate goal is to avoid the stop
machine dance. With a per CPU update it is actually doable without more
ill defined hacks all over the place.
1) Bring down the CPU to CPUHP_AP_SCHED_WAIT_EMPTY, which is the last
state before stop machine is invoked.
At that point:
- no user space thread is running on the CPU anymore
- everything related to this CPU has been shut down or moved
elsewhere
- interrupt managed device queues are quiesced if the CPU was
the last online one in the queue affinity mask. If not the
interrupt might still be affine to the CPU, but there is at
least one other CPU available in the mask.
2) Update the tick NOHZ handover
This can be done without going into stop machine by providing a
hotplug callback right between CPUHP_AP_SMPBOOT_THREADS and
CPUHP_AP_IRQ_AFFINITY_ONLINE.
That's trivial enough to achieve and can work independently of
NOHZ full.
3) Rework the affinity management, so that interrupt affinities can
be reassigned in the CPUHP_AP_IRQ_AFFINITY_ONLINE state.
That needs a lot of thoughts, but there is no real reason why it
can't work.
4) Flip the housekeeping CPU masks in sched_cpu_wait_empty() after
balance_hotplug_wait().
5) Bring the CPU online again.
For #2 and #3 to work you need a separate CPU mask which avoids touching
CPU online mask. For #3 this needs some more work to avoid reassigning the
interrupts once sparse_irq_lock is dropped, but the bulk is achieved
with the separate CPU mask.
No?
Thanks,
tglx
^ permalink raw reply
* Re: [PATCH net-deletions] net: remove ax25 and amateur radio (hamradio) subsystem
From: Dan Cross @ 2026-04-21 16:17 UTC (permalink / raw)
To: stephen
Cc: Jakub Kicinski, davem, netdev, edumazet, pabeni, andrew+netdev,
horms, corbet, skhan, federico.vaga, carlos.bilbao, avadhut.naik,
alexs, si.yanteng, dzm91, 2023002089, tsbogend, dsahern,
jani.nikula, mchehab+huawei, gregkh, jirislaby, tytso, herbert,
ebiggers, johannes.berg, geert, pablo, tglx, mashiro.chen, mingo,
dqfext, jreuter, sdf, pkshih, enelsonmoore, mkl, toke, kees,
jlayton, wangliang74, aha310510, takamitz, kuniyu, linux-doc,
linux-mips
In-Reply-To: <20260421065507.2c5e3ba7@phoenix.local>
On Tue, Apr 21, 2026 at 9:55 AM Stephen Hemminger
<stephen@networkplumber.org> wrote:
> On Mon, 20 Apr 2026 19:18:23 -0700
> Jakub Kicinski <kuba@kernel.org> wrote:
> > Remove the amateur radio (AX.25, NET/ROM, ROSE) protocol implementation
> > and all associated hamradio device drivers from the kernel tree.
> > This set of protocols has long been a huge bug/syzbot magnet,
> > and since nobody stepped up to help us deal with the influx
> > of the AI-generated bug reports we need to move it out of tree
> > to protect our sanity.
> >
> > The code is moved to an out-of-tree repo:
> > https://github.com/linux-netdev/mod-orphan
> > if it's cleaned up and reworked there we can accept it back.
>
> It would be good if these protocols could be done in userspace
> or with BPF?
Consensus for a userspace implementation is what folks on linux-hams
seem to be converging on.
The amateur radio protocols are more or less specific to low-speed
links, they are not particularly coupled to anything else that
requires running in the kernel, and the main coupling point (IP over
AX.25) can be implemented via TAP/TUN.
There are several popular packages that already implement AX.25 and
NET/ROM in user-space (for the interested, LinBPQ seems to be the
canonical example). The main missing piece is ROSE, but it is likely
easier to add that to an existing package, or potentially something
brand new, than keep it in the kernel.
There's no compelling reason to keep these protocols in the kernel,
whether in-tree or out-of-tree; at least, one has not been
articulated.
- Dan C.
^ permalink raw reply
* Re: [PATCH net] netconsole: avoid out-of-bounds access on empty string in trim_newline()
From: Simon Horman @ 2026-04-21 16:22 UTC (permalink / raw)
To: Breno Leitao
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Matthew Wood, netdev, linux-kernel, kernel-team,
stable
In-Reply-To: <20260420-netcons_trim_newline-v1-1-dc35889aeedf@debian.org>
On Mon, Apr 20, 2026 at 03:18:36AM -0700, Breno Leitao wrote:
> trim_newline() unconditionally dereferences s[len - 1] after computing
> len = strnlen(s, maxlen). When the string is empty, len is 0 and the
> expression underflows to s[(size_t)-1], reading (and potentially
> writing) one byte before the buffer.
>
> The two callers feed trim_newline() with the result of strscpy() from
> configfs store callbacks (dev_name_store, userdatum_value_store).
> configfs guarantees count >= 1 reaches the callback, but the byte
> itself can be NUL: a userspace write(fd, "\0", 1) leaves the
> destination empty after strscpy() and triggers the underflow. The OOB
> write only fires if the adjacent byte happens to be '\n', so this is
> not a security issue, but the access is undefined behaviour either way.
>
> This pattern is commonly flagged by LLM-based code reviewers. While it
> is not a security fix, the underlying access is undefined behaviour and
> the change is small and self-contained, so it is a reasonable candidate
> for the stable trees.
>
> Guard the dereference on a non-zero length.
>
> Fixes: ae001dc67907 ("net: netconsole: move newline trimming to function")
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Sashiko has provided some feedback on this patch.
I do not believe that should hold up progress of this patch.
But I'd appreciate it if you could look over that feedback
and see if any follow-up is warranted.
Thanks!
^ permalink raw reply
* Re: [PATCH bpf v2] bpf: Fix NULL pointer dereference in bpf_sk_storage_clone and diag paths
From: bot+bpf-ci @ 2026-04-21 16:24 UTC (permalink / raw)
To: bestswngs, davem, edumazet, kuba, pabeni
Cc: horms, martin.lau, ast, ameryhung, leon.hwang, kees, wangfushuai,
menglong8.dong, netdev, bpf, xmei5, bestswngs, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, clm, ihor.solodrai
In-Reply-To: <20260421155024.997077-2-bestswngs@gmail.com>
[-- Attachment #1: Type: text/plain, Size: 1555 bytes --]
> diff --git a/net/core/bpf_sk_storage.c b/net/core/bpf_sk_storage.c
> index 14eb7812bda4..b2cd2b0292cb 100644
> --- a/net/core/bpf_sk_storage.c
> +++ b/net/core/bpf_sk_storage.c
[ ... ]
> @@ -596,6 +596,8 @@ static int bpf_sk_storage_diag_put_all(struct sock *sk, struct sk_buff *skb,
> saved_len = skb->len;
> hlist_for_each_entry_rcu(selem, &sk_storage->list, snode) {
> smap = rcu_dereference(SDATA(selem)->smap);
> + if (!smap)
> + continue;
> diag_size += nla_value_size(smap->map.value_size);
>
> if (nla_stgs && diag_get(SDATA(selem), skb))
Martin KaFai Lau pointed out in the v1 review that the NULL check for
smap doesn't fully address the race condition. The issue is that
diag_get() will re-dereference sdata->smap independently:
net/core/bpf_sk_storage.c:diag_get() {
...
smap = rcu_dereference(sdata->smap);
...
}
Between the NULL check above and the call to diag_get(), smap could
become NULL again, leading to the same race condition.
Martin suggested passing smap as a parameter to diag_get() to avoid
the second dereference. This would ensure diag_get() uses the same
smap pointer that was already checked.
Could the fix be updated to pass smap to diag_get() as Martin
suggested?
Reference: https://lore.kernel.org/bpf/2026420182243.zAiN.martin.lau@linux.dev/
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/24732581685
^ permalink raw reply
* Re: [PATCH] netfilter: xt_HL: add pr_fmt, default case and NULL checks
From: kernel test robot @ 2026-04-21 16:24 UTC (permalink / raw)
To: Marino Dzalto, pablo, fw
Cc: llvm, oe-kbuild-all, netfilter-devel, coreteam, netdev,
linux-kernel, Marino Dzalto
In-Reply-To: <20260403193929.89449-1-marino.dzalto@gmail.com>
Hi Marino,
kernel test robot noticed the following build errors:
[auto build test ERROR on netfilter-nf/main]
[also build test ERROR on linus/master v7.0 next-20260420]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]
url: https://github.com/intel-lab-lkp/linux/commits/Marino-Dzalto/netfilter-xt_HL-add-pr_fmt-default-case-and-NULL-checks/20260420-185652
base: https://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf.git main
patch link: https://lore.kernel.org/r/20260403193929.89449-1-marino.dzalto%40gmail.com
patch subject: [PATCH] netfilter: xt_HL: add pr_fmt, default case and NULL checks
config: s390-allmodconfig (https://download.01.org/0day-ci/archive/20260422/202604220024.iITt6Hv8-lkp@intel.com/config)
compiler: clang version 18.1.8 (https://github.com/llvm/llvm-project 3b5b5c1ec4a3095ab096dd780e84d7ab81f3d7ff)
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20260422/202604220024.iITt6Hv8-lkp@intel.com/reproduce)
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202604220024.iITt6Hv8-lkp@intel.com/
All errors (new ones prefixed by >>):
>> net/netfilter/xt_hl.c:34:6: error: cannot assign to variable 'ttl' with const-qualified type 'const u8' (aka 'const unsigned char')
34 | ttl = ip_hdr(skb)->ttl;
| ~~~ ^
net/netfilter/xt_hl.c:29:11: note: variable 'ttl' declared const here
29 | const u8 ttl;
| ~~~~~~~~~^~~
1 error generated.
vim +34 net/netfilter/xt_hl.c
25
26 static bool ttl_mt(const struct sk_buff *skb, struct xt_action_param *par)
27 {
28 const struct ipt_ttl_info *info = par->matchinfo;
29 const u8 ttl;
30
31 if (!skb)
32 return false;
33
> 34 ttl = ip_hdr(skb)->ttl;
35
36 switch (info->mode) {
37 case IPT_TTL_EQ:
38 return ttl == info->ttl;
39 case IPT_TTL_NE:
40 return ttl != info->ttl;
41 case IPT_TTL_LT:
42 return ttl < info->ttl;
43 case IPT_TTL_GT:
44 return ttl > info->ttl;
45 default:
46 pr_warn("Unknown TTL match mode: %d\n", info->mode);
47 return false;
48 }
49 }
50
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply
* Re: [PATCH bpf v2] bpf: Fix NULL pointer dereference in bpf_sk_storage_clone and diag paths
From: Amery Hung @ 2026-04-21 16:29 UTC (permalink / raw)
To: bot+bpf-ci
Cc: bestswngs, davem, edumazet, kuba, pabeni, horms, martin.lau, ast,
leon.hwang, kees, wangfushuai, menglong8.dong, netdev, bpf, xmei5,
andrii, daniel, eddyz87, yonghong.song, clm, ihor.solodrai
In-Reply-To: <cc500964d83d95307126b05cc7557309b498bafce0c514fa3f1656dd6413ce71@mail.kernel.org>
On Tue, Apr 21, 2026 at 9:24 AM <bot+bpf-ci@kernel.org> wrote:
>
> > diff --git a/net/core/bpf_sk_storage.c b/net/core/bpf_sk_storage.c
> > index 14eb7812bda4..b2cd2b0292cb 100644
> > --- a/net/core/bpf_sk_storage.c
> > +++ b/net/core/bpf_sk_storage.c
>
> [ ... ]
>
> > @@ -596,6 +596,8 @@ static int bpf_sk_storage_diag_put_all(struct sock *sk, struct sk_buff *skb,
> > saved_len = skb->len;
> > hlist_for_each_entry_rcu(selem, &sk_storage->list, snode) {
> > smap = rcu_dereference(SDATA(selem)->smap);
> > + if (!smap)
> > + continue;
> > diag_size += nla_value_size(smap->map.value_size);
> >
> > if (nla_stgs && diag_get(SDATA(selem), skb))
>
> Martin KaFai Lau pointed out in the v1 review that the NULL check for
> smap doesn't fully address the race condition. The issue is that
> diag_get() will re-dereference sdata->smap independently:
>
> net/core/bpf_sk_storage.c:diag_get() {
> ...
> smap = rcu_dereference(sdata->smap);
> ...
> }
>
> Between the NULL check above and the call to diag_get(), smap could
> become NULL again, leading to the same race condition.
>
> Martin suggested passing smap as a parameter to diag_get() to avoid
> the second dereference. This would ensure diag_get() uses the same
> smap pointer that was already checked.
>
> Could the fix be updated to pass smap to diag_get() as Martin
> suggested?
>
> Reference: https://lore.kernel.org/bpf/2026420182243.zAiN.martin.lau@linux.dev/
>
CI is right. Please pass smap to diag_get() as Martin suggested. It
will be the caller's responsibility to do rcu_dereference() and
null-check it. Note that there are more than one caller to diag_get().
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/24732581685
^ permalink raw reply
* Re: [PATCH net v1] net: validate skb->napi_id in RX tracepoints
From: Simon Horman @ 2026-04-21 16:33 UTC (permalink / raw)
To: Kohei Enju
Cc: netdev, linux-trace-kernel, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Steven Rostedt, Masami Hiramatsu,
Mathieu Desnoyers
In-Reply-To: <20260420105427.162816-1-kohei@enjuk.jp>
On Mon, Apr 20, 2026 at 10:54:23AM +0000, Kohei Enju wrote:
> Since commit 2bd82484bb4c ("xps: fix xps for stacked devices"),
> skb->napi_id shares storage with sender_cpu. RX tracepoints using
> net_dev_rx_verbose_template read skb->napi_id directly and can therefore
> report sender_cpu values as if they were NAPI IDs.
>
> For example, on the loopback path this can report 1 as napi_id, where 1
> comes from raw_smp_processor_id() + 1 in the XPS path:
>
> # bpftrace -e 'tracepoint:net:netif_rx_entry{ print(args->napi_id); }'
> # taskset -c 0 ping -c 1 ::1
>
> Report only valid NAPI IDs in these tracepoints and use 0 otherwise.
>
> Fixes: 2bd82484bb4c ("xps: fix xps for stacked devices")
> Signed-off-by: Kohei Enju <kohei@enjuk.jp>
Reviewed-by: Simon Horman <horms@kernel.org>
> ---
> include/trace/events/net.h | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/include/trace/events/net.h b/include/trace/events/net.h
> index fdd9ad474ce3..dbc2c5598e35 100644
> --- a/include/trace/events/net.h
> +++ b/include/trace/events/net.h
> @@ -10,6 +10,7 @@
> #include <linux/if_vlan.h>
> #include <linux/ip.h>
> #include <linux/tracepoint.h>
> +#include <net/busy_poll.h>
>
> TRACE_EVENT(net_dev_start_xmit,
>
> @@ -208,7 +209,8 @@ DECLARE_EVENT_CLASS(net_dev_rx_verbose_template,
> TP_fast_assign(
> __assign_str(name);
> #ifdef CONFIG_NET_RX_BUSY_POLL
> - __entry->napi_id = skb->napi_id;
> + __entry->napi_id = napi_id_valid(skb->napi_id) ?
> + skb->napi_id : 0;
Note to self: they key is that if the storage at napi_id is
being used as a sender_cpu then napi_id_valid because
the valid values for a sender_cpu are disjoint from those
of a valid napi_id. This can be seen clearly in the
implementation of napi_id_valid() and the comment above it.
> #else
> __entry->napi_id = 0;
> #endif
> --
> 2.51.0
>
^ permalink raw reply
* Re: [PATCH net v4 0/5] net: mana: Fix probe/remove error path bugs
From: Simon Horman @ 2026-04-21 16:49 UTC (permalink / raw)
To: Erni Sri Satya Vennela
Cc: kys, haiyangz, wei.liu, decui, longli, andrew+netdev, davem,
edumazet, kuba, pabeni, ssengar, dipayanroy, gargaditya,
shirazsaleem, kees, kotaranov, leon, shacharr, stephen,
linux-hyperv, netdev, linux-kernel
In-Reply-To: <20260420124741.1056179-1-ernis@linux.microsoft.com>
On Mon, Apr 20, 2026 at 05:47:34AM -0700, Erni Sri Satya Vennela wrote:
> Fix five bugs in mana_probe()/mana_remove() error handling that can
> cause warnings on uninitialized work structs, NULL pointer dereferences,
> masked errors, and resource leaks when early probe steps fail.
>
> Patches 1-2 move work struct initialization (link_change_work and
> gf_stats_work) to before any error path that could trigger
> mana_remove(), preventing WARN_ON in __flush_work() or debug object
> warnings when sync cancellation runs on uninitialized work structs.
>
> Patch 3 guards mana_remove() against double invocation. If PM resume
> fails, mana_probe() calls mana_remove() which sets gdma_context and
> driver_data to NULL. A failed resume does not unbind the driver, so
> when the device is eventually unbound, mana_remove() is called again
> and dereferences NULL, causing a kernel panic. An early return on
> NULL gdma_context or driver_data makes the second call harmless.
>
> Patch 4 prevents add_adev() from overwriting a port probe error,
> which could leave the driver in a broken state with NULL ports while
> reporting success.
>
> Patch 5 changes 'goto out' to 'break' in mana_remove()'s port loop
> so that mana_destroy_eq() is always reached, preventing EQ leaks when
> a NULL port is encountered.
> ---
> Changes in v4:
> * Correct Fixes tag from ca9c54d2d6a5 to 635096a86edb
> * Correct Fixes tag from ced82fce77e9 to a69839d4327d
Thanks for the updates.
For the series:
Reviewed-by: Simon Horman <horms@kernel.org>
I see that Sashiko provided feedback on patch 4/5.
However, as it notes, the issue it flags is addressed in patch 5/5.
No further action required AFAICS.
^ permalink raw reply
* Re: [PATCH bpf-next v4 1/4] bpf: Reject TCP_NODELAY in TCP header option callbacks
From: bot+bpf-ci @ 2026-04-21 16:51 UTC (permalink / raw)
To: kafai.wan, ast, daniel, john.fastabend, andrii, martin.lau,
eddyz87, memxor, song, yonghong.song, jolsa, sdf, davem, edumazet,
kuba, pabeni, horms, dsahern, shuah, ihor.solodrai, kafai.wan,
jiayuan.chen, hoyeon.lee, ameryhung, bpf, linux-kernel, netdev,
linux-kselftest
Cc: 2022090917019, dddddd, M202472210, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, clm, ihor.solodrai
In-Reply-To: <20260421155804.135786-2-kafai.wan@linux.dev>
[-- Attachment #1: Type: text/plain, Size: 2641 bytes --]
> diff --git a/net/core/filter.c b/net/core/filter.c
> index 5fa9189eb772..96849f4c1fbc 100644
> --- a/net/core/filter.c
> +++ b/net/core/filter.c
> @@ -5833,6 +5833,12 @@ BPF_CALL_5(bpf_sock_ops_setsockopt, struct bpf_sock_ops_kern *, bpf_sock,
> if (!is_locked_tcp_sock_ops(bpf_sock))
> return -EOPNOTSUPP;
>
> + /* TCP_NODELAY triggers tcp_push_pending_frames() and re-enters these callbacks. */
> + if ((bpf_sock->op == BPF_SOCK_OPS_HDR_OPT_LEN_CB ||
> + bpf_sock->op == BPF_SOCK_OPS_WRITE_HDR_OPT_CB) &&
> + level == SOL_TCP && optname == TCP_NODELAY)
> + return -EOPNOTSUPP;
> +
Does this same recursion vulnerability exist for BPF TCP congestion
control algorithms using BPF_PROG_TYPE_STRUCT_OPS?
If a BPF congestion control algorithm invokes bpf_setsockopt(TCP_NODELAY)
from its cwnd_event callback when handling CA_EVENT_TX_START, could it
trigger unbounded recursion through this path:
BPF_PROG_TYPE_STRUCT_OPS cwnd_event callback (CA_EVENT_TX_START)
-> bpf_setsockopt(TCP_NODELAY)
-> tcp_push_pending_frames()
-> tcp_write_xmit()
-> tcp_transmit_skb()
-> tcp_event_data_sent()
-> cwnd_event()
-> BPF congestion control cwnd_event callback
The author confirmed this issue in the v3 discussion thread and Martin
KaFai Lau suggested creating a new func_proto
(bpf_sk_setsockopt_nodelay_proto) to reject TCP_NODELAY for all struct_ops
programs, but this patch only addresses BPF_SOCK_OPS callbacks.
Reference: https://lore.kernel.org/bpf/20260418092235.98444C19424@smtp.kernel.org/
Does this fix also need to block TCP_CORK when the value is 0?
When bpf_setsockopt() is called with SOL_TCP and TCP_CORK with value 0
(disabling cork), the call path goes through __tcp_sock_set_cork() in
net/ipv4/tcp.c, which unconditionally calls tcp_push_pending_frames(sk)
when cork is disabled:
__tcp_sock_set_cork()
-> tcp_push_pending_frames()
-> tcp_current_mss()
-> tcp_established_options()
-> bpf_skops_hdr_opt_len()
-> BPF_SOCK_OPS_HDR_OPT_LEN_CB
This triggers the same recursion chain described in the commit message,
but the patch only blocks TCP_NODELAY, not TCP_CORK with value 0.
Reference: https://lore.kernel.org/bpf/5fea6b5e9d1a13e72edf8222a8ebe50cbb660b37e62b87b65dab6f6f317943ec@mail.kernel.org/
> return _bpf_setsockopt(bpf_sock->sk, level, optname, optval, optlen);
> }
>
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/24733356810
^ permalink raw reply
* Re: [PATCH net] netconsole: avoid out-of-bounds access on empty string in trim_newline()
From: Breno Leitao @ 2026-04-21 16:55 UTC (permalink / raw)
To: Simon Horman
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Matthew Wood, netdev, linux-kernel, kernel-team,
stable
In-Reply-To: <20260421162219.GF651125@horms.kernel.org>
On Tue, Apr 21, 2026 at 05:22:19PM +0100, Simon Horman wrote:
> On Mon, Apr 20, 2026 at 03:18:36AM -0700, Breno Leitao wrote:
> > trim_newline() unconditionally dereferences s[len - 1] after computing
> > len = strnlen(s, maxlen). When the string is empty, len is 0 and the
> > expression underflows to s[(size_t)-1], reading (and potentially
> > writing) one byte before the buffer.
> >
> > The two callers feed trim_newline() with the result of strscpy() from
> > configfs store callbacks (dev_name_store, userdatum_value_store).
> > configfs guarantees count >= 1 reaches the callback, but the byte
> > itself can be NUL: a userspace write(fd, "\0", 1) leaves the
> > destination empty after strscpy() and triggers the underflow. The OOB
> > write only fires if the adjacent byte happens to be '\n', so this is
> > not a security issue, but the access is undefined behaviour either way.
> >
> > This pattern is commonly flagged by LLM-based code reviewers. While it
> > is not a security fix, the underlying access is undefined behaviour and
> > the change is small and self-contained, so it is a reasonable candidate
> > for the stable trees.
> >
> > Guard the dereference on a non-zero length.
> >
> > Fixes: ae001dc67907 ("net: netconsole: move newline trimming to function")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Breno Leitao <leitao@debian.org>
>
> Reviewed-by: Simon Horman <horms@kernel.org>
>
> Sashiko has provided some feedback on this patch.
> I do not believe that should hold up progress of this patch.
> But I'd appreciate it if you could look over that feedback
> and see if any follow-up is warranted.
Thanks for the review, I've had a quick look, and it is complaining
about problems are not regressions, but some other issues in the code,
which I will need to check more carefully tomorrow.
https://sashiko.dev/#/patchset/20260420-netcons_trim_newline-v1-1-dc35889aeedf%40debian.org
Thanks,
--breno
^ permalink raw reply
* Re: [syzbot] [kvm?] [net?] [virt?] BUG: sleeping function called from invalid context in vhost_get_avail_idx
From: Kohei Enju @ 2026-04-21 17:11 UTC (permalink / raw)
To: syzbot; +Cc: jasowang, linux-kernel, mst, netdev, syzkaller-bugs
In-Reply-To: <69e6a414.050a0220.24bfd3.002d.GAE@google.com>
On 04/20 15:09, syzbot wrote:
> Hello,
>
> syzbot found the following issue on:
>
> HEAD commit: 8541d8f725c6 Merge tag 'mtd/for-7.1' of git://git.kernel.o..
> git tree: upstream
> console output: https://syzkaller.appspot.com/x/log.txt?x=136454ce580000
> kernel config: https://syzkaller.appspot.com/x/.config?x=7e54da1916e8d11f
> dashboard link: https://syzkaller.appspot.com/bug?extid=6985cb8e543ea90ba8ee
> compiler: gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44
> syz repro: https://syzkaller.appspot.com/x/repro.syz?x=15d264ce580000
> C reproducer: https://syzkaller.appspot.com/x/repro.c?x=143ec1ba580000
>
> Downloadable assets:
> disk image (non-bootable): https://storage.googleapis.com/syzbot-assets/d900f083ada3/non_bootable_disk-8541d8f7.raw.xz
> vmlinux: https://storage.googleapis.com/syzbot-assets/22dfea2c37c2/vmlinux-8541d8f7.xz
> kernel image: https://storage.googleapis.com/syzbot-assets/e2f93ad68fe3/bzImage-8541d8f7.xz
>
> IMPORTANT: if you fix the issue, please add the following tag to the commit:
> Reported-by: syzbot+6985cb8e543ea90ba8ee@syzkaller.appspotmail.com
>
> BUG: sleeping function called from invalid context at drivers/vhost/vhost.c:1527
> in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 6110, name: vhost-6109
> preempt_count: 1, expected: 0
> RCU nest depth: 0, expected: 0
> 2 locks held by vhost-6109/6110:
> #0: ffff888055624cb0 (&vq->mutex/1){+.+.}-{4:4}, at: handle_tx+0x2d/0x160 drivers/vhost/net.c:971
> #1: ffff888055620248 (&vq->mutex){+.+.}-{4:4}, at: vhost_net_busy_poll+0x9c/0x730 drivers/vhost/net.c:554
> Preemption disabled at:
> [<ffffffff88f1a006>] vhost_net_busy_poll+0x1c6/0x730 drivers/vhost/net.c:563
I think the blamed commit may be commit 030881372460 ("vhost_net: basic
polling support"), since it introduced preempt_{disable,enable}() around
the busy-poll loop, which calls a sleepable function inside the loop.
Also, from the changelog of the series,
https://lore.kernel.org/netdev/1448435489-5949-4-git-send-email-jasowang@redhat.com/T/#u
Changes from RFC V1:
...
- Disable preemption during busy looping to make sure local_clock() was
correctly used.
So my understanding is that preempt_disable() was introduced to keep
local_clock() based timeout accounting on a single CPU, rather than as a
requirement of busy polling itself.
If my understanding is correct, migrate_disable() is sufficient here
instead of preempt_disable(), avoiding sleepable accesses from a
preempt-disabled context.
#syz test
diff --git a/drivers/vhost/net.c b/drivers/vhost/net.c
index 80965181920c..c6536cad9c4f 100644
--- a/drivers/vhost/net.c
+++ b/drivers/vhost/net.c
@@ -560,7 +560,7 @@ static void vhost_net_busy_poll(struct vhost_net *net,
busyloop_timeout = poll_rx ? rvq->busyloop_timeout:
tvq->busyloop_timeout;
- preempt_disable();
+ migrate_disable();
endtime = busy_clock() + busyloop_timeout;
while (vhost_can_busy_poll(endtime)) {
@@ -577,7 +577,7 @@ static void vhost_net_busy_poll(struct vhost_net *net,
cpu_relax();
}
- preempt_enable();
+ migrate_enable();
if (poll_rx || sock_has_rx_data(sock))
vhost_net_busy_poll_try_queue(net, vq);
> CPU: 0 UID: 0 PID: 6110 Comm: vhost-6109 Not tainted syzkaller #0 PREEMPT(full)
> Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
> Call Trace:
> <TASK>
> __dump_stack lib/dump_stack.c:94 [inline]
> dump_stack_lvl+0x100/0x190 lib/dump_stack.c:120
> __might_resched.cold+0x1ec/0x232 kernel/sched/core.c:9162
> __might_fault+0x8b/0x140 mm/memory.c:7322
> vhost_get_avail_idx+0x31c/0x4f0 drivers/vhost/vhost.c:1527
> vhost_vq_avail_empty drivers/vhost/vhost.c:3206 [inline]
> vhost_vq_avail_empty+0xa9/0xe0 drivers/vhost/vhost.c:3199
> vhost_net_busy_poll+0x297/0x730 drivers/vhost/net.c:574
> vhost_net_tx_get_vq_desc drivers/vhost/net.c:610 [inline]
> get_tx_bufs.constprop.0+0x338/0x600 drivers/vhost/net.c:650
> handle_tx_copy+0x28c/0x12e0 drivers/vhost/net.c:778
> handle_tx+0x139/0x160 drivers/vhost/net.c:985
> vhost_run_work_list+0x183/0x220 drivers/vhost/vhost.c:454
> vhost_task_fn+0x156/0x430 kernel/vhost_task.c:49
> ret_from_fork+0x72b/0xd50 arch/x86/kernel/process.c:158
> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
> </TASK>
>
>
> ---
> This report is generated by a bot. It may contain errors.
> See https://goo.gl/tpsmEJ for more information about syzbot.
> syzbot engineers can be reached at syzkaller@googlegroups.com.
>
> syzbot will keep track of this issue. See:
> https://goo.gl/tpsmEJ#status for how to communicate with syzbot.
>
> If the report is already addressed, let syzbot know by replying with:
> #syz fix: exact-commit-title
>
> If you want syzbot to run the reproducer, reply with:
> #syz test: git://repo/address.git branch-or-commit-hash
> If you attach or paste a git patch, syzbot will apply it before testing.
>
> If you want to overwrite report's subsystems, reply with:
> #syz set subsystems: new-subsystem
> (See the list of subsystem names on the web dashboard)
>
> If the report is a duplicate of another one, reply with:
> #syz dup: exact-subject-of-another-report
>
> If you want to undo deduplication, reply with:
> #syz undup
^ permalink raw reply related
* Re: [PATCH net-deletions] net: remove ISDN subsystem and Bluetooth CMTP
From: Randy Dunlap @ 2026-04-21 17:12 UTC (permalink / raw)
To: Luiz Augusto von Dentz, Jakub Kicinski
Cc: davem, netdev, edumazet, pabeni, andrew+netdev, horms, corbet,
skhan, marcel, mchehab+huawei, jani.nikula, gregkh, demarchi,
justonli, ivecera, jonathan.cameron, kees, marco.crivellari,
ferr.lambarginio, nihaal, mingo, tglx, linmq006, linux-doc,
linux-bluetooth
In-Reply-To: <CABBYNZ+yCH2hxbS32o6eDT7BDMLZd3YpjUZ=sfiw=z9XjMT6OQ@mail.gmail.com>
On 4/21/26 6:55 AM, Luiz Augusto von Dentz wrote:
> Hi Jakub,
>
> On Mon, Apr 20, 2026 at 10:21 PM Jakub Kicinski <kuba@kernel.org> wrote:
>>
>> Remove the ISDN (mISDN, CAPI) subsystem and Bluetooth CMTP protocol
>> from the kernel tree.
>>
>>
>
> Acked-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
Please don't send 1.7 MB emails for an Ack.
See https://people.kernel.org/tglx/,
especially "Trim replies".
--
~Randy
^ permalink raw reply
* Re: [PATCH net-deletions] net: remove ax25 and amateur radio (hamradio) subsystem
From: Stephen Hemminger @ 2026-04-21 17:14 UTC (permalink / raw)
To: Dan Cross
Cc: Jakub Kicinski, davem, netdev, edumazet, pabeni, andrew+netdev,
horms, corbet, skhan, federico.vaga, carlos.bilbao, avadhut.naik,
alexs, si.yanteng, dzm91, 2023002089, tsbogend, dsahern,
jani.nikula, mchehab+huawei, gregkh, jirislaby, tytso, herbert,
ebiggers, johannes.berg, geert, pablo, tglx, mashiro.chen, mingo,
dqfext, jreuter, sdf, pkshih, enelsonmoore, mkl, toke, kees,
jlayton, wangliang74, aha310510, takamitz, kuniyu, linux-doc,
linux-mips
In-Reply-To: <CAEoi9W6ZRw6aEh62Xbgkg-TW8URHbVp6dHTT9krFiTkotjTuTA@mail.gmail.com>
On Tue, 21 Apr 2026 12:17:23 -0400
Dan Cross <crossd@gmail.com> wrote:
> On Tue, Apr 21, 2026 at 9:55 AM Stephen Hemminger
> <stephen@networkplumber.org> wrote:
> > On Mon, 20 Apr 2026 19:18:23 -0700
> > Jakub Kicinski <kuba@kernel.org> wrote:
> > > Remove the amateur radio (AX.25, NET/ROM, ROSE) protocol implementation
> > > and all associated hamradio device drivers from the kernel tree.
> > > This set of protocols has long been a huge bug/syzbot magnet,
> > > and since nobody stepped up to help us deal with the influx
> > > of the AI-generated bug reports we need to move it out of tree
> > > to protect our sanity.
> > >
> > > The code is moved to an out-of-tree repo:
> > > https://github.com/linux-netdev/mod-orphan
> > > if it's cleaned up and reworked there we can accept it back.
> >
> > It would be good if these protocols could be done in userspace
> > or with BPF?
>
> Consensus for a userspace implementation is what folks on linux-hams
> seem to be converging on.
>
> The amateur radio protocols are more or less specific to low-speed
> links, they are not particularly coupled to anything else that
> requires running in the kernel, and the main coupling point (IP over
> AX.25) can be implemented via TAP/TUN.
>
> There are several popular packages that already implement AX.25 and
> NET/ROM in user-space (for the interested, LinBPQ seems to be the
> canonical example). The main missing piece is ROSE, but it is likely
> easier to add that to an existing package, or potentially something
> brand new, than keep it in the kernel.
>
> There's no compelling reason to keep these protocols in the kernel,
> whether in-tree or out-of-tree; at least, one has not been
> articulated.
>
> - Dan C.
Thanks, my other concern is carrying support for these in ip commands.
If not kernel based, then iproute2 doesn't need to worry.
^ permalink raw reply
* Re: [PATCH net v4 2/2] net: airoha: Add missing bits in airoha_qdma_cleanup_tx_queue()
From: Lorenzo Bianconi @ 2026-04-21 17:15 UTC (permalink / raw)
To: Paolo Abeni
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Simon Horman, linux-arm-kernel, linux-mediatek, netdev
In-Reply-To: <210f5d0b-6232-4c0f-adff-3a97d54159b3@redhat.com>
[-- Attachment #1: Type: text/plain, Size: 1904 bytes --]
> On 4/17/26 8:36 AM, Lorenzo Bianconi wrote:
> > @@ -1055,8 +1058,33 @@ static void airoha_qdma_cleanup_tx_queue(struct airoha_queue *q)
> > e->dma_addr = 0;
> > e->skb = NULL;
> > list_add_tail(&e->list, &q->tx_list);
> > +
> > + /* Reset DMA descriptor */
> > + WRITE_ONCE(desc->ctrl, 0);
> > + WRITE_ONCE(desc->addr, 0);
> > + WRITE_ONCE(desc->data, 0);
> > + WRITE_ONCE(desc->msg0, 0);
> > + WRITE_ONCE(desc->msg1, 0);
> > + WRITE_ONCE(desc->msg2, 0);
>
> Sashiko has some complains on this patch that look legit to me.
>
> Also the pre-existing issues mentioned WRT patch 1/2 makes such patch
> IMHO almost ineffective, I think you should address them in the same series.
>
> Note that you should have commented on sashiko review on the ML, it
> would have saved a significant amount of time on us.
Since this series is marked as 'Changes Requested', it is not clear to me what
next steps are. I guess we have two possible approach here:
1) - Post patch 1/2 ("net: airoha: Move ndesc initialization at
end of airoha_qdma_init_tx()") with the series available upstream
(not merged yet) in [0] where I am fixing similar issues for
airoha_qdma_init_rx_queue() and airoha_qdma_tx_irq_init().
- Post patch 2/2 ("net: airoha: Add missing bits in
airoha_qdma_cleanup_tx_queue()") with a fix for airoha_ndo_stop() waiting
for TX/RX DMA engine to complete before running
airoha_qdma_cleanup_tx_queue().
2) - Since all the issues rised by Sashiko are not strictly related to this
series and they are already fixed in pending patches, just apply the fixes
separately without the needs to repost this series.
Which approach do you prefer?
Regards,
Lorenzo
[0] https://patchwork.kernel.org/project/netdevbpf/cover/20260420-airoha_qdma_init_rx_queue-fix-v2-0-d99347e5c18d@kernel.org/
>
> /P
>
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 228 bytes --]
^ permalink raw reply
* Re: [PATCH 0/3] mptcp: add RECVERR and MSG_ERRQUEUE support
From: David CARLIER @ 2026-04-21 17:16 UTC (permalink / raw)
To: Matthieu Baerts
Cc: netdev, mptcp, martineau, geliang, davem, edumazet, kuba, pabeni,
horms
In-Reply-To: <e190409c-3dee-4de8-b1a0-7867213cb801@kernel.org>
Thanks for the head-up, cheers !
On Tue, 21 Apr 2026 at 17:07, Matthieu Baerts <matttbe@kernel.org> wrote:
>
> Hi David,
>
> On 21/04/2026 17:22, David Carlier wrote:
> > MPTCP already advertises IP_RECVERR/IPV6_RECVERR as supported, but the
> > parent socket does not currently provide usable MSG_ERRQUEUE handling.
> >
> > This series wires the MPTCP socket up to the IPv4/IPv6 error queue
> > paths. It propagates RECVERR-related sockopts to existing and future
> > subflows, makes poll() report pending errqueue activity through the
> > parent socket, and allows recvmsg(MSG_ERRQUEUE) on the MPTCP socket to
> > consume queued errors with the parent socket ABI.
> >
> > The series also handles mixed-family subflows by applying the matching
> > sockopt according to each subflow family, and avoids silently losing an
> > error skb if requeueing to the parent socket fails under rmem pressure.
> Thank you for this series!
>
> Even if I agree it would be good to have full MSG_ERRQUEUE support,
> net-next is currently closed, and only bug fixes are accepted, see:
>
> https://docs.kernel.org/process/maintainer-netdev.html
>
> pw-bot: defer
>
> Instead, I suggest switching the discussions only to the MPTCP ML if
> that's OK. If the CI is happy, someone will try to review it over there,
> when time permits. If not, please send the new versions only to the
> MPTCP ML, with the 'PATCH mptcp-next' prefix, and ideally on top of the
> 'export' (or 'for-review') branch of our tree. For more details:
>
> https://www.mptcp.dev/contributing.html#kernel-development
>
> Cheers,
> Matt
> --
> Sponsored by the NGI0 Core fund.
>
^ permalink raw reply
* Re: [PATCH 18/23] cpu/hotplug: Add a new cpuhp_offline_cb() API
From: Waiman Long @ 2026-04-21 17:29 UTC (permalink / raw)
To: Thomas Gleixner, Tejun Heo, Johannes Weiner, Michal Koutný,
Jonathan Corbet, Shuah Khan, Catalin Marinas, Will Deacon,
K. Y. Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
Guenter Roeck, Frederic Weisbecker, Paul E. McKenney,
Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
Uladzislau Rezki, Steven Rostedt, Mathieu Desnoyers,
Lai Jiangshan, Zqiang, Anna-Maria Behnsen, Ingo Molnar,
Chen Ridong, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman
Cc: cgroups, linux-doc, linux-kernel, linux-arm-kernel, linux-hyperv,
linux-hwmon, rcu, netdev, linux-kselftest, Costa Shulyupin,
Qiliang Yuan
In-Reply-To: <87o6jcb84w.ffs@tglx>
On 4/21/26 12:17 PM, Thomas Gleixner wrote:
> On Mon, Apr 20 2026 at 23:03, Waiman Long wrote:
>> Add a new cpuhp_offline_cb() API that allows us to offline a set of
>> CPUs one-by-one, run the given callback function and then bring those
>> CPUs back online again while inhibiting any concurrent CPU hotplug
>> operations from happening.
> Please provide a properly structured change log which explains the
> context, the problem and the solution in separate paragraphs and this
> order. This is not new. It's documented...
>
>> This new API can be used to enable runtime adjustment of nohz_full and
>> isolcpus boot command line options. A new cpuhp_offline_cb_mode flag
>> is also added to signal that the system is in this offline callback
>> transient state so that some hotplug operations can be optimized out
>> if we choose to.
> We chose nothing.
>
>> +#include <linux/cpumask_types.h>
> What for? This header only needs a 'struct cpumask' forward declaration
> so that the compiler can handle the pointer argument, no?
>
>> +typedef int (*cpuhp_cb_t)(void *arg);
> You couldn't come up with a more generic name for this, right?
>
>> struct device;
>>
>> extern int lockdep_is_cpus_held(void);
>> @@ -29,6 +31,8 @@ void clear_tasks_mm_cpumask(int cpu);
>> int remove_cpu(unsigned int cpu);
>> int cpu_device_down(struct device *dev);
>> void smp_shutdown_nonboot_cpus(unsigned int primary_cpu);
>> +int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg);
> Ditto.
>
>> +extern bool cpuhp_offline_cb_mode;
> Groan. The only users are in the cpusets code which invokes this muck
> and should therefore know what's going on, no?
>
>> #else /* CONFIG_HOTPLUG_CPU */
>>
>> @@ -43,6 +47,11 @@ static inline void cpu_hotplug_disable(void) { }
>> static inline void cpu_hotplug_enable(void) { }
>> static inline int remove_cpu(unsigned int cpu) { return -EPERM; }
>> static inline void smp_shutdown_nonboot_cpus(unsigned int primary_cpu) { }
>> +static inline int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg)
>> +{
>> + return -EPERM;
> -EPERM?
>
>> +/**
>> + * cpuhp_offline_cb - offline CPUs, invoke callback function & online CPUs afterward
>> + * @mask: A mask of CPUs to be taken offline and then online
>> + * @func: A callback function to be invoked while the given CPUs are offline
>> + * @arg: Argument to be passed back to the callback function
>> + *
>> + * Return: 0 if successful, an error code otherwise
>> + */
>> +int cpuhp_offline_cb(struct cpumask *mask, cpuhp_cb_t func, void *arg)
>> +{
>> + int off_cpu, on_cpu, ret, ret2 = 0;
>> +
>> + if (WARN_ON_ONCE(cpumask_empty(mask) ||
>> + !cpumask_subset(mask, cpu_online_mask)))
>> + return -EINVAL;
> No line break required. You have 100 characters.
>
> But what's worse is that the access to cpu_online_mask is not protected
> against a concurrent CPU hotplug operation.
>
>> +
>> + pr_debug("%s: begin (CPU list = %*pbl)\n", __func__, cpumask_pr_args(mask));
> Tracing?
>
>> + lock_device_hotplug();
>> + cpuhp_offline_cb_mode = true;
>> + /*
>> + * If all offline operations succeed, off_cpu should become nr_cpu_ids.
>> + */
>> + for_each_cpu(off_cpu, mask) {
>> + ret = device_offline(get_cpu_device(off_cpu));
>> + if (unlikely(ret))
>> + break;
>> + }
>> + if (!ret)
>> + ret = func(arg);
>> +
>> + /* Bring previously offline CPUs back online */
>> + for_each_cpu(on_cpu, mask) {
>> + int retries = 0;
>> +
>> + if (on_cpu == off_cpu)
>> + break;
>> +
>> +retry:
>> + ret2 = device_online(get_cpu_device(on_cpu));
>> +
>> + /*
>> + * With the unlikely event that CPU hotplug is disabled while
>> + * this operation is in progress, we will need to wait a bit
>> + * for hotplug to hopefully be re-enabled again. If not, print
>> + * a warning and return the error.
>> + *
>> + * cpu_hotplug_disabled is supposed to be accessed while
>> + * holding the cpu_add_remove_lock mutex. So we need to
>> + * use the data_race() macro to access it here.
>> + */
>> + while ((ret2 == -EBUSY) && data_race(cpu_hotplug_disabled) &&
>> + (++retries <= 5)) {
>> + msleep(20);
>> + if (!data_race(cpu_hotplug_disabled))
>> + goto retry;
>> + }
>> + if (ret2) {
>> + pr_warn("%s: Failed to bring CPU %d back online!\n",
>> + __func__, on_cpu);
> Provide a proper text and not this silly __func__ thing.
>
>> + break;
>> + }
>> + }
> TBH. This is unreviewable gunk and the whole 'unlikely event that CPU
> hotplug is disabled' is just a lazy hack.
>
> All of this can be avoided including this made up callback function.
>
> It's not rocket science to provide:
>
> 1) A function which serializes against any other CPU hotplug
> related action.
>
> 2) A function which brings the CPUs in a given CPU mask down
>
> 3) A function which brings the CPUs in a given CPU mask up
>
> 4) A function which undoes #1
>
> Yeah I know, it's more work and not convoluted enough. But see below.
>
> That brings me to that other hack namely cpuhp_offline_cb_mode, which
> you self described as such in patch 21/23:
>
>> + /*
>> + * Hack: In cpuhp_offline_cb_mode, pretend all partitions are empty
>> + * to prevent unnecessary partition invalidation.
>> + */
>> + if (cpuhp_offline_cb_mode)
>> + return false;
>> +
> We are not merging hacks. End of story. But you knew that already, no?
>
> Let's take a step back and see what you really need to achieve:
>
> 1) Update tick_nohz_full_mask
> 2) Update the managed interrupt mask
> 3) Update CPU sets
>
> Independent of the direction of this update you need to ensure that the
> affected functionality keeps working correctly.
>
> You achieve that by bulk offlining the affected CPUs, invoking a magic
> callback and then bulk onlining the affected CPUs again, which requires
> that ill defined cpuhp_offline_cb_mode hackery and probably some more
> hacks all over the place.
>
> You can achieve the same by doing CPU by CPU operations in the right
> order without this mode hack, when you establish proper limitations for
> this:
>
> At no point in time it's allowed to empty a CPU set or a affected CPU
> mask, except when you completely undo the isolation of CPUs.
>
> That can be computed upfront w/o changing anything at all. Once the
> validity is established, the update can proceed. Or you can leave it
> to user space which can keep the pieces if it gets it wrong.
>
> That's a reasonable limitation as there is absolutely zero justification
> to support something like:
>
> housekeeping_cpus = [CPU 0], isolated_cpus = [CPU 1]
> ---> housekeeping_cpus = [CPU 1], isolated_cpus = [CPU 0]
>
> just because we can with enough horrible hacks.
>
> If you get that out of the way, then a CPU by CPU update becomes the
> obvious and simplest solution. The ordering constraints can be computed
> in user space upfront and there is no reason to do any of this in the
> kernel itself except for an eventual validation step. It might be a tad
> slower, but this is all but a hotpath operation.
>
> Just for the record. I suggested exactly this more than a year ago and
> it's still the right thing to do.
>
> And of course neither your cover letter nor any of the patches give a
> proper rationale why you think that your bulk hackery is better. For the
> very simple reason that there is no rationale at all.
>
> This bulk muck is doomed when your ultimate goal is to avoid the stop
> machine dance. With a per CPU update it is actually doable without more
> ill defined hacks all over the place.
>
> 1) Bring down the CPU to CPUHP_AP_SCHED_WAIT_EMPTY, which is the last
> state before stop machine is invoked.
>
> At that point:
>
> - no user space thread is running on the CPU anymore
>
> - everything related to this CPU has been shut down or moved
> elsewhere
>
> - interrupt managed device queues are quiesced if the CPU was
> the last online one in the queue affinity mask. If not the
> interrupt might still be affine to the CPU, but there is at
> least one other CPU available in the mask.
>
> 2) Update the tick NOHZ handover
>
> This can be done without going into stop machine by providing a
> hotplug callback right between CPUHP_AP_SMPBOOT_THREADS and
> CPUHP_AP_IRQ_AFFINITY_ONLINE.
>
> That's trivial enough to achieve and can work independently of
> NOHZ full.
>
> 3) Rework the affinity management, so that interrupt affinities can
> be reassigned in the CPUHP_AP_IRQ_AFFINITY_ONLINE state.
>
> That needs a lot of thoughts, but there is no real reason why it
> can't work.
>
> 4) Flip the housekeeping CPU masks in sched_cpu_wait_empty() after
> balance_hotplug_wait().
>
> 5) Bring the CPU online again.
>
> For #2 and #3 to work you need a separate CPU mask which avoids touching
> CPU online mask. For #3 this needs some more work to avoid reassigning the
> interrupts once sparse_irq_lock is dropped, but the bulk is achieved
> with the separate CPU mask.
>
> No?
Thanks for the great suggestions. I will certainly look into that.
We actually have a cpu_active_mask that will be cleared early in
sched_cpu_deactivate(). In the CPUHP_AP_SCHED_WAIT_EMPTY state, the CPU
will still have online bit set but the active bit will be cleared. Or we
could add another cpumask that can be used to indicate CPUs that have
reached CPUHP_AP_SCHED_WAIT_EMPTY or below if necessary.
Cheers,
Longman
^ permalink raw reply
* Re: [PATCH net v4 2/2] net: airoha: Add missing bits in airoha_qdma_cleanup_tx_queue()
From: Paolo Abeni @ 2026-04-21 17:32 UTC (permalink / raw)
To: Lorenzo Bianconi
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Simon Horman, linux-arm-kernel, linux-mediatek, netdev
In-Reply-To: <aeewoVjN7EHLZTW_@lore-desk>
On 4/21/26 7:15 PM, Lorenzo Bianconi wrote:
>> On 4/17/26 8:36 AM, Lorenzo Bianconi wrote:
>>> @@ -1055,8 +1058,33 @@ static void airoha_qdma_cleanup_tx_queue(struct airoha_queue *q)
>>> e->dma_addr = 0;
>>> e->skb = NULL;
>>> list_add_tail(&e->list, &q->tx_list);
>>> +
>>> + /* Reset DMA descriptor */
>>> + WRITE_ONCE(desc->ctrl, 0);
>>> + WRITE_ONCE(desc->addr, 0);
>>> + WRITE_ONCE(desc->data, 0);
>>> + WRITE_ONCE(desc->msg0, 0);
>>> + WRITE_ONCE(desc->msg1, 0);
>>> + WRITE_ONCE(desc->msg2, 0);
>>
>> Sashiko has some complains on this patch that look legit to me.
>>
>> Also the pre-existing issues mentioned WRT patch 1/2 makes such patch
>> IMHO almost ineffective, I think you should address them in the same series.
>>
>> Note that you should have commented on sashiko review on the ML, it
>> would have saved a significant amount of time on us.
>
> Since this series is marked as 'Changes Requested', it is not clear to me what
> next steps are. I guess we have two possible approach here:
>
> 1) - Post patch 1/2 ("net: airoha: Move ndesc initialization at
> end of airoha_qdma_init_tx()") with the series available upstream
> (not merged yet) in [0] where I am fixing similar issues for
> airoha_qdma_init_rx_queue() and airoha_qdma_tx_irq_init().
> - Post patch 2/2 ("net: airoha: Add missing bits in
> airoha_qdma_cleanup_tx_queue()") with a fix for airoha_ndo_stop() waiting
> for TX/RX DMA engine to complete before running
> airoha_qdma_cleanup_tx_queue().
>
> 2) - Since all the issues rised by Sashiko are not strictly related to this
> series and they are already fixed in pending patches, just apply the fixes
> separately without the needs to repost this series.
Given the current flood on the ML I think option 2 could be the better.
Note that the tree is currently in Jakub's hands and he can very legitly
disagree.
/P
^ permalink raw reply
* Re: [PATCH net] net: ipv6: fix NOREF dst use in seg6 and rpl lwtunnels
From: Justin Iurman @ 2026-04-21 17:33 UTC (permalink / raw)
To: Andrea Mayer, davem, dsahern, edumazet, kuba, pabeni, horms
Cc: bigeasy, clrkwllms, rostedt, david.lebrun, alex.aring,
stefano.salsano, netdev, linux-rt-devel, linux-kernel, stable
In-Reply-To: <20260421094735.20997-1-andrea.mayer@uniroma2.it>
On 4/21/26 11:47, Andrea Mayer wrote:
> seg6_input_core() and rpl_input() call ip6_route_input() which sets a
> NOREF dst on the skb, then pass it to dst_cache_set_ip6() invoking
> dst_hold() unconditionally.
> On PREEMPT_RT, ksoftirqd is preemptible and a higher-priority task can
> release the underlying pcpu_rt between the lookup and the caching
> through a concurrent FIB lookup on a shared nexthop.
> Simplified race sequence:
>
> ksoftirqd/X higher-prio task (same CPU X)
> ----------- --------------------------------
> seg6_input_core(,skb)/rpl_input(skb)
> dst_cache_get()
> -> miss
> ip6_route_input(skb)
> -> ip6_pol_route(,skb,flags)
> [RT6_LOOKUP_F_DST_NOREF in flags]
> -> FIB lookup resolves fib6_nh
> [nhid=N route]
> -> rt6_make_pcpu_route()
> [creates pcpu_rt, refcount=1]
> pcpu_rt->sernum = fib6_sernum
> [fib6_sernum=W]
> -> cmpxchg(fib6_nh.rt6i_pcpu,
> NULL, pcpu_rt)
> [slot was empty, store succeeds]
> -> skb_dst_set_noref(skb, dst)
> [dst is pcpu_rt, refcount still 1]
>
> rt_genid_bump_ipv6()
> -> bumps fib6_sernum
> [fib6_sernum from W to Z]
> ip6_route_output()
> -> ip6_pol_route()
> -> FIB lookup resolves fib6_nh
> [nhid=N]
> -> rt6_get_pcpu_route()
> pcpu_rt->sernum != fib6_sernum
> [W <> Z, stale]
> -> prev = xchg(rt6i_pcpu, NULL)
> -> dst_release(prev)
> [prev is pcpu_rt,
> refcount 1->0, dead]
>
> dst = skb_dst(skb)
> [dst is the dead pcpu_rt]
> dst_cache_set_ip6(dst)
> -> dst_hold() on dead dst
> -> WARN / use-after-free
>
> For the race to occur, ksoftirqd must be preemptible (PREEMPT_RT without
> PREEMPT_RT_NEEDS_BH_LOCK) and a concurrent task must be able to release
> the pcpu_rt. Shared nexthop objects provide such a path, as two routes
> pointing to the same nhid share the same fib6_nh and its rt6i_pcpu
> entry.
>
> Fix seg6_input_core() and rpl_input() by calling skb_dst_force() after
> ip6_route_input() to force the NOREF dst into a refcounted one before
> caching.
> The output path is not affected as ip6_route_output() already returns a
> refcounted dst.
>
> Fixes: af4a2209b134 ("ipv6: sr: use dst_cache in seg6_input")
> Fixes: a7a29f9c361f ("net: ipv6: add rpl sr tunnel")
> Cc: stable@vger.kernel.org
> Signed-off-by: Andrea Mayer <andrea.mayer@uniroma2.it>
> ---
> net/ipv6/rpl_iptunnel.c | 9 +++++++++
> net/ipv6/seg6_iptunnel.c | 9 +++++++++
> 2 files changed, 18 insertions(+)
>
> diff --git a/net/ipv6/rpl_iptunnel.c b/net/ipv6/rpl_iptunnel.c
> index c7942cf65567..4e10adcd70e8 100644
> --- a/net/ipv6/rpl_iptunnel.c
> +++ b/net/ipv6/rpl_iptunnel.c
> @@ -287,7 +287,16 @@ static int rpl_input(struct sk_buff *skb)
>
> if (!dst) {
> ip6_route_input(skb);
> +
> + /* ip6_route_input() sets a NOREF dst; force a refcount on it
> + * before caching or further use.
> + */
> + skb_dst_force(skb);
> dst = skb_dst(skb);
> + if (unlikely(!dst)) {
> + err = -ENETUNREACH;
> + goto drop;
> + }
>
> /* cache only if we don't create a dst reference loop */
> if (!dst->error && lwtst != dst->lwtstate) {
> diff --git a/net/ipv6/seg6_iptunnel.c b/net/ipv6/seg6_iptunnel.c
> index 97b50d9b1365..94284b483be0 100644
> --- a/net/ipv6/seg6_iptunnel.c
> +++ b/net/ipv6/seg6_iptunnel.c
> @@ -515,7 +515,16 @@ static int seg6_input_core(struct net *net, struct sock *sk,
>
> if (!dst) {
> ip6_route_input(skb);
> +
> + /* ip6_route_input() sets a NOREF dst; force a refcount on it
> + * before caching or further use.
> + */
> + skb_dst_force(skb);
> dst = skb_dst(skb);
> + if (unlikely(!dst)) {
> + err = -ENETUNREACH;
> + goto drop;
> + }
>
> /* cache only if we don't create a dst reference loop */
> if (!dst->error && lwtst != dst->lwtstate) {
Thanks for taking care of this, Andrea! LGTM.
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
^ permalink raw reply
* Re: [PATCH net 1/2] net/mlx5e: psp: Fix invalid access on PSP dev registration fail
From: Cosmin Ratiu @ 2026-04-21 17:34 UTC (permalink / raw)
To: kuba@kernel.org
Cc: Boris Pismenny, willemdebruijn.kernel@gmail.com,
andrew+netdev@lunn.ch, daniel.zahka@gmail.com,
davem@davemloft.net, leon@kernel.org,
linux-kernel@vger.kernel.org, edumazet@google.com,
linux-rdma@vger.kernel.org, Rahul Rameshbabu, Raed Salem,
Dragos Tatulea, kees@kernel.org, Mark Bloch, pabeni@redhat.com,
Tariq Toukan, Saeed Mahameed, netdev@vger.kernel.org,
Gal Pressman
In-Reply-To: <20260421080951.570e6e49@kernel.org>
On Tue, 2026-04-21 at 08:09 -0700, Jakub Kicinski wrote:
> On Tue, 21 Apr 2026 14:33:51 +0000 Cosmin Ratiu wrote:
> > > > priv->psp and steering at the time of mlx5e_psp_register() is
> > > > inert
> > > > without the PSP device. Cleaning it on psp_dev_create() failure
> > > > would
> > > > be weird, it's cleaned up anyway on netdev teardown. The fact
> > > > that
> > > > only
> > > > memory allocations can fail inside psp_dev_create() is
> > > > irrelevant
> > > > here.
> > > > psp_dev_create() failing shouldn't bring down the whole
> > > > netdevice,
> > > > so
> > > > logging a message and continuing is ok (which is what is also
> > > > done
> > > > for
> > > > macsec and ktls).
> > >
> > > This is a misguided cargo cult. Or something motivated by OOT
> > > compatibility. Alex D sometimes tries to do the same thing with
> > > Meta
> > > drivers. I don't get it. Of course we want the device to be
> > > operational
> > > if some *device* init fails. The compatibility matrix with all
> > > device
> > > generations and fw versions could justify that. But continuing
> > > init
> > > when a single-page kmalloc failed is pure silliness.
> >
> > I am not sure about the wider context, but from the POV of the
> > driver,
> > it's calling $thing from the kernel which can fail and it needs to
> > do
> > something about it, either fail the entire netdev bringup or accept
> > that $thing won't be functional and continue without it. The driver
> > shouldn't need to know what $thing does inside and how it can fail,
> > which can change over time. Today it's a kmalloc(), tomorrow it may
> > be
> > something else.
>
> Like what?
The inner workings of $thing aren't and shouldn't be relevant, no?
Maybe tomorrow the kernel will lazy-init some TCP shenanigans for the
first PSP device being initialized or whatever, or maybe some other
moving parts inside can fail. It's an abstraction, why make it
unnecessarily leaky for the purpose of writing driver code?
>
> > It doesn't and shouldn't matter for the local decision
> > to continue or not without $thing working.
> >
> > Isn't this reasonable?
>
> No, the normal thing to do is to propagate errors.
> If you want to diverge from that _you_ should have a reason,
> a better reason than a vague "kernel can fail".
> I'd prefer for the driver to fail in an obvious way.
> Which will be immediately spotted by the operator, not 2 weeks
> later when 10% of the fleet is upgraded already.
> The only exception I'd make is to keep devlink registered in
> case the fix is to flash a different FW.
In this case, PSP not working would be spotted on the next PSP dev-get
op which produces zilch instead of working devices.
But I understand what you want. You'd like the netdevice to either be
fully initialized with all supported+configured protocols or fail the
open operation. No intermediate/partial states. This is a non-trivial
refactor for mlx5, because mlx5_nic_enable() returns nothing.
Refactoring seems possible though, its only caller is
mlx5e_attach_netdev(), which returns errors. It's certainly not
something that should be done for a net fix though.
I have a series pending for net-next where the PSP configuration is
hooked to mlx5e_psp_set_config(). I will look into implementing what
you propose there and propagate errors.
Meanwhile, do you want to take these fixes (1 and 2) or maybe just 2
for net or not?
Cosmin.
^ permalink raw reply
* Re: [PATCH v2 0/2] Bluetooth: ISO: Fix KCSAN data-races on iso_pi(sk)
From: patchwork-bot+bluetooth @ 2026-04-21 17:40 UTC (permalink / raw)
To: SeungJu Cheon
Cc: luiz.dentz, marcel, linux-bluetooth, netdev, linux-kernel, me,
skhan, linux-kernel-mentees
In-Reply-To: <20260421025122.55781-1-suunj1331@gmail.com>
Hello:
This series was applied to bluetooth/bluetooth-next.git (master)
by Luiz Augusto von Dentz <luiz.von.dentz@intel.com>:
On Tue, 21 Apr 2026 11:51:20 +0900 you wrote:
> Found while auditing iso_pi(sk) field accesses after a KCSAN report.
> Patch 1/2 is the reported race on iso_pi(sk)->dst in iso_sock_connect();
> patch 2/2 covers related races on other iso_pi(sk) fields accessed in
> iso_connect_{bis,cis}() and iso_connect_ind() that were found by
> inspection during the same audit.
>
> Changes in v2:
> - Patch 1/2: Use sa->iso_bdaddr directly instead of caching the
> bacmp() result in a local variable, as suggested by Luiz [1].
> This avoids reading from iso_pi(sk) entirely for the broadcast
> check.
>
> [...]
Here is the summary with links:
- [v2,1/2] Bluetooth: ISO: Fix data-race on dst in iso_sock_connect()
https://git.kernel.org/bluetooth/bluetooth-next/c/20ca2749b31a
- [v2,2/2] Bluetooth: ISO: Fix data-race on iso_pi(sk) in socket and HCI event paths
https://git.kernel.org/bluetooth/bluetooth-next/c/66d4d518020b
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply
* Re: [PATCH net-deletions] net: remove ax25 and amateur radio (hamradio) subsystem
From: Dan Cross @ 2026-04-21 17:47 UTC (permalink / raw)
To: Stephen Hemminger
Cc: Jakub Kicinski, davem, netdev, edumazet, pabeni, andrew+netdev,
horms, corbet, skhan, federico.vaga, carlos.bilbao, avadhut.naik,
alexs, si.yanteng, dzm91, 2023002089, tsbogend, dsahern,
jani.nikula, mchehab+huawei, gregkh, jirislaby, tytso, herbert,
ebiggers, johannes.berg, geert, pablo, tglx, mashiro.chen, mingo,
dqfext, jreuter, sdf, pkshih, enelsonmoore, mkl, toke, kees,
jlayton, wangliang74, aha310510, takamitz, kuniyu, linux-doc,
linux-mips
In-Reply-To: <20260421101400.67545b20@phoenix.local>
On Tue, Apr 21, 2026 at 1:14 PM Stephen Hemminger
<stephen@networkplumber.org> wrote:
> On Tue, 21 Apr 2026 12:17:23 -0400
> Dan Cross <crossd@gmail.com> wrote:
>
> > On Tue, Apr 21, 2026 at 9:55 AM Stephen Hemminger
> > <stephen@networkplumber.org> wrote:
> > > On Mon, 20 Apr 2026 19:18:23 -0700
> > > Jakub Kicinski <kuba@kernel.org> wrote:
> > > > Remove the amateur radio (AX.25, NET/ROM, ROSE) protocol implementation
> > > > and all associated hamradio device drivers from the kernel tree.
> > > > This set of protocols has long been a huge bug/syzbot magnet,
> > > > and since nobody stepped up to help us deal with the influx
> > > > of the AI-generated bug reports we need to move it out of tree
> > > > to protect our sanity.
> > > >
> > > > The code is moved to an out-of-tree repo:
> > > > https://github.com/linux-netdev/mod-orphan
> > > > if it's cleaned up and reworked there we can accept it back.
> > >
> > > It would be good if these protocols could be done in userspace
> > > or with BPF?
> >
> > Consensus for a userspace implementation is what folks on linux-hams
> > seem to be converging on.
> >
> > The amateur radio protocols are more or less specific to low-speed
> > links, they are not particularly coupled to anything else that
> > requires running in the kernel, and the main coupling point (IP over
> > AX.25) can be implemented via TAP/TUN.
> >
> > There are several popular packages that already implement AX.25 and
> > NET/ROM in user-space (for the interested, LinBPQ seems to be the
> > canonical example). The main missing piece is ROSE, but it is likely
> > easier to add that to an existing package, or potentially something
> > brand new, than keep it in the kernel.
> >
> > There's no compelling reason to keep these protocols in the kernel,
> > whether in-tree or out-of-tree; at least, one has not been
> > articulated.
>
> Thanks, my other concern is carrying support for these in ip commands.
> If not kernel based, then iproute2 doesn't need to worry.
Agreed.
If someone really wants mimic the existing output of those commands in
the context of a userspace implementation, they could write a wrapper
program that invokes the real thing, and extracts relevant information
from the ham protocol implementation, and interpolates it into the
output. It may be an imperfect simulation, but it's probably close
enough for most users.
- Dan C.
^ permalink raw reply
* [PATCH net] hv_sock: Return -EIO for malformed/short packets
From: Dexuan Cui @ 2026-04-21 17:49 UTC (permalink / raw)
To: kys, haiyangz, wei.liu, decui, longli, sgarzare, davem, edumazet,
kuba, pabeni, horms, niuxuewei.nxw, linux-hyperv, virtualization,
netdev, linux-kernel
Cc: stable
Commit f63152958994 fixes a regression, however it fails to report an
error for malformed/short packets -- normally we should never see such
packets, but let's report an error for them just in case.
Fixes: f63152958994 ("hv_sock: Report EOF instead of -EIO for FIN")
Cc: stable@vger.kernel.org
Signed-off-by: Dexuan Cui <decui@microsoft.com>
---
Commit f63152958994 is currently only in net.git's master branch.
net/vmw_vsock/hyperv_transport.c | 29 +++++++++++++++++++----------
1 file changed, 19 insertions(+), 10 deletions(-)
diff --git a/net/vmw_vsock/hyperv_transport.c b/net/vmw_vsock/hyperv_transport.c
index 76e78c83fdbc..8faaa14bccda 100644
--- a/net/vmw_vsock/hyperv_transport.c
+++ b/net/vmw_vsock/hyperv_transport.c
@@ -704,18 +704,27 @@ static s64 hvs_stream_has_data(struct vsock_sock *vsk)
if (hvs->recv_desc) {
/* Here hvs->recv_data_len is 0, so hvs->recv_desc must
* be NULL unless it points to the 0-byte-payload FIN
- * packet: see hvs_update_recv_data().
+ * packet or a malformed/short packet: see
+ * hvs_update_recv_data().
*
- * Here all the payload has been dequeued, but
- * hvs_channel_readable_payload() still returns 1,
- * because the VMBus ringbuffer's read_index is not
- * updated for the FIN packet: hvs_stream_dequeue() ->
- * hv_pkt_iter_next() updates the cached priv_read_index
- * but has no opportunity to update the read_index in
- * hv_pkt_iter_close() as hvs_stream_has_data() returns
- * 0 for the FIN packet, so it won't get dequeued.
+ * If hvs->recv_desc points to the FIN packet, here all
+ * the payload has been dequeued and the peer_shutdown
+ * flag is set, but hvs_channel_readable_payload() still
+ * returns 1, because the VMBus ringbuffer's read_index
+ * is not updated for the FIN packet:
+ * hvs_stream_dequeue() -> hv_pkt_iter_next() updates
+ * the cached priv_read_index but has no opportunity to
+ * update the read_index in hv_pkt_iter_close() as
+ * hvs_stream_has_data() returns 0 for the FIN packet,
+ * so it won't get dequeued.
+ *
+ * In case hvs->recv_desc points to a malformed/short
+ * packet, return -EIO.
*/
- return 0;
+ if (hvs->vsk->peer_shutdown & SEND_SHUTDOWN)
+ return 0;
+ else
+ return -EIO;
}
hvs->recv_desc = hv_pkt_iter_first(hvs->chan);
--
2.49.0
^ permalink raw reply related
* RE: [EXTERNAL] Re: [PATCH net v2] hv_sock: Report EOF instead of -EIO for FIN
From: Dexuan Cui @ 2026-04-21 17:54 UTC (permalink / raw)
To: Jakub Kicinski, Stefano Garzarella
Cc: patchwork-bot+netdevbpf@kernel.org, KY Srinivasan, Haiyang Zhang,
wei.liu@kernel.org, Long Li, davem@davemloft.net,
edumazet@google.com, pabeni@redhat.com, horms@kernel.org,
niuxuewei.nxw@antgroup.com, linux-hyperv@vger.kernel.org,
virtualization@lists.linux.dev, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, stable@vger.kernel.org, Ben Hillis,
levymitchell0@gmail.com
In-Reply-To: <20260421071839.30217a60@kernel.org>
> From: Jakub Kicinski <kuba@kernel.org>
> Sent: Tuesday, April 21, 2026 7:19 AM
> ...
> > Anyway, let's wait for Jakub's or other net maintainers' suggestions.
>
> Yes, you have to post an incremental fix
Thanks for the quick replies! I posted an incremental fix:
https://lore.kernel.org/linux-hyperv/20260421174931.1152238-1-decui@microsoft.com/T/#u
^ permalink raw reply
* RE: [PATCH net v3] hv_sock: Report EOF instead of -EIO for FIN
From: Dexuan Cui @ 2026-04-21 17:57 UTC (permalink / raw)
To: Dexuan Cui, KY Srinivasan, Haiyang Zhang, wei.liu@kernel.org,
Long Li, sgarzare@redhat.com, davem@davemloft.net,
edumazet@google.com, kuba@kernel.org, pabeni@redhat.com,
horms@kernel.org, niuxuewei.nxw@antgroup.com,
linux-hyperv@vger.kernel.org, virtualization@lists.linux.dev,
netdev@vger.kernel.org, linux-kernel@vger.kernel.org
Cc: stable@vger.kernel.org, Ben Hillis, Mitchell Levy
In-Reply-To: <20260421025950.1099495-1-decui@microsoft.com>
> From: Dexuan Cui
> Sent: Monday, April 20, 2026 8:00 PM
Please ignore the email, as I just posted an incremental patch here:
https://lore.kernel.org/linux-hyperv/20260421174931.1152238-1-decui@microsoft.com/T/#u
See the link for more context:
https://lore.kernel.org/linux-hyperv/177672238581.1802062.15838493180057695674.git-patchwork-notify@kernel.org/T/#t
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox