From: Breno Leitao <leitao@debian.org>
To: "Liam R. Howlett (Oracle)" <liam@infradead.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
maple-tree@lists.infradead.org
Subject: Re: [PATCH v2 03/19] maple_tree: Add write lock checking with lockdep sequence numbers
Date: Mon, 20 Jul 2026 09:04:59 -0700 [thread overview]
Message-ID: <al5F7sT-39J_wbbR@gmail.com> (raw)
In-Reply-To: <dzp5qpqjzc3j2jwnamgbv7j2blqcbp6xq5jnwvkm6hzx6fbjni@2vom3uipdkuo>
On Mon, Jul 20, 2026 at 10:03:51AM -0400, Liam R. Howlett (Oracle) wrote:
> On 26/07/08 08:44AM, Breno Leitao wrote:
> > Hello Liam,
>
> Hello!
>
> >
> > On Tue, Jun 30, 2026 at 03:08:27PM -0400, Liam R. Howlett (Oracle) wrote:
> > > +#ifdef CONFIG_RCU_STRICT_GRACE_PERIOD
> > > if (!mt_locked(mas->tree)) {
> > > if (mt_in_rcu(mas->tree))
> > > WARN_ON_ONCE(poll_state_synchronize_rcu(mas->rcu_gp));
> >
> > Testing next-20260706 on arm64 under a stress-ng + fork/exit/mprotect workload,
> > I'm hitting the new maple_tree write-lock sequence check:
> >
> > WARNING: lib/maple_tree.c:1188 at mas_store_prealloc+0x2a8/0xac0
>
> No, this is not expected and I didn't know about it. Does it still
> happen in next? I'll start looking at this while I wait to hear from
> your side.
Yes, I see it on my everyday's test.
>
> That's very odd. Does anything show up with a lockdep build?
Yes, but not sure if this is related. This is happening twice, thus, the
messy kernel message
01:37:12 perf: interrupt took too long (3127 > 3126), lowering kernel.perf_event_max_sample_rate to 63000
01:37:38
------------[ cut here ]------------
WARNING: lib/maple_tree.c:1188 at mas_find+0x620/0x810, CPU#45: stress-ng-sysfs/419993 ^[[36m[warn]^[[0m
Modules linked in: md4(E) vhost_vsock(E) vfio_iommu_type1(E) vfio(E) act_gact(E) tcp_diag(E) inet_diag(E)
------------[ cut here ]------------
cls_bpf(E)
evdev(E) ipmi_devintf(E) ipmi_msghandler(E) spi_nor(E) button(E) sch_fq_codel(E) vhost_net(E) tap(E) tun(E)
WARNING: lib/maple_tree.c:1188 at mas_store+0x2a8/0x678, CPU#7: stress-ng-sysfs/419987 ^[[36m[warn]^[[0m
mpls_iptunnel(E) vhost(E) tls(E)
Modules linked in:
vhost_iotlb(E) mpls_gso(E)
md4(E)
mpls_router(E)
vhost_vsock(E)
loop(E)
vfio_iommu_type1(E)
fou(E)
vfio(E)
efivarfs(E)
act_gact(E)
autofs4(E)
tcp_diag(E)
[last unloaded: test_bpf(E)]
inet_diag(E)
cls_bpf(E) evdev(E) ipmi_devintf(E) ipmi_msghandler(E) spi_nor(E) button(E) sch_fq_codel(E) vhost_net(E)
CPU: 45 UID: 0 PID: 419993 Comm: stress-ng-sysfs Kdump: loaded Tainted: G E 7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full)
tap(E) tun(E) mpls_iptunnel(E)
Tainted: [E]=UNSIGNED_MODULE
vhost(E)
tls(E)
vhost_iotlb(E) mpls_gso(E) mpls_router(E)
pstate: 03401009 (nzcv daif +PAN -UAO +TCO +DIT +SSBS BTYPE=--)
loop(E) fou(E)
pc : mas_find+0x620/0x810
efivarfs(E) autofs4(E) [last unloaded: test_bpf(E)]
lr : mas_find+0x1c8/0x810
sp : ffff80011fa579e0
x29: ffff80011fa579e0
CPU: 7 UID: 0 PID: 419987 Comm: stress-ng-sysfs Kdump: loaded Tainted: G E 7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full)
x28: 1ffff00023f4af5d x27: 1ffff00023f4af62
Tainted: [E]=UNSIGNED_MODULE
x26: ffff0000b65379a0
x25: 00000000000000e7 x24: 1ffff00023f4af6f
pstate: 03401009 (nzcv daif +PAN -UAO +TCO +DIT +SSBS BTYPE=--)
x23: dfff800000000000 x22: ffff00011aa90e78
pc : mas_store+0x2a8/0x678
x21: fffffffffffffffe
x20: ffff80011fa57b30
lr : mas_store+0x2a0/0x678
x19: ffff80011fa57b78
sp : ffff80011e9c7810
x18: 0000000000000000
x29: ffff80011e9c7890
x17: ffff800082ef24c4
x28: 1fffe000507cf8f2
x16: 0000000000000068 x15: 0000000000000001
x27: dfff800000000000
x14: 1fffe001310d8da4
x26: ffff000283e7c790
x13: 0000000000000001
x25: 0000000000000097 x24: 1ffff00023d38f40
x12: 0000aaf4db60a000
x11: ffff00015fb58bf0
x23: dfff800000000000 x22: ffff0009dd07f178
x10: 0000aaf4db60a000
x21: ffff000283e7c780
x9 : 0000000000000300
x20: ffff80011e9c7a00
x8 : 0000000000000000
x19: ffff80011e9c79b8 x18: 0000000000000118
x7 : ffff8000808bcf34
x6 : 0000000000000000
x17: ffff800082ef24c4
x5 : 0000000000000000
x16: 0000000000000068
x4 : 0000000000000001
x15: 000000000000000a
x3 : 0000aaf4db60a000
x14: 1ffff00023d38f07
x13: 0000000000000000
x2 : 0000000000000000
x12: 0000000000000000
x1 : 00000000ffffffff
x0 : 00000000ffffffff
x11: ffff700023d38f11
x10: dfff800000000000
Call trace:
x9 : 0000000000000300
x8 : 0000000000000000
mas_find+0x620/0x810 (P)
x7 : 0000000000000000 x6 : ffff8000808bf46c
x5 : 0000000000000000 x4 : 0000000000000008 x3 : 0000000000000000
unmap_vmas+0x180/0x320
x2 : 0000000000000008 x1 : 00000000ffffffff x0 : 00000000ffffffff
Call trace:
exit_mmap+0x134/0x8f8
mas_store+0x2a8/0x678 (P)
__mmput+0xb8/0x340
dup_mmap+0x698/0x1350
mmput+0x60/0x98
copy_mm+0xe4/0x3c8
do_exit+0x594/0x1a28
do_group_exit+0x180/0x220
copy_process+0x12f0/0x31e8
__arm64_sys_exit_group+0x50/0x58
kernel_clone+0x14c/0x768
__arm64_sys_clone+0x100/0x140
invoke_syscall+0x74/0x188
do_el0_svc+0x10c/0x198
invoke_syscall+0x74/0x188
do_el0_svc+0x10c/0x198
el0_svc+0x64/0x260
el0_svc+0x64/0x260
el0t_64_sync_handler+0x84/0x130
el0t_64_sync_handler+0x84/0x130
el0t_64_sync+0x198/0x1a0
el0t_64_sync+0x198/0x1a0
irq event stamp: 251
irq event stamp: 265
hardirqs last enabled at (251): [<ffff800080911648>] seqcount_lockdep_reader_access+0x68/0xc8
hardirqs last enabled at (265): [<ffff800082da75ec>] _raw_spin_unlock_irqrestore+0x44/0xb0
hardirqs last disabled at (250): [<ffff80008091160c>] seqcount_lockdep_reader_access+0x2c/0xc8
hardirqs last disabled at (264): [<ffff800082da73c0>] _raw_spin_lock_irqsave+0x38/0x90
softirqs last enabled at (136): [<ffff80008011dfb4>] handle_softirqs+0xd8c/0xf08
softirqs last disabled at (131): [<ffff8000800103d0>] __do_softirq+0x20/0x2c
softirqs last enabled at (178): [<ffff800080092ed4>] fpsimd_save_and_flush_current_state+0xac/0x100
---[ end trace 0000000000000000 ]---
softirqs last disabled at (176): [<ffff800080092e58>] fpsimd_save_and_flush_current_state+0x30/0x100
---[ end trace 0000000000000000 ]---
======================================================
WARNING: possible circular locking dependency detected ^[[36m[lockdep]^[[0m
7.2.0-rc2-next-20260710upstream-baseline #19 Tainted: G W E
------------------------------------------------------
stress-ng-sysfs/419983 is trying to acquire lock:
ffff000390315b78 (&mm->mmap_lock){++++}-{4:4}, at: mmap_read_lock_killable+0x20/0xe8
but task is already holding lock:
ffff0000a1747180 (&root->kernfs_rwsem){++++}-{4:4}, at: kernfs_fop_readdir+0x1e8/0x720
which lock already depends on the new lock.
the existing dependency chain (in reverse order) is:
-> #3 (&root->kernfs_rwsem){++++}-{4:4}:
down_read+0x68/0x320
kernfs_find_and_get_ns+0x3c/0xf8
sysfs_notify+0x90/0xd0
btrfs_exclop_finish+0x94/0xb8
btrfs_swap_activate+0xedc/0x1900
setup_swap_extents+0x13c/0x4a8
__arm64_sys_swapon+0xb20/0x1490
invoke_syscall+0x74/0x188
do_el0_svc+0x10c/0x198
el0_svc+0x64/0x260
el0t_64_sync_handler+0x84/0x130
el0t_64_sync+0x198/0x1a0
-> #2 (&ei->i_mmap_lock){++++}-{4:4}:
down_read+0x68/0x320
btrfs_page_mkwrite+0x414/0x1088
do_page_mkwrite+0x124/0x258
do_wp_page+0x52c/0x3270
handle_mm_fault+0xa24/0x1bb0
do_page_fault+0x2ac/0xa90
do_mem_abort+0x78/0x198
el0_da+0x78/0x258
el0t_64_sync_handler+0x90/0x130
el0t_64_sync+0x198/0x1a0
-> #1 (sb_pagefaults){.+.+}-{0:0}:
btrfs_page_mkwrite+0x200/0x1088
do_page_mkwrite+0x124/0x258
do_pte_missing+0xa7c/0x2160
handle_mm_fault+0x9c8/0x1bb0
do_page_fault+0x35c/0xa90
do_translation_fault+0x78/0x90
do_mem_abort+0x78/0x198
el0_da+0x78/0x258
el0t_64_sync_handler+0x90/0x130
el0t_64_sync+0x198/0x1a0
-> #0 (&mm->mmap_lock){++++}-{4:4}:
__lock_acquire+0x1b58/0x3418
lock_acquire+0x160/0x3a0
down_read_killable+0x70/0x358
mmap_read_lock_killable+0x20/0xe8
lock_mm_and_find_vma+0x1e4/0x1f0
do_page_fault+0x314/0xa90
do_translation_fault+0x78/0x90
do_mem_abort+0x78/0x198
el1_abort+0x50/0x78
el1h_64_sync_handler+0x50/0x100
el1h_64_sync+0x6c/0x70
filldir64+0x228/0x500
kernfs_fop_readdir+0x444/0x720
iterate_dir+0x1f4/0x370
__arm64_sys_getdents64+0xc0/0x1e0
invoke_syscall+0x74/0x188
do_el0_svc+0x10c/0x198
el0_svc+0x64/0x260
el0t_64_sync_handler+0x84/0x130
el0t_64_sync+0x198/0x1a0
other info that might help us debug this:
Chain exists of:
&mm->mmap_lock --> &ei->i_mmap_lock --> &root->kernfs_rwsem
Possible unsafe locking scenario:
CPU0 CPU1
---- ----
rlock(&root->kernfs_rwsem);
lock(&ei->i_mmap_lock);
lock(&root->kernfs_rwsem);
rlock(&mm->mmap_lock);
*** DEADLOCK ***
locks held by stress-ng-sysfs/419983: 3, last CPU#4:
#0: ffff0003919b1b30 (&f->f_pos_lock){+.+.}-{4:4}, at: fdget_pos+0x1b0/0x240
#1: ffff0000bfa8d200 (&type->i_mutex_dir_key#3){++++}-{4:4}, at: iterate_dir+0x164/0x370
#2: ffff0000a1747180 (&root->kernfs_rwsem){++++}-{4:4}, at: kernfs_fop_readdir+0x1e8/0x720
stack backtrace:
CPU: 4 UID: 0 PID: 419983 Comm: stress-ng-sysfs Kdump: loaded Tainted: G W E 7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full)
Tainted: [W]=WARN, [E]=UNSIGNED_MODULE
Call trace:
show_stack+0x24/0x38 (C)
__dump_stack+0x28/0x38
dump_stack_lvl+0x74/0xa8
dump_stack+0x18/0x24
print_circular_bug+0x328/0x330
check_noncircular+0x17c/0x1a0
__lock_acquire+0x1b58/0x3418
lock_acquire+0x160/0x3a0
down_read_killable+0x70/0x358
mmap_read_lock_killable+0x20/0xe8
lock_mm_and_find_vma+0x1e4/0x1f0
do_page_fault+0x314/0xa90
do_translation_fault+0x78/0x90
do_mem_abort+0x78/0x198
el1_abort+0x50/0x78
el1h_64_sync_handler+0x50/0x100
el1h_64_sync+0x6c/0x70
filldir64+0x228/0x500 (P)
kernfs_fop_readdir+0x444/0x720
iterate_dir+0x1f4/0x370
__arm64_sys_getdents64+0xc0/0x1e0
invoke_syscall+0x74/0x188
do_el0_svc+0x10c/0x198
el0_svc+0x64/0x260
el0t_64_sync_handler+0x84/0x130
el0t_64_sync+0x198/0x1a0
01:37:47 capability: warning: `stress-ng-cap' uses deprecated v2 capabilities in a way that may be insecure
> Your
> trace comment leads me to believe there is a lot of these triggering
> with CONFIG_RCU_STRICT_GRACE_PERIOD, is that correct?
I don't think this is set in the kernel I am seeing it Why do you think
CONFIG_RCU_STRICT_GRACE_PERIOD is set?
Thanks,
--breno
next prev parent reply other threads:[~2026-07-20 16:05 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-30 19:08 [PATCH v2 00/19] maple_tree: lock checking and clean ups Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 01/19] maple_tree: Add rcu locking check when LOCKDEP is enabled Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 02/19] locking/lockdep: Add sequence counter to held_lock Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 03/19] maple_tree: Add write lock checking with lockdep sequence numbers Liam R. Howlett (Oracle)
2026-07-08 15:44 ` Breno Leitao
2026-07-20 14:03 ` Liam R. Howlett (Oracle)
2026-07-20 16:04 ` Breno Leitao [this message]
2026-07-20 21:41 ` Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 04/19] maple_tree: Documentation fix Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 05/19] maple_tree: Drop dead code from mas_extend_spanning_null() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 06/19] maple_tree: Drop MAPLE_ALLOC_SLOTS Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 07/19] maple_tree: Clarify comments on mas_nomem() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 08/19] maple_tree: Use prefetched value in mas_wr_store_type() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 09/19] maple_tree: Optimise mas_wr_node_store() when not in rcu mode Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 10/19] maple_tree: micro optimisation of mas_wr_store_type() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 11/19] maple_tree: Add bulk parent set helper Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 12/19] maple_tree: Catch race in mas_alloc_cyclic() Liam R. Howlett (Oracle)
2026-06-30 19:41 ` Chuck Lever
2026-06-30 19:08 ` [PATCH v2 13/19] maple_tree: Document that erase may use GFP_KERNEL for allocations Liam R. Howlett (Oracle)
2026-06-30 19:32 ` Rik van Riel
2026-06-30 19:08 ` [PATCH v2 14/19] maple_tree: WARN_ON_ONCE when allocations fail Liam R. Howlett (Oracle)
2026-06-30 23:02 ` Andrew Morton
2026-06-30 19:08 ` [PATCH v2 15/19] maple_tree: Document erase and allocations better Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 16/19] maple_tree: Change two GFP flags in tests Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 17/19] maple_tree: Fix argument name in header Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 18/19] maple_tree: Avoid extra gap calculation Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 19/19] maple_tree: Add helper mas_make_walkable() Liam R. Howlett (Oracle)
2026-06-30 23:05 ` [PATCH v2 00/19] maple_tree: lock checking and clean ups Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=al5F7sT-39J_wbbR@gmail.com \
--to=leitao@debian.org \
--cc=akpm@linux-foundation.org \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=maple-tree@lists.infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox