Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Breno Leitao <leitao@debian.org>
To: "Liam R. Howlett (Oracle)" <liam@infradead.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	linux-mm@kvack.org,  linux-kernel@vger.kernel.org,
	maple-tree@lists.infradead.org
Subject: Re: [PATCH v2 03/19] maple_tree: Add write lock checking with lockdep sequence numbers
Date: Mon, 20 Jul 2026 09:04:59 -0700	[thread overview]
Message-ID: <al5F7sT-39J_wbbR@gmail.com> (raw)
In-Reply-To: <dzp5qpqjzc3j2jwnamgbv7j2blqcbp6xq5jnwvkm6hzx6fbjni@2vom3uipdkuo>

On Mon, Jul 20, 2026 at 10:03:51AM -0400, Liam R. Howlett (Oracle) wrote:
> On 26/07/08 08:44AM, Breno Leitao wrote:
> > Hello Liam,
> 
> Hello!
> 
> > 
> > On Tue, Jun 30, 2026 at 03:08:27PM -0400, Liam R. Howlett (Oracle) wrote:
> > > +#ifdef CONFIG_RCU_STRICT_GRACE_PERIOD
> > >  	if (!mt_locked(mas->tree)) {
> > >  		if (mt_in_rcu(mas->tree))
> > >  			WARN_ON_ONCE(poll_state_synchronize_rcu(mas->rcu_gp));
> > 
> > Testing next-20260706 on arm64 under a stress-ng + fork/exit/mprotect workload,
> > I'm hitting the new maple_tree write-lock sequence check:
> > 
> >     WARNING: lib/maple_tree.c:1188 at mas_store_prealloc+0x2a8/0xac0
> 
> No, this is not expected and I didn't know about it.  Does it still
> happen in next?  I'll start looking at this while I wait to hear from
> your side.

Yes, I see it on my everyday's test.

> 
> That's very odd.  Does anything show up with a lockdep build?

Yes, but not sure if this is related. This is happening twice, thus, the
messy kernel message

 01:37:12  perf: interrupt took too long (3127 > 3126), lowering kernel.perf_event_max_sample_rate to 63000
 01:37:38  
           ------------[ cut here ]------------
           WARNING: lib/maple_tree.c:1188 at mas_find+0x620/0x810, CPU#45: stress-ng-sysfs/419993  ^[[36m[warn]^[[0m
           Modules linked in: md4(E) vhost_vsock(E) vfio_iommu_type1(E) vfio(E) act_gact(E) tcp_diag(E) inet_diag(E)
           ------------[ cut here ]------------
            cls_bpf(E)
            evdev(E) ipmi_devintf(E) ipmi_msghandler(E) spi_nor(E) button(E) sch_fq_codel(E) vhost_net(E) tap(E) tun(E)
           WARNING: lib/maple_tree.c:1188 at mas_store+0x2a8/0x678, CPU#7: stress-ng-sysfs/419987  ^[[36m[warn]^[[0m
            mpls_iptunnel(E) vhost(E) tls(E)
           Modules linked in:
            vhost_iotlb(E) mpls_gso(E)
            md4(E)
            mpls_router(E)
            vhost_vsock(E)
            loop(E)
            vfio_iommu_type1(E)
            fou(E)
            vfio(E)
            efivarfs(E)
            act_gact(E)
            autofs4(E)
            tcp_diag(E)
            [last unloaded: test_bpf(E)]
            inet_diag(E)
           
            cls_bpf(E) evdev(E) ipmi_devintf(E) ipmi_msghandler(E) spi_nor(E) button(E) sch_fq_codel(E) vhost_net(E)
           CPU: 45 UID: 0 PID: 419993 Comm: stress-ng-sysfs Kdump: loaded Tainted: G            E       7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full) 
            tap(E) tun(E) mpls_iptunnel(E)
           Tainted: [E]=UNSIGNED_MODULE
            vhost(E)
            tls(E)
            vhost_iotlb(E) mpls_gso(E) mpls_router(E)
           pstate: 03401009 (nzcv daif +PAN -UAO +TCO +DIT +SSBS BTYPE=--)
            loop(E) fou(E)
           pc : mas_find+0x620/0x810
            efivarfs(E) autofs4(E) [last unloaded: test_bpf(E)]
           lr : mas_find+0x1c8/0x810
           
           sp : ffff80011fa579e0
           x29: ffff80011fa579e0
           CPU: 7 UID: 0 PID: 419987 Comm: stress-ng-sysfs Kdump: loaded Tainted: G            E       7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full) 
            x28: 1ffff00023f4af5d x27: 1ffff00023f4af62
           Tainted: [E]=UNSIGNED_MODULE
           x26: ffff0000b65379a0
            x25: 00000000000000e7 x24: 1ffff00023f4af6f
           pstate: 03401009 (nzcv daif +PAN -UAO +TCO +DIT +SSBS BTYPE=--)
           
           x23: dfff800000000000 x22: ffff00011aa90e78
           pc : mas_store+0x2a8/0x678
            x21: fffffffffffffffe
           x20: ffff80011fa57b30
           lr : mas_store+0x2a0/0x678
            x19: ffff80011fa57b78
           sp : ffff80011e9c7810
            x18: 0000000000000000
           x29: ffff80011e9c7890
           x17: ffff800082ef24c4
            x28: 1fffe000507cf8f2
            x16: 0000000000000068 x15: 0000000000000001
            x27: dfff800000000000
           
           
           x14: 1fffe001310d8da4
           x26: ffff000283e7c790
            x13: 0000000000000001
            x25: 0000000000000097 x24: 1ffff00023d38f40
            x12: 0000aaf4db60a000
           
           x11: ffff00015fb58bf0
           x23: dfff800000000000 x22: ffff0009dd07f178
            x10: 0000aaf4db60a000
            x21: ffff000283e7c780
            x9 : 0000000000000300
           
           
           x20: ffff80011e9c7a00
           x8 : 0000000000000000
            x19: ffff80011e9c79b8 x18: 0000000000000118
            x7 : ffff8000808bcf34
           
            x6 : 0000000000000000
           x17: ffff800082ef24c4
           
           x5 : 0000000000000000
            x16: 0000000000000068
            x4 : 0000000000000001
            x15: 000000000000000a
            x3 : 0000aaf4db60a000
           
           x14: 1ffff00023d38f07
           
            x13: 0000000000000000
           x2 : 0000000000000000
            x12: 0000000000000000
            x1 : 00000000ffffffff
           
            x0 : 00000000ffffffff
           x11: ffff700023d38f11
           
            x10: dfff800000000000
           Call trace:
            x9 : 0000000000000300
           x8 : 0000000000000000
            mas_find+0x620/0x810 (P)
            x7 : 0000000000000000 x6 : ffff8000808bf46c
           x5 : 0000000000000000 x4 : 0000000000000008 x3 : 0000000000000000
            unmap_vmas+0x180/0x320
           x2 : 0000000000000008 x1 : 00000000ffffffff x0 : 00000000ffffffff
           Call trace:
            exit_mmap+0x134/0x8f8
            mas_store+0x2a8/0x678 (P)
            __mmput+0xb8/0x340
            dup_mmap+0x698/0x1350
            mmput+0x60/0x98
            copy_mm+0xe4/0x3c8
            do_exit+0x594/0x1a28
            do_group_exit+0x180/0x220
            copy_process+0x12f0/0x31e8
            __arm64_sys_exit_group+0x50/0x58
            kernel_clone+0x14c/0x768
            __arm64_sys_clone+0x100/0x140
            invoke_syscall+0x74/0x188
            do_el0_svc+0x10c/0x198
            invoke_syscall+0x74/0x188
            do_el0_svc+0x10c/0x198
            el0_svc+0x64/0x260
            el0_svc+0x64/0x260
            el0t_64_sync_handler+0x84/0x130
            el0t_64_sync_handler+0x84/0x130
            el0t_64_sync+0x198/0x1a0
            el0t_64_sync+0x198/0x1a0
           irq event stamp: 251
           irq event stamp: 265
           hardirqs last  enabled at (251): [<ffff800080911648>] seqcount_lockdep_reader_access+0x68/0xc8
           hardirqs last  enabled at (265): [<ffff800082da75ec>] _raw_spin_unlock_irqrestore+0x44/0xb0
           hardirqs last disabled at (250): [<ffff80008091160c>] seqcount_lockdep_reader_access+0x2c/0xc8
           hardirqs last disabled at (264): [<ffff800082da73c0>] _raw_spin_lock_irqsave+0x38/0x90
           softirqs last  enabled at (136): [<ffff80008011dfb4>] handle_softirqs+0xd8c/0xf08
           softirqs last disabled at (131): [<ffff8000800103d0>] __do_softirq+0x20/0x2c
           softirqs last  enabled at (178): [<ffff800080092ed4>] fpsimd_save_and_flush_current_state+0xac/0x100
           ---[ end trace 0000000000000000 ]---
           softirqs last disabled at (176): [<ffff800080092e58>] fpsimd_save_and_flush_current_state+0x30/0x100
           ---[ end trace 0000000000000000 ]---
           ======================================================
           WARNING: possible circular locking dependency detected  ^[[36m[lockdep]^[[0m
           7.2.0-rc2-next-20260710upstream-baseline #19 Tainted: G        W   E      
           ------------------------------------------------------
           stress-ng-sysfs/419983 is trying to acquire lock:
           ffff000390315b78 (&mm->mmap_lock){++++}-{4:4}, at: mmap_read_lock_killable+0x20/0xe8
           
but task is already holding lock:
           ffff0000a1747180 (&root->kernfs_rwsem){++++}-{4:4}, at: kernfs_fop_readdir+0x1e8/0x720
           
which lock already depends on the new lock.

           
the existing dependency chain (in reverse order) is:
           
-> #3 (&root->kernfs_rwsem){++++}-{4:4}:
                  down_read+0x68/0x320
                  kernfs_find_and_get_ns+0x3c/0xf8
                  sysfs_notify+0x90/0xd0
                  btrfs_exclop_finish+0x94/0xb8
                  btrfs_swap_activate+0xedc/0x1900
                  setup_swap_extents+0x13c/0x4a8
                  __arm64_sys_swapon+0xb20/0x1490
                  invoke_syscall+0x74/0x188
                  do_el0_svc+0x10c/0x198
                  el0_svc+0x64/0x260
                  el0t_64_sync_handler+0x84/0x130
                  el0t_64_sync+0x198/0x1a0
           
-> #2 (&ei->i_mmap_lock){++++}-{4:4}:
                  down_read+0x68/0x320
                  btrfs_page_mkwrite+0x414/0x1088
                  do_page_mkwrite+0x124/0x258
                  do_wp_page+0x52c/0x3270
                  handle_mm_fault+0xa24/0x1bb0
                  do_page_fault+0x2ac/0xa90
                  do_mem_abort+0x78/0x198
                  el0_da+0x78/0x258
                  el0t_64_sync_handler+0x90/0x130
                  el0t_64_sync+0x198/0x1a0
           
-> #1 (sb_pagefaults){.+.+}-{0:0}:
                  btrfs_page_mkwrite+0x200/0x1088
                  do_page_mkwrite+0x124/0x258
                  do_pte_missing+0xa7c/0x2160
                  handle_mm_fault+0x9c8/0x1bb0
                  do_page_fault+0x35c/0xa90
                  do_translation_fault+0x78/0x90
                  do_mem_abort+0x78/0x198
                  el0_da+0x78/0x258
                  el0t_64_sync_handler+0x90/0x130
                  el0t_64_sync+0x198/0x1a0
           
-> #0 (&mm->mmap_lock){++++}-{4:4}:
                  __lock_acquire+0x1b58/0x3418
                  lock_acquire+0x160/0x3a0
                  down_read_killable+0x70/0x358
                  mmap_read_lock_killable+0x20/0xe8
                  lock_mm_and_find_vma+0x1e4/0x1f0
                  do_page_fault+0x314/0xa90
                  do_translation_fault+0x78/0x90
                  do_mem_abort+0x78/0x198
                  el1_abort+0x50/0x78
                  el1h_64_sync_handler+0x50/0x100
                  el1h_64_sync+0x6c/0x70
                  filldir64+0x228/0x500
                  kernfs_fop_readdir+0x444/0x720
                  iterate_dir+0x1f4/0x370
                  __arm64_sys_getdents64+0xc0/0x1e0
                  invoke_syscall+0x74/0x188
                  do_el0_svc+0x10c/0x198
                  el0_svc+0x64/0x260
                  el0t_64_sync_handler+0x84/0x130
                  el0t_64_sync+0x198/0x1a0
           
other info that might help us debug this:

           Chain exists of:
  &mm->mmap_lock --> &ei->i_mmap_lock --> &root->kernfs_rwsem

            Possible unsafe locking scenario:

                  CPU0                    CPU1
                  ----                    ----
             rlock(&root->kernfs_rwsem);
                                          lock(&ei->i_mmap_lock);
                                          lock(&root->kernfs_rwsem);
             rlock(&mm->mmap_lock);
           
 *** DEADLOCK ***

           locks held by stress-ng-sysfs/419983: 3, last CPU#4:
            #0: ffff0003919b1b30 (&f->f_pos_lock){+.+.}-{4:4}, at: fdget_pos+0x1b0/0x240
            #1: ffff0000bfa8d200 (&type->i_mutex_dir_key#3){++++}-{4:4}, at: iterate_dir+0x164/0x370
            #2: ffff0000a1747180 (&root->kernfs_rwsem){++++}-{4:4}, at: kernfs_fop_readdir+0x1e8/0x720
           
stack backtrace:
           CPU: 4 UID: 0 PID: 419983 Comm: stress-ng-sysfs Kdump: loaded Tainted: G        W   E       7.2.0-rc2-next-20260710upstream-baseline #19 PREEMPT(full) 
           Tainted: [W]=WARN, [E]=UNSIGNED_MODULE
           Call trace:
            show_stack+0x24/0x38 (C)
            __dump_stack+0x28/0x38
            dump_stack_lvl+0x74/0xa8
            dump_stack+0x18/0x24
            print_circular_bug+0x328/0x330
            check_noncircular+0x17c/0x1a0
            __lock_acquire+0x1b58/0x3418
            lock_acquire+0x160/0x3a0
            down_read_killable+0x70/0x358
            mmap_read_lock_killable+0x20/0xe8
            lock_mm_and_find_vma+0x1e4/0x1f0
            do_page_fault+0x314/0xa90
            do_translation_fault+0x78/0x90
            do_mem_abort+0x78/0x198
            el1_abort+0x50/0x78
            el1h_64_sync_handler+0x50/0x100
            el1h_64_sync+0x6c/0x70
            filldir64+0x228/0x500 (P)
            kernfs_fop_readdir+0x444/0x720
            iterate_dir+0x1f4/0x370
            __arm64_sys_getdents64+0xc0/0x1e0
            invoke_syscall+0x74/0x188
            do_el0_svc+0x10c/0x198
            el0_svc+0x64/0x260
            el0t_64_sync_handler+0x84/0x130
            el0t_64_sync+0x198/0x1a0
 01:37:47  capability: warning: `stress-ng-cap' uses deprecated v2 capabilities in a way that may be insecure

> Your
> trace comment leads me to believe there is a lot of these triggering
> with CONFIG_RCU_STRICT_GRACE_PERIOD, is that correct?

I don't think this is set in the kernel I am seeing it Why do you think
CONFIG_RCU_STRICT_GRACE_PERIOD is set?

Thanks,
--breno


  reply	other threads:[~2026-07-20 16:05 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-30 19:08 [PATCH v2 00/19] maple_tree: lock checking and clean ups Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 01/19] maple_tree: Add rcu locking check when LOCKDEP is enabled Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 02/19] locking/lockdep: Add sequence counter to held_lock Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 03/19] maple_tree: Add write lock checking with lockdep sequence numbers Liam R. Howlett (Oracle)
2026-07-08 15:44   ` Breno Leitao
2026-07-20 14:03     ` Liam R. Howlett (Oracle)
2026-07-20 16:04       ` Breno Leitao [this message]
2026-07-20 21:41         ` Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 04/19] maple_tree: Documentation fix Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 05/19] maple_tree: Drop dead code from mas_extend_spanning_null() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 06/19] maple_tree: Drop MAPLE_ALLOC_SLOTS Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 07/19] maple_tree: Clarify comments on mas_nomem() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 08/19] maple_tree: Use prefetched value in mas_wr_store_type() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 09/19] maple_tree: Optimise mas_wr_node_store() when not in rcu mode Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 10/19] maple_tree: micro optimisation of mas_wr_store_type() Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 11/19] maple_tree: Add bulk parent set helper Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 12/19] maple_tree: Catch race in mas_alloc_cyclic() Liam R. Howlett (Oracle)
2026-06-30 19:41   ` Chuck Lever
2026-06-30 19:08 ` [PATCH v2 13/19] maple_tree: Document that erase may use GFP_KERNEL for allocations Liam R. Howlett (Oracle)
2026-06-30 19:32   ` Rik van Riel
2026-06-30 19:08 ` [PATCH v2 14/19] maple_tree: WARN_ON_ONCE when allocations fail Liam R. Howlett (Oracle)
2026-06-30 23:02   ` Andrew Morton
2026-06-30 19:08 ` [PATCH v2 15/19] maple_tree: Document erase and allocations better Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 16/19] maple_tree: Change two GFP flags in tests Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 17/19] maple_tree: Fix argument name in header Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 18/19] maple_tree: Avoid extra gap calculation Liam R. Howlett (Oracle)
2026-06-30 19:08 ` [PATCH v2 19/19] maple_tree: Add helper mas_make_walkable() Liam R. Howlett (Oracle)
2026-06-30 23:05 ` [PATCH v2 00/19] maple_tree: lock checking and clean ups Andrew Morton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=al5F7sT-39J_wbbR@gmail.com \
    --to=leitao@debian.org \
    --cc=akpm@linux-foundation.org \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=maple-tree@lists.infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox