All of lore.kernel.org
 help / color / mirror / Atom feed
* Fwd: btrfs: OOM in split_item().
       [not found] <nka8fLGwxf2NcMXCcuLi4CkGSIKrKG4Sc9IbvTTBG5L6T63w6tImgOKYEmu-pWaoZZFO4Eqg8zDyMggihfvlFRTd1Vubz6E52PVEhU83gn8=@protonmail.com>
@ 2026-09-07 22:15 ` Qu Wenruo
  2026-09-08 13:14   ` David Sterba
  0 siblings, 1 reply; 6+ messages in thread
From: Qu Wenruo @ 2026-09-07 22:15 UTC (permalink / raw)
  To: linux-btrfs

[-- Attachment #1: Type: text/plain, Size: 611 bytes --]

Forwarded for archive purposes, as patchcheck requires a link: tag 
following reported-by: tag.

Meanwhile the original report is only a private mail to me.

-------- 转发的消息 --------
主题: 	btrfs: OOM in split_item().
日期: 	Mon, 07 Sep 2026 20:10:56 +0000
发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
收件人: 	wqu@suse.com <wqu@suse.com>



Hello.

Attached is a kernel log entry showing a strange out of memory issue 
that rarely seems to happen.
It's kind of annoying and makes me wish GFP_NOFS would go the way of the 
dinosaur.

Thanks and have a good day / evening.


[-- Attachment #2: btrfs_oom_splat.txt --]
[-- Type: text/plain, Size: 4999 bytes --]


[5558615.041527] kworker/u69:8: page allocation failure: order:4, mode:0x40c40(GFP_NOFS|__GFP_COMP), nodemask=(null)
[5558615.041540] CPU: 3 UID: 0 PID: 1154528 Comm: kworker/u69:8 Not tainted 7.0.2 #1 PREEMPTLAZY 
[5558615.041544] Workqueue: events_unbound btrfs_async_reclaim_metadata_space
[5558615.041550] Call Trace:
[5558615.041553]  <TASK>
[5558615.041554]  dump_stack_lvl+0x47/0x60
[5558615.041558]  warn_alloc.cold+0x67/0xec
[5558615.041561]  __alloc_pages_slowpath.constprop.0+0x9bf/0xed0
[5558615.041565]  __alloc_frozen_pages_noprof+0x1ac/0x1c0
[5558615.041567]  ___kmalloc_large_node+0x9d/0xc0
[5558615.041569]  __kmalloc_noprof+0x17b/0x1f0
[5558615.041571]  split_item+0x9e/0x2e0
[5558615.041574]  ? memset_extent_buffer+0xd3/0x120
[5558615.041576]  btrfs_del_csums+0x285/0x400
[5558615.041578]  ? release_extent_buffer+0x3c/0x100
[5558615.041579]  ? btrfs_csum_root+0x94/0xc0
[5558615.041581]  __btrfs_free_extent.isra.0+0x6de/0x12b0
[5558615.041583]  __btrfs_run_delayed_refs+0x522/0x10c0
[5558615.041586]  ? __write_extent_buffer+0x13d/0x1d0
[5558615.041587]  ? btrfs_qgroup_convert_reserved_meta+0x20/0x290
[5558615.041589]  ? btrfs_block_rsv_release+0x153/0x1f0
[5558615.041591]  ? __btrfs_update_delayed_inode+0xcb/0x450
[5558615.041593]  btrfs_run_delayed_refs+0x4d/0x1d0
[5558615.041595]  flush_space+0x34d/0x4e0
[5558615.041597]  ? btrfs_reduce_alloc_profile+0xa2/0x1a0
[5558615.041599]  ? calc_available_free_space.isra.0+0x58/0x90
[5558615.041601]  ? btrfs_calc_reclaim_metadata_size+0x20/0x60
[5558615.041603]  do_async_reclaim_metadata_space+0x89/0x1d0
[5558615.041605]  ? mempool_free+0x3b/0x60
[5558615.041609]  btrfs_async_reclaim_metadata_space+0x44/0x60
[5558615.041611]  process_one_work+0x145/0x230
[5558615.041614]  worker_thread+0x185/0x2e0
[5558615.041616]  ? rescuer_thread+0x4a0/0x4a0
[5558615.041617]  kthread+0xca/0x100
[5558615.041619]  ? kthreads_online_cpu+0x10/0x10
[5558615.041620]  ret_from_fork+0x14e/0x200
[5558615.041623]  ? kthreads_online_cpu+0x10/0x10
[5558615.041624]  ? kthreads_online_cpu+0x10/0x10
[5558615.041625]  ret_from_fork_asm+0x11/0x20
[5558615.041628]  </TASK>
[5558615.041629] Mem-Info:
[5558615.041635] active_anon:22695 inactive_anon:822896 isolated_anon:0
                  active_file:3274307 inactive_file:2571939 isolated_file:0
                  unevictable:582 dirty:20975 writeback:161089
                  slab_reclaimable:1024567 slab_unreclaimable:61165
                  mapped:153706 shmem:48899 pagetables:8758
                  sec_pagetables:0 bounce:0
                  kernel_misc_reclaimable:0
                  free:254495 free_pcp:104 free_cma:0
[5558615.041641] Node 0 active_anon:90780kB inactive_anon:3291584kB active_file:13097228kB inactive_file:10287756kB unevictable:2328kB isolated(anon):0kB isolated(file):0kB mapped:614824kB dirty:83900kB writeback:644356kB shmem:195596kB kernel_stack:19392kB pagetables:35032kB sec_pagetables:0kB all_unreclaimable? no Balloon:0kB
[5558615.041645] DMA32 free:255048kB boost:8192kB min:10336kB low:13400kB high:16464kB reserved_highatomic:4096KB free_highatomic:3976KB active_anon:13720kB inactive_anon:91944kB active_file:924484kB inactive_file:674012kB unevictable:244kB writepending:150372kB zspages:0kB present:3329364kB managed:3084408kB mlocked:244kB bounce:0kB free_pcp:0kB local_pcp:0kB free_cma:0kB
[5558615.041650] lowmem_reserve[]: 0 28864 28864 28864
[5558615.041654] Normal free:762932kB boost:65536kB min:86232kB low:115788kB high:145344kB reserved_highatomic:4096KB free_highatomic:4096KB active_anon:76968kB inactive_anon:3199640kB active_file:12172576kB inactive_file:9614396kB unevictable:2084kB writepending:578116kB zspages:0kB present:30133760kB managed:29556920kB mlocked:2084kB bounce:0kB free_pcp:416kB local_pcp:0kB free_cma:0kB
[5558615.041659] lowmem_reserve[]: 0 0 0 0
[5558615.041661] DMA32: 20347*4kB (UME) 12098*8kB (UMEH) 3960*16kB (UMEH) 293*32kB (MEH) 7*64kB (H) 1*128kB (H) 4*256kB (H) 0*512kB 2*1024kB (H) 0*2048kB 0*4096kB = 254556kB
[5558615.041675] Normal: 116705*4kB (UME) 28108*8kB (UME) 3330*16kB (UME) 379*32kB (UME) 0*64kB 0*128kB 0*256kB 0*512kB 0*1024kB 0*2048kB 1*4096kB (H) = 761188kB
[5558615.041686] 5801327 total pagecache pages
[5558615.041687] 8365781 pages RAM
[5558615.041688] 0 pages HighMem/MovableOnly
[5558615.041689] 205449 pages reserved
[5558615.041690] 0 pages hwpoisoned
[5558615.041691] BTRFS error (device dm-3 state A): Transaction aborted (error -12)
[5558615.041696] BTRFS: error (device dm-3 state A) in btrfs_del_csums:1053: errno=-12 Out of memory
[5558615.041698] BTRFS info (device dm-3 state EA): forced readonly
[5558615.041701] BTRFS: error (device dm-3 state EA) in do_free_extent_accounting:3168: errno=-12 Out of memory
[5558615.041704] BTRFS error (device dm-3 state EA): failed to run delayed ref for logical 1202913873920 num_bytes 274432 type 184 action 2 ref_mod 1: -12
[5558615.041708] BTRFS: error (device dm-3 state EA) in btrfs_run_delayed_refs:2247: errno=-12 Out of memory


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Fwd: btrfs: OOM in split_item().
  2026-09-07 22:15 ` Fwd: btrfs: OOM in split_item() Qu Wenruo
@ 2026-09-08 13:14   ` David Sterba
  2026-09-08 21:27     ` Qu Wenruo
  0 siblings, 1 reply; 6+ messages in thread
From: David Sterba @ 2026-09-08 13:14 UTC (permalink / raw)
  To: Qu Wenruo; +Cc: linux-btrfs

On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
> Forwarded for archive purposes, as patchcheck requires a link: tag 
> following reported-by: tag.
> 
> Meanwhile the original report is only a private mail to me.
> 
> -------- 转发的消息 --------
> 主题: 	btrfs: OOM in split_item().
> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
> 收件人: 	wqu@suse.com <wqu@suse.com>
> 
> 
> 
> Hello.
> 
> Attached is a kernel log entry showing a strange out of memory issue 
> that rarely seems to happen.
> It's kind of annoying and makes me wish GFP_NOFS would go the way of the 
> dinosaur.

What is the story behind that? One way or another we will get GPF_NOFS
semantics, either the flag or the scoped NOFS and this limits the
allocator.

Using 64k nodes is problematic and so I'd rather people not use it on 4k
systems.

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Fwd: btrfs: OOM in split_item().
  2026-09-08 13:14   ` David Sterba
@ 2026-09-08 21:27     ` Qu Wenruo
  2026-09-09  0:44       ` David Sterba
  0 siblings, 1 reply; 6+ messages in thread
From: Qu Wenruo @ 2026-09-08 21:27 UTC (permalink / raw)
  To: dsterba; +Cc: linux-btrfs



在 2026/9/8 22:44, David Sterba 写道:
> On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
>> Forwarded for archive purposes, as patchcheck requires a link: tag
>> following reported-by: tag.
>>
>> Meanwhile the original report is only a private mail to me.
>>
>> -------- 转发的消息 --------
>> 主题: 	btrfs: OOM in split_item().
>> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
>> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
>> 收件人: 	wqu@suse.com <wqu@suse.com>
>>
>>
>>
>> Hello.
>>
>> Attached is a kernel log entry showing a strange out of memory issue
>> that rarely seems to happen.
>> It's kind of annoying and makes me wish GFP_NOFS would go the way of the
>> dinosaur.
> 
> What is the story behind that? One way or another we will get GPF_NOFS
> semantics, either the flag or the scoped NOFS and this limits the
> allocator.

No extra follow up unfortunately.

> 
> Using 64k nodes is problematic and so I'd rather people not use it on 4k
> systems.

In fact the only problematic part in b-tree operation is exactly the 
vmalloc() I'm fixing.

Other than that I see no obvious problem related to 64K nodesize on 4K 
page systems.

The way we handle metadata is already vmalloc() like, we allocate page 
sized folios for 64K nodes anyway.

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Fwd: btrfs: OOM in split_item().
  2026-09-08 21:27     ` Qu Wenruo
@ 2026-09-09  0:44       ` David Sterba
  2026-09-09  1:06         ` Qu Wenruo
  0 siblings, 1 reply; 6+ messages in thread
From: David Sterba @ 2026-09-09  0:44 UTC (permalink / raw)
  To: Qu Wenruo; +Cc: linux-btrfs

On Wed, Sep 09, 2026 at 06:57:59AM +0930, Qu Wenruo wrote:
> 在 2026/9/8 22:44, David Sterba 写道:
> > On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
> >> Forwarded for archive purposes, as patchcheck requires a link: tag
> >> following reported-by: tag.
> >>
> >> Meanwhile the original report is only a private mail to me.
> >>
> >> -------- 转发的消息 --------
> >> 主题: 	btrfs: OOM in split_item().
> >> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
> >> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
> >> 收件人: 	wqu@suse.com <wqu@suse.com>
> >>
> >>
> >>
> >> Hello.
> >>
> >> Attached is a kernel log entry showing a strange out of memory issue
> >> that rarely seems to happen.
> >> It's kind of annoying and makes me wish GFP_NOFS would go the way of the
> >> dinosaur.
> > 
> > What is the story behind that? One way or another we will get GPF_NOFS
> > semantics, either the flag or the scoped NOFS and this limits the
> > allocator.
> 
> No extra follow up unfortunately.

For the record, as it was in a separate mail, that GFP_NOFS can
sometimes fail because of page fragmentation and limited options for the
allocator.

> > Using 64k nodes is problematic and so I'd rather people not use it on 4k
> > systems.
> 
> In fact the only problematic part in b-tree operation is exactly the 
> vmalloc() I'm fixing.
> 
> Other than that I see no obvious problem related to 64K nodesize on 4K 
> page systems.

Technically if the allocation has a fallback then there's no problem. In
practice the 64K contiguous memory needs to get shifted for almost every
metadata insertion. COW requires the whole 64K set of pages to be
allocated. So there's a lot of dead weight carried around.

It's a tradeoff, 4K nodesize would need taller b-tree for the same
amount of raw items. This is worse due to lock contention and concurrent
changes. I think the 16K is a reasonable middle ground.

> The way we handle metadata is already vmalloc() like, we allocate page 
> sized folios for 64K nodes anyway.

Yeah, that we've been heading towards large folios also means contiguous
ranges for nodes. Memory management is aware of that and allocator
should provides that by the means of compaction. I don't know how much
it could be improved.

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Fwd: btrfs: OOM in split_item().
  2026-09-09  0:44       ` David Sterba
@ 2026-09-09  1:06         ` Qu Wenruo
  2026-09-11 17:55           ` David Sterba
  0 siblings, 1 reply; 6+ messages in thread
From: Qu Wenruo @ 2026-09-09  1:06 UTC (permalink / raw)
  To: dsterba; +Cc: linux-btrfs



在 2026/9/9 10:14, David Sterba 写道:
> On Wed, Sep 09, 2026 at 06:57:59AM +0930, Qu Wenruo wrote:
>> 在 2026/9/8 22:44, David Sterba 写道:
>>> On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
>>>> Forwarded for archive purposes, as patchcheck requires a link: tag
>>>> following reported-by: tag.
>>>>
>>>> Meanwhile the original report is only a private mail to me.
>>>>
>>>> -------- 转发的消息 --------
>>>> 主题: 	btrfs: OOM in split_item().
>>>> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
>>>> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
>>>> 收件人: 	wqu@suse.com <wqu@suse.com>
>>>>
>>>>
>>>>
>>>> Hello.
>>>>
>>>> Attached is a kernel log entry showing a strange out of memory issue
>>>> that rarely seems to happen.
>>>> It's kind of annoying and makes me wish GFP_NOFS would go the way of the
>>>> dinosaur.
>>>
>>> What is the story behind that? One way or another we will get GPF_NOFS
>>> semantics, either the flag or the scoped NOFS and this limits the
>>> allocator.
>>
>> No extra follow up unfortunately.
> 
> For the record, as it was in a separate mail, that GFP_NOFS can
> sometimes fail because of page fragmentation and limited options for the
> allocator.
> 
>>> Using 64k nodes is problematic and so I'd rather people not use it on 4k
>>> systems.
>>
>> In fact the only problematic part in b-tree operation is exactly the
>> vmalloc() I'm fixing.
>>
>> Other than that I see no obvious problem related to 64K nodesize on 4K
>> page systems.
> 
> Technically if the allocation has a fallback then there's no problem. In
> practice the 64K contiguous memory needs to get shifted for almost every
> metadata insertion.

That is not true, at least not for all cases.

The 64K buffer is only needed when we got a huge item to split. 
Meanwhile the most common cases of btrfs items are all fixed sized.

Some have variable length like DIR items, but name length should not 
exceed 4K either.

The real exceptions are:

- XATTR items
   User can provide a huge XATTR, and the end user may even intentionally
   choose 64K nodesize just for such large XATTR support.

   This should be rare, but I still remember some users are using XATTR
   to store extra checksums.
   But on the other hand, such XATTR should not be split, so nothing to
   bother yet.

- Inline extents
   For 4K sectorsize it's never a problem, even for bs > ps support it's
   not a problem either as we do not allow inlined extent for extent
   larger than 4K.

   But it's still possible to hit a 64K inlined extent if it's created
   on 64K page systems, then mount on 4K page systems.

   It's not only very rare, but also we won't split the item, so again
   not to bother.

- Csum items
   This is the exact case we're hitting.

   The point here is, we can use as large as the whole leaf for a csum
   item.
   This makes split much harder, requiring a huge buffer for such split.

   I think we should introduce an artificial limit on the csum item size.
   Keep it 16K at max should greatly reduce the need for large buffer.
   And the cost is pretty minimal, just extra btrfs_items.
   At most it will be 3 * 25 bytes for 64K node size, I think it's
   definitely acceptable.

   By that, we also reduce the buffer requirement for such csum item
   split.

> COW requires the whole 64K set of pages to be
> allocated. So there's a lot of dead weight carried around.
> 
> It's a tradeoff, 4K nodesize would need taller b-tree for the same
> amount of raw items. This is worse due to lock contention and concurrent
> changes. I think the 16K is a reasonable middle ground.

It's a trade-off for nodesize selection, but for this particular csum 
item split, I think we should fix it no matter if you want to prevent 
64K node size or not.

Personally I do not want to discourage any valid nodesize/sectorsize 
combination.

> 
>> The way we handle metadata is already vmalloc() like, we allocate page
>> sized folios for 64K nodes anyway.
> 
> Yeah, that we've been heading towards large folios also means contiguous
> ranges for nodes. Memory management is aware of that and allocator
> should provides that by the means of compaction. I don't know how much
> it could be improved.


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Fwd: btrfs: OOM in split_item().
  2026-09-09  1:06         ` Qu Wenruo
@ 2026-09-11 17:55           ` David Sterba
  0 siblings, 0 replies; 6+ messages in thread
From: David Sterba @ 2026-09-11 17:55 UTC (permalink / raw)
  To: Qu Wenruo; +Cc: dsterba, linux-btrfs

On Wed, Sep 09, 2026 at 10:36:10AM +0930, Qu Wenruo wrote:
> 在 2026/9/9 10:14, David Sterba 写道:
> > On Wed, Sep 09, 2026 at 06:57:59AM +0930, Qu Wenruo wrote:
> >> 在 2026/9/8 22:44, David Sterba 写道:
> >>> On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
> >>>> Forwarded for archive purposes, as patchcheck requires a link: tag
> >>>> following reported-by: tag.
> >>>>
> >>>> Meanwhile the original report is only a private mail to me.
> >>>>
> >>>> -------- 转发的消息 --------
> >>>> 主题: 	btrfs: OOM in split_item().
> >>>> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
> >>>> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
> >>>> 收件人: 	wqu@suse.com <wqu@suse.com>
> >>>>
> >>>>
> >>>>
> >>>> Hello.
> >>>>
> >>>> Attached is a kernel log entry showing a strange out of memory issue
> >>>> that rarely seems to happen.
> >>>> It's kind of annoying and makes me wish GFP_NOFS would go the way of the
> >>>> dinosaur.
> >>>
> >>> What is the story behind that? One way or another we will get GPF_NOFS
> >>> semantics, either the flag or the scoped NOFS and this limits the
> >>> allocator.
> >>
> >> No extra follow up unfortunately.
> > 
> > For the record, as it was in a separate mail, that GFP_NOFS can
> > sometimes fail because of page fragmentation and limited options for the
> > allocator.
> > 
> >>> Using 64k nodes is problematic and so I'd rather people not use it on 4k
> >>> systems.
> >>
> >> In fact the only problematic part in b-tree operation is exactly the
> >> vmalloc() I'm fixing.
> >>
> >> Other than that I see no obvious problem related to 64K nodesize on 4K
> >> page systems.
> > 
> > Technically if the allocation has a fallback then there's no problem. In
> > practice the 64K contiguous memory needs to get shifted for almost every
> > metadata insertion.
> 
> That is not true, at least not for all cases.
> 
> The 64K buffer is only needed when we got a huge item to split. 
> Meanwhile the most common cases of btrfs items are all fixed sized.

We're talking about the whole node, not individual items and the
variable length types you listed below. Any insertion/deletion in the
middle of a 64k node needs to shift the bytes adjacent to the change. On
fuller nodes it means more data.

> - Csum items
>    This is the exact case we're hitting.
> 
>    The point here is, we can use as large as the whole leaf for a csum
>    item.
>    This makes split much harder, requiring a huge buffer for such split.
> 
>    I think we should introduce an artificial limit on the csum item size.
>    Keep it 16K at max should greatly reduce the need for large buffer.
>    And the cost is pretty minimal, just extra btrfs_items.
>    At most it will be 3 * 25 bytes for 64K node size, I think it's
>    definitely acceptable.
> 
>    By that, we also reduce the buffer requirement for such csum item
>    split.

For filesystems with 64k nodes we can do that, the default case of 16k
is unaffected.

> > COW requires the whole 64K set of pages to be
> > allocated. So there's a lot of dead weight carried around.
> > 
> > It's a tradeoff, 4K nodesize would need taller b-tree for the same
> > amount of raw items. This is worse due to lock contention and concurrent
> > changes. I think the 16K is a reasonable middle ground.
> 
> It's a trade-off for nodesize selection, but for this particular csum 
> item split, I think we should fix it no matter if you want to prevent 
> 64K node size or not.
> 
> Personally I do not want to discourage any valid nodesize/sectorsize 
> combination.

The option exists but is not common and brings an overhead, it should be
a conscious or benchmarked choice rather than a guess that bigger nodes
mean better.

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-11 17:55 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <nka8fLGwxf2NcMXCcuLi4CkGSIKrKG4Sc9IbvTTBG5L6T63w6tImgOKYEmu-pWaoZZFO4Eqg8zDyMggihfvlFRTd1Vubz6E52PVEhU83gn8=@protonmail.com>
2026-09-07 22:15 ` Fwd: btrfs: OOM in split_item() Qu Wenruo
2026-09-08 13:14   ` David Sterba
2026-09-08 21:27     ` Qu Wenruo
2026-09-09  0:44       ` David Sterba
2026-09-09  1:06         ` Qu Wenruo
2026-09-11 17:55           ` David Sterba

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.