Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: David Sterba <dsterba@suse.cz>
To: Qu Wenruo <wqu@suse.com>
Cc: dsterba@suse.cz, linux-btrfs <linux-btrfs@vger.kernel.org>
Subject: Re: Fwd: btrfs: OOM in split_item().
Date: Fri, 11 Sep 2026 19:55:26 +0200	[thread overview]
Message-ID: <20260911175526.GC54722@twin.jikos.cz> (raw)
In-Reply-To: <f6eed62a-c9f6-4f0f-848e-e99445f46538@suse.com>

On Wed, Sep 09, 2026 at 10:36:10AM +0930, Qu Wenruo wrote:
> 在 2026/9/9 10:14, David Sterba 写道:
> > On Wed, Sep 09, 2026 at 06:57:59AM +0930, Qu Wenruo wrote:
> >> 在 2026/9/8 22:44, David Sterba 写道:
> >>> On Tue, Sep 08, 2026 at 07:45:58AM +0930, Qu Wenruo wrote:
> >>>> Forwarded for archive purposes, as patchcheck requires a link: tag
> >>>> following reported-by: tag.
> >>>>
> >>>> Meanwhile the original report is only a private mail to me.
> >>>>
> >>>> -------- 转发的消息 --------
> >>>> 主题: 	btrfs: OOM in split_item().
> >>>> 日期: 	Mon, 07 Sep 2026 20:10:56 +0000
> >>>> 发件人: 	xavierbachmeyer182 <xavierbachmeyer182@protonmail.com>
> >>>> 收件人: 	wqu@suse.com <wqu@suse.com>
> >>>>
> >>>>
> >>>>
> >>>> Hello.
> >>>>
> >>>> Attached is a kernel log entry showing a strange out of memory issue
> >>>> that rarely seems to happen.
> >>>> It's kind of annoying and makes me wish GFP_NOFS would go the way of the
> >>>> dinosaur.
> >>>
> >>> What is the story behind that? One way or another we will get GPF_NOFS
> >>> semantics, either the flag or the scoped NOFS and this limits the
> >>> allocator.
> >>
> >> No extra follow up unfortunately.
> > 
> > For the record, as it was in a separate mail, that GFP_NOFS can
> > sometimes fail because of page fragmentation and limited options for the
> > allocator.
> > 
> >>> Using 64k nodes is problematic and so I'd rather people not use it on 4k
> >>> systems.
> >>
> >> In fact the only problematic part in b-tree operation is exactly the
> >> vmalloc() I'm fixing.
> >>
> >> Other than that I see no obvious problem related to 64K nodesize on 4K
> >> page systems.
> > 
> > Technically if the allocation has a fallback then there's no problem. In
> > practice the 64K contiguous memory needs to get shifted for almost every
> > metadata insertion.
> 
> That is not true, at least not for all cases.
> 
> The 64K buffer is only needed when we got a huge item to split. 
> Meanwhile the most common cases of btrfs items are all fixed sized.

We're talking about the whole node, not individual items and the
variable length types you listed below. Any insertion/deletion in the
middle of a 64k node needs to shift the bytes adjacent to the change. On
fuller nodes it means more data.

> - Csum items
>    This is the exact case we're hitting.
> 
>    The point here is, we can use as large as the whole leaf for a csum
>    item.
>    This makes split much harder, requiring a huge buffer for such split.
> 
>    I think we should introduce an artificial limit on the csum item size.
>    Keep it 16K at max should greatly reduce the need for large buffer.
>    And the cost is pretty minimal, just extra btrfs_items.
>    At most it will be 3 * 25 bytes for 64K node size, I think it's
>    definitely acceptable.
> 
>    By that, we also reduce the buffer requirement for such csum item
>    split.

For filesystems with 64k nodes we can do that, the default case of 16k
is unaffected.

> > COW requires the whole 64K set of pages to be
> > allocated. So there's a lot of dead weight carried around.
> > 
> > It's a tradeoff, 4K nodesize would need taller b-tree for the same
> > amount of raw items. This is worse due to lock contention and concurrent
> > changes. I think the 16K is a reasonable middle ground.
> 
> It's a trade-off for nodesize selection, but for this particular csum 
> item split, I think we should fix it no matter if you want to prevent 
> 64K node size or not.
> 
> Personally I do not want to discourage any valid nodesize/sectorsize 
> combination.

The option exists but is not common and brings an overhead, it should be
a conscious or benchmarked choice rather than a guess that bigger nodes
mean better.

      reply	other threads:[~2026-09-11 17:55 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <nka8fLGwxf2NcMXCcuLi4CkGSIKrKG4Sc9IbvTTBG5L6T63w6tImgOKYEmu-pWaoZZFO4Eqg8zDyMggihfvlFRTd1Vubz6E52PVEhU83gn8=@protonmail.com>
2026-09-07 22:15 ` Fwd: btrfs: OOM in split_item() Qu Wenruo
2026-09-08 13:14   ` David Sterba
2026-09-08 21:27     ` Qu Wenruo
2026-09-09  0:44       ` David Sterba
2026-09-09  1:06         ` Qu Wenruo
2026-09-11 17:55           ` David Sterba [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260911175526.GC54722@twin.jikos.cz \
    --to=dsterba@suse.cz \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=wqu@suse.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox