* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
[not found] <20250331074541.gK4N_A2Q@linutronix.de>
@ 2025-04-08 16:43 ` Darrick J. Wong
2025-04-08 17:06 ` Luis Chamberlain
0 siblings, 1 reply; 10+ messages in thread
From: Darrick J. Wong @ 2025-04-08 16:43 UTC (permalink / raw)
To: Luis Chamberlain
Cc: Jan Kara, Kefeng Wang, David Bueso, Tso Ted, Ritesh Harjani,
Johannes Weiner, Oliver Sang, Matthew Wilcox, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
Hi Luis,
I'm not sure if this is related, but I'm seeing the same "BUG: sleeping
function called from invalid context at mm/util.c:743" message when
running fstests on XFS. Nothing exciting with fstests here other than
the machine is arm64 with 64k basepages and 4k fsblock size:
MKFS_OPTIONS="-m metadir=1,autofsck=1,uquota,gquota,pquota"
--D
[18182.889554] run fstests generic/457 at 2025-04-07 23:06:25
[18182.973535] spectre-v4 mitigation disabled by command-line option
[18184.849467] XFS (sda3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18184.852941] XFS (sda3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18184.852962] XFS (sda3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18184.858065] XFS (sda3): Mounting V5 Filesystem 13d8c72d-ddac-4052-8d3c-a82c4ce0377d
[18184.900002] XFS (sda3): Ending clean mount
[18184.905990] XFS (sda3): Quotacheck needed: Please wait.
[18184.919801] XFS (sda3): Quotacheck: Done.
[18184.954170] XFS (sda3): Unmounting Filesystem 13d8c72d-ddac-4052-8d3c-a82c4ce0377d
[18186.165572] XFS (dm-4): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18186.165601] XFS (dm-4): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18186.165608] XFS (dm-4): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18186.169589] XFS (dm-4): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18187.121289] XFS (dm-4): Ending clean mount
[18187.131797] XFS (dm-4): Quotacheck needed: Please wait.
[18187.145700] XFS (dm-4): Quotacheck: Done.
[18187.393486] XFS (dm-4): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18190.592061] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18190.592083] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18190.592089] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18190.601815] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18190.744215] XFS (dm-3): Starting recovery (logdev: internal)
[18190.807553] XFS (dm-3): Ending recovery (logdev: internal)
[18190.818708] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18193.786621] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18193.788879] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18193.788882] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18193.790518] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18193.877969] XFS (dm-3): Starting recovery (logdev: internal)
[18193.917688] XFS (dm-3): Ending recovery (logdev: internal)
[18193.945675] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18196.985726] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18196.988868] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18196.988873] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18196.998845] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18197.193740] XFS (dm-3): Starting recovery (logdev: internal)
[18197.254119] XFS (dm-3): Ending recovery (logdev: internal)
[18197.280596] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18200.173003] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18200.176855] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18200.176859] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18200.185721] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18200.370893] XFS (dm-3): Starting recovery (logdev: internal)
[18200.430454] XFS (dm-3): Ending recovery (logdev: internal)
[18200.462036] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18203.311440] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18203.311454] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18203.311464] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18203.324374] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18203.437989] XFS (dm-3): Starting recovery (logdev: internal)
[18203.491993] XFS (dm-3): Ending recovery (logdev: internal)
[18203.517090] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18206.442639] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18206.444851] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18206.444854] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18206.455415] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18206.600488] XFS (dm-3): Starting recovery (logdev: internal)
[18206.642538] XFS (dm-3): Ending recovery (logdev: internal)
[18206.673822] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18209.666477] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18209.678778] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18209.678782] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18209.690805] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18209.859688] XFS (dm-3): Starting recovery (logdev: internal)
[18209.923426] XFS (dm-3): Ending recovery (logdev: internal)
[18209.947181] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18212.920991] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18212.921001] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18212.921012] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18212.925332] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18213.067578] XFS (dm-3): Starting recovery (logdev: internal)
[18213.138633] XFS (dm-3): Ending recovery (logdev: internal)
[18213.161827] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18216.154862] XFS (dm-3): EXPERIMENTAL metadata directory tree feature enabled. Use at your own risk!
[18216.156952] XFS (dm-3): EXPERIMENTAL exchange range feature enabled. Use at your own risk!
[18216.157070] XFS (dm-3): EXPERIMENTAL parent pointer feature enabled. Use at your own risk!
[18216.161145] XFS (dm-3): Mounting V5 Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18216.333087] XFS (dm-3): Starting recovery (logdev: internal)
[18216.389192] XFS (dm-3): Ending recovery (logdev: internal)
[18216.410647] XFS (dm-3): Unmounting Filesystem 6ade490d-15b0-43e5-9f17-db534769c746
[18217.949035] BUG: sleeping function called from invalid context at mm/util.c:743
[18217.949047] in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 35, name: kcompactd0
[18217.949056] preempt_count: 1, expected: 0
[18217.949058] RCU nest depth: 0, expected: 0
[18217.949060] Preemption disabled at:
[18217.949062] [<fffffe0080339c98>] __buffer_migrate_folio+0xb8/0x2d0
[18217.949070] CPU: 0 UID: 0 PID: 35 Comm: kcompactd0 Not tainted 6.15.0-rc1-acha #rc1 PREEMPT 92ec4d9d73adc951fe6bbe0d3f3b75d35d67fded
[18217.949074] Hardware name: QEMU KVM Virtual Machine, BIOS 1.6.6 08/22/2023
[18217.949075] Call trace:
[18217.949076] show_stack+0x20/0x38 (C)
[18217.949080] dump_stack_lvl+0x78/0x90
[18217.949083] dump_stack+0x18/0x28
[18217.949084] __might_resched+0x164/0x1d0
[18217.949086] folio_mc_copy+0x5c/0xa0
[18217.949089] __migrate_folio.constprop.0+0x70/0x1c8
[18217.949092] __buffer_migrate_folio+0x2bc/0x2d0
[18217.949094] buffer_migrate_folio_norefs+0x1c/0x30
[18217.949096] move_to_new_folio+0x70/0x1f0
[18217.949099] migrate_pages_batch+0x9c4/0xf20
[18217.949101] migrate_pages+0xb74/0xde8
[18217.949103] compact_zone+0x9ac/0xff0
[18217.949105] compact_node+0x9c/0x1a0
[18217.949107] kcompactd+0x38c/0x400
[18217.949108] kthread+0x144/0x210
[18217.949110] ret_from_fork+0x10/0x20
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 16:43 ` [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c Darrick J. Wong
@ 2025-04-08 17:06 ` Luis Chamberlain
2025-04-08 17:24 ` Luis Chamberlain
0 siblings, 1 reply; 10+ messages in thread
From: Luis Chamberlain @ 2025-04-08 17:06 UTC (permalink / raw)
To: Darrick J. Wong, David Bueso
Cc: Jan Kara, Kefeng Wang, David Bueso, Tso Ted, Ritesh Harjani,
Johannes Weiner, Oliver Sang, Matthew Wilcox, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 09:43:07AM -0700, Darrick J. Wong wrote:
> Hi Luis,
>
> I'm not sure if this is related, but I'm seeing the same "BUG: sleeping
> function called from invalid context at mm/util.c:743" message when
> running fstests on XFS. Nothing exciting with fstests here other than
> the machine is arm64 with 64k basepages and 4k fsblock size:
How exotic :D
> MKFS_OPTIONS="-m metadir=1,autofsck=1,uquota,gquota,pquota"
>
> --D
>
> [18182.889554] run fstests generic/457 at 2025-04-07 23:06:25
Me and Davidlohr have some fixes brewed up now, before we post we just
want to run one more test for metrics on success rate analysis for folio
migration. Other than that, given the exotic nature of your system we'll
Cc you on preliminary patches, in case you can test to see if it also
fixes your issue. It should given your splat is on the buffer-head side
of things! See _buffer_migrate_folio() reference on the splat. Fun
puzzle for the community is figuring out *why* oh why did a large folio
end up being used on buffer-heads for your use case *without* an LBS
device (logical block size) being present, as I assume you didn't have
one, ie say a nvme or virtio block device with logical block size >
PAGE_SIZE. The area in question would trigger on folio migration *only*
if you are migrating large buffer-head folios. We only create those if
you have an LBS device and are leveragin the block device cache or a
filesystem with buffer-heads with LBS (they don't exist yet other than
the block device cache).
Regardless, the patches we have brewed up should fix this, regardless
of the puzzle described above. We'll cc you for testing before we
post patches to address this.
Luis
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 17:06 ` Luis Chamberlain
@ 2025-04-08 17:24 ` Luis Chamberlain
2025-04-08 17:48 ` Darrick J. Wong
0 siblings, 1 reply; 10+ messages in thread
From: Luis Chamberlain @ 2025-04-08 17:24 UTC (permalink / raw)
To: Darrick J. Wong, David Bueso
Cc: Jan Kara, Kefeng Wang, Tso Ted, Ritesh Harjani, Johannes Weiner,
Oliver Sang, Matthew Wilcox, David Hildenbrand, Alistair Popple,
linux-mm, Christian Brauner, Hannes Reinecke, oe-lkp, lkp,
John Garry, linux-block, ltp, Pankaj Raghav, Daniel Gomez,
Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> Fun
> puzzle for the community is figuring out *why* oh why did a large folio
> end up being used on buffer-heads for your use case *without* an LBS
> device (logical block size) being present, as I assume you didn't have
> one, ie say a nvme or virtio block device with logical block size >
> PAGE_SIZE. The area in question would trigger on folio migration *only*
> if you are migrating large buffer-head folios. We only create those
To be clear, large folios for buffer-heads.
> if
> you have an LBS device and are leveraging the block device cache or a
> filesystem with buffer-heads with LBS (they don't exist yet other than
> the block device cache).
Luis
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 17:24 ` Luis Chamberlain
@ 2025-04-08 17:48 ` Darrick J. Wong
2025-04-08 17:51 ` Matthew Wilcox
2025-04-08 18:06 ` Luis Chamberlain
0 siblings, 2 replies; 10+ messages in thread
From: Darrick J. Wong @ 2025-04-08 17:48 UTC (permalink / raw)
To: Luis Chamberlain
Cc: David Bueso, Jan Kara, Kefeng Wang, Tso Ted, Ritesh Harjani,
Johannes Weiner, Oliver Sang, Matthew Wilcox, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > Fun
> > puzzle for the community is figuring out *why* oh why did a large folio
> > end up being used on buffer-heads for your use case *without* an LBS
> > device (logical block size) being present, as I assume you didn't have
> > one, ie say a nvme or virtio block device with logical block size >
> > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > if you are migrating large buffer-head folios. We only create those
>
> To be clear, large folios for buffer-heads.
> > if
> > you have an LBS device and are leveraging the block device cache or a
> > filesystem with buffer-heads with LBS (they don't exist yet other than
> > the block device cache).
My guess is that udev or something tries to read the disk label in
response to some uevent (mkfs, mount, unmount, etc), which creates a
large folio because min_order > 0, and attaches a buffer head. There's
a separate crash report that I'll cc you on.
--D
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 17:48 ` Darrick J. Wong
@ 2025-04-08 17:51 ` Matthew Wilcox
2025-04-08 18:02 ` Darrick J. Wong
2025-04-08 18:06 ` Luis Chamberlain
1 sibling, 1 reply; 10+ messages in thread
From: Matthew Wilcox @ 2025-04-08 17:51 UTC (permalink / raw)
To: Darrick J. Wong
Cc: Luis Chamberlain, David Bueso, Jan Kara, Kefeng Wang, Tso Ted,
Ritesh Harjani, Johannes Weiner, Oliver Sang, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 10:48:55AM -0700, Darrick J. Wong wrote:
> On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> > On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > > Fun
> > > puzzle for the community is figuring out *why* oh why did a large folio
> > > end up being used on buffer-heads for your use case *without* an LBS
> > > device (logical block size) being present, as I assume you didn't have
> > > one, ie say a nvme or virtio block device with logical block size >
> > > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > > if you are migrating large buffer-head folios. We only create those
> >
> > To be clear, large folios for buffer-heads.
> > > if
> > > you have an LBS device and are leveraging the block device cache or a
> > > filesystem with buffer-heads with LBS (they don't exist yet other than
> > > the block device cache).
>
> My guess is that udev or something tries to read the disk label in
> response to some uevent (mkfs, mount, unmount, etc), which creates a
> large folio because min_order > 0, and attaches a buffer head. There's
> a separate crash report that I'll cc you on.
But you said:
> the machine is arm64 with 64k basepages and 4k fsblock size:
so that shouldn't be using large folios because you should have set the
order to 0. Right? Or did you mis-speak and use a 4K PAGE_SIZE kernel
with a 64k fsblocksize?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 17:51 ` Matthew Wilcox
@ 2025-04-08 18:02 ` Darrick J. Wong
2025-04-08 18:51 ` Matthew Wilcox
0 siblings, 1 reply; 10+ messages in thread
From: Darrick J. Wong @ 2025-04-08 18:02 UTC (permalink / raw)
To: Matthew Wilcox
Cc: Luis Chamberlain, David Bueso, Jan Kara, Kefeng Wang, Tso Ted,
Ritesh Harjani, Johannes Weiner, Oliver Sang, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 06:51:14PM +0100, Matthew Wilcox wrote:
> On Tue, Apr 08, 2025 at 10:48:55AM -0700, Darrick J. Wong wrote:
> > On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> > > On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > > > Fun
> > > > puzzle for the community is figuring out *why* oh why did a large folio
> > > > end up being used on buffer-heads for your use case *without* an LBS
> > > > device (logical block size) being present, as I assume you didn't have
> > > > one, ie say a nvme or virtio block device with logical block size >
> > > > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > > > if you are migrating large buffer-head folios. We only create those
> > >
> > > To be clear, large folios for buffer-heads.
> > > > if
> > > > you have an LBS device and are leveraging the block device cache or a
> > > > filesystem with buffer-heads with LBS (they don't exist yet other than
> > > > the block device cache).
> >
> > My guess is that udev or something tries to read the disk label in
> > response to some uevent (mkfs, mount, unmount, etc), which creates a
> > large folio because min_order > 0, and attaches a buffer head. There's
> > a separate crash report that I'll cc you on.
>
> But you said:
>
> > the machine is arm64 with 64k basepages and 4k fsblock size:
>
> so that shouldn't be using large folios because you should have set the
> order to 0. Right? Or did you mis-speak and use a 4K PAGE_SIZE kernel
> with a 64k fsblocksize?
This particular kernel warning is arm64 with 64k base pages and a 4k
fsblock size, and my suspicion is that udev/libblkid are creating the
buffer heads or something weird like that.
On x64 with 4k base pages, xfs/032 creates a filesystem with 64k sector
size and there's an actual kernel crash resulting from a udev worker:
https://lore.kernel.org/linux-fsdevel/20250408175125.GL6266@frogsfrogsfrogs/T/#u
So I didn't misspeak, I just have two problems. I actually have four
problems, but the others are loop device behavior changes.
--D
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 17:48 ` Darrick J. Wong
2025-04-08 17:51 ` Matthew Wilcox
@ 2025-04-08 18:06 ` Luis Chamberlain
1 sibling, 0 replies; 10+ messages in thread
From: Luis Chamberlain @ 2025-04-08 18:06 UTC (permalink / raw)
To: Darrick J. Wong
Cc: David Bueso, Jan Kara, Kefeng Wang, Tso Ted, Ritesh Harjani,
Johannes Weiner, Oliver Sang, Matthew Wilcox, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 10:48:55AM -0700, Darrick J. Wong wrote:
> On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> > On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > > Fun
> > > puzzle for the community is figuring out *why* oh why did a large folio
> > > end up being used on buffer-heads for your use case *without* an LBS
> > > device (logical block size) being present, as I assume you didn't have
> > > one, ie say a nvme or virtio block device with logical block size >
> > > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > > if you are migrating large buffer-head folios. We only create those
> >
> > To be clear, large folios for buffer-heads.
> > > if
> > > you have an LBS device and are leveraging the block device cache or a
> > > filesystem with buffer-heads with LBS (they don't exist yet other than
> > > the block device cache).
>
> My guess is that udev or something tries to read the disk label in
> response to some uevent (mkfs, mount, unmount, etc), which creates a
> large folio because min_order > 0, and attaches a buffer head. There's
> a separate crash report that I'll cc you on.
OK so as willy pointed out I buy that for x86_64 *iff* we do already
have opportunistic large folio support for the buffer-head read/write
path. But also, I don't think we enable large folios yet on the block
device cache aops unless we have a min order block device... so what
gives?
Luis
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 18:02 ` Darrick J. Wong
@ 2025-04-08 18:51 ` Matthew Wilcox
2025-04-08 19:13 ` Luis Chamberlain
2025-04-08 19:13 ` Luis Chamberlain
0 siblings, 2 replies; 10+ messages in thread
From: Matthew Wilcox @ 2025-04-08 18:51 UTC (permalink / raw)
To: Darrick J. Wong
Cc: Luis Chamberlain, David Bueso, Jan Kara, Kefeng Wang, Tso Ted,
Ritesh Harjani, Johannes Weiner, Oliver Sang, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 11:02:40AM -0700, Darrick J. Wong wrote:
> On Tue, Apr 08, 2025 at 06:51:14PM +0100, Matthew Wilcox wrote:
> > On Tue, Apr 08, 2025 at 10:48:55AM -0700, Darrick J. Wong wrote:
> > > On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> > > > On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > > > > Fun
> > > > > puzzle for the community is figuring out *why* oh why did a large folio
> > > > > end up being used on buffer-heads for your use case *without* an LBS
> > > > > device (logical block size) being present, as I assume you didn't have
> > > > > one, ie say a nvme or virtio block device with logical block size >
> > > > > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > > > > if you are migrating large buffer-head folios. We only create those
> > > >
> > > > To be clear, large folios for buffer-heads.
> > > > > if
> > > > > you have an LBS device and are leveraging the block device cache or a
> > > > > filesystem with buffer-heads with LBS (they don't exist yet other than
> > > > > the block device cache).
> > >
> > > My guess is that udev or something tries to read the disk label in
> > > response to some uevent (mkfs, mount, unmount, etc), which creates a
> > > large folio because min_order > 0, and attaches a buffer head. There's
> > > a separate crash report that I'll cc you on.
> >
> > But you said:
> >
> > > the machine is arm64 with 64k basepages and 4k fsblock size:
> >
> > so that shouldn't be using large folios because you should have set the
> > order to 0. Right? Or did you mis-speak and use a 4K PAGE_SIZE kernel
> > with a 64k fsblocksize?
>
> This particular kernel warning is arm64 with 64k base pages and a 4k
> fsblock size, and my suspicion is that udev/libblkid are creating the
> buffer heads or something weird like that.
>
> On x64 with 4k base pages, xfs/032 creates a filesystem with 64k sector
> size and there's an actual kernel crash resulting from a udev worker:
> https://lore.kernel.org/linux-fsdevel/20250408175125.GL6266@frogsfrogsfrogs/T/#u
>
> So I didn't misspeak, I just have two problems. I actually have four
> problems, but the others are loop device behavior changes.
Right, but this warning only triggers for large folios. So somehow
we've got a multi-page folio in the bdev's page cache.
Ah. I see.
block/bdev.c: mapping_set_folio_min_order(BD_INODE(bdev)->i_mapping,
so we're telling the bdev that it can go up to MAX_PAGECACHE_ORDER.
And then we call readahead, which will happily put order-2 folios
in the pagecache because of my bug that we've never bothered fixing.
We should probably fix that now, but as a temporary measure if
you'd like to put:
mapping_set_folio_order_range(BD_INODE(bdev)->i_mapping, min, min)
instead of the mapping_set_folio_min_order(), that would make the bug
no longer appear for you.
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 18:51 ` Matthew Wilcox
@ 2025-04-08 19:13 ` Luis Chamberlain
2025-04-08 19:13 ` Luis Chamberlain
1 sibling, 0 replies; 10+ messages in thread
From: Luis Chamberlain @ 2025-04-08 19:13 UTC (permalink / raw)
To: Matthew Wilcox
Cc: Darrick J. Wong, David Bueso, Jan Kara, Kefeng Wang, Tso Ted,
Ritesh Harjani, Johannes Weiner, Oliver Sang, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 07:51:03PM +0100, Matthew Wilcox wrote:
> On Tue, Apr 08, 2025 at 11:02:40AM -0700, Darrick J. Wong wrote:
> > On Tue, Apr 08, 2025 at 06:51:14PM +0100, Matthew Wilcox wrote:
> > > On Tue, Apr 08, 2025 at 10:48:55AM -0700, Darrick J. Wong wrote:
> > > > On Tue, Apr 08, 2025 at 10:24:40AM -0700, Luis Chamberlain wrote:
> > > > > On Tue, Apr 8, 2025 at 10:06 AM Luis Chamberlain <mcgrof@kernel.org> wrote:
> > > > > > Fun
> > > > > > puzzle for the community is figuring out *why* oh why did a large folio
> > > > > > end up being used on buffer-heads for your use case *without* an LBS
> > > > > > device (logical block size) being present, as I assume you didn't have
> > > > > > one, ie say a nvme or virtio block device with logical block size >
> > > > > > PAGE_SIZE. The area in question would trigger on folio migration *only*
> > > > > > if you are migrating large buffer-head folios. We only create those
> > > > >
> > > > > To be clear, large folios for buffer-heads.
> > > > > > if
> > > > > > you have an LBS device and are leveraging the block device cache or a
> > > > > > filesystem with buffer-heads with LBS (they don't exist yet other than
> > > > > > the block device cache).
> > > >
> > > > My guess is that udev or something tries to read the disk label in
> > > > response to some uevent (mkfs, mount, unmount, etc), which creates a
> > > > large folio because min_order > 0, and attaches a buffer head. There's
> > > > a separate crash report that I'll cc you on.
> > >
> > > But you said:
> > >
> > > > the machine is arm64 with 64k basepages and 4k fsblock size:
> > >
> > > so that shouldn't be using large folios because you should have set the
> > > order to 0. Right? Or did you mis-speak and use a 4K PAGE_SIZE kernel
> > > with a 64k fsblocksize?
> >
> > This particular kernel warning is arm64 with 64k base pages and a 4k
> > fsblock size, and my suspicion is that udev/libblkid are creating the
> > buffer heads or something weird like that.
> >
> > On x64 with 4k base pages, xfs/032 creates a filesystem with 64k sector
> > size and there's an actual kernel crash resulting from a udev worker:
> > https://lore.kernel.org/linux-fsdevel/20250408175125.GL6266@frogsfrogsfrogs/T/#u
> >
> > So I didn't misspeak, I just have two problems. I actually have four
> > problems, but the others are loop device behavior changes.
>
> Right, but this warning only triggers for large folios. So somehow
> we've got a multi-page folio in the bdev's page cache.
>
> Ah. I see.
>
> block/bdev.c: mapping_set_folio_min_order(BD_INODE(bdev)->i_mapping,
>
> so we're telling the bdev that it can go up to MAX_PAGECACHE_ORDER.
Ah yes silly me that would explain the large folios without LBS devices.
> And then we call readahead, which will happily put order-2 folios
> in the pagecache because of my bug that we've never bothered fixing.
>
> We should probably fix that now, but as a temporary measure if
> you'd like to put:
>
> mapping_set_folio_order_range(BD_INODE(bdev)->i_mapping, min, min)
>
> instead of the mapping_set_folio_min_order(), that would make the bug
> no longer appear for you.
Agreed.
Luis
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c
2025-04-08 18:51 ` Matthew Wilcox
2025-04-08 19:13 ` Luis Chamberlain
@ 2025-04-08 19:13 ` Luis Chamberlain
1 sibling, 0 replies; 10+ messages in thread
From: Luis Chamberlain @ 2025-04-08 19:13 UTC (permalink / raw)
To: Matthew Wilcox
Cc: Darrick J. Wong, David Bueso, Jan Kara, Kefeng Wang, Tso Ted,
Ritesh Harjani, Johannes Weiner, Oliver Sang, David Hildenbrand,
Alistair Popple, linux-mm, Christian Brauner, Hannes Reinecke,
oe-lkp, lkp, John Garry, linux-block, ltp, Pankaj Raghav,
Daniel Gomez, Dave Chinner, gost.dev, linux-fsdevel
On Tue, Apr 08, 2025 at 07:51:03PM +0100, Matthew Wilcox wrote:
> And then we call readahead, which will happily put order-2 folios
> in the pagecache because of my bug that we've never bothered fixing.
What was that BTW?
Luis
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2025-04-08 19:13 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20250331074541.gK4N_A2Q@linutronix.de>
2025-04-08 16:43 ` [linux-next:master] [block/bdev] 3c20917120: BUG:sleeping_function_called_from_invalid_context_at_mm/util.c Darrick J. Wong
2025-04-08 17:06 ` Luis Chamberlain
2025-04-08 17:24 ` Luis Chamberlain
2025-04-08 17:48 ` Darrick J. Wong
2025-04-08 17:51 ` Matthew Wilcox
2025-04-08 18:02 ` Darrick J. Wong
2025-04-08 18:51 ` Matthew Wilcox
2025-04-08 19:13 ` Luis Chamberlain
2025-04-08 19:13 ` Luis Chamberlain
2025-04-08 18:06 ` Luis Chamberlain
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox