* [RFC] madvise: best-effort deferred writeback for shared file mappings
@ 2026-09-28 7:07 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
0 siblings, 1 reply; 20+ messages in thread
From: 天狼 @ 2026-09-28 7:07 UTC (permalink / raw)
To: linux-mm; +Cc: linux-fsdevel
Hello,
I would like to request feedback on a mapping-scoped, best-effort hint
to defer background writeback while an application is still constructing
data in a writable MAP_SHARED mapping of a regular disk-backed file.
This is a feature request, not a patch submission or a benchmark report.
Motivation
----------
Consider an application building a large data structure, for example
10 GiB, through multiple passes that repeatedly modify the same pages.
After construction, it persists the file and publishes it for other
processes to mmap and read.
Building directly in MAP_SHARED keeps the data in the file's page cache,
which also serves the consumers. However, background writeback can write
intermediate contents that subsequent construction passes dirty again.
The application knows when construction ends, but cannot express that
knowledge as a mapping-local writeback scheduling hint.
Building in private memory and then doing buffered writes adds a bulk
copy into the file's page cache. Direct I/O does not itself populate that
cache for subsequent file-mapped consumers. A memfd can share the build
buffer, but requires a separate object for persistence and changes the
consumer interface. Global dirty/writeback tuning is broader than this
particular mapping and construction phase.
Proposed interface and behavior
-------------------------------
Tentatively, a pair of madvise operations could express this:
madvise(addr, len, MADV_DEFER_WRITEBACK);
/* Construct and repeatedly update the shared file mapping. */
madvise(addr, len, MADV_ALLOW_WRITEBACK);
/* Use existing synchronization calls if durability is required. */
The names are placeholders for discussion, not existing constants.
The requested behavior is deliberately advisory:
* Prefer to defer ordinary background writeback for the corresponding
file pages while the hint is active, including pages faulted in later
within the advised range.
* The kernel may override or ignore the hint for memory pressure, dirty
limits, fairness, filesystem constraints, or forward progress. It must
not reserve or pin memory, exempt pages from dirty accounting, or
promise that intermediate data will never reach storage.
* Explicit synchronization, including msync(MS_SYNC), fsync, fdatasync,
syncfs, and sync, retains its existing semantics and must not wait for
the application to revoke the hint. In-flight I/O need not be canceled.
* Shared-memory visibility, data contents, and existing durability
guarantees remain unchanged. This is not an atomic publication,
transaction, or crash-consistency mechanism.
* Revoking the hint restores normal scheduling; it does not itself
require an immediate writeback or imply durability.
Automatic lifetime cleanup is an important part of the request:
* munmap removes this mapping's hint for the unmapped portion.
* Mapping teardown on process exit, including abnormal termination,
automatically removes its hints. No application cleanup handler is
required, and no persistent per-file setting should remain behind.
* Independently advised mappings may refer to the same file pages.
Removing one mapping must not remove another mapping's own hint.
Once no applicable hints remain, the pages have normal policy.
An initial scope could be ordinary page-cache-backed regular files with
writable MAP_SHARED mappings. Unsupported mappings could be rejected.
Details such as fork inheritance, VMA splits/merges, mremap, and interaction
with unadvised aliases would need explicit definitions; the lifetime rule
should apply consistently to every mapping that actually owns a hint.
Implementation and evaluation questions
---------------------------------------
Although the proposed API is mapping-scoped, writeback acts on file-cache
folios shared by mappings. An implementation would need an inexpensive
way to connect those scopes and retire the hint without leaving stale
writeback state. This request does not assume that adding a VMA flag
alone is sufficient, or mandate a particular accounting mechanism.
Useful evaluation would compare repeated-update construction with and
without the hint, followed by explicit synchronization and a consumer
mmap/read pass. Metrics should include device bytes written, total time
through final synchronization, dirty throttling, peak memory, and consumer
read I/O. Memory-pressure cases, competing writers, partial unmap, and
process termination should also be covered. There are no measured results
attached to this request; the expected benefit is a hypothesis to test.
Is there an existing per-mapping or per-range mechanism that already
provides this behavior? If not, would a best-effort madvise extension be
an appropriate interface, or would an fd/range advisory API fit the
writeback subsystem better while retaining automatic lifetime cleanup?
Pointers to prior discussions and feedback on whether this is worth
prototyping would be appreciated.
This proposal was developed with AI assistance from an application-use
case discussion; no implementation or performance results are claimed.
Regards,
Peng Lei
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-28 7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
@ 2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09 ` 天狼
0 siblings, 1 reply; 20+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-28 11:29 UTC (permalink / raw)
To: 天狼, linux-mm
Cc: linux-fsdevel, Lorenzo Stoakes (Arm), Liam R. Howlett,
Vlastimil Babka, Jann Horn
On 9/28/26 09:07, 天狼 wrote:
> Hello,
Hi,
>
> I would like to request feedback on a mapping-scoped, best-effort hint
> to defer background writeback while an application is still constructing
> data in a writable MAP_SHARED mapping of a regular disk-backed file.
> This is a feature request, not a patch submission or a benchmark report.
>
> Motivation
> ----------
>
> Consider an application building a large data structure, for example
> 10 GiB, through multiple passes that repeatedly modify the same pages.
> After construction, it persists the file and publishes it for other
> processes to mmap and read.
Is it an option to just construct the data using anonymous memory and then
write()'in it once ready in one operation?
Why exactly are you using writable MAP_SHARED mappings?
>
> Building directly in MAP_SHARED keeps the data in the file's page cache,
> which also serves the consumers. However, background writeback can write
> intermediate contents that subsequent construction passes dirty again.
> The application knows when construction ends, but cannot express that
> knowledge as a mapping-local writeback scheduling hint.
>
> Building in private memory and then doing buffered writes adds a bulk
> copy into the file's page cache. Direct I/O does not itself populate that
> cache for subsequent file-mapped consumers. A memfd can share the build
> buffer, but requires a separate object for persistence and changes the
> consumer interface. Global dirty/writeback tuning is broader than this
> particular mapping and construction phase.
>
> Proposed interface and behavior
> -------------------------------
>
> Tentatively, a pair of madvise operations could express this:
>
> madvise(addr, len, MADV_DEFER_WRITEBACK);
> /* Construct and repeatedly update the shared file mapping. */
> madvise(addr, len, MADV_ALLOW_WRITEBACK);
> /* Use existing synchronization calls if durability is required. */
>
> The names are placeholders for discussion, not existing constants.
I don't think this behavior should be part of madvise. Way too specific.
--
Cheers,
David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-28 11:29 ` David Hildenbrand (Arm)
@ 2026-09-28 13:09 ` 天狼
2026-09-28 19:12 ` David Hildenbrand (Arm)
0 siblings, 1 reply; 20+ messages in thread
From: 天狼 @ 2026-09-28 13:09 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: linux-mm, linux-fsdevel, Lorenzo Stoakes (Arm), Liam R. Howlett,
Vlastimil Babka, Jann Horn
Hi David,
The use case is a concurrent, lock-free, relocatable data structure that
uses offsets rather than absolute pointers. It implements copy-on-write
at the data-structure level and is designed to remain recoverable after
a process crash. We therefore build it directly in a file-backed
MAP_SHARED mapping, so its state is not lost when the constructing
process exits unexpectedly.
Building it in anonymous memory and calling write() only after
construction would lose that state if the process crashed before the
write. It would also require copying the entire structure into the
file's page cache.
During construction, writes are scattered throughout the mapping,
dirtying pages rapidly and repeatedly redirtying pages that may already
have been written back. The goal is to reduce writeback of these
intermediate states while retaining the same page-cache pages for
subsequent readers.
The hint would remain strictly best-effort: memory pressure and explicit
synchronization could override it. By "process crash," I mean termination
of the process while the kernel remains running; recovery after a system
crash or power loss is a separate durability concern.
Regarding the interface, madvise was only a suggestion. A file-level
advisory interface would also work for this use case: deferring background
writeback for the entire file, rather than individual mapped ranges, is
acceptable. The hint would remain best-effort, with memory pressure and
explicit synchronization taking precedence. It should also be cleared
automatically when the owning mapping or file descriptor is released,
including on process exit. Would that be a better fit for the writeback
subsystem?
Thanks,
Peng
David Hildenbrand (Arm) <david@kernel.org> 于2026年9月28日周一 19:29写道:
>
> On 9/28/26 09:07, 天狼 wrote:
> > Hello,
>
> Hi,
>
> >
> > I would like to request feedback on a mapping-scoped, best-effort hint
> > to defer background writeback while an application is still constructing
> > data in a writable MAP_SHARED mapping of a regular disk-backed file.
> > This is a feature request, not a patch submission or a benchmark report.
> >
> > Motivation
> > ----------
> >
> > Consider an application building a large data structure, for example
> > 10 GiB, through multiple passes that repeatedly modify the same pages.
> > After construction, it persists the file and publishes it for other
> > processes to mmap and read.
>
> Is it an option to just construct the data using anonymous memory and then
> write()'in it once ready in one operation?
>
> Why exactly are you using writable MAP_SHARED mappings?
>
> >
> > Building directly in MAP_SHARED keeps the data in the file's page cache,
> > which also serves the consumers. However, background writeback can write
> > intermediate contents that subsequent construction passes dirty again.
> > The application knows when construction ends, but cannot express that
> > knowledge as a mapping-local writeback scheduling hint.
> >
> > Building in private memory and then doing buffered writes adds a bulk
> > copy into the file's page cache. Direct I/O does not itself populate that
> > cache for subsequent file-mapped consumers. A memfd can share the build
> > buffer, but requires a separate object for persistence and changes the
> > consumer interface. Global dirty/writeback tuning is broader than this
> > particular mapping and construction phase.
> >
> > Proposed interface and behavior
> > -------------------------------
> >
> > Tentatively, a pair of madvise operations could express this:
> >
> > madvise(addr, len, MADV_DEFER_WRITEBACK);
> > /* Construct and repeatedly update the shared file mapping. */
> > madvise(addr, len, MADV_ALLOW_WRITEBACK);
> > /* Use existing synchronization calls if durability is required. */
> >
> > The names are placeholders for discussion, not existing constants.
>
> I don't think this behavior should be part of madvise. Way too specific.
>
> --
> Cheers,
>
> David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-28 13:09 ` 天狼
@ 2026-09-28 19:12 ` David Hildenbrand (Arm)
2026-09-29 4:06 ` 天狼
0 siblings, 1 reply; 20+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-28 19:12 UTC (permalink / raw)
To: 天狼
Cc: linux-mm, linux-fsdevel, Lorenzo Stoakes (Arm), Liam R. Howlett,
Vlastimil Babka, Jann Horn
On 9/28/26 15:09, 天狼 wrote:
> Hi David,
Hi,
>
> The use case is a concurrent, lock-free, relocatable data structure that
> uses offsets rather than absolute pointers. It implements copy-on-write
> at the data-structure level and is designed to remain recoverable after
> a process crash. We therefore build it directly in a file-backed
> MAP_SHARED mapping, so its state is not lost when the constructing
> process exits unexpectedly.
What is supposed to happen if the process crashes when updating the MAP_SHARED
region halfway through?
It would be helpful if the use case + data structure would be explained in a
bit more detail.
>
> Building it in anonymous memory and calling write() only after
> construction would lose that state if the process crashed before the
> write. It would also require copying the entire structure into the
> file's page cache.
Yes. Unless to would try to bypass the page cache of course (if possible for
your use case).
>
> During construction, writes are scattered throughout the mapping,
> dirtying pages rapidly and repeatedly redirtying pages that may already
> have been written back. The goal is to reduce writeback of these
> intermediate states while retaining the same page-cache pages for
> subsequent readers.
What exactly is the poblem with writeback here? Unnecessary I/O? Is writeback
the problem or actual reclaim after writeback?
(I recall that in a fuse server you can in theory delay the writeback request.
So maybe you could get something going by serving your file through fuse and
enlightening the fuse server about it. Just a random idea.)
>
> The hint would remain strictly best-effort: memory pressure and explicit
> synchronization could override it. By "process crash," I mean termination
> of the process while the kernel remains running; recovery after a system
> crash or power loss is a separate durability concern.
>
> Regarding the interface, madvise was only a suggestion. A file-level
> advisory interface would also work for this use case: deferring background
> writeback for the entire file, rather than individual mapped ranges, is
> acceptable.
So fadvise would be an option. However, this "defer mode" is really odd. It
sounds more like you would want to have a custom policy there, instead of
hardcoding something that really not a lot might want. Hmmm
--
Cheers,
David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-28 19:12 ` David Hildenbrand (Arm)
@ 2026-09-29 4:06 ` 天狼
2026-09-29 6:59 ` David Hildenbrand (Arm)
0 siblings, 1 reply; 20+ messages in thread
From: 天狼 @ 2026-09-29 4:06 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: linux-mm, linux-fsdevel, Lorenzo Stoakes (Arm), Liam R. Howlett,
Vlastimil Babka, Jann Horn
Hi David,
> What is supposed to happen if the process crashes when updating the
> MAP_SHARED region halfway through?
The data structure survives an interrupted update through its
application-level copy-on-write and lock-free concurrency design. A
file-backed MAP_SHARED mapping preserves the latest userspace writes after
the process dies, even if they have not yet reached disk, provided the
kernel remains running.
This does not depend on preventing writeback. Deferring writeback is purely
a performance optimization, not a correctness requirement.
> It would be helpful if the use case + data structure would be explained
> in a bit more detail.
The data structure originally used uint32_t offsets instead of pointers to
reduce pointer overhead. On that basis, we implemented lock-free concurrent
reads and writes, obtaining crash safety for free at the same time. We use
a file-backed MAP_SHARED mapping to turn that capability into an actual
feature that survives process crashes.
A concurrent skiplist using offsets as pointers is a simple example of such
a data structure.
During construction, writes are scattered throughout a large mapping,
dirtying pages rapidly and repeatedly modifying pages.
> Unless you would try to bypass the page cache of course (if possible for
> your use case).
The page cache is useful here: it preserves the latest shared-mapping
writes across a process crash, and subsequent consumer processes can map
the file and read the same pages. Building directly in MAP_SHARED gives us
both properties without a separate copy of the completed structure.
> What exactly is the problem with writeback here? Unnecessary I/O? Is
> writeback the problem or actual reclaim after writeback?
Unnecessary I/O. Reclaim after writeback is not the problem I am trying to
address.
During construction, background writeback can write intermediate contents
that the application will soon overwrite. Those repeated writes waste disk
bandwidth. Delaying writeback would avoid that waste where possible; it is
not required for correctness.
> In a fuse server you can in theory delay the writeback request.
Thanks for the suggestion. That sounds like a possible way to experiment
with delayed writeback. The feature I am requesting would make this
advisory behavior available to applications using ordinary file-backed
shared mappings.
> So fadvise would be an option. However, this "defer mode" is really odd.
> It sounds more like you would want to have a custom policy there [...]
madvise is not the only option, deferring writeback for the entire file
would also be acceptable.
The application only wants to indicate that it is still modifying the data
and would prefer background writeback to happen later. The kernel would
retain control over the actual timing. Memory pressure could override the
hint, and explicit synchronization would keep its normal semantics.
The hint should also expire automatically when the owning mapping or handle
is released, including on process exit.
Thanks,
Peng
David Hildenbrand (Arm) <david@kernel.org> 于2026年9月29日周二 03:12写道:
>
> On 9/28/26 15:09, 天狼 wrote:
> > Hi David,
>
> Hi,
>
> >
> > The use case is a concurrent, lock-free, relocatable data structure that
> > uses offsets rather than absolute pointers. It implements copy-on-write
> > at the data-structure level and is designed to remain recoverable after
> > a process crash. We therefore build it directly in a file-backed
> > MAP_SHARED mapping, so its state is not lost when the constructing
> > process exits unexpectedly.
>
> What is supposed to happen if the process crashes when updating the MAP_SHARED
> region halfway through?
>
> It would be helpful if the use case + data structure would be explained in a
> bit more detail.
>
> >
> > Building it in anonymous memory and calling write() only after
> > construction would lose that state if the process crashed before the
> > write. It would also require copying the entire structure into the
> > file's page cache.
>
> Yes. Unless to would try to bypass the page cache of course (if possible for
> your use case).
>
> >
> > During construction, writes are scattered throughout the mapping,
> > dirtying pages rapidly and repeatedly redirtying pages that may already
> > have been written back. The goal is to reduce writeback of these
> > intermediate states while retaining the same page-cache pages for
> > subsequent readers.
>
> What exactly is the poblem with writeback here? Unnecessary I/O? Is writeback
> the problem or actual reclaim after writeback?
>
> (I recall that in a fuse server you can in theory delay the writeback request.
> So maybe you could get something going by serving your file through fuse and
> enlightening the fuse server about it. Just a random idea.)
>
> >
> > The hint would remain strictly best-effort: memory pressure and explicit
> > synchronization could override it. By "process crash," I mean termination
> > of the process while the kernel remains running; recovery after a system
> > crash or power loss is a separate durability concern.
> >
> > Regarding the interface, madvise was only a suggestion. A file-level
> > advisory interface would also work for this use case: deferring background
> > writeback for the entire file, rather than individual mapped ranges, is
> > acceptable.
>
> So fadvise would be an option. However, this "defer mode" is really odd. It
> sounds more like you would want to have a custom policy there, instead of
> hardcoding something that really not a lot might want. Hmmm
>
> --
> Cheers,
>
> David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-29 4:06 ` 天狼
@ 2026-09-29 6:59 ` David Hildenbrand (Arm)
2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
0 siblings, 1 reply; 20+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-29 6:59 UTC (permalink / raw)
To: 天狼
Cc: linux-mm, linux-fsdevel, Lorenzo Stoakes (Arm), Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On 9/29/26 06:06, 天狼 wrote:
> Hi David,
>
Hi,
>> What is supposed to happen if the process crashes when updating the
>> MAP_SHARED region halfway through?
>
> The data structure survives an interrupted update through its
> application-level copy-on-write and lock-free concurrency design. A
> file-backed MAP_SHARED mapping preserves the latest userspace writes after
> the process dies, even if they have not yet reached disk, provided the
> kernel remains running.
>
> This does not depend on preventing writeback. Deferring writeback is purely
> a performance optimization, not a correctness requirement.
>
>> It would be helpful if the use case + data structure would be explained
>> in a bit more detail.
>
> The data structure originally used uint32_t offsets instead of pointers to
> reduce pointer overhead. On that basis, we implemented lock-free concurrent
> reads and writes, obtaining crash safety for free at the same time. We use
> a file-backed MAP_SHARED mapping to turn that capability into an actual
> feature that survives process crashes.
Ok, thanks. Just be sure: you really need an actual file on disk. Shared memory
/ shmem / memfd is not sufficient?
>> What exactly is the problem with writeback here? Unnecessary I/O? Is
>> writeback the problem or actual reclaim after writeback?
>
> Unnecessary I/O. Reclaim after writeback is not the problem I am trying to
> address.
Ok.
>
> During construction, background writeback can write intermediate contents
> that the application will soon overwrite. Those repeated writes waste disk
> bandwidth. Delaying writeback would avoid that waste where possible; it is
> not required for correctness.
Understood.
>
>> In a fuse server you can in theory delay the writeback request.
>
> Thanks for the suggestion. That sounds like a possible way to experiment
> with delayed writeback. The feature I am requesting would make this
> advisory behavior available to applications using ordinary file-backed
> shared mappings.
>
>> So fadvise would be an option. However, this "defer mode" is really odd.
>> It sounds more like you would want to have a custom policy there [...]
>
> madvise is not the only option, deferring writeback for the entire file
> would also be acceptable.
>
> The application only wants to indicate that it is still modifying the data
> and would prefer background writeback to happen later. The kernel would
> retain control over the actual timing. Memory pressure could override the
> hint, and explicit synchronization would keep its normal semantics.
>
> The hint should also expire automatically when the owning mapping or handle
> is released, including on process exit.
Ok, so while you are updating the large mmap'ed file concurrently, you don't
want writeback to go crazy, because you know that you will modify the memory
immediately anyway.
Let me CC some more people.
--
Cheers,
David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-29 6:59 ` David Hildenbrand (Arm)
@ 2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
2026-09-30 10:27 ` 天狼
0 siblings, 1 reply; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-29 8:59 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: 天狼, linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Tue, Sep 29, 2026 at 08:59:57AM +0200, David Hildenbrand (Arm) wrote:
> On 9/29/26 06:06, 天狼 wrote:
> > Hi David,
> >
>
> Hi,
>
> >> What is supposed to happen if the process crashes when updating the
> >> MAP_SHARED region halfway through?
> >
> > The data structure survives an interrupted update through its
> > application-level copy-on-write and lock-free concurrency design. A
> > file-backed MAP_SHARED mapping preserves the latest userspace writes after
> > the process dies, even if they have not yet reached disk, provided the
> > kernel remains running.
> >
> > This does not depend on preventing writeback. Deferring writeback is purely
> > a performance optimization, not a correctness requirement.
> >
> >> It would be helpful if the use case + data structure would be explained
> >> in a bit more detail.
> >
> > The data structure originally used uint32_t offsets instead of pointers to
> > reduce pointer overhead. On that basis, we implemented lock-free concurrent
> > reads and writes, obtaining crash safety for free at the same time. We use
> > a file-backed MAP_SHARED mapping to turn that capability into an actual
> > feature that survives process crashes.
>
> Ok, thanks. Just be sure: you really need an actual file on disk. Shared memory
> / shmem / memfd is not sufficient?
>
I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not sufficient
either?
> >> What exactly is the problem with writeback here? Unnecessary I/O? Is
> >> writeback the problem or actual reclaim after writeback?
> >
> > Unnecessary I/O. Reclaim after writeback is not the problem I am trying to
> > address.
>
> Ok.
I mean if you're mmap'ing it you're literally mapping the page cache folio and
dirtying that folio.
So the kernel really does have to do writeback and it might be quite problematic
trying to prevent the writeback algorithm from doing its work on a specific
range.
I think the shape of a viable solution really is either MAP_PRIVATE-mapping it
or using some anon shmem/memfd as David suggests.
I think it'd be problematic to effectively corrupt the page cache by saying 'hey
this is dirty but just clear dirty state and pretend this is what the disk has'
or something.
>
> >
> > During construction, background writeback can write intermediate contents
> > that the application will soon overwrite. Those repeated writes waste disk
> > bandwidth. Delaying writeback would avoid that waste where possible; it is
> > not required for correctness.
>
> Understood.
Yeah again delaying writeback could interfere with the writeback balancing logic
which tries to keep dirty page levels sane and is fair and balanced so processes
that writeback a lot get delayed in doing so under heavy dirtying.
I think anything like this would interfere with that.
>
> >
> >> In a fuse server you can in theory delay the writeback request.
> >
> > Thanks for the suggestion. That sounds like a possible way to experiment
> > with delayed writeback. The feature I am requesting would make this
> > advisory behavior available to applications using ordinary file-backed
> > shared mappings.
> >
> >> So fadvise would be an option. However, this "defer mode" is really odd.
> >> It sounds more like you would want to have a custom policy there [...]
> >
> > madvise is not the only option, deferring writeback for the entire file
> > would also be acceptable.
> >
> > The application only wants to indicate that it is still modifying the data
> > and would prefer background writeback to happen later. The kernel would
> > retain control over the actual timing. Memory pressure could override the
> > hint, and explicit synchronization would keep its normal semantics.
> >
> > The hint should also expire automatically when the owning mapping or handle
> > is released, including on process exit.
>
> Ok, so while you are updating the large mmap'ed file concurrently, you don't
> want writeback to go crazy, because you know that you will modify the memory
> immediately anyway.
Yeah see above, I really think the only sensible solution is a MAP_PRIVATE CoW'd
mapping or memfd etc.
>
> Let me CC some more people.
Christian and probably Willy also? But I'm not so sure there's anything sensible
to do here other than something-anon.
>
>
> --
> Cheers,
>
> David
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 10:27 ` 天狼
2026-09-30 11:38 ` David Hildenbrand (Arm)
2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
0 siblings, 2 replies; 20+ messages in thread
From: 天狼 @ 2026-09-30 10:27 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
Hi David, Lorenzo,
> You really need an actual file on disk. Shared memory / shmem / memfd is
> not sufficient?
Yes, an ordinary disk-backed file is the intended output, and subsequent
consumers access it through mmap.
Shared memory can preserve state across a process crash if the shared
object remains alive. However, using shmem/memfd would introduce a separate
staging and transfer step to put the data into the final file. Building
directly in a file-backed MAP_SHARED mapping lets us preserve the latest
userspace writes across process crashes and retain the same file-cache
pages for subsequent consumers.
> I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not
> sufficient either?
Correct. With MAP_PRIVATE, modifications remain private and are not carried
through to the underlying file. If the constructing process crashes before
transferring those modifications to the file, another process reopening the
file cannot access them. That would lose the process-crash survival
property we obtain from MAP_SHARED.
> It'd be problematic to effectively corrupt the page cache by saying "hey
> this is dirty but just clear dirty state and pretend this is what the
> disk has".
To clarify, the proposal does not involve clearing dirty state or
pretending that memory and disk contents match. The pages would remain
dirty and continue to count toward the existing dirty-page limits.
The requested hint only expresses that the application expects to modify
the data again soon and would prefer background writeback to happen later.
> Delaying writeback could interfere with the writeback balancing logic.
I understand that concern. The hint should not exempt the application from
dirty-page limits or throttling. If writeback is needed for balancing,
memory pressure, or explicit synchronization, the kernel should override
the hint.
The intended benefit is to avoid writing intermediate versions when the
kernel has room to defer that work. There is no requirement to prevent
writeback until construction finishes.
madvise is not the only option, deferring writeback for the entire file
would also be acceptable.
Could such an advisory preference be accommodated within the existing
writeback balancing policy while retaining its limits and fairness?
Thanks,
Peng
Lorenzo Stoakes (ARM) <ljs@kernel.org> 于2026年9月29日周二 16:59写道:
>
> On Tue, Sep 29, 2026 at 08:59:57AM +0200, David Hildenbrand (Arm) wrote:
> > On 9/29/26 06:06, 天狼 wrote:
> > > Hi David,
> > >
> >
> > Hi,
> >
> > >> What is supposed to happen if the process crashes when updating the
> > >> MAP_SHARED region halfway through?
> > >
> > > The data structure survives an interrupted update through its
> > > application-level copy-on-write and lock-free concurrency design. A
> > > file-backed MAP_SHARED mapping preserves the latest userspace writes after
> > > the process dies, even if they have not yet reached disk, provided the
> > > kernel remains running.
> > >
> > > This does not depend on preventing writeback. Deferring writeback is purely
> > > a performance optimization, not a correctness requirement.
> > >
> > >> It would be helpful if the use case + data structure would be explained
> > >> in a bit more detail.
> > >
> > > The data structure originally used uint32_t offsets instead of pointers to
> > > reduce pointer overhead. On that basis, we implemented lock-free concurrent
> > > reads and writes, obtaining crash safety for free at the same time. We use
> > > a file-backed MAP_SHARED mapping to turn that capability into an actual
> > > feature that survives process crashes.
> >
> > Ok, thanks. Just be sure: you really need an actual file on disk. Shared memory
> > / shmem / memfd is not sufficient?
> >
>
> I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not sufficient
> either?
>
> > >> What exactly is the problem with writeback here? Unnecessary I/O? Is
> > >> writeback the problem or actual reclaim after writeback?
> > >
> > > Unnecessary I/O. Reclaim after writeback is not the problem I am trying to
> > > address.
> >
> > Ok.
>
> I mean if you're mmap'ing it you're literally mapping the page cache folio and
> dirtying that folio.
>
> So the kernel really does have to do writeback and it might be quite problematic
> trying to prevent the writeback algorithm from doing its work on a specific
> range.
>
> I think the shape of a viable solution really is either MAP_PRIVATE-mapping it
> or using some anon shmem/memfd as David suggests.
>
> I think it'd be problematic to effectively corrupt the page cache by saying 'hey
> this is dirty but just clear dirty state and pretend this is what the disk has'
> or something.
>
>
> >
> > >
> > > During construction, background writeback can write intermediate contents
> > > that the application will soon overwrite. Those repeated writes waste disk
> > > bandwidth. Delaying writeback would avoid that waste where possible; it is
> > > not required for correctness.
> >
> > Understood.
>
> Yeah again delaying writeback could interfere with the writeback balancing logic
> which tries to keep dirty page levels sane and is fair and balanced so processes
> that writeback a lot get delayed in doing so under heavy dirtying.
>
> I think anything like this would interfere with that.
>
> >
> > >
> > >> In a fuse server you can in theory delay the writeback request.
> > >
> > > Thanks for the suggestion. That sounds like a possible way to experiment
> > > with delayed writeback. The feature I am requesting would make this
> > > advisory behavior available to applications using ordinary file-backed
> > > shared mappings.
> > >
> > >> So fadvise would be an option. However, this "defer mode" is really odd.
> > >> It sounds more like you would want to have a custom policy there [...]
> > >
> > > madvise is not the only option, deferring writeback for the entire file
> > > would also be acceptable.
> > >
> > > The application only wants to indicate that it is still modifying the data
> > > and would prefer background writeback to happen later. The kernel would
> > > retain control over the actual timing. Memory pressure could override the
> > > hint, and explicit synchronization would keep its normal semantics.
> > >
> > > The hint should also expire automatically when the owning mapping or handle
> > > is released, including on process exit.
> >
> > Ok, so while you are updating the large mmap'ed file concurrently, you don't
> > want writeback to go crazy, because you know that you will modify the memory
> > immediately anyway.
>
> Yeah see above, I really think the only sensible solution is a MAP_PRIVATE CoW'd
> mapping or memfd etc.
>
> >
> > Let me CC some more people.
>
> Christian and probably Willy also? But I'm not so sure there's anything sensible
> to do here other than something-anon.
>
> >
> >
> > --
> > Cheers,
> >
> > David
>
> --
> Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 10:27 ` 天狼
@ 2026-09-30 11:38 ` David Hildenbrand (Arm)
2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
1 sibling, 0 replies; 20+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-30 11:38 UTC (permalink / raw)
To: 天狼, Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-fsdevel, Liam R. Howlett, Vlastimil Babka,
Jann Horn, Matthew Wilcox, Jan Kara
On 9/30/26 12:27, 天狼 wrote:
> Hi David, Lorenzo,
>
Hi,
>> You really need an actual file on disk. Shared memory / shmem / memfd is
>> not sufficient?
>
> Yes, an ordinary disk-backed file is the intended output, and subsequent
> consumers access it through mmap.
>
> Shared memory can preserve state across a process crash if the shared
> object remains alive. However, using shmem/memfd would introduce a separate
> staging and transfer step to put the data into the final file. Building
> directly in a file-backed MAP_SHARED mapping lets us preserve the latest
> userspace writes across process crashes and retain the same file-cache
> pages for subsequent consumers.
Why do you really need the data on a file in disk?
What you could do is, load it once from the file into shmem, then let everybody
work on shmem, and have some background thread that periodically writes the
shmem content out to the real file on disk.
Then, you could completely control when I/O would happen.
[...]
>> It'd be problematic to effectively corrupt the page cache by saying "hey
>> this is dirty but just clear dirty state and pretend this is what the
>> disk has".
>
> To clarify, the proposal does not involve clearing dirty state or
> pretending that memory and disk contents match. The pages would remain
> dirty and continue to count toward the existing dirty-page limits.
>
IIUC, vmscan would still write out the folio if dirty, so I'd assume that at
least memory reclaim would not be affected, only background writeback.
> The requested hint only expresses that the application expects to modify
> the data again soon and would prefer background writeback to happen later.
>
>> Delaying writeback could interfere with the writeback balancing logic.
>
> I understand that concern. The hint should not exempt the application from
> dirty-page limits or throttling. If writeback is needed for balancing,
> memory pressure, or explicit synchronization, the kernel should override
> the hint.
>
> The intended benefit is to avoid writing intermediate versions when the
> kernel has room to defer that work. There is no requirement to prevent
> writeback until construction finishes.
>
> madvise is not the only option, deferring writeback for the entire file
> would also be acceptable.
>
> Could such an advisory preference be accommodated within the existing
> writeback balancing policy while retaining its limits and fairness?
When using MAP_SHARED with disk-based files for VM memory, people ran into
similar problems: the VM will constantly dirty many folios and background
writeback will just permanently write these out and wear storage. For VMs you
really only want the memory persisted in the file once you e.g., migrate the VM.
Essentially, once the VM is paused.
One of the reasons why we tell people to not use that combination and use
shmem/hugetlb instead. I do wonder whether there is some similarity ... we don't
want brackground writeback as long as there is heavy activity on these files.
--
Cheers,
David
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 10:27 ` 天狼
2026-09-30 11:38 ` David Hildenbrand (Arm)
@ 2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
1 sibling, 1 reply; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 12:08 UTC (permalink / raw)
To: 天狼
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Wed, Sep 30, 2026 at 06:27:14PM +0800, 天狼 wrote:
> > You really need an actual file on disk. Shared memory / shmem / memfd is
> > not sufficient?
>
> Yes, an ordinary disk-backed file is the intended output, and subsequent
> consumers access it through mmap.
>
> Shared memory can preserve state across a process crash if the shared
> object remains alive. However, using shmem/memfd would introduce a separate
> staging and transfer step to put the data into the final file. Building
> directly in a file-backed MAP_SHARED mapping lets us preserve the latest
> userspace writes across process crashes and retain the same file-cache
> pages for subsequent consumers.
In general, writing to files through MAP_SHARED mappings is discouraged.
>
> > I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not
> > sufficient either?
>
> Correct. With MAP_PRIVATE, modifications remain private and are not carried
> through to the underlying file. If the constructing process crashes before
> transferring those modifications to the file, another process reopening the
> file cannot access them. That would lose the process-crash survival
> property we obtain from MAP_SHARED.
But you can solve that with a memfd?
>
> > It'd be problematic to effectively corrupt the page cache by saying "hey
> > this is dirty but just clear dirty state and pretend this is what the
> > disk has".
>
> To clarify, the proposal does not involve clearing dirty state or
> pretending that memory and disk contents match. The pages would remain
> dirty and continue to count toward the existing dirty-page limits.
>
> The requested hint only expresses that the application expects to modify
> the data again soon and would prefer background writeback to happen later.
But that's the same as saying 'don't writeback'?
When there's a lot of dirtied pages on disk writes are delayed in proportion to
how dirtied with various other heuristics taken into account.
I don't see how this requested feature wouldn't just give programs a way to work
around that, meanwhile everything else doing dirty writeback are _impacted by
the dirtying that the requesting process has done_.
That seems really problematic.
>
> > Delaying writeback could interfere with the writeback balancing logic.
>
> I understand that concern. The hint should not exempt the application from
> dirty-page limits or throttling. If writeback is needed for balancing,
> memory pressure, or explicit synchronization, the kernel should override
> the hint.
Hmm.
The delay before background writeback happens is, by default, quite long,
until you have a lot of dirtied pages.
dirty_expire_centisecs defaults to 3000, i.e. 30s before it is considered
for writeback.
And then that only happens every dirty_writeback_centisecs, defaults to
500, i.e. every 5 seconds.
If you're not writing to the mapping within 5 seconds even, let alone 30
seconds then that seems like an issue with your program.
I guess you are re-dirtying again with later changes, but the _triggering_
of writeback is at an inode granularity on dirtying - so you want to delay
writing back some pages because later pages are not meaningful?
All the while you are adding to the total dirty pages but not paying the
price anywhere, nor allowing the dirty page count to reduce.
Unfortunately there's no real way to avoid the per-inode thing, and
delaying valid writeback there because SOME of it isn't 'valid' is not
really sensible.
Either that, or you are actually concerned when there are enough
dirty pages to trigger background writeback i.e. dirty_background_ratio is
exceeded (where the above limits don't apply).
But later you say you don't want the proposed change to impact how that
behaves (the balance_dirty_pages() logic), which contradicts that.
So are you taking ~35s to actually write to this properly?
And is it that you're wanting the _triggering_ of writeback to be
per-folio? That isn't really sensible.
And nor is it really to delay writeback for dirty folios because later ones
are 'invalid'.
>
> The intended benefit is to avoid writing intermediate versions when the
> kernel has room to defer that work. There is no requirement to prevent
> writeback until construction finishes.
>
> madvise is not the only option, deferring writeback for the entire file
> would also be acceptable.
>
> Could such an advisory preference be accommodated within the existing
> writeback balancing policy while retaining its limits and fairness?
I think the fundamental friction here is that you're charging a cost to a
global limit then doing something that prevents that cost being paid (you
can't writeback), and meanwhile the process can wait for arbitrary time
before having to write to it.
And the fact that it is taking you so long to actually write what is
intended there that this matters suggests to me there's something wrong
with what you're doing here.
I'd say in general the shape of the solution would be something
cgroups-ish, but since all of this is global stuff and the knobs are all
global I think you need to think about your design a little more.
Having a process that is kept around maybe with an appropriate OOM score
that issues a memfd would get you the sharing even if process dies stuff,
and then it can periodically write to a file when necessary.
I think that'd work a lot better, and I don't really see a kernel change
here that would make sense.
All in all my intuition is that you're looking for a kernel solution to
something that should be implemented differently in userland.
And in general we've seen a pattern of 'writes to MAP_SHARED, experiences
problems' before.
And the solution is generally - use something like memfd, have a process
that does writebacks manually for you as you need.
So I strongly suggest you do that! :)
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14 ` 天狼
0 siblings, 1 reply; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 12:11 UTC (permalink / raw)
To: 天狼
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Wed, Sep 30, 2026 at 01:08:33PM +0100, Lorenzo Stoakes (ARM) wrote:
> I'd say in general the shape of the solution would be something
> cgroups-ish, but since all of this is global stuff and the knobs are all
> global I think you need to think about your design a little more.
(Sorry to be clear, this is - if we accepted that this was a kernel issue that
truly needed something like this - which I do not. The solution here is to use
shared memory as per below.)
> And in general we've seen a pattern of 'writes to MAP_SHARED, experiences
> problems' before.
>
> And the solution is generally - use something like memfd, have a process
> that does writebacks manually for you as you need.
>
> So I strongly suggest you do that! :)
>
> --
> Cheers, Lorenzo
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 14:14 ` 天狼
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
2026-09-30 15:29 ` 天狼
0 siblings, 2 replies; 20+ messages in thread
From: 天狼 @ 2026-09-30 14:14 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
Hi David, Lorenzo,
We currently work around this issue by setting the global
vm.dirty_expire_centisecs to 60000.
We have also developed a kernel module as another option. It exposes an
ioctl that takes a file descriptor. For an already-dirty inode, it updates
dirtied_when to the current jiffies and moves the inode to the newest end
of the writeback dirty list, postponing its selection by periodic
writeback. It leaves the dirty state intact and does not remove an inode
from an already-queued flush.
To keep postponing periodic writeback, the application must call this
ioctl continuously at intervals shorter than vm.dirty_expire_centisecs,
expressed in centiseconds. Once the calls stop, the normal expiration
interval applies from the last call.
The key code, with checks and locking omitted, is:
inode->dirtied_when = jiffies;
if (!(inode->i_state & I_SYNC_QUEUED))
list_move(&inode->i_io_list, &wb->b_dirty);
However, neither changing a global setting nor maintaining a separate
kernel module is as elegant as having this supported directly in the
kernel.
Thanks,
Peng
Lorenzo Stoakes (ARM) <ljs@kernel.org> 于2026年9月30日周三 20:11写道:
>
> On Wed, Sep 30, 2026 at 01:08:33PM +0100, Lorenzo Stoakes (ARM) wrote:
> > I'd say in general the shape of the solution would be something
> > cgroups-ish, but since all of this is global stuff and the knobs are all
> > global I think you need to think about your design a little more.
>
> (Sorry to be clear, this is - if we accepted that this was a kernel issue that
> truly needed something like this - which I do not. The solution here is to use
> shared memory as per below.)
>
> > And in general we've seen a pattern of 'writes to MAP_SHARED, experiences
> > problems' before.
> >
> > And the solution is generally - use something like memfd, have a process
> > that does writebacks manually for you as you need.
> >
> > So I strongly suggest you do that! :)
> >
> > --
> > Cheers, Lorenzo
>
> --
> Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 14:14 ` 天狼
@ 2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
2026-09-30 15:33 ` Lorenzo Stoakes (ARM)
2026-09-30 15:29 ` 天狼
1 sibling, 1 reply; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 14:24 UTC (permalink / raw)
To: 天狼
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Wed, Sep 30, 2026 at 10:14:27PM +0800, 天狼 wrote:
> Hi David, Lorenzo,
>
> We currently work around this issue by setting the global
> vm.dirty_expire_centisecs to 60000.
>
> We have also developed a kernel module as another option. It exposes an
> ioctl that takes a file descriptor. For an already-dirty inode, it updates
> dirtied_when to the current jiffies and moves the inode to the newest end
> of the writeback dirty list, postponing its selection by periodic
> writeback. It leaves the dirty state intact and does not remove an inode
> from an already-queued flush.
>
> To keep postponing periodic writeback, the application must call this
> ioctl continuously at intervals shorter than vm.dirty_expire_centisecs,
> expressed in centiseconds. Once the calls stop, the normal expiration
> interval applies from the last call.
>
> The key code, with checks and locking omitted, is:
>
> inode->dirtied_when = jiffies;
> if (!(inode->i_state & I_SYNC_QUEUED))
> list_move(&inode->i_io_list, &wb->b_dirty);
>
Yeah that's absolutely violating how writeback is supposed to work.
> However, neither changing a global setting nor maintaining a separate
> kernel module is as elegant as having this supported directly in the
> kernel.
You're neglecting the question David and I have both raised with you here, which
is why you can't simply use shmem (e.g. a memfd) to achieve what you want to do
here?
Writing to a MAP_SHARED mapping is not recommended for a number of reasons, and
most software that does something like what you're doing uses shmem to achieve
it.
Overall I don't think there's a sensible kernel solution to your problem.
I strongly recommend you look at a shmem solution.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 14:14 ` 天狼
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 15:29 ` 天狼
2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
2026-09-30 16:38 ` Jan Kara
1 sibling, 2 replies; 20+ messages in thread
From: 天狼 @ 2026-09-30 15:29 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
Hi David, Lorenzo,
The distinction behind this request is between surviving a process crash
and surviving a power failure. After a process crash, the kernel and page
cache remain alive. A shared file mapping therefore retains the writes
already made by the process, even if they have not been written to disk.
ToplingDB has file-mmap MemTables whose data structures survive a process
crash. Its crash-safe recovery reuses those MemTables instead of rebuilding
them by replaying the entire WAL. A published sequence number aligns the
MemTable's per-KV crash-safe boundary with the database's WriteBatch
atomicity, so recovery can reuse the published data and replay only the
remaining WAL tail.
These MemTables also support cheap, millisecond-level ConvertToSST
(wrap-up and sync, i.e. LSM flush). The file-mmap path renames the existing
file and appends metadata, preserving the data structure already built in
it. This avoids copying the contents into a separately constructed SST.
Moving construction into shmem would add a transfer to the regular file
that this path currently avoids.
Avoiding full WAL replay makes much larger MemTables practical. While a
large MemTable is being populated, random updates can repeatedly dirty
pages that background writeback has already written out. Those writes of
intermediate states consume disk bandwidth without being necessary for
process-crash recovery. Deferring them is a performance optimization;
correctness does not depend on the kernel honoring the hint.
This is a general problem, not something specific to ToplingDB. ToplingDB
explicitly distinguishes process crashes from power failures and has
invested heavily in optimizing recovery from process crashes. By contrast,
LMDB faces the same problem but does not explicitly distinguish process
crashes from power failures. If sync is skipped when using LMDB, the
remaining writeback problem is the same as in ToplingDB.
The global setting and kernel module described in my previous email let us
work around this locally. The remaining problem is making the optimization
available to users of a database library. We cannot reasonably require
every deployment to change a system-wide setting or install and maintain
an out-of-tree kernel module. The module gets us from no solution to a
local solution, but it does not provide a broadly deployable interface.
That is why I am asking for kernel support: a supported, best-effort hint
that applications can use directly, while leaving the kernel free to write
back when necessary. madvise is not the only option, deferring writeback
for the entire file would also be acceptable.
Per-file deferred writeback would require minimal kernel changes and pose
a very low risk, while offering enormous benefits.
Thanks,
Peng
天狼 <rockeet@gmail.com> 于2026年9月30日周三 22:14写道:
>
> Hi David, Lorenzo,
>
> We currently work around this issue by setting the global
> vm.dirty_expire_centisecs to 60000.
>
> We have also developed a kernel module as another option. It exposes an
> ioctl that takes a file descriptor. For an already-dirty inode, it updates
> dirtied_when to the current jiffies and moves the inode to the newest end
> of the writeback dirty list, postponing its selection by periodic
> writeback. It leaves the dirty state intact and does not remove an inode
> from an already-queued flush.
>
> To keep postponing periodic writeback, the application must call this
> ioctl continuously at intervals shorter than vm.dirty_expire_centisecs,
> expressed in centiseconds. Once the calls stop, the normal expiration
> interval applies from the last call.
>
> The key code, with checks and locking omitted, is:
>
> inode->dirtied_when = jiffies;
> if (!(inode->i_state & I_SYNC_QUEUED))
> list_move(&inode->i_io_list, &wb->b_dirty);
>
> However, neither changing a global setting nor maintaining a separate
> kernel module is as elegant as having this supported directly in the
> kernel.
>
> Thanks,
> Peng
>
> Lorenzo Stoakes (ARM) <ljs@kernel.org> 于2026年9月30日周三 20:11写道:
> >
> > On Wed, Sep 30, 2026 at 01:08:33PM +0100, Lorenzo Stoakes (ARM) wrote:
> > > I'd say in general the shape of the solution would be something
> > > cgroups-ish, but since all of this is global stuff and the knobs are all
> > > global I think you need to think about your design a little more.
> >
> > (Sorry to be clear, this is - if we accepted that this was a kernel issue that
> > truly needed something like this - which I do not. The solution here is to use
> > shared memory as per below.)
> >
> > > And in general we've seen a pattern of 'writes to MAP_SHARED, experiences
> > > problems' before.
> > >
> > > And the solution is generally - use something like memfd, have a process
> > > that does writebacks manually for you as you need.
> > >
> > > So I strongly suggest you do that! :)
> > >
> > > --
> > > Cheers, Lorenzo
> >
> > --
> > Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 15:33 ` Lorenzo Stoakes (ARM)
0 siblings, 0 replies; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 15:33 UTC (permalink / raw)
To: 天狼
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Wed, Sep 30, 2026 at 03:24:53PM +0100, Lorenzo Stoakes (ARM) wrote:
> I strongly recommend you look at a shmem solution.
To expand upon that:
The general structure would be something like a minimal process, with an
oom_score_adj set such that it can't be killed (or is at least very unlikely).
Have it do something like allocate a memfd.
Then have that be as minimal as possible - it sets things up, it hands off
memfds over a UNIX socket for instance, and is set up to periodically write
back to disk.
That way you have something that should never crash/exit which gives you the
guarantees you need.
You can then also control how to synchronise things between processes, when to
write back and how much, etc.
You can keep it in a cgroup as well to isolate it from the rest of the system
too, potentially.
Then have anything that does anything non-trivial be clients of that.
Everything works as before, you can ensure that intermediate state doesn't get
written back through whatever mechanism you want to employ, and you avoid all of
the pitfalls of writing to a MAP_SHARED file-backed mapping.
If the memfd isn't enough for you, you could also have it establish a tmpfs file
which is then shared between all of the processes. That will survive even the
establishing process dying.
The main win here is that you can choose when and how to writeback using any
mechanism you like.
For instance, you could have a dirty bitmap be part of the shared memory and
atomically update the relevant bit when it's ready to be written back.
That kind of approach gives you total control over when and how things are
written back.
And avoids the known issues with writing to a MAP_SHARED file:
* Terrible error handling - random SIGBUS's, network file systems in particular
are very problematic, writeback errors are hard to obtain (fsync(), msync()
needed to even get them) - note the tmpfs compromise mentioned above has this
issue too.
* No control over writeback - Exactly your issue. This is what you're trying to
work around.
* Writeback stalls - as above. The dirty writeback balancing bites there.
* Reclaim - dirty pages can't be reclaimed until written back.
* Fault overhead - every fresh dirtying write is a page fault (that's how dirty
tracking works) and you have that overhead throughout. Once written back it's
cleaned and causes the same cost again next time.
The TL;DR is you're trying to manipulate logic whose whole job it is to
writeback to not do that.
So stop doing that :) and life is much easier.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 15:29 ` 天狼
@ 2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
2026-09-30 16:38 ` Jan Kara
1 sibling, 0 replies; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 15:43 UTC (permalink / raw)
To: 天狼
Cc: David Hildenbrand (Arm), linux-mm, linux-fsdevel, Liam R. Howlett,
Vlastimil Babka, Jann Horn, Matthew Wilcox, Jan Kara
On Wed, Sep 30, 2026 at 11:29:09PM +0800, 天狼 wrote:
> Hi David, Lorenzo,
...
> Per-file deferred writeback would require minimal kernel changes and pose
> a very low risk, while offering enormous benefits.
I'm afraid the opposite is true.
The proposed expiry hint does nothing as soon as you hit the dirty background
threshold and in any case violates kernel writeback as explained.
And nothing you've proposed here is upstreamable.
A locked down simple process handing out memfds (as explained in detail in
other mail) or a named tmpfs file survives process crash perfectly fine and is
the correct shape of the solution.
This is what you should pursue.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 15:29 ` 天狼
2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 16:38 ` Jan Kara
2026-10-01 4:28 ` 天狼
1 sibling, 1 reply; 20+ messages in thread
From: Jan Kara @ 2026-09-30 16:38 UTC (permalink / raw)
To: 天狼
Cc: Lorenzo Stoakes (ARM), David Hildenbrand (Arm), linux-mm,
linux-fsdevel, Liam R. Howlett, Vlastimil Babka, Jann Horn,
Matthew Wilcox, Jan Kara
Hello!
On Wed 30-09-26 23:29:09, 天狼 wrote:
> The distinction behind this request is between surviving a process crash
> and surviving a power failure. After a process crash, the kernel and page
> cache remain alive. A shared file mapping therefore retains the writes
> already made by the process, even if they have not been written to disk.
>
> ToplingDB has file-mmap MemTables whose data structures survive a process
> crash. Its crash-safe recovery reuses those MemTables instead of rebuilding
> them by replaying the entire WAL. A published sequence number aligns the
> MemTable's per-KV crash-safe boundary with the database's WriteBatch
> atomicity, so recovery can reuse the published data and replay only the
> remaining WAL tail.
>
> These MemTables also support cheap, millisecond-level ConvertToSST
> (wrap-up and sync, i.e. LSM flush). The file-mmap path renames the existing
> file and appends metadata, preserving the data structure already built in
> it. This avoids copying the contents into a separately constructed SST.
> Moving construction into shmem would add a transfer to the regular file
> that this path currently avoids.
So I think I understand your design and the issues you're facing. After all
I've seen very similar issues with Berkley DB which was also using
MAP_SHARED memory for the DB and under some circumstances the workload
performance was suffering from too frequent writeback.
Now what I miss from your problem description is why "a transfer to the
regular file" from shmem would matter to you. What is the problem with
that? Because a lot of databases do it like that and it works for them. If
you use AIO + direct IO (or io_uring) your performance is going to be even
somewhat superior (or at least on par) to current writeback workers in the
kernel. In terms of data integrity guarantees you are again as good as or
better than using in-kernel writeback because you are in full control when
you write and what you write. So the only reason I can come up with is that
you want to avoid the complexity of your writeback process in your DB. But
in the same spirit we as kernel people want to keep the complexity out of
the kernel :) and if we'd add a feature like this we'd have to maintain it
practically forever so we are very cautious with that.
> Avoiding full WAL replay makes much larger MemTables practical. While a
> large MemTable is being populated, random updates can repeatedly dirty
> pages that background writeback has already written out. Those writes of
> intermediate states consume disk bandwidth without being necessary for
> process-crash recovery. Deferring them is a performance optimization;
> correctness does not depend on the kernel honoring the hint.
>
> This is a general problem, not something specific to ToplingDB. ToplingDB
> explicitly distinguishes process crashes from power failures and has
> invested heavily in optimizing recovery from process crashes. By contrast,
> LMDB faces the same problem but does not explicitly distinguish process
> crashes from power failures. If sync is skipped when using LMDB, the
> remaining writeback problem is the same as in ToplingDB.
>
> The global setting and kernel module described in my previous email let us
> work around this locally. The remaining problem is making the optimization
> available to users of a database library. We cannot reasonably require
> every deployment to change a system-wide setting or install and maintain
> an out-of-tree kernel module. The module gets us from no solution to a
> local solution, but it does not provide a broadly deployable interface.
>
> That is why I am asking for kernel support: a supported, best-effort hint
> that applications can use directly, while leaving the kernel free to write
> back when necessary. madvise is not the only option, deferring writeback
> for the entire file would also be acceptable.
> Per-file deferred writeback would require minimal kernel changes and pose
> a very low risk, while offering enormous benefits.
The problem I see with your per-inode hint is that its semantics is very
vague (and practically tied to the implementation). And this practically
always backfires because people start to make assumptions about the
behavior, then something changes, people's expectations break and people
complain. Or they come with 100 and 1 ways how to tweak the vague behavior
to fit their special usecase and it quickly becomes a mess of conflicting
needs. We've been there many times :-|.
What I could imagine both technically and semantically feasible (but
haven't decided whether it would be used enough to be worth the hassle) is
to implement some say fcntl() / ioctl() with which an owner of the inode
could set / clear a file hint that the inode should be extempt from
periodic writeback - i.e., the writeback that runs once per 5s and flushes
inodes that were dirtied longer than 30s ago (by default). I.e., we will
writeback the inode whenever we see fit but we won't do it just because of
the file's age.
But again similarly as David or Lorenzo I think that implementing your own
writeback worker would be a superior solution for your DB.
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-09-30 16:38 ` Jan Kara
@ 2026-10-01 4:28 ` 天狼
2026-10-01 8:35 ` Lorenzo Stoakes (ARM)
0 siblings, 1 reply; 20+ messages in thread
From: 天狼 @ 2026-10-01 4:28 UTC (permalink / raw)
To: Jan Kara
Cc: Lorenzo Stoakes (ARM), David Hildenbrand (Arm), linux-mm,
linux-fsdevel, Liam R. Howlett, Vlastimil Babka, Jann Horn,
Matthew Wilcox
Hi Jan,
The reasons for requesting this hint are:
1. With buffered I/O, the data already resides in shmem, and writing
it to the regular file allocates another set of page-cache pages. When
both copies are resident, this roughly doubles the memory occupied by
the data. Traditional LSM MemTables are usually small, so duplicating
one is relatively inexpensive (they waste so much elsewhere that they
never even reach the scale where this becomes a concern). Our
MemTables can be very large, making duplication a substantial waste of
memory that significantly undermines the advantages of our design.
2. During the MemTable-to-SST transition, existing readers still
access the data through the MemTable, while new readers access it
through the SST. We therefore cannot immediately release the shmem
copy. Our current ConvertToSST instead renames the existing file and
appends metadata. The inode remains the same, so the existing MemTable
mapping and the new SST mapping share the same page-cache pages
throughout this overlap period.
3. AIO or io_uring with direct I/O avoids allocating the destination
page cache during the write, but it does not turn the source shmem
pages into page-cache pages of the SST file. Subsequent SST mmap reads
populate a separate page cache, while existing MemTable readers still
need the shmem pages. The duplicate memory therefore remains a problem
during the overlap period.
4. Temporary files are a much broader use case. Periodic automatic
writeback is unnecessary for temporary data; writeback for memory
reclaim is what serves a practical purpose. A per-file exemption from
age-based periodic writeback would benefit these workloads as well.
The per-file hint you described would meet our needs: exempting the
inode from periodic writeback based on its dirty age. The kernel would
remain free to write it back for other reasons, and explicit
synchronization would continue to work normally.
Thanks,
Peng
Jan Kara <jack@suse.cz> 于2026年10月1日周四 00:38写道:
>
> Hello!
>
> On Wed 30-09-26 23:29:09, 天狼 wrote:
> > The distinction behind this request is between surviving a process crash
> > and surviving a power failure. After a process crash, the kernel and page
> > cache remain alive. A shared file mapping therefore retains the writes
> > already made by the process, even if they have not been written to disk.
> >
> > ToplingDB has file-mmap MemTables whose data structures survive a process
> > crash. Its crash-safe recovery reuses those MemTables instead of rebuilding
> > them by replaying the entire WAL. A published sequence number aligns the
> > MemTable's per-KV crash-safe boundary with the database's WriteBatch
> > atomicity, so recovery can reuse the published data and replay only the
> > remaining WAL tail.
> >
> > These MemTables also support cheap, millisecond-level ConvertToSST
> > (wrap-up and sync, i.e. LSM flush). The file-mmap path renames the existing
> > file and appends metadata, preserving the data structure already built in
> > it. This avoids copying the contents into a separately constructed SST.
> > Moving construction into shmem would add a transfer to the regular file
> > that this path currently avoids.
>
> So I think I understand your design and the issues you're facing. After all
> I've seen very similar issues with Berkley DB which was also using
> MAP_SHARED memory for the DB and under some circumstances the workload
> performance was suffering from too frequent writeback.
>
> Now what I miss from your problem description is why "a transfer to the
> regular file" from shmem would matter to you. What is the problem with
> that? Because a lot of databases do it like that and it works for them. If
> you use AIO + direct IO (or io_uring) your performance is going to be even
> somewhat superior (or at least on par) to current writeback workers in the
> kernel. In terms of data integrity guarantees you are again as good as or
> better than using in-kernel writeback because you are in full control when
> you write and what you write. So the only reason I can come up with is that
> you want to avoid the complexity of your writeback process in your DB. But
> in the same spirit we as kernel people want to keep the complexity out of
> the kernel :) and if we'd add a feature like this we'd have to maintain it
> practically forever so we are very cautious with that.
>
> > Avoiding full WAL replay makes much larger MemTables practical. While a
> > large MemTable is being populated, random updates can repeatedly dirty
> > pages that background writeback has already written out. Those writes of
> > intermediate states consume disk bandwidth without being necessary for
> > process-crash recovery. Deferring them is a performance optimization;
> > correctness does not depend on the kernel honoring the hint.
> >
> > This is a general problem, not something specific to ToplingDB. ToplingDB
> > explicitly distinguishes process crashes from power failures and has
> > invested heavily in optimizing recovery from process crashes. By contrast,
> > LMDB faces the same problem but does not explicitly distinguish process
> > crashes from power failures. If sync is skipped when using LMDB, the
> > remaining writeback problem is the same as in ToplingDB.
> >
> > The global setting and kernel module described in my previous email let us
> > work around this locally. The remaining problem is making the optimization
> > available to users of a database library. We cannot reasonably require
> > every deployment to change a system-wide setting or install and maintain
> > an out-of-tree kernel module. The module gets us from no solution to a
> > local solution, but it does not provide a broadly deployable interface.
> >
> > That is why I am asking for kernel support: a supported, best-effort hint
> > that applications can use directly, while leaving the kernel free to write
> > back when necessary. madvise is not the only option, deferring writeback
> > for the entire file would also be acceptable.
> > Per-file deferred writeback would require minimal kernel changes and pose
> > a very low risk, while offering enormous benefits.
>
> The problem I see with your per-inode hint is that its semantics is very
> vague (and practically tied to the implementation). And this practically
> always backfires because people start to make assumptions about the
> behavior, then something changes, people's expectations break and people
> complain. Or they come with 100 and 1 ways how to tweak the vague behavior
> to fit their special usecase and it quickly becomes a mess of conflicting
> needs. We've been there many times :-|.
>
> What I could imagine both technically and semantically feasible (but
> haven't decided whether it would be used enough to be worth the hassle) is
> to implement some say fcntl() / ioctl() with which an owner of the inode
> could set / clear a file hint that the inode should be extempt from
> periodic writeback - i.e., the writeback that runs once per 5s and flushes
> inodes that were dirtied longer than 30s ago (by default). I.e., we will
> writeback the inode whenever we see fit but we won't do it just because of
> the file's age.
>
> But again similarly as David or Lorenzo I think that implementing your own
> writeback worker would be a superior solution for your DB.
>
>
> Honza
> --
> Jan Kara <jack@suse.com>
> SUSE Labs, CR
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-10-01 4:28 ` 天狼
@ 2026-10-01 8:35 ` Lorenzo Stoakes (ARM)
2026-10-01 13:21 ` 天狼
0 siblings, 1 reply; 20+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-01 8:35 UTC (permalink / raw)
To: 天狼
Cc: Jan Kara, David Hildenbrand (Arm), linux-mm, linux-fsdevel,
Liam R. Howlett, Vlastimil Babka, Jann Horn, Matthew Wilcox
Hi Peng,
(Point of etiquette - and you are perhaps not aware - in general in the
kernel we reply inline to people rather than sending big lists in
reply.)
>1. With buffered I/O, the data already resides in shmem, and writing
>it to the regular file allocates another set of page-cache pages. When
>both copies are resident, this roughly doubles the memory occupied by
>the data. Traditional LSM MemTables are usually small, so duplicating
>one is relatively inexpensive (they waste so much elsewhere that they
>never even reach the scale where this becomes a concern). Our
>MemTables can be very large, making duplication a substantial waste of
>memory that significantly undermines the advantages of our design.
Then use O_DIRECT :)
>
>2. During the MemTable-to-SST transition, existing readers still
>access the data through the MemTable, while new readers access it
>through the SST. We therefore cannot immediately release the shmem
>copy. Our current ConvertToSST instead renames the existing file and
>appends metadata. The inode remains the same, so the existing MemTable
>mapping and the new SST mapping share the same page-cache pages
>throughout this overlap period.
That's your choice, not a requirement. The SST is the MemTable plus
metadata, so append the metadata in shmem and serve new readers from there
too.
>
>3. AIO or io_uring with direct I/O avoids allocating the destination
>page cache during the write, but it does not turn the source shmem
>pages into page-cache pages of the SST file. Subsequent SST mmap reads
>populate a separate page cache, while existing MemTable readers still
>need the shmem pages. The duplicate memory therefore remains a problem
>during the overlap period.
It doesn't need to. Serve everyone from shmem until existing readers drain,
then drop it and switch to the SST. Only one copy is ever resident.
>
>4. Temporary files are a much broader use case. Periodic automatic
>writeback is unnecessary for temporary data; writeback for memory
>reclaim is what serves a practical purpose. A per-file exemption from
>age-based periodic writeback would benefit these workloads as well.
Temporary data that doesn't need to reach disk belongs in shmem :)
Overall - each objection so far has been to the cost of changing your
design rather than to the userspace approaches not working.
I agree with Jan that the inode hint is viable in shape, but we'd need a
good reason to add it, and avoiding a design change isn't one :)
David, Jan and I have all suggested a shmem-based userspace
approach. Please explore that properly first.
I took the time to write up the approach, in detail, at
https://lore.kernel.org/linux-mm/ar0mdJKLC0cHBoKM@gremlin/ hopefully that's
helpful.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
2026-10-01 8:35 ` Lorenzo Stoakes (ARM)
@ 2026-10-01 13:21 ` 天狼
0 siblings, 0 replies; 20+ messages in thread
From: 天狼 @ 2026-10-01 13:21 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: Jan Kara, David Hildenbrand (Arm), linux-mm, linux-fsdevel,
Liam R. Howlett, Vlastimil Babka, Jann Horn, Matthew Wilcox
Hi Lorenzo,
Thanks for explaining the shmem approach and for pointing out the
mailing-list etiquette.
> Overall - each objection so far has been to the cost of changing your
> design rather than to the userspace approaches not working.
I agree that userspace workarounds can solve the problem. In fact, our
current workarounds already do. However, they are awkward and much less
elegant than addressing the issue through the kernel interface we are
hoping for.
The underlying interface gap remains: an application knows when writeback
would be useful, but has no per-file way to ask the kernel to defer
age-based periodic writeback.
I believe this is a general issue, rather than something specific to our
design, and that we are unlikely to be the only ones who would welcome
such an interface.
The best-effort hint Jan described seems to address this directly, while
leaving memory-pressure writeback and explicit synchronization available
as usual. I hope we can consider its broader usefulness on that basis.
My proposal is only one possible approach. I would be very happy to see
the kernel provide an even more elegant solution than the one I have in
mind.
Thanks,
Peng
^ permalink raw reply [flat|nested] 20+ messages in thread
end of thread, other threads:[~2026-10-01 13:21 UTC | newest]
Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09 ` 天狼
2026-09-28 19:12 ` David Hildenbrand (Arm)
2026-09-29 4:06 ` 天狼
2026-09-29 6:59 ` David Hildenbrand (Arm)
2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
2026-09-30 10:27 ` 天狼
2026-09-30 11:38 ` David Hildenbrand (Arm)
2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14 ` 天狼
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
2026-09-30 15:33 ` Lorenzo Stoakes (ARM)
2026-09-30 15:29 ` 天狼
2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
2026-09-30 16:38 ` Jan Kara
2026-10-01 4:28 ` 天狼
2026-10-01 8:35 ` Lorenzo Stoakes (ARM)
2026-10-01 13:21 ` 天狼
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox