Linux filesystem development
 help / color / mirror / Atom feed
* [RFC] madvise: best-effort deferred writeback for shared file mappings
@ 2026-09-28  7:07 天狼
  2026-09-28 11:29 ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 20+ messages in thread
From: 天狼 @ 2026-09-28  7:07 UTC (permalink / raw)
  To: linux-mm; +Cc: linux-fsdevel

Hello,

I would like to request feedback on a mapping-scoped, best-effort hint
to defer background writeback while an application is still constructing
data in a writable MAP_SHARED mapping of a regular disk-backed file.
This is a feature request, not a patch submission or a benchmark report.

Motivation
----------

Consider an application building a large data structure, for example
10 GiB, through multiple passes that repeatedly modify the same pages.
After construction, it persists the file and publishes it for other
processes to mmap and read.

Building directly in MAP_SHARED keeps the data in the file's page cache,
which also serves the consumers. However, background writeback can write
intermediate contents that subsequent construction passes dirty again.
The application knows when construction ends, but cannot express that
knowledge as a mapping-local writeback scheduling hint.

Building in private memory and then doing buffered writes adds a bulk
copy into the file's page cache. Direct I/O does not itself populate that
cache for subsequent file-mapped consumers. A memfd can share the build
buffer, but requires a separate object for persistence and changes the
consumer interface. Global dirty/writeback tuning is broader than this
particular mapping and construction phase.

Proposed interface and behavior
-------------------------------

Tentatively, a pair of madvise operations could express this:

    madvise(addr, len, MADV_DEFER_WRITEBACK);
    /* Construct and repeatedly update the shared file mapping. */
    madvise(addr, len, MADV_ALLOW_WRITEBACK);
    /* Use existing synchronization calls if durability is required. */

The names are placeholders for discussion, not existing constants.

The requested behavior is deliberately advisory:

* Prefer to defer ordinary background writeback for the corresponding
  file pages while the hint is active, including pages faulted in later
  within the advised range.
* The kernel may override or ignore the hint for memory pressure, dirty
  limits, fairness, filesystem constraints, or forward progress. It must
  not reserve or pin memory, exempt pages from dirty accounting, or
  promise that intermediate data will never reach storage.
* Explicit synchronization, including msync(MS_SYNC), fsync, fdatasync,
  syncfs, and sync, retains its existing semantics and must not wait for
  the application to revoke the hint. In-flight I/O need not be canceled.
* Shared-memory visibility, data contents, and existing durability
  guarantees remain unchanged. This is not an atomic publication,
  transaction, or crash-consistency mechanism.
* Revoking the hint restores normal scheduling; it does not itself
  require an immediate writeback or imply durability.

Automatic lifetime cleanup is an important part of the request:

* munmap removes this mapping's hint for the unmapped portion.
* Mapping teardown on process exit, including abnormal termination,
  automatically removes its hints. No application cleanup handler is
  required, and no persistent per-file setting should remain behind.
* Independently advised mappings may refer to the same file pages.
  Removing one mapping must not remove another mapping's own hint.
  Once no applicable hints remain, the pages have normal policy.

An initial scope could be ordinary page-cache-backed regular files with
writable MAP_SHARED mappings. Unsupported mappings could be rejected.
Details such as fork inheritance, VMA splits/merges, mremap, and interaction
with unadvised aliases would need explicit definitions; the lifetime rule
should apply consistently to every mapping that actually owns a hint.

Implementation and evaluation questions
---------------------------------------

Although the proposed API is mapping-scoped, writeback acts on file-cache
folios shared by mappings. An implementation would need an inexpensive
way to connect those scopes and retire the hint without leaving stale
writeback state. This request does not assume that adding a VMA flag
alone is sufficient, or mandate a particular accounting mechanism.

Useful evaluation would compare repeated-update construction with and
without the hint, followed by explicit synchronization and a consumer
mmap/read pass. Metrics should include device bytes written, total time
through final synchronization, dirty throttling, peak memory, and consumer
read I/O. Memory-pressure cases, competing writers, partial unmap, and
process termination should also be covered. There are no measured results
attached to this request; the expected benefit is a hypothesis to test.

Is there an existing per-mapping or per-range mechanism that already
provides this behavior? If not, would a best-effort madvise extension be
an appropriate interface, or would an fd/range advisory API fit the
writeback subsystem better while retaining automatic lifetime cleanup?

Pointers to prior discussions and feedback on whether this is worth
prototyping would be appreciated.

This proposal was developed with AI assistance from an application-use
case discussion; no implementation or performance results are claimed.

Regards,
Peng Lei

^ permalink raw reply	[flat|nested] 20+ messages in thread

end of thread, other threads:[~2026-10-01 13:21 UTC | newest]

Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm)
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM)
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm)
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM)
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM)
2026-10-01 13:21                             ` 天狼

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox