* [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression
@ 2026-06-18 8:00 kernel test robot
2026-06-18 9:30 ` Jan Kara
2026-07-23 2:45 ` Bo Zhang
0 siblings, 2 replies; 17+ messages in thread
From: kernel test robot @ 2026-06-18 8:00 UTC (permalink / raw)
To: Frederick Mayle
Cc: oe-lkp, lkp, Andrew Morton, Jan Kara, Kalesh Singh,
David Hildenbrand, Lorenzo Stoakes, Matthew Wilcox,
Suren Baghdasaryan, linux-fsdevel, linux-mm, oliver.sang
Hello,
kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on:
commit: 7b32f64bc512b40b268776c5ac4d354b325b3197 ("mm: limit filemap_fault readahead to VMA boundaries")
https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
[still regression on linux-next/master ec039126b7fac4e3af35ebccaa7c6f9b6875ba81]
testcase: pts
config: x86_64-rhel-9.4
compiler: gcc-14
test machine: 256 threads 2 sockets Intel(R) Xeon(R) 6767P CPU @ 2.4GHz (Granite Rapids) with 256G memory
parameters:
test: svt-av1-2.11.1
option_a: 1
option_b: Bosphorus 4K
cpufreq_governor: performance
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <oliver.sang@intel.com>
| Closes: https://lore.kernel.org/oe-lkp/202606181547.617a6967-lkp@intel.com
Details are as below:
-------------------------------------------------------------------------------------------------->
The kernel config and materials to reproduce are available at:
https://download.01.org/0day-ci/archive/20260618/202606181547.617a6967-lkp@intel.com
=========================================================================================
compiler/cpufreq_governor/kconfig/option_a/option_b/rootfs/tbox_group/test/testcase:
gcc-14/performance/x86_64-rhel-9.4/1/Bosphorus 4K/debian-12-x86_64-phoronix/lkp-gnr-2sp3/svt-av1-2.11.1/pts
commit:
0b20c36c11 ("mm/madvise: reject invalid process_madvise() advice for zero-length vectors")
7b32f64bc5 ("mm: limit filemap_fault readahead to VMA boundaries")
0b20c36c118d2122 7b32f64bc512b40b268776c5ac4
---------------- ---------------------------
%stddev %change %stddev
\ | \
169.95 ± 6% -45.8% 92.15 ± 10% pts.svt-av1.Preset13.Bosphorus4K.frames_per_second
220.57 ± 3% +870.9% 2141 pts.time.major_page_faults
5898121 -3.7% 5682645 pts.time.maximum_resident_set_size
1408522 -1.3% 1390537 pts.time.minor_page_faults
370.43 ± 3% -12.1% 325.57 ± 6% pts.time.percent_of_cpu_this_job_got
645228 ± 2% +7.6% 694407 ± 3% pts.time.voluntary_context_switches
162542 ± 7% +22.4% 198946 ± 8% sched_debug.cpu.avg_idle.stddev
274573 ± 3% -9.3% 249101 ± 3% vmstat.io.bi
6.779e+09 ± 3% +11.3% 7.545e+09 ± 4% cpuidle..time
7639156 ± 3% +10.5% 8439210 ± 3% cpuidle..usage
3.73 ± 39% -2.2 1.52 ± 52% perf-profile.calltrace.cycles-pp.pv_native_safe_halt.acpi_safe_halt.acpi_idle_do_entry.acpi_idle_enter.cpuidle_enter_state
0.65 ± 87% +1.8 2.47 ± 60% perf-profile.children.cycles-pp.link_path_walk
0.24 ± 46% -80.9% 0.05 ±143% perf-sched.sch_delay.avg.ms.perf_trace_sched_switch.do_wait.kernel_wait4.do_syscall_64.entry_SYSCALL_64_after_hwframe
0.65 ± 18% -33.7% 0.43 ± 17% perf-sched.wait_and_delay.avg.ms.perf_trace_sched_switch.do_wait.kernel_wait4.do_syscall_64.entry_SYSCALL_64_after_hwframe
0.02 ± 7% +0.0 0.06 ± 8% mpstat.cpu.all.iowait%
1.08 ± 4% -0.2 0.93 ± 6% mpstat.cpu.all.usr%
22.00 ± 2% +14.3% 25.14 ± 2% mpstat.max_utilization.seconds
10.00 ± 3% -40.5% 5.94 ± 2% mpstat.max_utilization_pct
5916248 +6.9% 6324576 meminfo.Active
3659476 ± 2% +11.4% 4078190 ± 2% meminfo.Active(anon)
760816 ± 3% +16.9% 889730 ± 3% meminfo.AnonHugePages
2263661 ± 4% +18.5% 2682180 ± 3% meminfo.AnonPages
6611024 +11.9% 7400589 ± 2% meminfo.Committed_AS
11386404 +5.1% 11964152 meminfo.Memused
21554 +5.6% 22770 ± 2% meminfo.PageTables
916081 ± 2% +11.4% 1020691 ± 2% proc-vmstat.nr_active_anon
567129 ± 3% +18.4% 671694 ± 3% proc-vmstat.nr_anon_pages
370.88 ± 3% +17.2% 434.57 ± 3% proc-vmstat.nr_anon_transparent_hugepages
48477 +1.6% 49272 proc-vmstat.nr_kernel_stack
5399 +5.4% 5691 ± 2% proc-vmstat.nr_page_table_pages
916081 ± 2% +11.4% 1020691 ± 2% proc-vmstat.nr_zone_active_anon
4614742 -1.5% 4547741 proc-vmstat.pgalloc_normal
1667517 -1.0% 1650517 proc-vmstat.pgfault
2150014 -2.9% 2087812 proc-vmstat.pgfree
224.57 ± 3% +862.1% 2160 proc-vmstat.pgmajfault
1021 -7.2% 948.14 proc-vmstat.thp_fault_alloc
2.679e+09 ± 3% -7.5% 2.477e+09 ± 2% perf-stat.i.branch-instructions
0.80 +0.1 0.86 perf-stat.i.branch-miss-rate%
27722463 ± 3% -8.0% 25498433 ± 2% perf-stat.i.branch-misses
47056892 ± 6% -11.2% 41804370 ± 6% perf-stat.i.cache-misses
397.43 ± 3% -7.4% 368.14 ± 2% perf-stat.i.cpu-migrations
0.03 ± 4% +0.0 0.03 ± 4% perf-stat.i.dTLB-load-miss-rate%
3146082 ± 5% -13.9% 2707716 ± 5% perf-stat.i.dTLB-load-misses
1739613 ± 5% -12.5% 1521663 ± 4% perf-stat.i.dTLB-store-misses
9.97 ± 5% +712.5% 81.01 ± 3% perf-stat.i.major-faults
40.27 ± 4% -8.2% 36.99 ± 3% perf-stat.i.metric.M/sec
58347 ± 4% -11.5% 51627 ± 3% perf-stat.i.minor-faults
63.09 ± 6% -5.8 57.29 ± 8% perf-stat.i.node-load-miss-rate%
6474191 ± 5% -14.1% 5560267 ± 5% perf-stat.i.node-loads
58356 ± 4% -11.4% 51708 ± 3% perf-stat.i.page-faults
0.98 -0.0 0.96 perf-stat.overall.branch-miss-rate%
435.13 ± 3% +10.0% 478.43 ± 4% perf-stat.overall.cycles-between-cache-misses
0.06 ± 2% -0.0 0.06 ± 4% perf-stat.overall.dTLB-load-miss-rate%
0.08 ± 2% -0.0 0.07 ± 2% perf-stat.overall.dTLB-store-miss-rate%
2.646e+09 ± 3% -8.7% 2.415e+09 ± 2% perf-stat.ps.branch-instructions
25808992 ± 3% -10.2% 23188954 ± 2% perf-stat.ps.branch-misses
45057464 ± 5% -14.4% 38574053 ± 6% perf-stat.ps.cache-misses
1.428e+08 ± 3% -9.7% 1.289e+08 ± 2% perf-stat.ps.cache-references
367.11 ± 2% -7.4% 339.78 ± 2% perf-stat.ps.cpu-migrations
2972584 ± 5% -17.3% 2458144 ± 5% perf-stat.ps.dTLB-load-misses
4.801e+09 ± 3% -9.2% 4.359e+09 ± 2% perf-stat.ps.dTLB-loads
1715051 ± 5% -14.7% 1462728 ± 4% perf-stat.ps.dTLB-store-misses
2.193e+09 ± 3% -9.2% 1.992e+09 ± 2% perf-stat.ps.dTLB-stores
2.06e+10 ± 3% -9.2% 1.871e+10 ± 2% perf-stat.ps.instructions
8.54 ± 4% +768.3% 74.19 ± 2% perf-stat.ps.major-faults
59802 ± 3% -11.0% 53218 ± 3% perf-stat.ps.minor-faults
6145046 ± 4% -17.9% 5046537 ± 4% perf-stat.ps.node-loads
59810 ± 3% -10.9% 53292 ± 3% perf-stat.ps.page-faults
Disclaimer:
Results have been estimated based on internal Intel analysis and are provided
for informational purposes only. Any difference in system hardware or software
design or configuration may affect actual performance.
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 8:00 [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression kernel test robot @ 2026-06-18 9:30 ` Jan Kara 2026-06-18 14:30 ` Lorenzo Stoakes 2026-07-23 2:45 ` Bo Zhang 1 sibling, 1 reply; 17+ messages in thread From: Jan Kara @ 2026-06-18 9:30 UTC (permalink / raw) To: kernel test robot Cc: Frederick Mayle, oe-lkp, lkp, Andrew Morton, Jan Kara, Kalesh Singh, David Hildenbrand, Lorenzo Stoakes, Matthew Wilcox, Suren Baghdasaryan, linux-fsdevel, linux-mm On Thu 18-06-26 16:00:42, kernel test robot wrote: > Hello, > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: This one looks serious enough and real. It would be good to figure out what happens in this benchmark that it benefits from the readahead across VMA boundaries so much... Honza > commit: 7b32f64bc512b40b268776c5ac4d354b325b3197 ("mm: limit filemap_fault readahead to VMA boundaries") > https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master > > [still regression on linux-next/master ec039126b7fac4e3af35ebccaa7c6f9b6875ba81] > > testcase: pts > config: x86_64-rhel-9.4 > compiler: gcc-14 > test machine: 256 threads 2 sockets Intel(R) Xeon(R) 6767P CPU @ 2.4GHz (Granite Rapids) with 256G memory > parameters: > > test: svt-av1-2.11.1 > option_a: 1 > option_b: Bosphorus 4K > cpufreq_governor: performance > > > > If you fix the issue in a separate patch/commit (i.e. not just a new version of > the same patch/commit), kindly add following tags > | Reported-by: kernel test robot <oliver.sang@intel.com> > | Closes: https://lore.kernel.org/oe-lkp/202606181547.617a6967-lkp@intel.com > > > Details are as below: > --------------------------------------------------------------------------------------------------> > > > The kernel config and materials to reproduce are available at: > https://download.01.org/0day-ci/archive/20260618/202606181547.617a6967-lkp@intel.com > > ========================================================================================= > compiler/cpufreq_governor/kconfig/option_a/option_b/rootfs/tbox_group/test/testcase: > gcc-14/performance/x86_64-rhel-9.4/1/Bosphorus 4K/debian-12-x86_64-phoronix/lkp-gnr-2sp3/svt-av1-2.11.1/pts > > commit: > 0b20c36c11 ("mm/madvise: reject invalid process_madvise() advice for zero-length vectors") > 7b32f64bc5 ("mm: limit filemap_fault readahead to VMA boundaries") > > 0b20c36c118d2122 7b32f64bc512b40b268776c5ac4 > ---------------- --------------------------- > %stddev %change %stddev > \ | \ > 169.95 ± 6% -45.8% 92.15 ± 10% pts.svt-av1.Preset13.Bosphorus4K.frames_per_second > 220.57 ± 3% +870.9% 2141 pts.time.major_page_faults > 5898121 -3.7% 5682645 pts.time.maximum_resident_set_size > 1408522 -1.3% 1390537 pts.time.minor_page_faults > 370.43 ± 3% -12.1% 325.57 ± 6% pts.time.percent_of_cpu_this_job_got > 645228 ± 2% +7.6% 694407 ± 3% pts.time.voluntary_context_switches > 162542 ± 7% +22.4% 198946 ± 8% sched_debug.cpu.avg_idle.stddev > 274573 ± 3% -9.3% 249101 ± 3% vmstat.io.bi > 6.779e+09 ± 3% +11.3% 7.545e+09 ± 4% cpuidle..time > 7639156 ± 3% +10.5% 8439210 ± 3% cpuidle..usage > 3.73 ± 39% -2.2 1.52 ± 52% perf-profile.calltrace.cycles-pp.pv_native_safe_halt.acpi_safe_halt.acpi_idle_do_entry.acpi_idle_enter.cpuidle_enter_state > 0.65 ± 87% +1.8 2.47 ± 60% perf-profile.children.cycles-pp.link_path_walk > 0.24 ± 46% -80.9% 0.05 ±143% perf-sched.sch_delay.avg.ms.perf_trace_sched_switch.do_wait.kernel_wait4.do_syscall_64.entry_SYSCALL_64_after_hwframe > 0.65 ± 18% -33.7% 0.43 ± 17% perf-sched.wait_and_delay.avg.ms.perf_trace_sched_switch.do_wait.kernel_wait4.do_syscall_64.entry_SYSCALL_64_after_hwframe > 0.02 ± 7% +0.0 0.06 ± 8% mpstat.cpu.all.iowait% > 1.08 ± 4% -0.2 0.93 ± 6% mpstat.cpu.all.usr% > 22.00 ± 2% +14.3% 25.14 ± 2% mpstat.max_utilization.seconds > 10.00 ± 3% -40.5% 5.94 ± 2% mpstat.max_utilization_pct > 5916248 +6.9% 6324576 meminfo.Active > 3659476 ± 2% +11.4% 4078190 ± 2% meminfo.Active(anon) > 760816 ± 3% +16.9% 889730 ± 3% meminfo.AnonHugePages > 2263661 ± 4% +18.5% 2682180 ± 3% meminfo.AnonPages > 6611024 +11.9% 7400589 ± 2% meminfo.Committed_AS > 11386404 +5.1% 11964152 meminfo.Memused > 21554 +5.6% 22770 ± 2% meminfo.PageTables > 916081 ± 2% +11.4% 1020691 ± 2% proc-vmstat.nr_active_anon > 567129 ± 3% +18.4% 671694 ± 3% proc-vmstat.nr_anon_pages > 370.88 ± 3% +17.2% 434.57 ± 3% proc-vmstat.nr_anon_transparent_hugepages > 48477 +1.6% 49272 proc-vmstat.nr_kernel_stack > 5399 +5.4% 5691 ± 2% proc-vmstat.nr_page_table_pages > 916081 ± 2% +11.4% 1020691 ± 2% proc-vmstat.nr_zone_active_anon > 4614742 -1.5% 4547741 proc-vmstat.pgalloc_normal > 1667517 -1.0% 1650517 proc-vmstat.pgfault > 2150014 -2.9% 2087812 proc-vmstat.pgfree > 224.57 ± 3% +862.1% 2160 proc-vmstat.pgmajfault > 1021 -7.2% 948.14 proc-vmstat.thp_fault_alloc > 2.679e+09 ± 3% -7.5% 2.477e+09 ± 2% perf-stat.i.branch-instructions > 0.80 +0.1 0.86 perf-stat.i.branch-miss-rate% > 27722463 ± 3% -8.0% 25498433 ± 2% perf-stat.i.branch-misses > 47056892 ± 6% -11.2% 41804370 ± 6% perf-stat.i.cache-misses > 397.43 ± 3% -7.4% 368.14 ± 2% perf-stat.i.cpu-migrations > 0.03 ± 4% +0.0 0.03 ± 4% perf-stat.i.dTLB-load-miss-rate% > 3146082 ± 5% -13.9% 2707716 ± 5% perf-stat.i.dTLB-load-misses > 1739613 ± 5% -12.5% 1521663 ± 4% perf-stat.i.dTLB-store-misses > 9.97 ± 5% +712.5% 81.01 ± 3% perf-stat.i.major-faults > 40.27 ± 4% -8.2% 36.99 ± 3% perf-stat.i.metric.M/sec > 58347 ± 4% -11.5% 51627 ± 3% perf-stat.i.minor-faults > 63.09 ± 6% -5.8 57.29 ± 8% perf-stat.i.node-load-miss-rate% > 6474191 ± 5% -14.1% 5560267 ± 5% perf-stat.i.node-loads > 58356 ± 4% -11.4% 51708 ± 3% perf-stat.i.page-faults > 0.98 -0.0 0.96 perf-stat.overall.branch-miss-rate% > 435.13 ± 3% +10.0% 478.43 ± 4% perf-stat.overall.cycles-between-cache-misses > 0.06 ± 2% -0.0 0.06 ± 4% perf-stat.overall.dTLB-load-miss-rate% > 0.08 ± 2% -0.0 0.07 ± 2% perf-stat.overall.dTLB-store-miss-rate% > 2.646e+09 ± 3% -8.7% 2.415e+09 ± 2% perf-stat.ps.branch-instructions > 25808992 ± 3% -10.2% 23188954 ± 2% perf-stat.ps.branch-misses > 45057464 ± 5% -14.4% 38574053 ± 6% perf-stat.ps.cache-misses > 1.428e+08 ± 3% -9.7% 1.289e+08 ± 2% perf-stat.ps.cache-references > 367.11 ± 2% -7.4% 339.78 ± 2% perf-stat.ps.cpu-migrations > 2972584 ± 5% -17.3% 2458144 ± 5% perf-stat.ps.dTLB-load-misses > 4.801e+09 ± 3% -9.2% 4.359e+09 ± 2% perf-stat.ps.dTLB-loads > 1715051 ± 5% -14.7% 1462728 ± 4% perf-stat.ps.dTLB-store-misses > 2.193e+09 ± 3% -9.2% 1.992e+09 ± 2% perf-stat.ps.dTLB-stores > 2.06e+10 ± 3% -9.2% 1.871e+10 ± 2% perf-stat.ps.instructions > 8.54 ± 4% +768.3% 74.19 ± 2% perf-stat.ps.major-faults > 59802 ± 3% -11.0% 53218 ± 3% perf-stat.ps.minor-faults > 6145046 ± 4% -17.9% 5046537 ± 4% perf-stat.ps.node-loads > 59810 ± 3% -10.9% 53292 ± 3% perf-stat.ps.page-faults > > > > > Disclaimer: > Results have been estimated based on internal Intel analysis and are provided > for informational purposes only. Any difference in system hardware or software > design or configuration may affect actual performance. > > > -- > 0-DAY CI Kernel Test Service > https://github.com/intel/lkp-tests/wiki > -- Jan Kara <jack@suse.com> SUSE Labs, CR ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 9:30 ` Jan Kara @ 2026-06-18 14:30 ` Lorenzo Stoakes 2026-06-18 16:03 ` Suren Baghdasaryan 2026-06-29 10:35 ` Lorenzo Stoakes 0 siblings, 2 replies; 17+ messages in thread From: Lorenzo Stoakes @ 2026-06-18 14:30 UTC (permalink / raw) To: Jan Kara Cc: kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, Suren Baghdasaryan, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 11:30:47AM +0200, Jan Kara wrote: > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > Hello, > > > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > This one looks serious enough and real. It would be good to figure out what > happens in this benchmark that it benefits from the readahead across VMA > boundaries so much... I think a revert first no? This seems pretty huge for something that isn't key to the kernel, then a new attempt can be tried with this issue addressed perhaps? I seem to recall objecting to this in the past, had a look around and saw [0] which has some discussion I think specifically on the VMA boundaries thing. [0]:https://lore.kernel.org/linux-mm/CAC_TJvfG8GcwG_2w1o6GOTZS8tfEx2h9A91qsenYfYsX8Te=Bg@mail.gmail.com/ (Sorry to be naggy buuut :) - I see there wasn't maintainer [Matthew] signoff on this, we should really be in the habit of _requiring_ that for merge). Thanks, Lorenzo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 14:30 ` Lorenzo Stoakes @ 2026-06-18 16:03 ` Suren Baghdasaryan 2026-06-18 16:32 ` Pedro Falcato 2026-06-29 10:35 ` Lorenzo Stoakes 1 sibling, 1 reply; 17+ messages in thread From: Suren Baghdasaryan @ 2026-06-18 16:03 UTC (permalink / raw) To: Lorenzo Stoakes Cc: Jan Kara, kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 2:30 PM Lorenzo Stoakes <ljs@kernel.org> wrote: > > On Thu, Jun 18, 2026 at 11:30:47AM +0200, Jan Kara wrote: > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > > Hello, > > > > > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > This one looks serious enough and real. It would be good to figure out what > > happens in this benchmark that it benefits from the readahead across VMA > > boundaries so much... > > I think a revert first no? This seems pretty huge for something that isn't key > to the kernel, then a new attempt can be tried with this issue addressed > perhaps? A quick search yields: "The pts.svt-av1.Preset13.Bosphorus4K.frames_per_second is a benchmarking metric from the Phoronix Test Suite that measures how many frames per second a CPU can encode using the open-source SVT-AV1 video encoder." If this is a video encoding benchmark I would expect it to explicitly prefetch the data from the disk before measuring the encoding speed. If limiting readahead caused this regression, I suspect the benchmark doesn't explicitly prefetch the data... > > I seem to recall objecting to this in the past, had a look around and saw [0] > which has some discussion I think specifically on the VMA boundaries thing. > > [0]:https://lore.kernel.org/linux-mm/CAC_TJvfG8GcwG_2w1o6GOTZS8tfEx2h9A91qsenYfYsX8Te=Bg@mail.gmail.com/ > > (Sorry to be naggy buuut :) - I see there wasn't maintainer [Matthew] signoff on > this, we should really be in the habit of _requiring_ that for merge). > > Thanks, Lorenzo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 16:03 ` Suren Baghdasaryan @ 2026-06-18 16:32 ` Pedro Falcato 2026-06-18 16:51 ` Suren Baghdasaryan 2026-06-19 11:11 ` Lorenzo Stoakes 0 siblings, 2 replies; 17+ messages in thread From: Pedro Falcato @ 2026-06-18 16:32 UTC (permalink / raw) To: Suren Baghdasaryan Cc: Lorenzo Stoakes, Jan Kara, kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 04:03:43PM +0000, Suren Baghdasaryan wrote: > On Thu, Jun 18, 2026 at 2:30 PM Lorenzo Stoakes <ljs@kernel.org> wrote: > > > > On Thu, Jun 18, 2026 at 11:30:47AM +0200, Jan Kara wrote: > > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > > > Hello, > > > > > > > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > > > This one looks serious enough and real. It would be good to figure out what > > > happens in this benchmark that it benefits from the readahead across VMA > > > boundaries so much... > > > > I think a revert first no? This seems pretty huge for something that isn't key > > to the kernel, then a new attempt can be tried with this issue addressed > > perhaps? > > A quick search yields: "The > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second is a benchmarking > metric from the Phoronix Test Suite that measures how many frames per > second a CPU can encode using the open-source SVT-AV1 video encoder." > > If this is a video encoding benchmark I would expect it to explicitly > prefetch the data from the disk before measuring the encoding speed. > If limiting readahead caused this regression, I suspect the benchmark > doesn't explicitly prefetch the data... Well, commonly video data doesn't actually fit in memory :) A quick look at the code (I think it's https://gitlab.com/AOMediaCodec/SVT-AV1/-/blob/master/Source/App/app_process_cmd.c#L821) suggests it is progressively mapping the file data for a given frame (or frames?). So the old behavior would result in page faults for a given frame starting readahead for the next few frames. This looks reasonable. FWIW I suspected this was a really weird case regarding mprotect or something, and I'm happy it isn't; but at least I had a suggestion for that - for this, maybe dropping the change (for now?) is the best course of action. -- Pedro ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 16:32 ` Pedro Falcato @ 2026-06-18 16:51 ` Suren Baghdasaryan 2026-06-19 11:11 ` Lorenzo Stoakes 1 sibling, 0 replies; 17+ messages in thread From: Suren Baghdasaryan @ 2026-06-18 16:51 UTC (permalink / raw) To: Pedro Falcato Cc: Lorenzo Stoakes, Jan Kara, kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 9:32 AM Pedro Falcato <pfalcato@suse.de> wrote: > > On Thu, Jun 18, 2026 at 04:03:43PM +0000, Suren Baghdasaryan wrote: > > On Thu, Jun 18, 2026 at 2:30 PM Lorenzo Stoakes <ljs@kernel.org> wrote: > > > > > > On Thu, Jun 18, 2026 at 11:30:47AM +0200, Jan Kara wrote: > > > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > > > > Hello, > > > > > > > > > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > > > > > This one looks serious enough and real. It would be good to figure out what > > > > happens in this benchmark that it benefits from the readahead across VMA > > > > boundaries so much... > > > > > > I think a revert first no? This seems pretty huge for something that isn't key > > > to the kernel, then a new attempt can be tried with this issue addressed > > > perhaps? > > > > A quick search yields: "The > > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second is a benchmarking > > metric from the Phoronix Test Suite that measures how many frames per > > second a CPU can encode using the open-source SVT-AV1 video encoder." > > > > If this is a video encoding benchmark I would expect it to explicitly > > prefetch the data from the disk before measuring the encoding speed. > > If limiting readahead caused this regression, I suspect the benchmark > > doesn't explicitly prefetch the data... > > Well, commonly video data doesn't actually fit in memory :) > > A quick look at the code (I think it's https://gitlab.com/AOMediaCodec/SVT-AV1/-/blob/master/Source/App/app_process_cmd.c#L821) > suggests it is progressively mapping the file data for a given frame > (or frames?). So the old behavior would result in page faults for a given > frame starting readahead for the next few frames. This looks reasonable. Yeah, looks like it's mapping one frame at a time and even that is done with 3 mmap calls - it maps luma_read_size and then 2 chroma_read_size chunks which are consequitive and could have been mapped with one syscall. Seems very inefficient but I guess it's legit. > > FWIW I suspected this was a really weird case regarding mprotect or > something, and I'm happy it isn't; but at least I had a suggestion for that - > for this, maybe dropping the change (for now?) is the best course of action. > > -- > Pedro ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 16:32 ` Pedro Falcato 2026-06-18 16:51 ` Suren Baghdasaryan @ 2026-06-19 11:11 ` Lorenzo Stoakes 1 sibling, 0 replies; 17+ messages in thread From: Lorenzo Stoakes @ 2026-06-19 11:11 UTC (permalink / raw) To: Pedro Falcato Cc: Suren Baghdasaryan, Jan Kara, kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 05:32:31PM +0100, Pedro Falcato wrote: > On Thu, Jun 18, 2026 at 04:03:43PM +0000, Suren Baghdasaryan wrote: > > On Thu, Jun 18, 2026 at 2:30 PM Lorenzo Stoakes <ljs@kernel.org> wrote: > > > > > > On Thu, Jun 18, 2026 at 11:30:47AM +0200, Jan Kara wrote: > > > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > > > > Hello, > > > > > > > > > > kernel test robot noticed a 45.8% regression of pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > > > > > This one looks serious enough and real. It would be good to figure out what > > > > happens in this benchmark that it benefits from the readahead across VMA > > > > boundaries so much... > > > > > > I think a revert first no? This seems pretty huge for something that isn't key > > > to the kernel, then a new attempt can be tried with this issue addressed > > > perhaps? > > > > A quick search yields: "The > > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second is a benchmarking > > metric from the Phoronix Test Suite that measures how many frames per > > second a CPU can encode using the open-source SVT-AV1 video encoder." > > > > If this is a video encoding benchmark I would expect it to explicitly > > prefetch the data from the disk before measuring the encoding speed. > > If limiting readahead caused this regression, I suspect the benchmark > > doesn't explicitly prefetch the data... > > Well, commonly video data doesn't actually fit in memory :) > > A quick look at the code (I think it's https://gitlab.com/AOMediaCodec/SVT-AV1/-/blob/master/Source/App/app_process_cmd.c#L821) > suggests it is progressively mapping the file data for a given frame > (or frames?). So the old behavior would result in page faults for a given > frame starting readahead for the next few frames. This looks reasonable. > > FWIW I suspected this was a really weird case regarding mprotect or > something, and I'm happy it isn't; but at least I had a suggestion for that - > for this, maybe dropping the change (for now?) is the best course of action. Am sending a revert. > > -- > Pedro Thanks, Lorenzo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 14:30 ` Lorenzo Stoakes 2026-06-18 16:03 ` Suren Baghdasaryan @ 2026-06-29 10:35 ` Lorenzo Stoakes 1 sibling, 0 replies; 17+ messages in thread From: Lorenzo Stoakes @ 2026-06-29 10:35 UTC (permalink / raw) To: Jan Kara Cc: kernel test robot, Frederick Mayle, oe-lkp, lkp, Andrew Morton, Kalesh Singh, David Hildenbrand, Matthew Wilcox, Suren Baghdasaryan, linux-fsdevel, linux-mm On Thu, Jun 18, 2026 at 03:30:28PM +0100, Lorenzo Stoakes wrote: > (Sorry to be naggy buuut :) - I see there wasn't maintainer [Matthew] signoff on > this, we should really be in the habit of _requiring_ that for merge). To be clear this is no comment on importance or not of review - all feedback is taken very seriously and I respect Jan's opinions _very_ highly indeed in particular. This was merely an administrative point. Cheers, Lorenzo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-06-18 8:00 [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression kernel test robot 2026-06-18 9:30 ` Jan Kara @ 2026-07-23 2:45 ` Bo Zhang 2026-07-23 20:34 ` Frederick Mayle 2026-07-27 2:17 ` Oliver Sang 1 sibling, 2 replies; 17+ messages in thread From: Bo Zhang @ 2026-07-23 2:45 UTC (permalink / raw) To: oliver.sang Cc: jack, akpm, david, fmayle, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, surenb, willy, baohua, zhangbo56 From: zhangbo56 <zhangbo56@xiaomi.com> On Thu 18-06-26 16:00:42, kernel test robot wrote: > kernel test robot noticed a 45.8% regression of > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > commit: 7b32f64bc512b40b268776c5ac4d354b325b3197 > ("mm: limit filemap_fault readahead to VMA boundaries") > > 169.95 +/- 6% -45.8% 92.15 +/- 10% pts.svt-av1.Preset13.Bosphorus4K.frames_per_second > 220.57 +/- 3% +870.9% 2141 pts.time.major_page_faults Hi Oliver, Could you help test if the below patch fixes the regression? The 870% increase in major faults suggests that readahead is being cut short too aggressively. The current approach unconditionally sets _max_index on every fault, which likely prevents readahead from prefetching ahead effectively for sequential access patterns. The fix: only limit readahead when the fault is close to the VMA end -- if there are fewer pages remaining in the VMA than ra_pages, set _max_index. Otherwise, do nothing. For a 4MB VMA with ra_pages=32, only faults in the last 32 pages (128KB) trigger the limit. The other 99.9% of faults see no change at all. Similarly, only clamp the read-around start when the fault is near the VMA beginning. This applies on top of 7b32f64bc512 and can be applied with git am. Reported-by: kernel test robot <oliver.sang@intel.com> Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com> --- mm/filemap.c | 15 ++++++++++++--- 1 file changed, 12 insertions(+), 3 deletions(-) diff --git a/mm/filemap.c b/mm/filemap.c index 97772a05a18e..9f1e1c9ea6df 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -3313,8 +3313,11 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) vm_flags_t vm_flags = vmf->vma->vm_flags; bool force_thp_readahead = false; unsigned short mmap_miss; + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + vma_pages(vmf->vma); + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; + if (vma_pages_left < ra->ra_pages) + ractl._max_index = vma_end_pgoff - 1; /* Use the readahead code, even if readahead is disabled */ if (IS_ENABLED(CONFIG_TRANSPARENT_HUGEPAGE) && @@ -3398,7 +3401,8 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) * mmap read-around */ ra->start = max_t(long, 0, vmf->pgoff - ra->ra_pages / 2); - ra->start = max(ra->start, vmf->vma->vm_pgoff); + if (vmf->pgoff - vmf->vma->vm_pgoff < ra->ra_pages / 2) + ra->start = max(ra->start, vmf->vma->vm_pgoff); ra->size = ra->ra_pages; ra->async_size = ra->ra_pages / 4; ra->order = 0; @@ -3441,7 +3445,12 @@ static struct file *do_async_mmap_readahead(struct vm_fault *vmf, } if (folio_test_readahead(folio)) { - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + + vma_pages(vmf->vma); + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; + + if (vma_pages_left < ra->ra_pages) + ractl._max_index = vma_end_pgoff - 1; fpin = maybe_unlock_mmap_for_io(vmf, fpin); page_cache_async_ra(&ractl, folio, ra->ra_pages); } -- 2.34.1 ^ permalink raw reply related [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-23 2:45 ` Bo Zhang @ 2026-07-23 20:34 ` Frederick Mayle 2026-07-27 1:46 ` Bo Zhang 2026-07-28 5:25 ` Bo Zhang 2026-07-27 2:17 ` Oliver Sang 1 sibling, 2 replies; 17+ messages in thread From: Frederick Mayle @ 2026-07-23 20:34 UTC (permalink / raw) To: Bo Zhang Cc: oliver.sang, jack, akpm, david, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, surenb, willy, baohua, zhangbo56 On Wed, Jul 22, 2026 at 7:46 PM Bo Zhang <zhangbo0325@gmail.com> wrote: > > From: zhangbo56 <zhangbo56@xiaomi.com> > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > kernel test robot noticed a 45.8% regression of > > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > commit: 7b32f64bc512b40b268776c5ac4d354b325b3197 > > ("mm: limit filemap_fault readahead to VMA boundaries") > > > > 169.95 +/- 6% -45.8% 92.15 +/- 10% pts.svt-av1.Preset13.Bosphorus4K.frames_per_second > > 220.57 +/- 3% +870.9% 2141 pts.time.major_page_faults > > Hi Oliver, > > Could you help test if the below patch fixes the regression? > > The 870% increase in major faults suggests that readahead is being cut > short too aggressively. The current approach unconditionally sets > _max_index on every fault, which likely prevents readahead from > prefetching ahead effectively for sequential access patterns. > > The fix: only limit readahead when the fault is close to the VMA end -- > if there are fewer pages remaining in the VMA than ra_pages, set > _max_index. Otherwise, do nothing. > > For a 4MB VMA with ra_pages=32, only faults in the last 32 pages (128KB) > trigger the limit. The other 99.9% of faults see no change at all. > > Similarly, only clamp the read-around start when the fault is near the > VMA beginning. > > This applies on top of 7b32f64bc512 and can be applied with git am. > > Reported-by: kernel test robot <oliver.sang@intel.com> > Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com> > --- > mm/filemap.c | 15 ++++++++++++--- > 1 file changed, 12 insertions(+), 3 deletions(-) > > diff --git a/mm/filemap.c b/mm/filemap.c > index 97772a05a18e..9f1e1c9ea6df 100644 > --- a/mm/filemap.c > +++ b/mm/filemap.c > @@ -3313,8 +3313,11 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) > vm_flags_t vm_flags = vmf->vma->vm_flags; > bool force_thp_readahead = false; > unsigned short mmap_miss; > + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + vma_pages(vmf->vma); > + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; > > - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; > + if (vma_pages_left < ra->ra_pages) > + ractl._max_index = vma_end_pgoff - 1; > > /* Use the readahead code, even if readahead is disabled */ > if (IS_ENABLED(CONFIG_TRANSPARENT_HUGEPAGE) && > @@ -3398,7 +3401,8 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) > * mmap read-around > */ > ra->start = max_t(long, 0, vmf->pgoff - ra->ra_pages / 2); > - ra->start = max(ra->start, vmf->vma->vm_pgoff); > + if (vmf->pgoff - vmf->vma->vm_pgoff < ra->ra_pages / 2) > + ra->start = max(ra->start, vmf->vma->vm_pgoff); > ra->size = ra->ra_pages; > ra->async_size = ra->ra_pages / 4; > ra->order = 0; > @@ -3441,7 +3445,12 @@ static struct file *do_async_mmap_readahead(struct vm_fault *vmf, > } > > if (folio_test_readahead(folio)) { > - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; > + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + > + vma_pages(vmf->vma); > + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; > + > + if (vma_pages_left < ra->ra_pages) > + ractl._max_index = vma_end_pgoff - 1; > fpin = maybe_unlock_mmap_for_io(vmf, fpin); > page_cache_async_ra(&ractl, folio, ra->ra_pages); > } > -- > 2.34.1 > I maybe missing something subtle, but, I think you've changed it from "don't read beyond the VMA" to "if there is a chance we could read beyond the VMA, don't read beyond the VMA", which seems like a more complex expression of the same behavior. ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-23 20:34 ` Frederick Mayle @ 2026-07-27 1:46 ` Bo Zhang 2026-07-28 10:33 ` Pedro Falcato 2026-07-28 5:25 ` Bo Zhang 1 sibling, 1 reply; 17+ messages in thread From: Bo Zhang @ 2026-07-27 1:46 UTC (permalink / raw) To: fmayle Cc: jack, akpm, david, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, oliver.sang, surenb, willy, baohua, Bo Zhang On Thu, Jul 23, 2026 at 01:34:38PM -0700, Frederick Mayle wrote: > I maybe missing something subtle, but, I think you've changed it from "don't > read beyond the VMA" to "if there is a chance we could read beyond the VMA, > don't read beyond the VMA", which seems like a more complex expression of the > same behavior. Hi Frederick, The subtlety is in how async readahead interacts with _max_index. The two are not equivalent because async readahead sets ra->start to the end of the previous readahead window, not to vmf->pgoff. So even though the user is still faulting well inside the VMA, the readahead target (ra->start) has already moved past the VMA end, however _max_index blocks it. Here is the concrete scenario (ra_pages=128, VMA covers pages 0-1023): 1. Fault at page 0 triggers sync readahead: reads pages 0-127, places PG_readahead marker at page ~96. 2. Sequential faults continue. Fault at page 96 hits PG_readahead, triggers async readahead in do_async_mmap_readahead(). 3. page_cache_async_ra() detects sequential pattern (index == expected) and does: ra->start += ra->size; /* pushes start to end of previous window */ So ra->start = 128 (NOT vmf->pgoff=96), ra->size = 256 which is doubled 4. page_cache_ra_order() calculates: limit = min(file_end, ractl->_max_index) With the unconditional approach: _max_index = 1023 (always set) limit = 1023, ra->start(128) < limit (It works fine here) But after several ramp-ups, when the window reaches the boundary: 5. Fault at page ~896 hits PG_readahead, async readahead fires. page_cache_async_ra() does ra->start += ra->size: ra->start = 1024 (previous window ended at 1023) ra->size = 128 6. page_cache_ra_order(): limit = min(file_end, 1023) = 1023 ra->start(1024) > limit(1023) (Which will read NOTHING) Result: page 1024 (in VMA2) is never prefetched. When the process enters VMA2, it takes a major fault and readahead ramps up from scratch. With my conditional approach: At step 5, vmf->pgoff=896, vma_pages_left = 1024-896 = 128 128 < 128 (the condition is false), so _max_index stays ULONG_MAX ra->start(1024) < ULONG_MAX, and it will prefetch pages 1024-1151 normally. Page 1024 is already cached when VMA2 is entered, only minor fault happens. The key point: when mprotect splits a large file mapping into adjacent VMAs (common for ELF segments, or read-then-write patterns), there's no benefit in preventing readahead from crossing the boundary, so the data is still sequential in the file. The limit only helps when we're actually near the end of useful data (within ra_pages of the VMA end). I confirmed this with a test program that mmap's a 256MB file and uses mprotect to split it into 4MB segments (simulating mprotect-split VMAs): #include <stdio.h> #include <stdlib.h> #include <fcntl.h> #include <sys/mman.h> #include <time.h> #include <unistd.h> #define FILE_SIZE (256UL * 1024 * 1024) #define SEGMENT_SIZE (4UL * 1024 * 1024) int main() { int fd = open("/tmp/testfile", O_RDONLY); char *addr = mmap(NULL, FILE_SIZE, PROT_READ, MAP_PRIVATE, fd, 0); /* Split into 4MB VMAs via alternating mprotect */ for (size_t off = 0; off < FILE_SIZE; off += SEGMENT_SIZE * 2) if (off + SEGMENT_SIZE < FILE_SIZE) mprotect(addr + off + SEGMENT_SIZE, SEGMENT_SIZE, PROT_READ | PROT_WRITE); /* Sequential read across all VMA boundaries */ struct timespec start, end; volatile unsigned long sum = 0; clock_gettime(CLOCK_MONOTONIC, &start); for (size_t i = 0; i < FILE_SIZE; i += 4096) sum += addr[i]; clock_gettime(CLOCK_MONOTONIC, &end); double ms = (end.tv_sec - start.tv_sec) * 1000.0 + (end.tv_nsec - start.tv_nsec) / 1e6; printf("%.1f ms (%.1f MB/s)\n", ms, 256000.0 / ms); munmap(addr, FILE_SIZE); close(fd); } Results on Qualcomm SM8850 mobile phone (ra_pages=128, UFS 4.0), sequential read across 64 mprotect-split VMAs (256MB file, 4MB segments): Without VMA limit (baseline): ~158 ms, ~1620 MB/s With unconditional _max_index (7b32f64bc512): ~180 ms, ~1420 MB/s With conditional _max_index (this fix): ~158 ms, ~1620 MB/s The unconditional approach shows ~13% throughput regression for this workload due to readahead stalling at every 4MB VMA boundary. The conditional approach has zero measurable regression while still limiting readahead for the last ra_pages of each VMA (the original goal). Does this clarify the difference? The conditional check is "don't limit when the async readahead mechanism would set ra->start beyond the VMA boundary and get blocked, even though the actual fault is still well within the VMA." Thanks, Bo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-27 1:46 ` Bo Zhang @ 2026-07-28 10:33 ` Pedro Falcato 2026-07-28 14:09 ` Bo Zhang 0 siblings, 1 reply; 17+ messages in thread From: Pedro Falcato @ 2026-07-28 10:33 UTC (permalink / raw) To: Bo Zhang Cc: fmayle, jack, akpm, david, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, oliver.sang, surenb, willy, baohua, Bo Zhang On Mon, Jul 27, 2026 at 09:46:10AM +0800, Bo Zhang wrote: > On Thu, Jul 23, 2026 at 01:34:38PM -0700, Frederick Mayle wrote: > > I maybe missing something subtle, but, I think you've changed it from "don't > > read beyond the VMA" to "if there is a chance we could read beyond the VMA, > > don't read beyond the VMA", which seems like a more complex expression of the > > same behavior. > > Hi Frederick, > > The subtlety is in how async readahead interacts with _max_index. The two > are not equivalent because async readahead sets ra->start to the end of > the previous readahead window, not to vmf->pgoff. So even though the user > is still faulting well inside the VMA, the readahead target (ra->start) > has already moved past the VMA end, however _max_index blocks it. > > Here is the concrete scenario (ra_pages=128, VMA covers pages 0-1023): > > 1. Fault at page 0 triggers sync readahead: reads pages 0-127, > places PG_readahead marker at page ~96. > > 2. Sequential faults continue. Fault at page 96 hits PG_readahead, > triggers async readahead in do_async_mmap_readahead(). > > 3. page_cache_async_ra() detects sequential pattern (index == expected) > and does: > ra->start += ra->size; /* pushes start to end of previous window */ > So ra->start = 128 (NOT vmf->pgoff=96), ra->size = 256 which is doubled > > 4. page_cache_ra_order() calculates: > limit = min(file_end, ractl->_max_index) > > With the unconditional approach: > _max_index = 1023 (always set) > limit = 1023, ra->start(128) < limit (It works fine here) > > But after several ramp-ups, when the window reaches the boundary: > > 5. Fault at page ~896 hits PG_readahead, async readahead fires. > page_cache_async_ra() does ra->start += ra->size: > ra->start = 1024 (previous window ended at 1023) > ra->size = 128 > > 6. page_cache_ra_order(): > limit = min(file_end, 1023) = 1023 > ra->start(1024) > limit(1023) (Which will read NOTHING) > > Result: page 1024 (in VMA2) is never prefetched. When the process > enters VMA2, it takes a major fault and readahead ramps up from scratch. That's on purpose. > > With my conditional approach: > At step 5, vmf->pgoff=896, vma_pages_left = 1024-896 = 128 > 128 < 128 (the condition is false), so _max_index stays ULONG_MAX > ra->start(1024) < ULONG_MAX, and it will prefetch pages 1024-1151 normally. This is exactly what we're trying to solve. > Page 1024 is already cached when VMA2 is entered, only minor fault happens. > > The key point: when mprotect splits a large file mapping into adjacent > VMAs (common for ELF segments, or read-then-write patterns), there's no > benefit in preventing readahead from crossing the boundary, so the data is > still sequential in the file. The limit only helps when we're actually > near the end of useful data (within ra_pages of the VMA end). This is completely missing the point. This change tried to back off from reading _anything_ past the VMA _on purpose_, due to ELF file spillage. This patch reintroduces spillage in a very obfuscated way. -- Pedro ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-28 10:33 ` Pedro Falcato @ 2026-07-28 14:09 ` Bo Zhang 0 siblings, 0 replies; 17+ messages in thread From: Bo Zhang @ 2026-07-28 14:09 UTC (permalink / raw) To: pfalcato Cc: fmayle, jack, akpm, david, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, oliver.sang, surenb, willy, baohua, Bo Zhang On Tue, Jul 28, 2026 at 11:33:22AM +0100, Pedro Falcato wrote: > > With my conditional approach: > > At step 5, vmf->pgoff=896, vma_pages_left = 1024-896 = 128 > > 128 < 128 (the condition is false), so _max_index stays ULONG_MAX > > This is exactly what we're trying to solve. > > The key point: when mprotect splits a large file mapping into adjacent > > VMAs (common for ELF segments, or read-then-write patterns), there's no > > benefit in preventing readahead from crossing the boundary, so the data is > > still sequential in the file. The limit only helps when we're actually > > near the end of useful data (within ra_pages of the VMA end). > > This is completely missing the point. This change tried to back off from > reading _anything_ past the VMA _on purpose_, due to ELF file spillage. > This patch reintroduces spillage in a very obfuscated way. Hi Pedro, I understand the ELF spillage concern. However, I think it is better to distinguish two different scenarios: 1. ELF segment loading: .text is typically a few hundred KB to a few MB, access is random (e.g., instruction fetches jump around), and the data beyond the segment boundary (e.g., .rodata padding or another segment with different access patterns) is not useful to prefetch. 2. Sequential mmap of contiguous file data: SVT-AV1 maps frames sequentially from the same file, and the data beyond the current VMA boundary is the next thing to be accessed. Same applies to mprotect-split regions of a single large mapping. In your earlier message on Jun 18 [1], you noted that SVT-AV1 "progressively mapping the file data for a given frame" and that the old behavior of "page faults for a given frame starting readahead for the next few frames" looks "reasonable". I totally agree, and my patch tries to preserve exactly that behavior for the sequential case while still limiting spillage for the ELF case. Here's why the conditional check still handles ELF spillage: For a typical ELF .text segment (512KB = 128 pages, ra_pages=128): - First fault: vma_pages_left = 128, 128 < 128 is false, no limit (but ra->size starts small, won't reach boundary yet) - When readahead ramps up and approaches the end: vma_pages_left < 128, _max_index is set, spillage is blocked. For VM_EXEC VMAs, Frederick's patch already has a separate code path that unconditionally limits readahead to VMA boundaries (the exec folio order path). So ELF .text spillage is already handled regardless. The regression comes from the non-VM_EXEC sequential path (mmap read of data files), where blocking cross-VMA readahead has a real cost with no benefit, the data is contiguous and will be needed. The 45.8% fps regression and 870% major fault increase are not acceptable tradeoffs for blocking readahead that would have been useful. My patch preserves the anti-spillage behavior for the last ra_pages of each VMA, while allowing the readahead machinery to work as designed for the bulk of sequential access. [1] https://lore.kernel.org/oe-lkp/ajQbXthzbr9xgUIM@pedro-suse/ Bo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-23 20:34 ` Frederick Mayle 2026-07-27 1:46 ` Bo Zhang @ 2026-07-28 5:25 ` Bo Zhang 1 sibling, 0 replies; 17+ messages in thread From: Bo Zhang @ 2026-07-28 5:25 UTC (permalink / raw) To: fmayle Cc: jack, akpm, david, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, oliver.sang, surenb, willy, baohua, Bo Zhang On Thu, Jul 23, 2026 at 01:34:38PM -0700, Frederick Mayle wrote: > I maybe missing something subtle, but, I think you've changed it from "don't > read beyond the VMA" to "if there is a chance we could read beyond the VMA, > don't read beyond the VMA", which seems like a more complex expression of the > same behavior. Hi Frederick, I looked into the SVT-AV1 benchmark that triggered this regression and found why unconditional _max_index causes the 870% major fault increase. SVT-AV1 uses per-frame mmap on Linux to read its input video file. From Source/App/EncApp/EbAppProcessCmd.c: void read_input_frames(EbConfig *config, ...) { if (config->mmap.enable) { int64_t offset = get_mmap_offset(config, read_size); input_ptr->luma = svt_mmap(&config->mmap, offset, luma_read_size); offset += luma_read_size; input_ptr->cb = svt_mmap(&config->mmap, offset, chroma_read_size); offset += chroma_read_size; input_ptr->cr = svt_mmap(&config->mmap, offset, chroma_read_size); } ... svt_av1_enc_send_picture(component_handle, header_ptr); if (config->mmap.enable) release_memory_mapped_file(config, is_16bit, header_ptr); /* munmap */ } Where svt_mmap() is: void *svt_mmap(MemMapFile *h, int64_t offset, int64_t size) { uint8_t *base = mmap(NULL, size, PROT_READ, MAP_PRIVATE, h->fd, offset); return base + align; } So for Bosphorus 4K (3840x2160 YUV420 8bit), each frame creates 3 separate mmap regions on the SAME file, the layout is: Frame N: [luma 8MB][cb 2MB][cr 2MB] (file offsets contiguous) Frame N+1: [luma 8MB][cb 2MB][cr 2MB] (immediately follows) In page counts (ra_pages=32 assumed): luma VMA: ~2025 pages chroma VMA: ~506 pages each These become 3 independent VMAs per frame, but all map contiguous regions of the same file. The encoder accesses each sequentially, then munmaps before processing the next frame. Here is the problem with unconditional _max_index: 1. Encoder accesses luma VMA sequentially, readahead ramps up. Last async readahead window: ra->start=1993, ra->size=32. 2. Fault at page ~1993 hits PG_readahead, triggers async readahead: ra->start += ra->size; Then ra->start = 2025 3. In page_cache_ra_order(): limit = min(file_end, _max_index); so limit = 2024 while (index <= limit) ... so 2025 > 2024, reads NOTHING 4. Pages 2025+ (which are the cb plane data) are never prefetched. When encoder accesses the cb mmap, it takes a major fault. 5. Same thing happens at cb->cr boundary, and at frame N->N+1. Result: 3 major faults per frame × 120 frames = hundreds of extra major faults, explaining the 870% increase. Without _max_index (original kernel): At step 2, readahead proceeds past page 2025, prefetching the cb data into page cache. When the cb mmap is created and accessed, the pages are already cached, so minor fault happens only. With my conditional patch: At step 2, vmf->pgoff=1993, vma_pages_left = 2025-1993 = 32. The condition "32 < 32" is false, so _max_index stays ULONG_MAX. Readahead proceeds normally, prefetching cb data. Only when vma_pages_left < ra_pages (for example, the last few pages of the VMA), we will set the limit, at that point the readahead window has already pushed useful data into the page cache. For workloads that mmap sequential regions of the same file (SVT-AV1 frames, ELF segment loading, mprotect-split VMAs), the data beyond the VMA boundary is useful and will be accessed soon via a different mapping. Blocking readahead at the VMA boundary in these cases hurts performance. The conditional check preserves the benefit of VMA-bounded readahead (avoiding useless prefetch for small partial mappings) while allowing cross-VMA prefetch when the access is clearly sequential and far from the VMA end. Thanks to Oliver for reporting this regression and providing the test case details, looking into the actual SVT-AV1 code made the mechanism much clearer. Bo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-23 2:45 ` Bo Zhang 2026-07-23 20:34 ` Frederick Mayle @ 2026-07-27 2:17 ` Oliver Sang 2026-07-27 4:20 ` Bo Zhang 1 sibling, 1 reply; 17+ messages in thread From: Oliver Sang @ 2026-07-27 2:17 UTC (permalink / raw) To: Bo Zhang Cc: jack, akpm, david, fmayle, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, surenb, willy, baohua, oliver.sang@intel.com zhangbo56 hi, Bo, On Thu, Jul 23, 2026 at 10:45:27AM +0800, Bo Zhang wrote: > From: zhangbo56 <zhangbo56@xiaomi.com> > > On Thu 18-06-26 16:00:42, kernel test robot wrote: > > kernel test robot noticed a 45.8% regression of > > pts.svt-av1.Preset13.Bosphorus4K.frames_per_second on: > > > > commit: 7b32f64bc512b40b268776c5ac4d354b325b3197 > > ("mm: limit filemap_fault readahead to VMA boundaries") > > > > 169.95 +/- 6% -45.8% 92.15 +/- 10% pts.svt-av1.Preset13.Bosphorus4K.frames_per_second > > 220.57 +/- 3% +870.9% 2141 pts.time.major_page_faults > > Hi Oliver, > > Could you help test if the below patch fixes the regression? quite sorry that we lost the machine we reported this regression, and since pts running need some specific settings in our test framework, we still need some time to setup on other test machines, we cannot start the test soon. will keep you updated. thanks for your paticence and sorry for any inconvenience. > > The 870% increase in major faults suggests that readahead is being cut > short too aggressively. The current approach unconditionally sets > _max_index on every fault, which likely prevents readahead from > prefetching ahead effectively for sequential access patterns. > > The fix: only limit readahead when the fault is close to the VMA end -- > if there are fewer pages remaining in the VMA than ra_pages, set > _max_index. Otherwise, do nothing. > > For a 4MB VMA with ra_pages=32, only faults in the last 32 pages (128KB) > trigger the limit. The other 99.9% of faults see no change at all. > > Similarly, only clamp the read-around start when the fault is near the > VMA beginning. > > This applies on top of 7b32f64bc512 and can be applied with git am. > > Reported-by: kernel test robot <oliver.sang@intel.com> > Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com> > --- > mm/filemap.c | 15 ++++++++++++--- > 1 file changed, 12 insertions(+), 3 deletions(-) > > diff --git a/mm/filemap.c b/mm/filemap.c > index 97772a05a18e..9f1e1c9ea6df 100644 > --- a/mm/filemap.c > +++ b/mm/filemap.c > @@ -3313,8 +3313,11 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) > vm_flags_t vm_flags = vmf->vma->vm_flags; > bool force_thp_readahead = false; > unsigned short mmap_miss; > + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + vma_pages(vmf->vma); > + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; > > - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; > + if (vma_pages_left < ra->ra_pages) > + ractl._max_index = vma_end_pgoff - 1; > > /* Use the readahead code, even if readahead is disabled */ > if (IS_ENABLED(CONFIG_TRANSPARENT_HUGEPAGE) && > @@ -3398,7 +3401,8 @@ static struct file *do_sync_mmap_readahead(struct vm_fault *vmf) > * mmap read-around > */ > ra->start = max_t(long, 0, vmf->pgoff - ra->ra_pages / 2); > - ra->start = max(ra->start, vmf->vma->vm_pgoff); > + if (vmf->pgoff - vmf->vma->vm_pgoff < ra->ra_pages / 2) > + ra->start = max(ra->start, vmf->vma->vm_pgoff); > ra->size = ra->ra_pages; > ra->async_size = ra->ra_pages / 4; > ra->order = 0; > @@ -3441,7 +3445,12 @@ static struct file *do_async_mmap_readahead(struct vm_fault *vmf, > } > > if (folio_test_readahead(folio)) { > - ractl._max_index = vmf->vma->vm_pgoff + vma_pages(vmf->vma) - 1; > + unsigned long vma_end_pgoff = vmf->vma->vm_pgoff + > + vma_pages(vmf->vma); > + unsigned long vma_pages_left = vma_end_pgoff - vmf->pgoff; > + > + if (vma_pages_left < ra->ra_pages) > + ractl._max_index = vma_end_pgoff - 1; > fpin = maybe_unlock_mmap_for_io(vmf, fpin); > page_cache_async_ra(&ractl, folio, ra->ra_pages); > } > -- > 2.34.1 > ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-27 2:17 ` Oliver Sang @ 2026-07-27 4:20 ` Bo Zhang 2026-07-28 2:22 ` Oliver Sang 0 siblings, 1 reply; 17+ messages in thread From: Bo Zhang @ 2026-07-27 4:20 UTC (permalink / raw) To: oliver.sang Cc: jack, akpm, david, fmayle, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, surenb, willy, baohua, Bo Zhang On Mon, Jul 27, 2026 at 10:17:21AM +0800, Oliver Sang wrote: > quite sorry that we lost the machine we reported this regression, and since > pts running need some specific settings in our test framework, we still need > some time to setup on other test machines, we cannot start the test soon. will > keep you updated. Hi Oliver, No worries at all, take your time. I should mention that my patch may not fully fix the regression reported here. It is just based on my analysis of the code path, but I haven't been able to reproduce the exact SVT-AV1 workload. To better understand the issue, would you please share some information about the memory access pattern of the pts.svt-av1 test case? For example: - Does SVT-AV1 mmap the input video file, or use read()? - If mmap'd, are there multiple VMAs over the same file (e.g., from mprotect or partial mappings)? - What is the typical VMA size relative to the file size? - Is the access pattern sequential, random, or a mix? This would help us understand whether the regression comes from the async readahead blocking I described in my reply to Frederick, or from a different interaction with _max_index. Thanks, Bo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression 2026-07-27 4:20 ` Bo Zhang @ 2026-07-28 2:22 ` Oliver Sang 0 siblings, 0 replies; 17+ messages in thread From: Oliver Sang @ 2026-07-28 2:22 UTC (permalink / raw) To: Bo Zhang Cc: jack, akpm, david, fmayle, kaleshsingh, linux-fsdevel, linux-mm, ljs, lkp, oe-lkp, surenb, willy, baohua, oliver.sang@intel.com Bo Zhang hi, Bo, On Mon, Jul 27, 2026 at 12:20:22PM +0800, Bo Zhang wrote: > On Mon, Jul 27, 2026 at 10:17:21AM +0800, Oliver Sang wrote: > > quite sorry that we lost the machine we reported this regression, and since > > pts running need some specific settings in our test framework, we still need > > some time to setup on other test machines, we cannot start the test soon. will > > keep you updated. > > Hi Oliver, > > No worries at all, take your time. > > I should mention that my patch may not fully fix the regression reported > here. It is just based on my analysis of the code path, but I haven't been > able to reproduce the exact SVT-AV1 workload. > > To better understand the issue, would you please share some information > about the memory access pattern of the pts.svt-av1 test case? For example: unfortunately, we just simply integrate "Phoronix Test Suite" [1] into our test framework, and for now we still keep on 9.4.1 version [2] since we don't have enough bandwidth to do upgrade recently. and for svt-av1, we now use 2.11.1 version [3]. we don't have enough knowledge to dig deep for your questions. sorry about this. [1] http://www.phoronix-test-suite.com/ [2] http://phoronix-test-suite.com/releases/phoronix-test-suite-9.4.1.tar.gz [3] https://github.com/phoronix-test-suite/phoronix-test-suite/tree/master/ob-cache/test-profiles/pts/svt-av1-2.11.1 > > - Does SVT-AV1 mmap the input video file, or use read()? > - If mmap'd, are there multiple VMAs over the same file > (e.g., from mprotect or partial mappings)? > - What is the typical VMA size relative to the file size? > - Is the access pattern sequential, random, or a mix? > > This would help us understand whether the regression comes from the > async readahead blocking I described in my reply to Frederick, or from > a different interaction with _max_index. > > Thanks, > Bo ^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-07-28 14:10 UTC | newest] Thread overview: 17+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-06-18 8:00 [linux-next:master] [mm] 7b32f64bc5: pts.svt-av1.Preset13.Bosphorus4K.frames_per_second 45.8% regression kernel test robot 2026-06-18 9:30 ` Jan Kara 2026-06-18 14:30 ` Lorenzo Stoakes 2026-06-18 16:03 ` Suren Baghdasaryan 2026-06-18 16:32 ` Pedro Falcato 2026-06-18 16:51 ` Suren Baghdasaryan 2026-06-19 11:11 ` Lorenzo Stoakes 2026-06-29 10:35 ` Lorenzo Stoakes 2026-07-23 2:45 ` Bo Zhang 2026-07-23 20:34 ` Frederick Mayle 2026-07-27 1:46 ` Bo Zhang 2026-07-28 10:33 ` Pedro Falcato 2026-07-28 14:09 ` Bo Zhang 2026-07-28 5:25 ` Bo Zhang 2026-07-27 2:17 ` Oliver Sang 2026-07-27 4:20 ` Bo Zhang 2026-07-28 2:22 ` Oliver Sang
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox