All of lore.kernel.org
 help / color / mirror / Atom feed
* [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
@ 2024-11-07  4:55 kernel test robot
  2024-11-07 11:35 ` Jeff Layton
  2024-11-14 14:37 ` Chuck Lever
  0 siblings, 2 replies; 19+ messages in thread
From: kernel test robot @ 2024-11-07  4:55 UTC (permalink / raw)
  To: Jeff Layton; +Cc: Chuck Lever, Mike Snitzer, oe-lkp, lkp, oliver.sang


hi, Jeff Layton,

in commit message, it is mentioned the change is expected to solve the
"App Overhead" on the fs_mark test we reported in
https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/

however, in our tests, there is sill similar regression. at the same
time, there is still no performance difference for fsmark.files_per_sec

   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
     18.57            +0.0%      18.57        fsmark.files_per_sec


another thing is our bot bisect to this commit in repo/branch as below detail
information. if there is a more porper repo/branch to test the patch, could
you let us know? thanks a lot!



Hello,

kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:


commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5


testcase: fsmark
config: x86_64-rhel-8.3
compiler: gcc-12
test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
parameters:

	iterations: 1x
	nr_threads: 1t
	disk: 1HDD
	fs: btrfs
	fs2: nfsv4
	filesize: 4K
	test_size: 40M
	sync_method: fsyncBeforeClose
	nr_files_per_directory: 1fpd
	cpufreq_governor: performance




If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <oliver.sang@intel.com>
| Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com


Details are as below:
-------------------------------------------------------------------------------------------------->


The kernel config and materials to reproduce are available at:
https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com

=========================================================================================
compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
  gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark

commit: 
  37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
  b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")

37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
---------------- --------------------------- 
         %stddev     %change         %stddev
             \          |                \  
     97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
      3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
      2936            -6.2%       2755        vmstat.system.cs
   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
     18.57            +0.0%      18.57        fsmark.files_per_sec
     53420           -17.3%      44185        fsmark.time.voluntary_context_switches
      1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
      3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
   1756957            -2.1%    1720536        proc-vmstat.numa_hit
   1624496            -2.2%    1588039        proc-vmstat.numa_local
      1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
      2916            -6.2%       2735        perf-stat.i.context-switches
      1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
      2910            -6.2%       2729        perf-stat.ps.context-switches
      0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
      0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
      0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
      0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
      0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
      1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
      2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
      2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
      2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
      0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
      0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
      0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
      0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
      0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
      0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
      0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
      0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
      0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
      0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
      0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
      0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
      1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
      1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
      0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events




Disclaimer:
Results have been estimated based on internal Intel analysis and are provided
for informational purposes only. Any difference in system hardware or software
design or configuration may affect actual performance.


-- 
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki


^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-07  4:55 [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression kernel test robot
@ 2024-11-07 11:35 ` Jeff Layton
  2024-11-13 15:48   ` Chuck Lever
  2024-11-14 14:37 ` Chuck Lever
  1 sibling, 1 reply; 19+ messages in thread
From: Jeff Layton @ 2024-11-07 11:35 UTC (permalink / raw)
  To: kernel test robot; +Cc: Chuck Lever, Mike Snitzer, oe-lkp, lkp

On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> hi, Jeff Layton,
> 
> in commit message, it is mentioned the change is expected to solve the
> "App Overhead" on the fs_mark test we reported in
> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> 
> however, in our tests, there is sill similar regression. at the same
> time, there is still no performance difference for fsmark.files_per_sec
> 
>    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>      18.57            +0.0%      18.57        fsmark.files_per_sec
> 
> 
> another thing is our bot bisect to this commit in repo/branch as below detail
> information. if there is a more porper repo/branch to test the patch, could
> you let us know? thanks a lot!
> 
> 
> 
> Hello,
> 
> kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> 
> 
> commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> 
> 
> testcase: fsmark
> config: x86_64-rhel-8.3
> compiler: gcc-12
> test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> parameters:
> 
> 	iterations: 1x
> 	nr_threads: 1t
> 	disk: 1HDD
> 	fs: btrfs
> 	fs2: nfsv4
> 	filesize: 4K
> 	test_size: 40M
> 	sync_method: fsyncBeforeClose
> 	nr_files_per_directory: 1fpd
> 	cpufreq_governor: performance
> 
> 
> 
> 
> If you fix the issue in a separate patch/commit (i.e. not just a new version of
> the same patch/commit), kindly add following tags
> > Reported-by: kernel test robot <oliver.sang@intel.com>
> > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> 
> 
> Details are as below:
> -------------------------------------------------------------------------------------------------->
> 
> 
> The kernel config and materials to reproduce are available at:
> https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> 
> =========================================================================================
> compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
>   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> 
> commit: 
>   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
>   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> 
> 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> ---------------- --------------------------- 
>          %stddev     %change         %stddev
>              \          |                \  
>      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
>       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
>       2936            -6.2%       2755        vmstat.system.cs
>    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>      18.57            +0.0%      18.57        fsmark.files_per_sec
>      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
>       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>    1756957            -2.1%    1720536        proc-vmstat.numa_hit
>    1624496            -2.2%    1588039        proc-vmstat.numa_local
>       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
>       2916            -6.2%       2735        perf-stat.i.context-switches
>       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
>       2910            -6.2%       2729        perf-stat.ps.context-switches
>       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
>       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
>       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
>       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
>       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
>       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
>       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
>       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
>       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
>       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
>       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
>       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
>       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
>       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
>       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
>       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
>       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
>       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
>       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
>       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
>       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
>       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
>       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
>       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> 
> 
> 
> 
> Disclaimer:
> Results have been estimated based on internal Intel analysis and are provided
> for informational purposes only. Any difference in system hardware or software
> design or configuration may affect actual performance.
> 
> 

This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
a performance regression in the server. The only thing I can figure
here is that this test is causing the server to recall the delegation
that it hands out, and then the client has to go and establish a new
open stateid in order to return it. That would likely be slower than
just handing out both an open and delegation stateid in the first
place.

I don't think there is anything we can do about that though, since the
feature seems to is working as designed.
-- 
Jeff Layton <jlayton@kernel.org>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-07 11:35 ` Jeff Layton
@ 2024-11-13 15:48   ` Chuck Lever
  2024-11-13 16:10     ` Jeff Layton
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2024-11-13 15:48 UTC (permalink / raw)
  To: Jeff Layton; +Cc: kernel test robot, Mike Snitzer, oe-lkp, lkp

On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
> On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> > hi, Jeff Layton,
> > 
> > in commit message, it is mentioned the change is expected to solve the
> > "App Overhead" on the fs_mark test we reported in
> > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > 
> > however, in our tests, there is sill similar regression. at the same
> > time, there is still no performance difference for fsmark.files_per_sec
> > 
> >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > 
> > 
> > another thing is our bot bisect to this commit in repo/branch as below detail
> > information. if there is a more porper repo/branch to test the patch, could
> > you let us know? thanks a lot!
> > 
> > 
> > 
> > Hello,
> > 
> > kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> > 
> > 
> > commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> > 
> > 
> > testcase: fsmark
> > config: x86_64-rhel-8.3
> > compiler: gcc-12
> > test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> > parameters:
> > 
> > 	iterations: 1x
> > 	nr_threads: 1t
> > 	disk: 1HDD
> > 	fs: btrfs
> > 	fs2: nfsv4
> > 	filesize: 4K
> > 	test_size: 40M
> > 	sync_method: fsyncBeforeClose
> > 	nr_files_per_directory: 1fpd
> > 	cpufreq_governor: performance
> > 
> > 
> > 
> > 
> > If you fix the issue in a separate patch/commit (i.e. not just a new version of
> > the same patch/commit), kindly add following tags
> > > Reported-by: kernel test robot <oliver.sang@intel.com>
> > > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> > 
> > 
> > Details are as below:
> > -------------------------------------------------------------------------------------------------->
> > 
> > 
> > The kernel config and materials to reproduce are available at:
> > https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> > 
> > =========================================================================================
> > compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
> >   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> > 
> > commit: 
> >   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
> >   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > 
> > 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> > ---------------- --------------------------- 
> >          %stddev     %change         %stddev
> >              \          |                \  
> >      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
> >       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
> >       2936            -6.2%       2755        vmstat.system.cs
> >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> >      18.57            +0.0%      18.57        fsmark.files_per_sec
> >      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
> >       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> >       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> >    1756957            -2.1%    1720536        proc-vmstat.numa_hit
> >    1624496            -2.2%    1588039        proc-vmstat.numa_local
> >       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
> >       2916            -6.2%       2735        perf-stat.i.context-switches
> >       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
> >       2910            -6.2%       2729        perf-stat.ps.context-switches
> >       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
> >       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
> >       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
> >       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
> >       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
> >       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
> >       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> >       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
> >       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> >       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
> >       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
> >       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
> >       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
> >       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
> >       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
> >       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
> >       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
> >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
> >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
> >       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
> >       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
> >       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
> >       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
> >       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
> >       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
> >       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> > 
> > 
> > 
> > 
> > Disclaimer:
> > Results have been estimated based on internal Intel analysis and are provided
> > for informational purposes only. Any difference in system hardware or software
> > design or configuration may affect actual performance.
> > 
> > 
> 
> This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
> a performance regression in the server. The only thing I can figure
> here is that this test is causing the server to recall the delegation
> that it hands out, and then the client has to go and establish a new
> open stateid in order to return it. That would likely be slower than
> just handing out both an open and delegation stateid in the first
> place.
> 
> I don't think there is anything we can do about that though, since the
> feature seems to is working as designed.

We seem to have hit this problem before. That makes me wonder
whether it is actually worth supporting the XOR flag at all. After
all, the client sends the CLOSE asynchronously; applications are not
being held up while the unneeded state ID is returned.

Can XOR support be disabled for now? I don't want to add an
administrative interface for that, but also, "no regressions" is
ringing in my ears, and 92% is a mighty noticeable one.


-- 
Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-13 15:48   ` Chuck Lever
@ 2024-11-13 16:10     ` Jeff Layton
  2024-11-13 16:14       ` Chuck Lever
  0 siblings, 1 reply; 19+ messages in thread
From: Jeff Layton @ 2024-11-13 16:10 UTC (permalink / raw)
  To: Chuck Lever
  Cc: kernel test robot, Mike Snitzer, oe-lkp, lkp, Anna Schumaker,
	Trond Myklebust, Thomas Haynes, linux-nfs

On Wed, 2024-11-13 at 10:48 -0500, Chuck Lever wrote:
> On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
> > On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> > > hi, Jeff Layton,
> > > 
> > > in commit message, it is mentioned the change is expected to solve the
> > > "App Overhead" on the fs_mark test we reported in
> > > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > > 
> > > however, in our tests, there is sill similar regression. at the same
> > > time, there is still no performance difference for fsmark.files_per_sec
> > > 
> > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > > 
> > > 
> > > another thing is our bot bisect to this commit in repo/branch as below detail
> > > information. if there is a more porper repo/branch to test the patch, could
> > > you let us know? thanks a lot!
> > > 
> > > 
> > > 
> > > Hello,
> > > 
> > > kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> > > 
> > > 
> > > commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> > > 
> > > 
> > > testcase: fsmark
> > > config: x86_64-rhel-8.3
> > > compiler: gcc-12
> > > test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> > > parameters:
> > > 
> > > 	iterations: 1x
> > > 	nr_threads: 1t
> > > 	disk: 1HDD
> > > 	fs: btrfs
> > > 	fs2: nfsv4
> > > 	filesize: 4K
> > > 	test_size: 40M
> > > 	sync_method: fsyncBeforeClose
> > > 	nr_files_per_directory: 1fpd
> > > 	cpufreq_governor: performance
> > > 
> > > 
> > > 
> > > 
> > > If you fix the issue in a separate patch/commit (i.e. not just a new version of
> > > the same patch/commit), kindly add following tags
> > > > Reported-by: kernel test robot <oliver.sang@intel.com>
> > > > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > 
> > > 
> > > Details are as below:
> > > -------------------------------------------------------------------------------------------------->
> > > 
> > > 
> > > The kernel config and materials to reproduce are available at:
> > > https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > 
> > > =========================================================================================
> > > compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
> > >   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> > > 
> > > commit: 
> > >   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
> > >   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > 
> > > 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> > > ---------------- --------------------------- 
> > >          %stddev     %change         %stddev
> > >              \          |                \  
> > >      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
> > >       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
> > >       2936            -6.2%       2755        vmstat.system.cs
> > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > >      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
> > >       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > >       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > >    1756957            -2.1%    1720536        proc-vmstat.numa_hit
> > >    1624496            -2.2%    1588039        proc-vmstat.numa_local
> > >       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
> > >       2916            -6.2%       2735        perf-stat.i.context-switches
> > >       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
> > >       2910            -6.2%       2729        perf-stat.ps.context-switches
> > >       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
> > >       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > >       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
> > >       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > >       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
> > >       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
> > >       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > >       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
> > >       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > >       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
> > >       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
> > >       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
> > >       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
> > >       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
> > >       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
> > >       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
> > >       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
> > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
> > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
> > >       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
> > >       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
> > >       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
> > >       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
> > >       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
> > >       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
> > >       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> > > 
> > > 
> > > 
> > > 
> > > Disclaimer:
> > > Results have been estimated based on internal Intel analysis and are provided
> > > for informational purposes only. Any difference in system hardware or software
> > > design or configuration may affect actual performance.
> > > 
> > > 
> > 
> > This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
> > a performance regression in the server. The only thing I can figure
> > here is that this test is causing the server to recall the delegation
> > that it hands out, and then the client has to go and establish a new
> > open stateid in order to return it. That would likely be slower than
> > just handing out both an open and delegation stateid in the first
> > place.
> > 
> > I don't think there is anything we can do about that though, since the
> > feature seems to is working as designed.
> 
> We seem to have hit this problem before. That makes me wonder
> whether it is actually worth supporting the XOR flag at all. After
> all, the client sends the CLOSE asynchronously; applications are not
> being held up while the unneeded state ID is returned.
> 
> Can XOR support be disabled for now? I don't want to add an
> administrative interface for that, but also, "no regressions" is
> ringing in my ears, and 92% is a mighty noticeable one.
> 
> 

To be clear, this increase is for the "App Overhead" which is all of
the operations between the stuff that is being measured in this test. I
did run this test for a bit and got similar results, but was never able
to nail down where the overhead came from. My speculation is that it's
the recall and reestablishment of an open stateid that slows things
down, but I never could fully confirm it.

My issue with disabling this is that the decision of whether to set
OPEN_XOR_DELEGATION is up to the client. The server in this case is
just doing what the client asks. ISTM that if we were going to disable
(or throttle) this anywhere, that should be done by the client.

-- 
Jeff Layton <jlayton@kernel.org>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-13 16:10     ` Jeff Layton
@ 2024-11-13 16:14       ` Chuck Lever
  2024-11-13 16:20         ` Jeff Layton
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2024-11-13 16:14 UTC (permalink / raw)
  To: Jeff Layton
  Cc: kernel test robot, Mike Snitzer, oe-lkp, lkp, Anna Schumaker,
	Trond Myklebust, Thomas Haynes, linux-nfs

On Wed, Nov 13, 2024 at 11:10:35AM -0500, Jeff Layton wrote:
> On Wed, 2024-11-13 at 10:48 -0500, Chuck Lever wrote:
> > On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
> > > On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> > > > hi, Jeff Layton,
> > > > 
> > > > in commit message, it is mentioned the change is expected to solve the
> > > > "App Overhead" on the fs_mark test we reported in
> > > > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > > > 
> > > > however, in our tests, there is sill similar regression. at the same
> > > > time, there is still no performance difference for fsmark.files_per_sec
> > > > 
> > > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > 
> > > > 
> > > > another thing is our bot bisect to this commit in repo/branch as below detail
> > > > information. if there is a more porper repo/branch to test the patch, could
> > > > you let us know? thanks a lot!
> > > > 
> > > > 
> > > > 
> > > > Hello,
> > > > 
> > > > kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> > > > 
> > > > 
> > > > commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> > > > 
> > > > 
> > > > testcase: fsmark
> > > > config: x86_64-rhel-8.3
> > > > compiler: gcc-12
> > > > test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> > > > parameters:
> > > > 
> > > > 	iterations: 1x
> > > > 	nr_threads: 1t
> > > > 	disk: 1HDD
> > > > 	fs: btrfs
> > > > 	fs2: nfsv4
> > > > 	filesize: 4K
> > > > 	test_size: 40M
> > > > 	sync_method: fsyncBeforeClose
> > > > 	nr_files_per_directory: 1fpd
> > > > 	cpufreq_governor: performance
> > > > 
> > > > 
> > > > 
> > > > 
> > > > If you fix the issue in a separate patch/commit (i.e. not just a new version of
> > > > the same patch/commit), kindly add following tags
> > > > > Reported-by: kernel test robot <oliver.sang@intel.com>
> > > > > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > 
> > > > 
> > > > Details are as below:
> > > > -------------------------------------------------------------------------------------------------->
> > > > 
> > > > 
> > > > The kernel config and materials to reproduce are available at:
> > > > https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > 
> > > > =========================================================================================
> > > > compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
> > > >   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> > > > 
> > > > commit: 
> > > >   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
> > > >   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > 
> > > > 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> > > > ---------------- --------------------------- 
> > > >          %stddev     %change         %stddev
> > > >              \          |                \  
> > > >      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
> > > >       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
> > > >       2936            -6.2%       2755        vmstat.system.cs
> > > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > > >      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
> > > >       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > >       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > >    1756957            -2.1%    1720536        proc-vmstat.numa_hit
> > > >    1624496            -2.2%    1588039        proc-vmstat.numa_local
> > > >       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
> > > >       2916            -6.2%       2735        perf-stat.i.context-switches
> > > >       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
> > > >       2910            -6.2%       2729        perf-stat.ps.context-switches
> > > >       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
> > > >       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > >       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
> > > >       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > >       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
> > > >       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
> > > >       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > >       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
> > > >       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > >       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
> > > >       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
> > > >       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
> > > >       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
> > > >       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
> > > >       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
> > > >       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
> > > >       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
> > > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
> > > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
> > > >       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
> > > >       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
> > > >       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
> > > >       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
> > > >       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
> > > >       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
> > > >       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> > > > 
> > > > 
> > > > 
> > > > 
> > > > Disclaimer:
> > > > Results have been estimated based on internal Intel analysis and are provided
> > > > for informational purposes only. Any difference in system hardware or software
> > > > design or configuration may affect actual performance.
> > > > 
> > > > 
> > > 
> > > This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
> > > a performance regression in the server. The only thing I can figure
> > > here is that this test is causing the server to recall the delegation
> > > that it hands out, and then the client has to go and establish a new
> > > open stateid in order to return it. That would likely be slower than
> > > just handing out both an open and delegation stateid in the first
> > > place.
> > > 
> > > I don't think there is anything we can do about that though, since the
> > > feature seems to is working as designed.
> > 
> > We seem to have hit this problem before. That makes me wonder
> > whether it is actually worth supporting the XOR flag at all. After
> > all, the client sends the CLOSE asynchronously; applications are not
> > being held up while the unneeded state ID is returned.
> > 
> > Can XOR support be disabled for now? I don't want to add an
> > administrative interface for that, but also, "no regressions" is
> > ringing in my ears, and 92% is a mighty noticeable one.
> 
> To be clear, this increase is for the "App Overhead" which is all of
> the operations between the stuff that is being measured in this test. I
> did run this test for a bit and got similar results, but was never able
> to nail down where the overhead came from. My speculation is that it's
> the recall and reestablishment of an open stateid that slows things
> down, but I never could fully confirm it.
> 
> My issue with disabling this is that the decision of whether to set
> OPEN_XOR_DELEGATION is up to the client. The server in this case is
> just doing what the client asks. ISTM that if we were going to disable
> (or throttle) this anywhere, that should be done by the client.

If the client sets the XOR flag, NFSD could simply return only an
OPEN stateid, for instance.


-- 
Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-13 16:14       ` Chuck Lever
@ 2024-11-13 16:20         ` Jeff Layton
  2024-11-13 16:37           ` Chuck Lever III
  0 siblings, 1 reply; 19+ messages in thread
From: Jeff Layton @ 2024-11-13 16:20 UTC (permalink / raw)
  To: Chuck Lever
  Cc: kernel test robot, Mike Snitzer, oe-lkp, lkp, Anna Schumaker,
	Trond Myklebust, Thomas Haynes, linux-nfs

On Wed, 2024-11-13 at 11:14 -0500, Chuck Lever wrote:
> On Wed, Nov 13, 2024 at 11:10:35AM -0500, Jeff Layton wrote:
> > On Wed, 2024-11-13 at 10:48 -0500, Chuck Lever wrote:
> > > On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
> > > > On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> > > > > hi, Jeff Layton,
> > > > > 
> > > > > in commit message, it is mentioned the change is expected to solve the
> > > > > "App Overhead" on the fs_mark test we reported in
> > > > > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > > > > 
> > > > > however, in our tests, there is sill similar regression. at the same
> > > > > time, there is still no performance difference for fsmark.files_per_sec
> > > > > 
> > > > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > > 
> > > > > 
> > > > > another thing is our bot bisect to this commit in repo/branch as below detail
> > > > > information. if there is a more porper repo/branch to test the patch, could
> > > > > you let us know? thanks a lot!
> > > > > 
> > > > > 
> > > > > 
> > > > > Hello,
> > > > > 
> > > > > kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> > > > > 
> > > > > 
> > > > > commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > > https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> > > > > 
> > > > > 
> > > > > testcase: fsmark
> > > > > config: x86_64-rhel-8.3
> > > > > compiler: gcc-12
> > > > > test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> > > > > parameters:
> > > > > 
> > > > > 	iterations: 1x
> > > > > 	nr_threads: 1t
> > > > > 	disk: 1HDD
> > > > > 	fs: btrfs
> > > > > 	fs2: nfsv4
> > > > > 	filesize: 4K
> > > > > 	test_size: 40M
> > > > > 	sync_method: fsyncBeforeClose
> > > > > 	nr_files_per_directory: 1fpd
> > > > > 	cpufreq_governor: performance
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > If you fix the issue in a separate patch/commit (i.e. not just a new version of
> > > > > the same patch/commit), kindly add following tags
> > > > > > Reported-by: kernel test robot <oliver.sang@intel.com>
> > > > > > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > > 
> > > > > 
> > > > > Details are as below:
> > > > > -------------------------------------------------------------------------------------------------->
> > > > > 
> > > > > 
> > > > > The kernel config and materials to reproduce are available at:
> > > > > https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > > 
> > > > > =========================================================================================
> > > > > compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
> > > > >   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> > > > > 
> > > > > commit: 
> > > > >   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
> > > > >   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > > 
> > > > > 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> > > > > ---------------- --------------------------- 
> > > > >          %stddev     %change         %stddev
> > > > >              \          |                \  
> > > > >      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
> > > > >       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
> > > > >       2936            -6.2%       2755        vmstat.system.cs
> > > > >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > > >      18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > >      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
> > > > >       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > > >       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > > >    1756957            -2.1%    1720536        proc-vmstat.numa_hit
> > > > >    1624496            -2.2%    1588039        proc-vmstat.numa_local
> > > > >       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
> > > > >       2916            -6.2%       2735        perf-stat.i.context-switches
> > > > >       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
> > > > >       2910            -6.2%       2729        perf-stat.ps.context-switches
> > > > >       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
> > > > >       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > > >       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
> > > > >       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > > >       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
> > > > >       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
> > > > >       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > > >       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
> > > > >       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > > >       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
> > > > >       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
> > > > >       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
> > > > >       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
> > > > >       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
> > > > >       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
> > > > >       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
> > > > >       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
> > > > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
> > > > >       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
> > > > >       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
> > > > >       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
> > > > >       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
> > > > >       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
> > > > >       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
> > > > >       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
> > > > >       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > Disclaimer:
> > > > > Results have been estimated based on internal Intel analysis and are provided
> > > > > for informational purposes only. Any difference in system hardware or software
> > > > > design or configuration may affect actual performance.
> > > > > 
> > > > > 
> > > > 
> > > > This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
> > > > a performance regression in the server. The only thing I can figure
> > > > here is that this test is causing the server to recall the delegation
> > > > that it hands out, and then the client has to go and establish a new
> > > > open stateid in order to return it. That would likely be slower than
> > > > just handing out both an open and delegation stateid in the first
> > > > place.
> > > > 
> > > > I don't think there is anything we can do about that though, since the
> > > > feature seems to is working as designed.
> > > 
> > > We seem to have hit this problem before. That makes me wonder
> > > whether it is actually worth supporting the XOR flag at all. After
> > > all, the client sends the CLOSE asynchronously; applications are not
> > > being held up while the unneeded state ID is returned.
> > > 
> > > Can XOR support be disabled for now? I don't want to add an
> > > administrative interface for that, but also, "no regressions" is
> > > ringing in my ears, and 92% is a mighty noticeable one.
> > 
> > To be clear, this increase is for the "App Overhead" which is all of
> > the operations between the stuff that is being measured in this test. I
> > did run this test for a bit and got similar results, but was never able
> > to nail down where the overhead came from. My speculation is that it's
> > the recall and reestablishment of an open stateid that slows things
> > down, but I never could fully confirm it.
> > 
> > My issue with disabling this is that the decision of whether to set
> > OPEN_XOR_DELEGATION is up to the client. The server in this case is
> > just doing what the client asks. ISTM that if we were going to disable
> > (or throttle) this anywhere, that should be done by the client.
> 
> If the client sets the XOR flag, NFSD could simply return only an
> OPEN stateid, for instance.
> 
> 

Sure, but then it's disabled for everyone, for every workload. There
may be workloads the benefit too. What would really be nice is another
datapoint. To the HS guys:

If you run this test against the Anvil both with and without
OPEN_XOR_DELEGATION enabled, do you also see an increase in the "App
Overhead" with this enabled? It would be nice to know if this effect is
common among servers that implement this, or whether it's something
particular to knfsd.

-- 
Jeff Layton <jlayton@kernel.org>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-13 16:20         ` Jeff Layton
@ 2024-11-13 16:37           ` Chuck Lever III
  2024-11-13 17:02             ` Jeff Layton
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever III @ 2024-11-13 16:37 UTC (permalink / raw)
  To: Jeff Layton
  Cc: kernel test robot, Mike Snitzer, oe-lkp@lists.linux.dev,
	kernel test robot, Anna Schumaker, Trond Myklebust, Tom Haynes,
	Linux NFS Mailing List



> On Nov 13, 2024, at 11:20 AM, Jeff Layton <jlayton@kernel.org> wrote:
> 
> On Wed, 2024-11-13 at 11:14 -0500, Chuck Lever wrote:
>> On Wed, Nov 13, 2024 at 11:10:35AM -0500, Jeff Layton wrote:
>>> On Wed, 2024-11-13 at 10:48 -0500, Chuck Lever wrote:
>>>> On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
>>>>> On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
>>>>>> hi, Jeff Layton,
>>>>>> 
>>>>>> in commit message, it is mentioned the change is expected to solve the
>>>>>> "App Overhead" on the fs_mark test we reported in
>>>>>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
>>>>>> 
>>>>>> however, in our tests, there is sill similar regression. at the same
>>>>>> time, there is still no performance difference for fsmark.files_per_sec
>>>>>> 
>>>>>>   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>>>>>>     18.57            +0.0%      18.57        fsmark.files_per_sec
>>>>>> 
>>>>>> 
>>>>>> another thing is our bot bisect to this commit in repo/branch as below detail
>>>>>> information. if there is a more porper repo/branch to test the patch, could
>>>>>> you let us know? thanks a lot!
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> Hello,
>>>>>> 
>>>>>> kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
>>>>>> 
>>>>>> 
>>>>>> commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
>>>>>> https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
>>>>>> 
>>>>>> 
>>>>>> testcase: fsmark
>>>>>> config: x86_64-rhel-8.3
>>>>>> compiler: gcc-12
>>>>>> test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
>>>>>> parameters:
>>>>>> 
>>>>>> iterations: 1x
>>>>>> nr_threads: 1t
>>>>>> disk: 1HDD
>>>>>> fs: btrfs
>>>>>> fs2: nfsv4
>>>>>> filesize: 4K
>>>>>> test_size: 40M
>>>>>> sync_method: fsyncBeforeClose
>>>>>> nr_files_per_directory: 1fpd
>>>>>> cpufreq_governor: performance
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> If you fix the issue in a separate patch/commit (i.e. not just a new version of
>>>>>> the same patch/commit), kindly add following tags
>>>>>>> Reported-by: kernel test robot <oliver.sang@intel.com>
>>>>>>> Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
>>>>>> 
>>>>>> 
>>>>>> Details are as below:
>>>>>> -------------------------------------------------------------------------------------------------->
>>>>>> 
>>>>>> 
>>>>>> The kernel config and materials to reproduce are available at:
>>>>>> https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
>>>>>> 
>>>>>> =========================================================================================
>>>>>> compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
>>>>>>  gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
>>>>>> 
>>>>>> commit: 
>>>>>>  37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
>>>>>>  b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
>>>>>> 
>>>>>> 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
>>>>>> ---------------- --------------------------- 
>>>>>>         %stddev     %change         %stddev
>>>>>>             \          |                \  
>>>>>>     97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
>>>>>>      3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
>>>>>>      2936            -6.2%       2755        vmstat.system.cs
>>>>>>   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>>>>>>     18.57            +0.0%      18.57        fsmark.files_per_sec
>>>>>>     53420           -17.3%      44185        fsmark.time.voluntary_context_switches
>>>>>>      1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>>>>>>      3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>>>>>>   1756957            -2.1%    1720536        proc-vmstat.numa_hit
>>>>>>   1624496            -2.2%    1588039        proc-vmstat.numa_local
>>>>>>      1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
>>>>>>      2916            -6.2%       2735        perf-stat.i.context-switches
>>>>>>      1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
>>>>>>      2910            -6.2%       2729        perf-stat.ps.context-switches
>>>>>>      0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
>>>>>>      0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
>>>>>>      0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
>>>>>>      0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
>>>>>>      0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
>>>>>>      1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
>>>>>>      2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>>>>>>      2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
>>>>>>      2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>>>>>>      0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
>>>>>>      0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
>>>>>>      0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
>>>>>>      0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
>>>>>>      0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
>>>>>>      0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
>>>>>>      0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
>>>>>>      0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
>>>>>>      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
>>>>>>      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
>>>>>>      0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
>>>>>>      0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
>>>>>>      0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
>>>>>>      0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
>>>>>>      1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
>>>>>>      1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
>>>>>>      0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> Disclaimer:
>>>>>> Results have been estimated based on internal Intel analysis and are provided
>>>>>> for informational purposes only. Any difference in system hardware or software
>>>>>> design or configuration may affect actual performance.
>>>>>> 
>>>>>> 
>>>>> 
>>>>> This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
>>>>> a performance regression in the server. The only thing I can figure
>>>>> here is that this test is causing the server to recall the delegation
>>>>> that it hands out, and then the client has to go and establish a new
>>>>> open stateid in order to return it. That would likely be slower than
>>>>> just handing out both an open and delegation stateid in the first
>>>>> place.
>>>>> 
>>>>> I don't think there is anything we can do about that though, since the
>>>>> feature seems to is working as designed.
>>>> 
>>>> We seem to have hit this problem before. That makes me wonder
>>>> whether it is actually worth supporting the XOR flag at all. After
>>>> all, the client sends the CLOSE asynchronously; applications are not
>>>> being held up while the unneeded state ID is returned.
>>>> 
>>>> Can XOR support be disabled for now? I don't want to add an
>>>> administrative interface for that, but also, "no regressions" is
>>>> ringing in my ears, and 92% is a mighty noticeable one.
>>> 
>>> To be clear, this increase is for the "App Overhead" which is all of
>>> the operations between the stuff that is being measured in this test. I
>>> did run this test for a bit and got similar results, but was never able
>>> to nail down where the overhead came from. My speculation is that it's
>>> the recall and reestablishment of an open stateid that slows things
>>> down, but I never could fully confirm it.
>>> 
>>> My issue with disabling this is that the decision of whether to set
>>> OPEN_XOR_DELEGATION is up to the client. The server in this case is
>>> just doing what the client asks. ISTM that if we were going to disable
>>> (or throttle) this anywhere, that should be done by the client.
>> 
>> If the client sets the XOR flag, NFSD could simply return only an
>> OPEN stateid, for instance.
> 
> Sure, but then it's disabled for everyone, for every workload. There
> may be workloads the benefit too.

It doesn't impact NFSv4.[01] mounts, nor will it impact any
workload on a client that does not now support XOR.

I don't think we need bad press around this feature, nor do we
want complaints about "why didn't you add a switch to disable
this?". I'd also like to avoid a bug filed for this regression
if it gets merged.

Can you test to see if returning only the OPEN stateid makes the
regression go away? The behavior would be in place only until
there is a root cause and a way to avoid the regression.


--
Chuck Lever



^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-13 16:37           ` Chuck Lever III
@ 2024-11-13 17:02             ` Jeff Layton
  0 siblings, 0 replies; 19+ messages in thread
From: Jeff Layton @ 2024-11-13 17:02 UTC (permalink / raw)
  To: Chuck Lever III
  Cc: kernel test robot, Mike Snitzer, oe-lkp@lists.linux.dev,
	kernel test robot, Anna Schumaker, Trond Myklebust, Tom Haynes,
	Linux NFS Mailing List

On Wed, 2024-11-13 at 16:37 +0000, Chuck Lever III wrote:
> 
> > On Nov 13, 2024, at 11:20 AM, Jeff Layton <jlayton@kernel.org> wrote:
> > 
> > On Wed, 2024-11-13 at 11:14 -0500, Chuck Lever wrote:
> > > On Wed, Nov 13, 2024 at 11:10:35AM -0500, Jeff Layton wrote:
> > > > On Wed, 2024-11-13 at 10:48 -0500, Chuck Lever wrote:
> > > > > On Thu, Nov 07, 2024 at 06:35:11AM -0500, Jeff Layton wrote:
> > > > > > On Thu, 2024-11-07 at 12:55 +0800, kernel test robot wrote:
> > > > > > > hi, Jeff Layton,
> > > > > > > 
> > > > > > > in commit message, it is mentioned the change is expected to solve the
> > > > > > > "App Overhead" on the fs_mark test we reported in
> > > > > > > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > > > > > > 
> > > > > > > however, in our tests, there is sill similar regression. at the same
> > > > > > > time, there is still no performance difference for fsmark.files_per_sec
> > > > > > > 
> > > > > > >   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > > > > >     18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > > > > 
> > > > > > > 
> > > > > > > another thing is our bot bisect to this commit in repo/branch as below detail
> > > > > > > information. if there is a more porper repo/branch to test the patch, could
> > > > > > > you let us know? thanks a lot!
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > Hello,
> > > > > > > 
> > > > > > > kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> > > > > > > 
> > > > > > > 
> > > > > > > commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > > > > https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> > > > > > > 
> > > > > > > 
> > > > > > > testcase: fsmark
> > > > > > > config: x86_64-rhel-8.3
> > > > > > > compiler: gcc-12
> > > > > > > test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> > > > > > > parameters:
> > > > > > > 
> > > > > > > iterations: 1x
> > > > > > > nr_threads: 1t
> > > > > > > disk: 1HDD
> > > > > > > fs: btrfs
> > > > > > > fs2: nfsv4
> > > > > > > filesize: 4K
> > > > > > > test_size: 40M
> > > > > > > sync_method: fsyncBeforeClose
> > > > > > > nr_files_per_directory: 1fpd
> > > > > > > cpufreq_governor: performance
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > If you fix the issue in a separate patch/commit (i.e. not just a new version of
> > > > > > > the same patch/commit), kindly add following tags
> > > > > > > > Reported-by: kernel test robot <oliver.sang@intel.com>
> > > > > > > > Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > > > > 
> > > > > > > 
> > > > > > > Details are as below:
> > > > > > > -------------------------------------------------------------------------------------------------->
> > > > > > > 
> > > > > > > 
> > > > > > > The kernel config and materials to reproduce are available at:
> > > > > > > https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> > > > > > > 
> > > > > > > =========================================================================================
> > > > > > > compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
> > > > > > >  gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> > > > > > > 
> > > > > > > commit: 
> > > > > > >  37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
> > > > > > >  b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> > > > > > > 
> > > > > > > 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> > > > > > > ---------------- --------------------------- 
> > > > > > >         %stddev     %change         %stddev
> > > > > > >             \          |                \  
> > > > > > >     97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
> > > > > > >      3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
> > > > > > >      2936            -6.2%       2755        vmstat.system.cs
> > > > > > >   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > > > > >     18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > > > >     53420           -17.3%      44185        fsmark.time.voluntary_context_switches
> > > > > > >      1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > > > > >      3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
> > > > > > >   1756957            -2.1%    1720536        proc-vmstat.numa_hit
> > > > > > >   1624496            -2.2%    1588039        proc-vmstat.numa_local
> > > > > > >      1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
> > > > > > >      2916            -6.2%       2735        perf-stat.i.context-switches
> > > > > > >      1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
> > > > > > >      2910            -6.2%       2729        perf-stat.ps.context-switches
> > > > > > >      0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
> > > > > > >      0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > > > > >      0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
> > > > > > >      0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
> > > > > > >      0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
> > > > > > >      1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
> > > > > > >      2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > > > > >      2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
> > > > > > >      2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
> > > > > > >      0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
> > > > > > >      0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
> > > > > > >      0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
> > > > > > >      0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
> > > > > > >      0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
> > > > > > >      0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
> > > > > > >      0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
> > > > > > >      0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
> > > > > > >      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
> > > > > > >      0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
> > > > > > >      0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
> > > > > > >      0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
> > > > > > >      0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
> > > > > > >      0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
> > > > > > >      1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
> > > > > > >      1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
> > > > > > >      0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > Disclaimer:
> > > > > > > Results have been estimated based on internal Intel analysis and are provided
> > > > > > > for informational purposes only. Any difference in system hardware or software
> > > > > > > design or configuration may affect actual performance.
> > > > > > > 
> > > > > > > 
> > > > > > 
> > > > > > This patch (b4be3ccf1c) is exceedingly simple, so I doubt it's causing
> > > > > > a performance regression in the server. The only thing I can figure
> > > > > > here is that this test is causing the server to recall the delegation
> > > > > > that it hands out, and then the client has to go and establish a new
> > > > > > open stateid in order to return it. That would likely be slower than
> > > > > > just handing out both an open and delegation stateid in the first
> > > > > > place.
> > > > > > 
> > > > > > I don't think there is anything we can do about that though, since the
> > > > > > feature seems to is working as designed.
> > > > > 
> > > > > We seem to have hit this problem before. That makes me wonder
> > > > > whether it is actually worth supporting the XOR flag at all. After
> > > > > all, the client sends the CLOSE asynchronously; applications are not
> > > > > being held up while the unneeded state ID is returned.
> > > > > 
> > > > > Can XOR support be disabled for now? I don't want to add an
> > > > > administrative interface for that, but also, "no regressions" is
> > > > > ringing in my ears, and 92% is a mighty noticeable one.
> > > > 
> > > > To be clear, this increase is for the "App Overhead" which is all of
> > > > the operations between the stuff that is being measured in this test. I
> > > > did run this test for a bit and got similar results, but was never able
> > > > to nail down where the overhead came from. My speculation is that it's
> > > > the recall and reestablishment of an open stateid that slows things
> > > > down, but I never could fully confirm it.
> > > > 
> > > > My issue with disabling this is that the decision of whether to set
> > > > OPEN_XOR_DELEGATION is up to the client. The server in this case is
> > > > just doing what the client asks. ISTM that if we were going to disable
> > > > (or throttle) this anywhere, that should be done by the client.
> > > 
> > > If the client sets the XOR flag, NFSD could simply return only an
> > > OPEN stateid, for instance.
> > 
> > Sure, but then it's disabled for everyone, for every workload. There
> > may be workloads the benefit too.
> 
> It doesn't impact NFSv4.[01] mounts, nor will it impact any
> workload on a client that does not now support XOR.
> 
> I don't think we need bad press around this feature, nor do we
> want complaints about "why didn't you add a switch to disable
> this?". I'd also like to avoid a bug filed for this regression
> if it gets merged.
> 
> Can you test to see if returning only the OPEN stateid makes the
> regression go away? The behavior would be in place only until
> there is a root cause and a way to avoid the regression.
> 

I don't have a test rig set up for this at the moment as this sort of
testing requires real hw. I also don't have anything automated for
doing this.

That will probably take a while, so if you're that concerned, then just
drop the patch. Maybe we'll get a datapoint from the hammerspace folks
that will help clarify the problem in the meantime.

-- 
Jeff Layton <jlayton@kernel.org>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-07  4:55 [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression kernel test robot
  2024-11-07 11:35 ` Jeff Layton
@ 2024-11-14 14:37 ` Chuck Lever
  2024-11-18  7:53   ` Oliver Sang
  1 sibling, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2024-11-14 14:37 UTC (permalink / raw)
  To: kernel test robot; +Cc: Jeff Layton, Mike Snitzer, oe-lkp, lkp

On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
> 
> hi, Jeff Layton,
> 
> in commit message, it is mentioned the change is expected to solve the
> "App Overhead" on the fs_mark test we reported in
> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> 
> however, in our tests, there is sill similar regression. at the same
> time, there is still no performance difference for fsmark.files_per_sec
> 
>    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>      18.57            +0.0%      18.57        fsmark.files_per_sec

Because the most recent version of this commit does not appear to
address the reported regression, Jeff suggested dropping it. I've
done that in the branch below, which is being prepared for v6.13.


> another thing is our bot bisect to this commit in repo/branch as below detail
> information. if there is a more porper repo/branch to test the patch, could
> you let us know? thanks a lot!

The proper branch to test is:

https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next


> Hello,
> 
> kernel test robot noticed a 92.0% regression of fsmark.app_overhead on:
> 
> 
> commit: b4be3ccf1c251cbd3a3cf5391a80fe3a5f6f075e ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> https://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git cel-nfsd-next-6.12-rc5
> 
> 
> testcase: fsmark
> config: x86_64-rhel-8.3
> compiler: gcc-12
> test machine: 128 threads 2 sockets Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz (Ice Lake) with 128G memory
> parameters:
> 
> 	iterations: 1x
> 	nr_threads: 1t
> 	disk: 1HDD
> 	fs: btrfs
> 	fs2: nfsv4
> 	filesize: 4K
> 	test_size: 40M
> 	sync_method: fsyncBeforeClose
> 	nr_files_per_directory: 1fpd
> 	cpufreq_governor: performance
> 
> 
> 
> 
> If you fix the issue in a separate patch/commit (i.e. not just a new version of
> the same patch/commit), kindly add following tags
> | Reported-by: kernel test robot <oliver.sang@intel.com>
> | Closes: https://lore.kernel.org/oe-lkp/202411071017.ddd9e9e2-oliver.sang@intel.com
> 
> 
> Details are as below:
> -------------------------------------------------------------------------------------------------->
> 
> 
> The kernel config and materials to reproduce are available at:
> https://download.01.org/0day-ci/archive/20241107/202411071017.ddd9e9e2-oliver.sang@intel.com
> 
> =========================================================================================
> compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
>   gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-8.3/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark
> 
> commit: 
>   37f27b20cd ("nfsd: add support for FATTR4_OPEN_ARGUMENTS")
>   b4be3ccf1c ("nfsd: implement OPEN_ARGS_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION")
> 
> 37f27b20cd64e2e0 b4be3ccf1c251cbd3a3cf5391a8 
> ---------------- --------------------------- 
>          %stddev     %change         %stddev
>              \          |                \  
>      97.33 ±  9%     -16.3%      81.50 ±  9%  perf-c2c.HITM.local
>       3788 ±101%    +147.5%       9377 ±  6%  sched_debug.cfs_rq:/.load_avg.max
>       2936            -6.2%       2755        vmstat.system.cs
>    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>      18.57            +0.0%      18.57        fsmark.files_per_sec
>      53420           -17.3%      44185        fsmark.time.voluntary_context_switches
>       1.50 ±  7%     +13.4%       1.70 ±  3%  perf-sched.wait_time.avg.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>       3.00 ±  7%     +13.4%       3.40 ±  3%  perf-sched.wait_time.max.ms.devkmsg_read.vfs_read.ksys_read.do_syscall_64
>    1756957            -2.1%    1720536        proc-vmstat.numa_hit
>    1624496            -2.2%    1588039        proc-vmstat.numa_local
>       1.28 ±  4%      -8.2%       1.17 ±  3%  perf-stat.i.MPKI
>       2916            -6.2%       2735        perf-stat.i.context-switches
>       1529 ±  4%      +8.2%       1655 ±  3%  perf-stat.i.cycles-between-cache-misses
>       2910            -6.2%       2729        perf-stat.ps.context-switches
>       0.67 ± 15%      -0.4        0.27 ±100%  perf-profile.calltrace.cycles-pp._Fork
>       0.95 ± 15%      +0.3        1.26 ± 11%  perf-profile.calltrace.cycles-pp.__sched_setaffinity.sched_setaffinity.__x64_sys_sched_setaffinity.do_syscall_64.entry_SYSCALL_64_after_hwframe
>       0.70 ± 47%      +0.3        1.04 ±  9%  perf-profile.calltrace.cycles-pp._nohz_idle_balance.do_idle.cpu_startup_entry.start_secondary.common_startup_64
>       0.52 ± 45%      +0.3        0.86 ± 15%  perf-profile.calltrace.cycles-pp.seq_read_iter.vfs_read.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe
>       0.72 ± 50%      +0.4        1.12 ± 18%  perf-profile.calltrace.cycles-pp.link_path_walk.path_openat.do_filp_open.do_sys_openat2.__x64_sys_openat
>       1.22 ± 26%      +0.4        1.67 ± 12%  perf-profile.calltrace.cycles-pp.sched_setaffinity.evlist_cpu_iterator__next.read_counters.process_interval.dispatch_events
>       2.20 ± 11%      +0.6        2.78 ± 13%  perf-profile.calltrace.cycles-pp.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>       2.20 ± 11%      +0.6        2.82 ± 12%  perf-profile.calltrace.cycles-pp.entry_SYSCALL_64_after_hwframe.read
>       2.03 ± 13%      +0.6        2.67 ± 13%  perf-profile.calltrace.cycles-pp.ksys_read.do_syscall_64.entry_SYSCALL_64_after_hwframe.read
>       0.68 ± 15%      -0.2        0.47 ± 19%  perf-profile.children.cycles-pp._Fork
>       0.56 ± 15%      -0.1        0.42 ± 16%  perf-profile.children.cycles-pp.tcp_v6_do_rcv
>       0.46 ± 13%      -0.1        0.35 ± 16%  perf-profile.children.cycles-pp.alloc_pages_mpol_noprof
>       0.10 ± 75%      +0.2        0.29 ± 19%  perf-profile.children.cycles-pp.refresh_cpu_vm_stats
>       0.28 ± 33%      +0.2        0.47 ± 16%  perf-profile.children.cycles-pp.show_stat
>       0.34 ± 32%      +0.2        0.54 ± 16%  perf-profile.children.cycles-pp.fold_vm_numa_events
>       0.37 ± 11%      +0.3        0.63 ± 23%  perf-profile.children.cycles-pp.setup_items_for_insert
>       0.88 ± 15%      +0.3        1.16 ± 12%  perf-profile.children.cycles-pp.__set_cpus_allowed_ptr
>       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.btrfs_writepages
>       0.37 ± 14%      +0.3        0.67 ± 61%  perf-profile.children.cycles-pp.extent_write_cache_pages
>       0.64 ± 19%      +0.3        0.94 ± 15%  perf-profile.children.cycles-pp.btrfs_insert_empty_items
>       0.38 ± 12%      +0.3        0.68 ± 58%  perf-profile.children.cycles-pp.btrfs_fdatawrite_range
>       0.32 ± 23%      +0.3        0.63 ± 64%  perf-profile.children.cycles-pp.extent_writepage
>       0.97 ± 14%      +0.3        1.31 ± 10%  perf-profile.children.cycles-pp.__sched_setaffinity
>       1.07 ± 16%      +0.4        1.44 ± 10%  perf-profile.children.cycles-pp.__x64_sys_sched_setaffinity
>       1.39 ± 18%      +0.5        1.90 ± 12%  perf-profile.children.cycles-pp.seq_read_iter
>       0.34 ± 30%      +0.2        0.52 ± 16%  perf-profile.self.cycles-pp.fold_vm_numa_events
> 
> 
> 
> 
> Disclaimer:
> Results have been estimated based on internal Intel analysis and are provided
> for informational purposes only. Any difference in system hardware or software
> design or configuration may affect actual performance.
> 
> 
> -- 
> 0-DAY CI Kernel Test Service
> https://github.com/intel/lkp-tests/wiki
> 

-- 
Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-14 14:37 ` Chuck Lever
@ 2024-11-18  7:53   ` Oliver Sang
  2024-12-02 15:40     ` Chuck Lever III
  0 siblings, 1 reply; 19+ messages in thread
From: Oliver Sang @ 2024-11-18  7:53 UTC (permalink / raw)
  To: Chuck Lever; +Cc: Jeff Layton, Mike Snitzer, oe-lkp, lkp, oliver.sang

hi, Chuck Lever,

On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
> > 
> > hi, Jeff Layton,
> > 
> > in commit message, it is mentioned the change is expected to solve the
> > "App Overhead" on the fs_mark test we reported in
> > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > 
> > however, in our tests, there is sill similar regression. at the same
> > time, there is still no performance difference for fsmark.files_per_sec
> > 
> >    2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> >      18.57            +0.0%      18.57        fsmark.files_per_sec
> 
> Because the most recent version of this commit does not appear to
> address the reported regression, Jeff suggested dropping it. I've
> done that in the branch below, which is being prepared for v6.13.
> 
> 
> > another thing is our bot bisect to this commit in repo/branch as below detail
> > information. if there is a more porper repo/branch to test the patch, could
> > you let us know? thanks a lot!
> 
> The proper branch to test is:
> 
> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next

thanks for information! this repo is under bot's monitoring.
unfortunately, our cluster run into problems these days and need some time
to full back. so if there is still regression reports, maybe will delay to
a later time.

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-11-18  7:53   ` Oliver Sang
@ 2024-12-02 15:40     ` Chuck Lever III
  2024-12-03  5:01       ` Oliver Sang
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever III @ 2024-12-02 15:40 UTC (permalink / raw)
  To: Oliver Sang
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com



> On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> 
> hi, Chuck Lever,
> 
> On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
>> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
>>> 
>>> hi, Jeff Layton,
>>> 
>>> in commit message, it is mentioned the change is expected to solve the
>>> "App Overhead" on the fs_mark test we reported in
>>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
>>> 
>>> however, in our tests, there is sill similar regression. at the same
>>> time, there is still no performance difference for fsmark.files_per_sec
>>> 
>>>   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>>>     18.57            +0.0%      18.57        fsmark.files_per_sec
>> 
>> Because the most recent version of this commit does not appear to
>> address the reported regression, Jeff suggested dropping it. I've
>> done that in the branch below, which is being prepared for v6.13.
>> 
>> 
>>> another thing is our bot bisect to this commit in repo/branch as below detail
>>> information. if there is a more porper repo/branch to test the patch, could
>>> you let us know? thanks a lot!
>> 
>> The proper branch to test is:
>> 
>> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
> 
> thanks for information! this repo is under bot's monitoring.
> unfortunately, our cluster run into problems these days and need some time
> to full back. so if there is still regression reports, maybe will delay to
> a later time.

Hello Oliver -

To avoid regressing v6.13, we postponed this series.

The series is now in 

  https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing

And I would like the robot to include that branch in regular
test runs going forward. I will try to push new patches there
first before pushing them to nfsd-next.

--
Chuck Lever



^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-02 15:40     ` Chuck Lever III
@ 2024-12-03  5:01       ` Oliver Sang
  2024-12-03 14:54         ` Chuck Lever III
  0 siblings, 1 reply; 19+ messages in thread
From: Oliver Sang @ 2024-12-03  5:01 UTC (permalink / raw)
  To: Chuck Lever III
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com,
	oliver.sang

hi, Chuck Lever,

On Mon, Dec 02, 2024 at 03:40:49PM +0000, Chuck Lever III wrote:
> 
> 
> > On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> > 
> > hi, Chuck Lever,
> > 
> > On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
> >> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
> >>> 
> >>> hi, Jeff Layton,
> >>> 
> >>> in commit message, it is mentioned the change is expected to solve the
> >>> "App Overhead" on the fs_mark test we reported in
> >>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> >>> 
> >>> however, in our tests, there is sill similar regression. at the same
> >>> time, there is still no performance difference for fsmark.files_per_sec
> >>> 
> >>>   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> >>>     18.57            +0.0%      18.57        fsmark.files_per_sec
> >> 
> >> Because the most recent version of this commit does not appear to
> >> address the reported regression, Jeff suggested dropping it. I've
> >> done that in the branch below, which is being prepared for v6.13.
> >> 
> >> 
> >>> another thing is our bot bisect to this commit in repo/branch as below detail
> >>> information. if there is a more porper repo/branch to test the patch, could
> >>> you let us know? thanks a lot!
> >> 
> >> The proper branch to test is:
> >> 
> >> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
> > 
> > thanks for information! this repo is under bot's monitoring.
> > unfortunately, our cluster run into problems these days and need some time
> > to full back. so if there is still regression reports, maybe will delay to
> > a later time.
> 
> Hello Oliver -
> 
> To avoid regressing v6.13, we postponed this series.
> 
> The series is now in 
> 
>   https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing
> 
> And I would like the robot to include that branch in regular
> test runs going forward. I will try to push new patches there
> first before pushing them to nfsd-next.

got it. this nfsd-testing is already in our monitor.

07b2337dc158fb (cel/nfsd-testing) nfsd: add support for FATTR4_OPEN_ARGUMENTS
54b0e9c3727d2f nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
fcc756b92ce832 nfsd: fix handling of delegated change attr in CB_GETATTR
a354743d35b49b nfsd: trace: remove redundant stateid even deleg_recall
40384c840ea194 (tag: v6.13-rc1,

at the same time, I triggered same tests for original report upon 07b2337dc158fb
and v6.13-rc1, will update you results later.

> 
> --
> Chuck Lever
> 
> 

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-03  5:01       ` Oliver Sang
@ 2024-12-03 14:54         ` Chuck Lever III
  2024-12-05  7:27           ` Oliver Sang
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever III @ 2024-12-03 14:54 UTC (permalink / raw)
  To: Oliver Sang
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com



> On Dec 3, 2024, at 12:01 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> 
> hi, Chuck Lever,
> 
> On Mon, Dec 02, 2024 at 03:40:49PM +0000, Chuck Lever III wrote:
>> 
>> 
>>> On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
>>> 
>>> hi, Chuck Lever,
>>> 
>>> On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
>>>> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
>>>>> 
>>>>> hi, Jeff Layton,
>>>>> 
>>>>> in commit message, it is mentioned the change is expected to solve the
>>>>> "App Overhead" on the fs_mark test we reported in
>>>>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
>>>>> 
>>>>> however, in our tests, there is sill similar regression. at the same
>>>>> time, there is still no performance difference for fsmark.files_per_sec
>>>>> 
>>>>>  2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>>>>>    18.57            +0.0%      18.57        fsmark.files_per_sec
>>>> 
>>>> Because the most recent version of this commit does not appear to
>>>> address the reported regression, Jeff suggested dropping it. I've
>>>> done that in the branch below, which is being prepared for v6.13.
>>>> 
>>>> 
>>>>> another thing is our bot bisect to this commit in repo/branch as below detail
>>>>> information. if there is a more porper repo/branch to test the patch, could
>>>>> you let us know? thanks a lot!
>>>> 
>>>> The proper branch to test is:
>>>> 
>>>> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
>>> 
>>> thanks for information! this repo is under bot's monitoring.
>>> unfortunately, our cluster run into problems these days and need some time
>>> to full back. so if there is still regression reports, maybe will delay to
>>> a later time.
>> 
>> Hello Oliver -
>> 
>> To avoid regressing v6.13, we postponed this series.
>> 
>> The series is now in 
>> 
>>  https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing
>> 
>> And I would like the robot to include that branch in regular
>> test runs going forward. I will try to push new patches there
>> first before pushing them to nfsd-next.
> 
> got it. this nfsd-testing is already in our monitor.
> 
> 07b2337dc158fb (cel/nfsd-testing) nfsd: add support for FATTR4_OPEN_ARGUMENTS
> 54b0e9c3727d2f nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
> fcc756b92ce832 nfsd: fix handling of delegated change attr in CB_GETATTR
> a354743d35b49b nfsd: trace: remove redundant stateid even deleg_recall
> 40384c840ea194 (tag: v6.13-rc1,

That's not yet a complete set of revised patches. We'll get
there in a few days.


> at the same time, I triggered same tests for original report upon 07b2337dc158fb
> and v6.13-rc1, will update you results later.

Thanks!


--
Chuck Lever



^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-03 14:54         ` Chuck Lever III
@ 2024-12-05  7:27           ` Oliver Sang
  2024-12-05 14:15             ` Jeff Layton
  2024-12-05 15:37             ` Chuck Lever
  0 siblings, 2 replies; 19+ messages in thread
From: Oliver Sang @ 2024-12-05  7:27 UTC (permalink / raw)
  To: Chuck Lever III
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com,
	oliver.sang

hi, Chuck Lever,

On Tue, Dec 03, 2024 at 02:54:47PM +0000, Chuck Lever III wrote:
> 
> 
> > On Dec 3, 2024, at 12:01 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> > 
> > hi, Chuck Lever,
> > 
> > On Mon, Dec 02, 2024 at 03:40:49PM +0000, Chuck Lever III wrote:
> >> 
> >> 
> >>> On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> >>> 
> >>> hi, Chuck Lever,
> >>> 
> >>> On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
> >>>> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
> >>>>> 
> >>>>> hi, Jeff Layton,
> >>>>> 
> >>>>> in commit message, it is mentioned the change is expected to solve the
> >>>>> "App Overhead" on the fs_mark test we reported in
> >>>>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> >>>>> 
> >>>>> however, in our tests, there is sill similar regression. at the same
> >>>>> time, there is still no performance difference for fsmark.files_per_sec
> >>>>> 
> >>>>>  2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> >>>>>    18.57            +0.0%      18.57        fsmark.files_per_sec
> >>>> 
> >>>> Because the most recent version of this commit does not appear to
> >>>> address the reported regression, Jeff suggested dropping it. I've
> >>>> done that in the branch below, which is being prepared for v6.13.
> >>>> 
> >>>> 
> >>>>> another thing is our bot bisect to this commit in repo/branch as below detail
> >>>>> information. if there is a more porper repo/branch to test the patch, could
> >>>>> you let us know? thanks a lot!
> >>>> 
> >>>> The proper branch to test is:
> >>>> 
> >>>> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
> >>> 
> >>> thanks for information! this repo is under bot's monitoring.
> >>> unfortunately, our cluster run into problems these days and need some time
> >>> to full back. so if there is still regression reports, maybe will delay to
> >>> a later time.
> >> 
> >> Hello Oliver -
> >> 
> >> To avoid regressing v6.13, we postponed this series.
> >> 
> >> The series is now in 
> >> 
> >>  https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing
> >> 
> >> And I would like the robot to include that branch in regular
> >> test runs going forward. I will try to push new patches there
> >> first before pushing them to nfsd-next.
> > 
> > got it. this nfsd-testing is already in our monitor.
> > 
> > 07b2337dc158fb (cel/nfsd-testing) nfsd: add support for FATTR4_OPEN_ARGUMENTS
> > 54b0e9c3727d2f nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
> > fcc756b92ce832 nfsd: fix handling of delegated change attr in CB_GETATTR
> > a354743d35b49b nfsd: trace: remove redundant stateid even deleg_recall
> > 40384c840ea194 (tag: v6.13-rc1,
> 
> That's not yet a complete set of revised patches. We'll get
> there in a few days.

got it. thanks for information.

> 
> 
> > at the same time, I triggered same tests for original report upon 07b2337dc158fb
> > and v6.13-rc1, will update you results later.

if just comparing 07b2337dc158fb with v6.13-rc1, there is no big diff for
fsmark.app_overhead.

       v6.13-rc1 07b2337dc158fb33e7af288b74c
---------------- ---------------------------
         %stddev     %change         %stddev
             \          |                \
   3110649            -1.3%    3070867        fsmark.app_overhead
     18.58            -0.3%      18.52        fsmark.files_per_sec

but as above, seems the 07b2337dc158fb is not the correct commit to test?

we noticed it's now

b95256b0edb43 (cel/nfsd-testing) nfsd: prepare delegation code for handing out *_ATTRS_DELEG delegations
77152578a5c16 nfsd: rename NFS4_SHARE_WANT_* constants to OPEN4_SHARE_ACCESS_WANT_*
019cb8ea1e9d4 nfsd: switch to autogenerated definitions for open_delegation_type4
07b2337dc158f nfsd: add support for FATTR4_OPEN_ARGUMENTS
54b0e9c3727d2 nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
fcc756b92ce83 nfsd: fix handling of delegated change attr in CB_GETATTR
a354743d35b49 nfsd: trace: remove redundant stateid even deleg_recall
40384c840ea19 (tag: v6.13-rc1,

if you need us test b95256b0edb43, please let us know. or you could request us
when the branch is ready. thanks!

> 
> Thanks!
> 
> 
> --
> Chuck Lever
> 
> 

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd]  b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-05  7:27           ` Oliver Sang
@ 2024-12-05 14:15             ` Jeff Layton
  2024-12-05 15:37             ` Chuck Lever
  1 sibling, 0 replies; 19+ messages in thread
From: Jeff Layton @ 2024-12-05 14:15 UTC (permalink / raw)
  To: Oliver Sang, Chuck Lever III
  Cc: Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com

On Thu, 2024-12-05 at 15:27 +0800, Oliver Sang wrote:
> hi, Chuck Lever,
> 
> On Tue, Dec 03, 2024 at 02:54:47PM +0000, Chuck Lever III wrote:
> > 
> > 
> > > On Dec 3, 2024, at 12:01 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> > > 
> > > hi, Chuck Lever,
> > > 
> > > On Mon, Dec 02, 2024 at 03:40:49PM +0000, Chuck Lever III wrote:
> > > > 
> > > > 
> > > > > On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
> > > > > 
> > > > > hi, Chuck Lever,
> > > > > 
> > > > > On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
> > > > > > On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
> > > > > > > 
> > > > > > > hi, Jeff Layton,
> > > > > > > 
> > > > > > > in commit message, it is mentioned the change is expected to solve the
> > > > > > > "App Overhead" on the fs_mark test we reported in
> > > > > > > https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
> > > > > > > 
> > > > > > > however, in our tests, there is sill similar regression. at the same
> > > > > > > time, there is still no performance difference for fsmark.files_per_sec
> > > > > > > 
> > > > > > >  2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
> > > > > > >    18.57            +0.0%      18.57        fsmark.files_per_sec
> > > > > > 
> > > > > > Because the most recent version of this commit does not appear to
> > > > > > address the reported regression, Jeff suggested dropping it. I've
> > > > > > done that in the branch below, which is being prepared for v6.13.
> > > > > > 
> > > > > > 
> > > > > > > another thing is our bot bisect to this commit in repo/branch as below detail
> > > > > > > information. if there is a more porper repo/branch to test the patch, could
> > > > > > > you let us know? thanks a lot!
> > > > > > 
> > > > > > The proper branch to test is:
> > > > > > 
> > > > > > https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
> > > > > 
> > > > > thanks for information! this repo is under bot's monitoring.
> > > > > unfortunately, our cluster run into problems these days and need some time
> > > > > to full back. so if there is still regression reports, maybe will delay to
> > > > > a later time.
> > > > 
> > > > Hello Oliver -
> > > > 
> > > > To avoid regressing v6.13, we postponed this series.
> > > > 
> > > > The series is now in 
> > > > 
> > > >  https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing
> > > > 
> > > > And I would like the robot to include that branch in regular
> > > > test runs going forward. I will try to push new patches there
> > > > first before pushing them to nfsd-next.
> > > 
> > > got it. this nfsd-testing is already in our monitor.
> > > 
> > > 07b2337dc158fb (cel/nfsd-testing) nfsd: add support for FATTR4_OPEN_ARGUMENTS
> > > 54b0e9c3727d2f nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
> > > fcc756b92ce832 nfsd: fix handling of delegated change attr in CB_GETATTR
> > > a354743d35b49b nfsd: trace: remove redundant stateid even deleg_recall
> > > 40384c840ea194 (tag: v6.13-rc1,
> > 
> > That's not yet a complete set of revised patches. We'll get
> > there in a few days.
> 
> got it. thanks for information.
> 
> > 
> > 
> > > at the same time, I triggered same tests for original report upon 07b2337dc158fb
> > > and v6.13-rc1, will update you results later.
> 
> if just comparing 07b2337dc158fb with v6.13-rc1, there is no big diff for
> fsmark.app_overhead.
> 
>        v6.13-rc1 07b2337dc158fb33e7af288b74c
> ---------------- ---------------------------
>          %stddev     %change         %stddev
>              \          |                \
>    3110649            -1.3%    3070867        fsmark.app_overhead
>      18.58            -0.3%      18.52        fsmark.files_per_sec
> 
> but as above, seems the 07b2337dc158fb is not the correct commit to test?
> 

Yes, I wouldn't expect to see a regression with that patch. That commit
just adds support for FATTR4_OPEN_ARGUMENTS, without any real change in
functionality.  I think Chuck was planning to put the rest of the
patches back in and just disable
OPEN4_SHARE_ACCESS_WANT_OPEN_XOR_DELEGATION support. With that you
should be able to see the regression again, but I don't see those
patches in his tree just yet.


> we noticed it's now
> 
> b95256b0edb43 (cel/nfsd-testing) nfsd: prepare delegation code for handing out *_ATTRS_DELEG delegations
> 77152578a5c16 nfsd: rename NFS4_SHARE_WANT_* constants to OPEN4_SHARE_ACCESS_WANT_*
> 019cb8ea1e9d4 nfsd: switch to autogenerated definitions for open_delegation_type4
> 07b2337dc158f nfsd: add support for FATTR4_OPEN_ARGUMENTS
> 54b0e9c3727d2 nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
> fcc756b92ce83 nfsd: fix handling of delegated change attr in CB_GETATTR
> a354743d35b49 nfsd: trace: remove redundant stateid even deleg_recall
> 40384c840ea19 (tag: v6.13-rc1,
> 
> if you need us test b95256b0edb43, please let us know. or you could request us
> when the branch is ready. thanks!
> 
> > 
> > Thanks!
> > 
> > 
> > --
> > Chuck Lever
> > 
> > 

-- 
Jeff Layton <jlayton@kernel.org>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-05  7:27           ` Oliver Sang
  2024-12-05 14:15             ` Jeff Layton
@ 2024-12-05 15:37             ` Chuck Lever
  2024-12-06  2:33               ` Oliver Sang
  1 sibling, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2024-12-05 15:37 UTC (permalink / raw)
  To: Oliver Sang
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com

On 12/5/24 2:27 AM, Oliver Sang wrote:
> hi, Chuck Lever,
> 
> On Tue, Dec 03, 2024 at 02:54:47PM +0000, Chuck Lever III wrote:
>>
>>
>>> On Dec 3, 2024, at 12:01 AM, Oliver Sang <oliver.sang@intel.com> wrote:
>>>
>>> hi, Chuck Lever,
>>>
>>> On Mon, Dec 02, 2024 at 03:40:49PM +0000, Chuck Lever III wrote:
>>>>
>>>>
>>>>> On Nov 18, 2024, at 2:53 AM, Oliver Sang <oliver.sang@intel.com> wrote:
>>>>>
>>>>> hi, Chuck Lever,
>>>>>
>>>>> On Thu, Nov 14, 2024 at 09:37:30AM -0500, Chuck Lever wrote:
>>>>>> On Thu, Nov 07, 2024 at 12:55:02PM +0800, kernel test robot wrote:
>>>>>>>
>>>>>>> hi, Jeff Layton,
>>>>>>>
>>>>>>> in commit message, it is mentioned the change is expected to solve the
>>>>>>> "App Overhead" on the fs_mark test we reported in
>>>>>>> https://lore.kernel.org/oe-lkp/202409161645.d44bced5-oliver.sang@intel.com/
>>>>>>>
>>>>>>> however, in our tests, there is sill similar regression. at the same
>>>>>>> time, there is still no performance difference for fsmark.files_per_sec
>>>>>>>
>>>>>>>   2015880 ±  3%     +92.0%    3870164        fsmark.app_overhead
>>>>>>>     18.57            +0.0%      18.57        fsmark.files_per_sec
>>>>>>
>>>>>> Because the most recent version of this commit does not appear to
>>>>>> address the reported regression, Jeff suggested dropping it. I've
>>>>>> done that in the branch below, which is being prepared for v6.13.
>>>>>>
>>>>>>
>>>>>>> another thing is our bot bisect to this commit in repo/branch as below detail
>>>>>>> information. if there is a more porper repo/branch to test the patch, could
>>>>>>> you let us know? thanks a lot!
>>>>>>
>>>>>> The proper branch to test is:
>>>>>>
>>>>>> https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-next
>>>>>
>>>>> thanks for information! this repo is under bot's monitoring.
>>>>> unfortunately, our cluster run into problems these days and need some time
>>>>> to full back. so if there is still regression reports, maybe will delay to
>>>>> a later time.
>>>>
>>>> Hello Oliver -
>>>>
>>>> To avoid regressing v6.13, we postponed this series.
>>>>
>>>> The series is now in
>>>>
>>>>   https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git/log/?h=nfsd-testing
>>>>
>>>> And I would like the robot to include that branch in regular
>>>> test runs going forward. I will try to push new patches there
>>>> first before pushing them to nfsd-next.
>>>
>>> got it. this nfsd-testing is already in our monitor.
>>>
>>> 07b2337dc158fb (cel/nfsd-testing) nfsd: add support for FATTR4_OPEN_ARGUMENTS
>>> 54b0e9c3727d2f nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
>>> fcc756b92ce832 nfsd: fix handling of delegated change attr in CB_GETATTR
>>> a354743d35b49b nfsd: trace: remove redundant stateid even deleg_recall
>>> 40384c840ea194 (tag: v6.13-rc1,
>>
>> That's not yet a complete set of revised patches. We'll get
>> there in a few days.
> 
> got it. thanks for information.
> 
>>
>>
>>> at the same time, I triggered same tests for original report upon 07b2337dc158fb
>>> and v6.13-rc1, will update you results later.
> 
> if just comparing 07b2337dc158fb with v6.13-rc1, there is no big diff for
> fsmark.app_overhead.
> 
>         v6.13-rc1 07b2337dc158fb33e7af288b74c
> ---------------- ---------------------------
>           %stddev     %change         %stddev
>               \          |                \
>     3110649            -1.3%    3070867        fsmark.app_overhead
>       18.58            -0.3%      18.52        fsmark.files_per_sec
> 
> but as above, seems the 07b2337dc158fb is not the correct commit to test?

> we noticed it's now
> 
> b95256b0edb43 (cel/nfsd-testing) nfsd: prepare delegation code for handing out *_ATTRS_DELEG delegations
> 77152578a5c16 nfsd: rename NFS4_SHARE_WANT_* constants to OPEN4_SHARE_ACCESS_WANT_*
> 019cb8ea1e9d4 nfsd: switch to autogenerated definitions for open_delegation_type4
> 07b2337dc158f nfsd: add support for FATTR4_OPEN_ARGUMENTS
> 54b0e9c3727d2 nfs_common: make include/linux/nfs4.h include generated nfs4_1.h
> fcc756b92ce83 nfsd: fix handling of delegated change attr in CB_GETATTR
> a354743d35b49 nfsd: trace: remove redundant stateid even deleg_recall
> 40384c840ea19 (tag: v6.13-rc1,
> 
> if you need us test b95256b0edb43, please let us know. or you could request us
> when the branch is ready. thanks!

Earlier this week I thought I had the full set of patches to
test and could just apply them on v6.13-rc1, but it turns out
that I don't have them all yet. There are one or two that I
broke while trying to prepare for v6.12, and I've misplaced
Jeff's originals.

I will let you know when that branch is ready for a retest.
Sorry for the noise!


-- 
Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-05 15:37             ` Chuck Lever
@ 2024-12-06  2:33               ` Oliver Sang
  2024-12-16 22:23                 ` Chuck Lever
  0 siblings, 1 reply; 19+ messages in thread
From: Oliver Sang @ 2024-12-06  2:33 UTC (permalink / raw)
  To: Chuck Lever
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com,
	oliver.sang

hi, Chuck Lever,

On Thu, Dec 05, 2024 at 10:37:10AM -0500, Chuck Lever wrote:

[...]

> 
> Earlier this week I thought I had the full set of patches to
> test and could just apply them on v6.13-rc1, but it turns out
> that I don't have them all yet. There are one or two that I
> broke while trying to prepare for v6.12, and I've misplaced
> Jeff's originals.
> 
> I will let you know when that branch is ready for a retest.
> Sorry for the noise!

Not at all! always our great pleasure that we can supply any help.

got it. just wait your notification. thanks!

> 
> 
> -- 
> Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-06  2:33               ` Oliver Sang
@ 2024-12-16 22:23                 ` Chuck Lever
  2024-12-19  2:32                   ` Oliver Sang
  0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2024-12-16 22:23 UTC (permalink / raw)
  To: Oliver Sang, Jeff Layton
  Cc: Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com

On 12/5/24 9:33 PM, Oliver Sang wrote:
> hi, Chuck Lever,
> 
> On Thu, Dec 05, 2024 at 10:37:10AM -0500, Chuck Lever wrote:
> 
> [...]
> 
>>
>> Earlier this week I thought I had the full set of patches to
>> test and could just apply them on v6.13-rc1, but it turns out
>> that I don't have them all yet. There are one or two that I
>> broke while trying to prepare for v6.12, and I've misplaced
>> Jeff's originals.
>>
>> I will let you know when that branch is ready for a retest.
>> Sorry for the noise!
> 
> Not at all! always our great pleasure that we can supply any help.
> 
> got it. just wait your notification. thanks!

The series currently in the nfsd-testing branch is complete and stable
enough for you to benchmark.


-- 
Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression
  2024-12-16 22:23                 ` Chuck Lever
@ 2024-12-19  2:32                   ` Oliver Sang
  0 siblings, 0 replies; 19+ messages in thread
From: Oliver Sang @ 2024-12-19  2:32 UTC (permalink / raw)
  To: Chuck Lever
  Cc: Jeff Layton, Mike Snitzer, oe-lkp@lists.linux.dev, lkp@intel.com,
	oliver.sang

hi, Chuck Lever,

On Mon, Dec 16, 2024 at 05:23:48PM -0500, Chuck Lever wrote:
> On 12/5/24 9:33 PM, Oliver Sang wrote:
> > hi, Chuck Lever,
> > 
> > On Thu, Dec 05, 2024 at 10:37:10AM -0500, Chuck Lever wrote:
> > 
> > [...]
> > 
> > > 
> > > Earlier this week I thought I had the full set of patches to
> > > test and could just apply them on v6.13-rc1, but it turns out
> > > that I don't have them all yet. There are one or two that I
> > > broke while trying to prepare for v6.12, and I've misplaced
> > > Jeff's originals.
> > > 
> > > I will let you know when that branch is ready for a retest.
> > > Sorry for the noise!
> > 
> > Not at all! always our great pleasure that we can supply any help.
> > 
> > got it. just wait your notification. thanks!
> 
> The series currently in the nfsd-testing branch is complete and stable
> enough for you to benchmark.

FYI. we compared the branch tip with v6.13-rc3, the fsmark.app_overhead
regression are less now. while fsmark.files_per_sec keeps not being impacted.

if bisect can succeed to spot out one commit, we will also send out a report.

=========================================================================================
compiler/cpufreq_governor/disk/filesize/fs2/fs/iterations/kconfig/nr_files_per_directory/nr_threads/rootfs/sync_method/tbox_group/test_size/testcase:
  gcc-12/performance/1HDD/4K/nfsv4/btrfs/1x/x86_64-rhel-9.4/1fpd/1t/debian-12-x86_64-20240206.cgz/fsyncBeforeClose/lkp-icl-2sp6/40M/fsmark

commit:
  v6.13-rc3
  8d5b7358ea7c0 ("nfsd: add shrinker to reduce number of slots allocated per session")

       v6.13-rc3 8d5b7358ea7c07b69c44f0af21e
---------------- ---------------------------
         %stddev     %change         %stddev
             \          |                \
   2790150 ± 52%     -60.6%    1099116 ±120%  numa-meminfo.node1.FilePages
    697551 ± 52%     -60.6%     274783 ±120%  numa-vmstat.node1.nr_file_pages
      1176 ± 71%   +2210.3%      27185 ±247%  sched_debug.cpu.max_idle_balance_cost.stddev
      2882            -7.1%       2677        vmstat.system.cs
   3064384           +57.8%    4834169        fsmark.app_overhead
     18.58            -0.2%      18.54        fsmark.files_per_sec
     53429           -17.3%      44169        fsmark.time.voluntary_context_switches
      2805            -7.4%       2598        perf-stat.i.context-switches
      2800            -7.4%       2594        perf-stat.ps.context-switches
   1750049            -1.2%    1729157        proc-vmstat.numa_hit
   1617582            -1.3%    1596589        proc-vmstat.numa_local
      2.03 ±  2%      +5.5%       2.14        time.system_time
     53429           -17.3%      44169        time.voluntary_context_switches
      0.80 ± 11%      -0.3        0.52 ± 56%  perf-profile.calltrace.cycles-pp.tcp_write_xmit.__tcp_push_pending_frames.tcp_sendmsg_locked.tcp_sendmsg.sock_sendmsg
      0.80 ± 11%      -0.2        0.59 ± 39%  perf-profile.calltrace.cycles-pp.__tcp_push_pending_frames.tcp_sendmsg_locked.tcp_sendmsg.sock_sendmsg.svc_tcp_sendmsg
      0.39 ± 23%      -0.2        0.19 ± 53%  perf-profile.children.cycles-pp.__close
      0.39 ± 20%      -0.2        0.20 ± 47%  perf-profile.children.cycles-pp.__x64_sys_close
      0.14 ± 15%      -0.1        0.06 ± 79%  perf-profile.children.cycles-pp.syscall_return_via_sysret

> 
> 
> -- 
> Chuck Lever

^ permalink raw reply	[flat|nested] 19+ messages in thread

end of thread, other threads:[~2024-12-19  2:32 UTC | newest]

Thread overview: 19+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2024-11-07  4:55 [snitzer:cel-nfsd-next-6.12-rc5] [nfsd] b4be3ccf1c: fsmark.app_overhead 92.0% regression kernel test robot
2024-11-07 11:35 ` Jeff Layton
2024-11-13 15:48   ` Chuck Lever
2024-11-13 16:10     ` Jeff Layton
2024-11-13 16:14       ` Chuck Lever
2024-11-13 16:20         ` Jeff Layton
2024-11-13 16:37           ` Chuck Lever III
2024-11-13 17:02             ` Jeff Layton
2024-11-14 14:37 ` Chuck Lever
2024-11-18  7:53   ` Oliver Sang
2024-12-02 15:40     ` Chuck Lever III
2024-12-03  5:01       ` Oliver Sang
2024-12-03 14:54         ` Chuck Lever III
2024-12-05  7:27           ` Oliver Sang
2024-12-05 14:15             ` Jeff Layton
2024-12-05 15:37             ` Chuck Lever
2024-12-06  2:33               ` Oliver Sang
2024-12-16 22:23                 ` Chuck Lever
2024-12-19  2:32                   ` Oliver Sang

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.