The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: David Vernet <void@manifault.com>
To: Aaron Lu <aaron.lu@intel.com>
Cc: Peter Zijlstra <peterz@infradead.org>,
	linux-kernel@vger.kernel.org, mingo@redhat.com,
	juri.lelli@redhat.com, vincent.guittot@linaro.org,
	rostedt@goodmis.org, dietmar.eggemann@arm.com,
	bsegall@google.com, mgorman@suse.de, bristot@redhat.com,
	vschneid@redhat.com, joshdon@google.com,
	roman.gushchin@linux.dev, tj@kernel.org, kernel-team@meta.com
Subject: Re: [RFC PATCH 3/3] sched: Implement shared wakequeue in CFS
Date: Thu, 15 Jun 2023 18:26:05 -0500	[thread overview]
Message-ID: <20230615232605.GB2915572@maniforge> (raw)
In-Reply-To: <20230615073153.GA110814@ziqianlu-dell>

On Thu, Jun 15, 2023 at 03:31:53PM +0800, Aaron Lu wrote:
> On Thu, Jun 15, 2023 at 12:49:17PM +0800, Aaron Lu wrote:
> > I'll see if I can find a smaller machine and give it a run there too.
> 
> Found a Skylake with 18cores/36threads on each socket/LLC and with
> netperf, the contention is still serious.
> 
> "
> $ netserver
> $ sudo sh -c "echo SWQUEUE > /sys/kernel/debug/sched/features"
> $ for i in `seq 72`; do netperf -l 60 -n 72 -6 -t UDP_RR & done
> "
> 
>         53.61%    53.61%  [kernel.vmlinux]            [k] native_queued_spin_lock_slowpath            -      -            
>             |          
>             |--27.93%--sendto
>             |          entry_SYSCALL_64
>             |          do_syscall_64
>             |          |          
>             |           --27.93%--__x64_sys_sendto
>             |                     __sys_sendto
>             |                     sock_sendmsg
>             |                     inet6_sendmsg
>             |                     udpv6_sendmsg
>             |                     udp_v6_send_skb
>             |                     ip6_send_skb
>             |                     ip6_local_out
>             |                     ip6_output
>             |                     ip6_finish_output
>             |                     ip6_finish_output2
>             |                     __dev_queue_xmit
>             |                     __local_bh_enable_ip
>             |                     do_softirq.part.0
>             |                     __do_softirq
>             |                     net_rx_action
>             |                     __napi_poll
>             |                     process_backlog
>             |                     __netif_receive_skb
>             |                     __netif_receive_skb_one_core
>             |                     ipv6_rcv
>             |                     ip6_input
>             |                     ip6_input_finish
>             |                     ip6_protocol_deliver_rcu
>             |                     udpv6_rcv
>             |                     __udp6_lib_rcv
>             |                     udp6_unicast_rcv_skb
>             |                     udpv6_queue_rcv_skb
>             |                     udpv6_queue_rcv_one_skb
>             |                     __udp_enqueue_schedule_skb
>             |                     sock_def_readable
>             |                     __wake_up_sync_key
>             |                     __wake_up_common_lock
>             |                     |          
>             |                      --27.85%--__wake_up_common
>             |                                receiver_wake_function
>             |                                autoremove_wake_function
>             |                                default_wake_function
>             |                                try_to_wake_up
>             |                                |          
>             |                                 --27.85%--ttwu_do_activate
>             |                                           enqueue_task
>             |                                           enqueue_task_fair
>             |                                           |          
>             |                                            --27.85%--_raw_spin_lock_irqsave
>             |                                                      |          
>             |                                                       --27.85%--native_queued_spin_lock_slowpath
>             |          
>              --25.67%--recvfrom
>                        entry_SYSCALL_64
>                        do_syscall_64
>                        __x64_sys_recvfrom
>                        __sys_recvfrom
>                        sock_recvmsg
>                        inet6_recvmsg
>                        udpv6_recvmsg
>                        __skb_recv_udp
>                        |          
>                         --25.67%--__skb_wait_for_more_packets
>                                   schedule_timeout
>                                   schedule
>                                   __schedule
>                                   |          
>                                    --25.66%--pick_next_task_fair
>                                              |          
>                                               --25.65%--swqueue_remove_task
>                                                         |          
>                                                          --25.65%--_raw_spin_lock_irqsave
>                                                                    |          
>                                                                     --25.65%--native_queued_spin_lock_slowpath
> 
> I didn't aggregate the throughput(Trans. Rate per sec) from all these
> clients, but a glimpse from the result showed that the throughput of
> each client dropped from 4xxxx(NO_SWQUEUE) to 2xxxx(SWQUEUE).
> 
> Thanks,
> Aaron

Ok, it seems that the issue is that I wasn't creating enough netperf
clients. I assumed that -n $(nproc) was sufficient. I was able to repro
the contention on my 26 core / 52 thread skylake client as well:


    41.01%  netperf          [kernel.vmlinux]                                                 [k] queued_spin_lock_slowpath
            |          
             --41.01%--queued_spin_lock_slowpath
                       |          
                        --40.63%--_raw_spin_lock_irqsave
                                  |          
                                  |--21.18%--enqueue_task_fair
                                  |          |          
                                  |           --21.09%--default_wake_function
                                  |                     |          
                                  |                      --21.09%--autoremove_wake_function
                                  |                                |          
                                  |                                 --21.09%--__wake_up_sync_key
                                  |                                           sock_def_readable
                                  |                                           __udp_enqueue_schedule_skb
                                  |                                           udpv6_queue_rcv_one_skb
                                  |                                           __udp6_lib_rcv
                                  |                                           ip6_input
                                  |                                           ipv6_rcv
                                  |                                           process_backlog
                                  |                                           net_rx_action
                                  |                                           |          
                                  |                                            --21.09%--__softirqentry_text_start
                                  |                                                      __local_bh_enable_ip
                                  |                                                      ip6_output
                                  |                                                      ip6_local_out
                                  |                                                      ip6_send_skb
                                  |                                                      udp_v6_send_skb
                                  |                                                      udpv6_sendmsg
                                  |                                                      __sys_sendto
                                  |                                                      __x64_sys_sendto
                                  |                                                      do_syscall_64
                                  |                                                      entry_SYSCALL_64
                                  |          
                                   --19.44%--swqueue_remove_task
                                             |          
                                              --19.42%--pick_next_task_fair
                                                        |          
                                                         --19.42%--schedule
                                                                   |          
                                                                    --19.21%--schedule_timeout
                                                                              __skb_wait_for_more_packets
                                                                              __skb_recv_udp
                                                                              udpv6_recvmsg
                                                                              inet6_recvmsg
                                                                              __x64_sys_recvfrom
                                                                              do_syscall_64
                                                                              entry_SYSCALL_64
    40.87%  netserver        [kernel.vmlinux]                                                 [k] queued_spin_lock_slowpath
            |          
             --40.87%--queued_spin_lock_slowpath
                       |          
                        --40.51%--_raw_spin_lock_irqsave
                                  |          
                                  |--21.03%--enqueue_task_fair
                                  |          |          
                                  |           --20.94%--default_wake_function
                                  |                     |          
                                  |                      --20.94%--autoremove_wake_function
                                  |                                |          
                                  |                                 --20.94%--__wake_up_sync_key
                                  |                                           sock_def_readable
                                  |                                           __udp_enqueue_schedule_skb
                                  |                                           udpv6_queue_rcv_one_skb
                                  |                                           __udp6_lib_rcv
                                  |                                           ip6_input
                                  |                                           ipv6_rcv
                                  |                                           process_backlog
                                  |                                           net_rx_action
                                  |                                           |          
                                  |                                            --20.94%--__softirqentry_text_start
                                  |                                                      __local_bh_enable_ip
                                  |                                                      ip6_output
                                  |                                                      ip6_local_out
                                  |                                                      ip6_send_skb
                                  |                                                      udp_v6_send_skb
                                  |                                                      udpv6_sendmsg
                                  |                                                      __sys_sendto
                                  |                                                      __x64_sys_sendto
                                  |                                                      do_syscall_64
                                  |                                                      entry_SYSCALL_64
                                  |          
                                   --19.48%--swqueue_remove_task
                                             |          
                                              --19.47%--pick_next_task_fair
                                                        schedule
                                                        |          
                                                         --19.38%--schedule_timeout
                                                                   __skb_wait_for_more_packets
                                                                   __skb_recv_udp
                                                                   udpv6_recvmsg
                                                                   inet6_recvmsg
                                                                   __x64_sys_recvfrom
                                                                   do_syscall_64
                                                                   entry_SYSCALL_64

Thanks for the help in getting the repro on my end.

So yes, there is certainly a scalability concern to bear in mind for
swqueue for LLCs with a lot of cores. If you have a lot of tasks quickly
e.g. blocking and waking on futexes in a tight loop, I expect a similar
issue would be observed.

On the other hand, the issue did not occur on my 7950X. I also wasn't
able to repro the contention on the Skylake if I ran with the default
netperf workload rather than UDP_RR (even with the additional clients).
I didn't bother to take the mean of all of the throughput results
between NO_SWQUEUE and SWQUEUE, but they looked roughly equal.

So swqueue isn't ideal for every configuration, but I'll echo my
sentiment from [0] that this shouldn't on its own necessarily preclude
it from being merged given that it does help a large class of
configurations and workloads, and it's disabled by default.

[0]: https://lore.kernel.org/all/20230615000103.GC2883716@maniforge/

Thanks,
David

  reply	other threads:[~2023-06-15 23:27 UTC|newest]

Thread overview: 46+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-06-13  5:20 [RFC PATCH 0/3] sched: Implement shared wakequeue in CFS David Vernet
2023-06-13  5:20 ` [RFC PATCH 1/3] sched: Make migrate_task_to() take any task David Vernet
2023-06-21 13:04   ` Peter Zijlstra
2023-06-22  2:07     ` David Vernet
2023-06-13  5:20 ` [RFC PATCH 2/3] sched/fair: Add SWQUEUE sched feature and skeleton calls David Vernet
2023-06-21 12:49   ` Peter Zijlstra
2023-06-22 14:53     ` David Vernet
2023-06-13  5:20 ` [RFC PATCH 3/3] sched: Implement shared wakequeue in CFS David Vernet
2023-06-13  8:32   ` Peter Zijlstra
2023-06-14  4:35     ` Aaron Lu
2023-06-14  9:27       ` Peter Zijlstra
2023-06-15  0:01       ` David Vernet
2023-06-15  4:49         ` Aaron Lu
2023-06-15  7:31           ` Aaron Lu
2023-06-15 23:26             ` David Vernet [this message]
2023-06-16  0:53               ` Aaron Lu
2023-06-20 17:36                 ` David Vernet
2023-06-21  2:35                   ` Aaron Lu
2023-06-21  2:43                     ` David Vernet
2023-06-21  4:54                       ` Aaron Lu
2023-06-21  5:43                         ` David Vernet
2023-06-21  6:03                           ` Aaron Lu
2023-06-22 15:57                             ` Chris Mason
2023-06-13  8:41   ` Peter Zijlstra
2023-06-14 20:26     ` David Vernet
2023-06-16  8:08   ` Vincent Guittot
2023-06-20 19:54     ` David Vernet
2023-06-20 21:37       ` Roman Gushchin
2023-06-21 14:22       ` Peter Zijlstra
2023-06-19  6:13   ` Gautham R. Shenoy
2023-06-20 20:08     ` David Vernet
2023-06-21  8:17       ` Gautham R. Shenoy
2023-06-22  1:43         ` David Vernet
2023-06-22  9:11           ` Gautham R. Shenoy
2023-06-22 10:29             ` Peter Zijlstra
2023-06-23  9:50               ` Gautham R. Shenoy
2023-06-26  6:04                 ` Gautham R. Shenoy
2023-06-27  3:17                   ` David Vernet
2023-06-27 16:31                     ` Chris Mason
2023-06-21 14:20   ` Peter Zijlstra
2023-06-21 20:34     ` David Vernet
2023-06-22 10:58       ` Peter Zijlstra
2023-06-22 14:43         ` David Vernet
2023-07-10 11:57 ` [RFC PATCH 0/3] " K Prateek Nayak
2023-07-11  4:43   ` David Vernet
2023-07-11  5:06     ` K Prateek Nayak

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20230615232605.GB2915572@maniforge \
    --to=void@manifault.com \
    --cc=aaron.lu@intel.com \
    --cc=bristot@redhat.com \
    --cc=bsegall@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=joshdon@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kernel-team@meta.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=roman.gushchin@linux.dev \
    --cc=rostedt@goodmis.org \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox