All of lore.kernel.org
 help / color / mirror / Atom feed
From: Hillf Danton <hdanton@sina.com>
To: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: linux-kernel@vger.kernel.org, peterz@infradead.org,
	seanjc@google.com, kprateek.nayak@amd.com
Subject: Re: [PATCH 01/17] sched/docs: Document cpu_paravirt_mask and Paravirt CPU concept
Date: Fri, 21 Nov 2025 06:56:00 +0800	[thread overview]
Message-ID: <20251120225603.9460-1-hdanton@sina.com> (raw)
In-Reply-To: <d6fb199e-fdba-4583-8229-6e6834a85a1b@linux.ibm.com>

On Thu, 20 Nov 2025 20:24:13 +0530 Shrikanth Hegde wrote:
> On 11/20/25 3:18 AM, Hillf Danton wrote:
> > On Wed, 19 Nov 2025 18:14:33 +0530 Shrikanth Hegde wrote:
> >> Add documentation for new cpumask called cpu_paravirt_mask. This could
> >> help users in understanding what this mask and the concept behind it.
> >>
> >> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
> >> ---
> >>   Documentation/scheduler/sched-arch.rst | 37 ++++++++++++++++++++++++++
> >>   1 file changed, 37 insertions(+)
> >>
> >> diff --git a/Documentation/scheduler/sched-arch.rst b/Documentation/scheduler/sched-arch.rst
> >> index ed07efea7d02..6972c295013d 100644
> >> --- a/Documentation/scheduler/sched-arch.rst
> >> +++ b/Documentation/scheduler/sched-arch.rst
> >> @@ -62,6 +62,43 @@ Your cpu_idle routines need to obey the following rules:
> >>   arch/x86/kernel/process.c has examples of both polling and
> >>   sleeping idle functions.
> >>   
> >> +Paravirt CPUs
> >> +=============
> >> +
> >> +Under virtualised environments it is possible to overcommit CPU resources.
> >> +i.e sum of virtual CPU(vCPU) of all VM's is greater than number of physical
> >> +CPUs(pCPU). Under such conditions when all or many VM's have high utilization,
> >> +hypervisor won't be able to satisfy the CPU requirement and has to context
> >> +switch within or across VM. i.e hypervisor need to preempt one vCPU to run
> >> +another. This is called vCPU preemption. This is more expensive compared to
> >> +task context switch within a vCPU.
> >> +
> > What is missing is
> > 1) vCPU preemption is X% more expensive compared to task context switch within a vCPU.
> > 
> 
> This would change from arch to arch IMO. Will try to get numbers from PowerVM hypervisor.
> 
> >> +In such cases it is better that VM's co-ordinate among themselves and ask for
> >> +less CPU by not using some of the vCPUs. Such vCPUs where workload can be
> >> +avoided at the moment for less vCPU preemption are called as "Paravirt CPUs".
> >> +Note that when the pCPU contention goes away, these vCPUs can be used again
> >> +by the workload.
> >> +
> > 2) given X, how to work out Y, the number of Paravirt CPUs for the simple
> > scenario like 8 pCPUs and 16 vCPUs (8 vCPUs from VM1, 8 vCPUs from VM2)?
> > 
> 
> Y need not be dependent on X. Note CPUs are marked as paravirt only when both VM's
> end up consuming all the CPU resource.
> 
To check that dependence, the frequence of vCPU preemption can be set to
100HZ and the frequence of task context switch within a vCPU to 250HZ,
on top of __zero__ Y (actually what we can do before this work), to compare
with the result of whatever Y this work can select.

BTW workload on vCPU can be compiling linux kernel with -j 8.

> Different cases:
> 1. VM1 is idle and VM2 is idle - No vCPUs are marked as paravirt.
> 2. VM1 is 100% busy and VM2 is idle - No steal time seen - No vCPUs is marked as paravirt.
> 3. VM1 is idle and VM2 is 100% busy - No steal time seen - No vCPUs is marked as paravirt.
> 4. VM1 is 100% busy and VM2 is 100% busy - 50% steal time would be seen in each -
> 	Since there are only 8 pCPUs (assuming each VM1 is allocated equally), 4 vCPUs in
> 	each VM will be marked as paravirt. Workload consolidates to remaining 4 vCPUs and
> 	hence no steal time will seen. Benefit would seen since host doesn't need to change
> 	expensive VM context switches.

  reply	other threads:[~2025-11-20 22:57 UTC|newest]

Thread overview: 45+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-11-19 12:44 [PATCH 00/17] Paravirt CPUs and push task for less vCPU preemption Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 01/17] sched/docs: Document cpu_paravirt_mask and Paravirt CPU concept Shrikanth Hegde
2025-11-19 21:48   ` Hillf Danton
2025-11-20 14:54     ` Shrikanth Hegde
2025-11-20 22:56       ` Hillf Danton [this message]
2025-11-19 12:44 ` [PATCH 02/17] cpumask: Introduce cpu_paravirt_mask Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 03/17] sched/core: Dont allow to use CPU marked as paravirt Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 04/17] sched/debug: Remove unused schedstats Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 05/17] sched/fair: Add paravirt movements for proc sched file Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 06/17] sched/fair: Pass current cpu in select_idle_sibling Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 07/17] sched/fair: Don't consider paravirt CPUs for wakeup and load balance Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 08/17] sched/rt: Don't select paravirt CPU for wakeup and push/pull rt task Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 09/17] sched/core: Add support for nohz_full CPUs Shrikanth Hegde
2025-11-21  3:16   ` K Prateek Nayak
2025-11-21  4:40     ` Shrikanth Hegde
2025-11-24  4:36       ` K Prateek Nayak
2025-11-19 12:44 ` [PATCH 10/17] sched/core: Push current task from paravirt CPU Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 11/17] sysfs: Add paravirt CPU file Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 12/17] powerpc: method to initialize ec and vp cores Shrikanth Hegde
2025-11-21  8:29   ` kernel test robot
2025-11-21 10:14   ` kernel test robot
2025-11-19 12:44 ` [PATCH 13/17] powerpc: enable/disable paravirt CPUs based on steal time Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 14/17] powerpc: process steal values at fixed intervals Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 15/17] powerpc: add debugfs file for controlling handling on steal values Shrikanth Hegde
2025-11-19 12:44 ` [PATCH 16/17] sysfs: Provide write method for paravirt Shrikanth Hegde
2025-11-24 17:04   ` Greg KH
2025-11-24 17:24     ` Steven Rostedt
2025-11-25  2:49       ` Shrikanth Hegde
2025-11-25 15:52         ` Steven Rostedt
2025-11-25 16:02           ` Konstantin Ryabitsev
2025-11-25 16:08             ` Steven Rostedt
2025-11-19 12:44 ` [PATCH 17/17] sysfs: disable arch handling if paravirt file being written Shrikanth Hegde
2025-11-24 17:05 ` [PATCH 00/17] Paravirt CPUs and push task for less vCPU preemption Greg KH
2025-11-25  2:39   ` Shrikanth Hegde
2025-11-25  7:48     ` Christophe Leroy (CS GROUP)
2025-11-25  8:48       ` Shrikanth Hegde
2025-11-27 10:44 ` Shrikanth Hegde
2025-12-04 13:28 ` Ilya Leoshkevich
2025-12-05  5:30   ` Shrikanth Hegde
2025-12-15 17:39     ` Yury Norov
2025-12-18  5:22       ` Shrikanth Hegde
2025-12-08  4:47 ` K Prateek Nayak
2025-12-08  9:57   ` Shrikanth Hegde
2025-12-08 17:58     ` K Prateek Nayak
2026-03-26  6:11 ` Shrikanth Hegde

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20251120225603.9460-1-hdanton@sina.com \
    --to=hdanton@sina.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=peterz@infradead.org \
    --cc=seanjc@google.com \
    --cc=sshegde@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.