From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f200.google.com (mail-pf1-f200.google.com [209.85.210.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 49AAB4A653F for ; Thu, 6 Aug 2026 23:37:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.200 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786059423; cv=none; b=ra2brblrqknfirnxfX2VN3LGVykdqEz30+fjTEwWu10jf6iEJyeiHRLvEEhMU/5IIeDuP89EUbZ7gFCEeG7QEjW96+TO+MhC0LL4c6P3n9GIcu7t4371dYttAvltHuFIqyFclkWHdZV65EYrxT8C6h7xLxDxiXl+7dPRbPiJHLs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786059423; c=relaxed/simple; bh=rWbS/5EcUfIGw+MDz630kobNACYbTmDG/E889zv4mS0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=tEvkYpL5PCX5i+/5kohulcOivbtYJQFv5xtlM5mTbkwAAr2xiLi284Uvm4yGS9ngCxL8/a9n32gIATe4T0ltQoaDIfzKqFAmlULQOlUASHfqf6C0lwmRnvEWjAsVqzNYgUgGx+RwZrrkBxSh/bNT9dz2MA7uZjsXs2OBQM1RGec= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=FH6oJW9Q; arc=none smtp.client-ip=209.85.210.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="FH6oJW9Q" Received: by mail-pf1-f200.google.com with SMTP id d2e1a72fcca58-84a67b16217so4508396b3a.3 for ; Thu, 06 Aug 2026 16:37:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786059421; x=1786664221; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=o8F+mGATSkCnO7RB7Nrhalcr3db+GPA7pyXAMrOPMa0=; b=FH6oJW9Q9Y2BvkCSqOJ1mnEyLdt7Emx0DGlmxy6pyiJxb1y3PqniFwGdNpCFdI4crH Sjz/hN79/bxw42AiUXM8abq9WtKaRDNWAFBXzswFU+ORy/sKj7dlXE2rmrRA4Dk/8oVE fnwRlzs4rdBKMiNVLWR5yKIYWyql3dGwiqaJLjaIUZJvuHzd3SjEbyXwJNEVdPltRRHt DRMysF3qowb9p6s5mVLxBMIU6ikTbXrSTL+WnI/8/OqBzOTPwo0ui6eFHg/bfoTDKZyh riJBfTiISugE80comI7H7jjbnRXWsGWc9tF4q1RXehCNqY5EUuhgU2Uo1FNJYx1dkgwh ArzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786059421; x=1786664221; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=o8F+mGATSkCnO7RB7Nrhalcr3db+GPA7pyXAMrOPMa0=; b=oz+09lNcVTzpaIlyRNqhnFYGudYc0vJN0kw1vaiqbF+MsZQam/tJvl3v6y80ajx4pC Hxah62IZEvN3BePE8W+AoJlQmI+w+CHQyifL0xvM8m/8u4v7DaSZMESkDlJF+noARW3C CYpzAJSQ+zAtNVSs1/YQZ0JVAGNGhK3cgPUTiq3Uik1uYP7VOhsv9BUkXXAgHK66OHH3 PjCPBhd0GJyXtO+Y8bDIDDAz2Xk0H1/sGzGTh+cyEJHk1k1Myqyo01JRBfWmimv+5tlF QW6cwvlEajizSOXgMSA69pK/gX+pnlBSVZ70K7DICKrJChU+reDZsGYWNQcGli/E+Tn7 X6iw== X-Forwarded-Encrypted: i=1; AHgh+RpbhEaov0/SvoSK8YwRtDvV/V+aUd4ZcZ/2yO6fBhmuCFcMyoD7wGJHwZ4wM8HuVPndo+91eYzuOVN26ck=@vger.kernel.org X-Gm-Message-State: AOJu0YwAzcwtAgULjfzvt74B8LG9dLCfmbvk78TCScZs7cfIyT+2/AL/ L/0T3gnZ6MHijoEH5M0pQMfMcfj8q5ctexgSU1CIDWPtgvQ7qWg00LjeUl04TF96zvlqmzIQSNM fOd7IUw== X-Received: from pfwp49.prod.google.com ([2002:a05:6a00:26f1:b0:84a:894f:23f2]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:1daa:b0:848:401c:98e with SMTP id d2e1a72fcca58-84f2e0432a7mr20883725b3a.15.1786059420393; Thu, 06 Aug 2026 16:37:00 -0700 (PDT) Reply-To: Sean Christopherson Date: Thu, 6 Aug 2026 16:35:46 -0700 In-Reply-To: <20260806233609.212337-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260806233609.212337-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.679.g6767b8d81c-goog Message-ID: <20260806233609.212337-30-seanjc@google.com> Subject: [PATCH v6 29/51] x86/kvm: Don't disable kvmclock on BSP in syscore_suspend() From: Sean Christopherson To: Kiryl Shutsemau , Rick Edgecombe , Sean Christopherson , Paolo Bonzini , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Ajay Kaher , Alexey Makhalov , Jan Kiszka , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Juergen Gross , Daniel Lezcano , Thomas Gleixner , John Stultz Cc: Vitaly Kuznetsov , Broadcom internal kernel review list , Boris Ostrovsky , Stephen Boyd , Miroslav Lichvar , x86@kernel.org, linux-coco@lists.linux.dev, kvm@vger.kernel.org, linux-hyperv@vger.kernel.org, virtualization@lists.linux.dev, linux-kernel@vger.kernel.org, xen-devel@lists.xenproject.org, Michael Kelley , Tom Lendacky , Nikunj A Dadhania , David Woodhouse , David Woodhouse , Thomas Gleixner Content-Type: text/plain; charset="UTF-8" Don't disable kvmclock on the BSP during syscore_suspend(), as the BSP's clock is NOT restored during syscore_resume(), but is instead restored earlier via the sched_clock restore callback. If suspend is aborted, e.g. due to a late wakeup, the BSP will run without its clock enabled, which "works" only because KVM-the-hypervisor is kind enough to not clobber the shared memory when the clock is disabled. But over time, the BSP's view of time will drift from APs. Plumb in an "action" to KVM-as-a-guest and kvmclock code in preparation for additional cleanups to kvmclock's suspend/resume logic. Fixes: c02027b5742b ("x86/kvm: Disable kvmclock on all CPUs on shutdown") Cc: stable@vger.kernel.org Reviewed-by: David Woodhouse Signed-off-by: Sean Christopherson --- arch/x86/include/asm/kvm_para.h | 8 +++++++- arch/x86/kernel/kvm.c | 15 ++++++++------- arch/x86/kernel/kvmclock.c | 31 +++++++++++++++++++++++++------ 3 files changed, 40 insertions(+), 14 deletions(-) diff --git a/arch/x86/include/asm/kvm_para.h b/arch/x86/include/asm/kvm_para.h index 4a49fc286b4c..08686ff19caa 100644 --- a/arch/x86/include/asm/kvm_para.h +++ b/arch/x86/include/asm/kvm_para.h @@ -118,8 +118,14 @@ static inline long kvm_sev_hypercall3(unsigned int nr, unsigned long p1, } #ifdef CONFIG_KVM_GUEST +enum kvm_guest_cpu_action { + KVM_GUEST_BSP_SUSPEND, + KVM_GUEST_AP_OFFLINE, + KVM_GUEST_SHUTDOWN, +}; + void kvmclock_init(bool prefer_tsc); -void kvmclock_disable(void); +void kvmclock_cpu_action(enum kvm_guest_cpu_action action); bool kvm_para_available(void); unsigned int kvm_arch_para_features(void); unsigned int kvm_arch_para_hints(void); diff --git a/arch/x86/kernel/kvm.c b/arch/x86/kernel/kvm.c index 83b9351f2810..9c0b2fc98cc9 100644 --- a/arch/x86/kernel/kvm.c +++ b/arch/x86/kernel/kvm.c @@ -460,7 +460,7 @@ static void __init sev_map_percpu_data(void) } } -static void kvm_guest_cpu_offline(bool shutdown) +static void kvm_guest_cpu_offline(enum kvm_guest_cpu_action action) { kvm_disable_steal_time(); if (kvm_para_has_feature(KVM_FEATURE_PV_EOI)) @@ -468,9 +468,10 @@ static void kvm_guest_cpu_offline(bool shutdown) if (kvm_para_has_feature(KVM_FEATURE_MIGRATION_CONTROL)) wrmsrq(MSR_KVM_MIGRATION_CONTROL, 0); kvm_pv_disable_apf(); - if (!shutdown) + if (action != KVM_GUEST_SHUTDOWN) apf_task_wake_all(); - kvmclock_disable(); + + kvmclock_cpu_action(action); } static int kvm_cpu_online(unsigned int cpu) @@ -728,7 +729,7 @@ static int kvm_cpu_down_prepare(unsigned int cpu) unsigned long flags; local_irq_save(flags); - kvm_guest_cpu_offline(false); + kvm_guest_cpu_offline(KVM_GUEST_AP_OFFLINE); local_irq_restore(flags); return 0; } @@ -739,7 +740,7 @@ static int kvm_suspend(void *data) { u64 val = 0; - kvm_guest_cpu_offline(false); + kvm_guest_cpu_offline(KVM_GUEST_BSP_SUSPEND); #ifdef CONFIG_ARCH_CPUIDLE_HALTPOLL if (kvm_para_has_feature(KVM_FEATURE_POLL_CONTROL)) @@ -770,7 +771,7 @@ static struct syscore kvm_syscore = { static void kvm_pv_guest_cpu_reboot(void *unused) { - kvm_guest_cpu_offline(true); + kvm_guest_cpu_offline(KVM_GUEST_SHUTDOWN); } static int kvm_pv_reboot_notify(struct notifier_block *nb, @@ -794,7 +795,7 @@ static struct notifier_block kvm_pv_reboot_nb = { #ifdef CONFIG_CRASH_DUMP static void kvm_crash_shutdown(struct pt_regs *regs) { - kvm_guest_cpu_offline(true); + kvm_guest_cpu_offline(KVM_GUEST_SHUTDOWN); native_machine_crash_shutdown(regs); } #endif diff --git a/arch/x86/kernel/kvmclock.c b/arch/x86/kernel/kvmclock.c index b0c871ba8232..a3ec298d56d7 100644 --- a/arch/x86/kernel/kvmclock.c +++ b/arch/x86/kernel/kvmclock.c @@ -199,8 +199,22 @@ static void kvm_register_clock(char *txt) pr_debug("kvm-clock: cpu %d, msr %llx, %s", smp_processor_id(), pa, txt); } +static void kvmclock_disable(void) +{ + if (msr_kvm_system_time) + native_write_msr(msr_kvm_system_time, 0); +} + static void kvm_save_sched_clock_state(void) { + /* + * Stop host writes to kvmclock immediately prior to suspend/hibernate. + * If the system is hibernating, then kvmclock will likely reside at a + * different physical address when the system awakens, and host writes + * to the old address prior to reconfiguring kvmclock would clobber + * random memory. + */ + kvmclock_disable(); } static void kvm_restore_sched_clock_state(void) @@ -208,6 +222,17 @@ static void kvm_restore_sched_clock_state(void) kvm_register_clock("primary cpu clock, resume"); } +void kvmclock_cpu_action(enum kvm_guest_cpu_action action) +{ + /* + * Don't disable kvmclock on the BSP during suspend. If kvmclock is + * being used for sched_clock, then it needs to be kept alive until the + * last minute, and restored as quickly as possible after resume. + */ + if (action != KVM_GUEST_BSP_SUSPEND) + kvmclock_disable(); +} + #ifdef CONFIG_SMP static void kvm_setup_secondary_clock(void) { @@ -215,12 +240,6 @@ static void kvm_setup_secondary_clock(void) } #endif -void kvmclock_disable(void) -{ - if (msr_kvm_system_time) - native_write_msr(msr_kvm_system_time, 0); -} - static void __init kvmclock_init_mem(void) { unsigned long ncpus; -- 2.55.0.679.g6767b8d81c-goog