All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64
@ 2026-08-27 14:31 Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 01/13] xen/arm: Add suspend and resume timer helpers Mykola Kvach
                   ` (13 more replies)
  0 siblings, 14 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Jens Wiklander, Rahul Singh

This is part 2 of the ARM Xen system suspend/resume patch series, based
on earlier work by Mirela Simonovic and Mykyta Poturai.

Part 1, covering guest suspend functionality, is already in mainline.

NOTE: Host-wide suspend/resume support is guarded by CONFIG_SYSTEM_SUSPEND,
which can currently only be selected when UNSUPPORTED is set, and thus the
host suspend backend is neither enabled by default nor built in supported
configurations. The separate HAS_HWDOM_SYSTEM_SUSPEND policy bit only changes
how ARM treats SHUTDOWN_suspend from the hardware domain; it does not enable
the host-wide suspend backend by itself.

This version is ported to Xen master and includes extensive improvements
based on reviewer feedback. The patch series restructures code to improve
robustness and maintainability, and implements initial ARM64 host-wide
Suspend-to-RAM support driven by control-domain PSCI SYSTEM_SUSPEND
requests. vPSCI also exposes SYSTEM_SUSPEND as a domain suspend operation
for all domains; attempt-time host-suspend policy failures are reported as
PSCI_DENIED rather than hidden through PSCI_FEATURES.

Key updates in this series:
 - Introduced architecture-specific suspend/resume infrastructure
 - Integrated GICv2/GICv3 suspend and resume, including memory-backed context
   save/restore with error handling
 - Added time and IRQ suspend/resume hooks, ensuring correct timer/interrupt
   state across suspend cycles
 - Implemented proper PSCI SYSTEM_SUSPEND invocation and version checks
 - Added vPSCI SYSTEM_SUSPEND policy for domain suspend and host-wide
   control-domain sequencing
 - Improved state management and recovery in error cases during suspend/resume
 - Added support for IPMMU-VMSA/SMMUv3 context save/restore
 - Added support for GICv3 eSPI registers context save/restore
 - Added support for ITS registers context save/restore
---

Link to CI: https://gitlab.com/xen-project/people/mykola_kvach/xen/-/pipelines/2796733872
---

TODOs:
 - Enable "xl suspend" support on ARM
 - Add suspend/resume CI test for ARM (QEMU if feasible)
 - PCI suspend ?
---

Detailed changelogs can be found in each patch.

Changes in v12:
- Rebase onto the latest master.
- Updated patch 12; all other patches are unchanged.
  No functional changes.

Changes in v11:
- Keep SMMUv3 reset helpers in init text when CONFIG_SYSTEM_SUSPEND is
  disabled.
- Update host suspend policy blockers after review: make the runtime gate
  __ro_after_init, log the SMMUv3 MSI blocker only once, and wrap the Arm
  IOMMU blocker in CONFIG_SYSTEM_SUSPEND.

Changes in v10:
- Clarify the vPSCI SYSTEM_SUSPEND policy summary: keep SYSTEM_SUSPEND
  advertised once implemented and return PSCI_DENIED, rather than
  PSCI_NOT_SUPPORTED, for attempt-time host-suspend policy failures.
- Tighten GICv2/GICv3 suspend/resume based on review feedback: avoid
  reserved interrupt register ranges, check visible active-priority state,
  restore configuration before enable state, and re-enable the redistributor
  before restoring CPU/virtual interface state on abort paths.
- Refine ITS resume so MAPC is replayed only for ITS-backed collections and
  clarify the collection-ID assumptions.
- Rework IPMMU and SMMUv3 resume/suspend handling, including root-before-cache
  IPMMU restore ordering and disabling SMMU interrupt generation before
  suspend.
- Save and restore CNTHCTL_EL2 in the arm64 CPU resume context and simplify
  the resume trampoline/context hand-off.
- Re-apply boot CPU errata/workaround handling after SYSTEM_SUSPEND and move
  set_init_ttbr() declaration to asm/mmu/mm.h.
- Update patch 12 details: shorten SYSTEM_SUSPEND blocker logs, use %pd for
  control-domain logging, mark serial_suspend_available as __ro_after_init,
  and mention the xen/suspend.h struct domain forward declaration.

Changes in v9:
- Split the control-domain SYSTEM_SUSPEND flow so host availability,
  runtime blockers and domain-readiness checks are handled separately from
  the host suspend backend.
- Gate vPSCI SYSTEM_SUSPEND on cached host PSCI support and Xen runtime
  suspend blockers, and log firmware support during initialization.
- Fold the arm64 resume trampoline into the CPU context save/restore patch
  and use asm-offsets-generated RESUME_CTX_* definitions for the assembly
  save/restore path.
- Tighten the GICv2/GICv3/ITS/IPMMU/SMMUv3 suspend/resume paths based on
  review feedback, including state-save/restore fixes and safer failure
  handling.
- Reorder the host suspend/resume phases so timer and GIC state are
  handled with local IRQs disabled and restored before console/IOMMU
  resume.

Changes in v8:
- Rebased to latest master and refreshed the series accordingly.
- Added a new GICv3 patch to tolerate retained redistributor LPI state
  across CPU_OFF/CPU_ON.
- GICv2 suspend now disables the CPU interface and distributor before
  saving state.
- GICv3 suspend/resume fixes the redistributor base used for LPI state.
- ITS and SMMUv3 suspend/resume paths were tightened, with safer
  restore/rollback handling and stricter fatal-error handling.
- System suspend now checks that all domains are already in
  SHUTDOWN_suspend before proceeding, and renames the hardware-domain
  suspend capability/helper for clearer semantics.
- Fixed alignment/cleanup issues in the low-level suspend/resume code.

Changes in v7:
- Timer helper renamed/clarified; virtual/hyper/phys handling documented.
- GICv2 uses one context block; restore saved CTLR; panic on alloc failure.
- GICv3/eSPI/ITS always suspend/resume; restore LPI/eSPI; rdist timeout.
- IPMMU suspend context allocated before PCI setup.
- System suspend: control domain drives host suspend.
- Dropped v6 IRQ descriptor restore patches; use setup_irq and re-register
  local IRQs on resume instead.

For earlier changelogs, please refer to the previous cover letters.

Mirela Simonovic (5):
  xen/arm: Add suspend and resume timer helpers
  xen/arm: gic-v2: Implement GIC suspend/resume functions
  xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND
  xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface)
  xen/arm: Add host system suspend backend

Mykola Kvach (7):
  xen/arm: gic-v3: tolerate retained redistributor LPI state across
    CPU_OFF
  xen/arm: gic-v3: Implement GICv3 suspend/resume functions
  xen/arm: gic-v3: add ITS suspend/resume support
  xen/arm: tee: keep init_tee_secondary() for hotplug and resume
  xen/arm: ffa: fix notification SRI across CPU hotplug/suspend
  xen/arm: smmu-v3: add suspend/resume handlers
  xen/arm: Add vPSCI SYSTEM_SUSPEND policy

Oleksandr Tyshchenko (1):
  iommu/ipmmu-vmsa: Implement suspend/resume callbacks

 xen/arch/arm/Kconfig                     |   2 +
 xen/arch/arm/Makefile                    |   1 +
 xen/arch/arm/arm64/asm-offsets.c         |  21 +
 xen/arch/arm/arm64/head.S                | 122 ++++++
 xen/arch/arm/cpuerrata.c                 |   7 +-
 xen/arch/arm/gic-v2.c                    | 226 +++++++++++
 xen/arch/arm/gic-v3-its.c                | 146 ++++++-
 xen/arch/arm/gic-v3-lpi.c                |  80 +++-
 xen/arch/arm/gic-v3.c                    | 482 ++++++++++++++++++++++-
 xen/arch/arm/gic.c                       |  35 ++
 xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
 xen/arch/arm/include/asm/cpuerrata.h     |   1 +
 xen/arch/arm/include/asm/gic.h           |  16 +
 xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
 xen/arch/arm/include/asm/gic_v3_its.h    |  28 ++
 xen/arch/arm/include/asm/mmu/mm.h        |   2 +
 xen/arch/arm/include/asm/psci.h          |   4 +
 xen/arch/arm/include/asm/suspend.h       |  37 ++
 xen/arch/arm/include/asm/time.h          |   5 +
 xen/arch/arm/mmu/smpboot.c               |   2 +-
 xen/arch/arm/psci.c                      |  38 +-
 xen/arch/arm/suspend.c                   | 210 ++++++++++
 xen/arch/arm/tee/ffa_notif.c             |  63 ++-
 xen/arch/arm/tee/tee.c                   |   2 +-
 xen/arch/arm/time.c                      |  44 ++-
 xen/arch/arm/vpsci.c                     | 120 +++++-
 xen/common/Kconfig                       |   3 +
 xen/common/domain.c                      |   7 +-
 xen/drivers/char/serial.c                |  12 +
 xen/drivers/passthrough/arm/iommu.c      |   6 +
 xen/drivers/passthrough/arm/ipmmu-vmsa.c | 323 ++++++++++++++-
 xen/drivers/passthrough/arm/smmu-v3.c    | 203 ++++++++--
 xen/include/xen/list.h                   |  14 +
 xen/include/xen/serial.h                 |   1 +
 xen/include/xen/suspend.h                |   2 +
 35 files changed, 2169 insertions(+), 104 deletions(-)
 create mode 100644 xen/arch/arm/suspend.c

-- 
2.43.0



^ permalink raw reply	[flat|nested] 37+ messages in thread

* [PATCH v12 01/13] xen/arm: Add suspend and resume timer helpers
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions Mykola Kvach
                   ` (12 subsequent siblings)
  13 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Julien Grall, Luca Fancellu

From: Mirela Simonovic <mirela.simonovic@aggios.com>

Timer interrupts must be disabled while the system is suspended to prevent
spurious wake-ups. Suspending timers in Xen consists of disabling the
physical timer and the hypervisor timer on the current CPU. The virtual
timer does not need explicit handling here, as it is already disabled on
vCPU context switch and its state is restored per-vCPU on the next context
restore.

Resuming consists of raising TIMER_SOFTIRQ, which prompts the generic
timer code to reprogram the hypervisor timer with the correct timeout.

Xen does not use or expose the physical timer, so it remains disabled
across suspend/resume.

Introduce a new helper, disable_phys_hyp_timers(), to encapsulate disabling
of the physical and hypervisor timers.

Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Acked-by: Julien Grall <jgrall@amazon.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V7:
  - Dropped EL1/EL2 wording; use "physical timer" and "hypervisor timer"
  - Renamed helper to disable_phys_hyp_timers() to reflect its actual scope
  - Clarified virtual timer handling (disabled on vCPU switch-out, restored
    on context restore) and added comments in suspend/resume paths
  - Added resume comment explaining which timers are restored by
    TIMER_SOFTIRQ
---
 xen/arch/arm/include/asm/time.h |  5 ++++
 xen/arch/arm/time.c             | 44 ++++++++++++++++++++++++++++-----
 2 files changed, 43 insertions(+), 6 deletions(-)

diff --git a/xen/arch/arm/include/asm/time.h b/xen/arch/arm/include/asm/time.h
index c194dbb9f5..9313b157ea 100644
--- a/xen/arch/arm/include/asm/time.h
+++ b/xen/arch/arm/include/asm/time.h
@@ -105,6 +105,11 @@ void preinit_xen_time(void);
 
 void force_update_vcpu_system_time(struct vcpu *v);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+void time_suspend(void);
+void time_resume(void);
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 #endif /* __ARM_TIME_H__ */
 /*
  * Local variables:
diff --git a/xen/arch/arm/time.c b/xen/arch/arm/time.c
index be54b87438..680535a2ca 100644
--- a/xen/arch/arm/time.c
+++ b/xen/arch/arm/time.c
@@ -298,6 +298,14 @@ static void check_timer_irq_cfg(unsigned int irq, const char *which)
 static DEFINE_PER_CPU_READ_MOSTLY(struct irqaction, irq_hyp);
 static DEFINE_PER_CPU_READ_MOSTLY(struct irqaction, irq_virt);
 
+/* Disable physical and hypervisor timers on the current CPU */
+static inline void disable_phys_hyp_timers(void)
+{
+    WRITE_SYSREG(0, CNTP_CTL_EL0);    /* Physical timer disabled */
+    WRITE_SYSREG(0, CNTHP_CTL_EL2);   /* Hypervisor's timer disabled */
+    isb();
+}
+
 /* Set up the timer interrupt on this CPU */
 void init_timer_interrupt(void)
 {
@@ -308,9 +316,7 @@ void init_timer_interrupt(void)
     WRITE_SYSREG64(0, CNTVOFF_EL2);     /* No VM-specific offset */
     /* Do not let the VMs program the physical timer, only read the physical counter */
     WRITE_SYSREG(CNTHCTL_EL2_EL1PCTEN, CNTHCTL_EL2);
-    WRITE_SYSREG(0, CNTP_CTL_EL0);    /* Physical timer disabled */
-    WRITE_SYSREG(0, CNTHP_CTL_EL2);   /* Hypervisor's timer disabled */
-    isb();
+    disable_phys_hyp_timers();
 
     hyp_action->name = "hyptimer";
     hyp_action->handler = htimer_interrupt;
@@ -335,9 +341,7 @@ void init_timer_interrupt(void)
  */
 static void deinit_timer_interrupt(void)
 {
-    WRITE_SYSREG(0, CNTP_CTL_EL0);    /* Disable physical timer */
-    WRITE_SYSREG(0, CNTHP_CTL_EL2);   /* Disable hypervisor's timer */
-    isb();
+    disable_phys_hyp_timers();
 
     release_irq(timer_irq[TIMER_HYP_PPI], NULL);
     release_irq(timer_irq[TIMER_VIRT_PPI], NULL);
@@ -377,6 +381,34 @@ void domain_set_time_offset(struct domain *d, int64_t time_offset_seconds)
     /* XXX update guest visible wallclock time */
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+void time_suspend(void)
+{
+    /* CNTV already disabled by virt_timer_save() during vcpu context switch. */
+    disable_phys_hyp_timers();
+}
+
+void time_resume(void)
+{
+    /*
+     * Raising TIMER_SOFTIRQ triggers generic timer code to reprogram the
+     * hypervisor timer with the correct timeout (not known here).
+     *
+     * Xen doesn't use or expose the physical timer, so it remains disabled
+     * across suspend/resume.
+     *
+     * The virtual timer state is restored per-vCPU on the next context switch.
+     *
+     * No further action is needed to restore timekeeping after power down,
+     * since the system counter is unaffected. See ARM DDI 0487 L.a, D12.1.2
+     * "The system counter must be implemented in an always-on power domain."
+     */
+    raise_softirq(TIMER_SOFTIRQ);
+}
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 static int cpu_time_callback(struct notifier_block *nfb,
                              unsigned long action,
                              void *hcpu)
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 01/13] xen/arm: Add suspend and resume timer helpers Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-23 15:27   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF Mykola Kvach
                   ` (11 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

From: Mirela Simonovic <mirela.simonovic@aggios.com>

System suspend may lead to a state where GIC would be powered down.
Therefore, Xen should save/restore the context of GIC on suspend/resume.

Note that the context consists of states of registers which are
controlled by the hypervisor. Other GIC registers which are accessible
by guests are saved/restored on context switch.

Transient physical SGI pending state (GICD_CPENDSGIRn/GICD_SPENDSGIRn)
is intentionally excluded. CPU-interface active-priority state is also
not restored across suspend/resume. Xen reaches the final suspend path
at a quiescent point, so there is no active-priority execution context
to replay after resume. Enforce this with a runtime check after
disabling the CPU interface: if any implemented GICC_APRn word is still
non-zero, restore GICC_CTLR and abort suspend with -EBUSY.

This does not apply to distributor active state. With GICv2 EOImode==1,
EOIR only drops the interrupt priority; final deactivation is a separate
step. For guest-routed interrupts, Xen can have already EOIed the physical
IRQ while deactivation is still pending on the vGIC/GICV path. Therefore
GICD_ISACTIVER is preserved as architectural in-flight interrupt state.

Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V10:
- Limit GICC_APR<n> active-priority checks to APR bits visible from
  the Xen CPU-interface view.
- Avoid touching reserved GICD_IPRIORITYR/GICD_ITARGETSR words when the
  last implemented interrupt block is partial.
- Restore distributor configuration before restoring interrupt enable
  state, so GICD_ICFGR is written while the corresponding interrupts are
  disabled.

Changes in V9:
- Skip saving/restoring GICD_ITARGETSR0..7 because SGI/PPI target
  registers hold no state (read-only on MP, RAZ/WI on UP).
- Add a runtime GICC_APRn quiescence check after disabling the CPU
  interface, and restore GICC_CTLR before returning -EBUSY.

Changes in V8:
- disable cpu interface + distributor before suspend
- change 0xffffffff to GENMASK;
- cosmetic changes;

Changes in V7:
- Allocate one contiguous memory block for the GICv2 dist suspend context.
- gicv2_resume() no longer unconditionally re-enables the distributor/CPU
  interface; it now writes back the saved CTLR values as-is.
- gicv2_alloc_context() now returns 0 on success and panics on failure,
  since suspend context allocation is not recoverable.
---
 xen/arch/arm/gic-v2.c          | 226 +++++++++++++++++++++++++++++++++
 xen/arch/arm/gic.c             |  29 +++++
 xen/arch/arm/include/asm/gic.h |  12 ++
 3 files changed, 267 insertions(+)

diff --git a/xen/arch/arm/gic-v2.c b/xen/arch/arm/gic-v2.c
index 43a379fdda..a0ef6ffc7f 100644
--- a/xen/arch/arm/gic-v2.c
+++ b/xen/arch/arm/gic-v2.c
@@ -1108,6 +1108,223 @@ static int gicv2_iomem_deny_access(struct domain *d)
     return iomem_deny_access(d, mfn, mfn + nr - 1);
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+/* This struct represents block of 32 IRQs */
+struct irq_block {
+    uint32_t icfgr[2]; /* 2 registers of 16 IRQs each */
+    uint32_t ipriorityr[8];
+    uint32_t isenabler;
+    uint32_t isactiver;
+    uint32_t itargetsr[8];
+};
+
+/* GICv2 registers to be saved/restored on system suspend/resume */
+struct gicv2_context {
+    /* GICC context */
+    struct cpu_ctx {
+        uint32_t ctlr;
+        uint32_t pmr;
+        uint32_t bpr;
+    } cpu;
+
+    /* GICD context */
+    struct dist_ctx {
+        uint32_t ctlr;
+        /* Includes banked SGI/PPI state for the boot CPU. */
+        struct irq_block *irqs;
+    } dist;
+};
+
+static struct gicv2_context gic_ctx;
+
+#define GICV2_NR_APRS          4
+#define GICV2_APR_BITS_PER_REG 32U
+
+static int gicv2_check_active_priorities(uint32_t bpr)
+{
+    unsigned int i, apr_bits, nr_aprs;
+
+    /*
+     * Xen writes GICC_BPR to 0 during CPU init and does not change it. Per
+     * IHI0048B.b, a write below the implementation minimum reads back as the
+     * minimum supported BPR value. Table 4-47 maps that Xen-visible BPR value
+     * to the visible GICC_APR<n> bits. Avoid reading APR registers outside
+     * that visible range.
+     *
+     * This covers both GICv2 with and without Security Extensions.
+     */
+    apr_bits = 1U << (7 - (bpr & 0x7));
+    nr_aprs = DIV_ROUND_UP(apr_bits, GICV2_APR_BITS_PER_REG);
+
+    ASSERT(nr_aprs <= GICV2_NR_APRS);
+
+    for ( i = 0; i < nr_aprs; i++ )
+    {
+        unsigned int bits = min(GICV2_APR_BITS_PER_REG,
+                                apr_bits - i * GICV2_APR_BITS_PER_REG);
+        uint32_t mask = GENMASK(bits - 1, 0);
+        uint32_t apr = readl_gicc(GICC_APR + i * 4) & mask;
+
+        if ( !apr )
+            continue;
+
+        printk(XENLOG_ERR "GICv2: suspend aborted: GICC_APR%u=%#08x\n",
+               i, apr);
+        return -EBUSY;
+    }
+
+    return 0;
+}
+
+static int gicv2_suspend(void)
+{
+    unsigned int i, blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
+    int ret;
+
+    /* Save GICC_CTLR configuration. */
+    gic_ctx.cpu.ctlr = readl_gicc(GICC_CTLR);
+
+    /* Quiesce the GIC CPU interface before suspend. */
+    gicv2_cpu_disable();
+
+    gic_ctx.cpu.bpr = readl_gicc(GICC_BPR);
+
+    /*
+     * Check the active-priority state for the group Xen drives through the
+     * CPU interface. GICC_CTL_ENABLE enables Group 0 without SecurityExtn and
+     * Group 1 in Xen's Non-secure view with SecurityExtn, and in both cases
+     * the relevant state is visible through GICC_APRn. The APR layout is
+     * implementation-defined, so only test the bits visible from Xen's CPU
+     * interface view instead of reading every possible APR register.
+     */
+    ret = gicv2_check_active_priorities(gic_ctx.cpu.bpr);
+    if ( ret )
+    {
+        writel_gicc(gic_ctx.cpu.ctlr, GICC_CTLR);
+        return ret;
+    }
+
+    gic_ctx.cpu.pmr = readl_gicc(GICC_PMR);
+
+    /* Save GICD configuration */
+    gic_ctx.dist.ctlr = readl_gicd(GICD_CTLR);
+    writel_gicd(0, GICD_CTLR);
+
+    for ( i = 0; i < blocks; i++ )
+    {
+        struct irq_block *irqs = gic_ctx.dist.irqs + i;
+        size_t j, off = i * sizeof(irqs->isenabler);
+        size_t nr_regs = ARRAY_SIZE(irqs->ipriorityr);
+
+        if ( i == blocks - 1 )
+            nr_regs = DIV_ROUND_UP(gicv2_info.nr_lines - i * 32, 4);
+
+        irqs->isenabler = readl_gicd(GICD_ISENABLER + off);
+
+        /*
+         * Save distributor active state as part of the hypervisor-owned
+         * physical interrupt state. In GICv2 EOImode==1, EOIR only drops the
+         * priority; final deactivation is separate. For guest-routed
+         * interrupts, Xen may have EOIed the physical IRQ while the guest/vGIC
+         * side still owns the deactivate step. Therefore GICD_ISACTIVER can
+         * legitimately remain set even though transient SGI pending state and
+         * CPU-interface active-priority state are expected to be quiesced here.
+         */
+        irqs->isactiver = readl_gicd(GICD_ISACTIVER + off);
+
+        off = i * sizeof(irqs->ipriorityr);
+        for ( j = 0; j < nr_regs; j++ )
+            irqs->ipriorityr[j] = readl_gicd(GICD_IPRIORITYR + off + j * 4);
+
+        /*
+         * GICD_ITARGETSR0..7 cover SGIs/PPIs and hold no state to save:
+         * they are read-only on multiprocessor implementations and RAZ/WI
+         * on uniprocessor implementations.
+         */
+        if ( i )
+        {
+            off = i * sizeof(irqs->itargetsr);
+            for ( j = 0; j < nr_regs; j++ )
+                irqs->itargetsr[j] = readl_gicd(GICD_ITARGETSR + off + j * 4);
+        }
+
+        off = i * sizeof(irqs->icfgr);
+        for ( j = 0; j < ARRAY_SIZE(irqs->icfgr); j++ )
+            irqs->icfgr[j] = readl_gicd(GICD_ICFGR + off + j * 4);
+    }
+
+    return 0;
+}
+
+static void gicv2_resume(void)
+{
+    unsigned int i, blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
+
+    gicv2_cpu_disable();
+    /* Disable distributor */
+    writel_gicd(0, GICD_CTLR);
+
+    for ( i = 0; i < blocks; i++ )
+    {
+        struct irq_block *irqs = gic_ctx.dist.irqs + i;
+        size_t j, off = i * sizeof(irqs->isenabler);
+        size_t nr_regs = ARRAY_SIZE(irqs->ipriorityr);
+
+        if ( i == blocks - 1 )
+            nr_regs = DIV_ROUND_UP(gicv2_info.nr_lines - i * 32, 4);
+
+        writel_gicd(GENMASK(31, 0), GICD_ICENABLER + off);
+
+        off = i * sizeof(irqs->icfgr);
+        for ( j = 0; j < ARRAY_SIZE(irqs->icfgr); j++ )
+            writel_gicd(irqs->icfgr[j], GICD_ICFGR + off + j * 4);
+
+        off = i * sizeof(irqs->ipriorityr);
+        for ( j = 0; j < nr_regs; j++ )
+            writel_gicd(irqs->ipriorityr[j], GICD_IPRIORITYR + off + j * 4);
+
+        /*
+         * GICD_ITARGETSR0..7 cover SGIs/PPIs and hold no state to save:
+         * they are read-only on multiprocessor implementations and RAZ/WI
+         * on uniprocessor implementations.
+         */
+        if ( i )
+        {
+            off = i * sizeof(irqs->itargetsr);
+            for ( j = 0; j < nr_regs; j++ )
+                writel_gicd(irqs->itargetsr[j], GICD_ITARGETSR + off + j * 4);
+        }
+
+        off = i * sizeof(irqs->isenabler);
+        writel_gicd(irqs->isenabler, GICD_ISENABLER + off);
+
+        writel_gicd(GENMASK(31, 0), GICD_ICACTIVER + off);
+        writel_gicd(irqs->isactiver, GICD_ISACTIVER + off);
+    }
+
+    /* Restore distributor control state. */
+    writel_gicd(gic_ctx.dist.ctlr, GICD_CTLR);
+
+    /* Restore GIC CPU interface configuration */
+    writel_gicc(gic_ctx.cpu.pmr, GICC_PMR);
+    writel_gicc(gic_ctx.cpu.bpr, GICC_BPR);
+
+    /* Enable GIC CPU interface */
+    writel_gicc(gic_ctx.cpu.ctlr, GICC_CTLR);
+}
+
+static void __init gicv2_alloc_context(void)
+{
+    uint32_t blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
+
+    gic_ctx.dist.irqs = xzalloc_array(struct irq_block, blocks);
+    if ( !gic_ctx.dist.irqs )
+        panic("Failed to allocate memory for GICv2 suspend context\n");
+}
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 #ifdef CONFIG_ACPI
 static unsigned long gicv2_get_hwdom_extra_madt_size(const struct domain *d)
 {
@@ -1312,6 +1529,11 @@ static int __init gicv2_init(void)
 
     spin_unlock(&gicv2.lock);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+    /* Allocate memory to be used for saving GIC context during the suspend */
+    gicv2_alloc_context();
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
     return 0;
 }
 
@@ -1355,6 +1577,10 @@ static const struct gic_hw_operations gicv2_ops = {
     .map_hwdom_extra_mappings = gicv2_map_hwdom_extra_mappings,
     .iomem_deny_access   = gicv2_iomem_deny_access,
     .do_LPI              = gicv2_do_LPI,
+#ifdef CONFIG_SYSTEM_SUSPEND
+    .suspend             = gicv2_suspend,
+    .resume              = gicv2_resume,
+#endif /* CONFIG_SYSTEM_SUSPEND */
 };
 
 /* Set up the GIC */
diff --git a/xen/arch/arm/gic.c b/xen/arch/arm/gic.c
index 078049e741..ffc11f36a1 100644
--- a/xen/arch/arm/gic.c
+++ b/xen/arch/arm/gic.c
@@ -438,6 +438,35 @@ int gic_iomem_deny_access(struct domain *d)
     return gic_hw_ops->iomem_deny_access(d);
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+int gic_suspend(void)
+{
+    /* Must be called by boot CPU#0 with interrupts disabled */
+    ASSERT(!local_irq_is_enabled());
+    ASSERT(!smp_processor_id());
+
+    if ( !gic_hw_ops->suspend || !gic_hw_ops->resume )
+        return -ENOSYS;
+
+    return gic_hw_ops->suspend();
+}
+
+void gic_resume(void)
+{
+    /*
+     * Must be called by boot CPU#0 with interrupts disabled after gic_suspend
+     * has returned successfully.
+     */
+    ASSERT(!local_irq_is_enabled());
+    ASSERT(!smp_processor_id());
+    ASSERT(gic_hw_ops->resume);
+
+    gic_hw_ops->resume();
+}
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 static int cpu_gic_callback(struct notifier_block *nfb,
                             unsigned long action,
                             void *hcpu)
diff --git a/xen/arch/arm/include/asm/gic.h b/xen/arch/arm/include/asm/gic.h
index ee2c26adb4..29bb9a89a4 100644
--- a/xen/arch/arm/include/asm/gic.h
+++ b/xen/arch/arm/include/asm/gic.h
@@ -301,6 +301,12 @@ extern int gicv_setup(struct domain *d);
 extern void gic_save_state(struct vcpu *v);
 extern void gic_restore_state(struct vcpu *v);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+/* Suspend/resume */
+extern int gic_suspend(void);
+extern void gic_resume(void);
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 /* SGI (AKA IPIs) */
 enum gic_sgi {
     GIC_SGI_EVENT_CHECK,
@@ -444,6 +450,12 @@ struct gic_hw_operations {
     int (*iomem_deny_access)(struct domain *d);
     /* Handle LPIs, which require special handling */
     void (*do_LPI)(unsigned int lpi);
+#ifdef CONFIG_SYSTEM_SUSPEND
+    /* Save GIC configuration due to the system suspend */
+    int (*suspend)(void);
+    /* Restore GIC configuration due to the system resume */
+    void (*resume)(void);
+#endif /* CONFIG_SYSTEM_SUSPEND */
 };
 
 extern const struct gic_hw_operations *gic_hw_ops;
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 01/13] xen/arm: Add suspend and resume timer helpers Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-23 15:34   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions Mykola Kvach
                   ` (10 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

PSCI does not guarantee that a GICv3 redistributor is powered down across
CPU_OFF -> CPU_ON.

DEN0022F.b says CPU_OFF powers down the calling core (5.5) and CPU_ON
brings the core back with a defined initial CPU state (5.6, 6.4).
However, PSCI leaves interrupt migration and GIC re-initialization to the
supervisory software/firmware stack: the caller must migrate interrupts
away before CPU_OFF (5.5.2), and the execution context that is lost in a
powerdown state must be saved and restored by software (6.8). PSCI also
calls out GIC management explicitly in 6.8, including retargeting SPIs,
preventing PPIs/SGIs from targeting a powered down CPU, and reinitializing
the CPU interface after CPU_ON.

This matches the GIC architecture. IHI0069H.b Chapter 11.1 requires the PE
and CPU interface to share a power domain, but explicitly allows the
associated redistributor, distributor, and ITS to remain powered while the
PE and CPU interface are off. All other GIC power-management behavior is
IMPLEMENTATION DEFINED. DEN0050D Chapter 4.2, "Generic Interrupt
Controller (GIC)", says the GICv3 redistributor may live either in the AP
core power domain or in a relatively always-on parent domain. So after
CPU_OFF -> CPU_ON a secondary CPU can legitimately come back to a live
redistributor with GICR_CTLR.EnableLPIs still set.

Handle that case in the LPI setup path instead of assuming a fully reset
redistributor.

The LPI path needs special care because the GIC spec makes redistributor
LPI state sticky and partially implementation defined. IHI0069H.b 5.1.1
and 5.1.2 say that changing GICR_PROPBASER or GICR_PENDBASER while
GICR_CTLR.EnableLPIs == 1 is UNPREDICTABLE. After clearing EnableLPIs,
software must wait for GICR_CTLR.RWP == 0 before touching the pending
table. The architecture also permits implementations where, once
EnableLPIs has been set, clearing it again is not guaranteed to work.
Where an ITS is present, the spec strongly recommends moving LPIs to
another redistributor before clearing EnableLPIs.

Because of that, treat a retained EnableLPIs state as valid when the
redistributor still points at Xen's expected PROPBASER/PENDBASER tables.
Only try to clear EnableLPIs when the retained configuration does not
match Xen's state, and wait for RWP before reprogramming the tables.

This is also consistent with platform firmware reality: PSCI and the GIC
architecture allow platform-specific redistributor power handling, and not
all platform firmware implementations force a full redistributor power-off
through implementation-defined controls during CPU_OFF. Xen therefore needs
to tolerate retained redistributor state on secondary CPU bring-up.

Keep gicv3_populate_rdist() resident as well, because gicv3_cpu_init()
reuses it on secondary CPU bring-up after init.

Tested using Xen's non-boot CPU disable/enable path on Arm
FVP_Base_RevC-2xAEMvA, both with and without:
-C gic_distributor.allow-LPIEN-clear=1
-C gic_distributor.GICR-clear-enable-supported=1
and on Orange Pi 5.

Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in v10:
- Drop unrelated gicv3_populate_rdist() printk() format cleanups to keep
  the patch focused on retained redistributor LPI state.

Changes in v9:
- move gicv3_do_wait_for_rwp prototype from its related header to gic.h
- drop __init from gicv3_populate_rdist(), which is reused on secondary
  CPU bring-up after boot
- changed print format for smp_processor_id in gicv3_populate_rdist func
- cosmetic changes
---
 xen/arch/arm/gic-v3-lpi.c      | 77 +++++++++++++++++++++++++++++++++-
 xen/arch/arm/gic-v3.c          | 15 ++++---
 xen/arch/arm/include/asm/gic.h |  4 ++
 3 files changed, 90 insertions(+), 6 deletions(-)

diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
index 9ee338edc2..847da26ff7 100644
--- a/xen/arch/arm/gic-v3-lpi.c
+++ b/xen/arch/arm/gic-v3-lpi.c
@@ -81,6 +81,13 @@ static DEFINE_PER_CPU(struct lpi_redist_data, lpi_redist);
 #define MAX_NR_HOST_LPIS   (lpi_data.max_host_lpi_ids - LPI_OFFSET)
 #define HOST_LPIS_PER_PAGE      (PAGE_SIZE / sizeof(union host_lpi))
 
+#define GICR_PROPBASER_XEN_MASK  GENMASK_ULL(51, 12)
+/*
+ * For retained redistributor state, match the pending table by address only.
+ * Attribute bits such as PTZ may not read back with the programmed value.
+ */
+#define GICR_PENDBASER_XEN_MASK  GENMASK_ULL(51, 16)
+
 static union host_lpi *gic_get_host_lpi(uint32_t plpi)
 {
     union host_lpi *block;
@@ -296,6 +303,60 @@ static int gicv3_lpi_set_pendtable(void __iomem *rdist_base)
     return 0;
 }
 
+static uint64_t gicv3_lpi_expected_proptable(void)
+{
+    return virt_to_maddr(lpi_data.lpi_property);
+}
+
+static uint64_t gicv3_lpi_expected_pendtable(void)
+{
+    return virt_to_maddr(this_cpu(lpi_redist).pending_table);
+}
+
+static bool gicv3_lpi_tables_match(void __iomem *rdist_base)
+{
+    uint64_t propbase, pendbase;
+
+    if ( !lpi_data.lpi_property || !this_cpu(lpi_redist).pending_table )
+        return false;
+
+    propbase = readq_relaxed(rdist_base + GICR_PROPBASER);
+    pendbase = readq_relaxed(rdist_base + GICR_PENDBASER);
+
+    return ((propbase & GICR_PROPBASER_XEN_MASK) ==
+            (gicv3_lpi_expected_proptable() & GICR_PROPBASER_XEN_MASK)) &&
+           ((pendbase & GICR_PENDBASER_XEN_MASK) ==
+            (gicv3_lpi_expected_pendtable() & GICR_PENDBASER_XEN_MASK));
+}
+
+static int gicv3_lpi_disable_lpis(void __iomem *rdist_base)
+{
+    uint32_t reg = readl_relaxed(rdist_base + GICR_CTLR);
+    int ret;
+
+    if ( !(reg & GICR_CTLR_ENABLE_LPIS) )
+        return 0;
+
+    writel_relaxed(reg & ~GICR_CTLR_ENABLE_LPIS, rdist_base + GICR_CTLR);
+
+    /*
+     * The spec only guarantees programmability when we have observed the bit
+     * cleared. Where clearing is supported, RWP must reach 0 before touching
+     * PROPBASER/PENDBASER again.
+     */
+    wmb();
+
+    ret = gicv3_do_wait_for_rwp(rdist_base, GICR_CTLR_RWP);
+    if ( ret )
+        return ret;
+
+    reg = readl_relaxed(rdist_base + GICR_CTLR);
+    if ( reg & GICR_CTLR_ENABLE_LPIS )
+        return -EBUSY;
+
+    return 0;
+}
+
 /*
  * Tell a redistributor about the (shared) property table, allocating one
  * if not already done.
@@ -374,7 +435,21 @@ int gicv3_lpi_init_rdist(void __iomem * rdist_base)
     /* Make sure LPIs are disabled before setting up the tables. */
     reg = readl_relaxed(rdist_base + GICR_CTLR);
     if ( reg & GICR_CTLR_ENABLE_LPIS )
-        return -EBUSY;
+    {
+        if ( gicv3_lpi_tables_match(rdist_base) )
+            return -EBUSY;
+
+        ret = gicv3_lpi_disable_lpis(rdist_base);
+        if ( ret == -EBUSY )
+        {
+            printk(XENLOG_ERR
+                   "GICv3: CPU%u: LPIs still enabled with unexpected redistributor tables\n",
+                   smp_processor_id());
+            return -EINVAL;
+        }
+        if ( ret )
+            return ret;
+    }
 
     ret = gicv3_lpi_set_pendtable(rdist_base);
     if ( ret )
diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
index acdac22953..b16888ad84 100644
--- a/xen/arch/arm/gic-v3.c
+++ b/xen/arch/arm/gic-v3.c
@@ -275,7 +275,7 @@ static void gicv3_enable_sre(void)
 }
 
 /* Wait for completion of a distributor/redistributor change */
-static void gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
+int gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
 {
     uint32_t val;
     bool timeout = false;
@@ -299,17 +299,22 @@ static void gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
     } while ( 1 );
 
     if ( timeout )
+    {
         dprintk(XENLOG_ERR, "RWP timeout\n");
+        return -ETIMEDOUT;
+    }
+
+    return 0;
 }
 
 static void gicv3_dist_wait_for_rwp(void)
 {
-    gicv3_do_wait_for_rwp(GICD, GICD_CTLR_RWP);
+    (void)gicv3_do_wait_for_rwp(GICD, GICD_CTLR_RWP);
 }
 
 static void gicv3_redist_wait_for_rwp(void)
 {
-    gicv3_do_wait_for_rwp(GICD_RDIST_BASE, GICR_CTLR_RWP);
+    (void)gicv3_do_wait_for_rwp(GICD_RDIST_BASE, GICR_CTLR_RWP);
 }
 
 static void gicv3_wait_for_rwp(int irq)
@@ -863,7 +868,7 @@ static bool gicv3_enable_lpis(void)
     return true;
 }
 
-static int __init gicv3_populate_rdist(void)
+static int gicv3_populate_rdist(void)
 {
     int i;
     uint32_t aff;
@@ -931,7 +936,7 @@ static int __init gicv3_populate_rdist(void)
                     gicv3_set_redist_address(rdist_addr, procnum);
 
                     ret = gicv3_lpi_init_rdist(ptr);
-                    if ( ret && ret != -ENODEV )
+                    if ( ret && ret != -ENODEV && ret != -EBUSY )
                     {
                         printk("GICv3: CPU%d: Cannot initialize LPIs: %u\n",
                                smp_processor_id(), ret);
diff --git a/xen/arch/arm/include/asm/gic.h b/xen/arch/arm/include/asm/gic.h
index 29bb9a89a4..68003ab116 100644
--- a/xen/arch/arm/include/asm/gic.h
+++ b/xen/arch/arm/include/asm/gic.h
@@ -301,6 +301,10 @@ extern int gicv_setup(struct domain *d);
 extern void gic_save_state(struct vcpu *v);
 extern void gic_restore_state(struct vcpu *v);
 
+#ifdef CONFIG_GICV3
+int gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit);
+#endif
+
 #ifdef CONFIG_SYSTEM_SUSPEND
 /* Suspend/resume */
 extern int gic_suspend(void);
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (2 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-23 15:35   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support Mykola Kvach
                   ` (9 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

System suspend may lead to a state where GIC would be powered down.
Therefore, Xen should save/restore the context of GIC on suspend/resume.

Note that the context consists of states of registers which are
controlled by the hypervisor. Other GIC registers which are accessible
by guests are saved/restored on context switch.

Before continuing suspend, also verify that the physical CPU interface
has no Group 1 active-priority state left. Use ICC_CTLR_EL1.PRIbits to
decide which ICC_AP1R<n>_EL1 registers are implemented, so Xen does not
read an unimplemented AP1R register.

Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V10:
- abort suspend when the physical Group 1 active-priority state is still
  present, deriving accessible ICC_AP1R<n>_EL1 registers from
  ICC_CTLR_EL1.PRIbits;
- re-enable the redistributor before restoring CPU and virtual interface
  state on the suspend abort path;
- panic if the redistributor cannot be re-enabled on the suspend abort path;
- avoid saving/restoring reserved GICD_IPRIORITYR and GICD_IROUTER entries
  for a partially populated last SPI block;
- disable Distributor group forwarding while preserving affinity routing
  state before restoring Distributor configuration;
- disable SPI/eSPI forwarding and wait for RWP before restoring
  GICD_ICFGR<n>.Int_config.

Changes in V9:
- fix the suspend-context comment typo and split dist_ctx declarations;
- restore ICC_IGRPEN1_EL1 on the suspend error path;
- re-initialize GICD_IGROUPRnE during resume;
- restore GICD_IROUTER only after re-enabling ARE_NS during resume.

Changes in V8:
- use right rdist base for prop/pend baser and ctrl

Changes in V7:
- restore LPI regs on resume
- add timeout during redist disabling
- squash with suspend/resume handling for GICv3 eSPI registers
- drop ITS guard paths so suspend/resume always runs; switch missing ctx
  allocation to panic
- trim TODO comments; narrow redistributor storage to PPI icfgr
- keep distributor context allocation even without ITS; adjust resume
  to use GENMASK(31, 0) for clearing enables
- drop storage of the SGI configuration register, as SGIs are always
  edge-triggered
---
 xen/arch/arm/gic-v3-lpi.c                |   3 +
 xen/arch/arm/gic-v3.c                    | 458 ++++++++++++++++++++++-
 xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
 xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
 4 files changed, 466 insertions(+), 3 deletions(-)

diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
index 847da26ff7..a63c8c4979 100644
--- a/xen/arch/arm/gic-v3-lpi.c
+++ b/xen/arch/arm/gic-v3-lpi.c
@@ -467,6 +467,9 @@ static int cpu_callback(struct notifier_block *nfb, unsigned long action,
     switch ( action )
     {
     case CPU_UP_PREPARE:
+        if ( system_state == SYS_STATE_resume )
+            break;
+
         rc = gicv3_lpi_allocate_pendtable(cpu);
         if ( rc )
             printk(XENLOG_ERR "Unable to allocate the pendtable for CPU%lu\n",
diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
index b16888ad84..038bf41142 100644
--- a/xen/arch/arm/gic-v3.c
+++ b/xen/arch/arm/gic-v3.c
@@ -1078,12 +1078,12 @@ out:
     return res;
 }
 
-static void gicv3_hyp_disable(void)
+static void gicv3_hyp_enable(bool enable)
 {
     register_t hcr;
 
     hcr = READ_SYSREG(ICH_HCR_EL2);
-    hcr &= ~GICH_HCR_EN;
+    hcr = enable ? (hcr | GICH_HCR_EN) : (hcr & ~GICH_HCR_EN);
     WRITE_SYSREG(hcr, ICH_HCR_EL2);
     isb();
 }
@@ -1190,7 +1190,7 @@ static void gicv3_disable_interface(void)
     spin_lock(&gicv3.lock);
 
     gicv3_cpu_disable();
-    gicv3_hyp_disable();
+    gicv3_hyp_enable(false);
 
     spin_unlock(&gicv3.lock);
 }
@@ -1926,6 +1926,450 @@ static bool gic_dist_supports_lpis(void)
     return (readl_relaxed(GICD + GICD_TYPER) & GICD_TYPE_LPIS);
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+/* This struct represents a block of 32 IRQs */
+struct dist_irq_block {
+    uint32_t icfgr[2];
+    uint32_t ipriorityr[8];
+    uint64_t irouter[32];
+    uint32_t isactiver;
+    uint32_t isenabler;
+};
+
+struct redist_ctx {
+    uint32_t ctlr;
+    uint32_t icfgr; /* only PPIs stored */
+    uint32_t igroupr;
+    uint32_t ipriorityr[8];
+    uint32_t isactiver;
+    uint32_t isenabler;
+
+    uint64_t pendbase;
+    uint64_t propbase;
+};
+
+/* GICv3 registers to be saved/restored on system suspend/resume */
+struct gicv3_ctx {
+    struct dist_ctx {
+        uint32_t ctlr;
+        struct dist_irq_block *irqs;
+        struct dist_irq_block *espi_irqs;
+    } dist;
+
+    /* have only one rdist structure for last running CPU during suspend */
+    struct redist_ctx rdist;
+
+    struct cpu_ctx {
+        uint32_t ctlr;
+        uint32_t pmr;
+        uint32_t bpr;
+        uint32_t sre_el2;
+        uint32_t grpen;
+    } cpu;
+};
+
+static struct gicv3_ctx gicv3_ctx;
+
+static void __init gicv3_alloc_context(void)
+{
+    uint32_t blocks = DIV_ROUND_UP(gicv3_info.nr_lines, 32);
+
+    /* The spec allows for systems without any SPIs */
+    if ( blocks > 1 )
+    {
+        gicv3_ctx.dist.irqs = xzalloc_array(struct dist_irq_block, blocks - 1);
+        if ( !gicv3_ctx.dist.irqs )
+            panic("Failed to allocate memory for GICv3 suspend context\n");
+    }
+
+#ifdef CONFIG_GICV3_ESPI
+    if ( !gic_number_espis() )
+        return;
+
+    blocks = gic_number_espis() / 32;
+    gicv3_ctx.dist.espi_irqs = xzalloc_array(struct dist_irq_block, blocks);
+    if ( !gicv3_ctx.dist.espi_irqs )
+        panic("Failed to allocate memory for GICv3 eSPI suspend context\n");
+#endif
+}
+
+static int gicv3_disable_redist(void)
+{
+    void __iomem *waker = GICD_RDIST_BASE + GICR_WAKER;
+    s_time_t deadline;
+
+    /*
+     * Avoid infinite loop if Non-secure does not have access to GICR_WAKER.
+     * See Arm IHI 0069H.b, 12.11.42 GICR_WAKER:
+     *     When GICD_CTLR.DS == 0 and an access is Non-secure accesses to this
+     *     register are RAZ/WI.
+     */
+    if ( !(readl_relaxed(GICD + GICD_CTLR) & GICD_CTLR_DS) )
+        return 0;
+
+    deadline = NOW() + MILLISECS(1000);
+
+    writel_relaxed(readl_relaxed(waker) | GICR_WAKER_ProcessorSleep, waker);
+    while ( (readl_relaxed(waker) & GICR_WAKER_ChildrenAsleep) == 0 )
+    {
+        if ( NOW() > deadline )
+        {
+            printk("GICv3: Timeout waiting for redistributor to sleep\n");
+            return -ETIMEDOUT;
+        }
+        cpu_relax();
+        udelay(10);
+    }
+
+    return 0;
+}
+
+#define GET_SPI_REG_OFFSET(name, is_espi) \
+    ((is_espi) ? GICD_##name##nE : GICD_##name)
+
+static void gicv3_store_spi_irq_block(struct dist_irq_block *irqs,
+                                      unsigned int i, unsigned int nr_irqs,
+                                      bool is_espi)
+{
+    void __iomem *base;
+    unsigned int irq, nr_priority_regs;
+
+    ASSERT(nr_irqs && nr_irqs <= 32);
+    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
+
+    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
+    irqs->icfgr[0] = readl_relaxed(base);
+    irqs->icfgr[1] = readl_relaxed(base + 4);
+
+    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
+    base += i * sizeof(irqs->ipriorityr);
+    for ( irq = 0; irq < nr_priority_regs; irq++ )
+        irqs->ipriorityr[irq] = readl_relaxed(base + 4 * irq);
+
+    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
+    base += i * sizeof(irqs->irouter);
+    for ( irq = 0; irq < nr_irqs; irq++ )
+        irqs->irouter[irq] = readq_relaxed_non_atomic(base + 8 * irq);
+
+    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
+    base += i * sizeof(irqs->isactiver);
+    irqs->isactiver = readl_relaxed(base);
+
+    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
+    base += i * sizeof(irqs->isenabler);
+    irqs->isenabler = readl_relaxed(base);
+}
+
+static void gicv3_restore_spi_irq_config(struct dist_irq_block *irqs,
+                                         unsigned int i, unsigned int nr_irqs,
+                                         bool is_espi)
+{
+    void __iomem *base;
+    unsigned int irq, nr_priority_regs;
+
+    ASSERT(nr_irqs && nr_irqs <= 32);
+    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
+
+    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
+    writel_relaxed(irqs->icfgr[0], base);
+    writel_relaxed(irqs->icfgr[1], base + 4);
+
+    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
+    base += i * sizeof(irqs->ipriorityr);
+    for ( irq = 0; irq < nr_priority_regs; irq++ )
+        writel_relaxed(irqs->ipriorityr[irq], base + 4 * irq);
+}
+
+static void gicv3_restore_spi_irq_routing(struct dist_irq_block *irqs,
+                                          unsigned int i, unsigned int nr_irqs,
+                                          bool is_espi)
+{
+    void __iomem *base;
+    unsigned int irq;
+
+    ASSERT(nr_irqs && nr_irqs <= 32);
+
+    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
+    base += i * sizeof(irqs->irouter);
+    for ( irq = 0; irq < nr_irqs; irq++ )
+        writeq_relaxed_non_atomic(irqs->irouter[irq], base + 8 * irq);
+}
+
+static void gicv3_disable_spi_irq_block(unsigned int i, bool is_espi)
+{
+    void __iomem *base;
+
+    base = GICD + GET_SPI_REG_OFFSET(ICENABLER, is_espi) + i * 4;
+    writel_relaxed(GENMASK(31, 0), base);
+}
+
+static void gicv3_restore_spi_irq_state(struct dist_irq_block *irqs,
+                                        unsigned int i, bool is_espi)
+{
+    void __iomem *base;
+
+    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
+    base += i * sizeof(irqs->isenabler);
+    writel_relaxed(irqs->isenabler, base);
+
+    base = GICD + GET_SPI_REG_OFFSET(ICACTIVER, is_espi) + i * 4;
+    writel_relaxed(GENMASK(31, 0), base);
+
+    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
+    base += i * sizeof(irqs->isactiver);
+    writel_relaxed(irqs->isactiver, base);
+}
+
+static int gicv3_check_ap1r(unsigned int n, register_t apr)
+{
+    if ( !apr )
+        return 0;
+
+    printk(XENLOG_ERR "GICv3: suspend aborted: ICC_AP1R%u_EL1=%#"
+           PRIregister"\n", n, apr);
+
+    return -EBUSY;
+}
+
+static int gicv3_check_active_priorities(register_t ctlr)
+{
+    unsigned int pribits = MASK_EXTR(ctlr, ICC_CTLR_EL1_PRIBITS_MASK) + 1;
+    int ret;
+
+    /*
+     * Xen enables physical Group 1 interrupts through ICC_IGRPEN1_EL1,
+     * so only the physical Group 1 active-priority registers are relevant
+     * here. Use ICC_CTLR_EL1.PRIbits for the physical CPU interface, not
+     * ICH_VTR_EL2, which describes the virtual interface. ICC_AP1R1_EL1 is
+     * only implemented with at least 6 physical priority bits, and
+     * ICC_AP1R2_EL1/ICC_AP1R3_EL1 with at least 7.
+     */
+    switch ( pribits )
+    {
+    case 8:
+    case 7:
+        ret = gicv3_check_ap1r(3, READ_SYSREG(ICC_AP1R3_EL1));
+        if ( ret )
+            return ret;
+        ret = gicv3_check_ap1r(2, READ_SYSREG(ICC_AP1R2_EL1));
+        if ( ret )
+            return ret;
+        /* Fall through */
+    case 6:
+        ret = gicv3_check_ap1r(1, READ_SYSREG(ICC_AP1R1_EL1));
+        if ( ret )
+            return ret;
+        /* Fall through */
+    default:
+        return gicv3_check_ap1r(0, READ_SYSREG(ICC_AP1R0_EL1));
+    }
+}
+
+static int gicv3_suspend(void)
+{
+    unsigned int i, nr_irqs;
+    void __iomem *base;
+    int ret;
+    struct redist_ctx *rdist = &gicv3_ctx.rdist;
+
+    /* Save GICC configuration */
+    gicv3_ctx.cpu.ctlr     = READ_SYSREG(ICC_CTLR_EL1);
+    gicv3_ctx.cpu.pmr      = READ_SYSREG(ICC_PMR_EL1);
+    gicv3_ctx.cpu.bpr      = READ_SYSREG(ICC_BPR1_EL1);
+    gicv3_ctx.cpu.sre_el2  = READ_SYSREG(ICC_SRE_EL2);
+    gicv3_ctx.cpu.grpen    = READ_SYSREG(ICC_IGRPEN1_EL1);
+
+    gicv3_disable_interface();
+
+    ret = gicv3_check_active_priorities(gicv3_ctx.cpu.ctlr);
+    if ( ret )
+        goto out_enable_iface;
+
+    ret = gicv3_disable_redist();
+    if ( ret )
+        goto out_enable_iface;
+
+    /* Save GICR configuration */
+    gicv3_redist_wait_for_rwp();
+
+    base = GICD_RDIST_BASE;
+
+    rdist->ctlr = readl_relaxed(base + GICR_CTLR);
+
+    rdist->propbase = readq_relaxed(base + GICR_PROPBASER);
+    rdist->pendbase = readq_relaxed(base + GICR_PENDBASER);
+
+    base = GICD_RDIST_SGI_BASE;
+
+    /* Save priority on PPI and SGI interrupts */
+    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
+        rdist->ipriorityr[i] = readl_relaxed(base + GICR_IPRIORITYR0 + 4 * i);
+
+    rdist->isactiver = readl_relaxed(base + GICR_ISACTIVER0);
+    rdist->isenabler = readl_relaxed(base + GICR_ISENABLER0);
+    rdist->igroupr   = readl_relaxed(base + GICR_IGROUPR0);
+    rdist->icfgr     = readl_relaxed(base + GICR_ICFGR1);
+
+    /* Save GICD configuration */
+    gicv3_dist_wait_for_rwp();
+    gicv3_ctx.dist.ctlr = readl_relaxed(GICD + GICD_CTLR);
+
+    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
+    {
+        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
+        gicv3_store_spi_irq_block(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
+                                  false);
+    }
+
+#ifdef CONFIG_GICV3_ESPI
+    for ( i = 0; i < gic_number_espis() / 32; i++ )
+        gicv3_store_spi_irq_block(gicv3_ctx.dist.espi_irqs + i, i, 32, true);
+#endif
+
+    return 0;
+
+ out_enable_iface:
+    if ( gicv3_enable_redist() )
+        panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
+
+    gicv3_hyp_enable(true);
+    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
+    isb();
+
+    return ret;
+}
+
+static void gicv3_resume(void)
+{
+    int ret;
+    unsigned int i, nr_irqs;
+    uint32_t dist_ctlr;
+    void __iomem *base;
+    struct redist_ctx *rdist = &gicv3_ctx.rdist;
+
+    dist_ctlr = gicv3_ctx.dist.ctlr & GICD_CTLR_ARE_NS;
+
+    /* Disable group forwarding while preserving affinity routing state. */
+    writel_relaxed(dist_ctlr, GICD + GICD_CTLR);
+    gicv3_dist_wait_for_rwp();
+
+    /*
+     * IHI0069H.b 12.9.9 says changing GICD_ICFGR<n>.Int_config
+     * while the interrupt is individually enabled is UNPREDICTABLE.
+     * Disable SPIs first; 4.7.1 defines GICD_ICENABLER<n>, n > 0,
+     * as the per-SPI disable mechanism.
+     */
+    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
+        gicv3_disable_spi_irq_block(i, false);
+
+#ifdef CONFIG_GICV3_ESPI
+    for ( i = 0; i < gic_number_espis() / 32; i++ )
+        gicv3_disable_spi_irq_block(i, true);
+#endif
+
+    gicv3_dist_wait_for_rwp();
+
+    for ( i = NR_GIC_LOCAL_IRQS; i < gicv3_info.nr_lines; i += 32 )
+        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPR + (i / 32) * 4);
+
+    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
+    {
+        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
+        gicv3_restore_spi_irq_config(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
+                                     false);
+    }
+
+#ifdef CONFIG_GICV3_ESPI
+    for ( i = 0; i < gic_number_espis() / 32; i++ )
+    {
+        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPRnE + i * 4);
+        gicv3_restore_spi_irq_config(gicv3_ctx.dist.espi_irqs + i, i, 32,
+                                     true);
+    }
+#endif
+
+    if ( dist_ctlr )
+    {
+        for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
+        {
+            nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
+            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.irqs + i - 1, i,
+                                          nr_irqs, false);
+        }
+
+#ifdef CONFIG_GICV3_ESPI
+        for ( i = 0; i < gic_number_espis() / 32; i++ )
+            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.espi_irqs + i, i,
+                                          32, true);
+#endif
+    }
+
+    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
+        gicv3_restore_spi_irq_state(gicv3_ctx.dist.irqs + i - 1, i, false);
+
+#ifdef CONFIG_GICV3_ESPI
+    for ( i = 0; i < gic_number_espis() / 32; i++ )
+        gicv3_restore_spi_irq_state(gicv3_ctx.dist.espi_irqs + i, i, true);
+#endif
+
+    writel_relaxed(gicv3_ctx.dist.ctlr, GICD + GICD_CTLR);
+    gicv3_dist_wait_for_rwp();
+
+    ret = gicv3_lpi_init_rdist(GICD_RDIST_BASE);
+    /*
+     * If LPIs are already enabled, assume firmware or the still-powered
+     * redistributor has valid PROPBASER/PENDBASER and skip reprogramming.
+     * Return -EBUSY so callers can ignore this case.
+     */
+    if ( ret && ret != -ENODEV && ret != -EBUSY )
+        panic("GICv3: Failed to re-initialize LPIs during resume\n");
+    else if ( ret == -EBUSY ) /* extra checks, just to be sure */
+    {
+        base = GICD_RDIST_BASE;
+        if ( readq_relaxed(base + GICR_PROPBASER) != rdist->propbase ||
+             readq_relaxed(base + GICR_PENDBASER) != rdist->pendbase )
+            panic("GICv3: LPIs already enabled with unexpected PROPBASER/PENDBASER during resume\n");
+    }
+
+    /* Restore GICR (Redistributor) configuration */
+    if ( gicv3_enable_redist() )
+        panic("GICv3: Failed to re-enable redistributor during resume\n");
+
+    base = GICD_RDIST_SGI_BASE;
+
+    writel_relaxed(GENMASK(31, 0), base + GICR_ICENABLER0);
+    gicv3_redist_wait_for_rwp();
+
+    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
+        writel_relaxed(rdist->ipriorityr[i], base + GICR_IPRIORITYR0 + i * 4);
+
+    writel_relaxed(rdist->isactiver, base + GICR_ISACTIVER0);
+    writel_relaxed(rdist->igroupr,   base + GICR_IGROUPR0);
+    writel_relaxed(rdist->icfgr,     base + GICR_ICFGR1);
+
+    gicv3_redist_wait_for_rwp();
+
+    writel_relaxed(rdist->isenabler, base + GICR_ISENABLER0);
+    writel_relaxed(rdist->ctlr, GICD_RDIST_BASE + GICR_CTLR);
+
+    gicv3_redist_wait_for_rwp();
+
+    WRITE_SYSREG(gicv3_ctx.cpu.sre_el2, ICC_SRE_EL2);
+    isb();
+
+    /* Restore CPU interface (System registers) */
+    WRITE_SYSREG(gicv3_ctx.cpu.pmr,   ICC_PMR_EL1);
+    WRITE_SYSREG(gicv3_ctx.cpu.bpr,   ICC_BPR1_EL1);
+    WRITE_SYSREG(gicv3_ctx.cpu.ctlr,  ICC_CTLR_EL1);
+    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
+    isb();
+
+    gicv3_hyp_init();
+}
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 /* Set up the GIC */
 static int __init gicv3_init(void)
 {
@@ -2011,6 +2455,10 @@ static int __init gicv3_init(void)
 
     gicv3_hyp_init();
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+    gicv3_alloc_context();
+#endif
+
 out:
     spin_unlock(&gicv3.lock);
 
@@ -2050,6 +2498,10 @@ static const struct gic_hw_operations gicv3_ops = {
 #endif
     .iomem_deny_access   = gicv3_iomem_deny_access,
     .do_LPI              = gicv3_do_LPI,
+#ifdef CONFIG_SYSTEM_SUSPEND
+    .suspend             = gicv3_suspend,
+    .resume              = gicv3_resume,
+#endif
 };
 
 static int __init gicv3_dt_preinit(struct dt_device_node *node, const void *data)
diff --git a/xen/arch/arm/include/asm/arm64/sysregs.h b/xen/arch/arm/include/asm/arm64/sysregs.h
index f3c11d871e..2261620316 100644
--- a/xen/arch/arm/include/asm/arm64/sysregs.h
+++ b/xen/arch/arm/include/asm/arm64/sysregs.h
@@ -16,6 +16,11 @@
 #define ICC_SRE_EL1               S3_0_C12_C12_5
 #define ICC_IGRPEN1_EL1           S3_0_C12_C12_7
 
+#define ICC_AP1R0_EL1             S3_0_C12_C9_0
+#define ICC_AP1R1_EL1             S3_0_C12_C9_1
+#define ICC_AP1R2_EL1             S3_0_C12_C9_2
+#define ICC_AP1R3_EL1             S3_0_C12_C9_3
+
 #define ICH_VSEIR_EL2             S3_4_C12_C9_4
 #define ICC_SRE_EL2               S3_4_C12_C9_5
 #define ICH_HCR_EL2               S3_4_C12_C11_0
diff --git a/xen/arch/arm/include/asm/gic_v3_defs.h b/xen/arch/arm/include/asm/gic_v3_defs.h
index 3714cfeb7d..f741587322 100644
--- a/xen/arch/arm/include/asm/gic_v3_defs.h
+++ b/xen/arch/arm/include/asm/gic_v3_defs.h
@@ -94,12 +94,15 @@
 #define GICD_TYPE_LPIS               (1U << 17)
 
 #define GICD_CTLR_RWP                (1UL << 31)
+#define GICD_CTLR_DS                 (1U << 6)
 #define GICD_CTLR_ARE_NS             (1U << 4)
 #define GICD_CTLR_ENABLE_G1A         (1U << 1)
 #define GICD_CTLR_ENABLE_G1          (1U << 0)
 #define GICD_IROUTER_SPI_MODE_ANY    (1UL << 31)
 
 #define GICC_CTLR_EL1_EOImode_drop   (1U << 1)
+#define ICC_CTLR_EL1_PRIBITS_SHIFT   8
+#define ICC_CTLR_EL1_PRIBITS_MASK    (0x7U << ICC_CTLR_EL1_PRIBITS_SHIFT)
 
 #define GICR_WAKER_ProcessorSleep    (1U << 1)
 #define GICR_WAKER_ChildrenAsleep    (1U << 2)
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (3 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-23 15:36   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 06/13] xen/arm: tee: keep init_tee_secondary() for hotplug and resume Mykola Kvach
                   ` (8 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Luca Fancellu

Handle system suspend/resume for GICv3 with an ITS present so LPIs keep
working after firmware powers the GIC down.

Save and restore the ITS CTLR, CBASER and BASER registers. On resume,
re-establish the collection mapping only when the collection is held in
the ITS itself. Memory-backed collections are restored through the
restored GITS_BASER tables and must not be remapped unconditionally.

Add list_for_each_entry_continue_reverse() in list.h for the ITS suspend
error path that needs to roll back partially saved state.

Based on Linux commit dba0bc7b76dc:
"irqchip/gic-v3-its: Add ability to save/restore ITS state".
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V10:
- Replay MAPC on resume only for collections held in the ITS itself, as
  indicated by GITS_TYPER.HCC. Memory-backed collections are restored
  through GITS_BASER and are no longer remapped unconditionally.
- Make the current Xen col_id == cpu assumption explicit in the ITS
  resume path.
- Use "unpredictable" instead of "undefined" in the CBASER/BASER restore
  comment.

Changes in V9:
- fix the ITS suspend/resume coding-style nits;
- preserve the saved GITS_CTLR state while masking the read-only
  QUIESCENT bit.

Changes in V8:
- Reword the CBASER/CWRITER comment to match Xen and drop the stale Linux
  cmd_write reference.
- Clarify the list_for_each_entry_continue_reverse() comment.
- Factor out per-ITS helpers for collection setup and resume.
- Restore each ITS and re-establish its collection mapping in the same
  loop, so a failed ITS resume is not followed by MAPC/SYNC on that
  un-restored instance.
- panic in case when resume of an ITS failed
- cleanup baser cache during suspend
---
 xen/arch/arm/gic-v3-its.c             | 146 ++++++++++++++++++++++++--
 xen/arch/arm/gic-v3.c                 |  11 +-
 xen/arch/arm/include/asm/gic_v3_its.h |  28 +++++
 xen/include/xen/list.h                |  14 +++
 4 files changed, 189 insertions(+), 10 deletions(-)

diff --git a/xen/arch/arm/gic-v3-its.c b/xen/arch/arm/gic-v3-its.c
index 7560d46c6d..dd53209865 100644
--- a/xen/arch/arm/gic-v3-its.c
+++ b/xen/arch/arm/gic-v3-its.c
@@ -335,6 +335,22 @@ static int its_send_cmd_inv(struct host_its *its,
     return its_send_command(its, cmd);
 }
 
+static int gicv3_its_setup_collection_single(struct host_its *its,
+                                             unsigned int cpu)
+{
+    int ret;
+
+    ret = its_send_cmd_mapc(its, cpu, cpu);
+    if ( ret )
+        return ret;
+
+    ret = its_send_cmd_sync(its, cpu);
+    if ( ret )
+        return ret;
+
+    return gicv3_its_wait_commands(its);
+}
+
 /* Set up the (1:1) collection mapping for the given host CPU. */
 int gicv3_its_setup_collection(unsigned int cpu)
 {
@@ -343,15 +359,7 @@ int gicv3_its_setup_collection(unsigned int cpu)
 
     list_for_each_entry(its, &host_its_list, entry)
     {
-        ret = its_send_cmd_mapc(its, cpu, cpu);
-        if ( ret )
-            return ret;
-
-        ret = its_send_cmd_sync(its, cpu);
-        if ( ret )
-            return ret;
-
-        ret = gicv3_its_wait_commands(its);
+        ret = gicv3_its_setup_collection_single(its, cpu);
         if ( ret )
             return ret;
     }
@@ -1211,6 +1219,126 @@ int gicv3_its_init(void)
     return 0;
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+int gicv3_its_suspend(void)
+{
+    struct host_its *its;
+    int ret;
+
+    list_for_each_entry( its, &host_its_list, entry )
+    {
+        unsigned int i;
+        void __iomem *base = its->its_base;
+
+        /*
+         * By the time Xen reaches gic_suspend(), every domain is already in
+         * SHUTDOWN_suspend, so ITS-targeting interrupt sources are expected
+         * to have been quiesced by the owning OS before SYSTEM_SUSPEND.
+         */
+        /* Preserve saved GITS_CTLR state, excluding read-only QUIESCENT. */
+        its->suspend_ctx.ctlr = readl_relaxed(base + GITS_CTLR) &
+                                ~GITS_CTLR_QUIESCENT;
+        ret = gicv3_disable_its(its);
+        if ( ret )
+        {
+            writel_relaxed(its->suspend_ctx.ctlr, base + GITS_CTLR);
+            goto err;
+        }
+
+        its->suspend_ctx.cbaser = readq_relaxed(base + GITS_CBASER);
+
+        for ( i = 0; i < GITS_BASER_NR_REGS; i++ )
+        {
+            uint64_t baser = readq_relaxed(base + GITS_BASER0 + i * 8);
+
+            its->suspend_ctx.baser[i] = 0;
+
+            if ( !(baser & GITS_VALID_BIT) )
+                continue;
+
+            its->suspend_ctx.baser[i] = baser;
+        }
+    }
+
+    return 0;
+
+ err:
+    list_for_each_entry_continue_reverse( its, &host_its_list, entry )
+        writel_relaxed(its->suspend_ctx.ctlr, its->its_base + GITS_CTLR);
+
+    return ret;
+}
+
+static int gicv3_its_resume_single(struct host_its *its, unsigned int cpu)
+{
+    void __iomem *base = its->its_base;
+    unsigned int i;
+    int ret;
+    uint64_t typer;
+    unsigned int col_id = cpu; /* Xen currently uses col_id == cpu. */
+
+    /*
+     * Make sure that the ITS is disabled. If it fails to quiesce,
+     * don't restore it since writing to CBASER or BASER<n>
+     * registers is unpredictable according to the GIC v3 ITS
+     * Specification.
+     */
+    WARN_ON(readl_relaxed(base + GITS_CTLR) & GITS_CTLR_ENABLE);
+    ret = gicv3_disable_its(its);
+    if ( ret )
+        return ret;
+
+    writeq_relaxed(its->suspend_ctx.cbaser, base + GITS_CBASER);
+
+    /*
+     * Writing CBASER resets CREADR to 0, so reset CWRITER to
+     * keep the command queue pointers aligned.
+     */
+    writeq_relaxed(0, base + GITS_CWRITER);
+
+    /* Restore GITS_BASER from the value cache. */
+    for ( i = 0; i < GITS_BASER_NR_REGS; i++ )
+    {
+        uint64_t baser = its->suspend_ctx.baser[i];
+
+        if ( !(baser & GITS_VALID_BIT) )
+            continue;
+
+        writeq_relaxed(baser, base + GITS_BASER0 + i * 8);
+    }
+
+    writel_relaxed(its->suspend_ctx.ctlr, base + GITS_CTLR);
+
+    typer = readq_relaxed(base + GITS_TYPER);
+
+    /*
+     * Only collections with IDs below HCC are held in the ITS itself
+     * and lose their state across an ITS reset/power loss. Memory-backed
+     * collections are restored by restoring GITS_BASER and must not be
+     * remapped here.
+     */
+    if ( col_id < GITS_TYPER_HCC(typer) )
+        return gicv3_its_setup_collection_single(its, cpu);
+
+    return 0;
+}
+
+void gicv3_its_resume(void)
+{
+    struct host_its *its;
+    unsigned int cpu = smp_processor_id();
+    int ret;
+
+    list_for_each_entry( its, &host_its_list, entry )
+    {
+        ret = gicv3_its_resume_single(its, cpu);
+        if ( ret )
+            panic("GICv3: ITS@%"PRIpaddr": failed to restore during resume: %d\n",
+                   its->addr, ret);
+    }
+}
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
 
 /*
  * Local variables:
diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
index 038bf41142..6b025c023d 100644
--- a/xen/arch/arm/gic-v3.c
+++ b/xen/arch/arm/gic-v3.c
@@ -2186,10 +2186,14 @@ static int gicv3_suspend(void)
     if ( ret )
         goto out_enable_iface;
 
-    ret = gicv3_disable_redist();
+    ret = gicv3_its_suspend();
     if ( ret )
         goto out_enable_iface;
 
+    ret = gicv3_disable_redist();
+    if ( ret )
+        goto out_its_resume;
+
     /* Save GICR configuration */
     gicv3_redist_wait_for_rwp();
 
@@ -2229,6 +2233,9 @@ static int gicv3_suspend(void)
 
     return 0;
 
+ out_its_resume:
+    gicv3_its_resume();
+
  out_enable_iface:
     if ( gicv3_enable_redist() )
         panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
@@ -2355,6 +2362,8 @@ static void gicv3_resume(void)
 
     gicv3_redist_wait_for_rwp();
 
+    gicv3_its_resume();
+
     WRITE_SYSREG(gicv3_ctx.cpu.sre_el2, ICC_SRE_EL2);
     isb();
 
diff --git a/xen/arch/arm/include/asm/gic_v3_its.h b/xen/arch/arm/include/asm/gic_v3_its.h
index fc5a84892c..0f8cb16e41 100644
--- a/xen/arch/arm/include/asm/gic_v3_its.h
+++ b/xen/arch/arm/include/asm/gic_v3_its.h
@@ -43,6 +43,11 @@
 #define GITS_CTLR_QUIESCENT             BIT(31, UL)
 #define GITS_CTLR_ENABLE                BIT(0, UL)
 
+#define GITS_TYPER_HCC_SHIFT            24
+#define GITS_TYPER_HCC_MASK             0xffUL
+#define GITS_TYPER_HCC(r)               (((r) >> GITS_TYPER_HCC_SHIFT) & \
+                                                 GITS_TYPER_HCC_MASK)
+
 #define GITS_TYPER_PTA                  BIT(19, UL)
 #define GITS_TYPER_DEVIDS_SHIFT         13
 #define GITS_TYPER_DEVIDS_MASK          (0x1fUL << GITS_TYPER_DEVIDS_SHIFT)
@@ -129,6 +134,13 @@ struct host_its {
     spinlock_t cmd_lock;
     void *cmd_buf;
     unsigned int flags;
+#ifdef CONFIG_SYSTEM_SUSPEND
+    struct suspend_ctx {
+        uint32_t ctlr;
+        uint64_t cbaser;
+        uint64_t baser[GITS_BASER_NR_REGS];
+    } suspend_ctx;
+#endif
 };
 
 /* Map a collection for this host CPU to each host ITS. */
@@ -204,6 +216,11 @@ uint64_t gicv3_its_get_cacheability(void);
 uint64_t gicv3_its_get_shareability(void);
 unsigned int gicv3_its_get_memflags(void);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+int gicv3_its_suspend(void);
+void gicv3_its_resume(void);
+#endif
+
 #else
 
 #ifdef CONFIG_ACPI
@@ -271,6 +288,17 @@ static inline int gicv3_its_make_hwdom_dt_nodes(const struct domain *d,
     return 0;
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+static inline int gicv3_its_suspend(void)
+{
+    return 0;
+}
+
+static inline void gicv3_its_resume(void)
+{
+}
+#endif
+
 #endif /* CONFIG_HAS_ITS */
 
 #endif
diff --git a/xen/include/xen/list.h b/xen/include/xen/list.h
index 98d8482dab..2aab274157 100644
--- a/xen/include/xen/list.h
+++ b/xen/include/xen/list.h
@@ -535,6 +535,20 @@ static inline void list_splice_init(struct list_head *list,
          &(pos)->member != (head);                                        \
          (pos) = list_entry((pos)->member.next, typeof(*(pos)), member))
 
+/**
+ * list_for_each_entry_continue_reverse - iterate backwards from the given point
+ * @pos:    the type * to use as a loop cursor.
+ * @head:   the head for your list.
+ * @member: the name of the list_head within the struct.
+ *
+ * Iterate over list of given type backwards, starting from the element previous
+ * to the current one in list order.
+ */
+#define list_for_each_entry_continue_reverse(pos, head, member)           \
+    for ((pos) = list_entry((pos)->member.prev, typeof(*(pos)), member);  \
+         &(pos)->member != (head);                                        \
+         (pos) = list_entry((pos)->member.prev, typeof(*(pos)), member))
+
 /**
  * list_for_each_entry_from - iterate over list of given type from the
  *                            current point
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 06/13] xen/arm: tee: keep init_tee_secondary() for hotplug and resume
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (4 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 07/13] xen/arm: ffa: fix notification SRI across CPU hotplug/suspend Mykola Kvach
                   ` (7 subsequent siblings)
  13 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Volodymyr Babchuk, Bertrand Marquis, Jens Wiklander,
	Stefano Stabellini, Julien Grall, Michal Orzel, Luca Fancellu

init_tee_secondary() was marked __init and freed after boot. Calling it
from the CPU hotplug/resume path then executed discarded code, which
could crash Xen. Drop __init so the TEE mediator secondary init can run
safely on hotplugged and resumed CPUs.

Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
Reviewed-by: Bertrand Marquis <bertrand.marquis@arm.com>
Reviewed-by: Volodymyr Babchuk <volodymyr_babchuk@epam.com>
---
 xen/arch/arm/tee/tee.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/xen/arch/arm/tee/tee.c b/xen/arch/arm/tee/tee.c
index 8501443c8e..00e561fc78 100644
--- a/xen/arch/arm/tee/tee.c
+++ b/xen/arch/arm/tee/tee.c
@@ -128,7 +128,7 @@ static int __init tee_init(void)
 
 presmp_initcall(tee_init);
 
-void __init init_tee_secondary(void)
+void init_tee_secondary(void)
 {
     if ( cur_mediator && cur_mediator->ops->init_secondary )
         cur_mediator->ops->init_secondary();
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 07/13] xen/arm: ffa: fix notification SRI across CPU hotplug/suspend
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (5 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 06/13] xen/arm: tee: keep init_tee_secondary() for hotplug and resume Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks Mykola Kvach
                   ` (6 subsequent siblings)
  13 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Volodymyr Babchuk, Bertrand Marquis, Jens Wiklander,
	Stefano Stabellini, Julien Grall, Michal Orzel, Luca Fancellu

The FF-A notification SRI interrupt handler was not correctly tied to
CPU hotplug and suspend/resume. As a result, CPUs going offline and
back online could end up with stale or missing handlers, breaking
delivery of FF-A notifications.

Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
Reviewed-by: Bertrand Marquis <bertrand.marquis@arm.com>
---
 xen/arch/arm/tee/ffa_notif.c | 63 ++++++++++++++++++++++++++++--------
 1 file changed, 50 insertions(+), 13 deletions(-)

diff --git a/xen/arch/arm/tee/ffa_notif.c b/xen/arch/arm/tee/ffa_notif.c
index 186e726412..513c399594 100644
--- a/xen/arch/arm/tee/ffa_notif.c
+++ b/xen/arch/arm/tee/ffa_notif.c
@@ -360,10 +360,28 @@ static int32_t ffa_notification_bitmap_destroy(uint16_t vm_id)
     return ffa_simple_call(FFA_NOTIFICATION_BITMAP_DESTROY, vm_id, 0, 0, 0);
 }
 
-void ffa_notif_init_interrupt(void)
+static DEFINE_PER_CPU_READ_MOSTLY(struct irqaction, sri_irq);
+
+static int request_sri_irq(void)
 {
     int ret;
+    struct irqaction *sri_action = &this_cpu(sri_irq);
+
+    sri_action->name = "FF-A notif";
+    sri_action->handler = notif_irq_handler;
+    sri_action->dev_id = NULL;
+    sri_action->free_on_release = 0;
+
+    ret = setup_irq(notif_sri_irq, 0, sri_action);
+    if ( ret )
+        printk(XENLOG_ERR "ffa: setup_irq irq %u failed: error %d\n",
+               notif_sri_irq, ret);
 
+    return ret;
+}
+
+void ffa_notif_init_interrupt(void)
+{
     if ( fw_notif_enabled && notif_sri_irq < NR_GIC_SGI )
     {
         /*
@@ -376,14 +394,36 @@ void ffa_notif_init_interrupt(void)
          * pending, while the SPMC in the secure world will not notice that
          * the interrupt was lost.
          */
-        ret = request_irq(notif_sri_irq, 0, notif_irq_handler, "FF-A notif",
-                          NULL);
-        if ( ret )
-            printk(XENLOG_ERR "ffa: request_irq irq %u failed: error %d\n",
-                   notif_sri_irq, ret);
+        request_sri_irq();
     }
 }
 
+static void deinit_ffa_notif_interrupt(void)
+{
+    if ( fw_notif_enabled && notif_sri_irq < NR_GIC_SGI )
+        release_irq(notif_sri_irq, NULL);
+}
+
+static int cpu_ffa_notif_callback(struct notifier_block *nfb,
+                                  unsigned long action,
+                                  void *hcpu)
+{
+    switch ( action )
+    {
+    case CPU_DYING:
+        deinit_ffa_notif_interrupt();
+        break;
+    default:
+        break;
+    }
+
+    return NOTIFY_DONE;
+}
+
+static struct notifier_block cpu_ffa_notif_nfb = {
+    .notifier_call = cpu_ffa_notif_callback,
+};
+
 void ffa_notif_init(void)
 {
     const struct arm_smccc_1_2_regs arg = {
@@ -392,7 +432,6 @@ void ffa_notif_init(void)
     };
     struct arm_smccc_1_2_regs resp;
     unsigned int irq;
-    int ret;
 
     /* Only enable fw notification if all ABIs we need are supported */
     if ( ffa_fw_supports_fid(FFA_NOTIFICATION_BITMAP_CREATE) &&
@@ -408,13 +447,11 @@ void ffa_notif_init(void)
         notif_sri_irq = irq;
         if ( irq >= NR_GIC_SGI )
             irq_set_type(irq, IRQ_TYPE_EDGE_RISING);
-        ret = request_irq(irq, 0, notif_irq_handler, "FF-A notif", NULL);
-        if ( ret )
-        {
-            printk(XENLOG_ERR "ffa: request_irq irq %u failed: error %d\n",
-                   irq, ret);
+
+        if ( request_sri_irq() )
             return;
-        }
+
+        register_cpu_notifier(&cpu_ffa_notif_nfb);
         fw_notif_enabled = true;
     }
 }
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (6 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 07/13] xen/arm: ffa: fix notification SRI across CPU hotplug/suspend Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-28  8:01   ` Mykola Kvach
  2026-08-27 14:31 ` [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers Mykola Kvach
                   ` (5 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

From: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>

Store and restore active context and micro-TLB registers.

On resume, restore Root IPMMU context state before restoring Cache IPMMU
micro-TLB state. Cache IPMMUs select Root contexts through their micro-TLB
configuration, so restoring Cache micro-TLBs before the Root context
registers are restored can expose stale or uninitialized context state.

Tested on R-Car H3 Starter Kit.

Signed-off-by: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V10:
- Iterate over registered IPMMUs in reverse order during resume so Root IPMMU
  context state is restored before Cache IPMMU micro-TLB state.

Changes in V9:
- set dt_device_set_protected() only after ipmmu_alloc_ctx_suspend()
  succeeds, so DT devices do not remain protected on allocation failure.

Changes in V7:
- moved suspend context allocation before pci stuff
---
 xen/drivers/passthrough/arm/ipmmu-vmsa.c | 323 +++++++++++++++++++++--
 1 file changed, 308 insertions(+), 15 deletions(-)

diff --git a/xen/drivers/passthrough/arm/ipmmu-vmsa.c b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
index fa9ab9cb13..2e54fa63d6 100644
--- a/xen/drivers/passthrough/arm/ipmmu-vmsa.c
+++ b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
@@ -71,6 +71,8 @@
 })
 #endif
 
+#define dev_dbg(dev, fmt, ...)    \
+    dev_print(dev, XENLOG_DEBUG, fmt, ## __VA_ARGS__)
 #define dev_info(dev, fmt, ...)    \
     dev_print(dev, XENLOG_INFO, fmt, ## __VA_ARGS__)
 #define dev_warn(dev, fmt, ...)    \
@@ -130,6 +132,24 @@ struct ipmmu_features {
     unsigned int imuctr_ttsel_mask;
 };
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+struct ipmmu_reg_ctx {
+    unsigned int imttlbr0;
+    unsigned int imttubr0;
+    unsigned int imttbcr;
+    unsigned int imctr;
+};
+
+struct ipmmu_vmsa_backup {
+    struct device *dev;
+    unsigned int *utlbs_val;
+    unsigned int *asids_val;
+    struct list_head list;
+};
+
+#endif
+
 /* Root/Cache IPMMU device's information */
 struct ipmmu_vmsa_device {
     struct device *dev;
@@ -142,6 +162,9 @@ struct ipmmu_vmsa_device {
     struct ipmmu_vmsa_domain *domains[IPMMU_CTX_MAX];
     unsigned int utlb_refcount[IPMMU_UTLB_MAX];
     const struct ipmmu_features *features;
+#ifdef CONFIG_SYSTEM_SUSPEND
+    struct ipmmu_reg_ctx *reg_backup[IPMMU_CTX_MAX];
+#endif
 };
 
 /*
@@ -547,6 +570,249 @@ static void ipmmu_domain_free_context(struct ipmmu_vmsa_device *mmu,
     spin_unlock_irqrestore(&mmu->lock, flags);
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+static DEFINE_SPINLOCK(ipmmu_devices_backup_lock);
+static LIST_HEAD(ipmmu_devices_backup);
+
+static struct ipmmu_reg_ctx root_pgtable[IPMMU_CTX_MAX];
+
+static uint32_t ipmmu_imuasid_read(struct ipmmu_vmsa_device *mmu,
+                                   unsigned int utlb)
+{
+    return ipmmu_read(mmu, ipmmu_utlb_reg(mmu, IMUASID(utlb)));
+}
+
+static void ipmmu_utlbs_backup(struct ipmmu_vmsa_device *mmu)
+{
+    struct ipmmu_vmsa_backup *backup_data;
+
+    dev_dbg(mmu->dev, "Handle micro-TLBs backup\n");
+
+    spin_lock(&ipmmu_devices_backup_lock);
+
+    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
+    {
+        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
+        unsigned int i;
+
+        if ( to_ipmmu(backup_data->dev) != mmu )
+            continue;
+
+        for ( i = 0; i < fwspec->num_ids; i++ )
+        {
+            unsigned int utlb = fwspec->ids[i];
+
+            backup_data->asids_val[i] = ipmmu_imuasid_read(mmu, utlb);
+            backup_data->utlbs_val[i] = ipmmu_imuctr_read(mmu, utlb);
+        }
+    }
+
+    spin_unlock(&ipmmu_devices_backup_lock);
+}
+
+static void ipmmu_utlbs_restore(struct ipmmu_vmsa_device *mmu)
+{
+    struct ipmmu_vmsa_backup *backup_data;
+
+    dev_dbg(mmu->dev, "Handle micro-TLBs restore\n");
+
+    spin_lock(&ipmmu_devices_backup_lock);
+
+    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
+    {
+        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
+        unsigned int i;
+
+        if ( to_ipmmu(backup_data->dev) != mmu )
+            continue;
+
+        for ( i = 0; i < fwspec->num_ids; i++ )
+        {
+            unsigned int utlb = fwspec->ids[i];
+
+            ipmmu_imuasid_write(mmu, utlb, backup_data->asids_val[i]);
+            ipmmu_imuctr_write(mmu, utlb, backup_data->utlbs_val[i]);
+        }
+    }
+
+    spin_unlock(&ipmmu_devices_backup_lock);
+}
+
+static void ipmmu_domain_backup_context(struct ipmmu_vmsa_domain *domain)
+{
+    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
+    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
+
+    dev_dbg(mmu->dev, "Handle domain context %u backup\n", domain->context_id);
+
+    regs->imttlbr0 = ipmmu_ctx_read_root(domain, IMTTLBR0);
+    regs->imttubr0 = ipmmu_ctx_read_root(domain, IMTTUBR0);
+    regs->imttbcr  = ipmmu_ctx_read_root(domain, IMTTBCR);
+    regs->imctr    = ipmmu_ctx_read_root(domain, IMCTR);
+}
+
+static void ipmmu_domain_restore_context(struct ipmmu_vmsa_domain *domain)
+{
+    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
+    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
+
+    dev_dbg(mmu->dev, "Handle domain context %u restore\n", domain->context_id);
+
+    ipmmu_ctx_write_root(domain, IMTTLBR0, regs->imttlbr0);
+    ipmmu_ctx_write_root(domain, IMTTUBR0, regs->imttubr0);
+    ipmmu_ctx_write_root(domain, IMTTBCR,  regs->imttbcr);
+    ipmmu_ctx_write_all(domain,  IMCTR,    regs->imctr | IMCTR_FLUSH);
+}
+
+/*
+ * Xen: Unlike Linux implementation, Xen uses a single driver instance
+ * for handling all IPMMUs. There is no framework for ipmmu_suspend/resume
+ * callbacks to be invoked for each IPMMU device. So, we need to iterate
+ * through all registered IPMMUs performing required actions.
+ *
+ * Also take care of restoring special settings, such as translation
+ * table format, etc.
+ */
+static int __must_check ipmmu_suspend(void)
+{
+    struct ipmmu_vmsa_device *mmu;
+
+    if ( !iommu_enabled )
+        return 0;
+
+    printk(XENLOG_DEBUG "ipmmu: Suspending...\n");
+
+    spin_lock(&ipmmu_devices_lock);
+
+    list_for_each_entry( mmu, &ipmmu_devices, list )
+    {
+        if ( ipmmu_is_root(mmu) )
+        {
+            unsigned int i;
+
+            for ( i = 0; i < mmu->num_ctx; i++ )
+            {
+                if ( !mmu->domains[i] )
+                    continue;
+                ipmmu_domain_backup_context(mmu->domains[i]);
+            }
+        }
+        else
+            ipmmu_utlbs_backup(mmu);
+    }
+
+    spin_unlock(&ipmmu_devices_lock);
+
+    return 0;
+}
+
+static void ipmmu_resume(void)
+{
+    struct ipmmu_vmsa_device *mmu;
+
+    if ( !iommu_enabled )
+        return;
+
+    printk(XENLOG_DEBUG "ipmmu: Resuming...\n");
+
+    spin_lock(&ipmmu_devices_lock);
+
+    /*
+     * IPMMUs are registered with list_add(), with Root IPMMU probed first.
+     * Walk backwards to restore Root contexts before Cache micro-TLBs.
+     */
+    list_for_each_entry_reverse( mmu, &ipmmu_devices, list )
+    {
+        uint32_t reg;
+
+        /* Do not use security group function */
+        reg = IMSCTLR + mmu->features->control_offset_base;
+        ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) & ~IMSCTLR_USE_SECGRP);
+
+        if ( ipmmu_is_root(mmu) )
+        {
+            unsigned int i;
+
+            /* Use stage 2 translation table format */
+            reg = IMSAUXCTLR + mmu->features->control_offset_base;
+            ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) | IMSAUXCTLR_S2PTE);
+
+            for ( i = 0; i < mmu->num_ctx; i++ )
+            {
+                if ( !mmu->domains[i] )
+                    continue;
+                ipmmu_domain_restore_context(mmu->domains[i]);
+            }
+        }
+        else
+            ipmmu_utlbs_restore(mmu);
+    }
+
+    spin_unlock(&ipmmu_devices_lock);
+}
+
+static int ipmmu_alloc_ctx_suspend(struct device *dev)
+{
+    struct ipmmu_vmsa_backup *backup_data;
+    unsigned int *utlbs_val, *asids_val;
+    struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(dev);
+
+    utlbs_val = xzalloc_array(unsigned int, fwspec->num_ids);
+    if ( !utlbs_val )
+        return -ENOMEM;
+
+    asids_val = xzalloc_array(unsigned int, fwspec->num_ids);
+    if ( !asids_val )
+    {
+        xfree(utlbs_val);
+        return -ENOMEM;
+    }
+
+    backup_data = xzalloc(struct ipmmu_vmsa_backup);
+    if ( !backup_data )
+    {
+        xfree(utlbs_val);
+        xfree(asids_val);
+        return -ENOMEM;
+    }
+
+    backup_data->dev = dev;
+    backup_data->utlbs_val = utlbs_val;
+    backup_data->asids_val = asids_val;
+
+    spin_lock(&ipmmu_devices_backup_lock);
+    list_add(&backup_data->list, &ipmmu_devices_backup);
+    spin_unlock(&ipmmu_devices_backup_lock);
+
+    return 0;
+}
+
+#ifdef CONFIG_HAS_PCI
+static void ipmmu_free_ctx_suspend(struct device *dev)
+{
+    struct ipmmu_vmsa_backup *backup_data, *tmp;
+
+    spin_lock(&ipmmu_devices_backup_lock);
+
+    list_for_each_entry_safe( backup_data, tmp, &ipmmu_devices_backup, list )
+    {
+        if ( backup_data->dev == dev )
+        {
+            list_del(&backup_data->list);
+            xfree(backup_data->utlbs_val);
+            xfree(backup_data->asids_val);
+            xfree(backup_data);
+            break;
+        }
+    }
+
+    spin_unlock(&ipmmu_devices_backup_lock);
+}
+#endif /* CONFIG_HAS_PCI */
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
 {
     uint64_t ttbr;
@@ -559,6 +825,9 @@ static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
         return ret;
 
     domain->context_id = ret;
+#ifdef CONFIG_SYSTEM_SUSPEND
+    domain->mmu->root->reg_backup[ret] = &root_pgtable[ret];
+#endif
 
     /*
      * TTBR0
@@ -615,6 +884,9 @@ static void ipmmu_domain_destroy_context(struct ipmmu_vmsa_domain *domain)
     ipmmu_ctx_write_root(domain, IMCTR, IMCTR_FLUSH);
     ipmmu_tlb_sync(domain);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+    domain->mmu->root->reg_backup[domain->context_id] = NULL;
+#endif
     ipmmu_domain_free_context(domain->mmu->root, domain->context_id);
 }
 
@@ -1338,10 +1610,11 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
     struct iommu_fwspec *fwspec;
 
 #ifdef CONFIG_HAS_PCI
+    int ret;
+
     if ( dev_is_pci(dev) )
     {
         struct pci_dev *pdev = dev_to_pci(dev);
-        int ret;
 
         if ( devfn != pdev->devfn )
             return 0;
@@ -1358,17 +1631,24 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
     if ( !to_ipmmu(dev) )
         return -ENODEV;
 
-    if ( !dev_is_pci(dev) )
+    if ( !dev_is_pci(dev) && dt_device_is_protected(dev_to_dt(dev)) )
     {
-        if ( dt_device_is_protected(dev_to_dt(dev)) )
-        {
-            dev_err(dev, "Already added to IPMMU\n");
-            return -EEXIST;
-        }
+        dev_err(dev, "Already added to IPMMU\n");
+        return -EEXIST;
+    }
 
-        /* Let Xen know that the master device is protected by an IOMMU. */
-        dt_device_set_protected(dev_to_dt(dev));
+#ifdef CONFIG_SYSTEM_SUSPEND
+    if ( ipmmu_alloc_ctx_suspend(dev) )
+    {
+        dev_err(dev, "Failed to allocate context for suspend\n");
+        return -ENOMEM;
     }
+#endif
+
+    /* Let Xen know that the master device is protected by an IOMMU. */
+    if ( !dev_is_pci(dev) )
+        dt_device_set_protected(dev_to_dt(dev));
+
 #ifdef CONFIG_HAS_PCI
     if ( dev_is_pci(dev) )
     {
@@ -1377,26 +1657,28 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
         struct pci_host_bridge *bridge;
         struct iommu_fwspec *fwspec_bridge;
         unsigned int utlb_osid0 = 0;
-        int ret;
 
         bridge = pci_find_host_bridge(pdev->seg, pdev->bus);
         if ( !bridge )
         {
             dev_err(dev, "Failed to find host bridge\n");
-            return -ENODEV;
+            ret = -ENODEV;
+            goto free_suspend_ctx;
         }
 
         fwspec_bridge = dev_iommu_fwspec_get(dt_to_dev(bridge->dt_node));
         if ( fwspec_bridge->num_ids < 1 )
         {
             dev_err(dev, "Failed to find host bridge uTLB\n");
-            return -ENXIO;
+            ret = -ENXIO;
+            goto free_suspend_ctx;
         }
 
         if ( fwspec->num_ids < 1 )
         {
             dev_err(dev, "Failed to find uTLB");
-            return -ENXIO;
+            ret = -ENXIO;
+            goto free_suspend_ctx;
         }
 
         rcar4_pcie_osid_regs_init(bridge);
@@ -1405,7 +1687,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
         if ( ret < 0 )
         {
             dev_err(dev, "No unused OSID regs\n");
-            return ret;
+            goto free_suspend_ctx;
         }
         reg_id = ret;
 
@@ -1420,7 +1702,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
         {
             rcar4_pcie_osid_bdf_clear(bridge, reg_id);
             rcar4_pcie_osid_reg_free(bridge, reg_id);
-            return ret;
+            goto free_suspend_ctx;
         }
     }
 #endif
@@ -1429,6 +1711,13 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
              dev_name(fwspec->iommu_dev), fwspec->num_ids);
 
     return 0;
+#ifdef CONFIG_HAS_PCI
+ free_suspend_ctx:
+#ifdef CONFIG_SYSTEM_SUSPEND
+    ipmmu_free_ctx_suspend(dev);
+#endif
+    return ret;
+#endif
 }
 
 static int ipmmu_iommu_domain_init(struct domain *d)
@@ -1490,6 +1779,10 @@ static const struct iommu_ops ipmmu_iommu_ops =
     .unmap_page      = arm_iommu_unmap_page,
     .dt_xlate        = ipmmu_dt_xlate,
     .add_device      = ipmmu_add_device,
+#ifdef CONFIG_SYSTEM_SUSPEND
+    .suspend         = ipmmu_suspend,
+    .resume          = ipmmu_resume,
+#endif
 };
 
 static __init int ipmmu_init(struct dt_device_node *node, const void *data)
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (7 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-28 16:17   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND Mykola Kvach
                   ` (4 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Bertrand Marquis, Rahul Singh, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk,
	Pranjal Shrivastava, Luca Fancellu

Add system suspend/resume callbacks for the Arm SMMUv3 driver.

During suspend, configure GBPA to abort incoming transactions, disable the
translation interface while keeping CMDQ enabled, issue CMD_SYNC to ensure
all previously issued commands have completed, then disable the SMMU IRQs
and SMMU.

Resume uses arm_smmu_device_reset() to reprogram the SMMU and re-enable
translation and interrupt generation.

The IRQ setup split follows the approach from Pranjal Shrivastava's Linux
arm-smmu-v3 runtime/system sleep series: IRQ handlers are requested once
during probe, while reset/resume only restores SMMU hardware state and
re-enables IRQ_CTRL.

Only the pieces relevant to Xen's currently supported SMMUv3 path are
ported here. Xen documents SMMUv3 MSI and PCI ATS as unsupported and not
compiled/tested, so this patch does not restore SMMU MSI IRQ_CFGn registers
nor reinitialize ATS/PRI endpoints. If those paths become usable,
suspend/resume will need corresponding MSI restore and ATS/PRI
quiesce/reinit steps.

Link: https://lore.kernel.org/r/20260414194702.1229094-1-praan@google.com/
Based-on-patch-by: Pranjal Shrivastava <praan@google.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in V11:
- Keep arm_smmu_update_gbpa() and arm_smmu_device_reset() in init text when
  CONFIG_SYSTEM_SUSPEND is disabled.

Changes in V10:
- Disable SMMU interrupt generation during suspend before disabling the
  SMMU interface, matching the resume/reset path which re-enables IRQ_CTRL.

Changes in V9:
- Use CMD_SYNC in suspend instead of polling CMDQ_CONS, so the suspend
  path waits for command completion rather than only command consumption.
- Document that arm_smmu_setup_irqs() is probe-only and that future Xen
  SMMUv3 MSI support will need to restore SMMU IRQ_CFGn registers on
  resume.
- Restore the reference to Pranjal's Linux runtime/system sleep series and
  clarify that MSI/ATS/PRI resume handling is outside the supported Xen
  path.
- Prefix the subject with xen/arm for consistency with the rest of the
  Arm suspend/resume series.

Changes in V8:
- Honor ARM_SMMU_FEAT_SEV when draining the CMDQ during suspend, matching
  the existing runtime CMD_SYNC path.
- Fold the suspend rollback reset path into a helper and rename the error
  reporting to describe suspend rollback rather than resume.
- Treat SMMU reset failure during resume as fatal instead of logging and
  continuing with a potentially unusable IOMMU.
- cosmetic changes
---
 xen/drivers/passthrough/arm/smmu-v3.c | 194 +++++++++++++++++++++-----
 1 file changed, 158 insertions(+), 36 deletions(-)

diff --git a/xen/drivers/passthrough/arm/smmu-v3.c b/xen/drivers/passthrough/arm/smmu-v3.c
index bf153227db..7f1d00fb81 100644
--- a/xen/drivers/passthrough/arm/smmu-v3.c
+++ b/xen/drivers/passthrough/arm/smmu-v3.c
@@ -94,6 +94,12 @@
 
 #include "smmu-v3.h"
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+#define __init_or_smmu_suspend
+#else
+#define __init_or_smmu_suspend __init
+#endif
+
 #define ARM_SMMU_VTCR_SH_IS		3
 #define ARM_SMMU_VTCR_RGN_WBWA		1
 #define ARM_SMMU_VTCR_TG0_4K		0
@@ -1814,8 +1820,8 @@ static int arm_smmu_write_reg_sync(struct arm_smmu_device *smmu, u32 val,
 }
 
 /* GBPA is "special" */
-static int __init arm_smmu_update_gbpa(struct arm_smmu_device *smmu,
-                                       u32 set, u32 clr)
+static int __init_or_smmu_suspend
+arm_smmu_update_gbpa(struct arm_smmu_device *smmu, u32 set, u32 clr)
 {
 	int ret;
 	u32 reg, __iomem *gbpa = smmu->base + ARM_SMMU_GBPA;
@@ -1995,10 +2001,35 @@ err_free_evtq_irq:
 	return ret;
 }
 
+static int arm_smmu_enable_irqs(struct arm_smmu_device *smmu)
+{
+	int ret;
+	u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
+
+	if ( smmu->features & ARM_SMMU_FEAT_PRI )
+		irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
+
+	/* Enable interrupt generation on the SMMU */
+	ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
+				      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
+	if ( ret )
+	{
+		dev_warn(smmu->dev, "failed to enable irqs\n");
+		return ret;
+	}
+
+	return 0;
+}
+
+/*
+ * Probe-time only: request host IRQs and, when available, program the SMMU's
+ * MSI doorbells. Resume does not restore the SMMU *_IRQ_CFGn MSI registers,
+ * so any host suspend support must treat the active MSI IRQ path as
+ * unsupported until that restore path exists.
+ */
 static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
 {
 	int ret, irq;
-	u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
 
 	/* Disable IRQs first */
 	ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
@@ -2028,22 +2059,7 @@ static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
 		}
 	}
 
-	if (smmu->features & ARM_SMMU_FEAT_PRI)
-		irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
-
-	/* Enable interrupt generation on the SMMU */
-	ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
-				      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
-	if (ret) {
-		dev_warn(smmu->dev, "failed to enable irqs\n");
-		goto err_free_irqs;
-	}
-
 	return 0;
-
-err_free_irqs:
-	arm_smmu_free_irqs(smmu);
-	return ret;
 }
 
 static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
@@ -2057,7 +2073,8 @@ static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
 	return ret;
 }
 
-static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
+static int __init_or_smmu_suspend
+arm_smmu_device_reset(struct arm_smmu_device *smmu)
 {
 	int ret;
 	u32 reg, enables;
@@ -2163,17 +2180,9 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
 		}
 	}
 
-	ret = arm_smmu_setup_irqs(smmu);
-	if (ret) {
-		dev_err(smmu->dev, "failed to setup irqs\n");
+	ret = arm_smmu_enable_irqs(smmu);
+	if ( ret )
 		return ret;
-	}
-
-	/* Initialize tasklets for threaded IRQs*/
-	tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
-	tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
-	tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
-				 smmu);
 
 	/* Enable the SMMU interface, or ensure bypass */
 	if (disable_bypass) {
@@ -2181,20 +2190,16 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
 	} else {
 		ret = arm_smmu_update_gbpa(smmu, 0, GBPA_ABORT);
 		if (ret)
-			goto err_free_irqs;
+			return ret;
 	}
 	ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
 				      ARM_SMMU_CR0ACK);
 	if (ret) {
 		dev_err(smmu->dev, "failed to enable SMMU interface\n");
-		goto err_free_irqs;
+		return ret;
 	}
 
 	return 0;
-
-err_free_irqs:
-	arm_smmu_free_irqs(smmu);
-	return ret;
 }
 
 static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
@@ -2558,10 +2563,23 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
 	if (ret)
 		goto out_free;
 
+	ret = arm_smmu_setup_irqs(smmu);
+	if ( ret )
+	{
+		dev_err(smmu->dev, "failed to setup irqs\n");
+		goto out_free;
+	}
+
+	/* Initialize tasklets for threaded IRQs*/
+	tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
+	tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
+	tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
+				smmu);
+
 	/* Reset the device */
 	ret = arm_smmu_device_reset(smmu);
 	if (ret)
-		goto out_free;
+		goto out_free_irqs;
 
 	/*
 	 * Keep a list of all probed devices. This will be used to query
@@ -2575,6 +2593,8 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
 
 	return 0;
 
+out_free_irqs:
+	arm_smmu_free_irqs(smmu);
 
 out_free:
 	arm_smmu_free_structures(smmu);
@@ -2855,6 +2875,104 @@ static void arm_smmu_iommu_xen_domain_teardown(struct domain *d)
 	xfree(xen_domain);
 }
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+
+static void arm_smmu_reset_for_suspend_rollback(struct arm_smmu_device *smmu)
+{
+	int ret = arm_smmu_device_reset(smmu);
+
+	if ( ret )
+		dev_err(smmu->dev, "Failed to reset during suspend rollback: %d\n",
+				ret);
+}
+
+static int arm_smmu_suspend(void)
+{
+	struct arm_smmu_device *smmu;
+	int ret = 0;
+
+	list_for_each_entry(smmu, &arm_smmu_devices, devices)
+	{
+		/* Abort all transactions before disable to avoid spurious bypass */
+		ret = arm_smmu_update_gbpa(smmu, GBPA_ABORT, 0);
+		if ( ret )
+			goto fail;
+
+		ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
+					ARM_SMMU_IRQ_CTRLACK);
+		if ( ret )
+		{
+			dev_err(smmu->dev, "Timed-out while disabling SMMU irqs\n");
+			goto fail;
+		}
+
+		/* Disable the SMMU via CR0.EN and all queues except CMDQ */
+		ret = arm_smmu_write_reg_sync(smmu, CR0_CMDQEN, ARM_SMMU_CR0,
+					ARM_SMMU_CR0ACK);
+		if ( ret )
+		{
+			dev_err(smmu->dev, "Timed-out while disabling smmu\n");
+			goto fail;
+		}
+
+		/*
+		 * At this point the translation interface is disabled and the
+		 * SMMU won't access translation/config structures, even
+		 * speculatively, as per the IHI0070 spec (section 6.3.9.6).
+		 * CMDQ is still enabled so that a CMD_SYNC can complete any
+		 * previously issued commands.
+		 */
+
+		/* Ensure all previously issued commands have completed. */
+		ret = arm_smmu_cmdq_issue_sync(smmu);
+		if ( ret )
+		{
+			dev_err(smmu->dev, "Timed-out waiting for pending commands\n");
+			goto fail;
+		}
+
+		/* Disable everything */
+		ret = arm_smmu_device_disable(smmu);
+		if ( ret )
+			goto fail;
+
+		dev_dbg(smmu->dev, "Suspended smmu\n");
+	}
+
+	return 0;
+
+ fail:
+	/* Reset the device that failed as well as any already-suspended ones. */
+	arm_smmu_reset_for_suspend_rollback(smmu);
+
+	list_for_each_entry_continue_reverse(smmu, &arm_smmu_devices, devices)
+		arm_smmu_reset_for_suspend_rollback(smmu);
+
+	return ret;
+}
+
+static void arm_smmu_resume(void)
+{
+	int ret;
+	struct arm_smmu_device *smmu;
+
+	list_for_each_entry(smmu, &arm_smmu_devices, devices)
+	{
+		dev_dbg(smmu->dev, "Resuming device\n");
+
+		/*
+		 * The reset will re-initialize all the base addresses, queues,
+		 * prod and cons maintained within struct arm_smmu_device as well as
+		 * re-enable the interrupts.
+		 */
+		ret = arm_smmu_device_reset(smmu);
+		if ( ret )
+			panic("SMMUv3: %s: Failed to reset during resume: %d\n",
+			      dev_name(smmu->dev), ret);
+	}
+}
+#endif
+
 static const struct iommu_ops arm_smmu_iommu_ops = {
 	.page_sizes		= PAGE_SIZE_4K,
 	.init			= arm_smmu_iommu_xen_domain_init,
@@ -2867,6 +2985,10 @@ static const struct iommu_ops arm_smmu_iommu_ops = {
 	.unmap_page		= arm_iommu_unmap_page,
 	.dt_xlate		= arm_smmu_dt_xlate,
 	.add_device		= arm_smmu_add_device,
+#ifdef CONFIG_SYSTEM_SUSPEND
+	.suspend		= arm_smmu_suspend,
+	.resume			= arm_smmu_resume,
+#endif
 };
 
 static __init int arm_smmu_dt_init(struct dt_device_node *dev,
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (8 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-28 16:17   ` Bertrand Marquis
  2026-08-27 14:31 ` [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface) Mykola Kvach
                   ` (3 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Oleksandr Tyshchenko,
	Luca Fancellu

From: Mirela Simonovic <mirela.simonovic@aggios.com>

On wakeup from PSCI SYSTEM_SUSPEND, Xen re-enters EL2 with the MMU and
data cache disabled. The resume path must first switch back to Xen's
runtime page tables before it can access the saved CPU context using
virtual addresses.

Add an arm64 hyp_resume trampoline that reuses enable_secondary_cpu_mm()
to enable the data cache and MMU, switch to init_ttbr, and resume in the
runtime virtual mapping. The trampoline then restores the saved CPU
general-purpose and system-control register context.

prepare_resume_ctx() must be invoked just before the PSCI system suspend
call is issued to the platform firmware. It saves the current CPU context
and returns a non-zero value so that the caller enters the physical
SYSTEM_SUSPEND call.

On resume, hyp_resume restores the saved context, including the saved link
register. Control therefore returns to the place where prepare_resume_ctx()
was called. To avoid re-entering the suspend path, the restored path sees
prepare_resume_ctx() return zero.

The assembly save/restore code uses offsets generated by asm-offsets.c
from struct resume_cpu_context, keeping the assembly memory accesses in
sync with the C structure layout.

Support for ARM32 is not implemented. Instead, compilation fails with a
build-time error if suspend is enabled for ARM32.

Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in v10:
- Save and restore CNTHCTL_EL2 across SYSTEM_SUSPEND

Changes in v9:
- Drop the misleading prepare_resume_ctx() pointer argument and make both
  save/restore paths use the global resume_cpu_context.
- Squash the arm64 resume trampoline into the context save/restore patch.
- Document in code that hyp_resume relies on PSCI initial-state rules.
- Use generic platform firmware wording instead of ATF-specific wording.
- Rename the saved context type/storage to resume_cpu_context and rely on
  implicit zero-initialization for the file-scope object.
- Use asm-offsets.c-generated RESUME_CTX_* offsets to keep the assembly
  save/restore code in sync with struct resume_cpu_context.

Changes in v8:
- Fix alignments in code.

Changes in v7:
- No functional changes, just moved commit.
---
 xen/arch/arm/Makefile              |   1 +
 xen/arch/arm/arm64/asm-offsets.c   |  21 +++++
 xen/arch/arm/arm64/head.S          | 122 +++++++++++++++++++++++++++++
 xen/arch/arm/include/asm/suspend.h |  27 +++++++
 xen/arch/arm/suspend.c             |  14 ++++
 5 files changed, 185 insertions(+)
 create mode 100644 xen/arch/arm/suspend.c

diff --git a/xen/arch/arm/Makefile b/xen/arch/arm/Makefile
index b7afd3e58c..788db83ba9 100644
--- a/xen/arch/arm/Makefile
+++ b/xen/arch/arm/Makefile
@@ -51,6 +51,7 @@ obj-y += setup.o
 obj-y += shutdown.o
 obj-y += smp.o
 obj-y += smpboot.o
+obj-$(CONFIG_SYSTEM_SUSPEND) += suspend.o
 obj-$(CONFIG_SYSCTL) += sysctl.o
 obj-y += time.o
 obj-y += traps.o
diff --git a/xen/arch/arm/arm64/asm-offsets.c b/xen/arch/arm/arm64/asm-offsets.c
index 38a3894a3b..5d60406e9c 100644
--- a/xen/arch/arm/arm64/asm-offsets.c
+++ b/xen/arch/arm/arm64/asm-offsets.c
@@ -13,6 +13,7 @@
 #include <asm/mm.h>
 #include <asm/setup.h>
 #include <asm/smccc.h>
+#include <asm/suspend.h>
 
 #define DEFINE(_sym, _val)                                                 \
     asm volatile ( "\n.ascii\"==>#define " #_sym " %0 /* " #_val " */<==\""\
@@ -57,6 +58,26 @@ void __dummy__(void)
    OFFSET(INITINFO_stack, struct init_info, stack);
    BLANK();
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+   OFFSET(RESUME_CTX_X19, struct resume_cpu_context, callee_regs[0]);
+   OFFSET(RESUME_CTX_X21, struct resume_cpu_context, callee_regs[2]);
+   OFFSET(RESUME_CTX_X23, struct resume_cpu_context, callee_regs[4]);
+   OFFSET(RESUME_CTX_X25, struct resume_cpu_context, callee_regs[6]);
+   OFFSET(RESUME_CTX_X27, struct resume_cpu_context, callee_regs[8]);
+   OFFSET(RESUME_CTX_X29, struct resume_cpu_context, callee_regs[10]);
+   OFFSET(RESUME_CTX_SP, struct resume_cpu_context, sp);
+   OFFSET(RESUME_CTX_VBAR_EL2, struct resume_cpu_context, vbar_el2);
+   OFFSET(RESUME_CTX_VTCR_EL2, struct resume_cpu_context, vtcr_el2);
+   OFFSET(RESUME_CTX_VTTBR_EL2, struct resume_cpu_context, vttbr_el2);
+   OFFSET(RESUME_CTX_TPIDR_EL2, struct resume_cpu_context, tpidr_el2);
+   OFFSET(RESUME_CTX_MDCR_EL2, struct resume_cpu_context, mdcr_el2);
+   OFFSET(RESUME_CTX_HSTR_EL2, struct resume_cpu_context, hstr_el2);
+   OFFSET(RESUME_CTX_CPTR_EL2, struct resume_cpu_context, cptr_el2);
+   OFFSET(RESUME_CTX_HCR_EL2, struct resume_cpu_context, hcr_el2);
+   OFFSET(RESUME_CTX_CNTHCTL_EL2, struct resume_cpu_context, cnthctl_el2);
+   BLANK();
+#endif
+
    OFFSET(SMCCC_RES_a0, struct arm_smccc_res, a0);
    OFFSET(SMCCC_RES_a2, struct arm_smccc_res, a2);
    OFFSET(ARM_SMCCC_1_2_REGS_X0_OFFS, struct arm_smccc_1_2_regs, a0);
diff --git a/xen/arch/arm/arm64/head.S b/xen/arch/arm/arm64/head.S
index 72c7b24498..962be716ae 100644
--- a/xen/arch/arm/arm64/head.S
+++ b/xen/arch/arm/arm64/head.S
@@ -561,6 +561,128 @@ END(efi_xen_start)
 
 #endif /* CONFIG_ARM_EFI */
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+/*
+ * int prepare_resume_ctx(void)
+ *
+ * CPU context saved here will be restored on resume in hyp_resume function.
+ * prepare_resume_ctx shall return a non-zero value. Upon restoring context
+ * hyp_resume shall return value zero instead. From C code that invokes
+ * prepare_resume_ctx, the return value is interpreted to determine whether
+ * the context is saved (prepare_resume_ctx) or restored (hyp_resume).
+ */
+FUNC(prepare_resume_ctx)
+        ldr   x0, =resume_cpu_context
+
+        /* Store callee-saved registers */
+        stp   x19, x20, [x0, #RESUME_CTX_X19]
+        stp   x21, x22, [x0, #RESUME_CTX_X21]
+        stp   x23, x24, [x0, #RESUME_CTX_X23]
+        stp   x25, x26, [x0, #RESUME_CTX_X25]
+        stp   x27, x28, [x0, #RESUME_CTX_X27]
+        stp   x29, lr, [x0, #RESUME_CTX_X29]
+
+        /* Store stack-pointer */
+        mov   x2, sp
+        str   x2, [x0, #RESUME_CTX_SP]
+
+        /* Store system control registers */
+        mrs   x2, VBAR_EL2
+        str   x2, [x0, #RESUME_CTX_VBAR_EL2]
+        mrs   x2, VTCR_EL2
+        str   x2, [x0, #RESUME_CTX_VTCR_EL2]
+        mrs   x2, VTTBR_EL2
+        str   x2, [x0, #RESUME_CTX_VTTBR_EL2]
+        mrs   x2, TPIDR_EL2
+        str   x2, [x0, #RESUME_CTX_TPIDR_EL2]
+        mrs   x2, MDCR_EL2
+        str   x2, [x0, #RESUME_CTX_MDCR_EL2]
+        mrs   x2, HSTR_EL2
+        str   x2, [x0, #RESUME_CTX_HSTR_EL2]
+        mrs   x2, CPTR_EL2
+        str   x2, [x0, #RESUME_CTX_CPTR_EL2]
+        mrs   x2, HCR_EL2
+        str   x2, [x0, #RESUME_CTX_HCR_EL2]
+        mrs   x2, CNTHCTL_EL2
+        str   x2, [x0, #RESUME_CTX_CNTHCTL_EL2]
+
+        /* prepare_resume_ctx must return a non-zero value */
+        mov   x0, #1
+        ret
+END(prepare_resume_ctx)
+
+FUNC(hyp_resume)
+        /*
+         * PSCI states that SYSTEM_SUSPEND follows the CPU_SUSPEND initial
+         * state rules, so PSCI-compliant firmware must enter the return
+         * exception level with DAIF masked.
+         */
+
+        /* Initialize the UART if earlyprintk has been enabled. */
+#ifdef CONFIG_EARLY_PRINTK
+        bl    init_uart
+#endif
+        PRINT_ID("- Xen resuming -\r\n")
+
+        bl    check_cpu_mode
+        bl    cpu_init
+
+        ldr   x0, =start
+        adr   x20, start             /* x20 := paddr (start) */
+        sub   x20, x20, x0           /* x20 := phys-offset */
+        ldr   lr, =mmu_resumed
+        b     enable_secondary_cpu_mm
+
+mmu_resumed:
+        /* Now we can access the saved context, so restore it here. */
+        ldr   x0, =resume_cpu_context
+
+        /* Restore callee-saved registers */
+        ldp   x19, x20, [x0, #RESUME_CTX_X19]
+        ldp   x21, x22, [x0, #RESUME_CTX_X21]
+        ldp   x23, x24, [x0, #RESUME_CTX_X23]
+        ldp   x25, x26, [x0, #RESUME_CTX_X25]
+        ldp   x27, x28, [x0, #RESUME_CTX_X27]
+        ldp   x29, lr, [x0, #RESUME_CTX_X29]
+
+        /* Restore stack pointer */
+        ldr   x2, [x0, #RESUME_CTX_SP]
+        mov   sp, x2
+
+        /* Restore system control registers */
+        ldr   x2, [x0, #RESUME_CTX_VBAR_EL2]
+        msr   VBAR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_VTCR_EL2]
+        msr   VTCR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_VTTBR_EL2]
+        msr   VTTBR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_TPIDR_EL2]
+        msr   TPIDR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_MDCR_EL2]
+        msr   MDCR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_HSTR_EL2]
+        msr   HSTR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_CPTR_EL2]
+        msr   CPTR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_HCR_EL2]
+        msr   HCR_EL2, x2
+        ldr   x2, [x0, #RESUME_CTX_CNTHCTL_EL2]
+        msr   CNTHCTL_EL2, x2
+        isb
+
+        /*
+         * Since context is restored return from this function will appear
+         * as return from prepare_resume_ctx. To distinguish a return from
+         * prepare_resume_ctx which is called upon finalizing the suspend,
+         * as opposed to return from this function which executes on resume,
+         * we need to return zero value here.
+         */
+        mov   x0, #0
+        ret
+END(hyp_resume)
+
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 /*
  * Local variables:
  * mode: ASM
diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
index 31a98a1f1b..c848fc6340 100644
--- a/xen/arch/arm/include/asm/suspend.h
+++ b/xen/arch/arm/include/asm/suspend.h
@@ -3,6 +3,8 @@
 #ifndef ARM_SUSPEND_H
 #define ARM_SUSPEND_H
 
+#include <xen/types.h>
+
 struct domain;
 struct vcpu;
 struct vcpu_guest_context;
@@ -14,6 +16,31 @@ struct resume_info {
 
 void arch_domain_resume(struct domain *d);
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+#ifdef CONFIG_ARM_64
+struct resume_cpu_context {
+    register_t callee_regs[12];
+    register_t sp;
+    register_t vbar_el2;
+    register_t vtcr_el2;
+    register_t vttbr_el2;
+    register_t tpidr_el2;
+    register_t mdcr_el2;
+    register_t hstr_el2;
+    register_t cptr_el2;
+    register_t hcr_el2;
+    register_t cnthctl_el2;
+} __aligned(16);
+#else
+#error "Define resume_cpu_context structure for arm32"
+#endif
+
+extern struct resume_cpu_context resume_cpu_context;
+
+int prepare_resume_ctx(void);
+void hyp_resume(void);
+#endif /* CONFIG_SYSTEM_SUSPEND */
+
 #endif /* ARM_SUSPEND_H */
 
 /*
diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
new file mode 100644
index 0000000000..6ea4a0f9cc
--- /dev/null
+++ b/xen/arch/arm/suspend.c
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+
+#include <asm/suspend.h>
+
+struct resume_cpu_context resume_cpu_context;
+
+/*
+ * Local variables:
+ * mode: C
+ * c-file-style: "BSD"
+ * c-basic-offset: 4
+ * indent-tabs-mode: nil
+ * End:
+ */
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface)
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (9 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND Mykola Kvach
@ 2026-08-27 14:31 ` Mykola Kvach
  2026-09-28 16:18   ` Bertrand Marquis
  2026-08-27 14:32 ` [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy Mykola Kvach
                   ` (2 subsequent siblings)
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:31 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

From: Mirela Simonovic <mirela.simonovic@aggios.com>

Invoke PSCI SYSTEM_SUSPEND to finalize Xen's suspend sequence on ARM64
platforms. Pass the Xen resume entry point (hyp_resume) to EL3 together
with a zero context ID, matching Linux.

This patch wires up only the host-side PSCI SYSTEM_SUSPEND invocation.
The resume trampoline and context restore are provided by earlier patches
in the series.

Only enable this path when CONFIG_SYSTEM_SUSPEND is set and PSCI
advertises SYSTEM_SUSPEND via PSCI_FEATURES.

Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
---
Changes in v9:
- cache SYSTEM_SUSPEND support using PSCI_FEATURES and gate the host call
  on the cached capability
- keep the cached SYSTEM_SUSPEND capability read-only after init
- log whether firmware reports SYSTEM_SUSPEND support
- pass an explicit zero context ID in the SYSTEM_SUSPEND call
- drop the stale note claiming hyp_resume is still a stub
---
 xen/arch/arm/include/asm/psci.h |  1 +
 xen/arch/arm/psci.c             | 31 ++++++++++++++++++++++++++++++-
 2 files changed, 31 insertions(+), 1 deletion(-)

diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
index 48a93e6b79..bb3c73496e 100644
--- a/xen/arch/arm/include/asm/psci.h
+++ b/xen/arch/arm/include/asm/psci.h
@@ -23,6 +23,7 @@ int call_psci_cpu_on(int cpu);
 void call_psci_cpu_off(void);
 void call_psci_system_off(void);
 void call_psci_system_reset(void);
+int call_psci_system_suspend(void);
 
 /* Range of allocated PSCI function numbers */
 #define	PSCI_FNUM_MIN_VALUE                 _AC(0,U)
diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
index b6860a7760..e05dae1133 100644
--- a/xen/arch/arm/psci.c
+++ b/xen/arch/arm/psci.c
@@ -17,23 +17,27 @@
 #include <asm/cpufeature.h>
 #include <asm/psci.h>
 #include <asm/acpi.h>
+#include <asm/suspend.h>
 
 /*
  * While a 64-bit OS can make calls with SMC32 calling conventions, for
  * some calls it is necessary to use SMC64 to pass or return 64-bit values.
- * For such calls PSCI_0_2_FN_NATIVE(x) will choose the appropriate
+ * For such calls PSCI_*_FN_NATIVE(x) will choose the appropriate
  * (native-width) function ID.
  */
 #ifdef CONFIG_ARM_64
 #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN64_##name
+#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN64_##name
 #else
 #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN32_##name
+#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN32_##name
 #endif
 
 uint32_t psci_ver;
 uint32_t smccc_ver;
 
 static uint32_t psci_cpu_on_nr;
+static bool __ro_after_init has_psci_system_suspend;
 
 #define PSCI_RET(res)   ((int32_t)(res).a0)
 
@@ -60,6 +64,25 @@ void call_psci_cpu_off(void)
     }
 }
 
+int call_psci_system_suspend(void)
+{
+#ifdef CONFIG_SYSTEM_SUSPEND
+    struct arm_smccc_res res;
+
+    if ( !has_psci_system_suspend )
+        return PSCI_NOT_SUPPORTED;
+
+    /* Context ID is unused for the Xen resume path. */
+    arm_smccc_smc(PSCI_1_0_FN_NATIVE(SYSTEM_SUSPEND), __pa(hyp_resume), 0,
+                  &res);
+    return PSCI_RET(res);
+#else
+    dprintk(XENLOG_WARNING,
+            "SYSTEM_SUSPEND not supported (CONFIG_SYSTEM_SUSPEND disabled)\n");
+    return PSCI_NOT_SUPPORTED;
+#endif
+}
+
 void call_psci_system_off(void)
 {
     if ( psci_ver > PSCI_VERSION(0, 1) )
@@ -223,9 +246,15 @@ int __init psci_init(void)
 
     psci_init_smccc();
 
+    has_psci_system_suspend =
+        psci_features(PSCI_1_0_FN_NATIVE(SYSTEM_SUSPEND)) == 0;
+
     printk(XENLOG_INFO "Using PSCI v%u.%u\n",
            PSCI_VERSION_MAJOR(psci_ver), PSCI_VERSION_MINOR(psci_ver));
 
+    printk(XENLOG_DEBUG "PSCI SYSTEM_SUSPEND is %ssupported by firmware\n",
+           has_psci_system_suspend ? "" : "not ");
+
     return 0;
 }
 
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (10 preceding siblings ...)
  2026-08-27 14:31 ` [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface) Mykola Kvach
@ 2026-08-27 14:32 ` Mykola Kvach
  2026-09-28 16:18   ` Bertrand Marquis
  2026-08-27 14:32 ` [PATCH v12 13/13] xen/arm: Add host system suspend backend Mykola Kvach
  2026-09-22  7:04 ` Ping: [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
  13 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:32 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Rahul Singh,
	Oleksandr Tyshchenko

Introduce CONFIG_HAS_HWDOM_SYSTEM_SUSPEND as an architecture-selected
capability for platforms where the hardware domain can be parked with
SHUTDOWN_suspend without calling hwdom_shutdown().

Expose PSCI SYSTEM_SUSPEND as a vPSCI operation for all domains. For
non-control domains, including the hardware domain when it is not acting
as a control domain, the call is handled as a guest/domain suspend request
and parks the domain in SHUTDOWN_suspend.

Control domains need additional sequencing because their SYSTEM_SUSPEND
request is used to coordinate host-wide suspend. A non-last awake control
domain may be parked in SHUTDOWN_suspend without requiring the host
suspend path to be available. The last awake control domain is treated as
the point where the request becomes a host-suspend request, and it may
only proceed when all non-control domains are already in SHUTDOWN_suspend
and the host suspend path is available.

Keep the control-domain sequencing and domain-readiness checks out of
PSCI_FEATURES. They are per-attempt runtime conditions rather than stable
PSCI function availability. Advertise SYSTEM_SUSPEND as implemented by
vPSCI and report attempt-time policy failures as PSCI_DENIED.

Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
so that SHUTDOWN_suspend from the hardware domain can be treated as a
domain suspend state rather than as a hardware-domain initiated host
shutdown. This does not by itself imply that host-wide suspend is
available.

Add host_system_suspend_allowed() to combine the host PSCI SYSTEM_SUSPEND
capability with runtime blockers reported by Xen-owned subsystems. Add
runtime blockers for registered serial, IOMMU, GIC and SMMUv3 MSI IRQ
paths lacking suspend/resume support. These blockers are runtime based,
so they only apply to drivers or paths that Xen actually uses on the
platform. For SMMUv3, the blocker applies only when Xen actually uses the
MSI IRQ path, since resume does not restore the SMMU *_IRQ_CFGn MSI
registers yet.

Add a struct domain forward declaration to xen/suspend.h so the generic
header can expose arch_domain_resume() without requiring a full domain.h
include.

Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
Reviewed-by: Oleksandr Tyshchenko <Oleksandr_Tyshchenko@epam.com>
---
Changes in V12:
- handle missing is_shut_down, change checking to call of
  domain_shutdown_completed

Changes in V11:
- Mark host_system_suspend_runtime_allowed as __ro_after_init.
- Avoid printing the SMMUv3 MSI IRQ host suspend blocker more than once
  when multiple SMMUv3 instances use MSIs.
- Wrap the Arm IOMMU host suspend blocker in CONFIG_SYSTEM_SUSPEND to make
  its policy-only use explicit.

Changes in V10:
- Return PSCI_DENIED rather than PSCI_NOT_SUPPORTED when the last awake
  control domain cannot proceed to host suspend, keeping PSCI_FEATURES
  stable once SYSTEM_SUSPEND is advertised.
- Shorten SYSTEM_SUSPEND blocker messages and use %pd when logging the
  control domain.
- Mark serial_suspend_available as __ro_after_init.
- Mention the struct domain forward declaration added to xen/suspend.h.

Changes in V9:
- Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
  so that hardware-domain SHUTDOWN_suspend support is not tied to
  host-wide system suspend availability.
- Add runtime host suspend blockers for Xen-owned subsystems lacking
  suspend/resume support.
- Keep vPSCI SYSTEM_SUSPEND advertised through PSCI_FEATURES and enforce
  control-domain sequencing in the call handler.
---
 xen/arch/arm/Kconfig                  |   1 +
 xen/arch/arm/gic.c                    |   6 ++
 xen/arch/arm/include/asm/psci.h       |   3 +
 xen/arch/arm/include/asm/suspend.h    |  10 ++-
 xen/arch/arm/psci.c                   |   7 ++
 xen/arch/arm/suspend.c                |  40 +++++++++
 xen/arch/arm/vpsci.c                  | 114 +++++++++++++++++++++++---
 xen/common/Kconfig                    |   3 +
 xen/common/domain.c                   |   7 +-
 xen/drivers/char/serial.c             |  12 +++
 xen/drivers/passthrough/arm/iommu.c   |   6 ++
 xen/drivers/passthrough/arm/smmu-v3.c |   9 ++
 xen/include/xen/serial.h              |   1 +
 xen/include/xen/suspend.h             |   2 +
 14 files changed, 208 insertions(+), 13 deletions(-)

diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
index 843a43897e..9027aa17eb 100644
--- a/xen/arch/arm/Kconfig
+++ b/xen/arch/arm/Kconfig
@@ -19,6 +19,7 @@ config ARM
 	select HAS_ALTERNATIVE if HAS_VMAP
 	select HAS_DEVICE_TREE_DISCOVERY
 	select HAS_DOM0LESS
+	select HAS_HWDOM_SYSTEM_SUSPEND if !MPU
 	select HAS_GRANT_CACHE_FLUSH if GRANT_TABLE
 	select HAS_STACK_PROTECTOR
 	select HAS_STATIC_MEMORY
diff --git a/xen/arch/arm/gic.c b/xen/arch/arm/gic.c
index ffc11f36a1..0695474432 100644
--- a/xen/arch/arm/gic.c
+++ b/xen/arch/arm/gic.c
@@ -26,6 +26,7 @@
 #include <asm/device.h>
 #include <asm/io.h>
 #include <asm/gic.h>
+#include <asm/suspend.h>
 #include <asm/vgic.h>
 #include <asm/acpi.h>
 
@@ -44,6 +45,11 @@ static void __init __maybe_unused build_assertions(void)
 void register_gic_ops(const struct gic_hw_operations *ops)
 {
     gic_hw_ops = ops;
+
+#ifdef CONFIG_SYSTEM_SUSPEND
+    if ( !ops->suspend || !ops->resume )
+        host_system_suspend_disable("GIC driver lacks suspend support");
+#endif
 }
 
 static void clear_cpu_lr_mask(void)
diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
index bb3c73496e..142fa1bfe5 100644
--- a/xen/arch/arm/include/asm/psci.h
+++ b/xen/arch/arm/include/asm/psci.h
@@ -24,6 +24,9 @@ void call_psci_cpu_off(void);
 void call_psci_system_off(void);
 void call_psci_system_reset(void);
 int call_psci_system_suspend(void);
+#ifdef CONFIG_SYSTEM_SUSPEND
+bool psci_system_suspend_allowed(void);
+#endif
 
 /* Range of allocated PSCI function numbers */
 #define	PSCI_FNUM_MIN_VALUE                 _AC(0,U)
diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
index c848fc6340..50dc6e9fdf 100644
--- a/xen/arch/arm/include/asm/suspend.h
+++ b/xen/arch/arm/include/asm/suspend.h
@@ -39,7 +39,15 @@ extern struct resume_cpu_context resume_cpu_context;
 
 int prepare_resume_ctx(void);
 void hyp_resume(void);
-#endif /* CONFIG_SYSTEM_SUSPEND */
+bool host_system_suspend_allowed(void);
+void host_system_suspend_disable(const char *reason);
+
+#else /* !CONFIG_SYSTEM_SUSPEND */
+
+static inline bool host_system_suspend_allowed(void) { return false; }
+static inline void host_system_suspend_disable(const char *reason) {}
+
+#endif
 
 #endif /* ARM_SUSPEND_H */
 
diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
index e05dae1133..e9d78668fd 100644
--- a/xen/arch/arm/psci.c
+++ b/xen/arch/arm/psci.c
@@ -41,6 +41,13 @@ static bool __ro_after_init has_psci_system_suspend;
 
 #define PSCI_RET(res)   ((int32_t)(res).a0)
 
+#ifdef CONFIG_SYSTEM_SUSPEND
+bool psci_system_suspend_allowed(void)
+{
+    return has_psci_system_suspend;
+}
+#endif
+
 int call_psci_cpu_on(int cpu)
 {
     struct arm_smccc_res res;
diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
index 6ea4a0f9cc..c7c26bcf03 100644
--- a/xen/arch/arm/suspend.c
+++ b/xen/arch/arm/suspend.c
@@ -1,9 +1,49 @@
 /* SPDX-License-Identifier: GPL-2.0-only */
 
+#include <asm/psci.h>
 #include <asm/suspend.h>
 
+#include <xen/lib.h>
+#include <xen/serial.h>
+
 struct resume_cpu_context resume_cpu_context;
 
+/*
+ * Non-PSCI infrastructure can make host suspend impossible even when the PSCI
+ * SYSTEM_SUSPEND conduit is present, e.g. when a Xen-owned driver has no valid
+ * suspend/resume path.
+ *
+ * This gate is checked only when the last awake control domain attempts to
+ * turn a guest SYSTEM_SUSPEND request into a host-suspend request.
+ */
+static bool __ro_after_init host_system_suspend_runtime_allowed = true;
+
+static bool host_serial_suspend_allowed(void)
+{
+    if ( serial_suspend_supported() )
+        return true;
+
+    printk_once(XENLOG_INFO
+                "Host SYSTEM_SUSPEND blocked: serial unsupported\n");
+
+    return false;
+}
+
+bool host_system_suspend_allowed(void)
+{
+    return psci_system_suspend_allowed() &&
+           host_serial_suspend_allowed() &&
+           host_system_suspend_runtime_allowed;
+}
+
+void host_system_suspend_disable(const char *reason)
+{
+    host_system_suspend_runtime_allowed = false;
+
+    printk(XENLOG_INFO "Host SYSTEM_SUSPEND blocked: %s\n",
+           reason ? reason : "unsupported suspend/resume path");
+}
+
 /*
  * Local variables:
  * mode: C
diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
index ac6af6118f..a41355d75d 100644
--- a/xen/arch/arm/vpsci.c
+++ b/xen/arch/arm/vpsci.c
@@ -5,6 +5,7 @@
 
 #include <asm/current.h>
 #include <asm/domain.h>
+#include <asm/suspend.h>
 #include <asm/vgic.h>
 #include <asm/vpsci.h>
 #include <asm/event.h>
@@ -219,6 +220,89 @@ static void do_psci_0_2_system_reset(void)
     domain_shutdown(d,SHUTDOWN_reboot);
 }
 
+/*
+ * Serialise SYSTEM_SUSPEND policy decisions with the domain suspend transition,
+ * so multiple control domains cannot all observe each other as still awake.
+ */
+static DEFINE_SPINLOCK(vpsci_system_suspend_lock);
+
+static bool domain_in_suspend_state(struct domain *d)
+{
+    bool suspended;
+
+    spin_lock(&d->shutdown_lock);
+    suspended = domain_shutdown_completed(d) && (d->shutdown_code == SHUTDOWN_suspend);
+    spin_unlock(&d->shutdown_lock);
+
+    return suspended;
+}
+
+static int32_t domain_psci_system_suspend_policy(struct domain *d)
+{
+    struct domain *other;
+    bool last_awake_control_domain = true;
+    bool awake_non_control_domain = false;
+
+    /* Only control domains participate in sequencing policy. */
+    if ( !is_control_domain(d) )
+        return 0;
+
+    rcu_read_lock(&domlist_read_lock);
+
+    for_each_domain ( other )
+    {
+        bool suspended;
+
+        if ( other == d )
+            continue;
+
+        suspended = domain_in_suspend_state(other);
+        if ( suspended )
+            continue;
+
+        if ( is_control_domain(other) )
+        {
+            last_awake_control_domain = false;
+            break;
+        }
+
+        awake_non_control_domain = true;
+    }
+
+    rcu_read_unlock(&domlist_read_lock);
+
+    /*
+     * Another control domain is still awake. This request is only the first
+     * phase of the sequencing: park this control domain and leave the host
+     * running. Host-wide suspend gates must not block this intermediate state.
+     */
+    if ( !last_awake_control_domain )
+        return 0;
+
+    /*
+     * This is the last awake control domain. It must not be parked unless the
+     * request can proceed as a host-suspend request; otherwise Xen would lose
+     * the last domain that can coordinate the system suspend.
+     */
+    if ( awake_non_control_domain )
+    {
+        printk(XENLOG_DEBUG
+               "SYSTEM_SUSPEND denied for %pd: non-control domains awake\n",
+               d);
+        return PSCI_DENIED;
+    }
+
+    /*
+     * Host-wide gates are relevant only for the last-control-domain case. They
+     * must not block parking of a non-last control domain, but they must deny
+     * the last control domain when host suspend is not currently available.
+     */
+    if ( !host_system_suspend_allowed() )
+        return PSCI_DENIED;
+
+    return 0;
+}
+
 static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
 {
     int32_t rc;
@@ -232,10 +316,6 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
     if ( is_64bit_domain(d) && is_thumb )
         return PSCI_INVALID_ADDRESS;
 
-    /* SYSTEM_SUSPEND is not supported for the hardware domain yet */
-    if ( is_hardware_domain(d) )
-        return PSCI_NOT_SUPPORTED;
-
     /* Ensure that all CPUs other than the calling one are offline */
     domain_lock(d);
     for_each_vcpu ( d, v )
@@ -252,16 +332,29 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
     if ( rc )
         return PSCI_DENIED;
 
-    rc = domain_shutdown(d, SHUTDOWN_suspend);
+    spin_lock(&vpsci_system_suspend_lock);
+
+    rc = domain_psci_system_suspend_policy(d);
+    if ( !rc )
+    {
+        rc = domain_shutdown(d, SHUTDOWN_suspend);
+        if ( rc )
+            rc = PSCI_DENIED;
+        else
+        {
+            rctx->ctxt = ctxt;
+            rctx->wake_cpu = current;
+        }
+    }
+
+    spin_unlock(&vpsci_system_suspend_lock);
+
     if ( rc )
     {
         free_vcpu_guest_context(ctxt);
-        return PSCI_DENIED;
+        return rc;
     }
 
-    rctx->ctxt = ctxt;
-    rctx->wake_cpu = current;
-
     gprintk(XENLOG_DEBUG,
             "SYSTEM_SUSPEND requested, epoint=%#"PRIregister", cid=%#"PRIregister"\n",
             epoint, cid);
@@ -287,10 +380,9 @@ static int32_t do_psci_1_0_features(uint32_t psci_func_id)
     case PSCI_0_2_FN32_SYSTEM_RESET:
     case PSCI_1_0_FN32_PSCI_FEATURES:
     case ARM_SMCCC_VERSION_FID:
-        return 0;
     case PSCI_1_0_FN32_SYSTEM_SUSPEND:
     case PSCI_1_0_FN64_SYSTEM_SUSPEND:
-        return is_hardware_domain(current->domain) ? PSCI_NOT_SUPPORTED : 0;
+        return 0;
     default:
         return PSCI_NOT_SUPPORTED;
     }
diff --git a/xen/common/Kconfig b/xen/common/Kconfig
index da80fdba84..52bd98f7ad 100644
--- a/xen/common/Kconfig
+++ b/xen/common/Kconfig
@@ -140,6 +140,9 @@ config HAS_EX_TABLE
 config HAS_FAST_MULTIPLY
 	bool
 
+config HAS_HWDOM_SYSTEM_SUSPEND
+	bool
+
 config HAS_IOPORTS
 	bool
 
diff --git a/xen/common/domain.c b/xen/common/domain.c
index e16f1ac383..10c358c7aa 100644
--- a/xen/common/domain.c
+++ b/xen/common/domain.c
@@ -1377,6 +1377,11 @@ void __domain_crash(struct domain *d)
     domain_shutdown(d, SHUTDOWN_crash);
 }
 
+static inline bool want_hwdom_shutdown(uint8_t reason)
+{
+    return !IS_ENABLED(CONFIG_HAS_HWDOM_SYSTEM_SUSPEND) ||
+           reason != SHUTDOWN_suspend;
+}
 
 int domain_shutdown(struct domain *d, u8 reason)
 {
@@ -1393,7 +1398,7 @@ int domain_shutdown(struct domain *d, u8 reason)
         d->shutdown_code = reason;
     reason = d->shutdown_code;
 
-    if ( is_hardware_domain(d) )
+    if ( is_hardware_domain(d) && want_hwdom_shutdown(reason) )
         hwdom_shutdown(reason);
 
     if ( domain_shutting_down(d) )
diff --git a/xen/drivers/char/serial.c b/xen/drivers/char/serial.c
index cf0abf1893..1cdf4968ac 100644
--- a/xen/drivers/char/serial.c
+++ b/xen/drivers/char/serial.c
@@ -490,6 +490,8 @@ const struct vuart_info *serial_vuart_info(int idx)
 
 #ifdef CONFIG_SYSTEM_SUSPEND
 
+static bool __ro_after_init serial_suspend_available = true;
+
 void serial_suspend(void)
 {
     int i;
@@ -506,6 +508,11 @@ void serial_resume(void)
             com[i].driver->resume(&com[i]);
 }
 
+bool serial_suspend_supported(void)
+{
+    return serial_suspend_available;
+}
+
 #endif /* CONFIG_SYSTEM_SUSPEND */
 
 void __init serial_register_uart(int idx, struct uart_driver *driver,
@@ -514,6 +521,11 @@ void __init serial_register_uart(int idx, struct uart_driver *driver,
     /* Store UART-specific info. */
     com[idx].driver = driver;
     com[idx].uart   = uart;
+
+#ifdef CONFIG_SYSTEM_SUSPEND
+    if ( !driver->suspend || !driver->resume )
+        serial_suspend_available = false;
+#endif
 }
 
 void __init serial_async_transmit(struct serial_port *port)
diff --git a/xen/drivers/passthrough/arm/iommu.c b/xen/drivers/passthrough/arm/iommu.c
index 100545e23f..461e01703e 100644
--- a/xen/drivers/passthrough/arm/iommu.c
+++ b/xen/drivers/passthrough/arm/iommu.c
@@ -19,6 +19,7 @@
 #include <xen/device_tree.h>
 #include <xen/iommu.h>
 #include <xen/lib.h>
+#include <xen/suspend.h>
 
 #include <asm/device.h>
 
@@ -46,6 +47,11 @@ void __init iommu_set_ops(const struct iommu_ops *ops)
     }
 
     iommu_ops = ops;
+
+#ifdef CONFIG_SYSTEM_SUSPEND
+    if ( !ops->suspend || !ops->resume )
+        host_system_suspend_disable("IOMMU driver lacks suspend support");
+#endif
 }
 
 int __init iommu_hardware_setup(void)
diff --git a/xen/drivers/passthrough/arm/smmu-v3.c b/xen/drivers/passthrough/arm/smmu-v3.c
index 7f1d00fb81..16947a12f2 100644
--- a/xen/drivers/passthrough/arm/smmu-v3.c
+++ b/xen/drivers/passthrough/arm/smmu-v3.c
@@ -91,6 +91,7 @@
 #include <asm/io.h>
 #include <asm/iommu_fwspec.h>
 #include <asm/platform.h>
+#include <asm/suspend.h>
 
 #include "smmu-v3.h"
 
@@ -1866,6 +1867,7 @@ static void arm_smmu_write_msi_msg(struct msi_desc *desc, struct msi_msg *msg)
 
 static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
 {
+	static bool __ro_after_init host_suspend_blocked_by_msi;
 	struct msi_desc *desc;
 	int ret, nvec = ARM_SMMU_MAX_MSIS;
 	struct device *dev = smmu->dev;
@@ -1910,6 +1912,13 @@ static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
 		}
 	}
 
+	if ( !host_suspend_blocked_by_msi )
+	{
+		host_suspend_blocked_by_msi = true;
+		host_system_suspend_disable(
+			"SMMUv3 MSI IRQ path is unsupported for host suspend");
+	}
+
 	/* Add callback to free MSIs on teardown */
 	devm_add_action(dev, arm_smmu_free_msis, dev);
 }
diff --git a/xen/include/xen/serial.h b/xen/include/xen/serial.h
index 8e18445552..418b00ead0 100644
--- a/xen/include/xen/serial.h
+++ b/xen/include/xen/serial.h
@@ -137,6 +137,7 @@ const struct vuart_info* serial_vuart_info(int idx);
 /* Serial suspend/resume. */
 void serial_suspend(void);
 void serial_resume(void);
+bool serial_suspend_supported(void);
 #endif
 
 /*
diff --git a/xen/include/xen/suspend.h b/xen/include/xen/suspend.h
index 6f94fd53b0..a941331035 100644
--- a/xen/include/xen/suspend.h
+++ b/xen/include/xen/suspend.h
@@ -6,6 +6,8 @@
 #if __has_include(<asm/suspend.h>)
 #include <asm/suspend.h>
 #else
+struct domain;
+
 static inline void arch_domain_resume(struct domain *d) {}
 #endif
 
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* [PATCH v12 13/13] xen/arm: Add host system suspend backend
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (11 preceding siblings ...)
  2026-08-27 14:32 ` [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy Mykola Kvach
@ 2026-08-27 14:32 ` Mykola Kvach
  2026-08-27 21:59   ` Volodymyr Babchuk
  2026-09-28 16:19   ` Bertrand Marquis
  2026-09-22  7:04 ` Ping: [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
  13 siblings, 2 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-08-27 14:32 UTC (permalink / raw)
  To: xen-devel
  Cc: Mykola Kvach, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk

From: Mirela Simonovic <mirela.simonovic@aggios.com>

Add the Xen-wide suspend/resume backend used after a control-domain
vPSCI SYSTEM_SUSPEND request has been accepted. The vPSCI policy,
runtime driver blockers and control-domain sequencing checks are handled
by the preceding commit; this change adds the code that actually drives
the host suspend attempt.

The backend runs from a tasklet scheduled on pCPU0, because non-boot CPUs
are disabled during suspend. It freezes domains, disables the scheduler
and then disables non-boot CPUs.

Host-side suspend participants are handled in phases. IOMMU and console
state are suspended first. Local IRQs are then disabled before suspending
timer and GIC state. On resume or failure, the completed suspend phases
are unwound in reverse: GIC and timer state are restored while IRQs are
still disabled, local IRQs are restored, and then console and IOMMU state
are restored.

On boot, init_ttbr is normally initialized during secondary CPU hotplug.
On uniprocessor systems this can leave init_ttbr uninitialized, so set it
from the boot CPU before entering suspend.

Note: the code is behind CONFIG_HAS_SYSTEM_SUSPEND, which is currently
only selected when UNSUPPORTED is set and MPU is not set.

Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
---
Changes in V10:
- Re-apply boot CPU local errata/workaround handling after SYSTEM_SUSPEND,
  before resuming the rest of the host suspend path.
- Move set_init_ttbr() declaration to asm/mmu/mm.h, since it is
  MMU-specific.

Changes in V9:
- Split vPSCI availability policy, runtime host-suspend blockers and the
  domain-readiness precheck into the preceding commit.
- Trigger the host suspend backend from the control-domain SYSTEM_SUSPEND
  path.
- Reorder the host suspend/resume phases so the timer is suspended with
  local IRQs disabled and local IRQs are restored after the GIC and timer
  resume paths, before the console and IOMMU resume paths.
- Move HAS_HWDOM_SYSTEM_SUSPEND and related logic to policy patch.

Changes in V8:
- Add a pre-suspend check in system_suspend() after scheduler_disable() to
  require all domains to be in the shut down state with SHUTDOWN_suspend
  before proceeding with the global suspend flow.
- Drop the common-level depends on !ARM_64 || !SYSTEM_SUSPEND from
  CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND and model the ARM64 suspend case
  with an arch-selected capability instead.
- Rename CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND to
  CONFIG_HAS_HWDOM_SYSTEM_SUSPEND.
- Rename need_hwdom_shutdown() to want_hwdom_shutdown().

Changes in V7:
- Control domain is responsible for host suspend.
- Add an empty inline host_system_suspend() function when SYSTEM_SUSPEND
  config is disabled.
- Use IS_ENABLED() for config checking instead of #ifdef.
- Replace #ifdef checks in domain_shutdown() with IS_ENABLED() to simplify
  control flow.
- Factor hardware domain shutdown condition into a helper
  (need_hwdom_shutdown()) to avoid preprocessor directives inside the
  function.
- Squash with iommu suspend/resume commit.
---
 xen/arch/arm/Kconfig                 |   1 +
 xen/arch/arm/cpuerrata.c             |   7 +-
 xen/arch/arm/include/asm/cpuerrata.h |   1 +
 xen/arch/arm/include/asm/mmu/mm.h    |   2 +
 xen/arch/arm/include/asm/suspend.h   |   2 +
 xen/arch/arm/mmu/smpboot.c           |   2 +-
 xen/arch/arm/suspend.c               | 156 +++++++++++++++++++++++++++
 xen/arch/arm/vpsci.c                 |  10 +-
 8 files changed, 177 insertions(+), 4 deletions(-)

diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
index 9027aa17eb..da1585ec50 100644
--- a/xen/arch/arm/Kconfig
+++ b/xen/arch/arm/Kconfig
@@ -9,6 +9,7 @@ config ARM_64
 	select 64BIT
 	select HAS_DOMAIN_TYPE
 	select HAS_FAST_MULTIPLY
+	select HAS_SYSTEM_SUSPEND if !MPU && UNSUPPORTED
 	select HAS_VPCI_GUEST_SUPPORT if PCI_PASSTHROUGH
 
 config ARM
diff --git a/xen/arch/arm/cpuerrata.c b/xen/arch/arm/cpuerrata.c
index 3a32183618..e6499aaab3 100644
--- a/xen/arch/arm/cpuerrata.c
+++ b/xen/arch/arm/cpuerrata.c
@@ -782,6 +782,11 @@ void check_local_cpu_errata(void)
     update_cpu_capabilities(arm_errata, "enabled workaround for");
 }
 
+int enable_local_cpu_errata_workarounds(void)
+{
+    return enable_nonboot_cpu_caps(arm_errata);
+}
+
 void __init enable_errata_workarounds(void)
 {
     enable_cpu_capabilities(arm_errata);
@@ -818,7 +823,7 @@ static int cpu_errata_callback(struct notifier_block *nfb,
          * fixed to expect an error at CPU_STARTING phase.
          */
         ASSERT(system_state != SYS_STATE_boot);
-        rc = enable_nonboot_cpu_caps(arm_errata);
+        rc = enable_local_cpu_errata_workarounds();
         break;
     default:
         break;
diff --git a/xen/arch/arm/include/asm/cpuerrata.h b/xen/arch/arm/include/asm/cpuerrata.h
index 1799a16d7e..b93521326f 100644
--- a/xen/arch/arm/include/asm/cpuerrata.h
+++ b/xen/arch/arm/include/asm/cpuerrata.h
@@ -5,6 +5,7 @@
 #include <asm/alternative.h>
 
 void check_local_cpu_errata(void);
+int enable_local_cpu_errata_workarounds(void);
 void enable_errata_workarounds(void);
 
 #define CHECK_WORKAROUND_HELPER(erratum, feature, arch)         \
diff --git a/xen/arch/arm/include/asm/mmu/mm.h b/xen/arch/arm/include/asm/mmu/mm.h
index 7f4d59137d..ee73a77777 100644
--- a/xen/arch/arm/include/asm/mmu/mm.h
+++ b/xen/arch/arm/include/asm/mmu/mm.h
@@ -110,6 +110,8 @@ void dump_pt_walk(paddr_t ttbr, paddr_t addr,
 extern void switch_ttbr(uint64_t ttbr);
 extern void relocate_and_switch_ttbr(uint64_t ttbr);
 
+void set_init_ttbr(lpae_t *root);
+
 #endif /* __ARM_MMU_MM_H__ */
 
 /*
diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
index 50dc6e9fdf..889a6509d9 100644
--- a/xen/arch/arm/include/asm/suspend.h
+++ b/xen/arch/arm/include/asm/suspend.h
@@ -41,11 +41,13 @@ int prepare_resume_ctx(void);
 void hyp_resume(void);
 bool host_system_suspend_allowed(void);
 void host_system_suspend_disable(const char *reason);
+void host_system_suspend(struct domain *d);
 
 #else /* !CONFIG_SYSTEM_SUSPEND */
 
 static inline bool host_system_suspend_allowed(void) { return false; }
 static inline void host_system_suspend_disable(const char *reason) {}
+static inline void host_system_suspend(struct domain *d) {}
 
 #endif
 
diff --git a/xen/arch/arm/mmu/smpboot.c b/xen/arch/arm/mmu/smpboot.c
index 37e91d72b7..ff508ecf40 100644
--- a/xen/arch/arm/mmu/smpboot.c
+++ b/xen/arch/arm/mmu/smpboot.c
@@ -72,7 +72,7 @@ static void clear_boot_pagetables(void)
     clear_table(boot_third);
 }
 
-static void set_init_ttbr(lpae_t *root)
+void set_init_ttbr(lpae_t *root)
 {
     /*
      * init_ttbr is part of the identity mapping which is read-only. So
diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
index c7c26bcf03..3fe2ffa4fb 100644
--- a/xen/arch/arm/suspend.c
+++ b/xen/arch/arm/suspend.c
@@ -1,10 +1,18 @@
 /* SPDX-License-Identifier: GPL-2.0-only */
 
+#include <asm/cpuerrata.h>
+#include <asm/cpufeature.h>
+#include <asm/gic.h>
 #include <asm/psci.h>
 #include <asm/suspend.h>
 
+#include <xen/console.h>
+#include <xen/cpu.h>
+#include <xen/iommu.h>
 #include <xen/lib.h>
+#include <xen/sched.h>
 #include <xen/serial.h>
+#include <xen/tasklet.h>
 
 struct resume_cpu_context resume_cpu_context;
 
@@ -44,6 +52,154 @@ void host_system_suspend_disable(const char *reason)
            reason ? reason : "unsupported suspend/resume path");
 }
 
+/* Xen suspend. data identifies the domain that initiated suspend. */
+static void system_suspend(void *data)
+{
+    int status;
+    unsigned long flags;
+    struct domain *d = (struct domain *)data;
+
+    BUG_ON(system_state != SYS_STATE_active);
+
+    system_state = SYS_STATE_suspend;
+
+    printk("Xen suspending...\n");
+
+    freeze_domains();
+    scheduler_disable();
+
+    /*
+     * Non-boot CPUs have to be disabled on suspend and enabled on resume
+     * (hotplug-based mechanism). Disabling non-boot CPUs will lead to PSCI
+     * CPU_OFF to be called by each non-boot CPU. Depending on the underlying
+     * platform capabilities, this may lead to the physical powering down of
+     * CPUs.
+     */
+    status = disable_nonboot_cpus();
+    if ( status )
+    {
+        system_state = SYS_STATE_resume;
+        goto resume_nonboot_cpus;
+    }
+
+    console_start_sync();
+    status = iommu_suspend();
+    if ( status )
+    {
+        system_state = SYS_STATE_resume;
+        goto resume_end_sync;
+    }
+
+    status = console_suspend();
+    if ( status )
+    {
+        dprintk(XENLOG_ERR, "Failed to suspend the console, err=%d\n", status);
+        system_state = SYS_STATE_resume;
+        goto resume_iommu;
+    }
+
+    local_irq_save(flags);
+
+    time_suspend();
+
+    status = gic_suspend();
+    if ( status )
+    {
+        system_state = SYS_STATE_resume;
+        goto resume_time;
+    }
+
+    set_init_ttbr(xen_pgtable);
+
+    /*
+     * Enable identity mapping before entering suspend to simplify
+     * the resume path
+     */
+    update_boot_mapping(true);
+
+    if ( prepare_resume_ctx() )
+    {
+        status = call_psci_system_suspend();
+        /*
+         * If suspend is finalized properly by above system suspend PSCI call,
+         * the code below in this 'if' branch will never execute. Execution
+         * will continue from hyp_resume which is the hypervisor's resume point.
+         * In hyp_resume CPU context will be restored and since link-register is
+         * restored as well, it will appear to return from prepare_resume_ctx.
+         * The difference in returning from prepare_resume_ctx on system suspend
+         * versus resume is in function's return value: on suspend, the return
+         * value is a non-zero value, on resume it is zero. That is why the
+         * control flow will not re-enter this 'if' branch on resume.
+         */
+        if ( status )
+            dprintk(XENLOG_WARNING, "PSCI system suspend failed, err=%d\n",
+                    status);
+
+        system_state = SYS_STATE_resume;
+    }
+    else
+    {
+        system_state = SYS_STATE_resume;
+
+        /*
+         * CPU0 resumes directly from hyp_resume(), bypassing the CPU hotplug
+         * path that re-checks and re-enables errata workarounds for secondary
+         * CPUs.
+         */
+        check_local_cpu_errata();
+        check_local_cpu_features();
+        BUG_ON(enable_local_cpu_errata_workarounds());
+    }
+
+    update_boot_mapping(false);
+
+    gic_resume();
+
+ resume_time:
+    time_resume();
+
+    local_irq_restore(flags);
+
+    console_resume();
+
+ resume_iommu:
+    iommu_resume();
+
+ resume_end_sync:
+    console_end_sync();
+
+ resume_nonboot_cpus:
+    /*
+     * The rcu_barrier() has to be added to ensure that the per cpu area is
+     * freed before a non-boot CPU tries to initialize it (_free_percpu_area()
+     * has to be called before the init_percpu_area()). This scenario occurs
+     * when non-boot CPUs are hot-unplugged on suspend and hotplugged on resume.
+     */
+    rcu_barrier();
+    enable_nonboot_cpus();
+
+    scheduler_enable();
+    thaw_domains();
+
+    system_state = SYS_STATE_active;
+
+    printk("Resume (status %d)\n", status);
+
+    domain_resume(d);
+}
+
+static DECLARE_TASKLET(system_suspend_tasklet, system_suspend, NULL);
+
+void host_system_suspend(struct domain *d)
+{
+    system_suspend_tasklet.data = (void *)d;
+    /*
+     * The suspend procedure has to be finalized by the pCPU#0 (non-boot pCPUs
+     * will be disabled during the suspend).
+     */
+    tasklet_schedule_on_cpu(&system_suspend_tasklet, 0);
+}
+
 /*
  * Local variables:
  * mode: C
diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
index a41355d75d..5134e75c24 100644
--- a/xen/arch/arm/vpsci.c
+++ b/xen/arch/arm/vpsci.c
@@ -237,7 +237,8 @@ static bool domain_in_suspend_state(struct domain *d)
     return suspended;
 }
 
-static int32_t domain_psci_system_suspend_policy(struct domain *d)
+static int32_t domain_psci_system_suspend_policy(struct domain *d,
+                                                 bool *host_suspend)
 {
     struct domain *other;
     bool last_awake_control_domain = true;
@@ -300,6 +301,7 @@ static int32_t domain_psci_system_suspend_policy(struct domain *d)
     if ( !host_system_suspend_allowed() )
         return PSCI_DENIED;
 
+    *host_suspend = true;
     return 0;
 }
 
@@ -310,6 +312,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
     struct vcpu *v;
     struct domain *d = current->domain;
     bool is_thumb = epoint & 1;
+    bool host_suspend = false;
     struct resume_info *rctx = &d->arch.resume_ctx;
 
     /* THUMB set is not allowed with 64-bit domain */
@@ -334,7 +337,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
 
     spin_lock(&vpsci_system_suspend_lock);
 
-    rc = domain_psci_system_suspend_policy(d);
+    rc = domain_psci_system_suspend_policy(d, &host_suspend);
     if ( !rc )
     {
         rc = domain_shutdown(d, SHUTDOWN_suspend);
@@ -359,6 +362,9 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
             "SYSTEM_SUSPEND requested, epoint=%#"PRIregister", cid=%#"PRIregister"\n",
             epoint, cid);
 
+    if ( host_suspend )
+        host_system_suspend(d);
+
     return rc;
 }
 
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 13/13] xen/arm: Add host system suspend backend
  2026-08-27 14:32 ` [PATCH v12 13/13] xen/arm: Add host system suspend backend Mykola Kvach
@ 2026-08-27 21:59   ` Volodymyr Babchuk
  2026-09-28 16:19   ` Bertrand Marquis
  1 sibling, 0 replies; 37+ messages in thread
From: Volodymyr Babchuk @ 2026-08-27 21:59 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Bertrand Marquis, Michal Orzel

Hi Mykola,

Mykola Kvach <mykola_kvach@epam.com> writes:

> From: Mirela Simonovic <mirela.simonovic@aggios.com>
>
> Add the Xen-wide suspend/resume backend used after a control-domain
> vPSCI SYSTEM_SUSPEND request has been accepted. The vPSCI policy,
> runtime driver blockers and control-domain sequencing checks are handled
> by the preceding commit; this change adds the code that actually drives
> the host suspend attempt.
>
> The backend runs from a tasklet scheduled on pCPU0, because non-boot CPUs
> are disabled during suspend. It freezes domains, disables the scheduler
> and then disables non-boot CPUs.
>
> Host-side suspend participants are handled in phases. IOMMU and console
> state are suspended first. Local IRQs are then disabled before suspending
> timer and GIC state. On resume or failure, the completed suspend phases
> are unwound in reverse: GIC and timer state are restored while IRQs are
> still disabled, local IRQs are restored, and then console and IOMMU state
> are restored.
>
> On boot, init_ttbr is normally initialized during secondary CPU hotplug.
> On uniprocessor systems this can leave init_ttbr uninitialized, so set it
> from the boot CPU before entering suspend.
>
> Note: the code is behind CONFIG_HAS_SYSTEM_SUSPEND, which is currently
> only selected when UNSUPPORTED is set and MPU is not set.
>
> Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>

Reviewed-by: Volodymyr Babchuk <volodymyr_babchuk@epam.com>

> ---
> Changes in V10:
> - Re-apply boot CPU local errata/workaround handling after SYSTEM_SUSPEND,
>   before resuming the rest of the host suspend path.
> - Move set_init_ttbr() declaration to asm/mmu/mm.h, since it is
>   MMU-specific.
>
> Changes in V9:
> - Split vPSCI availability policy, runtime host-suspend blockers and the
>   domain-readiness precheck into the preceding commit.
> - Trigger the host suspend backend from the control-domain SYSTEM_SUSPEND
>   path.
> - Reorder the host suspend/resume phases so the timer is suspended with
>   local IRQs disabled and local IRQs are restored after the GIC and timer
>   resume paths, before the console and IOMMU resume paths.
> - Move HAS_HWDOM_SYSTEM_SUSPEND and related logic to policy patch.
>
> Changes in V8:
> - Add a pre-suspend check in system_suspend() after scheduler_disable() to
>   require all domains to be in the shut down state with SHUTDOWN_suspend
>   before proceeding with the global suspend flow.
> - Drop the common-level depends on !ARM_64 || !SYSTEM_SUSPEND from
>   CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND and model the ARM64 suspend case
>   with an arch-selected capability instead.
> - Rename CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND to
>   CONFIG_HAS_HWDOM_SYSTEM_SUSPEND.
> - Rename need_hwdom_shutdown() to want_hwdom_shutdown().
>
> Changes in V7:
> - Control domain is responsible for host suspend.
> - Add an empty inline host_system_suspend() function when SYSTEM_SUSPEND
>   config is disabled.
> - Use IS_ENABLED() for config checking instead of #ifdef.
> - Replace #ifdef checks in domain_shutdown() with IS_ENABLED() to simplify
>   control flow.
> - Factor hardware domain shutdown condition into a helper
>   (need_hwdom_shutdown()) to avoid preprocessor directives inside the
>   function.
> - Squash with iommu suspend/resume commit.
> ---
>  xen/arch/arm/Kconfig                 |   1 +
>  xen/arch/arm/cpuerrata.c             |   7 +-
>  xen/arch/arm/include/asm/cpuerrata.h |   1 +
>  xen/arch/arm/include/asm/mmu/mm.h    |   2 +
>  xen/arch/arm/include/asm/suspend.h   |   2 +
>  xen/arch/arm/mmu/smpboot.c           |   2 +-
>  xen/arch/arm/suspend.c               | 156 +++++++++++++++++++++++++++
>  xen/arch/arm/vpsci.c                 |  10 +-
>  8 files changed, 177 insertions(+), 4 deletions(-)
>
> diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
> index 9027aa17eb..da1585ec50 100644
> --- a/xen/arch/arm/Kconfig
> +++ b/xen/arch/arm/Kconfig
> @@ -9,6 +9,7 @@ config ARM_64
>  	select 64BIT
>  	select HAS_DOMAIN_TYPE
>  	select HAS_FAST_MULTIPLY
> +	select HAS_SYSTEM_SUSPEND if !MPU && UNSUPPORTED
>  	select HAS_VPCI_GUEST_SUPPORT if PCI_PASSTHROUGH
>  
>  config ARM
> diff --git a/xen/arch/arm/cpuerrata.c b/xen/arch/arm/cpuerrata.c
> index 3a32183618..e6499aaab3 100644
> --- a/xen/arch/arm/cpuerrata.c
> +++ b/xen/arch/arm/cpuerrata.c
> @@ -782,6 +782,11 @@ void check_local_cpu_errata(void)
>      update_cpu_capabilities(arm_errata, "enabled workaround for");
>  }
>  
> +int enable_local_cpu_errata_workarounds(void)
> +{
> +    return enable_nonboot_cpu_caps(arm_errata);
> +}
> +
>  void __init enable_errata_workarounds(void)
>  {
>      enable_cpu_capabilities(arm_errata);
> @@ -818,7 +823,7 @@ static int cpu_errata_callback(struct notifier_block *nfb,
>           * fixed to expect an error at CPU_STARTING phase.
>           */
>          ASSERT(system_state != SYS_STATE_boot);
> -        rc = enable_nonboot_cpu_caps(arm_errata);
> +        rc = enable_local_cpu_errata_workarounds();
>          break;
>      default:
>          break;
> diff --git a/xen/arch/arm/include/asm/cpuerrata.h b/xen/arch/arm/include/asm/cpuerrata.h
> index 1799a16d7e..b93521326f 100644
> --- a/xen/arch/arm/include/asm/cpuerrata.h
> +++ b/xen/arch/arm/include/asm/cpuerrata.h
> @@ -5,6 +5,7 @@
>  #include <asm/alternative.h>
>  
>  void check_local_cpu_errata(void);
> +int enable_local_cpu_errata_workarounds(void);
>  void enable_errata_workarounds(void);
>  
>  #define CHECK_WORKAROUND_HELPER(erratum, feature, arch)         \
> diff --git a/xen/arch/arm/include/asm/mmu/mm.h b/xen/arch/arm/include/asm/mmu/mm.h
> index 7f4d59137d..ee73a77777 100644
> --- a/xen/arch/arm/include/asm/mmu/mm.h
> +++ b/xen/arch/arm/include/asm/mmu/mm.h
> @@ -110,6 +110,8 @@ void dump_pt_walk(paddr_t ttbr, paddr_t addr,
>  extern void switch_ttbr(uint64_t ttbr);
>  extern void relocate_and_switch_ttbr(uint64_t ttbr);
>  
> +void set_init_ttbr(lpae_t *root);
> +
>  #endif /* __ARM_MMU_MM_H__ */
>  
>  /*
> diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> index 50dc6e9fdf..889a6509d9 100644
> --- a/xen/arch/arm/include/asm/suspend.h
> +++ b/xen/arch/arm/include/asm/suspend.h
> @@ -41,11 +41,13 @@ int prepare_resume_ctx(void);
>  void hyp_resume(void);
>  bool host_system_suspend_allowed(void);
>  void host_system_suspend_disable(const char *reason);
> +void host_system_suspend(struct domain *d);
>  
>  #else /* !CONFIG_SYSTEM_SUSPEND */
>  
>  static inline bool host_system_suspend_allowed(void) { return false; }
>  static inline void host_system_suspend_disable(const char *reason) {}
> +static inline void host_system_suspend(struct domain *d) {}
>  
>  #endif
>  
> diff --git a/xen/arch/arm/mmu/smpboot.c b/xen/arch/arm/mmu/smpboot.c
> index 37e91d72b7..ff508ecf40 100644
> --- a/xen/arch/arm/mmu/smpboot.c
> +++ b/xen/arch/arm/mmu/smpboot.c
> @@ -72,7 +72,7 @@ static void clear_boot_pagetables(void)
>      clear_table(boot_third);
>  }
>  
> -static void set_init_ttbr(lpae_t *root)
> +void set_init_ttbr(lpae_t *root)
>  {
>      /*
>       * init_ttbr is part of the identity mapping which is read-only. So
> diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> index c7c26bcf03..3fe2ffa4fb 100644
> --- a/xen/arch/arm/suspend.c
> +++ b/xen/arch/arm/suspend.c
> @@ -1,10 +1,18 @@
>  /* SPDX-License-Identifier: GPL-2.0-only */
>  
> +#include <asm/cpuerrata.h>
> +#include <asm/cpufeature.h>
> +#include <asm/gic.h>
>  #include <asm/psci.h>
>  #include <asm/suspend.h>
>  
> +#include <xen/console.h>
> +#include <xen/cpu.h>
> +#include <xen/iommu.h>
>  #include <xen/lib.h>
> +#include <xen/sched.h>
>  #include <xen/serial.h>
> +#include <xen/tasklet.h>
>  
>  struct resume_cpu_context resume_cpu_context;
>  
> @@ -44,6 +52,154 @@ void host_system_suspend_disable(const char *reason)
>             reason ? reason : "unsupported suspend/resume path");
>  }
>  
> +/* Xen suspend. data identifies the domain that initiated suspend. */
> +static void system_suspend(void *data)
> +{
> +    int status;
> +    unsigned long flags;
> +    struct domain *d = (struct domain *)data;
> +
> +    BUG_ON(system_state != SYS_STATE_active);
> +
> +    system_state = SYS_STATE_suspend;
> +
> +    printk("Xen suspending...\n");
> +
> +    freeze_domains();
> +    scheduler_disable();
> +
> +    /*
> +     * Non-boot CPUs have to be disabled on suspend and enabled on resume
> +     * (hotplug-based mechanism). Disabling non-boot CPUs will lead to PSCI
> +     * CPU_OFF to be called by each non-boot CPU. Depending on the underlying
> +     * platform capabilities, this may lead to the physical powering down of
> +     * CPUs.
> +     */
> +    status = disable_nonboot_cpus();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_nonboot_cpus;
> +    }
> +
> +    console_start_sync();
> +    status = iommu_suspend();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_end_sync;
> +    }
> +
> +    status = console_suspend();
> +    if ( status )
> +    {
> +        dprintk(XENLOG_ERR, "Failed to suspend the console, err=%d\n", status);
> +        system_state = SYS_STATE_resume;
> +        goto resume_iommu;
> +    }
> +
> +    local_irq_save(flags);
> +
> +    time_suspend();
> +
> +    status = gic_suspend();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_time;
> +    }
> +
> +    set_init_ttbr(xen_pgtable);
> +
> +    /*
> +     * Enable identity mapping before entering suspend to simplify
> +     * the resume path
> +     */
> +    update_boot_mapping(true);
> +
> +    if ( prepare_resume_ctx() )
> +    {
> +        status = call_psci_system_suspend();
> +        /*
> +         * If suspend is finalized properly by above system suspend PSCI call,
> +         * the code below in this 'if' branch will never execute. Execution
> +         * will continue from hyp_resume which is the hypervisor's resume point.
> +         * In hyp_resume CPU context will be restored and since link-register is
> +         * restored as well, it will appear to return from prepare_resume_ctx.
> +         * The difference in returning from prepare_resume_ctx on system suspend
> +         * versus resume is in function's return value: on suspend, the return
> +         * value is a non-zero value, on resume it is zero. That is why the
> +         * control flow will not re-enter this 'if' branch on resume.
> +         */
> +        if ( status )
> +            dprintk(XENLOG_WARNING, "PSCI system suspend failed, err=%d\n",
> +                    status);
> +
> +        system_state = SYS_STATE_resume;
> +    }
> +    else
> +    {
> +        system_state = SYS_STATE_resume;
> +
> +        /*
> +         * CPU0 resumes directly from hyp_resume(), bypassing the CPU hotplug
> +         * path that re-checks and re-enables errata workarounds for secondary
> +         * CPUs.
> +         */
> +        check_local_cpu_errata();
> +        check_local_cpu_features();
> +        BUG_ON(enable_local_cpu_errata_workarounds());
> +    }
> +
> +    update_boot_mapping(false);
> +
> +    gic_resume();
> +
> + resume_time:
> +    time_resume();
> +
> +    local_irq_restore(flags);
> +
> +    console_resume();
> +
> + resume_iommu:
> +    iommu_resume();
> +
> + resume_end_sync:
> +    console_end_sync();
> +
> + resume_nonboot_cpus:
> +    /*
> +     * The rcu_barrier() has to be added to ensure that the per cpu area is
> +     * freed before a non-boot CPU tries to initialize it (_free_percpu_area()
> +     * has to be called before the init_percpu_area()). This scenario occurs
> +     * when non-boot CPUs are hot-unplugged on suspend and hotplugged on resume.
> +     */
> +    rcu_barrier();
> +    enable_nonboot_cpus();
> +
> +    scheduler_enable();
> +    thaw_domains();
> +
> +    system_state = SYS_STATE_active;
> +
> +    printk("Resume (status %d)\n", status);
> +
> +    domain_resume(d);
> +}
> +
> +static DECLARE_TASKLET(system_suspend_tasklet, system_suspend, NULL);
> +
> +void host_system_suspend(struct domain *d)
> +{
> +    system_suspend_tasklet.data = (void *)d;
> +    /*
> +     * The suspend procedure has to be finalized by the pCPU#0 (non-boot pCPUs
> +     * will be disabled during the suspend).
> +     */
> +    tasklet_schedule_on_cpu(&system_suspend_tasklet, 0);
> +}
> +
>  /*
>   * Local variables:
>   * mode: C
> diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
> index a41355d75d..5134e75c24 100644
> --- a/xen/arch/arm/vpsci.c
> +++ b/xen/arch/arm/vpsci.c
> @@ -237,7 +237,8 @@ static bool domain_in_suspend_state(struct domain *d)
>      return suspended;
>  }
>  
> -static int32_t domain_psci_system_suspend_policy(struct domain *d)
> +static int32_t domain_psci_system_suspend_policy(struct domain *d,
> +                                                 bool *host_suspend)
>  {
>      struct domain *other;
>      bool last_awake_control_domain = true;
> @@ -300,6 +301,7 @@ static int32_t domain_psci_system_suspend_policy(struct domain *d)
>      if ( !host_system_suspend_allowed() )
>          return PSCI_DENIED;
>  
> +    *host_suspend = true;
>      return 0;
>  }
>  
> @@ -310,6 +312,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>      struct vcpu *v;
>      struct domain *d = current->domain;
>      bool is_thumb = epoint & 1;
> +    bool host_suspend = false;
>      struct resume_info *rctx = &d->arch.resume_ctx;
>  
>      /* THUMB set is not allowed with 64-bit domain */
> @@ -334,7 +337,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>  
>      spin_lock(&vpsci_system_suspend_lock);
>  
> -    rc = domain_psci_system_suspend_policy(d);
> +    rc = domain_psci_system_suspend_policy(d, &host_suspend);
>      if ( !rc )
>      {
>          rc = domain_shutdown(d, SHUTDOWN_suspend);
> @@ -359,6 +362,9 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>              "SYSTEM_SUSPEND requested, epoint=%#"PRIregister", cid=%#"PRIregister"\n",
>              epoint, cid);
>  
> +    if ( host_suspend )
> +        host_system_suspend(d);
> +
>      return rc;
>  }

-- 
WBR, Volodymyr

^ permalink raw reply	[flat|nested] 37+ messages in thread

* Ping: [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64
  2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
                   ` (12 preceding siblings ...)
  2026-08-27 14:32 ` [PATCH v12 13/13] xen/arm: Add host system suspend backend Mykola Kvach
@ 2026-09-22  7:04 ` Mykola Kvach
  13 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-09-22  7:04 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: Xen-devel, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Jens Wiklander, Rahul Singh

Hi all,

Just a gentle ping on this series.

I would appreciate any further review or feedback when you have a chance.

Thanks,
Mykola

On Thu, Aug 27, 2026 at 7:22 PM Mykola Kvach <mykola_kvach@epam.com> wrote:
>
> This is part 2 of the ARM Xen system suspend/resume patch series, based
> on earlier work by Mirela Simonovic and Mykyta Poturai.
>
> Part 1, covering guest suspend functionality, is already in mainline.
>
> NOTE: Host-wide suspend/resume support is guarded by CONFIG_SYSTEM_SUSPEND,
> which can currently only be selected when UNSUPPORTED is set, and thus the
> host suspend backend is neither enabled by default nor built in supported
> configurations. The separate HAS_HWDOM_SYSTEM_SUSPEND policy bit only changes
> how ARM treats SHUTDOWN_suspend from the hardware domain; it does not enable
> the host-wide suspend backend by itself.
>
> This version is ported to Xen master and includes extensive improvements
> based on reviewer feedback. The patch series restructures code to improve
> robustness and maintainability, and implements initial ARM64 host-wide
> Suspend-to-RAM support driven by control-domain PSCI SYSTEM_SUSPEND
> requests. vPSCI also exposes SYSTEM_SUSPEND as a domain suspend operation
> for all domains; attempt-time host-suspend policy failures are reported as
> PSCI_DENIED rather than hidden through PSCI_FEATURES.
>
> Key updates in this series:
>  - Introduced architecture-specific suspend/resume infrastructure
>  - Integrated GICv2/GICv3 suspend and resume, including memory-backed context
>    save/restore with error handling
>  - Added time and IRQ suspend/resume hooks, ensuring correct timer/interrupt
>    state across suspend cycles
>  - Implemented proper PSCI SYSTEM_SUSPEND invocation and version checks
>  - Added vPSCI SYSTEM_SUSPEND policy for domain suspend and host-wide
>    control-domain sequencing
>  - Improved state management and recovery in error cases during suspend/resume
>  - Added support for IPMMU-VMSA/SMMUv3 context save/restore
>  - Added support for GICv3 eSPI registers context save/restore
>  - Added support for ITS registers context save/restore
> ---
>
> Link to CI: https://gitlab.com/xen-project/people/mykola_kvach/xen/-/pipelines/2796733872
> ---
>
> TODOs:
>  - Enable "xl suspend" support on ARM
>  - Add suspend/resume CI test for ARM (QEMU if feasible)
>  - PCI suspend ?
> ---
>
> Detailed changelogs can be found in each patch.
>
> Changes in v12:
> - Rebase onto the latest master.
> - Updated patch 12; all other patches are unchanged.
>   No functional changes.
>
> Changes in v11:
> - Keep SMMUv3 reset helpers in init text when CONFIG_SYSTEM_SUSPEND is
>   disabled.
> - Update host suspend policy blockers after review: make the runtime gate
>   __ro_after_init, log the SMMUv3 MSI blocker only once, and wrap the Arm
>   IOMMU blocker in CONFIG_SYSTEM_SUSPEND.
>
> Changes in v10:
> - Clarify the vPSCI SYSTEM_SUSPEND policy summary: keep SYSTEM_SUSPEND
>   advertised once implemented and return PSCI_DENIED, rather than
>   PSCI_NOT_SUPPORTED, for attempt-time host-suspend policy failures.
> - Tighten GICv2/GICv3 suspend/resume based on review feedback: avoid
>   reserved interrupt register ranges, check visible active-priority state,
>   restore configuration before enable state, and re-enable the redistributor
>   before restoring CPU/virtual interface state on abort paths.
> - Refine ITS resume so MAPC is replayed only for ITS-backed collections and
>   clarify the collection-ID assumptions.
> - Rework IPMMU and SMMUv3 resume/suspend handling, including root-before-cache
>   IPMMU restore ordering and disabling SMMU interrupt generation before
>   suspend.
> - Save and restore CNTHCTL_EL2 in the arm64 CPU resume context and simplify
>   the resume trampoline/context hand-off.
> - Re-apply boot CPU errata/workaround handling after SYSTEM_SUSPEND and move
>   set_init_ttbr() declaration to asm/mmu/mm.h.
> - Update patch 12 details: shorten SYSTEM_SUSPEND blocker logs, use %pd for
>   control-domain logging, mark serial_suspend_available as __ro_after_init,
>   and mention the xen/suspend.h struct domain forward declaration.
>
> Changes in v9:
> - Split the control-domain SYSTEM_SUSPEND flow so host availability,
>   runtime blockers and domain-readiness checks are handled separately from
>   the host suspend backend.
> - Gate vPSCI SYSTEM_SUSPEND on cached host PSCI support and Xen runtime
>   suspend blockers, and log firmware support during initialization.
> - Fold the arm64 resume trampoline into the CPU context save/restore patch
>   and use asm-offsets-generated RESUME_CTX_* definitions for the assembly
>   save/restore path.
> - Tighten the GICv2/GICv3/ITS/IPMMU/SMMUv3 suspend/resume paths based on
>   review feedback, including state-save/restore fixes and safer failure
>   handling.
> - Reorder the host suspend/resume phases so timer and GIC state are
>   handled with local IRQs disabled and restored before console/IOMMU
>   resume.
>
> Changes in v8:
> - Rebased to latest master and refreshed the series accordingly.
> - Added a new GICv3 patch to tolerate retained redistributor LPI state
>   across CPU_OFF/CPU_ON.
> - GICv2 suspend now disables the CPU interface and distributor before
>   saving state.
> - GICv3 suspend/resume fixes the redistributor base used for LPI state.
> - ITS and SMMUv3 suspend/resume paths were tightened, with safer
>   restore/rollback handling and stricter fatal-error handling.
> - System suspend now checks that all domains are already in
>   SHUTDOWN_suspend before proceeding, and renames the hardware-domain
>   suspend capability/helper for clearer semantics.
> - Fixed alignment/cleanup issues in the low-level suspend/resume code.
>
> Changes in v7:
> - Timer helper renamed/clarified; virtual/hyper/phys handling documented.
> - GICv2 uses one context block; restore saved CTLR; panic on alloc failure.
> - GICv3/eSPI/ITS always suspend/resume; restore LPI/eSPI; rdist timeout.
> - IPMMU suspend context allocated before PCI setup.
> - System suspend: control domain drives host suspend.
> - Dropped v6 IRQ descriptor restore patches; use setup_irq and re-register
>   local IRQs on resume instead.
>
> For earlier changelogs, please refer to the previous cover letters.
>
> Mirela Simonovic (5):
>   xen/arm: Add suspend and resume timer helpers
>   xen/arm: gic-v2: Implement GIC suspend/resume functions
>   xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND
>   xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface)
>   xen/arm: Add host system suspend backend
>
> Mykola Kvach (7):
>   xen/arm: gic-v3: tolerate retained redistributor LPI state across
>     CPU_OFF
>   xen/arm: gic-v3: Implement GICv3 suspend/resume functions
>   xen/arm: gic-v3: add ITS suspend/resume support
>   xen/arm: tee: keep init_tee_secondary() for hotplug and resume
>   xen/arm: ffa: fix notification SRI across CPU hotplug/suspend
>   xen/arm: smmu-v3: add suspend/resume handlers
>   xen/arm: Add vPSCI SYSTEM_SUSPEND policy
>
> Oleksandr Tyshchenko (1):
>   iommu/ipmmu-vmsa: Implement suspend/resume callbacks
>
>  xen/arch/arm/Kconfig                     |   2 +
>  xen/arch/arm/Makefile                    |   1 +
>  xen/arch/arm/arm64/asm-offsets.c         |  21 +
>  xen/arch/arm/arm64/head.S                | 122 ++++++
>  xen/arch/arm/cpuerrata.c                 |   7 +-
>  xen/arch/arm/gic-v2.c                    | 226 +++++++++++
>  xen/arch/arm/gic-v3-its.c                | 146 ++++++-
>  xen/arch/arm/gic-v3-lpi.c                |  80 +++-
>  xen/arch/arm/gic-v3.c                    | 482 ++++++++++++++++++++++-
>  xen/arch/arm/gic.c                       |  35 ++
>  xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
>  xen/arch/arm/include/asm/cpuerrata.h     |   1 +
>  xen/arch/arm/include/asm/gic.h           |  16 +
>  xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
>  xen/arch/arm/include/asm/gic_v3_its.h    |  28 ++
>  xen/arch/arm/include/asm/mmu/mm.h        |   2 +
>  xen/arch/arm/include/asm/psci.h          |   4 +
>  xen/arch/arm/include/asm/suspend.h       |  37 ++
>  xen/arch/arm/include/asm/time.h          |   5 +
>  xen/arch/arm/mmu/smpboot.c               |   2 +-
>  xen/arch/arm/psci.c                      |  38 +-
>  xen/arch/arm/suspend.c                   | 210 ++++++++++
>  xen/arch/arm/tee/ffa_notif.c             |  63 ++-
>  xen/arch/arm/tee/tee.c                   |   2 +-
>  xen/arch/arm/time.c                      |  44 ++-
>  xen/arch/arm/vpsci.c                     | 120 +++++-
>  xen/common/Kconfig                       |   3 +
>  xen/common/domain.c                      |   7 +-
>  xen/drivers/char/serial.c                |  12 +
>  xen/drivers/passthrough/arm/iommu.c      |   6 +
>  xen/drivers/passthrough/arm/ipmmu-vmsa.c | 323 ++++++++++++++-
>  xen/drivers/passthrough/arm/smmu-v3.c    | 203 ++++++++--
>  xen/include/xen/list.h                   |  14 +
>  xen/include/xen/serial.h                 |   1 +
>  xen/include/xen/suspend.h                |   2 +
>  35 files changed, 2169 insertions(+), 104 deletions(-)
>  create mode 100644 xen/arch/arm/suspend.c
>
> --
> 2.43.0
>
>


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions
  2026-08-27 14:31 ` [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions Mykola Kvach
@ 2026-09-23 15:27   ` Bertrand Marquis
  2026-09-24 22:23     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-23 15:27 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

Sorry for the delay to review this serie.

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> From: Mirela Simonovic <mirela.simonovic@aggios.com>
> 
> System suspend may lead to a state where GIC would be powered down.
> Therefore, Xen should save/restore the context of GIC on suspend/resume.
> 
> Note that the context consists of states of registers which are
> controlled by the hypervisor. Other GIC registers which are accessible
> by guests are saved/restored on context switch.
> 
> Transient physical SGI pending state (GICD_CPENDSGIRn/GICD_SPENDSGIRn)
> is intentionally excluded. CPU-interface active-priority state is also
> not restored across suspend/resume. Xen reaches the final suspend path
> at a quiescent point, so there is no active-priority execution context
> to replay after resume. Enforce this with a runtime check after
> disabling the CPU interface: if any implemented GICC_APRn word is still
> non-zero, restore GICC_CTLR and abort suspend with -EBUSY.

You mention SGI pending state but you do not say what would happen for PPI/SPI
pending state, and the patch does not look at or save/restore GICD_ISPENDR.

Can you clarify what is expected for those?

Cheers
Bertrand

> 
> This does not apply to distributor active state. With GICv2 EOImode==1,
> EOIR only drops the interrupt priority; final deactivation is a separate
> step. For guest-routed interrupts, Xen can have already EOIed the physical
> IRQ while deactivation is still pending on the vGIC/GICV path. Therefore
> GICD_ISACTIVER is preserved as architectural in-flight interrupt state.
> 
> Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in V10:
> - Limit GICC_APR<n> active-priority checks to APR bits visible from
>  the Xen CPU-interface view.
> - Avoid touching reserved GICD_IPRIORITYR/GICD_ITARGETSR words when the
>  last implemented interrupt block is partial.
> - Restore distributor configuration before restoring interrupt enable
>  state, so GICD_ICFGR is written while the corresponding interrupts are
>  disabled.
> 
> Changes in V9:
> - Skip saving/restoring GICD_ITARGETSR0..7 because SGI/PPI target
>  registers hold no state (read-only on MP, RAZ/WI on UP).
> - Add a runtime GICC_APRn quiescence check after disabling the CPU
>  interface, and restore GICC_CTLR before returning -EBUSY.
> 
> Changes in V8:
> - disable cpu interface + distributor before suspend
> - change 0xffffffff to GENMASK;
> - cosmetic changes;
> 
> Changes in V7:
> - Allocate one contiguous memory block for the GICv2 dist suspend context.
> - gicv2_resume() no longer unconditionally re-enables the distributor/CPU
>  interface; it now writes back the saved CTLR values as-is.
> - gicv2_alloc_context() now returns 0 on success and panics on failure,
>  since suspend context allocation is not recoverable.
> ---
> xen/arch/arm/gic-v2.c          | 226 +++++++++++++++++++++++++++++++++
> xen/arch/arm/gic.c             |  29 +++++
> xen/arch/arm/include/asm/gic.h |  12 ++
> 3 files changed, 267 insertions(+)
> 
> diff --git a/xen/arch/arm/gic-v2.c b/xen/arch/arm/gic-v2.c
> index 43a379fdda..a0ef6ffc7f 100644
> --- a/xen/arch/arm/gic-v2.c
> +++ b/xen/arch/arm/gic-v2.c
> @@ -1108,6 +1108,223 @@ static int gicv2_iomem_deny_access(struct domain *d)
>     return iomem_deny_access(d, mfn, mfn + nr - 1);
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +/* This struct represents block of 32 IRQs */
> +struct irq_block {
> +    uint32_t icfgr[2]; /* 2 registers of 16 IRQs each */
> +    uint32_t ipriorityr[8];
> +    uint32_t isenabler;
> +    uint32_t isactiver;
> +    uint32_t itargetsr[8];
> +};
> +
> +/* GICv2 registers to be saved/restored on system suspend/resume */
> +struct gicv2_context {
> +    /* GICC context */
> +    struct cpu_ctx {
> +        uint32_t ctlr;
> +        uint32_t pmr;
> +        uint32_t bpr;
> +    } cpu;
> +
> +    /* GICD context */
> +    struct dist_ctx {
> +        uint32_t ctlr;
> +        /* Includes banked SGI/PPI state for the boot CPU. */
> +        struct irq_block *irqs;
> +    } dist;
> +};
> +
> +static struct gicv2_context gic_ctx;
> +
> +#define GICV2_NR_APRS          4
> +#define GICV2_APR_BITS_PER_REG 32U
> +
> +static int gicv2_check_active_priorities(uint32_t bpr)
> +{
> +    unsigned int i, apr_bits, nr_aprs;
> +
> +    /*
> +     * Xen writes GICC_BPR to 0 during CPU init and does not change it. Per
> +     * IHI0048B.b, a write below the implementation minimum reads back as the
> +     * minimum supported BPR value. Table 4-47 maps that Xen-visible BPR value
> +     * to the visible GICC_APR<n> bits. Avoid reading APR registers outside
> +     * that visible range.
> +     *
> +     * This covers both GICv2 with and without Security Extensions.
> +     */
> +    apr_bits = 1U << (7 - (bpr & 0x7));
> +    nr_aprs = DIV_ROUND_UP(apr_bits, GICV2_APR_BITS_PER_REG);
> +
> +    ASSERT(nr_aprs <= GICV2_NR_APRS);
> +
> +    for ( i = 0; i < nr_aprs; i++ )
> +    {
> +        unsigned int bits = min(GICV2_APR_BITS_PER_REG,
> +                                apr_bits - i * GICV2_APR_BITS_PER_REG);
> +        uint32_t mask = GENMASK(bits - 1, 0);
> +        uint32_t apr = readl_gicc(GICC_APR + i * 4) & mask;
> +
> +        if ( !apr )
> +            continue;
> +
> +        printk(XENLOG_ERR "GICv2: suspend aborted: GICC_APR%u=%#08x\n",
> +               i, apr);
> +        return -EBUSY;
> +    }
> +
> +    return 0;
> +}
> +
> +static int gicv2_suspend(void)
> +{
> +    unsigned int i, blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
> +    int ret;
> +
> +    /* Save GICC_CTLR configuration. */
> +    gic_ctx.cpu.ctlr = readl_gicc(GICC_CTLR);
> +
> +    /* Quiesce the GIC CPU interface before suspend. */
> +    gicv2_cpu_disable();
> +
> +    gic_ctx.cpu.bpr = readl_gicc(GICC_BPR);
> +
> +    /*
> +     * Check the active-priority state for the group Xen drives through the
> +     * CPU interface. GICC_CTL_ENABLE enables Group 0 without SecurityExtn and
> +     * Group 1 in Xen's Non-secure view with SecurityExtn, and in both cases
> +     * the relevant state is visible through GICC_APRn. The APR layout is
> +     * implementation-defined, so only test the bits visible from Xen's CPU
> +     * interface view instead of reading every possible APR register.
> +     */
> +    ret = gicv2_check_active_priorities(gic_ctx.cpu.bpr);
> +    if ( ret )
> +    {
> +        writel_gicc(gic_ctx.cpu.ctlr, GICC_CTLR);
> +        return ret;
> +    }
> +
> +    gic_ctx.cpu.pmr = readl_gicc(GICC_PMR);
> +
> +    /* Save GICD configuration */
> +    gic_ctx.dist.ctlr = readl_gicd(GICD_CTLR);
> +    writel_gicd(0, GICD_CTLR);
> +
> +    for ( i = 0; i < blocks; i++ )
> +    {
> +        struct irq_block *irqs = gic_ctx.dist.irqs + i;
> +        size_t j, off = i * sizeof(irqs->isenabler);
> +        size_t nr_regs = ARRAY_SIZE(irqs->ipriorityr);
> +
> +        if ( i == blocks - 1 )
> +            nr_regs = DIV_ROUND_UP(gicv2_info.nr_lines - i * 32, 4);
> +
> +        irqs->isenabler = readl_gicd(GICD_ISENABLER + off);
> +
> +        /*
> +         * Save distributor active state as part of the hypervisor-owned
> +         * physical interrupt state. In GICv2 EOImode==1, EOIR only drops the
> +         * priority; final deactivation is separate. For guest-routed
> +         * interrupts, Xen may have EOIed the physical IRQ while the guest/vGIC
> +         * side still owns the deactivate step. Therefore GICD_ISACTIVER can
> +         * legitimately remain set even though transient SGI pending state and
> +         * CPU-interface active-priority state are expected to be quiesced here.
> +         */
> +        irqs->isactiver = readl_gicd(GICD_ISACTIVER + off);
> +
> +        off = i * sizeof(irqs->ipriorityr);
> +        for ( j = 0; j < nr_regs; j++ )
> +            irqs->ipriorityr[j] = readl_gicd(GICD_IPRIORITYR + off + j * 4);
> +
> +        /*
> +         * GICD_ITARGETSR0..7 cover SGIs/PPIs and hold no state to save:
> +         * they are read-only on multiprocessor implementations and RAZ/WI
> +         * on uniprocessor implementations.
> +         */
> +        if ( i )
> +        {
> +            off = i * sizeof(irqs->itargetsr);
> +            for ( j = 0; j < nr_regs; j++ )
> +                irqs->itargetsr[j] = readl_gicd(GICD_ITARGETSR + off + j * 4);
> +        }
> +
> +        off = i * sizeof(irqs->icfgr);
> +        for ( j = 0; j < ARRAY_SIZE(irqs->icfgr); j++ )
> +            irqs->icfgr[j] = readl_gicd(GICD_ICFGR + off + j * 4);
> +    }
> +
> +    return 0;
> +}
> +
> +static void gicv2_resume(void)
> +{
> +    unsigned int i, blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
> +
> +    gicv2_cpu_disable();
> +    /* Disable distributor */
> +    writel_gicd(0, GICD_CTLR);
> +
> +    for ( i = 0; i < blocks; i++ )
> +    {
> +        struct irq_block *irqs = gic_ctx.dist.irqs + i;
> +        size_t j, off = i * sizeof(irqs->isenabler);
> +        size_t nr_regs = ARRAY_SIZE(irqs->ipriorityr);
> +
> +        if ( i == blocks - 1 )
> +            nr_regs = DIV_ROUND_UP(gicv2_info.nr_lines - i * 32, 4);
> +
> +        writel_gicd(GENMASK(31, 0), GICD_ICENABLER + off);
> +
> +        off = i * sizeof(irqs->icfgr);
> +        for ( j = 0; j < ARRAY_SIZE(irqs->icfgr); j++ )
> +            writel_gicd(irqs->icfgr[j], GICD_ICFGR + off + j * 4);
> +
> +        off = i * sizeof(irqs->ipriorityr);
> +        for ( j = 0; j < nr_regs; j++ )
> +            writel_gicd(irqs->ipriorityr[j], GICD_IPRIORITYR + off + j * 4);
> +
> +        /*
> +         * GICD_ITARGETSR0..7 cover SGIs/PPIs and hold no state to save:
> +         * they are read-only on multiprocessor implementations and RAZ/WI
> +         * on uniprocessor implementations.
> +         */
> +        if ( i )
> +        {
> +            off = i * sizeof(irqs->itargetsr);
> +            for ( j = 0; j < nr_regs; j++ )
> +                writel_gicd(irqs->itargetsr[j], GICD_ITARGETSR + off + j * 4);
> +        }
> +
> +        off = i * sizeof(irqs->isenabler);
> +        writel_gicd(irqs->isenabler, GICD_ISENABLER + off);
> +
> +        writel_gicd(GENMASK(31, 0), GICD_ICACTIVER + off);
> +        writel_gicd(irqs->isactiver, GICD_ISACTIVER + off);
> +    }
> +
> +    /* Restore distributor control state. */
> +    writel_gicd(gic_ctx.dist.ctlr, GICD_CTLR);
> +
> +    /* Restore GIC CPU interface configuration */
> +    writel_gicc(gic_ctx.cpu.pmr, GICC_PMR);
> +    writel_gicc(gic_ctx.cpu.bpr, GICC_BPR);
> +
> +    /* Enable GIC CPU interface */
> +    writel_gicc(gic_ctx.cpu.ctlr, GICC_CTLR);
> +}
> +
> +static void __init gicv2_alloc_context(void)
> +{
> +    uint32_t blocks = DIV_ROUND_UP(gicv2_info.nr_lines, 32);
> +
> +    gic_ctx.dist.irqs = xzalloc_array(struct irq_block, blocks);
> +    if ( !gic_ctx.dist.irqs )
> +        panic("Failed to allocate memory for GICv2 suspend context\n");
> +}
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> #ifdef CONFIG_ACPI
> static unsigned long gicv2_get_hwdom_extra_madt_size(const struct domain *d)
> {
> @@ -1312,6 +1529,11 @@ static int __init gicv2_init(void)
> 
>     spin_unlock(&gicv2.lock);
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    /* Allocate memory to be used for saving GIC context during the suspend */
> +    gicv2_alloc_context();
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
>     return 0;
> }
> 
> @@ -1355,6 +1577,10 @@ static const struct gic_hw_operations gicv2_ops = {
>     .map_hwdom_extra_mappings = gicv2_map_hwdom_extra_mappings,
>     .iomem_deny_access   = gicv2_iomem_deny_access,
>     .do_LPI              = gicv2_do_LPI,
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    .suspend             = gicv2_suspend,
> +    .resume              = gicv2_resume,
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> };
> 
> /* Set up the GIC */
> diff --git a/xen/arch/arm/gic.c b/xen/arch/arm/gic.c
> index 078049e741..ffc11f36a1 100644
> --- a/xen/arch/arm/gic.c
> +++ b/xen/arch/arm/gic.c
> @@ -438,6 +438,35 @@ int gic_iomem_deny_access(struct domain *d)
>     return gic_hw_ops->iomem_deny_access(d);
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +int gic_suspend(void)
> +{
> +    /* Must be called by boot CPU#0 with interrupts disabled */
> +    ASSERT(!local_irq_is_enabled());
> +    ASSERT(!smp_processor_id());
> +
> +    if ( !gic_hw_ops->suspend || !gic_hw_ops->resume )
> +        return -ENOSYS;
> +
> +    return gic_hw_ops->suspend();
> +}
> +
> +void gic_resume(void)
> +{
> +    /*
> +     * Must be called by boot CPU#0 with interrupts disabled after gic_suspend
> +     * has returned successfully.
> +     */
> +    ASSERT(!local_irq_is_enabled());
> +    ASSERT(!smp_processor_id());
> +    ASSERT(gic_hw_ops->resume);
> +
> +    gic_hw_ops->resume();
> +}
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> static int cpu_gic_callback(struct notifier_block *nfb,
>                             unsigned long action,
>                             void *hcpu)
> diff --git a/xen/arch/arm/include/asm/gic.h b/xen/arch/arm/include/asm/gic.h
> index ee2c26adb4..29bb9a89a4 100644
> --- a/xen/arch/arm/include/asm/gic.h
> +++ b/xen/arch/arm/include/asm/gic.h
> @@ -301,6 +301,12 @@ extern int gicv_setup(struct domain *d);
> extern void gic_save_state(struct vcpu *v);
> extern void gic_restore_state(struct vcpu *v);
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +/* Suspend/resume */
> +extern int gic_suspend(void);
> +extern void gic_resume(void);
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> /* SGI (AKA IPIs) */
> enum gic_sgi {
>     GIC_SGI_EVENT_CHECK,
> @@ -444,6 +450,12 @@ struct gic_hw_operations {
>     int (*iomem_deny_access)(struct domain *d);
>     /* Handle LPIs, which require special handling */
>     void (*do_LPI)(unsigned int lpi);
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    /* Save GIC configuration due to the system suspend */
> +    int (*suspend)(void);
> +    /* Restore GIC configuration due to the system resume */
> +    void (*resume)(void);
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> };
> 
> extern const struct gic_hw_operations *gic_hw_ops;
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF
  2026-08-27 14:31 ` [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF Mykola Kvach
@ 2026-09-23 15:34   ` Bertrand Marquis
  2026-09-24 23:22     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-23 15:34 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> PSCI does not guarantee that a GICv3 redistributor is powered down across
> CPU_OFF -> CPU_ON.
> 
> DEN0022F.b says CPU_OFF powers down the calling core (5.5) and CPU_ON
> brings the core back with a defined initial CPU state (5.6, 6.4).
> However, PSCI leaves interrupt migration and GIC re-initialization to the
> supervisory software/firmware stack: the caller must migrate interrupts
> away before CPU_OFF (5.5.2), and the execution context that is lost in a
> powerdown state must be saved and restored by software (6.8). PSCI also
> calls out GIC management explicitly in 6.8, including retargeting SPIs,
> preventing PPIs/SGIs from targeting a powered down CPU, and reinitializing
> the CPU interface after CPU_ON.
> 
> This matches the GIC architecture. IHI0069H.b Chapter 11.1 requires the PE
> and CPU interface to share a power domain, but explicitly allows the
> associated redistributor, distributor, and ITS to remain powered while the
> PE and CPU interface are off. All other GIC power-management behavior is
> IMPLEMENTATION DEFINED. DEN0050D Chapter 4.2, "Generic Interrupt
> Controller (GIC)", says the GICv3 redistributor may live either in the AP
> core power domain or in a relatively always-on parent domain. So after
> CPU_OFF -> CPU_ON a secondary CPU can legitimately come back to a live
> redistributor with GICR_CTLR.EnableLPIs still set.
> 
> Handle that case in the LPI setup path instead of assuming a fully reset
> redistributor.
> 
> The LPI path needs special care because the GIC spec makes redistributor
> LPI state sticky and partially implementation defined. IHI0069H.b 5.1.1
> and 5.1.2 say that changing GICR_PROPBASER or GICR_PENDBASER while
> GICR_CTLR.EnableLPIs == 1 is UNPREDICTABLE. After clearing EnableLPIs,
> software must wait for GICR_CTLR.RWP == 0 before touching the pending
> table. The architecture also permits implementations where, once
> EnableLPIs has been set, clearing it again is not guaranteed to work.
> Where an ITS is present, the spec strongly recommends moving LPIs to
> another redistributor before clearing EnableLPIs.
> 
> Because of that, treat a retained EnableLPIs state as valid when the
> redistributor still points at Xen's expected PROPBASER/PENDBASER tables.
> Only try to clear EnableLPIs when the retained configuration does not
> match Xen's state, and wait for RWP before reprogramming the tables.
> 
> This is also consistent with platform firmware reality: PSCI and the GIC
> architecture allow platform-specific redistributor power handling, and not
> all platform firmware implementations force a full redistributor power-off
> through implementation-defined controls during CPU_OFF. Xen therefore needs
> to tolerate retained redistributor state on secondary CPU bring-up.
> 
> Keep gicv3_populate_rdist() resident as well, because gicv3_cpu_init()
> reuses it on secondary CPU bring-up after init.
> 
> Tested using Xen's non-boot CPU disable/enable path on Arm
> FVP_Base_RevC-2xAEMvA, both with and without:
> -C gic_distributor.allow-LPIEN-clear=1
> -C gic_distributor.GICR-clear-enable-supported=1
> and on Orange Pi 5.
> 
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in v10:
> - Drop unrelated gicv3_populate_rdist() printk() format cleanups to keep
>  the patch focused on retained redistributor LPI state.
> 
> Changes in v9:
> - move gicv3_do_wait_for_rwp prototype from its related header to gic.h
> - drop __init from gicv3_populate_rdist(), which is reused on secondary
>  CPU bring-up after boot
> - changed print format for smp_processor_id in gicv3_populate_rdist func
> - cosmetic changes
> ---
> xen/arch/arm/gic-v3-lpi.c      | 77 +++++++++++++++++++++++++++++++++-
> xen/arch/arm/gic-v3.c          | 15 ++++---
> xen/arch/arm/include/asm/gic.h |  4 ++
> 3 files changed, 90 insertions(+), 6 deletions(-)
> 
> diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
> index 9ee338edc2..847da26ff7 100644
> --- a/xen/arch/arm/gic-v3-lpi.c
> +++ b/xen/arch/arm/gic-v3-lpi.c
> @@ -81,6 +81,13 @@ static DEFINE_PER_CPU(struct lpi_redist_data, lpi_redist);
> #define MAX_NR_HOST_LPIS   (lpi_data.max_host_lpi_ids - LPI_OFFSET)
> #define HOST_LPIS_PER_PAGE      (PAGE_SIZE / sizeof(union host_lpi))
> 
> +#define GICR_PROPBASER_XEN_MASK  GENMASK_ULL(51, 12)
> +/*
> + * For retained redistributor state, match the pending table by address only.
> + * Attribute bits such as PTZ may not read back with the programmed value.
> + */
> +#define GICR_PENDBASER_XEN_MASK  GENMASK_ULL(51, 16)
> +
> static union host_lpi *gic_get_host_lpi(uint32_t plpi)
> {
>     union host_lpi *block;
> @@ -296,6 +303,60 @@ static int gicv3_lpi_set_pendtable(void __iomem *rdist_base)
>     return 0;
> }
> 
> +static uint64_t gicv3_lpi_expected_proptable(void)
> +{
> +    return virt_to_maddr(lpi_data.lpi_property);
> +}
> +
> +static uint64_t gicv3_lpi_expected_pendtable(void)
> +{
> +    return virt_to_maddr(this_cpu(lpi_redist).pending_table);
> +}
> +
> +static bool gicv3_lpi_tables_match(void __iomem *rdist_base)
> +{
> +    uint64_t propbase, pendbase;
> +
> +    if ( !lpi_data.lpi_property || !this_cpu(lpi_redist).pending_table )
> +        return false;
> +
> +    propbase = readq_relaxed(rdist_base + GICR_PROPBASER);
> +    pendbase = readq_relaxed(rdist_base + GICR_PENDBASER);
> +
> +    return ((propbase & GICR_PROPBASER_XEN_MASK) ==
> +            (gicv3_lpi_expected_proptable() & GICR_PROPBASER_XEN_MASK)) &&
> +           ((pendbase & GICR_PENDBASER_XEN_MASK) ==
> +            (gicv3_lpi_expected_pendtable() & GICR_PENDBASER_XEN_MASK));
> +}
> +
> +static int gicv3_lpi_disable_lpis(void __iomem *rdist_base)
> +{
> +    uint32_t reg = readl_relaxed(rdist_base + GICR_CTLR);
> +    int ret;
> +
> +    if ( !(reg & GICR_CTLR_ENABLE_LPIS) )
> +        return 0;
> +
> +    writel_relaxed(reg & ~GICR_CTLR_ENABLE_LPIS, rdist_base + GICR_CTLR);
> +
> +    /*
> +     * The spec only guarantees programmability when we have observed the bit
> +     * cleared. Where clearing is supported, RWP must reach 0 before touching
> +     * PROPBASER/PENDBASER again.
> +     */
> +    wmb();
> +
> +    ret = gicv3_do_wait_for_rwp(rdist_base, GICR_CTLR_RWP);
> +    if ( ret )
> +        return ret;
> +
> +    reg = readl_relaxed(rdist_base + GICR_CTLR);
> +    if ( reg & GICR_CTLR_ENABLE_LPIS )
> +        return -EBUSY;
> +
> +    return 0;
> +}
> +
> /*
>  * Tell a redistributor about the (shared) property table, allocating one
>  * if not already done.
> @@ -374,7 +435,21 @@ int gicv3_lpi_init_rdist(void __iomem * rdist_base)
>     /* Make sure LPIs are disabled before setting up the tables. */
>     reg = readl_relaxed(rdist_base + GICR_CTLR);
>     if ( reg & GICR_CTLR_ENABLE_LPIS )
> -        return -EBUSY;
> +    {
> +        if ( gicv3_lpi_tables_match(rdist_base) )
> +            return -EBUSY;

I am wondering if there is a corner case when a CPU is unplugged and then
plugged back in. free_percpu_area() eventually frees the per-CPU area
containing lpi_redist.pending_table, but not the table itself. On the next
cpu_up(), gicv3_lpi_allocate_pendtable() allocates a new table, and I cannot
find where the old one is freed.

If the redistributor kept EnableLPIs=1 and GICR_PENDBASER pointing to the
old table, wouldn't gicv3_lpi_tables_match() fail? 

What will happen if EnableLPIs cannot be cleared ? (i think this is something
possible in the hardware). 

Cheers
Bertrand

> +
> +        ret = gicv3_lpi_disable_lpis(rdist_base);
> +        if ( ret == -EBUSY )
> +        {
> +            printk(XENLOG_ERR
> +                   "GICv3: CPU%u: LPIs still enabled with unexpected redistributor tables\n",
> +                   smp_processor_id());
> +            return -EINVAL;
> +        }
> +        if ( ret )
> +            return ret;
> +    }
> 
>     ret = gicv3_lpi_set_pendtable(rdist_base);
>     if ( ret )
> diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
> index acdac22953..b16888ad84 100644
> --- a/xen/arch/arm/gic-v3.c
> +++ b/xen/arch/arm/gic-v3.c
> @@ -275,7 +275,7 @@ static void gicv3_enable_sre(void)
> }
> 
> /* Wait for completion of a distributor/redistributor change */
> -static void gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
> +int gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
> {
>     uint32_t val;
>     bool timeout = false;
> @@ -299,17 +299,22 @@ static void gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit)
>     } while ( 1 );
> 
>     if ( timeout )
> +    {
>         dprintk(XENLOG_ERR, "RWP timeout\n");
> +        return -ETIMEDOUT;
> +    }
> +
> +    return 0;
> }
> 
> static void gicv3_dist_wait_for_rwp(void)
> {
> -    gicv3_do_wait_for_rwp(GICD, GICD_CTLR_RWP);
> +    (void)gicv3_do_wait_for_rwp(GICD, GICD_CTLR_RWP);
> }
> 
> static void gicv3_redist_wait_for_rwp(void)
> {
> -    gicv3_do_wait_for_rwp(GICD_RDIST_BASE, GICR_CTLR_RWP);
> +    (void)gicv3_do_wait_for_rwp(GICD_RDIST_BASE, GICR_CTLR_RWP);
> }
> 
> static void gicv3_wait_for_rwp(int irq)
> @@ -863,7 +868,7 @@ static bool gicv3_enable_lpis(void)
>     return true;
> }
> 
> -static int __init gicv3_populate_rdist(void)
> +static int gicv3_populate_rdist(void)
> {
>     int i;
>     uint32_t aff;
> @@ -931,7 +936,7 @@ static int __init gicv3_populate_rdist(void)
>                     gicv3_set_redist_address(rdist_addr, procnum);
> 
>                     ret = gicv3_lpi_init_rdist(ptr);
> -                    if ( ret && ret != -ENODEV )
> +                    if ( ret && ret != -ENODEV && ret != -EBUSY )
>                     {
>                         printk("GICv3: CPU%d: Cannot initialize LPIs: %u\n",
>                                smp_processor_id(), ret);
> diff --git a/xen/arch/arm/include/asm/gic.h b/xen/arch/arm/include/asm/gic.h
> index 29bb9a89a4..68003ab116 100644
> --- a/xen/arch/arm/include/asm/gic.h
> +++ b/xen/arch/arm/include/asm/gic.h
> @@ -301,6 +301,10 @@ extern int gicv_setup(struct domain *d);
> extern void gic_save_state(struct vcpu *v);
> extern void gic_restore_state(struct vcpu *v);
> 
> +#ifdef CONFIG_GICV3
> +int gicv3_do_wait_for_rwp(void __iomem *base, uint32_t rwp_bit);
> +#endif
> +
> #ifdef CONFIG_SYSTEM_SUSPEND
> /* Suspend/resume */
> extern int gic_suspend(void);
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions
  2026-08-27 14:31 ` [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions Mykola Kvach
@ 2026-09-23 15:35   ` Bertrand Marquis
  2026-09-25  0:03     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-23 15:35 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> System suspend may lead to a state where GIC would be powered down.
> Therefore, Xen should save/restore the context of GIC on suspend/resume.
> 
> Note that the context consists of states of registers which are
> controlled by the hypervisor. Other GIC registers which are accessible
> by guests are saved/restored on context switch.
> 
> Before continuing suspend, also verify that the physical CPU interface
> has no Group 1 active-priority state left. Use ICC_CTLR_EL1.PRIbits to
> decide which ICC_AP1R<n>_EL1 registers are implemented, so Xen does not
> read an unimplemented AP1R register.
> 
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in V10:
> - abort suspend when the physical Group 1 active-priority state is still
>  present, deriving accessible ICC_AP1R<n>_EL1 registers from
>  ICC_CTLR_EL1.PRIbits;
> - re-enable the redistributor before restoring CPU and virtual interface
>  state on the suspend abort path;
> - panic if the redistributor cannot be re-enabled on the suspend abort path;
> - avoid saving/restoring reserved GICD_IPRIORITYR and GICD_IROUTER entries
>  for a partially populated last SPI block;
> - disable Distributor group forwarding while preserving affinity routing
>  state before restoring Distributor configuration;
> - disable SPI/eSPI forwarding and wait for RWP before restoring
>  GICD_ICFGR<n>.Int_config.
> 
> Changes in V9:
> - fix the suspend-context comment typo and split dist_ctx declarations;
> - restore ICC_IGRPEN1_EL1 on the suspend error path;
> - re-initialize GICD_IGROUPRnE during resume;
> - restore GICD_IROUTER only after re-enabling ARE_NS during resume.
> 
> Changes in V8:
> - use right rdist base for prop/pend baser and ctrl
> 
> Changes in V7:
> - restore LPI regs on resume
> - add timeout during redist disabling
> - squash with suspend/resume handling for GICv3 eSPI registers
> - drop ITS guard paths so suspend/resume always runs; switch missing ctx
>  allocation to panic
> - trim TODO comments; narrow redistributor storage to PPI icfgr
> - keep distributor context allocation even without ITS; adjust resume
>  to use GENMASK(31, 0) for clearing enables
> - drop storage of the SGI configuration register, as SGIs are always
>  edge-triggered
> ---
> xen/arch/arm/gic-v3-lpi.c                |   3 +
> xen/arch/arm/gic-v3.c                    | 458 ++++++++++++++++++++++-
> xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
> xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
> 4 files changed, 466 insertions(+), 3 deletions(-)
> 
> diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
> index 847da26ff7..a63c8c4979 100644
> --- a/xen/arch/arm/gic-v3-lpi.c
> +++ b/xen/arch/arm/gic-v3-lpi.c
> @@ -467,6 +467,9 @@ static int cpu_callback(struct notifier_block *nfb, unsigned long action,
>     switch ( action )
>     {
>     case CPU_UP_PREPARE:
> +        if ( system_state == SYS_STATE_resume )
> +            break;
> +
>         rc = gicv3_lpi_allocate_pendtable(cpu);
>         if ( rc )
>             printk(XENLOG_ERR "Unable to allocate the pendtable for CPU%lu\n",
> diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
> index b16888ad84..038bf41142 100644
> --- a/xen/arch/arm/gic-v3.c
> +++ b/xen/arch/arm/gic-v3.c
> @@ -1078,12 +1078,12 @@ out:
>     return res;
> }
> 
> -static void gicv3_hyp_disable(void)
> +static void gicv3_hyp_enable(bool enable)
> {
>     register_t hcr;
> 
>     hcr = READ_SYSREG(ICH_HCR_EL2);
> -    hcr &= ~GICH_HCR_EN;
> +    hcr = enable ? (hcr | GICH_HCR_EN) : (hcr & ~GICH_HCR_EN);
>     WRITE_SYSREG(hcr, ICH_HCR_EL2);
>     isb();
> }
> @@ -1190,7 +1190,7 @@ static void gicv3_disable_interface(void)
>     spin_lock(&gicv3.lock);
> 
>     gicv3_cpu_disable();
> -    gicv3_hyp_disable();
> +    gicv3_hyp_enable(false);
> 
>     spin_unlock(&gicv3.lock);
> }
> @@ -1926,6 +1926,450 @@ static bool gic_dist_supports_lpis(void)
>     return (readl_relaxed(GICD + GICD_TYPER) & GICD_TYPE_LPIS);
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +/* This struct represents a block of 32 IRQs */
> +struct dist_irq_block {
> +    uint32_t icfgr[2];
> +    uint32_t ipriorityr[8];
> +    uint64_t irouter[32];
> +    uint32_t isactiver;
> +    uint32_t isenabler;
> +};
> +
> +struct redist_ctx {
> +    uint32_t ctlr;
> +    uint32_t icfgr; /* only PPIs stored */

Can you capitalize first comment letter ?
s/only/Only/

> +    uint32_t igroupr;
> +    uint32_t ipriorityr[8];
> +    uint32_t isactiver;
> +    uint32_t isenabler;
> +
> +    uint64_t pendbase;
> +    uint64_t propbase;
> +};
> +
> +/* GICv3 registers to be saved/restored on system suspend/resume */
> +struct gicv3_ctx {
> +    struct dist_ctx {
> +        uint32_t ctlr;
> +        struct dist_irq_block *irqs;
> +        struct dist_irq_block *espi_irqs;
> +    } dist;
> +
> +    /* have only one rdist structure for last running CPU during suspend */

Same here
s/have/Have/

> +    struct redist_ctx rdist;
> +
> +    struct cpu_ctx {
> +        uint32_t ctlr;
> +        uint32_t pmr;
> +        uint32_t bpr;
> +        uint32_t sre_el2;
> +        uint32_t grpen;
> +    } cpu;
> +};
> +
> +static struct gicv3_ctx gicv3_ctx;
> +
> +static void __init gicv3_alloc_context(void)
> +{
> +    uint32_t blocks = DIV_ROUND_UP(gicv3_info.nr_lines, 32);
> +
> +    /* The spec allows for systems without any SPIs */
> +    if ( blocks > 1 )
> +    {
> +        gicv3_ctx.dist.irqs = xzalloc_array(struct dist_irq_block, blocks - 1);
> +        if ( !gicv3_ctx.dist.irqs )
> +            panic("Failed to allocate memory for GICv3 suspend context\n");
> +    }
> +
> +#ifdef CONFIG_GICV3_ESPI
> +    if ( !gic_number_espis() )
> +        return;
> +
> +    blocks = gic_number_espis() / 32;
> +    gicv3_ctx.dist.espi_irqs = xzalloc_array(struct dist_irq_block, blocks);
> +    if ( !gicv3_ctx.dist.espi_irqs )
> +        panic("Failed to allocate memory for GICv3 eSPI suspend context\n");
> +#endif
> +}
> +
> +static int gicv3_disable_redist(void)
> +{
> +    void __iomem *waker = GICD_RDIST_BASE + GICR_WAKER;
> +    s_time_t deadline;
> +
> +    /*
> +     * Avoid infinite loop if Non-secure does not have access to GICR_WAKER.
> +     * See Arm IHI 0069H.b, 12.11.42 GICR_WAKER:
> +     *     When GICD_CTLR.DS == 0 and an access is Non-secure accesses to this
> +     *     register are RAZ/WI.
> +     */
> +    if ( !(readl_relaxed(GICD + GICD_CTLR) & GICD_CTLR_DS) )
> +        return 0;
> +
> +    deadline = NOW() + MILLISECS(1000);
> +
> +    writel_relaxed(readl_relaxed(waker) | GICR_WAKER_ProcessorSleep, waker);
> +    while ( (readl_relaxed(waker) & GICR_WAKER_ChildrenAsleep) == 0 )
> +    {
> +        if ( NOW() > deadline )
> +        {
> +            printk("GICv3: Timeout waiting for redistributor to sleep\n");
> +            return -ETIMEDOUT;
> +        }
> +        cpu_relax();
> +        udelay(10);
> +    }
> +
> +    return 0;
> +}
> +
> +#define GET_SPI_REG_OFFSET(name, is_espi) \
> +    ((is_espi) ? GICD_##name##nE : GICD_##name)
> +
> +static void gicv3_store_spi_irq_block(struct dist_irq_block *irqs,
> +                                      unsigned int i, unsigned int nr_irqs,
> +                                      bool is_espi)
> +{
> +    void __iomem *base;
> +    unsigned int irq, nr_priority_regs;
> +
> +    ASSERT(nr_irqs && nr_irqs <= 32);
> +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
> +    irqs->icfgr[0] = readl_relaxed(base);
> +    irqs->icfgr[1] = readl_relaxed(base + 4);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
> +    base += i * sizeof(irqs->ipriorityr);
> +    for ( irq = 0; irq < nr_priority_regs; irq++ )
> +        irqs->ipriorityr[irq] = readl_relaxed(base + 4 * irq);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
> +    base += i * sizeof(irqs->irouter);
> +    for ( irq = 0; irq < nr_irqs; irq++ )
> +        irqs->irouter[irq] = readq_relaxed_non_atomic(base + 8 * irq);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
> +    base += i * sizeof(irqs->isactiver);
> +    irqs->isactiver = readl_relaxed(base);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
> +    base += i * sizeof(irqs->isenabler);
> +    irqs->isenabler = readl_relaxed(base);
> +}
> +
> +static void gicv3_restore_spi_irq_config(struct dist_irq_block *irqs,
> +                                         unsigned int i, unsigned int nr_irqs,
> +                                         bool is_espi)
> +{
> +    void __iomem *base;
> +    unsigned int irq, nr_priority_regs;
> +
> +    ASSERT(nr_irqs && nr_irqs <= 32);
> +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
> +    writel_relaxed(irqs->icfgr[0], base);
> +    writel_relaxed(irqs->icfgr[1], base + 4);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
> +    base += i * sizeof(irqs->ipriorityr);
> +    for ( irq = 0; irq < nr_priority_regs; irq++ )
> +        writel_relaxed(irqs->ipriorityr[irq], base + 4 * irq);
> +}
> +
> +static void gicv3_restore_spi_irq_routing(struct dist_irq_block *irqs,
> +                                          unsigned int i, unsigned int nr_irqs,
> +                                          bool is_espi)
> +{
> +    void __iomem *base;
> +    unsigned int irq;
> +
> +    ASSERT(nr_irqs && nr_irqs <= 32);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
> +    base += i * sizeof(irqs->irouter);
> +    for ( irq = 0; irq < nr_irqs; irq++ )
> +        writeq_relaxed_non_atomic(irqs->irouter[irq], base + 8 * irq);
> +}
> +
> +static void gicv3_disable_spi_irq_block(unsigned int i, bool is_espi)
> +{
> +    void __iomem *base;
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ICENABLER, is_espi) + i * 4;
> +    writel_relaxed(GENMASK(31, 0), base);
> +}
> +
> +static void gicv3_restore_spi_irq_state(struct dist_irq_block *irqs,
> +                                        unsigned int i, bool is_espi)
> +{
> +    void __iomem *base;
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
> +    base += i * sizeof(irqs->isenabler);
> +    writel_relaxed(irqs->isenabler, base);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ICACTIVER, is_espi) + i * 4;
> +    writel_relaxed(GENMASK(31, 0), base);
> +
> +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
> +    base += i * sizeof(irqs->isactiver);
> +    writel_relaxed(irqs->isactiver, base);
> +}
> +
> +static int gicv3_check_ap1r(unsigned int n, register_t apr)
> +{
> +    if ( !apr )
> +        return 0;
> +
> +    printk(XENLOG_ERR "GICv3: suspend aborted: ICC_AP1R%u_EL1=%#"
> +           PRIregister"\n", n, apr);
> +
> +    return -EBUSY;
> +}
> +
> +static int gicv3_check_active_priorities(register_t ctlr)
> +{
> +    unsigned int pribits = MASK_EXTR(ctlr, ICC_CTLR_EL1_PRIBITS_MASK) + 1;
> +    int ret;
> +
> +    /*
> +     * Xen enables physical Group 1 interrupts through ICC_IGRPEN1_EL1,
> +     * so only the physical Group 1 active-priority registers are relevant
> +     * here. Use ICC_CTLR_EL1.PRIbits for the physical CPU interface, not
> +     * ICH_VTR_EL2, which describes the virtual interface. ICC_AP1R1_EL1 is
> +     * only implemented with at least 6 physical priority bits, and
> +     * ICC_AP1R2_EL1/ICC_AP1R3_EL1 with at least 7.
> +     */
> +    switch ( pribits )
> +    {
> +    case 8:
> +    case 7:
> +        ret = gicv3_check_ap1r(3, READ_SYSREG(ICC_AP1R3_EL1));
> +        if ( ret )
> +            return ret;
> +        ret = gicv3_check_ap1r(2, READ_SYSREG(ICC_AP1R2_EL1));
> +        if ( ret )
> +            return ret;
> +        /* Fall through */
> +    case 6:
> +        ret = gicv3_check_ap1r(1, READ_SYSREG(ICC_AP1R1_EL1));
> +        if ( ret )
> +            return ret;
> +        /* Fall through */
> +    default:
> +        return gicv3_check_ap1r(0, READ_SYSREG(ICC_AP1R0_EL1));
> +    }
> +}
> +
> +static int gicv3_suspend(void)
> +{
> +    unsigned int i, nr_irqs;
> +    void __iomem *base;
> +    int ret;
> +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
> +
> +    /* Save GICC configuration */
> +    gicv3_ctx.cpu.ctlr     = READ_SYSREG(ICC_CTLR_EL1);
> +    gicv3_ctx.cpu.pmr      = READ_SYSREG(ICC_PMR_EL1);
> +    gicv3_ctx.cpu.bpr      = READ_SYSREG(ICC_BPR1_EL1);
> +    gicv3_ctx.cpu.sre_el2  = READ_SYSREG(ICC_SRE_EL2);
> +    gicv3_ctx.cpu.grpen    = READ_SYSREG(ICC_IGRPEN1_EL1);
> +
> +    gicv3_disable_interface();
> +
> +    ret = gicv3_check_active_priorities(gicv3_ctx.cpu.ctlr);
> +    if ( ret )
> +        goto out_enable_iface;
> +
> +    ret = gicv3_disable_redist();
> +    if ( ret )
> +        goto out_enable_iface;

I am wondering about the timeout case here. 

gicv3_disable_redist() has set ProcessorSleep to 1, but returns while
ChildrenAsleep is still 0.
This new error path then calls gicv3_enable_redist(), which clears
ProcessorSleep. 

Could that happen before ChildrenAsleep reaches 1? 
The GIC specification says that transition is UNPREDICTABLE. 
How should we handle the timeout?

> +
> +    /* Save GICR configuration */
> +    gicv3_redist_wait_for_rwp();
> +
> +    base = GICD_RDIST_BASE;
> +
> +    rdist->ctlr = readl_relaxed(base + GICR_CTLR);
> +
> +    rdist->propbase = readq_relaxed(base + GICR_PROPBASER);
> +    rdist->pendbase = readq_relaxed(base + GICR_PENDBASER);
> +
> +    base = GICD_RDIST_SGI_BASE;
> +
> +    /* Save priority on PPI and SGI interrupts */
> +    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
> +        rdist->ipriorityr[i] = readl_relaxed(base + GICR_IPRIORITYR0 + 4 * i);
> +
> +    rdist->isactiver = readl_relaxed(base + GICR_ISACTIVER0);
> +    rdist->isenabler = readl_relaxed(base + GICR_ISENABLER0);
> +    rdist->igroupr   = readl_relaxed(base + GICR_IGROUPR0);
> +    rdist->icfgr     = readl_relaxed(base + GICR_ICFGR1);
> +
> +    /* Save GICD configuration */
> +    gicv3_dist_wait_for_rwp();
> +    gicv3_ctx.dist.ctlr = readl_relaxed(GICD + GICD_CTLR);
> +
> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> +    {
> +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> +        gicv3_store_spi_irq_block(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
> +                                  false);
> +    }
> +
> +#ifdef CONFIG_GICV3_ESPI
> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> +        gicv3_store_spi_irq_block(gicv3_ctx.dist.espi_irqs + i, i, 32, true);
> +#endif
> +
> +    return 0;
> +
> + out_enable_iface:
> +    if ( gicv3_enable_redist() )
> +        panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
> +
> +    gicv3_hyp_enable(true);
> +    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
> +    isb();
> +
> +    return ret;
> +}
> +
> +static void gicv3_resume(void)
> +{
> +    int ret;
> +    unsigned int i, nr_irqs;
> +    uint32_t dist_ctlr;
> +    void __iomem *base;
> +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
> +
> +    dist_ctlr = gicv3_ctx.dist.ctlr & GICD_CTLR_ARE_NS;
> +
> +    /* Disable group forwarding while preserving affinity routing state. */
> +    writel_relaxed(dist_ctlr, GICD + GICD_CTLR);
> +    gicv3_dist_wait_for_rwp();
> +
> +    /*
> +     * IHI0069H.b 12.9.9 says changing GICD_ICFGR<n>.Int_config
> +     * while the interrupt is individually enabled is UNPREDICTABLE.
> +     * Disable SPIs first; 4.7.1 defines GICD_ICENABLER<n>, n > 0,
> +     * as the per-SPI disable mechanism.
> +     */
> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> +        gicv3_disable_spi_irq_block(i, false);
> +
> +#ifdef CONFIG_GICV3_ESPI
> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> +        gicv3_disable_spi_irq_block(i, true);
> +#endif
> +
> +    gicv3_dist_wait_for_rwp();
> +
> +    for ( i = NR_GIC_LOCAL_IRQS; i < gicv3_info.nr_lines; i += 32 )
> +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPR + (i / 32) * 4);
> +
> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> +    {
> +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
> +                                     false);
> +    }
> +
> +#ifdef CONFIG_GICV3_ESPI
> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> +    {
> +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPRnE + i * 4);
> +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.espi_irqs + i, i, 32,
> +                                     true);
> +    }
> +#endif
> +
> +    if ( dist_ctlr )
> +    {
> +        for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> +        {
> +            nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.irqs + i - 1, i,
> +                                          nr_irqs, false);
> +        }
> +
> +#ifdef CONFIG_GICV3_ESPI
> +        for ( i = 0; i < gic_number_espis() / 32; i++ )
> +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.espi_irqs + i, i,
> +                                          32, true);
> +#endif
> +    }
> +
> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.irqs + i - 1, i, false);
> +
> +#ifdef CONFIG_GICV3_ESPI
> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.espi_irqs + i, i, true);
> +#endif
> +
> +    writel_relaxed(gicv3_ctx.dist.ctlr, GICD + GICD_CTLR);
> +    gicv3_dist_wait_for_rwp();
> +
> +    ret = gicv3_lpi_init_rdist(GICD_RDIST_BASE);
> +    /*
> +     * If LPIs are already enabled, assume firmware or the still-powered
> +     * redistributor has valid PROPBASER/PENDBASER and skip reprogramming.
> +     * Return -EBUSY so callers can ignore this case.
> +     */
> +    if ( ret && ret != -ENODEV && ret != -EBUSY )
> +        panic("GICv3: Failed to re-initialize LPIs during resume\n");
> +    else if ( ret == -EBUSY ) /* extra checks, just to be sure */

Comment first letter capitalize:
s/extra/Extra/

Cheers
Bertrand

> +    {
> +        base = GICD_RDIST_BASE;
> +        if ( readq_relaxed(base + GICR_PROPBASER) != rdist->propbase ||
> +             readq_relaxed(base + GICR_PENDBASER) != rdist->pendbase )
> +            panic("GICv3: LPIs already enabled with unexpected PROPBASER/PENDBASER during resume\n");
> +    }
> +
> +    /* Restore GICR (Redistributor) configuration */
> +    if ( gicv3_enable_redist() )
> +        panic("GICv3: Failed to re-enable redistributor during resume\n");
> +
> +    base = GICD_RDIST_SGI_BASE;
> +
> +    writel_relaxed(GENMASK(31, 0), base + GICR_ICENABLER0);
> +    gicv3_redist_wait_for_rwp();
> +
> +    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
> +        writel_relaxed(rdist->ipriorityr[i], base + GICR_IPRIORITYR0 + i * 4);
> +
> +    writel_relaxed(rdist->isactiver, base + GICR_ISACTIVER0);
> +    writel_relaxed(rdist->igroupr,   base + GICR_IGROUPR0);
> +    writel_relaxed(rdist->icfgr,     base + GICR_ICFGR1);
> +
> +    gicv3_redist_wait_for_rwp();
> +
> +    writel_relaxed(rdist->isenabler, base + GICR_ISENABLER0);
> +    writel_relaxed(rdist->ctlr, GICD_RDIST_BASE + GICR_CTLR);
> +
> +    gicv3_redist_wait_for_rwp();
> +
> +    WRITE_SYSREG(gicv3_ctx.cpu.sre_el2, ICC_SRE_EL2);
> +    isb();
> +
> +    /* Restore CPU interface (System registers) */
> +    WRITE_SYSREG(gicv3_ctx.cpu.pmr,   ICC_PMR_EL1);
> +    WRITE_SYSREG(gicv3_ctx.cpu.bpr,   ICC_BPR1_EL1);
> +    WRITE_SYSREG(gicv3_ctx.cpu.ctlr,  ICC_CTLR_EL1);
> +    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
> +    isb();
> +
> +    gicv3_hyp_init();
> +}
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> /* Set up the GIC */
> static int __init gicv3_init(void)
> {
> @@ -2011,6 +2455,10 @@ static int __init gicv3_init(void)
> 
>     gicv3_hyp_init();
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    gicv3_alloc_context();
> +#endif
> +
> out:
>     spin_unlock(&gicv3.lock);
> 
> @@ -2050,6 +2498,10 @@ static const struct gic_hw_operations gicv3_ops = {
> #endif
>     .iomem_deny_access   = gicv3_iomem_deny_access,
>     .do_LPI              = gicv3_do_LPI,
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    .suspend             = gicv3_suspend,
> +    .resume              = gicv3_resume,
> +#endif
> };
> 
> static int __init gicv3_dt_preinit(struct dt_device_node *node, const void *data)
> diff --git a/xen/arch/arm/include/asm/arm64/sysregs.h b/xen/arch/arm/include/asm/arm64/sysregs.h
> index f3c11d871e..2261620316 100644
> --- a/xen/arch/arm/include/asm/arm64/sysregs.h
> +++ b/xen/arch/arm/include/asm/arm64/sysregs.h
> @@ -16,6 +16,11 @@
> #define ICC_SRE_EL1               S3_0_C12_C12_5
> #define ICC_IGRPEN1_EL1           S3_0_C12_C12_7
> 
> +#define ICC_AP1R0_EL1             S3_0_C12_C9_0
> +#define ICC_AP1R1_EL1             S3_0_C12_C9_1
> +#define ICC_AP1R2_EL1             S3_0_C12_C9_2
> +#define ICC_AP1R3_EL1             S3_0_C12_C9_3
> +
> #define ICH_VSEIR_EL2             S3_4_C12_C9_4
> #define ICC_SRE_EL2               S3_4_C12_C9_5
> #define ICH_HCR_EL2               S3_4_C12_C11_0
> diff --git a/xen/arch/arm/include/asm/gic_v3_defs.h b/xen/arch/arm/include/asm/gic_v3_defs.h
> index 3714cfeb7d..f741587322 100644
> --- a/xen/arch/arm/include/asm/gic_v3_defs.h
> +++ b/xen/arch/arm/include/asm/gic_v3_defs.h
> @@ -94,12 +94,15 @@
> #define GICD_TYPE_LPIS               (1U << 17)
> 
> #define GICD_CTLR_RWP                (1UL << 31)
> +#define GICD_CTLR_DS                 (1U << 6)
> #define GICD_CTLR_ARE_NS             (1U << 4)
> #define GICD_CTLR_ENABLE_G1A         (1U << 1)
> #define GICD_CTLR_ENABLE_G1          (1U << 0)
> #define GICD_IROUTER_SPI_MODE_ANY    (1UL << 31)
> 
> #define GICC_CTLR_EL1_EOImode_drop   (1U << 1)
> +#define ICC_CTLR_EL1_PRIBITS_SHIFT   8
> +#define ICC_CTLR_EL1_PRIBITS_MASK    (0x7U << ICC_CTLR_EL1_PRIBITS_SHIFT)
> 
> #define GICR_WAKER_ProcessorSleep    (1U << 1)
> #define GICR_WAKER_ChildrenAsleep    (1U << 2)
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support
  2026-08-27 14:31 ` [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support Mykola Kvach
@ 2026-09-23 15:36   ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-23 15:36 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> Handle system suspend/resume for GICv3 with an ITS present so LPIs keep
> working after firmware powers the GIC down.
> 
> Save and restore the ITS CTLR, CBASER and BASER registers. On resume,
> re-establish the collection mapping only when the collection is held in
> the ITS itself. Memory-backed collections are restored through the
> restored GITS_BASER tables and must not be remapped unconditionally.
> 
> Add list_for_each_entry_continue_reverse() in list.h for the ITS suspend
> error path that needs to roll back partially saved state.
> 
> Based on Linux commit dba0bc7b76dc:
> "irqchip/gic-v3-its: Add ability to save/restore ITS state".
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>

Reviewed-by: Bertrand Marquis <bertrand.marquis@arm.com> # Arm code

Changes to list.h will need a review from the Others group but they look
good to me.


Cheers
Bertrand

> ---
> Changes in V10:
> - Replay MAPC on resume only for collections held in the ITS itself, as
>  indicated by GITS_TYPER.HCC. Memory-backed collections are restored
>  through GITS_BASER and are no longer remapped unconditionally.
> - Make the current Xen col_id == cpu assumption explicit in the ITS
>  resume path.
> - Use "unpredictable" instead of "undefined" in the CBASER/BASER restore
>  comment.
> 
> Changes in V9:
> - fix the ITS suspend/resume coding-style nits;
> - preserve the saved GITS_CTLR state while masking the read-only
>  QUIESCENT bit.
> 
> Changes in V8:
> - Reword the CBASER/CWRITER comment to match Xen and drop the stale Linux
>  cmd_write reference.
> - Clarify the list_for_each_entry_continue_reverse() comment.
> - Factor out per-ITS helpers for collection setup and resume.
> - Restore each ITS and re-establish its collection mapping in the same
>  loop, so a failed ITS resume is not followed by MAPC/SYNC on that
>  un-restored instance.
> - panic in case when resume of an ITS failed
> - cleanup baser cache during suspend
> ---
> xen/arch/arm/gic-v3-its.c             | 146 ++++++++++++++++++++++++--
> xen/arch/arm/gic-v3.c                 |  11 +-
> xen/arch/arm/include/asm/gic_v3_its.h |  28 +++++
> xen/include/xen/list.h                |  14 +++
> 4 files changed, 189 insertions(+), 10 deletions(-)
> 
> diff --git a/xen/arch/arm/gic-v3-its.c b/xen/arch/arm/gic-v3-its.c
> index 7560d46c6d..dd53209865 100644
> --- a/xen/arch/arm/gic-v3-its.c
> +++ b/xen/arch/arm/gic-v3-its.c
> @@ -335,6 +335,22 @@ static int its_send_cmd_inv(struct host_its *its,
>     return its_send_command(its, cmd);
> }
> 
> +static int gicv3_its_setup_collection_single(struct host_its *its,
> +                                             unsigned int cpu)
> +{
> +    int ret;
> +
> +    ret = its_send_cmd_mapc(its, cpu, cpu);
> +    if ( ret )
> +        return ret;
> +
> +    ret = its_send_cmd_sync(its, cpu);
> +    if ( ret )
> +        return ret;
> +
> +    return gicv3_its_wait_commands(its);
> +}
> +
> /* Set up the (1:1) collection mapping for the given host CPU. */
> int gicv3_its_setup_collection(unsigned int cpu)
> {
> @@ -343,15 +359,7 @@ int gicv3_its_setup_collection(unsigned int cpu)
> 
>     list_for_each_entry(its, &host_its_list, entry)
>     {
> -        ret = its_send_cmd_mapc(its, cpu, cpu);
> -        if ( ret )
> -            return ret;
> -
> -        ret = its_send_cmd_sync(its, cpu);
> -        if ( ret )
> -            return ret;
> -
> -        ret = gicv3_its_wait_commands(its);
> +        ret = gicv3_its_setup_collection_single(its, cpu);
>         if ( ret )
>             return ret;
>     }
> @@ -1211,6 +1219,126 @@ int gicv3_its_init(void)
>     return 0;
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +int gicv3_its_suspend(void)
> +{
> +    struct host_its *its;
> +    int ret;
> +
> +    list_for_each_entry( its, &host_its_list, entry )
> +    {
> +        unsigned int i;
> +        void __iomem *base = its->its_base;
> +
> +        /*
> +         * By the time Xen reaches gic_suspend(), every domain is already in
> +         * SHUTDOWN_suspend, so ITS-targeting interrupt sources are expected
> +         * to have been quiesced by the owning OS before SYSTEM_SUSPEND.
> +         */
> +        /* Preserve saved GITS_CTLR state, excluding read-only QUIESCENT. */
> +        its->suspend_ctx.ctlr = readl_relaxed(base + GITS_CTLR) &
> +                                ~GITS_CTLR_QUIESCENT;
> +        ret = gicv3_disable_its(its);
> +        if ( ret )
> +        {
> +            writel_relaxed(its->suspend_ctx.ctlr, base + GITS_CTLR);
> +            goto err;
> +        }
> +
> +        its->suspend_ctx.cbaser = readq_relaxed(base + GITS_CBASER);
> +
> +        for ( i = 0; i < GITS_BASER_NR_REGS; i++ )
> +        {
> +            uint64_t baser = readq_relaxed(base + GITS_BASER0 + i * 8);
> +
> +            its->suspend_ctx.baser[i] = 0;
> +
> +            if ( !(baser & GITS_VALID_BIT) )
> +                continue;
> +
> +            its->suspend_ctx.baser[i] = baser;
> +        }
> +    }
> +
> +    return 0;
> +
> + err:
> +    list_for_each_entry_continue_reverse( its, &host_its_list, entry )
> +        writel_relaxed(its->suspend_ctx.ctlr, its->its_base + GITS_CTLR);
> +
> +    return ret;
> +}
> +
> +static int gicv3_its_resume_single(struct host_its *its, unsigned int cpu)
> +{
> +    void __iomem *base = its->its_base;
> +    unsigned int i;
> +    int ret;
> +    uint64_t typer;
> +    unsigned int col_id = cpu; /* Xen currently uses col_id == cpu. */
> +
> +    /*
> +     * Make sure that the ITS is disabled. If it fails to quiesce,
> +     * don't restore it since writing to CBASER or BASER<n>
> +     * registers is unpredictable according to the GIC v3 ITS
> +     * Specification.
> +     */
> +    WARN_ON(readl_relaxed(base + GITS_CTLR) & GITS_CTLR_ENABLE);
> +    ret = gicv3_disable_its(its);
> +    if ( ret )
> +        return ret;
> +
> +    writeq_relaxed(its->suspend_ctx.cbaser, base + GITS_CBASER);
> +
> +    /*
> +     * Writing CBASER resets CREADR to 0, so reset CWRITER to
> +     * keep the command queue pointers aligned.
> +     */
> +    writeq_relaxed(0, base + GITS_CWRITER);
> +
> +    /* Restore GITS_BASER from the value cache. */
> +    for ( i = 0; i < GITS_BASER_NR_REGS; i++ )
> +    {
> +        uint64_t baser = its->suspend_ctx.baser[i];
> +
> +        if ( !(baser & GITS_VALID_BIT) )
> +            continue;
> +
> +        writeq_relaxed(baser, base + GITS_BASER0 + i * 8);
> +    }
> +
> +    writel_relaxed(its->suspend_ctx.ctlr, base + GITS_CTLR);
> +
> +    typer = readq_relaxed(base + GITS_TYPER);
> +
> +    /*
> +     * Only collections with IDs below HCC are held in the ITS itself
> +     * and lose their state across an ITS reset/power loss. Memory-backed
> +     * collections are restored by restoring GITS_BASER and must not be
> +     * remapped here.
> +     */
> +    if ( col_id < GITS_TYPER_HCC(typer) )
> +        return gicv3_its_setup_collection_single(its, cpu);
> +
> +    return 0;
> +}
> +
> +void gicv3_its_resume(void)
> +{
> +    struct host_its *its;
> +    unsigned int cpu = smp_processor_id();
> +    int ret;
> +
> +    list_for_each_entry( its, &host_its_list, entry )
> +    {
> +        ret = gicv3_its_resume_single(its, cpu);
> +        if ( ret )
> +            panic("GICv3: ITS@%"PRIpaddr": failed to restore during resume: %d\n",
> +                   its->addr, ret);
> +    }
> +}
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> 
> /*
>  * Local variables:
> diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
> index 038bf41142..6b025c023d 100644
> --- a/xen/arch/arm/gic-v3.c
> +++ b/xen/arch/arm/gic-v3.c
> @@ -2186,10 +2186,14 @@ static int gicv3_suspend(void)
>     if ( ret )
>         goto out_enable_iface;
> 
> -    ret = gicv3_disable_redist();
> +    ret = gicv3_its_suspend();
>     if ( ret )
>         goto out_enable_iface;
> 
> +    ret = gicv3_disable_redist();
> +    if ( ret )
> +        goto out_its_resume;
> +
>     /* Save GICR configuration */
>     gicv3_redist_wait_for_rwp();
> 
> @@ -2229,6 +2233,9 @@ static int gicv3_suspend(void)
> 
>     return 0;
> 
> + out_its_resume:
> +    gicv3_its_resume();
> +
>  out_enable_iface:
>     if ( gicv3_enable_redist() )
>         panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
> @@ -2355,6 +2362,8 @@ static void gicv3_resume(void)
> 
>     gicv3_redist_wait_for_rwp();
> 
> +    gicv3_its_resume();
> +
>     WRITE_SYSREG(gicv3_ctx.cpu.sre_el2, ICC_SRE_EL2);
>     isb();
> 
> diff --git a/xen/arch/arm/include/asm/gic_v3_its.h b/xen/arch/arm/include/asm/gic_v3_its.h
> index fc5a84892c..0f8cb16e41 100644
> --- a/xen/arch/arm/include/asm/gic_v3_its.h
> +++ b/xen/arch/arm/include/asm/gic_v3_its.h
> @@ -43,6 +43,11 @@
> #define GITS_CTLR_QUIESCENT             BIT(31, UL)
> #define GITS_CTLR_ENABLE                BIT(0, UL)
> 
> +#define GITS_TYPER_HCC_SHIFT            24
> +#define GITS_TYPER_HCC_MASK             0xffUL
> +#define GITS_TYPER_HCC(r)               (((r) >> GITS_TYPER_HCC_SHIFT) & \
> +                                                 GITS_TYPER_HCC_MASK)
> +
> #define GITS_TYPER_PTA                  BIT(19, UL)
> #define GITS_TYPER_DEVIDS_SHIFT         13
> #define GITS_TYPER_DEVIDS_MASK          (0x1fUL << GITS_TYPER_DEVIDS_SHIFT)
> @@ -129,6 +134,13 @@ struct host_its {
>     spinlock_t cmd_lock;
>     void *cmd_buf;
>     unsigned int flags;
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    struct suspend_ctx {
> +        uint32_t ctlr;
> +        uint64_t cbaser;
> +        uint64_t baser[GITS_BASER_NR_REGS];
> +    } suspend_ctx;
> +#endif
> };
> 
> /* Map a collection for this host CPU to each host ITS. */
> @@ -204,6 +216,11 @@ uint64_t gicv3_its_get_cacheability(void);
> uint64_t gicv3_its_get_shareability(void);
> unsigned int gicv3_its_get_memflags(void);
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +int gicv3_its_suspend(void);
> +void gicv3_its_resume(void);
> +#endif
> +
> #else
> 
> #ifdef CONFIG_ACPI
> @@ -271,6 +288,17 @@ static inline int gicv3_its_make_hwdom_dt_nodes(const struct domain *d,
>     return 0;
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +static inline int gicv3_its_suspend(void)
> +{
> +    return 0;
> +}
> +
> +static inline void gicv3_its_resume(void)
> +{
> +}
> +#endif
> +
> #endif /* CONFIG_HAS_ITS */
> 
> #endif
> diff --git a/xen/include/xen/list.h b/xen/include/xen/list.h
> index 98d8482dab..2aab274157 100644
> --- a/xen/include/xen/list.h
> +++ b/xen/include/xen/list.h
> @@ -535,6 +535,20 @@ static inline void list_splice_init(struct list_head *list,
>          &(pos)->member != (head);                                        \
>          (pos) = list_entry((pos)->member.next, typeof(*(pos)), member))
> 
> +/**
> + * list_for_each_entry_continue_reverse - iterate backwards from the given point
> + * @pos:    the type * to use as a loop cursor.
> + * @head:   the head for your list.
> + * @member: the name of the list_head within the struct.
> + *
> + * Iterate over list of given type backwards, starting from the element previous
> + * to the current one in list order.
> + */
> +#define list_for_each_entry_continue_reverse(pos, head, member)           \
> +    for ((pos) = list_entry((pos)->member.prev, typeof(*(pos)), member);  \
> +         &(pos)->member != (head);                                        \
> +         (pos) = list_entry((pos)->member.prev, typeof(*(pos)), member))
> +
> /**
>  * list_for_each_entry_from - iterate over list of given type from the
>  *                            current point
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions
  2026-09-23 15:27   ` Bertrand Marquis
@ 2026-09-24 22:23     ` Mykola Kvach
  2026-09-28  7:36       ` Bertrand Marquis
  0 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-09-24 22:23 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Bertrand,

Thank you for the review.

On Wed, Sep 23, 2026 at 6:28 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> Sorry for the delay to review this serie.
>
> > On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > From: Mirela Simonovic <mirela.simonovic@aggios.com>
> >
> > System suspend may lead to a state where GIC would be powered down.
> > Therefore, Xen should save/restore the context of GIC on suspend/resume.
> >
> > Note that the context consists of states of registers which are
> > controlled by the hypervisor. Other GIC registers which are accessible
> > by guests are saved/restored on context switch.
> >
> > Transient physical SGI pending state (GICD_CPENDSGIRn/GICD_SPENDSGIRn)
> > is intentionally excluded. CPU-interface active-priority state is also
> > not restored across suspend/resume. Xen reaches the final suspend path
> > at a quiescent point, so there is no active-priority execution context
> > to replay after resume. Enforce this with a runtime check after
> > disabling the CPU interface: if any implemented GICC_APRn word is still
> > non-zero, restore GICC_CTLR and abort suspend with -EBUSY.
>
> You mention SGI pending state but you do not say what would happen for PPI/SPI
> pending state, and the patch does not look at or save/restore GICD_ISPENDR.
>
> Can you clarify what is expected for those?

This patch was originally based on the Linux GICv2 suspend/resume
code, which also does not save PPI/SPI pending state. The comment
above gic_dist_restore() explains that level interrupts still
asserted after resume will be handled, while edge events during
suspend need to be handled by the platform-specific wakeup
mechanism.

Saving GICD_ISPENDR would preserve the pending state at the time of
each read. However, an interrupt could become pending after that
read and before the GIC loses power. Saving pending state alone
therefore does not cover the whole suspend transition.

The assumption here is that device drivers have stopped normal I/O
and quiesced non-wakeup interrupt sources before Xen suspends the
GIC. Earlier events must already have been handled, or their state
must be preserved outside the GIC. Only configured wakeup sources
are expected to generate new events at this point.

We rely on the platform wakeup mechanism throughout suspend entry
and sleep. If the GIC loses power, this mechanism must capture
wakeup events outside the GIC and keep them observable after
resume.

Not saving pending state depends on these assumptions. The race
after a register read does not, by itself, justify losing an event
that is already pending.
---

While checking the pending-state question, I also noticed a related
issue with disabling the Distributor during suspend entry.

My earlier reasoning relied on the Distributor being powered down
during system suspend. Not every platform is required to follow
BSA.

PSCI requires us to save the state that could be lost. Section 6.8
explicitly discusses Distributor power-down as a feature of some
systems. This does not require Xen to disable its interrupt group
before calling SYSTEM_SUSPEND.

The GICv2 pseudocode in section 3.7.2 shows that irq_wake and
fiq_wake depend on the Distributor group enables, but not on the
CPU interface group enables. Clearing the group enable in
GICD_CTLR can therefore block that group's wakeup path on a
platform that uses these signals.

I therefore propose leaving the Distributor group enabled during
suspend entry, while still saving its configuration in case it
loses power. The platform would handle any further shutdown and
the required wakeup configuration. The Distributor would still
be disabled while restoring its registers on resume.

Linux also leaves the GICv2 CPU interface enabled before the PSCI
call. Our early gicv2_cpu_disable() is another difference: it can
hide that group's interrupts from TF-A's early ISR_EL1 check.
I propose leaving that shutdown to firmware as well.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF
  2026-09-23 15:34   ` Bertrand Marquis
@ 2026-09-24 23:22     ` Mykola Kvach
  2026-09-28  7:37       ` Bertrand Marquis
  0 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-09-24 23:22 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Bertrand,

Thank you for the review.

On Wed, Sep 23, 2026 at 6:36 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > PSCI does not guarantee that a GICv3 redistributor is powered down across
> > CPU_OFF -> CPU_ON.
> >
> > DEN0022F.b says CPU_OFF powers down the calling core (5.5) and CPU_ON
> > brings the core back with a defined initial CPU state (5.6, 6.4).
> > However, PSCI leaves interrupt migration and GIC re-initialization to the
> > supervisory software/firmware stack: the caller must migrate interrupts
> > away before CPU_OFF (5.5.2), and the execution context that is lost in a
> > powerdown state must be saved and restored by software (6.8). PSCI also
> > calls out GIC management explicitly in 6.8, including retargeting SPIs,
> > preventing PPIs/SGIs from targeting a powered down CPU, and reinitializing
> > the CPU interface after CPU_ON.
> >
> > This matches the GIC architecture. IHI0069H.b Chapter 11.1 requires the PE
> > and CPU interface to share a power domain, but explicitly allows the
> > associated redistributor, distributor, and ITS to remain powered while the
> > PE and CPU interface are off. All other GIC power-management behavior is
> > IMPLEMENTATION DEFINED. DEN0050D Chapter 4.2, "Generic Interrupt
> > Controller (GIC)", says the GICv3 redistributor may live either in the AP
> > core power domain or in a relatively always-on parent domain. So after
> > CPU_OFF -> CPU_ON a secondary CPU can legitimately come back to a live
> > redistributor with GICR_CTLR.EnableLPIs still set.
> >
> > Handle that case in the LPI setup path instead of assuming a fully reset
> > redistributor.
> >
> > The LPI path needs special care because the GIC spec makes redistributor
> > LPI state sticky and partially implementation defined. IHI0069H.b 5.1.1
> > and 5.1.2 say that changing GICR_PROPBASER or GICR_PENDBASER while
> > GICR_CTLR.EnableLPIs == 1 is UNPREDICTABLE. After clearing EnableLPIs,
> > software must wait for GICR_CTLR.RWP == 0 before touching the pending
> > table. The architecture also permits implementations where, once
> > EnableLPIs has been set, clearing it again is not guaranteed to work.
> > Where an ITS is present, the spec strongly recommends moving LPIs to
> > another redistributor before clearing EnableLPIs.
> >
> > Because of that, treat a retained EnableLPIs state as valid when the
> > redistributor still points at Xen's expected PROPBASER/PENDBASER tables.
> > Only try to clear EnableLPIs when the retained configuration does not
> > match Xen's state, and wait for RWP before reprogramming the tables.
> >
> > This is also consistent with platform firmware reality: PSCI and the GIC
> > architecture allow platform-specific redistributor power handling, and not
> > all platform firmware implementations force a full redistributor power-off
> > through implementation-defined controls during CPU_OFF. Xen therefore needs
> > to tolerate retained redistributor state on secondary CPU bring-up.
> >
> > Keep gicv3_populate_rdist() resident as well, because gicv3_cpu_init()
> > reuses it on secondary CPU bring-up after init.
> >
> > Tested using Xen's non-boot CPU disable/enable path on Arm
> > FVP_Base_RevC-2xAEMvA, both with and without:
> > -C gic_distributor.allow-LPIEN-clear=1
> > -C gic_distributor.GICR-clear-enable-supported=1
> > and on Orange Pi 5.
> >
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> > ---
> > Changes in v10:
> > - Drop unrelated gicv3_populate_rdist() printk() format cleanups to keep
> >  the patch focused on retained redistributor LPI state.
> >
> > Changes in v9:
> > - move gicv3_do_wait_for_rwp prototype from its related header to gic.h
> > - drop __init from gicv3_populate_rdist(), which is reused on secondary
> >  CPU bring-up after boot
> > - changed print format for smp_processor_id in gicv3_populate_rdist func
> > - cosmetic changes
> > ---
> > xen/arch/arm/gic-v3-lpi.c      | 77 +++++++++++++++++++++++++++++++++-
> > xen/arch/arm/gic-v3.c          | 15 ++++---
> > xen/arch/arm/include/asm/gic.h |  4 ++
> > 3 files changed, 90 insertions(+), 6 deletions(-)
> >
> > diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
> > index 9ee338edc2..847da26ff7 100644
> > --- a/xen/arch/arm/gic-v3-lpi.c
> > +++ b/xen/arch/arm/gic-v3-lpi.c
> > @@ -81,6 +81,13 @@ static DEFINE_PER_CPU(struct lpi_redist_data, lpi_redist);
> > #define MAX_NR_HOST_LPIS   (lpi_data.max_host_lpi_ids - LPI_OFFSET)
> > #define HOST_LPIS_PER_PAGE      (PAGE_SIZE / sizeof(union host_lpi))
> >
> > +#define GICR_PROPBASER_XEN_MASK  GENMASK_ULL(51, 12)
> > +/*
> > + * For retained redistributor state, match the pending table by address only.
> > + * Attribute bits such as PTZ may not read back with the programmed value.
> > + */
> > +#define GICR_PENDBASER_XEN_MASK  GENMASK_ULL(51, 16)
> > +
> > static union host_lpi *gic_get_host_lpi(uint32_t plpi)
> > {
> >     union host_lpi *block;
> > @@ -296,6 +303,60 @@ static int gicv3_lpi_set_pendtable(void __iomem *rdist_base)
> >     return 0;
> > }
> >
> > +static uint64_t gicv3_lpi_expected_proptable(void)
> > +{
> > +    return virt_to_maddr(lpi_data.lpi_property);
> > +}
> > +
> > +static uint64_t gicv3_lpi_expected_pendtable(void)
> > +{
> > +    return virt_to_maddr(this_cpu(lpi_redist).pending_table);
> > +}
> > +
> > +static bool gicv3_lpi_tables_match(void __iomem *rdist_base)
> > +{
> > +    uint64_t propbase, pendbase;
> > +
> > +    if ( !lpi_data.lpi_property || !this_cpu(lpi_redist).pending_table )
> > +        return false;
> > +
> > +    propbase = readq_relaxed(rdist_base + GICR_PROPBASER);
> > +    pendbase = readq_relaxed(rdist_base + GICR_PENDBASER);
> > +
> > +    return ((propbase & GICR_PROPBASER_XEN_MASK) ==
> > +            (gicv3_lpi_expected_proptable() & GICR_PROPBASER_XEN_MASK)) &&
> > +           ((pendbase & GICR_PENDBASER_XEN_MASK) ==
> > +            (gicv3_lpi_expected_pendtable() & GICR_PENDBASER_XEN_MASK));
> > +}
> > +
> > +static int gicv3_lpi_disable_lpis(void __iomem *rdist_base)
> > +{
> > +    uint32_t reg = readl_relaxed(rdist_base + GICR_CTLR);
> > +    int ret;
> > +
> > +    if ( !(reg & GICR_CTLR_ENABLE_LPIS) )
> > +        return 0;
> > +
> > +    writel_relaxed(reg & ~GICR_CTLR_ENABLE_LPIS, rdist_base + GICR_CTLR);
> > +
> > +    /*
> > +     * The spec only guarantees programmability when we have observed the bit
> > +     * cleared. Where clearing is supported, RWP must reach 0 before touching
> > +     * PROPBASER/PENDBASER again.
> > +     */
> > +    wmb();
> > +
> > +    ret = gicv3_do_wait_for_rwp(rdist_base, GICR_CTLR_RWP);
> > +    if ( ret )
> > +        return ret;
> > +
> > +    reg = readl_relaxed(rdist_base + GICR_CTLR);
> > +    if ( reg & GICR_CTLR_ENABLE_LPIS )
> > +        return -EBUSY;
> > +
> > +    return 0;
> > +}
> > +
> > /*
> >  * Tell a redistributor about the (shared) property table, allocating one
> >  * if not already done.
> > @@ -374,7 +435,21 @@ int gicv3_lpi_init_rdist(void __iomem * rdist_base)
> >     /* Make sure LPIs are disabled before setting up the tables. */
> >     reg = readl_relaxed(rdist_base + GICR_CTLR);
> >     if ( reg & GICR_CTLR_ENABLE_LPIS )
> > -        return -EBUSY;
> > +    {
> > +        if ( gicv3_lpi_tables_match(rdist_base) )
> > +            return -EBUSY;
>
> I am wondering if there is a corner case when a CPU is unplugged and then
> plugged back in. free_percpu_area() eventually frees the per-CPU area
> containing lpi_redist.pending_table, but not the table itself. On the next
> cpu_up(), gicv3_lpi_allocate_pendtable() allocates a new table, and I cannot
> find where the old one is freed.
>
> If the redistributor kept EnableLPIs=1 and GICR_PENDBASER pointing to the
> old table, wouldn't gicv3_lpi_tables_match() fail?
>
> What will happen if EnableLPIs cannot be cleared ? (i think this is something
> possible in the hardware).

The sequence you describe would need to be handled when adding
CPU hotplug support on Arm. There is currently no runtime caller
for such an offline/online cycle outside system suspend/resume.

During system suspend, the common code preserves the per-CPU area.
The following GICv3 suspend/resume patch also skips pending-table
allocation on resume, so the existing table and pointer are
reused.

The pending-table lifetime across normal CPU hotplug should be
handled by the CPU hotplug series.
---

While checking CPU bring-up failures during resume, I found a
separate issue in the common cleanup code. Both CPU_UP_CANCELED
and CPU_RESUME_FAILED can call free_percpu_area() for the same CPU
during resume. The release metadata, including the rcu_head, is
itself stored in that CPU's per-CPU area.

If the first release has completed, another call would access
invalid per-CPU state. Otherwise, it can queue the same rcu_head
again. The timer and CPU-pool callbacks on CPU_RESUME_FAILED also
still need the per-CPU area.

On x86, park_offline_cpus prevents this particular release path.

I plan to address this in a separate preparatory patch in this
series, keeping the per-CPU area available until the final
CPU_RESUME_FAILED cleanup and releasing it only once.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions
  2026-09-23 15:35   ` Bertrand Marquis
@ 2026-09-25  0:03     ` Mykola Kvach
  2026-09-28  7:39       ` Bertrand Marquis
  0 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-09-25  0:03 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Bertrand,

Thank you for the review.

On Wed, Sep 23, 2026 at 6:58 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > System suspend may lead to a state where GIC would be powered down.
> > Therefore, Xen should save/restore the context of GIC on suspend/resume.
> >
> > Note that the context consists of states of registers which are
> > controlled by the hypervisor. Other GIC registers which are accessible
> > by guests are saved/restored on context switch.
> >
> > Before continuing suspend, also verify that the physical CPU interface
> > has no Group 1 active-priority state left. Use ICC_CTLR_EL1.PRIbits to
> > decide which ICC_AP1R<n>_EL1 registers are implemented, so Xen does not
> > read an unimplemented AP1R register.
> >
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> > ---
> > Changes in V10:
> > - abort suspend when the physical Group 1 active-priority state is still
> >  present, deriving accessible ICC_AP1R<n>_EL1 registers from
> >  ICC_CTLR_EL1.PRIbits;
> > - re-enable the redistributor before restoring CPU and virtual interface
> >  state on the suspend abort path;
> > - panic if the redistributor cannot be re-enabled on the suspend abort path;
> > - avoid saving/restoring reserved GICD_IPRIORITYR and GICD_IROUTER entries
> >  for a partially populated last SPI block;
> > - disable Distributor group forwarding while preserving affinity routing
> >  state before restoring Distributor configuration;
> > - disable SPI/eSPI forwarding and wait for RWP before restoring
> >  GICD_ICFGR<n>.Int_config.
> >
> > Changes in V9:
> > - fix the suspend-context comment typo and split dist_ctx declarations;
> > - restore ICC_IGRPEN1_EL1 on the suspend error path;
> > - re-initialize GICD_IGROUPRnE during resume;
> > - restore GICD_IROUTER only after re-enabling ARE_NS during resume.
> >
> > Changes in V8:
> > - use right rdist base for prop/pend baser and ctrl
> >
> > Changes in V7:
> > - restore LPI regs on resume
> > - add timeout during redist disabling
> > - squash with suspend/resume handling for GICv3 eSPI registers
> > - drop ITS guard paths so suspend/resume always runs; switch missing ctx
> >  allocation to panic
> > - trim TODO comments; narrow redistributor storage to PPI icfgr
> > - keep distributor context allocation even without ITS; adjust resume
> >  to use GENMASK(31, 0) for clearing enables
> > - drop storage of the SGI configuration register, as SGIs are always
> >  edge-triggered
> > ---
> > xen/arch/arm/gic-v3-lpi.c                |   3 +
> > xen/arch/arm/gic-v3.c                    | 458 ++++++++++++++++++++++-
> > xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
> > xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
> > 4 files changed, 466 insertions(+), 3 deletions(-)
> >
> > diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
> > index 847da26ff7..a63c8c4979 100644
> > --- a/xen/arch/arm/gic-v3-lpi.c
> > +++ b/xen/arch/arm/gic-v3-lpi.c
> > @@ -467,6 +467,9 @@ static int cpu_callback(struct notifier_block *nfb, unsigned long action,
> >     switch ( action )
> >     {
> >     case CPU_UP_PREPARE:
> > +        if ( system_state == SYS_STATE_resume )
> > +            break;
> > +
> >         rc = gicv3_lpi_allocate_pendtable(cpu);
> >         if ( rc )
> >             printk(XENLOG_ERR "Unable to allocate the pendtable for CPU%lu\n",
> > diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
> > index b16888ad84..038bf41142 100644
> > --- a/xen/arch/arm/gic-v3.c
> > +++ b/xen/arch/arm/gic-v3.c
> > @@ -1078,12 +1078,12 @@ out:
> >     return res;
> > }
> >
> > -static void gicv3_hyp_disable(void)
> > +static void gicv3_hyp_enable(bool enable)
> > {
> >     register_t hcr;
> >
> >     hcr = READ_SYSREG(ICH_HCR_EL2);
> > -    hcr &= ~GICH_HCR_EN;
> > +    hcr = enable ? (hcr | GICH_HCR_EN) : (hcr & ~GICH_HCR_EN);
> >     WRITE_SYSREG(hcr, ICH_HCR_EL2);
> >     isb();
> > }
> > @@ -1190,7 +1190,7 @@ static void gicv3_disable_interface(void)
> >     spin_lock(&gicv3.lock);
> >
> >     gicv3_cpu_disable();
> > -    gicv3_hyp_disable();
> > +    gicv3_hyp_enable(false);
> >
> >     spin_unlock(&gicv3.lock);
> > }
> > @@ -1926,6 +1926,450 @@ static bool gic_dist_supports_lpis(void)
> >     return (readl_relaxed(GICD + GICD_TYPER) & GICD_TYPE_LPIS);
> > }
> >
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +
> > +/* This struct represents a block of 32 IRQs */
> > +struct dist_irq_block {
> > +    uint32_t icfgr[2];
> > +    uint32_t ipriorityr[8];
> > +    uint64_t irouter[32];
> > +    uint32_t isactiver;
> > +    uint32_t isenabler;
> > +};
> > +
> > +struct redist_ctx {
> > +    uint32_t ctlr;
> > +    uint32_t icfgr; /* only PPIs stored */
>
> Can you capitalize first comment letter ?
> s/only/Only/
>
> > +    uint32_t igroupr;
> > +    uint32_t ipriorityr[8];
> > +    uint32_t isactiver;
> > +    uint32_t isenabler;
> > +
> > +    uint64_t pendbase;
> > +    uint64_t propbase;
> > +};
> > +
> > +/* GICv3 registers to be saved/restored on system suspend/resume */
> > +struct gicv3_ctx {
> > +    struct dist_ctx {
> > +        uint32_t ctlr;
> > +        struct dist_irq_block *irqs;
> > +        struct dist_irq_block *espi_irqs;
> > +    } dist;
> > +
> > +    /* have only one rdist structure for last running CPU during suspend */
>
> Same here
> s/have/Have/
>
> > +    struct redist_ctx rdist;
> > +
> > +    struct cpu_ctx {
> > +        uint32_t ctlr;
> > +        uint32_t pmr;
> > +        uint32_t bpr;
> > +        uint32_t sre_el2;
> > +        uint32_t grpen;
> > +    } cpu;
> > +};
> > +
> > +static struct gicv3_ctx gicv3_ctx;
> > +
> > +static void __init gicv3_alloc_context(void)
> > +{
> > +    uint32_t blocks = DIV_ROUND_UP(gicv3_info.nr_lines, 32);
> > +
> > +    /* The spec allows for systems without any SPIs */
> > +    if ( blocks > 1 )
> > +    {
> > +        gicv3_ctx.dist.irqs = xzalloc_array(struct dist_irq_block, blocks - 1);
> > +        if ( !gicv3_ctx.dist.irqs )
> > +            panic("Failed to allocate memory for GICv3 suspend context\n");
> > +    }
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +    if ( !gic_number_espis() )
> > +        return;
> > +
> > +    blocks = gic_number_espis() / 32;
> > +    gicv3_ctx.dist.espi_irqs = xzalloc_array(struct dist_irq_block, blocks);
> > +    if ( !gicv3_ctx.dist.espi_irqs )
> > +        panic("Failed to allocate memory for GICv3 eSPI suspend context\n");
> > +#endif
> > +}
> > +
> > +static int gicv3_disable_redist(void)
> > +{
> > +    void __iomem *waker = GICD_RDIST_BASE + GICR_WAKER;
> > +    s_time_t deadline;
> > +
> > +    /*
> > +     * Avoid infinite loop if Non-secure does not have access to GICR_WAKER.
> > +     * See Arm IHI 0069H.b, 12.11.42 GICR_WAKER:
> > +     *     When GICD_CTLR.DS == 0 and an access is Non-secure accesses to this
> > +     *     register are RAZ/WI.
> > +     */
> > +    if ( !(readl_relaxed(GICD + GICD_CTLR) & GICD_CTLR_DS) )
> > +        return 0;
> > +
> > +    deadline = NOW() + MILLISECS(1000);
> > +
> > +    writel_relaxed(readl_relaxed(waker) | GICR_WAKER_ProcessorSleep, waker);
> > +    while ( (readl_relaxed(waker) & GICR_WAKER_ChildrenAsleep) == 0 )
> > +    {
> > +        if ( NOW() > deadline )
> > +        {
> > +            printk("GICv3: Timeout waiting for redistributor to sleep\n");
> > +            return -ETIMEDOUT;
> > +        }
> > +        cpu_relax();
> > +        udelay(10);
> > +    }
> > +
> > +    return 0;
> > +}
> > +
> > +#define GET_SPI_REG_OFFSET(name, is_espi) \
> > +    ((is_espi) ? GICD_##name##nE : GICD_##name)
> > +
> > +static void gicv3_store_spi_irq_block(struct dist_irq_block *irqs,
> > +                                      unsigned int i, unsigned int nr_irqs,
> > +                                      bool is_espi)
> > +{
> > +    void __iomem *base;
> > +    unsigned int irq, nr_priority_regs;
> > +
> > +    ASSERT(nr_irqs && nr_irqs <= 32);
> > +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
> > +    irqs->icfgr[0] = readl_relaxed(base);
> > +    irqs->icfgr[1] = readl_relaxed(base + 4);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
> > +    base += i * sizeof(irqs->ipriorityr);
> > +    for ( irq = 0; irq < nr_priority_regs; irq++ )
> > +        irqs->ipriorityr[irq] = readl_relaxed(base + 4 * irq);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
> > +    base += i * sizeof(irqs->irouter);
> > +    for ( irq = 0; irq < nr_irqs; irq++ )
> > +        irqs->irouter[irq] = readq_relaxed_non_atomic(base + 8 * irq);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
> > +    base += i * sizeof(irqs->isactiver);
> > +    irqs->isactiver = readl_relaxed(base);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
> > +    base += i * sizeof(irqs->isenabler);
> > +    irqs->isenabler = readl_relaxed(base);
> > +}
> > +
> > +static void gicv3_restore_spi_irq_config(struct dist_irq_block *irqs,
> > +                                         unsigned int i, unsigned int nr_irqs,
> > +                                         bool is_espi)
> > +{
> > +    void __iomem *base;
> > +    unsigned int irq, nr_priority_regs;
> > +
> > +    ASSERT(nr_irqs && nr_irqs <= 32);
> > +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
> > +    writel_relaxed(irqs->icfgr[0], base);
> > +    writel_relaxed(irqs->icfgr[1], base + 4);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
> > +    base += i * sizeof(irqs->ipriorityr);
> > +    for ( irq = 0; irq < nr_priority_regs; irq++ )
> > +        writel_relaxed(irqs->ipriorityr[irq], base + 4 * irq);
> > +}
> > +
> > +static void gicv3_restore_spi_irq_routing(struct dist_irq_block *irqs,
> > +                                          unsigned int i, unsigned int nr_irqs,
> > +                                          bool is_espi)
> > +{
> > +    void __iomem *base;
> > +    unsigned int irq;
> > +
> > +    ASSERT(nr_irqs && nr_irqs <= 32);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
> > +    base += i * sizeof(irqs->irouter);
> > +    for ( irq = 0; irq < nr_irqs; irq++ )
> > +        writeq_relaxed_non_atomic(irqs->irouter[irq], base + 8 * irq);
> > +}
> > +
> > +static void gicv3_disable_spi_irq_block(unsigned int i, bool is_espi)
> > +{
> > +    void __iomem *base;
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ICENABLER, is_espi) + i * 4;
> > +    writel_relaxed(GENMASK(31, 0), base);
> > +}
> > +
> > +static void gicv3_restore_spi_irq_state(struct dist_irq_block *irqs,
> > +                                        unsigned int i, bool is_espi)
> > +{
> > +    void __iomem *base;
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
> > +    base += i * sizeof(irqs->isenabler);
> > +    writel_relaxed(irqs->isenabler, base);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ICACTIVER, is_espi) + i * 4;
> > +    writel_relaxed(GENMASK(31, 0), base);
> > +
> > +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
> > +    base += i * sizeof(irqs->isactiver);
> > +    writel_relaxed(irqs->isactiver, base);
> > +}
> > +
> > +static int gicv3_check_ap1r(unsigned int n, register_t apr)
> > +{
> > +    if ( !apr )
> > +        return 0;
> > +
> > +    printk(XENLOG_ERR "GICv3: suspend aborted: ICC_AP1R%u_EL1=%#"
> > +           PRIregister"\n", n, apr);
> > +
> > +    return -EBUSY;
> > +}
> > +
> > +static int gicv3_check_active_priorities(register_t ctlr)
> > +{
> > +    unsigned int pribits = MASK_EXTR(ctlr, ICC_CTLR_EL1_PRIBITS_MASK) + 1;
> > +    int ret;
> > +
> > +    /*
> > +     * Xen enables physical Group 1 interrupts through ICC_IGRPEN1_EL1,
> > +     * so only the physical Group 1 active-priority registers are relevant
> > +     * here. Use ICC_CTLR_EL1.PRIbits for the physical CPU interface, not
> > +     * ICH_VTR_EL2, which describes the virtual interface. ICC_AP1R1_EL1 is
> > +     * only implemented with at least 6 physical priority bits, and
> > +     * ICC_AP1R2_EL1/ICC_AP1R3_EL1 with at least 7.
> > +     */
> > +    switch ( pribits )
> > +    {
> > +    case 8:
> > +    case 7:
> > +        ret = gicv3_check_ap1r(3, READ_SYSREG(ICC_AP1R3_EL1));
> > +        if ( ret )
> > +            return ret;
> > +        ret = gicv3_check_ap1r(2, READ_SYSREG(ICC_AP1R2_EL1));
> > +        if ( ret )
> > +            return ret;
> > +        /* Fall through */
> > +    case 6:
> > +        ret = gicv3_check_ap1r(1, READ_SYSREG(ICC_AP1R1_EL1));
> > +        if ( ret )
> > +            return ret;
> > +        /* Fall through */
> > +    default:
> > +        return gicv3_check_ap1r(0, READ_SYSREG(ICC_AP1R0_EL1));
> > +    }
> > +}
> > +
> > +static int gicv3_suspend(void)
> > +{
> > +    unsigned int i, nr_irqs;
> > +    void __iomem *base;
> > +    int ret;
> > +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
> > +
> > +    /* Save GICC configuration */
> > +    gicv3_ctx.cpu.ctlr     = READ_SYSREG(ICC_CTLR_EL1);
> > +    gicv3_ctx.cpu.pmr      = READ_SYSREG(ICC_PMR_EL1);
> > +    gicv3_ctx.cpu.bpr      = READ_SYSREG(ICC_BPR1_EL1);
> > +    gicv3_ctx.cpu.sre_el2  = READ_SYSREG(ICC_SRE_EL2);
> > +    gicv3_ctx.cpu.grpen    = READ_SYSREG(ICC_IGRPEN1_EL1);
> > +
> > +    gicv3_disable_interface();
> > +
> > +    ret = gicv3_check_active_priorities(gicv3_ctx.cpu.ctlr);
> > +    if ( ret )
> > +        goto out_enable_iface;
> > +
> > +    ret = gicv3_disable_redist();
> > +    if ( ret )
> > +        goto out_enable_iface;
>
> I am wondering about the timeout case here.
>
> gicv3_disable_redist() has set ProcessorSleep to 1, but returns while
> ChildrenAsleep is still 0.
> This new error path then calls gicv3_enable_redist(), which clears
> ProcessorSleep.
>
> Could that happen before ChildrenAsleep reaches 1?
> The GIC specification says that transition is UNPREDICTABLE.
> How should we handle the timeout?

Yes, ChildrenAsleep may still be 0 when the error path calls
gicv3_enable_redist(). Clearing ProcessorSleep in that state is
UNPREDICTABLE.

I propose calling panic() if waiting for ChildrenAsleep to become
1 times out, before entering the rollback path. We cannot safely
restore the CPU interface without completing the redistributor
sleep/wake sequence.

>
> > +
> > +    /* Save GICR configuration */
> > +    gicv3_redist_wait_for_rwp();
> > +
> > +    base = GICD_RDIST_BASE;
> > +
> > +    rdist->ctlr = readl_relaxed(base + GICR_CTLR);
> > +
> > +    rdist->propbase = readq_relaxed(base + GICR_PROPBASER);
> > +    rdist->pendbase = readq_relaxed(base + GICR_PENDBASER);
> > +
> > +    base = GICD_RDIST_SGI_BASE;
> > +
> > +    /* Save priority on PPI and SGI interrupts */
> > +    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
> > +        rdist->ipriorityr[i] = readl_relaxed(base + GICR_IPRIORITYR0 + 4 * i);
> > +
> > +    rdist->isactiver = readl_relaxed(base + GICR_ISACTIVER0);
> > +    rdist->isenabler = readl_relaxed(base + GICR_ISENABLER0);
> > +    rdist->igroupr   = readl_relaxed(base + GICR_IGROUPR0);
> > +    rdist->icfgr     = readl_relaxed(base + GICR_ICFGR1);
> > +
> > +    /* Save GICD configuration */
> > +    gicv3_dist_wait_for_rwp();
> > +    gicv3_ctx.dist.ctlr = readl_relaxed(GICD + GICD_CTLR);
> > +
> > +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> > +    {
> > +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> > +        gicv3_store_spi_irq_block(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
> > +                                  false);
> > +    }
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> > +        gicv3_store_spi_irq_block(gicv3_ctx.dist.espi_irqs + i, i, 32, true);
> > +#endif
> > +
> > +    return 0;
> > +
> > + out_enable_iface:

I also noticed that this label should be immediately before
gicv3_hyp_enable(true).

> > +    if ( gicv3_enable_redist() )
> > +        panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
> > +
> > +    gicv3_hyp_enable(true);
> > +    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
> > +    isb();
> > +
> > +    return ret;
> > +}
> > +
> > +static void gicv3_resume(void)
> > +{
> > +    int ret;
> > +    unsigned int i, nr_irqs;
> > +    uint32_t dist_ctlr;
> > +    void __iomem *base;
> > +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
> > +
> > +    dist_ctlr = gicv3_ctx.dist.ctlr & GICD_CTLR_ARE_NS;
> > +
> > +    /* Disable group forwarding while preserving affinity routing state. */
> > +    writel_relaxed(dist_ctlr, GICD + GICD_CTLR);
> > +    gicv3_dist_wait_for_rwp();
> > +
> > +    /*
> > +     * IHI0069H.b 12.9.9 says changing GICD_ICFGR<n>.Int_config
> > +     * while the interrupt is individually enabled is UNPREDICTABLE.
> > +     * Disable SPIs first; 4.7.1 defines GICD_ICENABLER<n>, n > 0,
> > +     * as the per-SPI disable mechanism.
> > +     */
> > +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> > +        gicv3_disable_spi_irq_block(i, false);
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> > +        gicv3_disable_spi_irq_block(i, true);
> > +#endif
> > +
> > +    gicv3_dist_wait_for_rwp();
> > +
> > +    for ( i = NR_GIC_LOCAL_IRQS; i < gicv3_info.nr_lines; i += 32 )
> > +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPR + (i / 32) * 4);
> > +
> > +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> > +    {
> > +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> > +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
> > +                                     false);
> > +    }
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> > +    {
> > +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPRnE + i * 4);
> > +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.espi_irqs + i, i, 32,
> > +                                     true);
> > +    }
> > +#endif
> > +
> > +    if ( dist_ctlr )
> > +    {
> > +        for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> > +        {
> > +            nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
> > +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.irqs + i - 1, i,
> > +                                          nr_irqs, false);
> > +        }
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +        for ( i = 0; i < gic_number_espis() / 32; i++ )
> > +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.espi_irqs + i, i,
> > +                                          32, true);
> > +#endif
> > +    }
> > +
> > +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
> > +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.irqs + i - 1, i, false);
> > +
> > +#ifdef CONFIG_GICV3_ESPI
> > +    for ( i = 0; i < gic_number_espis() / 32; i++ )
> > +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.espi_irqs + i, i, true);
> > +#endif
> > +
> > +    writel_relaxed(gicv3_ctx.dist.ctlr, GICD + GICD_CTLR);
> > +    gicv3_dist_wait_for_rwp();
> > +
> > +    ret = gicv3_lpi_init_rdist(GICD_RDIST_BASE);
> > +    /*
> > +     * If LPIs are already enabled, assume firmware or the still-powered
> > +     * redistributor has valid PROPBASER/PENDBASER and skip reprogramming.
> > +     * Return -EBUSY so callers can ignore this case.
> > +     */
> > +    if ( ret && ret != -ENODEV && ret != -EBUSY )
> > +        panic("GICv3: Failed to re-initialize LPIs during resume\n");
> > +    else if ( ret == -EBUSY ) /* extra checks, just to be sure */
>
> Comment first letter capitalize:
> s/extra/Extra/

I will also fix all the capitalization issues you pointed out.

Thanks,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions
  2026-09-24 22:23     ` Mykola Kvach
@ 2026-09-28  7:36       ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28  7:36 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 25 Sep 2026, at 00:23, Mykola Kvach <xakep.amatop@gmail.com> wrote:
> 
> Hi Bertrand,
> 
> Thank you for the review.
> 
> On Wed, Sep 23, 2026 at 6:28 PM Bertrand Marquis
> <Bertrand.Marquis@arm.com> wrote:
>> 
>> Hi Mykola,
>> 
>> Sorry for the delay to review this serie.
>> 
>>> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
>>> 
>>> From: Mirela Simonovic <mirela.simonovic@aggios.com>
>>> 
>>> System suspend may lead to a state where GIC would be powered down.
>>> Therefore, Xen should save/restore the context of GIC on suspend/resume.
>>> 
>>> Note that the context consists of states of registers which are
>>> controlled by the hypervisor. Other GIC registers which are accessible
>>> by guests are saved/restored on context switch.
>>> 
>>> Transient physical SGI pending state (GICD_CPENDSGIRn/GICD_SPENDSGIRn)
>>> is intentionally excluded. CPU-interface active-priority state is also
>>> not restored across suspend/resume. Xen reaches the final suspend path
>>> at a quiescent point, so there is no active-priority execution context
>>> to replay after resume. Enforce this with a runtime check after
>>> disabling the CPU interface: if any implemented GICC_APRn word is still
>>> non-zero, restore GICC_CTLR and abort suspend with -EBUSY.
>> 
>> You mention SGI pending state but you do not say what would happen for PPI/SPI
>> pending state, and the patch does not look at or save/restore GICD_ISPENDR.
>> 
>> Can you clarify what is expected for those?
> 
> This patch was originally based on the Linux GICv2 suspend/resume
> code, which also does not save PPI/SPI pending state. The comment
> above gic_dist_restore() explains that level interrupts still
> asserted after resume will be handled, while edge events during
> suspend need to be handled by the platform-specific wakeup
> mechanism.
> 
> Saving GICD_ISPENDR would preserve the pending state at the time of
> each read. However, an interrupt could become pending after that
> read and before the GIC loses power. Saving pending state alone
> therefore does not cover the whole suspend transition.
> 
> The assumption here is that device drivers have stopped normal I/O
> and quiesced non-wakeup interrupt sources before Xen suspends the
> GIC. Earlier events must already have been handled, or their state
> must be preserved outside the GIC. Only configured wakeup sources
> are expected to generate new events at this point.
> 
> We rely on the platform wakeup mechanism throughout suspend entry
> and sleep. If the GIC loses power, this mechanism must capture
> wakeup events outside the GIC and keep them observable after
> resume.
> 
> Not saving pending state depends on these assumptions. The race
> after a register read does not, by itself, justify losing an event
> that is already pending.

I think this would deserve a bit more explanation in the commit message
and maybe a comment in the code so the next one does not wonder if
something is needed here or not and why.

> ---
> 
> While checking the pending-state question, I also noticed a related
> issue with disabling the Distributor during suspend entry.
> 
> My earlier reasoning relied on the Distributor being powered down
> during system suspend. Not every platform is required to follow
> BSA.
> 
> PSCI requires us to save the state that could be lost. Section 6.8
> explicitly discusses Distributor power-down as a feature of some
> systems. This does not require Xen to disable its interrupt group
> before calling SYSTEM_SUSPEND.
> 
> The GICv2 pseudocode in section 3.7.2 shows that irq_wake and
> fiq_wake depend on the Distributor group enables, but not on the
> CPU interface group enables. Clearing the group enable in
> GICD_CTLR can therefore block that group's wakeup path on a
> platform that uses these signals.
> 
> I therefore propose leaving the Distributor group enabled during
> suspend entry, while still saving its configuration in case it
> loses power. The platform would handle any further shutdown and
> the required wakeup configuration. The Distributor would still
> be disabled while restoring its registers on resume.
> 
> Linux also leaves the GICv2 CPU interface enabled before the PSCI
> call. Our early gicv2_cpu_disable() is another difference: it can
> hide that group's interrupts from TF-A's early ISR_EL1 check.
> I propose leaving that shutdown to firmware as well.

Yes that would make sense.
I will let you do that on next version once i will be done with this
version of the serie.

Cheers
Bertrand

> 
> Best regards,
> Mykola



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF
  2026-09-24 23:22     ` Mykola Kvach
@ 2026-09-28  7:37       ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28  7:37 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 25 Sep 2026, at 01:22, Mykola Kvach <xakep.amatop@gmail.com> wrote:
> 
> Hi Bertrand,
> 
> Thank you for the review.
> 
> On Wed, Sep 23, 2026 at 6:36 PM Bertrand Marquis
> <Bertrand.Marquis@arm.com> wrote:
>> 
>> Hi Mykola,
>> 
>>> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
>>> 
>>> PSCI does not guarantee that a GICv3 redistributor is powered down across
>>> CPU_OFF -> CPU_ON.
>>> 
>>> DEN0022F.b says CPU_OFF powers down the calling core (5.5) and CPU_ON
>>> brings the core back with a defined initial CPU state (5.6, 6.4).
>>> However, PSCI leaves interrupt migration and GIC re-initialization to the
>>> supervisory software/firmware stack: the caller must migrate interrupts
>>> away before CPU_OFF (5.5.2), and the execution context that is lost in a
>>> powerdown state must be saved and restored by software (6.8). PSCI also
>>> calls out GIC management explicitly in 6.8, including retargeting SPIs,
>>> preventing PPIs/SGIs from targeting a powered down CPU, and reinitializing
>>> the CPU interface after CPU_ON.
>>> 
>>> This matches the GIC architecture. IHI0069H.b Chapter 11.1 requires the PE
>>> and CPU interface to share a power domain, but explicitly allows the
>>> associated redistributor, distributor, and ITS to remain powered while the
>>> PE and CPU interface are off. All other GIC power-management behavior is
>>> IMPLEMENTATION DEFINED. DEN0050D Chapter 4.2, "Generic Interrupt
>>> Controller (GIC)", says the GICv3 redistributor may live either in the AP
>>> core power domain or in a relatively always-on parent domain. So after
>>> CPU_OFF -> CPU_ON a secondary CPU can legitimately come back to a live
>>> redistributor with GICR_CTLR.EnableLPIs still set.
>>> 
>>> Handle that case in the LPI setup path instead of assuming a fully reset
>>> redistributor.
>>> 
>>> The LPI path needs special care because the GIC spec makes redistributor
>>> LPI state sticky and partially implementation defined. IHI0069H.b 5.1.1
>>> and 5.1.2 say that changing GICR_PROPBASER or GICR_PENDBASER while
>>> GICR_CTLR.EnableLPIs == 1 is UNPREDICTABLE. After clearing EnableLPIs,
>>> software must wait for GICR_CTLR.RWP == 0 before touching the pending
>>> table. The architecture also permits implementations where, once
>>> EnableLPIs has been set, clearing it again is not guaranteed to work.
>>> Where an ITS is present, the spec strongly recommends moving LPIs to
>>> another redistributor before clearing EnableLPIs.
>>> 
>>> Because of that, treat a retained EnableLPIs state as valid when the
>>> redistributor still points at Xen's expected PROPBASER/PENDBASER tables.
>>> Only try to clear EnableLPIs when the retained configuration does not
>>> match Xen's state, and wait for RWP before reprogramming the tables.
>>> 
>>> This is also consistent with platform firmware reality: PSCI and the GIC
>>> architecture allow platform-specific redistributor power handling, and not
>>> all platform firmware implementations force a full redistributor power-off
>>> through implementation-defined controls during CPU_OFF. Xen therefore needs
>>> to tolerate retained redistributor state on secondary CPU bring-up.
>>> 
>>> Keep gicv3_populate_rdist() resident as well, because gicv3_cpu_init()
>>> reuses it on secondary CPU bring-up after init.
>>> 
>>> Tested using Xen's non-boot CPU disable/enable path on Arm
>>> FVP_Base_RevC-2xAEMvA, both with and without:
>>> -C gic_distributor.allow-LPIEN-clear=1
>>> -C gic_distributor.GICR-clear-enable-supported=1
>>> and on Orange Pi 5.
>>> 
>>> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
>>> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
>>> ---
>>> Changes in v10:
>>> - Drop unrelated gicv3_populate_rdist() printk() format cleanups to keep
>>> the patch focused on retained redistributor LPI state.
>>> 
>>> Changes in v9:
>>> - move gicv3_do_wait_for_rwp prototype from its related header to gic.h
>>> - drop __init from gicv3_populate_rdist(), which is reused on secondary
>>> CPU bring-up after boot
>>> - changed print format for smp_processor_id in gicv3_populate_rdist func
>>> - cosmetic changes
>>> ---
>>> xen/arch/arm/gic-v3-lpi.c      | 77 +++++++++++++++++++++++++++++++++-
>>> xen/arch/arm/gic-v3.c          | 15 ++++---
>>> xen/arch/arm/include/asm/gic.h |  4 ++
>>> 3 files changed, 90 insertions(+), 6 deletions(-)
>>> 
>>> diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
>>> index 9ee338edc2..847da26ff7 100644
>>> --- a/xen/arch/arm/gic-v3-lpi.c
>>> +++ b/xen/arch/arm/gic-v3-lpi.c
>>> @@ -81,6 +81,13 @@ static DEFINE_PER_CPU(struct lpi_redist_data, lpi_redist);
>>> #define MAX_NR_HOST_LPIS   (lpi_data.max_host_lpi_ids - LPI_OFFSET)
>>> #define HOST_LPIS_PER_PAGE      (PAGE_SIZE / sizeof(union host_lpi))
>>> 
>>> +#define GICR_PROPBASER_XEN_MASK  GENMASK_ULL(51, 12)
>>> +/*
>>> + * For retained redistributor state, match the pending table by address only.
>>> + * Attribute bits such as PTZ may not read back with the programmed value.
>>> + */
>>> +#define GICR_PENDBASER_XEN_MASK  GENMASK_ULL(51, 16)
>>> +
>>> static union host_lpi *gic_get_host_lpi(uint32_t plpi)
>>> {
>>>    union host_lpi *block;
>>> @@ -296,6 +303,60 @@ static int gicv3_lpi_set_pendtable(void __iomem *rdist_base)
>>>    return 0;
>>> }
>>> 
>>> +static uint64_t gicv3_lpi_expected_proptable(void)
>>> +{
>>> +    return virt_to_maddr(lpi_data.lpi_property);
>>> +}
>>> +
>>> +static uint64_t gicv3_lpi_expected_pendtable(void)
>>> +{
>>> +    return virt_to_maddr(this_cpu(lpi_redist).pending_table);
>>> +}
>>> +
>>> +static bool gicv3_lpi_tables_match(void __iomem *rdist_base)
>>> +{
>>> +    uint64_t propbase, pendbase;
>>> +
>>> +    if ( !lpi_data.lpi_property || !this_cpu(lpi_redist).pending_table )
>>> +        return false;
>>> +
>>> +    propbase = readq_relaxed(rdist_base + GICR_PROPBASER);
>>> +    pendbase = readq_relaxed(rdist_base + GICR_PENDBASER);
>>> +
>>> +    return ((propbase & GICR_PROPBASER_XEN_MASK) ==
>>> +            (gicv3_lpi_expected_proptable() & GICR_PROPBASER_XEN_MASK)) &&
>>> +           ((pendbase & GICR_PENDBASER_XEN_MASK) ==
>>> +            (gicv3_lpi_expected_pendtable() & GICR_PENDBASER_XEN_MASK));
>>> +}
>>> +
>>> +static int gicv3_lpi_disable_lpis(void __iomem *rdist_base)
>>> +{
>>> +    uint32_t reg = readl_relaxed(rdist_base + GICR_CTLR);
>>> +    int ret;
>>> +
>>> +    if ( !(reg & GICR_CTLR_ENABLE_LPIS) )
>>> +        return 0;
>>> +
>>> +    writel_relaxed(reg & ~GICR_CTLR_ENABLE_LPIS, rdist_base + GICR_CTLR);
>>> +
>>> +    /*
>>> +     * The spec only guarantees programmability when we have observed the bit
>>> +     * cleared. Where clearing is supported, RWP must reach 0 before touching
>>> +     * PROPBASER/PENDBASER again.
>>> +     */
>>> +    wmb();
>>> +
>>> +    ret = gicv3_do_wait_for_rwp(rdist_base, GICR_CTLR_RWP);
>>> +    if ( ret )
>>> +        return ret;
>>> +
>>> +    reg = readl_relaxed(rdist_base + GICR_CTLR);
>>> +    if ( reg & GICR_CTLR_ENABLE_LPIS )
>>> +        return -EBUSY;
>>> +
>>> +    return 0;
>>> +}
>>> +
>>> /*
>>> * Tell a redistributor about the (shared) property table, allocating one
>>> * if not already done.
>>> @@ -374,7 +435,21 @@ int gicv3_lpi_init_rdist(void __iomem * rdist_base)
>>>    /* Make sure LPIs are disabled before setting up the tables. */
>>>    reg = readl_relaxed(rdist_base + GICR_CTLR);
>>>    if ( reg & GICR_CTLR_ENABLE_LPIS )
>>> -        return -EBUSY;
>>> +    {
>>> +        if ( gicv3_lpi_tables_match(rdist_base) )
>>> +            return -EBUSY;
>> 
>> I am wondering if there is a corner case when a CPU is unplugged and then
>> plugged back in. free_percpu_area() eventually frees the per-CPU area
>> containing lpi_redist.pending_table, but not the table itself. On the next
>> cpu_up(), gicv3_lpi_allocate_pendtable() allocates a new table, and I cannot
>> find where the old one is freed.
>> 
>> If the redistributor kept EnableLPIs=1 and GICR_PENDBASER pointing to the
>> old table, wouldn't gicv3_lpi_tables_match() fail?
>> 
>> What will happen if EnableLPIs cannot be cleared ? (i think this is something
>> possible in the hardware).
> 
> The sequence you describe would need to be handled when adding
> CPU hotplug support on Arm. There is currently no runtime caller
> for such an offline/online cycle outside system suspend/resume.
> 
> During system suspend, the common code preserves the per-CPU area.
> The following GICv3 suspend/resume patch also skips pending-table
> allocation on resume, so the existing table and pointer are
> reused.
> 
> The pending-table lifetime across normal CPU hotplug should be
> handled by the CPU hotplug series.

Then we should at least leave a comment to make sure that this is
handled in that serie.

> ---
> 
> While checking CPU bring-up failures during resume, I found a
> separate issue in the common cleanup code. Both CPU_UP_CANCELED
> and CPU_RESUME_FAILED can call free_percpu_area() for the same CPU
> during resume. The release metadata, including the rcu_head, is
> itself stored in that CPU's per-CPU area.
> 
> If the first release has completed, another call would access
> invalid per-CPU state. Otherwise, it can queue the same rcu_head
> again. The timer and CPU-pool callbacks on CPU_RESUME_FAILED also
> still need the per-CPU area.
> 
> On x86, park_offline_cpus prevents this particular release path.
> 
> I plan to address this in a separate preparatory patch in this
> series, keeping the per-CPU area available until the final
> CPU_RESUME_FAILED cleanup and releasing it only once.

Ok then this will be an extra patch in the next version of the serie.

Cheers
Bertrand

> 
> Best regards,
> Mykola



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions
  2026-09-25  0:03     ` Mykola Kvach
@ 2026-09-28  7:39       ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28  7:39 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 25 Sep 2026, at 02:03, Mykola Kvach <xakep.amatop@gmail.com> wrote:
> 
> Hi Bertrand,
> 
> Thank you for the review.
> 
> On Wed, Sep 23, 2026 at 6:58 PM Bertrand Marquis
> <Bertrand.Marquis@arm.com> wrote:
>> 
>> Hi Mykola,
>> 
>>> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
>>> 
>>> System suspend may lead to a state where GIC would be powered down.
>>> Therefore, Xen should save/restore the context of GIC on suspend/resume.
>>> 
>>> Note that the context consists of states of registers which are
>>> controlled by the hypervisor. Other GIC registers which are accessible
>>> by guests are saved/restored on context switch.
>>> 
>>> Before continuing suspend, also verify that the physical CPU interface
>>> has no Group 1 active-priority state left. Use ICC_CTLR_EL1.PRIbits to
>>> decide which ICC_AP1R<n>_EL1 registers are implemented, so Xen does not
>>> read an unimplemented AP1R register.
>>> 
>>> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
>>> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
>>> ---
>>> Changes in V10:
>>> - abort suspend when the physical Group 1 active-priority state is still
>>> present, deriving accessible ICC_AP1R<n>_EL1 registers from
>>> ICC_CTLR_EL1.PRIbits;
>>> - re-enable the redistributor before restoring CPU and virtual interface
>>> state on the suspend abort path;
>>> - panic if the redistributor cannot be re-enabled on the suspend abort path;
>>> - avoid saving/restoring reserved GICD_IPRIORITYR and GICD_IROUTER entries
>>> for a partially populated last SPI block;
>>> - disable Distributor group forwarding while preserving affinity routing
>>> state before restoring Distributor configuration;
>>> - disable SPI/eSPI forwarding and wait for RWP before restoring
>>> GICD_ICFGR<n>.Int_config.
>>> 
>>> Changes in V9:
>>> - fix the suspend-context comment typo and split dist_ctx declarations;
>>> - restore ICC_IGRPEN1_EL1 on the suspend error path;
>>> - re-initialize GICD_IGROUPRnE during resume;
>>> - restore GICD_IROUTER only after re-enabling ARE_NS during resume.
>>> 
>>> Changes in V8:
>>> - use right rdist base for prop/pend baser and ctrl
>>> 
>>> Changes in V7:
>>> - restore LPI regs on resume
>>> - add timeout during redist disabling
>>> - squash with suspend/resume handling for GICv3 eSPI registers
>>> - drop ITS guard paths so suspend/resume always runs; switch missing ctx
>>> allocation to panic
>>> - trim TODO comments; narrow redistributor storage to PPI icfgr
>>> - keep distributor context allocation even without ITS; adjust resume
>>> to use GENMASK(31, 0) for clearing enables
>>> - drop storage of the SGI configuration register, as SGIs are always
>>> edge-triggered
>>> ---
>>> xen/arch/arm/gic-v3-lpi.c                |   3 +
>>> xen/arch/arm/gic-v3.c                    | 458 ++++++++++++++++++++++-
>>> xen/arch/arm/include/asm/arm64/sysregs.h |   5 +
>>> xen/arch/arm/include/asm/gic_v3_defs.h   |   3 +
>>> 4 files changed, 466 insertions(+), 3 deletions(-)
>>> 
>>> diff --git a/xen/arch/arm/gic-v3-lpi.c b/xen/arch/arm/gic-v3-lpi.c
>>> index 847da26ff7..a63c8c4979 100644
>>> --- a/xen/arch/arm/gic-v3-lpi.c
>>> +++ b/xen/arch/arm/gic-v3-lpi.c
>>> @@ -467,6 +467,9 @@ static int cpu_callback(struct notifier_block *nfb, unsigned long action,
>>>    switch ( action )
>>>    {
>>>    case CPU_UP_PREPARE:
>>> +        if ( system_state == SYS_STATE_resume )
>>> +            break;
>>> +
>>>        rc = gicv3_lpi_allocate_pendtable(cpu);
>>>        if ( rc )
>>>            printk(XENLOG_ERR "Unable to allocate the pendtable for CPU%lu\n",
>>> diff --git a/xen/arch/arm/gic-v3.c b/xen/arch/arm/gic-v3.c
>>> index b16888ad84..038bf41142 100644
>>> --- a/xen/arch/arm/gic-v3.c
>>> +++ b/xen/arch/arm/gic-v3.c
>>> @@ -1078,12 +1078,12 @@ out:
>>>    return res;
>>> }
>>> 
>>> -static void gicv3_hyp_disable(void)
>>> +static void gicv3_hyp_enable(bool enable)
>>> {
>>>    register_t hcr;
>>> 
>>>    hcr = READ_SYSREG(ICH_HCR_EL2);
>>> -    hcr &= ~GICH_HCR_EN;
>>> +    hcr = enable ? (hcr | GICH_HCR_EN) : (hcr & ~GICH_HCR_EN);
>>>    WRITE_SYSREG(hcr, ICH_HCR_EL2);
>>>    isb();
>>> }
>>> @@ -1190,7 +1190,7 @@ static void gicv3_disable_interface(void)
>>>    spin_lock(&gicv3.lock);
>>> 
>>>    gicv3_cpu_disable();
>>> -    gicv3_hyp_disable();
>>> +    gicv3_hyp_enable(false);
>>> 
>>>    spin_unlock(&gicv3.lock);
>>> }
>>> @@ -1926,6 +1926,450 @@ static bool gic_dist_supports_lpis(void)
>>>    return (readl_relaxed(GICD + GICD_TYPER) & GICD_TYPE_LPIS);
>>> }
>>> 
>>> +#ifdef CONFIG_SYSTEM_SUSPEND
>>> +
>>> +/* This struct represents a block of 32 IRQs */
>>> +struct dist_irq_block {
>>> +    uint32_t icfgr[2];
>>> +    uint32_t ipriorityr[8];
>>> +    uint64_t irouter[32];
>>> +    uint32_t isactiver;
>>> +    uint32_t isenabler;
>>> +};
>>> +
>>> +struct redist_ctx {
>>> +    uint32_t ctlr;
>>> +    uint32_t icfgr; /* only PPIs stored */
>> 
>> Can you capitalize first comment letter ?
>> s/only/Only/
>> 
>>> +    uint32_t igroupr;
>>> +    uint32_t ipriorityr[8];
>>> +    uint32_t isactiver;
>>> +    uint32_t isenabler;
>>> +
>>> +    uint64_t pendbase;
>>> +    uint64_t propbase;
>>> +};
>>> +
>>> +/* GICv3 registers to be saved/restored on system suspend/resume */
>>> +struct gicv3_ctx {
>>> +    struct dist_ctx {
>>> +        uint32_t ctlr;
>>> +        struct dist_irq_block *irqs;
>>> +        struct dist_irq_block *espi_irqs;
>>> +    } dist;
>>> +
>>> +    /* have only one rdist structure for last running CPU during suspend */
>> 
>> Same here
>> s/have/Have/
>> 
>>> +    struct redist_ctx rdist;
>>> +
>>> +    struct cpu_ctx {
>>> +        uint32_t ctlr;
>>> +        uint32_t pmr;
>>> +        uint32_t bpr;
>>> +        uint32_t sre_el2;
>>> +        uint32_t grpen;
>>> +    } cpu;
>>> +};
>>> +
>>> +static struct gicv3_ctx gicv3_ctx;
>>> +
>>> +static void __init gicv3_alloc_context(void)
>>> +{
>>> +    uint32_t blocks = DIV_ROUND_UP(gicv3_info.nr_lines, 32);
>>> +
>>> +    /* The spec allows for systems without any SPIs */
>>> +    if ( blocks > 1 )
>>> +    {
>>> +        gicv3_ctx.dist.irqs = xzalloc_array(struct dist_irq_block, blocks - 1);
>>> +        if ( !gicv3_ctx.dist.irqs )
>>> +            panic("Failed to allocate memory for GICv3 suspend context\n");
>>> +    }
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +    if ( !gic_number_espis() )
>>> +        return;
>>> +
>>> +    blocks = gic_number_espis() / 32;
>>> +    gicv3_ctx.dist.espi_irqs = xzalloc_array(struct dist_irq_block, blocks);
>>> +    if ( !gicv3_ctx.dist.espi_irqs )
>>> +        panic("Failed to allocate memory for GICv3 eSPI suspend context\n");
>>> +#endif
>>> +}
>>> +
>>> +static int gicv3_disable_redist(void)
>>> +{
>>> +    void __iomem *waker = GICD_RDIST_BASE + GICR_WAKER;
>>> +    s_time_t deadline;
>>> +
>>> +    /*
>>> +     * Avoid infinite loop if Non-secure does not have access to GICR_WAKER.
>>> +     * See Arm IHI 0069H.b, 12.11.42 GICR_WAKER:
>>> +     *     When GICD_CTLR.DS == 0 and an access is Non-secure accesses to this
>>> +     *     register are RAZ/WI.
>>> +     */
>>> +    if ( !(readl_relaxed(GICD + GICD_CTLR) & GICD_CTLR_DS) )
>>> +        return 0;
>>> +
>>> +    deadline = NOW() + MILLISECS(1000);
>>> +
>>> +    writel_relaxed(readl_relaxed(waker) | GICR_WAKER_ProcessorSleep, waker);
>>> +    while ( (readl_relaxed(waker) & GICR_WAKER_ChildrenAsleep) == 0 )
>>> +    {
>>> +        if ( NOW() > deadline )
>>> +        {
>>> +            printk("GICv3: Timeout waiting for redistributor to sleep\n");
>>> +            return -ETIMEDOUT;
>>> +        }
>>> +        cpu_relax();
>>> +        udelay(10);
>>> +    }
>>> +
>>> +    return 0;
>>> +}
>>> +
>>> +#define GET_SPI_REG_OFFSET(name, is_espi) \
>>> +    ((is_espi) ? GICD_##name##nE : GICD_##name)
>>> +
>>> +static void gicv3_store_spi_irq_block(struct dist_irq_block *irqs,
>>> +                                      unsigned int i, unsigned int nr_irqs,
>>> +                                      bool is_espi)
>>> +{
>>> +    void __iomem *base;
>>> +    unsigned int irq, nr_priority_regs;
>>> +
>>> +    ASSERT(nr_irqs && nr_irqs <= 32);
>>> +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
>>> +    irqs->icfgr[0] = readl_relaxed(base);
>>> +    irqs->icfgr[1] = readl_relaxed(base + 4);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
>>> +    base += i * sizeof(irqs->ipriorityr);
>>> +    for ( irq = 0; irq < nr_priority_regs; irq++ )
>>> +        irqs->ipriorityr[irq] = readl_relaxed(base + 4 * irq);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
>>> +    base += i * sizeof(irqs->irouter);
>>> +    for ( irq = 0; irq < nr_irqs; irq++ )
>>> +        irqs->irouter[irq] = readq_relaxed_non_atomic(base + 8 * irq);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
>>> +    base += i * sizeof(irqs->isactiver);
>>> +    irqs->isactiver = readl_relaxed(base);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
>>> +    base += i * sizeof(irqs->isenabler);
>>> +    irqs->isenabler = readl_relaxed(base);
>>> +}
>>> +
>>> +static void gicv3_restore_spi_irq_config(struct dist_irq_block *irqs,
>>> +                                         unsigned int i, unsigned int nr_irqs,
>>> +                                         bool is_espi)
>>> +{
>>> +    void __iomem *base;
>>> +    unsigned int irq, nr_priority_regs;
>>> +
>>> +    ASSERT(nr_irqs && nr_irqs <= 32);
>>> +    nr_priority_regs = DIV_ROUND_UP(nr_irqs, 4);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ICFGR, is_espi) + i * sizeof(irqs->icfgr);
>>> +    writel_relaxed(irqs->icfgr[0], base);
>>> +    writel_relaxed(irqs->icfgr[1], base + 4);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(IPRIORITYR, is_espi);
>>> +    base += i * sizeof(irqs->ipriorityr);
>>> +    for ( irq = 0; irq < nr_priority_regs; irq++ )
>>> +        writel_relaxed(irqs->ipriorityr[irq], base + 4 * irq);
>>> +}
>>> +
>>> +static void gicv3_restore_spi_irq_routing(struct dist_irq_block *irqs,
>>> +                                          unsigned int i, unsigned int nr_irqs,
>>> +                                          bool is_espi)
>>> +{
>>> +    void __iomem *base;
>>> +    unsigned int irq;
>>> +
>>> +    ASSERT(nr_irqs && nr_irqs <= 32);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(IROUTER, is_espi);
>>> +    base += i * sizeof(irqs->irouter);
>>> +    for ( irq = 0; irq < nr_irqs; irq++ )
>>> +        writeq_relaxed_non_atomic(irqs->irouter[irq], base + 8 * irq);
>>> +}
>>> +
>>> +static void gicv3_disable_spi_irq_block(unsigned int i, bool is_espi)
>>> +{
>>> +    void __iomem *base;
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ICENABLER, is_espi) + i * 4;
>>> +    writel_relaxed(GENMASK(31, 0), base);
>>> +}
>>> +
>>> +static void gicv3_restore_spi_irq_state(struct dist_irq_block *irqs,
>>> +                                        unsigned int i, bool is_espi)
>>> +{
>>> +    void __iomem *base;
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ISENABLER, is_espi);
>>> +    base += i * sizeof(irqs->isenabler);
>>> +    writel_relaxed(irqs->isenabler, base);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ICACTIVER, is_espi) + i * 4;
>>> +    writel_relaxed(GENMASK(31, 0), base);
>>> +
>>> +    base = GICD + GET_SPI_REG_OFFSET(ISACTIVER, is_espi);
>>> +    base += i * sizeof(irqs->isactiver);
>>> +    writel_relaxed(irqs->isactiver, base);
>>> +}
>>> +
>>> +static int gicv3_check_ap1r(unsigned int n, register_t apr)
>>> +{
>>> +    if ( !apr )
>>> +        return 0;
>>> +
>>> +    printk(XENLOG_ERR "GICv3: suspend aborted: ICC_AP1R%u_EL1=%#"
>>> +           PRIregister"\n", n, apr);
>>> +
>>> +    return -EBUSY;
>>> +}
>>> +
>>> +static int gicv3_check_active_priorities(register_t ctlr)
>>> +{
>>> +    unsigned int pribits = MASK_EXTR(ctlr, ICC_CTLR_EL1_PRIBITS_MASK) + 1;
>>> +    int ret;
>>> +
>>> +    /*
>>> +     * Xen enables physical Group 1 interrupts through ICC_IGRPEN1_EL1,
>>> +     * so only the physical Group 1 active-priority registers are relevant
>>> +     * here. Use ICC_CTLR_EL1.PRIbits for the physical CPU interface, not
>>> +     * ICH_VTR_EL2, which describes the virtual interface. ICC_AP1R1_EL1 is
>>> +     * only implemented with at least 6 physical priority bits, and
>>> +     * ICC_AP1R2_EL1/ICC_AP1R3_EL1 with at least 7.
>>> +     */
>>> +    switch ( pribits )
>>> +    {
>>> +    case 8:
>>> +    case 7:
>>> +        ret = gicv3_check_ap1r(3, READ_SYSREG(ICC_AP1R3_EL1));
>>> +        if ( ret )
>>> +            return ret;
>>> +        ret = gicv3_check_ap1r(2, READ_SYSREG(ICC_AP1R2_EL1));
>>> +        if ( ret )
>>> +            return ret;
>>> +        /* Fall through */
>>> +    case 6:
>>> +        ret = gicv3_check_ap1r(1, READ_SYSREG(ICC_AP1R1_EL1));
>>> +        if ( ret )
>>> +            return ret;
>>> +        /* Fall through */
>>> +    default:
>>> +        return gicv3_check_ap1r(0, READ_SYSREG(ICC_AP1R0_EL1));
>>> +    }
>>> +}
>>> +
>>> +static int gicv3_suspend(void)
>>> +{
>>> +    unsigned int i, nr_irqs;
>>> +    void __iomem *base;
>>> +    int ret;
>>> +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
>>> +
>>> +    /* Save GICC configuration */
>>> +    gicv3_ctx.cpu.ctlr     = READ_SYSREG(ICC_CTLR_EL1);
>>> +    gicv3_ctx.cpu.pmr      = READ_SYSREG(ICC_PMR_EL1);
>>> +    gicv3_ctx.cpu.bpr      = READ_SYSREG(ICC_BPR1_EL1);
>>> +    gicv3_ctx.cpu.sre_el2  = READ_SYSREG(ICC_SRE_EL2);
>>> +    gicv3_ctx.cpu.grpen    = READ_SYSREG(ICC_IGRPEN1_EL1);
>>> +
>>> +    gicv3_disable_interface();
>>> +
>>> +    ret = gicv3_check_active_priorities(gicv3_ctx.cpu.ctlr);
>>> +    if ( ret )
>>> +        goto out_enable_iface;
>>> +
>>> +    ret = gicv3_disable_redist();
>>> +    if ( ret )
>>> +        goto out_enable_iface;
>> 
>> I am wondering about the timeout case here.
>> 
>> gicv3_disable_redist() has set ProcessorSleep to 1, but returns while
>> ChildrenAsleep is still 0.
>> This new error path then calls gicv3_enable_redist(), which clears
>> ProcessorSleep.
>> 
>> Could that happen before ChildrenAsleep reaches 1?
>> The GIC specification says that transition is UNPREDICTABLE.
>> How should we handle the timeout?
> 
> Yes, ChildrenAsleep may still be 0 when the error path calls
> gicv3_enable_redist(). Clearing ProcessorSleep in that state is
> UNPREDICTABLE.
> 
> I propose calling panic() if waiting for ChildrenAsleep to become
> 1 times out, before entering the rollback path. We cannot safely
> restore the CPU interface without completing the redistributor
> sleep/wake sequence.

Yes it agree we should do that and have a proper log for this.

> 
>> 
>>> +
>>> +    /* Save GICR configuration */
>>> +    gicv3_redist_wait_for_rwp();
>>> +
>>> +    base = GICD_RDIST_BASE;
>>> +
>>> +    rdist->ctlr = readl_relaxed(base + GICR_CTLR);
>>> +
>>> +    rdist->propbase = readq_relaxed(base + GICR_PROPBASER);
>>> +    rdist->pendbase = readq_relaxed(base + GICR_PENDBASER);
>>> +
>>> +    base = GICD_RDIST_SGI_BASE;
>>> +
>>> +    /* Save priority on PPI and SGI interrupts */
>>> +    for ( i = 0; i < NR_GIC_LOCAL_IRQS / 4; i++ )
>>> +        rdist->ipriorityr[i] = readl_relaxed(base + GICR_IPRIORITYR0 + 4 * i);
>>> +
>>> +    rdist->isactiver = readl_relaxed(base + GICR_ISACTIVER0);
>>> +    rdist->isenabler = readl_relaxed(base + GICR_ISENABLER0);
>>> +    rdist->igroupr   = readl_relaxed(base + GICR_IGROUPR0);
>>> +    rdist->icfgr     = readl_relaxed(base + GICR_ICFGR1);
>>> +
>>> +    /* Save GICD configuration */
>>> +    gicv3_dist_wait_for_rwp();
>>> +    gicv3_ctx.dist.ctlr = readl_relaxed(GICD + GICD_CTLR);
>>> +
>>> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
>>> +    {
>>> +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
>>> +        gicv3_store_spi_irq_block(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
>>> +                                  false);
>>> +    }
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
>>> +        gicv3_store_spi_irq_block(gicv3_ctx.dist.espi_irqs + i, i, 32, true);
>>> +#endif
>>> +
>>> +    return 0;
>>> +
>>> + out_enable_iface:
> 
> I also noticed that this label should be immediately before
> gicv3_hyp_enable(true).

ack

> 
>>> +    if ( gicv3_enable_redist() )
>>> +        panic("GICv3: Failed to re-enable redistributor after suspend abort\n");
>>> +
>>> +    gicv3_hyp_enable(true);
>>> +    WRITE_SYSREG(gicv3_ctx.cpu.grpen, ICC_IGRPEN1_EL1);
>>> +    isb();
>>> +
>>> +    return ret;
>>> +}
>>> +
>>> +static void gicv3_resume(void)
>>> +{
>>> +    int ret;
>>> +    unsigned int i, nr_irqs;
>>> +    uint32_t dist_ctlr;
>>> +    void __iomem *base;
>>> +    struct redist_ctx *rdist = &gicv3_ctx.rdist;
>>> +
>>> +    dist_ctlr = gicv3_ctx.dist.ctlr & GICD_CTLR_ARE_NS;
>>> +
>>> +    /* Disable group forwarding while preserving affinity routing state. */
>>> +    writel_relaxed(dist_ctlr, GICD + GICD_CTLR);
>>> +    gicv3_dist_wait_for_rwp();
>>> +
>>> +    /*
>>> +     * IHI0069H.b 12.9.9 says changing GICD_ICFGR<n>.Int_config
>>> +     * while the interrupt is individually enabled is UNPREDICTABLE.
>>> +     * Disable SPIs first; 4.7.1 defines GICD_ICENABLER<n>, n > 0,
>>> +     * as the per-SPI disable mechanism.
>>> +     */
>>> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
>>> +        gicv3_disable_spi_irq_block(i, false);
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
>>> +        gicv3_disable_spi_irq_block(i, true);
>>> +#endif
>>> +
>>> +    gicv3_dist_wait_for_rwp();
>>> +
>>> +    for ( i = NR_GIC_LOCAL_IRQS; i < gicv3_info.nr_lines; i += 32 )
>>> +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPR + (i / 32) * 4);
>>> +
>>> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
>>> +    {
>>> +        nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
>>> +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.irqs + i - 1, i, nr_irqs,
>>> +                                     false);
>>> +    }
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
>>> +    {
>>> +        writel_relaxed(GENMASK(31, 0), GICD + GICD_IGROUPRnE + i * 4);
>>> +        gicv3_restore_spi_irq_config(gicv3_ctx.dist.espi_irqs + i, i, 32,
>>> +                                     true);
>>> +    }
>>> +#endif
>>> +
>>> +    if ( dist_ctlr )
>>> +    {
>>> +        for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
>>> +        {
>>> +            nr_irqs = min(32U, gicv3_info.nr_lines - i * 32);
>>> +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.irqs + i - 1, i,
>>> +                                          nr_irqs, false);
>>> +        }
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +        for ( i = 0; i < gic_number_espis() / 32; i++ )
>>> +            gicv3_restore_spi_irq_routing(gicv3_ctx.dist.espi_irqs + i, i,
>>> +                                          32, true);
>>> +#endif
>>> +    }
>>> +
>>> +    for ( i = 1; i < DIV_ROUND_UP(gicv3_info.nr_lines, 32); i++ )
>>> +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.irqs + i - 1, i, false);
>>> +
>>> +#ifdef CONFIG_GICV3_ESPI
>>> +    for ( i = 0; i < gic_number_espis() / 32; i++ )
>>> +        gicv3_restore_spi_irq_state(gicv3_ctx.dist.espi_irqs + i, i, true);
>>> +#endif
>>> +
>>> +    writel_relaxed(gicv3_ctx.dist.ctlr, GICD + GICD_CTLR);
>>> +    gicv3_dist_wait_for_rwp();
>>> +
>>> +    ret = gicv3_lpi_init_rdist(GICD_RDIST_BASE);
>>> +    /*
>>> +     * If LPIs are already enabled, assume firmware or the still-powered
>>> +     * redistributor has valid PROPBASER/PENDBASER and skip reprogramming.
>>> +     * Return -EBUSY so callers can ignore this case.
>>> +     */
>>> +    if ( ret && ret != -ENODEV && ret != -EBUSY )
>>> +        panic("GICv3: Failed to re-initialize LPIs during resume\n");
>>> +    else if ( ret == -EBUSY ) /* extra checks, just to be sure */
>> 
>> Comment first letter capitalize:
>> s/extra/Extra/
> 
> I will also fix all the capitalization issues you pointed out.

Ack

Cheers
Bertrand

> 
> Thanks,
> Mykola



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks
  2026-08-27 14:31 ` [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks Mykola Kvach
@ 2026-09-28  8:01   ` Mykola Kvach
  2026-09-28  9:33     ` Bertrand Marquis
  0 siblings, 1 reply; 37+ messages in thread
From: Mykola Kvach @ 2026-09-28  8:01 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel, Stefano Stabellini, Julien Grall, Bertrand Marquis,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi all,

I have one more thing to check here regarding the R-Car Gen4 PCIe
BDF-to-OSID mappings across system suspend/resume.

The mappings are programmed in the PCIe APP CNVID/CNVIDMSK registers,
while CNVOSIDCTRL provides the default OSID. If this state is lost
during SYSTEM_SUSPEND, PCI DMA could potentially use the default OSID
after resume instead of the OSID assigned to the device.

I also found a recent Linux rcar-gen4 PCIe PM patch which notes that
S4 and V4H use an always-on PCIe power domain, while on V4M the PCIe
controller loses state during suspend. However, I have not found an
explicit guarantee that these particular APP registers are retained
on S4.

The Linux BSP is not a direct reference for this case, as its current
R-Car S4 IPMMU setup does not use the per-BDF OSID mappings programmed
through these registers.

This series has not previously been tested on R-Car S4 Spider. I have
tested the suspend/resume path on several other platforms (R-Car H3ULCB,
R-Car Gen5 and Orange Pi 5), but not this particular PCIe state.

I will also try to check this on a real R-Car S4 board by comparing
the PCIe APP register state before and after SYSTEM_SUSPEND.

So please consider this patch as needing some further investigation
from my side for now.

Thanks,
Mykola

On Thu, Aug 27, 2026 at 6:45 PM Mykola Kvach <mykola_kvach@epam.com> wrote:
>
> From: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
>
> Store and restore active context and micro-TLB registers.
>
> On resume, restore Root IPMMU context state before restoring Cache IPMMU
> micro-TLB state. Cache IPMMUs select Root contexts through their micro-TLB
> configuration, so restoring Cache micro-TLBs before the Root context
> registers are restored can expose stale or uninitialized context state.
>
> Tested on R-Car H3 Starter Kit.
>
> Signed-off-by: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in V10:
> - Iterate over registered IPMMUs in reverse order during resume so Root IPMMU
>   context state is restored before Cache IPMMU micro-TLB state.
>
> Changes in V9:
> - set dt_device_set_protected() only after ipmmu_alloc_ctx_suspend()
>   succeeds, so DT devices do not remain protected on allocation failure.
>
> Changes in V7:
> - moved suspend context allocation before pci stuff
> ---
>  xen/drivers/passthrough/arm/ipmmu-vmsa.c | 323 +++++++++++++++++++++--
>  1 file changed, 308 insertions(+), 15 deletions(-)
>
> diff --git a/xen/drivers/passthrough/arm/ipmmu-vmsa.c b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
> index fa9ab9cb13..2e54fa63d6 100644
> --- a/xen/drivers/passthrough/arm/ipmmu-vmsa.c
> +++ b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
> @@ -71,6 +71,8 @@
>  })
>  #endif
>
> +#define dev_dbg(dev, fmt, ...)    \
> +    dev_print(dev, XENLOG_DEBUG, fmt, ## __VA_ARGS__)
>  #define dev_info(dev, fmt, ...)    \
>      dev_print(dev, XENLOG_INFO, fmt, ## __VA_ARGS__)
>  #define dev_warn(dev, fmt, ...)    \
> @@ -130,6 +132,24 @@ struct ipmmu_features {
>      unsigned int imuctr_ttsel_mask;
>  };
>
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +struct ipmmu_reg_ctx {
> +    unsigned int imttlbr0;
> +    unsigned int imttubr0;
> +    unsigned int imttbcr;
> +    unsigned int imctr;
> +};
> +
> +struct ipmmu_vmsa_backup {
> +    struct device *dev;
> +    unsigned int *utlbs_val;
> +    unsigned int *asids_val;
> +    struct list_head list;
> +};
> +
> +#endif
> +
>  /* Root/Cache IPMMU device's information */
>  struct ipmmu_vmsa_device {
>      struct device *dev;
> @@ -142,6 +162,9 @@ struct ipmmu_vmsa_device {
>      struct ipmmu_vmsa_domain *domains[IPMMU_CTX_MAX];
>      unsigned int utlb_refcount[IPMMU_UTLB_MAX];
>      const struct ipmmu_features *features;
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    struct ipmmu_reg_ctx *reg_backup[IPMMU_CTX_MAX];
> +#endif
>  };
>
>  /*
> @@ -547,6 +570,249 @@ static void ipmmu_domain_free_context(struct ipmmu_vmsa_device *mmu,
>      spin_unlock_irqrestore(&mmu->lock, flags);
>  }
>
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +static DEFINE_SPINLOCK(ipmmu_devices_backup_lock);
> +static LIST_HEAD(ipmmu_devices_backup);
> +
> +static struct ipmmu_reg_ctx root_pgtable[IPMMU_CTX_MAX];
> +
> +static uint32_t ipmmu_imuasid_read(struct ipmmu_vmsa_device *mmu,
> +                                   unsigned int utlb)
> +{
> +    return ipmmu_read(mmu, ipmmu_utlb_reg(mmu, IMUASID(utlb)));
> +}
> +
> +static void ipmmu_utlbs_backup(struct ipmmu_vmsa_device *mmu)
> +{
> +    struct ipmmu_vmsa_backup *backup_data;
> +
> +    dev_dbg(mmu->dev, "Handle micro-TLBs backup\n");
> +
> +    spin_lock(&ipmmu_devices_backup_lock);
> +
> +    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
> +    {
> +        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
> +        unsigned int i;
> +
> +        if ( to_ipmmu(backup_data->dev) != mmu )
> +            continue;
> +
> +        for ( i = 0; i < fwspec->num_ids; i++ )
> +        {
> +            unsigned int utlb = fwspec->ids[i];
> +
> +            backup_data->asids_val[i] = ipmmu_imuasid_read(mmu, utlb);
> +            backup_data->utlbs_val[i] = ipmmu_imuctr_read(mmu, utlb);
> +        }
> +    }
> +
> +    spin_unlock(&ipmmu_devices_backup_lock);
> +}
> +
> +static void ipmmu_utlbs_restore(struct ipmmu_vmsa_device *mmu)
> +{
> +    struct ipmmu_vmsa_backup *backup_data;
> +
> +    dev_dbg(mmu->dev, "Handle micro-TLBs restore\n");
> +
> +    spin_lock(&ipmmu_devices_backup_lock);
> +
> +    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
> +    {
> +        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
> +        unsigned int i;
> +
> +        if ( to_ipmmu(backup_data->dev) != mmu )
> +            continue;
> +
> +        for ( i = 0; i < fwspec->num_ids; i++ )
> +        {
> +            unsigned int utlb = fwspec->ids[i];
> +
> +            ipmmu_imuasid_write(mmu, utlb, backup_data->asids_val[i]);
> +            ipmmu_imuctr_write(mmu, utlb, backup_data->utlbs_val[i]);
> +        }
> +    }
> +
> +    spin_unlock(&ipmmu_devices_backup_lock);
> +}
> +
> +static void ipmmu_domain_backup_context(struct ipmmu_vmsa_domain *domain)
> +{
> +    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
> +    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
> +
> +    dev_dbg(mmu->dev, "Handle domain context %u backup\n", domain->context_id);
> +
> +    regs->imttlbr0 = ipmmu_ctx_read_root(domain, IMTTLBR0);
> +    regs->imttubr0 = ipmmu_ctx_read_root(domain, IMTTUBR0);
> +    regs->imttbcr  = ipmmu_ctx_read_root(domain, IMTTBCR);
> +    regs->imctr    = ipmmu_ctx_read_root(domain, IMCTR);
> +}
> +
> +static void ipmmu_domain_restore_context(struct ipmmu_vmsa_domain *domain)
> +{
> +    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
> +    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
> +
> +    dev_dbg(mmu->dev, "Handle domain context %u restore\n", domain->context_id);
> +
> +    ipmmu_ctx_write_root(domain, IMTTLBR0, regs->imttlbr0);
> +    ipmmu_ctx_write_root(domain, IMTTUBR0, regs->imttubr0);
> +    ipmmu_ctx_write_root(domain, IMTTBCR,  regs->imttbcr);
> +    ipmmu_ctx_write_all(domain,  IMCTR,    regs->imctr | IMCTR_FLUSH);
> +}
> +
> +/*
> + * Xen: Unlike Linux implementation, Xen uses a single driver instance
> + * for handling all IPMMUs. There is no framework for ipmmu_suspend/resume
> + * callbacks to be invoked for each IPMMU device. So, we need to iterate
> + * through all registered IPMMUs performing required actions.
> + *
> + * Also take care of restoring special settings, such as translation
> + * table format, etc.
> + */
> +static int __must_check ipmmu_suspend(void)
> +{
> +    struct ipmmu_vmsa_device *mmu;
> +
> +    if ( !iommu_enabled )
> +        return 0;
> +
> +    printk(XENLOG_DEBUG "ipmmu: Suspending...\n");
> +
> +    spin_lock(&ipmmu_devices_lock);
> +
> +    list_for_each_entry( mmu, &ipmmu_devices, list )
> +    {
> +        if ( ipmmu_is_root(mmu) )
> +        {
> +            unsigned int i;
> +
> +            for ( i = 0; i < mmu->num_ctx; i++ )
> +            {
> +                if ( !mmu->domains[i] )
> +                    continue;
> +                ipmmu_domain_backup_context(mmu->domains[i]);
> +            }
> +        }
> +        else
> +            ipmmu_utlbs_backup(mmu);
> +    }
> +
> +    spin_unlock(&ipmmu_devices_lock);
> +
> +    return 0;
> +}
> +
> +static void ipmmu_resume(void)
> +{
> +    struct ipmmu_vmsa_device *mmu;
> +
> +    if ( !iommu_enabled )
> +        return;
> +
> +    printk(XENLOG_DEBUG "ipmmu: Resuming...\n");
> +
> +    spin_lock(&ipmmu_devices_lock);
> +
> +    /*
> +     * IPMMUs are registered with list_add(), with Root IPMMU probed first.
> +     * Walk backwards to restore Root contexts before Cache micro-TLBs.
> +     */
> +    list_for_each_entry_reverse( mmu, &ipmmu_devices, list )
> +    {
> +        uint32_t reg;
> +
> +        /* Do not use security group function */
> +        reg = IMSCTLR + mmu->features->control_offset_base;
> +        ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) & ~IMSCTLR_USE_SECGRP);
> +
> +        if ( ipmmu_is_root(mmu) )
> +        {
> +            unsigned int i;
> +
> +            /* Use stage 2 translation table format */
> +            reg = IMSAUXCTLR + mmu->features->control_offset_base;
> +            ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) | IMSAUXCTLR_S2PTE);
> +
> +            for ( i = 0; i < mmu->num_ctx; i++ )
> +            {
> +                if ( !mmu->domains[i] )
> +                    continue;
> +                ipmmu_domain_restore_context(mmu->domains[i]);
> +            }
> +        }
> +        else
> +            ipmmu_utlbs_restore(mmu);
> +    }
> +
> +    spin_unlock(&ipmmu_devices_lock);
> +}
> +
> +static int ipmmu_alloc_ctx_suspend(struct device *dev)
> +{
> +    struct ipmmu_vmsa_backup *backup_data;
> +    unsigned int *utlbs_val, *asids_val;
> +    struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(dev);
> +
> +    utlbs_val = xzalloc_array(unsigned int, fwspec->num_ids);
> +    if ( !utlbs_val )
> +        return -ENOMEM;
> +
> +    asids_val = xzalloc_array(unsigned int, fwspec->num_ids);
> +    if ( !asids_val )
> +    {
> +        xfree(utlbs_val);
> +        return -ENOMEM;
> +    }
> +
> +    backup_data = xzalloc(struct ipmmu_vmsa_backup);
> +    if ( !backup_data )
> +    {
> +        xfree(utlbs_val);
> +        xfree(asids_val);
> +        return -ENOMEM;
> +    }
> +
> +    backup_data->dev = dev;
> +    backup_data->utlbs_val = utlbs_val;
> +    backup_data->asids_val = asids_val;
> +
> +    spin_lock(&ipmmu_devices_backup_lock);
> +    list_add(&backup_data->list, &ipmmu_devices_backup);
> +    spin_unlock(&ipmmu_devices_backup_lock);
> +
> +    return 0;
> +}
> +
> +#ifdef CONFIG_HAS_PCI
> +static void ipmmu_free_ctx_suspend(struct device *dev)
> +{
> +    struct ipmmu_vmsa_backup *backup_data, *tmp;
> +
> +    spin_lock(&ipmmu_devices_backup_lock);
> +
> +    list_for_each_entry_safe( backup_data, tmp, &ipmmu_devices_backup, list )
> +    {
> +        if ( backup_data->dev == dev )
> +        {
> +            list_del(&backup_data->list);
> +            xfree(backup_data->utlbs_val);
> +            xfree(backup_data->asids_val);
> +            xfree(backup_data);
> +            break;
> +        }
> +    }
> +
> +    spin_unlock(&ipmmu_devices_backup_lock);
> +}
> +#endif /* CONFIG_HAS_PCI */
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
>  static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
>  {
>      uint64_t ttbr;
> @@ -559,6 +825,9 @@ static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
>          return ret;
>
>      domain->context_id = ret;
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    domain->mmu->root->reg_backup[ret] = &root_pgtable[ret];
> +#endif
>
>      /*
>       * TTBR0
> @@ -615,6 +884,9 @@ static void ipmmu_domain_destroy_context(struct ipmmu_vmsa_domain *domain)
>      ipmmu_ctx_write_root(domain, IMCTR, IMCTR_FLUSH);
>      ipmmu_tlb_sync(domain);
>
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    domain->mmu->root->reg_backup[domain->context_id] = NULL;
> +#endif
>      ipmmu_domain_free_context(domain->mmu->root, domain->context_id);
>  }
>
> @@ -1338,10 +1610,11 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>      struct iommu_fwspec *fwspec;
>
>  #ifdef CONFIG_HAS_PCI
> +    int ret;
> +
>      if ( dev_is_pci(dev) )
>      {
>          struct pci_dev *pdev = dev_to_pci(dev);
> -        int ret;
>
>          if ( devfn != pdev->devfn )
>              return 0;
> @@ -1358,17 +1631,24 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>      if ( !to_ipmmu(dev) )
>          return -ENODEV;
>
> -    if ( !dev_is_pci(dev) )
> +    if ( !dev_is_pci(dev) && dt_device_is_protected(dev_to_dt(dev)) )
>      {
> -        if ( dt_device_is_protected(dev_to_dt(dev)) )
> -        {
> -            dev_err(dev, "Already added to IPMMU\n");
> -            return -EEXIST;
> -        }
> +        dev_err(dev, "Already added to IPMMU\n");
> +        return -EEXIST;
> +    }
>
> -        /* Let Xen know that the master device is protected by an IOMMU. */
> -        dt_device_set_protected(dev_to_dt(dev));
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    if ( ipmmu_alloc_ctx_suspend(dev) )
> +    {
> +        dev_err(dev, "Failed to allocate context for suspend\n");
> +        return -ENOMEM;
>      }
> +#endif
> +
> +    /* Let Xen know that the master device is protected by an IOMMU. */
> +    if ( !dev_is_pci(dev) )
> +        dt_device_set_protected(dev_to_dt(dev));
> +
>  #ifdef CONFIG_HAS_PCI
>      if ( dev_is_pci(dev) )
>      {
> @@ -1377,26 +1657,28 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>          struct pci_host_bridge *bridge;
>          struct iommu_fwspec *fwspec_bridge;
>          unsigned int utlb_osid0 = 0;
> -        int ret;
>
>          bridge = pci_find_host_bridge(pdev->seg, pdev->bus);
>          if ( !bridge )
>          {
>              dev_err(dev, "Failed to find host bridge\n");
> -            return -ENODEV;
> +            ret = -ENODEV;
> +            goto free_suspend_ctx;
>          }
>
>          fwspec_bridge = dev_iommu_fwspec_get(dt_to_dev(bridge->dt_node));
>          if ( fwspec_bridge->num_ids < 1 )
>          {
>              dev_err(dev, "Failed to find host bridge uTLB\n");
> -            return -ENXIO;
> +            ret = -ENXIO;
> +            goto free_suspend_ctx;
>          }
>
>          if ( fwspec->num_ids < 1 )
>          {
>              dev_err(dev, "Failed to find uTLB");
> -            return -ENXIO;
> +            ret = -ENXIO;
> +            goto free_suspend_ctx;
>          }
>
>          rcar4_pcie_osid_regs_init(bridge);
> @@ -1405,7 +1687,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>          if ( ret < 0 )
>          {
>              dev_err(dev, "No unused OSID regs\n");
> -            return ret;
> +            goto free_suspend_ctx;
>          }
>          reg_id = ret;
>
> @@ -1420,7 +1702,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>          {
>              rcar4_pcie_osid_bdf_clear(bridge, reg_id);
>              rcar4_pcie_osid_reg_free(bridge, reg_id);
> -            return ret;
> +            goto free_suspend_ctx;
>          }
>      }
>  #endif
> @@ -1429,6 +1711,13 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>               dev_name(fwspec->iommu_dev), fwspec->num_ids);
>
>      return 0;
> +#ifdef CONFIG_HAS_PCI
> + free_suspend_ctx:
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    ipmmu_free_ctx_suspend(dev);
> +#endif
> +    return ret;
> +#endif
>  }
>
>  static int ipmmu_iommu_domain_init(struct domain *d)
> @@ -1490,6 +1779,10 @@ static const struct iommu_ops ipmmu_iommu_ops =
>      .unmap_page      = arm_iommu_unmap_page,
>      .dt_xlate        = ipmmu_dt_xlate,
>      .add_device      = ipmmu_add_device,
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    .suspend         = ipmmu_suspend,
> +    .resume          = ipmmu_resume,
> +#endif
>  };
>
>  static __init int ipmmu_init(struct dt_device_node *node, const void *data)
> --
> 2.43.0
>
>


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks
  2026-09-28  8:01   ` Mykola Kvach
@ 2026-09-28  9:33     ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28  9:33 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 28 Sep 2026, at 10:01, Mykola Kvach <xakep.amatop@gmail.com> wrote:
> 
> Hi all,
> 
> I have one more thing to check here regarding the R-Car Gen4 PCIe
> BDF-to-OSID mappings across system suspend/resume.
> 
> The mappings are programmed in the PCIe APP CNVID/CNVIDMSK registers,
> while CNVOSIDCTRL provides the default OSID. If this state is lost
> during SYSTEM_SUSPEND, PCI DMA could potentially use the default OSID
> after resume instead of the OSID assigned to the device.
> 
> I also found a recent Linux rcar-gen4 PCIe PM patch which notes that
> S4 and V4H use an always-on PCIe power domain, while on V4M the PCIe
> controller loses state during suspend. However, I have not found an
> explicit guarantee that these particular APP registers are retained
> on S4.
> 
> The Linux BSP is not a direct reference for this case, as its current
> R-Car S4 IPMMU setup does not use the per-BDF OSID mappings programmed
> through these registers.
> 
> This series has not previously been tested on R-Car S4 Spider. I have
> tested the suspend/resume path on several other platforms (R-Car H3ULCB,
> R-Car Gen5 and Orange Pi 5), but not this particular PCIe state.
> 
> I will also try to check this on a real R-Car S4 board by comparing
> the PCIe APP register state before and after SYSTEM_SUSPEND.
> 
> So please consider this patch as needing some further investigation
> from my side for now.

Thanks for the information.

I will skip this patch for now (and anyway i will need some help on it as IPMMU and
in general Renesas drivers are not really something i know).

Cheers
Bertrand

> 
> Thanks,
> Mykola
> 
> On Thu, Aug 27, 2026 at 6:45 PM Mykola Kvach <mykola_kvach@epam.com> wrote:
>> 
>> From: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
>> 
>> Store and restore active context and micro-TLB registers.
>> 
>> On resume, restore Root IPMMU context state before restoring Cache IPMMU
>> micro-TLB state. Cache IPMMUs select Root contexts through their micro-TLB
>> configuration, so restoring Cache micro-TLBs before the Root context
>> registers are restored can expose stale or uninitialized context state.
>> 
>> Tested on R-Car H3 Starter Kit.
>> 
>> Signed-off-by: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
>> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
>> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
>> ---
>> Changes in V10:
>> - Iterate over registered IPMMUs in reverse order during resume so Root IPMMU
>>  context state is restored before Cache IPMMU micro-TLB state.
>> 
>> Changes in V9:
>> - set dt_device_set_protected() only after ipmmu_alloc_ctx_suspend()
>>  succeeds, so DT devices do not remain protected on allocation failure.
>> 
>> Changes in V7:
>> - moved suspend context allocation before pci stuff
>> ---
>> xen/drivers/passthrough/arm/ipmmu-vmsa.c | 323 +++++++++++++++++++++--
>> 1 file changed, 308 insertions(+), 15 deletions(-)
>> 
>> diff --git a/xen/drivers/passthrough/arm/ipmmu-vmsa.c b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
>> index fa9ab9cb13..2e54fa63d6 100644
>> --- a/xen/drivers/passthrough/arm/ipmmu-vmsa.c
>> +++ b/xen/drivers/passthrough/arm/ipmmu-vmsa.c
>> @@ -71,6 +71,8 @@
>> })
>> #endif
>> 
>> +#define dev_dbg(dev, fmt, ...)    \
>> +    dev_print(dev, XENLOG_DEBUG, fmt, ## __VA_ARGS__)
>> #define dev_info(dev, fmt, ...)    \
>>     dev_print(dev, XENLOG_INFO, fmt, ## __VA_ARGS__)
>> #define dev_warn(dev, fmt, ...)    \
>> @@ -130,6 +132,24 @@ struct ipmmu_features {
>>     unsigned int imuctr_ttsel_mask;
>> };
>> 
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +
>> +struct ipmmu_reg_ctx {
>> +    unsigned int imttlbr0;
>> +    unsigned int imttubr0;
>> +    unsigned int imttbcr;
>> +    unsigned int imctr;
>> +};
>> +
>> +struct ipmmu_vmsa_backup {
>> +    struct device *dev;
>> +    unsigned int *utlbs_val;
>> +    unsigned int *asids_val;
>> +    struct list_head list;
>> +};
>> +
>> +#endif
>> +
>> /* Root/Cache IPMMU device's information */
>> struct ipmmu_vmsa_device {
>>     struct device *dev;
>> @@ -142,6 +162,9 @@ struct ipmmu_vmsa_device {
>>     struct ipmmu_vmsa_domain *domains[IPMMU_CTX_MAX];
>>     unsigned int utlb_refcount[IPMMU_UTLB_MAX];
>>     const struct ipmmu_features *features;
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    struct ipmmu_reg_ctx *reg_backup[IPMMU_CTX_MAX];
>> +#endif
>> };
>> 
>> /*
>> @@ -547,6 +570,249 @@ static void ipmmu_domain_free_context(struct ipmmu_vmsa_device *mmu,
>>     spin_unlock_irqrestore(&mmu->lock, flags);
>> }
>> 
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +
>> +static DEFINE_SPINLOCK(ipmmu_devices_backup_lock);
>> +static LIST_HEAD(ipmmu_devices_backup);
>> +
>> +static struct ipmmu_reg_ctx root_pgtable[IPMMU_CTX_MAX];
>> +
>> +static uint32_t ipmmu_imuasid_read(struct ipmmu_vmsa_device *mmu,
>> +                                   unsigned int utlb)
>> +{
>> +    return ipmmu_read(mmu, ipmmu_utlb_reg(mmu, IMUASID(utlb)));
>> +}
>> +
>> +static void ipmmu_utlbs_backup(struct ipmmu_vmsa_device *mmu)
>> +{
>> +    struct ipmmu_vmsa_backup *backup_data;
>> +
>> +    dev_dbg(mmu->dev, "Handle micro-TLBs backup\n");
>> +
>> +    spin_lock(&ipmmu_devices_backup_lock);
>> +
>> +    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
>> +    {
>> +        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
>> +        unsigned int i;
>> +
>> +        if ( to_ipmmu(backup_data->dev) != mmu )
>> +            continue;
>> +
>> +        for ( i = 0; i < fwspec->num_ids; i++ )
>> +        {
>> +            unsigned int utlb = fwspec->ids[i];
>> +
>> +            backup_data->asids_val[i] = ipmmu_imuasid_read(mmu, utlb);
>> +            backup_data->utlbs_val[i] = ipmmu_imuctr_read(mmu, utlb);
>> +        }
>> +    }
>> +
>> +    spin_unlock(&ipmmu_devices_backup_lock);
>> +}
>> +
>> +static void ipmmu_utlbs_restore(struct ipmmu_vmsa_device *mmu)
>> +{
>> +    struct ipmmu_vmsa_backup *backup_data;
>> +
>> +    dev_dbg(mmu->dev, "Handle micro-TLBs restore\n");
>> +
>> +    spin_lock(&ipmmu_devices_backup_lock);
>> +
>> +    list_for_each_entry( backup_data, &ipmmu_devices_backup, list )
>> +    {
>> +        struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(backup_data->dev);
>> +        unsigned int i;
>> +
>> +        if ( to_ipmmu(backup_data->dev) != mmu )
>> +            continue;
>> +
>> +        for ( i = 0; i < fwspec->num_ids; i++ )
>> +        {
>> +            unsigned int utlb = fwspec->ids[i];
>> +
>> +            ipmmu_imuasid_write(mmu, utlb, backup_data->asids_val[i]);
>> +            ipmmu_imuctr_write(mmu, utlb, backup_data->utlbs_val[i]);
>> +        }
>> +    }
>> +
>> +    spin_unlock(&ipmmu_devices_backup_lock);
>> +}
>> +
>> +static void ipmmu_domain_backup_context(struct ipmmu_vmsa_domain *domain)
>> +{
>> +    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
>> +    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
>> +
>> +    dev_dbg(mmu->dev, "Handle domain context %u backup\n", domain->context_id);
>> +
>> +    regs->imttlbr0 = ipmmu_ctx_read_root(domain, IMTTLBR0);
>> +    regs->imttubr0 = ipmmu_ctx_read_root(domain, IMTTUBR0);
>> +    regs->imttbcr  = ipmmu_ctx_read_root(domain, IMTTBCR);
>> +    regs->imctr    = ipmmu_ctx_read_root(domain, IMCTR);
>> +}
>> +
>> +static void ipmmu_domain_restore_context(struct ipmmu_vmsa_domain *domain)
>> +{
>> +    struct ipmmu_vmsa_device *mmu = domain->mmu->root;
>> +    struct ipmmu_reg_ctx *regs = mmu->reg_backup[domain->context_id];
>> +
>> +    dev_dbg(mmu->dev, "Handle domain context %u restore\n", domain->context_id);
>> +
>> +    ipmmu_ctx_write_root(domain, IMTTLBR0, regs->imttlbr0);
>> +    ipmmu_ctx_write_root(domain, IMTTUBR0, regs->imttubr0);
>> +    ipmmu_ctx_write_root(domain, IMTTBCR,  regs->imttbcr);
>> +    ipmmu_ctx_write_all(domain,  IMCTR,    regs->imctr | IMCTR_FLUSH);
>> +}
>> +
>> +/*
>> + * Xen: Unlike Linux implementation, Xen uses a single driver instance
>> + * for handling all IPMMUs. There is no framework for ipmmu_suspend/resume
>> + * callbacks to be invoked for each IPMMU device. So, we need to iterate
>> + * through all registered IPMMUs performing required actions.
>> + *
>> + * Also take care of restoring special settings, such as translation
>> + * table format, etc.
>> + */
>> +static int __must_check ipmmu_suspend(void)
>> +{
>> +    struct ipmmu_vmsa_device *mmu;
>> +
>> +    if ( !iommu_enabled )
>> +        return 0;
>> +
>> +    printk(XENLOG_DEBUG "ipmmu: Suspending...\n");
>> +
>> +    spin_lock(&ipmmu_devices_lock);
>> +
>> +    list_for_each_entry( mmu, &ipmmu_devices, list )
>> +    {
>> +        if ( ipmmu_is_root(mmu) )
>> +        {
>> +            unsigned int i;
>> +
>> +            for ( i = 0; i < mmu->num_ctx; i++ )
>> +            {
>> +                if ( !mmu->domains[i] )
>> +                    continue;
>> +                ipmmu_domain_backup_context(mmu->domains[i]);
>> +            }
>> +        }
>> +        else
>> +            ipmmu_utlbs_backup(mmu);
>> +    }
>> +
>> +    spin_unlock(&ipmmu_devices_lock);
>> +
>> +    return 0;
>> +}
>> +
>> +static void ipmmu_resume(void)
>> +{
>> +    struct ipmmu_vmsa_device *mmu;
>> +
>> +    if ( !iommu_enabled )
>> +        return;
>> +
>> +    printk(XENLOG_DEBUG "ipmmu: Resuming...\n");
>> +
>> +    spin_lock(&ipmmu_devices_lock);
>> +
>> +    /*
>> +     * IPMMUs are registered with list_add(), with Root IPMMU probed first.
>> +     * Walk backwards to restore Root contexts before Cache micro-TLBs.
>> +     */
>> +    list_for_each_entry_reverse( mmu, &ipmmu_devices, list )
>> +    {
>> +        uint32_t reg;
>> +
>> +        /* Do not use security group function */
>> +        reg = IMSCTLR + mmu->features->control_offset_base;
>> +        ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) & ~IMSCTLR_USE_SECGRP);
>> +
>> +        if ( ipmmu_is_root(mmu) )
>> +        {
>> +            unsigned int i;
>> +
>> +            /* Use stage 2 translation table format */
>> +            reg = IMSAUXCTLR + mmu->features->control_offset_base;
>> +            ipmmu_write(mmu, reg, ipmmu_read(mmu, reg) | IMSAUXCTLR_S2PTE);
>> +
>> +            for ( i = 0; i < mmu->num_ctx; i++ )
>> +            {
>> +                if ( !mmu->domains[i] )
>> +                    continue;
>> +                ipmmu_domain_restore_context(mmu->domains[i]);
>> +            }
>> +        }
>> +        else
>> +            ipmmu_utlbs_restore(mmu);
>> +    }
>> +
>> +    spin_unlock(&ipmmu_devices_lock);
>> +}
>> +
>> +static int ipmmu_alloc_ctx_suspend(struct device *dev)
>> +{
>> +    struct ipmmu_vmsa_backup *backup_data;
>> +    unsigned int *utlbs_val, *asids_val;
>> +    struct iommu_fwspec *fwspec = dev_iommu_fwspec_get(dev);
>> +
>> +    utlbs_val = xzalloc_array(unsigned int, fwspec->num_ids);
>> +    if ( !utlbs_val )
>> +        return -ENOMEM;
>> +
>> +    asids_val = xzalloc_array(unsigned int, fwspec->num_ids);
>> +    if ( !asids_val )
>> +    {
>> +        xfree(utlbs_val);
>> +        return -ENOMEM;
>> +    }
>> +
>> +    backup_data = xzalloc(struct ipmmu_vmsa_backup);
>> +    if ( !backup_data )
>> +    {
>> +        xfree(utlbs_val);
>> +        xfree(asids_val);
>> +        return -ENOMEM;
>> +    }
>> +
>> +    backup_data->dev = dev;
>> +    backup_data->utlbs_val = utlbs_val;
>> +    backup_data->asids_val = asids_val;
>> +
>> +    spin_lock(&ipmmu_devices_backup_lock);
>> +    list_add(&backup_data->list, &ipmmu_devices_backup);
>> +    spin_unlock(&ipmmu_devices_backup_lock);
>> +
>> +    return 0;
>> +}
>> +
>> +#ifdef CONFIG_HAS_PCI
>> +static void ipmmu_free_ctx_suspend(struct device *dev)
>> +{
>> +    struct ipmmu_vmsa_backup *backup_data, *tmp;
>> +
>> +    spin_lock(&ipmmu_devices_backup_lock);
>> +
>> +    list_for_each_entry_safe( backup_data, tmp, &ipmmu_devices_backup, list )
>> +    {
>> +        if ( backup_data->dev == dev )
>> +        {
>> +            list_del(&backup_data->list);
>> +            xfree(backup_data->utlbs_val);
>> +            xfree(backup_data->asids_val);
>> +            xfree(backup_data);
>> +            break;
>> +        }
>> +    }
>> +
>> +    spin_unlock(&ipmmu_devices_backup_lock);
>> +}
>> +#endif /* CONFIG_HAS_PCI */
>> +
>> +#endif /* CONFIG_SYSTEM_SUSPEND */
>> +
>> static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
>> {
>>     uint64_t ttbr;
>> @@ -559,6 +825,9 @@ static int ipmmu_domain_init_context(struct ipmmu_vmsa_domain *domain)
>>         return ret;
>> 
>>     domain->context_id = ret;
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    domain->mmu->root->reg_backup[ret] = &root_pgtable[ret];
>> +#endif
>> 
>>     /*
>>      * TTBR0
>> @@ -615,6 +884,9 @@ static void ipmmu_domain_destroy_context(struct ipmmu_vmsa_domain *domain)
>>     ipmmu_ctx_write_root(domain, IMCTR, IMCTR_FLUSH);
>>     ipmmu_tlb_sync(domain);
>> 
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    domain->mmu->root->reg_backup[domain->context_id] = NULL;
>> +#endif
>>     ipmmu_domain_free_context(domain->mmu->root, domain->context_id);
>> }
>> 
>> @@ -1338,10 +1610,11 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>     struct iommu_fwspec *fwspec;
>> 
>> #ifdef CONFIG_HAS_PCI
>> +    int ret;
>> +
>>     if ( dev_is_pci(dev) )
>>     {
>>         struct pci_dev *pdev = dev_to_pci(dev);
>> -        int ret;
>> 
>>         if ( devfn != pdev->devfn )
>>             return 0;
>> @@ -1358,17 +1631,24 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>     if ( !to_ipmmu(dev) )
>>         return -ENODEV;
>> 
>> -    if ( !dev_is_pci(dev) )
>> +    if ( !dev_is_pci(dev) && dt_device_is_protected(dev_to_dt(dev)) )
>>     {
>> -        if ( dt_device_is_protected(dev_to_dt(dev)) )
>> -        {
>> -            dev_err(dev, "Already added to IPMMU\n");
>> -            return -EEXIST;
>> -        }
>> +        dev_err(dev, "Already added to IPMMU\n");
>> +        return -EEXIST;
>> +    }
>> 
>> -        /* Let Xen know that the master device is protected by an IOMMU. */
>> -        dt_device_set_protected(dev_to_dt(dev));
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    if ( ipmmu_alloc_ctx_suspend(dev) )
>> +    {
>> +        dev_err(dev, "Failed to allocate context for suspend\n");
>> +        return -ENOMEM;
>>     }
>> +#endif
>> +
>> +    /* Let Xen know that the master device is protected by an IOMMU. */
>> +    if ( !dev_is_pci(dev) )
>> +        dt_device_set_protected(dev_to_dt(dev));
>> +
>> #ifdef CONFIG_HAS_PCI
>>     if ( dev_is_pci(dev) )
>>     {
>> @@ -1377,26 +1657,28 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>         struct pci_host_bridge *bridge;
>>         struct iommu_fwspec *fwspec_bridge;
>>         unsigned int utlb_osid0 = 0;
>> -        int ret;
>> 
>>         bridge = pci_find_host_bridge(pdev->seg, pdev->bus);
>>         if ( !bridge )
>>         {
>>             dev_err(dev, "Failed to find host bridge\n");
>> -            return -ENODEV;
>> +            ret = -ENODEV;
>> +            goto free_suspend_ctx;
>>         }
>> 
>>         fwspec_bridge = dev_iommu_fwspec_get(dt_to_dev(bridge->dt_node));
>>         if ( fwspec_bridge->num_ids < 1 )
>>         {
>>             dev_err(dev, "Failed to find host bridge uTLB\n");
>> -            return -ENXIO;
>> +            ret = -ENXIO;
>> +            goto free_suspend_ctx;
>>         }
>> 
>>         if ( fwspec->num_ids < 1 )
>>         {
>>             dev_err(dev, "Failed to find uTLB");
>> -            return -ENXIO;
>> +            ret = -ENXIO;
>> +            goto free_suspend_ctx;
>>         }
>> 
>>         rcar4_pcie_osid_regs_init(bridge);
>> @@ -1405,7 +1687,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>         if ( ret < 0 )
>>         {
>>             dev_err(dev, "No unused OSID regs\n");
>> -            return ret;
>> +            goto free_suspend_ctx;
>>         }
>>         reg_id = ret;
>> 
>> @@ -1420,7 +1702,7 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>         {
>>             rcar4_pcie_osid_bdf_clear(bridge, reg_id);
>>             rcar4_pcie_osid_reg_free(bridge, reg_id);
>> -            return ret;
>> +            goto free_suspend_ctx;
>>         }
>>     }
>> #endif
>> @@ -1429,6 +1711,13 @@ static int ipmmu_add_device(u8 devfn, struct device *dev)
>>              dev_name(fwspec->iommu_dev), fwspec->num_ids);
>> 
>>     return 0;
>> +#ifdef CONFIG_HAS_PCI
>> + free_suspend_ctx:
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    ipmmu_free_ctx_suspend(dev);
>> +#endif
>> +    return ret;
>> +#endif
>> }
>> 
>> static int ipmmu_iommu_domain_init(struct domain *d)
>> @@ -1490,6 +1779,10 @@ static const struct iommu_ops ipmmu_iommu_ops =
>>     .unmap_page      = arm_iommu_unmap_page,
>>     .dt_xlate        = ipmmu_dt_xlate,
>>     .add_device      = ipmmu_add_device,
>> +#ifdef CONFIG_SYSTEM_SUSPEND
>> +    .suspend         = ipmmu_suspend,
>> +    .resume          = ipmmu_resume,
>> +#endif
>> };
>> 
>> static __init int ipmmu_init(struct dt_device_node *node, const void *data)
>> --
>> 2.43.0
>> 
>> 


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers
  2026-08-27 14:31 ` [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers Mykola Kvach
@ 2026-09-28 16:17   ` Bertrand Marquis
  2026-09-30 14:44     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28 16:17 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Rahul Singh, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk,
	Pranjal Shrivastava, Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> Add system suspend/resume callbacks for the Arm SMMUv3 driver.
> 
> During suspend, configure GBPA to abort incoming transactions, disable the
> translation interface while keeping CMDQ enabled, issue CMD_SYNC to ensure
> all previously issued commands have completed, then disable the SMMU IRQs
> and SMMU.
> 
> Resume uses arm_smmu_device_reset() to reprogram the SMMU and re-enable
> translation and interrupt generation.
> 
> The IRQ setup split follows the approach from Pranjal Shrivastava's Linux
> arm-smmu-v3 runtime/system sleep series: IRQ handlers are requested once
> during probe, while reset/resume only restores SMMU hardware state and
> re-enables IRQ_CTRL.
> 
> Only the pieces relevant to Xen's currently supported SMMUv3 path are
> ported here. Xen documents SMMUv3 MSI and PCI ATS as unsupported and not
> compiled/tested, so this patch does not restore SMMU MSI IRQ_CFGn registers
> nor reinitialize ATS/PRI endpoints. If those paths become usable,
> suspend/resume will need corresponding MSI restore and ATS/PRI
> quiesce/reinit steps.
> 
> Link: https://lore.kernel.org/r/20260414194702.1229094-1-praan@google.com/
> Based-on-patch-by: Pranjal Shrivastava <praan@google.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in V11:
> - Keep arm_smmu_update_gbpa() and arm_smmu_device_reset() in init text when
>  CONFIG_SYSTEM_SUSPEND is disabled.
> 
> Changes in V10:
> - Disable SMMU interrupt generation during suspend before disabling the
>  SMMU interface, matching the resume/reset path which re-enables IRQ_CTRL.
> 
> Changes in V9:
> - Use CMD_SYNC in suspend instead of polling CMDQ_CONS, so the suspend
>  path waits for command completion rather than only command consumption.
> - Document that arm_smmu_setup_irqs() is probe-only and that future Xen
>  SMMUv3 MSI support will need to restore SMMU IRQ_CFGn registers on
>  resume.
> - Restore the reference to Pranjal's Linux runtime/system sleep series and
>  clarify that MSI/ATS/PRI resume handling is outside the supported Xen
>  path.
> - Prefix the subject with xen/arm for consistency with the rest of the
>  Arm suspend/resume series.
> 
> Changes in V8:
> - Honor ARM_SMMU_FEAT_SEV when draining the CMDQ during suspend, matching
>  the existing runtime CMD_SYNC path.
> - Fold the suspend rollback reset path into a helper and rename the error
>  reporting to describe suspend rollback rather than resume.
> - Treat SMMU reset failure during resume as fatal instead of logging and
>  continuing with a potentially unusable IOMMU.
> - cosmetic changes
> ---
> xen/drivers/passthrough/arm/smmu-v3.c | 194 +++++++++++++++++++++-----
> 1 file changed, 158 insertions(+), 36 deletions(-)
> 
> diff --git a/xen/drivers/passthrough/arm/smmu-v3.c b/xen/drivers/passthrough/arm/smmu-v3.c
> index bf153227db..7f1d00fb81 100644
> --- a/xen/drivers/passthrough/arm/smmu-v3.c
> +++ b/xen/drivers/passthrough/arm/smmu-v3.c
> @@ -94,6 +94,12 @@
> 
> #include "smmu-v3.h"
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +#define __init_or_smmu_suspend
> +#else
> +#define __init_or_smmu_suspend __init
> +#endif
> +
> #define ARM_SMMU_VTCR_SH_IS 3
> #define ARM_SMMU_VTCR_RGN_WBWA 1
> #define ARM_SMMU_VTCR_TG0_4K 0
> @@ -1814,8 +1820,8 @@ static int arm_smmu_write_reg_sync(struct arm_smmu_device *smmu, u32 val,
> }
> 
> /* GBPA is "special" */
> -static int __init arm_smmu_update_gbpa(struct arm_smmu_device *smmu,
> -                                       u32 set, u32 clr)
> +static int __init_or_smmu_suspend
> +arm_smmu_update_gbpa(struct arm_smmu_device *smmu, u32 set, u32 clr)
> {
> int ret;
> u32 reg, __iomem *gbpa = smmu->base + ARM_SMMU_GBPA;
> @@ -1995,10 +2001,35 @@ err_free_evtq_irq:
> return ret;
> }
> 
> +static int arm_smmu_enable_irqs(struct arm_smmu_device *smmu)
> +{
> + int ret;
> + u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
> +
> + if ( smmu->features & ARM_SMMU_FEAT_PRI )
> + irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
> +
> + /* Enable interrupt generation on the SMMU */
> + ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
> +      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
> + if ( ret )
> + {
> + dev_warn(smmu->dev, "failed to enable irqs\n");
> + return ret;
> + }
> +
> + return 0;
> +}
> +
> +/*
> + * Probe-time only: request host IRQs and, when available, program the SMMU's
> + * MSI doorbells. Resume does not restore the SMMU *_IRQ_CFGn MSI registers,
> + * so any host suspend support must treat the active MSI IRQ path as
> + * unsupported until that restore path exists.
> + */
> static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
> {
> int ret, irq;
> - u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
> 
> /* Disable IRQs first */
> ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
> @@ -2028,22 +2059,7 @@ static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
> }
> }
> 
> - if (smmu->features & ARM_SMMU_FEAT_PRI)
> - irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
> -
> - /* Enable interrupt generation on the SMMU */
> - ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
> -      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
> - if (ret) {
> - dev_warn(smmu->dev, "failed to enable irqs\n");
> - goto err_free_irqs;
> - }
> -
> return 0;
> -
> -err_free_irqs:
> - arm_smmu_free_irqs(smmu);
> - return ret;
> }
> 
> static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
> @@ -2057,7 +2073,8 @@ static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
> return ret;
> }
> 
> -static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> +static int __init_or_smmu_suspend
> +arm_smmu_device_reset(struct arm_smmu_device *smmu)
> {
> int ret;
> u32 reg, enables;
> @@ -2163,17 +2180,9 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> }
> }
> 
> - ret = arm_smmu_setup_irqs(smmu);
> - if (ret) {
> - dev_err(smmu->dev, "failed to setup irqs\n");
> + ret = arm_smmu_enable_irqs(smmu);
> + if ( ret )
> return ret;
> - }
> -
> - /* Initialize tasklets for threaded IRQs*/
> - tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
> - tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
> - tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
> - smmu);
> 
> /* Enable the SMMU interface, or ensure bypass */
> if (disable_bypass) {
> @@ -2181,20 +2190,16 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> } else {
> ret = arm_smmu_update_gbpa(smmu, 0, GBPA_ABORT);
> if (ret)
> - goto err_free_irqs;
> + return ret;
> }
> ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
>      ARM_SMMU_CR0ACK);
> if (ret) {
> dev_err(smmu->dev, "failed to enable SMMU interface\n");
> - goto err_free_irqs;
> + return ret;
> }
> 
> return 0;
> -
> -err_free_irqs:
> - arm_smmu_free_irqs(smmu);
> - return ret;
> }
> 
> static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
> @@ -2558,10 +2563,23 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
> if (ret)
> goto out_free;
> 
> + ret = arm_smmu_setup_irqs(smmu);
> + if ( ret )
> + {
> + dev_err(smmu->dev, "failed to setup irqs\n");
> + goto out_free;
> + }
> +
> + /* Initialize tasklets for threaded IRQs*/
> + tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
> + tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
> + tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
> + smmu);
> +
> /* Reset the device */
> ret = arm_smmu_device_reset(smmu);
> if (ret)
> - goto out_free;
> + goto out_free_irqs;
> 
> /*
> * Keep a list of all probed devices. This will be used to query
> @@ -2575,6 +2593,8 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
> 
> return 0;
> 
> +out_free_irqs:
> + arm_smmu_free_irqs(smmu);
> 
> out_free:
> arm_smmu_free_structures(smmu);
> @@ -2855,6 +2875,104 @@ static void arm_smmu_iommu_xen_domain_teardown(struct domain *d)
> xfree(xen_domain);
> }
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +
> +static void arm_smmu_reset_for_suspend_rollback(struct arm_smmu_device *smmu)
> +{
> + int ret = arm_smmu_device_reset(smmu);
> +
> + if ( ret )
> + dev_err(smmu->dev, "Failed to reset during suspend rollback: %d\n",
> + ret);

If reset fails here, we only print an error.

Could the SMMU be left disabled with GBPA.ABORT cleared, allowing guest
DMA to bypass translation when the domains resume?


> +}
> +
> +static int arm_smmu_suspend(void)
> +{
> + struct arm_smmu_device *smmu;
> + int ret = 0;
> +
> + list_for_each_entry(smmu, &arm_smmu_devices, devices)
> + {
> + /* Abort all transactions before disable to avoid spurious bypass */
> + ret = arm_smmu_update_gbpa(smmu, GBPA_ABORT, 0);
> + if ( ret )
> + goto fail;
> +
> + ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
> + ARM_SMMU_IRQ_CTRLACK);
> + if ( ret )
> + {
> + dev_err(smmu->dev, "Timed-out while disabling SMMU irqs\n");
> + goto fail;
> + }
> +
> + /* Disable the SMMU via CR0.EN and all queues except CMDQ */
> + ret = arm_smmu_write_reg_sync(smmu, CR0_CMDQEN, ARM_SMMU_CR0,
> + ARM_SMMU_CR0ACK);
> + if ( ret )
> + {
> + dev_err(smmu->dev, "Timed-out while disabling smmu\n");
> + goto fail;
> + }
> +
> + /*
> + * At this point the translation interface is disabled and the
> + * SMMU won't access translation/config structures, even
> + * speculatively, as per the IHI0070 spec (section 6.3.9.6).
> + * CMDQ is still enabled so that a CMD_SYNC can complete any
> + * previously issued commands.
> + */
> +
> + /* Ensure all previously issued commands have completed. */
> + ret = arm_smmu_cmdq_issue_sync(smmu);
> + if ( ret )
> + {
> + dev_err(smmu->dev, "Timed-out waiting for pending commands\n");
> + goto fail;
> + }

Could we lose EVTQ events here because we do not check the queue after
stopping it?

Cheers
Bertrand

> +
> + /* Disable everything */
> + ret = arm_smmu_device_disable(smmu);
> + if ( ret )
> + goto fail;
> +
> + dev_dbg(smmu->dev, "Suspended smmu\n");
> + }
> +
> + return 0;
> +
> + fail:
> + /* Reset the device that failed as well as any already-suspended ones. */
> + arm_smmu_reset_for_suspend_rollback(smmu);
> +
> + list_for_each_entry_continue_reverse(smmu, &arm_smmu_devices, devices)
> + arm_smmu_reset_for_suspend_rollback(smmu);
> +
> + return ret;
> +}
> +
> +static void arm_smmu_resume(void)
> +{
> + int ret;
> + struct arm_smmu_device *smmu;
> +
> + list_for_each_entry(smmu, &arm_smmu_devices, devices)
> + {
> + dev_dbg(smmu->dev, "Resuming device\n");
> +
> + /*
> + * The reset will re-initialize all the base addresses, queues,
> + * prod and cons maintained within struct arm_smmu_device as well as
> + * re-enable the interrupts.
> + */
> + ret = arm_smmu_device_reset(smmu);
> + if ( ret )
> + panic("SMMUv3: %s: Failed to reset during resume: %d\n",
> +      dev_name(smmu->dev), ret);
> + }
> +}
> +#endif
> +
> static const struct iommu_ops arm_smmu_iommu_ops = {
> .page_sizes = PAGE_SIZE_4K,
> .init = arm_smmu_iommu_xen_domain_init,
> @@ -2867,6 +2985,10 @@ static const struct iommu_ops arm_smmu_iommu_ops = {
> .unmap_page = arm_iommu_unmap_page,
> .dt_xlate = arm_smmu_dt_xlate,
> .add_device = arm_smmu_add_device,
> +#ifdef CONFIG_SYSTEM_SUSPEND
> + .suspend = arm_smmu_suspend,
> + .resume = arm_smmu_resume,
> +#endif
> };
> 
> static __init int arm_smmu_dt_init(struct dt_device_node *dev,
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND
  2026-08-27 14:31 ` [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND Mykola Kvach
@ 2026-09-28 16:17   ` Bertrand Marquis
  0 siblings, 0 replies; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28 16:17 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Oleksandr Tyshchenko,
	Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> From: Mirela Simonovic <mirela.simonovic@aggios.com>
> 
> On wakeup from PSCI SYSTEM_SUSPEND, Xen re-enters EL2 with the MMU and
> data cache disabled. The resume path must first switch back to Xen's
> runtime page tables before it can access the saved CPU context using
> virtual addresses.
> 
> Add an arm64 hyp_resume trampoline that reuses enable_secondary_cpu_mm()
> to enable the data cache and MMU, switch to init_ttbr, and resume in the
> runtime virtual mapping. The trampoline then restores the saved CPU
> general-purpose and system-control register context.
> 
> prepare_resume_ctx() must be invoked just before the PSCI system suspend
> call is issued to the platform firmware. It saves the current CPU context
> and returns a non-zero value so that the caller enters the physical
> SYSTEM_SUSPEND call.
> 
> On resume, hyp_resume restores the saved context, including the saved link
> register. Control therefore returns to the place where prepare_resume_ctx()
> was called. To avoid re-entering the suspend path, the restored path sees
> prepare_resume_ctx() return zero.
> 
> The assembly save/restore code uses offsets generated by asm-offsets.c
> from struct resume_cpu_context, keeping the assembly memory accesses in
> sync with the C structure layout.
> 
> Support for ARM32 is not implemented. Instead, compilation fails with a
> build-time error if suspend is enabled for ARM32.
> 
> Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>

Reviewed-by: Bertrand Marquis <bertrand.marquis@arm.com>

Cheers
Bertrand

> ---
> Changes in v10:
> - Save and restore CNTHCTL_EL2 across SYSTEM_SUSPEND
> 
> Changes in v9:
> - Drop the misleading prepare_resume_ctx() pointer argument and make both
>  save/restore paths use the global resume_cpu_context.
> - Squash the arm64 resume trampoline into the context save/restore patch.
> - Document in code that hyp_resume relies on PSCI initial-state rules.
> - Use generic platform firmware wording instead of ATF-specific wording.
> - Rename the saved context type/storage to resume_cpu_context and rely on
>  implicit zero-initialization for the file-scope object.
> - Use asm-offsets.c-generated RESUME_CTX_* offsets to keep the assembly
>  save/restore code in sync with struct resume_cpu_context.
> 
> Changes in v8:
> - Fix alignments in code.
> 
> Changes in v7:
> - No functional changes, just moved commit.
> ---
> xen/arch/arm/Makefile              |   1 +
> xen/arch/arm/arm64/asm-offsets.c   |  21 +++++
> xen/arch/arm/arm64/head.S          | 122 +++++++++++++++++++++++++++++
> xen/arch/arm/include/asm/suspend.h |  27 +++++++
> xen/arch/arm/suspend.c             |  14 ++++
> 5 files changed, 185 insertions(+)
> create mode 100644 xen/arch/arm/suspend.c
> 
> diff --git a/xen/arch/arm/Makefile b/xen/arch/arm/Makefile
> index b7afd3e58c..788db83ba9 100644
> --- a/xen/arch/arm/Makefile
> +++ b/xen/arch/arm/Makefile
> @@ -51,6 +51,7 @@ obj-y += setup.o
> obj-y += shutdown.o
> obj-y += smp.o
> obj-y += smpboot.o
> +obj-$(CONFIG_SYSTEM_SUSPEND) += suspend.o
> obj-$(CONFIG_SYSCTL) += sysctl.o
> obj-y += time.o
> obj-y += traps.o
> diff --git a/xen/arch/arm/arm64/asm-offsets.c b/xen/arch/arm/arm64/asm-offsets.c
> index 38a3894a3b..5d60406e9c 100644
> --- a/xen/arch/arm/arm64/asm-offsets.c
> +++ b/xen/arch/arm/arm64/asm-offsets.c
> @@ -13,6 +13,7 @@
> #include <asm/mm.h>
> #include <asm/setup.h>
> #include <asm/smccc.h>
> +#include <asm/suspend.h>
> 
> #define DEFINE(_sym, _val)                                                 \
>     asm volatile ( "\n.ascii\"==>#define " #_sym " %0 /* " #_val " */<==\""\
> @@ -57,6 +58,26 @@ void __dummy__(void)
>    OFFSET(INITINFO_stack, struct init_info, stack);
>    BLANK();
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +   OFFSET(RESUME_CTX_X19, struct resume_cpu_context, callee_regs[0]);
> +   OFFSET(RESUME_CTX_X21, struct resume_cpu_context, callee_regs[2]);
> +   OFFSET(RESUME_CTX_X23, struct resume_cpu_context, callee_regs[4]);
> +   OFFSET(RESUME_CTX_X25, struct resume_cpu_context, callee_regs[6]);
> +   OFFSET(RESUME_CTX_X27, struct resume_cpu_context, callee_regs[8]);
> +   OFFSET(RESUME_CTX_X29, struct resume_cpu_context, callee_regs[10]);
> +   OFFSET(RESUME_CTX_SP, struct resume_cpu_context, sp);
> +   OFFSET(RESUME_CTX_VBAR_EL2, struct resume_cpu_context, vbar_el2);
> +   OFFSET(RESUME_CTX_VTCR_EL2, struct resume_cpu_context, vtcr_el2);
> +   OFFSET(RESUME_CTX_VTTBR_EL2, struct resume_cpu_context, vttbr_el2);
> +   OFFSET(RESUME_CTX_TPIDR_EL2, struct resume_cpu_context, tpidr_el2);
> +   OFFSET(RESUME_CTX_MDCR_EL2, struct resume_cpu_context, mdcr_el2);
> +   OFFSET(RESUME_CTX_HSTR_EL2, struct resume_cpu_context, hstr_el2);
> +   OFFSET(RESUME_CTX_CPTR_EL2, struct resume_cpu_context, cptr_el2);
> +   OFFSET(RESUME_CTX_HCR_EL2, struct resume_cpu_context, hcr_el2);
> +   OFFSET(RESUME_CTX_CNTHCTL_EL2, struct resume_cpu_context, cnthctl_el2);
> +   BLANK();
> +#endif
> +
>    OFFSET(SMCCC_RES_a0, struct arm_smccc_res, a0);
>    OFFSET(SMCCC_RES_a2, struct arm_smccc_res, a2);
>    OFFSET(ARM_SMCCC_1_2_REGS_X0_OFFS, struct arm_smccc_1_2_regs, a0);
> diff --git a/xen/arch/arm/arm64/head.S b/xen/arch/arm/arm64/head.S
> index 72c7b24498..962be716ae 100644
> --- a/xen/arch/arm/arm64/head.S
> +++ b/xen/arch/arm/arm64/head.S
> @@ -561,6 +561,128 @@ END(efi_xen_start)
> 
> #endif /* CONFIG_ARM_EFI */
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +/*
> + * int prepare_resume_ctx(void)
> + *
> + * CPU context saved here will be restored on resume in hyp_resume function.
> + * prepare_resume_ctx shall return a non-zero value. Upon restoring context
> + * hyp_resume shall return value zero instead. From C code that invokes
> + * prepare_resume_ctx, the return value is interpreted to determine whether
> + * the context is saved (prepare_resume_ctx) or restored (hyp_resume).
> + */
> +FUNC(prepare_resume_ctx)
> +        ldr   x0, =resume_cpu_context
> +
> +        /* Store callee-saved registers */
> +        stp   x19, x20, [x0, #RESUME_CTX_X19]
> +        stp   x21, x22, [x0, #RESUME_CTX_X21]
> +        stp   x23, x24, [x0, #RESUME_CTX_X23]
> +        stp   x25, x26, [x0, #RESUME_CTX_X25]
> +        stp   x27, x28, [x0, #RESUME_CTX_X27]
> +        stp   x29, lr, [x0, #RESUME_CTX_X29]
> +
> +        /* Store stack-pointer */
> +        mov   x2, sp
> +        str   x2, [x0, #RESUME_CTX_SP]
> +
> +        /* Store system control registers */
> +        mrs   x2, VBAR_EL2
> +        str   x2, [x0, #RESUME_CTX_VBAR_EL2]
> +        mrs   x2, VTCR_EL2
> +        str   x2, [x0, #RESUME_CTX_VTCR_EL2]
> +        mrs   x2, VTTBR_EL2
> +        str   x2, [x0, #RESUME_CTX_VTTBR_EL2]
> +        mrs   x2, TPIDR_EL2
> +        str   x2, [x0, #RESUME_CTX_TPIDR_EL2]
> +        mrs   x2, MDCR_EL2
> +        str   x2, [x0, #RESUME_CTX_MDCR_EL2]
> +        mrs   x2, HSTR_EL2
> +        str   x2, [x0, #RESUME_CTX_HSTR_EL2]
> +        mrs   x2, CPTR_EL2
> +        str   x2, [x0, #RESUME_CTX_CPTR_EL2]
> +        mrs   x2, HCR_EL2
> +        str   x2, [x0, #RESUME_CTX_HCR_EL2]
> +        mrs   x2, CNTHCTL_EL2
> +        str   x2, [x0, #RESUME_CTX_CNTHCTL_EL2]
> +
> +        /* prepare_resume_ctx must return a non-zero value */
> +        mov   x0, #1
> +        ret
> +END(prepare_resume_ctx)
> +
> +FUNC(hyp_resume)
> +        /*
> +         * PSCI states that SYSTEM_SUSPEND follows the CPU_SUSPEND initial
> +         * state rules, so PSCI-compliant firmware must enter the return
> +         * exception level with DAIF masked.
> +         */
> +
> +        /* Initialize the UART if earlyprintk has been enabled. */
> +#ifdef CONFIG_EARLY_PRINTK
> +        bl    init_uart
> +#endif
> +        PRINT_ID("- Xen resuming -\r\n")
> +
> +        bl    check_cpu_mode
> +        bl    cpu_init
> +
> +        ldr   x0, =start
> +        adr   x20, start             /* x20 := paddr (start) */
> +        sub   x20, x20, x0           /* x20 := phys-offset */
> +        ldr   lr, =mmu_resumed
> +        b     enable_secondary_cpu_mm
> +
> +mmu_resumed:
> +        /* Now we can access the saved context, so restore it here. */
> +        ldr   x0, =resume_cpu_context
> +
> +        /* Restore callee-saved registers */
> +        ldp   x19, x20, [x0, #RESUME_CTX_X19]
> +        ldp   x21, x22, [x0, #RESUME_CTX_X21]
> +        ldp   x23, x24, [x0, #RESUME_CTX_X23]
> +        ldp   x25, x26, [x0, #RESUME_CTX_X25]
> +        ldp   x27, x28, [x0, #RESUME_CTX_X27]
> +        ldp   x29, lr, [x0, #RESUME_CTX_X29]
> +
> +        /* Restore stack pointer */
> +        ldr   x2, [x0, #RESUME_CTX_SP]
> +        mov   sp, x2
> +
> +        /* Restore system control registers */
> +        ldr   x2, [x0, #RESUME_CTX_VBAR_EL2]
> +        msr   VBAR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_VTCR_EL2]
> +        msr   VTCR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_VTTBR_EL2]
> +        msr   VTTBR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_TPIDR_EL2]
> +        msr   TPIDR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_MDCR_EL2]
> +        msr   MDCR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_HSTR_EL2]
> +        msr   HSTR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_CPTR_EL2]
> +        msr   CPTR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_HCR_EL2]
> +        msr   HCR_EL2, x2
> +        ldr   x2, [x0, #RESUME_CTX_CNTHCTL_EL2]
> +        msr   CNTHCTL_EL2, x2
> +        isb
> +
> +        /*
> +         * Since context is restored return from this function will appear
> +         * as return from prepare_resume_ctx. To distinguish a return from
> +         * prepare_resume_ctx which is called upon finalizing the suspend,
> +         * as opposed to return from this function which executes on resume,
> +         * we need to return zero value here.
> +         */
> +        mov   x0, #0
> +        ret
> +END(hyp_resume)
> +
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> /*
>  * Local variables:
>  * mode: ASM
> diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> index 31a98a1f1b..c848fc6340 100644
> --- a/xen/arch/arm/include/asm/suspend.h
> +++ b/xen/arch/arm/include/asm/suspend.h
> @@ -3,6 +3,8 @@
> #ifndef ARM_SUSPEND_H
> #define ARM_SUSPEND_H
> 
> +#include <xen/types.h>
> +
> struct domain;
> struct vcpu;
> struct vcpu_guest_context;
> @@ -14,6 +16,31 @@ struct resume_info {
> 
> void arch_domain_resume(struct domain *d);
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +#ifdef CONFIG_ARM_64
> +struct resume_cpu_context {
> +    register_t callee_regs[12];
> +    register_t sp;
> +    register_t vbar_el2;
> +    register_t vtcr_el2;
> +    register_t vttbr_el2;
> +    register_t tpidr_el2;
> +    register_t mdcr_el2;
> +    register_t hstr_el2;
> +    register_t cptr_el2;
> +    register_t hcr_el2;
> +    register_t cnthctl_el2;
> +} __aligned(16);
> +#else
> +#error "Define resume_cpu_context structure for arm32"
> +#endif
> +
> +extern struct resume_cpu_context resume_cpu_context;
> +
> +int prepare_resume_ctx(void);
> +void hyp_resume(void);
> +#endif /* CONFIG_SYSTEM_SUSPEND */
> +
> #endif /* ARM_SUSPEND_H */
> 
> /*
> diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> new file mode 100644
> index 0000000000..6ea4a0f9cc
> --- /dev/null
> +++ b/xen/arch/arm/suspend.c
> @@ -0,0 +1,14 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +
> +#include <asm/suspend.h>
> +
> +struct resume_cpu_context resume_cpu_context;
> +
> +/*
> + * Local variables:
> + * mode: C
> + * c-file-style: "BSD"
> + * c-basic-offset: 4
> + * indent-tabs-mode: nil
> + * End:
> + */
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface)
  2026-08-27 14:31 ` [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface) Mykola Kvach
@ 2026-09-28 16:18   ` Bertrand Marquis
  2026-09-30 17:32     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28 16:18 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Luca Fancellu

Hi Mykola,

> On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> From: Mirela Simonovic <mirela.simonovic@aggios.com>
> 
> Invoke PSCI SYSTEM_SUSPEND to finalize Xen's suspend sequence on ARM64
> platforms. Pass the Xen resume entry point (hyp_resume) to EL3 together
> with a zero context ID, matching Linux.
> 
> This patch wires up only the host-side PSCI SYSTEM_SUSPEND invocation.
> The resume trampoline and context restore are provided by earlier patches
> in the series.
> 
> Only enable this path when CONFIG_SYSTEM_SUSPEND is set and PSCI
> advertises SYSTEM_SUSPEND via PSCI_FEATURES.
> 
> Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> ---
> Changes in v9:
> - cache SYSTEM_SUSPEND support using PSCI_FEATURES and gate the host call
>  on the cached capability
> - keep the cached SYSTEM_SUSPEND capability read-only after init
> - log whether firmware reports SYSTEM_SUSPEND support
> - pass an explicit zero context ID in the SYSTEM_SUSPEND call
> - drop the stale note claiming hyp_resume is still a stub
> ---
> xen/arch/arm/include/asm/psci.h |  1 +
> xen/arch/arm/psci.c             | 31 ++++++++++++++++++++++++++++++-
> 2 files changed, 31 insertions(+), 1 deletion(-)
> 
> diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
> index 48a93e6b79..bb3c73496e 100644
> --- a/xen/arch/arm/include/asm/psci.h
> +++ b/xen/arch/arm/include/asm/psci.h
> @@ -23,6 +23,7 @@ int call_psci_cpu_on(int cpu);
> void call_psci_cpu_off(void);
> void call_psci_system_off(void);
> void call_psci_system_reset(void);
> +int call_psci_system_suspend(void);
> 
> /* Range of allocated PSCI function numbers */
> #define PSCI_FNUM_MIN_VALUE                 _AC(0,U)
> diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
> index b6860a7760..e05dae1133 100644
> --- a/xen/arch/arm/psci.c
> +++ b/xen/arch/arm/psci.c
> @@ -17,23 +17,27 @@
> #include <asm/cpufeature.h>
> #include <asm/psci.h>
> #include <asm/acpi.h>
> +#include <asm/suspend.h>
> 
> /*
>  * While a 64-bit OS can make calls with SMC32 calling conventions, for
>  * some calls it is necessary to use SMC64 to pass or return 64-bit values.
> - * For such calls PSCI_0_2_FN_NATIVE(x) will choose the appropriate
> + * For such calls PSCI_*_FN_NATIVE(x) will choose the appropriate
>  * (native-width) function ID.
>  */
> #ifdef CONFIG_ARM_64
> #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN64_##name
> +#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN64_##name
> #else
> #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN32_##name
> +#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN32_##name
> #endif
> 
> uint32_t psci_ver;
> uint32_t smccc_ver;
> 
> static uint32_t psci_cpu_on_nr;
> +static bool __ro_after_init has_psci_system_suspend;
> 
> #define PSCI_RET(res)   ((int32_t)(res).a0)
> 
> @@ -60,6 +64,25 @@ void call_psci_cpu_off(void)
>     }
> }
> 
> +int call_psci_system_suspend(void)
> +{
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    struct arm_smccc_res res;
> +
> +    if ( !has_psci_system_suspend )
> +        return PSCI_NOT_SUPPORTED;
> +
> +    /* Context ID is unused for the Xen resume path. */
> +    arm_smccc_smc(PSCI_1_0_FN_NATIVE(SYSTEM_SUSPEND), __pa(hyp_resume), 0,
> +                  &res);

Since this series was sent, arm_smccc_smc() has been changed to return
the result directly instead of writing it through a result pointer.

You need to update this call.

Cheers
Bertrand


> +    return PSCI_RET(res);
> +#else
> +    dprintk(XENLOG_WARNING,
> +            "SYSTEM_SUSPEND not supported (CONFIG_SYSTEM_SUSPEND disabled)\n");
> +    return PSCI_NOT_SUPPORTED;
> +#endif
> +}
> +
> void call_psci_system_off(void)
> {
>     if ( psci_ver > PSCI_VERSION(0, 1) )
> @@ -223,9 +246,15 @@ int __init psci_init(void)
> 
>     psci_init_smccc();
> 
> +    has_psci_system_suspend =
> +        psci_features(PSCI_1_0_FN_NATIVE(SYSTEM_SUSPEND)) == 0;
> +
>     printk(XENLOG_INFO "Using PSCI v%u.%u\n",
>            PSCI_VERSION_MAJOR(psci_ver), PSCI_VERSION_MINOR(psci_ver));
> 
> +    printk(XENLOG_DEBUG "PSCI SYSTEM_SUSPEND is %ssupported by firmware\n",
> +           has_psci_system_suspend ? "" : "not ");
> +
>     return 0;
> }
> 
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy
  2026-08-27 14:32 ` [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy Mykola Kvach
@ 2026-09-28 16:18   ` Bertrand Marquis
  2026-09-30 20:42     ` Mykola Kvach
  0 siblings, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28 16:18 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk, Andrew Cooper, Anthony PERARD,
	Jan Beulich, Roger Pau Monné, Rahul Singh,
	Oleksandr Tyshchenko

Hi Mykola,

> On 27 Aug 2026, at 16:32, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> Introduce CONFIG_HAS_HWDOM_SYSTEM_SUSPEND as an architecture-selected
> capability for platforms where the hardware domain can be parked with
> SHUTDOWN_suspend without calling hwdom_shutdown().
> 
> Expose PSCI SYSTEM_SUSPEND as a vPSCI operation for all domains. For
> non-control domains, including the hardware domain when it is not acting
> as a control domain, the call is handled as a guest/domain suspend request
> and parks the domain in SHUTDOWN_suspend.
> 
> Control domains need additional sequencing because their SYSTEM_SUSPEND
> request is used to coordinate host-wide suspend. A non-last awake control
> domain may be parked in SHUTDOWN_suspend without requiring the host
> suspend path to be available. The last awake control domain is treated as
> the point where the request becomes a host-suspend request, and it may
> only proceed when all non-control domains are already in SHUTDOWN_suspend
> and the host suspend path is available.
> 
> Keep the control-domain sequencing and domain-readiness checks out of
> PSCI_FEATURES. They are per-attempt runtime conditions rather than stable
> PSCI function availability. Advertise SYSTEM_SUSPEND as implemented by
> vPSCI and report attempt-time policy failures as PSCI_DENIED.
> 
> Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
> so that SHUTDOWN_suspend from the hardware domain can be treated as a
> domain suspend state rather than as a hardware-domain initiated host
> shutdown. This does not by itself imply that host-wide suspend is
> available.
> 
> Add host_system_suspend_allowed() to combine the host PSCI SYSTEM_SUSPEND
> capability with runtime blockers reported by Xen-owned subsystems. Add
> runtime blockers for registered serial, IOMMU, GIC and SMMUv3 MSI IRQ
> paths lacking suspend/resume support. These blockers are runtime based,
> so they only apply to drivers or paths that Xen actually uses on the
> platform. For SMMUv3, the blocker applies only when Xen actually uses the
> MSI IRQ path, since resume does not restore the SMMU *_IRQ_CFGn MSI
> registers yet.
> 
> Add a struct domain forward declaration to xen/suspend.h so the generic
> header can expose arch_domain_resume() without requiring a full domain.h
> include.
> 
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> Reviewed-by: Oleksandr Tyshchenko <Oleksandr_Tyshchenko@epam.com>
> ---
> Changes in V12:
> - handle missing is_shut_down, change checking to call of
>  domain_shutdown_completed
> 
> Changes in V11:
> - Mark host_system_suspend_runtime_allowed as __ro_after_init.
> - Avoid printing the SMMUv3 MSI IRQ host suspend blocker more than once
>  when multiple SMMUv3 instances use MSIs.
> - Wrap the Arm IOMMU host suspend blocker in CONFIG_SYSTEM_SUSPEND to make
>  its policy-only use explicit.
> 
> Changes in V10:
> - Return PSCI_DENIED rather than PSCI_NOT_SUPPORTED when the last awake
>  control domain cannot proceed to host suspend, keeping PSCI_FEATURES
>  stable once SYSTEM_SUSPEND is advertised.
> - Shorten SYSTEM_SUSPEND blocker messages and use %pd when logging the
>  control domain.
> - Mark serial_suspend_available as __ro_after_init.
> - Mention the struct domain forward declaration added to xen/suspend.h.
> 
> Changes in V9:
> - Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
>  so that hardware-domain SHUTDOWN_suspend support is not tied to
>  host-wide system suspend availability.
> - Add runtime host suspend blockers for Xen-owned subsystems lacking
>  suspend/resume support.
> - Keep vPSCI SYSTEM_SUSPEND advertised through PSCI_FEATURES and enforce
>  control-domain sequencing in the call handler.
> ---
> xen/arch/arm/Kconfig                  |   1 +
> xen/arch/arm/gic.c                    |   6 ++
> xen/arch/arm/include/asm/psci.h       |   3 +
> xen/arch/arm/include/asm/suspend.h    |  10 ++-
> xen/arch/arm/psci.c                   |   7 ++
> xen/arch/arm/suspend.c                |  40 +++++++++
> xen/arch/arm/vpsci.c                  | 114 +++++++++++++++++++++++---
> xen/common/Kconfig                    |   3 +
> xen/common/domain.c                   |   7 +-
> xen/drivers/char/serial.c             |  12 +++
> xen/drivers/passthrough/arm/iommu.c   |   6 ++
> xen/drivers/passthrough/arm/smmu-v3.c |   9 ++
> xen/include/xen/serial.h              |   1 +
> xen/include/xen/suspend.h             |   2 +
> 14 files changed, 208 insertions(+), 13 deletions(-)
> 
> diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
> index 843a43897e..9027aa17eb 100644
> --- a/xen/arch/arm/Kconfig
> +++ b/xen/arch/arm/Kconfig
> @@ -19,6 +19,7 @@ config ARM
> select HAS_ALTERNATIVE if HAS_VMAP
> select HAS_DEVICE_TREE_DISCOVERY
> select HAS_DOM0LESS
> + select HAS_HWDOM_SYSTEM_SUSPEND if !MPU
> select HAS_GRANT_CACHE_FLUSH if GRANT_TABLE
> select HAS_STACK_PROTECTOR
> select HAS_STATIC_MEMORY
> diff --git a/xen/arch/arm/gic.c b/xen/arch/arm/gic.c
> index ffc11f36a1..0695474432 100644
> --- a/xen/arch/arm/gic.c
> +++ b/xen/arch/arm/gic.c
> @@ -26,6 +26,7 @@
> #include <asm/device.h>
> #include <asm/io.h>
> #include <asm/gic.h>
> +#include <asm/suspend.h>
> #include <asm/vgic.h>
> #include <asm/acpi.h>
> 
> @@ -44,6 +45,11 @@ static void __init __maybe_unused build_assertions(void)
> void register_gic_ops(const struct gic_hw_operations *ops)
> {
>     gic_hw_ops = ops;
> +
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    if ( !ops->suspend || !ops->resume )
> +        host_system_suspend_disable("GIC driver lacks suspend support");
> +#endif
> }
> 
> static void clear_cpu_lr_mask(void)
> diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
> index bb3c73496e..142fa1bfe5 100644
> --- a/xen/arch/arm/include/asm/psci.h
> +++ b/xen/arch/arm/include/asm/psci.h
> @@ -24,6 +24,9 @@ void call_psci_cpu_off(void);
> void call_psci_system_off(void);
> void call_psci_system_reset(void);
> int call_psci_system_suspend(void);
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +bool psci_system_suspend_allowed(void);
> +#endif
> 
> /* Range of allocated PSCI function numbers */
> #define PSCI_FNUM_MIN_VALUE                 _AC(0,U)
> diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> index c848fc6340..50dc6e9fdf 100644
> --- a/xen/arch/arm/include/asm/suspend.h
> +++ b/xen/arch/arm/include/asm/suspend.h
> @@ -39,7 +39,15 @@ extern struct resume_cpu_context resume_cpu_context;
> 
> int prepare_resume_ctx(void);
> void hyp_resume(void);
> -#endif /* CONFIG_SYSTEM_SUSPEND */
> +bool host_system_suspend_allowed(void);
> +void host_system_suspend_disable(const char *reason);
> +
> +#else /* !CONFIG_SYSTEM_SUSPEND */
> +
> +static inline bool host_system_suspend_allowed(void) { return false; }
> +static inline void host_system_suspend_disable(const char *reason) {}
> +
> +#endif
> 
> #endif /* ARM_SUSPEND_H */
> 
> diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
> index e05dae1133..e9d78668fd 100644
> --- a/xen/arch/arm/psci.c
> +++ b/xen/arch/arm/psci.c
> @@ -41,6 +41,13 @@ static bool __ro_after_init has_psci_system_suspend;
> 
> #define PSCI_RET(res)   ((int32_t)(res).a0)
> 
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +bool psci_system_suspend_allowed(void)
> +{
> +    return has_psci_system_suspend;
> +}
> +#endif
> +
> int call_psci_cpu_on(int cpu)
> {
>     struct arm_smccc_res res;
> diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> index 6ea4a0f9cc..c7c26bcf03 100644
> --- a/xen/arch/arm/suspend.c
> +++ b/xen/arch/arm/suspend.c
> @@ -1,9 +1,49 @@
> /* SPDX-License-Identifier: GPL-2.0-only */
> 
> +#include <asm/psci.h>
> #include <asm/suspend.h>
> 
> +#include <xen/lib.h>
> +#include <xen/serial.h>
> +
> struct resume_cpu_context resume_cpu_context;
> 
> +/*
> + * Non-PSCI infrastructure can make host suspend impossible even when the PSCI
> + * SYSTEM_SUSPEND conduit is present, e.g. when a Xen-owned driver has no valid
> + * suspend/resume path.
> + *
> + * This gate is checked only when the last awake control domain attempts to
> + * turn a guest SYSTEM_SUSPEND request into a host-suspend request.
> + */
> +static bool __ro_after_init host_system_suspend_runtime_allowed = true;
> +
> +static bool host_serial_suspend_allowed(void)
> +{
> +    if ( serial_suspend_supported() )
> +        return true;
> +
> +    printk_once(XENLOG_INFO
> +                "Host SYSTEM_SUSPEND blocked: serial unsupported\n");
> +
> +    return false;
> +}
> +
> +bool host_system_suspend_allowed(void)
> +{
> +    return psci_system_suspend_allowed() &&
> +           host_serial_suspend_allowed() &&
> +           host_system_suspend_runtime_allowed;
> +}
> +
> +void host_system_suspend_disable(const char *reason)
> +{
> +    host_system_suspend_runtime_allowed = false;
> +
> +    printk(XENLOG_INFO "Host SYSTEM_SUSPEND blocked: %s\n",
> +           reason ? reason : "unsupported suspend/resume path");
> +}
> +
> /*
>  * Local variables:
>  * mode: C
> diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
> index ac6af6118f..a41355d75d 100644
> --- a/xen/arch/arm/vpsci.c
> +++ b/xen/arch/arm/vpsci.c
> @@ -5,6 +5,7 @@
> 
> #include <asm/current.h>
> #include <asm/domain.h>
> +#include <asm/suspend.h>
> #include <asm/vgic.h>
> #include <asm/vpsci.h>
> #include <asm/event.h>
> @@ -219,6 +220,89 @@ static void do_psci_0_2_system_reset(void)
>     domain_shutdown(d,SHUTDOWN_reboot);
> }
> 
> +/*
> + * Serialise SYSTEM_SUSPEND policy decisions with the domain suspend transition,
> + * so multiple control domains cannot all observe each other as still awake.
> + */
> +static DEFINE_SPINLOCK(vpsci_system_suspend_lock);
> +
> +static bool domain_in_suspend_state(struct domain *d)
> +{
> +    bool suspended;
> +
> +    spin_lock(&d->shutdown_lock);
> +    suspended = domain_shutdown_completed(d) && (d->shutdown_code == SHUTDOWN_suspend);

This lines is over 80 chars and should be broken down.

> +    spin_unlock(&d->shutdown_lock);
> +
> +    return suspended;
> +}
> +
> +static int32_t domain_psci_system_suspend_policy(struct domain *d)
> +{
> +    struct domain *other;
> +    bool last_awake_control_domain = true;
> +    bool awake_non_control_domain = false;
> +
> +    /* Only control domains participate in sequencing policy. */
> +    if ( !is_control_domain(d) )
> +        return 0;

I am wondering what would happen in a dom0less setup with no control
domain if the domains request SYSTEM_SUSPEND.

Could you explain how they would be resumed?
Should we deny SYSTEM_SUSPEND requests in this case?

Cheers
Bertrand

> +
> +    rcu_read_lock(&domlist_read_lock);
> +
> +    for_each_domain ( other )
> +    {
> +        bool suspended;
> +
> +        if ( other == d )
> +            continue;
> +
> +        suspended = domain_in_suspend_state(other);
> +        if ( suspended )
> +            continue;
> +
> +        if ( is_control_domain(other) )
> +        {
> +            last_awake_control_domain = false;
> +            break;
> +        }
> +
> +        awake_non_control_domain = true;
> +    }
> +
> +    rcu_read_unlock(&domlist_read_lock);
> +
> +    /*
> +     * Another control domain is still awake. This request is only the first
> +     * phase of the sequencing: park this control domain and leave the host
> +     * running. Host-wide suspend gates must not block this intermediate state.
> +     */
> +    if ( !last_awake_control_domain )
> +        return 0;
> +
> +    /*
> +     * This is the last awake control domain. It must not be parked unless the
> +     * request can proceed as a host-suspend request; otherwise Xen would lose
> +     * the last domain that can coordinate the system suspend.
> +     */
> +    if ( awake_non_control_domain )
> +    {
> +        printk(XENLOG_DEBUG
> +               "SYSTEM_SUSPEND denied for %pd: non-control domains awake\n",
> +               d);
> +        return PSCI_DENIED;
> +    }
> +
> +    /*
> +     * Host-wide gates are relevant only for the last-control-domain case. They
> +     * must not block parking of a non-last control domain, but they must deny
> +     * the last control domain when host suspend is not currently available.
> +     */
> +    if ( !host_system_suspend_allowed() )
> +        return PSCI_DENIED;
> +
> +    return 0;
> +}
> +
> static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
> {
>     int32_t rc;
> @@ -232,10 +316,6 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>     if ( is_64bit_domain(d) && is_thumb )
>         return PSCI_INVALID_ADDRESS;
> 
> -    /* SYSTEM_SUSPEND is not supported for the hardware domain yet */
> -    if ( is_hardware_domain(d) )
> -        return PSCI_NOT_SUPPORTED;
> -
>     /* Ensure that all CPUs other than the calling one are offline */
>     domain_lock(d);
>     for_each_vcpu ( d, v )
> @@ -252,16 +332,29 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>     if ( rc )
>         return PSCI_DENIED;
> 
> -    rc = domain_shutdown(d, SHUTDOWN_suspend);
> +    spin_lock(&vpsci_system_suspend_lock);
> +
> +    rc = domain_psci_system_suspend_policy(d);
> +    if ( !rc )
> +    {
> +        rc = domain_shutdown(d, SHUTDOWN_suspend);
> +        if ( rc )
> +            rc = PSCI_DENIED;
> +        else
> +        {
> +            rctx->ctxt = ctxt;
> +            rctx->wake_cpu = current;
> +        }
> +    }
> +
> +    spin_unlock(&vpsci_system_suspend_lock);
> +
>     if ( rc )
>     {
>         free_vcpu_guest_context(ctxt);
> -        return PSCI_DENIED;
> +        return rc;
>     }
> 
> -    rctx->ctxt = ctxt;
> -    rctx->wake_cpu = current;
> -
>     gprintk(XENLOG_DEBUG,
>             "SYSTEM_SUSPEND requested, epoint=%#"PRIregister", cid=%#"PRIregister"\n",
>             epoint, cid);
> @@ -287,10 +380,9 @@ static int32_t do_psci_1_0_features(uint32_t psci_func_id)
>     case PSCI_0_2_FN32_SYSTEM_RESET:
>     case PSCI_1_0_FN32_PSCI_FEATURES:
>     case ARM_SMCCC_VERSION_FID:
> -        return 0;
>     case PSCI_1_0_FN32_SYSTEM_SUSPEND:
>     case PSCI_1_0_FN64_SYSTEM_SUSPEND:
> -        return is_hardware_domain(current->domain) ? PSCI_NOT_SUPPORTED : 0;
> +        return 0;
>     default:
>         return PSCI_NOT_SUPPORTED;
>     }
> diff --git a/xen/common/Kconfig b/xen/common/Kconfig
> index da80fdba84..52bd98f7ad 100644
> --- a/xen/common/Kconfig
> +++ b/xen/common/Kconfig
> @@ -140,6 +140,9 @@ config HAS_EX_TABLE
> config HAS_FAST_MULTIPLY
> bool
> 
> +config HAS_HWDOM_SYSTEM_SUSPEND
> + bool
> +
> config HAS_IOPORTS
> bool
> 
> diff --git a/xen/common/domain.c b/xen/common/domain.c
> index e16f1ac383..10c358c7aa 100644
> --- a/xen/common/domain.c
> +++ b/xen/common/domain.c
> @@ -1377,6 +1377,11 @@ void __domain_crash(struct domain *d)
>     domain_shutdown(d, SHUTDOWN_crash);
> }
> 
> +static inline bool want_hwdom_shutdown(uint8_t reason)
> +{
> +    return !IS_ENABLED(CONFIG_HAS_HWDOM_SYSTEM_SUSPEND) ||
> +           reason != SHUTDOWN_suspend;
> +}
> 
> int domain_shutdown(struct domain *d, u8 reason)
> {
> @@ -1393,7 +1398,7 @@ int domain_shutdown(struct domain *d, u8 reason)
>         d->shutdown_code = reason;
>     reason = d->shutdown_code;
> 
> -    if ( is_hardware_domain(d) )
> +    if ( is_hardware_domain(d) && want_hwdom_shutdown(reason) )
>         hwdom_shutdown(reason);
> 
>     if ( domain_shutting_down(d) )
> diff --git a/xen/drivers/char/serial.c b/xen/drivers/char/serial.c
> index cf0abf1893..1cdf4968ac 100644
> --- a/xen/drivers/char/serial.c
> +++ b/xen/drivers/char/serial.c
> @@ -490,6 +490,8 @@ const struct vuart_info *serial_vuart_info(int idx)
> 
> #ifdef CONFIG_SYSTEM_SUSPEND
> 
> +static bool __ro_after_init serial_suspend_available = true;
> +
> void serial_suspend(void)
> {
>     int i;
> @@ -506,6 +508,11 @@ void serial_resume(void)
>             com[i].driver->resume(&com[i]);
> }
> 
> +bool serial_suspend_supported(void)
> +{
> +    return serial_suspend_available;
> +}
> +
> #endif /* CONFIG_SYSTEM_SUSPEND */
> 
> void __init serial_register_uart(int idx, struct uart_driver *driver,
> @@ -514,6 +521,11 @@ void __init serial_register_uart(int idx, struct uart_driver *driver,
>     /* Store UART-specific info. */
>     com[idx].driver = driver;
>     com[idx].uart   = uart;
> +
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    if ( !driver->suspend || !driver->resume )
> +        serial_suspend_available = false;
> +#endif
> }
> 
> void __init serial_async_transmit(struct serial_port *port)
> diff --git a/xen/drivers/passthrough/arm/iommu.c b/xen/drivers/passthrough/arm/iommu.c
> index 100545e23f..461e01703e 100644
> --- a/xen/drivers/passthrough/arm/iommu.c
> +++ b/xen/drivers/passthrough/arm/iommu.c
> @@ -19,6 +19,7 @@
> #include <xen/device_tree.h>
> #include <xen/iommu.h>
> #include <xen/lib.h>
> +#include <xen/suspend.h>
> 
> #include <asm/device.h>
> 
> @@ -46,6 +47,11 @@ void __init iommu_set_ops(const struct iommu_ops *ops)
>     }
> 
>     iommu_ops = ops;
> +
> +#ifdef CONFIG_SYSTEM_SUSPEND
> +    if ( !ops->suspend || !ops->resume )
> +        host_system_suspend_disable("IOMMU driver lacks suspend support");
> +#endif
> }
> 
> int __init iommu_hardware_setup(void)
> diff --git a/xen/drivers/passthrough/arm/smmu-v3.c b/xen/drivers/passthrough/arm/smmu-v3.c
> index 7f1d00fb81..16947a12f2 100644
> --- a/xen/drivers/passthrough/arm/smmu-v3.c
> +++ b/xen/drivers/passthrough/arm/smmu-v3.c
> @@ -91,6 +91,7 @@
> #include <asm/io.h>
> #include <asm/iommu_fwspec.h>
> #include <asm/platform.h>
> +#include <asm/suspend.h>
> 
> #include "smmu-v3.h"
> 
> @@ -1866,6 +1867,7 @@ static void arm_smmu_write_msi_msg(struct msi_desc *desc, struct msi_msg *msg)
> 
> static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
> {
> + static bool __ro_after_init host_suspend_blocked_by_msi;
> struct msi_desc *desc;
> int ret, nvec = ARM_SMMU_MAX_MSIS;
> struct device *dev = smmu->dev;
> @@ -1910,6 +1912,13 @@ static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
> }
> }
> 
> + if ( !host_suspend_blocked_by_msi )
> + {
> + host_suspend_blocked_by_msi = true;
> + host_system_suspend_disable(
> + "SMMUv3 MSI IRQ path is unsupported for host suspend");
> + }
> +
> /* Add callback to free MSIs on teardown */
> devm_add_action(dev, arm_smmu_free_msis, dev);
> }
> diff --git a/xen/include/xen/serial.h b/xen/include/xen/serial.h
> index 8e18445552..418b00ead0 100644
> --- a/xen/include/xen/serial.h
> +++ b/xen/include/xen/serial.h
> @@ -137,6 +137,7 @@ const struct vuart_info* serial_vuart_info(int idx);
> /* Serial suspend/resume. */
> void serial_suspend(void);
> void serial_resume(void);
> +bool serial_suspend_supported(void);
> #endif
> 
> /*
> diff --git a/xen/include/xen/suspend.h b/xen/include/xen/suspend.h
> index 6f94fd53b0..a941331035 100644
> --- a/xen/include/xen/suspend.h
> +++ b/xen/include/xen/suspend.h
> @@ -6,6 +6,8 @@
> #if __has_include(<asm/suspend.h>)
> #include <asm/suspend.h>
> #else
> +struct domain;
> +
> static inline void arch_domain_resume(struct domain *d) {}
> #endif
> 
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 13/13] xen/arm: Add host system suspend backend
  2026-08-27 14:32 ` [PATCH v12 13/13] xen/arm: Add host system suspend backend Mykola Kvach
  2026-08-27 21:59   ` Volodymyr Babchuk
@ 2026-09-28 16:19   ` Bertrand Marquis
  2026-09-30 22:10     ` Mykola Kvach
  1 sibling, 1 reply; 37+ messages in thread
From: Bertrand Marquis @ 2026-09-28 16:19 UTC (permalink / raw)
  To: Mykola Kvach
  Cc: xen-devel@lists.xenproject.org, Stefano Stabellini, Julien Grall,
	Michal Orzel, Volodymyr Babchuk

Hi Mykola,

> On 27 Aug 2026, at 16:32, Mykola Kvach <mykola_kvach@epam.com> wrote:
> 
> From: Mirela Simonovic <mirela.simonovic@aggios.com>
> 
> Add the Xen-wide suspend/resume backend used after a control-domain
> vPSCI SYSTEM_SUSPEND request has been accepted. The vPSCI policy,
> runtime driver blockers and control-domain sequencing checks are handled
> by the preceding commit; this change adds the code that actually drives
> the host suspend attempt.
> 
> The backend runs from a tasklet scheduled on pCPU0, because non-boot CPUs
> are disabled during suspend. It freezes domains, disables the scheduler
> and then disables non-boot CPUs.
> 
> Host-side suspend participants are handled in phases. IOMMU and console
> state are suspended first. Local IRQs are then disabled before suspending
> timer and GIC state. On resume or failure, the completed suspend phases
> are unwound in reverse: GIC and timer state are restored while IRQs are
> still disabled, local IRQs are restored, and then console and IOMMU state
> are restored.
> 
> On boot, init_ttbr is normally initialized during secondary CPU hotplug.
> On uniprocessor systems this can leave init_ttbr uninitialized, so set it
> from the boot CPU before entering suspend.
> 
> Note: the code is behind CONFIG_HAS_SYSTEM_SUSPEND, which is currently
> only selected when UNSUPPORTED is set and MPU is not set.
> 
> Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> ---
> Changes in V10:
> - Re-apply boot CPU local errata/workaround handling after SYSTEM_SUSPEND,
>  before resuming the rest of the host suspend path.
> - Move set_init_ttbr() declaration to asm/mmu/mm.h, since it is
>  MMU-specific.
> 
> Changes in V9:
> - Split vPSCI availability policy, runtime host-suspend blockers and the
>  domain-readiness precheck into the preceding commit.
> - Trigger the host suspend backend from the control-domain SYSTEM_SUSPEND
>  path.
> - Reorder the host suspend/resume phases so the timer is suspended with
>  local IRQs disabled and local IRQs are restored after the GIC and timer
>  resume paths, before the console and IOMMU resume paths.
> - Move HAS_HWDOM_SYSTEM_SUSPEND and related logic to policy patch.
> 
> Changes in V8:
> - Add a pre-suspend check in system_suspend() after scheduler_disable() to
>  require all domains to be in the shut down state with SHUTDOWN_suspend
>  before proceeding with the global suspend flow.
> - Drop the common-level depends on !ARM_64 || !SYSTEM_SUSPEND from
>  CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND and model the ARM64 suspend case
>  with an arch-selected capability instead.
> - Rename CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND to
>  CONFIG_HAS_HWDOM_SYSTEM_SUSPEND.
> - Rename need_hwdom_shutdown() to want_hwdom_shutdown().
> 
> Changes in V7:
> - Control domain is responsible for host suspend.
> - Add an empty inline host_system_suspend() function when SYSTEM_SUSPEND
>  config is disabled.
> - Use IS_ENABLED() for config checking instead of #ifdef.
> - Replace #ifdef checks in domain_shutdown() with IS_ENABLED() to simplify
>  control flow.
> - Factor hardware domain shutdown condition into a helper
>  (need_hwdom_shutdown()) to avoid preprocessor directives inside the
>  function.
> - Squash with iommu suspend/resume commit.
> ---
> xen/arch/arm/Kconfig                 |   1 +
> xen/arch/arm/cpuerrata.c             |   7 +-
> xen/arch/arm/include/asm/cpuerrata.h |   1 +
> xen/arch/arm/include/asm/mmu/mm.h    |   2 +
> xen/arch/arm/include/asm/suspend.h   |   2 +
> xen/arch/arm/mmu/smpboot.c           |   2 +-
> xen/arch/arm/suspend.c               | 156 +++++++++++++++++++++++++++
> xen/arch/arm/vpsci.c                 |  10 +-
> 8 files changed, 177 insertions(+), 4 deletions(-)
> 
> diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
> index 9027aa17eb..da1585ec50 100644
> --- a/xen/arch/arm/Kconfig
> +++ b/xen/arch/arm/Kconfig
> @@ -9,6 +9,7 @@ config ARM_64
> select 64BIT
> select HAS_DOMAIN_TYPE
> select HAS_FAST_MULTIPLY
> + select HAS_SYSTEM_SUSPEND if !MPU && UNSUPPORTED
> select HAS_VPCI_GUEST_SUPPORT if PCI_PASSTHROUGH
> 
> config ARM
> diff --git a/xen/arch/arm/cpuerrata.c b/xen/arch/arm/cpuerrata.c
> index 3a32183618..e6499aaab3 100644
> --- a/xen/arch/arm/cpuerrata.c
> +++ b/xen/arch/arm/cpuerrata.c
> @@ -782,6 +782,11 @@ void check_local_cpu_errata(void)
>     update_cpu_capabilities(arm_errata, "enabled workaround for");
> }
> 
> +int enable_local_cpu_errata_workarounds(void)
> +{
> +    return enable_nonboot_cpu_caps(arm_errata);
> +}
> +
> void __init enable_errata_workarounds(void)
> {
>     enable_cpu_capabilities(arm_errata);
> @@ -818,7 +823,7 @@ static int cpu_errata_callback(struct notifier_block *nfb,
>          * fixed to expect an error at CPU_STARTING phase.
>          */
>         ASSERT(system_state != SYS_STATE_boot);
> -        rc = enable_nonboot_cpu_caps(arm_errata);
> +        rc = enable_local_cpu_errata_workarounds();
>         break;
>     default:
>         break;
> diff --git a/xen/arch/arm/include/asm/cpuerrata.h b/xen/arch/arm/include/asm/cpuerrata.h
> index 1799a16d7e..b93521326f 100644
> --- a/xen/arch/arm/include/asm/cpuerrata.h
> +++ b/xen/arch/arm/include/asm/cpuerrata.h
> @@ -5,6 +5,7 @@
> #include <asm/alternative.h>
> 
> void check_local_cpu_errata(void);
> +int enable_local_cpu_errata_workarounds(void);
> void enable_errata_workarounds(void);
> 
> #define CHECK_WORKAROUND_HELPER(erratum, feature, arch)         \
> diff --git a/xen/arch/arm/include/asm/mmu/mm.h b/xen/arch/arm/include/asm/mmu/mm.h
> index 7f4d59137d..ee73a77777 100644
> --- a/xen/arch/arm/include/asm/mmu/mm.h
> +++ b/xen/arch/arm/include/asm/mmu/mm.h
> @@ -110,6 +110,8 @@ void dump_pt_walk(paddr_t ttbr, paddr_t addr,
> extern void switch_ttbr(uint64_t ttbr);
> extern void relocate_and_switch_ttbr(uint64_t ttbr);
> 
> +void set_init_ttbr(lpae_t *root);
> +
> #endif /* __ARM_MMU_MM_H__ */
> 
> /*
> diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> index 50dc6e9fdf..889a6509d9 100644
> --- a/xen/arch/arm/include/asm/suspend.h
> +++ b/xen/arch/arm/include/asm/suspend.h
> @@ -41,11 +41,13 @@ int prepare_resume_ctx(void);
> void hyp_resume(void);
> bool host_system_suspend_allowed(void);
> void host_system_suspend_disable(const char *reason);
> +void host_system_suspend(struct domain *d);
> 
> #else /* !CONFIG_SYSTEM_SUSPEND */
> 
> static inline bool host_system_suspend_allowed(void) { return false; }
> static inline void host_system_suspend_disable(const char *reason) {}
> +static inline void host_system_suspend(struct domain *d) {}
> 
> #endif
> 
> diff --git a/xen/arch/arm/mmu/smpboot.c b/xen/arch/arm/mmu/smpboot.c
> index 37e91d72b7..ff508ecf40 100644
> --- a/xen/arch/arm/mmu/smpboot.c
> +++ b/xen/arch/arm/mmu/smpboot.c
> @@ -72,7 +72,7 @@ static void clear_boot_pagetables(void)
>     clear_table(boot_third);
> }
> 
> -static void set_init_ttbr(lpae_t *root)
> +void set_init_ttbr(lpae_t *root)
> {
>     /*
>      * init_ttbr is part of the identity mapping which is read-only. So
> diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> index c7c26bcf03..3fe2ffa4fb 100644
> --- a/xen/arch/arm/suspend.c
> +++ b/xen/arch/arm/suspend.c
> @@ -1,10 +1,18 @@
> /* SPDX-License-Identifier: GPL-2.0-only */
> 
> +#include <asm/cpuerrata.h>
> +#include <asm/cpufeature.h>
> +#include <asm/gic.h>
> #include <asm/psci.h>
> #include <asm/suspend.h>
> 
> +#include <xen/console.h>
> +#include <xen/cpu.h>
> +#include <xen/iommu.h>
> #include <xen/lib.h>
> +#include <xen/sched.h>
> #include <xen/serial.h>
> +#include <xen/tasklet.h>
> 
> struct resume_cpu_context resume_cpu_context;
> 
> @@ -44,6 +52,154 @@ void host_system_suspend_disable(const char *reason)
>            reason ? reason : "unsupported suspend/resume path");
> }
> 
> +/* Xen suspend. data identifies the domain that initiated suspend. */
> +static void system_suspend(void *data)
> +{
> +    int status;
> +    unsigned long flags;
> +    struct domain *d = (struct domain *)data;
> +
> +    BUG_ON(system_state != SYS_STATE_active);
> +
> +    system_state = SYS_STATE_suspend;
> +
> +    printk("Xen suspending...\n");
> +
> +    freeze_domains();
> +    scheduler_disable();
> +
> +    /*
> +     * Non-boot CPUs have to be disabled on suspend and enabled on resume
> +     * (hotplug-based mechanism). Disabling non-boot CPUs will lead to PSCI
> +     * CPU_OFF to be called by each non-boot CPU. Depending on the underlying
> +     * platform capabilities, this may lead to the physical powering down of
> +     * CPUs.
> +     */
> +    status = disable_nonboot_cpus();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_nonboot_cpus;
> +    }
> +
> +    console_start_sync();
> +    status = iommu_suspend();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_end_sync;
> +    }
> +
> +    status = console_suspend();
> +    if ( status )
> +    {
> +        dprintk(XENLOG_ERR, "Failed to suspend the console, err=%d\n", status);
> +        system_state = SYS_STATE_resume;
> +        goto resume_iommu;
> +    }
> +
> +    local_irq_save(flags);
> +
> +    time_suspend();
> +
> +    status = gic_suspend();
> +    if ( status )
> +    {
> +        system_state = SYS_STATE_resume;
> +        goto resume_time;
> +    }
> +
> +    set_init_ttbr(xen_pgtable);
> +
> +    /*
> +     * Enable identity mapping before entering suspend to simplify
> +     * the resume path
> +     */
> +    update_boot_mapping(true);
> +
> +    if ( prepare_resume_ctx() )
> +    {
> +        status = call_psci_system_suspend();
> +        /*
> +         * If suspend is finalized properly by above system suspend PSCI call,
> +         * the code below in this 'if' branch will never execute. Execution
> +         * will continue from hyp_resume which is the hypervisor's resume point.
> +         * In hyp_resume CPU context will be restored and since link-register is
> +         * restored as well, it will appear to return from prepare_resume_ctx.
> +         * The difference in returning from prepare_resume_ctx on system suspend
> +         * versus resume is in function's return value: on suspend, the return
> +         * value is a non-zero value, on resume it is zero. That is why the
> +         * control flow will not re-enter this 'if' branch on resume.
> +         */
> +        if ( status )
> +            dprintk(XENLOG_WARNING, "PSCI system suspend failed, err=%d\n",
> +                    status);
> +
> +        system_state = SYS_STATE_resume;
> +    }
> +    else
> +    {
> +        system_state = SYS_STATE_resume;
> +
> +        /*
> +         * CPU0 resumes directly from hyp_resume(), bypassing the CPU hotplug
> +         * path that re-checks and re-enables errata workarounds for secondary
> +         * CPUs.
> +         */
> +        check_local_cpu_errata();
> +        check_local_cpu_features();
> +        BUG_ON(enable_local_cpu_errata_workarounds());
> +    }
> +
> +    update_boot_mapping(false);
> +
> +    gic_resume();
> +
> + resume_time:
> +    time_resume();
> +
> +    local_irq_restore(flags);
> +
> +    console_resume();
> +
> + resume_iommu:
> +    iommu_resume();
> +
> + resume_end_sync:
> +    console_end_sync();
> +
> + resume_nonboot_cpus:
> +    /*
> +     * The rcu_barrier() has to be added to ensure that the per cpu area is
> +     * freed before a non-boot CPU tries to initialize it (_free_percpu_area()
> +     * has to be called before the init_percpu_area()). This scenario occurs
> +     * when non-boot CPUs are hot-unplugged on suspend and hotplugged on resume.

This line would need wrapping as it is over 80 chars.

> +     */
> +    rcu_barrier();

Could you clarify which per-CPU area needs freeing here?
  
system_state is SYS_STATE_suspend for every CPU taken down by
disable_nonboot_cpus(), so cpu_percpu_callback() does not queue
_free_percpu_area() for any of them.

So I do not quite get what your comment case actually is.
Could you explain ?

Cheers
Bertrand

> +    enable_nonboot_cpus();
> +
> +    scheduler_enable();
> +    thaw_domains();
> +
> +    system_state = SYS_STATE_active;
> +
> +    printk("Resume (status %d)\n", status);
> +
> +    domain_resume(d);
> +}
> +
> +static DECLARE_TASKLET(system_suspend_tasklet, system_suspend, NULL);
> +
> +void host_system_suspend(struct domain *d)
> +{
> +    system_suspend_tasklet.data = (void *)d;
> +    /*
> +     * The suspend procedure has to be finalized by the pCPU#0 (non-boot pCPUs
> +     * will be disabled during the suspend).
> +     */
> +    tasklet_schedule_on_cpu(&system_suspend_tasklet, 0);
> +}
> +
> /*
>  * Local variables:
>  * mode: C
> diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
> index a41355d75d..5134e75c24 100644
> --- a/xen/arch/arm/vpsci.c
> +++ b/xen/arch/arm/vpsci.c
> @@ -237,7 +237,8 @@ static bool domain_in_suspend_state(struct domain *d)
>     return suspended;
> }
> 
> -static int32_t domain_psci_system_suspend_policy(struct domain *d)
> +static int32_t domain_psci_system_suspend_policy(struct domain *d,
> +                                                 bool *host_suspend)
> {
>     struct domain *other;
>     bool last_awake_control_domain = true;
> @@ -300,6 +301,7 @@ static int32_t domain_psci_system_suspend_policy(struct domain *d)
>     if ( !host_system_suspend_allowed() )
>         return PSCI_DENIED;
> 
> +    *host_suspend = true;
>     return 0;
> }
> 
> @@ -310,6 +312,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>     struct vcpu *v;
>     struct domain *d = current->domain;
>     bool is_thumb = epoint & 1;
> +    bool host_suspend = false;
>     struct resume_info *rctx = &d->arch.resume_ctx;
> 
>     /* THUMB set is not allowed with 64-bit domain */
> @@ -334,7 +337,7 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
> 
>     spin_lock(&vpsci_system_suspend_lock);
> 
> -    rc = domain_psci_system_suspend_policy(d);
> +    rc = domain_psci_system_suspend_policy(d, &host_suspend);
>     if ( !rc )
>     {
>         rc = domain_shutdown(d, SHUTDOWN_suspend);
> @@ -359,6 +362,9 @@ static int32_t do_psci_1_0_system_suspend(register_t epoint, register_t cid)
>             "SYSTEM_SUSPEND requested, epoint=%#"PRIregister", cid=%#"PRIregister"\n",
>             epoint, cid);
> 
> +    if ( host_suspend )
> +        host_system_suspend(d);
> +
>     return rc;
> }
> 
> -- 
> 2.43.0
> 



^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers
  2026-09-28 16:17   ` Bertrand Marquis
@ 2026-09-30 14:44     ` Mykola Kvach
  0 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-09-30 14:44 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Rahul Singh,
	Stefano Stabellini, Julien Grall, Michal Orzel, Volodymyr Babchuk,
	Pranjal Shrivastava, Luca Fancellu

Hi Bertrand,

Thank you for the review.

On Mon, Sep 28, 2026 at 7:18 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > Add system suspend/resume callbacks for the Arm SMMUv3 driver.
> >
> > During suspend, configure GBPA to abort incoming transactions, disable the
> > translation interface while keeping CMDQ enabled, issue CMD_SYNC to ensure
> > all previously issued commands have completed, then disable the SMMU IRQs
> > and SMMU.
> >
> > Resume uses arm_smmu_device_reset() to reprogram the SMMU and re-enable
> > translation and interrupt generation.
> >
> > The IRQ setup split follows the approach from Pranjal Shrivastava's Linux
> > arm-smmu-v3 runtime/system sleep series: IRQ handlers are requested once
> > during probe, while reset/resume only restores SMMU hardware state and
> > re-enables IRQ_CTRL.
> >
> > Only the pieces relevant to Xen's currently supported SMMUv3 path are
> > ported here. Xen documents SMMUv3 MSI and PCI ATS as unsupported and not
> > compiled/tested, so this patch does not restore SMMU MSI IRQ_CFGn registers
> > nor reinitialize ATS/PRI endpoints. If those paths become usable,
> > suspend/resume will need corresponding MSI restore and ATS/PRI
> > quiesce/reinit steps.
> >
> > Link: https://lore.kernel.org/r/20260414194702.1229094-1-praan@google.com/
> > Based-on-patch-by: Pranjal Shrivastava <praan@google.com>
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> > ---
> > Changes in V11:
> > - Keep arm_smmu_update_gbpa() and arm_smmu_device_reset() in init text when
> >  CONFIG_SYSTEM_SUSPEND is disabled.
> >
> > Changes in V10:
> > - Disable SMMU interrupt generation during suspend before disabling the
> >  SMMU interface, matching the resume/reset path which re-enables IRQ_CTRL.
> >
> > Changes in V9:
> > - Use CMD_SYNC in suspend instead of polling CMDQ_CONS, so the suspend
> >  path waits for command completion rather than only command consumption.
> > - Document that arm_smmu_setup_irqs() is probe-only and that future Xen
> >  SMMUv3 MSI support will need to restore SMMU IRQ_CFGn registers on
> >  resume.
> > - Restore the reference to Pranjal's Linux runtime/system sleep series and
> >  clarify that MSI/ATS/PRI resume handling is outside the supported Xen
> >  path.
> > - Prefix the subject with xen/arm for consistency with the rest of the
> >  Arm suspend/resume series.
> >
> > Changes in V8:
> > - Honor ARM_SMMU_FEAT_SEV when draining the CMDQ during suspend, matching
> >  the existing runtime CMD_SYNC path.
> > - Fold the suspend rollback reset path into a helper and rename the error
> >  reporting to describe suspend rollback rather than resume.
> > - Treat SMMU reset failure during resume as fatal instead of logging and
> >  continuing with a potentially unusable IOMMU.
> > - cosmetic changes
> > ---
> > xen/drivers/passthrough/arm/smmu-v3.c | 194 +++++++++++++++++++++-----
> > 1 file changed, 158 insertions(+), 36 deletions(-)
> >
> > diff --git a/xen/drivers/passthrough/arm/smmu-v3.c b/xen/drivers/passthrough/arm/smmu-v3.c
> > index bf153227db..7f1d00fb81 100644
> > --- a/xen/drivers/passthrough/arm/smmu-v3.c
> > +++ b/xen/drivers/passthrough/arm/smmu-v3.c
> > @@ -94,6 +94,12 @@
> >
> > #include "smmu-v3.h"
> >
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +#define __init_or_smmu_suspend
> > +#else
> > +#define __init_or_smmu_suspend __init
> > +#endif
> > +
> > #define ARM_SMMU_VTCR_SH_IS 3
> > #define ARM_SMMU_VTCR_RGN_WBWA 1
> > #define ARM_SMMU_VTCR_TG0_4K 0
> > @@ -1814,8 +1820,8 @@ static int arm_smmu_write_reg_sync(struct arm_smmu_device *smmu, u32 val,
> > }
> >
> > /* GBPA is "special" */
> > -static int __init arm_smmu_update_gbpa(struct arm_smmu_device *smmu,
> > -                                       u32 set, u32 clr)
> > +static int __init_or_smmu_suspend
> > +arm_smmu_update_gbpa(struct arm_smmu_device *smmu, u32 set, u32 clr)
> > {
> > int ret;
> > u32 reg, __iomem *gbpa = smmu->base + ARM_SMMU_GBPA;
> > @@ -1995,10 +2001,35 @@ err_free_evtq_irq:
> > return ret;
> > }
> >
> > +static int arm_smmu_enable_irqs(struct arm_smmu_device *smmu)
> > +{
> > + int ret;
> > + u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
> > +
> > + if ( smmu->features & ARM_SMMU_FEAT_PRI )
> > + irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
> > +
> > + /* Enable interrupt generation on the SMMU */
> > + ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
> > +      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
> > + if ( ret )
> > + {
> > + dev_warn(smmu->dev, "failed to enable irqs\n");
> > + return ret;
> > + }
> > +
> > + return 0;
> > +}
> > +
> > +/*
> > + * Probe-time only: request host IRQs and, when available, program the SMMU's
> > + * MSI doorbells. Resume does not restore the SMMU *_IRQ_CFGn MSI registers,
> > + * so any host suspend support must treat the active MSI IRQ path as
> > + * unsupported until that restore path exists.
> > + */
> > static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
> > {
> > int ret, irq;
> > - u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
> >
> > /* Disable IRQs first */
> > ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
> > @@ -2028,22 +2059,7 @@ static int __init arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
> > }
> > }
> >
> > - if (smmu->features & ARM_SMMU_FEAT_PRI)
> > - irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
> > -
> > - /* Enable interrupt generation on the SMMU */
> > - ret = arm_smmu_write_reg_sync(smmu, irqen_flags,
> > -      ARM_SMMU_IRQ_CTRL, ARM_SMMU_IRQ_CTRLACK);
> > - if (ret) {
> > - dev_warn(smmu->dev, "failed to enable irqs\n");
> > - goto err_free_irqs;
> > - }
> > -
> > return 0;
> > -
> > -err_free_irqs:
> > - arm_smmu_free_irqs(smmu);
> > - return ret;
> > }
> >
> > static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
> > @@ -2057,7 +2073,8 @@ static int arm_smmu_device_disable(struct arm_smmu_device *smmu)
> > return ret;
> > }
> >
> > -static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> > +static int __init_or_smmu_suspend
> > +arm_smmu_device_reset(struct arm_smmu_device *smmu)
> > {
> > int ret;
> > u32 reg, enables;
> > @@ -2163,17 +2180,9 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> > }
> > }
> >
> > - ret = arm_smmu_setup_irqs(smmu);
> > - if (ret) {
> > - dev_err(smmu->dev, "failed to setup irqs\n");
> > + ret = arm_smmu_enable_irqs(smmu);
> > + if ( ret )
> > return ret;
> > - }
> > -
> > - /* Initialize tasklets for threaded IRQs*/
> > - tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
> > - tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
> > - tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
> > - smmu);
> >
> > /* Enable the SMMU interface, or ensure bypass */
> > if (disable_bypass) {
> > @@ -2181,20 +2190,16 @@ static int __init arm_smmu_device_reset(struct arm_smmu_device *smmu)
> > } else {
> > ret = arm_smmu_update_gbpa(smmu, 0, GBPA_ABORT);
> > if (ret)
> > - goto err_free_irqs;
> > + return ret;
> > }
> > ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
> >      ARM_SMMU_CR0ACK);
> > if (ret) {
> > dev_err(smmu->dev, "failed to enable SMMU interface\n");
> > - goto err_free_irqs;
> > + return ret;
> > }
> >
> > return 0;
> > -
> > -err_free_irqs:
> > - arm_smmu_free_irqs(smmu);
> > - return ret;
> > }
> >
> > static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
> > @@ -2558,10 +2563,23 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
> > if (ret)
> > goto out_free;
> >
> > + ret = arm_smmu_setup_irqs(smmu);
> > + if ( ret )
> > + {
> > + dev_err(smmu->dev, "failed to setup irqs\n");
> > + goto out_free;
> > + }
> > +
> > + /* Initialize tasklets for threaded IRQs*/
> > + tasklet_init(&smmu->evtq_irq_tasklet, arm_smmu_evtq_tasklet, smmu);
> > + tasklet_init(&smmu->priq_irq_tasklet, arm_smmu_priq_tasklet, smmu);
> > + tasklet_init(&smmu->combined_irq_tasklet, arm_smmu_combined_irq_tasklet,
> > + smmu);
> > +
> > /* Reset the device */
> > ret = arm_smmu_device_reset(smmu);
> > if (ret)
> > - goto out_free;
> > + goto out_free_irqs;
> >
> > /*
> > * Keep a list of all probed devices. This will be used to query
> > @@ -2575,6 +2593,8 @@ static int __init arm_smmu_device_probe(struct platform_device *pdev)
> >
> > return 0;
> >
> > +out_free_irqs:
> > + arm_smmu_free_irqs(smmu);
> >
> > out_free:
> > arm_smmu_free_structures(smmu);
> > @@ -2855,6 +2875,104 @@ static void arm_smmu_iommu_xen_domain_teardown(struct domain *d)
> > xfree(xen_domain);
> > }
> >
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +
> > +static void arm_smmu_reset_for_suspend_rollback(struct arm_smmu_device *smmu)
> > +{
> > + int ret = arm_smmu_device_reset(smmu);
> > +
> > + if ( ret )
> > + dev_err(smmu->dev, "Failed to reset during suspend rollback: %d\n",
> > + ret);
>
> If reset fails here, we only print an error.
>
> Could the SMMU be left disabled with GBPA.ABORT cleared, allowing guest
> DMA to bypass translation when the domains resume?

In Xen, disable_bypass is always true, so reset does not clear
GBPA.ABORT. After a successful ABORT update, the bit stays set
during rollback. If reset then fails with the SMMU disabled,
transactions will still be aborted.

Still, logging a rollback error is not enough. After suspend
fails, the host can let domains run again, so rollback must
restore the SMMU.

In the next version, if the first GBPA update times out, I will
leave the current SMMU enabled and restore only the previously
suspended ones.

The GBPA update may still complete later. This is safe here:
GBPA.ABORT controls transactions only when CR0.SMMUEN is clear
(IHI0070G.b, section 6.3.14). We have not written CR0 at this
point, so translation and queues remain enabled even if the
GBPA update completes after the timeout.

If a later suspend step fails, I will try to restore both the
current SMMU and the previously suspended ones. If any rollback
reset fails, I will call panic(), as we already do when reset
fails during resume.

I will also check the GBPA and both CMD_SYNC return values in
reset, so these errors are passed back to the caller.

A suspend failure will still be recoverable if rollback succeeds.

>
>
> > +}
> > +
> > +static int arm_smmu_suspend(void)
> > +{
> > + struct arm_smmu_device *smmu;
> > + int ret = 0;
> > +
> > + list_for_each_entry(smmu, &arm_smmu_devices, devices)
> > + {
> > + /* Abort all transactions before disable to avoid spurious bypass */
> > + ret = arm_smmu_update_gbpa(smmu, GBPA_ABORT, 0);
> > + if ( ret )
> > + goto fail;
> > +
> > + ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
> > + ARM_SMMU_IRQ_CTRLACK);
> > + if ( ret )
> > + {
> > + dev_err(smmu->dev, "Timed-out while disabling SMMU irqs\n");
> > + goto fail;
> > + }
> > +
> > + /* Disable the SMMU via CR0.EN and all queues except CMDQ */
> > + ret = arm_smmu_write_reg_sync(smmu, CR0_CMDQEN, ARM_SMMU_CR0,
> > + ARM_SMMU_CR0ACK);
> > + if ( ret )
> > + {
> > + dev_err(smmu->dev, "Timed-out while disabling smmu\n");
> > + goto fail;
> > + }
> > +
> > + /*
> > + * At this point the translation interface is disabled and the
> > + * SMMU won't access translation/config structures, even
> > + * speculatively, as per the IHI0070 spec (section 6.3.9.6).
> > + * CMDQ is still enabled so that a CMD_SYNC can complete any
> > + * previously issued commands.
> > + */
> > +
> > + /* Ensure all previously issued commands have completed. */
> > + ret = arm_smmu_cmdq_issue_sync(smmu);
> > + if ( ret )
> > + {
> > + dev_err(smmu->dev, "Timed-out waiting for pending commands\n");
> > + goto fail;
> > + }
>
> Could we lose EVTQ events here because we do not check the queue after
> stopping it?

Yes, the old code could lose unread EVTQ entries. The hardware
producer index may be ahead of our cached value. Reset would
then restore the old index and hide those entries.

I will save the producer index after the queue is disabled and
CR0ACK confirms that the change is complete. The queue memory
and consumer index will be kept across suspend.

I will also cover rollback after a timeout while stopping the
queue. I will wait again for the stop acknowledgement, then save
the producer index before reset overwrites it. If this wait also
times out, I will call panic().

After reset, I will schedule the EVTQ tasklet if unread entries
remain. Re-enabling interrupts does not report old events again
(IHI0070G.b, section 6.3.16).

This preserves entries already recorded in the queue. Section
6.3.9.4 allows uncommitted events from terminated faulting
transactions to be discarded when the queue is disabled.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface)
  2026-09-28 16:18   ` Bertrand Marquis
@ 2026-09-30 17:32     ` Mykola Kvach
  0 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-09-30 17:32 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Luca Fancellu

On Mon, Sep 28, 2026 at 7:19 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:31, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > From: Mirela Simonovic <mirela.simonovic@aggios.com>
> >
> > Invoke PSCI SYSTEM_SUSPEND to finalize Xen's suspend sequence on ARM64
> > platforms. Pass the Xen resume entry point (hyp_resume) to EL3 together
> > with a zero context ID, matching Linux.
> >
> > This patch wires up only the host-side PSCI SYSTEM_SUSPEND invocation.
> > The resume trampoline and context restore are provided by earlier patches
> > in the series.
> >
> > Only enable this path when CONFIG_SYSTEM_SUSPEND is set and PSCI
> > advertises SYSTEM_SUSPEND via PSCI_FEATURES.
> >
> > Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> > Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> > Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > Reviewed-by: Luca Fancellu <luca.fancellu@arm.com>
> > ---
> > Changes in v9:
> > - cache SYSTEM_SUSPEND support using PSCI_FEATURES and gate the host call
> >  on the cached capability
> > - keep the cached SYSTEM_SUSPEND capability read-only after init
> > - log whether firmware reports SYSTEM_SUSPEND support
> > - pass an explicit zero context ID in the SYSTEM_SUSPEND call
> > - drop the stale note claiming hyp_resume is still a stub
> > ---
> > xen/arch/arm/include/asm/psci.h |  1 +
> > xen/arch/arm/psci.c             | 31 ++++++++++++++++++++++++++++++-
> > 2 files changed, 31 insertions(+), 1 deletion(-)
> >
> > diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
> > index 48a93e6b79..bb3c73496e 100644
> > --- a/xen/arch/arm/include/asm/psci.h
> > +++ b/xen/arch/arm/include/asm/psci.h
> > @@ -23,6 +23,7 @@ int call_psci_cpu_on(int cpu);
> > void call_psci_cpu_off(void);
> > void call_psci_system_off(void);
> > void call_psci_system_reset(void);
> > +int call_psci_system_suspend(void);
> >
> > /* Range of allocated PSCI function numbers */
> > #define PSCI_FNUM_MIN_VALUE                 _AC(0,U)
> > diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
> > index b6860a7760..e05dae1133 100644
> > --- a/xen/arch/arm/psci.c
> > +++ b/xen/arch/arm/psci.c
> > @@ -17,23 +17,27 @@
> > #include <asm/cpufeature.h>
> > #include <asm/psci.h>
> > #include <asm/acpi.h>
> > +#include <asm/suspend.h>
> >
> > /*
> >  * While a 64-bit OS can make calls with SMC32 calling conventions, for
> >  * some calls it is necessary to use SMC64 to pass or return 64-bit values.
> > - * For such calls PSCI_0_2_FN_NATIVE(x) will choose the appropriate
> > + * For such calls PSCI_*_FN_NATIVE(x) will choose the appropriate
> >  * (native-width) function ID.
> >  */
> > #ifdef CONFIG_ARM_64
> > #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN64_##name
> > +#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN64_##name
> > #else
> > #define PSCI_0_2_FN_NATIVE(name)    PSCI_0_2_FN32_##name
> > +#define PSCI_1_0_FN_NATIVE(name)    PSCI_1_0_FN32_##name
> > #endif
> >
> > uint32_t psci_ver;
> > uint32_t smccc_ver;
> >
> > static uint32_t psci_cpu_on_nr;
> > +static bool __ro_after_init has_psci_system_suspend;
> >
> > #define PSCI_RET(res)   ((int32_t)(res).a0)
> >
> > @@ -60,6 +64,25 @@ void call_psci_cpu_off(void)
> >     }
> > }
> >
> > +int call_psci_system_suspend(void)
> > +{
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +    struct arm_smccc_res res;
> > +
> > +    if ( !has_psci_system_suspend )
> > +        return PSCI_NOT_SUPPORTED;
> > +
> > +    /* Context ID is unused for the Xen resume path. */
> > +    arm_smccc_smc(PSCI_1_0_FN_NATIVE(SYSTEM_SUSPEND), __pa(hyp_resume), 0,
> > +                  &res);
>
> Since this series was sent, arm_smccc_smc() has been changed to return
> the result directly instead of writing it through a result pointer.
>
> You need to update this call.

Ack.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy
  2026-09-28 16:18   ` Bertrand Marquis
@ 2026-09-30 20:42     ` Mykola Kvach
  0 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-09-30 20:42 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk, Andrew Cooper,
	Anthony PERARD, Jan Beulich, Roger Pau Monné, Rahul Singh,
	Oleksandr Tyshchenko

Hi Bertrand,

Thank you for the review.

On Mon, Sep 28, 2026 at 7:20 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:32, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > Introduce CONFIG_HAS_HWDOM_SYSTEM_SUSPEND as an architecture-selected
> > capability for platforms where the hardware domain can be parked with
> > SHUTDOWN_suspend without calling hwdom_shutdown().
> >
> > Expose PSCI SYSTEM_SUSPEND as a vPSCI operation for all domains. For
> > non-control domains, including the hardware domain when it is not acting
> > as a control domain, the call is handled as a guest/domain suspend request
> > and parks the domain in SHUTDOWN_suspend.
> >
> > Control domains need additional sequencing because their SYSTEM_SUSPEND
> > request is used to coordinate host-wide suspend. A non-last awake control
> > domain may be parked in SHUTDOWN_suspend without requiring the host
> > suspend path to be available. The last awake control domain is treated as
> > the point where the request becomes a host-suspend request, and it may
> > only proceed when all non-control domains are already in SHUTDOWN_suspend
> > and the host suspend path is available.
> >
> > Keep the control-domain sequencing and domain-readiness checks out of
> > PSCI_FEATURES. They are per-attempt runtime conditions rather than stable
> > PSCI function availability. Advertise SYSTEM_SUSPEND as implemented by
> > vPSCI and report attempt-time policy failures as PSCI_DENIED.
> >
> > Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
> > so that SHUTDOWN_suspend from the hardware domain can be treated as a
> > domain suspend state rather than as a hardware-domain initiated host
> > shutdown. This does not by itself imply that host-wide suspend is
> > available.
> >
> > Add host_system_suspend_allowed() to combine the host PSCI SYSTEM_SUSPEND
> > capability with runtime blockers reported by Xen-owned subsystems. Add
> > runtime blockers for registered serial, IOMMU, GIC and SMMUv3 MSI IRQ
> > paths lacking suspend/resume support. These blockers are runtime based,
> > so they only apply to drivers or paths that Xen actually uses on the
> > platform. For SMMUv3, the blocker applies only when Xen actually uses the
> > MSI IRQ path, since resume does not restore the SMMU *_IRQ_CFGn MSI
> > registers yet.
> >
> > Add a struct domain forward declaration to xen/suspend.h so the generic
> > header can expose arch_domain_resume() without requiring a full domain.h
> > include.
> >
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > Reviewed-by: Oleksandr Tyshchenko <Oleksandr_Tyshchenko@epam.com>
> > ---
> > Changes in V12:
> > - handle missing is_shut_down, change checking to call of
> >  domain_shutdown_completed
> >
> > Changes in V11:
> > - Mark host_system_suspend_runtime_allowed as __ro_after_init.
> > - Avoid printing the SMMUv3 MSI IRQ host suspend blocker more than once
> >  when multiple SMMUv3 instances use MSIs.
> > - Wrap the Arm IOMMU host suspend blocker in CONFIG_SYSTEM_SUSPEND to make
> >  its policy-only use explicit.
> >
> > Changes in V10:
> > - Return PSCI_DENIED rather than PSCI_NOT_SUPPORTED when the last awake
> >  control domain cannot proceed to host suspend, keeping PSCI_FEATURES
> >  stable once SYSTEM_SUSPEND is advertised.
> > - Shorten SYSTEM_SUSPEND blocker messages and use %pd when logging the
> >  control domain.
> > - Mark serial_suspend_available as __ro_after_init.
> > - Mention the struct domain forward declaration added to xen/suspend.h.
> >
> > Changes in V9:
> > - Select HAS_HWDOM_SYSTEM_SUSPEND independently from CONFIG_SYSTEM_SUSPEND
> >  so that hardware-domain SHUTDOWN_suspend support is not tied to
> >  host-wide system suspend availability.
> > - Add runtime host suspend blockers for Xen-owned subsystems lacking
> >  suspend/resume support.
> > - Keep vPSCI SYSTEM_SUSPEND advertised through PSCI_FEATURES and enforce
> >  control-domain sequencing in the call handler.
> > ---
> > xen/arch/arm/Kconfig                  |   1 +
> > xen/arch/arm/gic.c                    |   6 ++
> > xen/arch/arm/include/asm/psci.h       |   3 +
> > xen/arch/arm/include/asm/suspend.h    |  10 ++-
> > xen/arch/arm/psci.c                   |   7 ++
> > xen/arch/arm/suspend.c                |  40 +++++++++
> > xen/arch/arm/vpsci.c                  | 114 +++++++++++++++++++++++---
> > xen/common/Kconfig                    |   3 +
> > xen/common/domain.c                   |   7 +-
> > xen/drivers/char/serial.c             |  12 +++
> > xen/drivers/passthrough/arm/iommu.c   |   6 ++
> > xen/drivers/passthrough/arm/smmu-v3.c |   9 ++
> > xen/include/xen/serial.h              |   1 +
> > xen/include/xen/suspend.h             |   2 +
> > 14 files changed, 208 insertions(+), 13 deletions(-)
> >
> > diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
> > index 843a43897e..9027aa17eb 100644
> > --- a/xen/arch/arm/Kconfig
> > +++ b/xen/arch/arm/Kconfig
> > @@ -19,6 +19,7 @@ config ARM
> > select HAS_ALTERNATIVE if HAS_VMAP
> > select HAS_DEVICE_TREE_DISCOVERY
> > select HAS_DOM0LESS
> > + select HAS_HWDOM_SYSTEM_SUSPEND if !MPU
> > select HAS_GRANT_CACHE_FLUSH if GRANT_TABLE
> > select HAS_STACK_PROTECTOR
> > select HAS_STATIC_MEMORY
> > diff --git a/xen/arch/arm/gic.c b/xen/arch/arm/gic.c
> > index ffc11f36a1..0695474432 100644
> > --- a/xen/arch/arm/gic.c
> > +++ b/xen/arch/arm/gic.c
> > @@ -26,6 +26,7 @@
> > #include <asm/device.h>
> > #include <asm/io.h>
> > #include <asm/gic.h>
> > +#include <asm/suspend.h>
> > #include <asm/vgic.h>
> > #include <asm/acpi.h>
> >
> > @@ -44,6 +45,11 @@ static void __init __maybe_unused build_assertions(void)
> > void register_gic_ops(const struct gic_hw_operations *ops)
> > {
> >     gic_hw_ops = ops;
> > +
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +    if ( !ops->suspend || !ops->resume )
> > +        host_system_suspend_disable("GIC driver lacks suspend support");
> > +#endif
> > }
> >
> > static void clear_cpu_lr_mask(void)
> > diff --git a/xen/arch/arm/include/asm/psci.h b/xen/arch/arm/include/asm/psci.h
> > index bb3c73496e..142fa1bfe5 100644
> > --- a/xen/arch/arm/include/asm/psci.h
> > +++ b/xen/arch/arm/include/asm/psci.h
> > @@ -24,6 +24,9 @@ void call_psci_cpu_off(void);
> > void call_psci_system_off(void);
> > void call_psci_system_reset(void);
> > int call_psci_system_suspend(void);
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +bool psci_system_suspend_allowed(void);
> > +#endif
> >
> > /* Range of allocated PSCI function numbers */
> > #define PSCI_FNUM_MIN_VALUE                 _AC(0,U)
> > diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> > index c848fc6340..50dc6e9fdf 100644
> > --- a/xen/arch/arm/include/asm/suspend.h
> > +++ b/xen/arch/arm/include/asm/suspend.h
> > @@ -39,7 +39,15 @@ extern struct resume_cpu_context resume_cpu_context;
> >
> > int prepare_resume_ctx(void);
> > void hyp_resume(void);
> > -#endif /* CONFIG_SYSTEM_SUSPEND */
> > +bool host_system_suspend_allowed(void);
> > +void host_system_suspend_disable(const char *reason);
> > +
> > +#else /* !CONFIG_SYSTEM_SUSPEND */
> > +
> > +static inline bool host_system_suspend_allowed(void) { return false; }
> > +static inline void host_system_suspend_disable(const char *reason) {}
> > +
> > +#endif
> >
> > #endif /* ARM_SUSPEND_H */
> >
> > diff --git a/xen/arch/arm/psci.c b/xen/arch/arm/psci.c
> > index e05dae1133..e9d78668fd 100644
> > --- a/xen/arch/arm/psci.c
> > +++ b/xen/arch/arm/psci.c
> > @@ -41,6 +41,13 @@ static bool __ro_after_init has_psci_system_suspend;
> >
> > #define PSCI_RET(res)   ((int32_t)(res).a0)
> >
> > +#ifdef CONFIG_SYSTEM_SUSPEND
> > +bool psci_system_suspend_allowed(void)
> > +{
> > +    return has_psci_system_suspend;
> > +}
> > +#endif
> > +
> > int call_psci_cpu_on(int cpu)
> > {
> >     struct arm_smccc_res res;
> > diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> > index 6ea4a0f9cc..c7c26bcf03 100644
> > --- a/xen/arch/arm/suspend.c
> > +++ b/xen/arch/arm/suspend.c
> > @@ -1,9 +1,49 @@
> > /* SPDX-License-Identifier: GPL-2.0-only */
> >
> > +#include <asm/psci.h>
> > #include <asm/suspend.h>
> >
> > +#include <xen/lib.h>
> > +#include <xen/serial.h>
> > +
> > struct resume_cpu_context resume_cpu_context;
> >
> > +/*
> > + * Non-PSCI infrastructure can make host suspend impossible even when the PSCI
> > + * SYSTEM_SUSPEND conduit is present, e.g. when a Xen-owned driver has no valid
> > + * suspend/resume path.
> > + *
> > + * This gate is checked only when the last awake control domain attempts to
> > + * turn a guest SYSTEM_SUSPEND request into a host-suspend request.
> > + */
> > +static bool __ro_after_init host_system_suspend_runtime_allowed = true;
> > +
> > +static bool host_serial_suspend_allowed(void)
> > +{
> > +    if ( serial_suspend_supported() )
> > +        return true;
> > +
> > +    printk_once(XENLOG_INFO
> > +                "Host SYSTEM_SUSPEND blocked: serial unsupported\n");
> > +
> > +    return false;
> > +}
> > +
> > +bool host_system_suspend_allowed(void)
> > +{
> > +    return psci_system_suspend_allowed() &&
> > +           host_serial_suspend_allowed() &&
> > +           host_system_suspend_runtime_allowed;
> > +}
> > +
> > +void host_system_suspend_disable(const char *reason)
> > +{
> > +    host_system_suspend_runtime_allowed = false;
> > +
> > +    printk(XENLOG_INFO "Host SYSTEM_SUSPEND blocked: %s\n",
> > +           reason ? reason : "unsupported suspend/resume path");
> > +}
> > +
> > /*
> >  * Local variables:
> >  * mode: C
> > diff --git a/xen/arch/arm/vpsci.c b/xen/arch/arm/vpsci.c
> > index ac6af6118f..a41355d75d 100644
> > --- a/xen/arch/arm/vpsci.c
> > +++ b/xen/arch/arm/vpsci.c
> > @@ -5,6 +5,7 @@
> >
> > #include <asm/current.h>
> > #include <asm/domain.h>
> > +#include <asm/suspend.h>
> > #include <asm/vgic.h>
> > #include <asm/vpsci.h>
> > #include <asm/event.h>
> > @@ -219,6 +220,89 @@ static void do_psci_0_2_system_reset(void)
> >     domain_shutdown(d,SHUTDOWN_reboot);
> > }
> >
> > +/*
> > + * Serialise SYSTEM_SUSPEND policy decisions with the domain suspend transition,
> > + * so multiple control domains cannot all observe each other as still awake.
> > + */
> > +static DEFINE_SPINLOCK(vpsci_system_suspend_lock);
> > +
> > +static bool domain_in_suspend_state(struct domain *d)
> > +{
> > +    bool suspended;
> > +
> > +    spin_lock(&d->shutdown_lock);
> > +    suspended = domain_shutdown_completed(d) && (d->shutdown_code == SHUTDOWN_suspend);
>
> This lines is over 80 chars and should be broken down.

Ack.

>
> > +    spin_unlock(&d->shutdown_lock);
> > +
> > +    return suspended;
> > +}
> > +
> > +static int32_t domain_psci_system_suspend_policy(struct domain *d)
> > +{
> > +    struct domain *other;
> > +    bool last_awake_control_domain = true;
> > +    bool awake_non_control_domain = false;
> > +
> > +    /* Only control domains participate in sequencing policy. */
> > +    if ( !is_control_domain(d) )
> > +        return 0;
>
> I am wondering what would happen in a dom0less setup with no control
> domain if the domains request SYSTEM_SUSPEND.
>
> Could you explain how they would be resumed?
> Should we deny SYSTEM_SUSPEND requests in this case?

In the current implementation, a domain in SHUTDOWN_suspend needs
an explicit XEN_DOMCTL_resumedomain request to resume, for example
through "xl resume". A guest interrupt does not resume it.

This series relies on a control domain to coordinate resume,
including the order of backend and frontend domains. It does not
provide autonomous guest wakeup in a dom0less setup without a
control domain.

A separate policy could support dom0less systems with independent
guests. Each guest could resume on its own wake interrupt. Xen
could enter host suspend once all guests have completed suspend
and the platform is ready, then resume on a platform-supported
wake event. Such a mode could coexist with managed resume, but
it is not implemented in this series.

So yes, in v13 I will return PSCI_DENIED when a non-control domain
requests SYSTEM_SUSPEND and no control domain exists.
PSCI_FEATURES will remain unchanged.

This only checks that a control domain exists. It does not check
whether that domain has permission to resume this guest or will
remain available (it may be crashed, dying, or suspended).
The actual resumedomain request is still checked by XSM.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

* Re: [PATCH v12 13/13] xen/arm: Add host system suspend backend
  2026-09-28 16:19   ` Bertrand Marquis
@ 2026-09-30 22:10     ` Mykola Kvach
  0 siblings, 0 replies; 37+ messages in thread
From: Mykola Kvach @ 2026-09-30 22:10 UTC (permalink / raw)
  To: Bertrand Marquis
  Cc: Mykola Kvach, xen-devel@lists.xenproject.org, Stefano Stabellini,
	Julien Grall, Michal Orzel, Volodymyr Babchuk

Hi Bertrand,

On Mon, Sep 28, 2026 at 7:20 PM Bertrand Marquis
<Bertrand.Marquis@arm.com> wrote:
>
> Hi Mykola,
>
> > On 27 Aug 2026, at 16:32, Mykola Kvach <mykola_kvach@epam.com> wrote:
> >
> > From: Mirela Simonovic <mirela.simonovic@aggios.com>
> >
> > Add the Xen-wide suspend/resume backend used after a control-domain
> > vPSCI SYSTEM_SUSPEND request has been accepted. The vPSCI policy,
> > runtime driver blockers and control-domain sequencing checks are handled
> > by the preceding commit; this change adds the code that actually drives
> > the host suspend attempt.
> >
> > The backend runs from a tasklet scheduled on pCPU0, because non-boot CPUs
> > are disabled during suspend. It freezes domains, disables the scheduler
> > and then disables non-boot CPUs.
> >
> > Host-side suspend participants are handled in phases. IOMMU and console
> > state are suspended first. Local IRQs are then disabled before suspending
> > timer and GIC state. On resume or failure, the completed suspend phases
> > are unwound in reverse: GIC and timer state are restored while IRQs are
> > still disabled, local IRQs are restored, and then console and IOMMU state
> > are restored.
> >
> > On boot, init_ttbr is normally initialized during secondary CPU hotplug.
> > On uniprocessor systems this can leave init_ttbr uninitialized, so set it
> > from the boot CPU before entering suspend.
> >
> > Note: the code is behind CONFIG_HAS_SYSTEM_SUSPEND, which is currently
> > only selected when UNSUPPORTED is set and MPU is not set.
> >
> > Signed-off-by: Mirela Simonovic <mirela.simonovic@aggios.com>
> > Signed-off-by: Saeed Nowshadi <saeed.nowshadi@xilinx.com>
> > Signed-off-by: Mykyta Poturai <mykyta_poturai@epam.com>
> > Signed-off-by: Mykola Kvach <mykola_kvach@epam.com>
> > ---
> > Changes in V10:
> > - Re-apply boot CPU local errata/workaround handling after SYSTEM_SUSPEND,
> >  before resuming the rest of the host suspend path.
> > - Move set_init_ttbr() declaration to asm/mmu/mm.h, since it is
> >  MMU-specific.
> >
> > Changes in V9:
> > - Split vPSCI availability policy, runtime host-suspend blockers and the
> >  domain-readiness precheck into the preceding commit.
> > - Trigger the host suspend backend from the control-domain SYSTEM_SUSPEND
> >  path.
> > - Reorder the host suspend/resume phases so the timer is suspended with
> >  local IRQs disabled and local IRQs are restored after the GIC and timer
> >  resume paths, before the console and IOMMU resume paths.
> > - Move HAS_HWDOM_SYSTEM_SUSPEND and related logic to policy patch.
> >
> > Changes in V8:
> > - Add a pre-suspend check in system_suspend() after scheduler_disable() to
> >  require all domains to be in the shut down state with SHUTDOWN_suspend
> >  before proceeding with the global suspend flow.
> > - Drop the common-level depends on !ARM_64 || !SYSTEM_SUSPEND from
> >  CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND and model the ARM64 suspend case
> >  with an arch-selected capability instead.
> > - Rename CONFIG_HAS_HWDOM_SHUTDOWN_ON_SUSPEND to
> >  CONFIG_HAS_HWDOM_SYSTEM_SUSPEND.
> > - Rename need_hwdom_shutdown() to want_hwdom_shutdown().
> >
> > Changes in V7:
> > - Control domain is responsible for host suspend.
> > - Add an empty inline host_system_suspend() function when SYSTEM_SUSPEND
> >  config is disabled.
> > - Use IS_ENABLED() for config checking instead of #ifdef.
> > - Replace #ifdef checks in domain_shutdown() with IS_ENABLED() to simplify
> >  control flow.
> > - Factor hardware domain shutdown condition into a helper
> >  (need_hwdom_shutdown()) to avoid preprocessor directives inside the
> >  function.
> > - Squash with iommu suspend/resume commit.
> > ---
> > xen/arch/arm/Kconfig                 |   1 +
> > xen/arch/arm/cpuerrata.c             |   7 +-
> > xen/arch/arm/include/asm/cpuerrata.h |   1 +
> > xen/arch/arm/include/asm/mmu/mm.h    |   2 +
> > xen/arch/arm/include/asm/suspend.h   |   2 +
> > xen/arch/arm/mmu/smpboot.c           |   2 +-
> > xen/arch/arm/suspend.c               | 156 +++++++++++++++++++++++++++
> > xen/arch/arm/vpsci.c                 |  10 +-
> > 8 files changed, 177 insertions(+), 4 deletions(-)
> >
> > diff --git a/xen/arch/arm/Kconfig b/xen/arch/arm/Kconfig
> > index 9027aa17eb..da1585ec50 100644
> > --- a/xen/arch/arm/Kconfig
> > +++ b/xen/arch/arm/Kconfig
> > @@ -9,6 +9,7 @@ config ARM_64
> > select 64BIT
> > select HAS_DOMAIN_TYPE
> > select HAS_FAST_MULTIPLY
> > + select HAS_SYSTEM_SUSPEND if !MPU && UNSUPPORTED
> > select HAS_VPCI_GUEST_SUPPORT if PCI_PASSTHROUGH
> >
> > config ARM
> > diff --git a/xen/arch/arm/cpuerrata.c b/xen/arch/arm/cpuerrata.c
> > index 3a32183618..e6499aaab3 100644
> > --- a/xen/arch/arm/cpuerrata.c
> > +++ b/xen/arch/arm/cpuerrata.c
> > @@ -782,6 +782,11 @@ void check_local_cpu_errata(void)
> >     update_cpu_capabilities(arm_errata, "enabled workaround for");
> > }
> >
> > +int enable_local_cpu_errata_workarounds(void)
> > +{
> > +    return enable_nonboot_cpu_caps(arm_errata);
> > +}
> > +
> > void __init enable_errata_workarounds(void)
> > {
> >     enable_cpu_capabilities(arm_errata);
> > @@ -818,7 +823,7 @@ static int cpu_errata_callback(struct notifier_block *nfb,
> >          * fixed to expect an error at CPU_STARTING phase.
> >          */
> >         ASSERT(system_state != SYS_STATE_boot);
> > -        rc = enable_nonboot_cpu_caps(arm_errata);
> > +        rc = enable_local_cpu_errata_workarounds();
> >         break;
> >     default:
> >         break;
> > diff --git a/xen/arch/arm/include/asm/cpuerrata.h b/xen/arch/arm/include/asm/cpuerrata.h
> > index 1799a16d7e..b93521326f 100644
> > --- a/xen/arch/arm/include/asm/cpuerrata.h
> > +++ b/xen/arch/arm/include/asm/cpuerrata.h
> > @@ -5,6 +5,7 @@
> > #include <asm/alternative.h>
> >
> > void check_local_cpu_errata(void);
> > +int enable_local_cpu_errata_workarounds(void);
> > void enable_errata_workarounds(void);
> >
> > #define CHECK_WORKAROUND_HELPER(erratum, feature, arch)         \
> > diff --git a/xen/arch/arm/include/asm/mmu/mm.h b/xen/arch/arm/include/asm/mmu/mm.h
> > index 7f4d59137d..ee73a77777 100644
> > --- a/xen/arch/arm/include/asm/mmu/mm.h
> > +++ b/xen/arch/arm/include/asm/mmu/mm.h
> > @@ -110,6 +110,8 @@ void dump_pt_walk(paddr_t ttbr, paddr_t addr,
> > extern void switch_ttbr(uint64_t ttbr);
> > extern void relocate_and_switch_ttbr(uint64_t ttbr);
> >
> > +void set_init_ttbr(lpae_t *root);
> > +
> > #endif /* __ARM_MMU_MM_H__ */
> >
> > /*
> > diff --git a/xen/arch/arm/include/asm/suspend.h b/xen/arch/arm/include/asm/suspend.h
> > index 50dc6e9fdf..889a6509d9 100644
> > --- a/xen/arch/arm/include/asm/suspend.h
> > +++ b/xen/arch/arm/include/asm/suspend.h
> > @@ -41,11 +41,13 @@ int prepare_resume_ctx(void);
> > void hyp_resume(void);
> > bool host_system_suspend_allowed(void);
> > void host_system_suspend_disable(const char *reason);
> > +void host_system_suspend(struct domain *d);
> >
> > #else /* !CONFIG_SYSTEM_SUSPEND */
> >
> > static inline bool host_system_suspend_allowed(void) { return false; }
> > static inline void host_system_suspend_disable(const char *reason) {}
> > +static inline void host_system_suspend(struct domain *d) {}
> >
> > #endif
> >
> > diff --git a/xen/arch/arm/mmu/smpboot.c b/xen/arch/arm/mmu/smpboot.c
> > index 37e91d72b7..ff508ecf40 100644
> > --- a/xen/arch/arm/mmu/smpboot.c
> > +++ b/xen/arch/arm/mmu/smpboot.c
> > @@ -72,7 +72,7 @@ static void clear_boot_pagetables(void)
> >     clear_table(boot_third);
> > }
> >
> > -static void set_init_ttbr(lpae_t *root)
> > +void set_init_ttbr(lpae_t *root)
> > {
> >     /*
> >      * init_ttbr is part of the identity mapping which is read-only. So
> > diff --git a/xen/arch/arm/suspend.c b/xen/arch/arm/suspend.c
> > index c7c26bcf03..3fe2ffa4fb 100644
> > --- a/xen/arch/arm/suspend.c
> > +++ b/xen/arch/arm/suspend.c
> > @@ -1,10 +1,18 @@
> > /* SPDX-License-Identifier: GPL-2.0-only */
> >
> > +#include <asm/cpuerrata.h>
> > +#include <asm/cpufeature.h>
> > +#include <asm/gic.h>
> > #include <asm/psci.h>
> > #include <asm/suspend.h>
> >
> > +#include <xen/console.h>
> > +#include <xen/cpu.h>
> > +#include <xen/iommu.h>
> > #include <xen/lib.h>
> > +#include <xen/sched.h>
> > #include <xen/serial.h>
> > +#include <xen/tasklet.h>
> >
> > struct resume_cpu_context resume_cpu_context;
> >
> > @@ -44,6 +52,154 @@ void host_system_suspend_disable(const char *reason)
> >            reason ? reason : "unsupported suspend/resume path");
> > }
> >
> > +/* Xen suspend. data identifies the domain that initiated suspend. */
> > +static void system_suspend(void *data)
> > +{
> > +    int status;
> > +    unsigned long flags;
> > +    struct domain *d = (struct domain *)data;
> > +
> > +    BUG_ON(system_state != SYS_STATE_active);
> > +
> > +    system_state = SYS_STATE_suspend;
> > +
> > +    printk("Xen suspending...\n");
> > +
> > +    freeze_domains();
> > +    scheduler_disable();
> > +
> > +    /*
> > +     * Non-boot CPUs have to be disabled on suspend and enabled on resume
> > +     * (hotplug-based mechanism). Disabling non-boot CPUs will lead to PSCI
> > +     * CPU_OFF to be called by each non-boot CPU. Depending on the underlying
> > +     * platform capabilities, this may lead to the physical powering down of
> > +     * CPUs.
> > +     */
> > +    status = disable_nonboot_cpus();
> > +    if ( status )
> > +    {
> > +        system_state = SYS_STATE_resume;
> > +        goto resume_nonboot_cpus;
> > +    }
> > +
> > +    console_start_sync();
> > +    status = iommu_suspend();
> > +    if ( status )
> > +    {
> > +        system_state = SYS_STATE_resume;
> > +        goto resume_end_sync;
> > +    }
> > +
> > +    status = console_suspend();
> > +    if ( status )
> > +    {
> > +        dprintk(XENLOG_ERR, "Failed to suspend the console, err=%d\n", status);
> > +        system_state = SYS_STATE_resume;
> > +        goto resume_iommu;
> > +    }
> > +
> > +    local_irq_save(flags);
> > +
> > +    time_suspend();
> > +
> > +    status = gic_suspend();
> > +    if ( status )
> > +    {
> > +        system_state = SYS_STATE_resume;
> > +        goto resume_time;
> > +    }
> > +
> > +    set_init_ttbr(xen_pgtable);
> > +
> > +    /*
> > +     * Enable identity mapping before entering suspend to simplify
> > +     * the resume path
> > +     */
> > +    update_boot_mapping(true);
> > +
> > +    if ( prepare_resume_ctx() )
> > +    {
> > +        status = call_psci_system_suspend();
> > +        /*
> > +         * If suspend is finalized properly by above system suspend PSCI call,
> > +         * the code below in this 'if' branch will never execute. Execution
> > +         * will continue from hyp_resume which is the hypervisor's resume point.
> > +         * In hyp_resume CPU context will be restored and since link-register is
> > +         * restored as well, it will appear to return from prepare_resume_ctx.
> > +         * The difference in returning from prepare_resume_ctx on system suspend
> > +         * versus resume is in function's return value: on suspend, the return
> > +         * value is a non-zero value, on resume it is zero. That is why the
> > +         * control flow will not re-enter this 'if' branch on resume.
> > +         */
> > +        if ( status )
> > +            dprintk(XENLOG_WARNING, "PSCI system suspend failed, err=%d\n",
> > +                    status);
> > +
> > +        system_state = SYS_STATE_resume;
> > +    }
> > +    else
> > +    {
> > +        system_state = SYS_STATE_resume;
> > +
> > +        /*
> > +         * CPU0 resumes directly from hyp_resume(), bypassing the CPU hotplug
> > +         * path that re-checks and re-enables errata workarounds for secondary
> > +         * CPUs.
> > +         */
> > +        check_local_cpu_errata();
> > +        check_local_cpu_features();
> > +        BUG_ON(enable_local_cpu_errata_workarounds());
> > +    }
> > +
> > +    update_boot_mapping(false);
> > +
> > +    gic_resume();
> > +
> > + resume_time:
> > +    time_resume();
> > +
> > +    local_irq_restore(flags);
> > +
> > +    console_resume();
> > +
> > + resume_iommu:
> > +    iommu_resume();
> > +
> > + resume_end_sync:
> > +    console_end_sync();
> > +
> > + resume_nonboot_cpus:
> > +    /*
> > +     * The rcu_barrier() has to be added to ensure that the per cpu area is
> > +     * freed before a non-boot CPU tries to initialize it (_free_percpu_area()
> > +     * has to be called before the init_percpu_area()). This scenario occurs
> > +     * when non-boot CPUs are hot-unplugged on suspend and hotplugged on resume.
>
> This line would need wrapping as it is over 80 chars.
>
> > +     */
> > +    rcu_barrier();
>
> Could you clarify which per-CPU area needs freeing here?
>
> system_state is SYS_STATE_suspend for every CPU taken down by
> disable_nonboot_cpus(), so cpu_percpu_callback() does not queue
> _free_percpu_area() for any of them.
>
> So I do not quite get what your comment case actually is.
> Could you explain ?

You are right. I checked the original patch and the code it was
based on.

The barrier was added in the 2018 series. At that time, the Arm
CPU_DEAD callback queued _free_percpu_area() even during system
suspend. If resume reached init_percpu_area() before that callback
completed, it would return -EBUSY because the old area still
existed. This would trigger BUG_ON() in enable_nonboot_cpus().

The barrier ensured that the old area was freed before bringing
the CPU back up.

That reason no longer applies. The current code keeps the per-CPU
areas during suspend and reuses them on resume.

Also, since commit 540d4d60378c ("cpu: sync any remaining RCU
callbacks before CPU up/down"), cpu_up() already calls
rcu_barrier() through cpu_hotplug_begin(), before CPU_UP_PREPARE.

The separate barrier and its comment were carried over after these
changes. I will remove both in v13. This also removes the line
that needed wrapping.

Best regards,
Mykola


^ permalink raw reply	[flat|nested] 37+ messages in thread

end of thread, other threads:[~2026-09-30 22:10 UTC | newest]

Thread overview: 37+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-27 14:31 [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach
2026-08-27 14:31 ` [PATCH v12 01/13] xen/arm: Add suspend and resume timer helpers Mykola Kvach
2026-08-27 14:31 ` [PATCH v12 02/13] xen/arm: gic-v2: Implement GIC suspend/resume functions Mykola Kvach
2026-09-23 15:27   ` Bertrand Marquis
2026-09-24 22:23     ` Mykola Kvach
2026-09-28  7:36       ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 03/13] xen/arm: gic-v3: tolerate retained redistributor LPI state across CPU_OFF Mykola Kvach
2026-09-23 15:34   ` Bertrand Marquis
2026-09-24 23:22     ` Mykola Kvach
2026-09-28  7:37       ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 04/13] xen/arm: gic-v3: Implement GICv3 suspend/resume functions Mykola Kvach
2026-09-23 15:35   ` Bertrand Marquis
2026-09-25  0:03     ` Mykola Kvach
2026-09-28  7:39       ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 05/13] xen/arm: gic-v3: add ITS suspend/resume support Mykola Kvach
2026-09-23 15:36   ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 06/13] xen/arm: tee: keep init_tee_secondary() for hotplug and resume Mykola Kvach
2026-08-27 14:31 ` [PATCH v12 07/13] xen/arm: ffa: fix notification SRI across CPU hotplug/suspend Mykola Kvach
2026-08-27 14:31 ` [PATCH v12 08/13] iommu/ipmmu-vmsa: Implement suspend/resume callbacks Mykola Kvach
2026-09-28  8:01   ` Mykola Kvach
2026-09-28  9:33     ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 09/13] xen/arm: smmu-v3: add suspend/resume handlers Mykola Kvach
2026-09-28 16:17   ` Bertrand Marquis
2026-09-30 14:44     ` Mykola Kvach
2026-08-27 14:31 ` [PATCH v12 10/13] xen/arm64: Save/restore CPU context across SYSTEM_SUSPEND Mykola Kvach
2026-09-28 16:17   ` Bertrand Marquis
2026-08-27 14:31 ` [PATCH v12 11/13] xen/arm: Implement PSCI SYSTEM_SUSPEND call (host interface) Mykola Kvach
2026-09-28 16:18   ` Bertrand Marquis
2026-09-30 17:32     ` Mykola Kvach
2026-08-27 14:32 ` [PATCH v12 12/13] xen/arm: Add vPSCI SYSTEM_SUSPEND policy Mykola Kvach
2026-09-28 16:18   ` Bertrand Marquis
2026-09-30 20:42     ` Mykola Kvach
2026-08-27 14:32 ` [PATCH v12 13/13] xen/arm: Add host system suspend backend Mykola Kvach
2026-08-27 21:59   ` Volodymyr Babchuk
2026-09-28 16:19   ` Bertrand Marquis
2026-09-30 22:10     ` Mykola Kvach
2026-09-22  7:04 ` Ping: [PATCH v12 00/13] Add initial Xen Suspend-to-RAM support on ARM64 Mykola Kvach

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.