All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver
@ 2026-09-30 14:52 Georgi Djakov
  2026-09-30 14:52 ` [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding Georgi Djakov
                   ` (5 more replies)
  0 siblings, 6 replies; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

QPaCE (Qualcomm Page Compression Engine) is a hardware block found on
Qualcomm SoCs that accelerates page compression and decompression for
use with the zram block device. It is composed of multiple compression,
decompression, and copy engines.

This series uses QPaCE's "urgent" (synchronous) path, intended for low-
latency single-page operations. A command is issued by writing the
source and destination DMA addresses to one of several per-CPU urgent
command registers with a single store; completion is then detected by
polling the command's status register.

The hardware also provides a ring-based (asynchronous) path for high-
throughput batched compression, where transfer and event rings allow
multiple descriptors to be queued and submitted with a single doorbell
write, with completion signalled via an interrupt. This path is not
implemented in this series.

The driver integrates with zram through a new "qpace-lz4" zcomp backend.
Both compression and decompression are performed by the QPaCE hardware
via the urgent synchronous path: each request DMA-maps the source and
destination buffers, issues a blocking hardware command, and returns the
result size.

The driver also integrates with the Qualcomm Last Level Cache Controller
(LLCC) to activate dedicated cache slices for compression and decompression
data.

This is validated on next-20260929 with the following patch applied on
top to allow booting completely to the serial console:
https://lore.kernel.org/all/20260911-rsc_read-v2-1-98675c248278@oss.qualcomm.com/

Georgi Djakov (6):
  dt-bindings: soc: qcom: Add QPaCE binding
  soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  trace: qpace: Add tracepoints for QPaCE operations
  zram: Add QPaCE zcomp backend
  soc: qcom: qpace: Add LLCC slice support
  arm64: dts: qcom: hawi: Add QPaCE DT node

 .../bindings/soc/qcom/qcom,hawi-qpace.yaml    |  94 ++
 arch/arm64/boot/dts/qcom/hawi.dtsi            |  17 +
 drivers/block/zram/Kconfig                    |  11 +
 drivers/block/zram/Makefile                   |   1 +
 drivers/block/zram/backend_qpace.c            | 134 +++
 drivers/block/zram/backend_qpace.h            |  13 +
 drivers/block/zram/zcomp.c                    |  11 +-
 drivers/soc/qcom/Kconfig                      |  14 +
 drivers/soc/qcom/Makefile                     |   1 +
 drivers/soc/qcom/qpace.c                      | 830 ++++++++++++++++++
 drivers/soc/qcom/qpace_internal.h             |  84 ++
 include/linux/soc/qcom/llcc-qcom.h            |   2 +
 include/linux/soc/qcom/qpace.h                | 154 ++++
 include/trace/events/qpace.h                  |  99 +++
 14 files changed, 1463 insertions(+), 2 deletions(-)
 create mode 100644 Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml
 create mode 100644 drivers/block/zram/backend_qpace.c
 create mode 100644 drivers/block/zram/backend_qpace.h
 create mode 100644 drivers/soc/qcom/qpace.c
 create mode 100644 drivers/soc/qcom/qpace_internal.h
 create mode 100644 include/linux/soc/qcom/qpace.h
 create mode 100644 include/trace/events/qpace.h


^ permalink raw reply	[flat|nested] 25+ messages in thread

* [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-10-02  6:10   ` Krzysztof Kozlowski
  2026-09-30 14:52 ` [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver Georgi Djakov
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

Add device-tree binding documentation for the Qualcomm Page Compression
Engine (QPaCE), including the resources needed to describe it in SoC device
trees.

QPaCE accelerates page compression and decompression for use with the zram
block device.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 .../bindings/soc/qcom/qcom,hawi-qpace.yaml    | 94 +++++++++++++++++++
 1 file changed, 94 insertions(+)
 create mode 100644 Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml

diff --git a/Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml b/Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml
new file mode 100644
index 000000000000..b7ae50f34f16
--- /dev/null
+++ b/Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml
@@ -0,0 +1,94 @@
+# SPDX-License-Identifier: (GPL-2.0-only OR BSD-2-Clause)
+%YAML 1.2
+---
+$id: http://devicetree.org/schemas/soc/qcom/qcom,hawi-qpace.yaml#
+$schema: http://devicetree.org/meta-schemas/core.yaml#
+
+title: Qualcomm Page Compression Engine (QPaCE)
+
+maintainers:
+  - Georgi Djakov <georgi.djakov@oss.qualcomm.com>
+
+description:
+  QPaCE is a page-oriented hardware compression engine found on Qualcomm
+  SoCs. It is composed of a number of compression, decompression and copy
+  engines and exposes both a low-latency urgent path for synchronous
+  single-page operations and a high-throughput ring-based path for batched
+  asynchronous requests.
+
+properties:
+  compatible:
+    items:
+      - enum:
+          - qcom,hawi-qpace
+
+  reg:
+    maxItems: 1
+    description:
+      Single contiguous register window covering all QPaCE sub-regions
+      from the GEN registers base through the EE registers.
+
+  iommus:
+    items:
+      - description: IOMMU mapping for compression engine read DMA
+      - description: IOMMU mapping for compression engine write DMA
+      - description: IOMMU mapping for decompression engine read DMA
+      - description: IOMMU mapping for decompression engine write DMA
+
+  dma-coherent: true
+
+  interconnects:
+    maxItems: 1
+
+  interconnect-names:
+    items:
+      - const: qpace-mem
+
+  interrupts:
+    items:
+      - description: Ring completion and bus error interrupt
+      - description: Urgent command completion interrupt
+
+  interrupt-names:
+    items:
+      - const: ring-and-bus-err
+      - const: urgent
+
+required:
+  - compatible
+  - reg
+  - iommus
+  - interconnects
+  - interconnect-names
+  - interrupts
+  - interrupt-names
+
+additionalProperties: false
+
+examples:
+  - |
+    #include <dt-bindings/interrupt-controller/arm-gic.h>
+    #include <dt-bindings/interconnect/qcom,icc.h>
+    #include <dt-bindings/interconnect/qcom,hawi-rpmh.h>
+
+    soc {
+        #address-cells = <2>;
+        #size-cells = <2>;
+
+        compression@31400000 {
+            compatible = "qcom,hawi-qpace";
+            reg = <0x0 0x31400000 0x0 0x80000>;
+            iommus = <&apps_smmu 0x1c20 0x13>,
+                     <&apps_smmu 0x1c24 0x13>,
+                     <&apps_smmu 0x1c28 0x13>,
+                     <&apps_smmu 0x1c2c 0x13>;
+            dma-coherent;
+            interconnects = <&gem_noc MASTER_QPACE QCOM_ICC_TAG_ALWAYS
+                             &mc_virt SLAVE_EBI1 QCOM_ICC_TAG_ALWAYS>;
+            interconnect-names = "qpace-mem";
+            interrupts = <GIC_ESPI 72 IRQ_TYPE_LEVEL_HIGH>,
+                         <GIC_ESPI 71 IRQ_TYPE_LEVEL_HIGH>;
+            interrupt-names = "ring-and-bus-err",
+                              "urgent";
+        };
+    };

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
  2026-09-30 14:52 ` [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-09-30 15:04   ` sashiko-bot
  2026-10-01  8:50   ` Krzysztof Kozlowski
  2026-09-30 14:52 ` [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations Georgi Djakov
                   ` (3 subsequent siblings)
  5 siblings, 2 replies; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

Add a platform driver for the Qualcomm Page Compression Engine (QPaCE), a
hardware block that accelerates compression and decompression of memory
pages.

Provide the urgent command path for synchronous single-page compression and
decompression. This exposes the low-latency operations needed by
compressed-memory users such as zram, especially for page decompression on
the read path.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 drivers/soc/qcom/Kconfig          |  14 +
 drivers/soc/qcom/Makefile         |   1 +
 drivers/soc/qcom/qpace.c          | 764 ++++++++++++++++++++++++++++++
 drivers/soc/qcom/qpace_internal.h |  84 ++++
 include/linux/soc/qcom/qpace.h    | 154 ++++++
 5 files changed, 1017 insertions(+)
 create mode 100644 drivers/soc/qcom/qpace.c
 create mode 100644 drivers/soc/qcom/qpace_internal.h
 create mode 100644 include/linux/soc/qcom/qpace.h

diff --git a/drivers/soc/qcom/Kconfig b/drivers/soc/qcom/Kconfig
index 535c8619197b..6bcb86dcd726 100644
--- a/drivers/soc/qcom/Kconfig
+++ b/drivers/soc/qcom/Kconfig
@@ -288,6 +288,20 @@ config QCOM_PBS
 	  This module provides the APIs to the client drivers that wants to send the
 	  PBS trigger event to the PBS RAM.
 
+config QCOM_PAGE_COMPRESSION_ENGINE
+	tristate "Qualcomm Page Compression Engine (QPaCE)"
+	depends on ARM64
+	depends on ARCH_QCOM || COMPILE_TEST
+	depends on OF
+	depends on INTERCONNECT
+	help
+	  Enable support for the Qualcomm Page Compression Engine (QPaCE),
+	  a hardware accelerator that provides high-throughput page compression,
+	  decompression, and DMA copy operations.
+
+	  The engine is used as a hardware backend for compressed-memory
+	  subsystems such as zram. If unsure, say N.
+
 endif
 
 # Options selected by other drivers from different subsystems must be outside
diff --git a/drivers/soc/qcom/Makefile b/drivers/soc/qcom/Makefile
index bddd6f0a5836..45fe697cf2d4 100644
--- a/drivers/soc/qcom/Makefile
+++ b/drivers/soc/qcom/Makefile
@@ -42,3 +42,4 @@ qcom_ice-objs			+= ice.o
 obj-$(CONFIG_QCOM_INLINE_CRYPTO_ENGINE)	+= qcom_ice.o
 obj-$(CONFIG_QCOM_PBS) +=	qcom-pbs.o
 obj-$(CONFIG_QCOM_UBWC_CONFIG) += ubwc_config.o
+obj-$(CONFIG_QCOM_PAGE_COMPRESSION_ENGINE) += qpace.o
diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
new file mode 100644
index 000000000000..69aa14b7c477
--- /dev/null
+++ b/drivers/soc/qcom/qpace.c
@@ -0,0 +1,764 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#include <dt-bindings/interconnect/qcom,icc.h>
+#include <linux/delay.h>
+#include <linux/dma-mapping.h>
+#include <linux/export.h>
+#include <linux/interconnect.h>
+#include <linux/interrupt.h>
+#include <linux/io.h>
+#include <linux/iopoll.h>
+#include <linux/mm.h>
+#include <linux/module.h>
+#include <linux/mutex.h>
+#include <linux/of.h>
+#include <linux/of_platform.h>
+#include <linux/platform_device.h>
+#include <linux/pm_qos.h>
+#include <linux/soc/qcom/qpace.h>
+#include <linux/workqueue.h>
+
+#include "qpace_internal.h"
+
+#define QPACE_REG_PAGE_SIZE 4096
+#define QPACE_CORE_CACHEINDEX_SIZE 4
+
+struct qpace_priv {
+	void __iomem *gen_regs;
+	void __iomem *gen_core_regs;
+	void __iomem *comp_core_regs;
+	void __iomem *decomp_core_regs;
+	void __iomem *urg_regs;
+
+	struct device *dev;
+	u32 hw_version;
+	bool suspended;
+
+	struct icc_path *interconnect;
+	struct pm_qos_request qos_req;
+
+	int active_rings;
+	struct completion no_active_refs;
+
+	bool broken;
+	struct work_struct disable_work;
+};
+
+static inline u32 qpace_read_gen(struct qpace_priv *p, u32 offset)
+{
+	return readl(p->gen_regs + offset);
+}
+
+static inline void qpace_write_gen(struct qpace_priv *p, u32 offset, u32 val)
+{
+	writel(val, p->gen_regs + offset);
+}
+
+static inline u32 qpace_read_gen_core(struct qpace_priv *p, u32 offset)
+{
+	return readl(p->gen_core_regs + offset);
+}
+
+static inline void qpace_write_gen_core(struct qpace_priv *p, u32 offset, u32 val)
+{
+	writel(val, p->gen_core_regs + offset);
+}
+
+static inline u32 qpace_read_comp_core(struct qpace_priv *p, u32 offset)
+{
+	return readl(p->comp_core_regs + offset);
+}
+
+static inline void qpace_write_comp_core(struct qpace_priv *p, u32 offset, u32 val)
+{
+	writel(val, p->comp_core_regs + offset);
+}
+
+static inline u32 qpace_read_urg_cacheindex(struct qpace_priv *p, int idx, u32 offset)
+{
+	return readl(p->gen_regs + idx * QPACE_CORE_CACHEINDEX_SIZE + offset);
+}
+
+static inline void qpace_write_urg_cacheindex(struct qpace_priv *p, int idx, u32 offset, u32 val)
+{
+	writel(val, p->gen_regs + idx * QPACE_CORE_CACHEINDEX_SIZE + offset);
+}
+
+static inline u32 qpace_read_decomp_core(struct qpace_priv *p, u32 offset)
+{
+	return readl(p->decomp_core_regs + offset);
+}
+
+static inline void qpace_write_decomp_core(struct qpace_priv *p, u32 offset, u32 val)
+{
+	writel(val, p->decomp_core_regs + offset);
+}
+
+static inline u32 qpace_read_urg_cmd(struct qpace_priv *p, int idx, u32 offset)
+{
+	return readl(p->urg_regs + idx * QPACE_REG_PAGE_SIZE + offset);
+}
+
+static inline void qpace_write_urg_cmd(struct qpace_priv *p, int idx, u32 offset, u32 val)
+{
+	writel(val, p->urg_regs + idx * QPACE_REG_PAGE_SIZE + offset);
+}
+
+static inline void qpace_write_urg_cmd_ctx(struct qpace_priv *p, u32 offset,
+					   int urg_reg_num, int ctx_num, u32 val)
+{
+	writel(val, p->urg_regs + urg_reg_num * QPACE_REG_PAGE_SIZE +
+	       offset + ctx_num * URG_CMD_CONTEXT_SPACING);
+}
+
+static struct qpace_priv *qpace_priv;
+
+static DEFINE_STATIC_KEY_FALSE(qpace_drv_probed);
+
+static void qpace_disable_work_fn(struct work_struct *work)
+{
+	static_branch_disable(&qpace_drv_probed);
+}
+
+enum urg_reg_cxts {
+	LZ4_URG_COMP_CNTXT,
+	LZ4_URG_DECOMP_CNTXT,
+};
+
+const struct qpace_algorithm qpace_lz4_algorithm = {
+	.name = "qpace-lz4",
+	.comp_opcode = LZ4_COMP,
+	.decomp_opcode = LZ4_DECOMP,
+	.urg_comp_cntxt = LZ4_URG_COMP_CNTXT,
+	.urg_decomp_cntxt = LZ4_URG_DECOMP_CNTXT,
+};
+EXPORT_SYMBOL_GPL(qpace_lz4_algorithm);
+
+bool qpace_is_valid_algorithm(const char *algo_name)
+{
+	if (!static_branch_likely(&qpace_drv_probed))
+		return false;
+	if (READ_ONCE(qpace_priv->broken))
+		return false;
+
+	return algo_name && !strcmp(algo_name, qpace_lz4_algorithm.name);
+}
+EXPORT_SYMBOL_GPL(qpace_is_valid_algorithm);
+
+static inline void program_urg_comp_context(int urg_reg_num, const struct qpace_algorithm *algo)
+{
+	u32 urg_cmd_settings = 0;
+	int page_cnt = (1 << (PAGE_SHIFT - 12)) - 1;
+
+	/* Set the limit of the compression output size */
+	urg_cmd_settings = FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, PAGE_SIZE - 1);
+	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
+				urg_reg_num, algo->urg_comp_cntxt,
+				urg_cmd_settings);
+
+	urg_cmd_settings = FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_OPER, algo->comp_opcode);
+
+	/* Set page count for urgent compression input */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_PAGE_CNT, page_cnt);
+
+	/* Program part of the SMMU input and output SIDs */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_SCTX,
+				       DEFAULT_SMMU_CONTEXT);
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_DCTX,
+				       DEFAULT_SMMU_CONTEXT);
+
+	/* Cache accesses in the system cache during compression */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_WA, 1);
+
+	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_MISC_n_OFFSET,
+				urg_reg_num, algo->urg_comp_cntxt,
+				urg_cmd_settings);
+}
+
+static inline void program_urg_decomp_context(int urg_reg_num, const struct qpace_algorithm *algo)
+{
+	u32 urg_cmd_settings = 0;
+	int page_cnt = (1 << (PAGE_SHIFT - 12)) - 1;
+
+	/* The input size will be programmed by a requester */
+	urg_cmd_settings = FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, 0);
+	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
+				urg_reg_num, algo->urg_decomp_cntxt,
+				urg_cmd_settings);
+
+	urg_cmd_settings = FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_OPER, algo->decomp_opcode);
+
+	/* Configure page count which will limit decompression output size */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_PAGE_CNT, page_cnt);
+
+	/* Program part of the SMMU input and output SIDs */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_SCTX,
+				       DEFAULT_SMMU_CONTEXT);
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_DCTX,
+				       DEFAULT_SMMU_CONTEXT);
+
+	/* Cache accesses in the system cache during decompression */
+	urg_cmd_settings |= FIELD_PREP(URG_CMD_0_CFG_CNTXT_MISC_WA, 1);
+
+	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_MISC_n_OFFSET,
+				urg_reg_num, algo->urg_decomp_cntxt,
+				urg_cmd_settings);
+}
+
+static void program_urg_cmd_cacheindex(int urg_reg_num)
+{
+	u32 reg_val;
+
+	reg_val = qpace_read_urg_cacheindex(qpace_priv, urg_reg_num,
+					    QPACE_CORE_URG_CMD_n_CACHEINDEX_CFG_OFFSET);
+	reg_val = u32_replace_bits(reg_val, 0x24, CORE_URG_CMD_n_CACHEINDEX__ENGINE_CACHEINDEX);
+	reg_val = u32_replace_bits(reg_val, 0x1,
+				   CORE_URG_CMD_n_CACHEINDEX__ENGINE_CACHEINDEX_OVERRIDE);
+	qpace_write_urg_cacheindex(qpace_priv, urg_reg_num,
+				   QPACE_CORE_URG_CMD_n_CACHEINDEX_CFG_OFFSET, reg_val);
+}
+
+static void program_decomp_core_cfg(void)
+{
+	u32 reg_val;
+
+	reg_val = qpace_read_decomp_core(qpace_priv, QPACE_DECOMP_CORE_CFG_OFFSET);
+	reg_val = u32_replace_bits(reg_val, 0x8, DECOMP_CORE_CFG_DMA_RD_MAX_OT);
+	reg_val = u32_replace_bits(reg_val, 0x16, DECOMP_CORE_CFG_DMA_WR_MAX_OT);
+	/*
+	 * MEM_MNG_PAGE_SIZE controls the prefetch boundary:
+	 * 0: 4 KB, 1: 8 KB, 2: 16 KB, 3: 32 KB, 4: 64 KB
+	 */
+	reg_val = u32_replace_bits(reg_val, 0x0, DECOMP_CORE_CFG_DECOMP_DMA_MEM_MNG_PAGE_SIZE);
+	qpace_write_decomp_core(qpace_priv, QPACE_DECOMP_CORE_CFG_OFFSET, reg_val);
+}
+
+static void program_urg_command_contexts_v2(void)
+{
+	int urg_reg_num;
+
+	/* Configure the needed contexts for each of the urgent command registers */
+	for (urg_reg_num = 0; urg_reg_num < NUM_TRS_ERS_URG_CMD_REGS; urg_reg_num++) {
+		program_urg_cmd_cacheindex(urg_reg_num);
+		program_urg_comp_context(urg_reg_num, &qpace_lz4_algorithm);
+		program_urg_decomp_context(urg_reg_num, &qpace_lz4_algorithm);
+	}
+}
+
+static inline int qpace_urgent_command_trigger(dma_addr_t input_addr,
+					       dma_addr_t output_addr,
+					       int urg_reg_num,
+					       enum urg_reg_cxts command)
+{
+	void __iomem *td_dst_src_reg = qpace_priv->urg_regs +
+				       (urg_reg_num * QPACE_REG_PAGE_SIZE) +
+					QPACE_URG_CMD_0_TD_DST_ADDR_L_CFG_CNTXT_OFFSET;
+	u64 urg_addr_field_lower, urg_addr_field_upper;
+	u32 stat_reg;
+	unsigned long ret;
+
+	urg_addr_field_lower = FIELD_PREP(URG_CMD_0_TD_DST_ADDR_L__CMD_CFG_CNTXT,
+					  command);
+	urg_addr_field_lower |= GENMASK(63, 8) & output_addr;
+
+	urg_addr_field_upper = input_addr;
+
+	/* Clear a stale request-on-active error before ringing the doorbell. */
+	qpace_write_urg_cmd(qpace_priv, urg_reg_num, QPACE_URG_CMD_0_STAT_CLR_OFFSET,
+			    URG_CMD_0_STAT_CLR_REQ_ON_ACTIVE_ERR);
+
+	/* Ensure that preceding stores that QPaCE will depend on are done executing */
+	dma_wmb();
+
+	/*
+	 * Ring the doorbell with a single 128-bit atomic store. QPaCE
+	 * triggers processing when it observes the (dst_addr, src_addr)
+	 * pair land together.
+	 */
+	asm volatile("stp %0, %1, [%2]\n" :
+		     : "r" (urg_addr_field_lower), "r" (urg_addr_field_upper), "r" (td_dst_src_reg)
+		     : "memory");
+
+	ret = readl_poll_timeout_atomic(qpace_priv->urg_regs +
+					urg_reg_num * QPACE_REG_PAGE_SIZE +
+					QPACE_URG_CMD_0_ED_STAT_OFFSET,
+					stat_reg,
+					FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE,
+						  stat_reg) != OP_URG_ONGOING,
+					1, 100 * USEC_PER_MSEC);
+	if (ret) {
+		dev_err(qpace_priv->dev, "QPaCE urgent command timed out\n");
+		return -ETIMEDOUT;
+	}
+
+	return stat_reg;
+}
+
+int qpace_urgent_compress(dma_addr_t input_addr,
+			  dma_addr_t output_addr,
+			  struct qpace_algorithm *algo)
+{
+	int urg_reg_num;
+	int stat_reg;
+	u32 stat_reg_val;
+	int ret;
+
+	ret = qpace_get();
+	if (ret)
+		return ret;
+
+	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
+	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
+						algo->urg_comp_cntxt);
+	put_cpu();
+
+	qpace_put();
+
+	if (stat_reg < 0) {
+		ret = stat_reg;
+		goto out;
+	}
+
+	stat_reg_val = FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE, stat_reg);
+
+	if (stat_reg_val == OP_COMP_TOO_BIG) {
+		ret = -E2BIG;
+		goto out;
+	}
+
+	if (stat_reg_val != OP_OK) {
+		pr_err("%s: register %d failed with %u\n",
+		       __func__, urg_reg_num, stat_reg_val);
+		ret = -EINVAL;
+		goto out;
+	}
+
+	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
+out:
+	return ret;
+}
+EXPORT_SYMBOL_GPL(qpace_urgent_compress);
+
+int qpace_urgent_decompress(dma_addr_t input_addr,
+			    dma_addr_t output_addr,
+			    size_t input_size,
+			    struct qpace_algorithm *algo)
+{
+	int urg_reg_num;
+	int stat_reg;
+	u32 stat_reg_val;
+	int ret;
+
+	ret = qpace_get();
+	if (ret)
+		goto out;
+
+	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
+	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
+				urg_reg_num, algo->urg_decomp_cntxt,
+				FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, input_size));
+	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
+						algo->urg_decomp_cntxt);
+	put_cpu();
+
+	qpace_put();
+
+	if (stat_reg < 0) {
+		ret = stat_reg;
+		goto out;
+	}
+
+	stat_reg_val = FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE, stat_reg);
+	if (stat_reg_val != OP_OK) {
+		pr_err("%s: register %d failed with %u\n",
+		       __func__, urg_reg_num, stat_reg_val);
+		ret = -EINVAL;
+		goto out;
+	}
+
+	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
+out:
+	return ret;
+}
+EXPORT_SYMBOL_GPL(qpace_urgent_decompress);
+
+static DEFINE_MUTEX(qpace_ref_lock);
+
+static void _get_qpace(void)
+{
+	lockdep_assert_held(&qpace_ref_lock);
+	if (!qpace_priv->active_rings) {
+		reinit_completion(&qpace_priv->no_active_refs);
+		pm_stay_awake(qpace_priv->dev);
+		cpu_latency_qos_update_request(&qpace_priv->qos_req, 300);
+		program_urg_command_contexts_v2();
+		program_decomp_core_cfg();
+	}
+	qpace_priv->active_rings++;
+}
+
+static void _put_qpace(void)
+{
+	lockdep_assert_held(&qpace_ref_lock);
+	if (!--qpace_priv->active_rings) {
+		cpu_latency_qos_update_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
+		pm_relax(qpace_priv->dev);
+		complete(&qpace_priv->no_active_refs);
+	}
+}
+
+int qpace_get(void)
+{
+	int ret = 0;
+
+	mutex_lock(&qpace_ref_lock);
+	if (qpace_priv->suspended || READ_ONCE(qpace_priv->broken))
+		ret = -EBUSY;
+	else
+		_get_qpace();
+	mutex_unlock(&qpace_ref_lock);
+	return ret;
+}
+EXPORT_SYMBOL_GPL(qpace_get);
+
+void qpace_put(void)
+{
+	mutex_lock(&qpace_ref_lock);
+	_put_qpace();
+	mutex_unlock(&qpace_ref_lock);
+}
+EXPORT_SYMBOL_GPL(qpace_put);
+
+static irqreturn_t urgent_interrupt_handler(int irq, void *unused)
+{
+	pr_debug("Urgent interrupt handled\n");
+	return IRQ_HANDLED;
+}
+
+static int qpace_hw_init(void)
+{
+	u32 reg_val;
+
+	/* Select CPU SCID for our system cache slice. */
+	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET);
+	reg_val = u32_replace_bits(reg_val, 0x1, CORE_QNS4_CFG_CACHEINDEX);
+	qpace_write_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET, reg_val);
+
+	/* QMB2 register configurations. */
+	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET);
+	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_RD_OUTST_LIMIT);
+	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_WR_OUTST_LIMIT);
+	qpace_write_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET, reg_val);
+
+	/* DECOMP_CORE_CFG init steps. */
+	program_decomp_core_cfg();
+
+	/* Below settings help save power since all decomp cores are set to sync. */
+	reg_val = qpace_read_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET);
+	reg_val |= CORE_OPER_CFG_COMP_MEM_PWR_DWN_1;
+	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET, reg_val);
+
+	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET);
+	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_RD_MAX_OT);
+	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_WR_MAX_OT);
+	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET, reg_val);
+
+	/* Set all COMP engines to bulk mode. */
+	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET);
+	reg_val |= COMP_CORE_BULK_MODE_ALL_CORES;
+	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET, reg_val);
+
+	/* URG CMD register configurations. */
+	program_urg_command_contexts_v2();
+
+	return 0;
+}
+
+enum qpace_interrupts {
+	QPACE_IRQ_URGENT
+};
+
+static int qpace_register_interrupts(struct platform_device *pdev)
+{
+	struct device *dev = &pdev->dev;
+	int irq, ret;
+
+	irq = platform_get_irq(pdev, QPACE_IRQ_URGENT);
+	if (irq < 0)
+		return irq;
+
+	ret = devm_request_irq(dev, irq, urgent_interrupt_handler,
+			       0, "qpace-urgent-irq", NULL);
+	if (ret)
+		dev_err(dev, "failed to request urgent interrupt\n");
+
+	return ret;
+}
+
+static inline bool _qpace_power_on(void)
+{
+	u32 ready_status;
+
+	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_RUN);
+
+	if (readl_poll_timeout(qpace_priv->gen_core_regs +
+			       QPACE_CORE_OPER_CORE_READY_OFFSET,
+			       ready_status, ready_status,
+			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
+		pr_err("Timeout in waiting for QPaCE to turn on\n");
+		return false;
+	}
+
+	return true;
+}
+
+static int qpace_power_on(struct device *dev)
+{
+	int ret, ret2;
+
+	qpace_priv->interconnect = devm_of_icc_get(dev, "qpace-mem");
+	if (IS_ERR_OR_NULL(qpace_priv->interconnect)) {
+		ret = PTR_ERR_OR_ZERO(qpace_priv->interconnect);
+		pr_err("%s: devm_of_icc_get() failed with %d\n", __func__, ret);
+		return qpace_priv->interconnect ? ret : -EINVAL;
+	}
+
+	ret = device_init_wakeup(dev, true);
+	if (ret) {
+		pr_err("%s: device_init_wakeup() failed with %d\n", __func__, ret);
+		return ret;
+	}
+
+	cpu_latency_qos_add_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
+
+	icc_set_tag(qpace_priv->interconnect, QCOM_ICC_TAG_ACTIVE_ONLY);
+
+	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
+	if (ret) {
+		pr_err("Failed to turn on QPaCE VCD: %d\n", ret);
+		goto rm_qos;
+	}
+
+	if (!_qpace_power_on()) {
+		pr_err("Failed to start QPaCE\n");
+		ret = -EINVAL;
+		goto rm_bw;
+	}
+
+	return 0;
+
+rm_bw:
+	ret2 = icc_set_bw(qpace_priv->interconnect, 0, 0);
+	if (ret2)
+		pr_err("Failed to remove QPaCE VCD vote: %d\n", ret2);
+rm_qos:
+	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
+	device_init_wakeup(dev, false);
+
+	return ret;
+}
+
+static inline bool _qpace_power_off(void)
+{
+	u32 ready_status;
+
+	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_STOP);
+
+	if (readl_poll_timeout(qpace_priv->gen_core_regs +
+			       QPACE_CORE_OPER_CORE_READY_OFFSET,
+			       ready_status, !ready_status,
+			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
+		pr_err("Timeout in waiting for QPaCE to turn off\n");
+		return false;
+	}
+
+	return true;
+}
+
+static void qpace_power_off(struct device *dev)
+{
+	int ret;
+
+	/* If this fails we can still remove our vote for the VCD to turn QPaCE off */
+	if (!_qpace_power_off())
+		pr_err("Failed to stop QPaCE\n");
+
+	ret = icc_set_bw(qpace_priv->interconnect, 0, 0);
+	if (ret)
+		pr_err("Failed to turn off QPaCE VCD: %d\n", ret);
+
+	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
+
+	device_init_wakeup(dev, false);
+}
+
+static inline int qpace_register_ioremap(struct platform_device *pdev)
+{
+	qpace_priv->gen_regs = devm_platform_ioremap_resource(pdev, 0);
+	if (IS_ERR(qpace_priv->gen_regs))
+		return PTR_ERR(qpace_priv->gen_regs);
+
+	qpace_priv->gen_core_regs = qpace_priv->gen_regs + QPACE_GEN_CORE_REGS_OFFSET;
+	qpace_priv->comp_core_regs = qpace_priv->gen_regs + QPACE_COMP_CORE_REGS_OFFSET;
+	qpace_priv->decomp_core_regs = qpace_priv->gen_regs + QPACE_DECOMP_CORE_REGS_OFFSET;
+	qpace_priv->urg_regs = qpace_priv->gen_regs + QPACE_URG_REGS_OFFSET;
+
+	return 0;
+}
+
+bool qpace_is_dev_available(void)
+{
+	return static_branch_likely(&qpace_drv_probed) &&
+	       !READ_ONCE(qpace_priv->broken);
+}
+EXPORT_SYMBOL_GPL(qpace_is_dev_available);
+
+struct device *qpace_get_dma_dev(void)
+{
+	return qpace_priv->dev;
+}
+EXPORT_SYMBOL_GPL(qpace_get_dma_dev);
+
+static int qpace_probe(struct platform_device *pdev)
+{
+	struct device *dev = &pdev->dev;
+	struct qpace_priv *priv;
+	int ret;
+
+	priv = devm_kzalloc(dev, sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	priv->dev = dev;
+	/* Starts already complete since active_rings == 0 at init. */
+	init_completion(&priv->no_active_refs);
+	complete(&priv->no_active_refs);
+	INIT_WORK(&priv->disable_work, qpace_disable_work_fn);
+	qpace_priv = priv;
+	platform_set_drvdata(pdev, priv);
+
+	ret = dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
+	if (ret)
+		return dev_err_probe(dev, ret, "failed to set DMA mask\n");
+
+	ret = qpace_register_ioremap(pdev);
+	if (ret)
+		return dev_err_probe(dev, ret, "failed to map QPaCE registers\n");
+
+	ret = qpace_power_on(dev);
+	if (ret)
+		return dev_err_probe(dev, ret, "failed to power on QPaCE\n");
+
+	/* Get QPaCE HW version. */
+	qpace_priv->hw_version = qpace_read_gen(qpace_priv, QPACE_CORE_HW_VERSION_OFFSET);
+	if (qpace_priv->hw_version != QPACE_HW_VERSION_V2) {
+		dev_err(dev, "Unsupported QPaCE HW version returned: 0x%x\n",
+			qpace_priv->hw_version);
+		ret = -EINVAL;
+		goto power_off;
+	}
+
+	ret = qpace_hw_init();
+	if (ret) {
+		dev_err(dev, "init failed: (%d)\n", ret);
+		goto power_off;
+	}
+
+	ret = qpace_register_interrupts(pdev);
+	if (ret) {
+		dev_err(dev, "failed to register interrupts\n");
+		goto power_off;
+	}
+
+	static_branch_enable(&qpace_drv_probed);
+
+	return ret;
+
+power_off:
+	qpace_power_off(dev);
+
+	return ret;
+}
+
+static void qpace_remove(struct platform_device *pdev)
+{
+	mutex_lock(&qpace_ref_lock);
+	qpace_priv->suspended = true;
+	mutex_unlock(&qpace_ref_lock);
+
+	static_branch_disable(&qpace_drv_probed);
+
+	wait_for_completion(&qpace_priv->no_active_refs);
+
+	/* No callers remain; tear down the hardware. */
+	cancel_work_sync(&qpace_priv->disable_work);
+	qpace_power_off(&pdev->dev);
+}
+
+static const struct of_device_id qpace_match_table[] = {
+	{ .compatible = "qcom,hawi-qpace" },
+	{ }
+};
+MODULE_DEVICE_TABLE(of, qpace_match_table);
+
+static int qpace_suspend(struct device *dev)
+{
+	mutex_lock(&qpace_ref_lock);
+
+	qpace_priv->suspended = true;
+
+	if (qpace_priv->active_rings) {
+		dev_err(dev, "active_rings not 0, abort suspend\n");
+		qpace_priv->suspended = false;
+		mutex_unlock(&qpace_ref_lock);
+		return -EBUSY;
+	}
+
+	mutex_unlock(&qpace_ref_lock);
+	return 0;
+}
+
+static int qpace_resume(struct device *dev)
+{
+	int ret;
+
+	mutex_lock(&qpace_ref_lock);
+
+	if (qpace_priv->active_rings)
+		dev_err(dev, "active_rings not 0, unexpected case\n");
+
+	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
+	if (ret)
+		goto out_unlock;
+
+	program_urg_command_contexts_v2();
+	program_decomp_core_cfg();
+
+	qpace_priv->suspended = false;
+
+out_unlock:
+	if (ret)
+		dev_err(dev, "failed to resume QPaCE: %d\n", ret);
+	mutex_unlock(&qpace_ref_lock);
+	return ret;
+}
+
+static DEFINE_SIMPLE_DEV_PM_OPS(qpace_pm_ops, qpace_suspend, qpace_resume);
+
+static struct platform_driver qpace_driver = {
+	.probe = qpace_probe,
+	.remove = qpace_remove,
+	.driver = {
+		.name = "qcom-qpace",
+		.of_match_table = qpace_match_table,
+		.pm = pm_sleep_ptr(&qpace_pm_ops),
+	},
+};
+
+module_platform_driver(qpace_driver);
+
+MODULE_DESCRIPTION("Qualcomm Page Compression Engine driver");
+MODULE_LICENSE("GPL");
diff --git a/drivers/soc/qcom/qpace_internal.h b/drivers/soc/qcom/qpace_internal.h
new file mode 100644
index 000000000000..7e18d442e32f
--- /dev/null
+++ b/drivers/soc/qcom/qpace_internal.h
@@ -0,0 +1,84 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#ifndef __QPACE_INTERNAL_H__
+#define __QPACE_INTERNAL_H__
+
+#include <linux/bits.h>
+#include <linux/bitfield.h>
+
+/* hardware definitions and opcode values */
+#define NUM_TRS_ERS_URG_CMD_REGS 8
+#define QPACE_HW_VERSION_V2 0x2000000
+#define DEFAULT_SMMU_CONTEXT 0
+
+#define LZ4_COMP 0x9
+#define LZ4_DECOMP 0xa
+
+#define OP_OK 0x0
+#define OP_BUS_ERROR 0x1
+#define OP_COMP_TOO_BIG 0x2
+#define OP_TIMED_OUT 0x5
+#define OP_URG_ONGOING 0xf
+
+#define LLC_ALLOC 0x1
+
+#define URG_CMD_CONTEXT_SPACING 0x10
+
+#define QPACE_RUN 0x1
+#define QPACE_STOP 0x0
+#define QPACE_STATE_CHANGE_TIMEOUT_US 10000
+
+/* register offsets */
+#define QPACE_GEN_CORE_REGS_OFFSET 0x1000
+#define QPACE_COMP_CORE_REGS_OFFSET 0x2000
+#define QPACE_DECOMP_CORE_REGS_OFFSET 0x6000
+#define QPACE_URG_REGS_OFFSET 0x10000
+
+#define QPACE_CORE_HW_VERSION_OFFSET 0x0
+#define QPACE_CORE_QNS4_CFG_OFFSET 0x14
+#define CORE_QNS4_CFG_CACHEINDEX GENMASK(17, 12)
+
+#define QPACE_CORE_GEN_CFG_OFFSET 0x18
+#define CORE_GEN_CFG_QMB2_MAX_RD_OUTST_LIMIT GENMASK(13, 7)
+#define CORE_GEN_CFG_QMB2_MAX_WR_OUTST_LIMIT GENMASK(6, 0)
+
+#define QPACE_CORE_URG_CMD_n_CACHEINDEX_CFG_OFFSET 0x2fc
+#define CORE_URG_CMD_n_CACHEINDEX__ENGINE_CACHEINDEX_OVERRIDE BIT(6)
+#define CORE_URG_CMD_n_CACHEINDEX__ENGINE_CACHEINDEX GENMASK(5, 0)
+
+#define QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET 0x0
+#define QPACE_CORE_OPER_CORE_READY_OFFSET 0x4
+#define QPACE_CORE_OPER_CFG_OFFSET 0x8
+#define CORE_OPER_CFG_COMP_MEM_PWR_DWN_1 BIT(3)
+
+#define QPACE_COMP_CORE_CFG_OFFSET 0x0
+#define COMP_CORE_CFG_DMA_WR_MAX_OT GENMASK(17, 13)
+#define COMP_CORE_CFG_DMA_RD_MAX_OT GENMASK(12, 8)
+#define QPACE_COMP_CORE_BULK_MODE_OFFSET 0x4
+#define COMP_CORE_BULK_MODE_ALL_CORES GENMASK(5, 0)
+
+#define QPACE_DECOMP_CORE_CFG_OFFSET 0x0
+#define DECOMP_CORE_CFG_DECOMP_DMA_MEM_MNG_PAGE_SIZE GENMASK(12, 10)
+#define DECOMP_CORE_CFG_DMA_WR_MAX_OT GENMASK(9, 5)
+#define DECOMP_CORE_CFG_DMA_RD_MAX_OT GENMASK(4, 0)
+
+#define QPACE_URG_CMD_0_TD_DST_ADDR_L_CFG_CNTXT_OFFSET 0x0
+#define URG_CMD_0_TD_DST_ADDR_L__CMD_CFG_CNTXT GENMASK(3, 0)
+#define QPACE_URG_CMD_0_ED_STAT_OFFSET 0x18
+#define URG_CMD_0_ED_STAT_COMP_CODE GENMASK(23, 20)
+#define URG_CMD_0_ED_STAT_SIZE GENMASK(19, 0)
+#define QPACE_URG_CMD_0_STAT_CLR_OFFSET 0x28
+#define URG_CMD_0_STAT_CLR_REQ_ON_ACTIVE_ERR BIT(1)
+#define QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET 0x40
+#define URG_CMD_0_CFG_CNTXT_SIZE_SIZE GENMASK(19, 0)
+#define QPACE_URG_CMD_0_CFG_CNTXT_MISC_n_OFFSET 0x48
+#define URG_CMD_0_CFG_CNTXT_MISC_WA BIT(21)
+#define URG_CMD_0_CFG_CNTXT_MISC_DCTX GENMASK(17, 16)
+#define URG_CMD_0_CFG_CNTXT_MISC_SCTX GENMASK(15, 14)
+#define URG_CMD_0_CFG_CNTXT_MISC_PAGE_CNT GENMASK(13, 8)
+#define URG_CMD_0_CFG_CNTXT_MISC_OPER GENMASK(7, 4)
+
+#endif /* __QPACE_INTERNAL_H__ */
diff --git a/include/linux/soc/qcom/qpace.h b/include/linux/soc/qcom/qpace.h
new file mode 100644
index 000000000000..7ef7d12c4e66
--- /dev/null
+++ b/include/linux/soc/qcom/qpace.h
@@ -0,0 +1,154 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#ifndef __LINUX_SOC_QCOM_QPACE_H
+#define __LINUX_SOC_QCOM_QPACE_H
+
+#include <linux/device.h>
+#include <linux/dma-mapping.h>
+#include <linux/errno.h>
+#include <linux/kconfig.h>
+#include <linux/types.h>
+
+/**
+ * struct qpace_algorithm - Descriptor for a QPaCE compression algorithm.
+ * @name:            Algorithm name string (e.g. "qpace-lz4").
+ * @comp_opcode:     Hardware opcode for compression.
+ * @decomp_opcode:   Hardware opcode for decompression.
+ * @unused:          Reserved bits.
+ * @urg_comp_cntxt:  Urgent-path context index for compression.
+ * @urg_decomp_cntxt: Urgent-path context index for decompression.
+ */
+struct qpace_algorithm {
+	char *name;
+	u32 comp_opcode:4;
+	u32 decomp_opcode:4;
+	u32 unused:24;
+	int urg_comp_cntxt;
+	int urg_decomp_cntxt;
+};
+
+#if IS_ENABLED(CONFIG_QCOM_PAGE_COMPRESSION_ENGINE)
+
+/**
+ * qpace_is_valid_algorithm() - Check if a name identifies a supported QPaCE algorithm.
+ * @algo_name: Algorithm name string to check.
+ *
+ * Returns true if @algo_name matches a QPaCE algorithm and the driver has
+ * probed successfully, false otherwise.
+ */
+bool qpace_is_valid_algorithm(const char *algo_name);
+
+/**
+ * qpace_lz4_algorithm - Algorithm descriptor for QPaCE LZ4.
+ *
+ * Pass to qpace_urgent_compress() and qpace_urgent_decompress() to select
+ * the LZ4 algorithm.
+ */
+extern const struct qpace_algorithm qpace_lz4_algorithm;
+
+/**
+ * qpace_is_dev_available() - Check whether the QPaCE hardware is ready.
+ *
+ * Returns true if the driver has probed successfully and the hardware is
+ * available for use, false otherwise.
+ */
+bool qpace_is_dev_available(void);
+
+/**
+ * qpace_urgent_compress() - Compress a page synchronously via the urgent path.
+ * @input_addr:  DMA address of the page to compress.
+ * @output_addr: DMA address to write the compressed output to.
+ * @algo:        Algorithm descriptor to use.
+ *
+ * Blocks the calling CPU until compression completes. May sleep; must not be
+ * called from atomic or IRQ-disabled context.
+ *
+ * Return: Compressed size in bytes on success, negative error code on failure.
+ */
+int qpace_urgent_compress(dma_addr_t input_addr,
+			  dma_addr_t output_addr,
+			  struct qpace_algorithm *algo);
+
+/**
+ * qpace_urgent_decompress() - Decompress a page synchronously via the urgent path.
+ * @input_addr:  DMA address of the compressed input.
+ * @output_addr: DMA address to write the decompressed output to.
+ * @input_size:  Size of the compressed input in bytes.
+ * @algo:        Algorithm descriptor to use.
+ *
+ * Blocks the calling CPU until decompression completes. May sleep; must not
+ * be called from atomic or IRQ-disabled context.
+ *
+ * Return: Decompressed size in bytes on success, negative error code on failure.
+ */
+int qpace_urgent_decompress(dma_addr_t input_addr,
+			    dma_addr_t output_addr,
+			    size_t input_size,
+			    struct qpace_algorithm *algo);
+
+/**
+ * qpace_get() - Acquire a reference to QPaCE, preventing runtime power collapse.
+ *
+ * Must be paired with qpace_put() on success. Callers that hold a reference
+ * will keep the hardware awake and the urgent command contexts programmed.
+ *
+ * Return: 0 on success, -EBUSY if the device is suspended.
+ */
+int qpace_get(void);
+
+/**
+ * qpace_put() - Release a reference acquired with qpace_get().
+ *
+ * When the last reference is dropped, QPaCE may enter a low-power state.
+ */
+void qpace_put(void);
+
+/**
+ * qpace_get_dma_dev() - Return the device to use for QPaCE DMA mappings.
+ *
+ * Return: Pointer to the QPaCE platform device, or NULL if unavailable.
+ */
+struct device *qpace_get_dma_dev(void);
+
+#else /* !CONFIG_QCOM_PAGE_COMPRESSION_ENGINE */
+
+static inline bool qpace_is_valid_algorithm(const char *algo_name)
+{
+	return false;
+}
+
+static const struct qpace_algorithm qpace_lz4_algorithm;
+
+static inline bool qpace_is_dev_available(void)
+{
+	return false;
+}
+
+static inline int qpace_urgent_compress(dma_addr_t input_addr,
+					dma_addr_t output_addr,
+					struct qpace_algorithm *algo)
+{
+	return -EINVAL;
+}
+
+static inline int qpace_urgent_decompress(dma_addr_t input_addr,
+					  dma_addr_t output_addr,
+					  size_t input_size,
+					  struct qpace_algorithm *algo)
+{
+	return -EINVAL;
+}
+
+static inline int qpace_get(void) { return -ENODEV; }
+static inline void qpace_put(void) {}
+
+static inline struct device *qpace_get_dma_dev(void)
+{
+	return NULL;
+}
+
+#endif /* CONFIG_QCOM_PAGE_COMPRESSION_ENGINE */
+#endif /* __LINUX_SOC_QCOM_QPACE_H */

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
  2026-09-30 14:52 ` [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding Georgi Djakov
  2026-09-30 14:52 ` [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-09-30 15:00   ` sashiko-bot
  2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
                   ` (2 subsequent siblings)
  5 siblings, 1 reply; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

Add ftrace tracepoints for the Qualcomm Page Compression Engine
to allow performance analysis and debugging of QPaCE hardware
operations.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 drivers/soc/qcom/qpace.c     | 12 +++++
 include/trace/events/qpace.h | 99 ++++++++++++++++++++++++++++++++++++
 2 files changed, 111 insertions(+)
 create mode 100644 include/trace/events/qpace.h

diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
index 69aa14b7c477..f5b21ffbcbff 100644
--- a/drivers/soc/qcom/qpace.c
+++ b/drivers/soc/qcom/qpace.c
@@ -23,6 +23,9 @@
 
 #include "qpace_internal.h"
 
+#define CREATE_TRACE_POINTS
+#include <trace/events/qpace.h>
+
 #define QPACE_REG_PAGE_SIZE 4096
 #define QPACE_CORE_CACHEINDEX_SIZE 4
 
@@ -310,6 +313,8 @@ int qpace_urgent_compress(dma_addr_t input_addr,
 	if (ret)
 		return ret;
 
+	trace_start_qpace_urgent_compress((u64)input_addr, (u64)output_addr);
+
 	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
 	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
 						algo->urg_comp_cntxt);
@@ -338,6 +343,7 @@ int qpace_urgent_compress(dma_addr_t input_addr,
 
 	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
 out:
+	trace_end_qpace_urgent_compress((u64)input_addr, (u64)output_addr, ret);
 	return ret;
 }
 EXPORT_SYMBOL_GPL(qpace_urgent_compress);
@@ -356,6 +362,9 @@ int qpace_urgent_decompress(dma_addr_t input_addr,
 	if (ret)
 		goto out;
 
+	trace_start_qpace_urgent_decompress((u64)input_addr,
+					    (u64)output_addr, input_size);
+
 	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
 	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
 				urg_reg_num, algo->urg_decomp_cntxt,
@@ -381,6 +390,9 @@ int qpace_urgent_decompress(dma_addr_t input_addr,
 
 	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
 out:
+	trace_end_qpace_urgent_decompress((u64)input_addr,
+					  (u64)output_addr,
+					  input_size, ret);
 	return ret;
 }
 EXPORT_SYMBOL_GPL(qpace_urgent_decompress);
diff --git a/include/trace/events/qpace.h b/include/trace/events/qpace.h
new file mode 100644
index 000000000000..3f0ccdb8f206
--- /dev/null
+++ b/include/trace/events/qpace.h
@@ -0,0 +1,99 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#undef TRACE_SYSTEM
+#define TRACE_SYSTEM qpace
+
+#if !defined(_TRACE_QPACE_H) || defined(TRACE_HEADER_MULTI_READ)
+#define _TRACE_QPACE_H
+#include <linux/tracepoint.h>
+
+TRACE_EVENT(start_qpace_urgent_decompress,
+
+	TP_PROTO(u64 src, u64 dst, unsigned int size),
+	TP_ARGS(src, dst, size),
+
+	 TP_STRUCT__entry(
+		 __field(u64, src)
+		 __field(u64, dst)
+		 __field(unsigned int, size)
+	 ),
+
+	 TP_fast_assign(
+		 __entry->src = src;
+		 __entry->dst = dst;
+		 __entry->size = size;
+	 ),
+
+	 TP_printk("src=0x%llx dst=0x%llx size=%u",
+		   __entry->src, __entry->dst, __entry->size)
+);
+
+TRACE_EVENT(end_qpace_urgent_decompress,
+
+	 TP_PROTO(u64 src, u64 dst, unsigned int size, int ret),
+	 TP_ARGS(src, dst, size, ret),
+
+	 TP_STRUCT__entry(
+		 __field(u64, src)
+		 __field(u64, dst)
+		 __field(unsigned int, size)
+		 __field(int, ret)
+	 ),
+
+	 TP_fast_assign(
+		 __entry->src = src;
+		 __entry->dst = dst;
+		 __entry->size = size;
+		 __entry->ret = ret;
+	 ),
+
+	 TP_printk("src=0x%llx dst=0x%llx size=%u ret=%d",
+		   __entry->src, __entry->dst, __entry->size, __entry->ret)
+);
+
+TRACE_EVENT(start_qpace_urgent_compress,
+
+	 TP_PROTO(u64 src, u64 dst),
+	 TP_ARGS(src, dst),
+
+	 TP_STRUCT__entry(
+		 __field(u64, src)
+		 __field(u64, dst)
+	 ),
+
+	 TP_fast_assign(
+		 __entry->src = src;
+		 __entry->dst = dst;
+	 ),
+
+	 TP_printk("src=0x%llx dst=0x%llx",
+		   __entry->src, __entry->dst)
+);
+
+TRACE_EVENT(end_qpace_urgent_compress,
+
+	 TP_PROTO(u64 src, u64 dst, int ret),
+	 TP_ARGS(src, dst, ret),
+
+	 TP_STRUCT__entry(
+		 __field(u64, src)
+		 __field(u64, dst)
+		 __field(int, ret)
+	 ),
+
+	 TP_fast_assign(
+		 __entry->src = src;
+		 __entry->dst = dst;
+		 __entry->ret = ret;
+	 ),
+
+	 TP_printk("src=0x%llx dst=0x%llx ret=%d",
+		   __entry->src, __entry->dst, __entry->ret)
+);
+
+#endif /* _TRACE_QPACE_H */
+
+#include <trace/define_trace.h>

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
                   ` (2 preceding siblings ...)
  2026-09-30 14:52 ` [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-09-30 15:08   ` sashiko-bot
                     ` (2 more replies)
  2026-09-30 14:52 ` [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support Georgi Djakov
  2026-09-30 14:52 ` [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node Georgi Djakov
  5 siblings, 3 replies; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
zram can expose qpace-lz4 as a selectable compression algorithm when the
QPaCE driver is available.

Compress and decompress operations are handled via the QPaCE urgent
synchronous path: each request DMA-maps the source and destination
buffers, issues a blocking hardware command, and returns the result size.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 drivers/block/zram/Kconfig         |  11 +++
 drivers/block/zram/Makefile        |   1 +
 drivers/block/zram/backend_qpace.c | 134 +++++++++++++++++++++++++++++
 drivers/block/zram/backend_qpace.h |  13 +++
 drivers/block/zram/zcomp.c         |  11 ++-
 5 files changed, 168 insertions(+), 2 deletions(-)
 create mode 100644 drivers/block/zram/backend_qpace.c
 create mode 100644 drivers/block/zram/backend_qpace.h

diff --git a/drivers/block/zram/Kconfig b/drivers/block/zram/Kconfig
index 402b7b175863..eece67e7ea83 100644
--- a/drivers/block/zram/Kconfig
+++ b/drivers/block/zram/Kconfig
@@ -44,6 +44,17 @@ config ZRAM_BACKEND_842
 	select 842_COMPRESS
 	select 842_DECOMPRESS
 
+config ZRAM_BACKEND_QPACE
+	bool "QPaCE compression support"
+	depends on ZRAM
+	depends on QCOM_PAGE_COMPRESSION_ENGINE = y || \
+		(ZRAM = m && QCOM_PAGE_COMPRESSION_ENGINE = m)
+	help
+	  Enable Qualcomm Page Compression Engine (QPaCE) hardware
+	  acceleration as a zram backend. When selected, zram offloads
+	  synchronous compress and decompress operations to the QPaCE
+	  urgent command path, reducing CPU overhead.
+
 config ZRAM_BACKEND_FORCE_LZO
 	depends on ZRAM
 	def_bool !ZRAM_BACKEND_LZ4 && !ZRAM_BACKEND_LZ4HC && \
diff --git a/drivers/block/zram/Makefile b/drivers/block/zram/Makefile
index 0fdefd576691..9e1c35f942e0 100644
--- a/drivers/block/zram/Makefile
+++ b/drivers/block/zram/Makefile
@@ -8,5 +8,6 @@ zram-$(CONFIG_ZRAM_BACKEND_LZ4HC)	+= backend_lz4hc.o
 zram-$(CONFIG_ZRAM_BACKEND_ZSTD)	+= backend_zstd.o
 zram-$(CONFIG_ZRAM_BACKEND_DEFLATE)	+= backend_deflate.o
 zram-$(CONFIG_ZRAM_BACKEND_842)		+= backend_842.o
+zram-$(CONFIG_ZRAM_BACKEND_QPACE)	+= backend_qpace.o
 
 obj-$(CONFIG_ZRAM)	+=	zram.o
diff --git a/drivers/block/zram/backend_qpace.c b/drivers/block/zram/backend_qpace.c
new file mode 100644
index 000000000000..3af13727adef
--- /dev/null
+++ b/drivers/block/zram/backend_qpace.c
@@ -0,0 +1,134 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#include <linux/dma-mapping.h>
+#include <linux/slab.h>
+#include <linux/soc/qcom/qpace.h>
+
+#include "backend_qpace.h"
+
+struct qpace_ctx {
+	struct device *dev;
+	void *in_buf;
+	dma_addr_t in_dma;
+	void *out_buf;
+	dma_addr_t out_dma;
+};
+
+static int qpace_setup_lz4_params(struct zcomp_params *params)
+{
+	params->drv_data = (void *)&qpace_lz4_algorithm;
+	return 0;
+}
+
+static void qpace_release_params(struct zcomp_params *params)
+{
+}
+
+static int qpace_create_ctx(struct zcomp_params *params, struct zcomp_ctx *ctx)
+{
+	struct device *dev = qpace_get_dma_dev();
+	struct qpace_ctx *cc;
+
+	if (!dev)
+		return -ENODEV;
+
+	cc = kzalloc_obj(*cc, GFP_KERNEL);
+	if (!cc)
+		return -ENOMEM;
+
+	cc->dev = dev;
+	cc->in_buf = dma_alloc_coherent(dev, PAGE_SIZE, &cc->in_dma, GFP_KERNEL);
+	if (!cc->in_buf)
+		goto err_free;
+
+	cc->out_buf = dma_alloc_coherent(dev, PAGE_SIZE, &cc->out_dma, GFP_KERNEL);
+	if (!cc->out_buf)
+		goto err_free_in;
+
+	ctx->context = cc;
+	return 0;
+
+err_free_in:
+	dma_free_coherent(dev, PAGE_SIZE, cc->in_buf, cc->in_dma);
+err_free:
+	kfree(cc);
+	return -ENOMEM;
+}
+
+static void qpace_destroy_ctx(struct zcomp_ctx *ctx)
+{
+	struct qpace_ctx *cc = ctx->context;
+
+	if (!cc)
+		return;
+
+	dma_free_coherent(cc->dev, PAGE_SIZE, cc->out_buf, cc->out_dma);
+	dma_free_coherent(cc->dev, PAGE_SIZE, cc->in_buf, cc->in_dma);
+	kfree(cc);
+	ctx->context = NULL;
+}
+
+static int qpace_compress(struct zcomp_params *params, struct zcomp_ctx *ctx,
+			  struct zcomp_req *req)
+{
+	struct qpace_algorithm *algo = params->drv_data;
+	struct qpace_ctx *cc = ctx->context;
+	int ret;
+
+	memcpy(cc->in_buf, req->src, req->src_len);
+
+	ret = qpace_urgent_compress(cc->in_dma, cc->out_dma, algo);
+
+	if (ret == -E2BIG || ret >= PAGE_SIZE) {
+		req->dst_len = PAGE_SIZE;
+		return 0;
+	}
+	if (ret <= 0)
+		return ret < 0 ? ret : -EIO;
+
+	if (ret > req->dst_len)
+		return -EIO;
+
+	memcpy(req->dst, cc->out_buf, ret);
+	req->dst_len = ret;
+
+	return 0;
+}
+
+static int qpace_decompress(struct zcomp_params *params, struct zcomp_ctx *ctx,
+			    struct zcomp_req *req)
+{
+	struct qpace_algorithm *algo = params->drv_data;
+	struct qpace_ctx *cc = ctx->context;
+	int ret;
+
+	if (req->src_len > PAGE_SIZE)
+		return -EINVAL;
+
+	memcpy(cc->in_buf, req->src, req->src_len);
+
+	ret = qpace_urgent_decompress(cc->in_dma, cc->out_dma, req->src_len, algo);
+	if (ret <= 0)
+		return ret < 0 ? ret : -EIO;
+
+	if (ret > req->dst_len)
+		return -EIO;
+
+	memcpy(req->dst, cc->out_buf, ret);
+	req->dst_len = ret;
+
+	return 0;
+}
+
+const struct zcomp_ops backend_qpace_lz4 = {
+	.name			= "qpace-lz4",
+	.setup_params		= qpace_setup_lz4_params,
+	.release_params		= qpace_release_params,
+	.create_ctx		= qpace_create_ctx,
+	.destroy_ctx		= qpace_destroy_ctx,
+	.compress		= qpace_compress,
+	.decompress		= qpace_decompress,
+};
diff --git a/drivers/block/zram/backend_qpace.h b/drivers/block/zram/backend_qpace.h
new file mode 100644
index 000000000000..03747405edcf
--- /dev/null
+++ b/drivers/block/zram/backend_qpace.h
@@ -0,0 +1,13 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
+ */
+
+#ifndef __BACKEND_QPACE_H__
+#define __BACKEND_QPACE_H__
+
+#include "zcomp.h"
+
+extern const struct zcomp_ops backend_qpace_lz4;
+
+#endif /* __BACKEND_QPACE_H__ */
diff --git a/drivers/block/zram/zcomp.c b/drivers/block/zram/zcomp.c
index 028e2f0f587e..b5bd0b11a975 100644
--- a/drivers/block/zram/zcomp.c
+++ b/drivers/block/zram/zcomp.c
@@ -22,6 +22,10 @@
 #include "backend_deflate.h"
 #include "backend_842.h"
 
+#if IS_ENABLED(CONFIG_ZRAM_BACKEND_QPACE)
+#include "backend_qpace.h"
+#endif
+
 static const struct zcomp_ops *backends[] = {
 #if IS_ENABLED(CONFIG_ZRAM_BACKEND_LZO)
 	&backend_lzorle,
@@ -41,6 +45,9 @@ static const struct zcomp_ops *backends[] = {
 #endif
 #if IS_ENABLED(CONFIG_ZRAM_BACKEND_842)
 	&backend_842,
+#endif
+#if IS_ENABLED(CONFIG_ZRAM_BACKEND_QPACE)
+	&backend_qpace_lz4,
 #endif
 	NULL
 };
@@ -80,10 +87,10 @@ static const struct zcomp_ops *lookup_backend_ops(const char *comp)
 
 	while (backends[i]) {
 		if (sysfs_streq(comp, backends[i]->name))
-			break;
+			return backends[i];
 		i++;
 	}
-	return backends[i];
+	return NULL;
 }
 
 const char *zcomp_lookup_backend_name(const char *comp)

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
                   ` (3 preceding siblings ...)
  2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-09-30 15:08   ` sashiko-bot
  2026-09-30 14:52 ` [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node Georgi Djakov
  5 siblings, 1 reply; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

QPaCE can use dedicated LLCC slices for compression and decompression, but
the driver currently does not request or enable them. This leaves platforms
with QPaCE LLCC support running without the intended cache allocation.

Add the QPaCE LLCC slice IDs and make the driver request and activate the
slices.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 drivers/soc/qcom/qpace.c           | 58 ++++++++++++++++++++++++++++--
 include/linux/soc/qcom/llcc-qcom.h |  2 ++
 2 files changed, 58 insertions(+), 2 deletions(-)

diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
index f5b21ffbcbff..86a935ae2ca9 100644
--- a/drivers/soc/qcom/qpace.c
+++ b/drivers/soc/qcom/qpace.c
@@ -18,6 +18,7 @@
 #include <linux/of_platform.h>
 #include <linux/platform_device.h>
 #include <linux/pm_qos.h>
+#include <linux/soc/qcom/llcc-qcom.h>
 #include <linux/soc/qcom/qpace.h>
 #include <linux/workqueue.h>
 
@@ -42,6 +43,8 @@ struct qpace_priv {
 
 	struct icc_path *interconnect;
 	struct pm_qos_request qos_req;
+	struct llcc_slice_desc *llc_comp;
+	struct llcc_slice_desc *llc_decomp;
 
 	int active_rings;
 	struct completion no_active_refs;
@@ -660,9 +663,35 @@ static int qpace_probe(struct platform_device *pdev)
 	if (ret)
 		return dev_err_probe(dev, ret, "failed to map QPaCE registers\n");
 
+	priv->llc_comp = llcc_slice_getd(LLCC_QPACE_COMPRESSION);
+	if (IS_ERR(priv->llc_comp))
+		return dev_err_probe(dev, PTR_ERR(priv->llc_comp),
+				     "failed to get compression LLCC slice\n");
+
+	priv->llc_decomp = llcc_slice_getd(LLCC_QPACE_DECOMPRESSION);
+	if (IS_ERR(priv->llc_decomp)) {
+		ret = dev_err_probe(dev, PTR_ERR(priv->llc_decomp),
+				    "failed to get decompression LLCC slice\n");
+		priv->llc_decomp = NULL;
+		goto llc_put_comp;
+	}
+
+	ret = llcc_slice_activate(priv->llc_comp);
+	if (ret) {
+		dev_err_probe(dev, ret, "failed to activate compression LLCC slice\n");
+		goto llc_put_decomp;
+	}
+
+	ret = llcc_slice_activate(priv->llc_decomp);
+	if (ret) {
+		dev_err_probe(dev, ret, "failed to activate decompression LLCC slice\n");
+		goto llc_deactivate_comp;
+	}
 	ret = qpace_power_on(dev);
-	if (ret)
-		return dev_err_probe(dev, ret, "failed to power on QPaCE\n");
+	if (ret) {
+		dev_err(dev, "failed to power on QPaCE\n");
+		goto llc_deactivate_decomp;
+	}
 
 	/* Get QPaCE HW version. */
 	qpace_priv->hw_version = qpace_read_gen(qpace_priv, QPACE_CORE_HW_VERSION_OFFSET);
@@ -691,6 +720,14 @@ static int qpace_probe(struct platform_device *pdev)
 
 power_off:
 	qpace_power_off(dev);
+llc_deactivate_decomp:
+	llcc_slice_deactivate(priv->llc_decomp);
+llc_deactivate_comp:
+	llcc_slice_deactivate(priv->llc_comp);
+llc_put_decomp:
+	llcc_slice_putd(priv->llc_decomp);
+llc_put_comp:
+	llcc_slice_putd(priv->llc_comp);
 
 	return ret;
 }
@@ -708,6 +745,10 @@ static void qpace_remove(struct platform_device *pdev)
 	/* No callers remain; tear down the hardware. */
 	cancel_work_sync(&qpace_priv->disable_work);
 	qpace_power_off(&pdev->dev);
+	llcc_slice_deactivate(qpace_priv->llc_decomp);
+	llcc_slice_deactivate(qpace_priv->llc_comp);
+	llcc_slice_putd(qpace_priv->llc_decomp);
+	llcc_slice_putd(qpace_priv->llc_comp);
 }
 
 static const struct of_device_id qpace_match_table[] = {
@@ -729,6 +770,9 @@ static int qpace_suspend(struct device *dev)
 		return -EBUSY;
 	}
 
+	llcc_slice_deactivate(qpace_priv->llc_comp);
+	llcc_slice_deactivate(qpace_priv->llc_decomp);
+
 	mutex_unlock(&qpace_ref_lock);
 	return 0;
 }
@@ -746,6 +790,16 @@ static int qpace_resume(struct device *dev)
 	if (ret)
 		goto out_unlock;
 
+	ret = llcc_slice_activate(qpace_priv->llc_comp);
+	if (ret)
+		goto out_unlock;
+
+	ret = llcc_slice_activate(qpace_priv->llc_decomp);
+	if (ret) {
+		llcc_slice_deactivate(qpace_priv->llc_comp);
+		goto out_unlock;
+	}
+
 	program_urg_command_contexts_v2();
 	program_decomp_core_cfg();
 
diff --git a/include/linux/soc/qcom/llcc-qcom.h b/include/linux/soc/qcom/llcc-qcom.h
index 713cb0221b14..feddbebf1183 100644
--- a/include/linux/soc/qcom/llcc-qcom.h
+++ b/include/linux/soc/qcom/llcc-qcom.h
@@ -85,6 +85,8 @@
 #define LLCC_CAM_IPE_STROV	 92
 #define LLCC_CAM_OFE_STROV	 93
 #define LLCC_CPUSS_HEU	 94
+#define LLCC_QPACE_COMPRESSION 95
+#define LLCC_QPACE_DECOMPRESSION 96
 #define LLCC_PCIE_TCU	 97
 #define LLCC_GPUHTW_LITTLE	 98
 #define LLCC_MDM_PNG_FIXED	 100

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node
  2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
                   ` (4 preceding siblings ...)
  2026-09-30 14:52 ` [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support Georgi Djakov
@ 2026-09-30 14:52 ` Georgi Djakov
  2026-09-30 14:59   ` sashiko-bot
  5 siblings, 1 reply; 25+ messages in thread
From: Georgi Djakov @ 2026-09-30 14:52 UTC (permalink / raw)
  To: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky
  Cc: axboe, rostedt, mhiramat, mathieu.desnoyers, linux-arm-msm,
	devicetree, linux-kernel, linux-block, linux-trace-kernel, djakov

Add the device-tree node for the Qualcomm Page Compression Engine on
the Hawi platform.

Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
---
 arch/arm64/boot/dts/qcom/hawi.dtsi | 17 +++++++++++++++++
 1 file changed, 17 insertions(+)

diff --git a/arch/arm64/boot/dts/qcom/hawi.dtsi b/arch/arm64/boot/dts/qcom/hawi.dtsi
index 6b0a0ff0faa4..283a5acfa073 100644
--- a/arch/arm64/boot/dts/qcom/hawi.dtsi
+++ b/arch/arm64/boot/dts/qcom/hawi.dtsi
@@ -5677,6 +5677,23 @@ llclpi_noc: interconnect@31310000 {
 			qcom,bcm-voters = <&apps_bcm_voter>;
 		};
 
+		qpace: compression@31400000 {
+			compatible = "qcom,hawi-qpace";
+			reg = <0x0 0x31400000 0x0 0x80000>;
+			iommus = <&apps_smmu 0x1c20 0x13>,
+				 <&apps_smmu 0x1c24 0x13>,
+				 <&apps_smmu 0x1c28 0x13>,
+				 <&apps_smmu 0x1c2c 0x13>;
+			dma-coherent;
+			interconnects = <&gem_noc MASTER_QPACE QCOM_ICC_TAG_ALWAYS
+					 &mc_virt SLAVE_EBI1 QCOM_ICC_TAG_ALWAYS>;
+			interconnect-names = "qpace-mem";
+			interrupts = <GIC_ESPI 72 IRQ_TYPE_LEVEL_HIGH>,
+				     <GIC_ESPI 71 IRQ_TYPE_LEVEL_HIGH>;
+			interrupt-names = "ring-and-bus-err",
+					  "urgent";
+		};
+
 		system-cache-controller@31800000 {
 			compatible = "qcom,hawi-llcc";
 			reg = <0x0 0x31800000 0x0 0x200000>,

^ permalink raw reply related	[flat|nested] 25+ messages in thread

* Re: [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node
  2026-09-30 14:52 ` [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node Georgi Djakov
@ 2026-09-30 14:59   ` sashiko-bot
  0 siblings, 0 replies; 25+ messages in thread
From: sashiko-bot @ 2026-09-30 14:59 UTC (permalink / raw)
  To: Georgi Djakov; +Cc: devicetree, linux-trace-kernel, conor+dt, robh

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] arm64: dts: qcom: qpace interrupt index mismatch
--

commit d697c8786fb69e9e9786a688577373ec76f73ba9
Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>

arm64: dts: qcom: hawi: Add QPaCE DT node

Add the device-tree node for the Qualcomm Page Compression Engine on
the Hawi platform.

> diff --git a/arch/arm64/boot/dts/qcom/hawi.dtsi b/arch/arm64/boot/dts/qcom/hawi.dtsi
> index 6b0a0ff0faa45..283a5acfa0738 100644
> --- a/arch/arm64/boot/dts/qcom/hawi.dtsi
> +++ b/arch/arm64/boot/dts/qcom/hawi.dtsi
> @@ -5677,6 +5677,23 @@ llclpi_noc: interconnect@31310000 {
>  			qcom,bcm-voters = <&apps_bcm_voter>;
>  		};
>  
> +		qpace: compression@31400000 {
> +			compatible = "qcom,hawi-qpace";
> +			reg = <0x0 0x31400000 0x0 0x80000>;
> +			iommus = <&apps_smmu 0x1c20 0x13>,
> +				 <&apps_smmu 0x1c24 0x13>,
> +				 <&apps_smmu 0x1c28 0x13>,
> +				 <&apps_smmu 0x1c2c 0x13>;
> +			dma-coherent;
> +			interconnects = <&gem_noc MASTER_QPACE QCOM_ICC_TAG_ALWAYS
> +					 &mc_virt SLAVE_EBI1 QCOM_ICC_TAG_ALWAYS>;
> +			interconnect-names = "qpace-mem";
> +			interrupts = <GIC_ESPI 72 IRQ_TYPE_LEVEL_HIGH>,
> +				     <GIC_ESPI 71 IRQ_TYPE_LEVEL_HIGH>;
> +			interrupt-names = "ring-and-bus-err",
> +					  "urgent";

[Severity: Medium]
Will this interrupt mapping cause an IRQ storm?

Looking at drivers/soc/qcom/qpace.c in qpace_register_interrupts(), the
driver fetches the urgent interrupt by index 0 (QPACE_IRQ_URGENT = 0):

enum qpace_interrupts {
	QPACE_IRQ_URGENT
};

static int qpace_register_interrupts(struct platform_device *pdev)
{
	...
	irq = platform_get_irq(pdev, QPACE_IRQ_URGENT);

However, this device tree node places the urgent interrupt at index 1.

This would cause the driver to bind its urgent interrupt handler to the
ring-and-bus-err interrupt instead.

When an urgent completion IRQ fires (ESPI 71), it will be unhandled,
eventually causing an IRQ storm that the kernel must forcibly disable.
Meanwhile, error interrupts (ESPI 72) will incorrectly trigger the urgent
handler.

> +		};
> +
>  		system-cache-controller@31800000 {
>  			compatible = "qcom,hawi-llcc";
>  			reg = <0x0 0x31800000 0x0 0x200000>,

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930145215.3380977-1-georgi.djakov@oss.qualcomm.com?part=6

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations
  2026-09-30 14:52 ` [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations Georgi Djakov
@ 2026-09-30 15:00   ` sashiko-bot
  0 siblings, 0 replies; 25+ messages in thread
From: sashiko-bot @ 2026-09-30 15:00 UTC (permalink / raw)
  To: Georgi Djakov; +Cc: devicetree, linux-trace-kernel, robh, conor+dt

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] trace: unpaired end tracepoint in qpace_urgent_decompress() on error
- [Low] trace: qpace tracepoints violate subsystem_event naming convention
--

commit 95671997f765567ada636d922558170e766019a7
Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>

trace: qpace: Add tracepoints for QPaCE operations

Add ftrace tracepoints for the Qualcomm Page Compression Engine
to allow performance analysis and debugging of QPaCE hardware
operations.

> diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
> index 69aa14b7c4772..f5b21ffbcbff5 100644
> --- a/drivers/soc/qcom/qpace.c
> +++ b/drivers/soc/qcom/qpace.c

[ ... ]

> @@ -356,6 +362,9 @@ int qpace_urgent_decompress(dma_addr_t input_addr,
>  	if (ret)
>  		goto out;
>  
> +	trace_start_qpace_urgent_decompress((u64)input_addr,
> +					    (u64)output_addr, input_size);
> +
>  	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;

[ ... ]

>  	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
>  out:
> +	trace_end_qpace_urgent_decompress((u64)input_addr,
> +					  (u64)output_addr,
> +					  input_size, ret);
>  	return ret;
>  }

[Severity: Medium]
If qpace_get() fails at the beginning of qpace_urgent_decompress(), the code
jumps to the 'out' label, bypassing trace_start_qpace_urgent_decompress()
entirely.

Does this result in an unpaired end tracepoint being emitted unconditionally?

> diff --git a/include/trace/events/qpace.h b/include/trace/events/qpace.h
> new file mode 100644
> index 0000000000000..3f0ccdb8f2060
> --- /dev/null
> +++ b/include/trace/events/qpace.h
> @@ -0,0 +1,99 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +/*
> + * Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.
> + */
> +
> +#undef TRACE_SYSTEM
> +#define TRACE_SYSTEM qpace
> +
> +#if !defined(_TRACE_QPACE_H) || defined(TRACE_HEADER_MULTI_READ)
> +#define _TRACE_QPACE_H
> +#include <linux/tracepoint.h>
> +
> +TRACE_EVENT(start_qpace_urgent_decompress,

[Severity: Low]
Does this tracepoint name violate the mandatory subsystem_event naming
convention by using an action prefix instead of the subsystem prefix?

Could these be renamed to use the subsystem prefix, such as
qpace_start_urgent_decompress?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930145215.3380977-1-georgi.djakov@oss.qualcomm.com?part=3

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-09-30 14:52 ` [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver Georgi Djakov
@ 2026-09-30 15:04   ` sashiko-bot
  2026-10-03  0:06     ` Oreoluwa Babatunde
  2026-10-01  8:50   ` Krzysztof Kozlowski
  1 sibling, 1 reply; 25+ messages in thread
From: sashiko-bot @ 2026-09-30 15:04 UTC (permalink / raw)
  To: Georgi Djakov; +Cc: robh, linux-trace-kernel, conor+dt, devicetree

Thank you for your contribution! Sashiko AI review found 8 potential issue(s) to consider:
- [High] soc: qcom: qpace: concurrent MMIO access from CPU aliasing on >8 cores
- [High] soc: qcom: qpace: UAF of global qpace_priv after driver unbind
- [High] soc: qcom: qpace: 100ms busy wait with preemption disabled causes latency spikes
- [High] soc: qcom: qpace: interrupt storm due to uncleared level-triggered IRQ
- [High] soc: qcom: qpace: stray DMA writes on urgent command timeout
- [High] soc: qcom: qpace: hardware stuck in STOP mode after suspend/resume
- [Medium] soc: qcom: qpace: silent truncation of unaligned output_addr causes DMA corruption
- [Medium] soc: qcom: qpace: inline stp assembly breaks big-endian architectures
--

commit 5edbc123f3ebd977977e7d874e4ecebc5ed95e6a
Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>

soc: qcom: qpace: Add Qualcomm Page Compression Engine driver

This commit adds a platform driver for the Qualcomm Page Compression Engine
(QPaCE). It provides the urgent command path for synchronous single-page
compression and decompression, supporting low-latency operations for
compressed-memory users like zram.

> diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
> index 0000000000000..69aa14b7c4772
> --- /dev/null
> +++ b/drivers/soc/qcom/qpace.c
[ ... ]
> +static inline int qpace_urgent_command_trigger(dma_addr_t input_addr,
> +					       dma_addr_t output_addr,
> +					       int urg_reg_num,
> +					       enum urg_reg_cxts command)
> +{
> +	void __iomem *td_dst_src_reg = qpace_priv->urg_regs +
> +				       (urg_reg_num * QPACE_REG_PAGE_SIZE) +
> +					QPACE_URG_CMD_0_TD_DST_ADDR_L_CFG_CNTXT_OFFSET;
> +	u64 urg_addr_field_lower, urg_addr_field_upper;
> +	u32 stat_reg;
> +	unsigned long ret;
> +
> +	urg_addr_field_lower = FIELD_PREP(URG_CMD_0_TD_DST_ADDR_L__CMD_CFG_CNTXT,
> +					  command);
> +	urg_addr_field_lower |= GENMASK(63, 8) & output_addr;

[Severity: Medium]
Will this silently truncate the output_addr if the caller passes an address
that is not 256-byte aligned? In qpace_urgent_command_trigger(), this drops
the lowest 8 bits, meaning the hardware will write to the incorrectly
truncated address and overwrite up to 255 bytes of memory before the
intended buffer.

> +
> +	urg_addr_field_upper = input_addr;
> +
> +	/* Clear a stale request-on-active error before ringing the doorbell. */
> +	qpace_write_urg_cmd(qpace_priv, urg_reg_num, QPACE_URG_CMD_0_STAT_CLR_OFFSET,
> +			    URG_CMD_0_STAT_CLR_REQ_ON_ACTIVE_ERR);
> +
> +	/* Ensure that preceding stores that QPaCE will depend on are done executing */
> +	dma_wmb();
> +
> +	/*
> +	 * Ring the doorbell with a single 128-bit atomic store. QPaCE
> +	 * triggers processing when it observes the (dst_addr, src_addr)
> +	 * pair land together.
> +	 */
> +	asm volatile("stp %0, %1, [%2]\n" :
> +		     : "r" (urg_addr_field_lower), "r" (urg_addr_field_upper), "r" (td_dst_src_reg)
> +		     : "memory");

[Severity: Medium]
Does this inline assembly break big-endian architectures? By bypassing the
standard writeq() accessors, this stp instruction skips implicit cpu_to_le64
conversions, causing the hardware to receive byte-swapped, garbage memory
addresses and context configurations if the kernel is compiled with
CONFIG_CPU_BIG_ENDIAN.

> +
> +	ret = readl_poll_timeout_atomic(qpace_priv->urg_regs +
> +					urg_reg_num * QPACE_REG_PAGE_SIZE +
> +					QPACE_URG_CMD_0_ED_STAT_OFFSET,
> +					stat_reg,
> +					FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE,
> +						  stat_reg) != OP_URG_ONGOING,
> +					1, 100 * USEC_PER_MSEC);

[Severity: High]
Does this cause severe scheduling latency spikes or RCU stalls? The caller,
such as qpace_urgent_compress(), disables preemption via get_cpu() before
initiating this hardware poll. Executing a 100ms busy-wait loop in
qpace_urgent_command_trigger() while preemption is completely disabled
can freeze CPU cores for unacceptable durations under load.

> +	if (ret) {
> +		dev_err(qpace_priv->dev, "QPaCE urgent command timed out\n");
> +		return -ETIMEDOUT;
> +	}

[Severity: High]
Are we returning a timeout error here without stopping the hardware from
continuing the DMA operation? If the hardware takes longer than 100ms in
qpace_urgent_command_trigger(), this returns -ETIMEDOUT while the device
continues. If the caller (like zram) frees the output page, the hardware
could later finish and cause delayed asynchronous memory corruption.

> +
> +	return stat_reg;
> +}
> +
> +int qpace_urgent_compress(dma_addr_t input_addr,
> +			  dma_addr_t output_addr,
> +			  struct qpace_algorithm *algo)
> +{
> +	int urg_reg_num;
> +	int stat_reg;
> +	u32 stat_reg_val;
> +	int ret;
> +
> +	ret = qpace_get();
> +	if (ret)
> +		return ret;
> +
> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
> +						algo->urg_comp_cntxt);

[Severity: High]
Can this assign identical hardware registers to different CPUs running
concurrently when NR_CPUS is greater than 8 without cross-CPU mutual
exclusion? In qpace_urgent_compress(), get_cpu() disables preemption but
provides no mutual exclusion across different cores. If two threads initiate
compression on CPU 0 and CPU 8 simultaneously, they will both be assigned
urg_reg_num = 0 and race on the same MMIO block, leading to data corruption
and hardware timeouts.

> +	put_cpu();
> +
> +	qpace_put();
[ ... ]
> +int qpace_urgent_decompress(dma_addr_t input_addr,
> +			    dma_addr_t output_addr,
> +			    size_t input_size,
> +			    struct qpace_algorithm *algo)
> +{
> +	int urg_reg_num;
> +	int stat_reg;
> +	u32 stat_reg_val;
> +	int ret;
> +
> +	ret = qpace_get();
> +	if (ret)
> +		goto out;
> +
> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
> +	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
> +				urg_reg_num, algo->urg_decomp_cntxt,
> +				FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, input_size));

[Severity: High]
Does this dynamically overwrite the shared hardware context size register
concurrently across aliasing CPUs? In qpace_urgent_decompress(), with
more than 8 cores, aliasing CPUs will race on this MMIO block leading to
data corruption.

> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
> +						algo->urg_decomp_cntxt);
> +	put_cpu();
[ ... ]
> +int qpace_get(void)
> +{
> +	int ret = 0;
> +
> +	mutex_lock(&qpace_ref_lock);
> +	if (qpace_priv->suspended || READ_ONCE(qpace_priv->broken))
> +		ret = -EBUSY;

[Severity: High]
Can this exported API blindly dereference the global pointer after the driver
has been unbound? If the device is unbound via sysfs, qpace_priv is freed,
but the global pointer remains set. A subsequent call from a client like zram
will call qpace_get() which dereferences qpace_priv->suspended, resulting in
a use-after-free.

> +	else
> +		_get_qpace();
> +	mutex_unlock(&qpace_ref_lock);
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(qpace_get);
[ ... ]
> +static irqreturn_t urgent_interrupt_handler(int irq, void *unused)
> +{
> +	pr_debug("Urgent interrupt handled\n");
> +	return IRQ_HANDLED;
> +}

[Severity: High]
Will this cause a catastrophic interrupt storm? The
urgent_interrupt_handler() fails to acknowledge or clear the interrupt at
the device level. Because the GIC interrupt is level-triggered, returning
IRQ_HANDLED without clearing it on the hardware will cause it to
immediately re-fire forever, resulting in a hard CPU lockup if the
hardware ever asserts this line.

[ ... ]
> +static void qpace_remove(struct platform_device *pdev)
> +{
> +	mutex_lock(&qpace_ref_lock);
> +	qpace_priv->suspended = true;
> +	mutex_unlock(&qpace_ref_lock);
> +
> +	static_branch_disable(&qpace_drv_probed);
> +
> +	wait_for_completion(&qpace_priv->no_active_refs);
> +
> +	/* No callers remain; tear down the hardware. */
> +	cancel_work_sync(&qpace_priv->disable_work);
> +	qpace_power_off(&pdev->dev);
> +}

[Severity: High]
Does this leave the global qpace_priv pointer set? In qpace_remove(), the
device is torn down and devres will free the memory backing qpace_priv, but
the global pointer is never cleared, enabling the use-after-free mentioned
above.

[ ... ]
> +static int qpace_resume(struct device *dev)
> +{
> +	int ret;
> +
> +	mutex_lock(&qpace_ref_lock);
> +
> +	if (qpace_priv->active_rings)
> +		dev_err(dev, "active_rings not 0, unexpected case\n");
> +
> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
> +	if (ret)
> +		goto out_unlock;
> +
> +	program_urg_command_contexts_v2();
> +	program_decomp_core_cfg();
> +
> +	qpace_priv->suspended = false;

[Severity: High]
Will the hardware be stuck in STOP mode after suspend/resume? In
qpace_resume(), we restore contexts but fail to wake the hardware engine
up to RUN mode (omitting the QPACE_RUN write) and fail to reprogram
essential bulk mode and cache index settings from qpace_hw_init(). This
means all hardware operations will indefinitely time out after the system
resumes.

> +
> +out_unlock:
> +	if (ret)
> +		dev_err(dev, "failed to resume QPaCE: %d\n", ret);
> +	mutex_unlock(&qpace_ref_lock);
> +	return ret;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930145215.3380977-1-georgi.djakov@oss.qualcomm.com?part=2

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support
  2026-09-30 14:52 ` [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support Georgi Djakov
@ 2026-09-30 15:08   ` sashiko-bot
  0 siblings, 0 replies; 25+ messages in thread
From: sashiko-bot @ 2026-09-30 15:08 UTC (permalink / raw)
  To: Georgi Djakov; +Cc: robh, linux-trace-kernel, devicetree, conor+dt

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] qcom/qpace: unconditional probe failure if LLCC is disabled or absent
- [Medium] qcom/qpace: double deactivation of LLCC slices on failed resume
- [Low] qcom/qpace: missing Kconfig dependency causes linker error when LLCC=m
--

commit 91b78e77e6f189e8963165dc5c4bd1f6ed1995ec
Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>

soc: qcom: qpace: Add LLCC slice support

This commit adds support for QPaCE to use dedicated LLCC slices for compression
and decompression. It updates the driver to request and activate the LLCC slices
during probe and resume, and deactivate them during suspend and remove.

> diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
> index f5b21ffbcbff5..86a935ae2ca93 100644
> --- a/drivers/soc/qcom/qpace.c
> +++ b/drivers/soc/qcom/qpace.c
[ ... ]
> @@ -660,9 +663,35 @@ static int qpace_probe(struct platform_device *pdev)
>  	if (ret)
>  		return dev_err_probe(dev, ret, "failed to map QPaCE registers\n");
>  
> +	priv->llc_comp = llcc_slice_getd(LLCC_QPACE_COMPRESSION);
> +	if (IS_ERR(priv->llc_comp))
> +		return dev_err_probe(dev, PTR_ERR(priv->llc_comp),
> +				     "failed to get compression LLCC slice\n");

[Severity: High]
Does this unconditionally fail the probe if LLCC is disabled or missing from
the device tree? If CONFIG_QCOM_LLCC=n, llcc_slice_getd() returns NULL (which
passes the IS_ERR check). Then the subsequent llcc_slice_activate(NULL)
invokes a macro stub returning -EINVAL, unconditionally failing the probe.

Additionally, if the QPaCE slice is absent from the device tree, it returns
an error pointer, which also aborts probe. Should this be treated as an
optional optimization rather than a hard requirement?

[Severity: Low]
Can this cause a linker error if CONFIG_QCOM_PAGE_COMPRESSION_ENGINE=y but
CONFIG_QCOM_LLCC=m? In this configuration, the built-in QPaCE driver cannot
link against the exported symbols of the LLCC module, but
IS_ENABLED(CONFIG_QCOM_LLCC) evaluates to true, so it attempts to use the
header declarations rather than inline stubs. Is there a missing Kconfig
dependency?

> +
> +	priv->llc_decomp = llcc_slice_getd(LLCC_QPACE_DECOMPRESSION);
> +	if (IS_ERR(priv->llc_decomp)) {
> +		ret = dev_err_probe(dev, PTR_ERR(priv->llc_decomp),
> +				    "failed to get decompression LLCC slice\n");
> +		priv->llc_decomp = NULL;
> +		goto llc_put_comp;
> +	}
> +
> +	ret = llcc_slice_activate(priv->llc_comp);
> +	if (ret) {
> +		dev_err_probe(dev, ret, "failed to activate compression LLCC slice\n");
> +		goto llc_put_decomp;
> +	}
[ ... ]
> @@ -708,6 +745,10 @@ static void qpace_remove(struct platform_device *pdev)
>  	/* No callers remain; tear down the hardware. */
>  	cancel_work_sync(&qpace_priv->disable_work);
>  	qpace_power_off(&pdev->dev);
> +	llcc_slice_deactivate(qpace_priv->llc_decomp);
> +	llcc_slice_deactivate(qpace_priv->llc_comp);

[Severity: Medium]
Can this cause a double-deactivation if a previous qpace_resume() failed? If
an error occurs during qpace_resume() (e.g., if icc_set_bw() fails), the
driver leaves the slices deactivated and returns early.

If the device is subsequently unbound, qpace_remove() unconditionally calls
llcc_slice_deactivate() on both slices. Since their refcounts would already
be 0, wouldn't this underflow the refcount, triggering a
WARN_ON_ONCE(!refcount_dec_if_one(...)) in the LLCC core and sending an
erroneous deactivate command to the hardware?

> +	llcc_slice_putd(qpace_priv->llc_decomp);
> +	llcc_slice_putd(qpace_priv->llc_comp);
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930145215.3380977-1-georgi.djakov@oss.qualcomm.com?part=5

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
@ 2026-09-30 15:08   ` sashiko-bot
  2026-10-01  5:43   ` Sergey Senozhatsky
  2026-10-01  8:51   ` Krzysztof Kozlowski
  2 siblings, 0 replies; 25+ messages in thread
From: sashiko-bot @ 2026-09-30 15:08 UTC (permalink / raw)
  To: Georgi Djakov; +Cc: linux-trace-kernel, conor+dt, devicetree, robh

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] block: zram: signed/unsigned promotion masks hardware errors in qpace_compress
- [High] block: zram: missing qpace availability check causes NULL pointer dereference
- [High] block: zram: unbinding qpace driver causes Use-After-Free in backend
- [Medium] block: zram: qpace backend uses bounce buffers instead of claimed zero-copy DMA mapping
--

commit e4fba83b2c0230490d177ce709e7ea22807d84d4
Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>

zram: Add QPaCE zcomp backend

Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
zram can expose qpace-lz4 as a selectable compression algorithm when the
QPaCE driver is available.

> Compress and decompress operations are handled via the QPaCE urgent
> synchronous path: each request DMA-maps the source and destination
> buffers, issues a blocking hardware command, and returns the result size.

[Severity: Medium]
Does the implementation contradict the commit message's claim about DMA
mapping the request buffers? The code allocates static bounce buffers and
performs synchronous memory copies instead of mapping the actual request
memory directly.

> diff --git a/drivers/block/zram/backend_qpace.c b/drivers/block/zram/backend_qpace.c
> new file mode 100644
> index 0000000000000..3af13727adefc
> --- /dev/null
> +++ b/drivers/block/zram/backend_qpace.c
[ ... ]
> +static int qpace_create_ctx(struct zcomp_params *params, struct zcomp_ctx *ctx)
> +{
> +	struct device *dev = qpace_get_dma_dev();

[Severity: High]
Is it safe to unconditionally call qpace_get_dma_dev() here without
verifying if the hardware actually probed via qpace_is_dev_available()?

If the QPaCE driver hasn't probed successfully or is missing, the global
qpace_priv pointer remains NULL, and qpace_get_dma_dev() dereferences it:

drivers/soc/qcom/qpace.c:qpace_get_dma_dev() {
    return qpace_priv->dev;
}

Could this cause a kernel panic if a user writes 'qpace-lz4' to
/sys/block/zram0/comp_algorithm on an unsupported system?

[Severity: High]
Additionally, does this cache the device pointer without acquiring a robust
reference via get_device() or taking a module reference?

If the QPaCE driver is unbound via sysfs, qpace_priv is destroyed.
Subsequent compress operations or teardown might access the freed device
pointer or qpace_priv structure, leading to a use-after-free.

> +	struct qpace_ctx *cc;
> +
> +	if (!dev)
> +		return -ENODEV;
> +
> +	cc = kzalloc_obj(*cc, GFP_KERNEL);
> +	if (!cc)
> +		return -ENOMEM;
> +
> +	cc->dev = dev;
> +	cc->in_buf = dma_alloc_coherent(dev, PAGE_SIZE, &cc->in_dma, GFP_KERNEL);

[Severity: Medium]
This statically allocates a bounce buffer rather than preparing to map the
dynamic request buffers. Does this introduce CPU overhead by breaking the
zero-copy DMA design claimed in the commit message?

[ ... ]
> +static void qpace_destroy_ctx(struct zcomp_ctx *ctx)
> +{
> +	struct qpace_ctx *cc = ctx->context;
> +
> +	if (!cc)
> +		return;
> +
> +	dma_free_coherent(cc->dev, PAGE_SIZE, cc->out_buf, cc->out_dma);

[Severity: High]
If the device was unbound and qpace_priv was freed, could accessing cc->dev
here trigger a use-after-free during cleanup?

[ ... ]
> +static int qpace_compress(struct zcomp_params *params, struct zcomp_ctx *ctx,
> +			  struct zcomp_req *req)
> +{
> +	struct qpace_algorithm *algo = params->drv_data;
> +	struct qpace_ctx *cc = ctx->context;
> +	int ret;
> +
> +	memcpy(cc->in_buf, req->src, req->src_len);

[Severity: Medium]
Does copying data to the bounce buffer here confirm the data path is not
zero-copy?

> +
> +	ret = qpace_urgent_compress(cc->in_dma, cc->out_dma, algo);
> +
> +	if (ret == -E2BIG || ret >= PAGE_SIZE) {

[Severity: High]
Are hardware errors from the QPaCE compressor masked here?

If qpace_urgent_compress() returns a negative error code (e.g., -EINVAL),
comparing the signed integer 'ret' to PAGE_SIZE (typically unsigned long)
promotes 'ret' to a large unsigned value.

Would this cause the check to incorrectly pass, silently treating broken or
failed compressions as incompressible data instead of propagating the failure?

> +		req->dst_len = PAGE_SIZE;
> +		return 0;
> +	}
[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930145215.3380977-1-georgi.djakov@oss.qualcomm.com?part=4

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
  2026-09-30 15:08   ` sashiko-bot
@ 2026-10-01  5:43   ` Sergey Senozhatsky
  2026-10-06 23:53     ` Oreoluwa Babatunde
  2026-10-01  8:51   ` Krzysztof Kozlowski
  2 siblings, 1 reply; 25+ messages in thread
From: Sergey Senozhatsky @ 2026-10-01  5:43 UTC (permalink / raw)
  To: Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On (26/09/30 07:52), Georgi Djakov wrote:
[..]
> @@ -80,10 +87,10 @@ static const struct zcomp_ops *lookup_backend_ops(const char *comp)
>  
>  	while (backends[i]) {
>  		if (sysfs_streq(comp, backends[i]->name))
> -			break;
> +			return backends[i];
>  		i++;
>  	}
> -	return backends[i];
> +	return NULL;
>  }

This hunk looks unrelated to the patch in question.

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-09-30 14:52 ` [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver Georgi Djakov
  2026-09-30 15:04   ` sashiko-bot
@ 2026-10-01  8:50   ` Krzysztof Kozlowski
  2026-10-06 23:51     ` Oreoluwa Babatunde
  1 sibling, 1 reply; 25+ messages in thread
From: Krzysztof Kozlowski @ 2026-10-01  8:50 UTC (permalink / raw)
  To: Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On Wed, Sep 30, 2026 at 07:52:11AM -0700, Georgi Djakov wrote:
> Add a platform driver for the Qualcomm Page Compression Engine (QPaCE), a
> hardware block that accelerates compression and decompression of memory
> pages.
> 
> Provide the urgent command path for synchronous single-page compression and
> decompression. This exposes the low-latency operations needed by
> compressed-memory users such as zram, especially for page decompression on
> the read path.
> 
> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
> ---
>  drivers/soc/qcom/Kconfig          |  14 +
>  drivers/soc/qcom/Makefile         |   1 +
>  drivers/soc/qcom/qpace.c          | 764 ++++++++++++++++++++++++++++++
>  drivers/soc/qcom/qpace_internal.h |  84 ++++
>  include/linux/soc/qcom/qpace.h    | 154 ++++++
>  5 files changed, 1017 insertions(+)
>  create mode 100644 drivers/soc/qcom/qpace.c
>  create mode 100644 drivers/soc/qcom/qpace_internal.h
>  create mode 100644 include/linux/soc/qcom/qpace.h
> 
> diff --git a/drivers/soc/qcom/Kconfig b/drivers/soc/qcom/Kconfig
> index 535c8619197b..6bcb86dcd726 100644
> --- a/drivers/soc/qcom/Kconfig

Sorry, but no. Soc is not a dumping ground. This has clear function of
compression offload, so it should have some dedicated maintainers like
other offload engines.

> +++ b/drivers/soc/qcom/Kconfig
> @@ -288,6 +288,20 @@ config QCOM_PBS
>  	  This module provides the APIs to the client drivers that wants to send the
>  	  PBS trigger event to the PBS RAM.
>  
> +config QCOM_PAGE_COMPRESSION_ENGINE
> +	tristate "Qualcomm Page Compression Engine (QPaCE)"
> +	depends on ARM64

Why this can't be built on other archs? This is really odd and I do not
see any asm headers included.


> +	depends on ARCH_QCOM || COMPILE_TEST
> +	depends on OF
> +	depends on INTERCONNECT
> +	help
> +	  Enable support for the Qualcomm Page Compression Engine (QPaCE),
> +	  a hardware accelerator that provides high-throughput page compression,
> +	  decompression, and DMA copy operations.
> +
> +	  The engine is used as a hardware backend for compressed-memory
> +	  subsystems such as zram. If unsure, say N.
> +

...


> +	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
> +out:
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(qpace_urgent_compress);
> +
> +int qpace_urgent_decompress(dma_addr_t input_addr,
> +			    dma_addr_t output_addr,
> +			    size_t input_size,
> +			    struct qpace_algorithm *algo)

You need kerneldoc for every export.

> +{
> +	int urg_reg_num;
> +	int stat_reg;
> +	u32 stat_reg_val;
> +	int ret;
> +
> +	ret = qpace_get();
> +	if (ret)
> +		goto out;
> +
> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
> +	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
> +				urg_reg_num, algo->urg_decomp_cntxt,
> +				FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, input_size));
> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
> +						algo->urg_decomp_cntxt);
> +	put_cpu();
> +
> +	qpace_put();
> +
> +	if (stat_reg < 0) {
> +		ret = stat_reg;
> +		goto out;
> +	}
> +
> +	stat_reg_val = FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE, stat_reg);
> +	if (stat_reg_val != OP_OK) {
> +		pr_err("%s: register %d failed with %u\n",
> +		       __func__, urg_reg_num, stat_reg_val);
> +		ret = -EINVAL;
> +		goto out;
> +	}
> +
> +	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
> +out:
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(qpace_urgent_decompress);
> +

So singleton? For what reason exactly? Random drivers will be getting
the reference to compress something? If so, aren't you duplicating
existing infrastructure/API for in-kernel hardware offloaded
compression (e.g. drivers/crypto/)?

You miss proper comments (see checkpatch --strict) explaining lock
usage.


> +static DEFINE_MUTEX(qpace_ref_lock);
> +
> +static void _get_qpace(void)
> +{
> +	lockdep_assert_held(&qpace_ref_lock);
> +	if (!qpace_priv->active_rings) {
> +		reinit_completion(&qpace_priv->no_active_refs);
> +		pm_stay_awake(qpace_priv->dev);
> +		cpu_latency_qos_update_request(&qpace_priv->qos_req, 300);
> +		program_urg_command_contexts_v2();
> +		program_decomp_core_cfg();
> +	}
> +	qpace_priv->active_rings++;
> +}
> +
> +static void _put_qpace(void)
> +{
> +	lockdep_assert_held(&qpace_ref_lock);
> +	if (!--qpace_priv->active_rings) {
> +		cpu_latency_qos_update_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
> +		pm_relax(qpace_priv->dev);
> +		complete(&qpace_priv->no_active_refs);
> +	}
> +}
> +
> +int qpace_get(void)
> +{
> +	int ret = 0;
> +
> +	mutex_lock(&qpace_ref_lock);
> +	if (qpace_priv->suspended || READ_ONCE(qpace_priv->broken))
> +		ret = -EBUSY;
> +	else
> +		_get_qpace();
> +	mutex_unlock(&qpace_ref_lock);
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(qpace_get);
> +
> +void qpace_put(void)
> +{
> +	mutex_lock(&qpace_ref_lock);
> +	_put_qpace();
> +	mutex_unlock(&qpace_ref_lock);
> +}
> +EXPORT_SYMBOL_GPL(qpace_put);
> +
> +static irqreturn_t urgent_interrupt_handler(int irq, void *unused)
> +{
> +	pr_debug("Urgent interrupt handled\n");
> +	return IRQ_HANDLED;
> +}
> +
> +static int qpace_hw_init(void)
> +{
> +	u32 reg_val;
> +
> +	/* Select CPU SCID for our system cache slice. */
> +	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET);
> +	reg_val = u32_replace_bits(reg_val, 0x1, CORE_QNS4_CFG_CACHEINDEX);
> +	qpace_write_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET, reg_val);
> +
> +	/* QMB2 register configurations. */
> +	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET);
> +	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_RD_OUTST_LIMIT);
> +	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_WR_OUTST_LIMIT);
> +	qpace_write_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET, reg_val);
> +
> +	/* DECOMP_CORE_CFG init steps. */
> +	program_decomp_core_cfg();
> +
> +	/* Below settings help save power since all decomp cores are set to sync. */
> +	reg_val = qpace_read_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET);
> +	reg_val |= CORE_OPER_CFG_COMP_MEM_PWR_DWN_1;
> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET, reg_val);
> +
> +	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET);
> +	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_RD_MAX_OT);
> +	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_WR_MAX_OT);
> +	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET, reg_val);
> +
> +	/* Set all COMP engines to bulk mode. */
> +	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET);
> +	reg_val |= COMP_CORE_BULK_MODE_ALL_CORES;
> +	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET, reg_val);
> +
> +	/* URG CMD register configurations. */
> +	program_urg_command_contexts_v2();
> +
> +	return 0;
> +}
> +
> +enum qpace_interrupts {
> +	QPACE_IRQ_URGENT
> +};
> +
> +static int qpace_register_interrupts(struct platform_device *pdev)
> +{
> +	struct device *dev = &pdev->dev;
> +	int irq, ret;
> +
> +	irq = platform_get_irq(pdev, QPACE_IRQ_URGENT);
> +	if (irq < 0)
> +		return irq;
> +
> +	ret = devm_request_irq(dev, irq, urgent_interrupt_handler,
> +			       0, "qpace-urgent-irq", NULL);
> +	if (ret)
> +		dev_err(dev, "failed to request urgent interrupt\n");
> +
> +	return ret;
> +}
> +
> +static inline bool _qpace_power_on(void)
> +{
> +	u32 ready_status;
> +
> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_RUN);
> +
> +	if (readl_poll_timeout(qpace_priv->gen_core_regs +
> +			       QPACE_CORE_OPER_CORE_READY_OFFSET,
> +			       ready_status, ready_status,
> +			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
> +		pr_err("Timeout in waiting for QPaCE to turn on\n");
> +		return false;
> +	}
> +
> +	return true;
> +}
> +
> +static int qpace_power_on(struct device *dev)
> +{
> +	int ret, ret2;
> +
> +	qpace_priv->interconnect = devm_of_icc_get(dev, "qpace-mem");
> +	if (IS_ERR_OR_NULL(qpace_priv->interconnect)) {
> +		ret = PTR_ERR_OR_ZERO(qpace_priv->interconnect);
> +		pr_err("%s: devm_of_icc_get() failed with %d\n", __func__, ret);

use dev_err, not pr_err


> +		return qpace_priv->interconnect ? ret : -EINVAL;
> +	}
> +
> +	ret = device_init_wakeup(dev, true);
> +	if (ret) {
> +		pr_err("%s: device_init_wakeup() failed with %d\n", __func__, ret);
> +		return ret;
> +	}
> +
> +	cpu_latency_qos_add_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
> +
> +	icc_set_tag(qpace_priv->interconnect, QCOM_ICC_TAG_ACTIVE_ONLY);
> +
> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
> +	if (ret) {
> +		pr_err("Failed to turn on QPaCE VCD: %d\n", ret);
> +		goto rm_qos;
> +	}
> +
> +	if (!_qpace_power_on()) {
> +		pr_err("Failed to start QPaCE\n");
> +		ret = -EINVAL;
> +		goto rm_bw;
> +	}
> +
> +	return 0;
> +
> +rm_bw:
> +	ret2 = icc_set_bw(qpace_priv->interconnect, 0, 0);
> +	if (ret2)
> +		pr_err("Failed to remove QPaCE VCD vote: %d\n", ret2);
> +rm_qos:
> +	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
> +	device_init_wakeup(dev, false);
> +
> +	return ret;
> +}
> +
> +static inline bool _qpace_power_off(void)
> +{
> +	u32 ready_status;
> +
> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_STOP);
> +
> +	if (readl_poll_timeout(qpace_priv->gen_core_regs +
> +			       QPACE_CORE_OPER_CORE_READY_OFFSET,
> +			       ready_status, !ready_status,
> +			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
> +		pr_err("Timeout in waiting for QPaCE to turn off\n");
> +		return false;
> +	}
> +
> +	return true;
> +}
> +
> +static void qpace_power_off(struct device *dev)
> +{
> +	int ret;
> +
> +	/* If this fails we can still remove our vote for the VCD to turn QPaCE off */
> +	if (!_qpace_power_off())
> +		pr_err("Failed to stop QPaCE\n");
> +
> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 0);
> +	if (ret)
> +		pr_err("Failed to turn off QPaCE VCD: %d\n", ret);
> +
> +	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
> +
> +	device_init_wakeup(dev, false);
> +}
> +
> +static inline int qpace_register_ioremap(struct platform_device *pdev)
> +{
> +	qpace_priv->gen_regs = devm_platform_ioremap_resource(pdev, 0);
> +	if (IS_ERR(qpace_priv->gen_regs))
> +		return PTR_ERR(qpace_priv->gen_regs);
> +
> +	qpace_priv->gen_core_regs = qpace_priv->gen_regs + QPACE_GEN_CORE_REGS_OFFSET;
> +	qpace_priv->comp_core_regs = qpace_priv->gen_regs + QPACE_COMP_CORE_REGS_OFFSET;
> +	qpace_priv->decomp_core_regs = qpace_priv->gen_regs + QPACE_DECOMP_CORE_REGS_OFFSET;
> +	qpace_priv->urg_regs = qpace_priv->gen_regs + QPACE_URG_REGS_OFFSET;
> +
> +	return 0;
> +}
> +
> +bool qpace_is_dev_available(void)
> +{
> +	return static_branch_likely(&qpace_drv_probed) &&
> +	       !READ_ONCE(qpace_priv->broken);
> +}
> +EXPORT_SYMBOL_GPL(qpace_is_dev_available);
> +
> +struct device *qpace_get_dma_dev(void)
> +{
> +	return qpace_priv->dev;
> +}
> +EXPORT_SYMBOL_GPL(qpace_get_dma_dev);
> +
> +static int qpace_probe(struct platform_device *pdev)
> +{
> +	struct device *dev = &pdev->dev;
> +	struct qpace_priv *priv;
> +	int ret;
> +
> +	priv = devm_kzalloc(dev, sizeof(*priv), GFP_KERNEL);
> +	if (!priv)
> +		return -ENOMEM;
> +
> +	priv->dev = dev;
> +	/* Starts already complete since active_rings == 0 at init. */
> +	init_completion(&priv->no_active_refs);
> +	complete(&priv->no_active_refs);
> +	INIT_WORK(&priv->disable_work, qpace_disable_work_fn);
> +	qpace_priv = priv;
> +	platform_set_drvdata(pdev, priv);
> +
> +	ret = dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
> +	if (ret)
> +		return dev_err_probe(dev, ret, "failed to set DMA mask\n");
> +
> +	ret = qpace_register_ioremap(pdev);
> +	if (ret)
> +		return dev_err_probe(dev, ret, "failed to map QPaCE registers\n");
> +
> +	ret = qpace_power_on(dev);
> +	if (ret)
> +		return dev_err_probe(dev, ret, "failed to power on QPaCE\n");
> +
> +	/* Get QPaCE HW version. */
> +	qpace_priv->hw_version = qpace_read_gen(qpace_priv, QPACE_CORE_HW_VERSION_OFFSET);
> +	if (qpace_priv->hw_version != QPACE_HW_VERSION_V2) {
> +		dev_err(dev, "Unsupported QPaCE HW version returned: 0x%x\n",
> +			qpace_priv->hw_version);
> +		ret = -EINVAL;
> +		goto power_off;
> +	}
> +
> +	ret = qpace_hw_init();
> +	if (ret) {

How is this possible?

> +		dev_err(dev, "init failed: (%d)\n", ret);
> +		goto power_off;
> +	}
> +
> +	ret = qpace_register_interrupts(pdev);
> +	if (ret) {
> +		dev_err(dev, "failed to register interrupts\n");

Do not print same error multiple times.

> +		goto power_off;
> +	}
> +
> +	static_branch_enable(&qpace_drv_probed);
> +
> +	return ret;
> +
> +power_off:
> +	qpace_power_off(dev);
> +
> +	return ret;
> +}

Best regards,
Krzysztof


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
  2026-09-30 15:08   ` sashiko-bot
  2026-10-01  5:43   ` Sergey Senozhatsky
@ 2026-10-01  8:51   ` Krzysztof Kozlowski
  2026-10-01 10:27     ` Sergey Senozhatsky
  2026-10-07  0:03     ` Oreoluwa Babatunde
  2 siblings, 2 replies; 25+ messages in thread
From: Krzysztof Kozlowski @ 2026-10-01  8:51 UTC (permalink / raw)
  To: Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On Wed, Sep 30, 2026 at 07:52:13AM -0700, Georgi Djakov wrote:
> Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
> zram can expose qpace-lz4 as a selectable compression algorithm when the
> QPaCE driver is available.
> 
> Compress and decompress operations are handled via the QPaCE urgent
> synchronous path: each request DMA-maps the source and destination
> buffers, issues a blocking hardware command, and returns the result size.
> 
> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
> ---
>  drivers/block/zram/Kconfig         |  11 +++
>  drivers/block/zram/Makefile        |   1 +
>  drivers/block/zram/backend_qpace.c | 134 +++++++++++++++++++++++++++++

So entire qpace should go here, no? Why did you create this entire layer
of indirection, singleton management under drivers/soc?

>  drivers/block/zram/backend_qpace.h |  13 +++
>  drivers/block/zram/zcomp.c         |  11 ++-
>  5 files changed, 168 insertions(+), 2 deletions(-)
>  create mode 100644 drivers/block/zram/backend_qpace.c
>  create mode 100644 drivers/block/zram/backend_qpace.h

Best regards,
Krzysztof


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-10-01  8:51   ` Krzysztof Kozlowski
@ 2026-10-01 10:27     ` Sergey Senozhatsky
  2026-10-07  0:04       ` Oreoluwa Babatunde
  2026-10-07  0:03     ` Oreoluwa Babatunde
  1 sibling, 1 reply; 25+ messages in thread
From: Sergey Senozhatsky @ 2026-10-01 10:27 UTC (permalink / raw)
  To: Krzysztof Kozlowski
  Cc: Georgi Djakov, andersson, konradybcio, abelvesa, robh, krzk+dt,
	conor+dt, minchan, senozhatsky, axboe, rostedt, mhiramat,
	mathieu.desnoyers, linux-arm-msm, devicetree, linux-kernel,
	linux-block, linux-trace-kernel, djakov

On (26/10/01 10:51), Krzysztof Kozlowski wrote:
> On Wed, Sep 30, 2026 at 07:52:13AM -0700, Georgi Djakov wrote:
> > Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
> > zram can expose qpace-lz4 as a selectable compression algorithm when the
> > QPaCE driver is available.
> > 
> > Compress and decompress operations are handled via the QPaCE urgent
> > synchronous path: each request DMA-maps the source and destination
> > buffers, issues a blocking hardware command, and returns the result size.
> > 
> > Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
> > ---
> >  drivers/block/zram/Kconfig         |  11 +++
> >  drivers/block/zram/Makefile        |   1 +
> >  drivers/block/zram/backend_qpace.c | 134 +++++++++++++++++++++++++++++
> 
> So entire qpace should go here, no? Why did you create this entire layer
> of indirection, singleton management under drivers/soc?

qpace seems to be standalone/independent arch/soc specific, and can gain
users outside of zram (fs compression, zswap, etc.).  zram's backends
potentially can disappear all together, if we switch to acomp crypto API.

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding
  2026-09-30 14:52 ` [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding Georgi Djakov
@ 2026-10-02  6:10   ` Krzysztof Kozlowski
  0 siblings, 0 replies; 25+ messages in thread
From: Krzysztof Kozlowski @ 2026-10-02  6:10 UTC (permalink / raw)
  To: Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On Wed, Sep 30, 2026 at 07:52:10AM -0700, Georgi Djakov wrote:
> Add device-tree binding documentation for the Qualcomm Page Compression
> Engine (QPaCE), including the resources needed to describe it in SoC device
> trees.
> 
> QPaCE accelerates page compression and decompression for use with the zram
> block device.
> 
> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
> ---
>  .../bindings/soc/qcom/qcom,hawi-qpace.yaml    | 94 +++++++++++++++++++
>  1 file changed, 94 insertions(+)
>  create mode 100644 Documentation/devicetree/bindings/soc/qcom/qcom,hawi-qpace.yaml
> 

Reviewed-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>

Best regards,
Krzysztof


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-09-30 15:04   ` sashiko-bot
@ 2026-10-03  0:06     ` Oreoluwa Babatunde
  0 siblings, 0 replies; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-03  0:06 UTC (permalink / raw)
  To: sashiko-reviews, Georgi Djakov
  Cc: robh, linux-trace-kernel, conor+dt, devicetree

On 9/30/2026 8:04 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 8 potential issue(s) to consider:
> - [High] soc: qcom: qpace: concurrent MMIO access from CPU aliasing on >8 cores
> - [High] soc: qcom: qpace: UAF of global qpace_priv after driver unbind
> - [High] soc: qcom: qpace: 100ms busy wait with preemption disabled causes latency spikes
> - [High] soc: qcom: qpace: interrupt storm due to uncleared level-triggered IRQ
> - [High] soc: qcom: qpace: stray DMA writes on urgent command timeout
> - [High] soc: qcom: qpace: hardware stuck in STOP mode after suspend/resume
> - [Medium] soc: qcom: qpace: silent truncation of unaligned output_addr causes DMA corruption
> - [Medium] soc: qcom: qpace: inline stp assembly breaks big-endian architectures
> --
> 
> commit 5edbc123f3ebd977977e7d874e4ecebc5ed95e6a
> Author: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
> 
> soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
> 
> This commit adds a platform driver for the Qualcomm Page Compression Engine
> (QPaCE). It provides the urgent command path for synchronous single-page
> compression and decompression, supporting low-latency operations for
> compressed-memory users like zram.
> 
>> diff --git a/drivers/soc/qcom/qpace.c b/drivers/soc/qcom/qpace.c
>> index 0000000000000..69aa14b7c4772
>> --- /dev/null
>> +++ b/drivers/soc/qcom/qpace.c
> [ ... ]
>> +static inline int qpace_urgent_command_trigger(dma_addr_t input_addr,
>> +					       dma_addr_t output_addr,
>> +					       int urg_reg_num,
>> +					       enum urg_reg_cxts command)
>> +{
>> +	void __iomem *td_dst_src_reg = qpace_priv->urg_regs +
>> +				       (urg_reg_num * QPACE_REG_PAGE_SIZE) +
>> +					QPACE_URG_CMD_0_TD_DST_ADDR_L_CFG_CNTXT_OFFSET;
>> +	u64 urg_addr_field_lower, urg_addr_field_upper;
>> +	u32 stat_reg;
>> +	unsigned long ret;
>> +
>> +	urg_addr_field_lower = FIELD_PREP(URG_CMD_0_TD_DST_ADDR_L__CMD_CFG_CNTXT,
>> +					  command);
>> +	urg_addr_field_lower |= GENMASK(63, 8) & output_addr;
> 
> [Severity: Medium]
> Will this silently truncate the output_addr if the caller passes an address
> that is not 256-byte aligned? In qpace_urgent_command_trigger(), this drops
> the lowest 8 bits, meaning the hardware will write to the incorrectly
> truncated address and overwrite up to 255 bytes of memory before the
> intended buffer.

Current use case for qpace is zram (although it is not limited to only 
zram). ZRAM operates on a page granularity which means that all 
addresses submitted to qpace will be 4KB aligned which is also 256 byte 
aligned.

> 
>> +
>> +	urg_addr_field_upper = input_addr;
>> +
>> +	/* Clear a stale request-on-active error before ringing the doorbell. */
>> +	qpace_write_urg_cmd(qpace_priv, urg_reg_num, QPACE_URG_CMD_0_STAT_CLR_OFFSET,
>> +			    URG_CMD_0_STAT_CLR_REQ_ON_ACTIVE_ERR);
>> +
>> +	/* Ensure that preceding stores that QPaCE will depend on are done executing */
>> +	dma_wmb();
>> +
>> +	/*
>> +	 * Ring the doorbell with a single 128-bit atomic store. QPaCE
>> +	 * triggers processing when it observes the (dst_addr, src_addr)
>> +	 * pair land together.
>> +	 */
>> +	asm volatile("stp %0, %1, [%2]\n" :
>> +		     : "r" (urg_addr_field_lower), "r" (urg_addr_field_upper), "r" (td_dst_src_reg)
>> +		     : "memory");
> 
> [Severity: Medium]
> Does this inline assembly break big-endian architectures? By bypassing the
> standard writeq() accessors, this stp instruction skips implicit cpu_to_le64
> conversions, causing the hardware to receive byte-swapped, garbage memory
> addresses and context configurations if the kernel is compiled with
> CONFIG_CPU_BIG_ENDIAN.
> 

ACK. Qpace is used in only little endian arch so far. We can look at a 
handling for big endian configurations as well.

>> +
>> +	ret = readl_poll_timeout_atomic(qpace_priv->urg_regs +
>> +					urg_reg_num * QPACE_REG_PAGE_SIZE +
>> +					QPACE_URG_CMD_0_ED_STAT_OFFSET,
>> +					stat_reg,
>> +					FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE,
>> +						  stat_reg) != OP_URG_ONGOING,
>> +					1, 100 * USEC_PER_MSEC);
> 
> [Severity: High]
> Does this cause severe scheduling latency spikes or RCU stalls? The caller,
> such as qpace_urgent_compress(), disables preemption via get_cpu() before
> initiating this hardware poll. Executing a 100ms busy-wait loop in
> qpace_urgent_command_trigger() while preemption is completely disabled
> can freeze CPU cores for unacceptable durations under load.

Under normal operating conditions qpace will finish processing a request
much sooner than 100ms.We have not seen any RCU stalls as a result of 
this 100ms timeout loop.

> 
>> +	if (ret) {
>> +		dev_err(qpace_priv->dev, "QPaCE urgent command timed out\n");
>> +		return -ETIMEDOUT;
>> +	}
> 
> [Severity: High]
> Are we returning a timeout error here without stopping the hardware from
> continuing the DMA operation? If the hardware takes longer than 100ms in
> qpace_urgent_command_trigger(), this returns -ETIMEDOUT while the device
> continues. If the caller (like zram) frees the output page, the hardware
> could later finish and cause delayed asynchronous memory corruption.
> 
If we ever run into the timeout path, then that means a hang has 
happened and qpace all toegther is not making forward progress on the 
request. This means it will not be doing any more DMA accesses.
But regardless I think adding some error handling to stop Qpace from SW 
is a good idea and can be implemented. Thanks.
>> +
>> +	return stat_reg;
>> +}
>> +
>> +int qpace_urgent_compress(dma_addr_t input_addr,
>> +			  dma_addr_t output_addr,
>> +			  struct qpace_algorithm *algo)
>> +{
>> +	int urg_reg_num;
>> +	int stat_reg;
>> +	u32 stat_reg_val;
>> +	int ret;
>> +
>> +	ret = qpace_get();
>> +	if (ret)
>> +		return ret;
>> +
>> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
>> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
>> +						algo->urg_comp_cntxt);
> 
> [Severity: High]
> Can this assign identical hardware registers to different CPUs running
> concurrently when NR_CPUS is greater than 8 without cross-CPU mutual
> exclusion? In qpace_urgent_compress(), get_cpu() disables preemption but
> provides no mutual exclusion across different cores. If two threads initiate
> compression on CPU 0 and CPU 8 simultaneously, they will both be assigned
> urg_reg_num = 0 and race on the same MMIO block, leading to data corruption
> and hardware timeouts.
> 
Current qpace HW configuraion is to work on soc with 8 cores so that 
these urg_cmd registers match without overlap. Alternatively, we can add 
a check at probe time for cpu topology and abort device initialization 
if num cpus > 8.

>> +	put_cpu();
>> +
>> +	qpace_put();
> [ ... ]
>> +int qpace_urgent_decompress(dma_addr_t input_addr,
>> +			    dma_addr_t output_addr,
>> +			    size_t input_size,
>> +			    struct qpace_algorithm *algo)
>> +{
>> +	int urg_reg_num;
>> +	int stat_reg;
>> +	u32 stat_reg_val;
>> +	int ret;
>> +
>> +	ret = qpace_get();
>> +	if (ret)
>> +		goto out;
>> +
>> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
>> +	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
>> +				urg_reg_num, algo->urg_decomp_cntxt,
>> +				FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, input_size));
> 
> [Severity: High]
> Does this dynamically overwrite the shared hardware context size register
> concurrently across aliasing CPUs? In qpace_urgent_decompress(), with
> more than 8 cores, aliasing CPUs will race on this MMIO block leading to
> data corruption.

ACK. Same response as above comment about > 8 cpus.

> 
>> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
>> +						algo->urg_decomp_cntxt);
>> +	put_cpu();
> [ ... ]
>> +int qpace_get(void)
>> +{
>> +	int ret = 0;
>> +
>> +	mutex_lock(&qpace_ref_lock);
>> +	if (qpace_priv->suspended || READ_ONCE(qpace_priv->broken))
>> +		ret = -EBUSY;
> 
> [Severity: High]
> Can this exported API blindly dereference the global pointer after the driver
> has been unbound? If the device is unbound via sysfs, qpace_priv is freed,
> but the global pointer remains set. A subsequent call from a client like zram
> will call qpace_get() which dereferences qpace_priv->suspended, resulting in
> a use-after-free.

ACK. Possibly we can make use of qpace_drv_probed in exported functions
to avoid using qpace if it has been removed or not probed.

> 
>> +	else
>> +		_get_qpace();
>> +	mutex_unlock(&qpace_ref_lock);
>> +	return ret;
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_get);
> [ ... ]
>> +static irqreturn_t urgent_interrupt_handler(int irq, void *unused)
>> +{
>> +	pr_debug("Urgent interrupt handled\n");
>> +	return IRQ_HANDLED;
>> +}
> 
> [Severity: High]
> Will this cause a catastrophic interrupt storm? The
> urgent_interrupt_handler() fails to acknowledge or clear the interrupt at
> the device level. Because the GIC interrupt is level-triggered, returning
> IRQ_HANDLED without clearing it on the hardware will cause it to
> immediately re-fire forever, resulting in a hard CPU lockup if the
> hardware ever asserts this line.

Currently as seen above, SW is noified of a completion event by polling 
the HW register and not using interrupts. When other usecases are added 
which use interrupt, then you are right that the HW registers will be 
reset/cleared.

> 
> [ ... ]
>> +static void qpace_remove(struct platform_device *pdev)
>> +{
>> +	mutex_lock(&qpace_ref_lock);
>> +	qpace_priv->suspended = true;
>> +	mutex_unlock(&qpace_ref_lock);
>> +
>> +	static_branch_disable(&qpace_drv_probed);
>> +
>> +	wait_for_completion(&qpace_priv->no_active_refs);
>> +
>> +	/* No callers remain; tear down the hardware. */
>> +	cancel_work_sync(&qpace_priv->disable_work);
>> +	qpace_power_off(&pdev->dev);
>> +}
> 
> [Severity: High]
> Does this leave the global qpace_priv pointer set? In qpace_remove(), the
> device is torn down and devres will free the memory backing qpace_priv, but
> the global pointer is never cleared, enabling the use-after-free mentioned
> above.

ACK. The pointer can be set to NULL to be safe and qpace_drv_probed can 
be checked by exported functions to avoid using qpace if it has been 
removed or not probed.

> 
> [ ... ]
>> +static int qpace_resume(struct device *dev)
>> +{
>> +	int ret;
>> +
>> +	mutex_lock(&qpace_ref_lock);
>> +
>> +	if (qpace_priv->active_rings)
>> +		dev_err(dev, "active_rings not 0, unexpected case\n");
>> +
>> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
>> +	if (ret)
>> +		goto out_unlock;
>> +
>> +	program_urg_command_contexts_v2();
>> +	program_decomp_core_cfg();
>> +
>> +	qpace_priv->suspended = false;
> 
> [Severity: High]
> Will the hardware be stuck in STOP mode after suspend/resume? In
> qpace_resume(), we restore contexts but fail to wake the hardware engine
> up to RUN mode (omitting the QPACE_RUN write) and fail to reprogram
> essential bulk mode and cache index settings from qpace_hw_init(). This
> means all hardware operations will indefinitely time out after the system
> resumes.
> 

In suspend case the HW is in power collapse mode, but once woken up will 
retain it's run status.

>> +
>> +out_unlock:
>> +	if (ret)
>> +		dev_err(dev, "failed to resume QPaCE: %d\n", ret);
>> +	mutex_unlock(&qpace_ref_lock);
>> +	return ret;
>> +}
> 

Thanks,
Oreoluwa


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-10-01  8:50   ` Krzysztof Kozlowski
@ 2026-10-06 23:51     ` Oreoluwa Babatunde
  2026-10-07  7:51       ` Krzysztof Kozlowski
  0 siblings, 1 reply; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-06 23:51 UTC (permalink / raw)
  To: Krzysztof Kozlowski, Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 10/1/2026 1:50 AM, Krzysztof Kozlowski wrote:
> On Wed, Sep 30, 2026 at 07:52:11AM -0700, Georgi Djakov wrote:
>> Add a platform driver for the Qualcomm Page Compression Engine (QPaCE), a
>> hardware block that accelerates compression and decompression of memory
>> pages.
>>
>> Provide the urgent command path for synchronous single-page compression and
>> decompression. This exposes the low-latency operations needed by
>> compressed-memory users such as zram, especially for page decompression on
>> the read path.
>>
>> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
>> ---
>>   drivers/soc/qcom/Kconfig          |  14 +
>>   drivers/soc/qcom/Makefile         |   1 +
>>   drivers/soc/qcom/qpace.c          | 764 ++++++++++++++++++++++++++++++
>>   drivers/soc/qcom/qpace_internal.h |  84 ++++
>>   include/linux/soc/qcom/qpace.h    | 154 ++++++
>>   5 files changed, 1017 insertions(+)
>>   create mode 100644 drivers/soc/qcom/qpace.c
>>   create mode 100644 drivers/soc/qcom/qpace_internal.h
>>   create mode 100644 include/linux/soc/qcom/qpace.h
>>
>> diff --git a/drivers/soc/qcom/Kconfig b/drivers/soc/qcom/Kconfig
>> index 535c8619197b..6bcb86dcd726 100644
>> --- a/drivers/soc/qcom/Kconfig
> 
> Sorry, but no. Soc is not a dumping ground. This has clear function of
> compression offload, so it should have some dedicated maintainers like
> other offload engines.

The reason for putting this in soc/qcom is because this is a qcom HW 
block driver. As per your comments below we will check and see if we can 
make use of existing crypto framework and respond back on this.

>> +++ b/drivers/soc/qcom/Kconfig
>> @@ -288,6 +288,20 @@ config QCOM_PBS
>>   	  This module provides the APIs to the client drivers that wants to send the
>>   	  PBS trigger event to the PBS RAM.
>>   
>> +config QCOM_PAGE_COMPRESSION_ENGINE
>> +	tristate "Qualcomm Page Compression Engine (QPaCE)"
>> +	depends on ARM64
> 
> Why this can't be built on other archs? This is really odd and I do not
> see any asm headers included.
ACK. We will remove this so that it can be built on other architectures.

> 
>> +	depends on ARCH_QCOM || COMPILE_TEST
>> +	depends on OF
>> +	depends on INTERCONNECT
>> +	help
>> +	  Enable support for the Qualcomm Page Compression Engine (QPaCE),
>> +	  a hardware accelerator that provides high-throughput page compression,
>> +	  decompression, and DMA copy operations.
>> +
>> +	  The engine is used as a hardware backend for compressed-memory
>> +	  subsystems such as zram. If unsure, say N.
>> +
> 
> ...
> 
> 
>> +	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
>> +out:
>> +	return ret;
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_urgent_compress);
>> +
>> +int qpace_urgent_decompress(dma_addr_t input_addr,
>> +			    dma_addr_t output_addr,
>> +			    size_t input_size,
>> +			    struct qpace_algorithm *algo)
> 
> You need kerneldoc for every export.

ACK

>> +{
>> +	int urg_reg_num;
>> +	int stat_reg;
>> +	u32 stat_reg_val;
>> +	int ret;
>> +
>> +	ret = qpace_get();
>> +	if (ret)
>> +		goto out;
>> +
>> +	urg_reg_num = get_cpu() % NUM_TRS_ERS_URG_CMD_REGS;
>> +	qpace_write_urg_cmd_ctx(qpace_priv, QPACE_URG_CMD_0_CFG_CNTXT_SIZE_n_OFFSET,
>> +				urg_reg_num, algo->urg_decomp_cntxt,
>> +				FIELD_PREP(URG_CMD_0_CFG_CNTXT_SIZE_SIZE, input_size));
>> +	stat_reg = qpace_urgent_command_trigger(input_addr, output_addr, urg_reg_num,
>> +						algo->urg_decomp_cntxt);
>> +	put_cpu();
>> +
>> +	qpace_put();
>> +
>> +	if (stat_reg < 0) {
>> +		ret = stat_reg;
>> +		goto out;
>> +	}
>> +
>> +	stat_reg_val = FIELD_GET(URG_CMD_0_ED_STAT_COMP_CODE, stat_reg);
>> +	if (stat_reg_val != OP_OK) {
>> +		pr_err("%s: register %d failed with %u\n",
>> +		       __func__, urg_reg_num, stat_reg_val);
>> +		ret = -EINVAL;
>> +		goto out;
>> +	}
>> +
>> +	ret = FIELD_GET(URG_CMD_0_ED_STAT_SIZE, stat_reg);
>> +out:
>> +	return ret;
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_urgent_decompress);
>> +
> 
> So singleton? For what reason exactly? Random drivers will be getting
> the reference to compress something? If so, aren't you duplicating
> existing infrastructure/API for in-kernel hardware offloaded
> compression (e.g. drivers/crypto/)?
> 

We will check and see if we can use existing crypto framework and 
respond back on this.

> You miss proper comments (see checkpatch --strict) explaining lock
> usage.

ACK

> 
> 
>> +static DEFINE_MUTEX(qpace_ref_lock);
>> +
>> +static void _get_qpace(void)
>> +{
>> +	lockdep_assert_held(&qpace_ref_lock);
>> +	if (!qpace_priv->active_rings) {
>> +		reinit_completion(&qpace_priv->no_active_refs);
>> +		pm_stay_awake(qpace_priv->dev);
>> +		cpu_latency_qos_update_request(&qpace_priv->qos_req, 300);
>> +		program_urg_command_contexts_v2();
>> +		program_decomp_core_cfg();
>> +	}
>> +	qpace_priv->active_rings++;
>> +}
>> +
>> +static void _put_qpace(void)
>> +{
>> +	lockdep_assert_held(&qpace_ref_lock);
>> +	if (!--qpace_priv->active_rings) {
>> +		cpu_latency_qos_update_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
>> +		pm_relax(qpace_priv->dev);
>> +		complete(&qpace_priv->no_active_refs);
>> +	}
>> +}
>> +
>> +int qpace_get(void)
>> +{
>> +	int ret = 0;
>> +
>> +	mutex_lock(&qpace_ref_lock);
>> +	if (qpace_priv->suspended || READ_ONCE(qpace_priv->broken))
>> +		ret = -EBUSY;
>> +	else
>> +		_get_qpace();
>> +	mutex_unlock(&qpace_ref_lock);
>> +	return ret;
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_get);
>> +
>> +void qpace_put(void)
>> +{
>> +	mutex_lock(&qpace_ref_lock);
>> +	_put_qpace();
>> +	mutex_unlock(&qpace_ref_lock);
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_put);
>> +
>> +static irqreturn_t urgent_interrupt_handler(int irq, void *unused)
>> +{
>> +	pr_debug("Urgent interrupt handled\n");
>> +	return IRQ_HANDLED;
>> +}
>> +
>> +static int qpace_hw_init(void)
>> +{
>> +	u32 reg_val;
>> +
>> +	/* Select CPU SCID for our system cache slice. */
>> +	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET);
>> +	reg_val = u32_replace_bits(reg_val, 0x1, CORE_QNS4_CFG_CACHEINDEX);
>> +	qpace_write_gen(qpace_priv, QPACE_CORE_QNS4_CFG_OFFSET, reg_val);
>> +
>> +	/* QMB2 register configurations. */
>> +	reg_val = qpace_read_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET);
>> +	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_RD_OUTST_LIMIT);
>> +	reg_val = u32_replace_bits(reg_val, 0x48, CORE_GEN_CFG_QMB2_MAX_WR_OUTST_LIMIT);
>> +	qpace_write_gen(qpace_priv, QPACE_CORE_GEN_CFG_OFFSET, reg_val);
>> +
>> +	/* DECOMP_CORE_CFG init steps. */
>> +	program_decomp_core_cfg();
>> +
>> +	/* Below settings help save power since all decomp cores are set to sync. */
>> +	reg_val = qpace_read_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET);
>> +	reg_val |= CORE_OPER_CFG_COMP_MEM_PWR_DWN_1;
>> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CFG_OFFSET, reg_val);
>> +
>> +	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET);
>> +	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_RD_MAX_OT);
>> +	reg_val = u32_replace_bits(reg_val, 0x8, COMP_CORE_CFG_DMA_WR_MAX_OT);
>> +	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_CFG_OFFSET, reg_val);
>> +
>> +	/* Set all COMP engines to bulk mode. */
>> +	reg_val = qpace_read_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET);
>> +	reg_val |= COMP_CORE_BULK_MODE_ALL_CORES;
>> +	qpace_write_comp_core(qpace_priv, QPACE_COMP_CORE_BULK_MODE_OFFSET, reg_val);
>> +
>> +	/* URG CMD register configurations. */
>> +	program_urg_command_contexts_v2();
>> +
>> +	return 0;
>> +}
>> +
>> +enum qpace_interrupts {
>> +	QPACE_IRQ_URGENT
>> +};
>> +
>> +static int qpace_register_interrupts(struct platform_device *pdev)
>> +{
>> +	struct device *dev = &pdev->dev;
>> +	int irq, ret;
>> +
>> +	irq = platform_get_irq(pdev, QPACE_IRQ_URGENT);
>> +	if (irq < 0)
>> +		return irq;
>> +
>> +	ret = devm_request_irq(dev, irq, urgent_interrupt_handler,
>> +			       0, "qpace-urgent-irq", NULL);
>> +	if (ret)
>> +		dev_err(dev, "failed to request urgent interrupt\n");
>> +
>> +	return ret;
>> +}
>> +
>> +static inline bool _qpace_power_on(void)
>> +{
>> +	u32 ready_status;
>> +
>> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_RUN);
>> +
>> +	if (readl_poll_timeout(qpace_priv->gen_core_regs +
>> +			       QPACE_CORE_OPER_CORE_READY_OFFSET,
>> +			       ready_status, ready_status,
>> +			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
>> +		pr_err("Timeout in waiting for QPaCE to turn on\n");
>> +		return false;
>> +	}
>> +
>> +	return true;
>> +}
>> +
>> +static int qpace_power_on(struct device *dev)
>> +{
>> +	int ret, ret2;
>> +
>> +	qpace_priv->interconnect = devm_of_icc_get(dev, "qpace-mem");
>> +	if (IS_ERR_OR_NULL(qpace_priv->interconnect)) {
>> +		ret = PTR_ERR_OR_ZERO(qpace_priv->interconnect);
>> +		pr_err("%s: devm_of_icc_get() failed with %d\n", __func__, ret);
> 
> use dev_err, not pr_err

ACK.

> 
>> +		return qpace_priv->interconnect ? ret : -EINVAL;
>> +	}
>> +
>> +	ret = device_init_wakeup(dev, true);
>> +	if (ret) {
>> +		pr_err("%s: device_init_wakeup() failed with %d\n", __func__, ret);
>> +		return ret;
>> +	}
>> +
>> +	cpu_latency_qos_add_request(&qpace_priv->qos_req, PM_QOS_DEFAULT_VALUE);
>> +
>> +	icc_set_tag(qpace_priv->interconnect, QCOM_ICC_TAG_ACTIVE_ONLY);
>> +
>> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 1);
>> +	if (ret) {
>> +		pr_err("Failed to turn on QPaCE VCD: %d\n", ret);
>> +		goto rm_qos;
>> +	}
>> +
>> +	if (!_qpace_power_on()) {
>> +		pr_err("Failed to start QPaCE\n");
>> +		ret = -EINVAL;
>> +		goto rm_bw;
>> +	}
>> +
>> +	return 0;
>> +
>> +rm_bw:
>> +	ret2 = icc_set_bw(qpace_priv->interconnect, 0, 0);
>> +	if (ret2)
>> +		pr_err("Failed to remove QPaCE VCD vote: %d\n", ret2);
>> +rm_qos:
>> +	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
>> +	device_init_wakeup(dev, false);
>> +
>> +	return ret;
>> +}
>> +
>> +static inline bool _qpace_power_off(void)
>> +{
>> +	u32 ready_status;
>> +
>> +	qpace_write_gen_core(qpace_priv, QPACE_CORE_OPER_CORE_RUN_STOP_OFFSET, QPACE_STOP);
>> +
>> +	if (readl_poll_timeout(qpace_priv->gen_core_regs +
>> +			       QPACE_CORE_OPER_CORE_READY_OFFSET,
>> +			       ready_status, !ready_status,
>> +			       1000, 5 * QPACE_STATE_CHANGE_TIMEOUT_US)) {
>> +		pr_err("Timeout in waiting for QPaCE to turn off\n");
>> +		return false;
>> +	}
>> +
>> +	return true;
>> +}
>> +
>> +static void qpace_power_off(struct device *dev)
>> +{
>> +	int ret;
>> +
>> +	/* If this fails we can still remove our vote for the VCD to turn QPaCE off */
>> +	if (!_qpace_power_off())
>> +		pr_err("Failed to stop QPaCE\n");
>> +
>> +	ret = icc_set_bw(qpace_priv->interconnect, 0, 0);
>> +	if (ret)
>> +		pr_err("Failed to turn off QPaCE VCD: %d\n", ret);
>> +
>> +	cpu_latency_qos_remove_request(&qpace_priv->qos_req);
>> +
>> +	device_init_wakeup(dev, false);
>> +}
>> +
>> +static inline int qpace_register_ioremap(struct platform_device *pdev)
>> +{
>> +	qpace_priv->gen_regs = devm_platform_ioremap_resource(pdev, 0);
>> +	if (IS_ERR(qpace_priv->gen_regs))
>> +		return PTR_ERR(qpace_priv->gen_regs);
>> +
>> +	qpace_priv->gen_core_regs = qpace_priv->gen_regs + QPACE_GEN_CORE_REGS_OFFSET;
>> +	qpace_priv->comp_core_regs = qpace_priv->gen_regs + QPACE_COMP_CORE_REGS_OFFSET;
>> +	qpace_priv->decomp_core_regs = qpace_priv->gen_regs + QPACE_DECOMP_CORE_REGS_OFFSET;
>> +	qpace_priv->urg_regs = qpace_priv->gen_regs + QPACE_URG_REGS_OFFSET;
>> +
>> +	return 0;
>> +}
>> +
>> +bool qpace_is_dev_available(void)
>> +{
>> +	return static_branch_likely(&qpace_drv_probed) &&
>> +	       !READ_ONCE(qpace_priv->broken);
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_is_dev_available);
>> +
>> +struct device *qpace_get_dma_dev(void)
>> +{
>> +	return qpace_priv->dev;
>> +}
>> +EXPORT_SYMBOL_GPL(qpace_get_dma_dev);
>> +
>> +static int qpace_probe(struct platform_device *pdev)
>> +{
>> +	struct device *dev = &pdev->dev;
>> +	struct qpace_priv *priv;
>> +	int ret;
>> +
>> +	priv = devm_kzalloc(dev, sizeof(*priv), GFP_KERNEL);
>> +	if (!priv)
>> +		return -ENOMEM;
>> +
>> +	priv->dev = dev;
>> +	/* Starts already complete since active_rings == 0 at init. */
>> +	init_completion(&priv->no_active_refs);
>> +	complete(&priv->no_active_refs);
>> +	INIT_WORK(&priv->disable_work, qpace_disable_work_fn);
>> +	qpace_priv = priv;
>> +	platform_set_drvdata(pdev, priv);
>> +
>> +	ret = dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
>> +	if (ret)
>> +		return dev_err_probe(dev, ret, "failed to set DMA mask\n");
>> +
>> +	ret = qpace_register_ioremap(pdev);
>> +	if (ret)
>> +		return dev_err_probe(dev, ret, "failed to map QPaCE registers\n");
>> +
>> +	ret = qpace_power_on(dev);
>> +	if (ret)
>> +		return dev_err_probe(dev, ret, "failed to power on QPaCE\n");
>> +
>> +	/* Get QPaCE HW version. */
>> +	qpace_priv->hw_version = qpace_read_gen(qpace_priv, QPACE_CORE_HW_VERSION_OFFSET);
>> +	if (qpace_priv->hw_version != QPACE_HW_VERSION_V2) {
>> +		dev_err(dev, "Unsupported QPaCE HW version returned: 0x%x\n",
>> +			qpace_priv->hw_version);
>> +		ret = -EINVAL;
>> +		goto power_off;
>> +	}
>> +
>> +	ret = qpace_hw_init();
>> +	if (ret) {
> 
> How is this possible?

ACK. This can be removed.

> 
>> +		dev_err(dev, "init failed: (%d)\n", ret);
>> +		goto power_off;
>> +	}
>> +
>> +	ret = qpace_register_interrupts(pdev);
>> +	if (ret) {
>> +		dev_err(dev, "failed to register interrupts\n");
> 
> Do not print same error multiple times.

ACK.

> 
>> +		goto power_off;
>> +	}
>> +
>> +	static_branch_enable(&qpace_drv_probed);
>> +
>> +	return ret;
>> +
>> +power_off:
>> +	qpace_power_off(dev);
>> +
>> +	return ret;
>> +}
> 
> Best regards,
> Krzysztof


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-10-01  5:43   ` Sergey Senozhatsky
@ 2026-10-06 23:53     ` Oreoluwa Babatunde
  2026-10-08  3:52       ` Sergey Senozhatsky
  0 siblings, 1 reply; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-06 23:53 UTC (permalink / raw)
  To: Sergey Senozhatsky, Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 9/30/2026 10:43 PM, Sergey Senozhatsky wrote:
> On (26/09/30 07:52), Georgi Djakov wrote:
> [..]
>> @@ -80,10 +87,10 @@ static const struct zcomp_ops *lookup_backend_ops(const char *comp)
>>   
>>   	while (backends[i]) {
>>   		if (sysfs_streq(comp, backends[i]->name))
>> -			break;
>> +			return backends[i];
>>   		i++;
>>   	}
>> -	return backends[i];
>> +	return NULL;
>>   }
> 
> This hunk looks unrelated to the patch in question.
ACK. Will remove this.

Thanks,
Oreoluwa

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-10-01  8:51   ` Krzysztof Kozlowski
  2026-10-01 10:27     ` Sergey Senozhatsky
@ 2026-10-07  0:03     ` Oreoluwa Babatunde
  1 sibling, 0 replies; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-07  0:03 UTC (permalink / raw)
  To: Krzysztof Kozlowski, Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 10/1/2026 1:51 AM, Krzysztof Kozlowski wrote:
> On Wed, Sep 30, 2026 at 07:52:13AM -0700, Georgi Djakov wrote:
>> Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
>> zram can expose qpace-lz4 as a selectable compression algorithm when the
>> QPaCE driver is available.
>>
>> Compress and decompress operations are handled via the QPaCE urgent
>> synchronous path: each request DMA-maps the source and destination
>> buffers, issues a blocking hardware command, and returns the result size.
>>
>> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
>> ---
>>   drivers/block/zram/Kconfig         |  11 +++
>>   drivers/block/zram/Makefile        |   1 +
>>   drivers/block/zram/backend_qpace.c | 134 +++++++++++++++++++++++++++++
> 
> So entire qpace should go here, no? Why did you create this entire layer
> of indirection, singleton management under drivers/soc?
> 

Since the qpace driver is not specific to zram and may be utilized by 
other kernel components, it should remain independent of drivers/zram.
Consolidating everything under drivers/zram would couple the 
implementation to a single use case and reduce its reusability.

The backend_qpace implementation was intended to setup zram as a user of 
qpace and follow the structure of the zcomp frameowork. As per your 
comment in the other patches, we will check if this can be done through 
the crypto framework.

>>   drivers/block/zram/backend_qpace.h |  13 +++
>>   drivers/block/zram/zcomp.c         |  11 ++-
>>   5 files changed, 168 insertions(+), 2 deletions(-)
>>   create mode 100644 drivers/block/zram/backend_qpace.c
>>   create mode 100644 drivers/block/zram/backend_qpace.h

Thanks,
Oreoluwa

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-10-01 10:27     ` Sergey Senozhatsky
@ 2026-10-07  0:04       ` Oreoluwa Babatunde
  0 siblings, 0 replies; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-07  0:04 UTC (permalink / raw)
  To: Sergey Senozhatsky, Krzysztof Kozlowski
  Cc: Georgi Djakov, andersson, konradybcio, abelvesa, robh, krzk+dt,
	conor+dt, minchan, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 10/1/2026 3:27 AM, Sergey Senozhatsky wrote:
> On (26/10/01 10:51), Krzysztof Kozlowski wrote:
>> On Wed, Sep 30, 2026 at 07:52:13AM -0700, Georgi Djakov wrote:
>>> Add a zcomp backend for the Qualcomm Page Compression Engine (QPaCE) so
>>> zram can expose qpace-lz4 as a selectable compression algorithm when the
>>> QPaCE driver is available.
>>>
>>> Compress and decompress operations are handled via the QPaCE urgent
>>> synchronous path: each request DMA-maps the source and destination
>>> buffers, issues a blocking hardware command, and returns the result size.
>>>
>>> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
>>> ---
>>>   drivers/block/zram/Kconfig         |  11 +++
>>>   drivers/block/zram/Makefile        |   1 +
>>>   drivers/block/zram/backend_qpace.c | 134 +++++++++++++++++++++++++++++
>>
>> So entire qpace should go here, no? Why did you create this entire layer
>> of indirection, singleton management under drivers/soc?
> 
> qpace seems to be standalone/independent arch/soc specific, and can gain
> users outside of zram (fs compression, zswap, etc.).  zram's backends
> potentially can disappear all together, if we switch to acomp crypto API.

ACK

Thanks,
Oreoluwa

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-10-06 23:51     ` Oreoluwa Babatunde
@ 2026-10-07  7:51       ` Krzysztof Kozlowski
  2026-10-08  0:08         ` Oreoluwa Babatunde
  0 siblings, 1 reply; 25+ messages in thread
From: Krzysztof Kozlowski @ 2026-10-07  7:51 UTC (permalink / raw)
  To: Oreoluwa Babatunde, Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 07/10/2026 01:51, Oreoluwa Babatunde wrote:
> On 10/1/2026 1:50 AM, Krzysztof Kozlowski wrote:
>> On Wed, Sep 30, 2026 at 07:52:11AM -0700, Georgi Djakov wrote:
>>> Add a platform driver for the Qualcomm Page Compression Engine (QPaCE), a
>>> hardware block that accelerates compression and decompression of memory
>>> pages.
>>>
>>> Provide the urgent command path for synchronous single-page compression and
>>> decompression. This exposes the low-latency operations needed by
>>> compressed-memory users such as zram, especially for page decompression on
>>> the read path.
>>>
>>> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
>>> ---
>>>   drivers/soc/qcom/Kconfig          |  14 +
>>>   drivers/soc/qcom/Makefile         |   1 +
>>>   drivers/soc/qcom/qpace.c          | 764 ++++++++++++++++++++++++++++++
>>>   drivers/soc/qcom/qpace_internal.h |  84 ++++
>>>   include/linux/soc/qcom/qpace.h    | 154 ++++++
>>>   5 files changed, 1017 insertions(+)
>>>   create mode 100644 drivers/soc/qcom/qpace.c
>>>   create mode 100644 drivers/soc/qcom/qpace_internal.h
>>>   create mode 100644 include/linux/soc/qcom/qpace.h
>>>
>>> diff --git a/drivers/soc/qcom/Kconfig b/drivers/soc/qcom/Kconfig
>>> index 535c8619197b..6bcb86dcd726 100644
>>> --- a/drivers/soc/qcom/Kconfig
>>
>> Sorry, but no. Soc is not a dumping ground. This has clear function of
>> compression offload, so it should have some dedicated maintainers like
>> other offload engines.
> 
> The reason for putting this in soc/qcom is because this is a qcom HW 

Every qcom HW driver is a qcom HW driver and they DO NOT go to drivers/soc.

Again: soc is not a dumping ground.

Best regards,
Krzysztof

^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver
  2026-10-07  7:51       ` Krzysztof Kozlowski
@ 2026-10-08  0:08         ` Oreoluwa Babatunde
  0 siblings, 0 replies; 25+ messages in thread
From: Oreoluwa Babatunde @ 2026-10-08  0:08 UTC (permalink / raw)
  To: Krzysztof Kozlowski, Georgi Djakov
  Cc: andersson, konradybcio, abelvesa, robh, krzk+dt, conor+dt,
	minchan, senozhatsky, axboe, rostedt, mhiramat, mathieu.desnoyers,
	linux-arm-msm, devicetree, linux-kernel, linux-block,
	linux-trace-kernel, djakov

On 10/7/2026 12:51 AM, Krzysztof Kozlowski wrote:
> On 07/10/2026 01:51, Oreoluwa Babatunde wrote:
>> On 10/1/2026 1:50 AM, Krzysztof Kozlowski wrote:
>>> On Wed, Sep 30, 2026 at 07:52:11AM -0700, Georgi Djakov wrote:
>>>> Add a platform driver for the Qualcomm Page Compression Engine (QPaCE), a
>>>> hardware block that accelerates compression and decompression of memory
>>>> pages.
>>>>
>>>> Provide the urgent command path for synchronous single-page compression and
>>>> decompression. This exposes the low-latency operations needed by
>>>> compressed-memory users such as zram, especially for page decompression on
>>>> the read path.
>>>>
>>>> Signed-off-by: Georgi Djakov <georgi.djakov@oss.qualcomm.com>
>>>> ---
>>>>    drivers/soc/qcom/Kconfig          |  14 +
>>>>    drivers/soc/qcom/Makefile         |   1 +
>>>>    drivers/soc/qcom/qpace.c          | 764 ++++++++++++++++++++++++++++++
>>>>    drivers/soc/qcom/qpace_internal.h |  84 ++++
>>>>    include/linux/soc/qcom/qpace.h    | 154 ++++++
>>>>    5 files changed, 1017 insertions(+)
>>>>    create mode 100644 drivers/soc/qcom/qpace.c
>>>>    create mode 100644 drivers/soc/qcom/qpace_internal.h
>>>>    create mode 100644 include/linux/soc/qcom/qpace.h
>>>>
>>>> diff --git a/drivers/soc/qcom/Kconfig b/drivers/soc/qcom/Kconfig
>>>> index 535c8619197b..6bcb86dcd726 100644
>>>> --- a/drivers/soc/qcom/Kconfig
>>>
>>> Sorry, but no. Soc is not a dumping ground. This has clear function of
>>> compression offload, so it should have some dedicated maintainers like
>>> other offload engines.
>>
>> The reason for putting this in soc/qcom is because this is a qcom HW
> 
> Every qcom HW driver is a qcom HW driver and they DO NOT go to drivers/soc.
> 
> Again: soc is not a dumping ground.
> 
ACK. We are looking at the crypto driver to see if it fits our use of 
the qpace HW block.

Do you have other suggestions for where this would better fit besides 
the crypto framework?

Thanks,
Oreoluwa


^ permalink raw reply	[flat|nested] 25+ messages in thread

* Re: [PATCH 4/6] zram: Add QPaCE zcomp backend
  2026-10-06 23:53     ` Oreoluwa Babatunde
@ 2026-10-08  3:52       ` Sergey Senozhatsky
  0 siblings, 0 replies; 25+ messages in thread
From: Sergey Senozhatsky @ 2026-10-08  3:52 UTC (permalink / raw)
  To: Oreoluwa Babatunde
  Cc: Sergey Senozhatsky, Georgi Djakov, andersson, konradybcio,
	abelvesa, robh, krzk+dt, conor+dt, minchan, axboe, rostedt,
	mhiramat, mathieu.desnoyers, linux-arm-msm, devicetree,
	linux-kernel, linux-block, linux-trace-kernel, djakov

On (26/10/06 16:53), Oreoluwa Babatunde wrote:
> On 9/30/2026 10:43 PM, Sergey Senozhatsky wrote:
> > On (26/09/30 07:52), Georgi Djakov wrote:
> > [..]
> > > @@ -80,10 +87,10 @@ static const struct zcomp_ops *lookup_backend_ops(const char *comp)
> > >   	while (backends[i]) {
> > >   		if (sysfs_streq(comp, backends[i]->name))
> > > -			break;
> > > +			return backends[i];
> > >   		i++;
> > >   	}
> > > -	return backends[i];
> > > +	return NULL;
> > >   }
> > 
> > This hunk looks unrelated to the patch in question.
> ACK. Will remove this.

Thanks, if you are going to add qpace zcomp backend, please note
that zcomp API is undergoing some architectural changes [1] (shouldn't
be too much extra work to adapt.)

[1] https://lore.kernel.org/linux-kernel/20261005122036.718976-1-senozhatsky@chromium.org

^ permalink raw reply	[flat|nested] 25+ messages in thread

end of thread, other threads:[~2026-10-08  3:52 UTC | newest]

Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 14:52 [PATCH 0/6] soc: qcom: Add Qualcomm Page Compression Engine (QPaCE) driver Georgi Djakov
2026-09-30 14:52 ` [PATCH 1/6] dt-bindings: soc: qcom: Add QPaCE binding Georgi Djakov
2026-10-02  6:10   ` Krzysztof Kozlowski
2026-09-30 14:52 ` [PATCH 2/6] soc: qcom: qpace: Add Qualcomm Page Compression Engine driver Georgi Djakov
2026-09-30 15:04   ` sashiko-bot
2026-10-03  0:06     ` Oreoluwa Babatunde
2026-10-01  8:50   ` Krzysztof Kozlowski
2026-10-06 23:51     ` Oreoluwa Babatunde
2026-10-07  7:51       ` Krzysztof Kozlowski
2026-10-08  0:08         ` Oreoluwa Babatunde
2026-09-30 14:52 ` [PATCH 3/6] trace: qpace: Add tracepoints for QPaCE operations Georgi Djakov
2026-09-30 15:00   ` sashiko-bot
2026-09-30 14:52 ` [PATCH 4/6] zram: Add QPaCE zcomp backend Georgi Djakov
2026-09-30 15:08   ` sashiko-bot
2026-10-01  5:43   ` Sergey Senozhatsky
2026-10-06 23:53     ` Oreoluwa Babatunde
2026-10-08  3:52       ` Sergey Senozhatsky
2026-10-01  8:51   ` Krzysztof Kozlowski
2026-10-01 10:27     ` Sergey Senozhatsky
2026-10-07  0:04       ` Oreoluwa Babatunde
2026-10-07  0:03     ` Oreoluwa Babatunde
2026-09-30 14:52 ` [PATCH 5/6] soc: qcom: qpace: Add LLCC slice support Georgi Djakov
2026-09-30 15:08   ` sashiko-bot
2026-09-30 14:52 ` [PATCH 6/6] arm64: dts: qcom: hawi: Add QPaCE DT node Georgi Djakov
2026-09-30 14:59   ` sashiko-bot

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.