Linux s390 Architecture development
 help / color / mirror / Atom feed
* [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver
@ 2026-10-08  8:08 Julian Ruess
  2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
                   ` (3 more replies)
  0 siblings, 4 replies; 10+ messages in thread
From: Julian Ruess @ 2026-10-08  8:08 UTC (permalink / raw)
  To: schnelle, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, julianr,
	linux390-list, linux-s390

Hi all,

this series is based on net-next.

It introduces the concept of dibs peer devices as devices on a virtual
bus, called the dibs peer bus, implemented on top of the dibs
communication layer. These devices are based on the concept that one
peer on a dibs abstracted fabric can present virtual devices for
consumption by other peers on the same fabric. In this series two such
device drivers, a console/tty driver and a block device driver are
included. While the console/tty allows two peers to establish a virtual
TTY between them the block device driver allows read-only access to
block data, such as the contents of an image file.

This is a prerequisite for an in-development, open-source user space
driver stack (console server) that uses the s390 ISM (Internal Shared
Memory) device to implement and announce dibs peer devices to multiple
remote target systems. These peer devices are currently of type console
and block device. In addition to the console device which enables
console access, the block device support enables a target system to
access resources such as ISO images provided by the console server. Both
the console and block device functionality operate without requiring any
network connectivity. The console server is planned to be part of
s390-tools.

Thank you,
Julian and Tobias

Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
---
Julian Ruess (3):
      dibs: introduce dibs peer device bus driver
      dibs: introduce dibs console peer device driver
      dibs: introduce dibs block device peer device driver

Tobias Schumacher (1):
      dibs: introduce ring-buffer send/receive API

 arch/s390/configs/debug_defconfig |    3 +
 arch/s390/configs/defconfig       |    3 +
 drivers/dibs/Kconfig              |   32 +
 drivers/dibs/Makefile             |    5 +-
 drivers/dibs/dibs_blk.c           |  798 ++++++++++++++++++++++++
 drivers/dibs/dibs_console.c       |  388 ++++++++++++
 drivers/dibs/dibs_main.c          |  500 ++++++++++++++-
 drivers/dibs/dibs_peer_bus_drv.c  | 1220 +++++++++++++++++++++++++++++++++++++
 include/linux/dibs.h              |  177 +++++-
 9 files changed, 3122 insertions(+), 4 deletions(-)
---
base-commit: 8df0638138d3e0344fd1fb36cf2d1ca1cf5028f0
change-id: 20260330-console_server-fc7a9ced6481

Best regards,
-- 
Julian Ruess <julianr@linux.ibm.com>


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 1/4] dibs: introduce ring-buffer send/receive API
  2026-10-08  8:08 [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver Julian Ruess
@ 2026-10-08  8:08 ` Julian Ruess
  2026-10-08  8:21   ` sashiko-bot
  2026-10-09 20:47   ` Niklas Schnelle
  2026-10-08  8:08 ` [PATCH 2/4] dibs: introduce dibs peer device bus driver Julian Ruess
                   ` (2 subsequent siblings)
  3 siblings, 2 replies; 10+ messages in thread
From: Julian Ruess @ 2026-10-08  8:08 UTC (permalink / raw)
  To: schnelle, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, julianr,
	linux390-list, linux-s390

From: Tobias Schumacher <ts@linux.ibm.com>

Introduce a centralized API for sending and receiving data between dibs
peers via ring buffers. With this API, drivers no longer need to
interact directly with DMBs and manage read/write pointers.

A driver must first register its ring buffer using
dibs_register_ring_buffer() and set the rdmb_tok via
dibs_ring_set_rdmb_tok(). After that, data can be sent using
dibs_ring_send().

Receiving messages is split into two steps: dibs_ring_recv() returns a
pointer to the next DIBS message if one is available. After the message
is processed, dibs_ring_ack() must be called to free up the message
buffers and to advance the read pointer. If no message is available,
dibs_ring_recv() returns NULL.

Co-developed-by: Julian Ruess <julianr@linux.ibm.com>
Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
Signed-off-by: Tobias Schumacher <ts@linux.ibm.com>
---
 drivers/dibs/dibs_main.c | 493 ++++++++++++++++++++++++++++++++++++++++++++++-
 include/linux/dibs.h     |  99 ++++++++++
 2 files changed, 590 insertions(+), 2 deletions(-)

diff --git a/drivers/dibs/dibs_main.c b/drivers/dibs/dibs_main.c
index 20c50997a7cf..e3a33bc42fe8 100644
--- a/drivers/dibs/dibs_main.c
+++ b/drivers/dibs/dibs_main.c
@@ -13,6 +13,9 @@
 #include <linux/slab.h>
 #include <linux/err.h>
 #include <linux/dibs.h>
+#include <linux/circ_buf.h>
+#include <linux/log2.h>
+#include <asm/barrier.h>
 
 #include "dibs_loopback.h"
 
@@ -20,7 +23,7 @@ MODULE_DESCRIPTION("Direct Internal Buffer Sharing class");
 MODULE_LICENSE("GPL");
 
 static const struct class dibs_class = {
-	.name		= "dibs",
+	.name = "dibs",
 };
 
 /* use an array rather a list for fast mapping: */
@@ -93,7 +96,8 @@ int dibs_unregister_client(struct dibs_client *client)
 		max_dmbs = dibs->ops->max_dmbs();
 		for (int i = 0; i < max_dmbs; ++i) {
 			if (dibs->dmb_clientid_arr[i] == client->id) {
-				WARN(1, "%s: attempt to unregister '%s' with registered dmb(s)\n",
+				WARN(1,
+				     "%s: attempt to unregister '%s' with registered dmb(s)\n",
 				     __func__, client->name);
 				rc = -EBUSY;
 				goto err_reg_dmb;
@@ -245,6 +249,491 @@ void dibs_dev_del(struct dibs_dev *dibs)
 }
 EXPORT_SYMBOL_GPL(dibs_dev_del);
 
+/**
+ * dibs_ring_peer_meta() - take a snapshot of the peer's ring metadata
+ * @rb: pointer to the ring buffer
+ * @peer: snapshot to fill in
+ *
+ * The peer writes its metadata dmb at any time, so read each field once and
+ * copy to the snapshot.
+ */
+static void dibs_ring_peer_meta(struct dibs_ring_buffer *rb,
+				struct dibs_ring_meta *peer)
+{
+	peer->meta_dmb_tok = READ_ONCE(rb->rd_ring_meta->meta_dmb_tok);
+	peer->buff_dmb_tok = READ_ONCE(rb->rd_ring_meta->buff_dmb_tok);
+	peer->size = READ_ONCE(rb->rd_ring_meta->size);
+	peer->head = READ_ONCE(rb->rd_ring_meta->head);
+	peer->tail = READ_ONCE(rb->rd_ring_meta->tail);
+	peer->reserved = 0;
+}
+
+static int dibs_ring_send_hdr(struct dibs_ring_buffer *rb, u64 meta_dmb_tok)
+{
+	int res;
+
+	if (!meta_dmb_tok)
+		return 0;
+
+	res = rb->dibs->ops->move_data(rb->dibs, meta_dmb_tok, 0, true, 0,
+				       &rb->wr_ring_meta,
+				       sizeof(struct dibs_ring_meta));
+
+	if (!res)
+		rb->hdr_sent = true;
+
+	return res;
+}
+
+/**
+ * dibs_ring_register() - register a DIBS ring buffer
+ * @name: name of the ring buffer. This will be used in debug output to
+ *        support identifying the ring buffer.
+ * @rb: pointer to the ring buffer struct. The fields in the struct don't have
+ *      to be initialized.
+ * @client: pointer to the DIBS client
+ * @dibs: pointer to the DIBS device
+ * @rgid: the remote GID that will be allowed to write into this ring buffer
+ * @size: size of the ring buffer in bytes
+ *
+ * Register and initialize a DIBS ring buffer. Allocates the metadata DMB and
+ * the payload DMB, initializes the fields in dibs_ring_buffer.
+ *
+ * Return: 0 in case of success, an error code in case of failure
+ */
+int dibs_ring_register(char *name, struct dibs_ring_buffer *rb,
+		       struct dibs_client *client, struct dibs_dev *dibs,
+		       uuid_t rgid, u32 size)
+{
+	int ret;
+
+	if (!size || !is_power_of_2(size))
+		return -EINVAL;
+
+	strscpy(rb->name, name, sizeof(rb->name));
+	rb->dibs = dibs;
+	rb->dmb.rgid = rgid;
+
+	rb->rd_meta_dmb.dmb_len = ALIGN(sizeof(struct dibs_ring_meta), 4096);
+	rb->rd_meta_dmb.rgid = rgid;
+	ret = dibs->ops->register_dmb(dibs, &rb->rd_meta_dmb, client);
+	if (ret)
+		return ret;
+
+	rb->dmb.dmb_len = ALIGN(size, 4096);
+	ret = dibs->ops->register_dmb(dibs, &rb->dmb, client);
+	if (ret) {
+		dibs->ops->unregister_dmb(rb->dibs, &rb->rd_meta_dmb);
+		return ret;
+	}
+
+	rb->dmb_registered = 1;
+
+	rb->wr_ring_meta.meta_dmb_tok = rb->rd_meta_dmb.dmb_tok;
+	rb->wr_ring_meta.buff_dmb_tok = rb->dmb.dmb_tok;
+	rb->wr_ring_meta.size = size;
+	rb->wr_ring_meta.head = 0;
+	rb->wr_ring_meta.tail = 0;
+	rb->wr_ring_meta.reserved = 0;
+
+	rb->rd_ring_meta = rb->rd_meta_dmb.cpu_addr;
+	rb->peer_buff_dmb_tok = 0;
+	rb->msg_size = 0;
+	rb->hdr_sent = false;
+
+	return ret;
+}
+EXPORT_SYMBOL_GPL(dibs_ring_register);
+
+/**
+ * dibs_ring_unregister() - unregister a DIBS ring buffer
+ * @rb: pointer to the ring buffer to unregister
+ *
+ * Free the resources allocated by dibs_ring_register().
+ *
+ * Return: 0 in case of success, an error code in case of failure
+ */
+int dibs_ring_unregister(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+	int ret;
+
+	rb->wr_ring_meta.meta_dmb_tok = 0;
+	rb->wr_ring_meta.buff_dmb_tok = 0;
+	rb->wr_ring_meta.size = 0;
+	rb->wr_ring_meta.head = 0;
+	rb->wr_ring_meta.tail = 0;
+
+	dibs_ring_peer_meta(rb, &peer);
+	if (peer.meta_dmb_tok) {
+		ret = dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
+		if (ret)
+			pr_warn("%s: failed to send header update: %d\n",
+				__func__, ret);
+	}
+
+	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->dmb);
+	if (ret)
+		pr_warn("%s: failed to unregister payload dmb: %d\n", __func__,
+			ret);
+
+	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->rd_meta_dmb);
+	if (ret)
+		pr_warn("%s: failed to unregister metadata dmb: %d\n", __func__,
+			ret);
+
+	rb->dmb_registered = 0;
+	return ret;
+}
+EXPORT_SYMBOL_GPL(dibs_ring_unregister);
+
+/**
+ * dibs_ring_peer_replaced() - check if the peer has replaced its ring buffer
+ * @rb: pointer to the ring buffer
+ * @peer: snapshot of the peer's metadata
+ *
+ * The fabric hands out a new dmb token for every registration, so a buffer
+ * token that differs from the one this side is synchronized to means the peer
+ * ring this side writes into is a different one.
+ *
+ * Return: true if the peer ring has been replaced, false otherwise
+ */
+static bool dibs_ring_peer_replaced(struct dibs_ring_buffer *rb,
+				    const struct dibs_ring_meta *peer)
+{
+	return peer->buff_dmb_tok != rb->peer_buff_dmb_tok;
+}
+
+/**
+ * dibs_ring_sync_peer() - synchronize with a replaced peer ring buffer
+ * @rb: pointer to the ring buffer
+ * @peer: snapshot of the peer's metadata
+ *
+ * Restart both pointers: head is this side's write position in the peer's
+ * buffer, which is a new and empty one, and tail is this side's read position
+ * in its own buffer, which the peer has started to write at offset 0 again.
+ * Sends the updated metadata header to the peer and only then remembers the
+ * new token, so that a failed send is retried on the next call.
+ *
+ * A peer that has torn its ring down is not synchronized to, it is only
+ * forgotten - this side cannot send anything until the peer registers again.
+ */
+static void dibs_ring_sync_peer(struct dibs_ring_buffer *rb,
+				const struct dibs_ring_meta *peer)
+{
+	if (!peer->buff_dmb_tok) {
+		rb->wr_ring_meta.tail = 0;
+		rb->peer_buff_dmb_tok = 0;
+		rb->hdr_sent = false;
+		return;
+	}
+
+	rb->wr_ring_meta.head = 0;
+	rb->wr_ring_meta.tail = 0;
+
+	if (dibs_ring_send_hdr(rb, peer->meta_dmb_tok)) {
+		pr_warn("%s(%s): failed to send header after peer restart\n",
+			__func__, rb->name);
+		return;
+	}
+
+	rb->peer_buff_dmb_tok = peer->buff_dmb_tok;
+}
+
+/**
+ * dibs_ring_peer_ready() - check if the peer ring can be written to
+ * @peer: snapshot of the peer's metadata
+ *
+ * All of the peer's metadata is remote controlled. Its size is used as a mask
+ * for the head pointer and by the CIRC_* helpers, both of which require a
+ * power of two, so validate it before anything is derived from it.
+ *
+ * Return: true if the peer ring can be used, false otherwise
+ */
+static bool dibs_ring_peer_ready(const struct dibs_ring_meta *peer)
+{
+	if (!peer->meta_dmb_tok || !peer->buff_dmb_tok)
+		return false;
+
+	return is_power_of_2(peer->size);
+}
+
+static size_t __dibs_ring_space(struct dibs_ring_buffer *rb,
+				const struct dibs_ring_meta *peer)
+{
+	if (!dibs_ring_peer_ready(peer))
+		return 0;
+
+	return CIRC_SPACE(rb->wr_ring_meta.head, peer->tail, peer->size);
+}
+
+static size_t __dibs_ring_space_to_end(struct dibs_ring_buffer *rb,
+				       const struct dibs_ring_meta *peer)
+{
+	if (!dibs_ring_peer_ready(peer))
+		return 0;
+
+	return CIRC_SPACE_TO_END(rb->wr_ring_meta.head, peer->tail, peer->size);
+}
+
+static u32 __dibs_ring_cnt(struct dibs_ring_buffer *rb,
+			   const struct dibs_ring_meta *peer)
+{
+	return CIRC_CNT(peer->head, rb->wr_ring_meta.tail,
+			rb->wr_ring_meta.size);
+}
+
+static u32 __dibs_ring_cnt_to_end(struct dibs_ring_buffer *rb,
+				  const struct dibs_ring_meta *peer)
+{
+	return CIRC_CNT_TO_END(peer->head, rb->wr_ring_meta.tail,
+			       rb->wr_ring_meta.size);
+}
+
+/**
+ * dibs_ring_space() - query the ring buffer's free space for sending
+ * @rb: pointer to the ring buffer
+ *
+ * Get the free space in the ring buffer that can be used for sending data.
+ * Automatically resynchronizes if the peer has replaced its ring.
+ *
+ * Return: the free space in the ring buffer
+ */
+size_t dibs_ring_space(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (dibs_ring_peer_replaced(rb, &peer))
+		dibs_ring_sync_peer(rb, &peer);
+
+	return __dibs_ring_space(rb, &peer);
+}
+
+/**
+ * dibs_ring_space_to_end() - query the rb's space until the end of the DMB
+ * @rb: pointer to the ring buffer
+ *
+ * Get the free space in the ring buffer until the end of the underlying DMB.
+ * This is the amount of data that can be written without wrapping to the
+ * beginning of the DMB.
+ *
+ * Return: the free space in the ring buffer until the end of the DMB.
+ */
+size_t dibs_ring_space_to_end(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	return __dibs_ring_space_to_end(rb, &peer);
+}
+
+/**
+ * dibs_ring_cnt() - query the ring buffer's fill state
+ * @rb: pointer to the ring buffer
+ *
+ * Get the number of bytes available in the ring buffer for reading.
+ * Automatically resynchronizes if the peer has replaced its ring.
+ *
+ * Return: fill state of the ring buffer
+ */
+u32 dibs_ring_cnt(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (dibs_ring_peer_replaced(rb, &peer))
+		dibs_ring_sync_peer(rb, &peer);
+
+	return __dibs_ring_cnt(rb, &peer);
+}
+
+/**
+ * dibs_ring_cnt_to_end() - query the number of bytes in rb before wrap
+ * @rb: pointer to the ring buffer
+ *
+ * Get the number of bytes that can be read from the ring buffer without
+ * wrapping the tail pointer.
+ *
+ * Return: number of consecutive readable bytes
+ */
+u32 dibs_ring_cnt_to_end(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	return __dibs_ring_cnt_to_end(rb, &peer);
+}
+
+/**
+ * dibs_ring_set_rdmb_tok() - complete ring buffer token exchange
+ * @rb: pointer to the ring buffer
+ * @rdmb_tok: remote peer's metadata DMB token
+ *
+ * Completes the ring buffer setup by storing the remote peer's metadata DMB
+ * token and sending our local metadata to the peer. Normally, DMB tokens are
+ * read from rd_meta_dmb, but this creates a circular dependency during initial
+ * setup (need the token to read the metadata that contains the token). To break
+ * this, one peer must receive the remote token through an alternative channel
+ * and call this function to send its own header as the initial handshake step.
+ * After this, both peers can read each other's metadata.
+ */
+void dibs_ring_set_rdmb_tok(struct dibs_ring_buffer *rb, u64 rdmb_tok)
+{
+	WRITE_ONCE(rb->rd_ring_meta->meta_dmb_tok, rdmb_tok);
+	dibs_ring_send_hdr(rb, rdmb_tok);
+}
+EXPORT_SYMBOL_GPL(dibs_ring_set_rdmb_tok);
+
+/**
+ * dibs_ring_send() - Send DIBS message to peer
+ * @rb: Ring buffer to send message to
+ * @msg: Message to send
+ * @size: size of the message to be sent
+ * @notify: if true, update the remote's ring header and generate an IRQ after
+ *          sending.
+ *
+ * Send a DIBS message to the peer device using the specified ring buffer.
+ *
+ * Return: zero in case of success, an error code in case of failure
+ */
+int dibs_ring_send(struct dibs_ring_buffer *rb, void *msg, u16 size,
+		   bool notify)
+{
+	struct dibs_ring_meta peer;
+	u32 old_head;
+	int ret;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (!dibs_ring_peer_ready(&peer))
+		return -EAGAIN;
+
+	if (__dibs_ring_space_to_end(rb, &peer) < size)
+		return -ENOSPC;
+
+	ret = rb->dibs->ops->move_data(rb->dibs, peer.buff_dmb_tok, 0, false,
+				       rb->wr_ring_meta.head, msg, size);
+	if (ret)
+		return ret;
+
+	old_head = rb->wr_ring_meta.head;
+	rb->wr_ring_meta.head = (old_head + size) & (peer.size - 1);
+
+	if (notify) {
+		ret = dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
+		if (ret) {
+			rb->wr_ring_meta.head = old_head;
+			return ret;
+		}
+	}
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(dibs_ring_send);
+
+/**
+ * dibs_ring_send_padding() - send padding data to ring buffer
+ * @rb: pointer to the ring buffer
+ * @size: number of bytes to send
+ *
+ * Advance the ring buffer's head pointer by size byte without actually
+ * writing any data into the data buffer. This can be used by higher level
+ * APIs to send padding data which shall be ignored by the receiving peer.
+ *
+ * Return: zero in case of success, an error code in case of failure
+ */
+int dibs_ring_send_padding(struct dibs_ring_buffer *rb, u16 size)
+{
+	struct dibs_ring_meta peer;
+	u32 old_head;
+	int ret;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (!dibs_ring_peer_ready(&peer))
+		return -EAGAIN;
+
+	if (__dibs_ring_space(rb, &peer) < size)
+		return -ENOSPC;
+
+	old_head = rb->wr_ring_meta.head;
+	rb->wr_ring_meta.head = (old_head + size) & (peer.size - 1);
+
+	ret = dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
+	if (ret) {
+		rb->wr_ring_meta.head = old_head;
+		return ret;
+	}
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(dibs_ring_send_padding);
+
+/**
+ * dibs_ring_recv() - Receive message from DIBS ring buffer
+ * @rb: Ring buffer to get message from
+ *
+ * Get the next message from the ring buffer. After processing, the caller must
+ * call dibs_ring_ack() to update the ring buffer.
+ *
+ * Return: pointer to the DIBS message or NULL if the ring buffer is empty
+ */
+void *dibs_ring_recv(struct dibs_ring_buffer *rb)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (unlikely(!rb->hdr_sent))
+		dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
+
+	if (dibs_ring_peer_replaced(rb, &peer))
+		dibs_ring_sync_peer(rb, &peer);
+
+	if (!__dibs_ring_cnt(rb, &peer))
+		return NULL;
+
+	/* pairs with the peer filling the payload before publishing head */
+	dma_rmb();
+
+	return rb->dmb.cpu_addr + rb->wr_ring_meta.tail;
+}
+EXPORT_SYMBOL_GPL(dibs_ring_recv);
+
+/**
+ * dibs_ring_ack() - Acknowledge received message
+ * @rb: Ring buffer to operate on
+ *
+ * Acknowledges the message received by dibs_ring_recv(). This means, the read
+ * pointer in the ring buffer header is incremented and the updated header is
+ * sent to the remote side. After calling dibs_ring_ack(), the message structure
+ * returned by dibs_ring_recv() may be reused, so the driver must not use it
+ * anymore.
+ */
+int dibs_ring_ack(struct dibs_ring_buffer *rb, size_t size)
+{
+	struct dibs_ring_meta peer;
+
+	dibs_ring_peer_meta(rb, &peer);
+
+	if (size > __dibs_ring_cnt(rb, &peer)) {
+		pr_warn("%s(%s, %lx): error: tail would overtake head\n",
+			__func__, rb->name, size);
+		return -EINVAL;
+	}
+
+	rb->wr_ring_meta.tail = (rb->wr_ring_meta.tail + size) &
+				(rb->wr_ring_meta.size - 1);
+
+	return dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
+}
+EXPORT_SYMBOL_GPL(dibs_ring_ack);
+
 static int __init dibs_init(void)
 {
 	int rc;
diff --git a/include/linux/dibs.h b/include/linux/dibs.h
index 0c10c224bcca..30ae03484b58 100644
--- a/include/linux/dibs.h
+++ b/include/linux/dibs.h
@@ -67,6 +67,83 @@ struct dibs_dmb {
 	dma_addr_t dma_addr;
 };
 
+/* Ring Buffer Communication
+ * -------------------------
+ * Based on the DMBs, DIBS provides a ring buffer mechanism for bi-directional
+ * communication between two peers. For each ring-buffer connection, two DMBs
+ * are allocated per peer:
+ * - one DMB holding metadata about the ring and
+ * - one DMB holding the actual data to be transferred.
+ *
+ * Metadata consists of parameters like the DMB tokens, size, head pointer and
+ * tail pointer. Note that the metadata stored in the locally allocated
+ * metadata DMB does not necessarily corelate to the locally allocated payload
+ * DMB. As illustrated below, each peer has access to one wr_ring_meta DMB
+ * (the one the other peer has allocated) and one rd_ring_meta DMB (the one
+ * it has allocated itself). This is because the locally allocated DMB can only
+ * be written by the other peer and vice versa.
+ *
+ * Peer A                   Peer B
+ * ---------------          ---------------
+ * wr_ring_meta ----------> rd_ring_meta
+ * rd_ring_meta <---------- wr_ring_meta
+ * send_buffer  ----------> recv_buffer
+ * recv_buffer  <---------- send_buffer
+ *
+ * Consider for example that peer A wants to send data to peer B. It first has
+ * to check how much data it previously has written (head pointer) and how far
+ * peer B has read (tail pointer) and compare this to the size. Since for send
+ * operations head is updated by peer A, it is located in the wr_ring_meta DMB.
+ * Tail is updated by the receiving side, peer B in this case, so peer A will
+ * find the current value in rd_ring_meta. The payload DMB was allocated by
+ * peer B, so peer B also updated the size - peer A therefore finds the correct
+ * value in rd_ring_meta.
+ *
+ * Now peer a can write data into the send_buffer and update the head pointer
+ * in wr_ring_meta. Peer B then reads this head pointer from its rd_ring_meta
+ * DMB, process the data and update tail in its wr_ring_meta.
+ */
+struct dibs_ring_meta {
+	/* meta_dmb_tok - Token for the metadata dmb. */
+	u64 meta_dmb_tok;
+	/* buff_dmb_tok - Token for the buffer dmb */
+	u64 buff_dmb_tok;
+	/* size - size of the buffer for receiving data in number of byte */
+	u32 size;
+	/* head - points to the head of the ring buffer, i.e., the element
+	 * that will be written next
+	 */
+	u32 head;
+	/* tail - points to the end of the ring-buffer, i.e., the element
+	 * that will be read next
+	 */
+	u32 tail;
+	/* reserved - keeps the struct at its natural size, must be zero */
+	u32 reserved;
+};
+
+struct dibs_ring_buffer {
+	char name[256];
+	struct dibs_dev *dibs;
+	int dmb_registered;
+	bool hdr_sent;
+	struct dibs_dmb dmb;
+	/* dmb for the ring metadata this side can read, peer can write */
+	struct dibs_dmb rd_meta_dmb;
+	/* ring metadata this side can write, peer can read */
+	struct dibs_ring_meta wr_ring_meta;
+	/* ring metadata this side can read, peer can write */
+	struct dibs_ring_meta *rd_ring_meta;
+	/* size validated by dibs_pbd_msg_recv(), acked by dibs_pbd_msg_ack() */
+	u32 msg_size;
+	/* buff_dmb_tok of the peer ring this side is synchronized to. The
+	 * fabric hands out a new token for every registration, so a different
+	 * token means the peer has replaced its ring and the pointers into it
+	 * are stale.
+	 */
+	u64 peer_buff_dmb_tok;
+};
+
 /* DIBS events
  * -----------
  * Dibs devices can optionally notify dibs clients about events that happened
@@ -432,6 +509,22 @@ static inline void *dibs_get_priv(struct dibs_dev *dev,
 	return dev->priv[client->id];
 }
 
+int dibs_ring_register(char *name, struct dibs_ring_buffer *rb,
+		       struct dibs_client *client,
+		       struct dibs_dev *dibs,
+		       uuid_t rgid,
+		       u32 size);
+int dibs_ring_unregister(struct dibs_ring_buffer *rb);
+size_t dibs_ring_space(struct dibs_ring_buffer *rb);
+size_t dibs_ring_space_to_end(struct dibs_ring_buffer *rb);
+u32 dibs_ring_cnt(struct dibs_ring_buffer *rb);
+u32 dibs_ring_cnt_to_end(struct dibs_ring_buffer *rb);
+void dibs_ring_set_rdmb_tok(struct dibs_ring_buffer *rb, u64 rdmb_tok);
+int dibs_ring_send(struct dibs_ring_buffer *rb, void *msg, u16 size, bool notify);
+int dibs_ring_send_padding(struct dibs_ring_buffer *rb, u16 size);
+void *dibs_ring_recv(struct dibs_ring_buffer *rb);
+int dibs_ring_ack(struct dibs_ring_buffer *rb, size_t size);
+
 /* ------- End of client-only functions ----------- */
 
 /* Functions to be called by dibs device drivers:
@@ -461,4 +554,10 @@ int dibs_dev_add(struct dibs_dev *dibs);
  */
 void dibs_dev_del(struct dibs_dev *dibs);
 
+struct dibs_msg_hdr {
+	u8 version;
+	u8 type;
+	u16 datalen;
+} __packed;
+
 #endif	/* _DIBS_H */

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 2/4] dibs: introduce dibs peer device bus driver
  2026-10-08  8:08 [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver Julian Ruess
  2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
@ 2026-10-08  8:08 ` Julian Ruess
  2026-10-08  8:27   ` sashiko-bot
  2026-10-08  8:08 ` [PATCH 3/4] dibs: introduce dibs console peer device driver Julian Ruess
  2026-10-08  8:08 ` [PATCH 4/4] dibs: introduce dibs block device " Julian Ruess
  3 siblings, 1 reply; 10+ messages in thread
From: Julian Ruess @ 2026-10-08  8:08 UTC (permalink / raw)
  To: schnelle, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, julianr,
	linux390-list, linux-s390

Introduce a dibs peer device bus driver that represents remotely
emulated dibs peer devices as kernel devices. The driver handles
discovery and lifecycle management of remote peer devices and provides
the foundation for dibs peer device drivers that operate on top of dibs.

Two such peer device drivers, hvc_dibs and dibsblk, are added in
follow-up patches.

Dibs peer devices are provided by vfio user space tooling (Console
Server) and are discovered by the kernel of a target system.

Discovery works by setting e.g. the kernel parameter
'dibs.dibs_peer_address=gid:64000da3-78e8-3931-0000-000000000000'. The
UUID must be the GID used by the console server of the remote system.

Disovery can also be triggered manually by writing the GID of the remote
system to the discover attribute like e.g. 'echo
gid:04000da5-78e8-3931-0000-000000000000 > /sys/bus/dibs_peer/discover'.

Connections can be terminated by unbinding an dibs peer device from its
driver.

Co-developed-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
---
 arch/s390/configs/debug_defconfig |    1 +
 arch/s390/configs/defconfig       |    1 +
 drivers/dibs/Kconfig              |   11 +
 drivers/dibs/Makefile             |    3 +-
 drivers/dibs/dibs_main.c          |    7 +
 drivers/dibs/dibs_peer_bus_drv.c  | 1220 +++++++++++++++++++++++++++++++++++++
 include/linux/dibs.h              |   72 ++-
 7 files changed, 1313 insertions(+), 2 deletions(-)

diff --git a/arch/s390/configs/debug_defconfig b/arch/s390/configs/debug_defconfig
index 3dae71474333..23d6c59189cc 100644
--- a/arch/s390/configs/debug_defconfig
+++ b/arch/s390/configs/debug_defconfig
@@ -133,6 +133,7 @@ CONFIG_SMC=m
 CONFIG_SMC_DIAG=m
 CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
+CONFIG_DIBS_PEER_BUS_DRV=y
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/arch/s390/configs/defconfig b/arch/s390/configs/defconfig
index 6f5722634b4d..3bffcd016be8 100644
--- a/arch/s390/configs/defconfig
+++ b/arch/s390/configs/defconfig
@@ -124,6 +124,7 @@ CONFIG_SMC=m
 CONFIG_SMC_DIAG=m
 CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
+CONFIG_DIBS_PEER_BUS_DRV=y
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/drivers/dibs/Kconfig b/drivers/dibs/Kconfig
index 8f41e5622100..98068bc541e5 100644
--- a/drivers/dibs/Kconfig
+++ b/drivers/dibs/Kconfig
@@ -21,3 +21,14 @@ config DIBS_LO
 	  occurs within the same OS. This helps in convenient testing of
 	  dibs clients, since dibs loopback is independent of architecture or
 	  hardware.
+
+config DIBS_PEER_BUS_DRV
+	bool "dibs peer bus driver support"
+	depends on DIBS
+	default n
+	help
+	  This dibs peer bus driver allows to enumerate and use virtual dibs
+	  peer devices emulated on a dibs remote system. This bus driver is the
+	  foundation for dibs peer device drivers that operate on top of dibs.
+	  It provides device discovery and lifecycle management for peer
+	  devices across the dibs fabric.
diff --git a/drivers/dibs/Makefile b/drivers/dibs/Makefile
index 85805490c77f..b1418f5149c2 100644
--- a/drivers/dibs/Makefile
+++ b/drivers/dibs/Makefile
@@ -5,4 +5,5 @@
 
 dibs-y += dibs_main.o
 obj-$(CONFIG_DIBS) += dibs.o
-dibs-$(CONFIG_DIBS_LO) += dibs_loopback.o
\ No newline at end of file
+dibs-$(CONFIG_DIBS_LO) += dibs_loopback.o
+dibs-$(CONFIG_DIBS_PEER_BUS_DRV) += dibs_peer_bus_drv.o
diff --git a/drivers/dibs/dibs_main.c b/drivers/dibs/dibs_main.c
index e3a33bc42fe8..8fa05971912f 100644
--- a/drivers/dibs/dibs_main.c
+++ b/drivers/dibs/dibs_main.c
@@ -746,8 +746,14 @@ static int __init dibs_init(void)
 	if (rc)
 		goto err_unregister;
 
+	rc = dibs_pbd_init();
+	if (rc)
+		goto err_loopback_exit;
+
 	return rc;
 
+err_loopback_exit:
+	dibs_loopback_exit();
 err_unregister:
 	class_unregister(&dibs_class);
 err:
@@ -757,6 +763,7 @@ static int __init dibs_init(void)
 
 static void __exit dibs_exit(void)
 {
+	dibs_pbd_exit();
 	dibs_loopback_exit();
 	class_unregister(&dibs_class);
 }
diff --git a/drivers/dibs/dibs_peer_bus_drv.c b/drivers/dibs/dibs_peer_bus_drv.c
new file mode 100644
index 000000000000..99eb0a6f8164
--- /dev/null
+++ b/drivers/dibs/dibs_peer_bus_drv.c
@@ -0,0 +1,1220 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * dibs peer bus driver
+ *
+ * Copyright IBM Corp.
+ */
+
+#include <linux/device.h>
+#include <linux/slab.h>
+#include <linux/dibs.h>
+#include <linux/gfp_types.h>
+#include <linux/kref.h>
+#include <linux/mutex.h>
+#include <linux/string.h>
+#include <linux/workqueue.h>
+#include <linux/device/bus.h>
+#include <linux/export.h>
+#include <linux/init.h>
+#include <linux/kernel.h>
+#include <linux/module.h>
+#include <linux/printk.h>
+#include <linux/spinlock.h>
+#include <linux/types.h>
+#include <linux/uuid.h>
+#include <linux/utsname.h>
+#include <asm/sysinfo.h>
+#include <asm/ebcdic.h>
+
+static u8 dibs_pbd_boot_event_sent;
+
+static char *dibs_peer_address;
+module_param(dibs_peer_address, charp, 0000);
+MODULE_PARM_DESC(dibs_peer_address,
+		 "specify the address of the remote system that emulates the peer bus");
+
+static unsigned int dibs_pbd_max_timeout = 60;
+module_param(dibs_pbd_max_timeout, uint, 0644);
+MODULE_PARM_DESC(dibs_pbd_max_timeout,
+		 "Maximum timeout for dibs peer bus remote in seconds");
+
+struct dibs_pbd_dmb_state {
+	struct dibs_peer_device *peer_dev;
+	struct dibs_peer_dev_driver *peer_dev_driver;
+	struct dibs_pbd_remote_endpoint *remote_endpoint;
+	/*
+	 * Protects peer_dev_driver and peer_dev pointers in dmb_state array.
+	 * Synchronizes IRQ handler with driver bind/unbind to prevent races
+	 * where driver callbacks are invoked on unbound or freed drivers.
+	 */
+	spinlock_t dmb_state_lock;
+};
+
+struct dibs_pbd_event_work {
+	struct work_struct work;
+	struct dibs_dev *dibs;
+	struct dibs_event event;
+};
+
+struct dibs_pbd_dibs_dev_state {
+	struct list_head list;
+	struct dibs_dev *dibs;
+	/* index is the dmbno of the parent dibs device */
+	struct dibs_pbd_dmb_state dmb_state[];
+};
+
+static LIST_HEAD(dibs_pbd_dibs_dev_state_list);
+/*
+ * List is modified by add_dev/del_dev (called with dibs_dev_list.mutex)
+ * and read by event handlers and sysfs operations.
+ */
+static DEFINE_MUTEX(dibs_pbd_dibs_dev_state_list_mutex);
+
+static LIST_HEAD(dibs_pbd_remote_endpoint_list);
+/*
+ * Endpoints may be added from sysfs and recovery work items and
+ * removed asynchronously from dev_disabled work items. Readers
+ * traversing the list must hold this mutex to prevent races with
+ * concurrent list_add()/list_del() operations.
+ */
+static DEFINE_MUTEX(dibs_pbd_remote_endpoint_list_mutex);
+
+struct access_token {
+	struct list_head list;
+	u128 access_tok;
+};
+
+static LIST_HEAD(dibs_pbd_access_token_list);
+
+static DEFINE_MUTEX(dibs_pbd_access_token_list_mutex);
+
+enum { DIBS_PBD_HDR_V1 = 1 };
+enum {
+	DIBS_PBD_REMOTE_ENDPOINT_ACCESS_TOKEN_MSG = 1,
+	DIBS_PBD_SYSTEM_INFO_MSG = 2,
+	DIBS_PBD_DEV_MSG = 3,
+	DIBS_PBD_PADDING_MSG = 0xff
+};
+
+struct dibs_pbd_remote_endpoint_dev_msg {
+	u64 dmb_tok;
+	u8 peer_dev_type;
+	u8 id;
+	u128 peer_dev_tok;
+} __packed;
+
+struct dibs_pbd_re_sysinfo_msg {
+	u64 dmb_tok;
+	u8 system_name[16];
+	u8 uuid[16];
+	u8 hostname[128];
+} __packed;
+
+struct dibs_pbd_remote_endpoint_access_tok_msg {
+	u128 access_tok;
+} __packed;
+
+struct dibs_pbd_remote_endpoint_msg {
+	struct dibs_msg_hdr hdr;
+	union {
+		struct dibs_pbd_remote_endpoint_access_tok_msg atmsg;
+		struct dibs_pbd_remote_endpoint_dev_msg devmsg;
+		struct dibs_pbd_re_sysinfo_msg simsg;
+	};
+};
+
+#define DIBS_PBD_RING_SIZE (1024 * 1024)
+#define DIBS_PBD_MSG_ALIGNMENT 4
+
+static u32 dibs_pbd_msg_size(struct dibs_msg_hdr *hdr)
+{
+	u32 size;
+
+	if (check_add_overflow(sizeof(*hdr), READ_ONCE(hdr->datalen), &size))
+		return 0;
+	return ALIGN(size, DIBS_PBD_MSG_ALIGNMENT);
+}
+
+/**
+ * dibs_pbd_msg_max_payload() - Calculate maximum payload size for a message
+ * @rb: Ring buffer to check available space
+ *
+ * Calculates the maximum payload size that can be sent in a single message
+ * without having to wrap around.
+ *
+ * Return: Maximum payload size in bytes, or 0 if no space available
+ */
+size_t dibs_pbd_msg_max_payload(struct dibs_ring_buffer *rb)
+{
+	size_t space = dibs_ring_space_to_end(rb);
+
+	if (space == 0)
+		return 0;
+
+	if (space == sizeof(struct dibs_msg_hdr)) {
+		space = dibs_ring_space(rb) - space;
+		if (space == 0)
+			return 0;
+	}
+
+	return ALIGN_DOWN(space - sizeof(struct dibs_msg_hdr),
+			  DIBS_PBD_MSG_ALIGNMENT);
+}
+EXPORT_SYMBOL_GPL(dibs_pbd_msg_max_payload);
+
+static int dibs_pbd_send_padding_msg(struct dibs_ring_buffer *rb)
+{
+	int ret;
+	struct dibs_msg_hdr padding_hdr;
+
+	padding_hdr.version = DIBS_PBD_HDR_V1;
+	padding_hdr.type = DIBS_PBD_PADDING_MSG;
+	padding_hdr.datalen =
+		dibs_ring_space_to_end(rb) - sizeof(struct dibs_msg_hdr);
+
+	ret = dibs_ring_send(rb, &padding_hdr, sizeof(padding_hdr), false);
+	if (ret)
+		return ret;
+
+	return dibs_ring_send_padding(rb, padding_hdr.datalen);
+}
+
+/**
+ * dibs_pbd_msg_send() - Send a message through the ring buffer
+ * @rb: Ring buffer to send message through
+ * @hdr: Message header with payload to send
+ *
+ * Sends a message through the ring buffer, handling alignment and wrapping.
+ * If the message doesn't fit at the end of the ring buffer, a padding message
+ * is sent first to wrap to the beginning.
+ *
+ * Return: 0 on success, -EINVAL if message size overflows, -ENOSPC if
+ *         insufficient space in ring buffer, or error from ring buffer operations
+ */
+int dibs_pbd_msg_send(struct dibs_ring_buffer *rb, struct dibs_msg_hdr *hdr)
+{
+	u32 msg_size = dibs_pbd_msg_size(hdr);
+	int ret;
+
+	if (!msg_size)
+		return -EINVAL;
+
+	if (msg_size > dibs_ring_space(rb))
+		return -ENOSPC;
+
+	if (msg_size > dibs_ring_space_to_end(rb)) {
+		ret = dibs_pbd_send_padding_msg(rb);
+		if (ret)
+			return ret;
+
+		if (msg_size > dibs_ring_space(rb))
+			return -ENOSPC;
+	}
+
+	return dibs_ring_send(rb, hdr, msg_size, true);
+}
+EXPORT_SYMBOL_GPL(dibs_pbd_msg_send);
+
+/**
+ * dibs_pbd_msg_recv() - Receive a message from the ring buffer
+ * @rb: Ring buffer to receive message from
+ *
+ * Receives the next message from the ring buffer, automatically skipping
+ * any padding messages used for alignment.
+ *
+ * Return: Pointer to received message header, or NULL if no message available
+ *         or error occurred
+ */
+struct dibs_msg_hdr *dibs_pbd_msg_recv(struct dibs_ring_buffer *rb)
+{
+	struct dibs_msg_hdr *msg;
+	u32 msg_size;
+
+	while ((msg = dibs_ring_recv(rb))) {
+		msg_size = dibs_pbd_msg_size(msg);
+		if (!msg_size || msg_size > dibs_ring_cnt_to_end(rb)) {
+			pr_warn("%s(%s): discarding message with invalid size\n",
+				__func__, rb->name);
+			rb->msg_size = 0;
+			return NULL;
+		}
+
+		rb->msg_size = msg_size;
+
+		if (READ_ONCE(msg->type) != DIBS_PBD_PADDING_MSG)
+			break;
+		if (dibs_pbd_msg_ack(rb))
+			return NULL;
+	}
+
+	return msg;
+}
+EXPORT_SYMBOL_GPL(dibs_pbd_msg_recv);
+
+/**
+ * dibs_pbd_msg_ack() - Acknowledge a received message
+ * @rb: Ring buffer the message was received from
+ *
+ * Acknowledges the message dibs_pbd_msg_recv() handed out, freeing its space
+ * in the ring buffer for reuse. Must be called after processing each received
+ * message. Uses the size validated by dibs_pbd_msg_recv(), so that the peer
+ * cannot move the tail by rewriting its datalen afterwards.
+ *
+ * Return: 0 on success, -EINVAL if there is no message to acknowledge, or
+ *         error from ring buffer operations
+ */
+int dibs_pbd_msg_ack(struct dibs_ring_buffer *rb)
+{
+	u32 msg_size = rb->msg_size;
+
+	if (!msg_size)
+		return -EINVAL;
+
+	rb->msg_size = 0;
+
+	return dibs_ring_ack(rb, msg_size);
+}
+EXPORT_SYMBOL_GPL(dibs_pbd_msg_ack);
+
+struct dibs_pbd_remote_endpoint {
+	struct dibs_dev *dibs;
+	struct list_head list;
+	uuid_t rgid;
+	bool trusted;
+	struct dibs_ring_buffer rb;
+	struct work_struct handle_irq_work;
+	/*
+	 * Reference count protects this structure from being freed while
+	 * IRQ handler or work items are using it. IRQ handler takes reference
+	 * with kref_get_unless_zero() before scheduling work. Work handler
+	 * drops reference when done. Initial reference is dropped in
+	 * handle_event_work for DIBS_DEV_DISABLED events.
+	 */
+	struct kref kref;
+};
+
+static int dibs_pbd_match(struct device *dev, const struct device_driver *drv)
+{
+	struct dibs_peer_dev_driver *dpdd;
+	struct dibs_peer_device *dpd;
+
+	dpdd = to_dibs_peer_dev_driver(drv);
+	dpd = to_dibs_peer_device(dev);
+	if (dpdd->type && dpdd->type == dpd->type)
+		return 1;
+	return 0;
+};
+
+static int dibs_pbd_probe(struct device *dev)
+{
+	struct dibs_peer_dev_driver *dpdd;
+	struct dibs_peer_device *dpd;
+
+	if (dev->driver) {
+		dpd = to_dibs_peer_device(dev);
+		dpdd = to_dibs_peer_dev_driver(dev->driver);
+		return dpdd->probe(dpd);
+	}
+
+	return 0;
+}
+
+static void dibs_pbd_remove(struct device *dev)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+	struct dibs_peer_dev_driver *dpdd;
+
+	if (dev->driver) {
+		dpdd = to_dibs_peer_dev_driver(dev->driver);
+		dpdd->remove(dpd);
+	}
+}
+
+static int dibs_pbd_uevent(const struct device *dev,
+			   struct kobj_uevent_env *env)
+{
+	struct dibs_peer_device *dpd;
+
+	dpd = to_dibs_peer_device(dev);
+	add_uevent_var(env, "MODALIAS=dibs:peer_dev:%d", dpd->type);
+
+	return 0;
+}
+
+static int dibs_pbd_notifier(struct notifier_block *nb, unsigned long action,
+			     void *data)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(data);
+
+	switch (action) {
+	case BUS_NOTIFY_BOUND_DRIVER:
+		dpd->status = PEER_DEV_CONNECTED;
+		break;
+	case BUS_NOTIFY_UNBIND_DRIVER:
+		dpd->status = PEER_DEV_DISCONNECTED;
+		break;
+	}
+	return NOTIFY_OK;
+}
+
+static struct notifier_block dibs_pbd_nb = { .notifier_call =
+						     dibs_pbd_notifier };
+
+static int dibs_pbd_s390_get_system_info(struct dibs_system_info *sysinfo)
+{
+	struct sysinfo_2_2_2 *info222;
+	struct sysinfo_3_2_2 *info322;
+	void *info;
+	int level;
+
+	info = (void *)get_zeroed_page(GFP_KERNEL);
+	if (!info)
+		return -ENOMEM;
+
+	memset(sysinfo->system_name, 0, sizeof(sysinfo->system_name));
+
+	level = stsi(NULL, 0, 0, 0);
+	if (level >= 2) {
+		info222 = (struct sysinfo_2_2_2 *)info;
+		if (stsi(info222, 2, 2, 2) == 0) {
+			EBCASC(info222->name, sizeof(info222->name));
+			memcpy(&sysinfo->system_name, info222->name,
+			       sizeof(sysinfo->system_name));
+			memcpy(&sysinfo->uuid, info222->uuid.b, UUID_SIZE);
+		}
+	}
+	if (level >= 3) {
+		info322 = (struct sysinfo_3_2_2 *)info;
+		if (stsi(info322, 3, 2, 2) == 0) {
+			EBCASC(info322->vm[0].name,
+			       sizeof(info322->vm[0].name));
+			memcpy(&sysinfo->system_name, info322->vm[0].name,
+			       sizeof(sysinfo->system_name));
+			memcpy(&sysinfo->uuid, info322->vm[0].uuid.b,
+			       UUID_SIZE);
+		}
+	}
+
+	free_page((unsigned long)info);
+	return 0;
+}
+
+static int dibs_pbd_get_system_info(struct dibs_system_info *sysinfo)
+{
+	int ret = 0;
+
+	strscpy(sysinfo->hostname, utsname()->nodename);
+
+#ifdef CONFIG_S390
+	ret = dibs_pbd_s390_get_system_info(sysinfo);
+#endif
+	return ret;
+}
+
+static void dibs_pbd_re_access_tok_msg(struct dibs_msg_hdr *msg_hdr,
+				       struct dibs_pbd_remote_endpoint *re)
+{
+	struct dibs_pbd_remote_endpoint_access_tok_msg atmsg;
+	struct dibs_pbd_remote_endpoint_msg msg_to_send;
+	struct dibs_pbd_remote_endpoint_msg *msg;
+	struct dibs_system_info sysinfo;
+	struct access_token *at;
+	int ret;
+
+	msg = (struct dibs_pbd_remote_endpoint_msg *)msg_hdr;
+	atmsg = msg->atmsg;
+
+	if (re->trusted) {
+		at = kzalloc_obj(*at);
+		if (!at) {
+			pr_err("%s: failed to allocate access token\n",
+			       __func__);
+			return;
+		}
+		at->access_tok = atmsg.access_tok;
+		mutex_lock(&dibs_pbd_access_token_list_mutex);
+		list_add_tail(&at->list, &dibs_pbd_access_token_list);
+		mutex_unlock(&dibs_pbd_access_token_list_mutex);
+	} else {
+		mutex_lock(&dibs_pbd_access_token_list_mutex);
+		list_for_each_entry(at, &dibs_pbd_access_token_list, list) {
+			if (at->access_tok == atmsg.access_tok)
+				re->trusted = true;
+		}
+		mutex_unlock(&dibs_pbd_access_token_list_mutex);
+	}
+
+	if (re->trusted) {
+		ret = dibs_pbd_get_system_info(&sysinfo);
+		if (!ret) {
+			msg_to_send.hdr.version = DIBS_PBD_HDR_V1;
+			msg_to_send.hdr.type = DIBS_PBD_SYSTEM_INFO_MSG;
+			msg_to_send.hdr.datalen =
+				sizeof(struct dibs_pbd_re_sysinfo_msg);
+			strscpy(msg_to_send.simsg.system_name,
+				sysinfo.system_name);
+			memcpy(msg_to_send.simsg.uuid, sysinfo.uuid,
+			       sizeof(sysinfo.uuid));
+			strscpy(msg_to_send.simsg.hostname, sysinfo.hostname);
+
+			dibs_pbd_msg_send(&re->rb,
+					  (struct dibs_msg_hdr *)&msg_to_send);
+		}
+	}
+}
+
+static int dibs_pbd_peer_dev_match_id(struct device *dev, const void *data)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+	const u8 *id = data;
+
+	if (dpd && dpd->id == *id)
+		return 1;
+	return 0;
+}
+
+static void dibs_pbd_peer_dev_release(struct device *dev)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+
+	kfree(dpd);
+}
+
+static ssize_t type_show(struct device *dev, struct device_attribute *attr,
+			 char *buf)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+
+	return sysfs_emit(buf, "%u\n", dpd->type);
+}
+static DEVICE_ATTR_RO(type);
+
+static ssize_t id_show(struct device *dev, struct device_attribute *attr,
+		       char *buf)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+
+	return sysfs_emit(buf, "%u\n", dpd->id);
+}
+static DEVICE_ATTR_RO(id);
+
+static ssize_t dmb_tok_show(struct device *dev, struct device_attribute *attr,
+			    char *buf)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+
+	if (!dpd->rb.dmb.dmb_tok)
+		return -ENODEV;
+
+	return sysfs_emit(buf, "%pUb,%llx\n", &dpd->parent->gid,
+			  dpd->rb.dmb.dmb_tok);
+}
+static DEVICE_ATTR_RO(dmb_tok);
+
+static ssize_t peer_show(struct device *dev, struct device_attribute *attr,
+			 char *buf)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+
+	if (uuid_is_null(&dpd->rgid))
+		return -ENODEV;
+	return sysfs_emit(buf, "%pUb\n", &dpd->rgid);
+}
+static DEVICE_ATTR_RO(peer);
+
+static ssize_t state_show(struct device *dev, struct device_attribute *attr,
+			  char *buf)
+{
+	struct dibs_peer_device *dpd;
+
+	dpd = to_dibs_peer_device(dev);
+	return sysfs_emit(buf, "%u\n", dpd->status);
+}
+static DEVICE_ATTR_RO(state);
+
+static struct attribute *dibs_peer_dev_attrs[] = {
+	&dev_attr_dmb_tok.attr, &dev_attr_peer.attr, &dev_attr_state.attr,
+	&dev_attr_type.attr,	&dev_attr_id.attr,   NULL,
+};
+ATTRIBUTE_GROUPS(dibs_peer_dev);
+
+static void dibs_pbd_dibs_event_dev_disabled_work(struct work_struct *work)
+{
+	struct dibs_peer_dev_driver *dpdd;
+	struct dibs_peer_device *dpd;
+
+	dpd = container_of(work, struct dibs_peer_device,
+			   event_dev_disabled_work);
+
+	device_lock(&dpd->dev);
+	if (dpd->dev.driver) {
+		dpdd = to_dibs_peer_dev_driver(dpd->dev.driver);
+		dpdd->disconnect(dpd);
+	}
+	device_unlock(&dpd->dev);
+
+	schedule_delayed_work(&dpd->timeout_work, dpd->timeout);
+}
+
+static void dibs_pbd_peer_dev_destroy(struct dibs_peer_device *dpd)
+{
+	if (atomic_xchg(&dpd->destroying, 1))
+		return;
+
+	cancel_work_sync(&dpd->event_dev_disabled_work);
+
+	/* Don't cancel if we're running in the timeout_work context */
+	if (current_work() != &dpd->timeout_work.work)
+		cancel_delayed_work_sync(&dpd->timeout_work);
+
+	device_unregister(&dpd->dev);
+}
+
+static void dibs_pbd_dibs_dev_timeout_work(struct work_struct *work)
+{
+	struct dibs_peer_device *dpd;
+
+	dpd = container_of(to_delayed_work(work), struct dibs_peer_device,
+			   timeout_work);
+
+	dibs_pbd_peer_dev_destroy(dpd);
+}
+
+const struct bus_type dibs_pbd_type;
+
+static int dibs_pbd_peer_dev_create(struct dibs_dev *dibs, uuid_t *rgid,
+				    u64 rdmb_tok, u8 server_dev_id, u8 type,
+				    u128 tok)
+{
+	struct dibs_peer_device *dpd;
+	int ret;
+
+	dpd = kzalloc_obj(*dpd, GFP_KERNEL);
+	if (!dpd)
+		return -ENOMEM;
+	atomic_set(&dpd->destroying, 0);
+	dpd->id = server_dev_id;
+	dpd->rgid = *rgid;
+	dpd->type = type;
+	dpd->dev.bus = &dibs_pbd_type;
+	dpd->dev.release = dibs_pbd_peer_dev_release;
+	dpd->parent = dibs;
+	dpd->status = PEER_DEV_DISCONNECTED;
+	dpd->rdmb_tok = rdmb_tok;
+	dpd->dev.groups = dibs_peer_dev_groups;
+	dpd->tok = tok;
+	dev_set_name(&dpd->dev, "%pUb:%u", rgid, dpd->id);
+
+	INIT_WORK(&dpd->event_dev_disabled_work,
+		  dibs_pbd_dibs_event_dev_disabled_work);
+	INIT_DELAYED_WORK(&dpd->timeout_work, dibs_pbd_dibs_dev_timeout_work);
+
+	dpd->timeout = secs_to_jiffies(dibs_pbd_max_timeout);
+
+	ret = device_register(&dpd->dev);
+	if (ret) {
+		put_device(&dpd->dev);
+		return ret;
+	}
+	ret = sysfs_create_link(&dpd->dev.kobj, &dpd->parent->dev.kobj, "dibs");
+	if (ret) {
+		device_unregister(&dpd->dev);
+		return ret;
+	}
+	return 0;
+}
+
+static void dibs_pbd_re_dev_msg(struct dibs_msg_hdr *msg_hdr,
+				struct dibs_pbd_remote_endpoint *re)
+{
+	struct dibs_pbd_remote_endpoint_dev_msg redm;
+	struct dibs_pbd_remote_endpoint_msg *msg;
+	struct dibs_peer_dev_driver *dpdd;
+	struct dibs_peer_device *dpd;
+	struct device *dev;
+	int ret;
+
+	msg = (struct dibs_pbd_remote_endpoint_msg *)msg_hdr;
+	redm = msg->devmsg;
+
+	dev = bus_find_device(&dibs_pbd_type, NULL, &redm.id,
+			      dibs_pbd_peer_dev_match_id);
+	if (dev) {
+		dpd = to_dibs_peer_device(dev);
+		if (dpd->tok != redm.peer_dev_tok) {
+			put_device(dev);
+			return;
+		}
+		if (dpd->rdmb_tok != redm.dmb_tok)
+			dpd->rdmb_tok = redm.dmb_tok;
+
+		device_lock(&dpd->dev);
+		if (dpd->dev.driver) {
+			dpdd = to_dibs_peer_dev_driver(dpd->dev.driver);
+			ret = dpdd->reconnect(dpd);
+			if (!ret)
+				dpd->status = PEER_DEV_CONNECTED;
+		}
+		device_unlock(&dpd->dev);
+		put_device(dev);
+	} else {
+		dibs_pbd_peer_dev_create(re->dibs, &re->rgid, redm.dmb_tok,
+					 redm.id, redm.peer_dev_type,
+					 redm.peer_dev_tok);
+	}
+}
+
+static void dibs_pbd_remote_endpoint_release(struct kref *kref)
+{
+	struct dibs_pbd_remote_endpoint *re;
+
+	re = container_of(kref, struct dibs_pbd_remote_endpoint, kref);
+
+	dibs_ring_unregister(&re->rb);
+	kfree(re);
+}
+
+static struct dibs_client dibs_pbd_drv;
+static void dibs_pbd_handle_irq_work(struct work_struct *work)
+{
+	struct dibs_pbd_remote_endpoint *re;
+	struct dibs_ring_buffer *rb;
+	struct dibs_msg_hdr *msg;
+
+	re = container_of(work, struct dibs_pbd_remote_endpoint,
+			  handle_irq_work);
+	rb = &re->rb;
+
+	while ((msg = dibs_pbd_msg_recv(rb))) {
+		switch (msg->type) {
+		case DIBS_PBD_REMOTE_ENDPOINT_ACCESS_TOKEN_MSG:
+			dibs_pbd_re_access_tok_msg(msg, re);
+			break;
+		case DIBS_PBD_DEV_MSG:
+			if (re->trusted)
+				dibs_pbd_re_dev_msg(msg, re);
+			break;
+		default:
+			pr_warn("%s: Received unknown message type: %d\n",
+				__func__, msg->type);
+			kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
+			return;
+		}
+		dibs_pbd_msg_ack(rb);
+	}
+
+	kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
+}
+
+static void dibs_pbd_drv_handle_irq(struct dibs_dev *dibs, unsigned int dmbno,
+				    u16 dmbemask)
+{
+	struct dibs_pbd_dibs_dev_state *dds;
+	struct dibs_pbd_remote_endpoint *re;
+	struct dibs_peer_dev_driver *dpdd;
+	struct dibs_peer_device *dpd;
+	unsigned long flags;
+
+	dds = dibs_get_priv(dibs, &dibs_pbd_drv);
+
+	spin_lock_irqsave(&dds->dmb_state[dmbno].dmb_state_lock, flags);
+	dpdd = dds->dmb_state[dmbno].peer_dev_driver;
+	dpd = dds->dmb_state[dmbno].peer_dev;
+	if (dpdd && dpd)
+		dpdd->handle_irq(dmbno, dpd, 0);
+	re = dds->dmb_state[dmbno].remote_endpoint;
+	if (re && !kref_get_unless_zero(&re->kref))
+		re = NULL;
+	spin_unlock_irqrestore(&dds->dmb_state[dmbno].dmb_state_lock, flags);
+
+	if (!re)
+		return;
+
+	if (!queue_work(system_dfl_wq, &re->handle_irq_work)) {
+		/* Work already pending, drop the reference */
+		kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
+	}
+	/* If queue_work returned true, work owns the reference */
+}
+
+static int dibs_pbd_peer_dev_event(struct device *dev, void *data)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+	struct dibs_event *event = data;
+
+	switch (event->subtype) {
+	case DIBS_SW_EVENT_REDISCOVER:
+		dpd->rgid = event->gid;
+		cancel_delayed_work(&dpd->timeout_work);
+		break;
+	case DIBS_DEV_DISABLED:
+		if (uuid_equal(&dpd->rgid, &event->gid))
+			schedule_work(&dpd->event_dev_disabled_work);
+		break;
+	}
+
+	return 0;
+}
+
+static int dibs_pbd_remote_endpoint_create(struct dibs_dev *dibs,
+					   const uuid_t *rgid,
+					   const bool trusted)
+{
+	struct dibs_pbd_remote_endpoint *re, *tmp;
+	struct dibs_pbd_dibs_dev_state *dds;
+	bool existing_re_found = false;
+	unsigned long flags;
+	int ret;
+
+	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
+	list_for_each_entry_safe(re, tmp, &dibs_pbd_remote_endpoint_list,
+				 list) {
+		if (uuid_equal(&re->rgid, rgid))
+			existing_re_found = true;
+	}
+	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+
+	if (existing_re_found) {
+		pr_warn("%s: remote endpoint already exists for rgid: %pUb\n",
+			__func__, rgid);
+		return 0;
+	}
+
+	re = kzalloc_obj(*re, GFP_KERNEL);
+	if (!re)
+		return -ENOMEM;
+
+	re->dibs = dibs;
+	uuid_copy(&re->rgid, rgid);
+	re->trusted = trusted;
+
+	ret = dibs_ring_register("peer_bus", &re->rb, &dibs_pbd_drv, dibs,
+				 re->rgid, DIBS_PBD_RING_SIZE);
+
+	if (ret) {
+		pr_err("%s: failed to register ring buffer\n", __func__);
+		kfree(re);
+		return ret;
+	}
+
+	INIT_WORK(&re->handle_irq_work, dibs_pbd_handle_irq_work);
+	kref_init(&re->kref);
+
+	dds = dibs_get_priv(dibs, &dibs_pbd_drv);
+	spin_lock_irqsave(&dds->dmb_state[re->rb.dmb.idx].dmb_state_lock,
+			  flags);
+	dds->dmb_state[re->rb.dmb.idx].remote_endpoint = re;
+	dds->dmb_state[re->rb.rd_meta_dmb.idx].remote_endpoint = re;
+	spin_unlock_irqrestore(&dds->dmb_state[re->rb.dmb.idx].dmb_state_lock,
+			       flags);
+
+	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
+	list_add(&re->list, &dibs_pbd_remote_endpoint_list);
+	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+
+	re->dibs->ops->signal_event(re->dibs, &re->rgid, 1,
+				    DIBS_SW_EVENT_DISCOVER,
+				    re->rb.wr_ring_meta.meta_dmb_tok);
+	dibs_pbd_boot_event_sent = 1;
+
+	return 0;
+}
+
+static void dibs_pbd_drv_handle_event_work(struct work_struct *work)
+{
+	struct dibs_pbd_remote_endpoint *re, *tmp;
+	struct dibs_pbd_event_work *ew;
+
+	ew = container_of(work, struct dibs_pbd_event_work, work);
+
+	bus_for_each_dev(&dibs_pbd_type, NULL, &ew->event,
+			 dibs_pbd_peer_dev_event);
+
+	switch (ew->event.subtype) {
+	case DIBS_SW_EVENT_REDISCOVER:
+		if (!ew->dibs->ops->query_remote_gid(ew->dibs, &ew->event.gid,
+						     0, 0))
+			dibs_pbd_remote_endpoint_create(ew->dibs,
+							&ew->event.gid, false);
+		break;
+	case DIBS_DEV_DISABLED:
+		mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
+		list_for_each_entry_safe(re, tmp,
+					 &dibs_pbd_remote_endpoint_list, list) {
+			if (uuid_equal(&re->rgid, &ew->event.gid)) {
+				list_del(&re->list);
+				/* Drop initial reference - free when all other refs are gone */
+				kref_put(&re->kref,
+					 dibs_pbd_remote_endpoint_release);
+			}
+		}
+		mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+		break;
+	}
+
+	kfree(ew);
+}
+
+static void dibs_pbd_drv_handle_event(struct dibs_dev *dev,
+				      const struct dibs_event *event)
+{
+	struct dibs_pbd_event_work *ew;
+
+	ew = kzalloc_obj(*ew, GFP_ATOMIC);
+	if (!ew)
+		return;
+	ew->dibs = dev;
+	ew->event = *event;
+
+	INIT_WORK(&ew->work, dibs_pbd_drv_handle_event_work);
+	schedule_work(&ew->work);
+}
+
+static ssize_t dibs_pbd_discover_store_handle_gid_prefix(const char *buf,
+							 size_t count)
+{
+	struct dibs_pbd_dibs_dev_state *dds;
+	struct dibs_pbd_remote_endpoint *re;
+	bool remote_system_reachable = false;
+	const char *gid_str;
+	uuid_t rgid;
+	int ret;
+
+	gid_str = buf + strlen("gid:");
+	ret = uuid_parse(gid_str, &rgid);
+	if (ret) {
+		pr_err("%s: could not parse rgid\n", __func__);
+		return -EINVAL;
+	}
+	if (uuid_is_null(&rgid)) {
+		pr_err("%s: gid is null\n", __func__);
+		return -EINVAL;
+	}
+	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
+	list_for_each_entry(re, &dibs_pbd_remote_endpoint_list, list) {
+		if (!uuid_equal(&rgid, &re->rgid))
+			continue;
+		re->dibs->ops->signal_event(re->dibs, &re->rgid, 1,
+					    DIBS_SW_EVENT_DISCOVER,
+					    re->rb.wr_ring_meta.meta_dmb_tok);
+		mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+		return count;
+	}
+	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+
+	mutex_lock(&dibs_pbd_dibs_dev_state_list_mutex);
+	list_for_each_entry(dds, &dibs_pbd_dibs_dev_state_list, list) {
+		if (dds->dibs->ops->query_remote_gid(dds->dibs, &rgid, 0, 0))
+			continue;
+		remote_system_reachable = true;
+		ret = dibs_pbd_remote_endpoint_create(dds->dibs, &rgid, true);
+		if (ret) {
+			mutex_unlock(&dibs_pbd_dibs_dev_state_list_mutex);
+			return ret;
+		}
+		break;
+	}
+	mutex_unlock(&dibs_pbd_dibs_dev_state_list_mutex);
+	if (!remote_system_reachable) {
+		pr_err("%s: remote system not reachable\n", __func__);
+		return -ENODEV;
+	}
+
+	return count;
+}
+
+static ssize_t discover_store(const struct bus_type *bus, const char *buf,
+			      size_t count)
+{
+	if (str_has_prefix(buf, "gid:"))
+		return dibs_pbd_discover_store_handle_gid_prefix(buf, count);
+
+	pr_err("dibs_pbd: unknown discover identifier type\n");
+	return -EINVAL;
+}
+static BUS_ATTR_WO(discover);
+
+static struct attribute *dibs_pbd_attrs[] = {
+	&bus_attr_discover.attr,
+	NULL,
+};
+ATTRIBUTE_GROUPS(dibs_pbd);
+
+const struct bus_type dibs_pbd_type = {
+	.name = "dibs_peer",
+	.match = dibs_pbd_match,
+	.bus_groups = dibs_pbd_groups,
+	.uevent = dibs_pbd_uevent,
+	.probe = dibs_pbd_probe,
+	.remove = dibs_pbd_remove,
+};
+
+static void dibs_pbd_drv_add_dibs(struct dibs_dev *dibs)
+{
+	struct dibs_pbd_dibs_dev_state *dds;
+	const char *gid_str;
+	int max_dmbs;
+	uuid_t gid;
+	int ret;
+
+	max_dmbs = dibs->ops->max_dmbs();
+	dds = kzalloc(sizeof(*dds) +
+			      max_dmbs * sizeof(struct dibs_pbd_dmb_state),
+		      GFP_KERNEL);
+	if (!dds) {
+		dev_err(&dibs->dev,
+			"Failed to allocate memory for dibs dev state\n");
+		return;
+	}
+
+	dds->dibs = dibs;
+	for (int i = 0; i < max_dmbs; i++)
+		spin_lock_init(&dds->dmb_state[i].dmb_state_lock);
+	dibs_set_priv(dibs, &dibs_pbd_drv, dds);
+	mutex_lock(&dibs_pbd_dibs_dev_state_list_mutex);
+	list_add(&dds->list, &dibs_pbd_dibs_dev_state_list);
+	mutex_unlock(&dibs_pbd_dibs_dev_state_list_mutex);
+
+	/* Send boot event to remote system when module parameter rgid is set.
+	 * Send boot event with the first dibs device that is able to
+	 * communicate with the remote system.
+	 */
+
+	if (dibs_peer_address && str_has_prefix(dibs_peer_address, "gid:") &&
+	    !dibs_pbd_boot_event_sent) {
+		gid_str = dibs_peer_address + strlen("gid:");
+		if (strlen(gid_str) != UUID_STRING_LEN) {
+			pr_err("%s: gid is not valid\n", __func__);
+			return;
+		}
+		ret = uuid_parse(gid_str, &gid);
+		if (ret) {
+			pr_err("%s: Could not parse gid\n", __func__);
+			return;
+		}
+		if (uuid_is_null(&gid)) {
+			pr_err("%s: gid is null\n", __func__);
+			return;
+		}
+		if (!dibs->ops->query_remote_gid(dibs, &gid, 0, 0)) {
+			/* Create a peer device for the dibs base client
+			 * This peer device registers the discovery DMB
+			 */
+			dibs_pbd_remote_endpoint_create(dibs, &gid, true);
+		}
+	}
+}
+
+static int dibs_pbd_destroy_matching_peer_dev(struct device *dev, void *data)
+{
+	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
+	struct dibs_dev *removed = data;
+
+	if (dpd->parent != removed)
+		return 0;
+
+	dibs_pbd_peer_dev_destroy(dpd);
+
+	return 0;
+}
+
+/**
+ * dibs_pbd_driver_register() - Register a peer device driver
+ * @dpdd: Peer device driver to register
+ *
+ * Registers a driver for dibs peer devices on the peer bus.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+int dibs_pbd_driver_register(struct dibs_peer_dev_driver *dpdd)
+{
+	dpdd->drv.bus = &dibs_pbd_type;
+
+	return driver_register(&dpdd->drv);
+}
+EXPORT_SYMBOL(dibs_pbd_driver_register);
+
+/**
+ * dibs_pbd_driver_unregister() - Unregister a peer device driver
+ * @dpdd: Peer device driver to unregister
+ *
+ * Unregisters a previously registered peer device driver.
+ *
+ * Return: Always returns 0
+ */
+int dibs_pbd_driver_unregister(struct dibs_peer_dev_driver *dpdd)
+{
+	driver_unregister(&dpdd->drv);
+	return 0;
+}
+EXPORT_SYMBOL(dibs_pbd_driver_unregister);
+
+static void peer_bus_drv_remove_dibs(struct dibs_dev *dibs)
+{
+	struct dibs_pbd_remote_endpoint *re, *tmp;
+	struct dibs_pbd_dibs_dev_state *dds;
+
+	bus_for_each_dev(&dibs_pbd_type, NULL, dibs,
+			 dibs_pbd_destroy_matching_peer_dev);
+
+	dds = dibs_get_priv(dibs, &dibs_pbd_drv);
+
+	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
+	list_for_each_entry_safe(re, tmp, &dibs_pbd_remote_endpoint_list, list) {
+		if (re->dibs == dibs) {
+			list_del(&re->list);
+			kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
+		}
+	}
+	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
+
+	mutex_lock(&dibs_pbd_dibs_dev_state_list_mutex);
+	list_del_init(&dds->list);
+	mutex_unlock(&dibs_pbd_dibs_dev_state_list_mutex);
+
+	kfree(dds);
+}
+
+static struct dibs_client_ops dibs_pbd_drv_ops = {
+	.add_dev = dibs_pbd_drv_add_dibs,
+	.del_dev = peer_bus_drv_remove_dibs,
+	.handle_event = dibs_pbd_drv_handle_event,
+	.handle_irq = dibs_pbd_drv_handle_irq,
+};
+
+static struct dibs_client dibs_pbd_drv = {
+	.name = "peer_bus",
+	.ops = &dibs_pbd_drv_ops,
+};
+
+/**
+ * dibs_pbd_register_ring_buffer() - Register a ring buffer for a peer device
+ * @name: Name for the ring buffer
+ * @dpd: Peer device to register ring buffer for
+ * @size: Size of the ring buffer
+ *
+ * Registers a ring buffer for communication with a peer device and associates
+ * it with the peer device's driver. The device must be locked before calling.
+ *
+ * Return: 0 on success, -ENODEV if no driver bound, or error from ring
+ *         buffer registration
+ */
+int dibs_pbd_register_ring_buffer(char *name, struct dibs_peer_device *dpd,
+				  u32 size)
+{
+	struct dibs_pbd_dibs_dev_state *dds;
+	struct dibs_peer_dev_driver *dpdd;
+	unsigned long flags;
+	int ret;
+
+	device_lock_assert(&dpd->dev);
+
+	ret = dibs_ring_register(name, &dpd->rb, &dibs_pbd_drv, dpd->parent,
+				 dpd->rgid, size);
+	if (ret)
+		return ret;
+
+	if (!dpd->dev.driver) {
+		dibs_ring_unregister(&dpd->rb);
+		return -ENODEV;
+	}
+	dpdd = to_dibs_peer_dev_driver(dpd->dev.driver);
+	dds = dibs_get_priv(dpd->parent, &dibs_pbd_drv);
+
+	spin_lock_irqsave(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
+			  flags);
+	dds->dmb_state[dpd->rb.dmb.idx].peer_dev_driver = dpdd;
+	dds->dmb_state[dpd->rb.dmb.idx].peer_dev = dpd;
+	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev_driver = dpdd;
+	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev = dpd;
+	spin_unlock_irqrestore(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
+			       flags);
+
+	return 0;
+}
+EXPORT_SYMBOL(dibs_pbd_register_ring_buffer);
+
+/**
+ * dibs_pbd_unregister_ring_buffer() - Unregister a peer device's ring buffer
+ * @dpd: Peer device whose ring buffer to unregister
+ *
+ * Unregisters the ring buffer associated with a peer device and clears
+ * the driver associations.
+ *
+ * Return: 0 on success, or error from ring buffer unregistration
+ */
+int dibs_pbd_unregister_ring_buffer(struct dibs_peer_device *dpd)
+{
+	struct dibs_pbd_dibs_dev_state *dds;
+	unsigned long flags;
+	int ret;
+
+	device_lock_assert(&dpd->dev);
+
+	dds = dibs_get_priv(dpd->parent, &dibs_pbd_drv);
+
+	spin_lock_irqsave(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
+			  flags);
+	dds->dmb_state[dpd->rb.dmb.idx].peer_dev_driver = NULL;
+	dds->dmb_state[dpd->rb.dmb.idx].peer_dev = NULL;
+	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev_driver = NULL;
+	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev = NULL;
+	spin_unlock_irqrestore(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
+			       flags);
+
+	ret = dibs_ring_unregister(&dpd->rb);
+	dpd->rb.dmb.idx = 0;
+	dpd->rb.rd_meta_dmb.idx = 0;
+
+	return ret;
+}
+EXPORT_SYMBOL(dibs_pbd_unregister_ring_buffer);
+
+/**
+ * dibs_pbd_init() - Initialize the dibs peer bus driver
+ *
+ * Initializes the peer bus subsystem by registering the bus type,
+ * notifier, and dibs client. Must be called before any peer devices
+ * can be used.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+int dibs_pbd_init(void)
+{
+	int ret;
+
+	ret = bus_register(&dibs_pbd_type);
+	if (ret)
+		return ret;
+	ret = bus_register_notifier(&dibs_pbd_type, &dibs_pbd_nb);
+	if (ret)
+		goto bus_reg_not_err;
+	ret = dibs_register_client(&dibs_pbd_drv);
+	if (ret)
+		goto dibs_reg_client_err;
+	return 0;
+
+dibs_reg_client_err:
+	bus_unregister_notifier(&dibs_pbd_type, &dibs_pbd_nb);
+bus_reg_not_err:
+	bus_unregister(&dibs_pbd_type);
+	return ret;
+}
+EXPORT_SYMBOL(dibs_pbd_init);
+
+/**
+ * dibs_pbd_exit() - Clean up the dibs peer bus driver
+ *
+ * Cleans up the peer bus subsystem by unregistering the dibs client,
+ * notifier, and bus type. Should be called during module unload.
+ */
+void dibs_pbd_exit(void)
+{
+	dibs_unregister_client(&dibs_pbd_drv);
+
+	bus_unregister_notifier(&dibs_pbd_type, &dibs_pbd_nb);
+	bus_unregister(&dibs_pbd_type);
+}
+EXPORT_SYMBOL(dibs_pbd_exit);
+MODULE_LICENSE("GPL");
+MODULE_AUTHOR("IBM Corporation");
diff --git a/include/linux/dibs.h b/include/linux/dibs.h
index 30ae03484b58..8c24f7b33e6f 100644
--- a/include/linux/dibs.h
+++ b/include/linux/dibs.h
@@ -164,7 +164,9 @@ enum dibs_event_subtype {
 	DIBS_BUF_UNREGISTERED,
 	DIBS_DEV_DISABLED,
 	DIBS_DEV_ERR_STATE,
-	DIBS_OTHER_SUBTYPE
+	DIBS_OTHER_SUBTYPE,
+	DIBS_SW_EVENT_DISCOVER = 0x9,
+	DIBS_SW_EVENT_REDISCOVER = 0x10,
 };
 
 struct dibs_event {
@@ -269,6 +271,12 @@ int dibs_unregister_client(struct dibs_client *client);
 
 /* dibs clients can call dibs device ops. */
 
+struct dibs_system_info {
+	u8 system_name[16];
+	u8 uuid[16];
+	u8 hostname[128];
+};
+
 /* DIBS devices
  * ------------
  */
@@ -554,10 +562,72 @@ int dibs_dev_add(struct dibs_dev *dibs);
  */
 void dibs_dev_del(struct dibs_dev *dibs);
 
+/*
+ * Dibs Peer Bus
+ */
+#define to_dibs_peer_dev_driver(x) \
+	container_of((x), struct dibs_peer_dev_driver, drv)
+#define to_dibs_peer_device(x) container_of((x), struct dibs_peer_device, dev)
+
+#define PEER_DEV_DISCONNECTED 0
+#define PEER_DEV_CONNECTED 1
+
+struct dibs_peer_device {
+	u8 id;
+	uuid_t rgid;
+	struct dibs_ring_buffer rb;
+	u64 rdmb_tok;
+	u8 status;
+	u8 type;
+	u128 tok;
+	struct delayed_work timeout_work;
+	unsigned long timeout;
+	struct work_struct event_dev_disabled_work;
+	atomic_t destroying;
+	struct dibs_dev *parent;
+	struct device dev;
+	/* priv pointer for driver */
+	void *drv_priv;
+};
+
+struct dibs_peer_dev_driver {
+	u8 type;
+	int (*probe)(struct dibs_peer_device *peer_dev);
+	void (*remove)(struct dibs_peer_device *peer_dev);
+	int (*reconnect)(struct dibs_peer_device *peer_dev);
+	void (*disconnect)(struct dibs_peer_device *peer_dev);
+	void (*handle_irq)(unsigned int dmbno,
+			   struct dibs_peer_device *peer_dev, u16 dmbemask);
+	struct device_driver drv;
+};
+
+#ifdef CONFIG_DIBS_PEER_BUS_DRV
+int dibs_pbd_init(void);
+void dibs_pbd_exit(void);
+#else
+static inline int dibs_pbd_init(void)
+{
+	return 0;
+}
+
+static inline void dibs_pbd_exit(void) {}
+
+#endif
+
 struct dibs_msg_hdr {
 	u8 version;
 	u8 type;
 	u16 datalen;
 } __packed;
 
+int dibs_pbd_register_ring_buffer(char *name, struct dibs_peer_device *dpd,
+				  u32 size);
+int dibs_pbd_unregister_ring_buffer(struct dibs_peer_device *dpd);
+int dibs_pbd_driver_register(struct dibs_peer_dev_driver *driver);
+int dibs_pbd_driver_unregister(struct dibs_peer_dev_driver *ipdd);
+size_t dibs_pbd_msg_max_payload(struct dibs_ring_buffer *rb);
+int dibs_pbd_msg_send(struct dibs_ring_buffer *rb, struct dibs_msg_hdr *hdr);
+struct dibs_msg_hdr *dibs_pbd_msg_recv(struct dibs_ring_buffer *rb);
+int dibs_pbd_msg_ack(struct dibs_ring_buffer *rb);
+
 #endif	/* _DIBS_H */

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 3/4] dibs: introduce dibs console peer device driver
  2026-10-08  8:08 [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver Julian Ruess
  2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
  2026-10-08  8:08 ` [PATCH 2/4] dibs: introduce dibs peer device bus driver Julian Ruess
@ 2026-10-08  8:08 ` Julian Ruess
  2026-10-08  8:26   ` sashiko-bot
  2026-10-08  8:08 ` [PATCH 4/4] dibs: introduce dibs block device " Julian Ruess
  3 siblings, 1 reply; 10+ messages in thread
From: Julian Ruess @ 2026-10-08  8:08 UTC (permalink / raw)
  To: schnelle, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, julianr,
	linux390-list, linux-s390

Introduce a virtual console device driver that makes use of the newly
added dibs peer device bus.

This allows to make use of tty's that connect two systems by using dibs
devices on the same dibs fabric. This works on LPAR as well as z/VM
without requiring a network setup.

The vfio user space tooling on the remote system announces dibs peer
devices of type tty to this instance. If the discovery mechanism of the
dibs peer bus detects a hvc peer device, this driver will be probed via
uevent.

The number of the used hvc device (/dev/hvcX) can be accessed by reading
'/sys/bus/dibs_peer/devices/<rgid>/hvc'

Once the connection is set up, the tty can be accessed by using
'/dev/hvcX'.

Co-developed-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
---
 arch/s390/configs/debug_defconfig |   1 +
 arch/s390/configs/defconfig       |   1 +
 drivers/dibs/Kconfig              |  11 ++
 drivers/dibs/Makefile             |   1 +
 drivers/dibs/dibs_console.c       | 388 ++++++++++++++++++++++++++++++++++++++
 include/linux/dibs.h              |   4 +
 6 files changed, 406 insertions(+)

diff --git a/arch/s390/configs/debug_defconfig b/arch/s390/configs/debug_defconfig
index 23d6c59189cc..c4acf13aeaba 100644
--- a/arch/s390/configs/debug_defconfig
+++ b/arch/s390/configs/debug_defconfig
@@ -134,6 +134,7 @@ CONFIG_SMC_DIAG=m
 CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
 CONFIG_DIBS_PEER_BUS_DRV=y
+CONFIG_DIBS_CONSOLE=m
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/arch/s390/configs/defconfig b/arch/s390/configs/defconfig
index 3bffcd016be8..344724c56917 100644
--- a/arch/s390/configs/defconfig
+++ b/arch/s390/configs/defconfig
@@ -125,6 +125,7 @@ CONFIG_SMC_DIAG=m
 CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
 CONFIG_DIBS_PEER_BUS_DRV=y
+CONFIG_DIBS_CONSOLE=m
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/drivers/dibs/Kconfig b/drivers/dibs/Kconfig
index 98068bc541e5..a0ed49977274 100644
--- a/drivers/dibs/Kconfig
+++ b/drivers/dibs/Kconfig
@@ -32,3 +32,14 @@ config DIBS_PEER_BUS_DRV
 	  foundation for dibs peer device drivers that operate on top of dibs.
 	  It provides device discovery and lifecycle management for peer
 	  devices across the dibs fabric.
+
+config DIBS_CONSOLE
+	tristate "dibs device console support"
+	depends on DIBS && DIBS_PEER_BUS_DRV
+	select HVC_DRIVER
+	help
+	  Support for console access over dibs peer devices of type console.
+	  This peer device driver operates on the dibs peer bus and enables
+	  virtual console communication across systems in a dibs fabric. A
+	  device of type /dev/hvcN will be created for each peer device that is
+	  detected.
diff --git a/drivers/dibs/Makefile b/drivers/dibs/Makefile
index b1418f5149c2..a39a67020e05 100644
--- a/drivers/dibs/Makefile
+++ b/drivers/dibs/Makefile
@@ -7,3 +7,4 @@ dibs-y += dibs_main.o
 obj-$(CONFIG_DIBS) += dibs.o
 dibs-$(CONFIG_DIBS_LO) += dibs_loopback.o
 dibs-$(CONFIG_DIBS_PEER_BUS_DRV) += dibs_peer_bus_drv.o
+obj-$(CONFIG_DIBS_CONSOLE) += dibs_console.o
diff --git a/drivers/dibs/dibs_console.c b/drivers/dibs/dibs_console.c
new file mode 100644
index 000000000000..df53cbcf2fb5
--- /dev/null
+++ b/drivers/dibs/dibs_console.c
@@ -0,0 +1,388 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * dibs hypervisor console (HVC) device driver
+ *
+ * This hvc device driver provides terminal access using
+ * dibs peer devices of type tty.
+ *
+ * Copyright IBM Corp.
+ */
+
+#include "../tty/hvc/hvc_console.h"
+#include <linux/dibs.h>
+#include <linux/list.h>
+#include <linux/minmax.h>
+#include <linux/miscdevice.h>
+#include <linux/module.h>
+#include <linux/printk.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/spinlock_types.h>
+#include <linux/stringify.h>
+#include <linux/uuid.h>
+#include <linux/workqueue.h>
+
+MODULE_DESCRIPTION("dibs hypervisor console (HVC) device driver");
+MODULE_ALIAS("dibs:peer_dev:" __stringify(PEER_DEV_TYPE_CONSOLE));
+
+static unsigned int dibs_con_max_timeout = 60;
+module_param(dibs_con_max_timeout, uint, 0644);
+MODULE_PARM_DESC(dibs_con_max_timeout,
+		 "Maximum timeout for dibs remote in seconds");
+
+enum { DIBS_CON_MAX_PUT_CHARS = 256 };
+enum { DIBS_CON_MAX_VTERMNOS = 8 };
+
+enum { DIBS_CON_PEER_DEV_HDR_VERSION = 1 };
+enum {
+	DIBS_CON_CONNECT_MSG = 1,
+	DIBS_CON_DATA_MSG = 2,
+	DIBS_CON_RESIZE_MSG = 4
+};
+
+#define DIBS_CON_PEER_DEV_BUFFER_SIZE 2048
+#define DIBS_CON_PEER_DEV_RING_SIZE (64 * 1024)
+#define DIBS_CON_PEER_DEV_MAX_RBE_SIZE 512
+
+#define DIBS_CON_PEER_DEV_DATA_MSG_MAX_DATALEN \
+	(DIBS_CON_PEER_DEV_MAX_RBE_SIZE - sizeof(struct dibs_msg_hdr))
+
+struct dibs_con_tty_data_msg {
+	u8 data[DIBS_CON_PEER_DEV_DATA_MSG_MAX_DATALEN];
+} __packed;
+
+struct dibs_con_resize_msg {
+	struct winsize ws;
+} __packed;
+
+struct dibs_con_connect_msg {
+	u64 dmb_tok;
+} __packed;
+
+struct dibs_con_pd_msg {
+	struct dibs_msg_hdr hdr;
+	union {
+		struct dibs_con_connect_msg cm;
+		struct dibs_con_tty_data_msg tdm;
+		struct dibs_con_resize_msg rm;
+	};
+};
+
+struct dibs_con_peer_dev_rbe {
+	union {
+		struct dibs_con_pd_msg pdm;
+		u8 reserved[DIBS_CON_PEER_DEV_MAX_RBE_SIZE];
+	};
+};
+
+struct dibs_con_peer_dev {
+	u8 vtermno;
+	int connect_msg_enqueued;
+	bool ring_registered;
+	struct dibs_peer_device *dpd;
+	struct hvc_struct *hvc;
+	struct kfifo fifo;
+	spinlock_t mr_spinlock /* protects the message ring */;
+};
+
+struct dibs_con_peer_dev *hvc_peer_devs[DIBS_CON_MAX_VTERMNOS] = { NULL };
+
+static void dibs_con_peer_dev_send_connect_msg(struct dibs_peer_device *dpd)
+{
+	struct dibs_con_peer_dev_rbe dcpdr;
+	struct dibs_con_peer_dev *dcpd;
+	int ret;
+
+	dcpd = dpd->drv_priv;
+
+	/* Write own DMB token to the other side */
+	dcpdr.pdm.hdr.version = DIBS_CON_PEER_DEV_HDR_VERSION;
+	dcpdr.pdm.hdr.type = DIBS_CON_CONNECT_MSG;
+	dcpdr.pdm.hdr.datalen = sizeof(dpd->rb.dmb.dmb_tok);
+	dcpdr.pdm.cm.dmb_tok = dpd->rb.dmb.dmb_tok;
+
+	ret = dibs_pbd_msg_send(&dpd->rb, &dcpdr.pdm.hdr);
+	if (!ret)
+		dcpd->connect_msg_enqueued = 1;
+}
+
+static void dibs_con_peer_dev_handle_irq(unsigned int dmbno,
+					 struct dibs_peer_device *peer_dev,
+					 u16 dmbemask)
+{
+	struct dibs_con_peer_dev *dcpd;
+	struct dibs_con_pd_msg *msg;
+	struct winsize winsize;
+	size_t count;
+	u16 datalen;
+
+	dcpd = peer_dev->drv_priv;
+
+	if (!dcpd->connect_msg_enqueued)
+		dibs_con_peer_dev_send_connect_msg(peer_dev);
+
+	while ((msg = (struct dibs_con_pd_msg *)dibs_pbd_msg_recv(&peer_dev->rb))) {
+		switch (msg->hdr.type) {
+		case DIBS_CON_DATA_MSG:
+			datalen = min_t(u16, READ_ONCE(msg->hdr.datalen),
+					DIBS_CON_PEER_DEV_DATA_MSG_MAX_DATALEN);
+			count = kfifo_in(&dcpd->fifo, &msg->tdm, datalen);
+			if (count < datalen)
+				pr_warn("%s: dropping %zu bytes of data\n",
+					__func__, datalen - count);
+			hvc_kick();
+			break;
+
+		case DIBS_CON_RESIZE_MSG:
+			winsize = msg->rm.ws;
+			hvc_resize(dcpd->hvc, winsize);
+			break;
+
+		default:
+			break;
+		}
+		dibs_pbd_msg_ack(&peer_dev->rb);
+	}
+}
+
+static void dibs_con_peer_dev_disconnect(struct dibs_peer_device *dpd)
+{
+	struct dibs_con_peer_dev *dcpd = dpd->drv_priv;
+
+	if (dcpd->ring_registered) {
+		if (dibs_pbd_unregister_ring_buffer(dcpd->dpd))
+			pr_err("%s: dibs_pbd_unregister_ring_buffer() returned with error\n",
+			       __func__);
+		dcpd->ring_registered = false;
+	}
+}
+
+/*
+ * dibs_con_get_chars()
+ * This function is called by the HVC thread to receive tty data.
+ *
+ * @vtermmno: HVC virtual terminal number
+ * @buf: buffer offered by HVC to store data. By calling hvc_kick()
+ *	this function is called again with a new buffer to transfer more data
+ *	to HVC.
+ */
+static ssize_t dibs_con_get_chars(u32 vtermno, u8 *buf, size_t count)
+{
+	struct dibs_con_peer_dev *dcpd = NULL;
+
+	dcpd = hvc_peer_devs[vtermno];
+	if (!dcpd)
+		return 0;
+
+	spin_lock(&dcpd->mr_spinlock);
+	count = kfifo_out(&dcpd->fifo, buf, count);
+	if (!kfifo_is_empty(&dcpd->fifo))
+		hvc_kick();
+
+	spin_unlock(&dcpd->mr_spinlock);
+	return count;
+}
+
+/*
+ * dibs_con_put_chars()
+ *
+ * This function is called by the HVC framework when new tty
+ * data is available. It enqueues the data in the dibs dmb ring
+ * to transfer it to the other side.
+ */
+static ssize_t dibs_con_put_chars(u32 vtermno, const u8 *buf, size_t count)
+{
+	struct dibs_con_peer_dev *dcpd = NULL;
+	struct dibs_con_peer_dev_rbe dcpdr;
+	size_t max_cnt;
+
+	dcpd = hvc_peer_devs[vtermno];
+	if (!dcpd)
+		return 0;
+	if (!dcpd->connect_msg_enqueued)
+		return 0;
+
+	max_cnt = dibs_pbd_msg_max_payload(&dcpd->dpd->rb);
+	if (!max_cnt)
+		return 0;
+
+	count = min(count, max_cnt);
+	dcpdr.pdm.hdr.version = DIBS_CON_PEER_DEV_HDR_VERSION;
+	dcpdr.pdm.hdr.type = DIBS_CON_DATA_MSG;
+	dcpdr.pdm.hdr.datalen = count;
+	memcpy(&dcpdr.pdm.tdm.data, buf, count);
+	if (!dibs_pbd_msg_send(&dcpd->dpd->rb, &dcpdr.pdm.hdr))
+		return (ssize_t)count;
+	else
+		return 0;
+}
+
+static const struct hv_ops dibs_con_ops = {
+	.get_chars = dibs_con_get_chars,
+	.put_chars = dibs_con_put_chars,
+};
+
+static int dibs_con_alloc_hvc(struct dibs_con_peer_dev *dcpd)
+{
+	int vtermno;
+	int ret;
+
+	for (vtermno = 0; vtermno < DIBS_CON_MAX_VTERMNOS; vtermno++) {
+		ret = hvc_instantiate(vtermno, vtermno, &dibs_con_ops);
+		if (ret)
+			continue;
+		dcpd->hvc = hvc_alloc(vtermno, vtermno, &dibs_con_ops,
+				      DIBS_CON_MAX_PUT_CHARS);
+		if (IS_ERR(dcpd->hvc)) {
+			ret = PTR_ERR(dcpd->hvc);
+			pr_err("%s: Cannot allocate HVC devices\n", __func__);
+			dcpd->hvc = NULL;
+			return ret;
+		}
+
+		hvc_peer_devs[vtermno] = dcpd;
+		dcpd->vtermno = vtermno;
+		dcpd->hvc->irq_requested = 1;
+		return 0;
+	}
+	return -EBUSY;
+}
+
+static ssize_t termid_show(struct device *dev, struct device_attribute *attr,
+			   char *buf)
+{
+	struct dibs_peer_device *dpd;
+	struct dibs_con_peer_dev *dcpd;
+
+	dpd = to_dibs_peer_device(dev);
+	dcpd = dpd->drv_priv;
+	if (dcpd->dpd == dpd)
+		return sysfs_emit(buf, "hvc%d\n", dcpd->hvc->index);
+	return 0;
+}
+static DEVICE_ATTR_RO(termid);
+
+static int dibs_con_peer_dev_probe(struct dibs_peer_device *dpd)
+{
+	struct dibs_con_peer_dev *dcpd;
+	int ret;
+
+	dpd->type = PEER_DEV_TYPE_CONSOLE;
+
+	dcpd = kzalloc_obj(*dcpd, GFP_KERNEL);
+	if (!dcpd)
+		return -ENOMEM;
+
+	dpd->drv_priv = dcpd;
+	dcpd->dpd = dpd;
+
+	ret = kfifo_alloc(&dcpd->fifo, DIBS_CON_PEER_DEV_BUFFER_SIZE,
+			  GFP_KERNEL);
+	if (ret)
+		goto err_alloc_fifo;
+
+	ret = dibs_con_alloc_hvc(dcpd);
+	if (ret)
+		goto err_alloc_hvc;
+	ret = sysfs_create_file(&dpd->dev.kobj, &dev_attr_termid.attr);
+	if (ret)
+		pr_warn("Adding hvc number to sysfs failed.\n");
+
+	ret = dibs_pbd_register_ring_buffer("dibs_con", dpd,
+					    DIBS_CON_PEER_DEV_RING_SIZE);
+	if (ret)
+		goto err_register_rb;
+
+	dcpd->ring_registered = true;
+	dibs_ring_set_rdmb_tok(&dpd->rb, dpd->rdmb_tok);
+
+	return 0;
+
+err_register_rb:
+	sysfs_remove_file(&dpd->dev.kobj, &dev_attr_termid.attr);
+	hvc_peer_devs[dcpd->vtermno] = NULL;
+	hvc_remove(dcpd->hvc);
+err_alloc_hvc:
+	kfifo_free(&dcpd->fifo);
+err_alloc_fifo:
+	dpd->drv_priv = NULL;
+	kfree(dcpd);
+	return ret;
+}
+
+static void dibs_con_peer_dev_remove(struct dibs_peer_device *dpd)
+{
+	struct dibs_con_peer_dev *dcpd = dpd->drv_priv;
+
+	if (dcpd->hvc)
+		hvc_remove(dcpd->hvc);
+	hvc_peer_devs[dcpd->vtermno] = NULL;
+	kfifo_free(&dcpd->fifo);
+
+	if (dcpd->ring_registered) {
+		dibs_pbd_unregister_ring_buffer(dcpd->dpd);
+		dcpd->ring_registered = false;
+	}
+
+	sysfs_remove_file(&dpd->dev.kobj, &dev_attr_termid.attr);
+
+	kfree(dcpd);
+}
+
+static int dibs_con_peer_dev_reconnect(struct dibs_peer_device *dpd)
+{
+	struct dibs_con_peer_dev *dcpd;
+	int ret;
+
+	dcpd = dpd->drv_priv;
+
+	dcpd->connect_msg_enqueued = 0;
+	if (dcpd->ring_registered) {
+		dibs_pbd_unregister_ring_buffer(dcpd->dpd);
+		dcpd->ring_registered = false;
+	}
+
+	ret = dibs_pbd_register_ring_buffer("dibs_con", dcpd->dpd,
+					    DIBS_CON_PEER_DEV_RING_SIZE);
+	if (ret) {
+		pr_err("%s: failed to register ring buffer\n", __func__);
+		return ret;
+	}
+	dcpd->ring_registered = true;
+	dibs_ring_set_rdmb_tok(&dcpd->dpd->rb, dcpd->dpd->rdmb_tok);
+
+	return ret;
+}
+
+static struct dibs_peer_dev_driver dibs_con_pdd = {
+	.drv = {
+		.owner = THIS_MODULE,
+		.name = "dibs_con",
+	},
+	.type = DIBS_CLIENT_TYPE_DIBS_CONSOLE,
+	.probe = dibs_con_peer_dev_probe,
+	.remove = dibs_con_peer_dev_remove,
+	.reconnect = dibs_con_peer_dev_reconnect,
+	.disconnect = dibs_con_peer_dev_disconnect,
+
+	.handle_irq = dibs_con_peer_dev_handle_irq,
+};
+
+static int __init dibs_con_init(void)
+{
+	int ret;
+
+	ret = dibs_pbd_driver_register(&dibs_con_pdd);
+
+	return ret;
+}
+
+static void __exit dibs_con_exit(void)
+{
+	dibs_pbd_driver_unregister(&dibs_con_pdd);
+}
+
+module_init(dibs_con_init);
+module_exit(dibs_con_exit);
+MODULE_LICENSE("GPL");
+MODULE_AUTHOR("IBM Corporation");
diff --git a/include/linux/dibs.h b/include/linux/dibs.h
index 8c24f7b33e6f..64e1f5505de5 100644
--- a/include/linux/dibs.h
+++ b/include/linux/dibs.h
@@ -572,6 +572,10 @@ void dibs_dev_del(struct dibs_dev *dibs);
 #define PEER_DEV_DISCONNECTED 0
 #define PEER_DEV_CONNECTED 1
 
+#define DIBS_CLIENT_TYPE_DIBS_CONSOLE 1
+
+#define PEER_DEV_TYPE_CONSOLE 1
+
 struct dibs_peer_device {
 	u8 id;
 	uuid_t rgid;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 4/4] dibs: introduce dibs block device peer device driver
  2026-10-08  8:08 [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver Julian Ruess
                   ` (2 preceding siblings ...)
  2026-10-08  8:08 ` [PATCH 3/4] dibs: introduce dibs console peer device driver Julian Ruess
@ 2026-10-08  8:08 ` Julian Ruess
  2026-10-08  8:30   ` sashiko-bot
  3 siblings, 1 reply; 10+ messages in thread
From: Julian Ruess @ 2026-10-08  8:08 UTC (permalink / raw)
  To: schnelle, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, julianr,
	linux390-list, linux-s390

Introduce a virtual block device driver that makes use of the newly
added dibs peer device bus.

This allows to access resources like e.g. an ISO file hosted on another
system via dibs.

If a dibs peer device of type block is announced to this instance, this
driver is probed via uevenvt.

If a backing resource is provided by the console server, it can be
accessed via '/dev/dibsblkX'.

Co-developed-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Tobias Schumacher <ts@linux.ibm.com>
Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
---
 arch/s390/configs/debug_defconfig |   1 +
 arch/s390/configs/defconfig       |   1 +
 drivers/dibs/Kconfig              |  10 +
 drivers/dibs/Makefile             |   1 +
 drivers/dibs/dibs_blk.c           | 798 ++++++++++++++++++++++++++++++++++++++
 include/linux/dibs.h              |   2 +
 6 files changed, 813 insertions(+)

diff --git a/arch/s390/configs/debug_defconfig b/arch/s390/configs/debug_defconfig
index c4acf13aeaba..6d0a84fd8d30 100644
--- a/arch/s390/configs/debug_defconfig
+++ b/arch/s390/configs/debug_defconfig
@@ -135,6 +135,7 @@ CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
 CONFIG_DIBS_PEER_BUS_DRV=y
 CONFIG_DIBS_CONSOLE=m
+CONFIG_DIBS_BLOCK=m
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/arch/s390/configs/defconfig b/arch/s390/configs/defconfig
index 344724c56917..9eeecb5f59fa 100644
--- a/arch/s390/configs/defconfig
+++ b/arch/s390/configs/defconfig
@@ -126,6 +126,7 @@ CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
 CONFIG_DIBS_PEER_BUS_DRV=y
 CONFIG_DIBS_CONSOLE=m
+CONFIG_DIBS_BLOCK=m
 CONFIG_XDP_SOCKETS=y
 CONFIG_XDP_SOCKETS_DIAG=m
 CONFIG_INET=y
diff --git a/drivers/dibs/Kconfig b/drivers/dibs/Kconfig
index a0ed49977274..e99bb4ea0a9a 100644
--- a/drivers/dibs/Kconfig
+++ b/drivers/dibs/Kconfig
@@ -43,3 +43,13 @@ config DIBS_CONSOLE
 	  virtual console communication across systems in a dibs fabric. A
 	  device of type /dev/hvcN will be created for each peer device that is
 	  detected.
+
+config DIBS_BLOCK
+	def_tristate m
+	prompt "Support for dibs block peer devices"
+	depends on BLOCK && DIBS && DIBS_PEER_BUS_DRV
+	help
+	  Select this option if you want to enable the dibs block device
+	  support for dibs peer devices. This peer device driver operates
+	  on the dibs peer bus and provides block device access to remote
+	  storage across systems connected via the dibs fabric.
diff --git a/drivers/dibs/Makefile b/drivers/dibs/Makefile
index a39a67020e05..5327e4c24efa 100644
--- a/drivers/dibs/Makefile
+++ b/drivers/dibs/Makefile
@@ -7,4 +7,5 @@ dibs-y += dibs_main.o
 obj-$(CONFIG_DIBS) += dibs.o
 dibs-$(CONFIG_DIBS_LO) += dibs_loopback.o
 dibs-$(CONFIG_DIBS_PEER_BUS_DRV) += dibs_peer_bus_drv.o
+obj-$(CONFIG_DIBS_BLOCK) += dibs_blk.o
 obj-$(CONFIG_DIBS_CONSOLE) += dibs_console.o
diff --git a/drivers/dibs/dibs_blk.c b/drivers/dibs/dibs_blk.c
new file mode 100644
index 000000000000..94bbd8d37b8a
--- /dev/null
+++ b/drivers/dibs/dibs_blk.c
@@ -0,0 +1,798 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * dibs block device driver.
+ *
+ * This driver provides block device access using
+ * dibs peer devices of type blk.
+ *
+ * Copyright IBM Corp.
+ */
+
+#include <linux/blk-mq.h>
+#include <linux/blkdev.h>
+#include <linux/blkpg.h>
+#include <linux/dibs.h>
+#include <linux/fs.h>
+#include <linux/genalloc.h>
+#include <linux/init.h>
+#include <linux/kernel.h>
+#include <linux/module.h>
+#include <linux/printk.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/stringify.h>
+#include <linux/uaccess.h>
+#include <linux/vmalloc.h>
+
+MODULE_ALIAS("dibs:peer_dev:" __stringify(PEER_DEV_TYPE_BLK));
+
+#define DIBS_BLK_SECTOR_SIZE 512ULL
+
+#define DIBS_BLK_RQ_TIMEOUT (5 * HZ)
+
+static unsigned int dibs_blk_max_timeout = 5;
+module_param(dibs_blk_max_timeout, uint, 0644);
+MODULE_PARM_DESC(dibs_blk_max_timeout,
+		 "Maximum timeout for dibs remote in seconds");
+
+enum { MAX_GPT_PARTITIONS = 128 };
+static DEFINE_MUTEX(blkpd_mutex);
+static int dibs_blk_idx;
+static LIST_HEAD(blkpd_list);
+
+enum { DIBS_BLK_PEER_DEV_HDR_VERSION = 1 };
+enum {
+	DIBS_BLK_CONNECT_MSG = 1,
+	DIBS_BLK_REQ_MSG = 2,
+	DIBS_BLK_DATA_MSG = 3,
+	DIBS_BLK_MEDIA_CHANGE = 4,
+	DIBS_BLK_DISK_DETACH = 5
+};
+
+enum {
+	DIBS_BLK_MEDIA_RECONNECT = 0x1,
+};
+
+/* DMB_SIZE must be at least (RING_SIZE * MAX_RBE_SIZE) + MAX_RBE_SIZE */
+/* and a multiple of 4k */
+#define DIBS_BLK_PEER_DEV_BLK_RING_SIZE (512 * 1024)
+#define DIBS_BLK_PEER_DEV_DMB_RING_SIZE 32
+#define DIBS_BLK_PEER_DEV_MAX_RBE_SIZE (28 * 1024)
+
+/* default sector size: 512 byte */
+#define DIBS_BLK_MAX_HW_SECTORS 54
+#define DIBS_BLK_MAX_HW_SECTORS_BYTES (DIBS_BLK_MAX_HW_SECTORS * 512) /* 8k */
+
+static struct queue_limits dibs_blk_lim = {
+	/* 512 byte sector size to be compatible with ISO 9660. */
+	.logical_block_size = 1 << 9,
+	.max_hw_sectors = DIBS_BLK_MAX_HW_SECTORS,
+};
+
+struct dibs_blk_blk_req_msg {
+	u64 offset;
+	u64 len;
+} __packed;
+
+struct dibs_blk_data_msg {
+	u64 total_bytes_of_req;
+	u16 datalen;
+	u8 data[];
+} __packed;
+
+struct dibs_blk_connect_msg {
+	u64 dmb_tok;
+} __packed;
+
+struct dibs_blk_media_change_msg {
+	u64 size;
+	u64 hash;
+	u8 flags;
+	u8 data[];
+} __packed;
+
+struct dibs_blk_peer_dev_msg {
+	struct dibs_msg_hdr hdr;
+	union {
+		struct dibs_blk_connect_msg cm;
+		struct dibs_blk_blk_req_msg blkrm;
+		struct dibs_blk_data_msg blkdm;
+		struct dibs_blk_media_change_msg blkmcm;
+	};
+};
+
+struct dibs_blk_peer_dev_rbe {
+	union {
+		struct dibs_blk_peer_dev_msg pdm;
+		u8 reserved[DIBS_BLK_PEER_DEV_MAX_RBE_SIZE];
+	};
+};
+
+struct dibs_blk_blkdev {
+	struct blk_mq_tag_set tag_set;
+	struct gendisk *gd;
+	bool quiesced;
+	int major;
+};
+
+struct dibs_blk_peer_dev_data_state {
+	u64 already_copied_bytes;
+	u8 data[DIBS_BLK_MAX_HW_SECTORS_BYTES];
+};
+
+struct dibs_blk_peer_dev {
+	struct dibs_blk_peer_dev_rbe rbe_to_send;
+	struct dibs_blk_peer_dev_data_state ds;
+	struct work_struct media_change_work;
+	struct work_struct media_gone_work;
+	struct work_struct proc_recv_msg_work;
+	unsigned long first_timeout;
+	unsigned long last_completion;
+	unsigned long max_wait;
+	struct dibs_peer_device *dpd;
+	int connect_msg_enqueued;
+	bool ring_registered;
+	bool tag_set_allocated;
+	bool blkdev_registered;
+	bool list_added;
+	struct dibs_blk_blkdev blk;
+	struct list_head list;
+	u8 device_add_disk;
+	u64 media_size;
+	/*
+	 * protects current_request, ds.already_copied_bytes, first_timeout,
+	 * media_size, media_hash and rbe_to_send
+	 */
+	spinlock_t req_lock;
+	struct request *current_request;
+	u64 media_hash;
+	int device_index;
+};
+
+static blk_status_t dibs_blk_queue_rq(struct blk_mq_hw_ctx *hctx,
+				      const struct blk_mq_queue_data *bd)
+{
+	struct request *req = bd->rq;
+	struct dibs_blk_peer_dev *blkpd;
+	unsigned long flags;
+	u64 offset;
+	int rc;
+
+	offset = blk_rq_pos(req) * DIBS_BLK_SECTOR_SIZE;
+
+	blkpd = bd->rq->q->disk->private_data;
+
+	if (!blkpd->device_add_disk || !blkpd->ring_registered)
+		return BLK_STS_IOERR;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+
+	if (!blkpd->media_size) {
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		return BLK_STS_IOERR;
+	}
+
+	if (offset + blk_rq_bytes(req) > blkpd->media_size) {
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		pr_err("%s: out of bounds request\n", __func__);
+		return BLK_STS_IOERR;
+	}
+
+	/* Send request to remote side */
+	blkpd->rbe_to_send.pdm.hdr.version = DIBS_BLK_PEER_DEV_HDR_VERSION;
+	blkpd->rbe_to_send.pdm.hdr.type = DIBS_BLK_REQ_MSG;
+	blkpd->rbe_to_send.pdm.hdr.datalen =
+		sizeof(struct dibs_blk_blk_req_msg);
+	blkpd->rbe_to_send.pdm.blkrm.offset = offset;
+	blkpd->rbe_to_send.pdm.blkrm.len = blk_rq_bytes(req);
+
+	blkpd->current_request = req;
+	blk_mq_start_request(req);
+
+	rc = dibs_pbd_msg_send(&blkpd->dpd->rb,
+			       (struct dibs_msg_hdr *)&blkpd->rbe_to_send);
+	if (rc)
+		blkpd->current_request = NULL;
+
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	if (!rc)
+		return BLK_STS_OK;
+
+	if (rc == -ENOSPC || rc == -EAGAIN)
+		return BLK_STS_RESOURCE;
+
+	return BLK_STS_IOERR;
+}
+
+static struct request *dibs_blk_claim_request(struct dibs_blk_peer_dev *blkpd)
+{
+	unsigned long flags;
+	struct request *req;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	req = blkpd->current_request;
+	blkpd->current_request = NULL;
+	blkpd->ds.already_copied_bytes = 0;
+	blkpd->first_timeout = 0;
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	return req;
+}
+
+static void dibs_blk_requeue_inflight(struct dibs_blk_peer_dev *blkpd)
+{
+	struct request *req = dibs_blk_claim_request(blkpd);
+
+	if (!req)
+		return;
+
+	blk_mq_requeue_request(req, false);
+}
+
+static void dibs_blk_unquiesce(struct dibs_blk_peer_dev *blkpd)
+{
+	if (!blkpd->blk.quiesced)
+		return;
+
+	blkpd->blk.quiesced = false;
+	if (!blkpd->blk.gd)
+		return;
+
+	blk_mq_unquiesce_queue(blkpd->blk.gd->queue);
+	blk_mq_kick_requeue_list(blkpd->blk.gd->queue);
+}
+
+static void dibs_blk_rescan_partitions(struct dibs_blk_peer_dev *blkpd)
+{
+	struct gendisk *gd = blkpd->blk.gd;
+	int ret;
+
+	mutex_lock(&gd->open_mutex);
+	ret = bdev_disk_changed(gd, false);
+	mutex_unlock(&gd->open_mutex);
+
+	if (ret == -EBUSY)
+		set_bit(GD_NEED_PART_SCAN, &gd->state);
+	else if (ret)
+		pr_warn("%s: partition scan of %s failed: %d\n", __func__,
+			gd->disk_name, ret);
+}
+
+static void dibs_blk_update_capacity(struct dibs_blk_peer_dev *blkpd)
+{
+	unsigned long flags;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	set_capacity(blkpd->blk.gd, blkpd->media_size / DIBS_BLK_SECTOR_SIZE);
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+}
+
+static void dibs_blk_media_gone_work(struct work_struct *work)
+{
+	struct dibs_blk_peer_dev *blkpd =
+		container_of(work, struct dibs_blk_peer_dev, media_gone_work);
+
+	if (!blkpd->device_add_disk)
+		return;
+
+	dibs_blk_update_capacity(blkpd);
+
+	dibs_blk_unquiesce(blkpd);
+	disk_force_media_change(blkpd->blk.gd);
+}
+
+static enum blk_eh_timer_return dibs_blk_timeout(struct request *req)
+{
+	struct dibs_blk_peer_dev *blkpd = req->q->disk->private_data;
+	unsigned long flags;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+
+	if (blkpd->current_request != req) {
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		return BLK_EH_DONE;
+	}
+
+	if (!blkpd->first_timeout) {
+		blkpd->first_timeout = jiffies;
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		return BLK_EH_RESET_TIMER;
+	}
+
+	if (!time_after(jiffies, blkpd->first_timeout + blkpd->max_wait)) {
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		return BLK_EH_RESET_TIMER;
+	}
+
+	blkpd->current_request = NULL;
+	blkpd->ds.already_copied_bytes = 0;
+	blkpd->first_timeout = 0;
+
+	blkpd->media_size = 0;
+	blkpd->media_hash = 0;
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	blk_mq_end_request(req, BLK_STS_IOERR);
+	schedule_work(&blkpd->media_gone_work);
+
+	return BLK_EH_DONE;
+}
+
+static bool dibs_blk_cancel_request(struct request *req, void *data)
+{
+	struct dibs_blk_peer_dev *blkpd = data;
+	unsigned long flags;
+	bool claimed;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	claimed = blkpd->current_request == req;
+	if (claimed) {
+		blkpd->current_request = NULL;
+		blkpd->ds.already_copied_bytes = 0;
+		blkpd->first_timeout = 0;
+	}
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	if (claimed)
+		blk_mq_end_request(req, BLK_STS_IOERR);
+
+	return true;
+}
+
+static struct blk_mq_ops dibs_blk_mq_ops = {
+	.queue_rq = dibs_blk_queue_rq,
+	.timeout = dibs_blk_timeout,
+};
+
+static const struct block_device_operations dibs_blk_ops = {
+	.owner = THIS_MODULE,
+};
+
+static void dibs_blk_peer_dev_disconnect(struct dibs_peer_device *dpd)
+{
+	struct dibs_blk_peer_dev *blkpd = dpd->drv_priv;
+
+	if (blkpd->blk.gd && !blkpd->blk.quiesced) {
+		blk_mq_quiesce_queue(blkpd->blk.gd->queue);
+		blkpd->blk.quiesced = true;
+	}
+
+	if (blkpd->ring_registered) {
+		if (dibs_pbd_unregister_ring_buffer(blkpd->dpd))
+			pr_err("%s: dibs_pbd_unregister_ring_buffer() returned with error\n",
+			       __func__);
+		blkpd->ring_registered = false;
+	}
+
+	cancel_work_sync(&blkpd->proc_recv_msg_work);
+	cancel_work_sync(&blkpd->media_gone_work);
+	/* Peer will not respond after disconnect, requeue inflight request. */
+	dibs_blk_requeue_inflight(blkpd);
+}
+
+static int dibs_blk_peer_dev_reconnect(struct dibs_peer_device *dpd)
+{
+	struct dibs_blk_peer_dev *blkpd;
+	int ret;
+
+	blkpd = dpd->drv_priv;
+	blkpd->connect_msg_enqueued = 0;
+
+	if (blkpd->ring_registered) {
+		dibs_pbd_unregister_ring_buffer(blkpd->dpd);
+		blkpd->ring_registered = false;
+	}
+	cancel_work_sync(&blkpd->proc_recv_msg_work);
+	cancel_work_sync(&blkpd->media_gone_work);
+	dibs_blk_requeue_inflight(blkpd);
+	ret = dibs_pbd_register_ring_buffer("dibs_blk", blkpd->dpd,
+					    DIBS_BLK_PEER_DEV_BLK_RING_SIZE);
+	if (ret) {
+		pr_err("%s: failed to register ring buffer\n", __func__);
+		return ret;
+	}
+
+	blkpd->ring_registered = true;
+	dibs_ring_set_rdmb_tok(&blkpd->dpd->rb, blkpd->dpd->rdmb_tok);
+	dibs_blk_unquiesce(blkpd);
+	schedule_work(&blkpd->proc_recv_msg_work);
+
+	return 0;
+}
+
+static void dibs_blk_peer_dev_send_connect_msg(struct dibs_peer_device *dpd)
+{
+	struct dibs_blk_peer_dev *blkpd;
+	unsigned long flags;
+	int ret;
+
+	blkpd = dpd->drv_priv;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	blkpd->rbe_to_send.pdm.hdr.version = DIBS_BLK_PEER_DEV_HDR_VERSION;
+	blkpd->rbe_to_send.pdm.hdr.type = DIBS_BLK_CONNECT_MSG;
+	blkpd->rbe_to_send.pdm.hdr.datalen = sizeof(dpd->rb.dmb.dmb_tok);
+	blkpd->rbe_to_send.pdm.cm.dmb_tok = dpd->rb.dmb.dmb_tok;
+	ret = dibs_pbd_msg_send(&dpd->rb,
+				(struct dibs_msg_hdr *)&blkpd->rbe_to_send);
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	if (!ret)
+		blkpd->connect_msg_enqueued = 1;
+}
+
+static void dibs_blk_handle_data_msg(struct dibs_blk_peer_dev *blkpd,
+				     void *current_data, u64 data_len,
+				     u64 total_bytes_of_request)
+{
+	struct req_iterator iter;
+	size_t len, copied = 0;
+	struct bio_vec bvec;
+	unsigned long flags;
+	struct request *req;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+
+	req = blkpd->current_request;
+	if (!req) {
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		pr_warn("Received response for unknown request\n");
+		return;
+	}
+
+	blkpd->last_completion = jiffies;
+	blkpd->first_timeout = 0;
+
+	if (total_bytes_of_request != blk_rq_bytes(req)) {
+		pr_err("%s: total_bytes mismatch: got %llu, expected %u\n",
+		       __func__, total_bytes_of_request, blk_rq_bytes(req));
+		goto err;
+	}
+
+	if (data_len > DIBS_BLK_MAX_HW_SECTORS_BYTES)
+		goto err;
+
+	if (blkpd->ds.already_copied_bytes + data_len >
+	    total_bytes_of_request) {
+		pr_err("%s: data overflow: already %llu + new %llu > total %llu\n",
+		       __func__, blkpd->ds.already_copied_bytes, data_len,
+		       total_bytes_of_request);
+		goto err;
+	}
+
+	memcpy(blkpd->ds.data + blkpd->ds.already_copied_bytes, current_data,
+	       data_len);
+	blkpd->ds.already_copied_bytes += data_len;
+
+	if (blkpd->ds.already_copied_bytes < total_bytes_of_request) {
+		/* partial request */
+		spin_unlock_irqrestore(&blkpd->req_lock, flags);
+		return;
+	}
+
+	blkpd->ds.already_copied_bytes = 0;
+	blkpd->current_request = NULL;
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	rq_for_each_segment(bvec, req, iter) {
+		void *buffer = kmap_local_page(bvec.bv_page) + bvec.bv_offset;
+
+		len = min((size_t)bvec.bv_len, total_bytes_of_request - copied);
+		memcpy(buffer, blkpd->ds.data + copied, len);
+		kunmap_local(buffer);
+		copied += len;
+		if (copied >= total_bytes_of_request)
+			break;
+	}
+	blk_mq_end_request(req, BLK_STS_OK);
+	return;
+
+err:
+	blkpd->ds.already_copied_bytes = 0;
+	blkpd->current_request = NULL;
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+	blk_mq_end_request(req, BLK_STS_IOERR);
+}
+
+static bool dibs_blk_set_media(struct dibs_blk_peer_dev *blkpd, u64 media_size,
+			       u64 media_hash, u8 media_flags)
+{
+	unsigned long flags;
+	bool changed = true;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	if (blkpd->device_add_disk && media_size == blkpd->media_size &&
+	    media_hash == blkpd->media_hash &&
+	    (media_flags & DIBS_BLK_MEDIA_RECONNECT)) {
+		changed = false;
+	} else {
+		blkpd->media_size = media_size;
+		blkpd->media_hash = media_hash;
+	}
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	return changed;
+}
+
+static void dibs_blk_media_change_work(struct work_struct *work)
+{
+	struct dibs_blk_peer_dev *blkpd;
+	int ret;
+
+	blkpd = container_of(work, struct dibs_blk_peer_dev, media_change_work);
+
+	if (!blkpd->device_add_disk) {
+		blkpd->blk.gd = blk_mq_alloc_disk(&blkpd->blk.tag_set,
+						  &dibs_blk_lim, &blkpd->blk);
+		if (IS_ERR(blkpd->blk.gd)) {
+			pr_err("%s: failed to allocate gendisk: %pe\n",
+			       __func__, blkpd->blk.gd);
+			blkpd->blk.gd = NULL;
+			return;
+		}
+
+		blkpd->blk.gd->private_data = blkpd;
+		blkpd->blk.gd->fops = &dibs_blk_ops;
+		blkpd->blk.gd->major = blkpd->blk.major;
+		blkpd->blk.gd->first_minor =
+			blkpd->device_index * MAX_GPT_PARTITIONS;
+		blkpd->blk.gd->minors = MAX_GPT_PARTITIONS;
+		blkpd->blk.gd->events = DISK_EVENT_MEDIA_CHANGE;
+
+		set_disk_ro(blkpd->blk.gd, 1);
+		snprintf(blkpd->blk.gd->disk_name, 10, "dibs_blk%d",
+			 blkpd->device_index);
+
+		dibs_blk_update_capacity(blkpd);
+		blkpd->device_add_disk = 1;
+
+		ret = device_add_disk(&blkpd->dpd->dev, blkpd->blk.gd, NULL);
+		if (ret) {
+			pr_err("%s: device_add_disk failed: %d\n", __func__,
+			       ret);
+			blkpd->device_add_disk = 0;
+			put_disk(blkpd->blk.gd);
+			blkpd->blk.gd = NULL;
+			return;
+		}
+	}
+
+	dibs_blk_update_capacity(blkpd);
+	disk_force_media_change(blkpd->blk.gd);
+	dibs_blk_rescan_partitions(blkpd);
+}
+
+static void dibs_blk_handle_disk_detach(struct dibs_blk_peer_dev *blkpd)
+{
+	unsigned long flags;
+
+	if (!blkpd->device_add_disk)
+		return;
+
+	spin_lock_irqsave(&blkpd->req_lock, flags);
+	blkpd->media_size = 0;
+	blkpd->media_hash = 0;
+	spin_unlock_irqrestore(&blkpd->req_lock, flags);
+
+	dibs_blk_update_capacity(blkpd);
+
+	dibs_blk_unquiesce(blkpd);
+	blk_mq_tagset_busy_iter(&blkpd->blk.tag_set, dibs_blk_cancel_request,
+				blkpd);
+	disk_force_media_change(blkpd->blk.gd);
+}
+
+static void dibs_blk_proc_recv_msg_work(struct work_struct *work)
+{
+	struct dibs_peer_device *peer_dev;
+	struct dibs_blk_peer_dev_msg *msg;
+	struct dibs_blk_peer_dev *blkpd;
+	u64 total_bytes_of_request;
+	void *current_data;
+	u64 media_size;
+	u64 media_hash;
+	u64 data_len;
+	u8 media_flags;
+
+	blkpd = container_of(work, struct dibs_blk_peer_dev,
+			     proc_recv_msg_work);
+	peer_dev = blkpd->dpd;
+
+	if (!blkpd->connect_msg_enqueued)
+		dibs_blk_peer_dev_send_connect_msg(peer_dev);
+
+	while ((msg = (struct dibs_blk_peer_dev_msg *)dibs_pbd_msg_recv(&peer_dev->rb))) {
+		switch (msg->hdr.type) {
+		case DIBS_BLK_DATA_MSG:
+			current_data = msg->blkdm.data;
+			data_len = READ_ONCE(msg->blkdm.datalen);
+			total_bytes_of_request =
+				READ_ONCE(msg->blkdm.total_bytes_of_req);
+
+			if (sizeof(msg->blkdm) + data_len >
+			    READ_ONCE(msg->hdr.datalen))
+				break;
+			dibs_blk_handle_data_msg(blkpd, current_data, data_len,
+						 total_bytes_of_request);
+			break;
+
+		case DIBS_BLK_MEDIA_CHANGE:
+			media_size = READ_ONCE(msg->blkmcm.size);
+			media_hash = READ_ONCE(msg->blkmcm.hash);
+			media_flags = READ_ONCE(msg->blkmcm.flags);
+
+			if (dibs_blk_set_media(blkpd, media_size, media_hash,
+					       media_flags))
+				schedule_work(&blkpd->media_change_work);
+			break;
+
+		case DIBS_BLK_DISK_DETACH:
+			dibs_blk_handle_disk_detach(blkpd);
+			break;
+
+		default:
+			pr_warn("Received unknown message type: %d\n",
+				msg->hdr.type);
+			break;
+		}
+		dibs_pbd_msg_ack(&peer_dev->rb);
+	}
+}
+
+static int dibs_blk_peer_dev_probe(struct dibs_peer_device *dpd)
+{
+	struct dibs_blk_peer_dev *blkpd;
+	int ret;
+
+	dpd->type = PEER_DEV_TYPE_BLK;
+
+	blkpd = kzalloc_obj(*blkpd, GFP_KERNEL);
+	if (!blkpd)
+		return -ENOMEM;
+
+	blkpd->dpd = dpd;
+	dpd->drv_priv = blkpd;
+
+	spin_lock_init(&blkpd->req_lock);
+	INIT_LIST_HEAD(&blkpd->list);
+	INIT_WORK(&blkpd->media_change_work, dibs_blk_media_change_work);
+	INIT_WORK(&blkpd->media_gone_work, dibs_blk_media_gone_work);
+	INIT_WORK(&blkpd->proc_recv_msg_work, dibs_blk_proc_recv_msg_work);
+
+	ret = dibs_pbd_register_ring_buffer("dibs_blk", dpd,
+					    DIBS_BLK_PEER_DEV_BLK_RING_SIZE);
+	if (ret) {
+		pr_err("%s: failed to register ring buffer\n", __func__);
+		goto err_free;
+	}
+
+	blkpd->ring_registered = true;
+	dibs_ring_set_rdmb_tok(&dpd->rb, dpd->rdmb_tok);
+	blkpd->blk.major = register_blkdev(0, "dibs_blk");
+	if (blkpd->blk.major < 0) {
+		pr_err("%s: failed to register block device\n", __func__);
+		ret = blkpd->blk.major;
+		goto err_unregister_ring;
+	}
+	blkpd->blkdev_registered = true;
+
+	blkpd->blk.tag_set.ops = &dibs_blk_mq_ops;
+	blkpd->blk.tag_set.nr_hw_queues = 1;
+	blkpd->blk.tag_set.queue_depth = 1;
+	blkpd->blk.tag_set.numa_node = NUMA_NO_NODE;
+	blkpd->blk.tag_set.timeout = DIBS_BLK_RQ_TIMEOUT;
+	blkpd->blk.tag_set.cmd_size = 0;
+	blkpd->blk.tag_set.flags = BLK_MQ_F_TAG_QUEUE_SHARED;
+	blkpd->blk.tag_set.driver_data = &blkpd->blk;
+
+	ret = blk_mq_alloc_tag_set(&blkpd->blk.tag_set);
+	if (ret) {
+		pr_err("%s: blk_mq_alloc_tag_set failed: %d\n", __func__, ret);
+		if (blkpd->blkdev_registered)
+			unregister_blkdev(blkpd->blk.major, "dibs_blk");
+		goto err_unregister_ring;
+	}
+	blkpd->tag_set_allocated = true;
+
+	mutex_lock(&blkpd_mutex);
+	blkpd->device_index = dibs_blk_idx++;
+	list_add(&blkpd->list, &blkpd_list);
+	blkpd->list_added = true;
+	mutex_unlock(&blkpd_mutex);
+
+	blkpd->max_wait = secs_to_jiffies(dibs_blk_max_timeout);
+
+	return 0;
+
+err_unregister_ring:
+	if (blkpd->ring_registered)
+		dibs_pbd_unregister_ring_buffer(dpd);
+err_free:
+	dpd->drv_priv = NULL;
+	kfree(blkpd);
+	return ret;
+}
+
+static void dibs_blk_peer_dev_remove(struct dibs_peer_device *dpd)
+{
+	struct dibs_blk_peer_dev *blkpd = dpd->drv_priv;
+
+	if (blkpd->list_added) {
+		mutex_lock(&blkpd_mutex);
+		list_del(&blkpd->list);
+		mutex_unlock(&blkpd_mutex);
+		blkpd->list_added = false;
+	}
+
+	if (blkpd->ring_registered) {
+		dibs_pbd_unregister_ring_buffer(blkpd->dpd);
+		blkpd->ring_registered = false;
+	}
+
+	cancel_work_sync(&blkpd->proc_recv_msg_work);
+	cancel_work_sync(&blkpd->media_change_work);
+	cancel_work_sync(&blkpd->media_gone_work);
+
+	if (blkpd->device_add_disk) {
+		blk_mark_disk_dead(blkpd->blk.gd);
+		dibs_blk_unquiesce(blkpd);
+		blk_mq_tagset_busy_iter(&blkpd->blk.tag_set,
+					dibs_blk_cancel_request, blkpd);
+		del_gendisk(blkpd->blk.gd);
+		put_disk(blkpd->blk.gd);
+		blkpd->blk.gd = NULL;
+		blkpd->device_add_disk = 0;
+	}
+
+	if (blkpd->tag_set_allocated)
+		blk_mq_free_tag_set(&blkpd->blk.tag_set);
+
+	if (blkpd->blkdev_registered)
+		unregister_blkdev(blkpd->blk.major, "dibs_blk");
+
+	kfree(blkpd);
+}
+
+static void dibs_blk_peer_dev_handle_irq(unsigned int dmbno,
+					 struct dibs_peer_device *peer_dev,
+					 u16 dmbemask)
+{
+	struct dibs_blk_peer_dev *blkpd;
+
+	blkpd = peer_dev->drv_priv;
+	schedule_work(&blkpd->proc_recv_msg_work);
+}
+
+static struct dibs_peer_dev_driver dibs_blk_pdd = {
+	.drv = {
+		.owner = THIS_MODULE,
+		.name = "dibs_blk",
+	},
+	.type = DIBS_CLIENT_TYPE_DIBS_BLK,
+	.probe = dibs_blk_peer_dev_probe,
+	.remove = dibs_blk_peer_dev_remove,
+	.reconnect = dibs_blk_peer_dev_reconnect,
+	.disconnect = dibs_blk_peer_dev_disconnect,
+
+	.handle_irq = dibs_blk_peer_dev_handle_irq,
+};
+
+static int __init dibs_blk_init(void)
+{
+	int ret;
+
+	ret = dibs_pbd_driver_register(&dibs_blk_pdd);
+	return ret;
+}
+
+static void __exit dibs_blk_exit(void)
+{
+	dibs_pbd_driver_unregister(&dibs_blk_pdd);
+}
+
+module_init(dibs_blk_init);
+module_exit(dibs_blk_exit);
+MODULE_LICENSE("GPL");
+MODULE_DESCRIPTION("dibs peer device driver for block devices");
+MODULE_AUTHOR("IBM Corporation");
diff --git a/include/linux/dibs.h b/include/linux/dibs.h
index 64e1f5505de5..2f4ee27c0ea3 100644
--- a/include/linux/dibs.h
+++ b/include/linux/dibs.h
@@ -573,8 +573,10 @@ void dibs_dev_del(struct dibs_dev *dibs);
 #define PEER_DEV_CONNECTED 1
 
 #define DIBS_CLIENT_TYPE_DIBS_CONSOLE 1
+#define DIBS_CLIENT_TYPE_DIBS_BLK 2
 
 #define PEER_DEV_TYPE_CONSOLE 1
+#define PEER_DEV_TYPE_BLK 2
 
 struct dibs_peer_device {
 	u8 id;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* Re: [PATCH 1/4] dibs: introduce ring-buffer send/receive API
  2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
@ 2026-10-08  8:21   ` sashiko-bot
  2026-10-09 20:47   ` Niklas Schnelle
  1 sibling, 0 replies; 10+ messages in thread
From: sashiko-bot @ 2026-10-08  8:21 UTC (permalink / raw)
  To: Julian Ruess
  Cc: Heiko Carstens, linux-s390, Christian Borntraeger,
	Alexander Gordeev, Vasily Gorbik

Thank you for your contribution! Sashiko AI review found 10 potential issue(s) to consider:
- [Critical] dibs: integer underflow in __dibs_ring_space_to_end() bypasses bounds checks
- [High] dibs: data race on ring pointers during peer resynchronization
- [High] dibs: missing peer resynchronization check in dibs_ring_send() and dibs_ring_space_to_end()
- [High] dibs: torn read of peer metadata in dibs_ring_peer_meta() enables out-of-bounds writes
- [High] dibs: transient header send failure in dibs_ring_sync_peer() causes infinite message replay
- [High] dibs: dibs_ring_unregister() lacks state validation causing NULL pointer dereference
- [High] dibs: TOCTOU race in dibs_ring_ack() allows remote peer to block acknowledgments
- [High] dibs: missing memory barrier in dibs_ring_ack() before publishing tail
- [High] dibs: unsafe DMA buffer allocation for wr_ring_meta due to cache-line sharing
- [Medium] dibs: dibs_ring_unregister() masks payload DMB unregistration failure
--

commit c70d1b4a787d1f8149e96425815df0aabbcfd89f
Author: Tobias Schumacher <ts@linux.ibm.com>
Subject: dibs: introduce ring-buffer send/receive API

This commit introduces a centralized API for sending and receiving data between
dibs peers via ring buffers. It manages DMB tokens and read/write pointers,
allowing drivers to use dibs_ring_send(), dibs_ring_recv(), and dibs_ring_ack().

> diff --git a/drivers/dibs/dibs_main.c b/drivers/dibs/dibs_main.c
> index 20c50997a7cf2..e3a33bc42fe88 100644
> --- a/drivers/dibs/dibs_main.c
> +++ b/drivers/dibs/dibs_main.c

[ ... ]

> @@ -245,6 +249,491 @@ void dibs_dev_del(struct dibs_dev *dibs)
> +static void dibs_ring_peer_meta(struct dibs_ring_buffer *rb,
> +				struct dibs_ring_meta *peer)
> +{
> +	peer->meta_dmb_tok = READ_ONCE(rb->rd_ring_meta->meta_dmb_tok);
> +	peer->buff_dmb_tok = READ_ONCE(rb->rd_ring_meta->buff_dmb_tok);
> +	peer->size = READ_ONCE(rb->rd_ring_meta->size);
> +	peer->head = READ_ONCE(rb->rd_ring_meta->head);
> +	peer->tail = READ_ONCE(rb->rd_ring_meta->tail);

[Severity: High]
Does this lockless reading of fields allow torn reads?

If the remote peer updates this structure concurrently via DMA, the CPU can
observe a torn state.

For example, it could observe a newly allocated buff_dmb_tok but an old, larger
size.

This allows __dibs_ring_space_to_end() to compute an overly large available
space, which enables dibs_ring_send() to write past the end of the new remote
buffer.

> +	peer->reserved = 0;
> +}
> +
> +static int dibs_ring_send_hdr(struct dibs_ring_buffer *rb, u64 meta_dmb_tok)
> +{
> +	int res;
> +
> +	if (!meta_dmb_tok)
> +		return 0;
> +
> +	res = rb->dibs->ops->move_data(rb->dibs, meta_dmb_tok, 0, true, 0,
> +				       &rb->wr_ring_meta,
> +				       sizeof(struct dibs_ring_meta));

[Severity: High]
Is it safe to pass the embedded wr_ring_meta struct to the device layer as a
DMA data buffer?

Because struct dibs_ring_meta wr_ring_meta is embedded directly within
struct dibs_ring_buffer without ____cacheline_aligned, it shares cache lines
with frequently modified fields.

When move_data() maps it for DMA, the required cache invalidation or flushing
operations could corrupt the adjacent fields or lose concurrent CPU writes.

[ ... ]

> +int dibs_ring_unregister(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_ring_meta peer;
> +	int ret;
> +
> +	rb->wr_ring_meta.meta_dmb_tok = 0;
> +	rb->wr_ring_meta.buff_dmb_tok = 0;
> +	rb->wr_ring_meta.size = 0;
> +	rb->wr_ring_meta.head = 0;
> +	rb->wr_ring_meta.tail = 0;
> +
> +	dibs_ring_peer_meta(rb, &peer);

[Severity: High]
Does this code trigger a NULL pointer dereference if the ring buffer was not
successfully registered?

The function dibs_ring_peer_meta() unconditionally executes
READ_ONCE(rb->rd_ring_meta->meta_dmb_tok).

If dibs_ring_register() previously failed during its first register_dmb() call,
rb->rd_ring_meta will remain NULL.

> +	if (peer.meta_dmb_tok) {
> +		ret = dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
> +		if (ret)
> +			pr_warn("%s: failed to send header update: %d\n",
> +				__func__, ret);
> +	}
> +
> +	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->dmb);
> +	if (ret)
> +		pr_warn("%s: failed to unregister payload dmb: %d\n", __func__,
> +			ret);
> +
> +	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->rd_meta_dmb);

[Severity: Medium]
Does this unintentionally mask a failure to unregister the payload DMB?

The error code in 'ret' from the first unregister_dmb() operation is
immediately clobbered by the second unregister_dmb() call.

If the payload DMB fails to unregister, the function will still return 0.

[ ... ]

> +static void dibs_ring_sync_peer(struct dibs_ring_buffer *rb,
> +				const struct dibs_ring_meta *peer)
> +{
> +	if (!peer->buff_dmb_tok) {
> +		rb->wr_ring_meta.tail = 0;
> +		rb->peer_buff_dmb_tok = 0;
> +		rb->hdr_sent = false;
> +		return;
> +	}
> +
> +	rb->wr_ring_meta.head = 0;
> +	rb->wr_ring_meta.tail = 0;
> +
> +	if (dibs_ring_send_hdr(rb, peer->meta_dmb_tok)) {
> +		pr_warn("%s(%s): failed to send header after peer restart\n",
> +			__func__, rb->name);
> +		return;
> +	}

[Severity: High]
Can a transient header send failure here trap the receiver in an infinite
message replay loop?

If dibs_ring_send_hdr() fails, it returns early without updating
rb->peer_buff_dmb_tok, but the local tail pointer is already zeroed.

The caller (dibs_ring_recv) proceeds to evaluate data with the zeroed tail.

On the next call to dibs_ring_recv(), dibs_ring_peer_replaced() evaluates to
true again, zeroing the tail again and repeatedly processing the exact same
message.

[ ... ]

> +static size_t __dibs_ring_space_to_end(struct dibs_ring_buffer *rb,
> +				       const struct dibs_ring_meta *peer)
> +{
> +	if (!dibs_ring_peer_ready(peer))
> +		return 0;
> +
> +	return CIRC_SPACE_TO_END(rb->wr_ring_meta.head, peer->tail, peer->size);

[Severity: Critical]
Does this code suffer from an integer underflow that bypasses bounds checks?

The macro performs signed integer math.

If a remote peer maliciously shrinks its size parameter while the local head
pointer is at a large offset, the macro yields a negative integer result.

When cast to the function's size_t return type, this underflows into an
enormous unsigned integer, silently defeating the space check in
dibs_ring_send() and allowing massive out-of-bounds remote memory writes.

[ ... ]

> +size_t dibs_ring_space(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_ring_meta peer;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +
> +	if (dibs_ring_peer_replaced(rb, &peer))
> +		dibs_ring_sync_peer(rb, &peer);

[Severity: High]
Is there a data race on the ring pointers here?

If a remote peer replaces its ring buffer, both the sender thread calling
dibs_ring_space() and the receiver thread calling dibs_ring_cnt() will detect
this and simultaneously call dibs_ring_sync_peer().

Both threads will then concurrently zero out rb->wr_ring_meta.head and tail,
issuing conflicting dibs_ring_send_hdr() DMA operations.

[ ... ]

> +size_t dibs_ring_space_to_end(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_ring_meta peer;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +
> +	return __dibs_ring_space_to_end(rb, &peer);

[Severity: High]
Does this function miss a check for a replaced peer ring buffer?

Unlike dibs_ring_space(), this exported capacity checker silently operates on a
stale local head against a newly replaced peer buffer because it omits checking
dibs_ring_peer_replaced().

[ ... ]

> +u32 dibs_ring_cnt(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_ring_meta peer;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +
> +	if (dibs_ring_peer_replaced(rb, &peer))
> +		dibs_ring_sync_peer(rb, &peer);

[Severity: High]
As mentioned earlier, does this concurrently trigger the exact same
state-mutating synchronization as dibs_ring_space(), leading to a data race?

[ ... ]

> +int dibs_ring_send(struct dibs_ring_buffer *rb, void *msg, u16 size,
> +		   bool notify)
> +{
> +	struct dibs_ring_meta peer;
> +	u32 old_head;
> +	int ret;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +
> +	if (!dibs_ring_peer_ready(&peer))
> +		return -EAGAIN;

[Severity: High]
Does this function also miss a check to synchronize the local head pointer?

Similar to dibs_ring_space_to_end(), this code reads the peer metadata but
fails to check dibs_ring_peer_replaced() before proceeding to write.

It will write payload data into the newly allocated remote buffer at an
incorrect offset using the stale, un-reset local head pointer.

[ ... ]

> +int dibs_ring_ack(struct dibs_ring_buffer *rb, size_t size)
> +{
> +	struct dibs_ring_meta peer;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +
> +	if (size > __dibs_ring_cnt(rb, &peer)) {
> +		pr_warn("%s(%s, %lx): error: tail would overtake head\n",
> +			__func__, rb->name, size);
> +		return -EINVAL;
> +	}

[Severity: High]
Is there a Time-Of-Check to Time-Of-Use vulnerability here?

Dynamically re-evaluating the volatile remote head pointer to gate advancing
the local tail pointer allows a remote peer to block acknowledgments.

A remote peer can manipulate its head pointer to equal the local tail, forcing
__dibs_ring_cnt to 0.

The check fails without advancing the tail, trapping the local receiver in an
infinite loop repeatedly reading the same stale message.

> +
> +	rb->wr_ring_meta.tail = (rb->wr_ring_meta.tail + size) &
> +				(rb->wr_ring_meta.size - 1);
> +
> +	return dibs_ring_send_hdr(rb, peer.meta_dmb_tok);

[Severity: High]
Does this code miss a memory barrier before publishing the tail?

The tail pointer is updated and transmitted to the remote peer without a
preceding smp_mb() or dma_rmb() to enforce ordering against the CPU's payload
reads.

The CPU can reorder the payload memory reads to occur after the tail update is
visible to the remote peer, causing data corruption if the remote peer
overwrites the payload buffer before the reads execute.

> diff --git a/include/linux/dibs.h b/include/linux/dibs.h
> index 0c10c224bccad..30ae03484b58a 100644
> --- a/include/linux/dibs.h
> +++ b/include/linux/dibs.h

[ ... ]

> @@ -67,6 +67,83 @@ struct dibs_dmb {
>  	dma_addr_t dma_addr;
>  };
>  

[ ... ]

> +struct dibs_ring_buffer {
> +	char name[256];
> +	struct dibs_dev *dibs;
> +	int dmb_registered;
> +	bool hdr_sent;
> +	struct dibs_dmb dmb;
> +	/* dmb for the ring metadata this side can read, peer can write */
> +	struct dibs_dmb rd_meta_dmb;
> +	/* ring metadata this side can write, peer can read */
> +	struct dibs_ring_meta wr_ring_meta;

[Severity: High]
As mentioned earlier, does embedding this struct without cacheline alignment
make it unsafe for DMA mapping in dibs_ring_send_hdr()?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008-console_server-v1-0-636b11ebf191@linux.ibm.com?part=1

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH 3/4] dibs: introduce dibs console peer device driver
  2026-10-08  8:08 ` [PATCH 3/4] dibs: introduce dibs console peer device driver Julian Ruess
@ 2026-10-08  8:26   ` sashiko-bot
  0 siblings, 0 replies; 10+ messages in thread
From: sashiko-bot @ 2026-10-08  8:26 UTC (permalink / raw)
  To: Julian Ruess
  Cc: linux-s390, Heiko Carstens, Alexander Gordeev,
	Christian Borntraeger, Vasily Gorbik

Thank you for your contribution! Sashiko AI review found 8 potential issue(s) to consider:
- [High] tty: uninitialized spinlock mr_spinlock in dibs_console
- [High] tty: UAF in IRQ handler due to inverted teardown in dibs_con_peer_dev_remove()
- [High] tty: racy hvc_instantiate() usage during hotplug corrupts global state
- [High] tty: dibs_console: use-after-free of hvc in termid_show()
- [High] dibs: kernel stack information leak in dibs_con_put_chars()
- [High] dibs: TOCTOU double-fetch in IRQ handler leading to OOB read
- [High] dibs: UAF of dcpd in dibs_con_get_chars() racing with device remove
- [Low] tty: sysfs attribute named 'termid' instead of documented 'hvc'
--

commit 1fe9a43373bccffcc93521015b20907f33642f67
Author: Julian Ruess <julianr@linux.ibm.com>

dibs: introduce dibs console peer device driver

This introduces a virtual console device driver over the dibs peer device
bus, allowing tty access across systems in a dibs fabric without requiring
a network setup. A /dev/hvcN device is created for each detected peer.

> diff --git a/drivers/dibs/dibs_console.c b/drivers/dibs/dibs_console.c
> new file mode 100644
> --- /dev/null
> +++ b/drivers/dibs/dibs_console.c

[ ... ]

> +static void dibs_con_peer_dev_handle_irq(unsigned int dmbno,
> +					 struct dibs_peer_device *peer_dev,
> +					 u16 dmbemask)
> +{
> +	struct dibs_con_peer_dev *dcpd;
> +	struct dibs_con_pd_msg *msg;
> +	struct winsize winsize;
> +	size_t count;
> +	u16 datalen;
> +
> +	dcpd = peer_dev->drv_priv;
> +
> +	if (!dcpd->connect_msg_enqueued)
> +		dibs_con_peer_dev_send_connect_msg(peer_dev);
> +
> +	while ((msg = (struct dibs_con_pd_msg *)dibs_pbd_msg_recv(&peer_dev->rb))) {
> +		switch (msg->hdr.type) {
> +		case DIBS_CON_DATA_MSG:
> +			datalen = min_t(u16, READ_ONCE(msg->hdr.datalen),
> +					DIBS_CON_PEER_DEV_DATA_MSG_MAX_DATALEN);
> +			count = kfifo_in(&dcpd->fifo, &msg->tdm, datalen);

[Severity: High]
Can a malicious peer concurrently modify msg->hdr.datalen between the bounds
check in dibs_pbd_msg_recv() and this second read?

The double fetch could allow a larger size to be used, meaning kfifo_in()
would read past the end of the mapped ring buffer memory, resulting in an
out-of-bounds read.

This also applies to the DIBS_CON_RESIZE_MSG case below where msg->rm.ws
is read. If datalen was initially 0, accessing the 8-byte payload can cross
the ring buffer boundary.

[ ... ]

> +static ssize_t dibs_con_get_chars(u32 vtermno, u8 *buf, size_t count)
> +{
> +	struct dibs_con_peer_dev *dcpd = NULL;
> +
> +	dcpd = hvc_peer_devs[vtermno];
> +	if (!dcpd)
> +		return 0;
> +
> +	spin_lock(&dcpd->mr_spinlock);

[Severity: High]
Is the mr_spinlock properly initialized before it is acquired here?

In dibs_con_peer_dev_probe(), dcpd is dynamically allocated using kzalloc_obj().
This zero-initializes the memory, but the lock is never properly initialized
with spin_lock_init(). Acquiring a zero-initialized spinlock will cause kernel
crashes on PREEMPT_RT systems or trigger warnings when lock debugging is enabled.

[Severity: High]
Is this pointer dereference safe from concurrent teardown?

In dibs_con_peer_dev_remove(), dcpd is freed without waiting for the khvcd
polling thread to exit. The hvc_remove() function temporarily takes and drops
hp->lock to clear console metadata, but it does not clear the hp->ops pointers.

Meanwhile, __hvc_poll() inside khvcd drops hp->lock to call cond_resched(),
allowing hvc_remove() to proceed and dibs_con_peer_dev_remove() to free dcpd.
When __hvc_poll() reacquires the lock, it jumps back and unconditionally calls
hp->ops->get_chars(), leading to a use-after-free here.

[ ... ]

> +static ssize_t dibs_con_put_chars(u32 vtermno, const u8 *buf, size_t count)
> +{
> +	struct dibs_con_peer_dev *dcpd = NULL;
> +	struct dibs_con_peer_dev_rbe dcpdr;
> +	size_t max_cnt;
> +
> +	dcpd = hvc_peer_devs[vtermno];
> +	if (!dcpd)
> +		return 0;
> +	if (!dcpd->connect_msg_enqueued)
> +		return 0;
> +
> +	max_cnt = dibs_pbd_msg_max_payload(&dcpd->dpd->rb);
> +	if (!max_cnt)
> +		return 0;
> +
> +	count = min(count, max_cnt);
> +	dcpdr.pdm.hdr.version = DIBS_CON_PEER_DEV_HDR_VERSION;
> +	dcpdr.pdm.hdr.type = DIBS_CON_DATA_MSG;
> +	dcpdr.pdm.hdr.datalen = count;
> +	memcpy(&dcpdr.pdm.tdm.data, buf, count);
> +	if (!dibs_pbd_msg_send(&dcpd->dpd->rb, &dcpdr.pdm.hdr))

[Severity: High]
Does this leak uninitialized stack padding to the remote peer?

The dcpdr struct is allocated on the stack and is not zero-initialized.
The memcpy() only writes count bytes into the data buffer.

When dibs_pbd_msg_send() calculates the transmit size using dibs_pbd_msg_size(),
it rounds the total size up to a 4-byte boundary. This will cause
dibs_ring_send() to read and transmit up to 3 bytes of uninitialized stack
memory adjacent to the valid data over the dibs fabric.

[ ... ]

> +static int dibs_con_alloc_hvc(struct dibs_con_peer_dev *dcpd)
> +{
> +	int vtermno;
> +	int ret;
> +
> +	for (vtermno = 0; vtermno < DIBS_CON_MAX_VTERMNOS; vtermno++) {
> +		ret = hvc_instantiate(vtermno, vtermno, &dibs_con_ops);
> +		if (ret)
> +			continue;

[Severity: High]
Is it safe to use hvc_instantiate() dynamically during device probe?

This function is a lockless API designed strictly for sequential early boot
console initialization. If multiple dibs peer devices are probed concurrently,
they will race to read and write the global vtermnos and cons_ops arrays,
leading to corruption where multiple devices claim the same slot.

Additionally, if the subsequent hvc_alloc() call fails, the claimed slot
is never un-instantiated, permanently leaking the console slot.

[ ... ]

> +static ssize_t termid_show(struct device *dev, struct device_attribute *attr,
> +			   char *buf)
> +{
> +	struct dibs_peer_device *dpd;
> +	struct dibs_con_peer_dev *dcpd;
> +
> +	dpd = to_dibs_peer_device(dev);
> +	dcpd = dpd->drv_priv;
> +	if (dcpd->dpd == dpd)
> +		return sysfs_emit(buf, "hvc%d\n", dcpd->hvc->index);

[Severity: High]
Can this dereference a freed pointer during concurrent device teardown?

In dibs_con_peer_dev_remove(), hvc_remove(dcpd->hvc) is called and can
immediately free the hvc_struct if no references are held. The sysfs file is
not removed via sysfs_remove_file() until later, leaving a window where
concurrent readers to the sysfs file can access freed memory here.

> +	return 0;
> +}
> +static DEVICE_ATTR_RO(termid);

[Severity: Low]
This isn't a bug, but the commit message states the hvc device number can be
accessed by reading /sys/bus/dibs_peer/devices/<rgid>/hvc.

The macro DEVICE_ATTR_RO(termid) defines a file named termid instead,
creating an undocumented API mismatch.

[ ... ]

> +static void dibs_con_peer_dev_remove(struct dibs_peer_device *dpd)
> +{
> +	struct dibs_con_peer_dev *dcpd = dpd->drv_priv;
> +
> +	if (dcpd->hvc)
> +		hvc_remove(dcpd->hvc);
> +	hvc_peer_devs[dcpd->vtermno] = NULL;
> +	kfifo_free(&dcpd->fifo);
> +
> +	if (dcpd->ring_registered) {
> +		dibs_pbd_unregister_ring_buffer(dcpd->dpd);

[Severity: High]
Does this teardown ordering create a use-after-free window in the IRQ handler?

The kfifo resource is freed before the ring buffer (which is the interrupt
source) is unregistered. If a message arrives from the peer between these two
calls, the interrupt handler dibs_con_peer_dev_handle_irq() will fire and
call kfifo_in() on the already freed dcpd->fifo buffer, resulting in memory
corruption.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008-console_server-v1-0-636b11ebf191@linux.ibm.com?part=3

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH 2/4] dibs: introduce dibs peer device bus driver
  2026-10-08  8:08 ` [PATCH 2/4] dibs: introduce dibs peer device bus driver Julian Ruess
@ 2026-10-08  8:27   ` sashiko-bot
  0 siblings, 0 replies; 10+ messages in thread
From: sashiko-bot @ 2026-10-08  8:27 UTC (permalink / raw)
  To: Julian Ruess
  Cc: Heiko Carstens, Christian Borntraeger, linux-s390, Vasily Gorbik,
	Alexander Gordeev

Thank you for your contribution! Sashiko AI review found 19 potential issue(s) to consider:
- [High] dibs: NULL pointer dereference after dibs_pbd_drv_add_dibs allocation failure
- [High] dibs: UAF of remote_endpoint due to dangling pointers in dds->dmb_state
- [High] dibs: Unbounded memory leak of access_token objects leading to DoS
- [High] dibs: UAF of dibs_peer_device due to event handling race with teardown
- [High] dibs: UAF of parent dibs device due to uncancelled handle_irq_work
- [High] dibs: UAF of dibs_peer_device from re-arming timeout_work during teardown
- [High] dibs: UAF on module exit due to unflushed dibs_pbd_event_work items
- [High] dibs: Out-of-bounds read due to missing payload length validation
- [High] dibs: Kernel stack info leak via uninitialized msg_to_send struct
- [High] dibs: Integer underflow in dibs_pbd_msg_max_payload()
- [High] dibs: use-after-free of remote_endpoint in IRQ handler
- [High] dibs: unbounded memory leak of access_token objects (remote DoS)
- [High] dibs: reconnected peer devices are incorrectly destroyed by uncancelled timeout
- [Medium] dibs: Ring buffer permanently wedges on unknown message types
- [Medium] dibs: Data race from unlocked modification of rd_meta_dmb.idx state
- [Medium] dibs: Device matching ignores rgid, breaking discovery for identical IDs
- [Medium] dibs: Integer truncation in send_padding_msg wedges ring buffer
- [Medium] dibs: UAF in dibs_pbd_peer_dev_create error path due to uncancelled works
- [Low] dibs: TOCTOU race in remote_endpoint creation creates duplicate entries
--

commit c1ec605f331a2f74f0e89eba64870a407cc1570c
Author: Julian Ruess <julianr@linux.ibm.com>

    dibs: introduce dibs peer device bus driver

    Introduce a dibs peer device bus driver that represents remotely
    emulated dibs peer devices as kernel devices. The driver handles
    discovery and lifecycle management of remote peer devices and provides
    the foundation for dibs peer device drivers that operate on top of dibs.

> diff --git a/drivers/dibs/dibs_peer_bus_drv.c b/drivers/dibs/dibs_peer_bus_drv.c
> new file mode 100644
> index 0000000000000..99eb0a6f8164d
> --- /dev/null
> +++ b/drivers/dibs/dibs_peer_bus_drv.c

[ ... ]

> +size_t dibs_pbd_msg_max_payload(struct dibs_ring_buffer *rb)
> +{
> +	size_t space = dibs_ring_space_to_end(rb);
> +
> +	if (space == 0)
> +		return 0;
> +
> +	if (space == sizeof(struct dibs_msg_hdr)) {
> +		space = dibs_ring_space(rb) - space;
> +		if (space == 0)
> +			return 0;
> +	}
> +
> +	return ALIGN_DOWN(space - sizeof(struct dibs_msg_hdr),
> +			  DIBS_PBD_MSG_ALIGNMENT);
> +}

[Severity: High]
Does this calculation safely prevent integer underflow?

If the ring buffer is nearly full and `space` is less than
`sizeof(struct dibs_msg_hdr)` (for instance, 3 bytes), subtracting the header
size will underflow, causing this function to return an extremely large value
(e.g., `SIZE_MAX - 3`). This could lead to memory corruption or out-of-bounds
writes in callers that rely on this size.

[ ... ]

> +static int dibs_pbd_send_padding_msg(struct dibs_ring_buffer *rb)
> +{
> +	int ret;
> +	struct dibs_msg_hdr padding_hdr;
> +
> +	padding_hdr.version = DIBS_PBD_HDR_V1;
> +	padding_hdr.type = DIBS_PBD_PADDING_MSG;
> +	padding_hdr.datalen =
> +		dibs_ring_space_to_end(rb) - sizeof(struct dibs_msg_hdr);

[Severity: Medium]
Could this assignment truncate the padding size?

The result of `dibs_ring_space_to_end(rb) - sizeof(struct dibs_msg_hdr)` is a
`size_t` that could potentially be up to 1MB, but `padding_hdr.datalen` is
a `u16`. If the required padding space exceeds 65539 bytes, the assigned
value will be silently truncated. This could cause the sender to lose
synchronization with the physical ring boundaries, wedging the connection.

[ ... ]

> +struct dibs_msg_hdr *dibs_pbd_msg_recv(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_msg_hdr *msg;
> +	u32 msg_size;
> +
> +	while ((msg = dibs_ring_recv(rb))) {
> +		msg_size = dibs_pbd_msg_size(msg);
> +		if (!msg_size || msg_size > dibs_ring_cnt_to_end(rb)) {

[Severity: High]
Is it possible for a remote peer to send a message where the payload size
is too small for the actual message type?

This size validation verifies that the message fits within the ring buffer,
but it does not enforce a minimum structural size for specific message types.
If a peer sets `datalen` to 0, subsequent casts and assignments in handlers
like `dibs_pbd_re_dev_msg()` could read out-of-bounds memory.

[ ... ]

> +static void dibs_pbd_re_access_tok_msg(struct dibs_msg_hdr *msg_hdr,
> +				       struct dibs_pbd_remote_endpoint *re)
> +{
> +	struct dibs_pbd_remote_endpoint_access_tok_msg atmsg;
> +	struct dibs_pbd_remote_endpoint_msg msg_to_send;

[Severity: High]
Can uninitialized stack memory be leaked here?

The `msg_to_send` structure is allocated on the stack without `{0}`
initialization. Later, `strscpy()` is used to populate fields like
`system_name` and `hostname`, leaving the trailing bytes of these fixed-size
arrays uninitialized. When `dibs_pbd_msg_send()` transmits the struct, this
raw kernel stack memory is sent over the bus.

> +	struct dibs_pbd_remote_endpoint_msg *msg;
> +	struct dibs_system_info sysinfo;
> +	struct access_token *at;
> +	int ret;
> +
> +	msg = (struct dibs_pbd_remote_endpoint_msg *)msg_hdr;
> +	atmsg = msg->atmsg;
> +
> +	if (re->trusted) {
> +		at = kzalloc_obj(*at);
> +		if (!at) {
> +			pr_err("%s: failed to allocate access token\n",
> +			       __func__);
> +			return;
> +		}
> +		at->access_tok = atmsg.access_tok;
> +		mutex_lock(&dibs_pbd_access_token_list_mutex);
> +		list_add_tail(&at->list, &dibs_pbd_access_token_list);
> +		mutex_unlock(&dibs_pbd_access_token_list_mutex);

[Severity: High]
Is there a risk of an unbounded memory leak on this list?

Every received token is unconditionally appended to the global
`dibs_pbd_access_token_list`, but there appears to be no mechanism to ever
remove or free them. A remote endpoint could send these messages indefinitely,
leading to a memory leak and a potential Denial of Service.

[ ... ]

> +static int dibs_pbd_peer_dev_match_id(struct device *dev, const void *data)
> +{
> +	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
> +	const u8 *id = data;
> +
> +	if (dpd && dpd->id == *id)
> +		return 1;
> +	return 0;
> +}

[Severity: Medium]
Should this matching function also check the remote GID (`rgid`)?

By only checking the device ID, this ignores the endpoint it originated from.
If two different endpoints provide a device with the same ID, the search
will return the first one found regardless of endpoint.

[ ... ]

> +static void dibs_pbd_dibs_event_dev_disabled_work(struct work_struct *work)
> +{
> +	struct dibs_peer_dev_driver *dpdd;
> +	struct dibs_peer_device *dpd;
> +
> +	dpd = container_of(work, struct dibs_peer_device,
> +			   event_dev_disabled_work);
> +
> +	device_lock(&dpd->dev);
> +	if (dpd->dev.driver) {
> +		dpdd = to_dibs_peer_dev_driver(dpd->dev.driver);
> +		dpdd->disconnect(dpd);
> +	}
> +	device_unlock(&dpd->dev);
> +
> +	schedule_delayed_work(&dpd->timeout_work, dpd->timeout);
> +}

[Severity: High]
Does this safely interact with `dibs_pbd_peer_dev_destroy()`?

During teardown, `dibs_pbd_peer_dev_destroy()` cancels
`event_dev_disabled_work`. If `event_dev_disabled_work` is already running
concurrently, teardown waits. But as seen here, the work ends by
re-arming `timeout_work`. Since `dibs_pbd_peer_dev_destroy()` skips
cancelling `timeout_work` if it is the current work context, the newly
re-armed timer remains active after the device is freed, causing a
use-after-free when it fires.

[ ... ]

> +static int dibs_pbd_peer_dev_create(struct dibs_dev *dibs, uuid_t *rgid,
> +				    u64 rdmb_tok, u8 server_dev_id, u8 type,
> +				    u128 tok)
> +{

[ ... ]

> +	ret = device_register(&dpd->dev);
> +	if (ret) {
> +		put_device(&dpd->dev);
> +		return ret;
> +	}
> +	ret = sysfs_create_link(&dpd->dev.kobj, &dpd->parent->dev.kobj, "dibs");
> +	if (ret) {
> +		device_unregister(&dpd->dev);
> +		return ret;
> +	}

[Severity: Medium]
Are background work items properly cancelled on this error path?

Between `device_register()` and the potential `sysfs_create_link()` failure,
the device is visible on the bus and could have events dispatched to it (e.g.,
`event_dev_disabled_work`). If this path calls `device_unregister()` without
first cancelling the workqueue items, it could lead to a use-after-free.

[ ... ]

> +static void dibs_pbd_re_dev_msg(struct dibs_msg_hdr *msg_hdr,
> +				struct dibs_pbd_remote_endpoint *re)
> +{

[ ... ]

> +	if (dev) {
> +		dpd = to_dibs_peer_device(dev);
> +		if (dpd->tok != redm.peer_dev_tok) {
> +			put_device(dev);
> +			return;
> +		}
> +		if (dpd->rdmb_tok != redm.dmb_tok)
> +			dpd->rdmb_tok = redm.dmb_tok;
> +
> +		device_lock(&dpd->dev);
> +		if (dpd->dev.driver) {
> +			dpdd = to_dibs_peer_dev_driver(dpd->dev.driver);
> +			ret = dpdd->reconnect(dpd);
> +			if (!ret)
> +				dpd->status = PEER_DEV_CONNECTED;
> +		}
> +		device_unlock(&dpd->dev);
> +		put_device(dev);

[Severity: High]
Does a successful reconnect cancel pending teardown timers?

If a peer device was recently disabled, `event_dev_disabled_work` and
`timeout_work` may already be scheduled to destroy the device. If the device
reconnects here, those timers are never cancelled, meaning the newly
reconnected device will still be destroyed shortly after.

[ ... ]

> +static void dibs_pbd_remote_endpoint_release(struct kref *kref)
> +{
> +	struct dibs_pbd_remote_endpoint *re;
> +
> +	re = container_of(kref, struct dibs_pbd_remote_endpoint, kref);
> +
> +	dibs_ring_unregister(&re->rb);
> +	kfree(re);
> +}

[Severity: High]
Does this safely clean up all endpoint references?

When `re` is freed here, the pointers to this endpoint stored in
`dds->dmb_state[...].remote_endpoint` are never cleared. A subsequent or
concurrent IRQ might fetch this dangling pointer from the `dmb_state` array
and use it, causing a use-after-free.

Additionally, this function is called when the reference count drops. If
`peer_bus_drv_remove_dibs()` drops the initial reference while
`handle_irq_work` is still running, the endpoint won't be freed until the
work finishes. However, the parent `dibs_dev` may be freed immediately. When
the work eventually finishes and invokes this release function,
`dibs_ring_unregister(&re->rb)` will access the already freed parent device.

[ ... ]

> +static void dibs_pbd_handle_irq_work(struct work_struct *work)
> +{

[ ... ]

> +	while ((msg = dibs_pbd_msg_recv(rb))) {
> +		switch (msg->type) {
> +		case DIBS_PBD_REMOTE_ENDPOINT_ACCESS_TOKEN_MSG:
> +			dibs_pbd_re_access_tok_msg(msg, re);
> +			break;
> +		case DIBS_PBD_DEV_MSG:
> +			if (re->trusted)
> +				dibs_pbd_re_dev_msg(msg, re);
> +			break;
> +		default:
> +			pr_warn("%s: Received unknown message type: %d\n",
> +				__func__, msg->type);
> +			kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
> +			return;
> +		}
> +		dibs_pbd_msg_ack(rb);
> +	}

[Severity: Medium]
Does returning early on unknown messages wedge the ring buffer?

If an unknown message type is received, the early return skips calling
`dibs_pbd_msg_ack(rb)`. Since the message is never acknowledged, the read
pointer doesn't advance. The next IRQ will read the exact same message,
looping infinitely and permanently wedging the channel.

[ ... ]

> +static void dibs_pbd_drv_handle_irq(struct dibs_dev *dibs, unsigned int dmbno,
> +				    u16 dmbemask)
> +{
> +	struct dibs_pbd_dibs_dev_state *dds;
> +	struct dibs_pbd_remote_endpoint *re;
> +	struct dibs_peer_dev_driver *dpdd;
> +	struct dibs_peer_device *dpd;
> +	unsigned long flags;
> +
> +	dds = dibs_get_priv(dibs, &dibs_pbd_drv);
> +
> +	spin_lock_irqsave(&dds->dmb_state[dmbno].dmb_state_lock, flags);

[Severity: High]
Is it possible for `dds` to be NULL here?

In `dibs_pbd_drv_add_dibs()`, if the allocation for `dds` fails, it returns
early without setting `dibs->priv`. This leaves the context pointer NULL.
When an IRQ fires, `dibs_get_priv` returns NULL, and this unconditionally
dereferences it.

[ ... ]

> +static int dibs_pbd_peer_dev_event(struct device *dev, void *data)
> +{
> +	struct dibs_peer_device *dpd = to_dibs_peer_device(dev);
> +	struct dibs_event *event = data;
> +
> +	switch (event->subtype) {
> +	case DIBS_SW_EVENT_REDISCOVER:
> +		dpd->rgid = event->gid;
> +		cancel_delayed_work(&dpd->timeout_work);
> +		break;

[Severity: High]
Should this also cancel `event_dev_disabled_work`?

This attempts to abort teardown by cancelling `timeout_work`, but fails to
cancel `event_dev_disabled_work`. The pending disabled work will run and
simply reschedule the timeout anyway, defeating the recovery event.

> +	case DIBS_DEV_DISABLED:
> +		if (uuid_equal(&dpd->rgid, &event->gid))
> +			schedule_work(&dpd->event_dev_disabled_work);
> +		break;
> +	}

[Severity: High]
Is there a race condition with device destruction here?

This work is scheduled without checking if the device is currently being
destroyed (e.g., via `dpd->destroying`). If `dibs_pbd_peer_dev_destroy()`
runs concurrently, the `bus_for_each_dev()` iterator used to dispatch this
event might hold a transient reference. Once dropped, `dpd` is freed, and
the workqueue will execute on the freed memory.

[ ... ]

> +static int dibs_pbd_remote_endpoint_create(struct dibs_dev *dibs,
> +					   const uuid_t *rgid,
> +					   const bool trusted)
> +{
> +	struct dibs_pbd_remote_endpoint *re, *tmp;
> +	struct dibs_pbd_dibs_dev_state *dds;
> +	bool existing_re_found = false;
> +	unsigned long flags;
> +	int ret;
> +
> +	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
> +	list_for_each_entry_safe(re, tmp, &dibs_pbd_remote_endpoint_list,
> +				 list) {
> +		if (uuid_equal(&re->rgid, rgid))
> +			existing_re_found = true;
> +	}
> +	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);
> +
> +	if (existing_re_found) {
> +		pr_warn("%s: remote endpoint already exists for rgid: %pUb\n",
> +			__func__, rgid);
> +		return 0;
> +	}
> +
> +	re = kzalloc_obj(*re, GFP_KERNEL);

[ ... ]

> +	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
> +	list_add(&re->list, &dibs_pbd_remote_endpoint_list);
> +	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);

[Severity: Low]
Could dropping and re-acquiring the list mutex create duplicates?

Because the `dibs_pbd_remote_endpoint_list_mutex` is dropped after confirming
no endpoint exists, two concurrent invocations could both see an empty list
and proceed to allocate and add duplicate endpoints for the same `rgid`.

[ ... ]

> +static void dibs_pbd_drv_handle_event(struct dibs_dev *dev,
> +				      const struct dibs_event *event)
> +{
> +	struct dibs_pbd_event_work *ew;
> +
> +	ew = kzalloc_obj(*ew, GFP_ATOMIC);
> +	if (!ew)
> +		return;
> +	ew->dibs = dev;
> +	ew->event = *event;
> +
> +	INIT_WORK(&ew->work, dibs_pbd_drv_handle_event_work);
> +	schedule_work(&ew->work);
> +}

[Severity: High]
Are these queued works flushed during module exit?

`dibs_pbd_exit()` unregisters the driver but fails to flush `system_wq` to
wait for dynamically allocated items like `ew->work`. If an event is scheduled
just before module unload, it could execute after the driver's code is
unmapped, causing a panic.

[ ... ]

> +static void dibs_pbd_drv_add_dibs(struct dibs_dev *dibs)
> +{
> +	struct dibs_pbd_dibs_dev_state *dds;
> +	const char *gid_str;
> +	int max_dmbs;
> +	uuid_t gid;
> +	int ret;
> +
> +	max_dmbs = dibs->ops->max_dmbs();
> +	dds = kzalloc(sizeof(*dds) +
> +			      max_dmbs * sizeof(struct dibs_pbd_dmb_state),
> +		      GFP_KERNEL);
> +	if (!dds) {
> +		dev_err(&dibs->dev,
> +			"Failed to allocate memory for dibs dev state\n");
> +		return;
> +	}

[Severity: High]
Does this error path leave subsystem state inconsistent?

Returning early on allocation failure leaves `dibs->priv` set to NULL for
this client. A later IRQ delivery will trigger a NULL pointer dereference in
`dibs_pbd_drv_handle_irq()`.

[ ... ]

> +static void peer_bus_drv_remove_dibs(struct dibs_dev *dibs)
> +{
> +	struct dibs_pbd_remote_endpoint *re, *tmp;
> +	struct dibs_pbd_dibs_dev_state *dds;
> +
> +	bus_for_each_dev(&dibs_pbd_type, NULL, dibs,
> +			 dibs_pbd_destroy_matching_peer_dev);
> +
> +	dds = dibs_get_priv(dibs, &dibs_pbd_drv);
> +
> +	mutex_lock(&dibs_pbd_remote_endpoint_list_mutex);
> +	list_for_each_entry_safe(re, tmp, &dibs_pbd_remote_endpoint_list, list) {
> +		if (re->dibs == dibs) {
> +			list_del(&re->list);
> +			kref_put(&re->kref, dibs_pbd_remote_endpoint_release);
> +		}
> +	}
> +	mutex_unlock(&dibs_pbd_remote_endpoint_list_mutex);

[Severity: High]
Should active endpoint work items be cancelled before destruction?

This path drops the initial reference to the endpoints but does not explicitly
call `cancel_work_sync(&re->handle_irq_work)`. As discussed previously,
if the work is running, it will outlive the parent `dibs_dev` and eventually
cause a use-after-free when `dibs_pbd_remote_endpoint_release()` touches the
freed parent device.

[ ... ]

> +int dibs_pbd_register_ring_buffer(char *name, struct dibs_peer_device *dpd,
> +				  u32 size)
> +{

[ ... ]

> +	spin_lock_irqsave(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
> +			  flags);
> +	dds->dmb_state[dpd->rb.dmb.idx].peer_dev_driver = dpdd;
> +	dds->dmb_state[dpd->rb.dmb.idx].peer_dev = dpd;
> +	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev_driver = dpdd;
> +	dds->dmb_state[dpd->rb.rd_meta_dmb.idx].peer_dev = dpd;
> +	spin_unlock_irqrestore(&dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock,
> +			       flags);
> +
> +	return 0;
> +}

[Severity: Medium]
Does modifying `rd_meta_dmb.idx` state here create a data race?

The code acquires the spinlock specifically for the `dmb.idx` slot
(`dds->dmb_state[dpd->rb.dmb.idx].dmb_state_lock`), but proceeds to modify
the state for the `rd_meta_dmb.idx` slot. Since these indices can differ,
the modification to the `rd_meta_dmb.idx` slot is unprotected, which can race
with the IRQ handler.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008-console_server-v1-0-636b11ebf191@linux.ibm.com?part=2

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH 4/4] dibs: introduce dibs block device peer device driver
  2026-10-08  8:08 ` [PATCH 4/4] dibs: introduce dibs block device " Julian Ruess
@ 2026-10-08  8:30   ` sashiko-bot
  0 siblings, 0 replies; 10+ messages in thread
From: sashiko-bot @ 2026-10-08  8:30 UTC (permalink / raw)
  To: Julian Ruess
  Cc: Heiko Carstens, Vasily Gorbik, Christian Borntraeger, linux-s390,
	Alexander Gordeev

Thank you for your contribution! Sashiko AI review found 15 potential issue(s) to consider:
- [Critical] dibs: data corruption due to missing request sequence numbers
- [High] dibs: UAF in worker due to unregistering ring buffer before cancel_work_sync
- [High] block: UAF of tag_set and driver data on queue teardown
- [High] block: uninitialized tag_set usage due to early ring buffer registration
- [High] dibs: race condition on device_add_disk state causes NULL dereference
- [High] dibs: UAF of blkpd due to media_gone_work scheduled after cancellation
- [High] dibs: UAF / Race condition accessing unregistered ring buffer from queue_rq
- [High] dibs: TOCTOU bounds check bypass via shared memory mutation
- [High] block: deadlock in dibs_blk_peer_dev_remove during teardown
- [Medium] dibs: UAF of blkpd in worker due to missing cancel_work_sync on probe error
- [Medium] block: API misuse exhausts dynamic major numbers and overflows minor numbers
- [Medium] block: double-unquiesce warning triggered by unsynchronized state access
- [Medium] dibs: Out-of-bounds read accessing message payload length fields
- [Medium] block: disk_name truncation causes sysfs naming collisions
- [Low] block: snprintf truncation causes sysfs name collisions
--

commit c8b558951c6dbf245b6e9d1bfed2adc0596fda35
Author: Julian Ruess <julianr@linux.ibm.com>
Subject: dibs: introduce dibs block device peer device driver

This introduces a virtual block device driver that uses the newly added dibs
peer device bus. It allows access to resources like an ISO file hosted on
another system via dibs. A block device is created when a dibs peer device
of type block is announced.

Hello,

Thanks for the patch. I noticed a few potential issues with the teardown
and communication logic that might need a closer look.

> diff --git a/drivers/dibs/dibs_blk.c b/drivers/dibs/dibs_blk.c
> new file mode 100644
> --- /dev/null
> +++ b/drivers/dibs/dibs_blk.c
> @@ -0,0 +1,798 @@
[ ... ]
> +static blk_status_t dibs_blk_queue_rq(struct blk_mq_hw_ctx *hctx,
> +				      const struct blk_mq_queue_data *bd)
> +{
[ ... ]
> +	blkpd = bd->rq->q->disk->private_data;
> +
> +	if (!blkpd->device_add_disk || !blkpd->ring_registered)
> +		return BLK_STS_IOERR;

[Severity: High]
Could concurrent userspace I/O access the unregistered ring buffer here?

In dibs_blk_peer_dev_remove(), the driver calls
dibs_pbd_unregister_ring_buffer() to unmap and free the ring buffer memory
before calling blk_mark_disk_dead() and del_gendisk().

Because the queue is still active, concurrent userspace I/O can enter
dibs_blk_queue_rq() while the ring buffer is being unregistered. This
lockless check might pass before the ring_registered flag is cleared,
causing dibs_pbd_msg_send() to attempt writing to the freed memory.

[ ... ]
> +static void dibs_blk_unquiesce(struct dibs_blk_peer_dev *blkpd)
> +{
> +	if (!blkpd->blk.quiesced)
> +		return;
> +
> +	blkpd->blk.quiesced = false;

[Severity: Medium]
Could this unsynchronized state access trigger a double-unquiesce warning?

If dibs_blk_unquiesce() is called concurrently from media_gone_work and
handle_disk_detach, both threads might read quiesced as true and proceed
to call blk_mq_unquiesce_queue().

This second call would drop q->quiesce_depth below zero, triggering a
WARN_ON_ONCE(q->quiesce_depth <= 0) in blk_mq_unquiesce_queue().

[ ... ]
> +static enum blk_eh_timer_return dibs_blk_timeout(struct request *req)
> +{
[ ... ]
> +	blk_mq_end_request(req, BLK_STS_IOERR);
> +	schedule_work(&blkpd->media_gone_work);

[Severity: High]
Can media_gone_work be scheduled after it has been cancelled?

In dibs_blk_peer_dev_remove(), cancel_work_sync(&blkpd->media_gone_work)
is called before iterating the tags to cancel inflight requests. Before
those requests are cancelled, a timer could expire, triggering this
timeout handler.

This explicitly calls schedule_work(&blkpd->media_gone_work) after the
cancellation has finished. The newly scheduled work will execute later and
dereference the freed blkpd structure.

[ ... ]
> +static void dibs_blk_handle_data_msg(struct dibs_blk_peer_dev *blkpd,
> +				     void *current_data, u64 data_len,
> +				     u64 total_bytes_of_request)
> +{
[ ... ]
> +	req = blkpd->current_request;
> +	if (!req) {
> +		spin_unlock_irqrestore(&blkpd->req_lock, flags);
> +		pr_warn("Received response for unknown request\n");
> +		return;
> +	}
> +
> +	blkpd->last_completion = jiffies;
> +	blkpd->first_timeout = 0;
> +
> +	if (total_bytes_of_request != blk_rq_bytes(req)) {

[Severity: Critical]
How does the driver guarantee this response belongs to the current request?

The dibs block protocol appears to lack transaction IDs or sequence numbers.
If a block request times out, dibs_blk_timeout() clears current_request
and completes it with an error. The block layer can immediately dispatch a
new request, assigning it to current_request.

When the remote peer later sends the delayed data response for the original
request, this code blindly accepts it for the new current_request because
the lengths often match for identical block sizes. This would cause stale
data from the old request to be copied into the new request's bio buffers,
corrupting filesystem data.

[ ... ]
> +static void dibs_blk_media_change_work(struct work_struct *work)
> +{
> +	struct dibs_blk_peer_dev *blkpd;
> +	int ret;
> +
> +	blkpd = container_of(work, struct dibs_blk_peer_dev, media_change_work);
> +
> +	if (!blkpd->device_add_disk) {
> +		blkpd->blk.gd = blk_mq_alloc_disk(&blkpd->blk.tag_set,
> +						  &dibs_blk_lim, &blkpd->blk);
[ ... ]
> +		blkpd->blk.gd->first_minor =
> +			blkpd->device_index * MAX_GPT_PARTITIONS;

[Severity: Medium]
Will this multiplication lead to minor number overflow?

In probe(), the driver allocates a dynamic major number per device via
register_blkdev(0, "dibs_blk") instead of doing this once in module init,
which exhausts the small pool of available dynamic major numbers.

Additionally, device_index unconditionally increments via dibs_blk_idx++.
Because the minor number is calculated as device_index * MAX_GPT_PARTITIONS,
after 8192 probes the minor number exceeds the maximum limit, causing
device_add_disk() to fail or overwrite the major number.

[ ... ]
> +		blkpd->blk.gd->minors = MAX_GPT_PARTITIONS;
> +		blkpd->blk.gd->events = DISK_EVENT_MEDIA_CHANGE;
> +
> +		set_disk_ro(blkpd->blk.gd, 1);
> +		snprintf(blkpd->blk.gd->disk_name, 10, "dibs_blk%d",
> +			 blkpd->device_index);

[Severity: Medium]
Is the size parameter of 10 too small for double-digit indices?

The prefix dibs_blk consumes 8 characters. For any device_index >= 10,
the string requires at least 11 characters including the null terminator.
Because the size limit is hardcoded to 10 instead of DISK_NAME_LEN,
snprintf truncates the output to 9 characters plus the null terminator.

This means devices with an index >= 10 will be truncated and named
dibs_blk1, guaranteeing sysfs naming collisions and initialization
failures when multiple devices are attached or hotplugged.

[ ... ]
> +		dibs_blk_update_capacity(blkpd);
> +		blkpd->device_add_disk = 1;
> +
> +		ret = device_add_disk(&blkpd->dpd->dev, blkpd->blk.gd, NULL);
> +		if (ret) {
> +			pr_err("%s: device_add_disk failed: %d\n", __func__,
> +			       ret);
> +			blkpd->device_add_disk = 0;
> +			put_disk(blkpd->blk.gd);
> +			blkpd->blk.gd = NULL;
> +			return;
> +		}
> +	}
> +
> +	dibs_blk_update_capacity(blkpd);
> +	disk_force_media_change(blkpd->blk.gd);
> +	dibs_blk_rescan_partitions(blkpd);
> +}

[Severity: High]
Could exposing the device_add_disk flag early lead to a race condition?

If device_add_disk() fails, the error path resets the flag to 0 and frees
the gendisk. Concurrently, dibs_blk_handle_disk_detach() reads the flag
locklessly. If it sees the flag as 1 before the failure path resets it, it
calls disk_force_media_change(blkpd->blk.gd).

This would access the pointer while it is being freed or after it is
NULLed, causing a crash.

[ ... ]
> +static void dibs_blk_proc_recv_msg_work(struct work_struct *work)
> +{
[ ... ]
> +	while ((msg = (struct dibs_blk_peer_dev_msg *)dibs_pbd_msg_recv(&peer_dev->rb))) {
> +		switch (msg->hdr.type) {
> +		case DIBS_BLK_DATA_MSG:
> +			current_data = msg->blkdm.data;
> +			data_len = READ_ONCE(msg->blkdm.datalen);

[Severity: Medium]
Should the code validate the message bounds before accessing payload fields?

While dibs_pbd_msg_recv() ensures the ring buffer holds at least
msg->hdr.datalen bytes, this code accesses READ_ONCE(msg->blkdm.datalen)
and READ_ONCE(msg->blkmcm.size) without verifying that msg->hdr.datalen is
large enough to contain the full struct dibs_blk_data_msg or struct
dibs_blk_media_change_msg.

If a peer sends a valid message type but a tiny header length, reading
these payload fields accesses memory beyond the validated message bounds,
potentially causing a page fault.

[ ... ]
> +			total_bytes_of_request =
> +				READ_ONCE(msg->blkdm.total_bytes_of_req);
> +
> +			if (sizeof(msg->blkdm) + data_len >
> +			    READ_ONCE(msg->hdr.datalen))
> +				break;

[Severity: High]
Is there a TOCTOU vulnerability here when re-reading datalen from shared memory?

Because msg resides in the shared ring buffer mapped to the remote peer, a
malicious peer can concurrently mutate msg->hdr.datalen to a large value
after the initial boundary validation but during this check.

This bypasses the validation, causing dibs_blk_handle_data_msg() to copy
data_len bytes from the shared memory pointer, reading past the end of the
mapped buffer and leaking adjacent kernel memory.

[ ... ]
> +static int dibs_blk_peer_dev_probe(struct dibs_peer_device *dpd)
> +{
[ ... ]
> +	ret = dibs_pbd_register_ring_buffer("dibs_blk", dpd,
> +					    DIBS_BLK_PEER_DEV_BLK_RING_SIZE);
> +	if (ret) {
> +		pr_err("%s: failed to register ring buffer\n", __func__);
> +		goto err_free;
> +	}
> +
> +	blkpd->ring_registered = true;
> +	dibs_ring_set_rdmb_tok(&dpd->rb, dpd->rdmb_tok);

[Severity: High]
Does activating the remote communication channels here allow the remote
peer to access uninitialized block layer structures?

The remote device can immediately send a DIBS_BLK_MEDIA_CHANGE message,
triggering the IRQ handler which schedules proc_recv_msg_work and then
media_change_work.

Since media_change_work calls blk_mq_alloc_disk(&blkpd->blk.tag_set, ...),
if this runs before probe() reaches blk_mq_alloc_tag_set(), the tag_set is
uninitialized memory, leading to a crash in the block layer.

[ ... ]
> +	blkpd->blk.major = register_blkdev(0, "dibs_blk");
[ ... ]
> +err_unregister_ring:
> +	if (blkpd->ring_registered)
> +		dibs_pbd_unregister_ring_buffer(dpd);
> +err_free:
> +	dpd->drv_priv = NULL;
> +	kfree(blkpd);
> +	return ret;
> +}

[Severity: Medium]
Is there a missing cancel_work_sync() before freeing blkpd on this error path?

If a step after ring buffer registration fails (like register_blkdev), the
code jumps to err_unregister_ring and frees blkpd.

Because the IRQ handler could have already scheduled proc_recv_msg_work,
failing to call cancel_work_sync(&blkpd->proc_recv_msg_work) means the
work could execute later on the freed blkpd structure.

[ ... ]
> +static void dibs_blk_peer_dev_remove(struct dibs_peer_device *dpd)
> +{
> +	struct dibs_blk_peer_dev *blkpd = dpd->drv_priv;
> +
> +	if (blkpd->list_added) {
> +		mutex_lock(&blkpd_mutex);
> +		list_del(&blkpd->list);
> +		mutex_unlock(&blkpd_mutex);
> +		blkpd->list_added = false;
> +	}
> +
> +	if (blkpd->ring_registered) {
> +		dibs_pbd_unregister_ring_buffer(blkpd->dpd);
> +		blkpd->ring_registered = false;
> +	}
> +
> +	cancel_work_sync(&blkpd->proc_recv_msg_work);

[Severity: High]
Would tearing down the ring buffer before stopping the worker lead to a
use-after-free?

In both dibs_blk_peer_dev_remove() and dibs_blk_peer_dev_disconnect(), the
driver unregisters the ring buffer before calling
cancel_work_sync(&blkpd->proc_recv_msg_work).

Because the background worker accesses the ring buffer via
dibs_pbd_msg_recv(), if the worker is running concurrently, tearing down
the ring buffer from underneath it causes a use-after-free.

[ ... ]
> +	cancel_work_sync(&blkpd->media_change_work);

[Severity: High]
Can this removal path deadlock by waiting on a work item blocked on the
quiesced queue?

When the device is disconnected, dibs_blk_peer_dev_disconnect() quiesces
the queue but does not cancel media_change_work. If media_change_work
runs, it calls bdev_disk_changed() which issues synchronous read I/O to
scan partitions.

Since the queue is quiesced, these read requests block indefinitely. When
dibs_blk_peer_dev_remove() subsequently calls cancel_work_sync() for
media_change_work, it hangs forever waiting for the blocked work. The
queue is not unquiesced until later in this function, resulting in a
permanent deadlock.

[ ... ]
> +	if (blkpd->device_add_disk) {
> +		blk_mark_disk_dead(blkpd->blk.gd);
> +		dibs_blk_unquiesce(blkpd);
> +		blk_mq_tagset_busy_iter(&blkpd->blk.tag_set,
> +					dibs_blk_cancel_request, blkpd);
> +		del_gendisk(blkpd->blk.gd);
> +		put_disk(blkpd->blk.gd);
> +		blkpd->blk.gd = NULL;
> +		blkpd->device_add_disk = 0;
> +	}
> +
> +	if (blkpd->tag_set_allocated)
> +		blk_mq_free_tag_set(&blkpd->blk.tag_set);
> +
> +	if (blkpd->blkdev_registered)
> +		unregister_blkdev(blkpd->blk.major, "dibs_blk");
> +
> +	kfree(blkpd);
> +}

[Severity: High]
Is it safe to free tag_set and driver data while the block layer may still
retain a reference to the queue?

This code drops the driver's reference using put_disk(), followed
immediately by blk_mq_free_tag_set() and kfree(blkpd). If a userspace
process holds the block device open, the gendisk and request_queue remain
alive.

When the user later closes the device, blk_release_queue() will attempt to
access the queue's tag_set to free hardware queues. Since tag_set points
to blkpd->blk.tag_set inside the freed blkpd struct, this causes a
use-after-free. The tag set and private data must be managed via the
gendisk's .free_disk callback.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008-console_server-v1-0-636b11ebf191@linux.ibm.com?part=4

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH 1/4] dibs: introduce ring-buffer send/receive API
  2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
  2026-10-08  8:21   ` sashiko-bot
@ 2026-10-09 20:47   ` Niklas Schnelle
  1 sibling, 0 replies; 10+ messages in thread
From: Niklas Schnelle @ 2026-10-09 20:47 UTC (permalink / raw)
  To: Julian Ruess, wintera, ts, oberpar, gbayer
  Cc: mjrosato, alifm, raspl, hca, agordeev, gor, linux390-list,
	linux-s390

On Thu, 2026-10-08 at 10:08 +0200, Julian Ruess wrote:
> From: Tobias Schumacher <ts@linux.ibm.com>
> 
> Introduce a centralized API for sending and receiving data between dibs
> peers via ring buffers. With this API, drivers no longer need to
> interact directly with DMBs and manage read/write pointers.
> 
> A driver must first register its ring buffer using
> dibs_register_ring_buffer() and set the rdmb_tok via
> dibs_ring_set_rdmb_tok(). After that, data can be sent using
> dibs_ring_send().
> 
> Receiving messages is split into two steps: dibs_ring_recv() returns a
> pointer to the next DIBS message if one is available. After the message
> is processed, dibs_ring_ack() must be called to free up the message
> buffers and to advance the read pointer. If no message is available,
> dibs_ring_recv() returns NULL.

I think this commit message could provide a few more details on the
functioning of the ring, it's data layout and parts (e.g.
rd_dmb/wr_dmb/meta_dmb), constraints (like power of 2) and how the data
transfers work with respect to the head/tail pointers (i.e. refer to
the CIRC_… macros) and where what part resides. Quite a bit of this is
in the comments in the header but reading this patch from the start I
only noticed late and would really have liked more of it in the commit
message. Also I think there should be a section on how
recovery/reconnect works.

I'm still torn about whether it would make sense to put this in
Documentation/dibs/ where its existence is also easier to pick out from
the change stats. I believe in earlier internal discussions I was
favoring comments and commit messages but since this has grown in
complexity and I think a documentation file would be nice.

> 
> Co-developed-by: Julian Ruess <julianr@linux.ibm.com>
> Signed-off-by: Julian Ruess <julianr@linux.ibm.com>
> Signed-off-by: Tobias Schumacher <ts@linux.ibm.com>
> ---
>  drivers/dibs/dibs_main.c | 493 ++++++++++++++++++++++++++++++++++++++++++++++-
>  include/linux/dibs.h     |  99 ++++++++++
>  2 files changed, 590 insertions(+), 2 deletions(-)
> 
> diff --git a/drivers/dibs/dibs_main.c b/drivers/dibs/dibs_main.c
> index 20c50997a7cf..e3a33bc42fe8 100644
> --- a/drivers/dibs/dibs_main.c
> +++ b/drivers/dibs/dibs_main.c
> @@ -13,6 +13,9 @@
>  #include <linux/slab.h>
>  #include <linux/err.h>
>  #include <linux/dibs.h>
> +#include <linux/circ_buf.h>
> +#include <linux/log2.h>
> +#include <asm/barrier.h>
>  
>  #include "dibs_loopback.h"
>  
> @@ -20,7 +23,7 @@ MODULE_DESCRIPTION("Direct Internal Buffer Sharing class");
>  MODULE_LICENSE("GPL");
>  
>  static const struct class dibs_class = {
> -	.name		= "dibs",
> +	.name = "dibs",
>  };
>  
>  /* use an array rather a list for fast mapping: */
> @@ -93,7 +96,8 @@ int dibs_unregister_client(struct dibs_client *client)
>  		max_dmbs = dibs->ops->max_dmbs();
>  		for (int i = 0; i < max_dmbs; ++i) {
>  			if (dibs->dmb_clientid_arr[i] == client->id) {
> -				WARN(1, "%s: attempt to unregister '%s' with registered dmb(s)\n",
> +				WARN(1,
> +				     "%s: attempt to unregister '%s' with registered dmb(s)\n",

Nit: This line as well as the change in dibs_class look like overly
eager auto format instead of an actual change. This is one of the
reasons I personally only do auto format over selected lines in my
editor and no auto format on save. Also I'd keep the above literal in
one line with the WARN().

>  				     __func__, client->name);
>  				rc = -EBUSY;
>  				goto err_reg_dmb;
> @@ -245,6 +249,491 @@ void dibs_dev_del(struct dibs_dev *dibs)
>  }
>  EXPORT_SYMBOL_GPL(dibs_dev_del);
>  
> +/**
> + * dibs_ring_peer_meta() - take a snapshot of the peer's ring metadata
> + * @rb: pointer to the ring buffer
> + * @peer: snapshot to fill in
> + *
> + * The peer writes its metadata dmb at any time, so read each field once and
> + * copy to the snapshot.
> + */
> +static void dibs_ring_peer_meta(struct dibs_ring_buffer *rb,
> +				struct dibs_ring_meta *peer)
> +{
> +	peer->meta_dmb_tok = READ_ONCE(rb->rd_ring_meta->meta_dmb_tok);
> +	peer->buff_dmb_tok = READ_ONCE(rb->rd_ring_meta->buff_dmb_tok);
> +	peer->size = READ_ONCE(rb->rd_ring_meta->size);
> +	peer->head = READ_ONCE(rb->rd_ring_meta->head);
> +	peer->tail = READ_ONCE(rb->rd_ring_meta->tail);
> +	peer->reserved = 0;
> +}
> +
> +static int dibs_ring_send_hdr(struct dibs_ring_buffer *rb, u64 meta_dmb_tok)
> +{
> +	int res;
> +
> +	if (!meta_dmb_tok)
> +		return 0;
> +
> +	res = rb->dibs->ops->move_data(rb->dibs, meta_dmb_tok, 0, true, 0,
> +				       &rb->wr_ring_meta,
> +				       sizeof(struct dibs_ring_meta));
> +

Nit: Stray empty line

> +	if (!res)
> +		rb->hdr_sent = true;
> +
> +	return res;
> +}
> +
> +/**
> + * dibs_ring_register() - register a DIBS ring buffer
> + * @name: name of the ring buffer. This will be used in debug output to
> + *        support identifying the ring buffer.
> + * @rb: pointer to the ring buffer struct. The fields in the struct don't have
> + *      to be initialized.
> + * @client: pointer to the DIBS client
> + * @dibs: pointer to the DIBS device
> + * @rgid: the remote GID that will be allowed to write into this ring buffer
> + * @size: size of the ring buffer in bytes
> + *
> + * Register and initialize a DIBS ring buffer. Allocates the metadata DMB and
> + * the payload DMB, initializes the fields in dibs_ring_buffer.
> + *
> + * Return: 0 in case of success, an error code in case of failure
> + */
> +int dibs_ring_register(char *name, struct dibs_ring_buffer *rb,
> +		       struct dibs_client *client, struct dibs_dev *dibs,
> +		       uuid_t rgid, u32 size)
> +{
> +	int ret;
> +
> +	if (!size || !is_power_of_2(size))
> +		return -EINVAL;
> +
> +	strscpy(rb->name, name, sizeof(rb->name));
> +	rb->dibs = dibs;
> +	rb->dmb.rgid = rgid;
> +
> +	rb->rd_meta_dmb.dmb_len = ALIGN(sizeof(struct dibs_ring_meta), 4096);
> +	rb->rd_meta_dmb.rgid = rgid;
> +	ret = dibs->ops->register_dmb(dibs, &rb->rd_meta_dmb, client);
> +	if (ret)
> +		return ret;
> +
> +	rb->dmb.dmb_len = ALIGN(size, 4096);

Nit: The two DMB registrations would be easier to read if the dmb.rgid
was set here just like it is done for the rd_meta_dmb.rgid above and
how the dmb.dmb_len is only set here.

> +	ret = dibs->ops->register_dmb(dibs, &rb->dmb, client);
> +	if (ret) {
> +		dibs->ops->unregister_dmb(rb->dibs, &rb->rd_meta_dmb);
> +		return ret;
> +	}
> +
> +	rb->dmb_registered = 1;
> +
> +	rb->wr_ring_meta.meta_dmb_tok = rb->rd_meta_dmb.dmb_tok;
> +	rb->wr_ring_meta.buff_dmb_tok = rb->dmb.dmb_tok;
> +	rb->wr_ring_meta.size = size;
> +	rb->wr_ring_meta.head = 0;
> +	rb->wr_ring_meta.tail = 0;
> +	rb->wr_ring_meta.reserved = 0;
> +
> +	rb->rd_ring_meta = rb->rd_meta_dmb.cpu_addr;
> +	rb->peer_buff_dmb_tok = 0;
> +	rb->msg_size = 0;
> +	rb->hdr_sent = false;
> +
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(dibs_ring_register);
> +
> +/**
> + * dibs_ring_unregister() - unregister a DIBS ring buffer
> + * @rb: pointer to the ring buffer to unregister
> + *
> + * Free the resources allocated by dibs_ring_register().
> + *
> + * Return: 0 in case of success, an error code in case of failure
> + */
> +int dibs_ring_unregister(struct dibs_ring_buffer *rb)
> +{
> +	struct dibs_ring_meta peer;
> +	int ret;
> +
> +	rb->wr_ring_meta.meta_dmb_tok = 0;
> +	rb->wr_ring_meta.buff_dmb_tok = 0;
> +	rb->wr_ring_meta.size = 0;
> +	rb->wr_ring_meta.head = 0;
> +	rb->wr_ring_meta.tail = 0;
> +
> +	dibs_ring_peer_meta(rb, &peer);
> +	if (peer.meta_dmb_tok) {
> +		ret = dibs_ring_send_hdr(rb, peer.meta_dmb_tok);
> +		if (ret)
> +			pr_warn("%s: failed to send header update: %d\n",
> +				__func__, ret);
> +	}

I think the function doc should mention that this sends a final header
update and explain why this is and what it looks like,

> +
> +	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->dmb);
> +	if (ret)
> +		pr_warn("%s: failed to unregister payload dmb: %d\n", __func__,
> +			ret);
> +
> +	ret = rb->dibs->ops->unregister_dmb(rb->dibs, &rb->rd_meta_dmb);
> +	if (ret)
> +		pr_warn("%s: failed to unregister metadata dmb: %d\n", __func__,
> +			ret);
> +
> +	rb->dmb_registered = 0;
> +	return ret;
> +}
> +EXPORT_SYMBOL_GPL(dibs_ring_unregister);
> +
> +/**
> + * dibs_ring_peer_replaced() - check if the peer has replaced its ring buffer
> + * @rb: pointer to the ring buffer
> + * @peer: snapshot of the peer's metadata
> + *
> + * The fabric hands out a new dmb token for every registration, so a buffer
> + * token that differs from the one this side is synchronized to means the peer
> + * ring this side writes into is a different one.

This function doc should mention under what circumstances this may be
the case. I'm assuming this only happens during recovery after a peer
was restarted/reconnected?  For more details this should probably be
part of the documentation and then the comment can refer to that.

> + *
> + * Return: true if the peer ring has been replaced, false otherwise
> + */
> +static bool dibs_ring_peer_replaced(struct dibs_ring_buffer *rb,
> +				    const struct dibs_ring_meta *peer)
> +{
> +	return peer->buff_dmb_tok != rb->peer_buff_dmb_tok;
> +}
> +
> +/**
> + * dibs_ring_sync_peer() - synchronize with a replaced peer ring buffer
> + * @rb: pointer to the ring buffer
> + * @peer: snapshot of the peer's metadata
> + *
> + * Restart both pointers: head is this side's write position in the peer's
> + * buffer, which is a new and empty one, and tail is this side's read position
> + * in its own buffer, which the peer has started to write at offset 0 again.
> + * Sends the updated metadata header to the peer and only then remembers the
> + * new token, so that a failed send is retried on the next call.
> + *
> + * A peer that has torn its ring down is not synchronized to, it is only
> + * forgotten - this side cannot send anything until the peer registers again.

Same comment as above.

> + */
> +static void dibs_ring_sync_peer(struct dibs_ring_buffer *rb,
> +				const struct dibs_ring_meta *peer)
> +{
> +	if (!peer->buff_dmb_tok) {

Should the condition be a call to dibs_ring_peer_ready()?

> +		rb->wr_ring_meta.tail = 0;
> +		rb->peer_buff_dmb_tok = 0;
> +		rb->hdr_sent = false;
> +		return;
> +	}
> +
> +	rb->wr_ring_meta.head = 0;
> +	rb->wr_ring_meta.tail = 0;
> +
> +	if (dibs_ring_send_hdr(rb, peer->meta_dmb_tok)) {
> +		pr_warn("%s(%s): failed to send header after peer restart\n",
> +			__func__, rb->name);
> +		return;
> +	}
> +
> +	rb->peer_buff_dmb_tok = peer->buff_dmb_tok;
> +}
> +
--- snip ---
>  
> +/* Ring Buffer Communication
> + * -------------------------
> + * Based on the DMBs, DIBS provides a ring buffer mechanism for bi-directional
> + * communication between two peers. For each ring-buffer connection, two DMBs
> + * are allocated per peer:
> + * - one DMB holding metadata about the ring and
> + * - one DMB holding the actual data to be transferred.
> + *
> + * Metadata consists of parameters like the DMB tokens, size, head pointer and
> + * tail pointer. Note that the metadata stored in the locally allocated
> + * metadata DMB does not necessarily corelate to the locally allocated payload
> + * DMB. As illustrated below, each peer has access to one wr_ring_meta DMB
> + * (the one the other peer has allocated) and one rd_ring_meta DMB (the one
> + * it has allocated itself). This is because the locally allocated DMB can only
> + * be written by the other peer and vice versa.
> + *
> + * Peer A                   Peer B
> + * ---------------          ---------------
> + * wr_ring_meta ----------> rd_ring_meta
> + * rd_ring_meta <---------- wr_ring_meta
> + * send_buffer  ----------> recv_buffer
> + * recv_buffer  <---------- send_buffer
> + *
> + * Consider for example that peer A wants to send data to peer B. It first has
> + * to check how much data it previously has written (head pointer) and how far
> + * peer B has read (tail pointer) and compare this to the size. Since for send
> + * operations head is updated by peer A, it is located in the wr_ring_meta DMB.
> + * Tail is updated by the receiving side, peer B in this case, so peer A will
> + * find the current value in rd_ring_meta. The payload DMB was allocated by
> + * peer B, so peer B also updated the size - peer A therefore finds the correct
> + * value in rd_ring_meta.
> + *
> + * Now peer a can write data into the send_buffer and update the head pointer
> + * in wr_ring_meta. Peer B then reads this head pointer from its rd_ring_meta
> + * DMB, process the data and update tail in its wr_ring_meta.
> + */

I saw this way too late going through the patch from top to bottom.
Then remembered that I had read this in an internal version where I
think I also saw it late. I think Documentation changes always get
sorted to the top and stand out in the stats so that's an argument for
adding docs I guess ;)

> +struct dibs_ring_meta {
> +	/* meta_dmb_tok - Token for the metadata dmb. */
> +	u64 meta_dmb_tok;
> +	/* buff_dmb_tok - Token for the buffer dmb */
> +	u64 buff_dmb_tok;
> +	/* size - size of the buffer for receiving data in number of byte */
> +	u32 size;
> +	/* head - points to the head of the ring buffer, i.e., the element
> +	 * that will be written next
> +	 */
> +	u32 head;
> +	/* tail - points to the end of the ring-buffer, i.e., the element
> +	 * that will be read next
> +	 */
> +	u32 tail;
> +	/* reserved - keeps the struct at its natural size, must be zero */
> +	u32 reserved;
> +};
> +
--- snip ---
>  /* Functions to be called by dibs device drivers:
> @@ -461,4 +554,10 @@ int dibs_dev_add(struct dibs_dev *dibs);
>   */
>  void dibs_dev_del(struct dibs_dev *dibs);
>  
> +struct dibs_msg_hdr {
> +	u8 version;
> +	u8 type;
> +	u16 datalen;
> +} __packed;
> +
>  #endif	/* _DIBS_H */

I have so far only thoroughly read through a small part of the patch 
so regard this as a set of first comments. I feel like it is already
actionable though so wanted to reply quickly rather than wait. 

Overall I really like the abstraction and concept!

Thanks,
Niklas

^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-10-09 20:48 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08  8:08 [PATCH 0/4] Introduce dibs peer bus driver, dibs console and dibs block device driver Julian Ruess
2026-10-08  8:08 ` [PATCH 1/4] dibs: introduce ring-buffer send/receive API Julian Ruess
2026-10-08  8:21   ` sashiko-bot
2026-10-09 20:47   ` Niklas Schnelle
2026-10-08  8:08 ` [PATCH 2/4] dibs: introduce dibs peer device bus driver Julian Ruess
2026-10-08  8:27   ` sashiko-bot
2026-10-08  8:08 ` [PATCH 3/4] dibs: introduce dibs console peer device driver Julian Ruess
2026-10-08  8:26   ` sashiko-bot
2026-10-08  8:08 ` [PATCH 4/4] dibs: introduce dibs block device " Julian Ruess
2026-10-08  8:30   ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox