All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] examples/rpcapd: demo version of packet capture daemon
@ 2026-09-08 21:07 Stephen Hemminger
  2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
                   ` (2 more replies)
  0 siblings, 3 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-08 21:07 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Reshma Pattan

This example adds RPCAP support over localhost TCP
integrated with DPDK. It uses a secondary process that allows
connections from using tcpdump defacto protocol rpcap.

See: doc/guides/sample_app_ug/rpcap.rst for more info.
This does not preclude using pdump, dumpcap, or the wireshark
extcap integration.

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---

 doc/guides/rel_notes/release_26_11.rst |    4 +
 doc/guides/sample_app_ug/index.rst     |    1 +
 doc/guides/sample_app_ug/rpcapd.rst    |  206 +++++
 examples/meson.build                   |    1 +
 examples/rpcapd/main.c                 | 1014 ++++++++++++++++++++++++
 examples/rpcapd/meson.build            |   11 +
 examples/rpcapd/rpcap-protocol.h       |   96 +++
 7 files changed, 1333 insertions(+)
 create mode 100644 doc/guides/sample_app_ug/rpcapd.rst
 create mode 100644 examples/rpcapd/main.c
 create mode 100644 examples/rpcapd/meson.build
 create mode 100644 examples/rpcapd/rpcap-protocol.h

diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 87c7e81bde..9dabd80c56 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,10 @@ New Features
      Also, make sure to start the actual text at the margin.
      =======================================================
 
+* **Added an example of tcpdump remote pcap daemon.**
+
+  Added an example that implements rpcap to allow live capture in tcpdump.
+
 
 Removed Items
 -------------
diff --git a/doc/guides/sample_app_ug/index.rst b/doc/guides/sample_app_ug/index.rst
index f12623bb66..61ed870318 100644
--- a/doc/guides/sample_app_ug/index.rst
+++ b/doc/guides/sample_app_ug/index.rst
@@ -31,6 +31,7 @@ Sample Applications User Guides
     l3_forward_graph
     l3_forward_power_man
     link_status_intr
+    rpcapd
     server_node_efd
     service_cores
     multi_process
diff --git a/doc/guides/sample_app_ug/rpcapd.rst b/doc/guides/sample_app_ug/rpcapd.rst
new file mode 100644
index 0000000000..e899a4d991
--- /dev/null
+++ b/doc/guides/sample_app_ug/rpcapd.rst
@@ -0,0 +1,206 @@
+..  SPDX-License-Identifier: BSD-3-Clause
+    Copyright(c) 2026
+
+.. _rpcapd_app:
+
+dpdk-rpcapd Sample Application
+==============================
+
+The ``dpdk-rpcapd`` sample application is a Data Plane Development Kit
+(DPDK) implementation of the remote packet capture daemon protocol
+(``rpcap``) used by libpcap.  It runs as a DPDK secondary process and
+allows libpcap-aware tools such as ``tcpdump`` and Wireshark to capture
+packets from a DPDK primary process live, without writing to an
+intermediate file.
+
+The ``dpdk-rpcapd`` tool implements a subset of the protocol spoken by
+the libpcap project's ``rpcapd``.  See
+https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd for the
+reference implementation.  Clients connect to ``dpdk-rpcapd`` using a
+``rpcap://`` URL, request the list of available interfaces (which are
+the ports of the DPDK primary), open one, and stream packets from it.
+
+The intended workflow is one-step capture: start the primary, start
+``dpdk-rpcapd``, and point a familiar tool at it.  No intermediate files,
+no separate post-processing step.
+
+.. warning::
+
+   ``dpdk-rpcapd`` listens on an unauthenticated, unencrypted TCP port
+   (default 2002, bound to ``127.0.0.1``).  Any local user able to
+   reach the port can list DPDK ports and capture all traffic flowing
+   through them.  This is a sample application intended for
+   development, debugging, and demonstration use only.  **Do not run
+   ``dpdk-rpcapd`` on a production system.**
+
+   The default bind address is ``127.0.0.1`` so the listener is not
+   reachable from other hosts.  An operator may override this with
+   ``--bind <addr>`` but should expect that the resulting deployment
+   exposes captured traffic to anyone who can reach that address; do
+   not do this on an untrusted network.
+
+
+.. note::
+
+   * The ``dpdk-rpcapd`` tool can only be used in conjunction with a
+     primary application that has the packet capture framework
+     initialized already.  In DPDK, only ``dpdk-testpmd`` is modified to
+     initialize the packet capture framework; other applications must
+     be modified to call ``rte_pdump_init()`` if they are to be
+     capturable.
+
+   * ``dpdk-rpcapd`` does not replace ``dpdk-dumpcap``.  ``dpdk-dumpcap``
+     produces pcapng files; ``dpdk-rpcapd`` produces a live rpcap
+     stream.  The two tools serve different workflows and may be used
+     in parallel.
+
+   * For Wireshark users specifically, the Wireshark ``extcap`` plugin
+     interface is the preferred live-capture path; see :doc:`extcap`.
+     ``extcap`` integrates directly with Wireshark and is simpler to
+     deploy.  ``dpdk-rpcapd`` is intended for users who want to use
+     ``tcpdump`` or other rpcap-aware libpcap clients, where ``extcap``
+     does not apply.
+
+
+Running the Application
+-----------------------
+
+The application has a small set of command-line options:
+
+*   ``-p <port>``, ``--port <port>``
+
+    TCP port to listen on.  Default is 2002, the IANA-assigned rpcap
+    port.
+
+*   ``-b <addr>``, ``--bind <addr>``
+
+    IPv4 address to bind the listener to.  Default is ``127.0.0.1``
+    (loopback only).  Setting any other address exposes captured
+    traffic to the network and should not be done on untrusted
+    networks.
+
+*   ``-N <ring_size>``
+
+    Size of the per-session capture ring in packets.  Default is 2048.
+
+*   ``-h``, ``--help``
+
+    Print usage and exit.
+
+EAL options are supplied automatically; the application runs as a
+secondary process and does not need EAL options on its command line for
+typical use.
+
+
+Client Setup
+------------
+
+Most Linux distributions ship libpcap built without ``rpcap`` support
+because the libpcap project leaves ``--enable-remote`` off by default.
+To use ``dpdk-rpcapd`` from ``tcpdump`` or Wireshark on Linux, libpcap
+must be rebuilt with remote support enabled.  Approximate steps:
+
+.. code-block:: console
+
+    wget https://www.tcpdump.org/release/libpcap-1.10.6.tar.xz
+    tar xf libpcap-1.10.6.tar.xz
+    cd libpcap-1.10.6
+    ./configure --enable-remote
+    make
+    sudo make install
+    sudo ldconfig
+
+To verify that the resulting library has rpcap support:
+
+.. code-block:: console
+
+    nm -D /usr/local/lib/libpcap.so | grep ' T pcap_open$'
+
+The symbol ``pcap_open`` should be present.  If not, the ``--enable-remote``
+flag did not take effect.
+
+``tcpdump`` rebuilt against this libpcap can be used as a client without
+further changes.  Wireshark on Windows and macOS ships with rpcap support
+enabled by default.
+
+
+Example
+-------
+
+Start a primary application with the packet capture framework
+initialized.  ``dpdk-testpmd`` is the simplest:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-testpmd -- --no-mlockall --vdev=net_tap0
+
+In another window, start ``dpdk-rpcapd``:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-rpcapd
+    RPCAPD: listening on TCP port 2002
+
+In a third window, list available interfaces using a libpcap-based
+``tcpdump`` rebuilt with remote support:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump --list-remote-interfaces=rpcap://localhost:2002/
+    rpcap://localhost:2002/net_tap0  Network adapter 'DPDK port' on remote node localhost
+
+Capture live from a port:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -nn -c 20
+
+Or save to a file readable by any pcap consumer:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -w /tmp/capture.pcap
+
+
+Limitations
+-----------
+
+The following limitations apply to this initial version of
+``dpdk-rpcapd`` and are expected to be addressed in subsequent patches:
+
+*   **Single client.** Only one client may be connected at a time.
+    Subsequent clients are queued by the listening socket but not
+    serviced until the first disconnects.  Multi-client support
+    requires an event-driven main loop (planned).
+
+*   **No BPF filter support.** ``UPDATEFILTER`` requests are
+    acknowledged and ignored.  Capture-side filtering requires an
+    extension to ``rte_pdump`` to support filter updates on an active
+    callback.
+
+*   **No authentication.** ``AUTH`` requests are acknowledged with an
+    empty reply (libpcap "version 0, null auth" semantics).  This
+    sample application does not implement password authentication.
+
+*   **TCP transport only; not for production use.** The rpcap protocol
+    over TCP is unauthenticated and unencrypted; any client that can
+    reach the listening port has full access to captured traffic.
+    Binding to ``127.0.0.1`` by default mitigates remote exposure but
+    does not address local users on a shared host.  See the warning at
+    the top of this document.
+
+*   **Microsecond timestamp resolution.** The rpcap protocol carries
+    timestamps at microsecond resolution.
+
+
+See Also
+--------
+
+*   :doc:`extcap` -- Wireshark ``extcap`` plugin for direct integration
+    with Wireshark, without going through the rpcap protocol.
+
+*   :doc:`../tools/dumpcap` -- file-based capture writing pcapng
+    output.
+
+*   The libpcap project's ``rpcapd`` reference implementation:
+    https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
diff --git a/examples/meson.build b/examples/meson.build
index 25d9c88457..24b6184353 100644
--- a/examples/meson.build
+++ b/examples/meson.build
@@ -45,6 +45,7 @@ all_examples = [
         'ptpclient',
         'qos_meter',
         'qos_sched',
+        'rpcapd',
         'rxtx_callbacks',
         'server_node_efd/efd_node',
         'server_node_efd/efd_server',
diff --git a/examples/rpcapd/main.c b/examples/rpcapd/main.c
new file mode 100644
index 0000000000..e16f5aa3e0
--- /dev/null
+++ b/examples/rpcapd/main.c
@@ -0,0 +1,1014 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026
+ *
+ * Proof-of-concept DPDK rpcapd: the libpcap remote packet capture
+ * daemon, implemented on top of DPDK pdump.  A libpcap client (e.g.
+ * Wireshark or tcpdump using "rpcap://host[:port]/portname") can
+ * connect, list DPDK ports, open one, and stream live packets from it.
+ *
+ * Based on the DPDK dumpcap application and on rpcapd from libpcap:
+ *   https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
+ *
+ * Only the bits of the RPCAP protocol that are needed for an
+ * unauthenticated, passive-mode capture session are implemented.
+ * Configuration files, BPF filters, active mode, statistics, sampling,
+ * IPv6 and concurrent clients are intentionally omitted to keep the
+ * example small.
+ */
+
+#include <arpa/inet.h>
+#include <errno.h>
+#include <getopt.h>
+#include <netinet/in.h>
+#include <netdb.h>
+#include <poll.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/time.h>
+#include <sys/types.h>
+#include <sys/uio.h>
+#include <unistd.h>
+
+#include <rte_alarm.h>
+#include <rte_common.h>
+#include <rte_debug.h>
+#include <rte_eal.h>
+#include <rte_errno.h>
+#include <rte_ether.h>
+#include <rte_ethdev.h>
+#include <rte_lcore.h>
+#include <rte_log.h>
+#include <rte_mbuf.h>
+#include <rte_mempool.h>
+#include <rte_pdump.h>
+#include <rte_stdatomic.h>
+#include <rte_ring.h>
+#include <rte_version.h>
+
+#include "rpcap-protocol.h"
+
+#define BURST_SIZE                    32
+#define MBUF_CACHE_SIZE               32
+#define DEFAULT_RING_SIZE             2048
+#define DEFAULT_SNAPLEN               RTE_MBUF_DEFAULT_BUF_SIZE
+#define PRIMARY_MONITOR_INTERVAL_US   (500 * 1000)
+
+/* Logging.  Use --log-level=rpcapd:debug to enable debug output. */
+RTE_LOG_REGISTER(rpcapd_logtype, rpcapd, INFO);
+#define RTE_LOGTYPE_RPCAPD rpcapd_logtype
+
+/* Per-client capture session state. */
+struct session {
+	int      ctrl_fd;
+	int      data_fd;
+	uint16_t port;				/* DPDK ethdev port being captured */
+	char     name[RTE_ETH_NAME_MAX_LEN];
+	uint32_t snaplen;
+	uint32_t npkt;				/* packet sequence for rpcap_pkthdr */
+	bool     capture_on;
+	struct rte_ring    *ring;
+	struct rte_mempool *mp;
+};
+
+/* Command-line options */
+static uint16_t listen_port = RPCAP_DEFAULT_NETPORT;
+static uint32_t ring_size = DEFAULT_RING_SIZE;
+static const char *lcore_arg;
+static const char *file_prefix;
+static const char *bind_arg;		/* -b argument, resolved after option parsing */
+static const char *debug_file;		/* --debug-file argument */
+static bool ipv4_only;			/* -4: restrict to IPv4 */
+static bool debug_log;			/* -D: enable RPCAPD debug logging */
+
+/* Bind address for the listener and the per-session data port.
+ * Defaults to IPv4 loopback because the rpcap protocol is insecure.
+ * It exposes captured traffic to anyone who can reach the port.
+ * An operator who knowingly accepts that risk can override with
+ * --bind <addr>.  IPv4 and IPv6 numeric addresses are both accepted.
+ */
+static struct sockaddr_storage listen_addr;
+static socklen_t               listen_addrlen;
+
+static void stop_capture(struct session *s);
+
+static void
+set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port)
+{
+	if (ss->ss_family == AF_INET6)
+		((struct sockaddr_in6 *)ss)->sin6_port = htons(port);
+	else
+		((struct sockaddr_in *)ss)->sin_port = htons(port);
+}
+
+static uint16_t
+get_sockaddr_port(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET6)
+		return ntohs(((const struct sockaddr_in6 *)ss)->sin6_port);
+	return ntohs(((const struct sockaddr_in *)ss)->sin_port);
+}
+
+static bool
+is_loopback(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET) {
+		const struct sockaddr_in *sin = (const void *)ss;
+
+		return (ntohl(sin->sin_addr.s_addr) >> 24) == 127;
+	}
+	if (ss->ss_family == AF_INET6) {
+		const struct sockaddr_in6 *sin6 = (const void *)ss;
+
+		return IN6_IS_ADDR_LOOPBACK(&sin6->sin6_addr);
+	}
+	return false;
+}
+
+static void
+parse_bind_addr(const char *str, int family)
+{
+	struct addrinfo hints = {
+		.ai_family   = family,
+		.ai_socktype = SOCK_STREAM,
+		.ai_flags    = AI_NUMERICHOST | AI_PASSIVE,
+	};
+	struct addrinfo *res;
+	int rc;
+
+	rc = getaddrinfo(str, NULL, &hints, &res);
+	if (rc != 0)
+		rte_exit(EXIT_FAILURE, "Invalid bind address '%s': %s\n",
+			 str, gai_strerror(rc));
+	memcpy(&listen_addr, res->ai_addr, res->ai_addrlen);
+	listen_addrlen = res->ai_addrlen;
+	freeaddrinfo(res);
+}
+
+static RTE_ATOMIC(bool) quit_signal;
+
+static void
+signal_handler(int sig __rte_unused)
+{
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+/* Read exactly len bytes; return 0 on success, -1 on error or EOF. */
+static int
+recv_full(int fd, void *buf, size_t len)
+{
+	uint8_t *p = buf;
+
+	while (len > 0) {
+		ssize_t n = recv(fd, p, len, 0);
+		if (rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed))
+			return -1;
+
+		if (n < 0 && errno == EINTR)
+			continue;
+
+		if (n <= 0)
+			return -1;
+
+		p += n;
+		len -= n;
+	}
+	return 0;
+}
+
+static int
+send_iov_full(int fd, struct iovec *iov, int iovcnt, int flags)
+{
+	struct msghdr msg = {
+		.msg_iov    = iov,
+		.msg_iovlen = iovcnt,
+	};
+
+	while (sendmsg(fd, &msg, flags | MSG_NOSIGNAL) < 0) {
+		if (errno != EINTR)
+			return -1;
+	}
+	return 0;
+}
+
+static int
+rpcap_send_msg(int fd, uint8_t type, uint16_t value, const void *payload, uint32_t plen)
+{
+	struct rpcap_header hdr = {
+		.ver = RPCAP_VERSION,
+		.type = type,
+		.value = htons(value),
+		.plen = htonl(plen),
+	};
+	struct iovec iov[2] = {
+		{ .iov_base = &hdr,                        .iov_len = sizeof(hdr) },
+		{ .iov_base = (void *)(uintptr_t)payload,  .iov_len = plen },
+	};
+
+	return send_iov_full(fd, iov, plen > 0 ? 2 : 1, 0);
+}
+
+static int
+rpcap_send_error(int fd, uint16_t errcode, const char *msg)
+{
+	RTE_LOG(WARNING, RPCAPD, "sending error to client: %s\n", msg);
+	return rpcap_send_msg(fd, RPCAP_MSG_ERROR, errcode, msg, strlen(msg));
+}
+
+static int
+rpcap_recv_header(int fd, struct rpcap_header *hdr)
+{
+	if (recv_full(fd, hdr, sizeof(*hdr)) < 0)
+		return -1;
+	hdr->value = ntohs(hdr->value);
+	hdr->plen = ntohl(hdr->plen);
+	return 0;
+}
+
+/* Throw away plen bytes of payload we don't care about. */
+static int
+rpcap_discard(int fd, uint32_t plen)
+{
+	uint8_t buf[256];
+
+	while (plen > 0) {
+		size_t chunk = plen > sizeof(buf) ? sizeof(buf) : plen;
+
+		if (recv_full(fd, buf, chunk) < 0)
+			return -1;
+		plen -= chunk;
+	}
+	return 0;
+}
+
+/* Build and send the list of available DPDK ports. */
+static int
+handle_findallif(int fd)
+{
+	uint8_t *buf = NULL;
+	size_t buflen = 0;
+	uint16_t nif = 0;
+	uint16_t p;
+	int rc;
+
+	RTE_ETH_FOREACH_DEV(p) {
+		static const char desc[] = "DPDK port";
+		char name[RTE_ETH_NAME_MAX_LEN];
+		size_t namelen, desclen, entry;
+		uint8_t *nb;
+
+		if (rte_eth_dev_get_name_by_port(p, name) < 0) {
+			RTE_LOG(INFO, RPCAPD, "can not find name for port %u\n", p);
+			continue;
+		}
+
+		RTE_LOG(INFO, RPCAPD, "findallif: port %u -> '%s'\n", p, name);
+		namelen = strlen(name);
+		desclen = strlen(desc);
+		entry = sizeof(struct rpcap_findalldevs_if) + namelen + desclen;
+
+		nb = realloc(buf, buflen + entry);
+		if (nb == NULL) {
+			RTE_LOG(ERR, RPCAPD, "out of memory in findallif\n");
+			free(buf);
+			return rpcap_send_error(fd, 0, "out of memory");
+		}
+		buf = nb;
+
+		struct rpcap_findalldevs_if iface = {
+			.namelen = htons(namelen),
+			.desclen = htons(desclen),
+			.flags = htonl(PCAP_IF_UP | PCAP_IF_RUNNING),
+		};
+		memcpy(buf + buflen, &iface, sizeof(iface));
+		memcpy(buf + buflen + sizeof(iface), name, namelen);
+		memcpy(buf + buflen + sizeof(iface) + namelen, desc, desclen);
+		buflen += entry;
+		nif++;
+	}
+
+	RTE_LOG(INFO, RPCAPD, "findallif: %u interface(s)\n", nif);
+	rc = rpcap_send_msg(fd, RPCAP_MSG_FINDALLIF_REPLY, nif, buf, buflen);
+	free(buf);
+	return rc;
+}
+
+/* OPEN_REQ: payload is the interface name (no NUL). */
+static int
+handle_open(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_openreply reply = {
+		.linktype = htonl(DLT_EN10MB),
+	};
+	uint16_t port;
+
+	if (s->capture_on)
+		stop_capture(s);
+
+	if (plen >= sizeof(s->name)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "interface name too long");
+	}
+	if (recv_full(fd, s->name, plen) < 0)
+		return -1;
+	s->name[plen] = '\0';
+
+	if (rte_eth_dev_get_port_by_name(s->name, &port) < 0) {
+		RTE_LOG(WARNING, RPCAPD, "open: no such port '%s'\n", s->name);
+		return rpcap_send_error(fd, 0, "unknown interface");
+	}
+	s->port = port;
+
+	RTE_LOG(DEBUG, RPCAPD, "open: '%s' -> dpdk port %u\n", s->name, port);
+	return rpcap_send_msg(fd, RPCAP_MSG_OPEN_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Open an ephemeral TCP listening socket; return fd, set *port_out. */
+static int
+open_data_listener(uint16_t *port_out)
+{
+	struct sockaddr_storage addr = listen_addr;
+	socklen_t alen;
+	int fd;
+
+	set_sockaddr_port(&addr, 0);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0) {
+		RTE_LOG(ERR, RPCAPD, "data socket: %s\n", strerror(errno));
+		return -1;
+	}
+
+	alen = listen_addrlen;
+	if (bind(fd, (struct sockaddr *)&addr, alen) < 0 ||
+	    listen(fd, 1) < 0 ||
+	    getsockname(fd, (struct sockaddr *)&addr, &alen) < 0) {
+		RTE_LOG(ERR, RPCAPD, "data port bind/listen: %s\n", strerror(errno));
+		close(fd);
+		return -1;
+	}
+	*port_out = get_sockaddr_port(&addr);
+	return fd;
+}
+
+static struct rte_ring *
+create_capture_ring(uint16_t port)
+{
+	char name[RTE_RING_NAMESIZE];
+
+	snprintf(name, sizeof(name), "rpcapd_r_%u_%d", port, getpid());
+	return rte_ring_create(name, ring_size, rte_socket_id(), 0);
+}
+
+static struct rte_mempool *
+create_capture_mempool(uint16_t port, uint32_t snaplen)
+{
+	char name[RTE_MEMPOOL_NAMESIZE];
+	uint32_t mbuf_size = RTE_PKTMBUF_HEADROOM + snaplen;
+
+	snprintf(name, sizeof(name), "rpcapd_p_%u_%d", port, getpid());
+	return rte_pktmbuf_pool_create(name, ring_size * 2, MBUF_CACHE_SIZE, 0,
+				       mbuf_size, rte_socket_id());
+}
+
+/* Tear down anything that handle_startcap brought up.  Safe to call
+ * after partial setup as well as after a successful capture.
+ */
+static void
+stop_capture(struct session *s)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int n;
+
+	if (s->capture_on) {
+		rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, RTE_PDUMP_FLAG_RXTX);
+		RTE_LOG(INFO, RPCAPD, "capture stopped on %s (%u packets)\n",
+			s->name, s->npkt);
+	}
+	s->capture_on = false;
+
+	if (s->ring != NULL) {
+		while ((n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts,
+						      BURST_SIZE, NULL)) > 0)
+			rte_pktmbuf_free_bulk(pkts, n);
+		rte_ring_free(s->ring);
+		s->ring = NULL;
+	}
+	if (s->mp != NULL) {
+		rte_mempool_free(s->mp);
+		s->mp = NULL;
+	}
+	if (s->data_fd >= 0) {
+		close(s->data_fd);
+		s->data_fd = -1;
+	}
+}
+
+/*
+ * STARTCAP_REQ: open the data connection and arm the pdump callback.
+ * We use passive mode with the server-allocated data port:
+ *   - the server picks an ephemeral port and listens on it
+ *   - the server returns that port in startcapreply.portdata
+ *   - the client connects back to that port for the packet stream
+ */
+static int
+handle_startcap(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_startcapreq req;
+	uint16_t data_port;
+	int data_listen;
+	int data_fd;
+
+	if (s->capture_on)
+		stop_capture(s);
+
+	if (plen < sizeof(req)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "short startcap request");
+	}
+	if (recv_full(fd, &req, sizeof(req)) < 0)
+		return -1;
+	/* Skip any embedded BPF filter; not supported here. */
+	if (rpcap_discard(fd, plen - sizeof(req)) < 0)
+		return -1;
+
+	s->snaplen = ntohl(req.snaplen);
+	if (s->snaplen == 0 || s->snaplen > DEFAULT_SNAPLEN)
+		s->snaplen = DEFAULT_SNAPLEN;
+
+	s->ring = create_capture_ring(s->port);
+	s->mp = create_capture_mempool(s->port, s->snaplen);
+	if (s->ring == NULL || s->mp == NULL) {
+		RTE_LOG(ERR, RPCAPD, "ring/mempool alloc failed: %s\n",
+			rte_strerror(rte_errno));
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "DPDK alloc failed");
+	}
+
+	data_listen = open_data_listener(&data_port);
+	if (data_listen < 0) {
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "data port setup failed");
+	}
+
+	struct rpcap_startcapreply reply = {
+		.bufsize = htonl(s->snaplen * BURST_SIZE),
+		.portdata = htons(data_port),
+	};
+	if (rpcap_send_msg(fd, RPCAP_MSG_STARTCAP_REPLY, 0, &reply, sizeof(reply)) < 0) {
+		close(data_listen);
+		stop_capture(s);
+		return -1;
+	}
+
+	RTE_LOG(INFO, RPCAPD, "awaiting connection\n");
+
+	data_fd = accept(data_listen, NULL, NULL);
+	close(data_listen);
+	if (data_fd < 0) {
+		RTE_LOG(ERR, RPCAPD, "accept on data port: %s\n", strerror(errno));
+		stop_capture(s);
+		return -1;
+	}
+
+	s->data_fd = data_fd;
+
+	if (rte_pdump_enable(s->port, RTE_PDUMP_ALL_QUEUES, RTE_PDUMP_FLAG_RXTX,
+			     s->ring, s->mp, NULL) < 0) {
+		RTE_LOG(ERR, RPCAPD, "rte_pdump_enable port %u failed: %s\n",
+			s->port, rte_strerror(rte_errno));
+		stop_capture(s);
+		return -1;
+	}
+	s->capture_on = true;
+	s->npkt = 0;
+
+	RTE_LOG(INFO, RPCAPD,
+		"capture started on %s (snaplen %u, data port %u)\n",
+		s->name, s->snaplen, data_port);
+	return 0;
+}
+
+/*
+ * Pull a burst from the ring, frame each packet into an RPCAP_MSG_PACKET
+ * message, and send it on the data connection.  MSG_MORE on all but the
+ * last send tells the kernel to coalesce the burst into full segments.
+ */
+static int
+process_ring(struct session *s)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int i, n;
+
+	n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts, BURST_SIZE, NULL);
+	if (n == 0)
+		return 0;
+
+	for (i = 0; i < n; i++) {
+		struct rte_mbuf *m = pkts[i];
+		uint8_t buf[RTE_ETHER_MAX_JUMBO_FRAME_LEN];
+		uint32_t pktlen = rte_pktmbuf_pkt_len(m);
+		uint32_t caplen = pktlen < s->snaplen ? pktlen : s->snaplen;
+		const void *data;
+		struct timeval tv;
+
+		s->npkt++;
+
+		struct rpcap_header hdr = {
+			.ver = RPCAP_VERSION,
+			.type = RPCAP_MSG_PACKET,
+			.plen = htonl(sizeof(struct rpcap_pkthdr) + caplen),
+		};
+
+		gettimeofday(&tv, NULL);
+
+		struct rpcap_pkthdr pkthdr = {
+			.timestamp_sec = htonl((uint32_t)tv.tv_sec),
+			.timestamp_usec = htonl((uint32_t)tv.tv_usec),
+			.caplen = htonl(caplen),
+			.len = htonl(pktlen),
+			.npkt = htonl(s->npkt),
+		};
+
+		data = rte_pktmbuf_read(m, 0, caplen, buf);
+
+		struct iovec iov[3] = {
+			{ .iov_base = &hdr,                       .iov_len = sizeof(hdr) },
+			{ .iov_base = &pkthdr,                    .iov_len = sizeof(pkthdr) },
+			{ .iov_base = (void *)(uintptr_t)data,    .iov_len = caplen },
+		};
+		if (send_iov_full(s->data_fd, iov, 3,
+				  i + 1 < n ? MSG_MORE : 0) < 0) {
+			RTE_LOG(NOTICE, RPCAPD, "data connection closed: %s\n", strerror(errno));
+			goto error;
+		}
+		rte_pktmbuf_free(m);
+	}
+
+	return (int)n;
+
+error:
+	rte_pktmbuf_free_bulk(pkts + i, n - i);
+	return -1;
+}
+
+/*
+ * Stay in the capture loop until either:
+ *   - a control message arrives (typically ENDCAP),
+ *   - the data connection breaks, or
+ *   - a quit signal is delivered.
+ */
+static int
+capture_loop(int ctrl_fd, struct session *s)
+{
+	struct pollfd pfd = { .fd = ctrl_fd, .events = POLLIN };
+	unsigned int idle = 0;
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int n;
+
+		if (poll(&pfd, 1, 0) > 0 && (pfd.revents & POLLIN))
+			return 0;
+
+		n = process_ring(s);
+		if (n < 0)
+			return -1;
+		if (n == 0) {
+			if (idle++ < 1000)
+				continue;
+			usleep(1000);
+			idle = 0;
+		} else {
+			idle = 0;
+		}
+	}
+	return 0;
+}
+
+static int
+handle_endcap(int fd, uint32_t plen, struct session *s)
+{
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+	stop_capture(s);
+	return rpcap_send_msg(fd, RPCAP_MSG_ENDCAP_REPLY, 0, NULL, 0);
+}
+
+static int
+handle_stats(int fd, uint32_t plen, const struct session *s)
+{
+	struct rte_eth_stats es = { 0 };
+
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+
+	if (s->capture_on)
+		rte_eth_stats_get(s->port, &es);
+
+	struct rpcap_stats reply = {
+		.ifrecv   = htonl((uint32_t)es.ipackets),
+		.ifdrop   = htonl((uint32_t)es.ierrors),
+		.krnldrop = 0,
+		.svrcapt  = htonl(s->npkt),
+	};
+	return rpcap_send_msg(fd, RPCAP_MSG_STATS_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Service a single client until it disconnects. */
+static void
+handle_client(int ctrl_fd)
+{
+	struct sockaddr_storage peer;
+	socklen_t plen = sizeof(peer);
+	char host[NI_MAXHOST] = "?";
+	struct session s = { .ctrl_fd = ctrl_fd, .data_fd = -1 };
+
+	if (getpeername(ctrl_fd, (struct sockaddr *)&peer, &plen) == 0)
+		getnameinfo((struct sockaddr *)&peer, plen,
+			    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	RTE_LOG(INFO, RPCAPD, "client %s connected\n", host);
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		struct rpcap_header hdr;
+
+		if (rpcap_recv_header(ctrl_fd, &hdr) < 0)
+			break;
+
+		switch (hdr.type) {
+		case RPCAP_MSG_AUTH_REQ:
+			/* No auth: discard credentials, ack with empty reply.
+			 * libpcap treats a zero-length AUTH_REPLY as "version
+			 * 0 only, same byte order".
+			 */
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_msg(ctrl_fd, RPCAP_MSG_AUTH_REPLY, 0, NULL, 0) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_FINDALLIF_REQ:
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 || handle_findallif(ctrl_fd) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_OPEN_REQ:
+			if (handle_open(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STARTCAP_REQ:
+			if (handle_startcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			if (capture_loop(ctrl_fd, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_UPDATEFILTER_REQ:
+			/* Filters not implemented; ack and ignore. */
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_msg(ctrl_fd, RPCAP_MSG_UPDATEFILTER_REPLY,
+					   0, NULL, 0) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_ENDCAP_REQ:
+			if (handle_endcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STATS_REQ:
+			if (handle_stats(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_CLOSE:
+			rpcap_discard(ctrl_fd, hdr.plen);
+			goto done;
+		default:
+			RTE_LOG(WARNING, RPCAPD, "unsupported request type 0x%02x\n", hdr.type);
+			rpcap_discard(ctrl_fd, hdr.plen);
+			rpcap_send_error(ctrl_fd, 0, "unsupported request");
+			break;
+		}
+	}
+done:
+	stop_capture(&s);
+	close(ctrl_fd);
+	RTE_LOG(INFO, RPCAPD, "client %s disconnected\n", host);
+}
+
+static int
+open_listen_socket(uint16_t port)
+{
+	struct sockaddr_storage addr = listen_addr;
+	char host[NI_MAXHOST];
+	int fd, one = 1;
+
+	set_sockaddr_port(&addr, port);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0)
+		rte_exit(EXIT_FAILURE, "socket: %s\n", strerror(errno));
+	setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one));
+
+	if (bind(fd, (struct sockaddr *)&addr, listen_addrlen) < 0)
+		rte_exit(EXIT_FAILURE, "bind(%u): %s\n", port, strerror(errno));
+
+	int err = getnameinfo((struct sockaddr *)&listen_addr, listen_addrlen,
+			      host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	if (err != 0)
+		rte_exit(EXIT_FAILURE, "Listen address lookup failed: %s\n",
+			 gai_strerror(err));
+
+	RTE_LOG(INFO, RPCAPD, "listening on %s port %u\n", host, listen_port);
+
+	if (!is_loopback(&listen_addr))
+		RTE_LOG(WARNING, RPCAPD,
+			"bound to non-loopback address %s; "
+			"rpcap is unauthenticated and unencrypted, "
+			"captured traffic is exposed to the network\n",
+			host);
+
+	if (listen(fd, 1) < 0)
+		rte_exit(EXIT_FAILURE, "listen: %s\n", strerror(errno));
+
+	return fd;
+}
+
+static void
+usage(FILE *f, const char *progname)
+{
+	fprintf(f, "Usage: %s [options]\n", progname);
+	fprintf(f,
+		"  -p, --port <port>     listen port (default %u)\n"
+		"  -b, --bind <addr>     bind address (default 127.0.0.1)\n"
+		"  -4                    use only IPv4 (reject IPv6 bind addresses)\n"
+		"  -N <ring size>        ring size in packets (default %u)\n"
+		"  -D, --debug           enable rpcapd debug log messages\n"
+		"      --debug-file <f>  redirect log output to file <f> (append mode)\n"
+		"      --version         print version and exit\n"
+		"  -h, --help            print this help and exit\n"
+		"      --lcore=<core>    CPU core to run on (default: any)\n"
+		"      --file-prefix=<p> prefix to use for multi-process\n"
+		"\n"
+		"WARNING: rpcap is unauthenticated and unencrypted.  Binding to\n"
+		"any non-loopback address exposes captured traffic to the\n"
+		"network.  Sample application; not for production use.\n",
+		RPCAP_DEFAULT_NETPORT, DEFAULT_RING_SIZE);
+}
+
+static void
+print_version(void)
+{
+	printf("rpcapd, a remote packet capture daemon (DPDK pdump backend)\n"
+	       "Built against %s\n", rte_version());
+}
+
+static void
+parse_opts(int argc, char **argv)
+{
+	enum {
+		OPT_LONG_ONLY = 0x100,
+		OPT_DEBUG_FILE,
+		OPT_VERSION,
+	};
+	static const struct option long_options[] = {
+		{ "port",        required_argument, NULL, 'p' },
+		{ "bind",        required_argument, NULL, 'b' },
+		{ "debug",       no_argument,       NULL, 'D' },
+		{ "help",        no_argument,       NULL, 'h' },
+		{ "version",     no_argument,       NULL, OPT_VERSION },
+		{ "debug-file",  required_argument, NULL, OPT_DEBUG_FILE },
+		{ "file-prefix", required_argument, NULL, 0 },
+		{ "lcore",       required_argument, NULL, 0 },
+		{ NULL, 0, NULL, 0 },
+	};
+	int option_index, c;
+
+	while ((c = getopt_long(argc, argv, "hD4p:b:N:",
+				long_options, &option_index)) != -1) {
+		switch (c) {
+		case 'p': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			if (u == 0 || u > UINT16_MAX)
+				rte_exit(EXIT_FAILURE, "Invalid port: %s\n", optarg);
+			listen_port = (uint16_t)u;
+			break;
+		}
+		case 'b':
+			bind_arg = optarg;
+			break;
+		case '4':
+			ipv4_only = true;
+			break;
+		case 'N':
+			ring_size = strtoul(optarg, NULL, 0);
+			if (ring_size < 64)
+				rte_exit(EXIT_FAILURE, "Ring size too small\n");
+			break;
+		case 'D':
+			debug_log = true;
+			break;
+		case 'h':
+			usage(stdout, argv[0]);
+			exit(0);
+		case OPT_VERSION:
+			print_version();
+			exit(0);
+		case OPT_DEBUG_FILE:
+			debug_file = optarg;
+			break;
+		case 0: {
+			const char *longopt = long_options[option_index].name;
+
+			if (!strcmp(longopt, "lcore")) {
+				lcore_arg = optarg;
+				break;
+			} else if (!strcmp(longopt, "file-prefix")) {
+				file_prefix = optarg;
+				break;
+			}
+		}
+			/* fallthrough */
+		default:
+			usage(stderr, argv[0]);
+			exit(EXIT_FAILURE);
+		}
+	}
+
+	/* Resolve the bind address now that -4 has been seen. */
+	parse_bind_addr(bind_arg ? bind_arg : "127.0.0.1",
+			ipv4_only ? AF_INET : AF_UNSPEC);
+}
+
+/*
+ * Periodic check that the DPDK primary process is still alive.
+ * If it dies our shared-memory state (rings, mempools, pdump) becomes
+ * unsafe to touch, so we set quit_signal and let the main loop tear
+ * down cleanly on its next iteration.  The callback runs on the EAL
+ * interrupt thread; quit_signal is atomic so the read in the main
+ * loop is well-defined.
+ */
+static void
+monitor_primary(void *arg __rte_unused)
+{
+	if (rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed))
+		return;
+
+	if (rte_eal_primary_proc_alive(NULL)) {
+		rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL);
+		return;
+	}
+
+	RTE_LOG(NOTICE, RPCAPD, "primary process exited, shutting down\n");
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+static void
+enable_primary_monitor(void)
+{
+	if (rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL) < 0)
+		RTE_LOG(WARNING, RPCAPD, "failed to install primary process monitor\n");
+}
+
+static void
+disable_primary_monitor(void)
+{
+	rte_eal_alarm_cancel(monitor_primary, NULL);
+}
+
+/*
+ * Bring up EAL as a secondary process so that pdump can attach to a
+ * running primary DPDK application.  Mirrors dumpcap's approach: the
+ * RPCAP user sees a small set of options (port, ring size) rather
+ * than the full DPDK EAL command line.
+ */
+static int
+dpdk_init(void)
+{
+	static const char * const args[] = {
+		"rpcapd",
+		"--proc-type", "secondary",
+		"--log-level", "info",        /* EAL stays quiet */
+	};
+	int eal_argc = RTE_DIM(args);
+	rte_cpuset_t cpuset = { };
+	char **eal_argv;
+	unsigned int i;
+
+	if (file_prefix != NULL)
+		eal_argc += 2;
+
+	if (lcore_arg != NULL)
+		eal_argc += 2;
+
+	eal_argv = calloc(eal_argc + 1, sizeof(char *));
+	if (eal_argv == NULL)
+		return -1;
+
+	for (i = 0; i < RTE_DIM(args); i++) {
+		eal_argv[i] = strdup(args[i]);
+		if (eal_argv[i] == NULL)
+			return -1;
+	}
+
+	if (file_prefix != NULL && *file_prefix != '\0') {
+		eal_argv[i++] = strdup("--file-prefix");
+		eal_argv[i++] = strdup(file_prefix);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+
+	if (lcore_arg != NULL) {
+		eal_argv[i++] = strdup("--lcores");
+		eal_argv[i++] = strdup(lcore_arg);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+	eal_argc = i;
+
+	/*
+	 * Need to get the original cpuset, before EAL init changes
+	 * the affinity of this thread (main lcore).
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_get_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		rte_panic("rte_thread_getaffinity failed\n");
+
+	if (rte_eal_init(eal_argc, eal_argv) < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failed: is the primary process running?\n");
+
+	/*
+	 * If no lcore argument was specified,
+	 * then run this program as a normal process
+	 * which can be scheduled on any non-isolated CPU.
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_set_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		RTE_LOG(INFO, RPCAPD,
+			 "Can not restore original CPU affinity\n");
+
+	if (rte_pdump_init() < 0)
+		rte_exit(EXIT_FAILURE, "rte_pdump_init failed\n");
+
+	return 0;
+}
+
+int
+main(int argc, char **argv)
+{
+	struct sigaction action = {
+		.sa_handler = signal_handler,
+	};
+	int srv_fd;
+
+	parse_opts(argc, argv);
+
+	/*
+	 * Redirect log output before EAL init so EAL's own messages are
+	 * captured too.  The FILE handle is intentionally never closed:
+	 * the kernel reclaims it at process exit.
+	 */
+	if (debug_file != NULL) {
+		FILE *fp = fopen(debug_file, "a");
+
+		if (fp == NULL)
+			rte_exit(EXIT_FAILURE, "Cannot open debug file '%s': %s\n",
+				 debug_file, strerror(errno));
+		setvbuf(fp, NULL, _IOLBF, 0);
+		rte_openlog_stream(fp);
+	}
+
+	if (dpdk_init() < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failure\n");
+
+	if (debug_log)
+		rte_log_set_level(rpcapd_logtype, RTE_LOG_DEBUG);
+
+	if (rte_eth_dev_count_avail() == 0)
+		rte_exit(EXIT_FAILURE, "No Ethernet ports found\n");
+
+	sigaction(SIGTERM, &action, NULL);
+	sigaction(SIGINT, &action, NULL);
+
+	/* If peer closes, this detected in next recv() */
+	signal(SIGPIPE, SIG_IGN);
+
+	srv_fd = open_listen_socket(listen_port);
+
+	enable_primary_monitor();
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int cfd = accept(srv_fd, NULL, NULL);
+		if (cfd < 0) {
+			if (errno == EINTR)
+				continue;
+			RTE_LOG(ERR, RPCAPD, "accept: %s\n", strerror(errno));
+			break;
+		}
+		handle_client(cfd);
+	}
+
+	disable_primary_monitor();
+	RTE_LOG(INFO, RPCAPD, "shutting down\n");
+	close(srv_fd);
+	rte_pdump_uninit();
+	return rte_eal_cleanup() ? EXIT_FAILURE : 0;
+}
diff --git a/examples/rpcapd/meson.build b/examples/rpcapd/meson.build
new file mode 100644
index 0000000000..5eb1a8487c
--- /dev/null
+++ b/examples/rpcapd/meson.build
@@ -0,0 +1,11 @@
+# SPDX-License-Identifier: BSD-3-Clause
+# Copyright(c) 2026
+
+if is_windows
+    build = false
+    reason = 'not supported on Windows'
+    subdir_done()
+endif
+
+sources = files('main.c')
+deps += ['ethdev', 'pdump']
diff --git a/examples/rpcapd/rpcap-protocol.h b/examples/rpcapd/rpcap-protocol.h
new file mode 100644
index 0000000000..46d87928aa
--- /dev/null
+++ b/examples/rpcapd/rpcap-protocol.h
@@ -0,0 +1,96 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026
+ *
+ * On-the-wire RPCAP protocol definitions, transcribed from libpcap's
+ * rpcap-protocol.h.  See:
+ *   https://github.com/the-tcpdump-group/libpcap/blob/master/rpcap-protocol.h
+ *
+ * Only the subset needed by the DPDK rpcd POC is included here.  All
+ * multi-byte fields in the structures below are big-endian on the wire.
+ */
+
+#ifndef _RPCAP_PROTOCOL_H_
+#define _RPCAP_PROTOCOL_H_
+
+#include <stdint.h>
+
+#define RPCAP_VERSION              0
+#define RPCAP_DEFAULT_NETPORT      2002
+
+/* Message types */
+#define RPCAP_MSG_ERROR            0x01
+#define RPCAP_MSG_FINDALLIF_REQ    0x02
+#define RPCAP_MSG_OPEN_REQ         0x03
+#define RPCAP_MSG_STARTCAP_REQ     0x04
+#define RPCAP_MSG_UPDATEFILTER_REQ 0x05
+#define RPCAP_MSG_CLOSE            0x06
+#define RPCAP_MSG_PACKET           0x07
+#define RPCAP_MSG_AUTH_REQ         0x08
+#define RPCAP_MSG_STATS_REQ        0x09
+#define RPCAP_MSG_ENDCAP_REQ       0x0a
+#define RPCAP_MSG_IS_REPLY         0x80
+
+#define RPCAP_MSG_FINDALLIF_REPLY    (RPCAP_MSG_FINDALLIF_REQ    | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_OPEN_REPLY         (RPCAP_MSG_OPEN_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STARTCAP_REPLY     (RPCAP_MSG_STARTCAP_REQ     | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_UPDATEFILTER_REPLY (RPCAP_MSG_UPDATEFILTER_REQ | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_AUTH_REPLY         (RPCAP_MSG_AUTH_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_ENDCAP_REPLY       (RPCAP_MSG_ENDCAP_REQ       | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STATS_REPLY	     (RPCAP_MSG_STATS_REQ	 | RPCAP_MSG_IS_REPLY)
+
+/* Subset of pcap interface flags (pcap.h) */
+#define PCAP_IF_UP                 0x00000002
+#define PCAP_IF_RUNNING            0x00000004
+
+/* DLT_EN10MB - ethernet, the only link type we report */
+#define DLT_EN10MB                 1
+
+struct rpcap_header {
+	uint8_t  ver;
+	uint8_t  type;
+	uint16_t value;
+	uint32_t plen;
+};
+
+struct rpcap_findalldevs_if {
+	uint16_t namelen;
+	uint16_t desclen;
+	uint32_t flags;
+	uint16_t naddr;
+	uint16_t dummy;
+};
+
+struct rpcap_openreply {
+	int32_t  linktype;
+	int32_t  tzoff;
+};
+
+struct rpcap_startcapreq {
+	uint32_t snaplen;
+	uint32_t read_timeout;
+	uint16_t flags;
+	uint16_t portdata;
+};
+
+struct rpcap_startcapreply {
+	int32_t  bufsize;
+	uint16_t portdata;
+	uint16_t dummy;
+};
+
+struct rpcap_stats {
+	uint32_t ifrecv;
+	uint32_t ifdrop;
+	uint32_t krnldrop;
+	uint32_t svrcapt;
+};
+
+struct rpcap_pkthdr {
+	uint32_t timestamp_sec;
+	uint32_t timestamp_usec;
+	uint32_t caplen;
+	uint32_t len;
+	uint32_t npkt;
+};
+
+#endif /* _RPCAP_PROTOCOL_H_ */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-08 21:07 [PATCH] examples/rpcapd: demo version of packet capture daemon Stephen Hemminger
@ 2026-09-20 18:59 ` Stephen Hemminger
  2026-09-21  9:51   ` Marat Khalili
  2026-09-22 18:45   ` Stephen Hemminger
  2026-09-22 21:31 ` [PATCH v3] " Stephen Hemminger
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
  2 siblings, 2 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-20 18:59 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Thomas Monjalon, Reshma Pattan

This example adds RPCAP support over localhost TCP
integrated with DPDK. It uses a secondary process that allows
connections from using tcpdump defacto protocol rpcap.

See: doc/guides/sample_app_ug/rpcapd.rst for more info

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---
v2 - Build on Linux only
   - add filter support
   - fix lots of AI review feedback

 MAINTAINERS                            |    2 +
 doc/guides/rel_notes/release_26_11.rst |    4 +
 doc/guides/sample_app_ug/index.rst     |    1 +
 doc/guides/sample_app_ug/rpcapd.rst    |  216 ++++
 examples/meson.build                   |    1 +
 examples/rpcapd/main.c                 | 1430 ++++++++++++++++++++++++
 examples/rpcapd/meson.build            |   19 +
 examples/rpcapd/rpcap-protocol.h       |  127 +++
 8 files changed, 1800 insertions(+)
 create mode 100644 doc/guides/sample_app_ug/rpcapd.rst
 create mode 100644 examples/rpcapd/main.c
 create mode 100644 examples/rpcapd/meson.build
 create mode 100644 examples/rpcapd/rpcap-protocol.h

diff --git a/MAINTAINERS b/MAINTAINERS
index 8c50c52933..1eca09646d 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -1724,6 +1724,8 @@ F: app/pdump/
 F: doc/guides/tools/pdump.rst
 F: app/dumpcap/
 F: doc/guides/tools/dumpcap.rst
+F: examples/rpcapd/
+F: doc/guides/sample_app_ug/rpcapd.rst
 
 
 Packet Framework
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 4b3e5d995c..958f5ebc16 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,10 @@ New Features
      Also, make sure to start the actual text at the margin.
      =======================================================
 
+* **Added an example of tcpdump remote pcap daemon.**
+
+  Added an example that implements rpcap to allow live capture in tcpdump.
+
 
 Removed Items
 -------------
diff --git a/doc/guides/sample_app_ug/index.rst b/doc/guides/sample_app_ug/index.rst
index f12623bb66..61ed870318 100644
--- a/doc/guides/sample_app_ug/index.rst
+++ b/doc/guides/sample_app_ug/index.rst
@@ -31,6 +31,7 @@ Sample Applications User Guides
     l3_forward_graph
     l3_forward_power_man
     link_status_intr
+    rpcapd
     server_node_efd
     service_cores
     multi_process
diff --git a/doc/guides/sample_app_ug/rpcapd.rst b/doc/guides/sample_app_ug/rpcapd.rst
new file mode 100644
index 0000000000..6afbfa216c
--- /dev/null
+++ b/doc/guides/sample_app_ug/rpcapd.rst
@@ -0,0 +1,216 @@
+..  SPDX-License-Identifier: BSD-3-Clause
+    Copyright(c) 2026 Stephen Hemminger
+
+.. _rpcapd_app:
+
+dpdk-rpcapd Sample Application
+==============================
+
+The ``dpdk-rpcapd`` sample application is a Data Plane Development Kit
+(DPDK) implementation of the remote packet capture daemon protocol
+(``rpcap``) used by libpcap.  It runs as a DPDK secondary process and
+allows libpcap-aware tools such as ``tcpdump`` and Wireshark to capture
+packets from a DPDK primary process live, without writing to an
+intermediate file.
+
+The ``dpdk-rpcapd`` tool implements a subset of the protocol spoken by
+the libpcap project's ``rpcapd``.
+See
+https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd for the
+reference implementation.
+Clients connect to ``dpdk-rpcapd`` using a ``rpcap://`` URL,
+request the list of available interfaces(which are the ports of the DPDK primary),
+open one, and stream packets from it.
+
+The intended workflow is one-step capture: start the primary, start
+``dpdk-rpcapd``, and point a familiar tool at it.  No intermediate files,
+no separate post-processing step.
+
+.. warning::
+
+   ``dpdk-rpcapd`` listens on an unauthenticated, unencrypted TCP port
+   (default 2002, bound to ``127.0.0.1``).  Any local user able to
+   reach the port can list DPDK ports and capture all traffic flowing
+   through them.  This is a sample application intended for
+   development, debugging, and demonstration use only.  **Do not run
+   ``dpdk-rpcapd`` on a production system.**
+
+   The default bind address is ``127.0.0.1`` so the listener is not
+   reachable from other hosts.  An operator may override this with
+   ``--bind <addr>`` but should expect that the resulting deployment
+   exposes captured traffic to anyone who can reach that address; do
+   not do this on an untrusted network.
+
+
+.. note::
+
+   * ``dpdk-rpcapd`` is experimental and provided for demonstration purposes only.
+     It may change or be removed without notice, and it is not intended to be relied upon.
+
+
+Running the Application
+-----------------------
+
+The application has a small set of command-line options:
+
+*   ``-p <port>``, ``--port <port>``
+
+    TCP port to listen on.  Default is 2002, the IANA-assigned rpcap
+    port.
+
+*   ``-b <addr>``, ``--bind <addr>``
+
+    Numeric IPv4 or IPv6 address to bind the listener to.  Default is
+    ``127.0.0.1`` (loopback only).  Setting any other address exposes
+    captured traffic to the network and should not be done on untrusted
+    networks.
+
+*   ``-4``
+
+    Use only IPv4; an IPv6 argument to ``-b`` is rejected.
+
+*   ``-N <ring_size>``
+
+    Size of the per-session capture ring in packets.  Default is 2048.
+    Rounded up to a power of two if necessary.
+
+*   ``-D``, ``--debug``
+
+    Increase log verbosity.  By default only notices, warnings and
+    errors are printed.  A single ``-D`` adds session-level messages
+    (client connected, capture started and stopped); ``-DD`` adds
+    per-request protocol detail.
+
+*   ``--debug-file <file>``
+
+    Append log output to ``<file>`` instead of writing it to standard
+    error.
+
+*   ``--lcore <core>``
+
+    CPU core to run on.  By default the daemon runs as an ordinary
+    process on any non-isolated CPU.
+
+*   ``--file-prefix <prefix>``
+
+    EAL file prefix of the primary process to attach to.  Needed when
+    the primary was started with a non-default prefix.
+
+*   ``--version``
+
+    Print the version and exit.
+
+*   ``-h``, ``--help``
+
+    Print usage and exit.
+
+EAL options are supplied automatically; the application runs as a
+secondary process and does not need EAL options on its command line for
+typical use.
+
+
+Client Setup
+------------
+
+Most Linux distributions ship libpcap built without ``rpcap`` support
+because the libpcap project leaves ``--enable-remote`` off by default.
+To use ``dpdk-rpcapd`` from ``tcpdump`` or Wireshark on Linux, libpcap
+must be rebuilt with remote support enabled.  Approximate steps:
+
+.. code-block:: console
+
+    wget https://www.tcpdump.org/release/libpcap-1.10.7.tar.xz
+    tar xf libpcap-1.10.7.tar.xz
+    cd libpcap-1.10.7
+    ./configure --enable-remote
+    make
+    sudo make install
+
+Only the client side of ``rpcap`` is used for ``dpdk-rpcapd``.
+Do not run libpcap's version of ``rpcapd``.
+
+``tcpdump`` rebuilt against this libpcap can be used as a client without
+further changes.  Wireshark on Windows and macOS ships with rpcap support
+enabled by default.
+
+
+Example
+-------
+
+Start a primary application with the packet capture framework
+initialized.  ``dpdk-testpmd`` is the simplest:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-testpmd --vdev=net_tap0 -- -i
+
+In another window, start ``dpdk-rpcapd``:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/examples/dpdk-rpcapd
+    RPCAPD: open_listen_socket(): listening on 127.0.0.1 port 2002
+
+In a third window, list available interfaces using a libpcap-based
+``tcpdump`` rebuilt with remote support:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump --list-remote-interfaces=rpcap://localhost:2002/
+    rpcap://localhost:2002/net_tap0  Network adapter 'DPDK port' on remote node localhost
+
+Capture live from a port:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -nn -c 20
+
+Or save to a file readable by any pcap consumer:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -w /tmp/capture.pcap
+
+
+Limitations
+-----------
+
+The following limitations apply to this initial version of
+``dpdk-rpcapd`` and are expected to be addressed in subsequent patches:
+
+*   **Single client.** Only one client may be connected at a time.
+    Subsequent clients are queued by the listening socket but not
+    serviced until the first disconnects.  Multi-client support
+    requires an event-driven main loop (planned).
+
+*   **No authentication.** ``AUTH`` requests are acknowledged with an
+    empty reply (libpcap "version 0, null auth" semantics).  This
+    sample application does not implement password authentication.
+
+*   **TCP transport only; not for production use.** The rpcap protocol
+    over TCP is unauthenticated and unencrypted; any client that can
+    reach the listening port has full access to captured traffic.
+    Binding to ``127.0.0.1`` by default mitigates remote exposure but
+    does not address local users on a shared host.  See the warning at
+    the top of this document.
+
+*   **Microsecond timestamp resolution.** The rpcap protocol carries
+    timestamps at microsecond resolution.
+
+*   **Original length of truncated packets is not reported.** The
+    capture framework in the primary process copies only the snaplen
+    worth of bytes and does not carry the original frame length across
+    to the secondary, so a truncated packet is reported to the client
+    with its on-the-wire length equal to its captured length.  A frame
+    longer than the snaplen therefore appears to the client as a short
+    frame rather than as a truncated long one.
+
+
+See Also
+--------
+
+*   :doc:`../tools/dumpcap` -- file-based capture writing pcapng
+    output.
+
+*   The libpcap project's ``rpcapd`` reference implementation:
+    https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
diff --git a/examples/meson.build b/examples/meson.build
index 25d9c88457..24b6184353 100644
--- a/examples/meson.build
+++ b/examples/meson.build
@@ -45,6 +45,7 @@ all_examples = [
         'ptpclient',
         'qos_meter',
         'qos_sched',
+        'rpcapd',
         'rxtx_callbacks',
         'server_node_efd/efd_node',
         'server_node_efd/efd_server',
diff --git a/examples/rpcapd/main.c b/examples/rpcapd/main.c
new file mode 100644
index 0000000000..530543ce7e
--- /dev/null
+++ b/examples/rpcapd/main.c
@@ -0,0 +1,1430 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Demonstration server for the rpcap protocol for DPDK.
+ * This allows a libpcap client (e.g. Wireshark or tcpdump)
+ * to use "rpcap://host[:port]/portname" as capture device.
+ *
+ * Based on the DPDK dumpcap application and on rpcapd from libpcap:
+ *   https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
+ *
+ * Only the bits of the RPCAP protocol that are needed for an
+ * unauthenticated, passive-mode capture session are implemented.
+ * Configuration files, active mode, sampling and concurrent clients
+ * are intentionally omitted to keep the example small.
+ *
+ * A capture filter may be sent with the start-capture request:
+ * the client compiles it, so it arrives as cBPF which is converted to
+ * DPDK BPF and handed to pdump.  Filters cannot be changed once the
+ * capture is running; see the UPDATEFILTER handling.
+ */
+
+#include <arpa/inet.h>
+#include <errno.h>
+#include <getopt.h>
+#include <netinet/in.h>
+#include <netdb.h>
+#include <poll.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/time.h>
+#include <sys/types.h>
+#include <sys/uio.h>
+#include <unistd.h>
+
+#include <pcap/pcap.h>
+
+#include <rte_alarm.h>
+#include <rte_bpf.h>
+#include <rte_common.h>
+#include <rte_debug.h>
+#include <rte_eal.h>
+#include <rte_errno.h>
+#include <rte_ethdev.h>
+#include <rte_lcore.h>
+#include <rte_log.h>
+#include <rte_malloc.h>
+#include <rte_mbuf.h>
+#include <rte_mempool.h>
+#include <rte_pdump.h>
+#include <rte_stdatomic.h>
+#include <rte_ring.h>
+#include <rte_version.h>
+
+#include "rpcap-protocol.h"
+
+#define BURST_SIZE                    32
+#define MBUF_CACHE_SIZE               32
+#define DEFAULT_RING_SIZE             2048
+#define MAX_RING_SIZE                 (1U << 20)
+#define DEFAULT_SNAPLEN               RTE_MBUF_DEFAULT_DATAROOM
+#define PRIMARY_MONITOR_INTERVAL_US   (500 * 1000)
+#define SLEEP_THRESHOLD		      100
+#define SLEEP_US		      100
+
+#define DATA_ACCEPT_TIMEOUT_MS        10000
+#define POLL_INTERVAL_MS              500
+
+#define MAX_FILTER_INSNS              4096
+
+#define RTE_LOGTYPE_RPCAPD RTE_LOGTYPE_USER1
+#define RPCAPD_LOG(level, ...) \
+	RTE_LOG_LINE_PREFIX(level, RPCAPD, "%s(): ", __func__, __VA_ARGS__)
+
+/* Per-client capture session state. */
+struct session {
+	int      data_fd;
+	uint16_t port;				/* DPDK ethdev port being captured */
+	char     name[RTE_ETH_NAME_MAX_LEN];
+	uint32_t snaplen;
+	uint32_t npkt;				/* packet sequence for rpcap_pkthdr */
+	uint32_t pdump_flags;			/* RTE_PDUMP_FLAG_* in use */
+	bool     opened;			/* OPEN_REQ has selected a port */
+	bool     capture_on;
+	bool     promisc_set;			/* we enabled promiscuous mode */
+	struct rte_ring    *ring;
+	struct rte_mempool *mp;
+	struct rte_bpf_prm *prm;		/* capture filter, NULL if none */
+};
+
+/* Command-line options */
+static uint16_t listen_port = RPCAP_DEFAULT_NETPORT;
+static uint32_t ring_size = DEFAULT_RING_SIZE;
+static const char *lcore_arg;
+static const char *file_prefix;
+static const char *bind_arg;		/* -b argument, resolved after option parsing */
+static const char *debug_file;		/* --debug-file argument */
+static bool ipv4_only;			/* -4: restrict to IPv4 */
+static unsigned int debug_log;		/* -D count: raise RPCAPD log verbosity */
+
+static struct sockaddr_storage listen_addr;
+static socklen_t               listen_addrlen;
+
+static void stop_capture(struct session *s);
+
+static void
+set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port)
+{
+	if (ss->ss_family == AF_INET6)
+		((struct sockaddr_in6 *)ss)->sin6_port = htons(port);
+	else
+		((struct sockaddr_in *)ss)->sin_port = htons(port);
+}
+
+static uint16_t
+get_sockaddr_port(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET6)
+		return ntohs(((const struct sockaddr_in6 *)ss)->sin6_port);
+	return ntohs(((const struct sockaddr_in *)ss)->sin_port);
+}
+
+static bool
+is_loopback(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET) {
+		const struct sockaddr_in *sin = (const void *)ss;
+
+		return (ntohl(sin->sin_addr.s_addr) >> 24) == 127;
+	}
+	if (ss->ss_family == AF_INET6) {
+		const struct sockaddr_in6 *sin6 = (const void *)ss;
+
+		return IN6_IS_ADDR_LOOPBACK(&sin6->sin6_addr);
+	}
+	return false;
+}
+
+static void
+parse_bind_addr(const char *str, int family)
+{
+	struct addrinfo hints = {
+		.ai_family   = family,
+		.ai_socktype = SOCK_STREAM,
+		.ai_flags    = AI_NUMERICHOST | AI_PASSIVE,
+	};
+	struct addrinfo *res;
+	int rc;
+
+	rc = getaddrinfo(str, NULL, &hints, &res);
+	if (rc != 0)
+		rte_exit(EXIT_FAILURE, "Invalid bind address '%s': %s\n",
+			 str, gai_strerror(rc));
+	memcpy(&listen_addr, res->ai_addr, res->ai_addrlen);
+	listen_addrlen = res->ai_addrlen;
+	freeaddrinfo(res);
+}
+
+static RTE_ATOMIC(bool) quit_signal;
+
+static void
+signal_handler(int sig __rte_unused)
+{
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+/*
+ * Wait for fd to become readable, in POLL_INTERVAL_MS slices so that a
+ * quit signal (from SIGINT/SIGTERM or from the primary process dying)
+ * is noticed while blocked.  timeout_ms < 0 waits indefinitely.
+ *
+ * Returns 1 when readable, 0 on timeout, -1 on error or quit.
+ */
+static int
+wait_readable(int fd, int timeout_ms)
+{
+	struct pollfd pfd = { .fd = fd, .events = POLLIN };
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int wait_ms = POLL_INTERVAL_MS;
+		int rc;
+
+		if (timeout_ms >= 0) {
+			if (timeout_ms == 0)
+				return 0;
+			if (timeout_ms < wait_ms)
+				wait_ms = timeout_ms;
+			timeout_ms -= wait_ms;
+		}
+
+		rc = poll(&pfd, 1, wait_ms);
+		if (rc < 0) {
+			if (errno == EINTR)
+				continue;
+			RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+			return -1;
+		}
+		if (rc > 0)
+			return 1;
+	}
+	return -1;
+}
+
+/* accept() with a timeout, so a stalled client cannot wedge the daemon. */
+static int
+accept_timeout(int listen_fd, int timeout_ms)
+{
+	int fd;
+
+	switch (wait_readable(listen_fd, timeout_ms)) {
+	case 1:
+		break;
+	case 0:
+		RPCAPD_LOG(ERR, "timed out waiting for data connection");
+		return -1;
+	default:
+		return -1;
+	}
+
+	fd = accept(listen_fd, NULL, NULL);
+	if (fd < 0)
+		RPCAPD_LOG(ERR, "accept: %s", strerror(errno));
+	return fd;
+}
+
+/* Read exactly len bytes; return 0 on success, -1 on error or EOF. */
+static int
+recv_full(int fd, void *buf, size_t len)
+{
+	uint8_t *p = buf;
+
+	while (len > 0) {
+		ssize_t n;
+
+		/* Wait with a timeout rather than blocking in recv(), so a
+		 * quit signal or a dead primary is acted on promptly.
+		 */
+		if (wait_readable(fd, -1) != 1)
+			return -1;
+
+		n = recv(fd, p, len, 0);
+		if (n < 0 && errno == EINTR)
+			continue;
+
+		if (n <= 0)
+			return -1;
+
+		p += n;
+		len -= n;
+	}
+	return 0;
+}
+
+/*
+ * Send all of iov, resending the remainder if sendmsg() reports a short
+ * count (possible when the connection breaks or a signal arrives after
+ * some bytes were copied).  Consumes iov, so pass a scratch copy.
+ */
+static int
+send_iov_full(int fd, struct iovec *iov, int iovcnt, int flags)
+{
+	struct msghdr msg = {
+		.msg_iov    = iov,
+		.msg_iovlen = iovcnt,
+	};
+
+	while (msg.msg_iovlen > 0) {
+		ssize_t n = sendmsg(fd, &msg, flags | MSG_NOSIGNAL);
+
+		if (n < 0) {
+			if (errno == EINTR)
+				continue;
+			return -1;
+		}
+		if (n == 0)
+			return -1;
+
+		/* Drop whole iovecs that were fully sent, then trim the
+		 * partially sent one.
+		 */
+		while (msg.msg_iovlen > 0 && (size_t)n >= msg.msg_iov->iov_len) {
+			n -= msg.msg_iov->iov_len;
+			msg.msg_iov++;
+			msg.msg_iovlen--;
+		}
+		if (n > 0) {
+			msg.msg_iov->iov_base = (char *)msg.msg_iov->iov_base + n;
+			msg.msg_iov->iov_len -= n;
+		}
+	}
+	return 0;
+}
+
+static int
+rpcap_send_msg(int fd, uint8_t type, uint16_t value, const void *payload, uint32_t plen)
+{
+	struct rpcap_header hdr = {
+		.ver = RPCAP_VERSION,
+		.type = type,
+		.value = htons(value),
+		.plen = htonl(plen),
+	};
+	struct iovec iov[2] = {
+		{ .iov_base = &hdr,                        .iov_len = sizeof(hdr) },
+		{ .iov_base = (void *)(uintptr_t)payload,  .iov_len = plen },
+	};
+
+	return send_iov_full(fd, iov, plen > 0 ? 2 : 1, 0);
+}
+
+static int
+rpcap_send_error(int fd, uint16_t errcode, const char *msg)
+{
+	RPCAPD_LOG(WARNING, "sending error to client: %s", msg);
+	return rpcap_send_msg(fd, RPCAP_MSG_ERROR, errcode, msg, strlen(msg));
+}
+
+static int
+rpcap_recv_header(int fd, struct rpcap_header *hdr)
+{
+	if (recv_full(fd, hdr, sizeof(*hdr)) < 0)
+		return -1;
+	hdr->value = ntohs(hdr->value);
+	hdr->plen = ntohl(hdr->plen);
+	return 0;
+}
+
+/* Throw away plen bytes of payload we don't care about. */
+static int
+rpcap_discard(int fd, uint32_t plen)
+{
+	uint8_t buf[256];
+
+	while (plen > 0) {
+		size_t chunk = plen > sizeof(buf) ? sizeof(buf) : plen;
+
+		if (recv_full(fd, buf, chunk) < 0)
+			return -1;
+		plen -= chunk;
+	}
+	return 0;
+}
+
+/* Build and send the list of available DPDK ports. */
+static int
+handle_findallif(int fd)
+{
+	uint8_t *buf = NULL;
+	size_t buflen = 0;
+	uint16_t nif = 0;
+	uint16_t p;
+	int rc;
+
+	RTE_ETH_FOREACH_DEV(p) {
+		static const char desc[] = "DPDK port";
+		char name[RTE_ETH_NAME_MAX_LEN];
+		size_t namelen, desclen, entry;
+		uint8_t *nb;
+
+		if (rte_eth_dev_get_name_by_port(p, name) < 0) {
+			RPCAPD_LOG(DEBUG, "can not find name for port %u", p);
+			continue;
+		}
+
+		RPCAPD_LOG(DEBUG, "findallif: port %u -> '%s'", p, name);
+		namelen = strlen(name);
+		desclen = strlen(desc);
+		entry = sizeof(struct rpcap_findalldevs_if) + namelen + desclen;
+
+		nb = realloc(buf, buflen + entry);
+		if (nb == NULL) {
+			RPCAPD_LOG(ERR, "out of memory in findallif");
+			free(buf);
+			return rpcap_send_error(fd, 0, "out of memory");
+		}
+		buf = nb;
+
+		struct rpcap_findalldevs_if iface = {
+			.namelen = htons(namelen),
+			.desclen = htons(desclen),
+			.flags = htonl(PCAP_IF_UP | PCAP_IF_RUNNING),
+		};
+		memcpy(buf + buflen, &iface, sizeof(iface));
+		memcpy(buf + buflen + sizeof(iface), name, namelen);
+		memcpy(buf + buflen + sizeof(iface) + namelen, desc, desclen);
+		buflen += entry;
+		nif++;
+	}
+
+	RPCAPD_LOG(DEBUG, "findallif: %u interface(s)", nif);
+	rc = rpcap_send_msg(fd, RPCAP_MSG_FINDALLIF_REPLY, nif, buf, buflen);
+	free(buf);
+	return rc;
+}
+
+/* OPEN_REQ: payload is the interface name (no NUL). */
+static int
+handle_open(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_openreply reply = {
+		.linktype = htonl(DLT_EN10MB),
+	};
+	uint16_t port;
+
+	/* Unconditionally, not just when capture_on: a failed UPDATEFILTER
+	 * leaves the ring, mempool and data connection live with the capture
+	 * already disabled, and those must not survive into a new session.
+	 */
+	stop_capture(s);
+
+	if (plen >= sizeof(s->name)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "interface name too long");
+	}
+	if (recv_full(fd, s->name, plen) < 0)
+		return -1;
+	s->name[plen] = '\0';
+
+	if (rte_eth_dev_get_port_by_name(s->name, &port) < 0) {
+		RPCAPD_LOG(WARNING, "open: no such port '%s'", s->name);
+		/* s->name has already been overwritten; make sure a later
+		 * STARTCAP cannot capture the previously opened port.
+		 */
+		s->opened = false;
+		return rpcap_send_error(fd, 0, "unknown interface");
+	}
+	s->port = port;
+	s->opened = true;
+
+	RPCAPD_LOG(DEBUG, "open: '%s' -> dpdk port %u", s->name, port);
+	return rpcap_send_msg(fd, RPCAP_MSG_OPEN_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Open an ephemeral TCP listening socket; return fd, set *port_out. */
+static int
+open_data_listener(uint16_t *port_out)
+{
+	struct sockaddr_storage addr = listen_addr;
+	socklen_t alen;
+	int fd;
+
+	set_sockaddr_port(&addr, 0);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0) {
+		RPCAPD_LOG(ERR, "data socket: %s", strerror(errno));
+		return -1;
+	}
+
+	alen = listen_addrlen;
+	if (bind(fd, (struct sockaddr *)&addr, alen) < 0 ||
+	    listen(fd, 1) < 0 ||
+	    getsockname(fd, (struct sockaddr *)&addr, &alen) < 0) {
+		RPCAPD_LOG(ERR, "data port bind/listen: %s", strerror(errno));
+		close(fd);
+		return -1;
+	}
+	*port_out = get_sockaddr_port(&addr);
+	return fd;
+}
+
+static struct rte_ring *
+create_capture_ring(uint16_t port)
+{
+	char name[RTE_RING_NAMESIZE];
+
+	snprintf(name, sizeof(name), "rpcapd_r_%u_%d", port, getpid());
+	return rte_ring_create(name, ring_size, rte_socket_id(), 0);
+}
+
+static struct rte_mempool *
+create_capture_mempool(uint16_t port, uint32_t snaplen)
+{
+	char name[RTE_MEMPOOL_NAMESIZE];
+	uint32_t mbuf_size = RTE_PKTMBUF_HEADROOM + snaplen;
+
+	snprintf(name, sizeof(name), "rpcapd_p_%u_%d", port, getpid());
+	return rte_pktmbuf_pool_create(name, ring_size * 2, MBUF_CACHE_SIZE, 0,
+				       mbuf_size, rte_socket_id());
+}
+
+/*
+ * Read the optional capture filter that follows a start-capture request,
+ * and convert it for pdump. Client passes cBPF.
+ */
+static int
+read_filter(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_filterbpf_insn winsn;
+	struct rpcap_filter filter;
+	struct bpf_program bf;
+	struct bpf_insn *insns;
+	uint32_t i, nitems;
+
+	if (plen == 0)
+		return 0;		/* no filter: capture everything */
+
+	if (plen < sizeof(filter)) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "short filter header") < 0 ? -1 : 1;
+	}
+
+	if (recv_full(fd, &filter, sizeof(filter)) < 0)
+		return -1;
+	plen -= sizeof(filter);
+
+	if (ntohs(filter.filtertype) != RPCAP_UPDATEFILTER_BPF) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "unsupported filter type") < 0 ? -1 : 1;
+	}
+
+	/* nitems is client-supplied; bound it before trusting the length. */
+	nitems = ntohl(filter.nitems);
+	if (nitems == 0)
+		return rpcap_discard(fd, plen) < 0 ? -1 : 0;
+
+	if (nitems > MAX_FILTER_INSNS || plen < nitems * sizeof(winsn)) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "bad filter length") < 0 ? -1 : 1;
+	}
+
+	insns = calloc(nitems, sizeof(*insns));
+	if (insns == NULL) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "out of memory") < 0 ? -1 : 1;
+	}
+
+	for (i = 0; i < nitems; i++) {
+		if (recv_full(fd, &winsn, sizeof(winsn)) < 0) {
+			free(insns);
+			return -1;
+		}
+		insns[i].code = ntohs(winsn.code);
+		insns[i].jt   = winsn.jt;
+		insns[i].jf   = winsn.jf;
+		insns[i].k    = ntohl(winsn.k);
+	}
+	plen -= nitems * sizeof(winsn);
+
+	/* Anything after the instructions is padding we do not need. */
+	if (rpcap_discard(fd, plen) < 0) {
+		free(insns);
+		return -1;
+	}
+
+	bf.bf_len = nitems;
+	bf.bf_insns = insns;
+
+	/* Reject a malformed program here */
+	if (!bpf_validate(bf.bf_insns, bf.bf_len)) {
+		free(insns);
+		return rpcap_send_error(fd, 0, "invalid filter program") < 0 ? -1 : 1;
+	}
+
+	/* A filter recorded by an earlier UPDATEFILTER may still be here;
+	 * it is about to be replaced, so do not leak it.
+	 */
+	rte_free(s->prm);
+	s->prm = rte_bpf_convert(&bf);
+	free(insns);
+	if (s->prm == NULL) {
+		RPCAPD_LOG(ERR, "rte_bpf_convert failed: %s",
+			rte_strerror(rte_errno));
+		return rpcap_send_error(fd, 0, "cannot convert filter") < 0 ? -1 : 1;
+	}
+
+	RPCAPD_LOG(DEBUG, "capture filter: %u instructions", nitems);
+	return 0;
+}
+
+/* Tear down anything that handle_startcap brought up.  Safe to call
+ * after partial setup as well as after a successful capture.
+ */
+static void
+stop_capture(struct session *s)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int n;
+
+	if (s->capture_on) {
+		rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+		RPCAPD_LOG(INFO, "capture stopped on %s (%u packets)",
+			s->name, s->npkt);
+	}
+	s->capture_on = false;
+
+	if (s->promisc_set) {
+		rte_eth_promiscuous_disable(s->port);
+		s->promisc_set = false;
+	}
+
+	if (s->ring != NULL) {
+		while ((n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts,
+						      BURST_SIZE, NULL)) > 0)
+			rte_pktmbuf_free_bulk(pkts, n);
+		rte_ring_free(s->ring);
+		s->ring = NULL;
+	}
+	if (s->mp != NULL) {
+		rte_mempool_free(s->mp);
+		s->mp = NULL;
+	}
+
+	/* Only safe once pdump is disabled */
+	rte_free(s->prm);
+	s->prm = NULL;
+	if (s->data_fd >= 0) {
+		close(s->data_fd);
+		s->data_fd = -1;
+	}
+}
+
+/*
+ * STARTCAP_REQ: open the data connection and arm the pdump callback.
+ * We use passive mode with the server-allocated data port:
+ *   - the server picks an ephemeral port and listens on it
+ *   - the server returns that port in startcapreply.portdata
+ *   - the client connects back to that port for the packet stream
+ */
+static int
+handle_startcap(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_startcapreq req;
+	uint16_t data_port;
+	uint16_t flags;
+	struct rte_bpf_prm *recorded;
+	int data_listen;
+	int data_fd;
+	int ret;
+
+	recorded = s->prm;
+	s->prm = NULL;
+	stop_capture(s);
+	s->prm = recorded;
+
+	if (!s->opened) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "no interface open");
+	}
+
+	if (plen < sizeof(req)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "short startcap request");
+	}
+	if (recv_full(fd, &req, sizeof(req)) < 0)
+		return -1;
+
+	flags = ntohs(req.flags);
+	if (flags & RPCAP_STARTCAPREQ_FLAG_DGRAM) {
+		rpcap_discard(fd, plen - sizeof(req));
+		return rpcap_send_error(fd, 0, "UDP data transfer not supported");
+	}
+
+	ret = read_filter(fd, plen - sizeof(req), s);
+	if (ret != 0)
+		return ret < 0 ? -1 : 0;	/* error already reported to client */
+
+	/* Direction flags map onto pdump's RX/TX selection; neither (or both)
+	 * means capture in both directions.
+	 */
+	s->pdump_flags = RTE_PDUMP_FLAG_RXTX;
+	if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+		      RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+	    RPCAP_STARTCAPREQ_FLAG_INBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_RX;
+	else if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+			   RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+		 RPCAP_STARTCAPREQ_FLAG_OUTBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_TX;
+
+	s->snaplen = ntohl(req.snaplen);
+	if (s->snaplen == 0 || s->snaplen > DEFAULT_SNAPLEN)
+		s->snaplen = DEFAULT_SNAPLEN;
+
+	s->ring = create_capture_ring(s->port);
+	s->mp = create_capture_mempool(s->port, s->snaplen);
+	if (s->ring == NULL || s->mp == NULL) {
+		RPCAPD_LOG(ERR, "ring/mempool alloc failed: %s",
+			rte_strerror(rte_errno));
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "DPDK alloc failed");
+	}
+
+	data_listen = open_data_listener(&data_port);
+	if (data_listen < 0) {
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "data port setup failed");
+	}
+
+	/* Leave the port alone if it is already promiscuous: it belongs to
+	 * the primary process, and stop_capture() must not turn off
+	 * something this daemon did not turn on.
+	 */
+	if ((flags & RPCAP_STARTCAPREQ_FLAG_PROMISC) &&
+	    rte_eth_promiscuous_get(s->port) != 1) {
+		if (rte_eth_promiscuous_enable(s->port) == 0)
+			s->promisc_set = true;
+		else
+			RPCAPD_LOG(NOTICE, "cannot enable promiscuous mode on %s",
+				s->name);
+	}
+
+	/* Arm pdump before replying. */
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		close(data_listen);
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "cannot enable capture");
+	}
+	s->capture_on = true;
+	s->npkt = 0;
+
+	struct rpcap_startcapreply reply = {
+		.bufsize = htonl(s->snaplen * BURST_SIZE),
+		.portdata = htons(data_port),
+	};
+	if (rpcap_send_msg(fd, RPCAP_MSG_STARTCAP_REPLY, 0, &reply, sizeof(reply)) < 0) {
+		close(data_listen);
+		stop_capture(s);
+		return -1;
+	}
+
+	RPCAPD_LOG(DEBUG, "awaiting connection");
+
+	data_fd = accept_timeout(data_listen, DATA_ACCEPT_TIMEOUT_MS);
+	close(data_listen);
+	if (data_fd < 0) {
+		stop_capture(s);
+		return -1;
+	}
+
+	s->data_fd = data_fd;
+
+	RPCAPD_LOG(INFO,
+		"capture started on %s (snaplen %u, data port %u)",
+		s->name, s->snaplen, data_port);
+	return 0;
+}
+
+/*
+ * UPDATEFILTER_REQ: replace the capture filter.
+ *
+ * pdump takes its filter when the callback is armed and offers no way
+ * to replace it, so this disables and re-enables the callback with the
+ * new program.  Packets already in the ring are kept; only the brief
+ * gap between disable and enable is lost.  Refusing the request is not
+ * an option: libpcap sends UPDATEFILTER right after STARTCAP when the
+ * client was opened with PCAP_OPENFLAG_NOCAPTURE_RPCAP and aborts the
+ * capture if it fails, and Wireshark sets that flag by default.
+ *
+ * Before the capture starts this just records the filter for the
+ * eventual STARTCAP.
+ */
+static int
+handle_updatefilter(int fd, uint32_t plen, struct session *s)
+{
+	struct rte_bpf_prm *old = s->prm;
+	int ret;
+
+	s->prm = NULL;
+	ret = read_filter(fd, plen, s);
+	if (ret != 0) {
+		/* Malformed request: keep running with the old filter. */
+		rte_free(s->prm);
+		s->prm = old;
+		return ret < 0 ? -1 : 0;	/* error already reported */
+	}
+
+	if (!s->capture_on) {
+		rte_free(old);
+		return rpcap_send_msg(fd, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+	}
+
+	rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+	s->capture_on = false;
+
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		rte_free(old);
+		/* The capture cannot be resumed, so do not leave the ring,
+		 * mempool and data connection behind: the client has been
+		 * told the capture is over, and a session that is neither
+		 * capturing nor torn down has no way back.
+		 */
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "cannot apply filter");
+	}
+	s->capture_on = true;
+
+	/* Safe now that the old program is no longer referenced. */
+	rte_free(old);
+
+	RPCAPD_LOG(DEBUG, "capture filter updated on %s", s->name);
+	return rpcap_send_msg(fd, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+}
+
+/*
+ * Pull a burst from the ring, frame each packet into an RPCAP_MSG_PACKET
+ * message, and send it on the data connection.  MSG_MORE corks the
+ * socket until the ring drains, so a backlog coalesces into full
+ * segments instead of flushing every BURST_SIZE packets.
+ */
+static ssize_t
+process_ring(struct session *s, unsigned int *avail)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int i, n;
+	ssize_t written = 0;
+	struct timeval tv;
+
+	n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts, BURST_SIZE, avail);
+	if (n == 0)
+		return 0;
+
+	/* One timestamp for the whole burst */
+	gettimeofday(&tv, NULL);
+
+	for (i = 0; i < n; i++) {
+		struct rte_mbuf *m = pkts[i];
+		/* Sized from the same bound that clamps caplen below, so the
+		 * two cannot drift apart.
+		 */
+		uint8_t buf[DEFAULT_SNAPLEN];
+		uint32_t pktlen = rte_pktmbuf_pkt_len(m);
+		uint32_t caplen = pktlen < s->snaplen ? pktlen : s->snaplen;
+		const void *data;
+
+		s->npkt++;
+
+		struct rpcap_header hdr = {
+			.ver = RPCAP_VERSION,
+			.type = RPCAP_MSG_PACKET,
+			.plen = htonl(sizeof(struct rpcap_pkthdr) + caplen),
+		};
+
+		/*
+		 * pdump copies at most the snaplen into the capture mempool
+		 * and rte_pktmbuf_copy() counts only what it copied, so
+		 * pktlen is already clamped: a truncated packet is reported
+		 * with len == caplen.  The original wire length does not
+		 * reach this process.  See the Limitations section of
+		 * doc/guides/sample_app_ug/rpcapd.rst.
+		 */
+		struct rpcap_pkthdr pkthdr = {
+			.timestamp_sec = htonl((uint32_t)tv.tv_sec),
+			.timestamp_usec = htonl((uint32_t)tv.tv_usec),
+			.caplen = htonl(caplen),
+			.len = htonl(pktlen),
+			.npkt = htonl(s->npkt),
+		};
+
+		data = rte_pktmbuf_read(m, 0, caplen, buf);
+
+		struct iovec iov[3] = {
+			{ .iov_base = &hdr,                       .iov_len = sizeof(hdr) },
+			{ .iov_base = &pkthdr,                    .iov_len = sizeof(pkthdr) },
+			{ .iov_base = (void *)(uintptr_t)data,    .iov_len = caplen },
+		};
+
+		/* more to come in this burst, or still queued in the ring */
+		bool more = (i + 1 < n) || (*avail > 0);
+
+		if (send_iov_full(s->data_fd, iov, 3, more ? MSG_MORE : 0) < 0) {
+			if (errno == EPIPE || errno == ECONNRESET)
+				RPCAPD_LOG(DEBUG, "data connection closed by client");
+			else
+				RPCAPD_LOG(NOTICE, "send on data connection failed: %s",
+					   strerror(errno));
+			goto error;
+		}
+		rte_pktmbuf_free(m);
+		written += sizeof(hdr) + sizeof(pkthdr) + caplen;
+	}
+
+	return written;
+
+error:
+	rte_pktmbuf_free_bulk(pkts + i, n - i);
+	return -1;
+}
+
+/* Poll the control socket while idle.
+ * Returns 0 to keep capturing, 1 if a control message (typically
+ * ENDCAP) is pending, or -1 if the client has gone away.
+ */
+static int
+check_socket_status(int ctrl_fd)
+{
+	struct pollfd pfd = { .fd = ctrl_fd, .events = POLLIN };
+
+	if (poll(&pfd, 1, 0) < 0) {
+		if (errno == EINTR)
+			return 0;
+		RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+		return -1;
+	}
+	if (pfd.revents & (POLLERR | POLLHUP | POLLNVAL)) {
+		RPCAPD_LOG(DEBUG, "client closed control connection");
+		return -1;
+	}
+	if (pfd.revents & POLLIN)
+		return 1;
+	return 0;
+}
+
+/*
+ * Stay in the capture loop until either:
+ *   - a control message arrives (typically ENDCAP),
+ *   - the data connection breaks, or
+ *   - a quit signal is delivered.
+ *
+ * Returns 0 if the session should continue (the caller reads the
+ * pending control message), -1 if the client is gone.
+ *
+ * The control socket is polled once per iteration, not just when the
+ * ring runs dry.  A client that sends a request mid-capture blocks
+ * waiting for the reply without draining the data socket, so under
+ * sustained traffic a poll that only happens while idle never runs and
+ * both ends wedge once the socket buffers fill.
+ */
+static int
+capture_loop(int ctrl_fd, struct session *s)
+{
+	unsigned int empty_count = 0;
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		ssize_t written;
+		unsigned int avail = 0;
+
+		switch (check_socket_status(ctrl_fd)) {
+		case 1:
+			/* control message pending, let caller service it */
+			return 0;
+		case 0:
+			break;
+		default:
+			/* client is gone */
+			return -1;
+		}
+
+		written = process_ring(s, &avail);
+		if (written < 0) {
+			/* process_ring has already logged the reason */
+			return -1;
+		}
+
+		if (written > 0) {
+			/* are there more packets? */
+			empty_count = (avail == 0);
+			continue;
+		}
+
+		if (empty_count < SLEEP_THRESHOLD) {
+			/* spin a few times before checking */
+			++empty_count;
+			rte_pause();
+			continue;
+		}
+
+		/* ring has been empty for a while: stop spinning */
+		rte_delay_us_sleep(SLEEP_US);
+	}
+	return 0;
+}
+
+static int
+handle_endcap(int fd, uint32_t plen, struct session *s)
+{
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+	stop_capture(s);
+	return rpcap_send_msg(fd, RPCAP_MSG_ENDCAP_REPLY, 0, NULL, 0);
+}
+
+static int
+handle_stats(int fd, uint32_t plen, const struct session *s)
+{
+	struct rte_eth_stats es = { 0 };
+
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+
+	if (s->capture_on)
+		rte_eth_stats_get(s->port, &es);
+
+	struct rpcap_stats reply = {
+		.ifrecv   = htonl((uint32_t)es.ipackets),
+		.ifdrop   = htonl((uint32_t)es.ierrors),
+		.krnldrop = 0,
+		.svrcapt  = htonl(s->npkt),
+	};
+	return rpcap_send_msg(fd, RPCAP_MSG_STATS_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Service a single client until it disconnects. */
+static void
+handle_client(int ctrl_fd)
+{
+	struct sockaddr_storage peer;
+	socklen_t plen = sizeof(peer);
+	char host[NI_MAXHOST] = "?";
+	struct session s = { .data_fd = -1 };
+
+	if (getpeername(ctrl_fd, (struct sockaddr *)&peer, &plen) == 0)
+		getnameinfo((struct sockaddr *)&peer, plen,
+			    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	RPCAPD_LOG(INFO, "client %s connected", host);
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		struct rpcap_header hdr;
+
+		/* Drain the ring whenever a capture is running */
+		if (s.capture_on && capture_loop(ctrl_fd, &s) < 0)
+			goto done;
+
+		if (rpcap_recv_header(ctrl_fd, &hdr) < 0)
+			break;
+
+		/* Only version 0 is spoken here */
+		if (hdr.ver != RPCAP_VERSION) {
+			RPCAPD_LOG(WARNING, "unsupported protocol version %u",
+				hdr.ver);
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_error(ctrl_fd, PCAP_ERR_WRONGVER,
+					     "unsupported protocol version") < 0)
+				goto done;
+			continue;
+		}
+
+		switch (hdr.type) {
+		case RPCAP_MSG_AUTH_REQ:
+			/* No auth: discard credentials, ack with empty reply.
+			 * libpcap treats a zero-length AUTH_REPLY as "version
+			 * 0 only, same byte order".
+			 */
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_msg(ctrl_fd, RPCAP_MSG_AUTH_REPLY, 0, NULL, 0) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_FINDALLIF_REQ:
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 || handle_findallif(ctrl_fd) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_OPEN_REQ:
+			if (handle_open(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STARTCAP_REQ:
+			if (handle_startcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_UPDATEFILTER_REQ:
+			if (handle_updatefilter(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_ENDCAP_REQ:
+			if (handle_endcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STATS_REQ:
+			if (handle_stats(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_CLOSE:
+			rpcap_discard(ctrl_fd, hdr.plen);
+			goto done;
+		default:
+			RPCAPD_LOG(WARNING, "unsupported request type 0x%02x", hdr.type);
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_error(ctrl_fd, 0, "unsupported request") < 0)
+				goto done;
+			break;
+		}
+	}
+done:
+	stop_capture(&s);
+	close(ctrl_fd);
+	RPCAPD_LOG(INFO, "client %s disconnected", host);
+}
+
+static int
+open_listen_socket(uint16_t port)
+{
+	struct sockaddr_storage addr = listen_addr;
+	char host[NI_MAXHOST];
+	int fd, one = 1;
+
+	set_sockaddr_port(&addr, port);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0)
+		rte_exit(EXIT_FAILURE, "socket: %s\n", strerror(errno));
+	setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one));
+
+	if (bind(fd, (struct sockaddr *)&addr, listen_addrlen) < 0)
+		rte_exit(EXIT_FAILURE, "bind(%u): %s\n", port, strerror(errno));
+
+	int err = getnameinfo((struct sockaddr *)&listen_addr, listen_addrlen,
+			      host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	if (err != 0)
+		rte_exit(EXIT_FAILURE, "Listen address lookup failed: %s\n",
+			 gai_strerror(err));
+
+	RPCAPD_LOG(NOTICE, "listening on %s port %u", host, listen_port);
+
+	if (!is_loopback(&listen_addr))
+		RPCAPD_LOG(WARNING,
+			"non-loopback address %s; "
+			"rpcap is unauthenticated and unencrypted, captured traffic is exposed to the network",
+			host);
+
+	if (listen(fd, 1) < 0)
+		rte_exit(EXIT_FAILURE, "listen: %s\n", strerror(errno));
+
+	return fd;
+}
+
+static void
+usage(FILE *f, const char *progname)
+{
+	fprintf(f, "Usage: %s [options]\n", progname);
+	fprintf(f,
+		"  -p, --port <port>     listen port (default %u)\n"
+		"  -b, --bind <addr>     bind address (default 127.0.0.1)\n"
+		"  -4                    use only IPv4 (reject IPv6 bind addresses)\n"
+		"  -N <ring size>        ring size in packets (default %u)\n"
+		"  -D, --debug           increase log verbosity (-D info, -DD debug)\n"
+		"      --debug-file <f>  redirect log output to file <f> (append mode)\n"
+		"      --version         print version and exit\n"
+		"  -h, --help            print this help and exit\n"
+		"      --lcore=<core>    CPU core to run on (default: any)\n"
+		"      --file-prefix=<p> prefix to use for multi-process\n"
+		"\n"
+		"WARNING: rpcap is unauthenticated and unencrypted.  Binding to\n"
+		"any non-loopback address exposes captured traffic to the\n"
+		"network.  Sample application; not for production use.\n",
+		RPCAP_DEFAULT_NETPORT, DEFAULT_RING_SIZE);
+}
+
+static void
+print_version(void)
+{
+	printf("rpcapd, a remote packet capture daemon (DPDK pdump backend)\n"
+	       "Built against %s\n", rte_version());
+}
+
+static void
+parse_opts(int argc, char **argv)
+{
+	enum {
+		OPT_LONG_ONLY = 0x100,
+		OPT_DEBUG_FILE,
+		OPT_VERSION,
+	};
+	static const struct option long_options[] = {
+		{ "port",        required_argument, NULL, 'p' },
+		{ "bind",        required_argument, NULL, 'b' },
+		{ "debug",       no_argument,       NULL, 'D' },
+		{ "help",        no_argument,       NULL, 'h' },
+		{ "version",     no_argument,       NULL, OPT_VERSION },
+		{ "debug-file",  required_argument, NULL, OPT_DEBUG_FILE },
+		{ "file-prefix", required_argument, NULL, 0 },
+		{ "lcore",       required_argument, NULL, 0 },
+		{ NULL, 0, NULL, 0 },
+	};
+	int option_index, c;
+
+	while ((c = getopt_long(argc, argv, "hD4p:b:N:",
+				long_options, &option_index)) != -1) {
+		switch (c) {
+		case 'p': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			if (u == 0 || u > UINT16_MAX)
+				rte_exit(EXIT_FAILURE, "Invalid port: %s\n", optarg);
+			listen_port = (uint16_t)u;
+			break;
+		}
+		case 'b':
+			bind_arg = optarg;
+			break;
+		case '4':
+			ipv4_only = true;
+			break;
+		case 'N': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			/* Check the full value before narrowing it: an upper
+			 * bound is needed anyway because rte_align32pow2()
+			 * wraps to zero above 2^31, and that failure would
+			 * otherwise only surface in rte_ring_create() on the
+			 * first capture.
+			 */
+			if (u < 64 || u > MAX_RING_SIZE)
+				rte_exit(EXIT_FAILURE,
+					 "Ring size must be between 64 and %u\n",
+					 MAX_RING_SIZE);
+			ring_size = (uint32_t)u;
+			/* rte_ring_create() requires a power of two. */
+			if (!rte_is_power_of_2(ring_size)) {
+				ring_size = rte_align32pow2(ring_size);
+				RPCAPD_LOG(NOTICE, "ring size rounded up to %u",
+					ring_size);
+			}
+			break;
+		}
+		case 'D':
+			debug_log++;
+			break;
+		case 'h':
+			usage(stdout, argv[0]);
+			exit(0);
+		case OPT_VERSION:
+			print_version();
+			exit(0);
+		case OPT_DEBUG_FILE:
+			debug_file = optarg;
+			break;
+		case 0: {
+			const char *longopt = long_options[option_index].name;
+
+			if (!strcmp(longopt, "lcore")) {
+				lcore_arg = optarg;
+				break;
+			} else if (!strcmp(longopt, "file-prefix")) {
+				file_prefix = optarg;
+				break;
+			}
+		}
+			/* fallthrough */
+		default:
+			usage(stderr, argv[0]);
+			exit(EXIT_FAILURE);
+		}
+	}
+
+	/* Resolve the bind address now that -4 has been seen. */
+	parse_bind_addr(bind_arg ? bind_arg : "127.0.0.1",
+			ipv4_only ? AF_INET : AF_UNSPEC);
+}
+
+/*
+ * Periodic check that the DPDK primary process is still alive.
+ * If it dies our shared-memory state (rings, mempools, pdump) becomes
+ * unsafe to touch, so we set quit_signal and let the main loop tear
+ * down cleanly on its next iteration.  The callback runs on the EAL
+ * interrupt thread; quit_signal is atomic so the read in the main
+ * loop is well-defined.
+ */
+static void
+monitor_primary(void *arg __rte_unused)
+{
+	if (rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed))
+		return;
+
+	if (rte_eal_primary_proc_alive(NULL)) {
+		rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL);
+		return;
+	}
+
+	RPCAPD_LOG(NOTICE, "primary process exited, shutting down");
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+static void
+enable_primary_monitor(void)
+{
+	if (rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL) < 0)
+		RPCAPD_LOG(WARNING, "failed to install primary process monitor");
+}
+
+static void
+disable_primary_monitor(void)
+{
+	rte_eal_alarm_cancel(monitor_primary, NULL);
+}
+
+/*
+ * Bring up EAL as a secondary process so that pdump can attach to a
+ * running primary DPDK application.  Mirrors dumpcap's approach: the
+ * RPCAP user sees a small set of options (port, ring size) rather
+ * than the full DPDK EAL command line.
+ */
+static int
+dpdk_init(void)
+{
+	static const char * const args[] = {
+		"rpcapd",
+		"--proc-type", "secondary",
+		"--log-level", "info",        /* EAL stays quiet */
+	};
+	int eal_argc = RTE_DIM(args);
+	rte_cpuset_t cpuset = { };
+	char **eal_argv;
+	unsigned int i;
+
+	if (file_prefix != NULL)
+		eal_argc += 2;
+
+	if (lcore_arg != NULL)
+		eal_argc += 2;
+
+	eal_argv = calloc(eal_argc + 1, sizeof(char *));
+	if (eal_argv == NULL)
+		return -1;
+
+	for (i = 0; i < RTE_DIM(args); i++) {
+		eal_argv[i] = strdup(args[i]);
+		if (eal_argv[i] == NULL)
+			return -1;
+	}
+
+	if (file_prefix != NULL && *file_prefix != '\0') {
+		eal_argv[i++] = strdup("--file-prefix");
+		eal_argv[i++] = strdup(file_prefix);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+
+	if (lcore_arg != NULL) {
+		eal_argv[i++] = strdup("--lcores");
+		eal_argv[i++] = strdup(lcore_arg);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+	eal_argc = i;
+
+	/*
+	 * Need to get the original cpuset, before EAL init changes
+	 * the affinity of this thread (main lcore).
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_get_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		rte_panic("rte_thread_getaffinity failed\n");
+
+	if (rte_eal_init(eal_argc, eal_argv) < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failed: is the primary process running?\n");
+
+	/*
+	 * If no lcore argument was specified,
+	 * then run this program as a normal process
+	 * which can be scheduled on any non-isolated CPU.
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_set_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		RPCAPD_LOG(INFO, "Can not restore original CPU affinity");
+
+	if (rte_pdump_init() < 0)
+		rte_exit(EXIT_FAILURE, "rte_pdump_init failed\n");
+
+	return 0;
+}
+
+int
+main(int argc, char **argv)
+{
+	struct sigaction action = {
+		.sa_handler = signal_handler,
+	};
+	int srv_fd;
+
+	parse_opts(argc, argv);
+
+	/*
+	 * Redirect log output before EAL init so EAL's own messages are
+	 * captured too.  The FILE handle is intentionally never closed:
+	 * the kernel reclaims it at process exit.
+	 */
+	if (debug_file != NULL) {
+		FILE *fp = fopen(debug_file, "a");
+
+		if (fp == NULL)
+			rte_exit(EXIT_FAILURE, "Cannot open debug file '%s': %s\n",
+				 debug_file, strerror(errno));
+		setvbuf(fp, NULL, _IOLBF, 0);
+		rte_openlog_stream(fp);
+	}
+
+	if (dpdk_init() < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failure\n");
+
+	/* Default to NOTICE: only things the operator needs to see.
+	 * Each -D steps down one level, to INFO then DEBUG.
+	 */
+	rte_log_set_level(RTE_LOGTYPE_RPCAPD,
+			  debug_log >= 2 ? RTE_LOG_DEBUG :
+			  debug_log == 1 ? RTE_LOG_INFO : RTE_LOG_NOTICE);
+
+	if (rte_eth_dev_count_avail() == 0)
+		rte_exit(EXIT_FAILURE, "No Ethernet ports found\n");
+
+	sigaction(SIGTERM, &action, NULL);
+	sigaction(SIGINT, &action, NULL);
+
+	/* If peer closes, this detected in next recv() */
+	signal(SIGPIPE, SIG_IGN);
+
+	srv_fd = open_listen_socket(listen_port);
+
+	enable_primary_monitor();
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int cfd = accept_timeout(srv_fd, -1);
+
+		if (cfd < 0) {
+			if (errno == EINTR)
+				continue;
+			break;
+		}
+		handle_client(cfd);
+	}
+
+	disable_primary_monitor();
+	RPCAPD_LOG(NOTICE, "shutting down");
+	close(srv_fd);
+	rte_pdump_uninit();
+	return rte_eal_cleanup() ? EXIT_FAILURE : 0;
+}
diff --git a/examples/rpcapd/meson.build b/examples/rpcapd/meson.build
new file mode 100644
index 0000000000..320b262666
--- /dev/null
+++ b/examples/rpcapd/meson.build
@@ -0,0 +1,19 @@
+# SPDX-License-Identifier: BSD-3-Clause
+# Copyright(c) 2026 Stephen Hemminger
+
+# since it relies on primary/secondary process
+# this example is Linux only
+if not is_linux
+    build = false
+    subdir_done()
+endif
+
+if not dpdk_conf.has('RTE_HAS_LIBPCAP')
+    build = false
+    reason = 'missing dependency, "libpcap"'
+    subdir_done()
+endif
+
+sources = files('main.c')
+ext_deps += pcap_dep
+deps += ['ethdev', 'pdump', 'bpf']
diff --git a/examples/rpcapd/rpcap-protocol.h b/examples/rpcapd/rpcap-protocol.h
new file mode 100644
index 0000000000..b381271de7
--- /dev/null
+++ b/examples/rpcapd/rpcap-protocol.h
@@ -0,0 +1,127 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * On-the-wire RPCAP protocol definitions, transcribed from libpcap's
+ * rpcap-protocol.h.  See:
+ *   https://github.com/the-tcpdump-group/libpcap/blob/master/rpcap-protocol.h
+ *
+ * Only the subset needed by the DPDK rpcapd example is included here.  All
+ * multi-byte fields in the structures below are big-endian on the wire.
+ */
+
+#ifndef _RPCAP_PROTOCOL_H_
+#define _RPCAP_PROTOCOL_H_
+
+#include <stdint.h>
+
+#define RPCAP_VERSION              0
+#define RPCAP_DEFAULT_NETPORT      2002
+
+/* Message types */
+#define RPCAP_MSG_ERROR            0x01
+#define RPCAP_MSG_FINDALLIF_REQ    0x02
+#define RPCAP_MSG_OPEN_REQ         0x03
+#define RPCAP_MSG_STARTCAP_REQ     0x04
+#define RPCAP_MSG_UPDATEFILTER_REQ 0x05
+#define RPCAP_MSG_CLOSE            0x06
+#define RPCAP_MSG_PACKET           0x07
+#define RPCAP_MSG_AUTH_REQ         0x08
+#define RPCAP_MSG_STATS_REQ        0x09
+#define RPCAP_MSG_ENDCAP_REQ       0x0a
+#define RPCAP_MSG_IS_REPLY         0x80
+
+#define RPCAP_MSG_FINDALLIF_REPLY    (RPCAP_MSG_FINDALLIF_REQ    | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_OPEN_REPLY         (RPCAP_MSG_OPEN_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STARTCAP_REPLY     (RPCAP_MSG_STARTCAP_REQ     | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_UPDATEFILTER_REPLY (RPCAP_MSG_UPDATEFILTER_REQ | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_AUTH_REPLY         (RPCAP_MSG_AUTH_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_ENDCAP_REPLY       (RPCAP_MSG_ENDCAP_REQ       | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STATS_REPLY	     (RPCAP_MSG_STATS_REQ	 | RPCAP_MSG_IS_REPLY)
+
+/* Error codes carried in the 'value' field of RPCAP_MSG_ERROR */
+#define PCAP_ERR_WRONGVER          17
+
+/* Filter encoding: the filter is a BPF/NPF program */
+#define RPCAP_UPDATEFILTER_BPF     1
+
+/* Flags in rpcap_startcapreq.flags */
+#define RPCAP_STARTCAPREQ_FLAG_PROMISC     0x00000001	/* promiscuous mode */
+#define RPCAP_STARTCAPREQ_FLAG_DGRAM       0x00000002	/* use UDP for data */
+#define RPCAP_STARTCAPREQ_FLAG_SERVEROPEN  0x00000004	/* server connects out */
+#define RPCAP_STARTCAPREQ_FLAG_INBOUND     0x00000008	/* capture inbound only */
+#define RPCAP_STARTCAPREQ_FLAG_OUTBOUND    0x00000010	/* capture outbound only */
+
+/* Subset of pcap interface flags (pcap.h) */
+#define PCAP_IF_UP                 0x00000002
+#define PCAP_IF_RUNNING            0x00000004
+
+/* DLT_EN10MB - ethernet, the only link type we report */
+#define DLT_EN10MB                 1
+
+struct rpcap_header {
+	uint8_t  ver;
+	uint8_t  type;
+	uint16_t value;
+	uint32_t plen;
+};
+
+struct rpcap_findalldevs_if {
+	uint16_t namelen;
+	uint16_t desclen;
+	uint32_t flags;
+	uint16_t naddr;
+	uint16_t dummy;
+};
+
+struct rpcap_openreply {
+	int32_t  linktype;
+	int32_t  tzoff;
+};
+
+struct rpcap_startcapreq {
+	uint32_t snaplen;
+	uint32_t read_timeout;
+	uint16_t flags;
+	uint16_t portdata;
+};
+
+struct rpcap_startcapreply {
+	int32_t  bufsize;
+	uint16_t portdata;
+	uint16_t dummy;
+};
+
+/*
+ * A filter, sent either after rpcap_startcapreq or in an
+ * RPCAP_MSG_UPDATEFILTER_REQ, followed by nitems instructions.
+ */
+struct rpcap_filter {
+	uint16_t filtertype;
+	uint16_t dummy;
+	uint32_t nitems;
+};
+
+/* One cBPF instruction, repeated nitems times after rpcap_filter. */
+struct rpcap_filterbpf_insn {
+	uint16_t code;
+	uint8_t  jt;
+	uint8_t  jf;
+	int32_t  k;
+};
+
+struct rpcap_stats {
+	uint32_t ifrecv;
+	uint32_t ifdrop;
+	uint32_t krnldrop;
+	uint32_t svrcapt;
+};
+
+struct rpcap_pkthdr {
+	uint32_t timestamp_sec;
+	uint32_t timestamp_usec;
+	uint32_t caplen;
+	uint32_t len;
+	uint32_t npkt;
+};
+
+#endif /* _RPCAP_PROTOCOL_H_ */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
@ 2026-09-21  9:51   ` Marat Khalili
  2026-09-21 15:57     ` Stephen Hemminger
                       ` (2 more replies)
  2026-09-22 18:45   ` Stephen Hemminger
  1 sibling, 3 replies; 21+ messages in thread
From: Marat Khalili @ 2026-09-21  9:51 UTC (permalink / raw)
  To: dev

Very cool example, demonstrating many advanced DPDK features at once. In 
this regard, it really excels.

It could be however useful as a debug tool, and in this case having to 
deal with a secondary process is a hassle. Could we have this feature in 
dpdk-testpmd, instead or as well?

Otherwise, having seen the demonstration, I do not have many comments, 
except for a couple nits below.

On 20/09/2026 19:59, Stephen Hemminger wrote:
> This example adds RPCAP support over localhost TCP
> integrated with DPDK. It uses a secondary process that allows
> connections from using tcpdump defacto protocol rpcap.
>
> See: doc/guides/sample_app_ug/rpcapd.rst for more info
>
> Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
> ---
> v2 - Build on Linux only
>     - add filter support
>     - fix lots of AI review feedback
>
>   MAINTAINERS                            |    2 +
>   doc/guides/rel_notes/release_26_11.rst |    4 +
>   doc/guides/sample_app_ug/index.rst     |    1 +
>   doc/guides/sample_app_ug/rpcapd.rst    |  216 ++++
>   examples/meson.build                   |    1 +
>   examples/rpcapd/main.c                 | 1430 ++++++++++++++++++++++++

main.c is a bit large and handles many different functions. It would be 
more readable if different aspects were handled by different files, or 
at least clearly marked sections of the file, with overall structure 
explained somewhere at the top.

>   examples/rpcapd/meson.build            |   19 +
>   examples/rpcapd/rpcap-protocol.h       |  127 +++
Could we just use one from libpcap?

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21  9:51   ` Marat Khalili
@ 2026-09-21 15:57     ` Stephen Hemminger
  2026-09-21 15:58     ` Stephen Hemminger
  2026-09-21 16:17     ` Stephen Hemminger
  2 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-21 15:57 UTC (permalink / raw)
  To: Marat Khalili; +Cc: dev

On Mon, 21 Sep 2026 10:51:39 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> Very cool example, demonstrating many advanced DPDK features at once. In 
> this regard, it really excels.
> 
> It could be however useful as a debug tool, and in this case having to 
> deal with a secondary process is a hassle. Could we have this feature in 
> dpdk-testpmd, instead or as well?

No. There is dpdk-extcap which is better tool.
Testpmd doesn't need more bloat and opening a socket in testpmd would
also expose more attack surface

> 
> Otherwise, having seen the demonstration, I do not have many comments, 
> except for a couple nits below.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21  9:51   ` Marat Khalili
  2026-09-21 15:57     ` Stephen Hemminger
@ 2026-09-21 15:58     ` Stephen Hemminger
  2026-09-21 16:43       ` Marat Khalili
  2026-09-21 16:17     ` Stephen Hemminger
  2 siblings, 1 reply; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-21 15:58 UTC (permalink / raw)
  To: Marat Khalili; +Cc: dev

On Mon, 21 Sep 2026 10:51:39 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> >   examples/rpcapd/meson.build            |   19 +
> >   examples/rpcapd/rpcap-protocol.h       |  127 +++  
> Could we just use one from libpcap?

Could but since most distro's don't enable rpcapd the header
file is not exported. We could "vendor it".

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21  9:51   ` Marat Khalili
  2026-09-21 15:57     ` Stephen Hemminger
  2026-09-21 15:58     ` Stephen Hemminger
@ 2026-09-21 16:17     ` Stephen Hemminger
  2 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-21 16:17 UTC (permalink / raw)
  To: Marat Khalili; +Cc: dev

On Mon, 21 Sep 2026 10:51:39 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> >   examples/rpcapd/main.c                 | 1430 ++++++++++++++++++++++++  
> 
> main.c is a bit large and handles many different functions. It would be 
> more readable if different aspects were handled by different files, or 
> at least clearly marked sections of the file, with overall structure 
> explained somewhere at the top.

It is mostly comments.
Fits within the general sizes of examples.

source   blank   comments   comment-lines 
    2569      499       13      17  examples/fips_validation/main.c
    2302      411      272     151  examples/l3fwd-power/main.c
    2208      452      226     173  examples/l2fwd-crypto/main.c
    1542      343      200     102  examples/vhost/main.c
    1489      226      104      77  examples/l3fwd/main.c
    1254      216      101      99  examples/l2fwd-macsec/main.c
    1116      224       99      85  examples/l3fwd-graph/main.c
    1045      190      195      72  examples/rpcapd/main.c
     944      181       69      60  examples/bbdev_app/main.c
                            

> >   examples/rpcapd/meson.build            |   19 +
> >   examples/rpcapd/rpcap-protocol.h       |  127 +++  
> Could we just use one from libpcap?

The long answer:

- the file is internal to libpcap and not shipped.
- Would have to add SPDX
- It isn't just wire format but includes function prototypes
  for code in libpcap's implementation
- It is ugly style overall: lots of typedefs, windows-ism's, #pragma

What AI did was extract the parts that were used.


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21 15:58     ` Stephen Hemminger
@ 2026-09-21 16:43       ` Marat Khalili
  2026-09-21 17:31         ` Stephen Hemminger
  0 siblings, 1 reply; 21+ messages in thread
From: Marat Khalili @ 2026-09-21 16:43 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: dev

On 21/09/2026 16:58, Stephen Hemminger wrote:
> On Mon, 21 Sep 2026 10:51:39 +0100
> Marat Khalili <qm2k21@gmail.com> wrote:
>
>>>    examples/rpcapd/meson.build            |   19 +
>>>    examples/rpcapd/rpcap-protocol.h       |  127 +++
>> Could we just use one from libpcap?
> Could but since most distro's don't enable rpcapd the header
> file is not exported. We could "vendor it".
Is it possible to use your new app without having a full version of 
libpcap in the system (that would have this rpcap-protocol.h anyway)? In 
"Client Setup" you mention that libpcap must be rebuilt with remote 
support enabled, and at this moment our own abridged copy of 
rpcap-protocol.h becomes redundant. Our copy allows building the 
application in systems where rpcapd is missing or built without the 
remote support, but then the build result cannot be used. Am I 
misunderstanding something?

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21 16:43       ` Marat Khalili
@ 2026-09-21 17:31         ` Stephen Hemminger
  2026-09-21 17:53           ` Marat Khalili
  0 siblings, 1 reply; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-21 17:31 UTC (permalink / raw)
  To: Marat Khalili; +Cc: dev

On Mon, 21 Sep 2026 17:43:31 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> On 21/09/2026 16:58, Stephen Hemminger wrote:
> > On Mon, 21 Sep 2026 10:51:39 +0100
> > Marat Khalili <qm2k21@gmail.com> wrote:
> >  
> >>>    examples/rpcapd/meson.build            |   19 +
> >>>    examples/rpcapd/rpcap-protocol.h       |  127 +++  
> >> Could we just use one from libpcap?  
> > Could but since most distro's don't enable rpcapd the header
> > file is not exported. We could "vendor it".  
> Is it possible to use your new app without having a full version of 
> libpcap in the system (that would have this rpcap-protocol.h anyway)? In 
> "Client Setup" you mention that libpcap must be rebuilt with remote 
> support enabled, and at this moment our own abridged copy of 
> rpcap-protocol.h becomes redundant. Our copy allows building the 
> application in systems where rpcapd is missing or built without the 
> remote support, but then the build result cannot be used. Am I 
> misunderstanding something?

Read the docs, there are two issues:
 1. You need to rebuild libpcap since rpcap is a security problem,
    it is not enabled by default. Tcpdump et al will not support
    rpcap without rebuild.

 2. The rpcap-protocol.h file is not exported by libpcap and
    considered internal. The original is too damn messy to use
    in DPDK as a copy.


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-21 17:31         ` Stephen Hemminger
@ 2026-09-21 17:53           ` Marat Khalili
  0 siblings, 0 replies; 21+ messages in thread
From: Marat Khalili @ 2026-09-21 17:53 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: dev

>   2. The rpcap-protocol.h file is not exported by libpcap and
>      considered internal. The original is too damn messy to use
>      in DPDK as a copy.
Ok, I missed that part, Strange that they don't have a public header 
with protocol constants, but I guess there is little we can do, we have 
to have our own version. Maybe worth explaining it in the file header.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v2] examples/rpcapd: demo version of packet capture daemon
  2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
  2026-09-21  9:51   ` Marat Khalili
@ 2026-09-22 18:45   ` Stephen Hemminger
  1 sibling, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-22 18:45 UTC (permalink / raw)
  To: dev; +Cc: Thomas Monjalon, Reshma Pattan

On Sun, 20 Sep 2026 11:59:09 -0700
Stephen Hemminger <stephen@networkplumber.org> wrote:

> This example adds RPCAP support over localhost TCP
> integrated with DPDK. It uses a secondary process that allows
> connections from using tcpdump defacto protocol rpcap.
> 
> See: doc/guides/sample_app_ug/rpcapd.rst for more info
> 
> Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
> ---

Recheck-request: rebase=main, iol-unit-arm64-testing

^ permalink raw reply	[flat|nested] 21+ messages in thread

* [PATCH v3] examples/rpcapd: demo version of packet capture daemon
  2026-09-08 21:07 [PATCH] examples/rpcapd: demo version of packet capture daemon Stephen Hemminger
  2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
@ 2026-09-22 21:31 ` Stephen Hemminger
  2026-09-28 16:18   ` Marat Khalili
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
  2 siblings, 1 reply; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-22 21:31 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Thomas Monjalon, Reshma Pattan

This example adds RPCAP support over localhost TCP
integrated with DPDK. It uses a secondary process that allows
connections from using tcpdump defacto protocol rpcap.

See: doc/guides/sample_app_ug/rpcapd.rst for more info

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---

v3 - rebase to force retest

 MAINTAINERS                            |    2 +
 doc/guides/rel_notes/release_26_11.rst |    4 +
 doc/guides/sample_app_ug/index.rst     |    1 +
 doc/guides/sample_app_ug/rpcapd.rst    |  216 ++++
 examples/meson.build                   |    1 +
 examples/rpcapd/main.c                 | 1430 ++++++++++++++++++++++++
 examples/rpcapd/meson.build            |   19 +
 examples/rpcapd/rpcap-protocol.h       |  127 +++
 8 files changed, 1800 insertions(+)
 create mode 100644 doc/guides/sample_app_ug/rpcapd.rst
 create mode 100644 examples/rpcapd/main.c
 create mode 100644 examples/rpcapd/meson.build
 create mode 100644 examples/rpcapd/rpcap-protocol.h

diff --git a/MAINTAINERS b/MAINTAINERS
index 8c50c52933..1eca09646d 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -1724,6 +1724,8 @@ F: app/pdump/
 F: doc/guides/tools/pdump.rst
 F: app/dumpcap/
 F: doc/guides/tools/dumpcap.rst
+F: examples/rpcapd/
+F: doc/guides/sample_app_ug/rpcapd.rst
 
 
 Packet Framework
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index f5d10d3de4..ac57c19e2f 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -65,6 +65,10 @@ New Features
 
   * Added VF support on AMD Solarflare X45xx adapters.
 
+* **Added an example of tcpdump remote pcap daemon.**
+
+  Added an example that implements rpcap to allow live capture in tcpdump.
+
 
 Removed Items
 -------------
diff --git a/doc/guides/sample_app_ug/index.rst b/doc/guides/sample_app_ug/index.rst
index f12623bb66..61ed870318 100644
--- a/doc/guides/sample_app_ug/index.rst
+++ b/doc/guides/sample_app_ug/index.rst
@@ -31,6 +31,7 @@ Sample Applications User Guides
     l3_forward_graph
     l3_forward_power_man
     link_status_intr
+    rpcapd
     server_node_efd
     service_cores
     multi_process
diff --git a/doc/guides/sample_app_ug/rpcapd.rst b/doc/guides/sample_app_ug/rpcapd.rst
new file mode 100644
index 0000000000..6afbfa216c
--- /dev/null
+++ b/doc/guides/sample_app_ug/rpcapd.rst
@@ -0,0 +1,216 @@
+..  SPDX-License-Identifier: BSD-3-Clause
+    Copyright(c) 2026 Stephen Hemminger
+
+.. _rpcapd_app:
+
+dpdk-rpcapd Sample Application
+==============================
+
+The ``dpdk-rpcapd`` sample application is a Data Plane Development Kit
+(DPDK) implementation of the remote packet capture daemon protocol
+(``rpcap``) used by libpcap.  It runs as a DPDK secondary process and
+allows libpcap-aware tools such as ``tcpdump`` and Wireshark to capture
+packets from a DPDK primary process live, without writing to an
+intermediate file.
+
+The ``dpdk-rpcapd`` tool implements a subset of the protocol spoken by
+the libpcap project's ``rpcapd``.
+See
+https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd for the
+reference implementation.
+Clients connect to ``dpdk-rpcapd`` using a ``rpcap://`` URL,
+request the list of available interfaces(which are the ports of the DPDK primary),
+open one, and stream packets from it.
+
+The intended workflow is one-step capture: start the primary, start
+``dpdk-rpcapd``, and point a familiar tool at it.  No intermediate files,
+no separate post-processing step.
+
+.. warning::
+
+   ``dpdk-rpcapd`` listens on an unauthenticated, unencrypted TCP port
+   (default 2002, bound to ``127.0.0.1``).  Any local user able to
+   reach the port can list DPDK ports and capture all traffic flowing
+   through them.  This is a sample application intended for
+   development, debugging, and demonstration use only.  **Do not run
+   ``dpdk-rpcapd`` on a production system.**
+
+   The default bind address is ``127.0.0.1`` so the listener is not
+   reachable from other hosts.  An operator may override this with
+   ``--bind <addr>`` but should expect that the resulting deployment
+   exposes captured traffic to anyone who can reach that address; do
+   not do this on an untrusted network.
+
+
+.. note::
+
+   * ``dpdk-rpcapd`` is experimental and provided for demonstration purposes only.
+     It may change or be removed without notice, and it is not intended to be relied upon.
+
+
+Running the Application
+-----------------------
+
+The application has a small set of command-line options:
+
+*   ``-p <port>``, ``--port <port>``
+
+    TCP port to listen on.  Default is 2002, the IANA-assigned rpcap
+    port.
+
+*   ``-b <addr>``, ``--bind <addr>``
+
+    Numeric IPv4 or IPv6 address to bind the listener to.  Default is
+    ``127.0.0.1`` (loopback only).  Setting any other address exposes
+    captured traffic to the network and should not be done on untrusted
+    networks.
+
+*   ``-4``
+
+    Use only IPv4; an IPv6 argument to ``-b`` is rejected.
+
+*   ``-N <ring_size>``
+
+    Size of the per-session capture ring in packets.  Default is 2048.
+    Rounded up to a power of two if necessary.
+
+*   ``-D``, ``--debug``
+
+    Increase log verbosity.  By default only notices, warnings and
+    errors are printed.  A single ``-D`` adds session-level messages
+    (client connected, capture started and stopped); ``-DD`` adds
+    per-request protocol detail.
+
+*   ``--debug-file <file>``
+
+    Append log output to ``<file>`` instead of writing it to standard
+    error.
+
+*   ``--lcore <core>``
+
+    CPU core to run on.  By default the daemon runs as an ordinary
+    process on any non-isolated CPU.
+
+*   ``--file-prefix <prefix>``
+
+    EAL file prefix of the primary process to attach to.  Needed when
+    the primary was started with a non-default prefix.
+
+*   ``--version``
+
+    Print the version and exit.
+
+*   ``-h``, ``--help``
+
+    Print usage and exit.
+
+EAL options are supplied automatically; the application runs as a
+secondary process and does not need EAL options on its command line for
+typical use.
+
+
+Client Setup
+------------
+
+Most Linux distributions ship libpcap built without ``rpcap`` support
+because the libpcap project leaves ``--enable-remote`` off by default.
+To use ``dpdk-rpcapd`` from ``tcpdump`` or Wireshark on Linux, libpcap
+must be rebuilt with remote support enabled.  Approximate steps:
+
+.. code-block:: console
+
+    wget https://www.tcpdump.org/release/libpcap-1.10.7.tar.xz
+    tar xf libpcap-1.10.7.tar.xz
+    cd libpcap-1.10.7
+    ./configure --enable-remote
+    make
+    sudo make install
+
+Only the client side of ``rpcap`` is used for ``dpdk-rpcapd``.
+Do not run libpcap's version of ``rpcapd``.
+
+``tcpdump`` rebuilt against this libpcap can be used as a client without
+further changes.  Wireshark on Windows and macOS ships with rpcap support
+enabled by default.
+
+
+Example
+-------
+
+Start a primary application with the packet capture framework
+initialized.  ``dpdk-testpmd`` is the simplest:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-testpmd --vdev=net_tap0 -- -i
+
+In another window, start ``dpdk-rpcapd``:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/examples/dpdk-rpcapd
+    RPCAPD: open_listen_socket(): listening on 127.0.0.1 port 2002
+
+In a third window, list available interfaces using a libpcap-based
+``tcpdump`` rebuilt with remote support:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump --list-remote-interfaces=rpcap://localhost:2002/
+    rpcap://localhost:2002/net_tap0  Network adapter 'DPDK port' on remote node localhost
+
+Capture live from a port:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -nn -c 20
+
+Or save to a file readable by any pcap consumer:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -w /tmp/capture.pcap
+
+
+Limitations
+-----------
+
+The following limitations apply to this initial version of
+``dpdk-rpcapd`` and are expected to be addressed in subsequent patches:
+
+*   **Single client.** Only one client may be connected at a time.
+    Subsequent clients are queued by the listening socket but not
+    serviced until the first disconnects.  Multi-client support
+    requires an event-driven main loop (planned).
+
+*   **No authentication.** ``AUTH`` requests are acknowledged with an
+    empty reply (libpcap "version 0, null auth" semantics).  This
+    sample application does not implement password authentication.
+
+*   **TCP transport only; not for production use.** The rpcap protocol
+    over TCP is unauthenticated and unencrypted; any client that can
+    reach the listening port has full access to captured traffic.
+    Binding to ``127.0.0.1`` by default mitigates remote exposure but
+    does not address local users on a shared host.  See the warning at
+    the top of this document.
+
+*   **Microsecond timestamp resolution.** The rpcap protocol carries
+    timestamps at microsecond resolution.
+
+*   **Original length of truncated packets is not reported.** The
+    capture framework in the primary process copies only the snaplen
+    worth of bytes and does not carry the original frame length across
+    to the secondary, so a truncated packet is reported to the client
+    with its on-the-wire length equal to its captured length.  A frame
+    longer than the snaplen therefore appears to the client as a short
+    frame rather than as a truncated long one.
+
+
+See Also
+--------
+
+*   :doc:`../tools/dumpcap` -- file-based capture writing pcapng
+    output.
+
+*   The libpcap project's ``rpcapd`` reference implementation:
+    https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
diff --git a/examples/meson.build b/examples/meson.build
index 25d9c88457..24b6184353 100644
--- a/examples/meson.build
+++ b/examples/meson.build
@@ -45,6 +45,7 @@ all_examples = [
         'ptpclient',
         'qos_meter',
         'qos_sched',
+        'rpcapd',
         'rxtx_callbacks',
         'server_node_efd/efd_node',
         'server_node_efd/efd_server',
diff --git a/examples/rpcapd/main.c b/examples/rpcapd/main.c
new file mode 100644
index 0000000000..530543ce7e
--- /dev/null
+++ b/examples/rpcapd/main.c
@@ -0,0 +1,1430 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Demonstration server for the rpcap protocol for DPDK.
+ * This allows a libpcap client (e.g. Wireshark or tcpdump)
+ * to use "rpcap://host[:port]/portname" as capture device.
+ *
+ * Based on the DPDK dumpcap application and on rpcapd from libpcap:
+ *   https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
+ *
+ * Only the bits of the RPCAP protocol that are needed for an
+ * unauthenticated, passive-mode capture session are implemented.
+ * Configuration files, active mode, sampling and concurrent clients
+ * are intentionally omitted to keep the example small.
+ *
+ * A capture filter may be sent with the start-capture request:
+ * the client compiles it, so it arrives as cBPF which is converted to
+ * DPDK BPF and handed to pdump.  Filters cannot be changed once the
+ * capture is running; see the UPDATEFILTER handling.
+ */
+
+#include <arpa/inet.h>
+#include <errno.h>
+#include <getopt.h>
+#include <netinet/in.h>
+#include <netdb.h>
+#include <poll.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/time.h>
+#include <sys/types.h>
+#include <sys/uio.h>
+#include <unistd.h>
+
+#include <pcap/pcap.h>
+
+#include <rte_alarm.h>
+#include <rte_bpf.h>
+#include <rte_common.h>
+#include <rte_debug.h>
+#include <rte_eal.h>
+#include <rte_errno.h>
+#include <rte_ethdev.h>
+#include <rte_lcore.h>
+#include <rte_log.h>
+#include <rte_malloc.h>
+#include <rte_mbuf.h>
+#include <rte_mempool.h>
+#include <rte_pdump.h>
+#include <rte_stdatomic.h>
+#include <rte_ring.h>
+#include <rte_version.h>
+
+#include "rpcap-protocol.h"
+
+#define BURST_SIZE                    32
+#define MBUF_CACHE_SIZE               32
+#define DEFAULT_RING_SIZE             2048
+#define MAX_RING_SIZE                 (1U << 20)
+#define DEFAULT_SNAPLEN               RTE_MBUF_DEFAULT_DATAROOM
+#define PRIMARY_MONITOR_INTERVAL_US   (500 * 1000)
+#define SLEEP_THRESHOLD		      100
+#define SLEEP_US		      100
+
+#define DATA_ACCEPT_TIMEOUT_MS        10000
+#define POLL_INTERVAL_MS              500
+
+#define MAX_FILTER_INSNS              4096
+
+#define RTE_LOGTYPE_RPCAPD RTE_LOGTYPE_USER1
+#define RPCAPD_LOG(level, ...) \
+	RTE_LOG_LINE_PREFIX(level, RPCAPD, "%s(): ", __func__, __VA_ARGS__)
+
+/* Per-client capture session state. */
+struct session {
+	int      data_fd;
+	uint16_t port;				/* DPDK ethdev port being captured */
+	char     name[RTE_ETH_NAME_MAX_LEN];
+	uint32_t snaplen;
+	uint32_t npkt;				/* packet sequence for rpcap_pkthdr */
+	uint32_t pdump_flags;			/* RTE_PDUMP_FLAG_* in use */
+	bool     opened;			/* OPEN_REQ has selected a port */
+	bool     capture_on;
+	bool     promisc_set;			/* we enabled promiscuous mode */
+	struct rte_ring    *ring;
+	struct rte_mempool *mp;
+	struct rte_bpf_prm *prm;		/* capture filter, NULL if none */
+};
+
+/* Command-line options */
+static uint16_t listen_port = RPCAP_DEFAULT_NETPORT;
+static uint32_t ring_size = DEFAULT_RING_SIZE;
+static const char *lcore_arg;
+static const char *file_prefix;
+static const char *bind_arg;		/* -b argument, resolved after option parsing */
+static const char *debug_file;		/* --debug-file argument */
+static bool ipv4_only;			/* -4: restrict to IPv4 */
+static unsigned int debug_log;		/* -D count: raise RPCAPD log verbosity */
+
+static struct sockaddr_storage listen_addr;
+static socklen_t               listen_addrlen;
+
+static void stop_capture(struct session *s);
+
+static void
+set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port)
+{
+	if (ss->ss_family == AF_INET6)
+		((struct sockaddr_in6 *)ss)->sin6_port = htons(port);
+	else
+		((struct sockaddr_in *)ss)->sin_port = htons(port);
+}
+
+static uint16_t
+get_sockaddr_port(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET6)
+		return ntohs(((const struct sockaddr_in6 *)ss)->sin6_port);
+	return ntohs(((const struct sockaddr_in *)ss)->sin_port);
+}
+
+static bool
+is_loopback(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET) {
+		const struct sockaddr_in *sin = (const void *)ss;
+
+		return (ntohl(sin->sin_addr.s_addr) >> 24) == 127;
+	}
+	if (ss->ss_family == AF_INET6) {
+		const struct sockaddr_in6 *sin6 = (const void *)ss;
+
+		return IN6_IS_ADDR_LOOPBACK(&sin6->sin6_addr);
+	}
+	return false;
+}
+
+static void
+parse_bind_addr(const char *str, int family)
+{
+	struct addrinfo hints = {
+		.ai_family   = family,
+		.ai_socktype = SOCK_STREAM,
+		.ai_flags    = AI_NUMERICHOST | AI_PASSIVE,
+	};
+	struct addrinfo *res;
+	int rc;
+
+	rc = getaddrinfo(str, NULL, &hints, &res);
+	if (rc != 0)
+		rte_exit(EXIT_FAILURE, "Invalid bind address '%s': %s\n",
+			 str, gai_strerror(rc));
+	memcpy(&listen_addr, res->ai_addr, res->ai_addrlen);
+	listen_addrlen = res->ai_addrlen;
+	freeaddrinfo(res);
+}
+
+static RTE_ATOMIC(bool) quit_signal;
+
+static void
+signal_handler(int sig __rte_unused)
+{
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+/*
+ * Wait for fd to become readable, in POLL_INTERVAL_MS slices so that a
+ * quit signal (from SIGINT/SIGTERM or from the primary process dying)
+ * is noticed while blocked.  timeout_ms < 0 waits indefinitely.
+ *
+ * Returns 1 when readable, 0 on timeout, -1 on error or quit.
+ */
+static int
+wait_readable(int fd, int timeout_ms)
+{
+	struct pollfd pfd = { .fd = fd, .events = POLLIN };
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int wait_ms = POLL_INTERVAL_MS;
+		int rc;
+
+		if (timeout_ms >= 0) {
+			if (timeout_ms == 0)
+				return 0;
+			if (timeout_ms < wait_ms)
+				wait_ms = timeout_ms;
+			timeout_ms -= wait_ms;
+		}
+
+		rc = poll(&pfd, 1, wait_ms);
+		if (rc < 0) {
+			if (errno == EINTR)
+				continue;
+			RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+			return -1;
+		}
+		if (rc > 0)
+			return 1;
+	}
+	return -1;
+}
+
+/* accept() with a timeout, so a stalled client cannot wedge the daemon. */
+static int
+accept_timeout(int listen_fd, int timeout_ms)
+{
+	int fd;
+
+	switch (wait_readable(listen_fd, timeout_ms)) {
+	case 1:
+		break;
+	case 0:
+		RPCAPD_LOG(ERR, "timed out waiting for data connection");
+		return -1;
+	default:
+		return -1;
+	}
+
+	fd = accept(listen_fd, NULL, NULL);
+	if (fd < 0)
+		RPCAPD_LOG(ERR, "accept: %s", strerror(errno));
+	return fd;
+}
+
+/* Read exactly len bytes; return 0 on success, -1 on error or EOF. */
+static int
+recv_full(int fd, void *buf, size_t len)
+{
+	uint8_t *p = buf;
+
+	while (len > 0) {
+		ssize_t n;
+
+		/* Wait with a timeout rather than blocking in recv(), so a
+		 * quit signal or a dead primary is acted on promptly.
+		 */
+		if (wait_readable(fd, -1) != 1)
+			return -1;
+
+		n = recv(fd, p, len, 0);
+		if (n < 0 && errno == EINTR)
+			continue;
+
+		if (n <= 0)
+			return -1;
+
+		p += n;
+		len -= n;
+	}
+	return 0;
+}
+
+/*
+ * Send all of iov, resending the remainder if sendmsg() reports a short
+ * count (possible when the connection breaks or a signal arrives after
+ * some bytes were copied).  Consumes iov, so pass a scratch copy.
+ */
+static int
+send_iov_full(int fd, struct iovec *iov, int iovcnt, int flags)
+{
+	struct msghdr msg = {
+		.msg_iov    = iov,
+		.msg_iovlen = iovcnt,
+	};
+
+	while (msg.msg_iovlen > 0) {
+		ssize_t n = sendmsg(fd, &msg, flags | MSG_NOSIGNAL);
+
+		if (n < 0) {
+			if (errno == EINTR)
+				continue;
+			return -1;
+		}
+		if (n == 0)
+			return -1;
+
+		/* Drop whole iovecs that were fully sent, then trim the
+		 * partially sent one.
+		 */
+		while (msg.msg_iovlen > 0 && (size_t)n >= msg.msg_iov->iov_len) {
+			n -= msg.msg_iov->iov_len;
+			msg.msg_iov++;
+			msg.msg_iovlen--;
+		}
+		if (n > 0) {
+			msg.msg_iov->iov_base = (char *)msg.msg_iov->iov_base + n;
+			msg.msg_iov->iov_len -= n;
+		}
+	}
+	return 0;
+}
+
+static int
+rpcap_send_msg(int fd, uint8_t type, uint16_t value, const void *payload, uint32_t plen)
+{
+	struct rpcap_header hdr = {
+		.ver = RPCAP_VERSION,
+		.type = type,
+		.value = htons(value),
+		.plen = htonl(plen),
+	};
+	struct iovec iov[2] = {
+		{ .iov_base = &hdr,                        .iov_len = sizeof(hdr) },
+		{ .iov_base = (void *)(uintptr_t)payload,  .iov_len = plen },
+	};
+
+	return send_iov_full(fd, iov, plen > 0 ? 2 : 1, 0);
+}
+
+static int
+rpcap_send_error(int fd, uint16_t errcode, const char *msg)
+{
+	RPCAPD_LOG(WARNING, "sending error to client: %s", msg);
+	return rpcap_send_msg(fd, RPCAP_MSG_ERROR, errcode, msg, strlen(msg));
+}
+
+static int
+rpcap_recv_header(int fd, struct rpcap_header *hdr)
+{
+	if (recv_full(fd, hdr, sizeof(*hdr)) < 0)
+		return -1;
+	hdr->value = ntohs(hdr->value);
+	hdr->plen = ntohl(hdr->plen);
+	return 0;
+}
+
+/* Throw away plen bytes of payload we don't care about. */
+static int
+rpcap_discard(int fd, uint32_t plen)
+{
+	uint8_t buf[256];
+
+	while (plen > 0) {
+		size_t chunk = plen > sizeof(buf) ? sizeof(buf) : plen;
+
+		if (recv_full(fd, buf, chunk) < 0)
+			return -1;
+		plen -= chunk;
+	}
+	return 0;
+}
+
+/* Build and send the list of available DPDK ports. */
+static int
+handle_findallif(int fd)
+{
+	uint8_t *buf = NULL;
+	size_t buflen = 0;
+	uint16_t nif = 0;
+	uint16_t p;
+	int rc;
+
+	RTE_ETH_FOREACH_DEV(p) {
+		static const char desc[] = "DPDK port";
+		char name[RTE_ETH_NAME_MAX_LEN];
+		size_t namelen, desclen, entry;
+		uint8_t *nb;
+
+		if (rte_eth_dev_get_name_by_port(p, name) < 0) {
+			RPCAPD_LOG(DEBUG, "can not find name for port %u", p);
+			continue;
+		}
+
+		RPCAPD_LOG(DEBUG, "findallif: port %u -> '%s'", p, name);
+		namelen = strlen(name);
+		desclen = strlen(desc);
+		entry = sizeof(struct rpcap_findalldevs_if) + namelen + desclen;
+
+		nb = realloc(buf, buflen + entry);
+		if (nb == NULL) {
+			RPCAPD_LOG(ERR, "out of memory in findallif");
+			free(buf);
+			return rpcap_send_error(fd, 0, "out of memory");
+		}
+		buf = nb;
+
+		struct rpcap_findalldevs_if iface = {
+			.namelen = htons(namelen),
+			.desclen = htons(desclen),
+			.flags = htonl(PCAP_IF_UP | PCAP_IF_RUNNING),
+		};
+		memcpy(buf + buflen, &iface, sizeof(iface));
+		memcpy(buf + buflen + sizeof(iface), name, namelen);
+		memcpy(buf + buflen + sizeof(iface) + namelen, desc, desclen);
+		buflen += entry;
+		nif++;
+	}
+
+	RPCAPD_LOG(DEBUG, "findallif: %u interface(s)", nif);
+	rc = rpcap_send_msg(fd, RPCAP_MSG_FINDALLIF_REPLY, nif, buf, buflen);
+	free(buf);
+	return rc;
+}
+
+/* OPEN_REQ: payload is the interface name (no NUL). */
+static int
+handle_open(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_openreply reply = {
+		.linktype = htonl(DLT_EN10MB),
+	};
+	uint16_t port;
+
+	/* Unconditionally, not just when capture_on: a failed UPDATEFILTER
+	 * leaves the ring, mempool and data connection live with the capture
+	 * already disabled, and those must not survive into a new session.
+	 */
+	stop_capture(s);
+
+	if (plen >= sizeof(s->name)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "interface name too long");
+	}
+	if (recv_full(fd, s->name, plen) < 0)
+		return -1;
+	s->name[plen] = '\0';
+
+	if (rte_eth_dev_get_port_by_name(s->name, &port) < 0) {
+		RPCAPD_LOG(WARNING, "open: no such port '%s'", s->name);
+		/* s->name has already been overwritten; make sure a later
+		 * STARTCAP cannot capture the previously opened port.
+		 */
+		s->opened = false;
+		return rpcap_send_error(fd, 0, "unknown interface");
+	}
+	s->port = port;
+	s->opened = true;
+
+	RPCAPD_LOG(DEBUG, "open: '%s' -> dpdk port %u", s->name, port);
+	return rpcap_send_msg(fd, RPCAP_MSG_OPEN_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Open an ephemeral TCP listening socket; return fd, set *port_out. */
+static int
+open_data_listener(uint16_t *port_out)
+{
+	struct sockaddr_storage addr = listen_addr;
+	socklen_t alen;
+	int fd;
+
+	set_sockaddr_port(&addr, 0);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0) {
+		RPCAPD_LOG(ERR, "data socket: %s", strerror(errno));
+		return -1;
+	}
+
+	alen = listen_addrlen;
+	if (bind(fd, (struct sockaddr *)&addr, alen) < 0 ||
+	    listen(fd, 1) < 0 ||
+	    getsockname(fd, (struct sockaddr *)&addr, &alen) < 0) {
+		RPCAPD_LOG(ERR, "data port bind/listen: %s", strerror(errno));
+		close(fd);
+		return -1;
+	}
+	*port_out = get_sockaddr_port(&addr);
+	return fd;
+}
+
+static struct rte_ring *
+create_capture_ring(uint16_t port)
+{
+	char name[RTE_RING_NAMESIZE];
+
+	snprintf(name, sizeof(name), "rpcapd_r_%u_%d", port, getpid());
+	return rte_ring_create(name, ring_size, rte_socket_id(), 0);
+}
+
+static struct rte_mempool *
+create_capture_mempool(uint16_t port, uint32_t snaplen)
+{
+	char name[RTE_MEMPOOL_NAMESIZE];
+	uint32_t mbuf_size = RTE_PKTMBUF_HEADROOM + snaplen;
+
+	snprintf(name, sizeof(name), "rpcapd_p_%u_%d", port, getpid());
+	return rte_pktmbuf_pool_create(name, ring_size * 2, MBUF_CACHE_SIZE, 0,
+				       mbuf_size, rte_socket_id());
+}
+
+/*
+ * Read the optional capture filter that follows a start-capture request,
+ * and convert it for pdump. Client passes cBPF.
+ */
+static int
+read_filter(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_filterbpf_insn winsn;
+	struct rpcap_filter filter;
+	struct bpf_program bf;
+	struct bpf_insn *insns;
+	uint32_t i, nitems;
+
+	if (plen == 0)
+		return 0;		/* no filter: capture everything */
+
+	if (plen < sizeof(filter)) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "short filter header") < 0 ? -1 : 1;
+	}
+
+	if (recv_full(fd, &filter, sizeof(filter)) < 0)
+		return -1;
+	plen -= sizeof(filter);
+
+	if (ntohs(filter.filtertype) != RPCAP_UPDATEFILTER_BPF) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "unsupported filter type") < 0 ? -1 : 1;
+	}
+
+	/* nitems is client-supplied; bound it before trusting the length. */
+	nitems = ntohl(filter.nitems);
+	if (nitems == 0)
+		return rpcap_discard(fd, plen) < 0 ? -1 : 0;
+
+	if (nitems > MAX_FILTER_INSNS || plen < nitems * sizeof(winsn)) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "bad filter length") < 0 ? -1 : 1;
+	}
+
+	insns = calloc(nitems, sizeof(*insns));
+	if (insns == NULL) {
+		if (rpcap_discard(fd, plen) < 0)
+			return -1;
+		return rpcap_send_error(fd, 0, "out of memory") < 0 ? -1 : 1;
+	}
+
+	for (i = 0; i < nitems; i++) {
+		if (recv_full(fd, &winsn, sizeof(winsn)) < 0) {
+			free(insns);
+			return -1;
+		}
+		insns[i].code = ntohs(winsn.code);
+		insns[i].jt   = winsn.jt;
+		insns[i].jf   = winsn.jf;
+		insns[i].k    = ntohl(winsn.k);
+	}
+	plen -= nitems * sizeof(winsn);
+
+	/* Anything after the instructions is padding we do not need. */
+	if (rpcap_discard(fd, plen) < 0) {
+		free(insns);
+		return -1;
+	}
+
+	bf.bf_len = nitems;
+	bf.bf_insns = insns;
+
+	/* Reject a malformed program here */
+	if (!bpf_validate(bf.bf_insns, bf.bf_len)) {
+		free(insns);
+		return rpcap_send_error(fd, 0, "invalid filter program") < 0 ? -1 : 1;
+	}
+
+	/* A filter recorded by an earlier UPDATEFILTER may still be here;
+	 * it is about to be replaced, so do not leak it.
+	 */
+	rte_free(s->prm);
+	s->prm = rte_bpf_convert(&bf);
+	free(insns);
+	if (s->prm == NULL) {
+		RPCAPD_LOG(ERR, "rte_bpf_convert failed: %s",
+			rte_strerror(rte_errno));
+		return rpcap_send_error(fd, 0, "cannot convert filter") < 0 ? -1 : 1;
+	}
+
+	RPCAPD_LOG(DEBUG, "capture filter: %u instructions", nitems);
+	return 0;
+}
+
+/* Tear down anything that handle_startcap brought up.  Safe to call
+ * after partial setup as well as after a successful capture.
+ */
+static void
+stop_capture(struct session *s)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int n;
+
+	if (s->capture_on) {
+		rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+		RPCAPD_LOG(INFO, "capture stopped on %s (%u packets)",
+			s->name, s->npkt);
+	}
+	s->capture_on = false;
+
+	if (s->promisc_set) {
+		rte_eth_promiscuous_disable(s->port);
+		s->promisc_set = false;
+	}
+
+	if (s->ring != NULL) {
+		while ((n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts,
+						      BURST_SIZE, NULL)) > 0)
+			rte_pktmbuf_free_bulk(pkts, n);
+		rte_ring_free(s->ring);
+		s->ring = NULL;
+	}
+	if (s->mp != NULL) {
+		rte_mempool_free(s->mp);
+		s->mp = NULL;
+	}
+
+	/* Only safe once pdump is disabled */
+	rte_free(s->prm);
+	s->prm = NULL;
+	if (s->data_fd >= 0) {
+		close(s->data_fd);
+		s->data_fd = -1;
+	}
+}
+
+/*
+ * STARTCAP_REQ: open the data connection and arm the pdump callback.
+ * We use passive mode with the server-allocated data port:
+ *   - the server picks an ephemeral port and listens on it
+ *   - the server returns that port in startcapreply.portdata
+ *   - the client connects back to that port for the packet stream
+ */
+static int
+handle_startcap(int fd, uint32_t plen, struct session *s)
+{
+	struct rpcap_startcapreq req;
+	uint16_t data_port;
+	uint16_t flags;
+	struct rte_bpf_prm *recorded;
+	int data_listen;
+	int data_fd;
+	int ret;
+
+	recorded = s->prm;
+	s->prm = NULL;
+	stop_capture(s);
+	s->prm = recorded;
+
+	if (!s->opened) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "no interface open");
+	}
+
+	if (plen < sizeof(req)) {
+		rpcap_discard(fd, plen);
+		return rpcap_send_error(fd, 0, "short startcap request");
+	}
+	if (recv_full(fd, &req, sizeof(req)) < 0)
+		return -1;
+
+	flags = ntohs(req.flags);
+	if (flags & RPCAP_STARTCAPREQ_FLAG_DGRAM) {
+		rpcap_discard(fd, plen - sizeof(req));
+		return rpcap_send_error(fd, 0, "UDP data transfer not supported");
+	}
+
+	ret = read_filter(fd, plen - sizeof(req), s);
+	if (ret != 0)
+		return ret < 0 ? -1 : 0;	/* error already reported to client */
+
+	/* Direction flags map onto pdump's RX/TX selection; neither (or both)
+	 * means capture in both directions.
+	 */
+	s->pdump_flags = RTE_PDUMP_FLAG_RXTX;
+	if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+		      RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+	    RPCAP_STARTCAPREQ_FLAG_INBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_RX;
+	else if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+			   RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+		 RPCAP_STARTCAPREQ_FLAG_OUTBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_TX;
+
+	s->snaplen = ntohl(req.snaplen);
+	if (s->snaplen == 0 || s->snaplen > DEFAULT_SNAPLEN)
+		s->snaplen = DEFAULT_SNAPLEN;
+
+	s->ring = create_capture_ring(s->port);
+	s->mp = create_capture_mempool(s->port, s->snaplen);
+	if (s->ring == NULL || s->mp == NULL) {
+		RPCAPD_LOG(ERR, "ring/mempool alloc failed: %s",
+			rte_strerror(rte_errno));
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "DPDK alloc failed");
+	}
+
+	data_listen = open_data_listener(&data_port);
+	if (data_listen < 0) {
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "data port setup failed");
+	}
+
+	/* Leave the port alone if it is already promiscuous: it belongs to
+	 * the primary process, and stop_capture() must not turn off
+	 * something this daemon did not turn on.
+	 */
+	if ((flags & RPCAP_STARTCAPREQ_FLAG_PROMISC) &&
+	    rte_eth_promiscuous_get(s->port) != 1) {
+		if (rte_eth_promiscuous_enable(s->port) == 0)
+			s->promisc_set = true;
+		else
+			RPCAPD_LOG(NOTICE, "cannot enable promiscuous mode on %s",
+				s->name);
+	}
+
+	/* Arm pdump before replying. */
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		close(data_listen);
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "cannot enable capture");
+	}
+	s->capture_on = true;
+	s->npkt = 0;
+
+	struct rpcap_startcapreply reply = {
+		.bufsize = htonl(s->snaplen * BURST_SIZE),
+		.portdata = htons(data_port),
+	};
+	if (rpcap_send_msg(fd, RPCAP_MSG_STARTCAP_REPLY, 0, &reply, sizeof(reply)) < 0) {
+		close(data_listen);
+		stop_capture(s);
+		return -1;
+	}
+
+	RPCAPD_LOG(DEBUG, "awaiting connection");
+
+	data_fd = accept_timeout(data_listen, DATA_ACCEPT_TIMEOUT_MS);
+	close(data_listen);
+	if (data_fd < 0) {
+		stop_capture(s);
+		return -1;
+	}
+
+	s->data_fd = data_fd;
+
+	RPCAPD_LOG(INFO,
+		"capture started on %s (snaplen %u, data port %u)",
+		s->name, s->snaplen, data_port);
+	return 0;
+}
+
+/*
+ * UPDATEFILTER_REQ: replace the capture filter.
+ *
+ * pdump takes its filter when the callback is armed and offers no way
+ * to replace it, so this disables and re-enables the callback with the
+ * new program.  Packets already in the ring are kept; only the brief
+ * gap between disable and enable is lost.  Refusing the request is not
+ * an option: libpcap sends UPDATEFILTER right after STARTCAP when the
+ * client was opened with PCAP_OPENFLAG_NOCAPTURE_RPCAP and aborts the
+ * capture if it fails, and Wireshark sets that flag by default.
+ *
+ * Before the capture starts this just records the filter for the
+ * eventual STARTCAP.
+ */
+static int
+handle_updatefilter(int fd, uint32_t plen, struct session *s)
+{
+	struct rte_bpf_prm *old = s->prm;
+	int ret;
+
+	s->prm = NULL;
+	ret = read_filter(fd, plen, s);
+	if (ret != 0) {
+		/* Malformed request: keep running with the old filter. */
+		rte_free(s->prm);
+		s->prm = old;
+		return ret < 0 ? -1 : 0;	/* error already reported */
+	}
+
+	if (!s->capture_on) {
+		rte_free(old);
+		return rpcap_send_msg(fd, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+	}
+
+	rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+	s->capture_on = false;
+
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		rte_free(old);
+		/* The capture cannot be resumed, so do not leave the ring,
+		 * mempool and data connection behind: the client has been
+		 * told the capture is over, and a session that is neither
+		 * capturing nor torn down has no way back.
+		 */
+		stop_capture(s);
+		return rpcap_send_error(fd, 0, "cannot apply filter");
+	}
+	s->capture_on = true;
+
+	/* Safe now that the old program is no longer referenced. */
+	rte_free(old);
+
+	RPCAPD_LOG(DEBUG, "capture filter updated on %s", s->name);
+	return rpcap_send_msg(fd, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+}
+
+/*
+ * Pull a burst from the ring, frame each packet into an RPCAP_MSG_PACKET
+ * message, and send it on the data connection.  MSG_MORE corks the
+ * socket until the ring drains, so a backlog coalesces into full
+ * segments instead of flushing every BURST_SIZE packets.
+ */
+static ssize_t
+process_ring(struct session *s, unsigned int *avail)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int i, n;
+	ssize_t written = 0;
+	struct timeval tv;
+
+	n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts, BURST_SIZE, avail);
+	if (n == 0)
+		return 0;
+
+	/* One timestamp for the whole burst */
+	gettimeofday(&tv, NULL);
+
+	for (i = 0; i < n; i++) {
+		struct rte_mbuf *m = pkts[i];
+		/* Sized from the same bound that clamps caplen below, so the
+		 * two cannot drift apart.
+		 */
+		uint8_t buf[DEFAULT_SNAPLEN];
+		uint32_t pktlen = rte_pktmbuf_pkt_len(m);
+		uint32_t caplen = pktlen < s->snaplen ? pktlen : s->snaplen;
+		const void *data;
+
+		s->npkt++;
+
+		struct rpcap_header hdr = {
+			.ver = RPCAP_VERSION,
+			.type = RPCAP_MSG_PACKET,
+			.plen = htonl(sizeof(struct rpcap_pkthdr) + caplen),
+		};
+
+		/*
+		 * pdump copies at most the snaplen into the capture mempool
+		 * and rte_pktmbuf_copy() counts only what it copied, so
+		 * pktlen is already clamped: a truncated packet is reported
+		 * with len == caplen.  The original wire length does not
+		 * reach this process.  See the Limitations section of
+		 * doc/guides/sample_app_ug/rpcapd.rst.
+		 */
+		struct rpcap_pkthdr pkthdr = {
+			.timestamp_sec = htonl((uint32_t)tv.tv_sec),
+			.timestamp_usec = htonl((uint32_t)tv.tv_usec),
+			.caplen = htonl(caplen),
+			.len = htonl(pktlen),
+			.npkt = htonl(s->npkt),
+		};
+
+		data = rte_pktmbuf_read(m, 0, caplen, buf);
+
+		struct iovec iov[3] = {
+			{ .iov_base = &hdr,                       .iov_len = sizeof(hdr) },
+			{ .iov_base = &pkthdr,                    .iov_len = sizeof(pkthdr) },
+			{ .iov_base = (void *)(uintptr_t)data,    .iov_len = caplen },
+		};
+
+		/* more to come in this burst, or still queued in the ring */
+		bool more = (i + 1 < n) || (*avail > 0);
+
+		if (send_iov_full(s->data_fd, iov, 3, more ? MSG_MORE : 0) < 0) {
+			if (errno == EPIPE || errno == ECONNRESET)
+				RPCAPD_LOG(DEBUG, "data connection closed by client");
+			else
+				RPCAPD_LOG(NOTICE, "send on data connection failed: %s",
+					   strerror(errno));
+			goto error;
+		}
+		rte_pktmbuf_free(m);
+		written += sizeof(hdr) + sizeof(pkthdr) + caplen;
+	}
+
+	return written;
+
+error:
+	rte_pktmbuf_free_bulk(pkts + i, n - i);
+	return -1;
+}
+
+/* Poll the control socket while idle.
+ * Returns 0 to keep capturing, 1 if a control message (typically
+ * ENDCAP) is pending, or -1 if the client has gone away.
+ */
+static int
+check_socket_status(int ctrl_fd)
+{
+	struct pollfd pfd = { .fd = ctrl_fd, .events = POLLIN };
+
+	if (poll(&pfd, 1, 0) < 0) {
+		if (errno == EINTR)
+			return 0;
+		RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+		return -1;
+	}
+	if (pfd.revents & (POLLERR | POLLHUP | POLLNVAL)) {
+		RPCAPD_LOG(DEBUG, "client closed control connection");
+		return -1;
+	}
+	if (pfd.revents & POLLIN)
+		return 1;
+	return 0;
+}
+
+/*
+ * Stay in the capture loop until either:
+ *   - a control message arrives (typically ENDCAP),
+ *   - the data connection breaks, or
+ *   - a quit signal is delivered.
+ *
+ * Returns 0 if the session should continue (the caller reads the
+ * pending control message), -1 if the client is gone.
+ *
+ * The control socket is polled once per iteration, not just when the
+ * ring runs dry.  A client that sends a request mid-capture blocks
+ * waiting for the reply without draining the data socket, so under
+ * sustained traffic a poll that only happens while idle never runs and
+ * both ends wedge once the socket buffers fill.
+ */
+static int
+capture_loop(int ctrl_fd, struct session *s)
+{
+	unsigned int empty_count = 0;
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		ssize_t written;
+		unsigned int avail = 0;
+
+		switch (check_socket_status(ctrl_fd)) {
+		case 1:
+			/* control message pending, let caller service it */
+			return 0;
+		case 0:
+			break;
+		default:
+			/* client is gone */
+			return -1;
+		}
+
+		written = process_ring(s, &avail);
+		if (written < 0) {
+			/* process_ring has already logged the reason */
+			return -1;
+		}
+
+		if (written > 0) {
+			/* are there more packets? */
+			empty_count = (avail == 0);
+			continue;
+		}
+
+		if (empty_count < SLEEP_THRESHOLD) {
+			/* spin a few times before checking */
+			++empty_count;
+			rte_pause();
+			continue;
+		}
+
+		/* ring has been empty for a while: stop spinning */
+		rte_delay_us_sleep(SLEEP_US);
+	}
+	return 0;
+}
+
+static int
+handle_endcap(int fd, uint32_t plen, struct session *s)
+{
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+	stop_capture(s);
+	return rpcap_send_msg(fd, RPCAP_MSG_ENDCAP_REPLY, 0, NULL, 0);
+}
+
+static int
+handle_stats(int fd, uint32_t plen, const struct session *s)
+{
+	struct rte_eth_stats es = { 0 };
+
+	if (rpcap_discard(fd, plen) < 0)
+		return -1;
+
+	if (s->capture_on)
+		rte_eth_stats_get(s->port, &es);
+
+	struct rpcap_stats reply = {
+		.ifrecv   = htonl((uint32_t)es.ipackets),
+		.ifdrop   = htonl((uint32_t)es.ierrors),
+		.krnldrop = 0,
+		.svrcapt  = htonl(s->npkt),
+	};
+	return rpcap_send_msg(fd, RPCAP_MSG_STATS_REPLY, 0, &reply, sizeof(reply));
+}
+
+/* Service a single client until it disconnects. */
+static void
+handle_client(int ctrl_fd)
+{
+	struct sockaddr_storage peer;
+	socklen_t plen = sizeof(peer);
+	char host[NI_MAXHOST] = "?";
+	struct session s = { .data_fd = -1 };
+
+	if (getpeername(ctrl_fd, (struct sockaddr *)&peer, &plen) == 0)
+		getnameinfo((struct sockaddr *)&peer, plen,
+			    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	RPCAPD_LOG(INFO, "client %s connected", host);
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		struct rpcap_header hdr;
+
+		/* Drain the ring whenever a capture is running */
+		if (s.capture_on && capture_loop(ctrl_fd, &s) < 0)
+			goto done;
+
+		if (rpcap_recv_header(ctrl_fd, &hdr) < 0)
+			break;
+
+		/* Only version 0 is spoken here */
+		if (hdr.ver != RPCAP_VERSION) {
+			RPCAPD_LOG(WARNING, "unsupported protocol version %u",
+				hdr.ver);
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_error(ctrl_fd, PCAP_ERR_WRONGVER,
+					     "unsupported protocol version") < 0)
+				goto done;
+			continue;
+		}
+
+		switch (hdr.type) {
+		case RPCAP_MSG_AUTH_REQ:
+			/* No auth: discard credentials, ack with empty reply.
+			 * libpcap treats a zero-length AUTH_REPLY as "version
+			 * 0 only, same byte order".
+			 */
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_msg(ctrl_fd, RPCAP_MSG_AUTH_REPLY, 0, NULL, 0) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_FINDALLIF_REQ:
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 || handle_findallif(ctrl_fd) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_OPEN_REQ:
+			if (handle_open(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STARTCAP_REQ:
+			if (handle_startcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_UPDATEFILTER_REQ:
+			if (handle_updatefilter(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_ENDCAP_REQ:
+			if (handle_endcap(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STATS_REQ:
+			if (handle_stats(ctrl_fd, hdr.plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_CLOSE:
+			rpcap_discard(ctrl_fd, hdr.plen);
+			goto done;
+		default:
+			RPCAPD_LOG(WARNING, "unsupported request type 0x%02x", hdr.type);
+			if (rpcap_discard(ctrl_fd, hdr.plen) < 0 ||
+			    rpcap_send_error(ctrl_fd, 0, "unsupported request") < 0)
+				goto done;
+			break;
+		}
+	}
+done:
+	stop_capture(&s);
+	close(ctrl_fd);
+	RPCAPD_LOG(INFO, "client %s disconnected", host);
+}
+
+static int
+open_listen_socket(uint16_t port)
+{
+	struct sockaddr_storage addr = listen_addr;
+	char host[NI_MAXHOST];
+	int fd, one = 1;
+
+	set_sockaddr_port(&addr, port);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0)
+		rte_exit(EXIT_FAILURE, "socket: %s\n", strerror(errno));
+	setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one));
+
+	if (bind(fd, (struct sockaddr *)&addr, listen_addrlen) < 0)
+		rte_exit(EXIT_FAILURE, "bind(%u): %s\n", port, strerror(errno));
+
+	int err = getnameinfo((struct sockaddr *)&listen_addr, listen_addrlen,
+			      host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	if (err != 0)
+		rte_exit(EXIT_FAILURE, "Listen address lookup failed: %s\n",
+			 gai_strerror(err));
+
+	RPCAPD_LOG(NOTICE, "listening on %s port %u", host, listen_port);
+
+	if (!is_loopback(&listen_addr))
+		RPCAPD_LOG(WARNING,
+			"non-loopback address %s; "
+			"rpcap is unauthenticated and unencrypted, captured traffic is exposed to the network",
+			host);
+
+	if (listen(fd, 1) < 0)
+		rte_exit(EXIT_FAILURE, "listen: %s\n", strerror(errno));
+
+	return fd;
+}
+
+static void
+usage(FILE *f, const char *progname)
+{
+	fprintf(f, "Usage: %s [options]\n", progname);
+	fprintf(f,
+		"  -p, --port <port>     listen port (default %u)\n"
+		"  -b, --bind <addr>     bind address (default 127.0.0.1)\n"
+		"  -4                    use only IPv4 (reject IPv6 bind addresses)\n"
+		"  -N <ring size>        ring size in packets (default %u)\n"
+		"  -D, --debug           increase log verbosity (-D info, -DD debug)\n"
+		"      --debug-file <f>  redirect log output to file <f> (append mode)\n"
+		"      --version         print version and exit\n"
+		"  -h, --help            print this help and exit\n"
+		"      --lcore=<core>    CPU core to run on (default: any)\n"
+		"      --file-prefix=<p> prefix to use for multi-process\n"
+		"\n"
+		"WARNING: rpcap is unauthenticated and unencrypted.  Binding to\n"
+		"any non-loopback address exposes captured traffic to the\n"
+		"network.  Sample application; not for production use.\n",
+		RPCAP_DEFAULT_NETPORT, DEFAULT_RING_SIZE);
+}
+
+static void
+print_version(void)
+{
+	printf("rpcapd, a remote packet capture daemon (DPDK pdump backend)\n"
+	       "Built against %s\n", rte_version());
+}
+
+static void
+parse_opts(int argc, char **argv)
+{
+	enum {
+		OPT_LONG_ONLY = 0x100,
+		OPT_DEBUG_FILE,
+		OPT_VERSION,
+	};
+	static const struct option long_options[] = {
+		{ "port",        required_argument, NULL, 'p' },
+		{ "bind",        required_argument, NULL, 'b' },
+		{ "debug",       no_argument,       NULL, 'D' },
+		{ "help",        no_argument,       NULL, 'h' },
+		{ "version",     no_argument,       NULL, OPT_VERSION },
+		{ "debug-file",  required_argument, NULL, OPT_DEBUG_FILE },
+		{ "file-prefix", required_argument, NULL, 0 },
+		{ "lcore",       required_argument, NULL, 0 },
+		{ NULL, 0, NULL, 0 },
+	};
+	int option_index, c;
+
+	while ((c = getopt_long(argc, argv, "hD4p:b:N:",
+				long_options, &option_index)) != -1) {
+		switch (c) {
+		case 'p': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			if (u == 0 || u > UINT16_MAX)
+				rte_exit(EXIT_FAILURE, "Invalid port: %s\n", optarg);
+			listen_port = (uint16_t)u;
+			break;
+		}
+		case 'b':
+			bind_arg = optarg;
+			break;
+		case '4':
+			ipv4_only = true;
+			break;
+		case 'N': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			/* Check the full value before narrowing it: an upper
+			 * bound is needed anyway because rte_align32pow2()
+			 * wraps to zero above 2^31, and that failure would
+			 * otherwise only surface in rte_ring_create() on the
+			 * first capture.
+			 */
+			if (u < 64 || u > MAX_RING_SIZE)
+				rte_exit(EXIT_FAILURE,
+					 "Ring size must be between 64 and %u\n",
+					 MAX_RING_SIZE);
+			ring_size = (uint32_t)u;
+			/* rte_ring_create() requires a power of two. */
+			if (!rte_is_power_of_2(ring_size)) {
+				ring_size = rte_align32pow2(ring_size);
+				RPCAPD_LOG(NOTICE, "ring size rounded up to %u",
+					ring_size);
+			}
+			break;
+		}
+		case 'D':
+			debug_log++;
+			break;
+		case 'h':
+			usage(stdout, argv[0]);
+			exit(0);
+		case OPT_VERSION:
+			print_version();
+			exit(0);
+		case OPT_DEBUG_FILE:
+			debug_file = optarg;
+			break;
+		case 0: {
+			const char *longopt = long_options[option_index].name;
+
+			if (!strcmp(longopt, "lcore")) {
+				lcore_arg = optarg;
+				break;
+			} else if (!strcmp(longopt, "file-prefix")) {
+				file_prefix = optarg;
+				break;
+			}
+		}
+			/* fallthrough */
+		default:
+			usage(stderr, argv[0]);
+			exit(EXIT_FAILURE);
+		}
+	}
+
+	/* Resolve the bind address now that -4 has been seen. */
+	parse_bind_addr(bind_arg ? bind_arg : "127.0.0.1",
+			ipv4_only ? AF_INET : AF_UNSPEC);
+}
+
+/*
+ * Periodic check that the DPDK primary process is still alive.
+ * If it dies our shared-memory state (rings, mempools, pdump) becomes
+ * unsafe to touch, so we set quit_signal and let the main loop tear
+ * down cleanly on its next iteration.  The callback runs on the EAL
+ * interrupt thread; quit_signal is atomic so the read in the main
+ * loop is well-defined.
+ */
+static void
+monitor_primary(void *arg __rte_unused)
+{
+	if (rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed))
+		return;
+
+	if (rte_eal_primary_proc_alive(NULL)) {
+		rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL);
+		return;
+	}
+
+	RPCAPD_LOG(NOTICE, "primary process exited, shutting down");
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+static void
+enable_primary_monitor(void)
+{
+	if (rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL) < 0)
+		RPCAPD_LOG(WARNING, "failed to install primary process monitor");
+}
+
+static void
+disable_primary_monitor(void)
+{
+	rte_eal_alarm_cancel(monitor_primary, NULL);
+}
+
+/*
+ * Bring up EAL as a secondary process so that pdump can attach to a
+ * running primary DPDK application.  Mirrors dumpcap's approach: the
+ * RPCAP user sees a small set of options (port, ring size) rather
+ * than the full DPDK EAL command line.
+ */
+static int
+dpdk_init(void)
+{
+	static const char * const args[] = {
+		"rpcapd",
+		"--proc-type", "secondary",
+		"--log-level", "info",        /* EAL stays quiet */
+	};
+	int eal_argc = RTE_DIM(args);
+	rte_cpuset_t cpuset = { };
+	char **eal_argv;
+	unsigned int i;
+
+	if (file_prefix != NULL)
+		eal_argc += 2;
+
+	if (lcore_arg != NULL)
+		eal_argc += 2;
+
+	eal_argv = calloc(eal_argc + 1, sizeof(char *));
+	if (eal_argv == NULL)
+		return -1;
+
+	for (i = 0; i < RTE_DIM(args); i++) {
+		eal_argv[i] = strdup(args[i]);
+		if (eal_argv[i] == NULL)
+			return -1;
+	}
+
+	if (file_prefix != NULL && *file_prefix != '\0') {
+		eal_argv[i++] = strdup("--file-prefix");
+		eal_argv[i++] = strdup(file_prefix);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+
+	if (lcore_arg != NULL) {
+		eal_argv[i++] = strdup("--lcores");
+		eal_argv[i++] = strdup(lcore_arg);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+	eal_argc = i;
+
+	/*
+	 * Need to get the original cpuset, before EAL init changes
+	 * the affinity of this thread (main lcore).
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_get_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		rte_panic("rte_thread_getaffinity failed\n");
+
+	if (rte_eal_init(eal_argc, eal_argv) < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failed: is the primary process running?\n");
+
+	/*
+	 * If no lcore argument was specified,
+	 * then run this program as a normal process
+	 * which can be scheduled on any non-isolated CPU.
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_set_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		RPCAPD_LOG(INFO, "Can not restore original CPU affinity");
+
+	if (rte_pdump_init() < 0)
+		rte_exit(EXIT_FAILURE, "rte_pdump_init failed\n");
+
+	return 0;
+}
+
+int
+main(int argc, char **argv)
+{
+	struct sigaction action = {
+		.sa_handler = signal_handler,
+	};
+	int srv_fd;
+
+	parse_opts(argc, argv);
+
+	/*
+	 * Redirect log output before EAL init so EAL's own messages are
+	 * captured too.  The FILE handle is intentionally never closed:
+	 * the kernel reclaims it at process exit.
+	 */
+	if (debug_file != NULL) {
+		FILE *fp = fopen(debug_file, "a");
+
+		if (fp == NULL)
+			rte_exit(EXIT_FAILURE, "Cannot open debug file '%s': %s\n",
+				 debug_file, strerror(errno));
+		setvbuf(fp, NULL, _IOLBF, 0);
+		rte_openlog_stream(fp);
+	}
+
+	if (dpdk_init() < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failure\n");
+
+	/* Default to NOTICE: only things the operator needs to see.
+	 * Each -D steps down one level, to INFO then DEBUG.
+	 */
+	rte_log_set_level(RTE_LOGTYPE_RPCAPD,
+			  debug_log >= 2 ? RTE_LOG_DEBUG :
+			  debug_log == 1 ? RTE_LOG_INFO : RTE_LOG_NOTICE);
+
+	if (rte_eth_dev_count_avail() == 0)
+		rte_exit(EXIT_FAILURE, "No Ethernet ports found\n");
+
+	sigaction(SIGTERM, &action, NULL);
+	sigaction(SIGINT, &action, NULL);
+
+	/* If peer closes, this detected in next recv() */
+	signal(SIGPIPE, SIG_IGN);
+
+	srv_fd = open_listen_socket(listen_port);
+
+	enable_primary_monitor();
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int cfd = accept_timeout(srv_fd, -1);
+
+		if (cfd < 0) {
+			if (errno == EINTR)
+				continue;
+			break;
+		}
+		handle_client(cfd);
+	}
+
+	disable_primary_monitor();
+	RPCAPD_LOG(NOTICE, "shutting down");
+	close(srv_fd);
+	rte_pdump_uninit();
+	return rte_eal_cleanup() ? EXIT_FAILURE : 0;
+}
diff --git a/examples/rpcapd/meson.build b/examples/rpcapd/meson.build
new file mode 100644
index 0000000000..320b262666
--- /dev/null
+++ b/examples/rpcapd/meson.build
@@ -0,0 +1,19 @@
+# SPDX-License-Identifier: BSD-3-Clause
+# Copyright(c) 2026 Stephen Hemminger
+
+# since it relies on primary/secondary process
+# this example is Linux only
+if not is_linux
+    build = false
+    subdir_done()
+endif
+
+if not dpdk_conf.has('RTE_HAS_LIBPCAP')
+    build = false
+    reason = 'missing dependency, "libpcap"'
+    subdir_done()
+endif
+
+sources = files('main.c')
+ext_deps += pcap_dep
+deps += ['ethdev', 'pdump', 'bpf']
diff --git a/examples/rpcapd/rpcap-protocol.h b/examples/rpcapd/rpcap-protocol.h
new file mode 100644
index 0000000000..b381271de7
--- /dev/null
+++ b/examples/rpcapd/rpcap-protocol.h
@@ -0,0 +1,127 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * On-the-wire RPCAP protocol definitions, transcribed from libpcap's
+ * rpcap-protocol.h.  See:
+ *   https://github.com/the-tcpdump-group/libpcap/blob/master/rpcap-protocol.h
+ *
+ * Only the subset needed by the DPDK rpcapd example is included here.  All
+ * multi-byte fields in the structures below are big-endian on the wire.
+ */
+
+#ifndef _RPCAP_PROTOCOL_H_
+#define _RPCAP_PROTOCOL_H_
+
+#include <stdint.h>
+
+#define RPCAP_VERSION              0
+#define RPCAP_DEFAULT_NETPORT      2002
+
+/* Message types */
+#define RPCAP_MSG_ERROR            0x01
+#define RPCAP_MSG_FINDALLIF_REQ    0x02
+#define RPCAP_MSG_OPEN_REQ         0x03
+#define RPCAP_MSG_STARTCAP_REQ     0x04
+#define RPCAP_MSG_UPDATEFILTER_REQ 0x05
+#define RPCAP_MSG_CLOSE            0x06
+#define RPCAP_MSG_PACKET           0x07
+#define RPCAP_MSG_AUTH_REQ         0x08
+#define RPCAP_MSG_STATS_REQ        0x09
+#define RPCAP_MSG_ENDCAP_REQ       0x0a
+#define RPCAP_MSG_IS_REPLY         0x80
+
+#define RPCAP_MSG_FINDALLIF_REPLY    (RPCAP_MSG_FINDALLIF_REQ    | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_OPEN_REPLY         (RPCAP_MSG_OPEN_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STARTCAP_REPLY     (RPCAP_MSG_STARTCAP_REQ     | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_UPDATEFILTER_REPLY (RPCAP_MSG_UPDATEFILTER_REQ | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_AUTH_REPLY         (RPCAP_MSG_AUTH_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_ENDCAP_REPLY       (RPCAP_MSG_ENDCAP_REQ       | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STATS_REPLY	     (RPCAP_MSG_STATS_REQ	 | RPCAP_MSG_IS_REPLY)
+
+/* Error codes carried in the 'value' field of RPCAP_MSG_ERROR */
+#define PCAP_ERR_WRONGVER          17
+
+/* Filter encoding: the filter is a BPF/NPF program */
+#define RPCAP_UPDATEFILTER_BPF     1
+
+/* Flags in rpcap_startcapreq.flags */
+#define RPCAP_STARTCAPREQ_FLAG_PROMISC     0x00000001	/* promiscuous mode */
+#define RPCAP_STARTCAPREQ_FLAG_DGRAM       0x00000002	/* use UDP for data */
+#define RPCAP_STARTCAPREQ_FLAG_SERVEROPEN  0x00000004	/* server connects out */
+#define RPCAP_STARTCAPREQ_FLAG_INBOUND     0x00000008	/* capture inbound only */
+#define RPCAP_STARTCAPREQ_FLAG_OUTBOUND    0x00000010	/* capture outbound only */
+
+/* Subset of pcap interface flags (pcap.h) */
+#define PCAP_IF_UP                 0x00000002
+#define PCAP_IF_RUNNING            0x00000004
+
+/* DLT_EN10MB - ethernet, the only link type we report */
+#define DLT_EN10MB                 1
+
+struct rpcap_header {
+	uint8_t  ver;
+	uint8_t  type;
+	uint16_t value;
+	uint32_t plen;
+};
+
+struct rpcap_findalldevs_if {
+	uint16_t namelen;
+	uint16_t desclen;
+	uint32_t flags;
+	uint16_t naddr;
+	uint16_t dummy;
+};
+
+struct rpcap_openreply {
+	int32_t  linktype;
+	int32_t  tzoff;
+};
+
+struct rpcap_startcapreq {
+	uint32_t snaplen;
+	uint32_t read_timeout;
+	uint16_t flags;
+	uint16_t portdata;
+};
+
+struct rpcap_startcapreply {
+	int32_t  bufsize;
+	uint16_t portdata;
+	uint16_t dummy;
+};
+
+/*
+ * A filter, sent either after rpcap_startcapreq or in an
+ * RPCAP_MSG_UPDATEFILTER_REQ, followed by nitems instructions.
+ */
+struct rpcap_filter {
+	uint16_t filtertype;
+	uint16_t dummy;
+	uint32_t nitems;
+};
+
+/* One cBPF instruction, repeated nitems times after rpcap_filter. */
+struct rpcap_filterbpf_insn {
+	uint16_t code;
+	uint8_t  jt;
+	uint8_t  jf;
+	int32_t  k;
+};
+
+struct rpcap_stats {
+	uint32_t ifrecv;
+	uint32_t ifdrop;
+	uint32_t krnldrop;
+	uint32_t svrcapt;
+};
+
+struct rpcap_pkthdr {
+	uint32_t timestamp_sec;
+	uint32_t timestamp_usec;
+	uint32_t caplen;
+	uint32_t len;
+	uint32_t npkt;
+};
+
+#endif /* _RPCAP_PROTOCOL_H_ */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH v3] examples/rpcapd: demo version of packet capture daemon
  2026-09-22 21:31 ` [PATCH v3] " Stephen Hemminger
@ 2026-09-28 16:18   ` Marat Khalili
  2026-09-28 17:24     ` Stephen Hemminger
  0 siblings, 1 reply; 21+ messages in thread
From: Marat Khalili @ 2026-09-28 16:18 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: Thomas Monjalon, Reshma Pattan, dev

Summarizing the exchange to v2, our own copy of rpcap-protocol.h is 
unavoidable here since one included in libpcap is an internal header 
that is not installed by design (irrespective of package maintainers). I 
am ready to ack the patch after this is clarified in the code to prevent 
other people from being similarly confused since situation with the 
dependencies here is quite complex.

My request to better mark sections of main.c responsible for different 
concerns still stands, but is not a show stopper. Same for my request to 
make it a part of dpdk-testpmd instead of a secondary process. Anyone 
interested can probably do either or both locally in half an hour in a 
code agent anyway.


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3] examples/rpcapd: demo version of packet capture daemon
  2026-09-28 16:18   ` Marat Khalili
@ 2026-09-28 17:24     ` Stephen Hemminger
  0 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-09-28 17:24 UTC (permalink / raw)
  To: Marat Khalili; +Cc: Thomas Monjalon, Reshma Pattan, dev

On Mon, 28 Sep 2026 17:18:40 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> Summarizing the exchange to v2, our own copy of rpcap-protocol.h is 
> unavoidable here since one included in libpcap is an internal header 
> that is not installed by design (irrespective of package maintainers). I 
> am ready to ack the patch after this is clarified in the code to prevent 
> other people from being similarly confused since situation with the 
> dependencies here is quite complex.
> 
> My request to better mark sections of main.c responsible for different 
> concerns still stands, but is not a show stopper. Same for my request to 
> make it a part of dpdk-testpmd instead of a secondary process. Anyone 
> interested can probably do either or both locally in half an hour in a 
> code agent anyway.
> 

I have a new version which has more features and splits up the code.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* [PATCH v4 0/4] add rpcap remote capture daemon
  2026-09-08 21:07 [PATCH] examples/rpcapd: demo version of packet capture daemon Stephen Hemminger
  2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
  2026-09-22 21:31 ` [PATCH v3] " Stephen Hemminger
@ 2026-10-01  2:35 ` Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 1/4] pcapng: add API to read back capture mbuf header Stephen Hemminger
                     ` (4 more replies)
  2 siblings, 5 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01  2:35 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger

This series adds support for live capture using tcpdump and related
tools that use libpcap. See document doc/guides/tools/rpcapd.rst
for the details.

The dpdkr rpcap daemon runs as a secondary process similar to dpdk-pdump
and dpdk-dumpcap. The difference is instead of writing packets
directly to a file, it provides a standard API for transferring
packet capture over a socket. The command line arguments
are chosen to match the existing libpcap rpcapd, but not
all command flags are implemented.

The capture is only possible if the rpcap daemon is running.
And since it is a secondary process, it needs to be started
after the DPDK primary process. The daemon will also exit when
the primary process exits.

Since it is intended for debugging (and due to limitations in
the pdump API), only a single client can access rpcap at a time.

The dpdk-rpcapd uses existing pdump infrastructure to capture
and filter packets. One small additional API was needed to extract
the timestamp and original capture length from the mbuf
that was processed with pcapng mode. Since rpcap protocol
is limited to pcap legacy format, only microsecond timestamps
(and no metadata info) are visible in the stream.

By default dpdk-rpcapd can only be used over local TCP sockets.
This version also adds TLS and password authentication which
could be useful if dpdk-rpcapd is setup to allow non-local access.
An additional security feature was added based on what libpcap rpcapd has,
with the -l flag a list of allowed hosts can be used.

Changes since previous version:

  - Moved from examples/ to app/
  - Added TLS and password authentication (patch 3).
  - Added the -l host list option (patch 4).

Stephen Hemminger (4):
  pcapng: add API to read back capture mbuf header
  app/rpcapd: remote pcap daemon
  app/rpcapd: add TLS support
  app/rpcapd: add host list option

 MAINTAINERS                            |   2 +
 app/meson.build                        |   1 +
 app/rpcapd/capture.c                   | 541 ++++++++++++++++
 app/rpcapd/filter.c                    | 164 +++++
 app/rpcapd/main.c                      | 840 +++++++++++++++++++++++++
 app/rpcapd/meson.build                 |  40 ++
 app/rpcapd/rpcap-protocol.h            | 146 +++++
 app/rpcapd/rpcapd.h                    | 122 ++++
 app/rpcapd/session.c                   | 292 +++++++++
 app/rpcapd/sock.c                      | 365 +++++++++++
 app/rpcapd/tls.c                       | 273 ++++++++
 app/test/test_pcapng.c                 | 133 ++++
 doc/guides/rel_notes/release_26_11.rst |  13 +
 doc/guides/tools/index.rst             |   1 +
 doc/guides/tools/rpcapd.rst            | 288 +++++++++
 lib/pcapng/rte_pcapng.c                |  40 ++
 lib/pcapng/rte_pcapng.h                |  46 ++
 17 files changed, 3307 insertions(+)
 create mode 100644 app/rpcapd/capture.c
 create mode 100644 app/rpcapd/filter.c
 create mode 100644 app/rpcapd/main.c
 create mode 100644 app/rpcapd/meson.build
 create mode 100644 app/rpcapd/rpcap-protocol.h
 create mode 100644 app/rpcapd/rpcapd.h
 create mode 100644 app/rpcapd/session.c
 create mode 100644 app/rpcapd/sock.c
 create mode 100644 app/rpcapd/tls.c
 create mode 100644 doc/guides/tools/rpcapd.rst

-- 
2.53.0


^ permalink raw reply	[flat|nested] 21+ messages in thread

* [PATCH v4 1/4] pcapng: add API to read back capture mbuf header
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
@ 2026-10-01  2:35   ` Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 2/4] app/rpcapd: remote pcap daemon Stephen Hemminger
                     ` (3 subsequent siblings)
  4 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01  2:35 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Reshma Pattan

An mbuf from rte_pcapng_copy() starts with an enhanced packet block
holding the capture time, the length before truncation and the port.
The only way to get at that was to write the mbuf to a file, which
is no use to something forwarding captured packets elsewhere.

Add rte_pcapng_pkt_info() to decode that header in place, and check
it is well formed before trusting the lengths in it.

The capture time is reported as the raw TSC value: there is no
capture file here to take a reference point from, so converting it
to a time of day is left to the caller.

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---
 app/test/test_pcapng.c                 | 133 +++++++++++++++++++++++++
 doc/guides/rel_notes/release_26_11.rst |   5 +
 lib/pcapng/rte_pcapng.c                |  40 ++++++++
 lib/pcapng/rte_pcapng.h                |  46 +++++++++
 4 files changed, 224 insertions(+)

diff --git a/app/test/test_pcapng.c b/app/test/test_pcapng.c
index d14ea84f0d..cb36ea1d54 100644
--- a/app/test/test_pcapng.c
+++ b/app/test/test_pcapng.c
@@ -672,6 +672,138 @@ test_write_before_open(void)
 	return -1;
 }
 
+/*
+ * Check that rte_pcapng_pkt_info() reads back what rte_pcapng_copy()
+ * recorded.  The length before truncation and the capture time are
+ * only in the block header, so this is the only way a consumer that
+ * does not write a file can get at them.
+ */
+static int
+test_pkt_info(void)
+{
+	struct dummy_mbuf mbfs;
+	struct rte_mbuf *mc;
+	struct rte_pcapng_pkt pkt;
+	uint32_t pkt_len, snaplen, saved;
+	uint64_t before, after;
+	const uint8_t *data;
+	int ret;
+
+	mbuf1_prepare(&mbfs);
+	mbuf1_resize(&mbfs, 512);
+	pkt_len = rte_pktmbuf_pkt_len(&mbfs.mb[0]);
+
+	/* An untruncated copy reports the length it came in with. */
+	before = rte_get_tsc_cycles();
+	mc = rte_pcapng_copy(port_id, 0, &mbfs.mb[0], mp, pkt_len,
+			     RTE_PCAPNG_DIRECTION_IN, NULL);
+	TEST_ASSERT(mc != NULL, "rte_pcapng_copy failed");
+	after = rte_get_tsc_cycles();
+
+	ret = rte_pcapng_pkt_info(mc, &pkt);
+	TEST_ASSERT(ret == 0, "rte_pcapng_pkt_info failed: %d", ret);
+
+	TEST_ASSERT(pkt.original_len == pkt_len,
+		    "original_len is %u, expected %u", pkt.original_len, pkt_len);
+	TEST_ASSERT(pkt.captured_len == pkt_len,
+		    "captured_len is %u, expected %u", pkt.captured_len, pkt_len);
+	TEST_ASSERT(pkt.port == port_id,
+		    "port is %u, expected %u", pkt.port, port_id);
+
+	/* The copy was made between the two readings, so the recorded
+	 * cycle count has to fall between them.
+	 */
+	TEST_ASSERT(pkt.cycles >= before && pkt.cycles <= after,
+		    "cycles %"PRIu64" is outside [%"PRIu64", %"PRIu64"]",
+		    pkt.cycles, before, after);
+
+	/* data_offset points at the packet itself, not the block header. */
+	data = rte_pktmbuf_mtod_offset(mc, const uint8_t *, pkt.data_offset);
+	TEST_ASSERT(memcmp(data, rte_pktmbuf_mtod(&mbfs.mb[0], const void *),
+			   rte_pktmbuf_data_len(&mbfs.mb[0])) == 0,
+		    "packet data is not at data_offset");
+
+	/* A corrupt block is rejected rather than believed. */
+	{
+		struct pcapng_test_epb {
+			uint32_t block_type;
+			uint32_t block_length;
+		} *epb = rte_pktmbuf_mtod(mc, struct pcapng_test_epb *);
+
+		saved = epb->block_type;
+		epb->block_type = ~saved;
+		TEST_ASSERT(rte_pcapng_pkt_info(mc, &pkt) == -EINVAL,
+			    "bad block_type was accepted");
+		epb->block_type = saved;
+
+		saved = epb->block_length;
+		epb->block_length = saved + 1;
+		TEST_ASSERT(rte_pcapng_pkt_info(mc, &pkt) == -EINVAL,
+			    "bad block_length was accepted");
+		epb->block_length = saved;
+
+		/* and is fine again once put back */
+		TEST_ASSERT(rte_pcapng_pkt_info(mc, &pkt) == 0,
+			    "restored block was rejected");
+	}
+
+	TEST_ASSERT(rte_pcapng_pkt_info(NULL, &pkt) == -EINVAL,
+		    "NULL mbuf was accepted");
+	TEST_ASSERT(rte_pcapng_pkt_info(mc, NULL) == -EINVAL,
+		    "NULL result was accepted");
+
+	rte_pktmbuf_free(mc);
+
+	/*
+	 * Truncated copy.  This is the case that cannot be recovered
+	 * from the mbuf alone: captured_len shrinks to the snaplen
+	 * while original_len still describes the packet on the wire.
+	 */
+	snaplen = pkt_len / 2;
+	mc = rte_pcapng_copy(port_id, 0, &mbfs.mb[0], mp, snaplen,
+			     RTE_PCAPNG_DIRECTION_IN, NULL);
+	TEST_ASSERT(mc != NULL, "truncated rte_pcapng_copy failed");
+
+	ret = rte_pcapng_pkt_info(mc, &pkt);
+	TEST_ASSERT(ret == 0, "rte_pcapng_pkt_info failed on truncated: %d", ret);
+
+	TEST_ASSERT(pkt.captured_len == snaplen,
+		    "captured_len is %u, expected %u", pkt.captured_len, snaplen);
+	TEST_ASSERT(pkt.original_len == pkt_len,
+		    "original_len is %u, expected %u, truncation lost it",
+		    pkt.original_len, pkt_len);
+
+	rte_pktmbuf_free(mc);
+
+	/*
+	 * A stripped VLAN tag is put back by the copy, but is not
+	 * counted in the length reported by the hardware.  So the
+	 * captured packet is larger than the original, and a consumer
+	 * has to cope with that rather than assume it cannot happen.
+	 */
+	mbfs.mb[0].ol_flags |= RTE_MBUF_F_RX_VLAN_STRIPPED;
+	mbfs.mb[0].vlan_tci = 42;
+
+	mc = rte_pcapng_copy(port_id, 0, &mbfs.mb[0], mp, pkt_len,
+			     RTE_PCAPNG_DIRECTION_IN, NULL);
+	TEST_ASSERT(mc != NULL, "VLAN rte_pcapng_copy failed");
+
+	ret = rte_pcapng_pkt_info(mc, &pkt);
+	TEST_ASSERT(ret == 0, "rte_pcapng_pkt_info failed on VLAN: %d", ret);
+
+	TEST_ASSERT(pkt.captured_len == pkt_len + sizeof(struct rte_vlan_hdr),
+		    "captured_len is %u, expected %zu with the tag restored",
+		    pkt.captured_len, pkt_len + sizeof(struct rte_vlan_hdr));
+	TEST_ASSERT(pkt.original_len == pkt_len,
+		    "original_len is %u, expected %u", pkt.original_len, pkt_len);
+	TEST_ASSERT(pkt.captured_len > pkt.original_len,
+		    "restored VLAN tag did not make the capture longer");
+
+	rte_pktmbuf_free(mc);
+
+	return 0;
+}
+
 static void
 test_cleanup(void)
 {
@@ -688,6 +820,7 @@ unit_test_suite test_pcapng_suite  = {
 		TEST_CASE(test_add_interface),
 		TEST_CASE(test_write_packets),
 		TEST_CASE(test_write_before_open),
+		TEST_CASE(test_pkt_info),
 		TEST_CASES_END()
 	}
 };
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index e027c7a27f..5b5a9f006e 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
      Also, make sure to start the actual text at the margin.
      =======================================================
 
+* **Added pcapng API to read back a captured packet header.**
+
+  Added the experimental ``rte_pcapng_pkt_info()`` function to read back what
+  ``rte_pcapng_copy()`` records in a captured packet.
+
 * **Added API to get CPU socket ID.**
 
   Added the experimental ``rte_cpu_socket_id()`` function
diff --git a/lib/pcapng/rte_pcapng.c b/lib/pcapng/rte_pcapng.c
index b5d1026891..54a0aa1342 100644
--- a/lib/pcapng/rte_pcapng.c
+++ b/lib/pcapng/rte_pcapng.c
@@ -707,6 +707,46 @@ rte_pcapng_copy(uint16_t port_id, uint32_t queue,
 	return NULL;
 }
 
+/* Read back the block header put there by rte_pcapng_copy() */
+RTE_EXPORT_EXPERIMENTAL_SYMBOL(rte_pcapng_pkt_info, 26.11)
+int
+rte_pcapng_pkt_info(const struct rte_mbuf *m, struct rte_pcapng_pkt *pkt)
+{
+	const struct pcapng_enhance_packet_block *epb;
+	struct pcapng_enhance_packet_block ebuf;
+
+	if (unlikely(m == NULL || pkt == NULL))
+		return -EINVAL;
+
+	epb = rte_pktmbuf_read(m, 0, sizeof(*epb), &ebuf);
+	if (unlikely(epb == NULL))
+		return -EINVAL;
+
+	if (unlikely(epb->block_type != PCAPNG_ENHANCED_PACKET_BLOCK))
+		return -EINVAL;
+
+	/*
+	 * rte_pcapng_copy() sets block_length to the whole mbuf length, and
+	 * the packet data has to fit in what is left after the header.
+	 */
+	if (unlikely(epb->block_length != rte_pktmbuf_pkt_len(m)))
+		return -EINVAL;
+
+	if (unlikely(epb->capture_length >
+		     epb->block_length - sizeof(*epb)))
+		return -EINVAL;
+
+	pkt->cycles = (uint64_t)epb->timestamp_hi << 32;
+	pkt->cycles += epb->timestamp_lo;
+
+	pkt->captured_len = epb->capture_length;
+	pkt->original_len = epb->original_length;
+	pkt->data_offset = sizeof(*epb);
+	pkt->port = m->port;
+
+	return 0;
+}
+
 /* Write pre-formatted packets to file. */
 RTE_EXPORT_SYMBOL(rte_pcapng_write_packets)
 ssize_t
diff --git a/lib/pcapng/rte_pcapng.h b/lib/pcapng/rte_pcapng.h
index d8d328f710..055075e921 100644
--- a/lib/pcapng/rte_pcapng.h
+++ b/lib/pcapng/rte_pcapng.h
@@ -22,6 +22,8 @@
 #include <stdint.h>
 #include <sys/types.h>
 
+#include <rte_compat.h>
+#include <rte_mbuf.h>
 #include <rte_mempool.h>
 
 #ifdef __cplusplus
@@ -140,6 +142,50 @@ rte_pcapng_copy(uint16_t port_id, uint32_t queue,
 		uint32_t length,
 		enum rte_pcapng_direction direction, const char *comment);
 
+/**
+ * Decoded header of an mbuf produced by rte_pcapng_copy().
+ *
+ * @warning
+ * @b EXPERIMENTAL: this structure may change without prior notice.
+ */
+struct rte_pcapng_pkt {
+	uint64_t cycles;	/**< TSC value when the packet was captured */
+	uint32_t captured_len;	/**< bytes of packet data present */
+	uint32_t original_len;	/**< length of the packet on the wire */
+	uint32_t data_offset;	/**< offset of packet data in the mbuf */
+	uint16_t port;		/**< port recorded by rte_pcapng_copy() */
+};
+
+/**
+ * Extract info from mbuf created by rte_pcapng_copy().
+ *
+ * @warning
+ * @b EXPERIMENTAL: this API may change without prior notice.
+ *
+ * Only valid for packets created by rte_pcapng_copy().
+ * The mbuf is not modified.
+ * To reach the packet data, read *captured_len* bytes starting at *data_offset*.
+ *
+ * The capture time is reported as the raw TSC value recorded by
+ * rte_pcapng_copy(), since this has no capture file to take a reference
+ * point from.  To turn it into a time of day, sample rte_get_tsc_cycles()
+ * and the system clock together once, then scale the difference by
+ * rte_get_tsc_hz().
+ *
+ * @param m
+ *   An mbuf returned by rte_pcapng_copy().
+ * @param pkt
+ *   Filled in on success.
+ * @return
+ *   0 on success, -EINVAL if the mbuf is not a well formed enhanced
+ *   packet block.
+ *
+ * @note
+ *   Length may vary from the original because rte_pcapng_copy() inserts VLAN.
+ */
+__rte_experimental
+int
+rte_pcapng_pkt_info(const struct rte_mbuf *m, struct rte_pcapng_pkt *pkt);
 
 /**
  * Determine optimum mbuf data size.
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v4 2/4] app/rpcapd: remote pcap daemon
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 1/4] pcapng: add API to read back capture mbuf header Stephen Hemminger
@ 2026-10-01  2:35   ` Stephen Hemminger
  2026-10-01 18:50     ` Marat Khalili
  2026-10-01  2:35   ` [PATCH v4 3/4] app/rpcapd: add TLS support Stephen Hemminger
                     ` (2 subsequent siblings)
  4 siblings, 1 reply; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01  2:35 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Thomas Monjalon, Reshma Pattan

Add RPCAP support over localhost TCP integrated with DPDK.
It runs as a secondary process that allows connections from
tools using tcpdump's defacto protocol rpcap.

See: doc/guides/tools/rpcapd.rst for more info

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---
 MAINTAINERS                            |   2 +
 app/meson.build                        |   1 +
 app/rpcapd/capture.c                   | 526 ++++++++++++++++++++++
 app/rpcapd/filter.c                    | 164 +++++++
 app/rpcapd/main.c                      | 577 +++++++++++++++++++++++++
 app/rpcapd/meson.build                 |  25 ++
 app/rpcapd/rpcap-protocol.h            | 142 ++++++
 app/rpcapd/rpcapd.h                    | 105 +++++
 app/rpcapd/session.c                   | 136 ++++++
 app/rpcapd/sock.c                      | 315 ++++++++++++++
 doc/guides/rel_notes/release_26_11.rst |   5 +
 doc/guides/tools/index.rst             |   1 +
 doc/guides/tools/rpcapd.rst            | 199 +++++++++
 13 files changed, 2198 insertions(+)
 create mode 100644 app/rpcapd/capture.c
 create mode 100644 app/rpcapd/filter.c
 create mode 100644 app/rpcapd/main.c
 create mode 100644 app/rpcapd/meson.build
 create mode 100644 app/rpcapd/rpcap-protocol.h
 create mode 100644 app/rpcapd/rpcapd.h
 create mode 100644 app/rpcapd/session.c
 create mode 100644 app/rpcapd/sock.c
 create mode 100644 doc/guides/tools/rpcapd.rst

diff --git a/MAINTAINERS b/MAINTAINERS
index 482bc7df76..fd56b6b440 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -1722,6 +1722,8 @@ F: app/pdump/
 F: doc/guides/tools/pdump.rst
 F: app/dumpcap/
 F: doc/guides/tools/dumpcap.rst
+F: app/rpcapd/
+F: doc/guides/tools/rpcapd.rst
 
 
 Packet Framework
diff --git a/app/meson.build b/app/meson.build
index 4515688471..b9227f1fe4 100644
--- a/app/meson.build
+++ b/app/meson.build
@@ -17,6 +17,7 @@ apps = [
         'graph',
         'pdump',
         'proc-info',
+        'rpcapd',
         'test-acl',
         'test-bbdev',
         'test-cmdline',
diff --git a/app/rpcapd/capture.c b/app/rpcapd/capture.c
new file mode 100644
index 0000000000..17fce3fbca
--- /dev/null
+++ b/app/rpcapd/capture.c
@@ -0,0 +1,526 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Starting and stopping a capture, and streaming the captured packets
+ * to the client over the data connection.
+ */
+
+#include <errno.h>
+#include <poll.h>
+#include <stdio.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/time.h>
+#include <sys/uio.h>
+#include <time.h>
+#include <unistd.h>
+
+#include <rte_byteorder.h>
+#include <rte_common.h>
+#include <rte_cycles.h>
+#include <rte_errno.h>
+#include <rte_ethdev.h>
+#include <rte_ether.h>
+#include <rte_malloc.h>
+#include <rte_mbuf.h>
+#include <rte_mempool.h>
+#include <rte_pcapng.h>
+#include <rte_pdump.h>
+#include <rte_ring.h>
+#include <rte_stdatomic.h>
+#include <rte_time.h>
+
+#include "rpcap-protocol.h"
+#include "rpcapd.h"
+
+#define BURST_SIZE                    32
+#define MBUF_CACHE_SIZE               32
+#define SLEEP_THRESHOLD		      100
+#define SLEEP_US		      100
+#define DATA_ACCEPT_TIMEOUT_MS        10000
+
+/* Reference point for converting a captured TSC to a time of day.
+ * The TSC is the same counter in the primary that did the capture.
+ */
+static uint64_t tsc_base;
+static uint64_t ns_base;
+
+void
+timestamp_init(void)
+{
+	struct timespec ts;
+	uint64_t cycles;
+
+	cycles = rte_get_tsc_cycles();
+	clock_gettime(CLOCK_REALTIME, &ts);
+	ns_base = rte_timespec_to_ns(&ts);
+	tsc_base = (cycles + rte_get_tsc_cycles()) / 2;
+}
+
+/* Convert a captured TSC to nanoseconds since the Unix epoch.  Whole
+ * seconds come out first so scaling the remainder cannot overflow, and
+ * a packet copied before startup is behind the reference point.
+ */
+static uint64_t
+timestamp_to_ns(uint64_t cycles)
+{
+	const uint64_t hz = rte_get_tsc_hz();
+	uint64_t delta, secs, rem;
+	bool before;
+
+	before = cycles < tsc_base;
+	delta = before ? tsc_base - cycles : cycles - tsc_base;
+
+	secs = delta / hz;
+	rem = delta % hz;
+	delta = secs * NSEC_PER_SEC + (rem * NSEC_PER_SEC) / hz;
+
+	return before ? ns_base - delta : ns_base + delta;
+}
+
+
+/* Open an ephemeral TCP listening socket; return fd, set *port_out. */
+static int
+open_data_listener(uint16_t *port_out)
+{
+	struct sockaddr_storage addr = listen_addr;
+	socklen_t alen;
+	int fd;
+
+	set_sockaddr_port(&addr, 0);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0) {
+		RPCAPD_LOG(ERR, "data socket: %s", strerror(errno));
+		return -1;
+	}
+
+	alen = listen_addrlen;
+	if (bind(fd, (struct sockaddr *)&addr, alen) < 0 ||
+	    listen(fd, 1) < 0 ||
+	    getsockname(fd, (struct sockaddr *)&addr, &alen) < 0) {
+		RPCAPD_LOG(ERR, "data port bind/listen: %s", strerror(errno));
+		close(fd);
+		return -1;
+	}
+	*port_out = get_sockaddr_port(&addr);
+	return fd;
+}
+
+static struct rte_ring *
+create_capture_ring(uint16_t port)
+{
+	char name[RTE_RING_NAMESIZE];
+
+	snprintf(name, sizeof(name), "rpcapd_r_%u_%d", port, getpid());
+	return rte_ring_create(name, ring_size, rte_socket_id(), 0);
+}
+
+static struct rte_mempool *
+create_capture_mempool(uint16_t port, uint32_t snaplen)
+{
+	char name[RTE_MEMPOOL_NAMESIZE];
+	/* Leaves room for the pcapng block header, the options and the
+	 * trailer, as well as the packet itself.
+	 */
+	uint32_t mbuf_size = rte_pcapng_mbuf_size(snaplen);
+
+	snprintf(name, sizeof(name), "rpcapd_p_%u_%d", port, getpid());
+	return rte_pktmbuf_pool_create(name, ring_size * 2, MBUF_CACHE_SIZE, 0,
+				       mbuf_size, rte_socket_id());
+}
+
+
+/* Tear down anything that handle_startcap brought up.
+ * Safe to call after partial setup as well as after a successful capture.
+ */
+void
+stop_capture(struct session *s)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int n;
+
+	if (s->capture_on) {
+		rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+		RPCAPD_LOG(NOTICE, "capture stopped on %s (%u packets)",
+			s->name, s->npkt);
+	}
+	s->capture_on = false;
+
+	if (s->promisc_set) {
+		rte_eth_promiscuous_disable(s->port);
+		s->promisc_set = false;
+	}
+
+	if (s->ring != NULL) {
+		while ((n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts,
+						      BURST_SIZE, NULL)) > 0)
+			rte_pktmbuf_free_bulk(pkts, n);
+		rte_ring_free(s->ring);
+		s->ring = NULL;
+	}
+	if (s->mp != NULL) {
+		rte_mempool_free(s->mp);
+		s->mp = NULL;
+	}
+
+	/* Only safe once pdump is disabled */
+	rte_free(s->prm);
+	s->prm = NULL;
+	if (s->data.fd >= 0) {
+		close(s->data.fd);
+		s->data.fd = -1;
+	}
+}
+
+/*
+ * STARTCAP_REQ: open the data connection and arm the pdump callback.
+ * We use passive mode with the server-allocated data port:
+ *   - the server picks an ephemeral port and listens on it
+ *   - the server returns that port in startcapreply.portdata
+ *   - the client connects back to that port for the packet stream
+ */
+int
+handle_startcap(const struct conn *c, uint32_t plen, struct session *s)
+{
+	struct rpcap_startcapreq req;
+	uint16_t data_port;
+	uint16_t flags;
+	struct rte_bpf_prm *recorded;
+	int data_listen;
+	int data_fd;
+	int ret;
+
+	/* Keep a filter set before the capture started, drop one from a
+	 * capture being restarted: this request brings its own.
+	 */
+	recorded = s->capture_on ? NULL : s->prm;
+	if (recorded != NULL)
+		s->prm = NULL;
+	stop_capture(s);
+	s->prm = recorded;
+
+	if (!s->opened) {
+		rpcap_discard(c, plen);
+		return rpcap_send_error(c, 0, "no interface open");
+	}
+
+	if (plen < sizeof(req)) {
+		rpcap_discard(c, plen);
+		return rpcap_send_error(c, 0, "short startcap request");
+	}
+	if (recv_full(c, &req, sizeof(req)) < 0)
+		return -1;
+
+	flags = rte_be_to_cpu_16(req.flags);
+	if (flags & RPCAP_STARTCAPREQ_FLAG_DGRAM) {
+		rpcap_discard(c, plen - sizeof(req));
+		return rpcap_send_error(c, 0, "UDP data transfer not supported");
+	}
+
+	ret = read_filter(c, plen - sizeof(req), s);
+	if (ret != 0)
+		return ret < 0 ? -1 : 0;	/* error already reported to client */
+
+	/* Direction flags map onto pdump's RX/TX selection; neither (or both)
+	 * means capture in both directions.
+	 */
+	s->pdump_flags = RTE_PDUMP_FLAG_RXTX;
+	if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+		      RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+	    RPCAP_STARTCAPREQ_FLAG_INBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_RX;
+	else if ((flags & (RPCAP_STARTCAPREQ_FLAG_INBOUND |
+			   RPCAP_STARTCAPREQ_FLAG_OUTBOUND)) ==
+		 RPCAP_STARTCAPREQ_FLAG_OUTBOUND)
+		s->pdump_flags = RTE_PDUMP_FLAG_TX;
+
+	s->snaplen = rte_be_to_cpu_32(req.snaplen);
+	if (s->snaplen == 0 || s->snaplen > DEFAULT_SNAPLEN)
+		s->snaplen = DEFAULT_SNAPLEN;
+
+	s->ring = create_capture_ring(s->port);
+	s->mp = create_capture_mempool(s->port, s->snaplen);
+	if (s->ring == NULL || s->mp == NULL) {
+		RPCAPD_LOG(ERR, "ring/mempool alloc failed: %s",
+			rte_strerror(rte_errno));
+		stop_capture(s);
+		return rpcap_send_error(c, 0, "DPDK alloc failed");
+	}
+
+	data_listen = open_data_listener(&data_port);
+	if (data_listen < 0) {
+		stop_capture(s);
+		return rpcap_send_error(c, 0, "data port setup failed");
+	}
+
+	/* Leave the port alone if it is already promiscuous: it belongs to
+	 * the primary process, and stop_capture() must not turn off
+	 * something this daemon did not turn on.
+	 */
+	if ((flags & RPCAP_STARTCAPREQ_FLAG_PROMISC) &&
+	    rte_eth_promiscuous_get(s->port) != 1) {
+		if (rte_eth_promiscuous_enable(s->port) == 0)
+			s->promisc_set = true;
+		else
+			RPCAPD_LOG(NOTICE, "cannot enable promiscuous mode on %s",
+				s->name);
+	}
+
+	/* Setup packet capture callbacks. */
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES,
+				 s->pdump_flags | RTE_PDUMP_FLAG_PCAPNG,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		close(data_listen);
+		stop_capture(s);
+		return rpcap_send_error(c, 0, "cannot enable capture");
+	}
+	s->capture_on = true;
+	s->npkt = 0;
+
+	struct rpcap_startcapreply reply = {
+		.bufsize = rte_cpu_to_be_32(s->snaplen * BURST_SIZE),
+		.portdata = rte_cpu_to_be_16(data_port),
+	};
+	if (rpcap_send_msg(c, RPCAP_MSG_STARTCAP_REPLY, 0, &reply, sizeof(reply)) < 0) {
+		close(data_listen);
+		stop_capture(s);
+		return -1;
+	}
+
+	RPCAPD_LOG(DEBUG, "awaiting connection");
+
+	data_fd = accept_from(data_listen, &s->peer, DATA_ACCEPT_TIMEOUT_MS);
+	close(data_listen);
+	if (data_fd < 0) {
+		stop_capture(s);
+		return -1;
+	}
+
+	/* Bound how long a send can block. */
+	if (send_timeout > 0) {
+		struct timeval tv = {
+			.tv_sec = send_timeout,
+		};
+
+		if (setsockopt(data_fd, SOL_SOCKET, SO_SNDTIMEO, &tv, sizeof(tv)) < 0)
+			RPCAPD_LOG(NOTICE, "cannot set data send timeout: %s",
+				   strerror(errno));
+	}
+
+	s->data.fd = data_fd;
+
+	RPCAPD_LOG(NOTICE,
+		   "capture started on %s (snaplen %u, data port %u)",
+		   s->name, s->snaplen, data_port);
+	return 0;
+}
+
+/*
+ * Frame each packet from the ring into an RPCAP_MSG_PACKET message and
+ * send it on the data connection.  MSG_MORE corks the socket until the
+ * ring drains, so a backlog coalesces into full segments.  pdump wraps
+ * packets in a pcapng enhanced packet block, which carries the capture
+ * time and the pre-truncation length.
+ */
+static ssize_t
+process_ring(struct session *s, unsigned int *avail)
+{
+	struct rte_mbuf *pkts[BURST_SIZE];
+	unsigned int i, n;
+	ssize_t written = 0;
+
+	n = rte_ring_sc_dequeue_burst(s->ring, (void **)pkts, BURST_SIZE, avail);
+	if (n == 0)
+		return 0;
+
+	for (i = 0; i < n; i++) {
+		struct rte_mbuf *m = pkts[i];
+		uint8_t buf[MAX_CAPTURE_LEN];
+		struct rte_pcapng_pkt pkt;
+		uint32_t caplen, wirelen;
+		const void *data;
+
+		if (unlikely(rte_pcapng_pkt_info(m, &pkt) != 0)) {
+			RPCAPD_LOG(ERR, "malformed capture mbuf on %s", s->name);
+			goto error;
+		}
+
+		caplen = pkt.captured_len;
+		if (unlikely(caplen > sizeof(buf)))
+			caplen = sizeof(buf);
+
+		/* clients reject a packet whose len is below its caplen */
+		wirelen = RTE_MAX(pkt.original_len, caplen);
+		data = rte_pktmbuf_read(m, pkt.data_offset, caplen, buf);
+		if (unlikely(data == NULL)) {
+			RPCAPD_LOG(ERR, "short capture mbuf on %s", s->name);
+			goto error;
+		}
+
+		s->npkt++;
+
+		struct rpcap_header hdr = {
+			.ver = RPCAP_VERSION,
+			.type = RPCAP_MSG_PACKET,
+			.plen = rte_cpu_to_be_32(sizeof(struct rpcap_pkthdr) + caplen),
+		};
+
+		/* rpcap protocol has timestamp in microseconds. */
+		uint64_t us = timestamp_to_ns(pkt.cycles) / 1000;
+		struct rpcap_pkthdr pkthdr = {
+			.timestamp_sec = rte_cpu_to_be_32(us / US_PER_S),
+			.timestamp_usec = rte_cpu_to_be_32(us % US_PER_S),
+			.caplen = rte_cpu_to_be_32(caplen),
+			.len = rte_cpu_to_be_32(wirelen),
+			.npkt = rte_cpu_to_be_32(s->npkt),
+		};
+
+		struct iovec iov[3] = {
+			{
+				.iov_base = &hdr,
+				.iov_len = sizeof(hdr),
+			},
+			{
+				.iov_base = &pkthdr,
+				.iov_len = sizeof(pkthdr),
+			},
+			{
+				.iov_base = (void *)(uintptr_t)data,
+				.iov_len = caplen,
+			},
+		};
+
+		/* more to come in this burst, or still queued in the ring */
+		bool more = (i + 1 < n) || (*avail > 0);
+
+		if (send_iov_full(&s->data, iov, 3, more ? MSG_MORE : 0) < 0) {
+			if (errno == EPIPE || errno == ECONNRESET)
+				RPCAPD_LOG(DEBUG, "data connection closed by client");
+			else if (errno == EAGAIN || errno == EWOULDBLOCK)
+				RPCAPD_LOG(NOTICE,
+					   "client stopped reading data connection, closing");
+			else
+				RPCAPD_LOG(NOTICE, "send on data connection failed: %s",
+					   strerror(errno));
+			goto error;
+		}
+		rte_pktmbuf_free(m);
+		written += sizeof(hdr) + sizeof(pkthdr) + caplen;
+	}
+
+	return written;
+
+error:
+	rte_pktmbuf_free_bulk(pkts + i, n - i);
+	return -1;
+}
+
+/* Poll the control socket while idle.
+ * Returns 0 to keep capturing, 1 if a control message (typically
+ * ENDCAP) is pending, or -1 if the client has gone away.
+ */
+static int
+check_socket_status(const struct conn *ctrl)
+{
+	struct pollfd pfd = { .fd = ctrl->fd, .events = POLLIN };
+
+	if (poll(&pfd, 1, 0) < 0) {
+		if (errno == EINTR)
+			return 0;
+		RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+		return -1;
+	}
+	if (pfd.revents & (POLLERR | POLLHUP | POLLNVAL)) {
+		RPCAPD_LOG(DEBUG, "client closed control connection");
+		return -1;
+	}
+	if (pfd.revents & POLLIN)
+		return 1;
+	return 0;
+}
+
+/*
+ * Drain the ring until a control message arrives, the data connection
+ * breaks, or a quit signal is delivered.  Returns 0 if the session
+ * should continue, -1 if the client is gone.
+ *
+ * The control socket is polled every iteration, not only when the ring
+ * is empty: a client waiting for a reply stops draining the data
+ * socket, and both ends wedge once the buffers fill.
+ */
+int
+capture_loop(const struct conn *ctrl, struct session *s)
+{
+	unsigned int empty_count = 0;
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		ssize_t written;
+		unsigned int avail = 0;
+
+		switch (check_socket_status(ctrl)) {
+		case 1:
+			/* control message pending, let caller service it */
+			return 0;
+		case 0:
+			break;
+		default:
+			/* client is gone */
+			return -1;
+		}
+
+		written = process_ring(s, &avail);
+		if (written < 0) {
+			/* process_ring has already logged the reason */
+			return -1;
+		}
+
+		if (written > 0) {
+			/* are there more packets? */
+			empty_count = (avail == 0);
+			continue;
+		}
+
+		if (empty_count < SLEEP_THRESHOLD) {
+			/* spin a few times before checking */
+			++empty_count;
+			rte_pause();
+			continue;
+		}
+
+		/* ring has been empty for a while: stop spinning */
+		rte_delay_us_sleep(SLEEP_US);
+	}
+	return 0;
+}
+
+int
+handle_endcap(const struct conn *c, uint32_t plen, struct session *s)
+{
+	if (rpcap_discard(c, plen) < 0)
+		return -1;
+	stop_capture(s);
+	return rpcap_send_msg(c, RPCAP_MSG_ENDCAP_REPLY, 0, NULL, 0);
+}
+
+int
+handle_stats(const struct conn *c, uint32_t plen, const struct session *s)
+{
+	struct rte_eth_stats es = { 0 };
+
+	if (rpcap_discard(c, plen) < 0)
+		return -1;
+
+	if (s->capture_on)
+		rte_eth_stats_get(s->port, &es);
+
+	struct rpcap_stats reply = {
+		.ifrecv   = rte_cpu_to_be_32((uint32_t)es.ipackets),
+		.ifdrop   = rte_cpu_to_be_32((uint32_t)es.ierrors),
+		.krnldrop = 0,
+		.svrcapt  = rte_cpu_to_be_32(s->npkt),
+	};
+	return rpcap_send_msg(c, RPCAP_MSG_STATS_REPLY, 0, &reply, sizeof(reply));
+}
diff --git a/app/rpcapd/filter.c b/app/rpcapd/filter.c
new file mode 100644
index 0000000000..bdf7b8eaf1
--- /dev/null
+++ b/app/rpcapd/filter.c
@@ -0,0 +1,164 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Capture filters.  The client compiles the filter, so it arrives as
+ * cBPF and has to be converted to the DPDK form that pdump takes.
+ */
+
+#include <stdlib.h>
+
+#include <pcap/pcap.h>
+
+#include <rte_bpf.h>
+#include <rte_byteorder.h>
+#include <rte_errno.h>
+#include <rte_malloc.h>
+#include <rte_pdump.h>
+
+#include "rpcap-protocol.h"
+#include "rpcapd.h"
+
+#define MAX_FILTER_INSNS              4096
+
+/*
+ * Read the optional capture filter that follows a start-capture request,
+ * and convert it for pdump. Client passes cBPF.
+ */
+int
+read_filter(const struct conn *c, uint32_t plen, struct session *s)
+{
+	struct rpcap_filterbpf_insn winsn;
+	struct rpcap_filter filter;
+	struct bpf_program bf;
+	struct bpf_insn *insns;
+	uint32_t i, nitems;
+
+	if (plen == 0)
+		return 0;		/* no filter: capture everything */
+
+	if (plen < sizeof(filter)) {
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
+		return rpcap_send_error(c, 0, "short filter header") < 0 ? -1 : 1;
+	}
+
+	if (recv_full(c, &filter, sizeof(filter)) < 0)
+		return -1;
+	plen -= sizeof(filter);
+
+	if (rte_be_to_cpu_16(filter.filtertype) != RPCAP_UPDATEFILTER_BPF) {
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
+		return rpcap_send_error(c, 0, "unsupported filter type") < 0 ? -1 : 1;
+	}
+
+	/* nitems is client-supplied; bound it before trusting the length. */
+	nitems = rte_be_to_cpu_32(filter.nitems);
+	if (nitems == 0)
+		return rpcap_discard(c, plen) < 0 ? -1 : 0;
+
+	if (nitems > MAX_FILTER_INSNS || plen < nitems * sizeof(winsn)) {
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
+		return rpcap_send_error(c, 0, "bad filter length") < 0 ? -1 : 1;
+	}
+
+	insns = calloc(nitems, sizeof(*insns));
+	if (insns == NULL) {
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
+		return rpcap_send_error(c, 0, "out of memory") < 0 ? -1 : 1;
+	}
+
+	for (i = 0; i < nitems; i++) {
+		if (recv_full(c, &winsn, sizeof(winsn)) < 0) {
+			free(insns);
+			return -1;
+		}
+		insns[i].code = rte_be_to_cpu_16(winsn.code);
+		insns[i].jt   = winsn.jt;
+		insns[i].jf   = winsn.jf;
+		insns[i].k    = rte_be_to_cpu_32(winsn.k);
+	}
+	plen -= nitems * sizeof(winsn);
+
+	/* Anything after the instructions is padding we do not need. */
+	if (rpcap_discard(c, plen) < 0) {
+		free(insns);
+		return -1;
+	}
+
+	bf.bf_len = nitems;
+	bf.bf_insns = insns;
+
+	/* Reject a malformed program here */
+	if (!bpf_validate(bf.bf_insns, bf.bf_len)) {
+		free(insns);
+		return rpcap_send_error(c, 0, "invalid filter program") < 0 ? -1 : 1;
+	}
+
+	/* A filter recorded by an earlier UPDATEFILTER may still be here */
+	rte_free(s->prm);
+	s->prm = rte_bpf_convert(&bf);
+	free(insns);
+	if (s->prm == NULL) {
+		RPCAPD_LOG(ERR, "rte_bpf_convert failed: %s",
+			rte_strerror(rte_errno));
+		return rpcap_send_error(c, 0, "cannot convert filter") < 0 ? -1 : 1;
+	}
+
+	RPCAPD_LOG(DEBUG, "capture filter: %u instructions", nitems);
+	return 0;
+}
+
+/*
+ * UPDATEFILTER_REQ: replace the capture filter.
+ *
+ * pdump takes its filter when the callback is setup.
+ * To replace need to drop old callback and put in new one.
+ * Packets already in the ring are kept.
+ *
+ * Before the capture starts this just records the filter for the
+ * eventual STARTCAP.
+ */
+int
+handle_updatefilter(const struct conn *c, uint32_t plen, struct session *s)
+{
+	struct rte_bpf_prm *old = s->prm;
+	int ret;
+
+	s->prm = NULL;
+	ret = read_filter(c, plen, s);
+	if (ret != 0) {
+		/* Malformed request: keep running with the old filter. */
+		rte_free(s->prm);
+		s->prm = old;
+		return ret < 0 ? -1 : 0;	/* error already reported */
+	}
+
+	if (!s->capture_on) {
+		rte_free(old);
+		return rpcap_send_msg(c, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+	}
+
+	rte_pdump_disable(s->port, RTE_PDUMP_ALL_QUEUES, s->pdump_flags);
+	s->capture_on = false;
+
+	if (rte_pdump_enable_bpf(s->port, RTE_PDUMP_ALL_QUEUES,
+				 s->pdump_flags | RTE_PDUMP_FLAG_PCAPNG,
+				 s->snaplen, s->ring, s->mp, s->prm) < 0) {
+		RPCAPD_LOG(ERR, "rte_pdump_enable_bpf port %u failed: %s",
+			s->port, rte_strerror(rte_errno));
+		rte_free(old);
+		/* The capture cannot be resumed */
+		stop_capture(s);
+		return rpcap_send_error(c, 0, "cannot apply filter");
+	}
+	s->capture_on = true;
+
+	/* Safe now that the old program is no longer referenced. */
+	rte_free(old);
+
+	RPCAPD_LOG(DEBUG, "capture filter updated on %s", s->name);
+	return rpcap_send_msg(c, RPCAP_MSG_UPDATEFILTER_REPLY, 0, NULL, 0);
+}
diff --git a/app/rpcapd/main.c b/app/rpcapd/main.c
new file mode 100644
index 0000000000..7b7288785d
--- /dev/null
+++ b/app/rpcapd/main.c
@@ -0,0 +1,577 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Demonstration server for the rpcap protocol for DPDK.
+ * This allows a libpcap client (e.g. Wireshark or tcpdump)
+ * to use "rpcap://host[:port]/portname" as capture device.
+ *
+ * Based on the DPDK dumpcap application and on rpcapd from libpcap:
+ *   https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
+ *
+ * Only the bits of the RPCAP protocol that are needed for an
+ * unauthenticated, passive-mode capture session are implemented.
+ * Configuration files, active mode, sampling and concurrent clients
+ * are intentionally omitted.
+ *
+ * Options, startup and the control connection dispatcher live here; the
+ * request handlers are in session.c, capture.c and filter.c.
+ */
+
+#include <arpa/inet.h>
+#include <errno.h>
+#include <getopt.h>
+#include <netinet/in.h>
+#include <netdb.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/types.h>
+#include <unistd.h>
+
+#include <rte_alarm.h>
+#include <rte_byteorder.h>
+#include <rte_common.h>
+#include <rte_debug.h>
+#include <rte_eal.h>
+#include <rte_ethdev.h>
+#include <rte_lcore.h>
+#include <rte_log.h>
+#include <rte_pdump.h>
+#include <rte_stdatomic.h>
+#include <rte_version.h>
+
+#include "rpcap-protocol.h"
+#include "rpcapd.h"
+
+#define DEFAULT_RING_SIZE             2048
+#define MAX_RING_SIZE                 (1U << 20)
+#define PRIMARY_MONITOR_INTERVAL_US   (500 * 1000)
+#define DATA_SEND_TIMEOUT_SEC         10
+
+/* Command-line options */
+static uint16_t listen_port = RPCAP_DEFAULT_NETPORT;
+uint32_t ring_size = DEFAULT_RING_SIZE;
+static const char *lcore_arg;
+static const char *file_prefix;
+static const char *bind_addr;		/* -b argument; NULL means loopback */
+static int bind_family = AF_UNSPEC;
+static const char *debug_file;		/* --debug-file argument */
+static unsigned int debug_log;		/* -D count: raise RPCAPD log verbosity */
+uint32_t send_timeout = DATA_SEND_TIMEOUT_SEC;	/* 0 means no limit */
+
+struct sockaddr_storage listen_addr;
+socklen_t               listen_addrlen;
+
+RTE_ATOMIC(bool) quit_signal;
+
+static bool
+is_loopback(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET) {
+		const struct sockaddr_in *sin = (const void *)ss;
+
+		return (ntohl(sin->sin_addr.s_addr) >> 24) == 127;
+	}
+	if (ss->ss_family == AF_INET6) {
+		const struct sockaddr_in6 *sin6 = (const void *)ss;
+
+		/* A v4 client on a dual-stack socket arrives as
+		 * ::ffff:127.0.0.1, which is loopback too.
+		 */
+		if (IN6_IS_ADDR_V4MAPPED(&sin6->sin6_addr))
+			return sin6->sin6_addr.s6_addr[12] == 127;
+
+		return IN6_IS_ADDR_LOOPBACK(&sin6->sin6_addr);
+	}
+	return false;
+}
+
+static void
+parse_bind_addr(void)
+{
+	struct addrinfo hints = {
+		.ai_family   = bind_family,
+		.ai_socktype = SOCK_STREAM,
+		.ai_flags    = AI_NUMERICHOST | AI_PASSIVE,
+	};
+	struct addrinfo *res;
+	int rc;
+
+	/* Loopback by default; the wildcard address is not a safe default. */
+	if (bind_addr == NULL)
+		bind_addr = (bind_family == AF_INET6) ? "::1" : "127.0.0.1";
+
+	rc = getaddrinfo(bind_addr, NULL, &hints, &res);
+	if (rc != 0)
+		rte_exit(EXIT_FAILURE, "Invalid bind address '%s': %s\n",
+			 bind_addr, gai_strerror(rc));
+	memcpy(&listen_addr, res->ai_addr, res->ai_addrlen);
+	listen_addrlen = res->ai_addrlen;
+	freeaddrinfo(res);
+}
+
+
+static void
+signal_handler(int sig __rte_unused)
+{
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+/* Service a single client until it disconnects. */
+static void
+handle_client(int ctrl_fd)
+{
+	struct sockaddr_storage peer;
+	socklen_t peerlen = sizeof(peer);
+	char host[NI_MAXHOST] = "?";
+	struct conn ctrl = { .fd = ctrl_fd };
+	struct session s = { .data.fd = -1 };
+
+	/* Remembered so the data connection can be restricted to this peer. */
+	if (getpeername(ctrl_fd, (struct sockaddr *)&peer, &peerlen) != 0) {
+		RPCAPD_LOG(ERR, "getpeername: %s", strerror(errno));
+		return;
+	}
+	s.peer = peer;
+	getnameinfo((struct sockaddr *)&peer, peerlen,
+		    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	RPCAPD_LOG(NOTICE, "client %s connected", host);
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		struct rpcap_header hdr;
+		uint32_t plen;
+
+		/* Drain the ring whenever a capture is running */
+		if (s.capture_on && capture_loop(&ctrl, &s) < 0)
+			goto done;
+
+		if (recv_full(&ctrl, &hdr, sizeof(hdr)) < 0)
+			break;
+
+		plen = rte_be_to_cpu_32(hdr.plen);
+
+		/* Only version 0 is spoken here */
+		if (hdr.ver != RPCAP_VERSION) {
+			RPCAPD_LOG(WARNING, "unsupported protocol version %u",
+				hdr.ver);
+			if (rpcap_discard(&ctrl, plen) < 0 ||
+			    rpcap_send_error(&ctrl, PCAP_ERR_WRONGVER,
+					     "unsupported protocol version") < 0)
+				goto done;
+			continue;
+		}
+
+		switch (hdr.type) {
+		case RPCAP_MSG_AUTH_REQ:
+			/* libpcap treats a zero-length AUTH_REPLY as "version
+			 * 0 only, same byte order".
+			 */
+			if (handle_auth(&ctrl, plen) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_FINDALLIF_REQ:
+			if (rpcap_discard(&ctrl, plen) < 0 || handle_findallif(&ctrl) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_OPEN_REQ:
+			if (handle_open(&ctrl, plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STARTCAP_REQ:
+			if (handle_startcap(&ctrl, plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_UPDATEFILTER_REQ:
+			if (handle_updatefilter(&ctrl, plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_ENDCAP_REQ:
+			if (handle_endcap(&ctrl, plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_STATS_REQ:
+			if (handle_stats(&ctrl, plen, &s) < 0)
+				goto done;
+			break;
+		case RPCAP_MSG_CLOSE:
+			rpcap_discard(&ctrl, plen);
+			goto done;
+		default:
+			RPCAPD_LOG(WARNING, "unsupported request type 0x%02x", hdr.type);
+			if (rpcap_discard(&ctrl, plen) < 0 ||
+			    rpcap_send_error(&ctrl, 0, "unsupported request") < 0)
+				goto done;
+			break;
+		}
+	}
+done:
+	stop_capture(&s);
+	close(ctrl_fd);
+	RPCAPD_LOG(NOTICE, "client %s disconnected", host);
+}
+
+static int
+open_listen_socket(uint16_t port)
+{
+	struct sockaddr_storage addr = listen_addr;
+	char host[NI_MAXHOST];
+	int fd, one = 1;
+
+	set_sockaddr_port(&addr, port);
+
+	fd = socket(addr.ss_family, SOCK_STREAM, 0);
+	if (fd < 0)
+		rte_exit(EXIT_FAILURE, "socket: %s\n", strerror(errno));
+	setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one));
+
+	if (bind(fd, (struct sockaddr *)&addr, listen_addrlen) < 0)
+		rte_exit(EXIT_FAILURE, "bind(%u): %s\n", port, strerror(errno));
+
+	int err = getnameinfo((struct sockaddr *)&listen_addr, listen_addrlen,
+			      host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+	if (err != 0)
+		rte_exit(EXIT_FAILURE, "Listen address lookup failed: %s\n",
+			 gai_strerror(err));
+
+	RPCAPD_LOG(NOTICE, "listening on %s port %u", host, listen_port);
+
+	if (!is_loopback(&listen_addr))
+		RPCAPD_LOG(WARNING,
+			"non-loopback address %s; "
+			"rpcap is unauthenticated and unencrypted, captured traffic is exposed to the network",
+			host);
+
+	if (listen(fd, 1) < 0)
+		rte_exit(EXIT_FAILURE, "listen: %s\n", strerror(errno));
+
+	return fd;
+}
+
+static void
+usage(FILE *f, const char *progname)
+{
+	fprintf(f, "Usage: %s [options]\n", progname);
+	fprintf(f,
+		"  -p, --port <port>     listen port (default %u)\n"
+		"  -b, --bind <addr>     bind address (default 127.0.0.1, ::1 with -6)\n"
+		"  -4                    use only IPv4\n"
+		"  -6                    use only IPv6\n"
+		"  -N <ring size>        ring size in packets (default %u)\n"
+		"  -D, --debug           increase log verbosity (-D info, -DD debug)\n"
+		"      --debug-file <f>  redirect log output to file <f> (append mode)\n"
+		"      --send-timeout <s> seconds a data send may block before the\n"
+		"                        client is treated as dead (default %u, 0 waits\n"
+		"                        forever)\n"
+		"      --version         print version and exit\n"
+		"  -h, --help            print this help and exit\n"
+		"      --lcore=<core>    CPU core to run on (default: any)\n"
+		"      --file-prefix=<p> prefix to use for multi-process\n"
+		"\n"
+		"WARNING: rpcap is unauthenticated and unencrypted.  Binding to\n"
+		"any non-loopback address exposes captured traffic to the\n"
+		"network.  Not for production use.\n",
+		RPCAP_DEFAULT_NETPORT, DEFAULT_RING_SIZE,
+		DATA_SEND_TIMEOUT_SEC);
+}
+
+static void
+print_version(void)
+{
+	printf("rpcapd, a remote packet capture daemon (DPDK pdump backend)\n"
+	       "Built against %s\n", rte_version());
+}
+
+static void
+parse_opts(int argc, char **argv)
+{
+	enum {
+		OPT_LONG_ONLY = 0x100,
+		OPT_DEBUG_FILE,
+		OPT_VERSION,
+		OPT_SEND_TIMEOUT,
+	};
+	static const struct option long_options[] = {
+		{ "port",         required_argument, NULL, 'p' },
+		{ "bind",         required_argument, NULL, 'b' },
+		{ "debug",        no_argument,       NULL, 'D' },
+		{ "help",         no_argument,       NULL, 'h' },
+		{ "version",      no_argument,       NULL, OPT_VERSION },
+		{ "debug-file",   required_argument, NULL, OPT_DEBUG_FILE },
+		{ "send-timeout", required_argument, NULL, OPT_SEND_TIMEOUT },
+		{ "file-prefix",  required_argument, NULL, 0 },
+		{ "lcore",        required_argument, NULL, 0 },
+		{ NULL, 0, NULL, 0 },
+	};
+	int option_index, c;
+
+	while ((c = getopt_long(argc, argv, "hD46p:b:N:",
+				long_options, &option_index)) != -1) {
+		switch (c) {
+		case 'p': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			if (u == 0 || u > UINT16_MAX)
+				rte_exit(EXIT_FAILURE, "Invalid port: %s\n", optarg);
+			listen_port = (uint16_t)u;
+			break;
+		}
+		case 'b':
+			bind_addr = optarg;
+			break;
+		case '4':
+			bind_family = AF_INET;
+			break;
+		case '6':
+			bind_family = AF_INET6;
+			break;
+		case 'N': {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			/* Check the full value before narrowing it: an upper
+			 * bound is needed anyway because rte_align32pow2()
+			 * wraps to zero above 2^31, and that failure would
+			 * otherwise only surface in rte_ring_create() on the
+			 * first capture.
+			 */
+			if (u < 64 || u > MAX_RING_SIZE)
+				rte_exit(EXIT_FAILURE,
+					 "Ring size must be between 64 and %u\n",
+					 MAX_RING_SIZE);
+			ring_size = (uint32_t)u;
+			/* rte_ring_create() requires a power of two. */
+			if (!rte_is_power_of_2(ring_size)) {
+				ring_size = rte_align32pow2(ring_size);
+				RPCAPD_LOG(NOTICE, "ring size rounded up to %u",
+					ring_size);
+			}
+			break;
+		}
+		case 'D':
+			debug_log++;
+			break;
+		case 'h':
+			usage(stdout, argv[0]);
+			exit(0);
+		case OPT_VERSION:
+			print_version();
+			exit(0);
+		case OPT_DEBUG_FILE:
+			debug_file = optarg;
+			break;
+		case OPT_SEND_TIMEOUT: {
+			unsigned long u = strtoul(optarg, NULL, 0);
+
+			/* Zero means wait forever, which is what the socket
+			 * does without SO_SNDTIMEO.
+			 */
+			if (u > INT32_MAX)
+				rte_exit(EXIT_FAILURE,
+					 "Invalid send timeout: %s\n", optarg);
+			send_timeout = (uint32_t)u;
+			break;
+		}
+		case 0: {
+			const char *longopt = long_options[option_index].name;
+
+			if (!strcmp(longopt, "lcore")) {
+				lcore_arg = optarg;
+				break;
+			} else if (!strcmp(longopt, "file-prefix")) {
+				file_prefix = optarg;
+				break;
+			}
+		}
+			/* fallthrough */
+		default:
+			usage(stderr, argv[0]);
+			exit(EXIT_FAILURE);
+		}
+	}
+
+	/* Resolve the bind address now that -4/-6/-b have been seen. */
+	parse_bind_addr();
+}
+
+/*
+ * Periodic check that the DPDK primary process is still alive.
+ * If it dies our shared-memory state (rings, mempools, pdump) becomes
+ * unsafe to touch, so we set quit_signal and let the main loop tear
+ * down cleanly on its next iteration.  The callback runs on the EAL
+ * interrupt thread; quit_signal is atomic so the read in the main
+ * loop is well-defined.
+ */
+static void
+monitor_primary(void *arg __rte_unused)
+{
+	if (rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed))
+		return;
+
+	if (rte_eal_primary_proc_alive(NULL)) {
+		rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL);
+		return;
+	}
+
+	RPCAPD_LOG(NOTICE, "primary process exited, shutting down");
+	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
+}
+
+static void
+enable_primary_monitor(void)
+{
+	if (rte_eal_alarm_set(PRIMARY_MONITOR_INTERVAL_US, monitor_primary, NULL) < 0)
+		RPCAPD_LOG(WARNING, "failed to install primary process monitor");
+}
+
+static void
+disable_primary_monitor(void)
+{
+	rte_eal_alarm_cancel(monitor_primary, NULL);
+}
+
+/*
+ * Bring up EAL as a secondary process so that pdump can attach to a
+ * running primary DPDK application. Hide most of the EAL
+ * complexity and only show serious messages from EAL.
+ */
+static int
+dpdk_init(void)
+{
+	static const char * const args[] = {
+		"rpcapd",
+		"--proc-type", "secondary",
+		"--log-level", "lib.eal:warning",
+	};
+	int eal_argc = RTE_DIM(args);
+	rte_cpuset_t cpuset = { };
+	char **eal_argv;
+	unsigned int i;
+
+	if (file_prefix != NULL)
+		eal_argc += 2;
+
+	if (lcore_arg != NULL)
+		eal_argc += 2;
+
+	eal_argv = calloc(eal_argc + 1, sizeof(char *));
+	if (eal_argv == NULL)
+		return -1;
+
+	for (i = 0; i < RTE_DIM(args); i++) {
+		eal_argv[i] = strdup(args[i]);
+		if (eal_argv[i] == NULL)
+			return -1;
+	}
+
+	if (file_prefix != NULL && *file_prefix != '\0') {
+		eal_argv[i++] = strdup("--file-prefix");
+		eal_argv[i++] = strdup(file_prefix);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+
+	if (lcore_arg != NULL) {
+		eal_argv[i++] = strdup("--lcores");
+		eal_argv[i++] = strdup(lcore_arg);
+		if (eal_argv[i - 1] == NULL || eal_argv[i - 2] == NULL)
+			return -1;
+	}
+	eal_argc = i;
+
+	/*
+	 * Need to get the original cpuset, before EAL init changes
+	 * the affinity of this thread (main lcore).
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_get_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		rte_panic("rte_thread_getaffinity failed\n");
+
+	if (rte_eal_init(eal_argc, eal_argv) < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failed: is the primary process running?\n");
+
+	/*
+	 * If no lcore argument was specified,
+	 * then run this program as a normal process
+	 * which can be scheduled on any non-isolated CPU.
+	 */
+	if (lcore_arg == NULL &&
+	    rte_thread_set_affinity_by_id(rte_thread_self(), &cpuset) != 0)
+		RPCAPD_LOG(INFO, "Can not restore original CPU affinity");
+
+	if (rte_pdump_init() < 0)
+		rte_exit(EXIT_FAILURE, "rte_pdump_init failed\n");
+
+	/* Needs the TSC frequency, so must follow rte_eal_init(). */
+	timestamp_init();
+
+	return 0;
+}
+
+int
+main(int argc, char **argv)
+{
+	struct sigaction action = {
+		.sa_handler = signal_handler,
+	};
+	int srv_fd;
+
+	parse_opts(argc, argv);
+
+	/*
+	 * Redirect log output before EAL init so EAL's own messages are
+	 * captured too.  The FILE handle is intentionally never closed:
+	 * the kernel reclaims it at process exit.
+	 */
+	if (debug_file != NULL) {
+		FILE *fp = fopen(debug_file, "a");
+
+		if (fp == NULL)
+			rte_exit(EXIT_FAILURE, "Cannot open debug file '%s': %s\n",
+				 debug_file, strerror(errno));
+		setvbuf(fp, NULL, _IOLBF, 0);
+		rte_openlog_stream(fp);
+	}
+
+	if (dpdk_init() < 0)
+		rte_exit(EXIT_FAILURE, "EAL init failure\n");
+
+	/* Default to NOTICE: only things the operator needs to see.
+	 * Each -D steps down one level, to INFO then DEBUG.
+	 */
+	rte_log_set_level(RTE_LOGTYPE_RPCAPD,
+			  debug_log >= 2 ? RTE_LOG_DEBUG :
+			  debug_log == 1 ? RTE_LOG_INFO : RTE_LOG_NOTICE);
+
+	if (rte_eth_dev_count_avail() == 0)
+		rte_exit(EXIT_FAILURE, "No Ethernet ports found\n");
+
+	sigaction(SIGTERM, &action, NULL);
+	sigaction(SIGINT, &action, NULL);
+
+	/* If peer closes, this detected in next recv() */
+	signal(SIGPIPE, SIG_IGN);
+
+	srv_fd = open_listen_socket(listen_port);
+
+	enable_primary_monitor();
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int cfd = accept_timeout(srv_fd, -1);
+
+		if (cfd < 0) {
+			if (errno == EINTR)
+				continue;
+			break;
+		}
+		handle_client(cfd);
+	}
+
+	disable_primary_monitor();
+	RPCAPD_LOG(NOTICE, "shutting down");
+	close(srv_fd);
+	rte_pdump_uninit();
+	return rte_eal_cleanup() ? EXIT_FAILURE : 0;
+}
diff --git a/app/rpcapd/meson.build b/app/rpcapd/meson.build
new file mode 100644
index 0000000000..61f4dc0a95
--- /dev/null
+++ b/app/rpcapd/meson.build
@@ -0,0 +1,25 @@
+# SPDX-License-Identifier: BSD-3-Clause
+# Copyright(c) 2026 Stephen Hemminger
+
+# relies on primary/secondary process, so Linux only
+if not is_linux
+    build = false
+    reason = 'only supported on Linux'
+    subdir_done()
+endif
+
+if not dpdk_conf.has('RTE_HAS_LIBPCAP')
+    build = false
+    reason = 'missing dependency, "libpcap"'
+    subdir_done()
+endif
+
+sources = files(
+        'capture.c',
+        'filter.c',
+        'main.c',
+        'session.c',
+        'sock.c',
+)
+ext_deps += pcap_dep
+deps += ['ethdev', 'pdump', 'bpf', 'pcapng']
diff --git a/app/rpcapd/rpcap-protocol.h b/app/rpcapd/rpcap-protocol.h
new file mode 100644
index 0000000000..438fd8dd84
--- /dev/null
+++ b/app/rpcapd/rpcap-protocol.h
@@ -0,0 +1,142 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * On-the-wire RPCAP protocol definitions, transcribed from libpcap's
+ * rpcap-protocol.h which is an internal file and not exported.
+ * See:
+ *   https://github.com/the-tcpdump-group/libpcap/blob/master/rpcap-protocol.h
+ *
+ * Only the subset needed by dpdk-rpcapd is included here.
+ * All multi-byte fields in the structures below are big-endian on the wire.
+ */
+
+#ifndef _RPCAP_PROTOCOL_H_
+#define _RPCAP_PROTOCOL_H_
+
+#include <stdint.h>
+
+#include <rte_byteorder.h>
+
+#define RPCAP_VERSION              0
+#define RPCAP_DEFAULT_NETPORT      2002
+
+/* Message types */
+#define RPCAP_MSG_ERROR            0x01
+#define RPCAP_MSG_FINDALLIF_REQ    0x02
+#define RPCAP_MSG_OPEN_REQ         0x03
+#define RPCAP_MSG_STARTCAP_REQ     0x04
+#define RPCAP_MSG_UPDATEFILTER_REQ 0x05
+#define RPCAP_MSG_CLOSE            0x06
+#define RPCAP_MSG_PACKET           0x07
+#define RPCAP_MSG_AUTH_REQ         0x08
+#define RPCAP_MSG_STATS_REQ        0x09
+#define RPCAP_MSG_ENDCAP_REQ       0x0a
+#define RPCAP_MSG_IS_REPLY         0x80
+
+#define RPCAP_MSG_FINDALLIF_REPLY    (RPCAP_MSG_FINDALLIF_REQ    | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_OPEN_REPLY         (RPCAP_MSG_OPEN_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STARTCAP_REPLY     (RPCAP_MSG_STARTCAP_REQ     | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_UPDATEFILTER_REPLY (RPCAP_MSG_UPDATEFILTER_REQ | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_AUTH_REPLY         (RPCAP_MSG_AUTH_REQ         | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_ENDCAP_REPLY       (RPCAP_MSG_ENDCAP_REQ       | RPCAP_MSG_IS_REPLY)
+#define RPCAP_MSG_STATS_REPLY	     (RPCAP_MSG_STATS_REQ	 | RPCAP_MSG_IS_REPLY)
+
+/* Error codes carried in the 'value' field of RPCAP_MSG_ERROR */
+#define PCAP_ERR_WRONGVER          17
+#define PCAP_ERR_AUTH_TYPE_NOTSUP  20
+
+/* Authentication types in rpcap_auth.type */
+#define RPCAP_RMTAUTH_NULL         0	/* no credentials supplied */
+#define RPCAP_RMTAUTH_PWD          1	/* username and password follow */
+
+/* Filter encoding: the filter is a BPF/NPF program */
+#define RPCAP_UPDATEFILTER_BPF     1
+
+/* Flags in rpcap_startcapreq.flags */
+#define RPCAP_STARTCAPREQ_FLAG_PROMISC     0x00000001	/* promiscuous mode */
+#define RPCAP_STARTCAPREQ_FLAG_DGRAM       0x00000002	/* use UDP for data */
+#define RPCAP_STARTCAPREQ_FLAG_SERVEROPEN  0x00000004	/* server connects out */
+#define RPCAP_STARTCAPREQ_FLAG_INBOUND     0x00000008	/* capture inbound only */
+#define RPCAP_STARTCAPREQ_FLAG_OUTBOUND    0x00000010	/* capture outbound only */
+
+/* Subset of pcap interface flags (pcap.h) */
+#define PCAP_IF_UP                 0x00000002
+#define PCAP_IF_RUNNING            0x00000004
+
+/* DLT_EN10MB - ethernet, the only link type we report */
+#define DLT_EN10MB                 1
+
+struct rpcap_header {
+	uint8_t     ver;
+	uint8_t     type;
+	rte_be16_t  value;
+	rte_be32_t  plen;
+};
+
+struct rpcap_findalldevs_if {
+	rte_be16_t  namelen;
+	rte_be16_t  desclen;
+	rte_be32_t  flags;
+	rte_be16_t  naddr;
+	uint16_t    dummy;
+};
+
+struct rpcap_openreply {
+	rte_be32_t  linktype;
+	rte_be32_t  tzoff;
+};
+
+struct rpcap_auth {
+	rte_be16_t  type;	/* RPCAP_RMTAUTH_* */
+	uint16_t    dummy;
+	rte_be16_t  slen1;	/* length of username, if any */
+	rte_be16_t  slen2;	/* length of password, if any */
+};
+
+struct rpcap_startcapreq {
+	rte_be32_t  snaplen;
+	rte_be32_t  read_timeout;
+	rte_be16_t  flags;
+	rte_be16_t  portdata;
+};
+
+struct rpcap_startcapreply {
+	rte_be32_t  bufsize;
+	rte_be16_t  portdata;
+	uint16_t    dummy;
+};
+
+/*
+ * A filter, sent either after rpcap_startcapreq or in an
+ * RPCAP_MSG_UPDATEFILTER_REQ, followed by nitems instructions.
+ */
+struct rpcap_filter {
+	rte_be16_t  filtertype;
+	uint16_t    dummy;
+	rte_be32_t  nitems;
+};
+
+/* One cBPF instruction, repeated nitems times after rpcap_filter. */
+struct rpcap_filterbpf_insn {
+	rte_be16_t  code;
+	uint8_t     jt;
+	uint8_t     jf;
+	rte_be32_t  k;
+};
+
+struct rpcap_stats {
+	rte_be32_t  ifrecv;
+	rte_be32_t  ifdrop;
+	rte_be32_t  krnldrop;
+	rte_be32_t  svrcapt;
+};
+
+struct rpcap_pkthdr {
+	rte_be32_t  timestamp_sec;
+	rte_be32_t  timestamp_usec;
+	rte_be32_t  caplen;
+	rte_be32_t  len;
+	rte_be32_t  npkt;
+};
+
+#endif /* _RPCAP_PROTOCOL_H_ */
diff --git a/app/rpcapd/rpcapd.h b/app/rpcapd/rpcapd.h
new file mode 100644
index 0000000000..df38231bfb
--- /dev/null
+++ b/app/rpcapd/rpcapd.h
@@ -0,0 +1,105 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * State and helpers shared between the parts of the rpcap daemon.
+ */
+
+#ifndef _RPCAPD_H_
+#define _RPCAPD_H_
+
+#include <stdbool.h>
+#include <stdint.h>
+#include <sys/socket.h>
+#include <sys/uio.h>
+
+#include <rte_ethdev.h>
+#include <rte_ether.h>
+#include <rte_log.h>
+#include <rte_mbuf.h>
+#include <rte_stdatomic.h>
+
+struct rte_bpf_prm;
+struct rte_mempool;
+struct rte_ring;
+
+#define RTE_LOGTYPE_RPCAPD RTE_LOGTYPE_USER1
+#define RPCAPD_LOG(level, ...) \
+	RTE_LOG_LINE_PREFIX(level, RPCAPD, "%s(): ", __func__, __VA_ARGS__)
+
+/* Largest snaplen a client can be given. */
+#define DEFAULT_SNAPLEN		RTE_MBUF_DEFAULT_DATAROOM
+
+/*
+ * rte_pcapng_copy() truncates to the snaplen and then re-inserts any
+ * VLAN or QinQ tag the NIC stripped, so a capture can exceed the
+ * snaplen by up to two tags.
+ */
+#define MAX_CAPTURE_LEN		(DEFAULT_SNAPLEN + 2 * sizeof(struct rte_vlan_hdr))
+
+/* A connection to the client. */
+struct conn {
+	int fd;
+};
+
+/* Per-client capture session state. */
+struct session {
+	struct conn data;			/* data connection */
+	struct sockaddr_storage peer;		/* control connection peer */
+	uint16_t port;				/* DPDK ethdev port being captured */
+	char     name[RTE_ETH_NAME_MAX_LEN];
+	uint32_t snaplen;
+	uint32_t npkt;				/* packet sequence for rpcap_pkthdr */
+	uint32_t pdump_flags;			/* direction bits handed to pdump */
+	bool     opened;			/* OPEN_REQ has selected a port */
+	bool     capture_on;
+	bool     promisc_set;			/* we enabled promiscuous mode */
+	struct rte_ring    *ring;
+	struct rte_mempool *mp;
+	struct rte_bpf_prm *prm;		/* capture filter, NULL if none */
+};
+
+/* Set once by the signal handler to unwind the main and capture loops. */
+extern RTE_ATOMIC(bool) quit_signal;
+
+/* Command-line settings needed outside of main.c */
+extern uint32_t ring_size;
+extern uint32_t send_timeout;		/* seconds; 0 means no limit */
+
+/* Address the control socket is bound to; the data socket uses the same
+ * address with an ephemeral port.
+ */
+extern struct sockaddr_storage listen_addr;
+extern socklen_t               listen_addrlen;
+
+/* sock.c: transport and message framing */
+int wait_readable(const struct conn *c, int timeout_ms);
+int accept_timeout(int listen_fd, int timeout_ms);
+int accept_from(int listen_fd, const struct sockaddr_storage *want,
+		int timeout_ms);
+int recv_full(const struct conn *c, void *buf, size_t len);
+int send_iov_full(const struct conn *c, struct iovec *iov, int iovcnt, int flags);
+int rpcap_send_msg(const struct conn *c, uint8_t type, uint16_t value,
+		   const void *payload, uint32_t plen);
+int rpcap_send_error(const struct conn *c, uint16_t errcode, const char *msg);
+int rpcap_discard(const struct conn *c, uint32_t plen);
+void set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port);
+uint16_t get_sockaddr_port(const struct sockaddr_storage *ss);
+
+/* session.c: control requests handled before a capture starts */
+int handle_auth(const struct conn *c, uint32_t plen);
+int handle_findallif(const struct conn *c);
+int handle_open(const struct conn *c, uint32_t plen, struct session *s);
+
+/* filter.c */
+int read_filter(const struct conn *c, uint32_t plen, struct session *s);
+int handle_updatefilter(const struct conn *c, uint32_t plen, struct session *s);
+
+/* capture.c */
+void timestamp_init(void);
+int handle_startcap(const struct conn *c, uint32_t plen, struct session *s);
+int handle_endcap(const struct conn *c, uint32_t plen, struct session *s);
+int handle_stats(const struct conn *c, uint32_t plen, const struct session *s);
+void stop_capture(struct session *s);
+int capture_loop(const struct conn *ctrl, struct session *s);
+
+#endif /* _RPCAPD_H_ */
diff --git a/app/rpcapd/session.c b/app/rpcapd/session.c
new file mode 100644
index 0000000000..cc14f26335
--- /dev/null
+++ b/app/rpcapd/session.c
@@ -0,0 +1,136 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Control requests handled before a capture starts: authentication,
+ * the interface list, and selecting an interface.
+ */
+
+#include <stdlib.h>
+#include <string.h>
+
+#include <rte_byteorder.h>
+#include <rte_ethdev.h>
+
+#include "rpcap-protocol.h"
+#include "rpcapd.h"
+
+/* Build and send the list of available DPDK ports. */
+int
+handle_findallif(const struct conn *c)
+{
+	uint8_t *buf = NULL;
+	size_t buflen = 0;
+	uint16_t nif = 0;
+	uint16_t p;
+	int rc;
+
+	RTE_ETH_FOREACH_DEV(p) {
+		static const char desc[] = "DPDK port";
+		char name[RTE_ETH_NAME_MAX_LEN];
+		size_t namelen, desclen, entry;
+		uint8_t *nb;
+
+		if (rte_eth_dev_get_name_by_port(p, name) < 0) {
+			RPCAPD_LOG(DEBUG, "can not find name for port %u", p);
+			continue;
+		}
+
+		RPCAPD_LOG(DEBUG, "findallif: port %u -> '%s'", p, name);
+		namelen = strlen(name);
+		desclen = strlen(desc);
+		entry = sizeof(struct rpcap_findalldevs_if) + namelen + desclen;
+
+		nb = realloc(buf, buflen + entry);
+		if (nb == NULL) {
+			RPCAPD_LOG(ERR, "out of memory in findallif");
+			free(buf);
+			return rpcap_send_error(c, 0, "out of memory");
+		}
+		buf = nb;
+
+		struct rpcap_findalldevs_if iface = {
+			.namelen = rte_cpu_to_be_16(namelen),
+			.desclen = rte_cpu_to_be_16(desclen),
+			.flags = rte_cpu_to_be_32(PCAP_IF_UP | PCAP_IF_RUNNING),
+		};
+		memcpy(buf + buflen, &iface, sizeof(iface));
+		memcpy(buf + buflen + sizeof(iface), name, namelen);
+		memcpy(buf + buflen + sizeof(iface) + namelen, desc, desclen);
+		buflen += entry;
+		nif++;
+	}
+
+	RPCAPD_LOG(DEBUG, "findallif: %u interface(s)", nif);
+	rc = rpcap_send_msg(c, RPCAP_MSG_FINDALLIF_REPLY, nif, buf, buflen);
+	free(buf);
+	return rc;
+}
+
+/*
+ * AUTH_REQ: check the authentication type only.
+ *
+ * There is no credential store, so a username and password cannot be
+ * verified; refuse them rather than reply that they were accepted.
+ */
+int
+handle_auth(const struct conn *c, uint32_t plen)
+{
+	struct rpcap_auth auth;
+	uint16_t type;
+
+	if (plen < sizeof(auth)) {
+		rpcap_discard(c, plen);
+		return rpcap_send_error(c, 0, "short authentication request");
+	}
+
+	if (recv_full(c, &auth, sizeof(auth)) < 0)
+		return -1;
+
+	/* Discard any username and password that followed. */
+	if (rpcap_discard(c, plen - sizeof(auth)) < 0)
+		return -1;
+
+	type = rte_be_to_cpu_16(auth.type);
+	if (type != RPCAP_RMTAUTH_NULL) {
+		RPCAPD_LOG(NOTICE, "rejecting authentication type %u", type);
+		return rpcap_send_error(c, PCAP_ERR_AUTH_TYPE_NOTSUP,
+					"this server cannot check credentials; "
+					"connect without a username or password");
+	}
+
+	return rpcap_send_msg(c, RPCAP_MSG_AUTH_REPLY, 0, NULL, 0);
+}
+
+/* OPEN_REQ: payload is the interface name (no NUL). */
+int
+handle_open(const struct conn *c, uint32_t plen, struct session *s)
+{
+	struct rpcap_openreply reply = {
+		.linktype = rte_cpu_to_be_32(DLT_EN10MB),
+	};
+	uint16_t port;
+
+	stop_capture(s);
+
+	if (plen >= sizeof(s->name)) {
+		rpcap_discard(c, plen);
+		return rpcap_send_error(c, 0, "interface name too long");
+	}
+	if (recv_full(c, s->name, plen) < 0)
+		return -1;
+	s->name[plen] = '\0';
+
+	if (rte_eth_dev_get_port_by_name(s->name, &port) < 0) {
+		RPCAPD_LOG(WARNING, "open: no such port '%s'", s->name);
+		/* s->name has already been overwritten; make sure a later
+		 * STARTCAP cannot capture the previously opened port.
+		 */
+		s->opened = false;
+		return rpcap_send_error(c, 0, "unknown interface");
+	}
+	s->port = port;
+	s->opened = true;
+
+	RPCAPD_LOG(DEBUG, "open: '%s' -> dpdk port %u", s->name, port);
+	return rpcap_send_msg(c, RPCAP_MSG_OPEN_REPLY, 0, &reply, sizeof(reply));
+}
diff --git a/app/rpcapd/sock.c b/app/rpcapd/sock.c
new file mode 100644
index 0000000000..179c1a31c1
--- /dev/null
+++ b/app/rpcapd/sock.c
@@ -0,0 +1,315 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * Socket helpers and rpcap message framing, used by both the control
+ * connection and the data connection.
+ */
+
+#include <errno.h>
+#include <netdb.h>
+#include <netinet/in.h>
+#include <poll.h>
+#include <stdbool.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/uio.h>
+#include <time.h>
+#include <unistd.h>
+
+#include <rte_byteorder.h>
+#include <rte_common.h>
+#include <rte_stdatomic.h>
+
+#include "rpcap-protocol.h"
+#include "rpcapd.h"
+
+#define POLL_INTERVAL_MS              500
+
+/* Monotonic milliseconds, for timing out across repeated waits. */
+static int64_t
+get_monotonic_ms(void)
+{
+	struct timespec ts;
+
+	clock_gettime(CLOCK_MONOTONIC, &ts);
+	return (int64_t)ts.tv_sec * 1000 + ts.tv_nsec / 1000000;
+}
+
+void
+set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port)
+{
+	if (ss->ss_family == AF_INET6)
+		((struct sockaddr_in6 *)ss)->sin6_port = htons(port);
+	else
+		((struct sockaddr_in *)ss)->sin_port = htons(port);
+}
+
+uint16_t
+get_sockaddr_port(const struct sockaddr_storage *ss)
+{
+	if (ss->ss_family == AF_INET6)
+		return ntohs(((const struct sockaddr_in6 *)ss)->sin6_port);
+	return ntohs(((const struct sockaddr_in *)ss)->sin_port);
+}
+
+
+/* Wait for a connection to become readable with timeout */
+int
+wait_readable(const struct conn *c, int timeout_ms)
+{
+	struct pollfd pfd = { .fd = c->fd, .events = POLLIN };
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		int wait_ms = POLL_INTERVAL_MS;
+		int rc;
+
+		if (timeout_ms >= 0) {
+			if (timeout_ms == 0)
+				return 0;
+			if (timeout_ms < wait_ms)
+				wait_ms = timeout_ms;
+			timeout_ms -= wait_ms;
+		}
+
+		rc = poll(&pfd, 1, wait_ms);
+		if (rc < 0) {
+			if (errno == EINTR)
+				continue;
+			RPCAPD_LOG(ERR, "poll failed: %s", strerror(errno));
+			return -1;
+		}
+		if (rc > 0)
+			return 1;
+	}
+	return -1;
+}
+
+/* accept() with a timeout, so a stalled client cannot wedge the daemon. */
+int
+accept_timeout(int listen_fd, int timeout_ms)
+{
+	struct conn listener = { .fd = listen_fd };
+	int fd;
+
+	switch (wait_readable(&listener, timeout_ms)) {
+	case 1:
+		break;
+	case 0:
+		RPCAPD_LOG(ERR, "timed out waiting for data connection");
+		return -1;
+	default:
+		return -1;
+	}
+
+	fd = accept(listen_fd, NULL, NULL);
+	if (fd < 0)
+		RPCAPD_LOG(ERR, "accept: %s", strerror(errno));
+	return fd;
+}
+
+/* Compare the host part of two addresses, ignoring the port: the data
+ * connection comes from an ephemeral port, not the control one.
+ */
+static bool
+same_host(const struct sockaddr_storage *a, const struct sockaddr_storage *b)
+{
+	if (a->ss_family != b->ss_family)
+		return false;
+
+	if (a->ss_family == AF_INET) {
+		const struct sockaddr_in *sa = (const void *)a;
+		const struct sockaddr_in *sb = (const void *)b;
+
+		return sa->sin_addr.s_addr == sb->sin_addr.s_addr;
+	}
+	if (a->ss_family == AF_INET6) {
+		const struct sockaddr_in6 *sa = (const void *)a;
+		const struct sockaddr_in6 *sb = (const void *)b;
+
+		return IN6_ARE_ADDR_EQUAL(&sa->sin6_addr, &sb->sin6_addr);
+	}
+	return false;
+}
+
+/*
+ * Accept a data connection only from the control connection's peer;
+ * the port is handed to the client in the clear, so any local user
+ * could otherwise race for the stream.  A mismatch is rejected and the
+ * wait continues.
+ */
+int
+accept_from(int listen_fd, const struct sockaddr_storage *want, int timeout_ms)
+{
+	struct conn listener = { .fd = listen_fd };
+	int remaining = timeout_ms;
+
+	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
+		struct sockaddr_storage peer;
+		socklen_t peerlen = sizeof(peer);
+		char host[NI_MAXHOST] = "?";
+		int64_t start, waited;
+		int fd;
+
+		start = get_monotonic_ms();
+		switch (wait_readable(&listener, remaining)) {
+		case 1:
+			break;
+		case 0:
+			RPCAPD_LOG(ERR, "timed out waiting for data connection");
+			return -1;
+		default:
+			return -1;
+		}
+
+		fd = accept(listen_fd, (struct sockaddr *)&peer, &peerlen);
+		if (fd < 0) {
+			if (errno == EINTR || errno == ECONNABORTED)
+				goto next;
+			RPCAPD_LOG(ERR, "accept: %s", strerror(errno));
+			return -1;
+		}
+
+		if (same_host(&peer, want))
+			return fd;
+
+		getnameinfo((struct sockaddr *)&peer, peerlen,
+			    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
+		RPCAPD_LOG(WARNING,
+			   "rejected data connection from %s: does not match control peer",
+			   host);
+		close(fd);
+next:
+		if (remaining >= 0) {
+			waited = get_monotonic_ms() - start;
+			remaining -= (waited > 0) ? (int)waited : 0;
+			if (remaining <= 0) {
+				RPCAPD_LOG(ERR,
+					   "timed out waiting for data connection");
+				return -1;
+			}
+		}
+	}
+	return -1;
+}
+
+/* Read exactly len bytes; return 0 on success, -1 on error or EOF. */
+int
+recv_full(const struct conn *c, void *buf, size_t len)
+{
+	uint8_t *p = buf;
+
+	while (len > 0) {
+		ssize_t n;
+
+		/* Timed wait, so a quit signal or a dead primary is acted
+		 * on promptly.
+		 */
+		if (wait_readable(c, -1) != 1)
+			return -1;
+
+		n = recv(c->fd, p, len, 0);
+		if (n < 0 && errno == EINTR)
+			continue;
+
+		if (n <= 0)
+			return -1;
+
+		p += n;
+		len -= n;
+	}
+	return 0;
+}
+
+/*
+ * Send all of iov, resending the remainder if sendmsg() reports a short
+ * count (possible when the connection breaks or a signal arrives after
+ * some bytes were copied).  Consumes iov, so pass a scratch copy.
+ */
+int
+send_iov_full(const struct conn *c, struct iovec *iov, int iovcnt, int flags)
+{
+	struct msghdr msg = {
+		.msg_iov    = iov,
+		.msg_iovlen = iovcnt,
+	};
+
+	while (msg.msg_iovlen > 0) {
+		ssize_t n = sendmsg(c->fd, &msg, flags | MSG_NOSIGNAL);
+
+		if (n < 0) {
+			/*
+			 * Send blocks rather than polling first; the data
+			 * socket has a send timeout so a client that stops
+			 * reading fails with EAGAIN.
+			 */
+			if (errno == EINTR &&
+			    !rte_atomic_load_explicit(&quit_signal,
+						      rte_memory_order_relaxed))
+				continue;
+			return -1;
+		}
+		if (n == 0)
+			return -1;
+
+		/* Drop whole iovecs that were fully sent, then trim the
+		 * partially sent one.
+		 */
+		while (msg.msg_iovlen > 0 && (size_t)n >= msg.msg_iov->iov_len) {
+			n -= msg.msg_iov->iov_len;
+			msg.msg_iov++;
+			msg.msg_iovlen--;
+		}
+		if (n > 0) {
+			msg.msg_iov->iov_base = (char *)msg.msg_iov->iov_base + n;
+			msg.msg_iov->iov_len -= n;
+		}
+	}
+	return 0;
+}
+
+int
+rpcap_send_msg(const struct conn *c, uint8_t type, uint16_t value,
+	       const void *payload, uint32_t plen)
+{
+	struct rpcap_header hdr = {
+		.ver = RPCAP_VERSION,
+		.type = type,
+		.value = rte_cpu_to_be_16(value),
+		.plen = rte_cpu_to_be_32(plen),
+	};
+	struct iovec iov[2] = {
+		{
+			.iov_base = &hdr,
+			.iov_len = sizeof(hdr),
+		},
+		{
+			.iov_base = (void *)(uintptr_t)payload,
+			.iov_len = plen,
+		},
+	};
+
+	return send_iov_full(c, iov, plen > 0 ? 2 : 1, 0);
+}
+
+int
+rpcap_send_error(const struct conn *c, uint16_t errcode, const char *msg)
+{
+	RPCAPD_LOG(WARNING, "sending error to client: %s", msg);
+	return rpcap_send_msg(c, RPCAP_MSG_ERROR, errcode, msg, strlen(msg));
+}
+
+/* Throw away plen bytes of payload we don't care about. */
+int
+rpcap_discard(const struct conn *c, uint32_t plen)
+{
+	uint8_t buf[256];
+
+	while (plen > 0) {
+		size_t chunk = plen > sizeof(buf) ? sizeof(buf) : plen;
+
+		if (recv_full(c, buf, chunk) < 0)
+			return -1;
+		plen -= chunk;
+	}
+	return 0;
+}
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 5b5a9f006e..8e107b48b6 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -143,6 +143,11 @@ New Features
   Added ``rte_bbdev_queue_stats_get()`` function to retrieve statistics
   for a specific queue, complementing the existing device-level statistics API.
 
+* **Added libpcap remote capture daemon.**
+
+  Added the ``dpdk-rpcapd`` application, which implements the rpcap
+  protocol to allow live capture in tcpdump and Wireshark.
+
 
 Removed Items
 -------------
diff --git a/doc/guides/tools/index.rst b/doc/guides/tools/index.rst
index 13f75a5bc6..a23333f763 100644
--- a/doc/guides/tools/index.rst
+++ b/doc/guides/tools/index.rst
@@ -13,6 +13,7 @@ DPDK Tools User Guides
     proc_info
     pmdinfo
     dumpcap
+    rpcapd
     pdump
     telemetrywatcher
     dmaperf
diff --git a/doc/guides/tools/rpcapd.rst b/doc/guides/tools/rpcapd.rst
new file mode 100644
index 0000000000..a8b026a409
--- /dev/null
+++ b/doc/guides/tools/rpcapd.rst
@@ -0,0 +1,199 @@
+..  SPDX-License-Identifier: BSD-3-Clause
+    Copyright(c) 2026 Stephen Hemminger
+
+.. _rpcapd_tool:
+
+dpdk-rpcapd Application
+=======================
+
+The ``dpdk-rpcapd`` application is a Data Plane Development Kit
+(DPDK) implementation of the remote packet capture daemon protocol
+(``rpcap``) used by libpcap.  It runs as a DPDK secondary process and
+allows libpcap-aware tools such as ``tcpdump`` and Wireshark to capture
+packets from a DPDK primary process live, without writing to an
+intermediate file.
+
+The ``dpdk-rpcapd`` tool implements a subset of the protocol spoken by
+the libpcap project's ``rpcapd``.
+See
+https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
+for the reference implementation.
+Clients connect to ``dpdk-rpcapd`` using a ``rpcap://`` URL,
+request the list of available interfaces(which are the ports of the DPDK primary),
+open one, and stream packets from it.
+
+.. warning::
+
+   ``dpdk-rpcapd`` listens on an unauthenticated, unencrypted TCP port
+   (default 2002).  Anyone able to reach the port can list DPDK ports
+   and capture all traffic flowing through them.  The default bind
+   address is ``127.0.0.1``, so the listener is not reachable from
+   other hosts; overriding this with ``--bind`` exposes captured
+   traffic to anyone who can reach that address.  **Do not run
+   ``dpdk-rpcapd`` on a production system.**
+
+
+Running the Application
+-----------------------
+
+The application has a small set of command-line options:
+
+*   ``-p <port>``, ``--port <port>``
+
+    TCP port to listen on.  Default is 2002, the IANA-assigned rpcap
+    port.
+
+*   ``-b <addr>``, ``--bind <addr>``
+
+    Numeric IPv4 or IPv6 address to bind the listener to.  Default is
+    ``127.0.0.1``, or ``::1`` when ``-6`` is given (loopback only).
+    See the warning above before using any other address.
+
+*   ``-4``
+
+    Use only IPv4; an IPv6 argument to ``-b`` is rejected.
+
+*   ``-6``
+
+    Use only IPv6; an IPv4 argument to ``-b`` is rejected.  The default
+    bind address becomes ``::1``.
+
+*   ``-N <ring_size>``
+
+    Size of the per-session capture ring in packets.  Default is 2048.
+    Rounded up to a power of two if necessary.
+
+*   ``-D``, ``--debug``
+
+    Increase log verbosity.  A single ``-D`` adds informational
+    messages; ``-DD`` adds per-request protocol detail.
+
+*   ``--debug-file <file>``
+
+    Append log output to ``<file>`` instead of writing it to standard
+    error.
+
+*   ``--send-timeout <seconds>``
+
+    How long a send on the data connection may block before the client
+    is treated as dead and the capture stopped.  Default is 10 seconds;
+    zero waits forever.
+
+*   ``--lcore <core>``
+
+    CPU core to run on.  By default the daemon runs as an ordinary
+    process on any non-isolated CPU.
+
+*   ``--file-prefix <prefix>``
+
+    EAL file prefix of the primary process to attach to.  Needed when
+    the primary was started with a non-default prefix.
+
+*   ``--version``
+
+    Print the version and exit.
+
+*   ``-h``, ``--help``
+
+    Print usage and exit.
+
+EAL options are supplied automatically; the application runs as a
+secondary process and does not need EAL options on its command line for
+typical use.
+
+
+Client Setup
+------------
+
+Most Linux distributions ship libpcap built without ``rpcap`` support,
+since ``--enable-remote`` is off by default.  To use ``dpdk-rpcapd``
+from ``tcpdump`` or Wireshark on Linux, rebuild libpcap with it:
+
+.. code-block:: console
+
+    wget https://www.tcpdump.org/release/libpcap-1.10.7.tar.xz
+    tar xf libpcap-1.10.7.tar.xz
+    cd libpcap-1.10.7
+    ./configure --enable-remote
+    make
+    sudo make install
+
+Only the client side of ``rpcap`` is used for ``dpdk-rpcapd``.
+Do not run libpcap's version of ``rpcapd``.
+
+``tcpdump`` rebuilt against this libpcap can be used as a client without
+further changes.  Wireshark on Windows and macOS ships with rpcap support
+enabled by default.
+
+
+Example
+-------
+
+Start a primary application with the packet capture framework
+initialized.  ``dpdk-testpmd`` is the simplest:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-testpmd --vdev=net_tap0 -- -i
+
+In another window, start ``dpdk-rpcapd``:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-rpcapd
+    RPCAPD: open_listen_socket(): listening on 127.0.0.1 port 2002
+
+In a third window, list available interfaces using a libpcap-based
+``tcpdump`` rebuilt with remote support:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump --list-remote-interfaces=rpcap://localhost:2002/
+    rpcap://localhost:2002/net_tap0  Network adapter 'DPDK port' on remote node localhost
+
+Capture live from a port:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -nn -c 20
+
+Or save to a file readable by any pcap consumer:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcap://localhost:2002/net_tap0 -w /tmp/capture.pcap
+
+
+Limitations
+-----------
+
+The following features of the reference ``rpcapd`` are not implemented
+in this initial version:
+
+*   **Single client.** Only one client may be connected at a time.
+    Subsequent clients are queued by the listening socket but not
+    serviced until the first disconnects.
+
+*   **No authentication.** Password authentication is refused with
+    ``PCAP_ERR_AUTH_TYPE_NOTSUP``; clients must connect without
+    credentials, which is what a ``rpcap://`` URL with no userinfo does.
+    With the default loopback bind, reaching the port already requires
+    an account on the host.
+
+*   **No TLS.** The ``-S`` option of the reference ``rpcapd`` is not
+    implemented, so the connection is always in the clear.  This is
+    reasonable for the default loopback bind, where the traffic never
+    leaves the host, but means ``--bind`` to any other address sends
+    captured packets over the network unencrypted.
+
+*   **TCP data transport only.** A client requesting UDP is refused.
+
+
+See Also
+--------
+
+*   :doc:`dumpcap` -- file-based capture writing pcapng
+    output.
+
+*   The libpcap project's ``rpcapd`` reference implementation:
+    https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v4 3/4] app/rpcapd: add TLS support
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 1/4] pcapng: add API to read back capture mbuf header Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 2/4] app/rpcapd: remote pcap daemon Stephen Hemminger
@ 2026-10-01  2:35   ` Stephen Hemminger
  2026-10-01  2:35   ` [PATCH v4 4/4] app/rpcapd: add host list option Stephen Hemminger
  2026-10-01 18:50   ` [PATCH v4 0/4] add rpcap remote capture daemon Marat Khalili
  4 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01  2:35 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Reshma Pattan

Support remote packet capture over TLS.
This requires OpenSSL to be available.

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---
 app/rpcapd/capture.c                   |  15 ++
 app/rpcapd/main.c                      | 154 +++++++++++++-
 app/rpcapd/meson.build                 |  15 ++
 app/rpcapd/rpcap-protocol.h            |   3 +
 app/rpcapd/rpcapd.h                    |  20 +-
 app/rpcapd/session.c                   | 182 +++++++++++++++--
 app/rpcapd/sock.c                      |  52 ++++-
 app/rpcapd/tls.c                       | 273 +++++++++++++++++++++++++
 doc/guides/rel_notes/release_26_11.rst |   2 +
 doc/guides/tools/rpcapd.rst            | 119 +++++++++--
 10 files changed, 789 insertions(+), 46 deletions(-)
 create mode 100644 app/rpcapd/tls.c

diff --git a/app/rpcapd/capture.c b/app/rpcapd/capture.c
index 17fce3fbca..b168216246 100644
--- a/app/rpcapd/capture.c
+++ b/app/rpcapd/capture.c
@@ -168,6 +168,7 @@ stop_capture(struct session *s)
 	rte_free(s->prm);
 	s->prm = NULL;
 	if (s->data.fd >= 0) {
+		tls_close(&s->data);
 		close(s->data.fd);
 		s->data.fd = -1;
 	}
@@ -312,6 +313,14 @@ handle_startcap(const struct conn *c, uint32_t plen, struct session *s)
 
 	s->data.fd = data_fd;
 
+	/* The client starts its handshake as soon as it has connected,
+	 * so promote before anything is sent.
+	 */
+	if (use_tls && tls_accept(&s->data) < 0) {
+		stop_capture(s);
+		return -1;
+	}
+
 	RPCAPD_LOG(NOTICE,
 		   "capture started on %s (snaplen %u, data port %u)",
 		   s->name, s->snaplen, data_port);
@@ -427,6 +436,12 @@ check_socket_status(const struct conn *ctrl)
 {
 	struct pollfd pfd = { .fd = ctrl->fd, .events = POLLIN };
 
+	/* A request may already be decrypted and waiting out of sight
+	 * of poll(), sharing a TLS record with an earlier one.
+	 */
+	if (tls_pending(ctrl))
+		return 1;
+
 	if (poll(&pfd, 1, 0) < 0) {
 		if (errno == EINTR)
 			return 0;
diff --git a/app/rpcapd/main.c b/app/rpcapd/main.c
index 7b7288785d..0cf6c3f4ba 100644
--- a/app/rpcapd/main.c
+++ b/app/rpcapd/main.c
@@ -62,13 +62,17 @@ static int bind_family = AF_UNSPEC;
 static const char *debug_file;		/* --debug-file argument */
 static unsigned int debug_log;		/* -D count: raise RPCAPD log verbosity */
 uint32_t send_timeout = DATA_SEND_TIMEOUT_SEC;	/* 0 means no limit */
+bool use_tls;				/* -S */
+bool null_auth_ok;			/* -n */
+static const char *tls_certfile;	/* -X argument */
+static const char *tls_keyfile;		/* -K argument */
 
 struct sockaddr_storage listen_addr;
 socklen_t               listen_addrlen;
 
 RTE_ATOMIC(bool) quit_signal;
 
-static bool
+bool
 is_loopback(const struct sockaddr_storage *ss)
 {
 	if (ss->ss_family == AF_INET) {
@@ -121,6 +125,62 @@ signal_handler(int sig __rte_unused)
 	rte_atomic_store_explicit(&quit_signal, true, rte_memory_order_relaxed);
 }
 
+/*
+ * TLS is not negotiated in the rpcap protocol, so a mismatch has to be
+ * detected from the first byte: an rpcap message starts with the
+ * protocol version 0, a TLS handshake with content type 22.
+ */
+#define TLS_RECORD_TYPE_HANDSHAKE	22
+
+/* How long a client has to send its first byte. */
+#define FIRST_BYTE_TIMEOUT_MS		(10 * 1000)
+
+static int
+setup_tls(struct conn *ctrl)
+{
+	uint8_t first;
+
+	/* Bounded wait: only one client is served at a time, so a peer
+	 * that connects and says nothing must not hold the daemon.
+	 */
+	switch (wait_readable(ctrl, FIRST_BYTE_TIMEOUT_MS)) {
+	case 1:
+		break;
+	case 0:
+		RPCAPD_LOG(NOTICE, "client sent nothing within %u seconds, closing",
+			FIRST_BYTE_TIMEOUT_MS / 1000);
+		return -1;
+	default:
+		return -1;
+	}
+
+	if (recv(ctrl->fd, &first, 1, MSG_PEEK) != 1)
+		return -1;
+
+	if (!use_tls) {
+		if (first == TLS_RECORD_TYPE_HANDSHAKE) {
+			tls_reject_handshake(ctrl->fd);
+			return -1;
+		}
+		return 0;
+	}
+
+	if (first != TLS_RECORD_TYPE_HANDSHAKE) {
+		struct rpcap_header hdr;
+
+		/* Reply in the clear; it is all the client will understand. */
+		RPCAPD_LOG(WARNING, "rejecting plaintext client: server requires TLS");
+		if (recv_full(ctrl, &hdr, sizeof(hdr)) == 0)
+			rpcap_discard(ctrl, rte_be_to_cpu_32(hdr.plen));
+
+		rpcap_send_error(ctrl, PCAP_ERR_TLS_REQUIRED,
+				 "TLS is required by this server; use rpcaps://");
+		return -1;
+	}
+
+	return tls_accept(ctrl);
+}
+
 /* Service a single client until it disconnects. */
 static void
 handle_client(int ctrl_fd)
@@ -134,6 +194,7 @@ handle_client(int ctrl_fd)
 	/* Remembered so the data connection can be restricted to this peer. */
 	if (getpeername(ctrl_fd, (struct sockaddr *)&peer, &peerlen) != 0) {
 		RPCAPD_LOG(ERR, "getpeername: %s", strerror(errno));
+		close(ctrl_fd);
 		return;
 	}
 	s.peer = peer;
@@ -141,6 +202,15 @@ handle_client(int ctrl_fd)
 		    host, sizeof(host), NULL, 0, NI_NUMERICHOST);
 	RPCAPD_LOG(NOTICE, "client %s connected", host);
 
+	if (!use_tls && !is_loopback(&peer))
+		RPCAPD_LOG(ERR,
+			"remote client %s is connected without TLS; "
+			"captured traffic and any credentials are exposed to the network",
+			host);
+
+	if (setup_tls(&ctrl) < 0)
+		goto done;
+
 	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
 		struct rpcap_header hdr;
 		uint32_t plen;
@@ -165,12 +235,24 @@ handle_client(int ctrl_fd)
 			continue;
 		}
 
+		/* Nothing but authentication is served until it succeeds. */
+		if (!s.authenticated && hdr.type != RPCAP_MSG_AUTH_REQ &&
+		    hdr.type != RPCAP_MSG_CLOSE) {
+			RPCAPD_LOG(NOTICE, "request 0x%02x before authentication",
+				hdr.type);
+			if (rpcap_discard(&ctrl, plen) < 0 ||
+			    rpcap_send_error(&ctrl, PCAP_ERR_AUTH,
+					     "not authenticated") < 0)
+				goto done;
+			continue;
+		}
+
 		switch (hdr.type) {
 		case RPCAP_MSG_AUTH_REQ:
 			/* libpcap treats a zero-length AUTH_REPLY as "version
 			 * 0 only, same byte order".
 			 */
-			if (handle_auth(&ctrl, plen) < 0)
+			if (handle_auth(&ctrl, plen, &s) < 0)
 				goto done;
 			break;
 		case RPCAP_MSG_FINDALLIF_REQ:
@@ -210,6 +292,7 @@ handle_client(int ctrl_fd)
 	}
 done:
 	stop_capture(&s);
+	tls_close(&ctrl);
 	close(ctrl_fd);
 	RPCAPD_LOG(NOTICE, "client %s disconnected", host);
 }
@@ -239,10 +322,10 @@ open_listen_socket(uint16_t port)
 
 	RPCAPD_LOG(NOTICE, "listening on %s port %u", host, listen_port);
 
-	if (!is_loopback(&listen_addr))
-		RPCAPD_LOG(WARNING,
-			"non-loopback address %s; "
-			"rpcap is unauthenticated and unencrypted, captured traffic is exposed to the network",
+	if (!is_loopback(&listen_addr) && !use_tls)
+		RPCAPD_LOG(ERR,
+			"listening on non-loopback address %s without TLS; "
+			"captured traffic will be exposed to the network, use -S",
 			host);
 
 	if (listen(fd, 1) < 0)
@@ -261,6 +344,12 @@ usage(FILE *f, const char *progname)
 		"  -4                    use only IPv4\n"
 		"  -6                    use only IPv6\n"
 		"  -N <ring size>        ring size in packets (default %u)\n"
+#ifdef RTE_HAS_OPENSSL
+		"  -S, --tls             encrypt connections with TLS (rpcaps://)\n"
+		"  -X, --cert <file>     server certificate chain, PEM (needs -S)\n"
+		"  -K, --key <file>      server private key, PEM (needs -S)\n"
+#endif
+		"  -n, --null-auth       permit unauthenticated remote clients\n"
 		"  -D, --debug           increase log verbosity (-D info, -DD debug)\n"
 		"      --debug-file <f>  redirect log output to file <f> (append mode)\n"
 		"      --send-timeout <s> seconds a data send may block before the\n"
@@ -271,9 +360,9 @@ usage(FILE *f, const char *progname)
 		"      --lcore=<core>    CPU core to run on (default: any)\n"
 		"      --file-prefix=<p> prefix to use for multi-process\n"
 		"\n"
-		"WARNING: rpcap is unauthenticated and unencrypted.  Binding to\n"
-		"any non-loopback address exposes captured traffic to the\n"
-		"network.  Not for production use.\n",
+		"Remote clients must authenticate with a system username and\n"
+		"password, and must use TLS to send it.  Loopback clients may\n"
+		"connect unauthenticated.  Not for production use.\n",
 		RPCAP_DEFAULT_NETPORT, DEFAULT_RING_SIZE,
 		DATA_SEND_TIMEOUT_SEC);
 }
@@ -297,6 +386,12 @@ parse_opts(int argc, char **argv)
 	static const struct option long_options[] = {
 		{ "port",         required_argument, NULL, 'p' },
 		{ "bind",         required_argument, NULL, 'b' },
+		{ "null-auth",    no_argument,       NULL, 'n' },
+#ifdef RTE_HAS_OPENSSL
+		{ "tls",          no_argument,       NULL, 'S' },
+		{ "cert",         required_argument, NULL, 'X' },
+		{ "key",          required_argument, NULL, 'K' },
+#endif
 		{ "debug",        no_argument,       NULL, 'D' },
 		{ "help",         no_argument,       NULL, 'h' },
 		{ "version",      no_argument,       NULL, OPT_VERSION },
@@ -308,8 +403,11 @@ parse_opts(int argc, char **argv)
 	};
 	int option_index, c;
 
-	while ((c = getopt_long(argc, argv, "hD46p:b:N:",
-				long_options, &option_index)) != -1) {
+	while ((c = getopt_long(argc, argv, "hnD46p:b:N:"
+#ifdef RTE_HAS_OPENSSL
+				"SX:K:"
+#endif
+				, long_options, &option_index)) != -1) {
 		switch (c) {
 		case 'p': {
 			unsigned long u = strtoul(optarg, NULL, 0);
@@ -328,6 +426,20 @@ parse_opts(int argc, char **argv)
 		case '6':
 			bind_family = AF_INET6;
 			break;
+		case 'n':
+			null_auth_ok = true;
+			break;
+#ifdef RTE_HAS_OPENSSL
+		case 'S':
+			use_tls = true;
+			break;
+		case 'X':
+			tls_certfile = optarg;
+			break;
+		case 'K':
+			tls_keyfile = optarg;
+			break;
+#endif
 		case 'N': {
 			unsigned long u = strtoul(optarg, NULL, 0);
 
@@ -394,6 +506,22 @@ parse_opts(int argc, char **argv)
 
 	/* Resolve the bind address now that -4/-6/-b have been seen. */
 	parse_bind_addr();
+
+	/* There is no sensible default for either: libpcap's rpcapd looks
+	 * for cert.pem and key.pem in the current directory, which is not
+	 * something a daemon started as root should do.
+	 */
+	if (use_tls && (tls_certfile == NULL || tls_keyfile == NULL))
+		rte_exit(EXIT_FAILURE,
+			 "TLS needs both a certificate (-X) and a private key (-K)\n");
+
+	if (!use_tls && (tls_certfile != NULL || tls_keyfile != NULL))
+		rte_exit(EXIT_FAILURE,
+			 "A certificate or key was given without -S\n");
+
+	if (null_auth_ok && !is_loopback(&listen_addr))
+		RPCAPD_LOG(ERR,
+			"-n allows any client that can reach this port to capture traffic");
 }
 
 /*
@@ -548,6 +676,10 @@ main(int argc, char **argv)
 	if (rte_eth_dev_count_avail() == 0)
 		rte_exit(EXIT_FAILURE, "No Ethernet ports found\n");
 
+	/* Fail here rather than on the first client's handshake. */
+	if (use_tls && tls_init(tls_certfile, tls_keyfile) < 0)
+		rte_exit(EXIT_FAILURE, "TLS setup failed\n");
+
 	sigaction(SIGTERM, &action, NULL);
 	sigaction(SIGINT, &action, NULL);
 
diff --git a/app/rpcapd/meson.build b/app/rpcapd/meson.build
index 61f4dc0a95..42b9eec854 100644
--- a/app/rpcapd/meson.build
+++ b/app/rpcapd/meson.build
@@ -14,12 +14,27 @@ if not dpdk_conf.has('RTE_HAS_LIBPCAP')
     subdir_done()
 endif
 
+# password authentication uses crypt(3)
+libcrypt_dep = cc.find_library('crypt', required: false)
+if not libcrypt_dep.found()
+    build = false
+    reason = 'missing dependency, "libcrypt"'
+    subdir_done()
+endif
+
 sources = files(
         'capture.c',
         'filter.c',
         'main.c',
         'session.c',
         'sock.c',
+        'tls.c',
 )
 ext_deps += pcap_dep
+ext_deps += libcrypt_dep
 deps += ['ethdev', 'pdump', 'bpf', 'pcapng']
+
+# TLS support is optional; tls.c builds as stubs without it
+if dpdk_conf.has('RTE_HAS_OPENSSL')
+    ext_deps += openssl_dep
+endif
diff --git a/app/rpcapd/rpcap-protocol.h b/app/rpcapd/rpcap-protocol.h
index 438fd8dd84..7fc1ec5db5 100644
--- a/app/rpcapd/rpcap-protocol.h
+++ b/app/rpcapd/rpcap-protocol.h
@@ -42,7 +42,10 @@
 #define RPCAP_MSG_STATS_REPLY	     (RPCAP_MSG_STATS_REQ	 | RPCAP_MSG_IS_REPLY)
 
 /* Error codes carried in the 'value' field of RPCAP_MSG_ERROR */
+#define PCAP_ERR_AUTH               3	/* generic authentication error */
 #define PCAP_ERR_WRONGVER          17
+#define PCAP_ERR_AUTH_FAILED       18	/* credentials were not accepted */
+#define PCAP_ERR_TLS_REQUIRED      19	/* server will only speak TLS */
 #define PCAP_ERR_AUTH_TYPE_NOTSUP  20
 
 /* Authentication types in rpcap_auth.type */
diff --git a/app/rpcapd/rpcapd.h b/app/rpcapd/rpcapd.h
index df38231bfb..3eea5c08e2 100644
--- a/app/rpcapd/rpcapd.h
+++ b/app/rpcapd/rpcapd.h
@@ -21,6 +21,7 @@
 struct rte_bpf_prm;
 struct rte_mempool;
 struct rte_ring;
+struct ssl_st;
 
 #define RTE_LOGTYPE_RPCAPD RTE_LOGTYPE_USER1
 #define RPCAPD_LOG(level, ...) \
@@ -36,9 +37,10 @@ struct rte_ring;
  */
 #define MAX_CAPTURE_LEN		(DEFAULT_SNAPLEN + 2 * sizeof(struct rte_vlan_hdr))
 
-/* A connection to the client. */
+/* A connection to the client; ssl is NULL when not encrypted. */
 struct conn {
 	int fd;
+	struct ssl_st *ssl;
 };
 
 /* Per-client capture session state. */
@@ -50,6 +52,7 @@ struct session {
 	uint32_t snaplen;
 	uint32_t npkt;				/* packet sequence for rpcap_pkthdr */
 	uint32_t pdump_flags;			/* direction bits handed to pdump */
+	bool     authenticated;			/* AUTH_REQ has succeeded */
 	bool     opened;			/* OPEN_REQ has selected a port */
 	bool     capture_on;
 	bool     promisc_set;			/* we enabled promiscuous mode */
@@ -64,6 +67,8 @@ extern RTE_ATOMIC(bool) quit_signal;
 /* Command-line settings needed outside of main.c */
 extern uint32_t ring_size;
 extern uint32_t send_timeout;		/* seconds; 0 means no limit */
+extern bool use_tls;			/* -S: encrypt both connections */
+extern bool null_auth_ok;		/* -n: permit null auth off loopback */
 
 /* Address the control socket is bound to; the data socket uses the same
  * address with an ephemeral port.
@@ -71,6 +76,8 @@ extern uint32_t send_timeout;		/* seconds; 0 means no limit */
 extern struct sockaddr_storage listen_addr;
 extern socklen_t               listen_addrlen;
 
+bool is_loopback(const struct sockaddr_storage *ss);
+
 /* sock.c: transport and message framing */
 int wait_readable(const struct conn *c, int timeout_ms);
 int accept_timeout(int listen_fd, int timeout_ms);
@@ -85,8 +92,17 @@ int rpcap_discard(const struct conn *c, uint32_t plen);
 void set_sockaddr_port(struct sockaddr_storage *ss, uint16_t port);
 uint16_t get_sockaddr_port(const struct sockaddr_storage *ss);
 
+/* tls.c: TLS transport, stubbed out when built without OpenSSL */
+int tls_init(const char *certfile, const char *keyfile);
+int tls_accept(struct conn *c);
+void tls_close(struct conn *c);
+int tls_send(struct ssl_st *ssl, const void *buf, size_t len);
+int tls_recv(struct ssl_st *ssl, void *buf, size_t len);
+bool tls_pending(const struct conn *c);
+void tls_reject_handshake(int fd);
+
 /* session.c: control requests handled before a capture starts */
-int handle_auth(const struct conn *c, uint32_t plen);
+int handle_auth(const struct conn *c, uint32_t plen, struct session *s);
 int handle_findallif(const struct conn *c);
 int handle_open(const struct conn *c, uint32_t plen, struct session *s);
 
diff --git a/app/rpcapd/session.c b/app/rpcapd/session.c
index cc14f26335..61600cb286 100644
--- a/app/rpcapd/session.c
+++ b/app/rpcapd/session.c
@@ -5,8 +5,13 @@
  * the interface list, and selecting an interface.
  */
 
+#include <crypt.h>
+#include <errno.h>
+#include <pwd.h>
+#include <shadow.h>
 #include <stdlib.h>
 #include <string.h>
+#include <unistd.h>
 
 #include <rte_byteorder.h>
 #include <rte_ethdev.h>
@@ -14,6 +19,12 @@
 #include "rpcap-protocol.h"
 #include "rpcapd.h"
 
+/* Slow down a client working through a password list. */
+#define AUTH_FAIL_DELAY_SEC	1
+
+/* Bound what a client can make us allocate for credentials. */
+#define MAX_CREDENTIAL_LEN	256
+
 /* Build and send the list of available DPDK ports. */
 int
 handle_findallif(const struct conn *c)
@@ -67,37 +78,182 @@ handle_findallif(const struct conn *c)
 }
 
 /*
- * AUTH_REQ: check the authentication type only.
- *
- * There is no credential store, so a username and password cannot be
- * verified; refuse them rather than reply that they were accepted.
+ * Check credentials against the system password database, as libpcap's
+ * rpcapd does.  Privileges are not dropped afterwards, since that would
+ * break the capture, so this authenticates without authorising.
+ * Returns 0 if the credentials are good.
+ */
+static int
+check_password(const char *user, const char *password)
+{
+	const struct passwd *pw;
+	const struct spwd *sp;
+	const char *hash;
+	char *result;
+
+	pw = getpwnam(user);
+	if (pw == NULL) {
+		RPCAPD_LOG(NOTICE, "authentication failed: no such user");
+		return -1;
+	}
+
+	/* The password database only holds a placeholder when the real
+	 * hash lives in the shadow file.
+	 */
+	sp = getspnam(user);
+	hash = (sp != NULL) ? sp->sp_pwdp : pw->pw_passwd;
+
+	/* Not a hash: the account is locked ('!' or '*') or has no
+	 * password.  Either way there is nothing to check against.
+	 */
+	if (hash == NULL || *hash != '$') {
+		RPCAPD_LOG(NOTICE,
+			   "authentication failed: account has no usable password "
+			   "(is /etc/shadow readable?)");
+		return -1;
+	}
+
+	errno = 0;
+	result = crypt(password, hash);
+	if (result == NULL) {
+		RPCAPD_LOG(ERR, "crypt failed: %s",
+			   errno != 0 ? strerror(errno) : "unknown error");
+		return -1;
+	}
+
+	if (strcmp(result, hash) != 0) {
+		RPCAPD_LOG(NOTICE, "authentication failed: wrong password");
+		return -1;
+	}
+
+	return 0;
+}
+
+/*
+ * Wipe a credential before freeing it.  explicit_bzero() because the
+ * compiler may drop a memset() before free() as a dead store.
+ */
+static void
+free_credential(char *cred)
+{
+	if (cred != NULL) {
+		explicit_bzero(cred, strlen(cred));
+		free(cred);
+	}
+}
+
+/* Read a length-prefixed credential out of the AUTH_REQ payload. */
+static int
+recv_credential(const struct conn *c, uint32_t len, uint32_t *plen, char **out)
+{
+	char *buf;
+
+	if (len > *plen || len > MAX_CREDENTIAL_LEN)
+		return -1;
+
+	buf = malloc(len + 1);
+	if (buf == NULL)
+		return -1;
+
+	if (recv_full(c, buf, len) < 0) {
+		explicit_bzero(buf, len);
+		free(buf);
+		return -1;
+	}
+	buf[len] = '\0';
+	*plen -= len;
+	*out = buf;
+	return 0;
+}
+
+/*
+ * AUTH_REQ: null authentication is accepted from a loopback peer only;
+ * a remote client needs a username and password, unless -n was given.
  */
 int
-handle_auth(const struct conn *c, uint32_t plen)
+handle_auth(const struct conn *c, uint32_t plen, struct session *s)
 {
+	char *user = NULL, *password = NULL;
 	struct rpcap_auth auth;
 	uint16_t type;
+	int rc;
+
+	s->authenticated = false;
 
 	if (plen < sizeof(auth)) {
 		rpcap_discard(c, plen);
-		return rpcap_send_error(c, 0, "short authentication request");
+		return rpcap_send_error(c, PCAP_ERR_AUTH, "short authentication request");
 	}
 
 	if (recv_full(c, &auth, sizeof(auth)) < 0)
 		return -1;
-
-	/* Discard any username and password that followed. */
-	if (rpcap_discard(c, plen - sizeof(auth)) < 0)
-		return -1;
+	plen -= sizeof(auth);
 
 	type = rte_be_to_cpu_16(auth.type);
-	if (type != RPCAP_RMTAUTH_NULL) {
+	switch (type) {
+	case RPCAP_RMTAUTH_NULL:
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
+
+		if (!is_loopback(&s->peer) && !null_auth_ok) {
+			RPCAPD_LOG(NOTICE,
+				   "rejecting null authentication from remote client");
+			return rpcap_send_error(c, PCAP_ERR_AUTH_FAILED,
+						"this server requires a username and "
+						"password for remote clients");
+		}
+		break;
+
+	case RPCAP_RMTAUTH_PWD:
+		if (recv_credential(c, rte_be_to_cpu_16(auth.slen1), &plen, &user) < 0 ||
+		    recv_credential(c, rte_be_to_cpu_16(auth.slen2), &plen, &password) < 0) {
+			free_credential(user);
+			return -1;
+		}
+
+		if (rpcap_discard(c, plen) < 0) {
+			free_credential(user);
+			free_credential(password);
+			return -1;
+		}
+
+		/* Refuse before checking, so a rejected password has not
+		 * already crossed the network in the clear.
+		 */
+		if (c->ssl == NULL && !is_loopback(&s->peer)) {
+			free_credential(user);
+			free_credential(password);
+			RPCAPD_LOG(NOTICE,
+				   "refusing password authentication on an unencrypted connection");
+			return rpcap_send_error(c, PCAP_ERR_AUTH_FAILED,
+						"this server will not accept a password "
+						"over an unencrypted connection; "
+						"use rpcaps://");
+		}
+
+		rc = check_password(user, password);
+		free_credential(user);
+		free_credential(password);
+
+		if (rc != 0) {
+			/* Delay a guess, and do not say which of the two
+			 * was wrong.
+			 */
+			sleep(AUTH_FAIL_DELAY_SEC);
+			return rpcap_send_error(c, PCAP_ERR_AUTH_FAILED,
+						"authentication failed");
+		}
+		break;
+
+	default:
+		if (rpcap_discard(c, plen) < 0)
+			return -1;
 		RPCAPD_LOG(NOTICE, "rejecting authentication type %u", type);
 		return rpcap_send_error(c, PCAP_ERR_AUTH_TYPE_NOTSUP,
-					"this server cannot check credentials; "
-					"connect without a username or password");
+					"authentication type not supported");
 	}
 
+	s->authenticated = true;
 	return rpcap_send_msg(c, RPCAP_MSG_AUTH_REPLY, 0, NULL, 0);
 }
 
diff --git a/app/rpcapd/sock.c b/app/rpcapd/sock.c
index 179c1a31c1..49c30c5370 100644
--- a/app/rpcapd/sock.c
+++ b/app/rpcapd/sock.c
@@ -18,6 +18,7 @@
 
 #include <rte_byteorder.h>
 #include <rte_common.h>
+#include <rte_mbuf.h>
 #include <rte_stdatomic.h>
 
 #include "rpcap-protocol.h"
@@ -59,6 +60,10 @@ wait_readable(const struct conn *c, int timeout_ms)
 {
 	struct pollfd pfd = { .fd = c->fd, .events = POLLIN };
 
+	/* Decrypted bytes buffered in the SSL object are invisible to poll() */
+	if (tls_pending(c))
+		return 1;
+
 	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
 		int wait_ms = POLL_INTERVAL_MS;
 		int rc;
@@ -207,7 +212,11 @@ recv_full(const struct conn *c, void *buf, size_t len)
 		if (wait_readable(c, -1) != 1)
 			return -1;
 
-		n = recv(c->fd, p, len, 0);
+		if (c->ssl != NULL)
+			n = tls_recv(c->ssl, p, len);
+		else
+			n = recv(c->fd, p, len, 0);
+
 		if (n < 0 && errno == EINTR)
 			continue;
 
@@ -220,6 +229,44 @@ recv_full(const struct conn *c, void *buf, size_t len)
 	return 0;
 }
 
+/*
+ * No scatter/gather write in TLS, and SSL_write() gives each call its
+ * own record, so gather into one buffer rather than paying record
+ * overhead per piece.
+ */
+static int
+send_iov_tls(struct ssl_st *ssl, const struct iovec *iov, int iovcnt)
+{
+	uint8_t buf[sizeof(struct rpcap_header) + sizeof(struct rpcap_pkthdr) +
+		    MAX_CAPTURE_LEN];
+	const uint8_t *p = buf;
+	size_t len = 0;
+	int i;
+
+	for (i = 0; i < iovcnt; i++) {
+		if (len + iov[i].iov_len > sizeof(buf)) {
+			/* Cannot happen: buf is sized for both headers plus
+			 * MAX_CAPTURE_LEN.
+			 */
+			RPCAPD_LOG(ERR, "message too large for TLS buffer");
+			errno = EMSGSIZE;
+			return -1;
+		}
+		memcpy(buf + len, iov[i].iov_base, iov[i].iov_len);
+		len += iov[i].iov_len;
+	}
+
+	while (len > 0) {
+		int n = tls_send(ssl, p, len);
+
+		if (n <= 0)
+			return -1;
+		p += n;
+		len -= n;
+	}
+	return 0;
+}
+
 /*
  * Send all of iov, resending the remainder if sendmsg() reports a short
  * count (possible when the connection breaks or a signal arrives after
@@ -233,6 +280,9 @@ send_iov_full(const struct conn *c, struct iovec *iov, int iovcnt, int flags)
 		.msg_iovlen = iovcnt,
 	};
 
+	if (c->ssl != NULL)
+		return send_iov_tls(c->ssl, iov, iovcnt);
+
 	while (msg.msg_iovlen > 0) {
 		ssize_t n = sendmsg(c->fd, &msg, flags | MSG_NOSIGNAL);
 
diff --git a/app/rpcapd/tls.c b/app/rpcapd/tls.c
new file mode 100644
index 0000000000..f9671f63b4
--- /dev/null
+++ b/app/rpcapd/tls.c
@@ -0,0 +1,273 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Stephen Hemminger
+ *
+ * TLS transport for the rpcaps:// scheme.  Both the control and the
+ * data connection are promoted.  Built as stubs when DPDK was
+ * configured without OpenSSL.
+ */
+
+#include <errno.h>
+#include <stdbool.h>
+#include <stddef.h>
+#include <stdint.h>
+#include <string.h>
+#include <sys/socket.h>
+#include <sys/time.h>
+#include <unistd.h>
+
+#include "rpcapd.h"
+
+/* How long a peer may take to complete a handshake. */
+#define TLS_HANDSHAKE_TIMEOUT_SEC	10
+
+#ifdef RTE_HAS_OPENSSL
+
+#include <openssl/err.h>
+#include <openssl/ssl.h>
+
+static SSL_CTX *tls_ctx;
+
+static const char *
+tls_strerror(void)
+{
+	unsigned long e = ERR_get_error();
+
+	return (e != 0) ? ERR_reason_error_string(e) : "unknown error";
+}
+
+/*
+ * Build the server context at startup, so a bad certificate fails here
+ * rather than on the first client's handshake.
+ */
+int
+tls_init(const char *certfile, const char *keyfile)
+{
+	tls_ctx = SSL_CTX_new(TLS_server_method());
+	if (tls_ctx == NULL) {
+		RPCAPD_LOG(ERR, "cannot create TLS context: %s", tls_strerror());
+		return -1;
+	}
+
+	if (SSL_CTX_set_min_proto_version(tls_ctx, TLS1_2_VERSION) != 1) {
+		RPCAPD_LOG(ERR, "cannot set minimum TLS version: %s", tls_strerror());
+		return -1;
+	}
+
+	/* Hides a renegotiation from SSL_read()/SSL_write(). */
+	SSL_CTX_set_mode(tls_ctx, SSL_MODE_AUTO_RETRY);
+
+	if (SSL_CTX_use_certificate_chain_file(tls_ctx, certfile) != 1) {
+		RPCAPD_LOG(ERR, "cannot read certificate file '%s': %s",
+			   certfile, tls_strerror());
+		return -1;
+	}
+
+	if (SSL_CTX_use_PrivateKey_file(tls_ctx, keyfile, SSL_FILETYPE_PEM) != 1) {
+		RPCAPD_LOG(ERR, "cannot read private key file '%s': %s",
+			   keyfile, tls_strerror());
+		return -1;
+	}
+
+	if (SSL_CTX_check_private_key(tls_ctx) != 1) {
+		RPCAPD_LOG(ERR, "private key '%s' does not match certificate '%s'",
+			   keyfile, certfile);
+		return -1;
+	}
+
+	return 0;
+}
+
+/*
+ * SSL_accept() on a blocking socket waits indefinitely, and only one
+ * client is served at a time, so bound the handshake with socket
+ * timeouts.
+ */
+static int
+set_handshake_timeout(int fd, time_t seconds)
+{
+	struct timeval tv = { .tv_sec = seconds };
+
+	if (setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv)) < 0 ||
+	    setsockopt(fd, SOL_SOCKET, SO_SNDTIMEO, &tv, sizeof(tv)) < 0) {
+		RPCAPD_LOG(NOTICE, "cannot set TLS handshake timeout: %s",
+			   strerror(errno));
+		return -1;
+	}
+	return 0;
+}
+
+int
+tls_accept(struct conn *c)
+{
+	SSL *ssl = SSL_new(tls_ctx);
+	bool timed = set_handshake_timeout(c->fd, TLS_HANDSHAKE_TIMEOUT_SEC) == 0;
+
+	if (ssl == NULL) {
+		RPCAPD_LOG(ERR, "SSL_new: %s", tls_strerror());
+		return -1;
+	}
+
+	if (SSL_set_fd(ssl, c->fd) != 1) {
+		RPCAPD_LOG(ERR, "SSL_set_fd: %s", tls_strerror());
+		SSL_free(ssl);
+		return -1;
+	}
+
+	if (SSL_accept(ssl) != 1) {
+		/* A timeout surfaces as a syscall error on the read. */
+		if (errno == EAGAIN || errno == EWOULDBLOCK)
+			RPCAPD_LOG(ERR, "TLS handshake timed out after %u seconds",
+				   TLS_HANDSHAKE_TIMEOUT_SEC);
+		else
+			RPCAPD_LOG(ERR, "TLS handshake failed: %s", tls_strerror());
+		SSL_free(ssl);
+		return -1;
+	}
+
+	/* Back to blocking for the session. */
+	if (timed)
+		set_handshake_timeout(c->fd, 0);
+
+	RPCAPD_LOG(DEBUG, "TLS established: %s %s",
+		   SSL_get_version(ssl), SSL_get_cipher(ssl));
+	c->ssl = ssl;
+	return 0;
+}
+
+/* Send the close_notify alert so the client does not report a truncated
+ * stream.  The caller still owns the socket.
+ */
+void
+tls_close(struct conn *c)
+{
+	if (c->ssl == NULL)
+		return;
+
+	SSL_shutdown(c->ssl);
+	SSL_free(c->ssl);
+	c->ssl = NULL;
+}
+
+/*
+ * Map an SSL error onto the send()/recv() contract the callers expect:
+ * byte count on success, -1 with errno set on failure.
+ */
+static int
+tls_error(SSL *ssl, int ret, const char *what)
+{
+	int err = SSL_get_error(ssl, ret);
+
+	switch (err) {
+	case SSL_ERROR_ZERO_RETURN:
+		/* Clean shutdown by the peer: an orderly EOF. */
+		return 0;
+	case SSL_ERROR_SYSCALL:
+		/* errno is already set, unless the peer just vanished. */
+		if (errno == 0)
+			errno = ECONNRESET;
+		return -1;
+	case SSL_ERROR_WANT_READ:
+	case SSL_ERROR_WANT_WRITE:
+		errno = EAGAIN;
+		return -1;
+	default:
+		RPCAPD_LOG(DEBUG, "%s: %s", what, tls_strerror());
+		errno = EPROTO;
+		return -1;
+	}
+}
+
+int
+tls_send(struct ssl_st *ssl, const void *buf, size_t len)
+{
+	int ret = SSL_write(ssl, buf, len);
+
+	if (ret > 0)
+		return ret;
+	return tls_error(ssl, ret, "SSL_write");
+}
+
+int
+tls_recv(struct ssl_st *ssl, void *buf, size_t len)
+{
+	int ret = SSL_read(ssl, buf, len);
+
+	if (ret > 0)
+		return ret;
+	return tls_error(ssl, ret, "SSL_read");
+}
+
+/*
+ * One TLS record can hold several rpcap messages, and once read off the
+ * socket the rest sit in the SSL object where poll() cannot see them.
+ * Every wait must check this first.
+ */
+bool
+tls_pending(const struct conn *c)
+{
+	return c->ssl != NULL && SSL_pending(c->ssl) > 0;
+}
+
+#else /* !RTE_HAS_OPENSSL */
+
+int
+tls_init(const char *certfile __rte_unused, const char *keyfile __rte_unused)
+{
+	RPCAPD_LOG(ERR, "built without OpenSSL, TLS is not available");
+	return -1;
+}
+
+int
+tls_accept(struct conn *c __rte_unused)
+{
+	return -1;
+}
+
+void
+tls_close(struct conn *c __rte_unused)
+{
+}
+
+int
+tls_send(struct ssl_st *ssl __rte_unused, const void *buf __rte_unused,
+	 size_t len __rte_unused)
+{
+	errno = ENOTSUP;
+	return -1;
+}
+
+int
+tls_recv(struct ssl_st *ssl __rte_unused, void *buf __rte_unused,
+	 size_t len __rte_unused)
+{
+	errno = ENOTSUP;
+	return -1;
+}
+
+bool
+tls_pending(const struct conn *c __rte_unused)
+{
+	return false;
+}
+
+#endif /* RTE_HAS_OPENSSL */
+
+/*
+ * Turn away a handshake from a daemon without -S.  Written straight to
+ * the socket since there is no SSL context to generate it with.
+ */
+void
+tls_reject_handshake(int fd)
+{
+	static const uint8_t alert[] = {
+		21,	/* content type: alert */
+		3, 3,	/* legacy record version: TLS 1.2 */
+		0, 2,	/* payload length */
+		2,	/* level: fatal */
+		40,	/* description: handshake_failure */
+	};
+
+	RPCAPD_LOG(WARNING, "rejecting TLS handshake: server is not using TLS");
+	if (write(fd, alert, sizeof(alert)) != (ssize_t)sizeof(alert))
+		RPCAPD_LOG(DEBUG, "could not send TLS alert: %s", strerror(errno));
+}
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 8e107b48b6..b1ecaa3574 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -147,6 +147,8 @@ New Features
 
   Added the ``dpdk-rpcapd`` application, which implements the rpcap
   protocol to allow live capture in tcpdump and Wireshark.
+  Remote clients can be authenticated with a system username and password,
+  and connections encrypted with TLS when built with OpenSSL.
 
 
 Removed Items
diff --git a/doc/guides/tools/rpcapd.rst b/doc/guides/tools/rpcapd.rst
index a8b026a409..f5d292682c 100644
--- a/doc/guides/tools/rpcapd.rst
+++ b/doc/guides/tools/rpcapd.rst
@@ -18,19 +18,21 @@ the libpcap project's ``rpcapd``.
 See
 https://github.com/the-tcpdump-group/libpcap/tree/master/rpcapd
 for the reference implementation.
-Clients connect to ``dpdk-rpcapd`` using a ``rpcap://`` URL,
+Clients connect to ``dpdk-rpcapd`` using a ``rpcap://`` URL, or
+``rpcaps://`` for a TLS-encrypted connection,
 request the list of available interfaces(which are the ports of the DPDK primary),
 open one, and stream packets from it.
 
 .. warning::
 
-   ``dpdk-rpcapd`` listens on an unauthenticated, unencrypted TCP port
-   (default 2002).  Anyone able to reach the port can list DPDK ports
-   and capture all traffic flowing through them.  The default bind
-   address is ``127.0.0.1``, so the listener is not reachable from
-   other hosts; overriding this with ``--bind`` exposes captured
-   traffic to anyone who can reach that address.  **Do not run
-   ``dpdk-rpcapd`` on a production system.**
+   Anyone who can authenticate to ``dpdk-rpcapd`` can capture all
+   traffic flowing through the ports of the DPDK primary process.  The
+   default bind address is ``127.0.0.1``, so the listener is not
+   reachable from other hosts.  A client on the loopback address may
+   connect without credentials; a client from any other address must
+   authenticate with a system username and password and must use TLS to
+   send it, unless ``-n`` was given.  See `Authentication`_ and `TLS`_.
+   **Do not run ``dpdk-rpcapd`` on a production system.**
 
 
 Running the Application
@@ -63,6 +65,27 @@ The application has a small set of command-line options:
     Size of the per-session capture ring in packets.  Default is 2048.
     Rounded up to a power of two if necessary.
 
+*   ``-S``, ``--tls``
+
+    Encrypt both the control and the data connection with TLS.  Clients
+    must then use a ``rpcaps://`` URL.  Requires ``-X`` and ``-K``.
+    Available only when DPDK was built with OpenSSL.
+
+*   ``-X <file>``, ``--cert <file>``
+
+    Server certificate chain in PEM format.  Only meaningful with
+    ``-S``, and required by it.
+
+*   ``-K <file>``, ``--key <file>``
+
+    Server private key in PEM format.  Required with ``-S``; there is
+    no default.
+
+*   ``-n``, ``--null-auth``
+
+    Permit null authentication from any address, not just loopback.
+    Usually used with ``-l``.
+
 *   ``-D``, ``--debug``
 
     Increase log verbosity.  A single ``-D`` adds informational
@@ -102,6 +125,64 @@ secondary process and does not need EAL options on its command line for
 typical use.
 
 
+Authentication
+--------------
+
+A client on the loopback address may connect without credentials, which
+is what a ``rpcap://`` URL with no userinfo does.
+
+A client from any other address must supply a system username and
+password, checked against the host password database as the reference
+``rpcapd`` does.  Accounts without a usable password hash, such as
+locked accounts, are refused.
+
+A password is only accepted over an encrypted connection, so remote
+password authentication requires ``-S`` as well.  A password sent in
+the clear is refused without being checked.
+
+Credentials are checked but no privileges are dropped, so this
+authenticates a client without authorising it: any account that can log
+in has the same access to every port of the primary process.
+
+``-n`` waives the check and lets any client connect unauthenticated,
+from any address.
+
+
+TLS
+---
+
+``-S`` encrypts both the control and the data connection, and is
+available only when DPDK was built with OpenSSL.
+
+TLS is not negotiated in the rpcap protocol: the client decides from
+its URL scheme and the daemon from ``-S``, so the two have to be
+configured to agree.  A mismatch is reported rather than left to fail
+as a protocol error.
+
+A certificate and key can be generated for testing with:
+
+.. code-block:: console
+
+    openssl req -x509 -newkey rsa:2048 -nodes -days 30 \
+        -keyout key.pem -out cert.pem -subj /CN=localhost
+
+Start the daemon with them:
+
+.. code-block:: console
+
+    sudo ./<build_dir>/app/dpdk-rpcapd -S -X cert.pem -K key.pem
+
+Then connect with a ``rpcaps://`` URL:
+
+.. code-block:: console
+
+    sudo /usr/local/sbin/tcpdump -i rpcaps://localhost:2002/net_tap0 -nn -c 20
+
+A client does not validate a self-signed certificate unless told to
+trust it, so the connection is encrypted but the server is not
+authenticated.
+
+
 Client Setup
 ------------
 
@@ -174,17 +255,17 @@ in this initial version:
     Subsequent clients are queued by the listening socket but not
     serviced until the first disconnects.
 
-*   **No authentication.** Password authentication is refused with
-    ``PCAP_ERR_AUTH_TYPE_NOTSUP``; clients must connect without
-    credentials, which is what a ``rpcap://`` URL with no userinfo does.
-    With the default loopback bind, reaching the port already requires
-    an account on the host.
-
-*   **No TLS.** The ``-S`` option of the reference ``rpcapd`` is not
-    implemented, so the connection is always in the clear.  This is
-    reasonable for the default loopback bind, where the traffic never
-    leaves the host, but means ``--bind`` to any other address sends
-    captured packets over the network unencrypted.
+*   **TLS needs OpenSSL.** ``-S`` is only available when DPDK was built
+    with OpenSSL support.
+
+*   **Authentication does not restrict access.** Credentials are
+    checked, but the daemon keeps the root privileges it needs for
+    ``pdump`` instead of dropping to the authenticated user, so every
+    account that can log in has the same access to every port.
+
+*   **No client certificates.** TLS authenticates the server to the
+    client and encrypts the connection; the client is identified only
+    by its password.
 
 *   **TCP data transport only.** A client requesting UDP is refused.
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v4 4/4] app/rpcapd: add host list option
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
                     ` (2 preceding siblings ...)
  2026-10-01  2:35   ` [PATCH v4 3/4] app/rpcapd: add TLS support Stephen Hemminger
@ 2026-10-01  2:35   ` Stephen Hemminger
  2026-10-01 18:50   ` [PATCH v4 0/4] add rpcap remote capture daemon Marat Khalili
  4 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01  2:35 UTC (permalink / raw)
  To: dev; +Cc: Stephen Hemminger, Reshma Pattan

Add -l to restrict which clients may connect, as libpcap's rpcapd
does.  The argument is the list itself rather than a file naming one,
with hosts separated by a comma, semicolon or space, so a list can be
copied between the two daemons.

Names are resolved once at startup, so DNS stays out of the accept
path and an unresolvable name fails at startup rather than when a
client first arrives.  Every address a name resolves to is accepted.

The list is checked after any TLS handshake, since a client using
rpcaps:// can only read an error sent inside the session.  Only the
control connection is checked: the data connection is already
restricted to the address the control connection came from.

This bounds which hosts may connect but does not identify them, so it
complements the password authentication rather than replacing it.

Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
---
 app/rpcapd/main.c                      | 135 ++++++++++++++++++++++++-
 app/rpcapd/rpcap-protocol.h            |   1 +
 app/rpcapd/rpcapd.h                    |   1 +
 app/rpcapd/sock.c                      |   2 +-
 doc/guides/rel_notes/release_26_11.rst |   1 +
 doc/guides/tools/rpcapd.rst            |   8 ++
 6 files changed, 145 insertions(+), 3 deletions(-)

diff --git a/app/rpcapd/main.c b/app/rpcapd/main.c
index 0cf6c3f4ba..e354623cf0 100644
--- a/app/rpcapd/main.c
+++ b/app/rpcapd/main.c
@@ -66,6 +66,11 @@ bool use_tls;				/* -S */
 bool null_auth_ok;			/* -n */
 static const char *tls_certfile;	/* -X argument */
 static const char *tls_keyfile;		/* -K argument */
+static const char *host_list;		/* -l argument; NULL means any host */
+
+/* -l resolved at startup; empty means no restriction. */
+static struct sockaddr_storage *allowed_hosts;
+static unsigned int num_allowed_hosts;
 
 struct sockaddr_storage listen_addr;
 socklen_t               listen_addrlen;
@@ -94,6 +99,114 @@ is_loopback(const struct sockaddr_storage *ss)
 	return false;
 }
 
+/* Separators in the -l argument, same set as libpcap's rpcapd. */
+#define HOST_LIST_SEP " ,;"
+
+/*
+ * A dual-stack socket reports a v4 peer as ::ffff:a.b.c.d, which
+ * same_host() will not match against an AF_INET entry.  Normalize both
+ * the peer and the list, since a resolver can also return a mapped
+ * address.
+ */
+static void
+unmap_v4(struct sockaddr_storage *ss)
+{
+	const struct sockaddr_in6 *sin6 = (const void *)ss;
+	struct sockaddr_in sin = {
+		.sin_family = AF_INET,
+		.sin_port   = sin6->sin6_port,
+	};
+
+	if (ss->ss_family != AF_INET6 || !IN6_IS_ADDR_V4MAPPED(&sin6->sin6_addr))
+		return;
+
+	memcpy(&sin.sin_addr, &sin6->sin6_addr.s6_addr[12], sizeof(sin.sin_addr));
+	memset(ss, 0, sizeof(*ss));
+	memcpy(ss, &sin, sizeof(sin));
+}
+
+/*
+ * Resolve the -l list at startup, so an unresolvable name fails here
+ * rather than when a client connects.  A name can have several
+ * addresses; all of them are accepted.
+ */
+static void
+parse_host_list(void)
+{
+	static const struct addrinfo hints = {
+		.ai_family   = AF_UNSPEC,
+		.ai_socktype = SOCK_STREAM,
+	};
+	char *copy, *token, *saveptr;
+
+	if (host_list == NULL)
+		return;
+
+	copy = strdup(host_list);
+	if (copy == NULL)
+		rte_exit(EXIT_FAILURE, "Cannot copy host list: %s\n", strerror(errno));
+
+	for (token = strtok_r(copy, HOST_LIST_SEP, &saveptr); token != NULL;
+	     token = strtok_r(NULL, HOST_LIST_SEP, &saveptr)) {
+		struct addrinfo *res, *ai;
+		unsigned int n = 0;
+		void *tmp;
+		int rc;
+
+		rc = getaddrinfo(token, NULL, &hints, &res);
+		if (rc != 0)
+			rte_exit(EXIT_FAILURE, "Invalid host '%s' in host list: %s\n",
+				 token, gai_strerror(rc));
+
+		for (ai = res; ai != NULL; ai = ai->ai_next)
+			n++;
+
+		tmp = realloc(allowed_hosts,
+			      (num_allowed_hosts + n) * sizeof(*allowed_hosts));
+		if (tmp == NULL)
+			rte_exit(EXIT_FAILURE, "Cannot grow host list: %s\n",
+				 strerror(errno));
+		allowed_hosts = tmp;
+
+		for (ai = res; ai != NULL; ai = ai->ai_next) {
+			struct sockaddr_storage *slot =
+				&allowed_hosts[num_allowed_hosts++];
+
+			memset(slot, 0, sizeof(*slot));
+			memcpy(slot, ai->ai_addr, ai->ai_addrlen);
+			unmap_v4(slot);
+		}
+
+		freeaddrinfo(res);
+	}
+	free(copy);
+
+	/* An empty list is a typo, not a request to allow everyone. */
+	if (num_allowed_hosts == 0)
+		rte_exit(EXIT_FAILURE, "Host list '%s' contains no hosts\n", host_list);
+}
+
+/* Only the control connection is checked; accept_from() pins the data
+ * connection to the same peer.
+ */
+static bool
+host_allowed(const struct sockaddr_storage *ss)
+{
+	struct sockaddr_storage peer = *ss;
+	unsigned int i;
+
+	if (num_allowed_hosts == 0)
+		return true;
+
+	unmap_v4(&peer);
+
+	for (i = 0; i < num_allowed_hosts; i++)
+		if (same_host(&peer, &allowed_hosts[i]))
+			return true;
+
+	return false;
+}
+
 static void
 parse_bind_addr(void)
 {
@@ -211,6 +324,17 @@ handle_client(int ctrl_fd)
 	if (setup_tls(&ctrl) < 0)
 		goto done;
 
+	/* After the handshake: a TLS client can only read an error sent
+	 * inside the session.
+	 */
+	if (!host_allowed(&peer)) {
+		RPCAPD_LOG(WARNING, "rejected client %s: not in the allowed host list",
+			host);
+		rpcap_send_error(&ctrl, PCAP_ERR_HOSTNOAUTH,
+				 "this host is not allowed to connect to this server");
+		goto done;
+	}
+
 	while (!rte_atomic_load_explicit(&quit_signal, rte_memory_order_relaxed)) {
 		struct rpcap_header hdr;
 		uint32_t plen;
@@ -343,6 +467,8 @@ usage(FILE *f, const char *progname)
 		"  -b, --bind <addr>     bind address (default 127.0.0.1, ::1 with -6)\n"
 		"  -4                    use only IPv4\n"
 		"  -6                    use only IPv6\n"
+		"  -l, --hosts <list>    only accept clients from these hosts,\n"
+		"                        separated by ',' ';' or space\n"
 		"  -N <ring size>        ring size in packets (default %u)\n"
 #ifdef RTE_HAS_OPENSSL
 		"  -S, --tls             encrypt connections with TLS (rpcaps://)\n"
@@ -386,6 +512,7 @@ parse_opts(int argc, char **argv)
 	static const struct option long_options[] = {
 		{ "port",         required_argument, NULL, 'p' },
 		{ "bind",         required_argument, NULL, 'b' },
+		{ "hosts",        required_argument, NULL, 'l' },
 		{ "null-auth",    no_argument,       NULL, 'n' },
 #ifdef RTE_HAS_OPENSSL
 		{ "tls",          no_argument,       NULL, 'S' },
@@ -403,7 +530,7 @@ parse_opts(int argc, char **argv)
 	};
 	int option_index, c;
 
-	while ((c = getopt_long(argc, argv, "hnD46p:b:N:"
+	while ((c = getopt_long(argc, argv, "hnD46p:b:l:N:"
 #ifdef RTE_HAS_OPENSSL
 				"SX:K:"
 #endif
@@ -420,6 +547,9 @@ parse_opts(int argc, char **argv)
 		case 'b':
 			bind_addr = optarg;
 			break;
+		case 'l':
+			host_list = optarg;
+			break;
 		case '4':
 			bind_family = AF_INET;
 			break;
@@ -506,6 +636,7 @@ parse_opts(int argc, char **argv)
 
 	/* Resolve the bind address now that -4/-6/-b have been seen. */
 	parse_bind_addr();
+	parse_host_list();
 
 	/* There is no sensible default for either: libpcap's rpcapd looks
 	 * for cert.pem and key.pem in the current directory, which is not
@@ -519,7 +650,7 @@ parse_opts(int argc, char **argv)
 		rte_exit(EXIT_FAILURE,
 			 "A certificate or key was given without -S\n");
 
-	if (null_auth_ok && !is_loopback(&listen_addr))
+	if (null_auth_ok && !is_loopback(&listen_addr) && num_allowed_hosts == 0)
 		RPCAPD_LOG(ERR,
 			"-n allows any client that can reach this port to capture traffic");
 }
diff --git a/app/rpcapd/rpcap-protocol.h b/app/rpcapd/rpcap-protocol.h
index 7fc1ec5db5..787189eecc 100644
--- a/app/rpcapd/rpcap-protocol.h
+++ b/app/rpcapd/rpcap-protocol.h
@@ -43,6 +43,7 @@
 
 /* Error codes carried in the 'value' field of RPCAP_MSG_ERROR */
 #define PCAP_ERR_AUTH               3	/* generic authentication error */
+#define PCAP_ERR_HOSTNOAUTH        10	/* peer address is not in the allowed list */
 #define PCAP_ERR_WRONGVER          17
 #define PCAP_ERR_AUTH_FAILED       18	/* credentials were not accepted */
 #define PCAP_ERR_TLS_REQUIRED      19	/* server will only speak TLS */
diff --git a/app/rpcapd/rpcapd.h b/app/rpcapd/rpcapd.h
index 3eea5c08e2..aa4d01dba5 100644
--- a/app/rpcapd/rpcapd.h
+++ b/app/rpcapd/rpcapd.h
@@ -79,6 +79,7 @@ extern socklen_t               listen_addrlen;
 bool is_loopback(const struct sockaddr_storage *ss);
 
 /* sock.c: transport and message framing */
+bool same_host(const struct sockaddr_storage *a, const struct sockaddr_storage *b);
 int wait_readable(const struct conn *c, int timeout_ms);
 int accept_timeout(int listen_fd, int timeout_ms);
 int accept_from(int listen_fd, const struct sockaddr_storage *want,
diff --git a/app/rpcapd/sock.c b/app/rpcapd/sock.c
index 49c30c5370..0571e6c759 100644
--- a/app/rpcapd/sock.c
+++ b/app/rpcapd/sock.c
@@ -115,7 +115,7 @@ accept_timeout(int listen_fd, int timeout_ms)
 /* Compare the host part of two addresses, ignoring the port: the data
  * connection comes from an ephemeral port, not the control one.
  */
-static bool
+bool
 same_host(const struct sockaddr_storage *a, const struct sockaddr_storage *b)
 {
 	if (a->ss_family != b->ss_family)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index b1ecaa3574..9e35d5881a 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -149,6 +149,7 @@ New Features
   protocol to allow live capture in tcpdump and Wireshark.
   Remote clients can be authenticated with a system username and password,
   and connections encrypted with TLS when built with OpenSSL.
+  Access can also be restricted to a list of allowed client hosts.
 
 
 Removed Items
diff --git a/doc/guides/tools/rpcapd.rst b/doc/guides/tools/rpcapd.rst
index f5d292682c..9e041e3a0f 100644
--- a/doc/guides/tools/rpcapd.rst
+++ b/doc/guides/tools/rpcapd.rst
@@ -60,6 +60,14 @@ The application has a small set of command-line options:
     Use only IPv6; an IPv4 argument to ``-b`` is rejected.  The default
     bind address becomes ``::1``.
 
+*   ``-l <host_list>``, ``--hosts <host_list>``
+
+    Only allow the hosts in ``<host_list>`` to connect.  The list is
+    host names or addresses separated by commas, semicolons or spaces,
+    and includes loopback, so ``127.0.0.1`` must be listed for a local
+    client.  Names are resolved at startup.  By default any host that
+    can reach the port may connect.
+
 *   ``-N <ring_size>``
 
     Size of the per-session capture ring in packets.  Default is 2048.
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH v4 2/4] app/rpcapd: remote pcap daemon
  2026-10-01  2:35   ` [PATCH v4 2/4] app/rpcapd: remote pcap daemon Stephen Hemminger
@ 2026-10-01 18:50     ` Marat Khalili
  0 siblings, 0 replies; 21+ messages in thread
From: Marat Khalili @ 2026-10-01 18:50 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: Thomas Monjalon, Reshma Pattan, dev

Two non-critical nits inline.

On 01/10/2026 03:35, Stephen Hemminger wrote:
> +static void
> +parse_bind_addr(void)
> +{
> +	struct addrinfo hints = {
> +		.ai_family   = bind_family,
> +		.ai_socktype = SOCK_STREAM,
> +		.ai_flags    = AI_NUMERICHOST | AI_PASSIVE,
> +	};
> +	struct addrinfo *res;
> +	int rc;
> +
> +	/* Loopback by default; the wildcard address is not a safe default. */
> +	if (bind_addr == NULL)
> +		bind_addr = (bind_family == AF_INET6) ? "::1" : "127.0.0.1";
When neither `bind_addr` nor `bind_family` are specified, can we bind to 
all local addresses (127.0.0.1, ::1) that exist in the system? Or at 
least to the first of them that exists, to improve the user experience 
on IPv6-only systems a little?
> +
> +	rc = getaddrinfo(bind_addr, NULL, &hints, &res);
> +	if (rc != 0)
> +		rte_exit(EXIT_FAILURE, "Invalid bind address '%s': %s\n",
> +			 bind_addr, gai_strerror(rc));
> +	memcpy(&listen_addr, res->ai_addr, res->ai_addrlen);
> +	listen_addrlen = res->ai_addrlen;
> +	freeaddrinfo(res);
> +}
// ...
> +/* Largest snaplen a client can be given. */
> +#define DEFAULT_SNAPLEN		RTE_MBUF_DEFAULT_DATAROOM

The naming is not optimal, it's upper limit not default.


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v4 0/4] add rpcap remote capture daemon
  2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
                     ` (3 preceding siblings ...)
  2026-10-01  2:35   ` [PATCH v4 4/4] app/rpcapd: add host list option Stephen Hemminger
@ 2026-10-01 18:50   ` Marat Khalili
  2026-10-01 23:00     ` Stephen Hemminger
  4 siblings, 1 reply; 21+ messages in thread
From: Marat Khalili @ 2026-10-01 18:50 UTC (permalink / raw)
  To: Stephen Hemminger, dev

>    pcapng: add API to read back capture mbuf header
Analysis with AI tells that there are some opportunities to make the 
fast path of lib/pcapng both lighter and less .pcapng-centric to make 
life of consumers that prefer other formats easier. It does not have to 
(probably even shouldn't) be a part of the current patchset though.
>    app/rpcapd: remote pcap daemon
Ideally forwarding captured data could run in a separate thread with 
minimum unnecessary syscalls, although it's hard to tell without 
benchmarks how much current design is slowed down by them.
>    app/rpcapd: add TLS support
This patch simultaneously adds authentication, but it's only required 
for non-loopback users. This is a common security vulnerability, not all 
loopback users can be trusted. So if we care about authentication at 
all, it should probably also be required for localhost users, at least 
by default.
>    app/rpcapd: add host list option
>
>   MAINTAINERS                            |   2 +
>   app/meson.build                        |   1 +
>   app/rpcapd/capture.c                   | 541 ++++++++++++++++
>   app/rpcapd/filter.c                    | 164 +++++
>   app/rpcapd/main.c                      | 840 +++++++++++++++++++++++++
>   app/rpcapd/meson.build                 |  40 ++
>   app/rpcapd/rpcap-protocol.h            | 146 +++++
>   app/rpcapd/rpcapd.h                    | 122 ++++
>   app/rpcapd/session.c                   | 292 +++++++++
>   app/rpcapd/sock.c                      | 365 +++++++++++
>   app/rpcapd/tls.c                       | 273 ++++++++
>   app/test/test_pcapng.c                 | 133 ++++
>   doc/guides/rel_notes/release_26_11.rst |  13 +
>   doc/guides/tools/index.rst             |   1 +
>   doc/guides/tools/rpcapd.rst            | 288 +++++++++
>   lib/pcapng/rte_pcapng.c                |  40 ++
>   lib/pcapng/rte_pcapng.h                |  46 ++
>   17 files changed, 3307 insertions(+)
>   create mode 100644 app/rpcapd/capture.c
>   create mode 100644 app/rpcapd/filter.c
>   create mode 100644 app/rpcapd/main.c
>   create mode 100644 app/rpcapd/meson.build
>   create mode 100644 app/rpcapd/rpcap-protocol.h
>   create mode 100644 app/rpcapd/rpcapd.h
>   create mode 100644 app/rpcapd/session.c
>   create mode 100644 app/rpcapd/sock.c
>   create mode 100644 app/rpcapd/tls.c
>   create mode 100644 doc/guides/tools/rpcapd.rst

With or without addressing any of the comments above,

Series-Acked-by: Marat Khalili <qm2k21@gmail.com>


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v4 0/4] add rpcap remote capture daemon
  2026-10-01 18:50   ` [PATCH v4 0/4] add rpcap remote capture daemon Marat Khalili
@ 2026-10-01 23:00     ` Stephen Hemminger
  0 siblings, 0 replies; 21+ messages in thread
From: Stephen Hemminger @ 2026-10-01 23:00 UTC (permalink / raw)
  To: Marat Khalili; +Cc: dev

On Thu, 1 Oct 2026 19:50:34 +0100
Marat Khalili <qm2k21@gmail.com> wrote:

> >    pcapng: add API to read back capture mbuf header  
> Analysis with AI tells that there are some opportunities to make the 
> fast path of lib/pcapng both lighter and less .pcapng-centric to make 
> life of consumers that prefer other formats easier. It does not have to 
> (probably even shouldn't) be a part of the current patchset though.

Thought about that, and also looked at making rpcap another library
but it ends up with DPDK reimplementing libpcap...

The other option would be using dynamic fields. There already is timestamp,
but there is none for "original packet length". 

Third option is not using pdump callbacks at all and just having rpcap
specific callbacks. That is what the Wireshark extcap does. That is where
I think this will end up. Would like to not use pdump since it has
so many design problems baked in.


> >    app/rpcapd: remote pcap daemon  
> Ideally forwarding captured data could run in a separate thread with 
> minimum unnecessary syscalls, although it's hard to tell without 
> benchmarks how much current design is slowed down by them.

New version does use a data forwarding thread.
It has to write to a socket, which it does with sendmsg.
It could use liburing, to reduce syscalls further.

> >    app/rpcapd: add TLS support  
> This patch simultaneously adds authentication, but it's only required 
> for non-loopback users. This is a common security vulnerability, not all 
> loopback users can be trusted. So if we care about authentication at 
> all, it should probably also be required for localhost users, at least 
> by default.

Good point, but do not want to add DPDK specific user data base.
The problem is that TCP loopback has no way to identify who is on the
other end. Unix domain sockets do, but libpcap doesn't support that.

The real long term answer is having plugins in libpcap (Robin's proposal)
and make this one of the plugins. But that will take longer to reach
agreement on.

^ permalink raw reply	[flat|nested] 21+ messages in thread

end of thread, other threads:[~2026-10-01 23:00 UTC | newest]

Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-08 21:07 [PATCH] examples/rpcapd: demo version of packet capture daemon Stephen Hemminger
2026-09-20 18:59 ` [PATCH v2] " Stephen Hemminger
2026-09-21  9:51   ` Marat Khalili
2026-09-21 15:57     ` Stephen Hemminger
2026-09-21 15:58     ` Stephen Hemminger
2026-09-21 16:43       ` Marat Khalili
2026-09-21 17:31         ` Stephen Hemminger
2026-09-21 17:53           ` Marat Khalili
2026-09-21 16:17     ` Stephen Hemminger
2026-09-22 18:45   ` Stephen Hemminger
2026-09-22 21:31 ` [PATCH v3] " Stephen Hemminger
2026-09-28 16:18   ` Marat Khalili
2026-09-28 17:24     ` Stephen Hemminger
2026-10-01  2:35 ` [PATCH v4 0/4] add rpcap remote " Stephen Hemminger
2026-10-01  2:35   ` [PATCH v4 1/4] pcapng: add API to read back capture mbuf header Stephen Hemminger
2026-10-01  2:35   ` [PATCH v4 2/4] app/rpcapd: remote pcap daemon Stephen Hemminger
2026-10-01 18:50     ` Marat Khalili
2026-10-01  2:35   ` [PATCH v4 3/4] app/rpcapd: add TLS support Stephen Hemminger
2026-10-01  2:35   ` [PATCH v4 4/4] app/rpcapd: add host list option Stephen Hemminger
2026-10-01 18:50   ` [PATCH v4 0/4] add rpcap remote capture daemon Marat Khalili
2026-10-01 23:00     ` Stephen Hemminger

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.