Netdev List
 help / color / mirror / Atom feed
From: Stanislav Fomichev <sdf.kernel@gmail.com>
To: Mahe Tardy <mahe.tardy@gmail.com>
Cc: bpf@vger.kernel.org, andrew+netdev@lunn.ch, andrii@kernel.org,
	 ast@kernel.org, daniel@iogearbox.net, davem@davemloft.net,
	eddyz87@gmail.com,  edumazet@google.com,
	john.fastabend@gmail.com, kuba@kernel.org, liamwisehart@meta.com,
	 martin.lau@linux.dev, pabeni@redhat.com, song@kernel.org,
	netdev@vger.kernel.org,  ameryhung@gmail.com, kuniyu@google.com
Subject: Re: [PATCH bpf-next v2 2/5] bpf: Add ksock kfuncs
Date: Wed, 22 Jul 2026 11:27:32 -0700	[thread overview]
Message-ID: <amELAv1bgc9DP6WT@devvm7509.cco0.facebook.com> (raw)
In-Reply-To: <20260722104454.165911-3-mahe.tardy@gmail.com>

On 07/22, Mahe Tardy wrote:
> Add BPF kfuncs that allow BPF LSM programs to create and use sockets for
> sending data. This provides a mechanism for BPF programs to emit
> telemetry. For this first patch set, it's restricted to SOCK_DGRAM
> socket types with IPPROTO_UDP protocol but could be easily extended to
> SOCK_STREAM and IPPROTO_TCP in the future.
> 
> The API consists of five kfuncs:
> 
>   bpf_ksock_create()   - Create a socket (sleepable)
>   bpf_ksock_connect()  - Connect socket to remote address (sleepable)
>   bpf_ksock_send()     - Send data through the socket (sleepable)
>   bpf_ksock_acquire()  - Acquire a reference to a socket context
>   bpf_ksock_release()  - Release a reference (cleanup via
>                          queue_rcu_work since sock_release sleeps)
> 
> The setup kfuncs bpf_ksock_create, bpf_ksock_connect, can be called from
> SYSCALL programs only. While bpf_ksock_acquire, bpf_ksock_release and
> bpf_ksock_send can be called from SYSCALL and LSM programs.
> 
> The implementation follows the established kfunc lifecycle pattern
> (create/acquire/release with refcounting, kptr map storage, dtor
> registration). The kernel socket is wrapped in a refcounted bpf_ksock
> struct. Cleanup is deferred via queue_rcu_work() because sock_release()
> may sleep.
> 
> The kfuncs are only compiled when CONFIG_INET is enabled, as they
> specifically support AF_INET and AF_INET6 sockets.
> 
> The socket operations go through the expected LSM hooks instead of
> by-passing them like many kernel sockets since those are created by BPF
> programs and thus system users. Thus bpf_ksock_send() kfunc, which is
> exposed to LSM progs, has a re-entering protection to avoid recursion.
> Also, because of the LSM checks, we prevent the use of the kfuncs from
> asynchronous workqueue as the current value would then be invalid.
> 
> In bpf_ksock_create(), we copy the arg values to avoid TOCTOU races
> since the kfunc can sleep and the arg values could be stored in a map
> that could be re-written by BPF progs or even userspace programs if the
> map is mmaped.
> 
> Signed-off-by: Mahe Tardy <mahe.tardy@gmail.com>
> ---
>  include/linux/bpf_ksock.h |  46 +++++
>  include/linux/sched.h     |   4 +
>  kernel/bpf/verifier.c     |   3 +
>  net/core/Makefile         |   3 +
>  net/core/bpf_ksock.c      | 361 ++++++++++++++++++++++++++++++++++++++
>  5 files changed, 417 insertions(+)
>  create mode 100644 include/linux/bpf_ksock.h
>  create mode 100644 net/core/bpf_ksock.c
> 
> diff --git a/include/linux/bpf_ksock.h b/include/linux/bpf_ksock.h
> new file mode 100644
> index 000000000000..a2e06e3d604d
> --- /dev/null
> +++ b/include/linux/bpf_ksock.h
> @@ -0,0 +1,46 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +/* Copyright (c) 2026 Isovalent */
> +
> +#ifndef _BPF_KSOCK_H
> +#define _BPF_KSOCK_H
> +
> +#include <linux/types.h>
> +

[..]

> +/**
> + * struct bpf_ksock_create_opts - BPF kernel socket creation parameters
> + * @family:	Address family: AF_INET or AF_INET6.
> + * @type:	Socket type: only SOCK_DGRAM supported for now.
> + * @protocol:	Protocol number (e.g. IPPROTO_UDP), or 0 for the default protocol
> + *		of the given type.
> + * @reserved:	Must be zero. Reserved for future use.
> + */
> +struct bpf_ksock_create_opts {
> +	__u8 family;
> +	__u8 type;
> +	__u8 protocol;
> +	__u8 reserved;
> +};
> +
> +/**
> + * struct bpf_ksock_addr_opts - BPF kernel socket address parameters
> + * @family:	Address family: AF_INET or AF_INET6.
> + * @reserved:	Must be zero. Reserved for future use.
> + * @port:	Port in host byte order.
> + * @scope_id:	IPv6 scope ID for scoped AF_INET6 addresses, or zero.
> + *		Must be zero when family=AF_INET.
> + * @ipv4_addr:	IPv4 address in network byte order. Used when family=AF_INET.
> + * @ipv6_addr:	IPv6 address (16 bytes, network byte order). Used when family=AF_INET6.
> + */
> +struct bpf_ksock_addr_opts {
> +	__u8 family;
> +	__u8 reserved;
> +	__u16 port;
> +	__u32 scope_id;
> +
> +	union {
> +		__be32 ipv4_addr;
> +		__u32 ipv6_addr[4]; /* in6_addr; network order */
> +	};
> +};

Do you add these because of BTF type checking? Can we have something a bit
more orthodox like:

union bpf_sockaddr {
	struct sockaddr_in sin;
	struct sockaddr_in6 sin6;
};

?

I hope we can just pass these to sock_create and make it do all the input
checking..

  reply	other threads:[~2026-07-22 18:27 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-22 10:44 [PATCH bpf-next v2 0/5] Introduce bpf_ksock Mahe Tardy
2026-07-22 10:44 ` [PATCH bpf-next v2 1/5] net: Add __sys_connect_socket() helper Mahe Tardy
2026-07-22 10:44 ` [PATCH bpf-next v2 2/5] bpf: Add ksock kfuncs Mahe Tardy
2026-07-22 18:27   ` Stanislav Fomichev [this message]
2026-07-22 10:44 ` [PATCH bpf-next v2 3/5] selftests/bpf: Add ksock kfunc test Mahe Tardy
2026-07-22 10:44 ` [PATCH bpf-next v2 4/5] selftests/bpf: Add ksock LSM recursion test Mahe Tardy
2026-07-22 10:44 ` [PATCH bpf-next v2 5/5] selftests/bpf: Add ksock test for async callback guard Mahe Tardy

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amELAv1bgc9DP6WT@devvm7509.cco0.facebook.com \
    --to=sdf.kernel@gmail.com \
    --cc=ameryhung@gmail.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=eddyz87@gmail.com \
    --cc=edumazet@google.com \
    --cc=john.fastabend@gmail.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=liamwisehart@meta.com \
    --cc=mahe.tardy@gmail.com \
    --cc=martin.lau@linux.dev \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=song@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox