* [PATCH mptcp-next v2 1/2] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
2026-09-03 1:12 [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables Hangbin Liu
@ 2026-09-03 1:12 ` Hangbin Liu
2026-09-03 1:12 ` [PATCH mptcp-next v2 2/2] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Hangbin Liu
2026-09-04 14:59 ` [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables Matthieu Baerts
2 siblings, 0 replies; 4+ messages in thread
From: Hangbin Liu @ 2026-09-03 1:12 UTC (permalink / raw)
To: MPTCP Linux, Matthieu Baerts, Mat Martineau, Geliang Tang,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Shuah Khan
Cc: Hangbin Liu, netdev, linux-kselftest, linux-kernel, bpf,
Hangbin Liu
From: Hangbin Liu <liuhangbin@kylinos.cn>
During the conversion, we retain the same filter and chain names previously
used by iptables/ip6tables. Counters are not added to accept rules because
the test does not inspect them. After conversion, the generated output
matches the original iptables/ip6tables behavior.
Signed-off-by: Hangbin Liu <liuhangbin@kylinos.cn>
---
tools/testing/selftests/net/mptcp/mptcp_lib.sh | 2 +-
tools/testing/selftests/net/mptcp/mptcp_sockopt.sh | 51 ++++++++++++----------
2 files changed, 29 insertions(+), 24 deletions(-)
diff --git a/tools/testing/selftests/net/mptcp/mptcp_lib.sh b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
index b9d14647f401..41febb1bbbc7 100644
--- a/tools/testing/selftests/net/mptcp/mptcp_lib.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
@@ -528,7 +528,7 @@ mptcp_lib_check_tools() {
exit ${KSFT_SKIP}
fi
;;
- "iptables"* | "ip6tables"*)
+ "iptables"* | "ip6tables"* | "nft"*)
if ! "${tool}" -V &> /dev/null; then
mptcp_lib_pr_skip "Could not run all tests without ${tool}"
exit ${KSFT_SKIP}
diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
index e850a87429b6..a2c20483986d 100755
--- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
@@ -15,8 +15,6 @@ cin=""
cout=""
timeout_poll=30
timeout_test=$((timeout_poll * 2 + 1))
-iptables="iptables"
-ip6tables="ip6tables"
ns1=""
ns2=""
@@ -50,15 +48,25 @@ add_mark_rules()
local m=$2
local t
- for t in ${iptables} ${ip6tables}; do
+ for t in ip ip6; do
+ ip netns exec "$ns" nft add table "$t" filter
+ ip netns exec "$ns" nft add chain "$t" filter OUTPUT \
+ '{ type filter hook output priority 0; policy accept; }'
+
# just to debug: check we have multiple subflows connection requests
- ip netns exec $ns $t -A OUTPUT -p tcp --syn -m mark --mark $m -j ACCEPT
+ ip netns exec "$ns" nft add rule "$t" filter OUTPUT \
+ tcp flags \& \(fin \| syn \| rst \| ack\) == syn \
+ meta mark "$m" accept
# RST packets might be handled by a internal dummy socket
- ip netns exec $ns $t -A OUTPUT -p tcp --tcp-flags RST RST -m mark --mark 0 -j ACCEPT
+ ip netns exec "$ns" nft add rule "$t" filter OUTPUT \
+ tcp flags \& rst == rst meta mark 0x0 accept
+
+ ip netns exec "$ns" nft add rule "$t" filter OUTPUT \
+ meta l4proto tcp meta mark "$m" accept
+ ip netns exec "$ns" nft add rule "$t" filter OUTPUT \
+ meta l4proto tcp meta mark 0 counter drop
- ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark $m -j ACCEPT
- ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark 0 -j DROP
done
}
@@ -105,32 +113,29 @@ cleanup()
mptcp_lib_check_mptcp
mptcp_lib_check_kallsyms
-mptcp_lib_check_tools ip "${iptables}" "${ip6tables}"
+mptcp_lib_check_tools ip nft
check_mark()
{
local ns=$1
local af=$2
- local tables=${iptables}
+ local tables="ip"
if [ $af -eq 6 ];then
- tables=${ip6tables}
+ tables="ip6"
fi
- local counters values
- counters=$(ip netns exec $ns $tables -v -L OUTPUT | grep DROP)
- values=${counters%DROP*}
-
- local v
- for v in $values; do
- if [ $v -ne 0 ]; then
- mptcp_lib_pr_fail "got $tables $values in ns $ns," \
- "not 0 - not all expected packets marked"
- ret=${KSFT_FAIL}
- return 1
- fi
- done
+ local values
+ values=$(ip netns exec "$ns" nft list table "$tables" filter | \
+ grep -o "packets.*drop" | awk '{print $2}')
+
+ if [[ ! "$values" =~ ^[0-9]+$ ]] || [ "$values" -ne 0 ]; then
+ mptcp_lib_pr_fail "got $tables $values in ns $ns," \
+ "not 0 - not all expected packets marked"
+ ret=${KSFT_FAIL}
+ return 1
+ fi
return 0
}
--
2.55.0
^ permalink raw reply related [flat|nested] 4+ messages in thread* [PATCH mptcp-next v2 2/2] selftests: mptcp: convert iptables to nftables for mptcp_join.sh
2026-09-03 1:12 [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables Hangbin Liu
2026-09-03 1:12 ` [PATCH mptcp-next v2 1/2] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Hangbin Liu
@ 2026-09-03 1:12 ` Hangbin Liu
2026-09-04 14:59 ` [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables Matthieu Baerts
2 siblings, 0 replies; 4+ messages in thread
From: Hangbin Liu @ 2026-09-03 1:12 UTC (permalink / raw)
To: MPTCP Linux, Matthieu Baerts, Mat Martineau, Geliang Tang,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Shuah Khan
Cc: Hangbin Liu, netdev, linux-kselftest, linux-kernel, bpf,
Hangbin Liu
From: Hangbin Liu <liuhangbin@kylinos.cn>
During conversion, we retain the same table and chain names used by the
original iptables/ip6tables setup, so rule output is identical to the
former iptables/ip6tables output.
The BPF bytecode matching MPTCP add‑addr and remove‑addr suboptions
is replaced with native nft matching using "tcp option mptcp subtype".
Unlike iptables, nftables cannot match rules based on their full
specification. Rule handles are captured via "nft -e --handle" so rules
can be selectively removed during tests.
The config file adds CONFIG_NFT_NUMGEN (replaces iptables statistic nth),
CONFIG_NFT_REJECT and CONFIG_NFT_REJECT_IPV4 for reject‑related rules.
The iptables/ip6tables check within mptcp_lib.sh is preserved in case
some users still require it.
Signed-off-by: Hangbin Liu <liuhangbin@kylinos.cn>
---
tools/testing/selftests/net/mptcp/config | 3 +
tools/testing/selftests/net/mptcp/mptcp_join.sh | 142 +++++++++---------------
2 files changed, 55 insertions(+), 90 deletions(-)
diff --git a/tools/testing/selftests/net/mptcp/config b/tools/testing/selftests/net/mptcp/config
index 59051ee2a986..0d0a744c4ca8 100644
--- a/tools/testing/selftests/net/mptcp/config
+++ b/tools/testing/selftests/net/mptcp/config
@@ -30,6 +30,9 @@ CONFIG_NET_SCH_NETEM=m
CONFIG_NF_TABLES=m
CONFIG_NF_TABLES_INET=y
CONFIG_NFT_COMPAT=m
+CONFIG_NFT_NUMGEN=y
+CONFIG_NFT_REJECT=m
+CONFIG_NFT_REJECT_IPV4=m
CONFIG_NFT_SOCKET=m
CONFIG_NFT_TPROXY=m
CONFIG_SYN_COOKIES=y
diff --git a/tools/testing/selftests/net/mptcp/mptcp_join.sh b/tools/testing/selftests/net/mptcp/mptcp_join.sh
index 18ce7136a2b0..7c56680cc422 100755
--- a/tools/testing/selftests/net/mptcp/mptcp_join.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_join.sh
@@ -26,8 +26,6 @@ capout=""
cappid=""
ns1=""
ns2=""
-iptables="iptables"
-ip6tables="ip6tables"
timeout_poll=30
timeout_test=$((timeout_poll * 2 + 1))
capture=false
@@ -50,6 +48,7 @@ declare -A failed_tests
MPTCP_LIB_TEST_FORMAT="%03u %s\n"
TEST_NAME=""
nr_blank=6
+nft_handle=""
# These var are used only in some tests, make sure they are not already set
unset FAILING_LINKS
@@ -99,42 +98,6 @@ unset add_addr_tx_nr
unset add_addr_echo_tx_nr
unset add_addr_drop_tx_nr
-# generated using "nfbpf_compile '(ip && (ip[54] & 0xf0) == 0x30) ||
-# (ip6 && (ip6[74] & 0xf0) == 0x30)'"
-CBPF_MPTCP_SUBOPTION_ADD_ADDR="14,
- 48 0 0 0,
- 84 0 0 240,
- 21 0 3 64,
- 48 0 0 54,
- 84 0 0 240,
- 21 6 7 48,
- 48 0 0 0,
- 84 0 0 240,
- 21 0 4 96,
- 48 0 0 74,
- 84 0 0 240,
- 21 0 1 48,
- 6 0 0 65535,
- 6 0 0 0"
-
-# IPv4: TCP hdr of 48B, a first suboption of 12B (DACK8), the RM_ADDR suboption
-# generated using "nfbpf_compile '(ip[32] & 0xf0) == 0xc0 && ip[53] == 0x0c &&
-# (ip[66] & 0xf0) == 0x40'"
-CBPF_MPTCP_SUBOPTION_RM_ADDR="13,
- 48 0 0 0,
- 84 0 0 240,
- 21 0 9 64,
- 48 0 0 32,
- 84 0 0 240,
- 21 0 6 192,
- 48 0 0 53,
- 21 0 4 12,
- 48 0 0 66,
- 84 0 0 240,
- 21 0 1 64,
- 6 0 0 65535,
- 6 0 0 0"
-
init_partial()
{
capout=$(mktemp)
@@ -147,6 +110,18 @@ init_partial()
if $checksum; then
ip netns exec $netns sysctl -q net.mptcp.checksum_enabled=1
fi
+
+ for t in ip ip6; do
+ ip netns exec "$netns" nft add table "$t" filter
+ ip netns exec "$netns" nft add chain "$t" filter INPUT \
+ '{ type filter hook input priority filter; policy accept; }'
+ ip netns exec "$netns" nft add chain "$t" filter OUTPUT \
+ '{ type filter hook output priority filter; policy accept; }'
+
+ ip netns exec "$netns" nft add table "$t" mangle
+ ip netns exec "$netns" nft add chain "$t" mangle OUTPUT \
+ '{ type route hook output priority mangle; policy accept; }'
+ done
done
check_invert=0
@@ -196,7 +171,7 @@ init() {
mptcp_lib_check_mptcp
mptcp_lib_check_kallsyms
- mptcp_lib_check_tools ip tc ss "${iptables}" "${ip6tables}"
+ mptcp_lib_check_tools ip tc ss nft
sin=$(mktemp)
sout=$(mktemp)
@@ -380,24 +355,18 @@ reset_with_cookies()
# $1: test name
reset_with_add_addr_timeout()
{
- local ip="${2:-4}"
- local tables
+ local ip="${2:-}"
reset "${1}" || return 1
- tables="${iptables}"
- if [ $ip -eq 6 ]; then
- tables="${ip6tables}"
- fi
-
# set a maximum, to avoid too long timeout with exponential backoff
ip netns exec $ns1 sysctl -q net.mptcp.add_addr_timeout=1
- if ! ip netns exec $ns2 $tables -A OUTPUT -p tcp \
- -m tcp --tcp-option 30 \
- -m bpf --bytecode \
- "$CBPF_MPTCP_SUBOPTION_ADD_ADDR" \
- -j DROP; then
+ nft_handle=$(ip netns exec "$ns2" nft -e --handle add rule \
+ ip"$ip" filter OUTPUT meta l4proto tcp \
+ tcp option mptcp subtype add-addr \
+ drop | head -n1 | awk '{print $NF}')
+ if [ -z "$nft_handle" ]; then
mark_as_skipped "unable to set the 'add addr' rule"
return 1
fi
@@ -449,23 +418,17 @@ setup_fail_rules()
check_invert=1
validate_checksum=true
local i="$1"
- local ip="${2:-4}"
- local tables
+ local ip="${2:-}"
- tables="${iptables}"
- if [ $ip -eq 6 ]; then
- tables="${ip6tables}"
+ nft_handle=$(ip netns exec "$ns2" nft -e --handle add rule \
+ ip"$ip" mangle OUTPUT oifname ns2eth$i \
+ meta l4proto tcp \
+ meta length 150-9999 numgen inc mod 99999 1 \
+ meta mark set 42 | head -n1 | awk '{print $NF}')
+ if [ -z "$nft_handle" ]; then
+ return ${KSFT_SKIP}
fi
- ip netns exec $ns2 $tables \
- -t mangle \
- -A OUTPUT \
- -o ns2eth$i \
- -p tcp \
- -m length --length 150:9999 \
- -m statistic --mode nth --packet 1 --every 99999 \
- -j MARK --set-mark 42 || return ${KSFT_SKIP}
-
tc -n $ns2 qdisc add dev ns2eth$i clsact || return ${KSFT_SKIP}
tc -n $ns2 filter add dev ns2eth$i egress \
protocol ip prio 1000 \
@@ -515,11 +478,10 @@ reset_with_tcp_filter()
local target="${3}"
local chain="${4:-INPUT}"
- if ! ip netns exec "${ns}" ${iptables} \
- -A "${chain}" \
- -s "${src}" \
- -p tcp \
- -j "${target}"; then
+ nft_handle=$(ip netns exec "$ns" nft -e --handle add rule \
+ ip filter "${chain}" ip saddr "{ ${src} }" \
+ meta l4proto tcp "${target,,}" | head -n1 | awk '{print $NF}')
+ if [ -z "$nft_handle" ]; then
mark_as_skipped "unable to set the filter rules"
return 1
fi
@@ -4315,10 +4277,12 @@ userspace_tests()
# force quick loss
ip netns exec $ns2 sysctl -q net.ipv4.tcp_syn_retries=1
- if ip netns exec "${ns1}" ${iptables} -A INPUT -s "10.0.1.2" \
- -p tcp --tcp-option 30 -j REJECT --reject-with tcp-reset &&
- ip netns exec "${ns2}" ${iptables} -A INPUT -d "10.0.1.2" \
- -p tcp --tcp-option 30 -j REJECT --reject-with tcp-reset; then
+ if ip netns exec "${ns1}" nft add rule ip filter INPUT \
+ ip saddr "10.0.1.2" meta l4proto tcp \
+ tcp option mptcp exists reject with tcp reset &&
+ ip netns exec "${ns2}" nft add rule ip filter INPUT \
+ ip daddr "10.0.1.2" meta l4proto tcp \
+ tcp option mptcp exists reject with tcp reset; then
wait_event ns2 MPTCP_LIB_EVENT_SUB_CLOSED 1
wait_event ns1 MPTCP_LIB_EVENT_SUB_CLOSED 1
chk_subflows_total 1 1
@@ -4393,7 +4357,7 @@ endpoint_tests()
chk_subflow_nr "after new reject" 2
chk_mptcp_info subflows 1 subflows 1
- ip netns exec "${ns2}" ${iptables} -D OUTPUT -s "10.0.3.2" -p tcp -j REJECT
+ ip netns exec "${ns2}" nft delete rule ip filter OUTPUT handle "$nft_handle"
pm_nl_del_endpoint $ns2 3 10.0.3.2
pm_nl_add_endpoint $ns2 10.0.3.2 id 3 flags subflow
wait_mpj 3
@@ -4402,12 +4366,10 @@ endpoint_tests()
# To make sure RM_ADDR are sent over a different subflow, but
# allow the rest to quickly and cleanly close the subflow
- local ipt=1
- ip netns exec "${ns2}" ${iptables} -I OUTPUT -s "10.0.1.2" \
- -p tcp -m tcp --tcp-option 30 \
- -m bpf --bytecode \
- "$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
- -j DROP || ipt=0
+ nft_handle=$(ip netns exec "${ns2}" nft -e --handle insert rule \
+ ip filter OUTPUT ip saddr 10.0.1.2 meta l4proto tcp \
+ tcp option mptcp subtype remove-addr \
+ drop | head -n1 | awk '{print $NF}')
local i
for i in $(seq 3); do
pm_nl_del_endpoint $ns2 1 10.0.1.2
@@ -4420,7 +4382,8 @@ endpoint_tests()
chk_subflow_nr "after re-add id 0 ($i)" 3
chk_mptcp_info subflows 3 subflows 3
done
- [ ${ipt} = 1 ] && ip netns exec "${ns2}" ${iptables} -D OUTPUT 1
+ [ -n "${nft_handle}" ] && ip netns exec "${ns2}" nft delete rule \
+ ip filter OUTPUT handle "${nft_handle}"
mptcp_lib_kill_group_wait $tests_pid
@@ -4482,18 +4445,17 @@ endpoint_tests()
# To make sure RM_ADDR are sent over a different subflow, but
# allow the rest to quickly and cleanly close the subflow
- local ipt=1
- ip netns exec "${ns1}" ${iptables} -I OUTPUT -s "10.0.1.1" \
- -p tcp -m tcp --tcp-option 30 \
- -m bpf --bytecode \
- "$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
- -j DROP || ipt=0
+ nft_handle=$(ip netns exec "${ns1}" nft -e --handle insert rule \
+ ip filter OUTPUT ip saddr 10.0.1.1 meta l4proto tcp \
+ tcp option mptcp subtype remove-addr \
+ drop | head -n1 | awk '{print $NF}')
pm_nl_del_endpoint $ns1 42 10.0.1.1
sleep 0.5
chk_subflow_nr "after delete ID 0" 2
chk_mptcp_info subflows 2 subflows 2
chk_mptcp_info add_addr_signal 2 add_addr_accepted 2
- [ ${ipt} = 1 ] && ip netns exec "${ns1}" ${iptables} -D OUTPUT 1
+ [ -n "${nft_handle}" ] && ip netns exec "${ns1}" nft delete rule \
+ ip filter OUTPUT handle "${nft_handle}"
pm_nl_add_endpoint $ns1 10.0.1.1 id 42 flags signal
wait_mpj 4
@@ -4555,7 +4517,7 @@ endpoint_tests()
pm_nl_flush_endpoint $ns2
pm_nl_flush_endpoint $ns1
wait_rm_addr $ns2 0
- ip netns exec "${ns2}" ${iptables} -D OUTPUT -s "10.0.3.2" -p tcp -j REJECT
+ ip netns exec "${ns2}" nft delete rule ip filter OUTPUT handle "$nft_handle"
pm_nl_add_endpoint $ns2 10.0.3.2 id 3 flags subflow
wait_mpj 1
pm_nl_add_endpoint $ns1 10.0.3.1 id 2 flags signal
--
2.55.0
^ permalink raw reply related [flat|nested] 4+ messages in thread* Re: [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables
2026-09-03 1:12 [PATCH mptcp-next v2 0/2] selftests: mptcp: convert iptables to nftables Hangbin Liu
2026-09-03 1:12 ` [PATCH mptcp-next v2 1/2] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Hangbin Liu
2026-09-03 1:12 ` [PATCH mptcp-next v2 2/2] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Hangbin Liu
@ 2026-09-04 14:59 ` Matthieu Baerts
2 siblings, 0 replies; 4+ messages in thread
From: Matthieu Baerts @ 2026-09-04 14:59 UTC (permalink / raw)
To: Hangbin Liu, MPTCP Linux, Mat Martineau, Geliang Tang,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Shuah Khan
Cc: netdev, linux-kselftest, linux-kernel, bpf, Hangbin Liu
Hello,
On 03/09/2026 03:12, Hangbin Liu wrote:
> iptables has been deprecated for years. The Linux kernel has included
> nftables as the successor to iptables since 2014, and every major
> distribution uses nftables as the default packet filtering framework.
> The iptables command we run on modern systems is actually iptables‑nft,
> a compatibility layer that translates iptables syntax to nftables rules
> behind the scenes.
>
> There are also some features that can be set easily with nft, while we need
> to convert to BPF code under iptables, such as MPTCP add‑addr and
> remove‑addr suboptions. To make future work easier, convert iptables usage
> in mptcp to nftables.
>
> Tested with iptables-translate to make sure each nft conversion is the same
> with previous one. e.g. for mptcp_sockopt.sh, the ip6tables shows
FYI, Hangbin is working on a new version addressing my comments from v1.
The new version(s) will be sent to the MPTCP list only, and I will sent
these patches to Netdev when ready.
Updating here the PW status:
pw-bot: cr
Cheers,
Matt
--
Sponsored by the NGI0 Core fund.
^ permalink raw reply [flat|nested] 4+ messages in thread