Linux Netfilter development
 help / color / mirror / Atom feed
* Netfilter: match "tcp option" with the same type present multiple times
@ 2026-09-28 12:13 Matthieu Baerts
  2026-09-28 13:13 ` Florian Westphal
                   ` (3 more replies)
  0 siblings, 4 replies; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-28 12:13 UTC (permalink / raw)
  To: Netfilter Devel, Netfilter Coreteam

Hello Netfilter devs,

First, thank you for maintaining Netfilter in the kernel and the
userspace tools.

The MPTCP selftests are switching from IPTables to NFTables, and
Clashiko reported that this part of a rule wouldn't match anything:

  tcp option mptcp subtype remove-addr drop

When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
added to the TCP options, it is added after an MPTCP DSS option (type
0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
(type 30) options in the TCP options. It looks like Netfilter doesn't
handle that, because it stops processing other TCP options when the
expected type is found:

net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
    ...
	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
		optl = optlen(opt, i);

		if (priv->type != opt[i])
			continue;
		...
		return;  // <== it will look at the first MPTCP option
	}
    ...
}

It looks like it shouldn't stop if the wrong subtype is found, but the
subtype is not compared there if I'm not mistaken. Should there be a fix
to support this case?

For more details about the Clashiko report:

https://lore.kernel.org/179058242453.3145.1450534124183352369@kernel.org


BTW, I'm probably missing something, but adding the following doesn't
seem to have any effect on my side:

  nft add table ip filter
  nft add chain ip filter OUTPUT \
    { type filter hook output priority filter; policy accept; }
  nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop

Same with "mptcp subtype mp-capable", other subtypes or other types like
"timestamp". But if I use ...

  nft insert rule ip filter OUTPUT meta l4proto tcp counter drop

... then the drop is effective. Any idea what I'm missing here? :)


(Also, it looks like the MPTCP suboptions are not documented in the nft
man page :) )

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 12:13 Netfilter: match "tcp option" with the same type present multiple times Matthieu Baerts
@ 2026-09-28 13:13 ` Florian Westphal
  2026-09-28 20:26   ` Matthieu Baerts
  2026-09-28 20:56   ` Fernando Fernandez Mancera
  2026-09-28 20:54 ` Fernando Fernandez Mancera
                   ` (2 subsequent siblings)
  3 siblings, 2 replies; 22+ messages in thread
From: Florian Westphal @ 2026-09-28 13:13 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

Matthieu Baerts <matttbe@kernel.org> wrote:
> Hello Netfilter devs,
> 
> First, thank you for maintaining Netfilter in the kernel and the
> userspace tools.
> 
> The MPTCP selftests are switching from IPTables to NFTables, and
> Clashiko reported that this part of a rule wouldn't match anything:
> 
>   tcp option mptcp subtype remove-addr drop
> 
> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> added to the TCP options, it is added after an MPTCP DSS option (type
> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> (type 30) options in the TCP options. It looks like Netfilter doesn't
> handle that, because it stops processing other TCP options when the
> expected type is found:

Yes, this won't work. 'tcp option X' extracts the
option X.

I don't see how this could be fixed within the limitations of the
architecture.  Just use bpf.

> It looks like it shouldn't stop if the wrong subtype is found, but the
> subtype is not compared there if I'm not mistaken. Should there be a fix
> to support this case?

I don't know how, unless one would extend the kernel to make it aware of
mptcp, which also requires userspace to pass the suboption type to look
for in addition to 'mptcp option'.

> BTW, I'm probably missing something, but adding the following doesn't
> seem to have any effect on my side:
> 
>   nft add table ip filter
>   nft add chain ip filter OUTPUT \
>     { type filter hook output priority filter; policy accept; }
>   nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
> 
> Same with "mptcp subtype mp-capable", other subtypes or other types like
> "timestamp". But if I use ...
> 
>   nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
> 
> ... then the drop is effective. Any idea what I'm missing here? :)

No idea.
cd git/netfilter.org/nftables/tests/shell; ./run-tests.sh -k testcases/packetpath/tcp_options

passes here.
Even tried adding 'tcp option timestamp exists...' to the test, also
passes.

nft --debug=netlink list ruleset displays the rule like this:

inet t c 15 14
  [ exthdr load tcpopt 1b @ 8 + 0 present => reg 1 ]
  [ cmp eq reg 1 0x01 ]
  [ objref type 1 name tsc ]

Linux 7.3.0-rc3+
nftables v1.1.6 (Commodore Bullmoose #7)


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 13:13 ` Florian Westphal
@ 2026-09-28 20:26   ` Matthieu Baerts
  2026-09-28 21:08     ` Florian Westphal
  2026-09-28 20:56   ` Fernando Fernandez Mancera
  1 sibling, 1 reply; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-28 20:26 UTC (permalink / raw)
  To: Florian Westphal; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Florian,

Thank you for your reply!

On 28/09/2026 15:13, Florian Westphal wrote:
> Matthieu Baerts <matttbe@kernel.org> wrote:
>> Hello Netfilter devs,
>>
>> First, thank you for maintaining Netfilter in the kernel and the
>> userspace tools.
>>
>> The MPTCP selftests are switching from IPTables to NFTables, and
>> Clashiko reported that this part of a rule wouldn't match anything:
>>
>>   tcp option mptcp subtype remove-addr drop
>>
>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>> added to the TCP options, it is added after an MPTCP DSS option (type
>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>> handle that, because it stops processing other TCP options when the
>> expected type is found:
> 
> Yes, this won't work. 'tcp option X' extracts the
> option X.
> 
> I don't see how this could be fixed within the limitations of the
> architecture.

That's what I was thinking too :-/

> Just use bpf.

Yes, or with a raw payload expression with NFTables.

>> It looks like it shouldn't stop if the wrong subtype is found, but the
>> subtype is not compared there if I'm not mistaken. Should there be a fix
>> to support this case?
> 
> I don't know how, unless one would extend the kernel to make it aware of
> mptcp, which also requires userspace to pass the suboption type to look
> for in addition to 'mptcp option'.

Should there be a note somewhere in the doc about this limitation?
Because it looks like it will never be possible to match such subtype.

There are other MPTCP suboptions that can be used with DSS, e.g. MP_PRIO
and MP_FAIL. But also MP_RST that can be used with FAST_CLOSE.

>> BTW, I'm probably missing something, but adding the following doesn't
>> seem to have any effect on my side:
>>
>>   nft add table ip filter
>>   nft add chain ip filter OUTPUT \
>>     { type filter hook output priority filter; policy accept; }
>>   nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
>>
>> Same with "mptcp subtype mp-capable", other subtypes or other types like
>> "timestamp". But if I use ...
>>
>>   nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
>>
>> ... then the drop is effective. Any idea what I'm missing here? :)
> 
> No idea.
> cd git/netfilter.org/nftables/tests/shell; ./run-tests.sh -k testcases/packetpath/tcp_options
> 
> passes here.
> Even tried adding 'tcp option timestamp exists...' to the test, also
> passes.

Arf, my bad, I still had debugging code, sorry... :/

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 12:13 Netfilter: match "tcp option" with the same type present multiple times Matthieu Baerts
  2026-09-28 13:13 ` Florian Westphal
@ 2026-09-28 20:54 ` Fernando Fernandez Mancera
  2026-09-28 21:17   ` Fernando Fernandez Mancera
  2026-09-28 21:52 ` Pablo Neira Ayuso
  2026-09-28 21:54 ` Jan Engelhardt
  3 siblings, 1 reply; 22+ messages in thread
From: Fernando Fernandez Mancera @ 2026-09-28 20:54 UTC (permalink / raw)
  To: Matthieu Baerts, Netfilter Devel, Netfilter Coreteam

On 9/28/26 2:13 PM, Matthieu Baerts wrote:
> Hello Netfilter devs,
> 
> First, thank you for maintaining Netfilter in the kernel and the
> userspace tools.
> 
> The MPTCP selftests are switching from IPTables to NFTables, and
> Clashiko reported that this part of a rule wouldn't match anything:
> 
>    tcp option mptcp subtype remove-addr drop
> 
> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> added to the TCP options, it is added after an MPTCP DSS option (type
> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> (type 30) options in the TCP options. It looks like Netfilter doesn't
> handle that, because it stops processing other TCP options when the
> expected type is found:
> 
> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>      ...
> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> 		optl = optlen(opt, i);
> 
> 		if (priv->type != opt[i])
> 			continue;
> 		...
> 		return;  // <== it will look at the first MPTCP option
> 	}
>      ...
> }
> 
> It looks like it shouldn't stop if the wrong subtype is found, but the
> subtype is not compared there if I'm not mistaken. Should there be a fix
> to support this case?
> 
> For more details about the Clashiko report:
> 
> https://lore.kernel.org/179058242453.3145.1450534124183352369@kernel.org
> 
> 
> BTW, I'm probably missing something, but adding the following doesn't
> seem to have any effect on my side:
> 
>    nft add table ip filter
>    nft add chain ip filter OUTPUT \
>      { type filter hook output priority filter; policy accept; }
>    nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
> 
> Same with "mptcp subtype mp-capable", other subtypes or other types like
> "timestamp". But if I use ...
> 
>    nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
> 
> ... then the drop is effective. Any idea what I'm missing here? :)
> 
> 
> (Also, it looks like the MPTCP suboptions are not documented in the nft
> man page :) )
> 

Tomorrow there should be a patch in netfilter-devel adding it, we 
definitively should put more work into documentation.

> Cheers,
> Matt


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 13:13 ` Florian Westphal
  2026-09-28 20:26   ` Matthieu Baerts
@ 2026-09-28 20:56   ` Fernando Fernandez Mancera
  2026-09-28 21:22     ` Florian Westphal
  1 sibling, 1 reply; 22+ messages in thread
From: Fernando Fernandez Mancera @ 2026-09-28 20:56 UTC (permalink / raw)
  To: Florian Westphal, Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

On 9/28/26 3:13 PM, Florian Westphal wrote:
> Matthieu Baerts <matttbe@kernel.org> wrote:
>> Hello Netfilter devs,
>>
>> First, thank you for maintaining Netfilter in the kernel and the
>> userspace tools.
>>
>> The MPTCP selftests are switching from IPTables to NFTables, and
>> Clashiko reported that this part of a rule wouldn't match anything:
>>
>>    tcp option mptcp subtype remove-addr drop
>>
>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>> added to the TCP options, it is added after an MPTCP DSS option (type
>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>> handle that, because it stops processing other TCP options when the
>> expected type is found:
> 
> Yes, this won't work. 'tcp option X' extracts the
> option X.
> 
> I don't see how this could be fixed within the limitations of the
> architecture.  Just use bpf.
> 
>> It looks like it shouldn't stop if the wrong subtype is found, but the
>> subtype is not compared there if I'm not mistaken. Should there be a fix
>> to support this case?
> 
> I don't know how, unless one would extend the kernel to make it aware of
> mptcp, which also requires userspace to pass the suboption type to look
> for in addition to 'mptcp option'.
> 

This won't be easy neither fast but I added this to my TODO list. It 
sounds fun. Unless someone else does it first, I will take it.

>> BTW, I'm probably missing something, but adding the following doesn't
>> seem to have any effect on my side:
>>
>>    nft add table ip filter
>>    nft add chain ip filter OUTPUT \
>>      { type filter hook output priority filter; policy accept; }
>>    nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
>>
>> Same with "mptcp subtype mp-capable", other subtypes or other types like
>> "timestamp". But if I use ...
>>
>>    nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
>>
>> ... then the drop is effective. Any idea what I'm missing here? :)
> 
> No idea.
> cd git/netfilter.org/nftables/tests/shell; ./run-tests.sh -k testcases/packetpath/tcp_options
> 
> passes here.
> Even tried adding 'tcp option timestamp exists...' to the test, also
> passes.
> 
> nft --debug=netlink list ruleset displays the rule like this:
> 
> inet t c 15 14
>    [ exthdr load tcpopt 1b @ 8 + 0 present => reg 1 ]
>    [ cmp eq reg 1 0x01 ]
>    [ objref type 1 name tsc ]
> 
> Linux 7.3.0-rc3+
> nftables v1.1.6 (Commodore Bullmoose #7)


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 20:26   ` Matthieu Baerts
@ 2026-09-28 21:08     ` Florian Westphal
  0 siblings, 0 replies; 22+ messages in thread
From: Florian Westphal @ 2026-09-28 21:08 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

Matthieu Baerts <matttbe@kernel.org> wrote:
> Yes, or with a raw payload expression with NFTables.

If you expect identical layout every time, yes, that works too.

> Should there be a note somewhere in the doc about this limitation?
> Because it looks like it will never be possible to match such subtype.

Yes, not without extra code on the kernel side.

> There are other MPTCP suboptions that can be used with DSS, e.g. MP_PRIO
> and MP_FAIL. But also MP_RST that can be used with FAST_CLOSE.

Right.  I did not consider that this stops at first mptcp option
encountered.

> > passes here.
> > Even tried adding 'tcp option timestamp exists...' to the test, also
> > passes.
> 
> Arf, my bad, I still had debugging code, sorry... :/

Phew :-)

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 20:54 ` Fernando Fernandez Mancera
@ 2026-09-28 21:17   ` Fernando Fernandez Mancera
  2026-09-28 21:22     ` Matthieu Baerts
  0 siblings, 1 reply; 22+ messages in thread
From: Fernando Fernandez Mancera @ 2026-09-28 21:17 UTC (permalink / raw)
  To: Matthieu Baerts, Netfilter Devel, Netfilter Coreteam

On 9/28/26 10:54 PM, Fernando Fernandez Mancera wrote:
> On 9/28/26 2:13 PM, Matthieu Baerts wrote:
>> Hello Netfilter devs,
>>
>> First, thank you for maintaining Netfilter in the kernel and the
>> userspace tools.
>>
>> The MPTCP selftests are switching from IPTables to NFTables, and
>> Clashiko reported that this part of a rule wouldn't match anything:
>>
>>    tcp option mptcp subtype remove-addr drop
>>
>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>> added to the TCP options, it is added after an MPTCP DSS option (type
>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>> handle that, because it stops processing other TCP options when the
>> expected type is found:
>>
>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>      ...
>>     for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>>         optl = optlen(opt, i);
>>
>>         if (priv->type != opt[i])
>>             continue;
>>         ...
>>         return;  // <== it will look at the first MPTCP option
>>     }
>>      ...
>> }
>>
>> It looks like it shouldn't stop if the wrong subtype is found, but the
>> subtype is not compared there if I'm not mistaken. Should there be a fix
>> to support this case?
>>
>> For more details about the Clashiko report:
>>
>> https://lore.kernel.org/179058242453.3145.1450534124183352369@kernel.org
>>
>>
>> BTW, I'm probably missing something, but adding the following doesn't
>> seem to have any effect on my side:
>>
>>    nft add table ip filter
>>    nft add chain ip filter OUTPUT \
>>      { type filter hook output priority filter; policy accept; }
>>    nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
>>
>> Same with "mptcp subtype mp-capable", other subtypes or other types like
>> "timestamp". But if I use ...
>>
>>    nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
>>
>> ... then the drop is effective. Any idea what I'm missing here? :)
>>
>>
>> (Also, it looks like the MPTCP suboptions are not documented in the nft
>> man page :) )
>>
> 
> Tomorrow there should be a patch in netfilter-devel adding it, we 
> definitively should put more work into documentation.
> 

While at it, I will also document the existing limitation that was 
discussed above too.

>> Cheers,
>> Matt
> 


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 21:17   ` Fernando Fernandez Mancera
@ 2026-09-28 21:22     ` Matthieu Baerts
  0 siblings, 0 replies; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-28 21:22 UTC (permalink / raw)
  To: Fernando Fernandez Mancera, Netfilter Devel, Netfilter Coreteam

Hi Fernando,

On 28/09/2026 23:17, Fernando Fernandez Mancera wrote:
> On 9/28/26 10:54 PM, Fernando Fernandez Mancera wrote:
>> On 9/28/26 2:13 PM, Matthieu Baerts wrote:
>>> Hello Netfilter devs,
>>>
>>> First, thank you for maintaining Netfilter in the kernel and the
>>> userspace tools.
>>>
>>> The MPTCP selftests are switching from IPTables to NFTables, and
>>> Clashiko reported that this part of a rule wouldn't match anything:
>>>
>>>    tcp option mptcp subtype remove-addr drop
>>>
>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>>> added to the TCP options, it is added after an MPTCP DSS option (type
>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>>> handle that, because it stops processing other TCP options when the
>>> expected type is found:
>>>
>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>>      ...
>>>     for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>>>         optl = optlen(opt, i);
>>>
>>>         if (priv->type != opt[i])
>>>             continue;
>>>         ...
>>>         return;  // <== it will look at the first MPTCP option
>>>     }
>>>      ...
>>> }
>>>
>>> It looks like it shouldn't stop if the wrong subtype is found, but the
>>> subtype is not compared there if I'm not mistaken. Should there be a fix
>>> to support this case?
>>>
>>> For more details about the Clashiko report:
>>>
>>> https://lore.kernel.org/179058242453.3145.1450534124183352369@kernel.org
>>>
>>>
>>> BTW, I'm probably missing something, but adding the following doesn't
>>> seem to have any effect on my side:
>>>
>>>    nft add table ip filter
>>>    nft add chain ip filter OUTPUT \
>>>      { type filter hook output priority filter; policy accept; }
>>>    nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
>>>
>>> Same with "mptcp subtype mp-capable", other subtypes or other types like
>>> "timestamp". But if I use ...
>>>
>>>    nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
>>>
>>> ... then the drop is effective. Any idea what I'm missing here? :)
>>>
>>>
>>> (Also, it looks like the MPTCP suboptions are not documented in the nft
>>> man page :) )
>>>
>>
>> Tomorrow there should be a patch in netfilter-devel adding it, we
>> definitively should put more work into documentation.
>>
> 
> While at it, I will also document the existing limitation that was
> discussed above too.

Great, thank you for looking at all this :)

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 20:56   ` Fernando Fernandez Mancera
@ 2026-09-28 21:22     ` Florian Westphal
  2026-09-28 21:23       ` Fernando Fernandez Mancera
  0 siblings, 1 reply; 22+ messages in thread
From: Florian Westphal @ 2026-09-28 21:22 UTC (permalink / raw)
  To: Fernando Fernandez Mancera
  Cc: Matthieu Baerts, Netfilter Devel, Netfilter Coreteam

Fernando Fernandez Mancera <fmancera@suse.de> wrote:
> On 9/28/26 3:13 PM, Florian Westphal wrote:
> > > 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> > > (type 30) options in the TCP options. It looks like Netfilter doesn't
> > > handle that, because it stops processing other TCP options when the
> > > expected type is found:
> > 
> > Yes, this won't work. 'tcp option X' extracts the
> > option X.
> > 
> > I don't see how this could be fixed within the limitations of the
> > architecture.  Just use bpf.
> > 
> > > It looks like it shouldn't stop if the wrong subtype is found, but the
> > > subtype is not compared there if I'm not mistaken. Should there be a fix
> > > to support this case?
> > 
> > I don't know how, unless one would extend the kernel to make it aware of
> > mptcp, which also requires userspace to pass the suboption type to look
> > for in addition to 'mptcp option'.
> > 
> 
> This won't be easy neither fast but I added this to my TODO list. It sounds
> fun. Unless someone else does it first, I will take it.

Thanks Fernando.  I haven't looked at this at all.

I think the only sensible solution is to come up with a new syntax to clarify
that we want a specific mptcp subtype and not the first mptcp option.

The problem is that "tcp option mptcp subtype" really just tells kernel
"find the first mptcp option, if any, then place the subtype into dreg".

And kernel doesn't even know what a subtype is, it just extracts data
at given offset :-/

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 21:22     ` Florian Westphal
@ 2026-09-28 21:23       ` Fernando Fernandez Mancera
  0 siblings, 0 replies; 22+ messages in thread
From: Fernando Fernandez Mancera @ 2026-09-28 21:23 UTC (permalink / raw)
  To: Florian Westphal; +Cc: Matthieu Baerts, Netfilter Devel, Netfilter Coreteam

On 9/28/26 11:22 PM, Florian Westphal wrote:
> Fernando Fernandez Mancera <fmancera@suse.de> wrote:
>> On 9/28/26 3:13 PM, Florian Westphal wrote:
>>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>>>> handle that, because it stops processing other TCP options when the
>>>> expected type is found:
>>>
>>> Yes, this won't work. 'tcp option X' extracts the
>>> option X.
>>>
>>> I don't see how this could be fixed within the limitations of the
>>> architecture.  Just use bpf.
>>>
>>>> It looks like it shouldn't stop if the wrong subtype is found, but the
>>>> subtype is not compared there if I'm not mistaken. Should there be a fix
>>>> to support this case?
>>>
>>> I don't know how, unless one would extend the kernel to make it aware of
>>> mptcp, which also requires userspace to pass the suboption type to look
>>> for in addition to 'mptcp option'.
>>>
>>
>> This won't be easy neither fast but I added this to my TODO list. It sounds
>> fun. Unless someone else does it first, I will take it.
> 
> Thanks Fernando.  I haven't looked at this at all.
> 
> I think the only sensible solution is to come up with a new syntax to clarify
> that we want a specific mptcp subtype and not the first mptcp option.
> 
> The problem is that "tcp option mptcp subtype" really just tells kernel
> "find the first mptcp option, if any, then place the subtype into dreg".
> 
> And kernel doesn't even know what a subtype is, it just extracts data
> at given offset :-/

Yes, this will likely require both userspace and kernelspace code but I 
guess it is worth doing it :)

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 12:13 Netfilter: match "tcp option" with the same type present multiple times Matthieu Baerts
  2026-09-28 13:13 ` Florian Westphal
  2026-09-28 20:54 ` Fernando Fernandez Mancera
@ 2026-09-28 21:52 ` Pablo Neira Ayuso
  2026-09-29  9:21   ` Matthieu Baerts
  2026-09-28 21:54 ` Jan Engelhardt
  3 siblings, 1 reply; 22+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-28 21:52 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Mattieu,

On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
> Hello Netfilter devs,
> 
> First, thank you for maintaining Netfilter in the kernel and the
> userspace tools.
> 
> The MPTCP selftests are switching from IPTables to NFTables, and
> Clashiko reported that this part of a rule wouldn't match anything:
> 
>   tcp option mptcp subtype remove-addr drop
> 
> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> added to the TCP options, it is added after an MPTCP DSS option (type
> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> (type 30) options in the TCP options. It looks like Netfilter doesn't
> handle that, because it stops processing other TCP options when the
> expected type is found:
> 
> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>     ...
> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> 		optl = optlen(opt, i);
> 
> 		if (priv->type != opt[i])
> 			continue;
> 		...
> 		return;  // <== it will look at the first MPTCP option
> 	}
>     ...
> }
> 
> It looks like it shouldn't stop if the wrong subtype is found, but the
> subtype is not compared there if I'm not mistaken. Should there be a fix
> to support this case?

Would it work for you if you can specify what TCP option you want to
match? ie.
 
    tcp option[2] mptcp subtype remove-addr drop
               ^
               |
               |
       allow to specify a match at a given tcp option

Or would this be too strict for your use-case and you would prefer
that you can match it anywhere?

Thanks.

> For more details about the Clashiko report:
> 
> https://lore.kernel.org/179058242453.3145.1450534124183352369@kernel.org
> 
> 
> BTW, I'm probably missing something, but adding the following doesn't
> seem to have any effect on my side:
> 
>   nft add table ip filter
>   nft add chain ip filter OUTPUT \
>     { type filter hook output priority filter; policy accept; }
>   nft insert rule ip filter OUTPUT tcp option mptcp exists counter drop
> 
> Same with "mptcp subtype mp-capable", other subtypes or other types like
> "timestamp". But if I use ...
> 
>   nft insert rule ip filter OUTPUT meta l4proto tcp counter drop
> 
> ... then the drop is effective. Any idea what I'm missing here? :)
> 
> 
> (Also, it looks like the MPTCP suboptions are not documented in the nft
> man page :) )
> 
> Cheers,
> Matt
> -- 
> Sponsored by the NGI0 Core fund.
> 
> 

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 12:13 Netfilter: match "tcp option" with the same type present multiple times Matthieu Baerts
                   ` (2 preceding siblings ...)
  2026-09-28 21:52 ` Pablo Neira Ayuso
@ 2026-09-28 21:54 ` Jan Engelhardt
  2026-09-29  9:34   ` Matthieu Baerts
  3 siblings, 1 reply; 22+ messages in thread
From: Jan Engelhardt @ 2026-09-28 21:54 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

On Monday 2026-09-28 14:13, Matthieu Baerts wrote:

>The MPTCP selftests are switching from IPTables to NFTables, and
>Clashiko reported that this part of a rule wouldn't match anything:
>
>  tcp option mptcp subtype remove-addr drop
>
>When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>added to the TCP options, it is added after an MPTCP DSS option (type
>0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>(type 30) options in the TCP options. It looks like Netfilter doesn't
>handle that, because it stops processing other TCP options when the
>expected type is found:
>
>net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>    ...
>	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>		optl = optlen(opt, i);
>
>		if (		priv->type != opt[i])
>			continue;
>		...
>		return;  // <== it will look at the first MPTCP option
>	}

Hm, neither RFC 793 nor 9293 seem to define a general processing rule
or error-handling behavior for duplicate or repeated options within
the header of a TCP segment.

MPTCP RFC 8684 seems to acknowledge that, with:

"it may not be possible to combine all desired options .. on a single
packet"

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 21:52 ` Pablo Neira Ayuso
@ 2026-09-29  9:21   ` Matthieu Baerts
  2026-09-29 10:00     ` Pablo Neira Ayuso
  0 siblings, 1 reply; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-29  9:21 UTC (permalink / raw)
  To: Pablo Neira Ayuso; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Pablo,

Thank you for your reply!

On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
> Hi Mattieu,
> 
> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
>> Hello Netfilter devs,
>>
>> First, thank you for maintaining Netfilter in the kernel and the
>> userspace tools.
>>
>> The MPTCP selftests are switching from IPTables to NFTables, and
>> Clashiko reported that this part of a rule wouldn't match anything:
>>
>>   tcp option mptcp subtype remove-addr drop
>>
>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>> added to the TCP options, it is added after an MPTCP DSS option (type
>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>> handle that, because it stops processing other TCP options when the
>> expected type is found:
>>
>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>     ...
>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>> 		optl = optlen(opt, i);
>>
>> 		if (priv->type != opt[i])
>> 			continue;
>> 		...
>> 		return;  // <== it will look at the first MPTCP option
>> 	}
>>     ...
>> }
>>
>> It looks like it shouldn't stop if the wrong subtype is found, but the
>> subtype is not compared there if I'm not mistaken. Should there be a fix
>> to support this case?
> 
> Would it work for you if you can specify what TCP option you want to
> match? ie.
>  
>     tcp option[2] mptcp subtype remove-addr drop
>                ^
>                |
>                |
>        allow to specify a match at a given tcp option

In my specific use-case, yes it would work. Some MPTCP sub-options will
always be after another one... when the Linux stack is used.

> Or would this be too strict for your use-case and you would prefer
> that you can match it anywhere?

I think it would make more sense to match it anywhere. For example, an
MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
could reorder the options as there is no imposed order.

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-28 21:54 ` Jan Engelhardt
@ 2026-09-29  9:34   ` Matthieu Baerts
  0 siblings, 0 replies; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-29  9:34 UTC (permalink / raw)
  To: Jan Engelhardt; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Jan,

On 28/09/2026 23:54, Jan Engelhardt wrote:
> On Monday 2026-09-28 14:13, Matthieu Baerts wrote:
> 
>> The MPTCP selftests are switching from IPTables to NFTables, and
>> Clashiko reported that this part of a rule wouldn't match anything:
>>
>>  tcp option mptcp subtype remove-addr drop
>>
>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>> added to the TCP options, it is added after an MPTCP DSS option (type
>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>> handle that, because it stops processing other TCP options when the
>> expected type is found:
>>
>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>    ...
>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>> 		optl = optlen(opt, i);
>>
>> 		if (		priv->type != opt[i])
>> 			continue;
>> 		...
>> 		return;  // <== it will look at the first MPTCP option
>> 	}
> 
> Hm, neither RFC 793 nor 9293 seem to define a general processing rule
> or error-handling behavior for duplicate or repeated options within
> the header of a TCP segment.

From what I understood at the IETF, new TCP extensions can only be
assigned one "type". If multiple message types should be passed, the
subtype should be either extracted from the size, or by using another
TLV inside. (MPTCP is using the two techniques.)

(Old extensions like SACK and echo are exceptions, because they are old
ones.)

> MPTCP RFC 8684 seems to acknowledge that, with:
> 
> "it may not be possible to combine all desired options .. on a single
> packet"

The MPTCP RFC specifies this mainly because of the TCP options space
restriction (40 bytes). For example, the ADD_ADDR option can take up to
30 bytes (+2 for the alignment). If you add TCP Timestamp (10 + 2 for
the alignment) or SACK, you are over the limit.

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-29  9:21   ` Matthieu Baerts
@ 2026-09-29 10:00     ` Pablo Neira Ayuso
  2026-09-29 10:40       ` Matthieu Baerts
  0 siblings, 1 reply; 22+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-29 10:00 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
> Hi Pablo,
> 
> Thank you for your reply!
> 
> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
> > Hi Mattieu,
> > 
> > On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
> >> Hello Netfilter devs,
> >>
> >> First, thank you for maintaining Netfilter in the kernel and the
> >> userspace tools.
> >>
> >> The MPTCP selftests are switching from IPTables to NFTables, and
> >> Clashiko reported that this part of a rule wouldn't match anything:
> >>
> >>   tcp option mptcp subtype remove-addr drop
> >>
> >> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> >> added to the TCP options, it is added after an MPTCP DSS option (type
> >> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> >> (type 30) options in the TCP options. It looks like Netfilter doesn't
> >> handle that, because it stops processing other TCP options when the
> >> expected type is found:
> >>
> >> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
> >>     ...
> >> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> >> 		optl = optlen(opt, i);
> >>
> >> 		if (priv->type != opt[i])
> >> 			continue;
> >> 		...
> >> 		return;  // <== it will look at the first MPTCP option
> >> 	}
> >>     ...
> >> }
> >>
> >> It looks like it shouldn't stop if the wrong subtype is found, but the
> >> subtype is not compared there if I'm not mistaken. Should there be a fix
> >> to support this case?
> > 
> > Would it work for you if you can specify what TCP option you want to
> > match? ie.
> >  
> >     tcp option[2] mptcp subtype remove-addr drop
> >                ^
> >                |
> >                |
> >        allow to specify a match at a given tcp option
> 
> In my specific use-case, yes it would work. Some MPTCP sub-options will
> always be after another one... when the Linux stack is used.

This would require a smaller patch.

> > Or would this be too strict for your use-case and you would prefer
> > that you can match it anywhere?
> 
> I think it would make more sense to match it anywhere. For example, an
> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
> could reorder the options as there is no imposed order.

I think this would require a MPTCP subtype parser, ie. make kernel
aware of MPTCP subtypes.

I would support both of them eventually? Starting by the one above
that allows to match at a given position which looks simpler to
support to me. Then, look into adding a MPTCP subtype parser later?

Thanks.

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-29 10:00     ` Pablo Neira Ayuso
@ 2026-09-29 10:40       ` Matthieu Baerts
  2026-09-29 12:03         ` Pablo Neira Ayuso
  0 siblings, 1 reply; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-29 10:40 UTC (permalink / raw)
  To: Pablo Neira Ayuso; +Cc: Netfilter Devel, Netfilter Coreteam

On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
> On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
>> Hi Pablo,
>>
>> Thank you for your reply!
>>
>> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
>>> Hi Mattieu,
>>>
>>> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
>>>> Hello Netfilter devs,
>>>>
>>>> First, thank you for maintaining Netfilter in the kernel and the
>>>> userspace tools.
>>>>
>>>> The MPTCP selftests are switching from IPTables to NFTables, and
>>>> Clashiko reported that this part of a rule wouldn't match anything:
>>>>
>>>>   tcp option mptcp subtype remove-addr drop
>>>>
>>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>>>> added to the TCP options, it is added after an MPTCP DSS option (type
>>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>>>> handle that, because it stops processing other TCP options when the
>>>> expected type is found:
>>>>
>>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>>>     ...
>>>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>>>> 		optl = optlen(opt, i);
>>>>
>>>> 		if (priv->type != opt[i])
>>>> 			continue;
>>>> 		...
>>>> 		return;  // <== it will look at the first MPTCP option
>>>> 	}
>>>>     ...
>>>> }
>>>>
>>>> It looks like it shouldn't stop if the wrong subtype is found, but the
>>>> subtype is not compared there if I'm not mistaken. Should there be a fix
>>>> to support this case?
>>>
>>> Would it work for you if you can specify what TCP option you want to
>>> match? ie.
>>>  
>>>     tcp option[2] mptcp subtype remove-addr drop
>>>                ^
>>>                |
>>>                |
>>>        allow to specify a match at a given tcp option
>>
>> In my specific use-case, yes it would work. Some MPTCP sub-options will
>> always be after another one... when the Linux stack is used.
> 
> This would require a smaller patch.

I understand. But I guess that still means adding a new option.

In my case, it would work, but this might be confusing for people who
want to use this filter: they need to know if a suboption can be used
with another one. I guess they could have multiple rules to check the
presence of a suboption in the first matched type, and in a second one,
but that seems confusing. Or maybe not?

>>> Or would this be too strict for your use-case and you would prefer
>>> that you can match it anywhere?
>>
>> I think it would make more sense to match it anywhere. For example, an
>> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
>> could reorder the options as there is no imposed order.
> 
> I think this would require a MPTCP subtype parser, ie. make kernel
> aware of MPTCP subtypes.

In nft_exthdr_tcp_eval(), why can the comparison not be done there
instead of copying the data in a register, and compare later on? If the
comparison is done directly for a given type and is different, the code
could continue and check for the same type. (Or store in multiple
registers.)

I don't know if you need something specific to MPTCP, maybe just a way
to check all options with the same type.

> I would support both of them eventually? Starting by the one above
> that allows to match at a given position which looks simpler to
> support to me. Then, look into adding a MPTCP subtype parser later?

If the second one can be implemented in a simple way, perhaps the first
one is not needed? If not, yes, good idea to start with the first one,
which might be enough for most people.

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-29 10:40       ` Matthieu Baerts
@ 2026-09-29 12:03         ` Pablo Neira Ayuso
  2026-09-30 18:33           ` Matthieu Baerts
  0 siblings, 1 reply; 22+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-29 12:03 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
> On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
> > On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
> >> Hi Pablo,
> >>
> >> Thank you for your reply!
> >>
> >> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
> >>> Hi Mattieu,
> >>>
> >>> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
> >>>> Hello Netfilter devs,
> >>>>
> >>>> First, thank you for maintaining Netfilter in the kernel and the
> >>>> userspace tools.
> >>>>
> >>>> The MPTCP selftests are switching from IPTables to NFTables, and
> >>>> Clashiko reported that this part of a rule wouldn't match anything:
> >>>>
> >>>>   tcp option mptcp subtype remove-addr drop
> >>>>
> >>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> >>>> added to the TCP options, it is added after an MPTCP DSS option (type
> >>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> >>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
> >>>> handle that, because it stops processing other TCP options when the
> >>>> expected type is found:
> >>>>
> >>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
> >>>>     ...
> >>>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> >>>> 		optl = optlen(opt, i);
> >>>>
> >>>> 		if (priv->type != opt[i])
> >>>> 			continue;
> >>>> 		...
> >>>> 		return;  // <== it will look at the first MPTCP option
> >>>> 	}
> >>>>     ...
> >>>> }
> >>>>
> >>>> It looks like it shouldn't stop if the wrong subtype is found, but the
> >>>> subtype is not compared there if I'm not mistaken. Should there be a fix
> >>>> to support this case?
> >>>
> >>> Would it work for you if you can specify what TCP option you want to
> >>> match? ie.
> >>>  
> >>>     tcp option[2] mptcp subtype remove-addr drop
> >>>                ^
> >>>                |
> >>>                |
> >>>        allow to specify a match at a given tcp option
> >>
> >> In my specific use-case, yes it would work. Some MPTCP sub-options will
> >> always be after another one... when the Linux stack is used.
> > 
> > This would require a smaller patch.
> 
> I understand. But I guess that still means adding a new option.
> 
> In my case, it would work, but this might be confusing for people who
> want to use this filter: they need to know if a suboption can be used
> with another one. I guess they could have multiple rules to check the
> presence of a suboption in the first matched type, and in a second one,
> but that seems confusing. Or maybe not?

You mean, the might want to validate the entire list of suboptions in
a particular order, correct? ie. first suboption A then suboption B.

> >>> Or would this be too strict for your use-case and you would prefer
> >>> that you can match it anywhere?
> >>
> >> I think it would make more sense to match it anywhere. For example, an
> >> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
> >> could reorder the options as there is no imposed order.
> > 
> > I think this would require a MPTCP subtype parser, ie. make kernel
> > aware of MPTCP subtypes.
> 
> In nft_exthdr_tcp_eval(), why can the comparison not be done there
> instead of copying the data in a register, and compare later on? If the
> comparison is done directly for a given type and is different, the code
> could continue and check for the same type. (Or store in multiple
> registers.)

In nftables, the fetch then cmp instructions are splitted, ie. you
first retrieve then you compare (or make a set lookup with it).
Builtin comparison should be possible, but it defeats the integration
with the set infrastructure.

> I don't know if you need something specific to MPTCP, maybe just a way
> to check all options with the same type.

Going back to this:

"When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
added to the TCP options, it is added after an MPTCP DSS option (type
0x30, len >= 8, subtype 0x2)"

Would you like to check for this particular sequence, right? Then
position-based solution would work, because this would allow to match
strictly:

pos = 0, suboption DSS
pos = 1, suboption REMOVE_ADDR

This ruleset might break the protocol if there is ever a suboption
before DSS.

Loose mode is more protocol designer friendly as it does not constrain
future updates in the protocol:

"please find MP-TCP suboption REMOVE_ADDR for me"

but security folks might probably say this is ... loose, because it
does not validate DSS before and REMOVE_ADDR might come anywhere.
Protocol designer would say that this is good for them because this
firewall does not constrain future enhancements of the protocol (still
raw expressions are possible, then this last statement does not hold
anymore).

Which one would you pick? :-)

> > I would support both of them eventually? Starting by the one above
> > that allows to match at a given position which looks simpler to
> > support to me. Then, look into adding a MPTCP subtype parser later?
> 
> If the second one can be implemented in a simple way, perhaps the first
> one is not needed? If not, yes, good idea to start with the first one,
> which might be enough for most people.

Both solutions are feasible IMO, the second needs a bit more work,
that's all.

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-29 12:03         ` Pablo Neira Ayuso
@ 2026-09-30 18:33           ` Matthieu Baerts
  2026-10-06 23:29             ` Pablo Neira Ayuso
  0 siblings, 1 reply; 22+ messages in thread
From: Matthieu Baerts @ 2026-09-30 18:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Pablo,

Thank you for your reply!

On 29/09/2026 14:03, Pablo Neira Ayuso wrote:
> On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
>> On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
>>> On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
>>>> Hi Pablo,
>>>>
>>>> Thank you for your reply!
>>>>
>>>> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
>>>>> Hi Mattieu,
>>>>>
>>>>> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
>>>>>> Hello Netfilter devs,
>>>>>>
>>>>>> First, thank you for maintaining Netfilter in the kernel and the
>>>>>> userspace tools.
>>>>>>
>>>>>> The MPTCP selftests are switching from IPTables to NFTables, and
>>>>>> Clashiko reported that this part of a rule wouldn't match anything:
>>>>>>
>>>>>>   tcp option mptcp subtype remove-addr drop
>>>>>>
>>>>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>>>>>> added to the TCP options, it is added after an MPTCP DSS option (type
>>>>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>>>>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>>>>>> handle that, because it stops processing other TCP options when the
>>>>>> expected type is found:
>>>>>>
>>>>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>>>>>     ...
>>>>>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>>>>>> 		optl = optlen(opt, i);
>>>>>>
>>>>>> 		if (priv->type != opt[i])
>>>>>> 			continue;
>>>>>> 		...
>>>>>> 		return;  // <== it will look at the first MPTCP option
>>>>>> 	}
>>>>>>     ...
>>>>>> }
>>>>>>
>>>>>> It looks like it shouldn't stop if the wrong subtype is found, but the
>>>>>> subtype is not compared there if I'm not mistaken. Should there be a fix
>>>>>> to support this case?
>>>>>
>>>>> Would it work for you if you can specify what TCP option you want to
>>>>> match? ie.
>>>>>  
>>>>>     tcp option[2] mptcp subtype remove-addr drop
>>>>>                ^
>>>>>                |
>>>>>                |
>>>>>        allow to specify a match at a given tcp option
>>>>
>>>> In my specific use-case, yes it would work. Some MPTCP sub-options will
>>>> always be after another one... when the Linux stack is used.
>>>
>>> This would require a smaller patch.
>>
>> I understand. But I guess that still means adding a new option.
>>
>> In my case, it would work, but this might be confusing for people who
>> want to use this filter: they need to know if a suboption can be used
>> with another one. I guess they could have multiple rules to check the
>> presence of a suboption in the first matched type, and in a second one,
>> but that seems confusing. Or maybe not?
> 
> You mean, the might want to validate the entire list of suboptions in
> a particular order, correct? ie. first suboption A then suboption B.

Yes, that's one possibility. But initially, I was thinking about having
two rules (or a set?) to match "the MPTCP subtype is either in the first
MPTCP option, or the second one (if any)". (Or nft could always say with
MPTCP by default, check also for a second MPTCP option, if any)

>>>>> Or would this be too strict for your use-case and you would prefer
>>>>> that you can match it anywhere?
>>>>
>>>> I think it would make more sense to match it anywhere. For example, an
>>>> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
>>>> could reorder the options as there is no imposed order.
>>>
>>> I think this would require a MPTCP subtype parser, ie. make kernel
>>> aware of MPTCP subtypes.
>>
>> In nft_exthdr_tcp_eval(), why can the comparison not be done there
>> instead of copying the data in a register, and compare later on? If the
>> comparison is done directly for a given type and is different, the code
>> could continue and check for the same type. (Or store in multiple
>> registers.)
> 
> In nftables, the fetch then cmp instructions are splitted, ie. you
> first retrieve then you compare (or make a set lookup with it).
> Builtin comparison should be possible, but it defeats the integration
> with the set infrastructure.

OK, I see.

If someone gives ...

  tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop

... a bitmap could be used, but then that's specific to MPTCP I suppose.

>> I don't know if you need something specific to MPTCP, maybe just a way
>> to check all options with the same type.
> 
> Going back to this:
> 
> "When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> added to the TCP options, it is added after an MPTCP DSS option (type
> 0x30, len >= 8, subtype 0x2)"
> 
> Would you like to check for this particular sequence, right? Then

In my case, no particularly -- I just want to count/drop REMOVE_ADDR to
catch kernel regressions in the selftests -- but I can.

> position-based solution would work, because this would allow to match
> strictly:
> 
> pos = 0, suboption DSS
> pos = 1, suboption REMOVE_ADDR
> 
> This ruleset might break the protocol if there is ever a suboption
> before DSS.

Indeed. Technically, with MPTCP, we could send the REMOVE_ADDR alone
(without DSS), or before the DSS, or with another one in between if
there is room.

> Loose mode is more protocol designer friendly as it does not constrain
> future updates in the protocol:
> 
> "please find MP-TCP suboption REMOVE_ADDR for me"
> 
> but security folks might probably say this is ... loose, because it
> does not validate DSS before and REMOVE_ADDR might come anywhere.
> Protocol designer would say that this is good for them because this
> firewall does not constrain future enhancements of the protocol (still
> raw expressions are possible, then this last statement does not hold
> anymore).
> 
> Which one would you pick? :-)

I would first prefer if people don't drop packets with MPTCP options :-D

Honestly, I'm not sure what people would be interested in doing in
production, but I guess it would be more: "please match packet with
MPTCP suboption X, no matter the order". If people are interested in a
specific packet where MPTCP suboptions are in a specific order, with a
specific length, I *think* they can use a raw payload expression
instead, no?

In my case, I just wanted to drop packets with a specific MPTCP
suboption. It happens that I know exactly which packet I want to block,
and I can use a raw payload expression, but a simpler rule would be
welcome :)

>>> I would support both of them eventually? Starting by the one above
>>> that allows to match at a given position which looks simpler to
>>> support to me. Then, look into adding a MPTCP subtype parser later?
>>
>> If the second one can be implemented in a simple way, perhaps the first
>> one is not needed? If not, yes, good idea to start with the first one,
>> which might be enough for most people.
> 
> Both solutions are feasible IMO, the second needs a bit more work,
> that's all.

Understood!

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-09-30 18:33           ` Matthieu Baerts
@ 2026-10-06 23:29             ` Pablo Neira Ayuso
  2026-10-07  8:01               ` Fernando Fernandez Mancera
  2026-10-07  8:17               ` Matthieu Baerts
  0 siblings, 2 replies; 22+ messages in thread
From: Pablo Neira Ayuso @ 2026-10-06 23:29 UTC (permalink / raw)
  To: Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam

Hi Matthieu,

On Wed, Sep 30, 2026 at 08:33:36PM +0200, Matthieu Baerts wrote:
> Hi Pablo,
> 
> Thank you for your reply!
> 
> On 29/09/2026 14:03, Pablo Neira Ayuso wrote:
> > On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
> >> On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
> >>> On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
> >>>> Hi Pablo,
> >>>>
> >>>> Thank you for your reply!
> >>>>
> >>>> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
> >>>>> Hi Mattieu,
> >>>>>
> >>>>> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
> >>>>>> Hello Netfilter devs,
> >>>>>>
> >>>>>> First, thank you for maintaining Netfilter in the kernel and the
> >>>>>> userspace tools.
> >>>>>>
> >>>>>> The MPTCP selftests are switching from IPTables to NFTables, and
> >>>>>> Clashiko reported that this part of a rule wouldn't match anything:
> >>>>>>
> >>>>>>   tcp option mptcp subtype remove-addr drop
> >>>>>>
> >>>>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> >>>>>> added to the TCP options, it is added after an MPTCP DSS option (type
> >>>>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> >>>>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
> >>>>>> handle that, because it stops processing other TCP options when the
> >>>>>> expected type is found:
> >>>>>>
> >>>>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
> >>>>>>     ...
> >>>>>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> >>>>>> 		optl = optlen(opt, i);
> >>>>>>
> >>>>>> 		if (priv->type != opt[i])
> >>>>>> 			continue;
> >>>>>> 		...
> >>>>>> 		return;  // <== it will look at the first MPTCP option
> >>>>>> 	}
> >>>>>>     ...
> >>>>>> }
> >>>>>>
> >>>>>> It looks like it shouldn't stop if the wrong subtype is found, but the
> >>>>>> subtype is not compared there if I'm not mistaken. Should there be a fix
> >>>>>> to support this case?
> >>>>>
> >>>>> Would it work for you if you can specify what TCP option you want to
> >>>>> match? ie.
> >>>>>  
> >>>>>     tcp option[2] mptcp subtype remove-addr drop
> >>>>>                ^
> >>>>>                |
> >>>>>                |
> >>>>>        allow to specify a match at a given tcp option
> >>>>
> >>>> In my specific use-case, yes it would work. Some MPTCP sub-options will
> >>>> always be after another one... when the Linux stack is used.
> >>>
> >>> This would require a smaller patch.
> >>
> >> I understand. But I guess that still means adding a new option.
> >>
> >> In my case, it would work, but this might be confusing for people who
> >> want to use this filter: they need to know if a suboption can be used
> >> with another one. I guess they could have multiple rules to check the
> >> presence of a suboption in the first matched type, and in a second one,
> >> but that seems confusing. Or maybe not?
> > 
> > You mean, the might want to validate the entire list of suboptions in
> > a particular order, correct? ie. first suboption A then suboption B.
> 
> Yes, that's one possibility. But initially, I was thinking about having
> two rules (or a set?) to match "the MPTCP subtype is either in the first
> MPTCP option, or the second one (if any)". (Or nft could always say with
> MPTCP by default, check also for a second MPTCP option, if any)
>
> >>>>> Or would this be too strict for your use-case and you would prefer
> >>>>> that you can match it anywhere?
> >>>>
> >>>> I think it would make more sense to match it anywhere. For example, an
> >>>> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
> >>>> could reorder the options as there is no imposed order.
> >>>
> >>> I think this would require a MPTCP subtype parser, ie. make kernel
> >>> aware of MPTCP subtypes.
> >>
> >> In nft_exthdr_tcp_eval(), why can the comparison not be done there
> >> instead of copying the data in a register, and compare later on? If the
> >> comparison is done directly for a given type and is different, the code
> >> could continue and check for the same type. (Or store in multiple
> >> registers.)
> > 
> > In nftables, the fetch then cmp instructions are splitted, ie. you
> > first retrieve then you compare (or make a set lookup with it).
> > Builtin comparison should be possible, but it defeats the integration
> > with the set infrastructure.
> 
> OK, I see.
> 
> If someone gives ...
> 
>   tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop
> 
> ... a bitmap could be used, but then that's specific to MPTCP I suppose.

I think this makes sense from user perspective.

The nft_exthdr expression needs iterate over the whole list of tcp
options, then if type is 30, make a set lookup to check if the subtype
value is in the set.

Something like:

 - loop
 |  r1 <- exthdr     [ if no more tcp options, break loop ]
 |  r2 <- lookup(r1) [ if found, break loop ]

but exthdr need to be taugh to resume from the last visited tcp
option. A new loop expression would wrap these two exthdr and lookup
expression (similar to nft_inner) and store the iteration context (ie.
last visited tcp option to continue from there).

exthdr tcp option support needs advertise a new NFT_EXPR_LOOP flag, so
it can be used with within this new nft_loop expression.

Another question: Would you still like to validate that DSS comes
before REMOVE_ADDR?

> >> I don't know if you need something specific to MPTCP, maybe just a way
> >> to check all options with the same type.
> > 
> > Going back to this:
> > 
> > "When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> > added to the TCP options, it is added after an MPTCP DSS option (type
> > 0x30, len >= 8, subtype 0x2)"
> > 
> > Would you like to check for this particular sequence, right? Then
> 
> In my case, no particularly -- I just want to count/drop REMOVE_ADDR to
> catch kernel regressions in the selftests -- but I can.
> 
> > position-based solution would work, because this would allow to match
> > strictly:
> > 
> > pos = 0, suboption DSS
> > pos = 1, suboption REMOVE_ADDR
> > 
> > This ruleset might break the protocol if there is ever a suboption
> > before DSS.
> 
> Indeed. Technically, with MPTCP, we could send the REMOVE_ADDR alone
> (without DSS), or before the DSS, or with another one in between if
> there is room.

Position matching is too strict.

You can still make it with the raw expression, but I think the native
matching representation should be more flexible too as you describe.

> > Loose mode is more protocol designer friendly as it does not constrain
> > future updates in the protocol:
> > 
> > "please find MP-TCP suboption REMOVE_ADDR for me"
> > 
> > but security folks might probably say this is ... loose, because it
> > does not validate DSS before and REMOVE_ADDR might come anywhere.
> > Protocol designer would say that this is good for them because this
> > firewall does not constrain future enhancements of the protocol (still
> > raw expressions are possible, then this last statement does not hold
> > anymore).
> > 
> > Which one would you pick? :-)
> 
> I would first prefer if people don't drop packets with MPTCP options :-D
> 
> Honestly, I'm not sure what people would be interested in doing in
> production, but I guess it would be more: "please match packet with
> MPTCP suboption X, no matter the order".

Yes, as discussed, this is more protocol friendly too in terms of
extensibility.

> If people are interested in a specific packet where MPTCP suboptions
> are in a specific order, with a specific length, I *think* they can
> use a raw payload expression instead, no?
> 
> In my case, I just wanted to drop packets with a specific MPTCP
> suboption. It happens that I know exactly which packet I want to block,
> and I can use a raw payload expression, but a simpler rule would be
> welcome :)

Makes sense.

This would need an extension in the kernel as Florian has anticipated.

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-10-06 23:29             ` Pablo Neira Ayuso
@ 2026-10-07  8:01               ` Fernando Fernandez Mancera
  2026-10-07 10:47                 ` Pablo Neira Ayuso
  2026-10-07  8:17               ` Matthieu Baerts
  1 sibling, 1 reply; 22+ messages in thread
From: Fernando Fernandez Mancera @ 2026-10-07  8:01 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Matthieu Baerts; +Cc: Netfilter Devel, Netfilter Coreteam



On 10/7/26 1:29 AM, Pablo Neira Ayuso wrote:
> Hi Matthieu,
> 
> On Wed, Sep 30, 2026 at 08:33:36PM +0200, Matthieu Baerts wrote:
>> Hi Pablo,
>>
>> Thank you for your reply!
>>
>> On 29/09/2026 14:03, Pablo Neira Ayuso wrote:
>>> On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
>>>> On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
>>>>> On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
>>>>>> Hi Pablo,
>>>>>>
>>>>>> Thank you for your reply!
>>>>>>
>>>>>> On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
>>>>>>> Hi Mattieu,
>>>>>>>
>>>>>>> On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
>>>>>>>> Hello Netfilter devs,
>>>>>>>>
>>>>>>>> First, thank you for maintaining Netfilter in the kernel and the
>>>>>>>> userspace tools.
>>>>>>>>
>>>>>>>> The MPTCP selftests are switching from IPTables to NFTables, and
>>>>>>>> Clashiko reported that this part of a rule wouldn't match anything:
>>>>>>>>
>>>>>>>>    tcp option mptcp subtype remove-addr drop
>>>>>>>>
>>>>>>>> When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>>>>>>>> added to the TCP options, it is added after an MPTCP DSS option (type
>>>>>>>> 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
>>>>>>>> (type 30) options in the TCP options. It looks like Netfilter doesn't
>>>>>>>> handle that, because it stops processing other TCP options when the
>>>>>>>> expected type is found:
>>>>>>>>
>>>>>>>> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>>>>>>>>      ...
>>>>>>>> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
>>>>>>>> 		optl = optlen(opt, i);
>>>>>>>>
>>>>>>>> 		if (priv->type != opt[i])
>>>>>>>> 			continue;
>>>>>>>> 		...
>>>>>>>> 		return;  // <== it will look at the first MPTCP option
>>>>>>>> 	}
>>>>>>>>      ...
>>>>>>>> }
>>>>>>>>
>>>>>>>> It looks like it shouldn't stop if the wrong subtype is found, but the
>>>>>>>> subtype is not compared there if I'm not mistaken. Should there be a fix
>>>>>>>> to support this case?
>>>>>>>
>>>>>>> Would it work for you if you can specify what TCP option you want to
>>>>>>> match? ie.
>>>>>>>   
>>>>>>>      tcp option[2] mptcp subtype remove-addr drop
>>>>>>>                 ^
>>>>>>>                 |
>>>>>>>                 |
>>>>>>>         allow to specify a match at a given tcp option
>>>>>>
>>>>>> In my specific use-case, yes it would work. Some MPTCP sub-options will
>>>>>> always be after another one... when the Linux stack is used.
>>>>>
>>>>> This would require a smaller patch.
>>>>
>>>> I understand. But I guess that still means adding a new option.
>>>>
>>>> In my case, it would work, but this might be confusing for people who
>>>> want to use this filter: they need to know if a suboption can be used
>>>> with another one. I guess they could have multiple rules to check the
>>>> presence of a suboption in the first matched type, and in a second one,
>>>> but that seems confusing. Or maybe not?
>>>
>>> You mean, the might want to validate the entire list of suboptions in
>>> a particular order, correct? ie. first suboption A then suboption B.
>>
>> Yes, that's one possibility. But initially, I was thinking about having
>> two rules (or a set?) to match "the MPTCP subtype is either in the first
>> MPTCP option, or the second one (if any)". (Or nft could always say with
>> MPTCP by default, check also for a second MPTCP option, if any)
>>
>>>>>>> Or would this be too strict for your use-case and you would prefer
>>>>>>> that you can match it anywhere?
>>>>>>
>>>>>> I think it would make more sense to match it anywhere. For example, an
>>>>>> MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
>>>>>> could reorder the options as there is no imposed order.
>>>>>
>>>>> I think this would require a MPTCP subtype parser, ie. make kernel
>>>>> aware of MPTCP subtypes.
>>>>
>>>> In nft_exthdr_tcp_eval(), why can the comparison not be done there
>>>> instead of copying the data in a register, and compare later on? If the
>>>> comparison is done directly for a given type and is different, the code
>>>> could continue and check for the same type. (Or store in multiple
>>>> registers.)
>>>
>>> In nftables, the fetch then cmp instructions are splitted, ie. you
>>> first retrieve then you compare (or make a set lookup with it).
>>> Builtin comparison should be possible, but it defeats the integration
>>> with the set infrastructure.
>>
>> OK, I see.
>>
>> If someone gives ...
>>
>>    tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop
>>
>> ... a bitmap could be used, but then that's specific to MPTCP I suppose.
> 
> I think this makes sense from user perspective.
> 
> The nft_exthdr expression needs iterate over the whole list of tcp
> options, then if type is 30, make a set lookup to check if the subtype
> value is in the set.
> 
> Something like:
> 
>   - loop
>   |  r1 <- exthdr     [ if no more tcp options, break loop ]
>   |  r2 <- lookup(r1) [ if found, break loop ]
> 
> but exthdr need to be taugh to resume from the last visited tcp
> option. A new loop expression would wrap these two exthdr and lookup
> expression (similar to nft_inner) and store the iteration context (ie.
> last visited tcp option to continue from there).
> 
> exthdr tcp option support needs advertise a new NFT_EXPR_LOOP flag, so
> it can be used with within this new nft_loop expression.
> 

Wouldn't this be too complex? I was thinking in something a bit more 
generic and simpler. E.g a new NFTA_EXTHDR_INNER_TYPE which will work 
with PRESENT flag.

That way one can match MPTCP options exactly the same way TCP options 
are matched. The implementation would look first for the TCP option 30 
and then check the subtype, if it matches return 1 otherwise continue.

This is chained with a CMP_EQ expression.

I have a working patch, let me polish the code and will send it as RFC 
today so we can discuss it. I like this idea because it could be 
extended to match similar situations in IPv6 headers or SCTP.

> Another question: Would you still like to validate that DSS comes
> before REMOVE_ADDR?
> 
>>>> I don't know if you need something specific to MPTCP, maybe just a way
>>>> to check all options with the same type.
>>>
>>> Going back to this:
>>>
>>> "When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
>>> added to the TCP options, it is added after an MPTCP DSS option (type
>>> 0x30, len >= 8, subtype 0x2)"
>>>
>>> Would you like to check for this particular sequence, right? Then
>>
>> In my case, no particularly -- I just want to count/drop REMOVE_ADDR to
>> catch kernel regressions in the selftests -- but I can.
>>
>>> position-based solution would work, because this would allow to match
>>> strictly:
>>>
>>> pos = 0, suboption DSS
>>> pos = 1, suboption REMOVE_ADDR
>>>
>>> This ruleset might break the protocol if there is ever a suboption
>>> before DSS.
>>
>> Indeed. Technically, with MPTCP, we could send the REMOVE_ADDR alone
>> (without DSS), or before the DSS, or with another one in between if
>> there is room.
> 
> Position matching is too strict.
> 
> You can still make it with the raw expression, but I think the native
> matching representation should be more flexible too as you describe.
> 
>>> Loose mode is more protocol designer friendly as it does not constrain
>>> future updates in the protocol:
>>>
>>> "please find MP-TCP suboption REMOVE_ADDR for me"
>>>
>>> but security folks might probably say this is ... loose, because it
>>> does not validate DSS before and REMOVE_ADDR might come anywhere.
>>> Protocol designer would say that this is good for them because this
>>> firewall does not constrain future enhancements of the protocol (still
>>> raw expressions are possible, then this last statement does not hold
>>> anymore).
>>>
>>> Which one would you pick? :-)
>>
>> I would first prefer if people don't drop packets with MPTCP options :-D
>>
>> Honestly, I'm not sure what people would be interested in doing in
>> production, but I guess it would be more: "please match packet with
>> MPTCP suboption X, no matter the order".
> 
> Yes, as discussed, this is more protocol friendly too in terms of
> extensibility.
> 
>> If people are interested in a specific packet where MPTCP suboptions
>> are in a specific order, with a specific length, I *think* they can
>> use a raw payload expression instead, no?
>>
>> In my case, I just wanted to drop packets with a specific MPTCP
>> suboption. It happens that I know exactly which packet I want to block,
>> and I can use a raw payload expression, but a simpler rule would be
>> welcome :)
> 
> Makes sense.
> 
> This would need an extension in the kernel as Florian has anticipated.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-10-06 23:29             ` Pablo Neira Ayuso
  2026-10-07  8:01               ` Fernando Fernandez Mancera
@ 2026-10-07  8:17               ` Matthieu Baerts
  1 sibling, 0 replies; 22+ messages in thread
From: Matthieu Baerts @ 2026-10-07  8:17 UTC (permalink / raw)
  To: Pablo Neira Ayuso
  Cc: Netfilter Devel, Netfilter Coreteam, Fernando Fernandez Mancera

Hi Pablo,

Thank you for your reply!

On 07/10/2026 01:29, Pablo Neira Ayuso wrote:
> On Wed, Sep 30, 2026 at 08:33:36PM +0200, Matthieu Baerts wrote:
>> On 29/09/2026 14:03, Pablo Neira Ayuso wrote:
>>> On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
>>>> On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
>>>>> On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:

(...)

>> If someone gives ...
>>
>>   tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop
>>
>> ... a bitmap could be used, but then that's specific to MPTCP I suppose.
> 
> I think this makes sense from user perspective.
> 
> The nft_exthdr expression needs iterate over the whole list of tcp
> options, then if type is 30, make a set lookup to check if the subtype
> value is in the set.
> 
> Something like:
> 
>  - loop
>  |  r1 <- exthdr     [ if no more tcp options, break loop ]
>  |  r2 <- lookup(r1) [ if found, break loop ]
> 
> but exthdr need to be taugh to resume from the last visited tcp
> option. A new loop expression would wrap these two exthdr and lookup
> expression (similar to nft_inner) and store the iteration context (ie.
> last visited tcp option to continue from there).
> 
> exthdr tcp option support needs advertise a new NFT_EXPR_LOOP flag, so
> it can be used with within this new nft_loop expression.

Yes, this or Fernando's idea (BTW, thank you for looking at that!). Both
will need to expose a new capability to the userspace from what I
understand, so any of them are fine for me.

> Another question: Would you still like to validate that DSS comes
> before REMOVE_ADDR?

I don't think that's needed: the RFC doesn't impose an order nor a
specific combination. When multiple MPTCP options are used, they are
notifying two different things.

If people are interested in validating that option A comes before option
B, that's certainly because they are trying to find a specific pattern
-- e.g. to identify a specific stack or kernel version -- not to
identify packets of a certain protocol. For them, they can use other
techniques, but I don't think we need to increase the complexity to
support that case. (At least for MPTCP.)

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: Netfilter: match "tcp option" with the same type present multiple times
  2026-10-07  8:01               ` Fernando Fernandez Mancera
@ 2026-10-07 10:47                 ` Pablo Neira Ayuso
  0 siblings, 0 replies; 22+ messages in thread
From: Pablo Neira Ayuso @ 2026-10-07 10:47 UTC (permalink / raw)
  To: Fernando Fernandez Mancera
  Cc: Matthieu Baerts, Netfilter Devel, Netfilter Coreteam

Hi Fernando,

On Wed, Oct 07, 2026 at 10:01:51AM +0200, Fernando Fernandez Mancera wrote:
> 
> 
> On 10/7/26 1:29 AM, Pablo Neira Ayuso wrote:
> > Hi Matthieu,
> > 
> > On Wed, Sep 30, 2026 at 08:33:36PM +0200, Matthieu Baerts wrote:
> > > Hi Pablo,
> > > 
> > > Thank you for your reply!
> > > 
> > > On 29/09/2026 14:03, Pablo Neira Ayuso wrote:
> > > > On Tue, Sep 29, 2026 at 12:40:47PM +0200, Matthieu Baerts wrote:
> > > > > On 29/09/2026 12:00, Pablo Neira Ayuso wrote:
> > > > > > On Tue, Sep 29, 2026 at 11:21:15AM +0200, Matthieu Baerts wrote:
> > > > > > > Hi Pablo,
> > > > > > > 
> > > > > > > Thank you for your reply!
> > > > > > > 
> > > > > > > On 28/09/2026 23:52, Pablo Neira Ayuso wrote:
> > > > > > > > Hi Mattieu,
> > > > > > > > 
> > > > > > > > On Mon, Sep 28, 2026 at 02:13:16PM +0200, Matthieu Baerts wrote:
> > > > > > > > > Hello Netfilter devs,
> > > > > > > > > 
> > > > > > > > > First, thank you for maintaining Netfilter in the kernel and the
> > > > > > > > > userspace tools.
> > > > > > > > > 
> > > > > > > > > The MPTCP selftests are switching from IPTables to NFTables, and
> > > > > > > > > Clashiko reported that this part of a rule wouldn't match anything:
> > > > > > > > > 
> > > > > > > > >    tcp option mptcp subtype remove-addr drop
> > > > > > > > > 
> > > > > > > > > When an MPTCP REMOVE_ADDR suboption (type 0x30, len >=4, subtype 0x4) is
> > > > > > > > > added to the TCP options, it is added after an MPTCP DSS option (type
> > > > > > > > > 0x30, len >= 8, subtype 0x2). In other words, there will be two MPTCP
> > > > > > > > > (type 30) options in the TCP options. It looks like Netfilter doesn't
> > > > > > > > > handle that, because it stops processing other TCP options when the
> > > > > > > > > expected type is found:
> > > > > > > > > 
> > > > > > > > > net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
> > > > > > > > >      ...
> > > > > > > > > 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> > > > > > > > > 		optl = optlen(opt, i);
> > > > > > > > > 
> > > > > > > > > 		if (priv->type != opt[i])
> > > > > > > > > 			continue;
> > > > > > > > > 		...
> > > > > > > > > 		return;  // <== it will look at the first MPTCP option
> > > > > > > > > 	}
> > > > > > > > >      ...
> > > > > > > > > }
> > > > > > > > > 
> > > > > > > > > It looks like it shouldn't stop if the wrong subtype is found, but the
> > > > > > > > > subtype is not compared there if I'm not mistaken. Should there be a fix
> > > > > > > > > to support this case?
> > > > > > > > 
> > > > > > > > Would it work for you if you can specify what TCP option you want to
> > > > > > > > match? ie.
> > > > > > > >      tcp option[2] mptcp subtype remove-addr drop
> > > > > > > >                 ^
> > > > > > > >                 |
> > > > > > > >                 |
> > > > > > > >         allow to specify a match at a given tcp option
> > > > > > > 
> > > > > > > In my specific use-case, yes it would work. Some MPTCP sub-options will
> > > > > > > always be after another one... when the Linux stack is used.
> > > > > > 
> > > > > > This would require a smaller patch.
> > > > > 
> > > > > I understand. But I guess that still means adding a new option.
> > > > > 
> > > > > In my case, it would work, but this might be confusing for people who
> > > > > want to use this filter: they need to know if a suboption can be used
> > > > > with another one. I guess they could have multiple rules to check the
> > > > > presence of a suboption in the first matched type, and in a second one,
> > > > > but that seems confusing. Or maybe not?
> > > > 
> > > > You mean, the might want to validate the entire list of suboptions in
> > > > a particular order, correct? ie. first suboption A then suboption B.
> > > 
> > > Yes, that's one possibility. But initially, I was thinking about having
> > > two rules (or a set?) to match "the MPTCP subtype is either in the first
> > > MPTCP option, or the second one (if any)". (Or nft could always say with
> > > MPTCP by default, check also for a second MPTCP option, if any)
> > > 
> > > > > > > > Or would this be too strict for your use-case and you would prefer
> > > > > > > > that you can match it anywhere?
> > > > > > > 
> > > > > > > I think it would make more sense to match it anywhere. For example, an
> > > > > > > MP_RESET can be used alone, or after an MP_FASTCLOSE. Plus some stacks
> > > > > > > could reorder the options as there is no imposed order.
> > > > > > 
> > > > > > I think this would require a MPTCP subtype parser, ie. make kernel
> > > > > > aware of MPTCP subtypes.
> > > > > 
> > > > > In nft_exthdr_tcp_eval(), why can the comparison not be done there
> > > > > instead of copying the data in a register, and compare later on? If the
> > > > > comparison is done directly for a given type and is different, the code
> > > > > could continue and check for the same type. (Or store in multiple
> > > > > registers.)
> > > > 
> > > > In nftables, the fetch then cmp instructions are splitted, ie. you
> > > > first retrieve then you compare (or make a set lookup with it).
> > > > Builtin comparison should be possible, but it defeats the integration
> > > > with the set infrastructure.
> > > 
> > > OK, I see.
> > > 
> > > If someone gives ...
> > > 
> > >    tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop
> > > 
> > > ... a bitmap could be used, but then that's specific to MPTCP I suppose.
> > 
> > I think this makes sense from user perspective.
> > 
> > The nft_exthdr expression needs iterate over the whole list of tcp
> > options, then if type is 30, make a set lookup to check if the subtype
> > value is in the set.
> > 
> > Something like:
> > 
> >   - loop
> >   |  r1 <- exthdr     [ if no more tcp options, break loop ]
> >   |  r2 <- lookup(r1) [ if found, break loop ]
> > 
> > but exthdr need to be taugh to resume from the last visited tcp
> > option. A new loop expression would wrap these two exthdr and lookup
> > expression (similar to nft_inner) and store the iteration context (ie.
> > last visited tcp option to continue from there).
> > 
> > exthdr tcp option support needs advertise a new NFT_EXPR_LOOP flag, so
> > it can be used with within this new nft_loop expression.
> > 
> 
> Wouldn't this be too complex? I was thinking in something a bit more generic
> and simpler. E.g a new NFTA_EXTHDR_INNER_TYPE which will work with PRESENT
> flag.

That's fine, this is needed for matching the subtype.

> That way one can match MPTCP options exactly the same way TCP options are
> matched. The implementation would look first for the TCP option 30 and then
> check the subtype, if it matches return 1 otherwise continue.
> 
> This is chained with a CMP_EQ expression.

I think it would be great if we support this too:

    tcp option mptcp subtype { mp-capable, mp-join, remove-addr } drop

and even combine this new feature with maps.

IIUC, TCP option 30 might come several times, one for each subtype,
that is why I suggested the loop thing, because we need to have
control on the iteration over the list of options. But if you design
copes with the requirements that we have discussed or it can be
extended later on, that's fine.

> I have a working patch, let me polish the code and will send it as RFC today
> so we can discuss it. I like this idea because it could be extended to match
> similar situations in IPv6 headers or SCTP.

Yes, this should work with other extensions too.

^ permalink raw reply	[flat|nested] 22+ messages in thread

end of thread, other threads:[~2026-10-07 10:47 UTC | newest]

Thread overview: 22+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 12:13 Netfilter: match "tcp option" with the same type present multiple times Matthieu Baerts
2026-09-28 13:13 ` Florian Westphal
2026-09-28 20:26   ` Matthieu Baerts
2026-09-28 21:08     ` Florian Westphal
2026-09-28 20:56   ` Fernando Fernandez Mancera
2026-09-28 21:22     ` Florian Westphal
2026-09-28 21:23       ` Fernando Fernandez Mancera
2026-09-28 20:54 ` Fernando Fernandez Mancera
2026-09-28 21:17   ` Fernando Fernandez Mancera
2026-09-28 21:22     ` Matthieu Baerts
2026-09-28 21:52 ` Pablo Neira Ayuso
2026-09-29  9:21   ` Matthieu Baerts
2026-09-29 10:00     ` Pablo Neira Ayuso
2026-09-29 10:40       ` Matthieu Baerts
2026-09-29 12:03         ` Pablo Neira Ayuso
2026-09-30 18:33           ` Matthieu Baerts
2026-10-06 23:29             ` Pablo Neira Ayuso
2026-10-07  8:01               ` Fernando Fernandez Mancera
2026-10-07 10:47                 ` Pablo Neira Ayuso
2026-10-07  8:17               ` Matthieu Baerts
2026-09-28 21:54 ` Jan Engelhardt
2026-09-29  9:34   ` Matthieu Baerts

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox