Netdev List
 help / color / mirror / Atom feed
From: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: ayaka <ayaka@soulik.info>,
	 Willem de Bruijn <willemdebruijn.kernel@gmail.com>
Cc: Jason Wang <jasowang@redhat.com>,
	 netdev@vger.kernel.org,  davem@davemloft.net,
	 edumazet@google.com,  kuba@kernel.org,  pabeni@redhat.com,
	 linux-kernel@vger.kernel.org
Subject: Re: [PATCH] net: tuntap: add ioctl() TUNGETQUEUEINDX to fetch queue index
Date: Fri, 09 Aug 2024 10:55:48 -0400	[thread overview]
Message-ID: <66b62df442a85_3bec1229461@willemb.c.googlers.com.notmuch> (raw)
In-Reply-To: <9C79659E-2CB1-4959-B35C-9D397DF6F399@soulik.info>

ayaka wrote:
> 
> Sent from my iPad

Try to avoid ^^^
 
> > On Aug 9, 2024, at 2:49 AM, Willem de Bruijn <willemdebruijn.kernel@gmail.com> wrote:
> > 
> > 
> >> 
> >>> So I guess an application that owns all the queues could keep track of
> >>> the queue-id to FD mapping. But it is not trivial, nor defined ABI
> >>> behavior.
> >>> 
> >>> Querying the queue_id as in the proposed patch might not solve the
> >>> challenge, though. Since an FD's queue-id may change simply because
> >> Yes, when I asked about those eBPF thing, I thought I don’t need the queue id in those ebpf. It turns out a misunderstanding.
> >> Do we all agree that no matter which filter or steering method we used here, we need a method to query queue index assigned with a fd?
> > 
> > That depends how you intend to use it. And in particular how to work
> > around the issue of IDs not being stable. Without solving that, it
> > seems like an impractical and even dangerous -because easy to misuse-
> > interface.
> > 
> First of all, I need to figure out when the steering action happens.
> When I use multiq qdisc with skbedit, does it happens after the net_device_ops->ndo_select_queue() ?
> If it did, that will still generate unused rxhash and txhash and flow tracking. It sounds a big overhead.
> Is it the same path for tc-bpf solution ?

TC egress is called early in __dev_queue_xmit, the main entry point for
transmission, in sch_handle_egress.

A few lines below netdev_core_pick_tx selects the txq by setting
skb->queue_mapping. Either through netdev_pick_tx or through a device
specific callback ndo_select_queue if it exists.

For tun, tun_select_queue implements that callback. If
TUNSETSTEERINGEBPF is configured, then the BPF program is called. Else
it uses its own rx_hash based approach in tun_automq_select_queue.

There is a special case in between. If TC egress ran skbedit, then
this sets current->net_xmit.skip_txqueue. Which will read the txq
from the skb->queue_mapping set by skbedit, and skip netdev_pick_tx.

That seems more roundabout than I had expected. I thought the code
would just check whether skb->queue_mapping is set and if so skip
netdev_pick_tx.

I wonder if this now means that setting queue_mapping with any other
TC action than skbedit now gets ignored. Importantly, cls_bpf or
act_bpf.

> I would reply with my concern about violating IDs in your last question.
> >>> another queue was detached. So this would have to be queried on each
> >>> detach.
> >>> 
> >> Thank you Jason. That is why I mentioned I may need to submit another patch to bind the queue index with a flow.
> >> 
> >> I think here is a good chance to discuss about this.
> >> I think from the design, the number of queue was a fixed number in those hardware devices? Also for those remote processor type wireless device(I think those are the modem devices).
> >> The way invoked with hash in every packet could consume lots of CPU times. And it is not necessary to track every packet.
> > 
> > rxhash based steering is common. There needs to be a strong(er) reason
> > to implement an alternative.
> > 
> I have a few questions about this hash steering, which didn’t request any future filter invoked:
> 1. If a flow happens before wrote to the tun, how to filter it?

What do you mean?

> 2. Does such a hash operation happen to every packet passing through?

For packets with a local socket, the computation is cached in the
socket.

For these tunnel packets, see tun_automq_select_queue. Specifically,
the call to __skb_get_hash_symmetric.

I'm actually not entirely sure why tun has this, rather than defer
to netdev_pick_tx, which call skb_tx_hash.

> 3. Is rxhash based on the flow tracking record in the tun driver?
> Those CPU overhead may demolish the benefit of the multiple queues and filters in the kernel solution.

Keyword is "may". Avoid premature optimization in favor of data.

> Also the flow tracking has a limited to 4096 or 1024, for a IPv4 /24 subnet, if everyone opened 16 websites, are we run out of memory before some entries expired?
> 
> I want to  seek there is a modern way to implement VPN in Linux after so many features has been introduced to Linux. So far, I don’t find a proper way to make any advantage here than other platforms.
> >> Could I add another property in struct tun_file and steering program return wanted value. Then it is application’s work to keep this new property unique.
> > 
> > I don't entirely follow this suggestion?
> > 
> >>> I suppose one underlying question is how important is the mapping of
> >>> flows to specific queue-id's? Is it a problem if the destination queue
> >>> for a flow changes mid-stream?
> >> Yes, it matters. Or why I want to use this feature. From all the open source VPN I know, neither enabled this multiqueu feature nor create more than one queue for it.
> >> And virtual machine would use the tap at the most time(they want to emulate a real nic).
> >> So basically this multiple queue feature was kind of useless for the VPN usage.
> >> If the filter can’t work atomically here, which would lead to unwanted packets transmitted to the wrong thread.
> > 
> > What exactly is the issue if a flow migrates from one queue to
> > another? There may be some OOO arrival. But these configuration
> > changes are rare events.
> I don’t know what the OOO means here.

Out of order.

> If a flow would migrate from its supposed queue to another, that was against the pretension to use the multiple queues here.
> A queue presents a VPN node here. It means it would leak one’s data to the other.
> Also those data could be just garbage fragments costs bandwidth sending to a peer that can’t handle it.

MultiQ is normally just a scalability optimization. It does not matter
for correctness, bar the possibility of brief packet reordering when a
flow switches queues.

I now get that what you are trying to do is set up a 1:1 relationship
between VPN connections and multi queue tun FDs. What would be reading
these FDs? If a single process, then it can definitely handle flow
migration.


  reply	other threads:[~2024-08-09 14:55 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2024-07-31 11:19 [PATCH] net: tuntap: add ioctl() TUNGETQUEUEINDX to fetch queue index Randy Li
2024-07-31 14:12 ` Willem de Bruijn
2024-07-31 16:45   ` Randy Li
2024-07-31 21:57     ` Willem de Bruijn
2024-08-01  9:15       ` Randy Li
2024-08-01 13:04         ` Willem de Bruijn
2024-08-01 13:36           ` Randy Li
2024-08-01 14:17             ` Willem de Bruijn
2024-08-01 19:52               ` Randy Li
2024-08-02 15:10                 ` Willem de Bruijn
2024-08-07 18:54                   ` Randy Li
2024-08-08  2:10                     ` Willem de Bruijn
2024-08-08  2:49                       ` Jason Wang
2024-08-08  3:11                         ` Willem de Bruijn
2024-08-08  3:36                           ` Jason Wang
2024-08-08  4:16                           ` ayaka
2024-08-08 18:48                             ` Willem de Bruijn
2024-08-09  4:45                               ` ayaka
2024-08-09 14:55                                 ` Willem de Bruijn [this message]
2024-08-12  6:05                                   ` Jason Wang
2024-08-12 17:10                                     ` Willem de Bruijn
2024-08-13  3:53                                       ` Jason Wang
2024-08-13 13:22                                         ` Willem de Bruijn
  -- strict thread matches above, loose matches on Subject: below --
2024-08-09  4:50 ayaka

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=66b62df442a85_3bec1229461@willemb.c.googlers.com.notmuch \
    --to=willemdebruijn.kernel@gmail.com \
    --cc=ayaka@soulik.info \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=jasowang@redhat.com \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox