From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.1 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_PASS autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id E2EEEC43381 for ; Thu, 21 Mar 2019 14:16:40 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 9AAF2218D3 for ; Thu, 21 Mar 2019 14:16:40 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="Ey58nCqc" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1728286AbfCUOQj (ORCPT ); Thu, 21 Mar 2019 10:16:39 -0400 Received: from aserp2130.oracle.com ([141.146.126.79]:50020 "EHLO aserp2130.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1726551AbfCUOQj (ORCPT ); Thu, 21 Mar 2019 10:16:39 -0400 Received: from pps.filterd (aserp2130.oracle.com [127.0.0.1]) by aserp2130.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x2LEECqZ007540; Thu, 21 Mar 2019 14:16:24 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=content-type : mime-version : subject : from : in-reply-to : date : cc : content-transfer-encoding : message-id : references : to; s=corp-2018-07-02; bh=1BIOBq2T496aUOGDT1y+b/X0mwhwBg6HQWBAuoLtjSw=; b=Ey58nCqcS7kQ0EcxwevoZfrY/oFHP/sDSz8Do0stpo9mUWyc0/Se0h4LkT4PuzGvhxwh DBsDu8vDy8c1BItAMXsE4ICyRdgjskosMIuFSwSt2OC6UqjfhR0mQvzc3B4hz1+F/+Z/ jotF5TCGZSjqn2IFPIR4y/PDz+Ea8gXVDAHq80S+dl4vMAcrnPPgsjZkeUSdeEcAo7dL VMWZ8yq6q8cidOREaTJwXTPSXrqTxYSXaxK2eJLAe0jpWi1yBYY4s1w9nw0LjTZGQMEl ijlqq9YKy/wQlb1wyuV5dJxpX4wZTR6bVbEXe+DCY44v/2wI8DzuDdUqQzLav2D4AiBD Ig== Received: from userv0022.oracle.com (userv0022.oracle.com [156.151.31.74]) by aserp2130.oracle.com with ESMTP id 2r8pnf0wax-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 14:16:24 +0000 Received: from aserv0122.oracle.com (aserv0122.oracle.com [141.146.126.236]) by userv0022.oracle.com (8.14.4/8.14.4) with ESMTP id x2LEGMBx012813 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 14:16:22 GMT Received: from abhmp0018.oracle.com (abhmp0018.oracle.com [141.146.116.24]) by aserv0122.oracle.com (8.14.4/8.14.4) with ESMTP id x2LEGLp0007352; Thu, 21 Mar 2019 14:16:21 GMT Received: from [10.30.3.18] (/213.57.127.2) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Thu, 21 Mar 2019 07:16:21 -0700 Content-Type: text/plain; charset=utf-8 Mime-Version: 1.0 (Mac OS X Mail 11.1 \(3445.4.7\)) Subject: Re: [summary] virtio network device failover writeup From: Liran Alon In-Reply-To: <20190321094217-mutt-send-email-mst@kernel.org> Date: Thu, 21 Mar 2019 16:16:14 +0200 Cc: Stephen Hemminger , Si-Wei Liu , Sridhar Samudrala , Alexander Duyck , Jakub Kicinski , Jiri Pirko , David Miller , Netdev , virtualization@lists.linux-foundation.org, boris.ostrovsky@oracle.com, vijay.balakrishna@oracle.com, jfreimann@redhat.com, ogerlitz@mellanox.com, vuhuong@mellanox.com Content-Transfer-Encoding: quoted-printable Message-Id: References: <20190320180641-mutt-send-email-mst@kernel.org> <20190321044920-mutt-send-email-mst@kernel.org> <20190321082532-mutt-send-email-mst@kernel.org> <20190321085159-mutt-send-email-mst@kernel.org> <2939FB15-720A-4C9E-92B7-2DBA139DDE0F@oracle.com> <20190321090619-mutt-send-email-mst@kernel.org> <1B52153B-B968-4E5B-8959-E7E83CE7FEAF@oracle.com> <20190321094217-mutt-send-email-mst@kernel.org> To: "Michael S. Tsirkin" X-Mailer: Apple Mail (2.3445.4.7) X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9201 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1903210101 Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org > On 21 Mar 2019, at 15:51, Michael S. Tsirkin wrote: >=20 > On Thu, Mar 21, 2019 at 03:24:39PM +0200, Liran Alon wrote: >>=20 >>=20 >>> On 21 Mar 2019, at 15:12, Michael S. Tsirkin wrote: >>>=20 >>> On Thu, Mar 21, 2019 at 03:04:37PM +0200, Liran Alon wrote: >>>>=20 >>>>=20 >>>>> On 21 Mar 2019, at 14:57, Michael S. Tsirkin = wrote: >>>>>=20 >>>>> On Thu, Mar 21, 2019 at 02:47:50PM +0200, Liran Alon wrote: >>>>>>=20 >>>>>>=20 >>>>>>> On 21 Mar 2019, at 14:37, Michael S. Tsirkin = wrote: >>>>>>>=20 >>>>>>> On Thu, Mar 21, 2019 at 12:07:57PM +0200, Liran Alon wrote: >>>>>>>>>>>> 2) It brings non-intuitive customer experience. For = example, a customer may attempt to analyse connectivity issue by = checking the connectivity >>>>>>>>>>>> on a net-failover slave (e.g. the VF) but will see no = connectivity when in-fact checking the connectivity on the net-failover = master netdev shows correct connectivity. >>>>>>>>>>>>=20 >>>>>>>>>>>> The set of changes I vision to fix our issues are: >>>>>>>>>>>> 1) Hide net-failover slaves in a different netns created = and managed by the kernel. But that user can enter to it and manage the = netdevs there if wishes to do so explicitly. >>>>>>>>>>>> (E.g. Configure the net-failover VF slave in some special = way). >>>>>>>>>>>> 2) Match the virtio-net and the VF based on a PV attribute = instead of MAC. (Similar to as done in NetVSC). E.g. Provide a = virtio-net interface to get PCI slot where the matching VF will be = hot-plugged by hypervisor. >>>>>>>>>>>> 3) Have an explicit virtio-net control message to command = hypervisor to switch data-path from virtio-net to VF and vice-versa. = Instead of relying on intercepting the PCI master enable-bit >>>>>>>>>>>> as an indicator on when VF is about to be set up. (Similar = to as done in NetVSC). >>>>>>>>>>>>=20 >>>>>>>>>>>> Is there any clear issue we see regarding the above = suggestion? >>>>>>>>>>>>=20 >>>>>>>>>>>> -Liran >>>>>>>>>>>=20 >>>>>>>>>>> The issue would be this: how do we avoid conflicting with = namespaces >>>>>>>>>>> created by users? >>>>>>>>>>=20 >>>>>>>>>> This is kinda controversial, but maybe separate netns names = into 2 groups: hidden and normal. >>>>>>>>>> To reference a hidden netns, you need to do it explicitly.=20 >>>>>>>>>> Hidden and normal netns names can collide as they will be = maintained in different namespaces (Yes I=E2=80=99m overloading the term = namespace here=E2=80=A6). >>>>>>>>>=20 >>>>>>>>> Maybe it's an unnamed namespace. Hidden until userspace gives = it a name? >>>>>>>>=20 >>>>>>>> This is also a good idea that will solve the issue. Yes. >>>>>>>>=20 >>>>>>>>>=20 >>>>>>>>>> Does this seems reasonable? >>>>>>>>>>=20 >>>>>>>>>> -Liran >>>>>>>>>=20 >>>>>>>>> Reasonable I'd say yes, easy to implement probably no. But = maybe I >>>>>>>>> missed a trick or two. >>>>>>>>=20 >>>>>>>> BTW, from a practical point of view, I think that even until we = figure out a solution on how to implement this, >>>>>>>> it was better to create an kernel auto-generated name (e.g. = =E2=80=9Ckernel_net_failover_slaves") >>>>>>>> that will break only userspace workloads that by a very = rare-chance have a netns that collides with this then >>>>>>>> the breakage we have today for the various userspace = components. >>>>>>>>=20 >>>>>>>> -Liran >>>>>>>=20 >>>>>>> It seems quite easy to supply that as a module parameter. Do we = need two >>>>>>> namespaces though? Won't some userspace still be confused by the = two >>>>>>> slaves sharing the MAC address? >>>>>>=20 >>>>>> That=E2=80=99s one reasonable option. >>>>>> Another one is that we will indeed change the mechanism by which = we determine a VF should be bonded with a virtio-net device. >>>>>> i.e. Expose a new virtio-net property that specify the PCI slot = of the VF to be bonded with. >>>>>>=20 >>>>>> The second seems cleaner but I don=E2=80=99t have a strong = opinion on this. Both seem reasonable to me and your suggestion is = faster to implement from current state of things. >>>>>>=20 >>>>>> -Liran >>>>>=20 >>>>> OK. Now what happens if master is moved to another namespace? Do = we need >>>>> to move the slaves too? >>>>=20 >>>> No. Why would we move the slaves? >>>=20 >>>=20 >>> The reason we have 3 device model at all is so users can fine tune = the >>> slaves. >>=20 >> I Agree. >>=20 >>> I don't see why this applies to the root namespace but not >>> a container. If it has access to failover it should have access >>> to slaves. >>=20 >> Oh now I see your point. I haven=E2=80=99t thought about the = containers usage. >> My thinking was that customer can always just enter to the = =E2=80=9Chidden=E2=80=9D netns and configure there whatever he wants. >>=20 >> Do you have a suggestion how to handle this? >>=20 >> One option can be that every "visible" netns on system will have a = =E2=80=9Chidden=E2=80=9D unnamed netns where the net-failover slaves = reside in. >> If customer wishes to be able to enter to that netns and manage the = net-failover slaves explicitly, it will need to have an updated iproute2 >> that knows how to enter to that hidden netns. For most customers, = they won=E2=80=99t need to ever enter that netns and thus it is ok they = don=E2=80=99t >> have this updated iproute2. >=20 > Right so slaves need to be moved whenever master is moved. >=20 > Given the amount of mess involved, should we just teach > userspace to create the hidden netns and move slaves there? That=E2=80=99s a good question. However, I believe that it is easier and more suitable to happen in = kernel. This is because: 1) Implementation is generic across all various distros. 2) We seem to discover more and more issues with userspace as we keep = testing this on various distros, configurations and workloads. 3) It seems weird that kernel does some things automagically and some = things don=E2=80=99t. i.e. Kernel automatically binds the virtio-net and = VF to net-failover master and automatically opens the net-failover slave when the net-failover = master is opened, but it doesn=E2=80=99t care about the consequences = these actions have on userspace. Therefore, I propose let=E2=80=99s go =E2=80=9Call in=E2=80=9D: Kernel = should also be responsible for hiding it=E2=80=99s artefacts unless = customer userspace explicitly wants to view and manipulate them. >=20 >>>=20 >>>> The whole point is to make most customer ignore the net-failover = slaves and remain them =E2=80=9Chidden=E2=80=9D in their dedicated = netns. >>>=20 >>> So that makes the common case easy. That is good. My worry is it = might >>> make some uncommon cases impossible. >>>=20 >>>> We won=E2=80=99t prevent customer from explicitly moving the = net-failover slaves out of this netns, but we will not move them out of = there automatically. >>>>=20 >>>>>=20 >>>>> Also siwei's patch is then kind of extraneous right? >>>>> Attempts to rename a slave will now fail as it's in a namespace=E2=80= =A6 >>>>=20 >>>> I=E2=80=99m not sure actually. Isn't udev/systemd netns-aware? >>>> I would expect it to be able to provide names also to netdevs in = netns different than default netns. >>>=20 >>> I think most people move devices after they are renamed. >>=20 >> So? >> Si-Wei patch handles the issue that resolves from the fact the = net-failover master will be opened before the rename on the net-failover = slaves occur. >> This should happen (to my understanding) regardless of network = namespaces. >>=20 >> -Liran >=20 > My point was that any tool that moves devices after they > are renamed will be broken by kernel automatically putting > them in a namespace. I=E2=80=99m not sure I follow. How is this related to Si-Wei patch? Si-Wei patch (and the root-cause that leads to the issue it fixes) have = nothing to do with network namespaces. What do you mean tool that moves devices after they are renamed will be = broken by kernel? Care to give an example to clarify? -Liran >=20 >>>=20 >>>> If that=E2=80=99s the case, Si-Wei patch to be able to rename a = net-failover slave when it is already open is still required. As the = race-condition still exists. >>>>=20 >>>> -Liran >>>>=20 >>>>>=20 >>>>>>>=20 >>>>>>> --=20 >>>>>>> MST