From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.1 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_PASS autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id AF857C43381 for ; Thu, 21 Mar 2019 10:08:21 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 6AE0C218B0 for ; Thu, 21 Mar 2019 10:08:21 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="QDyeuEgn" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1728338AbfCUKIU (ORCPT ); Thu, 21 Mar 2019 06:08:20 -0400 Received: from userp2130.oracle.com ([156.151.31.86]:40786 "EHLO userp2130.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1728253AbfCUKIT (ORCPT ); Thu, 21 Mar 2019 06:08:19 -0400 Received: from pps.filterd (userp2130.oracle.com [127.0.0.1]) by userp2130.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x2LA4MJD161539; Thu, 21 Mar 2019 10:08:06 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=content-type : mime-version : subject : from : in-reply-to : date : cc : content-transfer-encoding : message-id : references : to; s=corp-2018-07-02; bh=FX17kfrjA8r1SY9qoA0Wus/PQwLERxM7377Mo5bCNtQ=; b=QDyeuEgnAnuHIaYD7sn20oOuVywivm2EhoSV0MjrDiV1yk+dc4424uqUHkpArbfDG3QP 7MFQQOl7ruNwMzyLmw4UgjjUlPOUHLqiYakQBSBOVcDBK+BpdhRsk6MRR4KfkG6xKbXP +CEdeOX+J9WO9BbAl7n82yDbJaPZluHxaRQAAUPlzcvwAzM3NNibKQ1IFwME8FGBppCM xN5G6hxHvbb4SalzAfs3a93MN5zNvIl/r7KiQ0gy1KRD9y34jGVLLKKrporuvtYL1+Ai X4xmcIvNkf9/l8DrrBXy7jSxvg8I9c6xqAU4IL5MU+nU4KIAz73sSLHl7qz8J7xa7KB5 Nw== Received: from userv0021.oracle.com (userv0021.oracle.com [156.151.31.71]) by userp2130.oracle.com with ESMTP id 2r8rjuykfy-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 10:08:06 +0000 Received: from userv0121.oracle.com (userv0121.oracle.com [156.151.31.72]) by userv0021.oracle.com (8.14.4/8.14.4) with ESMTP id x2LA85wh016530 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 10:08:05 GMT Received: from abhmp0001.oracle.com (abhmp0001.oracle.com [141.146.116.7]) by userv0121.oracle.com (8.14.4/8.13.8) with ESMTP id x2LA832M005122; Thu, 21 Mar 2019 10:08:03 GMT Received: from [192.168.14.112] (/79.179.229.177) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Thu, 21 Mar 2019 03:08:03 -0700 Content-Type: text/plain; charset=utf-8 Mime-Version: 1.0 (Mac OS X Mail 11.1 \(3445.4.7\)) Subject: Re: [summary] virtio network device failover writeup From: Liran Alon In-Reply-To: <20190321044920-mutt-send-email-mst@kernel.org> Date: Thu, 21 Mar 2019 12:07:57 +0200 Cc: Stephen Hemminger , Si-Wei Liu , Sridhar Samudrala , Alexander Duyck , Jakub Kicinski , Jiri Pirko , David Miller , Netdev , virtualization@lists.linux-foundation.org, boris.ostrovsky@oracle.com, vijay.balakrishna@oracle.com, jfreimann@redhat.com, ogerlitz@mellanox.com, vuhuong@mellanox.com Content-Transfer-Encoding: quoted-printable Message-Id: References: <54E7C3AF-C3C5-4AF2-86C9-AA50389F855F@oracle.com> <20190319084647.727f8dcf@shemminger-XPS-13-9360> <20190319171638-mutt-send-email-mst@kernel.org> <79F5D7C0-BBAA-4F78-9039-27A444970002@oracle.com> <20190320061632-mutt-send-email-mst@kernel.org> <20190320100747-mutt-send-email-mst@kernel.org> <36772E22-7A8F-4C42-A731-398E3204B418@oracle.com> <20190320180641-mutt-send-email-mst@kernel.org> <20190321044920-mutt-send-email-mst@kernel.org> To: "Michael S. Tsirkin" X-Mailer: Apple Mail (2.3445.4.7) X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9201 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1903210074 Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org > On 21 Mar 2019, at 10:58, Michael S. Tsirkin wrote: >=20 > On Thu, Mar 21, 2019 at 12:19:22AM +0200, Liran Alon wrote: >>=20 >>=20 >>> On 21 Mar 2019, at 0:10, Michael S. Tsirkin wrote: >>>=20 >>> On Wed, Mar 20, 2019 at 11:43:41PM +0200, Liran Alon wrote: >>>>=20 >>>>=20 >>>>> On 20 Mar 2019, at 16:09, Michael S. Tsirkin = wrote: >>>>>=20 >>>>> On Wed, Mar 20, 2019 at 02:23:36PM +0200, Liran Alon wrote: >>>>>>=20 >>>>>>=20 >>>>>>> On 20 Mar 2019, at 12:25, Michael S. Tsirkin = wrote: >>>>>>>=20 >>>>>>> On Wed, Mar 20, 2019 at 01:25:58AM +0200, Liran Alon wrote: >>>>>>>>=20 >>>>>>>>=20 >>>>>>>>> On 19 Mar 2019, at 23:19, Michael S. Tsirkin = wrote: >>>>>>>>>=20 >>>>>>>>> On Tue, Mar 19, 2019 at 08:46:47AM -0700, Stephen Hemminger = wrote: >>>>>>>>>> On Tue, 19 Mar 2019 14:38:06 +0200 >>>>>>>>>> Liran Alon wrote: >>>>>>>>>>=20 >>>>>>>>>>> b.3) cloud-init: If configured to perform = network-configuration, it attempts to configure all available netdevs. = It should avoid however doing so on net-failover slaves. >>>>>>>>>>> (Microsoft has handled this by adding a mechanism in = cloud-init to blacklist a netdev from being configured in case it is = owned by a specific PCI driver. Specifically, they blacklist Mellanox VF = driver. However, this technique doesn=E2=80=99t work for the = net-failover mechanism because both the net-failover netdev and the = virtio-net netdev are owned by the virtio-net PCI driver). >>>>>>>>>>=20 >>>>>>>>>> Cloud-init should really just ignore all devices that have a = master device. >>>>>>>>>> That would have been more general, and safer for other use = cases. >>>>>>>>>=20 >>>>>>>>> Given lots of userspace doesn't do this, I wonder whether it = would be >>>>>>>>> safer to just somehow pretend to userspace that the slave = links are >>>>>>>>> down? And add a special attribute for the actual link state. >>>>>>>>=20 >>>>>>>> I think this may be problematic as it would also break legit = use case >>>>>>>> of userspace attempt to set various config on VF slave. >>>>>>>> In general, lying to userspace usually leads to problems. >>>>>>>=20 >>>>>>> I hear you on this. So how about instead of lying, >>>>>>> we basically just fail some accesses to slaves >>>>>>> unless a flag is set e.g. in ethtool. >>>>>>>=20 >>>>>>> Some userspace will need to change to set it but in a minor way. >>>>>>> Arguably/hopefully failure to set config would generally be a = safer >>>>>>> failure. >>>>>>=20 >>>>>> Once userspace will set this new flag by ethtool, all operations = done by other userspace components will still work. >>>>>=20 >>>>> Sorry about being unclear, the idea would be to require the flag = on each ethtool operation. >>>>=20 >>>> Oh. I have indeed misunderstood your previous email then. :) >>>> Thanks for clarifying. >>>>=20 >>>>>=20 >>>>>> E.g. Running dhclient without parameters, after this flag was = set, will still attempt to perform DHCP on it and will now succeed. >>>>>=20 >>>>> I think sending/receiving should probably just fail = unconditionally. >>>>=20 >>>> You mean that you wish that somehow kernel will prevent Tx on = net-failover slave netdev >>>> unless skb is marked with some flag to indicate it has been sent = via the net-failover master? >>>=20 >>> We can maybe avoid binding a protocol socket to the device? >>=20 >> That is indeed another possibility that would work to avoid the DHCP = issues. >> And will still allow checking connectivity. So it is better. >> However, I still think it provides an non-intuitive customer = experience. >> In addition, I also want to take into account that most customers are = expected a 1:1 mapping between a vNIC and a netdev. >> i.e. A cloud instance should show 1-netdev if it has one vNIC = attached to it defined. >> Customers usually don=E2=80=99t care how they get accelerated = networking. They just care they do. >>=20 >>>=20 >>>> This indeed resolves the group of userspace issues around = performing DHCP on net-failover slaves directly (By dracut/initramfs, = dhclient and etc.). >>>>=20 >>>> However, I see a couple of down-sides to it: >>>> 1) It doesn=E2=80=99t resolve all userspace issues listed in this = email thread. For example, cloud-init will still attempt to perform = network config on net-failover slaves. >>>> It also doesn=E2=80=99t help with regard to Ubuntu=E2=80=99s = netplan issue that creates udev rules that match only by MAC. >>>=20 >>>=20 >>> How about we fail to retrieve mac from the slave? >>=20 >> That would work but I think it is cleaner to just not bind PV and VF = based on having the same MAC. >=20 > There's a reference to that under "Non-MAC based pairing". >=20 > I'll look into making it more explicit. Yes I know. I was referring to what you described in that section. >=20 >>>=20 >>>> 2) It brings non-intuitive customer experience. For example, a = customer may attempt to analyse connectivity issue by checking the = connectivity >>>> on a net-failover slave (e.g. the VF) but will see no connectivity = when in-fact checking the connectivity on the net-failover master netdev = shows correct connectivity. >>>>=20 >>>> The set of changes I vision to fix our issues are: >>>> 1) Hide net-failover slaves in a different netns created and = managed by the kernel. But that user can enter to it and manage the = netdevs there if wishes to do so explicitly. >>>> (E.g. Configure the net-failover VF slave in some special way). >>>> 2) Match the virtio-net and the VF based on a PV attribute instead = of MAC. (Similar to as done in NetVSC). E.g. Provide a virtio-net = interface to get PCI slot where the matching VF will be hot-plugged by = hypervisor. >>>> 3) Have an explicit virtio-net control message to command = hypervisor to switch data-path from virtio-net to VF and vice-versa. = Instead of relying on intercepting the PCI master enable-bit >>>> as an indicator on when VF is about to be set up. (Similar to as = done in NetVSC). >>>>=20 >>>> Is there any clear issue we see regarding the above suggestion? >>>>=20 >>>> -Liran >>>=20 >>> The issue would be this: how do we avoid conflicting with namespaces >>> created by users? >>=20 >> This is kinda controversial, but maybe separate netns names into 2 = groups: hidden and normal. >> To reference a hidden netns, you need to do it explicitly.=20 >> Hidden and normal netns names can collide as they will be maintained = in different namespaces (Yes I=E2=80=99m overloading the term namespace = here=E2=80=A6). >=20 > Maybe it's an unnamed namespace. Hidden until userspace gives it a = name? This is also a good idea that will solve the issue. Yes. >=20 >> Does this seems reasonable? >>=20 >> -Liran >=20 > Reasonable I'd say yes, easy to implement probably no. But maybe I > missed a trick or two. BTW, from a practical point of view, I think that even until we figure = out a solution on how to implement this, it was better to create an kernel auto-generated name (e.g. = =E2=80=9Ckernel_net_failover_slaves") that will break only userspace workloads that by a very rare-chance have = a netns that collides with this then the breakage we have today for the various userspace components. -Liran >=20 >>>=20 >>>>>=20 >>>>>> Therefore, this proposal just effectively delays when the = net-failover slave can be operated on by userspace. >>>>>> But what we actually want is to never allow a net-failover slave = to be operated by userspace unless it is explicitly stated >>>>>> by userspace that it wishes to perform a set of actions on the = net-failover slave. >>>>>>=20 >>>>>> Something that was achieved if, for example, the net-failover = slaves were in a different netns than default netns. >>>>>> This also aligns with expected customer experience that most = customers just want to see a 1:1 mapping between a vNIC and a visible = netdev. >>>>>> But of course maybe there are other ideas that can achieve = similar behaviour. >>>>>>=20 >>>>>> -Liran >>>>>>=20 >>>>>>>=20 >>>>>>> Which things to fail? Probably sending/receiving packets? = Getting MAC? >>>>>>> More? >>>>>>>=20 >>>>>>>> If we reach >>>>>>>> to a scenario where we try to avoid userspace issues = generically and >>>>>>>> not on a userspace component basis, I believe the right path = should be >>>>>>>> to hide the net-failover slaves such that explicit action is = required >>>>>>>> to actually manipulate them (As described in blog-post). E.g. >>>>>>>> Automatically move net-failover slaves by kernel to a different = netns. >>>>>>>>=20 >>>>>>>> -Liran >>>>>>>>=20 >>>>>>>>>=20 >>>>>>>>> --=20 >>>>>>>>> MST