From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.1 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_PASS autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 6689AC43381 for ; Thu, 21 Mar 2019 22:32:56 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 24D8C21917 for ; Thu, 21 Mar 2019 22:32:56 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="rTQ7X+ew" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1727151AbfCUWcy (ORCPT ); Thu, 21 Mar 2019 18:32:54 -0400 Received: from aserp2130.oracle.com ([141.146.126.79]:36952 "EHLO aserp2130.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1726681AbfCUWcu (ORCPT ); Thu, 21 Mar 2019 18:32:50 -0400 Received: from pps.filterd (aserp2130.oracle.com [127.0.0.1]) by aserp2130.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x2LMOHpK033861; Thu, 21 Mar 2019 22:32:35 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=subject : to : cc : references : from : message-id : date : mime-version : in-reply-to : content-type : content-transfer-encoding; s=corp-2018-07-02; bh=ILSreHlxgRLSjo1kz+apaIcQyRkIgpeMtBxtFw9JJ2w=; b=rTQ7X+eweQaSGZBs+fMlOTK+lskTt1I5TxHjav1nvEqCIZBFcKiti/49VIOH6cCw+fNy JvpNe4PujqZAgonpSnsEr/L+tWE5joXh+w91v+46BjYiJ/YmzkNVeMFDr9Ch1FKPcQk+ rJs79I3fFsnXVXRMKIWW++fg++AnoHUheKLSbuQaPeN4jaJNrFyY1q+yuicOeB4Pnrkl EKjhVT5z/HvAVwABck7nePqHMyy4hTzoKLLMeo7rXt58chq5eUa6s/n/19tlZN6gLC1I amlSuqOzmERi79jiK92EbTsrG8PxY9jKulwC8WI7hJXeU15XCNqVxy7Nmi+DJEflIiic IQ== Received: from aserv0022.oracle.com (aserv0022.oracle.com [141.146.126.234]) by aserp2130.oracle.com with ESMTP id 2r8pnf3afm-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 22:32:34 +0000 Received: from aserv0122.oracle.com (aserv0122.oracle.com [141.146.126.236]) by aserv0022.oracle.com (8.14.4/8.14.4) with ESMTP id x2LMWTc5021646 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 21 Mar 2019 22:32:29 GMT Received: from abhmp0014.oracle.com (abhmp0014.oracle.com [141.146.116.20]) by aserv0122.oracle.com (8.14.4/8.14.4) with ESMTP id x2LMWSnG009175; Thu, 21 Mar 2019 22:32:28 GMT Received: from [10.132.91.97] (/10.132.91.97) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Thu, 21 Mar 2019 15:32:28 -0700 Subject: Re: [summary] virtio network device failover writeup To: Stephen Hemminger , "Michael S. Tsirkin" Cc: Liran Alon , Sridhar Samudrala , Alexander Duyck , Jakub Kicinski , Jiri Pirko , David Miller , Netdev , virtualization@lists.linux-foundation.org, boris.ostrovsky@oracle.com, vijay.balakrishna@oracle.com, jfreimann@redhat.com, ogerlitz@mellanox.com, vuhuong@mellanox.com References: <20190320061632-mutt-send-email-mst@kernel.org> <20190320100747-mutt-send-email-mst@kernel.org> <36772E22-7A8F-4C42-A731-398E3204B418@oracle.com> <20190320180641-mutt-send-email-mst@kernel.org> <20190321044920-mutt-send-email-mst@kernel.org> <20190321082532-mutt-send-email-mst@kernel.org> <20190321085159-mutt-send-email-mst@kernel.org> <20190321084430.4750bd9b@shemminger-XPS-13-9360> From: si-wei liu Organization: Oracle Corporation Message-ID: Date: Thu, 21 Mar 2019 15:33:43 -0700 User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:60.0) Gecko/20100101 Thunderbird/60.4.0 MIME-Version: 1.0 In-Reply-To: <20190321084430.4750bd9b@shemminger-XPS-13-9360> Content-Type: text/plain; charset=utf-8; format=flowed Content-Transfer-Encoding: 8bit Content-Language: en-US X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9202 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1903210156 Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org On 3/21/2019 8:44 AM, Stephen Hemminger wrote: > On Thu, 21 Mar 2019 08:57:03 -0400 > "Michael S. Tsirkin" wrote: > >> On Thu, Mar 21, 2019 at 02:47:50PM +0200, Liran Alon wrote: >>> >>>> On 21 Mar 2019, at 14:37, Michael S. Tsirkin wrote: >>>> >>>> On Thu, Mar 21, 2019 at 12:07:57PM +0200, Liran Alon wrote: >>>>>>>>> 2) It brings non-intuitive customer experience. For example, a customer may attempt to analyse connectivity issue by checking the connectivity >>>>>>>>> on a net-failover slave (e.g. the VF) but will see no connectivity when in-fact checking the connectivity on the net-failover master netdev shows correct connectivity. >>>>>>>>> >>>>>>>>> The set of changes I vision to fix our issues are: >>>>>>>>> 1) Hide net-failover slaves in a different netns created and managed by the kernel. But that user can enter to it and manage the netdevs there if wishes to do so explicitly. >>>>>>>>> (E.g. Configure the net-failover VF slave in some special way). >>>>>>>>> 2) Match the virtio-net and the VF based on a PV attribute instead of MAC. (Similar to as done in NetVSC). E.g. Provide a virtio-net interface to get PCI slot where the matching VF will be hot-plugged by hypervisor. >>>>>>>>> 3) Have an explicit virtio-net control message to command hypervisor to switch data-path from virtio-net to VF and vice-versa. Instead of relying on intercepting the PCI master enable-bit >>>>>>>>> as an indicator on when VF is about to be set up. (Similar to as done in NetVSC). >>>>>>>>> >>>>>>>>> Is there any clear issue we see regarding the above suggestion? >>>>>>>>> >>>>>>>>> -Liran >>>>>>>> The issue would be this: how do we avoid conflicting with namespaces >>>>>>>> created by users? >>>>>>> This is kinda controversial, but maybe separate netns names into 2 groups: hidden and normal. >>>>>>> To reference a hidden netns, you need to do it explicitly. >>>>>>> Hidden and normal netns names can collide as they will be maintained in different namespaces (Yes I’m overloading the term namespace here…). >>>>>> Maybe it's an unnamed namespace. Hidden until userspace gives it a name? >>>>> This is also a good idea that will solve the issue. Yes. >>>>> >>>>>> >>>>>>> Does this seems reasonable? >>>>>>> >>>>>>> -Liran >>>>>> Reasonable I'd say yes, easy to implement probably no. But maybe I >>>>>> missed a trick or two. >>>>> BTW, from a practical point of view, I think that even until we figure out a solution on how to implement this, >>>>> it was better to create an kernel auto-generated name (e.g. “kernel_net_failover_slaves") >>>>> that will break only userspace workloads that by a very rare-chance have a netns that collides with this then >>>>> the breakage we have today for the various userspace components. >>>>> >>>>> -Liran >>>> It seems quite easy to supply that as a module parameter. Do we need two >>>> namespaces though? Won't some userspace still be confused by the two >>>> slaves sharing the MAC address? >>> That’s one reasonable option. >>> Another one is that we will indeed change the mechanism by which we determine a VF should be bonded with a virtio-net device. >>> i.e. Expose a new virtio-net property that specify the PCI slot of the VF to be bonded with. >>> >>> The second seems cleaner but I don’t have a strong opinion on this. Both seem reasonable to me and your suggestion is faster to implement from current state of things. >>> >>> -Liran >> OK. Now what happens if master is moved to another namespace? Do we need >> to move the slaves too? >> >> Also siwei's patch is then kind of extraneous right? >> Attempts to rename a slave will now fail as it's in a namespace... > I did try moving slave device into a namespace at one point. > The problem is that introduces all sorts of locking problems in the code > because you can't do it directly in the context of when the callback > happens that a new slave device is discovered. > > Since you can't safely change device namespace in the notifier, > it requires a work queue. Then you add more complexity and error cases > because the slave is exposed for a short period, and handling all the > state race unwinds... Thanks for your input, that's why I never got it started on the implementation before getting consensus here. I think we need to put the slave into kernel created netns otherwise it suffers from various userspace races. Userspace tool (such as udevd and ip) needs to specifically subscribe to events in those kernel created netns for rename and config. Locking still complicated, though there should be one way or another to work out. -Siwei > > Good idea but hard to implement