From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.1 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_PASS autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1412EC43381 for ; Wed, 20 Mar 2019 12:24:06 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id CE5EB2184D for ; Wed, 20 Mar 2019 12:24:05 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="RvZZw1dw" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1727606AbfCTMYF (ORCPT ); Wed, 20 Mar 2019 08:24:05 -0400 Received: from aserp2130.oracle.com ([141.146.126.79]:60590 "EHLO aserp2130.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1726366AbfCTMYC (ORCPT ); Wed, 20 Mar 2019 08:24:02 -0400 Received: from pps.filterd (aserp2130.oracle.com [127.0.0.1]) by aserp2130.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x2KCJQOc150269; Wed, 20 Mar 2019 12:23:45 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=content-type : mime-version : subject : from : in-reply-to : date : cc : content-transfer-encoding : message-id : references : to; s=corp-2018-07-02; bh=xNL0kLE5qCMZ1b2lF+fFzwD3EmKzSB7sWZh9PVRF3Q4=; b=RvZZw1dwcH3kzAU4xNlNh6T7d6X2bHeTx1zGqPXEd2FZCncBc8HID94MkccY5eGe+rI2 fHZnjLrCVdoPLkrRAmoqj8gzHc/A8KnhxMjcGHZPU+d+hFVR13RtCLcN1Uy0DDGGbSrR 6gNwwjvl7xshNeowQlA58sy6iSbzpX5b9E+L6AgBmSckt9VQLCW4qguQIM5dXiUt49H4 od36d9rSOGnTJPPzEMXnAbfOWFEdPGf+at2SAxHqpK4XIS6HkSg7YythU+qh7Zoxi/qm XaYj5dOUU5MnuNU5GlqK/RpzQU1Z1UyNc6O9+jllMVb3z0PJoSr+EXQ3ntYGf9x2hKOr yg== Received: from userv0022.oracle.com (userv0022.oracle.com [156.151.31.74]) by aserp2130.oracle.com with ESMTP id 2r8pnetktq-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 20 Mar 2019 12:23:45 +0000 Received: from userv0122.oracle.com (userv0122.oracle.com [156.151.31.75]) by userv0022.oracle.com (8.14.4/8.14.4) with ESMTP id x2KCNirx016682 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 20 Mar 2019 12:23:44 GMT Received: from abhmp0001.oracle.com (abhmp0001.oracle.com [141.146.116.7]) by userv0122.oracle.com (8.14.4/8.14.4) with ESMTP id x2KCNh24024928; Wed, 20 Mar 2019 12:23:43 GMT Received: from [10.0.5.57] (/213.57.127.10) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Wed, 20 Mar 2019 05:23:42 -0700 Content-Type: text/plain; charset=utf-8 Mime-Version: 1.0 (Mac OS X Mail 11.1 \(3445.4.7\)) Subject: Re: [summary] virtio network device failover writeup From: Liran Alon In-Reply-To: <20190320061632-mutt-send-email-mst@kernel.org> Date: Wed, 20 Mar 2019 14:23:36 +0200 Cc: Stephen Hemminger , Si-Wei Liu , Sridhar Samudrala , Alexander Duyck , Jakub Kicinski , Jiri Pirko , David Miller , Netdev , virtualization@lists.linux-foundation.org, boris.ostrovsky@oracle.com, vijay.balakrishna@oracle.com, jfreimann@redhat.com, ogerlitz@mellanox.com, vuhuong@mellanox.com Content-Transfer-Encoding: quoted-printable Message-Id: References: <20190317095052-mutt-send-email-mst@kernel.org> <54E7C3AF-C3C5-4AF2-86C9-AA50389F855F@oracle.com> <20190319084647.727f8dcf@shemminger-XPS-13-9360> <20190319171638-mutt-send-email-mst@kernel.org> <79F5D7C0-BBAA-4F78-9039-27A444970002@oracle.com> <20190320061632-mutt-send-email-mst@kernel.org> To: "Michael S. Tsirkin" X-Mailer: Apple Mail (2.3445.4.7) X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9200 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1903200097 Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org > On 20 Mar 2019, at 12:25, Michael S. Tsirkin wrote: >=20 > On Wed, Mar 20, 2019 at 01:25:58AM +0200, Liran Alon wrote: >>=20 >>=20 >>> On 19 Mar 2019, at 23:19, Michael S. Tsirkin wrote: >>>=20 >>> On Tue, Mar 19, 2019 at 08:46:47AM -0700, Stephen Hemminger wrote: >>>> On Tue, 19 Mar 2019 14:38:06 +0200 >>>> Liran Alon wrote: >>>>=20 >>>>> b.3) cloud-init: If configured to perform network-configuration, = it attempts to configure all available netdevs. It should avoid however = doing so on net-failover slaves. >>>>> (Microsoft has handled this by adding a mechanism in cloud-init to = blacklist a netdev from being configured in case it is owned by a = specific PCI driver. Specifically, they blacklist Mellanox VF driver. = However, this technique doesn=E2=80=99t work for the net-failover = mechanism because both the net-failover netdev and the virtio-net netdev = are owned by the virtio-net PCI driver). >>>>=20 >>>> Cloud-init should really just ignore all devices that have a master = device. >>>> That would have been more general, and safer for other use cases. >>>=20 >>> Given lots of userspace doesn't do this, I wonder whether it would = be >>> safer to just somehow pretend to userspace that the slave links are >>> down? And add a special attribute for the actual link state. >>=20 >> I think this may be problematic as it would also break legit use case >> of userspace attempt to set various config on VF slave. >> In general, lying to userspace usually leads to problems. >=20 > I hear you on this. So how about instead of lying, > we basically just fail some accesses to slaves > unless a flag is set e.g. in ethtool. >=20 > Some userspace will need to change to set it but in a minor way. > Arguably/hopefully failure to set config would generally be a safer > failure. Once userspace will set this new flag by ethtool, all operations done by = other userspace components will still work. E.g. Running dhclient without parameters, after this flag was set, will = still attempt to perform DHCP on it and will now succeed. Therefore, this proposal just effectively delays when the net-failover = slave can be operated on by userspace. But what we actually want is to never allow a net-failover slave to be = operated by userspace unless it is explicitly stated by userspace that it wishes to perform a set of actions on the = net-failover slave. Something that was achieved if, for example, the net-failover slaves = were in a different netns than default netns. This also aligns with expected customer experience that most customers = just want to see a 1:1 mapping between a vNIC and a visible netdev. But of course maybe there are other ideas that can achieve similar = behaviour. -Liran >=20 > Which things to fail? Probably sending/receiving packets? Getting = MAC? > More? >=20 >> If we reach >> to a scenario where we try to avoid userspace issues generically and >> not on a userspace component basis, I believe the right path should = be >> to hide the net-failover slaves such that explicit action is required >> to actually manipulate them (As described in blog-post). E.g. >> Automatically move net-failover slaves by kernel to a different = netns. >>=20 >> -Liran >>=20 >>>=20 >>> --=20 >>> MST