From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-4.1 required=3.0 tests=DKIM_SIGNED,DKIM_VALID, DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI,SIGNED_OFF_BY, SPF_PASS autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 58545C43381 for ; Wed, 27 Mar 2019 20:10:37 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 1C9CB2087C for ; Wed, 27 Mar 2019 20:10:37 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="AZfrLUme" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1729033AbfC0UKg (ORCPT ); Wed, 27 Mar 2019 16:10:36 -0400 Received: from aserp2130.oracle.com ([141.146.126.79]:45030 "EHLO aserp2130.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1727832AbfC0UKf (ORCPT ); Wed, 27 Mar 2019 16:10:35 -0400 Received: from pps.filterd (aserp2130.oracle.com [127.0.0.1]) by aserp2130.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x2RK8fiL184019; Wed, 27 Mar 2019 20:10:19 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=subject : to : references : cc : from : message-id : date : mime-version : in-reply-to : content-type : content-transfer-encoding; s=corp-2018-07-02; bh=km3eOXHCRu1zndC5sF1hT4dTvEMCY5MnU6hlvIZD82c=; b=AZfrLUmeNADJnn+wnP4aA4FGVlKk7BIezHIFr/ngrWbaprkuz3ndGv1NudzSYz3wzeg4 Aonc/vCpup3wKgTYPVSer3jj5jBk8m5n0Lov/JpcqURCkVC1hHq/Of4vZxpYoQcgIEfW Ev7HlduOkF15SkaolIm0MVljTVlQlBaspyCU3WeeX1wZGIw2FkW339GKR0w7EqeHtuSS FfjFtCHuMTe5alCu1gbFk0qNMEnMHZUdRN8zsf8iG9DPT4SG4FhZdoBK38i6MuHzoEqA NdwlePC2e3MWCPyCyQrHfwD7jwxJG2qe87F46MVT3pwz1aAFWDUZm2FDnCX6kGqXbeYc yw== Received: from aserv0022.oracle.com (aserv0022.oracle.com [141.146.126.234]) by aserp2130.oracle.com with ESMTP id 2re6g12uww-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 27 Mar 2019 20:10:19 +0000 Received: from aserv0122.oracle.com (aserv0122.oracle.com [141.146.126.236]) by aserv0022.oracle.com (8.14.4/8.14.4) with ESMTP id x2RKAIiN017049 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 27 Mar 2019 20:10:18 GMT Received: from abhmp0004.oracle.com (abhmp0004.oracle.com [141.146.116.10]) by aserv0122.oracle.com (8.14.4/8.14.4) with ESMTP id x2RKAGS3016741; Wed, 27 Mar 2019 20:10:16 GMT Received: from [10.159.236.24] (/10.159.236.24) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Wed, 27 Mar 2019 13:10:16 -0700 Subject: Re: [PATCH net v3] failover: allow name change on IFF_UP slave interfaces To: "Michael S. Tsirkin" , Stephen Hemminger References: <1553644093-10917-1-git-send-email-si-wei.liu@oracle.com> <20190326191342.11f0cb55@shemminger-XPS-13-9360> <20190327092424-mutt-send-email-mst@kernel.org> Cc: sridhar.samudrala@intel.com, davem@davemloft.net, kubakici@wp.pl, alexander.duyck@gmail.com, jiri@resnulli.us, netdev@vger.kernel.org, virtualization@lists.linux-foundation.org, liran.alon@oracle.com, boris.ostrovsky@oracle.com, vijay.balakrishna@oracle.com From: si-wei liu Organization: Oracle Corporation Message-ID: Date: Wed, 27 Mar 2019 13:10:10 -0700 User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:45.0) Gecko/20100101 Thunderbird/45.8.0 MIME-Version: 1.0 In-Reply-To: <20190327092424-mutt-send-email-mst@kernel.org> Content-Type: text/plain; charset=windows-1252; format=flowed Content-Transfer-Encoding: 7bit X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9208 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1903270140 Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org On 3/27/2019 6:25 AM, Michael S. Tsirkin wrote: > On Tue, Mar 26, 2019 at 07:13:42PM -0700, Stephen Hemminger wrote: >> On Tue, 26 Mar 2019 19:48:13 -0400 >> Si-Wei Liu wrote: >> >>> When a netdev appears through hot plug then gets enslaved by a failover >>> master that is already up and running, the slave will be opened >>> right away after getting enslaved. Today there's a race that userspace >>> (udev) may fail to rename the slave if the kernel (net_failover) >>> opens the slave earlier than when the userspace rename happens. >>> Unlike bond or team, the primary slave of failover can't be renamed by >>> userspace ahead of time, since the kernel initiated auto-enslavement is >>> unable to, or rather, is never meant to be synchronized with the rename >>> request from userspace. >>> >>> As the failover slave interfaces are not designed to be operated >>> directly by userspace apps: IP configuration, filter rules with >>> regard to network traffic passing and etc., should all be done on master >>> interface. In general, userspace apps only care about the >>> name of master interface, while slave names are less important as long >>> as admin users can see reliable names that may carry >>> other information describing the netdev. For e.g., they can infer that >>> "ens3nsby" is a standby slave of "ens3", while for a >>> name like "eth0" they can't tell which master it belongs to. >>> >>> Historically the name of IFF_UP interface can't be changed because >>> there might be admin script or management software that is already >>> relying on such behavior and assumes that the slave name can't be >>> changed once UP. But failover is special: with the in-kernel >>> auto-enslavement mechanism, the userspace expectation for device >>> enumeration and bring-up order is already broken. Previously initramfs >>> and various userspace config tools were modified to bypass failover >>> slaves because of auto-enslavement and duplicate MAC address. Similarly, >>> in case that users care about seeing reliable slave name, the new type >>> of failover slaves needs to be taken care of specifically in userspace >>> anyway. >>> >>> It's less risky to lift up the rename restriction on failover slave >>> which is already UP. Although it's possible this change may potentially >>> break userspace component (most likely configuration scripts or >>> management software) that assumes slave name can't be changed while >>> UP, it's relatively a limited and controllable set among all userspace >>> components, which can be fixed specifically to listen for the rename >>> and/or link down/up events on failover slaves. Userspace component >>> interacting with slaves is expected to be changed to operate on failover >>> master interface instead, as the failover slave is dynamic in nature >>> which may come and go at any point. The goal is to make the role of >>> failover slaves less relevant, and userspace components should only >>> deal with failover master in the long run. >>> >>> Fixes: 30c8bd5aa8b2 ("net: Introduce generic failover module") >>> Signed-off-by: Si-Wei Liu >>> Reviewed-by: Liran Alon >> >> Why do you need to do dev_close/dev_open which will bounce >> the link? > What we need is notify userspace that link went up/down. > close/open will do that but just sending notifications > would do that as well without playing with link states. > Since you were requesting to send fake link down/up events around rename, so as to keep existing userspace intact with this behavioral change, right? The thing is if you can't fake notification with just IFF_UP or ~IFF_UP then claim everything is done. If you look at rtnl_fill_ifinfo() where the notification payload is prepared, you'll find a lot of states and flags are correlated: ifi_flags IFLA_OPERSTATE IFLA_CARRIER IFLA_CARRIER_CHANGES which requires below states to be toggled or taken care of in between: operstate __LINK_STATE_START __LINK_STATE_NOCARRIER carrier_changes for e.g. user mostly treats IFF_RUNNING as the indication of link up/down as opposed to IFF_UP. That would require you to toggle __LINK_STATE_START (and operstate as well) without doing a full dev_close/open. Since __LINK_STATE_START is cleared, there's no sense to let CARRIER_OK remain set, and then you'd need to take care of carrier_changes... Since you don't really shutting down the device, the link watchdog keeps running and may race with inconsistent carrier state in between. dev_close/open may have done unneeded work, but it's the safest option IMHO, as apparently the cost and ugly complexity to fake link down/up events is not something worthwhile compared to simply bouncing the link state. Another point is kernel consumers of the NETDEV_CHANGENAME notifier might well assume the link is already taken down by dev_close() before the rename. I didn't check all those consumers in tree but thought it might be safe to keep the current convention. Now let me turn around and ask you what's your concerns if bouncing the link state. While I can tweak a lightweight version of dev_close/open to bypass ndo_stop and ndo_start while shutting down the link watchdog on behalf of drivers, it's far more involved than make me think if that's really what you had in mind. Another less safer option is that we just notify userspace anyway without sending down/up event around, as I don't see *any real application* cares about the link state or whatsoever when it attempts to detect rename. Given that the scope is limited to failover slave the chance of breaking userspace app would be extremely low in practice. Thanks, -Siwei