From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-5.5 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI,SIGNED_OFF_BY, SPF_PASS,URIBL_BLOCKED,USER_AGENT_MUTT autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id AB380C169C4 for ; Tue, 29 Jan 2019 15:19:48 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 7A5B521848 for ; Tue, 29 Jan 2019 15:19:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=default; t=1548775188; bh=gMUjzoLUShW+IisWyq0Zy3daLEXE+c9Pv85teha5rPA=; h=Date:From:To:Cc:Subject:References:In-Reply-To:List-ID:From; b=Qw3D15JZqLo1VvtIwZcmPOIsytimePs0YGBnULvgQellsb4ZFq7WvPUQMvaULvRgk i3jqrFj9bUupYLL3h0pqtB8uWgep8/6b4uyr66Pzu39L+38IgY73G2x5xKuf/HNyXe Luert7nXWu6BsDrFl859wOFWNZsgesLNd9CzleBU= Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1728227AbfA2PTr (ORCPT ); Tue, 29 Jan 2019 10:19:47 -0500 Received: from mail.kernel.org ([198.145.29.99]:52420 "EHLO mail.kernel.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1727977AbfA2PTq (ORCPT ); Tue, 29 Jan 2019 10:19:46 -0500 Received: from localhost (5356596B.cm-6-7b.dynamic.ziggo.nl [83.86.89.107]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by mail.kernel.org (Postfix) with ESMTPSA id 0F0F0214DA; Tue, 29 Jan 2019 15:19:44 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=default; t=1548775185; bh=gMUjzoLUShW+IisWyq0Zy3daLEXE+c9Pv85teha5rPA=; h=Date:From:To:Cc:Subject:References:In-Reply-To:From; b=nS3g38D2rRo+7v+pzPjmczyqoiLZgQNJ5VDiqp6vPb3hBdvYCpWmibwWoikNoBBiY biufv4+hF2l4S+sDvjn7jAZYfFHcWufbiLeAm57qPJzbfKOr8Lxuny4S9BrAWKUgV/ Evsm9Liw2CtgSvO23k5n8kou/p9M4WE9TBklDzG0= Date: Tue, 29 Jan 2019 16:19:42 +0100 From: Greg KH To: Guenter Roeck Cc: Zubin Mithra , "# v4 . 10+" , Guenter Roeck , benh@kernel.crashing.org, rafael@kernel.org Subject: Re: [PATCH v4.4.y] drivers: core: Remove glue dirs from sysfs earlier Message-ID: <20190129151942.GB15719@kroah.com> References: <20190128173130.22618-1-zsm@chromium.org> <20190129065836.GA24231@kroah.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.11.2 (2019-01-07) Sender: stable-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: stable@vger.kernel.org On Tue, Jan 29, 2019 at 06:35:48AM -0800, Guenter Roeck wrote: > On Mon, Jan 28, 2019 at 10:58 PM Greg KH wrote: > > > > On Mon, Jan 28, 2019 at 09:31:30AM -0800, Zubin Mithra wrote: > > > From: Benjamin Herrenschmidt > > > > > > commit 726e41097920a73e4c7c33385dcc0debb1281e18 upstream > > > > > > For devices with a class, we create a "glue" directory between > > > the parent device and the new device with the class name. > > > > > > This directory is never "explicitely" removed when empty however, > > > this is left to the implicit sysfs removal done by kobject_release() > > > when the object loses its last reference via kobject_put(). > > > > > > This is problematic because as long as it's not been removed from > > > sysfs, it is still present in the class kset and in sysfs directory > > > structure. > > > > > > The presence in the class kset exposes a use after free bug fixed > > > by the previous patch, but the presence in sysfs means that until > > > the kobject is released, which can take a while (especially with > > > kobject debugging), any attempt at re-creating such as binding a > > > new device for that class/parent pair, will result in a sysfs > > > duplicate file name error. > > > > > > This fixes it by instead doing an explicit kobject_del() when > > > the glue dir is empty, by keeping track of the number of > > > child devices of the gluedir. > > > > > > This is made easy by the fact that all glue dir operations are > > > done with a global mutex, and there's already a function > > > (cleanup_glue_dir) called in all the right places taking that > > > mutex that can be enhanced for this. It appears that this was > > > in fact the intent of the function, but the implementation was > > > wrong. > > > > > > Backport Note: kref_read() is not present in 4.4. Hence, > > > use atomic_read(&kref.refcount) instead of kref_read(&kref). > > > > > > Signed-off-by: Benjamin Herrenschmidt > > > Acked-by: Linus Torvalds > > > Signed-off-by: Greg Kroah-Hartman > > > Signed-off-by: Zubin Mithra > > > --- > > > drivers/base/core.c | 2 ++ > > > include/linux/kobject.h | 17 +++++++++++++++++ > > > 2 files changed, 19 insertions(+) > > > > Wait, why is this needed? > > > > We have syzcaller/syzbot hits, and there is a reproducer. Is this sufficient ? Nice, where? > > And why only for 4.4? What about 4.9 and 4.14? Do you want to upgrade > > and suddenly hit the same "bug" that you fixed before? > > > > Good point. Sorry, that got lost. We are, after all, oly human. I'd > ask Zubin to provide backports, or do it myself, but we'll have to > resolve the issue you bring up below first. > > > There was a reason that I did not backport this to the stable tree when > > it was submitted, and that was because this was an odd race to ever hit. > > Are you hitting this in the real world without kobject deferred > > release enabled? And if so, are you hitting the WARN_ON that is added > > here? > > > I think we may need updated rules for stable. Many bug fixes are > backported to stable releases without having been seen in the "real > world" (whatever that means). At the same time we do see many races in > the real world, many of them not fully understood. Our policy so far > is to fix as many problems as possible if they are understood, in the > hope that it fixes at least some of those problems seen in the field. I don't have a problem with backporting this, but it went in as a "fixes a theoritical issue, and let's WARN_ON if it really ever is hit". So I would like to see where we are hitting that WARN_ON, as it sounds like you found an easy-to-reproduce way to do it. > If there is a new rule that problems have to have been observed in the > real world (ie without debugging options enabled, and without specific > reproducer) before a patch is applied to stable, can we have that as > generic rule that is not only selectively applied ? It's not a specific rule, it's just that this specific patch went through a 30+ email chain of us arguing about it, so I really want to see why this is needed, and why Ben was right and I was wrong :) I'm not trying to be extra hard here, it's just that I have a lot of history with this patch... So if you all have a reproducer, I would love to see it. Oh, and the backports to the other kernel versions as well, you don't want to have to fix this again in a year when you all upgrade to a new release. thanks, greg k-h