From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail-wr0-f195.google.com ([209.85.128.195]:44757 "EHLO mail-wr0-f195.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753340AbeDRIc0 (ORCPT ); Wed, 18 Apr 2018 04:32:26 -0400 Subject: Re: kernel panics with 4.14.X versions To: Jan Kara Cc: Guillaume Morin , stable@vger.kernel.org, decui@microsoft.com, jack@suse.com, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, mszeredi@redhat.com References: <20180416132550.d25jtdntdvpy55l3@bender.morinfr.org> <20180416144041.t2mt7ugzwqr56ka3@quack2.suse.cz> <9b11cfba-4bdc-8a3e-cd33-2f7e8d513bdf@gmail.com> <20180417121207.cs7eijrndovbplgz@quack2.suse.cz> From: Pavlos Parissis Message-ID: <9cb08428-66ed-2306-d2f2-ae734863c68d@gmail.com> Date: Wed, 18 Apr 2018 10:32:21 +0200 MIME-Version: 1.0 In-Reply-To: <20180417121207.cs7eijrndovbplgz@quack2.suse.cz> Content-Type: multipart/signed; micalg=pgp-sha256; protocol="application/pgp-signature"; boundary="UFUmR4Vwbu75721t8O4xW35cTcncTHMH7" Sender: stable-owner@vger.kernel.org List-ID: This is an OpenPGP/MIME signed message (RFC 4880 and 3156) --UFUmR4Vwbu75721t8O4xW35cTcncTHMH7 Content-Type: multipart/mixed; boundary="OD0x0zwTmnxJLqQgUvkt3tsPYxkdgQ2gy"; protected-headers="v1" From: Pavlos Parissis To: Jan Kara Cc: Guillaume Morin , stable@vger.kernel.org, decui@microsoft.com, jack@suse.com, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, mszeredi@redhat.com Message-ID: <9cb08428-66ed-2306-d2f2-ae734863c68d@gmail.com> Subject: Re: kernel panics with 4.14.X versions References: <20180416132550.d25jtdntdvpy55l3@bender.morinfr.org> <20180416144041.t2mt7ugzwqr56ka3@quack2.suse.cz> <9b11cfba-4bdc-8a3e-cd33-2f7e8d513bdf@gmail.com> <20180417121207.cs7eijrndovbplgz@quack2.suse.cz> In-Reply-To: <20180417121207.cs7eijrndovbplgz@quack2.suse.cz> --OD0x0zwTmnxJLqQgUvkt3tsPYxkdgQ2gy Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: quoted-printable On 17/04/2018 02:12 =CE=BC=CE=BC, Jan Kara wrote: > On Tue 17-04-18 01:31:24, Pavlos Parissis wrote: >> On 16/04/2018 04:40 =CE=BC=CE=BC, Jan Kara wrote: >=20 > >=20 >>> How easily can you hit this? >> >> Very easily, I only need to wait 1-2 days for a crash to occur. >=20 > I wouldn't call that very easily but opinions may differ :). Anyway it'= s > good (at least for debugging) that it's reproducible. >=20 Unfortunately, I can't reproduce it, so waiting 1-2 days is the only opti= on I have. >>> Are you able to run debug kernels >> >> Well, I was under the impression I do as I have: >> grep -E 'DEBUG_KERNEL|DEBUG_INFO' /boot/config-4.14.32-1.el7.x86_64 >> CONFIG_DEBUG_INFO=3Dy >> # CONFIG_DEBUG_INFO_REDUCED is not set >> # CONFIG_DEBUG_INFO_SPLIT is not set >> # CONFIG_DEBUG_INFO_DWARF4 is not set >> CONFIG_DEBUG_KERNEL=3Dy >> >> Do you think that my kernel doesn't produce a proper crash dump? >> I have a production cluster where I can run any kernel we need, so if = I need >> to compile again with different settings I can certainly do that. >=20 > OK, good. So please try running 4.16 as you mention below to verify whe= ther > this is just a -stable regression or also a problem in the current upst= ream > kernel. Based on your results with 4.16 I'll prepare a debug patch for = you to > apply on top of 4.14.32 so that we can debug this further. >=20 >>> / inspect >>> crash dumps when the issue occurs? >> >> I can't do that as the server isn't responsive and I can only power cy= cle it. >=20 > Well, kernel crash dumps work in that situation as well - when the kern= el > panics, it will kexec into a new kernel and dump memory of the old kern= el > to disk. It can then be investigated with the 'crash' utility. But > obviously you don't have this set up and don't have experience with thi= s so > let's go via a standard 'debug patch' route. >=20 >>> Also testing with the latest mainline >>> kernel (4.16) would be welcome whether this isn't just an issue with = the >>> backport of fsnotify fixes from Miklos. >> >> I can try the kernel-ml-4.16.2 from elrepo (we use CentOS 7). >=20 > Yes, that would be good. >=20 I have production server running 4.16.2 and no kernel crash dumps yet. Let's wait another day before we say anything. Cheers, Pavlos --OD0x0zwTmnxJLqQgUvkt3tsPYxkdgQ2gy-- --UFUmR4Vwbu75721t8O4xW35cTcncTHMH7 Content-Type: application/pgp-signature; name="signature.asc" Content-Description: OpenPGP digital signature Content-Disposition: attachment; filename="signature.asc" -----BEGIN PGP SIGNATURE----- iQIzBAEBCAAdFiEEHZDZK6DBu+YUj6+dg/yS2h9xdrkFAlrXApYACgkQg/yS2h9x drnkLg/9E42gL7RHvAQTvCza/d89rB5HjVORXezDw66xfEEPKHFP+j/EzHxdd+QU hiW2gXWErFzLMpr8NhVgu4WDc7fOyWXJ58IxF5oneNgPHv1jjoDzkzCFoJzpG8Tg MdBsrEHajVwZZ1l9xdE1y8Qkhanb8U0lIeoprpBE0upHnT3EkhlWl6bhQDsYp4u8 5Xu15ZM3V+JB5mSgjpJZCiEkd3Xqk9oPgHc714X+MA4Qc7mCrHo6hhd5UpQV5HUz Iz/PMDjqufqlc8HQR/sc+XJtGufeOyvQa4KcKey8BBusYscc5XrTrODteg9Px8pa oGPd9BUec9mrxN9kPCsVXAmtcdFRp9/KTuiU9hQc3wfEL1Eo3imzOGVkr5wxvF0f xigUBOcraoak0eCUkvIEbMRphgtG7rEkTS6PLHZHCUh4KiXtYjnvmwb6Qk6aNhfy GP7vMTC47kw4lss2gC9jZQjVM9cq7ze7tljmmYRWnzf8k8rB2JmozcmLcTr5pHee KiC/sD0G0eDVtfnS3uP3GouvMRsQvI2FzjwIBZJI4ML4FSChcbklf0TI1fZbQW6I qmAsqnoUpencnJvxwvCNsA6x2WfwSdg+lUOC4IAw1Reeza8rNvzoRmF+hFOcwxQS jWKX8EkABVG/1212d5nZWUjYrDReCV5L7PiswXoHqdDfqQu0Hew= =8w/O -----END PGP SIGNATURE----- --UFUmR4Vwbu75721t8O4xW35cTcncTHMH7--