From mboxrd@z Thu Jan 1 00:00:00 1970 From: Michael Holzheu Subject: [PATCH v3] kdump: Fix crash_kexec - smp_send_stop race in panic Date: Tue, 29 Nov 2011 09:58:52 +0100 Message-ID: <1322557132.2751.6.camel@br98xy6r> References: <1319639649.3321.11.camel@br98xy6r> <20111028161143.e5ebf617.akpm@linux-foundation.org> <1320055036.2796.8.camel@br98xy6r> <20111031033948.a0edb7f3.akpm@linux-foundation.org> <1320314844.2989.6.camel@br98xy6r> <20111109160400.cc2d27d9.akpm@linux-foundation.org> <1320934932.16425.14.camel@br98xy6r> <4EBBE9B4.3040009@tilera.com> <1321014494.2745.7.camel@br98xy6r> <4EBD5536.7010806@tilera.com> Reply-To: holzheu-23VcF4HTsmIX0ybBhKVfKdBPR1lH4CV8@public.gmane.org Mime-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Return-path: In-Reply-To: <4EBD5536.7010806-kv+TWInifGbQT0dZR+AlfA@public.gmane.org> List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: kexec-bounces-IAPFreCvJWM7uuMidbF8XUB+6BGkLq7r@public.gmane.org Errors-To: kexec-bounces+glkk-kexec=m.gmane.org-IAPFreCvJWM7uuMidbF8XUB+6BGkLq7r@public.gmane.org To: Andrew Morton Cc: Fenghua Yu , Benjamin Herrenschmidt , Heiko Carstens , David Howells , Chen Liqin , Paul Mackerras , "H. Peter Anvin" , Guan Xuetao , Lennox Wu , Hans-Christian Egtvedt , Jonas Bonn , Jesper Nilsson , Russell King , Yoshinori Sato , Richard Weinberger , Helge Deller , "James E.J. Bottomley" , Ingo Molnar , Geert Uytterhoeven , Matt Turner , Vivek Goyal , Haavard Skinnemoen , Don Zickus List-Id: linux-arch.vger.kernel.org From: Michael Holzheu Subject: kdump: Fix crash_kexec()/smp_send_stop() race in panic When two CPUs call panic at the same time there is a possible race condition that can stop kdump. The first CPU calls crash_kexec() and the second CPU calls smp_send_stop() in panic() before crash_kexec() finished on the first CPU. So the second CPU stops the first CPU and therefore kdump fails: 1st CPU: panic()->crash_kexec()->mutex_trylock(&kexec_mutex)-> do kdump 2nd CPU: panic()->crash_kexec()->kexec_mutex already held by 1st CPU ->smp_send_stop()-> stop 1st CPU (stop kdump) This patch fixes the problem by introducing a spinlock in panic that allows only one CPU to process crash_kexec() and the subsequent panic code. All other CPUs call the weak function panic_smp_self_stop() that stops the CPU itself. This function can be overloaded by architecture code. For example "tile" can use their lower-power "nap" instruction for that. Acked-by: Chris Metcalf Signed-off-by: Michael Holzheu --- kernel/panic.c | 18 +++++++++++++++++- 1 file changed, 17 insertions(+), 1 deletion(-) --- a/kernel/panic.c +++ b/kernel/panic.c @@ -49,6 +49,15 @@ static long no_blink(int state) long (*panic_blink)(int state); EXPORT_SYMBOL(panic_blink); +/* + * Stop ourself in panic -- architecture code may override this + */ +void __weak panic_smp_self_stop(void) +{ + while (1) + cpu_relax(); +} + /** * panic - halt the system * @fmt: The text string to print @@ -59,6 +68,7 @@ EXPORT_SYMBOL(panic_blink); */ NORET_TYPE void panic(const char * fmt, ...) { + static DEFINE_SPINLOCK(panic_lock); static char buf[1024]; va_list args; long i, i_next = 0; @@ -68,8 +78,14 @@ NORET_TYPE void panic(const char * fmt, * It's possible to come here directly from a panic-assertion and * not have preempt disabled. Some functions called from here want * preempt to be disabled. No point enabling it later though... + * + * Only one CPU is allowed to execute the panic code from here. For + * multiple parallel invocations of panic, all other CPUs either + * stop themself or will wait until they are stopped by the 1st CPU + * with smp_send_stop(). */ - preempt_disable(); + if (!spin_trylock(&panic_lock)) + panic_smp_self_stop(); console_verbose(); bust_spinlocks(1);