From: Josh Poimboeuf <jpoimboe@redhat.com>
To: Fengguang Wu <fengguang.wu@intel.com>
Cc: Byungchul Park <byungchul.park@lge.com>,
Ingo Molnar <mingo@kernel.org>,
"Peter Zijlstra (Intel)" <peterz@infradead.org>,
linux-kernel@vger.kernel.org, LKP <lkp@01.org>,
Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>,
torvalds@linux-foundation.org, bp@alien8.de, x86@kernel.org,
"H. Peter Anvin" <hpa@zytor.com>,
Thomas Gleixner <tglx@linutronix.de>
Subject: Re: [lockdep] b09be676e0 BUG: unable to handle kernel NULL pointer dereference at 000001f2
Date: Tue, 3 Oct 2017 11:28:15 -0500 [thread overview]
Message-ID: <20171003162815.2eatamg4h3tpsypl@treble> (raw)
In-Reply-To: <20171003150538.57dq7xe6afikpzyp@treble>
On Tue, Oct 03, 2017 at 10:05:38AM -0500, Josh Poimboeuf wrote:
> I don't know the lockdep code, but one more comment from the peanut
> gallery. This code looks suspect to me:
>
>
> /*
> * Stop saving stack_trace if save_trace() was
> * called at least once:
> */
> if (save && ret == 2)
> save = NULL;
>
>
> From looking at check_prev_add(), a return value of 2 doesn't
> necessarily imply that save_trace() was called. If the
> check_redundant() call returns 0, then check_prev_add() can return 2,
> and the trace will still be uninitialized, but save will be set to NULL
> even though save_trace() hasn't been called. Then a subsequent call to
> check_prev_add() could add an uninitialized stack_trace struct to the
> dependency list.
>
> I could be wrong, but it's at least something the lockdep folks might
> want to look at.
[ Different manifestations of this bug have been discussed in several
different threads. Bringing partipants from those threads onto CC. ]
So, looking at the actual panic:
BUG: unable to handle kernel NULL pointer dereference at 000001f2
IP: update_stack_state+0xd4/0x340
*pde = 00000000
Oops: 0000 [#1] PREEMPT SMP
CPU: 0 PID: 18728 Comm: 01-cpu-hotplug Not tainted 4.13.0-rc4-00170-gb09be67 #592
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.9.3-20161025_171302-gandalf 04/01/2014
task: bb0b53c0 task.stack: bb3ac000
EIP: update_stack_state+0xd4/0x340
EFLAGS: 00010002 CPU: 0
EAX: 0000a570 EBX: bb3adccb ECX: 0000f401 EDX: 0000a570
ESI: 00000001 EDI: 000001ba EBP: bb3adc6b ESP: bb3adc3f
DS: 007b ES: 007b FS: 00d8 GS: 0000 SS: 0068
CR0: 80050033 CR2: 000001f2 CR3: 0b3a7000 CR4: 00140690
DR0: 00000000 DR1: 00000000 DR2: 00000000 DR3: 00000000
DR6: fffe0ff0 DR7: 00000400
Call Trace:
? unwind_next_frame+0xea/0x400
? __unwind_start+0xf5/0x180
? __save_stack_trace+0x81/0x160
? save_stack_trace+0x20/0x30
? __lock_acquire+0xfa5/0x12f0
? lock_acquire+0x1c2/0x230
? tick_periodic+0x3a/0xf0
? _raw_spin_lock+0x42/0x50
? tick_periodic+0x3a/0xf0
? tick_periodic+0x3a/0xf0
? debug_smp_processor_id+0x12/0x20
? tick_handle_periodic+0x23/0xc0
? local_apic_timer_interrupt+0x63/0x70
? smp_trace_apic_timer_interrupt+0x235/0x6a0
? trace_apic_timer_interrupt+0x37/0x3c
? strrchr+0x23/0x50
Code: 0f 95 c1 89 c7 89 45 e4 0f b6 c1 89 c6 89 45 dc 8b 04 85 98 cb 74 bc 88 4d e3 89 45 f0 83 c0 01 84 c9 89 04 b5 98 cb 74 bc 74 3b <8b> 47 38 8b 57 34 c6 43 1d 01 25 00 00 02 00 83 e2 03 09 d0 83
EIP: update_stack_state+0xd4/0x340 SS:ESP: 0068:bb3adc3f
CR2: 00000000000001f2
---[ end trace 0d147fd4aba8ff50 ]---
Kernel panic - not syncing: Fatal exception in interrupt
There are two bugs:
1) Somebody -- presumably lockdep -- is corrupting the stack. Need the
lockdep people to look at that.
2) The 32-bit FP unwinder isn't handling the corrupt stack very well,
It's blindly dereferencing untrusted data:
/* Is the next frame pointer an encoded pointer to pt_regs? */
regs = decode_frame_pointer(next_bp);
if (regs) {
frame = (unsigned long *)regs;
len = regs_size(regs);
state->got_irq = true;
On 32-bit, regs_size() dereferences the regs pointer before we know it
points to a valid stack. I'll fix that, along with the other unwinder
improvements I discussed with Linus.
--
Josh
next prev parent reply other threads:[~2017-10-03 16:28 UTC|newest]
Thread overview: 51+ messages / expand[flat|nested] mbox.gz Atom feed top
2017-10-03 14:06 [lockdep] b09be676e0 BUG: unable to handle kernel NULL pointer dereference at 000001f2 Fengguang Wu
2017-10-03 14:31 ` Josh Poimboeuf
2017-10-03 14:41 ` Josh Poimboeuf
2017-10-03 15:05 ` Josh Poimboeuf
2017-10-03 16:28 ` Josh Poimboeuf [this message]
2017-10-03 17:34 ` Josh Poimboeuf
2017-10-03 21:44 ` Tetsuo Handa
2017-10-04 21:06 ` Josh Poimboeuf
2017-10-04 21:30 ` Linus Torvalds
2017-10-04 22:15 ` Josh Poimboeuf
2017-10-04 22:40 ` Josh Poimboeuf
2017-10-05 11:02 ` Tetsuo Handa
2017-10-05 13:57 ` Josh Poimboeuf
2017-10-04 8:34 ` Peter Zijlstra
2017-10-10 5:57 ` Byungchul Park
2017-10-03 16:54 ` Linus Torvalds
2017-10-03 16:57 ` Linus Torvalds
2017-10-10 5:48 ` Byungchul Park
2017-10-10 16:22 ` Linus Torvalds
2017-10-10 16:56 ` Linus Torvalds
2017-10-10 18:14 ` Peter Zijlstra
2017-10-10 18:38 ` Linus Torvalds
2017-10-11 1:14 ` Byungchul Park
2017-10-11 2:36 ` Byungchul Park
2017-10-11 0:56 ` Byungchul Park
2017-10-11 1:02 ` Byungchul Park
2017-10-12 1:15 ` Byungchul Park
2017-10-03 17:18 ` Ingo Molnar
2017-10-04 9:20 ` Peter Zijlstra
2017-10-04 10:31 ` Ingo Molnar
2017-10-04 14:15 ` Josh Poimboeuf
2017-10-10 5:30 ` Byungchul Park
2017-10-05 13:01 ` Josh Poimboeuf
2017-10-05 14:54 ` Josh Poimboeuf
2017-10-09 10:50 ` Peter Zijlstra
2017-10-09 12:21 ` Fengguang Wu
2017-10-09 12:54 ` Peter Zijlstra
2017-10-09 12:59 ` Fengguang Wu
2017-10-09 13:03 ` Josh Poimboeuf
2017-10-09 12:55 ` Fengguang Wu
2017-10-09 13:26 ` Josh Poimboeuf
2017-10-09 14:17 ` Fengguang Wu
2017-10-09 15:28 ` Peter Zijlstra
2017-10-09 15:41 ` Fengguang Wu
2017-10-09 15:44 ` Peter Zijlstra
2017-10-09 15:47 ` Fengguang Wu
2017-10-10 5:08 ` Byungchul Park
2017-10-12 8:47 ` Peter Zijlstra
2017-10-12 9:21 ` Fengguang Wu
2017-10-12 9:28 ` Fengguang Wu
2017-10-12 11:45 ` Peter Zijlstra
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20171003162815.2eatamg4h3tpsypl@treble \
--to=jpoimboe@redhat.com \
--cc=bp@alien8.de \
--cc=byungchul.park@lge.com \
--cc=fengguang.wu@intel.com \
--cc=hpa@zytor.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lkp@01.org \
--cc=mingo@kernel.org \
--cc=penguin-kernel@I-love.SAKURA.ne.jp \
--cc=peterz@infradead.org \
--cc=tglx@linutronix.de \
--cc=torvalds@linux-foundation.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox