From: Eric Wong <normalperson@yhbt.net>
To: Eric Dumazet <eric.dumazet@gmail.com>
Cc: Al Viro <viro@ZenIV.linux.org.uk>,
netdev <netdev@vger.kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Re: strange crashes in tcp_poll() via epoll_wait
Date: Sat, 20 Jul 2013 02:03:18 +0000 [thread overview]
Message-ID: <20130720020318.GA12731@dcvr.yhbt.net> (raw)
In-Reply-To: <1374279005.26476.31.camel@edumazet-glaptop>
Eric Dumazet <eric.dumazet@gmail.com> wrote:
> On Fri, 2013-07-19 at 23:50 +0000, Eric Wong wrote:
> > Eric Dumazet <eric.dumazet@gmail.com> wrote:
> > > Hi Al
> > >
> > > I tried to debug strange crashes in tcp_poll() called from
> > > sys_epoll_wait() -> sock_poll()
> > >
> > > The symptom is that sock->sk is NULL and we therefore dereference a NULL
> > > pointer.
> > >
> > > It's really rare crashes but still, it would be nice to understand where
> > > is the bug. Presumably latest kernels would crash in sock_poll() because
> > > of the sk_can_busy_loop(sock->sk) call.
> > >
> > > We do test sock->sk being NULL in sock_fasync(), but epoll should be
> > > safe because of existing synchronization (epmutex) ?
> >
> > It should be safe because of ep->mtx, actually, as epmutex is not taken
> > in sys_epoll_wait.
>
> Hmm, it might be more complex than that for multi threaded programs :
>
> eventpoll_release_file()
>
> The problem might be because a thread closes a socket while an event
> was queued for it.
But ep->mtx is also held when traversing the ready list with
ep_send_events_proc.
Can sock->sk somehow be NULL before hitting eventpoll_release_file?
> > I took a look at this but have not found anything. I've yet to see this
> > this on my machines.
> >
> > When did you start noticing this?
>
> Hard to say, but we have these crashes on a 3.3+ based kernel.
So I don't think any of my epoll changes caused it. Phew!
> Probability of said crashes is very very low.
This still worries me since I rely heavily on multi-threaded epoll. I
don't have a lot of cores/CPUs, though, so maybe it's harder to trigger
any potential race as a result...
prev parent reply other threads:[~2013-07-20 2:03 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2013-07-19 16:24 strange crashes in tcp_poll() via epoll_wait Eric Dumazet
2013-07-19 23:50 ` Eric Wong
2013-07-20 0:10 ` Eric Dumazet
2013-07-20 2:03 ` Eric Wong [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20130720020318.GA12731@dcvr.yhbt.net \
--to=normalperson@yhbt.net \
--cc=eric.dumazet@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=viro@ZenIV.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox