The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Eric Dumazet <eric.dumazet@gmail.com>
To: Peter LaDow <petela@gocougs.wsu.edu>
Cc: linux-kernel@vger.kernel.org
Subject: Re: Process Hang in __read_seqcount_begin
Date: Mon, 22 Oct 2012 19:01:39 +0200	[thread overview]
Message-ID: <1350925299.8609.978.camel@edumazet-glaptop> (raw)
In-Reply-To: <CAN8Q1EdobyFuAvtajGRSqNsCG+XoPA+FkwN=2j_H126mmfEjTQ@mail.gmail.com>

On Mon, 2012-10-22 at 09:46 -0700, Peter LaDow wrote:
> I posted this problem some time back on the linux-rt-users and
> netfilter lists.  Since then, we thought we had a workaround to avoid
> this problem, so we dropped the issue.  But now 5 months later, the
> problem has reappeared.  And this time it is much more serious and
> much more difficult to re-create.  After perusing both those lists,
> I'm not sure if those were the proper places to post.  The netfilter
> list seems to be more focused on the user space side of things, and
> the RT page indicates that kernel side RT issues should go to lkml.
> 
> Anyway, here's a repost of that problem from July.  Perhaps somebody
> here can point us in the right direction.
> 
> We are running 3.0.36-rt57 on a powerpc box.  During some testing with
> heavy loads and interfaces coming up/going down (specifically PPP), we
> have run into a case where iptables hangs and cannot be killed.  It
> requires a reboot to fix the problem.
> 
> Connecting the BDI and debugging the kernel, we get:
> 
> #0  get_counters (t=0xdd5145a0, counters=0xe3458000)
>     at include/linux/seqlock.h:66
> #1  0xc026b4ac in do_ipt_get_ctl (sk=<value optimized out>,
>     cmd=<value optimized out>, user=0x10612078, len=<value optimized out>)
>     at net/ipv4/netfilter/ip_tables.c:918
> #2  0xc022226c in nf_sockopt (sk=<value optimized out>, pf=2 '\002',
>     val=<value optimized out>, opt=<value optimized out>, len=0xdd4c7d4c,
>     get=1) at net/netfilter/nf_sockopt.c:109
> #3  0xc0236b1c in ip_getsockopt (sk=0xdf071480, level=<value optimized out>,
>     optname=65, optval=0x10612078 <Address 0x10612078 out of bounds>,
>     optlen=0xbfbe0c2c) at net/ipv4/ip_sockglue.c:1308
> #4  0xc02522a8 in raw_getsockopt (sk=0xdf071480, level=<value optimized out>,
>     optname=<value optimized out>, optval=<value optimized out>,
>     optlen=<value optimized out>) at net/ipv4/raw.c:811
> #5  0xc01f4c38 in sock_common_getsockopt (sock=<value optimized out>,
>     level=<value optimized out>, optname=<value optimized out>,
>     optval=<value optimized out>, optlen=<value optimized out>)
>     at net/core/sock.c:2157
> #6  0xc01f2df8 in sys_getsockopt (fd=<value optimized out>, level=0,
>     optname=65, optval=0x10612078 <Address 0x10612078 out of bounds>,
>     optlen=0xbfbe0c2c) at net/socket.c:1839
> #7  0xc01f45b4 in sys_socketcall (call=15, args=<value optimized out>)
>     at net/socket.c:2421
> 
> It seems to be stuck in __read_seqcount_begin.  From include/linux/seqlock.h:
> 
> static inline unsigned __read_seqcount_begin(const seqcount_t *s)
> {
>         unsigned ret;
> 
> repeat:
>         ret = ACCESS_ONCE(s->sequence);
>         if (unlikely(ret & 1)) {
>                 cpu_relax();
>   <----- It is always here
>                 goto repeat;
>         }
>         return ret;
> }
> 
> I've been scouring the mailing lists and Google searches trying to
> find something, but thus far have come up with nothing.
> 
> Any tips would be appreciated.

This looks like a corruption of s->sequence, and is value is odd, even
if no writer is alive.

Does local_bh_disable() disables preemption on RT ?





  reply	other threads:[~2012-10-22 17:01 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2012-10-22 16:46 Process Hang in __read_seqcount_begin Peter LaDow
2012-10-22 17:01 ` Eric Dumazet [this message]
2012-10-22 17:25   ` Peter LaDow
2012-10-22 19:24     ` Peter LaDow
2012-10-24  0:15       ` Peter LaDow
2012-10-24  4:32         ` Eric Dumazet
2012-10-24 16:30           ` Peter LaDow
2012-10-24 16:44             ` Eric Dumazet
2012-10-26 16:15           ` Peter LaDow
2012-10-26 16:48             ` Eric Dumazet
2012-10-26 18:51               ` Peter LaDow
2012-10-26 20:53                 ` Thomas Gleixner
2012-10-26 21:05                 ` Eric Dumazet
2012-10-26 21:25                   ` Peter LaDow
2012-10-26 21:54                     ` Thomas Gleixner
2012-10-30 23:09                       ` Peter LaDow
2012-10-31  0:33                         ` Thomas Gleixner

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1350925299.8609.978.camel@edumazet-glaptop \
    --to=eric.dumazet@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=petela@gocougs.wsu.edu \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox