All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/4] mountd/exportd: exit signal handling fixups
@ 2026-09-24 15:29 Benjamin Coddington
  2026-09-24 15:29 ` [PATCH 1/4] support/export: take SIGINT/SIGTERM/SIGHUP via signalfd Benjamin Coddington
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Benjamin Coddington @ 2026-09-24 15:29 UTC (permalink / raw)
  To: Steve Dickson; +Cc: linux-nfs

mountd and exportd exit from inside their SIGTERM handler: killer()
unregisters with rpcbind, frees, logs and calls exit(), none of which is
async-signal-safe.  It mostly gets away with it because mountd spends
nearly all its time in select().

Setting rootdir in the [exports] section of nfs.conf makes it much
easier to hit.  nfsd_path_init() then starts a workqueue thread to do
path lookups in the chroot, and it does that after killer() is
installed.  pthread_create() allocates the new thread's TLS with
calloc(), so a SIGTERM there - say, from a reboot while mountd is still
coming up - runs killer() on top of a half-finished malloc.  We hit this
in a reboot test:

  malloc(): unsorted double linked list corrupted
  ...
  calloc
  clnt_vc_create
  local_rpcb
  rpcb_unset
  nfs_svc_unregister
  unregister_services
  killer
  <signal handler called>
  _int_malloc
  calloc
  allocate_dtv
  _dl_allocate_tls
  pthread_create
  main

Once that thread exists malloc takes the arena lock, so the same
re-entry later can deadlock instead of aborting.  Out of 400 SIGTERMs
sent during startup with rootdir set, one mountd hung on a futex, which
I think is that, though I didn't catch a stack.

These patches take HUP, INT and TERM through a signalfd in the cache
loop, so the shutdown work runs in normal context.  The parent of forked
workers keeps a handler, but all it does is kill(0, SIGTERM).  The last
patch holds signals while mountd registers with rpcbind rather than
ignoring them; a SIGTERM there used to just get dropped.

Tested in a private net namespace with its own rpcbind: foreground and
-t 4, HUP, TERM, an ha-callout, and 400 SIGTERMs at random points in
startup with rootdir set.  No aborts or hangs, and nothing dropped.

Benjamin Coddington (4):
  support/export: take SIGINT/SIGTERM/SIGHUP via signalfd
  mountd: stop exiting from a signal handler
  exportd: stop exiting from a signal handler
  mountd: hold SIGTERM during rpcbind registration instead of ignoring
    it

 support/export/cache.c       | 122 ++++++++++++++++++++++++++++++++++-
 support/export/export.h      |   2 +
 support/include/ha-callout.h |   5 +-
 utils/exportd/exportd.c      |  30 +++------
 utils/mountd/mountd.c        |  32 +++------
 utils/mountd/svc_run.c       |   2 +
 6 files changed, 147 insertions(+), 46 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-24 15:29 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 15:29 [PATCH 0/4] mountd/exportd: exit signal handling fixups Benjamin Coddington
2026-09-24 15:29 ` [PATCH 1/4] support/export: take SIGINT/SIGTERM/SIGHUP via signalfd Benjamin Coddington
2026-09-24 15:29 ` [PATCH 2/4] mountd: stop exiting from a signal handler Benjamin Coddington
2026-09-24 15:29 ` [PATCH 3/4] exportd: " Benjamin Coddington
2026-09-24 15:29 ` [PATCH 4/4] mountd: hold SIGTERM during rpcbind registration instead of ignoring it Benjamin Coddington

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.