Linux NFS development
 help / color / mirror / Atom feed
* [PATCH 0/4] mountd/exportd: exit signal handling fixups
@ 2026-09-24 15:29 Benjamin Coddington
  2026-09-24 15:29 ` [PATCH 1/4] support/export: take SIGINT/SIGTERM/SIGHUP via signalfd Benjamin Coddington
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Benjamin Coddington @ 2026-09-24 15:29 UTC (permalink / raw)
  To: Steve Dickson; +Cc: linux-nfs

mountd and exportd exit from inside their SIGTERM handler: killer()
unregisters with rpcbind, frees, logs and calls exit(), none of which is
async-signal-safe.  It mostly gets away with it because mountd spends
nearly all its time in select().

Setting rootdir in the [exports] section of nfs.conf makes it much
easier to hit.  nfsd_path_init() then starts a workqueue thread to do
path lookups in the chroot, and it does that after killer() is
installed.  pthread_create() allocates the new thread's TLS with
calloc(), so a SIGTERM there - say, from a reboot while mountd is still
coming up - runs killer() on top of a half-finished malloc.  We hit this
in a reboot test:

  malloc(): unsorted double linked list corrupted
  ...
  calloc
  clnt_vc_create
  local_rpcb
  rpcb_unset
  nfs_svc_unregister
  unregister_services
  killer
  <signal handler called>
  _int_malloc
  calloc
  allocate_dtv
  _dl_allocate_tls
  pthread_create
  main

Once that thread exists malloc takes the arena lock, so the same
re-entry later can deadlock instead of aborting.  Out of 400 SIGTERMs
sent during startup with rootdir set, one mountd hung on a futex, which
I think is that, though I didn't catch a stack.

These patches take HUP, INT and TERM through a signalfd in the cache
loop, so the shutdown work runs in normal context.  The parent of forked
workers keeps a handler, but all it does is kill(0, SIGTERM).  The last
patch holds signals while mountd registers with rpcbind rather than
ignoring them; a SIGTERM there used to just get dropped.

Tested in a private net namespace with its own rpcbind: foreground and
-t 4, HUP, TERM, an ha-callout, and 400 SIGTERMs at random points in
startup with rootdir set.  No aborts or hangs, and nothing dropped.

Benjamin Coddington (4):
  support/export: take SIGINT/SIGTERM/SIGHUP via signalfd
  mountd: stop exiting from a signal handler
  exportd: stop exiting from a signal handler
  mountd: hold SIGTERM during rpcbind registration instead of ignoring
    it

 support/export/cache.c       | 122 ++++++++++++++++++++++++++++++++++-
 support/export/export.h      |   2 +
 support/include/ha-callout.h |   5 +-
 utils/exportd/exportd.c      |  30 +++------
 utils/mountd/mountd.c        |  32 +++------
 utils/mountd/svc_run.c       |   2 +
 6 files changed, 147 insertions(+), 46 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-24 15:29 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 15:29 [PATCH 0/4] mountd/exportd: exit signal handling fixups Benjamin Coddington
2026-09-24 15:29 ` [PATCH 1/4] support/export: take SIGINT/SIGTERM/SIGHUP via signalfd Benjamin Coddington
2026-09-24 15:29 ` [PATCH 2/4] mountd: stop exiting from a signal handler Benjamin Coddington
2026-09-24 15:29 ` [PATCH 3/4] exportd: " Benjamin Coddington
2026-09-24 15:29 ` [PATCH 4/4] mountd: hold SIGTERM during rpcbind registration instead of ignoring it Benjamin Coddington

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox