Linux Kernel Selftest development
 help / color / mirror / Atom feed
* [PATCH] selftests/harness: Drain the test process group after a timeout
@ 2026-09-10 23:02 Alex Williamson
  0 siblings, 0 replies; only message in thread
From: Alex Williamson @ 2026-09-10 23:02 UTC (permalink / raw)
  To: kees
  Cc: Alex Williamson, luto, wad, shuah, linux-kselftest, linux-kernel,
	Alex Williamson, matlack, kvm

The harness runs each test in a child that it forks in __run_test().
That child's PID is saved as t->pid and the child calls setpgrp() to
make itself the leader of a new process group, PGID.  For a fixture
test, that child only runs the TEST_F() wrapper.  The wrapper forks a
*grandchild* to run the fixture body (ie., FIXTURE_SETUP -> test ->
FIXTURE_TEARDOWN), and that grandchild is what actually acquires
resources such as a vfio device fd.  The grandchild inherits the
group, so t->pid and the grandchild share PGID, t->pid.

On timeout __wait_for_test() kills that whole group with
kill(-(t->pid), SIGKILL), but at best only t->pid is reaped by
waitpid().  The grandchild exits asynchronously, while the harness
proceeds to the next tests.  Any resources required by those next
tests that are still owned exclusively by the grandchild process
result in a cascade of failures through those subsequent tests.

Instead, the harness needs to not only actively reap t->pid itself,
but it needs to poll the process group to provide time for the
grandchild, and any processes it may have created, to exit.

Additionally the harness itself needs to avoid getting blocked by an
uninterruptible test process, therefore process reaping is bounded to
5s, which allows completion of longer running release paths, such as
those including device resets in vfio-pci.

Assisted-by: LLM
Signed-off-by: Alex Williamson <alex.williamson@nvidia.com>
---
 tools/testing/selftests/kselftest_harness.h | 49 +++++++++++++++++----
 1 file changed, 41 insertions(+), 8 deletions(-)

diff --git a/tools/testing/selftests/kselftest_harness.h b/tools/testing/selftests/kselftest_harness.h
index 29a19bc87084..4e479c8cb49e 100644
--- a/tools/testing/selftests/kselftest_harness.h
+++ b/tools/testing/selftests/kselftest_harness.h
@@ -80,6 +80,7 @@ static inline void __kselftest_memset_safe(void *s, int c, size_t n)
 #define KSELFTEST_PRIO_XFAIL   20001
 
 #define TEST_TIMEOUT_DEFAULT 30
+#define TEST_TIMEOUT_DRAIN_MS 5000
 
 /* Utilities exposed to the test definitions */
 #ifndef TH_LOG_STREAM
@@ -981,8 +982,7 @@ static void __wait_for_test(struct __test_metadata *t)
 	 */
 	int status = KSFT_FAIL << 8;
 	struct pollfd poll_child;
-	int ret, child, childfd;
-	bool timed_out = false;
+	int ret, child = 0, childfd;
 
 	childfd = syscall(__NR_pidfd_open, t->pid, 0);
 	if (childfd == -1) {
@@ -1004,9 +1004,46 @@ static void __wait_for_test(struct __test_metadata *t)
 			t->name);
 		return;
 	} else if (ret == 0) {
-		timed_out = true;
+		int elapsed_ms = 0;
+
 		/* signal process group */
 		kill(-(t->pid), SIGKILL);
+
+		/*
+		 * wait(2): "A child that terminates, but has not been waited
+		 * for becomes a "zombie"...  As long as a zombie is not removed
+		 * from the system via a wait, it will consume a slot in the
+		 * kernel process table..."  Therefore, wait for the wrapper
+		 * process to exit and only then poll whether the process group
+		 * also still exists.  Only when the process group is no longer
+		 * found are all the processes exited.
+		 */
+		for (;;) {
+			if (child != t->pid) {
+				child = waitpid(t->pid, &status, WNOHANG);
+				if (child == -1 && errno != EINTR)
+					break;
+			}
+
+			if (child == t->pid &&
+			    kill(-(t->pid), 0) == -1 && errno == ESRCH)
+				break;
+
+			if (elapsed_ms >= TEST_TIMEOUT_DRAIN_MS) {
+				fprintf(TH_LOG_STREAM,
+					"# %s: process group not reaped %dms after timeout SIGKILL (task stuck in D state?); continuing\n",
+					t->name, TEST_TIMEOUT_DRAIN_MS);
+				break;
+			}
+
+			usleep(10 * 1000);
+			elapsed_ms += 10;
+		}
+
+		t->exit_code = KSFT_FAIL;
+		fprintf(TH_LOG_STREAM,
+			"# %s: Test terminated by timeout\n", t->name);
+		return;
 	}
 	child = waitpid(t->pid, &status, WNOHANG);
 	if (child == -1 && errno != EINTR) {
@@ -1017,11 +1054,7 @@ static void __wait_for_test(struct __test_metadata *t)
 		return;
 	}
 
-	if (timed_out) {
-		t->exit_code = KSFT_FAIL;
-		fprintf(TH_LOG_STREAM,
-			"# %s: Test terminated by timeout\n", t->name);
-	} else if (WIFEXITED(status)) {
+	if (WIFEXITED(status)) {
 		if (WEXITSTATUS(status) == KSFT_SKIP ||
 		    WEXITSTATUS(status) == KSFT_XPASS ||
 		    WEXITSTATUS(status) == KSFT_XFAIL) {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] only message in thread

only message in thread, other threads:[~2026-09-10 23:03 UTC | newest]

Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-10 23:02 [PATCH] selftests/harness: Drain the test process group after a timeout Alex Williamson

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox