From: Keith Busch <kbusch@meta.com>
To: <fio@vger.kernel.org>
Cc: <axboe@kernel.dk>, <vincentfu@gmail.com>,
Keith Busch <kbusch@kernel.org>
Subject: [PATCHv2 2/4] backend: don't mark a job as running before it has set up
Date: Thu, 10 Sep 2026 09:13:08 -0700 [thread overview]
Message-ID: <20260910161310.1478081-3-kbusch@meta.com> (raw)
In-Reply-To: <20260910161310.1478081-1-kbusch@meta.com>
From: Keith Busch <kbusch@kernel.org>
run_threads() moves a job from TD_INITIALIZED straight to TD_RUNNING and
only then releases it. The job still has all of its setup left to do:
exec_prerun, pre_read_files(), setup_files(), init_io_u(),
rate_submit_init() and finally set_epoch_time(). The job claims to be
running for that entire window when it has not issued any IO and has no
epoch to measure itself against.
Anything that reads the state during that window gets a wrong answer.
thread_eta() is the most readily observable one: it derives elapsed
from td->epoch, which is still zeroed, so a job with a three second
exec_prerun reports
Jobs: 1 (f=0): [R(1)][-.-%][eta 00m:00s]
Jobs: 1 (f=0): [R(1)][50.0%][eta 00m:03s]
before doing any work at all. Running, half done and no time left, none
of which is true.
Promote to TD_SETTING_UP instead, which is what that state is for, and
let the job promote itself to TD_RAMP or TD_RUNNING once setup is over
and it has recorded its epoch. The same job now reports
Jobs: 1 (f=1): [I(1)][0.0%][eta 00m:03s]
Widen the TERMINATE_STONEWALL check to match. It tests for runstate >=
TD_RUNNING to find jobs worth terminating, and a job in the setup window
used to satisfy that. Without this, exit_what=stonewall stops reaching a
job that is still setting up, and a test where the short job is reaped
while a longer one sits in exec_prerun goes from 5.5s to 20.5s. Note
TD_RAMP sorts below TD_SETTING_UP, so ramping jobs remain excluded from
that check exactly as before.
Signed-off-by: Keith Busch <kbusch@kernel.org>
---
backend.c | 21 ++++++++++++++++-----
libfio.c | 3 ++-
2 files changed, 18 insertions(+), 6 deletions(-)
diff --git a/backend.c b/backend.c
index 7f41bdfa..46bff828 100644
--- a/backend.c
+++ b/backend.c
@@ -2178,6 +2178,17 @@ static void *thread_main(void *data)
goto err;
set_epoch_time(td, o->log_alternate_epoch_clock_id, o->job_start_clock_id);
+
+ /*
+ * Setup is done and the job now has an epoch to measure itself
+ * against, so it is finally safe to call it running. Everything
+ * above this point ran as TD_SETTING_UP.
+ */
+ if (in_ramp_period(td))
+ td_set_runstate(td, TD_RAMP);
+ else
+ td_set_runstate(td, TD_RUNNING);
+
fio_getrusage(&td->ru_start);
memcpy(&td->bw_sample_time, &td->epoch, sizeof(td->epoch));
memcpy(&td->iops_sample_time, &td->epoch, sizeof(td->epoch));
@@ -2888,16 +2899,16 @@ reap:
}
/*
- * start created threads (TD_INITIALIZED -> TD_RUNNING).
+ * start created threads (TD_INITIALIZED -> TD_SETTING_UP).
+ * The job has plenty of setup left to do before it issues any
+ * IO, so it promotes itself to TD_RAMP or TD_RUNNING once that
+ * is done and it has recorded its epoch.
*/
for_each_td(td) {
if (td->runstate != TD_INITIALIZED)
continue;
- if (in_ramp_period(td))
- td_set_runstate(td, TD_RAMP);
- else
- td_set_runstate(td, TD_RUNNING);
+ td_set_runstate(td, TD_SETTING_UP);
nr_running++;
nr_started--;
m_rate += ddir_rw_sum(td->o.ratemin);
diff --git a/libfio.c b/libfio.c
index a57ede4f..322906c0 100644
--- a/libfio.c
+++ b/libfio.c
@@ -270,7 +270,8 @@ void fio_terminate_threads(unsigned int group_id, unsigned int terminate)
for_each_td(td) {
if ((terminate == TERMINATE_GROUP && group_id == TERMINATE_ALL) ||
(terminate == TERMINATE_GROUP && group_id == td->groupid) ||
- (terminate == TERMINATE_STONEWALL && td->runstate >= TD_RUNNING) ||
+ (terminate == TERMINATE_STONEWALL &&
+ td->runstate >= TD_SETTING_UP) ||
(terminate == TERMINATE_ALL)) {
dprint(FD_PROCESS, "setting terminate on %s/%d\n",
td->o.name, (int) td->pid);
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-10 16:13 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 16:13 [PATCHv2 0/4] fio: fix eta Keith Busch
2026-09-10 16:13 ` [PATCHv2 1/4] eta: count a job that is setting up as running Keith Busch
2026-09-10 16:13 ` Keith Busch [this message]
2026-09-11 19:49 ` [PATCHv2 2/4] backend: don't mark a job as running before it has set up Vincent Fu
2026-09-10 16:13 ` [PATCHv2 3/4] eta: cap the ETA at the job's own remaining runtime Keith Busch
2026-09-10 16:13 ` [PATCHv2 4/4] eta: remove now unused done_secs Keith Busch
2026-09-11 19:45 ` [PATCHv2 0/4] fio: fix eta Vincent Fu
2026-09-11 21:09 ` fiotestbot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910161310.1478081-3-kbusch@meta.com \
--to=kbusch@meta.com \
--cc=axboe@kernel.dk \
--cc=fio@vger.kernel.org \
--cc=kbusch@kernel.org \
--cc=vincentfu@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.