From: Keith Busch <kbusch@meta.com>
To: <fio@vger.kernel.org>
Cc: <axboe@kernel.dk>, <vincentfu@gmail.com>,
Keith Busch <kbusch@kernel.org>
Subject: [PATCH 2/3] eta: cap the ETA at the job's own remaining runtime
Date: Thu, 3 Sep 2026 10:21:00 -0700 [thread overview]
Message-ID: <20260903172101.1886315-3-kbusch@meta.com> (raw)
In-Reply-To: <20260903172101.1886315-1-kbusch@meta.com>
From: Keith Busch <kbusch@kernel.org>
The ETA of a running job is capped at "timeout + done_secs - elapsed".
done_secs is a global accumulator of the runtime of every job reaped so
far, so the cap grows every time a job finishes. Whenever the cap is
what actually gets reported, the ETA jumps back up by the runtime of
everything that has already completed.
The cap is what gets reported when the progress based estimate exceeds
the remaining runtime, that is when perc is below elapsed/timeout. Any
job that is behind on bytes relative to its runtime is in that state,
so this covers the common "size the job to the whole device, bound it
with runtime" pattern. A job that would complete its size early stays
at or above elapsed/timeout and never reaches the cap.
time_based makes no difference either way. It only lowers perc to
min(perc, elapsed/timeout), so a time_based job that cannot finish its
size within the runtime is affected exactly like a size based one.
The problem is most visible with stonewalled jobs, where the ETA climbs
back to the full run time at every batch boundary instead of counting
down. Start stonewalled io_uring jobs, size=10T runtime=5 time_based,
report:
22 21 20 24 23 22 21 20 24 23 22 21 20 24 23 22 21 20 24 ...
instead of counting down to 0.
It is not specific to stonewall. Two concurrent jobs with runtime=5 and
runtime=20 show the same jump when the short one is reaped at t=5:
19 18 17 16 15 19 18 17 16 15 14 ...
done_secs made sense when it was introduced: thread_eta() was handed
the global elapsed time back then, so "timeout + done_secs" was this
job's projected finish time relative to the start of the whole run.
b29ee5b3 switched elapsed to be per job, measured from td->epoch, but
left the done_secs term behind.
A job is terminated once utime_since(&td->epoch, now) reaches
td->o.timeout, so with a per job elapsed the cap is simply
"timeout - elapsed". Use that, and clamp at zero instead of relying on
the unsigned subtraction wrapping.
Fixes: b29ee5b3dee4 ("Update ramp_time")
Signed-off-by: Keith Busch <kbusch@kernel.org>
---
eta.c | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
diff --git a/eta.c b/eta.c
index 5dbf4d75..29e1bf41 100644
--- a/eta.c
+++ b/eta.c
@@ -263,9 +263,18 @@ static unsigned long thread_eta(struct thread_data *td)
eta_sec = (unsigned long) (elapsed * (1.0 / perc)) - elapsed;
}
- if (td->o.timeout &&
- eta_sec > (timeout + done_secs - elapsed))
- eta_sec = timeout + done_secs - elapsed;
+ /*
+ * A job never runs for longer than its own timeout, which is
+ * measured from its own epoch. Cap the estimate at whatever
+ * time this job has left.
+ */
+ if (td->o.timeout) {
+ unsigned long timeout_left;
+
+ timeout_left = timeout > elapsed ? timeout - elapsed : 0;
+ if (eta_sec > timeout_left)
+ eta_sec = timeout_left;
+ }
} else if (runstate == TD_NOT_CREATED || runstate == TD_CREATED
|| runstate == TD_INITIALIZED
|| runstate == TD_SETTING_UP
--
2.52.0
next prev parent reply other threads:[~2026-09-03 17:21 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 17:20 [PATCH 0/3] fio: fix eta Keith Busch
2026-09-03 17:20 ` [PATCH 1/3] eta: don't estimate runtime from an unset job epoch Keith Busch
2026-09-09 17:01 ` Vincent Fu
2026-09-09 20:15 ` Keith Busch
2026-09-03 17:21 ` Keith Busch [this message]
2026-09-03 17:21 ` [PATCH 3/3] eta: remove now unused done_secs Keith Busch
2026-09-03 20:01 ` [PATCH 0/3] fio: fix eta fiotestbot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260903172101.1886315-3-kbusch@meta.com \
--to=kbusch@meta.com \
--cc=axboe@kernel.dk \
--cc=fio@vger.kernel.org \
--cc=kbusch@kernel.org \
--cc=vincentfu@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox