From: sashiko-bot@kernel.org
To: "Anthony Krowiak" <akrowiak@linux.ibm.com>
Cc: "Christian Borntraeger" <borntraeger@linux.ibm.com>,
"Alexander Gordeev" <agordeev@linux.ibm.com>,
"Heiko Carstens" <hca@linux.ibm.com>,
"Vasily Gorbik" <gor@linux.ibm.com>,
linux-s390@vger.kernel.org, kvm@vger.kernel.org
Subject: Re: [PATCH v9 3/6] s390/vfio-ap: Fix unbounded loop in apq_reset_check()
Date: Tue, 29 Sep 2026 12:38:14 +0000 [thread overview]
Message-ID: <20260929123815.287871F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260929121837.2715710-4-akrowiak@linux.ibm.com>
Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [Critical] crypto: zeroization bypass and DMA vulnerability if ZAPQ returns BUSY
- [Critical] crypto: Use-After-Free of vdev pointer after reset timeout
- [High] crypto: resource leak and erroneous WARN when ZAPQ returns BUSY
- [High] crypto: zeroization bypass if non-reset state change completes at timeout
- [Medium] crypto: spurious ZAPQs from conflating TAPQ and ZAPQ BUSY responses
- [Low] crypto: dead code setting fallback reset status on timeout
--
commit 625818127ace62cbd33b251880fe0f1f6d732f69
Author: Anthony Krowiak <akrowiak@linux.ibm.com>
s390/vfio-ap: Fix unbounded loop in apq_reset_check()
This commit fixes an unbounded wait loop in the apq_reset_check() worker by
adding a 2000ms timeout. It also updates TAPQ status checks to accurately
verify asynchronous queue zeroization completion before freeing AQIC
resources.
> diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_ap_ops.c
> index 087e8474a34ab..07fbfa6f10152 100644
> --- a/drivers/s390/crypto/vfio_ap_ops.c
> +++ b/drivers/s390/crypto/vfio_ap_ops.c
> @@ -2164,6 +2164,12 @@ static int apq_status_check(int apqn, struct ap_queue_status *status)
> return -EBUSY;
>
> case AP_RESPONSE_BUSY:
> + /*
> + * The queue is busy with something unrelated to a reset and our
> + * ZAPQ was rejected outright. Re-issue the ZAPQ.
> + */
> + return -EAGAIN;
> +
[Severity: Medium]
Does returning `-EAGAIN` here trigger spurious ZAPQs by conflating TAPQ and
ZAPQ BUSY responses?
`apq_status_check()` evaluates the output of `ap_tapq()`. If `ap_tapq()`
returns `AP_RESPONSE_BUSY`, it means the read instruction itself was
transiently rejected. However, this code appears to incorrectly assume it
means the previous ZAPQ was rejected outright, returning `-EAGAIN`.
This forces the worker loop to immediately issue a new `ap_zapq()`, which
overwrites `q->reset_status` and destroys the correct queue state tracking,
likely leading to further failures.
[ ... ]
> @@ -2223,30 +2280,73 @@ static void apq_reset_check(struct work_struct *reset_work)
> */
> memcpy(&q->reset_status, &status, sizeof(status));
> return;
> - }
> - if (ret == -EBUSY) {
> + } else if (elapsed >= AP_RESET_MAX_WAIT) {
> + /*Timed out without being able to verify zapq completed */
> + if (!ret || ret == -ENODEV) {
[Severity: High]
Does evaluating the timeout condition before the retry logic allow an
unrelated state change completion to bypass the required zeroization step?
If the initial ZAPQ returns `AP_RESPONSE_STATE_CHANGE_IN_PROGRESS`, the worker
polls until it completes before retrying ZAPQ. If the state change completes
exactly at `elapsed >= AP_RESET_MAX_WAIT`, `ap_tapq()` indicates the queue is
idle, and `ret = 0`.
Because the loop evaluates the timeout condition `else if (elapsed >=
AP_RESET_MAX_WAIT)` before the ZAPQ retry logic, the timeout block executes
`apq_reset_finalize()`, forces the status to `AP_RESPONSE_NORMAL`, and returns.
The required `ap_zapq()` is entirely bypassed, leaving the queue unzeroized
while freeing AQIC resources, resulting in a DMA vulnerability.
> + /*
> + * Zeroization confirmed (ret == 0): the TAPQ status bits
> + * indicate the async portion of the ZAPQ completed
> + * successfully. Free AQIC resources and return.
> + *
> + * Queue non-operational (ret == -ENODEV): the queue is
> + * deconfigured or checkstopped; interrupts are not
> + * possible so AQIC resources can be safely freed.
> + * Zeroization cannot be confirmed in this state, but the
> + * queue cannot generate interrupts, so the NIB page is
> + * no longer a DMA target and it is safe to free it.
> + */
> + apq_reset_finalize(q, &status, ret);
> + return;
> + }
> +
> + report_apq_reset_check_timeout(q);
> +
> + /*
> + * Zeroization could not be confirmed; set
> + * reset_status to AP_RESPONSE_RESET_IN_PROGRESS.
> + * This is used internally to signal that the reset
> + * did not complete, and ensures that if the queue
> + * is reset again, the re-issue logic in
> + * apq_reset_check() will re-issue the ZAPQ.
> + */
> + q->reset_status.response_code = AP_RESPONSE_RESET_IN_PROGRESS;
[Severity: Low]
Is this assignment to `q->reset_status.response_code` dead code?
The timeout logic assigns `AP_RESPONSE_RESET_IN_PROGRESS` to theoretically
trigger a re-issue during the next reset. However, the next reset requires
calling `vfio_ap_mdev_reset_queue()`, which immediately executes
`status = ap_zapq(q->apqn, 0)` and unconditionally overwrites
`q->reset_status` via `memcpy`.
> +
> + return;
[Severity: Critical]
Does returning early here skip cleanup and leave `q->aqic_resources.vdev`
as a dangling pointer?
`q->aqic_resources.vdev` points to the `vdev` embedded in `ap_matrix_mdev`
without holding a reference count. When the new timeout mechanism returns
early without calling `apq_reset_finalize()`, it leaves
`q->aqic_resources.vdev` populated.
When the mediated device is removed, `vfio_ap_mdev_remove()` flushes the
worker and calls `vfio_put_device()`, which frees the `ap_matrix_mdev`.
This turns `q->aqic_resources.vdev` into a dangling pointer on a surviving
`vfio_ap_queue`.
Later, if the AP queue is unbound or reused, `vfio_ap_free_aqic_resources()`
unconditionally dereferences this pointer via
`vfio_unpin_pages(q->aqic_resources.vdev, ...)`, which could trigger a
use-after-free on the host.
> + } else if (ret == -EBUSY) {
> pr_notice_ratelimited(WAIT_MSG, elapsed,
> AP_QID_CARD(q->apqn),
> AP_QID_QUEUE(q->apqn),
> status.response_code,
> status.queue_empty,
> status.irq_enabled);
> + continue;
> } else {
> - if (q->reset_status.response_code == AP_RESPONSE_RESET_IN_PROGRESS ||
> - q->reset_status.response_code == AP_RESPONSE_BUSY ||
> - q->reset_status.response_code == AP_RESPONSE_STATE_CHANGE_IN_PROGRESS ||
> - ret == -EAGAIN) {
> + if (ret == -EAGAIN ||
> + q->reset_status.response_code == AP_RESPONSE_RESET_IN_PROGRESS ||
> + q->reset_status.response_code == AP_RESPONSE_STATE_CHANGE_IN_PROGRESS) {
[Severity: Critical]
Is it possible for `AP_RESPONSE_BUSY` to be missed in this retry check?
If `ap_zapq()` inside the worker loop is rejected with `AP_RESPONSE_BUSY`,
`q->reset_status.response_code` is set to `AP_RESPONSE_BUSY`.
On the next iteration, if the queue becomes idle, `ap_tapq()` returns
`AP_RESPONSE_NORMAL` (ret = 0).
Since `AP_RESPONSE_BUSY` is missing from the retry condition, the worker seems
to skip re-issuing the ZAPQ. It would fall through to `apq_reset_finalize()`,
falsely overriding the status to `AP_RESPONSE_NORMAL` and freeing the AQIC
resources. Can this leave secret keys intact and expose the host to wild
DMA writes?
> + status = ap_zapq(q->apqn, 0);
> + memcpy(&q->reset_status, &status, sizeof(status));
> continue;
> }
I also have a question about `vfio_ap_mdev_reset_queue()`:
drivers/s390/crypto/vfio_ap_ops.c:vfio_ap_mdev_reset_queue() {
...
switch (status.response_code) {
case AP_RESPONSE_NORMAL:
case AP_RESPONSE_RESET_IN_PROGRESS:
case AP_RESPONSE_STATE_CHANGE_IN_PROGRESS:
...
}
[Severity: High]
Does lacking `case AP_RESPONSE_BUSY:` here cause legitimately busy queues
to fail in the default case?
If the queue is transiently busy, ZAPQ returns `AP_RESPONSE_BUSY`. Execution
falls to the `default:` branch, triggering a host WARN (falsely claiming a
hardware bug) and intentionally leaking AQIC resources without queueing the
worker to retry. Could this allow a guest that keeps the queue busy to cause
resource leaks and kernel warnings on the host when the device is removed or
reset?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929121837.2715710-1-akrowiak@linux.ibm.com?part=3
next prev parent reply other threads:[~2026-09-29 12:38 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 12:18 [PATCH v9 0/6] s390/vfio-ap: Fix pre-existing bugs in vfio_ap device driver Anthony Krowiak
2026-09-29 12:18 ` [PATCH v9 1/6] s390/vfio-ap: Fix leaks of pinned NIB and registered GISC Anthony Krowiak
2026-09-29 12:33 ` sashiko-bot
2026-09-29 13:38 ` Anthony Krowiak
2026-09-29 12:18 ` [PATCH v9 2/6] s390/vfio-ap: Fix failure to release IRQ notification eventfd contexts Anthony Krowiak
2026-09-29 12:25 ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 3/6] s390/vfio-ap: Fix unbounded loop in apq_reset_check() Anthony Krowiak
2026-09-29 12:38 ` sashiko-bot [this message]
2026-09-29 16:57 ` Anthony Krowiak
2026-09-29 12:18 ` [PATCH v9 4/6] s390/vfio-ap: Use AP_DOMAINS for adm_add bitmap size in vfio_ap_mdev_cfg_add() Anthony Krowiak
2026-09-29 12:27 ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 5/6] s390/vfio-ap: fix queue state leakage to guest and host Anthony Krowiak
2026-09-29 12:33 ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 6/6] s390/vfio-ap: replace guest-reachable WARNs with ratelimited warnings Anthony Krowiak
2026-09-29 12:41 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929123815.287871F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=agordeev@linux.ibm.com \
--cc=akrowiak@linux.ibm.com \
--cc=borntraeger@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=kvm@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox