OpenSBI Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: bin.yao@ingenic.com
To: opensbi@lists.infradead.org
Subject: [Question] Possible tlb_sync hang when remote SFENCE.VMA races with HSM hart_stop
Date: $(date -R)	[thread overview]
Message-ID: <codex-opensbi-rfence-hsm-$(date +%s)@ingenic.com> (raw)

Hello OpenSBI maintainers,

We are seeing a hang during aging/stress testing and would like to check whether our analysis is correct.

Environment:
- OpenSBI base version: v1.8.1
- Base commit in our tree: 74434f255873d74e56cc50aa762d1caf24c099f8
- Downstream tree with platform-specific HSM hart_stop support
- Platform: downstream Ingenic platform

Observed symptom:
When one hart executes remote SFENCE.VMA while another hart is stopping, the sender can spin forever in tlb_sync() waiting for a completion that never arrives.

Our current understanding of the race is:

1. The source hart calls sbi_tlb_request() and then sbi_ipi_send_many().
2. sbi_ipi_send_many() builds the target mask using sbi_hsm_hart_interruptible_mask(), so a target hart can still be selected while it is in SBI_HSM_STATE_STARTED.
3. Before the source completes the send path, the target hart can execute sbi_hsm_hart_stop() and transition from STARTED to STOP_PENDING.
4. The source still queues the TLB request in tlb_update() and increments its local tlb_sync counter.
5. That counter is decremented only when the target hart later runs tlb_entry_process().
6. If the target hart has already passed the point in the stop path where pending IPIs are drained, the queued TLB request is never processed, so the source remains in tlb_sync() forever.

So this looks like a race between:
- the interruptible-mask snapshot in sbi_ipi_send_many(), and
- the STARTED -> STOP_PENDING transition in sbi_hsm_hart_stop()

In other words, tlb_sync itself does not appear to be wrong; the problem seems to be that a hart can be counted as a target, but then stop before it can process the queued TLB request.

To make the window easier to hit during debugging, we temporarily added an artificial delay in sbi_ipi_send_many() in our downstream tree. However, our understanding is that the delay only increases the reproduction probability, and the race exists in principle even without that modification.

Questions:
1. Does the analysis above look correct?
2. Is this race already known?
3. Is there an existing or preferred way to serialize remote rfence/IPI send against hart_stop?
4. Could this be related to issue #402 ("Lifelock condition under high load with many TLB shootdowns"), or is it better treated as a separate issue?

If useful, I can also send a minimal timeline or a proposed fix.

Thanks,
Bin Yao
Ingenic

-- 
opensbi mailing list
opensbi@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/opensbi

             reply	other threads:[~2026-07-03  8:22 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-03  8:22 bin.yao [this message]
     [not found] <6a477160.de7c5347.2799ad.72abSMTPIN_ADDED_BROKEN@mx.google.com>
2026-07-03 12:03 ` [Question] Possible tlb_sync hang when remote SFENCE.VMA races with HSM hart_stop Anup Patel
  -- strict thread matches above, loose matches on Subject: below --
2026-09-11  3:44 Alex Mao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to='codex-opensbi-rfence-hsm-$(date +%s)@ingenic.com' \
    --to=bin.yao@ingenic.com \
    --cc=opensbi@lists.infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox