All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback
@ 2026-08-09  8:49 Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 1/6] alpha: run check_mmu_context() from finish_arch_post_lock_switch() Magnus Lindholm
                   ` (5 more replies)
  0 siblings, 6 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

On Alpha, stale TLB translations can break copy-on-write and shared-mapping
writeback: a multi-threaded process can lose stores to its own private
memory, read data belonging to its own child, and lose data written through
a shared file mapping. The copy-on-write failures need more than one CPU;
the writeback failure also happens on a uniprocessor.

Three related problems are fixed here.

Patch 1 is stranded deferred-ASN bookkeeping. check_mmu_context() clears
asn_lock and acts on need_new_asn, but it runs only as the tail of
switch_to(), after alpha_switch_to() returns. A newly forked task never
gets there: its first context switch resumes at ret_from_fork instead.
asn_lock is left set and the task goes on to run user space with it set and
interrupts enabled, so a shootdown IPI arriving in that window takes the
deferred path, and the need_new_asn handshake meant to cover it never runs.
finish_task_switch() calls finish_arch_post_lock_switch() with preemption
disabled, on the CPU that ran switch_mm(), which is where that bookkeeping
can be completed.

  kthread_use_mm() and sched_force_init_mm() also reach the same hook,
  outside the scheduler's preemption-disabled switch tail.
  check_mmu_context() acts on per-CPU state, so it can only complete this
  bookkeeping while still on the CPU that ran switch_mm(). The hook
  therefore tests preemptible() directly: where migration is possible it
  does nothing, while where preemption remains disabled, or is not
  configured, completing the bookkeeping is safe.

The second problem is a targeted tbi() issued against the wrong context.
tbi() acts on the address space context currently loaded on a CPU, so it
reaches an mm's translations only when a thread of that mm is current.
current->active_mm is not sufficient: under lazy TLB an idle or kernel task
keeps an mm as its active_mm while a different ASN is loaded, so the
invalidate hits the wrong context, and nothing retires the old ASN -
mm->context[cpu] is still valid, so the resuming thread reuses it.  Patch 2
fixes the shootdown IPI handler, patch 3 the local side of
flush_tlb_page(), patch 4 the uniprocessor flush_tlb_page().

The third is an omitted caller-side invalidate, covered by patches 3, 5
and 6: when the target mm is not the calling CPU's active_mm, nothing
invalidates that CPU at all, because smp_call_function() does not call back
into the caller. The UP flush_tlb_mm() and flush_icache_user_page() already
contain exactly the missing branch. Their active_mm tests are left alone:
both load a fresh context through __load_new_mm_context() rather than a
targeted tbi() against whatever ASN happened to be loaded, and only the
targeted tbi() depends on which context is loaded.

No user-space data race is involved in the reproducers: every slot is
written and read by one thread only, and the main thread inspects them only
after joining the workers. The other writes come from forked children with
their own address space, so observing one of those in the parent is the bug.

Deferred-path behaviour after the fixes, counted with the counters reset
before each 6-second run of the lost-store reproducer:

	                       lost stores    runs entering the window
	  no fixes               11 of 40      13 of 40, 11 of them lost
	  patched                 0 of 400     15 of 400, none lost

Where the remaining patches are reached, counted the same way:

	                                  thread of mm     lazily
	                                  current          borrowing
	  writeback of a shared mapping       51577          50688
	  reclaim under memory pressure       17609          14827
	  anonymous COW / fork                144237             0

  About half the calls during writeback, and none at all on anonymous
  memory, which is why patches 1 and 2 did not cover it. flush_tlb_mm() was
  entered with the mm not this CPU's active_mm 2934 times over a fork-heavy
  run and 632 times while otherwise idle.

Reproducing it. Two self-contained tests were written for this. Source for
both can be made available on request.

  alpha-cow-smoketest.c covers patches 1 and 2 (pthreads only, ~40s, needs
  more than one CPU; confined to one with taskset -c 0 it does not fail). A
  failing run reports:

	stale read after COW fault       FAIL
	    thread 2 read 0xdeadbeefcafebabe, expected 0x1000002   <- the CHILD's value

  Two details in it matter, because getting either wrong hides the bug:
  slots are 128 bytes apart so several threads share a page, and a thread
  writes its slot once then reads it many times, since a thread that keeps
  writing refreshes its own translation. Its lost-stores check fires in at
  best a quarter of runs; the stale-read check is the reliable one.

  mkclean4.c covers patches 3 and 4 (needs root), and fails on both SMP and
  UP. A single thread writes a small MAP_SHARED file while background
  writeback cleans it, and the file is then compared against the mapping;
  every round loses data. It forces two conditions that are rare in normal
  operation on SMP: the flusher kworker and the writer on the same CPU, and
  a working set small enough to stay resident in the data TLB. That is also
  why this is hard to hit in the field - any faulting write from any CPU
  repairs the dirty state. On a uniprocessor the first condition holds by
  construction, so no pinning is needed there.

  Patches 5 and 6 have no reproducer for the bug they fix; they are
  justified by the contract of the functions, by the UP implementations
  already having the missing branch, and by the counts above. Patch 6's
  path was exercised for regressions by driving gdb to set and clear a
  breakpoint several hundred times, reaching copy_to_user_page() ->
  flush_icache_user_page().

  Originally found as intermittent heap corruption in glibc's
  malloc/tst-malloc-fork-deadlock-malloc-check. glibc is not at fault: with
  MALLOC_CHECK_=3 it is simply a very effective detector, and every fork()
  runs __malloc_fork_unlock_child() in the child, which writes to allocator
  state.

Testing. ES40, EV68AL (21264C) Tsunami, 3 CPUs. Generated against
v7.2-rc6; the results below are from v7.2-rc2, and the series was
re-tested on a clean v7.2-rc1 tree with the same outcome.

	                                   before          after
	  smoke test, stale-read check     7 of 9 rounds   0 of 9
	  lost stores                      11 of 40        0 of 400
	  writeback (mkclean4), SMP        10 of 10        0 of 10
	  writeback (mkclean4), UP         every round     0 of 8
	  tst-malloc-fork-deadlock-
	    malloc-check                   8 of 10 fail    25/25 pass
	  ptrace breakpoint exerciser      -               result matches

  Also 10/10 pass each for tst-malloc-fork-deadlock, tst-malloc-check,
  tst-tcfree1-malloc-check and tst-tcfree2-malloc-check. No measurable
  cost: 1470/1038/206 forks per second with 0/2/8 sibling threads against
  1472/1043/199 unpatched. Patch 5 changes a function none of the
  reproducers exercise, so it is covered for regressions only.

  Also run with CONFIG_COMPACTION and CONFIG_MIGRATION enabled, with
  compaction forced continuously underneath the tests: 368522 folios
  migrated during the run, no failures.

  Also built and tested with CONFIG_ALPHA_GENERIC and CONFIG_SMP=n.  That is
  how patch 4 was found: arch/alpha/kernel/smp.c is not built there, so
  patches 2, 3, 5 and 6 are absent and patch 1 is inert, and mkclean4.c
  failed on every round because the uniprocessor flush_tlb_page() carries
  the same defect. With patch 4 it passes 8 of 8, and the rest of the tests
  above pass there too.

Magnus Lindholm (6):
  alpha: run check_mmu_context() from finish_arch_post_lock_switch()
  alpha: only use a targeted tbi() when the target mm is really current
  alpha: fix the local TLB invalidate in flush_tlb_page()
  alpha: only use a targeted tbi() when the target mm is really current
    (UP)
  alpha: invalidate the local context in flush_tlb_mm()
  alpha: invalidate the local context in flush_icache_user_page()

 arch/alpha/include/asm/mmu_context.h | 28 +++++++++++++++++++
 arch/alpha/include/asm/tlbflush.h    |  9 +++++-
 arch/alpha/kernel/smp.c              | 41 ++++++++++++++++++++++++++--
 3 files changed, 75 insertions(+), 3 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH 1/6] alpha: run check_mmu_context() from finish_arch_post_lock_switch()
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 2/6] alpha: only use a targeted tbi() when the target mm is really current Magnus Lindholm
                   ` (4 subsequent siblings)
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

check_mmu_context() clears asn_lock and acts on need_new_asn, but it runs
only as the tail of switch_to(), after alpha_switch_to() returns. A newly
forked task never gets there: its first context switch resumes at
ret_from_fork, which goes to schedule_tail() and then to user space rather
than returning to the code following alpha_switch_to(). A new kernel
thread reaches schedule_tail() the same way, through
ret_from_kernel_thread().

asn_lock is left set on that CPU, so the forked task runs user space with
it set and interrupts enabled. A TLB shootdown IPI arriving in that window
takes the deferred path, and the need_new_asn handshake meant to cover that
never runs.

finish_task_switch() calls finish_arch_post_lock_switch() with preemption
disabled, on the CPU that ran switch_mm(), so hooking check_mmu_context()
there completes the deferred-ASN bookkeeping for both. Running it when
switch_to() has already done the work is harmless: check_mmu_context()
clears need_new_asn as it goes.

The same hook is also called from kthread_use_mm() and
sched_force_init_mm(), outside the scheduler's preemption-disabled switch
tail. check_mmu_context() acts on per-CPU state, so it can only complete
this bookkeeping while still on the CPU that ran switch_mm(). Testing
preemptible() expresses that condition directly rather than naming callers:
where those paths leave preemption enabled the CPU may already have changed
and nothing is done, and where preemption is disabled across the switch, or
not configured at all, no migration is possible and running it is correct.

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/include/asm/mmu_context.h | 28 ++++++++++++++++++++++++++++
 1 file changed, 28 insertions(+)

diff --git a/arch/alpha/include/asm/mmu_context.h b/arch/alpha/include/asm/mmu_context.h
index eee8fe836a59..f7afc64ca8dd 100644
--- a/arch/alpha/include/asm/mmu_context.h
+++ b/arch/alpha/include/asm/mmu_context.h
@@ -181,6 +181,34 @@ do {								\
 #define check_mmu_context()  do { } while(0)
 #endif
 
+/*
+ * check_mmu_context() clears asn_lock and acts on need_new_asn, but it runs
+ * only as the tail of switch_to(), which a newly forked task never reaches:
+ * it resumes at ret_from_fork instead.  asn_lock is left set and the task
+ * goes on to run user space with it set and interrupts enabled, so a
+ * shootdown IPI arriving in that window takes the deferred path and the
+ * need_new_asn handshake meant to cover it never runs.
+ *
+ * finish_task_switch() calls this hook with preemption disabled on the CPU
+ * that ran switch_mm(), which covers that case.  Running it when switch_to()
+ * has already done the work is harmless: check_mmu_context() clears
+ * need_new_asn as it goes.
+ *
+ * kthread_use_mm() and sched_force_init_mm() also call this hook, outside
+ * the scheduler's preemption-disabled switch tail.  check_mmu_context() acts
+ * on per-CPU state, so it can only complete this bookkeeping while still on
+ * the CPU that ran switch_mm(); preemptible() tests that directly.  Where
+ * those paths leave preemption enabled the CPU may already have changed and
+ * nothing is done; where preemption is disabled across the switch, or not
+ * configured at all, no migration is possible and running it is correct.
+ */
+#define finish_arch_post_lock_switch finish_arch_post_lock_switch
+static inline void finish_arch_post_lock_switch(void)
+{
+	if (!preemptible())
+		check_mmu_context();
+}
+
 __EXTERN_INLINE void
 ev5_activate_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm)
 {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 2/6] alpha: only use a targeted tbi() when the target mm is really current
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 1/6] alpha: run check_mmu_context() from finish_arch_post_lock_switch() Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 3/6] alpha: fix the local TLB invalidate in flush_tlb_page() Magnus Lindholm
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

A copy-on-write fault replaces the page and calls ptep_clear_flush(),
which ends up in flush_tlb_page(). For a non-executable vma the remote
IPI handler issues a targeted tbi(2, addr).

tbi() acts on the address space context currently loaded on that CPU, so
it means something only when that context belongs to the target mm.
current->active_mm is the wrong test: under lazy TLB an idle or kernel
task keeps the mm as its active_mm while a different ASN is loaded in the
PCB - enter_lazy_tlb() updates only the borrowing task's page table base,
not its ASN.  The tbi() then invalidates the wrong context and the stale
translation survives. Nothing forces the old ASN to be retired
afterwards either, because mm->context[cpu] is still valid, so the
resuming thread can reuse it along with the stale entry.

Use current->mm instead. When no thread of the mm is current, fall back
to the deferred invalidation, which forces a fresh ASN at the next switch
and is correct whatever is loaded now.

This does not make current->mm a guarantee that the mm's context is
loaded: kthread_use_mm() sets current->mm and reaches switch_mm_irqs_off()
directly, which on alpha only prepares the incoming PCB. That is a
separate problem in the switch path rather than in this handler, and it
is not addressed here.

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/kernel/smp.c | 15 ++++++++++++++-
 1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index ed06367ece57..0dfe29b59039 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -669,7 +669,20 @@ ipi_flush_tlb_page(void *x)
 	struct flush_tlb_page_struct *data = x;
 	struct mm_struct * mm = data->mm;
 
-	if (mm == current->active_mm && !asn_locked())
+	/*
+	 * tbi() acts on the address space context currently loaded on this
+	 * CPU, so it reaches MM's translations only when a thread of MM is
+	 * really running here.  current->active_mm is not sufficient: under
+	 * lazy TLB an idle or kernel task keeps MM as its active_mm while a
+	 * different ASN is loaded in the PCB, so the tbi() invalidates the
+	 * wrong context and the stale entry survives.  Nothing forces the
+	 * old ASN to be retired afterwards either, mm->context[cpu] still
+	 * being valid, so the resuming thread can reuse it.
+	 *
+	 * Otherwise fall back to invalidating the context, which forces a
+	 * fresh ASN at the next switch whatever is loaded now.
+	 */
+	if (mm == current->mm && !asn_locked())
 		flush_tlb_current_page(mm, data->vma, data->addr);
 	else
 		flush_tlb_other(mm);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 3/6] alpha: fix the local TLB invalidate in flush_tlb_page()
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 1/6] alpha: run check_mmu_context() from finish_arch_post_lock_switch() Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 2/6] alpha: only use a targeted tbi() when the target mm is really current Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 4/6] alpha: only use a targeted tbi() when the target mm is really current (UP) Magnus Lindholm
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

flush_tlb_page() invalidates the calling CPU itself before asking the
others, and gates that on current->active_mm. For a non-executable vma
that means a targeted tbi(2, addr), which acts on the context currently
loaded, so as in the IPI handler it reaches nothing when only active_mm
names the mm, and nothing forces the old ASN to be retired afterwards.

Test current->mm instead. When no thread of the mm is current, clear
mm->context[cpu] so a fresh ASN is taken at the next switch. That also
covers the case where the mm is not this CPU's active_mm at all: the CPU
may still hold translations for it, and smp_call_function() does not call
back into the caller.

Reached in practice by folio_mkclean() from the writeback flusher
kworker, which has no mm of its own: about half the calls during
writeback of a shared mapping, and none at all on anonymous memory.

flush_tlb_mm() does not need the same active_mm to current->mm change,
because its active_mm path loads a new context rather than issuing a
targeted tbi(). It does have the separate caller-CPU omission when the
target mm is not active_mm; that is fixed in the following patch.

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/kernel/smp.c | 12 +++++++++++-
 1 file changed, 11 insertions(+), 1 deletion(-)

diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index 0dfe29b59039..d167b3ba1303 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -696,7 +696,15 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
 
 	preempt_disable();
 
-	if (mm == current->active_mm) {
+	/*
+	 * As in ipi_flush_tlb_page(): the targeted tbi() reaches MM's
+	 * translations only when a thread of MM is current, so test
+	 * current->mm.  Otherwise - lazily borrowing MM, or not running it
+	 * at all - clear mm->context[cpu] so a fresh ASN is taken at the
+	 * next switch.  smp_call_function() below does not call back into
+	 * this CPU, so this is the only chance to retire what it holds.
+	 */
+	if (mm == current->mm) {
 		flush_tlb_current_page(mm, vma, addr);
 		if (atomic_read(&mm->mm_users) <= 1) {
 			int cpu, this_cpu = smp_processor_id();
@@ -709,6 +717,8 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
 			preempt_enable();
 			return;
 		}
+	} else {
+		flush_tlb_other(mm);
 	}
 
 	data.vma = vma;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 4/6] alpha: only use a targeted tbi() when the target mm is really current (UP)
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
                   ` (2 preceding siblings ...)
  2026-08-09  8:49 ` [PATCH 3/6] alpha: fix the local TLB invalidate in flush_tlb_page() Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 5/6] alpha: invalidate the local context in flush_tlb_mm() Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 6/6] alpha: invalidate the local context in flush_icache_user_page() Magnus Lindholm
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

The uniprocessor flush_tlb_page() has the same defect the previous patch
fixed for SMP:

	if (mm == current->active_mm)
		flush_tlb_current_page(mm, vma, addr);
	else
		flush_tlb_other(mm);

For a non-executable vma flush_tlb_current_page() issues tbi(2, addr),
which acts on the address space context currently loaded, so it reaches
the mm's translations only when that context belongs to it. Under lazy
TLB an idle or kernel task keeps the mm as its active_mm while a different
ASN is loaded, so the tbi() invalidates the wrong context and the stale
translation survives.

Use current->mm instead, as for SMP.

This is not theoretical on a uniprocessor. folio_mkclean() runs in the
writeback flusher kworker, which borrows the mm, and with one CPU that
kworker necessarily shares it with the thread holding the translation.  A
test that writes a small MAP_SHARED file while background writeback cleans
it loses data on every round: the mapping holds one value and the file
another.

flush_tlb_mm() and flush_icache_user_page() need no equivalent change
here.  Both use __load_new_mm_context(), which allocates and loads a fresh
context rather than relying on a targeted tbi() against whatever ASN
happened to be loaded.

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/include/asm/tlbflush.h | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)

diff --git a/arch/alpha/include/asm/tlbflush.h b/arch/alpha/include/asm/tlbflush.h
index 0c8529997f54..6593a64090f1 100644
--- a/arch/alpha/include/asm/tlbflush.h
+++ b/arch/alpha/include/asm/tlbflush.h
@@ -87,7 +87,14 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
 {
 	struct mm_struct *mm = vma->vm_mm;
 
-	if (mm == current->active_mm)
+	/*
+	 * tbi() acts on the address space context currently loaded, so it
+	 * reaches MM's translations only when a thread of MM is current.
+	 * Under lazy TLB an idle or kernel task keeps MM as its active_mm
+	 * with a different ASN loaded, and a targeted tbi() would then
+	 * invalidate the wrong context.
+	 */
+	if (mm == current->mm)
 		flush_tlb_current_page(mm, vma, addr);
 	else
 		flush_tlb_other(mm);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 5/6] alpha: invalidate the local context in flush_tlb_mm()
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
                   ` (3 preceding siblings ...)
  2026-08-09  8:49 ` [PATCH 4/6] alpha: only use a targeted tbi() when the target mm is really current (UP) Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  2026-08-09  8:49 ` [PATCH 6/6] alpha: invalidate the local context in flush_icache_user_page() Magnus Lindholm
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

The SMP flush_tlb_mm() only touches the calling CPU when the mm is its
active_mm:

	if (mm == current->active_mm) {
		flush_tlb_current(mm);
		...
	}

	smp_call_function(ipi_flush_tlb_mm, mm, 1);

If it is not, nothing happens locally at all: smp_call_function() does not
call back into the caller. mm->context[cpu] is left valid, so this CPU
may later reuse the old ASN, and any translations it still holds for MM
stay usable.

Callers reach this regularly - counting the branch gave 2934 such calls
over a fork-heavy workload and 632 while otherwise idle.

Add the missing else. The UP implementation in asm/tlbflush.h already has
exactly this shape:

	if (mm == current->active_mm)
		flush_tlb_current(mm);
	else
		flush_tlb_other(mm);

Unlike flush_tlb_page(), the active_mm test itself is valid here:
flush_tlb_current() calls __load_new_mm_context(), which allocates and
loads a fresh context rather than relying on a targeted tbi() to operate
on whatever ASN happened to be loaded beforehand.

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/kernel/smp.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index d167b3ba1303..f501a91001cc 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -649,6 +649,13 @@ flush_tlb_mm(struct mm_struct *mm)
 			preempt_enable();
 			return;
 		}
+	} else {
+		/*
+		 * smp_call_function() below does not call back into this
+		 * CPU, so this is the only chance to retire what it holds
+		 * for MM.  The UP implementation already does this.
+		 */
+		flush_tlb_other(mm);
 	}
 
 	smp_call_function(ipi_flush_tlb_mm, mm, 1);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 6/6] alpha: invalidate the local context in flush_icache_user_page()
  2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
                   ` (4 preceding siblings ...)
  2026-08-09  8:49 ` [PATCH 5/6] alpha: invalidate the local context in flush_tlb_mm() Magnus Lindholm
@ 2026-08-09  8:49 ` Magnus Lindholm
  5 siblings, 0 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-08-09  8:49 UTC (permalink / raw)
  To: richard.henderson, mattst88, linux-kernel, linux-alpha
  Cc: glaubitz, mcree, ink, macro, Magnus Lindholm

flush_icache_user_page() has the same caller-CPU omission that the
previous patch fixed in flush_tlb_mm():

	if (mm == current->active_mm) {
		__load_new_mm_context(mm);
		...
	}

	smp_call_function(ipi_flush_icache_page, mm, 1);

When the target mm is not the calling CPU's active_mm nothing happens
locally, and smp_call_function() handles only the other CPUs, so this CPU
may later reuse the old ASN together with the translations it still holds.

This matters here in particular because the function exists for operating
on another process's mappings: the comment above it describes setting
breakpoints through ptrace, and access_remote_vm() reaches it through
copy_to_user_page(). The calling CPU is therefore often running something
other than the target mm.

As in flush_tlb_mm(), the UP implementation in asm/cacheflush.h already
has the missing case:

	if (current->active_mm == mm)
		__load_new_mm_context(mm);
	else
		mm->context[smp_processor_id()] = 0;

Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
 arch/alpha/kernel/smp.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index f501a91001cc..13b86f7224de 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -780,6 +780,13 @@ flush_icache_user_page(struct vm_area_struct *vma, struct page *page,
 			preempt_enable();
 			return;
 		}
+	} else {
+		/*
+		 * As in flush_tlb_mm(): smp_call_function() does not call
+		 * back into this CPU, and this function is used precisely
+		 * when operating on another process's mappings.
+		 */
+		flush_tlb_other(mm);
 	}
 
 	smp_call_function(ipi_flush_icache_page, mm, 1);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-09  8:55 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-09  8:49 [PATCH 0/6] alpha: fix stale TLB translations breaking copy-on-write and writeback Magnus Lindholm
2026-08-09  8:49 ` [PATCH 1/6] alpha: run check_mmu_context() from finish_arch_post_lock_switch() Magnus Lindholm
2026-08-09  8:49 ` [PATCH 2/6] alpha: only use a targeted tbi() when the target mm is really current Magnus Lindholm
2026-08-09  8:49 ` [PATCH 3/6] alpha: fix the local TLB invalidate in flush_tlb_page() Magnus Lindholm
2026-08-09  8:49 ` [PATCH 4/6] alpha: only use a targeted tbi() when the target mm is really current (UP) Magnus Lindholm
2026-08-09  8:49 ` [PATCH 5/6] alpha: invalidate the local context in flush_tlb_mm() Magnus Lindholm
2026-08-09  8:49 ` [PATCH 6/6] alpha: invalidate the local context in flush_icache_user_page() Magnus Lindholm

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.