BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops
@ 2026-10-04 13:37 Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info Eduard Zingerman
                   ` (44 more replies)
  0 siblings, 45 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

This series implements the scalar evolution (SCEV) technique for
verification of loops.

Scalar evolution is a static analysis technique that infers algebraic
expressions describing how variables change within a loop body.
These expressions can then be used to estimate the number of loop body
executions. This estimate can be used to represent induction variables
as ranges instead of enumerating each possible value.

For example, the following loop:

  for (r0 = 0; r0 < 10; r0++) {
    ...
  }

would now be verified with the assumption that r0 is a scalar value
in the range [0..9] within its body.

See the following patches for a more technical overview:
- "bpf: compute scalar evolution expressions for loops"
- "bpf: use SCEV to widen bounded loops"

The series can be viewed as consisting of the following parts:
- Preparatory patches extending liveness analysis to collect
  additional information and adding various utility functions.
- Patches relaxing the verifier's current restrictions on accessing
  memory via varying offset pointers, specifically:
  - "bpf: allow subrange relations for PTR_TO_STACK in regsafe()"
  - "bpf: representation for intervals with steps"
  - "bpf: varying offset access support for PTR_TO_BTF_ID pointers"
- Patches to compute the immediate dominator tree and loop hierarchy.
- Patches implementing scalar evolution and integrating it with
  the main verification pass:
  - "bpf: avoid widening registers that hinder exact stack-slot tracking"
  - "bpf: use SCEV to widen bounded loops"
  - "bpf: compute scalar evolution expressions for loops"
- Tests.

Literature
==========

- "Symbolic Evaluation of Chains of Recurrences for Loop Optimization"
  Robert A. van Engelen, 2000
- "The CR# Algebra and its Application in Loop Analysis and Optimization"
  Robert A. van Engelen, 2004

Supported patterns
==================

As scalar evolution analysis attempts to infer recurring relationships
from algebraic expressions describing loop variables, there are some
limitations on what can be expressed.

Below is a list of patterns that are expected to work with the current
implementation. The examples are pseudo-code for the indicated BPF
instruction shapes.

1. Counted loops, including decreasing counters:

       u64 i = 0;
       do { i += 1; } while (i < 10);

       u64 j = 3;
       do { j += -1; } while (j != 0);

   Header ranges are i in [0,9] and j in [1,3].
   Pre-condition loops, such as while (i < 10) { i += 1; },
   are also supported. The counter can reside in a register or a fixed
   8-byte frame-pointer spill.

2. Strided accesses into a map value:

       /* bytes is a map-value array with at least 16 bytes. */
       u64 i = 0, off = 0;
       do {
           bytes[off] = 1;
           i += 1;
           off += 2;
       } while (i < 8);

   At the header, off is in [0,14] with step 2. The byte store stays
   within the map value even though i and off are tracked separately.

3. A nonconstant entry value with a separate counted induction variable:

       u64 x = bpf_get_prandom_u32() & 6;  // {0,2,4,6}
       u64 i = 0;
       do { x += 6; i += 1; } while (i < 3);

   The counter i supplies the bound. At the header, x widens to [0,18]
   with step 2, retaining the alignment common to its entry values and
   its slope; x itself need not start at a constant.

4. Conditional assignment from an invariant value:

       u64 x = 5;
       for (u64 i = 0; i < 3; i += 1)
           if (i == 1) x = 10;
       return x;

   Widening joins the pre-loop value 5 with the assigned value 10,
   giving x in [5,10] both inside and after the loop.

5. Early exits in addition to the counted backedge:

       u64 i = 0;
       do {
           if (bpf_get_prandom_u32() == 0) break;
           i += 1;
       } while (i != 3);

   The counter still bounds the loop. An additional exit reduces the
   available lower-bound information; it does not by itself prevent
   widening.

6. Array accesses through different memory types:

   Bounded varying offsets are supported for arrays in BTF-typed
   objects, stack memory, map values and keys, packet data and
   metadata, and memory buffers. The usual bounds, alignment and
   access restrictions still apply.

   For stack arrays indexed by a widened loop counter, only byte and
   half-word accesses are supported; see limitation 5 below.

   For an array in a BTF-typed object:

       struct inner { int a, b; };
       struct object { struct inner arr[8]; };
       /* o points to a live, non-NULL BTF-typed struct object. */
       u64 i = bpf_get_prandom_u32() & 7;
       value = o->arr[i].b;

   The 4-byte load has offsets 4 + 8*i, i in [0,7]. Bounds keep
   accesses within arr; base and step 8 ensure that each possible
   offset selects member b. Such accesses also work without a loop.

7. Nested loops with separate counters:

       u64 i = 0;
       do {
           u64 j = 0;
           do {
               /* body */
               j += 1;
           } while (j < 3);
           i += 1;
       } while (i < 4);

   Both loops can be widened: i is in [0,3] at the outer header,
   and j is in [0,2] at the inner header. The inner counter is reset
   on each outer iteration; i is unchanged by the inner loop.

8. Incrementing and comparing a pointer directly:

       /* bytes is a map-value array with at least 8 bytes. */
       u8 *p = bytes, *end = bytes + 8;
       do {
           *p = 1;
           p++;
       } while (p != end);

   The initial pointer and fixed end pointer give eight iterations,
   without a separate integer counter. At the loop header, p ranges
   from bytes to bytes + 7, so each byte store stays within the array.

Limitations
===========

Patterns that lose precision after widening
-------------------------------------------

Widening can lose information that the verifier gets by checking
each iteration separately. For example:

    r7 = 5;
    for (r6 = 0; r6 < 3; r6++)
        if (r6 == 1)
            r7 = 10;
    if (r7 != 10)
        invalid_stack_read();

R7 is always 10 at the real exit, but is represented as [5, 10] at
loop entry and after the loop. The verifier therefore cannot exclude
the invalid read and rejects this otherwise safe program.
It would be hard to avoid this limitation.

The same issue affects correlated linear counters at loop exit:

    u64 i = 0, off = 0;
    while (i < 8) {
        i++;
        off += 2;
    }
    if (off != 16)
        invalid_stack_read();

Enumeration establishes off == 16. Widening tracks i and off
independently, and the exit check on i does not reconstruct off's exact
final value. This can be lifted after analysis adjustments.

Patterns not widened yet
------------------------

If one of the variables modified in the loop cannot be widened, the
verifier checks each iteration instead. Examples include:

1. 32-bit recurrences and comparisons:

       for (u32 i = 0; i < 8; i++)      // ALU32 / JMP32
           use(i);

   Only 64-bit arithmetic and comparisons are supported.
   This limitation will be lifted.

2. Non-linear updates:

       for (u64 i = 1; i < 16; i *= 2)  // non-linear recurrence
           use(i);

   Recognition is currently limited to simple additive recurrences.
   Literature describes a way to represent this, but it is not
   considered a priority at the moment.

3. A controlling counter or bound that is not constant at loop entry:

       u64 limit = unknown_in_range(1, 8);
       for (u64 i = 0; i < limit; i++)
           use(i);

       u64 i = unknown_in_range(0, 7);
       while (i < 8)
           i++;

   Both loops have finite bounds, but iteration-count evaluation
   currently requires concrete entry values. Other induction variables
   may have nonconstant entry ranges once a separate counter establishes
   the iteration bound. This limitation can be lifted eventually.

4. Alternative values referring to another evolving register:

       r7 = 0;
       for (r6 = 0; r6 < 4; r6++)
           if (condition)
               r7 = r6;
       use(r7);

   R6 changes on each iteration. Widening currently supports
   conditional assignments only from constants or values that do not
   change in the loop. This limitation can be lifted.

5. Stack addresses whose widening would lose precise slot tracking:

       u64 slots[4];
       for (u64 i = 0; i < 4; i++)
           slots[i] = i;

       struct bpf_dynptr dptrs[4];
       for (u64 i = 0; i < 4; i++)
           bpf_dynptr_from_xdp(ctx, 0, &dptrs[i]);

   Keep such indices concrete for 4- and 8-byte stack loads/stores and
   calls requiring fixed-offset stack objects. Dependencies inside
   nested loops also constrain the outer loop.
   It would be hard to lift this limitation.

6. Values modified by a nested loop do not get a closed-form summary
   for the enclosing loop:

       u64 sum = 0;
       for (u64 i = 0; i < 4; i++)
           for (u64 j = 0; j < 4; j++)
               sum++;
       use(sum);

   The inner loop may widen, but the outer analysis forgets values
   modified by it rather than deriving the combined recurrence.
   This limitation can be lifted.

7. Multiple backedges and irreducible control flow:

       i = 0;
   H:  if (i >= 8) goto done;
       if (condition) { i++; goto H; }  // first backedge
       i++; goto H;                    // second backedge
   done:

       i = 0;
       if (condition) goto body;       // second entry into the loop
   H:  i++;
   body:
       if (i < 8) goto H;

   Widening requires a reducible loop with one backedge and a supported
   dominating exit condition. Multiple exits are supported, but their
   early-exit paths reduce what can be inferred about iteration counts.
   It is unclear whether this can be lifted at the moment.

8. Some equivalent latch forms and large or wrapping recurrences:

       if (bound > i) goto again;      // i is the BPF source operand

       for (u64 i = 0, x = 0; i < 4; i++) {
           use(x);
           x += 65536;                // slope outside signed 16 bits
       }

   Latch matching currently expects the induction variable as the
   destination operand. Widened slopes must be nonzero signed-16-bit
   constants, computed ranges must fit signed 64-bit arithmetic,
   and iteration counts must be representable and finite.

9. Some hard-coded constant limits:
   - Programs with more than 16 nested loops in one function are rejected.
   - Loops with more than 256 exits exceed the metadata limits and are
     not analyzed for widening.
   - Expression traversal is limited to depth 8. Only a small fixed number
     of alternative values can be joined.

Veristat changes
================

A is master and B is this patch-set. The comparison for 7,301
programs: 134 from Cilium, 1,222 from Meta, 375 from sched_ext,
and 5,570 from BPF selftests.

Current impact on selftests is small. Work is in progress to improve
this, for example, reducing pyperf600_nounroll from 460,062 to 1,588
processed instructions. This follow-up work is not included below.

The histogram includes all programs in the corpus, including failed
loads and programs below the table cutoff. This series does not
introduce any new rejections.

Insns change   Programs
-------------  --------
-100 .. -95  %: 12
 -95 .. -90  %: 10
 -90 .. -85  %: 8
 -85 .. -80  %: 6
 -80 .. -75  %: 10
 -75 .. -65  %: 18
 -65 .. -60  %: 32
 -60 .. -55  %: 20
 -55 .. -50  %: 17
 -50 .. -45  %: 35
 -45 .. -40  %: 9
 -40 .. -35  %: 11
 -35 .. -30  %: 6
 -30 .. -25  %: 6
 -25 .. -20  %: 12
 -20 .. -10  %: 7
 -10 .. -5   %: 61
  -5 .. 0    %: 9
   0 .. 5    %: 7001
   5 .. 15   %: 2
  35 .. 40   %: 7
  40 .. 50   %: 1
  60 .. 65   %: 1

Tables: success in both runs; baseline Insns >= 5,000.
Ranked by percentage change, then instruction-count change, file, program.
Meta names are anonymized; numbered objects are distinct files.

Top wins (50)
-------------

File                                Program                              Insns (A)  Insns (B)       Insns (DIFF)
----------------------------------  -----------------------------------  ---------  ---------  -----------------
verifier_iterating_callbacks.bpf.o  test1                                     7004         12    -6992 (-99.83%)
bpf_iter_tasks.bpf.o                dump_task_sleepable                      85522        556   -84966 (-99.35%)
bpf_iter_task_stack.bpf.o           dump_task_stack                           7209         65    -7144 (-99.10%)
bpf_iter_unix.bpf.o                 dump_unix                                 9026        226    -8800 (-97.50%)
security-monitor-62.bpf.o           lsm_bprm_creds_for_exec                  17538       1701   -15837 (-90.30%)
security-monitor-62.bpf.o           lsm_task_alloc                            5674        565    -5109 (-90.04%)
security-monitor-62.bpf.o           lsm_fo_task_update                       11399       1181   -10218 (-89.64%)
security-monitor-62.bpf.o           lsm_file_open                           154412      21168  -133244 (-86.29%)
security-sandbox-17.bpf.o           ptrace_traceme                           16036       3134   -12902 (-80.46%)
security-sandbox-10.bpf.o           fork                                     30760       6211   -24549 (-79.81%)
security-sandbox-17.bpf.o           ptrace_access_check                      21206       5258   -15948 (-75.21%)
security-sandbox-12.bpf.o           task_kill                                27823       7688   -20135 (-72.37%)
security-monitor-22.bpf.o           syscalls_kill                            97920      27794   -70126 (-71.62%)
security-sandbox-13.bpf.o           kernel_module_request                     7175       2461    -4714 (-65.70%)
security-sandbox-13.bpf.o           kernel_load_data                          7176       2462    -4714 (-65.69%)
security-sandbox-13.bpf.o           kernel_read_file                          7177       2463    -4714 (-65.68%)
security-monitor-32.bpf.o           syscall_setresuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setfsgid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setfsuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setgid                         122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setregid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setreuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setuid                         122082      43187   -78895 (-64.62%)
security-monitor-12.bpf.o           connect_security_socket_connect         216137      77825  -138312 (-63.99%)
security-monitor-16.bpf.o           kernel_modules_do_init                   92863      33648   -59215 (-63.77%)
security-monitor-20.bpf.o           enter_pivot_root                         93400      34193   -59207 (-63.39%)
security-monitor-08.bpf.o           bpf_prog_detect                          93266      34376   -58890 (-63.14%)
security-monitor-31.bpf.o           inode_create                             96265      37055   -59210 (-61.51%)
security-monitor-10.bpf.o           cgroup_mkdir                             96513      37525   -58988 (-61.12%)
security-monitor-03.bpf.o           net_block_bind                           64652      25231   -39421 (-60.97%)
security-monitor-03.bpf.o           net_block_recvmsg                        64652      25231   -39421 (-60.97%)
security-monitor-03.bpf.o           net_block_sendmsg                        64652      25231   -39421 (-60.97%)
security-monitor-04.bpf.o           action_proc_term_sched_process_exec      64739      25299   -39440 (-60.92%)
security-monitor-13.bpf.o           sock_iter                                64818      25350   -39468 (-60.89%)
file-system-monitor-03.bpf.o        proc_file_open                           64883      25437   -39446 (-60.80%)
security-monitor-34.bpf.o           udp_recv_v6                              65089      25623   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_send_v6                              65089      25623   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_recv_v4                              65092      25626   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_send_v4                              65093      25627   -39466 (-60.63%)
security-monitor-18.bpf.o           memfd_create                             65291      25845   -39446 (-60.42%)
security-monitor-17.bpf.o           lsm_file_open                           130665      51773   -78892 (-60.38%)
security-monitor-07.bpf.o           security_socket_listen                   65378      25932   -39446 (-60.34%)
security-monitor-05.bpf.o           accept_security_socket_accept            65435      25989   -39446 (-60.28%)
security-monitor-21.bpf.o           raw_tracepoint__sched_process_exec       65511      26045   -39466 (-60.24%)
security-monitor-06.bpf.o           bash_reader                              65601      26161   -39440 (-60.12%)
security-monitor-30.bpf.o           python3_armor                            65805      26339   -39466 (-59.97%)
security-monitor-32.bpf.o           syscalls_setgroups                       67083      27617   -39466 (-58.83%)
security-monitor-09.bpf.o           fexit_cap_capable                        68221      28162   -40059 (-58.72%)
security-sandbox-31.bpf.o           test_file_open                           15279       6331    -8948 (-58.56%)
security-monitor-03.bpf.o           net_block_init                           71217      31453   -39764 (-55.83%)

Top losses (10)
---------------

File               Program  Insns (A)  Insns (B)    Insns (DIFF)
-----------------  -------  ---------  ---------  --------------
firewall-06.bpf.o  ingress       9323       9624   +301 (+3.23%)
firewall-06.bpf.o  tc_in         9323       9624   +301 (+3.23%)
firewall-08.bpf.o  ingress       9323       9624   +301 (+3.23%)
firewall-08.bpf.o  tc_in         9323       9624   +301 (+3.23%)
firewall-06.bpf.o  egress        9334       9635   +301 (+3.22%)
firewall-08.bpf.o  egress        9334       9635   +301 (+3.22%)
firewall-06.bpf.o  tc_eg         9488       9789   +301 (+3.17%)
firewall-08.bpf.o  tc_eg         9488       9789   +301 (+3.17%)
firewall-05.bpf.o  tc_eg       156665     161269  +4604 (+2.94%)
firewall-07.bpf.o  tc_eg       156665     161269  +4604 (+2.94%)

Changelog
=========

v1 -> v2:
TLDR: a bunch of major correctness fixes.
- Introduce arg_access_info to distinguish read, may-write and
  must-write effects of helper/kfunc stack accesses. Handle map-specific
  argument semantics and writes through unknown callbacks.
- Fix base/step tracking across signed overflow, truncation,
  sign/zero extension and linked-register updates.
- Rework iteration-count calculation for signed/unsigned increasing
  and decreasing counters, including counter overflow checks.
  Evaluate the full latch base expression and consistently count
  header executions.
- Substitute all registers in instantiate_header_scevs().
- Track pointer origins when evaluating latch expressions.
  Require compatible operands before deriving iteration bounds.
- Preserve the correct congruence base when widening nonconstant
  entry values, assign fresh IDs to widened packet pointers,
  and exclude nullable pointers from widening.
- Account for values modified on nested-loop exit paths when computing
  invariants. Do not use an inner loop's latch to bound an outer loop.
- Restrict pruning against in-progress SCEV states to terminating
  innermost loops sharing the same entry checkpoint, preventing
  inner-loop pruning from hiding a nonterminating outer loop.
- Rework loop entry/exit handling on the main verification pass.
- Extend fixed-stack-offset widening suppression to ANY expressions
  and all dynptr arguments described by helper/kfunc prototypes.
- Fix BTF array handling: retain bounds checks for fixed arrays inside
  flexible-array elements and preserve array type IDs when element
  types contain typedefs.
- Fix SCEV handling of endian conversions and narrow fills on
  big-endian systems.
- Grow expression storage geometrically and scale the expression hash
  table with program size. Add cancellation/rescheduling checks to
  dominator computation and fix the postorder-number allocation leak.
- Pack bpf_reg_state/bpf_func_state for latter to fit back into
  the 1 KiB allocation bucket, measured SCEV memory consumption
  overhead is 3% now.
- Major additions in the selftests.

v1: https://lore.kernel.org/bpf/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com/
---
Eduard Zingerman (43):
      bpf: represent stack access effects with arg_access_info
      bpf: describe helper stack accesses with arg_access_info
      bpf: describe kfunc stack accesses with arg_access_info
      bpf: track may_write flags in liveness
      bpf: summarize may write stack slots in insn_aux_data
      bpf: summarize live stack slots in insn_aux_data
      bpf: summarize regs that may hold a frame pointer in insn_aux_data
      bpf: record write effects for atomic operations in liveness.c
      bpf: add tnum_alignment()
      bpf: add cnum{32,64}_union()
      bpf: add cnum64_intersect_linear()
      bpf: expose comparison opcode transformations
      bpf: allow subrange relations for PTR_TO_STACK in regsafe()
      bpf: representation for intervals with steps
      bpf: varying offset access support for PTR_TO_BTF_ID pointers
      bpf: save DFS postorder numbers for program instructions
      bpf: move the live-register and SCC printout to a standalone function
      bpf: compute immediate dominators
      bpf: compute loop hierarchy
      bpf: add bpf_set_reg_range()
      bpf: add bpf_mark_reg_known_scalar()
      bpf: add bpf_reg_union()
      bpf: add a min-heap for ordered analysis worklists
      bpf: record basic-block ends in insn_aux_data
      bpf: add bpf_split_cur_state()
      bpf: add bpf_same_memory_origin()
      bpf: allow precision backtracking between overlapping checkpoints
      bpf: compute scalar evolution expressions for loops
      bpf: use SCEV to widen bounded loops
      bpf: avoid widening registers that hinder exact stack-slot tracking
      bpf: bpf_func_state size optimization
      selftests/bpf: __msg_next tag for matching messages on consecutive lines
      selftests/bpf: bound UNIX socket path loops by sun_path size
      selftests/bpf: test for stack-pointer subrange pruning
      selftests/bpf: tests for may_write stack-liveness tracking
      selftests/bpf: tests for may_def marks of atomic RMW operations
      selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes()
      selftests/bpf: tests for register base/step arithmetic
      selftests/bpf: tests for register base/step state pruning
      selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
      selftests/bpf: tests for loop hierarchy computation
      selftests/bpf: tests for immediate dominator computation
      selftests/bpf: tests for SCEV analysis and loop widening

 include/linux/bpf_verifier.h                       |  185 +-
 include/linux/cnum.h                               |    3 +
 include/linux/tnum.h                               |    2 +
 kernel/bpf/Makefile                                |    2 +-
 kernel/bpf/backtrack.c                             |    5 +
 kernel/bpf/btf.c                                   |  159 +-
 kernel/bpf/cfg.c                                   |   56 +-
 kernel/bpf/cnum.c                                  |   36 +
 kernel/bpf/cnum_defs.h                             |   44 +
 kernel/bpf/const_fold.c                            |    2 +
 kernel/bpf/fixups.c                                |   18 +
 kernel/bpf/heap.c                                  |   87 +
 kernel/bpf/liveness.c                              |  372 ++-
 kernel/bpf/log.c                                   |    7 +
 kernel/bpf/loops.c                                 |  606 +++++
 kernel/bpf/ringbuf.c                               |    4 +-
 kernel/bpf/scev.c                                  | 2507 ++++++++++++++++++++
 kernel/bpf/states.c                                |  216 +-
 kernel/bpf/tnum.c                                  |   11 +
 kernel/bpf/verifier.c                              |  766 +++++-
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   10 +
 tools/testing/selftests/bpf/progs/bpf_iter_unix.c  |    2 +-
 tools/testing/selftests/bpf/progs/bpf_misc.h       |    6 +
 .../selftests/bpf/progs/test_skc_to_unix_sock.c    |    2 +-
 .../selftests/bpf/progs/verifier_bounds_step.c     |  473 ++++
 .../bpf/progs/verifier_btf_array_access.c          |  576 +++++
 tools/testing/selftests/bpf/progs/verifier_gotox.c |    6 +-
 tools/testing/selftests/bpf/progs/verifier_idoms.c |  390 +++
 .../selftests/bpf/progs/verifier_kfunc_uninit.c    |    6 +
 .../selftests/bpf/progs/verifier_live_stack.c      |  322 ++-
 .../selftests/bpf/progs/verifier_loop_hierarchy.c  |  346 +++
 tools/testing/selftests/bpf/progs/verifier_scev.c  | 2376 +++++++++++++++++++
 .../selftests/bpf/progs/verifier_stack_ptr.c       |   36 +
 tools/testing/selftests/bpf/test_loader.c          |   10 +
 34 files changed, 9305 insertions(+), 344 deletions(-)
---
base-commit: 99dc1ba542420db6b8df209744f55cc52466ad91
change-id: 20260925-scev-minimal-rebase-911799c95dff

^ permalink raw reply	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:37 ` [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses " Eduard Zingerman
                   ` (43 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

The following patches in the series add support for may_write property
inference. To properly track this property there is a need to
distinguish three ways the helpers/kfuncs affect stack:
- read;
- write and destroy all previous stack tracking information for
  affected slots;
- both read and write the stack slots, potentially preserving some of
  the tracking information.

The later property can't be encoded by signed access_byte convention.
Hence, introduce arg_access_info with a size and separate may_read,
may_write and must_write bits. Convert the liveness access recording
paths to use it. Proper may_write calculation and support will be
added in the following patches.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  8 ++++
 kernel/bpf/liveness.c        | 96 ++++++++++++++++++++++++++++----------------
 2 files changed, 69 insertions(+), 35 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c51083c761cf..d4f32acba236 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1723,6 +1723,14 @@ bool bpf_is_may_goto_insn(struct bpf_insn *insn);
 void bpf_verbose_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
 bool bpf_get_call_summary(struct bpf_verifier_env *env, struct bpf_insn *call,
 			  struct bpf_call_summary *cs);
+/* Stack effects used for state pruning; must_write implies may_write. */
+struct arg_access_info {
+	u32 size;		/* Maximum extent; U32_MAX if unknown. */
+	u8 may_read:1;		/* Incoming contents or initialization may be needed. */
+	u8 may_write:1;
+	u8 must_write:1;	/* Prior verifier state is destroyed throughout size. */
+};
+
 s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env,
 				  struct bpf_insn *insn, int arg,
 				  int insn_idx);
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..7b317fb5ca82 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1437,17 +1437,14 @@ static void arg_track_xfer(struct bpf_verifier_env *env, struct bpf_insn *insn,
 }
 
 /*
- * Record access_bytes from helper/kfunc or load/store insn.
- *   access_bytes > 0:      stack read
- *   access_bytes < 0:      stack write
- *   access_bytes == S64_MIN: unknown   — conservative, mark [0..slot] as read
- *   access_bytes == 0:      no access
- *
+ * Record reads for every touched half-slot. A definite
+ * write requires full coverage, a known size, and a single possible offset.
  */
 static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
-				   s64 access_bytes, u32 frame, u32 insn_idx)
+				   struct arg_access_info info, u32 frame, u32 insn_idx)
 {
 	s32 slot_hi, slot_lo;
+	int err;
 
 	if (fp_off >= 0)
 		/*
@@ -1456,49 +1453,53 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
 		 * by the main verifier pass later.
 		 */
 		return 0;
-	if (access_bytes == S64_MIN) {
+
+	if (info.may_read && info.size == U32_MAX) {
 		/* helper/kfunc read unknown amount of bytes from fp_off until fp+0 */
 		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
 		return mark_stack_read(instance, frame, insn_idx, 0, slot_hi);
 	}
-	if (access_bytes > 0) {
+	if (info.may_read) {
 		/* Mark any touched slot as use */
 		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
-		slot_lo = max_t(s32, (-fp_off - access_bytes) / STACK_SLOT_SZ, 0);
-		return mark_stack_read(instance, frame, insn_idx, slot_lo, slot_hi);
-	} else if (access_bytes < 0) {
+		slot_lo = max_t(s32, (-fp_off - info.size) / STACK_SLOT_SZ, 0);
+		err = mark_stack_read(instance, frame, insn_idx, slot_lo, slot_hi);
+		if (err)
+			return err;
+	}
+	if (info.must_write && info.size != U32_MAX) {
 		/* Mark only fully covered slots as def */
-		access_bytes = -access_bytes;
 		slot_hi = (-fp_off) / STACK_SLOT_SZ - 1;
-		slot_lo = max_t(s32, (-fp_off - access_bytes + STACK_SLOT_SZ - 1) / STACK_SLOT_SZ, 0);
+		slot_lo = max_t(s32, (-fp_off - info.size + STACK_SLOT_SZ - 1) / STACK_SLOT_SZ, 0);
 		return mark_stack_write(instance, frame, insn_idx, slot_lo, slot_hi);
 	}
 	return 0;
 }
 
-/*
- * 'arg' is FP-derived argument to helper/kfunc or load/store that
- * reads (positive) or writes (negative) 'access_bytes' into 'use' or 'def'.
- */
-static int record_stack_access(struct bpf_verifier_env *env, struct func_instance *instance,
+/* Record access through a pointer with a known frame and possibly known offsets. */
+static int record_stack_access(struct bpf_verifier_env *env,
+			       struct func_instance *instance,
 			       const struct arg_track *arg,
-			       s64 access_bytes, u32 frame, u32 insn_idx)
+			       struct arg_access_info info, u32 frame, u32 insn_idx)
 {
 	int i, err;
 
-	if (access_bytes == 0)
+	if (!info.size)
 		return 0;
 	if (arg->off_cnt == 0) {
-		if (access_bytes > 0 || access_bytes == S64_MIN)
-			return mark_stack_read_all(env, instance, frame, insn_idx);
+		if (info.may_read) {
+			err = mark_stack_read_all(env, instance, frame, insn_idx);
+			if (err)
+				return err;
+		}
 		return 0;
 	}
-	if (access_bytes != S64_MIN && access_bytes < 0 && arg->off_cnt != 1)
+	if (info.size != U32_MAX && info.must_write && arg->off_cnt != 1)
 		/* multi-offset write cannot set stack_def */
 		return 0;
 
 	for (i = 0; i < arg->off_cnt; i++) {
-		err = record_stack_access_off(instance, arg->off[i], access_bytes, frame, insn_idx);
+		err = record_stack_access_off(instance, arg->off[i], info, frame, insn_idx);
 		if (err)
 			return err;
 	}
@@ -1533,9 +1534,10 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 				    struct arg_track *at, int insn_idx)
 {
 	struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+	struct arg_access_info info = {};
 	int depth = instance->depth;
-	s32 sz = bpf_size_to_bytes(BPF_SIZE(insn->code));
 	u8 class = BPF_CLASS(insn->code);
+	bool read = false, write = false;
 	struct arg_track resolved, *ptr;
 	int oi;
 
@@ -1552,23 +1554,26 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 	switch (class) {
 	case BPF_LDX:
 		ptr = &at[insn->src_reg];
+		read = true;
 		break;
 	case BPF_STX:
 		if (BPF_MODE(insn->code) == BPF_ATOMIC) {
 			if (insn->imm == BPF_STORE_REL)
-				sz = -sz;
+				write = true;
+			else
+				read = true;
 			if (insn->imm == BPF_LOAD_ACQ)
 				ptr = &at[insn->src_reg];
 			else
 				ptr = &at[insn->dst_reg];
 		} else {
 			ptr = &at[insn->dst_reg];
-			sz = -sz;
+			write = true;
 		}
 		break;
 	case BPF_ST:
 		ptr = &at[insn->dst_reg];
-		sz = -sz;
+		write = true;
 		break;
 	default:
 		return 0;
@@ -1587,32 +1592,53 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 		ptr = &resolved;
 	}
 
+	info.may_read = read;
+	info.may_write = write;
+	info.must_write = write;
+	info.size = bpf_size_to_bytes(BPF_SIZE(insn->code));
 	if (ptr->frame >= 0 && ptr->frame <= depth)
-		return record_stack_access(env, instance, ptr, sz, ptr->frame, insn_idx);
+		return record_stack_access(env, instance, ptr, info, ptr->frame, insn_idx);
 	if (ptr->frame == ARG_IMPRECISE)
 		return record_imprecise(env, instance, ptr->mask, insn_idx);
 	/* ARG_NONE: not derived from any frame pointer, skip */
 	return 0;
 }
 
+/* Adapt the signed helper/kfunc access size until both producers are converted. */
+static struct arg_access_info stack_access_info(s64 bytes)
+{
+	bool write = bytes < 0 && bytes != S64_MIN && -bytes < U32_MAX;
+
+	return (struct arg_access_info) {
+		.size = bytes == S64_MIN
+			? U32_MAX
+			: min_t(u64, bytes < 0 ? -bytes : bytes, U32_MAX),
+		.may_read = bytes > 0 || bytes == S64_MIN,
+		.may_write = write,
+		.must_write = write,
+	};
+}
+
 static int record_arg_access(struct bpf_verifier_env *env,
 			     struct func_instance *instance,
 			     struct bpf_insn *insn,
 			     struct arg_track *at, int arg_idx,
 			     int insn_idx)
 {
+	struct arg_access_info info;
 	int depth = instance->depth;
 	int frame = at->frame;
 	int err = 0;
-	s64 bytes;
 
 	if (!arg_is_fp(at))
 		return 0;
 
 	if (bpf_helper_call(insn)) {
-		bytes = bpf_helper_stack_access_bytes(env, insn, arg_idx, insn_idx);
+		info = stack_access_info(bpf_helper_stack_access_bytes(env, insn,
+								       arg_idx, insn_idx));
 	} else if (bpf_pseudo_kfunc_call(insn)) {
-		bytes = bpf_kfunc_stack_access_bytes(env, insn, arg_idx, insn_idx);
+		info = stack_access_info(bpf_kfunc_stack_access_bytes(env, insn,
+								      arg_idx, insn_idx));
 	} else {
 		for (int f = 0; f <= depth; f++) {
 			err = mark_stack_read_all(env, instance, f, insn_idx);
@@ -1621,11 +1647,11 @@ static int record_arg_access(struct bpf_verifier_env *env,
 		}
 		return 0;
 	}
-	if (bytes == 0)
+	if (!info.size)
 		return 0;
 
 	if (frame >= 0 && frame <= depth)
-		err = record_stack_access(env, instance, at, bytes, frame, insn_idx);
+		err = record_stack_access(env, instance, at, info, frame, insn_idx);
 	else if (frame == ARG_IMPRECISE)
 		err = record_imprecise(env, instance, at->mask, insn_idx);
 	return err;

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses with arg_access_info
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 14:02   ` sashiko-bot
  2026-10-04 13:37 ` [PATCH bpf-next v2 03/43] bpf: describe kfunc " Eduard Zingerman
                   ` (42 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Adjust bpf_helper_stack_access_bytes() to return results as a struct
arg_access_info. The result is based on bpf_func_proto->argX_type
value as in the following table:

| Argument flags       | allow_uninit_stack | Meaning              |
|----------------------+--------------------+----------------------|
| MEM_WRITE            | true               | may_read, must_write |
| MEM_WRITE+MEM_UNINIT | true               | must_write           |
| MEM_WRITE            | false              | may_read, may_write  |
| MEM_WRITE+MEM_UNINIT | false              | may_read, may_write  |
| 0                    | true/false         | may_read             |

MEM_WRITE now produces must_write when allow_uninit_stack is true:
the verifier discards previous stack knowledge over the affected
range. Previously, liveness only treated MEM_UNINIT this way.

Also use resolve_map_arg_type() to account for bpf_map_peek_elem()
treating its buffer as an input for bloom filters.

Annotate ringbuf dynptr submit and discard arguments with MEM_WRITE.
Their existing OBJ_RELEASE handling already invalidates the dynptr, but
the prototypes need to describe the write as well as the release.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  6 ++--
 kernel/bpf/liveness.c        |  5 ++-
 kernel/bpf/ringbuf.c         |  4 +--
 kernel/bpf/verifier.c        | 81 +++++++++++++++++++++++---------------------
 4 files changed, 49 insertions(+), 47 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d4f32acba236..5c99c024c309 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1731,9 +1731,9 @@ struct arg_access_info {
 	u8 must_write:1;	/* Prior verifier state is destroyed throughout size. */
 };
 
-s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env,
-				  struct bpf_insn *insn, int arg,
-				  int insn_idx);
+struct arg_access_info
+bpf_helper_stack_access_bytes(struct bpf_verifier_env *env,
+			      struct bpf_insn *insn, int arg, int insn_idx);
 s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 				 struct bpf_insn *insn, int arg,
 				 int insn_idx);
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 7b317fb5ca82..353b52031577 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1604,7 +1604,7 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 	return 0;
 }
 
-/* Adapt the signed helper/kfunc access size until both producers are converted. */
+/* Adapt the signed kfunc access size until its producer is converted. */
 static struct arg_access_info stack_access_info(s64 bytes)
 {
 	bool write = bytes < 0 && bytes != S64_MIN && -bytes < U32_MAX;
@@ -1634,8 +1634,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
 		return 0;
 
 	if (bpf_helper_call(insn)) {
-		info = stack_access_info(bpf_helper_stack_access_bytes(env, insn,
-								       arg_idx, insn_idx));
+		info = bpf_helper_stack_access_bytes(env, insn, arg_idx, insn_idx);
 	} else if (bpf_pseudo_kfunc_call(insn)) {
 		info = stack_access_info(bpf_kfunc_stack_access_bytes(env, insn,
 								      arg_idx, insn_idx));
diff --git a/kernel/bpf/ringbuf.c b/kernel/bpf/ringbuf.c
index 3f1013d80544..447863ece26a 100644
--- a/kernel/bpf/ringbuf.c
+++ b/kernel/bpf/ringbuf.c
@@ -722,7 +722,7 @@ BPF_CALL_2(bpf_ringbuf_submit_dynptr, struct bpf_dynptr_kern *, ptr, u64, flags)
 const struct bpf_func_proto bpf_ringbuf_submit_dynptr_proto = {
 	.func		= bpf_ringbuf_submit_dynptr,
 	.ret_type	= RET_VOID,
-	.arg1_type	= ARG_PTR_TO_DYNPTR | DYNPTR_TYPE_RINGBUF | OBJ_RELEASE,
+	.arg1_type	= ARG_PTR_TO_DYNPTR | DYNPTR_TYPE_RINGBUF | OBJ_RELEASE | MEM_WRITE,
 	.arg2_type	= ARG_ANYTHING,
 };
 
@@ -741,7 +741,7 @@ BPF_CALL_2(bpf_ringbuf_discard_dynptr, struct bpf_dynptr_kern *, ptr, u64, flags
 const struct bpf_func_proto bpf_ringbuf_discard_dynptr_proto = {
 	.func		= bpf_ringbuf_discard_dynptr,
 	.ret_type	= RET_VOID,
-	.arg1_type	= ARG_PTR_TO_DYNPTR | DYNPTR_TYPE_RINGBUF | OBJ_RELEASE,
+	.arg1_type	= ARG_PTR_TO_DYNPTR | DYNPTR_TYPE_RINGBUF | OBJ_RELEASE | MEM_WRITE,
 	.arg2_type	= ARG_ANYTHING,
 };
 
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index fd3c0206bd67..b558399bdc08 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14344,35 +14344,32 @@ int bpf_fetch_kfunc_arg_meta(struct bpf_verifier_env *env,
 }
 
 /*
- * Determine how many bytes a helper accesses through a stack pointer at
- * argument position @arg (0-based, corresponding to R1-R5).
- *
- * Returns:
- *   > 0   known read access size in bytes
- *     0   doesn't read anything directly
- * S64_MIN unknown
- *   < 0   known write access of (-return) bytes
+ * Describe a helper's stack access through argument @arg (0-based, R1-R5).
+ * must_write means the verifier destroys prior state throughout size bytes.
+ * In other words, must_write follows the verifier's model of the
+ * function's behaviour, which is safe to use for is_state_visited() pruning.
  */
-s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *insn,
-				  int arg, int insn_idx)
+struct arg_access_info
+bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *insn,
+			      int arg, int insn_idx)
 {
 	struct bpf_insn_aux_data *aux = &env->insn_aux_data[insn_idx];
+	struct arg_access_info info = {
+		.size = U32_MAX,
+		.may_read = true,
+		.may_write = true,
+	};
 	const struct bpf_func_proto *fn;
+	enum bpf_access_type access_type;
 	enum bpf_arg_type at;
-	bool full_write;
-	s64 size;
+	bool exact_size = true;
+	u64 size = U32_MAX;
 
 	if (bpf_get_helper_proto(env, insn->imm, &fn) < 0)
-		return S64_MIN;
+		return info;
 
 	at = fn->arg_type[arg];
-	/*
-	 * Generic outputs may leave bytes untouched. Keep prior initialization
-	 * live when the caller cannot read uninitialized bytes. Constructors of
-	 * special objects, such as dynptrs, still define their storage.
-	 */
-	full_write = (at & MEM_UNINIT) &&
-		     (!arg_type_is_raw_mem(at) || env->allow_uninit_stack);
+	access_type = func_arg_access_type(at);
 
 	switch (base_type(at)) {
 	case ARG_PTR_TO_MAP_KEY:
@@ -14395,6 +14392,10 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 
 		i = aux->const_reg_vals[map_reg];
 		if (i < env->used_map_cnt) {
+			/* Bloom-filter peek reads the value buffer. */
+			if (!is_key && insn->imm == BPF_FUNC_map_peek_elem &&
+			    env->used_maps[i]->map_type == BPF_MAP_TYPE_BLOOM_FILTER)
+				access_type = BPF_READ;
 			size = is_key ? env->used_maps[i]->key_size
 				      : env->used_maps[i]->value_size;
 			goto out;
@@ -14404,7 +14405,12 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 		 * Map pointer is not known at this call site (e.g. different
 		 * maps on merged paths).  Conservatively return the largest
 		 * key_size or value_size across all maps used by the program.
+		 * This is only an upper bound, so it cannot establish a definite
+		 * write. Map-dependent argument types can also turn an output
+		 * into an input, so conservatively retain a read dependency.
 		 */
+		exact_size = false;
+		access_type |= BPF_READ;
 		val = 0;
 		for (i = 0; i < env->used_map_cnt; i++) {
 			struct bpf_map *map = env->used_maps[i];
@@ -14420,7 +14426,7 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 			}
 		}
 		if (!val)
-			return S64_MIN;
+			return info;
 		size = val;
 		goto out;
 	}
@@ -14434,18 +14440,12 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 			int size_reg = BPF_REG_1 + arg + 1;
 
 			if (aux->const_reg_mask & BIT(size_reg)) {
-				size = (s64)aux->const_reg_vals[size_reg];
+				size = aux->const_reg_vals[size_reg];
 				goto out;
 			}
-			/*
-			 * Size arg is const on each path but differs across merged
-			 * paths. Reads may extend anywhere up to the frame top.
-			 */
-			if (full_write)
-				return 0;
-			return S64_MIN;
 		}
-		return S64_MIN;
+		/* Preserve access directions even when the extent is unknown. */
+		goto out;
 	case ARG_PTR_TO_DYNPTR:
 		size = BPF_DYNPTR_SIZE;
 		break;
@@ -14455,18 +14455,21 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 		 * doesn't access stack. The callback subprog does and it's
 		 * analyzed separately.
 		 */
-		return 0;
+		return (struct arg_access_info) {};
 	default:
-		return S64_MIN;
+		return info;
 	}
 out:
-	/*
-	 * Other accesses keep the previous state live, including untouched bytes
-	 * of an unprivileged generic output.
-	 */
-	if (full_write)
-		return -size;
-	return size;
+	info.size = min_t(u64, size, U32_MAX);
+	info.may_read = !!(access_type & BPF_READ);
+	info.may_write = !!(access_type & BPF_WRITE);
+	info.must_write = info.may_write && exact_size && info.size != U32_MAX;
+	/* Generic unprivileged outputs retain prior initialization state. */
+	if (!env->allow_uninit_stack && arg_type_is_raw_mem(at)) {
+		info.may_read = true;
+		info.must_write = false;
+	}
+	return info;
 }
 
 /*

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 03/43] bpf: describe kfunc stack accesses with arg_access_info
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses " Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 04/43] bpf: track may_write flags in liveness Eduard Zingerman
                   ` (41 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Return kfunc stack effects through arg_access_info while preserving the
existing BTF size lookup and read/def classification. Retain __sz/__szk
handling and conservatively mark pointer arguments as possible writes.

Remove the temporary signed-size adapter from liveness. Add liveness
assertions for privileged and unprivileged __uninit outputs.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h                       |  6 +--
 kernel/bpf/liveness.c                              | 18 +--------
 kernel/bpf/verifier.c                              | 43 ++++++++++++----------
 .../selftests/bpf/progs/verifier_kfunc_uninit.c    |  6 +++
 4 files changed, 33 insertions(+), 40 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 5c99c024c309..eafe79929dfa 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1734,9 +1734,9 @@ struct arg_access_info {
 struct arg_access_info
 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env,
 			      struct bpf_insn *insn, int arg, int insn_idx);
-s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
-				 struct bpf_insn *insn, int arg,
-				 int insn_idx);
+struct arg_access_info
+bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
+			     struct bpf_insn *insn, int arg, int insn_idx);
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
 
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 353b52031577..d4089a488784 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1604,21 +1604,6 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 	return 0;
 }
 
-/* Adapt the signed kfunc access size until its producer is converted. */
-static struct arg_access_info stack_access_info(s64 bytes)
-{
-	bool write = bytes < 0 && bytes != S64_MIN && -bytes < U32_MAX;
-
-	return (struct arg_access_info) {
-		.size = bytes == S64_MIN
-			? U32_MAX
-			: min_t(u64, bytes < 0 ? -bytes : bytes, U32_MAX),
-		.may_read = bytes > 0 || bytes == S64_MIN,
-		.may_write = write,
-		.must_write = write,
-	};
-}
-
 static int record_arg_access(struct bpf_verifier_env *env,
 			     struct func_instance *instance,
 			     struct bpf_insn *insn,
@@ -1636,8 +1621,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
 	if (bpf_helper_call(insn)) {
 		info = bpf_helper_stack_access_bytes(env, insn, arg_idx, insn_idx);
 	} else if (bpf_pseudo_kfunc_call(insn)) {
-		info = stack_access_info(bpf_kfunc_stack_access_bytes(env, insn,
-								      arg_idx, insn_idx));
+		info = bpf_kfunc_stack_access_bytes(env, insn, arg_idx, insn_idx);
 	} else {
 		for (int f = 0; f <= depth; f++) {
 			err = mark_stack_read_all(env, instance, f, insn_idx);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index b558399bdc08..98e64a138921 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14473,28 +14473,29 @@ bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *ins
 }
 
 /*
- * Determine how many bytes a kfunc accesses through a stack pointer at
- * argument position @arg (0-based, corresponding to R1-R5).
- *
- * Returns:
- *   > 0      known read access size in bytes
- *     0      doesn't access memory through that argument (ex: not a pointer)
- *   S64_MIN  unknown
- *   < 0      known write access of (-return) bytes
+ * Describe a kfunc's stack access through argument slot @arg (0-based).
+ * As for helpers, must_write describes destruction of prior verifier state,
+ * rather than a guarantee that the callee writes every byte at runtime.
  */
-s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *insn,
-				 int arg, int insn_idx)
+struct arg_access_info
+bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *insn,
+			     int arg, int insn_idx)
 {
 	struct bpf_insn_aux_data *aux = &env->insn_aux_data[insn_idx];
+	struct arg_access_info info = {
+		.size = U32_MAX,
+		.may_read = true,
+		.may_write = true,
+	};
 	struct bpf_call_arg_meta meta;
 	const struct btf_param *args;
 	const struct btf_type *t, *ref_t;
 	const struct btf *btf;
 	u32 i, slot, nargs, type_size;
-	s64 size;
+	u64 size;
 
 	if (bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, &meta) < 0)
-		return S64_MIN;
+		return info;
 
 	btf = meta.btf;
 	args = btf_params(meta.func_proto);
@@ -14509,11 +14510,11 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
 	for (i = 0, slot = 0; i < nargs && slot < arg; i++)
 		slot += btf_arg_slots(btf_type_skip_modifiers(btf, args[i].type, NULL));
 	if (i >= nargs || slot != arg)
-		return 0;
+		return (struct arg_access_info) {};
 
 	t = btf_type_skip_modifiers(btf, args[i].type, NULL);
 	if (!btf_type_is_ptr(t))
-		return 0;
+		return (struct arg_access_info) {};
 
 	/* dynptr: fixed 16-byte on-stack representation */
 	if (is_kfunc_arg_dynptr(btf, &args[i])) {
@@ -14529,11 +14530,11 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
 
 		if (size_reg <= MAX_BPF_FUNC_REG_ARGS &&
 		    (aux->const_reg_mask & BIT(size_reg))) {
-			size = (s64)aux->const_reg_vals[size_reg];
+			size = aux->const_reg_vals[size_reg];
 			goto out;
 		}
 		/* Unknown size: the read may extend anywhere up to the frame top. */
-		return S64_MIN;
+		return info;
 	}
 
 	/* fixed-size pointed-to type: resolve via BTF */
@@ -14543,15 +14544,17 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
 		goto out;
 	}
 
-	return S64_MIN;
+	return info;
 out:
+	info.size = min_t(u64, size, U32_MAX);
 	/* KF_ITER_NEW kfuncs initialize the iterator state at arg 0 */
 	if (arg == 0 && meta.kfunc_flags & KF_ITER_NEW)
-		return -size;
+		info.may_read = false;
 	if (is_kfunc_arg_uninit(btf, &args[i]) &&
 	    (is_kfunc_arg_dynptr(btf, &args[i]) || env->allow_uninit_stack))
-		return -size;
-	return size;
+		info.may_read = false;
+	info.must_write = !info.may_read && info.size != U32_MAX;
+	return info;
 }
 
 /* check special kfuncs and return:
diff --git a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
index ff2fb36a0260..20ed71a54686 100644
--- a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
+++ b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
@@ -18,6 +18,9 @@ void __kfunc_btf_root(void)
 
 SEC("tc")
 __success __retval(10)
+__log_level(2)
+__msg("call bpf_kfunc_test_uninit_struct{{.*}}; def: fp0-8 fp0-16")
+__msg_unpriv("call bpf_kfunc_test_uninit_struct{{.*}}; use: fp0-8 fp0-16")
 __flag(BPF_F_TEST_STATE_FREQ)
 __caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
 __prepare_priv
@@ -44,6 +47,9 @@ __naked void struct_poisoned_at_checkpoint(void)
 
 SEC("tc")
 __success __retval(0x2a2a2a2a)
+__log_level(2)
+__msg("call bpf_kfunc_test_uninit_mem{{.*}}; def: fp0-8")
+__msg_unpriv("call bpf_kfunc_test_uninit_mem{{.*}}; use: fp0-8")
 __flag(BPF_F_TEST_STATE_FREQ)
 __caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
 __prepare_priv

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 04/43] bpf: track may_write flags in liveness
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (2 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 03/43] bpf: describe kfunc " Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 05/43] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
                   ` (40 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Track stack slots that instructions may modify, in order to account
for such writes when constructing SCEV expressions.

Consume the may_write flag in arg_access_info separately from reads and
definite writes. Include partially covered slots and all candidate
offsets; imprecise writes mark whole candidate frames. Possible writes
do not kill liveness.

Merge analyzed instances as:

  may_write(dst)  |= may_write(src)
  must_write(dst) &= must_write(src)

For each ancestor frame f, summarize callee writes at the callsite:

  may_write(callsite, f) |= OR_{i in callee} may_write(i, f)

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c                              | 207 ++++++++++++++-------
 .../selftests/bpf/progs/verifier_live_stack.c      |  38 ++--
 2 files changed, 168 insertions(+), 77 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index d4089a488784..e18d5d86c301 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -23,6 +23,7 @@ enum {
 	FM_MAY_READ,	/* stack slots that may be read by this instruction */
 	FM_MUST_WRITE,	/* stack slots written by this instruction */
 	FM_LIVE_BEFORE,	/* stack slots that may be read by this insn and its successors */
+	FM_MAY_WRITE,	/* stack slots that may be written by this instruction */
 	FM_MASK_CNT,
 };
 
@@ -262,6 +263,12 @@ static int mark_stack_write(struct func_instance *instance, u32 frame, u32 insn_
 	return mark_stack_range(instance, frame, insn_idx, FM_MUST_WRITE, lo, hi);
 }
 
+static int mark_stack_may_write(struct func_instance *instance, u32 frame, u32 insn_idx,
+				s32 lo, s32 hi)
+{
+	return mark_stack_range(instance, frame, insn_idx, FM_MAY_WRITE, lo, hi);
+}
+
 /*
  * Mark every half-slot of @frame as possibly read by @insn_idx. This widens
  * the masks to the program's stack budget: a full read recorded at a narrower
@@ -277,9 +284,16 @@ static int mark_stack_read_all(struct bpf_verifier_env *env, struct func_instanc
 			       env->stack_limit / BPF_HALF_REG_SIZE - 1);
 }
 
-/* Accumulate @src, a mask @src_words wide, into may_read of @frame at @insn_idx */
-static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 insn_idx,
-				const unsigned long *src, u32 src_words)
+static int mark_stack_may_write_all(struct bpf_verifier_env *env, struct func_instance *instance,
+				    u32 frame, u32 insn_idx)
+{
+	return mark_stack_may_write(instance, frame, insn_idx, 0,
+				    env->stack_limit / BPF_HALF_REG_SIZE - 1);
+}
+
+/* Accumulate @src, a mask @src_words wide, into @kind mask of @frame at @insn_idx */
+static int mark_stack_mask(struct func_instance *instance, u32 frame, u32 insn_idx, u32 kind,
+			   const unsigned long *src, u32 src_words)
 {
 	u32 nbits = src_words * BITS_PER_LONG;
 	struct frame_masks *fm;
@@ -292,7 +306,7 @@ static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 i
 	fm = widen_frame_masks(instance, frame, BITS_TO_LONGS(last + 1));
 	if (!fm)
 		return -ENOMEM;
-	dst = rel_mask(fm, relative_idx(instance, insn_idx), FM_MAY_READ);
+	dst = rel_mask(fm, relative_idx(instance, insn_idx), kind);
 	/* @src has no bits set past @last, hence none past @fm->words either */
 	src_words = min(src_words, fm->words);
 	for (w = 0; w < src_words; w++)
@@ -300,6 +314,12 @@ static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 i
 	return 0;
 }
 
+static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 insn_idx,
+				const unsigned long *src, u32 src_words)
+{
+	return mark_stack_mask(instance, frame, insn_idx, FM_MAY_READ, src, src_words);
+}
+
 int bpf_jmp_offset(struct bpf_insn *insn)
 {
 	u8 code = insn->code;
@@ -624,16 +644,41 @@ static char *fmt_spis_mask(struct bpf_verifier_env *env, int frame, bool first,
 	return env->tmp_str_buf;
 }
 
+/* Print mask @kind of the instruction at relative index @i for every frame, if any bit is set. */
+static bool print_mask(struct bpf_verifier_env *env, struct func_instance *instance, int i,
+		       const char *name, u32 kind)
+{
+	struct frame_masks *fm;
+	bool printed = false;
+	unsigned long *mask;
+	int frame;
+	u64 pos;
+
+	pos = env->log.end_pos;
+	verbose(env, "%s", name);
+	for (frame = instance->depth; frame >= 0; --frame) {
+		fm = instance->frames[frame];
+		if (!fm)
+			continue;
+		mask = rel_mask(fm, i, kind);
+		if (bitmap_empty(mask, frame_mask_bits(fm)))
+			continue;
+		verbose(env, "%s", fmt_spis_mask(env, frame, !printed, mask, fm->words));
+		printed = true;
+	}
+	if (!printed)
+		bpf_vlog_reset(&env->log, pos);
+	return printed;
+}
+
 static void print_instance(struct bpf_verifier_env *env, struct func_instance *instance)
 {
 	int start = env->subprog_info[instance->subprog].start;
 	struct bpf_insn *insns = env->prog->insnsi;
-	struct frame_masks *fm;
-	unsigned long *mask;
 	int len = instance->insn_cnt;
-	int insn_idx, frame, i;
-	bool has_use, has_def;
 	u64 pos, insn_pos;
+	int insn_idx, i;
+	bool printed;
 
 	if (!(env->log.level & BPF_LOG_LEVEL2))
 		return;
@@ -642,41 +687,17 @@ static void print_instance(struct bpf_verifier_env *env, struct func_instance *i
 	verbose(env, "%s:\n", fmt_instance(env, instance));
 	for (i = 0; i < len; i++) {
 		insn_idx = start + i;
-		has_use = false;
-		has_def = false;
 		pos = env->log.end_pos;
 		verbose(env, "%3d: ", insn_idx);
 		bpf_verbose_insn(env, &insns[insn_idx]);
 		insn_pos = env->log.end_pos;
 		verbose(env, "%*c;", bpf_vlog_alignment(insn_pos - pos), ' ');
-		pos = env->log.end_pos;
-		verbose(env, " use: ");
-		for (frame = instance->depth; frame >= 0; --frame) {
-			fm = instance->frames[frame];
-			if (!fm)
-				continue;
-			mask = rel_mask(fm, i, FM_MAY_READ);
-			if (bitmap_empty(mask, frame_mask_bits(fm)))
-				continue;
-			verbose(env, "%s", fmt_spis_mask(env, frame, !has_use, mask, fm->words));
-			has_use = true;
-		}
-		if (!has_use)
-			bpf_vlog_reset(&env->log, pos);
-		pos = env->log.end_pos;
-		verbose(env, " def: ");
-		for (frame = instance->depth; frame >= 0; --frame) {
-			fm = instance->frames[frame];
-			if (!fm)
-				continue;
-			mask = rel_mask(fm, i, FM_MUST_WRITE);
-			if (bitmap_empty(mask, frame_mask_bits(fm)))
-				continue;
-			verbose(env, "%s", fmt_spis_mask(env, frame, !has_def, mask, fm->words));
-			has_def = true;
-		}
-		if (!has_def)
-			bpf_vlog_reset(&env->log, has_use ? pos : insn_pos);
+		printed = false;
+		printed |= print_mask(env, instance, i, " use: ", FM_MAY_READ);
+		printed |= print_mask(env, instance, i, " def: ", FM_MUST_WRITE);
+		printed |= print_mask(env, instance, i, " may_def: ", FM_MAY_WRITE);
+		if (!printed)
+			bpf_vlog_reset(&env->log, insn_pos);
 		verbose(env, "\n");
 		if (bpf_is_ldimm64(&insns[insn_idx]))
 			i++;
@@ -1437,12 +1458,14 @@ static void arg_track_xfer(struct bpf_verifier_env *env, struct bpf_insn *insn,
 }
 
 /*
- * Record reads for every touched half-slot. A definite
+ * Record possible reads and writes for every touched half-slot. A definite
  * write requires full coverage, a known size, and a single possible offset.
  */
-static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
-				   struct arg_access_info info, u32 frame, u32 insn_idx)
+static int record_stack_access_off(struct func_instance *instance, const struct arg_track *arg,
+				   s64 off_idx, struct arg_access_info info,
+				   u32 frame, u32 insn_idx)
 {
+	s64 fp_off = arg->off[off_idx];
 	s32 slot_hi, slot_lo;
 	int err;
 
@@ -1454,20 +1477,25 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
 		 */
 		return 0;
 
-	if (info.may_read && info.size == U32_MAX) {
-		/* helper/kfunc read unknown amount of bytes from fp_off until fp+0 */
-		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
-		return mark_stack_read(instance, frame, insn_idx, 0, slot_hi);
-	}
+	/*
+	 * Read and may_write marks include partially covered slots.
+	 * An unknown access size may reach from fp_off to the frame top.
+	 */
+	slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
+	slot_lo = info.size == U32_MAX
+		  ? 0
+		  : max_t(s32, (-fp_off - info.size) / STACK_SLOT_SZ, 0);
 	if (info.may_read) {
-		/* Mark any touched slot as use */
-		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
-		slot_lo = max_t(s32, (-fp_off - info.size) / STACK_SLOT_SZ, 0);
 		err = mark_stack_read(instance, frame, insn_idx, slot_lo, slot_hi);
 		if (err)
 			return err;
 	}
-	if (info.must_write && info.size != U32_MAX) {
+	if (info.may_write) {
+		err = mark_stack_may_write(instance, frame, insn_idx, slot_lo, slot_hi);
+		if (err)
+			return err;
+	}
+	if (info.must_write && info.size != U32_MAX && arg->off_cnt == 1) {
 		/* Mark only fully covered slots as def */
 		slot_hi = (-fp_off) / STACK_SLOT_SZ - 1;
 		slot_lo = max_t(s32, (-fp_off - info.size + STACK_SLOT_SZ - 1) / STACK_SLOT_SZ, 0);
@@ -1492,26 +1520,24 @@ static int record_stack_access(struct bpf_verifier_env *env,
 			if (err)
 				return err;
 		}
+		if (info.may_write) {
+			err = mark_stack_may_write_all(env, instance, frame, insn_idx);
+			if (err)
+				return err;
+		}
 		return 0;
 	}
-	if (info.size != U32_MAX && info.must_write && arg->off_cnt != 1)
-		/* multi-offset write cannot set stack_def */
-		return 0;
-
 	for (i = 0; i < arg->off_cnt; i++) {
-		err = record_stack_access_off(instance, arg->off[i], info, frame, insn_idx);
+		err = record_stack_access_off(instance, arg, i, info, frame, insn_idx);
 		if (err)
 			return err;
 	}
 	return 0;
 }
 
-/*
- * When a pointer is ARG_IMPRECISE, conservatively mark every frame in
- * the bitmask as fully used.
- */
+/* Record possible effects on every candidate frame of an imprecise pointer. */
 static int record_imprecise(struct bpf_verifier_env *env, struct func_instance *instance,
-			    u32 mask, u32 insn_idx)
+			    struct arg_access_info info, u32 mask, u32 insn_idx)
 {
 	int depth = instance->depth;
 	int f, err;
@@ -1520,9 +1546,16 @@ static int record_imprecise(struct bpf_verifier_env *env, struct func_instance *
 		if (!(mask & 1))
 			continue;
 		if (f <= depth) {
-			err = mark_stack_read_all(env, instance, f, insn_idx);
-			if (err)
-				return err;
+			if (info.may_read) {
+				err = mark_stack_read_all(env, instance, f, insn_idx);
+				if (err)
+					return err;
+			}
+			if (info.may_write) {
+				err = mark_stack_may_write_all(env, instance, f, insn_idx);
+				if (err)
+					return err;
+			}
 		}
 	}
 	return 0;
@@ -1599,7 +1632,7 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 	if (ptr->frame >= 0 && ptr->frame <= depth)
 		return record_stack_access(env, instance, ptr, info, ptr->frame, insn_idx);
 	if (ptr->frame == ARG_IMPRECISE)
-		return record_imprecise(env, instance, ptr->mask, insn_idx);
+		return record_imprecise(env, instance, info, ptr->mask, insn_idx);
 	/* ARG_NONE: not derived from any frame pointer, skip */
 	return 0;
 }
@@ -1627,6 +1660,9 @@ static int record_arg_access(struct bpf_verifier_env *env,
 			err = mark_stack_read_all(env, instance, f, insn_idx);
 			if (err)
 				return err;
+			err = mark_stack_may_write_all(env, instance, f, insn_idx);
+			if (err)
+				return err;
 		}
 		return 0;
 	}
@@ -1636,7 +1672,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
 	if (frame >= 0 && frame <= depth)
 		err = record_stack_access(env, instance, at, info, frame, insn_idx);
 	else if (frame == ARG_IMPRECISE)
-		err = record_imprecise(env, instance, at->mask, insn_idx);
+		err = record_imprecise(env, instance, info, at->mask, insn_idx);
 	return err;
 }
 
@@ -1996,6 +2032,7 @@ static bool has_fp_args(struct arg_track *args)
 /*
  * Merge a freshly analyzed instance into the original.
  * may_read: union (any pass might read the slot).
+ * may_write: union (slots written on ANY pass).
  * must_write: intersection (only slots written on ALL passes are guaranteed).
  * live_before is recomputed by a subsequent update_instance() on @dst.
  *
@@ -2034,18 +2071,50 @@ static int merge_instances(struct func_instance *dst, struct func_instance *src)
 		for (i = 0; i < dst->insn_cnt; i++) {
 			unsigned long *dst_read = rel_mask(d, i, FM_MAY_READ);
 			unsigned long *dst_write = rel_mask(d, i, FM_MUST_WRITE);
+			unsigned long *dst_may_write = rel_mask(d, i, FM_MAY_WRITE);
 			unsigned long *src_read = rel_mask(s, i, FM_MAY_READ);
 			unsigned long *src_write = rel_mask(s, i, FM_MUST_WRITE);
+			unsigned long *src_may_write = rel_mask(s, i, FM_MAY_WRITE);
 
 			for (w = 0; w < d->words; w++) {
 				dst_read[w] |= w < s->words ? src_read[w] : 0;
 				dst_write[w] &= w < s->words ? src_write[w] : 0;
+				dst_may_write[w] |= w < s->words ? src_may_write[w] : 0;
 			}
 		}
 	}
 	return 0;
 }
 
+/*
+ * Fold a fully analyzed callee instance writes to upper frames as
+ * may_write marks at callsite in caller's frames.
+ */
+static int merge_may_write(struct func_instance *caller, struct func_instance *callee)
+{
+	DECLARE_BITMAP(acc, FRAME_HALF_SPIS);
+	u32 call_idx = callee->callsite;
+	struct frame_masks *fm;
+	u32 f, i, nbits;
+	int err;
+
+	for (f = 0; f < callee->depth; f++) {
+		fm = callee->frames[f];
+		if (!fm)
+			continue;
+		nbits = frame_mask_bits(fm);
+		bitmap_zero(acc, nbits);
+		for (i = 0; i < callee->insn_cnt; i++)
+			bitmap_or(acc, acc, rel_mask(fm, i, FM_MAY_WRITE), nbits);
+		if (bitmap_empty(acc, nbits))
+			continue;
+		err = mark_stack_mask(caller, f, call_idx, FM_MAY_WRITE, acc, fm->words);
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
 static struct func_instance *fresh_instance(struct func_instance *src)
 {
 	struct func_instance *f;
@@ -2161,6 +2230,9 @@ static int analyze_subprog(struct bpf_verifier_env *env,
 					err = mark_stack_read_all(env, instance, f, idx);
 					if (err)
 						goto out_free;
+					err = mark_stack_may_write_all(env, instance, f, idx);
+					if (err)
+						goto out_free;
 				}
 				continue;
 			}
@@ -2213,6 +2285,11 @@ static int analyze_subprog(struct bpf_verifier_env *env,
 					goto out_free;
 			}
 		}
+
+		/* Summarize callee's writes to ancestor frames onto the callsite */
+		err = merge_may_write(instance, callee_instance);
+		if (err)
+			goto out_free;
 	}
 
 	if (prev_instance) {
diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 16b2b1e57534..3ba21430bb75 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -61,7 +61,11 @@ __naked void read_write_join(void)
 SEC("socket")
 __log_level(2)
 __msg("stack use/def subprog#0 must_write_not_same_slot (d0,cs0):")
-__msg("6: (7b) *(u64 *)(r2 +0) = r0{{$}}")
+/*
+ * 'r2 += r1' adds a scalar, so the offset is lost (off_cnt == 0): no def,
+ * but the write conservatively marks the whole frame as may_def.
+ */
+__msg("6: (7b) *(u64 *)(r2 +0) = r0         ; may_def: fp0-8..-{{(512|2048)}}")
 __msg("Live regs before insn:")
 __naked void must_write_not_same_slot(void)
 {
@@ -106,9 +110,9 @@ __naked void must_write_not_same_type(void)
 
 SEC("socket")
 __log_level(2)
-/* Callee writes fp[0]-8: stack_use at call site has slots 0,1 live */
+/* Callee writes fp[0]-8: the def is summarized as may_def at the call site */
 __msg("stack use/def subprog#0 caller_stack_write (d0,cs0):")
-__msg("2: (85) call pc+1{{$}}")
+__msg("2: (85) call pc+1                    ; may_def: fp0-8")
 __msg("stack use/def subprog#1 write_first_param (d1,cs2):")
 __msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8")
 __naked void caller_stack_write(void)
@@ -804,9 +808,9 @@ void __kfunc_btf_root(void)
  */
 SEC("socket")
 __success __log_level(2)
-__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24{{$}}")
-__msg(" 9: (85) call bpf_iter_num_next{{.*}}         ; use: fp0-24{{$}}")
-__msg("14: (85) call bpf_iter_num_destroy{{.*}}      ; use: fp0-24{{$}}")
+__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24 may_def: fp0-24{{$}}")
+__msg(" 9: (85) call bpf_iter_num_next{{.*}}         ; use: fp0-24 may_def: fp0-24{{$}}")
+__msg("14: (85) call bpf_iter_num_destroy{{.*}}      ; use: fp0-24 may_def: fp0-24{{$}}")
 __naked void kfunc_iter_stack_liveness(void)
 {
 	asm volatile (
@@ -1008,7 +1012,8 @@ __naked void four_byte_read_upper_half(void)
 SEC("socket")
 __log_level(2)
 __msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
-__msg("1: (6a) *(u16 *)(r10 -4) = 0{{$}}")
+/* 2-byte write only partially covers the upper half: may_def, but no def. */
+__msg("1: (6a) *(u16 *)(r10 -4) = 0         ; may_def: fp0-4h")
 __msg("2: (61) r0 = *(u32 *)(r10 -4)        ; use: fp0-4h")
 __naked void two_byte_write_no_kill(void)
 {
@@ -1355,9 +1360,12 @@ __naked void fp_spill_loses_precision_kills_liveness(void)
  */
 SEC("socket")
 __log_level(2)
-/* fp-8 live at call (callee conditionally writes → slot not killed) */
+/*
+ * fp-8 live at call: callee conditionally writes it, so the slot is not killed
+ * (no def), but the conditional write surfaces as may_def at the call site.
+ */
 __msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8")
-__msg("4: (85) call pc+2{{$}}")
+__msg("4: (85) call pc+2                    ; may_def: fp0-8")
 __msg("5: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void conditional_stx_in_subprog(void)
 {
@@ -2386,7 +2394,12 @@ __msg("subprog#2 write_first_read_second:")
 __msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
 __msg("18: (79) r0 = *(u64 *)(r2 +0) // r1=fp0-8 r2=fp0-16{{$}}")
 __msg("stack use/def subprog#2 write_first_read_second (d2,cs15):")
-__msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
+/*
+ * Shared across two callsites with swapped args (r1 is fp-8 on one pass,
+ * fp-16 on the other): must_write intersects to empty (no def), may_write
+ * unions to both slots.
+ */
+__msg("17: (7a) *(u64 *)(r1 +0) = 42         ; may_def: fp0-8 fp0-16")
 __msg("18: (79) r0 = *(u64 *)(r2 +0)         ; use: fp0-8 fp0-16")
 __naked void shared_instance_must_write_overwrite(void)
 {
@@ -2848,8 +2861,9 @@ static __used __naked void imprecise_dst_spill_join_sub(void)
 SEC("socket")
 __log_level(2)
 __msg("0: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
-__msg("1: (73) *(u8 *)(r10 -1) = r0{{$}}")
-__msg("2: (6b) *(u16 *)(r10 -4) = r0{{$}}")
+/* narrow stores define nothing, but they may write the half-slot they touch */
+__msg("1: (73) *(u8 *)(r10 -1) = r0         ; may_def: fp0-4h")
+__msg("2: (6b) *(u16 *)(r10 -4) = r0        ; may_def: fp0-4h")
 __msg("3: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void narrow_store_defines_nothing(void)
 {

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 05/43] bpf: summarize may write stack slots in insn_aux_data
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (3 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 04/43] bpf: track may_write flags in liveness Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 06/43] bpf: summarize live " Eduard Zingerman
                   ` (39 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs may-write information to invalidate stack expressions after
indirect writes. Liveness tracks this per function instance,
but SCEV is not callchain-sensitive and needs a summary across
calling contexts.

For each instruction, union the current-frame may_write masks from all
analyzed instances. Represent 8-byte slots as a summary of constituent
halves. Store the summary in insn_aux_data and expose it via a helper.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  6 ++++++
 kernel/bpf/liveness.c        | 41 +++++++++++++++++++++++++++++++++++++++++
 2 files changed, 47 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index eafe79929dfa..4cb36ee3b6f7 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -662,6 +662,11 @@ struct bpf_insn_aux_data {
 	};
 	struct btf_struct_meta *kptr_struct_meta;
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
+	/*
+	 * Per-instruction summary of stack slots in the current frame
+	 * that this instruction may write to.
+	 */
+	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
@@ -1742,6 +1747,7 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
 int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
+const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx);
 bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 spi);
 int bpf_compute_live_registers(struct bpf_verifier_env *env);
 
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index e18d5d86c301..ec9c8beef556 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -718,6 +718,45 @@ static int cmp_instances(const void *pa, const void *pb)
 	return 0;
 }
 
+/* OR the 8-byte slots touched by a half-slot (4-byte) mask into @slots. */
+static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask, u32 nbits)
+{
+	u32 slot;
+
+	for (slot = 0; slot < MAX_BPF_STACK_SLOTS && slot * 2 + 1 < nbits; slot++)
+		if (test_bit(slot * 2, mask) || test_bit(slot * 2 + 1, mask))
+			__set_bit(slot, slots);
+}
+
+/*
+ * Precompute, for each instruction, the OR of may_write masks over its top
+ * frame across all func_instances reaching it, stash it in the insn_aux_data.
+ */
+static void compute_may_write_masks(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_liveness *liveness = env->liveness;
+	struct func_instance *instance;
+	struct frame_masks *fm;
+	u32 nbits;
+	int bkt, i;
+
+	hash_for_each(liveness->func_instances, bkt, instance, hl_node) {
+		fm = instance->frames[instance->depth];
+		if (!fm)
+			continue;
+		nbits = frame_mask_bits(fm);
+		for (i = 0; i < instance->insn_cnt; i++)
+			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
+					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
+	}
+}
+
+const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx)
+{
+	return env->insn_aux_data[insn_idx].may_write_mask;
+}
+
 /* print use/def slots for all instances ordered by callsite first, then by depth */
 static int print_instances(struct bpf_verifier_env *env)
 {
@@ -2352,6 +2391,8 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env)
 			goto out;
 	}
 
+	compute_may_write_masks(env);
+
 	if (env->log.level & BPF_LOG_LEVEL2)
 		err = print_instances(env);
 

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 06/43] bpf: summarize live stack slots in insn_aux_data
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (4 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 05/43] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 07/43] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
                   ` (38 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs live stack information in order to skip tracking dead
slots. Liveness tracks this per function instance, but SCEV is not
callchain-sensitive and needs a summary across calling contexts.

For each instruction, union the current-frame liveness masks from all
analyzed function instances. Represent 8-byte slots as a summary of
constituent halves. Store the summary in insn_aux_data.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/liveness.c        | 10 +++++++---
 2 files changed, 8 insertions(+), 3 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 4cb36ee3b6f7..5b615083ffa9 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -667,6 +667,7 @@ struct bpf_insn_aux_data {
 	 * that this instruction may write to.
 	 */
 	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
+	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index ec9c8beef556..004c1f48a065 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -729,8 +729,9 @@ static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask,
 }
 
 /*
- * Precompute, for each instruction, the OR of may_write masks over its top
- * frame across all func_instances reaching it, stash it in the insn_aux_data.
+ * Precompute, for each instruction, the OR of may_write and live_before masks
+ * over its top frame across all func_instances reaching it, stash them in the
+ * insn_aux_data.
  */
 static void compute_may_write_masks(struct bpf_verifier_env *env)
 {
@@ -746,9 +747,12 @@ static void compute_may_write_masks(struct bpf_verifier_env *env)
 		if (!fm)
 			continue;
 		nbits = frame_mask_bits(fm);
-		for (i = 0; i < instance->insn_cnt; i++)
+		for (i = 0; i < instance->insn_cnt; i++) {
 			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
 					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
+			half_spis_to_slots(aux[instance->subprog_start + i].live_stack_before,
+					   rel_mask(fm, i, FM_LIVE_BEFORE), nbits);
+		}
 	}
 }
 

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 07/43] bpf: summarize regs that may hold a frame pointer in insn_aux_data
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (5 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 06/43] bpf: summarize live " Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 08/43] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
                   ` (37 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs to know which registers might be stack pointers at a
particular instruction. This information is used to avoid widening
registers if that would hinder further verification.
See the patch [1] for additional information.

liveness.c:compute_subprog_args() already tracks this information.
This commit modifies it to save the information in insn_aux_data for
further usage.

[1] "bpf: avoid widening registers that hinder exact stack-slot tracking"

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/liveness.c        | 13 +++++++++++++
 2 files changed, 14 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 5b615083ffa9..d31c1a500958 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -668,6 +668,7 @@ struct bpf_insn_aux_data {
 	 */
 	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
 	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);
+	u16 stack_ptrs; /* bitmask of regs that may hold a frame pointer here (arg_track) */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 004c1f48a065..b9b76684906e 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1896,6 +1896,17 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
 	}
 }
 
+static void record_stack_ptrs(struct bpf_verifier_env *env, int idx, struct arg_track *at_in)
+{
+	u16 r, mask = 0;
+
+	for (r = 0; r < MAX_BPF_REG; r++)
+		if (arg_is_fp(&at_in[r]))
+			mask |= BIT(r);
+
+	env->insn_aux_data[idx].stack_ptrs |= mask;
+}
+
 /*
  * Compute arg tracking dataflow for a single subprog.
  * Runs forward fixed-point with arg_track_xfer(), then records
@@ -2045,6 +2056,8 @@ static int compute_subprog_args(struct bpf_verifier_env *env,
 			snap->slots = nslots;
 			memcpy(snap->at, &at_stack_in[(size_t)i * nslots], nslots * sizeof(*snap->at));
 		}
+
+		record_stack_ptrs(env, idx, at_in[i]);
 	}
 
 	info->at_in = at_in;

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 08/43] bpf: record write effects for atomic operations in liveness.c
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (6 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 07/43] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 09/43] bpf: add tnum_alignment() Eduard Zingerman
                   ` (36 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Atomic read-modify-write instructions are recorded as reads only,
omitting their stack write effects from the may_write masks needed
by SCEV.

Record both read and write accesses for atomic RMW operations. Keep
LOAD_ACQ read-only and STORE_REL write-only. Share the precise/imprecise
stack-access dispatch so both accesses use the same address handling.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c | 20 ++++++++++++++------
 1 file changed, 14 insertions(+), 6 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index b9b76684906e..ce74e3f16eef 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1634,14 +1634,22 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 		break;
 	case BPF_STX:
 		if (BPF_MODE(insn->code) == BPF_ATOMIC) {
-			if (insn->imm == BPF_STORE_REL)
-				write = true;
-			else
-				read = true;
-			if (insn->imm == BPF_LOAD_ACQ)
+			switch (insn->imm) {
+			case BPF_LOAD_ACQ:
 				ptr = &at[insn->src_reg];
-			else
+				read = true;
+				break;
+			case BPF_STORE_REL:
 				ptr = &at[insn->dst_reg];
+				write = true;
+				break;
+			default:
+				/* ADD/AND/OR/XOR(+FETCH), XCHG, CMPXCHG */
+				ptr = &at[insn->dst_reg];
+				read = true;
+				write = true;
+				break;
+			}
 		} else {
 			ptr = &at[insn->dst_reg];
 			write = true;

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 09/43] bpf: add tnum_alignment()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (7 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 08/43] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 10/43] bpf: add cnum{32,64}_union() Eduard Zingerman
                   ` (35 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Used by the SCEV widening logic later in the series.
The verifier checks alignment for some memory accesses.
When SCEV widens a pointer or an index variable, it must retain
alignment information in order to allow array access.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/tnum.h |  2 ++
 kernel/bpf/tnum.c    | 11 +++++++++++
 2 files changed, 13 insertions(+)

diff --git a/include/linux/tnum.h b/include/linux/tnum.h
index ca2cfec8de08..16682248f963 100644
--- a/include/linux/tnum.h
+++ b/include/linux/tnum.h
@@ -134,4 +134,6 @@ static inline bool tnum_subreg_is_const(struct tnum a)
 /* Returns the smallest member of t larger than z */
 u64 tnum_step(struct tnum t, u64 z);
 
+u32 tnum_alignment(struct tnum a);
+
 #endif /* _LINUX_TNUM_H */
diff --git a/kernel/bpf/tnum.c b/kernel/bpf/tnum.c
index ec9c310cf5d7..1f620facd1b7 100644
--- a/kernel/bpf/tnum.c
+++ b/kernel/bpf/tnum.c
@@ -317,3 +317,14 @@ u64 tnum_step(struct tnum t, u64 z)
 	inc = (filled + 1) & t.mask;
 	return t.value | inc;
 }
+
+/*
+ * Return the number of trailing bits known to be zero in a tnum.
+ * Return 64 for a tnum representing zero.
+ */
+u32 tnum_alignment(struct tnum a)
+{
+	u64 v = a.value | a.mask;
+
+	return v ? __ffs64(v) : 64;
+}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 10/43] bpf: add cnum{32,64}_union()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (8 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 09/43] bpf: add tnum_alignment() Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 11/43] bpf: add cnum64_intersect_linear() Eduard Zingerman
                   ` (34 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Add cnum32_union() and cnum64_union() to compute a smallest circular
range containing both inputs.

Do so by enumerating the following configurations in a rotated
frame (a.base at the origin):

0                                              UT_MAX
|---------------------------------------------------|
[= a ==============================]                |
[= b tail =]           [= b main ===================>
[= union ===========================================]

0                                              UT_MAX
|---------------------------------------------------|
[= a =====================]                         |
[= b tail ======]                    [= b main =====>
[= union tail ============]          [= union main =>

0                                              UT_MAX
|---------------------------------------------------|
[= a =======]                                       |
[= b tail =============]             [= b main =====>
[= union tail =========]             [= union main =>

0                                              UT_MAX
|---------------------------------------------------|
[= a =========================]                     |
|                 [= b ====================]        |
[= union ==================================]        |

0                                              UT_MAX
|---------------------------------------------------|
[= a =====================================]         |
|               [= b ========]                      |
[= union =================================]         |

0                                              UT_MAX
|---------------------------------------------------|
[= a ============]                                  |
|                               [= b =======]       |

Two possible covering arcs:

[= ab ======================================]       |
[= ba tail ======]              [= ba main =========>

Pick the smaller one.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/cnum.h   |  2 ++
 kernel/bpf/cnum_defs.h | 44 ++++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 46 insertions(+)

diff --git a/include/linux/cnum.h b/include/linux/cnum.h
index 49b7d0c7645d..ddf4726841e6 100644
--- a/include/linux/cnum.h
+++ b/include/linux/cnum.h
@@ -39,6 +39,7 @@ u32 cnum32_umin(struct cnum32 cnum);
 u32 cnum32_umax(struct cnum32 cnum);
 s32 cnum32_smin(struct cnum32 cnum);
 s32 cnum32_smax(struct cnum32 cnum);
+struct cnum32 cnum32_union(struct cnum32 a, struct cnum32 b);
 struct cnum32 cnum32_intersect(struct cnum32 a, struct cnum32 b);
 void cnum32_intersect_with(struct cnum32 *dst, struct cnum32 src);
 void cnum32_intersect_with_urange(struct cnum32 *dst, u32 min, u32 max);
@@ -65,6 +66,7 @@ u64 cnum64_umin(struct cnum64 cnum);
 u64 cnum64_umax(struct cnum64 cnum);
 s64 cnum64_smin(struct cnum64 cnum);
 s64 cnum64_smax(struct cnum64 cnum);
+struct cnum64 cnum64_union(struct cnum64 a, struct cnum64 b);
 struct cnum64 cnum64_intersect(struct cnum64 a, struct cnum64 b);
 void cnum64_intersect_with(struct cnum64 *dst, struct cnum64 src);
 void cnum64_intersect_with_urange(struct cnum64 *dst, u64 min, u64 max);
diff --git a/kernel/bpf/cnum_defs.h b/kernel/bpf/cnum_defs.h
index 30685e43de04..05b17a8b4814 100644
--- a/kernel/bpf/cnum_defs.h
+++ b/kernel/bpf/cnum_defs.h
@@ -198,6 +198,50 @@ static inline struct cnum_t FN(normalize)(struct cnum_t cnum)
 	return cnum;
 }
 
+/*
+ * Return a smallest arc containing both 'a' and 'b'.
+ * Break equal-size ties by choosing the smaller base.
+ */
+struct cnum_t FN(union)(struct cnum_t a, struct cnum_t b)
+{
+	struct cnum_t b1, ab, ba;
+	ut end;
+
+	if (FN(is_empty)(a))
+		return b;
+	if (FN(is_empty)(b))
+		return a;
+
+	/*
+	 * Rotate so that a1.base == 0 and a1.end == a.size.
+	 * Normalize b1 to preserve the full-circle representation.
+	 */
+	b1 = FN(normalize)((struct cnum_t){ b.base - a.base, b.size });
+	end = max(a.size, (ut)(b1.base + b1.size));
+
+	if (FN(urange_overflow)(b1)) {
+		/* a1 reaches b1's main arc: together they cover the circle. */
+		if (b1.base <= a.size)
+			return (struct cnum_t){ 0, UT_MAX };
+
+		/* Extend b1's tail through a1's end, then rotate back. */
+		return FN(normalize)((struct cnum_t){ b.base, end - b1.base });
+	}
+
+	/* ab, rotated back, covers both nonwrapping arcs. */
+	ab = (struct cnum_t){ a.base, end };
+	if (b1.base <= a.size)
+		return FN(normalize)(ab);
+
+	/* The arcs are disjoint; ba is the other possible covering arc. */
+	ba = (struct cnum_t){ b.base, a.size - b1.base };
+	if (ba.size < ab.size ||
+	    (ba.size == ab.size && ba.base < ab.base))
+		ab = ba;
+
+	return FN(normalize)(ab);
+}
+
 struct cnum_t FN(add)(struct cnum_t a, struct cnum_t b)
 {
 	if (FN(is_empty)(a) || FN(is_empty)(b))

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 11/43] bpf: add cnum64_intersect_linear()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (9 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 10/43] bpf: add cnum{32,64}_union() Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations Eduard Zingerman
                   ` (33 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Add cnum64_intersect_linear(), which intersects a cnum64 interval with the
set of integers congruent to 'base' modulo 'step', i.e. the points
base + step * k. It tightens the interval by rounding its signed minimum up
and its signed maximum down to the nearest such point, and returns an empty
cnum when no point falls within the interval.

The arithmetic is done on the signed bounds [smin, smax]. The line is defined
over the integers, while a cnum64 is an arc on the u64 circle, so a u64-modular
computation would misplace the line for negative values whenever 'step' does
not divide 2^64. Working from the signed bounds keeps the residues consistent,
e.g. it can reason about a range like [-9, 3] with step 3.

Also add imod() to bpf_verifier.h: a mathematical modulo returning an
always-non-negative residue (unlike C's '%', which follows the sign of the
dividend). It is used by cnum64_intersect_linear() and by the base/step
tracking added in subsequent patches.

This is a building block for tracking scalar registers whose value is known
to lie on a line base + step * k.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 14 ++++++++++++++
 include/linux/cnum.h         |  1 +
 kernel/bpf/cnum.c            | 36 ++++++++++++++++++++++++++++++++++++
 3 files changed, 51 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d31c1a500958..79ca8b28a83a 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -9,6 +9,20 @@
 #include <linux/filter.h> /* for MAX_BPF_STACK */
 #include <linux/tnum.h>
 #include <linux/cnum.h>
+#include <linux/math64.h> /* for div_s64_rem() */
+
+/*
+ * Mathematical modulo, the residue of 'v' modulo 'step', normalized to [0, step).
+ * This differs from C's '%' operator, which truncates the quotient toward zero
+ * and so returns a remainder with the sign of the dividend (e.g. -1 % 3 == -1, not 2).
+ */
+static inline u16 imod(s64 v, u16 step)
+{
+	s32 rem;
+
+	div_s64_rem(v, step, &rem);
+	return rem < 0 ? rem + step : rem;
+}
 
 /* Maximum variable offset umax_value permitted when resolving memory accesses.
  * In practice this is far bigger than any realistic pointer offset; this limit
diff --git a/include/linux/cnum.h b/include/linux/cnum.h
index ddf4726841e6..49160318adf9 100644
--- a/include/linux/cnum.h
+++ b/include/linux/cnum.h
@@ -80,5 +80,6 @@ bool cnum64_is_subset(struct cnum64 outer, struct cnum64 inner);
 
 struct cnum32 cnum32_from_cnum64(struct cnum64 cnum);
 struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b);
+struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step);
 
 #endif /* _LINUX_CNUM_H */
diff --git a/kernel/bpf/cnum.c b/kernel/bpf/cnum.c
index 86142cb2aee5..8ec5187d7c75 100644
--- a/kernel/bpf/cnum.c
+++ b/kernel/bpf/cnum.c
@@ -2,6 +2,7 @@
 /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
 
 #include <linux/bits.h>
+#include <linux/bpf_verifier.h> /* for imod() */
 
 #define T 32
 #include "cnum_defs.h"
@@ -118,3 +119,38 @@ struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b)
 	}
 	return t;
 }
+
+/* Intersect 'in' with the set of integers defined by equation 'base + step * k'. */
+struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step)
+{
+	s64 smin = cnum64_smin(in);
+	s64 smax = cnum64_smax(in);
+	s64 lo, hi;
+	u16 d;
+
+	if (step <= 1 || cnum64_is_empty(in) || cnum64_srange_overflow(in))
+		return in;
+	/*
+	 * Round smin up to the next value congruent to 'base' modulo 'step',
+	 * i.e. increase smin by d = (base - smin) mod step:
+	 *
+	 *                 |<---- d ---->|
+	 *     |-----------|=============|...
+	 * base+step*k    smin       base+step*(k+1)
+	 */
+	d = imod(base - imod(smin, step), step);
+	if ((u64)smax - (u64)smin < d)
+		return CNUM64_EMPTY;
+	lo = smin + d;
+	/*
+	 * Round smax down to the previous value congruent to 'base' modulo 'step',
+	 * i.e. decrease smax by d = (smax - base) mod step:
+	 *
+	 *     |<--- d --->|
+	 *  ...|===========|-------------|
+	 * base+step*k    smax       base+step*(k+1)
+	 */
+	d = imod(imod(smax, step) - base, step);
+	hi = smax - d;
+	return cnum64_from_srange(lo, hi);
+}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (10 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 11/43] bpf: add cnum64_intersect_linear() Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:37 ` [PATCH bpf-next v2 13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
                   ` (32 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Expose the existing helpers as bpf_rev_opcode() and bpf_flip_opcode()
and declare them in bpf_verifier.h. The helpers are used by SCEV
logic while analyzing loop conditions.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  3 +++
 kernel/bpf/verifier.c        | 16 +++++++---------
 2 files changed, 10 insertions(+), 9 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 79ca8b28a83a..4371fb3405ac 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1872,4 +1872,7 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env);
 int bpf_do_misc_fixups(struct bpf_verifier_env *env);
 int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 
+int bpf_flip_opcode(u32 opcode);
+u8 bpf_rev_opcode(u8 opcode);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 98e64a138921..4dd71bf7c1d9 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -17231,8 +17231,6 @@ static void find_good_pkt_pointers(struct bpf_verifier_state *vstate,
 
 static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state *reg2,
 				u8 opcode, bool is_jmp32);
-static u8 rev_opcode(u8 opcode);
-
 /*
  * Learn more information about live branches by simulating refinement on both branches.
  * regs_refine_cond_op() is sound, so producing ill-formed register bounds for the branch means
@@ -17241,7 +17239,7 @@ static u8 rev_opcode(u8 opcode);
 static int simulate_both_branches_taken(struct bpf_verifier_env *env, u8 opcode, bool is_jmp32)
 {
 	/* Fallthrough (FALSE) branch */
-	regs_refine_cond_op(&env->false_reg1, &env->false_reg2, rev_opcode(opcode), is_jmp32);
+	regs_refine_cond_op(&env->false_reg1, &env->false_reg2, bpf_rev_opcode(opcode), is_jmp32);
 	reg_bounds_sync(&env->false_reg1);
 	reg_bounds_sync(&env->false_reg2);
 	/*
@@ -17427,7 +17425,7 @@ static int is_scalar_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_s
 	return simulate_both_branches_taken(env, opcode, is_jmp32);
 }
 
-static int flip_opcode(u32 opcode)
+int bpf_flip_opcode(u32 opcode)
 {
 	/* How can we transform "a <op> b" into "b <op> a"? */
 	static const u8 opcode_flip[16] = {
@@ -17458,7 +17456,7 @@ static int is_pkt_ptr_branch_taken(struct bpf_reg_state *dst_reg,
 		pkt = dst_reg;
 	} else if (dst_reg->type == PTR_TO_PACKET_END) {
 		pkt = src_reg;
-		opcode = flip_opcode(opcode);
+		opcode = bpf_flip_opcode(opcode);
 	} else {
 		return -1;
 	}
@@ -17513,7 +17511,7 @@ static int is_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_state *r
 
 		/* arrange that reg2 is a scalar, and reg1 is a pointer */
 		if (!is_reg_const(reg2, is_jmp32)) {
-			opcode = flip_opcode(opcode);
+			opcode = bpf_flip_opcode(opcode);
 			swap(reg1, reg2);
 		}
 		/* and ensure that reg2 is a constant */
@@ -17547,7 +17545,7 @@ static int is_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_state *r
 /* Opcode that corresponds to a *false* branch condition.
  * E.g., if r1 < r2, then reverse (false) condition is r1 >= r2
  */
-static u8 rev_opcode(u8 opcode)
+u8 bpf_rev_opcode(u8 opcode)
 {
 	switch (opcode) {
 	case BPF_JEQ:		return BPF_JNE;
@@ -17582,7 +17580,7 @@ static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state
 	case BPF_JGT:
 	case BPF_JSGE:
 	case BPF_JSGT:
-		opcode = flip_opcode(opcode);
+		opcode = bpf_flip_opcode(opcode);
 		swap(reg1, reg2);
 		break;
 	default:
@@ -17649,7 +17647,7 @@ static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state
 			reg1->var_off = tnum_or(reg1->var_off, tnum_const(val));
 		}
 		break;
-	case BPF_JSET | BPF_X: /* reverse of BPF_JSET, see rev_opcode() */
+	case BPF_JSET | BPF_X: /* reverse of BPF_JSET, see bpf_rev_opcode() */
 		if (!is_reg_const(reg2, is_jmp32))
 			swap(reg1, reg2);
 		if (!is_reg_const(reg2, is_jmp32))

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (11 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps Eduard Zingerman
                   ` (31 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

regsafe() currently requires an exact match for PTR_TO_STACK registers.
This prevents an explored state from pruning a current state whose
variable-offset range is a subset of the explored range.

This change allows a checkpoint containing a widened stack pointer to
cover narrower offsets reached on subsequent loop iterations,
for example [-56, -8] within [-64, -8].

Merge the PTR_TO_STACK case with PTR_TO_MEM and others, as the set of
necessary checks is almost identical. The only difference is a map_uid
check, which is not needed for PTR_TO_STACK. All locations that
produce PTR_TO_STACK reset map_uid to zero.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/states.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 18bf7b660c2f..345079cf64d3 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -610,6 +610,7 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold,
 	case PTR_TO_MEM:
 	case PTR_TO_BUF:
 	case PTR_TO_TP_BUFFER:
+	case PTR_TO_STACK:
 		/* If the new min/max/var_off satisfy the old ones and
 		 * everything else matches, we are OK.
 		 */
@@ -643,8 +644,6 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold,
 		/* new val must satisfy old val knowledge */
 		return range_within(rold, rcur) &&
 		       tnum_in(rold->var_off, rcur->var_off);
-	case PTR_TO_STACK:
-		return regs_exact(rold, rcur, idmap);
 	case PTR_TO_ARENA:
 		return true;
 	case PTR_TO_INSN:

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (12 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 14:02   ` sashiko-bot
  2026-10-04 14:40   ` bot+bpf-ci
  2026-10-04 13:37 ` [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
                   ` (30 subsequent siblings)
  44 siblings, 2 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Extend scalar register tracking with a linear "base + step * k" description:
each scalar carries two u16 fields, base and step (invariant: base < step),
recording that its value is some point on the line base + step * k.
A fresh scalar starts as base=0, step=1, i.e. any integer.

This lets the verifier reason about strided values, e.g. an index multiplied
by an element size, which neither the min/max bounds nor the tnum can capture
for a non-power-of-2 stride. Such reasoning is necessary to allow
array access with a variable offset through PTR_TO_BTF_ID, e.g.
for expressions like p->arr[i].field.

reg_bounds_sync() uses the description in two ways:
- deduce_bounds_64_from_step() intersects the 64-bit range with the line via
  cnum64_intersect_linear(), snapping smin/smax to actual points on it;
- __reg_bound_offset() clears the low bits of var_off implied by the number
  of trailing zeros shared by base and step.

range_within() is extended to check whether the cur register's equation
describes a subset of the points described by the old register's equation.

Scalar ALU operations update base/step as follows:
- ADD of a constant (including pointer + scalar): keep step, shift base;
- MUL by a constant: scale step and base by the constant;
- LSH by a constant: scale step and base by the corresponding power of two;
- All other operations reset to base=0, step=1.

Branch (eq/neq) checks do not take into account base/step yet.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |   7 ++
 kernel/bpf/log.c             |   2 +
 kernel/bpf/states.c          |  26 ++++++-
 kernel/bpf/verifier.c        | 182 ++++++++++++++++++++++++++++++++++++++++++-
 4 files changed, 213 insertions(+), 4 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 4371fb3405ac..05914fcbe77c 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -178,6 +178,13 @@ struct bpf_reg_state {
 	 * during state comparisons.
 	 */
 	u32 map_uid;
+	/*
+	 * The value described by this register, interpreted as s64, lies on
+	 * a line described by a linear equation base + step * k.
+	 * Invariant: base < step.
+	 */
+	u16 base;
+	u16 step;
 	/* if (!precise && SCALAR_VALUE) min/max/tnum don't affect safety */
 	bool precise;
 };
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index d850a7863d2e..900d1bb1988b 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -692,6 +692,8 @@ static void print_reg_state(struct bpf_verifier_env *env,
 
 			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
 			verbose_a("var_off=%s", tn_buf);
+			if (reg->base != 0 || reg->step != 1)
+				verbose_a("step=%d+%d", reg->base, reg->step);
 		}
 	}
 	verbose(env, ")");
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 345079cf64d3..981328f5eb60 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -302,8 +302,30 @@ int bpf_update_branch_counts(struct bpf_verifier_env *env, struct bpf_verifier_s
 static bool range_within(const struct bpf_reg_state *old,
 			 const struct bpf_reg_state *cur)
 {
-	return cnum64_is_subset(old->r64, cur->r64) &&
-	       cnum32_is_subset(old->r32, cur->r32);
+	if (!cnum64_is_subset(old->r64, cur->r64) ||
+	    !cnum32_is_subset(old->r32, cur->r32))
+		return false;
+
+	if (old->step <= 1)
+		return true;
+
+	if (cnum64_is_const(cur->r64))
+		return imod((s64)cnum64_smin(cur->r64), old->step) == old->base;
+
+	/*
+	 * Both `old` and `cur` define some sets of points.
+	 * Return true, if points defined by `cur` are a subset of points defined by `old`:
+	 * - bounds for `cur` should be within bounds for `old`;
+	 * - cur->step should be dividable by old->step;
+	 * - cur->base should start at integer number of old->step
+	 *   steps from old->base.
+	 *
+	 * E.g. the following ranges are compatible:
+	 * - old [0, 10] step 2
+	 * - new [2, 10] step 4
+	 */
+	return cur->step % old->step == 0 &&
+	       (cur->base - old->base) % old->step == 0;
 }
 
 /* If in the old state two registers had the same id, then they need to have
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 4dd71bf7c1d9..f71778a75f30 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5,6 +5,7 @@
  */
 #include <uapi/linux/btf.h>
 #include <linux/bpf-cgroup.h>
+#include <linux/count_zeros.h>
 #include <linux/kernel.h>
 #include <linux/types.h>
 #include <linux/slab.h>
@@ -1910,12 +1911,128 @@ static void bpf_diag_record_caller_saved(struct bpf_verifier_env *env,
 	}
 }
 
+static void reg_step_reset(struct bpf_reg_state *reg)
+{
+	reg->base = 0;
+	reg->step = 1;
+}
+
+static void scalar_step_add(struct bpf_reg_state *dst_reg,
+			    const struct bpf_reg_state *a,
+			    const struct bpf_reg_state *b)
+{
+	const struct bpf_reg_state *reg;
+	s64 amount, tmp;
+
+	if (tnum_is_const(b->var_off)) {
+		reg = a;
+		amount = (s64)b->var_off.value;
+	} else if (tnum_is_const(a->var_off)) {
+		reg = b;
+		amount = (s64)a->var_off.value;
+	} else {
+		reg_step_reset(dst_reg);
+		return;
+	}
+
+	/*
+	 * Let x be a possible signed value of the register,
+	 * d = amount, s = reg->step, and b = reg->base.
+	 * Then x mod s = b, with s > 0 and 0 <= b < s.
+	 * If the addition causes signed underflow, the resulting value
+	 * is x + d + 2**64.
+	 *
+	 * Taking the result modulo s and substituting x mod s = b gives:
+	 *   (x + d + 2**64) mod s = (b + d + (2**64 mod s)) mod s      (1)
+	 * Keeping reg->step = s and updating reg->base to (b + d) mod s
+	 * requires (1) to equal (b + d) mod s, which holds if and only
+	 * if 2**64 mod s = 0.
+	 * Signed overflow subtracts 2**64 and has the same requirement.
+	 *
+	 * Conservatively reset the stride if either wrap is possible.
+	 */
+	if (check_add_overflow(reg_smin(reg), amount, &tmp) ||
+	    check_add_overflow(reg_smax(reg), amount, &tmp)) {
+		reg_step_reset(dst_reg);
+		return;
+	}
+
+	dst_reg->base = (reg->base + imod(amount, reg->step)) % reg->step;
+	dst_reg->step = reg->step;
+}
+
+static void scalar_step_scale(struct bpf_reg_state *dst_reg, u64 amount)
+{
+	u16 step;
+	s64 tmp;
+
+	/*
+	 * Using the same notation as scalar_step_add(), with d > 0,
+	 * the scaled value satisfies (x * d) mod (s * d) = b * d.
+	 * If the multiplication wraps, its signed result can be represented
+	 * as x * d - n * 2**64 for some nonzero integer n.
+	 *
+	 * Taking the result modulo s * d gives:
+	 *   (x * d - n * 2**64) mod (s * d) =
+	 *   (b * d - (n * 2**64 mod (s * d))) mod (s * d)              (1)
+	 * Updating dst_reg->base to b * d and dst_reg->step to s * d
+	 * requires (1) to equal b * d, which holds if and only if
+	 * n * 2**64 mod (s * d) = 0.
+	 *
+	 * Conservatively reset the stride if a wrap is possible.
+	 */
+	if (amount == 0 || check_mul_overflow(dst_reg->step, amount, &step) ||
+	    check_mul_overflow(reg_smin(dst_reg), (s64)amount, &tmp) ||
+	    check_mul_overflow(reg_smax(dst_reg), (s64)amount, &tmp)) {
+		reg_step_reset(dst_reg);
+		return;
+	}
+
+	dst_reg->base = (dst_reg->base * amount) % step;
+	dst_reg->step = step;
+}
+
+/*
+ * Let x be a possible signed value of the register,
+ * s = reg->step, and b = reg->base.
+ * Then x mod s = b, with s > 0 and 0 <= b < s.
+ * Truncating to N bits and then zero- or sign-extending x
+ * is equivalent to computing x + k * 2**N for some integer k.
+ * For example, sign-extending the low 8 bits of x = 255 can
+ * be expressed as 255 - 2**8 = -1.
+ *
+ * Taking the result modulo s and substituting x mod s = b gives:
+ *   (x + k * 2**N) mod s = (b + (k * 2**N mod s)) mod s
+ * Keeping reg->base = b and reg->step = s requires this to equal b,
+ * which holds if and only if k * 2**N mod s = 0.
+ * Continuing the example with s = 3 and b = 0, we have 255 mod 3 = 0,
+ * but (-1) mod 3 = 2, so the original base and step no longer hold.
+ *
+ * Conservatively reset the stride unless all possible values fit in
+ * the range where the conversion leaves them unchanged.
+ * Call before updating the register's bounds. Requires 0 < bits < 64.
+ */
+static void reg_step_check_unsigned(struct bpf_reg_state *reg, u32 bits)
+{
+	if (reg_umax(reg) >= (1ULL << bits))
+		reg_step_reset(reg);
+}
+
+static void reg_step_check_signed(struct bpf_reg_state *reg, u32 bits)
+{
+	s64 limit = 1LL << (bits - 1);
+
+	if (reg_smin(reg) < -limit || reg_smax(reg) >= limit)
+		reg_step_reset(reg);
+}
+
 /* This helper doesn't clear reg->id */
 static void ___mark_reg_known(struct bpf_reg_state *reg, u64 imm)
 {
 	reg->var_off = tnum_const(imm);
 	reg->r64 = cnum64_from_urange(imm, imm);
 	reg->r32 = cnum32_from_urange((u32)imm, (u32)imm);
+	reg_step_reset(reg);
 }
 
 /* Mark the unknown part of a register (variable offset or scalar value) as
@@ -2158,10 +2275,16 @@ static void deduce_bounds_64_from_32(struct bpf_reg_state *reg)
 	reg->r64 = cnum64_cnum32_intersect(reg->r64, reg->r32);
 }
 
+static void deduce_bounds_64_from_step(struct bpf_reg_state *reg)
+{
+	reg->r64 = cnum64_intersect_linear(reg->r64, reg->base, reg->step);
+}
+
 static void __reg_deduce_bounds(struct bpf_reg_state *reg)
 {
 	deduce_bounds_32_from_64(reg);
 	deduce_bounds_64_from_32(reg);
+	deduce_bounds_64_from_step(reg);
 }
 
 /* Attempts to improve var_off based on unsigned min/max information */
@@ -2173,8 +2296,18 @@ static void __reg_bound_offset(struct bpf_reg_state *reg)
 	struct tnum var32_off = tnum_intersect(tnum_subreg(var64_off),
 					       tnum_range(reg_u32_min(reg),
 							  reg_u32_max(reg)));
+	u32 trailing_zero_bits;
+	u16 base = reg->base;
+	u16 step = reg->step;
 
 	reg->var_off = tnum_or(tnum_clear_subreg(var64_off), var32_off);
+
+	if (base == 0)
+		trailing_zero_bits = count_trailing_zeros(step);
+	else
+		trailing_zero_bits = min(count_trailing_zeros(base),
+					 count_trailing_zeros(step));
+	reg->var_off = tnum_and(reg->var_off, tnum_const(~0ULL << trailing_zero_bits));
 }
 
 static bool range_bounds_violation(struct bpf_reg_state *reg);
@@ -2250,6 +2383,7 @@ static int reg_bounds_sanity_check(struct bpf_verifier_env *env,
 	if (env->test_reg_invariants)
 		return -EFAULT;
 	__mark_reg_unbounded(reg);
+	reg_step_reset(reg);
 	return 0;
 }
 
@@ -2260,6 +2394,7 @@ void bpf_mark_reg_unknown_imprecise(struct bpf_reg_state *reg)
 	reg->type = SCALAR_VALUE;
 	reg->var_off = tnum_unknown;
 	__mark_reg_unbounded(reg);
+	reg_step_reset(reg);
 }
 
 /* Mark a register as having a completely unknown (scalar) value,
@@ -5894,6 +6029,7 @@ static int check_buffer_access(struct bpf_verifier_env *env,
 /* BPF architecture zero extends alu32 ops into 64-bit registesr */
 static void zext_32_to_64(struct bpf_reg_state *reg)
 {
+	reg_step_check_unsigned(reg, 32);
 	reg->var_off = tnum_subreg(reg->var_off);
 	reg_set_urange64(reg, reg_u32_min(reg), reg_u32_max(reg));
 }
@@ -5903,13 +6039,14 @@ static void zext_32_to_64(struct bpf_reg_state *reg)
  */
 static void coerce_reg_to_size(struct bpf_reg_state *reg, int size)
 {
-	u64 mask;
+	u64 mask = (1ULL << (size * 8)) - 1;
+
+	reg_step_check_unsigned(reg, size * 8);
 
 	/* clear high bits in bit representation */
 	reg->var_off = tnum_cast(reg->var_off, size);
 
 	/* fix arithmetic bounds */
-	mask = ((u64)1 << (size * 8)) - 1;
 	if ((reg_umin(reg) & ~mask) == (reg_umax(reg) & ~mask))
 		reg_set_urange64(reg, reg_umin(reg) & mask, reg_umax(reg) & mask);
 	else
@@ -5947,6 +6084,8 @@ static void coerce_reg_to_size_sx(struct bpf_reg_state *reg, int size)
 	u64 top_smax_value, top_smin_value;
 	u64 num_bits = size * 8;
 
+	reg_step_check_signed(reg, num_bits);
+
 	if (tnum_is_const(reg->var_off)) {
 		u64_cval = reg->var_off.value;
 		if (size == 1)
@@ -6012,6 +6151,8 @@ static void coerce_subreg_to_size_sx(struct bpf_reg_state *reg, int size)
 	u32 top_smax_value, top_smin_value;
 	u32 num_bits = size * 8;
 
+	reg_step_check_unsigned(reg, num_bits - 1);
+
 	if (tnum_is_const(reg->var_off)) {
 		u32_val = reg->var_off.value;
 		if (size == 1)
@@ -6721,6 +6862,7 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
 	fake_reg.type = SCALAR_VALUE;
 	__mark_reg_known(&fake_reg, val);
 
+	scalar_step_add(dst_reg, dst_reg, &fake_reg);
 	scalar32_min_max_add(dst_reg, &fake_reg);
 	scalar_min_max_add(dst_reg, &fake_reg);
 	dst_reg->var_off = tnum_add(dst_reg->var_off, fake_reg.var_off);
@@ -15591,6 +15733,24 @@ static int sanitize_check_bounds(struct bpf_verifier_env *env,
 	return 0;
 }
 
+static void scalar_step_mul(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg)
+{
+	if (tnum_is_const(src_reg->var_off))
+		scalar_step_scale(dst_reg, src_reg->var_off.value);
+	else
+		reg_step_reset(dst_reg);
+}
+
+static void scalar_step_lsh(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg)
+{
+	u64 amount = src_reg->var_off.value;
+
+	if (tnum_is_const(src_reg->var_off) && amount < 64)
+		scalar_step_scale(dst_reg, 1ULL << amount);
+	else
+		reg_step_reset(dst_reg);
+}
+
 /* Handles arithmetic on a pointer and a scalar: computes new min/max and var_off.
  * Caller should also handle BPF_MOV case separately.
  * If we return -EACCES, caller may want to try again treating pointer as a
@@ -15759,6 +15919,7 @@ static int adjust_ptr_min_max_vals(struct bpf_verifier_env *env, struct bpf_insn
 		 * added into the variable offset, and we copy the fixed offset
 		 * from ptr_reg.
 		 */
+		scalar_step_add(dst_reg, ptr_reg, off_reg);
 		dst_reg->r64 = cnum64_add(ptr_reg->r64, off_reg->r64);
 		dst_reg->var_off = tnum_add(ptr_reg->var_off, off_reg->var_off);
 		dst_reg->raw = ptr_reg->raw;
@@ -15820,6 +15981,7 @@ static int adjust_ptr_min_max_vals(struct bpf_verifier_env *env, struct bpf_insn
 			if ((!known && smin_val < 0) || dst_reg->range < 0)
 				memset(&dst_reg->raw, 0, sizeof(dst_reg->raw));
 		}
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_AND:
 	case BPF_OR:
@@ -16527,6 +16689,7 @@ static void scalar_byte_swap(struct bpf_reg_state *dst_reg, struct bpf_insn *ins
 		 * Bounds will be re-derived from the new tnum later.
 		 */
 		__mark_reg_unbounded(dst_reg);
+		reg_step_reset(dst_reg);
 	}
 	/* For bswap16/32, truncate dst register to match the swapped size */
 	if (insn->imm == 16 || insn->imm == 32)
@@ -16653,6 +16816,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 	 */
 	switch (opcode) {
 	case BPF_ADD:
+		scalar_step_add(dst_reg, dst_reg, &src_reg);
 		scalar32_min_max_add(dst_reg, &src_reg);
 		scalar_min_max_add(dst_reg, &src_reg);
 		dst_reg->var_off = tnum_add(dst_reg->var_off, src_reg.var_off);
@@ -16661,6 +16825,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		scalar32_min_max_sub(dst_reg, &src_reg);
 		scalar_min_max_sub(dst_reg, &src_reg);
 		dst_reg->var_off = tnum_sub(dst_reg->var_off, src_reg.var_off);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_NEG:
 		env->fake_reg[0] = *dst_reg;
@@ -16668,8 +16833,10 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		scalar32_min_max_sub(dst_reg, &env->fake_reg[0]);
 		scalar_min_max_sub(dst_reg, &env->fake_reg[0]);
 		dst_reg->var_off = tnum_neg(env->fake_reg[0].var_off);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_MUL:
+		scalar_step_mul(dst_reg, &src_reg);
 		dst_reg->var_off = tnum_mul(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_mul(dst_reg, &src_reg);
 		scalar_min_max_mul(dst_reg, &src_reg);
@@ -16690,6 +16857,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 				scalar_min_max_sdiv(dst_reg, &src_reg);
 			else
 				scalar_min_max_udiv(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_MOD:
 		/* BPF mod specification: x % 0 = x */
@@ -16705,6 +16873,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 				scalar_min_max_smod(dst_reg, &src_reg);
 			else
 				scalar_min_max_umod(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_AND:
 		if (tnum_is_const(src_reg.var_off)) {
@@ -16715,6 +16884,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		dst_reg->var_off = tnum_and(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_and(dst_reg, &src_reg);
 		scalar_min_max_and(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_OR:
 		if (tnum_is_const(src_reg.var_off)) {
@@ -16725,13 +16895,16 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		dst_reg->var_off = tnum_or(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_or(dst_reg, &src_reg);
 		scalar_min_max_or(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_XOR:
 		dst_reg->var_off = tnum_xor(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_xor(dst_reg, &src_reg);
 		scalar_min_max_xor(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_LSH:
+		scalar_step_lsh(dst_reg, &src_reg);
 		if (alu32)
 			scalar32_min_max_lsh(dst_reg, &src_reg);
 		else
@@ -16742,15 +16915,18 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 			scalar32_min_max_rsh(dst_reg, &src_reg);
 		else
 			scalar_min_max_rsh(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_ARSH:
 		if (alu32)
 			scalar32_min_max_arsh(dst_reg, &src_reg);
 		else
 			scalar_min_max_arsh(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_END:
 		scalar_byte_swap(dst_reg, insn);
+		reg_step_reset(dst_reg);
 		break;
 	default:
 		break;
@@ -17657,6 +17833,7 @@ static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state
 		 * violations if we're on a dead branch.
 		 */
 		__mark_reg_unbounded(reg1);
+		reg_step_reset(reg1);
 		if (is_jmp32) {
 			t = tnum_and(tnum_subreg(reg1->var_off), tnum_const(~val));
 			reg1->var_off = tnum_with_subreg(reg1->var_off, t);
@@ -17979,6 +18156,7 @@ static void sync_linked_regs(struct bpf_verifier_env *env, struct bpf_verifier_s
 			reg->delta = saved_off;
 			reg->id = saved_id;
 
+			scalar_step_add(reg, reg, &fake_reg);
 			scalar32_min_max_add(reg, &fake_reg);
 			scalar_min_max_add(reg, &fake_reg);
 			reg->var_off = tnum_add(reg->var_off, fake_reg.var_off);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (13 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 15:11   ` bot+bpf-ci
  2026-10-04 13:37 ` [PATCH bpf-next v2 16/43] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
                   ` (29 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Relax the requirements for reads like `p->arr[i].field`, where `p` is a
PTR_TO_BTF_ID, by allowing `i` to be a non-constant value.
This is useful for loops over arrays accessed from a PTR_TO_BTF_ID
when the verifier widens the loop induction variable.

Check that the access stays within array bounds based on the access
register's min and max bounds. Check that the access is aligned with
array elements using the register's `base` and `step`.

Technically, modify btf_struct_access() such that:
- it identifies `field` by walking BTF from the minimum offset
  (the register's reg_smin + static instruction offset);
- it remembers which arrays contain `field`;
- it checks that offsets `k * step` stay within the array
  and point to the same `field`;
- it does not check upper bounds for flexible arrays
  (as is already the case for constant-offset access).

Support for flexible arrays nested inside trailing struct members
(e.g. p->nested.data[i], with data declared as int data[]) is left
for future work.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/btf.c      | 159 ++++++++++++++++++++++++++++++++++++++++++--------
 kernel/bpf/states.c   |   1 +
 kernel/bpf/verifier.c |  38 +++++++-----
 3 files changed, 159 insertions(+), 39 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 0630675377aa..f4d119f3d5ef 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -2079,7 +2079,8 @@ static const struct resolve_vertex *env_stack_peak(struct btf_verifier_env *env)
  * *elem_id: id of u32
  * *total_nelems: (x * y).  Hence, individual elem size is
  *                (*type_size / *total_nelems)
- * *type_id: id of type if it's changed within the function, 0 if not
+ * *type_id: ID of the returned type if leading modifiers were stripped;
+ *           otherwise left unchanged.
  *
  * type: is not an array (e.g. const struct X)
  * return type: type "struct X"
@@ -2087,7 +2088,8 @@ static const struct resolve_vertex *env_stack_peak(struct btf_verifier_env *env)
  * *elem_type: same as return type ("struct X")
  * *elem_id: 0
  * *total_nelems: 1
- * *type_id: id of type if it's changed within the function, 0 if not
+ * *type_id: ID of the returned type if leading modifiers were stripped;
+ *           otherwise left unchanged.
  */
 static const struct btf_type *
 __btf_resolve_size(const struct btf *btf, const struct btf_type *type,
@@ -2120,7 +2122,8 @@ __btf_resolve_size(const struct btf *btf, const struct btf_type *type,
 		case BTF_KIND_CONST:
 		case BTF_KIND_RESTRICT:
 		case BTF_KIND_TYPE_TAG:
-			id = type->type;
+			if (!array_type)
+				id = type->type;
 			type = btf_type_by_id(btf, type->type);
 			break;
 
@@ -7549,6 +7552,14 @@ bool btf_ctx_access(int off, int size, enum bpf_access_type type,
 }
 EXPORT_SYMBOL_GPL(btf_ctx_access);
 
+#define MAX_ARRAYS_WALK 16
+
+/* Description of an array crossed while walking a BTF access chain */
+struct array_access {
+	u32 id;		/* BTF_KIND_ARRAY type id */
+	u32 off;	/* offset of the access within the array */
+};
+
 enum bpf_struct_walk_result {
 	/* < 0 error */
 	WALK_SCALAR = 0,
@@ -7557,10 +7568,43 @@ enum bpf_struct_walk_result {
 	WALK_STRUCT,
 };
 
+static const struct btf_member *find_flex_member(const struct btf *btf, const struct btf_type *t)
+{
+	const struct btf_array *array;
+	const struct btf_type *mtype;
+	const struct btf_member *member;
+	u32 vlen = btf_type_vlen(t);
+
+	if (vlen == 0)
+		return NULL;
+
+	member = btf_type_member(t) + vlen - 1;
+	mtype = btf_type_skip_modifiers(btf, member->type, NULL);
+	if (!btf_type_is_array(mtype))
+		return NULL;
+
+	array = (const struct btf_array *)(mtype + 1);
+	if (array->nelems != 0)
+		return NULL;
+
+	return member;
+}
+
+static int push_array(struct array_access *arrays, u32 *arrays_cnt, struct array_access item)
+{
+	if (arrays == NULL)
+		return 0;
+	if (*arrays_cnt == MAX_ARRAYS_WALK)
+		return -1;
+	arrays[(*arrays_cnt)++] = item;
+	return 0;
+}
+
 static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 			   const struct btf_type *t, int off, int size,
 			   u32 *next_btf_id, enum bpf_type_flag *flag,
-			   const char **field_name, bool walk_flex_arrays)
+			   const char **field_name, bool walk_flex_arrays,
+			   struct array_access *arrays, u32 *arrays_cnt)
 {
 	u32 i, moff, mtrue_end, msize = 0, total_nelems = 0;
 	const struct btf_type *mtype, *elem_type = NULL;
@@ -7595,22 +7639,19 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 		/* If the last element is a variable size array, we may
 		 * need to relax the rule.
 		 */
-		if (vlen == 0)
+		member = find_flex_member(btf, t);
+		if (!member)
 			goto error;
 
-		member = btf_type_member(t) + vlen - 1;
-		mtype = btf_type_skip_modifiers(btf, member->type,
-						NULL);
-		if (!btf_type_is_array(mtype))
+		moff = __btf_member_bit_offset(t, member) / 8;
+		if (off < moff)
 			goto error;
 
+		mtype = btf_type_skip_modifiers(btf, member->type, &mid);
 		array_elem = (struct btf_array *)(mtype + 1);
-		if (array_elem->nelems != 0)
-			goto error;
 
-		moff = __btf_member_bit_offset(t, member) / 8;
-		if (off < moff)
-			goto error;
+		if (push_array(arrays, arrays_cnt, (struct array_access){ mid, off - moff }))
+			return -E2BIG;
 
 		/* allow structure and integer */
 		t = btf_type_skip_modifiers(btf, array_elem->type,
@@ -7733,6 +7774,9 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 			 *      the array's element as long as it is
 			 *      within the mtrue_end boundary.
 			 */
+			if (push_array(arrays, arrays_cnt,
+				       (struct array_access){ mid, off - moff }))
+				return -E2BIG;
 
 			/* skip empty array */
 			if (moff == mtrue_end)
@@ -7825,15 +7869,29 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 
 int btf_struct_access(struct bpf_verifier_log *log,
 		      const struct bpf_reg_state *reg,
-		      int off, int size, enum bpf_access_type atype __maybe_unused,
+		      int _off, int size, enum bpf_access_type atype __maybe_unused,
 		      u32 *next_btf_id, enum bpf_type_flag *flag,
 		      const char **field_name)
 {
 	const struct btf *btf = reg->btf;
 	enum bpf_type_flag tmp_flag = 0;
+	struct array_access arrays[MAX_ARRAYS_WALK];
+	const struct btf_member *member;
 	const struct btf_type *t;
+	u32 i, arrays_cnt = 0;
 	u32 id = reg->btf_id;
-	int err;
+	s64 min_off, max_off, flex_off;
+	int err, off, ret;
+
+	if (check_add_overflow(reg_smin(reg), (s64)_off, &min_off) ||
+	    check_add_overflow(reg_smax(reg), (s64)_off, &max_off) ||
+	    min_off != (int)min_off) {
+		bpf_log(log,
+			"offset computation overflows: register offset range is [%lld, %lld], instruction offset is %d\n",
+			reg_smin(reg), reg_smax(reg), _off);
+		return -EINVAL;
+	}
+	off = min_off;
 
 	while (type_is_alloc(reg->type)) {
 		struct btf_struct_meta *meta;
@@ -7860,26 +7918,30 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	t = btf_type_by_id(btf, id);
 	do {
 		err = btf_struct_walk(log, btf, t, off, size, &id, &tmp_flag,
-				      field_name, !type_is_alloc(reg->type));
-
+				      field_name, !type_is_alloc(reg->type), arrays, &arrays_cnt);
 		switch (err) {
 		case WALK_PTR:
 			/* For local types, the destination register cannot
 			 * become a pointer again.
 			 */
-			if (type_is_alloc(reg->type))
-				return SCALAR_VALUE;
+			if (type_is_alloc(reg->type)) {
+				ret = SCALAR_VALUE;
+				goto check_variable_offset;
+			}
 			/* If we found the pointer or scalar on t+off,
 			 * we're done.
 			 */
 			*next_btf_id = id;
 			*flag = tmp_flag;
-			return PTR_TO_BTF_ID;
+			ret = PTR_TO_BTF_ID;
+			goto check_variable_offset;
 		case WALK_PTR_UNTRUSTED:
 			*flag = MEM_RDONLY | PTR_UNTRUSTED;
-			return PTR_TO_MEM;
+			ret = PTR_TO_MEM;
+			goto check_variable_offset;
 		case WALK_SCALAR:
-			return SCALAR_VALUE;
+			ret = SCALAR_VALUE;
+			goto check_variable_offset;
 		case WALK_STRUCT:
 			/* We found nested struct, so continue the search
 			 * by diving in it. At this point the offset is
@@ -7899,6 +7961,55 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	} while (t);
 
 	return -EINVAL;
+
+check_variable_offset:
+	if (min_off == max_off)
+		return ret;
+
+	/* Find an offset at which access would go to a flexible array tail (if any). */
+	t = btf_type_skip_modifiers(btf, reg->btf_id, NULL);
+	member = find_flex_member(btf, t);
+	flex_off = member ? __btf_member_bit_offset(t, member) / 8 : S64_MAX;
+
+	/*
+	 * If this is a varying offset access, the step recorded within a register
+	 * should correspond to one of the arrays visited while walking.
+	 */
+	for (i = 0; i < arrays_cnt; i++) {
+		s64 array_start, array_end, access_end;
+		u32 asize, esize, elem_id;
+
+		/*
+		 * 'arrays' records top-level arrays only, use __btf_resolve_size()
+		 * to get flattened representation.
+		 */
+		t = btf_type_by_id(btf, arrays[i].id);
+		if (IS_ERR(__btf_resolve_size(btf, t, &asize, NULL, &elem_id, NULL, NULL)) ||
+		    IS_ERR(btf_resolve_size(btf, btf_type_by_id(btf, elem_id), &esize)))
+			continue;
+		/* Make sure every accessed offset lands on the same field of some element. */
+		if (esize == 0 || reg->step % esize != 0)
+			continue;
+
+		array_start = min_off - arrays[i].off;
+		/* An array within the flexible tail has no upper bound to exceed */
+		if (array_start >= flex_off && btf_type_is_array(t) &&
+		    btf_type_array(t)->nelems == 0)
+			return ret;
+
+		/* the 'size' bytes read at max_off must stay within the array */
+		if (check_add_overflow(array_start, (s64)asize, &array_end) ||
+		    check_add_overflow(max_off, (s64)size, &access_end) ||
+		    access_end > array_end)
+			continue;
+		return ret;
+	}
+
+	bpf_log(log,
+		"invalid variable offset access into struct %s: offsets [%lld, %lld] size %d step %d do not match any array\n",
+		btf_name_by_offset(btf, btf_type_by_id(btf, reg->btf_id)->name_off),
+		min_off, max_off, size, reg->step);
+	return -EINVAL;
 }
 
 /* Check that two BTF types, each specified as an BTF object + id, are exactly
@@ -7939,7 +8050,7 @@ bool btf_struct_ids_match(struct bpf_verifier_log *log,
 	if (!type)
 		return false;
 	err = btf_struct_walk(log, btf, type, off, 1, &id, &flag, NULL,
-			      walk_flex_arrays);
+			      walk_flex_arrays, NULL, NULL);
 	if (err != WALK_STRUCT)
 		return false;
 
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 981328f5eb60..2f6a7620eb72 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -633,6 +633,7 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold,
 	case PTR_TO_BUF:
 	case PTR_TO_TP_BUFFER:
 	case PTR_TO_STACK:
+	case PTR_TO_BTF_ID:
 		/* If the new min/max/var_off satisfy the old ones and
 		 * everything else matches, we are OK.
 		 */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f71778a75f30..27a53204696f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6521,6 +6521,7 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	const char *field_name = NULL;
 	enum bpf_type_flag flag = 0;
 	u32 btf_id = 0;
+	s64 min_off;
 	int ret;
 
 	if (!env->allow_ptr_leaks) {
@@ -6536,36 +6537,31 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 		return -EINVAL;
 	}
 
-	if (!tnum_is_const(reg->var_off)) {
-		char tn_buf[48];
-
-		tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
+	if (check_add_overflow(reg_smin(reg), off, &min_off)) {
 		verbose(env,
-			"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
-			reg_arg_name(env, argno), tname, off, tn_buf);
+			"%s is ptr_%s access, offset computation overflows: register's minimal offset is %lld, instruction offset is %d\n",
+			reg_arg_name(env, argno), tname, reg_smin(reg), off);
 		return -EACCES;
 	}
 
-	off += reg->var_off.value;
-
-	if (off < 0) {
+	if (min_off < 0) {
 		verbose(env,
-			"%s is ptr_%s invalid negative access: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s invalid negative access: off=%lld\n",
+			reg_arg_name(env, argno), tname, min_off);
 		return -EACCES;
 	}
 
 	if (reg->type & MEM_USER) {
 		verbose(env,
-			"%s is ptr_%s access user memory: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s access user memory\n",
+			reg_arg_name(env, argno), tname);
 		return -EACCES;
 	}
 
 	if (reg->type & MEM_PERCPU) {
 		verbose(env,
-			"%s is ptr_%s access percpu memory: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s access percpu memory\n",
+			reg_arg_name(env, argno), tname);
 		return -EACCES;
 	}
 
@@ -6575,6 +6571,18 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	}
 
 	if (env->ops->btf_struct_access && !type_is_alloc(reg->type) && atype == BPF_WRITE) {
+		if (!tnum_is_const(reg->var_off)) {
+			char tn_buf[48];
+
+			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
+			verbose(env,
+				"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
+				reg_arg_name(env, argno), tname, off, tn_buf);
+			return -EACCES;
+		}
+
+		off += reg->var_off.value;
+
 		if (!btf_is_kernel(reg->btf)) {
 			verifier_bug(env, "reg->btf must be kernel btf");
 			return -EFAULT;

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 16/43] bpf: save DFS postorder numbers for program instructions
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (14 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:37 ` [PATCH bpf-next v2 17/43] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
                   ` (28 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Dominator intersections and SCEV worklist scheduling need to compare
instructions by their DFS postorder rank. Save per-instruction
postorder numbers alongside the existing postorder sequence.

This commit modifies bpf_compute_postorder() to actually do the DFS
traversal, instead of scheduling traversal of all siblings at once.

Consider a graph:

  A -> B
  A -> C
  C -> B

Old result: C, B, A
DFS result: B, C, A

Fixes: efcda22aa541 ("bpf: compute instructions postorder per subprogram")
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/cfg.c             | 56 +++++++++++++++++++++++++++-----------------
 kernel/bpf/const_fold.c      |  2 ++
 kernel/bpf/verifier.c        |  1 +
 4 files changed, 39 insertions(+), 21 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 05914fcbe77c..14813040df84 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1036,6 +1036,7 @@ struct bpf_verifier_env {
 		 * see bpf_subprog_info->postorder_start.
 		 */
 		int *insn_postorder;
+		int *postorder_nums;
 		int cur_stack;
 		/* current position in the insn_postorder vector */
 		int cur_postorder;
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index d8a579680e5b..bce7b5abcfda 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -767,6 +767,11 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
 	return ret;
 }
 
+struct dfs_state {
+	u32 traversed:1;
+	u32 next_succ:31;
+};
+
 /*
  * For each subprogram 'i' fill array env->cfg.insn_subprogram sub-range
  * [env->subprog_info[i].postorder_start, env->subprog_info[i+1].postorder_start)
@@ -774,43 +779,52 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
  */
 int bpf_compute_postorder(struct bpf_verifier_env *env)
 {
-	u32 cur_postorder, i, top, stack_sz, s;
-	int *stack = NULL, *postorder = NULL, *state = NULL;
-	struct bpf_iarray *succ;
+	int *stack = NULL, *postorder = NULL, *postorder_nums = NULL;
+	int subprog_idx, stack_sz, cur, s, cur_postorder, start;
+	struct dfs_state *state = NULL;
+	struct bpf_iarray *succ = NULL;
 
+	postorder_nums = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
 	postorder = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
-	state = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
 	stack = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
-	if (!postorder || !state || !stack) {
+	state = kvzalloc_objs(struct dfs_state, env->prog->len, GFP_KERNEL_ACCOUNT);
+	if (!postorder_nums || !postorder || !stack || !state) {
+		kvfree(postorder_nums);
 		kvfree(postorder);
-		kvfree(state);
 		kvfree(stack);
+		kvfree(state);
 		return -ENOMEM;
 	}
 	cur_postorder = 0;
-	for (i = 0; i < env->subprog_cnt; i++) {
-		env->subprog_info[i].postorder_start = cur_postorder;
-		stack[0] = env->subprog_info[i].start;
+	for (subprog_idx = 0; subprog_idx < env->subprog_cnt; subprog_idx++) {
+		start = env->subprog_info[subprog_idx].start;
+		env->subprog_info[subprog_idx].postorder_start = cur_postorder;
+		stack[0] = start;
 		stack_sz = 1;
+		state[start].traversed = true;
 		do {
-			top = stack[stack_sz - 1];
-			state[top] |= DISCOVERED;
-			if (state[top] & EXPLORED) {
-				postorder[cur_postorder++] = top;
+			cur = stack[stack_sz - 1];
+			succ = bpf_insn_successors(env, cur);
+			if (state[cur].next_succ == succ->cnt) {
+				postorder_nums[cur] = cur_postorder;
+				postorder[cur_postorder] = cur;
+				cur_postorder++;
 				stack_sz--;
 				continue;
 			}
-			succ = bpf_insn_successors(env, top);
-			for (s = 0; s < succ->cnt; ++s) {
-				if (!state[succ->items[s]]) {
-					stack[stack_sz++] = succ->items[s];
-					state[succ->items[s]] |= DISCOVERED;
-				}
+			s = succ->items[state[cur].next_succ];
+			if (!state[s].traversed) {
+				state[s].traversed = true;
+				state[s].next_succ = 0;
+				stack[stack_sz] = s;
+				stack_sz++;
+				continue;
 			}
-			state[top] |= EXPLORED;
+			state[cur].next_succ++;
 		} while (stack_sz);
 	}
-	env->subprog_info[i].postorder_start = cur_postorder;
+	env->subprog_info[subprog_idx].postorder_start = cur_postorder;
+	env->cfg.postorder_nums = postorder_nums;
 	env->cfg.insn_postorder = postorder;
 	env->cfg.cur_postorder = cur_postorder;
 	kvfree(stack);
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index b1528adbeb79..8dd8c5958eb8 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -398,5 +398,7 @@ int bpf_prune_dead_branches(struct bpf_verifier_env *env)
 	/* recompute postorder, since CFG has changed */
 	kvfree(env->cfg.insn_postorder);
 	env->cfg.insn_postorder = NULL;
+	kvfree(env->cfg.postorder_nums);
+	env->cfg.postorder_nums = NULL;
 	return bpf_compute_postorder(env);
 }
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 27a53204696f..3298e2192239 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -23092,6 +23092,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
 	bpf_stack_liveness_free(env);
+	kvfree(env->cfg.postorder_nums);
 	kvfree(env->cfg.insn_postorder);
 	kvfree(env->scc_info);
 	kvfree(env->succ);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 17/43] bpf: move the live-register and SCC printout to a standalone function
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (15 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 16/43] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
@ 2026-10-04 13:37 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 18/43] bpf: compute immediate dominators Eduard Zingerman
                   ` (27 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:37 UTC (permalink / raw)
  To: bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman,
	Emil Tsalapatis

Tests for multiple analyses performed by the verifier need the
verifier log to contain analysis results alongside the program
disassembly. This patch moves the log-level-2 program dump from
bpf_compute_live_registers() to a standalone function called from
bpf_check(), in order to provide a common logging function for such
analyses (and thus avoid printing program disassembly multiple times).

Rename the dump header from "Live regs before insn:" to
"Program dump (scc? insn#: live_regs_before):" and print available
source-line information before the corresponding instructions.

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c                              | 28 +---------------
 kernel/bpf/verifier.c                              | 37 ++++++++++++++++++++++
 .../selftests/bpf/progs/verifier_live_stack.c      |  2 +-
 3 files changed, 39 insertions(+), 28 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index ce74e3f16eef..d5cee598f4c5 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -2624,8 +2624,7 @@ int bpf_compute_live_registers(struct bpf_verifier_env *env)
 	struct bpf_insn *insns = env->prog->insnsi;
 	struct insn_live_regs *state;
 	int insn_cnt = env->prog->len;
-	u64 pos, insn_pos;
-	int err = 0, i, j, subprog, start, end;
+	int err = 0, i, subprog, start, end;
 	bool changed, ret_reg_pair;
 
 	/* Use the following algorithm:
@@ -2704,31 +2703,6 @@ int bpf_compute_live_registers(struct bpf_verifier_env *env)
 		insn_aux[i].zext_dst = def32 >= 0 && (mask_hi(out) & BIT(def32));
 	}
 
-	if (env->log.level & BPF_LOG_LEVEL2) {
-		verbose(env, "Live regs before insn:\n");
-		for (i = 0; i < insn_cnt; ++i) {
-			if (env->insn_aux_data[i].scc)
-				verbose(env, "%3d ", env->insn_aux_data[i].scc);
-			else
-				verbose(env, "    ");
-			verbose(env, "%3d: ", i);
-			for (j = BPF_REG_0; j < BPF_REG_10; ++j)
-				if (insn_aux[i].live_regs_before & BIT(j))
-					verbose(env, "%d", j);
-				else
-					verbose(env, ".");
-			verbose(env, " ");
-			pos = env->log.end_pos;
-			bpf_verbose_insn(env, &insns[i]);
-			insn_pos = env->log.end_pos;
-			if (insn_aux[i].zext_dst)
-				verbose(env, "%*c; zext", bpf_vlog_alignment(insn_pos - pos), ' ');
-			verbose(env, "\n");
-			if (bpf_is_ldimm64(&insns[i]))
-				i++;
-		}
-	}
-
 out:
 	kvfree(state);
 	return err;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 3298e2192239..c5a3b0f75742 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22736,6 +22736,40 @@ static bool bpf_prog_reenters_datapath(const struct bpf_prog *prog)
 	return false;
 }
 
+/* Various log level 2 information about the program */
+static void log_program(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *insn_aux = env->insn_aux_data;
+	struct bpf_insn *insns = env->prog->insnsi;
+	u32 insn_cnt = env->prog->len;
+	u64 pos, insn_pos;
+	u32 i, j;
+
+	verbose(env, "Program dump (scc? insn#: live_regs_before):\n");
+	for (i = 0; i < insn_cnt; ++i) {
+		verbose_linfo(env, i, "    ; ");
+		if (env->insn_aux_data[i].scc)
+			verbose(env, "%3d ", env->insn_aux_data[i].scc);
+		else
+			verbose(env, "    ");
+		verbose(env, "%3d: ", i);
+		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
+			if (insn_aux[i].live_regs_before & BIT(j))
+				verbose(env, "%d", j);
+			else
+				verbose(env, ".");
+		verbose(env, " ");
+		pos = env->log.end_pos;
+		bpf_verbose_insn(env, &insns[i]);
+		insn_pos = env->log.end_pos;
+		if (insn_aux[i].zext_dst)
+			verbose(env, "%*c; zext", bpf_vlog_alignment(insn_pos - pos), ' ');
+		verbose(env, "\n");
+		if (bpf_is_ldimm64(&insns[i]))
+			i++;
+	}
+}
+
 int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	      struct bpf_log_attr *attr_log)
 {
@@ -22944,6 +22978,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	if (env->log.level & BPF_LOG_LEVEL2)
+		log_program(env);
+
 	ret = mark_fastcall_patterns(env);
 	if (ret < 0)
 		goto skip_full_check;
diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 3ba21430bb75..0e6f25a09b5d 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -66,7 +66,7 @@ __msg("stack use/def subprog#0 must_write_not_same_slot (d0,cs0):")
  * but the write conservatively marks the whole frame as may_def.
  */
 __msg("6: (7b) *(u64 *)(r2 +0) = r0         ; may_def: fp0-8..-{{(512|2048)}}")
-__msg("Live regs before insn:")
+__msg("Program dump")
 __naked void must_write_not_same_slot(void)
 {
 	asm volatile (

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 18/43] bpf: compute immediate dominators
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (16 preceding siblings ...)
  2026-10-04 13:37 ` [PATCH bpf-next v2 17/43] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:50   ` sashiko-bot
  2026-10-04 13:38 ` [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy Eduard Zingerman
                   ` (26 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs to find an exit condition that dominates a loop's backedge;
such conditions bound every loop iteration. Compute the immediate
dominator tree as a prerequisite for identifying such conditions.

Use the iterative reverse-postorder algorithm from:
"A Simple, Fast Dominance Algorithm" by Cooper et al.

Analyze each subprogram independently and exclude the second half of
ldimm64 instructions.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |   8 +++
 kernel/bpf/Makefile          |   2 +-
 kernel/bpf/loops.c           | 156 +++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |   8 ++-
 4 files changed, 172 insertions(+), 2 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 14813040df84..0a452de50b46 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -650,6 +650,11 @@ struct bpf_iarray {
 	u32 items[];
 };
 
+#define iarray_for_each(item, arr)						\
+	for (int ___idx = 0;							\
+	     ___idx < (arr)->cnt && ({ item = (arr)->items[___idx]; 1; });	\
+	     ___idx++)
+
 struct bpf_insn_aux_data {
 	union {
 		enum bpf_reg_type ptr_type;	/* pointer type for load/store insns */
@@ -1112,6 +1117,7 @@ struct bpf_verifier_env {
 	u32 scc_cnt;
 	struct bpf_iarray *succ;
 	struct bpf_iarray *gotox_tmp_buf;
+	int *idoms;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1883,4 +1889,6 @@ int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 int bpf_flip_opcode(u32 opcode);
 u8 bpf_rev_opcode(u8 opcode);
 
+int bpf_compute_idoms(struct bpf_verifier_env *env);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..6210601eb40b 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
new file mode 100644
index 000000000000..0a9e10edca87
--- /dev/null
+++ b/kernel/bpf/loops.c
@@ -0,0 +1,156 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/slab.h>
+#include <linux/sched/signal.h>
+#include <linux/bpf_verifier.h>
+
+static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
+{
+	struct bpf_iarray *succ, *preds, **result;
+	struct bpf_prog *prog = env->prog;
+	u32 *num_preds, i, s, sz, len = prog->len;
+	struct bpf_insn *insn;
+	void *tmp;
+
+	num_preds = kvcalloc(prog->len, sizeof(u32), GFP_KERNEL_ACCOUNT);
+	if (!num_preds)
+		return NULL;
+
+	/*
+	 * 'result' layout:
+	 *  - array of pointers (struct bpf_iarray *)[len]
+	 *  - struct bpf_iarray one after another
+	 */
+	sz = sizeof(struct bpf_iarray) * len;
+	sz += sizeof(struct bpf_iarray *) * len;
+	for (i = 0; i < len; i++) {
+		insn = env->prog->insnsi + i;
+		succ = bpf_insn_successors(env, i);
+		sz += sizeof(u32) * succ->cnt;
+		iarray_for_each(s, succ) {
+			num_preds[s]++;
+		}
+		if (bpf_is_ldimm64(insn))
+			i++;
+	}
+
+	result = kvzalloc(sz, GFP_KERNEL_ACCOUNT);
+	if (!result) {
+		kvfree(num_preds);
+		return NULL;
+	}
+
+	tmp = (void *)&result[len];
+	for (i = 0; i < len; i++) {
+		result[i] = tmp;
+		tmp += sizeof(struct bpf_iarray);
+		tmp += sizeof(u32) * num_preds[i];
+	}
+
+	for (i = 0; i < len; i++) {
+		insn = env->prog->insnsi + i;
+		succ = bpf_insn_successors(env, i);
+		iarray_for_each(s, succ) {
+			preds = result[s];
+			preds->items[preds->cnt++] = i;
+		}
+		if (bpf_is_ldimm64(insn))
+			i++;
+	}
+
+	kvfree(num_preds);
+	return result;
+}
+
+static int idoms_intersect(struct bpf_verifier_env *env, int a, int b)
+{
+	int *postorder_nums = env->cfg.postorder_nums;
+	int *idoms = env->idoms;
+
+	while (a != b) {
+		while (postorder_nums[a] < postorder_nums[b]) {
+			a = idoms[a];
+		}
+		while (postorder_nums[b] < postorder_nums[a]) {
+			b = idoms[b];
+		}
+	}
+	return a;
+}
+
+/* See "A Simple, Fast Dominance Algorithm" by Cooper et al. for details. */
+static int compute_subprog_idoms(struct bpf_verifier_env *env, struct bpf_iarray **preds,
+				int subprog_idx)
+{
+	struct bpf_subprog_info *subprog = &env->subprog_info[subprog_idx];
+	int start = subprog->start;
+	int po_first = subprog->postorder_start;
+	int po_last = (subprog + 1)->postorder_start - 1;
+	int *idoms = env->idoms;
+	int po_num, pred;
+	u32 work = 0;
+	bool changed;
+
+	idoms[start] = 0;
+	changed = true;
+	do {
+		changed = false;
+		/* iterate in reverse postorder */
+		for (po_num = po_last; po_num >= po_first; po_num--) {
+			int idx = env->cfg.insn_postorder[po_num];
+			int new_idom = -1;
+
+			iarray_for_each(pred, preds[idx]) {
+				if (++work % 1024 == 0) {
+					if (signal_pending(current))
+						return -EAGAIN;
+					cond_resched();
+				}
+
+				if (idoms[pred] == -1)
+					continue;
+				if (new_idom == -1)
+					new_idom = pred;
+				else
+					new_idom = idoms_intersect(env, pred, new_idom);
+			}
+			if (new_idom != -1 && idoms[idx] != new_idom) {
+				idoms[idx] = new_idom;
+				changed = true;
+			}
+		}
+	} while (changed);
+	idoms[start] = -1;
+	return 0;
+}
+
+int bpf_compute_idoms(struct bpf_verifier_env *env)
+{
+	struct bpf_iarray **preds;
+	u32 len = env->prog->len;
+	int *idoms, i, err = 0;
+
+	preds = compute_predecessors(env);
+	if (!preds)
+		return -ENOMEM;
+
+	idoms = kvcalloc(len, sizeof(*idoms), GFP_KERNEL_ACCOUNT);
+	if (!idoms) {
+		kvfree(preds);
+		return -ENOMEM;
+	}
+
+	env->idoms = idoms;
+	for (i = 0; i < len; i++)
+		idoms[i] = -1;
+
+	for (i = 0; i < env->subprog_cnt; i++) {
+		err = compute_subprog_idoms(env, preds, i);
+		if (err)
+			break;
+	}
+
+	kvfree(preds);
+	return err;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index c5a3b0f75742..770091306032 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22745,13 +22745,14 @@ static void log_program(struct bpf_verifier_env *env)
 	u64 pos, insn_pos;
 	u32 i, j;
 
-	verbose(env, "Program dump (scc? insn#: live_regs_before):\n");
+	verbose(env, "Program dump (scc? idom insn#: live_regs_before):\n");
 	for (i = 0; i < insn_cnt; ++i) {
 		verbose_linfo(env, i, "    ; ");
 		if (env->insn_aux_data[i].scc)
 			verbose(env, "%3d ", env->insn_aux_data[i].scc);
 		else
 			verbose(env, "    ");
+		verbose(env, "%3d ", env->idoms[i]);
 		verbose(env, "%3d: ", i);
 		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
 			if (insn_aux[i].live_regs_before & BIT(j))
@@ -22974,6 +22975,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	ret = bpf_compute_idoms(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = bpf_compute_live_registers(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -23137,6 +23142,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	kvfree(env->callx_edges);
 	kvfree(env->func_ptrs);
 	bpf_diag_free(env);
+	kvfree(env->idoms);
 	kvfree(env);
 	return ret;
 }

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (17 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 18/43] bpf: compute immediate dominators Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:40   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 20/43] bpf: add bpf_set_reg_range() Eduard Zingerman
                   ` (25 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs to process loops from innermost to outermost, summarizing
results for nested loops. Add a pass to compute the loop structure.

Use a non-recursive adaptation of the algorithm from:
"A New Algorithm for Identifying Loops in Decompilation" by Wei et al.

Record the loop hierarchy per instruction:
- The field insn_aux_data->loop_header records the innermost loop
  header containing it.
- The field insn_aux_data->loop records additional information about
  the loop if this instruction is a loop header (exits, backedges,
  irreducibility).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  42 +++++
 kernel/bpf/fixups.c          |  18 ++
 kernel/bpf/loops.c           | 431 +++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |  12 +-
 4 files changed, 502 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 0a452de50b46..367aad258050 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -655,6 +655,30 @@ struct bpf_iarray {
 	     ___idx < (arr)->cnt && ({ item = (arr)->items[___idx]; 1; });	\
 	     ___idx++)
 
+#define MAX_BACKEDGES 16
+#define MAX_LOOP_EXITS 256
+
+struct bpf_backedge {
+	int from;
+	int latch; /* -1 if no latch can be found */
+};
+
+struct bpf_loop_exit {
+	int from; /* instruction inside the loop */
+	int to; /* instruction outside the loop */
+};
+
+struct bpf_loop {
+	struct bpf_backedge backedges[MAX_BACKEDGES];
+	/* edges exiting from this loop, includes edges from nested loops */
+	struct bpf_loop_exit *exits;
+	int backedges_cnt;
+	int exits_cnt;
+	bool irreducible;
+	bool backedges_overflow;
+	bool exits_overflow;
+};
+
 struct bpf_insn_aux_data {
 	union {
 		enum bpf_reg_type ptr_type;	/* pointer type for load/store insns */
@@ -733,6 +757,13 @@ struct bpf_insn_aux_data {
 	u32 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
 	u32 jump_target:1;
+	/*
+	 * True if this instruction is a loop entry. For loops with a single entry
+	 * this bit will coincide with 'loop' pointer being non-NULL.
+	 * Irreducible loops have multiple entries, all of them will be marked as 'loop_entry',
+	 * but only the one at 'loop_header' will have a non-NULL 'loop' pointer.
+	 */
+	u32 loop_entry:1;
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
@@ -752,6 +783,13 @@ struct bpf_insn_aux_data {
 	u16 const_reg_map_mask;
 	u16 const_reg_subprog_mask;
 	u32 const_reg_vals[10];
+	/*
+	 * Header index of the innermost loop containing this instruction, -1 if none.
+	 * For a loop header, identifies the parent loop instead.
+	 */
+	s32 loop_header;
+	/* additional information about the loop if this instruction is a loop header */
+	struct bpf_loop *loop;
 };
 
 #define MAX_USED_MAPS 64 /* max number of maps accessed by one eBPF program */
@@ -1874,6 +1912,7 @@ int bpf_check_attach_btf_id_multi(struct btf *btf, struct bpf_prog *prog, u32 bt
 				  struct bpf_attach_target_info *tgt_info);
 
 /* Functions in fixups.c, called from bpf_check() */
+void bpf_clear_insn_aux_data(struct bpf_verifier_env *env, int start, int len);
 int bpf_remove_fastcall_spills_fills(struct bpf_verifier_env *env);
 int bpf_optimize_bpf_loop(struct bpf_verifier_env *env);
 void bpf_opt_hard_wire_dead_code_branches(struct bpf_verifier_env *env);
@@ -1890,5 +1929,8 @@ int bpf_flip_opcode(u32 opcode);
 u8 bpf_rev_opcode(u8 opcode);
 
 int bpf_compute_idoms(struct bpf_verifier_env *env);
+int bpf_compute_loops(struct bpf_verifier_env *env);
+int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx);
+bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header);
 
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 206cc9a614a9..b2fa7e8bedf3 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -590,6 +590,22 @@ static int bpf_adj_linfo_after_remove(struct bpf_verifier_env *env, u32 off,
 	return 0;
 }
 
+/* Clean up dynamically allocated fields of aux data for instructions [start, ...] */
+void bpf_clear_insn_aux_data(struct bpf_verifier_env *env, int start, int len)
+{
+	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
+	int end = start + len;
+	int i;
+
+	for (i = start; i < end; i++) {
+		if (aux_data[i].loop) {
+			kvfree(aux_data[i].loop->exits);
+			kvfree(aux_data[i].loop);
+			aux_data[i].loop = NULL;
+		}
+	}
+}
+
 static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 {
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -602,6 +618,8 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 	if (bpf_prog_is_offloaded(env->prog->aux))
 		bpf_prog_offload_remove_insns(env, off, cnt);
 
+	bpf_clear_insn_aux_data(env, off, cnt);
+
 	err = bpf_remove_insns(env->prog, off, cnt);
 	if (err)
 		return err;
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
index 0a9e10edca87..545cee0fd956 100644
--- a/kernel/bpf/loops.c
+++ b/kernel/bpf/loops.c
@@ -154,3 +154,434 @@ int bpf_compute_idoms(struct bpf_verifier_env *env)
 	kvfree(preds);
 	return err;
 }
+
+struct dfs_state {
+	u32 traversed:1;
+	u32 next_succ:31;
+};
+
+struct loops_dfs {
+	struct dfs_state *state;
+	int *dfs_pos;
+	int *stack;
+};
+
+static void mark_irreducible(struct bpf_verifier_env *env, int h)
+{
+	env->insn_aux_data[h].loop->irreducible = true;
+}
+
+static void mark_entry(struct bpf_verifier_env *env, int s)
+{
+	env->insn_aux_data[s].loop_entry = true;
+}
+
+static void add_backedge(struct bpf_verifier_env *env, int from, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop = aux[h].loop;
+	int cnt = loop->backedges_cnt;
+
+	if (cnt == MAX_BACKEDGES) {
+		loop->backedges_overflow = true;
+		return;
+	}
+	loop->backedges[cnt].from = from;
+	loop->backedges[cnt].latch = -1;
+	loop->backedges_cnt++;
+}
+
+static int add_exit(struct bpf_loop *loop, int from, int to)
+{
+	if (loop->exits_overflow)
+		return 0;
+	if (loop->exits_cnt == MAX_LOOP_EXITS) {
+		loop->exits_overflow = true;
+		return 0;
+	}
+	if (!loop->exits) {
+		loop->exits = kvcalloc(MAX_LOOP_EXITS, sizeof(*loop->exits), GFP_KERNEL_ACCOUNT);
+		if (!loop->exits)
+			return -ENOMEM;
+	}
+	loop->exits[loop->exits_cnt] = (struct bpf_loop_exit) {
+		.from = from,
+		.to = to,
+	};
+	loop->exits_cnt++;
+	return 0;
+}
+
+static int mark_as_header(struct bpf_verifier_env *env, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+
+	if (!aux[h].loop) {
+		mark_entry(env, h);
+		aux[h].loop = kvzalloc_obj(struct bpf_loop, GFP_KERNEL_ACCOUNT);
+		if (!aux[h].loop)
+			return -ENOMEM;
+	}
+	return 0;
+}
+
+static int assign_header(struct bpf_verifier_env *env, struct loops_dfs *dfs, int n, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int *dfs_pos = dfs->dfs_pos;
+	int err, nh;
+
+	err = mark_as_header(env, h);
+	if (err)
+		return err;
+
+	/* Don't encode self-loops, otherwise can't reflect loops nesting structure. */
+	if (n == h)
+		return 0;
+
+	/* Make sure that loop headers up the chain are sorted by dfs_pos. */
+	while (aux[n].loop_header != -1) {
+		nh = aux[n].loop_header;
+		if (nh == h)
+			return 0;
+		if (dfs_pos[nh] < dfs_pos[h]) {
+			aux[n].loop_header = h;
+			n = h;
+			h = nh;
+		} else {
+			n = nh;
+		}
+	}
+	aux[n].loop_header = h;
+	return 0;
+}
+
+static bool is_cond_jmp_insn(struct bpf_insn *insn)
+{
+	u8 class = BPF_CLASS(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+
+	if (class != BPF_JMP && class != BPF_JMP32)
+		return false;
+
+	switch (opcode) {
+	case BPF_JEQ:
+	case BPF_JGE:
+	case BPF_JGT:
+	case BPF_JLE:
+	case BPF_JLT:
+	case BPF_JNE:
+	case BPF_JSET:
+	case BPF_JSGE:
+	case BPF_JSGT:
+	case BPF_JSLE:
+	case BPF_JSLT:
+		return true;
+	default:
+		return false;
+	}
+}
+
+int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+
+	return aux[idx].loop ? idx : aux[idx].loop_header;
+}
+
+static int find_dominating_condition(struct bpf_verifier_env *env, int n, int top)
+{
+	struct bpf_insn *insns = env->prog->insnsi;
+	int common_dom, n_loop, t_loop, f_loop;
+	int *idoms = env->idoms;
+
+	n_loop = bpf_loop_at_index(env, n);
+	common_dom = idoms_intersect(env, n, top);
+	if (common_dom != top)
+		return -1;
+	while (n >= 0) {
+		if (is_cond_jmp_insn(&insns[n]) && bpf_loop_at_index(env, n) == n_loop) {
+			t_loop = bpf_loop_at_index(env, n + insns[n].off + 1);
+			f_loop = bpf_loop_at_index(env, n + 1);
+			if (f_loop != n_loop && !bpf_is_nested_loop(env, f_loop, n_loop))
+				return n;
+			if (t_loop != n_loop && !bpf_is_nested_loop(env, t_loop, n_loop))
+				return n;
+		}
+		if (n == top)
+			break;
+		n = idoms[n];
+	}
+	return -1;
+}
+
+/*
+ * As described in "A New Algorithm for Identifying Loops in Decompilation" by Wei et al,
+ * adapted to be non-recursive.
+ */
+static int compute_loops_in_subprog(struct bpf_verifier_env *env, struct loops_dfs *dfs,
+				    int subprog_idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct dfs_state *state = dfs->state;
+	int start = env->subprog_info[subprog_idx].start;
+	int *dfs_pos = dfs->dfs_pos;
+	int *stack = dfs->stack;
+	int s, h, err, cur, stack_sz;
+	struct bpf_iarray *succ;
+	u32 i;
+
+	stack[0] = start;
+	state[start].traversed = true;
+	state[start].next_succ = 0;
+	dfs_pos[start] = 1;
+	stack_sz = 1;
+	i = 0;
+	do {
+		/*
+		 * The algorithm should be very fast in practice,
+		 * guard against pathological inputs, just in case.
+		 */
+		if ((++i % 1024) == 0) {
+			if (signal_pending(current))
+				return -EAGAIN;
+			cond_resched();
+		}
+
+		cur = stack[stack_sz - 1];
+		succ = bpf_insn_successors(env, cur);
+		if (state[cur].next_succ == succ->cnt) {
+			dfs_pos[cur] = 0;
+			stack_sz--;
+			continue;
+		}
+		s = succ->items[state[cur].next_succ];
+		if (!state[s].traversed) {
+			/* Case A:  start -> ... -> cur -> s [unexplored] */
+			state[s].traversed = true;
+			state[s].next_succ = 0;
+			stack[stack_sz] = s;
+			dfs_pos[s] = stack_sz + 1;
+			stack_sz++;
+			continue;
+		}
+		/* 's' is fully explored at this point */
+		if (dfs_pos[s]) {
+			/*
+			 * start -> ... -> s -> cur --.
+			 *                 ^          |
+			 *                 '----------'
+			 * Case B: 's' is in the current DFS path.
+			 */
+			err = assign_header(env, dfs, cur, s);
+			if (err)
+				return err;
+			add_backedge(env, cur, s);
+		} else if (aux[s].loop_header == -1) {
+			/*
+			 * start -> ... -> ... -> s -> ... -> end
+			 *           |            ^
+			 *           '---> cur ---'
+			 * Case C: 's' is explored, not in the current DFS path,
+			 * and not a part of any loop.
+			 */
+		} else if (dfs_pos[aux[s].loop_header]) {
+			/*
+			 *                 .----------------------.
+			 *                 v                      |
+			 * start -> ... -> h -> ... -> ... -> s --'
+			 *                       |            ^
+			 *	                 '---> cur ---'
+			 * Case D: 's' is explored, not in current DFS path,
+			 * but its innermost loop header is.
+			 */
+			err = assign_header(env, dfs, cur, aux[s].loop_header);
+			if (err)
+				return err;
+		} else {
+			/*
+			 * case E: 's' is explored, not in current DFS path,
+			 * its innermost loop header is not in current DFS path,
+			 * hence 's' is another entry into the same loop.
+			 */
+			h = aux[s].loop_header;
+			mark_irreducible(env, h);
+			mark_entry(env, s);
+			while (aux[h].loop_header != -1) {
+				h = aux[h].loop_header;
+				if (dfs_pos[h]) {
+					err = assign_header(env, dfs, cur, h);
+					if (err)
+						return err;
+					break;
+				}
+				mark_irreducible(env, h);
+			}
+		}
+		state[cur].next_succ++;
+	} while (stack_sz);
+
+	return 0;
+}
+
+bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int idx;
+
+	for (idx = inner_header; idx >= 0; idx = aux[idx].loop_header)
+		if (aux[idx].loop_header == outer_header)
+			return true;
+
+	return false;
+}
+
+int bpf_compute_loops(struct bpf_verifier_env *env)
+{
+	int i, j, s, t, latch, iloop, sloop, err = 0, len = env->prog->len;
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct bpf_backedge *backedge;
+	struct loops_dfs dfs = {};
+	struct bpf_iarray *succ;
+	struct bpf_loop *loop;
+
+	dfs.dfs_pos = kvcalloc(len, sizeof(int), GFP_KERNEL_ACCOUNT);
+	dfs.state = kvcalloc(len, sizeof(struct dfs_state), GFP_KERNEL_ACCOUNT);
+	dfs.stack = kvcalloc(len, sizeof(int), GFP_KERNEL_ACCOUNT);
+	if (!dfs.dfs_pos || !dfs.state || !dfs.stack) {
+		err = -ENOMEM;
+		goto out;
+	}
+	for (i = 0; i < len; i++)
+		aux[i].loop_header = -1;
+	for (i = 0; i < env->subprog_cnt; i++) {
+		err = compute_loops_in_subprog(env, &dfs, i);
+		if (err)
+			goto out;
+	}
+	/* find latches */
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			backedge = &loop->backedges[j];
+			/*
+			 * In theory, the backedge->from and it's latch can reside in
+			 * an inner loop of `i`, e.g.:
+			 *
+			 *  1:  r7 += 1;
+			 *  2:  r6 += 1;
+			 *      if r7 == 2 goto 1b;
+			 *      if r6 < 2 goto 2b;
+			 *
+			 * For now, let's assume that latches are unknown for such cases.
+			 * (GCC/LLVM handle this by inserting artificial cfg nodes).
+			 */
+			if (bpf_loop_at_index(env, backedge->from) != i)
+				continue;
+			latch = find_dominating_condition(env, backedge->from, i);
+			if (latch < 0 || bpf_loop_at_index(env, latch) != i)
+				continue;
+			backedge->latch = latch;
+		}
+	}
+	/* find exits */
+	for (i = 0; i < len; i++) {
+		iloop = aux[i].loop ? i : aux[i].loop_header;
+		if (iloop < 0)
+			continue;
+		succ = bpf_insn_successors(env, i);
+		iarray_for_each(s, succ) {
+			/*
+			 * Nothing left to record once the innermost loop of 'i'
+			 * overflowed: the walk below would stop at it right away.
+			 */
+			if (aux[iloop].loop->exits_overflow)
+				break;
+			sloop = aux[s].loop ? s : aux[s].loop_header;
+			if (iloop == sloop)
+				continue;
+			if (bpf_is_nested_loop(env, sloop, iloop))
+				continue;
+			/*
+			 * At this point 'sloop' is either -1, an outer loop,
+			 * or a loop in another branch of the loop hierarchy.
+			 */
+			for (t = iloop; t >= 0 && t != sloop; t = aux[t].loop_header) {
+				/*
+				 * Account for the following configuration:
+				 *
+				 *   tloop {
+				 *     iloop {
+				 *       ... i: goto s;
+				 *     }
+				 *     sloop {
+				 *   s:
+				 *       ...
+				 *     }
+				 *   }
+				 */
+				if (bpf_is_nested_loop(env, sloop, t))
+					break;
+				/*
+				 * Record edges i -> s as exits from tloop when:
+				 *
+				 *   sloop {
+				 *     tloop {
+				 *       iloop {
+				 *         ... i: goto s;
+				 *       }
+				 *     }
+				 *  s: ...
+				 *   }
+				 */
+				/* Enclosing loops inherit the overflow, see below */
+				if (aux[t].loop->exits_overflow)
+					break;
+				err = add_exit(aux[t].loop, i, s);
+				if (err)
+					goto out;
+			}
+		}
+		if (bpf_is_ldimm64(env->prog->insnsi + i))
+			i++;
+	}
+	/* A nested loop with a truncated exit list can't be abstracted by SCEV */
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop || !loop->exits_overflow)
+			continue;
+		for (t = aux[i].loop_header; t >= 0; t = aux[t].loop_header)
+			aux[t].loop->exits_overflow = true;
+	}
+
+	if (env->log.level & BPF_LOG_LEVEL2) {
+		for (i = 0; i < len; i++) {
+			loop = aux[i].loop;
+			if (!loop)
+				continue;
+			bpf_log(log, "loop at %d", i);
+			if (aux[i].loop_header >= 0)
+				bpf_log(log, ", nested in %d", aux[i].loop_header);
+			if (loop->irreducible)
+				bpf_log(log, ", irreducible");
+			if (loop->exits_overflow)
+				bpf_log(log, ", too many exits");
+			bpf_log(log, "\n");
+			for (j = 0; j < loop->backedges_cnt; j++)
+				bpf_log(log, "  backedge from %d, latch at %d\n",
+					loop->backedges[j].from, loop->backedges[j].latch);
+			for (j = 0; j < loop->exits_cnt; j++)
+				bpf_log(log, "  exit from %d to %d\n",
+					loop->exits[j].from, loop->exits[j].to);
+		}
+	}
+
+out:
+	kvfree(dfs.dfs_pos);
+	kvfree(dfs.stack);
+	kvfree(dfs.state);
+	return err;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 770091306032..c0907fad744b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22745,13 +22745,17 @@ static void log_program(struct bpf_verifier_env *env)
 	u64 pos, insn_pos;
 	u32 i, j;
 
-	verbose(env, "Program dump (scc? idom insn#: live_regs_before):\n");
+	verbose(env, "Program dump (scc? loop_header? idom insn#: live_regs_before):\n");
 	for (i = 0; i < insn_cnt; ++i) {
 		verbose_linfo(env, i, "    ; ");
 		if (env->insn_aux_data[i].scc)
 			verbose(env, "%3d ", env->insn_aux_data[i].scc);
 		else
 			verbose(env, "    ");
+		if (env->insn_aux_data[i].loop_header >= 0)
+			verbose(env, "%3d ", env->insn_aux_data[i].loop_header);
+		else
+			verbose(env, "    ");
 		verbose(env, "%3d ", env->idoms[i]);
 		verbose(env, "%3d: ", i);
 		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
@@ -22979,6 +22983,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	ret = bpf_compute_loops(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = bpf_compute_live_registers(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -23131,6 +23139,8 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	release_btfs(env);
 err_free_env:
 	bpf_free_subprog_jts(env);
+	if (env->insn_aux_data)
+		bpf_clear_insn_aux_data(env, 0, env->insn_aux_data_len);
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
 	bpf_stack_liveness_free(env);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 20/43] bpf: add bpf_set_reg_range()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (18 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 21/43] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
                   ` (24 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

A utility function for setting a register's range and step information.
Used by SCEV widening and clamping logic further in the series.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/verifier.c        | 12 ++++++++++++
 2 files changed, 14 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 367aad258050..e1c0d0d3f1e1 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1811,6 +1811,8 @@ struct arg_access_info
 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 			     struct bpf_insn *insn, int arg, int insn_idx);
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
+int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
+		      struct cnum64 range, u16 base, u16 step);
 
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index c0907fad744b..22369b6f747d 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -17146,6 +17146,18 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 	return 0;
 }
 
+int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
+		      struct cnum64 range, u16 base, u16 step)
+{
+	reg->r64 = range;
+	reg->r32 = CNUM32_UNBOUNDED;
+	reg->step = step;
+	reg->base = base;
+	reg->var_off = tnum_unknown;
+	reg_bounds_sync(reg); /* this should infer the tnum alignment */
+	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
+}
+
 /* check validity of 32-bit and 64-bit arithmetic operations */
 static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 21/43] bpf: add bpf_mark_reg_known_scalar()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (19 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 20/43] bpf: add bpf_set_reg_range() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union() Eduard Zingerman
                   ` (23 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Provide a helper to set a register to a known scalar value for use outside
verifier.c. To be used by SCEV widening logic.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 1 +
 kernel/bpf/verifier.c        | 6 ++++++
 2 files changed, 7 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index e1c0d0d3f1e1..73bbd7f8f1cc 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1348,6 +1348,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
 void bpf_bt_sync_linked_regs(struct backtrack_state *bt, struct bpf_jmp_history_entry *hist);
 void bpf_mark_reg_not_init(const struct bpf_verifier_env *env,
 			   struct bpf_reg_state *reg);
+void bpf_mark_reg_known_scalar(struct bpf_reg_state *reg, u64 imm);
 void bpf_mark_reg_unknown_imprecise(struct bpf_reg_state *reg);
 void bpf_mark_all_scalars_precise(struct bpf_verifier_env *env,
 				  struct bpf_verifier_state *st);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 22369b6f747d..5fd6376da171 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -2049,6 +2049,12 @@ static void __mark_reg_known(struct bpf_reg_state *reg, u64 imm)
 	___mark_reg_known(reg, imm);
 }
 
+void bpf_mark_reg_known_scalar(struct bpf_reg_state *reg, u64 imm)
+{
+	__mark_reg_known(reg, imm);
+	reg->type = SCALAR_VALUE;
+}
+
 static void __mark_reg32_known(struct bpf_reg_state *reg, u64 imm)
 {
 	reg->var_off = tnum_const_subreg(reg->var_off, imm);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (20 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 21/43] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:59   ` sashiko-bot
  2026-10-04 13:38 ` [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
                   ` (22 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

A utility function to take the union of two registers' scalar values:
merge the circular 32-bit and 64-bit bounds, tnums and base/step
of two registers, then synchronize and validate the result.
Leave id invalidation and type validation to the caller.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/verifier.c        | 28 ++++++++++++++++++++++++++++
 2 files changed, 30 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 73bbd7f8f1cc..555a41c5660e 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1814,6 +1814,8 @@ bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
 int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 		      struct cnum64 range, u16 base, u16 step);
+int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
+		  const struct bpf_reg_state *src);
 
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 5fd6376da171..28c16e35b1f3 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -31,6 +31,7 @@
 #include <linux/module.h>
 #include <linux/cpumask.h>
 #include <linux/cnum.h>
+#include <linux/gcd.h>
 #include <linux/bpf_mem_alloc.h>
 #include <net/xdp.h>
 #include <linux/trace_events.h>
@@ -17164,6 +17165,33 @@ int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
 }
 
+/* acc := acc U src, matching types only. Caller must clear acc's scalar ID. */
+int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
+		  const struct bpf_reg_state *src)
+{
+	u16 base, step;
+
+	if (acc->type != src->type) {
+		verifier_bug(env, "union of registers with different types");
+		return -EFAULT;
+	}
+	acc->r64 = cnum64_union(acc->r64, src->r64);
+	acc->r32 = cnum32_union(acc->r32, src->r32);
+	acc->var_off = tnum_union(acc->var_off, src->var_off);
+
+	/* Retain a common congruence if the bases agree modulo the gcd. */
+	step = gcd(acc->step, src->step);
+	base = acc->base % step;
+	if (base != src->base % step) {
+		reg_step_reset(acc);
+	} else {
+		acc->base = base;
+		acc->step = step;
+	}
+	reg_bounds_sync(acc);
+	return reg_bounds_sanity_check(env, acc, "bpf_reg_union");
+}
+
 /* check validity of 32-bit and 64-bit arithmetic operations */
 static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (21 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:47   ` sashiko-bot
  2026-10-06 16:53   ` Alexei Starovoitov
  2026-10-04 13:38 ` [PATCH bpf-next v2 24/43] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
                   ` (21 subsequent siblings)
  44 siblings, 2 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV needs to process pending basic blocks in reverse postorder.
Provide a small integer min-heap whose comparator can order instruction
indices by their CFG ranks.

Grow the backing array on demand and pass caller context to the
comparator, so the heap need not encode the analysis-specific order.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 17 +++++++++
 kernel/bpf/Makefile          |  2 +-
 kernel/bpf/heap.c            | 87 ++++++++++++++++++++++++++++++++++++++++++++
 3 files changed, 105 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 555a41c5660e..8f5d35fa9aab 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1938,4 +1938,21 @@ int bpf_compute_loops(struct bpf_verifier_env *env);
 int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx);
 bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header);
 
+/*
+ * Simple binary heap implementation as described by
+ * https://en.wikipedia.org/wiki/Binary_heap
+ */
+struct bpf_min_heap {
+	int (*compare)(int, int, void *); /* ordering function for @elements */
+	int *elements; /* min-heap ordered by @compare */
+	void *arg; /* 3rd argument passed to @compare */
+	int capacity;
+	int count;
+};
+
+void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg);
+void bpf_min_heap_free(struct bpf_min_heap *heap);
+int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
+bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 6210601eb40b..7a2c179a9059 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/heap.c b/kernel/bpf/heap.c
new file mode 100644
index 000000000000..029f9217b872
--- /dev/null
+++ b/kernel/bpf/heap.c
@@ -0,0 +1,87 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_verifier.h>
+
+/* Indexes for binary tree encoded as an array */
+static inline int left_child(int i) { return 2 * i + 1; }
+static inline int right_child(int i) { return 2 * i + 2; }
+static inline int parent(int i) { return (i - 1) / 2; }
+
+static inline int greater(struct bpf_min_heap *heap, int a, int b)
+{
+	return heap->compare(a, b, heap->arg) > 0;
+}
+
+void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg)
+{
+	memset(heap, 0, sizeof(*heap));
+	heap->compare = compare;
+	heap->arg = arg;
+}
+
+void bpf_min_heap_free(struct bpf_min_heap *heap)
+{
+	kfree(heap->elements);
+	heap->elements = NULL;
+	heap->capacity = 0;
+	heap->count = 0;
+}
+
+int bpf_min_heap_push(struct bpf_min_heap *heap, int elt)
+{
+	int new_capacity, i;
+	int *elements;
+	void *tmp;
+
+	if (heap->count == heap->capacity) {
+		new_capacity = heap->capacity ? heap->capacity * 2 : 16;
+		tmp = krealloc(heap->elements,
+			       sizeof(*heap->elements) * new_capacity,
+			       GFP_KERNEL_ACCOUNT);
+		if (!tmp)
+			return -ENOMEM;
+		heap->elements = tmp;
+		heap->capacity = new_capacity;
+	}
+
+	elements = heap->elements;
+	i = heap->count;
+	elements[i] = elt;
+	heap->count++;
+	while (i != 0 && greater(heap, elements[parent(i)], elements[i])) {
+		swap(elements[i], elements[parent(i)]);
+		i = parent(i);
+	}
+	return 0;
+}
+
+static inline void sink_root(struct bpf_min_heap *heap)
+{
+	int *elements = heap->elements;
+	int i = 0;
+
+	while ((left_child(i)  < heap->count && greater(heap, elements[i], elements[left_child(i)])) ||
+	       (right_child(i) < heap->count && greater(heap, elements[i], elements[right_child(i)]))) {
+		if (right_child(i) >= heap->count || greater(heap, elements[right_child(i)], elements[left_child(i)])) {
+			swap(elements[i], elements[left_child(i)]);
+			i = left_child(i);
+		} else {
+			swap(elements[i], elements[right_child(i)]);
+			i = right_child(i);
+		}
+	}
+}
+
+bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt)
+{
+	if (heap->count == 0)
+		return false;
+
+	int *elements = heap->elements;
+	*elt = elements[0];
+	elements[0] = elements[heap->count - 1];
+	--heap->count;
+	sink_root(heap);
+	return true;
+}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 24/43] bpf: record basic-block ends in insn_aux_data
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (22 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 25/43] bpf: add bpf_split_cur_state() Eduard Zingerman
                   ` (20 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman,
	Eduard Zingerman

From: Eduard Zingerman <ezingerman@fb.com>

SCEV needs information about the program's basic-block structure.
Piggyback on compute_predecessors() and set
insn_aux_data->bb_end for an instruction if:
- it has multiple successors, or
- its sole successor has multiple predecessors, or
- it is a jump instruction.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/loops.c           | 19 +++++++++++++++++++
 2 files changed, 20 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 8f5d35fa9aab..8c00a78575a1 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -730,6 +730,7 @@ struct bpf_insn_aux_data {
 	bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
 	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
 	bool arena_scalar; /* ldx/stx/st/atomic through a number, it's an address in arena */
+	bool bb_end; /* last instruction of a basic block */
 	u8 alu_state; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
index 545cee0fd956..7a9167028012 100644
--- a/kernel/bpf/loops.c
+++ b/kernel/bpf/loops.c
@@ -5,8 +5,23 @@
 #include <linux/sched/signal.h>
 #include <linux/bpf_verifier.h>
 
+static bool is_cfg_jump(struct bpf_insn *insn)
+{
+	u8 class = BPF_CLASS(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+
+	switch (class) {
+	case BPF_JMP:
+	case BPF_JMP32:
+		return opcode != BPF_CALL;
+	default:
+		return false;
+	}
+}
+
 static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
 {
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
 	struct bpf_iarray *succ, *preds, **result;
 	struct bpf_prog *prog = env->prog;
 	u32 *num_preds, i, s, sz, len = prog->len;
@@ -55,6 +70,10 @@ static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
 			preds = result[s];
 			preds->items[preds->cnt++] = i;
 		}
+		if (succ->cnt > 1 ||
+		    (succ->cnt == 1 && num_preds[succ->items[0]] > 1) ||
+		    is_cfg_jump(insn))
+			aux[i].bb_end = true;
 		if (bpf_is_ldimm64(insn))
 			i++;
 	}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 25/43] bpf: add bpf_split_cur_state()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (23 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 24/43] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin() Eduard Zingerman
                   ` (19 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV widening needs an explicit checkpoint preserving the original
loop-entry state, independently of ordinary state-cache heuristics.
This commit extracts checkpoint creation logic from
bpf_is_state_visited() into a separate utility function.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/states.c          | 95 ++++++++++++++++++++++++--------------------
 2 files changed, 54 insertions(+), 42 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 8c00a78575a1..13c7825bb8b5 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1335,6 +1335,7 @@ void bpf_free_kfunc_btf_tab(struct bpf_kfunc_btf_tab *tab);
 int mark_chain_precision(struct bpf_verifier_env *env, int regno);
 
 int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx);
+int bpf_split_cur_state(struct bpf_verifier_env *env);
 int bpf_update_branch_counts(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
 
 void bpf_clear_jmp_history(struct bpf_verifier_state *state);
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 2f6a7620eb72..b1e2daefc264 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -1262,11 +1262,61 @@ static void mark_all_scalars_imprecise(struct bpf_verifier_env *env, struct bpf_
 	}
 }
 
-int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
+int bpf_split_cur_state(struct bpf_verifier_env *env)
 {
+	struct bpf_verifier_state *cur = env->cur_state, *new;
 	struct bpf_verifier_state_list *new_sl;
+	struct list_head *head;
+	int insn_idx = cur->insn_idx;
+	int err;
+
+	head = bpf_explored_state(env, insn_idx);
+	new_sl = kzalloc_obj(struct bpf_verifier_state_list, GFP_KERNEL_ACCOUNT);
+	if (!new_sl)
+		return -ENOMEM;
+	env->total_states++;
+	env->explored_states_size++;
+	update_peak_states(env);
+	env->prev_jmps_processed = env->jmps_processed;
+	env->prev_insn_processed = env->insn_processed;
+
+	/* forget precise markings we inherited, see __mark_chain_precision */
+	if (env->bpf_capable)
+		mark_all_scalars_imprecise(env, cur);
+
+	bpf_clear_singular_ids(env, cur);
+
+	/* add new state to the head of linked list */
+	new = &new_sl->state;
+	err = bpf_copy_verifier_state(new, cur);
+	if (err) {
+		bpf_free_verifier_state(new, false);
+		kfree(new_sl);
+		return err;
+	}
+	new->insn_idx = insn_idx;
+	verifier_bug_if(new->branches != 1, env,
+			"%s:branches_to_explore=%d insn %d",
+			__func__, new->branches, insn_idx);
+	err = maybe_enter_scc(env, new);
+	if (err) {
+		bpf_free_verifier_state(new, false);
+		kfree(new_sl);
+		return err;
+	}
+
+	cur->parent = new;
+	cur->first_insn_idx = insn_idx;
+	cur->dfs_depth = new->dfs_depth + 1;
+	bpf_clear_jmp_history(cur);
+	list_add(&new_sl->node, head);
+	return 0;
+}
+
+int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
+{
 	struct bpf_verifier_state_list *sl;
-	struct bpf_verifier_state *cur = env->cur_state, *new;
+	struct bpf_verifier_state *cur = env->cur_state;
 	bool force_new_state, add_new_state, loop;
 	int n, err, states_cnt = 0;
 	struct list_head *pos, *tmp, *head;
@@ -1594,44 +1644,5 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 	 * When looping the sl->state.branches will be > 0 and this state
 	 * will not be considered for equivalence until branches == 0.
 	 */
-	new_sl = kzalloc_obj(struct bpf_verifier_state_list, GFP_KERNEL_ACCOUNT);
-	if (!new_sl)
-		return -ENOMEM;
-	env->total_states++;
-	env->explored_states_size++;
-	update_peak_states(env);
-	env->prev_jmps_processed = env->jmps_processed;
-	env->prev_insn_processed = env->insn_processed;
-
-	/* forget precise markings we inherited, see __mark_chain_precision */
-	if (env->bpf_capable)
-		mark_all_scalars_imprecise(env, cur);
-
-	bpf_clear_singular_ids(env, cur);
-
-	/* add new state to the head of linked list */
-	new = &new_sl->state;
-	err = bpf_copy_verifier_state(new, cur);
-	if (err) {
-		bpf_free_verifier_state(new, false);
-		kfree(new_sl);
-		return err;
-	}
-	new->insn_idx = insn_idx;
-	verifier_bug_if(new->branches != 1, env,
-			"%s:branches_to_explore=%d insn %d",
-			__func__, new->branches, insn_idx);
-	err = maybe_enter_scc(env, new);
-	if (err) {
-		bpf_free_verifier_state(new, false);
-		kfree(new_sl);
-		return err;
-	}
-
-	cur->parent = new;
-	cur->first_insn_idx = insn_idx;
-	cur->dfs_depth = new->dfs_depth + 1;
-	bpf_clear_jmp_history(cur);
-	list_add(&new_sl->node, head);
-	return 0;
+	return bpf_split_cur_state(env);
 }

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (24 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 25/43] bpf: add bpf_split_cur_state() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 27/43] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
                   ` (18 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

When dealing with loops like:

  for (ptr = ...; ptr < end_ptr; ptr += 8)
    ...

Scalar evolution needs to know if ptr and end_ptr identify a same
memory location (possibly with different offsets). If they do,
it is sound to attempt to infer iteration bounds for such a loop.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/verifier.c        | 45 ++++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 47 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 13c7825bb8b5..7e858bb05dd1 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1814,6 +1814,8 @@ struct arg_access_info
 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 			     struct bpf_insn *insn, int arg, int insn_idx);
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
+bool bpf_same_memory_origin(const struct bpf_reg_state *reg_a,
+			    const struct bpf_reg_state *reg_b);
 int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 		      struct cnum64 range, u16 base, u16 step);
 int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 28c16e35b1f3..9e1c08caa0db 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -17153,6 +17153,51 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 	return 0;
 }
 
+/*
+ * Many checks done by this function are quite conservative.
+ * This is because main verification pass does not maintain
+ * enough information to track object identities for some
+ * of the interesting types, e.g. PTR_TO_MEM.
+ */
+bool bpf_same_memory_origin(const struct bpf_reg_state *reg_a,
+			    const struct bpf_reg_state *reg_b)
+{
+	if (reg_a == reg_b)
+		return true;
+	/* Require matching flags and base types. */
+	if (reg_a->type != reg_b->type)
+		return false;
+	/* NULL can't be compared to some base+offset pointer. */
+	if (type_may_be_null(reg_a->type))
+		return false;
+
+	switch (base_type(reg_a->type)) {
+	case PTR_TO_STACK:
+		return reg_a->frameno == reg_b->frameno;
+	case PTR_TO_MAP_VALUE:
+		if (reg_a->map_ptr != reg_b->map_ptr || reg_a->map_uid != reg_b->map_uid)
+			return false;
+		if (reg_a->id && reg_a->id == reg_b->id)
+			return true;
+		/* A plain single-element array has one stable value address. */
+		if (reg_a->map_ptr->map_type == BPF_MAP_TYPE_ARRAY &&
+		    reg_a->map_ptr->max_entries == 1)
+			return true;
+		/*
+		 * The rules above can be simplified / relaxed if:
+		 * - fresh IDs would always be assigned for map-value lookups;
+		 * - direct map value loads would always have and ID of zero.
+		 */
+		return false;
+	case PTR_TO_MEM:
+	case PTR_TO_BUF:
+		return reg_a->id && reg_a->id == reg_b->id;
+
+	default:
+		return false;
+	}
+}
+
 int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 		      struct cnum64 range, u16 base, u16 step)
 {

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 27/43] bpf: allow precision backtracking between overlapping checkpoints
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (25 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops Eduard Zingerman
                   ` (17 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

SCEV integration logic might create two consecutive checkpoints at the
same instruction w/o processing any instructions in between.

This happens for loop-entry checkpoint -> regular checkpoint
sequences, where the loop-entry checkpoint has to remain unchanged
but the regular checkpoint is used for widening.

This commit adapts bpf_mark_chain_precision() to support this.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 5 ++++-
 kernel/bpf/backtrack.c       | 5 +++++
 kernel/bpf/states.c          | 1 +
 3 files changed, 10 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 7e858bb05dd1..c0ed92a632d1 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -502,7 +502,10 @@ struct bpf_verifier_state {
 	bool speculative;
 	bool in_sleepable;
 
-	/* first and last insn idx of this verifier state */
+	/*
+	 * First and last insn idx of this verifier state.
+	 * last_insn_idx is -1 if no instructions have been executed yet.
+	 */
 	u32 first_insn_idx;
 	u32 last_insn_idx;
 	/* if this state is a backedge state then equal_state
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..43d5740034c0 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -888,6 +888,10 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 		}
 
 		if (last_idx < 0) {
+			/* Consecutive checkpoints can have no instructions between them. */
+			if (st->parent)
+				goto parent;
+
 			/* we are at the entry into subprog, which
 			 * is expected for global funcs, but only if
 			 * requested precise registers are R1-R5
@@ -952,6 +956,7 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 				return -EFAULT;
 			}
 		}
+parent:
 		st = st->parent;
 		if (!st)
 			break;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index b1e2daefc264..470bff8d8dc8 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -1306,6 +1306,7 @@ int bpf_split_cur_state(struct bpf_verifier_env *env)
 	}
 
 	cur->parent = new;
+	cur->last_insn_idx = -1;
 	cur->first_insn_idx = insn_idx;
 	cur->dfs_depth = new->dfs_depth + 1;
 	bpf_clear_jmp_history(cur);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (26 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 27/43] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:00   ` sashiko-bot
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops Eduard Zingerman
                   ` (16 subsequent siblings)
  44 siblings, 2 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Assign algebraic expressions to loop variables (registers and stack
spills), describing how their values evolve across iterations. These
summaries provide symbolic input for the subsequent widening patch.

The algorithm draws on ideas from the following papers:

- "Symbolic Evaluation of Chains of Recurrences for Loop Optimization"
  Robert A. van Engelen, 2000
- "The CR# Algebra and its Application in Loop Analysis and Optimization"
  Robert A. van Engelen, 2004

The analysis proceeds in two phases:

- Compute expressions for one symbolic iteration: start each variable
  with a reference to its input value, apply instruction effects and
  join expressions where control-flow paths merge.
- Convert the resulting backedge updates to recurrences at the header,
  then substitute these recurrences into expressions at the latch.

Analyze loops innermost first. Within each loop, visit blocks in
topological order (reverse postorder, with backedges ignored and nested
loops collapsed). Treat each nested loop as an opaque operation:
preserve its invariant values, forget the others and continue at its
exits, without expanding its iterations.

For example, using 64-bit arithmetic:

    0: r7 = 5;
    1: r6 = 0;
    2: do {                // header
    3:     if (r6 == 1)
    4:         r7 = 10;
    5:     r6 += 1;
    6: } while (r6 < 3);   // latch

Expressions use symbolic inputs r6 and r7; the concrete values from
lines 0-1 are supplied later by the main verifier.

First phase:
------------

At each instruction, an environment maps register numbers to their
current expressions. Transfer functions update this mapping, while
joins merge the environments arriving along different paths.

For illustration, follow each branch separately to the backedge:

- At 3, take the path skipping 4: r6 = r6, r7 = r7.
- At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
- At 2, the first path contributes r6 = (+ r6 1), r7 = r7.
- At 4, on the other path from 3, transfer(r7 = 10):
  r7 = r7 -> r7 = 10.
- At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
- At 2, join the two paths' contributions:
  join((+ r6 1), (+ r6 1)) = (+ r6 1) for r6;
  join(r7, 10) = (any r7 10) for r7.

Equal expressions stay unchanged; different expressions become ANY
alternatives. These summarize one iteration, without unrolling the loop.

Second phase:
-------------

The update r6 = (+ r6 1) says that each iteration adds 1 to r6.
Repeating this update n times gives r6_entry + n, a linear recurrence.

- At header 2, r6 = (+ r6 1) becomes (linear r6 1), meaning
  r6_entry + n, where n counts iterations from zero.
- At latch 6, substitute this recurrence into (+ r6 1):
  (+ (linear r6 1) 1) simplifies to (linear (+ r6 1) 1),
  meaning r6_entry + 1 + n.
- At both points, r7 retains (any r7 10): either its entry value
  or 10. ANY records alternatives, not numeric ranges or the path
  conditions selecting them. Unchanged variables retain their inputs.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |    7 +
 kernel/bpf/Makefile          |    2 +-
 kernel/bpf/scev.c            | 1583 ++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |    9 +
 4 files changed, 1600 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c0ed92a632d1..88afc639911b 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -734,6 +734,7 @@ struct bpf_insn_aux_data {
 	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
 	bool arena_scalar; /* ldx/stx/st/atomic through a number, it's an address in arena */
 	bool bb_end; /* last instruction of a basic block */
+	bool need_scev;
 	u8 alu_state; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
@@ -996,6 +997,7 @@ struct bpf_scc_info {
 };
 
 struct bpf_liveness;
+struct scev;
 
 struct bpf_fd_array {
 	union {
@@ -1160,6 +1162,7 @@ struct bpf_verifier_env {
 	struct bpf_iarray *succ;
 	struct bpf_iarray *gotox_tmp_buf;
 	int *idoms;
+	struct scev *scev;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1962,4 +1965,8 @@ void bpf_min_heap_free(struct bpf_min_heap *heap);
 int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
 bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
 
+int bpf_init_scev(struct bpf_verifier_env *env);
+void bpf_free_scev(struct bpf_verifier_env *env);
+int bpf_compute_scev(struct bpf_verifier_env *env);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 7a2c179a9059..fdff31d962d6 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o scev.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
new file mode 100644
index 000000000000..9ef28e0d8d48
--- /dev/null
+++ b/kernel/bpf/scev.c
@@ -0,0 +1,1583 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_verifier.h>
+#include <linux/jhash.h>
+#include <linux/log2.h>
+#include <linux/bug.h>
+
+#define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+#define UNKNOWN_EXPR_ID 0
+#define OPAQUE_EXPR_ID  1
+
+/*
+ * BPF instructions use 'code', 'src_reg', 'off' and 'imm' fields for instruction encoding.
+ * For scalar evolution purpose we want to reuse most of the opcode definitions,
+ * but also add a few custom operations (REG and IMM).
+ * Use the enum below to uniformly represent all operations relevant for SCEV.
+ */
+enum expr_op {
+	/* leave range 0..255 for standard bpf opcodes */
+	UNKNOWN = 256, /* start custom opcodes from the second byte */
+	REG,
+	IMM,
+	SDIV, SMOD,
+	SEXT8, SEXT16, SEXT32,
+	ZEXT8, ZEXT16, ZEXT32,
+	BSWAP16, BSWAP32, BSWAP64,
+	/* during the loop body execution the value can be either of param[0] or param[1] */
+	ANY,
+	/* some value that verifier is not going to track precisely */
+	OPAQUE,
+	/* '*(u8/16/32 *)(r10 + X) = Y' writes define slot contents only partially */
+	SPILL8, SPILL16, SPILL32,
+	/*
+	 * SCEV expression corresponding to linear equation 'param[0] + param[1] * k',
+	 * where k is a loop iteration number. Loop here refers to innermost loop
+	 * containing instruction associated with this expression, as returned by
+	 * bpf_loop_at_index().
+	 */
+	LINEAR_SCEV,
+};
+
+struct expr {
+	u32 op;
+	union {
+		u32 params[2];
+		s64 imm;
+	};
+};
+
+struct expr_bucket {
+	u32 cnt;
+	u32 cap;
+	u32 ids[];
+};
+
+struct env {
+	bool empty;
+	u32 reg2expr[REGS_NUM];
+	u32 reg2scev[REGS_NUM];
+};
+
+struct insn_envs {
+	u32 cnt;
+	struct {
+		int loop_header;
+		struct env *env;
+	} entries[];
+};
+
+#define EXPR_STACK_DEPTH 8
+
+struct expr_stack_elt {
+	u32 id:28;
+	u32 pre:1;
+	u32 next_param:2;
+};
+
+struct scev {
+	/*
+	 * Expressions are identified by id, exprs_ht ensures that
+         * each expression exists as a unique instance.
+	 * This allows for fast equivalence check: id1 === id2.
+	 */
+	struct expr_bucket **exprs_ht;
+	struct bpf_min_heap worklist;
+	struct insn_envs **envs;
+	struct expr *exprs;
+	/*
+	 * Loops are analyzed one by one, this array keeps track if a particular
+	 * instruction was visited on a current pass.
+	 */
+	u32 *discovered;
+	u32 exprs_ht_cnt;
+	int exprs_cnt;
+	int exprs_cap;
+	int envs_cnt;
+	int stack_sz;
+	struct expr_stack_elt expr_stack[EXPR_STACK_DEPTH];
+	u32 ids_buf[EXPR_STACK_DEPTH];
+};
+
+static u32 expr_hash(struct expr *e)
+{
+	return jhash_3words(e->op, e->params[0], e->params[1], 0);
+}
+
+static int expr_eq(struct expr *a, struct expr *b)
+{
+	return a->op == b->op && a->imm == b->imm;
+}
+
+static int add_expr(struct scev *scev, struct expr e)
+{
+	struct expr_bucket *bucket;
+	u32 i, id, hash, new_cap;
+	void *tmp;
+
+	hash = expr_hash(&e) & (scev->exprs_ht_cnt - 1);
+	bucket = scev->exprs_ht[hash];
+
+	if (bucket) {
+		for (i = 0; i < bucket->cnt; i++) {
+			id = bucket->ids[i];
+			if (expr_eq(&e, &scev->exprs[id]))
+				return id;
+		}
+	}
+
+	if (!bucket || bucket->cap == bucket->cnt) {
+		new_cap = bucket ? bucket->cap * 2 : 4;
+		bucket = kvrealloc(bucket, sizeof(*bucket) + sizeof(u32) * new_cap, GFP_KERNEL_ACCOUNT | __GFP_ZERO);
+		if (!bucket)
+			return -ENOMEM;
+		scev->exprs_ht[hash] = bucket;
+		bucket->cap = new_cap;
+	}
+
+	if (scev->exprs_cnt == scev->exprs_cap) {
+		new_cap = scev->exprs_cap ? scev->exprs_cap * 2 : 256;
+		tmp = kvrealloc(scev->exprs, sizeof(struct expr) * new_cap, GFP_KERNEL_ACCOUNT);
+		if (!tmp)
+			return -ENOMEM;
+		scev->exprs = tmp;
+		scev->exprs_cap = new_cap;
+	}
+
+	id = scev->exprs_cnt++;
+	scev->exprs[id] = e;
+	bucket->ids[bucket->cnt++] = id;
+	return id;
+}
+
+static int expr2(struct scev *scev, u32 op, int a, int b)
+{
+	if (a < 0)
+		return a;
+	if (b < 0)
+		return b;
+	return add_expr(scev, (struct expr){ .op = op, .params = {a, b} });
+}
+
+static int expr1(struct scev *scev, u32 op, int a)
+{
+	if (a < 0)
+		return a;
+	return expr2(scev, op, a, 0);
+}
+
+static int expr0(struct scev *scev, u32 op)
+{
+	return expr2(scev, op, 0, 0);
+}
+
+static int imm_expr(struct scev *scev, s64 value)
+{
+	return add_expr(scev, (struct expr){ .op = IMM, .imm = value });
+}
+
+static bool same_exprs(struct scev *scev, int id_a, int id_b)
+{
+	return id_a == id_b;
+}
+
+static bool is_imm(struct scev *scev, int id, s64 *imm)
+{
+	if (scev->exprs[id].op != IMM)
+		return false;
+	*imm = scev->exprs[id].imm;
+	return true;
+}
+
+static bool is_op(struct scev *scev, int id, u32 op)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	return true;
+}
+
+static bool is_unop(struct scev *scev, enum expr_op op, int id, u32 *p0)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	*p0 = scev->exprs[id].params[0];
+	return true;
+}
+
+static bool is_binop(struct scev *scev, enum expr_op op, int id, u32 *p0, u32 *p1)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	*p0 = scev->exprs[id].params[0];
+	*p1 = scev->exprs[id].params[1];
+	return true;
+}
+
+static bool is_reg(struct scev *scev, int id, u32 *reg)
+{
+	return is_unop(scev, REG, id, reg);
+}
+
+static bool is_add(struct scev *scev, int id, u32 *left, u32 *right)
+{
+	return is_binop(scev, BPF_ADD, id, left, right);
+}
+
+static bool is_zext32(struct scev *scev, int id, u32 *left)
+{
+	return is_unop(scev, ZEXT32, id, left);
+}
+
+static bool is_any(struct scev *scev, int id, u32 *left, u32 *right)
+{
+	return is_binop(scev, ANY, id, left, right);
+}
+
+static bool is_linear(struct scev *scev, int id, u32 *base, u32 *slope)
+{
+	return is_binop(scev, LINEAR_SCEV, id, base, slope);
+}
+
+static bool is_opaque(struct scev *scev, int id)
+{
+	return is_op(scev, id, OPAQUE);
+}
+
+static void log_reg(struct bpf_verifier_env *env, u32 reg)
+{
+	if (reg < MAX_BPF_REG)
+		bpf_log(&env->log, "r%d", reg);
+	else
+		bpf_log(&env->log, "*fp%d", (MAX_BPF_REG - reg - 1) * 8);
+}
+
+static const char *op_str(u32 op)
+{
+	switch (op) {
+	case BPF_ADD:  return "+";
+	case BPF_SUB:  return "-";
+	case BPF_MUL:  return "*";
+	case BPF_DIV:  return "/";
+	case SDIV:     return "s/";
+	case BPF_OR:   return "|";
+	case BPF_AND:  return "&";
+	case BPF_LSH:  return "<<";
+	case BPF_RSH:  return ">>";
+	case BPF_NEG:  return "-";
+	case BPF_MOD:  return "%";
+	case SMOD:     return "s%";
+	case BPF_XOR:  return "^";
+	case BPF_ARSH: return "s>>";
+	case SEXT8:    return "sext8";
+	case SEXT16:   return "sext16";
+	case SEXT32:   return "sext32";
+	case ZEXT8:    return "zext8";
+	case ZEXT16:   return "zext16";
+	case ZEXT32:   return "zext32";
+	case BSWAP16:  return "bswap16";
+	case BSWAP32:  return "bswap32";
+	case BSWAP64:  return "bswap64";
+	case SPILL8:   return "spill8";
+	case SPILL16:  return "spill16";
+	case SPILL32:  return "spill32";
+	case ANY:      return "any";
+	case LINEAR_SCEV:  return "linear";
+	}
+	return NULL;
+}
+
+static u32 op_params_num(u32 op)
+{
+	switch (op) {
+	case BPF_ADD:
+	case BPF_SUB:
+	case BPF_MUL:
+	case BPF_DIV:
+	case BPF_MOD:
+	case BPF_OR:
+	case BPF_XOR:
+	case BPF_AND:
+	case BPF_LSH:
+	case BPF_RSH:
+	case BPF_ARSH:
+	case SDIV:
+	case SMOD:
+	case ANY:
+	case LINEAR_SCEV:
+		return 2;
+	case BPF_NEG:
+	case SEXT8:
+	case SEXT16:
+	case SEXT32:
+	case ZEXT8:
+	case ZEXT16:
+	case ZEXT32:
+	case BSWAP16:
+	case BSWAP32:
+	case BSWAP64:
+	case SPILL8:
+	case SPILL16:
+	case SPILL32:
+		return 1;
+	case REG:
+	case IMM:
+	case OPAQUE:
+		return 0;
+	default:
+		return 0;
+	}
+}
+
+static bool expr_stack_push(struct scev *scev, u32 id)
+{
+	if (scev->stack_sz >= EXPR_STACK_DEPTH)
+		return false;
+	scev->expr_stack[scev->stack_sz].id = id;
+	scev->expr_stack[scev->stack_sz].pre = true;
+	scev->expr_stack[scev->stack_sz].next_param = 0;
+	scev->stack_sz++;
+	return true;
+}
+
+enum {
+	PRE = BIT(1), POST = BIT(2), DEPTH_LIMIT = BIT(3)
+};
+
+static bool expr_next(struct scev *scev, u32 *id, u32 *order)
+{
+	struct expr_stack_elt *elt;
+	struct expr *expr;
+	u32 num_params;
+
+	if (scev->stack_sz == 0)
+		return false;
+
+	elt = &scev->expr_stack[scev->stack_sz - 1];
+	*id = elt->id;
+	*order = 0;
+	expr = &scev->exprs[elt->id];
+	num_params = op_params_num(expr->op);
+	if (elt->pre) {
+		elt->pre = false;
+		*order = PRE;
+		return true;
+	}
+	if (elt->next_param == num_params) {
+		*order = POST;
+		scev->stack_sz--;
+		return true;
+	}
+	if (scev->stack_sz == EXPR_STACK_DEPTH) {
+		*order = POST | DEPTH_LIMIT;
+		scev->stack_sz--;
+		return true;
+	}
+	expr_stack_push(scev, expr->params[elt->next_param]);
+	elt->next_param++;
+	return expr_next(scev, id, order);
+}
+
+static void log_expr(struct bpf_verifier_env *env, u32 id)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct expr *expr;
+	const char *str;
+	u32 order;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if ((order & PRE) && scev->stack_sz > 1)
+			bpf_log(log, " ");
+		expr = &scev->exprs[id];
+		switch (expr->op) {
+		case UNKNOWN:
+			if (order & PRE)
+				bpf_log(log, "?");
+			break;
+		case OPAQUE:
+			if (order & PRE)
+				bpf_log(log, "_");
+			break;
+		case REG:
+			if (order & PRE)
+				log_reg(env, expr->params[0]);
+			break;
+		case IMM:
+			if (order & PRE)
+				bpf_log(log, "%lld", expr->imm);
+			break;
+		default:
+			if (order & PRE) {
+				str = op_str(expr->op);
+				bpf_log(log, "(");
+				if (str)
+					bpf_log(log, "%s", str);
+				else
+					bpf_log(log, "bad-expr-op %x", expr->op);
+			}
+			if (order & DEPTH_LIMIT)
+				bpf_log(log, "...");
+			if (order & POST)
+				bpf_log(log, ")");
+		}
+	}
+}
+
+enum print_env_flags {
+	PRINT_SCEV = BIT(1)
+};
+
+static bool reg_alive_at(struct bpf_verifier_env *env, u32 reg, u32 insn_idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	u16 live_regs = aux[insn_idx].live_regs_before;
+
+	return reg < MAX_BPF_REG
+	       ? (BIT(reg) & live_regs)
+	       : test_bit(reg - MAX_BPF_REG, aux[insn_idx].live_stack_before);
+}
+
+static void print_env(struct bpf_verifier_env *env, struct env *e, u32 insn_idx, u32 flags)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	bool printed_some = false;
+	bool print_scev = flags & PRINT_SCEV;
+	bool is_self_scev;
+	bool is_self_reg;
+	int i, r;
+
+	if (e->empty) {
+		bpf_log(log, "  <empty>\n");
+		return;
+	}
+
+	for (i = 0; i < REGS_NUM; i++) {
+		if (!reg_alive_at(env, i, insn_idx))
+			continue;
+		is_self_reg = is_reg(scev, e->reg2expr[i], &r) && i == r;
+		is_self_scev = is_reg(scev, e->reg2scev[i], &r) && i == r;
+		if (is_self_reg && (!print_scev || is_self_scev))
+			continue;
+		printed_some = true;
+		bpf_log(log, "  ");
+		log_reg(env, i);
+		bpf_log(log, "=");
+		log_expr(env, e->reg2expr[i]);
+		if (print_scev && e->reg2expr[i] != UNKNOWN_EXPR_ID) {
+			bpf_log(log, " / ");
+			log_expr(env, e->reg2scev[i]);
+		}
+		bpf_log(log, "\n");
+	}
+	if (!printed_some)
+		bpf_log(log, "  <all regs unchanged>\n");
+}
+static struct env *find_loop_env(struct scev *scev, int loop_header, int insn_idx)
+{
+	struct insn_envs *envs = scev->envs[insn_idx];
+	u32 i;
+
+	if (!envs)
+		return NULL;
+	for (i = 0; i < envs->cnt; i++)
+		if (envs->entries[i].loop_header == loop_header)
+			return envs->entries[i].env;
+	return NULL;
+}
+
+static struct env *find_header_env(struct scev *scev, int header)
+{
+	return find_loop_env(scev, header, header);
+}
+
+static struct env *get_loop_env(struct scev *scev, int loop_header, int insn_idx)
+{
+	struct insn_envs *envs = scev->envs[insn_idx];
+	u32 cnt = envs ? envs->cnt : 0;
+	struct insn_envs *tmp;
+	struct env *e;
+
+	e = find_loop_env(scev, loop_header, insn_idx);
+	if (e)
+		return e;
+
+	e = kzalloc(sizeof(*e), GFP_KERNEL_ACCOUNT);
+	if (!e)
+		return NULL;
+
+	tmp = krealloc(envs, struct_size(envs, entries, cnt + 1), GFP_KERNEL_ACCOUNT);
+	if (!tmp) {
+		kfree(e);
+		return NULL;
+	}
+
+	e->empty = true;
+	tmp->entries[cnt].loop_header = loop_header;
+	tmp->entries[cnt].env = e;
+	tmp->cnt = cnt + 1;
+	scev->envs[insn_idx] = tmp;
+	return e;
+}
+
+static int replace_reg(struct scev *scev, struct env *e, u32 reg, int id)
+{
+	if (id < 0)
+		return id;
+	e->reg2expr[reg] = id;
+	return 0;
+}
+
+static int setup_initial_loop_env(struct bpf_verifier_env *env, struct env *e, int insn_idx)
+{
+	struct scev *scev = env->scev;
+	int i, err;
+
+	for (i = 0; i < REGS_NUM; i++) {
+		err = replace_reg(scev, e, i, expr1(scev, REG, i));
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
+static void forget_call_regs(struct env *e)
+{
+	int i;
+
+	for (i = BPF_REG_0; i <= BPF_REG_5; i++)
+		e->reg2expr[i] = UNKNOWN_EXPR_ID;
+}
+
+static u32 spill_spi(struct bpf_insn *insn)
+{
+	return -insn->off / BPF_REG_SIZE - 1;
+}
+
+static u32 off_to_reg(struct bpf_insn *insn)
+{
+	return spill_spi(insn) + MAX_BPF_REG;
+}
+
+static bool is_spill_off(int off)
+{
+	return off % BPF_REG_SIZE == 0 &&
+	       off <= -BPF_REG_SIZE &&
+	       off >= -MAX_BPF_STACK_JIT;
+}
+
+static void mark_opaque(struct env *e, int off, int size)
+{
+	int b, spi;
+
+	for (b = off; b < off + size; b++) {
+		if (b >= 0 || b < -MAX_BPF_STACK_JIT)
+			continue;
+		spi = (-b - 1) / BPF_REG_SIZE;
+		e->reg2expr[MAX_BPF_REG + spi] = OPAQUE_EXPR_ID;
+	}
+}
+
+static int mk_spill(struct scev *scev, u8 size, int id)
+{
+	switch (size) {
+	case BPF_B:  return expr1(scev, SPILL8,  id);
+	case BPF_H:  return expr1(scev, SPILL16, id);
+	case BPF_W:  return expr1(scev, SPILL32, id);
+	case BPF_DW: return id;
+	}
+	return UNKNOWN_EXPR_ID;
+}
+
+static int mk_fill(struct scev *scev, int id, u8 code)
+{
+	bool sx = BPF_MODE(code) == BPF_MEMSX;
+
+#ifdef __BIG_ENDIAN
+	/* A narrow fill from a DW spill reads high bits on big-endian. */
+	if (BPF_SIZE(code) != BPF_DW)
+		return OPAQUE_EXPR_ID;
+#endif
+
+	switch (BPF_SIZE(code)) {
+	case BPF_B:  return expr1(scev, sx ? SEXT8  : ZEXT8,  id);
+	case BPF_H:  return expr1(scev, sx ? SEXT16 : ZEXT16, id);
+	case BPF_W:  return expr1(scev, sx ? SEXT32 : ZEXT32, id);
+	case BPF_DW: return sx ? UNKNOWN_EXPR_ID : id; /* DW sign-extended load is invalid */
+	}
+	return UNKNOWN_EXPR_ID;
+}
+
+static int maybe_store_fp(struct scev *scev, struct env *e, struct bpf_insn *insn, int id)
+{
+	u8 size = BPF_SIZE(insn->code);
+
+	if (insn->dst_reg != BPF_REG_FP)
+		return 0;
+	if (is_spill_off(insn->off))
+		return replace_reg(scev, e, off_to_reg(insn), mk_spill(scev, size, id));
+	mark_opaque(e, insn->off, bpf_size_to_bytes(size));
+	return 0;
+}
+
+static int maybe_load_fp(struct scev *scev, struct env *e, struct bpf_insn *insn)
+{
+	int id;
+
+	if (insn->src_reg == BPF_REG_FP && is_spill_off(insn->off))
+		id = mk_fill(scev, e->reg2expr[off_to_reg(insn)], insn->code);
+	else
+		id = OPAQUE_EXPR_ID;
+	return replace_reg(scev, e, insn->dst_reg, id);
+}
+
+static int transfer(struct bpf_verifier_env *env, struct env *e, int idx)
+{
+	const bool little_endian = htons(0x3412) == 0x1234;
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct scev *scev = env->scev;
+	u32 *reg2expr = e->reg2expr;
+	u8 class = BPF_CLASS(insn->code);
+	u8 x_or_k = BPF_SRC(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+	u8 mode = BPF_MODE(insn->code);
+	u32 dst = insn->dst_reg;
+	u32 src = insn->src_reg;
+	u32 op, sext;
+	bool need_bswap;
+	int i, id;
+	s64 imm;
+
+	switch (class) {
+	case BPF_ALU:
+	case BPF_ALU64:
+		switch (opcode) {
+		case BPF_MOV:
+			switch (insn->off) {
+			case 0: sext = 0; break;
+			case 8: sext = SEXT8; break;
+			case 16: sext = SEXT16; break;
+			case 32: sext = SEXT32; break;
+			default:
+				goto mark_dst_unknown;
+			}
+
+			if (x_or_k == BPF_X && insn->imm == 0)
+				id = reg2expr[src];
+			else if (x_or_k == BPF_K && src == 0 && insn->off == 0)
+				id = imm_expr(scev, insn->imm);
+			else
+				goto mark_dst_unknown;
+
+			if (sext)
+				id = expr1(scev, sext, id);
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_ADD:
+		case BPF_SUB:
+		case BPF_MUL:
+		case BPF_DIV:
+		case BPF_MOD:
+		case BPF_OR:
+		case BPF_XOR:
+		case BPF_AND:
+		case BPF_LSH:
+		case BPF_RSH:
+		case BPF_ARSH:
+			if (opcode == BPF_DIV && insn->off == 1)
+				op = SDIV;
+			else if (opcode == BPF_MOD && insn->off == 1)
+				op = SMOD;
+			else if (insn->off == 0)
+				op = opcode;
+			else
+				goto mark_dst_unknown;
+
+			if (x_or_k == BPF_X && insn->imm == 0)
+				id = reg2expr[src];
+			else if (x_or_k == BPF_K && src == 0)
+				id = imm_expr(scev, insn->imm);
+			else
+				goto mark_dst_unknown;
+
+			id = expr2(scev, op, reg2expr[dst], id);
+
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_NEG:
+			if (src == 0 && insn->off == 0 && insn->imm == 0)
+				id = expr1(scev, BPF_NEG, reg2expr[dst]);
+			else
+				goto mark_dst_unknown;
+
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_END:
+			if (src != 0 || insn->off != 0 || (class == BPF_ALU64 && x_or_k != 0))
+				goto mark_dst_unknown;
+
+			need_bswap = class == BPF_ALU64 || ((x_or_k == BPF_TO_LE) != little_endian);
+			switch (insn->imm) {
+			case 16: op = need_bswap ? BSWAP16 : ZEXT16; break;
+			case 32: op = need_bswap ? BSWAP32 : ZEXT32; break;
+			case 64: op = need_bswap ? BSWAP64 : 0; break;
+			default:
+				goto mark_dst_unknown;
+			}
+
+			id = reg2expr[dst];
+			if (op)
+				id = expr1(scev, op, reg2expr[dst]);
+
+			return replace_reg(scev, e, dst, id);
+		default:
+			goto mark_dst_unknown;
+		}
+		break;
+	case BPF_LDX:
+		switch (mode) {
+		case BPF_MEM:
+		case BPF_MEMSX:
+			return maybe_load_fp(scev, e, insn);
+		default:
+			goto mark_dst_unknown;
+		}
+	case BPF_STX:
+		switch (mode) {
+		case BPF_MEM:
+			return maybe_store_fp(scev, e, insn, reg2expr[src]);
+		case BPF_ATOMIC:
+			if (insn->imm == BPF_LOAD_ACQ)
+				return maybe_load_fp(scev, e, insn);
+			if (insn->imm == BPF_STORE_REL) {
+				return maybe_store_fp(scev, e, insn, reg2expr[src]);
+			}
+			/*
+			 * verifier does not track other atomic ops precisely,
+			 * hence mark the results as opaque.
+			 */
+			if (insn->dst_reg == BPF_REG_FP)
+				mark_opaque(e, insn->off, bpf_size_to_bytes(BPF_SIZE(insn->code)));
+			if (insn->imm == BPF_CMPXCHG)
+				return replace_reg(scev, e, BPF_REG_0, OPAQUE_EXPR_ID);
+			if (insn->imm & BPF_FETCH)
+				return replace_reg(scev, e, src, OPAQUE_EXPR_ID);
+			break;
+		}
+		break;
+	case BPF_ST:
+		if (insn->dst_reg == BPF_REG_FP) {
+			id = imm_expr(scev, insn->imm);
+			if (id < 0)
+				return id;
+			return maybe_store_fp(scev, e, insn, id);
+		}
+		break;
+	case BPF_JMP:
+	case BPF_JMP32:
+		if (opcode == BPF_CALL)
+			forget_call_regs(e);
+		/* for non-CALL there are no changes in register states */
+		break;
+	case BPF_LD:
+		switch (mode) {
+		case BPF_IMM:
+			/* rX = imm ll */
+			if (BPF_SIZE(insn->code) == BPF_DW && insn->src_reg == 0) {
+				imm = ((u64)(insn + 1)->imm << 32) | (u32)insn->imm;
+				return replace_reg(scev, e, dst, imm_expr(scev, imm));
+			}
+			/* map, map value, BTF id, function */
+			if (BPF_SIZE(insn->code) == BPF_DW)
+				return replace_reg(scev, e, dst, OPAQUE_EXPR_ID);
+			goto mark_dst_unknown;
+		case BPF_ABS:
+		case BPF_IND:
+			forget_call_regs(e);
+			break;
+		default:
+			goto mark_dst_unknown;
+		}
+		break;
+	default:
+		/* unknown instruction, nuke state */
+		for (i = 0; i < REGS_NUM; i++)
+			reg2expr[i] = UNKNOWN_EXPR_ID;
+		break;
+	}
+	return 0;
+
+mark_dst_unknown:
+	reg2expr[dst] = UNKNOWN_EXPR_ID;
+	return 0;
+}
+
+/*
+ * Construct minimal 'ANY' expression by traversing 'a' and 'b'
+ * and accumulating non-duplicated non-ANY entries.
+ */
+static int mk_any(struct scev *scev, u32 a, u32 b)
+{
+	struct expr_stack_elt *elt;
+	struct expr *expr;
+	u32 i, j, ids_buf_sz;
+	u32 roots[2] = {a, b};
+	int id;
+
+	ids_buf_sz = 0;
+	for (i = 0; i < ARRAY_SIZE(roots); i++) {
+		scev->stack_sz = 0;
+		expr_stack_push(scev, roots[i]);
+		while (scev->stack_sz) {
+			elt = &scev->expr_stack[--scev->stack_sz];
+			id = elt->id;
+			expr = &scev->exprs[elt->id];
+			if (expr->op == ANY) {
+				if (!expr_stack_push(scev, expr->params[0]) ||
+				    !expr_stack_push(scev, expr->params[1]))
+					return UNKNOWN_EXPR_ID;
+			} else {
+				for (j = 0; j < ids_buf_sz; j++) {
+					if (same_exprs(scev, scev->ids_buf[j], id))
+						goto next;
+				}
+				if (ids_buf_sz == ARRAY_SIZE(scev->ids_buf))
+					return UNKNOWN_EXPR_ID;
+				scev->ids_buf[ids_buf_sz++] = id;
+			}
+next:;
+		}
+	}
+	if (WARN_ON(ids_buf_sz == 0))
+		return -EFAULT;
+	id = scev->ids_buf[0];
+	for (i = 1; i < ids_buf_sz; i++) {
+		id = expr2(scev, ANY, id, scev->ids_buf[i]);
+		if (id < 0)
+			return id;
+	}
+	return id;
+}
+
+static int join(struct scev *scev, struct env *acc, struct env *cur)
+{
+	int i, id;
+
+	if (acc->empty) {
+		memcpy(acc, cur, sizeof(*acc));
+		acc->empty = false;
+		return 0;
+	}
+
+	for (i = 0; i < REGS_NUM; i++) {
+		if (!same_exprs(scev, acc->reg2expr[i], cur->reg2expr[i])) {
+			id = mk_any(scev, acc->reg2expr[i], cur->reg2expr[i]);
+			if (id < 0)
+				return id;
+			acc->reg2expr[i] = id;
+		}
+	}
+	return 0;
+}
+
+/* Mark any register modified in 'e' as unknown in 'acc'. */
+static void forget_modified_regs(struct scev *scev, struct env *acc, struct env *e)
+{
+	u32 i, reg;
+
+	for (i = 0; i < REGS_NUM; i++) {
+		if (e && is_reg(scev, e->reg2expr[i], &reg) && reg == i)
+			continue;
+		acc->reg2expr[i] = UNKNOWN_EXPR_ID;
+	}
+}
+
+/* For a loop H at `header`, mark any register modified as a result of H execution as unknown. */
+static void forget_non_invariants(struct bpf_verifier_env *env, struct env *acc, int header)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop = aux[header].loop;
+	struct scev *scev = env->scev;
+	int i, idx, h;
+
+	acc->empty = false;
+	forget_modified_regs(scev, acc, find_header_env(scev, header));
+	for (i = 0; i < loop->exits_cnt; i++) {
+		idx = loop->exits[i].from;
+		h = bpf_loop_at_index(env, idx);
+		/*
+		 * In order to account for situations like this:
+		 *
+		 * header:
+		 *   do {
+		 *	saved = rX;
+		 *	rX = -1000;
+		 *	do {				// H
+		 *		if (random())
+		 *			goto out;	// E
+		 *	} while (inner_again);
+		 *	rX = saved;
+		 *    } while (outer_again);
+		 *  out:
+		 *
+		 * For exits originating in nested loops check that rX is invariant
+		 * both at the header (H) and at exit (E).
+		 */
+		while (h != header) {
+			forget_modified_regs(scev, acc, find_header_env(scev, h));
+			forget_modified_regs(scev, acc, find_loop_env(scev, h, idx));
+			idx = h;
+			h = aux[h].loop_header;
+		}
+		/*
+		 * For exits originating at the same loop the header is already accounted for,
+		 * check only the exit environment.
+		 */
+		forget_modified_regs(scev, acc, find_loop_env(scev, header, idx));
+	}
+}
+
+static int worklist_push(struct scev *scev, int loop_header, int idx)
+{
+	u32 tag = (u32)loop_header + 1;
+	int err;
+
+	if (scev->discovered[idx] == tag)
+		return 0;
+
+	err = bpf_min_heap_push(&scev->worklist, idx);
+	if (err)
+		return err;
+
+	scev->discovered[idx] = tag;
+	return 0;
+}
+
+static bool reg_alive_at_succ(struct bpf_verifier_env *env, u32 r, u32 insn_idx)
+{
+	struct bpf_iarray *succ = bpf_insn_successors(env, insn_idx);
+	u32 i;
+
+	for (i = 0; i < succ->cnt; i++)
+		if (reg_alive_at(env, r, succ->items[i]))
+			return true;
+	return false;
+}
+
+enum changes_at {
+	LOG_AT_TRANSFER,
+	LOG_AT_JOIN,
+};
+
+static void log_env_changes(struct bpf_verifier_env *env, enum changes_at at,
+			    struct env *old, struct env *new, int insn_idx)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	u64 null_pos = log->end_pos;
+	bool any_changes = false;
+	u32 old_id, new_id;
+	u64 len;
+	int r;
+
+	bpf_log(log, "%s %4d: ", at == LOG_AT_TRANSFER ? "t" : "j", insn_idx);
+	bpf_verbose_insn(env, &env->prog->insnsi[insn_idx]);
+	len = log->end_pos - null_pos;
+	bpf_log(log, "%*s", max(37 - (int)len, 1), " ");
+	bpf_log(log, " ; ");
+	for (r = 0; r < REGS_NUM; r++) {
+		old_id = old->reg2expr[r];
+		new_id = new->reg2expr[r];
+		if (at == LOG_AT_TRANSFER && !reg_alive_at_succ(env, r, insn_idx))
+			continue;
+		if (at == LOG_AT_JOIN && !reg_alive_at(env, r, insn_idx))
+			continue;
+		if (same_exprs(scev, old_id, new_id))
+			continue;
+		if (any_changes)
+			bpf_log(log, ", ");
+		log_reg(env, r);
+		bpf_log(log, " ");
+		log_expr(env, old_id);
+		bpf_log(log, " -> ");
+		log_expr(env, new_id);
+		any_changes = true;
+	}
+	bpf_log(log, "\n");
+	if (!any_changes)
+		bpf_vlog_reset(log, null_pos);
+}
+
+static bool is_probe_read_helper(u32 func_id)
+{
+	return func_id == BPF_FUNC_probe_read ||
+	       func_id == BPF_FUNC_probe_read_kernel ||
+	       func_id == BPF_FUNC_probe_read_user ||
+	       func_id == BPF_FUNC_probe_read_str ||
+	       func_id == BPF_FUNC_probe_read_kernel_str ||
+	       func_id == BPF_FUNC_probe_read_user_str;
+}
+
+/*
+ * If instruction is an indirect write to stack, invalidate SCEVs for spi's
+ * that this instruction can touch.
+ */
+static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct env *cur_env, int idx)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct scev *scev = env->scev;
+	u8 class = BPF_CLASS(insn->code);
+	const unsigned long *mask;
+	u32 spi, reg;
+	bool opaque;
+
+	/* Direct fp stores are fine. */
+	if ((class == BPF_ST || class == BPF_STX) && insn->dst_reg == BPF_REG_FP)
+		return;
+
+	mask = bpf_may_write_mask(env, idx);
+	opaque = bpf_helper_call(insn) && is_probe_read_helper(insn->imm);
+	for_each_set_bit(spi, mask, MAX_BPF_STACK_SLOTS) {
+		reg = MAX_BPF_REG + spi;
+		if (cur_env->reg2expr[reg] == UNKNOWN_EXPR_ID)
+			continue;
+		replace_reg(scev, cur_env, reg, opaque ? OPAQUE_EXPR_ID : UNKNOWN_EXPR_ID);
+	}
+}
+
+/* Find the topmost loop header containing idx inside cur_header, or -1 if none. */
+static int topmost_nested_loop(struct bpf_verifier_env *env, int idx, int cur_header)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int header;
+
+	for (header = bpf_loop_at_index(env, idx); header >= 0;
+	     header = aux[header].loop_header)
+		if (aux[header].loop_header == cur_header)
+			return header;
+
+	return -1;
+}
+
+/*
+ * Join cur_env into the successor's environment and schedule their traversal.
+ * Map successors in nested loops to their topmost nested header.
+ * Ignore successors outside the current loop and its nested loops.
+ */
+static int join_successor(struct bpf_verifier_env *env, int cur_header, int succ_idx,
+			  struct env *old_env, struct env *cur_env)
+{
+	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
+	struct scev *scev = env->scev;
+	int err, nested_header;
+	struct env *succ_env;
+
+	/*
+	 * There are several possibilities for a successor:
+	 * - succ_idx can be a part of a loop outside of the cur_header's loop,
+	 *   such edges are ignored.
+	 * - succ_idx can be a part of the same loop as cur_header,
+	 *   for such edges succ_idx environment is updated:
+	 *     e[succ_idx] = join(e[succ_idx], cur_env)
+	 * - succ_idx can be a part of some loop inner to cur_header,
+	 *   in such a case there exists some loop header H,
+	 *   such that H.loop_header == cur_header
+	 *   and H is the same as succ_idx's loop or contains it.
+	 */
+	if (bpf_loop_at_index(env, succ_idx) != cur_header) {
+		/*
+		 * topmost_nested_loop() either finds H or returns -1,
+		 * in case if succ_idx is a part of a loop outer to cur_header.
+		 */
+		nested_header = topmost_nested_loop(env, succ_idx, cur_header);
+		if (nested_header < 0)
+			return 0;
+		succ_idx = nested_header;
+	}
+	succ_env = get_loop_env(scev, cur_header, succ_idx);
+	if (!succ_env)
+		return -ENOMEM;
+	if (log_level2)
+		memcpy(old_env, succ_env, sizeof(*old_env));
+	err = join(scev, succ_env, cur_env);
+	if (err)
+		return err;
+	err = worklist_push(scev, cur_header, succ_idx);
+	if (err)
+		return err;
+	if (log_level2)
+		log_env_changes(env, LOG_AT_JOIN, old_env, succ_env, succ_idx);
+	return 0;
+}
+
+/* Unsupported loops get only an unknown header environment. */
+static bool can_compute_loop_scev(const struct bpf_loop *loop)
+{
+	return !loop->irreducible && !loop->exits_overflow;
+}
+
+static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
+{
+	struct bpf_min_heap *worklist = &env->scev->worklist;
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *cur_loop = aux[cur_header].loop;
+	struct env *header_env;
+	struct scev *scev = env->scev;
+	struct bpf_loop *nested_loop;
+	struct env *cur_env = NULL;
+	struct env *old_env = NULL;
+	struct bpf_iarray *succ;
+	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
+	int s, i, err, idx, succ_idx;
+
+	if (log_level2)
+		bpf_log(&env->log, "Computing SCEV for loop at %d:\n", cur_header);
+
+	/*
+	 * For irreducible loops, and for loops nesting a loop with a truncated
+	 * exit list, just assume that everything is clobbered for now.
+	 */
+	if (!can_compute_loop_scev(cur_loop)) {
+		header_env = get_loop_env(scev, cur_header, cur_header);
+		if (!header_env)
+			return -ENOMEM;
+		/* The freshly allocated environment has all expressions unknown. */
+		header_env->empty = false;
+		return 0;
+	}
+
+	cur_env = kzalloc(sizeof(*cur_env), GFP_KERNEL_ACCOUNT);
+	if (!cur_env)
+		goto nomem;
+	if (log_level2) {
+		old_env = kzalloc(sizeof(*old_env), GFP_KERNEL_ACCOUNT);
+		if (!old_env)
+			goto nomem;
+	}
+	header_env = get_loop_env(scev, cur_header, cur_header);
+	if (!header_env)
+		goto nomem;
+	err = setup_initial_loop_env(env, header_env, cur_header);
+	if (err)
+		goto out;
+	err = worklist_push(scev, cur_header, cur_header);
+	if (err)
+		goto out;
+
+	for (;;) {
+		if (!bpf_min_heap_pop(worklist, &idx))
+			break;
+
+		/* join_successor() maps nested-loop entries to their representative header. */
+		nested_loop = idx != cur_header ? aux[idx].loop : NULL;
+		memcpy(cur_env, find_loop_env(scev, cur_header, idx), sizeof(*cur_env));
+		if (nested_loop) {
+			/*
+			 * Process nested loop as a single instruction by
+			 * forgetting anything non-invariant in the nested loop
+			 */
+			if (log_level2)
+				memcpy(old_env, cur_env, sizeof(*old_env));
+			forget_non_invariants(env, cur_env, idx);
+			if (log_level2)
+				log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
+			/*
+			 * Treat nested loop exits as successors,
+			 * join cur_env into successor's envs.
+			 */
+			for (i = 0; i < nested_loop->exits_cnt; i++) {
+				s = nested_loop->exits[i].to;
+				err = join_successor(env, cur_header, s, old_env, cur_env);
+				if (err)
+					goto out;
+			}
+		} else {
+			/*
+			 * Iterate instructions within a single basic block
+			 * starting at 'idx' mutating 'cur_env'.
+			 */
+			for (;;) {
+				if (log_level2)
+					memcpy(old_env, cur_env, sizeof(*old_env));
+				err = transfer(env, cur_env, idx);
+				if (err)
+					goto out;
+				reset_scevs_at_indirect_writes(env, cur_env, idx);
+				if (log_level2)
+					log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
+				succ = bpf_insn_successors(env, idx);
+				if (succ->cnt != 1)
+					break;
+				succ_idx = succ->items[0];
+				if (aux[idx].bb_end || aux[succ_idx].need_scev ||
+				    bpf_loop_at_index(env, succ_idx) != cur_header)
+					break;
+				idx = succ_idx;
+			}
+			/* Join cur_env into basic block successor's envs. */
+			iarray_for_each(s, succ) {
+				err = join_successor(env, cur_header, s, old_env, cur_env);
+				if (err)
+					goto out;
+			}
+		}
+	}
+
+	err = 0;
+out:
+	kfree(cur_env);
+	kfree(old_env);
+	return err;
+nomem:
+	err = -ENOMEM;
+	goto out;
+}
+
+static void mark_latches(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop;
+	int len = env->prog->len;
+	int i, j, latch;
+
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		aux[i].need_scev = true;
+		if (!can_compute_loop_scev(loop))
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			latch = loop->backedges[j].latch;
+			if (latch >= 0)
+				aux[latch].need_scev = true;
+		}
+		for (j = 0; j < loop->exits_cnt; j++)
+			aux[loop->exits[j].from].need_scev = true;
+	}
+}
+
+static bool is_any_imm_reg_opaque(struct scev *scev, u32 id)
+{
+	u32 l, r, reg, order;
+	s64 imm;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT)
+			return false;
+		if (!(order & PRE) ||
+		    is_imm(scev, id, &imm) ||
+		    is_reg(scev, id, &reg) ||
+		    is_opaque(scev, id) ||
+		    is_any(scev, id, &l, &r))
+			continue;
+		return false;
+	}
+	return true;
+}
+
+/*
+ * Can implement explicit stack version, but it is harder to read.
+ * Stick with recursive version for now.
+ */
+static int transform_expr_once(struct scev *scev, u32 lvl, u32 root, void *priv,
+			       int (*fn)(struct scev *scev, u32 id, void *priv))
+{
+	struct expr expr;
+	int p0, p1, id;
+
+	if (lvl >= EXPR_STACK_DEPTH)
+		return UNKNOWN_EXPR_ID;
+
+	expr = scev->exprs[root]; /* snapshot the expr before potential realloc */
+	switch (op_params_num(expr.op)) {
+	case 0:
+		id = root;
+		break;
+	case 1:
+		p0 = transform_expr_once(scev, lvl + 1, expr.params[0], priv, fn);
+		id = expr1(scev, expr.op, p0);
+		break;
+	case 2:
+		p0 = transform_expr_once(scev, lvl + 1, expr.params[0], priv, fn);
+		p1 = transform_expr_once(scev, lvl + 1, expr.params[1], priv, fn);
+		id = expr2(scev, expr.op, p0, p1);
+		break;
+	}
+	return id < 0 ? id : fn(scev, id, priv);
+}
+
+static int transform_expr(struct scev *scev, u32 root, void *priv,
+			  int (*fn)(struct scev *scev, u32 id, void *priv))
+{
+	int id_old, id_new = root;
+
+	do {
+		id_old = id_new;
+		id_new = transform_expr_once(scev, 0, id_old, priv, fn);
+		if (id_new < 0)
+			return id_new;
+	} while (id_old != id_new);
+	return id_new;
+}
+
+static int simplify(struct scev *scev, u32 id, void *priv)
+{
+	u32 l, r, base, slope;
+	s64 imm1, imm2;
+
+	/* (+ (linear base slope) imm) -> (linear (+ base imm) slope) */
+	if (is_add(scev, id, &l, &r) &&
+	    is_linear(scev, l, &base, &slope) &&
+	    is_imm(scev, r, &imm1))
+		return expr2(scev, LINEAR_SCEV, expr2(scev, BPF_ADD, base, r), slope);
+
+	/* (+ imm imm) -> imm */
+	if (is_add(scev, id, &l, &r) &&
+	    is_imm(scev, l, &imm1) &&
+	    is_imm(scev, r, &imm2))
+		return imm_expr(scev, imm1 + imm2);
+
+	if (is_zext32(scev, id, &l) && is_imm(scev, l, &imm1))
+		return imm_expr(scev, (u64)(u32)imm1);
+
+	return id;
+}
+
+static int compute_header_scevs(struct bpf_verifier_env *env, struct env *header_env)
+{
+	struct scev *scev = env->scev;
+	u32 ra, rb, l, r, ra_expr;
+	s64 imm;
+	int id;
+
+	for (ra = 0; ra < REGS_NUM; ra++) {
+		id = transform_expr(scev, header_env->reg2expr[ra], NULL, simplify);
+		if (id < 0)
+			return id;
+		header_env->reg2expr[ra] = id;
+		ra_expr = header_env->reg2expr[ra];
+		/* rA = (+ rA IMM) */
+		if (is_add(scev, ra_expr, &l, &r) &&
+		    is_reg(scev, l, &rb) &&
+		    is_imm(scev, r, &imm) &&
+		    ra == rb) {
+			id = expr2(scev, LINEAR_SCEV, l, r);
+			if (id < 0)
+				return id;
+			header_env->reg2scev[ra] = id;
+			continue;
+		}
+		/* rA = rA */
+		if (is_reg(scev, ra_expr, &rb) && ra == rb) {
+			header_env->reg2scev[ra] = ra_expr;
+			continue;
+		}
+		/* rA = (any 1 2 3 4 ...) */
+		if (is_any_imm_reg_opaque(scev, ra_expr)) {
+			header_env->reg2scev[ra] = ra_expr;
+			continue;
+		}
+
+	}
+	return 0;
+}
+static int instantiate_header_scevs(struct scev *scev, u32 id, void *priv)
+{
+	u32 reg, *reg2scev = priv;
+
+	/* (reg r) -> ... equation at header ... */
+	if (is_reg(scev, id, &reg))
+		return reg2scev[reg];
+	/* keep expression shapes */
+	return id;
+}
+
+static int compute_insn_scevs(struct bpf_verifier_env *env, struct env *eheader, struct env *einsn)
+{
+	struct scev *scev = env->scev;
+	int id, reg;
+
+	for (reg = 0; reg < REGS_NUM; reg++) {
+		id = einsn->reg2expr[reg];
+		id = transform_expr_once(scev, 0, id, eheader->reg2scev, instantiate_header_scevs);
+		if (id < 0)
+			return id;
+		id = transform_expr(scev, id, NULL, simplify);
+		if (id < 0)
+			return id;
+		einsn->reg2scev[reg] = id;
+	}
+
+	return 0;
+}
+
+static int log_scevs(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct env *header_env, *latch_env;
+	struct scev *scev = env->scev;
+	struct bpf_loop *loop;
+	int i, j, len, latch;
+
+	len = env->prog->len;
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		header_env = find_header_env(scev, i);
+		bpf_log(log, "scev at header %d:\n", i);
+		print_env(env, header_env, i, PRINT_SCEV);
+		if (!can_compute_loop_scev(loop))
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			latch = loop->backedges[j].latch;
+			if (latch < 0)
+				continue;
+			bpf_log(log, " scev at latch %d:\n", latch);
+			latch_env = find_loop_env(scev, i, latch);
+			if (verifier_bug_if(!latch_env, env, "i=%d, latch=%d", i, latch))
+				return -EFAULT;
+			print_env(env, latch_env, latch, PRINT_SCEV);
+		}
+	}
+	return 0;
+}
+
+int bpf_compute_scev(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct scev *scev = env->scev;
+	int *postorder = env->cfg.insn_postorder;
+	int cnt = env->cfg.cur_postorder;
+	int i, idx, err, header;
+
+	mark_latches(env);
+	/*
+	 * Visit loop headers in postorder, to guarantee that scevs
+	 * for innermost loops are computed first.
+	 */
+	for (i = 0; i < cnt; i++) {
+		idx = postorder[i];
+		if (!aux[idx].loop)
+			continue;
+		err = compute_scev_for_loop(env, idx);
+		if (err)
+			return err;
+	}
+
+	/*
+	 * Compute scevs from exprs collected on a previous step. Iterate instructions in
+	 * reverse post-order so that each loop header is processed before instructions
+	 * reachable from it.
+	 */
+	for (i = cnt - 1; i >= 0; i--) {
+		idx = postorder[i];
+		if (!aux[idx].need_scev)
+			continue;
+		header = bpf_loop_at_index(env, idx);
+		err = aux[idx].loop
+		      ? compute_header_scevs(env, find_header_env(scev, header))
+		      : compute_insn_scevs(env,
+					   find_header_env(scev, header),
+					   find_loop_env(scev, header, idx));
+		if (err)
+			return err;
+	}
+
+	for (i = 0; i < env->prog->len; i++)
+		if (aux[i].loop)
+			aux[i].prune_point = true;
+
+	if (env->log.level & BPF_LOG_LEVEL2) {
+		err = log_scevs(env);
+		if (err)
+			return err;
+	}
+
+	return 0;
+}
+
+static int reverse_ranked_compare(int a, int b, void *arg)
+{
+	int *rank = arg;
+
+	return rank[b] - rank[a];
+}
+
+void bpf_free_scev(struct bpf_verifier_env *env)
+{
+	struct scev *scev = env->scev;
+	struct insn_envs *envs;
+	int i;
+	u32 j;
+
+	if (!scev)
+		return;
+	for (i = 0; i < scev->envs_cnt; i++) {
+		envs = scev->envs[i];
+		if (!envs)
+			continue;
+		for (j = 0; j < envs->cnt; j++)
+			kfree(envs->entries[j].env);
+		kfree(envs);
+	}
+	if (scev->exprs_ht)
+		for (i = 0; i < scev->exprs_ht_cnt; i++)
+			kvfree(scev->exprs_ht[i]);
+	kvfree(scev->exprs_ht);
+	bpf_min_heap_free(&scev->worklist);
+	kvfree(scev->envs);
+	kvfree(scev->exprs);
+	kvfree(scev->discovered);
+	kfree(scev);
+	env->scev = NULL;
+}
+
+int bpf_init_scev(struct bpf_verifier_env *env)
+{
+	struct scev *scev;
+
+	scev = kzalloc(sizeof(struct scev), GFP_KERNEL_ACCOUNT);
+	if (!scev)
+		return -ENOMEM;
+	env->scev = scev;
+	/* Order worklist in reverse post-order. */
+	bpf_min_heap_init(&scev->worklist, reverse_ranked_compare, env->cfg.postorder_nums);
+	/* A power of two for hash & (exprs_ht_cnt - 1)` indexing. */
+	scev->exprs_ht_cnt = roundup_pow_of_two(max(256U, DIV_ROUND_UP(env->prog->len, 4)));
+	scev->exprs_ht = kvcalloc(scev->exprs_ht_cnt, sizeof(*scev->exprs_ht), GFP_KERNEL_ACCOUNT);
+	if (!scev->exprs_ht)
+		goto nomem;
+	if (expr0(scev, UNKNOWN) < 0 || expr0(scev, OPAQUE)  < 0)
+		goto nomem;
+	scev->envs = kvcalloc(env->prog->len, sizeof(*scev->envs), GFP_KERNEL_ACCOUNT);
+	scev->discovered = kvcalloc(env->prog->len, sizeof(*scev->discovered), GFP_KERNEL_ACCOUNT);
+	if (!scev->envs || !scev->discovered)
+		goto nomem;
+	/*
+	 * Remember original program length, in case bpf_free_scev()
+	 * is called after bpf program rewrites that increase program
+	 * length.
+	 */
+	scev->envs_cnt = env->prog->len;
+	return 0;
+nomem:
+	bpf_free_scev(env);
+	return -ENOMEM;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 9e1c08caa0db..1fc3e6e2387e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -23085,6 +23085,14 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (env->log.level & BPF_LOG_LEVEL2)
 		log_program(env);
 
+	ret = bpf_init_scev(env);
+	if (ret < 0)
+		goto skip_full_check;
+
+	ret = bpf_compute_scev(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = mark_fastcall_patterns(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -23234,6 +23242,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 		bpf_clear_insn_aux_data(env, 0, env->insn_aux_data_len);
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
+	bpf_free_scev(env);
 	bpf_stack_liveness_free(env);
 	kvfree(env->cfg.postorder_nums);
 	kvfree(env->cfg.insn_postorder);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (27 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
                   ` (15 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Loop-bound computation and register widening
===========================================

During the main verification pass, use SCEV expressions and register
values at the entry to the loop to estimate the number of iterations
the loop might execute. Use this number to widen loop induction variables.

For example:

    0: r7 = 5;
    1: r6 = 0;
    2: while (r6 < 3) {   // latch
    3:     if (r6 == 1)
    4:         r7 = 10;
    5:     r6 += 1;
    6: }
    8: use(r7);

- Scalar evolution expressions (SCEVs) at the latch are:
  - r6 = (linear r6 1)
  - r7 = (any r7 10)
- For this pre-condition loop, this corresponds to the linear inequality
  `r6_entry + n < 3`. With r6_entry = 0, the inequality is `n < 3`.
- Hence infer that r6's range at line 2 is [0,3], including the final
  failed test. The body starts with r6 in [0,2].
- r7's (any r7 10) expression means that its initial value 5 from line 0
  is joined with 10 from line 4: it is widened to [5,10] within
  and after the loop.

Currently such inference is supported for:
- Reducible loops with one backedge and complete backedge/exit lists.
- A linear latch dominating the backedge, with one successor exiting
  the loop. Additional exits are allowed; they can shorten execution.
- 64-bit arithmetic and comparisons whose continuation condition
  normalizes to signed/unsigned < or <=, or !=.
- Initial counter, latch offset, slope and bound evaluable as constants
  at loop entry, yielding a finite, representable iteration bound.
- Values with an eligible type at loop entry, e.g. SCALAR_VALUE or
  PTR_TO_STACK, but not PTR_TO_CTX.

In order for widening to proceed, every register live at the loop
entry has to have a SCEV that can be widened. Currently, the following
expressions are supported:
- (linear <self> <step>): entry value + iteration# * constant step.
- (reg <self>): loop invariant; keep its entry value unchanged.
- (any ...): join the entry value with immediates or invariant registers;
  types must match, and immediates require SCALAR_VALUE.

Main verification loop integration
==================================

Before proceeding with the usual instruction processing, do_check()
takes the following steps at loop entries (H):
- Checks whether the loop can be widened and obtains its
  iteration bound (scev.c:bpf_compute_loop_iters()).
- If it can, saves an unwidened checkpoint E0, which will be used as
  a base for deriving widened values throughout this invocation of
  the loop.
- Widens registers in the current state using E0 and SCEV expressions
  computed for H (scev.c:bpf_widen_scev_regs()).
- Pushes a per-frame loop-stack record to the current verifier state.
  For terminating loops, the record stores E0 and the iteration bounds.
- Lets is_state_visited() create a new checkpoint W at the same
  header, now containing the widened values, and proceeds with
  normal checking.
- On an exit edge, pops the loop record and verifies the exit
  path normally.

is_state_visited() treats a loop proven to terminate like an
iterator-based loop: RANGE_WITHIN comparison attempts to establish
convergence against a checkpoint whose exploration is still
in progress.

For example:

    r0 = 0;
    H: r0 += 1;
       if (r0 != 3) goto H;
       exit;

    first H:  prove H_count = 3; save E0(r0 = 0)
              widen r0 to [0,2]; push the loop record
              save W(r0 = [0,2])
    body:     r0 becomes [1,3]; fork at the conditional
    exit:     r0 = 3; pop the loop record and check exit
    backedge: r0 = [1,2]; clamp, then match with W and prune

Clamping
--------

Independent register ranges lose correlations between induction
variables. For example:

    r6 = 0; r7 = 0;
    do { r6 += 1; r7 += 2; } while (r6 < 3);
    use(r7);

    header:   r6 in [0,2], r7 in {0,2,4}
    body:     r6 in [1,3], r7 in {2,4,6}
    backedge: r6 in [1,2], r7 still in {2,4,6}

The latch narrows r6 but does not propagate this refinement to r7.
To account for this, before comparing states on a backedge,
scev.c:bpf_clamp_scev_regs() recomputes r7's range from E0 and the
saved iteration bound, then intersects the current range {2,4,6}
with the recomputed {0,2,4}, arriving at {2,4}.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h                       |  30 +
 kernel/bpf/log.c                                   |   5 +
 kernel/bpf/scev.c                                  | 840 +++++++++++++++++++++
 kernel/bpf/states.c                                |  88 ++-
 kernel/bpf/verifier.c                              | 203 ++++-
 tools/testing/selftests/bpf/progs/verifier_gotox.c |   6 +-
 6 files changed, 1160 insertions(+), 12 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 88afc639911b..d13f662dcabc 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -328,6 +328,20 @@ struct bpf_retval_range {
 	bool return_32bit;
 };
 
+struct bpf_loop_iters {
+	u32 min_header_count;   /* min number of times header is executed */
+	u32 max_header_count;   /* max number of times header is executed */
+};
+
+struct loop_stack_entry {
+	struct bpf_verifier_state *entry_state;
+	struct bpf_loop_iters iters;
+	u32 loop_id:31;
+	u32 terminates:1;
+};
+
+#define LOOP_STACK_SIZE 16
+
 /* state of the program:
  * type of all registers and stack info
  */
@@ -373,6 +387,11 @@ struct bpf_func_state {
 	u32 callback_depth;
 	/* Instructions processed in this frame and callees on the current path. */
 	u32 insns_subtotal;
+	/*
+	 * Control-flow loop nesting at the current insn within this frame's
+	 * subprogram (loops never cross subprogram boundaries).
+	 */
+	u32 loop_stack_cnt;
 
 	/* The following fields should be last. See copy_func_state() */
 	/* The state of the stack. Each element of the array describes BPF_REG_SIZE
@@ -390,6 +409,7 @@ struct bpf_func_state {
 
 	u16 out_stack_arg_cnt; /* Number of outgoing on-stack argument slots */
 	struct bpf_reg_state *stack_arg_regs; /* Outgoing on-stack arguments */
+	struct loop_stack_entry *loop_stack;
 };
 
 #define MAX_CALL_FRAMES 16
@@ -436,6 +456,7 @@ static_assert(MAX_BPF_STACK_SLOTS <= (1 << 12));
 #define MAX_STACK_ARG_SLOTS (MAX_BPF_FUNC_ARGS - MAX_BPF_FUNC_REG_ARGS)
 #define BPF_ID_MAP_SIZE ((MAX_BPF_REG + MAX_BPF_STACK_SLOTS + MAX_STACK_ARG_SLOTS) * \
 			 MAX_CALL_FRAMES)
+
 struct bpf_verifier_state {
 	/* call stack tracking */
 	struct bpf_func_state *frame[MAX_CALL_FRAMES];
@@ -1163,6 +1184,8 @@ struct bpf_verifier_env {
 	struct bpf_iarray *gotox_tmp_buf;
 	int *idoms;
 	struct scev *scev;
+	/* SCEV representation of stack slots not allocated yet. */
+	struct bpf_reg_state scev_not_init_reg;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1969,4 +1992,11 @@ int bpf_init_scev(struct bpf_verifier_env *env);
 void bpf_free_scev(struct bpf_verifier_env *env);
 int bpf_compute_scev(struct bpf_verifier_env *env);
 
+int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			   struct bpf_loop_iters *iters);
+int bpf_widen_scev_regs(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			struct bpf_verifier_state *loop_entry, struct bpf_loop_iters *iters);
+int bpf_clamp_scev_regs(struct bpf_verifier_env *env, struct bpf_func_state *st, u32 insn_idx,
+			struct bpf_verifier_state *entry_state, struct bpf_loop_iters *iters);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index 900d1bb1988b..6a5564f62669 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -792,6 +792,11 @@ void print_verifier_state(struct bpf_verifier_env *env, const struct bpf_verifie
 		verbose(env, " cb");
 	if (state->in_async_callback_fn)
 		verbose(env, " async_cb");
+	if (state->loop_stack_cnt) {
+		verbose(env, " loop_stack=");
+		for (i = 0; i < state->loop_stack_cnt; i++)
+			verbose(env, "%s%d", i ? "," : "", state->loop_stack[i].loop_id);
+	}
 	verbose(env, "\n");
 	if (!print_all)
 		mark_verifier_state_clean(env);
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
index 9ef28e0d8d48..a98f3fea52ec 100644
--- a/kernel/bpf/scev.c
+++ b/kernel/bpf/scev.c
@@ -1,10 +1,13 @@
 // SPDX-License-Identifier: GPL-2.0-only
 /* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
 
+#include "linux/cnum.h"
 #include <linux/bpf_verifier.h>
 #include <linux/jhash.h>
 #include <linux/log2.h>
 #include <linux/bug.h>
+#include <linux/tnum.h>
+#include <linux/overflow.h>
 
 #define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
 #define UNKNOWN_EXPR_ID 0
@@ -1557,6 +1560,7 @@ int bpf_init_scev(struct bpf_verifier_env *env)
 	if (!scev)
 		return -ENOMEM;
 	env->scev = scev;
+	bpf_mark_reg_not_init(env, &env->scev_not_init_reg);
 	/* Order worklist in reverse post-order. */
 	bpf_min_heap_init(&scev->worklist, reverse_ranked_compare, env->cfg.postorder_nums);
 	/* A power of two for hash & (exprs_ht_cnt - 1)` indexing. */
@@ -1581,3 +1585,839 @@ int bpf_init_scev(struct bpf_verifier_env *env)
 	bpf_free_scev(env);
 	return -ENOMEM;
 }
+
+static struct bpf_reg_state *scev_regno_to_reg(struct bpf_verifier_env *env,
+					    struct bpf_func_state *st, u32 r)
+{
+	int spi, slots_available;
+
+	if (r < MAX_BPF_REG)
+		return &st->regs[r];
+
+	slots_available = st->allocated_stack / BPF_REG_SIZE;
+	spi = r - MAX_BPF_REG;
+	if (spi < slots_available)
+		return &st->stack[spi].spilled_ptr;
+
+	return &env->scev_not_init_reg;
+}
+
+static bool scev_reg_alive(struct bpf_verifier_env *env, struct bpf_verifier_state *st, u32 r)
+{
+	int insn_idx = bpf_frame_insn_idx(st, st->curframe);
+	u16 live_regs = env->insn_aux_data[insn_idx].live_regs_before;
+	int spi;
+
+	if (r < MAX_BPF_REG) {
+		return BIT(r) & live_regs;
+	} else {
+		spi = r - MAX_BPF_REG;
+		return bpf_stack_slot_alive(env, st->curframe, spi * 2) ||
+		       bpf_stack_slot_alive(env, st->curframe, spi * 2 + 1);
+	}
+}
+
+/*
+ * Latch is a condition deciding if execution remains inside a loop.
+ * Linear latch represents a condition 'if <reg> <op> <loop invariant> goto <loop-header>',
+ * where equation '<base> + i * <step> <op> <bound>' describes values taken by register <reg>,
+ * 'i' is the loop iteration number, starting from 0.
+ */
+struct linear_latch {
+	u32 insn_idx;
+	u32 base_expr;
+	u32 step_expr;
+	u32 bound_expr;
+	u32 op;
+};
+
+static bool loop_invariant(struct scev *scev, u32 id)
+{
+	s64 imm;
+	u32 reg;
+
+	return is_reg(scev, id, &reg) || is_imm(scev, id, &imm);
+}
+
+static int match_linear_latch(struct bpf_verifier_env *env,
+			      u32 header,
+			      u32 latch_idx,
+			      struct linear_latch *latch)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[latch_idx];
+	struct scev *scev = env->scev;
+	struct env *latch_env;
+	u32 true_branch_tgt;
+	u32 src_reg_scev;
+	u32 dst_reg_scev;
+	u32 t, l, r, op;
+	int id;
+
+	/* 32-bit arithmetic is not handled yet */
+	if (BPF_CLASS(insn->code) != BPF_JMP)
+		return false;
+	op = BPF_OP(insn->code);
+	/* Flip the condition if true branch jumps out of the loop */
+	true_branch_tgt = latch_idx + bpf_jmp_offset(insn) + 1;
+	if (bpf_loop_at_index(env, true_branch_tgt) !=
+	    bpf_loop_at_index(env, latch_idx))
+		op = bpf_rev_opcode(op);
+	switch (op) {
+	case BPF_JSLT:
+	case BPF_JSLE:
+	case BPF_JSGT:
+	case BPF_JSGE:
+	case BPF_JLT:
+	case BPF_JLE:
+	case BPF_JGT:
+	case BPF_JGE:
+	case BPF_JNE:
+		break;
+	default:
+		return false;
+	}
+
+	latch->op = op;
+	latch->insn_idx = latch_idx;
+
+	latch_env = find_loop_env(scev, header, latch_idx);
+	if (verifier_bug_if(!latch_env, env, "header_idx=%d, latch_idx=%d\n", header, latch_idx))
+		return -EFAULT;
+	dst_reg_scev = latch_env->reg2scev[insn->dst_reg];
+	src_reg_scev = latch_env->reg2scev[insn->src_reg];
+	/*
+	 * simplify() produces two shapes for linear latches:
+	 * - (linear register step)
+	 *           ^----------------------------+-- base_expr
+	 *           base_expr                    |
+	 * - (linear (+ register imm) step)       |
+	 *           ^----------------------------'
+	 */
+	if (!is_linear(scev, dst_reg_scev, &latch->base_expr, &latch->step_expr))
+		return false;
+
+	if (BPF_SRC(insn->code) == BPF_K) {
+		id = imm_expr(scev, insn->imm);
+		if (id < 0)
+			return id;
+		latch->bound_expr = id;
+	} else {
+		latch->bound_expr = src_reg_scev;
+	}
+
+	if (!loop_invariant(scev, latch->step_expr) ||
+	    !loop_invariant(scev, latch->bound_expr))
+		return false;
+
+	/* (linear register step) */
+	if (is_reg(scev, latch->base_expr, &t))
+		return true;
+
+	/* (linear (+ register imm) step) */
+	if (is_add(scev, latch->base_expr, &l, &r) &&
+	    is_reg(scev, l, &t) &&
+	    loop_invariant(scev, r))
+		return true;
+
+	return false;
+}
+
+struct scev_value {
+	u64 value;	/* scalar value or a pointer offset */
+	int ptr_reg;	/* SCEV register identifying the pointer origin, or -1. */
+};
+
+static void log_scev_value(struct bpf_verifier_env *env, struct bpf_func_state *st,
+			   const struct scev_value *v)
+{
+	const struct bpf_reg_state *reg;
+
+	if (v->ptr_reg >= 0) {
+		reg = scev_regno_to_reg(env, st, v->ptr_reg);
+		log_reg(env, v->ptr_reg);
+		bpf_log(&env->log, " (%s)", reg_type_str(env, reg->type));
+	} else {
+		bpf_log(&env->log, "scalar %llu", v->value);
+	}
+}
+
+static void log_incompatible_scev_values(struct bpf_verifier_env *env, const char *reason,
+					struct bpf_func_state *st,
+					const struct scev_value *a, const struct scev_value *b)
+{
+	bpf_log(&env->log, "scev: %s ", reason);
+	log_scev_value(env, st, a);
+	bpf_log(&env->log, ", ");
+	log_scev_value(env, st, b);
+	bpf_log(&env->log, "\n");
+}
+
+static bool scev_value_compatible(struct bpf_verifier_env *env, struct bpf_func_state *st,
+				  const struct scev_value *a, const struct scev_value *b)
+{
+	const struct bpf_reg_state *ra, *rb;
+
+	if (a->ptr_reg < 0 || b->ptr_reg < 0) {
+		if (a->ptr_reg < 0 && b->ptr_reg < 0)
+			return true;
+		goto incompatible;
+	}
+	ra = scev_regno_to_reg(env, st, a->ptr_reg);
+	rb = scev_regno_to_reg(env, st, b->ptr_reg);
+	if (bpf_same_memory_origin(ra, rb))
+		return true;
+
+incompatible:
+	if (env->log.level & BPF_LOG_LEVEL2)
+		log_incompatible_scev_values(env, "incompatible latch operands", st, a, b);
+	return false;
+}
+
+static bool scev_value_add(struct bpf_verifier_env *env, struct bpf_func_state *st,
+			   const struct scev_value *a, const struct scev_value *b,
+			   struct scev_value *result)
+{
+	if (a->ptr_reg >= 0 && b->ptr_reg >= 0) {
+		if (env->log.level & BPF_LOG_LEVEL2)
+			log_incompatible_scev_values(env, "can't add two pointers", st, a, b);
+		return false;
+	}
+	result->value = a->value + b->value;
+	result->ptr_reg = a->ptr_reg >= 0 ? a->ptr_reg : b->ptr_reg;
+	return true;
+}
+
+static bool eval_expr(struct bpf_verifier_env *env, struct scev *scev,
+		      struct bpf_func_state *st, u32 id, struct scev_value *result)
+{
+	struct scev_value lval, rval;
+	struct bpf_reg_state *reg;
+	u32 l, r, regno;
+	s64 imm;
+
+	if (is_reg(scev, id, &regno)) {
+		reg = scev_regno_to_reg(env, st, regno);
+		if (reg->type == NOT_INIT)
+			return false;
+		if (tnum_is_const(reg->var_off)) {
+			result->value = reg->var_off.value;
+			result->ptr_reg = reg->type == SCALAR_VALUE ? -1 : (int)regno;
+			return true;
+		}
+	} else if (is_imm(scev, id, &imm)) {
+		result->value = imm;
+		result->ptr_reg = -1;
+		return true;
+	} else if (is_add(scev, id, &l, &r) &&
+		   eval_expr(env, scev, st, l, &lval) &&
+		   eval_expr(env, scev, st, r, &rval)) {
+		return scev_value_add(env, st, &lval, &rval, result);
+	}
+	return false;
+}
+
+/*
+ * Loop with post-condition:
+ *
+ *    r0 = 0
+ * l: ...                r0 ∈ [0,1,2] header executed 3 times
+ *    r0 += 1            r0 ∈ [0,1,2]
+ *    ...                r0 ∈ [1,2,3]
+ *    if r0 != 3 goto l  r0 ∈ [1,2,3] backedge taken 2 times
+ *    ...                r0 ∈ [3]
+ *
+ * SCEV at header: r0 = k
+ * SCEV at latch:  r0 = 1 + k
+ *
+ * Loop with pre-condition:
+ *
+ *    r0 = 0
+ * l: ...                r0 ∈ [0,1,2,3] header executed 4 times
+ *    if r0 == 3 goto e  r0 ∈ [0,1,2,3] backedge taken 3 times
+ *    r0 += 1            r0 ∈ [0,1,2]
+ *    ...                r0 ∈ [1,2,3]
+ *    goto l             r0 ∈ [1,2,3]
+ * e: ...                r0 ∈ [3]
+ *
+ * SCEV at header: r0 = k
+ * SCEV at latch:  r0 = k
+ */
+static bool compute_max_iters(struct bpf_verifier_env *env,
+			      struct bpf_func_state *st,
+			      struct linear_latch *latch,
+			      struct bpf_loop_iters *iters)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop = aux[bpf_loop_at_index(env, latch->insn_idx)].loop;
+	struct scev *scev = env->scev;
+	struct scev_value initial_val, step_val, bound_val;
+	u64 step, diff, bound, initial, backedge_taken_count;
+	u8 op = latch->op;
+	bool up, down;
+
+	if (!eval_expr(env, scev, st, latch->base_expr, &initial_val) ||
+	    !eval_expr(env, scev, st, latch->step_expr, &step_val) ||
+	    !eval_expr(env, scev, st, latch->bound_expr, &bound_val))
+		return false;
+
+	if (step_val.ptr_reg >= 0) {
+		if (env->log.level & BPF_LOG_LEVEL2) {
+			bpf_log(&env->log, "scev: latch step is a pointer ");
+			log_scev_value(env, st, &step_val);
+			bpf_log(&env->log, "\n");
+		}
+		return false;
+	}
+	if (!scev_value_compatible(env, st, &initial_val, &bound_val))
+		return false;
+
+	initial = initial_val.value;
+	step = step_val.value;
+	bound = bound_val.value;
+
+	if (step == 0)
+		return false;
+	down = (s64)step < 0;
+	up = (s64)step > 0;
+	diff = bound - initial;
+	switch (op) {
+	case BPF_JLT:
+		/*
+		 * E.g.: 3 + 2*i < 7 ⇔ 2*i < 4 ⇔ i ∈ [0,1].
+		 * Here the inequality describes a set of values `i` might take
+		 * while the latch condition remains true (the backedge is taken).
+		 * The backedge_taken_count is the size of this set.
+		 */
+		if (/*
+		     * If the condition is initially true, decreasing values won't
+		     * change that without underflow. If the condition is initially false,
+		     * the main pass would do a single iteration anyway.
+		     */
+		    down ||
+		    /* Is the latch condition false on the first iteration? */
+		    initial >= bound ||
+		    /* Count iterations satisfying the strict inequality, including i=0. */
+		    check_add_overflow((diff - 1) / step, 1, &backedge_taken_count) ||
+		    /*
+		     * Check if the last iteration overflows the counter:
+		     *   initial + backedge_taken_count * step > U64_MAX
+		     */
+		    backedge_taken_count > (U64_MAX - initial) / step)
+			return false;
+		break;
+	case BPF_JLE:
+		/* E.g.: 3 + 2*i ≤ 7 ⇔ 2*i ≤ 4 ⇔ i ∈ [0,1,2]. */
+		if (down || initial > bound ||
+		    /* Count iterations satisfying the non-strict inequality. */
+		    check_add_overflow(diff / step, 1, &backedge_taken_count) ||
+		    backedge_taken_count > (U64_MAX - initial) / step)
+			return false;
+		break;
+	case BPF_JGT:
+		/* E.g.: 7 - 2*i > 3 ⇔ 2*i < 4 ⇔ i ∈ [0,1]. */
+		if (up || initial <= bound ||
+		    check_add_overflow((-diff - 1) / -step, 1, &backedge_taken_count) ||
+		    /*
+		     * The last iteration underflows the counter if:
+		     *   initial + backedge_taken_count * step < 0 ⇔
+		     *             backedge_taken_count * step < -initial ⇔
+		     *                    backedge_taken_count > initial / -step
+		     * (direction flips because step < 0)
+		     */
+		    backedge_taken_count > initial / -step)
+			return false;
+		break;
+	case BPF_JGE:
+		/* E.g.: 7 - 2*i ≥ 3 ⇔ 2*i ≤ 4 ⇔ i ∈ [0,1,2]. */
+		if (up || initial < bound ||
+		    check_add_overflow(-diff / -step, 1, &backedge_taken_count) ||
+		    backedge_taken_count > initial / -step)
+			return false;
+		break;
+	case BPF_JSLT:
+		if (down || (s64)initial >= (s64)bound ||
+		    check_add_overflow((diff - 1) / step, 1, &backedge_taken_count) ||
+		    /*
+		     * Check if the last iteration overflows the signed counter:
+		     *   (s64)initial + backedge_taken_count * step > S64_MAX
+		     */
+		    backedge_taken_count > ((u64)S64_MAX - initial) / step)
+			return false;
+		break;
+	case BPF_JSLE:
+		if (down || (s64)initial > (s64)bound ||
+		    check_add_overflow(diff / step, 1, &backedge_taken_count) ||
+		    backedge_taken_count > ((u64)S64_MAX - initial) / step)
+			return false;
+		break;
+	case BPF_JSGT:
+		if (up || (s64)initial <= (s64)bound ||
+		    check_add_overflow((-diff - 1) / -step, 1, &backedge_taken_count) ||
+		    /*
+		     * The last iteration underflows the signed counter if:
+		     *   (s64)initial + backedge_taken_count * step < S64_MIN ⇔
+		     *                  backedge_taken_count * step < S64_MIN - (s64)initial ⇔
+		     *                         backedge_taken_count > initial - (u64)S64_MIN / -step
+		     *	 (direction flips because step < 0)
+		     */
+		    backedge_taken_count > (initial - (u64)S64_MIN) / -step)
+			return false;
+		break;
+	case BPF_JSGE:
+		if (up || (s64)initial < (s64)bound ||
+		    check_add_overflow(-diff / -step, 1, &backedge_taken_count) ||
+		    backedge_taken_count > (initial - (u64)S64_MIN) / -step)
+			return false;
+		break;
+	case BPF_JNE:
+		/*
+		 * For a decreasing counter flip the diff and step:
+		 *   7 - 2*i ≠ 3 ⇔ 2*i ≠ 7 - 3.
+		 */
+		diff = down ? -diff : diff;
+		step = down ? -step : step;
+		if (diff == 0 || diff % step)
+			return false;
+		backedge_taken_count = diff / step;
+		break;
+	default:
+		return false;
+	}
+
+	/*
+	 * max_header_count is u32 and U32_MAX means "infinite",
+	 * check to avoid overflow below.
+	 */
+	if (backedge_taken_count >= U32_MAX - 1)
+		return false;
+
+	/* The header executes once on entry and once for each taken backedge, hence +1. */
+	iters->max_header_count = backedge_taken_count + 1;
+	/*
+	 * The latch is a conditional jump with one jump target exiting the loop.
+	 * Linear latch is matched only if the loop has a single backedge.
+	 * The loop still, however can have multiple exits.
+	 * In such case, conservatively assume that non-latch exit can happen
+	 * at any iteration, thus setting minimal number of iterations as 0.
+	 */
+	iters->min_header_count = loop->exits_cnt == 1 ? iters->max_header_count : 0;
+	return true;
+}
+
+static void mark_scev_reg_scratched(struct bpf_verifier_env *env, u32 r)
+{
+	if (r < MAX_BPF_REG)
+		mark_reg_scratched(env, r);
+	else
+		mark_stack_slot_scratched(env, r - __MAX_BPF_REG);
+}
+
+/* Main logic in verifier.c forbids varying offsets for certain register types. */
+static bool is_widenable_reg_type(const struct bpf_reg_state *reg)
+{
+	if (type_may_be_null(reg->type))
+		return false;
+
+	switch (base_type(reg->type)) {
+	case SCALAR_VALUE:
+	case PTR_TO_MAP_VALUE:
+	case PTR_TO_MAP_KEY:
+	case PTR_TO_STACK:
+	case PTR_TO_PACKET:
+	case PTR_TO_PACKET_META:
+	case PTR_TO_MEM:
+	case PTR_TO_BUF:
+	case PTR_TO_BTF_ID:
+		return true;
+	default:
+		return false;
+	}
+}
+
+static void scratch_widened_reg_id(struct bpf_verifier_env *env, struct bpf_reg_state *reg)
+{
+	switch (base_type(reg->type)) {
+	case SCALAR_VALUE:
+		reg->id = 0;
+		break;
+	case PTR_TO_PACKET:
+	case PTR_TO_PACKET_META:
+		reg->id = ++env->id_gen;
+	default:
+	}
+}
+
+/* Check that ANY leaves can be unioned with r's loop-entry value. */
+static bool is_any_imm_reg(struct bpf_verifier_env *env, struct bpf_func_state *loop_entry,
+			   struct env *header_env, u32 r, u32 id)
+{
+	struct scev *scev = env->scev;
+	struct bpf_reg_state *reg, *leaf_reg;
+	u32 l, rr, leaf, ra, order;
+	s64 imm;
+
+	if (!is_any(scev, id, &l, &rr))
+		return false;
+	reg = scev_regno_to_reg(env, loop_entry, r);
+	if (!is_widenable_reg_type(reg))
+		return false;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT)
+			return false;
+		if (!(order & PRE) || is_any(scev, id, &l, &rr))
+			continue;
+		if (is_imm(scev, id, &imm)) {
+			if (reg->type != SCALAR_VALUE)
+				return false;
+		} else if (is_reg(scev, id, &leaf)) {
+			/* Other leaves must denote loop-invariant registers. */
+			if (leaf != r &&
+			    !(is_reg(scev, header_env->reg2scev[leaf], &ra) && ra == leaf))
+				return false;
+			leaf_reg = scev_regno_to_reg(env, loop_entry, leaf);
+			if (leaf_reg->type != reg->type)
+				return false;
+		} else {
+			return false;
+		}
+	}
+	return true;
+}
+
+struct bounds {
+	struct cnum64 range;
+	u16 base;
+	u16 step;
+};
+
+static bool is_simple_linear(struct scev *scev, u32 id, u32 *base_reg, s64 *slope_imm)
+{
+	u32 base, slope;
+
+	return is_linear(scev, id, &base, &slope) &&
+	       is_reg(scev, base, base_reg) &&
+	       is_imm(scev, slope, slope_imm) &&
+	       *slope_imm <= S16_MAX &&
+	       *slope_imm >= S16_MIN &&
+	       *slope_imm != 0;
+}
+
+/*
+ * Compute the range an induction variable in `reg` spans over the loop.
+ * Returns false if the computation overflows s64.
+ */
+static bool linear_bounds(struct bpf_reg_state *reg, struct bpf_loop_iters *iters, s64 slope,
+			  struct bounds *out)
+{
+	s64 slope_abs = slope < 0 ? -slope : slope;
+	s64 min_val = reg_smin(reg);
+	s64 max_val = reg_smax(reg);
+	s64 total_change;
+	u16 base, step;
+
+	if (check_mul_overflow(slope, (s64)iters->max_header_count - 1, &total_change))
+		return false;
+	if (slope > 0) {
+		if (check_add_overflow(max_val, total_change, &max_val))
+			return false;
+	} else {
+		if (check_add_overflow(min_val, total_change, &min_val))
+			return false;
+	}
+	/*
+	 * If the entry value is a single point the value set is 'v + slope * k',
+	 * so the step is |slope|. Otherwise, only the power-of-two alignment
+	 * shared by the entry value and the slope.
+	 */
+	if (cnum64_is_const(reg->r64)) {
+		step = slope_abs;
+		base = imod(min_val, step);
+	} else {
+		step = 1u << min_t(u32, tnum_alignment(reg->var_off), __ffs(slope_abs));
+		base = 0;
+	}
+	out->range = cnum64_from_srange(min_val, max_val);
+	out->base = base;
+	out->step = step;
+	return true;
+}
+
+int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			   struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *cur_func = st->frame[st->curframe];
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct env *header_env;
+	struct linear_latch latch;
+	struct bpf_loop *loop;
+	struct bounds bounds;
+	int insn_idx = st->insn_idx;
+	int linear_latch;
+	u32 r, base_reg;
+	int latch_idx;
+	s64 slope_imm;
+	int err;
+
+	/*
+	 * If insn_idx is a loop header for a reducible loop with a single backedge.
+	 * loop is NULL for secondary entries to irreducible loops.
+	 */
+	loop = aux[insn_idx].loop;
+	if (!loop || loop->irreducible || loop->backedges_cnt != 1 || loop->backedges_overflow ||
+	    loop->exits_overflow) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, unsupported loop:%s%s%s\n", insn_idx,
+				!loop || loop->irreducible ? " irreducible" : "",
+				loop && loop->backedges_cnt > 1 ? " multiple backedges" : "",
+				loop && loop->exits_overflow ? " too many exits" : "");
+		return 0;
+	}
+
+	/* If this backedge has a latch */
+	latch_idx = loop->backedges[0].latch;
+	if (latch_idx < 0) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, latch not identified\n", insn_idx);
+		return 0;
+	}
+
+	linear_latch = match_linear_latch(env, insn_idx, latch_idx, &latch);
+	if (linear_latch < 0)
+		return linear_latch;
+
+	if (!linear_latch) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, non-linear latch at %d\n",
+				insn_idx, latch_idx);
+		return 0;
+	}
+
+	if (!compute_max_iters(env, cur_func, &latch, iters)) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, can't compute iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (iters->max_header_count == 0) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, 0 iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (iters->max_header_count == U32_MAX) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, inf iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (log->level & BPF_LOG_LEVEL2) {
+		bpf_log(log, "loop header at %d, header_count is ", insn_idx);
+		if (iters->min_header_count == iters->max_header_count)
+			bpf_log(log, "%u ", iters->max_header_count);
+		else
+			bpf_log(log, "[%u..%u] ", iters->min_header_count, iters->max_header_count);
+		bpf_log(log, "\n");
+	}
+
+	err = bpf_live_stack_query_init(env, st);
+	if (err)
+		return err;
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		struct bpf_reg_state *reg;
+		u32 ra, r_scev, r_expr;
+
+		if (!scev_reg_alive(env, st, r))
+			continue;
+
+		r_scev = header_env->reg2scev[r];
+		reg = scev_regno_to_reg(env, cur_func, r);
+		/* If SCEV for r is (linear <reg> <slope>) */
+		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
+		    is_widenable_reg_type(reg)) {
+			/* Can't widen if the iteration range overflows */
+			if (base_reg == r &&
+			    !linear_bounds(reg, iters, slope_imm, &bounds))
+				goto cant_widen;
+			continue;
+		}
+
+		/* rA = rA, loop does not change this reg */
+		if (is_reg(scev, r_scev, &ra) && r == ra)
+			continue;
+
+		/* (any 1 (any 2 (any 3 4))) */
+		if (is_any_imm_reg(env, cur_func, header_env, r, r_scev))
+			continue;
+
+cant_widen:
+		if (log->level & BPF_LOG_LEVEL2) {
+			r_expr = header_env->reg2expr[r];
+			bpf_log(log, "loop header at %d, can't widen ", insn_idx);
+			log_reg(env, r);
+			bpf_log(log, ", expr is ");
+			log_expr(env, r_expr);
+			bpf_log(log, "\n");
+		}
+		return 0;
+	}
+
+	return 1;
+}
+
+/* bpf_compute_loop_iters() checked that every leaf can be unioned into acc. */
+static int union_any_reg(struct bpf_verifier_env *env, struct bpf_func_state *loop_entry,
+			 struct bpf_reg_state *acc, u32 id)
+{
+	struct bpf_reg_state *tmp = &env->fake_reg[0], *leaf_reg;
+	struct scev *scev = env->scev;
+	u32 l, r, leaf, order;
+	s64 imm;
+	int err;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT) {
+			verifier_bug(env, "scev ANY union exceeds expression depth limit");
+			return -EFAULT;
+		}
+		if (!(order & PRE) || is_any(scev, id, &l, &r))
+			continue;
+		if (is_imm(scev, id, &imm)) {
+			bpf_mark_reg_known_scalar(tmp, imm);
+			leaf_reg = tmp;
+		} else if (is_reg(scev, id, &leaf)) {
+			leaf_reg = scev_regno_to_reg(env, loop_entry, leaf);
+		} else {
+			verifier_bug(env, "scev ANY union has an unsupported leaf");
+			return -EFAULT;
+		}
+		err = bpf_reg_union(env, acc, leaf_reg);
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
+int bpf_widen_scev_regs(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			struct bpf_verifier_state *loop_entry, struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *cur_func = st->frame[st->curframe];
+	struct bpf_func_state *entry_func = loop_entry->frame[loop_entry->curframe];
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct bpf_reg_state *reg;
+	struct env *header_env;
+	struct bounds bounds;
+	u32 r, base_reg, r_expr, a, b;
+	int insn_idx = st->insn_idx;
+	s64 slope_imm;
+	int err;
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		if (!scev_reg_alive(env, st, r))
+			continue;
+
+		r_expr = header_env->reg2scev[r];
+		/* If SCEV for r is (linear <reg> <slope>)*/
+		if (is_simple_linear(scev, r_expr, &base_reg, &slope_imm) &&
+		    base_reg == r) {
+			reg = scev_regno_to_reg(env, cur_func, r);
+			/* bpf_compute_loop_iters() checked these bounds for overflow. */
+			if (!linear_bounds(reg, iters, slope_imm, &bounds)) {
+				verifier_bug(env, "scev widen bounds overflow for r%d", r);
+				return -EFAULT;
+			}
+			if (log->level & BPF_LOG_LEVEL2) {
+				bpf_log(log, "loop header at %d, widening ", insn_idx);
+				log_reg(env, r);
+				bpf_log(log, " to %lld..%lld step %u\n",
+					cnum64_smin(bounds.range), cnum64_smax(bounds.range),
+					bounds.step);
+			}
+			scratch_widened_reg_id(env, reg);
+			err = bpf_set_reg_range(env, reg, bounds.range, bounds.base, bounds.step);
+			if (err)
+				return err;
+			mark_scev_reg_scratched(env, r);
+		} else if (is_any(scev, r_expr, &a, &b)) {
+			reg = scev_regno_to_reg(env, cur_func, r);
+			err = union_any_reg(env, entry_func, reg, r_expr);
+			if (err)
+				return err;
+			if (log->level & BPF_LOG_LEVEL2) {
+				bpf_log(log, "loop header at %d, widening ", insn_idx);
+				log_reg(env, r);
+				bpf_log(log, " to %lld..%lld step %u\n",
+					reg_smin(reg), reg_smax(reg), reg->step);
+			}
+			scratch_widened_reg_id(env, reg);
+			mark_scev_reg_scratched(env, r);
+		}
+	}
+	return 1;
+}
+
+int bpf_clamp_scev_regs(struct bpf_verifier_env *env, struct bpf_func_state *cur_func_state, u32 insn_idx,
+			struct bpf_verifier_state *entry_state, struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *entry_st = entry_state->frame[entry_state->curframe];
+	struct bpf_verifier_log *log = &env->log;
+	struct bpf_reg_state *reg, *entry_reg;
+	struct scev *scev = env->scev;
+	struct env *header_env;
+	struct bounds bounds;
+	u32 r, base_reg;
+	s64 slope_imm;
+	int err;
+
+	if (entry_state->curframe != cur_func_state->frameno) {
+		verifier_bug(env, "clamping registers for a wrong frame: %d vs %d\n",
+			     entry_state->curframe, cur_func_state->frameno);
+		return -EFAULT;
+	}
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		/* If SCEV for r is (linear <reg> <slope>)*/
+		if (!is_simple_linear(scev, header_env->reg2scev[r], &base_reg, &slope_imm) ||
+		    base_reg != r)
+			continue;
+
+		entry_reg = scev_regno_to_reg(env, entry_st, r);
+		reg = scev_regno_to_reg(env, cur_func_state, r);
+		if (entry_reg->type == NOT_INIT || reg->type == NOT_INIT)
+			continue;
+
+		if (!linear_bounds(entry_reg, iters, slope_imm, &bounds)) {
+			verifier_bug(env, "scev clamp bounds overflow for r%d", r);
+			return -EFAULT;
+		}
+		bounds.range = cnum64_intersect(reg->r64, bounds.range);
+		if (cnum64_is_empty(bounds.range)) {
+			verifier_bug(env, "scev clamp produced empty range for r%d", r);
+			return -EFAULT;
+		}
+		if (log->level & BPF_LOG_LEVEL2) {
+			bpf_log(log, "loop header at %d, clamping ", insn_idx);
+			log_reg(env, r);
+			bpf_log(log, " to %lld..%lld step %u\n",
+				cnum64_smin(bounds.range), cnum64_smax(bounds.range),
+				bounds.step);
+		}
+		scratch_widened_reg_id(env, reg);
+		err = bpf_set_reg_range(env, reg, bounds.range, bounds.base, bounds.step);
+		if (err)
+			return err;
+		mark_scev_reg_scratched(env, r);
+	}
+	return 0;
+}
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 470bff8d8dc8..94339729f524 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -936,6 +936,27 @@ static bool refsafe(struct bpf_verifier_state *old, struct bpf_verifier_state *c
 	return true;
 }
 
+static bool loop_stack_safe(struct bpf_verifier_env *env, struct bpf_func_state *old,
+			    struct bpf_func_state *cur)
+{
+	struct loop_stack_entry *old_stack = old->loop_stack;
+	struct loop_stack_entry *cur_stack = cur->loop_stack;
+	u32 i;
+
+	if (old->loop_stack_cnt != cur->loop_stack_cnt)
+		return false;
+
+	for (i = 0; i < old->loop_stack_cnt; i++) {
+		if (old_stack[i].loop_id == cur_stack[i].loop_id &&
+		    old_stack[i].iters.max_header_count == cur_stack[i].iters.max_header_count &&
+		    old_stack[i].terminates == cur_stack[i].terminates)
+			continue;
+		return false;
+	}
+
+	return true;
+}
+
 /* compare two verifier states
  *
  * all states stored in state_list are known to be valid, since
@@ -986,6 +1007,9 @@ static bool func_states_equal(struct bpf_verifier_env *env, struct bpf_func_stat
 	if (!stack_arg_safe(env, old, cur, &env->idmap_scratch, exact))
 		return false;
 
+	if (!loop_stack_safe(env, old, cur))
+		return false;
+
 	return true;
 }
 
@@ -1032,6 +1056,7 @@ static bool states_equal(struct bpf_verifier_env *env,
 		if (!func_states_equal(env, old->frame[i], cur->frame[i], insn_idx, exact))
 			return false;
 	}
+
 	return true;
 }
 
@@ -1314,15 +1339,40 @@ int bpf_split_cur_state(struct bpf_verifier_env *env)
 	return 0;
 }
 
+/* Force a checkpoint upon reaching a loop header for a terminating loop. */
+static bool need_loop_checkpoint(struct bpf_verifier_env *env, int insn_idx)
+{
+	struct bpf_verifier_state *cur = env->cur_state;
+	struct bpf_func_state *frame = cur->frame[cur->curframe];
+	struct loop_stack_entry *top;
+
+	if (bpf_loop_at_index(env, insn_idx) != insn_idx || !frame->loop_stack_cnt)
+		return false;
+	top = &frame->loop_stack[frame->loop_stack_cnt - 1];
+	return top->loop_id == insn_idx && top->terminates;
+}
+
+static struct loop_stack_entry *current_loop(struct bpf_verifier_state *st)
+{
+	struct bpf_func_state *frame = st->frame[st->curframe];
+
+	if (frame->loop_stack_cnt == 0)
+		return NULL;
+	return &frame->loop_stack[frame->loop_stack_cnt - 1];
+}
+
 int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 {
+	struct loop_stack_entry *old_loop, *cur_loop;
 	struct bpf_verifier_state_list *sl;
 	struct bpf_verifier_state *cur = env->cur_state;
-	bool force_new_state, add_new_state, loop;
+	bool force_new_state, add_new_state, loop, loop_checkpoint;
 	int n, err, states_cnt = 0;
 	struct list_head *pos, *tmp, *head;
 
+	loop_checkpoint = need_loop_checkpoint(env, insn_idx);
 	force_new_state = env->test_state_freq || bpf_is_force_checkpoint(env, insn_idx) ||
+			  loop_checkpoint ||
 			  /* Avoid accumulating infinitely long jmp history */
 			  cur->jmp_history_cnt > 40;
 
@@ -1353,10 +1403,10 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 			continue;
 
 		if (sl->state.branches) {
-			struct bpf_func_state *frame = sl->state.frame[0];
+			struct bpf_func_state *old_top_frame = sl->state.frame[0];
 
-			if (frame->in_async_callback_fn &&
-			    frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
+			if (old_top_frame->in_async_callback_fn &&
+			    old_top_frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
 				/* Different async_entry_cnt means that the verifier is
 				 * processing another entry into async callback.
 				 * Seeing the same state is not an indication of infinite
@@ -1455,6 +1505,33 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 				}
 				goto skip_inf_loop_check;
 			}
+			/*
+			 * If old state belongs to a control flow loop that we know terminates,
+			 * it should be safe to prune current state. However, account for the
+			 * following situation:
+			 *
+			 *   for (;;) {
+			 *     for (i = 0; i < 10; i++)
+			 *       ...
+			 *   }
+			 *
+			 * Here a backedge to a non-terminating outer loop is dominated by an exit
+			 * from the inner loop. Pruning at a re-entry to the inner loop therefore
+			 * would hide the fact that outer loop is non-terminating.
+			 * Hence, only allow pruning states within the same entry state.
+			 */
+			old_loop = current_loop(&sl->state);
+			cur_loop = current_loop(cur);
+			if (old_loop && cur_loop &&
+			    old_loop->terminates && cur_loop->terminates &&
+			    old_loop->entry_state == cur_loop->entry_state) {
+				if (states_equal(env, &sl->state, cur, RANGE_WITHIN)) {
+					loop = true;
+					goto hit;
+				}
+				goto skip_inf_loop_check;
+			}
+
 			/* attempt to detect infinite loop to avoid unnecessary doomed work */
 			if (states_maybe_looping(&sl->state, cur) &&
 			    states_equal(env, &sl->state, cur, EXACT) &&
@@ -1613,7 +1690,8 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 		 * Use bigger 'n' for checkpoints because evicting checkpoint states
 		 * too early would hinder iterator convergence.
 		 */
-		n = bpf_is_force_checkpoint(env, insn_idx) && sl->state.branches > 0 ? 64 : 3;
+		n = (bpf_is_force_checkpoint(env, insn_idx) || loop_checkpoint) &&
+		    sl->state.branches > 0 ? 64 : 3;
 		if (sl->miss_cnt > sl->hit_cnt * n + n) {
 			/* the state is unlikely to be useful. Remove it to
 			 * speed up verification
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 1fc3e6e2387e..8e92c77b11a2 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -1669,6 +1669,7 @@ static void free_func_state(struct bpf_func_state *state)
 {
 	if (!state)
 		return;
+	kfree(state->loop_stack);
 	kfree(state->stack_arg_regs);
 	kfree(state->stack);
 	kfree(state);
@@ -1705,6 +1706,10 @@ static int copy_func_state(struct bpf_func_state *dst,
 	memcpy(dst, src, offsetof(struct bpf_func_state, stack));
 	/* Instruction accounting is path-local, not part of verifier state. */
 	dst->insns_subtotal = 0;
+	dst->loop_stack = copy_array(dst->loop_stack, src->loop_stack, src->loop_stack_cnt,
+				     sizeof(*src->loop_stack), GFP_KERNEL_ACCOUNT);
+	if (!dst->loop_stack)
+		return -ENOMEM;
 	return copy_stack_state(dst, src);
 }
 
@@ -19770,6 +19775,176 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 	return -EFAULT;
 }
 
+/*
+ * Push the loop entered at insn_idx onto the stack. Usually this adds a single
+ * entry on top of its enclosing loop. However, an edge may enter an inner loop
+ * directly, bypassing the header(s) of its enclosing loop(s), e.g.:
+ *
+ *   1: for (...):       // enclosing loop, header at 1
+ *   2:   for (...):     // inner loop, header at 2
+ *        ...
+ *   3: if ...:
+ *        goto 2b;       // enters loop 2 without going through header 1
+ *
+ * Such a bypass makes the enclosing loop irreducible, so its header is missing
+ * from the stack and both headers (1) and (2) need to be pushed onto stack.
+ * Only the innermost loop (the one actually entered at insn_idx) carries SCEV bounds;
+ * the bypassed ancestors are irreducible and pushed as non-terminating.
+ *
+ * Assumes loop_stack_pop() has already truncated the stack to the common ancestor
+ * of the bpf_loop_at_index(env->insn_idx) and whatever was at the top of the loop stack.
+ */
+static int loop_stack_push(struct bpf_verifier_env *env, bool *pushed)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_func_state *frame = cur_func(env);
+	struct loop_stack_entry *loop_stack = frame->loop_stack;
+	u32 missing_headers[LOOP_STACK_SIZE];
+	u32 cnt = frame->loop_stack_cnt;
+	u32 num_missing = 0;
+	int h;
+
+	*pushed = false;
+
+	for (h = bpf_loop_at_index(env, env->insn_idx); h >= 0; h = aux[h].loop_header, num_missing++) {
+		if (cnt && loop_stack[cnt - 1].loop_id == h)
+			break;
+		if (num_missing == LOOP_STACK_SIZE)
+			goto e2big;
+		missing_headers[num_missing] = h;
+	}
+
+	if (num_missing == 0)
+		return 0;
+	if (cnt + num_missing > LOOP_STACK_SIZE)
+		goto e2big;
+	frame->loop_stack = realloc_array(frame->loop_stack, cnt, cnt + num_missing,
+					 sizeof(*frame->loop_stack));
+	if (!frame->loop_stack)
+		return -ENOMEM;
+	loop_stack = frame->loop_stack;
+
+	for (; num_missing; num_missing--, cnt++) {
+		h = missing_headers[num_missing - 1];
+		loop_stack[cnt] = (struct loop_stack_entry){ .loop_id = h };
+		if (env->log.level & BPF_LOG_LEVEL2)
+			verbose(env, "entering loop %d\n", h);
+		*pushed = true;
+	}
+	frame->loop_stack_cnt = cnt;
+	return 0;
+
+e2big:
+	verbose(env, "Too many nested loops (%d/%d) at %d\n", cnt, num_missing, env->insn_idx);
+	return -E2BIG;
+}
+
+/*
+ * An exit from a loop can cross several nested loops, e.g.:
+ *
+ *   1: for (...):
+ *   2:   for (...):
+ *   3:     for (...):
+ *            if ...:    // before goto the loop stack is [1, 2, 3 <top>]
+ *              goto 1b; // after goto it should become [1]
+ */
+static int loop_stack_pop(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_func_state *frame = cur_func(env);
+	int new_cnt = 0;
+	int h, i;
+
+	/*
+	 * Walk the loop nest of insn_idx outwards (innermost first) and stop at
+	 * the first loop that is present on the stack - that loop is the common
+	 * ancestor and becomes the new top. If no enclosing loop is on the
+	 * stack, drop all loops.
+	 */
+	for (h = bpf_loop_at_index(env, env->insn_idx); h >= 0; h = aux[h].loop_header) {
+		for (i = frame->loop_stack_cnt - 1; i >= 0; i--) {
+			if (frame->loop_stack[i].loop_id == h) {
+				new_cnt = i + 1;
+				goto pop;
+			}
+		}
+	}
+
+pop:
+	if (env->log.level & BPF_LOG_LEVEL2)
+		for (i = frame->loop_stack_cnt - 1; i >= new_cnt; i--)
+			verbose(env, "exiting loop %d\n", frame->loop_stack[i].loop_id);
+	frame->loop_stack_cnt = new_cnt;
+	return 0;
+}
+
+static int maybe_clamp_scev_regs(struct bpf_verifier_env *env)
+{
+	struct bpf_func_state *frame = cur_func(env);
+	struct loop_stack_entry *entry;
+
+	if (frame->loop_stack_cnt == 0) {
+		verifier_bug(env, "%s: loop stack empty at %d", __FUNCTION__, env->insn_idx);
+		return -EFAULT;
+	}
+
+	entry = &frame->loop_stack[frame->loop_stack_cnt - 1];
+	if (!entry->terminates)
+		return 0;
+
+	return bpf_clamp_scev_regs(env, frame, env->insn_idx, entry->entry_state, &entry->iters);
+}
+
+static int handle_loop_entry_exit(struct bpf_verifier_env *env)
+{
+	struct bpf_verifier_state *cur = env->cur_state;
+	struct bpf_func_state *frame = cur_func(env);
+	struct loop_stack_entry *entry;
+	struct bpf_loop_iters iters;
+	bool pushed, terminates;
+	int err;
+
+	err = loop_stack_pop(env);
+	if (err)
+		return err;
+
+	err = loop_stack_push(env, &pushed);
+	if (err)
+		return err;
+
+	if (frame->loop_stack_cnt == 0)
+		return 0;
+
+	entry = &frame->loop_stack[frame->loop_stack_cnt - 1];
+	if (env->insn_idx != entry->loop_id)
+		return 0;
+
+	/* If nothing was pushed and env->insn_idx is a loop header, we've taken a backedge. */
+	if (!pushed)
+		return maybe_clamp_scev_regs(env);
+
+	err = bpf_compute_loop_iters(env, env->cur_state, &iters);
+	if (err < 0)
+		return err;
+
+	terminates = err == 1;
+	if (!terminates)
+		return 0;
+
+	/* Create a loop_entry checkpoint before widening */
+	err = bpf_split_cur_state(env);
+	if (err)
+		return err;
+
+	entry->iters = iters;
+	entry->entry_state = cur->parent;
+	entry->terminates = true;
+	err = bpf_widen_scev_regs(env, cur, cur->parent, &iters);
+	if (err < 0)
+		return err;
+	return 0;
+}
+
 static int do_check(struct bpf_verifier_env *env)
 {
 	bool pop_log = !(env->log.level & BPF_LOG_LEVEL2);
@@ -19784,6 +19959,12 @@ static int do_check(struct bpf_verifier_env *env)
 		struct bpf_insn_aux_data *insn_aux;
 		int err;
 
+		if (signal_pending(current))
+			return -EAGAIN;
+
+		if (need_resched())
+			cond_resched();
+
 		/* reset current history entry on each new instruction */
 		env->cur_hist_ent = NULL;
 
@@ -19830,6 +20011,15 @@ static int do_check(struct bpf_verifier_env *env)
 			}
 		}
 
+		/*
+		 * Possibly widen the registers before creating a checkpoint
+		 * in bpf_is_state_visited(). The next loop iteration will
+		 * have a chance to hit this checkpoint and converge.
+		 */
+		err = handle_loop_entry_exit(env);
+		if (err)
+			return err;
+
 		if (bpf_is_prune_point(env, env->insn_idx)) {
 			err = bpf_is_state_visited(env, env->insn_idx);
 			if (err < 0)
@@ -19855,12 +20045,6 @@ static int do_check(struct bpf_verifier_env *env)
 				return err;
 		}
 
-		if (signal_pending(current))
-			return -EAGAIN;
-
-		if (need_resched())
-			cond_resched();
-
 		if (env->log.level & BPF_LOG_LEVEL2 && do_print_state) {
 			verbose(env, "\nfrom %d to %d%s:",
 				env->prev_insn_idx, env->insn_idx,
@@ -19870,6 +20054,13 @@ static int do_check(struct bpf_verifier_env *env)
 			do_print_state = false;
 		}
 
+		if (bpf_loop_at_index(env, env->insn_idx) >= 0 &&
+		    cur_func(env)->loop_stack_cnt == 0) {
+			verifier_bug(env, "loop stack empty at %d, while inside the loop %d\n",
+				     env->insn_idx, bpf_loop_at_index(env, env->insn_idx));
+			return -EFAULT;
+		}
+
 		if (env->log.level & BPF_LOG_LEVEL) {
 			if (verifier_state_scratched(env))
 				print_insn_state(env, state, state->curframe);
diff --git a/tools/testing/selftests/bpf/progs/verifier_gotox.c b/tools/testing/selftests/bpf/progs/verifier_gotox.c
index dc7baa0f2458..84302a122ccf 100644
--- a/tools/testing/selftests/bpf/progs/verifier_gotox.c
+++ b/tools/testing/selftests/bpf/progs/verifier_gotox.c
@@ -511,7 +511,7 @@ __naked void too_many_gotox_edges(void)
 
 SEC("socket")
 __description("gotox-edges-at-limit")
-__success __retval(0)
+__failure __msg("Too many nested loops")
 __naked void gotox_edges_at_limit(void)
 {
 	asm volatile (
@@ -519,6 +519,10 @@ __naked void gotox_edges_at_limit(void)
 		 * 1000 gotox * 1000 targets = 1,000,000 CFG edges. At run
 		 * time, each block loads the next table entry and jumps to
 		 * the following block, so the program terminates.
+		 *
+		 * Every block can jump to every later block, so the CFG is
+		 * a nest of ~1000 irreducible loops, deeper than the
+		 * verifier's loop stack; the program is rejected.
 		 */
 		GOTOX_TABLE_BEGIN(1000)
 		".rept 1000;"

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (28 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:02   ` sashiko-bot
  2026-10-04 13:38 ` [PATCH bpf-next v2 31/43] bpf: bpf_func_state size optimization Eduard Zingerman
                   ` (14 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

For loops like this:

  int arr[10];
  for (int i = 0; i < 10; i++)
    arr[i] = i;
  use(arr[i]);

SCEV-based loop widening converts `i` from being enumerated 10 times
as a constant value 0, 1, ..., 9 to being enumerated once as a range
[0..9]. From the point of view of the main verification pass, this
replaces constant-offset stack access with varying-offset
stack access.

The verifier tracks stack values accessed through varying offsets much
less precisely than those accessed via constant offsets.
For example, the read at use(arr[i]) would yield STACK_MISC.
Precision might be lost for the following operations:
- spills (BPF_ST/STX to the stack): stored value can no longer be
  recovered when read back, e.g. after the loop;
- fills (BPF_LDX from the stack): a varying-offset read of a spilled
  pointer yields a SCALAR_VALUE, so the pointer is lost.
- calls that construct objects on the stack (dynptr, iter, irq_flag,
  res_spin_lock): the stack argument is required to have a
  constant offset.

This commit adds logic to avoid widening registers used by
such instructions:
- During the main SCEV computation stage, collect_store_base_regs():
  - inspects instructions accessing the stack (spill, fill, call to a
    function listed in bpf_needs_fixed_stack_off), and
  - if there is a SCEV expression for a stack base address register,
    records which registers the expression depends on in
    the bpf_loop->store_base_reg field.
- `store_base_reg`s accumulated for nested loops are propagated
  to outer loops as well.
- During the loop-widening stage, bpf_compute_loop_iters() refuses to
  widen a loop when some of the registers the loop changes appear in
  store_base_regs.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 10 +++++
 kernel/bpf/scev.c            | 96 +++++++++++++++++++++++++++++++++++++++++---
 kernel/bpf/verifier.c        | 49 ++++++++++++++++++++++
 3 files changed, 149 insertions(+), 6 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d13f662dcabc..d47d4613eb3c 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -692,6 +692,9 @@ struct bpf_loop_exit {
 	int to; /* instruction outside the loop */
 };
 
+/* SCEV/widening register space: r0..r10 plus every stack slot of a frame. */
+#define BPF_SCEV_REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+
 struct bpf_loop {
 	struct bpf_backedge backedges[MAX_BACKEDGES];
 	/* edges exiting from this loop, includes edges from nested loops */
@@ -701,6 +704,12 @@ struct bpf_loop {
 	bool irreducible;
 	bool backedges_overflow;
 	bool exits_overflow;
+	/*
+	 * Loop-entry registers used to compute stack addresses for loads,
+	 * stores and calls requiring fixed stack offsets, in this and nested
+	 * loops.
+	 */
+	unsigned long store_base_regs[BITS_TO_LONGS(BPF_SCEV_REGS_NUM)];
 };
 
 struct bpf_insn_aux_data {
@@ -1854,6 +1863,7 @@ int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
 int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
 const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx);
+bool bpf_needs_fixed_stack_off(struct bpf_verifier_env *env, int insn_idx);
 bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 spi);
 int bpf_compute_live_registers(struct bpf_verifier_env *env);
 
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
index a98f3fea52ec..a093e9bf84b9 100644
--- a/kernel/bpf/scev.c
+++ b/kernel/bpf/scev.c
@@ -9,7 +9,7 @@
 #include <linux/tnum.h>
 #include <linux/overflow.h>
 
-#define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+#define REGS_NUM BPF_SCEV_REGS_NUM
 #define UNKNOWN_EXPR_ID 0
 #define OPAQUE_EXPR_ID  1
 
@@ -1060,6 +1060,70 @@ static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct
 	}
 }
 
+/*
+ * Collect loop-entry registers referenced by expr 'id', best effort.
+ * DEPTH_LIMIT may leave dependencies unrecorded. This mask only avoids
+ * widening that would lose stack-access precision; missing dependencies
+ * may cause false rejections, but this is not a soundness issue.
+ */
+static void or_expr_regs(struct bpf_verifier_env *env, u32 id, unsigned long *mask)
+{
+	struct scev *scev = env->scev;
+	u32 order;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if ((order & PRE) && scev->exprs[id].op == REG)
+			__set_bit(scev->exprs[id].params[0], mask);
+	}
+}
+
+/* Mask of argument registers (R1..R5) a call at 'idx' passes by register. */
+static u16 call_params_mask(struct bpf_verifier_env *env, int idx)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct bpf_call_summary cs;
+	int n = bpf_get_call_summary(env, insn, &cs) ? cs.arg_slot_cnt : MAX_BPF_FUNC_REG_ARGS;
+
+	return n ? GENMASK(BPF_REG_1 + n - 1, BPF_REG_1) : 0;
+}
+
+/*
+ * For instructions like:
+ * - *(u64 *)(rBase + off) = rX
+ * - rX = *(u64 *)(rBase + off)
+ * - calls that construct objects on stack (e.g. dynptr_from_mem(rBase, ...))
+ * When 'rBase' can be a stack pointer and is derived from some registers Rs
+ * defined at loop entry, record Rs into 'mask'.
+ */
+static void collect_store_base_regs(struct bpf_verifier_env *env,
+				    struct env *cur_env, int idx, unsigned long *mask)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	u8 class = BPF_CLASS(insn->code);
+	u8 size = BPF_SIZE(insn->code);
+	u16 base_regs = 0;
+	u32 r;
+
+	if (size == BPF_W || size == BPF_DW) {
+		if ((class == BPF_STX || class == BPF_ST) && insn->dst_reg != BPF_REG_FP)
+			base_regs |= BIT(insn->dst_reg);
+		else if (class == BPF_LDX && insn->src_reg != BPF_REG_FP)
+			base_regs |= BIT(insn->src_reg);
+	}
+
+	if (class == BPF_JMP && BPF_OP(insn->code) == BPF_CALL &&
+	    bpf_needs_fixed_stack_off(env, idx))
+		base_regs |= call_params_mask(env, idx);
+
+	base_regs &= aux[idx].stack_ptrs;
+	for (r = 0; r < MAX_BPF_REG; r++)
+		if (base_regs & BIT(r))
+			or_expr_regs(env, cur_env->reg2expr[r], mask);
+}
+
 /* Find the topmost loop header containing idx inside cur_header, or -1 if none. */
 static int topmost_nested_loop(struct bpf_verifier_env *env, int idx, int cur_header)
 {
@@ -1144,6 +1208,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 	struct bpf_iarray *succ;
 	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
 	int s, i, err, idx, succ_idx;
+	u32 r;
 
 	if (log_level2)
 		bpf_log(&env->log, "Computing SCEV for loop at %d:\n", cur_header);
@@ -1157,6 +1222,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 		if (!header_env)
 			return -ENOMEM;
 		/* The freshly allocated environment has all expressions unknown. */
+		bitmap_fill(cur_loop->store_base_regs, REGS_NUM);
 		header_env->empty = false;
 		return 0;
 	}
@@ -1193,6 +1259,9 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 			 */
 			if (log_level2)
 				memcpy(old_env, cur_env, sizeof(*old_env));
+			/* Pull the nested loop's stack-store base dependencies up. */
+			for_each_set_bit(r, nested_loop->store_base_regs, BPF_SCEV_REGS_NUM)
+				or_expr_regs(env, cur_env->reg2expr[r], cur_loop->store_base_regs);
 			forget_non_invariants(env, cur_env, idx);
 			if (log_level2)
 				log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
@@ -1214,6 +1283,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 			for (;;) {
 				if (log_level2)
 					memcpy(old_env, cur_env, sizeof(*old_env));
+				collect_store_base_regs(env, cur_env, idx, cur_loop->store_base_regs);
 				err = transfer(env, cur_env, idx);
 				if (err)
 					goto out;
@@ -2232,6 +2302,7 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 	for (r = 0; r < REGS_NUM; r++) {
 		struct bpf_reg_state *reg;
 		u32 ra, r_scev, r_expr;
+		bool spill_base = false;
 
 		if (!scev_reg_alive(env, st, r))
 			continue;
@@ -2241,10 +2312,16 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 		/* If SCEV for r is (linear <reg> <slope>) */
 		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
 		    is_widenable_reg_type(reg)) {
-			/* Can't widen if the iteration range overflows */
-			if (base_reg == r &&
-			    !linear_bounds(reg, iters, slope_imm, &bounds))
-				goto cant_widen;
+			if (base_reg == r) {
+				/* Can't widen if the iteration range overflows */
+				if (!linear_bounds(reg, iters, slope_imm, &bounds))
+					goto cant_widen;
+				/* Spills at varying offsets lose precision */
+				if (test_bit(r, loop->store_base_regs)) {
+					spill_base = true;
+					goto cant_widen;
+				}
+			}
 			continue;
 		}
 
@@ -2253,8 +2330,13 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 			continue;
 
 		/* (any 1 (any 2 (any 3 4))) */
-		if (is_any_imm_reg(env, cur_func, header_env, r, r_scev))
+		if (is_any_imm_reg(env, cur_func, header_env, r, r_scev)) {
+			if (test_bit(r, loop->store_base_regs)) {
+				spill_base = true;
+				goto cant_widen;
+			}
 			continue;
+		}
 
 cant_widen:
 		if (log->level & BPF_LOG_LEVEL2) {
@@ -2263,6 +2345,8 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 			log_reg(env, r);
 			bpf_log(log, ", expr is ");
 			log_expr(env, r_expr);
+			if (spill_base)
+				bpf_log(log, ", requires exact stack-offset tracking");
 			bpf_log(log, "\n");
 		}
 		return 0;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 8e92c77b11a2..e4780c8bef13 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14211,6 +14211,55 @@ static bool kfunc_spin_allowed(struct bpf_verifier_env *env, s32 func_id, s16 of
 	return *kfunc.flags & KF_SPINLOCK_SAFE;
 }
 
+/*
+ * True if insn calls a helper/kfunc that requires one of its arguments to
+ * be a stack pointer with a constant offset.
+ */
+bool bpf_needs_fixed_stack_off(struct bpf_verifier_env *env, int insn_idx)
+{
+	const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+	const struct bpf_func_proto *fn;
+	struct bpf_kfunc_desc *desc;
+	u32 *flags, btf_id;
+	int i;
+
+	if (bpf_helper_call(insn)) {
+		if (bpf_get_helper_proto(env, insn->imm, &fn) < 0)
+			return false;
+	} else if (bpf_pseudo_kfunc_call(insn)) {
+		desc = find_kfunc_desc(env->prog, insn->imm, insn->off);
+		if (!desc)
+			return false;
+		fn = &desc->proto;
+	} else {
+		return false;
+	}
+
+	/* Both initialized and uninitialized stack dynptrs need a fixed offset. */
+	for (i = 0; i < ARRAY_SIZE(fn->arg_type); i++)
+		if (arg_type_is_dynptr(fn->arg_type[i]))
+			return true;
+
+	/* vmlinux kfuncs only */
+	if (!bpf_pseudo_kfunc_call(insn) || insn->off != 0)
+		return false;
+	btf_id = insn->imm;
+
+	flags = btf_kfunc_flags(btf_vmlinux, btf_id, env->prog);
+	if (flags && (*flags & (KF_ITER_NEW | KF_ITER_NEXT | KF_ITER_DESTROY)))
+		return true;
+
+	if (btf_id == special_kfunc_list[KF_bpf_res_spin_lock] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_unlock] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_lock_irqsave] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_unlock_irqrestore] ||
+	    btf_id == special_kfunc_list[KF_bpf_local_irq_save] ||
+	    btf_id == special_kfunc_list[KF_bpf_local_irq_restore])
+		return true;
+
+	return false;
+}
+
 static bool is_sync_callback_calling_kfunc(u32 btf_id)
 {
 	return is_bpf_rbtree_add_kfunc(btf_id);

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 31/43] bpf: bpf_func_state size optimization
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (29 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
                   ` (13 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Pack bpf_reg_state to bring it's size from 88 bytes to 80.
Consequently, bpf_func_state changes from 1048 to 960.
Thus bpf_func_state moves from 2K to 1K allocator bucket.
1K bucket is the one it fit before the SCEV series.

With this optimization SCEV increases memory consumption on a subset
of Meta internal programs by 3%, without this optimization the
increase is 48%.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 14 +++++++-------
 kernel/bpf/states.c          |  2 +-
 2 files changed, 8 insertions(+), 8 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d47d4613eb3c..a539c5f049ff 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -172,12 +172,6 @@ struct bpf_reg_state {
 	 * gets parent_id set to the dynptr's id.
 	 */
 	u32 parent_id;
-	/*
-	 * Distinguishes inner-map lookups and their keys and values. Zero for
-	 * other registers. Kept outside the metadata union for ID remapping
-	 * during state comparisons.
-	 */
-	u32 map_uid;
 	/*
 	 * The value described by this register, interpreted as s64, lies on
 	 * a line described by a linear equation base + step * k.
@@ -185,8 +179,14 @@ struct bpf_reg_state {
 	 */
 	u16 base;
 	u16 step;
+	/*
+	 * Distinguishes inner-map lookups and their keys and values. Zero for
+	 * other registers. Kept outside the metadata union for ID remapping
+	 * during state comparisons.
+	 */
+	u32 map_uid:31;
 	/* if (!precise && SCALAR_VALUE) min/max/tnum don't affect safety */
-	bool precise;
+	u32 precise:1;
 };
 
 static inline s64 reg_smin(const struct bpf_reg_state *reg)
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 94339729f524..862f7ceb1daf 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -1171,7 +1171,7 @@ static bool states_maybe_looping(struct bpf_verifier_state *old,
 	fcur = cur->frame[fr];
 	for (i = 0; i < MAX_BPF_REG; i++)
 		if (memcmp(&fold->regs[i], &fcur->regs[i],
-			   offsetof(struct bpf_reg_state, precise)))
+			   offsetofend(struct bpf_reg_state, step)))
 			return false;
 	return true;
 }

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (30 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 31/43] bpf: bpf_func_state size optimization Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 33/43] selftests/bpf: bound UNIX socket path loops by sun_path size Eduard Zingerman
                   ` (12 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Same as __msg, but expects the match on the line immediately after
the last match, not just any later line.
E.g.:

  __msg_next("a")
  __msg_next("c")

This would match consecutive output "a\n" "c\n",
but would not match output "a\n" "b\n" "c\n".

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/progs/bpf_misc.h |  6 ++++++
 tools/testing/selftests/bpf/test_loader.c    | 10 ++++++++++
 2 files changed, 16 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..d7f706c3bc4a 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -43,6 +43,10 @@
  * __msg_unpriv      Same as __msg but for unprivileged mode.
  * __not_msg_unpriv  Same as __not_msg but for unprivileged mode.
  *
+ * __msg_next        Same as __msg but expects the match to be on the next
+ * __msg_next_unpriv log line compared to the last match, not just some line
+ *                   after the last match.
+ *
  * __stderr          Message expected to be found in bpf stderr stream. The
  *                   same regex rules apply like __msg.
  * __stderr_unpriv   Same as __stderr but for unpriveleged mode.
@@ -145,6 +149,7 @@
 
 #define __msg(msg)		__test_tag("test_expect_msg=" msg)
 #define __not_msg(msg)		__test_tag("test_expect_not_msg=" msg)
+#define __msg_next(msg)		__test_tag("test_expect_msg_next=" msg)
 #define __xlated(msg)		__test_tag("test_expect_xlated=" msg)
 #define __jited(msg)		__test_tag("test_jited=" msg)
 #define __failure		__test_tag("test_expect_failure")
@@ -153,6 +158,7 @@
 #define __skip(reason)		__test_tag("test_skip=" reason)
 #define __msg_unpriv(msg)	__test_tag("test_expect_msg_unpriv=" msg)
 #define __not_msg_unpriv(msg)	__test_tag("test_expect_not_msg_unpriv=" msg)
+#define __msg_next_unpriv(msg)	__test_tag("test_expect_msg_next_unpriv=" msg)
 #define __xlated_unpriv(msg)	__test_tag("test_expect_xlated_unpriv=" msg)
 #define __jited_unpriv(msg)	__test_tag("test_jited_unpriv=" msg)
 #define __failure_unpriv	__test_tag("test_expect_failure_unpriv")
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index a89890cd56d8..6257bbedc4c0 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -508,6 +508,16 @@ static int parse_test_spec(struct test_loader *tester,
 			if (err)
 				goto cleanup;
 			spec->mode_mask |= UNPRIV;
+		} else if ((msg = str_has_pfx(s, "test_expect_msg_next="))) {
+			err = __push_msg(msg, true, false, &spec->priv.expect_msgs);
+			if (err)
+				goto cleanup;
+			spec->mode_mask |= PRIV;
+		} else if ((msg = str_has_pfx(s, "test_expect_msg_next_unpriv="))) {
+			err = __push_msg(msg, true, false, &spec->unpriv.expect_msgs);
+			if (err)
+				goto cleanup;
+			spec->mode_mask |= UNPRIV;
 		} else if ((msg = str_has_pfx(s, "test_jited="))) {
 			if (arch_mask == 0) {
 				PRINT_FAIL("__jited used before __arch_*");

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 33/43] selftests/bpf: bound UNIX socket path loops by sun_path size
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (31 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 34/43] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
                   ` (11 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

The bpf_iter/unix and skc_to_unix_sock tests index sun_path but bound
that index by sizeof(struct sockaddr_un). This allows indices 108 and
109 into the 108-byte array. With SCEV loop widening, the variable-offset
BTF access check rejects the resulting range.

Bound the index by sizeof(sun_path) in both tests. Runtime behavior for
valid UNIX addresses is unchanged.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/progs/bpf_iter_unix.c         | 2 +-
 tools/testing/selftests/bpf/progs/test_skc_to_unix_sock.c | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/tools/testing/selftests/bpf/progs/bpf_iter_unix.c b/tools/testing/selftests/bpf/progs/bpf_iter_unix.c
index df3d971b21e8..3061b148f338 100644
--- a/tools/testing/selftests/bpf/progs/bpf_iter_unix.c
+++ b/tools/testing/selftests/bpf/progs/bpf_iter_unix.c
@@ -70,7 +70,7 @@ int dump_unix(struct bpf_iter__unix *ctx)
 
 			for (i = 1; i < len; i++) {
 				/* unix_validate_addr() tests this upper bound. */
-				if (i >= sizeof(struct sockaddr_un))
+				if (i >= sizeof(unix_sk->addr->name->sun_path))
 					break;
 
 				BPF_SEQ_PRINTF(seq, "%c",
diff --git a/tools/testing/selftests/bpf/progs/test_skc_to_unix_sock.c b/tools/testing/selftests/bpf/progs/test_skc_to_unix_sock.c
index 4cfa42aa9436..c5c643064eb5 100644
--- a/tools/testing/selftests/bpf/progs/test_skc_to_unix_sock.c
+++ b/tools/testing/selftests/bpf/progs/test_skc_to_unix_sock.c
@@ -29,7 +29,7 @@ int BPF_PROG(unix_listen, struct socket *sock, int backlog)
 	len = unix_sk->addr->len - sizeof(short);
 	path[0] = '@';
 	for (i = 1; i < len; i++) {
-		if (i >= (int)sizeof(struct sockaddr_un))
+		if (i >= (int)sizeof(unix_sk->addr->name->sun_path))
 			break;
 
 		path[i] = unix_sk->addr->name->sun_path[i];

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 34/43] selftests/bpf: test for stack-pointer subrange pruning
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (32 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 33/43] selftests/bpf: bound UNIX socket path loops by sun_path size Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
                   ` (10 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Validate changes from:
"bpf: allow subrange relations for PTR_TO_STACK in regsafe()".
Add a test case that checks that a widened stack pointer with range
[-72, -16] is considered a superset of a stack pointer with range
[-64, -16].

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_stack_ptr.c       | 36 ++++++++++++++++++++++
 1 file changed, 36 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
index 3e0bea9819ca..3e9d8fc606ae 100644
--- a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
+++ b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
@@ -586,4 +586,40 @@ __naked void stack_check_size_512_with_may_goto(void)
 }
 #endif
 
+/*
+ * Verify that old PTR_TO_STACK state is considered a super-set of
+ * new PTR_TO_STACK state when new variable range is a sub-range
+ * of the old range, e.g. old [-72, -16] vs new [-64, -16].
+ */
+SEC("socket")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+/*
+ * r7 is widened to [-72, -16] at the loop header (insn 3),
+ * the loop body sees [-64, -8] after 'r7 += 8'
+ */
+__msg("loop header at 3, widening r7 to -72..-16 step 8")
+__msg("R7=fp(smin=smin32=-64,smax=smax32=-8")
+/*
+ * back-edge state is clamped to the remaining iterations and pruned
+ * at the header, because [-64, -16] is within the widened [-72, -16]
+ */
+__msg("loop header at 3, clamping r7 to -64..-16 step 8")
+__msg("from 5 to 3: safe")
+__msg("processed 9 insns")
+__naked void stack_ptr_subrange_in_loop(void)
+{
+	asm volatile ("					\
+	r7 = r10;					\
+	r7 += -72;					\
+	r6 = 0;						\
+	r7 += 8;					\
+	r6 += 1;					\
+	if r6 < 8 goto -3;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
 char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (33 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 34/43] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 36/43] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
                   ` (9 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Branches to cover:
  - per-insn writes (record_stack_access_off):
      * off_cnt==1, fully covered slots   -> def + may_def (equal)
      * partial coverage                  -> may_def only (no def)
      * off_cnt 2-4, precise offsets      -> may_def only, per offset
  - whole-frame writes:
      * off_cnt==0 (offset lost)          -> may_def SPIS_ALL
      * ARG_IMPRECISE (frame lost)        -> may_def SPIS_ALL per frame
  - merge_instances():
      * both passes touch a frame         -> may_def union
      * one pass leaves a frame untouched -> may_def preserved
  - merge_may_write(): a callee's writes into ancestor frames are
    summarized as may_def at the caller's call instruction
  - analyze_subprog()                     -> may_def stack slot range
                                             when callback is unknown.

Tests exercising them:
  - may_write_two_precise_offsets  - off_cnt 2-4 precise multi-offset
  - imprecise_frame_write          - cross-frame ARG_IMPRECISE write,
                                     plus its caller-side summarization
  - merge_preserves_may_write      - map-value vs stack pointer at two
                                     callsites: union is preserved when
                                     one pass does not touch the frame
  - must_write_not_same_slot       - off_cnt==0 whole-frame may_def
  - two_byte_write_no_kill         - partial write, may_def without def
  - kfunc_iter_stack_liveness      - full-slot def == may_def
  - helper_output_unknown_size     - output size unknown to liveness,
                                     may_def without use or def
  - caller_stack_write,
    conditional_stx_in_subprog,
    shared_instance_must_write_overwrite
                                   - callee writes summarized as may_def
                                     at the call site
  - bpf_loop_two_callbacks         - call to unknown callback marks full
                                     stack range as may_def.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_live_stack.c      | 180 ++++++++++++++++++++-
 1 file changed, 175 insertions(+), 5 deletions(-)

diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 0e6f25a09b5d..a06156c18c9e 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -83,6 +83,31 @@ __naked void must_write_not_same_slot(void)
 	: __clobber_all);
 }
 
+SEC("socket")
+__log_level(2)
+__msg("stack use/def subprog#0 may_write_two_precise_offsets (d0,cs0):")
+/*
+ * r1 is fp-8 or fp-16, with both offsets tracked precisely (off_cnt == 2).
+ * A write through it cannot set def, but marks both candidate slots as may_def.
+ */
+__msg("6: (7b) *(u64 *)(r1 +0) = r0         ; may_def: fp0-8 fp0-16")
+__naked void may_write_two_precise_offsets(void)
+{
+	asm volatile (
+	"call %[bpf_get_prandom_u32];"
+	"r1 = r10;"
+	"if r0 > 42 goto 1f;"
+	"r1 += -8;"
+	"goto 2f;"
+"1:"
+	"r1 += -16;"
+"2:"
+	"*(u64 *)(r1 + 0) = r0;"
+	"exit;"
+	:: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
 SEC("socket")
 __log_level(2)
 __msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
@@ -114,7 +139,7 @@ __log_level(2)
 __msg("stack use/def subprog#0 caller_stack_write (d0,cs0):")
 __msg("2: (85) call pc+1                    ; may_def: fp0-8")
 __msg("stack use/def subprog#1 write_first_param (d1,cs2):")
-__msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8")
+__msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8 may_def: fp0-8")
 __naked void caller_stack_write(void)
 {
 	asm volatile (
@@ -134,6 +159,49 @@ static __used __naked void write_first_param(void)
 	::: __clobber_all);
 }
 
+/*
+ * Cross-frame imprecise write: imprecise_frame_writer() receives a pointer
+ * into the caller's frame (frame 0) but conditionally replaces it with a
+ * pointer into its own frame (frame 1). At the store the two are joined into
+ * ARG_IMPRECISE (frame is unknown, mask = {0,1}), so the offset is dropped and
+ * the whole of both candidate frames is conservatively marked as may_def.
+ */
+SEC("socket")
+__log_level(2)
+/* The callee's write into the caller frame is summarized at the call site. */
+__msg("stack use/def subprog#0 imprecise_frame_write (d0,cs0):")
+__msg("2: (85) call pc+2                    ; may_def: fp0-8..-{{(512|2048)}}")
+__msg("stack use/def subprog#1 imprecise_frame_writer (d1,cs2):")
+__msg("12: (7b) *(u64 *)(r1 +0) = r2         ; may_def: fp1-8..-{{(512|2048)}} fp0-8..-{{(512|2048)}}")
+__naked void imprecise_frame_write(void)
+{
+	asm volatile (
+	"r1 = r10;"
+	"r1 += -8;"
+	"call imprecise_frame_writer;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked void imprecise_frame_writer(void)
+{
+	asm volatile (
+	"r6 = r1;"			/* save arg: pointer into frame 0 */
+	"call %[bpf_get_prandom_u32];"
+	"r1 = r6;"			/* r1 = frame0-8 (arg) */
+	"if r0 == 0 goto 1f;"
+	"r1 = r10;"
+	"r1 += -8;"			/* r1 = frame1-8 (own frame) */
+"1:"
+	"r2 = 0;"
+	"*(u64 *)(r1 + 0) = r2;"	/* join frame0-8 | frame1-8 -> imprecise */
+	"r0 = 0;"
+	"exit;"
+	:: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
 SEC("socket")
 __log_level(2)
 __msg("stack use/def subprog#0 caller_stack_read (d0,cs0):")
@@ -1011,7 +1079,7 @@ __naked void four_byte_read_upper_half(void)
  */
 SEC("socket")
 __log_level(2)
-__msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
+__msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8 may_def: fp0-8")
 /* 2-byte write only partially covers the upper half: may_def, but no def. */
 __msg("1: (6a) *(u16 *)(r10 -4) = 0         ; may_def: fp0-4h")
 __msg("2: (61) r0 = *(u32 *)(r10 -4)        ; use: fp0-4h")
@@ -1364,7 +1432,7 @@ __log_level(2)
  * fp-8 live at call: callee conditionally writes it, so the slot is not killed
  * (no def), but the conditional write surfaces as may_def at the call site.
  */
-__msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8")
+__msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8 may_def: fp0-8")
 __msg("4: (85) call pc+2                    ; may_def: fp0-8")
 __msg("5: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void conditional_stx_in_subprog(void)
@@ -2271,11 +2339,19 @@ static __used __naked void merge_leaf_read(void)
 	::: __clobber_all);
 }
 
-/* Same bpf_loop instruction calls different callbacks depending on branch. */
+/*
+ * Same bpf_loop instruction calls different callbacks depending on branch.
+ * One callback writes through ctx. With no unique callback, all frames must
+ * be marked as possibly read and written, including at enclosing call sites.
+ */
 SEC("socket")
 __log_level(2)
 __success
-__msg("call bpf_loop#181            ; use: fp2-8..-{{(512|2048)}} fp1-8..-{{(512|2048)}} fp0-8..-{{(512|2048)}}")
+__msg("stack use/def subprog#0 bpf_loop_two_callbacks (d0,cs0):")
+__msg("5: (85) call pc+{{.*}};{{.*}} may_def: fp0-8..-{{(512|2048)}}{{$}}")
+__msg("8: (85) call pc+{{.*}};{{.*}} may_def: fp0-8..-{{(512|2048)}}{{$}}")
+__msg("call bpf_loop#181            ; use: fp2-8..-{{(512|2048)}} fp1-8..-{{(512|2048)}} fp0-8..-{{(512|2048)}}"
+      " may_def: fp2-8..-{{(512|2048)}} fp1-8..-{{(512|2048)}} fp0-8..-{{(512|2048)}}{{$}}")
 __naked void bpf_loop_two_callbacks(void)
 {
 	asm volatile (
@@ -2393,6 +2469,20 @@ __flag(BPF_F_TEST_STATE_FREQ)
 __msg("subprog#2 write_first_read_second:")
 __msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
 __msg("18: (79) r0 = *(u64 *)(r2 +0) // r1=fp0-8 r2=fp0-16{{$}}")
+/*
+ * The callee's write into the caller frame is summarized as may_def at the
+ * call sites, transitively across both frames (depth 2 -> 1 -> 0). The two
+ * call sites differ because the shared write_first_read_second instance
+ * accumulates the cross-pass union ({fp-8, fp-16}) at different points
+ * relative to each forwarding_rw's summarization.
+ */
+__msg("stack use/def subprog#0 shared_instance_must_write_overwrite (d0,cs0):")
+__msg("7: (85) call pc+7                    ; use: fp0-8 fp0-16 may_def: fp0-8 fp0-16")
+__msg("12: (85) call pc+2                    ; use: fp0-8 may_def: fp0-16")
+__msg("stack use/def subprog#1 forwarding_rw (d1,cs7):")
+__msg("15: (85) call pc+1                    ; use: fp0-8 fp0-16 may_def: fp0-8 fp0-16")
+__msg("stack use/def subprog#1 forwarding_rw (d1,cs12):")
+__msg("15: (85) call pc+1                    ; use: fp0-8 may_def: fp0-16")
 __msg("stack use/def subprog#2 write_first_read_second (d2,cs15):")
 /*
  * Shared across two callsites with swapped args (r1 is fp-8 on one pass,
@@ -2441,6 +2531,65 @@ static __used __naked void write_first_read_second(void)
 	::: __clobber_all);
 }
 
+/*
+ * merge_instances() must preserve may_write when one pass doesn't touch a frame.
+ * mvs_leaf is reached through a single callsite (mvs_mid) from two outer chains,
+ * so it is one shared (depth,callsite) instance analyzed twice and merged:
+ * - chain A passes r2 = map value, mvs_leaf's doesn't touch stack;
+ * - chain B passes r2 = caller stack slot fp-8.
+ * The merge must keep the stack pass's may_write even though the map pass left
+ * frame 0 untouched.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("stack use/def subprog#2 mvs_leaf (d2,cs19):")
+__msg("21: (7a) *(u64 *)(r2 +0) = 42         ; may_def: fp0-8")
+__naked void merge_preserves_may_write(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 24) = r1;"		/* map key */
+	"r1 = %[map] ll;"
+	"r2 = r10;"
+	"r2 += -24;"
+	"call %[bpf_map_lookup_elem];"
+	"if r0 == 0 goto 1f;"
+	/* chain A: r2 = map value (ARG_NONE, not stack) */
+	"r1 = r10;"
+	"r1 += -16;"				/* unused fp anchor */
+	"r2 = r0;"
+	"call mvs_mid;"
+	/* chain B: r2 = caller stack slot fp-8 */
+	"r1 = r10;"
+	"r1 += -16;"
+	"r2 = r10;"
+	"r2 += -8;"
+	"call mvs_mid;"
+"1:"
+	"r0 = 0;"
+	"exit;"
+	: : __imm(bpf_map_lookup_elem), __imm_addr(map)
+	: __clobber_all);
+}
+
+static __used __naked void mvs_mid(void)
+{
+	asm volatile (
+	"call mvs_leaf;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked void mvs_leaf(void)
+{
+	asm volatile (
+	"*(u64 *)(r2 + 0) = 42;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
 /*
  * Shared must_write when (callsite, depth) instance is reused.
  * Main calls fwd_to_stale_wr at two sites. fwd_to_stale_wr calls
@@ -2949,3 +3098,24 @@ __naked void spill_below_512_stays_precise(void)
 	"exit;"
 	::: __clobber_all);
 }
+
+/* Unknown-size outputs still carry a possible write, without a read or kill. */
+SEC("socket")
+__success __log_level(2)
+__msg("call bpf_probe_read_kernel{{.*}}; may_def: fp0-8{{$}}")
+__naked void helper_output_unknown_size(void)
+{
+	asm volatile (
+	"call %[bpf_get_prandom_u32];"
+	"r0 &= 7;"
+	"r0 += 1;"
+	"r2 = r0;"
+	"r1 = r10;"
+	"r1 += -8;"
+	"r3 = 0;"
+	"call %[bpf_probe_read_kernel];"
+	"r0 = 0;"
+	"exit;"
+	:: __imm(bpf_get_prandom_u32), __imm(bpf_probe_read_kernel)
+	: __clobber_all);
+}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 36/43] selftests/bpf: tests for may_def marks of atomic RMW operations
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (34 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes() Eduard Zingerman
                   ` (8 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Liveness analysis should produce may_def marks for read-modify-write
atomic operations:
- lock *(u64 *)(r10 -8) += r1
- r1 = atomic64_fetch_add((u64 *)(r10 -8), r1)
- r1 = atomic64_xchg((u64 *)(r10 -8), r1)
- r0 = atomic64_cmpxchg((u64 *)(r10 -8), r0, r1)

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_live_stack.c      | 26 ++++++++++++++++++++++
 1 file changed, 26 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index a06156c18c9e..5b0fcef1f571 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -2674,6 +2674,32 @@ __naked void load_acquire_dont_clear_dst(void)
 
 #endif /* CAN_USE_LOAD_ACQ_STORE_REL */
 
+SEC("socket")
+__log_level(2)
+__success
+__msg("lock *(u64 *)(r10 -8) += r1{{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r1 = atomic64_fetch_add((u64 *)(r10 -8), r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r1 = atomic64_xchg((u64 *)(r10 -8), r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r0 = atomic64_cmpxchg((u64 *)(r10 -8), r0, r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__naked void atomic_rmw(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 8) = r1;"
+	".8byte %[atomic_add];"
+	".8byte %[atomic_fetch_add];"
+	".8byte %[atomic_xchg];"
+	"r0 = 0;"
+	".8byte %[atomic_cmpxchg];"
+	"exit;"
+	:
+	: __imm_insn(atomic_add, BPF_ATOMIC_OP(BPF_DW, BPF_ADD, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_fetch_add, BPF_ATOMIC_OP(BPF_DW, BPF_ADD | BPF_FETCH, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_xchg, BPF_ATOMIC_OP(BPF_DW, BPF_XCHG, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_cmpxchg, BPF_ATOMIC_OP(BPF_DW, BPF_CMPXCHG, BPF_REG_10, BPF_REG_1, -8))
+	: __clobber_all);
+}
+
 SEC("socket")
 __success
 __naked void imprecise_fill_loses_cross_frame(void)

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes()
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (35 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 36/43] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
                   ` (7 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Tests for helper arguments whose access size or read/write effects
depend on the map:

- helper_map_output_known_size checks that for a known map access
  MEM_WRITE | MEM_UNINIT produces a must_write mark.
- helper_map_output_merged_sizes checks that when maps with different
  value sizes merge, MEM_WRITE | MEM_UNINIT produces use and may_def
  marks over the maximum value size, without a def mark.
- helper_bloom_peek_input checks that for bpf_map_peek_elem()
  on a bloom filter the buffer is treated as an input.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_live_stack.c      | 76 ++++++++++++++++++++++
 1 file changed, 76 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 5b0fcef1f571..d93bb628f65e 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -3145,3 +3145,79 @@ __naked void helper_output_unknown_size(void)
 	:: __imm(bpf_get_prandom_u32), __imm(bpf_probe_read_kernel)
 	: __clobber_all);
 }
+
+struct {
+	__uint(type, BPF_MAP_TYPE_QUEUE);
+	__uint(max_entries, 1);
+	__type(value, __u64);
+} queue_8b SEC(".maps");
+
+struct {
+	__uint(type, BPF_MAP_TYPE_QUEUE);
+	__uint(max_entries, 1);
+	__type(value, __u64[2]);
+} queue_16b SEC(".maps");
+
+struct {
+	__uint(type, BPF_MAP_TYPE_BLOOM_FILTER);
+	__uint(max_entries, 16);
+	__type(value, __u64);
+} bloom_8b SEC(".maps");
+
+SEC("socket")
+__success __log_level(2)
+__msg("call bpf_map_pop_elem{{.*}}; def: fp0-16")
+__naked void helper_map_output_known_size(void)
+{
+	asm volatile (
+	"r1 = %[queue_8b] ll;"
+	"r2 = r10;"
+	"r2 += -16;"
+	"call %[bpf_map_pop_elem];"
+	"r0 = 0;"
+	"exit;"
+	:: __imm_addr(queue_8b), __imm(bpf_map_pop_elem)
+	: __clobber_all);
+}
+
+/* The maximum map value size is not a guaranteed output extent. */
+SEC("socket")
+__success __log_level(2)
+__msg("call bpf_map_pop_elem{{.*}}; use: fp0-16 fp0-24 may_def: fp0-16 fp0-24{{$}}")
+__naked void helper_map_output_merged_sizes(void)
+{
+	asm volatile (
+	"*(u64 *)(r10 - 16) = 0;"
+	"*(u64 *)(r10 - 24) = 0;"
+	"call %[bpf_get_prandom_u32];"
+	"r1 = %[queue_8b] ll;"
+	"if r0 == 0 goto 1f;"
+	"r1 = %[queue_16b] ll;"
+"1:"
+	"r2 = r10;"
+	"r2 += -24;"
+	"call %[bpf_map_pop_elem];"
+	"r0 = 0;"
+	"exit;"
+	:: __imm_addr(queue_8b), __imm_addr(queue_16b),
+	   __imm(bpf_get_prandom_u32), __imm(bpf_map_pop_elem)
+	: __clobber_all);
+}
+
+/* Bloom-filter peek treats buffer as an input despite the helper's output annotation. */
+SEC("socket")
+__success __log_level(2)
+__msg("call bpf_map_peek_elem{{.*}}; use: fp0-8{{$}}")
+__naked void helper_bloom_peek_input(void)
+{
+	asm volatile (
+	"*(u64 *)(r10 - 8) = 0;"
+	"r1 = %[bloom_8b] ll;"
+	"r2 = r10;"
+	"r2 += -8;"
+	"call %[bpf_map_peek_elem];"
+	"r0 = 0;"
+	"exit;"
+	:: __imm_addr(bloom_8b), __imm(bpf_map_peek_elem)
+	: __clobber_all);
+}

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (36 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes() Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
                   ` (6 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

- step_mul_non_pow2: Retain step 3 after multiplying a bounded scalar
  by a non-power-of-two constant.
- step_lsh: Record step 4 after shifting a bounded scalar left by two.
- step_add_const_base: Shift the base to 1 while preserving step 4
  after adding a constant.
- step_neg_value_range: Preserve the signed range of {-3, 0, 3, 6}
  when intersecting with a non-power-of-two step across zero.
- step_arith{32,64}_overflow: Reset the stride and preserve valid
  bounds when a {32,64}-bit operation wraps.
- step_mov32_truncate: Reset the stride when MOV32 discards nonzero
  upper bits from positive values.
- step_mov32_negative: Reset the stride when MOV32 zero-extends
  the low 32 bits of negative signed values.
- step_movsx: Reset the stride when sign extension changes the
  register value.
- step_linked_regs_add64: Translate both bounds and base when
  propagating a comparison's refinement through a scalar-ID delta.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../selftests/bpf/progs/verifier_bounds_step.c     | 290 +++++++++++++++++++++
 2 files changed, 292 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 460ad10ddc02..4f3af1cb998b 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -20,6 +20,7 @@
 #include "verifier_basic_stack.skel.h"
 #include "verifier_bitfield_write.skel.h"
 #include "verifier_bounds.skel.h"
+#include "verifier_bounds_step.skel.h"
 #include "verifier_bounds_deduction.skel.h"
 #include "verifier_bounds_deduction_non_const.skel.h"
 #include "verifier_bounds_mix_sign_unsign.skel.h"
@@ -205,6 +206,7 @@ void test_verifier_arena_globals2(void)       { RUN(verifier_arena_globals2); }
 void test_verifier_basic_stack(void)          { RUN(verifier_basic_stack); }
 void test_verifier_bitfield_write(void)       { RUN(verifier_bitfield_write); }
 void test_verifier_bounds(void)               { RUN(verifier_bounds); }
+void test_verifier_bounds_step(void)          { RUN(verifier_bounds_step); }
 void test_verifier_bounds_deduction(void)     { RUN(verifier_bounds_deduction); }
 void test_verifier_bounds_deduction_non_const(void)     { RUN(verifier_bounds_deduction_non_const); }
 void test_verifier_bounds_mix_sign_unsign(void) { RUN(verifier_bounds_mix_sign_unsign); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
new file mode 100644
index 000000000000..079e34820a8f
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
@@ -0,0 +1,290 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+/* Tests for the linear "base + step * k" description tracked per scalar register. */
+
+#define __no_step __not_msg("step=") __msg("\n")
+
+/*
+ * Safe ALU64 arithmetic preserves a non-power-of-two stride and scales its base:
+ * step=0+3 -> step=1+3 -> step=2+6 -> step=4+12.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 *= 3 {{.*}}step=0+3)")
+__msg("r0 += 1 {{.*}}step=1+3)")
+__msg("r0 *= 2 {{.*}}step=2+6)")
+__msg("r0 <<= 1 {{.*}}step=4+12)")
+__naked void step_arith64_no_overflow(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 0xff;					\
+	r0 *= 3;					\
+	r0 += 1;					\
+	r0 *= 2;					\
+	r0 <<= 1;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Safe ALU32 multiplication and addition preserve the stride and scale its base:
+ * step=0+3 -> step=1+3 -> step=2+6. MOV32 preserves the result.
+ * LSH32 currently resets the stride, but must preserve the correct upper bound.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("w0 *= 3 {{.*}}step=0+3)")
+__msg("w0 += 1 {{.*}}step=1+3)")
+__msg("w0 *= 2 {{.*}}step=2+6)")
+__msg("w1 = w0 {{.*}}step=2+6)")
+__msg("w0 <<= 1 {{.*}}smax=umax=smax32=umax32=3064,") __no_step
+__naked void step_arith32_no_overflow(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 0xff;					\
+	w0 *= 3;					\
+	w0 += 1;					\
+	w0 *= 2;					\
+	w1 = w0;					\
+	w0 <<= 1;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Sign-crossing range with a non-power-of-2 step. After "*= 3; += -3" the value
+ * set is {-3, 0, 3, 6}. The line description is tracked in signed space, so the
+ * intersection keeps smax=6. A u64-modular intersection would mis-place the
+ * line for the negative values and wrongly narrow smax to 4.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += -3 {{.*}}smax=smax32=6)")
+__naked void step_neg_value_range(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r0 *= 3;					\
+	r0 += -3;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * 0x49249249 * 7 wraps to U32_MAX. Keeping step=0+7 incorrectly caps the
+ * result at 0xfffffffc instead of U32_MAX.
+ * Adding 4 to {0xfffffffc, U32_MAX} wraps to {0, 3}; keeping step=1+3
+ * incorrectly narrows the result to 1.
+ * Shifting the same input left by 1 gives {0xfffffff8, 0xfffffffe};
+ * keeping step=0+6 incorrectly narrows the result to 0xfffffffa.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("w0 *= 7 {{.*}}umax=0xffffffff,") __no_step
+__msg("r2 = r0 {{.*}}R0=scalar({{.*}}step=0+3) R2=scalar({{.*}}step=0+3)")
+__msg("w0 += 4 {{.*}}smax=umax=smax32=umax32=3,") __no_step
+__msg("w2 <<= 1 {{.*}}smax=umax=umax32=0xfffffffe,") __no_step
+__naked void step_arith32_overflow(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	w0 *= 7;					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 1;					\
+	r0 *= 3;					\
+	r1 = 0xfffffffc ll;				\
+	r0 += r1;					\
+	r2 = r0;					\
+	w0 += 4;					\
+	w2 <<= 1;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Multiplying {0x2aaaaaaaaaaaaaab, 0x2aaaaaaaaaaaaaac} by 3 crosses S64_MAX
+ * and gives {S64_MIN + 1, S64_MIN + 4}. Keeping step=0+3 excludes both.
+ *
+ * Adding 5 to {S64_MAX - 4, S64_MAX - 1} gives {S64_MIN, S64_MIN + 3}.
+ * Updating the line to step=2+3 incorrectly narrows it to S64_MIN + 1.
+ *
+ * Adding -5 to {S64_MIN + 1, S64_MIN + 4} wraps to
+ * {S64_MAX - 3, S64_MAX}. The stale base 0 modulo 3 excludes S64_MAX.
+ *
+ * Shifting {0x4000000000000008, 0x400000000000000b} left by 1 gives
+ * {S64_MIN + 16, S64_MIN + 22}. Keeping step=0+6 excludes both.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 *= 3 {{.*}}smax=0x8000000000000004,") __no_step
+__msg("r0 += r1 {{.*}}step=0+3)")
+__msg("r0 += 5 {{.*}}smax=0x8000000000000003,") __no_step
+__msg("r0 += r1 {{.*}}step=2+3)")
+__msg("r0 += -5 {{.*}}umax=0x7fffffffffffffff,") __no_step
+__msg("r0 += r1 {{.*}}step=0+3)")
+__msg("r0 <<= 1 {{.*}}smax=0x8000000000000016,") __no_step
+__naked void step_arith64_overflow(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 1;					\
+	r2 = r0;					\
+	r1 = 0x2aaaaaaaaaaaaaab ll;			\
+	r0 += r1;					\
+	r0 *= 3;					\
+	r2 *= 3;					\
+	r0 = r2;					\
+	r1 = 0x7ffffffffffffffb ll;			\
+	r0 += r1;					\
+	r0 += 5;					\
+	r0 = r2;					\
+	r1 = 0x8000000000000001 ll;			\
+	r0 += r1;					\
+	r0 += -5;					\
+	r0 = r2;					\
+	r1 = 0x4000000000000008 ll;			\
+	r0 += r1;					\
+	r0 <<= 1;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * MOV32 maps {2^32, 2^32 + 3} to {0, 3}. The old base 1 modulo 3
+ * excludes both results.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += r1 {{.*}}step=1+3)")
+__msg("w0 = w0 {{.*}}smax=umax=smax32=umax32=3,") __no_step
+__naked void step_mov32_truncate(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 1;					\
+	r0 *= 3;					\
+	r1 = 0x100000000 ll;				\
+	r0 += r1;					\
+	w0 = w0;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * MOV32 maps {-4, -1} to {0xfffffffc, U32_MAX}. The old base 2
+ * modulo 3 excludes U32_MAX.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += -4 {{.*}}step=2+3)")
+__msg("w0 = w0 {{.*}}umax=0xffffffff,") __no_step
+__naked void step_mov32_negative(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 1;					\
+	r0 *= 3;					\
+	r0 += -4;					\
+	w0 = w0;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Sign-extending {128, 131} from s8 gives {-128, -125}.
+ * Keeping base 2 modulo 3 incorrectly narrows the result to -127.
+ *
+ * Sign-extending {32768, 32771} from s16 gives {-32768, -32765}.
+ * Keeping base 2 modulo 3 incorrectly narrows the result to -32767.
+ *
+ * Sign-extending {2^31, 2^31 + 3} from s32 gives {S32_MIN, S32_MIN + 3}.
+ * Keeping base 2 modulo 3 incorrectly narrows the result to S32_MIN + 1.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += 128 {{.*}}step=2+3)")
+__msg("r0 = (s8)r0 {{.*}}smax=smax32=-125,") __no_step
+__msg("r0 += 32768 {{.*}}step=2+3)")
+__msg("r0 = (s16)r0 {{.*}}smax=smax32=-32765,") __no_step
+__msg("r0 += r1 {{.*}}step=2+3)")
+__msg("r0 = (s32)r0 {{.*}}smax=0xffffffff80000003,") __no_step
+__naked void step_movsx(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 1;					\
+	r0 *= 3;					\
+	r2 = r0;					\
+	r0 += 128;					\
+	r0 = (s8)r0;					\
+	r0 = r2;					\
+	r0 += 32768;					\
+	r0 = (s16)r0;					\
+	r0 = r2;					\
+	r1 = 0x80000000 ll;				\
+	r0 += r1;					\
+	r0 = (s32)r0;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * R1 = R0 + 1 shares R0's scalar ID. Refining R0 to {0, 3, 6} must
+ * translate its line as well as its bounds when updating R1 to {1, 4, 7}.
+ * Copying R0's base 0 modulo 3 excludes the reachable value 7.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r1 += 1 {{.*}}step=1+3)")
+__msg("if r0 > 0x6 {{.*}}R1=scalar({{.*}}smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=7,{{.*}}step=1+3)")
+__naked void step_linked_regs_add64(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r0 *= 3;					\
+	r1 = r0;					\
+	r1 += 1;					\
+	if r0 > 6 goto l_out_%=;			\
+	r0 = r1;					\
+l_out_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (37 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:54   ` sashiko-bot
  2026-10-04 13:38 ` [PATCH bpf-next v2 40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
                   ` (5 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Tests for the base/step reasoning in states.c:range_within():
- step_prune_hit_multiple: prune a step-4 range contained in a cached
  step-2 range when the scalar is used to address the stack.
- step_prune_miss_non_multiple: do not let a cached step-2 range hide a
  step-3 path reaching division by zero.
- step_prune_hit_const_on_line: prune constant 6 against a cached
  base-0, step-3 range.
- step_prune_miss_const_off_line: do not prune constant 7 against that
  range even though its bounds and tnum admit the value.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_bounds_step.c     | 183 +++++++++++++++++++++
 1 file changed, 183 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
index 079e34820a8f..f1d8e36b39e3 100644
--- a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
+++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
@@ -287,4 +287,187 @@ l_out_%=:						\
 	: __clobber_all);
 }
 
+struct step_val {
+	__u8 data[1024];
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARRAY);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, struct step_val);
+} step_map SEC(".maps");
+
+/* Old register [4..130, step 2] should prune cur register [8..68, step 4]. */
+SEC("socket")
+__success __log_level(2)
+__msg("7: (27) r6 *= 4                       ; R6=scalar({{.*}}umin32=8,{{.*}}umax32=68,{{.*}},step=0+4)")
+__msg("10: (27) r7 *= 2                      ; R7=scalar({{.*}}umin32=4,{{.*}}umax32=130,{{.*}},step=0+2)")
+__msg("11: (25) if r0 > 0x2a goto pc+1")
+__msg("from 11 to 13: safe")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void step_prune_hit_multiple(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r6 &= 0x0f;					\
+	r6 += 2;					\
+	r6 *= 4;					\
+	r7 &= 0x3f;					\
+	r7 += 2;					\
+	r7 *= 2;					\
+	if r0 > 42 goto 1f;	/* can't predict */	\
+	r6 = r7;		/* step=2 explored first, step=4 explored next */ \
+1:	r0 = r10;					\
+	r6 = -r6;					\
+	r0 += r6;					\
+	*(u8 *)(r0 + 0) = 7;	/* force r6 precise */	\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
+/*
+ * Old register [0..189, step 3] should not prune cur register [0..30, step 2].
+ * Bounds and tnum containment allow pruning; step divisibility must prevent it.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("6: (27) r6 *= 2                       ; R6=scalar({{.*}}smin32=0,{{.*}}umax32=30,{{.*}},step=0+2)")
+__msg("8: (27) r7 *= 3                       ; R7=scalar({{.*}}smin32=0,{{.*}}umax32=189,{{.*}},step=0+3)")
+__msg("9: (25) if r0 > 0x2a goto pc+1")
+__msg("17: (95) exit")
+__msg("from 9 to 11: R6=scalar({{.*}},step=0+2)")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void step_prune_miss_non_multiple(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r6 &= 0x0f;					\
+	r6 *= 2;					\
+	r7 &= 0x3f;					\
+	r7 *= 3;					\
+	if r0 > 42 goto 1f;	/* can't predict */	\
+	r6 = r7;		/* step=3 explored first, step=2 explored next */ \
+1:							\
+	r2 = r10;					\
+	r6 += 1;					\
+	r6 = -r6;					\
+	r2 += r6;					\
+	*(u8 *)(r2 + 0) = 7;	/* force r6 precise */	\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Constant current register lying on the cached line: cached is {0,3,6,...}
+ * (step 3, base 0), current is the constant 6. range_within() takes the
+ * constant branch: imod(6, 3) == base 0, so cur is on the line and the
+ * (precise) state is pruned -> "safe".
+ */
+SEC("socket")
+__success __log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("14: (27) r1 *= 3")		/* cached path: step 3 line */
+__msg("16: (b7) r1 = 6")		/* current path: const 6, on the line */
+__msg("17: safe")			/* pruned at join: imod(6, 3) == 0 */
+__naked void step_prune_hit_const_on_line(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	r1 = 0;						\
+	*(u32*)(r10 - 4) = r1;				\
+	r2 = r10;					\
+	r2 += -4;					\
+	r1 = %[step_map] ll;				\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto l_out_%=;			\
+	r7 = r0;					\
+	r1 = r6;					\
+	r1 &= 0xff;					\
+	if r6 > 0 goto l_cur_%=;			\
+	r1 *= 3;			/* old: step 3 */	\
+	goto l_join_%=;					\
+l_cur_%=:						\
+	r1 = 6;				/* cur: const on line */	\
+l_join_%=:						\
+	r0 = r7;					\
+	r0 += r1;			/* r1 forced precise */	\
+	r2 = *(u8 *)(r0 + 0);				\
+l_out_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
+/*
+ * Constant current register NOT on the cached line: cached is {0,3,6,...}
+ * (step 3, base 0), current is the constant 7. imod(7, 3) == 1 != base 0,
+ * so range_within() fails and the join is traversed again. A power-of-two
+ * step is avoided on purpose: with step 3 the tnum is loose enough to admit
+ * 7, so imod() is
+ * the sole check that rejects it. No failure shape is possible here: 7 is
+ * within the line's bounds and tnum, and the cached line path already
+ * verifies the whole outro, so an eager prune could not miss an error.
+ */
+SEC("socket")
+__success __log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("14: (27) r1 *= 3")		/* cached path: step 3 line */
+__msg("16: (b7) r1 = 7")		/* current path: const 7, off the line */
+/* not pruned: current continues past the join with the constant offset 7 */
+__msg("19: R0=map_value({{.*}}imm=7) R1=7")
+__naked void step_prune_miss_const_off_line(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	r1 = 0;						\
+	*(u32*)(r10 - 4) = r1;				\
+	r2 = r10;					\
+	r2 += -4;					\
+	r1 = %[step_map] ll;				\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto l_out_%=;			\
+	r7 = r0;					\
+	r1 = r6;					\
+	r1 &= 0xff;					\
+	if r6 > 0 goto l_cur_%=;			\
+	r1 *= 3;			/* old: step 3 */	\
+	goto l_join_%=;					\
+l_cur_%=:						\
+	r1 = 7;				/* cur: const off line */	\
+l_join_%=:						\
+	r0 = r7;					\
+	r0 += r1;			/* r1 forced precise */	\
+	r2 = *(u8 *)(r0 + 0);				\
+l_out_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
 char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (38 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 13:38 ` [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
                   ` (4 subsequent siblings)
  44 siblings, 0 replies; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Verifier tests for accessing arrays at a varying offset through a
PTR_TO_BTF_ID pointer, e.g. a[i].b where 'i' is a register with a
varying value:
- bpf_obj_new() is used to obtain a PTR_TO_BTF_ID to a program-defined
  struct
- bpf_get_current_task_btf() is used to obtain a PTR_TO_BTF_ID from
  the kernel BTF.

The cases cover:
- in-bounds access to a field of an array of structs;
- out-of-bounds access where the maximal offset runs past the array;
- a scalar (long) array with a matching stride;
- a byte array (step 1);
- a 4-byte read whose maximal offset spills past the array end;
- an unaligned step that is not a whole number of elements;
- both dimensions of a 2D array;
- an array of pointers within a program-allocated object, where reading a
  pointer member yields a SCALAR_VALUE but the varying offset is still
  validated against the array bounds (in-bounds and out-of-bounds);
- a varying access into an array embedded in a kernel BTF type
  (task_struct's comm[]), covering the non-allocated, trusted-pointer path
  (in-bounds and out-of-bounds);
- an array nested below a struct member at a non-zero offset, exercising
  the offset telescoping done while walking across a struct boundary
  (in-bounds and out-of-bounds);
- a trailing flexible array member, whose unbounded tail makes a varying
  access have no upper bound to exceed. A program-defined struct cannot be
  used here because bpf_obj_new() forbids flexible array members, so a
  kernel type (struct vring_used, with its 'ring[]' tail) is used
  (in-bounds, and an unaligned step that is not a whole number of
   elements).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../bpf/progs/verifier_btf_array_access.c          | 576 +++++++++++++++++++++
 2 files changed, 578 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 4f3af1cb998b..2f05a0ce829f 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -27,6 +27,7 @@
 #include "verifier_bpf_get_stack.skel.h"
 #include "verifier_bpf_trap.skel.h"
 #include "verifier_bswap.skel.h"
+#include "verifier_btf_array_access.skel.h"
 #include "verifier_btf_ctx_access.skel.h"
 #include "verifier_btf_flex_array.skel.h"
 #include "verifier_btf_unreliable_prog.skel.h"
@@ -213,6 +214,7 @@ void test_verifier_bounds_mix_sign_unsign(void) { RUN(verifier_bounds_mix_sign_u
 void test_verifier_bpf_get_stack(void)        { RUN(verifier_bpf_get_stack); }
 void test_verifier_bpf_trap(void)             { RUN(verifier_bpf_trap); }
 void test_verifier_bswap(void)                { RUN(verifier_bswap); }
+void test_verifier_btf_array_access(void)     { RUN(verifier_btf_array_access); }
 void test_verifier_btf_ctx_access(void)       { RUN(verifier_btf_ctx_access); }
 void test_verifier_btf_flex_array(void)       { RUN(verifier_btf_flex_array); }
 void test_verifier_btf_unreliable_prog(void)  { RUN(verifier_btf_unreliable_prog); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
new file mode 100644
index 000000000000..be001cc77f81
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
@@ -0,0 +1,576 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_core_read.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+
+/*
+ * Variable-offset access into arrays within a BTF access chain, e.g. a[i].b,
+ * where 'i' is a register with a varying offset. The register's step has to
+ * match the stride of one of the arrays crossed while walking the access
+ * chain, and the maximal possible offset has to stay within that array.
+ *
+ * A program-allocated object (bpf_obj_new) is used as the source of a
+ * PTR_TO_BTF_ID with a shape we fully control.
+ */
+struct inner {
+	int a;
+	int b;
+};
+
+struct outer {
+	struct inner arr[8];	/* off 0,   stride 8, size 64 */
+	char bytes[32];		/* off 64,  stride 1, size 32 */
+	long longs[8];		/* off 96,  stride 8, size 64 */
+	int grid[4][4];		/* off 160, 16 ints,  size 64 */
+};
+
+/* arr[i].b, i in [0, 7], 4-byte read: stays within arr. */
+SEC("syscall")
+__success
+int arr_field_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* arr[i].b, i in [0, 15]: max offset runs past the end of arr. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int arr_field_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* longs[i], i in [0, 7], 8-byte read: scalar array with stride 8. */
+SEC("syscall")
+__success
+int long_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 96);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * bytes[i], i in [0, 31], read as 4 bytes: min_off is valid within the array,
+ * but the 4-byte read at max_off runs past the array end. Exercises the
+ * size-aware bound in btf_struct_access().
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int byte_array_size_spanning(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 64);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* bytes[i], i in [0, 31], 1-byte read: step 1 matches the char array stride. */
+SEC("syscall")
+__success
+int byte_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 64);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * arr[] accessed with step 4 (half of sizeof(struct inner)): the step is not
+ * a whole number of elements, so no crossed array has a matching stride.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int arr_unaligned_step(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * grid[i][0], i in [0, 3]: steps along the outer dimension of a 2D array
+ * (stride 16). __btf_resolve_size() linearizes the array; the step is a
+ * multiple of the innermost element size (4).
+ */
+SEC("syscall")
+__success
+int grid_outer_dim(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 3;						\
+	r2 *= 16;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 160);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * grid[0][i], i in [0, 3]: steps along the inner dimension of a 2D array
+ * (stride 4, the innermost element size).
+ */
+SEC("syscall")
+__success
+int grid_inner_dim(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 3;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 160);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * An array of pointers within a program-allocated object. Reading a pointer
+ * member of a local object yields a SCALAR_VALUE, but the varying offset still
+ * has to be validated against the array bounds.
+ */
+struct with_ptrs {
+	struct inner *parr[8];	/* off 0, stride 8 (sizeof ptr), size 64 */
+};
+
+/* parr[i], i in [0, 7], 8-byte read: stays within parr. */
+SEC("syscall")
+__success
+int ptr_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct with_ptrs *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * parr[i], i in [0, 15]: max offset runs past the end of parr. Without the
+ * variable-offset check on the WALK_PTR path this would be wrongly accepted.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct with_ptrs")
+int ptr_array_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct with_ptrs *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * The cases above all use bpf_obj_new() (PTR_TO_BTF_ID | MEM_ALLOC). Exercise
+ * the trusted kernel-BTF path too, using a task_struct from
+ * bpf_get_current_task_btf() and its embedded char comm[] array. The comm
+ * offset is folded into the pointer as a constant before the varying index.
+ */
+
+/* task->comm[i], i in [0, 15], 1-byte read: stays within comm. */
+SEC("syscall")
+__success
+int kernel_btf_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct task_struct *task;
+	int val = 0;
+
+	task = bpf_get_current_task_btf();
+	asm volatile ("						\
+	r1 = %[task];						\
+	r1 += %[off];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [task] "r"(task), [i] "r"(i),
+	  [off] "i"(offsetof(struct task_struct, comm))
+	: "r1", "r2");
+	return val;
+}
+
+/* task->comm[i], i in [0, 31]: max offset runs past comm. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct task_struct")
+int kernel_btf_array_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct task_struct *task;
+	int val = 0;
+
+	task = bpf_get_current_task_btf();
+	asm volatile ("						\
+	r1 = %[task];						\
+	r1 += %[off];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [task] "r"(task), [i] "r"(i),
+	  [off] "i"(offsetof(struct task_struct, comm))
+	: "r1", "r2");
+	return val;
+}
+
+/*
+ * Array nested below a struct member at a non-zero offset. The walk descends
+ * through 'm' (resetting the running offset) before reaching 'arr', exercising
+ * the array_start = min_off - arrays[i].off telescoping across a WALK_STRUCT
+ * dive.
+ */
+struct mid {
+	struct inner arr[8];	/* stride 8, size 64 */
+};
+
+struct nest {
+	long pad;		/* off 0 */
+	struct mid m;		/* off 8 */
+};
+
+/* m.arr[i].b, i in [0, 7], 4-byte read: stays within m.arr. */
+SEC("syscall")
+__success
+int nested_struct_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct nest *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 12);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* m.arr[i].b, i in [0, 15]: max offset runs past the end of m.arr. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct nest")
+int nested_struct_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct nest *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 12);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * A trailing flexible array member makes the struct tail an unbounded region,
+ * so a varying access into it has no upper bound to exceed (mirrors how
+ * unix_address.name[] is accessed via sun_path[i]).
+ *
+ * A custom (program) BTF struct can only be reached as a PTR_TO_BTF_ID through
+ * bpf_obj_new(), which does not allow flexible array members. So use a kernel
+ * type via an untrusted PTR_TO_BTF_ID (bpf_core_cast()). struct vring_used ends
+ * with a flexible array 'ring[]' of struct vring_used_elem { __virtio32 id, len; },
+ * mirroring a 'struct foo { int a; int b; }' flexible array.
+ */
+
+/* ring[i].len, i in [0, 63], 4-byte read: the flexible array has no upper bound. */
+SEC("syscall")
+__success
+int flex_array_field(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct vring_used *o;
+	int val = 0;
+
+	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 63;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 8);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	return val;
+}
+
+/*
+ * ring[] accessed with step 4 (half of sizeof(struct vring_used_elem)): the
+ * step is not a whole number of elements, so the flexible array stride does
+ * not match.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct vring_used")
+int flex_array_unaligned_step(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct vring_used *o;
+	int val = 0;
+
+	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	return val;
+}
+
+/*
+ * Access either name[0].sun_path[0] or one byte past sun_path.
+ * Step 108 does not match sizeof(struct sockaddr_un), which is 110.
+ * It matches the inner char array's stride, but exceeds its bounds.
+ * The flexible-tail exemption must not apply to that fixed array.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct unix_address")
+int flex_array_inner_array_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct unix_address *o;
+	int val = 0;
+
+	o = bpf_core_cast(bpf_get_current_task_btf(), struct unix_address);
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 1;						\
+	r2 *= %[len];						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + %[off]);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i),
+	  __imm_const(len, sizeof_field(struct sockaddr_un, sun_path)),
+	  __imm_const(off, offsetof(struct unix_address, name) +
+			   offsetof(struct sockaddr_un, sun_path))
+	: "r1", "r2");
+	return val;
+}
+
+typedef int array_element_t;
+typedef array_element_t array_type_t[4];
+
+struct typedef_array {
+	array_type_t values;
+};
+
+/*
+ * Resolving values' size crosses both the array and element typedefs.
+ * Record the array ID, not the int ID left by element-type resolution.
+ * values[i], i in [0, 3], must pass the variable-offset array check.
+ */
+SEC("syscall")
+__success
+int typedef_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct typedef_array *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 3;						\
+	r2 *= %[elem_size];					\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i),
+	  __imm_const(elem_size, sizeof(array_element_t))
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (39 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
                   ` (3 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman,
	Emil Tsalapatis

Test cases covering the following branches in bpf_compute_loops():
- Case B: simple backedge creating a loop (loop_single)
- Case B: two independent loops (loop_two_independent)
- Case B + D: nested loops where the inner header's loop_header points
  to the outer header (loop_nested)
- Case C: diamond CFG with no loops (fwd_edges_no_loop)
- Case D: sibling inner loops within one outer loop
  (loop_nested_siblings)
- Three levels of loop nesting (loop_three_levels)
- Loop with an if-else body containing forward branches
  (loop_with_if_else)
- Case E: An irreducible loop (loop_irreducible)
- A self-loop (loop_self)
- A test with sibling loops nested in outer loops
  (loop_nested_siblings_common_ancestors)
- A test case with outer loop's backedge originating
  from an inner loop (loop_outer_backedge_from_inner).

Cases B, C, D and E are described in loops.c:compute_loops_in_subprog().

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

dominators test case
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../selftests/bpf/progs/verifier_loop_hierarchy.c  | 346 +++++++++++++++++++++
 2 files changed, 348 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 2f05a0ce829f..6b60d7720e2a 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -74,6 +74,7 @@
 #include "verifier_live_stack.skel.h"
 #include "verifier_liveness_exp.skel.h"
 #include "verifier_load_acquire.skel.h"
+#include "verifier_loop_hierarchy.skel.h"
 #include "verifier_loops1.skel.h"
 #include "verifier_lwt.skel.h"
 #include "verifier_map_in_map.skel.h"
@@ -261,6 +262,7 @@ void test_verifier_leak_ptr(void)             { RUN(verifier_leak_ptr); }
 void test_verifier_linked_scalars(void)       { RUN(verifier_linked_scalars); }
 void test_verifier_live_stack(void)           { RUN(verifier_live_stack); }
 void test_verifier_liveness_exp(void)         { RUN(verifier_liveness_exp); }
+void test_verifier_loop_hierarchy(void)       { RUN(verifier_loop_hierarchy); }
 void test_verifier_loops1(void)               { RUN(verifier_loops1); }
 void test_verifier_lwt(void)                  { RUN(verifier_lwt); }
 void test_verifier_map_in_map(void)           { RUN(verifier_map_in_map); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
new file mode 100644
index 000000000000..0f49108c0e50
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
@@ -0,0 +1,346 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+/*
+ * kernel/bpf/loops.c:compute_loops() distinguish between
+ * the following cases:
+ * - B: backedge -> simple loop
+ * - C: cross edge to non-loop node -> no-op
+ * - D: edge to node whose header is in DFS path -> nested loop
+ * - E: edge to node whose header is NOT in DFS path -> irreducible
+ *
+ * Below test cases cover the above branches in various combinations.
+ */
+
+/* Case B: single bounded loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("          2   3: {{.*}} (95) exit")
+__naked void loop_single(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case B: two independent loops at the same nesting level. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 4{{$}}")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  2       0   1: {{.*}} (07) r0 += 1")
+__msg("  2   1   1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("          2   3: {{.*}} (b7) r1 = 0")
+__msg("  1       3   4: {{.*}} (07) r1 += 1")
+__msg("  1   4   4   5: {{.*}} (a5) if r1 < 0xa goto pc-2")
+__msg("          5   6: {{.*}} (95) exit")
+__naked void loop_two_independent(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 10 goto 2b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case B + D: nested loops. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")			/* outer loop header */
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")			/* outer loop insn */
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")			/* inner loop header */
+__msg("  1   3   3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")	/* inner loop insn */
+__msg("  1   1   4   5: {{.*}} (a5) if r0 < 0xa goto pc-5")	/* outer loop insn */
+__msg("          5   6: {{.*}} (95) exit")
+__naked void loop_nested(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case C: forward edges, no loops. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (85) call bpf_get_prandom_u32")
+__msg("          0   1: {{.*}} (25) if r0 > 0x0 goto pc+2")
+__msg("          1   2: {{.*}} (b7) r0 = 2")
+__msg("          2   3: {{.*}} (05) goto pc+1")
+__msg("          1   4: {{.*}} (b7) r0 = 3")
+__msg("          1   5: {{.*}} (95) exit")
+__naked void fwd_edges_no_loop(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r0 = 2;						\
+	goto 2f;					\
+1:	r0 = 3;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/* Case B + D: two sibling inner loops within one outer loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("loop at 6, nested in 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")
+__msg("  1   3   3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")
+__msg("  1   1   4   5: {{.*}} (b7) r2 = 0")
+__msg("  1   1   5   6: {{.*}} (07) r2 += 1")
+__msg("  1   6   6   7: {{.*}} (a5) if r2 < 0x5 goto pc-2")
+__msg("  1   1   7   8: {{.*}} (a5) if r0 < 0xa goto pc-8")
+__msg("          8   9: {{.*}} (95) exit")
+__naked void loop_nested_siblings(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 5 goto 3b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Three levels of nesting. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("loop at 5, nested in 3")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")
+__msg("  1   3   3   4: {{.*}} (b7) r2 = 0")
+__msg("  1   3   4   5: {{.*}} (07) r2 += 1")
+__msg("  1   5   5   6: {{.*}} (a5) if r2 < 0x3 goto pc-2")
+__msg("  1   3   6   7: {{.*}} (a5) if r1 < 0x5 goto pc-5")
+__msg("  1   1   7   8: {{.*}} (a5) if r0 < 0xa goto pc-8")
+__msg("          8   9: {{.*}} (95) exit")
+__naked void loop_three_levels(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 3 goto 3b;				\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Loop with an if-else body (forward branch inside loop, Case C). */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (bf) r1 = r0")
+__msg("  1   1   2   3: {{.*}} (25) if r1 > 0x5 goto pc+1")
+__msg("  1   1   3   4: {{.*}} (b7) r1 = 1")
+__msg("  1   1   3   5: {{.*}} (0f) r0 += r1")
+__msg("  1   1   5   6: {{.*}} (a5) if r0 < 0x64 goto pc-6")
+__msg("          6   7: {{.*}} (95) exit")
+__naked void loop_with_if_else(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = r0;					\
+	if r1 > 5 goto 2f;				\
+	r1 = 1;						\
+2:	r0 += r1;					\
+	if r0 < 100 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case E: irreducible loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 3, irreducible")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("          0   1: {{.*}} (b7) r1 = 0")
+__msg("          1   2: {{.*}} (25) if r0 > 0x5 goto pc+2")
+__msg("  1       2   3: {{.*}} (b7) r1 = 1")
+__msg("  1   3   3   4: {{.*}} (05) goto pc+1")
+__msg("          2   5: {{.*}} (b7) r1 = 2")
+__msg("  1   3   2   6: {{.*}} (0f) r0 += r1")
+__msg("  1   3   6   7: {{.*}} (a5) if r0 < 0x10 goto pc-5")
+__msg("          7   8: {{.*}} (95) exit")
+__naked void loop_irreducible(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+	if r0 > 5 goto 2f;				\
+1:	r1 = 1;						\
+	goto 3f;					\
+2:	r1 = 2;						\
+3:	r0 += r1;					\
+	if r0 < 16 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("socket")
+__failure
+__log_level(2)
+__msg("loop at 1")
+__msg("infinite loop detected at insn 1")
+__naked void loop_self(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+1:	if r0 < 10 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Loops at 5 and 7 are nested inside loops at 3 and 1.
+ * The edge 6 -> 7 exits loop at 5 and enters loop at 7.
+ * Check that it is logged only as an exit from loop at 5.
+ *
+ *  0: r0 = 0;
+ *     do {
+ *  1:     r0++;
+ *  2:     r1 = 0;
+ *         do {
+ *  3:         r1++;
+ *  4:         r2 = 0;
+ *             do {
+ *  5:             r2++;
+ *  6:         } while (r2 < 3);
+ *             do {
+ *  7:             r2--;
+ *  8:         } while (r2 != 0);
+ *  9:     } while (r1 < 4);
+ * 10: } while (r0 < 5);
+ * 11: return r0;
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg_next("  backedge from 10, latch at 10")
+__msg_next("  exit from 10 to 11")
+__msg_next("loop at 3, nested in 1")
+__msg_next("  backedge from 9, latch at 9")
+__msg_next("  exit from 9 to 10")
+__msg_next("loop at 5, nested in 3")
+__msg_next("  backedge from 6, latch at 6")
+__msg_next("  exit from 6 to 7")
+__msg_next("loop at 7, nested in 3")
+__msg_next("  backedge from 8, latch at 8")
+__msg_next("  exit from 8 to 9")
+__naked void loop_nested_siblings_common_ancestors(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 3 goto 3b;				\
+4:	r2 += -1;					\
+	if r2 != 0 goto 4b;				\
+	if r1 < 4 goto 2b;				\
+	if r0 < 5 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The outer loop backedge (4 -> 2) originates in an inner loop.
+ * For now, latch should not be identified in such cases.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 2{{$}}")
+__msg_next("  backedge from 4, latch at -1")
+__msg_next("  exit from 5 to 7")
+__msg_next("loop at 3, nested in 2")
+__msg_next("  backedge from 6, latch at 5")
+__msg_next("  exit from 4 to 2")
+__msg_next("  exit from 5 to 7")
+__naked void loop_outer_backedge_from_inner(void)
+{
+	asm volatile ("					\
+	r6 = 99;					\
+	r2 = 2;						\
+2:	r6 += 1;			/* loop at 2 */	\
+3:	r2 += 1;			/* loop at 3 */	\
+4:	if r2 == 3 goto 2b;				\
+5:	if r6 > 100 goto 7f;				\
+	goto 3b;					\
+7:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (40 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 13:38 ` [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening Eduard Zingerman
                   ` (2 subsequent siblings)
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

Coverage:
- straight-line code with an ldimm64 in the middle (exercises the
  bpf_is_ldimm64() skip in compute_predecessors());
- an asymmetric if-then-else diamond (unequal-depth idoms_intersect());
- a simple loop, a loop with an if-else body, and nested loops;
- a loop header with two back edges (three predecessors);
- an irreducible CFG (fixpoint convergence);
- immediate dominators across subprogram boundaries, including a callee
  whose entry instruction is itself a loop header;
- a self-loop;
- an indirect jump (gotox) with a two-entry jump table.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 tools/testing/selftests/bpf/progs/verifier_idoms.c | 390 +++++++++++++++++++++
 2 files changed, 392 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 6b60d7720e2a..b45916e04fef 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -58,6 +58,7 @@
 #include "verifier_helper_packet_access.skel.h"
 #include "verifier_helper_restricted.skel.h"
 #include "verifier_helper_value_access.skel.h"
+#include "verifier_idoms.skel.h"
 #include "verifier_int_ptr.skel.h"
 #include "verifier_iterating_callbacks.skel.h"
 #include "verifier_jeq_infer_not_null.skel.h"
@@ -246,6 +247,7 @@ void test_verifier_helper_access_var_len(void) { RUN(verifier_helper_access_var_
 void test_verifier_helper_packet_access(void) { RUN(verifier_helper_packet_access); }
 void test_verifier_helper_restricted(void)    { RUN(verifier_helper_restricted); }
 void test_verifier_helper_value_access(void)  { RUN(verifier_helper_value_access); }
+void test_verifier_idoms(void)                { RUN(verifier_idoms); }
 void test_verifier_int_ptr(void)              { RUN(verifier_int_ptr); }
 void test_verifier_iterating_callbacks(void)  { RUN(verifier_iterating_callbacks); }
 void test_verifier_jeq_infer_not_null(void)   { RUN(verifier_jeq_infer_not_null); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_idoms.c b/tools/testing/selftests/bpf/progs/verifier_idoms.c
new file mode 100644
index 000000000000..a29d6332baa4
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_idoms.c
@@ -0,0 +1,390 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+/*
+ * kernel/bpf/loops.c:bpf_compute_idoms() computes the immediate dominator
+ * of every instruction (Cooper et al, "A Simple, Fast Dominance Algorithm").
+ *
+ * The immediate dominator is printed by log_program() as the numeric column
+ * immediately before the "<insn#>:" field of every "Program dump" line:
+ *
+ *   Program dump (scc? loop_header? idom insn#: live_regs_before):
+ *            -1   0: ....... (b7) r0 = 0     <- idom(0) = -1 (subprog entry)
+ *             0   1: ....... (07) r0 += 1    <- idom(1) = 0
+ *             ^^^^^^
+ *             idom insn#
+ *
+ * The __msg() patterns below wildcard the scc/loop_header/live_regs columns and
+ * anchor on "<idom>   <insn#>:" followed by the disassembled instruction, so
+ * they assert the idom value of each instruction.
+ */
+
+/*
+ * Straight-line code, with an ldimm64 in the middle. Instruction index 2 is
+ * the second half of the ldimm64 at index 1 and is not a CFG node; the idom of
+ * index 3 must be 1, not 2. Exercises the bpf_is_ldimm64() skip in both passes
+ * of compute_predecessors().
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (18) r1 = 0x1122334455667788")
+__msg("{{.*}}  1   3: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  3   4: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  4   5: {{.*}} (95) exit")
+__naked void straight_line_ldimm64(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r1 = 0x1122334455667788 ll;			\
+	r0 += 1;					\
+	r0 += r1;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Asymmetric if-then-else diamond: the "then" arm is 5 instructions long, the
+ * "else" arm is a single instruction. The merge point (insn 8) is dominated by
+ * the branch (insn 1), not by either arm. This forces idoms_intersect() to walk
+ * the two predecessors up unequal postorder depths before they meet.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (25) if r0 > 0x0 goto pc+5")
+__msg("{{.*}}  1   2: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  2   3: {{.*}} (b7) r1 = 2")
+__msg("{{.*}}  3   4: {{.*}} (b7) r1 = 3")
+__msg("{{.*}}  4   5: {{.*}} (b7) r1 = 4")
+__msg("{{.*}}  5   6: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  1   7: {{.*}} (b7) r1 = 9")
+__msg("{{.*}}  1   8: {{.*}} (95) exit")
+__naked void asymmetric_diamond(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r1 = 1;						\
+	r1 = 2;						\
+	r1 = 3;						\
+	r1 = 4;						\
+	goto 2f;					\
+1:	r1 = 9;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Simple loop: the header (insn 1) is dominated by the pre-header (insn 0). The
+ * back edge (insn 2 -> insn 1) must not change idom(1); the back-edge
+ * predecessor folds up via idoms_intersect().
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("{{.*}}  2   3: {{.*}} (95) exit")
+__naked void simple_loop(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Loop with an if-else in its body. The in-loop merge point (insn 5) is
+ * dominated by the in-loop branch (insn 3), combining a back-edge intersect
+ * with a forward-diamond intersect.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (bf) r1 = r0")
+__msg("{{.*}}  2   3: {{.*}} (25) if r1 > 0x5 goto pc+1")
+__msg("{{.*}}  3   4: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  3   5: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  5   6: {{.*}} (a5) if r0 < 0x64 goto pc-6")
+__msg("{{.*}}  6   7: {{.*}} (95) exit")
+__naked void loop_with_if_else_body(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = r0;					\
+	if r1 > 5 goto 2f;				\
+	r1 = 1;						\
+2:	r0 += r1;					\
+	if r0 < 100 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Nested loops: the idom chain runs inner-header -> outer-body -> outer-header
+ * -> pre-header. Exercises intersect across back edges at two nesting depths.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  2   3: {{.*}} (07) r1 += 1")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")
+__msg("{{.*}}  4   5: {{.*}} (a5) if r0 < 0xa goto pc-5")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__naked void nested_loops(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Loop header (insn 1) with two back edges (from insn 4 and insn 5), i.e. three
+ * predecessors. Both back edges route through the incrementing header, so the
+ * loop is bounded and verifies. Repeated idoms_intersect() at the header must
+ * stay stable and keep idom(1) at the pre-header (insn 0).
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r6 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r6 += 1")
+__msg("{{.*}}  1   2: {{.*}} (25) if r6 > 0xa goto pc+3")
+__msg("{{.*}}  2   3: {{.*}} (bf) r7 = r6")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r7 < 0x5 goto pc-4")
+__msg("{{.*}}  4   5: {{.*}} (05) goto pc-5")
+__msg("{{.*}}  2   6: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  6   7: {{.*}} (95) exit")
+__naked void multi_backedge_header(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	r6 += 1;					\
+	if r6 > 10 goto 2f;				\
+	r7 = r6;					\
+	if r7 < 5 goto 1b;				\
+	goto 1b;					\
+2:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Irreducible CFG (loop with two entries, insn 3 and insn 5, reached from the
+ * insn 2 branch). Dominators remain well-defined; this is a convergence test
+ * for the fixpoint under irreducibility. Note insn 6 is dominated by the branch
+ * (insn 2), not by either arm.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  1   2: {{.*}} (25) if r0 > 0x5 goto pc+2")
+__msg("{{.*}}  2   3: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  3   4: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  2   5: {{.*}} (b7) r1 = 2")
+__msg("{{.*}}  2   6: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  6   7: {{.*}} (a5) if r0 < 0x10 goto pc-5")
+__msg("{{.*}}  7   8: {{.*}} (95) exit")
+__naked void irreducible(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+	if r0 > 5 goto 2f;				\
+1:	r1 = 1;						\
+	goto 3f;					\
+2:	r1 = 2;						\
+3:	r0 += r1;					\
+	if r0 < 16 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Idoms are computed per subprog. Layout (libbpf preorder) is:
+ *   main: 0..1, sub: 2..7
+ * The callee entry (insn 2) must have idom -1 (subprog reset), and idoms inside
+ * the callee must reference only the callee's instructions, never main's. The
+ * in-callee merge (insn 7) is dominated by the in-callee branch (insn 3). The
+ * branch is on a prandom value so it is not constant-folded away.
+ */
+static __naked __noinline __used
+unsigned long idoms_diamond_sub(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r0 = 1;						\
+	goto 2f;					\
+1:	r0 = 2;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call pc+1")
+__msg("{{.*}}  0   1: {{.*}} (95) exit")
+__msg("{{.*}} -1   2: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  2   3: {{.*}} (25) if r0 > 0x0 goto pc+2")
+__msg("{{.*}}  3   4: {{.*}} (b7) r0 = 1")
+__msg("{{.*}}  4   5: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  3   6: {{.*}} (b7) r0 = 2")
+__msg("{{.*}}  3   7: {{.*}} (95) exit")
+__naked void multi_subprog(void)
+{
+	asm volatile ("					\
+	call idoms_diamond_sub;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The first instruction of a subprog (insn 3) is itself a loop header, i.e. it
+ * has an incoming back edge. Its idom must still be -1 (the subprog entry has no
+ * dominator). Exercises the idoms[start]=0 ... idoms[start]=-1 handling in
+ * compute_subprog_idoms().
+ */
+static __naked __noinline __used
+unsigned long idoms_entry_header_sub(void)
+{
+	asm volatile ("					\
+1:	r1 += 1;					\
+	if r1 < 10 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  0   1: {{.*}} (85) call pc+1")
+__msg("{{.*}}  1   2: {{.*}} (95) exit")
+__msg("{{.*}} -1   3: {{.*}} (07) r1 += 1")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r1 < 0xa goto pc-2")
+__msg("{{.*}}  4   5: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__naked void entry_is_loop_header(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	call idoms_entry_header_sub;			\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Self-loop: insn 1 is its own predecessor. idom(1) must be the pre-header
+ * (insn 0); the self-edge is ignored (idoms[pred] == -1 on the first pass, then
+ * idoms_intersect(a == b) short-circuits). Program is rejected later for an
+ * infinite loop, but the dump (and idoms) are printed before that.
+ */
+SEC("socket")
+__failure
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (a5) if r0 < 0xa goto pc-1")
+__msg("{{.*}}  1   2: {{.*}} (95) exit")
+__msg("infinite loop detected")
+__naked void self_loop(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+1:	if r0 < 10 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64) || defined(__TARGET_ARCH_powerpc)
+/*
+ * Indirect jump (gotox) with a two-entry jump table. The gotox (insn 4) has two
+ * successors (insn 5 and insn 7), so both targets have the gotox as their only
+ * predecessor and idom. Also re-exercises the ldimm64 skip: insn 2's idom is 0,
+ * the ldimm64 at index 0 (index 1 is its second half).
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (18) r0 = {{0x[0-9a-f]+}}")
+__msg("{{.*}}  0   2: {{.*}} (07) r0 += 8")
+__msg("{{.*}}  2   3: {{.*}} (79) r0 = *(u64 *)(r0 +0)")
+__msg("{{.*}}  3   4: {{.*}} (0d) gotox r0")
+__msg("{{.*}}  4   5: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__msg("{{.*}}  4   7: {{.*}} (b7) r0 = 1")
+__msg("{{.*}}  7   8: {{.*}} (95) exit")
+__naked void gotox_jump_table(void)
+{
+	asm volatile ("						\
+	.pushsection .jumptables,\"\",@progbits;		\
+jt0_%=:								\
+	.quad ret0_%= - socket;					\
+	.quad ret1_%= - socket;					\
+	.size jt0_%=, 16;					\
+	.global jt0_%=;						\
+	.popsection;						\
+								\
+	r0 = jt0_%= ll;						\
+	r0 += 8;						\
+	r0 = *(u64 *)(r0 + 0);					\
+	.8byte %[gotox_r0];					\
+ret0_%=:							\
+	r0 = 0;							\
+	exit;							\
+ret1_%=:							\
+	r0 = 1;							\
+	exit;							\
+"	:
+	: __imm_insn(gotox_r0, BPF_RAW_INSN(BPF_JMP | BPF_JA | BPF_X, BPF_REG_0, 0, 0, 0))
+	: __clobber_all);
+}
+#endif
+
+char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (41 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
@ 2026-10-04 13:38 ` Eduard Zingerman
  2026-10-04 14:24   ` bot+bpf-ci
  2026-10-04 20:04 ` [syzbot ci] Re: bpf: use scalar evolution to widen bounded loops syzbot ci
  2026-10-06 16:40 ` [PATCH bpf-next v2 00/43] " patchwork-bot+netdevbpf
  44 siblings, 1 reply; 69+ messages in thread
From: Eduard Zingerman @ 2026-10-04 13:38 UTC (permalink / raw)
  To: bpf, ast; +Cc: andrii, daniel, kernel-team, yonghong.song, Eduard Zingerman

- Computed expression shapes: register and spilled-counter recurrences,
  substitution at latches, agreeing and conflicting backedge updates,
  and ALU, extension and byte-swap expression chains.

- Iteration bounds and header-visit counts for:
  - pre-condition loops using JEQ, JGE and JGT;
  - post-condition loops using JLT, JLE, JGE, JGT and JNE;
  - signed post-conditions using JSLT and JSLE;
  - increasing and decreasing counters, nonzero initial values,
    immediate and invariant-register bounds, and immediate exits;
  - additional exits that reduce the minimum iteration count.

- Widening, backedge clamping and state-pruning convergence,
  including multiple induction variables used for map-value accesses
  and nonconstant entry values requiring a common alignment.

- Nested and sibling loops, exits to another loop header, exits across
  multiple nesting levels, and side entries into loop nests.
  Check conservative summaries for irreducible children while still
  widening supported reducible descendants.

- Invalidation of stack-slot expressions after indirect stack writes.

- Avoid widening registers used to address stack spills, fills and
  dynptr constructor arguments, including dependencies through nested
  loops. Conversely, allow widening addresses of byte-sized stores.

- Conditional assignments joining scalar values and stack pointers,
  including preservation of a common base and gcd-derived step.

- Conservative fallback for unsupported recurrences, unavailable
  counter spills and zero computed header counts.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c |    2 +
 tools/testing/selftests/bpf/progs/verifier_scev.c | 2376 +++++++++++++++++++++
 2 files changed, 2378 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index b45916e04fef..c69c0f7e9cc8 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -108,6 +108,7 @@
 #include "verifier_ringbuf.skel.h"
 #include "verifier_runtime_jit.skel.h"
 #include "verifier_scalar_ids.skel.h"
+#include "verifier_scev.skel.h"
 #include "verifier_sdiv.skel.h"
 #include "verifier_search_pruning.skel.h"
 #include "verifier_sock.skel.h"
@@ -296,6 +297,7 @@ void test_verifier_regalloc(void)             { RUN(verifier_regalloc); }
 void test_verifier_ringbuf(void)              { RUN(verifier_ringbuf); }
 void test_verifier_runtime_jit(void)          { RUN(verifier_runtime_jit); }
 void test_verifier_scalar_ids(void)           { RUN(verifier_scalar_ids); }
+void test_verifier_scev(void)                 { RUN(verifier_scev); }
 void test_verifier_sdiv(void)                 { RUN(verifier_sdiv); }
 void test_verifier_search_pruning(void)       { RUN(verifier_search_pruning); }
 void test_verifier_sock(void)                 { RUN(verifier_sock); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_scev.c b/tools/testing/selftests/bpf/progs/verifier_scev.c
new file mode 100644
index 000000000000..7a1b5b8ec287
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_scev.c
@@ -0,0 +1,2376 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <stdbool.h>
+#include <bpf/bpf_helpers.h>
+#include "../../../include/linux/filter.h"
+#include "bpf_misc.h"
+#include "bpf_kfuncs.h"
+
+#define COND_TEST(name, _initial, _step, _bound, msg, body)                     \
+SEC("xdp")                                                                      \
+__success                                                                       \
+__log_level(2)                                                                  \
+__flag(BPF_F_TEST_STATE_FREQ)                                                   \
+__msg("loop header at {{[0-9]+}}, " msg)                                        \
+__naked void name(void)                                                         \
+{                                                                               \
+	asm volatile (                                                          \
+		"r6 = %[initial] ll;"                                           \
+		"r7 = %[bound] ll;"                                             \
+		"r9 = 0;"                                                       \
+		body                                                            \
+		"2: r0 = 0;"                                                    \
+		"exit;"                                                         \
+		:                                                               \
+		: __imm_const(initial, _initial),                               \
+		  __imm_const(step, _step),                                     \
+		  __imm_const(bound, _bound)                                    \
+		: __clobber_all);                                               \
+}
+
+/* Optional assembly runs before the tested condition, preserving its latch. */
+#define PRE__COND_TEST(name, op, initial, step, bound, msg, ...)                 \
+	COND_TEST(name, initial, step, bound, msg,                              \
+		"1: r8 = %[step] ll;" __VA_ARGS__                               \
+		"if r6 " op " r7 goto 2f;"                                      \
+		"r6 += r8;"                                                     \
+		"goto 1b;")
+
+#define POST_COND_TEST(name, op, initial, step, bound, msg, ...)                \
+	COND_TEST(name, initial, step, bound, msg,                              \
+		"1: r8 = %[step] ll;" __VA_ARGS__                               \
+		"r6 += r8;"                                                     \
+		"if r6 " op " r7 goto 1b;")
+
+/* Bound fallback verification of otherwise infinite or extremely long loops. */
+#define LIMIT_ITERATIONS "r9 += 1; if r9 > 8 goto 2f;"
+
+#define PRE__COND_TEST2(name, op, initial, step, bound, msg)                     \
+	PRE__COND_TEST(name, op, initial, step, bound, msg, LIMIT_ITERATIONS)
+
+#define ITER_U64_MAX (~0ULL)
+#define ITER_S64_MAX ((1ULL << 63) - 1)
+#define ITER_S64_MIN (1ULL << 63)
+#define ITER_U32_MAX ((1ULL << 32) - 1)
+
+struct map_val {
+	char foo[1024];
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_HASH);
+	__uint(max_entries, 1);
+	__type(key, int);
+	__type(value, struct map_val);
+} map SEC(".maps");
+
+typeof(map) other_map SEC(".maps");
+
+struct {
+	__uint(type, BPF_MAP_TYPE_RINGBUF);
+	__uint(max_entries, 4096);
+} ringbuf SEC(".maps");
+
+struct bpf_iter__bpf_map_elem {
+	struct bpf_iter_meta *meta;
+	struct bpf_map *map;
+	void *key;
+	void *value;
+};
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__naked void simple_loop1(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+loop_%=:						\
+	if r0 == 10 goto exit_%=;			\
+	r0 += 1;					\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("8: (7b) *(u64 *)(r1 +0) = r0     ; *fp-8 (+ *fp-8 1) -> ?")
+__naked void indirect_write_invalidates_scev(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 8) = r0;				\
+loop_%=:						\
+	r0 = *(u64 *)(r10 - 8);				\
+	if r0 == 10 goto exit_%=;			\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	r1 = r10;					\
+	r1 += -8;					\
+	*(u64 *)(r1 + 0) = r0;				\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 2:")
+__msg_next("  *fp-8=(+ *fp-8 1) / (linear *fp-8 1)")
+__msg_next(" scev at latch 3:")
+__msg_next("  r0=*fp-8 / (linear *fp-8 1)")
+__msg_next("  *fp-8=*fp-8 / (linear *fp-8 1)")
+__naked void simple_loop2(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 8) = r0;				\
+loop_%=:						\
+	r0 = *(u64 *)(r10 - 8);				\
+	if r0 == 10 goto exit_%=;			\
+	r0 = *(u64 *)(r10 - 8);				\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 2:")
+__msg_next("  r1=(+ r1 1) / (linear r1 1)")
+__naked void meet_agrees(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+1:							\
+	if r1 == 2 goto 3f;				\
+	if r0 == 7 goto 2f;				\
+	r1 += 1;					\
+	goto 1b;					\
+2:							\
+	r1 += 1;					\
+	goto 1b;					\
+3:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__failure
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(any r6 (+ r6 1)) / ?")
+__naked void meet_disagrees(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 3f;				\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 7 goto 1b;				\
+	r6 += 1;					\
+	goto 1b;					\
+3:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(bswap32 (bswap32 (zext32 (- (- (zext32 (+ (>>...) 1))))))) / ?")
+__naked void expr_chain(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 2f;				\
+	r6 += 1;					\
+	r6 <<= 32;					\
+	r6 >>= 32;					\
+	w6 += 1;					\
+	r6 = -r6;					\
+	w6 = -w6;					\
+	r6 = bswap32 r6;				\
+	r6 = bswap32 r6;				\
+	goto 1b;					\
+2:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+#define ALU_OP(insn) "r1 = r6; " insn "; r0 += r1;"
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("r1 += 2 {{.*}}; r1 r6 -> (+ r6 2)")
+__msg("r1 += r7 {{.*}}; r1 r6 -> (+ r6 r7)")
+__msg("w1 += 2 {{.*}}; r1 r6 -> (zext32 (+ r6 2))")
+__msg("w1 += w7 {{.*}}; r1 r6 -> (zext32 (+ r6 r7))")
+__msg("r1 -= 2 {{.*}}; r1 r6 -> (- r6 2)")
+__msg("r1 -= r7 {{.*}}; r1 r6 -> (- r6 r7)")
+__msg("w1 -= 2 {{.*}}; r1 r6 -> (zext32 (- r6 2))")
+__msg("w1 -= w7 {{.*}}; r1 r6 -> (zext32 (- r6 r7))")
+__msg("r1 *= 2 {{.*}}; r1 r6 -> (* r6 2)")
+__msg("r1 *= r7 {{.*}}; r1 r6 -> (* r6 r7)")
+__msg("w1 *= 2 {{.*}}; r1 r6 -> (zext32 (* r6 2))")
+__msg("w1 *= w7 {{.*}}; r1 r6 -> (zext32 (* r6 r7))")
+__msg("r1 /= 2 {{.*}}; r1 r6 -> (/ r6 2)")
+__msg("r1 /= r7 {{.*}}; r1 r6 -> (/ r6 r7)")
+__msg("w1 /= 2 {{.*}}; r1 r6 -> (zext32 (/ r6 2))")
+__msg("w1 /= w7 {{.*}}; r1 r6 -> (zext32 (/ r6 r7))")
+__msg("r1 s/= 2 {{.*}}; r1 r6 -> (s/ r6 2)")
+__msg("r1 s/= r7 {{.*}}; r1 r6 -> (s/ r6 r7)")
+__msg("w1 s/= 2 {{.*}}; r1 r6 -> (zext32 (s/ r6 2))")
+__msg("w1 s/= w7 {{.*}}; r1 r6 -> (zext32 (s/ r6 r7))")
+__msg("r1 %= 2 {{.*}}; r1 r6 -> (% r6 2)")
+__msg("r1 %= r7 {{.*}}; r1 r6 -> (% r6 r7)")
+__msg("w1 %= 2 {{.*}}; r1 r6 -> (zext32 (% r6 2))")
+__msg("w1 %= w7 {{.*}}; r1 r6 -> (zext32 (% r6 r7))")
+__msg("r1 s%= 2 {{.*}}; r1 r6 -> (s% r6 2)")
+__msg("r1 s%= r7 {{.*}}; r1 r6 -> (s% r6 r7)")
+__msg("w1 s%= 2 {{.*}}; r1 r6 -> (zext32 (s% r6 2))")
+__msg("w1 s%= w7 {{.*}}; r1 r6 -> (zext32 (s% r6 r7))")
+__msg("r1 |= 2 {{.*}}; r1 r6 -> (| r6 2)")
+__msg("r1 |= r7 {{.*}}; r1 r6 -> (| r6 r7)")
+__msg("w1 |= 2 {{.*}}; r1 r6 -> (zext32 (| r6 2))")
+__msg("w1 |= w7 {{.*}}; r1 r6 -> (zext32 (| r6 r7))")
+__msg("r1 &= 2 {{.*}}; r1 r6 -> (& r6 2)")
+__msg("r1 &= r7 {{.*}}; r1 r6 -> (& r6 r7)")
+__msg("w1 &= 2 {{.*}}; r1 r6 -> (zext32 (& r6 2))")
+__msg("w1 &= w7 {{.*}}; r1 r6 -> (zext32 (& r6 r7))")
+__msg("r1 ^= 2 {{.*}}; r1 r6 -> (^ r6 2)")
+__msg("r1 ^= r7 {{.*}}; r1 r6 -> (^ r6 r7)")
+__msg("w1 ^= 2 {{.*}}; r1 r6 -> (zext32 (^ r6 2))")
+__msg("w1 ^= w7 {{.*}}; r1 r6 -> (zext32 (^ r6 r7))")
+__msg("r1 <<= 2 {{.*}}; r1 r6 -> (<< r6 2)")
+__msg("r1 <<= r7 {{.*}}; r1 r6 -> (<< r6 r7)")
+__msg("w1 <<= 2 {{.*}}; r1 r6 -> (zext32 (<< r6 2))")
+__msg("w1 <<= w7 {{.*}}; r1 r6 -> (zext32 (<< r6 r7))")
+__msg("r1 >>= 2 {{.*}}; r1 r6 -> (>> r6 2)")
+__msg("r1 >>= r7 {{.*}}; r1 r6 -> (>> r6 r7)")
+__msg("w1 >>= 2 {{.*}}; r1 r6 -> (zext32 (>> r6 2))")
+__msg("w1 >>= w7 {{.*}}; r1 r6 -> (zext32 (>> r6 r7))")
+__msg("r1 s>>= 2 {{.*}}; r1 r6 -> (s>> r6 2)")
+__msg("r1 s>>= r7 {{.*}}; r1 r6 -> (s>> r6 r7)")
+__msg("w1 s>>= 2 {{.*}}; r1 r6 -> (zext32 (s>> r6 2))")
+__msg("w1 s>>= w7 {{.*}}; r1 r6 -> (zext32 (s>> r6 r7))")
+__msg("r1 = -r1 {{.*}}; r1 r6 -> (- r6)")
+__msg("w1 = -w1 {{.*}}; r1 r6 -> (zext32 (- r6))")
+__msg("r1 = 42 {{.*}}; r1 r6 -> 42")
+__msg("w1 = 42 {{.*}}; r1 r6 -> (zext32 42)")
+__msg("r1 = r7 {{.*}}; r1 r6 -> r7")
+__msg("w1 = w7 {{.*}}; r1 r6 -> (zext32 r7)")
+__msg("r1 = (s8)r6 {{.*}}; r1 r6 -> (sext8 r6)")
+__msg("r1 = (s16)r6 {{.*}}; r1 r6 -> (sext16 r6)")
+__msg("r1 = (s32)r6 {{.*}}; r1 r6 -> (sext32 r6)")
+__msg("w1 = (s8)w6 {{.*}}; r1 r6 -> (zext32 (sext8 r6))")
+__msg("w1 = (s16)w6 {{.*}}; r1 r6 -> (zext32 (sext16 r6))")
+__msg("r1 = bswap16 r1 {{.*}}; r1 r6 -> (bswap16 r6)")
+__msg("r1 = bswap32 r1 {{.*}}; r1 r6 -> (bswap32 r6)")
+__msg("r1 = bswap64 r1 {{.*}}; r1 r6 -> (bswap64 r6)")
+#if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__
+__msg("r1 = le16 r1 {{.*}}; r1 r6 -> (zext16 r6)")
+__msg("r1 = le32 r1 {{.*}}; r1 r6 -> (zext32 r6)")
+__msg("r1 += 1 {{.*}}; r1 r6 -> (+ r6 1)")
+__msg("r1 = be16 r1 {{.*}}; r1 r6 -> (bswap16 r6)")
+__msg("r1 = be32 r1 {{.*}}; r1 r6 -> (bswap32 r6)")
+__msg("r1 = be64 r1 {{.*}}; r1 r6 -> (bswap64 r6)")
+#else
+__msg("r1 = le16 r1 {{.*}}; r1 r6 -> (bswap16 r6)")
+__msg("r1 = le32 r1 {{.*}}; r1 r6 -> (bswap32 r6)")
+__msg("r1 = le64 r1 {{.*}}; r1 r6 -> (bswap64 r6)")
+__msg("r1 = be16 r1 {{.*}}; r1 r6 -> (zext16 r6)")
+__msg("r1 = be32 r1 {{.*}}; r1 r6 -> (zext32 r6)")
+__msg("r1 += 1 {{.*}}; r1 r6 -> (+ r6 1)")
+#endif
+__naked void alu_ops_exprs(void)
+{
+	asm volatile (
+	"r6 = 7;"
+	"r7 = 2;"
+	"r8 = 0;"
+"1:"
+	"r0 = 0;"
+	ALU_OP("r1 += 2")
+	ALU_OP("r1 += r7")
+	ALU_OP("w1 += 2")
+	ALU_OP("w1 += w7")
+	ALU_OP("r1 -= 2")
+	ALU_OP("r1 -= r7")
+	ALU_OP("w1 -= 2")
+	ALU_OP("w1 -= w7")
+	ALU_OP("r1 *= 2")
+	ALU_OP("r1 *= r7")
+	ALU_OP("w1 *= 2")
+	ALU_OP("w1 *= w7")
+	ALU_OP("r1 /= 2")
+	ALU_OP("r1 /= r7")
+	ALU_OP("w1 /= 2")
+	ALU_OP("w1 /= w7")
+	ALU_OP("r1 s/= 2")
+	ALU_OP("r1 s/= r7")
+	ALU_OP("w1 s/= 2")
+	ALU_OP("w1 s/= w7")
+	ALU_OP("r1 %%= 2")
+	ALU_OP("r1 %%= r7")
+	ALU_OP("w1 %%= 2")
+	ALU_OP("w1 %%= w7")
+	ALU_OP("r1 s%%= 2")
+	ALU_OP("r1 s%%= r7")
+	ALU_OP("w1 s%%= 2")
+	ALU_OP("w1 s%%= w7")
+	ALU_OP("r1 |= 2")
+	ALU_OP("r1 |= r7")
+	ALU_OP("w1 |= 2")
+	ALU_OP("w1 |= w7")
+	ALU_OP("r1 &= 2")
+	ALU_OP("r1 &= r7")
+	ALU_OP("w1 &= 2")
+	ALU_OP("w1 &= w7")
+	ALU_OP("r1 ^= 2")
+	ALU_OP("r1 ^= r7")
+	ALU_OP("w1 ^= 2")
+	ALU_OP("w1 ^= w7")
+	ALU_OP("r1 <<= 2")
+	ALU_OP("r1 <<= r7")
+	ALU_OP("w1 <<= 2")
+	ALU_OP("w1 <<= w7")
+	ALU_OP("r1 >>= 2")
+	ALU_OP("r1 >>= r7")
+	ALU_OP("w1 >>= 2")
+	ALU_OP("w1 >>= w7")
+	ALU_OP("r1 s>>= 2")
+	ALU_OP("r1 s>>= r7")
+	ALU_OP("w1 s>>= 2")
+	ALU_OP("w1 s>>= w7")
+	ALU_OP("r1 = -r1")
+	ALU_OP("w1 = -w1")
+	ALU_OP("r1 = 42")
+	ALU_OP("w1 = 42")
+	ALU_OP("r1 = r7")
+	ALU_OP("w1 = w7")
+	ALU_OP("r1 = (s8)r6")
+	ALU_OP("r1 = (s16)r6")
+	ALU_OP("r1 = (s32)r6")
+	ALU_OP("w1 = (s8)w6")
+	ALU_OP("w1 = (s16)w6")
+	ALU_OP("r1 = bswap16 r1")
+	ALU_OP("r1 = bswap32 r1")
+	ALU_OP("r1 = bswap64 r1")
+	ALU_OP("r1 = le16 r1")
+	ALU_OP("r1 = le32 r1")
+	/*
+	 * Expressions unchanged by transfer() do not appear in the log.
+	 * Add r1 += 1 to touch the expression and make it appear in the log.
+	 */
+	ALU_OP("r1 = le64 r1; r1 += 1")
+	ALU_OP("r1 = be16 r1")
+	ALU_OP("r1 = be32 r1")
+	ALU_OP("r1 = be64 r1; r1 += 1")
+	"r8 += 1;"
+	"if r8 < 2 goto 1b;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+#undef ALU_OP
+
+/* Reset operands and consume results to keep transfer changes live. */
+#define ST_OP(stmt)							\
+	"r0 = 0; r1 = 1; r2 = 2;"					\
+	"*(u64 *)(r10 - 8) = -8;"					\
+	stmt ";"							\
+	"r0 += r1; r0 += r2;"						\
+	"r3 = *(u64 *)(r10 - 8); r0 += r3;"
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("*(u64 *)(r10 -8) = r1 {{.*}}; *fp-8 -8 -> 1")
+__msg("*(u32 *)(r10 -8) = r1 {{.*}}; *fp-8 -8 -> (spill32 1)")
+__msg("*(u16 *)(r10 -8) = r1 {{.*}}; *fp-8 -8 -> (spill16 1)")
+__msg("*(u8 *)(r10 -8) = r1 {{.*}}; *fp-8 -8 -> (spill8 1)")
+__msg("r1 = *(u64 *)(r10 -8) {{.*}}; r1 1 -> -8")
+#if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__
+__msg("r1 = *(u32 *)(r10 -8) {{.*}}; r1 1 -> (zext32 -8)")
+__msg("r1 = *(u16 *)(r10 -8) {{.*}}; r1 1 -> (zext16 -8)")
+__msg("r1 = *(u8 *)(r10 -8) {{.*}}; r1 1 -> (zext8 -8)")
+__msg("r1 = *(s32 *)(r10 -8) {{.*}}; r1 1 -> (sext32 -8)")
+__msg("r1 = *(s16 *)(r10 -8) {{.*}}; r1 1 -> (sext16 -8)")
+__msg("r1 = *(s8 *)(r10 -8) {{.*}}; r1 1 -> (sext8 -8)")
+__msg("r1 = load_acquire((u8 *)(r10 -8)) {{.*}}; r1 1 -> (zext8 -8)")
+__msg("r1 = load_acquire((u16 *)(r10 -8)) {{.*}}; r1 1 -> (zext16 -8)")
+__msg("r1 = load_acquire((u32 *)(r10 -8)) {{.*}}; r1 1 -> (zext32 -8)")
+#else
+__msg("r1 = *(u32 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = *(u16 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = *(u8 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = *(s32 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = *(s16 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = *(s8 *)(r10 -8) {{.*}}; r1 1 -> _")
+__msg("r1 = load_acquire((u8 *)(r10 -8)) {{.*}}; r1 1 -> _")
+__msg("r1 = load_acquire((u16 *)(r10 -8)) {{.*}}; r1 1 -> _")
+__msg("r1 = load_acquire((u32 *)(r10 -8)) {{.*}}; r1 1 -> _")
+#endif
+__msg("r2 = load_acquire((u64 *)(r10 -8)) {{.*}}; r2 2 -> -8")
+__msg("store_release((u8 *)(r10 -8), r1) {{.*}}; *fp-8 -8 -> (spill8 1)")
+__msg("store_release((u16 *)(r10 -8), r1) {{.*}}; *fp-8 -8 -> (spill16 1)")
+__msg("store_release((u32 *)(r10 -8), r1) {{.*}}; *fp-8 -8 -> (spill32 1)")
+__msg("store_release((u64 *)(r10 -8), r1) {{.*}}; *fp-8 -8 -> 1")
+__msg("lock *(u64 *)(r10 -8) += r1 {{.*}}; *fp-8 -8 -> _")
+__msg("r1 = atomic64_fetch_add((u64 *)(r10 -8), r1) {{.*}}; r1 1 -> _, *fp-8 -8 -> _")
+__msg("r1 = atomic_fetch_add((u32 *)(r10 -8), r1) {{.*}}; r1 1 -> _, *fp-8 -8 -> _")
+__msg("r1 = atomic64_xchg((u64 *)(r10 -8), r1) {{.*}}; r1 1 -> _, *fp-8 -8 -> _")
+__msg("r0 = atomic64_cmpxchg((u64 *)(r10 -8), r0, r2) {{.*}}; r0 0 -> _, *fp-8 -8 -> _")
+__msg("*(u64 *)(r10 -8) = 42 {{.*}}; *fp-8 -8 -> 42")
+__msg("*(u8 *)(r10 -7) = r0 {{.*}}; *fp-8 -8 -> _")
+__naked void store_load_exprs(void)
+{
+	asm volatile (
+	"r7 = 0;"
+"1:"
+	ST_OP("*(u64 *)(r10 - 8) = r1")
+	ST_OP("*(u32 *)(r10 - 8) = r1")
+	ST_OP("*(u16 *)(r10 - 8) = r1")
+	ST_OP("*(u8 *)(r10 - 8) = r1")
+
+	ST_OP("r1 = *(u64 *)(r10 - 8)")
+	ST_OP("r1 = *(u32 *)(r10 - 8)")
+	ST_OP("r1 = *(u16 *)(r10 - 8)")
+	ST_OP("r1 = *(u8 *)(r10 - 8)")
+	ST_OP("r1 = *(s32 *)(r10 - 8)")
+	ST_OP("r1 = *(s16 *)(r10 - 8)")
+	ST_OP("r1 = *(s8 *)(r10 - 8)")
+
+	ST_OP(".8byte %[load_acquire8]")
+	ST_OP(".8byte %[load_acquire16]")
+	ST_OP(".8byte %[load_acquire32]")
+	ST_OP(".8byte %[load_acquire64]")
+	ST_OP(".8byte %[store_release8]")
+	ST_OP(".8byte %[store_release16]")
+	ST_OP(".8byte %[store_release32]")
+	ST_OP(".8byte %[store_release64]")
+
+	ST_OP(".8byte %[atomic_add64]")
+	ST_OP(".8byte %[atomic_fetch_add64]")
+	ST_OP(".8byte %[atomic_fetch_add32]")
+	ST_OP(".8byte %[atomic_xchg64]")
+	ST_OP(".8byte %[atomic_cmpxchg64]")
+
+	ST_OP("*(u64 *)(r10 - 8) = 42")
+	ST_OP("*(u8 *)(r10 - 7) = r0")
+
+	"r7 += 1;"
+	"if r7 < 2 goto 1b;"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm_insn(atomic_add64,       BPF_ATOMIC_OP(BPF_DW, BPF_ADD,       BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(atomic_xchg64,      BPF_ATOMIC_OP(BPF_DW, BPF_XCHG,      BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(load_acquire8,      BPF_ATOMIC_OP(BPF_B,  BPF_LOAD_ACQ,  BPF_REG_1,  BPF_REG_10, -8)),
+	  __imm_insn(load_acquire16,     BPF_ATOMIC_OP(BPF_H,  BPF_LOAD_ACQ,  BPF_REG_1,  BPF_REG_10, -8)),
+	  __imm_insn(load_acquire32,     BPF_ATOMIC_OP(BPF_W,  BPF_LOAD_ACQ,  BPF_REG_1,  BPF_REG_10, -8)),
+	  __imm_insn(load_acquire64,     BPF_ATOMIC_OP(BPF_DW, BPF_LOAD_ACQ,  BPF_REG_2,  BPF_REG_10, -8)),
+	  __imm_insn(store_release8,     BPF_ATOMIC_OP(BPF_B,  BPF_STORE_REL, BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(store_release16,    BPF_ATOMIC_OP(BPF_H,  BPF_STORE_REL, BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(store_release32,    BPF_ATOMIC_OP(BPF_W,  BPF_STORE_REL, BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(store_release64,    BPF_ATOMIC_OP(BPF_DW, BPF_STORE_REL, BPF_REG_10, BPF_REG_1,  -8)),
+	  __imm_insn(atomic_cmpxchg64,   BPF_ATOMIC_OP(BPF_DW, BPF_CMPXCHG,   BPF_REG_10, BPF_REG_2,  -8)),
+	  __imm_insn(atomic_fetch_add32, BPF_ATOMIC_OP(BPF_W,  BPF_ADD | BPF_FETCH, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_fetch_add64, BPF_ATOMIC_OP(BPF_DW, BPF_ADD | BPF_FETCH, BPF_REG_10, BPF_REG_1, -8))
+	: __clobber_all);
+}
+
+#undef ST_OP
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next(" scev at latch 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next("scev at header 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next(" scev at latch 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__naked void nested_loop1(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 2f;				\
+	r6 += 1;					\
+	r7 = 0;						\
+3:							\
+	if r7 == 2 goto 4f;				\
+	r7 += 1;					\
+	goto 3b;					\
+4:							\
+	goto 1b;					\
+2:							\
+	r0 = r7;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__msg("loop at 1")
+__msg_next("  backedge from 4, latch at -1")
+__msg_next("  exit from 1 to 7")
+__msg_next("loop at 4, nested in 1")
+__msg_next("  backedge from 6, latch at 4")
+__msg_next("  exit from 4 to 1")
+__msg("scev at header 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next("scev at header 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next(" scev at latch 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg("loop header at 1, latch not identified")
+__msg("loop header at 4, widening r7 to 0..2 step 1")
+__log_level(2)
+__naked void nested_loop_hdr_backedge1(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 3f;				\
+	r6 += 1;					\
+	r7 = 0;						\
+2:							\
+	  if r7 == 2 goto 1b;				\
+	  r7 += 1;					\
+	  goto 2b;					\
+3:							\
+	r0 = r7;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Outer loop {2, 3, 4, 5} contains inner loop {3, 4, 5}.
+ * The outer loop's only backedge is 4 -> 2, and its latch is also at 4,
+ * inside the inner loop. Outer loop's latch index is not computed in
+ * such a case, as there is no mechanism to build SCEVs for such latches
+ * at the moment.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop at 2{{$}}")
+__msg_next("  backedge from 4, latch at -1")
+__msg_next("  exit from 5 to 6")
+__msg_next("loop at 3, nested in 2")
+__msg_next("  backedge from 5, latch at 5")
+__msg_next("  exit from 4 to 2")
+__msg_next("  exit from 5 to 6")
+__msg("scev at header 2:")
+__msg_next("  r6=?")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next("scev at header 3:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next(" scev at latch 5:")
+__msg_next("  r6=(+ r6 1) / (linear (+ r6 1) 1)")
+__msg("loop header at 2, latch not identified")
+__msg("loop header at 3, header_count is [0..2] ")
+__naked void nested_loop_hdr_backedge2(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+	r7 = 0;						\
+1:	r7 += 1;					\
+2:	r6 += 1;					\
+	if r7 == 2 goto 1b;				\
+	if r6 < 2 goto 2b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Logical nest: for (r6 = 0; r6 < 3; r6++)
+ *                   for (r7 = 0; r7 < 4; r7++) r8++;
+ * Both backedges target header 3, so the CFG has one loop. r6 advances only
+ * on the outer backedge and r7 resets there; neither has a linear SCEV.
+ * r8 advances on both backedges and retains its linear SCEV.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop at 3{{$}}")
+__msg("scev at header 3:")
+__msg_next("  r6=(any r6 (+ r6 1)) / ?")
+__msg_next("  r7=(any (+ r7 1) 0) / ?")
+__msg_next("  r8=(+ r8 1) / (linear r8 1)")
+__msg("loop header at 3, unsupported loop: multiple backedges")
+__naked void nested_loops_shared_header(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+	r7 = 0;						\
+	r8 = 0;						\
+1:	r8 += 1;					\
+	r7 += 1;					\
+	if r7 < 4 goto 1b;				\
+	r7 = 0;						\
+	r6 += 1;					\
+	if r6 < 3 goto 1b;				\
+	r0 = r8;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Exact boundary hits for every continuation opcode.
+ *
+ * Example: pre_lt
+ *
+ *    u64 r6 = 1;
+ *    while (r6 < 7) {               // r6 ∈ [1,3,5,7]
+ *        r6 += 2;                   // r6 ∈ [3,5,7]
+ *    }
+ */
+/*              name               op              initial          step             bound  expected message */
+PRE__COND_TEST (pre_lt,            ">=",                 1,            2,                7, "header_count is 4 ")
+POST_COND_TEST (post_lt,           "<",                  1,            2,                7, "header_count is 3 ")
+PRE__COND_TEST (pre_le,            ">",                  1,            2,                7, "header_count is 5 ")
+POST_COND_TEST (post_le,           "<=",                 1,            2,                7, "header_count is 4 ")
+PRE__COND_TEST (pre_gt,            "<=",                 9,           -2,                3, "header_count is 4 ")
+POST_COND_TEST (post_gt,           ">",                  9,           -2,                3, "header_count is 3 ")
+PRE__COND_TEST (pre_ge,            "<",                  9,           -2,                3, "header_count is 5 ")
+POST_COND_TEST (post_ge,           ">=",                 9,           -2,                3, "header_count is 4 ")
+PRE__COND_TEST (pre_slt,           "s>=",               -5,            2,                1, "header_count is 4 ")
+POST_COND_TEST (post_slt,          "s<",                -5,            2,                1, "header_count is 3 ")
+PRE__COND_TEST (pre_sle,           "s>",                -5,            2,                1, "header_count is 5 ")
+POST_COND_TEST (post_sle,          "s<=",               -5,            2,                1, "header_count is 4 ")
+PRE__COND_TEST (pre_sgt,           "s<=",                5,           -2,               -1, "header_count is 4 ")
+POST_COND_TEST (post_sgt,          "s>",                 5,           -2,               -1, "header_count is 3 ")
+PRE__COND_TEST (pre_sge,           "s<",                 5,           -2,               -1, "header_count is 5 ")
+POST_COND_TEST (post_sge,          "s>=",                5,           -2,               -1, "header_count is 4 ")
+PRE__COND_TEST (pre_ne,            "==",                 1,            2,                7, "header_count is 4 ")
+POST_COND_TEST (post_ne,           "!=",                 1,            2,                7, "header_count is 3 ")
+PRE__COND_TEST (pre_ne_down,       "==",                 9,           -2,                3, "header_count is 4 ")
+POST_COND_TEST (post_ne_down,      "!=",                 9,           -2,                3, "header_count is 3 ")
+
+/*
+ * Bounds between two consecutive counter values.
+ *
+ * Example: pre_lt_round
+ *
+ *    u64 r6 = 1;
+ *    while (r6 < 8) {               // r6 ∈ [1,3,5,7,9]
+ *        r6 += 2;                   // r6 ∈ [3,5,7,9]
+ *    }
+ */
+PRE__COND_TEST (pre_lt_round,      ">=",                 1,            2,                8, "header_count is 5 ")
+PRE__COND_TEST (pre_le_round,      ">",                  1,            2,                8, "header_count is 5 ")
+PRE__COND_TEST (pre_gt_round,      "<=",                 9,           -2,                2, "header_count is 5 ")
+PRE__COND_TEST (pre_ge_round,      "<",                  9,           -2,                2, "header_count is 5 ")
+PRE__COND_TEST (pre_slt_round,     "s>=",               -5,            2,                2, "header_count is 5 ")
+PRE__COND_TEST (pre_sle_round,     "s>",                -5,            2,                2, "header_count is 5 ")
+PRE__COND_TEST (pre_sgt_round,     "s<=",                5,           -2,               -2, "header_count is 5 ")
+PRE__COND_TEST (pre_sge_round,     "s<",                 5,           -2,               -2, "header_count is 5 ")
+
+/*
+ * Equality on the first non-strict comparison still takes a backedge.
+ *
+ * Example: pre_le_eq
+ *
+ *    u64 r6 = 7;
+ *    while (r6 <= 7) {              // r6 ∈ [7,9]
+ *        r6 += 2;                   // r6 ∈ [9]
+ *    }
+ */
+PRE__COND_TEST (pre_le_eq,         ">",                  7,            2,                7, "header_count is 2 ")
+PRE__COND_TEST (pre_ge_eq,         "<",                  3,           -2,                3, "header_count is 2 ")
+PRE__COND_TEST (pre_sle_eq,        "s>",                 1,            2,                1, "header_count is 2 ")
+PRE__COND_TEST (pre_sge_eq,        "s<",                -1,           -2,               -1, "header_count is 2 ")
+
+/*
+ * Unsigned order across the sign bit.
+ *
+ * Example: pre_lt_sign
+ *
+ *    u64 H = 1ULL << 63, r6 = H - 2;
+ *    while (r6 < H + 1) {           // r6 ∈ [H-2,H-1,H,H+1]
+ *        r6 += 1;                   // r6 ∈ [H-1,H,H+1]
+ *    }
+ */
+PRE__COND_TEST (pre_lt_sign,       ">=",  ITER_S64_MAX - 1,            1, ITER_S64_MIN + 1, "header_count is 4 ")
+PRE__COND_TEST (pre_gt_sign,       "<=",  ITER_S64_MIN + 1,           -1, ITER_S64_MAX - 1, "header_count is 4 ")
+
+/*
+ * The first latch comparison is false.
+ *
+ * Example: pre_lt_false
+ *
+ *    u64 r6 = 7;
+ *    while (r6 < 7) {               // r6 ∈ [7]
+ *        r6 += 2;                   // unreachable
+ *    }
+ */
+PRE__COND_TEST (pre_lt_false,      ">=",                 7,            2,                7, "can't compute iterations count")
+PRE__COND_TEST (pre_le_false,      ">",                  8,            2,                7, "can't compute iterations count")
+PRE__COND_TEST (pre_gt_false,      "<=",                 3,           -2,                3, "can't compute iterations count")
+PRE__COND_TEST (pre_ge_false,      "<",                  2,           -2,                3, "can't compute iterations count")
+PRE__COND_TEST (pre_slt_false,     "s>=",                1,            2,                1, "can't compute iterations count")
+PRE__COND_TEST (pre_sle_false,     "s>",                 2,            2,                1, "can't compute iterations count")
+PRE__COND_TEST (pre_sgt_false,     "s<=",               -1,           -2,               -1, "can't compute iterations count")
+PRE__COND_TEST (pre_sge_false,     "s<",                -2,           -2,               -1, "can't compute iterations count")
+PRE__COND_TEST (pre_ne_false,      "==",                 7,            2,                7, "can't compute iterations count")
+
+/*
+ * The counter moves away from the exit; bound fallback verification.
+ *
+ * Example: pre_lt_dir
+ *
+ *    u64 r6 = 5, r9 = 0;
+ *    while (++r9 <= 8 && r6 < 7) {  // r6 ∈ [5,3,1,U64_MAX]
+ *        r6 -= 2;                   // r6 ∈ [3,1,U64_MAX]
+ *    }
+ */
+PRE__COND_TEST2(pre_lt_dir,        ">=",                 5,           -2,                7, "can't compute iterations count")
+PRE__COND_TEST2(pre_le_dir,        ">",                  5,           -2,                7, "can't compute iterations count")
+PRE__COND_TEST2(pre_gt_dir,        "<=",                 5,            2,                3, "can't compute iterations count")
+PRE__COND_TEST2(pre_ge_dir,        "<",                  5,            2,                3, "can't compute iterations count")
+PRE__COND_TEST2(pre_slt_dir,       "s>=",               -1,           -2,                1, "can't compute iterations count")
+PRE__COND_TEST2(pre_sle_dir,       "s>",                -1,           -2,                1, "can't compute iterations count")
+PRE__COND_TEST2(pre_sgt_dir,       "s<=",                1,            2,               -1, "can't compute iterations count")
+PRE__COND_TEST2(pre_sge_dir,       "s<",                 1,            2,               -1, "can't compute iterations count")
+
+/*
+ * The predicted exit would wrap the counter.
+ *
+ * Example: pre_lt_wrap
+ *
+ *    u64 M = U64_MAX, r6 = M - 1, r9 = 0;
+ *    while (++r9 <= 8 && r6 < M) {  // r6 ∈ [M-1,0,2,...,14]
+ *        r6 += 2;                   // r6 ∈ [0,2,...,14]
+ *    }
+ */
+PRE__COND_TEST2(pre_lt_wrap,       ">=",  ITER_U64_MAX - 1,            2,     ITER_U64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_le_wrap,       ">",       ITER_U64_MAX,            1,     ITER_U64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_gt_wrap,       "<=",                 1,           -2,                0, "can't compute iterations count")
+PRE__COND_TEST2(pre_ge_wrap,       "<",                  0,           -1,                0, "can't compute iterations count")
+PRE__COND_TEST2(pre_slt_wrap,      "s>=", ITER_S64_MAX - 1,            2,     ITER_S64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_sle_wrap,      "s>",      ITER_S64_MAX,            1,     ITER_S64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_sgt_wrap,      "s<=", ITER_S64_MIN + 1,           -2,     ITER_S64_MIN, "can't compute iterations count")
+PRE__COND_TEST2(pre_sge_wrap,      "s<",      ITER_S64_MIN,           -1,     ITER_S64_MIN, "can't compute iterations count")
+
+/*
+ * The non-strict count itself would overflow u64.
+ *
+ * Example: pre_le_count_ovf
+ *
+ *    u64 r6 = 0, r9 = 0;
+ *    while (++r9 <= 8 && r6 <= U64_MAX) {  // r6 ∈ [0..8]
+ *        r6 += 1;                         // r6 ∈ [1..8]
+ *    }
+ */
+PRE__COND_TEST2(pre_le_count_ovf,  ">",                  0,            1,     ITER_U64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_ge_count_ovf,  "<",       ITER_U64_MAX,           -1,                0, "can't compute iterations count")
+PRE__COND_TEST2(pre_sle_count_ovf, "s>",      ITER_S64_MIN,            1,     ITER_S64_MAX, "can't compute iterations count")
+PRE__COND_TEST2(pre_sge_count_ovf, "s<",      ITER_S64_MAX,           -1,     ITER_S64_MIN, "can't compute iterations count")
+
+/*
+ * JNE divisibility and equality across unsigned wraparound.
+ *
+ * Example: pre_ne_rem
+ *
+ *    u64 r6 = 1, r9 = 0;
+ *    while (++r9 <= 8 && r6 != 8) { // r6 ∈ [1,3,...,17]
+ *        r6 += 2;                   // r6 ∈ [3,5,...,17]
+ *    }
+ */
+PRE__COND_TEST2(pre_ne_rem,        "==",                 1,            2,                8, "can't compute iterations count")
+PRE__COND_TEST2(pre_ne_rem_down,   "==",                 9,           -2,                2, "can't compute iterations count")
+PRE__COND_TEST (pre_ne_wrap_up,    "==",  ITER_U64_MAX - 1,            2,                0, "header_count is 2 ")
+PRE__COND_TEST (pre_ne_wrap_down,  "==",                 1,           -2,     ITER_U64_MAX, "header_count is 2 ")
+
+/*
+ * Largest supported header count, sentinel, and u32 overflow.
+ *
+ * Example: pre_lt_max_hdrs
+ *
+ *    u64 r6 = 0, r9 = 0;
+ *    while (++r9 <= 8 && r6 < U32_MAX - 2) { // r6 ∈ [0..8]
+ *        r6 += 1;                            // r6 ∈ [1..8]
+ *    }
+ */
+PRE__COND_TEST2(pre_lt_max_hdrs,   ">=",                 0,            1, ITER_U32_MAX - 2, "header_count is [0..4294967294] ")
+PRE__COND_TEST2(pre_lt_sentinel,   ">=",                 0,            1, ITER_U32_MAX - 1, "can't compute iterations count")
+PRE__COND_TEST2(pre_lt_hdr_ovf,    ">=",                 0,            1,     ITER_U32_MAX, "can't compute iterations count")
+
+/*
+ * The extreme signed step values remain valid unsigned magnitudes.
+ *
+ * Example: pre_lt_step_max
+ *
+ *    u64 r6 = 0;
+ *    while (r6 < S64_MAX) {         // r6 ∈ [0,S64_MAX]
+ *        r6 += S64_MAX;             // r6 ∈ [S64_MAX]
+ *    }
+ */
+PRE__COND_TEST (pre_lt_step_max,   ">=",                 0, ITER_S64_MAX,     ITER_S64_MAX, "header_count is 2 ")
+PRE__COND_TEST (pre_gt_step_min,   "<=",      ITER_S64_MIN, ITER_S64_MIN,                0, "header_count is 2 ")
+PRE__COND_TEST (pre_ne_step_min,   "==",                 0, ITER_S64_MIN,     ITER_S64_MIN, "header_count is 2 ")
+
+/*
+ * A zero step cannot establish a finite iteration count.
+ *
+ * Example: pre_lt_zero_step
+ *
+ *    u64 r6 = 1, r9 = 0;
+ *    while (++r9 <= 8 && r6 < 7) {  // r6 ∈ [1], nine header visits
+ *        r6 += 0;                   // r6 ∈ [1]
+ *    }
+ */
+PRE__COND_TEST2(pre_lt_zero_step,  ">=",                 1,            0,                7, "can't compute iterations count")
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 10 ")
+__msg("loop header at 1, widening r0 to 0..9 step 1")
+__msg("processed 5 insns")
+__naked void post_cond_jlt(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 3, header_count is 8")
+__msg("loop header at 3, widening r0 to 2..9 step 1")
+__msg("3: R0=scalar(smin=umin=smin32=umin32=2,smax=umax=smax32=umax32=9,{{.*}})")
+__msg("processed 6 insns")
+__naked void post_cond_jlt_with_base(void)
+{
+	asm volatile ("					\
+	r7 = 10 ll;	/* ldimm64 for a twist */	\
+	r0 = 2;						\
+1:	r0 += 1;					\
+	if r0 < r7 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 4 ")
+__naked void latch_base_differs_from_step(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r1 = r0;					\
+	r1 += 1;					\
+	if r1 >= 10 goto 2f;				\
+	r0 += 4;					\
+	goto 1b;					\
+2:	r0 = r1;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 11")
+__msg("loop header at 1, widening r0 to 0..10 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=10,{{.*}})")
+__msg("processed 5 insns")
+__naked void post_cond_jle(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 <= 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 10")
+__msg("loop header at 1, widening r0 to 0..9 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=9,{{.*}})")
+__msg("processed 6 insns")
+__naked void post_cond_jge(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 >= 10 goto 2f;				\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 11")
+__msg("loop header at 1, widening r0 to 0..10 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=10,{{.*}})")
+__msg("processed 6 insns")
+__naked void pre_cond_jge(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	if r0 >= 10 goto 2f;				\
+	r0 += 1;					\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next(" scev at latch 2:")
+__msg_next("  r0=(+ r0 1) / (linear (+ r0 1) 1)")
+__msg("loop header at 1, widening r0")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,{{.*}})")
+__msg("1: (07) r0 += 1                       ; R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("2: (55) if r0 != 0x3 goto pc-2")
+__msg("3: (95) exit")
+__msg("loop header at 1, clamping r0")
+__msg("from 2 to 1: safe")
+__msg("processed 5 insns")
+__naked void post_cond_jne(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 != 3 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 -1) / (linear r0 -1)")
+__msg_next(" scev at latch 2:")
+__msg_next("  r0=(+ r0 -1) / (linear (+ r0 -1) -1)")
+__msg("loop header at 1, header_count is 3 ")
+__msg("loop header at 1, widening r0 to 1..3 step 1")
+__msg("1: R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,var_off=(0x0; 0x3)) loop_stack=1")
+__msg("loop header at 1, clamping r0 to 1..2 step 1")
+__msg("processed 5 insns")
+__naked void post_cond_jne_neg_step(void)
+{
+	asm volatile ("					\
+	r0 = 3;						\
+1:	r0 += -1;					\
+	if r0 != 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next(" scev at latch 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg("loop header at 1, widening r0")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("1: (15) if r0 == 0x3 goto pc+2        ; R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,{{.*}})")
+__msg("2: (07) r0 += 1                       ; R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("3: (05) goto pc-3")
+__msg("loop header at 1, clamping r0")
+__msg("1: safe")
+__msg("from 1 to 4: R0=3")
+__msg("4: R0=3")
+__msg("4: (95) exit")
+__msg("processed 6 insns")
+__naked void pre_cond_je1(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	if r0 == 3 goto 2f;				\
+	r0 += 1;					\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is [0..3] ")
+__naked void one_backedge_two_exits(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	call %[bpf_get_prandom_u32];			\
+	if r0 == 0 goto 2f;				\
+	r6 += 1;					\
+	if r6 != 3 goto 1b;				\
+2:	r0 = r6;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 11:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next("  r1=(+ r1 2) / (linear r1 2)")
+__msg_next(" scev at latch 16:")
+__msg_next("  r0=(+ r0 1) / (linear (+ r0 1) 1)")
+__msg_next("  r1=(+ r1 2) / (linear (+ r1 2) 2)")
+__msg("loop header at 11, widening r0")
+__msg("loop header at 11, widening r1")
+__msg("11: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=7,var_off=(0x0; 0x7)) R1=scalar(smin=smin32=0,smax=umax=smax32=umax32=14,var_off=(0x0; 0xe),step=0+2)")
+/* loop exit */
+__msg("16: (a5) if r0 < 0x8 goto pc-6")
+__msg("exiting loop 11")
+__msg("17: (95) exit")
+/* second iteration */
+__msg("loop header at 11, clamping r0")
+__msg("loop header at 11, clamping r1")
+/* iteration convergence */
+__msg("from 16 to 11: safe")
+__not_msg("{{^}}11:")
+/* map lookup error path */
+__msg("from 7 to 17: safe")
+__naked void correlated_regs(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	*(u64*)(r10 - 8) = r1;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 2f;				\
+	r6 = r0;					\
+	r0 = 0;						\
+	r1 = 0;						\
+1:	r2 = r6;					\
+	r2 += r1;					\
+	*(u8 *)(r2 + 0) = 1;				\
+	r0 += 1;					\
+	r1 += 2;					\
+	if r0 < 8 goto 1b;				\
+2:	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+/*
+ * k = 0
+ * for (i = 0; i < 4; i++):
+ *   for (j = 0; j < 4; j++):
+ *     k += 1
+ *     k <<= 1   // make SCEV construction not possible
+ *     k >>= 1
+ * map[k] = 1    // make k precise
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("19: (72) *(u8 *)(r4 +0) = 1           ; R4=map_value(id={{.*}},map=map,ks=4,vs=1024,imm=16)")
+__not_msg("19: ")
+__msg("processed 106 insns")
+__naked void nested_loops_precise_var1(void)
+{
+	asm volatile ("					\
+	*(u64*)(r10 - 8) = 0;				\
+	r1 = %[map] ll;					\
+	r2 = r10;					\
+	r2 += -8;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 3f;				\
+	r1 = 0;						\
+	r3 = 0;						\
+	/* outer loop */				\
+1:	r2 = 0;						\
+	/* inner loop */				\
+2:	r2 += 1;					\
+	r3 += 1;					\
+	r3 <<= 1;					\
+	r3 >>= 1;					\
+	if r2 < 4 goto 2b;				\
+	r1 += 1;					\
+	if r1 < 4 goto 1b;				\
+	r4 = r0;					\
+	r4 += r3;					\
+	*(u8 *)(r4 + 0) = 1;				\
+	r0 = 0;						\
+3:	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop header at 2, header_count is 100 ")
+__msg("loop header at 4, header_count is [0..100] ")
+__msg("loop header at 6, header_count is [0..100] ")
+__msg("processed 20 insns")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void nested_loop_with_two_exits(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = 0;						\
+1:	r6 += 1;					\
+	r7 = 0;						\
+2:	r7 += 1;					\
+	r8 = 0;						\
+3:	r8 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 42 goto +1;				\
+	goto 4f;					\
+	if r8 < 100 goto 3b;				\
+	if r7 < 100 goto 2b;				\
+4:	if r6 < 100 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__naked void exit_loop_into_loop_header(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	r2 = 0;						\
+loop_a_%=:						\
+	r1 += 1;					\
+	if r1 < 10 goto loop_a_%=;			\
+loop_b_%=:						\
+	r2 += 1;					\
+	if r2 < 10 goto loop_b_%=;			\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * This exercises verifier.c:loop_stack_{pop,push}() implementation,
+ * at 'goto d' loops 'b' and 'a' have to be popped from stack,
+ * while loops 'c' and 'd' have to be pushed to stack.
+ *
+ *   loop a:                  // header 5
+ *     loop b:                // header 6
+ *       if (rand) goto d;    // 8 -> 12, side entry into inner loop d
+ *       ...
+ *   loop c:                  // header 11
+ *     loop d:                // header 12
+ *       ...
+ */
+SEC("xdp")
+__log_level(2)
+__msg("loop at 5")
+__msg("loop at 6, nested in 5")
+__msg("loop at 11, irreducible")
+__msg("loop at 12, nested in 11")
+/* entry via if r0 == 5 goto d_%= false branch */
+__msg("loop header at 12, header_count is 3 ")
+__msg("loop header at 12, widening r9 to 0..2 step 1")
+__msg("12: R8=1 R9=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,var_off=(0x0; 0x3)) loop_stack=11,12")
+/* entry via if r0 == 5 goto d_%= true branch */
+__msg("loop header at 12, header_count is 3 ")
+__msg("loop header at 12, widening r9 to 0..2 step 1")
+__msg("from 8 to 12: R8=0 R9=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,var_off=(0x0; 0x3)) R10=fp0 loop_stack=11,12")
+__naked void enter_nested_loop_from_side(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = 0;						\
+	r7 = 0;						\
+	r8 = 0;						\
+	r9 = 0;						\
+a_%=:	r6 += 1;					\
+b_%=:	r7 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 5 goto d_%=;				\
+	if r7 < 3 goto b_%=;				\
+	if r6 < 3 goto a_%=;				\
+c_%=:	r8 += 1;					\
+d_%=:	r9 += 1;					\
+	if r9 < 3 goto d_%=;				\
+	if r8 < 3 goto c_%=;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Loop A (counter in r6) contains irreducible loop B (header 5),
+ * which contains C (counter in r9).
+ * Entering B at 'body' saves and restores r6, making it appear invariant.
+ * Entering B at 'alternate' skips the save and modifies r6 instead.
+ * Hence A must not infer a SCEV expression for r6.
+ * SCEV expression for r9 in C should still be computed.
+ *
+ *  0: r6 = 0;
+ *     do {                              // A
+ *  1:     r6++;
+ *  2:     r0 = bpf_get_prandom_u32();
+ *  3:     r7 = 0;
+ *  4:     if (r0 > 5) goto alternate;
+ *  5: B:  r8 = r6;                      // B
+ *  6:     goto body;
+ *  7: alternate:
+ *         r8 = r6;
+ *  8:     r8++;
+ *  9: body:
+ *         r6 = r8;
+ * 10:     r9 = 0;
+ *         do {                          // C
+ * 11:         r9++;
+ * 12:     } while (r9 < 3);
+ * 13:     r7++;
+ * 14:     if (r7 < 4) goto B;
+ * 15: } while (r6 < 4);
+ * 16: r0 = 0;
+ * 17: return r0;
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 5, nested in 1, irreducible")
+__msg("loop at 11, nested in 5")
+__msg("scev at header 1:")
+__msg_next("  r6=?")
+__msg("scev at header 11:")
+__msg_next("  r9=(+ r9 1) / (linear r9 1)")
+__msg("loop header at 11, widening r9 to 0..2 step 1")
+__not_msg("loop header at 1, widening r6")
+__naked void nested_irreducible_loop(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+a_%=:	r6 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = 0;						\
+	if r0 > 5 goto alternate_%=;			\
+b_%=:	r8 = r6;					\
+	goto body_%=;					\
+alternate_%=:						\
+	r8 = r6;					\
+	r8 += 1;					\
+body_%=:						\
+	r6 = r8;					\
+	r9 = 0;						\
+c_%=:	r9 += 1;					\
+	if r9 < 3 goto c_%=;				\
+	r7 += 1;					\
+	if r7 < 4 goto b_%=;				\
+	if r6 < 4 goto a_%=;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Induction variable seeded from a non-constant value. r7 enters the loop as a
+ * range aligned to 2 (prandom & 0x6 -> {0,2,4,6}) and is incremented by a
+ * non-power-of-2 slope of 6. Since the entry value is not a single point, only
+ * the power-of-two alignment shared by the entry value and the slope can be
+ * guaranteed, so the widened step is 2.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 4, widening r7 to 0..18 step 2")
+__msg("R7=scalar(smin=smin32=0,smax=umax=smax32=umax32=18,var_off=(0x0; 0x1e),step=0+2)")
+__naked void widen_nonconst_base(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	r7 &= 0x6;					\
+	r6 = 0;						\
+1:	r7 += 6;					\
+	r6 += 1;					\
+	if r6 < 3 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * After excluding zero from {0,4,8,12}, the interval starts at 1 while the
+ * values remain multiples of 4. Widening over three iterations must preserve
+ * base 0 and include 20.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 5, header_count is 3 ")
+__msg("5: R7=scalar({{.*}}smax=umax=smax32=umax32=20,{{.*}}step=0+4)")
+__naked void widen_nonconst_base_refined(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	r7 &= 0xc;					\
+	if r7 < 1 goto 2f;				\
+	r8 = 0;						\
+1:	r7 += 4;					\
+	r8 += 1;					\
+	if r8 < 3 goto 1b;				\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Two nested loops with (1) exiting directly to (2):
+ *
+ *   for (r6 = 0; r6 < 4; r6++) {
+ *     r7 = 0;
+ *     for (; r7 <  3; r7++) {}   // (1)
+ *     for (; r7 != 0; r7--) {}   // (2)
+ *   }
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop header at 1, widening r6 to 0..3 step 1")
+__msg("loop header at 2, widening r7 to 0..2 step 1")
+__msg("exiting loop 2")
+__msg("entering loop 4")
+__msg("loop header at 4, widening r7 to 1..3 step 1")
+__naked void sibling_inner_loops(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	r7 = 0;						\
+2:	r7 += 1;					\
+	if r7 < 3 goto 2b;				\
+3:	r7 += -1;					\
+	if r7 != 0 goto 3b;				\
+	r6 += 1;					\
+	if r6 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop header at 0, can't compute iterations count")
+__naked void uninit_slot_counter(void)
+{
+	asm volatile ("					\
+1:	r0 = *(u64 *)(r10 - 8);				\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	if r0 < 10 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The assignment happens only on the second iteration:
+ *
+ *   r7 = 5;
+ *   for (r6 = 0; r6 < 3; r6++)
+ *           if (r6 == 1)
+ *                   r7 = 10;
+ *   if (r7 != 10)
+ *           invalid_stack_read();
+ *
+ * R7 is always 10 at the real exit. Widening loses the correlation between
+ * R6 and R7, leaving R7 in [5, 10] both at the loop header and after the loop.
+ * The verifier therefore rejects the possible invalid stack read.
+ */
+SEC("xdp")
+__failure
+__log_level(2)
+__msg("2: {{.*}}R7=scalar(smin=umin=smin32=umin32=5,smax=umax=smax32=umax32=10,var_off=(0x0; 0xf))")
+__msg("7: R7=scalar(smin=umin=smin32=umin32=5,smax=umax=smax32=umax32=10,var_off=(0x0; 0xf))")
+__msg("invalid read from stack R10 off=0 size=8")
+__naked void conditional_assignment_on_second_iteration(void)
+{
+	asm volatile ("					\
+	r7 = 5;						\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = 10;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	if r7 == 10 goto 4f;				\
+	r0 = *(u64 *)(r10 + 0);				\
+4:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The latch bound has header SCEV (any r7 100), so it is not invariant:
+ *
+ *   u64 i = 0, n = 3;
+ *   for (;;) {
+ *           ++i;
+ *           if (i >= n)
+ *                   break;
+ *           if (i == 1)
+ *                   n = 100;
+ *   }
+ *
+ * The ANY expression must also appear at the latch, rather than a bare r7
+ * that would incorrectly be treated as a loop-invariant bound.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__msg("scev at header 2:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=(any r7 100) / (any r7 100)")
+__msg_next(" scev at latch 3:")
+__msg_next("  r6=(+ r6 1) / (linear (+ r6 1) 1)")
+__msg_next("  r7=r7 / (any r7 100)")
+__naked void no_widen_changing_latch_bound(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+	r7 = 3;						\
+loop_%=:						\
+	r6 += 1;					\
+	if r6 >= r7 goto exit_%=;			\
+	if r6 != 1 goto next_%=;			\
+	r7 = 100;					\
+next_%=:						\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Join two stack pointer offsets via (any r7 r8), with r8 loop-invariant.
+ * The joined range [-16, -8] still points to initialized stack memory,
+ * so dereferencing r7 after the loop is safe.
+ */
+SEC("xdp")
+__success __retval(0)
+__log_level(2)
+__msg("loop header at 7, widening r7 to -16..-8 step 1")
+__msg("7: {{.*}}R7=fp(smin=smin32=-16,smax=smax32=-8,")
+__msg("12: R7=fp(smin=smin32=-16,smax=smax32=-8,")
+__naked void conditional_stack_pointer_assignment(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 16) = 0;				\
+	*(u64 *)(r10 - 8) = 0;				\
+	r7 = r10;					\
+	r7 += -16;					\
+	r8 = r10;					\
+	r8 += -8;					\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = r8;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	r0 = *(u64 *)(r7 + 0);				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * r7 starts as 1 + 6*k and invariant r8 as 1 + 9*k (0 <= k <= 3).
+ * The (any r7 r8) union must preserve base 1 and step gcd(6, 9) = 3.
+ */
+SEC("xdp")
+__success __retval(1)
+__log_level(2)
+__msg("r7 += 1 {{.*}}step=1+6)")
+__msg("r8 += 1 {{.*}}step=1+9)")
+__msg("loop header at 9, widening r7 to 1..28 step 3")
+__msg("9: {{.*}}R7=scalar({{.*}},step=1+3)")
+__msg("14: R7=scalar({{.*}},step=1+3)")
+__naked void conditional_scalar_assignment_gcd(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r7 = r0;					\
+	r7 *= 6;					\
+	r7 += 1;					\
+	r8 = r0;					\
+	r8 *= 9;					\
+	r8 += 1;					\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = r8;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	r0 = r7;					\
+	r0 %%= 3;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * r7 = ctx->data;
+ * r8 = r7
+ * for (r6 = 0; r6 < 10 && random() != 42; r7++, r6++);
+ * if (r7 >= ctx->data_end)
+ *   return;
+ * *(r8 + 4);  // At this point the loop executed unknown number of times.
+ *	       // Hence, r7 range gives no information about r8.
+ */
+SEC("tc")
+__failure
+__msg("R8 min value is outside of the allowed memory range")
+__naked void break_pkt_pointers_id(void)
+{
+	asm volatile ("					\
+	r7 = *(u32*)(r1 + %[__sk_buff_data]);		\
+	r8 = r7;					\
+	r9 = *(u32*)(r1 + %[__sk_buff_data_end]);	\
+	r6 = 0;						\
+1:	call %[bpf_get_prandom_u32];			\
+	if r0 == 42 goto 3f;				\
+	r6 += 1;					\
+	r7 += 1;					\
+	if r6 < 10 goto 1b;				\
+3:	if r7 >= r9 goto 2f;				\
+	r0 = *(u8*)(r8 + 4);				\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm_const(__sk_buff_data, offsetof(struct __sk_buff, data)),
+	  __imm_const(__sk_buff_data_end, offsetof(struct __sk_buff, data_end)),
+	  __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * A loop induction variable used to compute the base address of a store to the
+ * stack must not be widened: the spill offset would become varying, which the
+ * verifier does not track.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 2, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__naked void no_widen_stack_spill(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u64 *)(r3 + 0) = r0;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+/*
+ * Same hazard across a loop nest: the outer induction variable r2 addresses a
+ * stack store performed inside the inner loop. The dependency is pulled up from
+ * the inner loop, so the outer loop must not widen r2.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 2, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__msg("loop header at 3, widening r1 to 0..1 step 1")
+__naked void no_widen_stack_spill_nested(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r1 = 0;						\
+2:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u64 *)(r3 + 0) = r1;				\
+	r1 += 1;					\
+	if r1 < 2 goto 2b;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The inner loop restores r6 on its backedge but exits with r6 = -1000 on break.
+ * Outer loop must account for the exit value instead of deriving r6 = r6_entry + n.
+ * On the first outer backedge r6 is -999, so the next map access is invalid.
+ *
+ *   u8 *r7 = map_value;
+ *   s64 r6 = 0;
+ *   do {
+ *   1:       r9 = 0;
+ *           r0 = r7[r6];
+ *           while (true) {
+ *   2:              r8 = r6;
+ *                   r6 = -1000;
+ *                   if (++r9 > 5)
+ *                           break;
+ *                   r6 = r8;
+ *   3:      }
+ *           r6++;
+ *   } while (r6 < 3);
+ */
+SEC("xdp")
+__failure
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev at header 9:")
+__msg_next("  r6=(+ ? 1) / ?")
+__msg_next(" scev at latch 20:")
+__msg_next("  r6=(+ ? 1) / (+ ? 1)")
+__msg_next("scev at header 13:")
+__msg_next("  r9=(+ r9 1) / (linear r9 1)")
+__msg_next(" scev at latch 16:")
+__msg_next("  r6=-1000 / -1000")
+__msg_next("  r8=r6 / r6")
+__msg_next("  r9=(+ r9 1) / (linear (+ r9 1) 1)")
+__msg("R6=-999")
+__msg("R1 min value is negative")
+__naked void nested_loop_exit_clobbers_reg(void)
+{
+	asm volatile (
+	"*(u64 *)(r10 - 8) = 0;"
+	"r2 = r10;"
+	"r2 += -8;"
+	"r1 = %[map] ll;"
+	"call %[bpf_map_lookup_elem];"
+	"if r0 == 0 goto 4f;"
+	"r7 = r0;"
+	"r6 = 0;"
+"1:"
+	"r9 = 0;"
+	"r1 = r7;"
+	"r1 += r6;"
+	"r0 = *(u8 *)(r1 + 0);"
+"2:"
+	"r8 = r6;"
+	"r6 = -1000;"
+	"r9 += 1;"
+	"if r9 > 5 goto 3f;"
+	"r6 = r8;"
+	"goto 2b;"
+"3:"
+	"r6 += 1;"
+	"if r6 s< 3 goto 1b;"
+"4:"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+/*
+ * The inner loop exits both nested loops with r6 unchanged, or exits only itself with r6 = -1000.
+ * The middle loop repairs the latter exit with a stack fill. It therefore preserves r6,
+ * so the outermost loop must retain r6_entry + n.
+ *
+ *   r6 = 0;
+ *   do {
+ *           r7 = 0;
+ *           do {
+ *                   spill = r6;
+ *                   r9 = 0;
+ *                   do {
+ *                           r8 = r6;
+ *                           if (random() & 1)
+ *                                   goto next;
+ *                           r6 = -1000;
+ *                           if (++r9 >= 2)
+ *                                   break;
+ *                           r6 = r8;
+ *                   } while (true);
+ *                   r6 = spill;
+ *           } while (++r7 < 2);
+ *   next:
+ *           r6++;
+ *   } while (r6 < 3);
+ */
+SEC("xdp")
+__success __retval(3)
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("r6 = *(u64 *)(r10 -8) {{.*}}; r6 ? -> r6")
+__msg("scev at header 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next(" scev at latch 16:")
+__msg_next("  r6=(+ r6 1) / (linear (+ r6 1) 1)")
+__msg("scev at header 2:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg("scev at header 4:")
+__msg_next("  r9=(+ r9 1) / (linear r9 1)")
+__msg(" scev at latch 9:")
+__msg_next("  r8=r6 / r6")
+__msg_next("  r9=(+ r9 1) / (linear (+ r9 1) 1)")
+__naked void nested_loop_exit_preserves_reg(void)
+{
+	asm volatile (
+	"r6 = 0;"
+"1:"
+	"r7 = 0;"
+"2:"
+	"*(u64 *)(r10 - 8) = r6;"
+	"r9 = 0;"
+"3:"
+	"r8 = r6;"
+	"call %[bpf_get_prandom_u32];"
+	"if r0 & 1 goto 5f;"
+	"r6 = -1000;"
+	"r9 += 1;"
+	"if r9 >= 2 goto 4f;"
+	"r6 = r8;"
+	"goto 3b;"
+"4:"
+	"r6 = *(u64 *)(r10 - 8);"
+	"r7 += 1;"
+	"if r7 < 2 goto 2b;"
+"5:"
+	"r6 += 1;"
+	"if r6 < 3 goto 1b;"
+	"r0 = r6;"
+	"exit;"
+	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Complement of no_widen_stack_spill: a sub-register (1-byte) store to the stack
+ * lands as STACK_MISC and carries no tracked value, so the induction variable
+ * addressing it (r2) is still widened.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("widening r2 to 0..24 step 8")
+__naked void widen_byte_stack_store(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u8 *)(r3 + 0) = r0;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * A fill (BPF_LDX) at a varying stack offset loses precision just like a spill,
+ * so the induction variable computing the load base (r2) must not be widened.
+ * The slots are initialized up front so the fill itself is a valid read.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 6, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__naked void no_widen_stack_fill(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 64) = r0;			\
+	*(u64 *)(r10 - 56) = r0;			\
+	*(u64 *)(r10 - 48) = r0;			\
+	*(u64 *)(r10 - 40) = r0;			\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	r4 = *(u64 *)(r3 + 0);				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * struct xdp_md *slots[2] = { ctx, ctx };
+ *
+ * r7 = 0;
+ * for (r8 = 0; r8 < 2; r8++) {
+ *         value = slots[r7]->data;
+ *         if (i == 0)
+ *                 r7 = 1;
+ * }
+ *
+ * SCEV expression for r7 is (any r7 8) and r7 is used to
+ * address the stack memory. Avoid widening it, otherwise
+ * verifier won't know the type of the slots[r7] expression.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev at header 4:")
+__msg_next("r7=(any r7 8) / (any r7 8)")
+__msg("loop header at 4, can't widen r7, expr is {{.*}}, requires exact stack-offset tracking")
+__naked void no_widen_stack_fill_any(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 16) = r1;			\
+	*(u64 *)(r10 - 8) = r1;				\
+	r7 = 0;						\
+	r8 = 0;						\
+1:	r2 = r10;					\
+	r2 += -16;					\
+	r2 += r7;					\
+	r3 = *(u64 *)(r2 + 0);	/* slots[r7] */		\
+	r0 = *(u32 *)(r3 + 0);	/* slots[r7]->data */	\
+	if r8 != 0 goto 2f;				\
+	r7 = 8;			/* conditionally update r7 */ \
+2:	r8 += 1;					\
+	if r8 < 2 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * A dynptr/iter/irq/res_spin_lock call initializes a stack object through a
+ * pointer argument, which acts like a spill base: the induction variable
+ * computing that argument's varying stack offset must not be widened, otherwise
+ * the slot can't be resolved. Here each iteration constructs an xdp dynptr at
+ * &dptrs[i].
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("can't widen {{.*}}, requires exact stack-offset tracking")
+#ifndef __clang__
+/* The issue is unrelated to SCEV */
+__skip("GCC emits a stack-pointer loop condition the verifier cannot resolve")
+#endif
+int no_widen_dynptr_kfunc_arg(struct xdp_md *ctx)
+{
+	struct bpf_dynptr dptrs[4];
+	int i;
+
+#ifdef __clang__
+#pragma clang loop unroll(disable)
+#else
+#pragma GCC unroll 0
+#endif
+	for (i = 0; i < 4; i++)
+		bpf_dynptr_from_xdp(ctx, 0, &dptrs[i]);
+
+	return 0;
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("header_count is 4 ")
+__msg("widening r6 to -16..-13 step 1")
+__naked void ptr_stack(void)
+{
+	asm volatile ("					\
+	r6 = r10;					\
+	r6 += -16;					\
+	r7 = r6;					\
+	r7 += 4;					\
+1:	r6 += 1;					\
+	if r6 < r7 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+__naked __noinline __used static int ptr_stack_callee(void)
+{
+	asm volatile ("					\
+	r6 = r10;					\
+	r6 += -16;					\
+	r7 = r1;					\
+	r0 = 0;						\
+1:	r6 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto 4f;	/* bound main pass iteration, but hide it from SCEV */	\
+	if r6 < r7 goto 1b;	/* r6 and r7 are from different frames, can't be widened */	\
+4:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* The end pointer belongs to the caller's frame. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r6 (fp), r7 (fp)")
+__not_msg("widening r6")
+__naked void ptr_stack_other_frame(void)
+{
+	asm volatile ("					\
+	r1 = r10;					\
+	r1 += -12;					\
+	call ptr_stack_callee;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Metadata and packet data have different origins. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r6 (pkt_meta), r7 (pkt)")
+__not_msg("widening r6")
+__naked void ptr_packet_meta_other_origin(void)
+{
+	asm volatile ("					\
+	r6 = *(u32 *)(r1 + %[data_meta]);		\
+	r7 = *(u32 *)(r1 + %[data]);			\
+	r7 += 4;					\
+	r0 = 0;						\
+1:	r6 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto 4f;				\
+	if r6 < r7 goto 1b;				\
+4:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm_const(data, offsetof(struct xdp_md, data)),
+	  __imm_const(data_meta, offsetof(struct xdp_md, data_meta))
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("header_count is 4 ")
+__msg("widening r6 to 0..3 step 1")
+__naked void ptr_map_value(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 8) = 0;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 2f;				\
+	r6 = r0;					\
+	r7 = r6;					\
+	r7 += 4;					\
+1:	r6 += 1;					\
+	if r6 < r7 goto 1b;				\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem), __imm_addr(map)
+	: __clobber_all);
+}
+
+/* Equal offsets into different maps do not establish a common origin. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r6 (map_value), r7 (map_value)")
+__not_msg("widening r6")
+__naked void ptr_map_value_other_map(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 8) = 0;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 2f;				\
+	r6 = r0;					\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[other_map] ll;				\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 2f;				\
+	r7 = r0;					\
+	r7 += 4;					\
+	r0 = 0;						\
+1:	r6 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto 4f;				\
+	if r6 < r7 goto 1b;				\
+4:							\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem), __imm_addr(map), __imm_addr(other_map)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("header_count is 4 ")
+__msg("widening r6 to 0..3 step 1")
+__naked void ptr_mem(void)
+{
+	asm volatile ("					\
+	r1 = %[ringbuf] ll;				\
+	r2 = 8;						\
+	r3 = 0;						\
+	call %[bpf_ringbuf_reserve];			\
+	if r0 == 0 goto 2f;				\
+	r8 = r0;					\
+	r6 = r0;					\
+	r7 = r6;					\
+	r7 += 4;					\
+1:	r6 += 1;					\
+	if r6 < r7 goto 1b;				\
+	r1 = r8;					\
+	r2 = 0;						\
+	call %[bpf_ringbuf_discard];			\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard), __imm_addr(ringbuf)
+	: __clobber_all);
+}
+
+/* Separate reservations have different pointer IDs. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r6 (ringbuf_mem), r7 (ringbuf_mem)")
+__not_msg("widening r6")
+__naked void ptr_mem_other_reservation(void)
+{
+	asm volatile ("					\
+	r1 = %[ringbuf] ll;				\
+	r2 = 8;						\
+	r3 = 0;						\
+	call %[bpf_ringbuf_reserve];			\
+	if r0 == 0 goto 3f;				\
+	r8 = r0;					\
+	r6 = r0;					\
+	r1 = %[ringbuf] ll;				\
+	r2 = 8;						\
+	r3 = 0;						\
+	call %[bpf_ringbuf_reserve];			\
+	if r0 == 0 goto 2f;				\
+	r9 = r0;					\
+	r7 = r0;					\
+	r7 += 4;					\
+	r0 = 0;						\
+1:	r6 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto 4f;				\
+	if r6 < r7 goto 1b;				\
+4:	r1 = r9;					\
+	r2 = 0;						\
+	call %[bpf_ringbuf_discard];			\
+2:	r1 = r8;					\
+	r2 = 0;						\
+	call %[bpf_ringbuf_discard];			\
+3:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard), __imm_addr(ringbuf)
+	: __clobber_all);
+}
+
+SEC("iter/bpf_map_elem")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("header_count is 4 ")
+__msg("widening r6 to 0..3 step 1")
+__naked void ptr_buf(void)
+{
+	asm volatile ("					\
+	r6 = *(u64 *)(r1 + %[value]);			\
+	if r6 == 0 goto 2f;				\
+	r7 = r6;					\
+	r7 += 4;					\
+1:	r6 += 1;					\
+	if r6 < r7 goto 1b;				\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm_const(value, offsetof(struct bpf_iter__bpf_map_elem, value))
+	: __clobber_all);
+}
+
+/* Separate context loads receive different pointer IDs. */
+SEC("iter/bpf_map_elem")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r6 (buf), r7 (buf)")
+__not_msg("widening r6")
+__naked void ptr_buf_other_load(void)
+{
+	asm volatile ("					\
+	r6 = *(u64 *)(r1 + %[value]);			\
+	if r6 == 0 goto 2f;				\
+	r7 = *(u64 *)(r1 + %[value]);			\
+	if r7 == 0 goto 2f;				\
+	r7 += 4;					\
+	r0 = 0;						\
+1:	r6 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto 4f;				\
+	if r6 < r7 goto 1b;				\
+4:							\
+2:	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm_const(value, offsetof(struct bpf_iter__bpf_map_elem, value))
+	: __clobber_all);
+}
+
+/* A pointer-versus-scalar latch must not establish an iteration count. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("scev: incompatible latch operands r1 (map_value), scalar 4")
+__not_msg("widening r1")
+__naked void ptr_vs_scalar_no_widen(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 8) = 0;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto out_%=;				\
+	r1 = r0;					\
+	r0 = 0;						\
+loop_%=:						\
+	r1 += 1;					\
+	r0 += 1;					\
+	if r0 & 8 goto out_%=;				\
+	if r1 < 4 goto loop_%=;				\
+out_%=:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+/* A nullable pointer must not be widened before its null check. */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 9, can't widen r1, expr is (+ r1 8)")
+__naked void no_widen_maybe_null(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	*(u64 *)(r10 - 8) = r1;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	r1 = r0;					\
+	r6 = 0;						\
+loop_%=:						\
+	if r1 == 0 goto exit_%=;			\
+	r1 += 8;					\
+	r6 += 1;					\
+	if r6 < 3 goto loop_%=;				\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+/*
+ * A terminating inner loop does not prove termination of the outer loop.
+ *
+ *    u64 r7 = 0;
+ *    while (bpf_get_prandom_u32() != 42) {
+ *        r7 = (r7 + 1) & 0xf;
+ *        u64 r8 = 0;
+ *        do {
+ *            r8++;
+ *        } while (r8 < 3);
+ *    }
+ *    return 0;
+ */
+SEC("xdp")
+__failure
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("infinite loop detected at insn 6")
+__naked void nested_terminating_prune(void)
+{
+	asm volatile ("					\
+	r7 = 0;						\
+outer_%=:						\
+	r7 += 1;					\
+	r7 &= 0xf;					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 42 goto out_%=;			\
+	r8 = 0;						\
+inner_%=:						\
+	r8 += 1;					\
+	if r8 < 3 goto inner_%=;			\
+	goto outer_%=;					\
+out_%=:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Returning directly from the inner loop to the outer header changes the
+ * active stack from [1, 4] to [1].
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop at 1{{$}}")
+__msg("loop at 4, nested in 1")
+__msg("entering loop 1")
+__msg("entering loop 4")
+__msg("exiting loop 4")
+__not_msg("entering loop 1")
+__msg("from 4 to 1:")
+__msg("loop_stack=1{{$}}")
+__naked void nested_loop_backedge_to_outer_header(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+outer_%=:						\
+	if r6 == 3 goto exit_%=;			\
+	r6 += 1;					\
+	r7 = 0;						\
+inner_%=:						\
+	if r7 == 2 goto outer_%=;			\
+	r7 += 1;					\
+	goto inner_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.53.0

^ permalink raw reply related	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists
  2026-10-04 13:38 ` [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
@ 2026-10-04 13:47   ` sashiko-bot
  2026-10-06 16:53   ` Alexei Starovoitov
  1 sibling, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 13:47 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: bpf_min_heap backing array uses krealloc leading to -ENOMEM
- [Low] bpf: custom min-heap reinventing generic include/linux/min_heap.h
--

commit 962cfb7ee3152595af40dbe59bd87156cd6e55d1
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: add a min-heap for ordered analysis worklists

This commit adds a small integer min-heap implementation for the BPF
verifier. It orders instruction indices by their CFG ranks for SCEV basic
block processing, growing a backing array on demand.

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
[ ... ]
> +/*
> + * Simple binary heap implementation as described by
> + * https://en.wikipedia.org/wiki/Binary_heap
> + */
> +struct bpf_min_heap {
> +	int (*compare)(int, int, void *); /* ordering function for @elements */
> +	int *elements; /* min-heap ordered by @compare */
> +	void *arg; /* 3rd argument passed to @compare */
> +	int capacity;
> +	int count;
> +};

[Severity: Low]
Does this custom structure reinvent the generic min-heap implementation already
available in include/linux/min_heap.h?

The kernel already provides a well-tested API. If dynamic reallocation is
necessary, could it be wrapped around the existing struct min_heap and its
callbacks rather than completely duplicating the core logic?

> +
> +void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg);
> +void bpf_min_heap_free(struct bpf_min_heap *heap);
> +int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
> +bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
> +
>  #endif /* _LINUX_BPF_VERIFIER_H */
> diff --git a/kernel/bpf/heap.c b/kernel/bpf/heap.c
> new file mode 100644
> --- /dev/null
> +++ b/kernel/bpf/heap.c
[ ... ]
> +void bpf_min_heap_free(struct bpf_min_heap *heap)
> +{
> +	kfree(heap->elements);

[Severity: Medium]
If the allocation in bpf_min_heap_push() below is changed to use kvrealloc(),
does this need to be updated to kvfree() to avoid crashes?

> +	heap->elements = NULL;
> +	heap->capacity = 0;
> +	heap->count = 0;
> +}
> +
> +int bpf_min_heap_push(struct bpf_min_heap *heap, int elt)
> +{
> +	int new_capacity, i;
> +	int *elements;
> +	void *tmp;
> +
> +	if (heap->count == heap->capacity) {
> +		new_capacity = heap->capacity ? heap->capacity * 2 : 16;
> +		tmp = krealloc(heap->elements,
> +			       sizeof(*heap->elements) * new_capacity,
> +			       GFP_KERNEL_ACCOUNT);

[Severity: Medium]
Could this krealloc() fail with -ENOMEM when loading large BPF programs on
fragmented systems?

Since the backing array scales up proportionally to the number of basic
blocks, this attempts to allocate physically contiguous memory. This can
frequently fail for high-order allocations when dealing with complex
programs that push the capacity to several megabytes.

Would it be more reliable to use kvrealloc() here to match the scalability
of the other SCEV verifier data structures?

> +		if (!tmp)
> +			return -ENOMEM;
> +		heap->elements = tmp;
> +		heap->capacity = new_capacity;
> +	}
[ ... ]
> +static inline void sink_root(struct bpf_min_heap *heap)
> +{
> +	int *elements = heap->elements;
> +	int i = 0;
> +
> +	while ((left_child(i)  < heap->count && greater(heap, elements[i], elements[left_child(i)])) ||
> +	       (right_child(i) < heap->count && greater(heap, elements[i], elements[right_child(i)]))) {

[Severity: Low]
Does this function unnecessarily duplicate the sift-down logic already
provided by min_heap_sift_down() in the core kernel API?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=23

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 18/43] bpf: compute immediate dominators
  2026-10-04 13:38 ` [PATCH bpf-next v2 18/43] bpf: compute immediate dominators Eduard Zingerman
@ 2026-10-04 13:50   ` sashiko-bot
  0 siblings, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 13:50 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] bpf: integer overflow in compute_predecessors size calculation
--

commit ce110737451057ca4a576ab4cd5332b75722cc64
Author: Eduard Zingerman <eddyz87@gmail.com>
bpf: compute immediate dominators

This commit introduces a new immediate dominator tree computation for the
BPF verifier. It analyzes each subprogram to find exit conditions that
dominate loop backedges.

> diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
> new file mode 100644
> index 0000000000000..0a9e10edca871
> --- /dev/null
> +++ b/kernel/bpf/loops.c
> @@ -0,0 +1,156 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +
> +#include <linux/slab.h>
> +#include <linux/sched/signal.h>
> +#include <linux/bpf_verifier.h>
> +
> +static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
> +{
> +	struct bpf_iarray *succ, *preds, **result;
> +	struct bpf_prog *prog = env->prog;
> +	u32 *num_preds, i, s, sz, len = prog->len;

[Severity: Critical]
Since sz is typed as a 32-bit integer, could it overflow during the size
calculations below?

> +	struct bpf_insn *insn;
> +	void *tmp;
> +
> +	num_preds = kvcalloc(prog->len, sizeof(u32), GFP_KERNEL_ACCOUNT);
> +	if (!num_preds)
> +		return NULL;
> +
> +	/*
> +	 * 'result' layout:
> +	 *  - array of pointers (struct bpf_iarray *)[len]
> +	 *  - struct bpf_iarray one after another
> +	 */
> +	sz = sizeof(struct bpf_iarray) * len;
> +	sz += sizeof(struct bpf_iarray *) * len;
> +	for (i = 0; i < len; i++) {
> +		insn = env->prog->insnsi + i;
> +		succ = bpf_insn_successors(env, i);
> +		sz += sizeof(u32) * succ->cnt;

[Severity: Critical]
Because sz is accumulated here without overflow checks, can a BPF program
with many gotox instructions sharing a large jump table repeatedly add a large
succ->cnt, causing sz to wrap around?

> +		iarray_for_each(s, succ) {
> +			num_preds[s]++;
> +		}
> +		if (bpf_is_ldimm64(insn))
> +			i++;
> +	}
> +
> +	result = kvzalloc(sz, GFP_KERNEL_ACCOUNT);
> +	if (!result) {
> +		kvfree(num_preds);
> +		return NULL;
> +	}
> +
> +	tmp = (void *)&result[len];
> +	for (i = 0; i < len; i++) {
> +		result[i] = tmp;

[Severity: Critical]
If sz wraps around and a small buffer is allocated, does tmp end up pointing
completely out-of-bounds?

Would this loop and subsequent array accesses then overwrite arbitrary kernel
heap memory, potentially leading to critical memory corruption when an
unprivileged user loads a crafted BPF program?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=18

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning
  2026-10-04 13:38 ` [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
@ 2026-10-04 13:54   ` sashiko-bot
  0 siblings, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 13:54 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: leaked stack pointer on exit in step_prune_hit_multiple()
- [Low] bpf: unused inline assembly operands in step_prune_hit_multiple()
--

commit 5e4048fc2e6297e557be720eceffc54196686ad8
Author: Eduard Zingerman <eddyz87@gmail.com>

selftests/bpf: tests for register base/step state pruning

This commit adds BPF selftests to verify base/step reasoning in
states.c:range_within(). It includes tests for step pruning hits and misses,
including multiple steps and constant boundary scenarios.

> diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> index 079e34820a8f1..f1d8e36b39e35 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
[ ... ]
> +SEC("socket")
> +__success __log_level(2)
> +__msg("7: (27) r6 *= 4                       ; R6=scalar({{.*}}umin32=8,{{.*}}umax32=68,{{.*}},step=0+4)")
> +__msg("10: (27) r7 *= 2                      ; R7=scalar({{.*}}umin32=4,{{.*}}umax32=130,{{.*}},step=0+2)")
> +__msg("11: (25) if r0 > 0x2a goto pc+1")
> +__msg("from 11 to 13: safe")
> +__flag(BPF_F_TEST_STATE_FREQ)
> +__naked void step_prune_hit_multiple(void)
> +{
> +	asm volatile ("					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r7 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 &= 0x0f;					\
> +	r6 += 2;					\
> +	r6 *= 4;					\
> +	r7 &= 0x3f;					\
> +	r7 += 2;					\
> +	r7 *= 2;					\
> +	if r0 > 42 goto 1f;	/* can't predict */	\
> +	r6 = r7;		/* step=2 explored first, step=4 explored next */ \
> +1:	r0 = r10;					\
> +	r6 = -r6;					\
> +	r0 += r6;					\
> +	*(u8 *)(r0 + 0) = 7;	/* force r6 precise */	\
> +	exit;						\

[Severity: Medium]
Does this code leak a stack pointer on exit?

It appears r0 is set to r10 (the stack pointer), modified, and then we hit
the exit instruction without resetting r0 to a scalar. While the verifier
permits this in privileged test code due to allow_ptr_leaks, returning a
stack pointer from a socket filter seems like an unintended test logic
defect.

Could r0 be reset to zero here before exiting, similar to how it is handled
in other tests?

> +"	:
> +	: __imm(bpf_get_prandom_u32),
> +	  __imm(bpf_map_lookup_elem),
> +	  __imm_addr(step_map)

[Severity: Low]
This isn't a bug, but do we need the bpf_map_lookup_elem and step_map
operands in this inline assembly block?

They are declared as input operands but appear to remain entirely unused
within the assembly code for step_prune_hit_multiple. Could they be removed
to simplify the test definition?

> +	: __clobber_all);
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=39

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union()
  2026-10-04 13:38 ` [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union() Eduard Zingerman
@ 2026-10-04 13:59   ` sashiko-bot
  0 siblings, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 13:59 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: division by zero in bpf_reg_union() due to step=0
- [Medium] bpf: bpf_reg_union loses alignment precision by missing optimal common congruence
--

commit f2da61545238e5504547c2ded4955ae49831e34d
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: add bpf_reg_union()

This commit adds a utility function bpf_reg_union() to merge the scalar values
of two registers. It combines the circular 32-bit and 64-bit bounds, tnums, and
base/step equations, synchronizing the result.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 5fd6376da1710..28c16e35b1f36 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -17164,6 +17165,33 @@ int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
>  	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
>  }
>  
> +/* acc := acc U src, matching types only. Caller must clear acc's scalar ID. */
> +int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
> +		  const struct bpf_reg_state *src)
> +{
> +	u16 base, step;
> +
> +	if (acc->type != src->type) {
> +		verifier_bug(env, "union of registers with different types");
> +		return -EFAULT;
> +	}
> +	acc->r64 = cnum64_union(acc->r64, src->r64);
> +	acc->r32 = cnum32_union(acc->r32, src->r32);
> +	acc->var_off = tnum_union(acc->var_off, src->var_off);
> +
> +	/* Retain a common congruence if the bases agree modulo the gcd. */
> +	step = gcd(acc->step, src->step);
> +	base = acc->base % step;

[Severity: Critical]
Does this code risk a division by zero exception and subsequent kernel panic?

BPF programs can create ANY nodes in Scalar Evolution (SCEV) loops. If a loop
combines two loop-invariant pointers (e.g., PTR_TO_MAP_VALUE from different
map lookups) in a phi-node, they are tracked by SCEV (since
is_widenable_reg_type() permits pointers) and passed to bpf_reg_union().

Because pointers are initialized via kzalloc or mark_reg_known_zero(), their
step field is 0. Calling gcd(0, 0) returns 0, leading to a modulo by zero on
this line.

> +	if (base != src->base % step) {
> +		reg_step_reset(acc);

[Severity: Medium]
Does this condition needlessly discard linear alignment information for valid
congruences?

If bpf_reg_union() merges registers with steps 15/base 10 and step 15/base 5,
it evaluates gcd(15,15)=15, sees 10!=5, and resets the step to 1.

The optimal step would be gcd(gcd(acc->step, src->step), abs(acc->base -
src->base)), which would correctly yield step 5/base 0. This unnecessary
precision loss might cause the verifier to incorrectly reject valid programs
due to a perceived loss of alignment constraints.

> +	} else {
> +		acc->base = base;
> +		acc->step = step;
> +	}
> +	reg_bounds_sync(acc);
> +	return reg_bounds_sanity_check(env, acc, "bpf_reg_union");
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=22

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops
  2026-10-04 13:38 ` [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops Eduard Zingerman
@ 2026-10-04 14:00   ` sashiko-bot
  2026-10-04 14:24   ` bot+bpf-ci
  1 sibling, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 14:00 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [Critical] bpf: SCEV analysis overwrites initial state, bypassing bounds checks on early exits
- [High] bpf: NULL pointer dereference in compute_insn_scevs for irreducible nested loops
- [Medium] bpf: log_scevs incorrectly rejects valid programs with nested latches under BPF_LOG_LEVEL2
--

commit 2d2f1f9c2cfacacbaf91e31cf373b1ae40a400ea
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: compute scalar evolution expressions for loops

This commit assigns algebraic expressions to loop variables to describe how
their values evolve across iterations without unrolling the loops. The analysis
proceeds in two phases, analyzing innermost loops first and treating nested
loops as opaque operations.

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> new file mode 100644
> index 0000000000000..9ef28e0d8d481
> --- /dev/null
> +++ b/kernel/bpf/scev.c
[ ... ]
> +static struct env *get_loop_env(struct scev *scev, int loop_header, int insn_idx)
> +{
[ ... ]
> +	e->empty = true;
> +	tmp->entries[cnt].loop_header = loop_header;
> +	tmp->entries[cnt].env = e;
> +	tmp->cnt = cnt + 1;
> +	scev->envs[insn_idx] = tmp;
> +	return e;
> +}
[ ... ]
> +static int setup_initial_loop_env(struct bpf_verifier_env *env, struct env *e, int insn_idx)
> +{
> +	struct scev *scev = env->scev;
> +	int i, err;
> +
> +	for (i = 0; i < REGS_NUM; i++) {
> +		err = replace_reg(scev, e, i, expr1(scev, REG, i));
> +		if (err)
> +			return err;
> +	}
> +	return 0;
> +}

[Severity: Critical]
Could this leave the loop environment marked as empty even after it has been
populated with the initial register state?

When get_loop_env() initializes the environment, it sets e->empty to true.
Since setup_initial_loop_env() does not set e->empty to false after filling
the registers, a subsequent call to join() on the backedge might completely
overwrite this initial state:

	if (acc->empty) {
		memcpy(acc, cur, sizeof(*acc));
		acc->empty = false;
		return 0;
	}

If the loop exits on the first iteration, could the verifier mistakenly
assume a constant from the backedge instead of the actual entry value,
potentially bypassing bounds checks?

[ ... ]
> +static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
> +{
[ ... ]
> +	/*
> +	 * For irreducible loops, and for loops nesting a loop with a truncated
> +	 * exit list, just assume that everything is clobbered for now.
> +	 */
> +	if (!can_compute_loop_scev(cur_loop)) {
> +		header_env = get_loop_env(scev, cur_header, cur_header);
> +		if (!header_env)
> +			return -ENOMEM;
> +		/* The freshly allocated environment has all expressions unknown. */
> +		header_env->empty = false;
> +		return 0;
> +	}
[ ... ]
> +static int compute_insn_scevs(struct bpf_verifier_env *env, struct env *eheader, struct env *einsn)
> +{
> +	struct scev *scev = env->scev;
> +	int id, reg;
> +
> +	for (reg = 0; reg < REGS_NUM; reg++) {
> +		id = einsn->reg2expr[reg];

[Severity: High]
Is it possible for einsn to be NULL here?

If a reducible outer loop has an exit instruction that is physically located
inside a nested irreducible loop, compute_scev_for_loop() will skip
allocating body environments for the inner loop.

When bpf_compute_scev() iterates over instructions and calls find_loop_env()
for this instruction, it might return NULL. This would result in passing a
NULL pointer into compute_insn_scevs(), leading to a dereference here.

> +		id = transform_expr_once(scev, 0, id, eheader->reg2scev, instantiate_header_scevs);
> +		if (id < 0)
> +			return id;
[ ... ]
> +static int log_scevs(struct bpf_verifier_env *env)
> +{
[ ... ]
> +		if (!can_compute_loop_scev(loop))
> +			continue;
> +		for (j = 0; j < loop->backedges_cnt; j++) {
> +			latch = loop->backedges[j].latch;
> +			if (latch < 0)
> +				continue;
> +			bpf_log(log, " scev at latch %d:\n", latch);
> +			latch_env = find_loop_env(scev, i, latch);
> +			if (verifier_bug_if(!latch_env, env, "i=%d, latch=%d", i, latch))
> +				return -EFAULT;

[Severity: Medium]
Could this trigger a verifier bug for valid programs when log level 2 is
enabled?

If an outer loop's latch is located inside a nested loop, its environment
might not be allocated since the nested loop's body is skipped during
evaluation. If so, find_loop_env() will return NULL, causing this check to
fail and abort verification for otherwise valid code.

> +			print_env(env, latch_env, latch, PRINT_SCEV);
> +		}
> +	}
> +	return 0;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=28

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses with arg_access_info
  2026-10-04 13:37 ` [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses " Eduard Zingerman
@ 2026-10-04 14:02   ` sashiko-bot
  0 siblings, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 14:02 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: must_write fails to discard liveness due to may_read mask
--

commit cbeee936583603e6c17220772e7f9ac86e513c50
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: describe helper stack accesses with arg_access_info

Adjust bpf_helper_stack_access_bytes() to return results as a struct
arg_access_info. The commit intends for MEM_WRITE to produce must_write
when allow_uninit_stack is true, allowing the verifier to discard
previous stack knowledge over the affected range.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -14455,18 +14455,21 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
>  		return (struct arg_access_info) {};
>  	default:
> -		return S64_MIN;
> +		return info;
>  	}
>  out:
> -	/*
> -	 * Other accesses keep the previous state live, including untouched bytes
> -	 * of an unprivileged generic output.
> -	 */
> -	if (full_write)
> -		return -size;
> -	return size;
> +	info.size = min_t(u64, size, U32_MAX);
> +	info.may_read = !!(access_type & BPF_READ);
> +	info.may_write = !!(access_type & BPF_WRITE);
> +	info.must_write = info.may_write && exact_size && info.size != U32_MAX;

[Severity: Medium]
When a helper uses MEM_WRITE without MEM_UNINIT, access_type evaluates to
BPF_READ | BPF_WRITE. This logic then sets both may_read and must_write to
true for the argument.

Does this accidentally defeat the intended liveness pruning optimization?

Looking at the liveness propagation equation in kernel/bpf/liveness.c
inside update_insn():

unsigned long new_before = (new_after[w] & ~must_write[w]) | may_read[w];

If may_read is true, the ~must_write clearing operation is overridden. As a
result, the bit is unconditionally set in new_before, and the verifier will
continue to track the state of these stack slots instead of discarding them.

> +	/* Generic unprivileged outputs retain prior initialization state. */
> +	if (!env->allow_uninit_stack && arg_type_is_raw_mem(at)) {
> +		info.may_read = true;
> +		info.must_write = false;
> +	}
> +	return info;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=2

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking
  2026-10-04 13:38 ` [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
@ 2026-10-04 14:02   ` sashiko-bot
  0 siblings, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 14:02 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: call_params_mask() overflows register mask for kfuncs with > 5 args
--

commit 2a8074d95fbbe20d412b57ca9a7f732dfbfaa51d
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: avoid widening registers that hinder exact stack-slot tracking

This commit adds logic to avoid widening registers used to compute stack
addresses for varying-offset stack accesses during SCEV loop bounds
computations. It collects base registers at loop entry and propagates
them to ensure accurate stack-slot tracking.

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> index a98f3fea52ec0..a093e9bf84b9d 100644
> --- a/kernel/bpf/scev.c
> +++ b/kernel/bpf/scev.c

[ ... ]

> @@ -1060,6 +1060,70 @@ static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct
>  	}
>  }
>  
> +/*
> + * Collect loop-entry registers referenced by expr 'id', best effort.
> + * DEPTH_LIMIT may leave dependencies unrecorded. This mask only avoids
> + * widening that would lose stack-access precision; missing dependencies
> + * may cause false rejections, but this is not a soundness issue.
> + */
> +static void or_expr_regs(struct bpf_verifier_env *env, u32 id, unsigned long *mask)
> +{
> +	struct scev *scev = env->scev;
> +	u32 order;
> +
> +	scev->stack_sz = 0;
> +	expr_stack_push(scev, id);
> +	while (expr_next(scev, &id, &order)) {
> +		if ((order & PRE) && scev->exprs[id].op == REG)
> +			__set_bit(scev->exprs[id].params[0], mask);
> +	}
> +}
> +
> +/* Mask of argument registers (R1..R5) a call at 'idx' passes by register. */
> +static u16 call_params_mask(struct bpf_verifier_env *env, int idx)
> +{
> +	struct bpf_insn *insn = &env->prog->insnsi[idx];
> +	struct bpf_call_summary cs;
> +	int n = bpf_get_call_summary(env, insn, &cs) ? cs.arg_slot_cnt : MAX_BPF_FUNC_REG_ARGS;
> +
> +	return n ? GENMASK(BPF_REG_1 + n - 1, BPF_REG_1) : 0;
> +}

[Severity: Medium]
Can this regression generate an incorrect register mask if a kfunc takes
more than 5 argument slots?

If bpf_get_call_summary() returns a kfunc with cs.arg_slot_cnt > 5 (since
extra arguments are passed on the stack), n will exceed MAX_BPF_FUNC_REG_ARGS.
Using GENMASK() with n > 5 will overflow into callee-saved registers like
BPF_REG_6 through BPF_REG_10, incorrectly treating them as argument
registers.

Could this pollute base_regs in collect_store_base_regs() and lead to the
verifier falsely rejecting valid BPF programs due to erroneous stack address
dependency tracking?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=30

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps
  2026-10-04 13:37 ` [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps Eduard Zingerman
@ 2026-10-04 14:02   ` sashiko-bot
  2026-10-04 14:40   ` bot+bpf-ci
  1 sibling, 0 replies; 69+ messages in thread
From: sashiko-bot @ 2026-10-04 14:02 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: unsound scalar pruning in explore_alu_limits due to ignored stride
--

commit d4d6630f513d526c40cb9f5ea524c5bd6dd31948
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: representation for intervals with steps

Extends scalar register tracking with base and step fields to represent
strided values (base + step * k). Updates verifier ALU operations to
propagate these bounds.

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 4371fb3405acb..05914fcbe77cd 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -178,6 +178,13 @@ struct bpf_reg_state {
>  	 * during state comparisons.
>  	 */
>  	u32 map_uid;
> +	/*
> +	 * The value described by this register, interpreted as s64, lies on
> +	 * a line described by a linear equation base + step * k.
> +	 * Invariant: base < step.
> +	 */
> +	u16 base;
> +	u16 step;

[Severity: High]
Does adding these fields after the id field bypass the exact state comparison
checks?

Functions like regs_exact() and scalars_exact_for_widen() use
offsetof(struct bpf_reg_state, id) for their memcmp() limit. Because
base and step are positioned after id in the struct, they are excluded from
the memcmp().

For example, in kernel/bpf/states.c:regsafe():

    if (env->explore_alu_limits) {
        return memcmp(rold, rcur, offsetof(struct bpf_reg_state, id)) == 0 &&
               check_scalar_ids(rold->id, rcur->id, idmap);
    }

When explore_alu_limits is enabled, could the verifier incorrectly prune
an unsafe execution branch (e.g. one with an unaligned scalar step) using
a cached safe branch that had an aligned step? This might allow unaligned
scalar offsets to bypass BTF array bounds alignment checks.

A similar omission occurs in regs_exact() which could cause false-positive
infinite loop detections, and in scalars_exact_for_widen() which might break
loop convergence expectations. Should these exact-match functions be updated
to explicitly compare the base and step fields?

>  	/* if (!precise && SCALAR_VALUE) min/max/tnum don't affect safety */
>  	bool precise;
>  };

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com?part=14

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations
  2026-10-04 13:37 ` [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1885 bytes --]

> bpf: expose comparison opcode transformations

TL;DR: Nit only: bpf_flip_opcode() is exported but nothing outside verifier.c
appears to use it, so it could stay static and the commit message be adjusted.

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 79ca8b28a83a..4371fb3405ac 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -1872,4 +1872,7 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env);
>  int bpf_do_misc_fixups(struct bpf_verifier_env *env);
>  int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
>
> +int bpf_flip_opcode(u32 opcode);
> +u8 bpf_rev_opcode(u8 opcode);
> +
>  #endif /* _LINUX_BPF_VERIFIER_H */

[ ... ]

> @@ -17427,7 +17425,7 @@ static int is_scalar_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_s
>  	return simulate_both_branches_taken(env, opcode, is_jmp32);
>  }
>
> -static int flip_opcode(u32 opcode)
> +int bpf_flip_opcode(u32 opcode)
>  {
>  	/* How can we transform "a <op> b" into "b <op> a"? */
>  	static const u8 opcode_flip[16] = {

This isn't a bug, but is there a user of bpf_flip_opcode() outside of
verifier.c?

Nothing else in the series (3ece537e8490..2c96fb13421f) calls it.  The only
external user added later is match_linear_latch() in kernel/bpf/scev.c, and
it calls just bpf_rev_opcode():

    op = bpf_rev_opcode(op);

The commit message says "The helpers are used by SCEV logic while analyzing
loop conditions", which appears to be true only of bpf_rev_opcode().

Could bpf_flip_opcode() stay static, with only bpf_rev_opcode() exported and
the commit message adjusted accordingly?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info
  2026-10-04 13:37 ` [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4071 bytes --]

> bpf: represent stack access effects with arg_access_info

TL;DR: In record_stack_access(), the early return for multi-offset must_write
pointers may skip read marks when an access is also may_read. This only
matters in later intermediate commits, where stack slots could look dead to
pruning.

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index cd9523f69298..7b317fb5ca82 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c

[ ... ]

> @@ -1456,49 +1453,53 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,

[ ... ]

> +	if (info.may_read) {
>  		/* Mark any touched slot as use */
>  		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
> -		slot_lo = max_t(s32, (-fp_off - access_bytes) / STACK_SLOT_SZ, 0);
> -		return mark_stack_read(instance, frame, insn_idx, slot_lo, slot_hi);
> -	} else if (access_bytes < 0) {
> +		slot_lo = max_t(s32, (-fp_off - info.size) / STACK_SLOT_SZ, 0);
> +		err = mark_stack_read(instance, frame, insn_idx, slot_lo, slot_hi);
> +		if (err)
> +			return err;
> +	}
> +	if (info.must_write && info.size != U32_MAX) {

[ ... ]

> +/* Record access through a pointer with a known frame and possibly known offsets. */
> +static int record_stack_access(struct bpf_verifier_env *env,
> +			       struct func_instance *instance,
>  			       const struct arg_track *arg,
> -			       s64 access_bytes, u32 frame, u32 insn_idx)
> +			       struct arg_access_info info, u32 frame, u32 insn_idx)
>  {
>  	int i, err;
>  
> -	if (access_bytes == 0)
> +	if (!info.size)
>  		return 0;
>  	if (arg->off_cnt == 0) {
> -		if (access_bytes > 0 || access_bytes == S64_MIN)
> -			return mark_stack_read_all(env, instance, frame, insn_idx);
> +		if (info.may_read) {
> +			err = mark_stack_read_all(env, instance, frame, insn_idx);
> +			if (err)
> +				return err;
> +		}
>  		return 0;
>  	}
> -	if (access_bytes != S64_MIN && access_bytes < 0 && arg->off_cnt != 1)
> +	if (info.size != U32_MAX && info.must_write && arg->off_cnt != 1)
>  		/* multi-offset write cannot set stack_def */
>  		return 0;
                ^^^^

If an access has both may_read and must_write set, and the pointer has
multiple possible offsets, this early return fires before the loop that
calls record_stack_access_off().  Doesn't that mean no mark_stack_read()
happens for the slots the access reads?

record_stack_access_off() was written so that a read can be followed by a
write (the may_read branch now falls through to the must_write branch), but
this check only looks at must_write:

    if (info.size != U32_MAX && info.must_write && arg->off_cnt != 1)

At this commit every producer (record_load_store_access() and
stack_access_info()) sets may_read and must_write mutually exclusively, so
the condition behaves like the old 'access_bytes < 0 && off_cnt != 1' check.

It becomes reachable in a3aabad37761 ("bpf: describe helper stack accesses
with arg_access_info"), where func_arg_access_type() returns BPF_READ |
BPF_WRITE for MEM_WRITE without MEM_UNINIT, and
bpf_helper_stack_access_bytes() then sets both may_read = 1 and
must_write = 1.  Examples are dynptr args such as the newly MEM_WRITE
annotated ringbuf submit/discard dynptr args (even unprivileged), and
raw-mem args when allow_uninit_stack is set.  Before that commit those
accesses were plain positive-size reads, so for a multi-offset stack
pointer their read marks are lost in a3aabad37761 and ac317625edcb.  The
slots can then look dead to state pruning.

The later commit 92670b9a02df ("bpf: track may_write flags in liveness")
removes this early return and moves the off_cnt == 1 check into the
must_write branch of record_stack_access_off().

Could this patch drop the early return, or add '!info.may_read' to it, so
the intermediate commits a3aabad37761 and ac317625edcb do not lose read
marks?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin()
  2026-10-04 13:38 ` [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin() Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4011 bytes --]

> bpf: add bpf_same_memory_origin()

TL;DR: bpf_same_memory_origin() may wrongly treat pointers into different
elements of an inner array map (BPF_F_INNER_MAP, template max_entries == 1) as
the same object, so SCEV loop bounds could be wrong and the verifier may
accept out-of-bounds accesses.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 28c16e35b1f3..9e1c08caa0db 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -17153,6 +17153,51 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
>  	return 0;
>  }
>
> +/*
> + * Many checks done by this function are quite conservative.
> + * This is because main verification pass does not maintain
> + * enough information to track object identities for some
> + * of the interesting types, e.g. PTR_TO_MEM.
> + */
> +bool bpf_same_memory_origin(const struct bpf_reg_state *reg_a,
> +			    const struct bpf_reg_state *reg_b)
> +{

[ ... ]

> +	switch (base_type(reg_a->type)) {

[ ... ]

> +	case PTR_TO_MAP_VALUE:
> +		if (reg_a->map_ptr != reg_b->map_ptr || reg_a->map_uid != reg_b->map_uid)
> +			return false;
> +		if (reg_a->id && reg_a->id == reg_b->id)
> +			return true;
> +		/* A plain single-element array has one stable value address. */
> +		if (reg_a->map_ptr->map_type == BPF_MAP_TYPE_ARRAY &&
> +		    reg_a->map_ptr->max_entries == 1)
> +			return true;

Can this shortcut return true for two pointers into different elements of
the same map when map_ptr is an inner map template?

bpf_map_meta_alloc() copies max_entries from the template inner map into
inner_map_meta, and array_map_meta_equal() does not compare max_entries
when BPF_F_INNER_MAP is set:

    return meta0->map_flags & BPF_F_INNER_MAP ? true :
           meta0->max_entries == meta1->max_entries;

So an ARRAY_OF_MAPS whose template is an array map with BPF_F_INNER_MAP
and max_entries == 1 can hold a runtime inner array with many more
entries.  The verifier already special-cases this elsewhere:
can_elide_value_nullness() returns false for BPF_F_INNER_MAP, and
array_map_gen_lookup() refuses to inline it.

Consider this program:

    /* map_ptr = inner_map_meta, map_uid = U */
    inner = bpf_map_lookup_elem(&outer, &zero);
    /* id = A, map_uid = U */
    p = bpf_map_lookup_elem(inner, &k0);
    /* id = B, map_uid = U */
    e = bpf_map_lookup_elem(inner, &k1);
    if (!p || !e)
        return 0;
    e += 64;
    i = 0;
    do { use(arr[i]); i++; p += 8; } while (p < e);

Here p and e share map_ptr and map_uid.  Their ids differ, so the id
check fails, and the code falls through to the map_type == BPF_MAP_TYPE_ARRAY
&& max_entries == 1 test on the template and returns true.  At runtime,
p and e point at elements k0 and k1, which are (k1 - k0) * elem_size bytes
apart.

Later in this series, fd7009a2fd2e ("bpf: use SCEV to widen bounded
loops") has scev_value_compatible() call this function before
compute_max_iters() turns (bound - initial) / step into
iters->max_header_count.  linear_bounds() and bpf_clamp_scev_regs() then
use that count to widen and clamp induction variables.

In the example above the verifier would bound i to about [0, 8], but at
runtime the loop runs about (k1 - k0) * elem_size / 8 times.  Can arr[i]
then read or write past the end of arr even though the verifier accepted
the access?

Comparing two pointers is allowed when env->allow_ptr_leaks is set, so a
program with CAP_BPF + CAP_PERFMON (or a BPF token) can reach this.

The opening comment says the checks are conservative, but this branch is
not.  Should the shortcut also require:

    !(reg_a->map_ptr->map_flags & BPF_F_INNER_MAP)

the same exclusion can_elide_value_nullness() uses?  No later commit in
2ac7d08257d4..2c96fb13421f changes this check.

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic
  2026-10-04 13:38 ` [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4846 bytes --]

> selftests/bpf: tests for register base/step arithmetic

TL;DR:
- Commit message lists tests (step_mul_non_pow2, step_lsh,
  step_add_const_base) that don't exist
- The "w1 = w0" selftest match may not check R1's stride
- Comment has a wrong value (0xfffffffa)

> selftests/bpf: tests for register base/step arithmetic
>
> - step_mul_non_pow2: Retain step 3 after multiplying a bounded scalar
>   by a non-power-of-two constant.
> - step_lsh: Record step 4 after shifting a bounded scalar left by two.
> - step_add_const_base: Shift the base to 1 while preserving step 4
>   after adding a constant.
> - step_neg_value_range: Preserve the signed range of {-3, 0, 3, 6}
>   when intersecting with a non-power-of-two step across zero.
> - step_arith{32,64}_overflow: Reset the stride and preserve valid
>   bounds when a {32,64}-bit operation wraps.
> - step_mov32_truncate: Reset the stride when MOV32 discards nonzero
>   upper bits from positive values.
> - step_mov32_negative: Reset the stride when MOV32 zero-extends
>   the low 32 bits of negative signed values.
> - step_movsx: Reset the stride when sign extension changes the
>   register value.
> - step_linked_regs_add64: Translate both bounds and base when
>   propagating a comparison's refinement through a scalar-ID delta.
>
> Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

The changelog lists step_mul_non_pow2, step_lsh and step_add_const_base, but
the patch doesn't define tests with those names, and no later patch in the
series does either.

It also doesn't mention step_arith64_no_overflow or step_arith32_no_overflow,
which are in the patch.  Those cover MUL by 3, ADD of 1, MUL by 2 and LSH by 1
(step=0+3 -> 1+3 -> 2+6 -> 4+12), not "step 4 after shifting a bounded scalar
left by two" or "shift the base to 1 while preserving step 4".

Should the bullet list be updated to match the tests that are actually added?

> diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> new file mode 100644
> index 000000000000..079e34820a8f
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> @@ -0,0 +1,290 @@

[ ... ]

> +/*
> + * Safe ALU32 multiplication and addition preserve the stride and scale its base:
> + * step=0+3 -> step=1+3 -> step=2+6. MOV32 preserves the result.
> + * LSH32 currently resets the stride, but must preserve the correct upper bound.
> + */
> +SEC("socket")
> +__success __log_level(2)
> +__msg("w0 *= 3 {{.*}}step=0+3)")
> +__msg("w0 += 1 {{.*}}step=1+3)")
> +__msg("w0 *= 2 {{.*}}step=2+6)")
> +__msg("w1 = w0 {{.*}}step=2+6)")

Can this match pass even if MOV32 doesn't preserve the stride?

The verifier log line for this instruction prints R0 first and then R1,
roughly:

    w1 = w0 ; R0=scalar(id=1,...,step=2+6) R1=scalar(id=1,...)

R0 still has step=2+6 from the previous instruction, so the regex can match
R0's state on its own.  If the MOV32 path (check_alu_op() -> zext_32_to_64())
started resetting the destination's stride, this assertion would still pass.

step_arith32_overflow in this same file anchors the destination register:

    __msg("r2 = r0 {{.*}}R0=scalar({{.*}}step=0+3) R2=scalar({{.*}}step=0+3)")

Would something like "w1 = w0 {{.*}}R1=scalar({{.*}}step=2+6)" be needed
here so the check actually tests R1?

[ ... ]

> +/*
> + * 0x49249249 * 7 wraps to U32_MAX. Keeping step=0+7 incorrectly caps the
> + * result at 0xfffffffc instead of U32_MAX.
> + * Adding 4 to {0xfffffffc, U32_MAX} wraps to {0, 3}; keeping step=1+3
> + * incorrectly narrows the result to 1.
> + * Shifting the same input left by 1 gives {0xfffffff8, 0xfffffffe};
> + * keeping step=0+6 incorrectly narrows the result to 0xfffffffa.
> + */

This isn't a bug, but is 0xfffffffa the right value in the last sentence of
this comment?

0xfffffffa is not on the line 0+6k (0xfffffffa mod 6 == 4).  If step=0+6 were
kept after w2 <<= 1, the range would be [0xfffffff8, 0xfffffffe], and
cnum64_intersect_linear() would round smin up by imod(0 - 2, 6) = 4 and smax
down by imod(2 - 0, 6) = 2, leaving the constant 0xfffffffc, the only
multiple of 6 in that range.

Should this read 0xfffffffc?  The test itself still passes, since it only
checks smax=umax=umax32=0xfffffffe and that step= is absent.

> +SEC("socket")
> +__success __log_level(2)
> +__msg("w0 *= 7 {{.*}}umax=0xffffffff,") __no_step
> +__msg("r2 = r0 {{.*}}R0=scalar({{.*}}step=0+3) R2=scalar({{.*}}step=0+3)")
> +__msg("w0 += 4 {{.*}}smax=umax=smax32=umax32=3,") __no_step
> +__msg("w2 <<= 1 {{.*}}smax=umax=umax32=0xfffffffe,") __no_step


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation
  2026-10-04 13:38 ` [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, emil, ast,
	andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 5202 bytes --]

> selftests/bpf: tests for loop hierarchy computation

TL;DR:
- Stray "dominators test case" line after Signed-off-by in commit message
- Test comments only: wrong function name (compute_loops) and possibly wrong
  "Case C" label in loop_with_if_else

>     selftests/bpf: tests for loop hierarchy computation
>
>     Test cases covering the following branches in bpf_compute_loops():
>     - Case B: simple backedge creating a loop (loop_single)
>     - Case B: two independent loops (loop_two_independent)
>     - Case B + D: nested loops where the inner header's loop_header points
>       to the outer header (loop_nested)
>     - Case C: diamond CFG with no loops (fwd_edges_no_loop)
>     - Case D: sibling inner loops within one outer loop
>       (loop_nested_siblings)
>     - Three levels of loop nesting (loop_three_levels)
>     - Loop with an if-else body containing forward branches
>       (loop_with_if_else)
>     - Case E: An irreducible loop (loop_irreducible)
>     - A self-loop (loop_self)
>     - A test with sibling loops nested in outer loops
>       (loop_nested_siblings_common_ancestors)
>     - A test case with outer loop's backedge originating
>       from an inner loop (loop_outer_backedge_from_inner).
>
>     Cases B, C, D and E are described in loops.c:compute_loops_in_subprog().
>
>     Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
>     Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
>
>     dominators test case

Should the trailing "dominators test case" line after the Signed-off-by
trailer be dropped?

It is not part of the changelog, and it looks like leftover text from a
squash or fixup.  The next patch in the series, "selftests/bpf: tests for
immediate dominator computation", is the dominators test patch.  As written,
this line would end up in git history after the trailer block.

> diff --git a/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
> new file mode 100644
> index 000000000000..0f49108c0e50
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
> @@ -0,0 +1,346 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +
> +#include <linux/bpf.h>
> +#include <bpf/bpf_helpers.h>
> +#include "bpf_misc.h"
> +
> +/*
> + * kernel/bpf/loops.c:compute_loops() distinguish between
> + * the following cases:
> + * - B: backedge -> simple loop
> + * - C: cross edge to non-loop node -> no-op
> + * - D: edge to node whose header is in DFS path -> nested loop
> + * - E: edge to node whose header is NOT in DFS path -> irreducible
> + *
> + * Below test cases cover the above branches in various combinations.
> + */

This isn't a bug, but there is no function named compute_loops() in
kernel/bpf/loops.c.

The B/C/D/E branches live in compute_loops_in_subprog(), which is called
from bpf_compute_loops().  The commit message already names
compute_loops_in_subprog() correctly.

Could this comment name compute_loops_in_subprog() so readers can find the
case labels it lists?  Also, "distinguish" should probably be
"distinguishes".

[ ... ]

> +/* Loop with an if-else body (forward branch inside loop, Case C). */
> +SEC("socket")
> +__success
> +__log_level(2)
> +__msg("loop at 1")
> +__msg("Program dump")
> +__msg("         -1   0: {{.*}} (b7) r0 = 0")
> +__msg("  1       0   1: {{.*}} (07) r0 += 1")
> +__msg("  1   1   1   2: {{.*}} (bf) r1 = r0")
> +__msg("  1   1   2   3: {{.*}} (25) if r1 > 0x5 goto pc+1")
> +__msg("  1   1   3   4: {{.*}} (b7) r1 = 1")
> +__msg("  1   1   3   5: {{.*}} (0f) r0 += r1")
> +__msg("  1   1   5   6: {{.*}} (a5) if r0 < 0x64 goto pc-6")
> +__msg("          6   7: {{.*}} (95) exit")
> +__naked void loop_with_if_else(void)
> +{
> +	asm volatile ("					\
> +	r0 = 0;						\
> +1:	r0 += 1;					\
> +	r1 = r0;					\
> +	if r1 > 5 goto 2f;				\
> +	r1 = 1;						\
> +2:	r0 += r1;					\
> +	if r0 < 100 goto 1b;				\
> +	exit;						\
> +"	::: __clobber_all);
> +}

This isn't a bug, but is "Case C" the right label in the comment above
loop_with_if_else()?

In compute_loops_in_subprog(), bpf_insn_successors() returns the fallthrough
first, so the DFS goes 3 -> 4 -> 5 -> 6 and, when it returns, sets
aux[5].loop_header = 1.  When the edge 3 -> 5 (the "if r1 > 5 goto 2f" jump)
is examined afterwards, insn 5 is fully explored and its loop_header is 1,
not -1.  Since dfs_pos[1] != 0, the "else if (dfs_pos[aux[s].loop_header])"
branch runs, which is Case D.

Case C only fires for aux[s].loop_header == -1, and the file header defines
it as a "cross edge to non-loop node".  Insn 5 is inside the loop at 1.  The
result is the same if the DFS order is reversed, because then 4 -> 5 would be
the cross edge into a loop node.

Only fwd_edges_no_loop() really covers Case C (the 4 -> 5 cross edge).  Should
this comment say Case D, or name no case, as the commit message does for this
test?

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation
  2026-10-04 13:38 ` [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3920 bytes --]

> selftests/bpf: tests for immediate dominator computation

TL;DR: The ldimm64 selftests in verifier_idoms.c may pass even if the
bpf_is_ldimm64() skip in compute_predecessors() is removed, so comments and
changelog overstate what they check.

> diff --git a/tools/testing/selftests/bpf/progs/verifier_idoms.c b/tools/testing/selftests/bpf/progs/verifier_idoms.c
> new file mode 100644
> index 000000000000..a29d6332baa4
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_idoms.c
> @@ -0,0 +1,390 @@

[ ... ]

> +
> +/*
> + * Straight-line code, with an ldimm64 in the middle. Instruction index 2 is
> + * the second half of the ldimm64 at index 1 and is not a CFG node; the idom of
> + * index 3 must be 1, not 2. Exercises the bpf_is_ldimm64() skip in both passes
> + * of compute_predecessors().
> + */
> +SEC("socket")
> +__success
> +__log_level(2)
> +__msg("Program dump")
> +__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
> +__msg("{{.*}}  0   1: {{.*}} (18) r1 = 0x1122334455667788")
> +__msg("{{.*}}  1   3: {{.*}} (07) r0 += 1")
> +__msg("{{.*}}  3   4: {{.*}} (0f) r0 += r1")
> +__msg("{{.*}}  4   5: {{.*}} (95) exit")
> +__naked void straight_line_ldimm64(void)
> +{
> +	asm volatile ("					\
> +	r0 = 0;						\
> +	r1 = 0x1122334455667788 ll;			\
> +	r0 += 1;					\
> +	r0 += r1;					\
> +	exit;						\
> +"	::: __clobber_all);
> +}

Can this test actually fail if the bpf_is_ldimm64() skip is missing from
compute_predecessors()?

The comment above and the commit message ("exercises the bpf_is_ldimm64()
skip in compute_predecessors()") present "the idom of index 3 must be 1,
not 2" as the failure being guarded against.

Without the skip, compute_predecessors() would call
bpf_insn_successors(env, 2) on the second half of the ldimm64 and record
insn 2 as an extra predecessor of insn 3.

But insn 2 is never in env->cfg.insn_postorder, because
bpf_compute_postorder() only reaches it through bpf_insn_successors(), which
advances by insn_sz == 2 for an ldimm64.  So idoms[2] keeps the -1 that
bpf_compute_idoms() set, and compute_subprog_idoms() drops that predecessor:

    if (idoms[pred] == -1)
        continue;

idom(3) still comes out as 1, so the __msg() lines above pass either way.
The only other effect is aux[].bb_end, which log_program() does not print.

A value of 2 would only appear if the ldimm64 second half were treated as a
CFG node in bpf_insn_successors() or bpf_compute_postorder(), not if the
compute_predecessors() skip were missing.

[ ... ]

> +#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64) || defined(__TARGET_ARCH_powerpc)
> +/*
> + * Indirect jump (gotox) with a two-entry jump table. The gotox (insn 4) has two
> + * successors (insn 5 and insn 7), so both targets have the gotox as their only
> + * predecessor and idom. Also re-exercises the ldimm64 skip: insn 2's idom is 0,
> + * the ldimm64 at index 0 (index 1 is its second half).
> + */
> +SEC("socket")
> +__success
> +__log_level(2)
> +__msg("Program dump")
> +__msg("{{.*}} -1   0: {{.*}} (18) r0 = {{0x[0-9a-f]+}}")
> +__msg("{{.*}}  0   2: {{.*}} (07) r0 += 8")
> +__msg("{{.*}}  2   3: {{.*}} (79) r0 = *(u64 *)(r0 +0)")
> +__msg("{{.*}}  3   4: {{.*}} (0d) gotox r0")
> +__msg("{{.*}}  4   5: {{.*}} (b7) r0 = 0")
> +__msg("{{.*}}  5   6: {{.*}} (95) exit")
> +__msg("{{.*}}  4   7: {{.*}} (b7) r0 = 1")
> +__msg("{{.*}}  7   8: {{.*}} (95) exit")
> +__naked void gotox_jump_table(void)

The gotox_jump_table comment ("Also re-exercises the ldimm64 skip: insn 2's
idom is 0") has the same limitation.

Should the comments and the changelog describe what is actually checked, or
should the test assert something that depends on the skip?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking
  2026-10-04 13:38 ` [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2236 bytes --]

> selftests/bpf: tests for may_write stack-liveness tracking

TL;DR: Test comment issue only: a new comment in verifier_live_stack.c says
one dynamic callback writes through ctx, but neither callback appears to
write. Comment or callback code should be fixed.

> diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
> index 0e6f25a09b5d..a06156c18c9e 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c

[ ... ]

> @@ -2271,11 +2339,19 @@ static __used __naked void merge_leaf_read(void)
>  	::: __clobber_all);
>  }
>  
> -/* Same bpf_loop instruction calls different callbacks depending on branch. */
> +/*
> + * Same bpf_loop instruction calls different callbacks depending on branch.
> + * One callback writes through ctx. With no unique callback, all frames must
> + * be marked as possibly read and written, including at enclosing call sites.
> + */

The new comment says "One callback writes through ctx", but do either of
the callbacks write?

The callbacks this test reaches through mid_dynamic_cb()'s bpf_loop call
only load through ctx:

    static __used __naked void dyn_cb_a(void)
    {
    	asm volatile (
    	"r0 = *(u64 *)(r2 + 0);"
    	"r0 = 0;"
    	"exit;"
    ...
    static __used __naked void dyn_cb_b(void)
    {
    	asm volatile (
    	"r0 = *(u64 *)(r2 + 0);"
    	...

The comment above them also describes these as "Callback A/B: read parent
stack through ctx", so the new sentence contradicts both the code and the
neighbouring comment.

The new may_def expectations do not depend on what the callbacks do, since
analyze_subprog() calls mark_stack_read_all() and
mark_stack_may_write_all() for every frame when find_callback_subprog()
returns -2, and never analyzes the callbacks.

Should dyn_cb_b() be changed to store through r2 (for example
"*(u64 *)(r2 + 0) = 0;"), or should that sentence be dropped?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening
  2026-10-04 13:38 ` [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 8855 bytes --]

> selftests/bpf: tests for SCEV analysis and loop widening

TL;DR:
- store_load_exprs test lacks CAN_USE_LOAD_ACQ_STORE_REL guard, so it may fail
  instead of skip on 32-bit or some JIT-only arches
- nested_irreducible_loop __not_msg may never fail
- comment inaccuracies only

> diff --git a/tools/testing/selftests/bpf/progs/verifier_scev.c b/tools/testing/selftests/bpf/progs/verifier_scev.c
> new file mode 100644
> index 000000000000..7a1b5b8ec287
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_scev.c
> @@ -0,0 +1,2376 @@

[ ... ]

> +__naked void store_load_exprs(void)
> +{
> +	asm volatile (

[ ... ]

> +	ST_OP(".8byte %[load_acquire8]")
> +	ST_OP(".8byte %[load_acquire16]")
> +	ST_OP(".8byte %[load_acquire32]")
> +	ST_OP(".8byte %[load_acquire64]")
> +	ST_OP(".8byte %[store_release8]")
> +	ST_OP(".8byte %[store_release16]")
> +	ST_OP(".8byte %[store_release32]")
> +	ST_OP(".8byte %[store_release64]")

A subsystem pattern flags this as potentially concerning: a test that depends
on an optional capability should be skipped, rather than fail, when the
capability is absent.

store_load_exprs() always emits BPF_LOAD_ACQ and BPF_STORE_REL instructions
and is tagged __success.  The other selftests that use these instructions
(verifier_load_acquire.c, verifier_store_release.c, verifier_precision.c,
verifier_arena_scalar.c, compute_live_registers.c and verifier_live_stack.c)
guard them with #ifdef CAN_USE_LOAD_ACQ_STORE_REL.

Without such a guard, could this test fail instead of being skipped on
configurations that lack support?  On 32-bit kernels, check_atomic() rejects
the BPF_DW form:

    case BPF_LOAD_ACQ:
        if (BPF_SIZE(insn->code) == BPF_DW && BITS_PER_LONG != 64) {
            verbose(env, "64-bit load-acquires are only supported on 64-bit arches\n");
            return -EOPNOTSUPP;
        }

On 64-bit arches whose JIT does not implement these instructions (mips64 and
sparc64, for example) and that set CONFIG_BPF_JIT_ALWAYS_ON,
bpf_prog_select_runtime() returns -ENOTSUPP.

In both cases test_loader would report unexpected_load_failure for the whole
store_load_exprs subtest, including its ordinary load/store expectations.
Without CONFIG_BPF_JIT_ALWAYS_ON the interpreter handles these instructions,
so the 64-bit arch case would still pass there.

Would it make sense to put the load_acquire/store_release ST_OP lines, their
__imm_insn operands and their __msg expectations under
#ifdef CAN_USE_LOAD_ACQ_STORE_REL, or move them into a separate guarded test?

[ ... ]

> +/*
> + * Loop A (counter in r6) contains irreducible loop B (header 5),
> + * which contains C (counter in r9).
> + * Entering B at 'body' saves and restores r6, making it appear invariant.
> + * Entering B at 'alternate' skips the save and modifies r6 instead.
> + * Hence A must not infer a SCEV expression for r6.
> + * SCEV expression for r9 in C should still be computed.
[ ... ]
> + *  5: B:  r8 = r6;                      // B
> + *  6:     goto body;
> + *  7: alternate:
> + *         r8 = r6;
> + *  8:     r8++;
> + *  9: body:
> + *         r6 = r8;

This isn't a bug, but does the comment assign the two entries of B the wrong
way round?

B is {5, 6, 9..14} in the pseudo-code above and has two entry points: its
header 'B' (insn 5, reached by fall-through from insn 4) and 'body' (insn 9,
reached from 'alternate', insns 7-8, which are outside B).

Entering at the header 'B' is the path that saves and restores r6:

    r8 = r6; goto body; ... r6 = r8;

Entering at 'body' via 'alternate' is the path that modifies r6:

    r8 = r6; r8++; r6 = r8;

The comment says "Entering B at 'body' saves and restores r6" and "Entering B
at 'alternate' skips the save and modifies r6".  That names 'body' as the
invariant entry, and names 'alternate' as an entry of B although it is not
part of B.

Would something like this be more accurate?

    Entering B at its header 'B' saves and restores r6, making it appear
    invariant.  Entering B at 'body' (via 'alternate') skips the save and
    modifies r6 instead.

[ ... ]

> +__msg("loop at 1{{$}}")
> +__msg("loop at 5, nested in 1, irreducible")
> +__msg("loop at 11, nested in 5")
> +__msg("scev at header 1:")
> +__msg_next("  r6=?")
> +__msg("scev at header 11:")
> +__msg_next("  r9=(+ r9 1) / (linear r9 1)")
> +__msg("loop header at 11, widening r9 to 0..2 step 1")
> +__not_msg("loop header at 1, widening r6")
> +__naked void nested_irreducible_loop(void)

Can this __not_msg() ever fail?  The comment says "Hence A must not infer a
SCEV expression for r6", but it looks like the check cannot catch widening of
r6 in loop A.

test_loader's match_negative_msgs() only searches for a __not_msg() pattern
in the span between the end of the previous positive match and the start of
the next one.  When no positive match follows, the span runs to the end of
the log.  Here the previous positive match is:

    __msg("loop header at 11, widening r9 to 0..2 step 1")

Widening is only logged when a loop is first pushed on the loop stack:
handle_loop_entry_exit() calls bpf_compute_loop_iters() and
bpf_widen_scev_regs() only when loop_stack_push() sets pushed = true.

Loop A (header 1) is pushed once, on the 0 -> 1 edge at the very start of
do_check(), which happens before the verifier first reaches insn 11 via
1..4 -> 5 -> 6 -> 9 -> 10 -> 11.  Every later visit to insn 1 arrives through
the 15 -> 1 backedge with A already on the loop stack.  That path goes to
maybe_clamp_scev_regs() and never logs "widening".

So any "loop header at 1, widening r6 ..." line would appear before the
matched r9 widening line, outside the searched span.  If
bpf_widen_scev_regs() started widening r6 for loop A, this check would still
pass.  Only the earlier __msg_next("  r6=?") would catch a related
regression, and that checks the SCEV printout rather than the widening
decision.

Would moving the __not_msg() before the "widening r9" __msg(), for example
right after:

    __msg_next("  r9=(+ r9 1) / (linear r9 1)")

make its span cover do_check()'s processing of header 1?

[ ... ]

> +/*
> + * Complement of no_widen_stack_spill: a sub-register (1-byte) store to the stack
> + * lands as STACK_MISC and carries no tracked value, so the induction variable
> + * addressing it (r2) is still widened.
> + */
[ ... ]
> +__naked void widen_byte_stack_store(void)

Is the stated reason for widening accurate here?  The comment says a 1-byte
store "lands as STACK_MISC and carries no tracked value", but that only holds
once r2 has been widened and the store has a varying offset
(check_stack_write_var_off() marks the bytes STACK_MISC).

At a fixed, 8-byte-aligned offset, which is what the un-widened loop would
produce here (fp-64, fp-56, fp-48, fp-40), check_stack_write_fixed_off()
saves a scalar byte store as a narrow spill whose value is tracked:

    if (reg && !(off % BPF_REG_SIZE) && reg->type == SCALAR_VALUE &&
        env->bpf_capable) {
        ...
        save_register_state(env, state, spi, reg, size);

check_stack_read_fixed_off() then returns the exact value for a 1-byte read
of the same slot.  So widening r2 does lose tracked values for byte stores
too.

The reason r2 is widened appears to be that collect_store_base_regs() only
treats BPF_W and BPF_DW stores and loads as stack bases:

    if (size == BPF_W || size == BPF_DW)

If so, the cause and effect in the comment are reversed: the store lands as
STACK_MISC because r2 was widened, not the other way round.  This also
misdescribes the trade-off made by commit 04dcb46b3aba.

Could the comment say that byte and half-word stores are intentionally not
considered stack bases (accepting the loss of narrow-spill precision), rather
than saying such stores carry no tracked value?

[ ... ]

> +/*
> + * struct xdp_md *slots[2] = { ctx, ctx };
> + *
> + * r7 = 0;
> + * for (r8 = 0; r8 < 2; r8++) {
> + *         value = slots[r7]->data;
> + *         if (i == 0)
> + *                 r7 = 1;

[ ... ]

> +	if r8 != 0 goto 2f;				\
> +	r7 = 8;			/* conditionally update r7 */ \

This isn't a bug, but the pseudo-code in the comment tests a variable i
that does not exist:

    if (i == 0)
            r7 = 1;

The loop counter is r8 (for (r8 = 0; r8 < 2; r8++)) and the assembly tests
r8:

    if r8 != 0 goto 2f;
    r7 = 8;

Should the condition read "if (r8 == 0)"?

The same comment also uses two units for r7: the pseudo-code sets r7 = 1 (an
index into slots[]), while the assembly and the following sentence ("SCEV
expression for r7 is (any r7 8)") use the byte offset 8.  Should the
pseudo-code say whether r7 is an index or a byte offset?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops
  2026-10-04 13:38 ` [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops Eduard Zingerman
@ 2026-10-04 14:24   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2003 bytes --]

> bpf: use SCEV to widen bounded loops

TL;DR: scratch_widened_reg_id() in kernel/bpf/scev.c ends a switch with a bare
"default:" label, which may fail to build on GCC older than 11 or with Clang
under CONFIG_WERROR; a "break;" is suggested.

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> index 9ef28e0d8d48..a98f3fea52ec 100644
> --- a/kernel/bpf/scev.c
> +++ b/kernel/bpf/scev.c

[ ... ]

> @@ -1581,3 +1585,839 @@ int bpf_init_scev(struct bpf_verifier_env *env)

[ ... ]

> +static void scratch_widened_reg_id(struct bpf_verifier_env *env, struct bpf_reg_state *reg)
> +{
> +	switch (base_type(reg->type)) {
> +	case SCALAR_VALUE:
> +		reg->id = 0;
> +		break;
> +	case PTR_TO_PACKET:
> +	case PTR_TO_PACKET_META:
> +		reg->id = ++env->id_gen;
> +	default:
> +	}
> +}

This switch ends with a bare "default:" label directly followed by the closing
brace.  A label at the end of a compound statement is only valid from C23
onward, so does this build with the toolchains the kernel supports?

With GCC older than 11 (the minimum GCC is 8.1.0 per
scripts/min-tool-version.sh) this is a hard error:

    error: label at end of compound statement

Clang accepts it with -Wc23-extensions, which turns into a build failure with
CONFIG_WERROR=y or W=e.  The same pattern has needed fixing before, for
example in 84f9288208df ("Bluetooth: hci_qca: Fix build error") and
d98a04dc1906 ("KVM: arm64: Add break to default case in
kvm_pgtable_stage2_pte_prot()").

The PTR_TO_PACKET/PTR_TO_PACKET_META case also falls into the default label
without a break.  Would adding "break;" after the "reg->id = ++env->id_gen;"
assignment and after "default:" address both?

The pattern is still present at the end of the series (2c96fb13421f), so no
later commit fixes it.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops
  2026-10-04 13:38 ` [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops Eduard Zingerman
  2026-10-04 14:00   ` sashiko-bot
@ 2026-10-04 14:24   ` bot+bpf-ci
  1 sibling, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:24 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4575 bytes --]

> bpf: compute scalar evolution expressions for loops

TL;DR: Commit message only: the worked example says r7 stays (any r7 10) at
the latch, but the code may produce a nested (any (any r7 10) 10) since
simplify() doesn't flatten it.

> bpf: compute scalar evolution expressions for loops
>
> Assign algebraic expressions to loop variables (registers and stack
> spills), describing how their values evolve across iterations. These
> summaries provide symbolic input for the subsequent widening patch.
>
> The algorithm draws on ideas from the following papers:
>
> - "Symbolic Evaluation of Chains of Recurrences for Loop Optimization"
>   Robert A. van Engelen, 2000
> - "The CR# Algebra and its Application in Loop Analysis and Optimization"
>   Robert A. van Engelen, 2004
>
> The analysis proceeds in two phases:
>
> - Compute expressions for one symbolic iteration: start each variable
>   with a reference to its input value, apply instruction effects and
>   join expressions where control-flow paths merge.
> - Convert the resulting backedge updates to recurrences at the header,
>   then substitute these recurrences into expressions at the latch.
>
> Analyze loops innermost first. Within each loop, visit blocks in
> topological order (reverse postorder, with backedges ignored and nested
> loops collapsed). Treat each nested loop as an opaque operation:
> preserve its invariant values, forget the others and continue at its
> exits, without expanding its iterations.
>
> For example, using 64-bit arithmetic:
>
>     0: r7 = 5;
>     1: r6 = 0;
>     2: do {                // header
>     3:     if (r6 == 1)
>     4:         r7 = 10;
>     5:     r6 += 1;
>     6: } while (r6 < 3);   // latch
>
> Expressions use symbolic inputs r6 and r7; the concrete values from
> lines 0-1 are supplied later by the main verifier.
>
> First phase:
> ------------
>
> At each instruction, an environment maps register numbers to their
> current expressions. Transfer functions update this mapping, while
> joins merge the environments arriving along different paths.
>
> For illustration, follow each branch separately to the backedge:
>
> - At 3, take the path skipping 4: r6 = r6, r7 = r7.
> - At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
> - At 2, the first path contributes r6 = (+ r6 1), r7 = r7.
> - At 4, on the other path from 3, transfer(r7 = 10):
>   r7 = r7 -> r7 = 10.
> - At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
> - At 2, join the two paths' contributions:
>   join((+ r6 1), (+ r6 1)) = (+ r6 1) for r6;
>   join(r7, 10) = (any r7 10) for r7.
>
> Equal expressions stay unchanged; different expressions become ANY
> alternatives. These summarize one iteration, without unrolling the loop.
>
> Second phase:
> -------------
>
> The update r6 = (+ r6 1) says that each iteration adds 1 to r6.
> Repeating this update n times gives r6_entry + n, a linear recurrence.
>
> - At header 2, r6 = (+ r6 1) becomes (linear r6 1), meaning
>   r6_entry + n, where n counts iterations from zero.
> - At latch 6, substitute this recurrence into (+ r6 1):
>   (+ (linear r6 1) 1) simplifies to (linear (+ r6 1) 1),
>   meaning r6_entry + 1 + n.
> - At both points, r7 retains (any r7 10): either its entry value
>   or 10. ANY records alternatives, not numeric ranges or the path
>   conditions selecting them. Unchanged variables retain their inputs.
>
> Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

Is the last bullet of the second phase accurate for latch 6?  It says that
at both points r7 retains (any r7 10).

In the worked example the header expression for r7 is (any r7 10), and
compute_header_scevs() keeps it as the header SCEV for r7.  At the latch,
compute_insn_scevs() substitutes that SCEV for (reg r7) and then runs
simplify():

    id = einsn->reg2expr[reg];
    id = transform_expr_once(scev, 0, id, eheader->reg2scev,
                             instantiate_header_scevs);
    ...
    id = transform_expr(scev, id, NULL, simplify);

simplify() has no rule that flattens a nested (any ...), so the expression
for r7 at latch 6 would come out as (any (any r7 10) 10) rather than
(any r7 10).  The set of values is the same, but the text doesn't match the
expression shape the code builds and logs.

Could the commit message describe the latch expression for r7 as it is
actually produced?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy
  2026-10-04 13:38 ` [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy Eduard Zingerman
@ 2026-10-04 14:40   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:40 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2187 bytes --]

> bpf: compute loop hierarchy

TL;DR: Nits only: insn_aux_data->loop_entry is written but never read, so
mark_entry() could be dropped; is_cond_jmp_insn() could reuse existing
helpers.

> diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
> index 0a9e10edca87..545cee0fd956 100644
> --- a/kernel/bpf/loops.c
> +++ b/kernel/bpf/loops.c
> @@ -154,3 +154,434 @@ int bpf_compute_idoms(struct bpf_verifier_env *env)

[ ... ]

> +static void mark_entry(struct bpf_verifier_env *env, int s)
> +{
> +	env->insn_aux_data[s].loop_entry = true;
> +}

This isn't a bug, but is insn_aux_data->loop_entry meant to have a reader
later?

mark_entry() is the only writer of the bit, and it is called from
mark_as_header() and from compute_loops_in_subprog().  Nothing in this
patch, or in the rest of the series, reads loop_entry.

If there is no reader planned, could mark_entry() and the bit be dropped
until something needs them?

[ ... ]

> +static bool is_cond_jmp_insn(struct bpf_insn *insn)
> +{
> +	u8 class = BPF_CLASS(insn->code);
> +	u8 opcode = BPF_OP(insn->code);
> +
> +	if (class != BPF_JMP && class != BPF_JMP32)
> +		return false;
> +
> +	switch (opcode) {
> +	case BPF_JEQ:
> +	case BPF_JGE:
[ ... ]
> +		return true;
> +	default:
> +		return false;
> +	}
> +}

This isn't a bug, but could is_cond_jmp_insn() be replaced with
bpf_insn_is_cond_jump(insn->code) && !bpf_is_may_goto_insn(insn), as
bpf_prune_dead_branches() does in const_fold.c?

[ ... ]

> +			 * case E: 's' is explored, not in current DFS path,
> +			 * its innermost loop header is not in current DFS path,
> +			 * hence 's' is another entry into the same loop.
> +			 */
> +			h = aux[s].loop_header;
> +			mark_irreducible(env, h);
> +			mark_entry(env, s);

This isn't a bug, but the same question as for mark_entry() above applies
to this call: is there a planned reader for insn_aux_data->loop_entry, or
could the mark_entry(env, s) call be dropped along with the bit?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps
  2026-10-04 13:37 ` [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps Eduard Zingerman
  2026-10-04 14:02   ` sashiko-bot
@ 2026-10-04 14:40   ` bot+bpf-ci
  1 sibling, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 14:40 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 7340 bytes --]

> bpf: representation for intervals with steps

TL;DR:
- base/step not compared with explore_alu_limits or regs_exact(), so states
  may be wrongly pruned, skipping verification
- step not logged when var_off is unknown
- 64-bit modulo in scalar_step_scale() may fail to link on 32-bit

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 4371fb3405ac..05914fcbe77c 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -178,6 +178,13 @@ struct bpf_reg_state {
>  	 * during state comparisons.
>  	 */
>  	u32 map_uid;
> +	/*
> +	 * The value described by this register, interpreted as s64, lies on
> +	 * a line described by a linear equation base + step * k.
> +	 * Invariant: base < step.
> +	 */
> +	u16 base;
> +	u16 step;

Are base and step covered by the strict scalar comparisons in states.c now
that they sit after id in struct bpf_reg_state?

With env->explore_alu_limits set, the SCALAR_VALUE case of regsafe() does
this:

kernel/bpf/states.c:regsafe() {
    ...
		if (env->explore_alu_limits) {
			/* explore_alu_limits disables tnum_in() and range_within()
			 * logic and requires everything to be strict
			 */
			return memcmp(rold, rcur, offsetof(struct bpf_reg_state, id)) == 0 &&
			       check_scalar_ids(rold->id, rcur->id, idmap);
		}
    ...
}

The memcmp stops at id, so base and step are never compared, and the
function returns before range_within() is reached.  Before this patch
var_off, r64 and r32 all sat ahead of id, so this comparison really was
strict for every field that drives value tracking.

With this patch step is value-relevant:

    reg_bounds_sync() -> __reg_deduce_bounds() -> deduce_bounds_64_from_step()

snaps r64 onto the line, and is_branch_taken() ->
simulate_both_branches_taken() treats an empty r64 after that snap as a dead
branch.  So step alone can decide that an eq/neq branch is never taken,
which seems to differ from the commit message statement "Branch (eq/neq)
checks do not take into account base/step yet".

explore_alu_limits is set by sanitize_ptr_alu() once a program loaded
without CAP_PERFMON (bypass_spec_v1 == false) does ptr += <non-const
scalar>.  After that, can this sequence be wrongly pruned?

    call bpf_get_prandom_u32
    r6 = r0
    if r6 & 1 goto l2
    r0 &= 3
    r0 *= 3            ; var_off=(0;0xf) r64=[0,9] base=0 step=3
    goto j
  l2:
    r0 &= 15
    if r0 > 9 goto out ; var_off=(0;0xf) r64=[0,9] base=0 step=1
  j:
    if r0 == 1 goto bad

The fallthrough path is explored first and its state is saved at j.  At
"if r0 == 1" simulate_both_branches_taken() intersects r64 with [1,1], and
deduce_bounds_64_from_step() turns that into an empty range, so the taken
branch is judged dead and bad is never verified for that state.

On the l2 path r0 arrives at j with byte-identical type, delta, raw,
var_off, r64 and r32 (tnum_mul((0;3),3) is (0;0xf), same as
tnum_range(0,9)).  Only step differs, so the explore_alu_limits memcmp
matches and the state is pruned.  At runtime r0 == 1 is possible on that
path, so the never-verified code at bad can run.

Without explore_alu_limits this is fine, since range_within() rejects
cur->step == 1 against old->step == 3.

regs_exact() has the same gap:

kernel/bpf/states.c:regs_exact() {
	return memcmp(rold, rcur, offsetof(struct bpf_reg_state, id)) == 0 &&
	       check_ids(rold->id, rcur->id, idmap) &&
    ...
}

It is used for EXACT comparisons and for the default pointer types in
regsafe().  scalars_exact_for_widen() ignores base/step in the same way.

base and step are still after id at the end of the series (2c96fb13421f).
The later commit 90d24cd54ac0 only switches states_maybe_looping() to
offsetofend(step), and nothing updates the explore_alu_limits memcmp or
regs_exact().

Should base/step be moved ahead of id (after r32), or compared explicitly in
the explore_alu_limits path and in regs_exact()?

>  	/* if (!precise && SCALAR_VALUE) min/max/tnum don't affect safety */
>  	bool precise;
>  };
> diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
> index d850a7863d2e..900d1bb1988b 100644
> --- a/kernel/bpf/log.c
> +++ b/kernel/bpf/log.c
> @@ -692,6 +692,8 @@ static void print_reg_state(struct bpf_verifier_env *env,
>
>  			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
>  			verbose_a("var_off=%s", tn_buf);
> +			if (reg->base != 0 || reg->step != 1)
> +				verbose_a("step=%d+%d", reg->base, reg->step);

Is it intentional that the step print sits inside the
!tnum_is_unknown(reg->var_off) block?

Whether base/step carry information does not depend on var_off.  A scalar
with an odd step and a signed range that crosses zero has a completely
unknown var_off, so its step is never printed.

For example, after "r0 s>>= 56" r64 is [-128,127] with an unknown tnum.
Then "r0 *= 3" makes scalar_step_scale() set base=0, step=3 (no overflow),
tnum_mul() keeps var_off unknown, and __reg_bound_offset() clears no bits
because count_trailing_zeros(3) is 0.  The log shows:

    scalar(smin=smin32=-384,smax=smax32=381)

with no step=0+3, although step affects branch outcomes (through
simulate_both_branches_taken()) and pruning (through range_within()).

Two states can print identically and still fail to prune each other, and an
"infinite loop detected" dump would hide the difference.

Would moving the step print out of the var_off block, right after
print_scalar_ranges(), be better?

>  		}
>  	}
>  	verbose(env, ")");
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 4dd71bf7c1d9..f71778a75f30 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -1910,12 +1911,128 @@ static void bpf_diag_record_caller_saved(struct bpf_verifier_env *env,

[ ... ]

> +static void scalar_step_scale(struct bpf_reg_state *dst_reg, u64 amount)
> +{
> +	u16 step;
> +	s64 tmp;
> +

[ ... ]

> +	if (amount == 0 || check_mul_overflow(dst_reg->step, amount, &step) ||
> +	    check_mul_overflow(reg_smin(dst_reg), (s64)amount, &tmp) ||
> +	    check_mul_overflow(reg_smax(dst_reg), (s64)amount, &tmp)) {
> +		reg_step_reset(dst_reg);
> +		return;
> +	}
> +
> +	dst_reg->base = (dst_reg->base * amount) % step;
> +	dst_reg->step = step;
> +}

Will this link on 32-bit architectures?  amount is a u64, so
(dst_reg->base * amount) % step is a u64 modulo with a non-constant divisor.
The compiler lowers that to a libgcc helper, __umoddi3 on i386 or
__aeabi_uldivmod on ARM32, and the kernel does not provide either one.

verifier.o is built whenever CONFIG_BPF_SYSCALL=y, which is set in ARM32
defconfigs such as aspeed_g5_defconfig and imx_v6_v7_defconfig, so those
builds look like they would fail at link time.  The rest of verifier.c uses
div64_u64()/div64_s64() for this reason.

The modulo also looks redundant.  The check_mul_overflow() calls above
guarantee that step == old_step * amount fits in u16, and base < old_step,
so base * amount < step always holds.

Could this become:

    dst_reg->base = dst_reg->base * amount;

or use div_u64_rem() if the modulo is kept?  The code is unchanged at the
end of the series (2c96fb13421f).


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers
  2026-10-04 13:37 ` [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
@ 2026-10-04 15:11   ` bot+bpf-ci
  0 siblings, 0 replies; 69+ messages in thread
From: bot+bpf-ci @ 2026-10-04 15:11 UTC (permalink / raw)
  To: eddyz87, bpf, ast
  Cc: andrii, daniel, kernel-team, yonghong.song, eddyz87, ast, andrii,
	daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4063 bytes --]

> bpf: varying offset access support for PTR_TO_BTF_ID pointers

TL;DR: Variable-offset BTF access may now reach allocated objects for writes
too; if a union overlays an array with a special field (spin lock, kptr), the
verifier may allow direct stores to it.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index f71778a75f30..27a53204696f 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -6536,36 +6537,31 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
>  		return -EINVAL;
>  	}
>  
> -	if (!tnum_is_const(reg->var_off)) {
> -		char tn_buf[48];
> -
> -		tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
> +	if (check_add_overflow(reg_smin(reg), off, &min_off)) {
>  		verbose(env,
> -			"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
> -			reg_arg_name(env, argno), tname, off, tn_buf);
> +			"%s is ptr_%s access, offset computation overflows: register's minimal offset is %lld, instruction offset is %d\n",
> +			reg_arg_name(env, argno), tname, reg_smin(reg), off);

[ ... ]

> @@ -6575,6 +6571,18 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
>  	}
>  
>  	if (env->ops->btf_struct_access && !type_is_alloc(reg->type) && atype == BPF_WRITE) {
> +		if (!tnum_is_const(reg->var_off)) {
> +			char tn_buf[48];
> +
> +			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
> +			verbose(env,
> +				"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
> +				reg_arg_name(env, argno), tname, off, tn_buf);
> +			return -EACCES;
> +		}
> +
> +		off += reg->var_off.value;
> +
>  		if (!btf_is_kernel(reg->btf)) {
>  			verifier_bug(env, "reg->btf must be kernel btf");
>  			return -EFAULT;

Can BPF_WRITE accesses to allocated objects now reach btf_struct_access()
with a non-constant var_off?

The old check at the top of check_ptr_to_btf_access() rejected every
non-constant var_off.  After this change the rejection only remains inside
the branch guarded by:

    if (env->ops->btf_struct_access && !type_is_alloc(reg->type) &&
        atype == BPF_WRITE)

For a PTR_TO_BTF_ID with MEM_ALLOC (type_is_ptr_alloc_obj()), that branch is
skipped, so the else path runs and the access goes on to btf_struct_access()
with a variable offset for both reads and writes.  The commit message only
talks about relaxing reads.

In btf_struct_access(), the special field check for allocated types is only
done at off = min_off:

    for (i = 0; i < rec->cnt; i++) {
            struct btf_field *field = &rec->fields[i];
            u32 offset = field->offset;
            if (off < offset + field->size && offset < off + size) {
                    bpf_log(log,
                            "direct access to %s is disallowed\n",
                            btf_field_type_name(field->type));
                    return -EACCES;
            }
    }

The code added under check_variable_offset: then only verifies that
[min_off, max_off + size) stays within one of the recorded arrays and that
reg->step is a multiple of the element size.  It does not look at the
btf_record fields again for offsets above min_off.

If a union overlays a scalar array with a bpf_spin_lock, a list or rbtree
node or head, a refcount, or a kptr (btf_find_struct_field() descends into
unions, so such a record is accepted), could an access at min_off pass the
overlap check above while a later offset in [min_off, max_off] lands on the
special field?  For a write, that would be a direct store into a field which
the constant offset check rejects.

Should check_ptr_to_btf_access() keep rejecting a non-constant var_off for
writes to allocated objects, or should the btf_record overlap check in
btf_struct_access() cover the whole [min_off, max_off + size) range?  If
writes are meant to be relaxed too, the commit message should say so.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37207249098

^ permalink raw reply	[flat|nested] 69+ messages in thread

* [syzbot ci] Re: bpf: use scalar evolution to widen bounded loops
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (42 preceding siblings ...)
  2026-10-04 13:38 ` [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening Eduard Zingerman
@ 2026-10-04 20:04 ` syzbot ci
  2026-10-06 16:40 ` [PATCH bpf-next v2 00/43] " patchwork-bot+netdevbpf
  44 siblings, 0 replies; 69+ messages in thread
From: syzbot ci @ 2026-10-04 20:04 UTC (permalink / raw)
  To: andrii, ast, bpf, daniel, eddyz87, emil, ezingerman, kernel-team,
	yonghong.song
  Cc: syzbot, syzkaller-bugs

syzbot ci has tested the following series

[v2] bpf: use scalar evolution to widen bounded loops
https://lore.kernel.org/all/20261004-scev-minimal-rebase-v2-0-f6254e29fdf2@gmail.com
* [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info
* [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses with arg_access_info
* [PATCH bpf-next v2 03/43] bpf: describe kfunc stack accesses with arg_access_info
* [PATCH bpf-next v2 04/43] bpf: track may_write flags in liveness
* [PATCH bpf-next v2 05/43] bpf: summarize may write stack slots in insn_aux_data
* [PATCH bpf-next v2 06/43] bpf: summarize live stack slots in insn_aux_data
* [PATCH bpf-next v2 07/43] bpf: summarize regs that may hold a frame pointer in insn_aux_data
* [PATCH bpf-next v2 08/43] bpf: record write effects for atomic operations in liveness.c
* [PATCH bpf-next v2 09/43] bpf: add tnum_alignment()
* [PATCH bpf-next v2 10/43] bpf: add cnum{32,64}_union()
* [PATCH bpf-next v2 11/43] bpf: add cnum64_intersect_linear()
* [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations
* [PATCH bpf-next v2 13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe()
* [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps
* [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers
* [PATCH bpf-next v2 16/43] bpf: save DFS postorder numbers for program instructions
* [PATCH bpf-next v2 17/43] bpf: move the live-register and SCC printout to a standalone function
* [PATCH bpf-next v2 18/43] bpf: compute immediate dominators
* [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy
* [PATCH bpf-next v2 20/43] bpf: add bpf_set_reg_range()
* [PATCH bpf-next v2 21/43] bpf: add bpf_mark_reg_known_scalar()
* [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union()
* [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists
* [PATCH bpf-next v2 24/43] bpf: record basic-block ends in insn_aux_data
* [PATCH bpf-next v2 25/43] bpf: add bpf_split_cur_state()
* [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin()
* [PATCH bpf-next v2 27/43] bpf: allow precision backtracking between overlapping checkpoints
* [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops
* [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops
* [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking
* [PATCH bpf-next v2 31/43] bpf: bpf_func_state size optimization
* [PATCH bpf-next v2 32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines
* [PATCH bpf-next v2 33/43] selftests/bpf: bound UNIX socket path loops by sun_path size
* [PATCH bpf-next v2 34/43] selftests/bpf: test for stack-pointer subrange pruning
* [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking
* [PATCH bpf-next v2 36/43] selftests/bpf: tests for may_def marks of atomic RMW operations
* [PATCH bpf-next v2 37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes()
* [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic
* [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning
* [PATCH bpf-next v2 40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
* [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation
* [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation
* [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening

and found the following issue:
general protection fault in bpf_compute_scev

Full report is available here:
https://ci.syzbot.org/series/ff81cb40-7952-4f8a-b785-1f9e875f9b1f

***

general protection fault in bpf_compute_scev

tree:      bpf-next
URL:       https://kernel.googlesource.com/pub/scm/linux/kernel/git/bpf/bpf-next.git
base:      99dc1ba542420db6b8df209744f55cc52466ad91
arch:      amd64
compiler:  Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
config:    https://ci.syzbot.org/builds/a6050598-290a-4e4f-bba8-55721490a3a8/config
syz repro: https://ci.syzbot.org/findings/63885d0d-bef7-464a-9e39-0d57d15cba85/syz_repro

Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000: 0000 [#1] SMP KASAN PTI
KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
CPU: 0 UID: 0 PID: 5853 Comm: syz.1.18 Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
RIP: 0010:compute_insn_scevs kernel/bpf/scev.c:1489 [inline]
RIP: 0010:bpf_compute_scev+0x5482/0x6050 kernel/bpf/scev.c:1569
Code: 24 08 30 04 00 00 49 81 c7 30 04 00 00 45 31 ed 48 8b 44 24 10 4e 8d 34 a8 4c 89 f0 48 c1 e8 03 48 b9 00 00 00 00 00 fc ff df <0f> b6 04 08 84 c0 0f 85 af 00 00 00 41 8b 16 48 89 df 31 f6 48 8b
RSP: 0018:ffffc9000372f6e0 EFLAGS: 00010247
RAX: 0000000000000000 RBX: ffff8881144bad00 RCX: dffffc0000000000
RDX: 0000000000000000 RSI: 0000000000000011 RDI: 0000000000000011
RBP: ffffc9000372f910 R08: ffff888101bd0000 R09: 0000000000000004
R10: 000000000000010f R11: 0000000000000000 R12: 0000000000000011
R13: 0000000000000000 R14: 0000000000000004 R15: 0000000000000430
FS:  00007fd0dc1f56c0(0000) GS:ffff88818d6b6000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007fd0db270050 CR3: 00000001bf29a000 CR4: 00000000000006f0
Call Trace:
 <TASK>
 bpf_check+0x2fbe/0x3550 kernel/bpf/verifier.c:23332
 bpf_prog_load+0x1594/0x1d60 kernel/bpf/syscall.c:3177
 __sys_bpf+0xd86/0xe00 kernel/bpf/syscall.c:6437
 __do_sys_bpf kernel/bpf/syscall.c:6559 [inline]
 __se_sys_bpf kernel/bpf/syscall.c:6556 [inline]
 __x64_sys_bpf+0xba/0xd0 kernel/bpf/syscall.c:6556
 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
 do_syscall_64+0x166/0x520 arch/x86/entry/syscall_64.c:84
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fd0db39e159
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 e8 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007fd0dc1f5028 EFLAGS: 00000246 ORIG_RAX: 0000000000000141
RAX: ffffffffffffffda RBX: 00007fd0db625fa0 RCX: 00007fd0db39e159
RDX: 0000000000000094 RSI: 0000200000000040 RDI: 0000000000000005
RBP: 00007fd0db43506b R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007fd0db626038 R14: 00007fd0db625fa0 R15: 00007fff2dcab378
 </TASK>
Modules linked in:
---[ end trace 0000000000000000 ]---
RIP: 0010:compute_insn_scevs kernel/bpf/scev.c:1489 [inline]
RIP: 0010:bpf_compute_scev+0x5482/0x6050 kernel/bpf/scev.c:1569
Code: 24 08 30 04 00 00 49 81 c7 30 04 00 00 45 31 ed 48 8b 44 24 10 4e 8d 34 a8 4c 89 f0 48 c1 e8 03 48 b9 00 00 00 00 00 fc ff df <0f> b6 04 08 84 c0 0f 85 af 00 00 00 41 8b 16 48 89 df 31 f6 48 8b
RSP: 0018:ffffc9000372f6e0 EFLAGS: 00010247
RAX: 0000000000000000 RBX: ffff8881144bad00 RCX: dffffc0000000000
RDX: 0000000000000000 RSI: 0000000000000011 RDI: 0000000000000011
RBP: ffffc9000372f910 R08: ffff888101bd0000 R09: 0000000000000004
R10: 000000000000010f R11: 0000000000000000 R12: 0000000000000011
R13: 0000000000000000 R14: 0000000000000004 R15: 0000000000000430
FS:  00007fd0dc1f56c0(0000) GS:ffff88818d6b6000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007fd0db270050 CR3: 00000001bf29a000 CR4: 00000000000006f0
----------------
Code disassembly (best guess), 4 bytes skipped:
   0:	00 00                	add    %al,(%rax)
   2:	49 81 c7 30 04 00 00 	add    $0x430,%r15
   9:	45 31 ed             	xor    %r13d,%r13d
   c:	48 8b 44 24 10       	mov    0x10(%rsp),%rax
  11:	4e 8d 34 a8          	lea    (%rax,%r13,4),%r14
  15:	4c 89 f0             	mov    %r14,%rax
  18:	48 c1 e8 03          	shr    $0x3,%rax
  1c:	48 b9 00 00 00 00 00 	movabs $0xdffffc0000000000,%rcx
  23:	fc ff df
* 26:	0f b6 04 08          	movzbl (%rax,%rcx,1),%eax <-- trapping instruction
  2a:	84 c0                	test   %al,%al
  2c:	0f 85 af 00 00 00    	jne    0xe1
  32:	41 8b 16             	mov    (%r14),%edx
  35:	48 89 df             	mov    %rbx,%rdi
  38:	31 f6                	xor    %esi,%esi
  3a:	48                   	rex.W
  3b:	8b                   	.byte 0x8b


***

If these findings have caused you to resend the series or submit a
separate fix, please add the following tag to your commit message:
  Tested-by: syzbot@syzkaller.appspotmail.com

---
This report is generated by a bot. It may contain errors.
syzbot ci engineers can be reached at syzkaller@googlegroups.com.

To test a fix for this bug, please reply with `#syz test`
(on a separate line) and attach the patch to the email.

Notes:
- The patch will be applied on top of the tested series (as an
  incremental fix).
- To test a new version of the whole series, please send it directly
  to syzbot@lists.linux.dev.
- Arguments like custom git repos and branches are not supported.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops
  2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (43 preceding siblings ...)
  2026-10-04 20:04 ` [syzbot ci] Re: bpf: use scalar evolution to widen bounded loops syzbot ci
@ 2026-10-06 16:40 ` patchwork-bot+netdevbpf
  2026-10-06 16:45   ` Alexei Starovoitov
  44 siblings, 1 reply; 69+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-06 16:40 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf, ast, andrii, daniel, kernel-team, yonghong.song

Hello:

This series was applied to bpf/bpf-next.git (master)
by Alexei Starovoitov <ast@kernel.org>:

On Sun,  4 Oct 2026 06:37:42 -0700 you wrote:
> This series implements the scalar evolution (SCEV) technique for
> verification of loops.
> 
> Scalar evolution is a static analysis technique that infers algebraic
> expressions describing how variables change within a loop body.
> These expressions can then be used to estimate the number of loop body
> executions. This estimate can be used to represent induction variables
> as ranges instead of enumerating each possible value.
> 
> [...]

Here is the summary with links:
  - [bpf-next,v2,01/43] bpf: represent stack access effects with arg_access_info
    https://git.kernel.org/bpf/bpf-next/c/0672c5556244
  - [bpf-next,v2,02/43] bpf: describe helper stack accesses with arg_access_info
    https://git.kernel.org/bpf/bpf-next/c/32aa356d6703
  - [bpf-next,v2,03/43] bpf: describe kfunc stack accesses with arg_access_info
    https://git.kernel.org/bpf/bpf-next/c/9aa0211998a7
  - [bpf-next,v2,04/43] bpf: track may_write flags in liveness
    https://git.kernel.org/bpf/bpf-next/c/8de0510cdb1c
  - [bpf-next,v2,05/43] bpf: summarize may write stack slots in insn_aux_data
    https://git.kernel.org/bpf/bpf-next/c/ccd41f1e7257
  - [bpf-next,v2,06/43] bpf: summarize live stack slots in insn_aux_data
    https://git.kernel.org/bpf/bpf-next/c/2ff0e602d931
  - [bpf-next,v2,07/43] bpf: summarize regs that may hold a frame pointer in insn_aux_data
    https://git.kernel.org/bpf/bpf-next/c/440ad035bbda
  - [bpf-next,v2,08/43] bpf: record write effects for atomic operations in liveness.c
    https://git.kernel.org/bpf/bpf-next/c/c695981052dc
  - [bpf-next,v2,09/43] bpf: add tnum_alignment()
    https://git.kernel.org/bpf/bpf-next/c/61f39d3292ca
  - [bpf-next,v2,10/43] bpf: add cnum{32,64}_union()
    https://git.kernel.org/bpf/bpf-next/c/7372e0fd2035
  - [bpf-next,v2,11/43] bpf: add cnum64_intersect_linear()
    https://git.kernel.org/bpf/bpf-next/c/575fbdd131db
  - [bpf-next,v2,12/43] bpf: expose comparison opcode transformations
    https://git.kernel.org/bpf/bpf-next/c/f244cbfba3ea
  - [bpf-next,v2,13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe()
    https://git.kernel.org/bpf/bpf-next/c/7ddae43ac5e7
  - [bpf-next,v2,14/43] bpf: representation for intervals with steps
    https://git.kernel.org/bpf/bpf-next/c/6499523b1636
  - [bpf-next,v2,15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers
    https://git.kernel.org/bpf/bpf-next/c/17cd31244f50
  - [bpf-next,v2,16/43] bpf: save DFS postorder numbers for program instructions
    https://git.kernel.org/bpf/bpf-next/c/6bab25c91a66
  - [bpf-next,v2,17/43] bpf: move the live-register and SCC printout to a standalone function
    https://git.kernel.org/bpf/bpf-next/c/ba536c460daa
  - [bpf-next,v2,18/43] bpf: compute immediate dominators
    https://git.kernel.org/bpf/bpf-next/c/695c19b3d36e
  - [bpf-next,v2,19/43] bpf: compute loop hierarchy
    https://git.kernel.org/bpf/bpf-next/c/62083ae4dd10
  - [bpf-next,v2,20/43] bpf: add bpf_set_reg_range()
    https://git.kernel.org/bpf/bpf-next/c/651f832d9b06
  - [bpf-next,v2,21/43] bpf: add bpf_mark_reg_known_scalar()
    https://git.kernel.org/bpf/bpf-next/c/8767ffd4e718
  - [bpf-next,v2,22/43] bpf: add bpf_reg_union()
    https://git.kernel.org/bpf/bpf-next/c/297788154c13
  - [bpf-next,v2,23/43] bpf: add a min-heap for ordered analysis worklists
    https://git.kernel.org/bpf/bpf-next/c/5048674597c7
  - [bpf-next,v2,24/43] bpf: record basic-block ends in insn_aux_data
    https://git.kernel.org/bpf/bpf-next/c/141bd53be1c8
  - [bpf-next,v2,25/43] bpf: add bpf_split_cur_state()
    https://git.kernel.org/bpf/bpf-next/c/0547bc753c30
  - [bpf-next,v2,26/43] bpf: add bpf_same_memory_origin()
    https://git.kernel.org/bpf/bpf-next/c/29a555f49246
  - [bpf-next,v2,27/43] bpf: allow precision backtracking between overlapping checkpoints
    https://git.kernel.org/bpf/bpf-next/c/57eb802c24bf
  - [bpf-next,v2,28/43] bpf: compute scalar evolution expressions for loops
    https://git.kernel.org/bpf/bpf-next/c/82bcc4f665cb
  - [bpf-next,v2,29/43] bpf: use SCEV to widen bounded loops
    https://git.kernel.org/bpf/bpf-next/c/408aff9a0159
  - [bpf-next,v2,30/43] bpf: avoid widening registers that hinder exact stack-slot tracking
    https://git.kernel.org/bpf/bpf-next/c/a6474f5f8464
  - [bpf-next,v2,31/43] bpf: bpf_func_state size optimization
    https://git.kernel.org/bpf/bpf-next/c/422ffa508a7b
  - [bpf-next,v2,32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines
    https://git.kernel.org/bpf/bpf-next/c/2db7911e0e7f
  - [bpf-next,v2,33/43] selftests/bpf: bound UNIX socket path loops by sun_path size
    https://git.kernel.org/bpf/bpf-next/c/6a697681f8ab
  - [bpf-next,v2,34/43] selftests/bpf: test for stack-pointer subrange pruning
    https://git.kernel.org/bpf/bpf-next/c/a8606df987d1
  - [bpf-next,v2,35/43] selftests/bpf: tests for may_write stack-liveness tracking
    https://git.kernel.org/bpf/bpf-next/c/95cf9a876448
  - [bpf-next,v2,36/43] selftests/bpf: tests for may_def marks of atomic RMW operations
    https://git.kernel.org/bpf/bpf-next/c/64a98c9510a0
  - [bpf-next,v2,37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes()
    https://git.kernel.org/bpf/bpf-next/c/22a9679beb90
  - [bpf-next,v2,38/43] selftests/bpf: tests for register base/step arithmetic
    https://git.kernel.org/bpf/bpf-next/c/a7d81f0f457d
  - [bpf-next,v2,39/43] selftests/bpf: tests for register base/step state pruning
    https://git.kernel.org/bpf/bpf-next/c/ff12031da64f
  - [bpf-next,v2,40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
    https://git.kernel.org/bpf/bpf-next/c/295ac48d3de6
  - [bpf-next,v2,41/43] selftests/bpf: tests for loop hierarchy computation
    https://git.kernel.org/bpf/bpf-next/c/09b89711604a
  - [bpf-next,v2,42/43] selftests/bpf: tests for immediate dominator computation
    https://git.kernel.org/bpf/bpf-next/c/6f7651bd3f68
  - [bpf-next,v2,43/43] selftests/bpf: tests for SCEV analysis and loop widening
    https://git.kernel.org/bpf/bpf-next/c/ca78d19fd106

You are awesome, thank you!
-- 
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html



^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops
  2026-10-06 16:40 ` [PATCH bpf-next v2 00/43] " patchwork-bot+netdevbpf
@ 2026-10-06 16:45   ` Alexei Starovoitov
  0 siblings, 0 replies; 69+ messages in thread
From: Alexei Starovoitov @ 2026-10-06 16:45 UTC (permalink / raw)
  To: patchwork-bot+netdevbpf
  Cc: Eduard Zingerman, bpf, Alexei Starovoitov, Andrii Nakryiko,
	Daniel Borkmann, Kernel Team, Yonghong Song

On Tue, Oct 6, 2026 at 9:40 AM <patchwork-bot+netdevbpf@kernel.org> wrote:
>
> Hello:
>
> This series was applied to bpf/bpf-next.git (master)
> by Alexei Starovoitov <ast@kernel.org>:

and reverted.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists
  2026-10-04 13:38 ` [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
  2026-10-04 13:47   ` sashiko-bot
@ 2026-10-06 16:53   ` Alexei Starovoitov
  1 sibling, 0 replies; 69+ messages in thread
From: Alexei Starovoitov @ 2026-10-06 16:53 UTC (permalink / raw)
  To: Eduard Zingerman, bpf; +Cc: andrii, daniel, kernel-team, yonghong.song

On Sun, Oct 04, 2026 at 06:38 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> SCEV needs to process pending basic blocks in reverse postorder.
> Provide a small integer min-heap whose comparator can order instruction
> indices by their CFG ranks.

Why not include/linux/min_heap.h ?
less() there takes 'void *args'.
worklist_push() checks discovered[] first, so an insn is pushed once
per loop. prog->len ints allocated in bpf_init_scev() are enough
and the push cannot fail.

^ permalink raw reply	[flat|nested] 69+ messages in thread

end of thread, other threads:[~2026-10-06 16:53 UTC | newest]

Thread overview: 69+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04 13:37 [PATCH bpf-next v2 00/43] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 01/43] bpf: represent stack access effects with arg_access_info Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:37 ` [PATCH bpf-next v2 02/43] bpf: describe helper stack accesses " Eduard Zingerman
2026-10-04 14:02   ` sashiko-bot
2026-10-04 13:37 ` [PATCH bpf-next v2 03/43] bpf: describe kfunc " Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 04/43] bpf: track may_write flags in liveness Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 05/43] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 06/43] bpf: summarize live " Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 07/43] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 08/43] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 09/43] bpf: add tnum_alignment() Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 10/43] bpf: add cnum{32,64}_union() Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 11/43] bpf: add cnum64_intersect_linear() Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 12/43] bpf: expose comparison opcode transformations Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:37 ` [PATCH bpf-next v2 13/43] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 14/43] bpf: representation for intervals with steps Eduard Zingerman
2026-10-04 14:02   ` sashiko-bot
2026-10-04 14:40   ` bot+bpf-ci
2026-10-04 13:37 ` [PATCH bpf-next v2 15/43] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
2026-10-04 15:11   ` bot+bpf-ci
2026-10-04 13:37 ` [PATCH bpf-next v2 16/43] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
2026-10-04 13:37 ` [PATCH bpf-next v2 17/43] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 18/43] bpf: compute immediate dominators Eduard Zingerman
2026-10-04 13:50   ` sashiko-bot
2026-10-04 13:38 ` [PATCH bpf-next v2 19/43] bpf: compute loop hierarchy Eduard Zingerman
2026-10-04 14:40   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 20/43] bpf: add bpf_set_reg_range() Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 21/43] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 22/43] bpf: add bpf_reg_union() Eduard Zingerman
2026-10-04 13:59   ` sashiko-bot
2026-10-04 13:38 ` [PATCH bpf-next v2 23/43] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
2026-10-04 13:47   ` sashiko-bot
2026-10-06 16:53   ` Alexei Starovoitov
2026-10-04 13:38 ` [PATCH bpf-next v2 24/43] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 25/43] bpf: add bpf_split_cur_state() Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 26/43] bpf: add bpf_same_memory_origin() Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 27/43] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 28/43] bpf: compute scalar evolution expressions for loops Eduard Zingerman
2026-10-04 14:00   ` sashiko-bot
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 29/43] bpf: use SCEV to widen bounded loops Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 30/43] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
2026-10-04 14:02   ` sashiko-bot
2026-10-04 13:38 ` [PATCH bpf-next v2 31/43] bpf: bpf_func_state size optimization Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 32/43] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 33/43] selftests/bpf: bound UNIX socket path loops by sun_path size Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 34/43] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 35/43] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 36/43] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 37/43] selftests/bpf: tests for map special cases in bpf_helper_stack_access_bytes() Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 38/43] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 39/43] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
2026-10-04 13:54   ` sashiko-bot
2026-10-04 13:38 ` [PATCH bpf-next v2 40/43] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
2026-10-04 13:38 ` [PATCH bpf-next v2 41/43] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 42/43] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 13:38 ` [PATCH bpf-next v2 43/43] selftests/bpf: tests for SCEV analysis and loop widening Eduard Zingerman
2026-10-04 14:24   ` bot+bpf-ci
2026-10-04 20:04 ` [syzbot ci] Re: bpf: use scalar evolution to widen bounded loops syzbot ci
2026-10-06 16:40 ` [PATCH bpf-next v2 00/43] " patchwork-bot+netdevbpf
2026-10-06 16:45   ` Alexei Starovoitov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox