BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops
@ 2026-09-26 14:19 Eduard Zingerman
  2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
                   ` (35 more replies)
  0 siblings, 36 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:19 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

This series implements the scalar evolution (SCEV) technique for
verification of loops.

Scalar evolution is a static analysis technique that infers algebraic
expressions describing how variables change within a loop body.
These expressions can then be used to estimate the number of loop body
executions. This estimate can be used to represent induction variables
as ranges instead of enumerating each possible value.

For example, the following loop:

  for (r0 = 0; r0 < 10; r0++) {
    ...
  }

would now be verified with the assumption that r0 is a scalar value
in the range [0..9] within its body.

See the following patches for a more technical overview:
- "bpf: compute scalar evolution expressions for loops"
- "bpf: use SCEV to widen bounded loops"

The series can be viewed as consisting of the following parts:
- Preparatory patches extending liveness analysis to collect
  additional information and adding various utility functions.
- Patches relaxing the verifier's current restrictions on accessing
  memory via varying offset pointers, specifically:
  - "bpf: allow subrange relations for PTR_TO_STACK in regsafe()"
  - "bpf: representation for intervals with steps"
  - "bpf: varying offset access support for PTR_TO_BTF_ID pointers"
- Patches to compute the immediate dominator tree and loop hierarchy.
- Patches implementing scalar evolution and integrating it with
  the main verification pass:
  - "bpf: avoid widening registers that hinder exact stack-slot tracking"
  - "bpf: use SCEV to widen bounded loops"
  - "bpf: compute scalar evolution expressions for loops"
- Tests.

Literature
==========

- "Symbolic Evaluation of Chains of Recurrences for Loop Optimization"
  Robert A. van Engelen, 2000
- "The CR# Algebra and its Application in Loop Analysis and Optimization"
  Robert A. van Engelen, 2004

Supported patterns
==================

As scalar evolution analysis attempts to infer recurring relationships
from algebraic expressions describing loop variables, there are some
limitations on what can be expressed.

Below is a list of patterns that are expected to work with the current
implementation. The examples are pseudo-code for the indicated BPF
instruction shapes.

1. Counted loops, including decreasing counters:

       u64 i = 0;
       do { i += 1; } while (i < 10);

       u64 j = 3;
       do { j += -1; } while (j != 0);

   Header ranges are i in [0,9] and j in [1,3].
   Pre-condition loops, such as while (i < 10) { i += 1; },
   are also supported. The counter can reside in a register or a fixed
   8-byte frame-pointer spill.

2. Strided accesses into a map value:

       /* bytes is a map-value array with at least 16 bytes. */
       u64 i = 0, off = 0;
       do {
           bytes[off] = 1;
           i += 1;
           off += 2;
       } while (i < 8);

   At the header, off is in [0,14] with step 2. The byte store stays
   within the map value even though i and off are tracked separately.

3. A nonconstant entry value with a separate counted induction variable:

       u64 x = bpf_get_prandom_u32() & 6;  // {0,2,4,6}
       u64 i = 0;
       do { x += 6; i += 1; } while (i < 3);

   The counter i supplies the bound. At the header, x widens to [0,18]
   with step 2, retaining the alignment common to its entry values and
   its slope; x itself need not start at a constant.

4. Conditional assignment from an invariant value:

       u64 x = 5;
       for (u64 i = 0; i < 3; i += 1)
           if (i == 1) x = 10;
       return x;

   Widening joins the pre-loop value 5 with the assigned value 10,
   giving x in [5,10] both inside and after the loop.

5. Early exits in addition to the counted backedge:

       u64 i = 0;
       do {
           if (bpf_get_prandom_u32() == 0) break;
           i += 1;
       } while (i != 3);

   The counter still bounds the loop. An additional exit reduces the
   available lower-bound information; it does not by itself prevent
   widening.

6. Array accesses through different memory types:

   Bounded varying offsets are supported for arrays in BTF-typed
   objects, stack memory, map values and keys, packet data and
   metadata, and memory buffers. The usual bounds, alignment and
   access restrictions still apply.

   For stack arrays indexed by a widened loop counter, only byte and
   half-word accesses are supported; see limitation 5 below.

   For an array in a BTF-typed object:

       struct inner { int a, b; };
       struct object { struct inner arr[8]; };
       /* o points to a live, non-NULL BTF-typed struct object. */
       u64 i = bpf_get_prandom_u32() & 7;
       value = o->arr[i].b;

   The 4-byte load has offsets 4 + 8*i, i in [0,7]. Bounds keep
   accesses within arr; base and step 8 ensure that each possible
   offset selects member b. Such accesses also work without a loop.

7. Nested loops with separate counters:

       u64 i = 0;
       do {
           u64 j = 0;
           do {
               /* body */
               j += 1;
           } while (j < 3);
           i += 1;
       } while (i < 4);

   Both loops can be widened: i is in [0,3] at the outer header,
   and j is in [0,2] at the inner header. The inner counter is reset
   on each outer iteration; i is unchanged by the inner loop.

8. Incrementing and comparing a pointer directly:

       /* bytes is a map-value array with at least 8 bytes. */
       u8 *p = bytes, *end = bytes + 8;
       do {
           *p = 1;
           p++;
       } while (p != end);

   The initial pointer and fixed end pointer give eight iterations,
   without a separate integer counter. At the loop header, p ranges
   from bytes to bytes + 7, so each byte store stays within the array.

Limitations
===========

Patterns that lose precision after widening
-------------------------------------------

Widening can lose information that the verifier gets by checking
each iteration separately. For example:

    r7 = 5;
    for (r6 = 0; r6 < 3; r6++)
        if (r6 == 1)
            r7 = 10;
    if (r7 != 10)
        invalid_stack_read();

R7 is always 10 at the real exit, but is represented as [5, 10] at
loop entry and after the loop. The verifier therefore cannot exclude
the invalid read and rejects this otherwise safe program.
It would be hard to avoid this limitation.

The same issue affects correlated linear counters at loop exit:

    u64 i = 0, off = 0;
    while (i < 8) {
        i++;
        off += 2;
    }
    if (off != 16)
        invalid_stack_read();

Enumeration establishes off == 16. Widening tracks i and off
independently, and the exit check on i does not reconstruct off's exact
final value. This can be lifted after analysis adjustments.

Patterns not widened yet
------------------------

If one of the variables modified in the loop cannot be widened, the
verifier checks each iteration instead. Examples include:

1. 32-bit recurrences and comparisons:

       for (u32 i = 0; i < 8; i++)      // ALU32 / JMP32
           use(i);

   Only 64-bit arithmetic and comparisons are supported.
   This limitation will be lifted.

2. Non-linear updates:

       for (u64 i = 1; i < 16; i *= 2)  // non-linear recurrence
           use(i);

   Recognition is currently limited to simple additive recurrences.
   Literature describes a way to represent this, but it is not
   considered a priority at the moment.

3. A controlling counter or bound that is not constant at loop entry:

       u64 limit = unknown_in_range(1, 8);
       for (u64 i = 0; i < limit; i++)
           use(i);

       u64 i = unknown_in_range(0, 7);
       while (i < 8)
           i++;

   Both loops have finite bounds, but iteration-count evaluation
   currently requires concrete entry values. Other induction variables
   may have nonconstant entry ranges once a separate counter establishes
   the iteration bound. This limitation can be lifted eventually.

4. Alternative values referring to another evolving register:

       r7 = 0;
       for (r6 = 0; r6 < 4; r6++)
           if (condition)
               r7 = r6;
       use(r7);

   R6 changes on each iteration. Widening currently supports
   conditional assignments only from constants or values that do not
   change in the loop. This limitation can be lifted.

5. Stack addresses whose widening would lose precise slot tracking:

       u64 slots[4];
       for (u64 i = 0; i < 4; i++)
           slots[i] = i;

       struct bpf_dynptr dptrs[4];
       for (u64 i = 0; i < 4; i++)
           bpf_dynptr_from_xdp(ctx, 0, &dptrs[i]);

   Keep such indices concrete for 4- and 8-byte stack loads/stores and
   calls requiring fixed-offset stack objects. Dependencies inside
   nested loops also constrain the outer loop.
   It would be hard to lift this limitation.

6. Values modified by a nested loop do not get a closed-form summary
   for the enclosing loop:

       u64 sum = 0;
       for (u64 i = 0; i < 4; i++)
           for (u64 j = 0; j < 4; j++)
               sum++;
       use(sum);

   The inner loop may widen, but the outer analysis forgets values
   modified by it rather than deriving the combined recurrence.
   This limitation can be lifted.

7. Multiple backedges and irreducible control flow:

       i = 0;
   H:  if (i >= 8) goto done;
       if (condition) { i++; goto H; }  // first backedge
       i++; goto H;                    // second backedge
   done:

       i = 0;
       if (condition) goto body;       // second entry into the loop
   H:  i++;
   body:
       if (i < 8) goto H;

   Widening requires a reducible loop with one backedge and a supported
   dominating exit condition. Multiple exits are supported, but their
   early-exit paths reduce what can be inferred about iteration counts.
   It is unclear whether this can be lifted at the moment.

8. Some equivalent latch forms and large or wrapping recurrences:

       if (bound > i) goto again;      // i is the BPF source operand

       for (u64 i = 0, x = 0; i < 4; i++) {
           use(x);
           x += 65536;                // slope outside signed 16 bits
       }

   Latch matching currently expects the induction variable as the
   destination operand. Widened slopes must be nonzero signed-16-bit
   constants, computed ranges must fit signed 64-bit arithmetic,
   and iteration counts must be representable and finite.

9. Some hard-coded constant limits:
   - Programs with more than 16 nested loops in one function are rejected.
   - Loops with more than 256 exits exceed the metadata limits and are
     not analyzed for widening.
   - Expression traversal is limited to depth 8. Only a small fixed number
     of alternative values can be joined.

Memory consumption
------------------

SCEV retains environments at analyzed block boundaries, including loop
headers and latch conditions, rather than at every straight-line
instruction. Branch-heavy loop bodies can nevertheless require many
environments, even when the loop is ultimately not widened:

    for (u64 i = 0; i < 8; i++) {
        if (p0) x += 1; else x += 2;
        if (p1) x += 3; else x += 4;
        ...                         // many more branches and joins
    }

Each environment uses 2,136 bytes for its register and stack arrays.
A loop with about 500,000 conditional jumps, each skipping one
instruction, can require roughly a million environments:
about 2 GiB for these arrays alone.

Veristat changes
================

A is master and B is this patch-set. The comparison for 7,152
programs: 134 from Cilium, 1,222 from Meta, 375 from sched_ext,
and 5,421 from BPF selftests.

Current impact on selftests is small. Work is in progress to improve
this, for example, reducing pyperf600_nounroll from 460,062 to 1,588
processed instructions. This follow-up work is not included below.

The histogram includes all programs in the corpus, including failed
loads and programs below the table cutoff. The 6,877 unchanged
programs are in the 0..5% bin. Veristat accepts 4,750 programs on A
and 4,758 on B: 10 new acceptances and two new rejections. Both new
rejections are expected by this series' selftests.
These are load verdicts, not test pass/fail counts.

Insns change   Programs
-------------  --------
-100 .. -95 %         8
-95 .. -90 %         10
-90 .. -85 %          8
-85 .. -75 %         15
-75 .. -65 %         17
-65 .. -60 %         32
-60 .. -55 %          7
-55 .. -50 %         16
-50 .. -45 %         23
-45 .. -40 %          8
-40 .. -35 %          6
-35 .. -25 %         10
-25 .. -20 %         11
-20 .. -10 %          8
-10 .. -5 %          55
-5 .. 0 %             9
0 .. 5 %           6898
5 .. 15 %             3
35 .. 40 %            6
40 .. 50 %            1
60 .. 65 %            1

The tables include only programs accepted in both runs with at least
5,000 baseline verifier-processed instructions.

Top wins (50)
-------------

File                                Program                              Insns (A)  Insns (B)       Insns (DIFF)
----------------------------------  -----------------------------------  ---------  ---------  -----------------
verifier_iterating_callbacks.bpf.o  test1                                     7004         12    -6992 (-99.83%)
bpf_iter_tasks.bpf.o                dump_task_sleepable                      85522        556   -84966 (-99.35%)
bpf_iter_task_stack.bpf.o           dump_task_stack                           7209         65    -7144 (-99.10%)
bpf_iter_unix.bpf.o                 dump_unix                                 9109        226    -8883 (-97.52%)
security-monitor-62.bpf.o           lsm_bprm_creds_for_exec                  17538       1701   -15837 (-90.30%)
security-monitor-62.bpf.o           lsm_task_alloc                            5674        565    -5109 (-90.04%)
security-monitor-62.bpf.o           lsm_fo_task_update                       11399       1181   -10218 (-89.64%)
security-monitor-62.bpf.o           lsm_file_open                           154412      21168  -133244 (-86.29%)
security-sandbox-17.bpf.o           ptrace_traceme                           16036       3134   -12902 (-80.46%)
security-sandbox-10.bpf.o           fork                                     30760       6211   -24549 (-79.81%)
security-sandbox-17.bpf.o           ptrace_access_check                      21206       5258   -15948 (-75.21%)
security-sandbox-12.bpf.o           task_kill                                27823       7688   -20135 (-72.37%)
security-monitor-22.bpf.o           syscalls_kill                            97920      27794   -70126 (-71.62%)
security-sandbox-13.bpf.o           kernel_module_request                     7175       2461    -4714 (-65.70%)
security-sandbox-13.bpf.o           kernel_load_data                          7176       2462    -4714 (-65.69%)
security-sandbox-13.bpf.o           kernel_read_file                          7177       2463    -4714 (-65.68%)
security-monitor-32.bpf.o           syscall_setresuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setfsgid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setfsuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setgid                         122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setregid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setreuid                       122082      43187   -78895 (-64.62%)
security-monitor-32.bpf.o           syscalls_setuid                         122082      43187   -78895 (-64.62%)
security-monitor-12.bpf.o           connect_security_socket_connect         216137      77825  -138312 (-63.99%)
security-monitor-16.bpf.o           kernel_modules_do_init                   92863      33648   -59215 (-63.77%)
security-monitor-20.bpf.o           enter_pivot_root                         93400      34193   -59207 (-63.39%)
security-monitor-08.bpf.o           bpf_prog_detect                          93266      34376   -58890 (-63.14%)
security-monitor-31.bpf.o           inode_create                             96265      37055   -59210 (-61.51%)
security-monitor-10.bpf.o           cgroup_mkdir                             96513      37525   -58988 (-61.12%)
security-monitor-03.bpf.o           net_block_bind                           64652      25231   -39421 (-60.97%)
security-monitor-03.bpf.o           net_block_recvmsg                        64652      25231   -39421 (-60.97%)
security-monitor-03.bpf.o           net_block_sendmsg                        64652      25231   -39421 (-60.97%)
security-monitor-04.bpf.o           action_proc_term_sched_process_exec      64739      25299   -39440 (-60.92%)
security-monitor-13.bpf.o           sock_iter                                64818      25350   -39468 (-60.89%)
file-system-monitor-03.bpf.o        proc_file_open                           64883      25437   -39446 (-60.80%)
security-monitor-34.bpf.o           udp_recv_v6                              65089      25623   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_send_v6                              65089      25623   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_recv_v4                              65092      25626   -39466 (-60.63%)
security-monitor-34.bpf.o           udp_send_v4                              65093      25627   -39466 (-60.63%)
security-monitor-18.bpf.o           memfd_create                             65291      25845   -39446 (-60.42%)
security-monitor-17.bpf.o           lsm_file_open                           130665      51773   -78892 (-60.38%)
security-monitor-07.bpf.o           security_socket_listen                   65378      25932   -39446 (-60.34%)
security-monitor-05.bpf.o           accept_security_socket_accept            65435      25989   -39446 (-60.28%)
security-monitor-21.bpf.o           raw_tracepoint__sched_process_exec       65511      26045   -39466 (-60.24%)
security-monitor-06.bpf.o           bash_reader                              65601      26161   -39440 (-60.12%)
security-monitor-30.bpf.o           python3_armor                            65805      26339   -39466 (-59.97%)
security-monitor-32.bpf.o           syscalls_setgroups                       67083      27617   -39466 (-58.83%)
security-monitor-09.bpf.o           fexit_cap_capable                        68221      28162   -40059 (-58.72%)
security-sandbox-31.bpf.o           test_file_open                           15279       6331    -8948 (-58.56%)
security-monitor-03.bpf.o           net_block_init                           71217      31453   -39764 (-55.83%)

Top losses (10)
---------------

File               Program  Insns (A)  Insns (B)    Insns (DIFF)
-----------------  -------  ---------  ---------  --------------
firewall-06.bpf.o  ingress       9323       9624   +301 (+3.23%)
firewall-06.bpf.o  tc_in         9323       9624   +301 (+3.23%)
firewall-08.bpf.o  ingress       9323       9624   +301 (+3.23%)
firewall-08.bpf.o  tc_in         9323       9624   +301 (+3.23%)
firewall-06.bpf.o  egress        9334       9635   +301 (+3.22%)
firewall-08.bpf.o  egress        9334       9635   +301 (+3.22%)
firewall-06.bpf.o  tc_eg         9488       9789   +301 (+3.17%)
firewall-08.bpf.o  tc_eg         9488       9789   +301 (+3.17%)
firewall-05.bpf.o  tc_eg       156665     161269  +4604 (+2.94%)
firewall-07.bpf.o  tc_eg       156665     161269  +4604 (+2.94%)

---
Eduard Zingerman (36):
      bpf: track may_write flags in liveness
      bpf: summarize may write stack slots in insn_aux_data
      bpf: summarize live stack slots in insn_aux_data
      bpf: summarize regs that may hold a frame pointer in insn_aux_data
      bpf: record write effects for atomic operations in liveness.c
      bpf: add tnum_alignment()
      bpf: add cnum{32,64}_union()
      bpf: add cnum64_intersect_linear()
      bpf: add bpf_set_reg_range()
      bpf: add bpf_mark_reg_known_scalar()
      bpf: add bpf_reg_union()
      bpf: expose comparison opcode transformations
      bpf: allow subrange relations for PTR_TO_STACK in regsafe()
      bpf: representation for intervals with steps
      bpf: varying offset access support for PTR_TO_BTF_ID pointers
      bpf: save DFS postorder numbers for program instructions
      bpf: move the live-register and SCC printout to a standalone function
      bpf: compute immediate dominators
      bpf: compute loop hierarchy
      bpf: add a min-heap for ordered analysis worklists
      bpf: record basic-block ends in insn_aux_data
      bpf: add bpf_split_cur_state()
      bpf: allow precision backtracking between overlapping checkpoints
      bpf: compute scalar evolution expressions for loops
      bpf: use SCEV to widen bounded loops
      bpf: avoid widening registers that hinder exact stack-slot tracking
      selftests/bpf: __msg_next tag for matching messages on consecutive lines
      selftests/bpf: test for stack-pointer subrange pruning
      selftests/bpf: tests for may_write stack-liveness tracking
      selftests/bpf: tests for may_def marks of atomic RMW operations
      selftests/bpf: tests for register base/step arithmetic
      selftests/bpf: tests for register base/step state pruning
      selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
      selftests/bpf: tests for loop hierarchy computation
      selftests/bpf: tests for immediate dominator computation
      selftests/bpf: cover SCEV analysis and loop widening

 include/linux/bpf_verifier.h                       |  156 +-
 include/linux/cnum.h                               |    3 +
 include/linux/tnum.h                               |    2 +
 kernel/bpf/Makefile                                |    2 +-
 kernel/bpf/backtrack.c                             |    5 +
 kernel/bpf/btf.c                                   |  149 +-
 kernel/bpf/cfg.c                                   |   56 +-
 kernel/bpf/cnum.c                                  |   36 +
 kernel/bpf/cnum_defs.h                             |   44 +
 kernel/bpf/fixups.c                                |   18 +
 kernel/bpf/heap.c                                  |   87 +
 kernel/bpf/liveness.c                              |  335 ++-
 kernel/bpf/log.c                                   |    7 +
 kernel/bpf/loops.c                                 |  576 +++++
 kernel/bpf/scev.c                                  | 2246 ++++++++++++++++++++
 kernel/bpf/states.c                                |  197 +-
 kernel/bpf/tnum.c                                  |   11 +
 kernel/bpf/verifier.c                              |  514 ++++-
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   10 +
 tools/testing/selftests/bpf/progs/bpf_misc.h       |    6 +
 .../selftests/bpf/progs/verifier_bounds_step.c     |  307 +++
 .../bpf/progs/verifier_btf_array_access.c          |  506 +++++
 tools/testing/selftests/bpf/progs/verifier_gotox.c |    6 +-
 tools/testing/selftests/bpf/progs/verifier_idoms.c |  390 ++++
 .../selftests/bpf/progs/verifier_live_stack.c      |  209 +-
 .../selftests/bpf/progs/verifier_loop_hierarchy.c  |  317 +++
 tools/testing/selftests/bpf/progs/verifier_scev.c  | 1108 ++++++++++
 .../selftests/bpf/progs/verifier_stack_ptr.c       |   33 +
 tools/testing/selftests/bpf/test_loader.c          |   10 +
 29 files changed, 7107 insertions(+), 239 deletions(-)
---
base-commit: 6ac134ec930642d5b574b63a72a0e999c2470c64
change-id: 20260925-scev-minimal-rebase-911799c95dff

^ permalink raw reply	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 01/36] bpf: track may_write flags in liveness
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
@ 2026-09-26 14:19 ` Eduard Zingerman
  2026-09-26 15:51   ` Alexei Starovoitov
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
                   ` (34 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:19 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Track stack slots that instructions may modify, in order to account
for such writes when constructing SCEV expressions.

Compared to must_write, may_write includes any possibly modified slot,
including partial writes, and does not kill liveness.
To accommodate this, record_stack_access_off() is now called from
record_stack_access() with any value of arg->off_cnt. It checks
`arg->off_cnt == 1` before updating must_write.

Merge analyzed instances as:

  may_write(dst)  |= may_write(src)
  must_write(dst) &= must_write(src)

For each ancestor frame f, summarize callee writes at the callsite:

  may_write(callsite, f) |= OR_{i in callee} may_write(i, f)

Imprecise writes mark whole candidate frames.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c                              | 196 +++++++++++++++------
 .../selftests/bpf/progs/verifier_live_stack.c      |  34 ++--
 2 files changed, 165 insertions(+), 65 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..cc3ad75aa1e2 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -23,6 +23,7 @@ enum {
 	FM_MAY_READ,	/* stack slots that may be read by this instruction */
 	FM_MUST_WRITE,	/* stack slots written by this instruction */
 	FM_LIVE_BEFORE,	/* stack slots that may be read by this insn and its successors */
+	FM_MAY_WRITE,	/* stack slots that may be written by this instruction */
 	FM_MASK_CNT,
 };
 
@@ -262,6 +263,12 @@ static int mark_stack_write(struct func_instance *instance, u32 frame, u32 insn_
 	return mark_stack_range(instance, frame, insn_idx, FM_MUST_WRITE, lo, hi);
 }
 
+static int mark_stack_may_write(struct func_instance *instance, u32 frame, u32 insn_idx,
+				s32 lo, s32 hi)
+{
+	return mark_stack_range(instance, frame, insn_idx, FM_MAY_WRITE, lo, hi);
+}
+
 /*
  * Mark every half-slot of @frame as possibly read by @insn_idx. This widens
  * the masks to the program's stack budget: a full read recorded at a narrower
@@ -277,9 +284,16 @@ static int mark_stack_read_all(struct bpf_verifier_env *env, struct func_instanc
 			       env->stack_limit / BPF_HALF_REG_SIZE - 1);
 }
 
-/* Accumulate @src, a mask @src_words wide, into may_read of @frame at @insn_idx */
-static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 insn_idx,
-				const unsigned long *src, u32 src_words)
+static int mark_stack_may_write_all(struct bpf_verifier_env *env, struct func_instance *instance,
+				    u32 frame, u32 insn_idx)
+{
+	return mark_stack_may_write(instance, frame, insn_idx, 0,
+				    env->stack_limit / BPF_HALF_REG_SIZE - 1);
+}
+
+/* Accumulate @src, a mask @src_words wide, into @kind mask of @frame at @insn_idx */
+static int mark_stack_mask(struct func_instance *instance, u32 frame, u32 insn_idx, u32 kind,
+			   const unsigned long *src, u32 src_words)
 {
 	u32 nbits = src_words * BITS_PER_LONG;
 	struct frame_masks *fm;
@@ -292,7 +306,7 @@ static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 i
 	fm = widen_frame_masks(instance, frame, BITS_TO_LONGS(last + 1));
 	if (!fm)
 		return -ENOMEM;
-	dst = rel_mask(fm, relative_idx(instance, insn_idx), FM_MAY_READ);
+	dst = rel_mask(fm, relative_idx(instance, insn_idx), kind);
 	/* @src has no bits set past @last, hence none past @fm->words either */
 	src_words = min(src_words, fm->words);
 	for (w = 0; w < src_words; w++)
@@ -300,6 +314,12 @@ static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 i
 	return 0;
 }
 
+static int mark_stack_read_mask(struct func_instance *instance, u32 frame, u32 insn_idx,
+				const unsigned long *src, u32 src_words)
+{
+	return mark_stack_mask(instance, frame, insn_idx, FM_MAY_READ, src, src_words);
+}
+
 int bpf_jmp_offset(struct bpf_insn *insn)
 {
 	u8 code = insn->code;
@@ -624,16 +644,41 @@ static char *fmt_spis_mask(struct bpf_verifier_env *env, int frame, bool first,
 	return env->tmp_str_buf;
 }
 
+/* Print mask @kind of the instruction at relative index @i for every frame, if any bit is set. */
+static bool print_mask(struct bpf_verifier_env *env, struct func_instance *instance, int i,
+		       const char *name, u32 kind)
+{
+	struct frame_masks *fm;
+	bool printed = false;
+	unsigned long *mask;
+	int frame;
+	u64 pos;
+
+	pos = env->log.end_pos;
+	verbose(env, "%s", name);
+	for (frame = instance->depth; frame >= 0; --frame) {
+		fm = instance->frames[frame];
+		if (!fm)
+			continue;
+		mask = rel_mask(fm, i, kind);
+		if (bitmap_empty(mask, frame_mask_bits(fm)))
+			continue;
+		verbose(env, "%s", fmt_spis_mask(env, frame, !printed, mask, fm->words));
+		printed = true;
+	}
+	if (!printed)
+		bpf_vlog_reset(&env->log, pos);
+	return printed;
+}
+
 static void print_instance(struct bpf_verifier_env *env, struct func_instance *instance)
 {
 	int start = env->subprog_info[instance->subprog].start;
 	struct bpf_insn *insns = env->prog->insnsi;
-	struct frame_masks *fm;
-	unsigned long *mask;
 	int len = instance->insn_cnt;
-	int insn_idx, frame, i;
-	bool has_use, has_def;
 	u64 pos, insn_pos;
+	int insn_idx, i;
+	bool printed;
 
 	if (!(env->log.level & BPF_LOG_LEVEL2))
 		return;
@@ -642,41 +687,17 @@ static void print_instance(struct bpf_verifier_env *env, struct func_instance *i
 	verbose(env, "%s:\n", fmt_instance(env, instance));
 	for (i = 0; i < len; i++) {
 		insn_idx = start + i;
-		has_use = false;
-		has_def = false;
 		pos = env->log.end_pos;
 		verbose(env, "%3d: ", insn_idx);
 		bpf_verbose_insn(env, &insns[insn_idx]);
 		insn_pos = env->log.end_pos;
 		verbose(env, "%*c;", bpf_vlog_alignment(insn_pos - pos), ' ');
-		pos = env->log.end_pos;
-		verbose(env, " use: ");
-		for (frame = instance->depth; frame >= 0; --frame) {
-			fm = instance->frames[frame];
-			if (!fm)
-				continue;
-			mask = rel_mask(fm, i, FM_MAY_READ);
-			if (bitmap_empty(mask, frame_mask_bits(fm)))
-				continue;
-			verbose(env, "%s", fmt_spis_mask(env, frame, !has_use, mask, fm->words));
-			has_use = true;
-		}
-		if (!has_use)
-			bpf_vlog_reset(&env->log, pos);
-		pos = env->log.end_pos;
-		verbose(env, " def: ");
-		for (frame = instance->depth; frame >= 0; --frame) {
-			fm = instance->frames[frame];
-			if (!fm)
-				continue;
-			mask = rel_mask(fm, i, FM_MUST_WRITE);
-			if (bitmap_empty(mask, frame_mask_bits(fm)))
-				continue;
-			verbose(env, "%s", fmt_spis_mask(env, frame, !has_def, mask, fm->words));
-			has_def = true;
-		}
-		if (!has_def)
-			bpf_vlog_reset(&env->log, has_use ? pos : insn_pos);
+		printed = false;
+		printed |= print_mask(env, instance, i, " use: ", FM_MAY_READ);
+		printed |= print_mask(env, instance, i, " def: ", FM_MUST_WRITE);
+		printed |= print_mask(env, instance, i, " may_def: ", FM_MAY_WRITE);
+		if (!printed)
+			bpf_vlog_reset(&env->log, insn_pos);
 		verbose(env, "\n");
 		if (bpf_is_ldimm64(&insns[insn_idx]))
 			i++;
@@ -1444,10 +1465,12 @@ static void arg_track_xfer(struct bpf_verifier_env *env, struct bpf_insn *insn,
  *   access_bytes == 0:      no access
  *
  */
-static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
-				   s64 access_bytes, u32 frame, u32 insn_idx)
+static int record_stack_access_off(struct func_instance *instance, const struct arg_track *arg,
+				   u32 off_idx, s64 access_bytes, u32 frame, u32 insn_idx)
 {
+	s64 fp_off = arg->off[off_idx];
 	s32 slot_hi, slot_lo;
+	int err;
 
 	if (fp_off >= 0)
 		/*
@@ -1459,7 +1482,8 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
 	if (access_bytes == S64_MIN) {
 		/* helper/kfunc read unknown amount of bytes from fp_off until fp+0 */
 		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
-		return mark_stack_read(instance, frame, insn_idx, 0, slot_hi);
+		err = mark_stack_read(instance, frame, insn_idx, 0, slot_hi);
+		return err ?: mark_stack_may_write(instance, frame, insn_idx, 0, slot_hi);
 	}
 	if (access_bytes > 0) {
 		/* Mark any touched slot as use */
@@ -1471,7 +1495,15 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
 		access_bytes = -access_bytes;
 		slot_hi = (-fp_off) / STACK_SLOT_SZ - 1;
 		slot_lo = max_t(s32, (-fp_off - access_bytes + STACK_SLOT_SZ - 1) / STACK_SLOT_SZ, 0);
-		return mark_stack_write(instance, frame, insn_idx, slot_lo, slot_hi);
+		if (arg->off_cnt == 1) {
+			err = mark_stack_write(instance, frame, insn_idx, slot_lo, slot_hi);
+			if (err)
+				return err;
+		}
+		/* Mark partially covered slots as may_def */
+		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
+		slot_lo = max_t(s32, (-fp_off - access_bytes) / STACK_SLOT_SZ, 0);
+		return mark_stack_may_write(instance, frame, insn_idx, slot_lo, slot_hi);
 	}
 	return 0;
 }
@@ -1489,16 +1521,21 @@ static int record_stack_access(struct bpf_verifier_env *env, struct func_instanc
 	if (access_bytes == 0)
 		return 0;
 	if (arg->off_cnt == 0) {
-		if (access_bytes > 0 || access_bytes == S64_MIN)
-			return mark_stack_read_all(env, instance, frame, insn_idx);
+		if (access_bytes > 0 || access_bytes == S64_MIN) {
+			err = mark_stack_read_all(env, instance, frame, insn_idx);
+			if (err)
+				return err;
+		}
+		if (access_bytes < 0 || access_bytes == S64_MIN) {
+			err = mark_stack_may_write_all(env, instance, frame, insn_idx);
+			if (err)
+				return err;
+		}
 		return 0;
 	}
-	if (access_bytes != S64_MIN && access_bytes < 0 && arg->off_cnt != 1)
-		/* multi-offset write cannot set stack_def */
-		return 0;
 
 	for (i = 0; i < arg->off_cnt; i++) {
-		err = record_stack_access_off(instance, arg->off[i], access_bytes, frame, insn_idx);
+		err = record_stack_access_off(instance, arg, i, access_bytes, frame, insn_idx);
 		if (err)
 			return err;
 	}
@@ -1507,10 +1544,11 @@ static int record_stack_access(struct bpf_verifier_env *env, struct func_instanc
 
 /*
  * When a pointer is ARG_IMPRECISE, conservatively mark every frame in
- * the bitmask as fully used.
+ * the bitmask as fully used. Same as in record_stack_access(),
+ * negative 'access_bytes' means stack write.
  */
 static int record_imprecise(struct bpf_verifier_env *env, struct func_instance *instance,
-			    u32 mask, u32 insn_idx)
+			    s64 access_bytes, u32 mask, u32 insn_idx)
 {
 	int depth = instance->depth;
 	int f, err;
@@ -1519,9 +1557,16 @@ static int record_imprecise(struct bpf_verifier_env *env, struct func_instance *
 		if (!(mask & 1))
 			continue;
 		if (f <= depth) {
-			err = mark_stack_read_all(env, instance, f, insn_idx);
-			if (err)
-				return err;
+			if (access_bytes > 0 || access_bytes == S64_MIN) {
+				err = mark_stack_read_all(env, instance, f, insn_idx);
+				if (err)
+					return err;
+			}
+			if (access_bytes < 0) {
+				err = mark_stack_may_write_all(env, instance, f, insn_idx);
+				if (err)
+					return err;
+			}
 		}
 	}
 	return 0;
@@ -1590,7 +1635,7 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 	if (ptr->frame >= 0 && ptr->frame <= depth)
 		return record_stack_access(env, instance, ptr, sz, ptr->frame, insn_idx);
 	if (ptr->frame == ARG_IMPRECISE)
-		return record_imprecise(env, instance, ptr->mask, insn_idx);
+		return record_imprecise(env, instance, sz, ptr->mask, insn_idx);
 	/* ARG_NONE: not derived from any frame pointer, skip */
 	return 0;
 }
@@ -1618,6 +1663,9 @@ static int record_arg_access(struct bpf_verifier_env *env,
 			err = mark_stack_read_all(env, instance, f, insn_idx);
 			if (err)
 				return err;
+			err = mark_stack_may_write_all(env, instance, f, insn_idx);
+			if (err)
+				return err;
 		}
 		return 0;
 	}
@@ -1627,7 +1675,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
 	if (frame >= 0 && frame <= depth)
 		err = record_stack_access(env, instance, at, bytes, frame, insn_idx);
 	else if (frame == ARG_IMPRECISE)
-		err = record_imprecise(env, instance, at->mask, insn_idx);
+		err = record_imprecise(env, instance, bytes, at->mask, insn_idx);
 	return err;
 }
 
@@ -1987,6 +2035,7 @@ static bool has_fp_args(struct arg_track *args)
 /*
  * Merge a freshly analyzed instance into the original.
  * may_read: union (any pass might read the slot).
+ * may_write: union (slots written on ANY pass).
  * must_write: intersection (only slots written on ALL passes are guaranteed).
  * live_before is recomputed by a subsequent update_instance() on @dst.
  *
@@ -2025,18 +2074,50 @@ static int merge_instances(struct func_instance *dst, struct func_instance *src)
 		for (i = 0; i < dst->insn_cnt; i++) {
 			unsigned long *dst_read = rel_mask(d, i, FM_MAY_READ);
 			unsigned long *dst_write = rel_mask(d, i, FM_MUST_WRITE);
+			unsigned long *dst_may_write = rel_mask(d, i, FM_MAY_WRITE);
 			unsigned long *src_read = rel_mask(s, i, FM_MAY_READ);
 			unsigned long *src_write = rel_mask(s, i, FM_MUST_WRITE);
+			unsigned long *src_may_write = rel_mask(s, i, FM_MAY_WRITE);
 
 			for (w = 0; w < d->words; w++) {
 				dst_read[w] |= w < s->words ? src_read[w] : 0;
 				dst_write[w] &= w < s->words ? src_write[w] : 0;
+				dst_may_write[w] |= w < s->words ? src_may_write[w] : 0;
 			}
 		}
 	}
 	return 0;
 }
 
+/*
+ * Fold a fully analyzed callee instance writes to upper frames as
+ * may_write marks at callsite in caller's frames.
+ */
+static int merge_may_write(struct func_instance *caller, struct func_instance *callee)
+{
+	DECLARE_BITMAP(acc, FRAME_HALF_SPIS);
+	u32 call_idx = callee->callsite;
+	struct frame_masks *fm;
+	u32 f, i, nbits;
+	int err;
+
+	for (f = 0; f < callee->depth; f++) {
+		fm = callee->frames[f];
+		if (!fm)
+			continue;
+		nbits = frame_mask_bits(fm);
+		bitmap_zero(acc, nbits);
+		for (i = 0; i < callee->insn_cnt; i++)
+			bitmap_or(acc, acc, rel_mask(fm, i, FM_MAY_WRITE), nbits);
+		if (bitmap_empty(acc, nbits))
+			continue;
+		err = mark_stack_mask(caller, f, call_idx, FM_MAY_WRITE, acc, fm->words);
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
 static struct func_instance *fresh_instance(struct func_instance *src)
 {
 	struct func_instance *f;
@@ -2204,6 +2285,11 @@ static int analyze_subprog(struct bpf_verifier_env *env,
 					goto out_free;
 			}
 		}
+
+		/* Summarize callee's writes to ancestor frames onto the callsite */
+		err = merge_may_write(instance, callee_instance);
+		if (err)
+			goto out_free;
 	}
 
 	if (prev_instance) {
diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 16b2b1e57534..7e2e165a7145 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -61,7 +61,11 @@ __naked void read_write_join(void)
 SEC("socket")
 __log_level(2)
 __msg("stack use/def subprog#0 must_write_not_same_slot (d0,cs0):")
-__msg("6: (7b) *(u64 *)(r2 +0) = r0{{$}}")
+/*
+ * 'r2 += r1' adds a scalar, so the offset is lost (off_cnt == 0): no def,
+ * but the write conservatively marks the whole frame as may_def.
+ */
+__msg("6: (7b) *(u64 *)(r2 +0) = r0         ; may_def: fp0-8..-{{(512|2048)}}")
 __msg("Live regs before insn:")
 __naked void must_write_not_same_slot(void)
 {
@@ -106,9 +110,9 @@ __naked void must_write_not_same_type(void)
 
 SEC("socket")
 __log_level(2)
-/* Callee writes fp[0]-8: stack_use at call site has slots 0,1 live */
+/* Callee writes fp[0]-8: the def is summarized as may_def at the call site */
 __msg("stack use/def subprog#0 caller_stack_write (d0,cs0):")
-__msg("2: (85) call pc+1{{$}}")
+__msg("2: (85) call pc+1                    ; may_def: fp0-8")
 __msg("stack use/def subprog#1 write_first_param (d1,cs2):")
 __msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8")
 __naked void caller_stack_write(void)
@@ -804,7 +808,7 @@ void __kfunc_btf_root(void)
  */
 SEC("socket")
 __success __log_level(2)
-__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24{{$}}")
+__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24 may_def: fp0-24{{$}}")
 __msg(" 9: (85) call bpf_iter_num_next{{.*}}         ; use: fp0-24{{$}}")
 __msg("14: (85) call bpf_iter_num_destroy{{.*}}      ; use: fp0-24{{$}}")
 __naked void kfunc_iter_stack_liveness(void)
@@ -1008,7 +1012,8 @@ __naked void four_byte_read_upper_half(void)
 SEC("socket")
 __log_level(2)
 __msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
-__msg("1: (6a) *(u16 *)(r10 -4) = 0{{$}}")
+/* 2-byte write only partially covers the upper half: may_def, but no def. */
+__msg("1: (6a) *(u16 *)(r10 -4) = 0         ; may_def: fp0-4h")
 __msg("2: (61) r0 = *(u32 *)(r10 -4)        ; use: fp0-4h")
 __naked void two_byte_write_no_kill(void)
 {
@@ -1355,9 +1360,12 @@ __naked void fp_spill_loses_precision_kills_liveness(void)
  */
 SEC("socket")
 __log_level(2)
-/* fp-8 live at call (callee conditionally writes → slot not killed) */
+/*
+ * fp-8 live at call: callee conditionally writes it, so the slot is not killed
+ * (no def), but the conditional write surfaces as may_def at the call site.
+ */
 __msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8")
-__msg("4: (85) call pc+2{{$}}")
+__msg("4: (85) call pc+2                    ; may_def: fp0-8")
 __msg("5: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void conditional_stx_in_subprog(void)
 {
@@ -2386,7 +2394,12 @@ __msg("subprog#2 write_first_read_second:")
 __msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
 __msg("18: (79) r0 = *(u64 *)(r2 +0) // r1=fp0-8 r2=fp0-16{{$}}")
 __msg("stack use/def subprog#2 write_first_read_second (d2,cs15):")
-__msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
+/*
+ * Shared across two callsites with swapped args (r1 is fp-8 on one pass,
+ * fp-16 on the other): must_write intersects to empty (no def), may_write
+ * unions to both slots.
+ */
+__msg("17: (7a) *(u64 *)(r1 +0) = 42         ; may_def: fp0-8 fp0-16")
 __msg("18: (79) r0 = *(u64 *)(r2 +0)         ; use: fp0-8 fp0-16")
 __naked void shared_instance_must_write_overwrite(void)
 {
@@ -2848,8 +2861,9 @@ static __used __naked void imprecise_dst_spill_join_sub(void)
 SEC("socket")
 __log_level(2)
 __msg("0: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
-__msg("1: (73) *(u8 *)(r10 -1) = r0{{$}}")
-__msg("2: (6b) *(u16 *)(r10 -4) = r0{{$}}")
+/* narrow stores define nothing, but they may write the half-slot they touch */
+__msg("1: (73) *(u8 *)(r10 -1) = r0         ; may_def: fp0-4h")
+__msg("2: (6b) *(u16 *)(r10 -4) = r0        ; may_def: fp0-4h")
 __msg("3: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void narrow_store_defines_nothing(void)
 {

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
  2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 15:51   ` Alexei Starovoitov
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 03/36] bpf: summarize live " Eduard Zingerman
                   ` (33 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV needs may-write information to invalidate stack expressions after
indirect writes. Liveness tracks this per function instance,
but SCEV is not callchain-sensitive and needs a summary across
calling contexts.

For each instruction, union the current-frame may_write masks from all
analyzed instances. Represent 8-byte slots as a summary of constituent
halves. Store the summary in insn_aux_data and expose it via a helper.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  6 ++++++
 kernel/bpf/liveness.c        | 41 +++++++++++++++++++++++++++++++++++++++++
 2 files changed, 47 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c775bd757706..c6d617581e84 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -662,6 +662,11 @@ struct bpf_insn_aux_data {
 	};
 	struct btf_struct_meta *kptr_struct_meta;
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
+	/*
+	 * Per-instruction summary of stack slots in the current frame
+	 * that this instruction may write to.
+	 */
+	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
@@ -1728,6 +1733,7 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
 int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
+const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx);
 bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 spi);
 int bpf_compute_live_registers(struct bpf_verifier_env *env);
 
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cc3ad75aa1e2..c871744ca5a8 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -718,6 +718,45 @@ static int cmp_instances(const void *pa, const void *pb)
 	return 0;
 }
 
+/* OR the 8-byte slots touched by a half-slot (4-byte) mask into @slots. */
+static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask, u32 nbits)
+{
+	u32 slot;
+
+	for (slot = 0; slot < MAX_BPF_STACK_SLOTS && slot * 2 + 1 < nbits; slot++)
+		if (test_bit(slot * 2, mask) || test_bit(slot * 2 + 1, mask))
+			__set_bit(slot, slots);
+}
+
+/*
+ * Precompute, for each instruction, the OR of may_write masks over its top
+ * frame across all func_instances reaching it, stash it in the insn_aux_data.
+ */
+static void compute_may_write_masks(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_liveness *liveness = env->liveness;
+	struct func_instance *instance;
+	struct frame_masks *fm;
+	u32 nbits;
+	int bkt, i;
+
+	hash_for_each(liveness->func_instances, bkt, instance, hl_node) {
+		fm = instance->frames[instance->depth];
+		if (!fm)
+			continue;
+		nbits = frame_mask_bits(fm);
+		for (i = 0; i < instance->insn_cnt; i++)
+			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
+					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
+	}
+}
+
+const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx)
+{
+	return env->insn_aux_data[insn_idx].may_write_mask;
+}
+
 /* print use/def slots for all instances ordered by callsite first, then by depth */
 static int print_instances(struct bpf_verifier_env *env)
 {
@@ -2352,6 +2391,8 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env)
 			goto out;
 	}
 
+	compute_may_write_masks(env);
+
 	if (env->log.level & BPF_LOG_LEVEL2)
 		err = print_instances(env);
 

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 03/36] bpf: summarize live stack slots in insn_aux_data
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
  2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:33   ` sashiko-bot
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
                   ` (32 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV needs live stack information in order to skip tracking dead
slots. Liveness tracks this per function instance, but SCEV is not
callchain-sensitive and needs a summary across calling contexts.

For each instruction, union the current-frame liveness masks from all
analyzed function instances. Represent 8-byte slots as a summary of
constituent halves. Store the summary in insn_aux_data and expose it
via a helper.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/liveness.c        | 10 +++++++---
 2 files changed, 8 insertions(+), 3 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c6d617581e84..7f31ce5ea6b7 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -667,6 +667,7 @@ struct bpf_insn_aux_data {
 	 * that this instruction may write to.
 	 */
 	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
+	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index c871744ca5a8..b6b7fd479569 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -729,8 +729,9 @@ static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask,
 }
 
 /*
- * Precompute, for each instruction, the OR of may_write masks over its top
- * frame across all func_instances reaching it, stash it in the insn_aux_data.
+ * Precompute, for each instruction, the OR of may_write and live_before masks
+ * over its top frame across all func_instances reaching it, stash them in the
+ * insn_aux_data.
  */
 static void compute_may_write_masks(struct bpf_verifier_env *env)
 {
@@ -746,9 +747,12 @@ static void compute_may_write_masks(struct bpf_verifier_env *env)
 		if (!fm)
 			continue;
 		nbits = frame_mask_bits(fm);
-		for (i = 0; i < instance->insn_cnt; i++)
+		for (i = 0; i < instance->insn_cnt; i++) {
 			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
 					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
+			half_spis_to_slots(aux[instance->subprog_start + i].live_stack_before,
+					   rel_mask(fm, i, FM_LIVE_BEFORE), nbits);
+		}
 	}
 }
 

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer in insn_aux_data
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (2 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 03/36] bpf: summarize live " Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:27   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
                   ` (31 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV needs to know which registers might be stack pointers at a
particular instruction. This information is used to invalidate SCEV
expressions for slots that might be overwritten by indirect writes.

liveness.c:compute_may_write_masks() already tracks this information.
This commit modifies it to save the information in insn_aux_data for
further usage.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/liveness.c        | 13 +++++++++++++
 2 files changed, 14 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 7f31ce5ea6b7..c9fb0a9985d4 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -668,6 +668,7 @@ struct bpf_insn_aux_data {
 	 */
 	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
 	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);
+	u16 stack_ptrs; /* bitmask of regs that may hold a frame pointer here (arg_track) */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
 	bool nospec; /* do not execute this instruction speculatively */
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index b6b7fd479569..1cce0c7406eb 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1899,6 +1899,17 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
 	}
 }
 
+static void record_stack_ptrs(struct bpf_verifier_env *env, int idx, struct arg_track *at_in)
+{
+	u16 r, mask = 0;
+
+	for (r = 0; r < MAX_BPF_REG; r++)
+		if (arg_is_fp(&at_in[r]))
+			mask |= BIT(r);
+
+	env->insn_aux_data[idx].stack_ptrs |= mask;
+}
+
 /*
  * Compute arg tracking dataflow for a single subprog.
  * Runs forward fixed-point with arg_track_xfer(), then records
@@ -2048,6 +2059,8 @@ static int compute_subprog_args(struct bpf_verifier_env *env,
 			snap->slots = nslots;
 			memcpy(snap->at, &at_stack_in[(size_t)i * nslots], nslots * sizeof(*snap->at));
 		}
+
+		record_stack_ptrs(env, idx, at_in[i]);
 	}
 
 	info->at_in = at_in;

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (3 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:27   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 06/36] bpf: add tnum_alignment() Eduard Zingerman
                   ` (30 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Atomic read-modify-write instructions are recorded as reads only,
omitting their stack write effects from the may_write masks needed
by SCEV.

Record both read and write accesses for atomic RMW operations. Keep
LOAD_ACQ read-only and STORE_REL write-only. Share the precise/imprecise
stack-access dispatch so both accesses use the same address handling.

Fixes: fed53dbcdb61 ("bpf: record arg tracking results in bpf_liveness masks")

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c | 57 ++++++++++++++++++++++++++++++++-------------------
 1 file changed, 36 insertions(+), 21 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 1cce0c7406eb..6b2f1bdbef7f 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -1555,9 +1555,10 @@ static int record_stack_access_off(struct func_instance *instance, const struct
  * 'arg' is FP-derived argument to helper/kfunc or load/store that
  * reads (positive) or writes (negative) 'access_bytes' into 'use' or 'def'.
  */
-static int record_stack_access(struct bpf_verifier_env *env, struct func_instance *instance,
-			       const struct arg_track *arg,
-			       s64 access_bytes, u32 frame, u32 insn_idx)
+static int record_precise(struct bpf_verifier_env *env,
+			  struct func_instance *instance,
+			  const struct arg_track *arg,
+			  s64 access_bytes, u32 frame, u32 insn_idx)
 {
 	int i, err;
 
@@ -1615,17 +1616,30 @@ static int record_imprecise(struct bpf_verifier_env *env, struct func_instance *
 	return 0;
 }
 
+static int record_stack_access(struct bpf_verifier_env *env,
+			       struct func_instance *instance,
+			       const struct arg_track *ptr,
+			       s64 access_bytes, u32 insn_idx)
+{
+	if (ptr->frame >= 0 && ptr->frame <= instance->depth)
+		return record_precise(env, instance, ptr, access_bytes, ptr->frame, insn_idx);
+	if (ptr->frame == ARG_IMPRECISE)
+		return record_imprecise(env, instance, access_bytes, ptr->mask, insn_idx);
+	/* ARG_NONE: not derived from any frame pointer, skip */
+	return 0;
+}
+
 /* Record load/store access for a given 'at' state of 'insn'. */
 static int record_load_store_access(struct bpf_verifier_env *env,
 				    struct func_instance *instance,
 				    struct arg_track *at, int insn_idx)
 {
 	struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
-	int depth = instance->depth;
 	s32 sz = bpf_size_to_bytes(BPF_SIZE(insn->code));
 	u8 class = BPF_CLASS(insn->code);
 	struct arg_track resolved, *ptr;
-	int oi;
+	bool rmw = false; /* atomic read-modify-write: both read and write */
+	int oi, err;
 
 	/*
 	 * Stack arg insns use dst_reg/src_reg=BPF_REG_PARAMS(11). Since at[]
@@ -1643,12 +1657,20 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 		break;
 	case BPF_STX:
 		if (BPF_MODE(insn->code) == BPF_ATOMIC) {
-			if (insn->imm == BPF_STORE_REL)
-				sz = -sz;
-			if (insn->imm == BPF_LOAD_ACQ)
+			switch (insn->imm) {
+			case BPF_LOAD_ACQ:
 				ptr = &at[insn->src_reg];
-			else
+				break;
+			case BPF_STORE_REL:
+				ptr = &at[insn->dst_reg];
+				sz = -sz;
+				break;
+			default:
+				/* ADD/AND/OR/XOR(+FETCH), XCHG, CMPXCHG */
 				ptr = &at[insn->dst_reg];
+				rmw = true;
+				break;
+			}
 		} else {
 			ptr = &at[insn->dst_reg];
 			sz = -sz;
@@ -1675,12 +1697,10 @@ static int record_load_store_access(struct bpf_verifier_env *env,
 		ptr = &resolved;
 	}
 
-	if (ptr->frame >= 0 && ptr->frame <= depth)
-		return record_stack_access(env, instance, ptr, sz, ptr->frame, insn_idx);
-	if (ptr->frame == ARG_IMPRECISE)
-		return record_imprecise(env, instance, sz, ptr->mask, insn_idx);
-	/* ARG_NONE: not derived from any frame pointer, skip */
-	return 0;
+	err = record_stack_access(env, instance, ptr, sz, insn_idx);
+	if (!err && rmw) /* also record the write half */
+		err = record_stack_access(env, instance, ptr, -sz, insn_idx);
+	return err;
 }
 
 static int record_arg_access(struct bpf_verifier_env *env,
@@ -1690,7 +1710,6 @@ static int record_arg_access(struct bpf_verifier_env *env,
 			     int insn_idx)
 {
 	int depth = instance->depth;
-	int frame = at->frame;
 	int err = 0;
 	s64 bytes;
 
@@ -1715,11 +1734,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
 	if (bytes == 0)
 		return 0;
 
-	if (frame >= 0 && frame <= depth)
-		err = record_stack_access(env, instance, at, bytes, frame, insn_idx);
-	else if (frame == ARG_IMPRECISE)
-		err = record_imprecise(env, instance, bytes, at->mask, insn_idx);
-	return err;
+	return record_stack_access(env, instance, at, bytes, insn_idx);
 }
 
 /* Record stack access for a given 'at' state of helper/kfunc 'insn' */

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 06/36] bpf: add tnum_alignment()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (4 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 07/36] bpf: add cnum{32,64}_union() Eduard Zingerman
                   ` (29 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Used by the SCEV widening logic later in the series.
The verifier checks alignment for some memory accesses.
When SCEV widens a pointer or an index variable, it must retain
alignment information in order to allow array access.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/tnum.h |  2 ++
 kernel/bpf/tnum.c    | 11 +++++++++++
 2 files changed, 13 insertions(+)

diff --git a/include/linux/tnum.h b/include/linux/tnum.h
index ca2cfec8de08..16682248f963 100644
--- a/include/linux/tnum.h
+++ b/include/linux/tnum.h
@@ -134,4 +134,6 @@ static inline bool tnum_subreg_is_const(struct tnum a)
 /* Returns the smallest member of t larger than z */
 u64 tnum_step(struct tnum t, u64 z);
 
+u32 tnum_alignment(struct tnum a);
+
 #endif /* _LINUX_TNUM_H */
diff --git a/kernel/bpf/tnum.c b/kernel/bpf/tnum.c
index ec9c310cf5d7..1f620facd1b7 100644
--- a/kernel/bpf/tnum.c
+++ b/kernel/bpf/tnum.c
@@ -317,3 +317,14 @@ u64 tnum_step(struct tnum t, u64 z)
 	inc = (filled + 1) & t.mask;
 	return t.value | inc;
 }
+
+/*
+ * Return the number of trailing bits known to be zero in a tnum.
+ * Return 64 for a tnum representing zero.
+ */
+u32 tnum_alignment(struct tnum a)
+{
+	u64 v = a.value | a.mask;
+
+	return v ? __ffs64(v) : 64;
+}

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 07/36] bpf: add cnum{32,64}_union()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (5 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 06/36] bpf: add tnum_alignment() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear() Eduard Zingerman
                   ` (28 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Add cnum32_union() and cnum64_union() to compute a smallest circular
range containing both inputs.

Do so by enumerating the following configurations in a rotated
frame (a.base at the origin):

0                                              UT_MAX
|---------------------------------------------------|
[= a ==============================]                |
[= b tail =]           [= b main ===================>
[= union ===========================================]

0                                              UT_MAX
|---------------------------------------------------|
[= a =====================]                         |
[= b tail ======]                    [= b main =====>
[= union tail ============]          [= union main =>

0                                              UT_MAX
|---------------------------------------------------|
[= a =======]                                       |
[= b tail =============]             [= b main =====>
[= union tail =========]             [= union main =>

0                                              UT_MAX
|---------------------------------------------------|
[= a =========================]                     |
|                 [= b ====================]        |
[= union ==================================]        |

0                                              UT_MAX
|---------------------------------------------------|
[= a =====================================]         |
|               [= b ========]                      |
[= union =================================]         |

0                                              UT_MAX
|---------------------------------------------------|
[= a ============]                                  |
|                               [= b =======]       |

Two possible covering arcs:

[= ab ======================================]       |
[= ba tail ======]              [= ba main =========>

Pick the smaller one.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/cnum.h   |  2 ++
 kernel/bpf/cnum_defs.h | 44 ++++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 46 insertions(+)

diff --git a/include/linux/cnum.h b/include/linux/cnum.h
index 49b7d0c7645d..ddf4726841e6 100644
--- a/include/linux/cnum.h
+++ b/include/linux/cnum.h
@@ -39,6 +39,7 @@ u32 cnum32_umin(struct cnum32 cnum);
 u32 cnum32_umax(struct cnum32 cnum);
 s32 cnum32_smin(struct cnum32 cnum);
 s32 cnum32_smax(struct cnum32 cnum);
+struct cnum32 cnum32_union(struct cnum32 a, struct cnum32 b);
 struct cnum32 cnum32_intersect(struct cnum32 a, struct cnum32 b);
 void cnum32_intersect_with(struct cnum32 *dst, struct cnum32 src);
 void cnum32_intersect_with_urange(struct cnum32 *dst, u32 min, u32 max);
@@ -65,6 +66,7 @@ u64 cnum64_umin(struct cnum64 cnum);
 u64 cnum64_umax(struct cnum64 cnum);
 s64 cnum64_smin(struct cnum64 cnum);
 s64 cnum64_smax(struct cnum64 cnum);
+struct cnum64 cnum64_union(struct cnum64 a, struct cnum64 b);
 struct cnum64 cnum64_intersect(struct cnum64 a, struct cnum64 b);
 void cnum64_intersect_with(struct cnum64 *dst, struct cnum64 src);
 void cnum64_intersect_with_urange(struct cnum64 *dst, u64 min, u64 max);
diff --git a/kernel/bpf/cnum_defs.h b/kernel/bpf/cnum_defs.h
index 30685e43de04..05b17a8b4814 100644
--- a/kernel/bpf/cnum_defs.h
+++ b/kernel/bpf/cnum_defs.h
@@ -198,6 +198,50 @@ static inline struct cnum_t FN(normalize)(struct cnum_t cnum)
 	return cnum;
 }
 
+/*
+ * Return a smallest arc containing both 'a' and 'b'.
+ * Break equal-size ties by choosing the smaller base.
+ */
+struct cnum_t FN(union)(struct cnum_t a, struct cnum_t b)
+{
+	struct cnum_t b1, ab, ba;
+	ut end;
+
+	if (FN(is_empty)(a))
+		return b;
+	if (FN(is_empty)(b))
+		return a;
+
+	/*
+	 * Rotate so that a1.base == 0 and a1.end == a.size.
+	 * Normalize b1 to preserve the full-circle representation.
+	 */
+	b1 = FN(normalize)((struct cnum_t){ b.base - a.base, b.size });
+	end = max(a.size, (ut)(b1.base + b1.size));
+
+	if (FN(urange_overflow)(b1)) {
+		/* a1 reaches b1's main arc: together they cover the circle. */
+		if (b1.base <= a.size)
+			return (struct cnum_t){ 0, UT_MAX };
+
+		/* Extend b1's tail through a1's end, then rotate back. */
+		return FN(normalize)((struct cnum_t){ b.base, end - b1.base });
+	}
+
+	/* ab, rotated back, covers both nonwrapping arcs. */
+	ab = (struct cnum_t){ a.base, end };
+	if (b1.base <= a.size)
+		return FN(normalize)(ab);
+
+	/* The arcs are disjoint; ba is the other possible covering arc. */
+	ba = (struct cnum_t){ b.base, a.size - b1.base };
+	if (ba.size < ab.size ||
+	    (ba.size == ab.size && ba.base < ab.base))
+		ab = ba;
+
+	return FN(normalize)(ab);
+}
+
 struct cnum_t FN(add)(struct cnum_t a, struct cnum_t b)
 {
 	if (FN(is_empty)(a) || FN(is_empty)(b))

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (6 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 07/36] bpf: add cnum{32,64}_union() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:34   ` sashiko-bot
  2026-09-26 14:20 ` [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range() Eduard Zingerman
                   ` (27 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Add cnum64_intersect_linear(), which intersects a cnum64 interval with the
set of integers congruent to 'base' modulo 'step', i.e. the points
base + step * k. It tightens the interval by rounding its signed minimum up
and its signed maximum down to the nearest such point, and returns an empty
cnum when no point falls within the interval.

The arithmetic is done on the signed bounds [smin, smax]. The line is defined
over the integers, while a cnum64 is an arc on the u64 circle, so a u64-modular
computation would misplace the line for negative values whenever 'step' does
not divide 2^64. Working from the signed bounds keeps the residues consistent,
e.g. it can reason about a range like [-9, 3] with step 3.

Also add imod() to bpf_verifier.h: a mathematical modulo returning an
always-non-negative residue (unlike C's '%', which follows the sign of the
dividend). It is used by cnum64_intersect_linear() and by the base/step
tracking added in subsequent patches.

This is a building block for tracking scalar registers whose value is known
to lie on a line base + step * k.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 14 ++++++++++++++
 include/linux/cnum.h         |  1 +
 kernel/bpf/cnum.c            | 36 ++++++++++++++++++++++++++++++++++++
 3 files changed, 51 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c9fb0a9985d4..b6cfac01ef3f 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -9,6 +9,20 @@
 #include <linux/filter.h> /* for MAX_BPF_STACK */
 #include <linux/tnum.h>
 #include <linux/cnum.h>
+#include <linux/math64.h> /* for div_s64_rem() */
+
+/*
+ * Mathematical modulo, the residue of 'v' modulo 'step', normalized to [0, step).
+ * This differs from C's '%' operator, which truncates the quotient toward zero
+ * and so returns a remainder with the sign of the dividend (e.g. -1 % 3 == -1, not 2).
+ */
+static inline u16 imod(s64 v, u16 step)
+{
+	s32 rem;
+
+	div_s64_rem(v, step, &rem);
+	return rem < 0 ? rem + step : rem;
+}
 
 /* Maximum variable offset umax_value permitted when resolving memory accesses.
  * In practice this is far bigger than any realistic pointer offset; this limit
diff --git a/include/linux/cnum.h b/include/linux/cnum.h
index ddf4726841e6..49160318adf9 100644
--- a/include/linux/cnum.h
+++ b/include/linux/cnum.h
@@ -80,5 +80,6 @@ bool cnum64_is_subset(struct cnum64 outer, struct cnum64 inner);
 
 struct cnum32 cnum32_from_cnum64(struct cnum64 cnum);
 struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b);
+struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step);
 
 #endif /* _LINUX_CNUM_H */
diff --git a/kernel/bpf/cnum.c b/kernel/bpf/cnum.c
index 86142cb2aee5..2bff2c7e0cdf 100644
--- a/kernel/bpf/cnum.c
+++ b/kernel/bpf/cnum.c
@@ -2,6 +2,7 @@
 /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
 
 #include <linux/bits.h>
+#include <linux/bpf_verifier.h> /* for imod() */
 
 #define T 32
 #include "cnum_defs.h"
@@ -118,3 +119,38 @@ struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b)
 	}
 	return t;
 }
+
+/* Intersect 'in' with the set of integers defined by equation 'base + step * k'. */
+struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step)
+{
+	s64 smin = cnum64_smin(in);
+	s64 smax = cnum64_smax(in);
+	s64 lo, hi;
+	u16 d;
+
+	if (step <= 1 || cnum64_is_empty(in))
+		return in;
+	/*
+	 * Round smin up to the next value congruent to 'base' modulo 'step',
+	 * i.e. increase smin by d = (base - smin) mod step:
+	 *
+	 *                 |<---- d ---->|
+	 *     |-----------|=============|...
+	 * base+step*k    smin       base+step*(k+1)
+	 */
+	d = imod(base - imod(smin, step), step);
+	if ((u64)smax - (u64)smin < d)
+		return CNUM64_EMPTY;
+	lo = smin + d;
+	/*
+	 * Round smax down to the previous value congruent to 'base' modulo 'step',
+	 * i.e. decrease smax by d = (smax - base) mod step:
+	 *
+	 *     |<--- d --->|
+	 *  ...|===========|-------------|
+	 * base+step*k    smax       base+step*(k+1)
+	 */
+	d = imod(imod(smax, step) - base, step);
+	hi = smax - d;
+	return cnum64_from_srange(lo, hi);
+}

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (7 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:36   ` sashiko-bot
  2026-09-26 14:20 ` [PATCH bpf-next 10/36] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
                   ` (26 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

A utility function for setting a register's range and step information.
Used by SCEV widening and clamping logic further in the series.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/verifier.c        | 12 ++++++++++++
 2 files changed, 14 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index b6cfac01ef3f..94c0dddaaaa7 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1745,6 +1745,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 				 struct bpf_insn *insn, int arg,
 				 int insn_idx);
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
+int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
+		      struct cnum64 range, u16 step);
 
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 03dbc0e00398..f2d44e026e4b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -16934,6 +16934,18 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 	return 0;
 }
 
+int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
+		      struct cnum64 range, u16 step)
+{
+	reg->r64 = range;
+	reg->r32 = CNUM32_UNBOUNDED;
+	reg->step = step;
+	reg->base = imod(cnum64_smin(range), step);
+	reg->var_off = tnum_unknown;
+	reg_bounds_sync(reg); /* this should infer the tnum alignment */
+	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
+}
+
 /* check validity of 32-bit and 64-bit arithmetic operations */
 static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 10/36] bpf: add bpf_mark_reg_known_scalar()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (8 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 11/36] bpf: add bpf_reg_union() Eduard Zingerman
                   ` (25 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Provide a helper to set a register to a known scalar value for use outside
verifier.c. To be used by SCEV widening logic.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 1 +
 kernel/bpf/verifier.c        | 6 ++++++
 2 files changed, 7 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 94c0dddaaaa7..719b7c7fd2f5 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1293,6 +1293,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
 void bpf_bt_sync_linked_regs(struct backtrack_state *bt, struct bpf_jmp_history_entry *hist);
 void bpf_mark_reg_not_init(const struct bpf_verifier_env *env,
 			   struct bpf_reg_state *reg);
+void bpf_mark_reg_known_scalar(struct bpf_reg_state *reg, u64 imm);
 void bpf_mark_reg_unknown_imprecise(struct bpf_reg_state *reg);
 void bpf_mark_all_scalars_precise(struct bpf_verifier_env *env,
 				  struct bpf_verifier_state *st);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f2d44e026e4b..b5f45cb46f14 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -1932,6 +1932,12 @@ static void __mark_reg_known(struct bpf_reg_state *reg, u64 imm)
 	___mark_reg_known(reg, imm);
 }
 
+void bpf_mark_reg_known_scalar(struct bpf_reg_state *reg, u64 imm)
+{
+	__mark_reg_known(reg, imm);
+	reg->type = SCALAR_VALUE;
+}
+
 static void __mark_reg32_known(struct bpf_reg_state *reg, u64 imm)
 {
 	reg->var_off = tnum_const_subreg(reg->var_off, imm);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 11/36] bpf: add bpf_reg_union()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (9 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 10/36] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations Eduard Zingerman
                   ` (24 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

A utility function to take the union of two registers' scalar values:
merge the circular 32-bit and 64-bit bounds and tnums of two registers,
then synchronize and validate the result. Leave scalar id invalidation
and type validation to the caller.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/verifier.c        | 28 ++++++++++++++++++++++++++++
 2 files changed, 30 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 719b7c7fd2f5..706fdefbc07a 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1748,6 +1748,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
 int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
 int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 		      struct cnum64 range, u16 step);
+int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
+		  const struct bpf_reg_state *src);
 
 int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index b5f45cb46f14..1ec44d72a346 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -30,6 +30,7 @@
 #include <linux/module.h>
 #include <linux/cpumask.h>
 #include <linux/cnum.h>
+#include <linux/gcd.h>
 #include <linux/bpf_mem_alloc.h>
 #include <net/xdp.h>
 #include <linux/trace_events.h>
@@ -16952,6 +16953,33 @@ int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
 }
 
+/* acc := acc U src, matching types only. Caller must clear acc's scalar ID. */
+int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
+		  const struct bpf_reg_state *src)
+{
+	u16 base, step;
+
+	if (acc->type != src->type) {
+		verifier_bug(env, "union of registers with different types");
+		return -EFAULT;
+	}
+	acc->r64 = cnum64_union(acc->r64, src->r64);
+	acc->r32 = cnum32_union(acc->r32, src->r32);
+	acc->var_off = tnum_union(acc->var_off, src->var_off);
+
+	/* Retain a common congruence if the bases agree modulo the gcd. */
+	step = gcd(acc->step, src->step);
+	base = acc->base % step;
+	if (base != src->base % step) {
+		reg_step_reset(acc);
+	} else {
+		acc->base = base;
+		acc->step = step;
+	}
+	reg_bounds_sync(acc);
+	return reg_bounds_sanity_check(env, acc, "bpf_reg_union");
+}
+
 /* check validity of 32-bit and 64-bit arithmetic operations */
 static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (10 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 11/36] bpf: add bpf_reg_union() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
                   ` (23 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Expose the existing helpers as bpf_rev_opcode() and bpf_flip_opcode()
and declare them in bpf_verifier.h. The helpers are used by SCEV
logic while analyzing loop conditions.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  3 +++
 kernel/bpf/verifier.c        | 14 ++++++--------
 2 files changed, 9 insertions(+), 8 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 706fdefbc07a..075a28de2a04 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1858,4 +1858,7 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env);
 int bpf_do_misc_fixups(struct bpf_verifier_env *env);
 int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 
+int bpf_flip_opcode(u32 opcode);
+u8 bpf_rev_opcode(u8 opcode);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 1ec44d72a346..79aa92486154 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -17249,8 +17249,6 @@ static void find_good_pkt_pointers(struct bpf_verifier_state *vstate,
 
 static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state *reg2,
 				u8 opcode, bool is_jmp32);
-static u8 rev_opcode(u8 opcode);
-
 /*
  * Learn more information about live branches by simulating refinement on both branches.
  * regs_refine_cond_op() is sound, so producing ill-formed register bounds for the branch means
@@ -17259,7 +17257,7 @@ static u8 rev_opcode(u8 opcode);
 static int simulate_both_branches_taken(struct bpf_verifier_env *env, u8 opcode, bool is_jmp32)
 {
 	/* Fallthrough (FALSE) branch */
-	regs_refine_cond_op(&env->false_reg1, &env->false_reg2, rev_opcode(opcode), is_jmp32);
+	regs_refine_cond_op(&env->false_reg1, &env->false_reg2, bpf_rev_opcode(opcode), is_jmp32);
 	reg_bounds_sync(&env->false_reg1);
 	reg_bounds_sync(&env->false_reg2);
 	/*
@@ -17445,7 +17443,7 @@ static int is_scalar_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_s
 	return simulate_both_branches_taken(env, opcode, is_jmp32);
 }
 
-static int flip_opcode(u32 opcode)
+int bpf_flip_opcode(u32 opcode)
 {
 	/* How can we transform "a <op> b" into "b <op> a"? */
 	static const u8 opcode_flip[16] = {
@@ -17476,7 +17474,7 @@ static int is_pkt_ptr_branch_taken(struct bpf_reg_state *dst_reg,
 		pkt = dst_reg;
 	} else if (dst_reg->type == PTR_TO_PACKET_END) {
 		pkt = src_reg;
-		opcode = flip_opcode(opcode);
+		opcode = bpf_flip_opcode(opcode);
 	} else {
 		return -1;
 	}
@@ -17531,7 +17529,7 @@ static int is_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_state *r
 
 		/* arrange that reg2 is a scalar, and reg1 is a pointer */
 		if (!is_reg_const(reg2, is_jmp32)) {
-			opcode = flip_opcode(opcode);
+			opcode = bpf_flip_opcode(opcode);
 			swap(reg1, reg2);
 		}
 		/* and ensure that reg2 is a constant */
@@ -17565,7 +17563,7 @@ static int is_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_state *r
 /* Opcode that corresponds to a *false* branch condition.
  * E.g., if r1 < r2, then reverse (false) condition is r1 >= r2
  */
-static u8 rev_opcode(u8 opcode)
+u8 bpf_rev_opcode(u8 opcode)
 {
 	switch (opcode) {
 	case BPF_JEQ:		return BPF_JNE;
@@ -17600,7 +17598,7 @@ static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state
 	case BPF_JGT:
 	case BPF_JSGE:
 	case BPF_JSGT:
-		opcode = flip_opcode(opcode);
+		opcode = bpf_flip_opcode(opcode);
 		swap(reg1, reg2);
 		break;
 	default:

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (11 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 14/36] bpf: representation for intervals with steps Eduard Zingerman
                   ` (22 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

regsafe() currently requires an exact match for PTR_TO_STACK registers.
This prevents an explored state from pruning a current state whose
variable-offset range is a subset of the explored range.

This change allows a checkpoint containing a widened stack pointer to
cover narrower offsets reached on subsequent loop iterations,
for example [-56, -8] within [-64, -8].

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/states.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 18bf7b660c2f..29d18520d27c 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -644,7 +644,12 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold,
 		return range_within(rold, rcur) &&
 		       tnum_in(rold->var_off, rcur->var_off);
 	case PTR_TO_STACK:
-		return regs_exact(rold, rcur, idmap);
+		return memcmp(rold, rcur, offsetof(struct bpf_reg_state, var_off)) == 0 &&
+		       range_within(rold, rcur) &&
+		       tnum_in(rold->var_off, rcur->var_off) &&
+		       check_ids(rold->id, rcur->id, idmap) &&
+		       check_ids(rold->parent_id, rcur->parent_id, idmap) &&
+		       rold->frameno == rcur->frameno;
 	case PTR_TO_ARENA:
 		return true;
 	case PTR_TO_INSN:

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 14/36] bpf: representation for intervals with steps
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (12 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:35   ` sashiko-bot
  2026-09-26 14:20 ` [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
                   ` (21 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Extend scalar register tracking with a linear "base + step * k" description:
each scalar carries two u16 fields, base and step (invariant: base < step),
recording that its value is some point on the line base + step * k.
A fresh scalar starts as base=0, step=1, i.e. any integer.

This lets the verifier reason about strided values, e.g. an index multiplied
by an element size, which neither the min/max bounds nor the tnum can capture
for a non-power-of-2 stride. Such reasoning is necessary to allow
array access with a variable offset through PTR_TO_BTF_ID, e.g.
for expressions like p->arr[i].field.

reg_bounds_sync() uses the description in two ways:
- deduce_bounds_64_from_step() intersects the 64-bit range with the line via
  cnum64_intersect_linear(), snapping smin/smax to actual points on it;
- __reg_bound_offset() clears the low bits of var_off implied by the number
  of trailing zeros shared by base and step.

range_within() is extended to check whether the cur register's equation
describes a subset of the points described by the old register's equation.

Scalar ALU operations update base/step as follows:
- ADD of a constant (including pointer + scalar): keep step, shift base;
- MUL by a constant: scale step and base by the constant;
- LSH by a constant: scale step and base by the corresponding power of two;
- All other operations reset to base=0, step=1.

Branch (eq/neq) checks do not take into account base/step yet.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  7 ++++
 kernel/bpf/log.c             |  2 +
 kernel/bpf/states.c          | 26 ++++++++++++-
 kernel/bpf/verifier.c        | 87 ++++++++++++++++++++++++++++++++++++++++++++
 4 files changed, 120 insertions(+), 2 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 075a28de2a04..3787de11f689 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -178,6 +178,13 @@ struct bpf_reg_state {
 	 * during state comparisons.
 	 */
 	u32 map_uid;
+	/*
+	 * The value described by this register is some point lying on
+	 * a line described by a linear equation base + step * k.
+	 * Invariant: base < step.
+	 */
+	u16 base;
+	u16 step;
 	/* if (!precise && SCALAR_VALUE) min/max/tnum don't affect safety */
 	bool precise;
 };
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index d850a7863d2e..900d1bb1988b 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -692,6 +692,8 @@ static void print_reg_state(struct bpf_verifier_env *env,
 
 			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
 			verbose_a("var_off=%s", tn_buf);
+			if (reg->base != 0 || reg->step != 1)
+				verbose_a("step=%d+%d", reg->base, reg->step);
 		}
 	}
 	verbose(env, ")");
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 29d18520d27c..d4fa98998543 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -302,8 +302,30 @@ int bpf_update_branch_counts(struct bpf_verifier_env *env, struct bpf_verifier_s
 static bool range_within(const struct bpf_reg_state *old,
 			 const struct bpf_reg_state *cur)
 {
-	return cnum64_is_subset(old->r64, cur->r64) &&
-	       cnum32_is_subset(old->r32, cur->r32);
+	if (!cnum64_is_subset(old->r64, cur->r64) ||
+	    !cnum32_is_subset(old->r32, cur->r32))
+		return false;
+
+	if (old->step <= 1)
+		return true;
+
+	if (cnum64_is_const(cur->r64))
+		return imod((s64)cnum64_smin(cur->r64), old->step) == old->base;
+
+	/*
+	 * Both `old` and `cur` define some sets of points.
+	 * Return true, if points defined by `cur` are a subset of points defined by `old`:
+	 * - bounds for `cur` should be within bounds for `old`;
+	 * - cur->step should be dividable by old->step;
+	 * - cur->base should start at integer number of old->step
+	 *   steps from old->base.
+	 *
+	 * E.g. the following ranges are compatible:
+	 * - old [0, 10] step 2
+	 * - new [2, 10] step 4
+	 */
+	return cur->step % old->step == 0 &&
+	       (cur->base - old->base) % old->step == 0;
 }
 
 /* If in the old state two registers had the same id, then they need to have
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 79aa92486154..89b1a0aa3a25 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5,6 +5,7 @@
  */
 #include <uapi/linux/btf.h>
 #include <linux/bpf-cgroup.h>
+#include <linux/count_zeros.h>
 #include <linux/kernel.h>
 #include <linux/types.h>
 #include <linux/slab.h>
@@ -1911,12 +1912,19 @@ static void bpf_diag_record_caller_saved(struct bpf_verifier_env *env,
 	}
 }
 
+static void reg_step_reset(struct bpf_reg_state *reg)
+{
+	reg->base = 0;
+	reg->step = 1;
+}
+
 /* This helper doesn't clear reg->id */
 static void ___mark_reg_known(struct bpf_reg_state *reg, u64 imm)
 {
 	reg->var_off = tnum_const(imm);
 	reg->r64 = cnum64_from_urange(imm, imm);
 	reg->r32 = cnum32_from_urange((u32)imm, (u32)imm);
+	reg_step_reset(reg);
 }
 
 /* Mark the unknown part of a register (variable offset or scalar value) as
@@ -2165,10 +2173,16 @@ static void deduce_bounds_64_from_32(struct bpf_reg_state *reg)
 	reg->r64 = cnum64_cnum32_intersect(reg->r64, reg->r32);
 }
 
+static void deduce_bounds_64_from_step(struct bpf_reg_state *reg)
+{
+	reg->r64 = cnum64_intersect_linear(reg->r64, reg->base, reg->step);
+}
+
 static void __reg_deduce_bounds(struct bpf_reg_state *reg)
 {
 	deduce_bounds_32_from_64(reg);
 	deduce_bounds_64_from_32(reg);
+	deduce_bounds_64_from_step(reg);
 }
 
 /* Attempts to improve var_off based on unsigned min/max information */
@@ -2180,8 +2194,18 @@ static void __reg_bound_offset(struct bpf_reg_state *reg)
 	struct tnum var32_off = tnum_intersect(tnum_subreg(var64_off),
 					       tnum_range(reg_u32_min(reg),
 							  reg_u32_max(reg)));
+	u32 trailing_zero_bits;
+	u16 base = reg->base;
+	u16 step = reg->step;
 
 	reg->var_off = tnum_or(tnum_clear_subreg(var64_off), var32_off);
+
+	if (base == 0)
+		trailing_zero_bits = count_trailing_zeros(step);
+	else
+		trailing_zero_bits = min(count_trailing_zeros(base),
+					 count_trailing_zeros(step));
+	reg->var_off = tnum_and(reg->var_off, tnum_const(~0ULL << trailing_zero_bits));
 }
 
 static bool range_bounds_violation(struct bpf_reg_state *reg);
@@ -2257,6 +2281,7 @@ static int reg_bounds_sanity_check(struct bpf_verifier_env *env,
 	if (env->test_reg_invariants)
 		return -EFAULT;
 	__mark_reg_unbounded(reg);
+	reg_step_reset(reg);
 	return 0;
 }
 
@@ -2267,6 +2292,7 @@ void bpf_mark_reg_unknown_imprecise(struct bpf_reg_state *reg)
 	reg->type = SCALAR_VALUE;
 	reg->var_off = tnum_unknown;
 	__mark_reg_unbounded(reg);
+	reg_step_reset(reg);
 }
 
 /* Mark a register as having a completely unknown (scalar) value,
@@ -15589,6 +15615,50 @@ static int sanitize_check_bounds(struct bpf_verifier_env *env,
 	return 0;
 }
 
+static void scalar_step_add(struct bpf_reg_state *dst_reg,
+			    const struct bpf_reg_state *a,
+			    const struct bpf_reg_state *b)
+{
+	u16 base, step;
+
+	/* If either 'a' or 'b' is a constant, update the base/step for the counterpart. */
+	if (tnum_is_const(b->var_off)) {
+		step = a->step;
+		base = imod((s64)a->base + (s64)b->var_off.value, step);
+	} else if (tnum_is_const(a->var_off)) {
+		step = b->step;
+		base = imod((s64)b->base + (s64)a->var_off.value, step);
+	} else {
+		step = 1;
+		base = 0;
+	}
+	dst_reg->base = base;
+	dst_reg->step = step;
+}
+
+static void scalar_step_mul(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg)
+{
+	u64 amount = src_reg->var_off.value;
+
+	if (tnum_is_const(src_reg->var_off) && (s64)amount >= 0 &&
+	    !check_mul_overflow(dst_reg->step, amount, &dst_reg->step) &&
+	    dst_reg->step != 0)
+		dst_reg->base = (dst_reg->base * amount) % dst_reg->step;
+	else
+		reg_step_reset(dst_reg);
+}
+
+static void scalar_step_lsh(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg)
+{
+	u64 amount = src_reg->var_off.value;
+
+	if (tnum_is_const(src_reg->var_off) && amount < 64 &&
+	    !check_mul_overflow(dst_reg->step, 1ull << amount, &dst_reg->step))
+		dst_reg->base = (dst_reg->base * (1ull << amount)) % dst_reg->step;
+	else
+		reg_step_reset(dst_reg);
+}
+
 /* Handles arithmetic on a pointer and a scalar: computes new min/max and var_off.
  * Caller should also handle BPF_MOV case separately.
  * If we return -EACCES, caller may want to try again treating pointer as a
@@ -15746,6 +15816,7 @@ static int adjust_ptr_min_max_vals(struct bpf_verifier_env *env, struct bpf_insn
 		 * added into the variable offset, and we copy the fixed offset
 		 * from ptr_reg.
 		 */
+		scalar_step_add(dst_reg, ptr_reg, off_reg);
 		dst_reg->r64 = cnum64_add(ptr_reg->r64, off_reg->r64);
 		dst_reg->var_off = tnum_add(ptr_reg->var_off, off_reg->var_off);
 		dst_reg->raw = ptr_reg->raw;
@@ -15807,6 +15878,7 @@ static int adjust_ptr_min_max_vals(struct bpf_verifier_env *env, struct bpf_insn
 			if ((!known && smin_val < 0) || dst_reg->range < 0)
 				memset(&dst_reg->raw, 0, sizeof(dst_reg->raw));
 		}
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_AND:
 	case BPF_OR:
@@ -16506,6 +16578,7 @@ static void scalar_byte_swap(struct bpf_reg_state *dst_reg, struct bpf_insn *ins
 		 * Bounds will be re-derived from the new tnum later.
 		 */
 		__mark_reg_unbounded(dst_reg);
+		reg_step_reset(dst_reg);
 	}
 	/* For bswap16/32, truncate dst register to match the swapped size */
 	if (insn->imm == 16 || insn->imm == 32)
@@ -16632,6 +16705,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 	 */
 	switch (opcode) {
 	case BPF_ADD:
+		scalar_step_add(dst_reg, dst_reg, &src_reg);
 		scalar32_min_max_add(dst_reg, &src_reg);
 		scalar_min_max_add(dst_reg, &src_reg);
 		dst_reg->var_off = tnum_add(dst_reg->var_off, src_reg.var_off);
@@ -16640,6 +16714,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		scalar32_min_max_sub(dst_reg, &src_reg);
 		scalar_min_max_sub(dst_reg, &src_reg);
 		dst_reg->var_off = tnum_sub(dst_reg->var_off, src_reg.var_off);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_NEG:
 		env->fake_reg[0] = *dst_reg;
@@ -16647,11 +16722,13 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		scalar32_min_max_sub(dst_reg, &env->fake_reg[0]);
 		scalar_min_max_sub(dst_reg, &env->fake_reg[0]);
 		dst_reg->var_off = tnum_neg(env->fake_reg[0].var_off);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_MUL:
 		dst_reg->var_off = tnum_mul(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_mul(dst_reg, &src_reg);
 		scalar_min_max_mul(dst_reg, &src_reg);
+		scalar_step_mul(dst_reg, &src_reg);
 		break;
 	case BPF_DIV:
 		/* BPF div specification: x / 0 = 0 */
@@ -16669,6 +16746,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 				scalar_min_max_sdiv(dst_reg, &src_reg);
 			else
 				scalar_min_max_udiv(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_MOD:
 		/* BPF mod specification: x % 0 = x */
@@ -16684,6 +16762,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 				scalar_min_max_smod(dst_reg, &src_reg);
 			else
 				scalar_min_max_umod(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_AND:
 		if (tnum_is_const(src_reg.var_off)) {
@@ -16694,6 +16773,7 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		dst_reg->var_off = tnum_and(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_and(dst_reg, &src_reg);
 		scalar_min_max_and(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_OR:
 		if (tnum_is_const(src_reg.var_off)) {
@@ -16704,32 +16784,38 @@ static int adjust_scalar_min_max_vals(struct bpf_verifier_env *env,
 		dst_reg->var_off = tnum_or(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_or(dst_reg, &src_reg);
 		scalar_min_max_or(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_XOR:
 		dst_reg->var_off = tnum_xor(dst_reg->var_off, src_reg.var_off);
 		scalar32_min_max_xor(dst_reg, &src_reg);
 		scalar_min_max_xor(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_LSH:
 		if (alu32)
 			scalar32_min_max_lsh(dst_reg, &src_reg);
 		else
 			scalar_min_max_lsh(dst_reg, &src_reg);
+		scalar_step_lsh(dst_reg, &src_reg);
 		break;
 	case BPF_RSH:
 		if (alu32)
 			scalar32_min_max_rsh(dst_reg, &src_reg);
 		else
 			scalar_min_max_rsh(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_ARSH:
 		if (alu32)
 			scalar32_min_max_arsh(dst_reg, &src_reg);
 		else
 			scalar_min_max_arsh(dst_reg, &src_reg);
+		reg_step_reset(dst_reg);
 		break;
 	case BPF_END:
 		scalar_byte_swap(dst_reg, insn);
+		reg_step_reset(dst_reg);
 		break;
 	default:
 		break;
@@ -17675,6 +17761,7 @@ static void regs_refine_cond_op(struct bpf_reg_state *reg1, struct bpf_reg_state
 		 * violations if we're on a dead branch.
 		 */
 		__mark_reg_unbounded(reg1);
+		reg_step_reset(reg1);
 		if (is_jmp32) {
 			t = tnum_and(tnum_subreg(reg1->var_off), tnum_const(~val));
 			reg1->var_off = tnum_with_subreg(reg1->var_off, t);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (13 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 14/36] bpf: representation for intervals with steps Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:37   ` sashiko-bot
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
                   ` (20 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Relax the requirements for reads like `p->arr[i].field`, where `p` is a
PTR_TO_BTF_ID, by allowing `i` to be a non-constant value.
This is useful for loops over arrays accessed from a PTR_TO_BTF_ID
when the verifier widens the loop induction variable.

Check that the access stays within array bounds based on the access
register's min and max bounds. Check that the access is aligned with
array elements using the register's `base` and `step`.

Technically, modify btf_struct_access() such that:
- it identifies `field` by walking BTF from the minimum offset
  (the register's reg_smin + static instruction offset);
- it remembers which arrays contain `field`;
- it checks that offsets `k * step` stay within the array
  and point to the same `field`;
- it does not check upper bounds for flexible arrays
  (as is already the case for constant-offset access).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/btf.c      | 149 +++++++++++++++++++++++++++++++++++++++++++-------
 kernel/bpf/states.c   |   1 +
 kernel/bpf/verifier.c |  38 ++++++++-----
 3 files changed, 152 insertions(+), 36 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 9bcfefdfb734..37437b93471d 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -7496,6 +7496,14 @@ bool btf_ctx_access(int off, int size, enum bpf_access_type type,
 }
 EXPORT_SYMBOL_GPL(btf_ctx_access);
 
+#define MAX_ARRAYS_WALK 16
+
+/* Description of an array crossed while walking a BTF access chain */
+struct array_access {
+	u32 id;		/* BTF_KIND_ARRAY type id */
+	u32 off;	/* offset of the access within the array */
+};
+
 enum bpf_struct_walk_result {
 	/* < 0 error */
 	WALK_SCALAR = 0,
@@ -7504,10 +7512,43 @@ enum bpf_struct_walk_result {
 	WALK_STRUCT,
 };
 
+static const struct btf_member *find_flex_member(const struct btf *btf, const struct btf_type *t)
+{
+	const struct btf_array *array;
+	const struct btf_type *mtype;
+	const struct btf_member *member;
+	u32 vlen = btf_type_vlen(t);
+
+	if (vlen == 0)
+		return NULL;
+
+	member = btf_type_member(t) + vlen - 1;
+	mtype = btf_type_skip_modifiers(btf, member->type, NULL);
+	if (!btf_type_is_array(mtype))
+		return NULL;
+
+	array = (const struct btf_array *)(mtype + 1);
+	if (array->nelems != 0)
+		return NULL;
+
+	return member;
+}
+
+static int push_array(struct array_access *arrays, u32 *arrays_cnt, struct array_access item)
+{
+	if (arrays == NULL)
+		return 0;
+	if (*arrays_cnt == MAX_ARRAYS_WALK)
+		return -1;
+	arrays[(*arrays_cnt)++] = item;
+	return 0;
+}
+
 static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 			   const struct btf_type *t, int off, int size,
 			   u32 *next_btf_id, enum bpf_type_flag *flag,
-			   const char **field_name, bool walk_flex_arrays)
+			   const char **field_name, bool walk_flex_arrays,
+			   struct array_access *arrays, u32 *arrays_cnt)
 {
 	u32 i, moff, mtrue_end, msize = 0, total_nelems = 0;
 	const struct btf_type *mtype, *elem_type = NULL;
@@ -7542,22 +7583,19 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 		/* If the last element is a variable size array, we may
 		 * need to relax the rule.
 		 */
-		if (vlen == 0)
+		member = find_flex_member(btf, t);
+		if (!member)
 			goto error;
 
-		member = btf_type_member(t) + vlen - 1;
-		mtype = btf_type_skip_modifiers(btf, member->type,
-						NULL);
-		if (!btf_type_is_array(mtype))
+		moff = __btf_member_bit_offset(t, member) / 8;
+		if (off < moff)
 			goto error;
 
+		mtype = btf_type_skip_modifiers(btf, member->type, &mid);
 		array_elem = (struct btf_array *)(mtype + 1);
-		if (array_elem->nelems != 0)
-			goto error;
 
-		moff = __btf_member_bit_offset(t, member) / 8;
-		if (off < moff)
-			goto error;
+		if (push_array(arrays, arrays_cnt, (struct array_access){ mid, off - moff }))
+			return -E2BIG;
 
 		/* allow structure and integer */
 		t = btf_type_skip_modifiers(btf, array_elem->type,
@@ -7680,6 +7718,9 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 			 *      the array's element as long as it is
 			 *      within the mtrue_end boundary.
 			 */
+			if (push_array(arrays, arrays_cnt,
+				       (struct array_access){ mid, off - moff }))
+				return -E2BIG;
 
 			/* skip empty array */
 			if (moff == mtrue_end)
@@ -7772,15 +7813,29 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
 
 int btf_struct_access(struct bpf_verifier_log *log,
 		      const struct bpf_reg_state *reg,
-		      int off, int size, enum bpf_access_type atype __maybe_unused,
+		      int _off, int size, enum bpf_access_type atype __maybe_unused,
 		      u32 *next_btf_id, enum bpf_type_flag *flag,
 		      const char **field_name)
 {
 	const struct btf *btf = reg->btf;
 	enum bpf_type_flag tmp_flag = 0;
+	struct array_access arrays[MAX_ARRAYS_WALK];
+	const struct btf_member *member;
 	const struct btf_type *t;
+	u32 i, arrays_cnt = 0;
 	u32 id = reg->btf_id;
-	int err;
+	s64 min_off, max_off, flex_off;
+	int err, off, ret;
+
+	if (check_add_overflow(reg_smin(reg), (s64)_off, &min_off) ||
+	    check_add_overflow(reg_smax(reg), (s64)_off, &max_off) ||
+	    min_off != (int)min_off) {
+		bpf_log(log,
+			"offset computation overflows: register offset range is [%lld, %lld], instruction offset is %d\n",
+			reg_smin(reg), reg_smax(reg), _off);
+		return -EINVAL;
+	}
+	off = min_off;
 
 	while (type_is_alloc(reg->type)) {
 		struct btf_struct_meta *meta;
@@ -7807,26 +7862,30 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	t = btf_type_by_id(btf, id);
 	do {
 		err = btf_struct_walk(log, btf, t, off, size, &id, &tmp_flag,
-				      field_name, !type_is_alloc(reg->type));
-
+				      field_name, !type_is_alloc(reg->type), arrays, &arrays_cnt);
 		switch (err) {
 		case WALK_PTR:
 			/* For local types, the destination register cannot
 			 * become a pointer again.
 			 */
-			if (type_is_alloc(reg->type))
-				return SCALAR_VALUE;
+			if (type_is_alloc(reg->type)) {
+				ret = SCALAR_VALUE;
+				goto check_variable_offset;
+			}
 			/* If we found the pointer or scalar on t+off,
 			 * we're done.
 			 */
 			*next_btf_id = id;
 			*flag = tmp_flag;
-			return PTR_TO_BTF_ID;
+			ret = PTR_TO_BTF_ID;
+			goto check_variable_offset;
 		case WALK_PTR_UNTRUSTED:
 			*flag = MEM_RDONLY | PTR_UNTRUSTED;
-			return PTR_TO_MEM;
+			ret = PTR_TO_MEM;
+			goto check_variable_offset;
 		case WALK_SCALAR:
-			return SCALAR_VALUE;
+			ret = SCALAR_VALUE;
+			goto check_variable_offset;
 		case WALK_STRUCT:
 			/* We found nested struct, so continue the search
 			 * by diving in it. At this point the offset is
@@ -7846,6 +7905,54 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	} while (t);
 
 	return -EINVAL;
+
+check_variable_offset:
+	if (min_off == max_off)
+		return ret;
+
+	/* Find an offset at which access would go to a flexible array tail (if any). */
+	t = btf_type_skip_modifiers(btf, reg->btf_id, NULL);
+	member = find_flex_member(btf, t);
+	flex_off = member ? __btf_member_bit_offset(t, member) / 8 : S64_MAX;
+
+	/*
+	 * If this is a varying offset access, the step recorded within a register
+	 * should correspond to one of the arrays visited while walking.
+	 */
+	for (i = 0; i < arrays_cnt; i++) {
+		s64 array_start, array_end, access_end;
+		u32 asize, esize, elem_id;
+
+		/*
+		 * 'arrays' records top-level arrays only, use __btf_resolve_size()
+		 * to get flattened representation.
+		 */
+		t = btf_type_by_id(btf, arrays[i].id);
+		if (IS_ERR(__btf_resolve_size(btf, t, &asize, NULL, &elem_id, NULL, NULL)) ||
+		    IS_ERR(btf_resolve_size(btf, btf_type_by_id(btf, elem_id), &esize)))
+			continue;
+		/* Make sure every accessed offset lands on the same field of some element. */
+		if (esize == 0 || reg->step % esize != 0)
+			continue;
+
+		array_start = min_off - arrays[i].off;
+		/* An array within the flexible tail has no upper bound to exceed */
+		if (array_start >= flex_off)
+			return ret;
+
+		/* the 'size' bytes read at max_off must stay within the array */
+		if (check_add_overflow(array_start, (s64)asize, &array_end) ||
+		    check_add_overflow(max_off, (s64)size, &access_end) ||
+		    access_end > array_end)
+			continue;
+		return ret;
+	}
+
+	bpf_log(log,
+		"invalid variable offset access into struct %s: offsets [%lld, %lld] size %d step %d do not match any array\n",
+		btf_name_by_offset(btf, btf_type_by_id(btf, reg->btf_id)->name_off),
+		min_off, max_off, size, reg->step);
+	return -EINVAL;
 }
 
 /* Check that two BTF types, each specified as an BTF object + id, are exactly
@@ -7886,7 +7993,7 @@ bool btf_struct_ids_match(struct bpf_verifier_log *log,
 	if (!type)
 		return false;
 	err = btf_struct_walk(log, btf, type, off, 1, &id, &flag, NULL,
-			      walk_flex_arrays);
+			      walk_flex_arrays, NULL, NULL);
 	if (err != WALK_STRUCT)
 		return false;
 
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index d4fa98998543..3eaae1452c54 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -632,6 +632,7 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold,
 	case PTR_TO_MEM:
 	case PTR_TO_BUF:
 	case PTR_TO_TP_BUFFER:
+	case PTR_TO_BTF_ID:
 		/* If the new min/max/var_off satisfy the old ones and
 		 * everything else matches, we are OK.
 		 */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 89b1a0aa3a25..a78a9924a610 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6372,6 +6372,7 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	const char *field_name = NULL;
 	enum bpf_type_flag flag = 0;
 	u32 btf_id = 0;
+	s64 min_off;
 	int ret;
 
 	if (!env->allow_ptr_leaks) {
@@ -6387,36 +6388,31 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 		return -EINVAL;
 	}
 
-	if (!tnum_is_const(reg->var_off)) {
-		char tn_buf[48];
-
-		tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
+	if (check_add_overflow(reg_smin(reg), off, &min_off)) {
 		verbose(env,
-			"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
-			reg_arg_name(env, argno), tname, off, tn_buf);
+			"%s is ptr_%s access, offset computation overflows: register's minimal offset is %lld, instruction offset is %d\n",
+			reg_arg_name(env, argno), tname, reg_smin(reg), off);
 		return -EACCES;
 	}
 
-	off += reg->var_off.value;
-
-	if (off < 0) {
+	if (min_off < 0) {
 		verbose(env,
-			"%s is ptr_%s invalid negative access: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s invalid negative access: off=%lld\n",
+			reg_arg_name(env, argno), tname, min_off);
 		return -EACCES;
 	}
 
 	if (reg->type & MEM_USER) {
 		verbose(env,
-			"%s is ptr_%s access user memory: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s access user memory\n",
+			reg_arg_name(env, argno), tname);
 		return -EACCES;
 	}
 
 	if (reg->type & MEM_PERCPU) {
 		verbose(env,
-			"%s is ptr_%s access percpu memory: off=%d\n",
-			reg_arg_name(env, argno), tname, off);
+			"%s is ptr_%s access percpu memory\n",
+			reg_arg_name(env, argno), tname);
 		return -EACCES;
 	}
 
@@ -6426,6 +6422,18 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	}
 
 	if (env->ops->btf_struct_access && !type_is_alloc(reg->type) && atype == BPF_WRITE) {
+		if (!tnum_is_const(reg->var_off)) {
+			char tn_buf[48];
+
+			tnum_strn(tn_buf, sizeof(tn_buf), reg->var_off);
+			verbose(env,
+				"%s is ptr_%s invalid variable offset: off=%d, var_off=%s\n",
+				reg_arg_name(env, argno), tname, off, tn_buf);
+			return -EACCES;
+		}
+
+		off += reg->var_off.value;
+
 		if (!btf_is_kernel(reg->btf)) {
 			verifier_bug(env, "reg->btf must be kernel btf");
 			return -EFAULT;

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (14 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:31   ` sashiko-bot
  2026-09-26 14:20 ` [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
                   ` (19 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Dominator intersections and SCEV worklist scheduling need to compare
instructions by their DFS postorder rank. Save per-instruction
postorder numbers alongside the existing postorder sequence.

This commit modifies bpf_compute_postorder() to actually do the DFS
traversal, instead of scheduling traversal of all siblings at once.

Consider a graph:

  A -> B
  A -> C
  C -> B

Old result: C, B, A
DFS result: B, C, A

Fixes: efcda22aa541 ("bpf: compute instructions postorder per subprogram")
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/cfg.c             | 56 +++++++++++++++++++++++++++-----------------
 kernel/bpf/verifier.c        |  1 +
 3 files changed, 37 insertions(+), 21 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 3787de11f689..6a8157e607ee 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1033,6 +1033,7 @@ struct bpf_verifier_env {
 		 * see bpf_subprog_info->postorder_start.
 		 */
 		int *insn_postorder;
+		int *postorder_nums;
 		int cur_stack;
 		/* current position in the insn_postorder vector */
 		int cur_postorder;
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index b0bd9ba951df..bd771efec66a 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -767,6 +767,11 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
 	return ret;
 }
 
+struct dfs_state {
+	u32 traversed:1;
+	u32 next_succ:31;
+};
+
 /*
  * For each subprogram 'i' fill array env->cfg.insn_subprogram sub-range
  * [env->subprog_info[i].postorder_start, env->subprog_info[i+1].postorder_start)
@@ -774,43 +779,52 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
  */
 int bpf_compute_postorder(struct bpf_verifier_env *env)
 {
-	u32 cur_postorder, i, top, stack_sz, s;
-	int *stack = NULL, *postorder = NULL, *state = NULL;
-	struct bpf_iarray *succ;
+	int *stack = NULL, *postorder = NULL, *postorder_nums = NULL;
+	int subprog_idx, stack_sz, cur, s, cur_postorder, start;
+	struct dfs_state *state = NULL;
+	struct bpf_iarray *succ = NULL;
 
+	postorder_nums = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
 	postorder = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
-	state = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
 	stack = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);
-	if (!postorder || !state || !stack) {
+	state = kvzalloc_objs(struct dfs_state, env->prog->len, GFP_KERNEL_ACCOUNT);
+	if (!postorder_nums || !postorder || !stack || !state) {
+		kvfree(postorder_nums);
 		kvfree(postorder);
-		kvfree(state);
 		kvfree(stack);
+		kvfree(state);
 		return -ENOMEM;
 	}
 	cur_postorder = 0;
-	for (i = 0; i < env->subprog_cnt; i++) {
-		env->subprog_info[i].postorder_start = cur_postorder;
-		stack[0] = env->subprog_info[i].start;
+	for (subprog_idx = 0; subprog_idx < env->subprog_cnt; subprog_idx++) {
+		start = env->subprog_info[subprog_idx].start;
+		env->subprog_info[subprog_idx].postorder_start = cur_postorder;
+		stack[0] = start;
 		stack_sz = 1;
+		state[start].traversed = true;
 		do {
-			top = stack[stack_sz - 1];
-			state[top] |= DISCOVERED;
-			if (state[top] & EXPLORED) {
-				postorder[cur_postorder++] = top;
+			cur = stack[stack_sz - 1];
+			succ = bpf_insn_successors(env, cur);
+			if (state[cur].next_succ == succ->cnt) {
+				postorder_nums[cur] = cur_postorder;
+				postorder[cur_postorder] = cur;
+				cur_postorder++;
 				stack_sz--;
 				continue;
 			}
-			succ = bpf_insn_successors(env, top);
-			for (s = 0; s < succ->cnt; ++s) {
-				if (!state[succ->items[s]]) {
-					stack[stack_sz++] = succ->items[s];
-					state[succ->items[s]] |= DISCOVERED;
-				}
+			s = succ->items[state[cur].next_succ];
+			if (!state[s].traversed) {
+				state[s].traversed = true;
+				state[s].next_succ = 0;
+				stack[stack_sz] = s;
+				stack_sz++;
+				continue;
 			}
-			state[top] |= EXPLORED;
+			state[cur].next_succ++;
 		} while (stack_sz);
 	}
-	env->subprog_info[i].postorder_start = cur_postorder;
+	env->subprog_info[subprog_idx].postorder_start = cur_postorder;
+	env->cfg.postorder_nums = postorder_nums;
 	env->cfg.insn_postorder = postorder;
 	env->cfg.cur_postorder = cur_postorder;
 	kvfree(stack);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a78a9924a610..40ed20693936 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22933,6 +22933,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
 	bpf_stack_liveness_free(env);
+	kvfree(env->cfg.postorder_nums);
 	kvfree(env->cfg.insn_postorder);
 	kvfree(env->scc_info);
 	kvfree(env->succ);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (15 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 18/36] bpf: compute immediate dominators Eduard Zingerman
                   ` (18 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	Emil Tsalapatis

Tests for multiple analyses performed by the verifier need the
verifier log to contain analysis results alongside the program
disassembly. This patch moves the log-level-2 program dump from
bpf_compute_live_registers() to a standalone function called from
bpf_check(), in order to provide a common logging function for such
analyses (and thus avoid printing program disassembly multiple times).

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 kernel/bpf/liveness.c                              | 28 +---------------
 kernel/bpf/verifier.c                              | 37 ++++++++++++++++++++++
 .../selftests/bpf/progs/verifier_live_stack.c      |  2 +-
 3 files changed, 39 insertions(+), 28 deletions(-)

diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 6b2f1bdbef7f..cba3a6c021b4 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -2631,8 +2631,7 @@ int bpf_compute_live_registers(struct bpf_verifier_env *env)
 	struct bpf_insn *insns = env->prog->insnsi;
 	struct insn_live_regs *state;
 	int insn_cnt = env->prog->len;
-	u64 pos, insn_pos;
-	int err = 0, i, j, subprog, start, end;
+	int err = 0, i, subprog, start, end;
 	bool changed, ret_reg_pair;
 
 	/* Use the following algorithm:
@@ -2711,31 +2710,6 @@ int bpf_compute_live_registers(struct bpf_verifier_env *env)
 		insn_aux[i].zext_dst = def32 >= 0 && (mask_hi(out) & BIT(def32));
 	}
 
-	if (env->log.level & BPF_LOG_LEVEL2) {
-		verbose(env, "Live regs before insn:\n");
-		for (i = 0; i < insn_cnt; ++i) {
-			if (env->insn_aux_data[i].scc)
-				verbose(env, "%3d ", env->insn_aux_data[i].scc);
-			else
-				verbose(env, "    ");
-			verbose(env, "%3d: ", i);
-			for (j = BPF_REG_0; j < BPF_REG_10; ++j)
-				if (insn_aux[i].live_regs_before & BIT(j))
-					verbose(env, "%d", j);
-				else
-					verbose(env, ".");
-			verbose(env, " ");
-			pos = env->log.end_pos;
-			bpf_verbose_insn(env, &insns[i]);
-			insn_pos = env->log.end_pos;
-			if (insn_aux[i].zext_dst)
-				verbose(env, "%*c; zext", bpf_vlog_alignment(insn_pos - pos), ' ');
-			verbose(env, "\n");
-			if (bpf_is_ldimm64(&insns[i]))
-				i++;
-		}
-	}
-
 out:
 	kvfree(state);
 	return err;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 40ed20693936..474caf8696fd 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22580,6 +22580,40 @@ static bool bpf_prog_reenters_datapath(const struct bpf_prog *prog)
 	return false;
 }
 
+/* Various log level 2 information about the program */
+static void log_program(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *insn_aux = env->insn_aux_data;
+	struct bpf_insn *insns = env->prog->insnsi;
+	u32 insn_cnt = env->prog->len;
+	u64 pos, insn_pos;
+	u32 i, j;
+
+	verbose(env, "Program dump (scc? insn#: live_regs_before):\n");
+	for (i = 0; i < insn_cnt; ++i) {
+		verbose_linfo(env, i, "    ; ");
+		if (env->insn_aux_data[i].scc)
+			verbose(env, "%3d ", env->insn_aux_data[i].scc);
+		else
+			verbose(env, "    ");
+		verbose(env, "%3d: ", i);
+		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
+			if (insn_aux[i].live_regs_before & BIT(j))
+				verbose(env, "%d", j);
+			else
+				verbose(env, ".");
+		verbose(env, " ");
+		pos = env->log.end_pos;
+		bpf_verbose_insn(env, &insns[i]);
+		insn_pos = env->log.end_pos;
+		if (insn_aux[i].zext_dst)
+			verbose(env, "%*c; zext", bpf_vlog_alignment(insn_pos - pos), ' ');
+		verbose(env, "\n");
+		if (bpf_is_ldimm64(&insns[i]))
+			i++;
+	}
+}
+
 int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	      struct bpf_log_attr *attr_log)
 {
@@ -22785,6 +22819,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	if (env->log.level & BPF_LOG_LEVEL2)
+		log_program(env);
+
 	ret = mark_fastcall_patterns(env);
 	if (ret < 0)
 		goto skip_full_check;
diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 7e2e165a7145..6b295b30936d 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -66,7 +66,7 @@ __msg("stack use/def subprog#0 must_write_not_same_slot (d0,cs0):")
  * but the write conservatively marks the whole frame as may_def.
  */
 __msg("6: (7b) *(u64 *)(r2 +0) = r0         ; may_def: fp0-8..-{{(512|2048)}}")
-__msg("Live regs before insn:")
+__msg("Program dump")
 __naked void must_write_not_same_slot(void)
 {
 	asm volatile (

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 18/36] bpf: compute immediate dominators
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (16 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 15:54   ` Alexei Starovoitov
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 19/36] bpf: compute loop hierarchy Eduard Zingerman
                   ` (17 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV needs to find an exit condition that dominates a loop's backedge;
such conditions bound every loop iteration. Compute the immediate
dominator tree as a prerequisite for identifying such conditions.

Use the iterative reverse-postorder algorithm from:
"A Simple, Fast Dominance Algorithm" by Cooper et al.

Analyze each subprogram independently and exclude the second half of
ldimm64 instructions.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |   8 +++
 kernel/bpf/Makefile          |   2 +-
 kernel/bpf/loops.c           | 143 +++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |   8 ++-
 4 files changed, 159 insertions(+), 2 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6a8157e607ee..5e6fa41a1353 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -650,6 +650,11 @@ struct bpf_iarray {
 	u32 items[];
 };
 
+#define iarray_for_each(item, arr)						\
+	for (int ___idx = 0;							\
+	     ___idx < (arr)->cnt && ({ item = (arr)->items[___idx]; 1; });	\
+	     ___idx++)
+
 struct bpf_insn_aux_data {
 	union {
 		enum bpf_reg_type ptr_type;	/* pointer type for load/store insns */
@@ -1109,6 +1114,7 @@ struct bpf_verifier_env {
 	u32 scc_cnt;
 	struct bpf_iarray *succ;
 	struct bpf_iarray *gotox_tmp_buf;
+	int *idoms;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1869,4 +1875,6 @@ int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 int bpf_flip_opcode(u32 opcode);
 u8 bpf_rev_opcode(u8 opcode);
 
+int bpf_compute_idoms(struct bpf_verifier_env *env);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..6210601eb40b 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
new file mode 100644
index 000000000000..e7ff02e9ab3a
--- /dev/null
+++ b/kernel/bpf/loops.c
@@ -0,0 +1,143 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/slab.h>
+#include <linux/bpf_verifier.h>
+
+static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
+{
+	struct bpf_iarray *succ, *preds, **result;
+	struct bpf_prog *prog = env->prog;
+	u32 *num_preds, i, s, sz, len = prog->len;
+	struct bpf_insn *insn;
+	void *tmp;
+
+	num_preds = kvcalloc(prog->len, sizeof(u32), GFP_KERNEL_ACCOUNT);
+	if (!num_preds)
+		return NULL;
+
+	/*
+	 * 'result' layout:
+	 *  - array of pointers (struct bpf_iarray *)[len]
+	 *  - struct bpf_iarray one after another
+	 */
+	sz = sizeof(struct bpf_iarray) * len;
+	sz += sizeof(struct bpf_iarray *) * len;
+	for (i = 0; i < len; i++) {
+		insn = env->prog->insnsi + i;
+		succ = bpf_insn_successors(env, i);
+		sz += sizeof(u32) * succ->cnt;
+		iarray_for_each(s, succ) {
+			num_preds[s]++;
+		}
+		if (bpf_is_ldimm64(insn))
+			i++;
+	}
+
+	result = kvzalloc(sz, GFP_KERNEL_ACCOUNT);
+	if (!result) {
+		kvfree(num_preds);
+		return NULL;
+	}
+
+	tmp = (void *)&result[len];
+	for (i = 0; i < len; i++) {
+		result[i] = tmp;
+		tmp += sizeof(struct bpf_iarray);
+		tmp += sizeof(u32) * num_preds[i];
+	}
+
+	for (i = 0; i < len; i++) {
+		insn = env->prog->insnsi + i;
+		succ = bpf_insn_successors(env, i);
+		iarray_for_each(s, succ) {
+			preds = result[s];
+			preds->items[preds->cnt++] = i;
+		}
+		if (bpf_is_ldimm64(insn))
+			i++;
+	}
+
+	kvfree(num_preds);
+	return result;
+}
+
+static int idoms_intersect(struct bpf_verifier_env *env, int a, int b)
+{
+	int *postorder_nums = env->cfg.postorder_nums;
+	int *idoms = env->idoms;
+
+	while (a != b) {
+		while (postorder_nums[a] < postorder_nums[b]) {
+			a = idoms[a];
+		}
+		while (postorder_nums[b] < postorder_nums[a]) {
+			b = idoms[b];
+		}
+	}
+	return a;
+}
+
+/* See "A Simple, Fast Dominance Algorithm" by Cooper et al. for details. */
+static void compute_subprog_idoms(struct bpf_verifier_env *env, struct bpf_iarray **preds, int subprog_idx)
+{
+	struct bpf_subprog_info *subprog = &env->subprog_info[subprog_idx];
+	int start = subprog->start;
+	int po_first = subprog->postorder_start;
+	int po_last = (subprog + 1)->postorder_start - 1;
+	int *idoms = env->idoms;
+	int po_num, pred;
+	bool changed;
+
+	idoms[start] = 0;
+	changed = true;
+	do {
+		changed = false;
+		/* iterate in reverse postorder */
+		for (po_num = po_last; po_num >= po_first; po_num--) {
+			int idx = env->cfg.insn_postorder[po_num];
+			int new_idom = -1;
+
+			iarray_for_each(pred, preds[idx]) {
+				if (idoms[pred] == -1)
+					continue;
+				if (new_idom == -1)
+					new_idom = pred;
+				else
+					new_idom = idoms_intersect(env, pred, new_idom);
+			}
+			if (new_idom != -1 && idoms[idx] != new_idom) {
+				idoms[idx] = new_idom;
+				changed = true;
+			}
+		}
+	} while (changed);
+	idoms[start] = -1;
+}
+
+int bpf_compute_idoms(struct bpf_verifier_env *env)
+{
+	struct bpf_iarray **preds;
+	u32 len = env->prog->len;
+	int *idoms, i;
+
+	preds = compute_predecessors(env);
+	if (!preds)
+		return -ENOMEM;
+
+	idoms = kvcalloc(len, sizeof(*idoms), GFP_KERNEL_ACCOUNT);
+	if (!idoms) {
+		kvfree(preds);
+		return -ENOMEM;
+	}
+
+	env->idoms = idoms;
+	for (i = 0; i < len; i++)
+		idoms[i] = -1;
+
+	for (i = 0; i < env->subprog_cnt; i++)
+		compute_subprog_idoms(env, preds, i);
+
+	kvfree(preds);
+	return 0;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 474caf8696fd..93100541475a 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22589,13 +22589,14 @@ static void log_program(struct bpf_verifier_env *env)
 	u64 pos, insn_pos;
 	u32 i, j;
 
-	verbose(env, "Program dump (scc? insn#: live_regs_before):\n");
+	verbose(env, "Program dump (scc? idom insn#: live_regs_before):\n");
 	for (i = 0; i < insn_cnt; ++i) {
 		verbose_linfo(env, i, "    ; ");
 		if (env->insn_aux_data[i].scc)
 			verbose(env, "%3d ", env->insn_aux_data[i].scc);
 		else
 			verbose(env, "    ");
+		verbose(env, "%3d ", env->idoms[i]);
 		verbose(env, "%3d: ", i);
 		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
 			if (insn_aux[i].live_regs_before & BIT(j))
@@ -22815,6 +22816,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	ret = bpf_compute_idoms(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = bpf_compute_live_registers(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -22978,6 +22983,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	kvfree(env->callx_edges);
 	kvfree(env->func_ptrs);
 	bpf_diag_free(env);
+	kvfree(env->idoms);
 	kvfree(env);
 	return ret;
 }

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 19/36] bpf: compute loop hierarchy
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (17 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 18/36] bpf: compute immediate dominators Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:43   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
                   ` (16 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV needs to process loops from innermost to outermost, summarizing
results for nested loops. Add a pass to compute the loop structure.

Use a non-recursive adaptation of the algorithm from:
"A New Algorithm for Identifying Loops in Decompilation" by Wei et al.

Record the loop hierarchy per instruction:
- The field insn_aux_data->loop_header records the innermost loop
  header containing it.
- The field insn_aux_data->loop records additional information about
  the loop if this instruction is a loop header (exits, backedges,
  irreducibility).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  39 ++++
 kernel/bpf/fixups.c          |  18 ++
 kernel/bpf/loops.c           | 414 +++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |  12 +-
 4 files changed, 482 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 5e6fa41a1353..d8142f1b2347 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -655,6 +655,30 @@ struct bpf_iarray {
 	     ___idx < (arr)->cnt && ({ item = (arr)->items[___idx]; 1; });	\
 	     ___idx++)
 
+#define MAX_BACKEDGES 16
+#define MAX_LOOP_EXITS 256
+
+struct bpf_backedge {
+	int from;
+	int latch; /* -1 if no latch can be found */
+};
+
+struct bpf_loop_exit {
+	int from; /* instruction inside the loop */
+	int to; /* instruction outside the loop */
+};
+
+struct bpf_loop {
+	struct bpf_backedge backedges[MAX_BACKEDGES];
+	/* edges exiting from this loop, includes edges from nested loops */
+	struct bpf_loop_exit *exits;
+	int backedges_cnt;
+	int exits_cnt;
+	bool irreducible;
+	bool backedges_overflow;
+	bool exits_overflow;
+};
+
 struct bpf_insn_aux_data {
 	union {
 		enum bpf_reg_type ptr_type;	/* pointer type for load/store insns */
@@ -732,6 +756,13 @@ struct bpf_insn_aux_data {
 	u32 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
 	u32 jump_target:1;
+	/*
+	 * True if this instruction is a loop entry. For loops with a single entry
+	 * this bit will coincide with 'loop' pointer being non-NULL.
+	 * Irreducible loops have multiple entries, all of them will be marked as 'loop_entry',
+	 * but only the one at 'loop_header' will have a non-NULL 'loop' pointer.
+	 */
+	u32 loop_entry:1;
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
@@ -751,6 +782,10 @@ struct bpf_insn_aux_data {
 	u16 const_reg_map_mask;
 	u16 const_reg_subprog_mask;
 	u32 const_reg_vals[10];
+	/* index of a loop header of the innermost loop containing this instruction, -1 if none */
+	s32 loop_header;
+	/* additional information about the loop if this instruction is a loop header */
+	struct bpf_loop *loop;
 };
 
 #define MAX_USED_MAPS 64 /* max number of maps accessed by one eBPF program */
@@ -1860,6 +1895,7 @@ int bpf_check_attach_btf_id_multi(struct btf *btf, struct bpf_prog *prog, u32 bt
 				  struct bpf_attach_target_info *tgt_info);
 
 /* Functions in fixups.c, called from bpf_check() */
+void bpf_clear_insn_aux_data(struct bpf_verifier_env *env, int start, int len);
 int bpf_remove_fastcall_spills_fills(struct bpf_verifier_env *env);
 int bpf_optimize_bpf_loop(struct bpf_verifier_env *env);
 void bpf_opt_hard_wire_dead_code_branches(struct bpf_verifier_env *env);
@@ -1876,5 +1912,8 @@ int bpf_flip_opcode(u32 opcode);
 u8 bpf_rev_opcode(u8 opcode);
 
 int bpf_compute_idoms(struct bpf_verifier_env *env);
+int bpf_compute_loops(struct bpf_verifier_env *env);
+int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx);
+bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header);
 
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 37cf130ebb57..ecea117f2848 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -614,6 +614,22 @@ static int bpf_adj_linfo_after_remove(struct bpf_verifier_env *env, u32 off,
 	return 0;
 }
 
+/* Clean up dynamically allocated fields of aux data for instructions [start, ...] */
+void bpf_clear_insn_aux_data(struct bpf_verifier_env *env, int start, int len)
+{
+	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
+	int end = start + len;
+	int i;
+
+	for (i = start; i < end; i++) {
+		if (aux_data[i].loop) {
+			kvfree(aux_data[i].loop->exits);
+			kvfree(aux_data[i].loop);
+			aux_data[i].loop = NULL;
+		}
+	}
+}
+
 static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 {
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -626,6 +642,8 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 	if (bpf_prog_is_offloaded(env->prog->aux))
 		bpf_prog_offload_remove_insns(env, off, cnt);
 
+	bpf_clear_insn_aux_data(env, off, cnt);
+
 	err = bpf_remove_insns(env->prog, off, cnt);
 	if (err)
 		return err;
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
index e7ff02e9ab3a..e269870f6a6e 100644
--- a/kernel/bpf/loops.c
+++ b/kernel/bpf/loops.c
@@ -141,3 +141,417 @@ int bpf_compute_idoms(struct bpf_verifier_env *env)
 	kvfree(preds);
 	return 0;
 }
+
+struct dfs_state {
+	u32 traversed:1;
+	u32 next_succ:31;
+};
+
+struct loops_dfs {
+	struct dfs_state *state;
+	int *dfs_pos;
+	int *stack;
+};
+
+static void mark_irreducible(struct bpf_verifier_env *env, int h)
+{
+	env->insn_aux_data[h].loop->irreducible = true;
+}
+
+static void mark_entry(struct bpf_verifier_env *env, int s)
+{
+	env->insn_aux_data[s].loop_entry = true;
+}
+
+static void add_backedge(struct bpf_verifier_env *env, int from, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop = aux[h].loop;
+	int cnt = loop->backedges_cnt;
+
+	if (cnt == MAX_BACKEDGES) {
+		loop->backedges_overflow = true;
+		return;
+	}
+	loop->backedges[cnt].from = from;
+	loop->backedges[cnt].latch = -1;
+	loop->backedges_cnt++;
+}
+
+static int add_exit(struct bpf_loop *loop, int from, int to)
+{
+	if (loop->exits_overflow)
+		return 0;
+	if (loop->exits_cnt == MAX_LOOP_EXITS) {
+		loop->exits_overflow = true;
+		return 0;
+	}
+	if (!loop->exits) {
+		loop->exits = kvcalloc(MAX_LOOP_EXITS, sizeof(*loop->exits), GFP_KERNEL_ACCOUNT);
+		if (!loop->exits)
+			return -ENOMEM;
+	}
+	loop->exits[loop->exits_cnt] = (struct bpf_loop_exit) {
+		.from = from,
+		.to = to,
+	};
+	loop->exits_cnt++;
+	return 0;
+}
+
+static int mark_as_header(struct bpf_verifier_env *env, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+
+	if (!aux[h].loop) {
+		mark_entry(env, h);
+		aux[h].loop = kvzalloc_obj(struct bpf_loop, GFP_KERNEL_ACCOUNT);
+		if (!aux[h].loop)
+			return -ENOMEM;
+	}
+	return 0;
+}
+
+static int assign_header(struct bpf_verifier_env *env, struct loops_dfs *dfs, int n, int h)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int *dfs_pos = dfs->dfs_pos;
+	int err, nh;
+
+	err = mark_as_header(env, h);
+	if (err)
+		return err;
+
+	/* Don't encode self-loops, otherwise can't reflect loops nesting structure. */
+	if (n == h)
+		return 0;
+
+	/* Make sure that loop headers up the chain are sorted by dfs_pos. */
+	while (aux[n].loop_header != -1) {
+		nh = aux[n].loop_header;
+		if (nh == h)
+			return 0;
+		if (dfs_pos[nh] < dfs_pos[h]) {
+			aux[n].loop_header = h;
+			n = h;
+			h = nh;
+		} else {
+			n = nh;
+		}
+	}
+	aux[n].loop_header = h;
+	return 0;
+}
+
+static bool is_cond_jmp_insn(struct bpf_insn *insn)
+{
+	u8 class = BPF_CLASS(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+
+	if (class != BPF_JMP && class != BPF_JMP32)
+		return false;
+
+	switch (opcode) {
+	case BPF_JEQ:
+	case BPF_JGE:
+	case BPF_JGT:
+	case BPF_JLE:
+	case BPF_JLT:
+	case BPF_JNE:
+	case BPF_JSET:
+	case BPF_JSGE:
+	case BPF_JSGT:
+	case BPF_JSLE:
+	case BPF_JSLT:
+		return true;
+	default:
+		return false;
+	}
+}
+
+int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+
+	return aux[idx].loop ? idx : aux[idx].loop_header;
+}
+
+static int find_dominating_condition(struct bpf_verifier_env *env, int n, int top)
+{
+	struct bpf_insn *insns = env->prog->insnsi;
+	int common_dom, n_loop, t_loop, f_loop;
+	int *idoms = env->idoms;
+
+	n_loop = bpf_loop_at_index(env, n);
+	common_dom = idoms_intersect(env, n, top);
+	if (common_dom != top)
+		return -1;
+	while (n >= 0) {
+		if (is_cond_jmp_insn(&insns[n]) && bpf_loop_at_index(env, n) == n_loop) {
+			t_loop = bpf_loop_at_index(env, n + insns[n].off + 1);
+			f_loop = bpf_loop_at_index(env, n + 1);
+			if (f_loop != n_loop && !bpf_is_nested_loop(env, f_loop, n_loop))
+				return n;
+			if (t_loop != n_loop && !bpf_is_nested_loop(env, t_loop, n_loop))
+				return n;
+		}
+		if (n == top)
+			break;
+		n = idoms[n];
+	}
+	return -1;
+}
+
+/*
+ * As described in "A New Algorithm for Identifying Loops in Decompilation" by Wei et al,
+ * adapted to be non-recursive.
+ */
+static int compute_loops_in_subprog(struct bpf_verifier_env *env, struct loops_dfs *dfs,
+				    int subprog_idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct dfs_state *state = dfs->state;
+	int start = env->subprog_info[subprog_idx].start;
+	int *dfs_pos = dfs->dfs_pos;
+	int *stack = dfs->stack;
+	int s, h, err, cur, stack_sz;
+	struct bpf_iarray *succ;
+	u32 i;
+
+	stack[0] = start;
+	state[start].traversed = true;
+	state[start].next_succ = 0;
+	dfs_pos[start] = 1;
+	stack_sz = 1;
+	i = 0;
+	do {
+		/*
+		 * The algorithm should be very fast in practice,
+		 * guard against pathological inputs, just in case.
+		 */
+		if ((++i % 1024) == 0) {
+			if (signal_pending(current))
+				return -EAGAIN;
+			cond_resched();
+		}
+
+		cur = stack[stack_sz - 1];
+		succ = bpf_insn_successors(env, cur);
+		if (state[cur].next_succ == succ->cnt) {
+			dfs_pos[cur] = 0;
+			stack_sz--;
+			continue;
+		}
+		s = succ->items[state[cur].next_succ];
+		if (!state[s].traversed) {
+			/* Case A:  start -> ... -> cur -> s [unexplored] */
+			state[s].traversed = true;
+			state[s].next_succ = 0;
+			stack[stack_sz] = s;
+			dfs_pos[s] = stack_sz + 1;
+			stack_sz++;
+			continue;
+		}
+		/* 's' is fully explored at this point */
+		if (dfs_pos[s]) {
+			/*
+			 * start -> ... -> s -> cur --.
+			 *                 ^          |
+			 *                 '----------'
+			 * Case B: 's' is in the current DFS path.
+			 */
+			err = assign_header(env, dfs, cur, s);
+			if (err)
+				return err;
+			add_backedge(env, cur, s);
+		} else if (aux[s].loop_header == -1) {
+			/*
+			 * start -> ... -> ... -> s -> ... -> end
+			 *           |            ^
+			 *           '---> cur ---'
+			 * Case C: 's' is explored, not in the current DFS path,
+			 * and not a part of any loop.
+			 */
+		} else if (dfs_pos[aux[s].loop_header]) {
+			/*
+			 *                 .----------------------.
+			 *                 v                      |
+			 * start -> ... -> h -> ... -> ... -> s --'
+			 *                       |            ^
+			 *	                 '---> cur ---'
+			 * Case D: 's' is explored, not in current DFS path,
+			 * but its innermost loop header is.
+			 */
+			err = assign_header(env, dfs, cur, aux[s].loop_header);
+			if (err)
+				return err;
+		} else {
+			/*
+			 * case E: 's' is explored, not in current DFS path,
+			 * its innermost loop header is not in current DFS path,
+			 * hence 's' is another entry into the same loop.
+			 */
+			h = aux[s].loop_header;
+			mark_irreducible(env, h);
+			mark_entry(env, s);
+			while (aux[h].loop_header != -1) {
+				h = aux[h].loop_header;
+				if (dfs_pos[h]) {
+					err = assign_header(env, dfs, cur, h);
+					if (err)
+						return err;
+					break;
+				}
+				mark_irreducible(env, h);
+			}
+		}
+		state[cur].next_succ++;
+	} while (stack_sz);
+
+	return 0;
+}
+
+bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int idx;
+
+	for (idx = inner_header; idx >= 0; idx = aux[idx].loop_header)
+		if (aux[idx].loop_header == outer_header)
+			return true;
+
+	return false;
+}
+
+int bpf_compute_loops(struct bpf_verifier_env *env)
+{
+	int i, j, s, t, iloop, sloop, err = 0, len = env->prog->len;
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct bpf_backedge *backedge;
+	struct loops_dfs dfs = {};
+	struct bpf_iarray *succ;
+	struct bpf_loop *loop;
+
+	dfs.dfs_pos = kvcalloc(len, sizeof(int), GFP_KERNEL_ACCOUNT);
+	dfs.state = kvcalloc(len, sizeof(struct dfs_state), GFP_KERNEL_ACCOUNT);
+	dfs.stack = kvcalloc(len, sizeof(int), GFP_KERNEL_ACCOUNT);
+	if (!dfs.dfs_pos || !dfs.state || !dfs.stack) {
+		err = -ENOMEM;
+		goto out;
+	}
+	for (i = 0; i < len; i++)
+		aux[i].loop_header = -1;
+	for (i = 0; i < env->subprog_cnt; i++) {
+		err = compute_loops_in_subprog(env, &dfs, i);
+		if (err)
+			goto out;
+	}
+	/* find latches */
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			backedge = &loop->backedges[j];
+			backedge->latch = find_dominating_condition(env, backedge->from, i);
+		}
+	}
+	/* find exits */
+	for (i = 0; i < len; i++) {
+		iloop = aux[i].loop ? i : aux[i].loop_header;
+		if (iloop < 0)
+			continue;
+		succ = bpf_insn_successors(env, i);
+		iarray_for_each(s, succ) {
+			/*
+			 * Nothing left to record once the innermost loop of 'i'
+			 * overflowed: the walk below would stop at it right away.
+			 */
+			if (aux[iloop].loop->exits_overflow)
+				break;
+			sloop = aux[s].loop ? s : aux[s].loop_header;
+			if (iloop == sloop)
+				continue;
+			if (bpf_is_nested_loop(env, sloop, iloop))
+				continue;
+			/*
+			 * At this point 'sloop' is either -1, an outer loop,
+			 * or a loop in another branch of the loop hierarchy.
+			 */
+			for (t = iloop; t >= 0 && t != sloop; t = aux[t].loop_header) {
+				/*
+				 * Account for the following configuration:
+				 *
+				 *   tloop {
+				 *     iloop {
+				 *       ... i: goto s;
+				 *     }
+				 *     sloop {
+				 *   s:
+				 *       ...
+				 *     }
+				 *   }
+				 */
+				if (bpf_is_nested_loop(env, sloop, t))
+					break;
+				/*
+				 * Record edges i -> s as exits from tloop when:
+				 *
+				 *   sloop {
+				 *     tloop {
+				 *       iloop {
+				 *         ... i: goto s;
+				 *       }
+				 *     }
+				 *  s: ...
+				 *   }
+				 */
+				/* Enclosing loops inherit the overflow, see below */
+				if (aux[t].loop->exits_overflow)
+					break;
+				err = add_exit(aux[t].loop, i, s);
+				if (err)
+					goto out;
+			}
+		}
+		if (bpf_is_ldimm64(env->prog->insnsi + i))
+			i++;
+	}
+	/* A nested loop with a truncated exit list can't be abstracted by SCEV */
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop || !loop->exits_overflow)
+			continue;
+		for (t = aux[i].loop_header; t >= 0; t = aux[t].loop_header)
+			aux[t].loop->exits_overflow = true;
+	}
+
+	if (env->log.level & BPF_LOG_LEVEL2) {
+		for (i = 0; i < len; i++) {
+			loop = aux[i].loop;
+			if (!loop)
+				continue;
+			bpf_log(log, "loop at %d", i);
+			if (aux[i].loop_header >= 0)
+				bpf_log(log, ", nested in %d", aux[i].loop_header);
+			if (loop->irreducible)
+				bpf_log(log, ", irreducible");
+			if (loop->exits_overflow)
+				bpf_log(log, ", too many exits");
+			bpf_log(log, "\n");
+			for (j = 0; j < loop->backedges_cnt; j++)
+				bpf_log(log, "  backedge from %d, latch at %d\n",
+					loop->backedges[j].from, loop->backedges[j].latch);
+			for (j = 0; j < loop->exits_cnt; j++)
+				bpf_log(log, "  exit from %d to %d\n",
+					loop->exits[j].from, loop->exits[j].to);
+		}
+	}
+
+out:
+	kvfree(dfs.dfs_pos);
+	kvfree(dfs.stack);
+	kvfree(dfs.state);
+	return err;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 93100541475a..0d18b3bb0865 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22589,13 +22589,17 @@ static void log_program(struct bpf_verifier_env *env)
 	u64 pos, insn_pos;
 	u32 i, j;
 
-	verbose(env, "Program dump (scc? idom insn#: live_regs_before):\n");
+	verbose(env, "Program dump (scc? loop_header? idom insn#: live_regs_before):\n");
 	for (i = 0; i < insn_cnt; ++i) {
 		verbose_linfo(env, i, "    ; ");
 		if (env->insn_aux_data[i].scc)
 			verbose(env, "%3d ", env->insn_aux_data[i].scc);
 		else
 			verbose(env, "    ");
+		if (env->insn_aux_data[i].loop_header >= 0)
+			verbose(env, "%3d ", env->insn_aux_data[i].loop_header);
+		else
+			verbose(env, "    ");
 		verbose(env, "%3d ", env->idoms[i]);
 		verbose(env, "%3d: ", i);
 		for (j = BPF_REG_0; j < BPF_REG_10; ++j)
@@ -22820,6 +22824,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	ret = bpf_compute_loops(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = bpf_compute_live_registers(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -22972,6 +22980,8 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	release_btfs(env);
 err_free_env:
 	bpf_free_subprog_jts(env);
+	if (env->insn_aux_data)
+		bpf_clear_insn_aux_data(env, 0, env->insn_aux_data_len);
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
 	bpf_stack_liveness_free(env);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (18 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 19/36] bpf: compute loop hierarchy Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
                   ` (15 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	Eduard Zingerman

From: Eduard Zingerman <ezingerman@fb.com>

SCEV needs to process pending basic blocks in reverse postorder.
Provide a small integer min-heap whose comparator can order instruction
indices by their CFG ranks.

Grow the backing array on demand and pass caller context to the
comparator, so the heap need not encode the analysis-specific order.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 17 +++++++++
 kernel/bpf/Makefile          |  2 +-
 kernel/bpf/heap.c            | 87 ++++++++++++++++++++++++++++++++++++++++++++
 3 files changed, 105 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d8142f1b2347..f2497236542c 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1916,4 +1916,21 @@ int bpf_compute_loops(struct bpf_verifier_env *env);
 int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx);
 bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header);
 
+/*
+ * Simple binary heap implementation as described by
+ * https://en.wikipedia.org/wiki/Binary_heap
+ */
+struct bpf_min_heap {
+	int (*compare)(int, int, void *); /* ordering function for @elements */
+	int *elements; /* min-heap ordered by @compare */
+	void *arg; /* 3rd argument passed to @compare */
+	int capacity;
+	int count;
+};
+
+void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg);
+void bpf_min_heap_free(struct bpf_min_heap *heap);
+int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
+bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 6210601eb40b..7a2c179a9059 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/heap.c b/kernel/bpf/heap.c
new file mode 100644
index 000000000000..029f9217b872
--- /dev/null
+++ b/kernel/bpf/heap.c
@@ -0,0 +1,87 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_verifier.h>
+
+/* Indexes for binary tree encoded as an array */
+static inline int left_child(int i) { return 2 * i + 1; }
+static inline int right_child(int i) { return 2 * i + 2; }
+static inline int parent(int i) { return (i - 1) / 2; }
+
+static inline int greater(struct bpf_min_heap *heap, int a, int b)
+{
+	return heap->compare(a, b, heap->arg) > 0;
+}
+
+void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg)
+{
+	memset(heap, 0, sizeof(*heap));
+	heap->compare = compare;
+	heap->arg = arg;
+}
+
+void bpf_min_heap_free(struct bpf_min_heap *heap)
+{
+	kfree(heap->elements);
+	heap->elements = NULL;
+	heap->capacity = 0;
+	heap->count = 0;
+}
+
+int bpf_min_heap_push(struct bpf_min_heap *heap, int elt)
+{
+	int new_capacity, i;
+	int *elements;
+	void *tmp;
+
+	if (heap->count == heap->capacity) {
+		new_capacity = heap->capacity ? heap->capacity * 2 : 16;
+		tmp = krealloc(heap->elements,
+			       sizeof(*heap->elements) * new_capacity,
+			       GFP_KERNEL_ACCOUNT);
+		if (!tmp)
+			return -ENOMEM;
+		heap->elements = tmp;
+		heap->capacity = new_capacity;
+	}
+
+	elements = heap->elements;
+	i = heap->count;
+	elements[i] = elt;
+	heap->count++;
+	while (i != 0 && greater(heap, elements[parent(i)], elements[i])) {
+		swap(elements[i], elements[parent(i)]);
+		i = parent(i);
+	}
+	return 0;
+}
+
+static inline void sink_root(struct bpf_min_heap *heap)
+{
+	int *elements = heap->elements;
+	int i = 0;
+
+	while ((left_child(i)  < heap->count && greater(heap, elements[i], elements[left_child(i)])) ||
+	       (right_child(i) < heap->count && greater(heap, elements[i], elements[right_child(i)]))) {
+		if (right_child(i) >= heap->count || greater(heap, elements[right_child(i)], elements[left_child(i)])) {
+			swap(elements[i], elements[left_child(i)]);
+			i = left_child(i);
+		} else {
+			swap(elements[i], elements[right_child(i)]);
+			i = right_child(i);
+		}
+	}
+}
+
+bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt)
+{
+	if (heap->count == 0)
+		return false;
+
+	int *elements = heap->elements;
+	*elt = elements[0];
+	elements[0] = elements[heap->count - 1];
+	--heap->count;
+	sink_root(heap);
+	return true;
+}

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (19 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 22/36] bpf: add bpf_split_cur_state() Eduard Zingerman
                   ` (14 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	Eduard Zingerman

From: Eduard Zingerman <ezingerman@fb.com>

SCEV needs information about the program's basic-block structure.
Piggyback on compute_predecessors() and set
insn_aux_data->bb_end for an instruction if:
- it has multiple successors, or
- it has multiple predecessors, or
- it is a jump instruction.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/loops.c           | 19 +++++++++++++++++++
 2 files changed, 20 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index f2497236542c..728567bde055 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -729,6 +729,7 @@ struct bpf_insn_aux_data {
 	bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
 	bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
 	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
+	bool bb_end;
 	u8 alu_state; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
index e269870f6a6e..5f52a88d5283 100644
--- a/kernel/bpf/loops.c
+++ b/kernel/bpf/loops.c
@@ -4,8 +4,23 @@
 #include <linux/slab.h>
 #include <linux/bpf_verifier.h>
 
+static bool is_cfg_jump(struct bpf_insn *insn)
+{
+	u8 class = BPF_CLASS(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+
+	switch (class) {
+	case BPF_JMP:
+	case BPF_JMP32:
+		return opcode != BPF_CALL;
+	default:
+		return false;
+	}
+}
+
 static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
 {
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
 	struct bpf_iarray *succ, *preds, **result;
 	struct bpf_prog *prog = env->prog;
 	u32 *num_preds, i, s, sz, len = prog->len;
@@ -54,6 +69,10 @@ static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
 			preds = result[s];
 			preds->items[preds->cnt++] = i;
 		}
+		if (succ->cnt > 1 ||
+		    (succ->cnt == 1 && num_preds[succ->items[0]] > 1) ||
+		    is_cfg_jump(insn))
+			aux[i].bb_end = true;
 		if (bpf_is_ldimm64(insn))
 			i++;
 	}

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 22/36] bpf: add bpf_split_cur_state()
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (20 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
                   ` (13 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV widening needs an explicit checkpoint preserving the original
loop-entry state, independently of ordinary state-cache heuristics.
This commit extracts checkpoint creation logic from
bpf_is_state_visited() into a separate utility function.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/states.c          | 95 ++++++++++++++++++++++++--------------------
 2 files changed, 54 insertions(+), 42 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 728567bde055..76833c2ad452 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1329,6 +1329,7 @@ void bpf_free_kfunc_btf_tab(struct bpf_kfunc_btf_tab *tab);
 int mark_chain_precision(struct bpf_verifier_env *env, int regno);
 
 int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx);
+int bpf_split_cur_state(struct bpf_verifier_env *env);
 int bpf_update_branch_counts(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
 
 void bpf_clear_jmp_history(struct bpf_verifier_state *state);
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 3eaae1452c54..d7f9f879bd87 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -1268,11 +1268,61 @@ static void mark_all_scalars_imprecise(struct bpf_verifier_env *env, struct bpf_
 	}
 }
 
-int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
+int bpf_split_cur_state(struct bpf_verifier_env *env)
 {
+	struct bpf_verifier_state *cur = env->cur_state, *new;
 	struct bpf_verifier_state_list *new_sl;
+	struct list_head *head;
+	int insn_idx = cur->insn_idx;
+	int err;
+
+	head = bpf_explored_state(env, insn_idx);
+	new_sl = kzalloc_obj(struct bpf_verifier_state_list, GFP_KERNEL_ACCOUNT);
+	if (!new_sl)
+		return -ENOMEM;
+	env->total_states++;
+	env->explored_states_size++;
+	update_peak_states(env);
+	env->prev_jmps_processed = env->jmps_processed;
+	env->prev_insn_processed = env->insn_processed;
+
+	/* forget precise markings we inherited, see __mark_chain_precision */
+	if (env->bpf_capable)
+		mark_all_scalars_imprecise(env, cur);
+
+	bpf_clear_singular_ids(env, cur);
+
+	/* add new state to the head of linked list */
+	new = &new_sl->state;
+	err = bpf_copy_verifier_state(new, cur);
+	if (err) {
+		bpf_free_verifier_state(new, false);
+		kfree(new_sl);
+		return err;
+	}
+	new->insn_idx = insn_idx;
+	verifier_bug_if(new->branches != 1, env,
+			"%s:branches_to_explore=%d insn %d",
+			__func__, new->branches, insn_idx);
+	err = maybe_enter_scc(env, new);
+	if (err) {
+		bpf_free_verifier_state(new, false);
+		kfree(new_sl);
+		return err;
+	}
+
+	cur->parent = new;
+	cur->first_insn_idx = insn_idx;
+	cur->dfs_depth = new->dfs_depth + 1;
+	bpf_clear_jmp_history(cur);
+	list_add(&new_sl->node, head);
+	return 0;
+}
+
+int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
+{
 	struct bpf_verifier_state_list *sl;
-	struct bpf_verifier_state *cur = env->cur_state, *new;
+	struct bpf_verifier_state *cur = env->cur_state;
 	bool force_new_state, add_new_state, loop;
 	int n, err, states_cnt = 0;
 	struct list_head *pos, *tmp, *head;
@@ -1600,44 +1650,5 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 	 * When looping the sl->state.branches will be > 0 and this state
 	 * will not be considered for equivalence until branches == 0.
 	 */
-	new_sl = kzalloc_obj(struct bpf_verifier_state_list, GFP_KERNEL_ACCOUNT);
-	if (!new_sl)
-		return -ENOMEM;
-	env->total_states++;
-	env->explored_states_size++;
-	update_peak_states(env);
-	env->prev_jmps_processed = env->jmps_processed;
-	env->prev_insn_processed = env->insn_processed;
-
-	/* forget precise markings we inherited, see __mark_chain_precision */
-	if (env->bpf_capable)
-		mark_all_scalars_imprecise(env, cur);
-
-	bpf_clear_singular_ids(env, cur);
-
-	/* add new state to the head of linked list */
-	new = &new_sl->state;
-	err = bpf_copy_verifier_state(new, cur);
-	if (err) {
-		bpf_free_verifier_state(new, false);
-		kfree(new_sl);
-		return err;
-	}
-	new->insn_idx = insn_idx;
-	verifier_bug_if(new->branches != 1, env,
-			"%s:branches_to_explore=%d insn %d",
-			__func__, new->branches, insn_idx);
-	err = maybe_enter_scc(env, new);
-	if (err) {
-		bpf_free_verifier_state(new, false);
-		kfree(new_sl);
-		return err;
-	}
-
-	cur->parent = new;
-	cur->first_insn_idx = insn_idx;
-	cur->dfs_depth = new->dfs_depth + 1;
-	bpf_clear_jmp_history(cur);
-	list_add(&new_sl->node, head);
-	return 0;
+	return bpf_split_cur_state(env);
 }

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (21 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 22/36] bpf: add bpf_split_cur_state() Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:27   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops Eduard Zingerman
                   ` (12 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

SCEV integration logic might create two consecutive checkpoints at the
same instruction w/o processing any instructions in between.

This happens for loop-entry checkpoint -> regular checkpoint
sequences, where the loop-entry checkpoint has to remain unchanged
but the regular checkpoint is used for widening.

This commit adapts bpf_mark_chain_precision() to support this.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 4 +++-
 kernel/bpf/backtrack.c       | 5 +++++
 kernel/bpf/states.c          | 1 +
 3 files changed, 9 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 76833c2ad452..08e0b49841df 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -502,7 +502,9 @@ struct bpf_verifier_state {
 	bool speculative;
 	bool in_sleepable;
 
-	/* first and last insn idx of this verifier state */
+	/* First and last insn idx of this verifier state.
+	 * last_insn_idx is -1 if no instructions have been executed yet.
+	 */
 	u32 first_insn_idx;
 	u32 last_insn_idx;
 	/* if this state is a backedge state then equal_state
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..43d5740034c0 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -888,6 +888,10 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 		}
 
 		if (last_idx < 0) {
+			/* Consecutive checkpoints can have no instructions between them. */
+			if (st->parent)
+				goto parent;
+
 			/* we are at the entry into subprog, which
 			 * is expected for global funcs, but only if
 			 * requested precise registers are R1-R5
@@ -952,6 +956,7 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 				return -EFAULT;
 			}
 		}
+parent:
 		st = st->parent;
 		if (!st)
 			break;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index d7f9f879bd87..68df27e34b60 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -1312,6 +1312,7 @@ int bpf_split_cur_state(struct bpf_verifier_env *env)
 	}
 
 	cur->parent = new;
+	cur->last_insn_idx = -1;
 	cur->first_insn_idx = insn_idx;
 	cur->dfs_depth = new->dfs_depth + 1;
 	bpf_clear_jmp_history(cur);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (22 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:38   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops Eduard Zingerman
                   ` (11 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Assign algebraic expressions to loop variables (registers and stack
spills), describing how their values evolve across iterations. These
summaries provide symbolic input for the subsequent widening patch.

The algorithm draws on ideas from the following papers:

- "Symbolic Evaluation of Chains of Recurrences for Loop Optimization"
  Robert A. van Engelen, 2000
- "The CR# Algebra and its Application in Loop Analysis and Optimization"
  Robert A. van Engelen, 2004

The analysis proceeds in two phases:

- Compute expressions for one symbolic iteration: start each variable
  with a reference to its input value, apply instruction effects and
  join expressions where control-flow paths merge.
- Convert the resulting backedge updates to recurrences at the header,
  then substitute these recurrences into expressions at the latch.

Analyze loops innermost first. Within each loop, visit blocks in
topological order (reverse postorder, with backedges ignored and nested
loops collapsed). Treat each nested loop as an opaque operation:
preserve its invariant values, forget the others and continue at its
exits, without expanding its iterations.

For example, using 64-bit arithmetic:

    0: r7 = 5;
    1: r6 = 0;
    2: do {                // header
    3:     if (r6 == 1)
    4:         r7 = 10;
    5:     r6 += 1;
    6: } while (r6 < 3);   // latch

Expressions use symbolic inputs r6 and r7; the concrete values from
lines 0-1 are supplied later by the main verifier.

First phase:
------------

At each instruction, an environment maps register numbers to their
current expressions. Transfer functions update this mapping, while
joins merge the environments arriving along different paths.

For illustration, follow each branch separately to the backedge:

- At 3, take the path skipping 4: r6 = r6, r7 = r7.
- At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
- At 2, the first path contributes r6 = (+ r6 1), r7 = r7.
- At 4, on the other path from 3, transfer(r7 = 10):
  r7 = r7 -> r7 = 10.
- At 5, transfer(r6 += 1): r6 = r6 -> r6 = (+ r6 1).
- At 2, join the two paths' contributions:
  join((+ r6 1), (+ r6 1)) = (+ r6 1) for r6;
  join(r7, 10) = (any r7 10) for r7.

Equal expressions stay unchanged; different expressions become ANY
alternatives. These summarize one iteration, without unrolling the loop.

Second phase:
-------------

The update r6 = (+ r6 1) says that each iteration adds 1 to r6.
Repeating this update n times gives r6_entry + n, a linear recurrence.

- At header 2, r6 = (+ r6 1) becomes (linear r6 1), meaning
  r6_entry + n, where n counts iterations from zero.
- At latch 6, substitute this recurrence into (+ r6 1):
  (+ (linear r6 1) 1) simplifies to (linear (+ r6 1) 1),
  meaning r6_entry + 1 + n.
- At both points, r7 retains (any r7 10): either its entry value
  or 10. ANY records alternatives, not numeric ranges or the path
  conditions selecting them. Unchanged variables retain their inputs.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h |    7 +
 kernel/bpf/Makefile          |    2 +-
 kernel/bpf/scev.c            | 1519 ++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c        |    9 +
 4 files changed, 1536 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 08e0b49841df..54a971800ea3 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -732,6 +732,7 @@ struct bpf_insn_aux_data {
 	bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
 	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
 	bool bb_end;
+	bool need_scev;
 	u8 alu_state; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
@@ -991,6 +992,7 @@ struct bpf_scc_info {
 };
 
 struct bpf_liveness;
+struct scev;
 
 struct bpf_fd_array {
 	union {
@@ -1153,6 +1155,7 @@ struct bpf_verifier_env {
 	struct bpf_iarray *succ;
 	struct bpf_iarray *gotox_tmp_buf;
 	int *idoms;
+	struct scev *scev;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1937,4 +1940,8 @@ void bpf_min_heap_free(struct bpf_min_heap *heap);
 int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
 bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
 
+int bpf_init_scev(struct bpf_verifier_env *env);
+void bpf_free_scev(struct bpf_verifier_env *env);
+int bpf_compute_scev(struct bpf_verifier_env *env);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 7a2c179a9059..fdff31d962d6 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -6,7 +6,7 @@ cflags-nogcse-$(CONFIG_X86)$(CONFIG_CC_IS_GCC) := -fno-gcse
 endif
 CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)
 
-obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o
+obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o heap.o scev.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
new file mode 100644
index 000000000000..396b0c776dfa
--- /dev/null
+++ b/kernel/bpf/scev.c
@@ -0,0 +1,1519 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_verifier.h>
+#include <linux/jhash.h>
+#include <linux/bug.h>
+
+#define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+#define UNKNOWN_EXPR_ID 0
+#define OPAQUE_EXPR_ID  1
+
+/*
+ * BPF instructions use 'code', 'src_reg', 'off' and 'imm' fields for instruction encoding.
+ * For scalar evolution purpose we want to reuse most of the opcode definitions,
+ * but also add a few custom operations (REG and IMM).
+ * Use the enum below to uniformly represent all operations relevant for SCEV.
+ */
+enum expr_op {
+	/* leave range 0..255 for standard bpf opcodes */
+	UNKNOWN = 256, /* start custom opcodes from the second byte */
+	REG,
+	IMM,
+	SDIV, SMOD,
+	SEXT8, SEXT16, SEXT32,
+	ZEXT8, ZEXT16, ZEXT32,
+	BSWAP16, BSWAP32, BSWAP64,
+	/* during the loop body execution the value can be either of param[0] or param[1] */
+	ANY,
+	/* some value that verifier is not going to track precisely */
+	OPAQUE,
+	/* '*(u8/16/32 *)(r10 + X) = Y' writes define slot contents only partially */
+	SPILL8, SPILL16, SPILL32,
+	/*
+	 * SCEV expression corresponding to linear equation 'param[0] + param[1] * k',
+	 * where k is a loop iteration number. Loop here refers to innermost loop
+	 * containing instruction associated with this expression, as returned by
+	 * bpf_loop_at_index().
+	 */
+	LINEAR_SCEV,
+};
+
+struct expr {
+	u32 op;
+	union {
+		u32 params[2];
+		s64 imm;
+	};
+};
+
+struct expr_bucket {
+	u32 cnt;
+	u32 cap;
+	u32 ids[];
+};
+
+struct env {
+	bool empty;
+	u32 reg2expr[REGS_NUM];
+	u32 reg2scev[REGS_NUM];
+};
+
+struct insn_envs {
+	u32 cnt;
+	struct {
+		int loop_header;
+		struct env *env;
+	} entries[];
+};
+
+#define NUM_BUCKETS 256
+#define EXPR_STACK_DEPTH 8
+
+struct expr_stack_elt {
+	u32 id:28;
+	u32 pre:1;
+	u32 next_param:2;
+};
+
+struct scev {
+	/*
+	 * Expressions are identified by id, exprs_ht ensures that
+         * each expression exists as a unique instance.
+	 * This allows for fast equivalence check: id1 === id2.
+	 */
+	struct expr_bucket *exprs_ht[NUM_BUCKETS]; // Don't want to add struct hlist_node to expr
+	struct bpf_min_heap worklist;
+	struct insn_envs **envs;
+	struct expr *exprs;
+	/*
+	 * Loops are analyzed one by one, this array keeps track if a particular
+	 * instruction was visited on a current pass.
+	 */
+	u32 *discovered;
+	int exprs_cnt;
+	int exprs_cap;
+	int envs_cnt;
+	int stack_sz;
+	struct expr_stack_elt expr_stack[EXPR_STACK_DEPTH];
+	u32 ids_buf[EXPR_STACK_DEPTH];
+};
+
+static u32 expr_hash(struct expr *e)
+{
+	return jhash_3words(e->op, e->params[0], e->params[1], 0);
+}
+
+static int expr_eq(struct expr *a, struct expr *b)
+{
+	return a->op == b->op && a->imm == b->imm;
+}
+
+static int add_expr(struct scev *scev, struct expr e)
+{
+	struct expr_bucket *bucket;
+	u32 i, id, hash, new_cap;
+	void *tmp;
+
+	hash = expr_hash(&e) % NUM_BUCKETS;
+	bucket = scev->exprs_ht[hash];
+
+	if (bucket) {
+		for (i = 0; i < bucket->cnt; i++) {
+			id = bucket->ids[i];
+			if (expr_eq(&e, &scev->exprs[id]))
+				return id;
+		}
+	}
+
+	if (!bucket || bucket->cap == bucket->cnt) {
+		new_cap = bucket ? bucket->cap * 2 : 32;
+		bucket = kvrealloc(bucket, sizeof(*bucket) + sizeof(u32) * new_cap, GFP_KERNEL_ACCOUNT | __GFP_ZERO);
+		if (!bucket)
+			return -ENOMEM;
+		scev->exprs_ht[hash] = bucket;
+		bucket->cap = new_cap;
+	}
+
+	if (scev->exprs_cnt == scev->exprs_cap) {
+		new_cap = scev->exprs_cap + 256;
+		tmp = kvrealloc(scev->exprs, sizeof(struct expr) * new_cap, GFP_KERNEL_ACCOUNT);
+		if (!tmp)
+			return -ENOMEM;
+		scev->exprs = tmp;
+		scev->exprs_cap = new_cap;
+	}
+
+	id = scev->exprs_cnt++;
+	scev->exprs[id] = e;
+	bucket->ids[bucket->cnt++] = id;
+	return id;
+}
+
+static int expr2(struct scev *scev, u32 op, int a, int b)
+{
+	if (a < 0)
+		return a;
+	if (b < 0)
+		return b;
+	return add_expr(scev, (struct expr){ .op = op, .params = {a, b} });
+}
+
+static int expr1(struct scev *scev, u32 op, int a)
+{
+	if (a < 0)
+		return a;
+	return expr2(scev, op, a, 0);
+}
+
+static int expr0(struct scev *scev, u32 op)
+{
+	return expr2(scev, op, 0, 0);
+}
+
+static bool is_expr1(struct expr *expr, u32 op, int a)
+{
+	return expr->op == op && expr->params[0] == a;
+}
+
+static int imm_expr(struct scev *scev, s64 value)
+{
+	return add_expr(scev, (struct expr){ .op = IMM, .imm = value });
+}
+
+static bool same_exprs(struct scev *scev, int id_a, int id_b)
+{
+	return id_a == id_b;
+}
+
+static bool is_imm(struct scev *scev, int id, s64 *imm)
+{
+	if (scev->exprs[id].op != IMM)
+		return false;
+	*imm = scev->exprs[id].imm;
+	return true;
+}
+
+static bool is_op(struct scev *scev, int id, u32 op)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	return true;
+}
+
+static bool is_unop(struct scev *scev, enum expr_op op, int id, u32 *p0)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	*p0 = scev->exprs[id].params[0];
+	return true;
+}
+
+static bool is_binop(struct scev *scev, enum expr_op op, int id, u32 *p0, u32 *p1)
+{
+	if (scev->exprs[id].op != op)
+		return false;
+	*p0 = scev->exprs[id].params[0];
+	*p1 = scev->exprs[id].params[1];
+	return true;
+}
+
+static bool is_reg(struct scev *scev, int id, u32 *reg)
+{
+	return is_unop(scev, REG, id, reg);
+}
+
+static bool is_add(struct scev *scev, int id, u32 *left, u32 *right)
+{
+	return is_binop(scev, BPF_ADD, id, left, right);
+}
+
+static bool is_zext32(struct scev *scev, int id, u32 *left)
+{
+	return is_unop(scev, ZEXT32, id, left);
+}
+
+static bool is_any(struct scev *scev, int id, u32 *left, u32 *right)
+{
+	return is_binop(scev, ANY, id, left, right);
+}
+
+static bool is_linear(struct scev *scev, int id, u32 *base, u32 *slope)
+{
+	return is_binop(scev, LINEAR_SCEV, id, base, slope);
+}
+
+static bool is_opaque(struct scev *scev, int id)
+{
+	return is_op(scev, id, OPAQUE);
+}
+
+static void log_reg(struct bpf_verifier_env *env, u32 reg)
+{
+	if (reg < MAX_BPF_REG)
+		bpf_log(&env->log, "r%d", reg);
+	else
+		bpf_log(&env->log, "*fp%d", (MAX_BPF_REG - reg - 1) * 8);
+}
+
+static const char *op_str(u32 op)
+{
+	switch (op) {
+	case BPF_ADD:  return "+";
+	case BPF_SUB:  return "-";
+	case BPF_MUL:  return "*";
+	case BPF_DIV:  return "/";
+	case SDIV:     return "s/";
+	case BPF_OR:   return "|";
+	case BPF_AND:  return "&";
+	case BPF_LSH:  return "<<";
+	case BPF_RSH:  return ">>";
+	case BPF_NEG:  return "-";
+	case BPF_MOD:  return "%";
+	case SMOD:     return "s%";
+	case BPF_XOR:  return "^";
+	case BPF_ARSH: return "s>>";
+	case SEXT8:    return "sext8";
+	case SEXT16:   return "sext16";
+	case SEXT32:   return "sext32";
+	case ZEXT8:    return "zext8";
+	case ZEXT16:   return "zext16";
+	case ZEXT32:   return "zext32";
+	case BSWAP16:  return "bswap16";
+	case BSWAP32:  return "bswap32";
+	case BSWAP64:  return "bswap64";
+	case SPILL8:   return "spill8";
+	case SPILL16:  return "spill16";
+	case SPILL32:  return "spill32";
+	case ANY:      return "any";
+	case LINEAR_SCEV:  return "linear";
+	}
+	return NULL;
+}
+
+static u32 op_params_num(u32 op)
+{
+	switch (op) {
+	case BPF_ADD:
+	case BPF_SUB:
+	case BPF_MUL:
+	case BPF_DIV:
+	case BPF_MOD:
+	case BPF_OR:
+	case BPF_XOR:
+	case BPF_AND:
+	case BPF_LSH:
+	case BPF_RSH:
+	case BPF_ARSH:
+	case SDIV:
+	case SMOD:
+	case ANY:
+	case LINEAR_SCEV:
+		return 2;
+	case BPF_NEG:
+	case SEXT8:
+	case SEXT16:
+	case SEXT32:
+	case ZEXT8:
+	case ZEXT16:
+	case ZEXT32:
+	case BSWAP16:
+	case BSWAP32:
+	case BSWAP64:
+	case SPILL8:
+	case SPILL16:
+	case SPILL32:
+		return 1;
+	case REG:
+	case IMM:
+	case OPAQUE:
+		return 0;
+	default:
+		return 0;
+	}
+}
+
+static bool expr_stack_push(struct scev *scev, u32 id)
+{
+	if (scev->stack_sz >= EXPR_STACK_DEPTH)
+		return false;
+	scev->expr_stack[scev->stack_sz].id = id;
+	scev->expr_stack[scev->stack_sz].pre = true;
+	scev->expr_stack[scev->stack_sz].next_param = 0;
+	scev->stack_sz++;
+	return true;
+}
+
+enum {
+	PRE = BIT(1), POST = BIT(2), DEPTH_LIMIT = BIT(3)
+};
+
+static bool expr_next(struct scev *scev, u32 *id, u32 *order)
+{
+	struct expr_stack_elt *elt;
+	struct expr *expr;
+	u32 num_params;
+
+	if (scev->stack_sz == 0)
+		return false;
+
+	elt = &scev->expr_stack[scev->stack_sz - 1];
+	*id = elt->id;
+	*order = 0;
+	expr = &scev->exprs[elt->id];
+	num_params = op_params_num(expr->op);
+	if (elt->pre) {
+		elt->pre = false;
+		*order = PRE;
+		return true;
+	}
+	if (elt->next_param == num_params) {
+		*order = POST;
+		scev->stack_sz--;
+		return true;
+	}
+	if (scev->stack_sz == EXPR_STACK_DEPTH) {
+		*order = POST | DEPTH_LIMIT;
+		scev->stack_sz--;
+		return true;
+	}
+	expr_stack_push(scev, expr->params[elt->next_param]);
+	elt->next_param++;
+	return expr_next(scev, id, order);
+}
+
+static void log_expr(struct bpf_verifier_env *env, u32 id)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct expr *expr;
+	const char *str;
+	u32 order;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if ((order & PRE) && scev->stack_sz > 1)
+			bpf_log(log, " ");
+		expr = &scev->exprs[id];
+		switch (expr->op) {
+		case UNKNOWN:
+			if (order & PRE)
+				bpf_log(log, "?");
+			break;
+		case OPAQUE:
+			if (order & PRE)
+				bpf_log(log, "_");
+			break;
+		case REG:
+			if (order & PRE)
+				log_reg(env, expr->params[0]);
+			break;
+		case IMM:
+			if (order & PRE)
+				bpf_log(log, "%lld", expr->imm);
+			break;
+		default:
+			if (order & PRE) {
+				str = op_str(expr->op);
+				bpf_log(log, "(");
+				if (str)
+					bpf_log(log, "%s", str);
+				else
+					bpf_log(log, "bad-expr-op %x", expr->op);
+			}
+			if (order & DEPTH_LIMIT)
+				bpf_log(log, "...");
+			if (order & POST)
+				bpf_log(log, ")");
+		}
+	}
+}
+
+enum print_env_flags {
+	PRINT_SCEV = BIT(1)
+};
+
+static bool reg_alive_at(struct bpf_verifier_env *env, u32 reg, u32 insn_idx)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	u16 live_regs = aux[insn_idx].live_regs_before;
+
+	return reg < MAX_BPF_REG
+	       ? (BIT(reg) & live_regs)
+	       : test_bit(reg - MAX_BPF_REG, aux[insn_idx].live_stack_before);
+}
+
+static void print_env(struct bpf_verifier_env *env, struct env *e, u32 insn_idx, u32 flags)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	bool printed_some = false;
+	bool print_scev = flags & PRINT_SCEV;
+	bool is_self_scev;
+	bool is_self_reg;
+	int i, r;
+
+	if (e->empty) {
+		bpf_log(log, "  <empty>\n");
+		return;
+	}
+
+	for (i = 0; i < REGS_NUM; i++) {
+		if (!reg_alive_at(env, i, insn_idx))
+			continue;
+		is_self_reg = is_reg(scev, e->reg2expr[i], &r) && i == r;
+		is_self_scev = is_reg(scev, e->reg2scev[i], &r) && i == r;
+		if (is_self_reg && (!print_scev || is_self_scev))
+			continue;
+		printed_some = true;
+		bpf_log(log, "  ");
+		log_reg(env, i);
+		bpf_log(log, "=");
+		log_expr(env, e->reg2expr[i]);
+		if (print_scev && e->reg2expr[i] != UNKNOWN_EXPR_ID) {
+			bpf_log(log, " / ");
+			log_expr(env, e->reg2scev[i]);
+		}
+		bpf_log(log, "\n");
+	}
+	if (!printed_some)
+		bpf_log(log, "  <all regs unchanged>\n");
+}
+static struct env *find_loop_env(struct scev *scev, int loop_header, int insn_idx)
+{
+	struct insn_envs *envs = scev->envs[insn_idx];
+	u32 i;
+
+	if (!envs)
+		return NULL;
+	for (i = 0; i < envs->cnt; i++)
+		if (envs->entries[i].loop_header == loop_header)
+			return envs->entries[i].env;
+	return NULL;
+}
+
+static struct env *find_header_env(struct scev *scev, int header)
+{
+	return find_loop_env(scev, header, header);
+}
+
+static struct env *get_loop_env(struct scev *scev, int loop_header, int insn_idx)
+{
+	struct insn_envs *envs = scev->envs[insn_idx];
+	u32 cnt = envs ? envs->cnt : 0;
+	struct insn_envs *tmp;
+	struct env *e;
+
+	e = find_loop_env(scev, loop_header, insn_idx);
+	if (e)
+		return e;
+
+	e = kzalloc(sizeof(*e), GFP_KERNEL_ACCOUNT);
+	if (!e)
+		return NULL;
+
+	tmp = krealloc(envs, struct_size(envs, entries, cnt + 1), GFP_KERNEL_ACCOUNT);
+	if (!tmp) {
+		kfree(e);
+		return NULL;
+	}
+
+	e->empty = true;
+	tmp->entries[cnt].loop_header = loop_header;
+	tmp->entries[cnt].env = e;
+	tmp->cnt = cnt + 1;
+	scev->envs[insn_idx] = tmp;
+	return e;
+}
+
+static void setup_initial_loop_env(struct bpf_verifier_env *env, struct env *e, int insn_idx)
+{
+	struct scev *scev = env->scev;
+	int i;
+
+	for (i = 0; i < REGS_NUM; i++)
+		e->reg2expr[i] = expr1(scev, REG, i);
+}
+
+static int replace_reg(struct scev *scev, struct env *e, u32 reg, int id)
+{
+	if (id < 0)
+		return id;
+	e->reg2expr[reg] = id;
+	return 0;
+}
+
+static void forget_call_regs(struct env *e)
+{
+	int i;
+
+	for (i = BPF_REG_0; i <= BPF_REG_5; i++)
+		e->reg2expr[i] = UNKNOWN_EXPR_ID;
+}
+
+static u32 spill_spi(struct bpf_insn *insn)
+{
+	return -insn->off / BPF_REG_SIZE - 1;
+}
+
+static u32 off_to_reg(struct bpf_insn *insn)
+{
+	return spill_spi(insn) + MAX_BPF_REG;
+}
+
+static bool is_spill_off(int off)
+{
+	return off % BPF_REG_SIZE == 0 &&
+	       off <= -BPF_REG_SIZE &&
+	       off >= -MAX_BPF_STACK_JIT;
+}
+
+static void mark_opaque(struct env *e, int off, int size)
+{
+	int b, spi;
+
+	for (b = off; b < off + size; b++) {
+		if (b >= 0 || b < -MAX_BPF_STACK_JIT)
+			continue;
+		spi = (-b - 1) / BPF_REG_SIZE;
+		e->reg2expr[MAX_BPF_REG + spi] = OPAQUE_EXPR_ID;
+	}
+}
+
+static int mk_spill(struct scev *scev, u8 size, int id)
+{
+	switch (size) {
+	case BPF_B:  return expr1(scev, SPILL8,  id);
+	case BPF_H:  return expr1(scev, SPILL16, id);
+	case BPF_W:  return expr1(scev, SPILL32, id);
+	case BPF_DW: return id;
+	}
+	return UNKNOWN_EXPR_ID;
+}
+
+static int mk_fill(struct scev *scev, int id, u8 code)
+{
+	bool sx = BPF_MODE(code) == BPF_MEMSX;
+
+	switch (BPF_SIZE(code)) {
+	case BPF_B:  return expr1(scev, sx ? SEXT8  : ZEXT8,  id);
+	case BPF_H:  return expr1(scev, sx ? SEXT16 : ZEXT16, id);
+	case BPF_W:  return expr1(scev, sx ? SEXT32 : ZEXT32, id);
+	case BPF_DW: return sx ? UNKNOWN_EXPR_ID : id; /* DW sign-extended load is invalid */
+	}
+	return UNKNOWN_EXPR_ID;
+}
+
+static int maybe_store_fp(struct scev *scev, struct env *e, struct bpf_insn *insn, int id)
+{
+	u8 size = BPF_SIZE(insn->code);
+
+	if (insn->dst_reg != BPF_REG_FP)
+		return 0;
+	if (is_spill_off(insn->off))
+		return replace_reg(scev, e, off_to_reg(insn), mk_spill(scev, size, id));
+	mark_opaque(e, insn->off, bpf_size_to_bytes(size));
+	return 0;
+}
+
+static int maybe_load_fp(struct scev *scev, struct env *e, struct bpf_insn *insn)
+{
+	int id;
+
+	if (insn->src_reg == BPF_REG_FP && is_spill_off(insn->off))
+		id = mk_fill(scev, e->reg2expr[off_to_reg(insn)], insn->code);
+	else
+		id = OPAQUE_EXPR_ID;
+	return replace_reg(scev, e, insn->dst_reg, id);
+}
+
+static int transfer(struct bpf_verifier_env *env, struct env *e, int idx)
+{
+	const bool little_endian = htons(0x3412) == 0x1234;
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct scev *scev = env->scev;
+	u32 *reg2expr = e->reg2expr;
+	u8 class = BPF_CLASS(insn->code);
+	u8 x_or_k = BPF_SRC(insn->code);
+	u8 opcode = BPF_OP(insn->code);
+	u8 mode = BPF_MODE(insn->code);
+	u32 dst = insn->dst_reg;
+	u32 src = insn->src_reg;
+	u32 op, sext;
+	int i, id;
+	s64 imm;
+
+	switch (class) {
+	case BPF_ALU:
+	case BPF_ALU64:
+		switch (opcode) {
+		case BPF_MOV:
+			switch (insn->off) {
+			case 0: sext = 0; break;
+			case 8: sext = SEXT8; break;
+			case 16: sext = SEXT16; break;
+			case 32: sext = SEXT32; break;
+			default:
+				goto mark_dst_unknown;
+			}
+
+			if (x_or_k == BPF_X && insn->imm == 0)
+				id = reg2expr[src];
+			else if (x_or_k == BPF_K && src == 0 && insn->off == 0)
+				id = imm_expr(scev, insn->imm);
+			else
+				goto mark_dst_unknown;
+
+			if (sext)
+				id = expr1(scev, sext, id);
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_ADD:
+		case BPF_SUB:
+		case BPF_MUL:
+		case BPF_DIV:
+		case BPF_MOD:
+		case BPF_OR:
+		case BPF_XOR:
+		case BPF_AND:
+		case BPF_LSH:
+		case BPF_RSH:
+		case BPF_ARSH:
+			if (opcode == BPF_DIV && insn->off == 1)
+				op = SDIV;
+			else if (opcode == BPF_MOD && insn->off == 1)
+				op = SMOD;
+			else if (insn->off == 0)
+				op = opcode;
+			else
+				goto mark_dst_unknown;
+
+			if (x_or_k == BPF_X && insn->imm == 0)
+				id = reg2expr[src];
+			else if (x_or_k == BPF_K && src == 0)
+				id = imm_expr(scev, insn->imm);
+			else
+				goto mark_dst_unknown;
+
+			id = expr2(scev, op, reg2expr[dst], id);
+
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_NEG:
+			if (src == 0 && insn->off == 0 && insn->imm == 0)
+				id = expr1(scev, BPF_NEG, reg2expr[dst]);
+			else
+				goto mark_dst_unknown;
+
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+
+		case BPF_END:
+			switch (insn->imm) {
+			case 16: op = BSWAP16; break;
+			case 32: op = BSWAP32; break;
+			case 64: op = BSWAP64; break;
+			default:
+				goto mark_dst_unknown;
+			}
+
+			if (class == BPF_ALU && x_or_k == BPF_TO_LE && insn->off == 0 && little_endian)
+				op = 0; /* little-endian to little-endian is noop */
+			else if (class == BPF_ALU && x_or_k == BPF_TO_BE && insn->off == 0 && !little_endian)
+				op = 0; /* big-endian to big-endian is noop */
+			else if (class == BPF_ALU64 && x_or_k == 0 && insn->off == 0)
+				/* always swap */;
+			else
+				goto mark_dst_unknown;
+
+			id = reg2expr[dst];
+			if (op)
+				id = expr1(scev, op, reg2expr[dst]);
+
+			if (class == BPF_ALU)
+				id = expr1(scev, ZEXT32, id);
+
+			return replace_reg(scev, e, dst, id);
+		default:
+			goto mark_dst_unknown;
+		}
+		break;
+	case BPF_LDX:
+		switch (mode) {
+		case BPF_MEM:
+		case BPF_MEMSX:
+			return maybe_load_fp(scev, e, insn);
+		default:
+			goto mark_dst_unknown;
+		}
+	case BPF_STX:
+		switch (mode) {
+		case BPF_MEM:
+			return maybe_store_fp(scev, e, insn, reg2expr[src]);
+		case BPF_ATOMIC:
+			if (insn->imm == BPF_LOAD_ACQ)
+				return maybe_load_fp(scev, e, insn);
+			if (insn->imm == BPF_STORE_REL) {
+				return maybe_store_fp(scev, e, insn, reg2expr[src]);
+			}
+			/*
+			 * verifier does not track other atomic ops precisely,
+			 * hence mark the results as opaque.
+			 */
+			if (insn->dst_reg == BPF_REG_FP)
+				mark_opaque(e, insn->off, bpf_size_to_bytes(BPF_SIZE(insn->code)));
+			if (insn->imm == BPF_CMPXCHG)
+				return replace_reg(scev, e, BPF_REG_0, OPAQUE_EXPR_ID);
+			if (insn->imm & BPF_FETCH)
+				return replace_reg(scev, e, src, OPAQUE_EXPR_ID);
+			break;
+		}
+		break;
+	case BPF_ST:
+		if (insn->dst_reg == BPF_REG_FP) {
+			id = imm_expr(scev, insn->imm);
+			if (id < 0)
+				return id;
+			return maybe_store_fp(scev, e, insn, id);
+		}
+		break;
+	case BPF_JMP:
+	case BPF_JMP32:
+		if (opcode == BPF_CALL)
+			forget_call_regs(e);
+		/* for non-CALL there are no changes in register states */
+		break;
+	case BPF_LD:
+		switch (mode) {
+		case BPF_IMM:
+			/* rX = imm ll */
+			if (BPF_SIZE(insn->code) == BPF_DW && insn->src_reg == 0) {
+				imm = ((u64)(insn + 1)->imm << 32) | (u32)insn->imm;
+				return replace_reg(scev, e, dst, imm_expr(scev, imm));
+			}
+			/* map, map value, BTF id, function */
+			if (BPF_SIZE(insn->code) == BPF_DW)
+				return replace_reg(scev, e, dst, OPAQUE_EXPR_ID);
+			goto mark_dst_unknown;
+		case BPF_ABS:
+		case BPF_IND:
+			forget_call_regs(e);
+			break;
+		default:
+			goto mark_dst_unknown;
+		}
+		break;
+	default:
+		/* unknown instruction, nuke state */
+		for (i = 0; i < REGS_NUM; i++)
+			reg2expr[i] = UNKNOWN_EXPR_ID;
+		break;
+	}
+	return 0;
+
+mark_dst_unknown:
+	reg2expr[dst] = UNKNOWN_EXPR_ID;
+	return 0;
+}
+
+/*
+ * Construct minimal 'ANY' expression by traversing 'a' and 'b'
+ * and accumulating non-duplicated non-ANY entries.
+ */
+static int mk_any(struct scev *scev, u32 a, u32 b)
+{
+	struct expr_stack_elt *elt;
+	struct expr *expr;
+	u32 i, j, ids_buf_sz;
+	u32 roots[2] = {a, b};
+	int id;
+
+	ids_buf_sz = 0;
+	for (i = 0; i < ARRAY_SIZE(roots); i++) {
+		scev->stack_sz = 0;
+		expr_stack_push(scev, roots[i]);
+		while (scev->stack_sz) {
+			elt = &scev->expr_stack[--scev->stack_sz];
+			id = elt->id;
+			expr = &scev->exprs[elt->id];
+			if (expr->op == ANY) {
+				if (!expr_stack_push(scev, expr->params[0]) ||
+				    !expr_stack_push(scev, expr->params[1]))
+					return UNKNOWN_EXPR_ID;
+			} else {
+				for (j = 0; j < ids_buf_sz; j++) {
+					if (same_exprs(scev, scev->ids_buf[j], id))
+						goto next;
+				}
+				if (ids_buf_sz == ARRAY_SIZE(scev->ids_buf))
+					return UNKNOWN_EXPR_ID;
+				scev->ids_buf[ids_buf_sz++] = id;
+			}
+next:;
+		}
+	}
+	if (WARN_ON(ids_buf_sz == 0))
+		return -EFAULT;
+	id = scev->ids_buf[0];
+	for (i = 1; i < ids_buf_sz; i++) {
+		id = expr2(scev, ANY, id, scev->ids_buf[i]);
+		if (id < 0)
+			return id;
+	}
+	return id;
+}
+
+static int join(struct scev *scev, struct env *acc, struct env *cur)
+{
+	int i, id;
+
+	if (acc->empty) {
+		memcpy(acc, cur, sizeof(*acc));
+		acc->empty = false;
+		return 0;
+	}
+
+	for (i = 0; i < REGS_NUM; i++) {
+		if (!same_exprs(scev, acc->reg2expr[i], cur->reg2expr[i])) {
+			id = mk_any(scev, acc->reg2expr[i], cur->reg2expr[i]);
+			if (id < 0)
+				return id;
+			acc->reg2expr[i] = id;
+		}
+	}
+	return 0;
+}
+
+/* Mark any register modified in 'header_env' as unknown in 'acc'. */
+static void forget_non_invariants(struct scev *scev, struct env *acc, struct env *header_env)
+{
+	struct expr *header_expr;
+	int i;
+
+	if (acc->empty)
+		acc->empty = false;
+
+	for (i = 0; i < REGS_NUM; i++) {
+		header_expr = &scev->exprs[header_env->reg2expr[i]];
+		if (is_expr1(header_expr, REG, i))
+			continue;
+		acc->reg2expr[i] = UNKNOWN_EXPR_ID;
+	}
+}
+
+static int worklist_push(struct scev *scev, int loop_header, int idx)
+{
+	u32 tag = (u32)loop_header + 1;
+	int err;
+
+	if (scev->discovered[idx] == tag)
+		return 0;
+
+	err = bpf_min_heap_push(&scev->worklist, idx);
+	if (err)
+		return err;
+
+	scev->discovered[idx] = tag;
+	return 0;
+}
+
+static bool reg_alive_at_succ(struct bpf_verifier_env *env, u32 r, u32 insn_idx)
+{
+	struct bpf_iarray *succ = bpf_insn_successors(env, insn_idx);
+	u32 i;
+
+	for (i = 0; i < succ->cnt; i++)
+		if (reg_alive_at(env, r, succ->items[i]))
+			return true;
+	return false;
+}
+
+enum changes_at {
+	LOG_AT_TRANSFER,
+	LOG_AT_JOIN,
+};
+
+static void log_env_changes(struct bpf_verifier_env *env, enum changes_at at,
+			    struct env *old, struct env *new, int insn_idx)
+{
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	u64 null_pos = log->end_pos;
+	bool any_changes = false;
+	u32 old_id, new_id;
+	u64 len;
+	int r;
+
+	bpf_log(log, "%s %4d: ", at == LOG_AT_TRANSFER ? "t" : "j", insn_idx);
+	bpf_verbose_insn(env, &env->prog->insnsi[insn_idx]);
+	len = log->end_pos - null_pos;
+	bpf_log(log, "%*s", max(37 - (int)len, 1), " ");
+	bpf_log(log, " ; ");
+	for (r = 0; r < REGS_NUM; r++) {
+		old_id = old->reg2expr[r];
+		new_id = new->reg2expr[r];
+		if (at == LOG_AT_TRANSFER && !reg_alive_at_succ(env, r, insn_idx))
+			continue;
+		if (at == LOG_AT_JOIN && !reg_alive_at(env, r, insn_idx))
+			continue;
+		if (same_exprs(scev, old_id, new_id))
+			continue;
+		if (any_changes)
+			bpf_log(log, ", ");
+		log_reg(env, r);
+		bpf_log(log, " ");
+		log_expr(env, old_id);
+		bpf_log(log, " -> ");
+		log_expr(env, new_id);
+		any_changes = true;
+	}
+	bpf_log(log, "\n");
+	if (!any_changes)
+		bpf_vlog_reset(log, null_pos);
+}
+
+static bool is_probe_read_helper(u32 func_id)
+{
+	return func_id == BPF_FUNC_probe_read ||
+	       func_id == BPF_FUNC_probe_read_kernel ||
+	       func_id == BPF_FUNC_probe_read_user ||
+	       func_id == BPF_FUNC_probe_read_str ||
+	       func_id == BPF_FUNC_probe_read_kernel_str ||
+	       func_id == BPF_FUNC_probe_read_user_str;
+}
+
+/*
+ * If instruction is an indirect write to stack, invalidate SCEVs for spi's
+ * that this instruction can touch.
+ */
+static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct env *cur_env, int idx)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct scev *scev = env->scev;
+	u8 class = BPF_CLASS(insn->code);
+	const unsigned long *mask;
+	u32 spi, reg;
+	bool opaque;
+
+	/* Direct fp stores are fine. */
+	if ((class == BPF_ST || class == BPF_STX) && insn->dst_reg == BPF_REG_FP)
+		return;
+
+	mask = bpf_may_write_mask(env, idx);
+	opaque = bpf_helper_call(insn) && is_probe_read_helper(insn->imm);
+	for_each_set_bit(spi, mask, MAX_BPF_STACK_SLOTS) {
+		reg = MAX_BPF_REG + spi;
+		if (cur_env->reg2expr[reg] == UNKNOWN_EXPR_ID)
+			continue;
+		replace_reg(scev, cur_env, reg, opaque ? OPAQUE_EXPR_ID : UNKNOWN_EXPR_ID);
+	}
+}
+
+/* Find the topmost loop header containing idx inside cur_header, or -1 if none. */
+static int topmost_nested_loop(struct bpf_verifier_env *env, int idx, int cur_header)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int header;
+
+	for (header = bpf_loop_at_index(env, idx); header >= 0;
+	     header = aux[header].loop_header)
+		if (aux[header].loop_header == cur_header)
+			return header;
+
+	return -1;
+}
+
+/*
+ * Join cur_env into the successor's environment and schedule their traversal.
+ * Map successors in nested loops to their topmost nested header.
+ * Ignore successors outside the current loop and its nested loops.
+ */
+static int join_successor(struct bpf_verifier_env *env, int cur_header, int succ_idx,
+			  struct env *old_env, struct env *cur_env)
+{
+	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
+	struct scev *scev = env->scev;
+	int err, nested_header;
+	struct env *succ_env;
+
+	/*
+	 * There are several possibilities for a successor:
+	 * - succ_idx can be a part of a loop outside of the cur_header's loop,
+	 *   such edges are ignored.
+	 * - succ_idx can be a part of the same loop as cur_header,
+	 *   for such edges succ_idx environment is updated:
+	 *     e[succ_idx] = join(e[succ_idx], cur_env)
+	 * - succ_idx can be a part of some loop inner to cur_header,
+	 *   in such a case there exists some loop header H,
+	 *   such that H.loop_header == cur_header
+	 *   and H is the same as succ_idx's loop or contains it.
+	 */
+	if (bpf_loop_at_index(env, succ_idx) != cur_header) {
+		/*
+		 * topmost_nested_loop() either finds H or returns -1,
+		 * in case if succ_idx is a part of a loop outer to cur_header.
+		 */
+		nested_header = topmost_nested_loop(env, succ_idx, cur_header);
+		if (nested_header < 0)
+			return 0;
+		succ_idx = nested_header;
+	}
+	succ_env = get_loop_env(scev, cur_header, succ_idx);
+	if (!succ_env)
+		return -ENOMEM;
+	if (log_level2)
+		memcpy(old_env, succ_env, sizeof(*old_env));
+	err = join(scev, succ_env, cur_env);
+	if (err)
+		return err;
+	err = worklist_push(scev, cur_header, succ_idx);
+	if (err)
+		return err;
+	if (log_level2)
+		log_env_changes(env, LOG_AT_JOIN, old_env, succ_env, succ_idx);
+	return 0;
+}
+
+static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
+{
+	struct bpf_min_heap *worklist = &env->scev->worklist;
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *cur_loop = aux[cur_header].loop;
+	struct env *header_env, *nested_header_env;
+	struct scev *scev = env->scev;
+	struct bpf_loop *nested_loop;
+	struct env *cur_env = NULL;
+	struct env *old_env = NULL;
+	struct bpf_iarray *succ;
+	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
+	int s, i, err, idx, succ_idx;
+
+	if (log_level2)
+		bpf_log(&env->log, "Computing SCEV for loop at %d:\n", cur_header);
+
+	/*
+	 * For irreducible loops, and for loops nesting a loop with a truncated
+	 * exit list, just assume that everything is clobbered for now.
+	 */
+	if (cur_loop->irreducible || cur_loop->exits_overflow) {
+		header_env = get_loop_env(scev, cur_header, cur_header);
+		if (!header_env)
+			return -ENOMEM;
+		/* The freshly allocated environment has all expressions unknown. */
+		header_env->empty = false;
+		return 0;
+	}
+
+	cur_env = kzalloc(sizeof(*cur_env), GFP_KERNEL_ACCOUNT);
+	if (!cur_env)
+		goto nomem;
+	if (log_level2) {
+		old_env = kzalloc(sizeof(*old_env), GFP_KERNEL_ACCOUNT);
+		if (!old_env)
+			goto nomem;
+	}
+	header_env = get_loop_env(scev, cur_header, cur_header);
+	if (!header_env)
+		goto nomem;
+	setup_initial_loop_env(env, header_env, cur_header);
+	err = worklist_push(scev, cur_header, cur_header);
+	if (err)
+		goto out;
+
+	for (;;) {
+		if (!bpf_min_heap_pop(worklist, &idx))
+			break;
+
+		/* join_successor() maps nested-loop entries to their representative header. */
+		nested_loop = idx != cur_header ? aux[idx].loop : NULL;
+		memcpy(cur_env, find_loop_env(scev, cur_header, idx), sizeof(*cur_env));
+		if (nested_loop) {
+			/*
+			 * Process nested loop as a single instruction by
+			 * forgetting anything non-invariant in the nested loop
+			 */
+			if (log_level2)
+				memcpy(old_env, cur_env, sizeof(*old_env));
+			nested_header_env = find_header_env(scev, idx);
+			forget_non_invariants(scev, cur_env, nested_header_env);
+			if (log_level2)
+				log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
+			/*
+			 * Treat nested loop exits as successors,
+			 * join cur_env into successor's envs.
+			 */
+			for (i = 0; i < nested_loop->exits_cnt; i++) {
+				s = nested_loop->exits[i].to;
+				err = join_successor(env, cur_header, s, old_env, cur_env);
+				if (err)
+					goto out;
+			}
+		} else {
+			/*
+			 * Iterate instructions within a single basic block
+			 * starting at 'idx' mutating 'cur_env'.
+			 */
+			for (;;) {
+				if (log_level2)
+					memcpy(old_env, cur_env, sizeof(*old_env));
+				err = transfer(env, cur_env, idx);
+				if (err)
+					goto out;
+				reset_scevs_at_indirect_writes(env, cur_env, idx);
+				if (log_level2)
+					log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
+				succ = bpf_insn_successors(env, idx);
+				if (succ->cnt != 1)
+					break;
+				succ_idx = succ->items[0];
+				if (aux[idx].bb_end || aux[succ_idx].need_scev ||
+				    bpf_loop_at_index(env, succ_idx) != cur_header)
+					break;
+				idx = succ_idx;
+			}
+			/* Join cur_env into basic block successor's envs. */
+			iarray_for_each(s, succ) {
+				err = join_successor(env, cur_header, s, old_env, cur_env);
+				if (err)
+					goto out;
+			}
+		}
+	}
+
+	err = 0;
+out:
+	kfree(cur_env);
+	kfree(old_env);
+	return err;
+nomem:
+	err = -ENOMEM;
+	goto out;
+}
+
+static void mark_latches(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop;
+	int len = env->prog->len;
+	int i, j, latch;
+
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		aux[i].need_scev = true;
+		if (loop->irreducible)
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			latch = loop->backedges[j].latch;
+			if (latch >= 0)
+				aux[latch].need_scev = true;
+		}
+	}
+}
+
+static bool is_any_imm_reg_opaque(struct scev *scev, u32 id)
+{
+	u32 l, r, reg, order;
+	s64 imm;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT)
+			return false;
+		if (!(order & PRE) ||
+		    is_imm(scev, id, &imm) ||
+		    is_reg(scev, id, &reg) ||
+		    is_opaque(scev, id) ||
+		    is_any(scev, id, &l, &r))
+			continue;
+		return false;
+	}
+	return true;
+}
+
+/*
+ * Can implement explicit stack version, but it is harder to read.
+ * Stick with recursive version for now.
+ */
+static int transform_expr_once(struct scev *scev, u32 lvl, u32 root, void *priv,
+                               int (*fn)(struct scev *scev, u32 id, void *priv))
+{
+	struct expr expr;
+	int p0, p1, id;
+
+	if (lvl >= EXPR_STACK_DEPTH)
+		return UNKNOWN_EXPR_ID;
+
+        expr = scev->exprs[root]; /* snapshot the expr before potential realloc */
+	switch (op_params_num(expr.op)) {
+	case 0:
+		id = root;
+		break;
+	case 1:
+		p0 = transform_expr_once(scev, lvl + 1, expr.params[0], priv, fn);
+		id = expr1(scev, expr.op, p0);
+		break;
+	case 2:
+		p0 = transform_expr_once(scev, lvl + 1, expr.params[0], priv, fn);
+		p1 = transform_expr_once(scev, lvl + 1, expr.params[1], priv, fn);
+		id = expr2(scev, expr.op, p0, p1);
+		break;
+	}
+	return id < 0 ? id : fn(scev, id, priv);
+}
+
+static int transform_expr(struct scev *scev, u32 root, void *priv,
+                          int (*fn)(struct scev *scev, u32 id, void *priv))
+{
+	int id_old, id_new = root;
+
+	do {
+		id_old = id_new;
+		id_new = transform_expr_once(scev, 0, id_old, priv, fn);
+		if (id_new < 0)
+			return id_new;
+	} while (id_old != id_new);
+	return id_new;
+}
+
+static int simplify(struct scev *scev, u32 id, void *priv)
+{
+	u32 l, r, base, slope;
+	s64 imm1, imm2;
+
+	/* (+ (linear base slope) imm) -> (linear (+ base imm) slope) */
+	if (is_add(scev, id, &l, &r) &&
+	    is_linear(scev, l, &base, &slope) &&
+	    is_imm(scev, r, &imm1))
+		return expr2(scev, LINEAR_SCEV, expr2(scev, BPF_ADD, base, r), slope);
+
+	/* (+ imm imm) -> imm */
+	if (is_add(scev, id, &l, &r) &&
+	    is_imm(scev, l, &imm1) &&
+	    is_imm(scev, r, &imm2))
+		return imm_expr(scev, imm1 + imm2);
+
+	if (is_zext32(scev, id, &l) && is_imm(scev, l, &imm1))
+		return imm_expr(scev, (u64)(u32)imm1);
+
+	return id;
+}
+
+static int compute_header_scevs(struct bpf_verifier_env *env, struct env *header_env)
+{
+	struct scev *scev = env->scev;
+	u32 ra, rb, l, r, ra_expr;
+	s64 imm;
+	int id;
+
+	for (ra = 0; ra < REGS_NUM; ra++) {
+		id = transform_expr(scev, header_env->reg2expr[ra], NULL, simplify);
+		if (id < 0)
+			return id;
+		header_env->reg2expr[ra] = id;
+		ra_expr = header_env->reg2expr[ra];
+		/* rA = (+ rA IMM) */
+		if (is_add(scev, ra_expr, &l, &r) &&
+		    is_reg(scev, l, &rb) &&
+		    is_imm(scev, r, &imm) &&
+		    ra == rb) {
+			id = expr2(scev, LINEAR_SCEV, l, r);
+			if (id < 0)
+				return id;
+			header_env->reg2scev[ra] = id;
+			continue;
+		}
+		/* rA = rA */
+		if (is_reg(scev, ra_expr, &rb) && ra == rb) {
+			header_env->reg2scev[ra] = ra_expr;
+			continue;
+		}
+		/* rA = (any 1 2 3 4 ...) */
+		if (is_any_imm_reg_opaque(scev, ra_expr)) {
+			header_env->reg2scev[ra] = ra_expr;
+			continue;
+		}
+
+	}
+	return 0;
+}
+static int instantiate_header_scevs(struct scev *scev, u32 id, void *priv)
+{
+	u32 reg, base, slope, *reg2scev = priv;
+
+	/* (reg r) -> (linear ...), where r is a LINEAR_SCEV in the header */
+	if (is_reg(scev, id, &reg) &&
+	    is_linear(scev, reg2scev[reg], &base, &slope))
+		return reg2scev[reg];
+	return id;
+}
+
+static int compute_insn_scevs(struct bpf_verifier_env *env, struct env *eheader, struct env *einsn)
+{
+	struct scev *scev = env->scev;
+	int id, reg;
+
+	for (reg = 0; reg < REGS_NUM; reg++) {
+		id = einsn->reg2expr[reg];
+		id = transform_expr_once(scev, 0, id, eheader->reg2scev, instantiate_header_scevs);
+		if (id < 0)
+			return id;
+		id = transform_expr(scev, id, NULL, simplify);
+		if (id < 0)
+			return id;
+		einsn->reg2scev[reg] = id;
+	}
+
+	return 0;
+}
+
+static void log_scevs(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct bpf_loop *loop;
+	int i, j, len, latch;
+
+	len = env->prog->len;
+	for (i = 0; i < len; i++) {
+		loop = aux[i].loop;
+		if (!loop)
+			continue;
+		bpf_log(log, "scev at header %d:\n", i);
+		print_env(env, find_header_env(scev, i), i, PRINT_SCEV);
+		if (loop->irreducible)
+			continue;
+		for (j = 0; j < loop->backedges_cnt; j++) {
+			latch = loop->backedges[j].latch;
+			if (latch < 0)
+				continue;
+			bpf_log(log, " scev at latch %d:\n", latch);
+			print_env(env, find_loop_env(scev, bpf_loop_at_index(env, latch), latch),
+				  latch, PRINT_SCEV);
+		}
+	}
+}
+
+int bpf_compute_scev(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct scev *scev = env->scev;
+	int *postorder = env->cfg.insn_postorder;
+	int cnt = env->cfg.cur_postorder;
+	int i, idx, err, header;
+
+	mark_latches(env);
+	/*
+	 * Visit loop headers in postorder, to guarantee that scevs
+         * for innermost loops are computed first.
+	 */
+	for (i = 0; i < cnt; i++) {
+		idx = postorder[i];
+		if (!aux[idx].loop)
+			continue;
+		err = compute_scev_for_loop(env, idx);
+		if (err)
+			return err;
+	}
+
+	/*
+	 * Compute scevs from exprs collected on a previous step. Iterate instructions in
+	 * reverse post-order so that each loop header is processed before instructions
+	 * reachable from it.
+	 */
+	for (i = cnt - 1; i >= 0; i--) {
+		idx = postorder[i];
+		if (!aux[idx].need_scev)
+			continue;
+		header = bpf_loop_at_index(env, idx);
+		err = aux[idx].loop
+		      ? compute_header_scevs(env, find_header_env(scev, header))
+		      : compute_insn_scevs(env, find_header_env(scev, header),
+				   find_loop_env(scev, header, idx));
+		if (err)
+			return err;
+	}
+
+	for (i = 0; i < env->prog->len; i++)
+		if (aux[i].loop)
+			aux[i].prune_point = true;
+
+	if (env->log.level & BPF_LOG_LEVEL2)
+		log_scevs(env);
+
+	return 0;
+}
+
+static int reverse_ranked_compare(int a, int b, void *arg)
+{
+	int *rank = arg;
+
+	return rank[b] - rank[a];
+}
+
+void bpf_free_scev(struct bpf_verifier_env *env)
+{
+	struct scev *scev = env->scev;
+	struct insn_envs *envs;
+	int i;
+	u32 j;
+
+	if (!scev)
+		return;
+	for (i = 0; i < scev->envs_cnt; i++) {
+		envs = scev->envs[i];
+		if (!envs)
+			continue;
+		for (j = 0; j < envs->cnt; j++)
+			kfree(envs->entries[j].env);
+		kfree(envs);
+	}
+	for (i = 0; i < ARRAY_SIZE(scev->exprs_ht); i++)
+		kvfree(scev->exprs_ht[i]);
+	bpf_min_heap_free(&scev->worklist);
+	kvfree(scev->envs);
+	kvfree(scev->exprs);
+	kvfree(scev->discovered);
+	kfree(scev);
+	env->scev = NULL;
+}
+
+int bpf_init_scev(struct bpf_verifier_env *env)
+{
+	struct scev *scev;
+
+	scev = kzalloc(sizeof(struct scev), GFP_KERNEL_ACCOUNT);
+	if (!scev)
+		return -ENOMEM;
+	env->scev = scev;
+	/* Order worklist in reverse post-order. */
+	bpf_min_heap_init(&scev->worklist, reverse_ranked_compare, env->cfg.postorder_nums);
+	if (expr0(scev, UNKNOWN) < 0 || expr0(scev, OPAQUE)  < 0)
+		goto nomem;
+	scev->envs = kvcalloc(env->prog->len, sizeof(*scev->envs), GFP_KERNEL_ACCOUNT);
+	scev->discovered = kvcalloc(env->prog->len, sizeof(*scev->discovered), GFP_KERNEL_ACCOUNT);
+	if (!scev->envs || !scev->discovered)
+		goto nomem;
+	/*
+	 * Remember original program length, in case bpf_free_scev()
+         * is called after bpf program rewrites that increase program
+         * length.
+	 */
+	scev->envs_cnt = env->prog->len;
+	return 0;
+nomem:
+	bpf_free_scev(env);
+	return -ENOMEM;
+}
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0d18b3bb0865..57085f1bf915 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22835,6 +22835,14 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (env->log.level & BPF_LOG_LEVEL2)
 		log_program(env);
 
+	ret = bpf_init_scev(env);
+	if (ret < 0)
+		goto skip_full_check;
+
+	ret = bpf_compute_scev(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	ret = mark_fastcall_patterns(env);
 	if (ret < 0)
 		goto skip_full_check;
@@ -22984,6 +22992,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 		bpf_clear_insn_aux_data(env, 0, env->insn_aux_data_len);
 	vfree(env->insn_aux_data);
 	kvfree(env->fd_array);
+	bpf_free_scev(env);
 	bpf_stack_liveness_free(env);
 	kvfree(env->cfg.postorder_nums);
 	kvfree(env->cfg.insn_postorder);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (23 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:42   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
                   ` (10 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Loop-bound computation and register widening
===========================================

During the main verification pass, use SCEV expressions and register
values at the entry to the loop to estimate the number of iterations
the loop might execute. Use this number to widen loop induction variables.

For example:

    0: r7 = 5;
    1: r6 = 0;
    2: while (r6 < 3) {   // latch
    3:     if (r6 == 1)
    4:         r7 = 10;
    5:     r6 += 1;
    6: }
    8: use(r7);

- Scalar evolution expressions (SCEVs) at the latch are:
  - r6 = (linear r6 1)
  - r7 = (any r7 10)
- For this pre-condition loop, this corresponds to the linear inequality
  `r6_entry + n < 3`. With r6_entry = 0, the inequality is `n < 3`.
- Hence infer that r6's range at line 2 is [0,3], including the final
  failed test. The body starts with r6 in [0,2].
- r7's (any r7 10) expression means that its initial value 5 from line 0
  is joined with 10 from line 4: it is widened to [5,10] within
  and after the loop.

Currently such inference is supported for:
- Reducible loops with one backedge and complete backedge/exit lists.
- A linear latch dominating the backedge, with one successor exiting
  the loop. Additional exits are allowed; they can shorten execution.
- 64-bit arithmetic and comparisons whose continuation condition
  normalizes to signed/unsigned < or <=, or !=.
- Initial counter, latch offset, slope and bound evaluable as constants
  at loop entry, yielding a finite, representable iteration bound.
- Values with an eligible type at loop entry, e.g. SCALAR_VALUE or
  PTR_TO_STACK, but not PTR_TO_CTX.

In order for widening to proceed, every register live at the loop
entry has to have a SCEV that can be widened. Currently, the following
expressions are supported:
- (linear <self> <step>): entry value + iteration# * constant step.
- (reg <self>): loop invariant; keep its entry value unchanged.
- (any ...): join the entry value with immediates or invariant registers;
  types must match, and immediates require SCALAR_VALUE.

Main verification loop integration
==================================

Before proceeding with the usual instruction processing, do_check()
takes the following steps at loop entries (H):
- Checks whether the loop can be widened and obtains its
  iteration bound (scev.c:bpf_compute_loop_iters()).
- If it can, saves an unwidened checkpoint E0, which will be used as
  a base for deriving widened values throughout this invocation of
  the loop.
- Widens registers in the current state using E0 and SCEV expressions
  computed for H (scev.c:bpf_widen_scev_regs()).
- Pushes a per-frame loop-stack record to the current verifier state.
  For terminating loops, the record stores E0 and the iteration bounds.
- Lets is_state_visited() create a new checkpoint W at the same
  header, now containing the widened values, and proceeds with
  normal checking.
- On an exit edge, pops the loop record and verifies the exit
  path normally.

is_state_visited() treats a loop proven to terminate like an
iterator-based loop: RANGE_WITHIN comparison attempts to establish
convergence against a checkpoint whose exploration is still
in progress.

For example:

    r0 = 0;
    H: r0 += 1;
       if (r0 != 3) goto H;
       exit;

    first H:  prove H_count = 3; save E0(r0 = 0)
              widen r0 to [0,2]; push the loop record
              save W(r0 = [0,2])
    body:     r0 becomes [1,3]; fork at the conditional
    exit:     r0 = 3; pop the loop record and check exit
    backedge: r0 = [1,2]; clamp, then match with W and prune

Clamping
--------

Independent register ranges lose correlations between induction
variables. For example:

    r6 = 0; r7 = 0;
    do { r6 += 1; r7 += 2; } while (r6 < 3);
    use(r7);

    header:   r6 in [0,2], r7 in {0,2,4}
    body:     r6 in [1,3], r7 in {2,4,6}
    backedge: r6 in [1,2], r7 still in {2,4,6}

The latch narrows r6 but does not propagate this refinement to r7.
To account for this, before comparing states on a backedge,
scev.c:bpf_clamp_scev_regs() recomputes r7's range from E0 and the
saved iteration bound, then intersects the current range {2,4,6}
with the recomputed {0,2,4}, arriving at {2,4}.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h                       |  31 +
 kernel/bpf/log.c                                   |   5 +
 kernel/bpf/scev.c                                  | 653 +++++++++++++++++++++
 kernel/bpf/states.c                                |  67 ++-
 kernel/bpf/verifier.c                              | 226 ++++++-
 tools/testing/selftests/bpf/progs/verifier_gotox.c |   6 +-
 6 files changed, 976 insertions(+), 12 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 54a971800ea3..c0eea472949d 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -328,6 +328,21 @@ struct bpf_retval_range {
 	bool return_32bit;
 };
 
+struct bpf_loop_iters {
+	u32 min_header_count;   /* min number of times header is executed */
+	u32 max_header_count;   /* max number of times header is executed */
+	bool pre_cond;
+};
+
+struct loop_stack_entry {
+	struct bpf_verifier_state *entry_state;
+	struct bpf_loop_iters iters;
+	u32 loop_id:31;
+	u32 terminates:1;
+};
+
+#define LOOP_STACK_SIZE 16
+
 /* state of the program:
  * type of all registers and stack info
  */
@@ -373,6 +388,12 @@ struct bpf_func_state {
 	u32 callback_depth;
 	/* Instructions processed in this frame and callees on the current path. */
 	u32 insns_subtotal;
+	/*
+	 * Control-flow loop nesting at the current insn within this frame's
+	 * subprogram (loops never cross subprogram boundaries).
+	 */
+	u32 loop_stack_cnt;
+	struct loop_stack_entry loop_stack[LOOP_STACK_SIZE];
 
 	/* The following fields should be last. See copy_func_state() */
 	/* The state of the stack. Each element of the array describes BPF_REG_SIZE
@@ -436,6 +457,7 @@ static_assert(MAX_BPF_STACK_SLOTS <= (1 << 12));
 #define MAX_STACK_ARG_SLOTS (MAX_BPF_FUNC_ARGS - MAX_BPF_FUNC_REG_ARGS)
 #define BPF_ID_MAP_SIZE ((MAX_BPF_REG + MAX_BPF_STACK_SLOTS + MAX_STACK_ARG_SLOTS) * \
 			 MAX_CALL_FRAMES)
+
 struct bpf_verifier_state {
 	/* call stack tracking */
 	struct bpf_func_state *frame[MAX_CALL_FRAMES];
@@ -1156,6 +1178,8 @@ struct bpf_verifier_env {
 	struct bpf_iarray *gotox_tmp_buf;
 	int *idoms;
 	struct scev *scev;
+	/* SCEV representation of stack slots not allocated yet. */
+	struct bpf_reg_state scev_not_init_reg;
 };
 
 static inline struct bpf_func_info_aux *subprog_aux(struct bpf_verifier_env *env, int subprog)
@@ -1944,4 +1968,11 @@ int bpf_init_scev(struct bpf_verifier_env *env);
 void bpf_free_scev(struct bpf_verifier_env *env);
 int bpf_compute_scev(struct bpf_verifier_env *env);
 
+int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			   struct bpf_loop_iters *iters);
+int bpf_widen_scev_regs(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			struct bpf_verifier_state *loop_entry, struct bpf_loop_iters *iters);
+int bpf_clamp_scev_regs(struct bpf_verifier_env *env, struct bpf_func_state *st, u32 insn_idx,
+			struct bpf_verifier_state *entry_state, struct bpf_loop_iters *iters);
+
 #endif /* _LINUX_BPF_VERIFIER_H */
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index 900d1bb1988b..6a5564f62669 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -792,6 +792,11 @@ void print_verifier_state(struct bpf_verifier_env *env, const struct bpf_verifie
 		verbose(env, " cb");
 	if (state->in_async_callback_fn)
 		verbose(env, " async_cb");
+	if (state->loop_stack_cnt) {
+		verbose(env, " loop_stack=");
+		for (i = 0; i < state->loop_stack_cnt; i++)
+			verbose(env, "%s%d", i ? "," : "", state->loop_stack[i].loop_id);
+	}
 	verbose(env, "\n");
 	if (!print_all)
 		mark_verifier_state_clean(env);
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
index 396b0c776dfa..205bf8684c3d 100644
--- a/kernel/bpf/scev.c
+++ b/kernel/bpf/scev.c
@@ -1,9 +1,12 @@
 // SPDX-License-Identifier: GPL-2.0-only
 /* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
 
+#include "linux/cnum.h"
 #include <linux/bpf_verifier.h>
 #include <linux/jhash.h>
 #include <linux/bug.h>
+#include <linux/tnum.h>
+#include <linux/overflow.h>
 
 #define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
 #define UNKNOWN_EXPR_ID 0
@@ -1498,6 +1501,7 @@ int bpf_init_scev(struct bpf_verifier_env *env)
 	if (!scev)
 		return -ENOMEM;
 	env->scev = scev;
+	bpf_mark_reg_not_init(env, &env->scev_not_init_reg);
 	/* Order worklist in reverse post-order. */
 	bpf_min_heap_init(&scev->worklist, reverse_ranked_compare, env->cfg.postorder_nums);
 	if (expr0(scev, UNKNOWN) < 0 || expr0(scev, OPAQUE)  < 0)
@@ -1517,3 +1521,652 @@ int bpf_init_scev(struct bpf_verifier_env *env)
 	bpf_free_scev(env);
 	return -ENOMEM;
 }
+
+static struct bpf_reg_state *scev_regno_to_reg(struct bpf_verifier_env *env,
+					    struct bpf_func_state *st, u32 r)
+{
+	int spi, slots_available;
+
+	if (r < MAX_BPF_REG)
+		return &st->regs[r];
+
+	slots_available = st->allocated_stack / BPF_REG_SIZE;
+	spi = r - MAX_BPF_REG;
+	if (spi < slots_available)
+		return &st->stack[spi].spilled_ptr;
+
+	return &env->scev_not_init_reg;
+}
+
+static bool scev_reg_alive(struct bpf_verifier_env *env, struct bpf_verifier_state *st, u32 r)
+{
+	int insn_idx = bpf_frame_insn_idx(st, st->curframe);
+	u16 live_regs = env->insn_aux_data[insn_idx].live_regs_before;
+	int spi;
+
+	if (r < MAX_BPF_REG) {
+		return BIT(r) & live_regs;
+	} else {
+		spi = r - MAX_BPF_REG;
+		return bpf_stack_slot_alive(env, st->curframe, spi * 2) ||
+		       bpf_stack_slot_alive(env, st->curframe, spi * 2 + 1);
+	}
+}
+
+/*
+ * Latch is a condition deciding if execution remains inside a loop.
+ * Linear latch represents a condition 'if <reg> <op> <loop invariant> goto <loop-header>',
+ * where equation '<base> + i * <step> <op> <bound>' describes values taken by register <reg>,
+ * 'i' is the loop iteration number, starting from 0.
+ */
+struct linear_latch {
+	u32 insn_idx;
+	u32 base_expr;
+	u32 step_expr;
+	u32 bound_expr;
+	u32 reg_expr;
+	u32 reg;
+	u32 op;
+};
+
+static bool loop_invariant(struct scev *scev, u32 id)
+{
+	s64 imm;
+	u32 reg;
+
+	return is_reg(scev, id, &reg) || is_imm(scev, id, &imm);
+}
+
+static int match_linear_latch(struct bpf_verifier_env *env,
+			      u32 latch_idx,
+			      struct linear_latch *latch)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[latch_idx];
+	struct scev *scev = env->scev;
+	struct env *latch_env = find_loop_env(scev, bpf_loop_at_index(env, latch_idx), latch_idx);
+	u32 true_branch_tgt;
+	u32 l, r, op, base;
+	u32 src_reg_scev;
+	u32 dst_reg_scev;
+	int id;
+
+	/* 32-bit arithmetic is not handled yet */
+	if (BPF_CLASS(insn->code) != BPF_JMP)
+		return false;
+	op = BPF_OP(insn->code);
+	/* Flip the condition if true branch jumps out of the loop */
+	true_branch_tgt = latch_idx + bpf_jmp_offset(insn) + 1;
+	if (bpf_loop_at_index(env, true_branch_tgt) != bpf_loop_at_index(env, latch_idx))
+		op = bpf_rev_opcode(op);
+	switch (op) {
+	case BPF_JSLT:
+	case BPF_JSLE:
+	case BPF_JSGT:
+	case BPF_JSGE:
+	case BPF_JLT:
+	case BPF_JLE:
+	case BPF_JGT:
+	case BPF_JGE:
+	case BPF_JNE:
+		break;
+	default:
+		return false;
+	}
+
+	latch->op = op;
+	latch->insn_idx = latch_idx;
+
+	dst_reg_scev = latch_env->reg2scev[insn->dst_reg];
+	src_reg_scev = latch_env->reg2scev[insn->src_reg];
+	if (!is_linear(scev, dst_reg_scev, &base, &latch->step_expr))
+		return false;
+
+	if (BPF_SRC(insn->code) == BPF_K) {
+		id = imm_expr(scev, insn->imm);
+		if (id < 0)
+			return id;
+		latch->bound_expr = id;
+	} else {
+		latch->bound_expr = src_reg_scev;
+	}
+
+	if (is_reg(scev, base, &latch->reg)) {
+		id = imm_expr(scev, 0);
+		if (id < 0)
+			return id;
+		latch->base_expr = id;
+		latch->reg_expr = base;
+	} else if (is_add(scev, base, &l, &r) &&
+		   is_reg(scev, l, &latch->reg)) {
+		latch->base_expr = r;
+		latch->reg_expr = l;
+	} else {
+		return false;
+	}
+
+	if (!loop_invariant(scev, latch->base_expr) ||
+	    !loop_invariant(scev, latch->step_expr) ||
+	    !loop_invariant(scev, latch->bound_expr))
+		return false;
+
+	return true;
+}
+
+static bool eval_expr(struct bpf_verifier_env *env, struct scev *scev, struct bpf_func_state *st, u32 id, u64 *result)
+{
+	struct bpf_reg_state *reg;
+	u32 regno;
+	s64 imm;
+
+	if (is_reg(scev, id, &regno)) {
+		reg = scev_regno_to_reg(env, st, regno);
+		if (reg->type == NOT_INIT)
+			return false;
+		if (tnum_is_const(reg->var_off)) {
+			*result = reg->var_off.value;
+			return true;
+		}
+	} else if (is_imm(scev, id, &imm)) {
+		*result = imm;
+		return true;
+	}
+	return false;
+}
+
+/* Like DIV_ROUND_UP() but overflow safe */
+static u64 div_round_up(u64 a, u64 b)
+{
+	return a / b + (a % b != 0);
+}
+
+/*
+ * Loop with post-condition:
+ *
+ *    r0 = 0
+ * l: ...                r0 ∈ [0,1,2] header executed 3 times
+ *    r0 += 1            r0 ∈ [0,1,2]
+ *    ...                r0 ∈ [1,2,3]
+ *    if r0 != 3 goto l  r0 ∈ [1,2,3] backedge taken 2 times
+ *    ...                r0 ∈ [3]
+ *
+ * SCEV at header: r0 = k
+ * SCEV at latch:  r0 = 1 + k
+ *
+ * Loop with pre-condition:
+ *
+ *    r0 = 0
+ * l: ...                r0 ∈ [0,1,2,3] header executed 4 times
+ *    if r0 == 3 goto e  r0 ∈ [0,1,2,3] backedge taken 3 times
+ *    r0 += 1            r0 ∈ [0,1,2]
+ *    ...                r0 ∈ [1,2,3]
+ *    goto l             r0 ∈ [1,2,3]
+ * e: ...                r0 ∈ [3]
+ *
+ * SCEV at header: r0 = k
+ * SCEV at latch:  r0 = k
+ */
+static bool compute_max_iters(struct bpf_verifier_env *env,
+			      struct bpf_func_state *st,
+			      struct linear_latch *latch,
+			      struct bpf_loop_iters *iters)
+{
+	u64 base, step, diff, bound, initial, max_iters;
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_loop *loop = aux[bpf_loop_at_index(env, latch->insn_idx)].loop;
+	struct scev *scev = env->scev;
+	u8 op = latch->op;
+
+	if (!eval_expr(env, scev, st, latch->base_expr, &base) ||
+	    !eval_expr(env, scev, st, latch->step_expr, &step) ||
+	    !eval_expr(env, scev, st, latch->bound_expr, &bound) ||
+	    !eval_expr(env, scev, st, latch->reg_expr, &initial))
+		return false;
+
+	if (step == 0)
+		return false;
+	if ((s64)step == S64_MIN)
+		return false;
+	if ((s64)step < 0) {
+		/* Multiply both sides of the equation by -1, e.g. -2*i > -3 becomes 2*i < 3 */
+		op = bpf_flip_opcode(op);
+		step = -step;
+		swap(bound, initial);
+	}
+	diff = bound - initial;
+	if (diff / step == U64_MAX)
+		return false;
+	switch (op) {
+	case BPF_JLT:
+		max_iters = (u64)initial >= (u64)bound ? 0 : div_round_up(diff, step);
+		break;
+	case BPF_JLE:
+		max_iters = (u64)initial >  (u64)bound ? 0 : diff / step + 1;
+		break;
+	case BPF_JSLT:
+		max_iters = (s64)initial >= (s64)bound ? 0 : div_round_up(diff, step);
+		break;
+	case BPF_JSLE:
+		max_iters = (s64)initial >  (s64)bound ? 0 : diff / step + 1;
+		break;
+	case BPF_JNE:
+		max_iters = diff % step ? U32_MAX : div_round_up(diff, step);
+		break;
+	default:
+		return false;
+	}
+
+	if (max_iters > U32_MAX)
+		return false;
+
+	/*
+	 * The latch is a conditional jump with one jump target exiting the loop.
+	 * Linear latch is matched only if the loop has a single backedge.
+	 * The loop still, however can have multiple exits.
+	 * In such case, conservatively assume that non-latch exit can happen
+	 * at any iteration, thus setting minimal number of iterations as 0.
+	 */
+	iters->max_header_count = max_iters + (base == 0 ? 1 : 0);
+	iters->min_header_count = loop->exits_cnt == 1 ? iters->max_header_count : 0;
+	iters->pre_cond = base == 0;
+	return true;
+}
+
+static void mark_scev_reg_scratched(struct bpf_verifier_env *env, u32 r)
+{
+	if (r < MAX_BPF_REG)
+		mark_reg_scratched(env, r);
+	else
+		mark_stack_slot_scratched(env, r - __MAX_BPF_REG);
+}
+
+/* Main logic in verifier.c forbids varying offsets for certain register types. */
+static bool is_widenable_reg_type(const struct bpf_reg_state *reg)
+{
+	switch (base_type(reg->type)) {
+	case SCALAR_VALUE:
+	case PTR_TO_MAP_VALUE:
+	case PTR_TO_MAP_KEY:
+	case PTR_TO_STACK:
+	case PTR_TO_PACKET:
+	case PTR_TO_PACKET_META:
+	case PTR_TO_MEM:
+	case PTR_TO_BUF:
+	case PTR_TO_BTF_ID:
+		return true;
+	default:
+		return false;
+	}
+}
+
+/* Check that ANY leaves can be unioned with r's loop-entry value. */
+static bool is_any_imm_reg(struct bpf_verifier_env *env, struct bpf_func_state *loop_entry,
+			   struct env *header_env, u32 r, u32 id)
+{
+	struct scev *scev = env->scev;
+	struct bpf_reg_state *reg, *leaf_reg;
+	u32 l, rr, leaf, ra, order;
+	s64 imm;
+
+	if (!is_any(scev, id, &l, &rr))
+		return false;
+	reg = scev_regno_to_reg(env, loop_entry, r);
+	if (!is_widenable_reg_type(reg))
+		return false;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT)
+			return false;
+		if (!(order & PRE) || is_any(scev, id, &l, &rr))
+			continue;
+		if (is_imm(scev, id, &imm)) {
+			if (reg->type != SCALAR_VALUE)
+				return false;
+		} else if (is_reg(scev, id, &leaf)) {
+			/* Other leaves must denote loop-invariant registers. */
+			if (leaf != r &&
+			    !(is_reg(scev, header_env->reg2scev[leaf], &ra) && ra == leaf))
+				return false;
+			leaf_reg = scev_regno_to_reg(env, loop_entry, leaf);
+			if (leaf_reg->type != reg->type)
+				return false;
+		} else {
+			return false;
+		}
+	}
+	return true;
+}
+
+struct bounds {
+	struct cnum64 range;
+	u16 step;
+};
+
+static bool is_simple_linear(struct scev *scev, u32 id, u32 *base_reg, s64 *slope_imm)
+{
+	u32 base, slope;
+
+	return is_linear(scev, id, &base, &slope) &&
+	       is_reg(scev, base, base_reg) &&
+	       is_imm(scev, slope, slope_imm) &&
+	       *slope_imm <= S16_MAX &&
+	       *slope_imm >= S16_MIN &&
+	       *slope_imm != 0;
+}
+
+/*
+ * Compute the range an induction variable in `reg` spans over the loop.
+ * Returns false if the computation overflows s64.
+ */
+static bool linear_bounds(struct bpf_reg_state *reg, struct bpf_loop_iters *iters, s64 slope,
+			  struct bounds *out)
+{
+	s64 slope_abs = slope < 0 ? -slope : slope;
+	s64 min_val = reg_smin(reg);
+	s64 max_val = reg_smax(reg);
+	s64 total_change;
+	u16 step;
+
+	if (check_mul_overflow(slope, (s64)iters->max_header_count - 1, &total_change))
+		return false;
+	if (slope > 0) {
+		if (check_add_overflow(max_val, total_change, &max_val))
+			return false;
+	} else {
+		if (check_add_overflow(min_val, total_change, &min_val))
+			return false;
+	}
+	/*
+	 * If the entry value is a single point the value set is 'v + slope * k',
+	 * so the step is |slope|. Otherwise, only the power-of-two alignment
+	 * shared by the entry value and the slope.
+	 */
+	if (cnum64_is_const(reg->r64))
+		step = slope_abs;
+	else
+		step = 1u << min_t(u32, tnum_alignment(reg->var_off), __ffs(slope_abs));
+	out->range = cnum64_from_srange(min_val, max_val);
+	out->step = step;
+	return true;
+}
+
+int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			   struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *cur_func = st->frame[st->curframe];
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct env *header_env;
+	struct linear_latch latch;
+	struct bpf_loop *loop;
+	struct bounds bounds;
+	int insn_idx = st->insn_idx;
+	int linear_latch;
+	u32 r, base_reg;
+	int latch_idx;
+	s64 slope_imm;
+	int err;
+
+	/*
+	 * If insn_idx is a loop header for a reducible loop with a single backedge.
+	 * loop is NULL for secondary entries to irreducible loops.
+	 */
+	loop = aux[insn_idx].loop;
+	if (!loop || loop->irreducible || loop->backedges_cnt != 1 || loop->backedges_overflow ||
+	    loop->exits_overflow) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, unsupported loop:%s%s%s\n", insn_idx,
+				!loop || loop->irreducible ? " irreducible" : "",
+				loop && loop->backedges_cnt > 1 ? " multiple backedges" : "",
+				loop && loop->exits_overflow ? " too many exits" : "");
+		return 0;
+	}
+
+	/* If this backedge has a latch */
+	latch_idx = loop->backedges[0].latch;
+	if (latch_idx < 0)
+		return 0;
+
+	linear_latch = match_linear_latch(env, latch_idx, &latch);
+	if (linear_latch < 0)
+		return linear_latch;
+
+	if (!linear_latch) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, non-linear latch at %d\n",
+				insn_idx, latch_idx);
+		return 0;
+	}
+
+	if (!compute_max_iters(env, cur_func, &latch, iters)) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, can't compute iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (iters->max_header_count == 0) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, 0 iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (iters->max_header_count == U32_MAX) {
+		if (log->level & BPF_LOG_LEVEL2)
+			bpf_log(log, "loop header at %d, inf iterations count\n", insn_idx);
+		return 0;
+	}
+
+	if (log->level & BPF_LOG_LEVEL2) {
+		bpf_log(log, "loop header at %d, header_count is ", insn_idx);
+		if (iters->min_header_count == iters->max_header_count)
+			bpf_log(log, "%u ", iters->max_header_count);
+		else
+			bpf_log(log, "[%u..%u] ", iters->min_header_count, iters->max_header_count);
+		bpf_log(log, "%s\n", iters->pre_cond ? "pre-cond" : "post-cond");
+	}
+
+	err = bpf_live_stack_query_init(env, st);
+	if (err)
+		return err;
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		struct bpf_reg_state *reg;
+		u32 ra, r_scev, r_expr;
+
+		if (!scev_reg_alive(env, st, r))
+			continue;
+
+		r_scev = header_env->reg2scev[r];
+		reg = scev_regno_to_reg(env, cur_func, r);
+		/* If SCEV for r is (linear <reg> <slope>) */
+		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
+		    is_widenable_reg_type(reg)) {
+			/* Can't widen if the iteration range overflows */
+			if (base_reg == r &&
+			    !linear_bounds(reg, iters, slope_imm, &bounds))
+				goto cant_widen;
+			continue;
+		}
+
+		/* rA = rA, loop does not change this reg */
+		if (is_reg(scev, r_scev, &ra) && r == ra)
+			continue;
+
+		/* (any 1 (any 2 (any 3 4))) */
+		if (is_any_imm_reg(env, cur_func, header_env, r, r_scev))
+			continue;
+
+cant_widen:
+		if (log->level & BPF_LOG_LEVEL2) {
+			r_expr = header_env->reg2expr[r];
+			bpf_log(log, "loop header at %d, can't widen ", insn_idx);
+			log_reg(env, r);
+			bpf_log(log, ", expr is ");
+			log_expr(env, r_expr);
+			bpf_log(log, "\n");
+		}
+		return 0;
+	}
+
+	return 1;
+}
+
+static void scratch_scalar_id(struct bpf_reg_state *reg)
+{
+	if (reg->type == SCALAR_VALUE)
+		reg->id = 0;
+}
+
+/* The filtering pass has checked that every leaf can be unioned into acc. */
+static int union_any_reg(struct bpf_verifier_env *env, struct bpf_func_state *loop_entry,
+			 struct bpf_reg_state *acc, u32 id)
+{
+	struct bpf_reg_state *tmp = &env->fake_reg[0], *leaf_reg;
+	struct scev *scev = env->scev;
+	u32 l, r, leaf, order;
+	s64 imm;
+	int err;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if (order & DEPTH_LIMIT) {
+			verifier_bug(env, "scev ANY union exceeds expression depth limit");
+			return -EFAULT;
+		}
+		if (!(order & PRE) || is_any(scev, id, &l, &r))
+			continue;
+		if (is_imm(scev, id, &imm)) {
+			bpf_mark_reg_known_scalar(tmp, imm);
+			leaf_reg = tmp;
+		} else if (is_reg(scev, id, &leaf)) {
+			leaf_reg = scev_regno_to_reg(env, loop_entry, leaf);
+		} else {
+			verifier_bug(env, "scev ANY union has an unsupported leaf");
+			return -EFAULT;
+		}
+		err = bpf_reg_union(env, acc, leaf_reg);
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
+int bpf_widen_scev_regs(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
+			struct bpf_verifier_state *loop_entry, struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *cur_func = st->frame[st->curframe];
+	struct bpf_func_state *entry_func = loop_entry->frame[loop_entry->curframe];
+	struct bpf_verifier_log *log = &env->log;
+	struct scev *scev = env->scev;
+	struct bpf_reg_state *reg;
+	struct env *header_env;
+	struct bounds bounds;
+	u32 r, base_reg, r_expr, a, b;
+	int insn_idx = st->insn_idx;
+	s64 slope_imm;
+	int err;
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		if (!scev_reg_alive(env, st, r))
+			continue;
+
+		r_expr = header_env->reg2scev[r];
+		/* If SCEV for r is (linear <reg> <slope>)*/
+		if (is_simple_linear(scev, r_expr, &base_reg, &slope_imm) &&
+		    base_reg == r) {
+			reg = scev_regno_to_reg(env, cur_func, r);
+			/* Feasibility was checked in the filtering pass above. */
+			if (!linear_bounds(reg, iters, slope_imm, &bounds)) {
+				verifier_bug(env, "scev widen bounds overflow for r%d", r);
+				return -EFAULT;
+			}
+			if (log->level & BPF_LOG_LEVEL2) {
+				bpf_log(log, "loop header at %d, widening ", insn_idx);
+				log_reg(env, r);
+				bpf_log(log, " to %lld..%lld step %u\n",
+					cnum64_smin(bounds.range), cnum64_smax(bounds.range),
+					bounds.step);
+			}
+			scratch_scalar_id(reg);
+			err = bpf_set_reg_range(env, reg, bounds.range, bounds.step);
+			if (err)
+				return err;
+			mark_scev_reg_scratched(env, r);
+		} else if (is_any(scev, r_expr, &a, &b)) {
+			reg = scev_regno_to_reg(env, cur_func, r);
+			err = union_any_reg(env, entry_func, reg, r_expr);
+			if (err)
+				return err;
+			if (log->level & BPF_LOG_LEVEL2) {
+				bpf_log(log, "loop header at %d, widening ", insn_idx);
+				log_reg(env, r);
+				bpf_log(log, " to %lld..%lld step %u\n",
+					reg_smin(reg), reg_smax(reg), reg->step);
+			}
+			scratch_scalar_id(reg);
+			mark_scev_reg_scratched(env, r);
+		}
+	}
+	return 1;
+}
+
+int bpf_clamp_scev_regs(struct bpf_verifier_env *env, struct bpf_func_state *cur_func_state, u32 insn_idx,
+			struct bpf_verifier_state *entry_state, struct bpf_loop_iters *iters)
+{
+	struct bpf_func_state *entry_st = entry_state->frame[entry_state->curframe];
+	struct bpf_verifier_log *log = &env->log;
+	struct bpf_reg_state *reg, *entry_reg;
+	struct scev *scev = env->scev;
+	struct env *header_env;
+	struct bounds bounds;
+	u32 r, base_reg;
+	s64 slope_imm;
+	int err;
+
+	if (entry_state->curframe != cur_func_state->frameno) {
+		verifier_bug(env, "clamping registers for a wrong frame: %d vs %d\n",
+			     entry_state->curframe, cur_func_state->frameno);
+		return -EFAULT;
+	}
+
+	header_env = find_header_env(scev, insn_idx);
+	for (r = 0; r < REGS_NUM; r++) {
+		/* If SCEV for r is (linear <reg> <slope>)*/
+		if (!is_simple_linear(scev, header_env->reg2scev[r], &base_reg, &slope_imm) ||
+		    base_reg != r)
+			continue;
+
+		entry_reg = scev_regno_to_reg(env, entry_st, r);
+		reg = scev_regno_to_reg(env, cur_func_state, r);
+		if (entry_reg->type == NOT_INIT || reg->type == NOT_INIT)
+			continue;
+
+		if (!linear_bounds(entry_reg, iters, slope_imm, &bounds)) {
+			verifier_bug(env, "scev clamp bounds overflow for r%d", r);
+			return -EFAULT;
+		}
+		bounds.range = cnum64_intersect(reg->r64, bounds.range);
+		if (cnum64_is_empty(bounds.range)) {
+			verifier_bug(env, "scev clamp produced empty range for r%d", r);
+			return -EFAULT;
+		}
+		if (log->level & BPF_LOG_LEVEL2) {
+			bpf_log(log, "loop header at %d, clamping ", insn_idx);
+			log_reg(env, r);
+			bpf_log(log, " to %lld..%lld step %u\n",
+				cnum64_smin(bounds.range), cnum64_smax(bounds.range),
+				bounds.step);
+		}
+		scratch_scalar_id(reg);
+		err = bpf_set_reg_range(env, reg, bounds.range, bounds.step);
+		if (err)
+			return err;
+		mark_scev_reg_scratched(env, r);
+	}
+	return 0;
+}
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 68df27e34b60..c1771a355e4e 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -942,6 +942,28 @@ static bool refsafe(struct bpf_verifier_state *old, struct bpf_verifier_state *c
 	return true;
 }
 
+static bool loop_stack_safe(struct bpf_verifier_env *env, struct bpf_func_state *old,
+			    struct bpf_func_state *cur)
+{
+	struct loop_stack_entry *old_stack = old->loop_stack;
+	struct loop_stack_entry *cur_stack = cur->loop_stack;
+	u32 i;
+
+	if (old->loop_stack_cnt != cur->loop_stack_cnt)
+		return false;
+
+	for (i = 0; i < old->loop_stack_cnt; i++) {
+		if (old_stack[i].loop_id == cur_stack[i].loop_id &&
+		    old_stack[i].iters.max_header_count == cur_stack[i].iters.max_header_count &&
+		    old_stack[i].iters.pre_cond == cur_stack[i].iters.pre_cond &&
+		    old_stack[i].terminates == cur_stack[i].terminates)
+			continue;
+		return false;
+	}
+
+	return true;
+}
+
 /* compare two verifier states
  *
  * all states stored in state_list are known to be valid, since
@@ -992,6 +1014,9 @@ static bool func_states_equal(struct bpf_verifier_env *env, struct bpf_func_stat
 	if (!stack_arg_safe(env, old, cur, &env->idmap_scratch, exact))
 		return false;
 
+	if (!loop_stack_safe(env, old, cur))
+		return false;
+
 	return true;
 }
 
@@ -1038,6 +1063,7 @@ static bool states_equal(struct bpf_verifier_env *env,
 		if (!func_states_equal(env, old->frame[i], cur->frame[i], insn_idx, exact))
 			return false;
 	}
+
 	return true;
 }
 
@@ -1320,15 +1346,30 @@ int bpf_split_cur_state(struct bpf_verifier_env *env)
 	return 0;
 }
 
+/* Force a checkpoint upon reaching a loop header for a terminating loop. */
+static bool need_loop_checkpoint(struct bpf_verifier_env *env, int insn_idx)
+{
+	struct bpf_verifier_state *cur = env->cur_state;
+	struct bpf_func_state *frame = cur->frame[cur->curframe];
+	struct loop_stack_entry *top;
+
+	if (bpf_loop_at_index(env, insn_idx) != insn_idx || !frame->loop_stack_cnt)
+		return false;
+	top = &frame->loop_stack[frame->loop_stack_cnt - 1];
+	return top->loop_id == insn_idx && top->terminates;
+}
+
 int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 {
 	struct bpf_verifier_state_list *sl;
 	struct bpf_verifier_state *cur = env->cur_state;
-	bool force_new_state, add_new_state, loop;
+	bool force_new_state, add_new_state, loop, loop_checkpoint;
 	int n, err, states_cnt = 0;
 	struct list_head *pos, *tmp, *head;
 
+	loop_checkpoint = need_loop_checkpoint(env, insn_idx);
 	force_new_state = env->test_state_freq || bpf_is_force_checkpoint(env, insn_idx) ||
+			  loop_checkpoint ||
 			  /* Avoid accumulating infinitely long jmp history */
 			  cur->jmp_history_cnt > 40;
 
@@ -1359,10 +1400,11 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 			continue;
 
 		if (sl->state.branches) {
-			struct bpf_func_state *frame = sl->state.frame[0];
+			struct bpf_func_state *old_top_frame = sl->state.frame[0];
+			struct bpf_func_state *old_cur_frame = sl->state.frame[sl->state.curframe];
 
-			if (frame->in_async_callback_fn &&
-			    frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
+			if (old_top_frame->in_async_callback_fn &&
+			    old_top_frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
 				/* Different async_entry_cnt means that the verifier is
 				 * processing another entry into async callback.
 				 * Seeing the same state is not an indication of infinite
@@ -1461,6 +1503,20 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 				}
 				goto skip_inf_loop_check;
 			}
+
+			/*
+			 * If old state belongs to a control flow loop that we know terminates,
+			 * it should be safe to prune current state.
+			 */
+			if (old_cur_frame->loop_stack_cnt &&
+			    old_cur_frame->loop_stack[old_cur_frame->loop_stack_cnt - 1].terminates) {
+				if (states_equal(env, &sl->state, cur, RANGE_WITHIN)) {
+					loop = true;
+					goto hit;
+				}
+				goto skip_inf_loop_check;
+			}
+
 			/* attempt to detect infinite loop to avoid unnecessary doomed work */
 			if (states_maybe_looping(&sl->state, cur) &&
 			    states_equal(env, &sl->state, cur, EXACT) &&
@@ -1619,7 +1675,8 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
 		 * Use bigger 'n' for checkpoints because evicting checkpoint states
 		 * too early would hinder iterator convergence.
 		 */
-		n = bpf_is_force_checkpoint(env, insn_idx) && sl->state.branches > 0 ? 64 : 3;
+		n = (bpf_is_force_checkpoint(env, insn_idx) || loop_checkpoint) &&
+		    sl->state.branches > 0 ? 64 : 3;
 		if (sl->miss_cnt > sl->hit_cnt * n + n) {
 			/* the state is unlikely to be useful. Remove it to
 			 * speed up verification
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 57085f1bf915..9f56ffe09eba 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19604,6 +19604,204 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 	return -EFAULT;
 }
 
+static bool has_entered_loop(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int prev_insn_idx = env->prev_insn_idx;
+	int insn_idx = env->insn_idx;
+
+	return aux[insn_idx].loop_entry &&
+	       (prev_insn_idx == -1 ||
+		bpf_loop_at_index(env, prev_insn_idx) != bpf_loop_at_index(env, insn_idx));
+}
+
+static bool is_next_loop_iteration(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int prev_insn_idx = env->prev_insn_idx;
+	int insn_idx = env->insn_idx;
+
+	return aux[insn_idx].loop_entry &&
+	       (prev_insn_idx != -1 &&
+		bpf_loop_at_index(env, prev_insn_idx) == bpf_loop_at_index(env, insn_idx));
+}
+
+static bool has_exited_loop(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	int prev_insn_idx = env->prev_insn_idx;
+	int insn_idx = env->insn_idx;
+	int prev_loop_idx, h;
+
+	if (prev_insn_idx == -1)
+		return false;
+
+	prev_loop_idx = bpf_loop_at_index(env, prev_insn_idx);
+	if (prev_loop_idx < 0)
+		return false;
+
+	/* Still inside the previous instruction's loop if that loop encloses insn_idx */
+	for (h = bpf_loop_at_index(env, insn_idx); h >= 0; h = aux[h].loop_header)
+		if (h == prev_loop_idx)
+			return false;
+
+	return true;
+}
+
+/*
+ * Push the loop entered at insn_idx onto the stack. Usually this adds a single
+ * entry on top of its enclosing loop. However, an edge may enter an inner loop
+ * directly, bypassing the header(s) of its enclosing loop(s), e.g.:
+ *
+ *   1: for (...):       // enclosing loop, header at 1
+ *   2:   for (...):     // inner loop, header at 2
+ *        ...
+ *   3: if ...:
+ *        goto 2b;       // enters loop 2 without going through header 1
+ *
+ * Such a bypass makes the enclosing loop irreducible, so its header is missing
+ * from the stack and both headers (1) and (2) need to be pushed onto stack.
+ * Only the innermost loop (the one actually entered at insn_idx) carries SCEV bounds;
+ * the bypassed ancestors are irreducible and pushed as non-terminating.
+ *
+ * Assumes loop_stack_pop() has already truncated the stack to the common ancestor
+ * of the bpf_loop_at_index(env->prev_insn_idx) and bpf_loop_at_index(env->insn_idx).
+ */
+static int loop_stack_push(struct bpf_verifier_env *env,
+			   struct bpf_verifier_state *entry_state,
+			   bool terminates,
+			   struct bpf_loop_iters *iters)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_func_state *frame = cur_func(env);
+	struct loop_stack_entry *loop_stack = frame->loop_stack;
+	u32 missing_headers[LOOP_STACK_SIZE];
+	u32 cnt = frame->loop_stack_cnt;
+	u32 num_missing = 0;
+	int h;
+
+	for (h = bpf_loop_at_index(env, env->insn_idx); h >= 0; h = aux[h].loop_header, num_missing++) {
+		if (cnt && loop_stack[cnt - 1].loop_id == h)
+			break;
+		if (num_missing == LOOP_STACK_SIZE)
+			goto e2big;
+		missing_headers[num_missing] = h;
+	}
+
+	for (; num_missing; num_missing--, cnt++) {
+		if (cnt == LOOP_STACK_SIZE)
+			goto e2big;
+		loop_stack[cnt].loop_id = missing_headers[num_missing - 1];
+		if (num_missing == 1) {
+			loop_stack[cnt].iters = terminates ? *iters : (struct bpf_loop_iters){};
+			loop_stack[cnt].terminates = terminates;
+			loop_stack[cnt].entry_state = entry_state;
+		} else {
+			loop_stack[cnt].iters = (struct bpf_loop_iters){};
+			loop_stack[cnt].terminates = false;
+			loop_stack[cnt].entry_state = NULL;
+		}
+	}
+	frame->loop_stack_cnt = cnt;
+	return 0;
+
+e2big:
+	verbose(env, "Too many nested loops (%d/%d) at %d\n", cnt, num_missing, env->insn_idx);
+	return -E2BIG;
+}
+
+/*
+ * An exit from a loop can cross several nested loops, e.g.:
+ *
+ *   1: for (...):
+ *   2:   for (...):
+ *   3:     for (...):
+ *            if ...:    // before goto the loop stack is [1, 2, 3 <top>]
+ *              goto 1b; // after goto it should become [1]
+ */
+static int loop_stack_pop(struct bpf_verifier_env *env)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_func_state *frame = cur_func(env);
+	int h, i;
+
+	/*
+	 * Walk the loop nest of insn_idx outwards (innermost first) and stop at
+	 * the first loop that is present on the stack - that loop is the common
+	 * ancestor and becomes the new top. If no enclosing loop is on the
+	 * stack, drop all loops.
+	 */
+	for (h = bpf_loop_at_index(env, env->insn_idx); h >= 0; h = aux[h].loop_header) {
+		for (i = frame->loop_stack_cnt - 1; i >= 0; i--) {
+			if (frame->loop_stack[i].loop_id == h) {
+				frame->loop_stack_cnt = i + 1;
+				return 0;
+			}
+		}
+	}
+
+	frame->loop_stack_cnt = 0;
+	return 0;
+}
+
+static int maybe_clamp_scev_regs(struct bpf_verifier_env *env)
+{
+	struct bpf_func_state *frame = cur_func(env);
+	struct loop_stack_entry *entry;
+
+	if (frame->loop_stack_cnt == 0) {
+		verifier_bug(env, "%s: loop stack empty at %d", __FUNCTION__, env->insn_idx);
+		return -EFAULT;
+	}
+
+	entry = &frame->loop_stack[frame->loop_stack_cnt - 1];
+	if (!entry->terminates)
+		return 0;
+
+	return bpf_clamp_scev_regs(env, frame, env->insn_idx, entry->entry_state, &entry->iters);
+}
+
+static int handle_loop_entry_exit(struct bpf_verifier_env *env)
+{
+	struct bpf_verifier_state *cur = env->cur_state;
+	struct bpf_loop_iters iters;
+	bool terminates;
+	int err;
+
+	if (is_next_loop_iteration(env)) {
+		err = maybe_clamp_scev_regs(env);
+		if (err)
+			return err;
+	}
+	if (has_exited_loop(env)) {
+		if (env->log.level & BPF_LOG_LEVEL2)
+			verbose(env, "exiting loop %d\n", bpf_loop_at_index(env, env->prev_insn_idx));
+		err = loop_stack_pop(env);
+		if (err)
+			return err;
+	}
+	if (has_entered_loop(env)) {
+		if (env->log.level & BPF_LOG_LEVEL2)
+			verbose(env, "entering loop %d\n", bpf_loop_at_index(env, env->insn_idx));
+		err = bpf_compute_loop_iters(env, env->cur_state, &iters);
+		if (err < 0)
+			return err;
+		terminates = err == 1;
+		if (terminates) {
+			err = bpf_split_cur_state(env);
+			if (err)
+				return err;
+			err = bpf_widen_scev_regs(env, cur, cur->parent, &iters);
+			if (err < 0)
+				return err;
+		}
+		err = loop_stack_push(env, env->cur_state->parent, terminates, &iters);
+		if (err)
+			return err;
+	}
+	return 0;
+}
+
 static int do_check(struct bpf_verifier_env *env)
 {
 	bool pop_log = !(env->log.level & BPF_LOG_LEVEL2);
@@ -19618,6 +19816,12 @@ static int do_check(struct bpf_verifier_env *env)
 		struct bpf_insn_aux_data *insn_aux;
 		int err;
 
+		if (signal_pending(current))
+			return -EAGAIN;
+
+		if (need_resched())
+			cond_resched();
+
 		/* reset current history entry on each new instruction */
 		env->cur_hist_ent = NULL;
 
@@ -19664,6 +19868,15 @@ static int do_check(struct bpf_verifier_env *env)
 			}
 		}
 
+		/*
+		 * Possibly widen the registers before creating a checkpoint
+		 * in bpf_is_state_visited(). The next loop iteration will
+		 * have a chance to hit this checkpoint and converge.
+		 */
+		err = handle_loop_entry_exit(env);
+		if (err)
+			return err;
+
 		if (bpf_is_prune_point(env, env->insn_idx)) {
 			err = bpf_is_state_visited(env, env->insn_idx);
 			if (err < 0)
@@ -19689,12 +19902,6 @@ static int do_check(struct bpf_verifier_env *env)
 				return err;
 		}
 
-		if (signal_pending(current))
-			return -EAGAIN;
-
-		if (need_resched())
-			cond_resched();
-
 		if (env->log.level & BPF_LOG_LEVEL2 && do_print_state) {
 			verbose(env, "\nfrom %d to %d%s:",
 				env->prev_insn_idx, env->insn_idx,
@@ -19704,6 +19911,13 @@ static int do_check(struct bpf_verifier_env *env)
 			do_print_state = false;
 		}
 
+		if (bpf_loop_at_index(env, env->insn_idx) >= 0 &&
+		    cur_func(env)->loop_stack_cnt == 0) {
+			verifier_bug(env, "loop stack empty at %d, while inside the loop %d\n",
+				     env->insn_idx, bpf_loop_at_index(env, env->insn_idx));
+			return -EFAULT;
+		}
+
 		if (env->log.level & BPF_LOG_LEVEL) {
 			if (verifier_state_scratched(env))
 				print_insn_state(env, state, state->curframe);
diff --git a/tools/testing/selftests/bpf/progs/verifier_gotox.c b/tools/testing/selftests/bpf/progs/verifier_gotox.c
index f5a9878c7b8d..03a3bee21236 100644
--- a/tools/testing/selftests/bpf/progs/verifier_gotox.c
+++ b/tools/testing/selftests/bpf/progs/verifier_gotox.c
@@ -509,7 +509,7 @@ __naked void too_many_gotox_edges(void)
 
 SEC("socket")
 __description("gotox-edges-at-limit")
-__success __retval(0)
+__failure __msg("Too many nested loops")
 __naked void gotox_edges_at_limit(void)
 {
 	asm volatile (
@@ -517,6 +517,10 @@ __naked void gotox_edges_at_limit(void)
 		 * 1000 gotox * 1000 targets = 1,000,000 CFG edges. At run
 		 * time, each block loads the next table entry and jumps to
 		 * the following block, so the program terminates.
+		 *
+		 * Every block can jump to every later block, so the CFG is
+		 * a nest of ~1000 irreducible loops, deeper than the
+		 * verifier's loop stack; the program is rejected.
 		 */
 		GOTOX_TABLE_BEGIN(1000)
 		".rept 1000;"

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (24 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:46   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 27/36] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
                   ` (9 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

For loops like this:

  int arr[10];
  for (int i = 0; i < 10; i++)
    arr[i] = i;
  use(arr[i]);

SCEV-based loop widening converts `i` from being enumerated 10 times
as a constant value 0, 1, ..., 9 to being enumerated once as a range
[0..9]. From the point of view of the main verification pass, this
replaces constant-offset stack access with varying-offset
stack access.

The verifier tracks stack values accessed through varying offsets much
less precisely than those accessed via constant offsets.
For example, the read at use(arr[i]) would yield STACK_MISC.
Precision might be lost for the following operations:
- spills (BPF_ST/STX to the stack): stored value can no longer be
  recovered when read back, e.g. after the loop;
- fills (BPF_LDX from the stack): a varying-offset read of a spilled
  pointer yields a SCALAR_VALUE, so the pointer is lost.
- calls that construct objects on the stack (dynptr, iter, irq_flag,
  res_spin_lock): the stack argument is required to have a
  constant offset.

This commit adds logic to avoid widening registers used by
such instructions:
- During the main SCEV computation stage, collect_store_base_regs():
  - inspects instructions accessing the stack (spill, fill, call to a
    function listed in bpf_needs_fixed_stack_off), and
  - if there is a SCEV expression for a stack base address register,
    records which registers the expression depends on in
    the bpf_loop->store_base_reg field.
- `store_base_reg`s accumulated for nested loops are propagated
  to outer loops as well.
- During the loop-widening stage, bpf_widen_scev_regs() refuses to widen a
  loop when some of the registers the loop changes appear in
  store_base_regs.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 include/linux/bpf_verifier.h | 10 ++++++
 kernel/bpf/scev.c            | 84 +++++++++++++++++++++++++++++++++++++++++---
 kernel/bpf/verifier.c        | 40 +++++++++++++++++++++
 3 files changed, 129 insertions(+), 5 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c0eea472949d..a5fa2b5cc0c0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -692,6 +692,9 @@ struct bpf_loop_exit {
 	int to; /* instruction outside the loop */
 };
 
+/* SCEV/widening register space: r0..r10 plus every stack slot of a frame. */
+#define BPF_SCEV_REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+
 struct bpf_loop {
 	struct bpf_backedge backedges[MAX_BACKEDGES];
 	/* edges exiting from this loop, includes edges from nested loops */
@@ -701,6 +704,12 @@ struct bpf_loop {
 	bool irreducible;
 	bool backedges_overflow;
 	bool exits_overflow;
+	/*
+	 * Loop-entry registers used to compute stack addresses for loads,
+	 * stores and calls requiring fixed stack offsets, in this and nested
+	 * loops.
+	 */
+	unsigned long store_base_regs[BITS_TO_LONGS(BPF_SCEV_REGS_NUM)];
 };
 
 struct bpf_insn_aux_data {
@@ -1835,6 +1844,7 @@ int bpf_stack_liveness_init(struct bpf_verifier_env *env);
 void bpf_stack_liveness_free(struct bpf_verifier_env *env);
 int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
 const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx);
+bool bpf_needs_fixed_stack_off(struct bpf_verifier_env *env, int insn_idx);
 bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 spi);
 int bpf_compute_live_registers(struct bpf_verifier_env *env);
 
diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
index 205bf8684c3d..1fa6b25a4f29 100644
--- a/kernel/bpf/scev.c
+++ b/kernel/bpf/scev.c
@@ -8,7 +8,7 @@
 #include <linux/tnum.h>
 #include <linux/overflow.h>
 
-#define REGS_NUM (MAX_BPF_REG + MAX_BPF_STACK_SLOTS)
+#define REGS_NUM BPF_SCEV_REGS_NUM
 #define UNKNOWN_EXPR_ID 0
 #define OPAQUE_EXPR_ID  1
 
@@ -1021,6 +1021,65 @@ static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct
 	}
 }
 
+/* OR the loop-entry registers referenced by expr 'id' into 'mask'. */
+static void or_expr_regs(struct bpf_verifier_env *env, u32 id, unsigned long *mask)
+{
+	struct scev *scev = env->scev;
+	u32 order;
+
+	scev->stack_sz = 0;
+	expr_stack_push(scev, id);
+	while (expr_next(scev, &id, &order)) {
+		if ((order & PRE) && scev->exprs[id].op == REG)
+			__set_bit(scev->exprs[id].params[0], mask);
+	}
+}
+
+/* Mask of argument registers (R1..R5) a call at 'idx' passes by register. */
+static u16 call_params_mask(struct bpf_verifier_env *env, int idx)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	struct bpf_call_summary cs;
+	int n = bpf_get_call_summary(env, insn, &cs) ? cs.arg_slot_cnt : MAX_BPF_FUNC_REG_ARGS;
+
+	return n ? GENMASK(BPF_REG_1 + n - 1, BPF_REG_1) : 0;
+}
+
+/*
+ * For instructions like:
+ * - *(u64 *)(rBase + off) = rX
+ * - rX = *(u64 *)(rBase + off)
+ * - calls that construct objects on stack (e.g. dynptr_from_mem(rBase, ...))
+ * When 'rBase' can be a stack pointer and is derived from some registers Rs
+ * defined at loop entry, record Rs into 'mask'.
+ */
+static void collect_store_base_regs(struct bpf_verifier_env *env,
+				    struct env *cur_env, int idx, unsigned long *mask)
+{
+	struct bpf_insn_aux_data *aux = env->insn_aux_data;
+	struct bpf_insn *insn = &env->prog->insnsi[idx];
+	u8 class = BPF_CLASS(insn->code);
+	u8 size = BPF_SIZE(insn->code);
+	u16 base_regs = 0;
+	u32 r;
+
+	if (size == BPF_W || size == BPF_DW) {
+		if ((class == BPF_STX || class == BPF_ST) && insn->dst_reg != BPF_REG_FP)
+			base_regs |= BIT(insn->dst_reg);
+		else if (class == BPF_LDX && insn->src_reg != BPF_REG_FP)
+			base_regs |= BIT(insn->src_reg);
+	}
+
+	if (class == BPF_JMP && BPF_OP(insn->code) == BPF_CALL &&
+	    bpf_needs_fixed_stack_off(env, idx))
+		base_regs |= call_params_mask(env, idx);
+
+	base_regs &= aux[idx].stack_ptrs;
+	for (r = 0; r < MAX_BPF_REG; r++)
+		if (base_regs & BIT(r))
+			or_expr_regs(env, cur_env->reg2expr[r], mask);
+}
+
 /* Find the topmost loop header containing idx inside cur_header, or -1 if none. */
 static int topmost_nested_loop(struct bpf_verifier_env *env, int idx, int cur_header)
 {
@@ -1099,6 +1158,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 	struct bpf_iarray *succ;
 	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
 	int s, i, err, idx, succ_idx;
+	u32 r;
 
 	if (log_level2)
 		bpf_log(&env->log, "Computing SCEV for loop at %d:\n", cur_header);
@@ -1112,6 +1172,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 		if (!header_env)
 			return -ENOMEM;
 		/* The freshly allocated environment has all expressions unknown. */
+		bitmap_fill(cur_loop->store_base_regs, REGS_NUM);
 		header_env->empty = false;
 		return 0;
 	}
@@ -1146,6 +1207,9 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 			 */
 			if (log_level2)
 				memcpy(old_env, cur_env, sizeof(*old_env));
+			/* Pull the nested loop's stack-store base dependencies up. */
+			for_each_set_bit(r, nested_loop->store_base_regs, BPF_SCEV_REGS_NUM)
+				or_expr_regs(env, cur_env->reg2expr[r], cur_loop->store_base_regs);
 			nested_header_env = find_header_env(scev, idx);
 			forget_non_invariants(scev, cur_env, nested_header_env);
 			if (log_level2)
@@ -1168,6 +1232,7 @@ static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
 			for (;;) {
 				if (log_level2)
 					memcpy(old_env, cur_env, sizeof(*old_env));
+				collect_store_base_regs(env, cur_env, idx, cur_loop->store_base_regs);
 				err = transfer(env, cur_env, idx);
 				if (err)
 					goto out;
@@ -1975,6 +2040,7 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 	for (r = 0; r < REGS_NUM; r++) {
 		struct bpf_reg_state *reg;
 		u32 ra, r_scev, r_expr;
+		bool spill_base = false;
 
 		if (!scev_reg_alive(env, st, r))
 			continue;
@@ -1984,10 +2050,16 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 		/* If SCEV for r is (linear <reg> <slope>) */
 		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
 		    is_widenable_reg_type(reg)) {
-			/* Can't widen if the iteration range overflows */
-			if (base_reg == r &&
-			    !linear_bounds(reg, iters, slope_imm, &bounds))
-				goto cant_widen;
+			if (base_reg == r) {
+				/* Can't widen if the iteration range overflows */
+				if (!linear_bounds(reg, iters, slope_imm, &bounds))
+					goto cant_widen;
+				/* Spills at varying offsets lose precision */
+				if (test_bit(r, loop->store_base_regs)) {
+					spill_base = true;
+					goto cant_widen;
+				}
+			}
 			continue;
 		}
 
@@ -2006,6 +2078,8 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
 			log_reg(env, r);
 			bpf_log(log, ", expr is ");
 			log_expr(env, r_expr);
+			if (spill_base)
+				bpf_log(log, ", requires exact stack-offset tracking");
 			bpf_log(log, "\n");
 		}
 		return 0;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 9f56ffe09eba..327bfc00da5f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14087,6 +14087,46 @@ static bool kfunc_spin_allowed(struct bpf_verifier_env *env, s32 func_id, s16 of
 	return *kfunc.flags & KF_SPINLOCK_SAFE;
 }
 
+/*
+ * True if insn calls a helper/kfunc that requires one of its arguments to
+ * be a stack pointer with a constant offset.
+ */
+bool bpf_needs_fixed_stack_off(struct bpf_verifier_env *env, int insn_idx)
+{
+	const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+	u32 *flags, btf_id;
+
+	if (bpf_helper_call(insn)) {
+		return insn->imm == BPF_FUNC_dynptr_from_mem ||
+		       insn->imm == BPF_FUNC_ringbuf_reserve_dynptr;
+	}
+
+	/* vmlinux kfuncs only */
+	if (!bpf_pseudo_kfunc_call(insn) || insn->off != 0)
+		return false;
+	btf_id = insn->imm;
+
+	flags = btf_kfunc_flags(btf_vmlinux, btf_id, env->prog);
+	if (flags && (*flags & (KF_ITER_NEW | KF_ITER_NEXT | KF_ITER_DESTROY)))
+		return true;
+
+	if (btf_id == special_kfunc_list[KF_bpf_res_spin_lock] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_unlock] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_lock_irqsave] ||
+	    btf_id == special_kfunc_list[KF_bpf_res_spin_unlock_irqrestore] ||
+	    btf_id == special_kfunc_list[KF_bpf_local_irq_save] ||
+	    btf_id == special_kfunc_list[KF_bpf_local_irq_restore])
+		return true;
+
+	if (btf_id == special_kfunc_list[KF_bpf_dynptr_from_skb] ||
+	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_xdp] ||
+	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_skb_meta] ||
+	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_file])
+		return true;
+
+	return false;
+}
+
 static bool is_sync_callback_calling_kfunc(u32 btf_id)
 {
 	return is_bpf_rbtree_add_kfunc(btf_id);

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 27/36] selftests/bpf: __msg_next tag for matching messages on consecutive lines
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (25 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
                   ` (8 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Same as __msg, but expects the match on the line immediately after
the last match, not just any later line.
E.g.:

  __msg_next("a")
  __msg_next("c")

This would match consecutive output "a\n" "c\n",
but would not match output "a\n" "b\n" "c\n".

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/progs/bpf_misc.h |  6 ++++++
 tools/testing/selftests/bpf/test_loader.c    | 10 ++++++++++
 2 files changed, 16 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..d7f706c3bc4a 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -43,6 +43,10 @@
  * __msg_unpriv      Same as __msg but for unprivileged mode.
  * __not_msg_unpriv  Same as __not_msg but for unprivileged mode.
  *
+ * __msg_next        Same as __msg but expects the match to be on the next
+ * __msg_next_unpriv log line compared to the last match, not just some line
+ *                   after the last match.
+ *
  * __stderr          Message expected to be found in bpf stderr stream. The
  *                   same regex rules apply like __msg.
  * __stderr_unpriv   Same as __stderr but for unpriveleged mode.
@@ -145,6 +149,7 @@
 
 #define __msg(msg)		__test_tag("test_expect_msg=" msg)
 #define __not_msg(msg)		__test_tag("test_expect_not_msg=" msg)
+#define __msg_next(msg)		__test_tag("test_expect_msg_next=" msg)
 #define __xlated(msg)		__test_tag("test_expect_xlated=" msg)
 #define __jited(msg)		__test_tag("test_jited=" msg)
 #define __failure		__test_tag("test_expect_failure")
@@ -153,6 +158,7 @@
 #define __skip(reason)		__test_tag("test_skip=" reason)
 #define __msg_unpriv(msg)	__test_tag("test_expect_msg_unpriv=" msg)
 #define __not_msg_unpriv(msg)	__test_tag("test_expect_not_msg_unpriv=" msg)
+#define __msg_next_unpriv(msg)	__test_tag("test_expect_msg_next_unpriv=" msg)
 #define __xlated_unpriv(msg)	__test_tag("test_expect_xlated_unpriv=" msg)
 #define __jited_unpriv(msg)	__test_tag("test_jited_unpriv=" msg)
 #define __failure_unpriv	__test_tag("test_expect_failure_unpriv")
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index 25eeb1c1248b..edeeb244e693 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -508,6 +508,16 @@ static int parse_test_spec(struct test_loader *tester,
 			if (err)
 				goto cleanup;
 			spec->mode_mask |= UNPRIV;
+		} else if ((msg = str_has_pfx(s, "test_expect_msg_next="))) {
+			err = __push_msg(msg, true, false, &spec->priv.expect_msgs);
+			if (err)
+				goto cleanup;
+			spec->mode_mask |= PRIV;
+		} else if ((msg = str_has_pfx(s, "test_expect_msg_next_unpriv="))) {
+			err = __push_msg(msg, true, false, &spec->unpriv.expect_msgs);
+			if (err)
+				goto cleanup;
+			spec->mode_mask |= UNPRIV;
 		} else if ((msg = str_has_pfx(s, "test_jited="))) {
 			if (arch_mask == 0) {
 				PRINT_FAIL("__jited used before __arch_*");

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (26 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 27/36] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:32   ` sashiko-bot
  2026-09-27 20:26   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 29/36] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
                   ` (7 subsequent siblings)
  35 siblings, 2 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Validate changes from:
"bpf: allow subrange relations for PTR_TO_STACK in regsafe()".
Add a test case that checks that a widened stack pointer with range
[-64, -8] is considered a superset of a stack pointer with range
[-56, -8].

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_stack_ptr.c       | 33 ++++++++++++++++++++++
 1 file changed, 33 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
index 3e0bea9819ca..5da8b9a6018c 100644
--- a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
+++ b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
@@ -586,4 +586,37 @@ __naked void stack_check_size_512_with_may_goto(void)
 }
 #endif
 
+/* Verify that old PTR_TO_STACK state is considered a super-set of
+ * new PTR_TO_STACK state when new variable range is a sub-range
+ * of the old range, e.g. old [-72, -16] vs new [-64, -16].
+ */
+SEC("socket")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+/* r7 is widened to [-72, -16] at the loop header (insn 3),
+ * the loop body sees [-64, -8] after 'r7 += 8'
+ */
+__msg("loop header at 3, widening r7 to -72..-16 step 8")
+__msg("R7=fp(smin=smin32=-64,smax=smax32=-8")
+/* back-edge state is clamped to the remaining iterations and pruned
+ * at the header, because [-64, -16] is within the widened [-72, -16]
+ */
+__msg("loop header at 3, clamping r7 to -64..-16 step 8")
+__msg("from 5 to 3: safe")
+__msg("processed 9 insns")
+__naked void stack_ptr_subrange_in_loop(void)
+{
+	asm volatile ("					\
+	r7 = r10;					\
+	r7 += -72;					\
+	r6 = 0;						\
+	r7 += 8;					\
+	r6 += 1;					\
+	if r6 < 8 goto -3;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
 char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 29/36] selftests/bpf: tests for may_write stack-liveness tracking
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (27 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 30/36] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
                   ` (6 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Branches to cover:
  - per-insn writes (record_stack_access_off):
      * off_cnt==1, fully covered slots   -> def + may_def (equal)
      * partial coverage                  -> may_def only (no def)
      * off_cnt 2-4, precise offsets      -> may_def only, per offset
  - whole-frame writes:
      * off_cnt==0 (offset lost)          -> may_def SPIS_ALL
      * ARG_IMPRECISE (frame lost)        -> may_def SPIS_ALL per frame
  - merge_instances():
      * both passes touch a frame         -> may_def union
      * one pass leaves a frame untouched -> may_def preserved
  - merge_may_write(): a callee's writes into ancestor frames are
    summarized as may_def at the caller's call instruction.

Tests exercising them:
  - may_write_two_precise_offsets  - off_cnt 2-4 precise multi-offset
  - imprecise_frame_write          - cross-frame ARG_IMPRECISE write,
                                     plus its caller-side summarization
  - merge_preserves_may_write      - map-value vs stack pointer at two
                                     callsites: union is preserved when
                                     one pass does not touch the frame
  - must_write_not_same_slot       - off_cnt==0 whole-frame may_def
  - two_byte_write_no_kill         - partial write, may_def without def
  - kfunc_iter_stack_liveness      - full-slot def == may_def
  - caller_stack_write,
    conditional_stx_in_subprog,
    shared_instance_must_write_overwrite
                                   - callee writes summarized as may_def
                                     at the call site.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_live_stack.c      | 147 ++++++++++++++++++++-
 1 file changed, 144 insertions(+), 3 deletions(-)

diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 6b295b30936d..41df0c393981 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -83,6 +83,31 @@ __naked void must_write_not_same_slot(void)
 	: __clobber_all);
 }
 
+SEC("socket")
+__log_level(2)
+__msg("stack use/def subprog#0 may_write_two_precise_offsets (d0,cs0):")
+/*
+ * r1 is fp-8 or fp-16, with both offsets tracked precisely (off_cnt == 2).
+ * A write through it cannot set def, but marks both candidate slots as may_def.
+ */
+__msg("6: (7b) *(u64 *)(r1 +0) = r0         ; may_def: fp0-8 fp0-16")
+__naked void may_write_two_precise_offsets(void)
+{
+	asm volatile (
+	"call %[bpf_get_prandom_u32];"
+	"r1 = r10;"
+	"if r0 > 42 goto 1f;"
+	"r1 += -8;"
+	"goto 2f;"
+"1:"
+	"r1 += -16;"
+"2:"
+	"*(u64 *)(r1 + 0) = r0;"
+	"exit;"
+	:: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
 SEC("socket")
 __log_level(2)
 __msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
@@ -114,7 +139,7 @@ __log_level(2)
 __msg("stack use/def subprog#0 caller_stack_write (d0,cs0):")
 __msg("2: (85) call pc+1                    ; may_def: fp0-8")
 __msg("stack use/def subprog#1 write_first_param (d1,cs2):")
-__msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8")
+__msg("4: (7a) *(u64 *)(r1 +0) = 7          ; def: fp0-8 may_def: fp0-8")
 __naked void caller_stack_write(void)
 {
 	asm volatile (
@@ -134,6 +159,49 @@ static __used __naked void write_first_param(void)
 	::: __clobber_all);
 }
 
+/*
+ * Cross-frame imprecise write: imprecise_frame_writer() receives a pointer
+ * into the caller's frame (frame 0) but conditionally replaces it with a
+ * pointer into its own frame (frame 1). At the store the two are joined into
+ * ARG_IMPRECISE (frame is unknown, mask = {0,1}), so the offset is dropped and
+ * the whole of both candidate frames is conservatively marked as may_def.
+ */
+SEC("socket")
+__log_level(2)
+/* The callee's write into the caller frame is summarized at the call site. */
+__msg("stack use/def subprog#0 imprecise_frame_write (d0,cs0):")
+__msg("2: (85) call pc+2                    ; may_def: fp0-8..-{{(512|2048)}}")
+__msg("stack use/def subprog#1 imprecise_frame_writer (d1,cs2):")
+__msg("12: (7b) *(u64 *)(r1 +0) = r2         ; may_def: fp1-8..-{{(512|2048)}} fp0-8..-{{(512|2048)}}")
+__naked void imprecise_frame_write(void)
+{
+	asm volatile (
+	"r1 = r10;"
+	"r1 += -8;"
+	"call imprecise_frame_writer;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked void imprecise_frame_writer(void)
+{
+	asm volatile (
+	"r6 = r1;"			/* save arg: pointer into frame 0 */
+	"call %[bpf_get_prandom_u32];"
+	"r1 = r6;"			/* r1 = frame0-8 (arg) */
+	"if r0 == 0 goto 1f;"
+	"r1 = r10;"
+	"r1 += -8;"			/* r1 = frame1-8 (own frame) */
+"1:"
+	"r2 = 0;"
+	"*(u64 *)(r1 + 0) = r2;"	/* join frame0-8 | frame1-8 -> imprecise */
+	"r0 = 0;"
+	"exit;"
+	:: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
 SEC("socket")
 __log_level(2)
 __msg("stack use/def subprog#0 caller_stack_read (d0,cs0):")
@@ -1011,7 +1079,7 @@ __naked void four_byte_read_upper_half(void)
  */
 SEC("socket")
 __log_level(2)
-__msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8")
+__msg("0: (7a) *(u64 *)(r10 -8) = 0         ; def: fp0-8 may_def: fp0-8")
 /* 2-byte write only partially covers the upper half: may_def, but no def. */
 __msg("1: (6a) *(u16 *)(r10 -4) = 0         ; may_def: fp0-4h")
 __msg("2: (61) r0 = *(u32 *)(r10 -4)        ; use: fp0-4h")
@@ -1364,7 +1432,7 @@ __log_level(2)
  * fp-8 live at call: callee conditionally writes it, so the slot is not killed
  * (no def), but the conditional write surfaces as may_def at the call site.
  */
-__msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8")
+__msg("1: (7b) *(u64 *)(r10 -8) = r1        ; def: fp0-8 may_def: fp0-8")
 __msg("4: (85) call pc+2                    ; may_def: fp0-8")
 __msg("5: (79) r0 = *(u64 *)(r10 -8)        ; use: fp0-8")
 __naked void conditional_stx_in_subprog(void)
@@ -2393,6 +2461,20 @@ __flag(BPF_F_TEST_STATE_FREQ)
 __msg("subprog#2 write_first_read_second:")
 __msg("17: (7a) *(u64 *)(r1 +0) = 42{{$}}")
 __msg("18: (79) r0 = *(u64 *)(r2 +0) // r1=fp0-8 r2=fp0-16{{$}}")
+/*
+ * The callee's write into the caller frame is summarized as may_def at the
+ * call sites, transitively across both frames (depth 2 -> 1 -> 0). The two
+ * call sites differ because the shared write_first_read_second instance
+ * accumulates the cross-pass union ({fp-8, fp-16}) at different points
+ * relative to each forwarding_rw's summarization.
+ */
+__msg("stack use/def subprog#0 shared_instance_must_write_overwrite (d0,cs0):")
+__msg("7: (85) call pc+7                    ; use: fp0-8 fp0-16 may_def: fp0-8 fp0-16")
+__msg("12: (85) call pc+2                    ; use: fp0-8 may_def: fp0-16")
+__msg("stack use/def subprog#1 forwarding_rw (d1,cs7):")
+__msg("15: (85) call pc+1                    ; use: fp0-8 fp0-16 may_def: fp0-8 fp0-16")
+__msg("stack use/def subprog#1 forwarding_rw (d1,cs12):")
+__msg("15: (85) call pc+1                    ; use: fp0-8 may_def: fp0-16")
 __msg("stack use/def subprog#2 write_first_read_second (d2,cs15):")
 /*
  * Shared across two callsites with swapped args (r1 is fp-8 on one pass,
@@ -2441,6 +2523,65 @@ static __used __naked void write_first_read_second(void)
 	::: __clobber_all);
 }
 
+/*
+ * merge_instances() must preserve may_write when one pass doesn't touch a frame.
+ * mvs_leaf is reached through a single callsite (mvs_mid) from two outer chains,
+ * so it is one shared (depth,callsite) instance analyzed twice and merged:
+ * - chain A passes r2 = map value, mvs_leaf's doesn't touch stack;
+ * - chain B passes r2 = caller stack slot fp-8.
+ * The merge must keep the stack pass's may_write even though the map pass left
+ * frame 0 untouched.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("stack use/def subprog#2 mvs_leaf (d2,cs19):")
+__msg("21: (7a) *(u64 *)(r2 +0) = 42         ; may_def: fp0-8")
+__naked void merge_preserves_may_write(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 24) = r1;"		/* map key */
+	"r1 = %[map] ll;"
+	"r2 = r10;"
+	"r2 += -24;"
+	"call %[bpf_map_lookup_elem];"
+	"if r0 == 0 goto 1f;"
+	/* chain A: r2 = map value (ARG_NONE, not stack) */
+	"r1 = r10;"
+	"r1 += -16;"				/* unused fp anchor */
+	"r2 = r0;"
+	"call mvs_mid;"
+	/* chain B: r2 = caller stack slot fp-8 */
+	"r1 = r10;"
+	"r1 += -16;"
+	"r2 = r10;"
+	"r2 += -8;"
+	"call mvs_mid;"
+"1:"
+	"r0 = 0;"
+	"exit;"
+	: : __imm(bpf_map_lookup_elem), __imm_addr(map)
+	: __clobber_all);
+}
+
+static __used __naked void mvs_mid(void)
+{
+	asm volatile (
+	"call mvs_leaf;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked void mvs_leaf(void)
+{
+	asm volatile (
+	"*(u64 *)(r2 + 0) = 42;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
 /*
  * Shared must_write when (callsite, depth) instance is reused.
  * Main calls fwd_to_stale_wr at two sites. fwd_to_stale_wr calls

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 30/36] selftests/bpf: tests for may_def marks of atomic RMW operations
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (28 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 29/36] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 31/36] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
                   ` (5 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Liveness analysis should produce may_def marks for read-modify-write
atomic operations:
- lock *(u64 *)(r10 -8) += r1
- r1 = atomic64_fetch_add((u64 *)(r10 -8), r1)
- r1 = atomic64_xchg((u64 *)(r10 -8), r1)
- r0 = atomic64_cmpxchg((u64 *)(r10 -8), r0, r1)

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_live_stack.c      | 26 ++++++++++++++++++++++
 1 file changed, 26 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index 41df0c393981..496229a9b284 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -2666,6 +2666,32 @@ __naked void load_acquire_dont_clear_dst(void)
 
 #endif /* CAN_USE_LOAD_ACQ_STORE_REL */
 
+SEC("socket")
+__log_level(2)
+__success
+__msg("lock *(u64 *)(r10 -8) += r1{{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r1 = atomic64_fetch_add((u64 *)(r10 -8), r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r1 = atomic64_xchg((u64 *)(r10 -8), r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__msg("r0 = atomic64_cmpxchg((u64 *)(r10 -8), r0, r1){{.*}}use: fp0-8{{.*}}may_def: fp0-8")
+__naked void atomic_rmw(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 8) = r1;"
+	".8byte %[atomic_add];"
+	".8byte %[atomic_fetch_add];"
+	".8byte %[atomic_xchg];"
+	"r0 = 0;"
+	".8byte %[atomic_cmpxchg];"
+	"exit;"
+	:
+	: __imm_insn(atomic_add, BPF_ATOMIC_OP(BPF_DW, BPF_ADD, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_fetch_add, BPF_ATOMIC_OP(BPF_DW, BPF_ADD | BPF_FETCH, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_xchg, BPF_ATOMIC_OP(BPF_DW, BPF_XCHG, BPF_REG_10, BPF_REG_1, -8)),
+	  __imm_insn(atomic_cmpxchg, BPF_ATOMIC_OP(BPF_DW, BPF_CMPXCHG, BPF_REG_10, BPF_REG_1, -8))
+	: __clobber_all);
+}
+
 SEC("socket")
 __success
 __naked void imprecise_fill_loses_cross_frame(void)

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 31/36] selftests/bpf: tests for register base/step arithmetic
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (29 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 30/36] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
                   ` (4 subsequent siblings)
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

- step_mul_non_pow2: Retain step 3 after multiplying a bounded scalar
  by a non-power-of-two constant.
- step_lsh: Record step 4 after shifting a bounded scalar left by two.
- step_add_const_base: Shift the base to 1 while preserving step 4
  after adding a constant.
- step_neg_value_range: Preserve the signed range of {-3, 0, 3, 6}
  when intersecting with a non-power-of-two step across zero.
- step_neg_value_sound: Reject a reachable invalid access that would
  be hidden by incorrectly narrowing that range's upper bound to 4.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../selftests/bpf/progs/verifier_bounds_step.c     | 124 +++++++++++++++++++++
 2 files changed, 126 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 8a6d341b754a..e1d559ba1271 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -19,6 +19,7 @@
 #include "verifier_basic_stack.skel.h"
 #include "verifier_bitfield_write.skel.h"
 #include "verifier_bounds.skel.h"
+#include "verifier_bounds_step.skel.h"
 #include "verifier_bounds_deduction.skel.h"
 #include "verifier_bounds_deduction_non_const.skel.h"
 #include "verifier_bounds_mix_sign_unsign.skel.h"
@@ -203,6 +204,7 @@ void test_verifier_arena_globals2(void)       { RUN(verifier_arena_globals2); }
 void test_verifier_basic_stack(void)          { RUN(verifier_basic_stack); }
 void test_verifier_bitfield_write(void)       { RUN(verifier_bitfield_write); }
 void test_verifier_bounds(void)               { RUN(verifier_bounds); }
+void test_verifier_bounds_step(void)          { RUN(verifier_bounds_step); }
 void test_verifier_bounds_deduction(void)     { RUN(verifier_bounds_deduction); }
 void test_verifier_bounds_deduction_non_const(void)     { RUN(verifier_bounds_deduction_non_const); }
 void test_verifier_bounds_mix_sign_unsign(void) { RUN(verifier_bounds_mix_sign_unsign); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
new file mode 100644
index 000000000000..a8358b625ce7
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
@@ -0,0 +1,124 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+/*
+ * Tests for the linear "base + step * k" description tracked per scalar
+ * register.
+ *
+ * The base/step computed by scalar arithmetic is observed directly in the
+ * level-2 register dump, printed as "step=<base>+<step>" when the register is
+ * not the trivial base=0, step=1. Range inference that depends on the line
+ * (e.g. a sign-crossing range with a non-power-of-2 step) is checked via the
+ * resulting smin/smax, which feed ordinary signed branch decisions.
+ *
+ * Note: eq/neq branch checks do not consult base/step, so impossible-value
+ * pruning is intentionally not relied upon here; the register dump is the
+ * reliable signal.
+ */
+
+/* a &= 0xff; a *= 3  =>  multiples of 3, step=0+3 (not representable as tnum) */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 *= 3 {{.*}}step=0+3)")
+__naked void step_mul_non_pow2(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 0xff;					\
+	r0 *= 3;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/* a &= 0xff; a <<= 2  =>  step=0+4 */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 <<= 2 {{.*}}step=0+4)")
+__naked void step_lsh(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 0xff;					\
+	r0 <<= 2;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/* a &= 0xff; a *= 4; a += 1  =>  base shifts to 1, step=1+4 */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += 1 {{.*}}step=1+4)")
+__naked void step_add_const_base(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 0xff;					\
+	r0 *= 4;					\
+	r0 += 1;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Sign-crossing range with a non-power-of-2 step. After "*= 3; += -3" the value
+ * set is {-3, 0, 3, 6}. The line description is tracked in signed space, so the
+ * intersection keeps smax=6. A u64-modular intersection would mis-place the
+ * line for the negative values and wrongly narrow smax to 4.
+ */
+SEC("socket")
+__success __log_level(2)
+__msg("r0 += -3 {{.*}}smax=smax32=6)")
+__naked void step_neg_value_range(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r0 *= 3;					\
+	r0 += -3;					\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Same {-3, 0, 3, 6} value set: 6 is reachable, so "if r0 s> 4" can be taken
+ * and the illegal scalar dereference behind it must be rejected. Guards against
+ * the unsound narrowing (smax=4) that would prune the branch and accept the
+ * program. This rides on the signed comparison, which uses smin/smax.
+ */
+SEC("socket")
+__failure __msg("invalid mem access 'scalar'")
+__naked void step_neg_value_sound(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r0 *= 3;					\
+	r0 += -3;					\
+	if r0 s> 4 goto l_bad_%=;			\
+	r0 = 0;						\
+	exit;						\
+l_bad_%=:						\
+	r1 = *(u8 *)(r0 + 0);				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (30 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 31/36] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:27   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
                   ` (3 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Tests for the base/step reasoning in states.c:range_within():
- step_prune_hit_multiple: prune a step-4 range contained in a cached
  step-2 range when the scalar is used to address the stack.
- step_prune_miss_non_multiple: do not let a cached step-2 range hide a
  step-3 path reaching division by zero.
- step_prune_hit_const_on_line: prune constant 6 against a cached
  base-0, step-3 range.
- step_prune_miss_const_off_line: do not prune constant 7 against that
  range even though its bounds and tnum admit the value.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 .../selftests/bpf/progs/verifier_bounds_step.c     | 183 +++++++++++++++++++++
 1 file changed, 183 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
index a8358b625ce7..cb55fc216bc8 100644
--- a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
+++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
@@ -121,4 +121,187 @@ l_bad_%=:						\
 	: __clobber_all);
 }
 
+struct step_val {
+	__u8 data[1024];
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARRAY);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, struct step_val);
+} step_map SEC(".maps");
+
+/* Old register [4..130, step 2] should prune cur register [8..64, step 4]. */
+SEC("socket")
+__success __log_level(2)
+__msg("7: (27) r6 *= 4                       ; R6=scalar({{.*}}umin32=8,{{.*}}umax32=68,{{.*}},step=0+4)")
+__msg("10: (27) r7 *= 2                      ; R7=scalar({{.*}}umin32=4,{{.*}}umax32=130,{{.*}},step=0+2)")
+__msg("11: (25) if r0 > 0x2a goto pc+1")
+__msg("from 11 to 13: safe")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void step_prune_hit_multiple(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r6 &= 0x0f;					\
+	r6 += 2;					\
+	r6 *= 4;					\
+	r7 &= 0x3f;					\
+	r7 += 2;					\
+	r7 *= 2;					\
+	if r0 > 42 goto 1f;	/* can't predict */	\
+	r6 = r7;		/* step=2 explored first, step=4 explored next */ \
+1:	r0 = r10;					\
+	r6 = -r6;					\
+	r0 += r6;					\
+	*(u8 *)(r0 + 0) = 7;	/* force r6 precise */	\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
+/* Old register [0..126, step 2] should not prune cur register [0..45, step 3]. */
+SEC("socket")
+__failure __log_level(2)
+__msg("6: (27) r6 *= 3                       ; R6=scalar({{.*}}smin32=0,{{.*}}umax32=45,{{.*}},step=0+3)")
+__msg("8: (27) r7 *= 2                       ; R7=scalar({{.*}}smin32=0,{{.*}}umax32=126,{{.*}},step=0+2)")
+__msg("9: (25) if r0 > 0x2a goto pc+1")
+__msg("11: (15) if r6 == 0x3 goto pc+2")
+__msg("11: R6=scalar({{.*}},step=0+2)")
+__msg("13: (95) exit")
+__msg("from 9 to 11: {{.*}} R6=scalar({{.*}},step=0+3)")
+__msg("from 11 to 14")
+__msg("div by zero")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void step_prune_miss_non_multiple(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	call %[bpf_get_prandom_u32];			\
+	r6 &= 0x0f;					\
+	r6 *= 3;					\
+	r7 &= 0x3f;					\
+	r7 *= 2;					\
+	if r0 > 42 goto 1f;	/* can't predict */	\
+	r6 = r7;		/* step=2 explored first, step=3 explored next */ \
+1:							\
+	if r6 == 3 goto 2f;	/* false if step=2, should not prune step=3 */ \
+	r0 = 0;						\
+	exit;						\
+2:							\
+	r0 /= 0;		/* trap */		\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Constant current register lying on the cached line: cached is {0,3,6,...}
+ * (step 3, base 0), current is the constant 6. range_within() takes the
+ * constant branch: imod(6, 3) == base 0, so cur is on the line and the
+ * (precise) state is pruned -> "safe".
+ */
+SEC("socket")
+__success __log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("14: (27) r1 *= 3")		/* cached path: step 3 line */
+__msg("16: (b7) r1 = 6")		/* current path: const 6, on the line */
+__msg("17: safe")			/* pruned at join: imod(6, 3) == 0 */
+__naked void step_prune_hit_const_on_line(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	r1 = 0;						\
+	*(u32*)(r10 - 4) = r1;				\
+	r2 = r10;					\
+	r2 += -4;					\
+	r1 = %[step_map] ll;				\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto l_out_%=;			\
+	r7 = r0;					\
+	r1 = r6;					\
+	r1 &= 0xff;					\
+	if r6 > 0 goto l_cur_%=;			\
+	r1 *= 3;			/* old: step 3 */	\
+	goto l_join_%=;					\
+l_cur_%=:						\
+	r1 = 6;				/* cur: const on line */	\
+l_join_%=:						\
+	r0 = r7;					\
+	r0 += r1;			/* r1 forced precise */	\
+	r2 = *(u8 *)(r0 + 0);				\
+l_out_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
+/*
+ * Constant current register NOT on the cached line: cached is {0,3,6,...}
+ * (step 3, base 0), current is the constant 7. imod(7, 3) == 1 != base 0,
+ * so range_within() fails and the join is traversed again. A power-of-two
+ * step is avoided on purpose: with step 3 the tnum is loose enough to admit
+ * 7, so imod() is
+ * the sole check that rejects it. No failure shape is possible here: 7 is
+ * within the line's bounds and tnum, and the cached line path already
+ * verifies the whole outro, so an eager prune could not miss an error.
+ */
+SEC("socket")
+__success __log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("14: (27) r1 *= 3")		/* cached path: step 3 line */
+__msg("16: (b7) r1 = 7")		/* current path: const 7, off the line */
+/* not pruned: current continues past the join with the constant offset 7 */
+__msg("19: R0=map_value({{.*}}imm=7) R1=7")
+__naked void step_prune_miss_const_off_line(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = r0;					\
+	r1 = 0;						\
+	*(u32*)(r10 - 4) = r1;				\
+	r2 = r10;					\
+	r2 += -4;					\
+	r1 = %[step_map] ll;				\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto l_out_%=;			\
+	r7 = r0;					\
+	r1 = r6;					\
+	r1 &= 0xff;					\
+	if r6 > 0 goto l_cur_%=;			\
+	r1 *= 3;			/* old: step 3 */	\
+	goto l_join_%=;					\
+l_cur_%=:						\
+	r1 = 7;				/* cur: const off line */	\
+l_join_%=:						\
+	r0 = r7;					\
+	r0 += r1;			/* r1 forced precise */	\
+	r2 = *(u8 *)(r0 + 0);				\
+l_out_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32),
+	  __imm(bpf_map_lookup_elem),
+	  __imm_addr(step_map)
+	: __clobber_all);
+}
+
 char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (31 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 34/36] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
                   ` (2 subsequent siblings)
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Verifier tests for accessing arrays at a varying offset through a
PTR_TO_BTF_ID pointer, e.g. a[i].b where 'i' is a register with a
varying value:
- bpf_obj_new() is used to obtain a PTR_TO_BTF_ID to a program-defined
  struct
- bpf_get_current_task_btf() is used to obtain a PTR_TO_BTF_ID from
  the kernel BTF.

The cases cover:
- in-bounds access to a field of an array of structs;
- out-of-bounds access where the maximal offset runs past the array;
- a scalar (long) array with a matching stride;
- a byte array (step 1);
- a 4-byte read whose maximal offset spills past the array end;
- an unaligned step that is not a whole number of elements;
- both dimensions of a 2D array;
- an array of pointers within a program-allocated object, where reading a
  pointer member yields a SCALAR_VALUE but the varying offset is still
  validated against the array bounds (in-bounds and out-of-bounds);
- a varying access into an array embedded in a kernel BTF type
  (task_struct's comm[]), covering the non-allocated, trusted-pointer path
  (in-bounds and out-of-bounds);
- an array nested below a struct member at a non-zero offset, exercising
  the offset telescoping done while walking across a struct boundary
  (in-bounds and out-of-bounds);
- a trailing flexible array member, whose unbounded tail makes a varying
  access have no upper bound to exceed. A program-defined struct cannot be
  used here because bpf_obj_new() forbids flexible array members, so a
  kernel type (struct vring_used, with its 'ring[]' tail) is used
  (in-bounds, and an unaligned step that is not a whole number of
   elements).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../bpf/progs/verifier_btf_array_access.c          | 506 +++++++++++++++++++++
 2 files changed, 508 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index e1d559ba1271..9f188ff9be89 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -26,6 +26,7 @@
 #include "verifier_bpf_get_stack.skel.h"
 #include "verifier_bpf_trap.skel.h"
 #include "verifier_bswap.skel.h"
+#include "verifier_btf_array_access.skel.h"
 #include "verifier_btf_ctx_access.skel.h"
 #include "verifier_btf_flex_array.skel.h"
 #include "verifier_btf_unreliable_prog.skel.h"
@@ -211,6 +212,7 @@ void test_verifier_bounds_mix_sign_unsign(void) { RUN(verifier_bounds_mix_sign_u
 void test_verifier_bpf_get_stack(void)        { RUN(verifier_bpf_get_stack); }
 void test_verifier_bpf_trap(void)             { RUN(verifier_bpf_trap); }
 void test_verifier_bswap(void)                { RUN(verifier_bswap); }
+void test_verifier_btf_array_access(void)     { RUN(verifier_btf_array_access); }
 void test_verifier_btf_ctx_access(void)       { RUN(verifier_btf_ctx_access); }
 void test_verifier_btf_flex_array(void)       { RUN(verifier_btf_flex_array); }
 void test_verifier_btf_unreliable_prog(void)  { RUN(verifier_btf_unreliable_prog); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
new file mode 100644
index 000000000000..ab6b74457365
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
@@ -0,0 +1,506 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_core_read.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+
+/*
+ * Variable-offset access into arrays within a BTF access chain, e.g. a[i].b,
+ * where 'i' is a register with a varying offset. The register's step has to
+ * match the stride of one of the arrays crossed while walking the access
+ * chain, and the maximal possible offset has to stay within that array.
+ *
+ * A program-allocated object (bpf_obj_new) is used as the source of a
+ * PTR_TO_BTF_ID with a shape we fully control.
+ */
+struct inner {
+	int a;
+	int b;
+};
+
+struct outer {
+	struct inner arr[8];	/* off 0,   stride 8, size 64 */
+	char bytes[32];		/* off 64,  stride 1, size 32 */
+	long longs[8];		/* off 96,  stride 8, size 64 */
+	int grid[4][4];		/* off 160, 16 ints,  size 64 */
+};
+
+/* arr[i].b, i in [0, 7], 4-byte read: stays within arr. */
+SEC("syscall")
+__success
+int arr_field_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* arr[i].b, i in [0, 15]: max offset runs past the end of arr. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int arr_field_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* longs[i], i in [0, 7], 8-byte read: scalar array with stride 8. */
+SEC("syscall")
+__success
+int long_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 96);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * bytes[i], i in [0, 31], read as 4 bytes: min_off is valid within the array,
+ * but the 4-byte read at max_off runs past the array end. Exercises the
+ * size-aware bound in btf_struct_access().
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int byte_array_size_spanning(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 64);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* bytes[i], i in [0, 31], 1-byte read: step 1 matches the char array stride. */
+SEC("syscall")
+__success
+int byte_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 64);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * arr[] accessed with step 4 (half of sizeof(struct inner)): the step is not
+ * a whole number of elements, so no crossed array has a matching stride.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct outer")
+int arr_unaligned_step(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * grid[i][0], i in [0, 3]: steps along the outer dimension of a 2D array
+ * (stride 16). __btf_resolve_size() linearizes the array; the step is a
+ * multiple of the innermost element size (4).
+ */
+SEC("syscall")
+__success
+int grid_outer_dim(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 3;						\
+	r2 *= 16;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 160);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * grid[0][i], i in [0, 3]: steps along the inner dimension of a 2D array
+ * (stride 4, the innermost element size).
+ */
+SEC("syscall")
+__success
+int grid_inner_dim(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct outer *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 3;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 160);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * An array of pointers within a program-allocated object. Reading a pointer
+ * member of a local object yields a SCALAR_VALUE, but the varying offset still
+ * has to be validated against the array bounds.
+ */
+struct with_ptrs {
+	struct inner *parr[8];	/* off 0, stride 8 (sizeof ptr), size 64 */
+};
+
+/* parr[i], i in [0, 7], 8-byte read: stays within parr. */
+SEC("syscall")
+__success
+int ptr_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct with_ptrs *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * parr[i], i in [0, 15]: max offset runs past the end of parr. Without the
+ * variable-offset check on the WALK_PTR path this would be wrongly accepted.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct with_ptrs")
+int ptr_array_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct with_ptrs *o;
+	long val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u64 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * The cases above all use bpf_obj_new() (PTR_TO_BTF_ID | MEM_ALLOC). Exercise
+ * the trusted kernel-BTF path too, using a task_struct from
+ * bpf_get_current_task_btf() and its embedded char comm[] array. The comm
+ * offset is folded into the pointer as a constant before the varying index.
+ */
+
+/* task->comm[i], i in [0, 15], 1-byte read: stays within comm. */
+SEC("syscall")
+__success
+int kernel_btf_array_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct task_struct *task;
+	int val = 0;
+
+	task = bpf_get_current_task_btf();
+	asm volatile ("						\
+	r1 = %[task];						\
+	r1 += %[off];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [task] "r"(task), [i] "r"(i),
+	  [off] "i"(offsetof(struct task_struct, comm))
+	: "r1", "r2");
+	return val;
+}
+
+/* task->comm[i], i in [0, 31]: max offset runs past comm. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct task_struct")
+int kernel_btf_array_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct task_struct *task;
+	int val = 0;
+
+	task = bpf_get_current_task_btf();
+	asm volatile ("						\
+	r1 = %[task];						\
+	r1 += %[off];						\
+	r2 = %[i];						\
+	r2 &= 31;						\
+	r1 += r2;						\
+	%[val] = *(u8 *)(r1 + 0);				\
+"	: [val] "=r"(val)
+	: [task] "r"(task), [i] "r"(i),
+	  [off] "i"(offsetof(struct task_struct, comm))
+	: "r1", "r2");
+	return val;
+}
+
+/*
+ * Array nested below a struct member at a non-zero offset. The walk descends
+ * through 'm' (resetting the running offset) before reaching 'arr', exercising
+ * the array_start = min_off - arrays[i].off telescoping across a WALK_STRUCT
+ * dive.
+ */
+struct mid {
+	struct inner arr[8];	/* stride 8, size 64 */
+};
+
+struct nest {
+	long pad;		/* off 0 */
+	struct mid m;		/* off 8 */
+};
+
+/* m.arr[i].b, i in [0, 7], 4-byte read: stays within m.arr. */
+SEC("syscall")
+__success
+int nested_struct_in_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct nest *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 12);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/* m.arr[i].b, i in [0, 15]: max offset runs past the end of m.arr. */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct nest")
+int nested_struct_out_of_bounds(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct nest *o;
+	int val = 0;
+
+	o = bpf_obj_new(typeof(*o));
+	if (!o)
+		return 0;
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 15;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 12);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	bpf_obj_drop(o);
+	return val;
+}
+
+/*
+ * A trailing flexible array member makes the struct tail an unbounded region,
+ * so a varying access into it has no upper bound to exceed (mirrors how
+ * unix_address.name[] is accessed via sun_path[i]).
+ *
+ * A custom (program) BTF struct can only be reached as a PTR_TO_BTF_ID through
+ * bpf_obj_new(), which does not allow flexible array members. So use a kernel
+ * type via an untrusted PTR_TO_BTF_ID (bpf_core_cast()). struct vring_used ends
+ * with a flexible array 'ring[]' of struct vring_used_elem { __virtio32 id, len; },
+ * mirroring a 'struct foo { int a; int b; }' flexible array.
+ */
+
+/* ring[i].len, i in [0, 63], 4-byte read: the flexible array has no upper bound. */
+SEC("syscall")
+__success
+int flex_array_field(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct vring_used *o;
+	int val = 0;
+
+	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 63;						\
+	r2 *= 8;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 8);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	return val;
+}
+
+/*
+ * ring[] accessed with step 4 (half of sizeof(struct vring_used_elem)): the
+ * step is not a whole number of elements, so the flexible array stride does
+ * not match.
+ */
+SEC("syscall")
+__failure
+__msg("invalid variable offset access into struct vring_used")
+int flex_array_unaligned_step(void *ctx)
+{
+	unsigned long i = bpf_get_prandom_u32();
+	struct vring_used *o;
+	int val = 0;
+
+	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
+	asm volatile ("						\
+	r1 = %[o];						\
+	r2 = %[i];						\
+	r2 &= 7;						\
+	r2 *= 4;						\
+	r1 += r2;						\
+	%[val] = *(u32 *)(r1 + 4);				\
+"	: [val] "=r"(val)
+	: [o] "r"(o), [i] "r"(i)
+	: "r1", "r2");
+	return val;
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 34/36] selftests/bpf: tests for loop hierarchy computation
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (32 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
  2026-09-26 14:20 ` [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening Eduard Zingerman
  35 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	Emil Tsalapatis

Test cases covering the following branches in bpf_compute_loops():
- Case B: simple backedge creating a loop (loop_single)
- Case B: two independent loops (loop_two_independent)
- Case B + D: nested loops where the inner header's loop_header points
  to the outer header (loop_nested)
- Case C: diamond CFG with no loops (fwd_edges_no_loop)
- Case D: sibling inner loops within one outer loop
  (loop_nested_siblings)
- Three levels of loop nesting (loop_three_levels)
- Loop with an if-else body containing forward branches
  (loop_with_if_else)
- Case E: An irreducible loop (loop_irreducible)
- A self-loop (loop_self)
- A test with sibling loops nested in outer loops
  (loop_nested_siblings_common_ancestors)

Cases B, C, D and E are described in loops.c:compute_loops_in_subprog().

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 .../selftests/bpf/progs/verifier_loop_hierarchy.c  | 317 +++++++++++++++++++++
 2 files changed, 319 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 9f188ff9be89..be4fd187ba5f 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -73,6 +73,7 @@
 #include "verifier_live_stack.skel.h"
 #include "verifier_liveness_exp.skel.h"
 #include "verifier_load_acquire.skel.h"
+#include "verifier_loop_hierarchy.skel.h"
 #include "verifier_loops1.skel.h"
 #include "verifier_lwt.skel.h"
 #include "verifier_map_in_map.skel.h"
@@ -259,6 +260,7 @@ void test_verifier_leak_ptr(void)             { RUN(verifier_leak_ptr); }
 void test_verifier_linked_scalars(void)       { RUN(verifier_linked_scalars); }
 void test_verifier_live_stack(void)           { RUN(verifier_live_stack); }
 void test_verifier_liveness_exp(void)         { RUN(verifier_liveness_exp); }
+void test_verifier_loop_hierarchy(void)       { RUN(verifier_loop_hierarchy); }
 void test_verifier_loops1(void)               { RUN(verifier_loops1); }
 void test_verifier_lwt(void)                  { RUN(verifier_lwt); }
 void test_verifier_map_in_map(void)           { RUN(verifier_map_in_map); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
new file mode 100644
index 000000000000..5c5bf5b44529
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_loop_hierarchy.c
@@ -0,0 +1,317 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+/*
+ * kernel/bpf/loops.c:compute_loops() distinguish between
+ * the following cases:
+ * - B: backedge -> simple loop
+ * - C: cross edge to non-loop node -> no-op
+ * - D: edge to node whose header is in DFS path -> nested loop
+ * - E: edge to node whose header is NOT in DFS path -> irreducible
+ *
+ * Below test cases cover the above branches in various combinations.
+ */
+
+/* Case B: single bounded loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("          2   3: {{.*}} (95) exit")
+__naked void loop_single(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case B: two independent loops at the same nesting level. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 4{{$}}")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  2       0   1: {{.*}} (07) r0 += 1")
+__msg("  2   1   1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("          2   3: {{.*}} (b7) r1 = 0")
+__msg("  1       3   4: {{.*}} (07) r1 += 1")
+__msg("  1   4   4   5: {{.*}} (a5) if r1 < 0xa goto pc-2")
+__msg("          5   6: {{.*}} (95) exit")
+__naked void loop_two_independent(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 10 goto 2b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case B + D: nested loops. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")			/* outer loop header */
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")			/* outer loop insn */
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")			/* inner loop header */
+__msg("  1   3   3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")	/* inner loop insn */
+__msg("  1   1   4   5: {{.*}} (a5) if r0 < 0xa goto pc-5")	/* outer loop insn */
+__msg("          5   6: {{.*}} (95) exit")
+__naked void loop_nested(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case C: forward edges, no loops. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (85) call bpf_get_prandom_u32")
+__msg("          0   1: {{.*}} (25) if r0 > 0x0 goto pc+2")
+__msg("          1   2: {{.*}} (b7) r0 = 2")
+__msg("          2   3: {{.*}} (05) goto pc+1")
+__msg("          1   4: {{.*}} (b7) r0 = 3")
+__msg("          1   5: {{.*}} (95) exit")
+__naked void fwd_edges_no_loop(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r0 = 2;						\
+	goto 2f;					\
+1:	r0 = 3;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/* Case B + D: two sibling inner loops within one outer loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("loop at 6, nested in 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")
+__msg("  1   3   3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")
+__msg("  1   1   4   5: {{.*}} (b7) r2 = 0")
+__msg("  1   1   5   6: {{.*}} (07) r2 += 1")
+__msg("  1   6   6   7: {{.*}} (a5) if r2 < 0x5 goto pc-2")
+__msg("  1   1   7   8: {{.*}} (a5) if r0 < 0xa goto pc-8")
+__msg("          8   9: {{.*}} (95) exit")
+__naked void loop_nested_siblings(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 5 goto 3b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Three levels of nesting. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 3, nested in 1")
+__msg("loop at 5, nested in 3")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (b7) r1 = 0")
+__msg("  1   1   2   3: {{.*}} (07) r1 += 1")
+__msg("  1   3   3   4: {{.*}} (b7) r2 = 0")
+__msg("  1   3   4   5: {{.*}} (07) r2 += 1")
+__msg("  1   5   5   6: {{.*}} (a5) if r2 < 0x3 goto pc-2")
+__msg("  1   3   6   7: {{.*}} (a5) if r1 < 0x5 goto pc-5")
+__msg("  1   1   7   8: {{.*}} (a5) if r0 < 0xa goto pc-8")
+__msg("          8   9: {{.*}} (95) exit")
+__naked void loop_three_levels(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 3 goto 3b;				\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Loop with an if-else body (forward branch inside loop, Case C). */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (b7) r0 = 0")
+__msg("  1       0   1: {{.*}} (07) r0 += 1")
+__msg("  1   1   1   2: {{.*}} (bf) r1 = r0")
+__msg("  1   1   2   3: {{.*}} (25) if r1 > 0x5 goto pc+1")
+__msg("  1   1   3   4: {{.*}} (b7) r1 = 1")
+__msg("  1   1   3   5: {{.*}} (0f) r0 += r1")
+__msg("  1   1   5   6: {{.*}} (a5) if r0 < 0x64 goto pc-6")
+__msg("          6   7: {{.*}} (95) exit")
+__naked void loop_with_if_else(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = r0;					\
+	if r1 > 5 goto 2f;				\
+	r1 = 1;						\
+2:	r0 += r1;					\
+	if r0 < 100 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/* Case E: irreducible loop. */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 3, irreducible")
+__msg("Program dump")
+__msg("         -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("          0   1: {{.*}} (b7) r1 = 0")
+__msg("          1   2: {{.*}} (25) if r0 > 0x5 goto pc+2")
+__msg("  1       2   3: {{.*}} (b7) r1 = 1")
+__msg("  1   3   3   4: {{.*}} (05) goto pc+1")
+__msg("          2   5: {{.*}} (b7) r1 = 2")
+__msg("  1   3   2   6: {{.*}} (0f) r0 += r1")
+__msg("  1   3   6   7: {{.*}} (a5) if r0 < 0x10 goto pc-5")
+__msg("          7   8: {{.*}} (95) exit")
+__naked void loop_irreducible(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+	if r0 > 5 goto 2f;				\
+1:	r1 = 1;						\
+	goto 3f;					\
+2:	r1 = 2;						\
+3:	r0 += r1;					\
+	if r0 < 16 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("socket")
+__failure
+__log_level(2)
+__msg("loop at 1")
+__msg("infinite loop detected at insn 1")
+__naked void loop_self(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+1:	if r0 < 10 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Loops at 5 and 7 are nested inside loops at 3 and 1.
+ * The edge 6 -> 7 exits loop at 5 and enters loop at 7.
+ * Check that it is logged only as an exit from loop at 5.
+ *
+ *  0: r0 = 0;
+ *     do {
+ *  1:     r0++;
+ *  2:     r1 = 0;
+ *         do {
+ *  3:         r1++;
+ *  4:         r2 = 0;
+ *             do {
+ *  5:             r2++;
+ *  6:         } while (r2 < 3);
+ *             do {
+ *  7:             r2--;
+ *  8:         } while (r2 != 0);
+ *  9:     } while (r1 < 4);
+ * 10: } while (r0 < 5);
+ * 11: return r0;
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg_next("  backedge from 10, latch at 10")
+__msg_next("  exit from 10 to 11")
+__msg_next("loop at 3, nested in 1")
+__msg_next("  backedge from 9, latch at 9")
+__msg_next("  exit from 9 to 10")
+__msg_next("loop at 5, nested in 3")
+__msg_next("  backedge from 6, latch at 6")
+__msg_next("  exit from 6 to 7")
+__msg_next("loop at 7, nested in 3")
+__msg_next("  backedge from 8, latch at 8")
+__msg_next("  exit from 8 to 9")
+__naked void loop_nested_siblings_common_ancestors(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	r2 = 0;						\
+3:	r2 += 1;					\
+	if r2 < 3 goto 3b;				\
+4:	r2 += -1;					\
+	if r2 != 0 goto 4b;				\
+	if r1 < 4 goto 2b;				\
+	if r0 < 5 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (33 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 34/36] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:27   ` bot+bpf-ci
  2026-09-26 14:20 ` [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening Eduard Zingerman
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

Coverage:
- straight-line code with an ldimm64 in the middle (exercises the
  bpf_is_ldimm64() skip in compute_predecessors());
- an asymmetric if-then-else diamond (unequal-depth idoms_intersect());
- a simple loop, a loop with an if-else body, and nested loops;
- a loop header with two back edges (three predecessors);
- an irreducible CFG (fixpoint convergence);
- immediate dominators across subprogram boundaries, including a callee
  whose entry instruction is itself a loop header;
- a self-loop;
- an indirect jump (gotox) with a two-entry jump table.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c  |   2 +
 tools/testing/selftests/bpf/progs/verifier_idoms.c | 390 +++++++++++++++++++++
 2 files changed, 392 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index be4fd187ba5f..82b21ec472ba 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -57,6 +57,7 @@
 #include "verifier_helper_packet_access.skel.h"
 #include "verifier_helper_restricted.skel.h"
 #include "verifier_helper_value_access.skel.h"
+#include "verifier_idoms.skel.h"
 #include "verifier_int_ptr.skel.h"
 #include "verifier_iterating_callbacks.skel.h"
 #include "verifier_jeq_infer_not_null.skel.h"
@@ -244,6 +245,7 @@ void test_verifier_helper_access_var_len(void) { RUN(verifier_helper_access_var_
 void test_verifier_helper_packet_access(void) { RUN(verifier_helper_packet_access); }
 void test_verifier_helper_restricted(void)    { RUN(verifier_helper_restricted); }
 void test_verifier_helper_value_access(void)  { RUN(verifier_helper_value_access); }
+void test_verifier_idoms(void)                { RUN(verifier_idoms); }
 void test_verifier_int_ptr(void)              { RUN(verifier_int_ptr); }
 void test_verifier_iterating_callbacks(void)  { RUN(verifier_iterating_callbacks); }
 void test_verifier_jeq_infer_not_null(void)   { RUN(verifier_jeq_infer_not_null); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_idoms.c b/tools/testing/selftests/bpf/progs/verifier_idoms.c
new file mode 100644
index 000000000000..a29d6332baa4
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_idoms.c
@@ -0,0 +1,390 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+/*
+ * kernel/bpf/loops.c:bpf_compute_idoms() computes the immediate dominator
+ * of every instruction (Cooper et al, "A Simple, Fast Dominance Algorithm").
+ *
+ * The immediate dominator is printed by log_program() as the numeric column
+ * immediately before the "<insn#>:" field of every "Program dump" line:
+ *
+ *   Program dump (scc? loop_header? idom insn#: live_regs_before):
+ *            -1   0: ....... (b7) r0 = 0     <- idom(0) = -1 (subprog entry)
+ *             0   1: ....... (07) r0 += 1    <- idom(1) = 0
+ *             ^^^^^^
+ *             idom insn#
+ *
+ * The __msg() patterns below wildcard the scc/loop_header/live_regs columns and
+ * anchor on "<idom>   <insn#>:" followed by the disassembled instruction, so
+ * they assert the idom value of each instruction.
+ */
+
+/*
+ * Straight-line code, with an ldimm64 in the middle. Instruction index 2 is
+ * the second half of the ldimm64 at index 1 and is not a CFG node; the idom of
+ * index 3 must be 1, not 2. Exercises the bpf_is_ldimm64() skip in both passes
+ * of compute_predecessors().
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (18) r1 = 0x1122334455667788")
+__msg("{{.*}}  1   3: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  3   4: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  4   5: {{.*}} (95) exit")
+__naked void straight_line_ldimm64(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r1 = 0x1122334455667788 ll;			\
+	r0 += 1;					\
+	r0 += r1;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Asymmetric if-then-else diamond: the "then" arm is 5 instructions long, the
+ * "else" arm is a single instruction. The merge point (insn 8) is dominated by
+ * the branch (insn 1), not by either arm. This forces idoms_intersect() to walk
+ * the two predecessors up unequal postorder depths before they meet.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (25) if r0 > 0x0 goto pc+5")
+__msg("{{.*}}  1   2: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  2   3: {{.*}} (b7) r1 = 2")
+__msg("{{.*}}  3   4: {{.*}} (b7) r1 = 3")
+__msg("{{.*}}  4   5: {{.*}} (b7) r1 = 4")
+__msg("{{.*}}  5   6: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  1   7: {{.*}} (b7) r1 = 9")
+__msg("{{.*}}  1   8: {{.*}} (95) exit")
+__naked void asymmetric_diamond(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r1 = 1;						\
+	r1 = 2;						\
+	r1 = 3;						\
+	r1 = 4;						\
+	goto 2f;					\
+1:	r1 = 9;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Simple loop: the header (insn 1) is dominated by the pre-header (insn 0). The
+ * back edge (insn 2 -> insn 1) must not change idom(1); the back-edge
+ * predecessor folds up via idoms_intersect().
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
+__msg("{{.*}}  2   3: {{.*}} (95) exit")
+__naked void simple_loop(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Loop with an if-else in its body. The in-loop merge point (insn 5) is
+ * dominated by the in-loop branch (insn 3), combining a back-edge intersect
+ * with a forward-diamond intersect.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (bf) r1 = r0")
+__msg("{{.*}}  2   3: {{.*}} (25) if r1 > 0x5 goto pc+1")
+__msg("{{.*}}  3   4: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  3   5: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  5   6: {{.*}} (a5) if r0 < 0x64 goto pc-6")
+__msg("{{.*}}  6   7: {{.*}} (95) exit")
+__naked void loop_with_if_else_body(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = r0;					\
+	if r1 > 5 goto 2f;				\
+	r1 = 1;						\
+2:	r0 += r1;					\
+	if r0 < 100 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Nested loops: the idom chain runs inner-header -> outer-body -> outer-header
+ * -> pre-header. Exercises intersect across back edges at two nesting depths.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
+__msg("{{.*}}  1   2: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  2   3: {{.*}} (07) r1 += 1")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r1 < 0x5 goto pc-2")
+__msg("{{.*}}  4   5: {{.*}} (a5) if r0 < 0xa goto pc-5")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__naked void nested_loops(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	r1 = 0;						\
+2:	r1 += 1;					\
+	if r1 < 5 goto 2b;				\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Loop header (insn 1) with two back edges (from insn 4 and insn 5), i.e. three
+ * predecessors. Both back edges route through the incrementing header, so the
+ * loop is bounded and verifies. Repeated idoms_intersect() at the header must
+ * stay stable and keep idom(1) at the pre-header (insn 0).
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r6 = 0")
+__msg("{{.*}}  0   1: {{.*}} (07) r6 += 1")
+__msg("{{.*}}  1   2: {{.*}} (25) if r6 > 0xa goto pc+3")
+__msg("{{.*}}  2   3: {{.*}} (bf) r7 = r6")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r7 < 0x5 goto pc-4")
+__msg("{{.*}}  4   5: {{.*}} (05) goto pc-5")
+__msg("{{.*}}  2   6: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  6   7: {{.*}} (95) exit")
+__naked void multi_backedge_header(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	r6 += 1;					\
+	if r6 > 10 goto 2f;				\
+	r7 = r6;					\
+	if r7 < 5 goto 1b;				\
+	goto 1b;					\
+2:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Irreducible CFG (loop with two entries, insn 3 and insn 5, reached from the
+ * insn 2 branch). Dominators remain well-defined; this is a convergence test
+ * for the fixpoint under irreducibility. Note insn 6 is dominated by the branch
+ * (insn 2), not by either arm.
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  1   2: {{.*}} (25) if r0 > 0x5 goto pc+2")
+__msg("{{.*}}  2   3: {{.*}} (b7) r1 = 1")
+__msg("{{.*}}  3   4: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  2   5: {{.*}} (b7) r1 = 2")
+__msg("{{.*}}  2   6: {{.*}} (0f) r0 += r1")
+__msg("{{.*}}  6   7: {{.*}} (a5) if r0 < 0x10 goto pc-5")
+__msg("{{.*}}  7   8: {{.*}} (95) exit")
+__naked void irreducible(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+	if r0 > 5 goto 2f;				\
+1:	r1 = 1;						\
+	goto 3f;					\
+2:	r1 = 2;						\
+3:	r0 += r1;					\
+	if r0 < 16 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Idoms are computed per subprog. Layout (libbpf preorder) is:
+ *   main: 0..1, sub: 2..7
+ * The callee entry (insn 2) must have idom -1 (subprog reset), and idoms inside
+ * the callee must reference only the callee's instructions, never main's. The
+ * in-callee merge (insn 7) is dominated by the in-callee branch (insn 3). The
+ * branch is on a prandom value so it is not constant-folded away.
+ */
+static __naked __noinline __used
+unsigned long idoms_diamond_sub(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 > 0 goto 1f;				\
+	r0 = 1;						\
+	goto 2f;					\
+1:	r0 = 2;						\
+2:	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call pc+1")
+__msg("{{.*}}  0   1: {{.*}} (95) exit")
+__msg("{{.*}} -1   2: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  2   3: {{.*}} (25) if r0 > 0x0 goto pc+2")
+__msg("{{.*}}  3   4: {{.*}} (b7) r0 = 1")
+__msg("{{.*}}  4   5: {{.*}} (05) goto pc+1")
+__msg("{{.*}}  3   6: {{.*}} (b7) r0 = 2")
+__msg("{{.*}}  3   7: {{.*}} (95) exit")
+__naked void multi_subprog(void)
+{
+	asm volatile ("					\
+	call idoms_diamond_sub;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The first instruction of a subprog (insn 3) is itself a loop header, i.e. it
+ * has an incoming back edge. Its idom must still be -1 (the subprog entry has no
+ * dominator). Exercises the idoms[start]=0 ... idoms[start]=-1 handling in
+ * compute_subprog_idoms().
+ */
+static __naked __noinline __used
+unsigned long idoms_entry_header_sub(void)
+{
+	asm volatile ("					\
+1:	r1 += 1;					\
+	if r1 < 10 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (b7) r1 = 0")
+__msg("{{.*}}  0   1: {{.*}} (85) call pc+1")
+__msg("{{.*}}  1   2: {{.*}} (95) exit")
+__msg("{{.*}} -1   3: {{.*}} (07) r1 += 1")
+__msg("{{.*}}  3   4: {{.*}} (a5) if r1 < 0xa goto pc-2")
+__msg("{{.*}}  4   5: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__naked void entry_is_loop_header(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	call idoms_entry_header_sub;			\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Self-loop: insn 1 is its own predecessor. idom(1) must be the pre-header
+ * (insn 0); the self-edge is ignored (idoms[pred] == -1 on the first pass, then
+ * idoms_intersect(a == b) short-circuits). Program is rejected later for an
+ * infinite loop, but the dump (and idoms) are printed before that.
+ */
+SEC("socket")
+__failure
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
+__msg("{{.*}}  0   1: {{.*}} (a5) if r0 < 0xa goto pc-1")
+__msg("{{.*}}  1   2: {{.*}} (95) exit")
+__msg("infinite loop detected")
+__naked void self_loop(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+1:	if r0 < 10 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64) || defined(__TARGET_ARCH_powerpc)
+/*
+ * Indirect jump (gotox) with a two-entry jump table. The gotox (insn 4) has two
+ * successors (insn 5 and insn 7), so both targets have the gotox as their only
+ * predecessor and idom. Also re-exercises the ldimm64 skip: insn 2's idom is 0,
+ * the ldimm64 at index 0 (index 1 is its second half).
+ */
+SEC("socket")
+__success
+__log_level(2)
+__msg("Program dump")
+__msg("{{.*}} -1   0: {{.*}} (18) r0 = {{0x[0-9a-f]+}}")
+__msg("{{.*}}  0   2: {{.*}} (07) r0 += 8")
+__msg("{{.*}}  2   3: {{.*}} (79) r0 = *(u64 *)(r0 +0)")
+__msg("{{.*}}  3   4: {{.*}} (0d) gotox r0")
+__msg("{{.*}}  4   5: {{.*}} (b7) r0 = 0")
+__msg("{{.*}}  5   6: {{.*}} (95) exit")
+__msg("{{.*}}  4   7: {{.*}} (b7) r0 = 1")
+__msg("{{.*}}  7   8: {{.*}} (95) exit")
+__naked void gotox_jump_table(void)
+{
+	asm volatile ("						\
+	.pushsection .jumptables,\"\",@progbits;		\
+jt0_%=:								\
+	.quad ret0_%= - socket;					\
+	.quad ret1_%= - socket;					\
+	.size jt0_%=, 16;					\
+	.global jt0_%=;						\
+	.popsection;						\
+								\
+	r0 = jt0_%= ll;						\
+	r0 += 8;						\
+	r0 = *(u64 *)(r0 + 0);					\
+	.8byte %[gotox_r0];					\
+ret0_%=:							\
+	r0 = 0;							\
+	exit;							\
+ret1_%=:							\
+	r0 = 1;							\
+	exit;							\
+"	:
+	: __imm_insn(gotox_r0, BPF_RAW_INSN(BPF_JMP | BPF_JA | BPF_X, BPF_REG_0, 0, 0, 0))
+	: __clobber_all);
+}
+#endif
+
+char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening
  2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
                   ` (34 preceding siblings ...)
  2026-09-26 14:20 ` [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
@ 2026-09-26 14:20 ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  35 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-26 14:20 UTC (permalink / raw)
  To: bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor

- Computed expression shapes: register and spilled-counter recurrences,
  substitution at latches, agreeing and conflicting backedge updates,
  and ALU, extension and byte-swap expression chains.

- Iteration bounds and header-visit counts for:
  - pre-condition loops using JEQ, JGE and JGT;
  - post-condition loops using JLT, JLE, JGE, JGT and JNE;
  - signed post-conditions using JSLT and JSLE;
  - increasing and decreasing counters, nonzero initial values,
    immediate and invariant-register bounds, and immediate exits;
  - additional exits that reduce the minimum iteration count.

- Widening, backedge clamping and state-pruning convergence,
  including multiple induction variables used for map-value accesses
  and nonconstant entry values requiring a common alignment.

- Nested and sibling loops, exits to another loop header, exits across
  multiple nesting levels, and side entries into loop nests.
  Check conservative summaries for irreducible children while still
  widening supported reducible descendants.

- Invalidation of stack-slot expressions after indirect stack writes.

- Avoid widening registers used to address stack spills, fills and
  dynptr constructor arguments, including dependencies through nested
  loops. Conversely, allow widening addresses of byte-sized stores.

- Conditional assignments joining scalar values and stack pointers,
  including preservation of a common base and gcd-derived step.

- Conservative fallback for unsupported recurrences, unavailable
  counter spills and zero computed header counts.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
---
 tools/testing/selftests/bpf/prog_tests/verifier.c |    2 +
 tools/testing/selftests/bpf/progs/verifier_scev.c | 1108 +++++++++++++++++++++
 2 files changed, 1110 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 82b21ec472ba..2bc1ef910c16 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -107,6 +107,7 @@
 #include "verifier_ringbuf.skel.h"
 #include "verifier_runtime_jit.skel.h"
 #include "verifier_scalar_ids.skel.h"
+#include "verifier_scev.skel.h"
 #include "verifier_sdiv.skel.h"
 #include "verifier_search_pruning.skel.h"
 #include "verifier_sock.skel.h"
@@ -294,6 +295,7 @@ void test_verifier_regalloc(void)             { RUN(verifier_regalloc); }
 void test_verifier_ringbuf(void)              { RUN(verifier_ringbuf); }
 void test_verifier_runtime_jit(void)          { RUN(verifier_runtime_jit); }
 void test_verifier_scalar_ids(void)           { RUN(verifier_scalar_ids); }
+void test_verifier_scev(void)                 { RUN(verifier_scev); }
 void test_verifier_sdiv(void)                 { RUN(verifier_sdiv); }
 void test_verifier_search_pruning(void)       { RUN(verifier_search_pruning); }
 void test_verifier_sock(void)                 { RUN(verifier_sock); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_scev.c b/tools/testing/selftests/bpf/progs/verifier_scev.c
new file mode 100644
index 000000000000..69e9062870ac
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_scev.c
@@ -0,0 +1,1108 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf.h>
+#include <stdbool.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "bpf_kfuncs.h"
+
+struct map_val {
+	char foo[1024];
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_HASH);
+	__uint(max_entries, 1);
+	__type(key, int);
+	__type(value, struct map_val);
+} map SEC(".maps");
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__naked void simple_loop1(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+loop_%=:						\
+	if r0 == 10 goto exit_%=;			\
+	r0 += 1;					\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("8: (7b) *(u64 *)(r1 +0) = r0     ; *fp-8 (+ *fp-8 1) -> ?")
+__naked void indirect_write_invalidates_scev(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 8) = r0;				\
+loop_%=:						\
+	r0 = *(u64 *)(r10 - 8);				\
+	if r0 == 10 goto exit_%=;			\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	r1 = r10;					\
+	r1 += -8;					\
+	*(u64 *)(r1 + 0) = r0;				\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 2:")
+__msg_next("  *fp-8=(+ *fp-8 1) / (linear *fp-8 1)")
+__msg_next(" scev at latch 3:")
+__msg_next("  r0=*fp-8 / (linear *fp-8 1)")
+__msg_next("  *fp-8=*fp-8 / (linear *fp-8 1)")
+__naked void simple_loop2(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 8) = r0;				\
+loop_%=:						\
+	r0 = *(u64 *)(r10 - 8);				\
+	if r0 == 10 goto exit_%=;			\
+	r0 = *(u64 *)(r10 - 8);				\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	goto loop_%=;					\
+exit_%=:						\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 2:")
+__msg_next("  r1=(+ r1 1) / (linear r1 1)")
+__naked void meet_agrees(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r1 = 0;						\
+1:							\
+	if r1 == 2 goto 3f;				\
+	if r0 == 7 goto 2f;				\
+	r1 += 1;					\
+	goto 1b;					\
+2:							\
+	r1 += 1;					\
+	goto 1b;					\
+3:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__failure
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(any r6 (+ r6 1)) / ?")
+__naked void meet_disagrees(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 3f;				\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 7 goto 1b;				\
+	r6 += 1;					\
+	goto 1b;					\
+3:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(bswap32 (bswap32 (zext32 (- (- (zext32 (+ (>>...) 1))))))) / ?")
+__naked void expr_chain(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 2f;				\
+	r6 += 1;					\
+	r6 <<= 32;					\
+	r6 >>= 32;					\
+	w6 += 1;					\
+	r6 = -r6;					\
+	w6 = -w6;					\
+	r6 = bswap32 r6;				\
+	r6 = bswap32 r6;				\
+	goto 1b;					\
+2:							\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg_next("scev at header 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next(" scev at latch 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next("scev at header 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next(" scev at latch 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__naked void nested_loop1(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 2f;				\
+	r6 += 1;					\
+	r7 = 0;						\
+3:							\
+	if r7 == 2 goto 4f;				\
+	r7 += 1;					\
+	goto 3b;					\
+4:							\
+	goto 1b;					\
+2:							\
+	r0 = r7;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__msg_next("scev at header 1:")
+__msg_next("  r6=(+ r6 1) / (linear r6 1)")
+__msg_next("  r7=?")
+__msg_next(" scev at latch 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next("scev at header 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg_next(" scev at latch 4:")
+__msg_next("  r7=(+ r7 1) / (linear r7 1)")
+__msg("loop header at 1, can't widen r7, expr is ?")
+__msg("loop header at 4, widening r7 to 0..2 step 1")
+__log_level(2)
+__naked void nested_loop_hdr_backedge1(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:							\
+	if r6 == 2 goto 3f;				\
+	r6 += 1;					\
+	  r7 = 0;					\
+2:							\
+	  if r7 == 2 goto 1b;				\
+	  r7 += 1;					\
+	  goto 2b;					\
+3:							\
+	r0 = r7;					\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 10 post-cond")
+__msg("loop header at 1, widening r0 to 0..9 step 1")
+__msg("processed 5 insns")
+__naked void post_cond_jlt(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 < 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 3, header_count is 8")
+__msg("loop header at 3, widening r0 to 2..9 step 1")
+__msg("3: R0=scalar(smin=umin=smin32=umin32=2,smax=umax=smax32=umax32=9,{{.*}})")
+__msg("processed 6 insns")
+__naked void post_cond_jlt_with_base(void)
+{
+	asm volatile ("					\
+	r7 = 10 ll;	/* ldimm64 for a twist */	\
+	r0 = 2;						\
+1:	r0 += 1;					\
+	if r0 < r7 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 11")
+__msg("loop header at 1, widening r0 to 0..10 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=10,{{.*}})")
+__msg("processed 5 insns")
+__naked void post_cond_jle(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 <= 10 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 10")
+__msg("loop header at 1, widening r0 to 0..9 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=9,{{.*}})")
+__msg("processed 6 insns")
+__naked void post_cond_jge(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 >= 10 goto 2f;				\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 11")
+__msg("loop header at 1, widening r0 to 0..10 step 1")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=10,{{.*}})")
+__msg("processed 6 insns")
+__naked void pre_cond_jge(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	if r0 >= 10 goto 2f;				\
+	r0 += 1;					\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 2 post-cond")
+__naked void post_cond_jgt(void)
+{
+	asm volatile ("					\
+	r0 = 2;						\
+1:	r0 += -1;					\
+	if r0 > 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next(" scev at latch 2:")
+__msg_next("  r0=(+ r0 1) / (linear (+ r0 1) 1)")
+__msg("loop header at 1, widening r0")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,{{.*}})")
+__msg("1: (07) r0 += 1                       ; R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("2: (55) if r0 != 0x3 goto pc-2")
+__msg("3: (95) exit")
+__msg("loop header at 1, clamping r0")
+__msg("from 2 to 1: safe")
+__msg("processed 5 insns")
+__naked void post_cond_jne(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 != 3 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 -1) / (linear r0 -1)")
+__msg_next(" scev at latch 2:")
+__msg_next("  r0=(+ r0 -1) / (linear (+ r0 -1) -1)")
+__msg("loop header at 1, header_count is 3 post-cond")
+__msg("loop header at 1, widening r0 to 1..3 step 1")
+__msg("1: R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,var_off=(0x0; 0x3)) loop_stack=1")
+__msg("loop header at 1, clamping r0 to 1..2 step 1")
+__msg("processed 5 insns")
+__naked void post_cond_jne_neg_step(void)
+{
+	asm volatile ("					\
+	r0 = 3;						\
+1:	r0 += -1;					\
+	if r0 != 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next(" scev at latch 1:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg("loop header at 1, widening r0")
+__msg("1: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("1: (15) if r0 == 0x3 goto pc+2        ; R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,{{.*}})")
+__msg("2: (07) r0 += 1                       ; R0=scalar(smin=umin=smin32=umin32=1,smax=umax=smax32=umax32=3,{{.*}})")
+__msg("3: (05) goto pc-3")
+__msg("loop header at 1, clamping r0")
+__msg("1: safe")
+__msg("from 1 to 4: R0=3")
+__msg("4: R0=3")
+__msg("4: (95) exit")
+__msg("processed 6 insns")
+__naked void pre_cond_je1(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	if r0 == 3 goto 2f;				\
+	r0 += 1;					\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is [0..3] post-cond")
+__naked void one_backedge_two_exits(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	call %[bpf_get_prandom_u32];			\
+	if r0 == 0 goto 2f;				\
+	r6 += 1;					\
+	if r6 != 3 goto 1b;				\
+2:	r0 = r6;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg_next("scev at header 11:")
+__msg_next("  r0=(+ r0 1) / (linear r0 1)")
+__msg_next("  r1=(+ r1 2) / (linear r1 2)")
+__msg_next(" scev at latch 16:")
+__msg_next("  r0=(+ r0 1) / (linear (+ r0 1) 1)")
+__msg_next("  r1=(+ r1 2) / (linear (+ r1 2) 2)")
+__msg("loop header at 11, widening r0")
+__msg("loop header at 11, widening r1")
+__msg("11: R0=scalar(smin=smin32=0,smax=umax=smax32=umax32=7,var_off=(0x0; 0x7)) R1=scalar(smin=smin32=0,smax=umax=smax32=umax32=14,var_off=(0x0; 0xe),step=0+2)")
+/* loop exit */
+__msg("16: (a5) if r0 < 0x8 goto pc-6")
+__msg("exiting loop 11")
+__msg("17: (95) exit")
+/* second iteration */
+__msg("loop header at 11, clamping r0")
+__msg("loop header at 11, clamping r1")
+/* iteration convergence */
+__msg("from 16 to 11: safe")
+__not_msg("{{^}}11:")
+/* map lookup error path */
+__msg("from 7 to 17: safe")
+__naked void correlated_regs(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	*(u64*)(r10 - 8) = r1;				\
+	r2 = r10;					\
+	r2 += -8;					\
+	r1 = %[map] ll;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 2f;				\
+	r6 = r0;					\
+	r0 = 0;						\
+	r1 = 0;						\
+1:	r2 = r6;					\
+	r2 += r1;					\
+	*(u8 *)(r2 + 0) = 1;				\
+	r0 += 1;					\
+	r1 += 2;					\
+	if r0 < 8 goto 1b;				\
+2:	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+/*
+ * k = 0
+ * for (i = 0; i < 4; i++):
+ *   for (j = 0; j < 4; j++):
+ *     k += 1
+ *     k <<= 1   // make SCEV construction not possible
+ *     k >>= 1
+ * map[k] = 1    // make k precise
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("19: (72) *(u8 *)(r4 +0) = 1           ; R4=map_value(id={{.*}},map=map,ks=4,vs=1024,imm=16)")
+__not_msg("19: ")
+__msg("processed 106 insns")
+__naked void nested_loops_precise_var1(void)
+{
+	asm volatile ("					\
+	*(u64*)(r10 - 8) = 0;				\
+	r1 = %[map] ll;					\
+	r2 = r10;					\
+	r2 += -8;					\
+	call %[bpf_map_lookup_elem];			\
+	if r0 == 0 goto 3f;				\
+	r1 = 0;						\
+	r3 = 0;						\
+	/* outer loop */				\
+1:	r2 = 0;						\
+	/* inner loop */				\
+2:	r2 += 1;					\
+	r3 += 1;					\
+	r3 <<= 1;					\
+	r3 >>= 1;					\
+	if r2 < 4 goto 2b;				\
+	r1 += 1;					\
+	if r1 < 4 goto 1b;				\
+	r4 = r0;					\
+	r4 += r3;					\
+	*(u8 *)(r4 + 0) = 1;				\
+	r0 = 0;						\
+3:	exit;						\
+"	:
+	: __imm(bpf_map_lookup_elem),
+	  __imm_addr(map)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 1 pre-cond")
+__naked void pre_cond_jgt(void)
+{
+	asm volatile ("					\
+	r0 = -1;					\
+1:	if r0 > 1 goto 2f;				\
+	r0 += 1;					\
+	goto 1b;					\
+2:	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 2 post-cond")
+__naked void loop_jslt_latch_post(void)
+{
+	asm volatile ("					\
+	r0 = -2;					\
+1:	r0 += 1;					\
+	if r0 s< 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 1 post-cond")
+__naked void loop_jslt_latch_post1(void)
+{
+	asm volatile ("					\
+	r0 = -1;					\
+1:	r0 += 1;					\
+	if r0 s< 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, 0 iterations count")
+__naked void loop_jslt_latch_post2(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+1:	r0 += 1;					\
+	if r0 s< 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 1, header_count is 3 post-cond")
+__naked void loop_jsle_latch_post(void)
+{
+	asm volatile ("					\
+	r0 = -2;					\
+1:	r0 += 1;					\
+	if r0 s<= 0 goto 1b;				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop header at 2, header_count is 100 post-cond")
+__msg("loop header at 4, header_count is [0..100] post-cond")
+__msg("loop header at 6, header_count is [0..100] post-cond")
+__msg("processed 20 insns")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void nested_loop_with_two_exits(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = 0;						\
+1:	r6 += 1;					\
+	r7 = 0;						\
+2:	r7 += 1;					\
+	r8 = 0;						\
+3:	r8 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 42 goto +1;				\
+	goto 4f;					\
+	if r8 < 100 goto 3b;				\
+	if r7 < 100 goto 2b;				\
+4:	if r6 < 100 goto 1b;				\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+SEC("xdp")
+__success
+__log_level(2)
+__naked void exit_loop_into_loop_header(void)
+{
+	asm volatile ("					\
+	r1 = 0;						\
+	r2 = 0;						\
+loop_a_%=:						\
+	r1 += 1;					\
+	if r1 < 10 goto loop_a_%=;			\
+loop_b_%=:						\
+	r2 += 1;					\
+	if r2 < 10 goto loop_b_%=;			\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * This exercises verifier.c:loop_stack_{pop,push}() implementation,
+ * at 'goto d' loops 'b' and 'a' have to be popped from stack,
+ * while loops 'c' and 'd' have to be pushed to stack.
+ *
+ *   loop a:                  // header 5
+ *     loop b:                // header 6
+ *       if (rand) goto d;    // 8 -> 12, side entry into inner loop d
+ *       ...
+ *   loop c:                  // header 11
+ *     loop d:                // header 12
+ *       ...
+ */
+SEC("xdp")
+__log_level(2)
+__msg("loop at 5")
+__msg("loop at 6, nested in 5")
+__msg("loop at 11, irreducible")
+__msg("loop at 12, nested in 11")
+/* entry via if r0 == 5 goto d_%= false branch */
+__msg("loop header at 12, header_count is 3 post-cond")
+__msg("loop header at 12, widening r9 to 0..2 step 1")
+__msg("12: R8=1 R9=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,var_off=(0x0; 0x3)) loop_stack=11,12")
+/* entry via if r0 == 5 goto d_%= true branch */
+__msg("loop header at 12, header_count is 3 post-cond")
+__msg("loop header at 12, widening r9 to 0..2 step 1")
+__msg("from 8 to 12: R8=0 R9=scalar(smin=smin32=0,smax=umax=smax32=umax32=2,var_off=(0x0; 0x3)) R10=fp0 loop_stack=11,12")
+__naked void enter_nested_loop_from_side(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r6 = 0;						\
+	r7 = 0;						\
+	r8 = 0;						\
+	r9 = 0;						\
+a_%=:	r6 += 1;					\
+b_%=:	r7 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	if r0 == 5 goto d_%=;				\
+	if r7 < 3 goto b_%=;				\
+	if r6 < 3 goto a_%=;				\
+c_%=:	r8 += 1;					\
+d_%=:	r9 += 1;					\
+	if r9 < 3 goto d_%=;				\
+	if r8 < 3 goto c_%=;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Loop A (counter in r6) contains irreducible loop B (header 5),
+ * which contains C (counter in r9).
+ * Entering B at 'body' saves and restores r6, making it appear invariant.
+ * Entering B at 'alternate' skips the save and modifies r6 instead.
+ * Hence A must not infer a SCEV expression for r6.
+ * SCEV expression for r9 in C should still be computed.
+ *
+ *  0: r6 = 0;
+ *     do {                              // A
+ *  1:     r6++;
+ *  2:     r0 = bpf_get_prandom_u32();
+ *  3:     r7 = 0;
+ *  4:     if (r0 > 5) goto alternate;
+ *  5: B:  r8 = r6;                      // B
+ *  6:     goto body;
+ *  7: alternate:
+ *         r8 = r6;
+ *  8:     r8++;
+ *  9: body:
+ *         r6 = r8;
+ * 10:     r9 = 0;
+ *         do {                          // C
+ * 11:         r9++;
+ * 12:     } while (r9 < 3);
+ * 13:     r7++;
+ * 14:     if (r7 < 4) goto B;
+ * 15: } while (r6 < 4);
+ * 16: r0 = 0;
+ * 17: return r0;
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop at 1{{$}}")
+__msg("loop at 5, nested in 1, irreducible")
+__msg("loop at 11, nested in 5")
+__msg("scev at header 1:")
+__msg_next("  r6=?")
+__msg("scev at header 11:")
+__msg_next("  r9=(+ r9 1) / (linear r9 1)")
+__msg("loop header at 11, widening r9 to 0..2 step 1")
+__not_msg("loop header at 1, widening r6")
+__naked void nested_irreducible_loop(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+a_%=:	r6 += 1;					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = 0;						\
+	if r0 > 5 goto alternate_%=;			\
+b_%=:	r8 = r6;					\
+	goto body_%=;					\
+alternate_%=:						\
+	r8 = r6;					\
+	r8 += 1;					\
+body_%=:						\
+	r6 = r8;					\
+	r9 = 0;						\
+c_%=:	r9 += 1;					\
+	if r9 < 3 goto c_%=;				\
+	r7 += 1;					\
+	if r7 < 4 goto b_%=;				\
+	if r6 < 4 goto a_%=;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Induction variable seeded from a non-constant value. r7 enters the loop as a
+ * range aligned to 2 (prandom & 0x6 -> {0,2,4,6}) and is incremented by a
+ * non-power-of-2 slope of 6. Since the entry value is not a single point, only
+ * the power-of-two alignment shared by the entry value and the slope can be
+ * guaranteed, so the widened step is 2 - not |slope|=6, which would be unsound
+ * here (a constant entry value would have allowed step 6).
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 4, widening r7 to 0..18 step 2")
+__msg("R7=scalar(smin=smin32=0,smax=umax=smax32=umax32=18,var_off=(0x0; 0x1e),step=0+2)")
+__naked void widen_nonconst_base(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r7 = r0;					\
+	r7 &= 0x6;					\
+	r6 = 0;						\
+1:	r7 += 6;					\
+	r6 += 1;					\
+	if r6 < 3 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * Two nested loops with (1) exiting directly to (2):
+ *
+ *   for (r6 = 0; r6 < 4; r6++) {
+ *     r7 = 0;
+ *     for (; r7 <  3; r7++) {}   // (1)
+ *     for (; r7 != 0; r7--) {}   // (2)
+ *   }
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__msg("loop header at 1, widening r6 to 0..3 step 1")
+__msg("loop header at 2, widening r7 to 0..2 step 1")
+__msg("loop header at 4, widening r7 to 1..3 step 1")
+__naked void sibling_inner_loops(void)
+{
+	asm volatile ("					\
+	r6 = 0;						\
+1:	r7 = 0;						\
+2:	r7 += 1;					\
+	if r7 < 3 goto 2b;				\
+3:	r7 += -1;					\
+	if r7 != 0 goto 3b;				\
+	r6 += 1;					\
+	if r6 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+SEC("socket")
+__success
+__log_level(2)
+__msg("loop header at 0, can't compute iterations count")
+__naked void uninit_slot_counter(void)
+{
+	asm volatile ("					\
+1:	r0 = *(u64 *)(r10 - 8);				\
+	r0 += 1;					\
+	*(u64 *)(r10 - 8) = r0;				\
+	if r0 < 10 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * The assignment happens only on the second iteration:
+ *
+ *   r7 = 5;
+ *   for (r6 = 0; r6 < 3; r6++)
+ *           if (r6 == 1)
+ *                   r7 = 10;
+ *   if (r7 != 10)
+ *           invalid_stack_read();
+ *
+ * R7 is always 10 at the real exit. Widening loses the correlation between
+ * R6 and R7, leaving R7 in [5, 10] both at the loop header and after the loop.
+ * The verifier therefore rejects the possible invalid stack read.
+ */
+SEC("xdp")
+__failure
+__log_level(2)
+__msg("2: {{.*}}R7=scalar(smin=umin=smin32=umin32=5,smax=umax=smax32=umax32=10,var_off=(0x0; 0xf))")
+__msg("7: R7=scalar(smin=umin=smin32=umin32=5,smax=umax=smax32=umax32=10,var_off=(0x0; 0xf))")
+__msg("invalid read from stack R10 off=0 size=8")
+__naked void conditional_assignment_on_second_iteration(void)
+{
+	asm volatile ("					\
+	r7 = 5;						\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = 10;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	if r7 == 10 goto 4f;				\
+	r0 = *(u64 *)(r10 + 0);				\
+4:	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Join two stack pointer offsets via (any r7 r8), with r8 loop-invariant.
+ * The joined range [-16, -8] still points to initialized stack memory,
+ * so dereferencing r7 after the loop is safe.
+ */
+SEC("xdp")
+__success __retval(0)
+__log_level(2)
+__msg("loop header at 7, widening r7 to -16..-8 step 1")
+__msg("7: {{.*}}R7=fp(smin=smin32=-16,smax=smax32=-8,")
+__msg("12: R7=fp(smin=smin32=-16,smax=smax32=-8,")
+__naked void conditional_stack_pointer_assignment(void)
+{
+	asm volatile ("					\
+	*(u64 *)(r10 - 16) = 0;				\
+	*(u64 *)(r10 - 8) = 0;				\
+	r7 = r10;					\
+	r7 += -16;					\
+	r8 = r10;					\
+	r8 += -8;					\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = r8;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	r0 = *(u64 *)(r7 + 0);				\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * r7 starts as 1 + 6*k and invariant r8 as 1 + 9*k (0 <= k <= 3).
+ * The (any r7 r8) union must preserve base 1 and step gcd(6, 9) = 3.
+ */
+SEC("xdp")
+__success __retval(1)
+__log_level(2)
+__msg("r7 += 1 {{.*}}step=1+6)")
+__msg("r8 += 1 {{.*}}step=1+9)")
+__msg("loop header at 9, widening r7 to 1..28 step 3")
+__msg("9: {{.*}}R7=scalar({{.*}},step=1+3)")
+__msg("14: R7=scalar({{.*}},step=1+3)")
+__naked void conditional_scalar_assignment_gcd(void)
+{
+	asm volatile ("					\
+	call %[bpf_get_prandom_u32];			\
+	r0 &= 3;					\
+	r7 = r0;					\
+	r7 *= 6;					\
+	r7 += 1;					\
+	r8 = r0;					\
+	r8 *= 9;					\
+	r8 += 1;					\
+	r6 = 0;						\
+1:	if r6 >= 3 goto 3f;				\
+	if r6 != 1 goto 2f;				\
+	r7 = r8;					\
+2:	r6 += 1;					\
+	goto 1b;					\
+3:	r0 = r7;					\
+	r0 %%= 3;					\
+	exit;						\
+"	:
+	: __imm(bpf_get_prandom_u32)
+	: __clobber_all);
+}
+
+/*
+ * A loop induction variable used to compute the base address of a store to the
+ * stack must not be widened: the spill offset would become varying, which the
+ * verifier does not track.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 2, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__naked void no_widen_stack_spill(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u64 *)(r3 + 0) = r0;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+/*
+ * Same hazard across a loop nest: the outer induction variable r2 addresses a
+ * stack store performed inside the inner loop. The dependency is pulled up from
+ * the inner loop, so the outer loop must not widen r2.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 2, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__msg("loop header at 3, widening r1 to 0..1 step 1")
+__naked void no_widen_stack_spill_nested(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r1 = 0;						\
+2:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u64 *)(r3 + 0) = r1;				\
+	r1 += 1;					\
+	if r1 < 2 goto 2b;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * Complement of no_widen_stack_spill: a sub-register (1-byte) store to the stack
+ * lands as STACK_MISC and carries no tracked value, so the induction variable
+ * addressing it (r2) is still widened.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("widening r2 to 0..24 step 8")
+__naked void widen_byte_stack_store(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	*(u8 *)(r3 + 0) = r0;				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * A fill (BPF_LDX) at a varying stack offset loses precision just like a spill,
+ * so the induction variable computing the load base (r2) must not be widened.
+ * The slots are initialized up front so the fill itself is a valid read.
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("loop header at 6, can't widen r2, expr is (+ r2 8), requires exact stack-offset tracking")
+__naked void no_widen_stack_fill(void)
+{
+	asm volatile ("					\
+	r0 = 0;						\
+	*(u64 *)(r10 - 64) = r0;			\
+	*(u64 *)(r10 - 56) = r0;			\
+	*(u64 *)(r10 - 48) = r0;			\
+	*(u64 *)(r10 - 40) = r0;			\
+	r2 = 0;						\
+1:	r3 = r10;					\
+	r3 += -64;					\
+	r3 += r2;					\
+	r4 = *(u64 *)(r3 + 0);				\
+	r0 += 1;					\
+	r2 += 8;					\
+	if r0 < 4 goto 1b;				\
+	r0 = 0;						\
+	exit;						\
+"	::: __clobber_all);
+}
+
+/*
+ * A dynptr/iter/irq/res_spin_lock call initializes a stack object through a
+ * pointer argument, which acts like a spill base: the induction variable
+ * computing that argument's varying stack offset must not be widened, otherwise
+ * the slot can't be resolved. Here each iteration constructs an xdp dynptr at
+ * &dptrs[i].
+ */
+SEC("xdp")
+__success
+__log_level(2)
+__flag(BPF_F_TEST_STATE_FREQ)
+__msg("can't widen {{.*}}, requires exact stack-offset tracking")
+int no_widen_dynptr_kfunc_arg(struct xdp_md *ctx)
+{
+	struct bpf_dynptr dptrs[4];
+	int i;
+
+#pragma clang loop unroll(disable)
+	for (i = 0; i < 4; i++)
+		bpf_dynptr_from_xdp(ctx, 0, &dptrs[i]);
+
+	return 0;
+}
+
+char _license[] SEC("license") = "GPL";

-- 
2.55.0

^ permalink raw reply related	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions
  2026-09-26 14:20 ` [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
@ 2026-09-26 14:31   ` sashiko-bot
  0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:31 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: memory leak of env->cfg.postorder_nums in bpf_prune_dead_branches()
--

commit 3333ceb14d7fa72313c9d88f2f43fc886dbdac3d
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: save DFS postorder numbers for program instructions

Dominator intersections and SCEV worklist scheduling need to compare
instructions by their DFS postorder rank. Save per-instruction
postorder numbers alongside the existing postorder sequence.

> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
> index b0bd9ba951dfb..bd771efec66a3 100644
> --- a/kernel/bpf/cfg.c
> +++ b/kernel/bpf/cfg.c

[ ... ]

> @@ -774,43 +779,52 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
>   */
>  int bpf_compute_postorder(struct bpf_verifier_env *env)
>  {
> -	u32 cur_postorder, i, top, stack_sz, s;
> -	int *stack = NULL, *postorder = NULL, *state = NULL;
> -	struct bpf_iarray *succ;
> +	int *stack = NULL, *postorder = NULL, *postorder_nums = NULL;
> +	int subprog_idx, stack_sz, cur, s, cur_postorder, start;
> +	struct dfs_state *state = NULL;
> +	struct bpf_iarray *succ = NULL;
>  
> +	postorder_nums = kvzalloc_objs(int, env->prog->len, GFP_KERNEL_ACCOUNT);

[Severity: High]
Does this code leak memory when bpf_compute_postorder() is called multiple
times?

If a user loads a BPF program that contains a dead branch determinable at load
time, it can cause bpf_prune_dead_branches() to successfully fold the branch
and call bpf_compute_postorder() a second time:

kernel/bpf/const_fold.c:bpf_prune_dead_branches() {
    ...
    kvfree(env->cfg.insn_postorder);
    env->cfg.insn_postorder = NULL;
    return bpf_compute_postorder(env);
}

The array allocated here unconditionally overwrites env->cfg.postorder_nums
without freeing the previously allocated one. Can we add a call to
kvfree(env->cfg.postorder_nums) in bpf_prune_dead_branches() to prevent a leak
of up to 4MB per program load?

[ ... ]

> -	env->subprog_info[i].postorder_start = cur_postorder;
> +	env->subprog_info[subprog_idx].postorder_start = cur_postorder;
> +	env->cfg.postorder_nums = postorder_nums;
>  	env->cfg.insn_postorder = postorder;
>  	env->cfg.cur_postorder = cur_postorder;
>  	kvfree(stack);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=16

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning
  2026-09-26 14:20 ` [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
@ 2026-09-26 14:32   ` sashiko-bot
  2026-09-27 20:26   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:32 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: incorrect multi-line comment style
--

commit 928bc76d46def360ecb11c2c10278100fdf80716
Author: Eduard Zingerman <eddyz87@gmail.com>
Subject: selftests/bpf: test for stack-pointer subrange pruning

This commit adds a test case to validate widened stack pointer subrange
relations in regsafe(). It checks that a widened stack pointer with range
[-64, -8] is treated as a superset of [-56, -8].

> diff --git a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
> index 3e0bea9819cab..5da8b9a6018cf 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_stack_ptr.c
> @@ -586,4 +586,37 @@ __naked void stack_check_size_512_with_may_goto(void)
>  }
>  #endif
>  
> +/* Verify that old PTR_TO_STACK state is considered a super-set of
> + * new PTR_TO_STACK state when new variable range is a sub-range
> + * of the old range, e.g. old [-72, -16] vs new [-64, -16].
> + */

[Severity: Low]
Does this violate the BPF subsystem multi-line comment style? The BPF 
subsystem requires multi-line comments to have the opening /* on its 
own line, with the text starting on the next line.

> +SEC("socket")
> +__success
> +__log_level(2)
> +__flag(BPF_F_TEST_STATE_FREQ)
> +/* r7 is widened to [-72, -16] at the loop header (insn 3),
> + * the loop body sees [-64, -8] after 'r7 += 8'
> + */

[Severity: Low]
Should this comment also be updated to start the text on the line following
the opening /* to match the subsystem standards?

> +__msg("loop header at 3, widening r7 to -72..-16 step 8")
> +__msg("R7=fp(smin=smin32=-64,smax=smax32=-8")
> +/* back-edge state is clamped to the remaining iterations and pruned
> + * at the header, because [-64, -16] is within the widened [-72, -16]
> + */

[Severity: Low]
Would it be better to adjust this comment block as well to ensure the opening
/* is placed on a separate line?

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=28

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 03/36] bpf: summarize live stack slots in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 03/36] bpf: summarize live " Eduard Zingerman
@ 2026-09-26 14:33   ` sashiko-bot
  2026-09-27 20:26   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:33 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: missing helper function promised in commit message
--

commit 86d6b186246ae90e5b5fd0f2e5d9c4148d983e50
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: summarize live stack slots in insn_aux_data

This commit precomputes the union of current-frame liveness masks from all
analyzed function instances for each instruction. It stores this summary in
insn_aux_data to help the scalar evolution analysis skip tracking dead stack
slots.

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index c6d617581e841..7f31ce5ea6b7b 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -667,6 +667,7 @@ struct bpf_insn_aux_data {
>  	 * that this instruction may write to.
>  	 */
>  	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
> +	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);

[Severity: Medium]
The commit message explicitly states: "Store the summary in insn_aux_data
and expose it via a helper."

Could the commit message be updated, or is there a missing helper function?
No helper function is introduced in this patch to expose this new field,
and it appears consumers in later patches access the bitmap directly via
test_bit rather than going through a helper.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=3

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear()
  2026-09-26 14:20 ` [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear() Eduard Zingerman
@ 2026-09-26 14:34   ` sashiko-bot
  2026-09-29 21:46     ` Eduard Zingerman
  0 siblings, 1 reply; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:34 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: cnum64_intersect_linear over-expands intervals that wrap the signed boundary
--

commit 07e53bba30487c590757dc2ec233d2811a607b9a
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: add cnum64_intersect_linear()

This commit introduces cnum64_intersect_linear() to intersect a cnum64
interval with integers congruent to 'base' modulo 'step'. It tightens the
interval using signed bounds to keep residues consistent. This is a building
block for tracking scalar registers with values known to lie on a line.

> diff --git a/kernel/bpf/cnum.c b/kernel/bpf/cnum.c
> index 86142cb2aee5f..2bff2c7e0cdf5 100644
> --- a/kernel/bpf/cnum.c
> +++ b/kernel/bpf/cnum.c
[ ... ]
> @@ -118,3 +119,38 @@ struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b)
>  	}
>  	return t;
>  }
> +
> +/* Intersect 'in' with the set of integers defined by equation 'base + step * k'. */
> +struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step)
> +{
> +	s64 smin = cnum64_smin(in);
> +	s64 smax = cnum64_smax(in);
> +	s64 lo, hi;
> +	u16 d;
> +
> +	if (step <= 1 || cnum64_is_empty(in))
> +		return in;
> +	/*
> +	 * Round smin up to the next value congruent to 'base' modulo 'step',
> +	 * i.e. increase smin by d = (base - smin) mod step:
> +	 *
> +	 *                 |<---- d ---->|
> +	 *     |-----------|=============|...
> +	 * base+step*k    smin       base+step*(k+1)
> +	 */
> +	d = imod(base - imod(smin, step), step);
> +	if ((u64)smax - (u64)smin < d)
> +		return CNUM64_EMPTY;
> +	lo = smin + d;
> +	/*
> +	 * Round smax down to the previous value congruent to 'base' modulo 'step',
> +	 * i.e. decrease smax by d = (smax - base) mod step:
> +	 *
> +	 *     |<--- d --->|
> +	 *  ...|===========|-------------|
> +	 * base+step*k    smax       base+step*(k+1)
> +	 */
> +	d = imod(imod(smax, step) - base, step);
> +	hi = smax - d;
> +	return cnum64_from_srange(lo, hi);

[Severity: Medium]
Does this code improperly expand intervals that wrap the signed boundary?

For such intervals, cnum64_smin() returns S64_MIN and cnum64_smax() returns
S64_MAX. The function then calculates 'lo' near S64_MIN and 'hi' near
S64_MAX, returning cnum64_from_srange(lo, hi).

This replaces a potentially small wrapping interval with a massive contiguous
range spanning nearly the entire 64-bit space, violating the intersection
property (A ∩ B ⊆ A). While this over-approximation is sound, it could cause
the verifier to lose precision and falsely reject valid BPF programs.

> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=8

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 14/36] bpf: representation for intervals with steps
  2026-09-26 14:20 ` [PATCH bpf-next 14/36] bpf: representation for intervals with steps Eduard Zingerman
@ 2026-09-26 14:35   ` sashiko-bot
  0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:35 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] bpf: verifier linear equation tracking ignores integer wrap-around leading to OOB access
--

commit e753188b2ce55741938a6a63ee79699c312fb712
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: representation for intervals with steps

Extend scalar register tracking with a linear base + step * k description
to allow the verifier to reason about strided values, updating scalar
ALU operations to maintain these constraints.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 79aa924861548..89b1a0aa3a25b 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -2165,10 +2173,16 @@ static void deduce_bounds_64_from_32(struct bpf_reg_state *reg)
>  	reg->r64 = cnum64_cnum32_intersect(reg->r64, reg->r32);
>  }
>  
> +static void deduce_bounds_64_from_step(struct bpf_reg_state *reg)
> +{
> +	reg->r64 = cnum64_intersect_linear(reg->r64, reg->base, reg->step);
> +}

[Severity: Critical]
If the base and step variables become desynchronized from the true runtime
value during ALU operations, could this intersection incorrectly shrink the
register bounds?

If the linear representation is mathematically flawed due to integer
wrap-around, the verifier might deduce bounds that exclude the actual runtime
value.

Is it possible this leads to out-of-bounds accesses bypassing checks?

[ ... ]

> @@ -15589,6 +15615,50 @@ static int sanitize_check_bounds(struct bpf_verifier_env *env,
>  	return 0;
>  }
>  
> +static void scalar_step_add(struct bpf_reg_state *dst_reg,
> +			    const struct bpf_reg_state *a,
> +			    const struct bpf_reg_state *b)
> +{
> +	u16 base, step;
> +
> +	/* If either 'a' or 'b' is a constant, update the base/step for the counterpart. */
> +	if (tnum_is_const(b->var_off)) {
> +		step = a->step;
> +		base = imod((s64)a->base + (s64)b->var_off.value, step);
> +	} else if (tnum_is_const(a->var_off)) {
> +		step = b->step;
> +		base = imod((s64)b->base + (s64)a->var_off.value, step);

[Severity: Critical]
Does this computation account for 32-bit or 64-bit integer wrap-around?

If a register is initialized with a non-power-of-2 stride (for example, base=0
and step=3), and then undergoes an addition that wraps around 32 bits, the
true remainder modulo 3 changes because 2^32 is not divisible by 3.

Since this computes imod() on the 64-bit sum, the tracking might blindly
ignore boundary wrap-around, leading to an inaccurate linear representation.

Furthermore, in adjust_scalar_min_max_vals(), when alu32 is true, the function
calls zext_32_to_64() which zero-extends the 32-bit result without
invalidating the 64-bit base and step tracking.

Could this allow a crafted BPF program to bypass bounds checks?

> +	} else {
> +		step = 1;
> +		base = 0;
> +	}
> +	dst_reg->base = base;
> +	dst_reg->step = step;
> +}
> +
> +static void scalar_step_mul(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg)
> +{
> +	u64 amount = src_reg->var_off.value;
> +
> +	if (tnum_is_const(src_reg->var_off) && (s64)amount >= 0 &&
> +	    !check_mul_overflow(dst_reg->step, amount, &dst_reg->step) &&

[Severity: Critical]
This code checks if the step value overflows a u16, but it does not appear to
verify if the register actual value multiplication wraps around 64 bits or
32 bits.

If the underlying integer multiplication wraps, the linear equation might
no longer accurately reflect the actual value of the register.

Could this also lead to an incorrect linear equation being propagated?

> +	    dst_reg->step != 0)
> +		dst_reg->base = (dst_reg->base * amount) % dst_reg->step;
> +	else
> +		reg_step_reset(dst_reg);
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=14

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range()
  2026-09-26 14:20 ` [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range() Eduard Zingerman
@ 2026-09-26 14:36   ` sashiko-bot
  0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:36 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: build failure due to missing step and base fields in struct bpf_reg_state
--

commit b21e96677bd5c8c118c6f3ee527144b5d83b0c82
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: add bpf_set_reg_range()

A utility function for setting a register's range and step information.
Used by SCEV widening and clamping logic further in the series.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 03dbc0e003986..f2d44e026e4bd 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -16934,6 +16934,18 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
>  	return 0;
>  }
>  
> +int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
> +		      struct cnum64 range, u16 step)
> +{
> +	reg->r64 = range;
> +	reg->r32 = CNUM32_UNBOUNDED;
> +	reg->step = step;
> +	reg->base = imod(cnum64_smin(range), step);

[Severity: Low]
Does this code cause a build failure?

It appears that the `step` and `base` fields do not exist in
`struct bpf_reg_state` at this point in the commit history, which would
break `git bisect`.

While this patch splitting error is resolved in the subsequent commit
'[PATCH bpf-next 14/36] bpf: representation for intervals with steps' which
adds the missing fields, each patch in a series must compile independently.

> +	reg->var_off = tnum_unknown;
> +	reg_bounds_sync(reg); /* this should infer the tnum alignment */
> +	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
> +}
> +
>  /* check validity of 32-bit and 64-bit arithmetic operations */

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=9

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers
  2026-09-26 14:20 ` [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
@ 2026-09-26 14:37   ` sashiko-bot
  2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:37 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: varying offset accesses to nested flexible arrays are incorrectly rejected
--

commit 9dc337511b42c543566061ccac0810b0e6e74dae
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: varying offset access support for PTR_TO_BTF_ID pointers

This commit relaxes the verifier requirements for reads like p->arr[i].field 
where p is a PTR_TO_BTF_ID, allowing i to be a non-constant value. It 
introduces bounds checking based on the access register's min and max bounds, 
along with alignment checks.

> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c
[ ... ]
> @@ -7846,6 +7905,54 @@ int btf_struct_access(struct bpf_verifier_log *log,
>  	} while (t);
>  
>  	return -EINVAL;
> +
> +check_variable_offset:
> +	if (min_off == max_off)
> +		return ret;
> +
> +	/* Find an offset at which access would go to a flexible array tail (if any). */
> +	t = btf_type_skip_modifiers(btf, reg->btf_id, NULL);
> +	member = find_flex_member(btf, t);
> +	flex_off = member ? __btf_member_bit_offset(t, member) / 8 : S64_MAX;

[Severity: Medium]
Does this code handle flexible arrays embedded within nested structs?

When a BPF program attempts a variable-offset access (for example,
ptr->nested_struct.flex_array[i]) on a PTR_TO_BTF_ID, find_flex_member()
is called strictly on the top-level struct type ID.

Because find_flex_member() does not recursively evaluate nested structs,
it returns NULL in this case. This causes flex_off to evaluate to S64_MAX,
which later makes the array bounds check fail since the flexible array
has a size of 0, resulting in -EINVAL being returned.

Can this be updated so the verifier doesn't incorrectly reject valid programs
performing varying-offset accesses to nested flexible arrays?

> +
> +	/*
> +	 * If this is a varying offset access, the step recorded within a register
> +	 * should correspond to one of the arrays visited while walking.
> +	 */
[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=15

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops
  2026-09-26 14:20 ` [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops Eduard Zingerman
@ 2026-09-26 14:38   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:38 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] bpf: NULL pointer dereference in compute_insn_scevs() for loops with exits_overflow
- [Medium] bpf: Incorrect SCEV modeling of BPF_END no-op conversions preserves truncated bits
--

commit 3c85c59367498a632c83889db787693adfac36e5
Author: Eduard Zingerman <eddyz87@gmail.com>
Subject: bpf: compute scalar evolution expressions for loops

This patch assigns algebraic expressions to loop variables (registers and
stack spills) to describe how their values evolve across iterations. The
analysis proceeds in two phases: computing expressions for one symbolic
iteration and converting backedge updates to recurrences.

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> new file mode 100644
> index 0000000000000..396b0c776dfae
> --- /dev/null
> +++ b/kernel/bpf/scev.c
[ ... ]
> @@ -713,29 +713,25 @@ static int transfer(struct bpf_verifier_env *env, struct env *e, int idx)
[ ... ]
> +		case BPF_END:
> +			switch (insn->imm) {
> +			case 16: op = BSWAP16; break;
> +			case 32: op = BSWAP32; break;
> +			case 64: op = BSWAP64; break;
> +			default:
> +				goto mark_dst_unknown;
> +			}
> +
> +			if (class == BPF_ALU && x_or_k == BPF_TO_LE && insn->off == 0 && little_endian)
> +				op = 0; /* little-endian to little-endian is noop */

[Severity: Medium]
Can this skip the required 16-bit truncation semantics?

When a BPF_END instruction is a 16-bit endianness conversion (imm == 16)
and acts as a no-op on the current architecture, op is set to 0.

Because op is 0, the explicit truncation operation (like ZEXT16) is skipped
below. Only ZEXT32 is applied to the register:

kernel/bpf/scev.c:transfer() {
    id = reg2expr[dst];
    if (op)
        id = expr1(scev, op, reg2expr[dst]);

    if (class == BPF_ALU)
        id = expr1(scev, ZEXT32, id);
    ...
}

The verifier and hardware will still truncate the register to 16 bits,
zeroing bits 16-63. By skipping this in the algebraic model, does it
incorrectly preserve bits 16-31 of the original register?

> +			else if (class == BPF_ALU && x_or_k == BPF_TO_BE && insn->off == 0 && !little_endian)
> +				op = 0; /* big-endian to big-endian is noop */
> +			else if (class == BPF_ALU64 && x_or_k == 0 && insn->off == 0)
> +				/* always swap */;
> +			else
> +				goto mark_dst_unknown;
> +
> +			id = reg2expr[dst];
> +			if (op)
> +				id = expr1(scev, op, reg2expr[dst]);
> +
> +			if (class == BPF_ALU)
> +				id = expr1(scev, ZEXT32, id);
> +
> +			return replace_reg(scev, e, dst, id);
[ ... ]
> @@ -1172,13 +1172,13 @@ static void mark_latches(struct bpf_verifier_env *env)
> +	for (i = 0; i < len; i++) {
> +		loop = aux[i].loop;
> +		if (!loop)
> +			continue;
> +		aux[i].need_scev = true;
> +		if (loop->irreducible)
> +			continue;

[Severity: High]
Does this code result in a NULL pointer dereference when a loop has an
exits_overflow condition?

Loops with exits_overflow are not skipped here (unlike irreducible loops),
so their latches are marked with aux[latch].need_scev = true.

However, compute_scev_for_loop() bails out early for exits_overflow loops
and does not allocate a loop environment for the latch:

kernel/bpf/scev.c:compute_scev_for_loop() {
    if (cur_loop->irreducible || cur_loop->exits_overflow) {
        header_env = get_loop_env(scev, cur_header, cur_header);
        ...
        return 0;
    }
    ...
}

Later, when bpf_compute_scev() visits the latch, it calls
compute_insn_scevs(), passing the result of find_loop_env(). Since the
environment was never created, find_loop_env() returns NULL.

This results in einsn being NULL in compute_insn_scevs():

kernel/bpf/scev.c:compute_insn_scevs() {
    ...
    for (reg = 0; reg < REGS_NUM; reg++) {
        id = einsn->reg2expr[reg];
}

Should exits_overflow loops be skipped here similarly to irreducible loops?

> +		for (j = 0; j < loop->backedges_cnt; j++) {
> +			latch = loop->backedges[j].latch;
> +			if (latch >= 0)
> +				aux[latch].need_scev = true;
> +		}
> +	}
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=24

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops
  2026-09-26 14:20 ` [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops Eduard Zingerman
@ 2026-09-26 14:42   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:42 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] bpf: verifier bypass due to signed algebraic inversion of unsigned loop conditions
--

commit 457e66d0b00c63117d606a731e93743130f1e826
Author: Eduard Zingerman <eddyz87@gmail.com>
Subject: bpf: use SCEV to widen bounded loops

This commit introduces scalar evolution (SCEV) to the BPF verifier to
estimate loop iteration bounds and widen induction variables. This allows
the verifier to support bounded loops with linear latches and bounds.

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> --- a/kernel/bpf/scev.c
> +++ b/kernel/bpf/scev.c
[ ... ]
> +static bool compute_max_iters(struct bpf_verifier_env *env,
> +			      struct bpf_func_state *st,
> +			      struct linear_latch *latch,
> +			      struct bpf_loop_iters *iters)
> +{
[ ... ]
> +	if (step == 0)
> +		return false;
> +	if ((s64)step == S64_MIN)
> +		return false;
> +	if ((s64)step < 0) {
> +		/* Multiply both sides of the equation by -1, e.g. -2*i > -3 becomes 2*i < 3 */
> +		op = bpf_flip_opcode(op);
> +		step = -step;
> +		swap(bound, initial);
> +	}

[Severity: Critical]
Does this logic improperly apply a signed algebraic identity to unsigned
comparisons?

Looking at compute_max_iters() in kernel/bpf/scev.c, if a BPF program
constructs a loop with an unsigned continuation condition, initializes
the register to a positive value, and decrements it, the algebraic
inversion treats the counter as if it crosses into negative space.

Could this incorrectly calculate a small iteration count, while at runtime
the unsigned subtraction wraps around to U64_MAX and allows the loop to
continue?

If this wraps, it appears the verifier might widen the register bounds to a
narrow range and prune branches targeting the wrapped values as dead code,
potentially allowing unverified payloads to execute.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=25

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking
  2026-09-26 14:20 ` [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
@ 2026-09-26 14:46   ` sashiko-bot
  2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-26 14:46 UTC (permalink / raw)
  To: Eduard Zingerman; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: bpf_needs_fixed_stack_off misses dynptr consumers, rejecting valid loops
--

commit a3ca958d64fafcbcf31deb9929c54958bdb72868
Author: Eduard Zingerman <eddyz87@gmail.com>

bpf: avoid widening registers that hinder exact stack-slot tracking

This commit adds logic to prevent loop-widening for registers that are
used as base addresses for stack access instructions that require exact
tracking (such as spills, fills, and calls constructing stack objects).

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 9f56ffe09ebac..327bfc00da5f5 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -14087,6 +14087,46 @@ static bool kfunc_spin_allowed(struct bpf_verifier_env *env, s32 func_id, s16 of
[ ... ]
> +bool bpf_needs_fixed_stack_off(struct bpf_verifier_env *env, int insn_idx)
> +{
> +	const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
> +	u32 *flags, btf_id;
> +
> +	if (bpf_helper_call(insn)) {
> +		return insn->imm == BPF_FUNC_dynptr_from_mem ||
> +		       insn->imm == BPF_FUNC_ringbuf_reserve_dynptr;
> +	}

[Severity: Medium]
Does this logic miss helper functions that consume dynptrs (like
BPF_FUNC_dynptr_read, BPF_FUNC_dynptr_write, or BPF_FUNC_ringbuf_submit_dynptr)?

Since process_dynptr_func() relies on dynptr_get_spi(), which in turn
requires constant stack offsets via stack_slot_obj_get_spi(), wouldn't
omitting these dynptr consumers here cause a regression? If the verifier
widens the loop variable used for a dynptr array index, the stack offset
becomes non-constant and the verifier will erroneously reject valid programs
that read or write stack-based dynptrs in a loop.

> +
> +	/* vmlinux kfuncs only */
> +	if (!bpf_pseudo_kfunc_call(insn) || insn->off != 0)
> +		return false;
> +	btf_id = insn->imm;
[ ... ]
> +	if (btf_id == special_kfunc_list[KF_bpf_dynptr_from_skb] ||
> +	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_xdp] ||
> +	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_skb_meta] ||
> +	    btf_id == special_kfunc_list[KF_bpf_dynptr_from_file])
> +		return true;

[Severity: Medium]
Similarly, does this miss kfuncs that consume or clone dynptrs (such as
KF_bpf_dynptr_clone and KF_bpf_dynptr_slice)?

If a loop variable is used to compute a dynptr address passed to these
kfuncs, the loop variable will be widened, leading to the same verification
regression when stack_slot_obj_get_spi() is called with a non-constant stack
offset.

> +
> +	return false;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926-scev-minimal-rebase-v1-0-c8e5ab5ba79f@gmail.com?part=26

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 01/36] bpf: track may_write flags in liveness
  2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
@ 2026-09-26 15:51   ` Alexei Starovoitov
  2026-09-27  8:43     ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-26 15:51 UTC (permalink / raw)
  To: Eduard Zingerman, bpf, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor

On Sat, Sep 26, 2026 at 07:19 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> Compared to must_write, may_write includes any possibly modified slot,
> including partial writes, and does not kill liveness.

It doesn't include the writes done by helpers and kfuncs.
bpf_helper_stack_access_bytes() returns -size only for MEM_UNINIT args.
bpf_fib_lookup() has ARG_PTR_TO_MEM | MEM_WRITE for params, so
record_arg_access() gets +size and the call has 'use' and no 'may_def',
though the helper writes into params.
Same for raw uninit mem when !allow_uninit_stack and for every
kfunc mem arg that is not __uninit.

> -__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24{{$}}")
> +__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24 may_def: fp0-24{{$}}")
>  __msg(" 9: (85) call bpf_iter_num_next{{.*}}         ; use: fp0-24{{$}}")

bpf_iter_num_next() is 'use' only too.

Returning a write as a read was the safe direction for liveness.
For may_write it's the opposite.
What does SCEV do with a slot that has no may_def inside the loop?

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
@ 2026-09-26 15:51   ` Alexei Starovoitov
  2026-09-29 20:12     ` Eduard Zingerman
  2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-26 15:51 UTC (permalink / raw)
  To: Eduard Zingerman, bpf, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor

On Sat, Sep 26, 2026 at 07:20 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> @@ -662,6 +662,11 @@ struct bpf_insn_aux_data {
>  	};
>  	struct btf_struct_meta *kptr_struct_meta;
>  	u64 map_key_state; /* constant (32 bit) key tracking for maps */
> +	/*
> +	 * Per-instruction summary of stack slots in the current frame
> +	 * that this instruction may write to.
> +	 */
> +	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);

That's 32 bytes per insn and the next patch adds 32 more,
for every prog whether it has loops or not.
FM_MAY_WRITE in patch 1 makes every frame_masks a third bigger too.
commit 481ceda77aeb sized the liveness masks by the stack the frame
uses to avoid exactly that.

bpf_may_write_mask() is the only accessor.
Can the summary stay in liveness.c, as wide as the subprog's stack,
and only for insns with scc != 0 ?

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 18/36] bpf: compute immediate dominators
  2026-09-26 14:20 ` [PATCH bpf-next 18/36] bpf: compute immediate dominators Eduard Zingerman
@ 2026-09-26 15:54   ` Alexei Starovoitov
  2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-26 15:54 UTC (permalink / raw)
  To: Eduard Zingerman, bpf, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor

On Sat, Sep 26, 2026 at 07:20 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> +static int idoms_intersect(struct bpf_verifier_env *env, int a, int b)
> +{
> +	int *postorder_nums = env->cfg.postorder_nums;
> +	int *idoms = env->idoms;
> +
> +	while (a != b) {
> +		while (postorder_nums[a] < postorder_nums[b]) {
> +			a = idoms[a];
> +		}

This walks the idom chain one insn at a time.
For a prog that is a long run of
  if (err) goto out;
'out' has N predecessors and compute_subprog_idoms() calls this
for each of them. The k-th call walks back over k branches.
That's N^2/2 steps, and the outer loop runs at least twice.
All of it before do_check(), for every prog, loops or not.
Same issue as Daniel's commit 3852e19cb954 fixed just yesterday.

SCEV needs dominators only inside loops.
Can it be computed for insns with scc != 0 only ?
and probably a bound for a big loop body.

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 01/36] bpf: track may_write flags in liveness
  2026-09-26 15:51   ` Alexei Starovoitov
@ 2026-09-27  8:43     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-27  8:43 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor

On Sat, 2026-09-26 at 15:51 +0000, Alexei Starovoitov wrote:
> On Sat, Sep 26, 2026 at 07:19 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> > Compared to must_write, may_write includes any possibly modified slot,
> > including partial writes, and does not kill liveness.
> 
> It doesn't include the writes done by helpers and kfuncs.
> bpf_helper_stack_access_bytes() returns -size only for MEM_UNINIT args.
> bpf_fib_lookup() has ARG_PTR_TO_MEM | MEM_WRITE for params, so
> record_arg_access() gets +size and the call has 'use' and no 'may_def',
> though the helper writes into params.
> Same for raw uninit mem when !allow_uninit_stack and for every
> kfunc mem arg that is not __uninit.
> 
> > -__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24{{$}}")
> > +__msg(" 6: (85) call bpf_iter_num_new{{.*}}          ; def: fp0-24 may_def: fp0-24{{$}}")
> >  __msg(" 9: (85) call bpf_iter_num_next{{.*}}         ; use: fp0-24{{$}}")
> 
> bpf_iter_num_next() is 'use' only too.
> 
> Returning a write as a read was the safe direction for liveness.
> For may_write it's the opposite.
> What does SCEV do with a slot that has no may_def inside the loop?

Missing may_write means that SCEV for the slot won't be invalidated on
indirect write that touches it. Which would lead to incorrect SCEV
expression.

This is a good point, I need to fix this bug and collect updated stats.
I think it's the main take-away from the reviews so far.

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
@ 2026-09-27 20:26   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ezingerman, ast, andrii, daniel, martin.lau, eddyz87,
	yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1940 bytes --]

>     bpf: record basic-block ends in insn_aux_data
>
>     SCEV needs information about the program's basic-block structure.
>     Piggyback on compute_predecessors() and set
>     insn_aux_data->bb_end for an instruction if:
>     - it has multiple successors, or
>     - it has multiple predecessors, or
>     - it is a jump instruction.
>
>     Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

The commit message says bb_end is set for an instruction if "it has
multiple predecessors", but the code does something different:

kernel/bpf/loops.c:compute_predecessors() {
    ...
    if (succ->cnt > 1 ||
        (succ->cnt == 1 && num_preds[succ->items[0]] > 1) ||
        is_cfg_jump(insn))
        aux[i].bb_end = true;
    ...
}

The second condition marks instruction i when i's single *successor* has
more than one predecessor. It does not look at num_preds[i].

An instruction with several predecessors is a join point, so it starts a
basic block rather than ending one. The code gets this right: the
instruction just before a join point is marked as the block end, and the
SCEV walker added later in the series depends on that, because it stops
at aux[idx].bb_end and sends the state through join_successor().

If the rule were implemented the way the changelog words it, the join
point would get bb_end instead, and the walker would enter the join
block from one predecessor without merging state from the others.

Could the wording be "its only successor has multiple predecessors"?

The new struct field bool bb_end in struct bpf_insn_aux_data has no
comment, so the changelog is currently the only description of what
bb_end means. A one-line comment on the field would also help.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations
  2026-09-26 14:20 ` [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations Eduard Zingerman
@ 2026-09-27 20:26   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1710 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 706fdefbc07a..075a28de2a04 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -1858,4 +1858,7 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env);
>  int bpf_do_misc_fixups(struct bpf_verifier_env *env);
>  int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
>
> +int bpf_flip_opcode(u32 opcode);
> +u8 bpf_rev_opcode(u8 opcode);
> +
>  #endif /* _LINUX_BPF_VERIFIER_H */
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 1ec44d72a346..79aa92486154 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -17565,7 +17563,7 @@ static int is_branch_taken(struct bpf_verifier_env *env, struct bpf_reg_state *r
>  /* Opcode that corresponds to a *false* branch condition.
>   * E.g., if r1 < r2, then reverse (false) condition is r1 >= r2
>   */
> -static u8 rev_opcode(u8 opcode)
> +u8 bpf_rev_opcode(u8 opcode)
>  {
>  	switch (opcode) {
>  	case BPF_JEQ:		return BPF_JNE;

The patch renames rev_opcode() to bpf_rev_opcode(), but there's still
a comment in regs_refine_cond_op() that references the old name:

kernel/bpf/verifier.c:regs_refine_cond_op() {
    ...
    case BPF_JSET | BPF_X: /* reverse of BPF_JSET, see rev_opcode() */
        if (!is_reg_const(reg2, is_jmp32))
            swap(reg1, reg2);
    ...
}

Should this comment be updated to reference bpf_rev_opcode() instead?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning
  2026-09-26 14:20 ` [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
  2026-09-26 14:32   ` sashiko-bot
@ 2026-09-27 20:26   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1687 bytes --]

> selftests/bpf: test for stack-pointer subrange pruning
>
> Validate changes from:
> "bpf: allow subrange relations for PTR_TO_STACK in regsafe()".
> Add a test case that checks that a widened stack pointer with range
> [-64, -8] is considered a superset of a stack pointer with range
> [-56, -8].
>
> Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

Do the ranges in the commit message match the test?

The changelog says the test checks that a widened stack pointer with range
[-64, -8] is a superset of one with range [-56, -8]. Those numbers appear to
be the example from the regsafe() commit 0516ad0f46c2 ("bpf: allow subrange
relations for PTR_TO_STACK in regsafe()").

The test itself starts r7 at fp-72, and the loop header runs 8 times (r6 =
0..7), so the verifier widens r7 at the header to [-72, -16]. The back-edge
state is clamped to [-64, -16] and pruned against that. The test's own
comment and __msg strings say this:

>  * of the old range, e.g. old [-72, -16] vs new [-64, -16].
>  __msg("loop header at 3, widening r7 to -72..-16 step 8")
>  __msg("loop header at 3, clamping r7 to -64..-16 step 8")

I traced this against bpf_widen_scev_regs()/linear_bounds() (-72 + 8*(8-1)
= -16) and bpf_clamp_scev_regs() (cnum64_intersect([-64,-8], [-72,-16]) =
[-64,-16]). The comparison the test exercises is [-72, -16] vs [-64, -16],
not [-64, -8] vs [-56, -8].

Could the changelog be updated to match the ranges the test checks?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists
  2026-09-26 14:20 ` [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
@ 2026-09-27 20:26   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ezingerman, ast, andrii, daniel, martin.lau, eddyz87,
	yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3083 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index d8142f1b2347..f2497236542c 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -1916,4 +1916,21 @@ int bpf_compute_loops(struct bpf_verifier_env *env);
>  int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx);
>  bool bpf_is_nested_loop(struct bpf_verifier_env *env, int inner_header, int outer_header);
>
> +/*
> + * Simple binary heap implementation as described by
> + * https://en.wikipedia.org/wiki/Binary_heap
> + */
> +struct bpf_min_heap {
> +	int (*compare)(int, int, void *); /* ordering function for @elements */
> +	int *elements; /* min-heap ordered by @compare */
> +	void *arg; /* 3rd argument passed to @compare */
> +	int capacity;
> +	int count;
> +};
> +
> +void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg);
> +void bpf_min_heap_free(struct bpf_min_heap *heap);
> +int bpf_min_heap_push(struct bpf_min_heap *heap, int elt);
> +bool bpf_min_heap_pop(struct bpf_min_heap *heap, int *elt);
> +
>  #endif /* _LINUX_BPF_VERIFIER_H */

[ ... ]

> diff --git a/kernel/bpf/heap.c b/kernel/bpf/heap.c
> new file mode 100644
> index 000000000000..029f9217b872
> --- /dev/null
> +++ b/kernel/bpf/heap.c
> @@ -0,0 +1,87 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
> +
> +#include <linux/bpf_verifier.h>
> +
> +/* Indexes for binary tree encoded as an array */
> +static inline int left_child(int i) { return 2 * i + 1; }
> +static inline int right_child(int i) { return 2 * i + 2; }
> +static inline int parent(int i) { return (i - 1) / 2; }
> +
> +static inline int greater(struct bpf_min_heap *heap, int a, int b)
> +{
> +	return heap->compare(a, b, heap->arg) > 0;
> +}
> +
> +void bpf_min_heap_init(struct bpf_min_heap *heap, int (*compare)(int, int, void *), void *arg)
> +{
> +	memset(heap, 0, sizeof(*heap));
> +	heap->compare = compare;
> +	heap->arg = arg;
> +}
> +
> +void bpf_min_heap_free(struct bpf_min_heap *heap)
> +{
> +	kfree(heap->elements);
> +	heap->elements = NULL;
> +	heap->capacity = 0;
> +	heap->count = 0;
> +}
> +
> +int bpf_min_heap_push(struct bpf_min_heap *heap, int elt)
> +{
> +	int new_capacity, i;
> +	int *elements;
> +	void *tmp;
> +
> +	if (heap->count == heap->capacity) {
> +		new_capacity = heap->capacity ? heap->capacity * 2 : 16;
> +		tmp = krealloc(heap->elements,
> +			       sizeof(*heap->elements) * new_capacity,
> +			       GFP_KERNEL_ACCOUNT);

This isn't a bug, but could the SCEV worklist use the existing
include/linux/min_heap.h helpers instead? The worklist is capped at
env->prog->len entries and min_heap's less() callback already takes a
context pointer. If there is a reason it doesn't fit, could the changelog
say what it is?

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function
  2026-09-26 14:20 ` [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
@ 2026-09-27 20:26   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	emil, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song,
	mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1463 bytes --]

> Tests for multiple analyses performed by the verifier need the
> verifier log to contain analysis results alongside the program
> disassembly. This patch moves the log-level-2 program dump from
> bpf_compute_live_registers() to a standalone function called from
> bpf_check(), in order to provide a common logging function for such
> analyses (and thus avoid printing program disassembly multiple times).

The commit message describes this as moving the dump, but log_program()
also changes the user-visible log output in two ways.

First, the header changes from "Live regs before insn:" to "Program dump
(scc? insn#: live_regs_before):". This is why the selftest in
verifier_live_stack.c needs the __msg update from "Live regs before insn:"
to "Program dump".

Second, log_program() adds a verbose_linfo(env, i, "    ; ") call that
interleaves BTF source-line annotations into the dump:

kernel/bpf/verifier.c:log_program() {
    ...
    for (i = 0; i < insn_cnt; ++i) {
        verbose_linfo(env, i, "    ; ");
        ...
    }
}

The removed code in bpf_compute_live_registers() did not print these
source-line annotations.

Could the commit message mention the header rename and the added source-line
annotations?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 03/36] bpf: summarize live stack slots in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 03/36] bpf: summarize live " Eduard Zingerman
  2026-09-26 14:33   ` sashiko-bot
@ 2026-09-27 20:26   ` bot+bpf-ci
  2026-09-29 18:16     ` Eduard Zingerman
  1 sibling, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:26 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2518 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index c6d617581e84..7f31ce5ea6b7 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -667,6 +667,7 @@ struct bpf_insn_aux_data {
>  	 * that this instruction may write to.
>  	 */
>  	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
> +	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);

This isn't a bug, but would a short comment on live_stack_before make it
clearer what the field holds?  Nearly every neighbouring field like
live_regs_before and scc has its own comment.

>  	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
>  	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
>  	bool nospec; /* do not execute this instruction speculatively */

[ ... ]

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index c871744ca5a8..b6b7fd479569 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
> @@ -729,8 +729,9 @@ static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask,
>  }
>
>  /*
> - * Precompute, for each instruction, the OR of may_write masks over its top
> - * frame across all func_instances reaching it, stash it in the insn_aux_data.
> + * Precompute, for each instruction, the OR of may_write and live_before masks
> + * over its top frame across all func_instances reaching it, stash them in the
> + * insn_aux_data.
>   */
>  static void compute_may_write_masks(struct bpf_verifier_env *env)
>  {
> @@ -746,9 +747,12 @@ static void compute_may_write_masks(struct bpf_verifier_env *env)
>  		if (!fm)
>  			continue;
>  		nbits = frame_mask_bits(fm);
> -		for (i = 0; i < instance->insn_cnt; i++)
> +		for (i = 0; i < instance->insn_cnt; i++) {
>  			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
>  					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
> +			half_spis_to_slots(aux[instance->subprog_start + i].live_stack_before,
> +					   rel_mask(fm, i, FM_LIVE_BEFORE), nbits);
> +		}
>  	}
>  }

This isn't a bug, but would a more general name for compute_may_write_masks()
make the function's purpose clearer?  The comment now describes computing both
may_write and live_before masks, but the name only mentions may_write.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints
  2026-09-26 14:20 ` [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
@ 2026-09-27 20:27   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:27 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1133 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 76833c2ad452..08e0b49841df 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -502,7 +502,9 @@ struct bpf_verifier_state {
>  	bool speculative;
>  	bool in_sleepable;
>
> -	/* first and last insn idx of this verifier state */
> +	/* First and last insn idx of this verifier state.
> +	 * last_insn_idx is -1 if no instructions have been executed yet.
> +	 */

This isn't a bug, but should the multi-line comment format follow the BPF
subsystem style for include/linux/bpf*.h files?  The opening /* should be
on a line by itself:

	/*
	 * First and last insn idx of this verifier state.
	 * last_insn_idx is -1 if no instructions have been executed yet.
	 */

>  	u32 first_insn_idx;
>  	u32 last_insn_idx;
>  	/* if this state is a backedge state then equal_state

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
@ 2026-09-27 20:27   ` bot+bpf-ci
  2026-09-29 20:21     ` Eduard Zingerman
  0 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:27 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2243 bytes --]

> bpf: summarize regs that may hold a frame pointer in insn_aux_data
>
> SCEV needs to know which registers might be stack pointers at a
> particular instruction. This information is used to invalidate SCEV
> expressions for slots that might be overwritten by indirect writes.
>
> liveness.c:compute_may_write_masks() already tracks this information.
> This commit modifies it to save the information in insn_aux_data for
> further usage.

Are these two claims accurate?

First, the changelog says compute_may_write_masks() already tracks which
registers hold frame pointers. Looking at that function, it only folds the
per-instance FM_MAY_WRITE and FM_LIVE_BEFORE frame masks into
aux->may_write_mask and aux->live_stack_before:

    half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
                       rel_mask(fm, i, FM_MAY_WRITE), nbits);

The register FP identity comes from the arg_track dataflow in
compute_subprog_args(), and that is where the new hook sits:

    record_stack_ptrs(env, idx, at_in[i]);

The patch does not touch compute_may_write_masks().

Second, the changelog states the information is used to invalidate SCEV
expressions for slots that might be overwritten by indirect writes. Looking
at the series up to commit 4a6395b27755, reset_scevs_at_indirect_writes()
in kernel/bpf/scev.c uses bpf_may_write_mask(env, idx) for that purpose,
not stack_ptrs.

The only reader of stack_ptrs is collect_store_base_regs(), added by
commit 7d5002bca1de ("bpf: avoid widening registers that hinder exact
stack-slot tracking"):

    base_regs &= aux[idx].stack_ptrs;

That code decides which loop-entry registers must not be widened because
they form stack base addresses for spills, fills and fixed-offset
stack-object calls. It does not invalidate SCEV expressions.

Should the changelog name compute_subprog_args() (the arg_track pass) as
the source and describe the widening-suppression use, rather than pointing
at compute_may_write_masks() and SCEV invalidation?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning
  2026-09-26 14:20 ` [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
@ 2026-09-27 20:27   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:27 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4510 bytes --]

> diff --git a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> index a8358b625ce7..cb55fc216bc8 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_bounds_step.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_bounds_step.c

[ ... ]

> +/* Old register [4..130, step 2] should prune cur register [8..64, step 4]. */
> +SEC("socket")
> +__success __log_level(2)
> +__msg("7: (27) r6 *= 4                       ; R6=scalar({{.*}}umin32=8,{{.*}}umax32=68,{{.*}},step=0+4)")
> +__msg("10: (27) r7 *= 2                      ; R7=scalar({{.*}}umin32=4,{{.*}}umax32=130,{{.*}},step=0+2)")
> +__msg("11: (25) if r0 > 0x2a goto pc+1")
> +__msg("from 11 to 13: safe")
> +__flag(BPF_F_TEST_STATE_FREQ)
> +__naked void step_prune_hit_multiple(void)
> +{
> +	asm volatile ("					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r7 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 &= 0x0f;					\
> +	r6 += 2;					\
> +	r6 *= 4;					\
                           ^^^^

The comment says the current register is [8..64, step 4], but the __msg
check expects umax32=68. Computing r6 &= 0x0f; r6 += 2; r6 *= 4 gives
[2..17] * 4 = [8..68]. Should the comment say [8..68, step 4]?

[ ... ]

> +/* Old register [0..126, step 2] should not prune cur register [0..45, step 3]. */
> +SEC("socket")
> +__failure __log_level(2)
> +__msg("6: (27) r6 *= 3                       ; R6=scalar({{.*}}smin32=0,{{.*}}umax32=45,{{.*}},step=0+3)")
> +__msg("8: (27) r7 *= 2                       ; R7=scalar({{.*}}smin32=0,{{.*}}umax32=126,{{.*}},step=0+2)")
> +__msg("9: (25) if r0 > 0x2a goto pc+1")
> +__msg("11: (15) if r6 == 0x3 goto pc+2")
> +__msg("11: R6=scalar({{.*}},step=0+2)")
> +__msg("13: (95) exit")
> +__msg("from 9 to 11: {{.*}} R6=scalar({{.*}},step=0+3)")
> +__msg("from 11 to 14")
> +__msg("div by zero")
> +__flag(BPF_F_TEST_STATE_FREQ)
> +__naked void step_prune_miss_non_multiple(void)
> +{
> +	asm volatile ("					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r7 = r0;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r6 &= 0x0f;					\
> +	r6 *= 3;					\
> +	r7 &= 0x3f;					\
> +	r7 *= 2;					\
> +	if r0 > 42 goto 1f;	/* can't predict */	\
> +	r6 = r7;		/* step=2 explored first, step=3 explored next */ \
> +1:							\
> +	if r6 == 3 goto 2f;	/* false if step=2, should not prune step=3 */ \
> +	r0 = 0;						\
> +	exit;						\
> +2:							\
> +	r0 /= 0;		/* trap */		\
> +	exit;						\
> +"	:
> +	: __imm(bpf_get_prandom_u32)
> +	: __clobber_all);
> +}

The commit message says this test checks that range_within() does not let
a cached step-2 range hide a step-3 path that reaches division by zero.
Does this test actually verify the step-divisibility check?

The cached r6 is r7 &= 0x3f; r7 *= 2, which is step 2, base 0.
__reg_bound_offset() clears the trailing zero bits of the step from
var_off:

    trailing_zero_bits = count_trailing_zeros(step);
    reg->var_off = tnum_and(reg->var_off,
                            tnum_const(~0ULL << trailing_zero_bits));

So bit 0 is known to be zero in the old var_off. That known-zero bit is
also what lets is_scalar_branch_taken() use tnum_overlap() to predict
'if r6 == 3' as not taken on the first path. The current r6 holds
{0, 3, ..., 45}, so any sound tnum for it must leave bit 0 unknown.
regsafe() then ends with:

    return range_within(rold, rcur) &&
           tnum_in(rold->var_off, rcur->var_off);

tnum_in() fails because (b.mask & ~a.mask) has bit 0 set. Suppose
range_within() had no 'cur->step % old->step == 0' check at all, or the
whole step block were missing as it was before this series. The state
would still not be pruned. The verifier would still reach 'r0 /= 0' and
print 'div by zero', and every __msg would still match.

When the old step is a power of two with base 0, the tnum already
describes the whole line. The comment on step_prune_miss_const_off_line
makes the same point about avoiding power-of-two steps so imod() is the
sole check. step_prune_miss_non_multiple does not follow that rule, so
could a non-power-of-two step for the old line let this test distinguish
a working step check from a missing one?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c
  2026-09-26 14:20 ` [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
@ 2026-09-27 20:27   ` bot+bpf-ci
  2026-09-29 20:26     ` Eduard Zingerman
  0 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:27 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2368 bytes --]

> bpf: record write effects for atomic operations in liveness.c
>
> Atomic read-modify-write instructions are recorded as reads only,
> omitting their stack write effects from the may_write masks needed
> by SCEV.
>
> Record both read and write accesses for atomic RMW operations. Keep
> LOAD_ACQ read-only and STORE_REL write-only. Share the precise/imprecise
> stack-access dispatch so both accesses use the same address handling.
>
> Fixes: fed53dbcdb61 ("bpf: record arg tracking results in bpf_liveness masks")

Does the Fixes: tag point at the right commit?

The commit message says the bug is that atomic RMW write effects are
missing from "the may_write masks needed by SCEV". Those masks
(FM_MAY_WRITE, mark_stack_may_write(), mark_stack_may_write_all()) do not
exist in fed53dbcdb61:

    $ git show fed53dbcdb61:kernel/bpf/liveness.c | grep -c may_write
    0

They were added by 2d21f2c85cb4 ("bpf: track may_write flags in
liveness"), which is four commits earlier in this same series.

In fed53dbcdb61 and in the pre-series base, the only write effect was
must_write. Liveness uses it only as:

    live_before = (live_after & ~must_write) | may_read

in update_insn(), and the must_write_acc transfer in fed53dbcdb61.

For an RMW recorded as a read, the half-slots a write would fully cover
are a subset of the half-slots the read touches at the same instruction:

    [ceil((F-n)/4), F/4-1] is contained in [floor((F-n)/4), (F-1)/4]

So leaving out the def never changed live_before; it was only more
conservative. fed53dbcdb61 therefore has no observable bug that this patch
fixes.

The actual omission was introduced by 2d21f2c85cb4. As written, the tag
would steer stable/AUTOSEL tooling to trees that contain fed53dbcdb61 but
not the may_write infrastructure. The patch would not apply there
(record_imprecise() has a different signature and mark_stack_may_write()
is missing), and it would change nothing there anyway.

Since the culprit is part of the same unmerged series, the usual fix is to
fold this change into 2d21f2c85cb4, or to drop the tag or point it at
2d21f2c85cb4 if the two land separately.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation
  2026-09-26 14:20 ` [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
@ 2026-09-27 20:27   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:27 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6817 bytes --]

> diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
> index be4fd187ba5f..82b21ec472ba 100644
> --- a/tools/testing/selftests/bpf/prog_tests/verifier.c
> +++ b/tools/testing/selftests/bpf/prog_tests/verifier.c

[ ... ]

> diff --git a/tools/testing/selftests/bpf/progs/verifier_idoms.c b/tools/testing/selftests/bpf/progs/verifier_idoms.c
> new file mode 100644
> index 000000000000..a29d6332baa4
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_idoms.c
> @@ -0,0 +1,390 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +
> +#include <linux/bpf.h>
> +#include <bpf/bpf_helpers.h>
> +#include "bpf_misc.h"
> +#include "../../../include/linux/filter.h"

[ ... ]

> +/*
> + * Simple loop: the header (insn 1) is dominated by the pre-header (insn 0). The
> + * back edge (insn 2 -> insn 1) must not change idom(1); the back-edge
> + * predecessor folds up via idoms_intersect().
> + */
> +SEC("socket")
> +__success
> +__log_level(2)
> +__msg("Program dump")
> +__msg("{{.*}} -1   0: {{.*}} (b7) r0 = 0")
> +__msg("{{.*}}  0   1: {{.*}} (07) r0 += 1")
> +__msg("{{.*}}  1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
> +__msg("{{.*}}  2   3: {{.*}} (95) exit")
> +__naked void simple_loop(void)
> +{
> +	asm volatile ("					\
> +	r0 = 0;						\
> +1:	r0 += 1;					\
> +	if r0 < 10 goto 1b;				\
> +	exit;						\
> +"	::: __clobber_all);
> +}

A subsystem pattern flags this as potentially concerning: four of the
new programs (simple_loop, nested_loops, loop_with_if_else_body, and
irreducible) are identical to programs that the previous commit in this
series (5be24395b013 "selftests/bpf: tests for loop hierarchy
computation") added to progs/verifier_loop_hierarchy.c. The earlier
tests already check the same idom column values, and they also check the
scc and loop_header columns.

For example, verifier_loop_hierarchy.c:loop_single already has the same
program as simple_loop above, including:

  __msg("         -1   0: {{.*}} (b7) r0 = 0")
  __msg("  1       0   1: {{.*}} (07) r0 += 1")
  __msg("  1   1   1   2: {{.*}} (a5) if r0 < 0xa goto pc-2")
  __msg("          2   3: {{.*}} (95) exit")

The same applies to nested_loops (== loop_nested),
loop_with_if_else_body (== loop_with_if_else), and irreducible
(== loop_irreducible). self_loop is also the same program as loop_self;
only the idom __msg lines are new there.

These four cases are covered, more strictly, by tests in the same
directory. Could they be dropped, or the missing idom __msg lines be
added to the existing verifier_loop_hierarchy.c cases, so the file only
keeps the cases that are new (ldimm64, asymmetric diamond, multiple back
edges, multiple subprogs, subprog entry as loop header, gotox)?

> +
> +/*
> + * Loop with an if-else in its body. The in-loop merge point (insn 5) is
> + * dominated by the in-loop branch (insn 3), combining a back-edge intersect
> + * with a forward-diamond intersect.
> + */

[ ... ]

> +/*
> + * Nested loops: the idom chain runs inner-header -> outer-body -> outer-header
> + * -> pre-header. Exercises intersect across back edges at two nesting depths.
> + */

[ ... ]

> +/*
> + * Irreducible CFG (loop with two entries, insn 3 and insn 5, reached from the
> + * insn 2 branch). Dominators remain well-defined; this is a convergence test
> + * for the fixpoint under irreducibility. Note insn 6 is dominated by the branch
> + * (insn 2), not by either arm.
> + */
> +SEC("socket")
> +__success
> +__log_level(2)
> +__msg("Program dump")
> +__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
> +__msg("{{.*}}  0   1: {{.*}} (b7) r1 = 0")
> +__msg("{{.*}}  1   2: {{.*}} (25) if r0 > 0x5 goto pc+2")
> +__msg("{{.*}}  2   3: {{.*}} (b7) r1 = 1")
> +__msg("{{.*}}  3   4: {{.*}} (05) goto pc+1")
> +__msg("{{.*}}  2   5: {{.*}} (b7) r1 = 2")
> +__msg("{{.*}}  2   6: {{.*}} (0f) r0 += r1")
> +__msg("{{.*}}  6   7: {{.*}} (a5) if r0 < 0x10 goto pc-5")
> +__msg("{{.*}}  7   8: {{.*}} (95) exit")
> +__naked void irreducible(void)
> +{
> +	asm volatile ("					\
> +	call %[bpf_get_prandom_u32];			\
> +	r1 = 0;						\
> +	if r0 > 5 goto 2f;				\
> +1:	r1 = 1;						\
> +	goto 3f;					\
> +2:	r1 = 2;						\
> +3:	r0 += r1;					\
> +	if r0 < 16 goto 1b;				\
> +	exit;						\
> +"	:
> +	: __imm(bpf_get_prandom_u32)
> +	: __clobber_all);
> +}

The comment says the loop's two entries are insn 3 and insn 5, but
looking at the program:

  2: if r0 > 5 goto pc+2   (-> 5)
  3: r1 = 1
  4: goto pc+1             (-> 6)
  5: r1 = 2
  6: r0 += r1
  7: if r0 < 16 goto pc-5  (-> 3)

The cycle is {3, 4, 6, 7}. Insn 5 is not in it: its only predecessor is
insn 2, and nothing in the loop reaches it. The two entry nodes are
insn 3 (from 2) and insn 6 (from 5). The previous commit's
loop_irreducible test checks exactly this for the same program: insn 5
has empty scc/loop_header columns ("          2   5:"), and insn 6 is in
loop 3 ("  1   3   2   6:").

Insns 3 and 5 are the targets of the insn 2 branch, not the loop
entries. Could the comment be corrected ("entries insn 3 and insn 6")?

[ ... ]

> +/*
> + * Self-loop: insn 1 is its own predecessor. idom(1) must be the pre-header
> + * (insn 0); the self-edge is ignored (idoms[pred] == -1 on the first pass, then
> + * idoms_intersect(a == b) short-circuits). Program is rejected later for an
> + * infinite loop, but the dump (and idoms) are printed before that.
> + */
> +SEC("socket")
> +__failure
> +__log_level(2)
> +__msg("Program dump")
> +__msg("{{.*}} -1   0: {{.*}} (85) call bpf_get_prandom_u32#7")
> +__msg("{{.*}}  0   1: {{.*}} (a5) if r0 < 0xa goto pc-1")
> +__msg("{{.*}}  1   2: {{.*}} (95) exit")
> +__msg("infinite loop detected")
> +__naked void self_loop(void)
> +{
> +	asm volatile ("					\
> +	call %[bpf_get_prandom_u32];			\
> +1:	if r0 < 10 goto 1b;				\
> +	exit;						\
> +"	:
> +	: __imm(bpf_get_prandom_u32)
> +	: __clobber_all);
> +}

The comment says "idoms_intersect(a == b) short-circuits", but preds[1]
is filled in ascending instruction order, so it is {0, 1}. On the second
pass of compute_subprog_idoms(), pred 0 gives new_idom = 0. Pred 1 then
has idoms[1] == 0, so the call is idoms_intersect(env, 1, 0), with
a != b. The loop walks a = idoms[1] = 0 once before the two values meet:

  while (a != b) {
          while (postorder_nums[a] < postorder_nums[b]) {
                  a = idoms[a];

The result, idom(1) = 0, is correct. Could the explanation be corrected
("idoms_intersect(1, 0) walks insn 1 up to its idom 0")?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 18/36] bpf: compute immediate dominators
  2026-09-26 14:20 ` [PATCH bpf-next 18/36] bpf: compute immediate dominators Eduard Zingerman
  2026-09-26 15:54   ` Alexei Starovoitov
@ 2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 7708 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 6a8157e607ee..5e6fa41a1353 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h

[ ... ]

> diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
> index c1f9b0d3468d..6210601eb40b 100644
> --- a/kernel/bpf/Makefile
> +++ b/kernel/bpf/Makefile

[ ... ]

> diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
> new file mode 100644
> index 000000000000..e7ff02e9ab3a
> --- /dev/null
> +++ b/kernel/bpf/loops.c
> @@ -0,0 +1,143 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +
> +#include <linux/slab.h>
> +#include <linux/bpf_verifier.h>
> +
> +static struct bpf_iarray **compute_predecessors(struct bpf_verifier_env *env)
> +{
> +	struct bpf_iarray *succ, *preds, **result;
> +	struct bpf_prog *prog = env->prog;
> +	u32 *num_preds, i, s, sz, len = prog->len;
> +	struct bpf_insn *insn;
> +	void *tmp;
> +
> +	num_preds = kvcalloc(prog->len, sizeof(u32), GFP_KERNEL_ACCOUNT);
> +	if (!num_preds)
> +		return NULL;
> +
> +	/*
> +	 * 'result' layout:
> +	 *  - array of pointers (struct bpf_iarray *)[len]
> +	 *  - struct bpf_iarray one after another
> +	 */
> +	sz = sizeof(struct bpf_iarray) * len;
> +	sz += sizeof(struct bpf_iarray *) * len;
> +	for (i = 0; i < len; i++) {
> +		insn = env->prog->insnsi + i;
> +		succ = bpf_insn_successors(env, i);
> +		sz += sizeof(u32) * succ->cnt;
                     ^^^^

Can sz overflow here?  It is a u32 with no overflow check, but for
programs with many gotox instructions the sum of edges can exceed
prog->len.  bpf_insn_successors() returns the whole shared jump table
for gotox (up to one entry per instruction in the subprogram), and
every gotox in that subprog adds that full count again.

For large privileged programs (prog->len up to BPF_COMPLEXITY_LIMIT_INSNS),
(number of gotox) * (jump table size) can make 4 * that product exceed
U32_MAX.

If sz wraps, kvzalloc(sz) returns a buffer smaller than what the next
loops write:

    result[0..len-1] pointers
    per-insn bpf_iarray headers
    preds->items[preds->cnt++] = i

and those writes run past the end of the allocation.

Even if sz does not wrap but exceeds INT_MAX, kvmalloc() WARNs and
returns NULL.

> +		iarray_for_each(s, succ) {
> +			num_preds[s]++;
> +		}
> +		if (bpf_is_ldimm64(insn))
> +			i++;
> +	}
> +
> +	result = kvzalloc(sz, GFP_KERNEL_ACCOUNT);
> +	if (!result) {
> +		kvfree(num_preds);
> +		return NULL;
> +	}

[ ... ]

> +static int idoms_intersect(struct bpf_verifier_env *env, int a, int b)
> +{
> +	int *postorder_nums = env->cfg.postorder_nums;
> +	int *idoms = env->idoms;
> +
> +	while (a != b) {
> +		while (postorder_nums[a] < postorder_nums[b]) {
> +			a = idoms[a];
> +		}
> +		while (postorder_nums[b] < postorder_nums[a]) {
> +			b = idoms[b];
> +		}
> +	}
> +	return a;
> +}

A review comment from v1 raised a performance concern here.  This walks
the idom chain one instruction at a time, and for a program with many
predecessors (for example, error handling with many "if (err) goto out;"
statements), the k-th call to compute_subprog_idoms() for an instruction
with N predecessors walks back over k branches.  That is N^2/2 steps,
and the outer loop runs at least twice, all before do_check() and for
every program regardless of whether it has loops.

The suggestion was to compute dominators only for instructions inside
loops (scc != 0) since SCEV only needs dominators inside loops, and
possibly to add a bound for large loop bodies.

Has this been addressed?

> +
> +/* See "A Simple, Fast Dominance Algorithm" by Cooper et al. for details. */
> +static void compute_subprog_idoms(struct bpf_verifier_env *env, struct bpf_iarray **preds, int subprog_idx)
> +{
> +	struct bpf_subprog_info *subprog = &env->subprog_info[subprog_idx];
> +	int start = subprog->start;
> +	int po_first = subprog->postorder_start;
> +	int po_last = (subprog + 1)->postorder_start - 1;
> +	int *idoms = env->idoms;
> +	int po_num, pred;
> +	bool changed;
> +
> +	idoms[start] = 0;
> +	changed = true;
> +	do {
> +		changed = false;
> +		/* iterate in reverse postorder */
> +		for (po_num = po_last; po_num >= po_first; po_num--) {
> +			int idx = env->cfg.insn_postorder[po_num];
> +			int new_idom = -1;
> +
> +			iarray_for_each(pred, preds[idx]) {
> +				if (idoms[pred] == -1)
> +					continue;
> +				if (new_idom == -1)
> +					new_idom = pred;
> +				else
> +					new_idom = idoms_intersect(env, pred, new_idom);
> +			}
> +			if (new_idom != -1 && idoms[idx] != new_idom) {
> +				idoms[idx] = new_idom;
> +				changed = true;
> +			}
> +		}
> +	} while (changed);
> +	idoms[start] = -1;
> +}

[ ... ]

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 474caf8696fd..93100541475a 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -22815,6 +22816,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
>  	if (ret < 0)
>  		goto skip_full_check;
>
> +	ret = bpf_compute_idoms(env);
> +	if (ret < 0)
> +		goto skip_full_check;
> +

bpf_compute_idoms() now runs on every program load before do_check(),
so none of the verifier's complexity limits apply yet.  The pass has a
quadratic worst case and nothing in it calls cond_resched() or checks
signal_pending().

The dominator tree is built per instruction, not per basic block, so a
node's depth in the tree is roughly the number of instructions before
it.  compute_subprog_idoms() folds every predecessor into new_idom
using idoms_intersect(), and the cost of each call is the tree distance
from pred up to new_idom.

For a program made of k instructions "if rX == 0 goto E" (fallthrough
to the next), followed by "E: exit", then idom(C_i) = C_{i-1} and
preds[E] = {C_0, C_1, ..., C_{k-1}}.  The call idoms_intersect(C_i, C_0)
walks i steps, so E alone costs about k^2/2 steps per pass.  The do/while
loop always makes at least one extra pass to confirm nothing changed, so
the total is about k^2.

With CAP_BPF, insn_cnt can be up to BPF_COMPLEXITY_LIMIT_INSNS (1M).
That gives about 1e12 pointer-chasing steps inside a single syscall.
The task cannot be killed during this time, and non-preemptible kernels
will report soft lockups.

The shape is not only adversarial: large programs where many checks jump
to one shared return block cost roughly (number of such jumps) x
(program length / 2) per pass.  For example, 50K checks in a 500K-
instruction program is about 1e10 steps per pass.  Unprivileged programs
are capped at BPF_MAXINSNS (4096), where the cost is small.

Other long-running verifier stages guard themselves.  do_check() checks
signal_pending() and calls cond_resched().  Liveness analysis has a
10000-call limit plus cond_resched().  The next commit in this series
adds "if (signal_pending(current)) return -EAGAIN; cond_resched();" to
compute_loops_in_subprog() "to guard against pathological inputs."

Should bpf_compute_idoms() have a similar guard?  Or could the tree be
built over basic blocks instead of instructions so the tree depth stays
small?

>  	ret = bpf_compute_live_registers(env);
>  	if (ret < 0)
>  		goto skip_full_check;

[ ... ]

> @@ -22978,6 +22983,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
>  	kvfree(env->callx_edges);
>  	kvfree(env->func_ptrs);
>  	bpf_diag_free(env);
> +	kvfree(env->idoms);
>  	kvfree(env);
>  	return ret;
>  }


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 11/36] bpf: add bpf_reg_union()
  2026-09-26 14:20 ` [PATCH bpf-next 11/36] bpf: add bpf_reg_union() Eduard Zingerman
@ 2026-09-27 20:42   ` bot+bpf-ci
  2026-09-30  0:09     ` Eduard Zingerman
  0 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6238 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 719b7c7fd2f5..706fdefbc07a 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -1748,6 +1748,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
>  int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
>  int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
>  		      struct cnum64 range, u16 step);
> +int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
> +		  const struct bpf_reg_state *src);
>
>  int bpf_stack_liveness_init(struct bpf_verifier_env *env);
>  void bpf_stack_liveness_free(struct bpf_verifier_env *env);
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index b5f45cb46f14..1ec44d72a346 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -30,6 +30,7 @@
>  #include <linux/module.h>
>  #include <linux/cpumask.h>
>  #include <linux/cnum.h>
> +#include <linux/gcd.h>
>  #include <linux/bpf_mem_alloc.h>
>  #include <net/xdp.h>
>  #include <linux/trace_events.h>

[ ... ]

> @@ -16952,6 +16953,33 @@ int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
>  	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
>  }
>
> +/* acc := acc U src, matching types only. Caller must clear acc's scalar ID. */
> +int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
> +		  const struct bpf_reg_state *src)
> +{
> +	u16 base, step;
> +
> +	if (acc->type != src->type) {
> +		verifier_bug(env, "union of registers with different types");
> +		return -EFAULT;
> +	}
> +	acc->r64 = cnum64_union(acc->r64, src->r64);
> +	acc->r32 = cnum32_union(acc->r32, src->r32);
> +	acc->var_off = tnum_union(acc->var_off, src->var_off);

Does this function handle pointer registers correctly? When bpf_reg_union()
widens a PTR_TO_PACKET's offset (r64/r32/var_off) but keeps its id, it can
break the id-tracking invariant.

For PTR_TO_PACKET, registers with the same id differ only by a known
constant. When adjust_ptr_min_max_vals() adds a non-constant to a packet
pointer, it gives it a fresh id:

  if (!known)
      dst_reg->id = ++env->id_gen;

find_good_pkt_pointers() depends on this rule. After a successful
'pkt_ptr > pkt_end' test, it uses reg_umax(dst_reg) as the new range
and sets reg->range = max(reg->range, new_range) on every PTR_TO_PACKET
with the same id, assuming they all track the same maximum offset.

Once bpf_reg_union() enlarges acc's umax but leaves acc in the same id
group as registers that were not widened, a comparison on acc can give
those other registers a range based on the widened umax that nothing at
runtime backs.

The comment says "Caller must clear acc's scalar ID", but for pointer
types the id has a different meaning. Later in the series when this is
called from union_any_reg() (1987b0e52b34), is_widenable_reg_type()
accepts PTR_TO_PACKET, and scratch_scalar_id() only clears the id when
reg->type == SCALAR_VALUE.

Here's a concrete example with a similar shape to the
conditional_stack_pointer_assignment selftest at the series end:

  r2 = *(u32 *)(r1 + 0)   // data: pkt, id=0
  r3 = *(u32 *)(r1 + 4)   // data_end
  r9 = r2                 // pkt, id=0, off 0
  r7 = r2
  r7 += 1000              // pkt, id=0 (constant add keeps id)
  r6 = 0
1: if r6 >= 3 goto 3f
  if r2 > r3 goto 3f      // fallthrough: find_good_pkt_pointers(r2)
  r0 = *(u8 *)(r9 + 999)
  if r6 != 1 goto 2f
  r2 = r7                 // r2 = (any r2 r7) at loop entry
2: r6 += 1
  goto 1b
3: r0 = 0
  exit

At loop entry, bpf_reg_union() turns r2 into pkt(id=0, [0,1000]), but r9
keeps id=0 and [0,0]. On the fallthrough of 'if r2 > r3',
find_good_pkt_pointers() sets r9->range = reg_umax(r2) = 1000. The load
from r9 + 999 then passes (0 + 999 + 1 <= 1000).

At runtime, on the first iteration r2 == data, so 'data > data_end'
is always false. The program then reads data[999] for any packet,
including one shorter than 1000 bytes: an out-of-bounds packet read.

Should widened PTR_TO_PACKET registers either get a fresh id (and zero
range), or should widening refuse packet pointers that share their id
with other live registers?

> +
> +	/* Retain a common congruence if the bases agree modulo the gcd. */
> +	step = gcd(acc->step, src->step);
> +	base = acc->base % step;
> +	if (base != src->base % step) {
> +		reg_step_reset(acc);
> +	} else {
> +		acc->base = base;
> +		acc->step = step;
> +	}
                ^^^^^^^^^^
                ^^^^^^^^^^

Does this commit build on its own? At ef75b765c289, struct bpf_reg_state
(include/linux/bpf_verifier.h:53) has no 'base' or 'step' members. The
struct ends with id, parent_id, map_uid, precise. Also reg_step_reset()
is not defined anywhere ('git grep reg_step_reset ef75b765c289' only
finds this call site).

So the references to acc->step, src->step, acc->base, src->base and the
reg_step_reset(acc) call all fail to build kernel/bpf/verifier.c when
CONFIG_BPF_SYSCALL=y, which breaks bisection.

The missing pieces are added later in the series by a89774dffa02
("bpf: representation for intervals with steps"), which adds 'u16 base;
u16 step;' to struct bpf_reg_state and defines reg_step_reset(). So the
tree builds again from that commit on, but ef75b765c289 and the commits
between it and a89774dffa02 (d18308307ed6, 0516ad0f46c2) do not build.
The earlier baecd299ec89 ("bpf: add bpf_set_reg_range()") has the same
dependency on reg->step/reg->base.

Would moving a89774dffa02 before these helper commits fix this?

Also, the commit message mentions merging "the circular 32-bit and 64-bit
bounds and tnums" but doesn't mention the base/step congruence merge that
the code does here.

> +	reg_bounds_sync(acc);
> +	return reg_bounds_sanity_check(env, acc, "bpf_reg_union");
> +}
> +
>  /* check validity of 32-bit and 64-bit arithmetic operations */
>  static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
>  {


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data
  2026-09-26 14:20 ` [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
  2026-09-26 15:51   ` Alexei Starovoitov
@ 2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 7300 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index c775bd757706..c6d617581e84 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -662,6 +662,11 @@ struct bpf_insn_aux_data {
>  	};
>  	struct btf_struct_meta *kptr_struct_meta;
>  	u64 map_key_state; /* constant (32 bit) key tracking for maps */
> +	/*
> +	 * Per-instruction summary of stack slots in the current frame
> +	 * that this instruction may write to.
> +	 */
> +	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);

Is the comment accurate that this covers every write? Looking at how the
mask is populated in compute_may_write_masks(), it only ORs in FM_MAY_WRITE
bits from liveness. When a bpf_loop()/for_each_map_elem()/find_vma()/
user_ringbuf_drain() call passes a frame pointer as ctx but the callback is
not a constant (R2 could be cb_a or cb_b on different branches), liveness
handles this case specially in analyze_subprog():

    if (callee == -2) {
        /*
         * same bpf_loop() calls two different callbacks and passes
         * stack pointer to them
         */
        if (info[subprog].at_in[j][caller_reg].frame == ARG_NONE)
            continue;
        for (int f = 0; f <= depth; f++) {
            err = mark_stack_read_all(env, instance, f, idx);
            ...
        }
        continue;
    }

That code calls mark_stack_read_all() but nothing marks may_write. The
callbacks themselves are analyzed standalone with R1-R5 as ARG_NONE, so a
store like `*(u64 *)(r2 + 0) = 0` in the callback is never attributed to
the caller's frame. This means FM_MAY_WRITE stays empty at the call site,
and compute_may_write_masks() stores an empty may_write_mask even when the
callback writes through ctx.

The verifier accepts this pattern: bpf_loop_two_callbacks in
tools/testing/selftests/bpf/progs/verifier_live_stack.c exercises the
callee == -2 branch (it passes fp but with two possible callbacks). If its
callbacks wrote through ctx instead of reading, aux[call].may_write_mask
would stay empty for those slots.

This matters because reset_scevs_at_indirect_writes() in
60fc0594170e ("bpf: compute scalar evolution expressions for loops") uses
bpf_may_write_mask() as the only way to invalidate stack-slot expressions
at a call. A stack-slot SCEV expression could survive a bpf_loop() whose
callback overwrites that slot, and then bpf_clamp_scev_regs() would
intersect the slot's range with bounds derived from that stale expression.

Should the callee == -2 branch also call mark_stack_may_write_all() for
frames 0..depth so the summary is conservative?

Alexei Starovoitov raised a concern about memory usage in the v1 thread
(https://lore.kernel.org/bpf/DLPD9Q4H9C62.2QMD510XIOOJR@gmail.com/):

> That's 32 bytes per insn and the next patch adds 32 more,
> for every prog whether it has loops or not.
> FM_MAY_WRITE in patch 1 makes every frame_masks a third bigger too.
> commit 481ceda77aeb sized the liveness masks by the stack the frame
> uses to avoid exactly that.
>
> bpf_may_write_mask() is the only accessor.
> Can the summary stay in liveness.c, as wide as the subprog's stack,
> and only for insns with scc != 0 ?

Could the summary be stored in liveness.c instead, sized by the subprog's
actual stack usage and only for instructions with scc != 0, to avoid the
fixed 32-byte overhead for every instruction?

>  	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
>  	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
>  	bool nospec; /* do not execute this instruction speculatively */
> @@ -1728,6 +1733,7 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
>  int bpf_stack_liveness_init(struct bpf_verifier_env *env);
>  void bpf_stack_liveness_free(struct bpf_verifier_env *env);
>  int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_state *st);
> +const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx);
>  bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 spi);
>  int bpf_compute_live_registers(struct bpf_verifier_env *env);

[ ... ]

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index cc3ad75aa1e2..c871744ca5a8 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
> @@ -718,6 +718,45 @@ static int cmp_instances(const void *pa, const void *pb)
>  	return 0;
>  }

[ ... ]

> +/*
> + * Precompute, for each instruction, the OR of may_write masks over its top
> + * frame across all func_instances reaching it, stash it in the insn_aux_data.
> + */
> +static void compute_may_write_masks(struct bpf_verifier_env *env)
> +{
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	struct bpf_liveness *liveness = env->liveness;
> +	struct func_instance *instance;
> +	struct frame_masks *fm;
> +	u32 nbits;
> +	int bkt, i;
> +
> +	hash_for_each(liveness->func_instances, bkt, instance, hl_node) {
> +		fm = instance->frames[instance->depth];
> +		if (!fm)
> +			continue;
> +		nbits = frame_mask_bits(fm);
> +		for (i = 0; i < instance->insn_cnt; i++)
> +			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
> +					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
> +	}
> +}

This function reads FM_MAY_WRITE from liveness and summarizes it into
may_write_mask. Does this capture all writes that the instruction might
perform?

As noted above for the header comment, when analyze_subprog() encounters
callee == -2 (same bpf_loop() or similar can invoke different callbacks),
it only marks reads and skips recording writes:

    if (callee == -2) {
        ...
        for (int f = 0; f <= depth; f++) {
            err = mark_stack_read_all(env, instance, f, idx);
            ...
        }
        continue;
    }

No callee instance is created for this case, so merge_may_write() never
runs. The helper's ARG_PTR_TO_STACK argument doesn't help either:
bpf_helper_stack_access_bytes() returns 0 for it, with the comment "The
callback subprog does and it's analyzed separately". But in the -2 case the
callbacks are analyzed only at depth 0 with ctx == ARG_NONE, so their
stores through ctx are never attributed to the caller's frame.

This means FM_MAY_WRITE stays empty at that call instruction even when a
callback writes through ctx into the caller's frame. Since this function
only ORs in FM_MAY_WRITE bits, it will produce an empty may_write_mask for
those slots.

This gap is still present at the end of the series at 4a6395b27755 - the
callee == -2 branch still has only mark_stack_read_all(). Should it also
call mark_stack_may_write_all() for frames 0..depth to keep the summary
conservative?

> +
> +const unsigned long *bpf_may_write_mask(struct bpf_verifier_env *env, u32 insn_idx)
> +{
> +	return env->insn_aux_data[insn_idx].may_write_mask;
> +}

[ ... ]

> @@ -2352,6 +2391,8 @@ int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env)
>  			goto out;
>  	}
>
> +	compute_may_write_masks(env);
> +
>  	if (env->log.level & BPF_LOG_LEVEL2)
>  		err = print_instances(env);


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening
  2026-09-26 14:20 ` [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening Eduard Zingerman
@ 2026-09-27 20:42   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 5811 bytes --]

> diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
> index 82b21ec472ba..2bc1ef910c16 100644
> --- a/tools/testing/selftests/bpf/prog_tests/verifier.c
> +++ b/tools/testing/selftests/bpf/prog_tests/verifier.c

[ ... ]

> diff --git a/tools/testing/selftests/bpf/progs/verifier_scev.c b/tools/testing/selftests/bpf/progs/verifier_scev.c
> new file mode 100644
> index 000000000000..69e9062870ac
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_scev.c

[ ... ]

> +/*
> + * Loop A (counter in r6) contains irreducible loop B (header 5),
> + * which contains C (counter in r9).
> + * Entering B at 'body' saves and restores r6, making it appear invariant.
> + * Entering B at 'alternate' skips the save and modifies r6 instead.
> + * Hence A must not infer a SCEV expression for r6.
> + * SCEV expression for r9 in C should still be computed.
> + *
> + *  0: r6 = 0;
> + *     do {                              // A
> + *  1:     r6++;
> + *  2:     r0 = bpf_get_prandom_u32();
> + *  3:     r7 = 0;
> + *  4:     if (r0 > 5) goto alternate;
> + *  5: B:  r8 = r6;                      // B
> + *  6:     goto body;
> + *  7: alternate:
> + *         r8 = r6;
> + *  8:     r8++;
> + *  9: body:
> + *         r6 = r8;
> + * 10:     r9 = 0;
> + *         do {                          // C
> + * 11:         r9++;
> + * 12:     } while (r9 < 3);
> + * 13:     r7++;
> + * 14:     if (r7 < 4) goto B;
> + * 15: } while (r6 < 4);
> + * 16: r0 = 0;
> + * 17: return r0;
> + */
> +SEC("xdp")
> +__success
> +__log_level(2)
> +__msg("loop at 1{{$}}")
> +__msg("loop at 5, nested in 1, irreducible")
> +__msg("loop at 11, nested in 5")
> +__msg("scev at header 1:")
> +__msg_next("  r6=?")
> +__msg("scev at header 11:")
> +__msg_next("  r9=(+ r9 1) / (linear r9 1)")
> +__msg("loop header at 11, widening r9 to 0..2 step 1")
> +__not_msg("loop header at 1, widening r6")
> +__naked void nested_irreducible_loop(void)
> +{
> +	asm volatile ("					\
> +	r6 = 0;						\
> +a_%=:	r6 += 1;					\
> +	call %[bpf_get_prandom_u32];			\
> +	r7 = 0;						\
> +	if r0 > 5 goto alternate_%=;			\
> +b_%=:	r8 = r6;					\
> +	goto body_%=;					\
> +alternate_%=:						\
> +	r8 = r6;					\
> +	r8 += 1;					\
> +body_%=:						\
> +	r6 = r8;					\
> +	r9 = 0;						\
> +c_%=:	r9 += 1;					\
> +	if r9 < 3 goto c_%=;				\
> +	r7 += 1;					\
> +	if r7 < 4 goto b_%=;				\
> +	if r6 < 4 goto a_%=;				\
> +	r0 = 0;						\
> +	exit;						\
> +"	:
> +	: __imm(bpf_get_prandom_u32)
> +	: __clobber_all);
> +}

Does the comment describe the entry points correctly? Looking at the
pseudo-code and the asm, loop B has its header at 'b_' (insn 5, the
backedge 14 -> 5), and two entry points: the header itself at 'b_' (reached
by falling through from insn 4), and 'body_' at insn 9 (reached from
'alternate_' at insns 7-8).

The path that saves and restores r6 enters at the header 'b_':

    b_:  r8 = r6;
         goto body_;
    ...
    body_: r6 = r8;

The path that modifies r6 goes through 'alternate_' (which is outside loop
B), then enters B at 'body_':

    alternate_:
         r8 = r6;
         r8 += 1;
    body_: r6 = r8;   // gives r6 + 1

The comment says "Entering B at 'body' saves and restores r6" but that
describes the header path, and "Entering B at 'alternate'" names a label
that is outside B. Could the comment be updated to say "Entering B at B
saves and restores r6 ... Entering B at body (via alternate) skips the save
and modifies r6 instead"?

[ ... ]

> +/*
> + * Complement of no_widen_stack_spill: a sub-register (1-byte) store to the stack
> + * lands as STACK_MISC and carries no tracked value, so the induction variable
> + * addressing it (r2) is still widened.
> + */
> +SEC("xdp")
> +__success
> +__log_level(2)
> +__flag(BPF_F_TEST_STATE_FREQ)
> +__msg("widening r2 to 0..24 step 8")
> +__naked void widen_byte_stack_store(void)
> +{
> +	asm volatile ("					\
> +	r0 = 0;						\
> +	r2 = 0;						\
> +1:	r3 = r10;					\
> +	r3 += -64;					\
> +	r3 += r2;					\
> +	*(u8 *)(r3 + 0) = r0;				\
> +	r0 += 1;					\
> +	r2 += 8;					\
> +	if r0 < 4 goto 1b;				\
> +	r0 = 0;						\
> +	exit;						\
> +"	::: __clobber_all);
> +}

Does the comment correctly describe the verifier's behavior here? In this
test, r3 = fp-64 + r2 with r2 in {0, 8, 16, 24}, so before widening each
store '*(u8 *)(r3 + 0) = r0' is at a fixed offset that is 8-byte-aligned
(-64, -56, -48, -40).

check_stack_write_fixed_off() handles that as a narrow spill when the value
fits:

    if (reg && !(off % BPF_REG_SIZE) && reg->type == SCALAR_VALUE &&
        env->bpf_capable) {
        ...
        save_register_state(env, state, spi, reg, size);

save_register_state() marks the low 'size' bytes STACK_SPILL and saves the
scalar. r0 is 0..3, so it fits in a byte and is fully tracked with its
scalar id. check_stack_read_fixed_off() then gives it back on a same-size
fill.

The store only lands as STACK_MISC after r2 is widened, because then the
write offset varies. That can happen with any size store, so it is not what
separates this case from no_widen_stack_spill.

The actual difference is that collect_store_base_regs() only counts BPF_W
and BPF_DW accesses:

    if (size == BPF_W || size == BPF_DW)

Byte and half-word spills lose precision under widening in the same way.

Could the comment describe the actual rule (byte and half-word accesses are
not treated as spill bases) rather than say a 1-byte store has no tracked
value? Also, BPF does not use the term "sub-register" for narrow
operations.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID
  2026-09-26 14:20 ` [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
@ 2026-09-27 20:42   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 8714 bytes --]

> diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
> index e1d559ba1271..9f188ff9be89 100644
> --- a/tools/testing/selftests/bpf/prog_tests/verifier.c
> +++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
> @@ -26,6 +26,7 @@
>  #include "verifier_bpf_get_stack.skel.h"
>  #include "verifier_bpf_trap.skel.h"
>  #include "verifier_bswap.skel.h"
> +#include "verifier_btf_array_access.skel.h"
>  #include "verifier_btf_ctx_access.skel.h"
>  #include "verifier_btf_flex_array.skel.h"
>  #include "verifier_btf_unreliable_prog.skel.h"
> @@ -211,6 +212,7 @@ void test_verifier_bounds_mix_sign_unsign(void) { RUN(verifier_bounds_mix_sign_u
>  void test_verifier_bpf_get_stack(void)        { RUN(verifier_bpf_get_stack); }
>  void test_verifier_bpf_trap(void)             { RUN(verifier_bpf_trap); }
>  void test_verifier_bswap(void)                { RUN(verifier_bswap); }
> +void test_verifier_btf_array_access(void)     { RUN(verifier_btf_array_access); }
>  void test_verifier_btf_ctx_access(void)       { RUN(verifier_btf_ctx_access); }
>  void test_verifier_btf_flex_array(void)       { RUN(verifier_btf_flex_array); }
>  void test_verifier_btf_unreliable_prog(void)  { RUN(verifier_btf_unreliable_prog); }
> diff --git a/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
> new file mode 100644
> index 000000000000..ab6b74457365
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_btf_array_access.c
> @@ -0,0 +1,506 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +
> +#include <vmlinux.h>
> +#include <bpf/bpf_helpers.h>
> +#include <bpf/bpf_core_read.h>
> +#include "bpf_experimental.h"
> +#include "bpf_misc.h"

[ ... ]

> +/*
> + * Array nested below a struct member at a non-zero offset. The walk descends
> + * through 'm' (resetting the running offset) before reaching 'arr', exercising
> + * the array_start = min_off - arrays[i].off telescoping across a WALK_STRUCT
> + * dive.
> + */
> +struct mid {
> +	struct inner arr[8];	/* stride 8, size 64 */
> +};
> +
> +struct nest {
> +	long pad;		/* off 0 */
> +	struct mid m;		/* off 8 */
> +};
> +
> +/* m.arr[i].b, i in [0, 7], 4-byte read: stays within m.arr. */
> +SEC("syscall")
> +__success
> +int nested_struct_in_bounds(void *ctx)
> +{
> +	unsigned long i = bpf_get_prandom_u32();
> +	struct nest *o;
> +	int val = 0;
> +
> +	o = bpf_obj_new(typeof(*o));
> +	if (!o)
> +		return 0;
> +	asm volatile ("						\
> +	r1 = %[o];						\
> +	r2 = %[i];						\
> +	r2 &= 7;						\
> +	r2 *= 8;						\
> +	r1 += r2;						\
> +	%[val] = *(u32 *)(r1 + 12);				\
> +"	: [val] "=r"(val)
> +	: [o] "r"(o), [i] "r"(i)
> +	: "r1", "r2");
> +	bpf_obj_drop(o);
> +	return val;
> +}

Does this test actually exercise the WALK_STRUCT path described in the
comment? The comment says the test exercises "offset telescoping across a
WALK_STRUCT dive," but with offset 12 (reading m.arr[i].b), the walk
path appears to skip the WALK_STRUCT branch.

Looking at btf_struct_walk() in the verifier, it only returns WALK_STRUCT
when the offset exactly equals a member's start:

    if (btf_type_is_struct(mtype)) {
        t = mtype;
        if (off == moff) {
            *next_btf_id = mid;
            return WALK_STRUCT;
        }
        off -= moff;
        goto again;
    }

With min_off at 12:
- At nest, m starts at moff 8 but off is 12, so it takes 'goto again'
  with off 4
- At mid, arr[0] starts at moff 0 but off is 4, so it takes 'goto again'
  again
- At inner, it finds b and returns WALK_SCALAR

The WALK_STRUCT branch in btf_struct_access() is never reached, so the
offset reset and arrays[] re-entry that the comment describes doesn't
occur.

Would reading m.arr[i].a (r1 + 8) instead put min_off at offsetof(struct
nest, m), causing the walk to go through WALK_STRUCT twice (nest->mid,
then mid->inner) and cover the telescoping path the comment describes?

> +
> +/* m.arr[i].b, i in [0, 15]: max offset runs past the end of m.arr. */
> +SEC("syscall")
> +__failure
> +__msg("invalid variable offset access into struct nest")
> +int nested_struct_out_of_bounds(void *ctx)
> +{
> +	unsigned long i = bpf_get_prandom_u32();
> +	struct nest *o;
> +	int val = 0;
> +
> +	o = bpf_obj_new(typeof(*o));
> +	if (!o)
> +		return 0;
> +	asm volatile ("						\
> +	r1 = %[o];						\
> +	r2 = %[i];						\
> +	r2 &= 15;						\
> +	r2 *= 8;						\
> +	r1 += r2;						\
> +	%[val] = *(u32 *)(r1 + 12);				\
> +"	: [val] "=r"(val)
> +	: [o] "r"(o), [i] "r"(i)
> +	: "r1", "r2");
> +	bpf_obj_drop(o);
> +	return val;
> +}
> +
> +/*
> + * A trailing flexible array member makes the struct tail an unbounded region,
> + * so a varying access into it has no upper bound to exceed (mirrors how
> + * unix_address.name[] is accessed via sun_path[i]).
> + *
> + * A custom (program) BTF struct can only be reached as a PTR_TO_BTF_ID through
> + * bpf_obj_new(), which does not allow flexible array members. So use a kernel
> + * type via an untrusted PTR_TO_BTF_ID (bpf_core_cast()). struct vring_used ends
> + * with a flexible array 'ring[]' of struct vring_used_elem { __virtio32 id, len; },
> + * mirroring a 'struct foo { int a; int b; }' flexible array.
> + */

Is the justification here accurate? The comment states that "bpf_obj_new()
does not allow flexible array members," which also appears in the commit
message:

    A program-defined struct cannot be used here because bpf_obj_new()
    forbids flexible array members

However, looking at the verifier, bpf_obj_new() itself doesn't reject
flexible array members. The obj_new path in check_kfunc_call() only
checks that the type is a struct and allocates ret_t->size bytes.

tools/testing/selftests/bpf/progs/linked_list_fail.c already has
obj_new_flex_array that allocates a struct with a flexible array member
using bpf_obj_new_impl():

    struct obj_new_flex {
        int hdr;
        struct obj_new_flex_elem cells[];
    };

It only fails later on access with "access beyond struct obj_new_flex."

The real restriction appears to be that btf_struct_access() passes
walk_flex_arrays = !type_is_alloc(reg->type), so flexible arrays aren't
walked for MEM_ALLOC registers.

Also, a program-BTF struct can be reached as a non-MEM_ALLOC PTR_TO_BTF_ID
through other means. Loading a __kptr field of a local type outside an RCU
critical section yields PTR_TO_BTF_ID | PTR_MAYBE_NULL | PTR_UNTRUSTED
without MEM_ALLOC (through btf_ld_kptr_type()). For such a register,
btf_struct_walk() does walk flexible arrays.

tools/testing/selftests/bpf/progs/verifier_btf_flex_array.c:stash_and_peek
already uses exactly this approach to reach the flexible array walk with a
program-defined struct in a SEC("syscall") program (which is outside RCU by
default).

Should the comment and commit message give the correct reason instead? The
current explanation also creates a dependency on struct vring_used being in
vmlinux BTF, which only exists when virtio is built in.

> +
> +/* ring[i].len, i in [0, 63], 4-byte read: the flexible array has no upper bound. */
> +SEC("syscall")
> +__success
> +int flex_array_field(void *ctx)
> +{
> +	unsigned long i = bpf_get_prandom_u32();
> +	struct vring_used *o;
> +	int val = 0;
> +
> +	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
> +	asm volatile ("						\
> +	r1 = %[o];						\
> +	r2 = %[i];						\
> +	r2 &= 63;						\
> +	r2 *= 8;						\
> +	r1 += r2;						\
> +	%[val] = *(u32 *)(r1 + 8);				\
> +"	: [val] "=r"(val)
> +	: [o] "r"(o), [i] "r"(i)
> +	: "r1", "r2");
> +	return val;
> +}
> +
> +/*
> + * ring[] accessed with step 4 (half of sizeof(struct vring_used_elem)): the
> + * step is not a whole number of elements, so the flexible array stride does
> + * not match.
> + */
> +SEC("syscall")
> +__failure
> +__msg("invalid variable offset access into struct vring_used")
> +int flex_array_unaligned_step(void *ctx)
> +{
> +	unsigned long i = bpf_get_prandom_u32();
> +	struct vring_used *o;
> +	int val = 0;
> +
> +	o = bpf_core_cast(bpf_get_current_task_btf(), struct vring_used);
> +	asm volatile ("						\
> +	r1 = %[o];						\
> +	r2 = %[i];						\
> +	r2 &= 7;						\
> +	r2 *= 4;						\
> +	r1 += r2;						\
> +	%[val] = *(u32 *)(r1 + 4);				\
> +"	: [val] "=r"(val)
> +	: [o] "r"(o), [i] "r"(i)
> +	: "r1", "r2");
> +	return val;
> +}
> +
> +char _license[] SEC("license") = "GPL";

---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe()
  2026-09-26 14:20 ` [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
@ 2026-09-27 20:42   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2007 bytes --]

> diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
> index 18bf7b660c2f..29d18520d27c 100644
> --- a/kernel/bpf/states.c
> +++ b/kernel/bpf/states.c
> @@ -608,12 +608,17 @@ static bool regsafe(...) {
>  	case PTR_TO_MAP_KEY:
>  	case PTR_TO_MAP_VALUE:
>  	case PTR_TO_MEM:
>  	case PTR_TO_BUF:
>  	case PTR_TO_TP_BUFFER:
>  		/* If the new min/max/var_off satisfy the old ones and
>  		 * everything else matches, we are OK.
>  		 */
>  		return memcmp(rold, rcur, offsetof(struct bpf_reg_state, var_off)) == 0 &&
>  		       range_within(rold, rcur) &&
>  		       tnum_in(rold->var_off, rcur->var_off) &&
>  		       check_ids(rold->id, rcur->id, idmap) &&
>  		       check_ids(rold->parent_id, rcur->parent_id, idmap) &&
>  		       check_ids(rold->map_uid, rcur->map_uid, idmap);

[ ... ]

>  	case PTR_TO_STACK:
> -		return regs_exact(rold, rcur, idmap);
> +		return memcmp(rold, rcur, offsetof(struct bpf_reg_state, var_off)) == 0 &&
> +		       range_within(rold, rcur) &&
> +		       tnum_in(rold->var_off, rcur->var_off) &&
> +		       check_ids(rold->id, rcur->id, idmap) &&
> +		       check_ids(rold->parent_id, rcur->parent_id, idmap) &&
> +		       rold->frameno == rcur->frameno;

This isn't a bug, but the new PTR_TO_STACK block copies the first five
conditions from the PTR_TO_MAP_VALUE/PTR_TO_MEM group above, with only
the final check differing.

Since frameno sits in the union before var_off, the memcmp() already
compares it, so the explicit frameno == frameno check at the end doesn't
change the result. The PTR_TO_MAP_* group's extra check_ids() on map_uid
wouldn't affect stack pointers (map_uid is always 0 for R10-derived
pointers).

Could PTR_TO_STACK simply join the existing case group above instead of
repeating its conditions?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 01/36] bpf: track may_write flags in liveness
  2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
  2026-09-26 15:51   ` Alexei Starovoitov
@ 2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 11196 bytes --]

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index cd9523f69298..cc3ad75aa1e2 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c

[ ... ]

> @@ -1590,7 +1635,7 @@ static int record_load_store_access(struct bpf_verifier_env *env,
>  	if (ptr->frame >= 0 && ptr->frame <= depth)
>  		return record_stack_access(env, instance, ptr, sz, ptr->frame, insn_idx);
>  	if (ptr->frame == ARG_IMPRECISE)
> -		return record_imprecise(env, instance, ptr->mask, insn_idx);
> +		return record_imprecise(env, instance, sz, ptr->mask, insn_idx);
>  	/* ARG_NONE: not derived from any frame pointer, skip */
>  	return 0;
>  }

For atomic read-modify-write instructions (BPF_ADD/AND/OR/XOR with or
without BPF_FETCH, BPF_XCHG, BPF_CMPXCHG), the size sz stays positive,
so record_imprecise() records them as reads only and never sets
FM_MAY_WRITE for the stack slot they modify.

Looking at the code path:

  record_load_store_access()
    case BPF_STX:
      if (BPF_MODE(insn->code) == BPF_ATOMIC) {
          if (insn->imm == BPF_STORE_REL)
              sz = -sz;
          ...
      }

Atomic RMW operations don't hit the BPF_STORE_REL path, so sz remains
positive and record_imprecise() marks only FM_MAY_READ.

Does this contradict the commit message's statement that may_write
includes "any possibly modified slot"?  A later commit in the series
(0428816ca270) fixes this, but at this commit the gap exists.

> @@ -1618,6 +1663,9 @@ static int record_arg_access(struct bpf_verifier_env *env,
>  			err = mark_stack_read_all(env, instance, f, insn_idx);
>  			if (err)
>  				return err;
> +			err = mark_stack_may_write_all(env, instance, f, insn_idx);
> +			if (err)
> +				return err;
>  		}
>  		return 0;
>  	}
> @@ -1627,7 +1675,7 @@ static int record_arg_access(struct bpf_verifier_env *env,
>  	if (frame >= 0 && frame <= depth)
>  		err = record_stack_access(env, instance, at, bytes, frame, insn_idx);
>  	else if (frame == ARG_IMPRECISE)
> -		err = record_imprecise(env, instance, at->mask, insn_idx);
> +		err = record_imprecise(env, instance, bytes, at->mask, insn_idx);
>  	return err;
>  }

The commit message says may_write covers "any possibly modified slot,
including partial writes", but for helper and kfunc calls, may_write
is derived only from the sign of bytes returned by
bpf_helper_stack_access_bytes() or bpf_kfunc_stack_access_bytes().

Looking at bpf_helper_stack_access_bytes() in verifier.c:

    full_write = (at & MEM_UNINIT) &&
                 (!arg_type_is_raw_mem(at) || env->allow_uninit_stack);
    if (full_write)
        return -size;
    return size;

It returns a negative value only for guaranteed full writes (MEM_UNINIT
with allow_uninit_stack).  Arguments that write to the stack but don't
meet the full_write condition return +size and are recorded as reads
only.

Some examples that write caller stack but never get FM_MAY_WRITE:

- Helper arguments with ARG_PTR_TO_MEM | MEM_WRITE (like
  bpf_fib_lookup params, bpf_skb_load_bytes buffer) when
  !env->allow_uninit_stack (CAP_BPF without CAP_PERFMON)

- Helper arguments with ARG_PTR_TO_FIXED_SIZE_MEM | MEM_WRITE |
  MEM_ALIGNED (like bpf_skb_check_mtu mtu_len), regardless of
  capabilities

- Kfunc output arguments that are not __uninit (like
  bpf_dynptr_slice buffer__szk)

Later in the series, fad299f2d710 exports the mask as
bpf_may_write_mask(), and 60fc0594170e uses it in
reset_scevs_at_indirect_writes() as the only way to invalidate
stack-spill SCEVs across calls.  A stale SCEV for a slot overwritten by
one of these helpers can produce wrong loop reasoning in
1987b0e52b34 (loop widening).

Should the may_write side be derived from the argument's
MEM_WRITE/MEM_UNINIT (and non-const kfunc mem args) rather than from
full_write, keeping the full_write distinction only for must_write?

> @@ -1444,10 +1465,12 @@ static void arg_track_xfer(struct bpf_verifier_env *env, struct bpf_insn *insn,
>   *   access_bytes == 0:      no access
>   *
>   */
> -static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
> -				   s64 access_bytes, u32 frame, u32 insn_idx)
> +static int record_stack_access_off(struct func_instance *instance, const struct arg_track *arg,
> +				   u32 off_idx, s64 access_bytes, u32 frame, u32 insn_idx)
>  {
> +	s64 fp_off = arg->off[off_idx];
>  	s32 slot_hi, slot_lo;
> +	int err;
>
>  	if (fp_off >= 0)
>  		/*

[ ... ]

> @@ -1471,7 +1495,15 @@ static int record_stack_access_off(struct func_instance *instance, s64 fp_off,
>  		access_bytes = -access_bytes;
>  		slot_hi = (-fp_off) / STACK_SLOT_SZ - 1;
>  		slot_lo = max_t(s32, (-fp_off - access_bytes + STACK_SLOT_SZ - 1) / STACK_SLOT_SZ, 0);
> -		return mark_stack_write(instance, frame, insn_idx, slot_lo, slot_hi);
> +		if (arg->off_cnt == 1) {
> +			err = mark_stack_write(instance, frame, insn_idx, slot_lo, slot_hi);
> +			if (err)
> +				return err;
> +		}
> +		/* Mark partially covered slots as may_def */
> +		slot_hi = (-fp_off - 1) / STACK_SLOT_SZ;
> +		slot_lo = max_t(s32, (-fp_off - access_bytes) / STACK_SLOT_SZ, 0);
> +		return mark_stack_may_write(instance, frame, insn_idx, slot_lo, slot_hi);
>  	}
>  	return 0;
>  }

record_stack_access_off() derives may_write purely from the sign of
access_bytes.  For helper and kfunc stack arguments, access_bytes comes
from bpf_helper_stack_access_bytes() or bpf_kfunc_stack_access_bytes(),
which return negative values only for guaranteed full writes:

    full_write = (at & MEM_UNINIT) &&
                 (!arg_type_is_raw_mem(at) || env->allow_uninit_stack);

Any other written argument is returned as +size and takes the
`access_bytes > 0` branch which marks only FM_MAY_READ.

Examples that write caller stack but never get FM_MAY_WRITE:

- ARG_PTR_TO_FIXED_SIZE_MEM | MEM_WRITE | MEM_ALIGNED (like
  bpf_skb_check_mtu mtu_len and bpf_xdp_check_mtu mtu_len)

- ARG_PTR_TO_MEM | MEM_WRITE (like bpf_skb_fib_lookup params,
  bpf_xdp_fib_lookup params, bpf_sock_ops_load_hdr_opt)

- ARG_PTR_TO_UNINIT_MEM raw buffers (like bpf_skb_load_bytes,
  bpf_map_pop_elem) when !env->allow_uninit_stack (CAP_BPF without
  CAP_PERFMON)

- Kfunc mem args that are written but not __uninit (like
  bpf_dynptr_slice buffer__opt)

The consumer introduced later in the series makes the omission matter.
scev.c:reset_scevs_at_indirect_writes() only invalidates stack-slot
SCEVs through bpf_may_write_mask(env, idx) for calls;
forget_call_regs() clears only R0-R5.

So a loop like this:

    r1 = *(u64 *)(r10 - 8); r1 += 1; *(u64 *)(r10 - 8) = r1;
    r1 = r6; r2 = 0; r3 = r10; r3 += -8; r4 = 8;
    call bpf_skb_load_bytes;
    r1 = *(u64 *)(r10 - 8); if r1 < 100 goto loop;

keeps fp-8 as a linear recurrence even though the helper overwrites it
with packet data each iteration.

bpf_compute_loop_iters() can then declare the loop terminating, and
states.c:is_state_visited() prunes such loops with RANGE_WITHIN and
skips the "infinite loop detected" check.  A loop that never terminates
at runtime could then be accepted.

This is not fixed in the range 2d21f2c85cb4..4a6395b27755.

Should the may_write side be derived from the argument's
MEM_WRITE/MEM_UNINIT (and non-const kfunc mem args) rather than from
full_write, keeping the full_write distinction only for must_write?

[ ... ]

> @@ -2025,18 +2074,50 @@ static int merge_instances(struct func_instance *dst, struct func_instance *src)
>  		for (i = 0; i < dst->insn_cnt; i++) {
>  			unsigned long *dst_read = rel_mask(d, i, FM_MAY_READ);
>  			unsigned long *dst_write = rel_mask(d, i, FM_MUST_WRITE);
> +			unsigned long *dst_may_write = rel_mask(d, i, FM_MAY_WRITE);
>  			unsigned long *src_read = rel_mask(s, i, FM_MAY_READ);
>  			unsigned long *src_write = rel_mask(s, i, FM_MUST_WRITE);
> +			unsigned long *src_may_write = rel_mask(s, i, FM_MAY_WRITE);
>
>  			for (w = 0; w < d->words; w++) {
>  				dst_read[w] |= w < s->words ? src_read[w] : 0;
>  				dst_write[w] &= w < s->words ? src_write[w] : 0;
> +				dst_may_write[w] |= w < s->words ? src_may_write[w] : 0;
>  			}
>  		}
>  	}
>  	return 0;
>  }
>
> +/*
> + * Fold a fully analyzed callee instance writes to upper frames as
> + * may_write marks at callsite in caller's frames.
> + */
> +static int merge_may_write(struct func_instance *caller, struct func_instance *callee)
> +{
> +	DECLARE_BITMAP(acc, FRAME_HALF_SPIS);
> +	u32 call_idx = callee->callsite;
> +	struct frame_masks *fm;
> +	u32 f, i, nbits;
> +	int err;
> +
> +	for (f = 0; f < callee->depth; f++) {
> +		fm = callee->frames[f];
> +		if (!fm)
> +			continue;
> +		nbits = frame_mask_bits(fm);
> +		bitmap_zero(acc, nbits);
> +		for (i = 0; i < callee->insn_cnt; i++)
> +			bitmap_or(acc, acc, rel_mask(fm, i, FM_MAY_WRITE), nbits);
> +		if (bitmap_empty(acc, nbits))
> +			continue;
> +		err = mark_stack_mask(caller, f, call_idx, FM_MAY_WRITE, acc, fm->words);
> +		if (err)
> +			return err;
> +	}
> +	return 0;
> +}
> +
>  static struct func_instance *fresh_instance(struct func_instance *src)
>  {
>  	struct func_instance *f;

[ ... ]

> @@ -2204,6 +2285,11 @@ static int analyze_subprog(struct bpf_verifier_env *env,
>  					goto out_free;
>  			}
>  		}
> +
> +		/* Summarize callee's writes to ancestor frames onto the callsite */
> +		err = merge_may_write(instance, callee_instance);
> +		if (err)
> +			goto out_free;
>  	}
>
>  	if (prev_instance) {

The commit adds merge_may_write() for known callees, but there is a
path earlier in analyze_subprog() that handles callback-calling helpers
(bpf_loop, bpf_for_each_map_elem, bpf_find_vma,
bpf_user_ringbuf_drain) whose callback is not a constant subprog.

When find_callback_subprog() returns -2 and the ctx argument points to
the stack, the code has this:

    if (callee == -2) {
        /*
         * same bpf_loop() calls two different callbacks and passes
         * stack pointer to them
         */
        if (info[subprog].at_in[j][caller_reg].frame == ARG_NONE)
            continue;
        for (int f = 0; f <= depth; f++) {
            err = mark_stack_read_all(env, instance, f, idx);
            if (err)
                goto out_free;
        }
        continue;
    }

This marks only mark_stack_read_all() for every frame and then does
`continue`, so merge_may_write() never runs and no
mark_stack_may_write_all() is recorded either.

The helper's ARG_PTR_TO_STACK ctx argument gets 0 from
bpf_helper_stack_access_bytes(), so record_arg_access() adds nothing
for it.

As a result, the call insn's may_write mask omits every caller slot the
unknown callback may write through ctx, even though the callback can do
'*(u64 *)ctx += 1' on a caller stack slot.

The unknown-callee path in record_arg_access() was updated in this
commit to add mark_stack_may_write_all().

The same stale-SCEV issue follows through
reset_scevs_at_indirect_writes()/bpf_may_write_mask().

This is still present at 4a6395b27755.

Should this loop also call mark_stack_may_write_all(env, instance, f,
idx)?

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers
  2026-09-26 14:20 ` [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
  2026-09-26 14:37   ` sashiko-bot
@ 2026-09-27 20:42   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:42 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 7230 bytes --]

> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> index 9bcfefdfb734..37437b93471d 100644
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c
> @@ -7496,6 +7496,14 @@ bool btf_ctx_access(int off, int size, enum bpf_access_type type,
>  }
>  EXPORT_SYMBOL_GPL(btf_ctx_access);
>
> +#define MAX_ARRAYS_WALK 16
> +
> +/* Description of an array crossed while walking a BTF access chain */
> +struct array_access {
> +	u32 id;		/* BTF_KIND_ARRAY type id */
> +	u32 off;	/* offset of the access within the array */
> +};

[ ... ]

> @@ -7542,22 +7583,19 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
>  		/* If the last element is a variable size array, we may
>  		 * need to relax the rule.
>  		 */
> -		if (vlen == 0)
> +		member = find_flex_member(btf, t);
> +		if (!member)
>  			goto error;
>
> -		member = btf_type_member(t) + vlen - 1;
> -		mtype = btf_type_skip_modifiers(btf, member->type,
> -						NULL);
> -		if (!btf_type_is_array(mtype))
> +		moff = __btf_member_bit_offset(t, member) / 8;
> +		if (off < moff)
>  			goto error;
>
> +		mtype = btf_type_skip_modifiers(btf, member->type, &mid);
>  		array_elem = (struct btf_array *)(mtype + 1);
> -		if (array_elem->nelems != 0)
> -			goto error;
>
> -		moff = __btf_member_bit_offset(t, member) / 8;
> -		if (off < moff)
> -			goto error;
> +		if (push_array(arrays, arrays_cnt, (struct array_access){ mid, off - moff }))
> +			return -E2BIG;

Does push_array() receive the correct BTF type id here? The comment on
struct array_access documents id as a 'BTF_KIND_ARRAY type id', but this
code stores mid after btf_type_skip_modifiers() has been called with
&mid as the output parameter.

btf_type_skip_modifiers() walks through modifiers and typedefs on the
array's element type, so for a member like 'u32 foo[N]', '__be32 addr[4]',
or 'atomic_t cnt[N]', mid would point to the resolved element type (INT or
STRUCT) rather than the ARRAY itself.

In btf_struct_access(), the check_variable_offset loop calls
__btf_resolve_size(btf, t, ...) where t comes from
btf_type_by_id(btf, arrays[i].id). If arrays[i].id is not an ARRAY type,
__btf_resolve_size() would leave elem_id as 0, causing btf_resolve_size()
to fail on btf_void, and the entry would be skipped with continue.

This would affect most kernel arrays since u8/u16/u32/u64/__be* types are
typedefs.

Should this store the array id before resolving the element type, perhaps
by calling btf_type_skip_modifiers(btf, member->type, &array_id) before
the existing call, or by capturing the id when btf_type_is_array(mtype)
first matches?


> @@ -7680,6 +7718,9 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
>  			 *      the array's element as long as it is
>  			 *      within the mtrue_end boundary.
>  			 */
> +			if (push_array(arrays, arrays_cnt,
> +				       (struct array_access){ mid, off - moff }))
> +				return -E2BIG;
>
>  			/* skip empty array */
>  			if (moff == mtrue_end)

[ ... ]

> @@ -7772,15 +7813,29 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf,
>
>  int btf_struct_access(struct bpf_verifier_log *log,
>  		      const struct bpf_reg_state *reg,
> -		      int off, int size, enum bpf_access_type atype __maybe_unused,
> +		      int _off, int size, enum bpf_access_type atype __maybe_unused,
>  		      u32 *next_btf_id, enum bpf_type_flag *flag,
>  		      const char **field_name)
>  {
>  	const struct btf *btf = reg->btf;
>  	enum bpf_type_flag tmp_flag = 0;
> +	struct array_access arrays[MAX_ARRAYS_WALK];
> +	const struct btf_member *member;
>  	const struct btf_type *t;
> +	u32 i, arrays_cnt = 0;
>  	u32 id = reg->btf_id;
> -	int err;
> +	s64 min_off, max_off, flex_off;
> +	int err, off, ret;

[ ... ]

> @@ -7846,6 +7905,54 @@ int btf_struct_access(struct bpf_verifier_log *log,
>  	} while (t);
>
>  	return -EINVAL;
> +
> +check_variable_offset:
> +	if (min_off == max_off)
> +		return ret;
> +
> +	/* Find an offset at which access would go to a flexible array tail (if any). */
> +	t = btf_type_skip_modifiers(btf, reg->btf_id, NULL);
> +	member = find_flex_member(btf, t);
> +	flex_off = member ? __btf_member_bit_offset(t, member) / 8 : S64_MAX;
> +
> +	/*
> +	 * If this is a varying offset access, the step recorded within a register
> +	 * should correspond to one of the arrays visited while walking.
> +	 */
> +	for (i = 0; i < arrays_cnt; i++) {
> +		s64 array_start, array_end, access_end;
> +		u32 asize, esize, elem_id;
> +
> +		/*
> +		 * 'arrays' records top-level arrays only, use __btf_resolve_size()
> +		 * to get flattened representation.
> +		 */
> +		t = btf_type_by_id(btf, arrays[i].id);
> +		if (IS_ERR(__btf_resolve_size(btf, t, &asize, NULL, &elem_id, NULL, NULL)) ||
> +		    IS_ERR(btf_resolve_size(btf, btf_type_by_id(btf, elem_id), &esize)))
> +			continue;
> +		/* Make sure every accessed offset lands on the same field of some element. */
> +		if (esize == 0 || reg->step % esize != 0)
> +			continue;
> +
> +		array_start = min_off - arrays[i].off;
> +		/* An array within the flexible tail has no upper bound to exceed */
> +		if (array_start >= flex_off)
> +			return ret;
>                         ^^^^^^^^^^

Does this shortcut accept inner arrays that should still be bounded?

The arrays list contains not just the flexible array itself, but also
fixed-size arrays nested inside a flexible array element. When
btf_struct_walk() pushes the flex array and then walks into the element
struct via 'goto again', it pushes the inner member arrays too.

For an inner array nested inside a flex element, array_start >= flex_off
is true, so this code returns success after only checking
reg->step % esize(inner) == 0. It never verifies that the step is a
multiple of the flex element stride, and never bounds the access by the
inner array's own size.

Consider this example:

    struct elem {
        struct foo *ptrs[2];
        long val;
    };                          /* sizeof 24 bytes */

    struct top {
        u32 n;
        struct elem items[];    /* flex_off = 8 */
    };

With a trusted pointer r1 to struct top and a scalar
r2 = (rand & 1023) * 8 (step 8, smin 0), the load *(u64 *)(r1 + 8)
walks at min_off 8 and records:

    arrays[0] = {items, 0}
    arrays[1] = {ptrs, 0}

For arrays[0] (the flex array), esize is 24 and 8 % 24 != 0, so the loop
continues. For arrays[1] (the inner ptrs array), esize is 8 and
8 % 8 == 0, and array_start = 8 >= flex_off, so the function returns
PTR_TO_BTF_ID(struct foo).

At runtime, offset 8 + 16 = 24 is items[0].val (a long), not a struct foo
pointer. Loads through PTR_TRUSTED/MEM_RCU/MEM_ALLOC are emitted without
PROBE_MEM, and the mistyped result can be passed to helpers expecting BTF
pointers.

Should the 'no upper bound' rule apply only to the flexible array entry
itself (the one whose array_start == flex_off or whose asize/nelems is 0),
while inner arrays must still pass the access_end <= array_start + asize
check?



---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops
  2026-09-26 14:20 ` [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops Eduard Zingerman
  2026-09-26 14:38   ` sashiko-bot
@ 2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:43 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 9096 bytes --]

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> new file mode 100644
> index 000000000000..396b0c776dfa
> --- /dev/null
> +++ b/kernel/bpf/scev.c

[ ... ]

> +static int add_expr(struct scev *scev, struct expr e)
> +{
> +	struct expr_bucket *bucket;
> +	u32 i, id, hash, new_cap;
> +	void *tmp;
> +
> +	hash = expr_hash(&e) % NUM_BUCKETS;
> +	bucket = scev->exprs_ht[hash];
> +
> +	if (bucket) {
> +		for (i = 0; i < bucket->cnt; i++) {
> +			id = bucket->ids[i];
> +			if (expr_eq(&e, &scev->exprs[id]))
> +				return id;
> +		}
> +	}
> +
> +	if (!bucket || bucket->cap == bucket->cnt) {
> +		new_cap = bucket ? bucket->cap * 2 : 32;
> +		bucket = kvrealloc(bucket, sizeof(*bucket) + sizeof(u32) * new_cap, GFP_KERNEL_ACCOUNT | __GFP_ZERO);
> +		if (!bucket)
> +			return -ENOMEM;
> +		scev->exprs_ht[hash] = bucket;
> +		bucket->cap = new_cap;
> +	}
> +
> +	if (scev->exprs_cnt == scev->exprs_cap) {
> +		new_cap = scev->exprs_cap + 256;
> +		tmp = kvrealloc(scev->exprs, sizeof(struct expr) * new_cap, GFP_KERNEL_ACCOUNT);
> +		if (!tmp)
> +			return -ENOMEM;
> +		scev->exprs = tmp;
> +		scev->exprs_cap = new_cap;
> +	}
> +
> +	id = scev->exprs_cnt++;
> +	scev->exprs[id] = e;
> +	bucket->ids[bucket->cnt++] = id;
> +	return id;
> +}

Could this cause excessive CPU time during verification for programs with
many loops?

The hash table has a fixed NUM_BUCKETS (256) and never resizes. The linear
search through each bucket grows quadratically with the number of distinct
expressions, which accumulates across all loops in a program. A program
near BPF_COMPLEXITY_LIMIT_INSNS (1M instructions) laid out as bounded
loops of ALU operations could generate around 3M expressions, leading to
roughly 1.7e10 bucket comparisons.

The exprs array also grows by a fixed 256 entries per realloc. Once past
the kmalloc range, kvrealloc falls back to vmalloc, and vrealloc copies
the entire array on every page-boundary growth.

There are no cond_resched or signal_pending checks in the SCEV computation
path, which runs from bpf_check before do_check_main. Could this cause an
uninterruptible CPU stall with soft-lockup warnings on systems without
CONFIG_PREEMPT?

[ ... ]

> +static void setup_initial_loop_env(struct bpf_verifier_env *env, struct env *e, int insn_idx)
> +{
> +	struct scev *scev = env->scev;
> +	int i;
> +
> +	for (i = 0; i < REGS_NUM; i++)
> +		e->reg2expr[i] = expr1(scev, REG, i);
> +}

Does this handle expr1 allocation failures?

expr1 calls add_expr, which returns a negative errno when kvrealloc fails,
for example under a memcg limit or fault injection. The loop fills REGS_NUM
(267) entries, so the exprs array is likely to grow during the loop. If
that fails, reg2expr[i] ends up as 0xFFFFFFF4 (-ENOMEM cast to u32) for
one or more high stack-slot registers.

The bogus id is then used as an index into scev->exprs:

kernel/bpf/scev.c:transform_expr_once() does expr = scev->exprs[root],
with root being u32 0xFFFFFFF4, when reached from compute_header_scevs.

kernel/bpf/scev.c:forget_non_invariants() does header_expr =
&scev->exprs[header_env->reg2expr[i]] when the loop is nested in an outer
one.

That is an out-of-bounds read tens of GiB past the array. Should
setup_initial_loop_env return int and fail on id < 0, with
compute_scev_for_loop propagating that error?

[ ... ]

> +static int mk_fill(struct scev *scev, int id, u8 code)
> +{
> +	bool sx = BPF_MODE(code) == BPF_MEMSX;
> +
> +	switch (BPF_SIZE(code)) {
> +	case BPF_B:  return expr1(scev, sx ? SEXT8  : ZEXT8,  id);
> +	case BPF_H:  return expr1(scev, sx ? SEXT16 : ZEXT16, id);
> +	case BPF_W:  return expr1(scev, sx ? SEXT32 : ZEXT32, id);
> +	case BPF_DW: return sx ? UNKNOWN_EXPR_ID : id; /* DW sign-extended load is invalid */
> +	}
> +	return UNKNOWN_EXPR_ID;
> +}

[ ... ]

> +static int maybe_load_fp(struct scev *scev, struct env *e, struct bpf_insn *insn)
> +{
> +	int id;
> +
> +	if (insn->src_reg == BPF_REG_FP && is_spill_off(insn->off))
> +		id = mk_fill(scev, e->reg2expr[off_to_reg(insn)], insn->code);
> +	else
> +		id = OPAQUE_EXPR_ID;
> +	return replace_reg(scev, e, insn->dst_reg, id);
> +}

Does mk_fill account for endianness?

maybe_load_fp calls mk_fill for any 8-byte-aligned fp offset, and mk_fill
models a narrow fill as the low-order bits of the spilled 64-bit value
(ZEXTn or SEXTn of the slot expression). That only holds on little-endian.

On a big-endian host such as s390, a B/H/W load at the start of an 8-byte
slot (fp-8, fp-16, etc.) reads the high-order bytes of a DW spill, not the
low-order ones. For BPF_W, simplify then folds ZEXT32(IMM c) into
IMM((u32)c), producing an immediate that does not match the runtime value.

The later commit 1987b0e52b34 in this series uses these expressions in
bpf_clamp_scev_regs for backedge clamping and in bpf_widen_scev_regs for
widening. Wrong immediates there can lead to incorrect register ranges.

Should mk_fill special-case big-endian layout or return OPAQUE or UNKNOWN
for narrow fills when little_endian is false?

[ ... ]

> +static void forget_non_invariants(struct scev *scev, struct env *acc, struct env *header_env)
> +{
> +	struct expr *header_expr;
> +	int i;
> +
> +	if (acc->empty)
> +		acc->empty = false;
> +
> +	for (i = 0; i < REGS_NUM; i++) {
> +		header_expr = &scev->exprs[header_env->reg2expr[i]];
> +		if (is_expr1(header_expr, REG, i))
> +			continue;
> +		acc->reg2expr[i] = UNKNOWN_EXPR_ID;
> +	}
> +}

[ ... ]

> +static int compute_scev_for_loop(struct bpf_verifier_env *env, int cur_header)
> +{
> +	struct bpf_min_heap *worklist = &env->scev->worklist;
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	struct bpf_loop *cur_loop = aux[cur_header].loop;
> +	struct env *header_env, *nested_header_env;
> +	struct scev *scev = env->scev;
> +	struct bpf_loop *nested_loop;
> +	struct env *cur_env = NULL;
> +	struct env *old_env = NULL;
> +	struct bpf_iarray *succ;
> +	bool log_level2 = env->log.level & BPF_LOG_LEVEL2;
> +	int s, i, err, idx, succ_idx;

[ ... ]

> +		memcpy(cur_env, find_loop_env(scev, cur_header, idx), sizeof(*cur_env));
> +		if (nested_loop) {
> +			/*
> +			 * Process nested loop as a single instruction by
> +			 * forgetting anything non-invariant in the nested loop
> +			 */
> +			if (log_level2)
> +				memcpy(old_env, cur_env, sizeof(*old_env));
> +			nested_header_env = find_header_env(scev, idx);
> +			forget_non_invariants(scev, cur_env, nested_header_env);
> +			if (log_level2)
> +				log_env_changes(env, LOG_AT_TRANSFER, old_env, cur_env, idx);
> +			/*
> +			 * Treat nested loop exits as successors,
> +			 * join cur_env into successor's envs.
> +			 */
> +			for (i = 0; i < nested_loop->exits_cnt; i++) {
> +				s = nested_loop->exits[i].to;
> +				err = join_successor(env, cur_header, s, old_env, cur_env);
> +				if (err)
> +					goto out;
> +			}

Does forget_non_invariants correctly identify invariants across nested
loops?

The code decides that register i is invariant across a nested loop only by
checking that the nested loop's header env maps i to (reg i). The header
env is the join of the backedge contributions only. But a register can hold
a different value on an exit edge and still be restored before the latch,
either by an exact register copy (MOV copies the expression id) or by a
spill and fill.

For example, using 64-bit arithmetic with r7 as a map value pointer and r6
starting at 0:

    2: r9 = 0                 // outer header
    3: r8 = r6                // inner header
    4: r6 = -1000
    5: r9 += 1
    6: if r9 > 5 goto 9       // inner exit, r6 == -1000 here
    7: r6 = r8                // restores r6: expr is (reg r6) again
    8: goto 3                 // inner backedge
    9: r6 += 1
   10: if r6 s< 3 goto 2      // outer latch

The inner header env has r6 = (reg r6), so forget_non_invariants keeps r6
and edge 6->9 receives r6 = (reg r6) instead of UNKNOWN. The outer header
SCEV becomes r6 = (linear r6 1), and the latch at 10 is matched as a linear
latch. At runtime, though, r6 is -999 when the backedge is first taken.

The later commit 1987b0e52b34 in this series consumes these SCEVs.
bpf_compute_loop_iters computes max_header_count = 3 for the loop above,
and bpf_widen_scev_regs widens r6 at the header to [0..2]. On the
backedge, bpf_clamp_scev_regs intersects the verifier's actual r6 range
with linear_bounds(entry r6 = 0) = [0..2]. With r6 = -999 the intersection
is empty, which hits verifier_bug and fails with -EFAULT. If the modified
value is a range that overlaps [0..2], the clamp silently narrows r6 to
[0..2], the widened header checkpoint then prunes the next iteration, and a
later r6-indexed map value access is accepted even though r6 can be
negative at runtime.

Should forget_non_invariants check that a register is unchanged along every
path from the nested header to each exit, not only along the paths to the
latch?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking
  2026-09-26 14:20 ` [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
  2026-09-26 14:46   ` sashiko-bot
@ 2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:43 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4834 bytes --]

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> index 205bf8684c3d..1fa6b25a4f29 100644
> --- a/kernel/bpf/scev.c
> +++ b/kernel/bpf/scev.c

[ ... ]

> @@ -1021,6 +1021,65 @@ static void reset_scevs_at_indirect_writes(struct bpf_verifier_env *env, struct
>  	}
>  }
>
> +/* OR the loop-entry registers referenced by expr 'id' into 'mask'. */
> +static void or_expr_regs(struct bpf_verifier_env *env, u32 id, unsigned long *mask)
> +{
> +	struct scev *scev = env->scev;
> +	u32 order;
> +
> +	scev->stack_sz = 0;
> +	expr_stack_push(scev, id);
> +	while (expr_next(scev, &id, &order)) {
> +		if ((order & PRE) && scev->exprs[id].op == REG)
> +			__set_bit(scev->exprs[id].params[0], mask);
> +	}
> +}

Does or_expr_regs() handle the DEPTH_LIMIT case? expr_next() stops
descending at EXPR_STACK_DEPTH (8) and returns POST | DEPTH_LIMIT without
visiting children. REG leaves below level 8 would be dropped from the mask,
but or_expr_regs() needs to collect every loop-entry register the stack
base expression depends on because the result blocks widening. Deep
expressions are easy to produce since transfer() wraps every ALU32 result
in ZEXT32:

  r6 = 0
  w1 = w6          ; zext(r6)
  w1 += 1          ; zext(add(zext(r6),1))
  w1 &= 1          ; zext(and(...))
  w1 <<= 3         ; zext(lsh(...))
  r3 = r10
  r3 += -16
  r3 += r1         ; REG r6 is at level 9 of r3's expr
  r4 = *(u64 *)(r3 + 0)   ; fill of a spilled ctx pointer
  r0 = *(u32 *)(r4 + 0)
  r6 += 1
  if r6 < 4 goto 1b

In collect_store_base_regs(), r3 would be selected for the LDX since r3
is in stack_ptrs. or_expr_regs() would record r10 but not r6, so r6 gets
widened to [0..3]. The fill then has a variable offset and yields a
SCALAR, rejecting the ctx dereference. The same program verifies when the
loop is enumerated.

Other traversals in this file handle this conservatively:
is_any_imm_reg() and is_any_imm_reg_opaque() return false on DEPTH_LIMIT,
and union_any_reg() treats it as a verifier_bug. Setting all bits
of mask (bitmap_fill) when DEPTH_LIMIT is seen would keep the analysis
conservative.

[ ... ]

> @@ -1975,6 +2040,7 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
>  	for (r = 0; r < REGS_NUM; r++) {
>  		struct bpf_reg_state *reg;
>  		u32 ra, r_scev, r_expr;
> +		bool spill_base = false;
>
>  		if (!scev_reg_alive(env, st, r))
>  			continue;
> @@ -1984,10 +2050,16 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta
>  		/* If SCEV for r is (linear <reg> <slope>) */
>  		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
>  		    is_widenable_reg_type(reg)) {
> -			/* Can't widen if the iteration range overflows */
> -			if (base_reg == r &&
> -			    !linear_bounds(reg, iters, slope_imm, &bounds))
> -				goto cant_widen;
> +			if (base_reg == r) {
> +				/* Can't widen if the iteration range overflows */
> +				if (!linear_bounds(reg, iters, slope_imm, &bounds))
> +					goto cant_widen;
> +				/* Spills at varying offsets lose precision */
> +				if (test_bit(r, loop->store_base_regs)) {
> +					spill_base = true;
> +					goto cant_widen;
> +				}
> +			}
>  			continue;
>  		}
>
> @@ -2006,6 +2078,8 @@ int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_sta

Does the store_base_regs check apply to all register types? The check
only runs for registers whose SCEV is (linear r slope). Registers accepted
by is_any_imm_reg() skip it and are widened via union_any_reg() later in
bpf_widen_scev_regs(), again without consulting store_base_regs.

A register j that is 0 on entry and set to 1 at the end of each iteration
has the header SCEV (any j 1). If the body fills a spilled pointer from
ptrs[j] (LDX with base fp-16 + j*8), or_expr_regs() puts j into
store_base_regs. j is still widened to [0..1], the fill has a variable
offset, and the pointer comes back as a SCALAR. The loop enumerates fine
without widening, so this is the same precision loss the commit is meant
to prevent.

The commit message states: "During the loop-widening stage,
bpf_widen_scev_regs() refuses to widen a loop when some of the registers
the loop changes appear in store_base_regs." The refusal is actually in
bpf_compute_loop_iters(), not bpf_widen_scev_regs(), and it covers only
self-based linear registers, not every register the loop changes. The
message also calls the field bpf_loop->store_base_reg instead of
store_base_regs.

Either apply the test_bit(r, loop->store_base_regs) check before the
is_any_imm_reg() acceptance too, or reword the message to state the
narrower rule.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops
  2026-09-26 14:20 ` [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops Eduard Zingerman
  2026-09-26 14:42   ` sashiko-bot
@ 2026-09-27 20:43   ` bot+bpf-ci
  1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:43 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 26250 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 54a971800ea3..c0eea472949d 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -328,6 +328,21 @@ struct bpf_retval_range {
>  	bool return_32bit;
>  };
>
> +struct bpf_loop_iters {
> +	u32 min_header_count;   /* min number of times header is executed */
> +	u32 max_header_count;   /* max number of times header is executed */
> +	bool pre_cond;
> +};
> +
> +struct loop_stack_entry {
> +	struct bpf_verifier_state *entry_state;
> +	struct bpf_loop_iters iters;
> +	u32 loop_id:31;
> +	u32 terminates:1;
> +};
> +
> +#define LOOP_STACK_SIZE 16
                               ^^

The fixed-size per-frame loop stack makes loop nesting depth a hard
verifier limit, so programs that were accepted before this commit are now
rejected. handle_loop_entry_exit() is called unconditionally from
do_check() and pushes an entry for every loop header on the path from the
current insn to the root of the loop tree, whether or not SCEV can prove
the loop terminates. loop_stack_push() fails verification with -E2BIG when
that depth exceeds 16:

kernel/bpf/verifier.c:loop_stack_push() {
    for (h = bpf_loop_at_index(env, env->insn_idx); h >= 0; h = aux[h].loop_header, num_missing++) {
        if (cnt && loop_stack[cnt - 1].loop_id == h) break;
        if (num_missing == LOOP_STACK_SIZE) goto e2big;
        ...
    }
    for (; num_missing; num_missing--, cnt++) {
        if (cnt == LOOP_STACK_SIZE) goto e2big;
        ...
    e2big:
        verbose(env, "Too many nested loops (%d/%d) at %d\n", cnt, num_missing, env->insn_idx);
        return -E2BIG;
}

The loop hierarchy has no nesting-depth limit of its own, and irreducible
or over-approximated CFGs (eg gotox jump tables where every block may jump
to every later block) produce very deep static nests even when the program
terminates trivially at run time.

This is visible in tools/testing/selftests/bpf/progs/verifier_gotox.c:
gotox_edges_at_limit flips from __success __retval(0) to __failure
__msg("Too many nested loops"), with the comment "the CFG is a nest of
~1000 irreducible loops, deeper than the verifier's loop stack; the
program is rejected." Deeply nested bpf_for()/can_loop loops (more than
16 levels in one subprogram) are rejected the same way.

The commit message ("use SCEV to widen bounded loops") does not mention
that loop nesting deeper than 16 is now a rejection condition, and nothing
later in the series changes this.

Since widening is an optimisation, could the logic stop tracking loops
beyond the limit and treat them as non-widenable (the pre-patch behaviour)
instead of failing the load?

[ ... ]

> diff --git a/kernel/bpf/scev.c b/kernel/bpf/scev.c
> index 396b0c776dfa..205bf8684c3d 100644
> --- a/kernel/bpf/scev.c
> +++ b/kernel/bpf/scev.c

[ ... ]

> +static bool eval_expr(struct bpf_verifier_env *env, struct scev *scev, struct bpf_func_state *st, u32 id, u64 *result)
> +{
> +	struct bpf_reg_state *reg;
> +	u32 regno;
> +	s64 imm;
> +
> +	if (is_reg(scev, id, &regno)) {
> +		reg = scev_regno_to_reg(env, st, regno);
> +		if (reg->type == NOT_INIT)
> +			return false;
> +		if (tnum_is_const(reg->var_off)) {
> +			*result = reg->var_off.value;
                             ^^^^^^^^^^^^^^^^^^^^
> +			return true;
> +		}
> +	} else if (is_imm(scev, id, &imm)) {
> +		*result = imm;
> +		return true;
> +	}
> +	return false;
> +}

eval_expr() returns reg->var_off.value for any initialized register type.
For pointer types (this tree folds the fixed offset into var_off), that is
the offset from the pointed-to object, not the runtime value the latch
compares.

When allow_ptr_leaks permits comparing a pointer against a scalar (or
against a pointer with a different base), the iteration count is computed
from offsets while the runtime comparison uses absolute addresses, so the
result has no relation to the real trip count.

Pointer types such as PTR_TO_STACK and PTR_TO_MAP_VALUE are accepted by
is_widenable_reg_type(), so this count can feed widening and clamping.

Should eval_expr() require SCALAR_VALUE for the latch register, bound and
initial values, or require both sides of the comparison to be the same
pointer base?

[ ... ]

> +static bool compute_max_iters(struct bpf_verifier_env *env,
> +			      struct bpf_func_state *st,
> +			      struct linear_latch *latch,
> +			      struct bpf_loop_iters *iters)
> +{
> +	u64 base, step, diff, bound, initial, max_iters;
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	struct bpf_loop *loop = aux[bpf_loop_at_index(env, latch->insn_idx)].loop;
> +	struct scev *scev = env->scev;
> +	u8 op = latch->op;
> +
> +	if (!eval_expr(env, scev, st, latch->base_expr, &base) ||
> +	    !eval_expr(env, scev, st, latch->step_expr, &step) ||
> +	    !eval_expr(env, scev, st, latch->bound_expr, &bound) ||
> +	    !eval_expr(env, scev, st, latch->reg_expr, &initial))
> +		return false;
> +
> +	if (step == 0)
> +		return false;
> +	if ((s64)step == S64_MIN)
> +		return false;
> +	if ((s64)step < 0) {
> +		/* Multiply both sides of the equation by -1, e.g. -2*i > -3 becomes 2*i < 3 */
> +		op = bpf_flip_opcode(op);
> +		step = -step;
> +		swap(bound, initial);
> +	}
> +	diff = bound - initial;
> +	if (diff / step == U64_MAX)
> +		return false;
> +	switch (op) {
> +	case BPF_JLT:
> +		max_iters = (u64)initial >= (u64)bound ? 0 : div_round_up(diff, step);
> +		break;
> +	case BPF_JLE:
> +		max_iters = (u64)initial >  (u64)bound ? 0 : diff / step + 1;
> +		break;
> +	case BPF_JSLT:
> +		max_iters = (s64)initial >= (s64)bound ? 0 : div_round_up(diff, step);
> +		break;
> +	case BPF_JSLE:
> +		max_iters = (s64)initial >  (s64)bound ? 0 : diff / step + 1;
> +		break;
> +	case BPF_JNE:
> +		max_iters = diff % step ? U32_MAX : div_round_up(diff, step);
> +		break;
> +	default:
> +		return false;
> +	}
> +
> +	if (max_iters > U32_MAX)
> +		return false;
> +
> +	/*
> +	 * The latch is a conditional jump with one jump target exiting the loop.
> +	 * Linear latch is matched only if the loop has a single backedge.
> +	 * The loop still, however can have multiple exits.
> +	 * In such case, conservatively assume that non-latch exit can happen
> +	 * at any iteration, thus setting minimal number of iterations as 0.
> +	 */
> +	iters->max_header_count = max_iters + (base == 0 ? 1 : 0);
                                   ^^^^^^^^^   ^^^^^^^^^^^^^^^^^^^^
> +	iters->min_header_count = loop->exits_cnt == 1 ? iters->max_header_count : 0;
> +	iters->pre_cond = base == 0;
                           ^^^^^^^^^^^^
> +	return true;
> +}

compute_max_iters() evaluates latch->base_expr but only uses it in a
"base == 0" test. The latch condition is "initial + base + k*step <op>
bound" (see the struct linear_latch comment), but max_iters is derived
from "bound - initial" alone.

The result is right only when base is 0 (pre-condition) or exactly equal
to step (the usual post-increment). match_linear_latch() also accepts any
"(linear (+ reg imm) step)" latch, eg when the compared register is a copy
of the induction variable plus a constant taken before the increment.

For a positive step, whenever base is below step and non-zero (a negative
constant, or 0 < base < step), the real number of header executions is
1 + ceil((bound - initial - base) / step), which is larger than the value
computed here.

For JNE, a base that is not a multiple of step can mean the equality is
never reached even though diff % step == 0.

Because bpf_widen_scev_regs() and bpf_clamp_scev_regs() treat
max_header_count as an upper bound, under-counting lets clamping throw
away induction-variable values the loop really reaches.

Should the arithmetic use (bound - (initial + base)) and add 1 for the
header entry, or should match_linear_latch() reject latches where base is
not 0 or step?

[ ... ]

> +/*
> + * Compute the range an induction variable in `reg` spans over the loop.
> + * Returns false if the computation overflows s64.
> + */
> +static bool linear_bounds(struct bpf_reg_state *reg, struct bpf_loop_iters *iters, s64 slope,
> +			  struct bounds *out)
> +{
> +	s64 slope_abs = slope < 0 ? -slope : slope;
> +	s64 min_val = reg_smin(reg);
> +	s64 max_val = reg_smax(reg);
> +	s64 total_change;
> +	u16 step;
> +
> +	if (check_mul_overflow(slope, (s64)iters->max_header_count - 1, &total_change))
> +		return false;
> +	if (slope > 0) {
> +		if (check_add_overflow(max_val, total_change, &max_val))
> +			return false;
> +	} else {
> +		if (check_add_overflow(min_val, total_change, &min_val))
> +			return false;
> +	}
> +	/*
> +	 * If the entry value is a single point the value set is 'v + slope * k',
> +	 * so the step is |slope|. Otherwise, only the power-of-two alignment
> +	 * shared by the entry value and the slope.
> +	 */
> +	if (cnum64_is_const(reg->r64))
> +		step = slope_abs;
> +	else
> +		step = 1u << min_t(u32, tnum_alignment(reg->var_off), __ffs(slope_abs));
> +	out->range = cnum64_from_srange(min_val, max_val);
> +	out->step = step;
> +	return true;
> +}

When the entry value is not a constant, linear_bounds() picks
step = 2^min(tnum_alignment(var_off), ctz(slope)). All real values are
then 0 mod step, because every tnum member and the slope are multiples of
2^k. But linear_bounds() returns only a range and a step, not a base.

bpf_widen_scev_regs() and bpf_clamp_scev_regs() pass that range and step
to bpf_set_reg_range(), which does:

    reg->base = imod(cnum64_smin(range), step);
    reg->var_off = tnum_unknown;
    reg_bounds_sync(reg);

The base therefore comes from smin of the range, not from the tnum
lattice. smin is not always a tnum member: __update_reg64_bounds() only
intersects r64 with the tnum's min/max and does not round umin up to the
next member. After a range refinement, smin can be off the lattice,
bpf_set_reg_range() records the wrong congruence class, and
cnum64_intersect_linear() in deduce_bounds_64_from_step() then shrinks the
range to that wrong class.

Example:

    r1 = map_lookup_elem(map with value_size 18); if r1 == 0 goto out
    r7 = *(u64 *)(r1 + 0)
    r7 &= 0xc           // var_off (0; 0xc), step reset to 1, r64 [0,12]
    if r7 < 1 goto out  // r64 [1,12]; real values {4,8,12}
    r6 = 0
  H:
    r2 = r1
    r2 += r7
    r0 = *(u8 *)(r2 + 0)
    r7 += 4
    r6 += 1
    if r6 < 3 goto H

At the header:

- max_header_count is 3.
- linear_bounds() returns range [1, 1 + 4*2] = [1, 20] with step
  1 << min(2, 2) = 4.
- bpf_set_reg_range() sets base = imod(1, 4) = 1. cnum64_intersect_linear()
  then narrows r7 to {1,5,9,13,17}, so umax is 17.

At runtime r7 at H is in {4,...,20}, which is 0 mod 4. That set is
disjoint from the verifier's set, and the real maximum 20 is above the
believed maximum 17.

The body's map access passes check_map_access() because 17 + 1 <= 18, but
at runtime it reads offset 20 of an 18-byte value: an out-of-bounds map
value access.

On the backedge, bpf_clamp_scev_regs() does the same. [5,21] intersected
with [1,20] is [5,20]. With base imod(5,4) = 1 this becomes [5,17], which
is within W, so the state is pruned and the error is never found.

The slope < 0 case has the same problem, since
min_val = reg_smin + slope*(n-1) is used as the base source.

Should the widened register base come from linear_bounds() and be passed
to bpf_set_reg_range()? For example, 0 in the tnum-alignment case, or
smin(entry) mod |slope| in the constant case. Alternatively, round min/max
onto the 0 mod step lattice before building the range. Later commits in
the series do not change this logic.

[ ... ]

> +int bpf_compute_loop_iters(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
> +			   struct bpf_loop_iters *iters)
> +{
> +	struct bpf_func_state *cur_func = st->frame[st->curframe];
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	struct bpf_verifier_log *log = &env->log;
> +	struct scev *scev = env->scev;
> +	struct env *header_env;
> +	struct linear_latch latch;
> +	struct bpf_loop *loop;
> +	struct bounds bounds;
> +	int insn_idx = st->insn_idx;
> +	int linear_latch;
> +	u32 r, base_reg;
> +	int latch_idx;
> +	s64 slope_imm;
> +	int err;

[ ... ]

> +	header_env = find_header_env(scev, insn_idx);
> +	for (r = 0; r < REGS_NUM; r++) {
> +		struct bpf_reg_state *reg;
> +		u32 ra, r_scev, r_expr;
> +
> +		if (!scev_reg_alive(env, st, r))
> +			continue;
> +
> +		r_scev = header_env->reg2scev[r];
> +		reg = scev_regno_to_reg(env, cur_func, r);
> +		/* If SCEV for r is (linear <reg> <slope>) */
> +		if (is_simple_linear(scev, r_scev, &base_reg, &slope_imm) &&
> +		    is_widenable_reg_type(reg)) {
                     ^^^^^^^^^^^^^^^^^^^^^
> +			/* Can't widen if the iteration range overflows */
> +			if (base_reg == r &&
> +			    !linear_bounds(reg, iters, slope_imm, &bounds))
> +				goto cant_widen;
> +			continue;
> +		}

is_widenable_reg_type() switches on base_type(reg->type), so
PTR_MAYBE_NULL variants are accepted, for example
PTR_TO_MAP_VALUE | PTR_MAYBE_NULL straight from bpf_map_lookup_elem().

The SCEV expressions are computed from the instructions alone, without
type information. So a maybe-null pointer that is null-checked and then
advanced inside the loop gets a (linear r 8) SCEV:

    r1 = map_lookup_elem(...)   // PTR_TO_MAP_VALUE_OR_NULL, id N
    r6 = 0
  H:
    if r1 == 0 goto out         // extra loop exit
    r1 += 8
    r6 += 1
    if r6 < 3 goto H
  out:
    ...

At the first visit of H, bpf_widen_scev_regs() calls bpf_set_reg_range()
on r1, which is still maybe-null. This gives r1 a variable offset
([0,16], step 8, non-constant var_off) while keeping the PTR_MAYBE_NULL
flag and id N, because scratch_scalar_id() only clears scalar ids.

The verifier normally never allows this state, because
adjust_ptr_min_max_vals() rejects arithmetic on maybe-null pointers.

The widened state then reaches the null check at H. check_cond_jmp_op()
calls mark_ptr_or_null_regs(), and mark_ptr_or_null_reg() fires:

kernel/bpf/verifier.c:mark_ptr_or_null_reg() {
    if (!(type_is_ptr_alloc_obj(reg->type) || type_is_non_owning_ref(reg->type)) &&
        WARN_ON_ONCE(!tnum_equals_const(reg->var_off, 0)))
        return;
}

So a kernel WARN can be triggered by any program load that is allowed to
use bounded loops (with panic_on_warn, the machine panics). The program is
later rejected at r1 += 8, but only after the WARN.

Should is_widenable_reg_type() reject type_may_be_null(reg->type), or at
least on the linear path? Later commits in the series do not change
is_widenable_reg_type().

[ ... ]

> +int bpf_widen_scev_regs(struct bpf_verifier_env *env, struct bpf_verifier_state *st,
> +			struct bpf_verifier_state *loop_entry, struct bpf_loop_iters *iters)
> +{
> +	struct bpf_func_state *cur_func = st->frame[st->curframe];
> +	struct bpf_func_state *entry_func = loop_entry->frame[loop_entry->curframe];
> +	struct bpf_verifier_log *log = &env->log;
> +	struct scev *scev = env->scev;
> +	struct bpf_reg_state *reg;
> +	struct env *header_env;
> +	struct bounds bounds;
> +	u32 r, base_reg, r_expr, a, b;
> +	int insn_idx = st->insn_idx;
> +	s64 slope_imm;
> +	int err;
> +
> +	header_env = find_header_env(scev, insn_idx);
> +	for (r = 0; r < REGS_NUM; r++) {
> +		if (!scev_reg_alive(env, st, r))
> +			continue;
> +
> +		r_expr = header_env->reg2scev[r];
> +		/* If SCEV for r is (linear <reg> <slope>)*/
> +		if (is_simple_linear(scev, r_expr, &base_reg, &slope_imm) &&
> +		    base_reg == r) {
> +			reg = scev_regno_to_reg(env, cur_func, r);
> +			/* Feasibility was checked in the filtering pass above. */
                             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> +			if (!linear_bounds(reg, iters, slope_imm, &bounds)) {
> +				verifier_bug(env, "scev widen bounds overflow for r%d", r);
> +				return -EFAULT;
> +			}

The comment refers to a "filtering pass above", but bpf_widen_scev_regs()
has no such pass. The overflow/feasibility filtering is done earlier in a
different function, bpf_compute_loop_iters() (the "Can't widen if the
iteration range overflows" check), which handle_loop_entry_exit() calls
before this one.

Could the comment name bpf_compute_loop_iters() so readers know where the
invariant that makes the verifier_bug() unreachable is established?

The comment above union_any_reg() has the same wording ("The filtering
pass has checked ...").

[ ... ]

> diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
> index 68df27e34b60..c1771a355e4e 100644
> --- a/kernel/bpf/states.c
> +++ b/kernel/bpf/states.c

[ ... ]

> @@ -1359,10 +1400,11 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
>  			continue;
>
>  		if (sl->state.branches) {
> -			struct bpf_func_state *frame = sl->state.frame[0];
> +			struct bpf_func_state *old_top_frame = sl->state.frame[0];
> +			struct bpf_func_state *old_cur_frame = sl->state.frame[sl->state.curframe];
>
> -			if (frame->in_async_callback_fn &&
> -			    frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
> +			if (old_top_frame->in_async_callback_fn &&
> +			    old_top_frame->async_entry_cnt != cur->frame[0]->async_entry_cnt) {
>  				/* Different async_entry_cnt means that the verifier is
>  				 * processing another entry into async callback.
>  				 * Seeing the same state is not an indication of infinite
> @@ -1461,6 +1503,20 @@ int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)
>  				}
>  				goto skip_inf_loop_check;
>  			}
> +
> +			/*
> +			 * If old state belongs to a control flow loop that we know terminates,
> +			 * it should be safe to prune current state.
> +			 */
> +			if (old_cur_frame->loop_stack_cnt &&
> +			    old_cur_frame->loop_stack[old_cur_frame->loop_stack_cnt - 1].terminates) {
> +				if (states_equal(env, &sl->state, cur, RANGE_WITHIN)) {
> +					loop = true;
> +					goto hit;
> +				}
> +				goto skip_inf_loop_check;
> +			}
> +

Pruning against a checkpoint that is still being explored (sl->state.branches > 0)
is only safe if every cycle leading from that checkpoint back to the
current state is proven to terminate.

The new check looks only at whether the innermost loop on the old state's
loop stack terminates. loop_stack_safe() (called from func_states_equal())
compares loop_id, iters.max_header_count, iters.pre_cond and terminates,
but it does not compare loop_stack_entry.entry_state, which is the thing
that identifies a particular invocation of the loop.

If branches > 0, cur descends from the old checkpoint, so some cycle joins
them. That cycle does not have to be the inner loop's backedge. It can
leave the terminating inner loop through an exit edge (loop_stack_pop()),
go around an enclosing loop that was pushed with terminates == false, and
enter the inner loop again.

The re-entry pushes a new entry with a new entry_state but, when the entry
values are the same, with the same loop_id/max_header_count/pre_cond/
terminates. The widened state W from the earlier invocation is still an
ancestor with branches > 0, so its top-of-stack entry says "terminates".
RANGE_WITHIN then prunes cur against W with loop = true.

Before this change, the same states went to the states_maybe_looping()/
EXACT infinite-loop check or ran into the complexity limit. Now the
non-terminating enclosing loop is accepted as converged, and its
termination is never established.

The iterator path just above has the same problem and handles it with an
identity check: "The loop converged only if it's the same iterator. The
one that was destroyed and created again starts from the beginning."
(iter_state->id == old_iter->id).

The SCEV path has no matching check, such as requiring the same
entry_state for the top entry, or requiring every enclosing entry on the
stack to be terminating. Nothing later in the series changes
kernel/bpf/states.c.

Could this also require the invocation to match, for example compare
entry_state of the top entry, or entry_state for every terminating entry
in loop_stack_safe() when used for in-progress pruning? The goal is that
cycles through an enclosing, unproven loop fall through to the existing
infinite-loop detection.

[ ... ]

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 57085f1bf915..9f56ffe09eba 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> +static bool has_entered_loop(struct bpf_verifier_env *env)
> +{
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	int prev_insn_idx = env->prev_insn_idx;
> +	int insn_idx = env->insn_idx;
> +
> +	return aux[insn_idx].loop_entry &&
> +	       (prev_insn_idx == -1 ||
> +		bpf_loop_at_index(env, prev_insn_idx) != bpf_loop_at_index(env, insn_idx));
> +}
> +
> +static bool is_next_loop_iteration(struct bpf_verifier_env *env)
> +{
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	int prev_insn_idx = env->prev_insn_idx;
> +	int insn_idx = env->insn_idx;
> +
> +	return aux[insn_idx].loop_entry &&
> +	       (prev_insn_idx != -1 &&
> +		bpf_loop_at_index(env, prev_insn_idx) == bpf_loop_at_index(env, insn_idx));
> +}

is_next_loop_iteration() and has_entered_loop() classify an edge as a
backedge only by comparing bpf_loop_at_index(prev_insn_idx) with
bpf_loop_at_index(insn_idx). They assume that if prev_insn_idx is in the
same loop, it is also in the same frame.

That does not hold for async callbacks, which may reference themselves
(check_helper_call/push_callback_call: "there is no real recursion here.
timer and workqueue callbacks are async").

The async state comes from push_async_cb(env,
env->subprog_info[subprog].start, insn_idx, ...). So when it is popped,
env->insn_idx is the callback's first instruction, env->prev_insn_idx is
the bpf_timer_set_callback()/bpf_wq_set_callback() call instruction, and
the frame is freshly allocated with loop_stack_cnt == 0.

Trigger: subprog S's first instruction is a loop header H (for example a
may_goto-bounded or otherwise bounded loop starting at S's entry), and S
calls bpf_timer_set_callback(timer, S) from inside loop H.

- aux[H].loop_entry is set, bpf_loop_at_index(call_insn) == H and
  bpf_loop_at_index(H) == H.
- So is_next_loop_iteration() returns true and handle_loop_entry_exit()
  calls maybe_clamp_scev_regs().
- On the empty loop stack of the fresh callback frame, that hits
  verifier_bug() and returns -EFAULT.
- has_entered_loop() returns false for the same edge, so H would never be
  pushed anyway. Without the first check, the new do_check() check "loop
  stack empty at %d, while inside the loop %d" would fire as well.

Before this patch, such a program verified normally. Now it is rejected
with -EFAULT, and a user-supplied program can trigger WARN_ONCE
(BPF_WARN_ONCE under CONFIG_DEBUG_KERNEL), which fuzzers report as a
kernel warning.

Could this treat any prev_insn_idx that belongs to a different frame or
callback entry as a loop entry? For example, check that the current
frame's loop stack contains bpf_loop_at_index(insn_idx), or push async and
callback states with prev_insn_idx = -1.

Alternatively, could has_entered_loop() become true whenever the header is
not on cur_func()->loop_stack?

[ ... ]

> +static bool has_exited_loop(struct bpf_verifier_env *env)
> +{
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +	int prev_insn_idx = env->prev_insn_idx;
> +	int insn_idx = env->insn_idx;
> +	int prev_loop_idx, h;
> +
> +	if (prev_insn_idx == -1)
> +		return false;
> +
> +	prev_loop_idx = bpf_loop_at_index(env, prev_insn_idx);
> +	if (prev_loop_idx < 0)
> +		return false;
> +
> +	/* Still inside the previous instruction's loop if that loop encloses insn_idx */
> +	for (h = bpf_loop_at_index(env, insn_idx); h >= 0; h = aux[h].loop_header)
> +		if (h == prev_loop_idx)
> +			return false;
> +
> +	return true;
> +}

has_entered_loop() and is_next_loop_iteration() test innermost-loop
equality, bpf_loop_at_index(prev) == bpf_loop_at_index(insn). They do not
test whether the loop at insn_idx encloses prev_insn_idx. has_exited_loop()
does walk the aux[h].loop_header chain, so the three helpers disagree.

Two kinds of real backedges into header H are therefore classified as new
loop entries:

- An edge whose source is inside a nested inner loop. For example, a
  jump-to-condition layout where the inner loop's latch falls through or
  jumps straight to the outer header: bpf_loop_at_index(prev) is the inner
  header, bpf_loop_at_index(insn) is H.
- A return from a bpf-to-bpf call whose call instruction is the loop's
  last body instruction and falls through to H: prev_insn_idx is the
  callee's exit instruction, so bpf_loop_at_index(prev) == -1.

What happens on such an edge in handle_loop_entry_exit():

- is_next_loop_iteration() is false, so bpf_clamp_scev_regs() is never
  applied on that backedge.
- has_entered_loop() is true, so bpf_compute_loop_iters() runs again. If
  it succeeds, the code does another bpf_split_cur_state() and
  bpf_widen_scev_regs() against the mid-loop state.
- loop_stack_push() then pushes nothing, because H is already the top of
  the stack.

Consequence: lockstep induction variables (registers that advance with the
counter but are not bounded by the loop condition) are widened at loop
entry but never clamped back to their SCEV bounds. Their ranges drift
further each iteration, so the state does not become RANGE_WITHIN the
widened checkpoint W.

The loop is then walked again with growing ranges instead of exact
constants. A program whose pre-patch walk used exact per-iteration
constants (for example j used as a bounded map-value or stack offset) can
now be falsely rejected, or can use many more states, plus one extra
checkpoint per misclassified backedge.

This is not a soundness problem: the widened values are supersets.
Suggested fix: classify the edge as a next iteration when
bpf_loop_at_index(insn_idx) is the header and it is on the current frame's
loop stack, or when that header encloses bpf_loop_at_index(prev_insn_idx)
through the aux[].loop_header chain, as has_exited_loop() already does.

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 19/36] bpf: compute loop hierarchy
  2026-09-26 14:20 ` [PATCH bpf-next 19/36] bpf: compute loop hierarchy Eduard Zingerman
@ 2026-09-27 20:43   ` bot+bpf-ci
  0 siblings, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:43 UTC (permalink / raw)
  To: eddyz87, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, eddyz87, memxor,
	ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 5725 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 5e6fa41a1353..d8142f1b2347 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h

[ ... ]

> @@ -751,6 +782,10 @@ struct bpf_insn_aux_data {
>  	u16 const_reg_map_mask;
>  	u16 const_reg_subprog_mask;
>  	u32 const_reg_vals[10];
> +	/* index of a loop header of the innermost loop containing this instruction, -1 if none */
> +	s32 loop_header;
> +	/* additional information about the loop if this instruction is a loop header */
> +	struct bpf_loop *loop;
>  };

Does the comment on loop_header match what the code stores for loop
header instructions?

Looking at assign_header() in kernel/bpf/loops.c, for a loop header
instruction, loop_header holds the header of the enclosing (parent)
loop, or -1 for an outermost loop:

    /* Don't encode self-loops, otherwise can't reflect loops nesting structure. */
    if (n == h)
        return 0;

The field does not hold the instruction's own index, even though the
header is itself part of the innermost loop that contains it.

Every consumer has to special-case headers with the pattern
'aux[idx].loop ? idx : aux[idx].loop_header' (seen in
bpf_loop_at_index() and twice in bpf_compute_loops()). bpf_is_nested_loop()
walks loop_header from a header to find its ancestors.

A reader who trusts the comment and uses aux[i].loop_header directly as
"the loop this insn belongs to" gets the wrong loop for every header
instruction.

Could the comment clarify that for loop headers, this field holds the
header of the enclosing loop? Something like: "for loop headers, the
header of the enclosing loop; for other instructions, the header of the
innermost loop containing them; -1 if none. Use bpf_loop_at_index() to
get the innermost loop."

[ ... ]

> diff --git a/kernel/bpf/loops.c b/kernel/bpf/loops.c
> index e7ff02e9ab3a..e269870f6a6e 100644
> --- a/kernel/bpf/loops.c
> +++ b/kernel/bpf/loops.c

[ ... ]

> +int bpf_loop_at_index(struct bpf_verifier_env *env, u32 idx)
> +{
> +	struct bpf_insn_aux_data *aux = env->insn_aux_data;
> +
> +	return aux[idx].loop ? idx : aux[idx].loop_header;
> +}
> +
> +static int find_dominating_condition(struct bpf_verifier_env *env, int n, int top)
> +{
> +	struct bpf_insn *insns = env->prog->insnsi;
> +	int common_dom, n_loop, t_loop, f_loop;
> +	int *idoms = env->idoms;
> +
> +	n_loop = bpf_loop_at_index(env, n);
> +	common_dom = idoms_intersect(env, n, top);
> +	if (common_dom != top)
> +		return -1;
> +	while (n >= 0) {
> +		if (is_cond_jmp_insn(&insns[n]) && bpf_loop_at_index(env, n) == n_loop) {
> +			t_loop = bpf_loop_at_index(env, n + insns[n].off + 1);
> +			f_loop = bpf_loop_at_index(env, n + 1);
> +			if (f_loop != n_loop && !bpf_is_nested_loop(env, f_loop, n_loop))
> +				return n;
> +			if (t_loop != n_loop && !bpf_is_nested_loop(env, t_loop, n_loop))
> +				return n;
> +		}
> +		if (n == top)
> +			break;
> +		n = idoms[n];
> +	}
> +	return -1;
> +}

Does find_dominating_condition() return the correct latch when the
backedge to 'top' comes from inside a nested loop?

The reference loop is taken from the backedge source:

    n_loop = bpf_loop_at_index(env, n);

When the backedge to 'top' comes from inside a nested loop (for example
a 'continue' of the outer loop issued from the inner loop body), n_loop
is the inner loop. The function then returns a conditional jump that
only leaves the inner loop, and stays inside 'top' on both of its
targets. Example:

    2: r6 += 1                ; H (outer header)
    3: r2 += 1                ; L2 (inner header)
    4: if r2 == 3 goto 2      ; H's only backedge, 4 -> 2
    5: if r6 > 100 goto 7     ; the only exit of H
    6: goto 3                 ; L2 backedge, 6 -> 3
    7: exit

The loop pass computes loop_header[3] = 2 and loop_header[4..6] = 3.

find_dominating_condition(env, 4, 2) sets n_loop = 3 and
idoms_intersect(4, 2) == 2.

At n = 4:
  f_loop = loop_at(5) = 3 == n_loop
  t_loop = loop_at(2) = 2
  bpf_is_nested_loop(env, 2, 3) is false

So it returns 4 as H's latch. Neither successor of insn 4 (2 and 5)
leaves H, so insn 4 does not decide whether execution stays in H.

With 'top' as the reference loop, the walk would visit 4, 3 and 2
without finding an exiting condition and return -1.

In this commit the wrong latch only appears in the verbose log. The SCEV
code added later in the same series trusts it, though.

bpf_compute_loop_iters() takes 'latch_idx = loop->backedges[0].latch'
for the header without checking that bpf_loop_at_index(latch_idx) ==
header. match_linear_latch() then evaluates the condition in the inner
loop's environment and flips the opcode because the true target is
"outside" the inner loop. compute_max_iters() documents the assumption
that "The latch is a conditional jump with one jump target exiting the
loop." As a result, the outer loop's max_header_count can be derived
from an induction variable of the inner loop that does not bound the
outer loop. Registers at the outer header are then widened and clamped
to that bound, which risks verifier unsoundness.

Should n_loop use 'top' as the reference loop instead?

[ ... ]

> @@ -22820,6 +22824,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
>  	if (ret < 0)
>  		goto skip_full_check;
>
> +	ret = bpf_compute_loops(env);
> +	if (ret < 0)
> +		goto skip_full_check;
> +
>  	ret = bpf_compute_live_registers(env);
>  	if (ret < 0)
>  		goto skip_full_check;

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 03/36] bpf: summarize live stack slots in insn_aux_data
  2026-09-27 20:26   ` bot+bpf-ci
@ 2026-09-29 18:16     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-29 18:16 UTC (permalink / raw)
  To: bot+bpf-ci, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor,
	martin.lau, mason, ihor.solodrai

On Sun, 2026-09-27 at 20:26 +0000, bot+bpf-ci@kernel.org wrote:
> > diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> > index c6d617581e84..7f31ce5ea6b7 100644
> > --- a/include/linux/bpf_verifier.h
> > +++ b/include/linux/bpf_verifier.h
> > @@ -667,6 +667,7 @@ struct bpf_insn_aux_data {
> >  	 * that this instruction may write to.
> >  	 */
> >  	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
> > +	DECLARE_BITMAP(live_stack_before, MAX_BPF_STACK_SLOTS);
> 
> This isn't a bug, but would a short comment on live_stack_before make it
> clearer what the field holds?  Nearly every neighbouring field like
> live_regs_before and scc has its own comment.

Naming is clear enough.

> 
> >  	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
> >  	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
> >  	bool nospec; /* do not execute this instruction speculatively */
> 
> [ ... ]
> 
> > diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> > index c871744ca5a8..b6b7fd479569 100644
> > --- a/kernel/bpf/liveness.c
> > +++ b/kernel/bpf/liveness.c
> > @@ -729,8 +729,9 @@ static void half_spis_to_slots(unsigned long *slots, const unsigned long *mask,
> >  }
> > 
> >  /*
> > - * Precompute, for each instruction, the OR of may_write masks over its top
> > - * frame across all func_instances reaching it, stash it in the insn_aux_data.
> > + * Precompute, for each instruction, the OR of may_write and live_before masks
> > + * over its top frame across all func_instances reaching it, stash them in the
> > + * insn_aux_data.
> >   */
> >  static void compute_may_write_masks(struct bpf_verifier_env *env)
> >  {
> > @@ -746,9 +747,12 @@ static void compute_may_write_masks(struct bpf_verifier_env *env)
> >  		if (!fm)
> >  			continue;
> >  		nbits = frame_mask_bits(fm);
> > -		for (i = 0; i < instance->insn_cnt; i++)
> > +		for (i = 0; i < instance->insn_cnt; i++) {
> >  			half_spis_to_slots(aux[instance->subprog_start + i].may_write_mask,
> >  					   rel_mask(fm, i, FM_MAY_WRITE), nbits);
> > +			half_spis_to_slots(aux[instance->subprog_start + i].live_stack_before,
> > +					   rel_mask(fm, i, FM_LIVE_BEFORE), nbits);
> > +		}
> >  	}
> >  }
> 
> This isn't a bug, but would a more general name for compute_may_write_masks()
> make the function's purpose clearer?  The comment now describes computing both
> may_write and live_before masks, but the name only mentions may_write.

Nope.

> 
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
> 
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346543957

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data
  2026-09-26 15:51   ` Alexei Starovoitov
@ 2026-09-29 20:12     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-29 20:12 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor

On Sat, 2026-09-26 at 15:51 +0000, Alexei Starovoitov wrote:
> On Sat, Sep 26, 2026 at 07:20 AM Eduard Zingerman <eddyz87@gmail.com> wrote:
> > @@ -662,6 +662,11 @@ struct bpf_insn_aux_data {
> >  	};
> >  	struct btf_struct_meta *kptr_struct_meta;
> >  	u64 map_key_state; /* constant (32 bit) key tracking for maps */
> > +	/*
> > +	 * Per-instruction summary of stack slots in the current frame
> > +	 * that this instruction may write to.
> > +	 */
> > +	DECLARE_BITMAP(may_write_mask, MAX_BPF_STACK_SLOTS);
> 
> That's 32 bytes per insn and the next patch adds 32 more,
> for every prog whether it has loops or not.
> FM_MAY_WRITE in patch 1 makes every frame_masks a third bigger too.
> commit 481ceda77aeb sized the liveness masks by the stack the frame
> uses to avoid exactly that.
> 
> bpf_may_write_mask() is the only accessor.
> Can the summary stay in liveness.c, as wide as the subprog's stack,
> and only for insns with scc != 0 ?

I did some profiling and may_write_mask does not show as a big hog in
total memory consumption increase. (+50% is the biggest observed
increase). The biggest increase is from a loop_stack_entry array in
bpf_func_state, and from a fact that I pushed bpf_func_state to a 2KB
allocation bucket. I'll see what can be done there.

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer in insn_aux_data
  2026-09-27 20:27   ` bot+bpf-ci
@ 2026-09-29 20:21     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-29 20:21 UTC (permalink / raw)
  To: bot+bpf-ci, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor,
	martin.lau, mason, ihor.solodrai

On Sun, 2026-09-27 at 20:27 +0000, bot+bpf-ci@kernel.org wrote:
> > bpf: summarize regs that may hold a frame pointer in insn_aux_data
> > 
> > SCEV needs to know which registers might be stack pointers at a
> > particular instruction. This information is used to invalidate SCEV
> > expressions for slots that might be overwritten by indirect writes.
> > 
> > liveness.c:compute_may_write_masks() already tracks this information.
> > This commit modifies it to save the information in insn_aux_data for
> > further usage.
> 
> Are these two claims accurate?

Commit message needs to be fixed.

...

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c
  2026-09-27 20:27   ` bot+bpf-ci
@ 2026-09-29 20:26     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-29 20:26 UTC (permalink / raw)
  To: bot+bpf-ci, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor,
	martin.lau, mason, ihor.solodrai

On Sun, 2026-09-27 at 20:27 +0000, bot+bpf-ci@kernel.org wrote:
> > bpf: record write effects for atomic operations in liveness.c
> > 
> > Atomic read-modify-write instructions are recorded as reads only,
> > omitting their stack write effects from the may_write masks needed
> > by SCEV.
> > 
> > Record both read and write accesses for atomic RMW operations. Keep
> > LOAD_ACQ read-only and STORE_REL write-only. Share the precise/imprecise
> > stack-access dispatch so both accesses use the same address handling.
> > 
> > Fixes: fed53dbcdb61 ("bpf: record arg tracking results in bpf_liveness masks")
> 
> Does the Fixes: tag point at the right commit?

Hm, fixes tag is not needed.
The original code returns negative size only for STORE_REL,
which is fine.

...

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear()
  2026-09-26 14:34   ` sashiko-bot
@ 2026-09-29 21:46     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-29 21:46 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf

On Sat, 2026-09-26 at 14:34 +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: cnum64_intersect_linear over-expands intervals that wrap the signed boundary
> --
> 
> commit 07e53bba30487c590757dc2ec233d2811a607b9a
> Author: Eduard Zingerman <eddyz87@gmail.com>
> 
> bpf: add cnum64_intersect_linear()
> 
> This commit introduces cnum64_intersect_linear() to intersect a cnum64
> interval with integers congruent to 'base' modulo 'step'. It tightens the
> interval using signed bounds to keep residues consistent. This is a building
> block for tracking scalar registers with values known to lie on a line.
> 
> > diff --git a/kernel/bpf/cnum.c b/kernel/bpf/cnum.c
> > index 86142cb2aee5f..2bff2c7e0cdf5 100644
> > --- a/kernel/bpf/cnum.c
> > +++ b/kernel/bpf/cnum.c
> [ ... ]
> > @@ -118,3 +119,38 @@ struct cnum64 cnum64_cnum32_intersect(struct cnum64 a, struct cnum32 b)
> >  	}
> >  	return t;
> >  }
> > +
> > +/* Intersect 'in' with the set of integers defined by equation 'base + step * k'. */
> > +struct cnum64 cnum64_intersect_linear(struct cnum64 in, u16 base, u16 step)
> > +{
> > +	s64 smin = cnum64_smin(in);
> > +	s64 smax = cnum64_smax(in);
> > +	s64 lo, hi;
> > +	u16 d;
> > +
> > +	if (step <= 1 || cnum64_is_empty(in))
> > +		return in;
> > +	/*
> > +	 * Round smin up to the next value congruent to 'base' modulo 'step',
> > +	 * i.e. increase smin by d = (base - smin) mod step:
> > +	 *
> > +	 *                 |<---- d ---->|
> > +	 *     |-----------|=============|...
> > +	 * base+step*k    smin       base+step*(k+1)
> > +	 */
> > +	d = imod(base - imod(smin, step), step);
> > +	if ((u64)smax - (u64)smin < d)
> > +		return CNUM64_EMPTY;
> > +	lo = smin + d;
> > +	/*
> > +	 * Round smax down to the previous value congruent to 'base' modulo 'step',
> > +	 * i.e. decrease smax by d = (smax - base) mod step:
> > +	 *
> > +	 *     |<--- d --->|
> > +	 *  ...|===========|-------------|
> > +	 * base+step*k    smax       base+step*(k+1)
> > +	 */
> > +	d = imod(imod(smax, step) - base, step);
> > +	hi = smax - d;
> > +	return cnum64_from_srange(lo, hi);
> 
> [Severity: Medium]
> Does this code improperly expand intervals that wrap the signed boundary?
> 
> For such intervals, cnum64_smin() returns S64_MIN and cnum64_smax() returns
> S64_MAX. The function then calculates 'lo' near S64_MIN and 'hi' near
> S64_MAX, returning cnum64_from_srange(lo, hi).

I'll add `|| cnum64_srange_overflow(in)` to the entry condition.

> This replaces a potentially small wrapping interval with a massive contiguous
> range spanning nearly the entire 64-bit space, violating the intersection
> property (A ∩ B ⊆ A). While this over-approximation is sound, it could cause
> the verifier to lose precision and falsely reject valid BPF programs.

> > +}

^ permalink raw reply	[flat|nested] 80+ messages in thread

* Re: [PATCH bpf-next 11/36] bpf: add bpf_reg_union()
  2026-09-27 20:42   ` bot+bpf-ci
@ 2026-09-30  0:09     ` Eduard Zingerman
  0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-30  0:09 UTC (permalink / raw)
  To: bot+bpf-ci, bpf, ast, andrii
  Cc: daniel, martin.lau, kernel-team, yonghong.song, memxor,
	martin.lau, mason, ihor.solodrai

On Sun, 2026-09-27 at 20:42 +0000, bot+bpf-ci@kernel.org wrote:
> > diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> > index 719b7c7fd2f5..706fdefbc07a 100644
> > --- a/include/linux/bpf_verifier.h
> > +++ b/include/linux/bpf_verifier.h
> > @@ -1748,6 +1748,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env,
> >  int bpf_compute_subprog_arg_access(struct bpf_verifier_env *env);
> >  int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
> >  		      struct cnum64 range, u16 step);
> > +int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
> > +		  const struct bpf_reg_state *src);
> >
> >  int bpf_stack_liveness_init(struct bpf_verifier_env *env);
> >  void bpf_stack_liveness_free(struct bpf_verifier_env *env);
> > diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> > index b5f45cb46f14..1ec44d72a346 100644
> > --- a/kernel/bpf/verifier.c
> > +++ b/kernel/bpf/verifier.c
> > @@ -30,6 +30,7 @@
> >  #include <linux/module.h>
> >  #include <linux/cpumask.h>
> >  #include <linux/cnum.h>
> > +#include <linux/gcd.h>
> >  #include <linux/bpf_mem_alloc.h>
> >  #include <net/xdp.h>
> >  #include <linux/trace_events.h>
>
> [ ... ]
>
> > @@ -16952,6 +16953,33 @@ int bpf_set_reg_range(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
> >  	return reg_bounds_sanity_check(env, reg, "bpf_set_reg_range");
> >  }
> >
> > +/* acc := acc U src, matching types only. Caller must clear acc's scalar ID. */
> > +int bpf_reg_union(struct bpf_verifier_env *env, struct bpf_reg_state *acc,
> > +		  const struct bpf_reg_state *src)
> > +{
> > +	u16 base, step;
> > +
> > +	if (acc->type != src->type) {
> > +		verifier_bug(env, "union of registers with different types");
> > +		return -EFAULT;
> > +	}
> > +	acc->r64 = cnum64_union(acc->r64, src->r64);
> > +	acc->r32 = cnum32_union(acc->r32, src->r32);
> > +	acc->var_off = tnum_union(acc->var_off, src->var_off);
>
> Does this function handle pointer registers correctly? When bpf_reg_union()
> widens a PTR_TO_PACKET's offset (r64/r32/var_off) but keeps its id, it can
> break the id-tracking invariant.
>
> For PTR_TO_PACKET, registers with the same id differ only by a known
> constant. When adjust_ptr_min_max_vals() adds a non-constant to a packet
> pointer, it gives it a fresh id:
>
>   if (!known)
>       dst_reg->id = ++env->id_gen;

The example is an utter hallucinated garbage, but the point is valid.
Here is a valid example that should be rejected:

  r7 = ctx->data;
  r8 = r7
  for (r6 = 0; r6 < 10 && random() != 42; r7++, r6++);
  if (r7 >= ctx->data_end)
    return;
  // at this point the loop executed unknown number of times
  // hence r7 range gives no information about r8
  *(r8 + 4);

There are several other types in is_widenable_reg_type() to consider.
So far I think that it is necessary to reset the ids for the following
types: SCALAR_VALUE, PTR_TO_PACKET, PTR_TO_PACKET_META

> find_good_pkt_pointers() depends on this rule. After a successful
> 'pkt_ptr > pkt_end' test, it uses reg_umax(dst_reg) as the new range
> and sets reg->range = max(reg->range, new_range) on every PTR_TO_PACKET
> with the same id, assuming they all track the same maximum offset.
>
> Once bpf_reg_union() enlarges acc's umax but leaves acc in the same id
> group as registers that were not widened, a comparison on acc can give
> those other registers a range based on the widened umax that nothing at
> runtime backs.
>
> The comment says "Caller must clear acc's scalar ID", but for pointer
> types the id has a different meaning. Later in the series when this is
> called from union_any_reg() (1987b0e52b34), is_widenable_reg_type()
> accepts PTR_TO_PACKET, and scratch_scalar_id() only clears the id when
> reg->type == SCALAR_VALUE.
>
> Here's a concrete example with a similar shape to the
> conditional_stack_pointer_assignment selftest at the series end:
>
>   r2 = *(u32 *)(r1 + 0)   // data: pkt, id=0
>   r3 = *(u32 *)(r1 + 4)   // data_end
>   r9 = r2                 // pkt, id=0, off 0
>   r7 = r2
>   r7 += 1000              // pkt, id=0 (constant add keeps id)
>   r6 = 0
> 1: if r6 >= 3 goto 3f
>   if r2 > r3 goto 3f      // fallthrough: find_good_pkt_pointers(r2)
>   r0 = *(u8 *)(r9 + 999)
>   if r6 != 1 goto 2f
>   r2 = r7                 // r2 = (any r2 r7) at loop entry
> 2: r6 += 1
>   goto 1b
> 3: r0 = 0
>   exit
>
> At loop entry, bpf_reg_union() turns r2 into pkt(id=0, [0,1000]), but r9
> keeps id=0 and [0,0]. On the fallthrough of 'if r2 > r3',
> find_good_pkt_pointers() sets r9->range = reg_umax(r2) = 1000. The load
> from r9 + 999 then passes (0 + 999 + 1 <= 1000).

^ permalink raw reply	[flat|nested] 80+ messages in thread

end of thread, other threads:[~2026-09-30  0:09 UTC | newest]

Thread overview: 80+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-26 14:19 [PATCH bpf-next 00/36] bpf: use scalar evolution to widen bounded loops Eduard Zingerman
2026-09-26 14:19 ` [PATCH bpf-next 01/36] bpf: track may_write flags in liveness Eduard Zingerman
2026-09-26 15:51   ` Alexei Starovoitov
2026-09-27  8:43     ` Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 02/36] bpf: summarize may write stack slots in insn_aux_data Eduard Zingerman
2026-09-26 15:51   ` Alexei Starovoitov
2026-09-29 20:12     ` Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 03/36] bpf: summarize live " Eduard Zingerman
2026-09-26 14:33   ` sashiko-bot
2026-09-27 20:26   ` bot+bpf-ci
2026-09-29 18:16     ` Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 04/36] bpf: summarize regs that may hold a frame pointer " Eduard Zingerman
2026-09-27 20:27   ` bot+bpf-ci
2026-09-29 20:21     ` Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 05/36] bpf: record write effects for atomic operations in liveness.c Eduard Zingerman
2026-09-27 20:27   ` bot+bpf-ci
2026-09-29 20:26     ` Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 06/36] bpf: add tnum_alignment() Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 07/36] bpf: add cnum{32,64}_union() Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 08/36] bpf: add cnum64_intersect_linear() Eduard Zingerman
2026-09-26 14:34   ` sashiko-bot
2026-09-29 21:46     ` Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 09/36] bpf: add bpf_set_reg_range() Eduard Zingerman
2026-09-26 14:36   ` sashiko-bot
2026-09-26 14:20 ` [PATCH bpf-next 10/36] bpf: add bpf_mark_reg_known_scalar() Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 11/36] bpf: add bpf_reg_union() Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci
2026-09-30  0:09     ` Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 12/36] bpf: expose comparison opcode transformations Eduard Zingerman
2026-09-27 20:26   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 13/36] bpf: allow subrange relations for PTR_TO_STACK in regsafe() Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 14/36] bpf: representation for intervals with steps Eduard Zingerman
2026-09-26 14:35   ` sashiko-bot
2026-09-26 14:20 ` [PATCH bpf-next 15/36] bpf: varying offset access support for PTR_TO_BTF_ID pointers Eduard Zingerman
2026-09-26 14:37   ` sashiko-bot
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 16/36] bpf: save DFS postorder numbers for program instructions Eduard Zingerman
2026-09-26 14:31   ` sashiko-bot
2026-09-26 14:20 ` [PATCH bpf-next 17/36] bpf: move the live-register and SCC printout to a standalone function Eduard Zingerman
2026-09-27 20:26   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 18/36] bpf: compute immediate dominators Eduard Zingerman
2026-09-26 15:54   ` Alexei Starovoitov
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 19/36] bpf: compute loop hierarchy Eduard Zingerman
2026-09-27 20:43   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 20/36] bpf: add a min-heap for ordered analysis worklists Eduard Zingerman
2026-09-27 20:26   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 21/36] bpf: record basic-block ends in insn_aux_data Eduard Zingerman
2026-09-27 20:26   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 22/36] bpf: add bpf_split_cur_state() Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 23/36] bpf: allow precision backtracking between overlapping checkpoints Eduard Zingerman
2026-09-27 20:27   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 24/36] bpf: compute scalar evolution expressions for loops Eduard Zingerman
2026-09-26 14:38   ` sashiko-bot
2026-09-27 20:43   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 25/36] bpf: use SCEV to widen bounded loops Eduard Zingerman
2026-09-26 14:42   ` sashiko-bot
2026-09-27 20:43   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 26/36] bpf: avoid widening registers that hinder exact stack-slot tracking Eduard Zingerman
2026-09-26 14:46   ` sashiko-bot
2026-09-27 20:43   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 27/36] selftests/bpf: __msg_next tag for matching messages on consecutive lines Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 28/36] selftests/bpf: test for stack-pointer subrange pruning Eduard Zingerman
2026-09-26 14:32   ` sashiko-bot
2026-09-27 20:26   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 29/36] selftests/bpf: tests for may_write stack-liveness tracking Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 30/36] selftests/bpf: tests for may_def marks of atomic RMW operations Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 31/36] selftests/bpf: tests for register base/step arithmetic Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 32/36] selftests/bpf: tests for register base/step state pruning Eduard Zingerman
2026-09-27 20:27   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 33/36] selftests/bpf: tests for varying offset access to PTR_TO_BTF_ID Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 34/36] selftests/bpf: tests for loop hierarchy computation Eduard Zingerman
2026-09-26 14:20 ` [PATCH bpf-next 35/36] selftests/bpf: tests for immediate dominator computation Eduard Zingerman
2026-09-27 20:27   ` bot+bpf-ci
2026-09-26 14:20 ` [PATCH bpf-next 36/36] selftests/bpf: cover SCEV analysis and loop widening Eduard Zingerman
2026-09-27 20:42   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox