Git development
 help / color / mirror / Atom feed
* [PATCH 0/4] faster SHA-1 collision detection
@ 2026-09-29 11:25 Scott Chacon
  2026-09-29 11:25 ` [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX Scott Chacon
                   ` (5 more replies)
  0 siblings, 6 replies; 21+ messages in thread
From: Scott Chacon @ 2026-09-29 11:25 UTC (permalink / raw)
  To: git

So, spoiler alert, the code in this patch series is mainly AI generated.
I would try to fool you, but too many of you are far too aware of my
actual C skills. That being said, I thought maybe someone here (especially
those of you working on server optimization stuff) would be interested
in the speed increases for both the server and client in making sha1dc
quite a bit faster.

This series ports the approach of Sam Reis's sha1dc Rust crate [1], 
which gitoxide recently switched to [2], to C. 

The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.

Hashing throughput on the Xeon, in MiB/s:

                                16KiB    1MiB   vs OpenSSL
  OpenSSL SHA-1 (no detection)   1234    1129      1.00x
  sha1dc/ (today)                 435     450      2.67x
  shani+avx2 (default here)      1002     901      1.24x
  shani+sse2                     1075    1008      1.13x
  portable+avx2                   553     654      1.96x
  portable+sse2                   603     681      1.84x
  portable                        466     565      2.29x

In other words, currently collision detection costs about 1.5–2.5x on
top of the hashing itself today, but only about 0.2x with the series. 

The patches are:

  [1/4]: sha1dc-accel: add a block loop for sha1dc's SHA1_CTX

    Just groundwork: our own block loop around sha1dc's context and DV
    table, with the same results and a few percent slower, plus tests
    that compare against sha1dc/ directly, including on real collisions
    in every mode.

  [2/4]: sha1dc-accel: vectorize the unavoidable-bitconditions check

    The UBC filter rewritten as SSE2, AVX2, NEON, and new scalar forms, 
    using the conditions the crate's solver picks for each. They're 
    carried as tables, with a short loop per form to run them. 
    1.29x on the Xeon, 1.27x on the M5 Max.

  [3/4]: sha1dc-accel: compress with SHA-NI on x86-64

    Hardware compression, with the schedule spilled, and recompression
    of flagged blocks in hardware, too. Another 2.07x on the Xeon.

  [4/4]: sha1dc-accel: compress with the ARMv8 SHA-1 instructions

    The same for arm64. Another 2.4x on the M5 Max.

The x86 numbers are from a 4-vCPU Xeon VM with SHA-NI and AVX2 (GCC 13,
Linux), which is unfortunately rather noisy; the per-patch hyperfine
output has the spread. The arm64 numbers are medians of 9 runs on an
Apple M5 Max (Apple clang, macOS). The full test suite passes on both.

[1] https://sam.dev/blog/faster-sha1-collision-detection
[2] https://github.com/GitoxideLabs/gitoxide/pull/3008

Scott Chacon (4):
  sha1dc-accel: add a block loop for sha1dc's SHA1_CTX
  sha1dc-accel: vectorize the unavoidable-bitconditions check
  sha1dc-accel: compress with SHA-NI on x86-64
  sha1dc-accel: compress with the ARMv8 SHA-1 instructions

 Makefile                            |   14 +
 contrib/buildsystems/CMakeLists.txt |    2 +-
 meson.build                         |    4 +
 sha1dc-accel/arm.c                  |  274 ++++
 sha1dc-accel/internal.h             |  126 ++
 sha1dc-accel/sha1.c                 |  498 ++++++++
 sha1dc-accel/sha1.h                 |   31 +
 sha1dc-accel/ubc_check.c            | 1789 +++++++++++++++++++++++++++
 sha1dc-accel/x86.c                  |  260 ++++
 sha1dc_git.c                        |   18 +
 t/.gitattributes                    |    1 +
 t/helper/test-sha1.c                |   95 ++
 t/helper/test-tool.c                |    2 +
 t/helper/test-tool.h                |    2 +
 t/meson.build                       |    1 +
 t/t0013-sha1dc.sh                   |   47 +
 t/t0013/sha-mbles-1.bin             |  Bin 0 -> 640 bytes
 t/t0013/sha1-reduced-round.bin      |  Bin 0 -> 128 bytes
 t/unit-tests/u-sha1dc.c             |  366 ++++++
 19 files changed, 3529 insertions(+), 1 deletion(-)
 create mode 100644 sha1dc-accel/arm.c
 create mode 100644 sha1dc-accel/internal.h
 create mode 100644 sha1dc-accel/sha1.c
 create mode 100644 sha1dc-accel/sha1.h
 create mode 100644 sha1dc-accel/ubc_check.c
 create mode 100644 sha1dc-accel/x86.c
 create mode 100644 t/t0013/sha-mbles-1.bin
 create mode 100644 t/t0013/sha1-reduced-round.bin
 create mode 100644 t/unit-tests/u-sha1dc.c


base-commit: a018953688f1b10bddf91bff8747068f5f4746a4
-- 
2.50.1 (Apple Git-155)



^ permalink raw reply	[flat|nested] 21+ messages in thread

* [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
@ 2026-09-29 11:25 ` Scott Chacon
  2026-10-07 12:17   ` Johannes Schindelin
  2026-09-29 11:25 ` [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check Scott Chacon
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 21+ messages in thread
From: Scott Chacon @ 2026-09-29 11:25 UTC (permalink / raw)
  To: git

Unless you build with some other SHA-1 implementation, every object we
hash goes through sha1collisiondetection. That's what we want for
safety, but it's not free. On the machine I'm testing with (a 4-vCPU
Xeon VM, which has SHA-NI and AVX2), sha1dc hashes at about 400 MiB/s,
while OpenSSL's SHA-1 (which does no detection) does about 1100 MiB/s.
On an Apple M5 Max, sha1dc manages about 700 MiB/s.
That matters for anything that hashes a lot of data we don't trust,
like index-pack.

Sam Reis recently wrote about closing that gap for gitoxide [1], with a
from-scratch rewrite of the detection code as a Rust crate [2]. There
are really three separate tricks in there:

  1. Compress with the CPU's SHA-1 instructions, and spill the expanded
     message schedule as we go. That's all the detection needs from the
     vast majority of blocks.

  2. The "unavoidable bitconditions" (UBC) filter, which rules out ~95%
     of blocks before we get to the expensive part, is most of what
     detection costs. Its conditions can be rewritten into an
     equivalent set that lines up in vector lanes, and the crate has a
     solver that does that for each instruction set.

  3. Recompress the few blocks the filter can't rule out with the SHA-1
     instructions, too.

We could in theory use the crate itself through our Rust support. But
that support is optional, the crate needs a much newer Rust than we
allow, and none of the three is all that much code once you know what
it's doing. So let's do it in C, which helps every build.

We can't really do any of that inside sha1dc/ itself. It's third-party
code (optionally a submodule), and its block loop is built around its
scalar compression, which stores intermediate states as it goes. So this
patch adds sha1dc-accel/, which has its own block loop but works on
sha1dc's SHA1_CTX and uses its table of disturbance vectors (DVs). For
each block it asks a "backend" to compress it (spilling the schedule and
the two states that recompression starts from), runs the UBC filter,
and then recompresses the block for each DV the filter didn't rule out.

For now there's only one backend: a portable C compression, plus
sha1dc's own ubc_check(). So this is just a reshuffling of what sha1dc/
already does, and the speedups come in the following patches. It comes
out a little slower. With 256MB of random data on the Xeon (GCC 13,
pinned to one CPU):

  Benchmark 1: test-tool.base sha1
    Time (mean ± σ):     727.3 ms ±  60.8 ms    [User: 677.2 ms, System: 40.7 ms]
    Range (min … max):   607.5 ms … 837.0 ms    30 runs

  Benchmark 2: test-tool.new sha1
    Time (mean ± σ):     797.2 ms ± 115.5 ms    [User: 745.4 ms, System: 41.8 ms]
    Range (min … max):   630.5 ms … 1078.0 ms    30 runs

  Summary
    test-tool.base sha1 ran
      1.10 ± 0.18 times faster than test-tool.new sha1

That's within the noise of this machine (which is quite noisy, sadly),
but callgrind counts 2.3% more instructions for the new code, and on the
M5 Max (Apple clang), hashing 1GiB goes from 1.49s to 1.55s (median of
9 runs), about 4% slower. That's a price I'm willing to pay for what
comes next.

A few other notes:

  - The results must be identical to sha1dc/'s: the same digest, and
    the same verdict on collisions. That includes the modes that Git
    itself doesn't use (the "safe hash" mitigation, disabling the UBC
    filter, and detecting reduced-round collisions). The new unit test
    checks exactly that against sha1dc/ for every backend the machine
    can run, on random data. There's only one so far, but there will be
    more. It also checks that sha1dc/'s table of DVs is laid out the way
    we expect, so that a change there fails loudly.

  - Likewise, "test-tool sha1dc-backends" lists the backends, and
    GIT_TEST_SHA1DC_BACKEND picks one for "test-tool sha1", which lets
    t0013 run its collision test against each of them. While I was
    there I added a second collision, the SHA-mbles chosen-prefix one
    [3]. It trips the same disturbance vector as SHAttered, II(52,0),
    but it's a different kind of attack, and gets detected in the
    tenth block rather than the fifth.

  - "test-tool sha1dc-compare" hashes a file both ways, in every
    combination of those modes and with the input split around each
    block, and dies if the two ever disagree. t0013 runs it for each
    backend on SHAttered, SHA-mbles and a reduced-round collision (from
    the crate's tests), which is the only way to exercise what happens
    after a collision is found. The last has no NUL byte, so
    t/.gitattributes marks t0013's .bin files as binary.

  - The backend is picked once, on first use. Threads can race to pick
    it, but they'd all store the same value; with GCC and clang we use
    relaxed atomics for that. Other compilers only get the portable
    backend, so they have nothing to pick, and no shared state.

  - DC_SHA1_EXTERNAL builds are left alone, since we don't know what
    their context looks like. For everybody else, DC_SHA1_NO_ACCEL
    goes back to using sha1dc/ on its own.

[1] https://sam.dev/blog/faster-sha1-collision-detection
[2] https://github.com/srijs/sha1dc
[3] https://sha-mbles.github.io/

Signed-off-by: Scott Chacon <scott@gitbutler.net>
Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
---
 Makefile                            |  10 +
 contrib/buildsystems/CMakeLists.txt |   2 +-
 meson.build                         |   1 +
 sha1dc-accel/internal.h             |  25 ++
 sha1dc-accel/sha1.c                 | 434 ++++++++++++++++++++++++++++
 sha1dc-accel/sha1.h                 |  31 ++
 sha1dc_git.c                        |  18 ++
 t/.gitattributes                    |   1 +
 t/helper/test-sha1.c                |  95 ++++++
 t/helper/test-tool.c                |   2 +
 t/helper/test-tool.h                |   2 +
 t/meson.build                       |   1 +
 t/t0013-sha1dc.sh                   |  47 +++
 t/t0013/sha-mbles-1.bin             | Bin 0 -> 640 bytes
 t/t0013/sha1-reduced-round.bin      | Bin 0 -> 128 bytes
 t/unit-tests/u-sha1dc.c             | 145 ++++++++++
 16 files changed, 813 insertions(+), 1 deletion(-)
 create mode 100644 sha1dc-accel/internal.h
 create mode 100644 sha1dc-accel/sha1.c
 create mode 100644 sha1dc-accel/sha1.h
 create mode 100644 t/t0013/sha-mbles-1.bin
 create mode 100644 t/t0013/sha1-reduced-round.bin
 create mode 100644 t/unit-tests/u-sha1dc.c

diff --git a/Makefile b/Makefile
index c649c93c51..9ab13ca2ab 100644
--- a/Makefile
+++ b/Makefile
@@ -567,6 +567,10 @@ include shared.mak
 # by the git project to migrate to using sha1collisiondetection as a
 # submodule.
 #
+# Unless DC_SHA1_EXTERNAL is defined, the built-in code is driven by the
+# block loop in sha1dc-accel/, which gives the same results.
+# Define DC_SHA1_NO_ACCEL to use the sha1collisiondetection code alone.
+#
 # === SHA-256 backend ===
 #
 # ==== Security ====
@@ -1555,6 +1559,7 @@ CLAR_TEST_SUITES += u-reftable-readwrite
 CLAR_TEST_SUITES += u-reftable-stack
 CLAR_TEST_SUITES += u-reftable-table
 CLAR_TEST_SUITES += u-reftable-tree
+CLAR_TEST_SUITES += u-sha1dc
 CLAR_TEST_SUITES += u-strbuf
 CLAR_TEST_SUITES += u-strcmp-offset
 CLAR_TEST_SUITES += u-string-list
@@ -2167,6 +2172,11 @@ ifdef DC_SHA1_SUBMODULE
 else
 	LIB_OBJS += sha1dc/sha1.o
 	LIB_OBJS += sha1dc/ubc_check.o
+endif
+ifdef DC_SHA1_NO_ACCEL
+	BASIC_CFLAGS += -DDC_SHA1_NO_ACCEL
+else
+	LIB_OBJS += sha1dc-accel/sha1.o
 endif
 	BASIC_CFLAGS += \
 		-DSHA1DC_NO_STANDARD_INCLUDES \
diff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt
index 7874e5a326..3c0ea2a27c 100644
--- a/contrib/buildsystems/CMakeLists.txt
+++ b/contrib/buildsystems/CMakeLists.txt
@@ -218,7 +218,7 @@ add_compile_definitions(NO_OPENSSL SHA1_DC SHA1DC_NO_STANDARD_INCLUDES
 			SHA1DC_INIT_SAFE_HASH_DEFAULT=0
 			SHA1DC_CUSTOM_INCLUDE_SHA1_C="git-compat-util.h"
 			SHA1DC_CUSTOM_INCLUDE_UBC_CHECK_C="git-compat-util.h" )
-list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
+list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c sha1dc-accel/sha1.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
 
 
 add_compile_definitions(PAGER_ENV="LESS=FRX LV=-c"
diff --git a/meson.build b/meson.build
index 0a95d90d21..a821b85f30 100644
--- a/meson.build
+++ b/meson.build
@@ -1631,6 +1631,7 @@ if sha1_backend == 'sha1dc'
     'sha1dc_git.c',
     'sha1dc/sha1.c',
     'sha1dc/ubc_check.c',
+    'sha1dc-accel/sha1.c',
   ]
 endif
 if sha1_backend == 'CommonCrypto' or sha1_unsafe_backend == 'CommonCrypto'
diff --git a/sha1dc-accel/internal.h b/sha1dc-accel/internal.h
new file mode 100644
index 0000000000..bf3513a1e3
--- /dev/null
+++ b/sha1dc-accel/internal.h
@@ -0,0 +1,25 @@
+#ifndef SHA1DC_ACCEL_INTERNAL_H
+#define SHA1DC_ACCEL_INTERNAL_H
+
+/*
+ * Shared between the files of sha1dc-accel/. See sha1.c for an overview.
+ *
+ * The schedule `w` is always the 80 expanded message words in step order,
+ * w[t] at index t, which is also what sha1dc/ keeps in SHA1_CTX.m1.
+ *
+ * A "state" is the five working words [a, b, c, d, e] before a step.
+ */
+
+#if defined(__GNUC__)
+# define SHA1DC_NOINLINE __attribute__((noinline))
+#else
+# define SHA1DC_NOINLINE
+#endif
+
+/* The step a DV's recompression starts from. */
+enum sha1dc_from {
+	SHA1DC_FROM_58 = 58,
+	SHA1DC_FROM_65 = 65
+};
+
+#endif /* SHA1DC_ACCEL_INTERNAL_H */
diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
new file mode 100644
index 0000000000..fd75289997
--- /dev/null
+++ b/sha1dc-accel/sha1.c
@@ -0,0 +1,434 @@
+/*
+ * SHA-1 with collision detection.
+ *
+ * This computes exactly what sha1dc/ computes: the SHA-1 digest of the
+ * input, and whether any block of it looks like one half of a collision
+ * made by one of the 32 known disturbance vectors (DVs) of Stevens and
+ * Shumow. It works on sha1dc's SHA1_CTX, and uses its table of DVs, but
+ * has its own block loop, which gives the following patches room to
+ * follow the approach of the "sha1dc" Rust crate by Sam Reis
+ * (https://github.com/srijs/sha1dc), which gitoxide uses.
+ *
+ * Each block is compressed by a "backend", which also spills the expanded
+ * message schedule, and the two intermediate states that recompression
+ * starts from (at steps 58 and 65). The unavoidable-bitconditions (UBC)
+ * filter then rules out about 95% of blocks; the rest are recompressed,
+ * once for each DV the filter could not rule out.
+ */
+
+#include "../git-compat-util.h"
+#include "../sha1dc_git.h"
+#if defined(DC_SHA1_SUBMODULE)
+#include "../sha1collisiondetection/lib/ubc_check.h"
+#else
+#include "../sha1dc/ubc_check.h"
+#endif
+#include "sha1.h"
+#include "internal.h"
+
+#define ROL(x, n) (((x) << (n)) | ((x) >> (32 - (n))))
+
+#define F_CH(b, c, d) ((d) ^ ((b) & ((c) ^ (d))))
+#define F_PARITY(b, c, d) ((b) ^ (c) ^ (d))
+#define F_MAJ(b, c, d) (((b) & (c)) | ((d) & ((b) | (c))))
+
+#define K0 0x5A827999
+#define K1 0x6ED9EBA1
+#define K2 0x8F1BBCDC
+#define K3 0xCA62C1D6
+
+/*
+ * One step, on names that rotate: after it, (e, a, b, c, d) are the new
+ * (a, b, c, d, e).
+ */
+#define STEP(f, k, a, b, c, d, e, x) \
+	do { \
+		e += ROL(a, 5) + f(b, c, d) + (k) + (x); \
+		b = ROL(b, 30); \
+	} while (0)
+
+#define LOAD(t) (w[t] = get_be32(block + 4 * (t)))
+#define EXPAND(t) (w[t] = ROL(w[(t) - 3] ^ w[(t) - 8] ^ w[(t) - 14] ^ w[(t) - 16], 1))
+
+#define FIVE_LOAD(f, k, a, b, c, d, e, t) \
+	do { \
+		STEP(f, k, a, b, c, d, e, LOAD(t)); \
+		STEP(f, k, e, a, b, c, d, LOAD((t) + 1)); \
+		STEP(f, k, d, e, a, b, c, LOAD((t) + 2)); \
+		STEP(f, k, c, d, e, a, b, LOAD((t) + 3)); \
+		STEP(f, k, b, c, d, e, a, LOAD((t) + 4)); \
+	} while (0)
+
+#define FIVE_EXPAND(f, k, a, b, c, d, e, t) \
+	do { \
+		STEP(f, k, a, b, c, d, e, EXPAND(t)); \
+		STEP(f, k, e, a, b, c, d, EXPAND((t) + 1)); \
+		STEP(f, k, d, e, a, b, c, EXPAND((t) + 2)); \
+		STEP(f, k, c, d, e, a, b, EXPAND((t) + 3)); \
+		STEP(f, k, b, c, d, e, a, EXPAND((t) + 4)); \
+	} while (0)
+
+/*
+ * The portable compression. Like the hardware ones it spills the schedule,
+ * but it writes the states at steps 58 and 65 directly, on the way past.
+ */
+static void compress_portable(uint32_t ihv[5], const unsigned char *block,
+			      uint32_t w[80], uint32_t state_58[5],
+			      uint32_t state_65[5])
+{
+	uint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];
+
+	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 0);
+	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 5);
+	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 10);
+	STEP(F_CH, K0, a, b, c, d, e, LOAD(15));
+	STEP(F_CH, K0, e, a, b, c, d, EXPAND(16));
+	STEP(F_CH, K0, d, e, a, b, c, EXPAND(17));
+	STEP(F_CH, K0, c, d, e, a, b, EXPAND(18));
+	STEP(F_CH, K0, b, c, d, e, a, EXPAND(19));
+
+	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 20);
+	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 25);
+	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 30);
+	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 35);
+
+	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 40);
+	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 45);
+	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 50);
+	STEP(F_MAJ, K2, a, b, c, d, e, EXPAND(55));
+	STEP(F_MAJ, K2, e, a, b, c, d, EXPAND(56));
+	STEP(F_MAJ, K2, d, e, a, b, c, EXPAND(57));
+	state_58[0] = c;
+	state_58[1] = d;
+	state_58[2] = e;
+	state_58[3] = a;
+	state_58[4] = b;
+	STEP(F_MAJ, K2, c, d, e, a, b, EXPAND(58));
+	STEP(F_MAJ, K2, b, c, d, e, a, EXPAND(59));
+
+	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 60);
+	state_65[0] = a;
+	state_65[1] = b;
+	state_65[2] = c;
+	state_65[3] = d;
+	state_65[4] = e;
+	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 65);
+	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 70);
+	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 75);
+
+	ihv[0] += a;
+	ihv[1] += b;
+	ihv[2] += c;
+	ihv[3] += d;
+	ihv[4] += e;
+}
+
+/*
+ * The rest runs only for blocks the filter flags, about one in twenty, so
+ * it favours being short over being fast.
+ */
+
+/*
+ * Moves `s` from the state before step `from` to the one before step `to`,
+ * either way, on the schedule `m1 ^ dm` (`dm` may be NULL). One loop per
+ * round function keeps the steps free of branches.
+ */
+#define WORD(t) (dm ? m1[t] ^ dm[t] : m1[t])
+#define FORWARD(f, k, end) \
+	for (; t < to && t < (end); t++) { \
+		x = ROL(a, 5) + f(b, c, d) + (k) + e + WORD(t); \
+		e = d; \
+		d = c; \
+		c = ROL(b, 30); \
+		b = a; \
+		a = x; \
+	}
+#define BACKWARD(f, k, start) \
+	for (; t > to && t > (start); t--) { \
+		x = a; \
+		a = b; \
+		b = ROL(c, 2); \
+		c = d; \
+		d = e; \
+		e = x - (ROL(a, 5) + f(b, c, d) + (k) + WORD(t - 1)); \
+	}
+
+static void walk(uint32_t s[5], const uint32_t *m1, const uint32_t *dm,
+		 unsigned from, unsigned to)
+{
+	uint32_t a = s[0], b = s[1], c = s[2], d = s[3], e = s[4], x;
+	unsigned t = from;
+
+	FORWARD(F_CH, K0, 20);
+	FORWARD(F_PARITY, K1, 40);
+	FORWARD(F_MAJ, K2, 60);
+	FORWARD(F_PARITY, K3, 80);
+
+	BACKWARD(F_PARITY, K3, 60);
+	BACKWARD(F_MAJ, K2, 40);
+	BACKWARD(F_PARITY, K1, 20);
+	BACKWARD(F_CH, K0, 0);
+
+	s[0] = a;
+	s[1] = b;
+	s[2] = c;
+	s[3] = d;
+	s[4] = e;
+}
+
+/*
+ * The recompression, the way sha1dc/ does it: from this block's state at
+ * `from`, the partner block's chaining value on the way in (`ihv_in`) and
+ * on the way out (`ihv_out`).
+ */
+static void recompress_portable(enum sha1dc_from from, const uint32_t m1[80],
+				const uint32_t dm[80], const uint32_t state[5],
+				uint32_t ihv_in[5], uint32_t ihv_out[5])
+{
+	uint32_t fwd[5];
+	int i;
+
+	memcpy(ihv_in, state, 5 * sizeof(*ihv_in));
+	walk(ihv_in, m1, dm, from, 0);
+	memcpy(fwd, state, sizeof(fwd));
+	walk(fwd, m1, dm, from, 80);
+	for (i = 0; i < 5; i++)
+		ihv_out[i] = ihv_in[i] + fwd[i];
+}
+
+/* For a mitigated ("safe") hash: one more compression of the schedule. */
+static void compress_schedule(uint32_t ihv[5], const uint32_t w[80])
+{
+	uint32_t s[5];
+	int i;
+
+	memcpy(s, ihv, sizeof(s));
+	walk(s, w, NULL, 0, 80);
+	for (i = 0; i < 5; i++)
+		ihv[i] += s[i];
+}
+
+struct backend {
+	const char *name;
+	/*
+	 * Compresses a block, spilling its schedule and the states at steps
+	 * 58 and 65.
+	 */
+	void (*compress)(uint32_t ihv[5], const unsigned char *block,
+			 uint32_t w[80], uint32_t state_58[5],
+			 uint32_t state_65[5]);
+	uint32_t (*ubc_check)(const uint32_t w[80]);
+	int (*available)(void);
+};
+
+/* sha1dc's own UBC check. */
+static uint32_t ubc_check_sha1dc(const uint32_t w[80])
+{
+	uint32_t mask;
+
+	ubc_check(w, &mask);
+	return mask;
+}
+
+/* In order of preference. */
+static const struct backend backends[] = {
+	{ "portable", compress_portable, ubc_check_sha1dc, NULL },
+};
+
+static int usable(const struct backend *be)
+{
+	return !be->available || be->available();
+}
+
+#if defined(__GNUC__)
+/*
+ * The backend in use, chosen on first use. Threads that race to choose it
+ * all store the same value, with relaxed atomics so that they may.
+ */
+static const struct backend *selected;
+
+static const struct backend *backend(void)
+{
+	const struct backend *be = __atomic_load_n(&selected, __ATOMIC_RELAXED);
+	size_t i;
+
+	if (be)
+		return be;
+	/* The last one, "portable", is always usable. */
+	for (i = 0; !usable(&backends[i]); i++)
+		;
+	be = &backends[i];
+	__atomic_store_n(&selected, be, __ATOMIC_RELAXED);
+	return be;
+}
+
+static void select_backend(const struct backend *be)
+{
+	__atomic_store_n(&selected, be, __ATOMIC_RELAXED);
+}
+#else
+/*
+ * Other compilers only get the portable backend (see internal.h), so
+ * there is nothing to choose, and no shared state to guard.
+ */
+static const struct backend *backend(void)
+{
+	return &backends[0];
+}
+
+static void select_backend(const struct backend *be UNUSED)
+{
+}
+#endif
+
+const char *sha1dc_accel_backend(void)
+{
+	return backend()->name;
+}
+
+int sha1dc_accel_select(const char *name)
+{
+	size_t i;
+
+	for (i = 0; i < ARRAY_SIZE(backends); i++) {
+		if (!strcmp(backends[i].name, name) && usable(&backends[i])) {
+			select_backend(&backends[i]);
+			return 0;
+		}
+	}
+	return -1;
+}
+
+const char *const *sha1dc_accel_backends(void)
+{
+	static const char *names[ARRAY_SIZE(backends) + 1];
+	size_t i, n = 0;
+
+	for (i = 0; i < ARRAY_SIZE(backends); i++)
+		if (usable(&backends[i]))
+			names[n++] = backends[i].name;
+	names[n] = NULL;
+	return names;
+}
+
+/*
+ * Whether any DV in `candidates` makes this block half of a collision.
+ * `ihv_in` and `ihv_out` are the chaining values before and after it.
+ */
+static SHA1DC_NOINLINE int attacked(SHA1_CTX *ctx, uint32_t candidates,
+				    const uint32_t w[80],
+				    const uint32_t state_58[5],
+				    const uint32_t state_65[5],
+				    const uint32_t ihv_in[5],
+				    const uint32_t ihv_out[5])
+{
+	int i;
+
+	for (i = 0; sha1_dvs[i].dvType != 0; i++) {
+		const dv_info_t *dv = &sha1_dvs[i];
+		enum sha1dc_from from;
+		const uint32_t *state;
+		uint32_t ihv2_in[5], ihv2_out[5];
+
+		if (!(candidates & ((uint32_t)1 << dv->maskb)))
+			continue;
+		switch (dv->testt) {
+		case 58:
+			from = SHA1DC_FROM_58;
+			state = state_58;
+			break;
+		case 65:
+			from = SHA1DC_FROM_65;
+			state = state_65;
+			break;
+		default:
+			BUG("sha1dc DV %d(%d,%d) recompresses from step %d, "
+			    "which sha1dc-accel does not save",
+			    dv->dvType, dv->dvK, dv->dvB, dv->testt);
+		}
+
+		recompress_portable(from, w, dv->dm, state, ihv2_in, ihv2_out);
+		if (!memcmp(ihv2_out, ihv_out, sizeof(ihv2_out)) ||
+		    (ctx->reduced_round_coll &&
+		     !memcmp(ihv2_in, ihv_in, sizeof(ihv2_in))))
+			return 1;
+	}
+	return 0;
+}
+
+static inline void process(const struct backend *be, SHA1_CTX *ctx,
+			   const unsigned char *block)
+{
+	uint32_t w[80], state_58[5], state_65[5], ihv_in[5], candidates;
+
+	memcpy(ihv_in, ctx->ihv, sizeof(ihv_in));
+	be->compress(ctx->ihv, block, w, state_58, state_65);
+	if (!ctx->detect_coll)
+		return;
+
+	candidates = ctx->ubc_check ? be->ubc_check(w) : 0xFFFFFFFF;
+	if (candidates &&
+	    attacked(ctx, candidates, w, state_58, state_65, ihv_in,
+		     ctx->ihv)) {
+		ctx->found_collision = 1;
+		/*
+		 * Two more compressions of this block give a digest that the
+		 * other block of the pair does not share.
+		 */
+		if (ctx->safe_hash) {
+			compress_schedule(ctx->ihv, w);
+			compress_schedule(ctx->ihv, w);
+		}
+	}
+}
+
+void sha1dc_accel_update(SHA1_CTX *ctx, const void *data, size_t len)
+{
+	const struct backend *be = backend();
+	const unsigned char *buf = data;
+	unsigned left, fill;
+
+	if (!len)
+		return;
+
+	left = ctx->total & 63;
+	fill = 64 - left;
+
+	if (left && len >= fill) {
+		ctx->total += fill;
+		memcpy(ctx->buffer + left, buf, fill);
+		process(be, ctx, ctx->buffer);
+		buf += fill;
+		len -= fill;
+		left = 0;
+	}
+	while (len >= 64) {
+		ctx->total += 64;
+		process(be, ctx, buf);
+		buf += 64;
+		len -= 64;
+	}
+	if (len) {
+		ctx->total += len;
+		memcpy(ctx->buffer + left, buf, len);
+	}
+}
+
+int sha1dc_accel_final(unsigned char hash[20], SHA1_CTX *ctx)
+{
+	static const unsigned char padding[64] = { 0x80 };
+	uint32_t last = ctx->total & 63;
+	uint32_t padn = last < 56 ? 56 - last : 120 - last;
+	uint64_t bits;
+	int i;
+
+	sha1dc_accel_update(ctx, padding, padn);
+	bits = (ctx->total - padn) << 3;
+	put_be32(ctx->buffer + 56, (uint32_t)(bits >> 32));
+	put_be32(ctx->buffer + 60, (uint32_t)bits);
+	process(backend(), ctx, ctx->buffer);
+
+	for (i = 0; i < 5; i++)
+		put_be32(hash + 4 * i, ctx->ihv[i]);
+	return ctx->found_collision;
+}
diff --git a/sha1dc-accel/sha1.h b/sha1dc-accel/sha1.h
new file mode 100644
index 0000000000..53dda518f0
--- /dev/null
+++ b/sha1dc-accel/sha1.h
@@ -0,0 +1,31 @@
+#ifndef SHA1DC_ACCEL_SHA1_H
+#define SHA1DC_ACCEL_SHA1_H
+
+#if defined(DC_SHA1_SUBMODULE)
+#include "../sha1collisiondetection/lib/sha1.h"
+#else
+#include "../sha1dc/sha1.h"
+#endif
+
+/*
+ * A faster implementation of SHA-1 with collision detection, working on
+ * the SHA1_CTX of sha1dc/ (or of the sha1collisiondetection submodule).
+ * SHA1DCInit() and the SHA1DCSet*() functions set it up as usual; these
+ * then take the place of SHA1DCUpdate() and SHA1DCFinal(), with the same
+ * results. See sha1.c.
+ */
+
+void sha1dc_accel_update(SHA1_CTX *ctx, const void *data, size_t len);
+int sha1dc_accel_final(unsigned char hash[20], SHA1_CTX *ctx);
+
+/*
+ * The implementation in use, such as "shani+avx2" or "portable". For tests,
+ * sha1dc_accel_select() switches to another one by that name; it returns -1
+ * if this build or CPU does not have it. Neither is thread-safe.
+ */
+const char *sha1dc_accel_backend(void);
+int sha1dc_accel_select(const char *name);
+/* The names this build and CPU can use, NULL-terminated. */
+const char *const *sha1dc_accel_backends(void);
+
+#endif /* SHA1DC_ACCEL_SHA1_H */
diff --git a/sha1dc_git.c b/sha1dc_git.c
index fe58d7962a..03958762cf 100644
--- a/sha1dc_git.c
+++ b/sha1dc_git.c
@@ -2,6 +2,15 @@
 #include "sha1dc_git.h"
 #include "hex.h"
 
+/*
+ * Unless told otherwise, hash with sha1dc-accel/, which gives the same
+ * results as sha1dc/ on the same SHA1_CTX.
+ */
+#if !defined(DC_SHA1_EXTERNAL) && !defined(DC_SHA1_NO_ACCEL)
+#include "sha1dc-accel/sha1.h"
+#define USE_SHA1DC_ACCEL
+#endif
+
 #ifdef DC_SHA1_EXTERNAL
 /*
  * Same as SHA1DCInit, but with default save_hash=0
@@ -18,8 +27,13 @@ void git_SHA1DCInit(SHA1_CTX *ctx)
  */
 void git_SHA1DCFinal(unsigned char hash[20], SHA1_CTX *ctx)
 {
+#ifdef USE_SHA1DC_ACCEL
+	if (!sha1dc_accel_final(hash, ctx))
+		return;
+#else
 	if (!SHA1DCFinal(hash, ctx))
 		return;
+#endif
 	die("SHA-1 appears to be part of a collision attack: %s",
 	    hash_to_hex_algop(hash, &hash_algos[GIT_HASH_SHA1]));
 }
@@ -29,6 +43,9 @@ void git_SHA1DCFinal(unsigned char hash[20], SHA1_CTX *ctx)
  */
 void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *vdata, size_t len)
 {
+#ifdef USE_SHA1DC_ACCEL
+	sha1dc_accel_update(ctx, vdata, len);
+#else
 	const char *data = vdata;
 	while (len > INT_MAX) {
 		SHA1DCUpdate(ctx, data, INT_MAX);
@@ -36,4 +53,5 @@ void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *vdata, size_t len)
 		len -= INT_MAX;
 	}
 	SHA1DCUpdate(ctx, data, len);
+#endif
 }
diff --git a/t/.gitattributes b/t/.gitattributes
index e867f38c71..f9437159ae 100644
--- a/t/.gitattributes
+++ b/t/.gitattributes
@@ -2,6 +2,7 @@ t[0-9][0-9][0-9][0-9]/* -whitespace
 /chainlint/*.expect eol=lf -whitespace
 /greplint/*.expect eol=lf -whitespace
 /greplint/*.test eol=lf -whitespace
+/t0013/*.bin binary
 /t0110/url-* binary
 /t3206/* eol=lf
 /t3900/*.txt eol=lf
diff --git a/t/helper/test-sha1.c b/t/helper/test-sha1.c
index 349540c4df..7a0e9cc716 100644
--- a/t/helper/test-sha1.c
+++ b/t/helper/test-sha1.c
@@ -1,8 +1,20 @@
 #include "test-tool.h"
 #include "hash.h"
+#include "strbuf.h"
+
+#if defined(SHA1_DC) && !defined(DC_SHA1_EXTERNAL) && !defined(DC_SHA1_NO_ACCEL)
+#include "sha1dc-accel/sha1.h"
+#define HAVE_SHA1DC_ACCEL
+#endif
 
 int cmd__sha1(int ac, const char **av)
 {
+#ifdef HAVE_SHA1DC_ACCEL
+	const char *backend = getenv("GIT_TEST_SHA1DC_BACKEND");
+
+	if (backend && sha1dc_accel_select(backend))
+		die("sha1dc backend '%s' is not available", backend);
+#endif
 	return cmd_hash_impl(ac, av, GIT_HASH_SHA1, 0);
 }
 
@@ -14,6 +26,89 @@ int cmd__sha1_is_sha1dc(int argc UNUSED, const char **argv UNUSED)
 	return 1;
 }
 
+/*
+ * Lists the implementations of collision-detecting SHA-1 this build and
+ * CPU can use, which GIT_TEST_SHA1DC_BACKEND selects for "test-tool sha1".
+ */
+int cmd__sha1dc_backends(int argc UNUSED, const char **argv UNUSED)
+{
+#ifdef HAVE_SHA1DC_ACCEL
+	const char *const *names;
+
+	for (names = sha1dc_accel_backends(); *names; names++)
+		puts(*names);
+#endif
+	return 0;
+}
+
+/*
+ * Hashes a file with sha1dc/ and with sha1dc-accel/ (the backend that
+ * GIT_TEST_SHA1DC_BACKEND selects), under every combination of settings,
+ * and dies if they ever disagree on the digest or on whether there is a
+ * collision. sha1dc-accel/ sees the file split at and around every block
+ * boundary, with an empty update in between. Prints what sha1dc/ found for
+ * each combination.
+ */
+int cmd__sha1dc_compare(int argc MAYBE_UNUSED, const char **argv MAYBE_UNUSED)
+{
+#ifdef HAVE_SHA1DC_ACCEL
+	const char *backend = getenv("GIT_TEST_SHA1DC_BACKEND");
+	struct strbuf buf = STRBUF_INIT;
+	int mode;
+
+	if (argc != 2)
+		die("usage: test-tool sha1dc-compare <file>");
+	if (backend && sha1dc_accel_select(backend))
+		die("sha1dc backend '%s' is not available", backend);
+	if (strbuf_read_file(&buf, argv[1], 0) < 0)
+		die_errno("could not read '%s'", argv[1]);
+
+	for (mode = 0; mode < 16; mode++) {
+		int detect = !!(mode & 8), safe = !!(mode & 4);
+		int ubc = !!(mode & 2), reduced = !!(mode & 1);
+		unsigned char want[20], got[20];
+		SHA1_CTX init, ctx;
+		int want_coll, got_coll, i;
+		size_t split;
+
+		SHA1DCInit(&init);
+		SHA1DCSetUseDetectColl(&init, detect);
+		SHA1DCSetSafeHash(&init, safe);
+		SHA1DCSetUseUBC(&init, ubc);
+		SHA1DCSetDetectReducedRoundCollision(&init, reduced);
+
+		ctx = init;
+		SHA1DCUpdate(&ctx, buf.buf, buf.len);
+		want_coll = SHA1DCFinal(want, &ctx);
+
+		for (split = 0; split <= buf.len; split++) {
+			if (split % 64 > 1 && split % 64 < 63 && split != buf.len)
+				continue;
+			ctx = init;
+			sha1dc_accel_update(&ctx, buf.buf, split);
+			sha1dc_accel_update(&ctx, buf.buf + split, 0);
+			sha1dc_accel_update(&ctx, buf.buf + split, buf.len - split);
+			got_coll = sha1dc_accel_final(got, &ctx);
+			if (got_coll != want_coll || memcmp(got, want, sizeof(got)))
+				die("%s disagrees with sha1dc/ on '%s' with detect=%d "
+				    "safe=%d ubc=%d reduced=%d, split at %"PRIuMAX,
+				    sha1dc_accel_backend(), argv[1], detect, safe,
+				    ubc, reduced, (uintmax_t)split);
+		}
+
+		printf("detect=%d safe=%d ubc=%d reduced=%d collision=%d ",
+		       detect, safe, ubc, reduced, want_coll);
+		for (i = 0; i < 20; i++)
+			printf("%02x", want[i]);
+		putchar('\n');
+	}
+	strbuf_release(&buf);
+	return 0;
+#else
+	die("test-tool sha1dc-compare: not built with sha1dc-accel");
+#endif
+}
+
 int cmd__sha1_unsafe(int ac, const char **av)
 {
 	return cmd_hash_impl(ac, av, GIT_HASH_SHA1, 1);
diff --git a/t/helper/test-tool.c b/t/helper/test-tool.c
index b71a22b43b..c459218f0b 100644
--- a/t/helper/test-tool.c
+++ b/t/helper/test-tool.c
@@ -73,6 +73,8 @@ static struct test_cmd cmds[] = {
 	{ "serve-v2", cmd__serve_v2 },
 	{ "sha1", cmd__sha1 },
 	{ "sha1-is-sha1dc", cmd__sha1_is_sha1dc },
+	{ "sha1dc-backends", cmd__sha1dc_backends },
+	{ "sha1dc-compare", cmd__sha1dc_compare },
 	{ "sha1-unsafe", cmd__sha1_unsafe },
 	{ "sha256", cmd__sha256 },
 	{ "sigchain", cmd__sigchain },
diff --git a/t/helper/test-tool.h b/t/helper/test-tool.h
index f2885b33d5..f0cdbddcdf 100644
--- a/t/helper/test-tool.h
+++ b/t/helper/test-tool.h
@@ -66,6 +66,8 @@ int cmd__scrap_cache_tree(int argc, const char **argv);
 int cmd__serve_v2(int argc, const char **argv);
 int cmd__sha1(int argc, const char **argv);
 int cmd__sha1_is_sha1dc(int argc, const char **argv);
+int cmd__sha1dc_backends(int argc, const char **argv);
+int cmd__sha1dc_compare(int argc, const char **argv);
 int cmd__sha1_unsafe(int argc, const char **argv);
 int cmd__sha256(int argc, const char **argv);
 int cmd__sigchain(int argc, const char **argv);
diff --git a/t/meson.build b/t/meson.build
index 3ca7b27104..ec25f0de40 100644
--- a/t/meson.build
+++ b/t/meson.build
@@ -20,6 +20,7 @@ clar_test_suites = [
   'unit-tests/u-reftable-stack.c',
   'unit-tests/u-reftable-table.c',
   'unit-tests/u-reftable-tree.c',
+  'unit-tests/u-sha1dc.c',
   'unit-tests/u-strbuf.c',
   'unit-tests/u-strcmp-offset.c',
   'unit-tests/u-string-list.c',
diff --git a/t/t0013-sha1dc.sh b/t/t0013-sha1dc.sh
index 3ea3169d92..c3c4b9d951 100755
--- a/t/t0013-sha1dc.sh
+++ b/t/t0013-sha1dc.sh
@@ -19,4 +19,51 @@ test_expect_success 'test-sha1 detects shattered pdf' '
 	test_grep 38762cf7f55934b34d179ae6a4c80cadccbb7f0a err
 '
 
+test_expect_success 'test-sha1 detects SHA-mbles chosen-prefix collision' '
+	test_must_fail test-tool sha1 <"$TEST_DATA/sha-mbles-1.bin" 2>err &&
+	test_grep collision err &&
+	test_grep 8ac60ba76f1999a1ab70223f225aefdc78d4ddc0 err
+'
+
+# Each implementation of the detection this build and CPU can use.
+for backend in $(test-tool sha1dc-backends)
+do
+	test_expect_success "$backend: detects collisions" '
+		test_must_fail env GIT_TEST_SHA1DC_BACKEND=$backend \
+			test-tool sha1 <"$TEST_DATA/shattered-1.pdf" 2>err &&
+		test_grep 38762cf7f55934b34d179ae6a4c80cadccbb7f0a err &&
+		test_must_fail env GIT_TEST_SHA1DC_BACKEND=$backend \
+			test-tool sha1 <"$TEST_DATA/sha-mbles-1.bin" 2>err &&
+		test_grep 8ac60ba76f1999a1ab70223f225aefdc78d4ddc0 err
+	'
+
+	# In every combination of settings, and with the input split around
+	# each block, it must find what sha1dc/ finds and give its digest.
+	test_expect_success "$backend: agrees with sha1dc/ on collisions" '
+		test_copy_bytes 320 <"$TEST_DATA/shattered-1.pdf" >shattered &&
+		for f in shattered "$TEST_DATA/sha-mbles-1.bin"
+		do
+			GIT_TEST_SHA1DC_BACKEND=$backend \
+				test-tool sha1dc-compare "$f" >out &&
+			test_grep ! "detect=1 .* collision=0" out &&
+			test_grep ! "detect=0 .* collision=1" out || return 1
+		done &&
+
+		# A reduced-round collision counts only when asked for.
+		GIT_TEST_SHA1DC_BACKEND=$backend test-tool sha1dc-compare \
+			"$TEST_DATA/sha1-reduced-round.bin" >out &&
+		grep "collision=1" out >actual &&
+		grep "detect=1 .* reduced=1 collision=1" out >expect &&
+		test_line_count = 4 expect &&
+		test_cmp expect actual
+	'
+
+	test_expect_success "$backend: hashes like the others" '
+		test-tool genrandom "$backend" 100000 >data &&
+		GIT_TEST_SHA1DC_BACKEND=$backend test-tool sha1 <data >actual &&
+		test-tool sha1 <data >expect &&
+		test_cmp expect actual
+	'
+done
+
 test_done
diff --git a/t/t0013/sha-mbles-1.bin b/t/t0013/sha-mbles-1.bin
new file mode 100644
index 0000000000000000000000000000000000000000..5a7c30e97646c66422abe0a9793a5fcb9f1cf8d6
GIT binary patch
literal 640
zcmV-`0)PFP1Pug#=of$iAOQbMWqBZJb0BbGa&#bXW*}i8V{dG1X>)0BZXqB^bSHBl
zVIXvJVQ?XN#v1Ui%mqai*(XkO2VzSd$NM9gi@4s4S6#Y$o~tpzXG?9DLwKksb1(IU
z9Co7S2XeKfeBtWE3%QfQEsSvDN>7bn&F#UnES&M4F|Q;kb)7=w-`g>9pICMy?o}x{
zw%puBpUP8JJ8<}Z-Y}v^>N@tvS)%d_G7WYOwomkV2v5_@v(40lV%ch(Lk1Vh|7<qK
zH|0OxC_#T>Z|qd<c|)XbUso{lyEywD_TUKs5YP@Jt$4qZWEqoSj*S(Hc%L-HZ{g+w
ze>J4b`+{(G#SY32i+svyyDTet0$KULm2lmSL^q=mU$6JW%D|e^Qf38QClE(f7mlv~
zf?6!9D$ljvWX^U$+*zeTsr;OEXIA3kJ;xKs!c3Qts%s87r}bYHMJgPkg$>=6V*Q#J
ztwKp^sc;DQMsoI!^kM6Wu$eQ~Cban&bezB^{oQOrU&JA3HP91H6)0P)EVqQD_sg{V
zQB6zm_9J}o3Z9=6E1CvwZ_$5jLYQ=TSa0@Gua<Owv?jTSE1HPp20vN5Gfcn+Q2084
z#3xa=8FbSC{3scs=<(w$8&S&`=D)<-o38eC)T;HdS4sqbk8RTI6*`kaB9oU*l8=ba
z*)}}>`F!H%LccW0YmW1WR(AfS%&6t}-k`d&K|M|24(A@=9~LXyZ62@LCFZWWu4*+-
z@qF?Hqy+oh68uF?LH*fW@<o<pqOAih9i|F%CO~!9@!;0MKsx83*kRv4<#2I`-ChUL
aSeu`VW-wJhkHb>4;KF^V3*EX*WC9JhR6N@N

literal 0
HcmV?d00001

diff --git a/t/t0013/sha1-reduced-round.bin b/t/t0013/sha1-reduced-round.bin
new file mode 100644
index 0000000000000000000000000000000000000000..4623336222bd5c9e7b1b0e244a5897430c1b5c12
GIT binary patch
literal 128
zcmV-`0Du3yemOb>aQ1}Yq=eq3R)<>6-}%Tb0s(7=4(It1;e;4*zrXPYaFxmJM6d36
z5+n(uvg<Auz|X=4#ULmUI6NzJ=Hkdhf3ZGJO<lI*gWw%|>Le^HwlGv^MX^H+A(X58
iQZ~LT$sQRU5x<XSUiqt^kK<}UEWbI|d>^ztun2PMcs$<#

literal 0
HcmV?d00001

diff --git a/t/unit-tests/u-sha1dc.c b/t/unit-tests/u-sha1dc.c
new file mode 100644
index 0000000000..31fcd68451
--- /dev/null
+++ b/t/unit-tests/u-sha1dc.c
@@ -0,0 +1,145 @@
+#include "unit-test.h"
+#include "hash.h"
+
+/*
+ * Tests sha1dc-accel/ against sha1dc/, which it must agree with exactly.
+ */
+#if defined(SHA1_DC) && !defined(DC_SHA1_EXTERNAL) && !defined(DC_SHA1_NO_ACCEL)
+#define HAVE_SHA1DC_ACCEL
+
+#if defined(DC_SHA1_SUBMODULE)
+#include "sha1collisiondetection/lib/ubc_check.h"
+#else
+#include "sha1dc/ubc_check.h"
+#endif
+#include "sha1dc-accel/sha1.h"
+#include "sha1dc-accel/internal.h"
+
+static uint64_t rng_state;
+
+static void rng_seed(uint64_t seed)
+{
+	rng_state = seed * 2 + 1;
+}
+
+static uint32_t rng(void)
+{
+	/* xorshift64* */
+	rng_state ^= rng_state >> 12;
+	rng_state ^= rng_state << 25;
+	rng_state ^= rng_state >> 27;
+	return (uint32_t)((rng_state * 0x2545F4914F6CDD1DULL) >> 32);
+}
+
+/*
+ * Hashes `buf` with sha1dc/ and with sha1dc-accel/ under the same settings
+ * and checks that both agree. Returns whether they found a collision.
+ */
+static int check_same(const unsigned char *buf, size_t len, int split,
+		      int safe_hash, int ubc, int reduced_round)
+{
+	SHA1_CTX want_ctx, got_ctx;
+	unsigned char want[20], got[20];
+	int want_coll, got_coll;
+
+	SHA1DCInit(&want_ctx);
+	SHA1DCSetSafeHash(&want_ctx, safe_hash);
+	SHA1DCSetUseUBC(&want_ctx, ubc);
+	SHA1DCSetDetectReducedRoundCollision(&want_ctx, reduced_round);
+	got_ctx = want_ctx;
+
+	SHA1DCUpdate(&want_ctx, (const char *)buf, len);
+	want_coll = SHA1DCFinal(want, &want_ctx);
+
+	sha1dc_accel_update(&got_ctx, buf, split);
+	sha1dc_accel_update(&got_ctx, buf + split, len - split);
+	got_coll = sha1dc_accel_final(got, &got_ctx);
+
+	cl_assert_equal_i(got_coll, want_coll);
+	cl_assert(!memcmp(got, want, sizeof(want)));
+	return got_coll;
+}
+
+/*
+ * sha1dc-accel/ relies on sha1dc/'s table of DVs having the 32 DVs below,
+ * with DV n at bit n of the UBC mask, and each recompressing from a state
+ * it saves (step 58 or 65). A change there must fail here, rather than
+ * quietly check the wrong DV or start from the wrong state.
+ */
+static void dv_table_is_as_expected(void)
+{
+	static const struct { int type, k, b; } want[] = {
+		{ 1, 43, 0 }, { 1, 44, 0 }, { 1, 45, 0 }, { 1, 46, 0 },
+		{ 1, 46, 2 }, { 1, 47, 0 }, { 1, 47, 2 }, { 1, 48, 0 },
+		{ 1, 48, 2 }, { 1, 49, 0 }, { 1, 49, 2 }, { 1, 50, 0 },
+		{ 1, 50, 2 }, { 1, 51, 0 }, { 1, 51, 2 }, { 1, 52, 0 },
+		{ 2, 45, 0 }, { 2, 46, 0 }, { 2, 46, 2 }, { 2, 47, 0 },
+		{ 2, 48, 0 }, { 2, 49, 0 }, { 2, 49, 2 }, { 2, 50, 0 },
+		{ 2, 50, 2 }, { 2, 51, 0 }, { 2, 51, 2 }, { 2, 52, 0 },
+		{ 2, 53, 0 }, { 2, 54, 0 }, { 2, 55, 0 }, { 2, 56, 0 },
+	};
+	int i;
+
+	for (i = 0; sha1_dvs[i].dvType != 0; i++) {
+		const dv_info_t *dv = &sha1_dvs[i];
+
+		cl_assert(i < (int)ARRAY_SIZE(want));
+		cl_assert_equal_i(dv->dvType, want[i].type);
+		cl_assert_equal_i(dv->dvK, want[i].k);
+		cl_assert_equal_i(dv->dvB, want[i].b);
+		cl_assert_equal_i(dv->maski, 0);
+		cl_assert_equal_i(dv->maskb, i);
+		cl_assert(dv->testt == SHA1DC_FROM_58 || dv->testt == SHA1DC_FROM_65);
+	}
+	cl_assert_equal_i(i, ARRAY_SIZE(want));
+}
+
+static void every_backend_agrees_with_sha1dc(void)
+{
+	const char *const *names = sha1dc_accel_backends();
+	unsigned char buf[4200];
+	const char *orig = sha1dc_accel_backend();
+
+	for (; *names; names++) {
+		size_t i;
+
+		cl_assert_equal_i(sha1dc_accel_select(*names), 0);
+		rng_seed(1);
+		for (i = 0; i < sizeof(buf); i++)
+			buf[i] = rng();
+
+		for (i = 0; i < 600; i++) {
+			size_t len = i < 200 ? i : rng() % (sizeof(buf) - 16);
+			size_t off = rng() % 16;
+			size_t split = len ? rng() % (len + 1) : 0;
+			/*
+			 * Without the filter, every block is recompressed
+			 * for all 32 DVs, which must not find anything in
+			 * random data either.
+			 */
+			int ubc = i % 5 != 4;
+
+			cl_assert_equal_i(check_same(buf + off, len, split,
+						     i & 1, ubc, i % 3 == 0), 0);
+		}
+	}
+	cl_assert_equal_i(sha1dc_accel_select(orig), 0);
+}
+
+#endif /* HAVE_SHA1DC_ACCEL */
+
+#ifdef HAVE_SHA1DC_ACCEL
+#define RUN_OR_SKIP(fn) fn()
+#else
+#define RUN_OR_SKIP(fn) cl_skip()
+#endif
+
+void test_sha1dc__dv_table_is_as_expected(void)
+{
+	RUN_OR_SKIP(dv_table_is_as_expected);
+}
+
+void test_sha1dc__every_backend_agrees_with_sha1dc(void)
+{
+	RUN_OR_SKIP(every_backend_agrees_with_sha1dc);
+}
-- 
2.50.1 (Apple Git-155)



^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
  2026-09-29 11:25 ` [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX Scott Chacon
@ 2026-09-29 11:25 ` Scott Chacon
  2026-10-07 12:17   ` Johannes Schindelin
  2026-09-29 11:25 ` [PATCH 3/4] sha1dc-accel: compress with SHA-NI on x86-64 Scott Chacon
                   ` (3 subsequent siblings)
  5 siblings, 1 reply; 21+ messages in thread
From: Scott Chacon @ 2026-09-29 11:25 UTC (permalink / raw)
  To: git

After compressing a block, sha1dc checks its expanded message schedule
against the "unavoidable bitconditions" (UBCs) of each of its 32
disturbance vectors (DVs). Each condition says that one bit of W[i]
XORed with one bit of W[j] must have a particular value if an attack
along that DV is in progress. A failed condition rules out its DVs, and
if all 32 are ruled out, which happens for about 95% of blocks, we can
skip the expensive recompression.

That makes the check a cheap filter, but it's cheap only compared to
recompressing. It's 111 scalar statements, of which 47 run for every
block, and the rest sit behind tests of whether their DVs are still
alive. On this machine it takes about half as long as the compression
itself.

The conditions for any one DV aren't a fixed list, though. They're
equations over the bits of the schedule, and equations chain: if bit A
must equal bit B, and B must equal C, then A must equal C. So each DV
really has a whole space of equivalent condition sets. The generator
that wrote sha1dc's ubc_check.c picked conditions that are shared by
many DVs, which is the right thing for scalar code. But a vector unit
can check 4 (or 8) conditions in a single pair of loads if they have the
same shape: the same distance between the two words, the same bit
positions, at consecutive values of i. That wants a different choice.

The sha1dc Rust crate has a solver that makes that choice for each
instruction set. It picks a "prefix" of vector groups that runs on every
block, sized so that it rules out as many blocks as possible for what
it costs, and leaves the remaining conditions to a scalar "tail" that
only runs for DVs the prefix left alive. Its plans work out roughly like
this (the last column is the solver's estimate for random blocks):

  form    prefix               tail   blocks reaching tail
  scalar  70 statements        86     21%
  sse2    26 groups of 4       58      9%
  avx2    16 groups of 8       44      7%
  neon    20 groups of 4       77     17%

The crate only emits Rust, so ubc_check.c carries its plans over as
tables: for each form, the conditions of its prefix (as one entry per
vector group, giving the two words, their shifts, and the bit and DVs
of each lane), and the remaining conditions of each DV for its tail.
The code that runs them is a short loop per form. We ask the compiler to
unroll the prefix loops fully, so that each table entry turns back into
immediates; without that, the forms are 2.2x to 2.4x slower with GCC
and 1.3x to 3.6x slower with clang. With it, GCC's build of the tables
runs the same number of instructions as the crate's plans written out as
straight-line code, at the same speed. The exception so far is clang 18
on x86-64, where the scalar form runs at half that speed (68ns a block
instead of 34ns), since it spills partial masks to the stack. That form
is only the fallback for CPUs without SSE2 or NEON, and Apple clang 21
does not do this.

The tables come from the crate as of commit 426b4afd. The crate is MIT
or Apache-2.0 at your option (sha1collisiondetection itself is MIT); we
use the tables under the MIT license, and the file carries its notice.
Note that the crate stores its schedule backwards on x86 (to save a
shuffle in its SHA-NI code), where we keep it in step order everywhere,
so our SSE2 and AVX2 tables list their lanes in the opposite order from
the crate's.

We now have these backends, in order of preference:

  - portable+avx2: if the CPU has AVX2 (and the OS saves the YMM
    registers, which we check with xgetbv)

  - portable+sse2: always on x86-64, where SSE2 is part of the ABI

  - portable+neon: always on arm64

  - portable: everywhere else, which still gets the new scalar form

The vector forms need target attributes and intrinsics, so for now
they're only built with GCC 5 or newer and clang. Other compilers (and
32-bit x86, where SSE2 isn't a given) get only the portable one.

With 256MB of random data again, the Xeon picks portable+avx2:

  Benchmark 1: test-tool.old sha1
    Time (mean ± σ):     801.1 ms ±  82.9 ms    [User: 750.1 ms, System: 40.3 ms]
    Range (min … max):   673.1 ms … 1006.1 ms    30 runs

  Benchmark 2: test-tool.new sha1
    Time (mean ± σ):     620.4 ms ± 111.2 ms    [User: 571.0 ms, System: 41.9 ms]
    Range (min … max):   486.4 ms … 857.2 ms    30 runs

  Summary
    test-tool.new sha1 ran
      1.29 ± 0.27 times faster than test-tool.old sha1

and callgrind counts 16% fewer instructions than before this patch. The
M5 Max picks portable+neon, and hashes 1GiB in 1.22s instead of 1.55s
(median of 9 runs), 1.27x faster.

For testing, the unit test compares every form against sha1dc's own
ubc_check(). Random schedules aren't enough, though: the prefix rules
out most of them, so the tail checks for any one DV would hardly ever
run. So for each DV, the test also generates random schedules until one
keeps that DV alive, and then compares all forms on it with each of its
2560 bits flipped in turn. That's the same approach the crate takes (and
sha1collisiondetection's own tools before it).

Outside of the test suite, I also checked every form against ubc_check()
on a million random schedules and 256 of those witnesses: built with
GCC 13 and clang 18 on x86-64, with Apple clang on arm64 (and with it
targeting x86-64, running the SSE2 form under Rosetta and the AVX2 one
translated to NEON by SIMDe), and with GCC and clang for aarch64 under
qemu.

Signed-off-by: Scott Chacon <scott@gitbutler.net>
Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
---
 Makefile                            |    5 +-
 contrib/buildsystems/CMakeLists.txt |    2 +-
 meson.build                         |    2 +
 sha1dc-accel/internal.h             |   49 +
 sha1dc-accel/sha1.c                 |   47 +-
 sha1dc-accel/ubc_check.c            | 1789 +++++++++++++++++++++++++++
 sha1dc-accel/x86.c                  |   38 +
 t/unit-tests/u-sha1dc.c             |   75 ++
 8 files changed, 1985 insertions(+), 22 deletions(-)
 create mode 100644 sha1dc-accel/ubc_check.c
 create mode 100644 sha1dc-accel/x86.c

diff --git a/Makefile b/Makefile
index 9ab13ca2ab..3ad8a7fc92 100644
--- a/Makefile
+++ b/Makefile
@@ -568,7 +568,8 @@ include shared.mak
 # submodule.
 #
 # Unless DC_SHA1_EXTERNAL is defined, the built-in code is driven by the
-# block loop in sha1dc-accel/, which gives the same results.
+# faster implementation in sha1dc-accel/, which gives the same results
+# using the CPU's vector units where it has them.
 # Define DC_SHA1_NO_ACCEL to use the sha1collisiondetection code alone.
 #
 # === SHA-256 backend ===
@@ -2177,6 +2178,8 @@ ifdef DC_SHA1_NO_ACCEL
 	BASIC_CFLAGS += -DDC_SHA1_NO_ACCEL
 else
 	LIB_OBJS += sha1dc-accel/sha1.o
+	LIB_OBJS += sha1dc-accel/ubc_check.o
+	LIB_OBJS += sha1dc-accel/x86.o
 endif
 	BASIC_CFLAGS += \
 		-DSHA1DC_NO_STANDARD_INCLUDES \
diff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt
index 3c0ea2a27c..67b96d601b 100644
--- a/contrib/buildsystems/CMakeLists.txt
+++ b/contrib/buildsystems/CMakeLists.txt
@@ -218,7 +218,7 @@ add_compile_definitions(NO_OPENSSL SHA1_DC SHA1DC_NO_STANDARD_INCLUDES
 			SHA1DC_INIT_SAFE_HASH_DEFAULT=0
 			SHA1DC_CUSTOM_INCLUDE_SHA1_C="git-compat-util.h"
 			SHA1DC_CUSTOM_INCLUDE_UBC_CHECK_C="git-compat-util.h" )
-list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c sha1dc-accel/sha1.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
+list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c sha1dc-accel/sha1.c sha1dc-accel/ubc_check.c sha1dc-accel/x86.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
 
 
 add_compile_definitions(PAGER_ENV="LESS=FRX LV=-c"
diff --git a/meson.build b/meson.build
index a821b85f30..47a60526e9 100644
--- a/meson.build
+++ b/meson.build
@@ -1632,6 +1632,8 @@ if sha1_backend == 'sha1dc'
     'sha1dc/sha1.c',
     'sha1dc/ubc_check.c',
     'sha1dc-accel/sha1.c',
+    'sha1dc-accel/ubc_check.c',
+    'sha1dc-accel/x86.c',
   ]
 endif
 if sha1_backend == 'CommonCrypto' or sha1_unsafe_backend == 'CommonCrypto'
diff --git a/sha1dc-accel/internal.h b/sha1dc-accel/internal.h
index bf3513a1e3..03427224da 100644
--- a/sha1dc-accel/internal.h
+++ b/sha1dc-accel/internal.h
@@ -10,10 +10,38 @@
  * A "state" is the five working words [a, b, c, d, e] before a step.
  */
 
+/*
+ * Which forms this build can have. The vector and hardware forms need
+ * GCC-compatible target attributes and intrinsics; anything else gets the
+ * portable form, which every build has.
+ */
+#if defined(__GNUC__) && !defined(SHA1DC_ACCEL_PORTABLE_ONLY)
+# if defined(__x86_64__) && (defined(__clang__) || __GNUC__ >= 5)
+#  define SHA1DC_HAVE_SSE2 1
+#  define SHA1DC_HAVE_AVX2 1
+#  include <immintrin.h>
+#  define SHA1DC_TARGET_SSE2
+#  define SHA1DC_TARGET_AVX2 __attribute__((target("avx2")))
+# elif defined(__aarch64__) && defined(__ARM_NEON)
+#  define SHA1DC_HAVE_NEON 1
+#  include <arm_neon.h>
+# endif
+#endif
+
 #if defined(__GNUC__)
 # define SHA1DC_NOINLINE __attribute__((noinline))
+# define sha1dc_ctz(x) ((unsigned)__builtin_ctz(x))
 #else
 # define SHA1DC_NOINLINE
+static inline unsigned sha1dc_ctz(uint32_t x)
+{
+	unsigned n = 0;
+	while (!(x & 1)) {
+		x >>= 1;
+		n++;
+	}
+	return n;
+}
 #endif
 
 /* The step a DV's recompression starts from. */
@@ -22,4 +50,25 @@ enum sha1dc_from {
 	SHA1DC_FROM_65 = 65
 };
 
+/*
+ * The UBC check, one form per instruction set. Each returns the same mask
+ * as ubc_check() in sha1dc/: a set bit names a DV that is still possible.
+ * See ubc_check.c.
+ */
+uint32_t sha1dc_ubc_check_scalar(const uint32_t w[80]);
+#ifdef SHA1DC_HAVE_SSE2
+uint32_t sha1dc_ubc_check_sse2(const uint32_t w[80]);
+#endif
+#ifdef SHA1DC_HAVE_AVX2
+uint32_t sha1dc_ubc_check_avx2(const uint32_t w[80]);
+#endif
+#ifdef SHA1DC_HAVE_NEON
+uint32_t sha1dc_ubc_check_neon(const uint32_t w[80]);
+#endif
+
+/* Whether the CPU (and OS) can run the AVX2 form. In x86.c. */
+#ifdef SHA1DC_HAVE_AVX2
+int sha1dc_avx2_available(void);
+#endif
+
 #endif /* SHA1DC_ACCEL_INTERNAL_H */
diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
index fd75289997..1b3d82b4e4 100644
--- a/sha1dc-accel/sha1.c
+++ b/sha1dc-accel/sha1.c
@@ -1,19 +1,23 @@
 /*
- * SHA-1 with collision detection.
+ * SHA-1 with collision detection, faster.
  *
  * This computes exactly what sha1dc/ computes: the SHA-1 digest of the
  * input, and whether any block of it looks like one half of a collision
  * made by one of the 32 known disturbance vectors (DVs) of Stevens and
- * Shumow. It works on sha1dc's SHA1_CTX, and uses its table of DVs, but
- * has its own block loop, which gives the following patches room to
- * follow the approach of the "sha1dc" Rust crate by Sam Reis
- * (https://github.com/srijs/sha1dc), which gitoxide uses.
+ * Shumow. It is a port to C of the approach of the "sha1dc" Rust crate by
+ * Sam Reis (https://github.com/srijs/sha1dc), which gitoxide uses:
  *
- * Each block is compressed by a "backend", which also spills the expanded
- * message schedule, and the two intermediate states that recompression
- * starts from (at steps 58 and 65). The unavoidable-bitconditions (UBC)
- * filter then rules out about 95% of blocks; the rest are recompressed,
- * once for each DV the filter could not rule out.
+ *  - The unavoidable-bitconditions (UBC) filter, which rules out about 95%
+ *    of blocks and is most of what detection costs, has one form per
+ *    instruction set (SSE2, AVX2, NEON and portable C). For each, the
+ *    crate's solver picked conditions equivalent to the published ones
+ *    that fill vector lanes well, for a prefix run on every block, and left
+ *    the rest to a scalar tail that few blocks reach. Those choices are
+ *    kept as tables, run by a short loop per form. See ubc_check.c.
+ *
+ * The compression is portable C, and spills the expanded message schedule
+ * and the two states that recompression starts from (at steps 58 and 65)
+ * as it goes.
  */
 
 #include "../git-compat-util.h"
@@ -221,18 +225,21 @@ struct backend {
 	int (*available)(void);
 };
 
-/* sha1dc's own UBC check. */
-static uint32_t ubc_check_sha1dc(const uint32_t w[80])
-{
-	uint32_t mask;
-
-	ubc_check(w, &mask);
-	return mask;
-}
-
 /* In order of preference. */
 static const struct backend backends[] = {
-	{ "portable", compress_portable, ubc_check_sha1dc, NULL },
+#ifdef SHA1DC_HAVE_AVX2
+	{ "portable+avx2", compress_portable, sha1dc_ubc_check_avx2,
+	  sha1dc_avx2_available },
+#endif
+#ifdef SHA1DC_HAVE_SSE2
+	{ "portable+sse2", compress_portable, sha1dc_ubc_check_sse2,
+	  NULL },
+#endif
+#ifdef SHA1DC_HAVE_NEON
+	{ "portable+neon", compress_portable, sha1dc_ubc_check_neon,
+	  NULL },
+#endif
+	{ "portable", compress_portable, sha1dc_ubc_check_scalar, NULL },
 };
 
 static int usable(const struct backend *be)
diff --git a/sha1dc-accel/ubc_check.c b/sha1dc-accel/ubc_check.c
new file mode 100644
index 0000000000..f95b799f9d
--- /dev/null
+++ b/sha1dc-accel/ubc_check.c
@@ -0,0 +1,1789 @@
+/*
+ * The unavoidable-bitconditions (UBC) check of SHA-1 collision detection,
+ * in one form per instruction set.
+ *
+ * Every form returns the same mask as ubc_check() in sha1dc/ubc_check.c:
+ * one bit per disturbance vector (DV) that the expanded message w[] has not
+ * ruled out. A DV is ruled out as soon as one of its bitconditions fails;
+ * each condition says that bit a of w[i] XOR bit b of w[j] equals c.
+ *
+ * Each form runs in two parts. The prefix tests a fixed set of conditions,
+ * chosen and packed into vector lanes for that instruction set, on every
+ * block; it rules out every DV for almost all blocks. The few blocks that
+ * survive it run the tail, which tests the remaining conditions of each DV
+ * still alive.
+ *
+ * The tables in this file were derived from the output of the solver in
+ * the "sha1dc" Rust crate by Sam Reis (https://github.com/srijs/sha1dc,
+ * commit 426b4afd), which picks the conditions and their packing. They are
+ * used under the MIT license:
+ *
+ *   Copyright (c) 2017 Marc Stevens (Cryptology Group, Centrum Wiskunde &
+ *   Informatica)
+ *   Copyright (c) 2017 Dan Shumow (Microsoft Research)
+ *   Copyright (c) 2026 Sam Reis
+ *
+ *   Permission is hereby granted, free of charge, to any person obtaining a
+ *   copy of this software and associated documentation files (the
+ *   "Software"), to deal in the Software without restriction, including
+ *   without limitation the rights to use, copy, modify, merge, publish,
+ *   distribute, sublicense, and/or sell copies of the Software, and to
+ *   permit persons to whom the Software is furnished to do so, subject to
+ *   the following conditions:
+ *
+ *   The above copyright notice and this permission notice shall be included
+ *   in all copies or substantial portions of the Software.
+ *
+ *   THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
+ *   OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
+ *   MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
+ *   IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY
+ *   CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT,
+ *   TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
+ *   SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
+ */
+
+#include "../git-compat-util.h"
+#include "internal.h"
+
+#define DV_I_43_0_BIT ((uint32_t)1 << 0)
+#define DV_I_44_0_BIT ((uint32_t)1 << 1)
+#define DV_I_45_0_BIT ((uint32_t)1 << 2)
+#define DV_I_46_0_BIT ((uint32_t)1 << 3)
+#define DV_I_46_2_BIT ((uint32_t)1 << 4)
+#define DV_I_47_0_BIT ((uint32_t)1 << 5)
+#define DV_I_47_2_BIT ((uint32_t)1 << 6)
+#define DV_I_48_0_BIT ((uint32_t)1 << 7)
+#define DV_I_48_2_BIT ((uint32_t)1 << 8)
+#define DV_I_49_0_BIT ((uint32_t)1 << 9)
+#define DV_I_49_2_BIT ((uint32_t)1 << 10)
+#define DV_I_50_0_BIT ((uint32_t)1 << 11)
+#define DV_I_50_2_BIT ((uint32_t)1 << 12)
+#define DV_I_51_0_BIT ((uint32_t)1 << 13)
+#define DV_I_51_2_BIT ((uint32_t)1 << 14)
+#define DV_I_52_0_BIT ((uint32_t)1 << 15)
+#define DV_II_45_0_BIT ((uint32_t)1 << 16)
+#define DV_II_46_0_BIT ((uint32_t)1 << 17)
+#define DV_II_46_2_BIT ((uint32_t)1 << 18)
+#define DV_II_47_0_BIT ((uint32_t)1 << 19)
+#define DV_II_48_0_BIT ((uint32_t)1 << 20)
+#define DV_II_49_0_BIT ((uint32_t)1 << 21)
+#define DV_II_49_2_BIT ((uint32_t)1 << 22)
+#define DV_II_50_0_BIT ((uint32_t)1 << 23)
+#define DV_II_50_2_BIT ((uint32_t)1 << 24)
+#define DV_II_51_0_BIT ((uint32_t)1 << 25)
+#define DV_II_51_2_BIT ((uint32_t)1 << 26)
+#define DV_II_52_0_BIT ((uint32_t)1 << 27)
+#define DV_II_53_0_BIT ((uint32_t)1 << 28)
+#define DV_II_54_0_BIT ((uint32_t)1 << 29)
+#define DV_II_55_0_BIT ((uint32_t)1 << 30)
+#define DV_II_56_0_BIT ((uint32_t)1 << 31)
+
+/*
+ * The prefixes loop over constant tables. Unrolling those loops fully lets
+ * the compiler fold each entry into the code as immediates, which is what
+ * makes them fast; without it they are two to three times slower.
+ */
+#if defined(__clang__) || (defined(__GNUC__) && __GNUC__ >= 8)
+#define UNROLL_TABLE _Pragma("GCC unroll 128")
+#else
+#define UNROLL_TABLE
+#endif
+
+/*
+ * The tail loops run over a few conditions at a time, too few to gain from
+ * vectorizing; clang would otherwise vectorize them when targeting AVX2.
+ */
+#ifdef __clang__
+#define NO_VECTORIZE _Pragma("clang loop vectorize(disable)")
+#else
+#define NO_VECTORIZE
+#endif
+
+/* A bitcondition: bit a of w[i] XOR bit b of w[j] must equal c. */
+struct ubc_cond {
+	uint8_t i, a, j, b, c;
+};
+
+static inline uint32_t cond_fails(const uint32_t *w, const struct ubc_cond *c)
+{
+	return (((w[c->i] >> c->a) ^ (w[c->j] >> c->b) ^ c->c) & 1);
+}
+
+/* A condition of the scalar prefix, and the DVs it rules out if it fails. */
+struct ubc_prefix_cond {
+	struct ubc_cond cond;
+	uint32_t dvs;
+};
+
+/*
+ * A group of conditions for a vector prefix, one per lane. In lane k, the
+ * bit selected by test[k] of (w[lo + k] >> lo_shift) ^ (w[hi + k] >> hi_shift)
+ * must equal want, or the DVs in dvs[k] are ruled out. The lanes share
+ * everything but test[] and dvs[]; an unused lane has both 0.
+ *
+ * No group reads past w[66].
+ */
+struct ubc_group4 {
+	uint8_t lo, lo_shift, hi, hi_shift, want;
+	uint32_t test[4];
+	uint32_t dvs[4];
+};
+
+struct ubc_group8 {
+	uint8_t lo, lo_shift, hi, hi_shift, want;
+	uint32_t test[8];
+	uint32_t dvs[8];
+};
+
+/* The tail conditions of DV d are checks[spans[d].start] onwards. */
+struct tail_span {
+	uint16_t start;
+	uint8_t len;
+};
+
+/*
+ * Runs the tail conditions of each DV still alive in `mask`, and clears
+ * the DVs with a failing one. Each condition fails about half the time, so
+ * it tests all of a DV's conditions rather than branch on each one.
+ */
+static inline uint32_t run_tail(const uint32_t *w, uint32_t mask,
+				const struct ubc_cond *checks,
+				const struct tail_span *spans)
+{
+	uint32_t out = mask;
+	while (mask) {
+		unsigned d = sha1dc_ctz(mask);
+		const struct ubc_cond *c = checks + spans[d].start;
+		const struct ubc_cond *end = c + spans[d].len;
+		uint32_t fail = 0;
+
+		NO_VECTORIZE
+		for (; c < end; c++)
+			fail |= cond_fails(w, c);
+		out &= ~(fail << d);
+		mask &= mask - 1;
+	}
+	return out;
+}
+
+/* scalar form */
+
+static const struct ubc_prefix_cond scalar_prefix_conds[] = {
+	{ { 44, 29, 45, 29, 0 },
+	  DV_I_48_0_BIT | DV_I_51_0_BIT | DV_I_52_0_BIT | DV_II_45_0_BIT |
+	  DV_II_46_0_BIT | DV_II_50_0_BIT | DV_II_51_0_BIT },
+	{ { 46, 29, 47, 29, 0 },
+	  DV_I_43_0_BIT | DV_I_50_0_BIT | DV_II_47_0_BIT | DV_II_48_0_BIT |
+	  DV_II_52_0_BIT | DV_II_53_0_BIT },
+	{ { 45, 4, 48, 29, 0 },
+	  DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT |
+	  DV_II_49_0_BIT | DV_II_54_0_BIT },
+	{ { 49, 29, 50, 29, 0 },
+	  DV_I_46_0_BIT | DV_II_45_0_BIT | DV_II_50_0_BIT | DV_II_51_0_BIT |
+	  DV_II_55_0_BIT | DV_II_56_0_BIT },
+	{ { 40, 29, 41, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_47_0_BIT | DV_I_48_0_BIT | DV_II_46_0_BIT |
+	  DV_II_47_0_BIT | DV_II_56_0_BIT },
+	{ { 36, 1, 37, 6, 1 },
+	  DV_I_47_2_BIT | DV_I_50_2_BIT | DV_II_46_2_BIT },
+	{ { 47, 29, 48, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_51_0_BIT | DV_II_48_0_BIT | DV_II_49_0_BIT |
+	  DV_II_53_0_BIT | DV_II_54_0_BIT },
+	{ { 39, 1, 40, 6, 1 },
+	  DV_I_46_2_BIT | DV_I_50_2_BIT | DV_II_49_2_BIT },
+	{ { 40, 1, 41, 6, 1 },
+	  DV_I_47_2_BIT | DV_I_51_2_BIT | DV_II_50_2_BIT },
+	{ { 41, 1, 42, 6, 1 },
+	  DV_I_48_2_BIT | DV_II_46_2_BIT | DV_II_51_2_BIT },
+	{ { 43, 4, 46, 29, 0 },
+	  DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT |
+	  DV_II_47_0_BIT | DV_II_52_0_BIT },
+	{ { 46, 4, 49, 29, 0 },
+	  DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT |
+	  DV_II_50_0_BIT | DV_II_55_0_BIT },
+	{ { 45, 6, 47, 6, 0 },
+	  DV_I_47_2_BIT | DV_I_49_2_BIT | DV_I_51_2_BIT },
+	{ { 45, 29, 46, 29, 0 },
+	  DV_I_49_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT | DV_II_47_0_BIT |
+	  DV_II_51_0_BIT | DV_II_52_0_BIT },
+	{ { 44, 4, 47, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT |
+	  DV_II_48_0_BIT | DV_II_53_0_BIT },
+	{ { 48, 29, 49, 29, 0 },
+	  DV_I_45_0_BIT | DV_I_52_0_BIT | DV_II_49_0_BIT | DV_II_50_0_BIT |
+	  DV_II_54_0_BIT | DV_II_55_0_BIT },
+	{ { 44, 6, 46, 6, 0 },
+	  DV_I_46_2_BIT | DV_I_48_2_BIT | DV_I_50_2_BIT },
+	{ { 47, 4, 50, 29, 0 },
+	  DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	  DV_II_51_0_BIT | DV_II_56_0_BIT },
+	{ { 35, 1, 36, 6, 1 },
+	  DV_I_46_2_BIT | DV_I_49_2_BIT },
+	{ { 44, 1, 45, 6, 1 },
+	  DV_I_51_2_BIT | DV_II_49_2_BIT },
+	{ { 42, 6, 43, 1, 0 },
+	  DV_II_46_2_BIT | DV_II_51_2_BIT },
+	{ { 37, 4, 40, 29, 0 },
+	  DV_I_43_0_BIT | DV_I_47_0_BIT | DV_II_46_0_BIT | DV_II_53_0_BIT |
+	  DV_II_55_0_BIT },
+	{ { 41, 6, 42, 1, 0 },
+	  DV_I_51_2_BIT | DV_II_50_2_BIT },
+	{ { 40, 4, 43, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_50_0_BIT | DV_II_49_0_BIT |
+	  DV_II_56_0_BIT },
+	{ { 41, 4, 44, 29, 0 },
+	  DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_51_0_BIT |
+	  DV_II_45_0_BIT | DV_II_50_0_BIT },
+	{ { 52, 29, 53, 29, 0 },
+	  DV_I_49_0_BIT | DV_II_45_0_BIT | DV_II_48_0_BIT | DV_II_53_0_BIT |
+	  DV_II_54_0_BIT },
+	{ { 40, 6, 41, 1, 0 },
+	  DV_I_50_2_BIT | DV_II_49_2_BIT },
+	{ { 46, 6, 47, 1, 0 },
+	  DV_I_46_2_BIT | DV_II_50_2_BIT },
+	{ { 47, 6, 48, 1, 0 },
+	  DV_I_47_2_BIT | DV_II_51_2_BIT },
+	{ { 42, 4, 45, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_52_0_BIT |
+	  DV_II_46_0_BIT | DV_II_51_0_BIT },
+	{ { 37, 1, 38, 6, 1 },
+	  DV_I_48_2_BIT | DV_I_51_2_BIT },
+	{ { 43, 6, 45, 6, 0 },
+	  DV_I_47_2_BIT | DV_I_49_2_BIT },
+	{ { 53, 29, 54, 29, 0 },
+	  DV_I_50_0_BIT | DV_II_46_0_BIT | DV_II_49_0_BIT | DV_II_54_0_BIT |
+	  DV_II_55_0_BIT },
+	{ { 41, 29, 42, 29, 0 },
+	  DV_I_45_0_BIT | DV_I_48_0_BIT | DV_I_49_0_BIT | DV_II_47_0_BIT |
+	  DV_II_48_0_BIT },
+	{ { 50, 29, 51, 29, 0 },
+	  DV_I_47_0_BIT | DV_II_46_0_BIT | DV_II_51_0_BIT | DV_II_52_0_BIT |
+	  DV_II_56_0_BIT },
+	{ { 61, 2, 62, 7, 1 },
+	  DV_I_46_2_BIT | DV_II_46_2_BIT },
+	{ { 46, 6, 48, 6, 0 },
+	  DV_I_48_2_BIT | DV_I_50_2_BIT },
+	{ { 39, 4, 42, 29, 0 },
+	  DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_49_0_BIT | DV_II_48_0_BIT |
+	  DV_II_55_0_BIT },
+	{ { 43, 29, 44, 29, 0 },
+	  DV_I_47_0_BIT | DV_I_50_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	  DV_II_49_0_BIT | DV_II_50_0_BIT },
+	{ { 47, 6, 49, 6, 0 },
+	  DV_I_49_2_BIT | DV_I_51_2_BIT },
+	{ { 51, 29, 52, 29, 0 },
+	  DV_I_48_0_BIT | DV_II_47_0_BIT | DV_II_52_0_BIT | DV_II_53_0_BIT },
+	{ { 36, 0, 37, 5, 1 },
+	  DV_II_49_2_BIT },
+	{ { 37, 0, 38, 5, 1 },
+	  DV_II_50_2_BIT },
+	{ { 38, 0, 39, 5, 1 },
+	  DV_II_51_2_BIT },
+	{ { 38, 4, 41, 29, 0 },
+	  DV_I_44_0_BIT | DV_I_48_0_BIT | DV_II_47_0_BIT | DV_II_54_0_BIT |
+	  DV_II_56_0_BIT },
+	{ { 50, 6, 51, 1, 0 },
+	  DV_I_50_2_BIT | DV_II_46_2_BIT },
+	{ { 42, 6, 44, 6, 0 },
+	  DV_I_46_2_BIT | DV_I_48_2_BIT },
+	{ { 48, 4, 51, 29, 0 },
+	  DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	  DV_II_52_0_BIT },
+	{ { 42, 29, 43, 29, 0 },
+	  DV_I_46_0_BIT | DV_I_49_0_BIT | DV_I_50_0_BIT | DV_II_48_0_BIT |
+	  DV_II_49_0_BIT },
+	{ { 54, 29, 55, 29, 0 },
+	  DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_50_0_BIT | DV_II_55_0_BIT |
+	  DV_II_56_0_BIT },
+	{ { 38, 1, 39, 6, 1 },
+	  DV_I_49_2_BIT },
+	{ { 45, 1, 46, 6, 1 },
+	  DV_II_50_2_BIT },
+	{ { 46, 1, 47, 6, 1 },
+	  DV_II_51_2_BIT },
+	{ { 50, 1, 51, 6, 1 },
+	  DV_II_49_2_BIT },
+	{ { 39, 4, 40, 29, 1 },
+	  DV_I_43_0_BIT | DV_II_53_0_BIT | DV_II_55_0_BIT },
+	{ { 55, 29, 56, 29, 0 },
+	  DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_51_0_BIT | DV_II_56_0_BIT },
+	{ { 48, 6, 50, 6, 0 },
+	  DV_I_50_2_BIT | DV_II_46_2_BIT },
+	{ { 49, 4, 52, 29, 0 },
+	  DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT | DV_II_47_0_BIT |
+	  DV_II_53_0_BIT },
+	{ { 40, 4, 41, 29, 1 },
+	  DV_I_44_0_BIT | DV_II_54_0_BIT | DV_II_56_0_BIT },
+	{ { 41, 4, 42, 29, 1 },
+	  DV_I_43_0_BIT | DV_I_45_0_BIT | DV_II_55_0_BIT },
+	{ { 42, 1, 43, 6, 1 },
+	  DV_I_49_2_BIT },
+	{ { 51, 1, 52, 6, 1 },
+	  DV_II_50_2_BIT },
+	{ { 52, 1, 53, 6, 1 },
+	  DV_II_51_2_BIT },
+	{ { 62, 2, 63, 7, 1 },
+	  DV_I_47_2_BIT },
+	{ { 63, 2, 64, 7, 1 },
+	  DV_I_48_2_BIT },
+	{ { 45, 6, 46, 1, 0 },
+	  DV_II_49_2_BIT },
+	{ { 36, 4, 40, 29, 0 },
+	  DV_I_46_0_BIT | DV_I_49_0_BIT | DV_II_45_0_BIT | DV_II_48_0_BIT },
+	{ { 50, 4, 53, 29, 0 },
+	  DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT | DV_II_48_0_BIT |
+	  DV_II_54_0_BIT },
+	{ { 56, 29, 57, 29, 0 },
+	  DV_II_49_0_BIT | DV_II_52_0_BIT },
+	{ { 43, 4, 44, 29, 1 },
+	  DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT }
+};
+
+static const struct ubc_cond scalar_tail_checks[] = {
+	/* DV_I_43_0 */
+	{ 58, 0, 59, 5, 1 },
+	{ 58, 0, 63, 30, 1 },
+	{ 61, 1, 62, 6, 1 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_I_44_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 42, 4, 44, 4, 1 },
+	{ 59, 0, 60, 5, 1 },
+	{ 59, 0, 64, 30, 1 },
+	{ 62, 1, 63, 6, 1 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_I_45_0 */
+	{ 43, 4, 45, 4, 1 },
+	{ 60, 0, 61, 5, 1 },
+	{ 63, 1, 64, 6, 1 },
+	{ 35, 4, 39, 29, 0 },
+	/* DV_I_46_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 42, 4, 44, 4, 1 },
+	{ 44, 4, 46, 4, 1 },
+	{ 61, 0, 62, 5, 1 },
+	/* DV_I_46_2 */
+	{ 39, 1, 42, 6, 1 },
+	/* DV_I_47_0 */
+	{ 43, 4, 45, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 62, 0, 63, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	/* DV_I_47_2 */
+	{ 40, 1, 43, 6, 1 },
+	/* DV_I_48_0 */
+	{ 42, 4, 44, 4, 1 },
+	{ 44, 4, 46, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	/* DV_I_48_2 */
+	{ 41, 1, 49, 1, 1 },
+	/* DV_I_49_0 */
+	{ 43, 4, 45, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 47, 4, 49, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	/* DV_I_49_2 */
+	{ 38, 1, 40, 1, 1 },
+	{ 42, 1, 50, 1, 1 },
+	/* DV_I_50_0 */
+	{ 36, 4, 37, 4, 1 },
+	{ 44, 4, 46, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 48, 4, 50, 4, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	/* DV_I_50_2 */
+	{ 43, 1, 44, 6, 1 },
+	/* DV_I_51_0 */
+	{ 37, 4, 38, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 47, 4, 49, 4, 1 },
+	{ 52, 29, 55, 29, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	/* DV_I_51_2 */
+	{ 44, 1, 51, 6, 1 },
+	{ 44, 1, 52, 1, 1 },
+	{ 35, 5, 39, 30, 0 },
+	{ 37, 1, 37, 6, 0 },
+	/* DV_I_52_0 */
+	{ 38, 4, 39, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 48, 4, 50, 4, 1 },
+	{ 53, 29, 56, 29, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	/* DV_II_45_0 */
+	{ 47, 4, 49, 4, 1 },
+	{ 60, 0, 61, 5, 1 },
+	{ 63, 1, 64, 6, 1 },
+	{ 41, 4, 45, 29, 0 },
+	/* DV_II_46_0 */
+	{ 48, 4, 50, 4, 1 },
+	{ 61, 0, 62, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	/* DV_II_46_2 */
+	{ 47, 1, 48, 6, 1 },
+	/* DV_II_47_0 */
+	{ 52, 29, 55, 29, 1 },
+	{ 62, 0, 63, 5, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	/* DV_II_48_0 */
+	{ 35, 30, 36, 3, 1 },
+	{ 35, 30, 40, 28, 1 },
+	{ 52, 29, 55, 29, 1 },
+	{ 53, 29, 56, 29, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	/* DV_II_49_0 */
+	{ 36, 30, 37, 3, 1 },
+	{ 36, 30, 41, 28, 1 },
+	{ 53, 29, 56, 29, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	{ 53, 4, 56, 29, 0 },
+	/* DV_II_49_2 */
+	{ 36, 0, 41, 30, 1 },
+	{ 50, 1, 53, 6, 1 },
+	{ 50, 1, 54, 1, 1 },
+	/* DV_II_50_0 */
+	{ 37, 30, 38, 3, 1 },
+	{ 37, 30, 42, 28, 1 },
+	{ 55, 29, 58, 29, 1 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	{ 57, 29, 58, 29, 0 },
+	/* DV_II_50_2 */
+	{ 37, 0, 42, 30, 1 },
+	{ 51, 1, 54, 6, 1 },
+	{ 51, 1, 55, 1, 1 },
+	/* DV_II_51_0 */
+	{ 38, 30, 39, 3, 1 },
+	{ 38, 30, 43, 28, 1 },
+	{ 55, 29, 58, 29, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 53, 4, 56, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	/* DV_II_51_2 */
+	{ 38, 0, 43, 30, 1 },
+	{ 52, 1, 55, 6, 1 },
+	{ 52, 1, 56, 1, 1 },
+	/* DV_II_52_0 */
+	{ 36, 4, 38, 4, 1 },
+	{ 39, 30, 40, 3, 1 },
+	{ 39, 30, 44, 28, 1 },
+	{ 54, 4, 60, 29, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 40, 4, 44, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	{ 56, 4, 59, 29, 0 },
+	/* DV_II_53_0 */
+	{ 55, 4, 57, 4, 1 },
+	{ 55, 4, 61, 29, 1 },
+	{ 41, 3, 45, 28, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	{ 57, 29, 58, 29, 0 },
+	/* DV_II_54_0 */
+	{ 36, 4, 38, 4, 1 },
+	{ 42, 3, 46, 28, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 56, 4, 59, 29, 0 },
+	{ 56, 4, 58, 29, 0 },
+	{ 58, 4, 62, 29, 0 },
+	/* DV_II_55_0 */
+	{ 43, 3, 47, 28, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	{ 57, 4, 59, 29, 0 },
+	{ 59, 4, 63, 29, 0 },
+	/* DV_II_56_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 44, 3, 48, 28, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	{ 60, 4, 64, 29, 0 }
+};
+
+static const struct tail_span scalar_tail_spans[32] = {
+	{ 0, 4 },	/* DV_I_43_0 */
+	{ 4, 6 },	/* DV_I_44_0 */
+	{ 10, 4 },	/* DV_I_45_0 */
+	{ 14, 4 },	/* DV_I_46_0 */
+	{ 18, 1 },	/* DV_I_46_2 */
+	{ 19, 4 },	/* DV_I_47_0 */
+	{ 23, 1 },	/* DV_I_47_2 */
+	{ 24, 6 },	/* DV_I_48_0 */
+	{ 30, 1 },	/* DV_I_48_2 */
+	{ 31, 4 },	/* DV_I_49_0 */
+	{ 35, 2 },	/* DV_I_49_2 */
+	{ 37, 6 },	/* DV_I_50_0 */
+	{ 43, 1 },	/* DV_I_50_2 */
+	{ 44, 8 },	/* DV_I_51_0 */
+	{ 52, 4 },	/* DV_I_51_2 */
+	{ 56, 7 },	/* DV_I_52_0 */
+	{ 63, 4 },	/* DV_II_45_0 */
+	{ 67, 4 },	/* DV_II_46_0 */
+	{ 71, 1 },	/* DV_II_46_2 */
+	{ 72, 7 },	/* DV_II_47_0 */
+	{ 79, 8 },	/* DV_II_48_0 */
+	{ 87, 7 },	/* DV_II_49_0 */
+	{ 94, 3 },	/* DV_II_49_2 */
+	{ 97, 8 },	/* DV_II_50_0 */
+	{ 105, 3 },	/* DV_II_50_2 */
+	{ 108, 8 },	/* DV_II_51_0 */
+	{ 116, 3 },	/* DV_II_51_2 */
+	{ 119, 9 },	/* DV_II_52_0 */
+	{ 128, 7 },	/* DV_II_53_0 */
+	{ 135, 6 },	/* DV_II_54_0 */
+	{ 141, 5 },	/* DV_II_55_0 */
+	{ 146, 5 }	/* DV_II_56_0 */
+};
+
+static uint32_t scalar_prefix(const uint32_t *w)
+{
+	uint32_t mask = 0xFFFFFFFF;
+	size_t i;
+
+	UNROLL_TABLE
+	for (i = 0; i < ARRAY_SIZE(scalar_prefix_conds); i++) {
+		const struct ubc_prefix_cond *p = &scalar_prefix_conds[i];
+		mask &= ~(p->dvs & (0 - cond_fails(w, &p->cond)));
+	}
+	return mask;
+}
+
+uint32_t sha1dc_ubc_check_scalar(const uint32_t w[80])
+{
+	uint32_t mask = scalar_prefix(w);
+	/* Every check only clears bits, so an empty mask settles it. */
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, scalar_tail_checks, scalar_tail_spans);
+}
+
+/* neon form */
+
+#ifdef SHA1DC_HAVE_NEON
+
+static const struct ubc_group4 neon_groups[] = {
+	{ 35, 0, 36, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 0 },
+	  { DV_I_46_2_BIT | DV_I_49_2_BIT,
+	    DV_I_47_2_BIT | DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_48_2_BIT | DV_I_51_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 36, 0, 37, 5, 1,
+	  { 1u << 0, 1u << 0, 1u << 1, 1u << 1 },
+	  { DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_46_2_BIT | DV_I_50_2_BIT | DV_II_49_2_BIT } },
+	{ 36, 0, 38, 0, 1,
+	  { 1u << 4, 1u << 4, 1u << 1, 1u << 4 },
+	  { DV_II_52_0_BIT | DV_II_54_0_BIT,
+	    DV_I_43_0_BIT | DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_II_55_0_BIT } },
+	{ 36, 0, 40, 25, 0,
+	  { 1u << 4, 1u << 5, 1u << 5, 1u << 5 },
+	  { DV_I_46_0_BIT | DV_I_49_0_BIT | DV_II_45_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 38, 0, 40, 0, 1,
+	  { 1u << 4, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_44_0_BIT | DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_50_2_BIT | DV_II_49_2_BIT,
+	    DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_II_46_2_BIT | DV_II_51_2_BIT } },
+	{ 37, 0, 40, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_43_0_BIT | DV_I_47_0_BIT | DV_II_46_0_BIT |
+	    DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_48_0_BIT | DV_II_47_0_BIT |
+	    DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_49_0_BIT | DV_II_48_0_BIT |
+	    DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_50_0_BIT | DV_II_49_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 40, 0, 41, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_47_2_BIT | DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_I_48_2_BIT | DV_II_46_2_BIT | DV_II_51_2_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_50_2_BIT } },
+	{ 40, 0, 41, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_44_0_BIT | DV_I_47_0_BIT | DV_I_48_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_56_0_BIT,
+	    DV_I_45_0_BIT | DV_I_48_0_BIT | DV_I_49_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_I_46_0_BIT | DV_I_49_0_BIT | DV_I_50_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT,
+	    DV_I_47_0_BIT | DV_I_50_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_49_0_BIT | DV_II_50_0_BIT } },
+	{ 41, 0, 43, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_47_2_BIT,
+	    DV_I_46_2_BIT | DV_I_48_2_BIT,
+	    DV_I_47_2_BIT | DV_I_49_2_BIT,
+	    DV_I_46_2_BIT | DV_I_48_2_BIT | DV_I_50_2_BIT } },
+	{ 41, 0, 44, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_51_0_BIT |
+	    DV_II_45_0_BIT | DV_II_50_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_52_0_BIT |
+	    DV_II_46_0_BIT | DV_II_51_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT |
+	    DV_II_47_0_BIT | DV_II_52_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT |
+	    DV_II_48_0_BIT | DV_II_53_0_BIT } },
+	{ 44, 0, 45, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_51_2_BIT | DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    DV_II_46_2_BIT } },
+	{ 44, 0, 45, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_48_0_BIT | DV_I_51_0_BIT | DV_I_52_0_BIT | DV_II_45_0_BIT |
+	    DV_II_46_0_BIT | DV_II_50_0_BIT | DV_II_51_0_BIT,
+	    DV_I_49_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_51_0_BIT | DV_II_52_0_BIT,
+	    DV_I_43_0_BIT | DV_I_50_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT | DV_II_52_0_BIT | DV_II_53_0_BIT,
+	    DV_I_44_0_BIT | DV_I_51_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT | DV_II_53_0_BIT | DV_II_54_0_BIT } },
+	{ 45, 5, 46, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_II_49_2_BIT,
+	    DV_I_46_2_BIT | DV_II_50_2_BIT,
+	    DV_I_47_2_BIT | DV_II_51_2_BIT,
+	    DV_I_48_2_BIT } },
+	{ 45, 0, 47, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_47_2_BIT | DV_I_49_2_BIT | DV_I_51_2_BIT,
+	    DV_I_48_2_BIT | DV_I_50_2_BIT,
+	    DV_I_49_2_BIT | DV_I_51_2_BIT,
+	    DV_I_50_2_BIT | DV_II_46_2_BIT } },
+	{ 45, 0, 47, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 4, 1u << 4 },
+	  { DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT,
+	    DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT,
+	    DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT,
+	    DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT } },
+	{ 45, 0, 48, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT |
+	    DV_II_49_0_BIT | DV_II_54_0_BIT,
+	    DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT |
+	    DV_II_50_0_BIT | DV_II_55_0_BIT,
+	    DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_51_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_52_0_BIT } },
+	{ 48, 0, 49, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_45_0_BIT | DV_I_52_0_BIT | DV_II_49_0_BIT |
+	    DV_II_50_0_BIT | DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_46_0_BIT | DV_II_45_0_BIT | DV_II_50_0_BIT |
+	    DV_II_51_0_BIT | DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_47_0_BIT | DV_II_46_0_BIT | DV_II_51_0_BIT |
+	    DV_II_52_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_II_47_0_BIT | DV_II_52_0_BIT |
+	    DV_II_53_0_BIT } },
+	{ 52, 0, 53, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_49_0_BIT | DV_II_45_0_BIT | DV_II_48_0_BIT |
+	    DV_II_53_0_BIT | DV_II_54_0_BIT,
+	    DV_I_50_0_BIT | DV_II_46_0_BIT | DV_II_49_0_BIT |
+	    DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_50_0_BIT |
+	    DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_51_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 52, 0, 55, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_48_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_49_0_BIT,
+	    DV_II_49_0_BIT | DV_II_50_0_BIT,
+	    DV_II_50_0_BIT | DV_II_51_0_BIT } },
+	{ 60, 0, 61, 5, 1,
+	  { 1u << 0, 1u << 2, 1u << 2, 1u << 2 },
+	  { DV_I_45_0_BIT | DV_II_45_0_BIT,
+	    DV_I_46_2_BIT | DV_II_46_2_BIT,
+	    DV_I_47_2_BIT,
+	    DV_I_48_2_BIT } }
+};
+
+static const struct ubc_cond neon_tail_checks[] = {
+	/* DV_I_43_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 58, 0, 59, 5, 1 },
+	{ 58, 0, 63, 30, 1 },
+	{ 61, 1, 62, 6, 1 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_I_44_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 59, 0, 60, 5, 1 },
+	{ 59, 0, 64, 30, 1 },
+	{ 62, 1, 63, 6, 1 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_I_45_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 63, 1, 64, 6, 1 },
+	{ 35, 4, 39, 29, 0 },
+	/* DV_I_46_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 44, 4, 46, 4, 1 },
+	{ 61, 0, 62, 5, 1 },
+	/* DV_I_46_2 */
+	{ 39, 1, 42, 6, 1 },
+	/* DV_I_47_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 62, 0, 63, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	/* DV_I_48_0 */
+	{ 44, 4, 46, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	/* DV_I_49_0 */
+	{ 45, 4, 47, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 49, 4, 52, 29, 0 },
+	/* DV_I_49_2 */
+	{ 42, 1, 50, 1, 1 },
+	/* DV_I_50_0 */
+	{ 36, 4, 37, 4, 1 },
+	{ 44, 4, 46, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	{ 50, 4, 53, 29, 0 },
+	/* DV_I_50_2 */
+	{ 48, 6, 51, 1, 0 },
+	/* DV_I_51_0 */
+	{ 37, 4, 38, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 49, 4, 52, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	/* DV_I_51_2 */
+	{ 44, 1, 51, 6, 1 },
+	{ 44, 1, 52, 1, 1 },
+	{ 35, 5, 39, 30, 0 },
+	{ 37, 1, 37, 6, 0 },
+	/* DV_I_52_0 */
+	{ 38, 4, 39, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 50, 4, 53, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	/* DV_II_45_0 */
+	{ 63, 1, 64, 6, 1 },
+	{ 41, 4, 45, 29, 0 },
+	{ 49, 4, 52, 29, 0 },
+	/* DV_II_46_0 */
+	{ 61, 0, 62, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 50, 4, 53, 29, 0 },
+	/* DV_II_46_2 */
+	{ 48, 6, 51, 1, 0 },
+	/* DV_II_47_0 */
+	{ 62, 0, 63, 5, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 49, 4, 52, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	/* DV_II_48_0 */
+	{ 35, 30, 36, 3, 1 },
+	{ 35, 30, 40, 28, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 50, 4, 53, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	/* DV_II_49_0 */
+	{ 36, 30, 37, 3, 1 },
+	{ 36, 30, 41, 28, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	{ 53, 4, 56, 29, 0 },
+	/* DV_II_49_2 */
+	{ 50, 1, 51, 6, 1 },
+	{ 50, 1, 53, 6, 1 },
+	{ 50, 1, 54, 1, 1 },
+	/* DV_II_50_0 */
+	{ 37, 30, 38, 3, 1 },
+	{ 37, 30, 42, 28, 1 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	/* DV_II_50_2 */
+	{ 51, 1, 52, 6, 1 },
+	{ 51, 1, 54, 6, 1 },
+	{ 51, 1, 55, 1, 1 },
+	/* DV_II_51_0 */
+	{ 38, 30, 39, 3, 1 },
+	{ 38, 30, 43, 28, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 53, 4, 56, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	/* DV_II_51_2 */
+	{ 52, 1, 53, 6, 1 },
+	{ 52, 1, 55, 6, 1 },
+	{ 52, 1, 56, 1, 1 },
+	/* DV_II_52_0 */
+	{ 39, 30, 40, 3, 1 },
+	{ 39, 30, 44, 28, 1 },
+	{ 54, 4, 56, 4, 1 },
+	{ 54, 4, 60, 29, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 40, 4, 44, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	{ 56, 4, 59, 29, 0 },
+	/* DV_II_53_0 */
+	{ 55, 4, 57, 4, 1 },
+	{ 55, 4, 61, 29, 1 },
+	{ 41, 3, 45, 28, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 49, 4, 52, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	{ 55, 4, 57, 29, 0 },
+	/* DV_II_54_0 */
+	{ 42, 3, 46, 28, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 50, 4, 53, 29, 0 },
+	{ 56, 4, 59, 29, 0 },
+	{ 56, 4, 58, 29, 0 },
+	{ 58, 4, 62, 29, 0 },
+	/* DV_II_55_0 */
+	{ 43, 3, 47, 28, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 51, 4, 54, 29, 0 },
+	{ 57, 4, 59, 29, 0 },
+	{ 59, 4, 63, 29, 0 },
+	/* DV_II_56_0 */
+	{ 40, 4, 42, 4, 1 },
+	{ 44, 3, 48, 28, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 52, 4, 55, 29, 0 },
+	{ 60, 4, 64, 29, 0 }
+};
+
+static const struct tail_span neon_tail_spans[32] = {
+	{ 0, 5 },	/* DV_I_43_0 */
+	{ 5, 5 },	/* DV_I_44_0 */
+	{ 10, 3 },	/* DV_I_45_0 */
+	{ 13, 3 },	/* DV_I_46_0 */
+	{ 16, 1 },	/* DV_I_46_2 */
+	{ 17, 4 },	/* DV_I_47_0 */
+	{ 21, 0 },	/* DV_I_47_2 */
+	{ 21, 5 },	/* DV_I_48_0 */
+	{ 26, 0 },	/* DV_I_48_2 */
+	{ 26, 3 },	/* DV_I_49_0 */
+	{ 29, 1 },	/* DV_I_49_2 */
+	{ 30, 6 },	/* DV_I_50_0 */
+	{ 36, 1 },	/* DV_I_50_2 */
+	{ 37, 7 },	/* DV_I_51_0 */
+	{ 44, 4 },	/* DV_I_51_2 */
+	{ 48, 6 },	/* DV_I_52_0 */
+	{ 54, 3 },	/* DV_II_45_0 */
+	{ 57, 4 },	/* DV_II_46_0 */
+	{ 61, 1 },	/* DV_II_46_2 */
+	{ 62, 7 },	/* DV_II_47_0 */
+	{ 69, 7 },	/* DV_II_48_0 */
+	{ 76, 6 },	/* DV_II_49_0 */
+	{ 82, 3 },	/* DV_II_49_2 */
+	{ 85, 6 },	/* DV_II_50_0 */
+	{ 91, 3 },	/* DV_II_50_2 */
+	{ 94, 7 },	/* DV_II_51_0 */
+	{ 101, 3 },	/* DV_II_51_2 */
+	{ 104, 9 },	/* DV_II_52_0 */
+	{ 113, 8 },	/* DV_II_53_0 */
+	{ 121, 6 },	/* DV_II_54_0 */
+	{ 127, 5 },	/* DV_II_55_0 */
+	{ 132, 5 }	/* DV_II_56_0 */
+};
+
+static uint32_t neon_prefix(const uint32_t *w)
+{
+	uint32x4_t acc = vdupq_n_u32(0);
+	uint32x2_t folded;
+	size_t i;
+
+	UNROLL_TABLE
+	for (i = 0; i < ARRAY_SIZE(neon_groups); i++) {
+		const struct ubc_group4 *g = &neon_groups[i];
+		uint32x4_t lo = vld1q_u32(w + g->lo);
+		uint32x4_t hi = vld1q_u32(w + g->hi);
+		uint32x4_t dvs = vld1q_u32(g->dvs);
+		uint32x4_t set, fail;
+
+		lo = vshlq_u32(lo, vdupq_n_s32(-(int32_t)g->lo_shift));
+		hi = vshlq_u32(hi, vdupq_n_s32(-(int32_t)g->hi_shift));
+		set = vtstq_u32(veorq_u32(lo, hi), vld1q_u32(g->test));
+		/*
+		 * The DVs of the lanes where the bit is not g->want. Each lane
+		 * of set is all ones or zero, so a saturating subtraction keeps
+		 * dvs where the bit is clear, and min keeps it where it is set.
+		 */
+		fail = g->want ? vqsubq_u32(dvs, set) : vminq_u32(set, dvs);
+		acc = vorrq_u32(acc, fail);
+	}
+
+	folded = vorr_u32(vget_low_u32(acc), vget_high_u32(acc));
+	return ~vget_lane_u32(vorr_u32(folded, vdup_lane_u32(folded, 1)), 0);
+}
+
+uint32_t sha1dc_ubc_check_neon(const uint32_t w[80])
+{
+	uint32_t mask = neon_prefix(w);
+	/* Every check only clears bits, so an empty mask settles it. */
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, neon_tail_checks, neon_tail_spans);
+}
+
+#endif /* SHA1DC_HAVE_NEON */
+
+/* sse2 form */
+
+#ifdef SHA1DC_HAVE_SSE2
+
+static const struct ubc_group4 sse2_groups[] = {
+	{ 35, 0, 36, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 0 },
+	  { DV_I_46_2_BIT | DV_I_49_2_BIT,
+	    DV_I_47_2_BIT | DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_48_2_BIT | DV_I_51_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 36, 0, 37, 5, 1,
+	  { 1u << 0, 1u << 0, 1u << 1, 1u << 1 },
+	  { DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_46_2_BIT | DV_I_50_2_BIT | DV_II_49_2_BIT } },
+	{ 36, 0, 38, 0, 1,
+	  { 1u << 4, 1u << 4, 1u << 1, 1u << 4 },
+	  { DV_II_52_0_BIT | DV_II_54_0_BIT,
+	    DV_I_43_0_BIT | DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_II_55_0_BIT } },
+	{ 37, 0, 40, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_43_0_BIT | DV_I_47_0_BIT | DV_II_46_0_BIT |
+	    DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_48_0_BIT | DV_II_47_0_BIT |
+	    DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_49_0_BIT | DV_II_48_0_BIT |
+	    DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_50_0_BIT | DV_II_49_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 36, 0, 40, 25, 0,
+	  { 1u << 4, 1u << 5, 1u << 5, 1u << 5 },
+	  { DV_I_46_0_BIT | DV_I_49_0_BIT | DV_II_45_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 38, 0, 40, 0, 1,
+	  { 1u << 4, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_44_0_BIT | DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_50_2_BIT | DV_II_49_2_BIT,
+	    DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_II_46_2_BIT | DV_II_51_2_BIT } },
+	{ 40, 0, 41, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_47_2_BIT | DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_I_48_2_BIT | DV_II_46_2_BIT | DV_II_51_2_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_50_2_BIT } },
+	{ 40, 0, 41, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_44_0_BIT | DV_I_47_0_BIT | DV_I_48_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_56_0_BIT,
+	    DV_I_45_0_BIT | DV_I_48_0_BIT | DV_I_49_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_I_46_0_BIT | DV_I_49_0_BIT | DV_I_50_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT,
+	    DV_I_47_0_BIT | DV_I_50_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_49_0_BIT | DV_II_50_0_BIT } },
+	{ 41, 0, 44, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_51_0_BIT |
+	    DV_II_45_0_BIT | DV_II_50_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_52_0_BIT |
+	    DV_II_46_0_BIT | DV_II_51_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT |
+	    DV_II_47_0_BIT | DV_II_52_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT |
+	    DV_II_48_0_BIT | DV_II_53_0_BIT } },
+	{ 42, 0, 44, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_46_2_BIT | DV_I_48_2_BIT,
+	    DV_I_47_2_BIT | DV_I_49_2_BIT,
+	    DV_I_46_2_BIT | DV_I_48_2_BIT | DV_I_50_2_BIT,
+	    DV_I_47_2_BIT | DV_I_49_2_BIT | DV_I_51_2_BIT } },
+	{ 40, 0, 44, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 29, 1u << 29 },
+	  { DV_I_46_2_BIT,
+	    DV_I_47_2_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT } },
+	{ 44, 0, 45, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_51_2_BIT | DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    DV_II_46_2_BIT } },
+	{ 44, 0, 45, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_48_0_BIT | DV_I_51_0_BIT | DV_I_52_0_BIT | DV_II_45_0_BIT |
+	    DV_II_46_0_BIT | DV_II_50_0_BIT | DV_II_51_0_BIT,
+	    DV_I_49_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_51_0_BIT | DV_II_52_0_BIT,
+	    DV_I_43_0_BIT | DV_I_50_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT | DV_II_52_0_BIT | DV_II_53_0_BIT,
+	    DV_I_44_0_BIT | DV_I_51_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT | DV_II_53_0_BIT | DV_II_54_0_BIT } },
+	{ 45, 5, 46, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_II_49_2_BIT,
+	    DV_I_46_2_BIT | DV_II_50_2_BIT,
+	    DV_I_47_2_BIT | DV_II_51_2_BIT,
+	    DV_I_48_2_BIT } },
+	{ 45, 0, 47, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 4 },
+	  { DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT,
+	    DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT,
+	    DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT,
+	    DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT } },
+	{ 46, 0, 48, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_48_2_BIT | DV_I_50_2_BIT,
+	    DV_I_49_2_BIT | DV_I_51_2_BIT,
+	    DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_51_2_BIT } },
+	{ 45, 0, 48, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT |
+	    DV_II_49_0_BIT | DV_II_54_0_BIT,
+	    DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT |
+	    DV_II_50_0_BIT | DV_II_55_0_BIT,
+	    DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_51_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_52_0_BIT } },
+	{ 48, 0, 49, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_45_0_BIT | DV_I_52_0_BIT | DV_II_49_0_BIT |
+	    DV_II_50_0_BIT | DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_46_0_BIT | DV_II_45_0_BIT | DV_II_50_0_BIT |
+	    DV_II_51_0_BIT | DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_47_0_BIT | DV_II_46_0_BIT | DV_II_51_0_BIT |
+	    DV_II_52_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_II_47_0_BIT | DV_II_52_0_BIT |
+	    DV_II_53_0_BIT } },
+	{ 49, 5, 50, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_49_2_BIT,
+	    DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_51_2_BIT,
+	    0 } },
+	{ 49, 0, 52, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_47_0_BIT | DV_II_53_0_BIT,
+	    DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_48_0_BIT | DV_II_54_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_49_0_BIT |
+	    DV_II_55_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_50_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 49, 0, 53, 0, 1,
+	  { 1u << 29, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_II_45_0_BIT,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 52, 0, 53, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_49_0_BIT | DV_II_45_0_BIT | DV_II_48_0_BIT |
+	    DV_II_53_0_BIT | DV_II_54_0_BIT,
+	    DV_I_50_0_BIT | DV_II_46_0_BIT | DV_II_49_0_BIT |
+	    DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_50_0_BIT |
+	    DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_51_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 53, 5, 54, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    0 } },
+	{ 52, 0, 55, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_48_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_49_0_BIT,
+	    DV_II_49_0_BIT | DV_II_50_0_BIT,
+	    DV_II_50_0_BIT | DV_II_51_0_BIT } },
+	{ 53, 0, 56, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_II_49_0_BIT | DV_II_51_0_BIT,
+	    DV_II_50_0_BIT | DV_II_52_0_BIT,
+	    DV_II_51_0_BIT | DV_II_53_0_BIT,
+	    DV_II_52_0_BIT | DV_II_54_0_BIT } },
+	{ 60, 0, 61, 5, 1,
+	  { 1u << 0, 1u << 2, 1u << 2, 1u << 2 },
+	  { DV_I_45_0_BIT | DV_II_45_0_BIT,
+	    DV_I_46_2_BIT | DV_II_46_2_BIT,
+	    DV_I_47_2_BIT,
+	    DV_I_48_2_BIT } }
+};
+
+static const struct ubc_cond sse2_tail_checks[] = {
+	/* DV_I_43_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 58, 0, 59, 5, 1 },
+	{ 58, 0, 63, 30, 1 },
+	{ 61, 1, 62, 6, 1 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_I_44_0 */
+	{ 59, 0, 60, 5, 1 },
+	{ 59, 0, 64, 30, 1 },
+	{ 62, 1, 63, 6, 1 },
+	{ 38, 4, 42, 4, 0 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_I_45_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 63, 1, 64, 6, 1 },
+	{ 35, 4, 39, 29, 0 },
+	/* DV_I_46_0 */
+	{ 61, 0, 62, 5, 1 },
+	/* DV_I_47_0 */
+	{ 41, 4, 43, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 62, 0, 63, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	/* DV_I_48_0 */
+	{ 46, 4, 48, 4, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	/* DV_I_49_0 */
+	{ 45, 4, 47, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 45, 4, 49, 4, 0 },
+	/* DV_I_50_0 */
+	{ 36, 4, 37, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	/* DV_I_51_0 */
+	{ 37, 4, 38, 4, 1 },
+	{ 45, 4, 47, 4, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 45, 4, 49, 4, 0 },
+	/* DV_I_51_2 */
+	{ 35, 5, 39, 30, 0 },
+	{ 37, 1, 37, 6, 0 },
+	/* DV_I_52_0 */
+	{ 38, 4, 39, 4, 1 },
+	{ 46, 4, 48, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	/* DV_II_45_0 */
+	{ 63, 1, 64, 6, 1 },
+	{ 41, 4, 45, 29, 0 },
+	/* DV_II_46_0 */
+	{ 61, 0, 62, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	/* DV_II_47_0 */
+	{ 62, 0, 63, 5, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_II_48_0 */
+	{ 35, 30, 36, 3, 1 },
+	{ 35, 30, 40, 28, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_II_49_0 */
+	{ 36, 30, 37, 3, 1 },
+	{ 36, 30, 41, 28, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	/* DV_II_49_2 */
+	{ 50, 1, 51, 6, 1 },
+	/* DV_II_50_0 */
+	{ 37, 30, 38, 3, 1 },
+	{ 37, 30, 42, 28, 1 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	/* DV_II_50_2 */
+	{ 51, 1, 52, 6, 1 },
+	/* DV_II_51_0 */
+	{ 38, 30, 39, 3, 1 },
+	{ 38, 30, 43, 28, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 42, 4, 46, 29, 0 },
+	/* DV_II_51_2 */
+	{ 52, 1, 53, 6, 1 },
+	/* DV_II_52_0 */
+	{ 39, 30, 40, 3, 1 },
+	{ 39, 30, 44, 28, 1 },
+	{ 54, 4, 56, 4, 1 },
+	{ 54, 4, 60, 29, 1 },
+	{ 56, 29, 59, 29, 1 },
+	{ 40, 4, 44, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_II_53_0 */
+	{ 55, 4, 57, 4, 1 },
+	{ 55, 4, 61, 29, 1 },
+	{ 41, 3, 45, 28, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 55, 4, 57, 29, 0 },
+	/* DV_II_54_0 */
+	{ 42, 3, 46, 28, 0 },
+	{ 42, 4, 46, 29, 0 },
+	{ 56, 4, 58, 29, 0 },
+	{ 58, 4, 62, 29, 0 },
+	/* DV_II_55_0 */
+	{ 43, 3, 47, 28, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 57, 4, 59, 29, 0 },
+	{ 59, 4, 63, 29, 0 },
+	/* DV_II_56_0 */
+	{ 38, 4, 42, 4, 0 },
+	{ 44, 3, 48, 28, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 60, 4, 64, 29, 0 }
+};
+
+static const struct tail_span sse2_tail_spans[32] = {
+	{ 0, 5 },	/* DV_I_43_0 */
+	{ 5, 5 },	/* DV_I_44_0 */
+	{ 10, 3 },	/* DV_I_45_0 */
+	{ 13, 1 },	/* DV_I_46_0 */
+	{ 14, 0 },	/* DV_I_46_2 */
+	{ 14, 4 },	/* DV_I_47_0 */
+	{ 18, 0 },	/* DV_I_47_2 */
+	{ 18, 4 },	/* DV_I_48_0 */
+	{ 22, 0 },	/* DV_I_48_2 */
+	{ 22, 3 },	/* DV_I_49_0 */
+	{ 25, 0 },	/* DV_I_49_2 */
+	{ 25, 4 },	/* DV_I_50_0 */
+	{ 29, 0 },	/* DV_I_50_2 */
+	{ 29, 6 },	/* DV_I_51_0 */
+	{ 35, 2 },	/* DV_I_51_2 */
+	{ 37, 4 },	/* DV_I_52_0 */
+	{ 41, 2 },	/* DV_II_45_0 */
+	{ 43, 3 },	/* DV_II_46_0 */
+	{ 46, 0 },	/* DV_II_46_2 */
+	{ 46, 5 },	/* DV_II_47_0 */
+	{ 51, 5 },	/* DV_II_48_0 */
+	{ 56, 4 },	/* DV_II_49_0 */
+	{ 60, 1 },	/* DV_II_49_2 */
+	{ 61, 4 },	/* DV_II_50_0 */
+	{ 65, 1 },	/* DV_II_50_2 */
+	{ 66, 5 },	/* DV_II_51_0 */
+	{ 71, 1 },	/* DV_II_51_2 */
+	{ 72, 7 },	/* DV_II_52_0 */
+	{ 79, 6 },	/* DV_II_53_0 */
+	{ 85, 4 },	/* DV_II_54_0 */
+	{ 89, 4 },	/* DV_II_55_0 */
+	{ 93, 4 }	/* DV_II_56_0 */
+};
+
+SHA1DC_TARGET_SSE2
+static uint32_t sse2_prefix(const uint32_t *w)
+{
+	const __m128i zero = _mm_setzero_si128();
+	__m128i acc = zero;
+	size_t i;
+
+	UNROLL_TABLE
+	for (i = 0; i < ARRAY_SIZE(sse2_groups); i++) {
+		const struct ubc_group4 *g = &sse2_groups[i];
+		__m128i lo = _mm_loadu_si128((const __m128i *)(w + g->lo));
+		__m128i hi = _mm_loadu_si128((const __m128i *)(w + g->hi));
+		__m128i test = _mm_loadu_si128((const __m128i *)g->test);
+		__m128i dvs = _mm_loadu_si128((const __m128i *)g->dvs);
+		__m128i clear, fail;
+
+		lo = _mm_srl_epi32(lo, _mm_cvtsi32_si128(g->lo_shift));
+		hi = _mm_srl_epi32(hi, _mm_cvtsi32_si128(g->hi_shift));
+		clear = _mm_and_si128(_mm_xor_si128(lo, hi), test);
+		clear = _mm_cmpeq_epi32(clear, zero);
+		/* The DVs of the lanes where the bit is not g->want. */
+		fail = g->want ? _mm_and_si128(clear, dvs) :
+				 _mm_andnot_si128(clear, dvs);
+		acc = _mm_or_si128(acc, fail);
+	}
+
+	acc = _mm_or_si128(acc, _mm_shuffle_epi32(acc, 0x4E));
+	acc = _mm_or_si128(acc, _mm_shuffle_epi32(acc, 0xB1));
+	return ~(uint32_t)_mm_cvtsi128_si32(acc);
+}
+
+SHA1DC_TARGET_SSE2
+uint32_t sha1dc_ubc_check_sse2(const uint32_t w[80])
+{
+	uint32_t mask = sse2_prefix(w);
+	/* Every check only clears bits, so an empty mask settles it. */
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, sse2_tail_checks, sse2_tail_spans);
+}
+
+#endif /* SHA1DC_HAVE_SSE2 */
+
+/* avx2 form */
+
+#ifdef SHA1DC_HAVE_AVX2
+
+static const struct ubc_group8 avx2_groups[] = {
+	{ 35, 0, 36, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 0,
+	    1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_46_2_BIT | DV_I_49_2_BIT,
+	    DV_I_47_2_BIT | DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_48_2_BIT | DV_I_51_2_BIT,
+	    DV_II_51_2_BIT,
+	    DV_I_46_2_BIT | DV_I_50_2_BIT | DV_II_49_2_BIT,
+	    DV_I_47_2_BIT | DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_I_48_2_BIT | DV_II_46_2_BIT | DV_II_51_2_BIT,
+	    DV_I_49_2_BIT } },
+	{ 36, 0, 38, 0, 1,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4,
+	    1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_II_52_0_BIT | DV_II_54_0_BIT,
+	    DV_I_43_0_BIT | DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT,
+	    DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT } },
+	{ 37, 0, 38, 5, 1,
+	  { 1u << 0, 1u << 1, 0, 0,
+	    0, 0, 1u << 1, 1u << 1 },
+	  { DV_II_50_2_BIT,
+	    DV_I_49_2_BIT,
+	    0,
+	    0,
+	    0,
+	    0,
+	    DV_I_50_2_BIT,
+	    DV_I_51_2_BIT | DV_II_49_2_BIT } },
+	{ 35, 0, 39, 25, 0,
+	  { 1u << 5, 1u << 4, 1u << 5, 1u << 5,
+	    1u << 5, 1u << 3, 1u << 3, 1u << 4 },
+	  { DV_I_51_2_BIT,
+	    DV_I_46_0_BIT | DV_I_49_0_BIT | DV_II_45_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    DV_II_52_0_BIT,
+	    DV_II_53_0_BIT,
+	    DV_I_52_0_BIT | DV_II_46_0_BIT | DV_II_51_0_BIT |
+	    DV_II_54_0_BIT } },
+	{ 37, 0, 40, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4,
+	    1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_43_0_BIT | DV_I_47_0_BIT | DV_II_46_0_BIT |
+	    DV_II_53_0_BIT | DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_48_0_BIT | DV_II_47_0_BIT |
+	    DV_II_54_0_BIT | DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_49_0_BIT | DV_II_48_0_BIT |
+	    DV_II_55_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_50_0_BIT | DV_II_49_0_BIT |
+	    DV_II_56_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_51_0_BIT |
+	    DV_II_45_0_BIT | DV_II_50_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_52_0_BIT |
+	    DV_II_46_0_BIT | DV_II_51_0_BIT,
+	    DV_I_43_0_BIT | DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT |
+	    DV_II_47_0_BIT | DV_II_52_0_BIT,
+	    DV_I_44_0_BIT | DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT |
+	    DV_II_48_0_BIT | DV_II_53_0_BIT } },
+	{ 39, 5, 40, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1,
+	    1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_49_2_BIT,
+	    DV_I_50_2_BIT | DV_II_49_2_BIT,
+	    DV_I_51_2_BIT | DV_II_50_2_BIT,
+	    DV_II_46_2_BIT | DV_II_51_2_BIT,
+	    0,
+	    0,
+	    DV_II_49_2_BIT,
+	    DV_I_46_2_BIT | DV_II_50_2_BIT } },
+	{ 40, 0, 41, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29,
+	    1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_44_0_BIT | DV_I_47_0_BIT | DV_I_48_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_56_0_BIT,
+	    DV_I_45_0_BIT | DV_I_48_0_BIT | DV_I_49_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT,
+	    DV_I_46_0_BIT | DV_I_49_0_BIT | DV_I_50_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT,
+	    DV_I_47_0_BIT | DV_I_50_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_49_0_BIT | DV_II_50_0_BIT,
+	    DV_I_48_0_BIT | DV_I_51_0_BIT | DV_I_52_0_BIT | DV_II_45_0_BIT |
+	    DV_II_46_0_BIT | DV_II_50_0_BIT | DV_II_51_0_BIT,
+	    DV_I_49_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_47_0_BIT | DV_II_51_0_BIT | DV_II_52_0_BIT,
+	    DV_I_43_0_BIT | DV_I_50_0_BIT | DV_II_47_0_BIT |
+	    DV_II_48_0_BIT | DV_II_52_0_BIT | DV_II_53_0_BIT,
+	    DV_I_44_0_BIT | DV_I_51_0_BIT | DV_II_48_0_BIT |
+	    DV_II_49_0_BIT | DV_II_53_0_BIT | DV_II_54_0_BIT } },
+	{ 40, 0, 42, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6,
+	    1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_46_2_BIT,
+	    DV_I_47_2_BIT,
+	    DV_I_46_2_BIT | DV_I_48_2_BIT,
+	    DV_I_47_2_BIT | DV_I_49_2_BIT,
+	    DV_I_46_2_BIT | DV_I_48_2_BIT | DV_I_50_2_BIT,
+	    DV_I_47_2_BIT | DV_I_49_2_BIT | DV_I_51_2_BIT,
+	    DV_I_48_2_BIT | DV_I_50_2_BIT,
+	    DV_I_49_2_BIT | DV_I_51_2_BIT } },
+	{ 45, 0, 46, 5, 1,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1,
+	    1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    DV_II_46_2_BIT,
+	    0,
+	    0,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT } },
+	{ 45, 0, 48, 25, 0,
+	  { 1u << 4, 1u << 4, 1u << 4, 1u << 4,
+	    1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_45_0_BIT | DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT |
+	    DV_II_49_0_BIT | DV_II_54_0_BIT,
+	    DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT |
+	    DV_II_50_0_BIT | DV_II_55_0_BIT,
+	    DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_51_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_52_0_BIT,
+	    DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT |
+	    DV_II_47_0_BIT | DV_II_53_0_BIT,
+	    DV_I_50_0_BIT | DV_I_52_0_BIT | DV_II_46_0_BIT |
+	    DV_II_48_0_BIT | DV_II_54_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_49_0_BIT |
+	    DV_II_55_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_50_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 47, 5, 48, 0, 0,
+	  { 1u << 1, 1u << 1, 1u << 1, 1u << 1,
+	    1u << 1, 1u << 1, 1u << 1, 1u << 1 },
+	  { DV_I_47_2_BIT | DV_II_51_2_BIT,
+	    DV_I_48_2_BIT,
+	    DV_I_49_2_BIT,
+	    DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_51_2_BIT,
+	    0,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT } },
+	{ 47, 0, 49, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29,
+	    1u << 4, 1u << 4, 1u << 4, 1u << 4 },
+	  { DV_I_46_0_BIT | DV_I_48_0_BIT | DV_I_50_0_BIT,
+	    DV_I_47_0_BIT | DV_I_49_0_BIT | DV_I_51_0_BIT,
+	    DV_I_48_0_BIT | DV_I_50_0_BIT | DV_I_52_0_BIT,
+	    DV_I_49_0_BIT | DV_I_51_0_BIT | DV_II_45_0_BIT,
+	    DV_II_49_0_BIT,
+	    DV_II_50_0_BIT,
+	    DV_II_51_0_BIT,
+	    DV_II_52_0_BIT } },
+	{ 48, 0, 49, 0, 0,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29,
+	    1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_45_0_BIT | DV_I_52_0_BIT | DV_II_49_0_BIT |
+	    DV_II_50_0_BIT | DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_46_0_BIT | DV_II_45_0_BIT | DV_II_50_0_BIT |
+	    DV_II_51_0_BIT | DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_47_0_BIT | DV_II_46_0_BIT | DV_II_51_0_BIT |
+	    DV_II_52_0_BIT | DV_II_56_0_BIT,
+	    DV_I_48_0_BIT | DV_II_47_0_BIT | DV_II_52_0_BIT |
+	    DV_II_53_0_BIT,
+	    DV_I_49_0_BIT | DV_II_45_0_BIT | DV_II_48_0_BIT |
+	    DV_II_53_0_BIT | DV_II_54_0_BIT,
+	    DV_I_50_0_BIT | DV_II_46_0_BIT | DV_II_49_0_BIT |
+	    DV_II_54_0_BIT | DV_II_55_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_50_0_BIT |
+	    DV_II_55_0_BIT | DV_II_56_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_51_0_BIT |
+	    DV_II_56_0_BIT } },
+	{ 48, 0, 50, 0, 0,
+	  { 1u << 6, 1u << 6, 1u << 6, 1u << 6,
+	    1u << 6, 1u << 6, 1u << 6, 1u << 6 },
+	  { DV_I_50_2_BIT | DV_II_46_2_BIT,
+	    DV_I_51_2_BIT,
+	    0,
+	    DV_II_49_2_BIT,
+	    DV_II_50_2_BIT,
+	    DV_II_51_2_BIT,
+	    0,
+	    0 } },
+	{ 51, 0, 54, 0, 1,
+	  { 1u << 29, 1u << 29, 1u << 29, 1u << 29,
+	    1u << 29, 1u << 29, 1u << 29, 1u << 29 },
+	  { DV_I_50_0_BIT | DV_II_46_0_BIT | DV_II_47_0_BIT,
+	    DV_I_51_0_BIT | DV_II_47_0_BIT | DV_II_48_0_BIT,
+	    DV_I_52_0_BIT | DV_II_48_0_BIT | DV_II_49_0_BIT,
+	    DV_II_49_0_BIT | DV_II_50_0_BIT,
+	    DV_II_50_0_BIT | DV_II_51_0_BIT,
+	    DV_II_51_0_BIT | DV_II_52_0_BIT,
+	    DV_II_52_0_BIT,
+	    DV_II_53_0_BIT } },
+	{ 58, 0, 59, 5, 1,
+	  { 1u << 0, 1u << 0, 1u << 0, 1u << 2,
+	    1u << 2, 1u << 2, 0, 0 },
+	  { DV_I_43_0_BIT,
+	    DV_I_44_0_BIT,
+	    DV_I_45_0_BIT | DV_II_45_0_BIT,
+	    DV_I_46_2_BIT | DV_II_46_2_BIT,
+	    DV_I_47_2_BIT,
+	    DV_I_48_2_BIT,
+	    0,
+	    0 } }
+};
+
+static const struct ubc_cond avx2_tail_checks[] = {
+	/* DV_I_43_0 */
+	{ 58, 0, 63, 30, 1 },
+	{ 61, 1, 62, 6, 1 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_I_44_0 */
+	{ 59, 0, 64, 30, 1 },
+	{ 62, 1, 63, 6, 1 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_I_45_0 */
+	{ 63, 1, 64, 6, 1 },
+	{ 35, 4, 39, 29, 0 },
+	/* DV_I_46_0 */
+	{ 61, 0, 62, 5, 1 },
+	/* DV_I_47_0 */
+	{ 62, 0, 63, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	/* DV_I_48_0 */
+	{ 63, 0, 64, 5, 1 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	/* DV_I_49_0 */
+	{ 39, 4, 43, 29, 0 },
+	/* DV_I_50_0 */
+	{ 36, 4, 37, 4, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	{ 46, 4, 50, 4, 0 },
+	/* DV_I_51_0 */
+	{ 37, 4, 38, 4, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	/* DV_I_51_2 */
+	{ 37, 1, 37, 6, 0 },
+	/* DV_I_52_0 */
+	{ 38, 4, 39, 4, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 46, 4, 50, 4, 0 },
+	/* DV_II_45_0 */
+	{ 63, 1, 64, 6, 1 },
+	{ 41, 4, 45, 29, 0 },
+	/* DV_II_46_0 */
+	{ 61, 0, 62, 5, 1 },
+	{ 37, 4, 41, 29, 0 },
+	/* DV_II_47_0 */
+	{ 62, 0, 63, 5, 1 },
+	{ 35, 3, 39, 28, 0 },
+	{ 35, 4, 39, 29, 0 },
+	{ 38, 4, 42, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	/* DV_II_48_0 */
+	{ 35, 30, 36, 3, 1 },
+	{ 35, 30, 40, 28, 1 },
+	{ 63, 0, 64, 5, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	/* DV_II_49_0 */
+	{ 36, 30, 37, 3, 1 },
+	{ 36, 30, 41, 28, 1 },
+	{ 37, 4, 41, 29, 0 },
+	{ 40, 4, 44, 29, 0 },
+	/* DV_II_49_2 */
+	{ 36, 0, 37, 5, 1 },
+	/* DV_II_50_0 */
+	{ 37, 30, 38, 3, 1 },
+	{ 37, 30, 42, 28, 1 },
+	{ 38, 4, 42, 29, 0 },
+	{ 41, 4, 45, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	/* DV_II_51_0 */
+	{ 38, 30, 39, 3, 1 },
+	{ 38, 30, 43, 28, 1 },
+	{ 39, 4, 43, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	/* DV_II_51_2 */
+	{ 52, 1, 56, 1, 1 },
+	/* DV_II_52_0 */
+	{ 39, 30, 40, 3, 1 },
+	{ 40, 4, 44, 29, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 54, 4, 57, 29, 0 },
+	{ 56, 4, 59, 29, 0 },
+	/* DV_II_53_0 */
+	{ 55, 4, 57, 4, 1 },
+	{ 41, 4, 45, 29, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 55, 4, 58, 29, 0 },
+	{ 55, 4, 57, 29, 0 },
+	/* DV_II_54_0 */
+	{ 42, 3, 46, 28, 0 },
+	{ 56, 4, 59, 29, 0 },
+	{ 56, 4, 58, 29, 0 },
+	{ 58, 4, 62, 29, 0 },
+	/* DV_II_55_0 */
+	{ 43, 3, 47, 28, 0 },
+	{ 43, 4, 47, 29, 0 },
+	{ 57, 4, 59, 29, 0 },
+	{ 59, 4, 63, 29, 0 },
+	/* DV_II_56_0 */
+	{ 44, 3, 48, 28, 0 },
+	{ 44, 4, 48, 29, 0 },
+	{ 60, 4, 64, 29, 0 }
+};
+
+static const struct tail_span avx2_tail_spans[32] = {
+	{ 0, 3 },	/* DV_I_43_0 */
+	{ 3, 3 },	/* DV_I_44_0 */
+	{ 6, 2 },	/* DV_I_45_0 */
+	{ 8, 1 },	/* DV_I_46_0 */
+	{ 9, 0 },	/* DV_I_46_2 */
+	{ 9, 2 },	/* DV_I_47_0 */
+	{ 11, 0 },	/* DV_I_47_2 */
+	{ 11, 3 },	/* DV_I_48_0 */
+	{ 14, 0 },	/* DV_I_48_2 */
+	{ 14, 1 },	/* DV_I_49_0 */
+	{ 15, 0 },	/* DV_I_49_2 */
+	{ 15, 4 },	/* DV_I_50_0 */
+	{ 19, 0 },	/* DV_I_50_2 */
+	{ 19, 4 },	/* DV_I_51_0 */
+	{ 23, 1 },	/* DV_I_51_2 */
+	{ 24, 3 },	/* DV_I_52_0 */
+	{ 27, 2 },	/* DV_II_45_0 */
+	{ 29, 2 },	/* DV_II_46_0 */
+	{ 31, 0 },	/* DV_II_46_2 */
+	{ 31, 5 },	/* DV_II_47_0 */
+	{ 36, 5 },	/* DV_II_48_0 */
+	{ 41, 4 },	/* DV_II_49_0 */
+	{ 45, 1 },	/* DV_II_49_2 */
+	{ 46, 5 },	/* DV_II_50_0 */
+	{ 51, 0 },	/* DV_II_50_2 */
+	{ 51, 4 },	/* DV_II_51_0 */
+	{ 55, 1 },	/* DV_II_51_2 */
+	{ 56, 5 },	/* DV_II_52_0 */
+	{ 61, 5 },	/* DV_II_53_0 */
+	{ 66, 4 },	/* DV_II_54_0 */
+	{ 70, 4 },	/* DV_II_55_0 */
+	{ 74, 3 }	/* DV_II_56_0 */
+};
+
+SHA1DC_TARGET_AVX2
+static uint32_t avx2_prefix(const uint32_t *w)
+{
+	const __m256i zero = _mm256_setzero_si256();
+	__m256i acc = zero;
+	__m128i half;
+	size_t i;
+
+	UNROLL_TABLE
+	for (i = 0; i < ARRAY_SIZE(avx2_groups); i++) {
+		const struct ubc_group8 *g = &avx2_groups[i];
+		__m256i lo = _mm256_loadu_si256((const __m256i *)(w + g->lo));
+		__m256i hi = _mm256_loadu_si256((const __m256i *)(w + g->hi));
+		__m256i test = _mm256_loadu_si256((const __m256i *)g->test);
+		__m256i dvs = _mm256_loadu_si256((const __m256i *)g->dvs);
+		__m256i clear, fail;
+
+		lo = _mm256_srl_epi32(lo, _mm_cvtsi32_si128(g->lo_shift));
+		hi = _mm256_srl_epi32(hi, _mm_cvtsi32_si128(g->hi_shift));
+		clear = _mm256_and_si256(_mm256_xor_si256(lo, hi), test);
+		clear = _mm256_cmpeq_epi32(clear, zero);
+		/* The DVs of the lanes where the bit is not g->want. */
+		fail = g->want ? _mm256_and_si256(clear, dvs) :
+				 _mm256_andnot_si256(clear, dvs);
+		acc = _mm256_or_si256(acc, fail);
+	}
+
+	half = _mm_or_si128(_mm256_castsi256_si128(acc),
+			    _mm256_extracti128_si256(acc, 1));
+	half = _mm_or_si128(half, _mm_shuffle_epi32(half, 0x4E));
+	half = _mm_or_si128(half, _mm_shuffle_epi32(half, 0xB1));
+	return ~(uint32_t)_mm_cvtsi128_si32(half);
+}
+
+SHA1DC_TARGET_AVX2
+uint32_t sha1dc_ubc_check_avx2(const uint32_t w[80])
+{
+	uint32_t mask = avx2_prefix(w);
+	/* Every check only clears bits, so an empty mask settles it. */
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, avx2_tail_checks, avx2_tail_spans);
+}
+
+#endif /* SHA1DC_HAVE_AVX2 */
diff --git a/sha1dc-accel/x86.c b/sha1dc-accel/x86.c
new file mode 100644
index 0000000000..7c942fe4de
--- /dev/null
+++ b/sha1dc-accel/x86.c
@@ -0,0 +1,38 @@
+/*
+ * The x86-64 parts of sha1.c: detecting what the CPU has.
+ */
+
+#include "../git-compat-util.h"
+#include "internal.h"
+
+#ifdef SHA1DC_HAVE_SSE2
+
+#include <cpuid.h>
+
+/* Leaf 7 of CPUID, subleaf 0, if the CPU has it. */
+static int cpuid_7(unsigned int *ebx)
+{
+	unsigned int eax, ecx, edx;
+
+	if (__get_cpuid_max(0, NULL) < 7)
+		return 0;
+	__cpuid_count(7, 0, eax, *ebx, ecx, edx);
+	return 1;
+}
+
+int sha1dc_avx2_available(void)
+{
+	unsigned int eax, ebx, ecx, edx, xcr0_lo, xcr0_hi;
+
+	if (!__get_cpuid(1, &eax, &ebx, &ecx, &edx))
+		return 0;
+	/* OSXSAVE and AVX, then whether the OS saves the YMM registers. */
+	if ((ecx & (1u << 27 | 1u << 28)) != (1u << 27 | 1u << 28))
+		return 0;
+	__asm__("xgetbv" : "=a"(xcr0_lo), "=d"(xcr0_hi) : "c"(0));
+	if ((xcr0_lo & 6) != 6)
+		return 0;
+	return cpuid_7(&ebx) && (ebx & (1u << 5));
+}
+
+#endif /* SHA1DC_HAVE_SSE2 */
diff --git a/t/unit-tests/u-sha1dc.c b/t/unit-tests/u-sha1dc.c
index 31fcd68451..15e71fb023 100644
--- a/t/unit-tests/u-sha1dc.c
+++ b/t/unit-tests/u-sha1dc.c
@@ -126,6 +126,76 @@ static void every_backend_agrees_with_sha1dc(void)
 	cl_assert_equal_i(sha1dc_accel_select(orig), 0);
 }
 
+/* Asserts only on a mismatch; clar's assertions are too slow to run 10^5 times. */
+#define check_form(got, want) do { \
+		uint32_t got_ = (got); \
+		if (got_ != (want)) \
+			cl_assert_equal_i(got_, (want)); \
+	} while (0)
+
+/*
+ * Checks every form of the UBC check against sha1dc/'s on `w`, and next to
+ * it, with each bit flipped in turn if `flip`.
+ */
+static void check_ubc_forms(uint32_t w[80], int flip, int avx2)
+{
+	int k;
+
+	for (k = 0; k <= (flip ? 80 * 32 : 0); k++) {
+		uint32_t want;
+
+		if (k)
+			w[(k - 1) / 32] ^= 1u << ((k - 1) % 32);
+		ubc_check(w, &want);
+		check_form(sha1dc_ubc_check_scalar(w), want);
+#ifdef SHA1DC_HAVE_SSE2
+		check_form(sha1dc_ubc_check_sse2(w), want);
+#endif
+#ifdef SHA1DC_HAVE_AVX2
+		if (avx2)
+			check_form(sha1dc_ubc_check_avx2(w), want);
+#endif
+#ifdef SHA1DC_HAVE_NEON
+		check_form(sha1dc_ubc_check_neon(w), want);
+#endif
+		(void)avx2;
+		if (k)
+			w[(k - 1) / 32] ^= 1u << ((k - 1) % 32);
+	}
+}
+
+static void ubc_forms_agree_with_sha1dc(void)
+{
+	uint32_t w[80];
+	int i, k, dv, avx2 = 0;
+
+#ifdef SHA1DC_HAVE_AVX2
+	/* Once: CPUID can take microseconds under a hypervisor. */
+	avx2 = sha1dc_avx2_available();
+#endif
+	rng_seed(2);
+	for (i = 0; i < 20000; i++) {
+		for (k = 0; k < 80; k++)
+			w[k] = rng();
+		check_ubc_forms(w, 0, avx2);
+	}
+
+	/*
+	 * Random schedules rarely get past the vector prefix of a form, so
+	 * look for one that keeps each DV alive, and check around it.
+	 */
+	for (dv = 0; dv < 32; dv++) {
+		uint32_t mask;
+
+		do {
+			for (k = 0; k < 80; k++)
+				w[k] = rng();
+			ubc_check(w, &mask);
+		} while (!(mask & (1u << dv)));
+		check_ubc_forms(w, 1, avx2);
+	}
+}
+
 #endif /* HAVE_SHA1DC_ACCEL */
 
 #ifdef HAVE_SHA1DC_ACCEL
@@ -143,3 +213,8 @@ void test_sha1dc__every_backend_agrees_with_sha1dc(void)
 {
 	RUN_OR_SKIP(every_backend_agrees_with_sha1dc);
 }
+
+void test_sha1dc__ubc_forms_agree_with_sha1dc(void)
+{
+	RUN_OR_SKIP(ubc_forms_agree_with_sha1dc);
+}
-- 
2.50.1 (Apple Git-155)



^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH 3/4] sha1dc-accel: compress with SHA-NI on x86-64
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
  2026-09-29 11:25 ` [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX Scott Chacon
  2026-09-29 11:25 ` [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check Scott Chacon
@ 2026-09-29 11:25 ` Scott Chacon
  2026-09-29 11:25 ` [PATCH 4/4] sha1dc-accel: compress with the ARMv8 SHA-1 instructions Scott Chacon
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 21+ messages in thread
From: Scott Chacon @ 2026-09-29 11:25 UTC (permalink / raw)
  To: git

With the UBC check out of the way, most of what's left is the SHA-1
compression itself, which we still do in portable C. Most x86-64 CPUs
from the last several years (Intel since Ice Lake and some Atoms, AMD
since Zen) have instructions for that: sha1rnds4 does four rounds at a
time, and sha1msg1/sha1msg2 do the message expansion.

There are two catches. The first is that the detection needs the
expanded schedule, which the instructions keep to themselves. That's
easy enough: as each group of four words is used, we store it (with a
shuffle, since SHA-NI holds them in reverse order). Words 16 to 31 come
from sha1msg1/sha1msg2, and the rest from the recurrence

  W[t] = (W[t-6] ^ W[t-16] ^ W[t-28] ^ W[t-32]) <<< 2

in plain SSE, in which no word of a group depends on another. That's
what the crate does, too; it notes that sha1msg2 is microcoded on
Sapphire Rapids, where this runs about 10% faster.

The second catch is that the recompression starts from the working
state before step 58 or step 65, and the instructions only let us see
the state between groups of four steps. But the state at step 60 is
just two steps away from 58, and 64 is one step from 65. So we save
those two for every block, and only for a block that the filter flags
(about one in twenty) do we walk them the rest of the way.

While we're here, we can do the recompression of those flagged blocks
with the SHA-1 instructions, too. The portable code runs the partner
block backwards from its state at 58 or 65 to find the chaining value
it started from, then forwards to find the one it ends on, and compares
that with ours. But the instructions only go forwards. So instead we run
the partner forwards from step 60 or 64 to step 80, which tells us the
only input chaining value that could produce our output (it's the output
minus the state at 80, since the feed-forward is just addition). Then we
run forwards from that input to step 60 or 64 and see if we arrive back
where we started. Since each step is a bijection, that's the same test.

The exception is reduced-round detection, which also wants to know the
partner's input chaining value for its own sake. Git never turns that
on, so it just keeps using the portable recompression.

We check for SHA-NI (plus SSSE3 and SSE4.1, which the code also needs)
with cpuid when we pick a backend, and compile with target attributes,
so there's nothing to configure. That gives two new backends in front of
the others, shani+avx2 and shani+sse2, which differ only in the UBC
check.

The unit test gains a check of the compression against SHA-1 steps
written out in the test itself, and of the recompression, which must
accept the partner's real output and reject any other.

With the usual 256MB of random data (the Xeon picks shani+avx2):

  Benchmark 1: test-tool.old sha1
    Time (mean ± σ):     590.3 ms ±  83.8 ms    [User: 541.0 ms, System: 42.6 ms]
    Range (min … max):   440.2 ms … 750.7 ms    30 runs

  Benchmark 2: test-tool.new sha1
    Time (mean ± σ):     285.2 ms ±  26.0 ms    [User: 238.9 ms, System: 42.1 ms]
    Range (min … max):   223.9 ms … 341.5 ms    30 runs

  Summary
    test-tool.new sha1 ran
      2.07 ± 0.35 times faster than test-tool.old sha1

That makes the whole series so far 2.70 ± 0.35 times faster than plain
sha1dc/ on this machine. Or, for index-pack on the 208MB pack of a clone
of git.git, single-threaded:

  Benchmark 1: git.old index-pack --threads=1
    Time (mean ± σ):     24.301 s ±  1.102 s    [User: 23.739 s, System: 0.219 s]
    Range (min … max):   23.557 s … 26.215 s    5 runs

  Benchmark 2: git.new index-pack --threads=1
    Time (mean ± σ):     12.689 s ±  0.385 s    [User: 12.293 s, System: 0.197 s]
    Range (min … max):   12.274 s … 13.226 s    5 runs

  Summary
    git.new index-pack --threads=1 ran
      1.92 ± 0.10 times faster than git.old index-pack --threads=1

and with 4 threads:

  Benchmark 1: git.old index-pack --threads=4
    Time (mean ± σ):     10.450 s ±  0.445 s    [User: 27.147 s, System: 0.497 s]
    Range (min … max):    9.853 s … 10.782 s    5 runs

  Benchmark 2: git.new index-pack --threads=4
    Time (mean ± σ):      5.924 s ±  0.214 s    [User: 12.693 s, System: 0.527 s]
    Range (min … max):    5.591 s …  6.185 s    5 runs

  Summary
    git.new index-pack --threads=4 ran
      1.76 ± 0.10 times faster than git.old index-pack --threads=4

Hashing in-process, shani+avx2 runs at 900 to 1050 MiB/s for messages
of 1KiB and up, against 1130 to 1230 MiB/s for OpenSSL's SHA-1 (which
doesn't detect collisions at all) and 420 to 450 MiB/s for sha1dc/. The
full test suite passes on this machine, as do t0013 and the unit tests
for each of its five backends.

Signed-off-by: Scott Chacon <scott@gitbutler.net>
Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
---
 Makefile                |   2 +-
 sha1dc-accel/internal.h |  28 +++++
 sha1dc-accel/sha1.c     | 105 ++++++++++++++-----
 sha1dc-accel/x86.c      | 224 +++++++++++++++++++++++++++++++++++++++-
 t/unit-tests/u-sha1dc.c | 141 ++++++++++++++++++++++++-
 5 files changed, 471 insertions(+), 29 deletions(-)

diff --git a/Makefile b/Makefile
index 3ad8a7fc92..8181ea5692 100644
--- a/Makefile
+++ b/Makefile
@@ -569,7 +569,7 @@ include shared.mak
 #
 # Unless DC_SHA1_EXTERNAL is defined, the built-in code is driven by the
 # faster implementation in sha1dc-accel/, which gives the same results
-# using the CPU's vector units where it has them.
+# using the CPU's SHA-1 instructions and vector units where it has them.
 # Define DC_SHA1_NO_ACCEL to use the sha1collisiondetection code alone.
 #
 # === SHA-256 backend ===
diff --git a/sha1dc-accel/internal.h b/sha1dc-accel/internal.h
index 03427224da..a27fda19d1 100644
--- a/sha1dc-accel/internal.h
+++ b/sha1dc-accel/internal.h
@@ -19,9 +19,11 @@
 # if defined(__x86_64__) && (defined(__clang__) || __GNUC__ >= 5)
 #  define SHA1DC_HAVE_SSE2 1
 #  define SHA1DC_HAVE_AVX2 1
+#  define SHA1DC_HAVE_SHANI 1
 #  include <immintrin.h>
 #  define SHA1DC_TARGET_SSE2
 #  define SHA1DC_TARGET_AVX2 __attribute__((target("avx2")))
+#  define SHA1DC_TARGET_SHANI __attribute__((target("sha,sse4.1,ssse3")))
 # elif defined(__aarch64__) && defined(__ARM_NEON)
 #  define SHA1DC_HAVE_NEON 1
 #  include <arm_neon.h>
@@ -71,4 +73,30 @@ uint32_t sha1dc_ubc_check_neon(const uint32_t w[80]);
 int sha1dc_avx2_available(void);
 #endif
 
+/*
+ * The state the partner block of a DV reaches at the group boundary next to
+ * where its recompression starts: step 60 from step 58, step 64 from step
+ * 65. `state` is this block's state at `from`. In sha1.c.
+ */
+void sha1dc_partner_boundary(enum sha1dc_from from, const uint32_t m1[80],
+			     const uint32_t dm[80], const uint32_t state[5],
+			     uint32_t out[5]);
+
+/*
+ * The hardware compression. It compresses one 64-byte block into `ihv`,
+ * writes the expanded schedule to `w`, and this block's own states at
+ * steps 60 and 64 to `at_60` and `at_64`.
+ *
+ * Its recompression answers whether a DV candidate is really an attack,
+ * running the partner block forwards from `state` (this block's state at
+ * `from`) with the SHA-1 instructions.
+ */
+#ifdef SHA1DC_HAVE_SHANI
+int sha1dc_shani_available(void);
+void sha1dc_compress_shani(uint32_t ihv[5], const unsigned char *block,
+			   uint32_t w[80], uint32_t at_60[5], uint32_t at_64[5]);
+int sha1dc_recompress_shani(enum sha1dc_from from, const uint32_t m1[80],
+			    const uint32_t dm[80], const uint32_t state[5],
+			    const uint32_t ihv_out[5]);
+#endif
 #endif /* SHA1DC_ACCEL_INTERNAL_H */
diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
index 1b3d82b4e4..2f30b207db 100644
--- a/sha1dc-accel/sha1.c
+++ b/sha1dc-accel/sha1.c
@@ -7,6 +7,14 @@
  * Shumow. It is a port to C of the approach of the "sha1dc" Rust crate by
  * Sam Reis (https://github.com/srijs/sha1dc), which gitoxide uses:
  *
+ *  - The compression runs on the CPU's SHA-1 instructions where it has them
+ *    (SHA-NI on x86-64), and
+ *    spills the expanded message schedule as it goes, which is all that
+ *    detection needs from an ordinary block. The hardware keeps no
+ *    intermediate states around, so the two states that recompression
+ *    starts from (at steps 58 and 65) are recovered from the ones at steps
+ *    60 and 64, and only for the rare block that needs them.
+ *
  *  - The unavoidable-bitconditions (UBC) filter, which rules out about 95%
  *    of blocks and is most of what detection costs, has one form per
  *    instruction set (SSE2, AVX2, NEON and portable C). For each, the
@@ -15,9 +23,12 @@
  *    the rest to a scalar tail that few blocks reach. Those choices are
  *    kept as tables, run by a short loop per form. See ubc_check.c.
  *
- * The compression is portable C, and spills the expanded message schedule
- * and the two states that recompression starts from (at steps 58 and 65)
- * as it goes.
+ *  - The recompression of a flagged block also runs on the SHA-1
+ *    instructions, forwards from the partner block's state at step 60 or
+ *    64, since that is the only direction they go.
+ *
+ * Without SHA-1 instructions the compression is portable C, and the
+ * best-suited form of the filter still applies.
  */
 
 #include "../git-compat-util.h"
@@ -180,6 +191,14 @@ static void walk(uint32_t s[5], const uint32_t *m1, const uint32_t *dm,
 	s[4] = e;
 }
 
+void sha1dc_partner_boundary(enum sha1dc_from from, const uint32_t m1[80],
+			     const uint32_t dm[80], const uint32_t state[5],
+			     uint32_t out[5])
+{
+	memcpy(out, state, 5 * sizeof(*out));
+	walk(out, m1, dm, from, from == SHA1DC_FROM_58 ? 60 : 64);
+}
+
 /*
  * The recompression, the way sha1dc/ does it: from this block's state at
  * `from`, the partner block's chaining value on the way in (`ihv_in`) and
@@ -215,31 +234,49 @@ static void compress_schedule(uint32_t ihv[5], const uint32_t w[80])
 struct backend {
 	const char *name;
 	/*
-	 * Compresses a block, spilling its schedule and the states at steps
-	 * 58 and 65.
+	 * Compresses a block and spills its schedule. If `at_60_64` is set,
+	 * it leaves the states at steps 60 and 64 where the others leave the
+	 * ones at 58 and 65.
 	 */
 	void (*compress)(uint32_t ihv[5], const unsigned char *block,
-			 uint32_t w[80], uint32_t state_58[5],
-			 uint32_t state_65[5]);
+			 uint32_t w[80], uint32_t s1[5], uint32_t s2[5]);
+	int at_60_64;
 	uint32_t (*ubc_check)(const uint32_t w[80]);
+	/* Whether a candidate is an attack; NULL for recompress_portable(). */
+	int (*recompress)(enum sha1dc_from from, const uint32_t m1[80],
+			  const uint32_t dm[80], const uint32_t state[5],
+			  const uint32_t ihv_out[5]);
 	int (*available)(void);
 };
 
+#ifdef SHA1DC_HAVE_SHANI
+static int shani_avx2_available(void)
+{
+	return sha1dc_shani_available() && sha1dc_avx2_available();
+}
+#endif
+
 /* In order of preference. */
 static const struct backend backends[] = {
+#ifdef SHA1DC_HAVE_SHANI
+	{ "shani+avx2", sha1dc_compress_shani, 1, sha1dc_ubc_check_avx2,
+	  sha1dc_recompress_shani, shani_avx2_available },
+	{ "shani+sse2", sha1dc_compress_shani, 1, sha1dc_ubc_check_sse2,
+	  sha1dc_recompress_shani, sha1dc_shani_available },
+#endif
 #ifdef SHA1DC_HAVE_AVX2
-	{ "portable+avx2", compress_portable, sha1dc_ubc_check_avx2,
+	{ "portable+avx2", compress_portable, 0, sha1dc_ubc_check_avx2, NULL,
 	  sha1dc_avx2_available },
 #endif
 #ifdef SHA1DC_HAVE_SSE2
-	{ "portable+sse2", compress_portable, sha1dc_ubc_check_sse2,
+	{ "portable+sse2", compress_portable, 0, sha1dc_ubc_check_sse2, NULL,
 	  NULL },
 #endif
 #ifdef SHA1DC_HAVE_NEON
-	{ "portable+neon", compress_portable, sha1dc_ubc_check_neon,
+	{ "portable+neon", compress_portable, 0, sha1dc_ubc_check_neon, NULL,
 	  NULL },
 #endif
-	{ "portable", compress_portable, sha1dc_ubc_check_scalar, NULL },
+	{ "portable", compress_portable, 0, sha1dc_ubc_check_scalar, NULL, NULL },
 };
 
 static int usable(const struct backend *be)
@@ -320,22 +357,27 @@ const char *const *sha1dc_accel_backends(void)
 
 /*
  * Whether any DV in `candidates` makes this block half of a collision.
- * `ihv_in` and `ihv_out` are the chaining values before and after it.
+ * `s1` and `s2` are what the compression left, `ihv_in` and `ihv_out` the
+ * chaining values before and after it.
  */
-static SHA1DC_NOINLINE int attacked(SHA1_CTX *ctx, uint32_t candidates,
-				    const uint32_t w[80],
-				    const uint32_t state_58[5],
-				    const uint32_t state_65[5],
+static SHA1DC_NOINLINE int attacked(const struct backend *be, SHA1_CTX *ctx,
+				    uint32_t candidates, const uint32_t w[80],
+				    uint32_t s1[5], uint32_t s2[5],
 				    const uint32_t ihv_in[5],
 				    const uint32_t ihv_out[5])
 {
+	const uint32_t *state_58 = s1, *state_65 = s2;
 	int i;
 
+	if (be->at_60_64) {
+		walk(s1, w, NULL, 60, 58);
+		walk(s2, w, NULL, 64, 65);
+	}
+
 	for (i = 0; sha1_dvs[i].dvType != 0; i++) {
 		const dv_info_t *dv = &sha1_dvs[i];
 		enum sha1dc_from from;
 		const uint32_t *state;
-		uint32_t ihv2_in[5], ihv2_out[5];
 
 		if (!(candidates & ((uint32_t)1 << dv->maskb)))
 			continue;
@@ -354,11 +396,23 @@ static SHA1DC_NOINLINE int attacked(SHA1_CTX *ctx, uint32_t candidates,
 			    dv->dvType, dv->dvK, dv->dvB, dv->testt);
 		}
 
-		recompress_portable(from, w, dv->dm, state, ihv2_in, ihv2_out);
-		if (!memcmp(ihv2_out, ihv_out, sizeof(ihv2_out)) ||
-		    (ctx->reduced_round_coll &&
-		     !memcmp(ihv2_in, ihv_in, sizeof(ihv2_in))))
-			return 1;
+		/*
+		 * Reduced-round collisions are recognized by the partner
+		 * block's chaining value on the way in, which only the
+		 * portable recompression computes.
+		 */
+		if (be->recompress && !ctx->reduced_round_coll) {
+			if (be->recompress(from, w, dv->dm, state, ihv_out))
+				return 1;
+		} else {
+			uint32_t ihv2_in[5], ihv2_out[5];
+
+			recompress_portable(from, w, dv->dm, state, ihv2_in, ihv2_out);
+			if (!memcmp(ihv2_out, ihv_out, sizeof(ihv2_out)) ||
+			    (ctx->reduced_round_coll &&
+			     !memcmp(ihv2_in, ihv_in, sizeof(ihv2_in))))
+				return 1;
+		}
 	}
 	return 0;
 }
@@ -366,17 +420,16 @@ static SHA1DC_NOINLINE int attacked(SHA1_CTX *ctx, uint32_t candidates,
 static inline void process(const struct backend *be, SHA1_CTX *ctx,
 			   const unsigned char *block)
 {
-	uint32_t w[80], state_58[5], state_65[5], ihv_in[5], candidates;
+	uint32_t w[80], s1[5], s2[5], ihv_in[5], candidates;
 
 	memcpy(ihv_in, ctx->ihv, sizeof(ihv_in));
-	be->compress(ctx->ihv, block, w, state_58, state_65);
+	be->compress(ctx->ihv, block, w, s1, s2);
 	if (!ctx->detect_coll)
 		return;
 
 	candidates = ctx->ubc_check ? be->ubc_check(w) : 0xFFFFFFFF;
 	if (candidates &&
-	    attacked(ctx, candidates, w, state_58, state_65, ihv_in,
-		     ctx->ihv)) {
+	    attacked(be, ctx, candidates, w, s1, s2, ihv_in, ctx->ihv)) {
 		ctx->found_collision = 1;
 		/*
 		 * Two more compressions of this block give a digest that the
diff --git a/sha1dc-accel/x86.c b/sha1dc-accel/x86.c
index 7c942fe4de..c2a0c71f17 100644
--- a/sha1dc-accel/x86.c
+++ b/sha1dc-accel/x86.c
@@ -1,5 +1,6 @@
 /*
- * The x86-64 parts of sha1.c: detecting what the CPU has.
+ * The x86-64 parts of sha1.c: detecting what the CPU has, and the SHA-1
+ * compression and recompression on SHA-NI.
  */
 
 #include "../git-compat-util.h"
@@ -35,4 +36,225 @@ int sha1dc_avx2_available(void)
 	return cpuid_7(&ebx) && (ebx & (1u << 5));
 }
 
+#ifdef SHA1DC_HAVE_SHANI
+
+int sha1dc_shani_available(void)
+{
+	unsigned int eax, ebx, ecx, edx;
+
+	if (!__get_cpuid(1, &eax, &ebx, &ecx, &edx))
+		return 0;
+	/* SSSE3 and SSE4.1 */
+	if ((ecx & (1u << 9 | 1u << 19)) != (1u << 9 | 1u << 19))
+		return 0;
+	return cpuid_7(&ebx) && (ebx & (1u << 29));
+}
+
+/*
+ * `sha1rnds4` does four steps at a time, on `abcd` held with A in the top
+ * lane; `sha1nexte` derives the fifth working word for the next group from
+ * the previous `abcd`, which is why two registers alternate as `live` and
+ * `held`. The group of schedule words for steps t..t+3 is held reversed too.
+ * Words 16 to 31 come from `sha1msg1`/`sha1msg2`, the rest from plain SSE
+ * by the recurrence W[t] = (W[t-6] ^ W[t-16] ^ W[t-28] ^ W[t-32]) <<< 2, in
+ * which no word of a group depends on another.
+ */
+
+/* Turns a group round, between step order and the order SHA-NI holds. */
+#define REVERSE 0x1B
+
+#define LOADU(p) _mm_loadu_si128((const __m128i *)(const void *)(p))
+#define STOREU(p, v) _mm_storeu_si128((__m128i *)(void *)(p), (v))
+
+/* Writes group `v` (steps t..t+3, held reversed) to w[t..t+3]. */
+#define SPILL(t, v) STOREU(w + (t), _mm_shuffle_epi32((v), REVERSE))
+
+/* Schedule words 4k..4k+3 for k from 8 on, from groups k-8, k-7, k-4, k-2, k-1. */
+SHA1DC_TARGET_SHANI
+static inline __m128i expand_rol2(__m128i v8, __m128i v7, __m128i v4,
+				  __m128i v2, __m128i v1)
+{
+	__m128i x = _mm_xor_si128(_mm_xor_si128(v8, v7), v4);
+	/* Words t-6 to t-3: the last two of group k-2, the first two of k-1. */
+	x = _mm_xor_si128(x, _mm_alignr_epi8(v2, v1, 8));
+	return _mm_or_si128(_mm_slli_epi32(x, 2), _mm_srli_epi32(x, 30));
+}
+
+/* Schedule words 4k..4k+3 for k from 4 to 7, with the SHA-NI instructions. */
+#define EXPAND_NI(a, b, c, d) \
+	_mm_sha1msg2_epu32(_mm_xor_si128(_mm_sha1msg1_epu32((a), (b)), (c)), (d))
+
+/* One group of four steps on schedule group `t / 4`, spilling it first. */
+#define ROUNDS(t, k, live, held) \
+	do { \
+		SPILL(t, v[(t) / 4]); \
+		live = _mm_sha1nexte_epu32(live, v[(t) / 4]); \
+		held = abcd; \
+		abcd = _mm_sha1rnds4_epu32(abcd, live, k); \
+	} while (0)
+
+/* The same, also storing the state [A, B, C, D, E] before it in `at`. */
+#define ROUNDS_AT(t, k, live, held, at) \
+	do { \
+		SPILL(t, v[(t) / 4]); \
+		live = _mm_sha1nexte_epu32(live, v[(t) / 4]); \
+		STOREU(at, _mm_shuffle_epi32(abcd, REVERSE)); \
+		/* `live` holds E + W[t] in its top lane. */ \
+		at[4] = (uint32_t)_mm_extract_epi32(_mm_sub_epi32(live, v[(t) / 4]), 3); \
+		held = abcd; \
+		abcd = _mm_sha1rnds4_epu32(abcd, live, k); \
+	} while (0)
+
+#define ROL2(k) v[k] = expand_rol2(v[(k) - 8], v[(k) - 7], v[(k) - 4], v[(k) - 2], v[(k) - 1])
+
+SHA1DC_TARGET_SHANI
+void sha1dc_compress_shani(uint32_t ihv[5], const unsigned char *block,
+			   uint32_t w[80], uint32_t at_60[5], uint32_t at_64[5])
+{
+	/* Big-endian words, and the four of a group reversed. */
+	const __m128i swap = _mm_set_epi64x(0x0001020304050607LL,
+					    0x08090A0B0C0D0E0FLL);
+	__m128i abcd = _mm_shuffle_epi32(LOADU(ihv), REVERSE);
+	const __m128i abcd_in = abcd;
+	const __m128i e_in = _mm_set_epi32((int)ihv[4], 0, 0, 0);
+	__m128i e0, e1, v[20];
+
+	v[0] = _mm_shuffle_epi8(LOADU(block), swap);
+	v[1] = _mm_shuffle_epi8(LOADU(block + 16), swap);
+	v[2] = _mm_shuffle_epi8(LOADU(block + 32), swap);
+	v[3] = _mm_shuffle_epi8(LOADU(block + 48), swap);
+
+	SPILL(0, v[0]);
+	e0 = _mm_add_epi32(e_in, v[0]);
+	e1 = abcd;
+	abcd = _mm_sha1rnds4_epu32(abcd, e0, 0);
+	v[4] = EXPAND_NI(v[0], v[1], v[2], v[3]);
+
+	ROUNDS(4, 0, e1, e0);
+	v[5] = EXPAND_NI(v[1], v[2], v[3], v[4]);
+	ROUNDS(8, 0, e0, e1);
+	v[6] = EXPAND_NI(v[2], v[3], v[4], v[5]);
+	ROUNDS(12, 0, e1, e0);
+	v[7] = EXPAND_NI(v[3], v[4], v[5], v[6]);
+	ROUNDS(16, 0, e0, e1);
+	ROL2(8);
+	ROUNDS(20, 1, e1, e0);
+	ROL2(9);
+	ROUNDS(24, 1, e0, e1);
+	ROL2(10);
+	ROUNDS(28, 1, e1, e0);
+	ROL2(11);
+	ROUNDS(32, 1, e0, e1);
+	ROL2(12);
+	ROUNDS(36, 1, e1, e0);
+	ROL2(13);
+	ROUNDS(40, 2, e0, e1);
+	ROL2(14);
+	ROUNDS(44, 2, e1, e0);
+	ROL2(15);
+	ROUNDS(48, 2, e0, e1);
+	ROL2(16);
+	ROUNDS(52, 2, e1, e0);
+	ROL2(17);
+	ROUNDS(56, 2, e0, e1);
+	ROL2(18);
+	ROUNDS_AT(60, 3, e1, e0, at_60);
+	ROL2(19);
+	ROUNDS_AT(64, 3, e0, e1, at_64);
+	ROUNDS(68, 3, e1, e0);
+	ROUNDS(72, 3, e0, e1);
+	ROUNDS(76, 3, e1, e0);
+
+	/* Feed-forward. */
+	e0 = _mm_sha1nexte_epu32(e0, e_in);
+	abcd = _mm_add_epi32(abcd, abcd_in);
+	STOREU(ihv, _mm_shuffle_epi32(abcd, REVERSE));
+	ihv[4] = (uint32_t)_mm_extract_epi32(e0, 3);
+}
+
+/*
+ * The partner block's schedule words for group `g`, in the order the
+ * rounds take them.
+ */
+#define WORDS(g) \
+	_mm_shuffle_epi32(_mm_xor_si128(LOADU(m1 + 4 * (g)), LOADU(dm + 4 * (g))), REVERSE)
+
+/* Four steps. The first of a run adds the fifth word itself. */
+#define GROUP_FIRST(e, g, k) \
+	do { \
+		__m128i live_ = _mm_add_epi32((e), WORDS(g)); \
+		held = abcd; \
+		abcd = _mm_sha1rnds4_epu32(abcd, live_, k); \
+	} while (0)
+#define GROUP(g, k) \
+	do { \
+		__m128i live_ = _mm_sha1nexte_epu32(held, WORDS(g)); \
+		held = abcd; \
+		abcd = _mm_sha1rnds4_epu32(abcd, live_, k); \
+	} while (0)
+
+/*
+ * Whether the partner block, whose state at `from` is `state`, ends on
+ * `ihv_out`. From its state at step 60 or 64, it runs out to step 80,
+ * which gives the chaining value an attack would have had to start from,
+ * and then in from that value, which must arrive back at the same state.
+ */
+SHA1DC_TARGET_SHANI
+int sha1dc_recompress_shani(enum sha1dc_from from, const uint32_t m1[80],
+			    const uint32_t dm[80], const uint32_t state[5],
+			    const uint32_t ihv_out[5])
+{
+	uint32_t at[5], reached[5];
+	__m128i abcd, held, e_at, e_80, e_in;
+
+	sha1dc_partner_boundary(from, m1, dm, state, at);
+	abcd = _mm_shuffle_epi32(LOADU(at), REVERSE);
+
+	/* Out to step 80, from 60 or from 64. */
+	e_at = _mm_set_epi32((int)at[4], 0, 0, 0);
+	if (from == SHA1DC_FROM_58) {
+		GROUP_FIRST(e_at, 15, 3);
+		GROUP(16, 3);
+	} else {
+		GROUP_FIRST(e_at, 16, 3);
+	}
+	GROUP(17, 3);
+	GROUP(18, 3);
+	GROUP(19, 3);
+
+	/*
+	 * The feed-forward adds the input to the state at 80, so the only
+	 * input that gives this block's output is the output less that state.
+	 */
+	e_80 = _mm_sha1nexte_epu32(held, _mm_setzero_si128());
+	abcd = _mm_sub_epi32(_mm_shuffle_epi32(LOADU(ihv_out), REVERSE), abcd);
+	e_in = _mm_sub_epi32(_mm_set_epi32((int)ihv_out[4], 0, 0, 0), e_80);
+
+	/* In from there, as far as the state the way out started from. */
+	GROUP_FIRST(e_in, 0, 0);
+	GROUP(1, 0);
+	GROUP(2, 0);
+	GROUP(3, 0);
+	GROUP(4, 0);
+	GROUP(5, 1);
+	GROUP(6, 1);
+	GROUP(7, 1);
+	GROUP(8, 1);
+	GROUP(9, 1);
+	GROUP(10, 2);
+	GROUP(11, 2);
+	GROUP(12, 2);
+	GROUP(13, 2);
+	GROUP(14, 2);
+	if (from == SHA1DC_FROM_65)
+		GROUP(15, 3);
+
+	STOREU(reached, _mm_shuffle_epi32(abcd, REVERSE));
+	reached[4] = (uint32_t)_mm_extract_epi32(
+		_mm_sha1nexte_epu32(held, _mm_setzero_si128()), 3);
+	return !memcmp(reached, at, sizeof(at));
+}
+
+#endif /* SHA1DC_HAVE_SHANI */
+
 #endif /* SHA1DC_HAVE_SSE2 */
diff --git a/t/unit-tests/u-sha1dc.c b/t/unit-tests/u-sha1dc.c
index 15e71fb023..5948466cb9 100644
--- a/t/unit-tests/u-sha1dc.c
+++ b/t/unit-tests/u-sha1dc.c
@@ -2,7 +2,8 @@
 #include "hash.h"
 
 /*
- * Tests sha1dc-accel/ against sha1dc/, which it must agree with exactly.
+ * Tests sha1dc-accel/ against sha1dc/, which it must agree with exactly,
+ * and its parts against the plain SHA-1 step function written out here.
  */
 #if defined(SHA1_DC) && !defined(DC_SHA1_EXTERNAL) && !defined(DC_SHA1_NO_ACCEL)
 #define HAVE_SHA1DC_ACCEL
@@ -196,6 +197,139 @@ static void ubc_forms_agree_with_sha1dc(void)
 	}
 }
 
+#ifdef SHA1DC_HAVE_SHANI
+
+static uint32_t rol(uint32_t x, int n)
+{
+	return (x << n) | (x >> (32 - n));
+}
+
+static uint32_t f_k(int t, uint32_t b, uint32_t c, uint32_t d)
+{
+	if (t < 20)
+		return ((b & c) | (~b & d)) + 0x5A827999;
+	if (t < 40)
+		return (b ^ c ^ d) + 0x6ED9EBA1;
+	if (t < 60)
+		return ((b & c) | (b & d) | (c & d)) + 0x8F1BBCDC;
+	return (b ^ c ^ d) + 0xCA62C1D6;
+}
+
+/* One SHA-1 step forwards, from the state before step t. */
+static void step(uint32_t s[5], int t, uint32_t w)
+{
+	uint32_t a = rol(s[0], 5) + f_k(t, s[1], s[2], s[3]) + s[4] + w;
+	s[4] = s[3];
+	s[3] = s[2];
+	s[2] = rol(s[1], 30);
+	s[1] = s[0];
+	s[0] = a;
+}
+
+/* One SHA-1 step backwards, to the state before step t. */
+static void unstep(uint32_t s[5], int t, uint32_t w)
+{
+	uint32_t a = s[1], b = rol(s[2], 2), c = s[3], d = s[4];
+	s[4] = s[0] - (rol(a, 5) + f_k(t, b, c, d) + w);
+	s[0] = a;
+	s[1] = b;
+	s[2] = c;
+	s[3] = d;
+}
+
+typedef void (*compress_fn)(uint32_t ihv[5], const unsigned char *block,
+			    uint32_t w[80], uint32_t at_60[5], uint32_t at_64[5]);
+typedef int (*recompress_fn)(enum sha1dc_from from, const uint32_t m1[80],
+			     const uint32_t dm[80], const uint32_t state[5],
+			     const uint32_t ihv_out[5]);
+
+static void check_compress(compress_fn compress)
+{
+	int i, t;
+
+	rng_seed(3);
+	for (i = 0; i < 2000; i++) {
+		unsigned char block[64];
+		uint32_t ihv[5], got[5], w[80], at_60[5], at_64[5], s[5];
+
+		for (t = 0; t < 64; t++)
+			block[t] = rng();
+		for (t = 0; t < 5; t++)
+			ihv[t] = got[t] = rng();
+		compress(got, block, w, at_60, at_64);
+
+		memcpy(s, ihv, sizeof(s));
+		for (t = 0; t < 80; t++) {
+			uint32_t want_w = t < 16 ? get_be32(block + 4 * t) :
+				rol(w[t - 3] ^ w[t - 8] ^ w[t - 14] ^ w[t - 16], 1);
+			cl_assert_equal_i(w[t], want_w);
+			if (t == 60)
+				cl_assert(!memcmp(at_60, s, sizeof(s)));
+			if (t == 64)
+				cl_assert(!memcmp(at_64, s, sizeof(s)));
+			step(s, t, w[t]);
+		}
+		for (t = 0; t < 5; t++)
+			cl_assert_equal_i(got[t], ihv[t] + s[t]);
+	}
+}
+
+/*
+ * A recompression must accept exactly the chaining value the partner
+ * block ends on, and nothing else.
+ */
+static void check_recompress(recompress_fn recompress)
+{
+	int i, t, d;
+
+	rng_seed(4);
+	for (i = 0; i < 50; i++) {
+		uint32_t w[80], state[5];
+
+		for (t = 0; t < 80; t++)
+			w[t] = rng();
+		for (t = 0; t < 5; t++)
+			state[t] = rng();
+
+		for (d = 0; sha1_dvs[d].dvType; d++) {
+			const dv_info_t *dv = &sha1_dvs[d];
+			enum sha1dc_from from = dv->testt == 58 ? SHA1DC_FROM_58 : SHA1DC_FROM_65;
+			uint32_t in[5], out[5];
+			unsigned nudge = rng();
+
+			memcpy(in, state, sizeof(in));
+			for (t = dv->testt - 1; t >= 0; t--)
+				unstep(in, t, w[t] ^ dv->dm[t]);
+			memcpy(out, state, sizeof(out));
+			for (t = dv->testt; t < 80; t++)
+				step(out, t, w[t] ^ dv->dm[t]);
+			for (t = 0; t < 5; t++)
+				out[t] += in[t];
+
+			cl_assert(recompress(from, w, dv->dm, state, out));
+			out[nudge % 5] ^= 1u << (nudge / 5 % 32);
+			cl_assert(!recompress(from, w, dv->dm, state, out));
+		}
+	}
+}
+
+#endif
+
+static void hardware_compression(void)
+{
+	int tested = 0;
+
+#ifdef SHA1DC_HAVE_SHANI
+	if (sha1dc_shani_available()) {
+		check_compress(sha1dc_compress_shani);
+		check_recompress(sha1dc_recompress_shani);
+		tested = 1;
+	}
+#endif
+	if (!tested)
+		cl_skip();
+}
+
 #endif /* HAVE_SHA1DC_ACCEL */
 
 #ifdef HAVE_SHA1DC_ACCEL
@@ -218,3 +352,8 @@ void test_sha1dc__ubc_forms_agree_with_sha1dc(void)
 {
 	RUN_OR_SKIP(ubc_forms_agree_with_sha1dc);
 }
+
+void test_sha1dc__hardware_compression(void)
+{
+	RUN_OR_SKIP(hardware_compression);
+}
-- 
2.50.1 (Apple Git-155)



^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH 4/4] sha1dc-accel: compress with the ARMv8 SHA-1 instructions
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
                   ` (2 preceding siblings ...)
  2026-09-29 11:25 ` [PATCH 3/4] sha1dc-accel: compress with SHA-NI on x86-64 Scott Chacon
@ 2026-09-29 11:25 ` Scott Chacon
  2026-10-07 12:17 ` [PATCH 0/4] faster SHA-1 collision detection Johannes Schindelin
  2026-10-07 17:23 ` Junio C Hamano
  5 siblings, 0 replies; 21+ messages in thread
From: Scott Chacon @ 2026-09-29 11:25 UTC (permalink / raw)
  To: git

Most arm64 CPUs, including Apple's and the Graviton line, have SHA-1
instructions, too, and we can use them the same way as SHA-NI in the
previous patch. It's a little simpler here: the instructions keep the
working state in the natural order, and vsha1h gives us the fifth word
directly, so there's no shuffling to be done.

Knowing whether we can use them is a little trickier:

  - If the compiler targets them already (__ARM_FEATURE_SHA2 or
    __ARM_FEATURE_CRYPTO, which is the case with Apple's clang on
    macOS), we just use them.

  - Otherwise we build the two functions that need them with a target
    attribute, and check at run time. On Linux that's getauxval(); on
    other platforms we don't know how to ask yet, and so don't use them.

  - Older compilers only declare the intrinsics when the whole file
    targets the instructions, so a target attribute doesn't help there.
    We only try the attribute with GCC 9+ and clang 17+.

The one new backend, armv8+neon, goes to the front of the list.

On an Apple M5 Max (macOS, Apple clang), which picks armv8+neon, hashing
1GiB of random data with "test-tool sha1" (median of 9 runs) goes:

  sha1dc, before this series        1.46s   (~700 MiB/s)
  block loop, after 1/4             1.55s
  portable+neon, after 2/4          1.22s
  armv8+neon, this patch            0.51s   (~2 GiB/s)

which is 2.85x overall. For index-pack of the 319MB pack of a clone of
git.git, it goes from 16.1s to 8.7s single-threaded (1.86x), and from
6.7s to 5.9s with all 18 threads. The full test suite passes there, too.
For comparison, the crate reports 81% of plain SHA-1's throughput on an
Apple M4 and 80% on a Graviton4, up from 29% [1].

The run-time check on Linux is only tested under qemu-user so far,
built with GCC 13 and clang 18, in a harness that compares every backend
and UBC form against sha1dc/ (the same comparisons as the unit tests).
I'd be happy to see numbers from anybody with a Graviton or similar
machine.

[1] https://sam.dev/blog/faster-sha1-collision-detection

Signed-off-by: Scott Chacon <scott@gitbutler.net>
Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
---
 Makefile                            |   1 +
 contrib/buildsystems/CMakeLists.txt |   2 +-
 meson.build                         |   1 +
 sha1dc-accel/arm.c                  | 274 ++++++++++++++++++++++++++++
 sha1dc-accel/internal.h             |  28 ++-
 sha1dc-accel/sha1.c                 |   6 +-
 t/unit-tests/u-sha1dc.c             |   9 +-
 7 files changed, 316 insertions(+), 5 deletions(-)
 create mode 100644 sha1dc-accel/arm.c

diff --git a/Makefile b/Makefile
index 8181ea5692..731a5f60e8 100644
--- a/Makefile
+++ b/Makefile
@@ -2177,6 +2177,7 @@ endif
 ifdef DC_SHA1_NO_ACCEL
 	BASIC_CFLAGS += -DDC_SHA1_NO_ACCEL
 else
+	LIB_OBJS += sha1dc-accel/arm.o
 	LIB_OBJS += sha1dc-accel/sha1.o
 	LIB_OBJS += sha1dc-accel/ubc_check.o
 	LIB_OBJS += sha1dc-accel/x86.o
diff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt
index 67b96d601b..37ebb67a8d 100644
--- a/contrib/buildsystems/CMakeLists.txt
+++ b/contrib/buildsystems/CMakeLists.txt
@@ -218,7 +218,7 @@ add_compile_definitions(NO_OPENSSL SHA1_DC SHA1DC_NO_STANDARD_INCLUDES
 			SHA1DC_INIT_SAFE_HASH_DEFAULT=0
 			SHA1DC_CUSTOM_INCLUDE_SHA1_C="git-compat-util.h"
 			SHA1DC_CUSTOM_INCLUDE_UBC_CHECK_C="git-compat-util.h" )
-list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c sha1dc-accel/sha1.c sha1dc-accel/ubc_check.c sha1dc-accel/x86.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
+list(APPEND compat_SOURCES sha1dc_git.c sha1dc/sha1.c sha1dc/ubc_check.c sha1dc-accel/arm.c sha1dc-accel/sha1.c sha1dc-accel/ubc_check.c sha1dc-accel/x86.c block-sha1/sha1.c sha256/block/sha256.c compat/qsort_s.c)
 
 
 add_compile_definitions(PAGER_ENV="LESS=FRX LV=-c"
diff --git a/meson.build b/meson.build
index 47a60526e9..7aed4db1ba 100644
--- a/meson.build
+++ b/meson.build
@@ -1631,6 +1631,7 @@ if sha1_backend == 'sha1dc'
     'sha1dc_git.c',
     'sha1dc/sha1.c',
     'sha1dc/ubc_check.c',
+    'sha1dc-accel/arm.c',
     'sha1dc-accel/sha1.c',
     'sha1dc-accel/ubc_check.c',
     'sha1dc-accel/x86.c',
diff --git a/sha1dc-accel/arm.c b/sha1dc-accel/arm.c
new file mode 100644
index 0000000000..0d18354e56
--- /dev/null
+++ b/sha1dc-accel/arm.c
@@ -0,0 +1,274 @@
+/*
+ * The SHA-1 compression and recompression on the ARMv8 SHA-1 instructions,
+ * for sha1.c. `vsha1{c,p,m}q_u32` do four steps at a time on `abcd`, with
+ * the fifth working word carried apart and derived by `vsha1h_u32`.
+ */
+
+#include "../git-compat-util.h"
+#include "internal.h"
+
+#ifdef SHA1DC_HAVE_ARMV8
+
+#if !defined(__ARM_FEATURE_SHA2) && !defined(__ARM_FEATURE_CRYPTO) && \
+	defined(__linux__)
+#include <sys/auxv.h>
+#ifndef HWCAP_SHA1
+#define HWCAP_SHA1 (1 << 5)
+#endif
+#endif
+
+int sha1dc_armv8_available(void)
+{
+#if defined(__ARM_FEATURE_SHA2) || defined(__ARM_FEATURE_CRYPTO) || \
+	defined(__APPLE__)
+	return 1;
+#elif defined(__linux__)
+	return !!(getauxval(AT_HWCAP) & HWCAP_SHA1);
+#else
+	return 0;
+#endif
+}
+
+#define K0 0x5A827999
+#define K1 0x6ED9EBA1
+#define K2 0x8F1BBCDC
+#define K3 0xCA62C1D6
+
+/* Big-endian message words t..t+3. */
+#define MSG(t) vreinterpretq_u32_u8(vrev32q_u8(vld1q_u8(block + 4 * (t))))
+
+/*
+ * Four steps of a compression that keeps the schedule four groups ahead:
+ * `f` the round instruction, `e` the fifth word in, `next` where the one
+ * for the next group goes, `wk` the schedule words plus the constant.
+ */
+#define QUAD(f, e, next, wk) \
+	do { \
+		next = vsha1h_u32(vgetq_lane_u32(abcd, 0)); \
+		abcd = f(abcd, e, wk); \
+	} while (0)
+
+SHA1DC_TARGET_ARMV8
+void sha1dc_compress_armv8(uint32_t ihv[5], const unsigned char *block,
+			   uint32_t w[80], uint32_t at_60[5], uint32_t at_64[5])
+{
+	uint32x4_t abcd = vld1q_u32(ihv);
+	const uint32x4_t abcd_in = abcd;
+	uint32_t e0 = ihv[4], e1;
+	const uint32_t e_in = e0;
+	const uint32x4_t k0 = vdupq_n_u32(K0), k1 = vdupq_n_u32(K1);
+	const uint32x4_t k2 = vdupq_n_u32(K2), k3 = vdupq_n_u32(K3);
+	uint32x4_t msg0 = MSG(0), msg1 = MSG(4), msg2 = MSG(8), msg3 = MSG(12);
+	uint32x4_t tmp0, tmp1;
+
+	vst1q_u32(w + 0, msg0);
+	vst1q_u32(w + 4, msg1);
+	vst1q_u32(w + 8, msg2);
+	vst1q_u32(w + 12, msg3);
+
+	tmp0 = vaddq_u32(msg0, k0);
+	tmp1 = vaddq_u32(msg1, k0);
+
+	/* Steps 0-3 */
+	QUAD(vsha1cq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg2, k0);
+	msg0 = vsha1su0q_u32(msg0, msg1, msg2);
+
+	/* Steps 4-7 */
+	QUAD(vsha1cq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg3, k0);
+	msg0 = vsha1su1q_u32(msg0, msg3);
+	vst1q_u32(w + 16, msg0);
+	msg1 = vsha1su0q_u32(msg1, msg2, msg3);
+
+	/* Steps 8-11 */
+	QUAD(vsha1cq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg0, k0);
+	msg1 = vsha1su1q_u32(msg1, msg0);
+	vst1q_u32(w + 20, msg1);
+	msg2 = vsha1su0q_u32(msg2, msg3, msg0);
+
+	/* Steps 12-15 */
+	QUAD(vsha1cq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg1, k1);
+	msg2 = vsha1su1q_u32(msg2, msg1);
+	vst1q_u32(w + 24, msg2);
+	msg3 = vsha1su0q_u32(msg3, msg0, msg1);
+
+	/* Steps 16-19 */
+	QUAD(vsha1cq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg2, k1);
+	msg3 = vsha1su1q_u32(msg3, msg2);
+	vst1q_u32(w + 28, msg3);
+	msg0 = vsha1su0q_u32(msg0, msg1, msg2);
+
+	/* Steps 20-23 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg3, k1);
+	msg0 = vsha1su1q_u32(msg0, msg3);
+	vst1q_u32(w + 32, msg0);
+	msg1 = vsha1su0q_u32(msg1, msg2, msg3);
+
+	/* Steps 24-27 */
+	QUAD(vsha1pq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg0, k1);
+	msg1 = vsha1su1q_u32(msg1, msg0);
+	vst1q_u32(w + 36, msg1);
+	msg2 = vsha1su0q_u32(msg2, msg3, msg0);
+
+	/* Steps 28-31 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg1, k1);
+	msg2 = vsha1su1q_u32(msg2, msg1);
+	vst1q_u32(w + 40, msg2);
+	msg3 = vsha1su0q_u32(msg3, msg0, msg1);
+
+	/* Steps 32-35 */
+	QUAD(vsha1pq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg2, k2);
+	msg3 = vsha1su1q_u32(msg3, msg2);
+	vst1q_u32(w + 44, msg3);
+	msg0 = vsha1su0q_u32(msg0, msg1, msg2);
+
+	/* Steps 36-39 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg3, k2);
+	msg0 = vsha1su1q_u32(msg0, msg3);
+	vst1q_u32(w + 48, msg0);
+	msg1 = vsha1su0q_u32(msg1, msg2, msg3);
+
+	/* Steps 40-43 */
+	QUAD(vsha1mq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg0, k2);
+	msg1 = vsha1su1q_u32(msg1, msg0);
+	vst1q_u32(w + 52, msg1);
+	msg2 = vsha1su0q_u32(msg2, msg3, msg0);
+
+	/* Steps 44-47 */
+	QUAD(vsha1mq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg1, k2);
+	msg2 = vsha1su1q_u32(msg2, msg1);
+	vst1q_u32(w + 56, msg2);
+	msg3 = vsha1su0q_u32(msg3, msg0, msg1);
+
+	/* Steps 48-51 */
+	QUAD(vsha1mq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg2, k2);
+	msg3 = vsha1su1q_u32(msg3, msg2);
+	vst1q_u32(w + 60, msg3);
+	msg0 = vsha1su0q_u32(msg0, msg1, msg2);
+
+	/* Steps 52-55 */
+	QUAD(vsha1mq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg3, k3);
+	msg0 = vsha1su1q_u32(msg0, msg3);
+	vst1q_u32(w + 64, msg0);
+	msg1 = vsha1su0q_u32(msg1, msg2, msg3);
+
+	/* Steps 56-59 */
+	QUAD(vsha1mq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg0, k3);
+	msg1 = vsha1su1q_u32(msg1, msg0);
+	vst1q_u32(w + 68, msg1);
+	msg2 = vsha1su0q_u32(msg2, msg3, msg0);
+
+	/* The state at step 60, for recompression to start from. */
+	vst1q_u32(at_60, abcd);
+	at_60[4] = e1;
+
+	/* Steps 60-63 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg1, k3);
+	msg2 = vsha1su1q_u32(msg2, msg1);
+	vst1q_u32(w + 72, msg2);
+	msg3 = vsha1su0q_u32(msg3, msg0, msg1);
+
+	/* And at step 64. */
+	vst1q_u32(at_64, abcd);
+	at_64[4] = e0;
+
+	/* Steps 64-67 */
+	QUAD(vsha1pq_u32, e0, e1, tmp0);
+	tmp0 = vaddq_u32(msg2, k3);
+	msg3 = vsha1su1q_u32(msg3, msg2);
+	vst1q_u32(w + 76, msg3);
+
+	/* Steps 68-71 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+	tmp1 = vaddq_u32(msg3, k3);
+
+	/* Steps 72-75 */
+	QUAD(vsha1pq_u32, e0, e1, tmp0);
+
+	/* Steps 76-79 */
+	QUAD(vsha1pq_u32, e1, e0, tmp1);
+
+	vst1q_u32(ihv, vaddq_u32(abcd_in, abcd));
+	ihv[4] = e0 + e_in;
+}
+
+/* Four steps of the partner block, on schedule group `g`. */
+#define GROUP(f, k, g) \
+	do { \
+		uint32x4_t wk_ = vaddq_u32(veorq_u32(vld1q_u32(m1 + 4 * (g)), \
+						     vld1q_u32(dm + 4 * (g))), \
+					   vdupq_n_u32(k)); \
+		uint32_t next_ = vsha1h_u32(vgetq_lane_u32(abcd, 0)); \
+		abcd = f(abcd, e, wk_); \
+		e = next_; \
+	} while (0)
+
+/*
+ * Whether the partner block, whose state at `from` is `state`, ends on
+ * `ihv_out`: out from its state at step 60 or 64 to step 80, then in from
+ * the chaining value that implies, which must arrive back at that state.
+ */
+SHA1DC_TARGET_ARMV8
+int sha1dc_recompress_armv8(enum sha1dc_from from, const uint32_t m1[80],
+			    const uint32_t dm[80], const uint32_t state[5],
+			    const uint32_t ihv_out[5])
+{
+	uint32_t at[5], reached[5], e;
+	uint32x4_t abcd;
+
+	sha1dc_partner_boundary(from, m1, dm, state, at);
+	abcd = vld1q_u32(at);
+	e = at[4];
+
+	/* Out to step 80, from 60 or from 64. */
+	if (from == SHA1DC_FROM_58)
+		GROUP(vsha1pq_u32, K3, 15);
+	GROUP(vsha1pq_u32, K3, 16);
+	GROUP(vsha1pq_u32, K3, 17);
+	GROUP(vsha1pq_u32, K3, 18);
+	GROUP(vsha1pq_u32, K3, 19);
+
+	/* The only input that gives this output is the output less the state. */
+	abcd = vsubq_u32(vld1q_u32(ihv_out), abcd);
+	e = ihv_out[4] - e;
+
+	/* In from there, as far as the state the way out started from. */
+	GROUP(vsha1cq_u32, K0, 0);
+	GROUP(vsha1cq_u32, K0, 1);
+	GROUP(vsha1cq_u32, K0, 2);
+	GROUP(vsha1cq_u32, K0, 3);
+	GROUP(vsha1cq_u32, K0, 4);
+	GROUP(vsha1pq_u32, K1, 5);
+	GROUP(vsha1pq_u32, K1, 6);
+	GROUP(vsha1pq_u32, K1, 7);
+	GROUP(vsha1pq_u32, K1, 8);
+	GROUP(vsha1pq_u32, K1, 9);
+	GROUP(vsha1mq_u32, K2, 10);
+	GROUP(vsha1mq_u32, K2, 11);
+	GROUP(vsha1mq_u32, K2, 12);
+	GROUP(vsha1mq_u32, K2, 13);
+	GROUP(vsha1mq_u32, K2, 14);
+	if (from == SHA1DC_FROM_65)
+		GROUP(vsha1pq_u32, K3, 15);
+
+	vst1q_u32(reached, abcd);
+	reached[4] = e;
+	return !memcmp(reached, at, sizeof(at));
+}
+
+#endif /* SHA1DC_HAVE_ARMV8 */
diff --git a/sha1dc-accel/internal.h b/sha1dc-accel/internal.h
index a27fda19d1..a50b294b96 100644
--- a/sha1dc-accel/internal.h
+++ b/sha1dc-accel/internal.h
@@ -27,6 +27,21 @@
 # elif defined(__aarch64__) && defined(__ARM_NEON)
 #  define SHA1DC_HAVE_NEON 1
 #  include <arm_neon.h>
+/*
+ * The SHA-1 instructions, where the compiler targets them already or can
+ * enable them for one function: older compilers declare their intrinsics
+ * only when the whole file targets them.
+ */
+#  if defined(__ARM_FEATURE_SHA2) || defined(__ARM_FEATURE_CRYPTO)
+#   define SHA1DC_HAVE_ARMV8 1
+#   define SHA1DC_TARGET_ARMV8
+#  elif defined(__clang__) && __clang_major__ >= 17
+#   define SHA1DC_HAVE_ARMV8 1
+#   define SHA1DC_TARGET_ARMV8 __attribute__((target("sha2")))
+#  elif !defined(__clang__) && __GNUC__ >= 9
+#   define SHA1DC_HAVE_ARMV8 1
+#   define SHA1DC_TARGET_ARMV8 __attribute__((target("+crypto")))
+#  endif
 # endif
 #endif
 
@@ -83,11 +98,11 @@ void sha1dc_partner_boundary(enum sha1dc_from from, const uint32_t m1[80],
 			     uint32_t out[5]);
 
 /*
- * The hardware compression. It compresses one 64-byte block into `ihv`,
+ * The hardware compressions. Each compresses one 64-byte block into `ihv`,
  * writes the expanded schedule to `w`, and this block's own states at
  * steps 60 and 64 to `at_60` and `at_64`.
  *
- * Its recompression answers whether a DV candidate is really an attack,
+ * Their recompressions answer whether a DV candidate is really an attack,
  * running the partner block forwards from `state` (this block's state at
  * `from`) with the SHA-1 instructions.
  */
@@ -99,4 +114,13 @@ int sha1dc_recompress_shani(enum sha1dc_from from, const uint32_t m1[80],
 			    const uint32_t dm[80], const uint32_t state[5],
 			    const uint32_t ihv_out[5]);
 #endif
+#ifdef SHA1DC_HAVE_ARMV8
+int sha1dc_armv8_available(void);
+void sha1dc_compress_armv8(uint32_t ihv[5], const unsigned char *block,
+			   uint32_t w[80], uint32_t at_60[5], uint32_t at_64[5]);
+int sha1dc_recompress_armv8(enum sha1dc_from from, const uint32_t m1[80],
+			    const uint32_t dm[80], const uint32_t state[5],
+			    const uint32_t ihv_out[5]);
+#endif
+
 #endif /* SHA1DC_ACCEL_INTERNAL_H */
diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
index 2f30b207db..1f1133169b 100644
--- a/sha1dc-accel/sha1.c
+++ b/sha1dc-accel/sha1.c
@@ -8,7 +8,7 @@
  * Sam Reis (https://github.com/srijs/sha1dc), which gitoxide uses:
  *
  *  - The compression runs on the CPU's SHA-1 instructions where it has them
- *    (SHA-NI on x86-64), and
+ *    (SHA-NI on x86-64, the ARMv8 cryptography extension on arm64), and
  *    spills the expanded message schedule as it goes, which is all that
  *    detection needs from an ordinary block. The hardware keeps no
  *    intermediate states around, so the two states that recompression
@@ -264,6 +264,10 @@ static const struct backend backends[] = {
 	{ "shani+sse2", sha1dc_compress_shani, 1, sha1dc_ubc_check_sse2,
 	  sha1dc_recompress_shani, sha1dc_shani_available },
 #endif
+#ifdef SHA1DC_HAVE_ARMV8
+	{ "armv8+neon", sha1dc_compress_armv8, 1, sha1dc_ubc_check_neon,
+	  sha1dc_recompress_armv8, sha1dc_armv8_available },
+#endif
 #ifdef SHA1DC_HAVE_AVX2
 	{ "portable+avx2", compress_portable, 0, sha1dc_ubc_check_avx2, NULL,
 	  sha1dc_avx2_available },
diff --git a/t/unit-tests/u-sha1dc.c b/t/unit-tests/u-sha1dc.c
index 5948466cb9..f629f59d1d 100644
--- a/t/unit-tests/u-sha1dc.c
+++ b/t/unit-tests/u-sha1dc.c
@@ -197,7 +197,7 @@ static void ubc_forms_agree_with_sha1dc(void)
 	}
 }
 
-#ifdef SHA1DC_HAVE_SHANI
+#if defined(SHA1DC_HAVE_SHANI) || defined(SHA1DC_HAVE_ARMV8)
 
 static uint32_t rol(uint32_t x, int n)
 {
@@ -325,6 +325,13 @@ static void hardware_compression(void)
 		check_recompress(sha1dc_recompress_shani);
 		tested = 1;
 	}
+#endif
+#ifdef SHA1DC_HAVE_ARMV8
+	if (sha1dc_armv8_available()) {
+		check_compress(sha1dc_compress_armv8);
+		check_recompress(sha1dc_recompress_armv8);
+		tested = 1;
+	}
 #endif
 	if (!tested)
 		cl_skip();
-- 
2.50.1 (Apple Git-155)



^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX
  2026-09-29 11:25 ` [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX Scott Chacon
@ 2026-10-07 12:17   ` Johannes Schindelin
  0 siblings, 0 replies; 21+ messages in thread
From: Johannes Schindelin @ 2026-10-07 12:17 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

Hi Scott,

On Tue, 29 Sep 2026, Scott Chacon wrote:

> diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
> new file mode 100644
> index 0000000000..fd75289997
> --- /dev/null
> +++ b/sha1dc-accel/sha1.c
> @@ -0,0 +1,434 @@
> +/*
> + * SHA-1 with collision detection.
> + *
> + * This computes exactly what sha1dc/ computes: the SHA-1 digest of the
> + * input, and whether any block of it looks like one half of a collision
> + * made by one of the 32 known disturbance vectors (DVs) of Stevens and
> + * Shumow. It works on sha1dc's SHA1_CTX, and uses its table of DVs, but
> + * has its own block loop, which gives the following patches room to
> + * follow the approach of the "sha1dc" Rust crate by Sam Reis
> + * (https://github.com/srijs/sha1dc), which gitoxide uses.
> + *
> + * Each block is compressed by a "backend", which also spills the expanded
> + * message schedule, and the two intermediate states that recompression
> + * starts from (at steps 58 and 65). The unavoidable-bitconditions (UBC)
> + * filter then rules out about 95% of blocks; the rest are recompressed,
> + * once for each DV the filter could not rule out.
> + */
> +
> +#include "../git-compat-util.h"
> +#include "../sha1dc_git.h"
> +#if defined(DC_SHA1_SUBMODULE)
> +#include "../sha1collisiondetection/lib/ubc_check.h"
> +#else
> +#include "../sha1dc/ubc_check.h"
> +#endif
> +#include "sha1.h"
> +#include "internal.h"
> +
> +#define ROL(x, n) (((x) << (n)) | ((x) >> (32 - (n))))
> +
> +#define F_CH(b, c, d) ((d) ^ ((b) & ((c) ^ (d))))
> +#define F_PARITY(b, c, d) ((b) ^ (c) ^ (d))
> +#define F_MAJ(b, c, d) (((b) & (c)) | ((d) & ((b) | (c))))
> +
> +#define K0 0x5A827999
> +#define K1 0x6ED9EBA1
> +#define K2 0x8F1BBCDC
> +#define K3 0xCA62C1D6
> +

The following lines, including the definition of `compress_portable()`,
duplicate the functionality implemented in `sha1_compression_states()` in
sha1dc/. Maybe we could use that latter function here, too, to make the
code DRYer?

> +/*
> + * One step, on names that rotate: after it, (e, a, b, c, d) are the new
> + * (a, b, c, d, e).
> + */
> +#define STEP(f, k, a, b, c, d, e, x) \
> +	do { \
> +		e += ROL(a, 5) + f(b, c, d) + (k) + (x); \
> +		b = ROL(b, 30); \
> +	} while (0)
> +
> +#define LOAD(t) (w[t] = get_be32(block + 4 * (t)))
> +#define EXPAND(t) (w[t] = ROL(w[(t) - 3] ^ w[(t) - 8] ^ w[(t) - 14] ^ w[(t) - 16], 1))
> +
> +#define FIVE_LOAD(f, k, a, b, c, d, e, t) \
> +	do { \
> +		STEP(f, k, a, b, c, d, e, LOAD(t)); \
> +		STEP(f, k, e, a, b, c, d, LOAD((t) + 1)); \
> +		STEP(f, k, d, e, a, b, c, LOAD((t) + 2)); \
> +		STEP(f, k, c, d, e, a, b, LOAD((t) + 3)); \
> +		STEP(f, k, b, c, d, e, a, LOAD((t) + 4)); \
> +	} while (0)
> +
> +#define FIVE_EXPAND(f, k, a, b, c, d, e, t) \
> +	do { \
> +		STEP(f, k, a, b, c, d, e, EXPAND(t)); \
> +		STEP(f, k, e, a, b, c, d, EXPAND((t) + 1)); \
> +		STEP(f, k, d, e, a, b, c, EXPAND((t) + 2)); \
> +		STEP(f, k, c, d, e, a, b, EXPAND((t) + 3)); \
> +		STEP(f, k, b, c, d, e, a, EXPAND((t) + 4)); \
> +	} while (0)
> +
> +/*
> + * The portable compression. Like the hardware ones it spills the schedule,
> + * but it writes the states at steps 58 and 65 directly, on the way past.
> + */
> +static void compress_portable(uint32_t ihv[5], const unsigned char *block,
> +			      uint32_t w[80], uint32_t state_58[5],
> +			      uint32_t state_65[5])
> +{
> +	uint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];
> +
> +	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 0);
> +	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 5);
> +	FIVE_LOAD(F_CH, K0, a, b, c, d, e, 10);
> +	STEP(F_CH, K0, a, b, c, d, e, LOAD(15));
> +	STEP(F_CH, K0, e, a, b, c, d, EXPAND(16));
> +	STEP(F_CH, K0, d, e, a, b, c, EXPAND(17));
> +	STEP(F_CH, K0, c, d, e, a, b, EXPAND(18));
> +	STEP(F_CH, K0, b, c, d, e, a, EXPAND(19));
> +
> +	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 20);
> +	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 25);
> +	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 30);
> +	FIVE_EXPAND(F_PARITY, K1, a, b, c, d, e, 35);
> +
> +	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 40);
> +	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 45);
> +	FIVE_EXPAND(F_MAJ, K2, a, b, c, d, e, 50);
> +	STEP(F_MAJ, K2, a, b, c, d, e, EXPAND(55));
> +	STEP(F_MAJ, K2, e, a, b, c, d, EXPAND(56));
> +	STEP(F_MAJ, K2, d, e, a, b, c, EXPAND(57));
> +	state_58[0] = c;
> +	state_58[1] = d;
> +	state_58[2] = e;
> +	state_58[3] = a;
> +	state_58[4] = b;
> +	STEP(F_MAJ, K2, c, d, e, a, b, EXPAND(58));
> +	STEP(F_MAJ, K2, b, c, d, e, a, EXPAND(59));
> +
> +	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 60);
> +	state_65[0] = a;
> +	state_65[1] = b;
> +	state_65[2] = c;
> +	state_65[3] = d;
> +	state_65[4] = e;
> +	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 65);
> +	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 70);
> +	FIVE_EXPAND(F_PARITY, K3, a, b, c, d, e, 75);
> +
> +	ihv[0] += a;
> +	ihv[1] += b;
> +	ihv[2] += c;
> +	ihv[3] += d;
> +	ihv[4] += e;
> +}

Ciao,
Johannes

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check
  2026-09-29 11:25 ` [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check Scott Chacon
@ 2026-10-07 12:17   ` Johannes Schindelin
  2026-10-07 21:32     ` Junio C Hamano
  0 siblings, 1 reply; 21+ messages in thread
From: Johannes Schindelin @ 2026-10-07 12:17 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

Hi Scott,

On Tue, 29 Sep 2026, Scott Chacon wrote:

>  sha1dc-accel/ubc_check.c            | 1789 +++++++++++++++++++++++++++

This is quite large. And of course it's essentially a machine-assisted
translation of the Rust code, which itself is the output of the solver.

Assuming that we will not only want to be able to confirm the translation
easily, but also be able to adapt to future improvements in the Rust code
(e.g. if Sam finds another neat trick to dismiss even more candidates even
earlier), here is a Perl script to do precisely that:

-- snip --
#!/usr/bin/perl
# Convert the generated src/ubc_check/{scalar,sse2,avx2,neon}.rs files at
# https://github.com/srijs/sha1dc/tree/426b4afd to C table definitions.
# Usage: perl sha1dc-accel/generate-ubc-tables.pl sse2 < path/to/sse2.rs
# Output is tables only; retain ubc_check.c's declarations and MIT notice.
use strict;
use warnings;

my $form = shift // die "usage: $0 scalar|sse2|avx2|neon < form.rs\n";
my %prefix_count = (scalar => 70, sse2 => 26, avx2 => 16, neon => 20);
my %tail_count = (scalar => 151, sse2 => 97, avx2 => 77, neon => 137);
exists $prefix_count{$form} && !@ARGV or die "unknown form: $form\n";
local $/;
my $src = <STDIN> // die "empty input\n";
my ($prefix) = $src =~ /^fn prefix\([^\n]*\) -> u32 \{(.*?)^\}/ms;
defined $prefix or die "missing prefix()\n";
$prefix =~ s/\(([^()]*)\)\s+as\s+i32/$1/g;
$prefix =~ s/\b(\d+)u32\b/$1/g;
my (@out, @rows);

sub table {
	my ($type, $name, $rows) = @_;
	push @out, "static const struct $type ${form}_$name\[\] = {\n",
		map("\t{ $_ },\n", @$rows), "};\n\n";
}

sub condition {
	my ($expr) = @_;
	$expr =~ s/[\s()]//g;
	$expr =~ s/&1$//;
	$expr =~ /\Aw\[(\d+)\](?:>>(\d+))?\^w\[(\d+)\]
		(?:>>(\d+))?(?:\^([01]))?\z/x
		or die "unknown condition: $expr\n";
	return [$1, $2 // 0, $3, $4 // 0, $5 // 0];
}

sub vector {
	my ($expr, $width) = @_;
	my @v;
	if ($expr =~ /(?:_mm(?:256)?_set1_epi32|vdupq_n_u32)\(([^()]*)\)/) {
		@v = ($1) x $width;
	} elsif ($expr =~ /(?:_mm(?:256)?_set_epi32)\(([^()]*)\)/
		|| $expr =~ /splat\(\[([^\]]*)\]\)/s) {
		@v = split /\s*,\s*/, $1;
	} else {
		die "unknown vector: $expr\n";
	}
	@v == $width or die "wrong lane count: $expr\n";
	for (@v) {
		s/^\s+|\s+$//g;
		/\A(?:0|1\s*<<\s*\d+|DV_I{1,2}_\d+_\d+_BIT
			(?:\s*\|\s*DV_I{1,2}_\d+_\d+_BIT)*)\z/x
			or die "unknown lane: $_\n";
		s/\b1\s*<</1u <</;
	}
	# The pinned x86 source lists logical step order, despite set_epi32.
	return "{ " . join(", ", @v) . " }";
}

if ($form eq "scalar") {
	while ($prefix =~ /mask\s*&=\s*(.*?);/sg) {
		my ($expr, $dvs) = $1 =~ /\A(.*?)\s*\|\s*!\(([^()]*)\)\s*\z/s;
		defined $dvs or die "unknown scalar prefix\n";
		$expr =~ s/\s//g;
		my ($bits, $want);
		if ($expr =~ /\A(.*)\.wrapping_sub\(1\)\z/s) {
			($bits, $want) = ($1, 0);
		} elsif ($expr =~ /\A\(0\)\.wrapping_sub\((.*)\)\z/s) {
			($bits, $want) = ($1, 1);
		} else {
			die "unknown scalar mask: $expr\n";
		}
		my $c = condition($bits);
		$c->[4] = $want;
		$dvs =~ s/\s+/ /g;
		push @rows, "{ " . join(", ", @$c) . " }, $dvs";
	}
	table("ubc_prefix_cond", "prefix_conds", \@rows);
} else {
	my $width = $form eq "avx2" ? 8 : 4;
	my %want = (
		"_mm_and_si128(miss,bits)" => 1,
		"_mm_andnot_si128(miss,bits)" => 0,
		"_mm256_and_si256(miss,bits)" => 1,
		"_mm256_andnot_si256(miss,bits)" => 0,
		"vqsubq_u32(bits,set)" => 1,
		"vminq_u32(set,bits)" => 0,
	);
	while ($prefix =~ /\{\s*(let near =.*?)\}/sg) {
		my $g = $1;
		$g =~ s/\s//g;
		my ($lo, $hi, $x, $test, $dvs, $acc) = $g =~
			/\Alet near = load::<(\d+)>\(w\);\s*
			let far = load::<(\d+)>\(w\);\s*let x = (.*?);\s*
			let (?:tested|set) = (.*?);\s*
			(?:let miss = [^;]*;\s*)?let bits = (.*?);\s*
			acc[01] = \w+\(acc[01],\s*(.*?)\);\s*\z/sx;
		defined $acc or die "unknown vector group: $g\n";
		$x =~ s/(?:_mm(?:256)?_xor_si(?:128|256)|veorq_u32)/xor/g;
		$x =~ s/(?:_mm(?:256)?_srli_epi32|vshrq_n_u32)/shr/g;
		$x =~ s/\s//g;
		$x =~ /\Axor\((?:shr\(near,(\d+)\)|near),
			(?:shr\(far,(\d+)\)|far)\)\z/x
			or die "unknown vector XOR: $x\n";
		my ($ls, $hs) = ($1 // 0, $2 // 0);
		$acc =~ s/\s//g;
		exists $want{$acc} or die "unknown predicate: $acc\n";
		push @rows, join(", ", $lo, $ls // 0, $hi, $hs // 0,
			$want{$acc}, vector($test, $width),
			vector($dvs, $width));
	}
	table("ubc_group$width", "groups", \@rows);
}
@rows == $prefix_count{$form} or die "wrong prefix count\n";

if ($form ne "neon") {
	for my $key ("CHECKS", "SPANS") {
		my ($n, $body) = $src =~ /static\s+TAIL_$key:[^\n]*;
			\s*(\d+)\]\s*=\s*\[(.*?)^\];/msx;
		defined $body or die "missing TAIL_$key\n";
		my @tuples = $body =~ /\((\d+(?:\s*,\s*\d+)*)\)/g;
		@tuples == $n && $n == ($key eq "CHECKS" ?
			$tail_count{$form} : 32) or die "wrong tail count\n";
		table($key eq "CHECKS" ? "ubc_cond" : "tail_span",
			"tail_" . lc($key), \@tuples);
	}
} else {
	my @dv;
	my ($tail) = $src =~ /^fn tail\([^\n]*\) -> u32 \{(.*?)^\}/ms;
	defined $tail or die "missing tail()\n";
	$tail =~ s/\s//g;
	while ($tail =~ /if mask & DV_I{1,2}_\d+_\d+_BIT != 0 \{\s*
		let fail = (.*?);\s*out &= !(?:fail|\(fail << (\d+)\));/sgx) {
		my ($expr, $d) = ($1, $2 // 0);
		$d < 32 && !defined $dv[$d] or die "duplicate/invalid DV\n";
		$expr =~ s/[\s()]//g;
		$expr =~ s/&1$// or die "unknown tail mask\n";
		$dv[$d] = [map { join(", ", @{condition($_)}) }
			split /\|/, $expr];
	}
	my (@checks, @spans);
	for my $d (0 .. 31) {
		my $c = $dv[$d] // [];
		push @spans, scalar(@checks) . ", " . scalar(@$c);
		push @checks, @$c;
	}
	@checks == $tail_count{$form} or die "wrong NEON tail count\n";
	table("ubc_cond", "tail_checks", \@checks);
	table("tail_span", "tail_spans", \@spans);
}
print @out;
-- snap --

This Perl script reproduces the tables (although with different
formatting, and without the inline comments, I verified it with
`--patience --color-words="[A-Za-z0-9_]+|."`).

As is my rule, I only offer code that I wrote with AI assistance if the
output is close enough to what I would have written myself if I had the
time (and wouldn't need to take care of my arm muscles' health), and this
Perl script is no exception. My first draft would probably have used less
informative (or no) error messages, and I only learned about that `//`
operator during this session.

With all that out of the way, I would like to ask to include this script
in the patch (or in a follow-up patch) so that the lengthy `ubc_check.c`
file's tables can be validated/regenerated independently.

> diff --git a/sha1dc-accel/ubc_check.c b/sha1dc-accel/ubc_check.c
> new file mode 100644
> index 0000000000..f95b799f9d
> --- /dev/null
> +++ b/sha1dc-accel/ubc_check.c
> @@ -0,0 +1,1789 @@
> [...]
> +static uint32_t neon_prefix(const uint32_t *w)
> +{
> +	uint32x4_t acc = vdupq_n_u32(0);
> +	uint32x2_t folded;
> +	size_t i;
> +
> +	UNROLL_TABLE
> +	for (i = 0; i < ARRAY_SIZE(neon_groups); i++) {
> +		const struct ubc_group4 *g = &neon_groups[i];
> +		uint32x4_t lo = vld1q_u32(w + g->lo);
> +		uint32x4_t hi = vld1q_u32(w + g->hi);
> +		uint32x4_t dvs = vld1q_u32(g->dvs);
> +		uint32x4_t set, fail;
> +
> +		lo = vshlq_u32(lo, vdupq_n_s32(-(int32_t)g->lo_shift));
> +		hi = vshlq_u32(hi, vdupq_n_s32(-(int32_t)g->hi_shift));
> +		set = vtstq_u32(veorq_u32(lo, hi), vld1q_u32(g->test));
> +		/*
> +		 * The DVs of the lanes where the bit is not g->want. Each lane
> +		 * of set is all ones or zero, so a saturating subtraction keeps
> +		 * dvs where the bit is clear, and min keeps it where it is set.
> +		 */
> +		fail = g->want ? vqsubq_u32(dvs, set) : vminq_u32(set, dvs);

While this code is correct, I think it is slightly misleading: depending
on `want`, it either subtracts `set` from `dvs`, or takes the minimum. But
that only happens to be what is desired because each lane of `set` is all
ones or all zero. What we actually want is to mask either those lanes or
everything but those lanes, i.e. `dvs & ~set` or `dvs & set`,
respectively. That would be:

		fail = g->want ? vbicq_u32(dvs, set) : vandq_u32(dvs, set);

This has no speed impact nor does it produce a "more correct" result, but
it might improve readability a bit.

I haven't looked very closely whether there are similar issues elsewhere
(it is relatively tedious for me to learn all this NEON stuff on the go,
this is all new to me). If you're familiar with NEON, it might be
worthwhile looking for similarly "correct but misleading" statements.

But then, the proof lies in the pudding, as they say, and the code is
probably good enough as-is.

Ciao,
Johannes


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
                   ` (3 preceding siblings ...)
  2026-09-29 11:25 ` [PATCH 4/4] sha1dc-accel: compress with the ARMv8 SHA-1 instructions Scott Chacon
@ 2026-10-07 12:17 ` Johannes Schindelin
  2026-10-07 17:23 ` Junio C Hamano
  5 siblings, 0 replies; 21+ messages in thread
From: Johannes Schindelin @ 2026-10-07 12:17 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

[-- Attachment #1: Type: text/plain, Size: 20278 bytes --]

Hi Scott,

On Tue, 29 Sep 2026, Scott Chacon wrote:

> So, spoiler alert, the code in this patch series is mainly AI generated.
> I would try to fool you, but too many of you are far too aware of my
> actual C skills. That being said, I thought maybe someone here (especially
> those of you working on server optimization stuff) would be interested
> in the speed increases for both the server and client in making sha1dc
> quite a bit faster.

Hah, I beat you by almost a full day with my "competing" series at
https://lore.kernel.org/git/pull.2240.git.1790610691.gitgitgadget@gmail.com/.
I put the "competing" in double quotes because I think that both patch
series have merit, and I would love to see both merged.

Performance-wise, in my tests the Rust `sha1dc` was a tad faster that this
C-accell code on my Ryzen, something like 3.3% for my favorite
`index-pack` benchmark with 30 randomized pairings. Naturally, I tried to
figure out where that difference comes from, but I haven't been able to
finish that analysis to my satisfaction, although I would like to offer
this patch to avoid reordering the message schedule, which closes the gap
from 3.3% to 1%:

-- snip --
From: Johannes Schindelin <johannes.schindelin@gmx.de>
Date: Sun, 4 Oct 2026 12:53:56 +0200
Subject: [PATCH] sha1dc-accel: bring SHA-NI performance closer to Rust

The C SHA-NI path spends instructions reordering the message schedule,
while Rust uses a mirrored schedule. Avoid that overhead without changing
the collision checks.

Across 30 randomized, AC-powered triplets on the same pack, C's mean
fell from 8.201s to 8.053s. The paired mean improvement was 0.147s
(95% CI: 0.083-0.211s), leaving C about 1% behind Rust.

Assisted-by: GPT-6 Sol
Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>
---
 sha1dc-accel/internal.h  |  7 +++-
 sha1dc-accel/sha1.c      | 22 +++++++++--
 sha1dc-accel/ubc_check.c | 82 ++++++++++++++++++++++++++++++----------
 sha1dc-accel/x86.c       |  4 +-
 t/unit-tests/u-sha1dc.c  | 34 +++++++++++++----
 5 files changed, 112 insertions(+), 37 deletions(-)

diff --git a/sha1dc-accel/internal.h b/sha1dc-accel/internal.h
index a50b294b96b..bcb60811a8f 100644
--- a/sha1dc-accel/internal.h
+++ b/sha1dc-accel/internal.h
@@ -4,8 +4,9 @@
 /*
  * Shared between the files of sha1dc-accel/. See sha1.c for an overview.
  *
- * The schedule `w` is always the 80 expanded message words in step order,
- * w[t] at index t, which is also what sha1dc/ keeps in SHA1_CTX.m1.
+ * The schedule `w` has 80 expanded message words. Normally w[t] is at
+ * index t, as in sha1dc/. SHA-NI spills w[t] at index 79 - t; a candidate
+ * is put back in step order before recompression.
  *
  * A "state" is the five working words [a, b, c, d, e] before a step.
  */
@@ -75,9 +76,11 @@ enum sha1dc_from {
 uint32_t sha1dc_ubc_check_scalar(const uint32_t w[80]);
 #ifdef SHA1DC_HAVE_SSE2
 uint32_t sha1dc_ubc_check_sse2(const uint32_t w[80]);
+uint32_t sha1dc_ubc_check_sse2_mirrored(const uint32_t w[80]);
 #endif
 #ifdef SHA1DC_HAVE_AVX2
 uint32_t sha1dc_ubc_check_avx2(const uint32_t w[80]);
+uint32_t sha1dc_ubc_check_avx2_mirrored(const uint32_t w[80]);
 #endif
 #ifdef SHA1DC_HAVE_NEON
 uint32_t sha1dc_ubc_check_neon(const uint32_t w[80]);
diff --git a/sha1dc-accel/sha1.c b/sha1dc-accel/sha1.c
index 1f1133169b5..bddc3bcdf2d 100644
--- a/sha1dc-accel/sha1.c
+++ b/sha1dc-accel/sha1.c
@@ -259,9 +259,11 @@ static int shani_avx2_available(void)
 /* In order of preference. */
 static const struct backend backends[] = {
 #ifdef SHA1DC_HAVE_SHANI
-	{ "shani+avx2", sha1dc_compress_shani, 1, sha1dc_ubc_check_avx2,
+	{ "shani+avx2", sha1dc_compress_shani, 1,
+	  sha1dc_ubc_check_avx2_mirrored,
 	  sha1dc_recompress_shani, shani_avx2_available },
-	{ "shani+sse2", sha1dc_compress_shani, 1, sha1dc_ubc_check_sse2,
+	{ "shani+sse2", sha1dc_compress_shani, 1,
+	  sha1dc_ubc_check_sse2_mirrored,
 	  sha1dc_recompress_shani, sha1dc_shani_available },
 #endif
 #ifdef SHA1DC_HAVE_ARMV8
@@ -432,8 +434,20 @@ static inline void process(const struct backend *be, SHA1_CTX *ctx,
 		return;
 
 	candidates = ctx->ubc_check ? be->ubc_check(w) : 0xFFFFFFFF;
-	if (candidates &&
-	    attacked(be, ctx, candidates, w, s1, s2, ihv_in, ctx->ihv)) {
+	if (!candidates)
+		return;
+#ifdef SHA1DC_HAVE_SHANI
+	if (be->compress == sha1dc_compress_shani) {
+		int i;
+
+		for (i = 0; i < 40; i++) {
+			uint32_t tmp = w[i];
+			w[i] = w[79 - i];
+			w[79 - i] = tmp;
+		}
+	}
+#endif
+	if (attacked(be, ctx, candidates, w, s1, s2, ihv_in, ctx->ihv)) {
 		ctx->found_collision = 1;
 		/*
 		 * Two more compressions of this block give a digest that the
diff --git a/sha1dc-accel/ubc_check.c b/sha1dc-accel/ubc_check.c
index f95b799f9d4..cbd21d33e68 100644
--- a/sha1dc-accel/ubc_check.c
+++ b/sha1dc-accel/ubc_check.c
@@ -105,9 +105,13 @@ struct ubc_cond {
 	uint8_t i, a, j, b, c;
 };
 
-static inline uint32_t cond_fails(const uint32_t *w, const struct ubc_cond *c)
+static inline uint32_t cond_fails(const uint32_t *w, const struct ubc_cond *c,
+				  int mirrored)
 {
-	return (((w[c->i] >> c->a) ^ (w[c->j] >> c->b) ^ c->c) & 1);
+	unsigned i = mirrored ? 79 - c->i : c->i;
+	unsigned j = mirrored ? 79 - c->j : c->j;
+
+	return (((w[i] >> c->a) ^ (w[j] >> c->b) ^ c->c) & 1);
 }
 
 /* A condition of the scalar prefix, and the DVs it rules out if it fails. */
@@ -149,7 +153,7 @@ struct tail_span {
  */
 static inline uint32_t run_tail(const uint32_t *w, uint32_t mask,
 				const struct ubc_cond *checks,
-				const struct tail_span *spans)
+				const struct tail_span *spans, int mirrored)
 {
 	uint32_t out = mask;
 	while (mask) {
@@ -160,7 +164,7 @@ static inline uint32_t run_tail(const uint32_t *w, uint32_t mask,
 
 		NO_VECTORIZE
 		for (; c < end; c++)
-			fail |= cond_fails(w, c);
+			fail |= cond_fails(w, c, mirrored);
 		out &= ~(fail << d);
 		mask &= mask - 1;
 	}
@@ -569,7 +573,7 @@ static uint32_t scalar_prefix(const uint32_t *w)
 	UNROLL_TABLE
 	for (i = 0; i < ARRAY_SIZE(scalar_prefix_conds); i++) {
 		const struct ubc_prefix_cond *p = &scalar_prefix_conds[i];
-		mask &= ~(p->dvs & (0 - cond_fails(w, &p->cond)));
+		mask &= ~(p->dvs & (0 - cond_fails(w, &p->cond, 0)));
 	}
 	return mask;
 }
@@ -580,7 +584,7 @@ uint32_t sha1dc_ubc_check_scalar(const uint32_t w[80])
 	/* Every check only clears bits, so an empty mask settles it. */
 	if (!mask)
 		return 0;
-	return run_tail(w, mask, scalar_tail_checks, scalar_tail_spans);
+	return run_tail(w, mask, scalar_tail_checks, scalar_tail_spans, 0);
 }
 
 /* neon form */
@@ -980,7 +984,7 @@ uint32_t sha1dc_ubc_check_neon(const uint32_t w[80])
 	/* Every check only clears bits, so an empty mask settles it. */
 	if (!mask)
 		return 0;
-	return run_tail(w, mask, neon_tail_checks, neon_tail_spans);
+	return run_tail(w, mask, neon_tail_checks, neon_tail_spans, 0);
 }
 
 #endif /* SHA1DC_HAVE_NEON */
@@ -1343,7 +1347,7 @@ static const struct tail_span sse2_tail_spans[32] = {
 };
 
 SHA1DC_TARGET_SSE2
-static uint32_t sse2_prefix(const uint32_t *w)
+static inline uint32_t sse2_prefix(const uint32_t *w, int mirrored)
 {
 	const __m128i zero = _mm_setzero_si128();
 	__m128i acc = zero;
@@ -1352,10 +1356,18 @@ static uint32_t sse2_prefix(const uint32_t *w)
 	UNROLL_TABLE
 	for (i = 0; i < ARRAY_SIZE(sse2_groups); i++) {
 		const struct ubc_group4 *g = &sse2_groups[i];
-		__m128i lo = _mm_loadu_si128((const __m128i *)(w + g->lo));
-		__m128i hi = _mm_loadu_si128((const __m128i *)(w + g->hi));
-		__m128i test = _mm_loadu_si128((const __m128i *)g->test);
-		__m128i dvs = _mm_loadu_si128((const __m128i *)g->dvs);
+		const uint32_t *lo_w = w + (mirrored ? 76 - g->lo : g->lo);
+		const uint32_t *hi_w = w + (mirrored ? 76 - g->hi : g->hi);
+		__m128i lo = _mm_loadu_si128((const __m128i *)lo_w);
+		__m128i hi = _mm_loadu_si128((const __m128i *)hi_w);
+		__m128i test = mirrored ?
+			_mm_set_epi32(g->test[0], g->test[1],
+				      g->test[2], g->test[3]) :
+			_mm_loadu_si128((const __m128i *)g->test);
+		__m128i dvs = mirrored ?
+			_mm_set_epi32(g->dvs[0], g->dvs[1],
+				      g->dvs[2], g->dvs[3]) :
+			_mm_loadu_si128((const __m128i *)g->dvs);
 		__m128i clear, fail;
 
 		lo = _mm_srl_epi32(lo, _mm_cvtsi32_si128(g->lo_shift));
@@ -1376,11 +1388,20 @@ static uint32_t sse2_prefix(const uint32_t *w)
 SHA1DC_TARGET_SSE2
 uint32_t sha1dc_ubc_check_sse2(const uint32_t w[80])
 {
-	uint32_t mask = sse2_prefix(w);
+	uint32_t mask = sse2_prefix(w, 0);
 	/* Every check only clears bits, so an empty mask settles it. */
 	if (!mask)
 		return 0;
-	return run_tail(w, mask, sse2_tail_checks, sse2_tail_spans);
+	return run_tail(w, mask, sse2_tail_checks, sse2_tail_spans, 0);
+}
+
+SHA1DC_TARGET_SSE2
+uint32_t sha1dc_ubc_check_sse2_mirrored(const uint32_t w[80])
+{
+	uint32_t mask = sse2_prefix(w, 1);
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, sse2_tail_checks, sse2_tail_spans, 1);
 }
 
 #endif /* SHA1DC_HAVE_SSE2 */
@@ -1743,7 +1764,7 @@ static const struct tail_span avx2_tail_spans[32] = {
 };
 
 SHA1DC_TARGET_AVX2
-static uint32_t avx2_prefix(const uint32_t *w)
+static inline uint32_t avx2_prefix(const uint32_t *w, int mirrored)
 {
 	const __m256i zero = _mm256_setzero_si256();
 	__m256i acc = zero;
@@ -1753,10 +1774,20 @@ static uint32_t avx2_prefix(const uint32_t *w)
 	UNROLL_TABLE
 	for (i = 0; i < ARRAY_SIZE(avx2_groups); i++) {
 		const struct ubc_group8 *g = &avx2_groups[i];
-		__m256i lo = _mm256_loadu_si256((const __m256i *)(w + g->lo));
-		__m256i hi = _mm256_loadu_si256((const __m256i *)(w + g->hi));
-		__m256i test = _mm256_loadu_si256((const __m256i *)g->test);
-		__m256i dvs = _mm256_loadu_si256((const __m256i *)g->dvs);
+		const uint32_t *lo_w = w + (mirrored ? 72 - g->lo : g->lo);
+		const uint32_t *hi_w = w + (mirrored ? 72 - g->hi : g->hi);
+		__m256i lo = _mm256_loadu_si256((const __m256i *)lo_w);
+		__m256i hi = _mm256_loadu_si256((const __m256i *)hi_w);
+		__m256i test = mirrored ?
+			_mm256_set_epi32(g->test[0], g->test[1], g->test[2],
+					 g->test[3], g->test[4], g->test[5],
+					 g->test[6], g->test[7]) :
+			_mm256_loadu_si256((const __m256i *)g->test);
+		__m256i dvs = mirrored ?
+			_mm256_set_epi32(g->dvs[0], g->dvs[1], g->dvs[2],
+					 g->dvs[3], g->dvs[4], g->dvs[5],
+					 g->dvs[6], g->dvs[7]) :
+			_mm256_loadu_si256((const __m256i *)g->dvs);
 		__m256i clear, fail;
 
 		lo = _mm256_srl_epi32(lo, _mm_cvtsi32_si128(g->lo_shift));
@@ -1779,11 +1810,20 @@ static uint32_t avx2_prefix(const uint32_t *w)
 SHA1DC_TARGET_AVX2
 uint32_t sha1dc_ubc_check_avx2(const uint32_t w[80])
 {
-	uint32_t mask = avx2_prefix(w);
+	uint32_t mask = avx2_prefix(w, 0);
 	/* Every check only clears bits, so an empty mask settles it. */
 	if (!mask)
 		return 0;
-	return run_tail(w, mask, avx2_tail_checks, avx2_tail_spans);
+	return run_tail(w, mask, avx2_tail_checks, avx2_tail_spans, 0);
+}
+
+SHA1DC_TARGET_AVX2
+uint32_t sha1dc_ubc_check_avx2_mirrored(const uint32_t w[80])
+{
+	uint32_t mask = avx2_prefix(w, 1);
+	if (!mask)
+		return 0;
+	return run_tail(w, mask, avx2_tail_checks, avx2_tail_spans, 1);
 }
 
 #endif /* SHA1DC_HAVE_AVX2 */
diff --git a/sha1dc-accel/x86.c b/sha1dc-accel/x86.c
index c2a0c71f17f..dfdc1fc3b62 100644
--- a/sha1dc-accel/x86.c
+++ b/sha1dc-accel/x86.c
@@ -66,8 +66,8 @@ int sha1dc_shani_available(void)
 #define LOADU(p) _mm_loadu_si128((const __m128i *)(const void *)(p))
 #define STOREU(p, v) _mm_storeu_si128((__m128i *)(void *)(p), (v))
 
-/* Writes group `v` (steps t..t+3, held reversed) to w[t..t+3]. */
-#define SPILL(t, v) STOREU(w + (t), _mm_shuffle_epi32((v), REVERSE))
+/* The group is already reversed; store it in the mirrored schedule. */
+#define SPILL(t, v) STOREU(w + 76 - (t), (v))
 
 /* Schedule words 4k..4k+3 for k from 8 on, from groups k-8, k-7, k-4, k-2, k-1. */
 SHA1DC_TARGET_SHANI
diff --git a/t/unit-tests/u-sha1dc.c b/t/unit-tests/u-sha1dc.c
index f629f59d1d6..c658c219a3a 100644
--- a/t/unit-tests/u-sha1dc.c
+++ b/t/unit-tests/u-sha1dc.c
@@ -151,6 +151,21 @@ static void check_ubc_forms(uint32_t w[80], int flip, int avx2)
 		check_form(sha1dc_ubc_check_scalar(w), want);
 #ifdef SHA1DC_HAVE_SSE2
 		check_form(sha1dc_ubc_check_sse2(w), want);
+		{
+			uint32_t mirrored[80];
+			int t;
+
+			for (t = 0; t < 80; t++)
+				mirrored[t] = w[79 - t];
+			check_form(sha1dc_ubc_check_sse2_mirrored(mirrored),
+				   want);
+#ifdef SHA1DC_HAVE_AVX2
+			if (avx2)
+				check_form(
+					sha1dc_ubc_check_avx2_mirrored(
+						mirrored), want);
+#endif
+		}
 #endif
 #ifdef SHA1DC_HAVE_AVX2
 		if (avx2)
@@ -243,14 +258,15 @@ typedef int (*recompress_fn)(enum sha1dc_from from, const uint32_t m1[80],
 			     const uint32_t dm[80], const uint32_t state[5],
 			     const uint32_t ihv_out[5]);
 
-static void check_compress(compress_fn compress)
+static void check_compress(compress_fn compress, int mirrored)
 {
 	int i, t;
 
 	rng_seed(3);
 	for (i = 0; i < 2000; i++) {
 		unsigned char block[64];
-		uint32_t ihv[5], got[5], w[80], at_60[5], at_64[5], s[5];
+		uint32_t ihv[5], got[5], w[80], expected[80];
+		uint32_t at_60[5], at_64[5], s[5];
 
 		for (t = 0; t < 64; t++)
 			block[t] = rng();
@@ -260,14 +276,16 @@ static void check_compress(compress_fn compress)
 
 		memcpy(s, ihv, sizeof(s));
 		for (t = 0; t < 80; t++) {
-			uint32_t want_w = t < 16 ? get_be32(block + 4 * t) :
-				rol(w[t - 3] ^ w[t - 8] ^ w[t - 14] ^ w[t - 16], 1);
-			cl_assert_equal_i(w[t], want_w);
+			expected[t] = t < 16 ? get_be32(block + 4 * t) :
+				rol(expected[t - 3] ^ expected[t - 8] ^
+				    expected[t - 14] ^ expected[t - 16], 1);
+			cl_assert_equal_i(w[mirrored ? 79 - t : t],
+					  expected[t]);
 			if (t == 60)
 				cl_assert(!memcmp(at_60, s, sizeof(s)));
 			if (t == 64)
 				cl_assert(!memcmp(at_64, s, sizeof(s)));
-			step(s, t, w[t]);
+			step(s, t, expected[t]);
 		}
 		for (t = 0; t < 5; t++)
 			cl_assert_equal_i(got[t], ihv[t] + s[t]);
@@ -321,14 +339,14 @@ static void hardware_compression(void)
 
 #ifdef SHA1DC_HAVE_SHANI
 	if (sha1dc_shani_available()) {
-		check_compress(sha1dc_compress_shani);
+		check_compress(sha1dc_compress_shani, 1);
 		check_recompress(sha1dc_recompress_shani);
 		tested = 1;
 	}
 #endif
 #ifdef SHA1DC_HAVE_ARMV8
 	if (sha1dc_armv8_available()) {
-		check_compress(sha1dc_compress_armv8);
+		check_compress(sha1dc_compress_armv8, 0);
 		check_recompress(sha1dc_recompress_armv8);
 		tested = 1;
 	}
-- snap --

This is admittedly a bit gnarly, and really, really hard to understand
unless you immersed yourself in Sam's work. But it _does_ accelerate SHA-1
computation with my Ryzen 7, and I'd be interested to hear whether it has
an equivalent effect with your Xeon.

> This series ports the approach of Sam Reis's sha1dc Rust crate [1], 
> which gitoxide recently switched to [2], to C. 
> 
> The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
> on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
> 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
> 
> Hashing throughput on the Xeon, in MiB/s:
> 
>                                 16KiB    1MiB   vs OpenSSL
>   OpenSSL SHA-1 (no detection)   1234    1129      1.00x
>   sha1dc/ (today)                 435     450      2.67x
>   shani+avx2 (default here)      1002     901      1.24x
>   shani+sse2                     1075    1008      1.13x
>   portable+avx2                   553     654      1.96x
>   portable+sse2                   603     681      1.84x
>   portable                        466     565      2.29x
> 
> In other words, currently collision detection costs about 1.5–2.5x on
> top of the hashing itself today, but only about 0.2x with the series. 
> 
> The patches are:
> 
>   [1/4]: sha1dc-accel: add a block loop for sha1dc's SHA1_CTX
> 
>     Just groundwork: our own block loop around sha1dc's context and DV
>     table, with the same results and a few percent slower, plus tests
>     that compare against sha1dc/ directly, including on real collisions
>     in every mode.
> 
>   [2/4]: sha1dc-accel: vectorize the unavoidable-bitconditions check
> 
>     The UBC filter rewritten as SSE2, AVX2, NEON, and new scalar forms, 
>     using the conditions the crate's solver picks for each. They're 
>     carried as tables, with a short loop per form to run them. 
>     1.29x on the Xeon, 1.27x on the M5 Max.
> 
>   [3/4]: sha1dc-accel: compress with SHA-NI on x86-64
> 
>     Hardware compression, with the schedule spilled, and recompression
>     of flagged blocks in hardware, too. Another 2.07x on the Xeon.
> 
>   [4/4]: sha1dc-accel: compress with the ARMv8 SHA-1 instructions
> 
>     The same for arm64. Another 2.4x on the M5 Max.

I really like this structure.

BTW I have run the entire test suite both on a Ryzen 7 (using WSL) and on
a Windows/ARM64 Cloud PC (using straight Windows because that Cloud PC
does not support WSL), with a slightly patched version: running the
original (slow) sha1dc, the sha1dc-accel and the Rust sha1dc. There were 0
discrepancies, which meshes with the Sol-assisted analysis of the code,
comparing it against the paper and Sam's detailed blog post.

I also wanted to compare the speed, but unfortunately, my tests always
suffer the noisy neighbor problem (working on a laptop with parallel work
going on, or Cloud PC that is of course a VM in a rack somewhere in
Virginia). So all I can really claim is that in my tests, the speed was
comparable (with the patch above).

All that said, I am very much in favor of accepting your patch series into
Git, with or without my suggested changes.

Thanks!
Johannes

> 
> The x86 numbers are from a 4-vCPU Xeon VM with SHA-NI and AVX2 (GCC 13,
> Linux), which is unfortunately rather noisy; the per-patch hyperfine
> output has the spread. The arm64 numbers are medians of 9 runs on an
> Apple M5 Max (Apple clang, macOS). The full test suite passes on both.
> 
> [1] https://sam.dev/blog/faster-sha1-collision-detection
> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
> 
> Scott Chacon (4):
>   sha1dc-accel: add a block loop for sha1dc's SHA1_CTX
>   sha1dc-accel: vectorize the unavoidable-bitconditions check
>   sha1dc-accel: compress with SHA-NI on x86-64
>   sha1dc-accel: compress with the ARMv8 SHA-1 instructions
> 
>  Makefile                            |   14 +
>  contrib/buildsystems/CMakeLists.txt |    2 +-
>  meson.build                         |    4 +
>  sha1dc-accel/arm.c                  |  274 ++++
>  sha1dc-accel/internal.h             |  126 ++
>  sha1dc-accel/sha1.c                 |  498 ++++++++
>  sha1dc-accel/sha1.h                 |   31 +
>  sha1dc-accel/ubc_check.c            | 1789 +++++++++++++++++++++++++++
>  sha1dc-accel/x86.c                  |  260 ++++
>  sha1dc_git.c                        |   18 +
>  t/.gitattributes                    |    1 +
>  t/helper/test-sha1.c                |   95 ++
>  t/helper/test-tool.c                |    2 +
>  t/helper/test-tool.h                |    2 +
>  t/meson.build                       |    1 +
>  t/t0013-sha1dc.sh                   |   47 +
>  t/t0013/sha-mbles-1.bin             |  Bin 0 -> 640 bytes
>  t/t0013/sha1-reduced-round.bin      |  Bin 0 -> 128 bytes
>  t/unit-tests/u-sha1dc.c             |  366 ++++++
>  19 files changed, 3529 insertions(+), 1 deletion(-)
>  create mode 100644 sha1dc-accel/arm.c
>  create mode 100644 sha1dc-accel/internal.h
>  create mode 100644 sha1dc-accel/sha1.c
>  create mode 100644 sha1dc-accel/sha1.h
>  create mode 100644 sha1dc-accel/ubc_check.c
>  create mode 100644 sha1dc-accel/x86.c
>  create mode 100644 t/t0013/sha-mbles-1.bin
>  create mode 100644 t/t0013/sha1-reduced-round.bin
>  create mode 100644 t/unit-tests/u-sha1dc.c
> 
> 
> base-commit: a018953688f1b10bddf91bff8747068f5f4746a4
> -- 
> 2.50.1 (Apple Git-155)
> 
> 
> 
> 

^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
                   ` (4 preceding siblings ...)
  2026-10-07 12:17 ` [PATCH 0/4] faster SHA-1 collision detection Johannes Schindelin
@ 2026-10-07 17:23 ` Junio C Hamano
  2026-10-07 18:13   ` Scott Chacon
  5 siblings, 1 reply; 21+ messages in thread
From: Junio C Hamano @ 2026-10-07 17:23 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

Scott Chacon <scott@gitbutler.net> writes:

> So, spoiler alert, the code in this patch series is mainly AI generated.
> I would try to fool you, but too many of you are far too aware of my
> actual C skills. That being said, I thought maybe someone here (especially
> those of you working on server optimization stuff) would be interested
> in the speed increases for both the server and client in making sha1dc
> quite a bit faster.
>
> This series ports the approach of Sam Reis's sha1dc Rust crate [1], 
> which gitoxide recently switched to [2], to C. 

Which means license-wise the original is compatible with us, I
presume, as they are "Apache2 or MIT, your choice".

How can you/we be sure, with respect to the current AI policy in
SubmittingPatches (which by the way was vetted by SFC lawyers), that
your "AI generated" code did not "borrow" from places that gets
you/us into trouble?

> The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
> on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
> 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
>
> Hashing throughput on the Xeon, in MiB/s:
>
>                                 16KiB    1MiB   vs OpenSSL
>   OpenSSL SHA-1 (no detection)   1234    1129      1.00x
>   sha1dc/ (today)                 435     450      2.67x
>   shani+avx2 (default here)      1002     901      1.24x
>   shani+sse2                     1075    1008      1.13x
>   portable+avx2                   553     654      1.96x
>   portable+sse2                   603     681      1.84x
>   portable                        466     565      2.29x
>
> In other words, currently collision detection costs about 1.5–2.5x on
> top of the hashing itself today, but only about 0.2x with the series. 

Thanks for these numbers.

> [1] https://sam.dev/blog/faster-sha1-collision-detection
> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008

And the pointers to the original sources.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-07 17:23 ` Junio C Hamano
@ 2026-10-07 18:13   ` Scott Chacon
  2026-10-08  6:21     ` Sebastian Thiel
  0 siblings, 1 reply; 21+ messages in thread
From: Scott Chacon @ 2026-10-07 18:13 UTC (permalink / raw)
  To: Junio C Hamano; +Cc: Scott Chacon, git, Sam Reis, Sebastian Thiel

Hey,

On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
> > This series ports the approach of Sam Reis's sha1dc Rust crate [1],
> > which gitoxide recently switched to [2], to C.
>
> Which means license-wise the original is compatible with us, I
> presume, as they are "Apache2 or MIT, your choice".
>
> How can you/we be sure, with respect to the current AI policy in
> SubmittingPatches (which by the way was vetted by SFC lawyers), that
> your "AI generated" code did not "borrow" from places that gets
> you/us into trouble?

It's a good question. I actually just submitted a proposed update to
that policy based on SFC's updated guidelines, but either way, I
learned about this from Sam and have talked to him about the port and
he seemed excited about it. I can triple check, but I'm fairly
confident that he's fine with this and I am fine signing off on it
under the terms of the DCO language.

Of course, he in turn used AI tooling to produce _his_ library, but
within the guidelines of the updated SFC guidelines. Johannes's
alternative series is the original Rust code of Sam that my agent
looked at to produce this (in addition to his blog post explaining
it), so I'm not sure how that might be materially different.

> > The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
> > on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
> > 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
> >
> > Hashing throughput on the Xeon, in MiB/s:
> >
> >                                 16KiB    1MiB   vs OpenSSL
> >   OpenSSL SHA-1 (no detection)   1234    1129      1.00x
> >   sha1dc/ (today)                 435     450      2.67x
> >   shani+avx2 (default here)      1002     901      1.24x
> >   shani+sse2                     1075    1008      1.13x
> >   portable+avx2                   553     654      1.96x
> >   portable+sse2                   603     681      1.84x
> >   portable                        466     565      2.29x
> >
> > In other words, currently collision detection costs about 1.5–2.5x on
> > top of the hashing itself today, but only about 0.2x with the series.
>
> Thanks for these numbers.

It would have been better had I provided the same relative scale (it
should be 1.5-2.5x vs 1.2x, but whatever, you probably get it. It's
20% overhead here vs 50%-150% overhead previously).

> > [1] https://sam.dev/blog/faster-sha1-collision-detection
> > [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
>
> And the pointers to the original sources.

CC'ing Sam (sha1dc rust guy) and Sebastian (Gitoxide) on this, just in
case they have an opinion but I'm pretty sure they would be more than
happy for this to be integrated.

Scott

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check
  2026-10-07 12:17   ` Johannes Schindelin
@ 2026-10-07 21:32     ` Junio C Hamano
  0 siblings, 0 replies; 21+ messages in thread
From: Junio C Hamano @ 2026-10-07 21:32 UTC (permalink / raw)
  To: Johannes Schindelin; +Cc: Scott Chacon, git

Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:

> This Perl script reproduces the tables (although with different
> formatting, and without the inline comments, I verified it with
> `--patience --color-words="[A-Za-z0-9_]+|."`).
> ...
> With all that out of the way, I would like to ask to include this script
> in the patch (or in a follow-up patch) so that the lengthy `ubc_check.c`
> file's tables can be validated/regenerated independently.
> ...
> While this code is correct, I think it is slightly misleading: depending
> on `want`, it either subtracts `set` from `dvs`, or takes the minimum. But
> that only happens to be what is desired because each lane of `set` is all
> ones or all zero. What we actually want is to mask either those lanes or
> everything but those lanes, i.e. `dvs & ~set` or `dvs & set`,
> respectively. That would be:
>
> 		fail = g->want ? vbicq_u32(dvs, set) : vandq_u32(dvs, set);
>
> This has no speed impact nor does it produce a "more correct" result, but
> it might improve readability a bit.
>
> I haven't looked very closely whether there are similar issues elsewhere
> (it is relatively tedious for me to learn all this NEON stuff on the go,
> this is all new to me). If you're familiar with NEON, it might be
> worthwhile looking for similarly "correct but misleading" statements.

Thanks for offering a very thoughtful help and offering to work well
together.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-07 18:13   ` Scott Chacon
@ 2026-10-08  6:21     ` Sebastian Thiel
  2026-10-08 11:20       ` Sam Reis
  0 siblings, 1 reply; 21+ messages in thread
From: Sebastian Thiel @ 2026-10-08  6:21 UTC (permalink / raw)
  To: Scott Chacon, Junio C Hamano; +Cc: Scott Chacon, git, Sam Reis

Thanks for reeling me in, Scott!

First of all, I am very happy to see that overall, everyone here is
making an effort to find a way to speed up SHA-1dc again.
It's so impactful!


It really did hurt when I finally had to add SHA-1dc to Gitoxide and see 
the performance of clones plummet. And it still hurts me knowing that
an incredible amount of CPU time is wasted doing something that we now
know can be done much faster. At GitHub scale, this must be more than
a blip.

While it's my dream to one day have a GitHub action that uses `gix` to
clone and safe even more power, I think Git is in a far better spot
to achieve significant savings much sooner.

On 07.10.26 20:13, Scott Chacon wrote:
> Hey,
> 
> On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
>>> This series ports the approach of Sam Reis's sha1dc Rust crate [1],
>>> which gitoxide recently switched to [2], to C.
>>
>> Which means license-wise the original is compatible with us, I
>> presume, as they are "Apache2 or MIT, your choice".
>>
>> How can you/we be sure, with respect to the current AI policy in
>> SubmittingPatches (which by the way was vetted by SFC lawyers), that
>> your "AI generated" code did not "borrow" from places that gets
>> you/us into trouble?
> 
> It's a good question. I actually just submitted a proposed update to
> that policy based on SFC's updated guidelines, but either way, I
> learned about this from Sam and have talked to him about the port and
> he seemed excited about it. I can triple check, but I'm fairly
> confident that he's fine with this and I am fine signing off on it
> under the terms of the DCO language.
> 
> Of course, he in turn used AI tooling to produce _his_ library, but
> within the guidelines of the updated SFC guidelines. Johannes's
> alternative series is the original Rust code of Sam that my agent
> looked at to produce this (in addition to his blog post explaining
> it), so I'm not sure how that might be materially different.
> 
>>> The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
>>> on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
>>> 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
>>>
>>> Hashing throughput on the Xeon, in MiB/s:
>>>
>>>                                  16KiB    1MiB   vs OpenSSL
>>>    OpenSSL SHA-1 (no detection)   1234    1129      1.00x
>>>    sha1dc/ (today)                 435     450      2.67x
>>>    shani+avx2 (default here)      1002     901      1.24x
>>>    shani+sse2                     1075    1008      1.13x
>>>    portable+avx2                   553     654      1.96x
>>>    portable+sse2                   603     681      1.84x
>>>    portable                        466     565      2.29x
>>>
>>> In other words, currently collision detection costs about 1.5–2.5x on
>>> top of the hashing itself today, but only about 0.2x with the series.
>>
>> Thanks for these numbers.
> 
> It would have been better had I provided the same relative scale (it
> should be 1.5-2.5x vs 1.2x, but whatever, you probably get it. It's
> 20% overhead here vs 50%-150% overhead previously).
> 
>>> [1] https://sam.dev/blog/faster-sha1-collision-detection
>>> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
>>
>> And the pointers to the original sources.
> 
> CC'ing Sam (sha1dc rust guy) and Sebastian (Gitoxide) on this, just in
> case they have an opinion but I'm pretty sure they would be more than
> happy for this to be integrated.
> 
> Scott


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08  6:21     ` Sebastian Thiel
@ 2026-10-08 11:20       ` Sam Reis
  2026-10-08 13:25         ` D. Ben Knoble
  2026-10-08 15:55         ` Junio C Hamano
  0 siblings, 2 replies; 21+ messages in thread
From: Sam Reis @ 2026-10-08 11:20 UTC (permalink / raw)
  To: Sebastian Thiel; +Cc: Scott Chacon, Junio C Hamano, Scott Chacon, git

Hey everyone. Just for the avoidance of doubt, very happy to see
Scott's patch here land and for git to benefit from faster sha1dc
hashing. Let me know if I can do anything to support!


On Thu, Oct 8, 2026 at 8:21 AM Sebastian Thiel
<sebastian.thiel@icloud.com> wrote:
>
> Thanks for reeling me in, Scott!
>
> First of all, I am very happy to see that overall, everyone here is
> making an effort to find a way to speed up SHA-1dc again.
> It's so impactful!
>
>
> It really did hurt when I finally had to add SHA-1dc to Gitoxide and see
> the performance of clones plummet. And it still hurts me knowing that
> an incredible amount of CPU time is wasted doing something that we now
> know can be done much faster. At GitHub scale, this must be more than
> a blip.
>
> While it's my dream to one day have a GitHub action that uses `gix` to
> clone and safe even more power, I think Git is in a far better spot
> to achieve significant savings much sooner.
>
> On 07.10.26 20:13, Scott Chacon wrote:
> > Hey,
> >
> > On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
> >>> This series ports the approach of Sam Reis's sha1dc Rust crate [1],
> >>> which gitoxide recently switched to [2], to C.
> >>
> >> Which means license-wise the original is compatible with us, I
> >> presume, as they are "Apache2 or MIT, your choice".
> >>
> >> How can you/we be sure, with respect to the current AI policy in
> >> SubmittingPatches (which by the way was vetted by SFC lawyers), that
> >> your "AI generated" code did not "borrow" from places that gets
> >> you/us into trouble?
> >
> > It's a good question. I actually just submitted a proposed update to
> > that policy based on SFC's updated guidelines, but either way, I
> > learned about this from Sam and have talked to him about the port and
> > he seemed excited about it. I can triple check, but I'm fairly
> > confident that he's fine with this and I am fine signing off on it
> > under the terms of the DCO language.
> >
> > Of course, he in turn used AI tooling to produce _his_ library, but
> > within the guidelines of the updated SFC guidelines. Johannes's
> > alternative series is the original Rust code of Sam that my agent
> > looked at to produce this (in addition to his blog post explaining
> > it), so I'm not sure how that might be materially different.
> >
> >>> The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
> >>> on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
> >>> 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
> >>>
> >>> Hashing throughput on the Xeon, in MiB/s:
> >>>
> >>>                                  16KiB    1MiB   vs OpenSSL
> >>>    OpenSSL SHA-1 (no detection)   1234    1129      1.00x
> >>>    sha1dc/ (today)                 435     450      2.67x
> >>>    shani+avx2 (default here)      1002     901      1.24x
> >>>    shani+sse2                     1075    1008      1.13x
> >>>    portable+avx2                   553     654      1.96x
> >>>    portable+sse2                   603     681      1.84x
> >>>    portable                        466     565      2.29x
> >>>
> >>> In other words, currently collision detection costs about 1.5–2.5x on
> >>> top of the hashing itself today, but only about 0.2x with the series.
> >>
> >> Thanks for these numbers.
> >
> > It would have been better had I provided the same relative scale (it
> > should be 1.5-2.5x vs 1.2x, but whatever, you probably get it. It's
> > 20% overhead here vs 50%-150% overhead previously).
> >
> >>> [1] https://sam.dev/blog/faster-sha1-collision-detection
> >>> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
> >>
> >> And the pointers to the original sources.
> >
> > CC'ing Sam (sha1dc rust guy) and Sebastian (Gitoxide) on this, just in
> > case they have an opinion but I'm pretty sure they would be more than
> > happy for this to be integrated.
> >
> > Scott
>

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 11:20       ` Sam Reis
@ 2026-10-08 13:25         ` D. Ben Knoble
  2026-10-08 13:59           ` Sam Reis
                             ` (3 more replies)
  2026-10-08 15:55         ` Junio C Hamano
  1 sibling, 4 replies; 21+ messages in thread
From: D. Ben Knoble @ 2026-10-08 13:25 UTC (permalink / raw)
  To: Sam Reis; +Cc: Sebastian Thiel, Scott Chacon, Junio C Hamano, Scott Chacon, git

Hi Sam,

On Thu, Oct 8, 2026 at 7:23 AM Sam Reis <sam@opencanopy.dev> wrote:
>
> Hey everyone. Just for the avoidance of doubt, very happy to see
> Scott's patch here land and for git to benefit from faster sha1dc
> hashing. Let me know if I can do anything to support!

[We bottom-post here]

> > On 07.10.26 20:13, Scott Chacon wrote:
> > > Hey,
> > >
> > > On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
> > >>> This series ports the approach of Sam Reis's sha1dc Rust crate [1],
> > >>> which gitoxide recently switched to [2], to C.
> > >>
> > >> Which means license-wise the original is compatible with us, I
> > >> presume, as they are "Apache2 or MIT, your choice".

Leaving aside the question below about AI policy, I think the
important question for Sam and Sebastien is license compatibility?

The sha1collisiondetection submodule and the sha1dc code (extracted
from that submodule's upstream, if I'm reading 28dc98e343 (sha1dc: add
collision-detecting sha1 implementation, 2017-03-16) correctly?) are
MIT licensed, too, so there is some precedent for Git here. I skimmed
what I could find of the original threads:

- https://lore.kernel.org/git/20170223195753.ppsat2gwd3jq22by@sigill.intra.peff.net/
- https://lore.kernel.org/git/?q=sha1dc%3A+add+collision-detecting+sha1+implementation

but I didn't see a discussion of licensing at that time. Perhaps the
idea is that we are clear that such code carries a different license
from Git?

Anyway, I suppose the fair thing would then be for Scott's code to be
MIT (and/or Apache2), in which case it would need similar
clarifications? (Or are we prepared to take the stance that de nouveau
code based on existing code can be license-washed, in this case to
GPL-2?)

Interestingly, Gentoo claims Git's license is only GPL-2, but I think
they compile in the sha1dc code since it's the default in meson.
Should we be claiming the Git package (with sha1dc) is actually GPL-2
and MIT?

(This is complex territory and I'm sure to have gotten it wrong;
pointers to past discussions, esp. those by copyright and licensing
professionals, welcome.)

> > >> How can you/we be sure, with respect to the current AI policy in
> > >> SubmittingPatches (which by the way was vetted by SFC lawyers), that
> > >> your "AI generated" code did not "borrow" from places that gets
> > >> you/us into trouble?
> > >
> > > It's a good question. I actually just submitted a proposed update to
> > > that policy based on SFC's updated guidelines, but either way, I
> > > learned about this from Sam and have talked to him about the port and
> > > he seemed excited about it. I can triple check, but I'm fairly
> > > confident that he's fine with this and I am fine signing off on it
> > > under the terms of the DCO language.
> > >
> > > Of course, he in turn used AI tooling to produce _his_ library, but
> > > within the guidelines of the updated SFC guidelines. Johannes's
> > > alternative series is the original Rust code of Sam that my agent
> > > looked at to produce this (in addition to his blog post explaining
> > > it), so I'm not sure how that might be materially different.

[snip]

> > >>> [1] https://sam.dev/blog/faster-sha1-collision-detection
> > >>> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008


-- 
D. Ben Knoble

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 13:25         ` D. Ben Knoble
@ 2026-10-08 13:59           ` Sam Reis
  2026-10-08 17:10           ` Junio C Hamano
                             ` (2 subsequent siblings)
  3 siblings, 0 replies; 21+ messages in thread
From: Sam Reis @ 2026-10-08 13:59 UTC (permalink / raw)
  To: D. Ben Knoble
  Cc: Sebastian Thiel, Scott Chacon, Junio C Hamano, Scott Chacon, git

Hi Ben,

On Thu, Oct 8, 2026 at 3:25 PM D. Ben Knoble <ben.knoble@gmail.com> wrote:
>
> Hi Sam,
>
> On Thu, Oct 8, 2026 at 7:23 AM Sam Reis <sam@opencanopy.dev> wrote:
> >
> > Hey everyone. Just for the avoidance of doubt, very happy to see
> > Scott's patch here land and for git to benefit from faster sha1dc
> > hashing. Let me know if I can do anything to support!
>
> [We bottom-post here]

No problem!

>
> > > On 07.10.26 20:13, Scott Chacon wrote:
> > > > Hey,
> > > >
> > > > On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
> > > >>> This series ports the approach of Sam Reis's sha1dc Rust crate [1],
> > > >>> which gitoxide recently switched to [2], to C.
> > > >>
> > > >> Which means license-wise the original is compatible with us, I
> > > >> presume, as they are "Apache2 or MIT, your choice".
>
> Leaving aside the question below about AI policy, I think the
> important question for Sam and Sebastien is license compatibility?

I can try answering as best as possible from my perspective. I think
it's important
to note, insofar this wasn't clear before, that the actual code for
the ubc checks is
fully autogenerated rather than written by hand.

The generator for it was written from scratch by me with LLM
assistance, and based
on the methodology described in the 2017 paper from Marc Stevens and Dan Shumow.
Nonetheless I ended up crediting Stevens and Shumow in the copyright notice
to be on the safe side.

>
> The sha1collisiondetection submodule and the sha1dc code (extracted
> from that submodule's upstream, if I'm reading 28dc98e343 (sha1dc: add
> collision-detecting sha1 implementation, 2017-03-16) correctly?) are
> MIT licensed, too, so there is some precedent for Git here. I skimmed
> what I could find of the original threads:
>
> - https://lore.kernel.org/git/20170223195753.ppsat2gwd3jq22by@sigill.intra.peff.net/
> - https://lore.kernel.org/git/?q=sha1dc%3A+add+collision-detecting+sha1+implementation
>
> but I didn't see a discussion of licensing at that time. Perhaps the
> idea is that we are clear that such code carries a different license
> from Git?
>
> Anyway, I suppose the fair thing would then be for Scott's code to be
> MIT (and/or Apache2), in which case it would need similar
> clarifications? (Or are we prepared to take the stance that de nouveau
> code based on existing code can be license-washed, in this case to
> GPL-2?)
>
> Interestingly, Gentoo claims Git's license is only GPL-2, but I think
> they compile in the sha1dc code since it's the default in meson.
> Should we be claiming the Git package (with sha1dc) is actually GPL-2
> and MIT?

Again IANAL, but my understanding was always that MIT is GPL-compatible.
From https://en.wikipedia.org/wiki/License_compatibility#GPL_compatibility:

"Many of the most common free-software licenses [...] are GPL-compatible.
That is, their code can be combined with a program under the GPL without
conflict, and the new combination would have the GPL applied to the whole
(but the other license would not so apply)."

So as long as it's correctly attributed, that might not be an issue?

>
> (This is complex territory and I'm sure to have gotten it wrong;
> pointers to past discussions, esp. those by copyright and licensing
> professionals, welcome.)
>
> > > >> How can you/we be sure, with respect to the current AI policy in
> > > >> SubmittingPatches (which by the way was vetted by SFC lawyers), that
> > > >> your "AI generated" code did not "borrow" from places that gets
> > > >> you/us into trouble?
> > > >
> > > > It's a good question. I actually just submitted a proposed update to
> > > > that policy based on SFC's updated guidelines, but either way, I
> > > > learned about this from Sam and have talked to him about the port and
> > > > he seemed excited about it. I can triple check, but I'm fairly
> > > > confident that he's fine with this and I am fine signing off on it
> > > > under the terms of the DCO language.
> > > >
> > > > Of course, he in turn used AI tooling to produce _his_ library, but
> > > > within the guidelines of the updated SFC guidelines. Johannes's
> > > > alternative series is the original Rust code of Sam that my agent
> > > > looked at to produce this (in addition to his blog post explaining
> > > > it), so I'm not sure how that might be materially different.
>
> [snip]
>
> > > >>> [1] https://sam.dev/blog/faster-sha1-collision-detection
> > > >>> [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
>
>
> --
> D. Ben Knoble

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 11:20       ` Sam Reis
  2026-10-08 13:25         ` D. Ben Knoble
@ 2026-10-08 15:55         ` Junio C Hamano
  1 sibling, 0 replies; 21+ messages in thread
From: Junio C Hamano @ 2026-10-08 15:55 UTC (permalink / raw)
  To: Sam Reis; +Cc: Sebastian Thiel, Scott Chacon, Scott Chacon, git

Sam Reis <sam@opencanopy.dev> writes:

> Hey everyone. Just for the avoidance of doubt, very happy to see
> Scott's patch here land and for git to benefit from faster sha1dc
> hashing. Let me know if I can do anything to support!

Thanks.  Just to make sure I understand, are you endorsing the idea
of your work geting ported to help Git, or are you also happy with
the actual code Scott submitted?

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 13:25         ` D. Ben Knoble
  2026-10-08 13:59           ` Sam Reis
@ 2026-10-08 17:10           ` Junio C Hamano
  2026-10-08 17:46             ` D. Ben Knoble
  2026-10-09 20:37           ` Todd Zullinger
  2026-10-09 23:41           ` Junio C Hamano
  3 siblings, 1 reply; 21+ messages in thread
From: Junio C Hamano @ 2026-10-08 17:10 UTC (permalink / raw)
  To: D. Ben Knoble; +Cc: Sam Reis, Sebastian Thiel, Scott Chacon, Scott Chacon, git

"D. Ben Knoble" <ben.knoble@gmail.com> writes:

> but I didn't see a discussion of licensing at that time. Perhaps the
> idea is that we are clear that such code carries a different license
> from Git?

We are GPL-2 only, which means we can incorporate BSD-licensed
software as long as we satisfy its license and copyright notice
requirements.

The above is not an AI-bot-supplied answer, but what one learns when
talking to copyright lawyers or reading books on software licensing.

But sometimes asking LLM gives sufficiently useful answer.  I typed

"Is GPLv2 compatible with BSD?"

in the search bar of a browser, and here is the early part of what I
got, which is not too bad.

    * AI summary

    Yes, GPLv2 is generally compatible with modern BSD (2-clause and
    3-clause) licenses, though the direction of the combination
    matters.

    Compatibility Details

    • BSD inside GPLv2: You can include a 2-clause or 3-clause
      BSD-licensed library or code snippet inside a GPLv2-licensed
      project. The resulting combined work must be distributed under
      the terms of the GPLv2.

    • GPLv2 inside BSD: You cannot take GPLv2-licensed code and
      place it into a purely BSD-licensed project. Because the GPLv2
      is a strong copyleft license, it forces the entire combined
      work to be covered by the GPLv2.


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 17:10           ` Junio C Hamano
@ 2026-10-08 17:46             ` D. Ben Knoble
  2026-10-08 21:03               ` Junio C Hamano
  0 siblings, 1 reply; 21+ messages in thread
From: D. Ben Knoble @ 2026-10-08 17:46 UTC (permalink / raw)
  To: Junio C Hamano; +Cc: Sam Reis, Sebastian Thiel, Scott Chacon, Scott Chacon, git

On Thu, Oct 8, 2026 at 1:10 PM Junio C Hamano <gitster@pobox.com> wrote:
>
> "D. Ben Knoble" <ben.knoble@gmail.com> writes:
>
> > but I didn't see a discussion of licensing at that time. Perhaps the
> > idea is that we are clear that such code carries a different license
> > from Git?
>
> We are GPL-2 only, which means we can incorporate BSD-licensed
> software as long as we satisfy its license and copyright notice
> requirements.
>
> The above is not an AI-bot-supplied answer, but what one learns when
> talking to copyright lawyers or reading books on software licensing.

Thanks, that's good to know---but in this case I thought we were
talking about the MIT license?

Assuming a similar analysis applies (not clear to me, but not
implausible either), that might also answer my question about Gentoo's
license descriptor. The product is GPL-2 even if one input was MIT
(though it feels strange to effectively "re-license" someone else's
code this way).

-- 
D. Ben Knoble

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 17:46             ` D. Ben Knoble
@ 2026-10-08 21:03               ` Junio C Hamano
  0 siblings, 0 replies; 21+ messages in thread
From: Junio C Hamano @ 2026-10-08 21:03 UTC (permalink / raw)
  To: D. Ben Knoble; +Cc: Sam Reis, Sebastian Thiel, Scott Chacon, Scott Chacon, git

"D. Ben Knoble" <ben.knoble@gmail.com> writes:

>> We are GPL-2 only, which means we can incorporate BSD-licensed
>> software as long as we satisfy its license and copyright notice
>> requirements.
>>
>> The above is not an AI-bot-supplied answer, but what one learns when
>> talking to copyright lawyers or reading books on software licensing.
>
> Thanks, that's good to know---but in this case I thought we were
> talking about the MIT license?

Yup, the story is the same.  Also BSD and MIT also fully compatible
in either direction.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 13:25         ` D. Ben Knoble
  2026-10-08 13:59           ` Sam Reis
  2026-10-08 17:10           ` Junio C Hamano
@ 2026-10-09 20:37           ` Todd Zullinger
  2026-10-09 23:41           ` Junio C Hamano
  3 siblings, 0 replies; 21+ messages in thread
From: Todd Zullinger @ 2026-10-09 20:37 UTC (permalink / raw)
  To: D. Ben Knoble
  Cc: Sam Reis, Sebastian Thiel, Scott Chacon, Junio C Hamano,
	Scott Chacon, git

D. Ben Knoble wrote:
> Interestingly, Gentoo claims Git's license is only GPL-2, but I think
> they compile in the sha1dc code since it's the default in meson.
> Should we be claiming the Git package (with sha1dc) is actually GPL-2
> and MIT?

I _think_ that that depends on whether Gentoo's license tag
is meant to be the "effective" license they are distributing
their Git package or attempting to encompass the license of
all of the code which goes into the Git package.

Fedora's Git package has:

    BSD-3-Clause AND GPL-2.0-only AND GPL-2.0-or-later AND LGPL-2.1-or-later AND MIT

I think I was the last to touch that, in ef75bcd (update
license data and convert to SPDX format, 2022-11-07).  I
attempted to capture all of the licensing, but I most
certainly could have missed some things.  And things may
have changed since then too.

Fedora's guidelines have a lengthy page on the License tag².
It includes a 'No "effective license" analysis' section
discussing this.

I don't know if any of that helps. :)

¹ https://src.fedoraproject.org/rpms/git/c/ef75bcd
² https://docs.fedoraproject.org/en-US/legal/license-field/

-- 
Todd

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH 0/4] faster SHA-1 collision detection
  2026-10-08 13:25         ` D. Ben Knoble
                             ` (2 preceding siblings ...)
  2026-10-09 20:37           ` Todd Zullinger
@ 2026-10-09 23:41           ` Junio C Hamano
  3 siblings, 0 replies; 21+ messages in thread
From: Junio C Hamano @ 2026-10-09 23:41 UTC (permalink / raw)
  To: D. Ben Knoble; +Cc: Sam Reis, Sebastian Thiel, Scott Chacon, Scott Chacon, git

"D. Ben Knoble" <ben.knoble@gmail.com> writes:

> The sha1collisiondetection submodule and the sha1dc code (extracted
> from that submodule's upstream, if I'm reading 28dc98e343 (sha1dc: add
> collision-detecting sha1 implementation, 2017-03-16) correctly?) are
> MIT licensed, too, so there is some precedent for Git here. I skimmed
> what I could find of the original threads:
>
> - https://lore.kernel.org/git/20170223195753.ppsat2gwd3jq22by@sigill.intra.peff.net/
> - https://lore.kernel.org/git/?q=sha1dc%3A+add+collision-detecting+sha1+implementation
>
> but I didn't see a discussion of licensing at that time. Perhaps the
> idea is that we are clear that such code carries a different license
> from Git?
>
> Anyway, I suppose the fair thing would then be for Scott's code to be
> MIT (and/or Apache2), in which case it would need similar
> clarifications? (Or are we prepared to take the stance that de nouveau
> code based on existing code can be license-washed, in this case to
> GPL-2?)
>
> Interestingly, Gentoo claims Git's license is only GPL-2, but I think
> they compile in the sha1dc code since it's the default in meson.
> Should we be claiming the Git package (with sha1dc) is actually GPL-2
> and MIT?

In the abov, Gentoo's mention is about "Git package" as a whole.
Git package as a whole can be distributed under GPLv2 only.

MIT, BSD-2 or BSD-3 are permissive and essentially says "you can do
whatever you want with the code (including combining with other code
or making it proprietary), as long as you keep our copyright notice,
keep our disclaimer, and (in the case of BSD-3) do not use our names
for endorsement".  Specifically, they do not forbid us from
incorporating their ware into our project that is licensed
differently, as long as we honor their licensing terms on the source
files we got from them.

Because we have mixed "permissive" code into GPLv2 code to form a
single "work based on the Program", GPLv2 Section 2(b) dictates that
the entire combined work must be distributed under the terms of the
GPLv2 (and again, the permissiveness of "other" licenses is what
allows us to do so).  You cannot distribute the finished binary or
the combined sources under a permissive license, because doing so
would violate the GPLv2's copyleft requirement.

The original "permissively licensed" files (and any modifications
made purely to those files) still maintain their original copyright
headers and original "permissive" license text.  This is because the
original copyright holder of the code granted a license to use their
files under the original "permissive" licensing terms, which
requires us to keep their copyright notice.  We do not own the
copyright to the original "permissive" code.  We are only licensed
to use them.  So we have no legal authority to strip these
"permissive" licenses or unilaterally "relicense" those files into
GPLv2.

So to answer your question in the last sentence, we should say "Git
package as a whole is GPLv2 only, but parts are borrowed from
copyright holders who licensed them under different terms, and these
parts can be used under these different parts.  For example, sha1dc
can be copied from our source tree to your non GPLv2 project as long
as you honor their MIT license".


^ permalink raw reply	[flat|nested] 21+ messages in thread

end of thread, other threads:[~2026-10-09 23:41 UTC | newest]

Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 11:25 [PATCH 0/4] faster SHA-1 collision detection Scott Chacon
2026-09-29 11:25 ` [PATCH 1/4] sha1dc-accel: add a block loop for sha1dc's SHA1_CTX Scott Chacon
2026-10-07 12:17   ` Johannes Schindelin
2026-09-29 11:25 ` [PATCH 2/4] sha1dc-accel: vectorize the unavoidable-bitconditions check Scott Chacon
2026-10-07 12:17   ` Johannes Schindelin
2026-10-07 21:32     ` Junio C Hamano
2026-09-29 11:25 ` [PATCH 3/4] sha1dc-accel: compress with SHA-NI on x86-64 Scott Chacon
2026-09-29 11:25 ` [PATCH 4/4] sha1dc-accel: compress with the ARMv8 SHA-1 instructions Scott Chacon
2026-10-07 12:17 ` [PATCH 0/4] faster SHA-1 collision detection Johannes Schindelin
2026-10-07 17:23 ` Junio C Hamano
2026-10-07 18:13   ` Scott Chacon
2026-10-08  6:21     ` Sebastian Thiel
2026-10-08 11:20       ` Sam Reis
2026-10-08 13:25         ` D. Ben Knoble
2026-10-08 13:59           ` Sam Reis
2026-10-08 17:10           ` Junio C Hamano
2026-10-08 17:46             ` D. Ben Knoble
2026-10-08 21:03               ` Junio C Hamano
2026-10-09 20:37           ` Todd Zullinger
2026-10-09 23:41           ` Junio C Hamano
2026-10-08 15:55         ` Junio C Hamano

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox