From: "brian m. carlson" <sandals@crustytoothpaste.net>
To: Christian Couder <christian.couder@gmail.com>
Cc: Scott Chacon <scott@gitbutler.net>,
git@vger.kernel.org, Patrick Steinhardt <ps@pks.im>
Subject: Re: [RFC PATCH 0/4] sign a SHA-256 digest of the tree in commits and tags
Date: Tue, 6 Oct 2026 22:26:40 +0000 [thread overview]
Message-ID: <asV1oB_avuEbgRVe@fruit.crustytoothpaste.net> (raw)
In-Reply-To: <CAP8UFD096CdR9MXd+VHk7Zf9rCJEnGTEiBhCc0mJdMmE3U_gOg@mail.gmail.com>
[-- Attachment #1: Type: text/plain, Size: 7052 bytes --]
On 2026-10-06 at 09:00:45, Christian Couder wrote:
> On Fri, Oct 2, 2026 at 9:15 PM brian m. carlson
> <sandals@crustytoothpaste.net> wrote:
>
> > The thing you really want is the interoperability work, which can
> > automatically rewrite repositories from one hash algorithm to another
> > during a clone or fetch operation. Yes, it isn't quite that simple for
> > submodules, but if you recursively clone the repository and all its
> > submodules, it should be possible to rewrite it in place, although that
> > hasn't been written yet. That work has not yet been sent upstream
> > because some of it was written at $DAYJOB, which requires that we use
> > Outlook and we all know that Outlook corrupts patches. However, there
> > is some intention for another company to handle the polishing and
> > sending, so it should be available sooner or later.
>
> Sorry for the possibly stupid following questions, but I think the
> answers might help us get a better idea of what might be needed to get
> a smoother transition.
>
> And yeah, I know that many people have said that merging all your
> interoperability work should not block Git 3.0. But if it can ensure a
> smoother transition, we might want to get at least part of it merged
> soon, and the rest in a good shape, anyway.
>
> Is the current state of the work publicly available somewhere? Or
> could you make it publicly available somewhere? (Fine if it's only as
> patches in a tarball.)
Yeah, it's at https://github.com/bk2204/git.git as `sha256-interop`.
> Is the submodule work the only missing part of the interoperability work?
Not quite. The limitations are outlined in
https://lore.kernel.org/git/ajCWBG9RHBrm8jMZ@fruit.crustytoothpaste.net/
and in `Documentation/gitformat-hash.adoc` in the branch.
There's no in-place migration tooling, which would be required for
dealing with submodules, since those would need to be migrated before
the main repository. There are essentially two cases for in-place
migration: adding the other hash as an extra mapping and rewriting all
the objects to the other hash with our hash as a mapping, much like `git
refs migrate` does.
I've started on some of the Rust pieces that I was planning to use for
in-place migration, but someone who wanted to work on a C-based
implementation could also do that.
There's some missing features that I've outlined, like a lack of support
for multi-pack index. Those can be added, but they're not immediately
essential. And there are other things which need to be actually
polished and fixed, like the fact that delta resolution is recursive
rather than iterative (which would allow a malicious server to cause
stack exhaustion). The good news is that most of those things are
labelled with `WIP` and a short description of what needs to be fixed in
the commit message.
One other thing that needs fixing is that we have many pieces of code
that write large numbers of loose objects into the ODB and may randomly
die in places; `git add` is a great example of this. The object maps
which are used for storing loose objects should ideally be written with
all of those objects at once in a batch and then committed. However, if
we die at any point, we've written the loose objects, but not the
batched object map, so the repository is then corrupt since it can't
perform mappings. If we write the objects into the object map one at a
time, then we end up with N object maps and `git gc` runs all the time
to repack those into a smaller set of data. We therefore need to either
fix the die-die-die behaviour or use ODB transactions to write both
loose objects and the object maps into a temporary directory. This is a
case where it technically works but it performs awfully, so we do need
to fix it before non-experimental use.
There is a partial rebase of the early entries in the series converted
to use the pluggable ODB work in the `sha256-interop-part-2` branch. The
pluggable ODB work has caused a lot of conflicts in the interop because,
unsurprisingly, both series are intimately involved in the object
database.
> How much work is this? (At one point it seemed to me that it was
> around 200 patches.)
It's presently about 212 patches.
> If you were to work full time on upstreaming it, how long would you
> expect it would take you?
Probably four release cycles, assuming release cycles are 6 weeks. The
reason is that reviews will be needed and those will take time.
If we wanted to write an in-place migration helper, I'd expect another
two cycles. Writing one that preserved the existing algorithm and just
added the mapping for the other algorithm would be easier because it
wouldn't require rewriting the `objects` directory.
> If some of us could help you, how could we best help?
I would love someone to start picking up patches from the early part of
the series, rebasing them onto `master`, fixing up any conflicts, and
polishing them, and then sending them in. The `sha256-interop-part-2`
series would be great for that.
Just let me know if you want to do this and then we won't conflict.
> Could you say which company is interested in helping with this? Would
> that company be willing to work openly with others on this?
I'd rather not disclose that without the permission of the person making
the offer. I'll just say that a respected contributor and member of the
community offered to have some of the work done on their company's dime
by a person who is also known to the list. They did note that there
would be a delay before starting, so it wouldn't happen right away.
I feel confident that the contributor in question would be willing to
collaborate with others in getting this work done because that's the
kind of person they are and obviously it would be in everyone's interest
to do that.
> Are there some tests or kinds of automated ways to check that things
> work as expected under realistic conditions like:
>
> - using real world repos (large ones, old ones, with submodules, etc),
> - mixing a number of new and old clients and servers,
> - interacting with other implementations (JGit, libgit2, gitoxide,
> forges, CI, etc)?
What I have done to test this is `git clone --object-format=sha256:sha1
https://github.com/bk2204/lawn.git` and then pushed to a SHA-256
repository. That's just a personal project of mine that doesn't contain
submodules, but it's what we've used for testing at work and we have
several internal copies of the SHA-256 version of that repo. It clearly
interoperates with GitHub on a SHA-1-only and SHA-256-only basis, but
there's no support for the interoperability on the server-side yet.
There are also tests for the interoperability code in t1017 which are
reasonably comprehensive.
Once we have in-place migration, I would like to test it with git.git,
since I think that would be a great real-world testcase.
--
brian m. carlson (they/them)
Toronto, Ontario, CA
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]
next prev parent reply other threads:[~2026-10-06 22:26 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 8:18 [RFC PATCH 0/4] sign a SHA-256 digest of the tree in commits and tags Scott Chacon
2026-10-02 8:18 ` [RFC PATCH 1/4] tree-sha256: hash the contents of a tree with SHA-256 Scott Chacon
2026-10-02 15:45 ` Junio C Hamano
2026-10-02 8:18 ` [RFC PATCH 2/4] tag: add --hash=sha256 to sign a tree-sha256 header Scott Chacon
2026-10-02 15:49 ` Junio C Hamano
2026-10-02 8:18 ` [RFC PATCH 3/4] commit: " Scott Chacon
2026-10-02 8:18 ` [RFC PATCH 4/4] gpg: add gpg.treeHash to sign a tree-sha256 header by default Scott Chacon
2026-10-02 15:52 ` [RFC PATCH 0/4] sign a SHA-256 digest of the tree in commits and tags Junio C Hamano
2026-10-02 19:06 ` brian m. carlson
2026-10-05 9:32 ` Scott Chacon
2026-10-05 12:41 ` Patrick Steinhardt
2026-10-05 14:16 ` Scott Chacon
2026-10-05 22:57 ` brian m. carlson
2026-10-06 13:36 ` Johannes Schindelin
2026-10-06 16:16 ` Kristoffer Haugsbakk
2026-10-06 21:55 ` brian m. carlson
2026-10-06 22:38 ` Junio C Hamano
2026-10-06 23:40 ` brian m. carlson
2026-10-06 9:00 ` Christian Couder
2026-10-06 22:26 ` brian m. carlson [this message]
2026-10-07 12:26 ` Christian Couder
2026-10-07 21:07 ` brian m. carlson
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asV1oB_avuEbgRVe@fruit.crustytoothpaste.net \
--to=sandals@crustytoothpaste.net \
--cc=christian.couder@gmail.com \
--cc=git@vger.kernel.org \
--cc=ps@pks.im \
--cc=scott@gitbutler.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox