From: Taylor Blau <ttaylorr@openai.com>
To: git@vger.kernel.org
Subject: [NOTES 07/07] Protocol v2 for pushes
Date: Tue, 6 Oct 2026 11:06:38 -0700 [thread overview]
Message-ID: <summit-2026.94e33e9ddf234334.07@ttaylorr.com> (raw)
In-Reply-To: <summit-2026.94e33e9ddf234334.00@ttaylorr.com>
Topic: Protocol v2 for pushes
Notetaker: Emily
* brian: Some customers have millions of refs. Reftable has helped us
push up to 100k at a time, which was a big increase. AI means more
branches and refs, so the problem keeps getting worse. One customer
was unhappy about an 896 MB ref advertisement. Protocol v2 for pushes
would help. Much of the code already exists and needs wiring up.
* Peff: We were waiting for somebody with a use case they care about.
Protocol v2 has a capability advertisement that leaves room for
further extensions.
* Elijah: With fetch v2, it would be useful to request only a handful of
heads rather than advertise all refs and tags.
* Jonathan Tan: We have prefix filtering for fetch v2. A client can
request refs/heads/somename, and the server sends refs matching that
prefix.
* Peff: The client needs a better refspec, and the user interface is
hard. How do we specify the mainline? Even excluding HEAD can speed
things up. This is an optimization; a push is correct as long as the
objects are reachable.
* Peff: Could we use multiple passes, trying a smaller set and falling
back to a larger one if some objects are still unreachable? Many
advertised refs are not useful, such as abandoned development
branches.
* [There was a related discussion about showing fewer refs in GitHub's
UI.]
* Peff: Could a server-side configuration limit the advertisement and UI
to a smaller set of refs, with the client ignoring the others?
* brian: It is useful to control what the client requests. During a
push, we want to say that we care about only a handful of branches.
* Peff: That still sounds like a refspec problem; the default asks for
too much.
* Jonathan Tan: This is push, though.
* Peff: The client could suggest useful branches more intelligently.
Guessing the main branches is a heuristic, but upstream tracking
branches might be useful.
* brian: A project setup script, perhaps linked from its documentation,
could provide useful initial configuration.
* Jonathan Tan: Instead of a ref advertisement during push, it would be
useful to negotiate. Restricting the advertised refs may cause the
client to send everything if the optimization does not work. Relying
on the remote to know what matters is error-prone. We would be happier
with a few round trips: advertise capabilities, exchange haves and
acknowledgments, negotiate the push, and send the pack.
* Elijah: During a repack, the server might give a false negative. This
may be relevant to shallow clones.
* Peff: We could race and miss a common commit, then be unable to look
further back because the client is shallow.
* Elijah: Geometric repacking may help because it is less likely to miss
something. A push to the wrong place should fail quickly. Let us leave
shallow clones to develop their own story.
* Peff: Successful negotiation could also make connectivity checks
easier. We could cache information in receive-pack and pass it to
check-connected.
* brian: Shallow clones can spend much more time trying to minimize the
wire transfer. One push went from two seconds to 35 seconds because of
that optimization. `objects-edge-aggressive` is relevant here.
* Jonathan Tan: We might get substantial savings from simply disabling
ref advertisements with a configuration option.
* Peff: Even without a protocol change, the server could decide which
refs are interesting and limit its advertisement. That might be easier
than introducing v2 for pushes.
* Jonathan Tan: The client needs to tell the server that it wants to
negotiate.
* Peff: It might not need to negotiate.
* Jonathan Tan: Would that not be bad?
* Peff: In practice the drawback may be small, though completely
unrelated histories would be a bad case.
* Elijah: That can happen when appending a shallow commit.
* brian: I have seen that happen too.
* Peff: In that degenerate case, do these optimizations already have
problems?
* Jonathan Tan: Negotiation does not have as much trouble because it
starts at the tip the client is pushing.
* [General agreement that putting unrelated histories into one
repository is best avoided.]
* Peff: Nobody objects to push v2. We have the hooks and the fetch-v2
infrastructure. There may be an easier place to start, though.
* brian: Another benefit of push v2 is interoperability. Currently a
client must push using the server's hash algorithm; it cannot
negotiate a different one. Fetch can do some negotiation.
* Peff: If v2 makes interoperability easier, go for it.
* Emily: What is the failure mode?
* brian: Without it, the client must know the server's main hash
algorithm and the mapping from object IDs in that algorithm to
content.
* Martin Fick: I would like push v2 for automated replication, where a
forge pushes to mirrors. Gerrit does this with thousands of targets.
Ref advertisements are expensive when checking all of them. We need a
way to know whether an advertisement differs from ours, so we can skip
targets that are already up to date.
* Emily: Push negotiation?
* Martin Fick: The ref advertisement happens before negotiation.
* Peff: You want an answer in tens of bytes rather than thousands. A
checksum could provide a small initial step.
* brian: With reftable, generation numbers make this easy.
* Martin Fick: That might work for this case, but assumes the
repositories are in lockstep and covers fewer cases than a full
protocol change.
* Peff: Would the rest of the push not fix it?
* Martin Fick: Parallelism makes that assumption harder.
* Peff: We discussed an ETag-style approach with reftable at GitHub
years ago. It would be useful to see an implementation, and it could
fit as an option in the existing protocol.
* Martin Fick: The server could also send a diff or a leaner pack.
* [Caching was also mentioned.]
* Patrick: Should we shrink the advertisement format?
* Peff: Perhaps compress it with zlib.
* Patrick: Let us explore options and benchmark them. We need v2 on the
push side. We also proposed sending reftable directly.
* Martin Fick: A reftable could contain extra information.
* Patrick: A compressed format may be simpler.
* brian: We also do not want to send hidden refs.
* Patrick: I mean using the reftable format, rather than sending the
whole reftable.
prev parent reply other threads:[~2026-10-06 18:06 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 18:05 Notes from the Git Contributor's Summit, 2026 Taylor Blau
2026-10-06 18:06 ` [NOTES 01/07] Security mailing list and security process Taylor Blau
2026-10-06 18:06 ` [NOTES 02/07] Git 3.0 Taylor Blau
2026-10-06 18:06 ` [NOTES 03/07] Documentation Taylor Blau
2026-10-07 4:49 ` Todd Zullinger
2026-10-07 17:38 ` Junio C Hamano
2026-10-06 18:06 ` [NOTES 04/07] Outreachy sponsorship Taylor Blau
2026-10-06 18:06 ` [NOTES 05/07] What can we do next with pluggable ODB? Taylor Blau
2026-10-06 18:06 ` [NOTES 06/07] AI contribution policy Taylor Blau
2026-10-06 18:06 ` Taylor Blau [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=summit-2026.94e33e9ddf234334.07@ttaylorr.com \
--to=ttaylorr@openai.com \
--cc=git@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox