Git development
 help / color / mirror / Atom feed
* [PATCH] index-pack: speed up promisor link recording
@ 2026-08-02 21:33 Arijit Banerjee via GitGitGadget
  2026-08-02 21:51 ` brian m. carlson
  0 siblings, 1 reply; 9+ messages in thread
From: Arijit Banerjee via GitGitGadget @ 2026-08-02 21:33 UTC (permalink / raw)
  To: git
  Cc: Jonathan Tan, Patrick Steinhardt, Junio C Hamano, Arijit Banerjee,
	Arijit Banerjee

From: Arijit Banerjee <arijit@effectiveailabs.com>

When indexing a promisor pack, index-pack parses every reconstructed
non-blob object into the shared object model to record its outgoing links.
Since parse_object_buffer() runs under read_mutex, worker threads serialize
while allocating persistent tree, commit, and tag structures that are only
needed to enumerate those links.

Read the links directly from the reconstructed object buffers instead. Keep
the strict and fsck paths unchanged, use worker-local typed oidmaps during
normal promisor indexing, and merge them after the workers exit. Transfer
entries during the merge so that it does not temporarily duplicate the
complete link set.

The typed entries preserve checks previously performed as a side effect of
object parsing. Reject malformed commit and tag headers, conflicting
expected types, and targets whose actual type disagrees when the target is
present in the pack. Preserve commit-graft handling and the existing policy
of recording only subtree entries from trees.

With three runs per version on Debian 12, median end-to-end wall-clock time
for a --filter=blob:none clone of linux.git decreased from 156 seconds to
133 seconds (15%). Trace2 attributed the change to the initial index-pack
--promisor phase, whose median duration decreased from 121 seconds to 98
seconds (19%). System CPU time decreased by 46%.

Two paired spot checks against GitHub showed end-to-end reductions of 18%
and 26%. These measurements include network and server variability and are
therefore corroborating rather than controlled results. A third pair was not
interpretable because the baseline request encountered a transport stall.

A full-clone control showed no material change, taking approximately 256
seconds with either version. This is expected because full clones do not
exercise promisor-link recording.

t5302-pack-index.sh passed with both SHA-1 and SHA-256, while
t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage
checks malformed commit headers, conflicting link types, and mismatched tag
target types.

Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
---
    index-pack: speed up promisor link recording
    
    AI assistance: OpenAI Codex was used to identify the bottleneck and
    assist with the implementation, testing, and benchmark analysis. I
    reviewed the resulting change and take responsibility for this
    submission.

Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-2191%2Farijit91%2Findex-pack-promisor-link-recording-v1
Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2191/arijit91/index-pack-promisor-link-recording-v1
Pull-Request: https://github.com/gitgitgadget/git/pull/2191

 builtin/index-pack.c  | 234 ++++++++++++++++++++++++++++++++++++++----
 t/t5302-pack-index.sh |  50 +++++++++
 2 files changed, 265 insertions(+), 19 deletions(-)

diff --git a/builtin/index-pack.c b/builtin/index-pack.c
index bc86925ad0..a973146757 100644
--- a/builtin/index-pack.c
+++ b/builtin/index-pack.c
@@ -23,7 +23,8 @@
 #include "odb.h"
 #include "odb/streaming.h"
 #include "oid-array.h"
-#include "oidset.h"
+#include "hash-lookup.h"
+#include "oidmap.h"
 #include "path.h"
 #include "replace-object.h"
 #include "tree-walk.h"
@@ -105,6 +106,12 @@ static size_t base_cache_limit;
 struct thread_local_data {
 	pthread_t thread;
 	int pack_fd;
+	struct oidmap outgoing_links;
+};
+
+struct outgoing_link {
+	struct oidmap_entry entry;
+	enum object_type type;
 };
 
 /* Remember to update object flag allocation in object.h */
@@ -155,11 +162,8 @@ static uint32_t input_crc32;
 static int input_fd, output_fd;
 static const char *curr_pack;
 
-/*
- * outgoing_links is guarded by read_mutex, and record_outgoing_links is
- * read-only in a thread.
- */
-static struct oidset outgoing_links = OIDSET_INIT;
+/* Worker-local maps are merged after all workers have exited. */
+static struct oidmap outgoing_links = OIDMAP_INIT;
 static int record_outgoing_links;
 
 static struct thread_local_data *thread_data;
@@ -196,6 +200,55 @@ static inline void unlock_mutex(pthread_mutex_t *mutex)
 		pthread_mutex_unlock(mutex);
 }
 
+static void record_outgoing_link_to(struct oidmap *map,
+				    const struct object_id *oid,
+				    enum object_type type)
+{
+	struct outgoing_link *link = oidmap_get(map, oid);
+
+	if (link) {
+		if (type != OBJ_ANY && link->type != OBJ_ANY &&
+		    type != link->type)
+			die(_("object %s is referred to as both a %s and a %s"),
+			    oid_to_hex(oid), type_name(link->type),
+			    type_name(type));
+		if (link->type == OBJ_ANY)
+			link->type = type;
+		return;
+	}
+
+	CALLOC_ARRAY(link, 1);
+	oidcpy(&link->entry.oid, oid);
+	link->type = type;
+	if (oidmap_put(map, link))
+		BUG("duplicate outgoing link");
+}
+
+static void merge_outgoing_links(struct oidmap *dest, struct oidmap *src)
+{
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
+
+	while ((link = oidmap_iter_first(src, &iter))) {
+		struct outgoing_link *old = oidmap_get(dest, &link->entry.oid);
+
+		oidmap_remove(src, &link->entry.oid);
+		if (old) {
+			if (link->type != OBJ_ANY && old->type != OBJ_ANY &&
+			    link->type != old->type)
+				die(_("object %s is referred to as both a %s and a %s"),
+				    oid_to_hex(&link->entry.oid),
+				    type_name(old->type), type_name(link->type));
+			if (old->type == OBJ_ANY)
+				old->type = link->type;
+			free(link);
+		} else if (oidmap_put(dest, link)) {
+			BUG("duplicate outgoing link");
+		}
+	}
+	oidmap_clear(src, 0);
+}
+
 /*
  * Mutex and conditional variable can't be statically-initialized on Windows.
  */
@@ -211,6 +264,7 @@ static void init_thread(void)
 	CALLOC_ARRAY(thread_data, nr_threads);
 	for (i = 0; i < nr_threads; i++) {
 		thread_data[i].pack_fd = xopen(curr_pack, O_RDONLY);
+		oidmap_init(&thread_data[i].outgoing_links, 0);
 	}
 
 	threads_active = 1;
@@ -221,14 +275,17 @@ static void cleanup_thread(void)
 	int i;
 	if (!threads_active)
 		return;
-	threads_active = 0;
 	pthread_mutex_destroy(&read_mutex);
 	pthread_mutex_destroy(&counter_mutex);
 	pthread_mutex_destroy(&work_mutex);
 	if (show_stat)
 		pthread_mutex_destroy(&deepest_delta_mutex);
-	for (i = 0; i < nr_threads; i++)
+	for (i = 0; i < nr_threads; i++) {
+		merge_outgoing_links(&outgoing_links,
+				     &thread_data[i].outgoing_links);
 		close(thread_data[i].pack_fd);
+	}
+	threads_active = 0;
 	pthread_key_delete(key);
 	free(thread_data);
 }
@@ -818,9 +875,14 @@ static int check_collison(struct object_entry *entry)
 	return 0;
 }
 
-static void record_outgoing_link(const struct object_id *oid)
+static void record_outgoing_link(const struct object_id *oid,
+				 enum object_type type)
 {
-	oidset_insert(&outgoing_links, oid);
+	struct oidmap *map = &outgoing_links;
+
+	if (threads_active && !strict && !do_fsck_object)
+		map = &get_thread_data()->outgoing_links;
+	record_outgoing_link_to(map, oid, type);
 }
 
 static void maybe_record_name_entry(const struct name_entry *entry)
@@ -849,7 +911,97 @@ static void maybe_record_name_entry(const struct name_entry *entry)
 	 * pack, so it won't be GC-ed, the tradeoff seems worth it.
 	*/
 	if (S_ISDIR(entry->mode))
-		record_outgoing_link(&entry->oid);
+		record_outgoing_link(&entry->oid, OBJ_ANY);
+}
+
+static int parse_outgoing_link_oid(const char **buf, const char *tail,
+				   const char *header, struct object_id *oid)
+{
+	const char *end;
+	size_t header_len = strlen(header);
+
+	if (tail - *buf <= header_len + the_hash_algo->hexsz ||
+	    memcmp(*buf, header, header_len) ||
+	    parse_oid_hex_algop(*buf + header_len, oid, &end,
+				the_hash_algo) ||
+	    end >= tail || *end != '\n')
+		return -1;
+	*buf = end + 1;
+	return 0;
+}
+
+static void record_outgoing_links_from_data(const void *data,
+					    unsigned long size,
+					    enum object_type type,
+					    const struct object_id *oid)
+{
+	const char *buf = data;
+	const char *tail = buf + size;
+
+	if (type == OBJ_TREE) {
+		struct tree_desc desc;
+		struct name_entry entry;
+
+		if (init_tree_desc_gently(&desc, oid, data, size, 0))
+			return;
+		while (tree_entry_gently(&desc, &entry))
+			maybe_record_name_entry(&entry);
+	} else if (type == OBJ_COMMIT) {
+		struct object_id link;
+		struct commit_graft *graft;
+		int i;
+
+		if (threads_active &&
+		    !the_repository->parsed_objects->commit_graft_prepared)
+			BUG("commit grafts were not prepared before resolving deltas");
+		graft = lookup_commit_graft(the_repository, oid);
+
+		if (parse_outgoing_link_oid(&buf, tail, "tree ", &link))
+			die(_("invalid tree line in commit %s"),
+			    oid_to_hex(oid));
+		if (buf >= tail)
+			die(_("truncated commit %s after tree line"),
+			    oid_to_hex(oid));
+		record_outgoing_link(&link, OBJ_TREE);
+
+		while (tail - buf > 7 + the_hash_algo->hexsz &&
+		       starts_with(buf, "parent ")) {
+			if (parse_outgoing_link_oid(&buf, tail, "parent ",
+						    &link))
+				die(_("invalid parent line in commit %s"),
+				    oid_to_hex(oid));
+			if (buf >= tail)
+				die(_("truncated commit %s after parent line"),
+				    oid_to_hex(oid));
+			if (!graft ||
+			    (graft->nr_parent >= 0 && grafts_keep_true_parents))
+				record_outgoing_link(&link, OBJ_COMMIT);
+		}
+		if (graft)
+			for (i = 0; i < graft->nr_parent; i++)
+				record_outgoing_link(&graft->parent[i],
+						     OBJ_COMMIT);
+	} else if (type == OBJ_TAG) {
+		struct object_id link;
+		const char *line_end;
+		enum object_type target_type;
+
+		if (size < the_hash_algo->hexsz + 24 ||
+		    parse_outgoing_link_oid(&buf, tail, "object ", &link))
+			die(_("invalid object line in tag %s"), oid_to_hex(oid));
+		if (!skip_prefix(buf, "type ", &buf) ||
+		    !(line_end = memchr(buf, '\n', tail - buf)))
+			die(_("invalid type line in tag %s"), oid_to_hex(oid));
+		target_type = type_from_string_gently(buf, line_end - buf, 1);
+		if (target_type < 0)
+			die(_("invalid type line in tag %s"), oid_to_hex(oid));
+		buf = line_end + 1;
+		if (buf + 4 >= tail || !skip_prefix(buf, "tag ", &buf) ||
+		    !memchr(buf, '\n', tail - buf))
+			die(_("invalid tag name line in tag %s"),
+			    oid_to_hex(oid));
+		record_outgoing_link(&link, target_type);
+	}
 }
 
 static void do_record_outgoing_links(struct object *obj)
@@ -871,12 +1023,13 @@ static void do_record_outgoing_links(struct object *obj)
 		struct commit *commit = (struct commit *) obj;
 		struct commit_list *parents = commit->parents;
 
-		record_outgoing_link(get_commit_tree_oid(commit));
+		record_outgoing_link(get_commit_tree_oid(commit), OBJ_TREE);
 		for (; parents; parents = parents->next)
-			record_outgoing_link(&parents->item->object.oid);
+			record_outgoing_link(&parents->item->object.oid,
+					     OBJ_COMMIT);
 	} else if (obj->type == OBJ_TAG) {
 		struct tag *tag = (struct tag *) obj;
-		record_outgoing_link(get_tagged_oid(tag));
+		record_outgoing_link(get_tagged_oid(tag), tag->tagged->type);
 	}
 }
 
@@ -925,6 +1078,12 @@ static void sha1_object(const void *data, struct object_entry *obj_entry,
 		free(has_data);
 	}
 
+	if (record_outgoing_links && !strict && !do_fsck_object) {
+		if (type != OBJ_BLOB)
+			record_outgoing_links_from_data(data, size, type, oid);
+		goto out;
+	}
+
 	if (strict || do_fsck_object || record_outgoing_links) {
 		read_lock();
 		if (type == OBJ_BLOB) {
@@ -975,6 +1134,7 @@ static void sha1_object(const void *data, struct object_entry *obj_entry,
 		read_unlock();
 	}
 
+out:
 	free(new_data);
 }
 
@@ -1811,20 +1971,53 @@ static void show_pack_info(int stat_only)
 	free(chain_histogram);
 }
 
+static const struct object_id *idx_object_oid(size_t pos, const void *table)
+{
+	struct pack_idx_entry * const *entries = table;
+
+	return &entries[pos]->oid;
+}
+
+static void validate_outgoing_link_types(struct pack_idx_entry **sorted,
+					 int nr)
+{
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
+
+	oidmap_iter_init(&outgoing_links, &iter);
+	while ((link = oidmap_iter_next(&iter))) {
+		int pos;
+		struct object_entry *actual;
+
+		if (link->type == OBJ_ANY)
+			continue;
+		pos = oid_pos(&link->entry.oid, sorted, nr, idx_object_oid);
+		if (pos < 0)
+			continue;
+		actual = container_of(sorted[pos], struct object_entry, idx);
+		if (actual->real_type != link->type)
+			die(_("object %s is a %s, but was referred to as a %s"),
+			    oid_to_hex(&link->entry.oid),
+			    type_name(actual->real_type),
+			    type_name(link->type));
+	}
+}
+
 static void repack_local_links(void)
 {
 	struct child_process cmd = CHILD_PROCESS_INIT;
 	FILE *out;
 	struct strbuf line = STRBUF_INIT;
-	struct oidset_iter iter;
-	struct object_id *oid;
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
 	char *base_name = NULL;
 
-	if (!oidset_size(&outgoing_links))
+	if (!oidmap_get_size(&outgoing_links))
 		return;
 
-	oidset_iter_init(&outgoing_links, &iter);
-	while ((oid = oidset_iter_next(&iter))) {
+	oidmap_iter_init(&outgoing_links, &iter);
+	while ((link = oidmap_iter_next(&iter))) {
+		const struct object_id *oid = &link->entry.oid;
 		struct odb_source_info source_info;
 		struct object_info info = {
 			.source_infop = &source_info,
@@ -1919,6 +2112,7 @@ int cmd_index_pack(int argc,
 	fsck_options.walk = mark_link;
 
 	reset_pack_idx_option(&opts);
+	oidmap_init(&outgoing_links, 0);
 	opts.flags |= WRITE_REV;
 	repo_config(the_repository, git_index_pack_config, &opts);
 	if (prefix && chdir(prefix))
@@ -2102,6 +2296,7 @@ int cmd_index_pack(int argc,
 		idx_objects[i] = &objects[i].idx;
 	curr_index = write_idx_file(the_repository, index_name, idx_objects,
 				    nr_objects, &opts, pack_hash);
+	validate_outgoing_link_types(idx_objects, nr_objects);
 	if (rev_index)
 		curr_rev_index = write_rev_file(the_repository, rev_index_name,
 						idx_objects, nr_objects,
@@ -2146,6 +2341,7 @@ int cmd_index_pack(int argc,
 	free(curr_rev_index);
 
 	repack_local_links();
+	oidmap_clear(&outgoing_links, 1);
 
 	/*
 	 * Let the caller know this pack is not self contained
diff --git a/t/t5302-pack-index.sh b/t/t5302-pack-index.sh
index 735de1023e..2f99a49d9d 100755
--- a/t/t5302-pack-index.sh
+++ b/t/t5302-pack-index.sh
@@ -309,4 +309,54 @@ test_expect_success DEFAULT_HASH_ALGORITHM 'index-pack --fsck-objects outside of
 	)
 '
 
+test_expect_success 'index-pack --promisor rejects malformed commits' '
+	test_when_finished "rm -rf malformed-src malformed-dst malformed.pack bad-commit" &&
+	test_create_repo malformed-src &&
+	test_create_repo malformed-dst &&
+	printf "tree not-an-object\n\nmessage\n" >bad-commit &&
+	bad_oid=$(git -C malformed-src hash-object --literally -t commit \
+		-w --stdin <bad-commit) &&
+	printf "%s\n" "$bad_oid" |
+		git -C malformed-src pack-objects --stdout >malformed.pack &&
+	test_must_fail git -C malformed-dst index-pack --stdin --promisor \
+		<malformed.pack 2>err &&
+	test_grep "invalid tree line in commit $bad_oid" err
+'
+
+test_expect_success 'index-pack --promisor rejects conflicting link types' '
+	test_when_finished "rm -rf conflict-src conflict-dst conflict.pack bad-commit" &&
+	test_create_repo conflict-src &&
+	test_create_repo conflict-dst &&
+	tree_oid=$(git -C conflict-src mktree </dev/null) &&
+	{
+		printf "tree %s\n" "$tree_oid" &&
+		printf "parent %s\n\nmessage\n" "$tree_oid"
+	} >bad-commit &&
+	commit_oid=$(git -C conflict-src hash-object --literally -t commit \
+		-w --stdin <bad-commit) &&
+	printf "%s\n%s\n" "$tree_oid" "$commit_oid" |
+		git -C conflict-src pack-objects --stdout >conflict.pack &&
+	test_must_fail git -C conflict-dst index-pack --stdin --promisor \
+		<conflict.pack 2>err &&
+	test_grep "object $tree_oid is referred to as both a tree and a commit" err
+'
+
+test_expect_success 'index-pack --promisor verifies tag target types' '
+	test_when_finished "rm -rf tag-src tag-dst tag.pack bad-tag" &&
+	test_create_repo tag-src &&
+	test_create_repo tag-dst &&
+	tree_oid=$(git -C tag-src mktree </dev/null) &&
+	{
+		printf "object %s\n" "$tree_oid" &&
+		printf "type commit\ntag wrong-type\n\nmessage\n"
+	} >bad-tag &&
+	tag_oid=$(git -C tag-src hash-object --literally -t tag \
+		-w --stdin <bad-tag) &&
+	printf "%s\n%s\n" "$tree_oid" "$tag_oid" |
+		git -C tag-src pack-objects --stdout >tag.pack &&
+	test_must_fail git -C tag-dst index-pack --stdin --promisor \
+		<tag.pack 2>err &&
+	test_grep "object $tree_oid is a tree, but was referred to as a commit" err
+'
+
 test_done

base-commit: a97fcc37c2bc6340a8d7ce78dedf227aac4e9aa7
-- 
gitgitgadget

^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-02 21:33 [PATCH] index-pack: speed up promisor link recording Arijit Banerjee via GitGitGadget
@ 2026-08-02 21:51 ` brian m. carlson
  2026-08-02 22:20   ` Arijit Banerjee
                     ` (2 more replies)
  0 siblings, 3 replies; 9+ messages in thread
From: brian m. carlson @ 2026-08-02 21:51 UTC (permalink / raw)
  To: Arijit Banerjee via GitGitGadget
  Cc: git, Jonathan Tan, Patrick Steinhardt, Junio C Hamano,
	Arijit Banerjee, Arijit Banerjee

[-- Attachment #1: Type: text/plain, Size: 3539 bytes --]

On 2026-08-02 at 21:33:15, Arijit Banerjee via GitGitGadget wrote:
> From: Arijit Banerjee <arijit@effectiveailabs.com>
> 
> When indexing a promisor pack, index-pack parses every reconstructed
> non-blob object into the shared object model to record its outgoing links.
> Since parse_object_buffer() runs under read_mutex, worker threads serialize
> while allocating persistent tree, commit, and tag structures that are only
> needed to enumerate those links.
> 
> Read the links directly from the reconstructed object buffers instead. Keep
> the strict and fsck paths unchanged, use worker-local typed oidmaps during
> normal promisor indexing, and merge them after the workers exit. Transfer
> entries during the merge so that it does not temporarily duplicate the
> complete link set.
> 
> The typed entries preserve checks previously performed as a side effect of
> object parsing. Reject malformed commit and tag headers, conflicting
> expected types, and targets whose actual type disagrees when the target is
> present in the pack. Preserve commit-graft handling and the existing policy
> of recording only subtree entries from trees.
> 
> With three runs per version on Debian 12, median end-to-end wall-clock time
> for a --filter=blob:none clone of linux.git decreased from 156 seconds to
> 133 seconds (15%). Trace2 attributed the change to the initial index-pack
> --promisor phase, whose median duration decreased from 121 seconds to 98
> seconds (19%). System CPU time decreased by 46%.
> 
> Two paired spot checks against GitHub showed end-to-end reductions of 18%
> and 26%. These measurements include network and server variability and are
> therefore corroborating rather than controlled results. A third pair was not
> interpretable because the baseline request encountered a transport stall.
> 
> A full-clone control showed no material change, taking approximately 256
> seconds with either version. This is expected because full clones do not
> exercise promisor-link recording.
> 
> t5302-pack-index.sh passed with both SHA-1 and SHA-256, while
> t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage
> checks malformed commit headers, conflicting link types, and mismatched tag
> target types.
> 
> Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
> ---
>     index-pack: speed up promisor link recording
> 
>     AI assistance: OpenAI Codex was used to identify the bottleneck and
>     assist with the implementation, testing, and benchmark analysis. I
>     reviewed the resulting change and take responsibility for this
>     submission.

I don't think SubmittingPatches really allows more than trivial changes
written by AI:

    The Developer's Certificate of Origin requires contributors to certify
    that they know the origin of their contributions to the project and
    that they have the right to submit it under the project's license.
    It's not yet clear that this can be legally satisfied when submitting
    significant amount of content that has been generated by AI tools.

    [...]

    To avoid these issues, we will reject anything that looks AI
    generated, that sounds overly formal or bloated, that looks like AI
    slop, that looks good on the surface but makes no sense, or that
    senders don’t understand or cannot explain.

This doesn't look like it's a trivial change, so I don't believe this
patch can be accepted.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-02 21:51 ` brian m. carlson
@ 2026-08-02 22:20   ` Arijit Banerjee
       [not found]   ` <CAFwoC-7wUzce_XvuviXZe=5eTxJ5yyCpz=vsOheWKPCnz9Kr4A@mail.gmail.com>
  2026-08-02 22:46   ` Junio C Hamano
  2 siblings, 0 replies; 9+ messages in thread
From: Arijit Banerjee @ 2026-08-02 22:20 UTC (permalink / raw)
  To: brian m. carlson, Arijit Banerjee via GitGitGadget, git,
	Jonathan Tan, Patrick Steinhardt, Junio C Hamano, Arijit Banerjee,
	Arijit Banerjee

On Sun, Aug 2, 2026, brian m. carlson wrote:
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.

Thanks, Brian. I am not trying to bypass the project's policy.

I do not claim to be an expert on this topic, but Codex appears to have
found a material performance improvement of about 15% on end-to-end
blobless clone times. Would it be appropriate to treat the current
submission as an RFC for maintainers before deciding if the optimization is
worth getting in? It seems worth trying to preserve the technical
result.

Thanks,
Arijit


On Sun, Aug 2, 2026 at 2:52 PM brian m. carlson
<sandals@crustytoothpaste.net> wrote:
>
> On 2026-08-02 at 21:33:15, Arijit Banerjee via GitGitGadget wrote:
> > From: Arijit Banerjee <arijit@effectiveailabs.com>
> >
> > When indexing a promisor pack, index-pack parses every reconstructed
> > non-blob object into the shared object model to record its outgoing links.
> > Since parse_object_buffer() runs under read_mutex, worker threads serialize
> > while allocating persistent tree, commit, and tag structures that are only
> > needed to enumerate those links.
> >
> > Read the links directly from the reconstructed object buffers instead. Keep
> > the strict and fsck paths unchanged, use worker-local typed oidmaps during
> > normal promisor indexing, and merge them after the workers exit. Transfer
> > entries during the merge so that it does not temporarily duplicate the
> > complete link set.
> >
> > The typed entries preserve checks previously performed as a side effect of
> > object parsing. Reject malformed commit and tag headers, conflicting
> > expected types, and targets whose actual type disagrees when the target is
> > present in the pack. Preserve commit-graft handling and the existing policy
> > of recording only subtree entries from trees.
> >
> > With three runs per version on Debian 12, median end-to-end wall-clock time
> > for a --filter=blob:none clone of linux.git decreased from 156 seconds to
> > 133 seconds (15%). Trace2 attributed the change to the initial index-pack
> > --promisor phase, whose median duration decreased from 121 seconds to 98
> > seconds (19%). System CPU time decreased by 46%.
> >
> > Two paired spot checks against GitHub showed end-to-end reductions of 18%
> > and 26%. These measurements include network and server variability and are
> > therefore corroborating rather than controlled results. A third pair was not
> > interpretable because the baseline request encountered a transport stall.
> >
> > A full-clone control showed no material change, taking approximately 256
> > seconds with either version. This is expected because full clones do not
> > exercise promisor-link recording.
> >
> > t5302-pack-index.sh passed with both SHA-1 and SHA-256, while
> > t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage
> > checks malformed commit headers, conflicting link types, and mismatched tag
> > target types.
> >
> > Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
> > ---
> >     index-pack: speed up promisor link recording
> >
> >     AI assistance: OpenAI Codex was used to identify the bottleneck and
> >     assist with the implementation, testing, and benchmark analysis. I
> >     reviewed the resulting change and take responsibility for this
> >     submission.
>
> I don't think SubmittingPatches really allows more than trivial changes
> written by AI:
>
>     The Developer's Certificate of Origin requires contributors to certify
>     that they know the origin of their contributions to the project and
>     that they have the right to submit it under the project's license.
>     It's not yet clear that this can be legally satisfied when submitting
>     significant amount of content that has been generated by AI tools.
>
>     [...]
>
>     To avoid these issues, we will reject anything that looks AI
>     generated, that sounds overly formal or bloated, that looks like AI
>     slop, that looks good on the surface but makes no sense, or that
>     senders don’t understand or cannot explain.
>
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.
> --
> brian m. carlson (they/them)
> Toronto, Ontario, CA

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
       [not found]   ` <CAFwoC-7wUzce_XvuviXZe=5eTxJ5yyCpz=vsOheWKPCnz9Kr4A@mail.gmail.com>
@ 2026-08-02 22:32     ` brian m. carlson
  2026-08-02 22:52       ` Junio C Hamano
  0 siblings, 1 reply; 9+ messages in thread
From: brian m. carlson @ 2026-08-02 22:32 UTC (permalink / raw)
  To: Arijit Banerjee
  Cc: Arijit Banerjee via GitGitGadget, git, Jonathan Tan,
	Patrick Steinhardt, Junio C Hamano, Arijit Banerjee, ttaylorr

[-- Attachment #1: Type: text/plain, Size: 2223 bytes --]

On 2026-08-02 at 22:12:16, Arijit Banerjee wrote:
> Thanks, Brian. I am not trying to bypass the project's policy.
> 
> The investigation is in the same general spirit as the Git performance work
> being tracked here:
> https://openai-git-upstream.openai.chatgpt.site/
> 
> I do not claim to be an expert on this topic, but Codex appears to have found
> a material performance improvement of about 15% on end-to-end blobless clone
> times.
> 
> Would it be appropriate to treat the current submission as an RFC? It seems
> worth trying to preserve the technical result.

I don't think the project's policy prevents you from doing analysis and
investigation with an LLM, although it does require you to verify the
correctness of the results and be accountable for them.  If, based on
the analysis of the performance impact, you write some code without the
use of an LLM that improves things, I think that would be allowed and
probably welcome, assuming it is otherwise acceptable.  Some
contributors will be willing to review such a contribution and others
will not, but it is not outside of the policy.

However, writing substantial code with an LLM doesn't appear to be
allowed.  The kinds of trivial changes that I think would be allowed to
be generated would be things like fixing spelling errors or adding
include guards to header files that lack them.  Of course, these are
also the kinds of things you could mostly fix with a small script, which
is why they are generally considered so trivial as to be
uncopyrightable.

So I think to have a patch accepted in this case, you would need to
totally discard the existing patch and rewrite it by hand without
recourse to the generated code.

I understand that the SubmittingPatches documentation is a bit long, but
I do suggest giving it at least a glance so you know what to expect.  I
think reading this sort of contributing documentation is more important
than ever since, in the era of LLMs, projects tend to have strong
opinions on what is and is not acceptable, not only just in terms of LLM
usage, but in how code and documentation are to be written and
formatted.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-02 21:51 ` brian m. carlson
  2026-08-02 22:20   ` Arijit Banerjee
       [not found]   ` <CAFwoC-7wUzce_XvuviXZe=5eTxJ5yyCpz=vsOheWKPCnz9Kr4A@mail.gmail.com>
@ 2026-08-02 22:46   ` Junio C Hamano
       [not found]     ` <CAFwoC-6EvoD-u7oceETi90MJ-FQA2zihdkn1i1wckKfoYRTKOw@mail.gmail.com>
  2 siblings, 1 reply; 9+ messages in thread
From: Junio C Hamano @ 2026-08-02 22:46 UTC (permalink / raw)
  To: brian m. carlson
  Cc: Arijit Banerjee via GitGitGadget, git, Jonathan Tan,
	Patrick Steinhardt, Arijit Banerjee, Arijit Banerjee

"brian m. carlson" <sandals@crustytoothpaste.net> writes:

>>     index-pack: speed up promisor link recording
>> 
>>     AI assistance: OpenAI Codex was used to identify the bottleneck and
>>     assist with the implementation, testing, and benchmark analysis. I
>>     reviewed the resulting change and take responsibility for this
>>     submission.
>
> I don't think SubmittingPatches really allows more than trivial changes
> written by AI:
>
>     The Developer's Certificate of Origin requires contributors to certify
>     that they know the origin of their contributions to the project and
>     that they have the right to submit it under the project's license.
>     It's not yet clear that this can be legally satisfied when submitting
>     significant amount of content that has been generated by AI tools.
>
>     [...]
>
>     To avoid these issues, we will reject anything that looks AI
>     generated, that sounds overly formal or bloated, that looks like AI
>     slop, that looks good on the surface but makes no sense, or that
>     senders don’t understand or cannot explain.
>
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.

The project we borrowed DCO from has this to say on this topic:

 https://docs.kernel.org/process/coding-assistants.html

 * All contributions must comply with licensing reuqirements.
 * Only humans can attest DCO by Siging off their patches.

   The human submitter is responsible for reviewing all AI generated
   code, ensuring compliance with licensing requirements, certify
   DCO with their sign-off, and taking full responsibility for the
   contribution.

Now we are *not* kernel, but I think there is a general concensus in
the community that, while we do not want to outright ban machine
assisted contributions, we generally want to tread very carefully,
especially in the DCO area.

It is very easy for anybody and their dog to say "I reviewed X" and
it is very hard for others to assess how trustworthy such a
statement is, so I am unsure how the kernel project is enforcing the
"human submitter is responsible for these", and more importantly,
even assuming that we would take a similar policy for ourselves, I
am not sure what mechanism we can put in place to detect cases where
these expectations are violated.

So...

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-02 22:32     ` brian m. carlson
@ 2026-08-02 22:52       ` Junio C Hamano
  2026-08-02 23:19         ` Arijit Banerjee
  0 siblings, 1 reply; 9+ messages in thread
From: Junio C Hamano @ 2026-08-02 22:52 UTC (permalink / raw)
  To: brian m. carlson
  Cc: Arijit Banerjee, Arijit Banerjee via GitGitGadget, git,
	Jonathan Tan, Patrick Steinhardt, Arijit Banerjee, ttaylorr

"brian m. carlson" <sandals@crustytoothpaste.net> writes:

> I understand that the SubmittingPatches documentation is a bit long, but
> I do suggest giving it at least a glance so you know what to expect.  I
> think reading this sort of contributing documentation is more important
> than ever since, in the era of LLMs, projects tend to have strong
> opinions on what is and is not acceptable, not only just in terms of LLM
> usage, but in how code and documentation are to be written and
> formatted.

Amen.

Since we seem to be drawn into the AI policy discussion, are there
things that we should consider borrowing from policies battle-tested
by other projects?  I kind of like what LLVM has as "extractive
contributions are rejected (whether it is AI or not AI)", as we are
severely review-bandwidth limited these days.


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-02 22:52       ` Junio C Hamano
@ 2026-08-02 23:19         ` Arijit Banerjee
  0 siblings, 0 replies; 9+ messages in thread
From: Arijit Banerjee @ 2026-08-02 23:19 UTC (permalink / raw)
  To: Junio C Hamano
  Cc: brian m. carlson, Arijit Banerjee via GitGitGadget, git,
	Jonathan Tan, Patrick Steinhardt, Arijit Banerjee, ttaylorr

Two suggestions:
- Maybe a karma system?
- A special release train that is more indulgent towards AI generated
code? Brave users can try out features and they get baked into stable
releases after enough soak time

Definitely not trying to be extractive, this one appears to be a
decent size perf improvement!

Still holding on to a SHA1-DC hardware acceleration change, that one
would admittedly be much harder to review ;)


On Sun, Aug 2, 2026 at 3:52 PM Junio C Hamano <gitster@pobox.com> wrote:
>
> "brian m. carlson" <sandals@crustytoothpaste.net> writes:
>
> > I understand that the SubmittingPatches documentation is a bit long, but
> > I do suggest giving it at least a glance so you know what to expect.  I
> > think reading this sort of contributing documentation is more important
> > than ever since, in the era of LLMs, projects tend to have strong
> > opinions on what is and is not acceptable, not only just in terms of LLM
> > usage, but in how code and documentation are to be written and
> > formatted.
>
> Amen.
>
> Since we seem to be drawn into the AI policy discussion, are there
> things that we should consider borrowing from policies battle-tested
> by other projects?  I kind of like what LLVM has as "extractive
> contributions are rejected (whether it is AI or not AI)", as we are
> severely review-bandwidth limited these days.
>

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
       [not found]     ` <CAFwoC-6EvoD-u7oceETi90MJ-FQA2zihdkn1i1wckKfoYRTKOw@mail.gmail.com>
@ 2026-08-03  0:31       ` brian m. carlson
  2026-08-03  0:55         ` Collin Funk
  0 siblings, 1 reply; 9+ messages in thread
From: brian m. carlson @ 2026-08-03  0:31 UTC (permalink / raw)
  To: Arijit Banerjee
  Cc: Junio C Hamano, Arijit Banerjee via GitGitGadget, git,
	Jonathan Tan, Patrick Steinhardt, Arijit Banerjee

[-- Attachment #1: Type: text/plain, Size: 3424 bytes --]

[please avoid top-posting]

On 2026-08-02 at 22:54:27, Arijit Banerjee wrote:
> Maybe software can ship experimental versions which is more indulging
> towards AI generated patches? Brave users get to try the features and they
> can baked into stable releases once there's enough soak time.
> 
> I have been holding onto my patch around hardware acceleration for SHA1-DC
> :)

The rationale, as Junio said, is based on the Developer's Certificate of
Origin.  That is a legal statement that a person has the legal right to
contribute those changes under the license and if they make a false or
misleading statement to that effect, they are responsible—legally and
otherwise—for it.

Considering the extensive litigation over LLM output at the moment, I
don't think anyone can clearly make that assertion.  Most of the
arguments I've heard are that it's fair use, which is a U.S. legal
concept.  That does not exist in Canada or the U.K., where there is fair
dealing, which is much more restricted.

If Company X includes LLM-generated code in their proprietary product
and it's found to be infringing in say, Germany, then they can simply
not distribute their code in Germany.  Git cannot do that: it's
distributed in Linux distributions around the planet, even in countries
subject to sanctions, such as Russia[0].  We must comply with the
license and the law everywhere in every country or we risk liability for
our contributors and distributors.  I, for one, am not willing to be
sued over this project and the project does not have the financial means
to deal with extensive litigation.

There are also concerns about the quality of the code and whether
submitters adequately understand the code well enough to have evaluated
and reviewed it thoroughly.  It's well known that when creating code
becomes cheap, the burden shifts to review and review becomes extremely
important.  That has been seen in lots of places, but we are an open
source project and we can't force contributors to do review like a
company can.  As Junio says, we already have trouble getting reviews
through and we don't want to make the problem worse.

LLM-generated content also has a negative quality reputation (see the
reaction to AI content in video games and books for an example) and
while all software has bugs, I appreciate the reputation that Git has
for quality and wish to retain that.

Those alone are reason enough for the policy, but there are other
concerns about the environmental impact, the impact on electricity and
hardware prices, the ethics of incorporating open source code without so
much as a credit[1], and a lot more.

So I don't think it's likely we're going to accept nontrivial
LLM-generated content in any capacity anytime soon and I don't think
trying to argue this or persuade us to accept it is going to be
productive or well received.  Of course, anyone can distribute their own
fork of Git with additional patches if they prefer, but we won't include
them.

[0] Debian, which distributes Git, has mirrors in Russia and Belarus:
https://www.debian.org/mirror/list
[1] For instance, as a member of ACM, I have to follow § 1.5 (Respect
the work required to produce new ideas, inventions, creative works, and
computing artifacts) of the Code of Ethics:
https://www.acm.org/code-of-ethics.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] index-pack: speed up promisor link recording
  2026-08-03  0:31       ` brian m. carlson
@ 2026-08-03  0:55         ` Collin Funk
  0 siblings, 0 replies; 9+ messages in thread
From: Collin Funk @ 2026-08-03  0:55 UTC (permalink / raw)
  To: brian m. carlson
  Cc: Arijit Banerjee, Junio C Hamano, Arijit Banerjee via GitGitGadget,
	git, Jonathan Tan, Patrick Steinhardt, Arijit Banerjee

"brian m. carlson" <sandals@crustytoothpaste.net> writes:

> If Company X includes LLM-generated code in their proprietary product
> and it's found to be infringing in say, Germany, then they can simply
> not distribute their code in Germany.  Git cannot do that: it's
> distributed in Linux distributions around the planet, even in countries
> subject to sanctions, such as Russia[0].  We must comply with the
> license and the law everywhere in every country or we risk liability for
> our contributors and distributors.  I, for one, am not willing to be
> sued over this project and the project does not have the financial means
> to deal with extensive litigation.

I generally agree. But some countries have more respected legal systems
than others, to put it mildly. I certainly hope that being extradited to
one of the sanctioned countries isn't too large of a concern.

Collin

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-08-03  0:55 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-02 21:33 [PATCH] index-pack: speed up promisor link recording Arijit Banerjee via GitGitGadget
2026-08-02 21:51 ` brian m. carlson
2026-08-02 22:20   ` Arijit Banerjee
     [not found]   ` <CAFwoC-7wUzce_XvuviXZe=5eTxJ5yyCpz=vsOheWKPCnz9Kr4A@mail.gmail.com>
2026-08-02 22:32     ` brian m. carlson
2026-08-02 22:52       ` Junio C Hamano
2026-08-02 23:19         ` Arijit Banerjee
2026-08-02 22:46   ` Junio C Hamano
     [not found]     ` <CAFwoC-6EvoD-u7oceETi90MJ-FQA2zihdkn1i1wckKfoYRTKOw@mail.gmail.com>
2026-08-03  0:31       ` brian m. carlson
2026-08-03  0:55         ` Collin Funk

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox