Git development
 help / color / mirror / Atom feed
* [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance
@ 2026-10-07 14:29 Scott Chacon
  2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
  2026-10-08 18:10 ` [RFC PATCH 0/1] " D. Ben Knoble
  0 siblings, 2 replies; 12+ messages in thread
From: Scott Chacon @ 2026-10-07 14:29 UTC (permalink / raw)
  To: git

After the AI discussion at the contributors' summit [1], I'd like to submit
a concrete alternative for allowing AI generated work responsibly. This moves
us closer to Linux's approach: use the tools you find helpful, but take 
responsibility for what you send.

Our current policy encourages careful use of AI, then says we'll reject
anything that looks AI generated. That leaves someone with a useful,
reviewed patch wondering whether telling us how they made it will get it
rejected. I believe that it would be better to allow valuable, reviewed
series and simply include disclosure (again, how Linux does it).

SFC's legal advice came up at the summit. The original policy credits
Rick Sanders [2] of the SFC and there was some talk of the policy being sound
because it had legal review. However, the SFC's public guidance has changed
since then. They date a change in strategy to November 2025 (1 month after
reviewing the Git policy), when they concluded that relying only on bans was
no longer a good approach [3].  

Their June 2026 recommendations describe how to use these tools responsibly:
review the output, disclose the assistance, and keep records [4]. The change
proposed with this patch is in line with their current recommendations.

There are also several other prominant example projects:

* Linux accepts tool-generated contributions under the existing DCO,
  asks for disclosure, and leaves maintainers free to request more
  testing or reject a patch [5][6].
* Xen is another GPLv2 project using human sign-off and an Assisted-by
  trailer [7][8]. Its documentation change explicitly followed Linux
  [9].
* Debian now allows responsible AI use too, though disclosure is
  optional there [10].

They all agree that we don't need to change the DCO to do this. Clause (a)
already covers work created "in whole or in part" by the contributor, and (b)
covers changes to appropriately licensed existing work [11]. Both still
require the right to submit the contribution. Neither requires the
submitter to have personally written every line [12].

So this updated version of the submission guidelines asks contributors sending
AI assisted patches (code included) to:

* Review and understand the whole submission, test it appropriately, and
  answer review comments.
* Meet the existing DCO and license requirements.
* Add an Assisted-by trailer for substantial assistance and briefly
  explain what the tool did and how they checked it.

I'd like us to give people a clear way to submit good work with these
tools, while keeping the expectations that make patches worth reviewing.

[1] Git Contributors' Summit 2026, AI contribution policy discussion:
    https://lore.kernel.org/git/summit-2026.94e33e9ddf234334.06@ttaylorr.com/
[2] Original Git policy commit:
    https://github.com/git/git/commit/7b0c37953d2e9198309ca6b6faf10bb5deeb4837
[3] SFC's account of its November 2025 strategic reassessment:
    https://sfconservancy.org/llm-gen-ai/
[4] SFC's recommendations, particularly points 4-8 and 11:
    https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html
[5] Linux tool-generated content guidelines:
    https://docs.kernel.org/process/generated-content.html
[6] Linux AI coding assistant requirements:
    https://docs.kernel.org/process/coding-assistants.html
[7] Xen contribution guidance, Assisted-by and Signed-off-by:
    https://xenbits.xen.org/docs/unstable/process/sending-patches.html#assisted-by
[8] Xen licensing:
    https://github.com/xen-project/xen/blob/master/COPYING
[9] Xen's Linux-inspired documentation patch, 2026-06-15:
    https://lists.xenproject.org/archives/html/xen-devel/2026-06/msg00882.html
[10] Debian GR 2026/002, winning option 5:
     https://www.debian.org/vote/2026/vote_002
[11] Developer Certificate of Origin 1.1:
     https://developercertificate.org/
[12] Red Hat's DCO analysis, 2025-10-15:
     https://www.redhat.com/en/blog/ai-assisted-development-and-open-source-navigating-legal-issues

Scott Chacon (1):
  SubmittingPatches: allow responsible AI assistance

 Documentation/SubmittingPatches | 73 ++++++++++++++++++++++-----------
 1 file changed, 49 insertions(+), 24 deletions(-)


base-commit: a018953688f1b10bddf91bff8747068f5f4746a4
-- 
2.50.1 (Apple Git-155)


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 14:29 [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance Scott Chacon
@ 2026-10-07 14:29 ` Scott Chacon
  2026-10-07 21:42   ` brian m. carlson
  2026-10-07 22:44   ` Junio C Hamano
  2026-10-08 18:10 ` [RFC PATCH 0/1] " D. Ben Knoble
  1 sibling, 2 replies; 12+ messages in thread
From: Scott Chacon @ 2026-10-07 14:29 UTC (permalink / raw)
  To: git

The AI section encourages careful use of AI tools, but also says we
will reject anything that looks AI generated. That leaves contributors
without a clear path for submitting useful, reviewed, understood work and
can discourage disclosure of the assistance they received.

Allow AI-assisted contributions under the usual quality and licensing
requirements. Require human understanding, appropriate testing, and
disclosure of substantial assistance. Retain the DCO without changing
its terms, and require contributors to consider provenance and meet
applicable license obligations. Reviewers can ask for further evidence
or decline work they cannot confidently assess.

Replace the appearance-based rejection rule with these concrete
expectations. AI assistance neither excuses an inadequate submission
nor prevents an otherwise acceptable one from being considered.

As an example, an OpenAI model was used to help me research, compare and
craft the appropriate legal language for this policy change to help us
match the modern, legally reviewed approaches now taken by peer GPL
projects such as the Linux kernel [1].

[1] https://docs.kernel.org/process/coding-assistants.html

Assisted-by: OpenAI GPT-6 Astra
Signed-off-by: Scott Chacon <scott@gitbutler.net>
---
 Documentation/SubmittingPatches | 73 ++++++++++++++++++++++-----------
 1 file changed, 49 insertions(+), 24 deletions(-)

diff --git a/Documentation/SubmittingPatches b/Documentation/SubmittingPatches
index c60855f706..f703f96667 100644
--- a/Documentation/SubmittingPatches
+++ b/Documentation/SubmittingPatches
@@ -571,30 +571,55 @@ the patches.
 [[ai]]
 === Use of Artificial Intelligence (AI)
 
-The Developer's Certificate of Origin requires contributors to certify
-that they know the origin of their contributions to the project and
-that they have the right to submit it under the project's license.
-It's not yet clear that this can be legally satisfied when submitting
-significant amount of content that has been generated by AI tools.
-
-Another issue with AI generated content is that AIs still often
-hallucinate or just produce bad code, commit messages, documentation
-or output, even when you point out their mistakes.
-
-To avoid these issues, we will reject anything that looks AI
-generated, that sounds overly formal or bloated, that looks like AI
-slop, that looks good on the surface but makes no sense, or that
-senders don’t understand or cannot explain.
-
-We strongly recommend using AI tools carefully and responsibly.
-
-Contributors would often benefit more from AI by using it to guide and
-help them step by step towards producing a solution by themselves
-rather than by asking for a full solution that they would then mostly
-copy-paste. They can also use AI to help with debugging, or with
-checking for obvious mistakes, things that can be improved, things
-that don’t match our style, guidelines or our feedback, before sending
-it to us.
+AI tools may be used to help prepare contributions, including code,
+tests, documentation, and commit messages. AI assistance does not by
+itself disqualify a contribution. The same requirements for correctness,
+maintainability, licensing, and review apply regardless of the tools
+used.
+
+You are responsible for the entire contribution. Before submitting it,
+review and understand the changes, check factual claims, and perform
+the testing appropriate to the change. Be prepared to explain your
+decisions and respond to review comments. Do not pass unreviewed tool
+output on to reviewers, including in commit messages or mailing list
+replies. Keep explanations concise and relevant to the change.
+
+The <<dco,Developer's Certificate of Origin>> applies unchanged. Only a
+human can make that certification; an AI tool cannot sign off on your
+behalf. Consider the origin and licensing of generated material,
+including any third-party material it reproduces, and comply with
+applicable license and attribution requirements. A tool's assurance
+that its output is original or compatible with our license is not a
+substitute for checking those requirements. If you cannot certify the
+DCO for a contribution, do not submit it.
+
+Disclose substantial AI assistance in each affected commit with an
+`Assisted-by:` trailer naming the tool and, when available, its model
+or version. For example:
+
+....
+	Assisted-by: ExampleTool version 1.2
+....
+
+In the accompanying explanation, briefly describe how the tool helped,
+which parts of the contribution it affected, and how you checked the
+result. This can go in the cover letter or below the `---` line in the
+patch email. Disclose substantial assistance with mailing list replies
+in those replies as well. Trivial spelling corrections, formatting,
+and identifier completion do not need disclosure. When in doubt,
+disclose the assistance.
+
+Keep relevant prompts and outputs to help answer questions during
+review. Include short prompts, or a summary of longer sessions, when
+they help explain the change. Do not include credentials or private
+material in these records when sharing them.
+
+Maintainers may request more explanation, testing, or information about
+provenance, and may decline contributions they cannot confidently
+assess. Tool use does not entitle a contribution to review or
+acceptance. Discuss plans for large-scale automated submissions on the
+mailing list before sending them; generating patches faster does not
+increase the project's capacity to review them.
 
 [[git-tools]]
 === Generate your patch using Git tools out of your commits.
-- 
2.50.1 (Apple Git-155)


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
@ 2026-10-07 21:42   ` brian m. carlson
  2026-10-08  4:49     ` Scott Chacon
  2026-10-07 22:44   ` Junio C Hamano
  1 sibling, 1 reply; 12+ messages in thread
From: brian m. carlson @ 2026-10-07 21:42 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

[-- Attachment #1: Type: text/plain, Size: 3605 bytes --]

On 2026-10-07 at 14:29:54, Scott Chacon wrote:
> The AI section encourages careful use of AI tools, but also says we
> will reject anything that looks AI generated. That leaves contributors
> without a clear path for submitting useful, reviewed, understood work and
> can discourage disclosure of the assistance they received.
> 
> Allow AI-assisted contributions under the usual quality and licensing
> requirements. Require human understanding, appropriate testing, and
> disclosure of substantial assistance. Retain the DCO without changing
> its terms, and require contributors to consider provenance and meet
> applicable license obligations. Reviewers can ask for further evidence
> or decline work they cannot confidently assess.
> 
> Replace the appearance-based rejection rule with these concrete
> expectations. AI assistance neither excuses an inadequate submission
> nor prevents an otherwise acceptable one from being considered.
> 
> As an example, an OpenAI model was used to help me research, compare and
> craft the appropriate legal language for this policy change to help us
> match the modern, legally reviewed approaches now taken by peer GPL
> projects such as the Linux kernel [1].

I don't think I'm in favour of this policy.  All the major models have
been trained on a large variety of code from a large variety of sources,
including sources such as news reports or personal websites that do not
allow copying, modification, or distribution.  Given that LLMs are known
to reproduce portions of their training set or craft code or text which
is very similar to items in the training set, how can anyone honestly
assert the DCO without knowing all of the sources that were used to
create it?

Even if the model were, for instance, trained only on MIT-licensed code,
the license still requires a copyright and permission notice on every
copy, so the fact that the code generated from an LLM doesn't contain
that would seem to violate the license and prohibit us from using it.

The DCO was created to help us unambiguously state that the code is
acceptable to be included to avoid any later claims that the code was
copied from somewhere that it shouldn't have been.  Given the fact that
nobody knows what the sources are with a current LLM, it doesn't seem
that a reasonable person could make such an assertion.

I'm a distributor of Git and I don't want to be sued or arrested because
I end up distributing code that I don't have the right to distribute.
Large companies may have lawyers and lots of money to fight those
claims, but I do not (nor does the Git project) and I don't want to
spend my resources fighting allegations of copyright infringement or
have my reputation besmirched for that reason.  Just because other
projects think it's okay to do legally and ethically questionable things
doesn't mean we should as well.

I'll add that if Git were to include a portion of my MIT- or
BSD-licensed code without including a copyright or permission notice
because it was laundered through an LLM, I would absolutely file a
copyright complaint, and rightfully so.

I refer you to policies from other major open source projects that cover
this exact provenance issue:

* Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy
* NetBSD: https://www.netbsd.org/developers/commit-guidelines.html

I also will point out the notes from the Contributor Summit where we
discussed this issue in some depth and proposed an approach for further
discussion.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
  2026-10-07 21:42   ` brian m. carlson
@ 2026-10-07 22:44   ` Junio C Hamano
  2026-10-08 13:53     ` Scott Chacon
  1 sibling, 1 reply; 12+ messages in thread
From: Junio C Hamano @ 2026-10-07 22:44 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

Scott Chacon <scott@gitbutler.net> writes:

> As an example, an OpenAI model was used to help me research, compare and
> craft the appropriate legal language for this policy change to help us
> match the modern, legally reviewed approaches now taken by peer GPL
> projects such as the Linux kernel [1].
>
> [1] https://docs.kernel.org/process/coding-assistants.html

That makes it sound as if this is just as legally sound as what the
kernel project uses.  However, the only assurance we get (unless you
are willing to act as our lawyer, and I do not know if you are one)
is that an OpenAI model produced plausible-sounding utterances.

Indeed, the proposed text seems to instruct developers and reviewers
to do quite different things from the rules I see in the above URL.
Note that in the following, I will be playing devil's advocate for
much of the time, so please accept my apologies in advance if I
sound too skeptical.

> +AI tools may be used to help prepare contributions, including code,
> +tests, documentation, and commit messages. AI assistance does not by
> +itself disqualify a contribution. The same requirements for correctness,
> +maintainability, licensing, and review apply regardless of the tools
> +used.
> +
> +You are responsible for the entire contribution. Before submitting it,
> +review and understand the changes, check factual claims, and perform
> +the testing appropriate to the change. Be prepared to explain your
> +decisions and respond to review comments. Do not pass unreviewed tool
> +output on to reviewers, including in commit messages or mailing list
> +replies. Keep explanations concise and relevant to the change.

It looks, at least to me, that there is not much that can be
meaningfully enforced by reviewers and followed by contributors in
the above text.  It seems to be little more than "the world would be
a wonderful place if everybody behaved this way."

A violation of "concise and relevant" seems to be the recent trend
of much AI-generated slop, so it may be a good suggestion to give
today.  But would we need to update it once the trend of text
generated by AI tools becomes "concise and relevant" nonsense that
merely sounds plausible?  What if an "AI-assisted" contributor lacks
common sense to tell between plausible-sounding nonsense and a
well-written description?  What if reviewers get too many such
"contributions" and cannot allocate enough review bandwidth to sift
good contributions from plausible-sounding nonsense?

> +The <<dco,Developer's Certificate of Origin>> applies unchanged. Only a
> +human can make that certification; an AI tool cannot sign off on your
> +behalf. Consider the origin and licensing of generated material,
> +including any third-party material it reproduces, and comply with
> +applicable license and attribution requirements. A tool's assurance
> +that its output is original or compatible with our license is not a
> +substitute for checking those requirements. If you cannot certify the
> +DCO for a contribution, do not submit it.

Again, this is a good aspiration to have, but I doubt that anyone
can practically certify that the output of an LLM is devoid of
content borrowed from problematic sources under the rule the text
above gives.  Would it not be more useful to help contributors by
defining what not to do more clearly?  Our current text says as much
more directly: you cannot practically certify, so do not send in
AI-generated slop, period.

> +Disclose substantial AI assistance in each affected commit with an
> +`Assisted-by:` trailer naming the tool and, when available, its model
> +or version. For example:
> +
> +....
> +	Assisted-by: ExampleTool version 1.2
> +....

I thought the kernel guidelines instructed us to say only "LLM"
these days, to avoid giving free advertising.  On the other hand,
they ask contributors to also list non-LLM tools, like coccinelle
and clang-tidy, that were used in their machine-assisted
contributions.  I am undecided on the merit of specifying the
exact model and version, but listing non-LLM tools alongside
materials for independent reproduction looks like a good idea.

> +Maintainers may request more explanation, testing, or information about
> +provenance, and may decline contributions they cannot confidently
> +assess.

The text of the kernel guidelines appears to give maintainers more
latitude (cf. https://docs.kernel.org/process/generated-content.html).
They can treat it just like any other contribution, reject it
outright, or choose any approach in between.  The proposed text above
does not account for cases where reviewers simply lack the bandwidth
to even think about what explanation and proof to request, and it
makes it sound as if declining a submission in such a case an unfair
rejection.



^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 21:42   ` brian m. carlson
@ 2026-10-08  4:49     ` Scott Chacon
  2026-10-08  5:29       ` Luca Milanesio
                         ` (2 more replies)
  0 siblings, 3 replies; 12+ messages in thread
From: Scott Chacon @ 2026-10-08  4:49 UTC (permalink / raw)
  To: brian m. carlson, Scott Chacon, git

On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson
<sandals@crustytoothpaste.net> wrote:
>
> On 2026-10-07 at 14:29:54, Scott Chacon wrote:
> I don't think I'm in favour of this policy.  All the major models have
> been trained on a large variety of code from a large variety of sources,
> including sources such as news reports or personal websites that do not
> allow copying, modification, or distribution.  Given that LLMs are known
> to reproduce portions of their training set or craft code or text which
> is very similar to items in the training set, how can anyone honestly
> assert the DCO without knowing all of the sources that were used to
> create it?

I agree that generated output can reproduce material we don't have
permission to distribute (though I think this is incredibly rare for
anything complex). What I question is whether that possibility means
knowing every source in the training set is necessary to make any DCO
certification.

Human contributors have also read code under many different licenses
(and news articles and blogs) . We don't ask them to account for
everything they've ever read before signing off on a patch. We do
expect them to have the right to submit the actual contribution and to
respect the licenses of material they incorporate. Again, Red Hat, the
Linux Foundation and the SFC all now state that LLM generated code is
acceptable and compatible with DCO requirements.

> I'm a distributor of Git and I don't want to be sued or arrested because
> I end up distributing code that I don't have the right to distribute.
> Large companies may have lawyers and lots of money to fight those
> claims, but I do not (nor does the Git project) and I don't want to
> spend my resources fighting allegations of copyright infringement or
> have my reputation besmirched for that reason.  Just because other
> projects think it's okay to do legally and ethically questionable things
> doesn't mean we should as well.

Again, a lot of my argumentation here was directly taken from the
SFC's recommendations [1], which Git is a member project of.

https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html

> I'll add that if Git were to include a portion of my MIT- or
> BSD-licensed code without including a copyright or permission notice
> because it was laundered through an LLM, I would absolutely file a
> copyright complaint, and rightfully so.
>
> I refer you to policies from other major open source projects that cover
> this exact provenance issue:
>
> * Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy
> * NetBSD: https://www.netbsd.org/developers/commit-guidelines.html

I'm not saying that other projects haven't taken positions as
conservative as Git's current policy. I'm saying that much larger
projects with much larger legal surface area such as Linux have
adopted more progressive ones. Linus is fine with it on a project with
the same DCO, the same license, and honestly, a lot more legal
scrutiny.

> I also will point out the notes from the Contributor Summit where we
> discussed this issue in some depth and proposed an approach for further
> discussion.

It was unclear from the notes what a "vote" meant exactly. It looked
like Taylor was roughly tasked with providing a Linux-style position,
so I thought I would help out by providing a first version of that
position.

This was written to follow the current SFC guidelines. I would
encourage Junio to have them read it, but I tried my best to follow
it's guidelines and advice.

Scott

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-08  4:49     ` Scott Chacon
@ 2026-10-08  5:29       ` Luca Milanesio
  2026-10-08  5:53       ` Kristoffer Haugsbakk
  2026-10-08 16:15       ` brian m. carlson
  2 siblings, 0 replies; 12+ messages in thread
From: Luca Milanesio @ 2026-10-08  5:29 UTC (permalink / raw)
  To: git; +Cc: Luca Milanesio



> On 8 Oct 2026, at 05:49, Scott Chacon <schacon@gmail.com> wrote:
> 
> On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson
> <sandals@crustytoothpaste.net> wrote:
>> 
>> On 2026-10-07 at 14:29:54, Scott Chacon wrote:
>> I don't think I'm in favour of this policy.  All the major models have
>> been trained on a large variety of code from a large variety of sources,
>> including sources such as news reports or personal websites that do not
>> allow copying, modification, or distribution.  Given that LLMs are known
>> to reproduce portions of their training set or craft code or text which
>> is very similar to items in the training set, how can anyone honestly
>> assert the DCO without knowing all of the sources that were used to
>> create it?
> 
> I agree that generated output can reproduce material we don't have
> permission to distribute (though I think this is incredibly rare for
> anything complex). What I question is whether that possibility means
> knowing every source in the training set is necessary to make any DCO
> certification.
> 
> Human contributors have also read code under many different licenses
> (and news articles and blogs) . We don't ask them to account for
> everything they've ever read before signing off on a patch.

There is a difference between “learning from existing code” which builds up
experience and professional capability and “copy & pasting” code from different
sources and putting it together.

The first (learning from existing code) is a product of whoever writes the
code, based on its mental model and experience developed, the second is a simple
violation of the contribution guidelines.

Where AI stays? LLMs are a simple processing of a large amount of code
for statistically detecting which part of the copy need to be copy&pasted
and blended together. Even though they had a “learning” path in terms of
developing an understanding and building a mental model (they don’t, at
least at the moment), *THEY* would be the author of the code, not the
person that just “pressed the button” on the AI model.

Should we use fully-generated AI contributions? When “fully-generated”
means code that has been totally written and reviewed by agents autonomously,
then I believe the answer should be a sound no, even just for the
lack of a legal framework around it on who is taking the responsibility
of what has been contributed.

> We do
> expect them to have the right to submit the actual contribution and to
> respect the licenses of material they incorporate. Again, Red Hat, the
> Linux Foundation and the SFC all now state that LLM generated code is
> acceptable and compatible with DCO requirements.

Do you have a link to their exact statement?
Do they really allow *fully AI generated* code contributions where the
human has no understanding of the lines of code generated?

If that was the case, how can you manage a review of code that wasn’t
fully guided and understood by the author? Are you foreseeing an agent
answering the mailing list to the comments of the reviewers?

Even though we would allow *fully* AI-generated code, who is really
the contribution from? The “assumed human author”? The LLM?
The data that LLM is trained on? The company that developed the LLM?
Nobody?

Also, imagine that the code generated *did contain* malicious
backdoors, who is liable for it?

There are currently lawsuit in progress against companies that have
trained LLMs on allegedly copyrighted material, we don’t know yet
where they’ll end up to.

> 
>> I'm a distributor of Git and I don't want to be sued or arrested because
>> I end up distributing code that I don't have the right to distribute.
>> Large companies may have lawyers and lots of money to fight those
>> claims, but I do not (nor does the Git project) and I don't want to
>> spend my resources fighting allegations of copyright infringement or
>> have my reputation besmirched for that reason.  Just because other
>> projects think it's okay to do legally and ethically questionable things
>> doesn't mean we should as well.
> 
> Again, a lot of my argumentation here was directly taken from the
> SFC's recommendations [1], which Git is a member project of.
> 
> https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html

I am fully aligned and in agreement with SFC’s recommendations, which is
very much aligned with my thinking.
If you are the author of the code, even if was created with the assistance of 
of LLMs, then you are fully responsible for each and every line of it and you
must understand it and approve it yourself.

I personally use LLMs *a lot* for giving me ideas and reviewing my code
or learning about something new that I never explored before. The end goal
of LLMs for me is to improve my capabilities as a developer, not to generate
contributions on my behalf.

My success criteria after having used LLMs is: I am a better developer now?
Am I more capable of writing better software and having a deeper understanding of 
existing code?

P.S. My comment was 100% human-generated, including spelling mistakes
and awkward expressions (I’m not mother tongue).

Luca.

> 
>> I'll add that if Git were to include a portion of my MIT- or
>> BSD-licensed code without including a copyright or permission notice
>> because it was laundered through an LLM, I would absolutely file a
>> copyright complaint, and rightfully so.
>> 
>> I refer you to policies from other major open source projects that cover
>> this exact provenance issue:
>> 
>> * Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy
>> * NetBSD: https://www.netbsd.org/developers/commit-guidelines.html
> 
> I'm not saying that other projects haven't taken positions as
> conservative as Git's current policy. I'm saying that much larger
> projects with much larger legal surface area such as Linux have
> adopted more progressive ones. Linus is fine with it on a project with
> the same DCO, the same license, and honestly, a lot more legal
> scrutiny.
> 
>> I also will point out the notes from the Contributor Summit where we
>> discussed this issue in some depth and proposed an approach for further
>> discussion.
> 
> It was unclear from the notes what a "vote" meant exactly. It looked
> like Taylor was roughly tasked with providing a Linux-style position,
> so I thought I would help out by providing a first version of that
> position.
> 
> This was written to follow the current SFC guidelines. I would
> encourage Junio to have them read it, but I tried my best to follow
> it's guidelines and advice.
> 
> Scott
> 


^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-08  4:49     ` Scott Chacon
  2026-10-08  5:29       ` Luca Milanesio
@ 2026-10-08  5:53       ` Kristoffer Haugsbakk
  2026-10-08 16:15       ` brian m. carlson
  2 siblings, 0 replies; 12+ messages in thread
From: Kristoffer Haugsbakk @ 2026-10-08  5:53 UTC (permalink / raw)
  To: Scott Chacon, brian m. carlson, Scott Chacon, git

On Thu, Oct 8, 2026, at 06:49, Scott Chacon wrote:
> On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson
> <sandals@crustytoothpaste.net> wrote:
>>[snip]
>
> I agree that generated output can reproduce material we don't have
> permission to distribute (though I think this is incredibly rare for
> anything complex). What I question is whether that possibility means
> knowing every source in the training set is necessary to make any DCO
> certification.
>
> Human contributors have also read code under many different licenses
> (and news articles and blogs) . We don't ask them to account for
> everything they've ever read before signing off on a patch.

Metaphors gone amok. You don’t regulate how submarines and human bodies
can operate in territorial waters based on the fact that they both swim.[1]

Corporations already have non-compete clauses in order to keep knowledge
workers from applying their braincraft to competing businesses.

>[snip]
>
>> I'm a distributor of Git and I don't want to be sued or arrested because
>> I end up distributing code that I don't have the right to distribute.
>>[snip]
>
>[snip]
>
>>[snip]
>
> I'm not saying that other projects haven't taken positions as
> conservative as Git's current policy. I'm saying that much larger
> projects with much larger legal surface area such as Linux have
> adopted more progressive ones. Linus is fine with it on a project with
> the same DCO, the same license, and honestly, a lot more legal
> scrutiny.

Honestly, if an individual wants to avoid the risk of getting sued over
copyright then that’s their prerogative.

>[snip]

[1] I’m not a lawyer so maybe one does.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 22:44   ` Junio C Hamano
@ 2026-10-08 13:53     ` Scott Chacon
  2026-10-09 10:22       ` Patrick Steinhardt
  0 siblings, 1 reply; 12+ messages in thread
From: Scott Chacon @ 2026-10-08 13:53 UTC (permalink / raw)
  To: Junio C Hamano; +Cc: Scott Chacon, git

On Thu, Oct 8, 2026 at 12:44 AM Junio C Hamano <gitster@pobox.com> wrote:
> Scott Chacon <scott@gitbutler.net> writes:
> > As an example, an OpenAI model was used to help me research, compare and
> > craft the appropriate legal language for this policy change to help us
> > match the modern, legally reviewed approaches now taken by peer GPL
> > projects such as the Linux kernel [1].
> >
> > [1] https://docs.kernel.org/process/coding-assistants.html
>
> That makes it sound as if this is just as legally sound as what the
> kernel project uses.  However, the only assurance we get (unless you
> are willing to act as our lawyer, and I do not know if you are one)
> is that an OpenAI model produced plausible-sounding utterances.

Well, honestly, most lawyers I know only barely produce plausible
sounding utterances.

I read the change and edited it where I thought clarification was
needed. But this is different from, say, a blog post or emails, which
I personally never write via LLM because I don't like the voice.
SubmittingPatches is supposed to be dry and factual. Legal guidance is
supposed to be neutral. I've found LLMs quite good at producing
correct legal documents. Again, unjokingly this time, better than most
human lawyers (and I deal with a lot of them).

It's not unlike a code-based LLM contribution. I had it generated with
specific guidance and context of what I wanted the change to be,
because it's faster, and I spent my time reviewing and editing the
result to get the text I was looking for.

I would love to have you pass this by the SFC, because most of the
guidance I gave my agent was _their_ guidelines.

> It looks, at least to me, that there is not much that can be
> meaningfully enforced by reviewers and followed by contributors in
> the above text.  It seems to be little more than "the world would be
> a wonderful place if everybody behaved this way."
>
> A violation of "concise and relevant" seems to be the recent trend
> of much AI-generated slop, so it may be a good suggestion to give
> today.  But would we need to update it once the trend of text
> generated by AI tools becomes "concise and relevant" nonsense that
> merely sounds plausible?  What if an "AI-assisted" contributor lacks
> common sense to tell between plausible-sounding nonsense and a
> well-written description?  What if reviewers get too many such
> "contributions" and cannot allocate enough review bandwidth to sift
> good contributions from plausible-sounding nonsense?

This is a fair point, but you'll get unreviewed crap either way. I'm
sure you already are. However, I don't think people who submit
complete bullshit are reading the SubmittingPatches file in the first
place, so I'm not sure that opening this wording up a little is going
to make much of a difference here.

My recent patch series converting the sha1dc is a possible example. I
don't understand all of the code it wrote. I read through it, but
there are some crazy tables and complex math in there. The first pass
did a weird Rust to C machine translation rather than reimplement it
in more idiomatic C, so I had it rewrite that - so there was some
approach guidance, but again, I wasn't hand crafting the code. I did,
however, spend a lot of time and resources testing and benchmarking it
on multiple architectures so that I was reasonably confident that it
was fast and correct.

But I hesitated to submit it at all because I knew the policy. I only
sent it so that if someone at GitHub or OpenAI or whatever wanted to
use it in an internal fork so they could save a ton of CPU, this would
be a way to get the implementation. I was aware that, although I
believe the patch is quite reasonable and valuable, due to the
conservative AI policies of this project, it would not seriously be
considered no matter what.

My point with this change is to open the possibility for AI assisted
change that is reasonable, similar to the Linux kernel's approach.

> > +The <<dco,Developer's Certificate of Origin>> applies unchanged. Only a
> > +human can make that certification; an AI tool cannot sign off on your
> > +behalf. Consider the origin and licensing of generated material,
> > +including any third-party material it reproduces, and comply with
> > +applicable license and attribution requirements. A tool's assurance
> > +that its output is original or compatible with our license is not a
> > +substitute for checking those requirements. If you cannot certify the
> > +DCO for a contribution, do not submit it.
>
> Again, this is a good aspiration to have, but I doubt that anyone
> can practically certify that the output of an LLM is devoid of
> content borrowed from problematic sources under the rule the text
> above gives.  Would it not be more useful to help contributors by
> defining what not to do more clearly?  Our current text says as much
> more directly: you cannot practically certify, so do not send in
> AI-generated slop, period.

There is a good section on this DCO issue in the Red Hat article on
navigating legal issues around AI [1] where they state that "the DCO
has never been interpreted to require that every line of a
contribution must be the personal creative expression of the
contributor or another human developer". I think that this suggested
paragraph in my patch is a fairly clear interpretation of this stance,
but I can give it another pass if there are more specifics you would
like covered.

[1] https://www.redhat.com/en/blog/ai-assisted-development-and-open-source-navigating-legal-issues

> > +Disclose substantial AI assistance in each affected commit with an
> > +`Assisted-by:` trailer naming the tool and, when available, its model
> > +or version. For example:
> > +
> > +....
> > +     Assisted-by: ExampleTool version 1.2
> > +....
>
> I thought the kernel guidelines instructed us to say only "LLM"
> these days, to avoid giving free advertising.  On the other hand,
> they ask contributors to also list non-LLM tools, like coccinelle
> and clang-tidy, that were used in their machine-assisted
> contributions.  I am undecided on the merit of specifying the
> exact model and version, but listing non-LLM tools alongside
> materials for independent reproduction looks like a good idea.

The kernel guidelines do say "LLM" only, but the SFC guidelines (part 5) says:

"Part of the contribution process should (at least) include a
disclosure of what LLM-gen-AI system was used, its version (as these
system change over time), and a brief description of how the system
assisted the contributor. This information should be included in a
machine-readable format in commit logs."

I'm happy to go back to the kernel guidelines, but since Git is an SFC
project, I figured this is what they would be more comfortable with.

> > +Maintainers may request more explanation, testing, or information about
> > +provenance, and may decline contributions they cannot confidently
> > +assess.
>
> The text of the kernel guidelines appears to give maintainers more
> latitude (cf. https://docs.kernel.org/process/generated-content.html).
> They can treat it just like any other contribution, reject it
> outright, or choose any approach in between.  The proposed text above
> does not account for cases where reviewers simply lack the bandwidth
> to even think about what explanation and proof to request, and it
> makes it sound as if declining a submission in such a case an unfair
> rejection.

I mean, no matter what this says, you can pretty much do whatever you
want. I'm happy to change it to any amount of latitude the maintainer
has that you want to convey. Or remove it entirely.

Scott

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-08  4:49     ` Scott Chacon
  2026-10-08  5:29       ` Luca Milanesio
  2026-10-08  5:53       ` Kristoffer Haugsbakk
@ 2026-10-08 16:15       ` brian m. carlson
  2 siblings, 0 replies; 12+ messages in thread
From: brian m. carlson @ 2026-10-08 16:15 UTC (permalink / raw)
  To: Scott Chacon, git

[-- Attachment #1: Type: text/plain, Size: 3306 bytes --]

On 2026-10-08 at 04:49:27, Scott Chacon wrote:
> Again, a lot of my argumentation here was directly taken from the
> SFC's recommendations [1], which Git is a member project of.
> 
> https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html

I'm in agreement with most of those policies.  I'm just not in agreement
that we should accept LLM-generated contributions and that document
doesn't say we should.

What it does say is that we shouldn't shun people who submit
LLM-generated contributions even if that violates our policies, and I
think we've respected that.  Every time this comes up—and it comes up
more often than it should, given that we have a documented policy and
that people should know to look for one—we've handled this graciously.
As far as I know, nobody has been blocked or excluded for having sent an
LLM-generated patch to Git and we usually explain the policy in a calm,
rational way.

Section 8 says we should avoid jumping to legal conclusions.  I agree; I
have said consistently on the list that one of the reasons we should
reject LLM-generated contributions is because the legal status is
unclear.  I have strong views that LLMs are unethical because of the way
they've been trained and the lack of credit, among other reasons, but I
have been very clear that the legality is uncertain.

The final thing that it says that kind of supports your argument is §11.
However, it's not the case that accepting LLM-generated contributions
would massively accelerate improvements to our codebase.  Git has a
reputation for high quality and we perform thorough reviews.  Those are
already a bottleneck for us even with only human-generated code and
welcoming LLM-generated code and documentation would submerge us under a
deluge of patches.  As we discussed at the Contributor's Summit, we're
already underwater on the security list due to the flood of LLM-assisted
bug reports, a number of which are of dubious quality, and we shouldn't
replicate that on the public list as well.

One thing that supports my argument is that we should support people who
"outright reject LLM-gen-AI systems."  Because of the way this list
works, if we accept LLM-generated code, contributors who don't want to
work with that content are going to receive unwanted patches that are
CC'd to them and then have to deal with those, whereas in a project like
Rust, one can simply block the LLM bots and then never have to deal with
that content at all.  So I don't think we can honour that term while we
allow LLM-generated content unless we change our development approach.

I would also say that we are not obligated to follow SFC's guidance at
all.  SFC also recommends that we leave GitHub[0], which obviously
neither of us are following, nor are most of our contributors, and many
people would disagree.  Regardless of their guidance, our project can
set our own policies and structure as we see fit.  QEMU is also an SFC
project and has adopted a policy very similar to ours[1], so there is
clearly precedent here.

[0] https://sfconservancy.org/GiveUpGitHub/
[1] https://gitlab.com/qemu-project/qemu/-/blob/master/docs/devel/code-provenance.rst?ref_type=heads
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance
  2026-10-07 14:29 [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance Scott Chacon
  2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
@ 2026-10-08 18:10 ` D. Ben Knoble
  1 sibling, 0 replies; 12+ messages in thread
From: D. Ben Knoble @ 2026-10-08 18:10 UTC (permalink / raw)
  To: Scott Chacon; +Cc: git

On Wed, Oct 7, 2026 at 10:35 AM Scott Chacon <scott@gitbutler.net> wrote:
>
> After the AI discussion at the contributors' summit [1], I'd like to submit
> a concrete alternative for allowing AI generated work responsibly. This moves
> us closer to Linux's approach: use the tools you find helpful, but take
> responsibility for what you send.

If you'll forgive me going a *bit* off-topic, I'd like to point
interested readers towards some starting points for fascinating
historical reading. It turns out, folks have questioned the use of AI
and tools for a long time, and we have a body of writing that examines
the impact tools have on us and we on them. So there's really no such
thing as "just a tool", in an important sense :)

Some starting points for going deeper are linked in the following
articles (on my personal blog, but most of the thinking and references
belong to others):

- https://benknoble.github.io/blog/2026/09/02/tool-dialogue/
- https://benknoble.github.io/blog/2026/09/20/turkle/
- https://benknoble.github.io/blog/2026/09/21/once-more/

And relatedly, some reading notes on folks studying what it's like to
actually *use* LLMs (albeit not exactly in our context):
https://benknoble.github.io/blog/2026/10/02/reading-notes-configuration-work/

I don't think this has much bearing on the conversation about Git's
policy other than to say: I don't think we ought to justify our use by
saying it's "just a tool like any other"---not because it's unlike
other tools so much as because the tools we choose and use matter;
they affect us and the people around us.

(I do appreciate that Scott's proposal has elements of taking
responsibility for what you produce. At RacketCon last weekend, a
maintainer put it a bit differently: "Review scales less than code;
manage your own backpressure.")

Somewhat more on the policy side, there is also the (polemically
written, but valuable nonetheless?) "AI Pascal's wager"
(https://ploum.net/2026-10-01-pascal_wager.html). To quote the
conclusion (note: "abandon your project" *also* seems like a
fear-mongering the stance the author tries to reject from AI boosters,
but let's try to take the rest of the argument in good faith in spite
of that):

> While mass marketing is trying to instil a Fear of Missing Out hysteria, the most rational and pragmatic approach is to strongly reject all AI-generated contributions to your projects. For now.
>
> Someday, we might realise that LLMs are doing good in the world, that they are evolving toward ethical, reliable, sustainable solutions, and that people who use them are happier (try to read that sentence again without rolling your eyes). If that really happens, you could always change your AI policy. It will cost you nothing.
>
> But if you let the slop in now, you may regret it forever… You may be forced to abandon your project.
>
> On the other hand, if you refuse AI-generated contributions to your project right now, the worst very hypothetical regret you could ever have is "I should probably have done it sooner".
>
> The conclusion is simple: If you are AI-agnostic, the pragmatic course of action is to strongly refuse any AI-generated contribution to your project.

"Wish we'd done it sooner" is definitely the boat I'd, personally,
rather be in down the line.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-08 13:53     ` Scott Chacon
@ 2026-10-09 10:22       ` Patrick Steinhardt
  2026-10-09 20:39         ` Junio C Hamano
  0 siblings, 1 reply; 12+ messages in thread
From: Patrick Steinhardt @ 2026-10-09 10:22 UTC (permalink / raw)
  To: Scott Chacon; +Cc: Junio C Hamano, Scott Chacon, git

On Thu, Oct 08, 2026 at 03:53:11PM +0200, Scott Chacon wrote:
> On Thu, Oct 8, 2026 at 12:44 AM Junio C Hamano <gitster@pobox.com> wrote:
> > Scott Chacon <scott@gitbutler.net> writes:
[snip]
> > It looks, at least to me, that there is not much that can be
> > meaningfully enforced by reviewers and followed by contributors in
> > the above text.  It seems to be little more than "the world would be
> > a wonderful place if everybody behaved this way."
> >
> > A violation of "concise and relevant" seems to be the recent trend
> > of much AI-generated slop, so it may be a good suggestion to give
> > today.  But would we need to update it once the trend of text
> > generated by AI tools becomes "concise and relevant" nonsense that
> > merely sounds plausible?  What if an "AI-assisted" contributor lacks
> > common sense to tell between plausible-sounding nonsense and a
> > well-written description?  What if reviewers get too many such
> > "contributions" and cannot allocate enough review bandwidth to sift
> > good contributions from plausible-sounding nonsense?
> 
> This is a fair point, but you'll get unreviewed crap either way. I'm
> sure you already are. However, I don't think people who submit
> complete bullshit are reading the SubmittingPatches file in the first
> place, so I'm not sure that opening this wording up a little is going
> to make much of a difference here.
> 
> My recent patch series converting the sha1dc is a possible example. I
> don't understand all of the code it wrote. I read through it, but
> there are some crazy tables and complex math in there. The first pass
> did a weird Rust to C machine translation rather than reimplement it
> in more idiomatic C, so I had it rewrite that - so there was some
> approach guidance, but again, I wasn't hand crafting the code. I did,
> however, spend a lot of time and resources testing and benchmarking it
> on multiple architectures so that I was reasonably confident that it
> was fast and correct.
> 
> But I hesitated to submit it at all because I knew the policy. I only
> sent it so that if someone at GitHub or OpenAI or whatever wanted to
> use it in an internal fork so they could save a ton of CPU, this would
> be a way to get the implementation. I was aware that, although I
> believe the patch is quite reasonable and valuable, due to the
> conservative AI policies of this project, it would not seriously be
> considered no matter what.

I would argue that this is a good thing though. We want people to thing
twice before submitting code that they don't fully understand, don't we?
Otherwise we will get even more slop than we already get.

> My point with this change is to open the possibility for AI assisted
> change that is reasonable, similar to the Linux kernel's approach.

We already are accepting AI-generated code. So if the change would have
the effect that people _don't_ think twice anymore about sending their
AI generated code to the mailing list then I think that's a net-negative
change.

The bottleneck of the Git project has never really been the amount of
code that people can write, but the number of developers that we have
reviewing it. And you can feel that this bottleneck is getting tighter
now with AI -- over the last couple months I have spent way more time
reviewing stuff, and the number of times that I noticed too late that
I'm reviewing slop is going up steadily. I guess for Junio that must be
even worse.

So I think loosening our AI policy shouldn't go without finding
solutions for this problem first, because otherwise I feel like we are
just going to make a preexisting problem significantly worse.

Patrick

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
  2026-10-09 10:22       ` Patrick Steinhardt
@ 2026-10-09 20:39         ` Junio C Hamano
  0 siblings, 0 replies; 12+ messages in thread
From: Junio C Hamano @ 2026-10-09 20:39 UTC (permalink / raw)
  To: Patrick Steinhardt; +Cc: Scott Chacon, Scott Chacon, git

Patrick Steinhardt <ps@pks.im> writes:

> So I think loosening our AI policy shouldn't go without finding
> solutions for this problem first, because otherwise I feel like we are
> just going to make a preexisting problem significantly worse.

Well articulated.  I have nothing more to add.  Thanks.

^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-10-09 20:39 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 14:29 [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance Scott Chacon
2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
2026-10-07 21:42   ` brian m. carlson
2026-10-08  4:49     ` Scott Chacon
2026-10-08  5:29       ` Luca Milanesio
2026-10-08  5:53       ` Kristoffer Haugsbakk
2026-10-08 16:15       ` brian m. carlson
2026-10-07 22:44   ` Junio C Hamano
2026-10-08 13:53     ` Scott Chacon
2026-10-09 10:22       ` Patrick Steinhardt
2026-10-09 20:39         ` Junio C Hamano
2026-10-08 18:10 ` [RFC PATCH 0/1] " D. Ben Knoble

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox