Git development
 help / color / mirror / Atom feed
From: Luca Milanesio <luca.milanesio@gmail.com>
To: git@vger.kernel.org
Cc: Luca Milanesio <luca.milanesio@gmail.com>
Subject: Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
Date: Thu, 8 Oct 2026 06:29:48 +0100	[thread overview]
Message-ID: <6CCE2DB2-E2E2-48A0-B443-95B0AABFE83B@gmail.com> (raw)
In-Reply-To: <CAP2yMa+kgphMe-cpcZSvPSqwm-npUDVp=HaNRW+MPmPzZ_aOXw@mail.gmail.com>



> On 8 Oct 2026, at 05:49, Scott Chacon <schacon@gmail.com> wrote:
> 
> On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson
> <sandals@crustytoothpaste.net> wrote:
>> 
>> On 2026-10-07 at 14:29:54, Scott Chacon wrote:
>> I don't think I'm in favour of this policy.  All the major models have
>> been trained on a large variety of code from a large variety of sources,
>> including sources such as news reports or personal websites that do not
>> allow copying, modification, or distribution.  Given that LLMs are known
>> to reproduce portions of their training set or craft code or text which
>> is very similar to items in the training set, how can anyone honestly
>> assert the DCO without knowing all of the sources that were used to
>> create it?
> 
> I agree that generated output can reproduce material we don't have
> permission to distribute (though I think this is incredibly rare for
> anything complex). What I question is whether that possibility means
> knowing every source in the training set is necessary to make any DCO
> certification.
> 
> Human contributors have also read code under many different licenses
> (and news articles and blogs) . We don't ask them to account for
> everything they've ever read before signing off on a patch.

There is a difference between “learning from existing code” which builds up
experience and professional capability and “copy & pasting” code from different
sources and putting it together.

The first (learning from existing code) is a product of whoever writes the
code, based on its mental model and experience developed, the second is a simple
violation of the contribution guidelines.

Where AI stays? LLMs are a simple processing of a large amount of code
for statistically detecting which part of the copy need to be copy&pasted
and blended together. Even though they had a “learning” path in terms of
developing an understanding and building a mental model (they don’t, at
least at the moment), *THEY* would be the author of the code, not the
person that just “pressed the button” on the AI model.

Should we use fully-generated AI contributions? When “fully-generated”
means code that has been totally written and reviewed by agents autonomously,
then I believe the answer should be a sound no, even just for the
lack of a legal framework around it on who is taking the responsibility
of what has been contributed.

> We do
> expect them to have the right to submit the actual contribution and to
> respect the licenses of material they incorporate. Again, Red Hat, the
> Linux Foundation and the SFC all now state that LLM generated code is
> acceptable and compatible with DCO requirements.

Do you have a link to their exact statement?
Do they really allow *fully AI generated* code contributions where the
human has no understanding of the lines of code generated?

If that was the case, how can you manage a review of code that wasn’t
fully guided and understood by the author? Are you foreseeing an agent
answering the mailing list to the comments of the reviewers?

Even though we would allow *fully* AI-generated code, who is really
the contribution from? The “assumed human author”? The LLM?
The data that LLM is trained on? The company that developed the LLM?
Nobody?

Also, imagine that the code generated *did contain* malicious
backdoors, who is liable for it?

There are currently lawsuit in progress against companies that have
trained LLMs on allegedly copyrighted material, we don’t know yet
where they’ll end up to.

> 
>> I'm a distributor of Git and I don't want to be sued or arrested because
>> I end up distributing code that I don't have the right to distribute.
>> Large companies may have lawyers and lots of money to fight those
>> claims, but I do not (nor does the Git project) and I don't want to
>> spend my resources fighting allegations of copyright infringement or
>> have my reputation besmirched for that reason.  Just because other
>> projects think it's okay to do legally and ethically questionable things
>> doesn't mean we should as well.
> 
> Again, a lot of my argumentation here was directly taken from the
> SFC's recommendations [1], which Git is a member project of.
> 
> https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html

I am fully aligned and in agreement with SFC’s recommendations, which is
very much aligned with my thinking.
If you are the author of the code, even if was created with the assistance of 
of LLMs, then you are fully responsible for each and every line of it and you
must understand it and approve it yourself.

I personally use LLMs *a lot* for giving me ideas and reviewing my code
or learning about something new that I never explored before. The end goal
of LLMs for me is to improve my capabilities as a developer, not to generate
contributions on my behalf.

My success criteria after having used LLMs is: I am a better developer now?
Am I more capable of writing better software and having a deeper understanding of 
existing code?

P.S. My comment was 100% human-generated, including spelling mistakes
and awkward expressions (I’m not mother tongue).

Luca.

> 
>> I'll add that if Git were to include a portion of my MIT- or
>> BSD-licensed code without including a copyright or permission notice
>> because it was laundered through an LLM, I would absolutely file a
>> copyright complaint, and rightfully so.
>> 
>> I refer you to policies from other major open source projects that cover
>> this exact provenance issue:
>> 
>> * Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy
>> * NetBSD: https://www.netbsd.org/developers/commit-guidelines.html
> 
> I'm not saying that other projects haven't taken positions as
> conservative as Git's current policy. I'm saying that much larger
> projects with much larger legal surface area such as Linux have
> adopted more progressive ones. Linus is fine with it on a project with
> the same DCO, the same license, and honestly, a lot more legal
> scrutiny.
> 
>> I also will point out the notes from the Contributor Summit where we
>> discussed this issue in some depth and proposed an approach for further
>> discussion.
> 
> It was unclear from the notes what a "vote" meant exactly. It looked
> like Taylor was roughly tasked with providing a Linux-style position,
> so I thought I would help out by providing a first version of that
> position.
> 
> This was written to follow the current SFC guidelines. I would
> encourage Junio to have them read it, but I tried my best to follow
> it's guidelines and advice.
> 
> Scott
> 


  reply	other threads:[~2026-10-08  5:30 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 14:29 [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance Scott Chacon
2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
2026-10-07 21:42   ` brian m. carlson
2026-10-08  4:49     ` Scott Chacon
2026-10-08  5:29       ` Luca Milanesio [this message]
2026-10-08  5:53       ` Kristoffer Haugsbakk
2026-10-08 16:15       ` brian m. carlson
2026-10-07 22:44   ` Junio C Hamano
2026-10-08 13:53     ` Scott Chacon
2026-10-09 10:22       ` Patrick Steinhardt
2026-10-09 20:39         ` Junio C Hamano
2026-10-08 18:10 ` [RFC PATCH 0/1] " D. Ben Knoble

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=6CCE2DB2-E2E2-48A0-B443-95B0AABFE83B@gmail.com \
    --to=luca.milanesio@gmail.com \
    --cc=git@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox