git.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: "brian m. carlson" <sandals@crustytoothpaste.net>
To: Scott Chacon <scott@gitbutler.net>
Cc: git@vger.kernel.org
Subject: Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
Date: Wed, 7 Oct 2026 21:42:35 +0000	[thread overview]
Message-ID: <asa8ymCv4hoRJcZM@fruit.crustytoothpaste.net> (raw)
In-Reply-To: <20261007142954.31761-2-scott@gitbutler.net>

[-- Attachment #1: Type: text/plain, Size: 3605 bytes --]

On 2026-10-07 at 14:29:54, Scott Chacon wrote:
> The AI section encourages careful use of AI tools, but also says we
> will reject anything that looks AI generated. That leaves contributors
> without a clear path for submitting useful, reviewed, understood work and
> can discourage disclosure of the assistance they received.
> 
> Allow AI-assisted contributions under the usual quality and licensing
> requirements. Require human understanding, appropriate testing, and
> disclosure of substantial assistance. Retain the DCO without changing
> its terms, and require contributors to consider provenance and meet
> applicable license obligations. Reviewers can ask for further evidence
> or decline work they cannot confidently assess.
> 
> Replace the appearance-based rejection rule with these concrete
> expectations. AI assistance neither excuses an inadequate submission
> nor prevents an otherwise acceptable one from being considered.
> 
> As an example, an OpenAI model was used to help me research, compare and
> craft the appropriate legal language for this policy change to help us
> match the modern, legally reviewed approaches now taken by peer GPL
> projects such as the Linux kernel [1].

I don't think I'm in favour of this policy.  All the major models have
been trained on a large variety of code from a large variety of sources,
including sources such as news reports or personal websites that do not
allow copying, modification, or distribution.  Given that LLMs are known
to reproduce portions of their training set or craft code or text which
is very similar to items in the training set, how can anyone honestly
assert the DCO without knowing all of the sources that were used to
create it?

Even if the model were, for instance, trained only on MIT-licensed code,
the license still requires a copyright and permission notice on every
copy, so the fact that the code generated from an LLM doesn't contain
that would seem to violate the license and prohibit us from using it.

The DCO was created to help us unambiguously state that the code is
acceptable to be included to avoid any later claims that the code was
copied from somewhere that it shouldn't have been.  Given the fact that
nobody knows what the sources are with a current LLM, it doesn't seem
that a reasonable person could make such an assertion.

I'm a distributor of Git and I don't want to be sued or arrested because
I end up distributing code that I don't have the right to distribute.
Large companies may have lawyers and lots of money to fight those
claims, but I do not (nor does the Git project) and I don't want to
spend my resources fighting allegations of copyright infringement or
have my reputation besmirched for that reason.  Just because other
projects think it's okay to do legally and ethically questionable things
doesn't mean we should as well.

I'll add that if Git were to include a portion of my MIT- or
BSD-licensed code without including a copyright or permission notice
because it was laundered through an LLM, I would absolutely file a
copyright complaint, and rightfully so.

I refer you to policies from other major open source projects that cover
this exact provenance issue:

* Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy
* NetBSD: https://www.netbsd.org/developers/commit-guidelines.html

I also will point out the notes from the Contributor Summit where we
discussed this issue in some depth and proposed an approach for further
discussion.
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 325 bytes --]

  reply	other threads:[~2026-10-07 21:42 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 14:29 [RFC PATCH 0/1] SubmittingPatches: allow responsible AI assistance Scott Chacon
2026-10-07 14:29 ` [RFC PATCH 1/1] " Scott Chacon
2026-10-07 21:42   ` brian m. carlson [this message]
2026-10-08  4:49     ` Scott Chacon
2026-10-08  5:29       ` Luca Milanesio
2026-10-08  5:53       ` Kristoffer Haugsbakk
2026-10-08 16:15       ` brian m. carlson
2026-10-07 22:44   ` Junio C Hamano
2026-10-08 13:53     ` Scott Chacon
2026-10-09 10:22       ` Patrick Steinhardt
2026-10-09 20:39         ` Junio C Hamano
2026-10-08 18:10 ` [RFC PATCH 0/1] " D. Ben Knoble

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asa8ymCv4hoRJcZM@fruit.crustytoothpaste.net \
    --to=sandals@crustytoothpaste.net \
    --cc=git@vger.kernel.org \
    --cc=scott@gitbutler.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).