All of lore.kernel.org
 help / color / mirror / Atom feed
From: Paolo Bonzini <pbonzini@redhat.com>
To: "Daniel P. Berrangé" <berrange@redhat.com>
Cc: qemu-devel@nongnu.org, "Michael S . Tsirkin" <mst@redhat.com>,
	"Alex Bennée" <alex.bennee@linaro.org>,
	"Alistair Francis" <alistair.francis@wdc.com>,
	"BALATON Zoltan" <balaton@eik.bme.hu>,
	"Fabiano Rosas" <farosas@suse.de>,
	"Kevin Wolf" <kwolf@redhat.com>,
	"Peter Maydell" <peter.maydell@linaro.org>,
	"Warner Losh" <imp@bsdimp.com>,
	"Philippe Mathieu-Daudé" <philmd@linaro.org>,
	"Paolo Bonzini" <bonzini@gnu.org>
Subject: Re: [PATCH 1/4] docs/devel: introduce a new policy on AI-generated contributions
Date: Thu, 3 Sep 2026 12:05:36 +0200	[thread overview]
Message-ID: <67db7357-c1b9-4391-9b03-fdcc1fbb257f@redhat.com> (raw)
In-Reply-To: <aphKOs7k92Pue40q@redhat.com>

On 9/2/26 18:09, Daniel P. Berrangé wrote:
> On Tue, Sep 01, 2026 at 06:14:19PM +0200, Paolo Bonzini wrote:
>> On top of this, several maintainers have pointed out that contributions
>> that can be reasonably suspected to be AI-assisted or more have been
>> posted and included.
> 
> I tend to view this problem as a consequence of our failure to provide
> an AGENTS.md file describing our policy, rather than fully attributed
> to the policy itself being wrong.  We were naive in thinking that a
> page buried in our docs was sufficient to make people & and agents
> aware of our policy.

It is also giving a measure of what people find AI to be useful for, in 
the context of QEMU.  Certainly an underestimation of how much they 
would use it if the policy was changed, but it's something.

>> Since the policy has been written, other projects have discussed and
>> taken their stance on AI contributions.  These vary from full prohibition
>> (though typically motivated by social reasons rather than legal, e.g.
>> for Zig) to limited use (e.g. GCC, allowing small contributions and
>> tests to use AI), to cautious experimentation.
>>
>> This proposed policy seeks to implement the cautious experimentation
>> approach, inspired mostly by the Rust project's policy and by Software
>> Freedom Conservancy's own recommendations on LLMs and generative AI.
> 
> IMHO the "cautious experimentation" phrasing is effectively
> marketing words for a policy that is "AI for anything".

I don't think it is.  Any policy builds on trust of the maintainers, and 
we know the community is not going to vibe code a rewrite of QEMU in Rust.

There are multiple axes, and this policy+AGENTS.md combo is not the most 
liberal on any axis.  For example GStreamer has a stricter AGENTS.md but 
basically no policy; that's more "AI for anything" than this proposal is.

Yes, it's intentionally leaving out any legal risks unlike the GCC 
policy.  I'm trusting Conservancy on that - they do not just have 
lawyers, they are in some sense "our" lawyers and I don't think they 
wrote their recommendations lightly.

> The "limited use" scenario from GCC is meaningfully different
> as it is attempting to limit the legal liability by restricting
> the scope of work to things that are unlikely to meet the
> threshold for copyright / licensing / legal concerns.

>> Conservancy in particular provides this point to alleviate the concern
>> that motivated the policy, about whether the submitter has the legal
>> right to contribute the code and about unintentional reproduction of
>> copyrighted code:
>>
>>     "Copyleft Everything" remains the best viable and safest approach
>>     Certainly those who want to release FOSS under non-copyleft licenses
>>     have more to worry about when using these tools.
> 
> My best interpretation is that it is trying to give reassurance
> that if the AI output were to be deemed a derived work of part
> of the training material, then projects are safer if they are
> Copyleft. [...]
> 
> That rationalization only works if the set of training material
> licenses forms a linear progression of restrictions with copyleft
> (GPL) at one end. The real training materials is such a jumble
> of licenses that there's no "root" and there are a huge set of
> copyleft variants. [...] Even the GPL has plain GPL vs LGPL vs
> AGPL, and v2-only vs v2-or-later vs v3-only. Even before the
> days of AI this was a compliance minefield

"Most viable and safest" is certainly not a 100% guarantee.  IANAL but 
what you want is the legal equivalent of the swiss cheese model where 
copyleft is only one of several mitigating factors.  These include the 
fact that there are a lot of QEMU forks in training materials, as you 
pointed out when discussing mechanical changes; the AGENTS.md 
instructions to involve the user in the design; "de minimis" and fair 
use exceptions (you probably don't want to get there but they exist); 
and so on.

> My concern with "pre-arrangement" is how we end up applying the
> rule in practice and whether the community dynamics that result
> from that are positive or negative ?
> 
> My own historical experiences with communities or processes where
> contributions requires pre-arrangement or scheduling were really
> quite negative. It frequently kills opportunistic or spontaneous
> work, and can result in a closed club which is hard to newcomers
> to break into.

I don't think putting pre-arrangement in an AI policy is saying anything 
new, it only makes it explicit in the area of highest risk.

Personal example: I did feel bad for including *my* implementation of 
AVX over the previous two, just because my employer didn't need one at a 
time and the review effort would have been massive (higher than writing 
my own when Red Hat did want one).

Pre-arrangement would have helped, and *now* I could say "hey, ask the 
AI to sketch a new x86 decoder with this and this characteristic, and 
let's see where that takes us".

Now *I* wouldn't use AI today to rewrite the x86 decoder either, but 
pre-arrangement can change not just the balance but also the dynamic 
between maintainer and contributor.  Maintainers have *a lot* more power 
to ask for changes if the effort to do them is comparably lower. 
Ideally that filters for AI users that are curious and interested in 
learning the underlying choices.

Of course, maybe I am wrong.

> I also conceptually dislike a policy which will lead to a situation
> where different rules will apply to different subsystems, depending
> on the preferences of individual maintainers. Work is also not
> always easily contained to subsystems, prerequisite refactoring
> can quickly spread it tentacles out.
> 
> Consider hypothetically a net subsystem maintainer agrees to
> use of AI for generating a large piece of code, and something in
> that work requires a change to QOM or QDev or QAPI. This quickly
> ends up exposing multiple other maintainers to TODO items from
> the AI generated contribution.

Indeed, it's not hypothetical even - see my reply to Peter about Alex's 
qdev/QOM refactoring.  But those large cases have *already* been done 
with pre-arrangement and one maintainer vouching for them, so we have a 
precedent.

> So again, IMHO, "cautious experimentation" with "pre-arrangement"
> is effectively "AI for anything" and all maintainers exposed to
> varying levels, but contributors need to get into a club first.

It's a possible outcome, it's not the only one though (or if it is, the 
club already exists and does not even include all maintainers---which is 
a problem in and of itself).

>> In any case, use of AI does not relax any other contribution requirement:
>> authors still comply with the DCO and take responsibility for the whole
>> patch via Signed-off-by.
> 
> When agents output are involved the DCO just rubber stamp exercise,
> as there's no practical way any contributor can understand whether
> there are legal concerns with the code the agent spat out unless it
> is so short as to not meet the threshold for copyright.

It still acknowledges the fact that, for example, the person's employer 
does not forbid him for contributing to QEMU.  And while it's weaker, 
AI-user-for + Signed-off-by protects QEMU more than "don't ask, don't tell".

[I won't rehash the same arguments below; I understand why you made them 
in the context of both the cover letter and the actual text]

>> +The following items **MUST** be written by humans:
> 
> This is enumerating three concrete examples, of a more
> general concept of "The QEMU community is a collaboration
> between humans".  IOW, we don't want AITM (AI In The Middle)
> for our communications. Can we say this explicitly
> 
> "The QEMU community is a collaboration between humans
>   and thus communications must NOT be directed through
>   an AI agents facade. This implies that the following
>   items MUST be written by humans:"

Sure.

>> +- use of LLMs to help generating parts of a larger patch---a test case, a
>> +  parser, boilerplate code for a new API, a tool to help performing
>> +  mechanical changes, etc. These are generally allowed, but disclosure
>> +  is recommended.
> 
> I don't think disclosure should be optional here, most especially for
> tests cases it needs to be mandatory IMHO.

What about "disclosure is highly recommended for non-trivial, functional 
code"?  The idea here was to avoid lowering the AI-used-for SNR and avoid

AI-used-for: code (turning CSV data into an array)

>> +There is no requirement to include your prompts or summarize the
>> +conversation in the commit message or cover letter.
> 
> I would be stronger and say we explicitly do NOT want the prompts
> or conversation history. If there was info in the prompts that is
> relevant to the reviewer, then include that info as natural language
> in the commit message, not a cut+paste of the prompts.
> 
>> +QEMU does *not* use ``Assisted-by``, ``Co-authored-by`` or ``Generated-by``
>> +trailers to indicate AI usage.  In particular, it is not necessary to
>> +specify the exact AI model or tool used to create the commit.
> 
> This says they're not required, but also doesn't forbid them,
> which leaves rather a gray zone. If we don't want to require
> them (which I think is correct, as this is just free advertizing
> for largely commercial tools), then IMHO its preferable to make
> checkpatch.pl explicitly reject them.

Sure.

> This (and SFC's recommendations) comes across as trying to
> square-the-circle.
> 
> Effectively the policy is saying that we have no choice but to accept
> them, despite reservations people might have (legal or social or
> environmental), as they've become too commonplace in the industry to
> decline.
I think this is a bit oversimplifying, but I can't deny that there's a 
kernel of truth in there.

Paolo



  reply	other threads:[~2026-09-03 10:06 UTC|newest]

Thread overview: 36+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-01 16:14 [RFC PATCH v2 0/4] docs/devel, AGENTS.md: a completely new policy on AI-generated contributions Paolo Bonzini
2026-09-01 16:14 ` [PATCH 1/4] docs/devel: introduce a " Paolo Bonzini
2026-09-02  1:38   ` Alistair Francis
2026-09-02  7:09     ` Paolo Bonzini
2026-09-02  9:41       ` Michael S. Tsirkin
2026-09-02 10:08       ` Daniel P. Berrangé
2026-09-02 12:01   ` Peter Maydell
2026-09-03  1:39     ` Alistair Francis
2026-09-03  7:09     ` Paolo Bonzini
2026-09-02 16:09   ` Daniel P. Berrangé
2026-09-03 10:05     ` Paolo Bonzini [this message]
2026-09-01 16:14 ` [PATCH 2/4] AGENTS.md: Let agents know what the participation rules are Paolo Bonzini
2026-09-02  1:44   ` Alistair Francis
2026-09-02 16:26     ` Daniel P. Berrangé
2026-09-02 17:45       ` Paolo Bonzini
2026-09-02 13:42   ` Alex Bennée
2026-09-02 13:47     ` Daniel P. Berrangé
2026-09-02 13:56       ` Michael S. Tsirkin
2026-09-02 14:08         ` Daniel P. Berrangé
2026-09-02 15:15           ` Michael S. Tsirkin
2026-09-02 13:48     ` Michael S. Tsirkin
2026-09-01 16:14 ` [PATCH 3/4] AGENTS.md: make the human part of the design process Paolo Bonzini
2026-09-02  1:44   ` Alistair Francis
2026-09-01 16:14 ` [PATCH 4/4] agents: add CLAUDE.md and GEMINI.md Paolo Bonzini
2026-09-02  1:47   ` Alistair Francis
2026-09-02  4:45     ` Philippe Mathieu-Daudé
2026-09-02  9:46   ` Michael S. Tsirkin
2026-09-02  9:56     ` Paolo Bonzini
2026-09-02 10:06       ` Michael S. Tsirkin
2026-09-02 10:12         ` Daniel P. Berrangé
2026-09-02 13:04           ` Michael S. Tsirkin
2026-09-02 14:30             ` Paolo Bonzini
2026-09-19  0:25           ` Mohamed Mediouni
2026-09-02 15:54 ` [RFC PATCH v2 0/4] docs/devel, AGENTS.md: a completely new policy on AI-generated contributions Peter Maydell
2026-09-02 17:50   ` Alex Bennée
2026-09-03 11:29   ` BALATON Zoltan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=67db7357-c1b9-4391-9b03-fdcc1fbb257f@redhat.com \
    --to=pbonzini@redhat.com \
    --cc=alex.bennee@linaro.org \
    --cc=alistair.francis@wdc.com \
    --cc=balaton@eik.bme.hu \
    --cc=berrange@redhat.com \
    --cc=bonzini@gnu.org \
    --cc=farosas@suse.de \
    --cc=imp@bsdimp.com \
    --cc=kwolf@redhat.com \
    --cc=mst@redhat.com \
    --cc=peter.maydell@linaro.org \
    --cc=philmd@linaro.org \
    --cc=qemu-devel@nongnu.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.