All of lore.kernel.org
 help / color / mirror / Atom feed
From: Roman Gushchin <roman.gushchin@linux.dev>
To: Jakub Kicinski <kuba@kernel.org>
Cc: "Chris Mason" <mason@kernel.org>,
	 sashiko@lists.linux.dev, ihor.solodrai@linux.dev,
	 ast@kernel.org
Subject: Re: [RFC] reworking the review-prompts subsystem guide
Date: Mon, 05 Oct 2026 21:14:12 +0000	[thread overview]
Message-ID: <7ia4y0cbsvdn.fsf@castle.c.googlers.com> (raw)
In-Reply-To: <20261005125859.3ed00b33@kernel.org> (Jakub Kicinski's message of "Mon, 5 Oct 2026 12:58:59 -0700")

Jakub Kicinski <kuba@kernel.org> writes:

> On Fri, 02 Oct 2026 15:04:00 -0400 Chris Mason wrote:
>> Hi everyone,
>> 
>> I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
>> 
>> https://github.com/masoncl/review-prompts.git subsystem-build
>> 
>> This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion.
>> 
>> The review prompts started with a lot of framework to explain how
>> the kernel works, and the prompts have focused on documenting
>> subsystems in a way that both LLMs and people can use for reviews
>> and writing code.
>> 
>> As models have progressed, they actually need different information
>> than they used to.  Recent models already know how most of the
>> kernel works, but they have some gaps because the kernel is always
>> changing, or because they have some bad assumptions baked in.
>> 
>> My new branch builds subsystem guides by asking a long list of
>> questions, and then extensively reading the sources to find the
>> right answers.  The delta between the LLM's answers and the right
>> answers is the new guide.
>> 
>> This is both much less useful to human readers and much longer.  I'm
>> not sure what to say about the human reader part, but instead of
>> having LLMs read the whole subsystem guide, I shifted to an index
>> where they search for symbols.  This is a better fit for more
>> advanced models, which mostly need updates on how the kernel has
>> changed since they were trained.
>> 
>> The subsystem-build branch was built with both sonnet-5.5 and
>> opus-5.5.  It's the union of all the things either model got wrong,
>> and the tree is setup so we can add builds for other kernel versions
>> (ex: stable kernel series).
>> 
>> I started down this path convinced that I needed to get the
>> subsystem guides into the kernel git tree.  My idea was this was
>> documentation, and it should land somewhere in the kernel for that
>> reason.  I ended up with something that probably shouldn't be in the
>> tree...different models need different things on top of different
>> kernels, and so I'm trying out the build idea.
>> 
>> What I know for sure is the existing review prompts have drifted
>> from mainline Linus.  It's impacting the quality of the reviews, so
>> I plan on working out something in the near future.
>
> I didn't read the whole thing, but makes sense AFAIU.
>
> Closing the loop with Sashiko and/or ML would be great :(
> If the bots read the ML they could both learn false positives and false
> negatives automatically. I suspect most of us thought about this by now.
> If we're doing a redesign should the ingest of ML be part of it?

I plan to do this (and had a prototype in the past), but it's tricky if
we take security seriously. And we absolutely should!
Obviously just giving an agent an access to lore archive is opening a
can of worms in terms of possible prompt injections. So I think we
should do it really carefully. And outside of security considerations,
humans are simple wrong too. So my plan with sashiko is to force it to
try to verify the feedback against the codebase and if it's not possible
trust only maintainers or people with a long history of meaningful
contributions.
And sashiko can (and does in my prototype) generate prompts based on the
feedback and automatically verify that it helps by re-reviewing original
patches. This should mostly eliminate a need for human-generated
prompts.
This is all doable, but probably will take few more week to build and
roll out.

  reply	other threads:[~2026-10-05 21:14 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
2026-10-02 21:18 ` Chuck Lever
2026-10-04  9:17   ` Chris Mason
2026-10-05 19:58 ` Jakub Kicinski
2026-10-05 21:14   ` Roman Gushchin [this message]
2026-10-05 21:53     ` Jakub Kicinski
2026-10-05 23:37     ` Ihor Solodrai
2026-10-06 21:22 ` Ihor Solodrai
2026-10-07  9:00   ` Chris Mason
2026-10-09 14:17 ` Fuad Tabba
2026-10-09 17:57   ` Chris Mason

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7ia4y0cbsvdn.fsf@castle.c.googlers.com \
    --to=roman.gushchin@linux.dev \
    --cc=ast@kernel.org \
    --cc=ihor.solodrai@linux.dev \
    --cc=kuba@kernel.org \
    --cc=mason@kernel.org \
    --cc=sashiko@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.