* [RFC] reworking the review-prompts subsystem guide
@ 2026-10-02 19:04 Chris Mason
2026-10-02 21:18 ` Chuck Lever
` (3 more replies)
0 siblings, 4 replies; 11+ messages in thread
From: Chris Mason @ 2026-10-02 19:04 UTC (permalink / raw)
To: sashiko, Roman Gushchin, ihor.solodrai, ast, kuba
Hi everyone,
I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
https://github.com/masoncl/review-prompts.git subsystem-build
This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion.
The review prompts started with a lot of framework to explain how the kernel works, and the prompts have focused on documenting subsystems in a way that both LLMs and people can use for reviews and writing code.
As models have progressed, they actually need different information than they used to. Recent models already know how most of the kernel works, but they have some gaps because the kernel is always changing, or because they have some bad assumptions baked in.
My new branch builds subsystem guides by asking a long list of questions, and then extensively reading the sources to find the right answers. The delta between the LLM's answers and the right answers is the new guide.
This is both much less useful to human readers and much longer. I'm not sure what to say about the human reader part, but instead of having LLMs read the whole subsystem guide, I shifted to an index where they search for symbols. This is a better fit for more advanced models, which mostly need updates on how the kernel has changed since they were trained.
The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. It's the union of all the things either model got wrong, and the tree is setup so we can add builds for other kernel versions (ex: stable kernel series).
I started down this path convinced that I needed to get the subsystem guides into the kernel git tree. My idea was this was documentation, and it should land somewhere in the kernel for that reason. I ended up with something that probably shouldn't be in the tree...different models need different things on top of different kernels, and so I'm trying out the build idea.
What I know for sure is the existing review prompts have drifted from mainline Linus. It's impacting the quality of the reviews, so I plan on working out something in the near future.
-chris
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
@ 2026-10-02 21:18 ` Chuck Lever
2026-10-04 9:17 ` Chris Mason
2026-10-05 19:58 ` Jakub Kicinski
` (2 subsequent siblings)
3 siblings, 1 reply; 11+ messages in thread
From: Chuck Lever @ 2026-10-02 21:18 UTC (permalink / raw)
To: Chris Mason, sashiko, Roman Gushchin, ihor.solodrai, ast,
Jakub Kicinski
On Fri, Oct 2, 2026, at 12:04 PM, Chris Mason wrote:
> The subsystem-build branch was built with both sonnet-5.5 and opus-5.5.
> It's the union of all the things either model got wrong, and the tree
> is setup so we can add builds for other kernel versions (ex: stable
> kernel series).
>
> I started down this path convinced that I needed to get the subsystem
> guides into the kernel git tree. My idea was this was documentation,
> and it should land somewhere in the kernel for that reason. I ended up
> with something that probably shouldn't be in the tree...different
> models need different things on top of different kernels, and so I'm
> trying out the build idea.
One thing that has been wiggling around in the back of my brain is
the review findings delta between the different models. Different
issues are found, different issues are skipped, different false
positives.
I'm sure some of that is unavoidable, but even in the simpler cases,
some models are simply blind to certain types of correctness issues.
Is there any way to compare the benchmark results between the models
(say, between gemini-3.1-pro, gpt-5.6-luna, and claude-sonnet-5) and
prompt one of the more advanced models (eg, gpt-6-astra) to figure out
why the review prompts should have guided these models to disparate
review findings? The goal being to achieve better consistency between
sashiko results for each model.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-02 21:18 ` Chuck Lever
@ 2026-10-04 9:17 ` Chris Mason
0 siblings, 0 replies; 11+ messages in thread
From: Chris Mason @ 2026-10-04 9:17 UTC (permalink / raw)
To: Chuck Lever, sashiko, Roman Gushchin, ihor.solodrai, ast,
Jakub Kicinski
On Fri, Oct 2, 2026, at 11:18 PM, Chuck Lever wrote:
> On Fri, Oct 2, 2026, at 12:04 PM, Chris Mason wrote:
>> The subsystem-build branch was built with both sonnet-5.5 and opus-5.5.
>> It's the union of all the things either model got wrong, and the tree
>> is setup so we can add builds for other kernel versions (ex: stable
>> kernel series).
>>
>> I started down this path convinced that I needed to get the subsystem
>> guides into the kernel git tree. My idea was this was documentation,
>> and it should land somewhere in the kernel for that reason. I ended up
>> with something that probably shouldn't be in the tree...different
>> models need different things on top of different kernels, and so I'm
>> trying out the build idea.
>
> One thing that has been wiggling around in the back of my brain is
> the review findings delta between the different models. Different
> issues are found, different issues are skipped, different false
> positives.
>
> I'm sure some of that is unavoidable, but even in the simpler cases,
> some models are simply blind to certain types of correctness issues.
>
> Is there any way to compare the benchmark results between the models
> (say, between gemini-3.1-pro, gpt-5.6-luna, and claude-sonnet-5) and
> prompt one of the more advanced models (eg, gpt-6-astra) to figure out
> why the review prompts should have guided these models to disparate
> review findings? The goal being to achieve better consistency between
> sashiko results for each model.
The short answer is yes, but it's kind of tricky to build it with existing kernel bugs. You have to pick recent bugs that no model has ever seen, or invent entirely new bugs.
When the reason for the failure is a lack of context, you can absolutely build better systems to pull in the right context. But when the model just can't wrap its head around finding a particular kind of bug, it's better to wait for the next model. You can see it with the early results from the review prompts, which were mostly error handling failures with smaller surprises mixed in. Newer models give much better results.
I'd love to try and build a benchmark that shows this clearly, but I don't think we can guide the earlier models into understanding bugs beyond their reasoning skills. Or at least I consistently fail to do so, I'd love to be wrong.
-chris
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
2026-10-02 21:18 ` Chuck Lever
@ 2026-10-05 19:58 ` Jakub Kicinski
2026-10-05 21:14 ` Roman Gushchin
2026-10-06 21:22 ` Ihor Solodrai
2026-10-09 14:17 ` Fuad Tabba
3 siblings, 1 reply; 11+ messages in thread
From: Jakub Kicinski @ 2026-10-05 19:58 UTC (permalink / raw)
To: Chris Mason; +Cc: sashiko, Roman Gushchin, ihor.solodrai, ast
On Fri, 02 Oct 2026 15:04:00 -0400 Chris Mason wrote:
> Hi everyone,
>
> I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
>
> https://github.com/masoncl/review-prompts.git subsystem-build
>
> This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion.
>
> The review prompts started with a lot of framework to explain how the kernel works, and the prompts have focused on documenting subsystems in a way that both LLMs and people can use for reviews and writing code.
>
> As models have progressed, they actually need different information than they used to. Recent models already know how most of the kernel works, but they have some gaps because the kernel is always changing, or because they have some bad assumptions baked in.
>
> My new branch builds subsystem guides by asking a long list of questions, and then extensively reading the sources to find the right answers. The delta between the LLM's answers and the right answers is the new guide.
>
> This is both much less useful to human readers and much longer. I'm not sure what to say about the human reader part, but instead of having LLMs read the whole subsystem guide, I shifted to an index where they search for symbols. This is a better fit for more advanced models, which mostly need updates on how the kernel has changed since they were trained.
>
> The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. It's the union of all the things either model got wrong, and the tree is setup so we can add builds for other kernel versions (ex: stable kernel series).
>
> I started down this path convinced that I needed to get the subsystem guides into the kernel git tree. My idea was this was documentation, and it should land somewhere in the kernel for that reason. I ended up with something that probably shouldn't be in the tree...different models need different things on top of different kernels, and so I'm trying out the build idea.
>
> What I know for sure is the existing review prompts have drifted from mainline Linus. It's impacting the quality of the reviews, so I plan on working out something in the near future.
I didn't read the whole thing, but makes sense AFAIU.
Closing the loop with Sashiko and/or ML would be great :(
If the bots read the ML they could both learn false positives and false
negatives automatically. I suspect most of us thought about this by now.
If we're doing a redesign should the ingest of ML be part of it?
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-05 19:58 ` Jakub Kicinski
@ 2026-10-05 21:14 ` Roman Gushchin
2026-10-05 21:53 ` Jakub Kicinski
2026-10-05 23:37 ` Ihor Solodrai
0 siblings, 2 replies; 11+ messages in thread
From: Roman Gushchin @ 2026-10-05 21:14 UTC (permalink / raw)
To: Jakub Kicinski; +Cc: Chris Mason, sashiko, ihor.solodrai, ast
Jakub Kicinski <kuba@kernel.org> writes:
> On Fri, 02 Oct 2026 15:04:00 -0400 Chris Mason wrote:
>> Hi everyone,
>>
>> I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
>>
>> https://github.com/masoncl/review-prompts.git subsystem-build
>>
>> This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion.
>>
>> The review prompts started with a lot of framework to explain how
>> the kernel works, and the prompts have focused on documenting
>> subsystems in a way that both LLMs and people can use for reviews
>> and writing code.
>>
>> As models have progressed, they actually need different information
>> than they used to. Recent models already know how most of the
>> kernel works, but they have some gaps because the kernel is always
>> changing, or because they have some bad assumptions baked in.
>>
>> My new branch builds subsystem guides by asking a long list of
>> questions, and then extensively reading the sources to find the
>> right answers. The delta between the LLM's answers and the right
>> answers is the new guide.
>>
>> This is both much less useful to human readers and much longer. I'm
>> not sure what to say about the human reader part, but instead of
>> having LLMs read the whole subsystem guide, I shifted to an index
>> where they search for symbols. This is a better fit for more
>> advanced models, which mostly need updates on how the kernel has
>> changed since they were trained.
>>
>> The subsystem-build branch was built with both sonnet-5.5 and
>> opus-5.5. It's the union of all the things either model got wrong,
>> and the tree is setup so we can add builds for other kernel versions
>> (ex: stable kernel series).
>>
>> I started down this path convinced that I needed to get the
>> subsystem guides into the kernel git tree. My idea was this was
>> documentation, and it should land somewhere in the kernel for that
>> reason. I ended up with something that probably shouldn't be in the
>> tree...different models need different things on top of different
>> kernels, and so I'm trying out the build idea.
>>
>> What I know for sure is the existing review prompts have drifted
>> from mainline Linus. It's impacting the quality of the reviews, so
>> I plan on working out something in the near future.
>
> I didn't read the whole thing, but makes sense AFAIU.
>
> Closing the loop with Sashiko and/or ML would be great :(
> If the bots read the ML they could both learn false positives and false
> negatives automatically. I suspect most of us thought about this by now.
> If we're doing a redesign should the ingest of ML be part of it?
I plan to do this (and had a prototype in the past), but it's tricky if
we take security seriously. And we absolutely should!
Obviously just giving an agent an access to lore archive is opening a
can of worms in terms of possible prompt injections. So I think we
should do it really carefully. And outside of security considerations,
humans are simple wrong too. So my plan with sashiko is to force it to
try to verify the feedback against the codebase and if it's not possible
trust only maintainers or people with a long history of meaningful
contributions.
And sashiko can (and does in my prototype) generate prompts based on the
feedback and automatically verify that it helps by re-reviewing original
patches. This should mostly eliminate a need for human-generated
prompts.
This is all doable, but probably will take few more week to build and
roll out.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-05 21:14 ` Roman Gushchin
@ 2026-10-05 21:53 ` Jakub Kicinski
2026-10-05 23:37 ` Ihor Solodrai
1 sibling, 0 replies; 11+ messages in thread
From: Jakub Kicinski @ 2026-10-05 21:53 UTC (permalink / raw)
To: Roman Gushchin; +Cc: Chris Mason, sashiko, ihor.solodrai, ast
On Mon, 05 Oct 2026 21:14:12 +0000 Roman Gushchin wrote:
> I plan to do this (and had a prototype in the past), but it's tricky if
> we take security seriously. And we absolutely should!
> Obviously just giving an agent an access to lore archive is opening a
> can of worms in terms of possible prompt injections. So I think we
> should do it really carefully.
Hm, is there really much extra attack surface here beyond what people
can already put in a commit message of a fake patch?
> And outside of security considerations,
> humans are simple wrong too. So my plan with sashiko is to force it to
> try to verify the feedback against the codebase and if it's not possible
> trust only maintainers or people with a long history of meaningful
> contributions.
Definitely, I was also wondering about doing some rough estimate
of "trustworthiness" of the person based on how long they have been
contributing + MAINTAINERS status. In practice, tho, noobs rarely
provide feedback. Closing the loop with patchwork is a better idea
(if patch go rejected == maintainer must have judged some feedback
as important)
> And sashiko can (and does in my prototype) generate prompts based on the
> feedback and automatically verify that it helps by re-reviewing original
> patches. This should mostly eliminate a need for human-generated
> prompts.
> This is all doable, but probably will take few more week to build and
> roll out.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-05 21:14 ` Roman Gushchin
2026-10-05 21:53 ` Jakub Kicinski
@ 2026-10-05 23:37 ` Ihor Solodrai
1 sibling, 0 replies; 11+ messages in thread
From: Ihor Solodrai @ 2026-10-05 23:37 UTC (permalink / raw)
To: Roman Gushchin, Jakub Kicinski; +Cc: Chris Mason, sashiko, ast
On 10/5/26 2:14 PM, Roman Gushchin wrote:
> Jakub Kicinski <kuba@kernel.org> writes:
>
>> [...]
>>
>> Closing the loop with Sashiko and/or ML would be great :(
>> If the bots read the ML they could both learn false positives and false
>> negatives automatically. I suspect most of us thought about this by now.
>> If we're doing a redesign should the ingest of ML be part of it?
FWIW the BPF bot has been reading lore when reviewing patches through
semcode since February:
https://github.com/kernel-patches/vmtest/pull/442
Basically, the container in which the AI session is running has access
to local semcode via MCP, and local semcode db has the full (bpf) lore
archive.
It definitely helps with quality to give AI more context, but it also
has side effects. For example, a bot may repeat an issue raised by
another bot in a previous revision. In many cases this is noise, but
sometimes it is appropriate: the review may have been inappropriately
ignored. Currently this is mostly hidden by rules in the prompts to
not repeat issues that have already been discussed, because it appears
that people ignore many of the AI reviews anyways.
>
> I plan to do this (and had a prototype in the past), but it's tricky if
> we take security seriously. And we absolutely should!
> Obviously just giving an agent an access to lore archive is opening a
> can of worms in terms of possible prompt injections. So I think we
> should do it really carefully. [...]
Re prompt injections, this is of course always a concern, and my yolo
attitude may not appeal to everyone. But I have to say since almost a
year of running AI reviews on BPF this hasn't been a problem once.
This of course depends on how the infrastructure is set up. BPF bot is
quite well isolated:
* dedicated AWS account
* each job runs on an ephemeral runner: not only new container, but
also ephemeral hosts (we use CodeBuild)
* the API access to inference is limited to 1h in a job
* the AWS and github access tokens have limited permissions
With all that, I sleep well.
One can imagine a sci-fi scenario with a prompt injection tiggering a
bot to take over the AWS account, but given corporate monitoring of
activity there, we'd detect it and turn everything off immediately.
I also think that as the end-users of the LLM APIs we automatically
benefit from all the protections on the provider side. Sometimes this
is even annoying, because some models refuse to investigate certain
issues due to the guardrails on the backend.
> And outside of security considerations,
> humans are simple wrong too. So my plan with sashiko is to force it to
> try to verify the feedback against the codebase and if it's not possible
> trust only maintainers or people with a long history of meaningful
> contributions.
I have small local datasets generated from mailing lists, and one way
I tried to extract signal is to check whether the feedback from the
bot was incorporated in the next revision of the patches. Some people
do that silently without crediting the bots. So yeah, code is the
authority. Reputation plays some role of course, but in the end the
bots should check the claims against the code and discussion history.
> And sashiko can (and does in my prototype) generate prompts based on the
> feedback and automatically verify that it helps by re-reviewing original
> patches. This should mostly eliminate a need for human-generated
> prompts.
I did this manually a few times over the year: analyse lore -> come up
with prompt improvements -> test them -> merge.
This is definitely automatable and can close the loop to a large
degree. Prompting and context management is not that far from actual
machine learning, as it turns out :)
> This is all doable, but probably will take few more week to build and
> roll out.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
2026-10-02 21:18 ` Chuck Lever
2026-10-05 19:58 ` Jakub Kicinski
@ 2026-10-06 21:22 ` Ihor Solodrai
2026-10-07 9:00 ` Chris Mason
2026-10-09 14:17 ` Fuad Tabba
3 siblings, 1 reply; 11+ messages in thread
From: Ihor Solodrai @ 2026-10-06 21:22 UTC (permalink / raw)
To: Chris Mason, sashiko, Roman Gushchin, ast, kuba
On 10/2/26 12:04 PM, Chris Mason wrote:
> Hi everyone,
>
> I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
>
> https://github.com/masoncl/review-prompts.git subsystem-build
Hi Chris,
I've set up "shadow" AI reviews for BPF, taking advantage of our
existing bpf-rc infra. It's a copy of the entire BPF CI pipeline that
we occasionally use for testing of the CI itself.
The shadow reviews are using your new prompts from the subsystem-build
branch. The reviews are posted on kernel-patches/bpf-rc PRs as
comments and are not mailed to the list. I configured the reviews to
run on every 4th series (so 25% of prod). I'll let it run for some
time (couple of weeks?) and then we'll be able to analyze the
difference and either switch or iterate.
One quirk I added: a datetime cutoff of the semcode's lore archive, so
that the rc reviewer doesn't see prod reviewer's comments. Not sure
how reliable that'll be, because AI may potentially figure out to read
them from github or something. But we should notice if it doesn't
work.
You can find the relevant PRs like this:
$ gh pr list -R kernel-patches/bpf-rc --label ai-review-shadow:subsystem-build --state all --limit 50 \
--json url,headRefName,title --jq '.[] | "\(.url)\t\(.headRefName)\t\(.title)"'
Note the "ai-review-shadow:subsystem-build" PR label.
The reviews run with debug output, and also the claude session logs
are captured as artifacts. Ask your agent to write a script to pull
those in with github cli, should be easy.
I just pushed, it'll take a few days before we get any real samples.
Thank you for working on the prompts!
>
> [...]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-06 21:22 ` Ihor Solodrai
@ 2026-10-07 9:00 ` Chris Mason
0 siblings, 0 replies; 11+ messages in thread
From: Chris Mason @ 2026-10-07 9:00 UTC (permalink / raw)
To: Ihor Solodrai, sashiko, Roman Gushchin, ast, Jakub Kicinski
On Tue, Oct 6, 2026, at 11:22 PM, Ihor Solodrai wrote:
> On 10/2/26 12:04 PM, Chris Mason wrote:
>> Hi everyone,
>>
>> I've been updating the review prompts, and you can find my current
>> work here in the subsystem-build branch:
>>
>> https://github.com/masoncl/review-prompts.git subsystem-build
>
> Hi Chris,
>
> I've set up "shadow" AI reviews for BPF, taking advantage of our
> existing bpf-rc infra. It's a copy of the entire BPF CI pipeline that
> we occasionally use for testing of the CI itself.
>
> The shadow reviews are using your new prompts from the subsystem-build
> branch. The reviews are posted on kernel-patches/bpf-rc PRs as
> comments and are not mailed to the list. I configured the reviews to
> run on every 4th series (so 25% of prod). I'll let it run for some
> time (couple of weeks?) and then we'll be able to analyze the
> difference and either switch or iterate.
>
> One quirk I added: a datetime cutoff of the semcode's lore archive,
> so that the rc reviewer doesn't see prod reviewer's comments. Not
> sure how reliable that'll be, because AI may potentially figure out
> to read them from github or something. But we should notice if it
> doesn't work.
>
> You can find the relevant PRs like this:
>
> $ gh pr list -R kernel-patches/bpf-rc --label ai-review-shadow:subsystem-
> build --state all --limit 50 \ --json url,headRefName,title --jq
> '.[] | "\(.url)\t\(.headRefName)\t\(.title)"'
>
> Note the "ai-review-shadow:subsystem-build" PR label.
>
> The reviews run with debug output, and also the claude session logs
> are captured as artifacts. Ask your agent to write a script to pull
> those in with github cli, should be easy.
>
> I just pushed, it'll take a few days before we get any real samples.
Oh this is great, thanks so much.
-chris
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
` (2 preceding siblings ...)
2026-10-06 21:22 ` Ihor Solodrai
@ 2026-10-09 14:17 ` Fuad Tabba
2026-10-09 17:57 ` Chris Mason
3 siblings, 1 reply; 11+ messages in thread
From: Fuad Tabba @ 2026-10-09 14:17 UTC (permalink / raw)
To: Chris Mason; +Cc: sashiko, Roman Gushchin, ihor.solodrai, ast, kuba, Fuad Tabba
Hi Chris,
On Fri, 02 Oct 2026 20:04:00 +0100, "Chris Mason" <mason@kernel.org> wrote:
[...]
> My new branch builds subsystem guides by asking a long list of
> questions, and then extensively reading the sources to find the
> right answers. The delta between the LLM's answers and the right
> answers is the new guide.
I went through the KVM/arm64 and protected KVM (pKVM) guides on the
branch, and checked a sample of their claims against the tree. They're
much more accurate and detailed than the hand-written ones, and the
checking against the source is what makes the difference.
Some of what a reviewer needs isn't in the code, though: which bugs
matter and how much, and which reports are false alarms. The old pKVM
guide said that a hypervisor crash only the host kernel can trigger is a
hardening item, while one a guest can trigger is a real bug. The design
deliberately keeps that kind of instruction out of the guides, so it's
gone from the built one. Where should it live instead? The "Conventions
for new code" files look closest: they already hold what maintainers ask
for and no code states, and a review finds them by the directory a patch
touches. Something like that for KVM/arm64 would work, and I can write
it.
Leaving out what the tested models already knew also tunes the guide to
those models, and the model doing the review may not be one of them.
Hand-written guides have the same problem, and I don't see an easy fix.
Your mail points at builds per model, so one option is for each project
to build for the model it runs. Another is for the build to keep the
full checked answer rather than only the difference, now that the index
makes length matter less. Which way are you leaning?
> This is both much less useful to human readers and much longer. I'm
> not sure what to say about the human reader part, but instead of
> having LLMs read the whole subsystem guide, I shifted to an index
> where they search for symbols. This is a better fit for more
> advanced models, which mostly need updates on how the kernel has
> changed since they were trained.
A common KVM patch adds a new hypercall. A symbol search finds the
answers about the existing calls the patch uses, but not the rule for
where a new one goes in the list: that answer is filed under a marker
the diff never touches, and the header the diff edits isn't the source
file of any index line. Could some answers be keyed to the directory as
well, like the short table you kept for a few guides?
[...]
> What I know for sure is the existing review prompts have drifted
> from mainline Linus. It's impacting the quality of the reviews, so
> I plan on working out something in the near future.
Built guides drift too, just in a different way. These were built at
7.3-rc5, and one fix in rc6 made four of the KVM/arm64 claims wrong
without renaming anything, so looking names up in the tree doesn't catch
it. Most index lines already record a source file, so a review on a
newer tree could check whether that file changed since the build, and
re-check or flag the answer if it did.
Related: when a maintainer finds a wrong answer, the only fix is to
change the question and rebuild, which needs model access and can come
out differently each time. Who looks after the question files, and could
there be a small override that a maintainer edits between rebuilds?
Feedback from the list, as Jakub and Roman discussed, could take the
same route: suggested questions that go through the same checks against
the tree, rather than guide text.
Separately, some pKVM questions were dropped for size, including the one
on protected guest system registers, and the measurement notes say
they're the first to bring back if the size limit is raised. There's no
length limit now, so I could send a patch to bring them back.
Cheers,
/fuad
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide
2026-10-09 14:17 ` Fuad Tabba
@ 2026-10-09 17:57 ` Chris Mason
0 siblings, 0 replies; 11+ messages in thread
From: Chris Mason @ 2026-10-09 17:57 UTC (permalink / raw)
To: Fuad Tabba
Cc: sashiko, Roman Gushchin, Ihor Solodrai, ast, Jakub Kicinski,
Fuad Tabba
On Fri, Oct 9, 2026, at 10:17 AM, Fuad Tabba wrote:
> Hi Chris,
>
> On Fri, 02 Oct 2026 20:04:00 +0100, "Chris Mason" <mason@kernel.org> wrote:
> [...]
>> My new branch builds subsystem guides by asking a long list of
>> questions, and then extensively reading the sources to find the
>> right answers. The delta between the LLM's answers and the right
>> answers is the new guide.
>
> I went through the KVM/arm64 and protected KVM (pKVM) guides on the
> branch, and checked a sample of their claims against the tree. They're
> much more accurate and detailed than the hand-written ones, and the
> checking against the source is what makes the difference.
Great, I only did limited verification in that area, so its good to hear.
>
> Some of what a reviewer needs isn't in the code, though: which bugs
> matter and how much, and which reports are false alarms. The old pKVM
> guide said that a hypervisor crash only the host kernel can trigger is a
> hardening item, while one a guest can trigger is a real bug. The design
> deliberately keeps that kind of instruction out of the guides, so it's
> gone from the built one. Where should it live instead? The "Conventions
> for new code" files look closest: they already hold what maintainers ask
> for and no code states, and a review finds them by the directory a patch
> touches. Something like that for KVM/arm64 would work, and I can write
> it.
There's already a way to pass specific parts verbatim, which feels more
correct than under a "new code" label. We can play around with a few
ideas though.
>
> Leaving out what the tested models already knew also tunes the guide to
> those models, and the model doing the review may not be one of them.
Yes. I generated the current built prompts using both sonnet and opus
with the assumption that other models would get roughly the same
things wrong as one of the two of them. This is obviously flawed, but
the prompts have a way to pick the best build directory based on kernel
version, so we could extend that to the model doing the
review (more below).
> Hand-written guides have the same problem, and I don't see an easy fix.
> Your mail points at builds per model, so one option is for each project
> to build for the model it runs. Another is for the build to keep the
> full checked answer rather than only the difference, now that the index
> makes length matter less. Which way are you leaning?
My original plan was to get the built guides small enough to be included
whole, in the same way the original guides were. Adding the keyword
index was basically me admitting defeat, and it just didn't occur to me
that we might be able to index our way to victory on the full output.
IOW, great idea, I didn't think of that at all ;)
I'll do another run and put the full build up. We can see if it's too big
to be useful or if we can make the index strong enough to get around
needing the per-model analysis.
>
>> This is both much less useful to human readers and much longer. I'm
>> not sure what to say about the human reader part, but instead of
>> having LLMs read the whole subsystem guide, I shifted to an index
>> where they search for symbols. This is a better fit for more
>> advanced models, which mostly need updates on how the kernel has
>> changed since they were trained.
>
> A common KVM patch adds a new hypercall. A symbol search finds the
> answers about the existing calls the patch uses, but not the rule for
> where a new one goes in the list: that answer is filed under a marker
> the diff never touches, and the header the diff edits isn't the source
> file of any index line. Could some answers be keyed to the directory as
> well, like the short table you kept for a few guides?
>
Absolutely. The index is pretty dumb, we can do a lot better.
> [...]
>> What I know for sure is the existing review prompts have drifted
>> from mainline Linus. It's impacting the quality of the reviews, so
>> I plan on working out something in the near future.
>
> Built guides drift too, just in a different way. These were built at
> 7.3-rc5, and one fix in rc6 made four of the KVM/arm64 claims wrong
> without renaming anything, so looking names up in the tree doesn't catch
> it. Most index lines already record a source file, so a review on a
> newer tree could check whether that file changed since the build, and
> re-check or flag the answer if it did.
A related question is how often do I need to rebuild in order for
the guides to be useful? I'd assume the absolute minimum is every
final release, but every RC is also reasonable.
>
> Related: when a maintainer finds a wrong answer, the only fix is to
> change the question and rebuild, which needs model access and can come
> out differently each time. Who looks after the question files, and could
> there be a small override that a maintainer edits between rebuilds?
> Feedback from the list, as Jakub and Roman discussed, could take the
> same route: suggested questions that go through the same checks against
> the tree, rather than guide text.
The build script can rebuild a single guide, or people can just have
their agents hand edit the build? It's a good point, we should have
AGENTS.md record some best practices.
>
> Separately, some pKVM questions were dropped for size, including the one
> on protected guest system registers, and the measurement notes say
> they're the first to bring back if the size limit is raised. There's no
> length limit now, so I could send a patch to bring them back.
Great, please do.
-chris
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-10-09 17:58 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason
2026-10-02 21:18 ` Chuck Lever
2026-10-04 9:17 ` Chris Mason
2026-10-05 19:58 ` Jakub Kicinski
2026-10-05 21:14 ` Roman Gushchin
2026-10-05 21:53 ` Jakub Kicinski
2026-10-05 23:37 ` Ihor Solodrai
2026-10-06 21:22 ` Ihor Solodrai
2026-10-07 9:00 ` Chris Mason
2026-10-09 14:17 ` Fuad Tabba
2026-10-09 17:57 ` Chris Mason
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.