* [RFC] reworking the review-prompts subsystem guide
@ 2026-10-02 19:04 Chris Mason
2026-10-02 21:18 ` Chuck Lever
` (3 more replies)
0 siblings, 4 replies; 12+ messages in thread
From: Chris Mason @ 2026-10-02 19:04 UTC (permalink / raw)
To: sashiko, Roman Gushchin, ihor.solodrai, ast, kuba
Hi everyone,
I've been updating the review prompts, and you can find my current work here in the subsystem-build branch:
https://github.com/masoncl/review-prompts.git subsystem-build
This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion.
The review prompts started with a lot of framework to explain how the kernel works, and the prompts have focused on documenting subsystems in a way that both LLMs and people can use for reviews and writing code.
As models have progressed, they actually need different information than they used to. Recent models already know how most of the kernel works, but they have some gaps because the kernel is always changing, or because they have some bad assumptions baked in.
My new branch builds subsystem guides by asking a long list of questions, and then extensively reading the sources to find the right answers. The delta between the LLM's answers and the right answers is the new guide.
This is both much less useful to human readers and much longer. I'm not sure what to say about the human reader part, but instead of having LLMs read the whole subsystem guide, I shifted to an index where they search for symbols. This is a better fit for more advanced models, which mostly need updates on how the kernel has changed since they were trained.
The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. It's the union of all the things either model got wrong, and the tree is setup so we can add builds for other kernel versions (ex: stable kernel series).
I started down this path convinced that I needed to get the subsystem guides into the kernel git tree. My idea was this was documentation, and it should land somewhere in the kernel for that reason. I ended up with something that probably shouldn't be in the tree...different models need different things on top of different kernels, and so I'm trying out the build idea.
What I know for sure is the existing review prompts have drifted from mainline Linus. It's impacting the quality of the reviews, so I plan on working out something in the near future.
-chris
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason @ 2026-10-02 21:18 ` Chuck Lever 2026-10-04 9:17 ` Chris Mason 2026-10-05 19:58 ` Jakub Kicinski ` (2 subsequent siblings) 3 siblings, 1 reply; 12+ messages in thread From: Chuck Lever @ 2026-10-02 21:18 UTC (permalink / raw) To: Chris Mason, sashiko, Roman Gushchin, ihor.solodrai, ast, Jakub Kicinski On Fri, Oct 2, 2026, at 12:04 PM, Chris Mason wrote: > The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. > It's the union of all the things either model got wrong, and the tree > is setup so we can add builds for other kernel versions (ex: stable > kernel series). > > I started down this path convinced that I needed to get the subsystem > guides into the kernel git tree. My idea was this was documentation, > and it should land somewhere in the kernel for that reason. I ended up > with something that probably shouldn't be in the tree...different > models need different things on top of different kernels, and so I'm > trying out the build idea. One thing that has been wiggling around in the back of my brain is the review findings delta between the different models. Different issues are found, different issues are skipped, different false positives. I'm sure some of that is unavoidable, but even in the simpler cases, some models are simply blind to certain types of correctness issues. Is there any way to compare the benchmark results between the models (say, between gemini-3.1-pro, gpt-5.6-luna, and claude-sonnet-5) and prompt one of the more advanced models (eg, gpt-6-astra) to figure out why the review prompts should have guided these models to disparate review findings? The goal being to achieve better consistency between sashiko results for each model. -- Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org) ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-02 21:18 ` Chuck Lever @ 2026-10-04 9:17 ` Chris Mason 0 siblings, 0 replies; 12+ messages in thread From: Chris Mason @ 2026-10-04 9:17 UTC (permalink / raw) To: Chuck Lever, sashiko, Roman Gushchin, ihor.solodrai, ast, Jakub Kicinski On Fri, Oct 2, 2026, at 11:18 PM, Chuck Lever wrote: > On Fri, Oct 2, 2026, at 12:04 PM, Chris Mason wrote: >> The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. >> It's the union of all the things either model got wrong, and the tree >> is setup so we can add builds for other kernel versions (ex: stable >> kernel series). >> >> I started down this path convinced that I needed to get the subsystem >> guides into the kernel git tree. My idea was this was documentation, >> and it should land somewhere in the kernel for that reason. I ended up >> with something that probably shouldn't be in the tree...different >> models need different things on top of different kernels, and so I'm >> trying out the build idea. > > One thing that has been wiggling around in the back of my brain is > the review findings delta between the different models. Different > issues are found, different issues are skipped, different false > positives. > > I'm sure some of that is unavoidable, but even in the simpler cases, > some models are simply blind to certain types of correctness issues. > > Is there any way to compare the benchmark results between the models > (say, between gemini-3.1-pro, gpt-5.6-luna, and claude-sonnet-5) and > prompt one of the more advanced models (eg, gpt-6-astra) to figure out > why the review prompts should have guided these models to disparate > review findings? The goal being to achieve better consistency between > sashiko results for each model. The short answer is yes, but it's kind of tricky to build it with existing kernel bugs. You have to pick recent bugs that no model has ever seen, or invent entirely new bugs. When the reason for the failure is a lack of context, you can absolutely build better systems to pull in the right context. But when the model just can't wrap its head around finding a particular kind of bug, it's better to wait for the next model. You can see it with the early results from the review prompts, which were mostly error handling failures with smaller surprises mixed in. Newer models give much better results. I'd love to try and build a benchmark that shows this clearly, but I don't think we can guide the earlier models into understanding bugs beyond their reasoning skills. Or at least I consistently fail to do so, I'd love to be wrong. -chris ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason 2026-10-02 21:18 ` Chuck Lever @ 2026-10-05 19:58 ` Jakub Kicinski 2026-10-05 21:14 ` Roman Gushchin 2026-10-06 21:22 ` Ihor Solodrai 2026-10-09 14:17 ` Fuad Tabba 3 siblings, 1 reply; 12+ messages in thread From: Jakub Kicinski @ 2026-10-05 19:58 UTC (permalink / raw) To: Chris Mason; +Cc: sashiko, Roman Gushchin, ihor.solodrai, ast On Fri, 02 Oct 2026 15:04:00 -0400 Chris Mason wrote: > Hi everyone, > > I've been updating the review prompts, and you can find my current work here in the subsystem-build branch: > > https://github.com/masoncl/review-prompts.git subsystem-build > > This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion. > > The review prompts started with a lot of framework to explain how the kernel works, and the prompts have focused on documenting subsystems in a way that both LLMs and people can use for reviews and writing code. > > As models have progressed, they actually need different information than they used to. Recent models already know how most of the kernel works, but they have some gaps because the kernel is always changing, or because they have some bad assumptions baked in. > > My new branch builds subsystem guides by asking a long list of questions, and then extensively reading the sources to find the right answers. The delta between the LLM's answers and the right answers is the new guide. > > This is both much less useful to human readers and much longer. I'm not sure what to say about the human reader part, but instead of having LLMs read the whole subsystem guide, I shifted to an index where they search for symbols. This is a better fit for more advanced models, which mostly need updates on how the kernel has changed since they were trained. > > The subsystem-build branch was built with both sonnet-5.5 and opus-5.5. It's the union of all the things either model got wrong, and the tree is setup so we can add builds for other kernel versions (ex: stable kernel series). > > I started down this path convinced that I needed to get the subsystem guides into the kernel git tree. My idea was this was documentation, and it should land somewhere in the kernel for that reason. I ended up with something that probably shouldn't be in the tree...different models need different things on top of different kernels, and so I'm trying out the build idea. > > What I know for sure is the existing review prompts have drifted from mainline Linus. It's impacting the quality of the reviews, so I plan on working out something in the near future. I didn't read the whole thing, but makes sense AFAIU. Closing the loop with Sashiko and/or ML would be great :( If the bots read the ML they could both learn false positives and false negatives automatically. I suspect most of us thought about this by now. If we're doing a redesign should the ingest of ML be part of it? ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-05 19:58 ` Jakub Kicinski @ 2026-10-05 21:14 ` Roman Gushchin 2026-10-05 21:53 ` Jakub Kicinski 2026-10-05 23:37 ` Ihor Solodrai 0 siblings, 2 replies; 12+ messages in thread From: Roman Gushchin @ 2026-10-05 21:14 UTC (permalink / raw) To: Jakub Kicinski; +Cc: Chris Mason, sashiko, ihor.solodrai, ast Jakub Kicinski <kuba@kernel.org> writes: > On Fri, 02 Oct 2026 15:04:00 -0400 Chris Mason wrote: >> Hi everyone, >> >> I've been updating the review prompts, and you can find my current work here in the subsystem-build branch: >> >> https://github.com/masoncl/review-prompts.git subsystem-build >> >> This is a pretty big change, and since a few projects are syncing automatically, I didn't want to just throw it in without discussion. >> >> The review prompts started with a lot of framework to explain how >> the kernel works, and the prompts have focused on documenting >> subsystems in a way that both LLMs and people can use for reviews >> and writing code. >> >> As models have progressed, they actually need different information >> than they used to. Recent models already know how most of the >> kernel works, but they have some gaps because the kernel is always >> changing, or because they have some bad assumptions baked in. >> >> My new branch builds subsystem guides by asking a long list of >> questions, and then extensively reading the sources to find the >> right answers. The delta between the LLM's answers and the right >> answers is the new guide. >> >> This is both much less useful to human readers and much longer. I'm >> not sure what to say about the human reader part, but instead of >> having LLMs read the whole subsystem guide, I shifted to an index >> where they search for symbols. This is a better fit for more >> advanced models, which mostly need updates on how the kernel has >> changed since they were trained. >> >> The subsystem-build branch was built with both sonnet-5.5 and >> opus-5.5. It's the union of all the things either model got wrong, >> and the tree is setup so we can add builds for other kernel versions >> (ex: stable kernel series). >> >> I started down this path convinced that I needed to get the >> subsystem guides into the kernel git tree. My idea was this was >> documentation, and it should land somewhere in the kernel for that >> reason. I ended up with something that probably shouldn't be in the >> tree...different models need different things on top of different >> kernels, and so I'm trying out the build idea. >> >> What I know for sure is the existing review prompts have drifted >> from mainline Linus. It's impacting the quality of the reviews, so >> I plan on working out something in the near future. > > I didn't read the whole thing, but makes sense AFAIU. > > Closing the loop with Sashiko and/or ML would be great :( > If the bots read the ML they could both learn false positives and false > negatives automatically. I suspect most of us thought about this by now. > If we're doing a redesign should the ingest of ML be part of it? I plan to do this (and had a prototype in the past), but it's tricky if we take security seriously. And we absolutely should! Obviously just giving an agent an access to lore archive is opening a can of worms in terms of possible prompt injections. So I think we should do it really carefully. And outside of security considerations, humans are simple wrong too. So my plan with sashiko is to force it to try to verify the feedback against the codebase and if it's not possible trust only maintainers or people with a long history of meaningful contributions. And sashiko can (and does in my prototype) generate prompts based on the feedback and automatically verify that it helps by re-reviewing original patches. This should mostly eliminate a need for human-generated prompts. This is all doable, but probably will take few more week to build and roll out. ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-05 21:14 ` Roman Gushchin @ 2026-10-05 21:53 ` Jakub Kicinski 2026-10-05 23:37 ` Ihor Solodrai 1 sibling, 0 replies; 12+ messages in thread From: Jakub Kicinski @ 2026-10-05 21:53 UTC (permalink / raw) To: Roman Gushchin; +Cc: Chris Mason, sashiko, ihor.solodrai, ast On Mon, 05 Oct 2026 21:14:12 +0000 Roman Gushchin wrote: > I plan to do this (and had a prototype in the past), but it's tricky if > we take security seriously. And we absolutely should! > Obviously just giving an agent an access to lore archive is opening a > can of worms in terms of possible prompt injections. So I think we > should do it really carefully. Hm, is there really much extra attack surface here beyond what people can already put in a commit message of a fake patch? > And outside of security considerations, > humans are simple wrong too. So my plan with sashiko is to force it to > try to verify the feedback against the codebase and if it's not possible > trust only maintainers or people with a long history of meaningful > contributions. Definitely, I was also wondering about doing some rough estimate of "trustworthiness" of the person based on how long they have been contributing + MAINTAINERS status. In practice, tho, noobs rarely provide feedback. Closing the loop with patchwork is a better idea (if patch go rejected == maintainer must have judged some feedback as important) > And sashiko can (and does in my prototype) generate prompts based on the > feedback and automatically verify that it helps by re-reviewing original > patches. This should mostly eliminate a need for human-generated > prompts. > This is all doable, but probably will take few more week to build and > roll out. ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-05 21:14 ` Roman Gushchin 2026-10-05 21:53 ` Jakub Kicinski @ 2026-10-05 23:37 ` Ihor Solodrai 1 sibling, 0 replies; 12+ messages in thread From: Ihor Solodrai @ 2026-10-05 23:37 UTC (permalink / raw) To: Roman Gushchin, Jakub Kicinski; +Cc: Chris Mason, sashiko, ast On 10/5/26 2:14 PM, Roman Gushchin wrote: > Jakub Kicinski <kuba@kernel.org> writes: > >> [...] >> >> Closing the loop with Sashiko and/or ML would be great :( >> If the bots read the ML they could both learn false positives and false >> negatives automatically. I suspect most of us thought about this by now. >> If we're doing a redesign should the ingest of ML be part of it? FWIW the BPF bot has been reading lore when reviewing patches through semcode since February: https://github.com/kernel-patches/vmtest/pull/442 Basically, the container in which the AI session is running has access to local semcode via MCP, and local semcode db has the full (bpf) lore archive. It definitely helps with quality to give AI more context, but it also has side effects. For example, a bot may repeat an issue raised by another bot in a previous revision. In many cases this is noise, but sometimes it is appropriate: the review may have been inappropriately ignored. Currently this is mostly hidden by rules in the prompts to not repeat issues that have already been discussed, because it appears that people ignore many of the AI reviews anyways. > > I plan to do this (and had a prototype in the past), but it's tricky if > we take security seriously. And we absolutely should! > Obviously just giving an agent an access to lore archive is opening a > can of worms in terms of possible prompt injections. So I think we > should do it really carefully. [...] Re prompt injections, this is of course always a concern, and my yolo attitude may not appeal to everyone. But I have to say since almost a year of running AI reviews on BPF this hasn't been a problem once. This of course depends on how the infrastructure is set up. BPF bot is quite well isolated: * dedicated AWS account * each job runs on an ephemeral runner: not only new container, but also ephemeral hosts (we use CodeBuild) * the API access to inference is limited to 1h in a job * the AWS and github access tokens have limited permissions With all that, I sleep well. One can imagine a sci-fi scenario with a prompt injection tiggering a bot to take over the AWS account, but given corporate monitoring of activity there, we'd detect it and turn everything off immediately. I also think that as the end-users of the LLM APIs we automatically benefit from all the protections on the provider side. Sometimes this is even annoying, because some models refuse to investigate certain issues due to the guardrails on the backend. > And outside of security considerations, > humans are simple wrong too. So my plan with sashiko is to force it to > try to verify the feedback against the codebase and if it's not possible > trust only maintainers or people with a long history of meaningful > contributions. I have small local datasets generated from mailing lists, and one way I tried to extract signal is to check whether the feedback from the bot was incorporated in the next revision of the patches. Some people do that silently without crediting the bots. So yeah, code is the authority. Reputation plays some role of course, but in the end the bots should check the claims against the code and discussion history. > And sashiko can (and does in my prototype) generate prompts based on the > feedback and automatically verify that it helps by re-reviewing original > patches. This should mostly eliminate a need for human-generated > prompts. I did this manually a few times over the year: analyse lore -> come up with prompt improvements -> test them -> merge. This is definitely automatable and can close the loop to a large degree. Prompting and context management is not that far from actual machine learning, as it turns out :) > This is all doable, but probably will take few more week to build and > roll out. ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason 2026-10-02 21:18 ` Chuck Lever 2026-10-05 19:58 ` Jakub Kicinski @ 2026-10-06 21:22 ` Ihor Solodrai 2026-10-07 9:00 ` Chris Mason 2026-10-09 14:17 ` Fuad Tabba 3 siblings, 1 reply; 12+ messages in thread From: Ihor Solodrai @ 2026-10-06 21:22 UTC (permalink / raw) To: Chris Mason, sashiko, Roman Gushchin, ast, kuba On 10/2/26 12:04 PM, Chris Mason wrote: > Hi everyone, > > I've been updating the review prompts, and you can find my current work here in the subsystem-build branch: > > https://github.com/masoncl/review-prompts.git subsystem-build Hi Chris, I've set up "shadow" AI reviews for BPF, taking advantage of our existing bpf-rc infra. It's a copy of the entire BPF CI pipeline that we occasionally use for testing of the CI itself. The shadow reviews are using your new prompts from the subsystem-build branch. The reviews are posted on kernel-patches/bpf-rc PRs as comments and are not mailed to the list. I configured the reviews to run on every 4th series (so 25% of prod). I'll let it run for some time (couple of weeks?) and then we'll be able to analyze the difference and either switch or iterate. One quirk I added: a datetime cutoff of the semcode's lore archive, so that the rc reviewer doesn't see prod reviewer's comments. Not sure how reliable that'll be, because AI may potentially figure out to read them from github or something. But we should notice if it doesn't work. You can find the relevant PRs like this: $ gh pr list -R kernel-patches/bpf-rc --label ai-review-shadow:subsystem-build --state all --limit 50 \ --json url,headRefName,title --jq '.[] | "\(.url)\t\(.headRefName)\t\(.title)"' Note the "ai-review-shadow:subsystem-build" PR label. The reviews run with debug output, and also the claude session logs are captured as artifacts. Ask your agent to write a script to pull those in with github cli, should be easy. I just pushed, it'll take a few days before we get any real samples. Thank you for working on the prompts! > > [...] ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-06 21:22 ` Ihor Solodrai @ 2026-10-07 9:00 ` Chris Mason 0 siblings, 0 replies; 12+ messages in thread From: Chris Mason @ 2026-10-07 9:00 UTC (permalink / raw) To: Ihor Solodrai, sashiko, Roman Gushchin, ast, Jakub Kicinski On Tue, Oct 6, 2026, at 11:22 PM, Ihor Solodrai wrote: > On 10/2/26 12:04 PM, Chris Mason wrote: >> Hi everyone, >> >> I've been updating the review prompts, and you can find my current >> work here in the subsystem-build branch: >> >> https://github.com/masoncl/review-prompts.git subsystem-build > > Hi Chris, > > I've set up "shadow" AI reviews for BPF, taking advantage of our > existing bpf-rc infra. It's a copy of the entire BPF CI pipeline that > we occasionally use for testing of the CI itself. > > The shadow reviews are using your new prompts from the subsystem-build > branch. The reviews are posted on kernel-patches/bpf-rc PRs as > comments and are not mailed to the list. I configured the reviews to > run on every 4th series (so 25% of prod). I'll let it run for some > time (couple of weeks?) and then we'll be able to analyze the > difference and either switch or iterate. > > One quirk I added: a datetime cutoff of the semcode's lore archive, > so that the rc reviewer doesn't see prod reviewer's comments. Not > sure how reliable that'll be, because AI may potentially figure out > to read them from github or something. But we should notice if it > doesn't work. > > You can find the relevant PRs like this: > > $ gh pr list -R kernel-patches/bpf-rc --label ai-review-shadow:subsystem- > build --state all --limit 50 \ --json url,headRefName,title --jq > '.[] | "\(.url)\t\(.headRefName)\t\(.title)"' > > Note the "ai-review-shadow:subsystem-build" PR label. > > The reviews run with debug output, and also the claude session logs > are captured as artifacts. Ask your agent to write a script to pull > those in with github cli, should be easy. > > I just pushed, it'll take a few days before we get any real samples. Oh this is great, thanks so much. -chris ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason ` (2 preceding siblings ...) 2026-10-06 21:22 ` Ihor Solodrai @ 2026-10-09 14:17 ` Fuad Tabba 2026-10-09 17:57 ` Chris Mason 3 siblings, 1 reply; 12+ messages in thread From: Fuad Tabba @ 2026-10-09 14:17 UTC (permalink / raw) To: Chris Mason; +Cc: sashiko, Roman Gushchin, ihor.solodrai, ast, kuba, Fuad Tabba Hi Chris, On Fri, 02 Oct 2026 20:04:00 +0100, "Chris Mason" <mason@kernel.org> wrote: [...] > My new branch builds subsystem guides by asking a long list of > questions, and then extensively reading the sources to find the > right answers. The delta between the LLM's answers and the right > answers is the new guide. I went through the KVM/arm64 and protected KVM (pKVM) guides on the branch, and checked a sample of their claims against the tree. They're much more accurate and detailed than the hand-written ones, and the checking against the source is what makes the difference. Some of what a reviewer needs isn't in the code, though: which bugs matter and how much, and which reports are false alarms. The old pKVM guide said that a hypervisor crash only the host kernel can trigger is a hardening item, while one a guest can trigger is a real bug. The design deliberately keeps that kind of instruction out of the guides, so it's gone from the built one. Where should it live instead? The "Conventions for new code" files look closest: they already hold what maintainers ask for and no code states, and a review finds them by the directory a patch touches. Something like that for KVM/arm64 would work, and I can write it. Leaving out what the tested models already knew also tunes the guide to those models, and the model doing the review may not be one of them. Hand-written guides have the same problem, and I don't see an easy fix. Your mail points at builds per model, so one option is for each project to build for the model it runs. Another is for the build to keep the full checked answer rather than only the difference, now that the index makes length matter less. Which way are you leaning? > This is both much less useful to human readers and much longer. I'm > not sure what to say about the human reader part, but instead of > having LLMs read the whole subsystem guide, I shifted to an index > where they search for symbols. This is a better fit for more > advanced models, which mostly need updates on how the kernel has > changed since they were trained. A common KVM patch adds a new hypercall. A symbol search finds the answers about the existing calls the patch uses, but not the rule for where a new one goes in the list: that answer is filed under a marker the diff never touches, and the header the diff edits isn't the source file of any index line. Could some answers be keyed to the directory as well, like the short table you kept for a few guides? [...] > What I know for sure is the existing review prompts have drifted > from mainline Linus. It's impacting the quality of the reviews, so > I plan on working out something in the near future. Built guides drift too, just in a different way. These were built at 7.3-rc5, and one fix in rc6 made four of the KVM/arm64 claims wrong without renaming anything, so looking names up in the tree doesn't catch it. Most index lines already record a source file, so a review on a newer tree could check whether that file changed since the build, and re-check or flag the answer if it did. Related: when a maintainer finds a wrong answer, the only fix is to change the question and rebuild, which needs model access and can come out differently each time. Who looks after the question files, and could there be a small override that a maintainer edits between rebuilds? Feedback from the list, as Jakub and Roman discussed, could take the same route: suggested questions that go through the same checks against the tree, rather than guide text. Separately, some pKVM questions were dropped for size, including the one on protected guest system registers, and the measurement notes say they're the first to bring back if the size limit is raised. There's no length limit now, so I could send a patch to bring them back. Cheers, /fuad ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-09 14:17 ` Fuad Tabba @ 2026-10-09 17:57 ` Chris Mason 2026-10-10 15:53 ` Fuad Tabba 0 siblings, 1 reply; 12+ messages in thread From: Chris Mason @ 2026-10-09 17:57 UTC (permalink / raw) To: Fuad Tabba Cc: sashiko, Roman Gushchin, Ihor Solodrai, ast, Jakub Kicinski, Fuad Tabba On Fri, Oct 9, 2026, at 10:17 AM, Fuad Tabba wrote: > Hi Chris, > > On Fri, 02 Oct 2026 20:04:00 +0100, "Chris Mason" <mason@kernel.org> wrote: > [...] >> My new branch builds subsystem guides by asking a long list of >> questions, and then extensively reading the sources to find the >> right answers. The delta between the LLM's answers and the right >> answers is the new guide. > > I went through the KVM/arm64 and protected KVM (pKVM) guides on the > branch, and checked a sample of their claims against the tree. They're > much more accurate and detailed than the hand-written ones, and the > checking against the source is what makes the difference. Great, I only did limited verification in that area, so its good to hear. > > Some of what a reviewer needs isn't in the code, though: which bugs > matter and how much, and which reports are false alarms. The old pKVM > guide said that a hypervisor crash only the host kernel can trigger is a > hardening item, while one a guest can trigger is a real bug. The design > deliberately keeps that kind of instruction out of the guides, so it's > gone from the built one. Where should it live instead? The "Conventions > for new code" files look closest: they already hold what maintainers ask > for and no code states, and a review finds them by the directory a patch > touches. Something like that for KVM/arm64 would work, and I can write > it. There's already a way to pass specific parts verbatim, which feels more correct than under a "new code" label. We can play around with a few ideas though. > > Leaving out what the tested models already knew also tunes the guide to > those models, and the model doing the review may not be one of them. Yes. I generated the current built prompts using both sonnet and opus with the assumption that other models would get roughly the same things wrong as one of the two of them. This is obviously flawed, but the prompts have a way to pick the best build directory based on kernel version, so we could extend that to the model doing the review (more below). > Hand-written guides have the same problem, and I don't see an easy fix. > Your mail points at builds per model, so one option is for each project > to build for the model it runs. Another is for the build to keep the > full checked answer rather than only the difference, now that the index > makes length matter less. Which way are you leaning? My original plan was to get the built guides small enough to be included whole, in the same way the original guides were. Adding the keyword index was basically me admitting defeat, and it just didn't occur to me that we might be able to index our way to victory on the full output. IOW, great idea, I didn't think of that at all ;) I'll do another run and put the full build up. We can see if it's too big to be useful or if we can make the index strong enough to get around needing the per-model analysis. > >> This is both much less useful to human readers and much longer. I'm >> not sure what to say about the human reader part, but instead of >> having LLMs read the whole subsystem guide, I shifted to an index >> where they search for symbols. This is a better fit for more >> advanced models, which mostly need updates on how the kernel has >> changed since they were trained. > > A common KVM patch adds a new hypercall. A symbol search finds the > answers about the existing calls the patch uses, but not the rule for > where a new one goes in the list: that answer is filed under a marker > the diff never touches, and the header the diff edits isn't the source > file of any index line. Could some answers be keyed to the directory as > well, like the short table you kept for a few guides? > Absolutely. The index is pretty dumb, we can do a lot better. > [...] >> What I know for sure is the existing review prompts have drifted >> from mainline Linus. It's impacting the quality of the reviews, so >> I plan on working out something in the near future. > > Built guides drift too, just in a different way. These were built at > 7.3-rc5, and one fix in rc6 made four of the KVM/arm64 claims wrong > without renaming anything, so looking names up in the tree doesn't catch > it. Most index lines already record a source file, so a review on a > newer tree could check whether that file changed since the build, and > re-check or flag the answer if it did. A related question is how often do I need to rebuild in order for the guides to be useful? I'd assume the absolute minimum is every final release, but every RC is also reasonable. > > Related: when a maintainer finds a wrong answer, the only fix is to > change the question and rebuild, which needs model access and can come > out differently each time. Who looks after the question files, and could > there be a small override that a maintainer edits between rebuilds? > Feedback from the list, as Jakub and Roman discussed, could take the > same route: suggested questions that go through the same checks against > the tree, rather than guide text. The build script can rebuild a single guide, or people can just have their agents hand edit the build? It's a good point, we should have AGENTS.md record some best practices. > > Separately, some pKVM questions were dropped for size, including the one > on protected guest system registers, and the measurement notes say > they're the first to bring back if the size limit is raised. There's no > length limit now, so I could send a patch to bring them back. Great, please do. -chris ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC] reworking the review-prompts subsystem guide 2026-10-09 17:57 ` Chris Mason @ 2026-10-10 15:53 ` Fuad Tabba 0 siblings, 0 replies; 12+ messages in thread From: Fuad Tabba @ 2026-10-10 15:53 UTC (permalink / raw) To: Chris Mason; +Cc: sashiko, Roman Gushchin, ihor.solodrai, ast, kuba, Fuad Tabba Hi Chris, On Fri, 09 Oct 2026 18:57:53 +0100, "Chris Mason" <mason@kernel.org> wrote: [...] > There's already a way to pass specific parts verbatim, which feels more > correct than under a "new code" label. We can play around with a few > ideas though. A verbatim file works. I'll send a KVM/arm64 one, policy only, with the patch that brings back the dropped questions. [...] > I'll do another run and put the full build up. We can see if it's too big > to be useful or if we can make the index strong enough to get around > needing the per-model analysis. When it's up I can put numbers on it: I have a test set of 115 buggy KVM/arm64 commits, fixed since January, each paired with what fixed it, plus the Sashiko findings on KVM/arm64 patches since July that a human answered. The hand guides are running on the bugs now; the current build and the full one are next. [...] > Absolutely. The index is pretty dumb, we can do a lot better. Most index lines already name the source file of their answer, so one cheap step is to search for the files the patch touches as well as its symbols. A directory table is still the place for the rest: the placement rule for a new hypercall was answered from arch/arm64/kvm/pkvm.c, which such a patch doesn't have to touch. [...] > A related question is how often do I need to rebuild in order for > the guides to be useful? I'd assume the absolute minimum is every > final release, but every RC is also reasonable. I think every rc is enough for mainline. What it can't reach is a fix that lands between rcs, or a review against -next, and the source file on the index line could catch part of that: re-check an answer whose file changed since the build's commit before relying on it. [...] > The build script can rebuild a single guide, or people can just have > their agents hand edit the build? It's a good point, we should have > AGENTS.md record some best practices. A hand edit to the guide fails check-built-guide.py, and a rebuild from the unchanged question can bring the wrong answer back. For AGENTS.md, how about: edit the answer file (an exception to the never-edit-build rule, but the check still passes), re-render and re-index with no model, and sharpen the question in the same change so the fix isn't only in the output. Cheers, /fuad ^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-10-10 15:53 UTC | newest] Thread overview: 12+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-10-02 19:04 [RFC] reworking the review-prompts subsystem guide Chris Mason 2026-10-02 21:18 ` Chuck Lever 2026-10-04 9:17 ` Chris Mason 2026-10-05 19:58 ` Jakub Kicinski 2026-10-05 21:14 ` Roman Gushchin 2026-10-05 21:53 ` Jakub Kicinski 2026-10-05 23:37 ` Ihor Solodrai 2026-10-06 21:22 ` Ihor Solodrai 2026-10-07 9:00 ` Chris Mason 2026-10-09 14:17 ` Fuad Tabba 2026-10-09 17:57 ` Chris Mason 2026-10-10 15:53 ` Fuad Tabba
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.