* [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
[not found] <b5a33433-b0b9-4231-9998-23e2a2202091@kernel.dk>
@ 2026-10-08 17:06 ` Hengyu Liang
2026-10-08 19:11 ` Jens Axboe
2026-10-09 11:23 ` Pavel Begunkov
0 siblings, 2 replies; 4+ messages in thread
From: Hengyu Liang @ 2026-10-08 17:06 UTC (permalink / raw)
To: Jens Axboe; +Cc: Pavel Begunkov, David Wei, io-uring, linux-kernel, netdev
Commit 8078486e1d53 ("io_uring: use region api for SQ") and commit
81a4058e0cd0 ("io_uring: use region api for CQ") made io_uring_setup()
create the rings with io_create_region().
However, io_create_region() charges user provided memory to
RLIMIT_MEMLOCK, and the rings of an IORING_SETUP_NO_MMAP ring were not
charged before those commits. As of now, a user without CAP_IPC_LOCK
gets ENOMEM from io_uring_queue_init_mem() when their rings exceed the
limit, which is 8 MiB by default. PostgreSQL 18 creates its rings with
this function [1].
The issue can be reproduced with a simple liburing program, run as an
unprivileged user:
#include <liburing.h>
#include <stdio.h>
#include <sys/mman.h>
int main(void)
{
static struct io_uring ring[64];
char *mem = mmap(NULL, 64 << 20, PROT_READ | PROT_WRITE,
MAP_SHARED | MAP_ANONYMOUS, -1, 0);
int i;
for (i = 0; i < 64; i++) {
struct io_uring_params p = { };
if (io_uring_queue_init_mem(4096, &ring[i], &p,
mem + (i << 20),
1 << 20) < 0)
break;
}
printf("%d rings\n", i);
return 0;
}
Before those commits (v6.13), it prints "64 rings". After those commits
(v6.14), it prints "21 rings".
This patch makes io_create_region() take the user to charge, like
io_free_region() does, and passes no user for the SQ/CQ rings.
Link: https://github.com/postgres/postgres/blob/REL_18_STABLE/src/backend/storage/aio/method_io_uring.c [1]
Fixes: 8078486e1d53 ("io_uring: use region api for SQ")
Fixes: 81a4058e0cd0 ("io_uring: use region api for CQ")
Cc: stable@vger.kernel.org
Signed-off-by: Hengyu Liang <hengyul@cs.unc.edu>
---
This goes on top of io_uring-7.3 (c746673517c6), as discussed in [2].
Tested on that branch as a user without CAP_IPC_LOCK and the default
8 MiB limit:
before after
program above 21 64
the same with io_uring_queue_init() 64 64
30 processes x 142 NO_MMAP rings of 64 entries 7 30
The last row is what 30 PostgreSQL 18 clusters with default settings
create. The number is how many of them got all their rings.
The per-user locked_vm count stays balanced over ring create, close and
resize. Provided buffer rings and IORING_REGISTER_MEM_REGION regions on
user memory are charged as before. Every liburing test gives the same
result with and without the patch.
Not changed here: provided buffer rings on user memory, which is what
io_uring_setup_buf_ring() registers, were not charged in v6.13 either
(64 of 64 rings of 32768 entries then, 16 now). I can send a patch for
those as well if you want them handled the same way.
[2] https://lore.kernel.org/io-uring/b5a33433-b0b9-4231-9998-23e2a2202091@kernel.dk/
io_uring/io_uring.c | 8 ++++----
io_uring/kbuf.c | 2 +-
io_uring/memmap.c | 7 ++++---
io_uring/memmap.h | 2 +-
io_uring/register.c | 10 +++++-----
io_uring/zcrx.c | 2 +-
6 files changed, 16 insertions(+), 15 deletions(-)
diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c
index c2ce83c7c1f1..2a16369894d4 100644
--- a/io_uring/io_uring.c
+++ b/io_uring/io_uring.c
@@ -2070,8 +2070,8 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned int nr)
static void io_rings_free(struct io_ring_ctx *ctx)
{
- io_free_region(ctx->user, &ctx->sq_region);
- io_free_region(ctx->user, &ctx->ring_region);
+ io_free_region(NULL, &ctx->sq_region);
+ io_free_region(NULL, &ctx->ring_region);
ctx->rings = NULL;
RCU_INIT_POINTER(ctx->rings_rcu, NULL);
ctx->sq_sqes = NULL;
@@ -2735,7 +2735,7 @@ static __cold int io_allocate_scq_urings(struct io_ring_ctx *ctx,
rd.user_addr = p->cq_off.user_addr;
rd.flags |= IORING_MEM_REGION_TYPE_USER;
}
- ret = io_create_region(ctx, &ctx->ring_region, &rd, IORING_OFF_CQ_RING);
+ ret = io_create_region(NULL, &ctx->ring_region, &rd, IORING_OFF_CQ_RING);
if (ret)
return ret;
ctx->rings = rings = io_region_get_ptr(&ctx->ring_region);
@@ -2749,7 +2749,7 @@ static __cold int io_allocate_scq_urings(struct io_ring_ctx *ctx,
rd.user_addr = p->sq_off.user_addr;
rd.flags |= IORING_MEM_REGION_TYPE_USER;
}
- ret = io_create_region(ctx, &ctx->sq_region, &rd, IORING_OFF_SQES);
+ ret = io_create_region(NULL, &ctx->sq_region, &rd, IORING_OFF_SQES);
if (ret) {
io_rings_free(ctx);
return ret;
diff --git a/io_uring/kbuf.c b/io_uring/kbuf.c
index 7c309173dd19..7c59ab9cfc8b 100644
--- a/io_uring/kbuf.c
+++ b/io_uring/kbuf.c
@@ -676,7 +676,7 @@ int io_register_pbuf_ring(struct io_ring_ctx *ctx, void __user *arg)
rd.user_addr = reg.ring_addr;
rd.flags |= IORING_MEM_REGION_TYPE_USER;
}
- ret = io_create_region(ctx, &bl->region, &rd, mmap_offset);
+ ret = io_create_region(ctx->user, &bl->region, &rd, mmap_offset);
if (ret)
goto fail;
br = io_region_get_ptr(&bl->region);
diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index da2328b52b38..51afe5d40b1d 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -192,7 +192,7 @@ static int io_region_allocate_pages(struct io_mapped_region *mr,
return 0;
}
-int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
+int io_create_region(struct user_struct *user, struct io_mapped_region *mr,
struct io_uring_region_desc *reg,
unsigned long mmap_offset)
{
@@ -223,9 +223,10 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
/*
* Only pinned user memory counts against RLIMIT_MEMLOCK, kernel
* allocated regions are memcg accounted through GFP_KERNEL_ACCOUNT.
+ * The SQ/CQ rings are never charged, their callers pass a NULL user.
*/
if (reg->flags & IORING_MEM_REGION_TYPE_USER)
- ret = io_region_pin_pages(mr, reg, ctx->user);
+ ret = io_region_pin_pages(mr, reg, user);
else
ret = io_region_allocate_pages(mr, reg, mmap_offset);
if (ret)
@@ -236,7 +237,7 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
goto out_free;
return 0;
out_free:
- io_free_region(ctx->user, mr);
+ io_free_region(user, mr);
return ret;
}
diff --git a/io_uring/memmap.h b/io_uring/memmap.h
index f4cfbb6b9a1f..0714bb5a9616 100644
--- a/io_uring/memmap.h
+++ b/io_uring/memmap.h
@@ -18,7 +18,7 @@ unsigned long io_uring_get_unmapped_area(struct file *file, unsigned long addr,
int io_uring_mmap(struct file *file, struct vm_area_struct *vma);
void io_free_region(struct user_struct *user, struct io_mapped_region *mr);
-int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
+int io_create_region(struct user_struct *user, struct io_mapped_region *mr,
struct io_uring_region_desc *reg,
unsigned long mmap_offset);
diff --git a/io_uring/register.c b/io_uring/register.c
index 02bc103bcc9d..d79971c57ace 100644
--- a/io_uring/register.c
+++ b/io_uring/register.c
@@ -480,8 +480,8 @@ struct io_ring_ctx_rings {
static void io_register_free_rings(struct io_ring_ctx *ctx,
struct io_ring_ctx_rings *r)
{
- io_free_region(ctx->user, &r->sq_region);
- io_free_region(ctx->user, &r->ring_region);
+ io_free_region(NULL, &r->sq_region);
+ io_free_region(NULL, &r->ring_region);
}
#define swap_old(ctx, o, n, field) \
@@ -529,7 +529,7 @@ static int io_register_resize_rings(struct io_ring_ctx *ctx, void __user *arg)
rd.user_addr = p->cq_off.user_addr;
rd.flags |= IORING_MEM_REGION_TYPE_USER;
}
- ret = io_create_region(ctx, &n.ring_region, &rd, IORING_OFF_CQ_RING);
+ ret = io_create_region(NULL, &n.ring_region, &rd, IORING_OFF_CQ_RING);
if (ret)
return ret;
@@ -559,7 +559,7 @@ static int io_register_resize_rings(struct io_ring_ctx *ctx, void __user *arg)
rd.user_addr = p->sq_off.user_addr;
rd.flags |= IORING_MEM_REGION_TYPE_USER;
}
- ret = io_create_region(ctx, &n.sq_region, &rd, IORING_OFF_SQES);
+ ret = io_create_region(NULL, &n.sq_region, &rd, IORING_OFF_SQES);
if (ret) {
io_register_free_rings(ctx, &n);
return ret;
@@ -730,7 +730,7 @@ static int io_register_mem_region(struct io_ring_ctx *ctx, void __user *uarg)
!(ctx->flags & IORING_SETUP_R_DISABLED))
return -EINVAL;
- ret = io_create_region(ctx, ®ion, &rd, IORING_MAP_OFF_PARAM_REGION);
+ ret = io_create_region(ctx->user, ®ion, &rd, IORING_MAP_OFF_PARAM_REGION);
if (ret)
return ret;
if (copy_to_user(rd_uptr, &rd, sizeof(rd))) {
diff --git a/io_uring/zcrx.c b/io_uring/zcrx.c
index beb35077f3ce..939885985333 100644
--- a/io_uring/zcrx.c
+++ b/io_uring/zcrx.c
@@ -432,7 +432,7 @@ static int io_allocate_rbuf_ring(struct io_ring_ctx *ctx,
mmap_offset = IORING_MAP_OFF_ZCRX_REGION;
mmap_offset += (u64)id << IORING_OFF_ZCRX_SHIFT;
- ret = io_create_region(ctx, &ifq->rq_region, rd, mmap_offset);
+ ret = io_create_region(ctx->user, &ifq->rq_region, rd, mmap_offset);
if (ret < 0)
return ret;
base-commit: c746673517c6ce9f5400cc0ea23e10ef5382eddb
--
2.53.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
2026-10-08 17:06 ` [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK Hengyu Liang
@ 2026-10-08 19:11 ` Jens Axboe
2026-10-09 4:14 ` Hengyu Liang
2026-10-09 11:23 ` Pavel Begunkov
1 sibling, 1 reply; 4+ messages in thread
From: Jens Axboe @ 2026-10-08 19:11 UTC (permalink / raw)
To: Hengyu Liang; +Cc: Pavel Begunkov, David Wei, io-uring, linux-kernel, netdev
On 10/8/26 11:06 AM, Hengyu Liang wrote:
> This goes on top of io_uring-7.3 (c746673517c6), as discussed in [2].
>
> Tested on that branch as a user without CAP_IPC_LOCK and the default
> 8 MiB limit:
>
> before after
> program above 21 64
> the same with io_uring_queue_init() 64 64
> 30 processes x 142 NO_MMAP rings of 64 entries 7 30
>
> The last row is what 30 PostgreSQL 18 clusters with default settings
> create. The number is how many of them got all their rings.
>
> The per-user locked_vm count stays balanced over ring create, close and
> resize. Provided buffer rings and IORING_REGISTER_MEM_REGION regions on
> user memory are charged as before. Every liburing test gives the same
> result with and without the patch.
>
> Not changed here: provided buffer rings on user memory, which is what
> io_uring_setup_buf_ring() registers, were not charged in v6.13 either
> (64 of 64 rings of 32768 entries then, 16 now). I can send a patch for
> those as well if you want them handled the same way.
Honestly, after taking a closer look at this, I think we're better off
with your original patch and one on top for pbuf rings. Please check:
https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git/log/?h=io_uring-7.3
for the top 2 commits. If you can re-test one more time, that'd be
great...
--
Jens Axboe
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
2026-10-08 19:11 ` Jens Axboe
@ 2026-10-09 4:14 ` Hengyu Liang
0 siblings, 0 replies; 4+ messages in thread
From: Hengyu Liang @ 2026-10-09 4:14 UTC (permalink / raw)
To: Jens Axboe; +Cc: Pavel Begunkov, David Wei, io-uring, linux-kernel, netdev
On 10/8/26 1:11 PM, Jens Axboe wrote:
> Honestly, after taking a closer look at this, I think we're better off
> with your original patch and one on top for pbuf rings. Please check:
>
> https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git/log/?h=io_uring-7.3
>
> for the top 2 commits. If you can re-test one more time, that'd be
> great...
That works for me, and the two commits look good.
I tested io_uring-7.3 at 9f0c88b7f884 as a user without CAP_IPC_LOCK and
the default 8 MiB limit. Number of objects created out of 64, with 4096
entries per ring and 32768 entries per buffer ring:
v6.13 v7.3-rc4 9f0c88b7f884
rings, io_uring_queue_init() 64 16 64
rings, io_uring_queue_init_mem() 64 21 64
buffer rings, IOU_PBUF_RING_MMAP 64 15 64
buffer rings, user memory 64 15 64
30 processes with 142 NO_MMAP rings of 64 entries each, which is what 30
PostgreSQL 18 clusters with default settings create, all get their rings.
On v7.3-rc4 7 of them do.
The per-user locked_vm count stays balanced over ring create, close and
resize and over buffer ring register and unregister. Registered buffers
and IORING_REGISTER_MEM_REGION regions are charged and refused past the
limit as before. Every liburing test gives the same result as on the
previous tip of the branch (c746673517c6).
For the buffer ring patch:
Tested-by: Hengyu Liang <hengyul@cs.unc.edu>
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
2026-10-08 17:06 ` [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK Hengyu Liang
2026-10-08 19:11 ` Jens Axboe
@ 2026-10-09 11:23 ` Pavel Begunkov
1 sibling, 0 replies; 4+ messages in thread
From: Pavel Begunkov @ 2026-10-09 11:23 UTC (permalink / raw)
To: Hengyu Liang, Jens Axboe; +Cc: David Wei, io-uring, linux-kernel, netdev
On 10/8/26 18:06, Hengyu Liang wrote:
> Commit 8078486e1d53 ("io_uring: use region api for SQ") and commit
> 81a4058e0cd0 ("io_uring: use region api for CQ") made io_uring_setup()
> create the rings with io_create_region().
>
> However, io_create_region() charges user provided memory to
> RLIMIT_MEMLOCK, and the rings of an IORING_SETUP_NO_MMAP ring were not
> charged before those commits. As of now, a user without CAP_IPC_LOCK
> gets ENOMEM from io_uring_queue_init_mem() when their rings exceed the
> limit, which is 8 MiB by default. PostgreSQL 18 creates its rings with
> this function [1].
That patch you mentioned 26bfa89e25f4 ("io_uring: place ring SQ/CQ
arrays under memcg memory limits") has always been a delayed time bomb,
though memlock is quite a nasty limit. Makes me wonder if there is a
way to migrate it to cgroups completely.
...> The issue can be reproduced with a simple liburing program, run as an
> unprivileged user:
>
> [2] https://lore.kernel.org/io-uring/b5a33433-b0b9-4231-9998-23e2a2202091@kernel.dk/
>
> io_uring/io_uring.c | 8 ++++----
> io_uring/kbuf.c | 2 +-
> io_uring/memmap.c | 7 ++++---
> io_uring/memmap.h | 2 +-
> io_uring/register.c | 10 +++++-----
> io_uring/zcrx.c | 2 +-
> 6 files changed, 16 insertions(+), 15 deletions(-)
>
> diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c
> index c2ce83c7c1f1..2a16369894d4 100644
> --- a/io_uring/io_uring.c
> +++ b/io_uring/io_uring.c
> @@ -2070,8 +2070,8 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned int nr)
>
> static void io_rings_free(struct io_ring_ctx *ctx)
> {
> - io_free_region(ctx->user, &ctx->sq_region);
> - io_free_region(ctx->user, &ctx->ring_region);
1. It's not perfect to make all these changes for a fix because of
backporting. Let's simplify it, add a wrapper and use the "__" version
only where needed.
__io_create_region(ctx, bool account, ...) {
if (!ctx->user)
account = false;
...
}
io_create_region(ctx, ...) {
return __io_create_region(ctx, true, ...);
}
2. The need to match alloc and free arguments has a high chance to
eventually blow up. It'd be better to turn it into a flag.
__io_create_region() {
if (account) {
...
mr->flags |= IO_REGION_F_ACCOUNTED;
}
}
io_free_region() {
if ((mr->flags & IO_REGION_F_ACCOUNTED)) {
WARN_ON_ONCE(!user);
unaccount(user);
}
}
--
Pavel Begunkov
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-09 11:23 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <b5a33433-b0b9-4231-9998-23e2a2202091@kernel.dk>
2026-10-08 17:06 ` [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK Hengyu Liang
2026-10-08 19:11 ` Jens Axboe
2026-10-09 4:14 ` Hengyu Liang
2026-10-09 11:23 ` Pavel Begunkov
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox