* [PATCH v3] migration: Stop migration immediately in RDMA error paths
@ 2023-10-24 16:39 Peter Xu
2023-10-24 16:45 ` Juan Quintela
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Peter Xu @ 2023-10-24 16:39 UTC (permalink / raw)
To: qemu-devel
Cc: peterx, Zhijian Li, Markus Armbruster, Juan Quintela,
Fabiano Rosas, Thomas Huth
In multiple places, RDMA errors are handled in a strange way, where it only
sets qemu_file_set_error() but not stop the migration immediately.
It's not obvious what will happen later if there is already an error. Make
all such failures stop migration immediately.
Cc: Zhijian Li (Fujitsu) <lizhijian@fujitsu.com>
Cc: Markus Armbruster <armbru@redhat.com>
Cc: Juan Quintela <quintela@redhat.com>
Cc: Fabiano Rosas <farosas@suse.de>
Reported-by: Thomas Huth <thuth@redhat.com>
Signed-off-by: Peter Xu <peterx@redhat.com>
---
v3:
- in ram_save_complete() return directly with retval, drop the "ret<0"
check after the loop [Juan]
This patch is based on Thomas's patch:
[PATCH v2] migration/ram: Fix compilation with -Wshadow=local
https://lore.kernel.org/r/20231024092220.55305-1-thuth@redhat.com
Above patch should have been queued by both Markus and Juan.
---
migration/ram.c | 21 ++++++++++-----------
1 file changed, 10 insertions(+), 11 deletions(-)
diff --git a/migration/ram.c b/migration/ram.c
index 212add4481..6cb8b5cd2f 100644
--- a/migration/ram.c
+++ b/migration/ram.c
@@ -3034,11 +3034,13 @@ static int ram_save_setup(QEMUFile *f, void *opaque)
ret = rdma_registration_start(f, RAM_CONTROL_SETUP);
if (ret < 0) {
qemu_file_set_error(f, ret);
+ return ret;
}
ret = rdma_registration_stop(f, RAM_CONTROL_SETUP);
if (ret < 0) {
qemu_file_set_error(f, ret);
+ return ret;
}
migration_ops = g_malloc0(sizeof(MigrationOps));
@@ -3104,6 +3106,7 @@ static int ram_save_iterate(QEMUFile *f, void *opaque)
ret = rdma_registration_start(f, RAM_CONTROL_ROUND);
if (ret < 0) {
qemu_file_set_error(f, ret);
+ goto out;
}
t0 = qemu_clock_get_ns(QEMU_CLOCK_REALTIME);
@@ -3208,8 +3211,6 @@ static int ram_save_complete(QEMUFile *f, void *opaque)
rs->last_stage = !migration_in_colo_state();
WITH_RCU_READ_LOCK_GUARD() {
- int rdma_reg_ret;
-
if (!migration_in_postcopy()) {
migration_bitmap_sync_precopy(rs, true);
}
@@ -3217,6 +3218,7 @@ static int ram_save_complete(QEMUFile *f, void *opaque)
ret = rdma_registration_start(f, RAM_CONTROL_FINISH);
if (ret < 0) {
qemu_file_set_error(f, ret);
+ return ret;
}
/* try transferring iterative blocks of memory */
@@ -3232,24 +3234,21 @@ static int ram_save_complete(QEMUFile *f, void *opaque)
break;
}
if (pages < 0) {
- ret = pages;
- break;
+ qemu_mutex_unlock(&rs->bitmap_mutex);
+ return pages;
}
}
qemu_mutex_unlock(&rs->bitmap_mutex);
ram_flush_compressed_data(rs);
- rdma_reg_ret = rdma_registration_stop(f, RAM_CONTROL_FINISH);
- if (rdma_reg_ret < 0) {
- qemu_file_set_error(f, rdma_reg_ret);
+ ret = rdma_registration_stop(f, RAM_CONTROL_FINISH);
+ if (ret < 0) {
+ qemu_file_set_error(f, ret);
+ return ret;
}
}
- if (ret < 0) {
- return ret;
- }
-
ret = multifd_send_sync_main(rs->pss[RAM_CHANNEL_PRECOPY].pss_channel);
if (ret < 0) {
return ret;
--
2.41.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH v3] migration: Stop migration immediately in RDMA error paths
2023-10-24 16:39 [PATCH v3] migration: Stop migration immediately in RDMA error paths Peter Xu
@ 2023-10-24 16:45 ` Juan Quintela
2023-10-24 17:09 ` Fabiano Rosas
2023-10-25 6:25 ` Markus Armbruster
2 siblings, 0 replies; 4+ messages in thread
From: Juan Quintela @ 2023-10-24 16:45 UTC (permalink / raw)
To: Peter Xu
Cc: qemu-devel, Zhijian Li, Markus Armbruster, Fabiano Rosas,
Thomas Huth
Peter Xu <peterx@redhat.com> wrote:
> In multiple places, RDMA errors are handled in a strange way, where it only
> sets qemu_file_set_error() but not stop the migration immediately.
>
> It's not obvious what will happen later if there is already an error. Make
> all such failures stop migration immediately.
>
> Cc: Zhijian Li (Fujitsu) <lizhijian@fujitsu.com>
> Cc: Markus Armbruster <armbru@redhat.com>
> Cc: Juan Quintela <quintela@redhat.com>
> Cc: Fabiano Rosas <farosas@suse.de>
> Reported-by: Thomas Huth <thuth@redhat.com>
> Signed-off-by: Peter Xu <peterx@redhat.com>
Reviewed-by: Juan Quintela <quintela@redhat.com>
queued.
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] migration: Stop migration immediately in RDMA error paths
2023-10-24 16:39 [PATCH v3] migration: Stop migration immediately in RDMA error paths Peter Xu
2023-10-24 16:45 ` Juan Quintela
@ 2023-10-24 17:09 ` Fabiano Rosas
2023-10-25 6:25 ` Markus Armbruster
2 siblings, 0 replies; 4+ messages in thread
From: Fabiano Rosas @ 2023-10-24 17:09 UTC (permalink / raw)
To: Peter Xu, qemu-devel
Cc: peterx, Zhijian Li, Markus Armbruster, Juan Quintela, Thomas Huth
Peter Xu <peterx@redhat.com> writes:
> In multiple places, RDMA errors are handled in a strange way, where it only
> sets qemu_file_set_error() but not stop the migration immediately.
>
> It's not obvious what will happen later if there is already an error. Make
> all such failures stop migration immediately.
>
> Cc: Zhijian Li (Fujitsu) <lizhijian@fujitsu.com>
> Cc: Markus Armbruster <armbru@redhat.com>
> Cc: Juan Quintela <quintela@redhat.com>
> Cc: Fabiano Rosas <farosas@suse.de>
> Reported-by: Thomas Huth <thuth@redhat.com>
> Signed-off-by: Peter Xu <peterx@redhat.com>
Reviewed-by: Fabiano Rosas <farosas@suse.de>
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] migration: Stop migration immediately in RDMA error paths
2023-10-24 16:39 [PATCH v3] migration: Stop migration immediately in RDMA error paths Peter Xu
2023-10-24 16:45 ` Juan Quintela
2023-10-24 17:09 ` Fabiano Rosas
@ 2023-10-25 6:25 ` Markus Armbruster
2 siblings, 0 replies; 4+ messages in thread
From: Markus Armbruster @ 2023-10-25 6:25 UTC (permalink / raw)
To: Peter Xu
Cc: qemu-devel, Zhijian Li, Juan Quintela, Fabiano Rosas, Thomas Huth
Peter Xu <peterx@redhat.com> writes:
> In multiple places, RDMA errors are handled in a strange way, where it only
> sets qemu_file_set_error() but not stop the migration immediately.
>
> It's not obvious what will happen later if there is already an error. Make
> all such failures stop migration immediately.
>
> Cc: Zhijian Li (Fujitsu) <lizhijian@fujitsu.com>
> Cc: Markus Armbruster <armbru@redhat.com>
> Cc: Juan Quintela <quintela@redhat.com>
> Cc: Fabiano Rosas <farosas@suse.de>
> Reported-by: Thomas Huth <thuth@redhat.com>
> Signed-off-by: Peter Xu <peterx@redhat.com>
Good move! Since this already has competent review, I'll leave it
there.
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2023-10-25 6:26 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2023-10-24 16:39 [PATCH v3] migration: Stop migration immediately in RDMA error paths Peter Xu
2023-10-24 16:45 ` Juan Quintela
2023-10-24 17:09 ` Fabiano Rosas
2023-10-25 6:25 ` Markus Armbruster
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).