From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.19]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EDF4E486650 for ; Thu, 6 Aug 2026 17:23:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.19 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786036985; cv=none; b=As1gWUEook3KWPv6SN7I33hWqQBZMH4BgGCA2rdDjhQwHEzt94XbsvYsBQytCwYE2hPxzXG7KLNE9ETk3Yhr3tPqEYqsYPdtKToicREu2rUC5IwUCdsT7wSTsPeZgfelUbT+TwaoYWNjD/qEkRPXIo0Moxazg2Bmvxk+/Tn0qDQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786036985; c=relaxed/simple; bh=fA3jpZHfe5h3RJ10UdpcihyeBJwHh7wLoaj6V6h6ZSQ=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=g3j7xTSvKy0XX1m8Y1Lo4chkr42FXno3uX9bS0o2BL6ChXiPDlFqzgCK8NuwAnaSC6+2Mu++VLTYmNnbae6uBjt78s5NGoCZIxWpZ1FfRFpNyOhAMyrt61rfP1dg5z+fjgpLYTKBBomH+15Oh20cCpcWh55IvuBSPP715gqP8yM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=HcDJHYSp; arc=none smtp.client-ip=192.198.163.19 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="HcDJHYSp" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1786036982; x=1817572982; h=message-id:subject:from:to:cc:date:in-reply-to: references:content-transfer-encoding:mime-version; bh=fA3jpZHfe5h3RJ10UdpcihyeBJwHh7wLoaj6V6h6ZSQ=; b=HcDJHYSppbKJWBrdI7j3aqRctG/vnSedHoHa7IqhUP25e7ECCSfpOFdm xM90Db3WZviNuA9B2as8WvTaVuJB/Saue42nkiaHUPX/KbIEc4nA1l/NT SIPir04QpTox8aFNf7CJmhTmEHeIPiWAwGdyNGULZz+GbHp/t8TtTH3Gn A5iTosKLQRzZpV0REemhQtGhYQlaILlNM9zDCNqKTk/46Db0E29xWhzDp i1hBrNG1mzLPRCF/jeqnzl+voshiVy0cDpsbD07f/id23m8khfW15fDbM byh0xrg3wcNkEeTUZIcns0cXkZNPJHSZMjE5XJPUKbYfArnsv6gFhImbL w==; X-CSE-ConnectionGUID: hfkAYPqTTTK5ZmVdcsMhWQ== X-CSE-MsgGUID: BLwcjOk8ShK90nrSbeh6xA== X-IronPort-AV: E=McAfee;i="6800,10657,11867"; a="85609585" X-IronPort-AV: E=Sophos;i="6.25,208,1779174000"; d="scan'208";a="85609585" Received: from fmviesa006.fm.intel.com ([10.60.135.146]) by fmvoesa113.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 06 Aug 2026 10:22:59 -0700 X-CSE-ConnectionGUID: vhZFyp6KRkqYx8wN//7ZZA== X-CSE-MsgGUID: uXh9nmkqSVukL2HDHil1RA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,208,1779174000"; d="scan'208";a="257862271" Received: from schen9-mobl4.amr.corp.intel.com (HELO [10.125.109.116]) ([10.125.109.116]) by fmviesa006-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 06 Aug 2026 10:22:58 -0700 Message-ID: Subject: Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance From: Tim Chen To: "Chen, Yu C" , Lu Wang Cc: peterz@infradead.org, mingo@redhat.com, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, linux-kernel@vger.kernel.org, "chen.yu@linux.dev" Date: Thu, 06 Aug 2026 10:22:57 -0700 In-Reply-To: <59db4420-7995-4261-89d5-03aa306c1370@intel.com> References: <20260801122252.2476258-1-wanglu.priv@gmail.com> <59db4420-7995-4261-89d5-03aa306c1370@intel.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.58.1 (3.58.1-1.fc43) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Thu, 2026-08-06 at 23:35 +0800, Chen, Yu C wrote: > Hi Lu Wang, Tim, >=20 > On 8/1/2026 8:22 PM, Lu Wang wrote: > > A passive load-balance pass marks group_llc_balance as migrate_llc_task > > and queues active balance when it cannot move a task. The CPU stopper > > callback constructs a fresh lb_env, so preserve the migration type on > > the runqueue across the asynchronous boundary. > >=20 > > For CAS-directed active balance, reject a candidate whose preferred LLC > > does not match the destination LLC. This keeps the fallback from moving > > a task away from its preferred LLC. > >=20 >=20 > It looks like this proposal provides fine-grain control on per-task base > migration strategy is promising. >=20 > > +static inline bool > > +migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env) > > +{ > > + return sched_cache_enabled() && > > + env->migration_type =3D=3D migrate_llc_task && > > + READ_ONCE(p->preferred_llc) !=3D llc_id(env->dst_cpu); > > +} > > + >=20 > [ ... ] >=20 > > @@ -13654,6 +13672,7 @@ static int active_load_balance_cpu_stop(void *d= ata) > > .src_rq =3D busiest_rq, > > .idle =3D CPU_IDLE, > > .flags =3D LBF_ACTIVE_LB, > > + .migration_type =3D (enum migration_type)busiest_rq->active_balance= _type, >=20 > If we overwrite migration_type for ALB (default is 0, i.e. migrate_load), > then in can_migrate_task() a delayed task might not be migrated in ALB: >=20 > if ((p->se.sched_delayed) && (env->migration_type !=3D migrate_load)) This is a good catch. It may be easier to create a migrate_llc_task_alb type and pass that in migration type. Then modify the above as=20 if ((p->se.sched_delayed) && env->migration_type !=3D migrate_load=C2=A0 && env->migration_type !=3D migrate_llc_task_alb) return 0 That avoids creating two cpu stop functions. Tim > return 0; > So an enhanced approach I'm thinking of is to pass > migrate_llc_task information via env->flags: > > }; > > =20 > > schedstat_inc(sd->alb_count); > > diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h > > index 56acf502b..82084d405 100644 > > --- a/kernel/sched/sched.h > > +++ b/kernel/sched/sched.h > > @@ -1266,6 +1266,7 @@ struct rq { > > /* For active balancing */ > > int active_balance; > > int push_cpu; > > + int active_balance_type; /* enum migration_type */ >=20 > We can set env->flags by invoking different callbacks of=20 > stop_one_cpu_nowait(), > thus avoiding the need to introduce active_balance_type into rq - which= =20 > could > cause false sharing if cache-line alignment is broken. >=20 > something like: >=20 > #define LBF_ACTIVE_LB_LLC 0x40 >=20 > -static int active_load_balance_cpu_stop(void *data) > +static int __active_load_balance_cpu_stop(void *data, unsigned int=20 > lb_flags) > -static int active_load_balance_cpu_stop(void *data) > +static int __active_load_balance_cpu_stop(void *data, unsigned int=20 > lb_flags) > { > struct rq *busiest_rq =3D data; > int busiest_cpu =3D cpu_of(busiest_rq); > @@ -13659,7 +13687,7 @@ static int active_load_balance_cpu_stop(void *dat= a) > .src_cpu =3D busiest_rq->cpu, > .src_rq =3D busiest_rq, > .idle =3D CPU_IDLE, > - .flags =3D LBF_ACTIVE_LB, > + .flags =3D LBF_ACTIVE_LB | lb_flags, > }; >=20 > and in migrate_llc_task_wrong_dst(), we check env->flags & LBF_ACTIVE_LB_= LLC > instead. >=20 > static int active_load_balance_cpu_stop(void *data) > { > return __active_load_balance_cpu_stop(data, 0); > } >=20 > static int active_load_balance_llc_cpu_stop(void *data) > { > return __active_load_balance_cpu_stop(data, LBF_ACTIVE_LB_LLC);= =20 > <-- new flag > } >=20 > static inline cpu_stop_fn_t alb_stop_fn(struct lb_env *env) > { > if (env->migration_type =3D=3D migrate_llc_task) > return active_load_balance_llc_cpu_stop; >=20 > return active_load_balance_cpu_stop; > } >=20 > if (active_balance) { > stop_one_cpu_nowait(cpu_of(busiest), > alb_stop_fn(&env), busiest, > &busiest->active_balance_work); > } >=20 >=20 > thanks, > Chenyu >=20 > > struct cpu_stop_work active_balance_work; > > =20 > > /* CPU of this runqueue: */