From mboxrd@z Thu Jan 1 00:00:00 1970 From: Kees Cook Subject: Re: [PATCH v15 00/17] arm64: untag user pointers passed to the kernel Date: Wed, 22 May 2019 12:21:27 -0700 Message-ID: <201905221157.A9BAB1F296@keescook> References: <20190517144931.GA56186@arrakis.emea.arm.com> <20190521182932.sm4vxweuwo5ermyd@mbp> <201905211633.6C0BF0C2@keescook> <20190522101110.m2stmpaj7seezveq@mbp> Mime-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: base64 Return-path: Content-Disposition: inline In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: enh Cc: Mark Rutland , kvm@vger.kernel.org, Szabolcs Nagy , Catalin Marinas , Will Deacon , dri-devel@lists.freedesktop.org, Linux Memory Management List , Khalid Aziz , "open list:KERNEL SELFTEST FRAMEWORK" , Vincenzo Frascino , Jacob Bramley , Leon Romanovsky , linux-rdma@vger.kernel.org, amd-gfx@lists.freedesktop.org, Dmitry Vyukov , Dave Martin , Evgenii Stepanov , linux-media@vger.kernel.org, Kevin Brodsky , Ruben Ayrapetyan , Andrey Konovalov , Ramana Radhakrishnan , Alex List-Id: amd-gfx.lists.freedesktop.org T24gV2VkLCBNYXkgMjIsIDIwMTkgYXQgMDg6MzA6MjFBTSAtMDcwMCwgZW5oIHdyb3RlOgo+IE9u IFdlZCwgTWF5IDIyLCAyMDE5IGF0IDM6MTEgQU0gQ2F0YWxpbiBNYXJpbmFzIDxjYXRhbGluLm1h cmluYXNAYXJtLmNvbT4gd3JvdGU6Cj4gPiBPbiBUdWUsIE1heSAyMSwgMjAxOSBhdCAwNTowNDoz OVBNIC0wNzAwLCBLZWVzIENvb2sgd3JvdGU6Cj4gPiA+IEkganVzdCB3YW50IHRvIG1ha2Ugc3Vy ZSBJIGZ1bGx5IHVuZGVyc3RhbmQgeW91ciBjb25jZXJuIGFib3V0IHRoaXMKPiA+ID4gYmVpbmcg YW4gQUJJIGJyZWFrLCBhbmQgSSB3b3JrIGJlc3Qgd2l0aCBleGFtcGxlcy4gVGhlIGNsb3Nlc3Qg c2l0dWF0aW9uCj4gPiA+IEkgY2FuIHNlZSB3b3VsZCBiZToKPiA+ID4KPiA+ID4gLSBzb21lIHBy b2dyYW0gaGFzIG5vIGlkZWEgYWJvdXQgTVRFCj4gPgo+ID4gQXBhcnQgZnJvbSBzb21lIGxpYnJh cmllcyBsaWtlIGxpYmMgKGFuZCBtYXliZSB0aG9zZSB0aGF0IGhhbmRsZQo+ID4gc3BlY2lmaWMg ZGV2aWNlIGlvY3RscyksIEkgdGhpbmsgbW9zdCBwcm9ncmFtcyBzaG91bGQgaGF2ZSBubyBpZGVh IGFib3V0Cj4gPiBNVEUuIEkgd291bGRuJ3QgZXhwZWN0IHByb2dyYW1tZXJzIHRvIGhhdmUgdG8g Y2hhbmdlIHRoZWlyIGFwcCBqdXN0Cj4gPiBiZWNhdXNlIHdlIGhhdmUgYSBuZXcgZmVhdHVyZSB0 aGF0IGNvbG91cnMgaGVhcCBhbGxvY2F0aW9ucy4KClJpZ2h0IC0tIHRoaW5ncyBzaG91bGQgSnVz dCBXb3JrIGZyb20gdGhlIGFwcGxpY2F0aW9uIHBlcnNwZWN0aXZlLgoKPiBvYnZpb3VzbHkgaSdt IGJpYXNlZCBhcyBhIGxpYmMgbWFpbnRhaW5lciwgYnV0Li4uCj4gCj4gaSBkb24ndCB0aGluayBp dCBoZWxwcyB0byBtb3ZlIHRoaXMgdG8gbGliYyAtLS0gbm93IHlvdSBqdXN0IGhhdmUgYW4KPiBl eHRyYSBkZXBlbmRlbmN5IHdoZXJlIHRvIGhhdmUgYSBndWFyYW50ZWVkIHdvcmtpbmcgc3lzdGVt IHlvdSBuZWVkIHRvCj4gdXBkYXRlIHlvdXIga2VybmVsIGFuZCBsaWJjIHRvZ2V0aGVyLiAob3Ig YXQgbGVhc3QgdXBkYXRlIHlvdXIgbGliYyB0bwo+IHVuZGVyc3RhbmQgbmV3IGlvY3RscyBldGMg X2JlZm9yZV8geW91IGNhbiB1cGRhdGUgeW91ciBrZXJuZWwuKQoKSSB0aGluayAoaG9wZT8pIHdl J3ZlIGFsbCBhZ3JlZWQgdGhhdCB3ZSBzaG91bGRuJ3QgcGFzcyB0aGlzIG9mZiB0bwp1c2Vyc3Bh Y2UuIEF0IHRoZSB2ZXJ5IGxlYXN0LCBpdCByZWR1Y2VzIHRoZSB1dGlsaXR5IG9mIE1URSwgYW5k IGF0IHdvcnN0Cml0IGNvbXBsaWNhdGVzIHVzZXJzcGFjZSB3aGVuIHRoaXMgaXMgY2xlYXJseSBh IGtlcm5lbC9hcmNoaXRlY3R1cmUgaXNzdWUuCgo+IAo+ID4gPiAtIG1hbGxvYygpIHN0YXJ0cyBy ZXR1cm5pbmcgTVRFLXRhZ2dlZCBhZGRyZXNzZXMKPiA+ID4gLSBwcm9ncmFtIGRvZXNuJ3QgYnJl YWsgZnJvbSB0aGF0IGNoYW5nZQo+ID4gPiAtIHByb2dyYW0gdXNlcyBzb21lIHN5c2NhbGwgdGhh dCBpcyBtaXNzaW5nIHVudGFnZ2VkX2FkZHIoKSBhbmQgZmFpbHMKPiA+ID4gLSBrZXJuZWwgaGFz IG5vdyBicm9rZW4gdXNlcnNwYWNlIHRoYXQgdXNlZCB0byB3b3JrCj4gPgo+ID4gVGhhdCdzIG9u ZSBhc3BlY3QgdGhvdWdoIHByb2JhYmx5IG1vcmUgb2YgYSBjYXNlIG9mIHBsdWdnaW5nIGluIGEg bmV3Cj4gPiBkZXZpY2UgKGdyYXBoaWNzIGNhcmQsIG5ldHdvcmsgZXRjLikgYW5kIHRoZSBpb2N0 bCB0byB0aGUgbmV3IGRldmljZQo+ID4gZG9lc24ndCB3b3JrLgoKSSB0aGluayBNVEUgd2lsbCBs aWtlbHkgYmUgcmF0aGVyIGxpa2UgTlgvUFhOIGFuZCBTTUFQL1BBTjogdGhlcmUgd2lsbApiZSBn bGl0Y2hlcywgYW5kIHdlIGNhbiBkaXNhYmxlIHN0dWZmIGVpdGhlciB2aWEgQ09ORklHIG9yIChh cyBpcyBtb3JlCmNvbW1vbiBub3cpIHZpYSBhIGtlcm5lbCBjb21tYW5kbGluZSB3aXRoIHVudGFn Z2VkX2FkZHIoKSBjb250YWluaW5nIGEKc3RhdGljIGJyYW5jaCwgZXRjLiBCdXQgSSBhY3R1YWxs eSBkb24ndCB0aGluayB3ZSBuZWVkIHRvIGdvIHRoaXMgcm91dGUKKHNlZSBiZWxvdy4uLikKCj4g PiBUaGUgb3RoZXIgaXMgdGhhdCwgYXNzdW1pbmcgd2UgcmVhY2ggYSBwb2ludCB3aGVyZSB0aGUg a2VybmVsIGVudGlyZWx5Cj4gPiBzdXBwb3J0cyB0aGlzIHJlbGF4ZWQgQUJJLCBjYW4gd2UgZ3Vh cmFudGVlIHRoYXQgaXQgd29uJ3QgYnJlYWsgaW4gdGhlCj4gPiBmdXR1cmUuIExldCdzIHNheSBz b21lIHN1YnNlcXVlbnQga2VybmVsIGNoYW5nZSAoc29tZSByZWZhY3RvcmluZykKPiA+IG1pc3Nl cyBvdXQgYW4gdW50YWdnZWRfYWRkcigpLiBUaGlzIHJlbmRlcnMgYSBwcmV2aW91c2x5IFRCSS9N VEUtY2FwYWJsZQo+ID4gc3lzY2FsbCB1bnVzYWJsZS4gQ2FuIHdlIHJlbHkgb25seSBvbiB0ZXN0 aW5nPwo+ID4KPiA+ID4gVGhlIHRyb3VibGUgSSBzZWUgd2l0aCB0aGlzIGlzIHRoYXQgaXQgaXMg bGFyZ2VseSB0aGVvcmV0aWNhbCBhbmQKPiA+ID4gcmVxdWlyZXMgcGFydCBvZiB1c2Vyc3BhY2Ug dG8gY29sbHVkZSB0byBzdGFydCB1c2luZyBhIG5ldyBDUFUgZmVhdHVyZQo+ID4gPiB0aGF0IHRp Y2tsZXMgYSBidWcgaW4gdGhlIGtlcm5lbC4gQXMgSSB1bmRlcnN0YW5kIHRoZSBnb2xkZW4gcnVs ZSwKPiA+ID4gdGhpcyBpcyBhIGJ1ZyBpbiB0aGUga2VybmVsIChhIG1pc3NlZCBpb2N0bCgpIG9y IHN1Y2gpIHRvIGJlIGZpeGVkLAo+ID4gPiBub3QgYSBnbG9iYWwgYnJlYWtpbmcgb2Ygc29tZSB1 c2Vyc3BhY2UgYmVoYXZpb3IuCj4gPgo+ID4gWWVzLCB3ZSBzaG91bGQgZm9sbG93IHRoZSBydWxl IHRoYXQgaXQncyBhIGtlcm5lbCBidWcgYnV0IGl0IGRvZXNuJ3QKPiA+IGhlbHAgdGhlIHVzZXIg dGhhdCBhIG5ld2x5IGluc3RhbGxlZCBrZXJuZWwgY2F1c2VzIHVzZXIgc3BhY2UgdG8gbm8KPiA+ IGxvbmdlciByZWFjaCBhIHByb21wdC4gSGVuY2UgdGhlIHByb3Bvc2FsIG9mIGFuIG9wdC1pbiB2 aWEgcGVyc29uYWxpdHkKPiA+IChmb3IgTVRFIHdlIHdvdWxkIG5lZWQgYW4gZXhwbGljaXQgb3B0 LWluIGJ5IHRoZSB1c2VyIGFueXdheSBzaW5jZSB0aGUKPiA+IHRvcCBieXRlIGlzIG5vIGxvbmdl ciBpZ25vcmVkIGJ1dCBjaGVja2VkIGFnYWluc3QgdGhlIGFsbG9jYXRpb24gdGFnKS4KPiAKPiBi dXQgcmVhbGlzdGljYWxseSB3b3VsZCB0aGlzIGFjdHVhbGx5IGdldCB1c2VkIGluIHRoaXMgd2F5 PyBvciB3b3VsZAo+IGFueSBnaXZlbiBzeXN0ZW0gZWl0aGVyIGJlIE1URSBvciBub24tTVRFLiBp biB3aGljaCBjYXNlIGEga2VybmVsCj4gY29uZmlndXJhdGlvbiBvcHRpb24gd291bGQgc2VlbSB0 byBtYWtlIG1vcmUgc2Vuc2UuIChiZWNhdXNlIGVpdGhlcgo+IHdheSwgdGhlIGh5cG90aGV0aWNh bCB1c2VyIGJhc2ljYWxseSBuZWVkcyB0byByZWNvbXBpbGUgdGhlIGtlcm5lbCB0bwo+IGdldCBi YWNrIG9uIHRoZWlyIGZlZXQuIG9yIGFsbCBvZiB1c2Vyc3BhY2UuKQoKUmlnaHQ6IHRoZSBwb2lu dCBpcyB0byBkZXNpZ24gdGhpbmdzIHNvIHRoYXQgd2UgZG8gb3VyIGJlc3QgdG8gbm90IGJyZWFr CnVzZXJzcGFjZSB0aGF0IGlzIHVzaW5nIHRoZSBuZXcgZmVhdHVyZSAod2hpY2ggSSB0aGluayB0 aGlzIHNlcmllcyBoYXMKZG9uZSB3ZWxsKS4gQnV0IHN1cHBvcnRpbmcgTVRFL1RCSSBpcyBqdXN0 IGxpa2Ugc3VwcG9ydGluZyBQQU46IGlmIHNvbWVvbmUKcmVmYWN0b3JzIGEgZHJpdmVyIGFuZCBz d2FwcyBhIGNvcHlfZnJvbV91c2VyKCkgdG8gYSBtZW1jcHkoKSwgaXQncyBnb2luZwp0byBicmVh ayB1bmRlciBQQU4uIFRoZXJlIHdpbGwgYmUgdGhlIHNhbWUgbG9uZyB0YWlsIG9mIHRoZXNlIGJ1 Z3MgbGlrZQphbnkgb3RoZXIsIGJ1dCBteSBzZW5zZSBpcyB0aGF0IHRoZXkgYXJlIHNtYWxsIGFu ZCByYXJlLiBCdXQgSSBhZ3JlZToKdGhleSdyZSBnb2luZyB0byBiZSBwcmV0dHkgd2VpcmQgYnVn cyB0byB0cmFjayBkb3duLiBUaGUgZmluYWwgcmVzdWx0LApob3dldmVyLCB3aWxsIGJlIGV4Y2Vs bGVudCBhbm5vdGF0aW9uIGluIHRoZSBrZXJuZWwgZm9yIHdoZXJlIHVzZXJzcGFjZQphZGRyZXNz ZXMgZ2V0IHVzZWQgYW5kIHBlb3BsZSBtYWtlIGFzc3VtcHRpb25zIGFib3V0IHRoZW0uCgpUaGUg c29vbmVyIHdlIGdldCB0aGUgc2VyaWVzIGxhbmRlZCBhbmQgZ2FpbiBRRU1VIHN1cHBvcnQgKG9y IHJlYWwKaGFyZHdhcmUpLCB0aGUgZmFzdGVyIHdlIGNhbiBoYW1tZXIgb3V0IHRoZXNlIG1pc3Nl ZCBjb3JuZXItY2FzZXMuCldoYXQncyB0aGUgdGltZWxpbmUgZm9yIGVpdGhlciBvZiB0aG9zZSB0 aGluZ3MsIEJUVz8KCj4gPiA+IEkgZmVlbCBsaWtlIEknbSBtaXNzaW5nIHNvbWV0aGluZyBhYm91 dCB0aGlzIGJlaW5nIHNlZW4gYXMgYW4gQUJJCj4gPiA+IGJyZWFrLiBUaGUga2VybmVsIGFscmVh ZHkgZmFpbHMgb24gdXNlcnNwYWNlIGFkZHJlc3NlcyB0aGF0IGhhdmUgaGlnaAo+ID4gPiBiaXRz IHNldCAtLSBhcmUgdGhlcmUgdGhpbmdzIHRoYXQgX2RlcGVuZF8gb24gdGhpcyBmYWlsdXJlIHRv IG9wZXJhdGU/Cj4gPgo+ID4gSXQncyBhYm91dCBwcm92aWRpbmcgYSByZWxheGVkIEFCSSB3aGlj aCBhbGxvd3Mgbm9uLXplcm8gdG9wIGJ5dGUgYW5kCj4gPiBicmVha2luZyBpdCBsYXRlciBpbmFk dmVydGVudGx5IHdpdGhvdXQgaGF2aW5nIHNvbWV0aGluZyBiZXR0ZXIgaW4gcGxhY2UKPiA+IHRv IGFuYWx5c2UgdGhlIGtlcm5lbCBjaGFuZ2VzLgoKSXQgc291bmRzIGxpa2UgdGhlIHF1ZXN0aW9u IGlzIGhvdyB0byBzd2l0Y2ggYSBwcm9jZXNzIGluIG9yIG91dCBvZiB0aGlzCkFCSSAoYnV0IEkg ZG9uJ3QgdGhpbmsgdGhhdCdzIHRoZSByZWFsIGlzc3VlOiBJIHRoaW5rIGl0J3MganVzdCBhIG1h dHRlcgpvZiB3aGV0aGVyIG9yIG5vdCBhIHByb2Nlc3MgdXNlcyB0YWdzIGF0IGFsbCkuIERvaW5n IGl0IGF0IHRoZSBwcmN0bCgpCmxldmVsIGRvZXNuJ3QgbWFrZSBzZW5zZSB0byBtZSwgZXhjZXB0 IG1heWJlIHRvIGRldGVjdCBNVEUgc3VwcG9ydCBvcgpzb21ldGhpbmcuICgiU2hvdWxkIEkgdGFn IGFsbG9jYXRpb25zPyIpIEFuZCB0aGF0IHN0YXRlIGlzIGNvbnRyb2xsZWQKYnkgdGhlIGtlcm5l bDogdGhlIGtlcm5lbCBkb2VzIGl0IG9yIGl0IGRvZXNuJ3QuCgpJZiBhIHByb2Nlc3Mgd2FudHMg dG8gbm90IHRhZywgdGhhdCdzIGFsc28gdXAgdG8gdGhlIGFsbG9jYXRvciB3aGVyZQppdCBjYW4g ZGVjaWRlIG5vdCB0byBhc2sgdGhlIGtlcm5lbCwgYW5kIGp1c3Qgbm90IHRhZy4gTm90aGluZyBi cmVha3MgaW4KdXNlcnNwYWNlIGlmIGEgcHJvY2VzcyBpcyBOT1QgdGFnZ2luZyBhbmQgdW50YWdn ZWRfYWRkcigpIGV4aXN0cyBvciBpcwptaXNzaW5nLiBUaGlzLCBJIHRoaW5rLCBpcyB0aGUgY29y ZSB3YXkgdGhpcyBkb2Vzbid0IHRyaXAgb3ZlciB0aGUKZ29sZGVuIHJ1bGU6IGFuIG9sZCBzeXN0 ZW0gaW1hZ2Ugd2lsbCBydW4gZmluZSAoYmVjYXVzZSBpdCdzIG5vdAp0YWdnaW5nKS4gQSAqbmV3 KiBzeXN0ZW0gbWF5IGVuY291bnRlciBidWdzIHdpdGggdGFnZ2luZyBiZWNhdXNlIGl0J3MgYQpu ZXcgZmVhdHVyZTogdGhpcyBpcyBUaGUgV2F5IE9mIFRoaW5ncy4gQnV0IHdlIGRvbid0IGJyZWFr IG9sZCB1c2Vyc3BhY2UKYmVjYXVzZSBvbGQgdXNlcnNwYWNlIGlzbid0IHVzaW5nIHRhZ3MuCgpT byB0aGUgYWdyZWVtZW50IGFwcGVhcnMgdG8gYmUgYmV0d2VlbiB0aGUga2VybmVsIGFuZCB0aGUg YWxsb2NhdG9yLgpLZXJuZWwgc2F5cyAiSSBzdXBwb3J0IHRoaXMiIG9yIG5vdC4gVGVsbGluZyB0 aGUgYWxsb2NhdG9yIHRvIG5vdCB0YWcgaWYKc29tZXRoaW5nIGJyZWFrcyBzb3VuZHMgbGlrZSBh biBlbnRpcmVseSB1c2Vyc3BhY2UgZGVjaXNpb24sIHllcz8KCi0tIApLZWVzIENvb2sKX19fX19f X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVsIG1haWxp bmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlzdHMuZnJl ZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVs From mboxrd@z Thu Jan 1 00:00:00 1970 From: keescook at chromium.org (Kees Cook) Date: Wed, 22 May 2019 12:21:27 -0700 Subject: [PATCH v15 00/17] arm64: untag user pointers passed to the kernel In-Reply-To: References: <20190517144931.GA56186@arrakis.emea.arm.com> <20190521182932.sm4vxweuwo5ermyd@mbp> <201905211633.6C0BF0C2@keescook> <20190522101110.m2stmpaj7seezveq@mbp> Message-ID: <201905221157.A9BAB1F296@keescook> On Wed, May 22, 2019 at 08:30:21AM -0700, enh wrote: > On Wed, May 22, 2019 at 3:11 AM Catalin Marinas wrote: > > On Tue, May 21, 2019 at 05:04:39PM -0700, Kees Cook wrote: > > > I just want to make sure I fully understand your concern about this > > > being an ABI break, and I work best with examples. The closest situation > > > I can see would be: > > > > > > - some program has no idea about MTE > > > > Apart from some libraries like libc (and maybe those that handle > > specific device ioctls), I think most programs should have no idea about > > MTE. I wouldn't expect programmers to have to change their app just > > because we have a new feature that colours heap allocations. Right -- things should Just Work from the application perspective. > obviously i'm biased as a libc maintainer, but... > > i don't think it helps to move this to libc --- now you just have an > extra dependency where to have a guaranteed working system you need to > update your kernel and libc together. (or at least update your libc to > understand new ioctls etc _before_ you can update your kernel.) I think (hope?) we've all agreed that we shouldn't pass this off to userspace. At the very least, it reduces the utility of MTE, and at worst it complicates userspace when this is clearly a kernel/architecture issue. > > > > - malloc() starts returning MTE-tagged addresses > > > - program doesn't break from that change > > > - program uses some syscall that is missing untagged_addr() and fails > > > - kernel has now broken userspace that used to work > > > > That's one aspect though probably more of a case of plugging in a new > > device (graphics card, network etc.) and the ioctl to the new device > > doesn't work. I think MTE will likely be rather like NX/PXN and SMAP/PAN: there will be glitches, and we can disable stuff either via CONFIG or (as is more common now) via a kernel commandline with untagged_addr() containing a static branch, etc. But I actually don't think we need to go this route (see below...) > > The other is that, assuming we reach a point where the kernel entirely > > supports this relaxed ABI, can we guarantee that it won't break in the > > future. Let's say some subsequent kernel change (some refactoring) > > misses out an untagged_addr(). This renders a previously TBI/MTE-capable > > syscall unusable. Can we rely only on testing? > > > > > The trouble I see with this is that it is largely theoretical and > > > requires part of userspace to collude to start using a new CPU feature > > > that tickles a bug in the kernel. As I understand the golden rule, > > > this is a bug in the kernel (a missed ioctl() or such) to be fixed, > > > not a global breaking of some userspace behavior. > > > > Yes, we should follow the rule that it's a kernel bug but it doesn't > > help the user that a newly installed kernel causes user space to no > > longer reach a prompt. Hence the proposal of an opt-in via personality > > (for MTE we would need an explicit opt-in by the user anyway since the > > top byte is no longer ignored but checked against the allocation tag). > > but realistically would this actually get used in this way? or would > any given system either be MTE or non-MTE. in which case a kernel > configuration option would seem to make more sense. (because either > way, the hypothetical user basically needs to recompile the kernel to > get back on their feet. or all of userspace.) Right: the point is to design things so that we do our best to not break userspace that is using the new feature (which I think this series has done well). But supporting MTE/TBI is just like supporting PAN: if someone refactors a driver and swaps a copy_from_user() to a memcpy(), it's going to break under PAN. There will be the same long tail of these bugs like any other, but my sense is that they are small and rare. But I agree: they're going to be pretty weird bugs to track down. The final result, however, will be excellent annotation in the kernel for where userspace addresses get used and people make assumptions about them. The sooner we get the series landed and gain QEMU support (or real hardware), the faster we can hammer out these missed corner-cases. What's the timeline for either of those things, BTW? > > > I feel like I'm missing something about this being seen as an ABI > > > break. The kernel already fails on userspace addresses that have high > > > bits set -- are there things that _depend_ on this failure to operate? > > > > It's about providing a relaxed ABI which allows non-zero top byte and > > breaking it later inadvertently without having something better in place > > to analyse the kernel changes. It sounds like the question is how to switch a process in or out of this ABI (but I don't think that's the real issue: I think it's just a matter of whether or not a process uses tags at all). Doing it at the prctl() level doesn't make sense to me, except maybe to detect MTE support or something. ("Should I tag allocations?") And that state is controlled by the kernel: the kernel does it or it doesn't. If a process wants to not tag, that's also up to the allocator where it can decide not to ask the kernel, and just not tag. Nothing breaks in userspace if a process is NOT tagging and untagged_addr() exists or is missing. This, I think, is the core way this doesn't trip over the golden rule: an old system image will run fine (because it's not tagging). A *new* system may encounter bugs with tagging because it's a new feature: this is The Way Of Things. But we don't break old userspace because old userspace isn't using tags. So the agreement appears to be between the kernel and the allocator. Kernel says "I support this" or not. Telling the allocator to not tag if something breaks sounds like an entirely userspace decision, yes? -- Kees Cook From mboxrd@z Thu Jan 1 00:00:00 1970 From: keescook@chromium.org (Kees Cook) Date: Wed, 22 May 2019 12:21:27 -0700 Subject: [PATCH v15 00/17] arm64: untag user pointers passed to the kernel In-Reply-To: References: <20190517144931.GA56186@arrakis.emea.arm.com> <20190521182932.sm4vxweuwo5ermyd@mbp> <201905211633.6C0BF0C2@keescook> <20190522101110.m2stmpaj7seezveq@mbp> Message-ID: <201905221157.A9BAB1F296@keescook> Content-Type: text/plain; charset="UTF-8" Message-ID: <20190522192127.nqfGMlLE34WljQA4QrEr1iV3jC1SrWwfXbW_2Jjse-c@z> On Wed, May 22, 2019@08:30:21AM -0700, enh wrote: > On Wed, May 22, 2019@3:11 AM Catalin Marinas wrote: > > On Tue, May 21, 2019@05:04:39PM -0700, Kees Cook wrote: > > > I just want to make sure I fully understand your concern about this > > > being an ABI break, and I work best with examples. The closest situation > > > I can see would be: > > > > > > - some program has no idea about MTE > > > > Apart from some libraries like libc (and maybe those that handle > > specific device ioctls), I think most programs should have no idea about > > MTE. I wouldn't expect programmers to have to change their app just > > because we have a new feature that colours heap allocations. Right -- things should Just Work from the application perspective. > obviously i'm biased as a libc maintainer, but... > > i don't think it helps to move this to libc --- now you just have an > extra dependency where to have a guaranteed working system you need to > update your kernel and libc together. (or at least update your libc to > understand new ioctls etc _before_ you can update your kernel.) I think (hope?) we've all agreed that we shouldn't pass this off to userspace. At the very least, it reduces the utility of MTE, and at worst it complicates userspace when this is clearly a kernel/architecture issue. > > > > - malloc() starts returning MTE-tagged addresses > > > - program doesn't break from that change > > > - program uses some syscall that is missing untagged_addr() and fails > > > - kernel has now broken userspace that used to work > > > > That's one aspect though probably more of a case of plugging in a new > > device (graphics card, network etc.) and the ioctl to the new device > > doesn't work. I think MTE will likely be rather like NX/PXN and SMAP/PAN: there will be glitches, and we can disable stuff either via CONFIG or (as is more common now) via a kernel commandline with untagged_addr() containing a static branch, etc. But I actually don't think we need to go this route (see below...) > > The other is that, assuming we reach a point where the kernel entirely > > supports this relaxed ABI, can we guarantee that it won't break in the > > future. Let's say some subsequent kernel change (some refactoring) > > misses out an untagged_addr(). This renders a previously TBI/MTE-capable > > syscall unusable. Can we rely only on testing? > > > > > The trouble I see with this is that it is largely theoretical and > > > requires part of userspace to collude to start using a new CPU feature > > > that tickles a bug in the kernel. As I understand the golden rule, > > > this is a bug in the kernel (a missed ioctl() or such) to be fixed, > > > not a global breaking of some userspace behavior. > > > > Yes, we should follow the rule that it's a kernel bug but it doesn't > > help the user that a newly installed kernel causes user space to no > > longer reach a prompt. Hence the proposal of an opt-in via personality > > (for MTE we would need an explicit opt-in by the user anyway since the > > top byte is no longer ignored but checked against the allocation tag). > > but realistically would this actually get used in this way? or would > any given system either be MTE or non-MTE. in which case a kernel > configuration option would seem to make more sense. (because either > way, the hypothetical user basically needs to recompile the kernel to > get back on their feet. or all of userspace.) Right: the point is to design things so that we do our best to not break userspace that is using the new feature (which I think this series has done well). But supporting MTE/TBI is just like supporting PAN: if someone refactors a driver and swaps a copy_from_user() to a memcpy(), it's going to break under PAN. There will be the same long tail of these bugs like any other, but my sense is that they are small and rare. But I agree: they're going to be pretty weird bugs to track down. The final result, however, will be excellent annotation in the kernel for where userspace addresses get used and people make assumptions about them. The sooner we get the series landed and gain QEMU support (or real hardware), the faster we can hammer out these missed corner-cases. What's the timeline for either of those things, BTW? > > > I feel like I'm missing something about this being seen as an ABI > > > break. The kernel already fails on userspace addresses that have high > > > bits set -- are there things that _depend_ on this failure to operate? > > > > It's about providing a relaxed ABI which allows non-zero top byte and > > breaking it later inadvertently without having something better in place > > to analyse the kernel changes. It sounds like the question is how to switch a process in or out of this ABI (but I don't think that's the real issue: I think it's just a matter of whether or not a process uses tags at all). Doing it at the prctl() level doesn't make sense to me, except maybe to detect MTE support or something. ("Should I tag allocations?") And that state is controlled by the kernel: the kernel does it or it doesn't. If a process wants to not tag, that's also up to the allocator where it can decide not to ask the kernel, and just not tag. Nothing breaks in userspace if a process is NOT tagging and untagged_addr() exists or is missing. This, I think, is the core way this doesn't trip over the golden rule: an old system image will run fine (because it's not tagging). A *new* system may encounter bugs with tagging because it's a new feature: this is The Way Of Things. But we don't break old userspace because old userspace isn't using tags. So the agreement appears to be between the kernel and the allocator. Kernel says "I support this" or not. Telling the allocator to not tag if something breaks sounds like an entirely userspace decision, yes? -- Kees Cook From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.0 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI,SPF_HELO_NONE, SPF_PASS,URIBL_BLOCKED autolearn=unavailable autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 023B6C282DD for ; Wed, 22 May 2019 19:21:46 +0000 (UTC) Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by mail.kernel.org (Postfix) with ESMTPS id C0CD421473 for ; Wed, 22 May 2019 19:21:45 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=lists.infradead.org header.i=@lists.infradead.org header.b="r+htNUMv"; dkim=fail reason="signature verification failed" (1024-bit key) header.d=chromium.org header.i=@chromium.org header.b="GFshwxVj" DMARC-Filter: OpenDMARC Filter v1.3.2 mail.kernel.org C0CD421473 Authentication-Results: mail.kernel.org; dmarc=fail (p=none dis=none) header.from=chromium.org Authentication-Results: mail.kernel.org; spf=none smtp.mailfrom=linux-arm-kernel-bounces+infradead-linux-arm-kernel=archiver.kernel.org@lists.infradead.org DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20170209; h=Sender: Content-Transfer-Encoding:Content-Type:Cc:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=PR5ZY/VR/B1Q978HpbkjO2+iTDF2FSjF12nKu4UlvLE=; b=r+htNUMvWucK8i J98UqdD6sa85btS5ZeGpBkKFDy3NhBrHMDdFFb9m+6AJ1riLMq69B3ygO4L9jbMlVj9IhhmUUbiSE wVx6LeXpg9xmbua2sw+BfNGBhF9OXrxXfYgxp5k5G/AnmnaoWxMozPbDFZfMicV9q0z148u9xbdMI +MKYR1yUQ6/Y5E+2OAQLGZ6H/Om+GNuRA4aG2ttWgoW/FXua4YRl0gHjh4KYrgYxrpAk/Y59UzJqf 5pkyaDiRhnwnHpwV41IcGTfEC0Ro+wmfg/WWyv9wukUx8ymjeNCpCQHEGORA6OKyl0dc7hIYlb5pz n+aX/4wWWtaDG2GFsB7g==; Received: from localhost ([127.0.0.1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.90_1 #2 (Red Hat Linux)) id 1hTWnt-0001H3-S9; Wed, 22 May 2019 19:21:33 +0000 Received: from mail-pg1-x544.google.com ([2607:f8b0:4864:20::544]) by bombadil.infradead.org with esmtps (Exim 4.90_1 #2 (Red Hat Linux)) id 1hTWnq-0001Fv-Pb for linux-arm-kernel@lists.infradead.org; Wed, 22 May 2019 19:21:32 +0000 Received: by mail-pg1-x544.google.com with SMTP id f25so1789795pgv.10 for ; Wed, 22 May 2019 12:21:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=chromium.org; s=google; h=date:from:to:cc:subject:message-id:references:mime-version :content-disposition:in-reply-to; bh=NP4JWkHqbCIs5rs8MQbZ/rJv31PHu50+VP85AcMzbbU=; b=GFshwxVjgFyTBuVsTzhm7IUt1UjnhWPD0MinKNguPB3el83BFNud8bHr5plz3emii+ cTc7Cjy2KfiVySMa4vl3S45MUDV6ughBv+eP5o9C9KIg6ntk97RzOYluREumaUsuEfOH SjJgP3hh0ohPvQPPbSdKkfRYeqkQoVWPqO5e4= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:from:to:cc:subject:message-id:references :mime-version:content-disposition:in-reply-to; bh=NP4JWkHqbCIs5rs8MQbZ/rJv31PHu50+VP85AcMzbbU=; b=eSxOX9+VFZ/VljfAM/swo9eE4jw/XXKGZpmPdQz1qyAyfw/XjdowmA+xb1jfv+i/R/ jMtPiROlYvl6ED9GoJx5FY++NXFC9WSiqOhR3UdV5nrh0Q5kToTvAmOr7GlRwlWwRqd+ EqGi/e/KEC/d915HmM60Uvip0SIhDM90PYkwqiNVSaEj61I3h2lB13RxSRFXIIv6HSaR GimbfKjXJy0Kc1GZkaJbLjxXUIMdgWD1fDoWyBGSWJmAJQ33NVh51EkO5ZZD65MDqGyo rVw8J3ZZFDKMlWopX6wbwTK+Ua64uZ8iO7hFFdvsuvkL5eXCabOI+rCc0ybyZF1RzEOt fbLQ== X-Gm-Message-State: APjAAAVlVX6YgXn0vNG50jzuHBDObZYT7Srja2g9Io0zGas79b6ftvm6 j4yPdF/eiptTPhPD3d8jhwojhw== X-Google-Smtp-Source: APXvYqzNjksdiGtBA2K0Yn7lD7hXz0a3WOc+KaCHe5Lfb///NUHl50P+zldmTZ4Tu8MFGQvOw+1qnw== X-Received: by 2002:a62:6456:: with SMTP id y83mr32990581pfb.71.1558552889913; Wed, 22 May 2019 12:21:29 -0700 (PDT) Received: from www.outflux.net (smtp.outflux.net. [198.145.64.163]) by smtp.gmail.com with ESMTPSA id x10sm37135797pfj.136.2019.05.22.12.21.28 (version=TLS1_2 cipher=ECDHE-RSA-CHACHA20-POLY1305 bits=256/256); Wed, 22 May 2019 12:21:28 -0700 (PDT) Date: Wed, 22 May 2019 12:21:27 -0700 From: Kees Cook To: enh Subject: Re: [PATCH v15 00/17] arm64: untag user pointers passed to the kernel Message-ID: <201905221157.A9BAB1F296@keescook> References: <20190517144931.GA56186@arrakis.emea.arm.com> <20190521182932.sm4vxweuwo5ermyd@mbp> <201905211633.6C0BF0C2@keescook> <20190522101110.m2stmpaj7seezveq@mbp> MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20190522_122130_851747_35F04525 X-CRM114-Status: GOOD ( 39.90 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.21 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Cc: Mark Rutland , kvm@vger.kernel.org, Szabolcs Nagy , Catalin Marinas , Will Deacon , dri-devel@lists.freedesktop.org, Linux Memory Management List , Khalid Aziz , "open list:KERNEL SELFTEST FRAMEWORK" , Vincenzo Frascino , Jacob Bramley , Leon Romanovsky , linux-rdma@vger.kernel.org, amd-gfx@lists.freedesktop.org, Dmitry Vyukov , Dave Martin , Evgenii Stepanov , linux-media@vger.kernel.org, Kevin Brodsky , Ruben Ayrapetyan , Andrey Konovalov , Ramana Radhakrishnan , Alex Williamson , Yishai Hadas , Mauro Carvalho Chehab , Linux ARM , Kostya Serebryany , Greg Kroah-Hartman , Felix Kuehling , LKML , Jens Wiklander , Lee Smith , Alexander Deucher , Andrew Morton , Robin Murphy , Christian Koenig , Luc Van Oostenryck Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+infradead-linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Wed, May 22, 2019 at 08:30:21AM -0700, enh wrote: > On Wed, May 22, 2019 at 3:11 AM Catalin Marinas wrote: > > On Tue, May 21, 2019 at 05:04:39PM -0700, Kees Cook wrote: > > > I just want to make sure I fully understand your concern about this > > > being an ABI break, and I work best with examples. The closest situation > > > I can see would be: > > > > > > - some program has no idea about MTE > > > > Apart from some libraries like libc (and maybe those that handle > > specific device ioctls), I think most programs should have no idea about > > MTE. I wouldn't expect programmers to have to change their app just > > because we have a new feature that colours heap allocations. Right -- things should Just Work from the application perspective. > obviously i'm biased as a libc maintainer, but... > > i don't think it helps to move this to libc --- now you just have an > extra dependency where to have a guaranteed working system you need to > update your kernel and libc together. (or at least update your libc to > understand new ioctls etc _before_ you can update your kernel.) I think (hope?) we've all agreed that we shouldn't pass this off to userspace. At the very least, it reduces the utility of MTE, and at worst it complicates userspace when this is clearly a kernel/architecture issue. > > > > - malloc() starts returning MTE-tagged addresses > > > - program doesn't break from that change > > > - program uses some syscall that is missing untagged_addr() and fails > > > - kernel has now broken userspace that used to work > > > > That's one aspect though probably more of a case of plugging in a new > > device (graphics card, network etc.) and the ioctl to the new device > > doesn't work. I think MTE will likely be rather like NX/PXN and SMAP/PAN: there will be glitches, and we can disable stuff either via CONFIG or (as is more common now) via a kernel commandline with untagged_addr() containing a static branch, etc. But I actually don't think we need to go this route (see below...) > > The other is that, assuming we reach a point where the kernel entirely > > supports this relaxed ABI, can we guarantee that it won't break in the > > future. Let's say some subsequent kernel change (some refactoring) > > misses out an untagged_addr(). This renders a previously TBI/MTE-capable > > syscall unusable. Can we rely only on testing? > > > > > The trouble I see with this is that it is largely theoretical and > > > requires part of userspace to collude to start using a new CPU feature > > > that tickles a bug in the kernel. As I understand the golden rule, > > > this is a bug in the kernel (a missed ioctl() or such) to be fixed, > > > not a global breaking of some userspace behavior. > > > > Yes, we should follow the rule that it's a kernel bug but it doesn't > > help the user that a newly installed kernel causes user space to no > > longer reach a prompt. Hence the proposal of an opt-in via personality > > (for MTE we would need an explicit opt-in by the user anyway since the > > top byte is no longer ignored but checked against the allocation tag). > > but realistically would this actually get used in this way? or would > any given system either be MTE or non-MTE. in which case a kernel > configuration option would seem to make more sense. (because either > way, the hypothetical user basically needs to recompile the kernel to > get back on their feet. or all of userspace.) Right: the point is to design things so that we do our best to not break userspace that is using the new feature (which I think this series has done well). But supporting MTE/TBI is just like supporting PAN: if someone refactors a driver and swaps a copy_from_user() to a memcpy(), it's going to break under PAN. There will be the same long tail of these bugs like any other, but my sense is that they are small and rare. But I agree: they're going to be pretty weird bugs to track down. The final result, however, will be excellent annotation in the kernel for where userspace addresses get used and people make assumptions about them. The sooner we get the series landed and gain QEMU support (or real hardware), the faster we can hammer out these missed corner-cases. What's the timeline for either of those things, BTW? > > > I feel like I'm missing something about this being seen as an ABI > > > break. The kernel already fails on userspace addresses that have high > > > bits set -- are there things that _depend_ on this failure to operate? > > > > It's about providing a relaxed ABI which allows non-zero top byte and > > breaking it later inadvertently without having something better in place > > to analyse the kernel changes. It sounds like the question is how to switch a process in or out of this ABI (but I don't think that's the real issue: I think it's just a matter of whether or not a process uses tags at all). Doing it at the prctl() level doesn't make sense to me, except maybe to detect MTE support or something. ("Should I tag allocations?") And that state is controlled by the kernel: the kernel does it or it doesn't. If a process wants to not tag, that's also up to the allocator where it can decide not to ask the kernel, and just not tag. Nothing breaks in userspace if a process is NOT tagging and untagged_addr() exists or is missing. This, I think, is the core way this doesn't trip over the golden rule: an old system image will run fine (because it's not tagging). A *new* system may encounter bugs with tagging because it's a new feature: this is The Way Of Things. But we don't break old userspace because old userspace isn't using tags. So the agreement appears to be between the kernel and the allocator. Kernel says "I support this" or not. Telling the allocator to not tag if something breaks sounds like an entirely userspace decision, yes? -- Kees Cook _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-1.1 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_HELO_NONE,SPF_PASS,URIBL_BLOCKED autolearn=unavailable autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 41FE5C282CE for ; Wed, 22 May 2019 20:04:05 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 05FD920863 for ; Wed, 22 May 2019 20:04:05 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (1024-bit key) header.d=chromium.org header.i=@chromium.org header.b="GFshwxVj" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1729962AbfEVTVb (ORCPT ); Wed, 22 May 2019 15:21:31 -0400 Received: from mail-pg1-f196.google.com ([209.85.215.196]:35776 "EHLO mail-pg1-f196.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1729933AbfEVTVa (ORCPT ); Wed, 22 May 2019 15:21:30 -0400 Received: by mail-pg1-f196.google.com with SMTP id t1so1809422pgc.2 for ; Wed, 22 May 2019 12:21:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=chromium.org; s=google; h=date:from:to:cc:subject:message-id:references:mime-version :content-disposition:in-reply-to; bh=NP4JWkHqbCIs5rs8MQbZ/rJv31PHu50+VP85AcMzbbU=; b=GFshwxVjgFyTBuVsTzhm7IUt1UjnhWPD0MinKNguPB3el83BFNud8bHr5plz3emii+ cTc7Cjy2KfiVySMa4vl3S45MUDV6ughBv+eP5o9C9KIg6ntk97RzOYluREumaUsuEfOH SjJgP3hh0ohPvQPPbSdKkfRYeqkQoVWPqO5e4= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:from:to:cc:subject:message-id:references :mime-version:content-disposition:in-reply-to; bh=NP4JWkHqbCIs5rs8MQbZ/rJv31PHu50+VP85AcMzbbU=; b=CgXmUmz2Jt2nVqJolFMKn0TfqQnv/iuWsa6wFBsfQWnHwkJyZBxTtdFI5qoX9FP6h/ ugrlJNwm6YILZ/SSk3OPnk2wmz2sEEucWADS30pfKWjfKJoZ/fl76/YK3NUOQTKsYZmZ uQx26aoQRcjF1Rn/+JmwyhUlsS2e0gCTa3lxzGl7voeb5qTkjmCWBYv7QpH6y4nS+y4f ExGQLKcGxSBac+gHvDkOqW9Fzo4xYzTtrDxJlr8H3m7SnQIwWI3hH5JTcoowrKww9FuQ sKQRISzvtH461f8LyoloMHSXwMtz1i31vWjo5KT64/XZ3pI6U4PggJv6TJk1MdPu1MeE 3Luw== X-Gm-Message-State: APjAAAUTRn68MLfcngQDbUBDeM3rJDKRqUQQ7WSTTkQ3Jhan2A1TajkW a9pjtz0BbW3JdVfyyEaZ6rs+bA== X-Google-Smtp-Source: APXvYqzNjksdiGtBA2K0Yn7lD7hXz0a3WOc+KaCHe5Lfb///NUHl50P+zldmTZ4Tu8MFGQvOw+1qnw== X-Received: by 2002:a62:6456:: with SMTP id y83mr32990581pfb.71.1558552889913; Wed, 22 May 2019 12:21:29 -0700 (PDT) Received: from www.outflux.net (smtp.outflux.net. [198.145.64.163]) by smtp.gmail.com with ESMTPSA id x10sm37135797pfj.136.2019.05.22.12.21.28 (version=TLS1_2 cipher=ECDHE-RSA-CHACHA20-POLY1305 bits=256/256); Wed, 22 May 2019 12:21:28 -0700 (PDT) Date: Wed, 22 May 2019 12:21:27 -0700 From: Kees Cook To: enh Cc: Catalin Marinas , Evgenii Stepanov , Andrey Konovalov , Khalid Aziz , Linux ARM , Linux Memory Management List , LKML , amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-rdma@vger.kernel.org, linux-media@vger.kernel.org, kvm@vger.kernel.org, "open list:KERNEL SELFTEST FRAMEWORK" , Vincenzo Frascino , Will Deacon , Mark Rutland , Andrew Morton , Greg Kroah-Hartman , Yishai Hadas , Felix Kuehling , Alexander Deucher , Christian Koenig , Mauro Carvalho Chehab , Jens Wiklander , Alex Williamson , Leon Romanovsky , Dmitry Vyukov , Kostya Serebryany , Lee Smith , Ramana Radhakrishnan , Jacob Bramley , Ruben Ayrapetyan , Robin Murphy , Luc Van Oostenryck , Dave Martin , Kevin Brodsky , Szabolcs Nagy Subject: Re: [PATCH v15 00/17] arm64: untag user pointers passed to the kernel Message-ID: <201905221157.A9BAB1F296@keescook> References: <20190517144931.GA56186@arrakis.emea.arm.com> <20190521182932.sm4vxweuwo5ermyd@mbp> <201905211633.6C0BF0C2@keescook> <20190522101110.m2stmpaj7seezveq@mbp> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: Sender: kvm-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: kvm@vger.kernel.org On Wed, May 22, 2019 at 08:30:21AM -0700, enh wrote: > On Wed, May 22, 2019 at 3:11 AM Catalin Marinas wrote: > > On Tue, May 21, 2019 at 05:04:39PM -0700, Kees Cook wrote: > > > I just want to make sure I fully understand your concern about this > > > being an ABI break, and I work best with examples. The closest situation > > > I can see would be: > > > > > > - some program has no idea about MTE > > > > Apart from some libraries like libc (and maybe those that handle > > specific device ioctls), I think most programs should have no idea about > > MTE. I wouldn't expect programmers to have to change their app just > > because we have a new feature that colours heap allocations. Right -- things should Just Work from the application perspective. > obviously i'm biased as a libc maintainer, but... > > i don't think it helps to move this to libc --- now you just have an > extra dependency where to have a guaranteed working system you need to > update your kernel and libc together. (or at least update your libc to > understand new ioctls etc _before_ you can update your kernel.) I think (hope?) we've all agreed that we shouldn't pass this off to userspace. At the very least, it reduces the utility of MTE, and at worst it complicates userspace when this is clearly a kernel/architecture issue. > > > > - malloc() starts returning MTE-tagged addresses > > > - program doesn't break from that change > > > - program uses some syscall that is missing untagged_addr() and fails > > > - kernel has now broken userspace that used to work > > > > That's one aspect though probably more of a case of plugging in a new > > device (graphics card, network etc.) and the ioctl to the new device > > doesn't work. I think MTE will likely be rather like NX/PXN and SMAP/PAN: there will be glitches, and we can disable stuff either via CONFIG or (as is more common now) via a kernel commandline with untagged_addr() containing a static branch, etc. But I actually don't think we need to go this route (see below...) > > The other is that, assuming we reach a point where the kernel entirely > > supports this relaxed ABI, can we guarantee that it won't break in the > > future. Let's say some subsequent kernel change (some refactoring) > > misses out an untagged_addr(). This renders a previously TBI/MTE-capable > > syscall unusable. Can we rely only on testing? > > > > > The trouble I see with this is that it is largely theoretical and > > > requires part of userspace to collude to start using a new CPU feature > > > that tickles a bug in the kernel. As I understand the golden rule, > > > this is a bug in the kernel (a missed ioctl() or such) to be fixed, > > > not a global breaking of some userspace behavior. > > > > Yes, we should follow the rule that it's a kernel bug but it doesn't > > help the user that a newly installed kernel causes user space to no > > longer reach a prompt. Hence the proposal of an opt-in via personality > > (for MTE we would need an explicit opt-in by the user anyway since the > > top byte is no longer ignored but checked against the allocation tag). > > but realistically would this actually get used in this way? or would > any given system either be MTE or non-MTE. in which case a kernel > configuration option would seem to make more sense. (because either > way, the hypothetical user basically needs to recompile the kernel to > get back on their feet. or all of userspace.) Right: the point is to design things so that we do our best to not break userspace that is using the new feature (which I think this series has done well). But supporting MTE/TBI is just like supporting PAN: if someone refactors a driver and swaps a copy_from_user() to a memcpy(), it's going to break under PAN. There will be the same long tail of these bugs like any other, but my sense is that they are small and rare. But I agree: they're going to be pretty weird bugs to track down. The final result, however, will be excellent annotation in the kernel for where userspace addresses get used and people make assumptions about them. The sooner we get the series landed and gain QEMU support (or real hardware), the faster we can hammer out these missed corner-cases. What's the timeline for either of those things, BTW? > > > I feel like I'm missing something about this being seen as an ABI > > > break. The kernel already fails on userspace addresses that have high > > > bits set -- are there things that _depend_ on this failure to operate? > > > > It's about providing a relaxed ABI which allows non-zero top byte and > > breaking it later inadvertently without having something better in place > > to analyse the kernel changes. It sounds like the question is how to switch a process in or out of this ABI (but I don't think that's the real issue: I think it's just a matter of whether or not a process uses tags at all). Doing it at the prctl() level doesn't make sense to me, except maybe to detect MTE support or something. ("Should I tag allocations?") And that state is controlled by the kernel: the kernel does it or it doesn't. If a process wants to not tag, that's also up to the allocator where it can decide not to ask the kernel, and just not tag. Nothing breaks in userspace if a process is NOT tagging and untagged_addr() exists or is missing. This, I think, is the core way this doesn't trip over the golden rule: an old system image will run fine (because it's not tagging). A *new* system may encounter bugs with tagging because it's a new feature: this is The Way Of Things. But we don't break old userspace because old userspace isn't using tags. So the agreement appears to be between the kernel and the allocator. Kernel says "I support this" or not. Telling the allocator to not tag if something breaks sounds like an entirely userspace decision, yes? -- Kees Cook