From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from cuda.sgi.com (cuda2.sgi.com [192.48.176.25]) by oss.sgi.com (8.14.3/8.14.3/SuSE Linux 0.8) with ESMTP id p5271Oc2180485 for ; Thu, 2 Jun 2011 02:01:24 -0500 Received: from ipmail06.adl2.internode.on.net (localhost [127.0.0.1]) by cuda.sgi.com (Spam Firewall) with ESMTP id 125274963B5 for ; Thu, 2 Jun 2011 00:01:21 -0700 (PDT) Received: from ipmail06.adl2.internode.on.net (ipmail06.adl2.internode.on.net [150.101.137.129]) by cuda.sgi.com with ESMTP id g5jfBVQm8sHAbx1e for ; Thu, 02 Jun 2011 00:01:21 -0700 (PDT) From: Dave Chinner Subject: [PATCH 0/12] Per superblock cache reclaim Date: Thu, 2 Jun 2011 17:00:55 +1000 Message-Id: <1306998067-27659-1-git-send-email-david@fromorbit.com> MIME-Version: 1.0 List-Id: XFS Filesystem from SGI List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: base64 Sender: xfs-bounces@oss.sgi.com Errors-To: xfs-bounces@oss.sgi.com To: linux-fsdevel@vger.kernel.org Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, xfs@oss.sgi.com VGhpcyBzZXJpZXMgY29udmVydHMgdGhlIFZGUyBjYWNoZSBzaHJpbmtlcnMgdG8gYSBwZXItc3Vw ZXJibG9jawpzaHJpbmtlciwgYW5kIHByb3ZpZGVzIGEgY2FsbG91dCBmcm9tIHRoZSBzdXBlcmJs b2NrIHNocmlua2VyIHRvCmFsbG93IHRoZSBmaWxlc3lzdGVtIHRvIHNocmluayBpbnRlcm5hbCBj YWNoZXMgcHJvcG9ydGlvbmFsbHkgdG8gdGhlCmFtb3VudCBvZiByZWNsYWltIGRvbmUgdG8gdGhl IFZGUyBjYWNoZXMuCgpUaGUgbW90aXZhdGlvbiBmb3IgdGhpcyB3b3JrIGlzIHRoYXQgdGhlIFZG UyBjYWNoZXMgYXJlIGRlcGVuZGVudApjYWNoZXMgLSBkZW50cmllcyBwaW4gaW5vZGVzLCBhbmQg aW5vZGVzIG9mdGVuIHBpbiBvdGhlciBmaWxlc3lzdGVtCnNwZWNpZmljIHN0cnVjdHVyZXMuICBU aGUgY2FjaGVzIGNhbiBncm93IHF1aXRlIGxhcmdlIGFuZCBpdCBpcyBlYXN5CmZvciB0aGVtIHRv IGdldCB1bmJhbGFuY2VkIHdoZW4gdGhleSBhcmUgc2hydW5rIGluZGVwZW5kZW50bHkuCgpSZWNs YWltIGlzIGFsc28gZm9jdXNzZWQgb24gc2hhcmluZyByZWNsYWltIGJhdGNoZXMgYWNyb3NzIGFs bApzdXBlcmJsb2NrcyByYXRoZXIgdGhhbiB3aXRoaW4gYSBzdXBlcmJsb2NrLCBzbyBvZnRlbiBy ZWNsYWltIGNhbGxzCm9ubHkgcmVtb3ZlIGEgZmV3IG9iamVjdHMgZnJvbSBlYWNoIHN1cGVyYmxv Y2sgYXQgYSB0aW1lLiBUaGlzIG1lYW5zCnRoYXQgd2UgdG91Y2ggbG90cyBvZiBzdXBlcmJsb2Nr cyBhbmQgTFJVcyBvbmUgZXZlcnkgc2hyaW5rZXIgY2FsbCwKYW5kIHdlIGhhdmUgdG8gdHJhdmVy c2UgdGhlIHN1cGVyYmxvY2sgbGlzdCBhbGwgdGhlIHRpbWUuCgpUaGlzIGxlYWRzIHRvIGxpZmUt Y3ljbGUgaXNzdWVzIC0gd2UgaGF2ZSB0byBlbnN1cmUgdGhhdCB0aGUKc3VwZXJibG9jayB3ZSBh cmUgdHJ5aW5nIHRvIHdvcmsgb24gaXMgYWN0aXZlIGFuZCB3b24ndCBnbyBhd2F5LCBhbmQKYWxz byBlbnN1cmUgdGhhdCB0aGUgdW5tb3VudCBwcm9jZXNzIHN5bmNocm9uaXNlcyBjb3JyZWN0bHkg d2l0aAphY3RpdmUgc2hyaW5rZXJzLiBUaGlzIGlzIGNvbXBsZXggYW5kIHRoZSBsb2NrcyBpbnZv bHZlZCBjYXVzZQppc3N1ZXMgd2l0aCBsb2NrZGVwIHJlZnVsYXJseSByZXBvcnRpbmcgZmFsc2Ug cG9zaXRpdmUgbG9jawppbnZlcnNpb25zLgoKRmlyc3RseSwgaG93ZXZlciwgdGhlcmUgYXJlIHNl dmVyYWwgbG9uZ3N0YW5kaW5nIGJ1Z3MgaW4gdGhlClZNIHNocmlua2VyIGluZnJhc3RydWN0dXJl IHRoYXQgbmVlZCB0byBiZSBmaXhlZC4gRmlyc3RseSwgd2UgbmVlZAp0byBhZGQgdHJhY2Vwb2lu dHMgc28gd2UgY2FuIG9ic2VydmUgdGhlIGJlaGF2aW91ciBvZiB0aGUgc2hyaW5rZXIKY2FsY3Vs YXRpb25zLiBTZWNvbmRseSwgdGhlIHNocmlua2VyIHNjYW4gY2FsY3VsYXRpb25zIGFyZSBub3Qg U01QCnNhZmUgYW5kIHRoYXQgaXMgY2F1c2luZyBzaHJpbmtlcnMgdG8gZWl0aGVyIG1pc3Mgd29y ayB0aGV5IHNob3VsZApiZSBkb2luZywgb3IgZG9pbmcgYSBsb3QgbW9yZSB3b3JrIHRoYW4gdGhl eSBzaG91bGQuCgpXaXRoIHRoZXNlIGZpeGVzIGluIHBsYWNlLCBJIGZvdW5kIHRoZSByZWFzb24g dGhhdCBJIHdhcyBub3QgYWJsZSB0bwpiYWxhbmNlIHN5c3RlbSBiZWhhdmlvdXIgb24gbXkgZmly c3QgYXR0ZW1wdCBhdCBwZXItc2Igc2hyaW5rZXJzLgpXaGVuIGEgc2hyaW5rZXIgcmVwZWF0ZWRs eSByZXR1cm5zICItMSIgdG8gYXZvaWQgZGVhZGxvY2tzLCBsaWtlCndpbGwgaGFwcGVuIHdoZW4g YSBmaWxlc3lzdGVtIGlzIGRvaW5nIEdGUF9OT0ZTIG1lbW9yeSBhbGxvY2F0aW9ucwpkdXJpbmcg dHJhbnNhY3Rpb25zIChhbmQgdGhhdCBoYXBwZW5zICphIGxvdCogZHVyaW5nIGZpbGVzeXN0ZW0K aW50ZW5zaXZlIHdvcmtsb2FkcyksIHRoZW4gdGhlIHdvcmsgaXMgZGVsYXllZCBieSBhZGRpbmcg aXQgdG8Kc2hyaW5rZXItPm5yIGZvciB0aGUgbmV4dCBzaHJpbmtlciBjYWxsIHRvIGRvLgoKVGhp cyBjYXVzZXMgdGhlIHNocmlua2VyLT5uciB0byBpbmNyZWFzZSB1bnRpbCBpdCBpcyAyeCB0aGUg bnVtYmVyCm9mIG9iamVjdHMgaW4gdGhlIGNhY2hlLCBhbmQgc28gd2hlbiB0aGUgc2hyaW5rZXIg aXMgZmluYWxseSBhYmxlIHRvCmRvIHdvcmssIGl0IGlzIGVmZmVjdGl2ZWx5IHRvbGQgdG8gc2hy aW5rIHRoZSBlbnRpcmUgY2FjaGUgdG8gemVyby4KVHdpY2Ugb3Zlci4gWW91J2xsIG5ldmVyIGd1 ZXNzIGhvdyBJIGZvdW5kIGl0IC0gdGhlIHRyYWNlcG9pbnRzIEkKYWRkZWQsIHBlcmhhcHM/IFRo aXMgcHJvYmxlbSBpcyBmaXhlZCBieSBvbmx5IGFsbG93aW5nIHRoZQpzaHJpbmtlci0+bnIgdG8g d2luZCB1cCB0byBoYWxmIHRoZSBzaXplIG9mIHRoZSBjYWNoZSB3aGVuIHRoZXJlIGFyZQpsb3Rz IG9mIGxpdHRsZSBhZGRpdGlvbnMgY2F1c2VkIGJ5IGRlYWRsb2NrIGF2b2lkYW5jZS4gVGhpcyBp cwpzdWZmaWNpZW50IHRvIG1haW50YWluIGN1cnJlbnQgbGV2ZWxzIG9mIHBlcmZvcm1hbmNlIHdo aWxzdCBhdm9pZGluZwp0aGUgY2FjaGUgdHJhc2hpbmcgcHJvYmxlbS4KClNvLCBiYWNrIHRvIHRo ZSBWRlMgY2FjaGUgc2hyaW5rZXJzLiAgVG8gYXZvaWQgYWxsIHRoZSBhYm92ZQpwcm9ibGVtcywg d2UgY2FuIHVzZSB0aGUgcGVyLXNocmlua2VyIGNvbnRleHQgaW5mcmFzdHJ1Y3R1cmUgdGhhdAp3 YXMgaW50cm9kdWNlZCByZWNlbnRseSBmb3IgWEZTLiBCeSBhZGRpbmcgYSBzaHJpbmtlciBjb250 ZXh0IHRvCmVhY2ggc3VwZXJibG9jayBhbmQgcmVnaXN0ZXJpbmcgdGhlIHNocmlua2VyIGFmdGVy IHRoZSBzdXBlcmJsb2NrIGlzCmNyZWF0ZWQgYW5kIHVucmVnaXN0ZXJpbmcgaXQgZWFybHkgaW4g dGhlIHVubW91bnQgcHJvY2VzcyB3ZSBhdm9pZAp0aGUgbmVlZCBmb3Igc3BlY2lmaWMgdW5tb3Vu dCBzeW5jaHJvbmlzYXRpb24gYmV0d2VlbiB0aGUgc2hyaW5rZXIKYW5kIHRoZSB1bm1vdW50IHBy b2Nlc3MuICBHb29kYnllIGlwcnVuZV9zZW0uCgpGdXJ0aGVyLCBieSBoYXZpbmcgcGVyLXN1cGVy YmxvY2sgc2hyaW5rZXIgY2FsbG91dHMsIHdlIG5vIGxvbmdlcgpuZWVkIHRvIHdhbGsgdGhlIHN1 cGVyYmxvY2sgbGlzdCBvbiBldmVyeSBzaHJpbmtlciBjYWxsIGZvciBib3RoIHRoZQpkZW50cnkg YW5kIGlub2RlIGNhY2hlcywgbm9yIGRvIHdlIG5lZWQgdG8gcHJvcG9ydGlvbiByZWNsYWltCmJl dHdlZW4gc3VwZXJibG9ja3MuIFRoYXQgc2ltcGxpZmllcyB0aGUgY2FjaGUgc2hyaW5raW5nCmlt cGxlbWVudGF0aW9uIHNpZ25pZmljYW50bHkuCgpIb3dldmVyLCB0byB0YWtlIGFkdmFudGFnZSBv ZiB0aGlzLCB0aGUgZmlyc3QgdGhpbmcgd2UgbmVlZCB0byBkbyBpcwpjb252ZXJ0IHRoZSBpbm9k ZSBjYWNoZSBMUlUgdG8gYSBwZXItc3VwZXJibG9jayBMUlUuIFRoaXMgaXMgdHJpdmlhbAp0byBk byAtIGl0J3MganVzdCBhIGNvcHkgb2YgdGhlIGRlbnRyeSBjYWNoZSBpbmZyYXN0cnVjdHVyZS4g VGhlCmlub2RlIGNhY2hlIExSVSBjYW4gYWxzbyBiZSB0cml2aWFsbHkgY29udmVydGVkIHRvIGEg bG9jayBwZXIKc3VwZXJibG9jayBhcyB3ZWxsLCBzbyB0aGF0IGlzIGRvbmUgYXQgdGhlIHNhbWUg dGltZS4KClsgTm90ZSB0aGF0IGl0IGxvb2tzIGxpa2UgdGhlIHNhbWUgY2hhbmdlIGNhbiBiZSBt YWRlIHRvIHRoZSBkZW50cnkKY2FjaGUgTFJVLCBidXQgdGhlIHNpbXBsZSBjb252ZXJzaW9uIGZy b20gdGhlIGdsb2JhbCBkY2FjaGVfbHJ1X2xvY2sKdG8gcGVyLXNiIGxvY2tzIHJlc3VsdHMgaW4g b2NjYXNpb25hbCwgc3RyYW5nZSBFTk9FTlQgZXJyb3JzIGR1cmluZwpwYXRoIGxvb2t1cHMuIFNv IHRoYXQgcGF0Y2ggaXMgb24gaG9sZC4gXQoKV2l0aCBhIHNpbmdsZSBzaHJpbmtlciAtIHBydW5l X3N1cGVyKCkgLSB0aGF0IGNhbiBhZGRyZXNzIGJvdGggIHRoZQpwZXItc2IgZGVudHJ5IGFuZCBp bm9kZSBMUlVzLCBpdCBpcyBhIHNpbXBsZSBtYXR0ZXIgb2YgcHJvcG9ydGlvbmluZwp0aGUgcmVj bGFpbSBiYXRjaCBiZXR3ZWVuIHRoZW0uIFRoaXMgaXMgZG9uZSBzaW1wbHkgYnkgdGhlIHJhdGlv IG9mCm9iamVjdHMgaW4gdGhlIHR3byBjYWNoZXMsIGFuZCB0aGUgZGVudHJ5IGNhY2hlIGlzIHBy dW5lZCBmaXJzdCBzbwp0aGF0IGl0IHVucGlucyBpbm9kZXMgYmVmb3JlIHRoZSBpbm9kZSBjYWNo ZSBpcyBwcnVuZWQuCgpOb3cgdGhhdCB3ZSBoYXZlIHBydW5lX3N1cGVyKCksIHJlY2xhaW1pbmcg aHVuZHJlZHMgb2YKdGhvdXNhbmRzIG9yIG1pbGxpb25zIG9mIGRlbnRyaWVzIGFuZCBpbm9kZXMg aW4gYmF0Y2hlcyBvZiAxMjgKb2JqZWN0cyBkb2VzIG5vdCBtYWtlIG11Y2ggc2Vuc2UuIFRoZSBW TSBzaHJpbmtlciBpbmZyYXN0cnVjdHVyZQp1c2VzIGEgYmF0Y2ggc2l6ZSBvZiAxMjggc28gdGhh dCBpdCBjYW4gcmVndWxhcmx5IHJlc2NoZWR1bGUgaWYKbmVjZXNzYXJ5LiBUaGUgZGVudHJ5IGNh Y2hlIHBydW5lciBhbHJlYWR5IGhhcyByZXNjaGVkdWxlIGNoZWNrcywKYW5kIGl0IGlzIHRyaXZp YWwgdG8gYWRkIHRoZW0gdG8gdGhlIFZGUyBhbmQgWEZTIGlub2RlIGNhY2hlCnBydW5lcnMuIFdp dGggdGhhdCBkb25lLCB0aGVyZSBpcyBubyByZWFzb24gd2h5IHdlIGNhbid0IHVzZSBhIG11Y2gK bGFyZ2VyIHJlY2xhaW0gYmF0Y2ggc2l6ZSBhbmQgcmVtb3ZlIG1vcmUgb2JqZWN0cyBmcm9tIGVh Y2ggY2FjaGUgb24KZWFjaCB2aXNpdCB0byB0aGVtLgoKVG8gZG8gdGhpcywgYWRkIGEgcGVyLXNo cmlua2VyIGJhdGNoIHNpemUgY29uZmlndXJhdGlvbiBmaWVsZCwgYW5kCmNvbmZpZ3VyZSBwcnVu ZV9zdXBlcigpIHRvIHVzZSBhIGxhcmdlciBiYXRjaCBzaXplIG9mIDEwMjQKb2JqZWN0cy4gVGhp cyByZWR1Y2VzIHRoZSBudW1iZXIgb2YgdGltZXMgd2UgbmVlZCB0byBtYWtlCmNhbGN1bGF0aW9u cywgdHJhZmZpYyBsb2NrcyBhbmQgc3RydWN0dXJlcywgYW5kIG1lYW5zIHdlIHNwZW5kIG1vcmUK dGltZSBpbiBjYWNoZSBzcGVjaWZpYyBsb29wcyB0aGFuIHdlIHdvdWxkIHdpdGggYSBzbWFsbGVy IGJhdGNoCnNpemUuIFRoaXMgcmVkdWNlcyB0aGUgb3ZlcmhlYWQgb2YgY2FjaGUgc2hyaW5raW5n LgoKT3ZlcmFsbCwgdGhlIGNoYW5nZXMgcmVzdWx0IGluIHN0ZWFkeSBzdGF0ZSBjYWNoZSByYXRp b3Mgb24gWEZTLApleHQ0IGFuZCBidHJmcyBvZiAxIGRlbnRyeSA6IDMgaW5vZGVzLiBUaGUgc3Rh dGUgcmF0aW8gaXMgMSBpbnVzZWQKaW5vZGUgOiAyIGZyZWUgaW5vZGVzICh0aGUgaW4tdXNlIGlu b2RlIGlzIHBpbm5lZCBieSB0aGUgZGVudHJ5KS4KVGhlIGZvbGxvd2luZyBjaGFydCBkZW1vbnN0 cmF0ZdGVIGV4dDQgKGxlZnQpIGFuZCBidHJmcyAocmlnaHQpIGNhY2hlCnJhdGlvcyB1bmRlciBz dGVhZHkgc3RhdGUgOC13YXkgZmlsZSBjcmVhdGlvbiBjb25kaXRpb25zLgoKaHR0cDovL3VzZXJ3 ZWIua2VybmVsLm9yZy9+ZGdjL3Nocmlua2VyL2V4dDQtYnRyZnMtY2FjaGUtcmF0aW8ucG5nCgpG b3IgWEZTLCBob3dldmVyLCB0aGUgc2l0dWF0aW9uIGlzIHNsaWdodGx5IG1vcmUgY29tcGxleC4g WEZTCm1haW50YWlucyBpdCdzIG93biBpbm9kZSBjYWNoZSAodGhlIFZGUyBpbm9kZSBjYWNoZSBp cyBhIHN1YnNldCBvZgp0aGUgWEZTIGNhY2hlKSwgYW5kIHNvIG5lZWRzIHRvIGJlIGFibGUgdG8g a2VlcCB0aGF0IHN5bmNocm9uaXNlZAp3aXRoIHRoZSBWRlMgY2FjaGVzLiBIZW5jZSBhIGZpbGVz eXN0ZW0gc3BlY2lmaWMgY2FsbG91dCBpcyBhZGRlZAp0byB0aGUgc3VwZXJibG9jayBwcnVuaW5n IG1ldGhvZCB0aGF0IGlzIHByb3BvcnRpb25lZCB3aXRoIHRoZQpWRlMgZGVudHJ5IGFuZCBpbm9k ZSBjYWNoZXMuIEltcGxlbWVudGluZyB0aGVzZSBtZXRob2RzIGlzIG9wdGlvbmFsLAphbmQgdGhp cyBpcyBkb25lIGZvciBYRlMgaW4gdGhlIGxhc3QgcGF0Y2ggaW4gdGhlIHNlcmllcy4KClhGUyBi ZWhhdmlvdXIgYXQgZGlmZmVyZW50IHN0YWdlcyBvZiB0aGUgcGF0Y2ggc2VyaWVzIGNhbiBiZSBz ZWVuIGluCnRoZSBmb2xsb3dpbmcgY2hhcnQ6CgpodHRwOi8vdXNlcndlYi5rZXJuZWwub3JnL35k Z2Mvc2hyaW5rZXIvcGVyLXNiLXNocmlua2VyLWNvbXBhcmlzb24ucG5nCgpUaGUgbGVmdC1tb3N0 IHRyYWNlcyBhcmUgZnJvbSBhIGtlcm5lbCB3aXRoIGp1c3QgdGhlIFZNCnNocmlua19zbGFiKCkg Zml4ZXMuIFRoZSBtaWRkbGUgdHJhY2UgaXMgdGhlIHNhbWUgOC13YXkgY3JlYXRlCndvcmtsb2Fk LCBidXQgd2l0aCB0aGUgaW5vZGUgY2FjaGUgTFJVIGNoYW5nZXMgYW5kIHRoZSBwZXItc2IKc3Vw ZXJibG9jayBzaHJpbmtlciBhZGRyZXNzaW5nIGp1c3QgdGhlIFZGUyBkZW50cnkgYW5kIGlub2Rl CmNhY2hlcy4gVGhlIHJpZ2h0LW1vc3QgKHBhcnRpYWwpIHdvcmtsb2FkIHRyYWNlIGlzIHRoZSBm dWxsIHNlcmllcwp3aXRoIHRoZSBYRlMgaW5vZGUgY2FjaGUgc2hyaW5rZXIgYmVpbmcgY2FsbGVk IGZyb20gcHJ1bmVfc3VwZXIoKS4KCllvdSBjYW4gc2VlIGZyb20gdGhlIHRvcCBjaGFydCB0aGF0 IHRoZSBjYWNoZSBiZWhhdmlvdXIgaGFzIG11Y2gKbGVzcyB2YXJpYW5jZSBpbiB0aGUgbWlkZGxl IHRyYWNlIHdpdGggdGhlIHBlci1zYiBzaHJpbmtlcnMgY29tcGFyZWQKdG8gdGhlIGxlZnQtbW9z dCB0cmFjZS4gQWxzbywgeW91IGNhbiBzZWUgdGhhdCB0aGUgWEZTIGlub2RlIGNhY2hlCnNpemUg Zm9sbG93cyB0aGUgVkZTIGlub2RlIGNhY2hlIHJlc2lkZW5jeSBtdWNoIG1vcmUgY2xvc2VseSBp biB0aGUKcmlnaHQtbW9zdCB0cmFjZSBhcyBhIHJlc3VsdCBvZiB1c2luZyB0aGUgcHJ1bmVfc3Vw ZXIoKSBmaWxlc3lzdGVtCmNhbGxvdXQuCgpZZXMsIHRoZXNlIFhGUyB0cmFjZXMgYXJlIG11Y2gg bW9yZSB2YXJpYWJsZSB0aGF0IHRoZSBleHQ0IGFuZCBidHJmcwpjaGFydHMsIGJ1dCBYRlMgaXMg cHV0dGluZyBzaWduaWZpY2FudGx5IG1vcmUgcHJlc3N1cmUgb24gdGhlIGNhY2hlcwphbmQgbW9z dCBhbGxvY2F0aW9ucyBhcmUgR0ZQX05PRlMsIGhlbmNlIHRyaWdnZXJpbmcgdGhlIHdpbmQtdXAK cHJvYmxlbXMgZGVzY3JpYmVkIGFib3ZlLiBJdCBpcywgaG93ZXZlciwgbXVjaCBiZXR0ZXIgYmVo YXZlZCB0aGFuCnRoZSBleGlzdGluZyBzaHJpbmtlciBiZWhhdmlvdXIgKHdvcnNlIHRoYW4gdGhl IGxlZnQtbW9zdCB0cmFjZSB3aXRoCnRoZSBWTSBmaXhlcykgYW5kIG11Y2ggYmV0dGVyIHRoYW4g dGhlIHByZXZpb3VzIChhYm9ydGVkKSBwZXItc2IKc2hyaW5rZXIgYXR0ZW1wdHM6CgpodHRwOi8v dXNlcndlYi5rZXJuZWwub3JnL35kZ2Mvc2hyaW5rZXItMi42LjM2L2ZzX21hcmstMi42LjM1LXJj NC1wZXItc2ItYmFzaWMtMTZ4NTAwLXhmcy5wbmcKaHR0cDovL3VzZXJ3ZWIua2VybmVsLm9yZy9+ ZGdjL3Nocmlua2VyLTIuNi4zNi9mc19tYXJrLTIuNi4zNS1yYzQtcGVyLXNiLWJhbGFuY2UtMTZ4 NTAwLXhmcy5wbmcKaHR0cDovL3VzZXJ3ZWIua2VybmVsLm9yZy9+ZGdjL3Nocmlua2VyLTIuNi4z Ni9mc19tYXJrLTIuNi4zNS1yYzQtcGVyLXNiLXByb3BvcnRpb25hbC0xNng1MDAteGZzLnBuZwoK LS0tCgpUaGUgZm9sbG93aW5nIGNoYW5nZXMgc2luY2UgY29tbWl0IGM3NDI3ZDIzZjdlZDY5NWFj MjI2ZGJlM2E4NGQ3ZjE5MDkxZDM0Y2U6CgogIGF1dG9mczQ6IGJvZ3VzIGRlbnRyeV91bmhhc2go KSBhZGRlZCBpbiAtPnVubGluaygpICgyMDExLTA1LTMwIDAxOjUwOjUzIC0wNDAwKQoKYXJlIGF2 YWlsYWJsZSBpbiB0aGUgZ2l0IHJlcG9zaXRvcnkgYXQ6CiAgZ2l0Oi8vZ2l0Lmtlcm5lbC5vcmcv cHViL3NjbS9saW51eC9wZW9wbGUvZGdjL3hmc2Rldi5naXQgcGVyLXNiLXNocmlua2VyCgpEYXZl IENoaW5uZXIgKDEyKToKICAgICAgdm1zY2FuOiBhZGQgc2hyaW5rX3NsYWIgdHJhY2Vwb2ludHMK ICAgICAgdm1zY2FuOiBzaHJpbmtlci0+bnIgdXBkYXRlcyByYWNlIGFuZCBnbyB3cm9uZwogICAg ICB2bXNjYW46IHJlZHVjZSB3aW5kIHVwIHNocmlua2VyLT5uciB3aGVuIHNocmlua2VyIGNhbid0 IGRvIHdvcmsKICAgICAgdm1zY2FuOiBhZGQgY3VzdG9taXNhYmxlIHNocmlua2VyIGJhdGNoIHNp emUKICAgICAgaW5vZGU6IGNvbnZlcnQgaW5vZGVfc3RhdC5ucl91bnVzZWQgdG8gcGVyLWNwdSBj b3VudGVycwogICAgICBpbm9kZTogTWFrZSB1bnVzZWQgaW5vZGUgTFJVIHBlciBzdXBlcmJsb2Nr CiAgICAgIGlub2RlOiBtb3ZlIHRvIHBlci1zYiBMUlUgbG9ja3MKICAgICAgc3VwZXJibG9jazog aW50cm9kdWNlIHBlci1zYiBjYWNoZSBzaHJpbmtlciBpbmZyYXN0cnVjdHVyZQogICAgICBpbm9k ZTogcmVtb3ZlIGlwcnVuZV9zZW0KICAgICAgc3VwZXJibG9jazogYWRkIGZpbGVzeXN0ZW0gc2hy aW5rZXIgb3BlcmF0aW9ucwogICAgICB2ZnM6IGluY3JlYXNlIHNocmlua2VyIGJhdGNoIHNpemUK ICAgICAgeGZzOiBtYWtlIHVzZSBvZiBuZXcgc2hyaW5rZXIgY2FsbG91dCBmb3IgdGhlIGlub2Rl IGNhY2hlCgogRG9jdW1lbnRhdGlvbi9maWxlc3lzdGVtcy92ZnMudHh0IHwgICAyMSArKysrKysK IGZzL2RjYWNoZS5jICAgICAgICAgICAgICAgICAgICAgICB8ICAxMjEgKysrKy0tLS0tLS0tLS0t LS0tLS0tLS0tLS0tLS0tLS0tLS0tCiBmcy9pbm9kZS5jICAgICAgICAgICAgICAgICAgICAgICAg fCAgMTI0ICsrKysrKysrKysrKy0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0KIGZzL3N1cGVyLmMg ICAgICAgICAgICAgICAgICAgICAgICB8ICAgNzkgKysrKysrKysrKysrKysrKysrKysrKystCiBm cy94ZnMvbGludXgtMi42L3hmc19zdXBlci5jICAgICAgfCAgIDI2ICsrKysrLS0tCiBmcy94ZnMv bGludXgtMi42L3hmc19zeW5jLmMgICAgICAgfCAgIDcxICsrKysrKysrLS0tLS0tLS0tLS0tLQog ZnMveGZzL2xpbnV4LTIuNi94ZnNfc3luYy5oICAgICAgIHwgICAgNSArLQogaW5jbHVkZS9saW51 eC9mcy5oICAgICAgICAgICAgICAgIHwgICAxNCArKysrCiBpbmNsdWRlL2xpbnV4L21tLmggICAg ICAgICAgICAgICAgfCAgICAxICsKIGluY2x1ZGUvdHJhY2UvZXZlbnRzL3Ztc2Nhbi5oICAgICB8 ICAgNzEgKysrKysrKysrKysrKysrKysrKysrCiBtbS92bXNjYW4uYyAgICAgICAgICAgICAgICAg ICAgICAgfCAgIDcwICsrKysrKysrKysrKysrKystLS0tLQogMTEgZmlsZXMgY2hhbmdlZCwgMzM3 IGluc2VydGlvbnMoKyksIDI2NiBkZWxldGlvbnMoLSkKCl9fX19fX19fX19fX19fX19fX19fX19f X19fX19fX19fX19fX19fX19fX19fX19fCnhmcyBtYWlsaW5nIGxpc3QKeGZzQG9zcy5zZ2kuY29t Cmh0dHA6Ly9vc3Muc2dpLmNvbS9tYWlsbWFuL2xpc3RpbmZvL3hmcwo= From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932933Ab1FBHB2 (ORCPT ); Thu, 2 Jun 2011 03:01:28 -0400 Received: from ipmail06.adl2.internode.on.net ([150.101.137.129]:27278 "EHLO ipmail06.adl2.internode.on.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932744Ab1FBHBX (ORCPT ); Thu, 2 Jun 2011 03:01:23 -0400 X-IronPort-Anti-Spam-Filtered: true X-IronPort-Anti-Spam-Result: ArYDAAky5015LCoegWdsb2JhbABTG4QuoWcVAQEWJiW2bpBigSuDbIEKBKAp From: Dave Chinner To: linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com Subject: [PATCH 0/12] Per superblock cache reclaim Date: Thu, 2 Jun 2011 17:00:55 +1000 Message-Id: <1306998067-27659-1-git-send-email-david@fromorbit.com> X-Mailer: git-send-email 1.7.5.1 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org This series converts the VFS cache shrinkers to a per-superblock shrinker, and provides a callout from the superblock shrinker to allow the filesystem to shrink internal caches proportionally to the amount of reclaim done to the VFS caches. The motivation for this work is that the VFS caches are dependent caches - dentries pin inodes, and inodes often pin other filesystem specific structures. The caches can grow quite large and it is easy for them to get unbalanced when they are shrunk independently. Reclaim is also focussed on sharing reclaim batches across all superblocks rather than within a superblock, so often reclaim calls only remove a few objects from each superblock at a time. This means that we touch lots of superblocks and LRUs one every shrinker call, and we have to traverse the superblock list all the time. This leads to life-cycle issues - we have to ensure that the superblock we are trying to work on is active and won't go away, and also ensure that the unmount process synchronises correctly with active shrinkers. This is complex and the locks involved cause issues with lockdep refularly reporting false positive lock inversions. Firstly, however, there are several longstanding bugs in the VM shrinker infrastructure that need to be fixed. Firstly, we need to add tracepoints so we can observe the behaviour of the shrinker calculations. Secondly, the shrinker scan calculations are not SMP safe and that is causing shrinkers to either miss work they should be doing, or doing a lot more work than they should. With these fixes in place, I found the reason that I was not able to balance system behaviour on my first attempt at per-sb shrinkers. When a shrinker repeatedly returns "-1" to avoid deadlocks, like will happen when a filesystem is doing GFP_NOFS memory allocations during transactions (and that happens *a lot* during filesystem intensive workloads), then the work is delayed by adding it to shrinker->nr for the next shrinker call to do. This causes the shrinker->nr to increase until it is 2x the number of objects in the cache, and so when the shrinker is finally able to do work, it is effectively told to shrink the entire cache to zero. Twice over. You'll never guess how I found it - the tracepoints I added, perhaps? This problem is fixed by only allowing the shrinker->nr to wind up to half the size of the cache when there are lots of little additions caused by deadlock avoidance. This is sufficient to maintain current levels of performance whilst avoiding the cache trashing problem. So, back to the VFS cache shrinkers. To avoid all the above problems, we can use the per-shrinker context infrastructure that was introduced recently for XFS. By adding a shrinker context to each superblock and registering the shrinker after the superblock is created and unregistering it early in the unmount process we avoid the need for specific unmount synchronisation between the shrinker and the unmount process. Goodbye iprune_sem. Further, by having per-superblock shrinker callouts, we no longer need to walk the superblock list on every shrinker call for both the dentry and inode caches, nor do we need to proportion reclaim between superblocks. That simplifies the cache shrinking implementation significantly. However, to take advantage of this, the first thing we need to do is convert the inode cache LRU to a per-superblock LRU. This is trivial to do - it's just a copy of the dentry cache infrastructure. The inode cache LRU can also be trivially converted to a lock per superblock as well, so that is done at the same time. [ Note that it looks like the same change can be made to the dentry cache LRU, but the simple conversion from the global dcache_lru_lock to per-sb locks results in occasional, strange ENOENT errors during path lookups. So that patch is on hold. ] With a single shrinker - prune_super() - that can address both the per-sb dentry and inode LRUs, it is a simple matter of proportioning the reclaim batch between them. This is done simply by the ratio of objects in the two caches, and the dentry cache is pruned first so that it unpins inodes before the inode cache is pruned. Now that we have prune_super(), reclaiming hundreds of thousands or millions of dentries and inodes in batches of 128 objects does not make much sense. The VM shrinker infrastructure uses a batch size of 128 so that it can regularly reschedule if necessary. The dentry cache pruner already has reschedule checks, and it is trivial to add them to the VFS and XFS inode cache pruners. With that done, there is no reason why we can't use a much larger reclaim batch size and remove more objects from each cache on each visit to them. To do this, add a per-shrinker batch size configuration field, and configure prune_super() to use a larger batch size of 1024 objects. This reduces the number of times we need to make calculations, traffic locks and structures, and means we spend more time in cache specific loops than we would with a smaller batch size. This reduces the overhead of cache shrinking. Overall, the changes result in steady state cache ratios on XFS, ext4 and btrfs of 1 dentry : 3 inodes. The state ratio is 1 inused inode : 2 free inodes (the in-use inode is pinned by the dentry). The following chart demonstrateѕ ext4 (left) and btrfs (right) cache ratios under steady state 8-way file creation conditions. http://userweb.kernel.org/~dgc/shrinker/ext4-btrfs-cache-ratio.png For XFS, however, the situation is slightly more complex. XFS maintains it's own inode cache (the VFS inode cache is a subset of the XFS cache), and so needs to be able to keep that synchronised with the VFS caches. Hence a filesystem specific callout is added to the superblock pruning method that is proportioned with the VFS dentry and inode caches. Implementing these methods is optional, and this is done for XFS in the last patch in the series. XFS behaviour at different stages of the patch series can be seen in the following chart: http://userweb.kernel.org/~dgc/shrinker/per-sb-shrinker-comparison.png The left-most traces are from a kernel with just the VM shrink_slab() fixes. The middle trace is the same 8-way create workload, but with the inode cache LRU changes and the per-sb superblock shrinker addressing just the VFS dentry and inode caches. The right-most (partial) workload trace is the full series with the XFS inode cache shrinker being called from prune_super(). You can see from the top chart that the cache behaviour has much less variance in the middle trace with the per-sb shrinkers compared to the left-most trace. Also, you can see that the XFS inode cache size follows the VFS inode cache residency much more closely in the right-most trace as a result of using the prune_super() filesystem callout. Yes, these XFS traces are much more variable that the ext4 and btrfs charts, but XFS is putting significantly more pressure on the caches and most allocations are GFP_NOFS, hence triggering the wind-up problems described above. It is, however, much better behaved than the existing shrinker behaviour (worse than the left-most trace with the VM fixes) and much better than the previous (aborted) per-sb shrinker attempts: http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-basic-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-balance-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-proportional-16x500-xfs.png --- The following changes since commit c7427d23f7ed695ac226dbe3a84d7f19091d34ce: autofs4: bogus dentry_unhash() added in ->unlink() (2011-05-30 01:50:53 -0400) are available in the git repository at: git://git.kernel.org/pub/scm/linux/people/dgc/xfsdev.git per-sb-shrinker Dave Chinner (12): vmscan: add shrink_slab tracepoints vmscan: shrinker->nr updates race and go wrong vmscan: reduce wind up shrinker->nr when shrinker can't do work vmscan: add customisable shrinker batch size inode: convert inode_stat.nr_unused to per-cpu counters inode: Make unused inode LRU per superblock inode: move to per-sb LRU locks superblock: introduce per-sb cache shrinker infrastructure inode: remove iprune_sem superblock: add filesystem shrinker operations vfs: increase shrinker batch size xfs: make use of new shrinker callout for the inode cache Documentation/filesystems/vfs.txt | 21 ++++++ fs/dcache.c | 121 ++++-------------------------------- fs/inode.c | 124 ++++++++++++------------------------- fs/super.c | 79 +++++++++++++++++++++++- fs/xfs/linux-2.6/xfs_super.c | 26 +++++--- fs/xfs/linux-2.6/xfs_sync.c | 71 ++++++++------------- fs/xfs/linux-2.6/xfs_sync.h | 5 +- include/linux/fs.h | 14 ++++ include/linux/mm.h | 1 + include/trace/events/vmscan.h | 71 +++++++++++++++++++++ mm/vmscan.c | 70 ++++++++++++++++----- 11 files changed, 337 insertions(+), 266 deletions(-) From mboxrd@z Thu Jan 1 00:00:00 1970 From: Dave Chinner Subject: [PATCH 0/12] Per superblock cache reclaim Date: Thu, 2 Jun 2011 17:00:55 +1000 Message-ID: <1306998067-27659-1-git-send-email-david@fromorbit.com> Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com To: linux-fsdevel@vger.kernel.org Return-path: Sender: owner-linux-mm@kvack.org List-Id: linux-fsdevel.vger.kernel.org This series converts the VFS cache shrinkers to a per-superblock shrinker, and provides a callout from the superblock shrinker to allow the filesystem to shrink internal caches proportionally to the amount of reclaim done to the VFS caches. The motivation for this work is that the VFS caches are dependent caches - dentries pin inodes, and inodes often pin other filesystem specific structures. The caches can grow quite large and it is easy for them to get unbalanced when they are shrunk independently. Reclaim is also focussed on sharing reclaim batches across all superblocks rather than within a superblock, so often reclaim calls only remove a few objects from each superblock at a time. This means that we touch lots of superblocks and LRUs one every shrinker call, and we have to traverse the superblock list all the time. This leads to life-cycle issues - we have to ensure that the superblock we are trying to work on is active and won't go away, and also ensure that the unmount process synchronises correctly with active shrinkers. This is complex and the locks involved cause issues with lockdep refularly reporting false positive lock inversions. Firstly, however, there are several longstanding bugs in the VM shrinker infrastructure that need to be fixed. Firstly, we need to add tracepoints so we can observe the behaviour of the shrinker calculations. Secondly, the shrinker scan calculations are not SMP safe and that is causing shrinkers to either miss work they should be doing, or doing a lot more work than they should. With these fixes in place, I found the reason that I was not able to balance system behaviour on my first attempt at per-sb shrinkers. When a shrinker repeatedly returns "-1" to avoid deadlocks, like will happen when a filesystem is doing GFP_NOFS memory allocations during transactions (and that happens *a lot* during filesystem intensive workloads), then the work is delayed by adding it to shrinker->nr for the next shrinker call to do. This causes the shrinker->nr to increase until it is 2x the number of objects in the cache, and so when the shrinker is finally able to do work, it is effectively told to shrink the entire cache to zero. Twice over. You'll never guess how I found it - the tracepoints I added, perhaps? This problem is fixed by only allowing the shrinker->nr to wind up to half the size of the cache when there are lots of little additions caused by deadlock avoidance. This is sufficient to maintain current levels of performance whilst avoiding the cache trashing problem. So, back to the VFS cache shrinkers. To avoid all the above problems, we can use the per-shrinker context infrastructure that was introduced recently for XFS. By adding a shrinker context to each superblock and registering the shrinker after the superblock is created and unregistering it early in the unmount process we avoid the need for specific unmount synchronisation between the shrinker and the unmount process. Goodbye iprune_sem. Further, by having per-superblock shrinker callouts, we no longer need to walk the superblock list on every shrinker call for both the dentry and inode caches, nor do we need to proportion reclaim between superblocks. That simplifies the cache shrinking implementation significantly. However, to take advantage of this, the first thing we need to do is convert the inode cache LRU to a per-superblock LRU. This is trivial to do - it's just a copy of the dentry cache infrastructure. The inode cache LRU can also be trivially converted to a lock per superblock as well, so that is done at the same time. [ Note that it looks like the same change can be made to the dentry cache LRU, but the simple conversion from the global dcache_lru_lock to per-sb locks results in occasional, strange ENOENT errors during path lookups. So that patch is on hold. ] With a single shrinker - prune_super() - that can address both the per-sb dentry and inode LRUs, it is a simple matter of proportioning the reclaim batch between them. This is done simply by the ratio of objects in the two caches, and the dentry cache is pruned first so that it unpins inodes before the inode cache is pruned. Now that we have prune_super(), reclaiming hundreds of thousands or millions of dentries and inodes in batches of 128 objects does not make much sense. The VM shrinker infrastructure uses a batch size of 128 so that it can regularly reschedule if necessary. The dentry cache pruner already has reschedule checks, and it is trivial to add them to the VFS and XFS inode cache pruners. With that done, there is no reason why we can't use a much larger reclaim batch size and remove more objects from each cache on each visit to them. To do this, add a per-shrinker batch size configuration field, and configure prune_super() to use a larger batch size of 1024 objects. This reduces the number of times we need to make calculations, traffic locks and structures, and means we spend more time in cache specific loops than we would with a smaller batch size. This reduces the overhead of cache shrinking. Overall, the changes result in steady state cache ratios on XFS, ext4 and btrfs of 1 dentry : 3 inodes. The state ratio is 1 inused inode : 2 free inodes (the in-use inode is pinned by the dentry). The following chart demonstrate=D1=95 ext4 (left) and btrfs (right) cache ratios under steady state 8-way file creation conditions. http://userweb.kernel.org/~dgc/shrinker/ext4-btrfs-cache-ratio.png For XFS, however, the situation is slightly more complex. XFS maintains it's own inode cache (the VFS inode cache is a subset of the XFS cache), and so needs to be able to keep that synchronised with the VFS caches. Hence a filesystem specific callout is added to the superblock pruning method that is proportioned with the VFS dentry and inode caches. Implementing these methods is optional, and this is done for XFS in the last patch in the series. XFS behaviour at different stages of the patch series can be seen in the following chart: http://userweb.kernel.org/~dgc/shrinker/per-sb-shrinker-comparison.png The left-most traces are from a kernel with just the VM shrink_slab() fixes. The middle trace is the same 8-way create workload, but with the inode cache LRU changes and the per-sb superblock shrinker addressing just the VFS dentry and inode caches. The right-most (partial) workload trace is the full series with the XFS inode cache shrinker being called from prune_super(). You can see from the top chart that the cache behaviour has much less variance in the middle trace with the per-sb shrinkers compared to the left-most trace. Also, you can see that the XFS inode cache size follows the VFS inode cache residency much more closely in the right-most trace as a result of using the prune_super() filesystem callout. Yes, these XFS traces are much more variable that the ext4 and btrfs charts, but XFS is putting significantly more pressure on the caches and most allocations are GFP_NOFS, hence triggering the wind-up problems described above. It is, however, much better behaved than the existing shrinker behaviour (worse than the left-most trace with the VM fixes) and much better than the previous (aborted) per-sb shrinker attempts: http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-= basic-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-= balance-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-= proportional-16x500-xfs.png --- The following changes since commit c7427d23f7ed695ac226dbe3a84d7f19091d34= ce: autofs4: bogus dentry_unhash() added in ->unlink() (2011-05-30 01:50:53= -0400) are available in the git repository at: git://git.kernel.org/pub/scm/linux/people/dgc/xfsdev.git per-sb-shrinke= r Dave Chinner (12): vmscan: add shrink_slab tracepoints vmscan: shrinker->nr updates race and go wrong vmscan: reduce wind up shrinker->nr when shrinker can't do work vmscan: add customisable shrinker batch size inode: convert inode_stat.nr_unused to per-cpu counters inode: Make unused inode LRU per superblock inode: move to per-sb LRU locks superblock: introduce per-sb cache shrinker infrastructure inode: remove iprune_sem superblock: add filesystem shrinker operations vfs: increase shrinker batch size xfs: make use of new shrinker callout for the inode cache Documentation/filesystems/vfs.txt | 21 ++++++ fs/dcache.c | 121 ++++---------------------------= ----- fs/inode.c | 124 ++++++++++++-------------------= ------ fs/super.c | 79 +++++++++++++++++++++++- fs/xfs/linux-2.6/xfs_super.c | 26 +++++--- fs/xfs/linux-2.6/xfs_sync.c | 71 ++++++++------------- fs/xfs/linux-2.6/xfs_sync.h | 5 +- include/linux/fs.h | 14 ++++ include/linux/mm.h | 1 + include/trace/events/vmscan.h | 71 +++++++++++++++++++++ mm/vmscan.c | 70 ++++++++++++++++----- 11 files changed, 337 insertions(+), 266 deletions(-) -- To unsubscribe, send a message with 'unsubscribe linux-mm' in the body to majordomo@kvack.org. For more info on Linux MM, see: http://www.linux-mm.org/ . Fight unfair telecom internet charges in Canada: sign http://stopthemeter= .ca/ Don't email: email@kvack.org From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail138.messagelabs.com (mail138.messagelabs.com [216.82.249.35]) by kanga.kvack.org (Postfix) with SMTP id C404F6B0078 for ; Thu, 2 Jun 2011 03:01:23 -0400 (EDT) From: Dave Chinner Subject: [PATCH 0/12] Per superblock cache reclaim Date: Thu, 2 Jun 2011 17:00:55 +1000 Message-Id: <1306998067-27659-1-git-send-email-david@fromorbit.com> MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sender: owner-linux-mm@kvack.org List-ID: To: linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com This series converts the VFS cache shrinkers to a per-superblock shrinker, and provides a callout from the superblock shrinker to allow the filesystem to shrink internal caches proportionally to the amount of reclaim done to the VFS caches. The motivation for this work is that the VFS caches are dependent caches - dentries pin inodes, and inodes often pin other filesystem specific structures. The caches can grow quite large and it is easy for them to get unbalanced when they are shrunk independently. Reclaim is also focussed on sharing reclaim batches across all superblocks rather than within a superblock, so often reclaim calls only remove a few objects from each superblock at a time. This means that we touch lots of superblocks and LRUs one every shrinker call, and we have to traverse the superblock list all the time. This leads to life-cycle issues - we have to ensure that the superblock we are trying to work on is active and won't go away, and also ensure that the unmount process synchronises correctly with active shrinkers. This is complex and the locks involved cause issues with lockdep refularly reporting false positive lock inversions. Firstly, however, there are several longstanding bugs in the VM shrinker infrastructure that need to be fixed. Firstly, we need to add tracepoints so we can observe the behaviour of the shrinker calculations. Secondly, the shrinker scan calculations are not SMP safe and that is causing shrinkers to either miss work they should be doing, or doing a lot more work than they should. With these fixes in place, I found the reason that I was not able to balance system behaviour on my first attempt at per-sb shrinkers. When a shrinker repeatedly returns "-1" to avoid deadlocks, like will happen when a filesystem is doing GFP_NOFS memory allocations during transactions (and that happens *a lot* during filesystem intensive workloads), then the work is delayed by adding it to shrinker->nr for the next shrinker call to do. This causes the shrinker->nr to increase until it is 2x the number of objects in the cache, and so when the shrinker is finally able to do work, it is effectively told to shrink the entire cache to zero. Twice over. You'll never guess how I found it - the tracepoints I added, perhaps? This problem is fixed by only allowing the shrinker->nr to wind up to half the size of the cache when there are lots of little additions caused by deadlock avoidance. This is sufficient to maintain current levels of performance whilst avoiding the cache trashing problem. So, back to the VFS cache shrinkers. To avoid all the above problems, we can use the per-shrinker context infrastructure that was introduced recently for XFS. By adding a shrinker context to each superblock and registering the shrinker after the superblock is created and unregistering it early in the unmount process we avoid the need for specific unmount synchronisation between the shrinker and the unmount process. Goodbye iprune_sem. Further, by having per-superblock shrinker callouts, we no longer need to walk the superblock list on every shrinker call for both the dentry and inode caches, nor do we need to proportion reclaim between superblocks. That simplifies the cache shrinking implementation significantly. However, to take advantage of this, the first thing we need to do is convert the inode cache LRU to a per-superblock LRU. This is trivial to do - it's just a copy of the dentry cache infrastructure. The inode cache LRU can also be trivially converted to a lock per superblock as well, so that is done at the same time. [ Note that it looks like the same change can be made to the dentry cache LRU, but the simple conversion from the global dcache_lru_lock to per-sb locks results in occasional, strange ENOENT errors during path lookups. So that patch is on hold. ] With a single shrinker - prune_super() - that can address both the per-sb dentry and inode LRUs, it is a simple matter of proportioning the reclaim batch between them. This is done simply by the ratio of objects in the two caches, and the dentry cache is pruned first so that it unpins inodes before the inode cache is pruned. Now that we have prune_super(), reclaiming hundreds of thousands or millions of dentries and inodes in batches of 128 objects does not make much sense. The VM shrinker infrastructure uses a batch size of 128 so that it can regularly reschedule if necessary. The dentry cache pruner already has reschedule checks, and it is trivial to add them to the VFS and XFS inode cache pruners. With that done, there is no reason why we can't use a much larger reclaim batch size and remove more objects from each cache on each visit to them. To do this, add a per-shrinker batch size configuration field, and configure prune_super() to use a larger batch size of 1024 objects. This reduces the number of times we need to make calculations, traffic locks and structures, and means we spend more time in cache specific loops than we would with a smaller batch size. This reduces the overhead of cache shrinking. Overall, the changes result in steady state cache ratios on XFS, ext4 and btrfs of 1 dentry : 3 inodes. The state ratio is 1 inused inode : 2 free inodes (the in-use inode is pinned by the dentry). The following chart demonstrateN? ext4 (left) and btrfs (right) cache ratios under steady state 8-way file creation conditions. http://userweb.kernel.org/~dgc/shrinker/ext4-btrfs-cache-ratio.png For XFS, however, the situation is slightly more complex. XFS maintains it's own inode cache (the VFS inode cache is a subset of the XFS cache), and so needs to be able to keep that synchronised with the VFS caches. Hence a filesystem specific callout is added to the superblock pruning method that is proportioned with the VFS dentry and inode caches. Implementing these methods is optional, and this is done for XFS in the last patch in the series. XFS behaviour at different stages of the patch series can be seen in the following chart: http://userweb.kernel.org/~dgc/shrinker/per-sb-shrinker-comparison.png The left-most traces are from a kernel with just the VM shrink_slab() fixes. The middle trace is the same 8-way create workload, but with the inode cache LRU changes and the per-sb superblock shrinker addressing just the VFS dentry and inode caches. The right-most (partial) workload trace is the full series with the XFS inode cache shrinker being called from prune_super(). You can see from the top chart that the cache behaviour has much less variance in the middle trace with the per-sb shrinkers compared to the left-most trace. Also, you can see that the XFS inode cache size follows the VFS inode cache residency much more closely in the right-most trace as a result of using the prune_super() filesystem callout. Yes, these XFS traces are much more variable that the ext4 and btrfs charts, but XFS is putting significantly more pressure on the caches and most allocations are GFP_NOFS, hence triggering the wind-up problems described above. It is, however, much better behaved than the existing shrinker behaviour (worse than the left-most trace with the VM fixes) and much better than the previous (aborted) per-sb shrinker attempts: http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-basic-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-balance-16x500-xfs.png http://userweb.kernel.org/~dgc/shrinker-2.6.36/fs_mark-2.6.35-rc4-per-sb-proportional-16x500-xfs.png --- The following changes since commit c7427d23f7ed695ac226dbe3a84d7f19091d34ce: autofs4: bogus dentry_unhash() added in ->unlink() (2011-05-30 01:50:53 -0400) are available in the git repository at: git://git.kernel.org/pub/scm/linux/people/dgc/xfsdev.git per-sb-shrinker Dave Chinner (12): vmscan: add shrink_slab tracepoints vmscan: shrinker->nr updates race and go wrong vmscan: reduce wind up shrinker->nr when shrinker can't do work vmscan: add customisable shrinker batch size inode: convert inode_stat.nr_unused to per-cpu counters inode: Make unused inode LRU per superblock inode: move to per-sb LRU locks superblock: introduce per-sb cache shrinker infrastructure inode: remove iprune_sem superblock: add filesystem shrinker operations vfs: increase shrinker batch size xfs: make use of new shrinker callout for the inode cache Documentation/filesystems/vfs.txt | 21 ++++++ fs/dcache.c | 121 ++++-------------------------------- fs/inode.c | 124 ++++++++++++------------------------- fs/super.c | 79 +++++++++++++++++++++++- fs/xfs/linux-2.6/xfs_super.c | 26 +++++--- fs/xfs/linux-2.6/xfs_sync.c | 71 ++++++++------------- fs/xfs/linux-2.6/xfs_sync.h | 5 +- include/linux/fs.h | 14 ++++ include/linux/mm.h | 1 + include/trace/events/vmscan.h | 71 +++++++++++++++++++++ mm/vmscan.c | 70 ++++++++++++++++----- 11 files changed, 337 insertions(+), 266 deletions(-) -- To unsubscribe, send a message with 'unsubscribe linux-mm' in the body to majordomo@kvack.org. For more info on Linux MM, see: http://www.linux-mm.org/ . Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/ Don't email: email@kvack.org