From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from cuda.sgi.com (cuda3.sgi.com [192.48.176.15]) by oss.sgi.com (8.14.3/8.14.3/SuSE Linux 0.8) with ESMTP id p5271nqA180604 for ; Thu, 2 Jun 2011 02:01:49 -0500 Received: from ipmail06.adl2.internode.on.net (localhost [127.0.0.1]) by cuda.sgi.com (Spam Firewall) with ESMTP id 682ED1ED190D for ; Thu, 2 Jun 2011 00:01:46 -0700 (PDT) Received: from ipmail06.adl2.internode.on.net (ipmail06.adl2.internode.on.net [150.101.137.129]) by cuda.sgi.com with ESMTP id tk2FGOLHYEq8HN7u for ; Thu, 02 Jun 2011 00:01:46 -0700 (PDT) From: Dave Chinner Subject: =?UTF-8?q?=5BPATCH=2008/12=5D=20superblock=3A=20introduce=20per-sb=20cache=20shrinker=20infrastructure?= Date: Thu, 2 Jun 2011 17:01:03 +1000 Message-Id: <1306998067-27659-9-git-send-email-david@fromorbit.com> In-Reply-To: <1306998067-27659-1-git-send-email-david@fromorbit.com> References: <1306998067-27659-1-git-send-email-david@fromorbit.com> MIME-Version: 1.0 List-Id: XFS Filesystem from SGI List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: base64 Sender: xfs-bounces@oss.sgi.com Errors-To: xfs-bounces@oss.sgi.com To: linux-fsdevel@vger.kernel.org Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, xfs@oss.sgi.com RnJvbTogRGF2ZSBDaGlubmVyIDxkY2hpbm5lckByZWRoYXQuY29tPgoKV2l0aCBjb250ZXh0IGJh c2VkIHNocmlua2Vycywgd2UgY2FuIGltcGxlbWVudCBhIHBlci1zdXBlcmJsb2NrCnNocmlua2Vy IHRoYXQgc2hyaW5rcyB0aGUgY2FjaGVzIGF0dGFjaGVkIHRvIHRoZSBzdXBlcmJsb2NrLiBXZQpj dXJyZW50bHkgaGF2ZSBnbG9iYWwgc2hyaW5rZXJzIGZvciB0aGUgaW5vZGUgYW5kIGRlbnRyeSBj YWNoZXMgdGhhdApzcGxpdCB1cCBpbnRvIHBlci1zdXBlcmJsb2NrIG9wZXJhdGlvbnMgdmlhIGEg Y29hcnNlIHByb3BvcnRpb25pbmcKbWV0aG9kIHRoYXQgZG9lcyBub3QgYmF0Y2ggdmVyeSB3ZWxs LiAgVGhlIGdsb2JhbCBzaHJpbmtlcnMgYWxzbwpoYXZlIGEgZGVwZW5kZW5jeSAtIGRlbnRyaWVz IHBpbiBpbm9kZXMgLSBzbyB3ZSBoYXZlIHRvIGJlIHZlcnkKY2FyZWZ1bCBhYm91dCBob3cgd2Ug cmVnaXN0ZXIgdGhlIGdsb2JhbCBzaHJpbmtlcnMgc28gdGhhdCB0aGUKaW1wbGljaXQgY2FsbCBv cmRlciBpcyBhbHdheXMgY29ycmVjdC4KCldpdGggYSBwZXItc2Igc2hyaW5rZXIgY2FsbG91dCwg d2UgY2FuIGVuY29kZSB0aGlzIGRlcGVuZGVuY3kKZGlyZWN0bHkgaW50byB0aGUgcGVyLXNiIHNo cmlua2VyLCBoZW5jZSBhdm9pZGluZyB0aGUgbmVlZCBmb3IKc3RyaWN0bHkgb3JkZXJpbmcgc2hy aW5rZXIgcmVnaXN0cmF0aW9ucy4gV2UgYWxzbyBoYXZlIG5vIG5lZWQgZm9yCmFueSBwcm9wb3J0 aW9uaW5nIGNvZGUgZm9yIHRoZSBzaHJpbmtlciBzdWJzeXN0ZW0gYWxyZWFkeSBwcm92aWRlcwp0 aGlzIGZ1bmN0aW9uYWxpdHkgYWNyb3NzIGFsbCBzaHJpbmtlcnMuIEFsbG93aW5nIHRoZSBzaHJp bmtlciB0bwpvcGVyYXRlIG9uIGEgc2luZ2xlIHN1cGVyYmxvY2sgYXQgYSB0aW1lIG1lYW5zIHRo YXQgd2UgZG8gbGVzcwpzdXBlcmJsb2NrIGxpc3QgdHJhdmVyc2FscyBhbmQgbG9ja2luZyBhbmQg cmVjbGFpbSBzaG91bGQgYmF0Y2ggbW9yZQplZmZlY3RpdmVseS4gVGhpcyBzaG91bGQgcmVzdWx0 IGluIGxlc3MgQ1BVIG92ZXJoZWFkIGZvciByZWNsYWltIGFuZApwb3RlbnRpYWxseSBmYXN0ZXIg cmVjbGFpbSBvZiBpdGVtcyBmcm9tIGVhY2ggZmlsZXN5c3RlbS4KClNpZ25lZC1vZmYtYnk6IERh dmUgQ2hpbm5lciA8ZGNoaW5uZXJAcmVkaGF0LmNvbT4KLS0tCiBmcy9kY2FjaGUuYyAgICAgICAg fCAgMTIxICsrKysrLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0t LQogZnMvaW5vZGUuYyAgICAgICAgIHwgIDExNyArKysrLS0tLS0tLS0tLS0tLS0tLS0tLS0tLS0t LS0tLS0tLS0tLS0tLS0tLS0tLS0tLQogZnMvc3VwZXIuYyAgICAgICAgIHwgICA1NSArKysrKysr KysrKysrKysrKysrKysrKy0KIGluY2x1ZGUvbGludXgvZnMuaCB8ICAgIDcgKysrCiA0IGZpbGVz IGNoYW5nZWQsIDgyIGluc2VydGlvbnMoKyksIDIxOCBkZWxldGlvbnMoLSkKCmRpZmYgLS1naXQg YS9mcy9kY2FjaGUuYyBiL2ZzL2RjYWNoZS5jCmluZGV4IDM3ZjcyZWUuLmY3M2VmMjMgMTAwNjQ0 Ci0tLSBhL2ZzL2RjYWNoZS5jCisrKyBiL2ZzL2RjYWNoZS5jCkBAIC03MjAsMTMgKzcyMCwxMSBA QCBzdGF0aWMgdm9pZCBzaHJpbmtfZGVudHJ5X2xpc3Qoc3RydWN0IGxpc3RfaGVhZCAqbGlzdCkK ICAqCiAgKiBJZiBmbGFncyBjb250YWlucyBEQ0FDSEVfUkVGRVJFTkNFRCByZWZlcmVuY2UgZGVu dHJpZXMgd2lsbCBub3QgYmUgcHJ1bmVkLgogICovCi1zdGF0aWMgdm9pZCBfX3Nocmlua19kY2Fj aGVfc2Ioc3RydWN0IHN1cGVyX2Jsb2NrICpzYiwgaW50ICpjb3VudCwgaW50IGZsYWdzKQorc3Rh dGljIHZvaWQgX19zaHJpbmtfZGNhY2hlX3NiKHN0cnVjdCBzdXBlcl9ibG9jayAqc2IsIGludCBj b3VudCwgaW50IGZsYWdzKQogewotCS8qIGNhbGxlZCBmcm9tIHBydW5lX2RjYWNoZSgpIGFuZCBz aHJpbmtfZGNhY2hlX3BhcmVudCgpICovCiAJc3RydWN0IGRlbnRyeSAqZGVudHJ5OwogCUxJU1Rf SEVBRChyZWZlcmVuY2VkKTsKIAlMSVNUX0hFQUQodG1wKTsKLQlpbnQgY250ID0gKmNvdW50Owog CiByZWxvY2s6CiAJc3Bpbl9sb2NrKCZkY2FjaGVfbHJ1X2xvY2spOwpAQCAtNzU0LDcgKzc1Miw3 IEBAIHJlbG9jazoKIAkJfSBlbHNlIHsKIAkJCWxpc3RfbW92ZV90YWlsKCZkZW50cnktPmRfbHJ1 LCAmdG1wKTsKIAkJCXNwaW5fdW5sb2NrKCZkZW50cnktPmRfbG9jayk7Ci0JCQlpZiAoIS0tY250 KQorCQkJaWYgKCEtLWNvdW50KQogCQkJCWJyZWFrOwogCQl9CiAJCWNvbmRfcmVzY2hlZF9sb2Nr KCZkY2FjaGVfbHJ1X2xvY2spOwpAQCAtNzY0LDgzICs3NjIsMjIgQEAgcmVsb2NrOgogCXNwaW5f dW5sb2NrKCZkY2FjaGVfbHJ1X2xvY2spOwogCiAJc2hyaW5rX2RlbnRyeV9saXN0KCZ0bXApOwot Ci0JKmNvdW50ID0gY250OwogfQogCiAvKioKLSAqIHBydW5lX2RjYWNoZSAtIHNocmluayB0aGUg ZGNhY2hlCi0gKiBAY291bnQ6IG51bWJlciBvZiBlbnRyaWVzIHRvIHRyeSB0byBmcmVlCisgKiBw cnVuZV9kY2FjaGVfc2IgLSBzaHJpbmsgdGhlIGRjYWNoZQorICogQG5yX3RvX3NjYW46IG51bWJl ciBvZiBlbnRyaWVzIHRvIHRyeSB0byBmcmVlCiAgKgotICogU2hyaW5rIHRoZSBkY2FjaGUuIFRo aXMgaXMgZG9uZSB3aGVuIHdlIG5lZWQgbW9yZSBtZW1vcnksIG9yIHNpbXBseSB3aGVuIHdlCi0g KiBuZWVkIHRvIHVubW91bnQgc29tZXRoaW5nIChhdCB3aGljaCBwb2ludCB3ZSBuZWVkIHRvIHVu dXNlIGFsbCBkZW50cmllcykuCisgKiBBdHRlbXB0IHRvIHNocmluayB0aGUgc3VwZXJibG9jayBk Y2FjaGUgTFJVIGJ5IEBucl90b19zY2FuIGVudHJpZXMuIFRoaXMgaXMKKyAqIGRvbmUgd2hlbiB3 ZSBuZWVkIG1vcmUgbWVtb3J5IGFuIGNhbGxlZCBmcm9tIHRoZSBzdXBlcmJsb2NrIHNocmlua2Vy CisgKiBmdW5jdGlvbi4KICAqCi0gKiBUaGlzIGZ1bmN0aW9uIG1heSBmYWlsIHRvIGZyZWUgYW55 IHJlc291cmNlcyBpZiBhbGwgdGhlIGRlbnRyaWVzIGFyZSBpbiB1c2UuCisgKiBUaGlzIGZ1bmN0 aW9uIG1heSBmYWlsIHRvIGZyZWUgYW55IHJlc291cmNlcyBpZiBhbGwgdGhlIGRlbnRyaWVzIGFy ZSBpbgorICogdXNlLgogICovCi1zdGF0aWMgdm9pZCBwcnVuZV9kY2FjaGUoaW50IGNvdW50KQor dm9pZCBwcnVuZV9kY2FjaGVfc2Ioc3RydWN0IHN1cGVyX2Jsb2NrICpzYiwgaW50IG5yX3RvX3Nj YW4pCiB7Ci0Jc3RydWN0IHN1cGVyX2Jsb2NrICpzYiwgKnAgPSBOVUxMOwotCWludCB3X2NvdW50 OwotCWludCB1bnVzZWQgPSBkZW50cnlfc3RhdC5ucl91bnVzZWQ7Ci0JaW50IHBydW5lX3JhdGlv OwotCWludCBwcnVuZWQ7Ci0KLQlpZiAodW51c2VkID09IDAgfHwgY291bnQgPT0gMCkKLQkJcmV0 dXJuOwotCWlmIChjb3VudCA+PSB1bnVzZWQpCi0JCXBydW5lX3JhdGlvID0gMTsKLQllbHNlCi0J CXBydW5lX3JhdGlvID0gdW51c2VkIC8gY291bnQ7Ci0Jc3Bpbl9sb2NrKCZzYl9sb2NrKTsKLQls aXN0X2Zvcl9lYWNoX2VudHJ5KHNiLCAmc3VwZXJfYmxvY2tzLCBzX2xpc3QpIHsKLQkJaWYgKGxp c3RfZW1wdHkoJnNiLT5zX2luc3RhbmNlcykpCi0JCQljb250aW51ZTsKLQkJaWYgKHNiLT5zX25y X2RlbnRyeV91bnVzZWQgPT0gMCkKLQkJCWNvbnRpbnVlOwotCQlzYi0+c19jb3VudCsrOwotCQkv KiBOb3csIHdlIHJlY2xhaW0gdW51c2VkIGRlbnRyaW5zIHdpdGggZmFpcm5lc3MuCi0JCSAqIFdl IHJlY2xhaW0gdGhlbSBzYW1lIHBlcmNlbnRhZ2UgZnJvbSBlYWNoIHN1cGVyYmxvY2suCi0JCSAq IFdlIGNhbGN1bGF0ZSBudW1iZXIgb2YgZGVudHJpZXMgdG8gc2NhbiBvbiB0aGlzIHNiCi0JCSAq IGFzIGZvbGxvd3MsIGJ1dCB0aGUgaW1wbGVtZW50YXRpb24gaXMgYXJyYW5nZWQgdG8gYXZvaWQK LQkJICogb3ZlcmZsb3dzOgotCQkgKiBudW1iZXIgb2YgZGVudHJpZXMgdG8gc2NhbiBvbiB0aGlz IHNiID0KLQkJICogY291bnQgKiAobnVtYmVyIG9mIGRlbnRyaWVzIG9uIHRoaXMgc2IgLwotCQkg KiBudW1iZXIgb2YgZGVudHJpZXMgaW4gdGhlIG1hY2hpbmUpCi0JCSAqLwotCQlzcGluX3VubG9j aygmc2JfbG9jayk7Ci0JCWlmIChwcnVuZV9yYXRpbyAhPSAxKQotCQkJd19jb3VudCA9IChzYi0+ c19ucl9kZW50cnlfdW51c2VkIC8gcHJ1bmVfcmF0aW8pICsgMTsKLQkJZWxzZQotCQkJd19jb3Vu dCA9IHNiLT5zX25yX2RlbnRyeV91bnVzZWQ7Ci0JCXBydW5lZCA9IHdfY291bnQ7Ci0JCS8qCi0J CSAqIFdlIG5lZWQgdG8gYmUgc3VyZSB0aGlzIGZpbGVzeXN0ZW0gaXNuJ3QgYmVpbmcgdW5tb3Vu dGVkLAotCQkgKiBvdGhlcndpc2Ugd2UgY291bGQgcmFjZSB3aXRoIGdlbmVyaWNfc2h1dGRvd25f c3VwZXIoKSwgYW5kCi0JCSAqIGVuZCB1cCBob2xkaW5nIGEgcmVmZXJlbmNlIHRvIGFuIGlub2Rl IHdoaWxlIHRoZSBmaWxlc3lzdGVtCi0JCSAqIGlzIHVubW91bnRlZC4gIFNvIHdlIHRyeSB0byBn ZXQgc191bW91bnQsIGFuZCBtYWtlIHN1cmUKLQkJICogc19yb290IGlzbid0IE5VTEwuCi0JCSAq LwotCQlpZiAoZG93bl9yZWFkX3RyeWxvY2soJnNiLT5zX3Vtb3VudCkpIHsKLQkJCWlmICgoc2It PnNfcm9vdCAhPSBOVUxMKSAmJgotCQkJICAgICghbGlzdF9lbXB0eSgmc2ItPnNfZGVudHJ5X2xy dSkpKSB7Ci0JCQkJX19zaHJpbmtfZGNhY2hlX3NiKHNiLCAmd19jb3VudCwKLQkJCQkJCURDQUNI RV9SRUZFUkVOQ0VEKTsKLQkJCQlwcnVuZWQgLT0gd19jb3VudDsKLQkJCX0KLQkJCXVwX3JlYWQo JnNiLT5zX3Vtb3VudCk7Ci0JCX0KLQkJc3Bpbl9sb2NrKCZzYl9sb2NrKTsKLQkJaWYgKHApCi0J CQlfX3B1dF9zdXBlcihwKTsKLQkJY291bnQgLT0gcHJ1bmVkOwotCQlwID0gc2I7Ci0JCS8qIG1v cmUgd29yayBsZWZ0IHRvIGRvPyAqLwotCQlpZiAoY291bnQgPD0gMCkKLQkJCWJyZWFrOwotCX0K LQlpZiAocCkKLQkJX19wdXRfc3VwZXIocCk7Ci0Jc3Bpbl91bmxvY2soJnNiX2xvY2spOworCV9f c2hyaW5rX2RjYWNoZV9zYihzYiwgbnJfdG9fc2NhbiwgRENBQ0hFX1JFRkVSRU5DRUQpOwogfQog CiAvKioKQEAgLTEyMTUsNDIgKzExNTIsMTAgQEAgdm9pZCBzaHJpbmtfZGNhY2hlX3BhcmVudChz dHJ1Y3QgZGVudHJ5ICogcGFyZW50KQogCWludCBmb3VuZDsKIAogCXdoaWxlICgoZm91bmQgPSBz ZWxlY3RfcGFyZW50KHBhcmVudCkpICE9IDApCi0JCV9fc2hyaW5rX2RjYWNoZV9zYihzYiwgJmZv dW5kLCAwKTsKKwkJX19zaHJpbmtfZGNhY2hlX3NiKHNiLCBmb3VuZCwgMCk7CiB9CiBFWFBPUlRf U1lNQk9MKHNocmlua19kY2FjaGVfcGFyZW50KTsKIAotLyoKLSAqIFNjYW4gYHNjLT5ucl9zbGFi X3RvX3JlY2xhaW0nIGRlbnRyaWVzIGFuZCByZXR1cm4gdGhlIG51bWJlciB3aGljaCByZW1haW4u Ci0gKgotICogV2UgbmVlZCB0byBhdm9pZCByZWVudGVyaW5nIHRoZSBmaWxlc3lzdGVtIGlmIHRo ZSBjYWxsZXIgaXMgcGVyZm9ybWluZyBhCi0gKiBHRlBfTk9GUyBhbGxvY2F0aW9uIGF0dGVtcHQu ICBPbmUgZXhhbXBsZSBkZWFkbG9jayBpczoKLSAqCi0gKiBleHQyX25ld19ibG9jay0+Z2V0Ymxr LT5HRlAtPnNocmlua19kY2FjaGVfbWVtb3J5LT5wcnVuZV9kY2FjaGUtPgotICogcHJ1bmVfb25l X2RlbnRyeS0+ZHB1dC0+ZGVudHJ5X2lwdXQtPmlwdXQtPmlub2RlLT5pX3NiLT5zX29wLT5wdXRf aW5vZGUtPgotICogZXh0Ml9kaXNjYXJkX3ByZWFsbG9jLT5leHQyX2ZyZWVfYmxvY2tzLT5sb2Nr X3N1cGVyLT5ERUFETE9DSy4KLSAqCi0gKiBJbiB0aGlzIGNhc2Ugd2UgcmV0dXJuIC0xIHRvIHRl bGwgdGhlIGNhbGxlciB0aGF0IHdlIGJhbGVkLgotICovCi1zdGF0aWMgaW50IHNocmlua19kY2Fj aGVfbWVtb3J5KHN0cnVjdCBzaHJpbmtlciAqc2hyaW5rLAotCQkJCXN0cnVjdCBzaHJpbmtfY29u dHJvbCAqc2MpCi17Ci0JaW50IG5yID0gc2MtPm5yX3RvX3NjYW47Ci0JZ2ZwX3QgZ2ZwX21hc2sg PSBzYy0+Z2ZwX21hc2s7Ci0KLQlpZiAobnIpIHsKLQkJaWYgKCEoZ2ZwX21hc2sgJiBfX0dGUF9G UykpCi0JCQlyZXR1cm4gLTE7Ci0JCXBydW5lX2RjYWNoZShucik7Ci0JfQotCi0JcmV0dXJuIChk ZW50cnlfc3RhdC5ucl91bnVzZWQgLyAxMDApICogc3lzY3RsX3Zmc19jYWNoZV9wcmVzc3VyZTsK LX0KLQotc3RhdGljIHN0cnVjdCBzaHJpbmtlciBkY2FjaGVfc2hyaW5rZXIgPSB7Ci0JLnNocmlu ayA9IHNocmlua19kY2FjaGVfbWVtb3J5LAotCS5zZWVrcyA9IERFRkFVTFRfU0VFS1MsCi19Owot CiAvKioKICAqIGRfYWxsb2MJLQlhbGxvY2F0ZSBhIGRjYWNoZSBlbnRyeQogICogQHBhcmVudDog cGFyZW50IG9mIGVudHJ5IHRvIGFsbG9jYXRlCkBAIC0zMDMwLDggKzI5MzUsNiBAQCBzdGF0aWMg dm9pZCBfX2luaXQgZGNhY2hlX2luaXQodm9pZCkKIAkgKi8KIAlkZW50cnlfY2FjaGUgPSBLTUVN X0NBQ0hFKGRlbnRyeSwKIAkJU0xBQl9SRUNMQUlNX0FDQ09VTlR8U0xBQl9QQU5JQ3xTTEFCX01F TV9TUFJFQUQpOwotCQotCXJlZ2lzdGVyX3Nocmlua2VyKCZkY2FjaGVfc2hyaW5rZXIpOwogCiAJ LyogSGFzaCBtYXkgaGF2ZSBiZWVuIHNldCB1cCBpbiBkY2FjaGVfaW5pdF9lYXJseSAqLwogCWlm ICghaGFzaGRpc3QpCmRpZmYgLS1naXQgYS9mcy9pbm9kZS5jIGIvZnMvaW5vZGUuYwppbmRleCA2 NjdhMjljLi44OTBkOTVlIDEwMDY0NAotLS0gYS9mcy9pbm9kZS5jCisrKyBiL2ZzL2lub2RlLmMK QEAgLTczLDcgKzczLDcgQEAgX19jYWNoZWxpbmVfYWxpZ25lZF9pbl9zbXAgREVGSU5FX1NQSU5M T0NLKGlub2RlX3diX2xpc3RfbG9jayk7CiAgKgogICogV2UgZG9uJ3QgYWN0dWFsbHkgbmVlZCBp dCB0byBwcm90ZWN0IGFueXRoaW5nIGluIHRoZSB1bW91bnQgcGF0aCwKICAqIGJ1dCBvbmx5IG5l ZWQgdG8gY3ljbGUgdGhyb3VnaCBpdCB0byBtYWtlIHN1cmUgYW55IGlub2RlIHRoYXQKLSAqIHBy dW5lX2ljYWNoZSB0b29rIG9mZiB0aGUgTFJVIGxpc3QgaGFzIGJlZW4gZnVsbHkgdG9ybiBkb3du IGJ5IHRoZQorICogcHJ1bmVfaWNhY2hlX3NiIHRvb2sgb2ZmIHRoZSBMUlUgbGlzdCBoYXMgYmVl biBmdWxseSB0b3JuIGRvd24gYnkgdGhlCiAgKiB0aW1lIHdlIGFyZSBwYXN0IGV2aWN0X2lub2Rl cy4KICAqLwogc3RhdGljIERFQ0xBUkVfUldTRU0oaXBydW5lX3NlbSk7CkBAIC01MzcsNyArNTM3 LDcgQEAgdm9pZCBldmljdF9pbm9kZXMoc3RydWN0IHN1cGVyX2Jsb2NrICpzYikKIAlkaXNwb3Nl X2xpc3QoJmRpc3Bvc2UpOwogCiAJLyoKLQkgKiBDeWNsZSB0aHJvdWdoIGlwcnVuZV9zZW0gdG8g bWFrZSBzdXJlIGFueSBpbm9kZSB0aGF0IHBydW5lX2ljYWNoZQorCSAqIEN5Y2xlIHRocm91Z2gg aXBydW5lX3NlbSB0byBtYWtlIHN1cmUgYW55IGlub2RlIHRoYXQgcHJ1bmVfaWNhY2hlX3NiCiAJ ICogbW92ZWQgb2ZmIHRoZSBsaXN0IGJlZm9yZSB3ZSB0b29rIHRoZSBsb2NrIGhhcyBiZWVuIGZ1 bGx5IHRvcm4KIAkgKiBkb3duLgogCSAqLwpAQCAtNjA1LDkgKzYwNSwxMCBAQCBzdGF0aWMgaW50 IGNhbl91bnVzZShzdHJ1Y3QgaW5vZGUgKmlub2RlKQogfQogCiAvKgotICogU2NhbiBgZ29hbCcg aW5vZGVzIG9uIHRoZSB1bnVzZWQgbGlzdCBmb3IgZnJlZWFibGUgb25lcy4gVGhleSBhcmUgbW92 ZWQgdG8gYQotICogdGVtcG9yYXJ5IGxpc3QgYW5kIHRoZW4gYXJlIGZyZWVkIG91dHNpZGUgc2It PnNfaW5vZGVfbHJ1X2xvY2sgYnkKLSAqIGRpc3Bvc2VfbGlzdCgpLgorICogV2FsayB0aGUgc3Vw ZXJibG9jayBpbm9kZSBMUlUgZm9yIGZyZWVhYmxlIGlub2RlcyBhbmQgYXR0ZW1wdCB0byBmcmVl IHRoZW0uCisgKiBUaGlzIGlzIGNhbGxlZCBmcm9tIHRoZSBzdXBlcmJsb2NrIHNocmlua2VyIGZ1 bmN0aW9uIHdpdGggYSBudW1iZXIgb2YgaW5vZGVzCisgKiB0byB0cmltIGZyb20gdGhlIExSVS4g SW5vZGVzIHRvIGJlIGZyZWVkIGFyZSBtb3ZlZCB0byBhIHRlbXBvcmFyeSBsaXN0IGFuZAorICog dGhlbiBhcmUgZnJlZWQgb3V0c2lkZSBpbm9kZV9sb2NrIGJ5IGRpc3Bvc2VfbGlzdCgpLgogICoK ICAqIEFueSBpbm9kZXMgd2hpY2ggYXJlIHBpbm5lZCBwdXJlbHkgYmVjYXVzZSBvZiBhdHRhY2hl ZCBwYWdlY2FjaGUgaGF2ZSB0aGVpcgogICogcGFnZWNhY2hlIHJlbW92ZWQuICBJZiB0aGUgaW5v ZGUgaGFzIG1ldGFkYXRhIGJ1ZmZlcnMgYXR0YWNoZWQgdG8KQEAgLTYyMSwxNCArNjIyLDE1IEBA IHN0YXRpYyBpbnQgY2FuX3VudXNlKHN0cnVjdCBpbm9kZSAqaW5vZGUpCiAgKiBMUlUgZG9lcyBu b3QgaGF2ZSBzdHJpY3Qgb3JkZXJpbmcuIEhlbmNlIHdlIGRvbid0IHdhbnQgdG8gcmVjbGFpbSBp bm9kZXMKICAqIHdpdGggdGhpcyBmbGFnIHNldCBiZWNhdXNlIHRoZXkgYXJlIHRoZSBpbm9kZXMg dGhhdCBhcmUgb3V0IG9mIG9yZGVyLgogICovCi1zdGF0aWMgdm9pZCBzaHJpbmtfaWNhY2hlX3Ni KHN0cnVjdCBzdXBlcl9ibG9jayAqc2IsIGludCAqbnJfdG9fc2NhbikKK3ZvaWQgcHJ1bmVfaWNh Y2hlX3NiKHN0cnVjdCBzdXBlcl9ibG9jayAqc2IsIGludCBucl90b19zY2FuKQogewogCUxJU1Rf SEVBRChmcmVlYWJsZSk7CiAJaW50IG5yX3NjYW5uZWQ7CiAJdW5zaWduZWQgbG9uZyByZWFwID0g MDsKIAorCWRvd25fcmVhZCgmaXBydW5lX3NlbSk7CiAJc3Bpbl9sb2NrKCZzYi0+c19pbm9kZV9s cnVfbG9jayk7Ci0JZm9yIChucl9zY2FubmVkID0gKm5yX3RvX3NjYW47IG5yX3NjYW5uZWQgPj0g MDsgbnJfc2Nhbm5lZC0tKSB7CisJZm9yIChucl9zY2FubmVkID0gbnJfdG9fc2NhbjsgbnJfc2Nh bm5lZCA+PSAwOyBucl9zY2FubmVkLS0pIHsKIAkJc3RydWN0IGlub2RlICppbm9kZTsKIAogCQlp ZiAobGlzdF9lbXB0eSgmc2ItPnNfaW5vZGVfbHJ1KSkKQEAgLTcwMCwxMTEgKzcwMiwxMSBAQCBz dGF0aWMgdm9pZCBzaHJpbmtfaWNhY2hlX3NiKHN0cnVjdCBzdXBlcl9ibG9jayAqc2IsIGludCAq bnJfdG9fc2NhbikKIAllbHNlCiAJCV9fY291bnRfdm1fZXZlbnRzKFBHSU5PREVTVEVBTCwgcmVh cCk7CiAJc3Bpbl91bmxvY2soJnNiLT5zX2lub2RlX2xydV9sb2NrKTsKLQkqbnJfdG9fc2NhbiA9 IG5yX3NjYW5uZWQ7CiAKIAlkaXNwb3NlX2xpc3QoJmZyZWVhYmxlKTsKLX0KLQotc3RhdGljIHZv aWQgcHJ1bmVfaWNhY2hlKGludCBjb3VudCkKLXsKLQlzdHJ1Y3Qgc3VwZXJfYmxvY2sgKnNiLCAq cCA9IE5VTEw7Ci0JaW50IHdfY291bnQ7Ci0JaW50IHVudXNlZCA9IGlub2Rlc19zdGF0Lm5yX3Vu dXNlZDsKLQlpbnQgcHJ1bmVfcmF0aW87Ci0JaW50IHBydW5lZDsKLQotCWlmICh1bnVzZWQgPT0g MCB8fCBjb3VudCA9PSAwKQotCQlyZXR1cm47Ci0JZG93bl9yZWFkKCZpcHJ1bmVfc2VtKTsKLQlp ZiAoY291bnQgPj0gdW51c2VkKQotCQlwcnVuZV9yYXRpbyA9IDE7Ci0JZWxzZQotCQlwcnVuZV9y YXRpbyA9IHVudXNlZCAvIGNvdW50OwotCXNwaW5fbG9jaygmc2JfbG9jayk7Ci0JbGlzdF9mb3Jf ZWFjaF9lbnRyeShzYiwgJnN1cGVyX2Jsb2Nrcywgc19saXN0KSB7Ci0JCWlmIChsaXN0X2VtcHR5 KCZzYi0+c19pbnN0YW5jZXMpKQotCQkJY29udGludWU7Ci0JCWlmIChzYi0+c19ucl9pbm9kZXNf dW51c2VkID09IDApCi0JCQljb250aW51ZTsKLQkJc2ItPnNfY291bnQrKzsKLQkJLyogTm93LCB3 ZSByZWNsYWltIHVudXNlZCBkZW50cmlucyB3aXRoIGZhaXJuZXNzLgotCQkgKiBXZSByZWNsYWlt IHRoZW0gc2FtZSBwZXJjZW50YWdlIGZyb20gZWFjaCBzdXBlcmJsb2NrLgotCQkgKiBXZSBjYWxj dWxhdGUgbnVtYmVyIG9mIGRlbnRyaWVzIHRvIHNjYW4gb24gdGhpcyBzYgotCQkgKiBhcyBmb2xs b3dzLCBidXQgdGhlIGltcGxlbWVudGF0aW9uIGlzIGFycmFuZ2VkIHRvIGF2b2lkCi0JCSAqIG92 ZXJmbG93czoKLQkJICogbnVtYmVyIG9mIGRlbnRyaWVzIHRvIHNjYW4gb24gdGhpcyBzYiA9Ci0J CSAqIGNvdW50ICogKG51bWJlciBvZiBkZW50cmllcyBvbiB0aGlzIHNiIC8KLQkJICogbnVtYmVy IG9mIGRlbnRyaWVzIGluIHRoZSBtYWNoaW5lKQotCQkgKi8KLQkJc3Bpbl91bmxvY2soJnNiX2xv Y2spOwotCQlpZiAocHJ1bmVfcmF0aW8gIT0gMSkKLQkJCXdfY291bnQgPSAoc2ItPnNfbnJfaW5v ZGVzX3VudXNlZCAvIHBydW5lX3JhdGlvKSArIDE7Ci0JCWVsc2UKLQkJCXdfY291bnQgPSBzYi0+ c19ucl9pbm9kZXNfdW51c2VkOwotCQlwcnVuZWQgPSB3X2NvdW50OwotCQkvKgotCQkgKiBXZSBu ZWVkIHRvIGJlIHN1cmUgdGhpcyBmaWxlc3lzdGVtIGlzbid0IGJlaW5nIHVubW91bnRlZCwKLQkJ ICogb3RoZXJ3aXNlIHdlIGNvdWxkIHJhY2Ugd2l0aCBnZW5lcmljX3NodXRkb3duX3N1cGVyKCks IGFuZAotCQkgKiBlbmQgdXAgaG9sZGluZyBhIHJlZmVyZW5jZSB0byBhbiBpbm9kZSB3aGlsZSB0 aGUgZmlsZXN5c3RlbQotCQkgKiBpcyB1bm1vdW50ZWQuICBTbyB3ZSB0cnkgdG8gZ2V0IHNfdW1v dW50LCBhbmQgbWFrZSBzdXJlCi0JCSAqIHNfcm9vdCBpc24ndCBOVUxMLgotCQkgKi8KLQkJaWYg KGRvd25fcmVhZF90cnlsb2NrKCZzYi0+c191bW91bnQpKSB7Ci0JCQlpZiAoKHNiLT5zX3Jvb3Qg IT0gTlVMTCkgJiYKLQkJCSAgICAoIWxpc3RfZW1wdHkoJnNiLT5zX2RlbnRyeV9scnUpKSkgewot CQkJCXNocmlua19pY2FjaGVfc2Ioc2IsICZ3X2NvdW50KTsKLQkJCQlwcnVuZWQgLT0gd19jb3Vu dDsKLQkJCX0KLQkJCXVwX3JlYWQoJnNiLT5zX3Vtb3VudCk7Ci0JCX0KLQkJc3Bpbl9sb2NrKCZz Yl9sb2NrKTsKLQkJaWYgKHApCi0JCQlfX3B1dF9zdXBlcihwKTsKLQkJY291bnQgLT0gcHJ1bmVk OwotCQlwID0gc2I7Ci0JCS8qIG1vcmUgd29yayBsZWZ0IHRvIGRvPyAqLwotCQlpZiAoY291bnQg PD0gMCkKLQkJCWJyZWFrOwotCX0KLQlpZiAocCkKLQkJX19wdXRfc3VwZXIocCk7Ci0Jc3Bpbl91 bmxvY2soJnNiX2xvY2spOwogCXVwX3JlYWQoJmlwcnVuZV9zZW0pOwogfQogCi0vKgotICogc2hy aW5rX2ljYWNoZV9tZW1vcnkoKSB3aWxsIGF0dGVtcHQgdG8gcmVjbGFpbSBzb21lIHVudXNlZCBp bm9kZXMuICBIZXJlLAotICogInVudXNlZCIgbWVhbnMgdGhhdCBubyBkZW50cmllcyBhcmUgcmVm ZXJyaW5nIHRvIHRoZSBpbm9kZXM6IHRoZSBmaWxlcyBhcmUKLSAqIG5vdCBvcGVuIGFuZCB0aGUg ZGNhY2hlIHJlZmVyZW5jZXMgdG8gdGhvc2UgaW5vZGVzIGhhdmUgYWxyZWFkeSBiZWVuCi0gKiBy ZWNsYWltZWQuCi0gKgotICogVGhpcyBmdW5jdGlvbiBpcyBwYXNzZWQgdGhlIG51bWJlciBvZiBp bm9kZXMgdG8gc2NhbiwgYW5kIGl0IHJldHVybnMgdGhlCi0gKiB0b3RhbCBudW1iZXIgb2YgcmVt YWluaW5nIHBvc3NpYmx5LXJlY2xhaW1hYmxlIGlub2Rlcy4KLSAqLwotc3RhdGljIGludCBzaHJp bmtfaWNhY2hlX21lbW9yeShzdHJ1Y3Qgc2hyaW5rZXIgKnNocmluaywKLQkJCQlzdHJ1Y3Qgc2hy aW5rX2NvbnRyb2wgKnNjKQotewotCWludCBuciA9IHNjLT5ucl90b19zY2FuOwotCWdmcF90IGdm cF9tYXNrID0gc2MtPmdmcF9tYXNrOwotCi0JaWYgKG5yKSB7Ci0JCS8qCi0JCSAqIE5hc3R5IGRl YWRsb2NrIGF2b2lkYW5jZS4gIFdlIG1heSBob2xkIHZhcmlvdXMgRlMgbG9ja3MsCi0JCSAqIGFu ZCB3ZSBkb24ndCB3YW50IHRvIHJlY3Vyc2UgaW50byB0aGUgRlMgdGhhdCBjYWxsZWQgdXMKLQkJ ICogaW4gY2xlYXJfaW5vZGUoKSBhbmQgZnJpZW5kcy4uCi0JCSAqLwotCQlpZiAoIShnZnBfbWFz ayAmIF9fR0ZQX0ZTKSkKLQkJCXJldHVybiAtMTsKLQkJcHJ1bmVfaWNhY2hlKG5yKTsKLQl9Ci0J cmV0dXJuIChnZXRfbnJfaW5vZGVzX3VudXNlZCgpIC8gMTAwKSAqIHN5c2N0bF92ZnNfY2FjaGVf cHJlc3N1cmU7Ci19Ci0KLXN0YXRpYyBzdHJ1Y3Qgc2hyaW5rZXIgaWNhY2hlX3Nocmlua2VyID0g ewotCS5zaHJpbmsgPSBzaHJpbmtfaWNhY2hlX21lbW9yeSwKLQkuc2Vla3MgPSBERUZBVUxUX1NF RUtTLAotfTsKLQogc3RhdGljIHZvaWQgX193YWl0X29uX2ZyZWVpbmdfaW5vZGUoc3RydWN0IGlu b2RlICppbm9kZSk7CiAvKgogICogQ2FsbGVkIHdpdGggdGhlIGlub2RlIGxvY2sgaGVsZC4KQEAg LTE2ODQsNyArMTU4Niw2IEBAIHZvaWQgX19pbml0IGlub2RlX2luaXQodm9pZCkKIAkJCQkJIChT TEFCX1JFQ0xBSU1fQUNDT1VOVHxTTEFCX1BBTklDfAogCQkJCQkgU0xBQl9NRU1fU1BSRUFEKSwK IAkJCQkJIGluaXRfb25jZSk7Ci0JcmVnaXN0ZXJfc2hyaW5rZXIoJmljYWNoZV9zaHJpbmtlcik7 CiAKIAkvKiBIYXNoIG1heSBoYXZlIGJlZW4gc2V0IHVwIGluIGlub2RlX2luaXRfZWFybHkgKi8K IAlpZiAoIWhhc2hkaXN0KQpkaWZmIC0tZ2l0IGEvZnMvc3VwZXIuYyBiL2ZzL3N1cGVyLmMKaW5k ZXggOWMzZmExZi4uZjQ2MzBkOSAxMDA2NDQKLS0tIGEvZnMvc3VwZXIuYworKysgYi9mcy9zdXBl ci5jCkBAIC0zOCw2ICszOCw1MCBAQAogTElTVF9IRUFEKHN1cGVyX2Jsb2Nrcyk7CiBERUZJTkVf U1BJTkxPQ0soc2JfbG9jayk7CiAKK3N0YXRpYyBpbnQgcHJ1bmVfc3VwZXIoc3RydWN0IHNocmlu a2VyICpzaHJpbmssIHN0cnVjdCBzaHJpbmtfY29udHJvbCAqc2MpCit7CisJc3RydWN0IHN1cGVy X2Jsb2NrICpzYjsKKwlpbnQgY291bnQ7CisKKwlzYiA9IGNvbnRhaW5lcl9vZihzaHJpbmssIHN0 cnVjdCBzdXBlcl9ibG9jaywgc19zaHJpbmspOworCisJLyoKKwkgKiBEZWFkbG9jayBhdm9pZGFu Y2UuICBXZSBtYXkgaG9sZCB2YXJpb3VzIEZTIGxvY2tzLCBhbmQgd2UgZG9uJ3Qgd2FudAorCSAq IHRvIHJlY3Vyc2UgaW50byB0aGUgRlMgdGhhdCBjYWxsZWQgdXMgaW4gY2xlYXJfaW5vZGUoKSBh bmQgZnJpZW5kcy4uCisJICovCisJaWYgKHNjLT5ucl90b19zY2FuICYmICEoc2MtPmdmcF9tYXNr ICYgX19HRlBfRlMpKQorCQlyZXR1cm4gLTE7CisKKwkvKgorCSAqIGlmIHdlIGNhbid0IGdldCB0 aGUgdW1vdW50IGxvY2ssIHRoZW4gdGhlcmUncyBubyBwb2ludCBoYXZpbmcgdGhlCisJICogc2hy aW5rZXIgdHJ5IGFnYWluIGJlY2F1c2UgdGhlIHNiIGlzIGJlaW5nIHRvcm4gZG93bi4KKwkgKi8K KwlpZiAoIWRvd25fcmVhZF90cnlsb2NrKCZzYi0+c191bW91bnQpKQorCQlyZXR1cm4gLTE7CisK KwlpZiAoIXNiLT5zX3Jvb3QpIHsKKwkJdXBfcmVhZCgmc2ItPnNfdW1vdW50KTsKKwkJcmV0dXJu IC0xOworCX0KKworCWlmIChzYy0+bnJfdG9fc2NhbikgeworCQkvKiBwcm9wb3J0aW9uIHRoZSBz Y2FuIGJldHdlZW4gdGhlIHR3byBjYWNoZdGVICovCisJCWludCB0b3RhbDsKKworCQl0b3RhbCA9 IHNiLT5zX25yX2RlbnRyeV91bnVzZWQgKyBzYi0+c19ucl9pbm9kZXNfdW51c2VkICsgMTsKKwkJ Y291bnQgPSAoc2MtPm5yX3RvX3NjYW4gKiBzYi0+c19ucl9kZW50cnlfdW51c2VkKSAvIHRvdGFs OworCisJCS8qIHBydW5lIGRjYWNoZSBmaXJzdCBhcyBpY2FjaGUgaXMgcGlubmVkIGJ5IGl0ICov CisJCXBydW5lX2RjYWNoZV9zYihzYiwgY291bnQpOworCQlwcnVuZV9pY2FjaGVfc2Ioc2IsIHNj LT5ucl90b19zY2FuIC0gY291bnQpOworCX0KKworCWNvdW50ID0gKChzYi0+c19ucl9kZW50cnlf dW51c2VkICsgc2ItPnNfbnJfaW5vZGVzX3VudXNlZCkgLyAxMDApCisJCQkJCQkqIHN5c2N0bF92 ZnNfY2FjaGVfcHJlc3N1cmU7CisJdXBfcmVhZCgmc2ItPnNfdW1vdW50KTsKKwlyZXR1cm4gY291 bnQ7Cit9CisKIC8qKgogICoJYWxsb2Nfc3VwZXIJLQljcmVhdGUgbmV3IHN1cGVyYmxvY2sKICAq CUB0eXBlOglmaWxlc3lzdGVtIHR5cGUgc3VwZXJibG9jayBzaG91bGQgYmVsb25nIHRvCkBAIC0x MTYsNiArMTYwLDkgQEAgc3RhdGljIHN0cnVjdCBzdXBlcl9ibG9jayAqYWxsb2Nfc3VwZXIoc3Ry dWN0IGZpbGVfc3lzdGVtX3R5cGUgKnR5cGUpCiAJCXMtPnNfb3AgPSAmZGVmYXVsdF9vcDsKIAkJ cy0+c190aW1lX2dyYW4gPSAxMDAwMDAwMDAwOwogCQlzLT5jbGVhbmNhY2hlX3Bvb2xpZCA9IC0x OworCisJCXMtPnNfc2hyaW5rLnNlZWtzID0gREVGQVVMVF9TRUVLUzsKKwkJcy0+c19zaHJpbmsu c2hyaW5rID0gcHJ1bmVfc3VwZXI7CiAJfQogb3V0OgogCXJldHVybiBzOwpAQCAtMjc4LDcgKzMy NSwxMiBAQCB2b2lkIGdlbmVyaWNfc2h1dGRvd25fc3VwZXIoc3RydWN0IHN1cGVyX2Jsb2NrICpz YikKIHsKIAljb25zdCBzdHJ1Y3Qgc3VwZXJfb3BlcmF0aW9ucyAqc29wID0gc2ItPnNfb3A7CiAK LQorCS8qCisJICogc2h1dCBkb3duIHRoZSBzaHJpbmtlciBmaXJzdCBzbyB3ZSBrbm93IHRoYXQg dGhlcmUgYXJlIG5vIHBvc3NpYmxlCisJICogcmFjZXMgd2hlbiBzaHJpbmtpbmcgdGhlIGRjYWNo ZSBvciBpY2FjaGUuIFJlbW92ZXMgdGhlIG5lZWQgZm9yCisJICogZXh0ZXJuYWwgbG9ja2luZyB0 byBwcmV2ZW50IHN1Y2ggcmFjZXMuCisJICovCisJdW5yZWdpc3Rlcl9zaHJpbmtlcigmc2ItPnNf c2hyaW5rKTsKIAlpZiAoc2ItPnNfcm9vdCkgewogCQlzaHJpbmtfZGNhY2hlX2Zvcl91bW91bnQo c2IpOwogCQlzeW5jX2ZpbGVzeXN0ZW0oc2IpOwpAQCAtMzY2LDYgKzQxOCw3IEBAIHJldHJ5Ogog CWxpc3RfYWRkKCZzLT5zX2luc3RhbmNlcywgJnR5cGUtPmZzX3N1cGVycyk7CiAJc3Bpbl91bmxv Y2soJnNiX2xvY2spOwogCWdldF9maWxlc3lzdGVtKHR5cGUpOworCXJlZ2lzdGVyX3Nocmlua2Vy KCZzLT5zX3Nocmluayk7CiAJcmV0dXJuIHM7CiB9CiAKZGlmZiAtLWdpdCBhL2luY2x1ZGUvbGlu dXgvZnMuaCBiL2luY2x1ZGUvbGludXgvZnMuaAppbmRleCBiYmQ0NzhlLi5jM2IzNDYyIDEwMDY0 NAotLS0gYS9pbmNsdWRlL2xpbnV4L2ZzLmgKKysrIGIvaW5jbHVkZS9saW51eC9mcy5oCkBAIC0z OTEsNiArMzkxLDcgQEAgc3RydWN0IGlub2Rlc19zdGF0X3QgewogI2luY2x1ZGUgPGxpbnV4L3Nl bWFwaG9yZS5oPgogI2luY2x1ZGUgPGxpbnV4L2ZpZW1hcC5oPgogI2luY2x1ZGUgPGxpbnV4L3Jj dWxpc3RfYmwuaD4KKyNpbmNsdWRlIDxsaW51eC9tbS5oPgogCiAjaW5jbHVkZSA8YXNtL2F0b21p Yy5oPgogI2luY2x1ZGUgPGFzbS9ieXRlb3JkZXIuaD4KQEAgLTE0NDAsOCArMTQ0MSwxNCBAQCBz dHJ1Y3Qgc3VwZXJfYmxvY2sgewogCSAqIFNhdmVkIHBvb2wgaWRlbnRpZmllciBmb3IgY2xlYW5j YWNoZSAoLTEgbWVhbnMgbm9uZSkKIAkgKi8KIAlpbnQgY2xlYW5jYWNoZV9wb29saWQ7CisKKwlz dHJ1Y3Qgc2hyaW5rZXIgc19zaHJpbms7CS8qIHBlci1zYiBzaHJpbmtlciBoYW5kbGUgKi8KIH07 CiAKKy8qIHN1cGVyYmxvY2sgY2FjaGUgcHJ1bmluZyBmdW5jdGlvbnMgKi8KK2V4dGVybiB2b2lk IHBydW5lX2ljYWNoZV9zYihzdHJ1Y3Qgc3VwZXJfYmxvY2sgKnNiLCBpbnQgbnJfdG9fc2Nhbik7 CitleHRlcm4gdm9pZCBwcnVuZV9kY2FjaGVfc2Ioc3RydWN0IHN1cGVyX2Jsb2NrICpzYiwgaW50 IG5yX3RvX3NjYW4pOworCiBleHRlcm4gc3RydWN0IHRpbWVzcGVjIGN1cnJlbnRfZnNfdGltZShz dHJ1Y3Qgc3VwZXJfYmxvY2sgKnNiKTsKIAogLyoKLS0gCjEuNy41LjEKCl9fX19fX19fX19fX19f X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fCnhmcyBtYWlsaW5nIGxpc3QKeGZzQG9z cy5zZ2kuY29tCmh0dHA6Ly9vc3Muc2dpLmNvbS9tYWlsbWFuL2xpc3RpbmZvL3hmcwo= From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S933322Ab1FBHC5 (ORCPT ); Thu, 2 Jun 2011 03:02:57 -0400 Received: from ipmail06.adl2.internode.on.net ([150.101.137.129]:27367 "EHLO ipmail06.adl2.internode.on.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S933020Ab1FBHBt (ORCPT ); Thu, 2 Jun 2011 03:01:49 -0400 X-IronPort-Anti-Spam-Filtered: true X-IronPort-Anti-Spam-Result: ArYDAAky5015LCoegWdsb2JhbABThEmhZxUBARYmJbZukGKBK4NsgQoEoCk From: Dave Chinner To: linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com Subject: =?UTF-8?q?=5BPATCH=2008/12=5D=20superblock=3A=20introduce=20per-sb=20cache=20shrinker=20infrastructure?= Date: Thu, 2 Jun 2011 17:01:03 +1000 Message-Id: <1306998067-27659-9-git-send-email-david@fromorbit.com> X-Mailer: git-send-email 1.7.5.1 In-Reply-To: <1306998067-27659-1-git-send-email-david@fromorbit.com> References: <1306998067-27659-1-git-send-email-david@fromorbit.com> MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org From: Dave Chinner With context based shrinkers, we can implement a per-superblock shrinker that shrinks the caches attached to the superblock. We currently have global shrinkers for the inode and dentry caches that split up into per-superblock operations via a coarse proportioning method that does not batch very well. The global shrinkers also have a dependency - dentries pin inodes - so we have to be very careful about how we register the global shrinkers so that the implicit call order is always correct. With a per-sb shrinker callout, we can encode this dependency directly into the per-sb shrinker, hence avoiding the need for strictly ordering shrinker registrations. We also have no need for any proportioning code for the shrinker subsystem already provides this functionality across all shrinkers. Allowing the shrinker to operate on a single superblock at a time means that we do less superblock list traversals and locking and reclaim should batch more effectively. This should result in less CPU overhead for reclaim and potentially faster reclaim of items from each filesystem. Signed-off-by: Dave Chinner --- fs/dcache.c | 121 +++++---------------------------------------------- fs/inode.c | 117 ++++---------------------------------------------- fs/super.c | 55 +++++++++++++++++++++++- include/linux/fs.h | 7 +++ 4 files changed, 82 insertions(+), 218 deletions(-) diff --git a/fs/dcache.c b/fs/dcache.c index 37f72ee..f73ef23 100644 --- a/fs/dcache.c +++ b/fs/dcache.c @@ -720,13 +720,11 @@ static void shrink_dentry_list(struct list_head *list) * * If flags contains DCACHE_REFERENCED reference dentries will not be pruned. */ -static void __shrink_dcache_sb(struct super_block *sb, int *count, int flags) +static void __shrink_dcache_sb(struct super_block *sb, int count, int flags) { - /* called from prune_dcache() and shrink_dcache_parent() */ struct dentry *dentry; LIST_HEAD(referenced); LIST_HEAD(tmp); - int cnt = *count; relock: spin_lock(&dcache_lru_lock); @@ -754,7 +752,7 @@ relock: } else { list_move_tail(&dentry->d_lru, &tmp); spin_unlock(&dentry->d_lock); - if (!--cnt) + if (!--count) break; } cond_resched_lock(&dcache_lru_lock); @@ -764,83 +762,22 @@ relock: spin_unlock(&dcache_lru_lock); shrink_dentry_list(&tmp); - - *count = cnt; } /** - * prune_dcache - shrink the dcache - * @count: number of entries to try to free + * prune_dcache_sb - shrink the dcache + * @nr_to_scan: number of entries to try to free * - * Shrink the dcache. This is done when we need more memory, or simply when we - * need to unmount something (at which point we need to unuse all dentries). + * Attempt to shrink the superblock dcache LRU by @nr_to_scan entries. This is + * done when we need more memory an called from the superblock shrinker + * function. * - * This function may fail to free any resources if all the dentries are in use. + * This function may fail to free any resources if all the dentries are in + * use. */ -static void prune_dcache(int count) +void prune_dcache_sb(struct super_block *sb, int nr_to_scan) { - struct super_block *sb, *p = NULL; - int w_count; - int unused = dentry_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused == 0 || count == 0) - return; - if (count >= unused) - prune_ratio = 1; - else - prune_ratio = unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_dentry_unused == 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb = - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio != 1) - w_count = (sb->s_nr_dentry_unused / prune_ratio) + 1; - else - w_count = sb->s_nr_dentry_unused; - pruned = w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root != NULL) && - (!list_empty(&sb->s_dentry_lru))) { - __shrink_dcache_sb(sb, &w_count, - DCACHE_REFERENCED); - pruned -= w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -= pruned; - p = sb; - /* more work left to do? */ - if (count <= 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); + __shrink_dcache_sb(sb, nr_to_scan, DCACHE_REFERENCED); } /** @@ -1215,42 +1152,10 @@ void shrink_dcache_parent(struct dentry * parent) int found; while ((found = select_parent(parent)) != 0) - __shrink_dcache_sb(sb, &found, 0); + __shrink_dcache_sb(sb, found, 0); } EXPORT_SYMBOL(shrink_dcache_parent); -/* - * Scan `sc->nr_slab_to_reclaim' dentries and return the number which remain. - * - * We need to avoid reentering the filesystem if the caller is performing a - * GFP_NOFS allocation attempt. One example deadlock is: - * - * ext2_new_block->getblk->GFP->shrink_dcache_memory->prune_dcache-> - * prune_one_dentry->dput->dentry_iput->iput->inode->i_sb->s_op->put_inode-> - * ext2_discard_prealloc->ext2_free_blocks->lock_super->DEADLOCK. - * - * In this case we return -1 to tell the caller that we baled. - */ -static int shrink_dcache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr = sc->nr_to_scan; - gfp_t gfp_mask = sc->gfp_mask; - - if (nr) { - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_dcache(nr); - } - - return (dentry_stat.nr_unused / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker dcache_shrinker = { - .shrink = shrink_dcache_memory, - .seeks = DEFAULT_SEEKS, -}; - /** * d_alloc - allocate a dcache entry * @parent: parent of entry to allocate @@ -3030,8 +2935,6 @@ static void __init dcache_init(void) */ dentry_cache = KMEM_CACHE(dentry, SLAB_RECLAIM_ACCOUNT|SLAB_PANIC|SLAB_MEM_SPREAD); - - register_shrinker(&dcache_shrinker); /* Hash may have been set up in dcache_init_early */ if (!hashdist) diff --git a/fs/inode.c b/fs/inode.c index 667a29c..890d95e 100644 --- a/fs/inode.c +++ b/fs/inode.c @@ -73,7 +73,7 @@ __cacheline_aligned_in_smp DEFINE_SPINLOCK(inode_wb_list_lock); * * We don't actually need it to protect anything in the umount path, * but only need to cycle through it to make sure any inode that - * prune_icache took off the LRU list has been fully torn down by the + * prune_icache_sb took off the LRU list has been fully torn down by the * time we are past evict_inodes. */ static DECLARE_RWSEM(iprune_sem); @@ -537,7 +537,7 @@ void evict_inodes(struct super_block *sb) dispose_list(&dispose); /* - * Cycle through iprune_sem to make sure any inode that prune_icache + * Cycle through iprune_sem to make sure any inode that prune_icache_sb * moved off the list before we took the lock has been fully torn * down. */ @@ -605,9 +605,10 @@ static int can_unuse(struct inode *inode) } /* - * Scan `goal' inodes on the unused list for freeable ones. They are moved to a - * temporary list and then are freed outside sb->s_inode_lru_lock by - * dispose_list(). + * Walk the superblock inode LRU for freeable inodes and attempt to free them. + * This is called from the superblock shrinker function with a number of inodes + * to trim from the LRU. Inodes to be freed are moved to a temporary list and + * then are freed outside inode_lock by dispose_list(). * * Any inodes which are pinned purely because of attached pagecache have their * pagecache removed. If the inode has metadata buffers attached to @@ -621,14 +622,15 @@ static int can_unuse(struct inode *inode) * LRU does not have strict ordering. Hence we don't want to reclaim inodes * with this flag set because they are the inodes that are out of order. */ -static void shrink_icache_sb(struct super_block *sb, int *nr_to_scan) +void prune_icache_sb(struct super_block *sb, int nr_to_scan) { LIST_HEAD(freeable); int nr_scanned; unsigned long reap = 0; + down_read(&iprune_sem); spin_lock(&sb->s_inode_lru_lock); - for (nr_scanned = *nr_to_scan; nr_scanned >= 0; nr_scanned--) { + for (nr_scanned = nr_to_scan; nr_scanned >= 0; nr_scanned--) { struct inode *inode; if (list_empty(&sb->s_inode_lru)) @@ -700,111 +702,11 @@ static void shrink_icache_sb(struct super_block *sb, int *nr_to_scan) else __count_vm_events(PGINODESTEAL, reap); spin_unlock(&sb->s_inode_lru_lock); - *nr_to_scan = nr_scanned; dispose_list(&freeable); -} - -static void prune_icache(int count) -{ - struct super_block *sb, *p = NULL; - int w_count; - int unused = inodes_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused == 0 || count == 0) - return; - down_read(&iprune_sem); - if (count >= unused) - prune_ratio = 1; - else - prune_ratio = unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_inodes_unused == 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb = - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio != 1) - w_count = (sb->s_nr_inodes_unused / prune_ratio) + 1; - else - w_count = sb->s_nr_inodes_unused; - pruned = w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root != NULL) && - (!list_empty(&sb->s_dentry_lru))) { - shrink_icache_sb(sb, &w_count); - pruned -= w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -= pruned; - p = sb; - /* more work left to do? */ - if (count <= 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); up_read(&iprune_sem); } -/* - * shrink_icache_memory() will attempt to reclaim some unused inodes. Here, - * "unused" means that no dentries are referring to the inodes: the files are - * not open and the dcache references to those inodes have already been - * reclaimed. - * - * This function is passed the number of inodes to scan, and it returns the - * total number of remaining possibly-reclaimable inodes. - */ -static int shrink_icache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr = sc->nr_to_scan; - gfp_t gfp_mask = sc->gfp_mask; - - if (nr) { - /* - * Nasty deadlock avoidance. We may hold various FS locks, - * and we don't want to recurse into the FS that called us - * in clear_inode() and friends.. - */ - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_icache(nr); - } - return (get_nr_inodes_unused() / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker icache_shrinker = { - .shrink = shrink_icache_memory, - .seeks = DEFAULT_SEEKS, -}; - static void __wait_on_freeing_inode(struct inode *inode); /* * Called with the inode lock held. @@ -1684,7 +1586,6 @@ void __init inode_init(void) (SLAB_RECLAIM_ACCOUNT|SLAB_PANIC| SLAB_MEM_SPREAD), init_once); - register_shrinker(&icache_shrinker); /* Hash may have been set up in inode_init_early */ if (!hashdist) diff --git a/fs/super.c b/fs/super.c index 9c3fa1f..f4630d9 100644 --- a/fs/super.c +++ b/fs/super.c @@ -38,6 +38,50 @@ LIST_HEAD(super_blocks); DEFINE_SPINLOCK(sb_lock); +static int prune_super(struct shrinker *shrink, struct shrink_control *sc) +{ + struct super_block *sb; + int count; + + sb = container_of(shrink, struct super_block, s_shrink); + + /* + * Deadlock avoidance. We may hold various FS locks, and we don't want + * to recurse into the FS that called us in clear_inode() and friends.. + */ + if (sc->nr_to_scan && !(sc->gfp_mask & __GFP_FS)) + return -1; + + /* + * if we can't get the umount lock, then there's no point having the + * shrinker try again because the sb is being torn down. + */ + if (!down_read_trylock(&sb->s_umount)) + return -1; + + if (!sb->s_root) { + up_read(&sb->s_umount); + return -1; + } + + if (sc->nr_to_scan) { + /* proportion the scan between the two cacheѕ */ + int total; + + total = sb->s_nr_dentry_unused + sb->s_nr_inodes_unused + 1; + count = (sc->nr_to_scan * sb->s_nr_dentry_unused) / total; + + /* prune dcache first as icache is pinned by it */ + prune_dcache_sb(sb, count); + prune_icache_sb(sb, sc->nr_to_scan - count); + } + + count = ((sb->s_nr_dentry_unused + sb->s_nr_inodes_unused) / 100) + * sysctl_vfs_cache_pressure; + up_read(&sb->s_umount); + return count; +} + /** * alloc_super - create new superblock * @type: filesystem type superblock should belong to @@ -116,6 +160,9 @@ static struct super_block *alloc_super(struct file_system_type *type) s->s_op = &default_op; s->s_time_gran = 1000000000; s->cleancache_poolid = -1; + + s->s_shrink.seeks = DEFAULT_SEEKS; + s->s_shrink.shrink = prune_super; } out: return s; @@ -278,7 +325,12 @@ void generic_shutdown_super(struct super_block *sb) { const struct super_operations *sop = sb->s_op; - + /* + * shut down the shrinker first so we know that there are no possible + * races when shrinking the dcache or icache. Removes the need for + * external locking to prevent such races. + */ + unregister_shrinker(&sb->s_shrink); if (sb->s_root) { shrink_dcache_for_umount(sb); sync_filesystem(sb); @@ -366,6 +418,7 @@ retry: list_add(&s->s_instances, &type->fs_supers); spin_unlock(&sb_lock); get_filesystem(type); + register_shrinker(&s->s_shrink); return s; } diff --git a/include/linux/fs.h b/include/linux/fs.h index bbd478e..c3b3462 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -391,6 +391,7 @@ struct inodes_stat_t { #include #include #include +#include #include #include @@ -1440,8 +1441,14 @@ struct super_block { * Saved pool identifier for cleancache (-1 means none) */ int cleancache_poolid; + + struct shrinker s_shrink; /* per-sb shrinker handle */ }; +/* superblock cache pruning functions */ +extern void prune_icache_sb(struct super_block *sb, int nr_to_scan); +extern void prune_dcache_sb(struct super_block *sb, int nr_to_scan); + extern struct timespec current_fs_time(struct super_block *sb); /* -- 1.7.5.1 From mboxrd@z Thu Jan 1 00:00:00 1970 From: Dave Chinner Subject: =?UTF-8?q?=5BPATCH=2008/12=5D=20superblock=3A=20introduce=20per-sb=20cache=20shrinker=20infrastructure?= Date: Thu, 2 Jun 2011 17:01:03 +1000 Message-ID: <1306998067-27659-9-git-send-email-david@fromorbit.com> References: <1306998067-27659-1-git-send-email-david@fromorbit.com> Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com To: linux-fsdevel@vger.kernel.org Return-path: In-Reply-To: <1306998067-27659-1-git-send-email-david@fromorbit.com> Sender: owner-linux-mm@kvack.org List-Id: linux-fsdevel.vger.kernel.org From: Dave Chinner With context based shrinkers, we can implement a per-superblock shrinker that shrinks the caches attached to the superblock. We currently have global shrinkers for the inode and dentry caches that split up into per-superblock operations via a coarse proportioning method that does not batch very well. The global shrinkers also have a dependency - dentries pin inodes - so we have to be very careful about how we register the global shrinkers so that the implicit call order is always correct. With a per-sb shrinker callout, we can encode this dependency directly into the per-sb shrinker, hence avoiding the need for strictly ordering shrinker registrations. We also have no need for any proportioning code for the shrinker subsystem already provides this functionality across all shrinkers. Allowing the shrinker to operate on a single superblock at a time means that we do less superblock list traversals and locking and reclaim should batch more effectively. This should result in less CPU overhead for reclaim and potentially faster reclaim of items from each filesystem. Signed-off-by: Dave Chinner --- fs/dcache.c | 121 +++++-----------------------------------------= ----- fs/inode.c | 117 ++++------------------------------------------= ---- fs/super.c | 55 +++++++++++++++++++++++- include/linux/fs.h | 7 +++ 4 files changed, 82 insertions(+), 218 deletions(-) diff --git a/fs/dcache.c b/fs/dcache.c index 37f72ee..f73ef23 100644 --- a/fs/dcache.c +++ b/fs/dcache.c @@ -720,13 +720,11 @@ static void shrink_dentry_list(struct list_head *li= st) * * If flags contains DCACHE_REFERENCED reference dentries will not be pr= uned. */ -static void __shrink_dcache_sb(struct super_block *sb, int *count, int f= lags) +static void __shrink_dcache_sb(struct super_block *sb, int count, int fl= ags) { - /* called from prune_dcache() and shrink_dcache_parent() */ struct dentry *dentry; LIST_HEAD(referenced); LIST_HEAD(tmp); - int cnt =3D *count; =20 relock: spin_lock(&dcache_lru_lock); @@ -754,7 +752,7 @@ relock: } else { list_move_tail(&dentry->d_lru, &tmp); spin_unlock(&dentry->d_lock); - if (!--cnt) + if (!--count) break; } cond_resched_lock(&dcache_lru_lock); @@ -764,83 +762,22 @@ relock: spin_unlock(&dcache_lru_lock); =20 shrink_dentry_list(&tmp); - - *count =3D cnt; } =20 /** - * prune_dcache - shrink the dcache - * @count: number of entries to try to free + * prune_dcache_sb - shrink the dcache + * @nr_to_scan: number of entries to try to free * - * Shrink the dcache. This is done when we need more memory, or simply w= hen we - * need to unmount something (at which point we need to unuse all dentri= es). + * Attempt to shrink the superblock dcache LRU by @nr_to_scan entries. T= his is + * done when we need more memory an called from the superblock shrinker + * function. * - * This function may fail to free any resources if all the dentries are = in use. + * This function may fail to free any resources if all the dentries are = in + * use. */ -static void prune_dcache(int count) +void prune_dcache_sb(struct super_block *sb, int nr_to_scan) { - struct super_block *sb, *p =3D NULL; - int w_count; - int unused =3D dentry_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused =3D=3D 0 || count =3D=3D 0) - return; - if (count >=3D unused) - prune_ratio =3D 1; - else - prune_ratio =3D unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_dentry_unused =3D=3D 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb =3D - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio !=3D 1) - w_count =3D (sb->s_nr_dentry_unused / prune_ratio) + 1; - else - w_count =3D sb->s_nr_dentry_unused; - pruned =3D w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root !=3D NULL) && - (!list_empty(&sb->s_dentry_lru))) { - __shrink_dcache_sb(sb, &w_count, - DCACHE_REFERENCED); - pruned -=3D w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -=3D pruned; - p =3D sb; - /* more work left to do? */ - if (count <=3D 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); + __shrink_dcache_sb(sb, nr_to_scan, DCACHE_REFERENCED); } =20 /** @@ -1215,42 +1152,10 @@ void shrink_dcache_parent(struct dentry * parent) int found; =20 while ((found =3D select_parent(parent)) !=3D 0) - __shrink_dcache_sb(sb, &found, 0); + __shrink_dcache_sb(sb, found, 0); } EXPORT_SYMBOL(shrink_dcache_parent); =20 -/* - * Scan `sc->nr_slab_to_reclaim' dentries and return the number which re= main. - * - * We need to avoid reentering the filesystem if the caller is performin= g a - * GFP_NOFS allocation attempt. One example deadlock is: - * - * ext2_new_block->getblk->GFP->shrink_dcache_memory->prune_dcache-> - * prune_one_dentry->dput->dentry_iput->iput->inode->i_sb->s_op->put_ino= de-> - * ext2_discard_prealloc->ext2_free_blocks->lock_super->DEADLOCK. - * - * In this case we return -1 to tell the caller that we baled. - */ -static int shrink_dcache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr =3D sc->nr_to_scan; - gfp_t gfp_mask =3D sc->gfp_mask; - - if (nr) { - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_dcache(nr); - } - - return (dentry_stat.nr_unused / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker dcache_shrinker =3D { - .shrink =3D shrink_dcache_memory, - .seeks =3D DEFAULT_SEEKS, -}; - /** * d_alloc - allocate a dcache entry * @parent: parent of entry to allocate @@ -3030,8 +2935,6 @@ static void __init dcache_init(void) */ dentry_cache =3D KMEM_CACHE(dentry, SLAB_RECLAIM_ACCOUNT|SLAB_PANIC|SLAB_MEM_SPREAD); -=09 - register_shrinker(&dcache_shrinker); =20 /* Hash may have been set up in dcache_init_early */ if (!hashdist) diff --git a/fs/inode.c b/fs/inode.c index 667a29c..890d95e 100644 --- a/fs/inode.c +++ b/fs/inode.c @@ -73,7 +73,7 @@ __cacheline_aligned_in_smp DEFINE_SPINLOCK(inode_wb_lis= t_lock); * * We don't actually need it to protect anything in the umount path, * but only need to cycle through it to make sure any inode that - * prune_icache took off the LRU list has been fully torn down by the + * prune_icache_sb took off the LRU list has been fully torn down by the * time we are past evict_inodes. */ static DECLARE_RWSEM(iprune_sem); @@ -537,7 +537,7 @@ void evict_inodes(struct super_block *sb) dispose_list(&dispose); =20 /* - * Cycle through iprune_sem to make sure any inode that prune_icache + * Cycle through iprune_sem to make sure any inode that prune_icache_sb * moved off the list before we took the lock has been fully torn * down. */ @@ -605,9 +605,10 @@ static int can_unuse(struct inode *inode) } =20 /* - * Scan `goal' inodes on the unused list for freeable ones. They are mov= ed to a - * temporary list and then are freed outside sb->s_inode_lru_lock by - * dispose_list(). + * Walk the superblock inode LRU for freeable inodes and attempt to free= them. + * This is called from the superblock shrinker function with a number of= inodes + * to trim from the LRU. Inodes to be freed are moved to a temporary lis= t and + * then are freed outside inode_lock by dispose_list(). * * Any inodes which are pinned purely because of attached pagecache have= their * pagecache removed. If the inode has metadata buffers attached to @@ -621,14 +622,15 @@ static int can_unuse(struct inode *inode) * LRU does not have strict ordering. Hence we don't want to reclaim ino= des * with this flag set because they are the inodes that are out of order. */ -static void shrink_icache_sb(struct super_block *sb, int *nr_to_scan) +void prune_icache_sb(struct super_block *sb, int nr_to_scan) { LIST_HEAD(freeable); int nr_scanned; unsigned long reap =3D 0; =20 + down_read(&iprune_sem); spin_lock(&sb->s_inode_lru_lock); - for (nr_scanned =3D *nr_to_scan; nr_scanned >=3D 0; nr_scanned--) { + for (nr_scanned =3D nr_to_scan; nr_scanned >=3D 0; nr_scanned--) { struct inode *inode; =20 if (list_empty(&sb->s_inode_lru)) @@ -700,111 +702,11 @@ static void shrink_icache_sb(struct super_block *s= b, int *nr_to_scan) else __count_vm_events(PGINODESTEAL, reap); spin_unlock(&sb->s_inode_lru_lock); - *nr_to_scan =3D nr_scanned; =20 dispose_list(&freeable); -} - -static void prune_icache(int count) -{ - struct super_block *sb, *p =3D NULL; - int w_count; - int unused =3D inodes_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused =3D=3D 0 || count =3D=3D 0) - return; - down_read(&iprune_sem); - if (count >=3D unused) - prune_ratio =3D 1; - else - prune_ratio =3D unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_inodes_unused =3D=3D 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb =3D - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio !=3D 1) - w_count =3D (sb->s_nr_inodes_unused / prune_ratio) + 1; - else - w_count =3D sb->s_nr_inodes_unused; - pruned =3D w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root !=3D NULL) && - (!list_empty(&sb->s_dentry_lru))) { - shrink_icache_sb(sb, &w_count); - pruned -=3D w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -=3D pruned; - p =3D sb; - /* more work left to do? */ - if (count <=3D 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); up_read(&iprune_sem); } =20 -/* - * shrink_icache_memory() will attempt to reclaim some unused inodes. H= ere, - * "unused" means that no dentries are referring to the inodes: the file= s are - * not open and the dcache references to those inodes have already been - * reclaimed. - * - * This function is passed the number of inodes to scan, and it returns = the - * total number of remaining possibly-reclaimable inodes. - */ -static int shrink_icache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr =3D sc->nr_to_scan; - gfp_t gfp_mask =3D sc->gfp_mask; - - if (nr) { - /* - * Nasty deadlock avoidance. We may hold various FS locks, - * and we don't want to recurse into the FS that called us - * in clear_inode() and friends.. - */ - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_icache(nr); - } - return (get_nr_inodes_unused() / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker icache_shrinker =3D { - .shrink =3D shrink_icache_memory, - .seeks =3D DEFAULT_SEEKS, -}; - static void __wait_on_freeing_inode(struct inode *inode); /* * Called with the inode lock held. @@ -1684,7 +1586,6 @@ void __init inode_init(void) (SLAB_RECLAIM_ACCOUNT|SLAB_PANIC| SLAB_MEM_SPREAD), init_once); - register_shrinker(&icache_shrinker); =20 /* Hash may have been set up in inode_init_early */ if (!hashdist) diff --git a/fs/super.c b/fs/super.c index 9c3fa1f..f4630d9 100644 --- a/fs/super.c +++ b/fs/super.c @@ -38,6 +38,50 @@ LIST_HEAD(super_blocks); DEFINE_SPINLOCK(sb_lock); =20 +static int prune_super(struct shrinker *shrink, struct shrink_control *s= c) +{ + struct super_block *sb; + int count; + + sb =3D container_of(shrink, struct super_block, s_shrink); + + /* + * Deadlock avoidance. We may hold various FS locks, and we don't want + * to recurse into the FS that called us in clear_inode() and friends.. + */ + if (sc->nr_to_scan && !(sc->gfp_mask & __GFP_FS)) + return -1; + + /* + * if we can't get the umount lock, then there's no point having the + * shrinker try again because the sb is being torn down. + */ + if (!down_read_trylock(&sb->s_umount)) + return -1; + + if (!sb->s_root) { + up_read(&sb->s_umount); + return -1; + } + + if (sc->nr_to_scan) { + /* proportion the scan between the two cache=D1=95 */ + int total; + + total =3D sb->s_nr_dentry_unused + sb->s_nr_inodes_unused + 1; + count =3D (sc->nr_to_scan * sb->s_nr_dentry_unused) / total; + + /* prune dcache first as icache is pinned by it */ + prune_dcache_sb(sb, count); + prune_icache_sb(sb, sc->nr_to_scan - count); + } + + count =3D ((sb->s_nr_dentry_unused + sb->s_nr_inodes_unused) / 100) + * sysctl_vfs_cache_pressure; + up_read(&sb->s_umount); + return count; +} + /** * alloc_super - create new superblock * @type: filesystem type superblock should belong to @@ -116,6 +160,9 @@ static struct super_block *alloc_super(struct file_sy= stem_type *type) s->s_op =3D &default_op; s->s_time_gran =3D 1000000000; s->cleancache_poolid =3D -1; + + s->s_shrink.seeks =3D DEFAULT_SEEKS; + s->s_shrink.shrink =3D prune_super; } out: return s; @@ -278,7 +325,12 @@ void generic_shutdown_super(struct super_block *sb) { const struct super_operations *sop =3D sb->s_op; =20 - + /* + * shut down the shrinker first so we know that there are no possible + * races when shrinking the dcache or icache. Removes the need for + * external locking to prevent such races. + */ + unregister_shrinker(&sb->s_shrink); if (sb->s_root) { shrink_dcache_for_umount(sb); sync_filesystem(sb); @@ -366,6 +418,7 @@ retry: list_add(&s->s_instances, &type->fs_supers); spin_unlock(&sb_lock); get_filesystem(type); + register_shrinker(&s->s_shrink); return s; } =20 diff --git a/include/linux/fs.h b/include/linux/fs.h index bbd478e..c3b3462 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -391,6 +391,7 @@ struct inodes_stat_t { #include #include #include +#include =20 #include #include @@ -1440,8 +1441,14 @@ struct super_block { * Saved pool identifier for cleancache (-1 means none) */ int cleancache_poolid; + + struct shrinker s_shrink; /* per-sb shrinker handle */ }; =20 +/* superblock cache pruning functions */ +extern void prune_icache_sb(struct super_block *sb, int nr_to_scan); +extern void prune_dcache_sb(struct super_block *sb, int nr_to_scan); + extern struct timespec current_fs_time(struct super_block *sb); =20 /* --=20 1.7.5.1 -- To unsubscribe, send a message with 'unsubscribe linux-mm' in the body to majordomo@kvack.org. For more info on Linux MM, see: http://www.linux-mm.org/ . Fight unfair telecom internet charges in Canada: sign http://stopthemeter= .ca/ Don't email: email@kvack.org From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail137.messagelabs.com (mail137.messagelabs.com [216.82.249.19]) by kanga.kvack.org (Postfix) with SMTP id 199DF90010B for ; Thu, 2 Jun 2011 03:01:49 -0400 (EDT) From: Dave Chinner Subject: =?UTF-8?q?=5BPATCH=2008/12=5D=20superblock=3A=20introduce=20per-sb=20cache=20shrinker=20infrastructure?= Date: Thu, 2 Jun 2011 17:01:03 +1000 Message-Id: <1306998067-27659-9-git-send-email-david@fromorbit.com> In-Reply-To: <1306998067-27659-1-git-send-email-david@fromorbit.com> References: <1306998067-27659-1-git-send-email-david@fromorbit.com> MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sender: owner-linux-mm@kvack.org List-ID: To: linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, xfs@oss.sgi.com From: Dave Chinner With context based shrinkers, we can implement a per-superblock shrinker that shrinks the caches attached to the superblock. We currently have global shrinkers for the inode and dentry caches that split up into per-superblock operations via a coarse proportioning method that does not batch very well. The global shrinkers also have a dependency - dentries pin inodes - so we have to be very careful about how we register the global shrinkers so that the implicit call order is always correct. With a per-sb shrinker callout, we can encode this dependency directly into the per-sb shrinker, hence avoiding the need for strictly ordering shrinker registrations. We also have no need for any proportioning code for the shrinker subsystem already provides this functionality across all shrinkers. Allowing the shrinker to operate on a single superblock at a time means that we do less superblock list traversals and locking and reclaim should batch more effectively. This should result in less CPU overhead for reclaim and potentially faster reclaim of items from each filesystem. Signed-off-by: Dave Chinner --- fs/dcache.c | 121 +++++---------------------------------------------- fs/inode.c | 117 ++++---------------------------------------------- fs/super.c | 55 +++++++++++++++++++++++- include/linux/fs.h | 7 +++ 4 files changed, 82 insertions(+), 218 deletions(-) diff --git a/fs/dcache.c b/fs/dcache.c index 37f72ee..f73ef23 100644 --- a/fs/dcache.c +++ b/fs/dcache.c @@ -720,13 +720,11 @@ static void shrink_dentry_list(struct list_head *list) * * If flags contains DCACHE_REFERENCED reference dentries will not be pruned. */ -static void __shrink_dcache_sb(struct super_block *sb, int *count, int flags) +static void __shrink_dcache_sb(struct super_block *sb, int count, int flags) { - /* called from prune_dcache() and shrink_dcache_parent() */ struct dentry *dentry; LIST_HEAD(referenced); LIST_HEAD(tmp); - int cnt = *count; relock: spin_lock(&dcache_lru_lock); @@ -754,7 +752,7 @@ relock: } else { list_move_tail(&dentry->d_lru, &tmp); spin_unlock(&dentry->d_lock); - if (!--cnt) + if (!--count) break; } cond_resched_lock(&dcache_lru_lock); @@ -764,83 +762,22 @@ relock: spin_unlock(&dcache_lru_lock); shrink_dentry_list(&tmp); - - *count = cnt; } /** - * prune_dcache - shrink the dcache - * @count: number of entries to try to free + * prune_dcache_sb - shrink the dcache + * @nr_to_scan: number of entries to try to free * - * Shrink the dcache. This is done when we need more memory, or simply when we - * need to unmount something (at which point we need to unuse all dentries). + * Attempt to shrink the superblock dcache LRU by @nr_to_scan entries. This is + * done when we need more memory an called from the superblock shrinker + * function. * - * This function may fail to free any resources if all the dentries are in use. + * This function may fail to free any resources if all the dentries are in + * use. */ -static void prune_dcache(int count) +void prune_dcache_sb(struct super_block *sb, int nr_to_scan) { - struct super_block *sb, *p = NULL; - int w_count; - int unused = dentry_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused == 0 || count == 0) - return; - if (count >= unused) - prune_ratio = 1; - else - prune_ratio = unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_dentry_unused == 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb = - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio != 1) - w_count = (sb->s_nr_dentry_unused / prune_ratio) + 1; - else - w_count = sb->s_nr_dentry_unused; - pruned = w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root != NULL) && - (!list_empty(&sb->s_dentry_lru))) { - __shrink_dcache_sb(sb, &w_count, - DCACHE_REFERENCED); - pruned -= w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -= pruned; - p = sb; - /* more work left to do? */ - if (count <= 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); + __shrink_dcache_sb(sb, nr_to_scan, DCACHE_REFERENCED); } /** @@ -1215,42 +1152,10 @@ void shrink_dcache_parent(struct dentry * parent) int found; while ((found = select_parent(parent)) != 0) - __shrink_dcache_sb(sb, &found, 0); + __shrink_dcache_sb(sb, found, 0); } EXPORT_SYMBOL(shrink_dcache_parent); -/* - * Scan `sc->nr_slab_to_reclaim' dentries and return the number which remain. - * - * We need to avoid reentering the filesystem if the caller is performing a - * GFP_NOFS allocation attempt. One example deadlock is: - * - * ext2_new_block->getblk->GFP->shrink_dcache_memory->prune_dcache-> - * prune_one_dentry->dput->dentry_iput->iput->inode->i_sb->s_op->put_inode-> - * ext2_discard_prealloc->ext2_free_blocks->lock_super->DEADLOCK. - * - * In this case we return -1 to tell the caller that we baled. - */ -static int shrink_dcache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr = sc->nr_to_scan; - gfp_t gfp_mask = sc->gfp_mask; - - if (nr) { - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_dcache(nr); - } - - return (dentry_stat.nr_unused / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker dcache_shrinker = { - .shrink = shrink_dcache_memory, - .seeks = DEFAULT_SEEKS, -}; - /** * d_alloc - allocate a dcache entry * @parent: parent of entry to allocate @@ -3030,8 +2935,6 @@ static void __init dcache_init(void) */ dentry_cache = KMEM_CACHE(dentry, SLAB_RECLAIM_ACCOUNT|SLAB_PANIC|SLAB_MEM_SPREAD); - - register_shrinker(&dcache_shrinker); /* Hash may have been set up in dcache_init_early */ if (!hashdist) diff --git a/fs/inode.c b/fs/inode.c index 667a29c..890d95e 100644 --- a/fs/inode.c +++ b/fs/inode.c @@ -73,7 +73,7 @@ __cacheline_aligned_in_smp DEFINE_SPINLOCK(inode_wb_list_lock); * * We don't actually need it to protect anything in the umount path, * but only need to cycle through it to make sure any inode that - * prune_icache took off the LRU list has been fully torn down by the + * prune_icache_sb took off the LRU list has been fully torn down by the * time we are past evict_inodes. */ static DECLARE_RWSEM(iprune_sem); @@ -537,7 +537,7 @@ void evict_inodes(struct super_block *sb) dispose_list(&dispose); /* - * Cycle through iprune_sem to make sure any inode that prune_icache + * Cycle through iprune_sem to make sure any inode that prune_icache_sb * moved off the list before we took the lock has been fully torn * down. */ @@ -605,9 +605,10 @@ static int can_unuse(struct inode *inode) } /* - * Scan `goal' inodes on the unused list for freeable ones. They are moved to a - * temporary list and then are freed outside sb->s_inode_lru_lock by - * dispose_list(). + * Walk the superblock inode LRU for freeable inodes and attempt to free them. + * This is called from the superblock shrinker function with a number of inodes + * to trim from the LRU. Inodes to be freed are moved to a temporary list and + * then are freed outside inode_lock by dispose_list(). * * Any inodes which are pinned purely because of attached pagecache have their * pagecache removed. If the inode has metadata buffers attached to @@ -621,14 +622,15 @@ static int can_unuse(struct inode *inode) * LRU does not have strict ordering. Hence we don't want to reclaim inodes * with this flag set because they are the inodes that are out of order. */ -static void shrink_icache_sb(struct super_block *sb, int *nr_to_scan) +void prune_icache_sb(struct super_block *sb, int nr_to_scan) { LIST_HEAD(freeable); int nr_scanned; unsigned long reap = 0; + down_read(&iprune_sem); spin_lock(&sb->s_inode_lru_lock); - for (nr_scanned = *nr_to_scan; nr_scanned >= 0; nr_scanned--) { + for (nr_scanned = nr_to_scan; nr_scanned >= 0; nr_scanned--) { struct inode *inode; if (list_empty(&sb->s_inode_lru)) @@ -700,111 +702,11 @@ static void shrink_icache_sb(struct super_block *sb, int *nr_to_scan) else __count_vm_events(PGINODESTEAL, reap); spin_unlock(&sb->s_inode_lru_lock); - *nr_to_scan = nr_scanned; dispose_list(&freeable); -} - -static void prune_icache(int count) -{ - struct super_block *sb, *p = NULL; - int w_count; - int unused = inodes_stat.nr_unused; - int prune_ratio; - int pruned; - - if (unused == 0 || count == 0) - return; - down_read(&iprune_sem); - if (count >= unused) - prune_ratio = 1; - else - prune_ratio = unused / count; - spin_lock(&sb_lock); - list_for_each_entry(sb, &super_blocks, s_list) { - if (list_empty(&sb->s_instances)) - continue; - if (sb->s_nr_inodes_unused == 0) - continue; - sb->s_count++; - /* Now, we reclaim unused dentrins with fairness. - * We reclaim them same percentage from each superblock. - * We calculate number of dentries to scan on this sb - * as follows, but the implementation is arranged to avoid - * overflows: - * number of dentries to scan on this sb = - * count * (number of dentries on this sb / - * number of dentries in the machine) - */ - spin_unlock(&sb_lock); - if (prune_ratio != 1) - w_count = (sb->s_nr_inodes_unused / prune_ratio) + 1; - else - w_count = sb->s_nr_inodes_unused; - pruned = w_count; - /* - * We need to be sure this filesystem isn't being unmounted, - * otherwise we could race with generic_shutdown_super(), and - * end up holding a reference to an inode while the filesystem - * is unmounted. So we try to get s_umount, and make sure - * s_root isn't NULL. - */ - if (down_read_trylock(&sb->s_umount)) { - if ((sb->s_root != NULL) && - (!list_empty(&sb->s_dentry_lru))) { - shrink_icache_sb(sb, &w_count); - pruned -= w_count; - } - up_read(&sb->s_umount); - } - spin_lock(&sb_lock); - if (p) - __put_super(p); - count -= pruned; - p = sb; - /* more work left to do? */ - if (count <= 0) - break; - } - if (p) - __put_super(p); - spin_unlock(&sb_lock); up_read(&iprune_sem); } -/* - * shrink_icache_memory() will attempt to reclaim some unused inodes. Here, - * "unused" means that no dentries are referring to the inodes: the files are - * not open and the dcache references to those inodes have already been - * reclaimed. - * - * This function is passed the number of inodes to scan, and it returns the - * total number of remaining possibly-reclaimable inodes. - */ -static int shrink_icache_memory(struct shrinker *shrink, - struct shrink_control *sc) -{ - int nr = sc->nr_to_scan; - gfp_t gfp_mask = sc->gfp_mask; - - if (nr) { - /* - * Nasty deadlock avoidance. We may hold various FS locks, - * and we don't want to recurse into the FS that called us - * in clear_inode() and friends.. - */ - if (!(gfp_mask & __GFP_FS)) - return -1; - prune_icache(nr); - } - return (get_nr_inodes_unused() / 100) * sysctl_vfs_cache_pressure; -} - -static struct shrinker icache_shrinker = { - .shrink = shrink_icache_memory, - .seeks = DEFAULT_SEEKS, -}; - static void __wait_on_freeing_inode(struct inode *inode); /* * Called with the inode lock held. @@ -1684,7 +1586,6 @@ void __init inode_init(void) (SLAB_RECLAIM_ACCOUNT|SLAB_PANIC| SLAB_MEM_SPREAD), init_once); - register_shrinker(&icache_shrinker); /* Hash may have been set up in inode_init_early */ if (!hashdist) diff --git a/fs/super.c b/fs/super.c index 9c3fa1f..f4630d9 100644 --- a/fs/super.c +++ b/fs/super.c @@ -38,6 +38,50 @@ LIST_HEAD(super_blocks); DEFINE_SPINLOCK(sb_lock); +static int prune_super(struct shrinker *shrink, struct shrink_control *sc) +{ + struct super_block *sb; + int count; + + sb = container_of(shrink, struct super_block, s_shrink); + + /* + * Deadlock avoidance. We may hold various FS locks, and we don't want + * to recurse into the FS that called us in clear_inode() and friends.. + */ + if (sc->nr_to_scan && !(sc->gfp_mask & __GFP_FS)) + return -1; + + /* + * if we can't get the umount lock, then there's no point having the + * shrinker try again because the sb is being torn down. + */ + if (!down_read_trylock(&sb->s_umount)) + return -1; + + if (!sb->s_root) { + up_read(&sb->s_umount); + return -1; + } + + if (sc->nr_to_scan) { + /* proportion the scan between the two cacheN? */ + int total; + + total = sb->s_nr_dentry_unused + sb->s_nr_inodes_unused + 1; + count = (sc->nr_to_scan * sb->s_nr_dentry_unused) / total; + + /* prune dcache first as icache is pinned by it */ + prune_dcache_sb(sb, count); + prune_icache_sb(sb, sc->nr_to_scan - count); + } + + count = ((sb->s_nr_dentry_unused + sb->s_nr_inodes_unused) / 100) + * sysctl_vfs_cache_pressure; + up_read(&sb->s_umount); + return count; +} + /** * alloc_super - create new superblock * @type: filesystem type superblock should belong to @@ -116,6 +160,9 @@ static struct super_block *alloc_super(struct file_system_type *type) s->s_op = &default_op; s->s_time_gran = 1000000000; s->cleancache_poolid = -1; + + s->s_shrink.seeks = DEFAULT_SEEKS; + s->s_shrink.shrink = prune_super; } out: return s; @@ -278,7 +325,12 @@ void generic_shutdown_super(struct super_block *sb) { const struct super_operations *sop = sb->s_op; - + /* + * shut down the shrinker first so we know that there are no possible + * races when shrinking the dcache or icache. Removes the need for + * external locking to prevent such races. + */ + unregister_shrinker(&sb->s_shrink); if (sb->s_root) { shrink_dcache_for_umount(sb); sync_filesystem(sb); @@ -366,6 +418,7 @@ retry: list_add(&s->s_instances, &type->fs_supers); spin_unlock(&sb_lock); get_filesystem(type); + register_shrinker(&s->s_shrink); return s; } diff --git a/include/linux/fs.h b/include/linux/fs.h index bbd478e..c3b3462 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -391,6 +391,7 @@ struct inodes_stat_t { #include #include #include +#include #include #include @@ -1440,8 +1441,14 @@ struct super_block { * Saved pool identifier for cleancache (-1 means none) */ int cleancache_poolid; + + struct shrinker s_shrink; /* per-sb shrinker handle */ }; +/* superblock cache pruning functions */ +extern void prune_icache_sb(struct super_block *sb, int nr_to_scan); +extern void prune_dcache_sb(struct super_block *sb, int nr_to_scan); + extern struct timespec current_fs_time(struct super_block *sb); /* -- 1.7.5.1 -- To unsubscribe, send a message with 'unsubscribe linux-mm' in the body to majordomo@kvack.org. For more info on Linux MM, see: http://www.linux-mm.org/ . Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/ Don't email: email@kvack.org