* cleanerd
@ 2008-11-05 21:45 John Huttley
[not found] ` <491213E2.4010905-jE24nFfhqzU3hwNNidygWXTaI6DYlTYJ@public.gmane.org>
0 siblings, 1 reply; 9+ messages in thread
From: John Huttley @ 2008-11-05 21:45 UTC (permalink / raw)
To: NILFS Users mailing list
Hi,
my cleanerd stopped again, though I don't see anything in the logs
Is there a way of getting better logging for debugging purposes?
--John
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <491213E2.4010905-jE24nFfhqzU3hwNNidygWXTaI6DYlTYJ@public.gmane.org>
@ 2008-11-06 2:50 ` Ryusuke Konishi
0 siblings, 0 replies; 9+ messages in thread
From: Ryusuke Konishi @ 2008-11-06 2:50 UTC (permalink / raw)
To: users-JrjvKiOkagjYtjvyW6yDsg,
John-jE24nFfhqzU3hwNNidygWXTaI6DYlTYJ
Hi,
On Thu, 06 Nov 2008 10:45:06 +1300, John Huttley wrote:
> Hi,
> my cleanerd stopped again, though I don't see anything in the logs
>
> Is there a way of getting better logging for debugging purposes?
>
> --John
Sorry for inconvenience.
To get better cleanerd logs, change log_priority level specified in
/etc/nilfs_cleanerd.conf as follows:
log_priority info
|
v
log_priority debug
This will make the cleanerd write debug information in syslog.
Note that a HUP signal must be sent to the cleanerd or doing
umount/mount the nilfs partition is required to reflect the change.
Then, if nilfs2 seems to hang during GC, please get a stack dump by
sending the following system request from an active terminal or
console.
# echo t > /proc/sysrq-trigger
This will dump stack information of every kernel thread. The stack
information that contains `nilfs' keywords would give the hint about
which function causes the hang problem.
With regards,
Ryusuke Konishi
^ permalink raw reply [flat|nested] 9+ messages in thread
* cleanerd
@ 2009-10-10 22:18 Jan de Kruyf
[not found] ` <ee5afd760910101518u4a85fd0fn6e5539a327b2a876-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
0 siblings, 1 reply; 9+ messages in thread
From: Jan de Kruyf @ 2009-10-10 22:18 UTC (permalink / raw)
To: users-JrjvKiOkagjYtjvyW6yDsg
[-- Attachment #1.1: Type: text/plain, Size: 2196 bytes --]
Hallo,
Here is an interesting disaster with nilfs2. Lightening went down the
chimney on the other side of the wall. Computer was off, but plugged in.
Motherboard died.
So I now try to fix partitions. at the moment it is /var. /home will come
later.
Here is the problem:
/var data is readable but the cleanerd has become confused after the
disaster and has filled up the partition
partition size 5.8 Gig, data size 2.56 Gig, but usage is now 100%
Normal mounting:
cleanerd cannot be stopped anymore, neither with TERM or KILL signals.
lscp shows only cleaner checkpoints now.
When mounting without the cleanerd running there are no dmesg's of interest
on mounting or on execution of any command
To rescue the partition do I copy the data to another disk and reformat? Or
is there a simpler solution?
Best Regard,
Jan de Kruyf.
------------------
ps.
running debian lenny, kernel 2.6.26-1-686
branch 'master' of http://git.nilfs.org/nilfs2-utils dd. 11 july 2009
branch 'master' of http://git.nilfs.org/nilfs2-module dd. 11 july 2009
tag 'v2.0.15'
and here is the superblock of the partition:
00000000: 0200 0000 0000 3434 0001 0000 8422 95d1 ......44....."..
00000010: 80b7 90e7 0200 0000 ca02 0000 0000 0000 ................
00000020: 00b4 6665 0100 0000 0100 0000 0000 0000 ..fe............
00000030: 0008 0000 0500 0000 e1e6 1400 0000 0000 ................
00000040: d69f 0d00 0000 0000 c181 0900 0000 0000 ................
00000050: 0000 0000 0000 0000 4815 5a4a 0000 0000 ........H.ZJ....
00000060: 3713 d04a 0000 0000 3713 d04a 0000 0000 7..J....7..J....
00000070: 4000 3200 0000 0100 4815 5a4a 0000 0000 @.2.....H.ZJ....
00000080: 004e ed00 0000 0000 0000 0000 0b00 0000 .N..............
00000090: 8000 2000 c000 1000 470b 1351 09a6 4016 .. .....G..Q..@.
000000a0: 9397 70f2 82f5 d61b 7661 7200 0000 0000 ..p.....var.....
000000b0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
000000c0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
000000d0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
000000e0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
000000f0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
[-- Attachment #1.2: Type: text/html, Size: 2534 bytes --]
[-- Attachment #2: Type: text/plain, Size: 158 bytes --]
_______________________________________________
users mailing list
users-JrjvKiOkagjYtjvyW6yDsg@public.gmane.org
https://www.nilfs.org/mailman/listinfo/users
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <ee5afd760910101518u4a85fd0fn6e5539a327b2a876-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2009-10-11 3:29 ` Ryusuke Konishi
[not found] ` <20091011.122911.131278678.ryusuke-sG5X7nlA6pw@public.gmane.org>
0 siblings, 1 reply; 9+ messages in thread
From: Ryusuke Konishi @ 2009-10-11 3:29 UTC (permalink / raw)
To: users-JrjvKiOkagjYtjvyW6yDsg, jan.de.kruyf-Re5JQEeQqe8AvxtiuMwx3w
Hi,
On Sun, 11 Oct 2009 00:18:10 +0200, Jan de Kruyf wrote:
> Hallo,
> Here is an interesting disaster with nilfs2. Lightening went down the
> chimney on the other side of the wall. Computer was off, but plugged in.
> Motherboard died.
>
> So I now try to fix partitions. at the moment it is /var. /home will come
> later.
>
> Here is the problem:
> /var data is readable but the cleanerd has become confused after the
> disaster and has filled up the partition
>
> partition size 5.8 Gig, data size 2.56 Gig, but usage is now 100%
> Normal mounting:
> cleanerd cannot be stopped anymore, neither with TERM or KILL signals.
Seems like it went into an infinite loop.
Can you change ``log priority'' written in /etc/nilfs_cleanerd.conf ?
Cleanerd may log additional messages if you change it ``debug'' level:
log_priority debug
The syslog is written out to the /var directory, so you should
mount the /var data on /mnt or other.
> lscp shows only cleaner checkpoints now.
>
> When mounting without the cleanerd running there are no dmesg's of interest
> on mounting or on execution of any command
>
> To rescue the partition do I copy the data to another disk and reformat? Or
> is there a simpler solution?
Well, I recommend you to upgrade the module to the latest version
'v2.0.17' because former versions may cause file system corruption or
kernel oopses if things turn out bad. Two maintenance releases were
made after you pulled nilfs2-module.git.
nilfs-utils also has new version (v2.0.14).
If you will see the same problem for these new versions, maybe you
should make a backup copy and reformat the partition.
But, I would appreciate it if you could help me to find the cause of
the busy loop in cleanerd before the reformat.
> Best Regard,
>
> Jan de Kruyf.
> ------------------
Thank you for reporting the issue.
With regards,
Ryusuke Konishi
> ps.
> running debian lenny, kernel 2.6.26-1-686
>
> branch 'master' of http://git.nilfs.org/nilfs2-utils dd. 11 july 2009
> branch 'master' of http://git.nilfs.org/nilfs2-module dd. 11 july 2009
> tag 'v2.0.15'
>
> and here is the superblock of the partition:
>
>
> 00000000: 0200 0000 0000 3434 0001 0000 8422 95d1 ......44....."..
> 00000010: 80b7 90e7 0200 0000 ca02 0000 0000 0000 ................
> 00000020: 00b4 6665 0100 0000 0100 0000 0000 0000 ..fe............
> 00000030: 0008 0000 0500 0000 e1e6 1400 0000 0000 ................
> 00000040: d69f 0d00 0000 0000 c181 0900 0000 0000 ................
> 00000050: 0000 0000 0000 0000 4815 5a4a 0000 0000 ........H.ZJ....
> 00000060: 3713 d04a 0000 0000 3713 d04a 0000 0000 7..J....7..J....
> 00000070: 4000 3200 0000 0100 4815 5a4a 0000 0000 @.2.....H.ZJ....
> 00000080: 004e ed00 0000 0000 0000 0000 0b00 0000 .N..............
> 00000090: 8000 2000 c000 1000 470b 1351 09a6 4016 .. .....G..Q..@.
> 000000a0: 9397 70f2 82f5 d61b 7661 7200 0000 0000 ..p.....var.....
> 000000b0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> 000000c0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> 000000d0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> 000000e0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> 000000f0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <20091011.122911.131278678.ryusuke-sG5X7nlA6pw@public.gmane.org>
@ 2009-10-11 5:32 ` Jan de Kruyf
[not found] ` <ee5afd760910102232o5f17c50cxb5024f6f76ecbcf5-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
0 siblings, 1 reply; 9+ messages in thread
From: Jan de Kruyf @ 2009-10-11 5:32 UTC (permalink / raw)
To: Ryusuke Konishi, users-JrjvKiOkagjYtjvyW6yDsg
[-- Attachment #1.1: Type: text/plain, Size: 5710 bytes --]
Hallo,
Sorry the detail was a little bit scant last night.
The nilfs versions running on the machine at the time of the disaster were
the latest versions.
This is the maintenance hard-drive running, I will update today.
The loop is (as far as I can see now from the logs)
-
-------kern.log-------------------------------------------
Oct 10 06:53:11 debianLenny kernel: [44514.982086] segctord starting.
Construction interval = 5 seconds, CP frequency < 30 seconds
Oct 10 06:53:11 debianLenny kernel: [44515.115227] NILFS warning: mounting
unchecked fs
Oct 10 06:53:11 debianLenny kernel: [44515.398152] NILFS: recovery complete.
Oct 10 06:53:28 debianLenny kernel: [44535.631729] NILFS warning (device
hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
Oct 10 06:53:33 debianLenny kernel: [44542.849960] NILFS warning (device
hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
Oct 10 06:53:38 debianLenny kernel: [44550.592403] NILFS warning (device
hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
---------------------------------------------------------------
this is from the maintenance versions on /dev/hda (still to be updated),
with /dev/hdb not functional.
It will fill up the log until the partition is full.
I will dd the partion into a loop-mountable file once I have updated, so
diagnostics
may continue.
I was aware of this problem before: when you overfill a partition cleanerd
goes into this loop.
The interesting part is: how did cleanerd get confused this time. since the
partition was only half full
and cleanerd was running more or less regularly. I am only aware that I
stopped cleanerd a few times that day with TERM
since it was in the way of other work. It made the /home partition so slow
that I could not write a dvd anymore.
But this is a separate issue.
So I will do a low-level check on the sick disc to check for media format
failures
and I will try to do some log reading over the next few days to see if I can
find
the log of the first mount after the accident.
Regards
Jan de Kruyf.
"Let us sing, the Lord is on his Throne and the earth is full of his Glory."
enjoy the rest of your day.
On Sun, Oct 11, 2009 at 5:29 AM, Ryusuke Konishi <ryusuke-sG5X7nlA6pw@public.gmane.org> wrote:
> Hi,
> On Sun, 11 Oct 2009 00:18:10 +0200, Jan de Kruyf wrote:
> > Hallo,
> > Here is an interesting disaster with nilfs2. Lightening went down the
> > chimney on the other side of the wall. Computer was off, but plugged in.
> > Motherboard died.
> >
> > So I now try to fix partitions. at the moment it is /var. /home will come
> > later.
> >
> > Here is the problem:
> > /var data is readable but the cleanerd has become confused after the
> > disaster and has filled up the partition
> >
> > partition size 5.8 Gig, data size 2.56 Gig, but usage is now 100%
> > Normal mounting:
> > cleanerd cannot be stopped anymore, neither with TERM or KILL signals.
>
> Seems like it went into an infinite loop.
>
> Can you change ``log priority'' written in /etc/nilfs_cleanerd.conf ?
>
> Cleanerd may log additional messages if you change it ``debug'' level:
>
> log_priority debug
>
> The syslog is written out to the /var directory, so you should
> mount the /var data on /mnt or other.
>
> > lscp shows only cleaner checkpoints now.
> >
> > When mounting without the cleanerd running there are no dmesg's of
> interest
> > on mounting or on execution of any command
> >
> > To rescue the partition do I copy the data to another disk and reformat?
> Or
> > is there a simpler solution?
>
> Well, I recommend you to upgrade the module to the latest version
> 'v2.0.17' because former versions may cause file system corruption or
> kernel oopses if things turn out bad. Two maintenance releases were
> made after you pulled nilfs2-module.git.
>
> nilfs-utils also has new version (v2.0.14).
>
> If you will see the same problem for these new versions, maybe you
> should make a backup copy and reformat the partition.
>
> But, I would appreciate it if you could help me to find the cause of
> the busy loop in cleanerd before the reformat.
>
> > Best Regard,
> >
> > Jan de Kruyf.
> > ------------------
>
> Thank you for reporting the issue.
>
> With regards,
> Ryusuke Konishi
>
> > ps.
> > running debian lenny, kernel 2.6.26-1-686
> >
> > branch 'master' of http://git.nilfs.org/nilfs2-utils dd. 11 july 2009
> > branch 'master' of http://git.nilfs.org/nilfs2-module dd. 11 july 2009
> > tag 'v2.0.15'
> >
> > and here is the superblock of the partition:
> >
> >
> > 00000000: 0200 0000 0000 3434 0001 0000 8422 95d1 ......44....."..
> > 00000010: 80b7 90e7 0200 0000 ca02 0000 0000 0000 ................
> > 00000020: 00b4 6665 0100 0000 0100 0000 0000 0000 ..fe............
> > 00000030: 0008 0000 0500 0000 e1e6 1400 0000 0000 ................
> > 00000040: d69f 0d00 0000 0000 c181 0900 0000 0000 ................
> > 00000050: 0000 0000 0000 0000 4815 5a4a 0000 0000 ........H.ZJ....
> > 00000060: 3713 d04a 0000 0000 3713 d04a 0000 0000 7..J....7..J....
> > 00000070: 4000 3200 0000 0100 4815 5a4a 0000 0000 @.2.....H.ZJ....
> > 00000080: 004e ed00 0000 0000 0000 0000 0b00 0000 .N..............
> > 00000090: 8000 2000 c000 1000 470b 1351 09a6 4016 .. .....G..Q..@.
> > 000000a0: 9397 70f2 82f5 d61b 7661 7200 0000 0000 ..p.....var.....
> > 000000b0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> > 000000c0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> > 000000d0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> > 000000e0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
> > 000000f0: 0000 0000 0000 0000 0000 0000 0000 0000 ................
>
[-- Attachment #1.2: Type: text/html, Size: 6890 bytes --]
[-- Attachment #2: Type: text/plain, Size: 158 bytes --]
_______________________________________________
users mailing list
users-JrjvKiOkagjYtjvyW6yDsg@public.gmane.org
https://www.nilfs.org/mailman/listinfo/users
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <ee5afd760910102232o5f17c50cxb5024f6f76ecbcf5-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2009-10-11 6:49 ` Ryusuke Konishi
[not found] ` <20091011.154916.07997858.ryusuke-sG5X7nlA6pw@public.gmane.org>
0 siblings, 1 reply; 9+ messages in thread
From: Ryusuke Konishi @ 2009-10-11 6:49 UTC (permalink / raw)
To: jan.de.kruyf-Re5JQEeQqe8AvxtiuMwx3w; +Cc: users-JrjvKiOkagjYtjvyW6yDsg
Hi,
On Sun, 11 Oct 2009 07:32:50 +0200, Jan de Kruyf wrote:
> Hallo,
> Sorry the detail was a little bit scant last night.
> The nilfs versions running on the machine at the time of the disaster were
> the latest versions.
> This is the maintenance hard-drive running, I will update today.
>
> The loop is (as far as I can see now from the logs)
> -
> -------kern.log-------------------------------------------
> Oct 10 06:53:11 debianLenny kernel: [44514.982086] segctord starting.
> Construction interval = 5 seconds, CP frequency < 30 seconds
> Oct 10 06:53:11 debianLenny kernel: [44515.115227] NILFS warning: mounting
> unchecked fs
> Oct 10 06:53:11 debianLenny kernel: [44515.398152] NILFS: recovery complete.
> Oct 10 06:53:28 debianLenny kernel: [44535.631729] NILFS warning (device
> hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> Oct 10 06:53:33 debianLenny kernel: [44542.849960] NILFS warning (device
> hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> Oct 10 06:53:38 debianLenny kernel: [44550.592403] NILFS warning (device
> hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> ---------------------------------------------------------------
>
> this is from the maintenance versions on /dev/hda (still to be updated),
> with /dev/hdb not functional.
> It will fill up the log until the partition is full.
According to the log, the error was repeatedly detected in a retry
loop in the nilfs_clean_segments kernel function which cleanerd calls
via ioctl. The err=-28 means ENOSPC (no space left on the device).
Yeah, if cleanerd falls into this state, it cannot handle any signals.
And, it doesn't return to userspace until the error is removed.
As you are pointing out, I wonder why this error is generated on the
device having enough free space.
I'll attach a patch to identify which function returns ENOSPC. Could
you try the patch ?
Thanks,
Ryusuke Konishi
> I will dd the partion into a loop-mountable file once I have updated, so
> diagnostics
> may continue.
>
> I was aware of this problem before: when you overfill a partition cleanerd
> goes into this loop.
>
> The interesting part is: how did cleanerd get confused this time. since the
> partition was only half full
> and cleanerd was running more or less regularly. I am only aware that I
> stopped cleanerd a few times that day with TERM
> since it was in the way of other work. It made the /home partition so slow
> that I could not write a dvd anymore.
> But this is a separate issue.
>
> So I will do a low-level check on the sick disc to check for media format
> failures
> and I will try to do some log reading over the next few days to see if I can
> find
> the log of the first mount after the accident.
>
> Regards
> Jan de Kruyf.
>
> "Let us sing, the Lord is on his Throne and the earth is full of his Glory."
> enjoy the rest of your day.
diff --git a/fs/alloc.c b/fs/alloc.c
index 1c76c38..fdec249 100644
--- a/fs/alloc.c
+++ b/fs/alloc.c
@@ -235,6 +235,8 @@ static int nilfs_palloc_find_available_slot(struct inode *inode,
return pos;
}
}
+ printk(KERN_ERR "%s: disk full\n", __func__);
+ dump_stack();
return -ENOSPC;
}
@@ -320,6 +322,8 @@ int nilfs_palloc_prepare_alloc_entry(struct inode *inode,
}
/* no entries left */
+ printk(KERN_ERR "%s: disk full\n", __func__);
+ dump_stack();
return -ENOSPC;
out_desc:
diff --git a/fs/sufile.c b/fs/sufile.c
index 47ad9a4..e109a7e 100644
--- a/fs/sufile.c
+++ b/fs/sufile.c
@@ -322,6 +322,8 @@ int nilfs_sufile_alloc(struct inode *sufile, __u64 *segnump)
}
/* no segments left */
+ printk(KERN_ERR "%s: disk full\n", __func__);
+ dump_stack();
ret = -ENOSPC;
out_header:
^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <20091011.154916.07997858.ryusuke-sG5X7nlA6pw@public.gmane.org>
@ 2009-10-13 19:57 ` Jan de Kruyf
2009-10-17 20:47 ` cleanerd Jan de Kruyf
1 sibling, 0 replies; 9+ messages in thread
From: Jan de Kruyf @ 2009-10-13 19:57 UTC (permalink / raw)
To: Ryusuke Konishi, users-JrjvKiOkagjYtjvyW6yDsg
[-- Attachment #1.1: Type: text/plain, Size: 4922 bytes --]
Hallo,
Have not done the patch yet. I was lost in EXT3, some data loss problem.
The HD media test ok with seatools.
I think I found the checkpoints and the segments written just before and one
after the disaster when I tried
to run the hd on another computer. At that time already the /var reported
full.
The system then made an emergency /var somewhere I do not know
Presumably in RAM.
But I cannot mount a snapshot on a loop mounted image and I cannot change a
cp to ss on a full drive.
Is this correct or am I confused?
From the log data on the broken partition I seem to think that nilfs does
not mount the latest checkpoint but I might be mistaken
that is why I wanted to mount the latest in the lscp list. and see if there
are differences.
Regards
Jan de Kruyf.
On Sun, Oct 11, 2009 at 8:49 AM, Ryusuke Konishi <ryusuke-sG5X7nlA6pw@public.gmane.org> wrote:
> Hi,
> On Sun, 11 Oct 2009 07:32:50 +0200, Jan de Kruyf wrote:
> > Hallo,
> > Sorry the detail was a little bit scant last night.
> > The nilfs versions running on the machine at the time of the disaster
> were
> > the latest versions.
> > This is the maintenance hard-drive running, I will update today.
> >
> > The loop is (as far as I can see now from the logs)
> > -
> > -------kern.log-------------------------------------------
> > Oct 10 06:53:11 debianLenny kernel: [44514.982086] segctord starting.
> > Construction interval = 5 seconds, CP frequency < 30 seconds
> > Oct 10 06:53:11 debianLenny kernel: [44515.115227] NILFS warning:
> mounting
> > unchecked fs
> > Oct 10 06:53:11 debianLenny kernel: [44515.398152] NILFS: recovery
> complete.
> > Oct 10 06:53:28 debianLenny kernel: [44535.631729] NILFS warning (device
> > hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> > Oct 10 06:53:33 debianLenny kernel: [44542.849960] NILFS warning (device
> > hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> > Oct 10 06:53:38 debianLenny kernel: [44550.592403] NILFS warning (device
> > hdb9): nilfs_clean_segments: segment construction failed. (err=-28)
> > ---------------------------------------------------------------
> >
> > this is from the maintenance versions on /dev/hda (still to be updated),
> > with /dev/hdb not functional.
> > It will fill up the log until the partition is full.
>
> According to the log, the error was repeatedly detected in a retry
> loop in the nilfs_clean_segments kernel function which cleanerd calls
> via ioctl. The err=-28 means ENOSPC (no space left on the device).
>
> Yeah, if cleanerd falls into this state, it cannot handle any signals.
> And, it doesn't return to userspace until the error is removed.
>
> As you are pointing out, I wonder why this error is generated on the
> device having enough free space.
>
> I'll attach a patch to identify which function returns ENOSPC. Could
> you try the patch ?
>
> Thanks,
> Ryusuke Konishi
>
> > I will dd the partion into a loop-mountable file once I have updated, so
> > diagnostics
> > may continue.
> >
> > I was aware of this problem before: when you overfill a partition
> cleanerd
> > goes into this loop.
> >
> > The interesting part is: how did cleanerd get confused this time. since
> the
> > partition was only half full
> > and cleanerd was running more or less regularly. I am only aware that I
> > stopped cleanerd a few times that day with TERM
> > since it was in the way of other work. It made the /home partition so
> slow
> > that I could not write a dvd anymore.
> > But this is a separate issue.
> >
> > So I will do a low-level check on the sick disc to check for media format
> > failures
> > and I will try to do some log reading over the next few days to see if I
> can
> > find
> > the log of the first mount after the accident.
> >
> > Regards
> > Jan de Kruyf.
> >
> > "Let us sing, the Lord is on his Throne and the earth is full of his
> Glory."
> > enjoy the rest of your day.
>
>
> diff --git a/fs/alloc.c b/fs/alloc.c
> index 1c76c38..fdec249 100644
> --- a/fs/alloc.c
> +++ b/fs/alloc.c
> @@ -235,6 +235,8 @@ static int nilfs_palloc_find_available_slot(struct
> inode *inode,
> return pos;
> }
> }
> + printk(KERN_ERR "%s: disk full\n", __func__);
> + dump_stack();
> return -ENOSPC;
> }
>
> @@ -320,6 +322,8 @@ int nilfs_palloc_prepare_alloc_entry(struct inode
> *inode,
> }
>
> /* no entries left */
> + printk(KERN_ERR "%s: disk full\n", __func__);
> + dump_stack();
> return -ENOSPC;
>
> out_desc:
> diff --git a/fs/sufile.c b/fs/sufile.c
> index 47ad9a4..e109a7e 100644
> --- a/fs/sufile.c
> +++ b/fs/sufile.c
> @@ -322,6 +322,8 @@ int nilfs_sufile_alloc(struct inode *sufile, __u64
> *segnump)
> }
>
> /* no segments left */
> + printk(KERN_ERR "%s: disk full\n", __func__);
> + dump_stack();
> ret = -ENOSPC;
>
> out_header:
>
[-- Attachment #1.2: Type: text/html, Size: 5860 bytes --]
[-- Attachment #2: Type: text/plain, Size: 158 bytes --]
_______________________________________________
users mailing list
users-JrjvKiOkagjYtjvyW6yDsg@public.gmane.org
https://www.nilfs.org/mailman/listinfo/users
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <20091011.154916.07997858.ryusuke-sG5X7nlA6pw@public.gmane.org>
2009-10-13 19:57 ` cleanerd Jan de Kruyf
@ 2009-10-17 20:47 ` Jan de Kruyf
[not found] ` <ee5afd760910171347x4ad27199ka59d0e76f3271050-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
1 sibling, 1 reply; 9+ messages in thread
From: Jan de Kruyf @ 2009-10-17 20:47 UTC (permalink / raw)
To: Ryusuke Konishi; +Cc: users-JrjvKiOkagjYtjvyW6yDsg
[-- Attachment #1.1: Type: text/plain, Size: 947 bytes --]
On Sun, Oct 11, 2009 at 8:49 AM, Ryusuke Konishi <ryusuke-sG5X7nlA6pw@public.gmane.org> wrote:
>
> According to the log, the error was repeatedly detected in a retry
> loop in the nilfs_clean_segments kernel function which cleanerd calls
> via ioctl. The err=-28 means ENOSPC (no space left on the device).
>
> Yeah, if cleanerd falls into this state, it cannot handle any signals.
> And, it doesn't return to userspace until the error is removed.
>
> As you are pointing out, I wonder why this error is generated on the
> device having enough free space.
>
> I'll attach a patch to identify which function returns ENOSPC. Could
> you try the patch ?
>
Hallo,
I have done the patch. The output, with debug switched on, is in the
attached file.
I have done some more analizing. Everything that looks helpful is included
in the file as wel.
I hope it is useful
Regards
Jan de Kruyf.
This day the Lord creates,
so let us celebrate with joy!
[-- Attachment #1.2: Type: text/html, Size: 1336 bytes --]
[-- Attachment #2: nilfsdebugVar.txt.gz --]
[-- Type: application/x-gzip, Size: 44410 bytes --]
[-- Attachment #3: Type: text/plain, Size: 158 bytes --]
_______________________________________________
users mailing list
users-JrjvKiOkagjYtjvyW6yDsg@public.gmane.org
https://www.nilfs.org/mailman/listinfo/users
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: cleanerd
[not found] ` <ee5afd760910171347x4ad27199ka59d0e76f3271050-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2009-10-19 2:07 ` Ryusuke Konishi
0 siblings, 0 replies; 9+ messages in thread
From: Ryusuke Konishi @ 2009-10-19 2:07 UTC (permalink / raw)
To: NILFS Users mailing list, Jan de Kruyf
Hi,
2009/10/18 Jan de Kruyf <jan.de.kruyf-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>:
> On Sun, Oct 11, 2009 at 8:49 AM, Ryusuke Konishi <ryusuke-sG5X7nlA6pw@public.gmane.org> wrote:
>>
>> According to the log, the error was repeatedly detected in a retry
>> loop in the nilfs_clean_segments kernel function which cleanerd calls
>> via ioctl. The err=-28 means ENOSPC (no space left on the device).
>>
>> Yeah, if cleanerd falls into this state, it cannot handle any signals.
>> And, it doesn't return to userspace until the error is removed.
>>
>> As you are pointing out, I wonder why this error is generated on the
>> device having enough free space.
>>
>> I'll attach a patch to identify which function returns ENOSPC. Could
>> you try the patch ?
>
> Hallo,
> I have done the patch. The output, with debug switched on, is in the
> attached file.
>
> I have done some more analizing. Everything that looks helpful is included
> in the file as wel.
>
> I hope it is useful
>
> Regards
>
> Jan de Kruyf.
>
> This day the Lord creates,
> so let us celebrate with joy!
Thank you for the detail report and your cooperation!
The log shows that GC got trapped due to real disk full on the /var directory.
Although the following lssu output shows there is one free segment, it's not
writable because the log to be written in the segment needs to know another
segment which comes next.
434 2009-09-26 12:31:31 -d- 2048
435 2009-09-28 20:42:54 ad- 2038
436 ---------- --:--:-- ad- 0
437 2009-09-25 16:25:30 -d- 2048
The problem is why this stuck situation happened.
If you have ever run an older version of nilfs (v2.0.13 or prior) for
the partition,
it may arise from erroneous metadata written by the old version.
Otherwise, the logic to predict disk full condition likely has some issue.
If you can reformat the partition, please try "-m" option of mkfs.nilfs2.
With the option, you can increase the number of segment reserved for
garbage collection.
By default, 5 percent of segments are reserved.
If the same problem happens for a higher ratio (e.g. 20% or so),
some kind of leakage seems to be in the disk full determination.
With regards,
Ryusuke Konishi
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2009-10-19 2:07 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2008-11-05 21:45 cleanerd John Huttley
[not found] ` <491213E2.4010905-jE24nFfhqzU3hwNNidygWXTaI6DYlTYJ@public.gmane.org>
2008-11-06 2:50 ` cleanerd Ryusuke Konishi
-- strict thread matches above, loose matches on Subject: below --
2009-10-10 22:18 cleanerd Jan de Kruyf
[not found] ` <ee5afd760910101518u4a85fd0fn6e5539a327b2a876-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2009-10-11 3:29 ` cleanerd Ryusuke Konishi
[not found] ` <20091011.122911.131278678.ryusuke-sG5X7nlA6pw@public.gmane.org>
2009-10-11 5:32 ` cleanerd Jan de Kruyf
[not found] ` <ee5afd760910102232o5f17c50cxb5024f6f76ecbcf5-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2009-10-11 6:49 ` cleanerd Ryusuke Konishi
[not found] ` <20091011.154916.07997858.ryusuke-sG5X7nlA6pw@public.gmane.org>
2009-10-13 19:57 ` cleanerd Jan de Kruyf
2009-10-17 20:47 ` cleanerd Jan de Kruyf
[not found] ` <ee5afd760910171347x4ad27199ka59d0e76f3271050-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2009-10-19 2:07 ` cleanerd Ryusuke Konishi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox