* HPMC bus timeout on C3600
@ 2008-12-20 12:39 Guy Martin
2008-12-20 22:16 ` Kyle McMartin
2008-12-21 8:10 ` Grant Grundler
0 siblings, 2 replies; 20+ messages in thread
From: Guy Martin @ 2008-12-20 12:39 UTC (permalink / raw)
To: linux-parisc
[-- Attachment #1: Type: text/plain, Size: 1189 bytes --]
Hi team,
I'm recently running into HPMC bus timeout problem. While this never
caused any problem before, I've noticed a few strange things. I've
attached the output I get from "pim hpmc" in the PDC.
This BUS timeout occurs after one or two days of idling when I plug a
specific card reader in the USB PCI card I plugged in. I think this
card reader causes problems because it's being polled every 2-3 seconds.
Nevertheless, looking at the pdc output, it seems that the OS HPMC
handler is not kicking off according to the chassis code CBF2 and CBFC.
CBF2 : bas OS HPMC len
CBFC : OS HPMC br err
I've check GR02 and it points to inw() which makes sens. Nothing
interesting there.
The failing device appears to be the built-in NIC according the path
provided in the output (10/0/12/0) and the C3600 service manual (figure
5-2).
Even if this look HW pb, I've had a few reports recently about HPMC PCI
timeout issues and I have a few doubts.
Now my questions are :
- does this look like an HW or SW problems ?
- why isn't the HPMC handler kicked off ?
- could the HPMC handler recover this ?
- what debug can I enable to get more info about this if applicable ?
Cheers,
Guy
[-- Attachment #2: c3600-hpmc --]
[-- Type: application/octet-stream, Size: 25894 bytes --]
Service Menu: Enter command > pim hpmc
PROCESSOR PIM INFORMATION
----------------- Processor 0 HPMC Information ------------------
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc
General Registers 0 - 31
00-03 0000000000000000 00000000104dc150 00000000101176f0 000000008f31e360
04-07 0000000000000000 000000000000000f 000000008f31e000 0000000000000000
08-11 0000000000000000 0000000010582950 00000000104dc150 000000001057ede8
12-15 00000000f0000000 0000000000000081 000000001040c0e4 00000000f0400004
16-19 00000000f00008c4 00000000f000017c 00000000f0000174 0000000000020000
20-23 00000000000000d4 00000000030a9fcb 000000001027a7b8 0000000000000000
24-27 0000000000000001 0000000000001040 000000008f804580 00000000104dc150
28-31 00000000104fb778 0000000002584fd6 000000008f8e0380 00000000101176f0
<Press any key to continue (q to quit)>
Control Registers 0 - 31
00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000
04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000
08-11 0000000000007ad2 0000000000000000 00000000000000c0 000000000000003f
12-15 0000000000000000 0000000000000000 000000000010f000 00000000ff800000
16-19 000055958963ba84 0000000000000000 000000001027a7c4 000000000e79009c
20-23 00000000a627fffb 0000000080421040 000000ff0006f90e 0000000080000000
24-27 000000000052a000 000000004250e000 0000000000044021 00000000f0412000
28-31 0000000055555555 0000000055555555 000000008f8e0000 0000000011111111
Space Registers 0 - 7
00-03 00000000 00000000 00000000 00003d69
04-07 00000000 00000000 00000000 00000000
IIA Space = 0x0000000000000000
IIA Offset = 0x000000001027a7c8
Check Type = 0x20000000
CPU State = 0x9e000004
Cache Check = 0x00000000
TLB Check = 0x00000000
Bus Check = 0x0030103b
Assists Check = 0x00000000
Assist State = 0x00000000
Path Info = 0x00000000
System Responder Address = 0x000000fffee01040
System Requestor Address = 0xfffffffffffa0000
Floating-Point Registers 0 - 31
00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000
04-07 3fd3e672c6670f8c 41d215ed21800000 0000000654000000 5ea496f000000000
08-11 00000003b26de560 3ff0000000000000 3ff0000000000000 0000000000000000
12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555
16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555
20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000
24-27 0000888630000000 0000ce9a00000000 000000006e2df49c 4024cb5ecf0a9480
28-31 3ff0000000000000 0000000000000000 4024cb5ecf0a9480 0000000000000000
'9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes:
Check Summary = 0xcb81041008000000
Available Memory = 0x0000000080000000
CPU Diagnose Register 2 = 0x0300000000000004
CPU Status Register 0 = 0x2420c20000000000
CPU Status Register 1 = 0x8002000000000000
SADD LOG = 0xd3081bfb486511b8
Read Short LOG = 0xc1af00fffee01040
ERROR_STATUS = 0x0000000000100010
MEM_ADDR = 0x000001ff3fffffff
MEM_SYND = 0x0000000000000000
MEM_ADDR_CORR = 0x000001ff3fffffff
MEM_SYND_CORR = 0x0000000000000000
RUN_DATA_HIGH = 0xc1bff0fffed08040
RUN_DATA_LOW = 0xc1bff0fffed08040
RUN_CTRL = 0x0000021c00001418
RUN_ADDR = 0xc1bff0fffed08040
System Responder Path = 0x00ffffff0a000c00
HPMC PIM Analysis Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes:
A Data I/O Fetch Timeout occurred while CPU 0 was
requesting information from a device at the path 10/0/12/0 (built-in PCI device).
Memory/IO Controller Error Analysis Information:
The Memory/IO Controller only observed the Broadcast Error. It did not log
any additional information about the HPMC.
Memory Error Log Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes:
No memory errors logged
I/O Module Error Log Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes:
Rope Word1 Word2 Word3
------ ------------ ------------
0 0x00000000 0x0e0cc2a9 0x00000000fed30048
1 0x00000000 0x1e0cc009 0x00000000fed32048
2 ---------- 0x2e0cc009 ------------------
3 ---------- 0x3e0cc009 ------------------
4 0x00000000 0x4e0cc009 0x00000000fed38048
5 ---------- 0x5e0cc009 ------------------
6 0x00000000 0x6e0cc009 0x00000000fed3c048
7 ---------- 0x7e0cc009 ------------------
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin
@ 2008-12-20 22:16 ` Kyle McMartin
2008-12-20 22:21 ` Kyle McMartin
2008-12-21 8:10 ` Grant Grundler
1 sibling, 1 reply; 20+ messages in thread
From: Kyle McMartin @ 2008-12-20 22:16 UTC (permalink / raw)
To: Guy Martin; +Cc: linux-parisc
On Sat, Dec 20, 2008 at 01:39:15PM +0100, Guy Martin wrote:
> I'm recently running into HPMC bus timeout problem. While this never
> caused any problem before, I've noticed a few strange things. I've
> attached the output I get from "pim hpmc" in the PDC.
>
> This BUS timeout occurs after one or two days of idling when I plug a
> specific card reader in the USB PCI card I plugged in. I think this
> card reader causes problems because it's being polled every 2-3 seconds.
>
> Nevertheless, looking at the pdc output, it seems that the OS HPMC
> handler is not kicking off according to the chassis code CBF2 and CBFC.
> CBF2 : bas OS HPMC len
> CBFC : OS HPMC br err
>
> I've check GR02 and it points to inw() which makes sens. Nothing
> interesting there.
>
> The failing device appears to be the built-in NIC according the path
> provided in the output (10/0/12/0) and the C3600 service manual (figure
> 5-2).
>
> Even if this look HW pb, I've had a few reports recently about HPMC PCI
> timeout issues and I have a few doubts.
>
> Now my questions are :
> - does this look like an HW or SW problems ?
> - why isn't the HPMC handler kicked off ?
> - could the HPMC handler recover this ?
> - what debug can I enable to get more info about this if applicable ?
>
"bus timeout" usually means we tried to read an address that doesn't
respond. that is, nothing on the bus accepted the transaction for it,
so it timed out and HPMC'd the box.
what you really need is the IIR, and the address it tried to access
(both the kernel vaddr which will be in the register, and the "system
requester address" from the hpmc dump which will be the physical address
mapped.
not sure why the hpmc handler is getting skipped, that's a little weird.
you can try hacking elroy to set softfail mode on that bus, which will
result in a timeout on the pci bus to return -1 (like what x86 and most
other architectures do) rather than hang the box, but it really likely
means a driver bug.
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-20 22:16 ` Kyle McMartin
@ 2008-12-20 22:21 ` Kyle McMartin
2008-12-21 8:12 ` Grant Grundler
0 siblings, 1 reply; 20+ messages in thread
From: Kyle McMartin @ 2008-12-20 22:21 UTC (permalink / raw)
To: Guy Martin; +Cc: linux-parisc
On Sat, Dec 20, 2008 at 05:16:53PM -0500, Kyle McMartin wrote:
> you can try hacking elroy to set softfail mode on that bus, which will
> result in a timeout on the pci bus to return -1 (like what x86 and most
> other architectures do) rather than hang the box, but it really likely
> means a driver bug.
untested and all that jazz.
diff --git a/drivers/parisc/lba_pci.c b/drivers/parisc/lba_pci.c
index a28c894..a34f759 100644
--- a/drivers/parisc/lba_pci.c
+++ b/drivers/parisc/lba_pci.c
@@ -1341,7 +1341,7 @@ lba_hw_init(struct lba_device *d)
/* Set HF mode as the default (vs. -1 mode). */
stat = READ_REG32(d->hba.base_addr + LBA_STAT_CTL);
- WRITE_REG32(stat | HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL);
+ WRITE_REG32(stat & ~HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL);
/*
** Writing a zero to STAT_CTL.rf (bit 0) will clear reset signal
^ permalink raw reply related [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin
2008-12-20 22:16 ` Kyle McMartin
@ 2008-12-21 8:10 ` Grant Grundler
1 sibling, 0 replies; 20+ messages in thread
From: Grant Grundler @ 2008-12-21 8:10 UTC (permalink / raw)
To: Guy Martin; +Cc: linux-parisc
On Sat, Dec 20, 2008 at 01:39:15PM +0100, Guy Martin wrote:
> I'm recently running into HPMC bus timeout problem. While this never
> caused any problem before, I've noticed a few strange things. I've
> attached the output I get from "pim hpmc" in the PDC.
>
> This BUS timeout occurs after one or two days of idling when I plug a
> specific card reader in the USB PCI card I plugged in. I think this
> card reader causes problems because it's being polled every 2-3 seconds.
>
> Nevertheless, looking at the pdc output, it seems that the OS HPMC
> handler is not kicking off according to the chassis code CBF2 and CBFC.
> CBF2 : bas OS HPMC len
> CBFC : OS HPMC br err
>
> I've check GR02 and it points to inw() which makes sens. Nothing
> interesting there.
>
> The failing device appears to be the built-in NIC according the path
> provided in the output (10/0/12/0) and the C3600 service manual (figure
> 5-2).
It's likely the NIC (tulip driver) is just more active than the other
USB controller and thus becomes the "victim" of the USB misbehavior.
This is very common for DMA programming bugs.
"word2" from this bit of the HPMC needs to be decoded:
I/O Module Error Log Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes:
Rope Word1 Word2 Word3
------ ------------ ------------
0 0x00000000 0x0e0cc2a9 0x00000000fed30048
...
Both the USB controller and NIC are on Rope 0. If the USB controller
falls over, it will take all devices on that bus with it.
We've seen "0x0e0cc2a9" before but it wasn't decoded then either.
> Even if this look HW pb, I've had a few reports recently about HPMC PCI
> timeout issues and I have a few doubts.
>
> Now my questions are :
> - does this look like an HW or SW problems ?
More likely SW.
> - why isn't the HPMC handler kicked off ?
No Idea. Probably not hooked in correctly.
But this might also explain why you haven't seen many HPMC bug reports.
> - could the HPMC handler recover this ?
No.
> - what debug can I enable to get more info about this if applicable ?
I'd look carefully at USB DMA streams. This will likely require
adding debug code to USB drivers. Mostly to dump when a USB buffer is
mapped and when it's unmap.
hth,
grant
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-20 22:21 ` Kyle McMartin
@ 2008-12-21 8:12 ` Grant Grundler
2008-12-22 21:40 ` Guy Martin
0 siblings, 1 reply; 20+ messages in thread
From: Grant Grundler @ 2008-12-21 8:12 UTC (permalink / raw)
To: Kyle McMartin; +Cc: Guy Martin, linux-parisc
On Sat, Dec 20, 2008 at 05:21:00PM -0500, Kyle McMartin wrote:
> On Sat, Dec 20, 2008 at 05:16:53PM -0500, Kyle McMartin wrote:
> > you can try hacking elroy to set softfail mode on that bus, which will
> > result in a timeout on the pci bus to return -1 (like what x86 and most
> > other architectures do) rather than hang the box, but it really likely
> > means a driver bug.
>
> untested and all that jazz.
>
> diff --git a/drivers/parisc/lba_pci.c b/drivers/parisc/lba_pci.c
> index a28c894..a34f759 100644
> --- a/drivers/parisc/lba_pci.c
> +++ b/drivers/parisc/lba_pci.c
> @@ -1341,7 +1341,7 @@ lba_hw_init(struct lba_device *d)
>
> /* Set HF mode as the default (vs. -1 mode). */
> stat = READ_REG32(d->hba.base_addr + LBA_STAT_CTL);
> - WRITE_REG32(stat | HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL);
> + WRITE_REG32(stat & ~HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL);
I'm not sure how helpful this will be. It will be unpredictable how
the tulip driver will handle getting -1's back from inw().
thanks,
grant
>
> /*
> ** Writing a zero to STAT_CTL.rf (bit 0) will clear reset signal
> --
> To unsubscribe from this list: send the line "unsubscribe linux-parisc" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-21 8:12 ` Grant Grundler
@ 2008-12-22 21:40 ` Guy Martin
2008-12-22 21:47 ` Kyle McMartin
0 siblings, 1 reply; 20+ messages in thread
From: Guy Martin @ 2008-12-22 21:40 UTC (permalink / raw)
To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc
[-- Attachment #1: Type: text/plain, Size: 1928 bytes --]
Kyle, Grant,
Thanks for your help but being unable to reproduce the issue at will
makes it very difficult to make any progress. The fault didn't occurred
those past two days.
So I decided to go another way. I've plugged in one of my well working
sata card and did a lot of IO using "dd if=/dev/zero of=/dev/sdc".
I'm not sure if this would trigger the same problem but it does trigger
an HPMC Bus timeout too after a few hundreds megs. This time, the PCI
slot pointed by the output is the sata card itself in slot 4.
Regarding GR02, it seems to point somewhere in user space.
Regarding Kyle's patch, it didn't changed anything but since the fault
isn't at inw(), I'm not sure it still even makes sens to apply it.
I tried to find out why the HPMC os handler wasn't triggered.
As documented in pdc20-v1.1-Ch3-pdce.pdf, the interruption vector table
is setup correctly. By this I mean that CR 14 points correctly to the
beginning of the fault_vector_20 (physical addr), the checksum is ok
and the length seems about right. I used a quickly coded module to
check all of this at run time.
The only thing that doesn't seem to be done is calling PDC_INSTR but I
still need to investigate that.
As far as I understand it's should provide instr to jump to the
os_hpmc handler. Those instr are already harcoded in the fault vector
so it's shouldn't be needed, unless calling PDC_INSTR actually does
tell the PDC that os_hpmc has been installed.
So again, I'm looking for help on going further with this issue.
Unfortunately, I don't have a lot of knowledge about PCI and the HP
side of the elroy chip. I wanted to look at elroy's documentation but
it seems that a large part was stripped down.
Any pointer to usefull doc is more than welcome !
I've attached two HPMC dumps generated using the same command. Hopefull
they'll provide more details.
This is on an C3600 running latest parisc-2.6 git tree.
Regards,
Guy
[-- Attachment #2: c3600-sata-pci-hpmc2 --]
[-- Type: application/octet-stream, Size: 5212 bytes --]
PROCESSOR PIM INFORMATION
----------------- Processor 0 HPMC Information ------------------
Timestamp =
Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13)
HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc
General Registers 0 - 31
00-03 0000000000000000 0000000000650000 00000000006500f4 000000008cc2c000
04-07 000000008dfb8828 000000008dfb8834 0000000000000001 000000008cc2c000
08-11 000000000000000f 000000008dfb8834 000000008dfb8844 000000000000001e
12-15 000000008dfb8800 000000008cc2de28 0000000000000005 00000000000000d0
16-19 0000000000002002 00000000f000017c 00000000f0000174 0000000000400000
20-23 0000000000000000 0000000000682000 0000000000679b9c 000000000000002f
24-27 000000006b719d0a 000000008cc2ddd0 000000008cc2c000 00000000104c78b0
28-31 0000000000679d3c 000000000001377c 000000008cc1c240 00000000006500f4
<Press any key to continue (q to quit)>
Control Registers 0 - 31
00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000
04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000
08-11 0000000000000e00 0000000000000000 00000000000000c0 000000000000003f
12-15 0000000000000000 0000000000000000 000000000010e000 00000000ffc00000
16-19 00000037fe47b62e 0000000000000000 0000000000679bd8 000000004abc0090
20-23 00000000a627ffee 0000000001e82048 000000000004000e 0000000080000000
24-27 000000000051a000 000000007de58000 0000000000044021 00000000f0412000
28-31 0000000055555555 0000000055555555 000000008cc1c000 0000000011111111
Space Registers 0 - 7
00-03 00000000 00000000 00000000 00000700
04-07 00000000 00000000 00000000 00000000
IIA Space = 0x0000000000000000
IIA Offset = 0x0000000000679bdc
Check Type = 0x20000000
CPU State = 0x9e000004
Cache Check = 0x00000000
TLB Check = 0x00000000
Bus Check = 0x0030103b
Assists Check = 0x00000000
Assist State = 0x00000000
Path Info = 0x00000000
System Responder Address = 0x000000fffb807048
System Requestor Address = 0xfffffffffffa0000
Floating-Point Registers 0 - 31
00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000
04-07 3febb8de8e377300 3febb8de8e377300 0000000100000000 0000000000000000
08-11 0000000000000000 000000028f812bc8 8f812bc000000002 1056f0b000000003
12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555
16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555
20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000
24-27 0000000000000000 0000ff9700000000 0000000000000000 10282858102b7368
28-31 ffffffff000137f2 104e0da4101871a0 0000000100000228 8f820200101165f4
'9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes:
Check Summary = 0xcb81041008000000
Available Memory = 0x0000000080000000
CPU Diagnose Register 2 = 0x0300000000000004
CPU Status Register 0 = 0x2420c20000000000
CPU Status Register 1 = 0x8002000000000000
SADD LOG = 0xc10f00fffb807048
Read Short LOG = 0xc1af00fffb807048
ERROR_STATUS = 0x0000000000100010
MEM_ADDR = 0x000001ff3fffffff
MEM_SYND = 0x0000000000000000
MEM_ADDR_CORR = 0x000001ff3fffffff
MEM_SYND_CORR = 0x0000000000000000
RUN_DATA_HIGH = 0xc1bff0fffed08040
RUN_DATA_LOW = 0xc1bff0fffed08040
RUN_CTRL = 0x0000021c00001418
RUN_ADDR = 0xc1bff0fffed08040
System Responder Path = 0x00ffffff0a010400
HPMC PIM Analysis Information:
Timestamp =
Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13)
'9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes:
A Data I/O Fetch Timeout occurred while CPU 0 was
requesting information from a device at the path 10/1/4/0 (PCI slot 4).
Memory/IO Controller Error Analysis Information:
The Memory/IO Controller only observed the Broadcast Error. It did not log
any additional information about the HPMC.
Memory Error Log Information:
Timestamp =
Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13)
'9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes:
No memory errors logged
I/O Module Error Log Information:
Timestamp =
Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13)
'9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes:
Rope Word1 Word2 Word3
------ ------------ ------------
0 0x00000000 0x0e0cc009 0x00000000fed30048
1 0x00000000 0x1e0cc289 0x00000000fed32048
2 ---------- 0x2e0cc009 ------------------
3 ---------- 0x3e0cc009 ------------------
4 0x00000000 0x4e0cc009 0x00000000fed38048
5 ---------- 0x5e0cc009 ------------------
6 0x00000000 0x6e0cc009 0x00000000fed3c048
7 ---------- 0x7e0cc009 ------------------
[-- Attachment #3: c3600-sata-pci-hpmc --]
[-- Type: application/octet-stream, Size: 5211 bytes --]
PROCESSOR PIM INFORMATION
----------------- Processor 0 HPMC Information ------------------
Timestamp =
Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48)
HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc
General Registers 0 - 31
00-03 0000000000000000 0000000000650000 00000000006500f4 000000002f084000
04-07 000000008cc1e828 000000008cc1e834 0000000000000001 000000002f084000
08-11 000000000000000f 000000008cc1e834 000000008cc1e844 000000000000001e
12-15 000000008cc1e800 000000002f085e28 0000000000000005 00000000000000d0
16-19 0000000000002002 00000000f000017c 00000000f0000174 0000000000400000
20-23 0000000000000000 0000000000682000 0000000000679b9c 000000000000004c
24-27 000000000a0ee5b0 000000002f085dd0 000000002f084000 00000000104c78b0
28-31 0000000000679d3c 000000000021f42e 000000002f284240 00000000006500f4
<Press any key to continue (q to quit)>
Control Registers 0 - 31
00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000
04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000
08-11 0000000000000f9a 0000000000000000 00000000000000c0 000000000000003f
12-15 0000000000000000 0000000000000000 000000000010e000 00000000ffc00000
16-19 000003a2c566501d 0000000000000000 0000000000679bd8 000000004abc0090
20-23 00000000a627ffee 0000000001e82048 000000000006000e 0000000080000000
24-27 000000000051a000 000000007df68000 0000000000044021 00000000f0412000
28-31 0000000055555555 0000000055555555 000000002f284000 0000000011111111
Space Registers 0 - 7
00-03 00000000 000007cd 00000000 000007cd
04-07 00000000 00000000 00000000 00000000
IIA Space = 0x0000000000000000
IIA Offset = 0x0000000000679bdc
Check Type = 0x20000000
CPU State = 0x9e000004
Cache Check = 0x00000000
TLB Check = 0x00000000
Bus Check = 0x0030103b
Assists Check = 0x00000000
Assist State = 0x00000000
Path Info = 0x00000000
System Responder Address = 0x000000fffb807048
System Requestor Address = 0xfffffffffffa0000
Floating-Point Registers 0 - 31
00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000
04-07 3e5058af0f4dd252 bf4e68a0d349be90 0000000000000000 bf4e6eea80000000
08-11 0000000538400000 000000028f812bc8 8f812bc000000002 1056f0b000000003
12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555
16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555
20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000
24-27 0000000000000000 0000ff97b22a09b6 0000000000000000 3ff0000000000000
28-31 00100000000137f2 104e0da4101871a0 0000000100000228 8f820200101165f4
'9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes:
Check Summary = 0xcb81041008000000
Available Memory = 0x0000000080000000
CPU Diagnose Register 2 = 0x0300000000000004
CPU Status Register 0 = 0x2420c20000000000
CPU Status Register 1 = 0x8002000000000000
SADD LOG = 0xc10f00fffb807048
Read Short LOG = 0xc1af00fffb807048
ERROR_STATUS = 0x0000000000100010
MEM_ADDR = 0x000001ff3fffffff
MEM_SYND = 0x0000000000000000
MEM_ADDR_CORR = 0x000001ff3fffffff
MEM_SYND_CORR = 0x0000000000000000
RUN_DATA_HIGH = 0xc1bff0fffed08040
RUN_DATA_LOW = 0xc1bff0fffed08040
RUN_CTRL = 0x0000021c00001418
RUN_ADDR = 0xc1bff0fffed08040
System Responder Path = 0x00ffffff0a010400
HPMC PIM Analysis Information:
Timestamp =
Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48)
'9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes:
A Data I/O Fetch Timeout occurred while CPU 0 was
requesting information from a device at the path 10/1/4/0 (PCI slot 4).
Memory/IO Controller Error Analysis Information:
The Memory/IO Controller only observed the Broadcast Error. It did not log
any additional information about the HPMC.
Memory Error Log Information:
Timestamp =
Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48)
'9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes:
No memory errors logged
I/O Module Error Log Information:
Timestamp =
Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48)
'9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes:
Rope Word1 Word2 Word3
------ ------------ ------------
0 0x00000000 0x0e0cc009 0x00000000fed30048
1 0x0001c800 0x1e0cc009 0x00000000fb807000
2 ---------- 0x2e0cc009 ------------------
3 ---------- 0x3e0cc009 ------------------
4 0x00000000 0x4e0cc009 0x00000000fed38048
5 ---------- 0x5e0cc009 ------------------
6 0x00000000 0x6e0cc009 0x00000000fed3c048
7 ---------- 0x7e0cc009 ------------------
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-22 21:40 ` Guy Martin
@ 2008-12-22 21:47 ` Kyle McMartin
2008-12-22 22:49 ` Guy Martin
0 siblings, 1 reply; 20+ messages in thread
From: Kyle McMartin @ 2008-12-22 21:47 UTC (permalink / raw)
To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc
On Mon, Dec 22, 2008 at 10:40:23PM +0100, Guy Martin wrote:
>
> Kyle, Grant,
>
> Thanks for your help but being unable to reproduce the issue at will
> makes it very difficult to make any progress. The fault didn't occurred
> those past two days.
>
> So I decided to go another way. I've plugged in one of my well working
> sata card and did a lot of IO using "dd if=/dev/zero of=/dev/sdc".
>
> I'm not sure if this would trigger the same problem but it does trigger
> an HPMC Bus timeout too after a few hundreds megs. This time, the PCI
> slot pointed by the output is the sata card itself in slot 4.
> Regarding GR02, it seems to point somewhere in user space.
>
> Regarding Kyle's patch, it didn't changed anything but since the fault
> isn't at inw(), I'm not sure it still even makes sens to apply it.
>
you won't know for sure, since readl/etc are inlines into the caller...
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-22 21:47 ` Kyle McMartin
@ 2008-12-22 22:49 ` Guy Martin
2008-12-25 8:06 ` Grant Grundler
0 siblings, 1 reply; 20+ messages in thread
From: Guy Martin @ 2008-12-22 22:49 UTC (permalink / raw)
To: Kyle McMartin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc
[-- Attachment #1: Type: text/plain, Size: 721 bytes --]
On Mon, 22 Dec 2008 16:47:43 -0500
Kyle McMartin <kyle@infradead.org> wrote:
> > Regarding Kyle's patch, it didn't changed anything but since the
> > fault isn't at inw(), I'm not sure it still even makes sens to
> > apply it.
> >
>
> you won't know for sure, since readl/etc are inlines into the
> caller...
It seems that I forgot to include your patch when I switched to the
git tree. I've been doing tests again with your patch and things are
different :
The command "dd if=/dev/zero of=/dev/sdc bs=1M" didn't trigger HPMC
anymore after a reasonable amount of time.
I then tried a NFS copy from a NFS share to the disk. The kernel did
panic with the attached output.
Hopefully this will be more useful.
Guy
[-- Attachment #2: c3600-sata-pci-panic --]
[-- Type: application/octet-stream, Size: 5875 bytes --]
[ 1577.999842] ata1.00: exception Emask 0x52 SAct 0x0 SErr 0xffffffff action 0xe
[ 1577.999842] ata1: SError: { RecovData RecovComm UnrecovData Persist Proto HostInt PHYRdyChg PHYInt CommWake 10B8B Dispar BadCRC Handshk LinkSeq TrStaTrns UnrecFIS DevExch }
[ 1577.999842] ata1.00: cmd 35/00:00:bf:1f:2f/00:04:00:00:00/e0 tag 0 dma 524288 out
[ 1577.999842] res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x72 (host bus error)
[ 1577.999842] ata1.00: status: { DRDY }
[ 1577.999842] ata1: hard resetting link
[ 1578.719842] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF)
[ 1583.719841] ata1: hard resetting link
[ 1584.039841] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF)
[ 1589.039841] ata1: hard resetting link
[ 1589.359841] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF)
[ 1589.359841] ata1.00: disabled
[ 1589.359841] sd 2:0:0:0: [sdc] Result: hostbyte=0x00 driverbyte=0x08
[ 1589.359841] sd 2:0:0:0: [sdc] Sense Key : 0xb [current] [descriptor]
[ 1589.359841] Descriptor sense data with sense descriptors (in hex):
[ 1589.359841] 72 0b 00 00 00 00 00 0c 00 0a 80 00 00 00 00 00
[ 1589.359841] 00 00 00 00
[ 1589.359841] sd 2:0:0:0: [sdc] ASC=0x0 ASCQ=0x0
[ 1589.359841] end_request: I/O error, dev sdc, sector 3088319
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386032
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386033
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386034
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386035
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386036
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386037
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386038
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386039
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386040
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] Buffer I/O error on device sdc1, logical block 386041
[ 1589.359841] lost page write due to I/O error on sdc1
[ 1589.359841] sd 2:0:0:0: rejecting I/O to offline device
[ 1589.359841] ata1: EH complete
[ 1589.359841] ata1.00: detaching (SCSI 2:0:0:0)
[ 1589.359841] sd 2:0:0:0: [sdc] Result: hostbyte=0x01 driverbyte=0x00
[ 1589.359841] end_request: I/O error, dev sdc, sector 3089343
[ 1589.403174] EXT3-fs error (device sdc1): read_block_bitmap: Cannot read block bitmap - block_group = 13, block_bitmap = 425984
[ 1589.409841] EXT3-fs error (device sdc1): read_block_bitmap: Cannot read block bitmap - block_group = 13, block_bitmap = 425984
[ 1589.409841] ------------[ cut here ]------------
[ 1589.466507] Badness at fs/buffer.c:1186
[ 1589.466507]
[ 1589.466507] YZrvWESTHLNXBCVMcbcbcbcbOGFRQPDI
[ 1589.466507] PSW: 00000000000001000000001100001111 Not tainted
[ 1589.466507] r00-03 0004030f 1056f8b0 101d0ccc 8e908400
[ 1589.466507] r04-07 3cf97400 8f510afc 00000001 3cf961a0
[ 1589.466507] r08-11 8e908400 7226c7d4 8ded3f00 00000000
[ 1589.466507] r12-15 8eb813c0 00000001 492ecb30 492ecb30
[ 1589.466507] r16-19 7226c8c8 5195f3ec 00068000 dfd101d1
[ 1589.466507] r20-23 3cf6f000 0000001d 00040000 00000e8e
[ 1589.466507] r24-27 00000000 00000e8f 8f510afc 104c78b0
[ 1589.466507] r28-31 00000000 00000006 7226cb00 fffffff7
[ 1589.466507] sr00-03 00000000 000015ef 00000000 00000035
[ 1589.466507] sr04-07 00000000 00000000 00000000 00000000
[ 1589.466507]
[ 1589.466507] IASQ: 00000000 00000000 IAOQ: 1019cc88 1019cc8c
[ 1589.466507] IIR: 03ffe01f ISR: 00000000 IOR: 00000000
[ 1589.466507] CPU: 0 CR30: 7226c000 CR31: 11111111
[ 1589.466507] ORIG_R28: 492ecb30
[ 1589.466507] IAOQ[0]: mark_buffer_dirty+0x84/0xa0
[ 1589.466507] IAOQ[1]: mark_buffer_dirty+0x88/0xa0
[ 1589.466507] RP(r2): ext3_commit_super+0x7c/0xb8
[ 1589.466507] Backtrace:
[ 1589.466507] [<101d0ccc>] ext3_commit_super+0x7c/0xb8
[ 1589.466507] [<101d1cc0>] ext3_handle_error+0x7c/0xc4
[ 1589.466507] [<101d1df8>] ext3_error+0x64/0x78
[ 1589.466507] [<101c4168>] read_block_bitmap+0xb4/0x1d4
[ 1589.466507] [<101c5294>] ext3_new_blocks+0x3bc/0x6d8
[ 1589.466507] [<101c8ecc>] ext3_get_blocks_handle+0x380/0xa48
[ 1589.466507] [<101c97f0>] ext3_get_block+0x7c/0x10c
[ 1589.466507] [<1019dc68>] __block_prepare_write+0x1e0/0x3d0
[ 1589.466507] [<1019df10>] block_write_begin+0x68/0x138
[ 1589.466507] [<101cb238>] ext3_write_begin+0x100/0x270
[ 1589.466507] [<10154b9c>] generic_file_buffered_write+0x12c/0x314
[ 1589.466507] [<101552d8>] __generic_file_aio_write_nolock+0x2c4/0x4a4
[ 1589.466507] [<10155538>] generic_file_aio_write+0x80/0x10c
[ 1589.466507] [<101c6524>] ext3_file_write+0x34/0xd8
[ 1589.466507] [<1017b354>] do_sync_write+0xec/0x144
[ 1589.466507] [<1017bd28>] vfs_write+0x98/0x13c
[ 1589.466507]
[ 1589.473174] JBD: Detected IO errors while flushing file data on sdc1
[ 1589.473174] Aborting journal on device sdc1.
[ 1589.526507] ext3_abort called.
[ 1589.526507] EXT3-fs error (device sdc1): ext3_journal_start_sb: Detected aborted journal
[ 1589.526507] Remounting filesystem read-only
[ 1589.939841] JBD: Detected IO errors while flushing file data on sdc1
[ 1589.939841] __journal_remove_journal_head: freeing b_committed_data
[ 1590.556507] sd 2:0:0:0: [sdc] Stopping disk
[ 1590.563174] sd 2:0:0:0: [sdc] START_STOP FAILED
[ 1590.563174] sd 2:0:0:0: [sdc] Result: hostbyte=0x04 driverbyte=0x00
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-22 22:49 ` Guy Martin
@ 2008-12-25 8:06 ` Grant Grundler
2008-12-25 15:57 ` Guy Martin
0 siblings, 1 reply; 20+ messages in thread
From: Grant Grundler @ 2008-12-25 8:06 UTC (permalink / raw)
To: Guy Martin; +Cc: Kyle McMartin, Grant Grundler, linux-parisc
On Mon, Dec 22, 2008 at 11:49:48PM +0100, Guy Martin wrote:
> The command "dd if=/dev/zero of=/dev/sdc bs=1M" didn't trigger HPMC
> anymore after a reasonable amount of time.
>
> I then tried a NFS copy from a NFS share to the disk. The kernel did
> panic with the attached output.
[ 1589.466507] Badness at fs/buffer.c:1186
Technically, this isn't a panic. This is a "WARN_ON".
I didn't see any panic's after that either.
The file system warning us that the sata disk it was
talking failed an IO.
However, this is likely to be some other issue with the SATA controller.
Can you post more details about the config?
o "lspci -v"
o hdparm -i /dev/sd<X>
I don't expect I'll know what's wrong, but just want to document the bug.
thanks,
grant
> Hopefully this will be more useful.
>
> Guy
>
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-25 8:06 ` Grant Grundler
@ 2008-12-25 15:57 ` Guy Martin
2008-12-25 21:36 ` Grant Grundler
0 siblings, 1 reply; 20+ messages in thread
From: Guy Martin @ 2008-12-25 15:57 UTC (permalink / raw)
To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc
On Thu, 25 Dec 2008 01:06:06 -0700
Grant Grundler <grundler@parisc-linux.org> wrote:
> [ 1589.466507] Badness at fs/buffer.c:1186
>
> Technically, this isn't a panic. This is a "WARN_ON".
> I didn't see any panic's after that either.
Yes sorry. The system is still usable after the WARN_ON occurs. So no
panic.
Because of this, wouldn't it be a good idea to commit Kyle's patch ?
To me it seems better to have a running system with failure messages
being log rather than an ugly and barely understandable HPMC.
> The file system warning us that the sata disk it was
> talking failed an IO.
>
> However, this is likely to be some other issue with the SATA
> controller. Can you post more details about the config?
> o "lspci -v"
> o hdparm -i /dev/sd<X>
Here it is :
01:04.0 Mass storage controller: Silicon Image, Inc. SiI 3112 [SATALink/SATARaid] Serial ATA Controller (rev 02)
Subsystem: Silicon Image, Inc. SiI 3112 SATALink Controller
Flags: bus master, 66MHz, medium devsel, latency 240, IRQ 21
I/O ports at 12400 [size=8]
I/O ports at 12300 [size=4]
I/O ports at 12200 [size=8]
I/O ports at 12100 [size=4]
I/O ports at 12000 [size=16]
Memory at fb807000 (32-bit, non-prefetchable) [size=512]
Expansion ROM at fb880000 [disabled] [size=512K]
Capabilities: [60] Power Management version 2
Kernel modules: sata_sil
/dev/sdc:
Model=WDC WD5000AAKS-00TMA0 , FwRev=12.01C01, SerialNo= WD-WMAPW1390228
Config={ HardSect NotMFM HdSw>15uSec SpinMotCtl Fixed DTR>5Mbs FmtGapReq }
RawCHS=16383/16/63, TrkSize=0, SectSize=0, ECCbytes=50
BuffType=unknown, BuffSize=16384kB, MaxMultSect=16, MultSect=?0?
CurCHS=16383/16/63, CurSects=16514064, LBA=yes, LBAsects=976773168
IORDY=on/off, tPIO={min:120,w/IORDY:120}, tDMA={min:120,rec:120}
PIO modes: pio0 pio3 pio4
DMA modes: mdma0 mdma1 mdma2
UDMA modes: udma0 udma1 udma2 udma3 udma4 *udma5 udma6
AdvancedPM=no WriteCache=disabled
Drive conforms to: Unspecified: ATA/ATAPI-1,2,3,4,5,6,7
* signifies the current active mode
I'd like to add that I've been using the exact same card and hard drive on one of my x86 box for month without any issue.
Also I've reproduce the problem again and this time I've had this message right before the "end_request" line :
[ 163.039983] timer_interrupt(CPU 0): delayed! cycles EB5EE1E6 rem 36AE6 next/now 706F7328/5BCE550E
Anything else I can do/provide to troubleshoot this ?
Cheers,
Guy
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-25 15:57 ` Guy Martin
@ 2008-12-25 21:36 ` Grant Grundler
2009-01-02 20:46 ` Guy Martin
0 siblings, 1 reply; 20+ messages in thread
From: Grant Grundler @ 2008-12-25 21:36 UTC (permalink / raw)
To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc
On Thu, Dec 25, 2008 at 04:57:50PM +0100, Guy Martin wrote:
> On Thu, 25 Dec 2008 01:06:06 -0700
> Grant Grundler <grundler@parisc-linux.org> wrote:
>
> > [ 1589.466507] Badness at fs/buffer.c:1186
> >
> > Technically, this isn't a panic. This is a "WARN_ON".
> > I didn't see any panic's after that either.
>
> Yes sorry. The system is still usable after the WARN_ON occurs. So no
> panic.
>
> Because of this, wouldn't it be a good idea to commit Kyle's patch ?
I don't think kyle intended to commit that patch. He just wanted to
collect more info.
HPMC is 99% (for parisc-linux at least) of the time a driver bug.
So having the system crash on IO errors (including DMA map/unmap bugs)
is good for getting bugs reported and the state of the IOMMU and PCI
Host controller to debug the problem.
On the other hand, users often don't care about those details even if they
see (like you did) the device isn't working right. They can more easily
track down which device/drivers are having problems and remove them from
the config.
> To me it seems better to have a running system with failure messages
> being log rather than an ugly and barely understandable HPMC.
While I agree with your characterization of HPMCs, I don't want to trade
HPMC dumps for kernel logs. HPMC provides info we can't otherwise get.
HPMCs might be more useful if the symbol name of every kernel address
in the dump were printed. Given the System.map, it should be
possible to do something like:
hpmc_symbols System.map < HPMC_dump.txt > HPMC_symbols.txt
The program "a.c" already exists to do a symbol lookup given an kernel
address and System.map file:
http://cvs.parisc-linux.org/build-tools/
> > The file system warning us that the sata disk it was
> > talking failed an IO.
> >
> > However, this is likely to be some other issue with the SATA
> > controller. Can you post more details about the config?
> > o "lspci -v"
> > o hdparm -i /dev/sd<X>
>
> Here it is :
>
> 01:04.0 Mass storage controller: Silicon Image, Inc. SiI 3112 [SATALink/SATARaid] Serial ATA Controller (rev 02)
Ok...so this is the sata_sil driver.
> Subsystem: Silicon Image, Inc. SiI 3112 SATALink Controller
> Flags: bus master, 66MHz, medium devsel, latency 240, IRQ 21
> I/O ports at 12400 [size=8]
> I/O ports at 12300 [size=4]
> I/O ports at 12200 [size=8]
> I/O ports at 12100 [size=4]
> I/O ports at 12000 [size=16]
> Memory at fb807000 (32-bit, non-prefetchable) [size=512]
> Expansion ROM at fb880000 [disabled] [size=512K]
> Capabilities: [60] Power Management version 2
> Kernel modules: sata_sil
>
>
> /dev/sdc:
>
> Model=WDC WD5000AAKS-00TMA0 , FwRev=12.01C01, SerialNo= WD-WMAPW1390228
> Config={ HardSect NotMFM HdSw>15uSec SpinMotCtl Fixed DTR>5Mbs FmtGapReq }
> RawCHS=16383/16/63, TrkSize=0, SectSize=0, ECCbytes=50
> BuffType=unknown, BuffSize=16384kB, MaxMultSect=16, MultSect=?0?
> CurCHS=16383/16/63, CurSects=16514064, LBA=yes, LBAsects=976773168
> IORDY=on/off, tPIO={min:120,w/IORDY:120}, tDMA={min:120,rec:120}
> PIO modes: pio0 pio3 pio4
> DMA modes: mdma0 mdma1 mdma2
> UDMA modes: udma0 udma1 udma2 udma3 udma4 *udma5 udma6
> AdvancedPM=no WriteCache=disabled
> Drive conforms to: Unspecified: ATA/ATAPI-1,2,3,4,5,6,7
> * signifies the current active mode
Ok - looks normal.
> I'd like to add that I've been using the exact same card and hard drive
> on one of my x86 box for month without any issue.
That doesn't mean the driver and chip operate 100% correctly.
Has anyone run exhaustive tests to detect data corruption with this card?
Looking at the sil_interrupt() code, it seems a PCI Master abort was
sometimes expected when reading the bmdma2 register:
u32 bmdma2 = readl(mmio_base + sil_port[ap->port_no].bmdma2);
...
if (bmdma2 == 0xffffffff ||
!(bmdma2 & (SIL_DMA_COMPLETE | SIL_DMA_SATA_IRQ)))
continue;
Also, older X86 platforms generally don't have an IOMMU (newer ones will)
and thus can't validate DMA transactions. I don't have the impression that's
the problem here though. But someone needs to decode the "Word2" of the
HPMC dump that you already provided.
> Also I've reproduce the problem again and this time I've had this message right before the "end_request" line :
> [ 163.039983] timer_interrupt(CPU 0): delayed! cycles EB5EE1E6 rem 36AE6 next/now 706F7328/5BCE550E
Error handlers sometimes don't play nicely with interrupts.
I don't know enough about the error handling cases to track this down.
> Anything else I can do/provide to troubleshoot this ?
Two things:
o consider posting some of the original findings on linux-ide and see if
anyone has tested this controller on PPC or IA64. I'm looking for any
other architecture that has "hard fail" behavior like parisc does.
Testing on any other Big Endian HW would be worth hearing about too.
o write a quick and dirty "hpmc_symbols" script as described above and
run it on the HPMC you provided earlier.
Debugging this further wil probably require modifying the sata_sil driver
to log (e.g. ktrace) it's activities while under test.
hth,
grant
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2008-12-25 21:36 ` Grant Grundler
@ 2009-01-02 20:46 ` Guy Martin
2009-01-02 21:01 ` Kyle McMartin
0 siblings, 1 reply; 20+ messages in thread
From: Guy Martin @ 2009-01-02 20:46 UTC (permalink / raw)
To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc
[-- Attachment #1: Type: text/plain, Size: 1040 bytes --]
On Thu, 25 Dec 2008 14:36:25 -0700
Grant Grundler <grundler@parisc-linux.org> wrote:
> > Anything else I can do/provide to troubleshoot this ?
>
> Two things:
> o consider posting some of the original findings on linux-ide and see
> if anyone has tested this controller on PPC or IA64. I'm looking for
> any other architecture that has "hard fail" behavior like parisc does.
> Testing on any other Big Endian HW would be worth hearing about too.
>
> o write a quick and dirty "hpmc_symbols" script as described above
> and run it on the HPMC you provided earlier.
>
> Debugging this further wil probably require modifying the sata_sil
> driver to log (e.g. ktrace) it's activities while under test.
Will do, just haven't had time yet.
I quickly looked into why the OS HPMC handler wasn't triggered and the
issue comes from the fact that the length provided to the PDC is
incorrect. The doc states that it should be in bytes while we provide it
in words.
I've attached a small patch that solves this little issue.
Cheers,
Guy
[-- Attachment #2: fix-hpmc-handler-length.diff --]
[-- Type: text/x-patch, Size: 433 bytes --]
OS HPMC handler length needs to be in bytes, not in words.
Signed-off-by: Guy Martin <gmsoft@tuxicoman.be>
--- /root/traps.c.orig 2009-01-02 09:20:41.000000000 +0100
+++ arch/parisc/kernel/traps.c 2009-01-02 09:23:47.000000000 +0100
@@ -840,7 +840,7 @@
/* Compute Checksum for HPMC handler */
- length = os_hpmc_end - os_hpmc;
+ length = (os_hpmc_end - os_hpmc) * sizeof(u32);
ivap[7] = length;
hpmcp = (u32 *)os_hpmc;
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-02 20:46 ` Guy Martin
@ 2009-01-02 21:01 ` Kyle McMartin
2009-01-02 21:14 ` John David Anglin
0 siblings, 1 reply; 20+ messages in thread
From: Kyle McMartin @ 2009-01-02 21:01 UTC (permalink / raw)
To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc
On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote:
> OS HPMC handler length needs to be in bytes, not in words.
> Signed-off-by: Guy Martin <gmsoft@tuxicoman.be>
>
> - length = os_hpmc_end - os_hpmc;
> + length = (os_hpmc_end - os_hpmc) * sizeof(u32);
> ivap[7] = length;
Ick. I wonder if I screwed that up when I fixed the function descriptor
breakage...
Will apply this fix, thanks Guy.
cheers, Kyle
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-02 21:01 ` Kyle McMartin
@ 2009-01-02 21:14 ` John David Anglin
2009-01-02 21:18 ` Kyle McMartin
0 siblings, 1 reply; 20+ messages in thread
From: John David Anglin @ 2009-01-02 21:14 UTC (permalink / raw)
To: Kyle McMartin; +Cc: gmsoft, grundler, kyle, linux-parisc
> On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote:
> > OS HPMC handler length needs to be in bytes, not in words.
> > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be>
> >
> > - length = os_hpmc_end - os_hpmc;
> > + length = (os_hpmc_end - os_hpmc) * sizeof(u32);
> > ivap[7] = length;
>
> Ick. I wonder if I screwed that up when I fixed the function descriptor
> breakage...
The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are
pointing to function descriptors.
Dave
--
J. David Anglin dave.anglin@nrc-cnrc.gc.ca
National Research Council of Canada (613) 990-0752 (FAX: 952-6602)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-02 21:14 ` John David Anglin
@ 2009-01-02 21:18 ` Kyle McMartin
2009-01-02 21:36 ` John David Anglin
2009-01-03 0:09 ` John David Anglin
0 siblings, 2 replies; 20+ messages in thread
From: Kyle McMartin @ 2009-01-02 21:18 UTC (permalink / raw)
To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc
On Fri, Jan 02, 2009 at 04:14:12PM -0500, John David Anglin wrote:
> > On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote:
> > > OS HPMC handler length needs to be in bytes, not in words.
> > > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be>
> > >
> > > - length = os_hpmc_end - os_hpmc;
> > > + length = (os_hpmc_end - os_hpmc) * sizeof(u32);
> > > ivap[7] = length;
> >
> > Ick. I wonder if I screwed that up when I fixed the function descriptor
> > breakage...
>
> The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are
> pointing to function descriptors.
>
I think we've been over this before ;-)
commit c3d4ed4e3e5aa8d9e6b4b795f004a7028ce780e9
Author: Kyle McMartin <kyle@parisc-linux.org>
Date: Mon Jun 4 02:26:52 2007 -0400
[PARISC] Fix kernel panic in check_ivt
check_ivt had some seriously broken code wrt function pointers on
parisc64. Instead of referencing the hpmc code via a function
pointer,
export symbols and reference it as a const array.
Thanks to jda for pointing out the broken 64-bit func ptr handling.
Signed-off-by: Kyle McMartin <kyle@parisc-linux.org>
cheers, Kyle
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-02 21:18 ` Kyle McMartin
@ 2009-01-02 21:36 ` John David Anglin
2009-01-03 0:09 ` John David Anglin
1 sibling, 0 replies; 20+ messages in thread
From: John David Anglin @ 2009-01-02 21:36 UTC (permalink / raw)
To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc
> On Fri, Jan 02, 2009 at 04:14:12PM -0500, John David Anglin wrote:
> > > On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote:
> > > > OS HPMC handler length needs to be in bytes, not in words.
> > > > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be>
> > > >
> > > > - length = os_hpmc_end - os_hpmc;
> > > > + length = (os_hpmc_end - os_hpmc) * sizeof(u32);
> > > > ivap[7] = length;
> > >
> > > Ick. I wonder if I screwed that up when I fixed the function descriptor
> > > breakage...
> >
> > The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are
> > pointing to function descriptors.
> >
>
> I think we've been over this before ;-)
Looking at traps.o for a 64-bit build, I see the references are:
0000000000000048 R_PARISC_DLTIND21L os_hpmc_end
0000000000000050 R_PARISC_DLTIND14R os_hpmc_end
0000000000000058 R_PARISC_DLTIND21L os_hpmc
000000000000005c R_PARISC_DLTIND14R os_hpmc
Wonder what the DLT is pointing to.
Dave
--
J. David Anglin dave.anglin@nrc-cnrc.gc.ca
National Research Council of Canada (613) 990-0752 (FAX: 952-6602)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-02 21:18 ` Kyle McMartin
2009-01-02 21:36 ` John David Anglin
@ 2009-01-03 0:09 ` John David Anglin
2009-01-03 0:44 ` Kyle McMartin
1 sibling, 1 reply; 20+ messages in thread
From: John David Anglin @ 2009-01-03 0:09 UTC (permalink / raw)
To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc
> I think we've been over this before ;-)
One safe way to compute this difference is using a difference
of local labels in .data.
Dave
--
J. David Anglin dave.anglin@nrc-cnrc.gc.ca
National Research Council of Canada (613) 990-0752 (FAX: 952-6602)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-03 0:09 ` John David Anglin
@ 2009-01-03 0:44 ` Kyle McMartin
2009-01-03 1:24 ` John David Anglin
0 siblings, 1 reply; 20+ messages in thread
From: Kyle McMartin @ 2009-01-03 0:44 UTC (permalink / raw)
To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc
On Fri, Jan 02, 2009 at 07:09:57PM -0500, John David Anglin wrote:
> > I think we've been over this before ;-)
>
> One safe way to compute this difference is using a difference
> of local labels in .data.
>
Something like this? Sorry, I've almost completely forgotten gas syntax
and all that jazz...
I seem to recall writing a patch to dump the first 16 insns of os_hpmc
from the check_ivt call, and found that it was working sensibly on
64-bit.
regards, Kyle
diff --git a/arch/parisc/kernel/hpmc.S b/arch/parisc/kernel/hpmc.S
index 2cbf13b..02e41da 100644
--- a/arch/parisc/kernel/hpmc.S
+++ b/arch/parisc/kernel/hpmc.S
@@ -295,5 +295,9 @@ os_hpmc_6:
b .
nop
ENDPROC(os_hpmc)
-ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */
+ENTRY(os_hpmc_end)
nop
+.data
+ .export os_hpmc_size
+os_hpmc_size:
+ .word os_hpmc_end-os_hpmc
diff --git a/arch/parisc/kernel/traps.c b/arch/parisc/kernel/traps.c
index 4c771cd..5cb66ac 100644
--- a/arch/parisc/kernel/traps.c
+++ b/arch/parisc/kernel/traps.c
@@ -821,8 +821,8 @@ void handle_interruption(int code, struct pt_regs *regs)
int __init check_ivt(void *iva)
{
- extern const u32 os_hpmc[];
- extern const u32 os_hpmc_end[];
+ extern u32 os_hpmc_size;
+ extern u32 os_hpmc[];
int i;
u32 check = 0;
@@ -839,8 +839,7 @@ int __init check_ivt(void *iva)
*ivap++ = 0;
/* Compute Checksum for HPMC handler */
^ permalink raw reply related [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-03 0:44 ` Kyle McMartin
@ 2009-01-03 1:24 ` John David Anglin
2009-01-03 1:27 ` Kyle McMartin
0 siblings, 1 reply; 20+ messages in thread
From: John David Anglin @ 2009-01-03 1:24 UTC (permalink / raw)
To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc
> Something like this? Sorry, I've almost completely forgotten gas syntax
> and all that jazz...
Yes. The gas syntax looks correct except you probably should add a
.align 4. I would use local labels (that's what's what is done for
debug and exception info). Just place the local labels adjacent to
os_hpmc and os_hpmc_end. This ensures the diff is absolute.
> -ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */
> +ENTRY(os_hpmc_end)
Probably, you don't need to export os_hpmc_end.
> - length = os_hpmc_end - os_hpmc;
> + length = os_hpmc_size * sizeof(u32);
I believe the size will be correct without multiplying by sizeof(u32).
Dave
--
J. David Anglin dave.anglin@nrc-cnrc.gc.ca
National Research Council of Canada (613) 990-0752 (FAX: 952-6602)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600
2009-01-03 1:24 ` John David Anglin
@ 2009-01-03 1:27 ` Kyle McMartin
0 siblings, 0 replies; 20+ messages in thread
From: Kyle McMartin @ 2009-01-03 1:27 UTC (permalink / raw)
To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc
On Fri, Jan 02, 2009 at 08:24:50PM -0500, John David Anglin wrote:
> > Something like this? Sorry, I've almost completely forgotten gas syntax
> > and all that jazz...
>
> Yes. The gas syntax looks correct except you probably should add a
> .align 4. I would use local labels (that's what's what is done for
> debug and exception info). Just place the local labels adjacent to
> os_hpmc and os_hpmc_end. This ensures the diff is absolute.
>
Ahh, excellent. Thanks.
> > -ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */
> > +ENTRY(os_hpmc_end)
>
> Probably, you don't need to export os_hpmc_end.
>
Indeed.
> > - length = os_hpmc_end - os_hpmc;
> > + length = os_hpmc_size * sizeof(u32);
>
> I believe the size will be correct without multiplying by sizeof(u32).
>
Ah, indeed. Good point.
cheers, Kyle
^ permalink raw reply [flat|nested] 20+ messages in thread
end of thread, other threads:[~2009-01-03 1:27 UTC | newest]
Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin
2008-12-20 22:16 ` Kyle McMartin
2008-12-20 22:21 ` Kyle McMartin
2008-12-21 8:12 ` Grant Grundler
2008-12-22 21:40 ` Guy Martin
2008-12-22 21:47 ` Kyle McMartin
2008-12-22 22:49 ` Guy Martin
2008-12-25 8:06 ` Grant Grundler
2008-12-25 15:57 ` Guy Martin
2008-12-25 21:36 ` Grant Grundler
2009-01-02 20:46 ` Guy Martin
2009-01-02 21:01 ` Kyle McMartin
2009-01-02 21:14 ` John David Anglin
2009-01-02 21:18 ` Kyle McMartin
2009-01-02 21:36 ` John David Anglin
2009-01-03 0:09 ` John David Anglin
2009-01-03 0:44 ` Kyle McMartin
2009-01-03 1:24 ` John David Anglin
2009-01-03 1:27 ` Kyle McMartin
2008-12-21 8:10 ` Grant Grundler
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.