* HPMC bus timeout on C3600
@ 2008-12-20 12:39 Guy Martin
2008-12-20 22:16 ` Kyle McMartin
2008-12-21 8:10 ` Grant Grundler
0 siblings, 2 replies; 20+ messages in thread
From: Guy Martin @ 2008-12-20 12:39 UTC (permalink / raw)
To: linux-parisc
[-- Attachment #1: Type: text/plain, Size: 1189 bytes --]
Hi team,
I'm recently running into HPMC bus timeout problem. While this never
caused any problem before, I've noticed a few strange things. I've
attached the output I get from "pim hpmc" in the PDC.
This BUS timeout occurs after one or two days of idling when I plug a
specific card reader in the USB PCI card I plugged in. I think this
card reader causes problems because it's being polled every 2-3 seconds.
Nevertheless, looking at the pdc output, it seems that the OS HPMC
handler is not kicking off according to the chassis code CBF2 and CBFC.
CBF2 : bas OS HPMC len
CBFC : OS HPMC br err
I've check GR02 and it points to inw() which makes sens. Nothing
interesting there.
The failing device appears to be the built-in NIC according the path
provided in the output (10/0/12/0) and the C3600 service manual (figure
5-2).
Even if this look HW pb, I've had a few reports recently about HPMC PCI
timeout issues and I have a few doubts.
Now my questions are :
- does this look like an HW or SW problems ?
- why isn't the HPMC handler kicked off ?
- could the HPMC handler recover this ?
- what debug can I enable to get more info about this if applicable ?
Cheers,
Guy
[-- Attachment #2: c3600-hpmc --]
[-- Type: application/octet-stream, Size: 25894 bytes --]
Service Menu: Enter command > pim hpmc
PROCESSOR PIM INFORMATION
----------------- Processor 0 HPMC Information ------------------
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc
General Registers 0 - 31
00-03 0000000000000000 00000000104dc150 00000000101176f0 000000008f31e360
04-07 0000000000000000 000000000000000f 000000008f31e000 0000000000000000
08-11 0000000000000000 0000000010582950 00000000104dc150 000000001057ede8
12-15 00000000f0000000 0000000000000081 000000001040c0e4 00000000f0400004
16-19 00000000f00008c4 00000000f000017c 00000000f0000174 0000000000020000
20-23 00000000000000d4 00000000030a9fcb 000000001027a7b8 0000000000000000
24-27 0000000000000001 0000000000001040 000000008f804580 00000000104dc150
28-31 00000000104fb778 0000000002584fd6 000000008f8e0380 00000000101176f0
<Press any key to continue (q to quit)>
Control Registers 0 - 31
00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000
04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000
08-11 0000000000007ad2 0000000000000000 00000000000000c0 000000000000003f
12-15 0000000000000000 0000000000000000 000000000010f000 00000000ff800000
16-19 000055958963ba84 0000000000000000 000000001027a7c4 000000000e79009c
20-23 00000000a627fffb 0000000080421040 000000ff0006f90e 0000000080000000
24-27 000000000052a000 000000004250e000 0000000000044021 00000000f0412000
28-31 0000000055555555 0000000055555555 000000008f8e0000 0000000011111111
Space Registers 0 - 7
00-03 00000000 00000000 00000000 00003d69
04-07 00000000 00000000 00000000 00000000
IIA Space = 0x0000000000000000
IIA Offset = 0x000000001027a7c8
Check Type = 0x20000000
CPU State = 0x9e000004
Cache Check = 0x00000000
TLB Check = 0x00000000
Bus Check = 0x0030103b
Assists Check = 0x00000000
Assist State = 0x00000000
Path Info = 0x00000000
System Responder Address = 0x000000fffee01040
System Requestor Address = 0xfffffffffffa0000
Floating-Point Registers 0 - 31
00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000
04-07 3fd3e672c6670f8c 41d215ed21800000 0000000654000000 5ea496f000000000
08-11 00000003b26de560 3ff0000000000000 3ff0000000000000 0000000000000000
12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555
16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555
20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000
24-27 0000888630000000 0000ce9a00000000 000000006e2df49c 4024cb5ecf0a9480
28-31 3ff0000000000000 0000000000000000 4024cb5ecf0a9480 0000000000000000
'9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes:
Check Summary = 0xcb81041008000000
Available Memory = 0x0000000080000000
CPU Diagnose Register 2 = 0x0300000000000004
CPU Status Register 0 = 0x2420c20000000000
CPU Status Register 1 = 0x8002000000000000
SADD LOG = 0xd3081bfb486511b8
Read Short LOG = 0xc1af00fffee01040
ERROR_STATUS = 0x0000000000100010
MEM_ADDR = 0x000001ff3fffffff
MEM_SYND = 0x0000000000000000
MEM_ADDR_CORR = 0x000001ff3fffffff
MEM_SYND_CORR = 0x0000000000000000
RUN_DATA_HIGH = 0xc1bff0fffed08040
RUN_DATA_LOW = 0xc1bff0fffed08040
RUN_CTRL = 0x0000021c00001418
RUN_ADDR = 0xc1bff0fffed08040
System Responder Path = 0x00ffffff0a000c00
HPMC PIM Analysis Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes:
A Data I/O Fetch Timeout occurred while CPU 0 was
requesting information from a device at the path 10/0/12/0 (built-in PCI device).
Memory/IO Controller Error Analysis Information:
The Memory/IO Controller only observed the Broadcast Error. It did not log
any additional information about the HPMC.
Memory Error Log Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes:
No memory errors logged
I/O Module Error Log Information:
Timestamp =
Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47)
'9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes:
Rope Word1 Word2 Word3
------ ------------ ------------
0 0x00000000 0x0e0cc2a9 0x00000000fed30048
1 0x00000000 0x1e0cc009 0x00000000fed32048
2 ---------- 0x2e0cc009 ------------------
3 ---------- 0x3e0cc009 ------------------
4 0x00000000 0x4e0cc009 0x00000000fed38048
5 ---------- 0x5e0cc009 ------------------
6 0x00000000 0x6e0cc009 0x00000000fed3c048
7 ---------- 0x7e0cc009 ------------------
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: HPMC bus timeout on C3600 2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin @ 2008-12-20 22:16 ` Kyle McMartin 2008-12-20 22:21 ` Kyle McMartin 2008-12-21 8:10 ` Grant Grundler 1 sibling, 1 reply; 20+ messages in thread From: Kyle McMartin @ 2008-12-20 22:16 UTC (permalink / raw) To: Guy Martin; +Cc: linux-parisc On Sat, Dec 20, 2008 at 01:39:15PM +0100, Guy Martin wrote: > I'm recently running into HPMC bus timeout problem. While this never > caused any problem before, I've noticed a few strange things. I've > attached the output I get from "pim hpmc" in the PDC. > > This BUS timeout occurs after one or two days of idling when I plug a > specific card reader in the USB PCI card I plugged in. I think this > card reader causes problems because it's being polled every 2-3 seconds. > > Nevertheless, looking at the pdc output, it seems that the OS HPMC > handler is not kicking off according to the chassis code CBF2 and CBFC. > CBF2 : bas OS HPMC len > CBFC : OS HPMC br err > > I've check GR02 and it points to inw() which makes sens. Nothing > interesting there. > > The failing device appears to be the built-in NIC according the path > provided in the output (10/0/12/0) and the C3600 service manual (figure > 5-2). > > Even if this look HW pb, I've had a few reports recently about HPMC PCI > timeout issues and I have a few doubts. > > Now my questions are : > - does this look like an HW or SW problems ? > - why isn't the HPMC handler kicked off ? > - could the HPMC handler recover this ? > - what debug can I enable to get more info about this if applicable ? > "bus timeout" usually means we tried to read an address that doesn't respond. that is, nothing on the bus accepted the transaction for it, so it timed out and HPMC'd the box. what you really need is the IIR, and the address it tried to access (both the kernel vaddr which will be in the register, and the "system requester address" from the hpmc dump which will be the physical address mapped. not sure why the hpmc handler is getting skipped, that's a little weird. you can try hacking elroy to set softfail mode on that bus, which will result in a timeout on the pci bus to return -1 (like what x86 and most other architectures do) rather than hang the box, but it really likely means a driver bug. ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-20 22:16 ` Kyle McMartin @ 2008-12-20 22:21 ` Kyle McMartin 2008-12-21 8:12 ` Grant Grundler 0 siblings, 1 reply; 20+ messages in thread From: Kyle McMartin @ 2008-12-20 22:21 UTC (permalink / raw) To: Guy Martin; +Cc: linux-parisc On Sat, Dec 20, 2008 at 05:16:53PM -0500, Kyle McMartin wrote: > you can try hacking elroy to set softfail mode on that bus, which will > result in a timeout on the pci bus to return -1 (like what x86 and most > other architectures do) rather than hang the box, but it really likely > means a driver bug. untested and all that jazz. diff --git a/drivers/parisc/lba_pci.c b/drivers/parisc/lba_pci.c index a28c894..a34f759 100644 --- a/drivers/parisc/lba_pci.c +++ b/drivers/parisc/lba_pci.c @@ -1341,7 +1341,7 @@ lba_hw_init(struct lba_device *d) /* Set HF mode as the default (vs. -1 mode). */ stat = READ_REG32(d->hba.base_addr + LBA_STAT_CTL); - WRITE_REG32(stat | HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL); + WRITE_REG32(stat & ~HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL); /* ** Writing a zero to STAT_CTL.rf (bit 0) will clear reset signal ^ permalink raw reply related [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-20 22:21 ` Kyle McMartin @ 2008-12-21 8:12 ` Grant Grundler 2008-12-22 21:40 ` Guy Martin 0 siblings, 1 reply; 20+ messages in thread From: Grant Grundler @ 2008-12-21 8:12 UTC (permalink / raw) To: Kyle McMartin; +Cc: Guy Martin, linux-parisc On Sat, Dec 20, 2008 at 05:21:00PM -0500, Kyle McMartin wrote: > On Sat, Dec 20, 2008 at 05:16:53PM -0500, Kyle McMartin wrote: > > you can try hacking elroy to set softfail mode on that bus, which will > > result in a timeout on the pci bus to return -1 (like what x86 and most > > other architectures do) rather than hang the box, but it really likely > > means a driver bug. > > untested and all that jazz. > > diff --git a/drivers/parisc/lba_pci.c b/drivers/parisc/lba_pci.c > index a28c894..a34f759 100644 > --- a/drivers/parisc/lba_pci.c > +++ b/drivers/parisc/lba_pci.c > @@ -1341,7 +1341,7 @@ lba_hw_init(struct lba_device *d) > > /* Set HF mode as the default (vs. -1 mode). */ > stat = READ_REG32(d->hba.base_addr + LBA_STAT_CTL); > - WRITE_REG32(stat | HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL); > + WRITE_REG32(stat & ~HF_ENABLE, d->hba.base_addr + LBA_STAT_CTL); I'm not sure how helpful this will be. It will be unpredictable how the tulip driver will handle getting -1's back from inw(). thanks, grant > > /* > ** Writing a zero to STAT_CTL.rf (bit 0) will clear reset signal > -- > To unsubscribe from this list: send the line "unsubscribe linux-parisc" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-21 8:12 ` Grant Grundler @ 2008-12-22 21:40 ` Guy Martin 2008-12-22 21:47 ` Kyle McMartin 0 siblings, 1 reply; 20+ messages in thread From: Guy Martin @ 2008-12-22 21:40 UTC (permalink / raw) To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc [-- Attachment #1: Type: text/plain, Size: 1928 bytes --] Kyle, Grant, Thanks for your help but being unable to reproduce the issue at will makes it very difficult to make any progress. The fault didn't occurred those past two days. So I decided to go another way. I've plugged in one of my well working sata card and did a lot of IO using "dd if=/dev/zero of=/dev/sdc". I'm not sure if this would trigger the same problem but it does trigger an HPMC Bus timeout too after a few hundreds megs. This time, the PCI slot pointed by the output is the sata card itself in slot 4. Regarding GR02, it seems to point somewhere in user space. Regarding Kyle's patch, it didn't changed anything but since the fault isn't at inw(), I'm not sure it still even makes sens to apply it. I tried to find out why the HPMC os handler wasn't triggered. As documented in pdc20-v1.1-Ch3-pdce.pdf, the interruption vector table is setup correctly. By this I mean that CR 14 points correctly to the beginning of the fault_vector_20 (physical addr), the checksum is ok and the length seems about right. I used a quickly coded module to check all of this at run time. The only thing that doesn't seem to be done is calling PDC_INSTR but I still need to investigate that. As far as I understand it's should provide instr to jump to the os_hpmc handler. Those instr are already harcoded in the fault vector so it's shouldn't be needed, unless calling PDC_INSTR actually does tell the PDC that os_hpmc has been installed. So again, I'm looking for help on going further with this issue. Unfortunately, I don't have a lot of knowledge about PCI and the HP side of the elroy chip. I wanted to look at elroy's documentation but it seems that a large part was stripped down. Any pointer to usefull doc is more than welcome ! I've attached two HPMC dumps generated using the same command. Hopefull they'll provide more details. This is on an C3600 running latest parisc-2.6 git tree. Regards, Guy [-- Attachment #2: c3600-sata-pci-hpmc2 --] [-- Type: application/octet-stream, Size: 5212 bytes --] PROCESSOR PIM INFORMATION ----------------- Processor 0 HPMC Information ------------------ Timestamp = Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13) HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc General Registers 0 - 31 00-03 0000000000000000 0000000000650000 00000000006500f4 000000008cc2c000 04-07 000000008dfb8828 000000008dfb8834 0000000000000001 000000008cc2c000 08-11 000000000000000f 000000008dfb8834 000000008dfb8844 000000000000001e 12-15 000000008dfb8800 000000008cc2de28 0000000000000005 00000000000000d0 16-19 0000000000002002 00000000f000017c 00000000f0000174 0000000000400000 20-23 0000000000000000 0000000000682000 0000000000679b9c 000000000000002f 24-27 000000006b719d0a 000000008cc2ddd0 000000008cc2c000 00000000104c78b0 28-31 0000000000679d3c 000000000001377c 000000008cc1c240 00000000006500f4 <Press any key to continue (q to quit)> Control Registers 0 - 31 00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000 04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000 08-11 0000000000000e00 0000000000000000 00000000000000c0 000000000000003f 12-15 0000000000000000 0000000000000000 000000000010e000 00000000ffc00000 16-19 00000037fe47b62e 0000000000000000 0000000000679bd8 000000004abc0090 20-23 00000000a627ffee 0000000001e82048 000000000004000e 0000000080000000 24-27 000000000051a000 000000007de58000 0000000000044021 00000000f0412000 28-31 0000000055555555 0000000055555555 000000008cc1c000 0000000011111111 Space Registers 0 - 7 00-03 00000000 00000000 00000000 00000700 04-07 00000000 00000000 00000000 00000000 IIA Space = 0x0000000000000000 IIA Offset = 0x0000000000679bdc Check Type = 0x20000000 CPU State = 0x9e000004 Cache Check = 0x00000000 TLB Check = 0x00000000 Bus Check = 0x0030103b Assists Check = 0x00000000 Assist State = 0x00000000 Path Info = 0x00000000 System Responder Address = 0x000000fffb807048 System Requestor Address = 0xfffffffffffa0000 Floating-Point Registers 0 - 31 00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000 04-07 3febb8de8e377300 3febb8de8e377300 0000000100000000 0000000000000000 08-11 0000000000000000 000000028f812bc8 8f812bc000000002 1056f0b000000003 12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555 16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555 20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000 24-27 0000000000000000 0000ff9700000000 0000000000000000 10282858102b7368 28-31 ffffffff000137f2 104e0da4101871a0 0000000100000228 8f820200101165f4 '9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes: Check Summary = 0xcb81041008000000 Available Memory = 0x0000000080000000 CPU Diagnose Register 2 = 0x0300000000000004 CPU Status Register 0 = 0x2420c20000000000 CPU Status Register 1 = 0x8002000000000000 SADD LOG = 0xc10f00fffb807048 Read Short LOG = 0xc1af00fffb807048 ERROR_STATUS = 0x0000000000100010 MEM_ADDR = 0x000001ff3fffffff MEM_SYND = 0x0000000000000000 MEM_ADDR_CORR = 0x000001ff3fffffff MEM_SYND_CORR = 0x0000000000000000 RUN_DATA_HIGH = 0xc1bff0fffed08040 RUN_DATA_LOW = 0xc1bff0fffed08040 RUN_CTRL = 0x0000021c00001418 RUN_ADDR = 0xc1bff0fffed08040 System Responder Path = 0x00ffffff0a010400 HPMC PIM Analysis Information: Timestamp = Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13) '9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes: A Data I/O Fetch Timeout occurred while CPU 0 was requesting information from a device at the path 10/1/4/0 (PCI slot 4). Memory/IO Controller Error Analysis Information: The Memory/IO Controller only observed the Broadcast Error. It did not log any additional information about the HPMC. Memory Error Log Information: Timestamp = Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13) '9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes: No memory errors logged I/O Module Error Log Information: Timestamp = Mon Dec 22 21:00:13 GMT 2008 (20:08:12:22:21:00:13) '9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes: Rope Word1 Word2 Word3 ------ ------------ ------------ 0 0x00000000 0x0e0cc009 0x00000000fed30048 1 0x00000000 0x1e0cc289 0x00000000fed32048 2 ---------- 0x2e0cc009 ------------------ 3 ---------- 0x3e0cc009 ------------------ 4 0x00000000 0x4e0cc009 0x00000000fed38048 5 ---------- 0x5e0cc009 ------------------ 6 0x00000000 0x6e0cc009 0x00000000fed3c048 7 ---------- 0x7e0cc009 ------------------ [-- Attachment #3: c3600-sata-pci-hpmc --] [-- Type: application/octet-stream, Size: 5211 bytes --] PROCESSOR PIM INFORMATION ----------------- Processor 0 HPMC Information ------------------ Timestamp = Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48) HPMC Chassis Codes = 2cbf0 2500b 2cbf2 2cbfc General Registers 0 - 31 00-03 0000000000000000 0000000000650000 00000000006500f4 000000002f084000 04-07 000000008cc1e828 000000008cc1e834 0000000000000001 000000002f084000 08-11 000000000000000f 000000008cc1e834 000000008cc1e844 000000000000001e 12-15 000000008cc1e800 000000002f085e28 0000000000000005 00000000000000d0 16-19 0000000000002002 00000000f000017c 00000000f0000174 0000000000400000 20-23 0000000000000000 0000000000682000 0000000000679b9c 000000000000004c 24-27 000000000a0ee5b0 000000002f085dd0 000000002f084000 00000000104c78b0 28-31 0000000000679d3c 000000000021f42e 000000002f284240 00000000006500f4 <Press any key to continue (q to quit)> Control Registers 0 - 31 00-03 0000000000000000 0000000000000000 0000000000000000 0000000000000000 04-07 0000000000000000 0000000000000000 0000000000000000 0000000000000000 08-11 0000000000000f9a 0000000000000000 00000000000000c0 000000000000003f 12-15 0000000000000000 0000000000000000 000000000010e000 00000000ffc00000 16-19 000003a2c566501d 0000000000000000 0000000000679bd8 000000004abc0090 20-23 00000000a627ffee 0000000001e82048 000000000006000e 0000000080000000 24-27 000000000051a000 000000007df68000 0000000000044021 00000000f0412000 28-31 0000000055555555 0000000055555555 000000002f284000 0000000011111111 Space Registers 0 - 7 00-03 00000000 000007cd 00000000 000007cd 04-07 00000000 00000000 00000000 00000000 IIA Space = 0x0000000000000000 IIA Offset = 0x0000000000679bdc Check Type = 0x20000000 CPU State = 0x9e000004 Cache Check = 0x00000000 TLB Check = 0x00000000 Bus Check = 0x0030103b Assists Check = 0x00000000 Assist State = 0x00000000 Path Info = 0x00000000 System Responder Address = 0x000000fffb807048 System Requestor Address = 0xfffffffffffa0000 Floating-Point Registers 0 - 31 00-03 0000001f00000000 0000000000000000 0000000000000000 0000000000000000 04-07 3e5058af0f4dd252 bf4e68a0d349be90 0000000000000000 bf4e6eea80000000 08-11 0000000538400000 000000028f812bc8 8f812bc000000002 1056f0b000000003 12-15 5555555555555555 5555555555555555 5555555555555555 5555555555555555 16-19 5555555555555555 5555555555555555 5555555555555555 5555555555555555 20-23 5555555555555555 5555555555555555 0000000000000000 0000000000000000 24-27 0000000000000000 0000ff97b22a09b6 0000000000000000 3ff0000000000000 28-31 00100000000137f2 104e0da4101871a0 0000000100000228 8f820200101165f4 '9000/785 B,C,J Workstation Unarchitected (per-CPU)', rev 1, 140 bytes: Check Summary = 0xcb81041008000000 Available Memory = 0x0000000080000000 CPU Diagnose Register 2 = 0x0300000000000004 CPU Status Register 0 = 0x2420c20000000000 CPU Status Register 1 = 0x8002000000000000 SADD LOG = 0xc10f00fffb807048 Read Short LOG = 0xc1af00fffb807048 ERROR_STATUS = 0x0000000000100010 MEM_ADDR = 0x000001ff3fffffff MEM_SYND = 0x0000000000000000 MEM_ADDR_CORR = 0x000001ff3fffffff MEM_SYND_CORR = 0x0000000000000000 RUN_DATA_HIGH = 0xc1bff0fffed08040 RUN_DATA_LOW = 0xc1bff0fffed08040 RUN_CTRL = 0x0000021c00001418 RUN_ADDR = 0xc1bff0fffed08040 System Responder Path = 0x00ffffff0a010400 HPMC PIM Analysis Information: Timestamp = Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48) '9000/785 B,C,J Workstation HPMC PIM Analysis (per-CPU)', rev 0, 1304 bytes: A Data I/O Fetch Timeout occurred while CPU 0 was requesting information from a device at the path 10/1/4/0 (PCI slot 4). Memory/IO Controller Error Analysis Information: The Memory/IO Controller only observed the Broadcast Error. It did not log any additional information about the HPMC. Memory Error Log Information: Timestamp = Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48) '9000/785 B,C,J Workstation Memory Error Log', rev 0, 64 bytes: No memory errors logged I/O Module Error Log Information: Timestamp = Mon Dec 22 17:51:48 GMT 2008 (20:08:12:22:17:51:48) '9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes: Rope Word1 Word2 Word3 ------ ------------ ------------ 0 0x00000000 0x0e0cc009 0x00000000fed30048 1 0x0001c800 0x1e0cc009 0x00000000fb807000 2 ---------- 0x2e0cc009 ------------------ 3 ---------- 0x3e0cc009 ------------------ 4 0x00000000 0x4e0cc009 0x00000000fed38048 5 ---------- 0x5e0cc009 ------------------ 6 0x00000000 0x6e0cc009 0x00000000fed3c048 7 ---------- 0x7e0cc009 ------------------ ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-22 21:40 ` Guy Martin @ 2008-12-22 21:47 ` Kyle McMartin 2008-12-22 22:49 ` Guy Martin 0 siblings, 1 reply; 20+ messages in thread From: Kyle McMartin @ 2008-12-22 21:47 UTC (permalink / raw) To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc On Mon, Dec 22, 2008 at 10:40:23PM +0100, Guy Martin wrote: > > Kyle, Grant, > > Thanks for your help but being unable to reproduce the issue at will > makes it very difficult to make any progress. The fault didn't occurred > those past two days. > > So I decided to go another way. I've plugged in one of my well working > sata card and did a lot of IO using "dd if=/dev/zero of=/dev/sdc". > > I'm not sure if this would trigger the same problem but it does trigger > an HPMC Bus timeout too after a few hundreds megs. This time, the PCI > slot pointed by the output is the sata card itself in slot 4. > Regarding GR02, it seems to point somewhere in user space. > > Regarding Kyle's patch, it didn't changed anything but since the fault > isn't at inw(), I'm not sure it still even makes sens to apply it. > you won't know for sure, since readl/etc are inlines into the caller... ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-22 21:47 ` Kyle McMartin @ 2008-12-22 22:49 ` Guy Martin 2008-12-25 8:06 ` Grant Grundler 0 siblings, 1 reply; 20+ messages in thread From: Guy Martin @ 2008-12-22 22:49 UTC (permalink / raw) To: Kyle McMartin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc [-- Attachment #1: Type: text/plain, Size: 721 bytes --] On Mon, 22 Dec 2008 16:47:43 -0500 Kyle McMartin <kyle@infradead.org> wrote: > > Regarding Kyle's patch, it didn't changed anything but since the > > fault isn't at inw(), I'm not sure it still even makes sens to > > apply it. > > > > you won't know for sure, since readl/etc are inlines into the > caller... It seems that I forgot to include your patch when I switched to the git tree. I've been doing tests again with your patch and things are different : The command "dd if=/dev/zero of=/dev/sdc bs=1M" didn't trigger HPMC anymore after a reasonable amount of time. I then tried a NFS copy from a NFS share to the disk. The kernel did panic with the attached output. Hopefully this will be more useful. Guy [-- Attachment #2: c3600-sata-pci-panic --] [-- Type: application/octet-stream, Size: 5875 bytes --] [ 1577.999842] ata1.00: exception Emask 0x52 SAct 0x0 SErr 0xffffffff action 0xe [ 1577.999842] ata1: SError: { RecovData RecovComm UnrecovData Persist Proto HostInt PHYRdyChg PHYInt CommWake 10B8B Dispar BadCRC Handshk LinkSeq TrStaTrns UnrecFIS DevExch } [ 1577.999842] ata1.00: cmd 35/00:00:bf:1f:2f/00:04:00:00:00/e0 tag 0 dma 524288 out [ 1577.999842] res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x72 (host bus error) [ 1577.999842] ata1.00: status: { DRDY } [ 1577.999842] ata1: hard resetting link [ 1578.719842] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF) [ 1583.719841] ata1: hard resetting link [ 1584.039841] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF) [ 1589.039841] ata1: hard resetting link [ 1589.359841] ata1: SATA link down (SStatus FFFFFFFF SControl FFFFFFFF) [ 1589.359841] ata1.00: disabled [ 1589.359841] sd 2:0:0:0: [sdc] Result: hostbyte=0x00 driverbyte=0x08 [ 1589.359841] sd 2:0:0:0: [sdc] Sense Key : 0xb [current] [descriptor] [ 1589.359841] Descriptor sense data with sense descriptors (in hex): [ 1589.359841] 72 0b 00 00 00 00 00 0c 00 0a 80 00 00 00 00 00 [ 1589.359841] 00 00 00 00 [ 1589.359841] sd 2:0:0:0: [sdc] ASC=0x0 ASCQ=0x0 [ 1589.359841] end_request: I/O error, dev sdc, sector 3088319 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386032 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386033 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386034 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386035 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386036 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386037 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386038 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386039 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386040 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] Buffer I/O error on device sdc1, logical block 386041 [ 1589.359841] lost page write due to I/O error on sdc1 [ 1589.359841] sd 2:0:0:0: rejecting I/O to offline device [ 1589.359841] ata1: EH complete [ 1589.359841] ata1.00: detaching (SCSI 2:0:0:0) [ 1589.359841] sd 2:0:0:0: [sdc] Result: hostbyte=0x01 driverbyte=0x00 [ 1589.359841] end_request: I/O error, dev sdc, sector 3089343 [ 1589.403174] EXT3-fs error (device sdc1): read_block_bitmap: Cannot read block bitmap - block_group = 13, block_bitmap = 425984 [ 1589.409841] EXT3-fs error (device sdc1): read_block_bitmap: Cannot read block bitmap - block_group = 13, block_bitmap = 425984 [ 1589.409841] ------------[ cut here ]------------ [ 1589.466507] Badness at fs/buffer.c:1186 [ 1589.466507] [ 1589.466507] YZrvWESTHLNXBCVMcbcbcbcbOGFRQPDI [ 1589.466507] PSW: 00000000000001000000001100001111 Not tainted [ 1589.466507] r00-03 0004030f 1056f8b0 101d0ccc 8e908400 [ 1589.466507] r04-07 3cf97400 8f510afc 00000001 3cf961a0 [ 1589.466507] r08-11 8e908400 7226c7d4 8ded3f00 00000000 [ 1589.466507] r12-15 8eb813c0 00000001 492ecb30 492ecb30 [ 1589.466507] r16-19 7226c8c8 5195f3ec 00068000 dfd101d1 [ 1589.466507] r20-23 3cf6f000 0000001d 00040000 00000e8e [ 1589.466507] r24-27 00000000 00000e8f 8f510afc 104c78b0 [ 1589.466507] r28-31 00000000 00000006 7226cb00 fffffff7 [ 1589.466507] sr00-03 00000000 000015ef 00000000 00000035 [ 1589.466507] sr04-07 00000000 00000000 00000000 00000000 [ 1589.466507] [ 1589.466507] IASQ: 00000000 00000000 IAOQ: 1019cc88 1019cc8c [ 1589.466507] IIR: 03ffe01f ISR: 00000000 IOR: 00000000 [ 1589.466507] CPU: 0 CR30: 7226c000 CR31: 11111111 [ 1589.466507] ORIG_R28: 492ecb30 [ 1589.466507] IAOQ[0]: mark_buffer_dirty+0x84/0xa0 [ 1589.466507] IAOQ[1]: mark_buffer_dirty+0x88/0xa0 [ 1589.466507] RP(r2): ext3_commit_super+0x7c/0xb8 [ 1589.466507] Backtrace: [ 1589.466507] [<101d0ccc>] ext3_commit_super+0x7c/0xb8 [ 1589.466507] [<101d1cc0>] ext3_handle_error+0x7c/0xc4 [ 1589.466507] [<101d1df8>] ext3_error+0x64/0x78 [ 1589.466507] [<101c4168>] read_block_bitmap+0xb4/0x1d4 [ 1589.466507] [<101c5294>] ext3_new_blocks+0x3bc/0x6d8 [ 1589.466507] [<101c8ecc>] ext3_get_blocks_handle+0x380/0xa48 [ 1589.466507] [<101c97f0>] ext3_get_block+0x7c/0x10c [ 1589.466507] [<1019dc68>] __block_prepare_write+0x1e0/0x3d0 [ 1589.466507] [<1019df10>] block_write_begin+0x68/0x138 [ 1589.466507] [<101cb238>] ext3_write_begin+0x100/0x270 [ 1589.466507] [<10154b9c>] generic_file_buffered_write+0x12c/0x314 [ 1589.466507] [<101552d8>] __generic_file_aio_write_nolock+0x2c4/0x4a4 [ 1589.466507] [<10155538>] generic_file_aio_write+0x80/0x10c [ 1589.466507] [<101c6524>] ext3_file_write+0x34/0xd8 [ 1589.466507] [<1017b354>] do_sync_write+0xec/0x144 [ 1589.466507] [<1017bd28>] vfs_write+0x98/0x13c [ 1589.466507] [ 1589.473174] JBD: Detected IO errors while flushing file data on sdc1 [ 1589.473174] Aborting journal on device sdc1. [ 1589.526507] ext3_abort called. [ 1589.526507] EXT3-fs error (device sdc1): ext3_journal_start_sb: Detected aborted journal [ 1589.526507] Remounting filesystem read-only [ 1589.939841] JBD: Detected IO errors while flushing file data on sdc1 [ 1589.939841] __journal_remove_journal_head: freeing b_committed_data [ 1590.556507] sd 2:0:0:0: [sdc] Stopping disk [ 1590.563174] sd 2:0:0:0: [sdc] START_STOP FAILED [ 1590.563174] sd 2:0:0:0: [sdc] Result: hostbyte=0x04 driverbyte=0x00 ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-22 22:49 ` Guy Martin @ 2008-12-25 8:06 ` Grant Grundler 2008-12-25 15:57 ` Guy Martin 0 siblings, 1 reply; 20+ messages in thread From: Grant Grundler @ 2008-12-25 8:06 UTC (permalink / raw) To: Guy Martin; +Cc: Kyle McMartin, Grant Grundler, linux-parisc On Mon, Dec 22, 2008 at 11:49:48PM +0100, Guy Martin wrote: > The command "dd if=/dev/zero of=/dev/sdc bs=1M" didn't trigger HPMC > anymore after a reasonable amount of time. > > I then tried a NFS copy from a NFS share to the disk. The kernel did > panic with the attached output. [ 1589.466507] Badness at fs/buffer.c:1186 Technically, this isn't a panic. This is a "WARN_ON". I didn't see any panic's after that either. The file system warning us that the sata disk it was talking failed an IO. However, this is likely to be some other issue with the SATA controller. Can you post more details about the config? o "lspci -v" o hdparm -i /dev/sd<X> I don't expect I'll know what's wrong, but just want to document the bug. thanks, grant > Hopefully this will be more useful. > > Guy > ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-25 8:06 ` Grant Grundler @ 2008-12-25 15:57 ` Guy Martin 2008-12-25 21:36 ` Grant Grundler 0 siblings, 1 reply; 20+ messages in thread From: Guy Martin @ 2008-12-25 15:57 UTC (permalink / raw) To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc On Thu, 25 Dec 2008 01:06:06 -0700 Grant Grundler <grundler@parisc-linux.org> wrote: > [ 1589.466507] Badness at fs/buffer.c:1186 > > Technically, this isn't a panic. This is a "WARN_ON". > I didn't see any panic's after that either. Yes sorry. The system is still usable after the WARN_ON occurs. So no panic. Because of this, wouldn't it be a good idea to commit Kyle's patch ? To me it seems better to have a running system with failure messages being log rather than an ugly and barely understandable HPMC. > The file system warning us that the sata disk it was > talking failed an IO. > > However, this is likely to be some other issue with the SATA > controller. Can you post more details about the config? > o "lspci -v" > o hdparm -i /dev/sd<X> Here it is : 01:04.0 Mass storage controller: Silicon Image, Inc. SiI 3112 [SATALink/SATARaid] Serial ATA Controller (rev 02) Subsystem: Silicon Image, Inc. SiI 3112 SATALink Controller Flags: bus master, 66MHz, medium devsel, latency 240, IRQ 21 I/O ports at 12400 [size=8] I/O ports at 12300 [size=4] I/O ports at 12200 [size=8] I/O ports at 12100 [size=4] I/O ports at 12000 [size=16] Memory at fb807000 (32-bit, non-prefetchable) [size=512] Expansion ROM at fb880000 [disabled] [size=512K] Capabilities: [60] Power Management version 2 Kernel modules: sata_sil /dev/sdc: Model=WDC WD5000AAKS-00TMA0 , FwRev=12.01C01, SerialNo= WD-WMAPW1390228 Config={ HardSect NotMFM HdSw>15uSec SpinMotCtl Fixed DTR>5Mbs FmtGapReq } RawCHS=16383/16/63, TrkSize=0, SectSize=0, ECCbytes=50 BuffType=unknown, BuffSize=16384kB, MaxMultSect=16, MultSect=?0? CurCHS=16383/16/63, CurSects=16514064, LBA=yes, LBAsects=976773168 IORDY=on/off, tPIO={min:120,w/IORDY:120}, tDMA={min:120,rec:120} PIO modes: pio0 pio3 pio4 DMA modes: mdma0 mdma1 mdma2 UDMA modes: udma0 udma1 udma2 udma3 udma4 *udma5 udma6 AdvancedPM=no WriteCache=disabled Drive conforms to: Unspecified: ATA/ATAPI-1,2,3,4,5,6,7 * signifies the current active mode I'd like to add that I've been using the exact same card and hard drive on one of my x86 box for month without any issue. Also I've reproduce the problem again and this time I've had this message right before the "end_request" line : [ 163.039983] timer_interrupt(CPU 0): delayed! cycles EB5EE1E6 rem 36AE6 next/now 706F7328/5BCE550E Anything else I can do/provide to troubleshoot this ? Cheers, Guy ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-25 15:57 ` Guy Martin @ 2008-12-25 21:36 ` Grant Grundler 2009-01-02 20:46 ` Guy Martin 0 siblings, 1 reply; 20+ messages in thread From: Grant Grundler @ 2008-12-25 21:36 UTC (permalink / raw) To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc On Thu, Dec 25, 2008 at 04:57:50PM +0100, Guy Martin wrote: > On Thu, 25 Dec 2008 01:06:06 -0700 > Grant Grundler <grundler@parisc-linux.org> wrote: > > > [ 1589.466507] Badness at fs/buffer.c:1186 > > > > Technically, this isn't a panic. This is a "WARN_ON". > > I didn't see any panic's after that either. > > Yes sorry. The system is still usable after the WARN_ON occurs. So no > panic. > > Because of this, wouldn't it be a good idea to commit Kyle's patch ? I don't think kyle intended to commit that patch. He just wanted to collect more info. HPMC is 99% (for parisc-linux at least) of the time a driver bug. So having the system crash on IO errors (including DMA map/unmap bugs) is good for getting bugs reported and the state of the IOMMU and PCI Host controller to debug the problem. On the other hand, users often don't care about those details even if they see (like you did) the device isn't working right. They can more easily track down which device/drivers are having problems and remove them from the config. > To me it seems better to have a running system with failure messages > being log rather than an ugly and barely understandable HPMC. While I agree with your characterization of HPMCs, I don't want to trade HPMC dumps for kernel logs. HPMC provides info we can't otherwise get. HPMCs might be more useful if the symbol name of every kernel address in the dump were printed. Given the System.map, it should be possible to do something like: hpmc_symbols System.map < HPMC_dump.txt > HPMC_symbols.txt The program "a.c" already exists to do a symbol lookup given an kernel address and System.map file: http://cvs.parisc-linux.org/build-tools/ > > The file system warning us that the sata disk it was > > talking failed an IO. > > > > However, this is likely to be some other issue with the SATA > > controller. Can you post more details about the config? > > o "lspci -v" > > o hdparm -i /dev/sd<X> > > Here it is : > > 01:04.0 Mass storage controller: Silicon Image, Inc. SiI 3112 [SATALink/SATARaid] Serial ATA Controller (rev 02) Ok...so this is the sata_sil driver. > Subsystem: Silicon Image, Inc. SiI 3112 SATALink Controller > Flags: bus master, 66MHz, medium devsel, latency 240, IRQ 21 > I/O ports at 12400 [size=8] > I/O ports at 12300 [size=4] > I/O ports at 12200 [size=8] > I/O ports at 12100 [size=4] > I/O ports at 12000 [size=16] > Memory at fb807000 (32-bit, non-prefetchable) [size=512] > Expansion ROM at fb880000 [disabled] [size=512K] > Capabilities: [60] Power Management version 2 > Kernel modules: sata_sil > > > /dev/sdc: > > Model=WDC WD5000AAKS-00TMA0 , FwRev=12.01C01, SerialNo= WD-WMAPW1390228 > Config={ HardSect NotMFM HdSw>15uSec SpinMotCtl Fixed DTR>5Mbs FmtGapReq } > RawCHS=16383/16/63, TrkSize=0, SectSize=0, ECCbytes=50 > BuffType=unknown, BuffSize=16384kB, MaxMultSect=16, MultSect=?0? > CurCHS=16383/16/63, CurSects=16514064, LBA=yes, LBAsects=976773168 > IORDY=on/off, tPIO={min:120,w/IORDY:120}, tDMA={min:120,rec:120} > PIO modes: pio0 pio3 pio4 > DMA modes: mdma0 mdma1 mdma2 > UDMA modes: udma0 udma1 udma2 udma3 udma4 *udma5 udma6 > AdvancedPM=no WriteCache=disabled > Drive conforms to: Unspecified: ATA/ATAPI-1,2,3,4,5,6,7 > * signifies the current active mode Ok - looks normal. > I'd like to add that I've been using the exact same card and hard drive > on one of my x86 box for month without any issue. That doesn't mean the driver and chip operate 100% correctly. Has anyone run exhaustive tests to detect data corruption with this card? Looking at the sil_interrupt() code, it seems a PCI Master abort was sometimes expected when reading the bmdma2 register: u32 bmdma2 = readl(mmio_base + sil_port[ap->port_no].bmdma2); ... if (bmdma2 == 0xffffffff || !(bmdma2 & (SIL_DMA_COMPLETE | SIL_DMA_SATA_IRQ))) continue; Also, older X86 platforms generally don't have an IOMMU (newer ones will) and thus can't validate DMA transactions. I don't have the impression that's the problem here though. But someone needs to decode the "Word2" of the HPMC dump that you already provided. > Also I've reproduce the problem again and this time I've had this message right before the "end_request" line : > [ 163.039983] timer_interrupt(CPU 0): delayed! cycles EB5EE1E6 rem 36AE6 next/now 706F7328/5BCE550E Error handlers sometimes don't play nicely with interrupts. I don't know enough about the error handling cases to track this down. > Anything else I can do/provide to troubleshoot this ? Two things: o consider posting some of the original findings on linux-ide and see if anyone has tested this controller on PPC or IA64. I'm looking for any other architecture that has "hard fail" behavior like parisc does. Testing on any other Big Endian HW would be worth hearing about too. o write a quick and dirty "hpmc_symbols" script as described above and run it on the HPMC you provided earlier. Debugging this further wil probably require modifying the sata_sil driver to log (e.g. ktrace) it's activities while under test. hth, grant ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-25 21:36 ` Grant Grundler @ 2009-01-02 20:46 ` Guy Martin 2009-01-02 21:01 ` Kyle McMartin 0 siblings, 1 reply; 20+ messages in thread From: Guy Martin @ 2009-01-02 20:46 UTC (permalink / raw) To: Grant Grundler; +Cc: Kyle McMartin, linux-parisc [-- Attachment #1: Type: text/plain, Size: 1040 bytes --] On Thu, 25 Dec 2008 14:36:25 -0700 Grant Grundler <grundler@parisc-linux.org> wrote: > > Anything else I can do/provide to troubleshoot this ? > > Two things: > o consider posting some of the original findings on linux-ide and see > if anyone has tested this controller on PPC or IA64. I'm looking for > any other architecture that has "hard fail" behavior like parisc does. > Testing on any other Big Endian HW would be worth hearing about too. > > o write a quick and dirty "hpmc_symbols" script as described above > and run it on the HPMC you provided earlier. > > Debugging this further wil probably require modifying the sata_sil > driver to log (e.g. ktrace) it's activities while under test. Will do, just haven't had time yet. I quickly looked into why the OS HPMC handler wasn't triggered and the issue comes from the fact that the length provided to the PDC is incorrect. The doc states that it should be in bytes while we provide it in words. I've attached a small patch that solves this little issue. Cheers, Guy [-- Attachment #2: fix-hpmc-handler-length.diff --] [-- Type: text/x-patch, Size: 433 bytes --] OS HPMC handler length needs to be in bytes, not in words. Signed-off-by: Guy Martin <gmsoft@tuxicoman.be> --- /root/traps.c.orig 2009-01-02 09:20:41.000000000 +0100 +++ arch/parisc/kernel/traps.c 2009-01-02 09:23:47.000000000 +0100 @@ -840,7 +840,7 @@ /* Compute Checksum for HPMC handler */ - length = os_hpmc_end - os_hpmc; + length = (os_hpmc_end - os_hpmc) * sizeof(u32); ivap[7] = length; hpmcp = (u32 *)os_hpmc; ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-02 20:46 ` Guy Martin @ 2009-01-02 21:01 ` Kyle McMartin 2009-01-02 21:14 ` John David Anglin 0 siblings, 1 reply; 20+ messages in thread From: Kyle McMartin @ 2009-01-02 21:01 UTC (permalink / raw) To: Guy Martin; +Cc: Grant Grundler, Kyle McMartin, linux-parisc On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote: > OS HPMC handler length needs to be in bytes, not in words. > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be> > > - length = os_hpmc_end - os_hpmc; > + length = (os_hpmc_end - os_hpmc) * sizeof(u32); > ivap[7] = length; Ick. I wonder if I screwed that up when I fixed the function descriptor breakage... Will apply this fix, thanks Guy. cheers, Kyle ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-02 21:01 ` Kyle McMartin @ 2009-01-02 21:14 ` John David Anglin 2009-01-02 21:18 ` Kyle McMartin 0 siblings, 1 reply; 20+ messages in thread From: John David Anglin @ 2009-01-02 21:14 UTC (permalink / raw) To: Kyle McMartin; +Cc: gmsoft, grundler, kyle, linux-parisc > On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote: > > OS HPMC handler length needs to be in bytes, not in words. > > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be> > > > > - length = os_hpmc_end - os_hpmc; > > + length = (os_hpmc_end - os_hpmc) * sizeof(u32); > > ivap[7] = length; > > Ick. I wonder if I screwed that up when I fixed the function descriptor > breakage... The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are pointing to function descriptors. Dave -- J. David Anglin dave.anglin@nrc-cnrc.gc.ca National Research Council of Canada (613) 990-0752 (FAX: 952-6602) ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-02 21:14 ` John David Anglin @ 2009-01-02 21:18 ` Kyle McMartin 2009-01-02 21:36 ` John David Anglin 2009-01-03 0:09 ` John David Anglin 0 siblings, 2 replies; 20+ messages in thread From: Kyle McMartin @ 2009-01-02 21:18 UTC (permalink / raw) To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc On Fri, Jan 02, 2009 at 04:14:12PM -0500, John David Anglin wrote: > > On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote: > > > OS HPMC handler length needs to be in bytes, not in words. > > > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be> > > > > > > - length = os_hpmc_end - os_hpmc; > > > + length = (os_hpmc_end - os_hpmc) * sizeof(u32); > > > ivap[7] = length; > > > > Ick. I wonder if I screwed that up when I fixed the function descriptor > > breakage... > > The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are > pointing to function descriptors. > I think we've been over this before ;-) commit c3d4ed4e3e5aa8d9e6b4b795f004a7028ce780e9 Author: Kyle McMartin <kyle@parisc-linux.org> Date: Mon Jun 4 02:26:52 2007 -0400 [PARISC] Fix kernel panic in check_ivt check_ivt had some seriously broken code wrt function pointers on parisc64. Instead of referencing the hpmc code via a function pointer, export symbols and reference it as a const array. Thanks to jda for pointing out the broken 64-bit func ptr handling. Signed-off-by: Kyle McMartin <kyle@parisc-linux.org> cheers, Kyle ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-02 21:18 ` Kyle McMartin @ 2009-01-02 21:36 ` John David Anglin 2009-01-03 0:09 ` John David Anglin 1 sibling, 0 replies; 20+ messages in thread From: John David Anglin @ 2009-01-02 21:36 UTC (permalink / raw) To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc > On Fri, Jan 02, 2009 at 04:14:12PM -0500, John David Anglin wrote: > > > On Fri, Jan 02, 2009 at 09:46:44PM +0100, Guy Martin wrote: > > > > OS HPMC handler length needs to be in bytes, not in words. > > > > Signed-off-by: Guy Martin <gmsoft@tuxicoman.be> > > > > > > > > - length = os_hpmc_end - os_hpmc; > > > > + length = (os_hpmc_end - os_hpmc) * sizeof(u32); > > > > ivap[7] = length; > > > > > > Ick. I wonder if I screwed that up when I fixed the function descriptor > > > breakage... > > > > The fix doesn't seem right. Probably, os_hpmc and os_hpmc_end are > > pointing to function descriptors. > > > > I think we've been over this before ;-) Looking at traps.o for a 64-bit build, I see the references are: 0000000000000048 R_PARISC_DLTIND21L os_hpmc_end 0000000000000050 R_PARISC_DLTIND14R os_hpmc_end 0000000000000058 R_PARISC_DLTIND21L os_hpmc 000000000000005c R_PARISC_DLTIND14R os_hpmc Wonder what the DLT is pointing to. Dave -- J. David Anglin dave.anglin@nrc-cnrc.gc.ca National Research Council of Canada (613) 990-0752 (FAX: 952-6602) ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-02 21:18 ` Kyle McMartin 2009-01-02 21:36 ` John David Anglin @ 2009-01-03 0:09 ` John David Anglin 2009-01-03 0:44 ` Kyle McMartin 1 sibling, 1 reply; 20+ messages in thread From: John David Anglin @ 2009-01-03 0:09 UTC (permalink / raw) To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc > I think we've been over this before ;-) One safe way to compute this difference is using a difference of local labels in .data. Dave -- J. David Anglin dave.anglin@nrc-cnrc.gc.ca National Research Council of Canada (613) 990-0752 (FAX: 952-6602) ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-03 0:09 ` John David Anglin @ 2009-01-03 0:44 ` Kyle McMartin 2009-01-03 1:24 ` John David Anglin 0 siblings, 1 reply; 20+ messages in thread From: Kyle McMartin @ 2009-01-03 0:44 UTC (permalink / raw) To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc On Fri, Jan 02, 2009 at 07:09:57PM -0500, John David Anglin wrote: > > I think we've been over this before ;-) > > One safe way to compute this difference is using a difference > of local labels in .data. > Something like this? Sorry, I've almost completely forgotten gas syntax and all that jazz... I seem to recall writing a patch to dump the first 16 insns of os_hpmc from the check_ivt call, and found that it was working sensibly on 64-bit. regards, Kyle diff --git a/arch/parisc/kernel/hpmc.S b/arch/parisc/kernel/hpmc.S index 2cbf13b..02e41da 100644 --- a/arch/parisc/kernel/hpmc.S +++ b/arch/parisc/kernel/hpmc.S @@ -295,5 +295,9 @@ os_hpmc_6: b . nop ENDPROC(os_hpmc) -ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */ +ENTRY(os_hpmc_end) nop +.data + .export os_hpmc_size +os_hpmc_size: + .word os_hpmc_end-os_hpmc diff --git a/arch/parisc/kernel/traps.c b/arch/parisc/kernel/traps.c index 4c771cd..5cb66ac 100644 --- a/arch/parisc/kernel/traps.c +++ b/arch/parisc/kernel/traps.c @@ -821,8 +821,8 @@ void handle_interruption(int code, struct pt_regs *regs) int __init check_ivt(void *iva) { - extern const u32 os_hpmc[]; - extern const u32 os_hpmc_end[]; + extern u32 os_hpmc_size; + extern u32 os_hpmc[]; int i; u32 check = 0; @@ -839,8 +839,7 @@ int __init check_ivt(void *iva) *ivap++ = 0; /* Compute Checksum for HPMC handler */ ^ permalink raw reply related [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-03 0:44 ` Kyle McMartin @ 2009-01-03 1:24 ` John David Anglin 2009-01-03 1:27 ` Kyle McMartin 0 siblings, 1 reply; 20+ messages in thread From: John David Anglin @ 2009-01-03 1:24 UTC (permalink / raw) To: Kyle McMartin; +Cc: kyle, gmsoft, grundler, linux-parisc > Something like this? Sorry, I've almost completely forgotten gas syntax > and all that jazz... Yes. The gas syntax looks correct except you probably should add a .align 4. I would use local labels (that's what's what is done for debug and exception info). Just place the local labels adjacent to os_hpmc and os_hpmc_end. This ensures the diff is absolute. > -ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */ > +ENTRY(os_hpmc_end) Probably, you don't need to export os_hpmc_end. > - length = os_hpmc_end - os_hpmc; > + length = os_hpmc_size * sizeof(u32); I believe the size will be correct without multiplying by sizeof(u32). Dave -- J. David Anglin dave.anglin@nrc-cnrc.gc.ca National Research Council of Canada (613) 990-0752 (FAX: 952-6602) ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2009-01-03 1:24 ` John David Anglin @ 2009-01-03 1:27 ` Kyle McMartin 0 siblings, 0 replies; 20+ messages in thread From: Kyle McMartin @ 2009-01-03 1:27 UTC (permalink / raw) To: John David Anglin; +Cc: Kyle McMartin, gmsoft, grundler, linux-parisc On Fri, Jan 02, 2009 at 08:24:50PM -0500, John David Anglin wrote: > > Something like this? Sorry, I've almost completely forgotten gas syntax > > and all that jazz... > > Yes. The gas syntax looks correct except you probably should add a > .align 4. I would use local labels (that's what's what is done for > debug and exception info). Just place the local labels adjacent to > os_hpmc and os_hpmc_end. This ensures the diff is absolute. > Ahh, excellent. Thanks. > > -ENTRY(os_hpmc_end) /* this label used to compute os_hpmc checksum */ > > +ENTRY(os_hpmc_end) > > Probably, you don't need to export os_hpmc_end. > Indeed. > > - length = os_hpmc_end - os_hpmc; > > + length = os_hpmc_size * sizeof(u32); > > I believe the size will be correct without multiplying by sizeof(u32). > Ah, indeed. Good point. cheers, Kyle ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: HPMC bus timeout on C3600 2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin 2008-12-20 22:16 ` Kyle McMartin @ 2008-12-21 8:10 ` Grant Grundler 1 sibling, 0 replies; 20+ messages in thread From: Grant Grundler @ 2008-12-21 8:10 UTC (permalink / raw) To: Guy Martin; +Cc: linux-parisc On Sat, Dec 20, 2008 at 01:39:15PM +0100, Guy Martin wrote: > I'm recently running into HPMC bus timeout problem. While this never > caused any problem before, I've noticed a few strange things. I've > attached the output I get from "pim hpmc" in the PDC. > > This BUS timeout occurs after one or two days of idling when I plug a > specific card reader in the USB PCI card I plugged in. I think this > card reader causes problems because it's being polled every 2-3 seconds. > > Nevertheless, looking at the pdc output, it seems that the OS HPMC > handler is not kicking off according to the chassis code CBF2 and CBFC. > CBF2 : bas OS HPMC len > CBFC : OS HPMC br err > > I've check GR02 and it points to inw() which makes sens. Nothing > interesting there. > > The failing device appears to be the built-in NIC according the path > provided in the output (10/0/12/0) and the C3600 service manual (figure > 5-2). It's likely the NIC (tulip driver) is just more active than the other USB controller and thus becomes the "victim" of the USB misbehavior. This is very common for DMA programming bugs. "word2" from this bit of the HPMC needs to be decoded: I/O Module Error Log Information: Timestamp = Tue Dec 16 21:23:47 GMT 2008 (20:08:12:16:21:23:47) '9000/785 B,C,J Workstation IO Error Log', rev 0, 228 bytes: Rope Word1 Word2 Word3 ------ ------------ ------------ 0 0x00000000 0x0e0cc2a9 0x00000000fed30048 ... Both the USB controller and NIC are on Rope 0. If the USB controller falls over, it will take all devices on that bus with it. We've seen "0x0e0cc2a9" before but it wasn't decoded then either. > Even if this look HW pb, I've had a few reports recently about HPMC PCI > timeout issues and I have a few doubts. > > Now my questions are : > - does this look like an HW or SW problems ? More likely SW. > - why isn't the HPMC handler kicked off ? No Idea. Probably not hooked in correctly. But this might also explain why you haven't seen many HPMC bug reports. > - could the HPMC handler recover this ? No. > - what debug can I enable to get more info about this if applicable ? I'd look carefully at USB DMA streams. This will likely require adding debug code to USB drivers. Mostly to dump when a USB buffer is mapped and when it's unmap. hth, grant ^ permalink raw reply [flat|nested] 20+ messages in thread
end of thread, other threads:[~2009-01-03 1:27 UTC | newest] Thread overview: 20+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2008-12-20 12:39 HPMC bus timeout on C3600 Guy Martin 2008-12-20 22:16 ` Kyle McMartin 2008-12-20 22:21 ` Kyle McMartin 2008-12-21 8:12 ` Grant Grundler 2008-12-22 21:40 ` Guy Martin 2008-12-22 21:47 ` Kyle McMartin 2008-12-22 22:49 ` Guy Martin 2008-12-25 8:06 ` Grant Grundler 2008-12-25 15:57 ` Guy Martin 2008-12-25 21:36 ` Grant Grundler 2009-01-02 20:46 ` Guy Martin 2009-01-02 21:01 ` Kyle McMartin 2009-01-02 21:14 ` John David Anglin 2009-01-02 21:18 ` Kyle McMartin 2009-01-02 21:36 ` John David Anglin 2009-01-03 0:09 ` John David Anglin 2009-01-03 0:44 ` Kyle McMartin 2009-01-03 1:24 ` John David Anglin 2009-01-03 1:27 ` Kyle McMartin 2008-12-21 8:10 ` Grant Grundler
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.