From mboxrd@z Thu Jan 1 00:00:00 1970 From: "Wiles, Keith" Subject: Re: Occasional instability in RSS Hashes/Queues from X540 NIC Date: Thu, 4 May 2017 16:34:58 +0000 Message-ID: <7B7539B0-8CB0-4DDB-B329-D11F295A2604@intel.com> References: Mime-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: quoted-printable Cc: "dev@dpdk.org" To: Matt Laswell Return-path: Received: from mga01.intel.com (mga01.intel.com [192.55.52.88]) by dpdk.org (Postfix) with ESMTP id 9B9B57D0D for ; Thu, 4 May 2017 18:35:01 +0200 (CEST) In-Reply-To: Content-Language: en-US Content-ID: List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org Sender: "dev" > On May 4, 2017, at 8:04 AM, Matt Laswell wrote: >=20 > Hey Folks, >=20 > I'm seeing some strange behavior with regard to the RSS hash values in my > applications and was hoping somebody might have some pointers on where to > look. In my application, I'm using RSS to divide work among multiple > cores, each of which services a single RX queue. When dealing with a > single long-lived TCP connection, I occasionally see packets going to the > wrong core. That is, almost all of the packets in the connection go to > core 5 in this case, but every once in a while, one goes to core 0 instea= d. >=20 > Upon further investigation, I find two problems are occurring. The first > is that problem packets have the RSS hash value in the mbuf incorrectly s= et > to zero. They are therefore put in queue zero, where they are read by co= re > zero. Other packets from the same connection that occur immediately befo= re > and after the packet in question have the correct hash value and therefor= e > go to a different core. The second problem is that we sometimes see > packets in which the RSS hash in the mbuf appears correct, but the packet= s > are incorrectly put into queue zero. As with the first, this results in > the wrong core getting the packet. Either one of these confuses the stat= e > tracking we're doing per-core. >=20 > A few details: >=20 > - Using an Intel X540-AT2 NIC and the igb_uio driver > - DPDK 16.04 > - A particular packet in our workflow always encounters this problem. > - Retransmissions of the packet in question also encounter the problem > - The packet is IPv4, with header length of 20 (so no options), no > fragmentation. > - The only differences I can see in the IP header between packets that > get the right hash value and those that get the wrong one are in the IP= ID, > total length, and checksum fields. > - Using ETH_RSS_IPV4 > - The packet is TCP with about 100 bytes of payload - it's not a jumbo > or a runt > - We fill the key in with 0x6d5a to get symmetric hashing of both sides > of the connection > - We only configure RSS information at boot; things like the key or > header fields are not being changed dynamically > - Traffic load is light when the problem occurs >=20 > Is anybody aware of an errata, either in the NIC or the PMD's configurati= on > of it that might explain something like this? Failing that, if you ran > into this sort of behavior, how would you approach finding the reason for > the error? Every failure mode I can think of would tend to affect all of > the packets in the connection consistently, even if incorrectly. Just to add more information to this email, can you provide hexdumps of the= packets to help someone maybe spot the problem? Need the previous OK packet plus the one after it and the failing packets y= ou are seeing. I do not know why this is happening as I do not know of any errata to expla= in this issue. >=20 > Thanks in advance for any ideas. >=20 > -- > Matt Laswell > laswell@infinite.io Regards, Keith