From mboxrd@z Thu Jan 1 00:00:00 1970 From: "Wiles, Keith" Subject: Re: Question about DPDK hugepage fd change Date: Tue, 5 Feb 2019 20:29:12 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: quoted-printable Cc: "dev@dpdk.org" , "edwin.leung@oracle.com" To: Iain Barker Return-path: Received: from mga03.intel.com (mga03.intel.com [134.134.136.65]) by dpdk.org (Postfix) with ESMTP id 6A5775F19 for ; Tue, 5 Feb 2019 21:29:15 +0100 (CET) In-Reply-To: Content-Language: en-US Content-ID: <6D06B580B82E5B4683C26D3F66304775@intel.com> List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org Sender: "dev" > On Feb 5, 2019, at 12:56 PM, Iain Barker wrote: >=20 > Hi everyone, >=20 > We just updated our application from DPDK 17.11.4 (LTS) to DPDK 18.11 (LT= S) and we noticed a regression. >=20 > Our host platform is providing 2MB huge pages, so for 8GB reservation thi= s means 4000 pages are allocated. >=20 > This worked fine in the prior LTS, but after upgrading DPDK what we are s= eeing is that select() on an fd is failing. >=20 > select() works fine when the process starts up, but does not work after D= PDK has been initialized. >=20 > We did some investigation and found in the DPDK patches linked below, the= hugepage tracking mechanism was changed from mmap to an array of file desc= riptors, and the rlimit for fd's is raised from the default to allow more f= d's to be open. >=20 > https://mails.dpdk.org/archives/dev/2018-September/110890.html > https://mails.dpdk.org/archives/dev/2018-September/110889.html >=20 > The problem is that the GNU C library (glibc) has a limit for the maximum= fd passed to select(), and is hard-coded in the POSIX header file and libc= at 1024 (and probably many other OS libraries too as a result). >=20 > Raising the rlimit for fd >1024 has undefined results, per the manpage: >=20 > http://man7.org/linux/man-pages/man2/select.2.html > An fd_set is a fixed size buffer. Executing FD_CLR() or FD_SET() > with a value of fd that is negative or is equal to or larger than > FD_SETSIZE will result in undefined behavior. Moreover, POSIX > requires fd to be a valid file descriptor. >=20 > The Linux kernel allows file descriptor sets of arbitrary size, > determining the length of the sets to be checked from the value of > nfds. However, in the glibc implementation, the fd_set type is fixed > in size. >=20 > Specifically, libc's header include/sys/select.h has an array of fd's whi= ch is FD_SETSIZE deep. > __fd_mask fds_bits[__FD_SETSIZE / __NFDBITS]; >=20 > and usr/include/linux/posix_types.h is hard-coded with > #define __FD_SETSIZE 1024 >=20 > As this define and array are in libc, they are used in many libraries on = a Linux system. So to use setsize >1024 means recompiling OS libraries and = any other package that needs to use FDs, or ensuring that no library used b= y the application ever calls select() on an fd set. That seems an unreasona= ble burden... >=20 > Any thoughts? Would poll work here instead? >=20 > thanks, > Iain Regards, Keith