Linux Manual Pages development
 help / color / mirror / Atom feed
From: Alejandro Colomar <alx@kernel.org>
To: Mark Harris <mark.hsj@gmail.com>
Cc: sergeh@kernel.org, linux-man@vger.kernel.org,
	 "Serge E. Hallyn" <serge@hallyn.com>,
	"G. Branden Robinson" <g.branden.robinson@gmail.com>,
	 Douglas McIlroy <douglas.mcilroy@dartmouth.edu>
Subject: Re: alx-0096r1 - string (and nonstring) copying
Date: Thu, 30 Jul 2026 22:02:08 +0200	[thread overview]
Message-ID: <amus_s5gwmZ69UU7@devuan> (raw)
In-Reply-To: <amupqzN_X57Du4AJ@devuan>

[-- Attachment #1: Type: text/plain, Size: 8295 bytes --]

> Date: 2026-07-30 21:51:16+0200
> From: Alejandro Colomar <alx@kernel.org>
>
> Hi Mark,
> 
> > Date: 2026-07-30 10:58:26-0700
> > From: Mark Harris <mark.hsj@gmail.com>
> >
> [...]
> > > It would be good to know if those systems that have a larger ptrdiff_t
> > > are still alive or can be ignored/forced to adapt.
> > 
> > SDCC (https://sdcc.sourceforge.net/), when building for mcs51
> > (-mmcs51) or ds390 (-mds390), uses a 16-bit size_t and 32-bit
> > ptrdiff_t.  It appears that Open Watcom C/16 v2
> > (https://github.com/open-watcom/open-watcom-v2) with the huge memory
> > model also uses a 16-bit size_t and a 32-bit ptrdiff_t, although I
> > haven't verified that one.  Both compilers have active development and
> > support some of the new features in C23, but do not have full C23 or
> > POSIX support.  I don't know about any others with ptrdiff_t larger
> > than size_t.
> 
> Thanks; I guess that can't be done.
> 
> [...]
> > > > https://github.com/illumos/illumos-gate/blob/4b44494cae63d4bf0c7d3f34828503e68d1c0e69/usr/src/lib/libc/port/stdio/fgets.c#L44
> > >
> > > Thanks!
> > >
> > > Let's now consider an alternative world where there was a function with
> > > semantics similar to memccpy(3) with one difference: it wouldn't copy
> > > the delimiter.  Let's call such a function memcpyc().  Here's a naive
> > > implementation of memccpy(3):
> > >
> > >         void *
> > >         memccpy(void *dst, const void *src, int c, size_t n)
> > >         {
> > >                 unsigned char  *d = dst;
> > >                 unsigned char  *s = src;
> > >
> > >                 for (size_t i = 0; i < n; i++) {
> > >                         *d++ = *s++;
> > >                         if (*s == c)
> > >                                 return d;
> > >                 }
> > >                 return NULL;
> > >         }
> > >
> > > Here's a naive implementation of the hypothetical memcpyc():
> > >
> > >         void *
> > >         memcpyc(void *dst, const void *src, int c, size_t n)
> > >         {
> > >                 unsigned char  *d = dst;
> > >                 unsigned char  *s = src;
> > >
> > >                 for (size_t i = 0; i < n; i++) {
> > >                         if (*s == c)
> > >                                 return d;
> > >                         *d++ = *s++;
> > >                 }
> > >                 return NULL;
> > >         }
> > >
> > > (I've written both of the above quickly, without compiling nor testing;
> > >  I might have made a mistake.)
> > >
> > > With such a function, the fgets(3) implementation from Illumos gate
> > > would change in this manner:
> > >
> > >         diff --git i/usr/src/lib/libc/port/stdio/fgets.c w/usr/src/lib/libc/port/stdio/fgets.c
> > >         index 9401217410ca..ac1f330ea46e 100644
> > >         --- i/usr/src/lib/libc/port/stdio/fgets.c
> > >         +++ w/usr/src/lib/libc/port/stdio/fgets.c
> > >         @@ -82,9 +82,12 @@ fgets(char *buf, int size, FILE *iop)
> > >                                         break;          /* nothing left to read */
> > >                         }
> > >                         n = (int)(size < iop->_cnt ? size : iop->_cnt);
> > >         -               if ((p = memccpy(ptr, (char *)iop->_ptr, '\n',
> > >         +               if ((p = memcpyc(ptr, (char *)iop->_ptr, '\n',
> > >                             (size_t)n)) != NULL)
> > >         +               {
> > >         +                       p = stpcpy(p, "\n");
> > >                                 n = (int)(p - ptr);
> > >         +               }
> > >                         ptr += n;
> > >                         iop->_cnt -= n;
> > >                         iop->_ptr += n;
> > >
> > > IMO, this wouldn't make it more complex.  However, such a hypothetical
> > > memcpyc() would be much more useful elsewhere.  For example, it would
> > > simplify the two unique uses of memccpy(3) in FreeBSD:
> > >
> > >         diff --git i/bin/sh/parser.c w/bin/sh/parser.c
> > >         index 0c1b7a91c257..713a35ad8cd8 100644
> > >         --- i/bin/sh/parser.c
> > >         +++ w/bin/sh/parser.c
> > >         @@ -2116,7 +2116,7 @@ getprompt(void *unused __unused)
> > >                                 if (fmt[0] != '}') {
> > >                                         char *end;
> > >
> > >         -                               end = memccpy(tfmt, fmt, '}', sizeof(tfmt));
> > >         +                               end = memcpyc(tfmt, fmt, '}', sizeof(tfmt));
> > >                                         if (end == NULL) {
> > >                                                 /*
> > >                                                  * Format too long or no '}', so
> > >         @@ -2129,7 +2129,7 @@ getprompt(void *unused __unused)
> > >                                                 fmt--;
> > >                                                 break;
> > >                                         }
> > >         -                               *--end = '\0'; /* Ignore the copy of '}'. */
> > >         +                               *end = '\0'; /* Ignore the copy of '}'. */
> > >                                         fmt += end - tfmt;
> > >                                 }
> > >                                 now = localtime(&(time_t){time(NULL)});
> > >         diff --git i/lib/libc/amd64/string/strncat.c w/lib/libc/amd64/string/strncat.c
> > >         index 2c63ab50b3c3..4a1e8c0119ac 100644
> > >         --- i/lib/libc/amd64/string/strncat.c
> > >         +++ w/lib/libc/amd64/string/strncat.c
> > >         @@ -19,13 +19,13 @@ strncat(char *dest, const char *src, size_t n)
> > >                 char *endptr;
> > >
> > >                 len = strlen(dest);
> > >         -       endptr = __memccpy(dest + len, src, '\0', n);
> > >         +       endptr = memcpyc(dest + len, src, '\0', n);
> > >
> > >                 /* avoid an extra branch */
> > >                 if (endptr == NULL)
> > >         -               endptr = dest + len + n + 1;
> > >         +               endptr = dest + len + n;
> > >
> > >         -       endptr[-1] = '\0';
> > >         +       *endptr = '\0';
> > >
> > >                 return (dest);
> > >          }
> > >
> > > memccpy(3) is a bad design, which forces complex code, prone to
> > > off-by-one bugs.
> > 
> > It seems like you've already made up your mind that memccpy() is bad
> > and are now just looking to justify that decision.
> 
> Not necessarily.  I can be convinced with good arguments.  :-)
> 
> > Your modification of fgets() is obviously more complex; it adds a
> > whole additional function call, and instead of copying all of the

FWIW, that additional function call is optimized out by GCC even on low
optimization levels, since it's a two-byte copy.  You could perfectly
use a single-byte manual copy:

        diff --git i/usr/src/lib/libc/port/stdio/fgets.c w/usr/src/lib/libc/port/stdio/fgets.c
        index 9401217410ca..ac1f330ea46e 100644
        --- i/usr/src/lib/libc/port/stdio/fgets.c
        +++ w/usr/src/lib/libc/port/stdio/fgets.c
        @@ -82,9 +82,12 @@ fgets(char *buf, int size, FILE *iop)
                                        break;          /* nothing left to read */
                        }
                        n = (int)(size < iop->_cnt ? size : iop->_cnt);
        -               if ((p = memccpy(ptr, (char *)iop->_ptr, '\n',
        +               if ((p = memcpyc(ptr, (char *)iop->_ptr, '\n',
                            (size_t)n)) != NULL)
        +               {
	+			*p++ = '\n';
                                n = (int)(p - ptr);
        +               }
                        ptr += n;
                        iop->_cnt -= n;
                        iop->_ptr += n;

Nevertheless, with your examples of fgets(3) and strncat(3), I'm now
convinced that memccpy(3) is good for this niche use case.

> > bytes at once from a single source it does two copies from two
> > different sources, specifying the delimiter twice making it more
> > error-prone and difficult to change.

This is true, though.


Cheers,
Alex

-- 
<https://www.alejandro-colomar.es>

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 833 bytes --]

  reply	other threads:[~2026-07-30 20:02 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-02 12:33 [PATCH v1 1/2] man/man3/str*.3: NAME: Explain the names Alejandro Colomar
2026-07-02 12:33 ` [PATCH v1 2/2] man/man3/strpbrk.3: BUGS: Clarify the NAME Alejandro Colomar
2026-07-03  5:03 ` [PATCH v1 1/2] man/man3/str*.3: NAME: Explain the names Mark Harris
2026-07-03 11:16   ` Alejandro Colomar
2026-07-08 15:09 ` [PATCH v2 0/4] str*.3, mem*.3: " Alejandro Colomar
2026-07-08 15:09   ` [PATCH v2 1/4] man/man3/str*.3: " Alejandro Colomar
2026-07-24 18:52     ` sergeh
2026-07-25 22:09       ` Alejandro Colomar
2026-07-26 23:14         ` alx-0096r1 - string (and nonstring) copying Alejandro Colomar
2026-07-28 12:05           ` Mark Harris
2026-07-28 15:04             ` Alejandro Colomar
2026-07-28 15:58               ` Alejandro Colomar
2026-07-29 17:27               ` Mark Harris
2026-07-29 19:23                 ` Alejandro Colomar
2026-07-30 17:58                   ` Mark Harris
2026-07-30 19:51                     ` Alejandro Colomar
2026-07-30 20:02                       ` Alejandro Colomar [this message]
2026-07-27  2:39         ` [PATCH v2 1/4] man/man3/str*.3: NAME: Explain the names Serge E. Hallyn
2026-07-27  7:51           ` Alejandro Colomar
2026-07-08 15:09   ` [PATCH v2 2/4] man/man3/strpbrk.3: BUGS: Clarify the NAME Alejandro Colomar
2026-07-24 18:53     ` sergeh
2026-07-25 22:10       ` Alejandro Colomar
2026-07-08 15:09   ` [PATCH v2 3/4] man/man3/mem*.3: NAME: Explain the names Alejandro Colomar
2026-07-24 18:55     ` sergeh
2026-07-08 15:09   ` [PATCH v2 4/4] man/man3/[b]string.3: Rewrite and merge Alejandro Colomar
2026-07-25 23:49 ` [PATCH v3 0/4] str*.3, mem*.3: NAME: Explain the names Alejandro Colomar
2026-07-25 23:49   ` [PATCH v3 1/4] man/man3/str*.3: " Alejandro Colomar
2026-07-25 23:50   ` [PATCH v3 2/4] man/man3/strpbrk.3: BUGS: Clarify the NAME Alejandro Colomar
2026-07-25 23:50   ` [PATCH v3 3/4] man/man3/mem*.3: NAME: Explain the names Alejandro Colomar
2026-07-25 23:50   ` [PATCH v3 4/4] man/man3/[b]string.3: Rewrite and merge Alejandro Colomar
2026-07-29 15:15   ` [PATCH v3 0/4] str*.3, mem*.3: NAME: Explain the names Alejandro Colomar

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amus_s5gwmZ69UU7@devuan \
    --to=alx@kernel.org \
    --cc=douglas.mcilroy@dartmouth.edu \
    --cc=g.branden.robinson@gmail.com \
    --cc=linux-man@vger.kernel.org \
    --cc=mark.hsj@gmail.com \
    --cc=serge@hallyn.com \
    --cc=sergeh@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox