• Re: Support of Unicode characters in various places (was Re: c needs this sign)

    From Keith Thompson@Keith.S.Thompson+u@gmail.com to comp.lang.c on Thu Sep 10 14:54:37 2026
    From Newsgroup: comp.lang.c

    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    I've yet to hear that anyone actually used trigraphs. - Weren't they,
    because of that, even removed from recent "C" (or C++) standards?

    I've used trigraphs, but only in deliberately obfuscated code. I'm not
    aware of any code that uses trigraphs because they're useful.

    Part of the point of trigraphs was to enable the use of C on
    EBCDIC-based systems that don't have all the characters C requires,
    including '{' and '}'. (There are EBCDIC code pages that do have those characters, but they're outside the "invariant subset".)

    It's possible that trigraphs have been used accidentally more often
    than they've been used intentionally:

    fprintf(stderr, "What the heck just happened??!\n");

    Even more fun (though this is deliberately contrived, not accidental)
    (??/ expands to \) :

    #include <stdio.h>
    int main(void) {
    // Is this a multi-line comment ??/
    if (1) puts("No, it isn't."); else
    puts("Yes, it is.");
    }

    Yes, trigraphs were removed in the C23 standard. (C++ removed them
    in C++17.)

    A <iso646.h> style is better, where someone lacking a ^ key can
    write "xor" instead.

    I don't know what the <iso646.h> header actually is. - The "problem"
    with "ISO 646" is that there's many (national) variants of it. - If
    we're speaking about the ISO 646 IRV (International Reference Version) there's a '^' available at least. (Which doesn't mean it's available
    on specific national keyboards, though.)

    <iso646.h> is a standard header, introduced in the 1995 amendment
    to the C90 standard. It defines 11 macros, intended for use on
    systems that use character sets that *aren't* compatible with ISO
    646 (basically 7-bit ASCII), or at least that make it difficult
    to use certain ASCII characters. The macros are:

    #define and &&
    #define and_eq &=
    #define bitand &
    #define bitor |
    #define compl ~
    #define not !
    #define not_eq !=
    #define or ||
    #define or_eq |=
    #define xor ^
    #define xor_eq ^=

    In my experience, they're rarely used.

    C95 also added *digraphs*, alternate spellings of certain tokens:

    <: [
    ]
    <% {
    }
    %: #
    %:%: ##

    <iso646.h> and digraphs operate on the token level. Trigraphs are
    (were) replaced in translation phase 1, so they can be used in
    character constants, string literals, header names, comments,
    and so forth.

    Presumably there are still C programmers who develop on ECBDIC-based
    systems. I don't know what solutions they typically use for required characters that aren't supported.

    [...]
    --
    Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
    void Void(void) { Void(); } /* The recursive call of the void */
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.lang.c on Thu Sep 10 22:17:04 2026
    From Newsgroup: comp.lang.c

    Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    I've yet to hear that anyone actually used trigraphs. - Weren't they,
    because of that, even removed from recent "C" (or C++) standards?

    I've used trigraphs, but only in deliberately obfuscated code. I'm not
    aware of any code that uses trigraphs because they're useful.

    Part of the point of trigraphs was to enable the use of C on
    EBCDIC-based systems that don't have all the characters C requires,
    including '{' and '}'. (There are EBCDIC code pages that do have those >characters, but they're outside the "invariant subset".)

    The other part may have been for those still using ASR-33 teletypes
    which were missing several characters required in C, including
    curly braces, vertical bar, backtick and tilde.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Keith Thompson@Keith.S.Thompson+u@gmail.com to comp.lang.c on Thu Sep 10 15:48:22 2026
    From Newsgroup: comp.lang.c

    scott@slp53.sl.home (Scott Lurndal) writes:
    Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    I've yet to hear that anyone actually used trigraphs. - Weren't they,
    because of that, even removed from recent "C" (or C++) standards?

    I've used trigraphs, but only in deliberately obfuscated code. I'm not >>aware of any code that uses trigraphs because they're useful.

    Part of the point of trigraphs was to enable the use of C on
    EBCDIC-based systems that don't have all the characters C requires, >>including '{' and '}'. (There are EBCDIC code pages that do have those >>characters, but they're outside the "invariant subset".)

    The other part may have been for those still using ASR-33 teletypes
    which were missing several characters required in C, including
    curly braces, vertical bar, backtick and tilde.

    Apparently the ASR-33 generated 7-bit ASCII, but didn't support
    lowercase letters.

    The stty command had (and still has, at least in some versions)
    options to map uppercase characters to lowercase, and to generate
    uppercase with a '\' prefix, so Hello would be entered as \HELLO.
    I suppose a C programmer would likely be using a line editor like
    "ed" to edit code, which would be subject to the current tty
    settings, and since C is mostly lowercase the translation would
    make it tolerable.

    I wonder how many ASR-33s were still in use for C development by
    the time <iso646.h> and trigraphs were introduced in 1995.
    --
    Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
    void Void(void) { Void(); } /* The recursive call of the void */
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Janis Papanagnou@janis_papanagnou+ng@hotmail.com to comp.lang.c on Fri Sep 11 02:03:53 2026
    From Newsgroup: comp.lang.c

    On 2026-09-10 23:54, Keith Thompson wrote:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]

    It's possible that trigraphs have been used accidentally more often
    than they've been used intentionally:

    fprintf(stderr, "What the heck just happened??!\n");

    LOL! - That's really great; you made my day! 8-)

    It never occurred to me that you may stumble into such a case.
    (Luckily the default settings of my compiler warns about it.)

    (BTW, in my youth I had specified some detailed semantics (and
    connotations) for the textual multiplication of '?' and '!' up
    to a length of 3 - it had not been common to express more than
    expressible with the single signs. I considered it worthy since
    the expressive power of plain '!' and '?' wasn't sufficient for
    the finer tones.)


    Even more fun (though this is deliberately contrived, not accidental)
    (??/ expands to \) :

    #include <stdio.h>
    int main(void) {
    // Is this a multi-line comment ??/
    if (1) puts("No, it isn't."); else
    puts("Yes, it is.");
    }

    Yeah, some features just invite to do fancy things. :-)


    Yes, trigraphs were removed in the C23 standard. (C++ removed them
    in C++17.)

    A <iso646.h> style is better, where someone lacking a ^ key can
    write "xor" instead.

    I don't know what the <iso646.h> header actually is. - [...]

    <iso646.h> is a standard header, introduced in the 1995 amendment
    to the C90 standard. It defines 11 macros, intended for use on
    systems that use character sets that *aren't* compatible with ISO
    646 (basically 7-bit ASCII), or at least that make it difficult
    to use certain ASCII characters. [...]

    Aha. Thanks.

    [...]

    <iso646.h> and digraphs operate on the token level. Trigraphs are
    (were) replaced in translation phase 1, so they can be used in
    character constants, string literals, header names, comments,
    and so forth.

    Interesting. Thanks.

    Janis

    [...]
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Keith Thompson@Keith.S.Thompson+u@gmail.com to comp.lang.c on Thu Sep 10 17:13:57 2026
    From Newsgroup: comp.lang.c

    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    (BTW, in my youth I had specified some detailed semantics (and
    connotations) for the textual multiplication of '?' and '!' up
    to a length of 3 - it had not been common to express more than
    expressible with the single signs. I considered it worthy since
    the expressive power of plain '!' and '?' wasn't sufficient for
    the finer tones.)

    Really??? Tell me more!!!

    r+yr+yr+yOr maybe notrC+rC+rC+
    --
    Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
    void Void(void) { Void(); } /* The recursive call of the void */
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Janis Papanagnou@janis_papanagnou+ng@hotmail.com to comp.lang.c on Fri Sep 11 02:24:40 2026
    From Newsgroup: comp.lang.c

    On 2026-09-11 02:13, Keith Thompson wrote:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    (BTW, in my youth I had specified some detailed semantics (and
    connotations) for the textual multiplication of '?' and '!' up
    to a length of 3 - it had not been common to express more than
    expressible with the single signs. I considered it worthy since
    the expressive power of plain '!' and '?' wasn't sufficient for
    the finer tones.)

    Really??? Tell me more!!!

    Yeah, these two are the simple ones; plain amplifications, IMO.
    It's getting more interesting in the difference of mixed forms.
    But, honestly, I don't recall the details any more. Back these
    days I had no computer and the pencil-and-paper-storage-device
    I used has long decayed, I fear.

    r+yr+yr+yOr maybe notrC+rC+rC+

    Better not. (Not here.) :-)

    Janis

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c on Fri Sep 11 09:42:45 2026
    From Newsgroup: comp.lang.c

    On 11/09/2026 02:13, Keith Thompson wrote:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    (BTW, in my youth I had specified some detailed semantics (and
    connotations) for the textual multiplication of '?' and '!' up
    to a length of 3 - it had not been common to express more than
    expressible with the single signs. I considered it worthy since
    the expressive power of plain '!' and '?' wasn't sufficient for
    the finer tones.)

    Really??? Tell me more!!!

    r+yr+yr+yOr maybe notrC+rC+rC+


    PHP has =, ==, and === as different operators, as well as ??. Maybe the
    ??? and !!! are not as unrealistic as they sound !!!


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From bart@bc@freeuk.com to comp.lang.c on Fri Sep 11 10:51:45 2026
    From Newsgroup: comp.lang.c

    On 10/09/2026 23:48, Keith Thompson wrote:
    scott@slp53.sl.home (Scott Lurndal) writes:
    Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    I've yet to hear that anyone actually used trigraphs. - Weren't they,
    because of that, even removed from recent "C" (or C++) standards?

    I've used trigraphs, but only in deliberately obfuscated code. I'm not
    aware of any code that uses trigraphs because they're useful.

    Part of the point of trigraphs was to enable the use of C on
    EBCDIC-based systems that don't have all the characters C requires,
    including '{' and '}'. (There are EBCDIC code pages that do have those
    characters, but they're outside the "invariant subset".)

    The other part may have been for those still using ASR-33 teletypes
    which were missing several characters required in C, including
    curly braces, vertical bar, backtick and tilde.


    (Where does C require backtick?)

    Apparently the ASR-33 generated 7-bit ASCII, but didn't support
    lowercase letters.

    The stty command had (and still has, at least in some versions)
    options to map uppercase characters to lowercase, and to generate
    uppercase with a '\' prefix, so Hello would be entered as \HELLO.
    I suppose a C programmer would likely be using a line editor like
    "ed" to edit code, which would be subject to the current tty
    settings, and since C is mostly lowercase the translation would
    make it tolerable.

    I wonder how many ASR-33s were still in use for C development by
    the time <iso646.h> and trigraphs were introduced in 1995.


    I used ASR33s and other terminals with various limitations from 1976.
    The first computer I /made/ used 6-bit (I think, ASCII codes 32 to 95
    offset to be 0 to 63, so upper-case only), in 1981.

    How did people even write C on such systems? Given that much of it /had/
    to be lower-case. It must have looked appalling. (I didn't think it
    looked that great anyway!)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c on Fri Sep 11 12:49:33 2026
    From Newsgroup: comp.lang.c

    On 11/09/2026 11:51, bart wrote:
    On 10/09/2026 23:48, Keith Thompson wrote:
    scott@slp53.sl.home (Scott Lurndal) writes:
    Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    The other part may have been for those still using ASR-33 teletypes
    which were missing several characters required in C, including
    curly braces, vertical bar, backtick and tilde.


    (Where does C require backtick?)

    It was added to the required source and execution character set in C23
    (along with @ and $), though I don't believe it has any specific use in
    the language. @ has no specific use either, while $ may be an
    additional implementation-defined "letter" in identifiers. I guess it
    was just added to the list of required characters while tidying up with
    the removal of trigraphs, as these are characters that people can find
    useful and which are present in any real-world C implementation.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.lang.c on Fri Sep 11 16:28:06 2026
    From Newsgroup: comp.lang.c

    bart <bc@freeuk.com> writes:
    On 10/09/2026 23:48, Keith Thompson wrote:
    scott@slp53.sl.home (Scott Lurndal) writes:
    Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
    [...]
    I've yet to hear that anyone actually used trigraphs. - Weren't they, >>>>> because of that, even removed from recent "C" (or C++) standards?

    I've used trigraphs, but only in deliberately obfuscated code. I'm not >>>> aware of any code that uses trigraphs because they're useful.

    Part of the point of trigraphs was to enable the use of C on
    EBCDIC-based systems that don't have all the characters C requires,
    including '{' and '}'. (There are EBCDIC code pages that do have those >>>> characters, but they're outside the "invariant subset".)

    The other part may have been for those still using ASR-33 teletypes
    which were missing several characters required in C, including
    curly braces, vertical bar, backtick and tilde.


    (Where does C require backtick?)

    Apparently the ASR-33 generated 7-bit ASCII, but didn't support
    lowercase letters.

    The stty command had (and still has, at least in some versions)
    options to map uppercase characters to lowercase, and to generate
    uppercase with a '\' prefix, so Hello would be entered as \HELLO.
    I suppose a C programmer would likely be using a line editor like
    "ed" to edit code, which would be subject to the current tty
    settings, and since C is mostly lowercase the translation would
    make it tolerable.

    I wonder how many ASR-33s were still in use for C development by
    the time <iso646.h> and trigraphs were introduced in 1995.


    I used ASR33s and other terminals with various limitations from 1976.
    The first computer I /made/ used 6-bit (I think, ASCII codes 32 to 95
    offset to be 0 to 63, so upper-case only), in 1981.

    How did people even write C on such systems?

    Did you read what you replied to? The Unix systems that most
    C programmers were using had a terminal driver that would
    translate upper-case to lower-case automatically, if so
    configured. The stty(1) command was used to change terminal
    driver confgurations. The login(1) command would automatically
    set that flag if the username was typed in upper-case.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.lang.c on Tue Sep 15 07:23:21 2026
    From Newsgroup: comp.lang.c

    On Wed, 9 Sep 2026 23:44:48 +0100, bart wrote:

    It does work, but I have to use Notepad to view it

    ArenrCOt you using an OS (Windows NT) which was just about the first to
    embrace Unicode right into its core APIs, back in 1993?

    The fact that it still doesnrCOt make it easy for you to work natively
    with Unicode text ... just a teentsy weentsy bit of an irony there,
    donrCOt you think?
    --- Synchronet 3.22a-Linux NewsLink 1.2