• Working on UTF-8 and counting bytes

    From Janis Papanagnou@janis_papanagnou+ng@hotmail.com to comp.lang.awk on Tue Aug 25 15:57:47 2026
    From Newsgroup: comp.lang.awk

    In GNU Awk I'm working with UTF-8 data and want to count _bytes_.

    $ awk '{print length}'
    aaa
    3
    |n|n|n
    3

    Resetting the locale fixes this...

    $ LC_ALL=C awk '{print length}'
    aaa
    3
    |n|n|n
    6

    ...but spoils that...

    $ LC_ALL=C awk '{print length, ">" substr ($0, 2) "<" }'
    aaa
    3 >aa<
    |n|n|n
    6 >N++|n|n<

    So fiddling with the locale seems inappropriate.

    This one seems to works, but appears to be a bit clumsy...

    $ awk '{print length (sprintf ("%s", $0)), ">" substr ($0, 2) "<" }'
    aaa
    3 >aa<
    |n|n|n
    3 >|n|n<

    Is there any simpler (still native) way to keep string processing
    logic to UTF-8 data intact but have the _byte count_ available?

    Janis

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Janis Papanagnou@janis_papanagnou+ng@hotmail.com to comp.lang.awk on Tue Aug 25 16:01:18 2026
    From Newsgroup: comp.lang.awk

    On 2026-08-25 15:57, Janis Papanagnou wrote:
    In GNU Awk I'm working with UTF-8 data and want to count _bytes_.
    [...]
    This one seems to works, but appears to be a bit clumsy...

    $ awk '{print length (sprintf ("%s", $0)), ">" substr ($0, 2) "<" }'
    aaa
    3 >aa<
    |n|n|n
    3 >|n|n<

    Darn, no! (Need some coffee.) - It should of course have been

    6 >|n|n<

    So the original question remains:


    Is there any simpler (still native) way to keep string processing
    logic to UTF-8 data intact but have the _byte count_ available?

    Janis


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From mack@mack@the-knife.org (Mack The Knife) to comp.lang.awk on Tue Aug 25 14:26:39 2026
    From Newsgroup: comp.lang.awk

    In article <116k77e$32eog$3@dont-email.me>,
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> wrote:
    So the original question remains:

    Is there any simpler (still native) way to keep string processing
    logic to UTF-8 data intact but have the _byte count_ available?

    The "mbs" extension in gawkextlib has an mbs_length() function
    that will do the job:

    ------------------
    @load "mbs"

    { print length($0), mbs_length($0) }
    ------------------

    Try it out.

    Otherwise, no, there is no way within gawk to get what you want.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Janis Papanagnou@janis_papanagnou+ng@hotmail.com to comp.lang.awk on Tue Aug 25 17:12:10 2026
    From Newsgroup: comp.lang.awk

    On 2026-08-25 16:26, Mack The Knife wrote:
    In article <116k77e$32eog$3@dont-email.me>,
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> wrote:
    So the original question remains:

    Is there any simpler (still native) way to keep string processing
    logic to UTF-8 data intact but have the _byte count_ available?

    The "mbs" extension in gawkextlib has an mbs_length() function
    that will do the job:

    ------------------
    @load "mbs"

    { print length($0), mbs_length($0) }
    ------------------

    Try it out.

    cannot open shared library `mbs' for reading: No such file or directory

    Hmm.. - I'm getting an error (seems I need more information about that;
    usually I don't use extensions!); I'll see how to fix my installation.


    (Or maybe I'll just invoke a 'getline' on a Unix "wc -c" co-process.)

    Otherwise, no, there is no way within gawk to get what you want.

    Thanks for the confirmation!

    Janis

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From mack@mack@the-knife.org (Mack The Knife) to comp.lang.awk on Tue Aug 25 18:28:10 2026
    From Newsgroup: comp.lang.awk

    In article <116kbca$32eog$4@dont-email.me>,
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> wrote:
    On 2026-08-25 16:26, Mack The Knife wrote:
    In article <116k77e$32eog$3@dont-email.me>,
    Janis Papanagnou <janis_papanagnou+ng@hotmail.com> wrote:
    So the original question remains:

    Is there any simpler (still native) way to keep string processing
    logic to UTF-8 data intact but have the _byte count_ available?

    The "mbs" extension in gawkextlib has an mbs_length() function
    that will do the job:

    ------------------
    @load "mbs"

    { print length($0), mbs_length($0) }
    ------------------

    Try it out.

    cannot open shared library `mbs' for reading: No such file or directory

    Hmm.. - I'm getting an error (seems I need more information about that; >usually I don't use extensions!); I'll see how to fix my installation.


    (Or maybe I'll just invoke a 'getline' on a Unix "wc -c" co-process.)

    Otherwise, no, there is no way within gawk to get what you want.

    Thanks for the confirmation!

    You're welcome.

    The gawkextlib project is separate from gawk itself. There are pointers
    to it in the gawk manual. See gawkextlib's documentation for building
    and installation.
    --- Synchronet 3.22a-Linux NewsLink 1.2