• Re: endian history

    From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sat Sep 5 08:17:34 2026
    From Newsgroup: comp.arch

    On Wed, 3 Jun 2026 15:33:53 -0000 (UTC), quadi wrote:

    On Wed, 03 Jun 2026 13:54:01 +0000, Thomas Koenig wrote:

    It causes problems with badly-written software.

    I don't see that as a fault of big-endian.

    Think of the 3 important numberings of memory elements in a machine architecture:

    1. Ordering of numbering of bytes within a multibyte quantity
    2. Ordering of numbering of bits within a byte
    3. Ordering of place values for binary digits within an integer (of
    whatever length)

    Only little-endian can keep all 3 consistent. With big-endian, you
    inevitably end up with situations like the ordering 3 being 7 minus
    ordering 2, or some even more complicated inversion. If you try to
    keep 2 and 3 consistent, then you lose a simple relationship between
    bit and byte numbers, like the byte number being the bit number
    right-shifted by 3.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 5 13:52:32 2026
    From Newsgroup: comp.arch

    On Sat, 5 Sep 2026 08:17:34 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Think of the 3 important numberings of memory elements in a machine >architecture:

    1. Ordering of numbering of bytes within a multibyte quantity
    2. Ordering of numbering of bits within a byte
    3. Ordering of place values for binary digits within an integer (of
    whatever length)

    Only little-endian can keep all 3 consistent. With big-endian, you
    inevitably end up with situations like the ordering 3 being 7 minus
    ordering 2, or some even more complicated inversion. If you try to
    keep 2 and 3 consistent, then you lose a simple relationship between
    bit and byte numbers, like the byte number being the bit number
    right-shifted by 3.

    Keeping 1 and 2 consistent is important.

    3, of course, can only be little-endian. But consistency with 3 is
    only a stylistic preference; one might think it nice that the bit
    representing 2^0 be bit 0, but it's not important for any actual
    purpose. It isn't in any way involved in actual computer arithmetic,
    and so it doesn't save any cycles.

    My argument for big-endian is that only big-endian can keep 1, 2, and
    4 consistent, where 4 is:

    4. Ordering of numbering of digits within a number represented in
    character form.

    Then you can have 1, 2, 4, and 5 consistent... where 5 is

    5. Ordering of numbering of digits within a packed decimal quantity.

    Then:

    A number in character form is written in the same direction as a
    number in packed decimal form, so they can be easily converted.

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    Plus, now numbers are written in the computers memory, laid out in
    locations of increasing addresses, the same way we would write them on
    a piece of paper, leading to less confusion in writing programs.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Sat Sep 5 15:58:38 2026
    From Newsgroup: comp.arch

    John Savard wrote:
    On Sat, 5 Sep 2026 08:17:34 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Think of the 3 important numberings of memory elements in a machine
    architecture:

    1. Ordering of numbering of bytes within a multibyte quantity
    2. Ordering of numbering of bits within a byte
    3. Ordering of place values for binary digits within an integer (of
    whatever length)

    Only little-endian can keep all 3 consistent. With big-endian, you
    inevitably end up with situations like the ordering 3 being 7 minus
    ordering 2, or some even more complicated inversion. If you try to
    keep 2 and 3 consistent, then you lose a simple relationship between
    bit and byte numbers, like the byte number being the bit number
    right-shifted by 3.

    Keeping 1 and 2 consistent is important.

    3, of course, can only be little-endian. But consistency with 3 is
    only a stylistic preference; one might think it nice that the bit representing 2^0 be bit 0, but it's not important for any actual
    purpose. It isn't in any way involved in actual computer arithmetic,
    and so it doesn't save any cycles.

    My argument for big-endian is that only big-endian can keep 1, 2, and
    4 consistent, where 4 is:

    4. Ordering of numbering of digits within a number represented in
    character form.

    Then you can have 1, 2, 4, and 5 consistent... where 5 is

    5. Ordering of numbering of digits within a packed decimal quantity.

    Then:

    A number in character form is written in the same direction as a
    number in packed decimal form, so they can be easily converted.

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    Plus, now numbers are written in the computers memory, laid out in
    locations of increasing addresses, the same way we would write them on
    a piece of paper, leading to less confusion in writing programs.


    Please just forget about it.

    Big endian is just as dead now as bytes that are not 8-bits wide.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 5 14:55:23 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    I do not have access to any machine where bytes are not eight bits :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Sat Sep 5 17:24:53 2026
    From Newsgroup: comp.arch

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new architecture, please just forget about this particular issue.

    I do not have access to any machine where bytes are not eight bits :-)

    My very first personal piece of programming was on a Unisys 1100 with
    36-bit words and either 6 or 9-bit bytes.

    The program was my own arbitrary precision library which I used to
    calculate as many digits of pi as I could manage inside a minute (the
    student max) of CPU runtime.

    I did not know that I had 36 bits, only that the docs said there was
    room for 10-11 decimal digits, so I wrote my code with lots and lots of
    modulo operations, over 10-digit chunks and temporary 20-digit
    multiplication results.

    I'm guessing I used the same 72-bit temp variables instead of carries,
    or the classic "is the sum smaller than either input, then it must have overflowed" approach.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 5 16:57:35 2026
    From Newsgroup: comp.arch


    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer architecture
    shall be LE. That fact that one can make a coherent BE machine is what
    should be forgotten.


    I do not have access to any machine where bytes are not eight bits :-)

    My very first personal piece of programming was on a Unisys 1100 with
    36-bit words and either 6 or 9-bit bytes.

    The program was my own arbitrary precision library which I used to
    calculate as many digits of pi as I could manage inside a minute (the student max) of CPU runtime.

    We only got 30 seconds...lucky you--but in WATFIV if you could put
    your calculation inside a WRITE statement, you got unlimited time.
    I calculated and printed 1!..to..200! and it took several minutes
    on the 360/67.

    I did not know that I had 36 bits, only that the docs said there was
    room for 10-11 decimal digits, so I wrote my code with lots and lots of modulo operations, over 10-digit chunks and temporary 20-digit multiplication results.

    I'm guessing I used the same 72-bit temp variables instead of carries,
    or the classic "is the sum smaller than either input, then it must have overflowed" approach.

    Terje

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sat Sep 5 21:16:17 2026
    From Newsgroup: comp.arch

    On Sat, 5 Sep 2026 14:55:23 -0000 (UTC), Thomas Koenig wrote:

    I have login to several machines which are big-endian (and we
    receive bug reports about them for gfortran).

    Fun fact: even on big-endian architectures, registers behave as though
    they were little-endian.

    Add that to the list of inconsistencies that only little-endian
    architectures can resolve. ;)

    I do not have access to any machine where bytes are not eight bits
    :-)

    Remember, the original meaning of rCLbyterCY (from the STRETCH project, wasnrCOt it) was not rCLminimum directly-addressable memory unitrCY but rCLvariable-length bitfieldrCY.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 6 01:55:24 2026
    From Newsgroup: comp.arch

    On Sat, 5 Sep 2026 21:16:17 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Fun fact: even on big-endian architectures, registers behave as though
    they were little-endian.

    I didn't know that registers behaved as if they had *any* endianness
    at all. Unless you're talking about registers for short vectors - like
    MMX, AltiVec, AVX - the issue of endianness doesn't even arise.

    Although there's a *sort* of endianness involving registers, but
    integer registers and floating-point registers do it the opposite way.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 6 02:06:36 2026
    From Newsgroup: comp.arch

    On Sat, 05 Sep 2026 16:57:35 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    Do not forget:: remember that LE won and any new computer architecture
    shall be LE. That fact that one can make a coherent BE machine is what
    should be forgotten.

    I don't see how it would even be possible *to* forget "the fact that
    one can make a consistent big-endian machine". In order not to realize
    this, one would have to become ignorant of such large chunks of the
    basic principles of designing computers from digital logic... that no
    one would be able to build any kind of a computer at all!

    At least that was my immediate reaction. But then I remembered that
    before the PDP-11 came along, people were building lots of computers
    just fine without even thinking of the possibility of a consistent little-endian machine. So I guess it is possible.

    But I don't think it's at all likely, unless all of us switch to
    speaking Arabic.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 02:09:32 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 01:55:24 GMT, John Savard wrote:

    On Sat, 5 Sep 2026 21:16:17 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    Fun fact: even on big-endian architectures, registers behave as
    though they were little-endian.

    I didn't know that registers behaved as if they had *any* endianness
    at all.

    It is quite common to have instructions that operate on partial
    register contents, is it not?

    Consider a sequence like this (for rCLmoverCY, read rCLloadrCY or rCLstorerCY as
    appropriate, and for rCLwordrCY read rCLsome convenient multi-byte quantityrCY):

    move.word A, B
    move.byte B, C

    On typical modern (i.e. RISC) architectures, A and C might be memory
    locations and B is a register, or vice versa.

    The question is, which byte of B ends up in C -- is it the
    most-significant or least-significant byte?

    On a little-endian architecture, itrCOs always the least-significant byte.

    On a big-endian architecture, it depends on whether B is a register or
    memory -- least-significant for register, most-significant for memory.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 07:04:02 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 02:06:36 GMT, John Savard wrote:

    At least that was my immediate reaction. But then I remembered that
    before the PDP-11 came along, people were building lots of computers
    just fine without even thinking of the possibility of a consistent little-endian machine. So I guess it is possible.

    Endianness only matters on byte-addressable machines, which were still
    fairly unusual at the time the PDP-11 came along. The idea was really popularized by microprocessor-based machines.

    But I don't think it's at all likely, unless all of us switch to
    speaking Arabic.

    Endianness has nothing to do with writing direction.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 07:07:08 2026
    From Newsgroup: comp.arch

    On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    But that limits the precision with which you can do basic addition and subtraction of decimal numbers, unless you go backwards in memory.

    In what order did the IBM 1401 and 1620 order the digits? Because I
    believe they could do arbitrary-precision addition and subtraction.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Sun Sep 6 13:57:42 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer architecture
    shall be LE. That fact that one can make a coherent BE machine is what
    should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    I don't consider that sufficient reason for a CPU to implement
    native arithmetic using big-endian containers, however. While
    ARMv8 included the capability to provide big-endian support,
    in the architecture, it hasn't been widely implemented and may
    in the future be eliminated.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Sun Sep 6 14:00:27 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    But that limits the precision with which you can do basic addition and >subtraction of decimal numbers, unless you go backwards in memory.

    In what order did the IBM 1401 and 1620 order the digits? Because I
    believe they could do arbitrary-precision addition and subtraction.

    The B3500 BCD digits were numbered with the highest magnitude digit
    having the lowest ordered address. As were earlier Burroughs BCD systems.

    The bits within the digit were ordered 8,4,2 and 1.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 15:39:03 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Sat, 5 Sep 2026 21:16:17 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Fun fact: even on big-endian architectures, registers behave as though
    they were little-endian.

    I didn't know that registers behaved as if they had *any* endianness
    at all. Unless you're talking about registers for short vectors - like
    MMX, AltiVec, AVX - the issue of endianness doesn't even arise.

    He is talking about how a byte is loaded/stored from the register field
    of least significance, how a halfword is loaded/stored from the lesser significance, how words are loaded/stored at low significance.

    Although there's a *sort* of endianness involving registers, but
    integer registers and floating-point registers do it the opposite way.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sun Sep 6 12:41:36 2026
    From Newsgroup: comp.arch

    On 9/6/2026 8:57 AM, Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer architecture
    shall be LE. That fact that one can make a coherent BE machine is what
    should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    I don't consider that sufficient reason for a CPU to implement
    native arithmetic using big-endian containers, however. While
    ARMv8 included the capability to provide big-endian support,
    in the architecture, it hasn't been widely implemented and may
    in the future be eliminated.


    Can usually be addressed well enough if the ISA has byte swap
    instructions. It is a little bit of a pain if it has to be done with
    shifts and masks, but in this case a byte swap does well enough.

    Could debate though whether to provide all of the byte swap cases (say,
    5 unique instruction), or simply an 8 byte swap to be followed by a
    right shift as-needed.

    ...



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 6 18:28:09 2026
    From Newsgroup: comp.arch

    On Sun, 6 Sep 2026 02:09:32 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    It is quite common to have instructions that operate on partial
    register contents, is it not?

    Consider a sequence like this (for rCLmoverCY, read rCLloadrCY or rCLstorerCY as
    appropriate, and for rCLwordrCY read rCLsome convenient multi-byte >quantityrCY):

    move.word A, B
    move.byte B, C

    On typical modern (i.e. RISC) architectures, A and C might be memory >locations and B is a register, or vice versa.

    The question is, which byte of B ends up in C -- is it the
    most-significant or least-significant byte?

    On a little-endian architecture, itrCOs always the least-significant byte.

    Oh. I don't think of registers like that.

    One has...

    load long
    load
    load halfword
    load byte

    and except for load long, yes, they put the memory item in the least significant part of the register - but they're also sign-extending it.
    So the register is a dimensionless point, containing a single number,
    not something with parts.

    You would need the special INSERT instruction to put a short integer
    into the least significant portion of a general register without
    affecting the more significant parts. And there is no way of directly addressing these more significant parts separately, which makes the
    concept of endianness not applicable to registers.

    The floating-point instructions sort of go the opposite way. But in
    the new IEEE 754 floating-point, since the exponent field is a
    different width for the different sizes of float, things aren't as
    simple as with the older hexadecimal floating-point of the 360.

    On a big-endian architecture, it depends on whether B is a register or
    memory -- least-significant for register, most-significant for memory.

    While that's true, registers aren't memory. A register is an available
    point for calculating with numbers. Memory is where numbers are
    written.

    So, although your point is valid in a way, the way in which the
    assembly languages of at least some big-endian machines are designed
    tends to obscure the issue; there isn't a single "move" instruction
    which suddenly changes to little-endian semantics if the destination
    is a register.

    One has load to go from memory to a register, and store to go from a
    register to memory. And then there's move which only ever moves from
    memory to memory.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 6 18:38:54 2026
    From Newsgroup: comp.arch

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly. But as it happens, the digits we use to write numbers in
    decimal notation are often referred to as "Arabic numerals". So it
    should not be surprising to learn that the Arabs also write numbers
    using decimal place-value notation. We changed the shapes of the
    digits we got from them, though.

    But while the Arabs write their language from right to left, when
    writing a number, they put the most significant digit on the left, and
    the units digit on the right. Just like we do. But now the least
    significant digit comes _first_ in the ordinary direction of reading.

    So the layout of numbers in memory in a little-endian computer seems
    natural and intuitive to someone whose language is Arabic - and
    strange and confusing to someone whose language is English.

    Before IBM came up with the consistently big-endian IBM System/360,
    their biggest-selling computer was the IBM 1401.

    This computer was designed for commercial data processing, not
    scientific work. Thus, it did all its arithmetic in decimal.

    Not only that, because it stored data in six-bit characters, with one
    extra bit for a word mark, it didn't even pack the decimal before
    doing arithmetic. It just did arithmetic on character strings.

    So the manual for the 1401 had illustrations of how a payroll record
    might look on punched cards or in a computer's memory. Someone's name, someone's employee number, someone's monthly or weekly salary. All
    written in the same direction.

    On a 360, this is only changed slightly. The numbers get squashed,
    with two digits in every eight-bit character cell (byte).

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 21:50:15 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 18:38:54 GMT, John Savard wrote:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Not at all.

    But while the Arabs write their language from right to left, when
    writing a number, they put the most significant digit on the left,
    and the units digit on the right. Just like we do. But now the least significant digit comes _first_ in the ordinary direction of
    reading.

    So the layout of numbers in memory in a little-endian computer seems
    natural and intuitive to someone whose language is Arabic - and
    strange and confusing to someone whose language is English.

    The usual convention is to store text strings in memory/files etc in
    reading order, not rendering order. This is a key aspect of the
    Unicode bidirectional layout algorithm
    <http://www.unicode.org/reports/tr9/>.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 21:53:00 2026
    From Newsgroup: comp.arch

    On Sun, 6 Sep 2026 12:41:36 -0500, BGB wrote:

    Can usually be addressed well enough if the ISA has byte swap
    instructions. It is a little bit of a pain if it has to be done with
    shifts and masks, but in this case a byte swap does well enough.

    Could debate though whether to provide all of the byte swap cases
    (say, 5 unique instruction), or simply an 8 byte swap to be followed
    by a right shift as-needed.

    ItrCOs only a small step from there to full-on rCLswizzlingrCY, which is a common part of SIMD instruction sets, isnrCOt it.

    Is it true that x86 instructions containing integer literals put them
    in big-endian order?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 23:57:39 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    But as it happens, the digits we use to write numbers in
    decimal notation are often referred to as "Arabic numerals". So it
    should not be surprising to learn that the Arabs also write numbers
    using decimal place-value notation. We changed the shapes of the
    digits we got from them, though.

    But while the Arabs write their language from right to left, when
    writing a number, they put the most significant digit on the left, and
    the units digit on the right. Just like we do. But now the least
    significant digit comes _first_ in the ordinary direction of reading.

    So the layout of numbers in memory in a little-endian computer seems
    natural and intuitive to someone whose language is Arabic - and
    strange and confusing to someone whose language is English.

    Before IBM came up with the consistently big-endian IBM System/360,
    their biggest-selling computer was the IBM 1401.

    This computer was designed for commercial data processing, not
    scientific work. Thus, it did all its arithmetic in decimal.

    Not only that, because it stored data in six-bit characters, with one
    extra bit for a word mark, it didn't even pack the decimal before
    doing arithmetic. It just did arithmetic on character strings.

    So the manual for the 1401 had illustrations of how a payroll record
    might look on punched cards or in a computer's memory. Someone's name, someone's employee number, someone's monthly or weekly salary. All
    written in the same direction.

    On a 360, this is only changed slightly. The numbers get squashed,
    with two digits in every eight-bit character cell (byte).

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 23:54:40 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:
    -----------------
    One has...

    load long
    load
    load halfword
    load byte

    Load Double
    Load Unsigned Word
    load Signed Word
    Load Unsigned Half
    Load Signed Half
    Load Unsigned byte
    Load Signed Byte
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 00:12:18 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Sun, 6 Sep 2026 12:41:36 -0500, BGB wrote:

    Can usually be addressed well enough if the ISA has byte swap
    instructions. It is a little bit of a pain if it has to be done with
    shifts and masks, but in this case a byte swap does well enough.

    Could debate though whether to provide all of the byte swap cases
    (say, 5 unique instruction), or simply an 8 byte swap to be followed
    by a right shift as-needed.

    ItrCOs only a small step from there to full-on rCLswizzlingrCY, which is a common part of SIMD instruction sets, isnrCOt it.

    Permute can rearrange all 8-bytes of a register in any order desired.

    Is it true that x86 instructions containing integer literals put them
    in big-endian order?

    Since 8088 was a byte sized buss and read things in use order, I would bet against that.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Sun Sep 6 17:31:43 2026
    From Newsgroup: comp.arch

    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to use
    long scrolls. :-)

    But ISTM that the choice of horizontal direction is independent of the
    choice of vertical direction, thus yielding four potential choices.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Mon Sep 7 02:35:05 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Another one: rCLboustrophedonicrCY -- rightraAleft and leftraAright on alternate lines.

    Used for example on Easter Island/Rapa Nui. (Not that anybody knows
    how to read it ...)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Mon Sep 7 02:44:26 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 18:28:09 GMT, John Savard wrote:

    While that's true, registers aren't memory. A register is an
    available point for calculating with numbers. Memory is where
    numbers are written.

    ThatrCOs not really a meaningful distinction. Calculations happen inside function units in the CPU chip, so architectural registers are just as
    much rCLwhere numbers are writtenrCY -- they are on-chip memory, separate
    from the function units, just with more limited forms of addressing
    available (for the sake of speed), compared to off-chip memory.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sun Sep 6 23:34:09 2026
    From Newsgroup: comp.arch

    On 9/6/2026 7:12 PM, MitchAlsup wrote:

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Sun, 6 Sep 2026 12:41:36 -0500, BGB wrote:

    Can usually be addressed well enough if the ISA has byte swap
    instructions. It is a little bit of a pain if it has to be done with
    shifts and masks, but in this case a byte swap does well enough.

    Could debate though whether to provide all of the byte swap cases
    (say, 5 unique instruction), or simply an 8 byte swap to be followed
    by a right shift as-needed.

    ItrCOs only a small step from there to full-on rCLswizzlingrCY, which is a >> common part of SIMD instruction sets, isnrCOt it.

    Permute can rearrange all 8-bytes of a register in any order desired.


    In my case, I lack such an instruction...

    Can note that there are 4 element shuffles for both 8 and 16 bit
    elements, and it is possible to combine them to get an arbitrary 8-byte shuffle; though the harder part is the lack of a computationally
    efficient way to map an arbitrary 8-element shuffle mask onto a minimal sequence of 4 element shuffles (it us generally possible to so so within
    ~ 3 instructions; just not with any simple algorithm that I am aware of).

    It is possible to build a lookup table for this, but this lookup table
    is quite large (I didn't usually distribute it with my compiler due to
    its immense size, but annoyingly without it there is seemingly no good
    way to do 8 element shuffles).


    However, the most common case of these (byte order reversal) can be
    expressed as a one-off 2R instruction...


    Is it true that x86 instructions containing integer literals put them
    in big-endian order?

    Since 8088 was a byte sized buss and read things in use order, I would bet against that.

    Yeah, x86 is pretty solidly little endian in these areas...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Sun Sep 6 21:46:44 2026
    From Newsgroup: comp.arch

    On 9/6/2026 5:31 PM, Stephen Fuld wrote:
    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to use long scrolls.-a :-)

    But ISTM that the choice of horizontal direction is independent of the choice of vertical direction, thus yielding four potential choices.

    Sorry for the self followup, but actually eight choices. For each of
    the four corners, proceed in the horizontal direction or the vertical direction.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 7 05:05:19 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    There is also https://en.wikipedia.org/wiki/Boustrophedon , writing
    "as the ox plows" (alternating left to right and right to left).
    Good if you're reading a monumental instription and do not want
    to walk back to the start of the next line.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Johann 'Myrkraverk' Oskarsson@johann@myrkraverk.invalid to comp.arch,alt.history.ancient-egypt on Mon Sep 7 13:57:26 2026
    From Newsgroup: comp.arch

    On 07/09/2026 12:46 PM, Stephen Fuld wrote:
    On 9/6/2026 5:31 PM, Stephen Fuld wrote:
    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to
    use long scrolls.-a :-)

    But ISTM that the choice of horizontal direction is independent of the
    choice of vertical direction, thus yielding four potential choices.

    Sorry for the self followup, but actually eight choices.-a For each of
    the four corners, proceed in the horizontal direction or the vertical direction.



    Dear Stephen,

    It seems you have forgotten your hieroglyphs. Please direct your atten-
    tion to Gardiner's /Egyptian Grammar/, section 16, page 25. It has a
    wonderful section on /direction of writing/. I'm sure you still have
    your copy from the days of juniour high, when it's natural for children
    to choose Egyptian Hieroglyphs for extra credits.

    As the Egyptians went by the beaks of the birds, or rather, all the
    mouths point in the direction against the reading, it's quite possible
    to switch direction mid line, for stylistic purposes, and you'll see
    this in quite a lot of temple inscriptions.

    As this has absolutely nothing at all to do with big vs. little endian,
    I have set the follow-up to alt.history.ancient-egypt; after all, the
    Egyptians did not have our modern digital computers, and instead made
    do with -- at best -- variations of the Antikythera mechanism for their automagic calculations. One has to wonder if they really counted the
    pyramids on their fingers, or if they used abacuses, or even some
    stranger counting systems?


    Best wishes, and happy studies of hieroglyphs.
    --
    Johann | email: invalid -> com | http://www.myrkraverk.com/blog/
    I'm not from the Internet, I just work there. | via Easynews.com https://bsky.app/profile/myrkraverk.bsky.social | for ( ;; ) _:;
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 7 10:32:34 2026
    From Newsgroup: comp.arch

    On 07/09/2026 01:57, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??


    (Historical note - They used several more orders, though others have
    died out except for niche uses (like decorative texts). For example, a
    number of ancient texts were written (or carved) with alternate lines in alternating directions - sometimes with letters also reversed, sometimes
    not).

    But as it happens, the digits we use to write numbers in
    decimal notation are often referred to as "Arabic numerals". So it
    should not be surprising to learn that the Arabs also write numbers
    using decimal place-value notation. We changed the shapes of the
    digits we got from them, though.

    The common shapes of the digits changed over time, both in the Arabic
    world and in Europe, from the shapes developed by Indian mathematicians
    from the 3rd century BC, through Arabic mathematicians in about the
    600's (the Indian addition of a zero as a place value symbol lead to a
    big jump in popularity), and gradually into Europe over the next few centuries. Modern European-style digits are no more dissimilar to the
    "Hindi" numerals first used in the Arab Caliphate than those used in
    modern written Arabic. Standardisation was heavily influenced by the
    printing press.


    But while the Arabs write their language from right to left, when
    writing a number, they put the most significant digit on the left, and
    the units digit on the right. Just like we do. But now the least
    significant digit comes _first_ in the ordinary direction of reading.

    So the layout of numbers in memory in a little-endian computer seems
    natural and intuitive to someone whose language is Arabic - and
    strange and confusing to someone whose language is English.


    I would disagree entirely. Numbers have always been read and written in
    a variety of ways, and this is not a direct match for word writing
    order. In many current languages, "123" is read as the equivalent of
    "one hundred three and twenty". "1024" may be read as "a thousand and
    twenty four", while "1066" may be read as "ten sixty six", or perhaps
    "one zero double six" if it is part of a phone number. For numbers in general, you do not read left-to-right or right-to-left - you look at
    the entire number, and use its value. Only then do you know that a
    number starting with the digit "1" is thirteen or a million.

    And once you start doing arithmetic, you do most of it in little-endian
    order - except division, which will be in big-endian order.

    Your theory here initially sounds like it makes sense, but it does not
    stand up to closer scrutiny.


    In the computing world, there is maybe a bit more clarity - certainly
    the choice of big or little endian can make certain operations simpler,
    while disadvantaging other operations. The same applies to the choice
    of binary or decimal encoding.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Mon Sep 7 13:56:05 2026
    From Newsgroup: comp.arch

    Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits wide.

    I have login to several machines which are big-endian (and we receive
    bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer architecture
    shall be LE. That fact that one can make a coherent BE machine is what
    should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    These protocols are a sufficient reason to provide BSWAP type capability
    in all CPU architectures. (Including BE ones since there are other LE protocols/encodings).

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 18:14:10 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Another one: rCLboustrophedonicrCY -- rightraAleft and leftraAright on alternate lines.

    Used for example on Easter Island/Rapa Nui. (Not that anybody knows
    how to read it ...)

    {Spok mode= On} Fascinating
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.arch on Mon Sep 7 20:24:26 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
    On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    But that limits the precision with which you can do basic addition and subtraction of decimal numbers, unless you go backwards in memory.

    In what order did the IBM 1401 and 1620 order the digits? Because I
    believe they could do arbitrary-precision addition and subtraction.

    On IBM 1401 least significant digit was at highest address. Arithmetic hardware was walking towards lower addresses.

    Good way to think about 1401 is to consider it as an electronic
    accounting machine. Smallest 1401 (with memory size 1400 charactes)
    was probably able to do all what contemporary accounting machines
    could do in a similar way. Early 1401 programmers frequently worked
    previously with accounting machines and could build on their past
    experience. IBM salesmen could sell 1401 as a replacement
    for accounting machines.

    So 1401 was a point in evolutinary path from accounting machines to
    modern computers. 1401 started with natural decimal memory
    addresses. But need for sligtly bigger memory lead to decimal
    based but rather unnatural scheme. IIUC slightly later
    Honeywell 200 had very similar logical structure, but switched
    to binary addresses. IBM 360 replaced character representation
    by BCD. 8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Mon Sep 7 23:54:31 2026
    From Newsgroup: comp.arch

    On Mon, 7 Sep 2026 13:56:05 +0200
    Terje Mathisen <terje.mathisen@tmsw.no> wrote:

    Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits
    wide.

    I have login to several machines which are big-endian (and we
    receive bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer
    architecture shall be LE. That fact that one can make a coherent
    BE machine is what should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    These protocols are a sufficient reason to provide BSWAP type
    capability in all CPU architectures. (Including BE ones since there
    are other LE protocols/encodings).

    Terje


    One 2-byte BE field in IP header. Another one 2-byte field in UDP
    header. The rest is either non-endian (single-byte or less) or
    endian-neutral (IP address, ports, checksums). I didn't look at TCP
    header but would think that it's approximately the same.
    It does not sound as sufficient reason to have BSWAP.




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Mon Sep 7 15:23:31 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/6/2026 5:31 PM, Stephen Fuld wrote:

    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to
    use long scrolls. :-)

    But ISTM that the choice of horizontal direction is independent of
    the choice of vertical direction, thus yielding four potential
    choices.

    Sorry for the self followup, but actually eight choices. For each of
    the four corners, proceed in the horizontal direction or the vertical direction.

    How about, for each of the four corners, proceed in a clockwise or anti-clockwise spiral? That's eight more choices.

    Let's not forget residents of Lilliput in Gulliver's Travels, who
    wrote diagonally.

    And what do we do about people who want to write on Mobius strips?
    Where can they start when there are no corners? Pursuing this idea
    could help protect against overflows, where every memory location
    holds two 8-bit quantities, and wrap-around writes to "the other
    side" of a byte rather than overwriting low memory.

    I suppose some sort of apology is in order, but trying to write one
    on a Mobius strip I'm at a loss as to where to begin. ;)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Mon Sep 7 15:25:39 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro <ldo@nz.invalid> writes:

    Endianness only matters on byte-addressable machines, which were
    still fairly unusual at the time the PDP-11 came along.

    False.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 7 16:58:47 2026
    From Newsgroup: comp.arch

    On 9/7/2026 1:24 PM, Waldek Hebisch wrote:
    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
    On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:

    A number in packed decimal form is oriented in the same way as an
    integer in binary form, so they can share registers and ALUs.

    But that limits the precision with which you can do basic addition and
    subtraction of decimal numbers, unless you go backwards in memory.

    In what order did the IBM 1401 and 1620 order the digits? Because I
    believe they could do arbitrary-precision addition and subtraction.

    On IBM 1401 least significant digit was at highest address. Arithmetic hardware was walking towards lower addresses.

    Good way to think about 1401 is to consider it as an electronic
    accounting machine. Smallest 1401 (with memory size 1400 charactes)
    was probably able to do all what contemporary accounting machines
    could do in a similar way. Early 1401 programmers frequently worked previously with accounting machines and could build on their past
    experience. IBM salesmen could sell 1401 as a replacement
    for accounting machines.

    So 1401 was a point in evolutinary path from accounting machines to
    modern computers. 1401 started with natural decimal memory
    addresses. But need for sligtly bigger memory lead to decimal
    based but rather unnatural scheme. IIUC slightly later
    Honeywell 200 had very similar logical structure, but switched
    to binary addresses. IBM 360 replaced character representation
    by BCD.

    Slight correction. S/360 implemented "extended" BCD. BCD is a six bit
    per character code, EBCDIC (the IC stands for "interchange code") is an
    eight bit per character code.


    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.


    And, of course, S/360 had an instruction to "pack" a string of numeric
    EBCDIC characters to two, four bit, digits per byte, an unpack
    instruction to convert back, and a set of arithmetic instructions that
    worked on packed data. These are still all supported on Z series.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 8 02:33:17 2026
    From Newsgroup: comp.arch

    On Sun, 6 Sep 2026 21:50:15 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    The usual convention is to store text strings in memory/files etc in
    reading order, not rendering order. This is a key aspect of the
    Unicode bidirectional layout algorithm
    <http://www.unicode.org/reports/tr9/>.

    Yes. So the ordering of Arabic letters in memory will be the same as
    the ordering of Latin letters in memory. However, since in Arab texts,
    the reading order of Arabic numerals is in the same physical direction
    as the reading order of the text, in a computer designed from the
    ground up by people who may not even know that languages other than
    Arabic *exist*...

    Substitute "English" for "Arabic", and look at the United States to be reassured that this is possible...

    now character strings representing numbers will be stored with the
    least significant digit in the lowest memory address.

    In little-endian form.

    Numbers are a normal native element of Arabic text. So while there
    would be a fancy switchover of directions if an English text were
    quoted in an Arabic document, the same way as if a Hebrew text were
    quoted in an English document, what do you mean that Arabs have to do
    the same switchover every time a number appears in a document, as if
    Arabic numerals were foreign to Arabic?

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 8 02:36:55 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    I should think the answer is obvious. Left and right are fundamentally equivalent; they may be opposites, but there's no reason to prefer one
    over the other. (Star Trek reference: "Let That Be Your Last
    Battlefield".)

    Top, however, is clearly and obviously and intrinsically a beginning
    spot for writing.

    Thus, languages written from top to bottom may have successive lines
    going from right to left (as in traditional Chinese) or possibly from
    left to right... but never will you find a language written either
    from right to left or left to right where the first line is at the
    bottom of the page.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 8 02:45:38 2026
    From Newsgroup: comp.arch

    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    And they're certainly normal machines. They're the ones you need to
    run regular Windows programs.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 8 02:47:29 2026
    From Newsgroup: comp.arch

    On Mon, 7 Sep 2026 16:58:47 -0700, Stephen Fuld
    <sfuld@alumni.cmu.edu.invalid> wrote:

    And, of course, S/360 had an instruction to "pack" a string of numeric >EBCDIC characters to two, four bit, digits per byte, an unpack
    instruction to convert back, and a set of arithmetic instructions that >worked on packed data. These are still all supported on Z series.

    I'm willing to accept that he doesn't consider z/Architecture
    mainframes to meet his definition of 'normal machines', since although
    they are still popular in their niche behind the scenes, they're not
    ubiquitous in people's homes.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 7 20:15:57 2026
    From Newsgroup: comp.arch

    On 9/7/2026 3:23 PM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/6/2026 5:31 PM, Stephen Fuld wrote:

    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to
    use long scrolls. :-)

    But ISTM that the choice of horizontal direction is independent of
    the choice of vertical direction, thus yielding four potential
    choices.

    Sorry for the self followup, but actually eight choices. For each of
    the four corners, proceed in the horizontal direction or the vertical
    direction.

    How about, for each of the four corners, proceed in a clockwise or anti-clockwise spiral? That's eight more choices.

    Now we are really going down the rabbit hole. There is an issue with
    "spiral" directions. Let's say you start at the upper left corner, then proceed to the right. When you hit the right edge, you are going to go "down". But what is the orientation of the letters, i.e. should you
    turn the paper 90 degrees to keep reading left to right, or should the orientation of the letters change so you can now read top to bottom
    without reorientating the paper? Both have problems.

    If you say you will change the orientation of the paper, what if the
    "paper" is actually a bronze plaque mounted on a monument. Kinda hard
    to change its orientation. And when you get to the bottom, do you stand
    on your head to read it? :-)

    On the other hand, say you change the orientation of the letters. That
    works for the right side - you just read it top to bottom. But now what
    do you do when you get to the bottom? Do you print the letters upside
    down, in which case you have the same problem as the above solution.
    You could keep the letters "right side up:, but now the words are
    spelled backwards. That's awkward.

    So I reject the spiral solutions, and other, even more bizarre ones.
    But still, eight possibilities is more than enough. :-)
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 7 20:21:11 2026
    From Newsgroup: comp.arch

    On 9/7/2026 7:36 PM, John Savard wrote:
    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    I should think the answer is obvious. Left and right are fundamentally equivalent; they may be opposites, but there's no reason to prefer one
    over the other.

    Perhaps not quite true. Since the vast majority of people are right
    handed, it is easier for them to write left to right, as it means you
    don't run the risk of smudging what you have already wrote as left
    handers do with left to right order (think old, pre-ball point pens.).
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 7 20:23:53 2026
    From Newsgroup: comp.arch

    On 9/7/2026 7:47 PM, John Savard wrote:
    On Mon, 7 Sep 2026 16:58:47 -0700, Stephen Fuld <sfuld@alumni.cmu.edu.invalid> wrote:

    And, of course, S/360 had an instruction to "pack" a string of numeric
    EBCDIC characters to two, four bit, digits per byte, an unpack
    instruction to convert back, and a set of arithmetic instructions that
    worked on packed data. These are still all supported on Z series.

    I'm willing to accept that he doesn't consider z/Architecture
    mainframes to meet his definition of 'normal machines', since although
    they are still popular in their niche behind the scenes, they're not ubiquitous in people's homes.

    Well, he brought S/360 up. I was just correcting a minor error.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Tue Sep 8 03:44:36 2026
    From Newsgroup: comp.arch

    On Mon, 7 Sep 2026 20:15:57 -0700, Stephen Fuld wrote:

    Now we are really going down the rabbit hole. There is an issue with
    "spiral" directions. Let's say you start at the upper left corner,
    then proceed to the right. When you hit the right edge, you are
    going to go "down". But what is the orientation of the letters, i.e.
    should you turn the paper 90 degrees to keep reading left to right,
    or should the orientation of the letters change so you can now read
    top to bottom without reorientating the paper? Both have problems.

    Been done, though. Look up the famous rCLPhaistos DiscrCY ...
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Tue Sep 8 03:46:48 2026
    From Newsgroup: comp.arch

    On Tue, 08 Sep 2026 02:33:17 GMT, John Savard wrote:

    Numbers are a normal native element of Arabic text. So while there
    would be a fancy switchover of directions if an English text were
    quoted in an Arabic document, the same way as if a Hebrew text were
    quoted in an English document, what do you mean that Arabs have to
    do the same switchover every time a number appears in a document, as
    if Arabic numerals were foreign to Arabic?

    Yup. ThatrCOs why a script like Arabic is called rCLbidirectionalrCY, not rCLright-to-leftrCY.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Tue Sep 8 03:49:21 2026
    From Newsgroup: comp.arch

    On Tue, 08 Sep 2026 02:47:29 GMT, John Savard wrote:

    I'm willing to accept that he doesn't consider z/Architecture
    mainframes to meet his definition of 'normal machines', since
    although they are still popular in their niche behind the scenes,
    they're not ubiquitous in people's homes.

    They are easy enough to instantiate in virtual form in the homes of
    ordinary people (like you or me), by way of the Hercules emulator.

    <https://packages.debian.org/trixie/hercules>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 8 05:07:38 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> schrieb:
    On Mon, 7 Sep 2026 13:56:05 +0200
    Terje Mathisen <terje.mathisen@tmsw.no> wrote:

    Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits
    wide.

    I have login to several machines which are big-endian (and we
    receive bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer
    architecture shall be LE. That fact that one can make a coherent
    BE machine is what should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    These protocols are a sufficient reason to provide BSWAP type
    capability in all CPU architectures. (Including BE ones since there
    are other LE protocols/encodings).

    Terje


    One 2-byte BE field in IP header. Another one 2-byte field in UDP
    header. The rest is either non-endian (single-byte or less) or
    endian-neutral (IP address, ports, checksums). I didn't look at TCP
    header but would think that it's approximately the same.

    Takes less then half a minute to look at the Wikipedia page...

    Source port, destination port (ok, these can be pre-computed).
    window are 16-bit. Sequence number and acknowledgement
    numbe are 32-bit. Checksum (the most compute-intensive task)
    adds up header and data in 16-bit one's complement, but that
    may be done by specialized haredware.

    It does not sound as sufficient reason to have BSWAP.

    It does, but only if you want performance in networking.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Tue Sep 8 08:38:29 2026
    From Newsgroup: comp.arch

    According to John Savard <quadibloc@invalid.com>:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    x86-64 has 32 bit and 64 bit modes. AAA is still there in 32 bit mode
    but not in 64 bit mode.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Tue Sep 8 11:32:29 2026
    From Newsgroup: comp.arch

    On 08/09/2026 05:21, Stephen Fuld wrote:
    On 9/7/2026 7:36 PM, John Savard wrote:
    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    I should think the answer is obvious. Left and right are fundamentally
    equivalent; they may be opposites, but there's no reason to prefer one
    over the other.

    Perhaps not quite true.-a Since the vast majority of people are right handed, it is easier for them to write left to right, as it means you
    don't run the risk of smudging what you have already wrote as left
    handers do with left to right order (think old, pre-ball point pens.).


    I think if you were chiselling the letters into stone, you'd have the
    chisel in your left hand and hammar in your right hand, and it would be
    easier to see what you are doing if you went right to left. So the
    manner of writing is also relevant.

    Bottom to top (or horizontal lines, starting at the bottom rather than
    the top) might also be a convenient choice for carving letters on a
    monument. If the text is short enough, it will save you getting a ladder.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 8 10:29:20 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:
    Is it true that x86 instructions containing integer literals put them
    in big-endian order?

    Since 8088 was a byte sized buss and read things in use order, I would bet >against that.

    The Datapoint 2200 (8088 ancestor) put immediate arguments that are
    larger than 8 bits (IIRC used in control-flow instructions) in
    little-endian order. And that was the only hardware thing in the
    Datapoint 2200 (and the 8008, its successor), where the byte order
    mattered.

    Concerning the 8088, I write in
    <2026Jun1.103622@mips.complang.tuwien.ac.at>:

    |For the 8088, in theory little-endian might provide an
    |advantage when it comes to addressing modes such as disp16[BX], but
    |AFAIK in practice the 8088 was internally mostly an 8086, with a
    |16-bit adder, so it loaded the whole 16-bit number anyway before doing
    |the full 16-bit add (am I wrong?). Likewise for the 386SX.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 8 10:56:59 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:
    But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    IA-32 has AAA. The AMD64 ISA has not. See <https://www.felixcloutier.com/x86/aaa>.

    No such instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
    Alpha, IA-64, ARM A64, RISC-V.

    And I daresay that IA-32 would not have had it if it had not been in
    the 80286. Whether the 80286 would have had it if the 8086 did not
    have it is less clear.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Tue Sep 8 13:34:12 2026
    From Newsgroup: comp.arch

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >Alpha, IA-64, ARM A64, RISC-V.

    And I daresay that IA-32 would not have had it if it had not been in
    the 80286. Whether the 80286 would have had it if the 8086 did not
    have it is less clear.

    I'm sure it wouldn't. AAA was extremely slow because nobody used it so
    nobody cared.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Tue Sep 8 08:31:54 2026
    From Newsgroup: comp.arch

    On 9/8/2026 2:32 AM, David Brown wrote:
    On 08/09/2026 05:21, Stephen Fuld wrote:
    On 9/7/2026 7:36 PM, John Savard wrote:
    On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    I should think the answer is obvious. Left and right are fundamentally
    equivalent; they may be opposites, but there's no reason to prefer one
    over the other.

    Perhaps not quite true.-a Since the vast majority of people are right
    handed, it is easier for them to write left to right, as it means you
    don't run the risk of smudging what you have already wrote as left
    handers do with left to right order (think old, pre-ball point pens.).


    I think if you were chiselling the letters into stone, you'd have the
    chisel in your left hand and hammar in your right hand, and it would be easier to see what you are doing if you went right to left.-a So the
    manner of writing is also relevant.

    Bottom to top (or horizontal lines, starting at the bottom rather than
    the top) might also be a convenient choice for carving letters on a monument.-a If the text is short enough, it will save you getting a ladder.


    Right! :-)
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Tue Sep 8 16:00:30 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/7/2026 1:24 PM, Waldek Hebisch wrote:



    And, of course, S/360 had an instruction to "pack" a string of numeric >EBCDIC characters to two, four bit, digits per byte, an unpack
    instruction to convert back, and a set of arithmetic instructions that >worked on packed data. These are still all supported on Z series.


    The Burroughs B3500 did this by simply adding, removing or preserving the
    zone digit (high-order 4 bits) of the EBCDIC when moving
    data from memory to memory. The ASCII flag indicated whether
    the zone digit inserted was 0b0011 or 0b1111.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Tue Sep 8 16:01:37 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    x86-64 repurposed the opcode formerly assigned to the AAA instruction.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 8 17:16:36 2026
    From Newsgroup: comp.arch


    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
    ------------

    And what do we do about people who want to write on Mobius strips?

    Best response ever.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 8 17:23:47 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
    --------------------

    Now we are really going down the rabbit hole. There is an issue with "spiral" directions. Let's say you start at the upper left corner, then proceed to the right. When you hit the right edge, you are going to go "down". But what is the orientation of the letters, i.e. should you
    turn the paper 90 degrees to keep reading left to right, or should the orientation of the letters change so you can now read top to bottom
    without reorientating the paper? Both have problems.

    Consider the "seal of the President" it is written clockwise around the
    outer margin of the seal but since it is but one line it is not spiral.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 8 17:31:51 2026
    From Newsgroup: comp.arch

    On Tue, 8 Sep 2026 03:46:48 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    On Tue, 08 Sep 2026 02:33:17 GMT, John Savard wrote:

    Numbers are a normal native element of Arabic text. So while there
    would be a fancy switchover of directions if an English text were
    quoted in an Arabic document, the same way as if a Hebrew text were
    quoted in an English document, what do you mean that Arabs have to
    do the same switchover every time a number appears in a document, as
    if Arabic numerals were foreign to Arabic?

    Yup. ThatrCOs why a script like Arabic is called rCLbidirectionalrCY, not >rCLright-to-leftrCY.

    Called by whom? I mean, the linguistic classification of the Arabic
    script would have been determined long before Unicode was invented.
    Before dot-matrix printers were invented. Before computers were
    invented.

    So why would its classification derive from Unicode rules?

    I'm beginning to think we are living in different worlds here.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 8 21:14:59 2026
    From Newsgroup: comp.arch

    Michael S wrote:
    On Mon, 7 Sep 2026 13:56:05 +0200
    Terje Mathisen <terje.mathisen@tmsw.no> wrote:

    Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Thomas Koenig wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    Big endian is just as dead now as bytes that are not 8-bits
    wide.

    I have login to several machines which are big-endian (and we
    receive bug reports about them for gfortran).

    Yeah, I know that they stil exist, but for anyone working on a new
    architecture, please just forget about this particular issue.

    Do not forget:: remember that LE won and any new computer
    architecture shall be LE. That fact that one can make a coherent
    BE machine is what should be forgotten.

    One might also remember that there are external protocols
    that will forever require big-endian ordering (IP/TCP et alia).

    These protocols are a sufficient reason to provide BSWAP type
    capability in all CPU architectures. (Including BE ones since there
    are other LE protocols/encodings).

    Terje


    One 2-byte BE field in IP header. Another one 2-byte field in UDP
    header. The rest is either non-endian (single-byte or less) or
    endian-neutral (IP address, ports, checksums). I didn't look at TCP
    header but would think that it's approximately the same.
    It does not sound as sufficient reason to have BSWAP.

    NTP works with 64-bit timstamps, with defined byte order: If your CPU
    does not agree with that, then you must have an 8-byte BSWAP, or general swizzle.

    I am sure there are many, many others!

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 8 21:17:50 2026
    From Newsgroup: comp.arch

    John Savard wrote:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    And they're certainly normal machines. They're the ones you need to
    run regular Windows programs.

    Lots of rarely used instruction went away when AMD defined x64.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 8 21:23:20 2026
    From Newsgroup: comp.arch

    John Levine wrote:
    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
    Alpha, IA-64, ARM A64, RISC-V.

    And I daresay that IA-32 would not have had it if it had not been in
    the 80286. Whether the 80286 would have had it if the 8086 did not
    have it is less clear.

    I'm sure it wouldn't. AAA was extremely slow because nobody used it so nobody cared.

    It wasn't just slow, it was intrinsically Single Instruction, Single
    Data. On any 32-bit or wider CPU, doing the same style of decimal math
    with 4, 8 or 16 digits in a register was far more efficient.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 8 19:30:56 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    quadibloc@invalid.com (John Savard) writes:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek >>Hebisch) wrote:
    But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    IA-32 has AAA. The AMD64 ISA has not. See
    <https://www.felixcloutier.com/x86/aaa>.

    No such instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
    Alpha, IA-64, ARM A64, RISC-V.

    Your information is off. HP-PA has the DCOR instruction. Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 8 20:13:40 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    /ERROR "unexpected byte sequence starting at index 572: '\xE2'" while decoding/:

    On Tue, 8 Sep 2026 03:46:48 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    On Tue, 08 Sep 2026 02:33:17 GMT, John Savard wrote:

    Numbers are a normal native element of Arabic text. So while there
    would be a fancy switchover of directions if an English text were
    quoted in an Arabic document, the same way as if a Hebrew text were
    quoted in an English document, what do you mean that Arabs have to
    do the same switchover every time a number appears in a document, as
    if Arabic numerals were foreign to Arabic?

    Yup. That|o-C-Os why a script like Arabic is called |o-C-Lbidirectional|o-C-Y, not
    |o-C-Lright-to-left|o-C-Y.

    Called by whom? I mean, the linguistic classification of the Arabic
    script would have been determined long before Unicode was invented.
    Before dot-matrix printers were invented. Before computers were
    invented.

    At NCR back in 1978, the company increased the character map to contain
    both Arabic and Hebrew scripts (different story) so, we had to write the display driver so that when Arabic characters were being input, the display added them heading left, and when Arabic numerals were being typed, the
    numeric field was added shifting them towards the left. So typing 1,234
    would see the following sequence: "1previous text", "1,previous text", "1,2previous text", 1,23previous text", and "1,234previous text" where
    the last t of text was at a fixed position on the display.

    So why would its classification derive from Unicode rules?

    I'm beginning to think we are living in different worlds here.

    The problem is that Americans and Europeans are so self absorbed they
    don't recognize other cultural preferences--then design rules such
    that those preferences can't be allowed.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 8 20:15:06 2026
    From Newsgroup: comp.arch


    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    John Savard wrote:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    And they're certainly normal machines. They're the ones you need to
    run regular Windows programs.

    Lots of rarely used instruction went away when AMD defined x64.

    We needed the OpCode space.

    Terje

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 00:22:37 2026
    From Newsgroup: comp.arch

    On Tue, 08 Sep 2026 17:31:51 GMT, John Savard wrote:

    So why would its classification derive from Unicode rules?

    No-one said it did. ItrCOs the other way round: Unicode has to cater to
    the way actual writing systems work.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Wed Sep 9 01:10:45 2026
    From Newsgroup: comp.arch

    On Wed, 9 Sep 2026 00:22:37 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
    On Tue, 08 Sep 2026 17:31:51 GMT, John Savard wrote:

    So why would its classification derive from Unicode rules?

    No-one said it did. ItrCOs the other way round: Unicode has to cater to
    the way actual writing systems work.

    Before Unicode came along, I thought that Arabs wrote from right to
    left, full stop. This included numbers. So, since the most significant
    digit was on the left for them, the same way it was for us, I assumed innocently that the least significant digit was the one that they
    wrote first and read first.

    Which is what would make little-endian byte order seem natural and
    intuitive for them - as it matches the way they read and write
    numbers.

    If the Arabic-style digits, unlike the letters of the Arabic alphabet,
    are indeed classified as left-to-right characters, so that they don't
    behave differently from other characters representing decimal digits,
    then the behavior you have described would indeed happen. But I view
    that behavior strictly as a consequence of Unicode. People wouldn't
    have waited until the technology of handling mixed-direction
    characters was available to build printers for use in the Arab world
    if computing had gotten up to an early start there.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Wed Sep 9 01:12:45 2026
    From Newsgroup: comp.arch

    On Tue, 8 Sep 2026 08:38:29 -0000 (UTC), John Levine <johnl@taugh.com>
    wrote:
    According to John Savard <quadibloc@invalid.com>:

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    x86-64 has 32 bit and 64 bit modes. AAA is still there in 32 bit mode
    but not in 64 bit mode.

    Thank you, I've learned something today.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 01:33:00 2026
    From Newsgroup: comp.arch

    On Wed, 09 Sep 2026 01:10:45 GMT, John Savard wrote:

    If the Arabic-style digits, unlike the letters of the Arabic
    alphabet, are indeed classified as left-to-right characters, so that
    they don't behave differently from other characters representing
    decimal digits, then the behavior you have described would indeed
    happen. But I view that behavior strictly as a consequence of
    Unicode.

    The term was always properly rCLbidirectionalrCY, never simplistically rCLright-to-leftrCY. I first came across the term before I even knew what Unicode was (Arabic support in old MacOS, rCLWorldScriptrCY I and II).

    Looking it up on Wikipedia
    <https://en.wikipedia.org/wiki/Bidirectional_text> makes that clear:

    Some so-called right-to-left scripts such as the Persian script
    and Arabic are mostly, but not exclusively,
    right-to-left--mathematical expressions, numeric dates and numbers
    bearing units are embedded from left to right.

    Also note the mention of Egyptian hieroglyphs and Chinese text.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 9 08:04:00 2026
    From Newsgroup: comp.arch

    John Levine <johnl@taugh.com> writes:
    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >>Alpha, IA-64, ARM A64, RISC-V.

    And I daresay that IA-32 would not have had it if it had not been in
    the 80286. Whether the 80286 would have had it if the 8086 did not
    have it is less clear.

    I'm sure it wouldn't. AAA was extremely slow because nobody used it so >nobody cared.

    But it seems to me that at the time it was still fashionable to add instructions to "close the semantic gap" and to have slow microcoded instructions instead of sequences of simpler instructions for code
    density. They added PUSHA, POPA, ENTER and BOUND in the 80186 <https://www.eeeguide.com/instruction-set-of-80186/>, and LEAVE in the
    80286 <https://www.eeeguide.com/80286-instruction-set/>, none of which
    are known for their speed. These CPUs were designed when VAX was in
    full swing, before RISC became a thing.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 9 08:22:44 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    IA-32 has AAA. The AMD64 ISA has not. See >><https://www.felixcloutier.com/x86/aaa>.

    No such instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
    Alpha, IA-64, ARM A64, RISC-V.

    Your information is off. HP-PA has the DCOR instruction. Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    Last I looked, the HPPA stuff worked on packed BCD (4 bits per digit),
    8 digits at a time (possibly 16 on HPPA 2.0, but I have not looked at
    that). By contrast, AAA works on unpacked BCD (8 bits per digit) and
    on a single byte. IA-32 has a similar instruction for packed BCD: DAA
    (also not in AMD64), also working on a single byte (but at least
    that's two digits).

    One could argue that HPPA has something like DAA, but AAA?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 9 11:06:40 2026
    From Newsgroup: comp.arch

    On 09/09/2026 03:10, John Savard wrote:
    On Wed, 9 Sep 2026 00:22:37 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
    On Tue, 08 Sep 2026 17:31:51 GMT, John Savard wrote:

    So why would its classification derive from Unicode rules?

    No-one said it did. ItrCOs the other way round: Unicode has to cater to
    the way actual writing systems work.

    Before Unicode came along, I thought that Arabs wrote from right to
    left, full stop. This included numbers. So, since the most significant
    digit was on the left for them, the same way it was for us, I assumed innocently that the least significant digit was the one that they
    wrote first and read first.

    Which is what would make little-endian byte order seem natural and
    intuitive for them - as it matches the way they read and write
    numbers.

    If the Arabic-style digits, unlike the letters of the Arabic alphabet,
    are indeed classified as left-to-right characters, so that they don't
    behave differently from other characters representing decimal digits,
    then the behavior you have described would indeed happen. But I view
    that behavior strictly as a consequence of Unicode. People wouldn't
    have waited until the technology of handling mixed-direction
    characters was available to build printers for use in the Arab world
    if computing had gotten up to an early start there.

    John Savard

    In Unicode, the Arabic numerals +a+i+o+u+n+N+a+o+?+- have left-to-right order, while most other Arabic characters are right-to-left. It would be
    interesting to know if native Arabic writers actually change direction
    when writing by hand, or if they simply write the least significant
    digit first.

    It is actually normal for some parts of text to be viewed in different directions or given different orders, if the parts of the text are from
    a writing system in the other direction. If I were quoting something in Hebrew in the middle of a sentence in English, I'd use normal
    right-to-left ordering for the Hebrew. It's just that it is a lot more
    common for Arabic or Hebrew speakers to write bits in Latin alphabets
    than vice versa, so us "westerners" are not as familiar with
    bi-direction writing.

    Ordering of information, such as numbers, and order of writing, are not necessarily the same thing. For phonetic writing systems, where letters represent sounds, the order of writing is the order of spoken sounds.
    But other written information is not necessarily like that. Digits and numerical information can mean many different things, pronounced in
    different ways, and ordered in different ways. "87" can be pronounced
    "four score and seven", and numbers can be written as "big-endian", "little-endian", or "muddle-endian" (like "XIV", or American dates), and non-trivial maths formula are often completely unpronounceable.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 9 11:08:32 2026
    From Newsgroup: comp.arch

    On 08/09/2026 22:15, MitchAlsup wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    John Savard wrote:
    On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
    Hebisch) wrote:

    8086 and 8088 while mainly binary had special instructions
    to support BCD arithmetic and arithmetic on string of ASCII digits
    (first instruction in 8088 instruction list is AAA, that is
    ASCII adjust after addition). But time changed and special support
    of this sort is gone from normal machines.

    What? If the 8086 had an AAA instruction, then wouldn't all subsequent
    x86 and x86-64 architecture machines also have that instruction?

    And they're certainly normal machines. They're the ones you need to
    run regular Windows programs.

    Lots of rarely used instruction went away when AMD defined x64.

    We needed the OpCode space.


    Fair enough. It was a good opportunity to clear out some old baggage.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Tue Sep 8 23:57:15 2026
    From Newsgroup: comp.arch

    Yup. That's why a script like Arabic is called bidirectional, not >>right-to-left
    Called by whom?

    What's happening here is that people like me who grew up in a "latin
    script" world are only exposed to left-to-right text, which makes it
    difficult for us to imagine anything else. It already takes an effort
    to imagine "right-to-left" or "top-to-bottom".

    In contrast, people who grow up around arabic or hebraic scripts are
    inevitably exposed to chunks of text going left-to-right and other
    chunks going right-to-left, all mixed in the same document.
    Whether it's because of math elements or elements in other languages or whatnot. In such environments, it's hard to find social bubbles where
    all text goes in the same direction. And it's been that way since
    before computers were invented.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 9 23:21:29 2026
    From Newsgroup: comp.arch

    Anton Ertl wrote:
    John Levine <johnl@taugh.com> writes:
    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >>> Alpha, IA-64, ARM A64, RISC-V.

    And I daresay that IA-32 would not have had it if it had not been in
    the 80286. Whether the 80286 would have had it if the 8086 did not
    have it is less clear.

    I'm sure it wouldn't. AAA was extremely slow because nobody used it so
    nobody cared.

    But it seems to me that at the time it was still fashionable to add instructions to "close the semantic gap" and to have slow microcoded instructions instead of sequences of simpler instructions for code
    density. They added PUSHA, POPA, ENTER and BOUND in the 80186 <https://www.eeeguide.com/instruction-set-of-80186/>, and LEAVE in the
    80286 <https://www.eeeguide.com/80286-instruction-set/>, none of which
    are known for their speed. These CPUs were designed when VAX was in
    full swing, before RISC became a thing.

    Please don't blame POPA!

    This was the final key ASCII opcode that made it possible to write an executable text program, using noting but those 70+ 7-bit character
    codes that are blessed in the MIME standard because they never need quoting.

    The problem was that without POPA, there was no way to get a predictable
    value loaded into any memory-addresing register: POP BX, SI, DI or BP
    are all using single-byte opcodes, but all of them are in the restricted character set, like '[', or ']'.

    What I did was to start by POP'ing the zero which is guaranteed to be on
    the top of the stack, PUSH it back, then DECrement the value to create
    0xFFF, push this value then DEC SP (leaving an unaligned stack pointer).
    Now I could POP a mixed value (0x00FF or 255).

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From George Neuner@gneuner2@comcast.net to comp.arch on Wed Sep 9 20:22:21 2026
    From Newsgroup: comp.arch

    On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
    <tr.17687@z991.linuxsc.com> wrote:


    ... trying to write on a Mobius strip I'm at a loss as to
    where to begin. ;)

    You start anywhere and keep going. 8-)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 10 01:25:37 2026
    From Newsgroup: comp.arch


    George Neuner <gneuner2@comcast.net> posted:

    On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
    <tr.17687@z991.linuxsc.com> wrote:


    ... trying to write on a Mobius strip I'm at a loss as to
    where to begin. ;)

    You start anywhere and keep going. 8-)

    when you run into yourself do you step up or down ??
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Thu Sep 10 12:09:37 2026
    From Newsgroup: comp.arch

    MitchAlsup wrote:

    George Neuner <gneuner2@comcast.net> posted:

    On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
    <tr.17687@z991.linuxsc.com> wrote:


    ... trying to write on a Mobius strip I'm at a loss as to
    where to begin. ;)

    You start anywhere and keep going. 8-)

    when you run into yourself do you step up or down ??


    If the strip is very long so you could run out of ink, you can instead
    send your twice as fast brother to run in front of you: When he catches
    back up with you, you know how long the strip was.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 14 10:29:19 2026
    From Newsgroup: comp.arch

    On 9/13/2026 11:55 AM, Paul Clayton wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    -aPower's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    Of course, if humans had standardized on base 16, this would not
    have been an issue. While base 60 simplifies division by 2, 3, 4, 5, 6,
    10, and 12, base 16 provides reasonable approximations
    for a third (5/16) and a fifth (3/16) without requiring such a
    large digit. 16 also has the nice binary quality of being the
    square of two squared.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|

    As the late great Tom Lehrer said decades ago, "Base eight is just like
    base ten if you don't have any thumbs."
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 14 17:32:31 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> schrieb:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15.

    Would this have made BCD easier, and if so, how?

    (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    EBCDIC is an extension of the six-bit characters that IBM used
    previously. ASCII was designed from scratch (with IBM, but the
    /360 came out too early for that).

    Of course, if humans had standardized on base 16, this would not
    have been an issue. While base 60 simplifies division by 2, 3,
    4, 5, 6, 10, and 12, base 16 provides reasonable approximations
    for a third (5/16) and a fifth (3/16) without requiring such a
    large digit. 16 also has the nice binary quality of being the
    square of two squared.

    The earliest numbering systems were more like base 60, in Sumer;
    I'd have to look up what Egyptians used. But they used only
    reciprocals for fractions, which is an interesting approach.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Mon Sep 14 23:29:03 2026
    From Newsgroup: comp.arch

    On Sun, 13 Sep 2026 14:55:25 -0400, Paul Clayton wrote:

    While base 60 simplifies division by 2, 3, 4, 5, 6, 10, and 12, base
    16 provides reasonable approximations for a third (5/16) and a fifth
    (3/16) without requiring such a large digit.

    Base 30 gives you terminating fractional representations for all the
    same divisors as base 60, and powers and combinations thereof.

    You only need one occurrence of a prime factor to be able to deal in a
    finite fashion with all combinations involving that factor.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Mon Sep 14 12:39:31 2026
    From Newsgroup: comp.arch

    Please don't blame POPA!

    This was the final key ASCII opcode that made it possible to write an executable text program, using noting but those 70+ 7-bit character codes that are blessed in the MIME standard because they never need quoting.

    I think John's Concertina would be the ideal ISA to extend with
    a "base64" subset where the encoding is careful to ensure all the bytes
    are properly preserved when sent via email (bonus points if it tolerates LF<=>CRLF conversions).


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 11:26:32 2026
    From Newsgroup: comp.arch

    Stefan Monnier wrote:
    Please don't blame POPA!

    This was the final key ASCII opcode that made it possible to write an
    executable text program, using noting but those 70+ 7-bit character codes
    that are blessed in the MIME standard because they never need quoting.

    I think John's Concertina would be the ideal ISA to extend with
    a "base64" subset where the encoding is careful to ensure all the bytes
    are properly preserved when sent via email (bonus points if it tolerates LF<=>CRLF conversions).

    Nice!

    My maketext actually handles any zero, one or two-byte line terminator
    setup, I wrote it with sufficient 'EEEE' targets for all cross-line
    jumps, so that it would work anyway. 'E' was my NOP.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 11:32:15 2026
    From Newsgroup: comp.arch

    Stephen Fuld wrote:
    On 9/13/2026 11:55 AM, Paul Clayton wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    |e-aPower's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    Of course, if humans had standardized on base 16, this would not
    have been an issue. While base 60 simplifies division by 2, 3, 4, 5,
    6, 10, and 12, base 16 provides reasonable approximations
    for a third (5/16) and a fifth (3/16) without requiring such a
    large digit. 16 also has the nice binary quality of being the
    square of two squared.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.|o-L-|

    As the late great Tom Lehrer said decades ago, "Base eight is just like
    base ten if you don't have any thumbs."
    As my also late great High School math teacher did:
    He brought his cherished Tom Lehrer LP to math class and played Tom's
    "New Math" song as the introduction to alternate base math, writing down the problems on the blackboard and following along with the song.
    A few months later a substitute teacher gave me an A- when grading the
    part of a test where I had done a direct base-12 to base-2 conversion,
    without going via base-10 as the book told us to. :-)
    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 12:49:35 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro wrote:
    On Sun, 13 Sep 2026 14:55:25 -0400, Paul Clayton wrote:

    While base 60 simplifies division by 2, 3, 4, 5, 6, 10, and 12, base
    16 provides reasonable approximations for a third (5/16) and a fifth
    (3/16) without requiring such a large digit.

    Base 30 gives you terminating fractional representations for all the
    same divisors as base 60, and powers and combinations thereof.

    You only need one occurrence of a prime factor to be able to deal in a
    finite fashion with all combinations involving that factor.

    Base 30 is perfect for a fast prime sieve, since there are exactly 8
    possible locations for a prime larger than 30.
    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 15 13:01:36 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    Lawrence D=E2=80=99Oliveiro wrote:
    You only need one occurrence of a prime factor to be able to deal in a
    finite fashion with all combinations involving that factor.
    =20
    Base 30 is perfect for a fast prime sieve, since there are exactly 8=20 >possible locations for a prime larger than 30.

    30 is also pretty close to a power of 2, so you can store it in binary
    without too much waste. Continuing along this line with further prime
    factors:

    base ld(base)
    2 1
    * 3 = 6 2.584962500721156
    * 5 = 30 4.906890595608519
    * 7 = 210 7.714245517666122
    * 11 = 2310 11.17367713630342
    * 13 = 30030 14.874116854444512
    * 17 = 510510 18.96157969569485

    510510 looks like another good base for this kind of thing. If it
    should be smaller, another good one is 2*3*5*17=510
    (ld(510)=8.994...).

    However, like for decimal FP I have my doubts that the benefits of
    non-binary bases are enough to justify their costs.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 07:49:49 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/7/2026 3:23 PM, Tim Rentsch wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/6/2026 5:31 PM, Stephen Fuld wrote:

    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to
    use long scrolls. :-)

    But ISTM that the choice of horizontal direction is independent of
    the choice of vertical direction, thus yielding four potential
    choices.

    Sorry for the self followup, but actually eight choices. For each of
    the four corners, proceed in the horizontal direction or the vertical
    direction.

    How about, for each of the four corners, proceed in a clockwise or
    anti-clockwise spiral? That's eight more choices.

    Now we are really going down the rabbit hole. There is an issue with "spiral" directions. Let's say you start at the upper left corner,
    then proceed to the right. When you hit the right edge, you are going
    to go "down". But what is the orientation of the letters, i.e. should
    you turn the paper 90 degrees to keep reading left to right, or should
    the orientation of the letters change so you can now read top to
    bottom without reorientating the paper? Both have problems.

    If you say you will change the orientation of the paper, what if the
    "paper" is actually a bronze plaque mounted on a monument. Kinda hard
    to change its orientation. And when you get to the bottom, do you
    stand on your head to read it? :-)

    On the other hand, say you change the orientation of the letters.
    That works for the right side - you just read it top to bottom. But
    now what do you do when you get to the bottom? Do you print the
    letters upside down, in which case you have the same problem as the
    above solution. You could keep the letters "right side up:, but now
    the words are spelled backwards. That's awkward.

    So I reject the spiral solutions, and other, even more bizarre
    ones. But still, eight possibilities is more than enough. :-)

    I'm sorry you took my comments so seriously. They were very much
    meant tongue in cheek.

    That said, let me address your comments.

    If we are considering a Latin alphabet, which incidentally is a
    phonetic alphabet, the only ordering that really works is
    left-to-right then top-to-bottom. This follows from the shapes of
    different letters: they have different heights and widths, and also
    are stylized in a way to facilitate left-to-right scanning. The
    idea that the aforementioned eight possibilities are all viable
    doesn't hold up to serious scrutiny.

    If on the other hand we are considering a word-based alphabet, such
    as Chinese, things are very different. Characters typically have
    similarly sized bounding boxes, and are not stylized to prefer a
    particular reading direction: left-to-right, or right-to-left, or top-to-bottom -- all are workable. But even Chinese characters have
    a preferred up/down distinction, clearly favoring a top-to-bottom
    orientation over a bottom-to-top orientation. I'm pretty sure I
    have seen all of left-to-right, right-to-left, and top-to-bottom,
    for word-based alphabets, in regular texts.

    More exotic alphabets, as for example hieroglyphics or emoji or
    traffic signs, may occupy different categories. And of course
    squeezing all of these things into tiny pictograms on cell phones
    has distorted the notion of what we consider "language". (Side
    comment: some cell phones have the bad habit of spontaneously
    rotating from portrait mode to landscape mode, which I find
    thoroughly annoying.)

    In light of the above, IMO your eight possibilities are no more
    serious than the other possibilities I mentioned. It seems
    disingenuous to pretend otherwise.

    (My apologies for posting comments that have nothing to do with
    computer architecture.)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 16:51:39 2026
    From Newsgroup: comp.arch

    Anton Ertl wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    Lawrence D=E2=80=99Oliveiro wrote:
    You only need one occurrence of a prime factor to be able to deal in a
    finite fashion with all combinations involving that factor.
    =20
    Base 30 is perfect for a fast prime sieve, since there are exactly 8=20
    possible locations for a prime larger than 30.

    30 is also pretty close to a power of 2, so you can store it in binary without too much waste. Continuing along this line with further prime factors:

    base ld(base)
    2 1
    * 3 = 6 2.584962500721156
    * 5 = 30 4.906890595608519
    * 7 = 210 7.714245517666122
    * 11 = 2310 11.17367713630342
    * 13 = 30030 14.874116854444512
    * 17 = 510510 18.96157969569485

    510510 looks like another good base for this kind of thing. If it
    should be smaller, another good one is 2*3*5*17=510
    (ld(510)=8.994...).

    However, like for decimal FP I have my doubts that the benefits of
    non-binary bases are enough to justify their costs.

    I did consider higher multiples, like 210, 2310 and 30030, the problem
    is that you want the number of potential primes to be less or equal to
    8, 16, 24 or 32.

    Base 30 is in fact a significant speedup for Sieve work:

    For each prime to be crossed out I calculated 8 pairs of
    [byte_increment, next_bit], so that stepping along was simply

    numbers[bytenr] |= 1 << bit;
    bytenr += byte_increment[bit];
    bit = next_bit[bit];

    I loaded up a bit less than last level cache worth of numbers, before
    crossing out all the primes found so far.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 08:53:47 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    As the late great Tom Lehrer said decades ago, "Base eight is just
    like base ten if you don't have any thumbs."

    Quoting:

    Now that actually is not the answer that I had in mind,
    because book I got this problem out of wants you to do
    it in base eight.

    But don't panic.

    Base eight is just like base ten really ...

    ... if you're missing two fingers.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 09:19:32 2026
    From Newsgroup: comp.arch

    George Neuner <gneuner2@comcast.net> writes:

    On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
    <tr.17687@z991.linuxsc.com> wrote:


    ... trying to write on a Mobius strip I'm at a loss as to
    where to begin. ;)

    You start anywhere and keep going. 8-)

    So, pretty much the same then as how some people choose
    to post in comp.arch? :)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Tue Sep 15 10:09:21 2026
    From Newsgroup: comp.arch

    On 9/15/2026 7:49 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/7/2026 3:23 PM, Tim Rentsch wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/6/2026 5:31 PM, Stephen Fuld wrote:

    On 9/6/2026 4:57 PM, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    Perhaps it has to do with the problems it would cause when trying to >>>>> use long scrolls. :-)

    But ISTM that the choice of horizontal direction is independent of
    the choice of vertical direction, thus yielding four potential
    choices.

    Sorry for the self followup, but actually eight choices. For each of
    the four corners, proceed in the horizontal direction or the vertical
    direction.

    How about, for each of the four corners, proceed in a clockwise or
    anti-clockwise spiral? That's eight more choices.

    Now we are really going down the rabbit hole. There is an issue with
    "spiral" directions. Let's say you start at the upper left corner,
    then proceed to the right. When you hit the right edge, you are going
    to go "down". But what is the orientation of the letters, i.e. should
    you turn the paper 90 degrees to keep reading left to right, or should
    the orientation of the letters change so you can now read top to
    bottom without reorientating the paper? Both have problems.

    If you say you will change the orientation of the paper, what if the
    "paper" is actually a bronze plaque mounted on a monument. Kinda hard
    to change its orientation. And when you get to the bottom, do you
    stand on your head to read it? :-)

    On the other hand, say you change the orientation of the letters.
    That works for the right side - you just read it top to bottom. But
    now what do you do when you get to the bottom? Do you print the
    letters upside down, in which case you have the same problem as the
    above solution. You could keep the letters "right side up:, but now
    the words are spelled backwards. That's awkward.

    So I reject the spiral solutions, and other, even more bizarre
    ones. But still, eight possibilities is more than enough. :-)

    I'm sorry you took my comments so seriously. They were very much
    meant tongue in cheek.

    I did assume your comments weren't serious, and I apologize if my "down
    the rabbit hole" didn't make that clear. :-(


    That said, let me address your comments.

    Big snip. Thank you. I found your comments interesting. My only disagreement is that there are a fair number of instances of at least
    English being written top to bottom, usually in signs organized
    vertically for space reasons. But there, even if they have multiple
    lines, are not "continuous" from one line to the next. I think most
    people have no trouble reading them

    (My apologies for posting comments that have nothing to do with
    computer architecture.)

    Yes. This will be my last post on this topic. I hope you have no hard feelings.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.arch on Wed Sep 16 04:41:19 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    IIUC some digital systems used "excess 3" code, that is started
    from 3 and went up to 12. It had advantage that using 4 bit
    biary adder you got binary carry exactly when there was decimal
    carry. But after that you had to do correction: if there were
    no carry subtract 3, if there was carry add 3.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|

    I read supposedly serious proposal to switch to base 8. Motivation
    was smaller and consequently easier to learn multiplication table.
    IIRC this proposal was in 19-th century Sweden. But it seems that
    they are still using decimal, so probably it did not pass.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 16 05:53:24 2026
    From Newsgroup: comp.arch

    Waldek Hebisch <antispam@fricas.org> schrieb:
    Paul Clayton <paaronclayton@gmail.com> wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    IIUC some digital systems used "excess 3" code, that is started
    from 3 and went up to 12. It had advantage that using 4 bit
    biary adder you got binary carry exactly when there was decimal
    carry. But after that you had to do correction: if there were
    no carry subtract 3, if there was carry add 3.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|

    I read supposedly serious proposal to switch to base 8. Motivation
    was smaller and consequently easier to learn multiplication table.
    IIRC this proposal was in 19-th century Sweden. But it seems that
    they are still using decimal, so probably it did not pass.

    Larry Niven has a few science fiction stories where aliens use
    base eight. He may have been influenced (however remotely) by
    the DEC machines.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 16 07:52:39 2026
    From Newsgroup: comp.arch

    On Wed, 16 Sep 2026 05:53:24 -0000 (UTC), Thomas Koenig wrote:

    Larry Niven has a few science fiction stories where aliens use base
    eight. He may have been influenced (however remotely) by the DEC
    machines.

    All the DEC machines (prior to the PDP-11) had word sizes which were
    multiples of 3: 12, 18, 36. Even the 16-bit PDP-11 gave meanings to
    3-bit fields in its instruction layout that mapped naturally to octal
    digits of the integer value of the instruction word.

    Was it IBM that popularized hexadecimal?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Jeroen Belleman@jeroen@nospam.please to comp.arch on Wed Sep 16 10:48:47 2026
    From Newsgroup: comp.arch

    On 9/16/26 06:41, Waldek Hebisch wrote:
    Paul Clayton <paaronclayton@gmail.com> wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    IIUC some digital systems used "excess 3" code, that is started
    from 3 and went up to 12. It had advantage that using 4 bit
    biary adder you got binary carry exactly when there was decimal
    carry. But after that you had to do correction: if there were
    no carry subtract 3, if there was carry add 3.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|

    I read supposedly serious proposal to switch to base 8. Motivation
    was smaller and consequently easier to learn multiplication table.
    IIRC this proposal was in 19-th century Sweden. But it seems that
    they are still using decimal, so probably it did not pass.


    I'd really have liked a system with digit values symmetrical
    around zero, maybe base 10, or even better, base 12. This
    makes arithmetic delightfully simple.

    There was an article about it in New Scientist of April 22,
    1982. I loved it.

    Jeroen Belleman
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 16 16:21:05 2026
    From Newsgroup: comp.arch

    Jeroen Belleman wrote:
    On 9/16/26 06:41, Waldek Hebisch wrote:
    Paul Clayton <paaronclayton@gmail.com> wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    -a Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    IIUC some digital systems used "excess 3" code, that is started
    from 3 and went up to 12.-a It had advantage that using 4 bit
    biary adder you got binary carry exactly when there was decimal
    carry.-a But after that you had to do correction: if there were
    no carry subtract 3, if there was carry add 3.
    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.|o-L-|

    I read supposedly serious proposal to switch to base 8.-a Motivation
    was smaller and consequently easier to learn multiplication table.
    IIRC this proposal was in 19-th century Sweden.-a But it seems that
    they are still using decimal, so probably it did not pass.


    I'd really have liked a system with digit values symmetrical
    around zero, maybe base 10, or even better, base 12. This
    makes arithmetic delightfully simple.

    There was an article about it in New Scientist of April 22,
    1982. I loved it.
    There was an Advent of Code puzzle in 2022 (last day) where you had to
    work in symmetrical base-5, using '-' and '=' for -1 and -2, along with
    '0', '1', '2'.
    You had to parse 113 such numbers, all happened to be positive, but they did of course not need to be. With symmetrical encoding, there is no
    extra minus sign in front, just a negative first digit.
    The fast solutions did each input line/number in about 4-5 us, then you
    had to return the (decimal) sum, before converting that back to the same base-5 system.
    Just like decimal, it is a lot easier to convert ascii to binary than
    binary to ascii, the difference between div/mod every digit and a single mul. Another reason I liked it so much was because it reminded me of Nov/Dec
    28 years previously, when Cleve Moler & I decided to write a workaround
    for the Pentium FDIV bug.
    The Pentium SRT divider circuit extracted two bits/iteration by looking
    up an approximation from -2 to 2. With 5 choices for two bits, that 25% redundancy was enough that the guess didn't need to be exact since any
    errors would be incorporated in the next round. On the original Pentium
    mask, used for 66/60 and 90/100 MHz chips, the lookup table was missing
    five entries.
    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 16 07:29:02 2026
    From Newsgroup: comp.arch

    On 9/16/2026 12:52 AM, Lawrence DrCOOliveiro wrote:
    On Wed, 16 Sep 2026 05:53:24 -0000 (UTC), Thomas Koenig wrote:

    Larry Niven has a few science fiction stories where aliens use base
    eight. He may have been influenced (however remotely) by the DEC
    machines.

    All the DEC machines (prior to the PDP-11) had word sizes which were multiples of 3: 12, 18, 36. Even the 16-bit PDP-11 gave meanings to
    3-bit fields in its instruction layout that mapped naturally to octal
    digits of the integer value of the instruction word.

    Was it IBM that popularized hexadecimal?

    Yes. Due to the cost of circuitry and memory, Prior to S/360, almost
    all systems were either decimal, or binary with a six bit character set.
    Six bit characters fit nicely into two octal "digits" of three bits
    each (Binary Coded Decimal, or BCD for short, or similar codes). Hence systems like those you referred to, and most others had word sizes that
    were multiples of six, and octal was the defacto standard.

    With the S/360, IBM defined the "byte" as eight bits, and "extended" BCD
    into EBCDIC. With that, expressing each byte as two four bit "nibbles" encouraged each nibble to be expressed as a four bit (i.e. hexadecimal)
    value.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 16 15:49:51 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/16/2026 12:52 AM, Lawrence DrCOOliveiro wrote:
    On Wed, 16 Sep 2026 05:53:24 -0000 (UTC), Thomas Koenig wrote:

    Larry Niven has a few science fiction stories where aliens use base
    eight. He may have been influenced (however remotely) by the DEC
    machines.

    All the DEC machines (prior to the PDP-11) had word sizes which were
    multiples of 3: 12, 18, 36. Even the 16-bit PDP-11 gave meanings to
    3-bit fields in its instruction layout that mapped naturally to octal
    digits of the integer value of the instruction word.

    Was it IBM that popularized hexadecimal?

    Yes. Due to the cost of circuitry and memory, Prior to S/360, almost
    all systems were either decimal, or binary with a six bit character set.
    Six bit characters fit nicely into two octal "digits" of three bits
    each (Binary Coded Decimal, or BCD for short, or similar codes). Hence >systems like those you referred to, and most others had word sizes that
    were multiples of six, and octal was the defacto standard.

    With the S/360, IBM defined the "byte" as eight bits, and "extended" BCD >into EBCDIC. With that, expressing each byte as two four bit "nibbles" >encouraged each nibble to be expressed as a four bit (i.e. hexadecimal) >value.

    The B3500 was contemporaneous with the S/360, and also defined the
    "byte" as eight bits (two 4-bit digits). Undigits (values 0b1010 - 0b1111)
    in the B3500 would give odd results for arithmetic operations (later
    versions of the architecture would fault instead if undigits were used
    in an arithmetic operation).

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 16 16:17:26 2026
    From Newsgroup: comp.arch

    Jeroen Belleman <jeroen@nospam.please> schrieb:
    On 9/16/26 06:41, Waldek Hebisch wrote:
    Paul Clayton <paaronclayton@gmail.com> wrote:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15. (I
    suspect there were reasonable historical reasons for both EBDIC
    and ASCII layouts.)

    IIUC some digital systems used "excess 3" code, that is started
    from 3 and went up to 12. It had advantage that using 4 bit
    biary adder you got binary carry exactly when there was decimal
    carry. But after that you had to do correction: if there were
    no carry subtract 3, if there was carry add 3.

    If one uses one's thumb as a place holder, finger counting
    gives four per hand, so base 8 might have been very natural and
    base 16 might not have been too awkward.ry|

    I read supposedly serious proposal to switch to base 8. Motivation
    was smaller and consequently easier to learn multiplication table.
    IIRC this proposal was in 19-th century Sweden. But it seems that
    they are still using decimal, so probably it did not pass.


    I'd really have liked a system with digit values symmetrical
    around zero, maybe base 10, or even better, base 12. This
    makes arithmetic delightfully simple.

    I am a fan of balanced ternary, and have a few OEIS sequences
    to prove it :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Andreas Eder@a_eder_muc@web.de to comp.arch on Thu Sep 17 10:23:08 2026
    From Newsgroup: comp.arch

    On So 06 Sep 2026 at 23:57, MitchAlsup wrote:

    quadibloc@invalid.com (John Savard) posted:

    On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
    =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:

    Endianness has nothing to do with writing direction.

    Not directly.

    Humans developed 3 writing orders {Left-Right, Right->left, and
    Top->bottom}. I wonder why nobody standardized on bottom->top ??

    There is also boustrophedon (+#++-a-a-a-U++-a+++|-i++) where alternate lines of writing are
    reversed, with letters also written in reverse, mirror-style.

    'Andreas
    --
    ceterum censeo redmondinem esse delendam
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Fri Sep 18 00:38:39 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had to
    work in symmetrical base-5, using '-' and '=' for -1 and -2, along
    with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding, there is
    no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.

    Just like decimal, it is a lot easier to convert ascii to binary than
    binary to ascii, the difference between div/mod every digit and a
    single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Fri Sep 18 17:29:11 2026
    From Newsgroup: comp.arch

    On Fri, 18 Sep 2026 00:38:39 -0700
    Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had
    to work in symmetrical base-5, using '-' and '=' for -1 and -2,
    along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding,
    there is no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.


    May be, numbers were big? May be, bigger than what fits in 64 bits, or
    even in 128?

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Fri Sep 18 16:47:29 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had to
    work in symmetrical base-5, using '-' and '=' for -1 and -2, along
    with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding, there is
    no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.

    Oops!

    I meant 4-5 ns not microseconds.

    3 orders of magnitude does make a small difference...

    Just like decimal, it is a lot easier to convert ascii to binary than
    binary to ascii, the difference between div/mod every digit and a
    single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?


    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    In order to be really fast, the compiler has to recognize that both the
    div and mod should use reciprocal mul (standard), and if possible, only
    do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Fri Sep 18 16:48:48 2026
    From Newsgroup: comp.arch

    Michael S wrote:
    On Fri, 18 Sep 2026 00:38:39 -0700
    Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had
    to work in symmetrical base-5, using '-' and '=' for -1 and -2,
    along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding,
    there is no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.


    May be, numbers were big? May be, bigger than what fits in 64 bits, or
    even in 128?


    See my reply to Tim, I was off by three orders of magnitude since each conversion ran in 4-5 ns. :-)

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Fri Sep 18 15:36:15 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had to
    work in symmetrical base-5, using '-' and '=' for -1 and -2, along
    with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding, there is
    no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.

    Oops!

    I meant 4-5 ns not microseconds.

    3 orders of magnitude does make a small difference...

    Just a little... :)

    Just like decimal, it is a lot easier to convert ascii to binary than
    binary to ascii, the difference between div/mod every digit and a
    single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    In order to be really fast, the compiler has to recognize that both
    the div and mod should use reciprocal mul (standard), and if possible,
    only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Sun Sep 20 07:14:04 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/15/2026 7:49 AM, Tim Rentsch wrote:
    [...]
    I'm sorry you took my comments so seriously. They were very much
    meant tongue in cheek.

    I did assume your comments weren't serious, and I apologize if my
    "down the rabbit hole" didn't make that clear. :-(

    Unfortunately it didn't. If you guessed my comments weren't meant
    to be taken seriously I don't understand why you gave what looked
    like a serious response.

    That said, let me address your comments.

    Big snip. Thank you. I found your comments interesting. My only disagreement is that there are a fair number of instances of at least
    English being written top to bottom, usually in signs organized
    vertically for space reasons. But there, even if they have multiple
    lines, are not "continuous" from one line to the next. I think most
    people have no trouble reading them

    I expect such signs are usually written in all capitals, and also
    in a sans-serif font, both of which makes them easier to read in a top-to-bottom format.

    (My apologies for posting comments that have nothing to do with
    computer architecture.)

    Yes. This will be my last post on this topic. I hope you have no
    hard feelings.

    To me it looks like you want to have it both ways. That's
    disappointing.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 22 19:51:14 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you had to >>>> work in symmetrical base-5, using '-' and '=' for -1 and -2, along
    with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, but
    they did of course not need to be. With symmetrical encoding, there is >>>> no extra minus sign in front, just a negative first digit.

    The fast solutions did each input line/number in about 4-5 us, then
    you had to return the (decimal) sum, before converting that back to
    the same base-5 system.

    I'm wondering why a seemingly simple problem took as much time as
    this. Maybe I'm misunderstanding the problem statement. Is there
    one (base negative 5) input number per line? Surely converting such
    an input to binary can be done more quickly than 4 microseconds.

    Oops!

    I meant 4-5 ns not microseconds.

    3 orders of magnitude does make a small difference...

    Just a little... :)

    Just like decimal, it is a lot easier to convert ascii to binary than
    binary to ascii, the difference between div/mod every digit and a
    single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a slice of
    bytes into a Rust String, the rest is simply what it seems like for a C programmer.


    In order to be really fast, the compiler has to recognize that both
    the div and mod should use reciprocal mul (standard), and if possible,
    only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.


    No, it was actually doing two IMULs by the reciprocal, the multipy by 5
    is just a regular LEA reg,[reg+reg*4] which only takes a cycle.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 22 18:16:35 2026
    From Newsgroup: comp.arch


    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    -----------------------
    In order to be really fast, the compiler has to recognize that both
    the div and mod should use reciprocal mul (standard), and if possible,
    only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    What is wrong with the standard way::

    quotient = n / BASE; // 1 divide
    remainder = n % BASE; // no additional instruction

    ??

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 22 20:27:19 2026
    From Newsgroup: comp.arch

    MitchAlsup wrote:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    -----------------------
    In order to be really fast, the compiler has to recognize that both
    the div and mod should use reciprocal mul (standard), and if possible,
    only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    What is wrong with the standard way::

    quotient = n / BASE; // 1 divide
    remainder = n % BASE; // no additional instruction

    This is the way I typically wrote it, unless I went the further step of back-multiply and subtract, but most ocmpilers would see both DIV and
    MOD and only do it once.

    However, when MUL is at least 3-5 x faster than DIV, then reciprocal
    MUL, back-multiply and SUB is about twice as fast as a single DIV.

    The latest Apple (ARM) hardware, like the M4, has gotten a much, much
    faster DIV, it is so fast that using reciprocal mul, back-mul, sub is significantly slower, so clang does not do that any longer.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 22 16:55:50 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    -----------------------

    In order to be really fast, the compiler has to recognize that both
    the div and mod should use reciprocal mul (standard), and if possible,
    only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should work for
    both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    What is wrong with the standard way::

    quotient = n / BASE; // 1 divide
    remainder = n % BASE; // no additional instruction

    ??

    I didn't mean to suggest a way to write the code. I was trying to
    explain how the compiler might have transformed the code, so as to
    need two multiplies rather than one.

    As it turns out, since I posted the earlier message, I have looked
    into this matter a bit more, and this proposed explanation isn't
    right. I think I now understand why there were two multiplies,
    but I haven't investigated that further because of possible
    differences between C and Rust.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Wed Sep 23 08:28:30 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you
    had to work in symmetrical base-5, using '-' and '=' for -1 and
    -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive,
    but they did of course not need to be. With symmetrical encoding,
    there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary
    than binary to ascii, the difference between div/mod every digit
    and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after
    giving r the value of n % BASE, r is used to index an array. That
    doesn't work if the value of r is negative (to be clear, taking
    the code as C code). So something is off.

    In order to be really fast, the compiler has to recognize that
    both the div and mod should use reciprocal mul (standard), and
    if possible, only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should
    work for both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.

    I realized later this idea is wrong. The two IMULs are a result
    of the two numerators being different, in one case being just n,
    and in the other case being n-<something>. An IMUL is needed for
    each of the two numerators (with one being implied for n%BASE).

    No, it was actually doing two IMULs by the reciprocal, the multipy
    by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
    cycle.

    Here is a way of converting that needs only one mul, although it
    does have an additional multiply-by-five LEA in the first loop
    (the name UL is a typedef for unsigned long):

    char *
    string_for_balanced_base_5( long x, char *p ){
    UL u = x < 0 ? -x : x;
    UL v = 0;
    UL k = -1;

    do k++, v = v*5 + 2; while( v < u );
    v += x;

    *--p = 0;
    do {
    *--p = "=-012"[ v%5 ];
    v /= 5;
    } while( k-- > 0 );

    return p;
    }
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 23 18:01:25 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you
    had to work in symmetrical base-5, using '-' and '=' for -1 and
    -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive,
    but they did of course not need to be. With symmetrical encoding, >>>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary
    than binary to ascii, the difference between div/mod every digit
    and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    As my original comments stated, in both Perl and Rust n % BASE will
    always return a positive remainder when the BASE is positive.

    I.e not like C where I believe this used to be implementation-defined.

    giving r the value of n % BASE, r is used to index an array. That
    doesn't work if the value of r is negative (to be clear, taking
    the code as C code). So something is off.

    In order to be really fast, the compiler has to recognize that
    both the div and mod should use reciprocal mul (standard), and
    if possible, only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should
    work for both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.

    I realized later this idea is wrong. The two IMULs are a result
    of the two numerators being different, in one case being just n,
    and in the other case being n-<something>. An IMUL is needed for
    each of the two numerators (with one being implied for n%BASE).

    Right!

    I know that both divisions should return the same result, with the
    second intentionally have zero remainder, but that does leave the door
    open for an optimized version. Thanks!

    No, it was actually doing two IMULs by the reciprocal, the multipy
    by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
    cycle.

    Here is a way of converting that needs only one mul, although it
    does have an additional multiply-by-five LEA in the first loop
    (the name UL is a typedef for unsigned long):

    char *
    string_for_balanced_base_5( long x, char *p ){
    UL u = x < 0 ? -x : x;
    UL v = 0;
    UL k = -1;

    do k++, v = v*5 + 2; while( v < u );
    v += x;

    *--p = 0;
    do {
    *--p = "=-012"[ v%5 ];
    v /= 5;
    } while( k-- > 0 );

    return p;
    }

    I'll see if this translates, and if it is faster...

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Wed Sep 23 10:47:09 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you
    had to work in symmetrical base-5, using '-' and '=' for -1 and
    -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive,
    but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    As my original comments stated, in both Perl and Rust n % BASE will
    always return a positive remainder when the BASE is positive.

    The original post did have a comment saying that the remainder
    would be positive (i.e., non-negative), but it didn't mention
    either Perl or Rust. It's because of this kind of possible
    incongruities that I asked for a C version.

    I.e not like C where I believe this used to be implementation-defined.

    In the original C standard division and remainder were indeed implementation-defined for negative operands. That was changed
    in C99.

    giving r the value of n % BASE, r is used to index an array. That
    doesn't work if the value of r is negative (to be clear, taking
    the code as C code). So something is off.

    In order to be really fast, the compiler has to recognize that
    both the div and mod should use reciprocal mul (standard), and
    if possible, only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should
    work for both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.

    I realized later this idea is wrong. The two IMULs are a result
    of the two numerators being different, in one case being just n,
    and in the other case being n-<something>. An IMUL is needed for
    each of the two numerators (with one being implied for n%BASE).

    Right!

    I know that both divisions should return the same result, with the
    second intentionally have zero remainder, but that does leave the door
    open for an optimized version. Thanks!

    The point is that the two divisions can give different results, even
    when both operands are positive -- one rounds down, the other might
    round up.

    No, it was actually doing two IMULs by the reciprocal, the multipy
    by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
    cycle.

    Here is a way of converting that needs only one mul, although it
    does have an additional multiply-by-five LEA in the first loop
    (the name UL is a typedef for unsigned long):

    char *
    string_for_balanced_base_5( long x, char *p ){
    UL u = x < 0 ? -x : x;
    UL v = 0;
    UL k = -1;

    do k++, v = v*5 + 2; while( v < u );
    v += x;

    *--p = 0;
    do {
    *--p = "=-012"[ v%5 ];
    v /= 5;
    } while( k-- > 0 );

    return p;
    }

    I'll see if this translates, and if it is faster...

    Please tweak it up to improve any platform-specific weaknesses
    that might be there. I'm sure you are better at doing that than
    I am (plus you are running on hardware with different performance characteristics).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 23 18:41:03 2026
    From Newsgroup: comp.arch


    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you
    had to work in symmetrical base-5, using '-' and '=' for -1 and
    -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive,
    but they did of course not need to be. With symmetrical encoding, >>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary
    than binary to ascii, the difference between div/mod every digit
    and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part
    of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    An obvious case where n should be unsigned not signed.

    giving r the value of n % BASE, r is used to index an array. That
    doesn't work if the value of r is negative (to be clear, taking
    the code as C code). So something is off.

    In order to be really fast, the compiler has to recognize that
    both the div and mod should use reciprocal mul (standard), and
    if possible, only do it once.
    ...
    Checking Godbolt:

    Turns out it is actually generating two IMULs, but it should
    work for both positive and negative inputs...

    I suspect the two multiplys are a consequence of the usual way for
    producing a quotient and remainder:

    quotient = n / BASE;
    remainder = n - quotient*BASE

    So one of the multiplys is doing a divide by BASE, and the other
    multiply is doing a multiply by BASE. That is plausible anyway.

    I realized later this idea is wrong. The two IMULs are a result
    of the two numerators being different, in one case being just n,
    and in the other case being n-<something>. An IMUL is needed for
    each of the two numerators (with one being implied for n%BASE).

    No, it was actually doing two IMULs by the reciprocal, the multipy
    by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
    cycle.

    Here is a way of converting that needs only one mul, although it
    does have an additional multiply-by-five LEA in the first loop
    (the name UL is a typedef for unsigned long):

    char *
    string_for_balanced_base_5( long x, char *p ){
    UL u = x < 0 ? -x : x;
    UL v = 0;
    UL k = -1;

    do k++, v = v*5 + 2; while( v < u );
    v += x;

    *--p = 0;
    do {
    *--p = "=-012"[ v%5 ];
    v /= 5;
    } while( k-- > 0 );

    return p;
    }
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Wed Sep 23 13:15:06 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you
    had to work in symmetrical base-5, using '-' and '=' for -1 and
    -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive,
    but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];

    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    An obvious case where n should be unsigned not signed.

    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Thu Sep 24 15:09:40 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you >>>>>>>> had to work in symmetrical base-5, using '-' and '=' for -1 and >>>>>>>> -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, >>>>>>>> but they did of course not need to be. With symmetrical encoding, >>>>>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary >>>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>>> and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part >>>>>>> of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, >>>>>> 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; >>>>>>
    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string() >>>>>> }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    An obvious case where n should be unsigned not signed.

    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.


    When using a balanced encoding like I was asked to use here, any digit
    can be negative, for both positive and negative inputs: Only the first6
    digit must have the same sign as the input value to be converted.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 17:00:40 2026
    From Newsgroup: comp.arch


    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [...]

    There was an Advent of Code puzzle in 2022 (last day) where you >>>>>>> had to work in symmetrical base-5, using '-' and '=' for -1 and >>>>>>> -2, along with '0', '1', '2'.

    You had to parse 113 such numbers, all happened to be positive, >>>>>>> but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
    digit.

    [...]

    Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.

    Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?

    Sure:

    const BASE:i64 = 5;
    const VALUE:[i8;256] = [
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, >>>>> 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
    0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; >>>>>
    const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];

    fn tobase(n:i64) -> String
    {
    let mut n = n;
    let mut buf:[u8;32] = [0;32];
    let mut p = buf.len();
    loop {
    let r = n % BASE; // positive remainder for positive $BASE
    let d = DIGITS[r as usize];
    n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
    p -= 1;
    buf[p] = d;
    if n == 0 {break;}
    }
    std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
    }

    I'm guessing this code is in Rust. Can I ask you to translate it
    to C for me? Mostly I can guess at the meaning, but as I am only
    a novice with Rust I worry that my guessing would be wrong in some
    cases, and wrong functionality would result.

    The only non-C part here is the final library call to convert a
    slice of bytes into a Rust String, the rest is simply what it seems
    like for a C programmer.

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative. But after

    An obvious case where n should be unsigned not signed.

    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.

    In my own code, 90%-odd of integer variables are unsigned. So,
    perhaps, over the decades, I have developed a way of coding that
    works better in unsigned space than in signed space. {{Almost
    every field in almost every CPU control register is inherently
    unsigned, as are the use of these as memory indexes, method-
    call indexes, and switch variables.}}

    I got to the point where if a negative value is not absolutely
    necessary, the variable is simply unsigned.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 24 18:35:35 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:



    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.

    In my own code, 90%-odd of integer variables are unsigned. So,
    perhaps, over the decades, I have developed a way of coding that
    works better in unsigned space than in signed space. {{Almost
    every field in almost every CPU control register is inherently
    unsigned, as are the use of these as memory indexes, method-
    call indexes, and switch variables.}}

    I got to the point where if a negative value is not absolutely
    necessary, the variable is simply unsigned.

    That's true in my code as well[*]. The only place where FP is
    used is when emulating ARM64 floating point instructions, and only
    then if the host FP instructions will provide identical results to similar ARM64
    FP instructions.

    [*] Although I prefer not to use unadorned 'unsigned' - I use
    the stdint.h types (uint64_t, uint32_t, uint16_t and uint8_t).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Fri Sep 25 02:47:10 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    [converting a signed value to a balanced base 5 representation]

    An obvious case where n should be unsigned not signed.

    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.

    When using a balanced encoding like I was asked to use here, any
    digit can be negative, for both positive and negative inputs:
    Only the first digit must have the same sign as the input value
    to be converted.

    I get that. But to represent all the values in the range
    we must deal with the possibility that other digits may
    also be negative.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Fri Sep 25 03:19:50 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> posted:

    [considering a function to convert a signed argument, n, to
    a balanced base 5 representation]

    It looks like this statement cannot be right. In C, if n is
    negative and not a multiple of 5, then n%5 is negative.

    An obvious case where n should be unsigned not signed.

    Yes and no. It is possible to solve this problem where the
    divisions and remainders are always done on non-negative values
    (and so could be unsigned rather than signed). I did in fact
    code up a solution that has this property. But such code can be
    harder to write and harder to understand. In contrast, it's
    fairly easy to solve this problem using signed quantities and
    simply deals with the negative remainders appropriately. Which
    approach is better? As usual that can depend on other factors,
    including performance.

    First I misunderstood what you were suggesting; sorry about that.

    The parameter n cannot be unsigned, because the purpose of the
    function to produce a representation of a range of values, with
    roughly half of them being negative.

    In my own code, 90%-odd of integer variables are unsigned. So,
    perhaps, over the decades, I have developed a way of coding that
    works better in unsigned space than in signed space. {{Almost
    every field in almost every CPU control register is inherently
    unsigned, as are the use of these as memory indexes, method-
    call indexes, and switch variables.}}

    I got to the point where if a negative value is not absolutely
    necessary, the variable is simply unsigned.

    I get that. I'm pretty much the same in terms of preferring
    unsigned types for integer values, and see similar ratios.

    The point here is that this problem is inherently signed in at
    least one aspect, because what is being converted is a signed
    value. There is no point in using a balanced base encoding if all
    the values being represented are non-negative.

    Given that the input argument will be signed, the question is how
    should the code be written to carry out the conversion? It can be
    written using signed variables, like what Terje did. It also can be
    written using unsigned variables, like the example function that I
    posted. Which one is better? I wouldn't choose an unsigned-only
    writing just because it doesn't used signed types (not counting the
    signed argument). In most cases I give a higher weight to which
    writing is easier to understand and easier to see how it works.
    Other things being equal, I prefer using unsigned types. Here I am
    not yet convinced that other factors are indeed equal.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 26 07:41:04 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my
    attempt at introducing them into the standard was passed by J3
    (the US standards body which does most of the work) but not taken
    up by WG5, the overarching ISO standards body. I went ahead and
    implemented it in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 27 02:48:24 2026
    From Newsgroup: comp.arch

    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a
    difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Robert Finch@robfi680@gmail.com to comp.arch on Sat Sep 26 23:49:18 2026
    From Newsgroup: comp.arch

    On 2026-09-26 10:48 p.m., Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    I have a version of TinyBasic that only supports floats. Floats are all
    you need Efye

    Signed vs unsigned is just a way of expressing commonly used number
    sets. I wonder if it is worth it to define a "prime" number type
    indicating the var only holds prime numbers. Types are shortcuts
    representing the rules used to form numbers. May be useful sometimes to
    have a way to define the rules for the numbers?

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Sun Sep 27 16:56:35 2026
    From Newsgroup: comp.arch

    On 27/09/2026 04:48, Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    My university degree is in mathematics and computer science - there is
    not a difference in thinking here. Concrete types in a programming
    language model subsets of mathematical objects, and different types do
    so in different ways, obeying different rules and restrictions, and
    useful for different things. Sometimes you want modular arithmetic.
    Sometimes you want to pretend that your types don't have limits.
    Sometimes you want bitfields, or to be thinking in terms of bits and
    bytes. Sometimes you are interested in "higher level" mathematical properties, sometimes you are interested in efficient implementations.

    Older languages tended to have a more limited range of types - in very
    old C, everything was an int or a char, and in TCL everything is a
    string. Later compiled languages (like later versions of C) had lots of different integer types for different use-cases. (Some of these are
    muddled, unfortunately.) Still later higher level languages have fewer
    again - Python has only one integer type that grows as needed.

    If you are using a language that has limited type support, you make do
    with what you have.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 27 17:47:48 2026
    From Newsgroup: comp.arch


    Robert Finch <robfi680@gmail.com> posted:

    On 2026-09-26 10:48 p.m., Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps? Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    I have a version of TinyBasic that only supports floats. Floats are all
    you need Efye

    As proven by APL !! even labels are floating point.

    Signed vs unsigned is just a way of expressing commonly used number
    sets.

    Cardinal versus ordinal.

    I wonder if it is worth it to define a "prime" number type
    indicating the var only holds prime numbers. Types are shortcuts representing the rules used to form numbers. May be useful sometimes to
    have a way to define the rules for the numbers?

    We can't even agree on detection of overflow (or not:: mostly)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Sun Sep 27 18:21:36 2026
    From Newsgroup: comp.arch

    According to MitchAlsup <user5857@newsgrouper.org.invalid>:
    I have a version of TinyBasic that only supports floats. Floats are all
    you need Efye

    As proven by APL !! even labels are floating point.

    You misspelled JOSS. Or maybe FOCAL.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 27 23:57:20 2026
    From Newsgroup: comp.arch

    On Sun, 27 Sep 2026 06:14:24 -0400, Robert Finch wrote:

    Writing code to support number types seems non-mathematical. But I
    suppose it is just a different way of expressing mathematical rules.

    Rules are code.

    Testing for a prime could just be a lookup table if the prime < 16
    bits.

    The primality of the first one million positive integers can be easily
    and efficiently encoded into (i.e. fast to do lookups on) a block of
    data of about 33K bytes in size.

    I remember going to an early talk by Stephen Wolfram where he was
    introducing this whizzy new maths program called rCLMathematicarCY; he mentioned that they used a Cray super to generate the table. Nowadays
    you can do it with a bit of Python code on your own PC.

    (My own program took a bit under 9 minutes to find all 78498 primes.)

    And then there are the Gaussian primes.

    >>> a = 3 + 2j
    >>> a * a.conjugate() == 13
    True
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Sun Sep 27 20:01:49 2026
    From Newsgroup: comp.arch

    On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    If you go back to the early to mid 1960s, in the mainframe era, at least
    the IBM S/360 series and the Univac 1100 series only had signed
    numbers. I don't know about the other mainframe architectures of the
    day (the BNCH). Of the most popular languages of the day, Fortran
    didn't support unsigned integers, and in COBOL, while you could
    essentially specify unsigned in the Picture clause (by not including a
    leading "S"), I am not sure what happened if you did. So unsigned was a relatively recent innovation.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 28 06:14:01 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> schrieb:
    On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a
    difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    If you go back to the early to mid 1960s, in the mainframe era, at least
    the IBM S/360 series and the Univac 1100 series only had signed
    numbers.

    I don't know about Univac, but the S/360 certainly supported unsigned arithmetic. They even had two versions of add instructions, which seems
    strange for a two's complement machine, but they set flags differently,
    of which the /360 had too few.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 28 10:38:30 2026
    From Newsgroup: comp.arch

    On 27/09/2026 12:14, Robert Finch wrote:

    rCLTinyrCYBasic also supports strings and local variables. There are floating-point instructions in hardware so it was easy to modify the
    parser from ints to floats. Floats are 96-bit triple precision decimal floats.


    Most "home computers" of the 1980's had BASIC interpreters integral to
    the system, with support for strings and numbers that were
    integer/floating point hybrids. On the ZX Spectrum these were 3 bytes,
    IIRC, where one byte was the exponent and sign, the other two were the mantissa. One exponent value was used to indicate that the mantissa
    word was an integer.

    Exceptions include the BBC Micro where you could also have integer
    variables that were significantly faster, and of course the Jupiter Ace
    that had Forth instead of BASIC as its built-in language.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 08:50:56 2026
    From Newsgroup: comp.arch

    David Brown <david.brown@hesbynett.no> writes:
    On 27/09/2026 12:14, Robert Finch wrote:

    rCLTinyrCYBasic also supports strings and local variables. There are
    floating-point instructions in hardware so it was easy to modify the
    parser from ints to floats. Floats are 96-bit triple precision decimal
    floats.


    Most "home computers" of the 1980's had BASIC interpreters integral to
    the system, with support for strings and numbers that were
    integer/floating point hybrids. On the ZX Spectrum these were 3 bytes, >IIRC, where one byte was the exponent and sign, the other two were the >mantissa. One exponent value was used to indicate that the mantissa
    word was an integer.

    Exceptions include the BBC Micro where you could also have integer
    variables that were significantly faster,

    Microsoft Basic (as experienced on the C64 by me) has no integer/FP
    hybrids. It has FP variables (no suffix), integer variables (IIRC
    suffix: %), and string variables (suffix: $). I have rarely seen code
    that used the integer variables, and have not used them much myself,
    for no reason I remember. I guess that the performance advantages, if
    any, were small, and there was the memory cost of an additional byte
    per use in the source code.

    Microsoft Basic was pretty widespread on computers designed in the USA
    (e.g., Apple, Commodore, Tandy, IBM), but obviously the UK companies
    (Acorn, Sinclair) preferred to roll their own.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 28 11:23:41 2026
    From Newsgroup: comp.arch

    On 28/09/2026 10:50, Anton Ertl wrote:
    David Brown <david.brown@hesbynett.no> writes:
    On 27/09/2026 12:14, Robert Finch wrote:

    rCLTinyrCYBasic also supports strings and local variables. There are
    floating-point instructions in hardware so it was easy to modify the
    parser from ints to floats. Floats are 96-bit triple precision decimal
    floats.


    Most "home computers" of the 1980's had BASIC interpreters integral to
    the system, with support for strings and numbers that were
    integer/floating point hybrids. On the ZX Spectrum these were 3 bytes,
    IIRC, where one byte was the exponent and sign, the other two were the
    mantissa. One exponent value was used to indicate that the mantissa
    word was an integer.

    Exceptions include the BBC Micro where you could also have integer
    variables that were significantly faster,

    Microsoft Basic (as experienced on the C64 by me) has no integer/FP
    hybrids. It has FP variables (no suffix), integer variables (IIRC
    suffix: %), and string variables (suffix: $). I have rarely seen code
    that used the integer variables, and have not used them much myself,
    for no reason I remember. I guess that the performance advantages, if
    any, were small, and there was the memory cost of an additional byte
    per use in the source code.


    BBC basic used the same suffixes. I think the integer variables were
    quite common in programming - certainly I used them a lot myself. There
    were also some special ones - A%, X% and Y% - that were used as a
    connection between BASIC and machine code. (BBC had a built-in
    assembler, making it a lot easier for assembly programming than many
    home computers.)

    Microsoft Basic was pretty widespread on computers designed in the USA
    (e.g., Apple, Commodore, Tandy, IBM), but obviously the UK companies
    (Acorn, Sinclair) preferred to roll their own.


    That could well be - my experience at the time was very biased!

    The TI-99/4A - an American computer that was fairly popular in the UK at
    that time - had its own version of BASIC. But I have no idea if it had
    any kind of "integer shortcut" in its numerical variables - it was
    certainly extraordinarily slow (an order of magnitude at least, compared
    to the ZX Spectrum, or BBC - despite being fast for ROM cartridge games).

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Mon Sep 28 14:33:28 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro wrote:
    On Sun, 27 Sep 2026 06:14:24 -0400, Robert Finch wrote:

    Writing code to support number types seems non-mathematical. But I
    suppose it is just a different way of expressing mathematical rules.

    Rules are code.

    Testing for a prime could just be a lookup table if the prime < 16
    bits.

    The primality of the first one million positive integers can be easily
    and efficiently encoded into (i.e. fast to do lookups on) a block of
    data of about 33K bytes in size.
    Encode it mod 30, one bit per possible prime location, so 33 kB.
    This is the method I used in my last Sieve program, which I wrote almost 3 decades ago.

    I remember going to an early talk by Stephen Wolfram where he was
    introducing this whizzy new maths program called |ore4+oMathematica|ore4-Y; he
    mentioned that they used a Cray super to generate the table. Nowadays
    you can do it with a bit of Python code on your own PC.

    (My own program took a bit under 9 minutes to find all 78498 primes.)
    That's stupidly slow!
    I just ressurrected that old (Feb 1998) 32-bit C code and ran the
    original binary.
    I had to go to 1E9 just to get something that took more than a second: h:\terjem\c2\PRIME>SIEVE0.EXE 1000000000
    Searching prime numbers to : 1000000000
    50847534 prime numbers foundn 4 secs.
    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:09:27 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> writes:
    On 9/14/26 1:32 PM, Thomas Koenig wrote:
    Paul Clayton <paaronclayton@gmail.com> schrieb:
    On 9/8/26 3:30 PM, Thomas Koenig wrote:
    [snip]
    Power's
    addg6s is also available for this purpose (but just generates
    the zero / sixes for the correction).

    It seems BCD arithmetic might have been a little simpler if BCD
    (and ASCII) had placed the numbers to be binary 5 to 15.

    Would this have made BCD easier, and if so, how?

    I was thinking that such would simplify the carry handling
    between digits, partially inspired by the previously mentioned
    addg6s and the explanation in the Power ISA documentation that a
    preceding addition of 0x6666_6666 was used. (By the way, it
    should have been 6 to 15, i.e., aligning the decimal digit to
    most significant edge of the binary nibble)

    My conception (such as it was) was that each digit would
    naturally carry into the next (where for the standard low value
    encoding, adding anything less than 7 to a 9 digit would not
    carry into the next nibble/digit). I failed to recognize that
    although a carry-in or carry-generate within a nibble would
    always generate a carry-out when appropriate (if I am thinking
    correctly now) it would not properly wrap the digit around 6
    but rather around 0.

    The Burroughs B3500 handled BCD addition and subtraction
    from most-significant-digit to least-significant digit
    (which made detecting overflow possible before storing
    the first result digit).

    It would zero extend the shorter operand (length between 1
    and 100 digits) to the length of the longer operand and
    start adding from the MSD rather than the normal
    schoolkid paper methods that start with the LSD. The
    algorithm wouldn't start writing result digits
    by counting 9's until there was no chance of overflow.

    https://bitsavers.org/pdf/burroughs/MediumSystems/B2500_B3500/1025475_B2500_B3500_RefMan_Oct69.pdf

    Flowchart on page 51.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:13:42 2026
    From Newsgroup: comp.arch

    John Levine <johnl@taugh.com> writes:
    According to MitchAlsup <user5857@newsgrouper.org.invalid>:
    I have a version of TinyBasic that only supports floats. Floats are all >>> you need Efye

    As proven by APL !! even labels are floating point.

    You misspelled JOSS. Or maybe FOCAL.

    One might categorize FOCAL labels as fixed point rather
    than floating point.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:19:36 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a
    difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    If you go back to the early to mid 1960s, in the mainframe era, at least
    the IBM S/360 series and the Univac 1100 series only had signed
    numbers. I don't know about the other mainframe architectures of the
    day (the BNCH).

    The B3500 had four data types:
    UA (Unsigned Alphanumeric)
    UN (Unsigned Numeric)
    SN (Signed Numeric)
    IA (Indirect Address)

    Financial code generally used fixed point[*] SN,
    OS and utility code used UN (or UA since the
    arithmetic instructions operated on UA data by
    ignoring the zone digit).

    There was a sort of kludge that allowed UA data
    to be signed by changing the zone digit in the
    most significant byte of the data.

    [*] The PIC clause defined the decimal point location;
    the UA, UN or SN data in memory had no concept of decimal point.

    There was a parallel floating point implementation (100 digit
    mantissa with two digit signed mantissa), but it was not
    widely used and was removed from the B4700 and successors.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 28 08:45:39 2026
    From Newsgroup: comp.arch

    On 9/27/2026 11:14 PM, Thomas Koenig wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> schrieb:
    On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
    On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    In my own code, 90%-odd of integer variables are unsigned.

    Standard Fortran has ony signed integers. Unfortunately, my attempt
    at introducing them into the standard was passed by J3 (the US
    standards body which does most of the work) but not taken up by WG5,
    the overarching ISO standards body. I went ahead and implemented it
    in gfortran anyway, in cooperation with flang.

    So, in Fortran code, the existing use of unsigned integers is
    extremely low. I use it privately, so it is not zero :-)

    Good for you. ;)

    IrCOm trying to understand the mindset that believes that signed
    integers are all you need, you donrCOt need unsigned ones; is this a
    difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    If you go back to the early to mid 1960s, in the mainframe era, at least
    the IBM S/360 series and the Univac 1100 series only had signed
    numbers.

    I don't know about Univac, but the S/360 certainly supported unsigned arithmetic.

    You may be right. However, I relied on the S/360 Principles of
    Operation manual, where the section on fixed point arithmetic (page 24)
    says

    "In general, both operands are signed and 32 bits long."

    They even had two versions of add instructions, which seems
    strange for a two's complement machine, but they set flags differently,
    of which the /360 had too few.


    I guess you are referring to the Add Logical, and similar. In my brief
    time as a S/360 assembler programmer, I never had occasion to use them,
    and never figured out what they were intended for. But the description
    of "regular" Add, etc. make it clear that it used signed arithmetic.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Mon Sep 28 13:05:39 2026
    From Newsgroup: comp.arch

    Robert Finch <robfi680@gmail.com> writes:

    On 2026-09-26 10:48 p.m., Lawrence D'Oliveiro wrote:
    [...]
    I'm trying to understand the mindset that believes that signed
    integers are all you need, you don't need unsigned ones; is this a
    difference between mathematicians and computer scientists, perhaps?
    Computer scientists are accustomed to thinking in terms of bitfields,
    where sign-extension is often a nuisance; mathematicians are not.

    I have a version of TinyBasic that only supports floats. Floats are
    all you need

    Sometimes true, but certainly not always true.

    Signed vs unsigned is just a way of expressing commonly used number
    sets.

    Not so. Besides the value sets being different, what the various
    operations do is different for different kinds of numbers. The
    expression 3-5 gives a different result than 3u-5u. The
    expression 5/3*3 gives a different result from the expression
    5./3.*3.

    I wonder if it is worth it to define a "prime" number type
    indicating the var only holds prime numbers. Types are shortcuts representing the rules used to form numbers. [...]

    You're conflating the aspect of representation with the idea of a
    value predicate. In languages like C, the type of an arithmetic
    variable is used mainly to convey information about representation,
    and not about dynamic properties such as being prime. In other
    cases types, as for example struct types and union types, can carry
    more information than just representation. The notion of "type" in
    computer languages is very different from the idea of "type" in
    mathematics. It's a mistake to think of these different kinds of
    "type" as being equivalent.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Mon Sep 28 22:50:26 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 08:50:56 GMT, Anton Ertl wrote:

    Microsoft Basic (as experienced on the C64 by me) has no integer/FP
    hybrids. It has FP variables (no suffix), integer variables (IIRC
    suffix: %), and string variables (suffix: $). I have rarely seen
    code that used the integer variables, and have not used them much
    myself, for no reason I remember.

    How do you define bitmask operations on floats?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Tue Sep 29 03:15:47 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 06:14:01 -0000 (UTC), Thomas Koenig wrote:

    ... the S/360 certainly supported unsigned arithmetic. They even had
    two versions of add instructions, which seems strange for a two's
    complement machine, but they set flags differently, of which the
    /360 had too few.

    rCLOverflowrCY versus rCLcarry/borrowrCY? (CanrCOt think of anything else.)

    Even the lowly PDP-11 could afford to have separate flags for those
    two functions. ;)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Mon Sep 28 14:37:34 2026
    From Newsgroup: comp.arch

    MitchAlsup [2026-09-24 17:00:40] wrote:
    In my own code, 90%-odd of integer variables are unsigned. So,
    perhaps, over the decades, I have developed a way of coding that
    works better in unsigned space than in signed space. {{Almost
    every field in almost every CPU control register is inherently
    unsigned, as are the use of these as memory indexes, method-
    call indexes, and switch variables.}}

    IME, in C both unsigned and signed integers work well (in the sense that
    I'm able to write code that works without too many pitfalls), but things
    become delicate when you mix the two. So I usually care more about
    using the same kind of integers as is used in the surrounding code.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 29 06:11:25 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Mon, 28 Sep 2026 08:50:56 GMT, Anton Ertl wrote:

    Microsoft Basic (as experienced on the C64 by me) has no integer/FP
    hybrids. It has FP variables (no suffix), integer variables (IIRC
    suffix: %), and string variables (suffix: $). I have rarely seen
    code that used the integer variables, and have not used them much
    myself, for no reason I remember.

    How do you define bitmask operations on floats?

    Does Microsoft Basic of the 1970s and early 1980s have bitmask
    operations? If it had, my guess is that the numbers were converted to
    integer, the operation performed, and then converted back to FP.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Tue Sep 29 22:47:35 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 14:37:34 -0400, Stefan Monnier wrote:

    IME, in C both unsigned and signed integers work well (in the sense
    that I'm able to write code that works without too many pitfalls),
    but things become delicate when you mix the two. So I usually care
    more about using the same kind of integers as is used in the
    surrounding code.

    ThatrCOs usually the case. Nobody can be bothered to remember the
    precise details of the integer-promotion rules. ;)

    Now, can anybody solve the mystery of why the Java language designers
    thought it would be a good idea to leave out unsigned integers?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 30 08:39:45 2026
    From Newsgroup: comp.arch

    On 30/09/2026 00:47, Lawrence DrCOOliveiro wrote:
    On Mon, 28 Sep 2026 14:37:34 -0400, Stefan Monnier wrote:

    IME, in C both unsigned and signed integers work well (in the sense
    that I'm able to write code that works without too many pitfalls),
    but things become delicate when you mix the two. So I usually care
    more about using the same kind of integers as is used in the
    surrounding code.

    ThatrCOs usually the case. Nobody can be bothered to remember the
    precise details of the integer-promotion rules. ;)

    Now, can anybody solve the mystery of why the Java language designers
    thought it would be a good idea to leave out unsigned integers?

    Unsigned integers in C have two important use-cases - they are needed if
    you want wrapping semantics (it's extremely rarely that it's actually
    useful, but some people think they need it all the time), and they are appropriate for hardware register interaction and low-level
    bit-twiddling. The rest of the time, you could just as well use a
    signed integer type as long as the range is big enough. (You might
    still /prefer/ unsigned types for some C code, but it is not actually necessary.)

    In Java, signed integer types wrap, so you don't need a special wrapping
    type. And Java does not really target low-level coding (though they
    tried for a bit). So maybe they thought it was unnecessary and
    confusing? I have virtually no Java experience, so this is just
    speculation.

    (For comparison, I believe in Zig unsigned types are not wrapping either
    - arithmetic overflow is UB. If you want wrapping arithmetic, you have
    to use the specific "add-modulo" style operators that puts the behaviour
    where it really belongs, on the operations and not the type.)

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From aph@aph@littlepinkcloud.invalid to comp.arch on Wed Sep 30 08:52:11 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:

    Now, can anybody solve the mystery of why the Java language designers
    thought it would be a good idea to leave out unsigned integers?

    It's unnecessary, and the designer(s) wanted to get away from the zoo of
    C's integer types.

    With wrapping signed arithmetic, all of the semantics of unsigned
    arithmetic is the same, except for right shift and division
    operations. Unsigned right shift has its own operator (>>>) and
    unsigned division is rare enough to be handed off to a method call.
    Bytecodes are precious, and it would have been hard to justify
    unsigned division.

    Andrew.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 30 22:26:55 2026
    From Newsgroup: comp.arch

    On Wed, 30 Sep 2026 08:52:11 +0000, aph wrote:

    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:

    Now, can anybody solve the mystery of why the Java language
    designers thought it would be a good idea to leave out unsigned
    integers?

    It's unnecessary ...

    I once had to write a glue program to interface a clientrCOs online shop
    system to their payment processor. The payment processor provided a
    blob of Java code for talking to their server, so I had to write a
    Java wrapper around that code which acted as an intermediary between
    that and our shop server back-end.

    Things worked mostly well ... except every week or two, a payment
    would fail to go through.

    I finally tracked it down to the encoding/decoding of a length field
    in a protocol packet. The shift/mask operations were getting screwed
    up by sign extension every time a byte value exceeded 127.

    Adding extra code to mask out the extended sign bits fixed the
    problem.

    But now you see why other programming languages keep the unsigned
    integers. It just saves a certain amount of irritation and occasional
    outright pain.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Oct 1 07:34:11 2026
    From Newsgroup: comp.arch

    aph@littlepinkcloud.invalid writes:
    With wrapping signed arithmetic, all of the semantics of unsigned
    arithmetic is the same, except for right shift and division
    operations.

    The most common operations in my experience where there is a
    difference between signed and unsigned are different are unequality comparisons. Apparently the Java designers did not think that
    comparing for unsigned <, >, <=, >= is something the Java programmers
    would miss, and they probably are right.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From aph@aph@littlepinkcloud.invalid to comp.arch on Thu Oct 1 10:44:32 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
    On Wed, 30 Sep 2026 08:52:11 +0000, aph wrote:

    Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:

    Now, can anybody solve the mystery of why the Java language
    designers thought it would be a good idea to leave out unsigned
    integers?

    It's unnecessary ...

    I once had to write a glue program to interface a clientrCOs online shop system to their payment processor. The payment processor provided a
    blob of Java code for talking to their server, so I had to write a
    Java wrapper around that code which acted as an intermediary between
    that and our shop server back-end.

    Things worked mostly well ... except every week or two, a payment
    would fail to go through.

    I finally tracked it down to the encoding/decoding of a length field
    in a protocol packet. The shift/mask operations were getting screwed
    up by sign extension every time a byte value exceeded 127.

    Adding extra code to mask out the extended sign bits fixed the
    problem.

    But now you see why other programming languages keep the unsigned
    integers. It just saves a certain amount of irritation and
    occasional outright pain.

    Some languages do, some don't. In this case, it's a matter of knowing
    the language: in C you have to use unsigned char because char might be
    signed or unsiged. In C, there are *three* byte types. I'm not aware
    of any language not derived from C which does that. Whatever you do
    with bytes, someone will make mistakes because it's not what they
    expect.

    I guess the java compiler itself could have done the masking, but the
    question is whether it was worth complexifying the language by
    replicating C's zoo of numeric types.

    Andrew.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Thu Oct 1 07:00:50 2026
    From Newsgroup: comp.arch

    aph@littlepinkcloud.invalid writes:

    Lawrence D?Oliveiro <ldo@nz.invalid> wrote:

    Now, can anybody solve the mystery of why the Java language designers
    thought it would be a good idea to leave out unsigned integers?

    It's unnecessary, and the designer(s) wanted to get away from the zoo of
    C's integer types.

    With wrapping signed arithmetic, all of the semantics of unsigned
    arithmetic is the same, except for right shift and division
    operations.

    Also left shift and relational comparisons.
    --- Synchronet 3.22a-Linux NewsLink 1.2