On Wed, 03 Jun 2026 13:54:01 +0000, Thomas Koenig wrote:
It causes problems with badly-written software.
I don't see that as a fault of big-endian.
Think of the 3 important numberings of memory elements in a machine >architecture:
1. Ordering of numbering of bytes within a multibyte quantity
2. Ordering of numbering of bits within a byte
3. Ordering of place values for binary digits within an integer (of
whatever length)
Only little-endian can keep all 3 consistent. With big-endian, you
inevitably end up with situations like the ordering 3 being 7 minus
ordering 2, or some even more complicated inversion. If you try to
keep 2 and 3 consistent, then you lose a simple relationship between
bit and byte numbers, like the byte number being the bit number
right-shifted by 3.
On Sat, 5 Sep 2026 08:17:34 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Think of the 3 important numberings of memory elements in a machine
architecture:
1. Ordering of numbering of bytes within a multibyte quantity
2. Ordering of numbering of bits within a byte
3. Ordering of place values for binary digits within an integer (of
whatever length)
Only little-endian can keep all 3 consistent. With big-endian, you
inevitably end up with situations like the ordering 3 being 7 minus
ordering 2, or some even more complicated inversion. If you try to
keep 2 and 3 consistent, then you lose a simple relationship between
bit and byte numbers, like the byte number being the bit number
right-shifted by 3.
Keeping 1 and 2 consistent is important.
3, of course, can only be little-endian. But consistency with 3 is
only a stylistic preference; one might think it nice that the bit representing 2^0 be bit 0, but it's not important for any actual
purpose. It isn't in any way involved in actual computer arithmetic,
and so it doesn't save any cycles.
My argument for big-endian is that only big-endian can keep 1, 2, and
4 consistent, where 4 is:
4. Ordering of numbering of digits within a number represented in
character form.
Then you can have 1, 2, 4, and 5 consistent... where 5 is
5. Ordering of numbering of digits within a packed decimal quantity.
Then:
A number in character form is written in the same direction as a
number in packed decimal form, so they can be easily converted.
A number in packed decimal form is oriented in the same way as an
integer in binary form, so they can share registers and ALUs.
Plus, now numbers are written in the computers memory, laid out in
locations of increasing addresses, the same way we would write them on
a piece of paper, leading to less confusion in writing programs.
Big endian is just as dead now as bytes that are not 8-bits wide.
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits wide.
I have login to several machines which are big-endian (and we receive
bug reports about them for gfortran).
I do not have access to any machine where bytes are not eight bits :-)
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits wide.
I have login to several machines which are big-endian (and we receive
bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new architecture, please just forget about this particular issue.
I do not have access to any machine where bytes are not eight bits :-)
My very first personal piece of programming was on a Unisys 1100 with
36-bit words and either 6 or 9-bit bytes.
The program was my own arbitrary precision library which I used to
calculate as many digits of pi as I could manage inside a minute (the student max) of CPU runtime.
I did not know that I had 36 bits, only that the docs said there was
room for 10-11 decimal digits, so I wrote my code with lots and lots of modulo operations, over 10-digit chunks and temporary 20-digit multiplication results.
I'm guessing I used the same 72-bit temp variables instead of carries,
or the classic "is the sum smaller than either input, then it must have overflowed" approach.
Terje
I have login to several machines which are big-endian (and we
receive bug reports about them for gfortran).
I do not have access to any machine where bytes are not eight bits
:-)
Fun fact: even on big-endian architectures, registers behave as though
they were little-endian.
Do not forget:: remember that LE won and any new computer architecture
shall be LE. That fact that one can make a coherent BE machine is what
should be forgotten.
On Sat, 5 Sep 2026 21:16:17 -0000 (UTC), Lawrence DrCOOliveiro wrote:
Fun fact: even on big-endian architectures, registers behave as
though they were little-endian.
I didn't know that registers behaved as if they had *any* endianness
at all.
At least that was my immediate reaction. But then I remembered that
before the PDP-11 came along, people were building lots of computers
just fine without even thinking of the possibility of a consistent little-endian machine. So I guess it is possible.
But I don't think it's at all likely, unless all of us switch to
speaking Arabic.
A number in packed decimal form is oriented in the same way as an
integer in binary form, so they can share registers and ALUs.
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits wide.
I have login to several machines which are big-endian (and we receive
bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer architecture
shall be LE. That fact that one can make a coherent BE machine is what
should be forgotten.
On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:
A number in packed decimal form is oriented in the same way as an
integer in binary form, so they can share registers and ALUs.
But that limits the precision with which you can do basic addition and >subtraction of decimal numbers, unless you go backwards in memory.
In what order did the IBM 1401 and 1620 order the digits? Because I
believe they could do arbitrary-precision addition and subtraction.
On Sat, 5 Sep 2026 21:16:17 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Fun fact: even on big-endian architectures, registers behave as though
they were little-endian.
I didn't know that registers behaved as if they had *any* endianness
at all. Unless you're talking about registers for short vectors - like
MMX, AltiVec, AVX - the issue of endianness doesn't even arise.
Although there's a *sort* of endianness involving registers, but--- Synchronet 3.22a-Linux NewsLink 1.2
integer registers and floating-point registers do it the opposite way.
John Savard
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits wide.
I have login to several machines which are big-endian (and we receive
bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer architecture
shall be LE. That fact that one can make a coherent BE machine is what
should be forgotten.
One might also remember that there are external protocols
that will forever require big-endian ordering (IP/TCP et alia).
I don't consider that sufficient reason for a CPU to implement
native arithmetic using big-endian containers, however. While
ARMv8 included the capability to provide big-endian support,
in the architecture, it hasn't been widely implemented and may
in the future be eliminated.
It is quite common to have instructions that operate on partial
register contents, is it not?
Consider a sequence like this (for rCLmoverCY, read rCLloadrCY or rCLstorerCY as
appropriate, and for rCLwordrCY read rCLsome convenient multi-byte >quantityrCY):
move.word A, B
move.byte B, C
On typical modern (i.e. RISC) architectures, A and C might be memory >locations and B is a register, or vice versa.
The question is, which byte of B ends up in C -- is it the
most-significant or least-significant byte?
On a little-endian architecture, itrCOs always the least-significant byte.
On a big-endian architecture, it depends on whether B is a register or
memory -- least-significant for register, most-significant for memory.
Endianness has nothing to do with writing direction.
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence DrCOOliveiro wrote:
Endianness has nothing to do with writing direction.
Not directly.
But while the Arabs write their language from right to left, when
writing a number, they put the most significant digit on the left,
and the units digit on the right. Just like we do. But now the least significant digit comes _first_ in the ordinary direction of
reading.
So the layout of numbers in memory in a little-endian computer seems
natural and intuitive to someone whose language is Arabic - and
strange and confusing to someone whose language is English.
Can usually be addressed well enough if the ISA has byte swap
instructions. It is a little bit of a pain if it has to be done with
shifts and masks, but in this case a byte swap does well enough.
Could debate though whether to provide all of the byte swap cases
(say, 5 unique instruction), or simply an 8 byte swap to be followed
by a right shift as-needed.
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
But as it happens, the digits we use to write numbers in--- Synchronet 3.22a-Linux NewsLink 1.2
decimal notation are often referred to as "Arabic numerals". So it
should not be surprising to learn that the Arabs also write numbers
using decimal place-value notation. We changed the shapes of the
digits we got from them, though.
But while the Arabs write their language from right to left, when
writing a number, they put the most significant digit on the left, and
the units digit on the right. Just like we do. But now the least
significant digit comes _first_ in the ordinary direction of reading.
So the layout of numbers in memory in a little-endian computer seems
natural and intuitive to someone whose language is Arabic - and
strange and confusing to someone whose language is English.
Before IBM came up with the consistently big-endian IBM System/360,
their biggest-selling computer was the IBM 1401.
This computer was designed for commercial data processing, not
scientific work. Thus, it did all its arithmetic in decimal.
Not only that, because it stored data in six-bit characters, with one
extra bit for a word mark, it didn't even pack the decimal before
doing arithmetic. It just did arithmetic on character strings.
So the manual for the 1401 had illustrations of how a payroll record
might look on punched cards or in a computer's memory. Someone's name, someone's employee number, someone's monthly or weekly salary. All
written in the same direction.
On a 360, this is only changed slightly. The numbers get squashed,
with two digits in every eight-bit character cell (byte).
John Savard
One has...
load long
load
load halfword
load byte
On Sun, 6 Sep 2026 12:41:36 -0500, BGB wrote:
Can usually be addressed well enough if the ISA has byte swap
instructions. It is a little bit of a pain if it has to be done with
shifts and masks, but in this case a byte swap does well enough.
Could debate though whether to provide all of the byte swap cases
(say, 5 unique instruction), or simply an 8 byte swap to be followed
by a right shift as-needed.
ItrCOs only a small step from there to full-on rCLswizzlingrCY, which is a common part of SIMD instruction sets, isnrCOt it.
Is it true that x86 instructions containing integer literals put them
in big-endian order?
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
While that's true, registers aren't memory. A register is an
available point for calculating with numbers. Memory is where
numbers are written.
Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:
On Sun, 6 Sep 2026 12:41:36 -0500, BGB wrote:
Can usually be addressed well enough if the ISA has byte swap
instructions. It is a little bit of a pain if it has to be done with
shifts and masks, but in this case a byte swap does well enough.
Could debate though whether to provide all of the byte swap cases
(say, 5 unique instruction), or simply an 8 byte swap to be followed
by a right shift as-needed.
ItrCOs only a small step from there to full-on rCLswizzlingrCY, which is a >> common part of SIMD instruction sets, isnrCOt it.
Permute can rearrange all 8-bytes of a register in any order desired.
Is it true that x86 instructions containing integer literals put them
in big-endian order?
Since 8088 was a byte sized buss and read things in use order, I would bet against that.
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to use long scrolls.-a :-)
But ISTM that the choice of horizontal direction is independent of the choice of vertical direction, thus yielding four potential choices.
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
On 9/6/2026 5:31 PM, Stephen Fuld wrote:
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to
use long scrolls.-a :-)
But ISTM that the choice of horizontal direction is independent of the
choice of vertical direction, thus yielding four potential choices.
Sorry for the self followup, but actually eight choices.-a For each of
the four corners, proceed in the horizontal direction or the vertical direction.
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
But as it happens, the digits we use to write numbers in
decimal notation are often referred to as "Arabic numerals". So it
should not be surprising to learn that the Arabs also write numbers
using decimal place-value notation. We changed the shapes of the
digits we got from them, though.
But while the Arabs write their language from right to left, when
writing a number, they put the most significant digit on the left, and
the units digit on the right. Just like we do. But now the least
significant digit comes _first_ in the ordinary direction of reading.
So the layout of numbers in memory in a little-endian computer seems
natural and intuitive to someone whose language is Arabic - and
strange and confusing to someone whose language is English.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits wide.
I have login to several machines which are big-endian (and we receive
bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer architecture
shall be LE. That fact that one can make a coherent BE machine is what
should be forgotten.
One might also remember that there are external protocols
that will forever require big-endian ordering (IP/TCP et alia).
On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup wrote:
Humans developed 3 writing orders {Left-Right, Right->left, and Top->bottom}. I wonder why nobody standardized on bottom->top ??
Another one: rCLboustrophedonicrCY -- rightraAleft and leftraAright on alternate lines.
Used for example on Easter Island/Rapa Nui. (Not that anybody knows
how to read it ...)
On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:
A number in packed decimal form is oriented in the same way as an
integer in binary form, so they can share registers and ALUs.
But that limits the precision with which you can do basic addition and subtraction of decimal numbers, unless you go backwards in memory.
In what order did the IBM 1401 and 1620 order the digits? Because I
believe they could do arbitrary-precision addition and subtraction.
Scott Lurndal wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits
wide.
I have login to several machines which are big-endian (and we
receive bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer
architecture shall be LE. That fact that one can make a coherent
BE machine is what should be forgotten.
One might also remember that there are external protocols
that will forever require big-endian ordering (IP/TCP et alia).
These protocols are a sufficient reason to provide BSWAP type
capability in all CPU architectures. (Including BE ones since there
are other LE protocols/encodings).
Terje
On 9/6/2026 5:31 PM, Stephen Fuld wrote:
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to
use long scrolls. :-)
But ISTM that the choice of horizontal direction is independent of
the choice of vertical direction, thus yielding four potential
choices.
Sorry for the self followup, but actually eight choices. For each of
the four corners, proceed in the horizontal direction or the vertical direction.
Endianness only matters on byte-addressable machines, which were
still fairly unusual at the time the PDP-11 came along.
Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
On Sat, 05 Sep 2026 13:52:32 GMT, John Savard wrote:
A number in packed decimal form is oriented in the same way as an
integer in binary form, so they can share registers and ALUs.
But that limits the precision with which you can do basic addition and
subtraction of decimal numbers, unless you go backwards in memory.
In what order did the IBM 1401 and 1620 order the digits? Because I
believe they could do arbitrary-precision addition and subtraction.
On IBM 1401 least significant digit was at highest address. Arithmetic hardware was walking towards lower addresses.
Good way to think about 1401 is to consider it as an electronic
accounting machine. Smallest 1401 (with memory size 1400 charactes)
was probably able to do all what contemporary accounting machines
could do in a similar way. Early 1401 programmers frequently worked previously with accounting machines and could build on their past
experience. IBM salesmen could sell 1401 as a replacement
for accounting machines.
So 1401 was a point in evolutinary path from accounting machines to
modern computers. 1401 started with natural decimal memory
addresses. But need for sligtly bigger memory lead to decimal
based but rather unnatural scheme. IIUC slightly later
Honeywell 200 had very similar logical structure, but switched
to binary addresses. IBM 360 replaced character representation
by BCD.
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
The usual convention is to store text strings in memory/files etc in
reading order, not rendering order. This is a key aspect of the
Unicode bidirectional layout algorithm
<http://www.unicode.org/reports/tr9/>.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
And, of course, S/360 had an instruction to "pack" a string of numeric >EBCDIC characters to two, four bit, digits per byte, an unpack
instruction to convert back, and a set of arithmetic instructions that >worked on packed data. These are still all supported on Z series.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/6/2026 5:31 PM, Stephen Fuld wrote:
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to
use long scrolls. :-)
But ISTM that the choice of horizontal direction is independent of
the choice of vertical direction, thus yielding four potential
choices.
Sorry for the self followup, but actually eight choices. For each of
the four corners, proceed in the horizontal direction or the vertical
direction.
How about, for each of the four corners, proceed in a clockwise or anti-clockwise spiral? That's eight more choices.
On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
I should think the answer is obvious. Left and right are fundamentally equivalent; they may be opposites, but there's no reason to prefer one
over the other.
On Mon, 7 Sep 2026 16:58:47 -0700, Stephen Fuld <sfuld@alumni.cmu.edu.invalid> wrote:
And, of course, S/360 had an instruction to "pack" a string of numeric
EBCDIC characters to two, four bit, digits per byte, an unpack
instruction to convert back, and a set of arithmetic instructions that
worked on packed data. These are still all supported on Z series.
I'm willing to accept that he doesn't consider z/Architecture
mainframes to meet his definition of 'normal machines', since although
they are still popular in their niche behind the scenes, they're not ubiquitous in people's homes.
Now we are really going down the rabbit hole. There is an issue with
"spiral" directions. Let's say you start at the upper left corner,
then proceed to the right. When you hit the right edge, you are
going to go "down". But what is the orientation of the letters, i.e.
should you turn the paper 90 degrees to keep reading left to right,
or should the orientation of the letters change so you can now read
top to bottom without reorientating the paper? Both have problems.
Numbers are a normal native element of Arabic text. So while there
would be a fancy switchover of directions if an English text were
quoted in an Arabic document, the same way as if a Hebrew text were
quoted in an English document, what do you mean that Arabs have to
do the same switchover every time a number appears in a document, as
if Arabic numerals were foreign to Arabic?
I'm willing to accept that he doesn't consider z/Architecture
mainframes to meet his definition of 'normal machines', since
although they are still popular in their niche behind the scenes,
they're not ubiquitous in people's homes.
On Mon, 7 Sep 2026 13:56:05 +0200
Terje Mathisen <terje.mathisen@tmsw.no> wrote:
Scott Lurndal wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits
wide.
I have login to several machines which are big-endian (and we
receive bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer
architecture shall be LE. That fact that one can make a coherent
BE machine is what should be forgotten.
One might also remember that there are external protocols
that will forever require big-endian ordering (IP/TCP et alia).
These protocols are a sufficient reason to provide BSWAP type
capability in all CPU architectures. (Including BE ones since there
are other LE protocols/encodings).
Terje
One 2-byte BE field in IP header. Another one 2-byte field in UDP
header. The rest is either non-endian (single-byte or less) or
endian-neutral (IP address, ports, checksums). I didn't look at TCP
header but would think that it's approximately the same.
It does not sound as sufficient reason to have BSWAP.
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
Hebisch) wrote:
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
On 9/7/2026 7:36 PM, John Savard wrote:
On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup
<user5857@newsgrouper.org.invalid> wrote:
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
I should think the answer is obvious. Left and right are fundamentally
equivalent; they may be opposites, but there's no reason to prefer one
over the other.
Perhaps not quite true.-a Since the vast majority of people are right handed, it is easier for them to write left to right, as it means you
don't run the risk of smudging what you have already wrote as left
handers do with left to right order (think old, pre-ball point pens.).
Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:
Is it true that x86 instructions containing integer literals put them
in big-endian order?
Since 8088 was a byte sized buss and read things in use order, I would bet >against that.
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
Hebisch) wrote:
But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >Alpha, IA-64, ARM A64, RISC-V.
And I daresay that IA-32 would not have had it if it had not been in
the 80286. Whether the 80286 would have had it if the 8086 did not
have it is less clear.
On 08/09/2026 05:21, Stephen Fuld wrote:
On 9/7/2026 7:36 PM, John Savard wrote:
On Sun, 06 Sep 2026 23:57:39 GMT, MitchAlsup
<user5857@newsgrouper.org.invalid> wrote:
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
I should think the answer is obvious. Left and right are fundamentally
equivalent; they may be opposites, but there's no reason to prefer one
over the other.
Perhaps not quite true.-a Since the vast majority of people are right
handed, it is easier for them to write left to right, as it means you
don't run the risk of smudging what you have already wrote as left
handers do with left to right order (think old, pre-ball point pens.).
I think if you were chiselling the letters into stone, you'd have the
chisel in your left hand and hammar in your right hand, and it would be easier to see what you are doing if you went right to left.-a So the
manner of writing is also relevant.
Bottom to top (or horizontal lines, starting at the bottom rather than
the top) might also be a convenient choice for carving letters on a monument.-a If the text is short enough, it will save you getting a ladder.
On 9/7/2026 1:24 PM, Waldek Hebisch wrote:
And, of course, S/360 had an instruction to "pack" a string of numeric >EBCDIC characters to two, four bit, digits per byte, an unpack
instruction to convert back, and a set of arithmetic instructions that >worked on packed data. These are still all supported on Z series.
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
Hebisch) wrote:
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
And what do we do about people who want to write on Mobius strips?
Now we are really going down the rabbit hole. There is an issue with "spiral" directions. Let's say you start at the upper left corner, then proceed to the right. When you hit the right edge, you are going to go "down". But what is the orientation of the letters, i.e. should you
turn the paper 90 degrees to keep reading left to right, or should the orientation of the letters change so you can now read top to bottom
without reorientating the paper? Both have problems.
On Tue, 08 Sep 2026 02:33:17 GMT, John Savard wrote:
Numbers are a normal native element of Arabic text. So while there
would be a fancy switchover of directions if an English text were
quoted in an Arabic document, the same way as if a Hebrew text were
quoted in an English document, what do you mean that Arabs have to
do the same switchover every time a number appears in a document, as
if Arabic numerals were foreign to Arabic?
Yup. ThatrCOs why a script like Arabic is called rCLbidirectionalrCY, not >rCLright-to-leftrCY.
On Mon, 7 Sep 2026 13:56:05 +0200
Terje Mathisen <terje.mathisen@tmsw.no> wrote:
Scott Lurndal wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Terje Mathisen <terje.mathisen@tmsw.no> posted:
Thomas Koenig wrote:
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
Big endian is just as dead now as bytes that are not 8-bits
wide.
I have login to several machines which are big-endian (and we
receive bug reports about them for gfortran).
Yeah, I know that they stil exist, but for anyone working on a new
architecture, please just forget about this particular issue.
Do not forget:: remember that LE won and any new computer
architecture shall be LE. That fact that one can make a coherent
BE machine is what should be forgotten.
One might also remember that there are external protocols
that will forever require big-endian ordering (IP/TCP et alia).
These protocols are a sufficient reason to provide BSWAP type
capability in all CPU architectures. (Including BE ones since there
are other LE protocols/encodings).
Terje
One 2-byte BE field in IP header. Another one 2-byte field in UDP
header. The rest is either non-endian (single-byte or less) or
endian-neutral (IP address, ports, checksums). I didn't look at TCP
header but would think that it's approximately the same.
It does not sound as sufficient reason to have BSWAP.
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
Hebisch) wrote:
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
And they're certainly normal machines. They're the ones you need to
run regular Windows programs.
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
Alpha, IA-64, ARM A64, RISC-V.
And I daresay that IA-32 would not have had it if it had not been in
the 80286. Whether the 80286 would have had it if the 8086 did not
have it is less clear.
I'm sure it wouldn't. AAA was extremely slow because nobody used it so nobody cared.
quadibloc@invalid.com (John Savard) writes:
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek >>Hebisch) wrote:
But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
IA-32 has AAA. The AMD64 ISA has not. See
<https://www.felixcloutier.com/x86/aaa>.
No such instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
Alpha, IA-64, ARM A64, RISC-V.
/ERROR "unexpected byte sequence starting at index 572: '\xE2'" while decoding/:
On Tue, 8 Sep 2026 03:46:48 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
On Tue, 08 Sep 2026 02:33:17 GMT, John Savard wrote:
Numbers are a normal native element of Arabic text. So while there
would be a fancy switchover of directions if an English text were
quoted in an Arabic document, the same way as if a Hebrew text were
quoted in an English document, what do you mean that Arabs have to
do the same switchover every time a number appears in a document, as
if Arabic numerals were foreign to Arabic?
Yup. That|o-C-Os why a script like Arabic is called |o-C-Lbidirectional|o-C-Y, not
|o-C-Lright-to-left|o-C-Y.
Called by whom? I mean, the linguistic classification of the Arabic
script would have been determined long before Unicode was invented.
Before dot-matrix printers were invented. Before computers were
invented.
So why would its classification derive from Unicode rules?
I'm beginning to think we are living in different worlds here.
John Savard--- Synchronet 3.22a-Linux NewsLink 1.2
John Savard wrote:
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek Hebisch) wrote:
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
And they're certainly normal machines. They're the ones you need to
run regular Windows programs.
Lots of rarely used instruction went away when AMD defined x64.
Terje
So why would its classification derive from Unicode rules?
On Tue, 08 Sep 2026 17:31:51 GMT, John Savard wrote:
So why would its classification derive from Unicode rules?
No-one said it did. ItrCOs the other way round: Unicode has to cater to
the way actual writing systems work.
According to John Savard <quadibloc@invalid.com>:
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
x86-64 has 32 bit and 64 bit modes. AAA is still there in 32 bit mode
but not in 64 bit mode.
If the Arabic-style digits, unlike the letters of the Arabic
alphabet, are indeed classified as left-to-right characters, so that
they don't behave differently from other characters representing
decimal digits, then the behavior you have described would indeed
happen. But I view that behavior strictly as a consequence of
Unicode.
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >>Alpha, IA-64, ARM A64, RISC-V.
And I daresay that IA-32 would not have had it if it had not been in
the 80286. Whether the 80286 would have had it if the 8086 did not
have it is less clear.
I'm sure it wouldn't. AAA was extremely slow because nobody used it so >nobody cared.
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
IA-32 has AAA. The AMD64 ISA has not. See >><https://www.felixcloutier.com/x86/aaa>.
No such instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC,
Alpha, IA-64, ARM A64, RISC-V.
Your information is off. HP-PA has the DCOR instruction. Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
On Wed, 9 Sep 2026 00:22:37 -0000 (UTC), Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
On Tue, 08 Sep 2026 17:31:51 GMT, John Savard wrote:
So why would its classification derive from Unicode rules?
No-one said it did. ItrCOs the other way round: Unicode has to cater to
the way actual writing systems work.
Before Unicode came along, I thought that Arabs wrote from right to
left, full stop. This included numbers. So, since the most significant
digit was on the left for them, the same way it was for us, I assumed innocently that the least significant digit was the one that they
wrote first and read first.
Which is what would make little-endian byte order seem natural and
intuitive for them - as it matches the way they read and write
numbers.
If the Arabic-style digits, unlike the letters of the Arabic alphabet,
are indeed classified as left-to-right characters, so that they don't
behave differently from other characters representing decimal digits,
then the behavior you have described would indeed happen. But I view
that behavior strictly as a consequence of Unicode. People wouldn't
have waited until the technology of handling mixed-direction
characters was available to build printers for use in the Arab world
if computing had gotten up to an early start there.
John Savard
Terje Mathisen <terje.mathisen@tmsw.no> posted:
John Savard wrote:
On Mon, 7 Sep 2026 20:24:26 -0000 (UTC), antispam@fricas.org (Waldek
Hebisch) wrote:
8086 and 8088 while mainly binary had special instructions
to support BCD arithmetic and arithmetic on string of ASCII digits
(first instruction in 8088 instruction list is AAA, that is
ASCII adjust after addition). But time changed and special support
of this sort is gone from normal machines.
What? If the 8086 had an AAA instruction, then wouldn't all subsequent
x86 and x86-64 architecture machines also have that instruction?
And they're certainly normal machines. They're the ones you need to
run regular Windows programs.
Lots of rarely used instruction went away when AMD defined x64.
We needed the OpCode space.
Yup. That's why a script like Arabic is called bidirectional, not >>right-to-leftCalled by whom?
John Levine <johnl@taugh.com> writes:
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
No such AAA instruction in ARM A32 (nor T32), HPPA, MIPS, SPARC, PowerPC, >>> Alpha, IA-64, ARM A64, RISC-V.
And I daresay that IA-32 would not have had it if it had not been in
the 80286. Whether the 80286 would have had it if the 8086 did not
have it is less clear.
I'm sure it wouldn't. AAA was extremely slow because nobody used it so
nobody cared.
But it seems to me that at the time it was still fashionable to add instructions to "close the semantic gap" and to have slow microcoded instructions instead of sequences of simpler instructions for code
density. They added PUSHA, POPA, ENTER and BOUND in the 80186 <https://www.eeeguide.com/instruction-set-of-80186/>, and LEAVE in the
80286 <https://www.eeeguide.com/80286-instruction-set/>, none of which
are known for their speed. These CPUs were designed when VAX was in
full swing, before RISC became a thing.
... trying to write on a Mobius strip I'm at a loss as to
where to begin. ;)
On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
<tr.17687@z991.linuxsc.com> wrote:
... trying to write on a Mobius strip I'm at a loss as to
where to begin. ;)
You start anywhere and keep going. 8-)
George Neuner <gneuner2@comcast.net> posted:
On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
<tr.17687@z991.linuxsc.com> wrote:
... trying to write on a Mobius strip I'm at a loss as to
where to begin. ;)
You start anywhere and keep going. 8-)
when you run into yourself do you step up or down ??
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
-aPower's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
Of course, if humans had standardized on base 16, this would not
have been an issue. While base 60 simplifies division by 2, 3, 4, 5, 6,
10, and 12, base 16 provides reasonable approximations
for a third (5/16) and a fifth (3/16) without requiring such a
large digit. 16 also has the nice binary quality of being the
square of two squared.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15.
(I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
Of course, if humans had standardized on base 16, this would not
have been an issue. While base 60 simplifies division by 2, 3,
4, 5, 6, 10, and 12, base 16 provides reasonable approximations
for a third (5/16) and a fifth (3/16) without requiring such a
large digit. 16 also has the nice binary quality of being the
square of two squared.
If one uses one's thumb as a place holder, finger counting--
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
While base 60 simplifies division by 2, 3, 4, 5, 6, 10, and 12, base
16 provides reasonable approximations for a third (5/16) and a fifth
(3/16) without requiring such a large digit.
Please don't blame POPA!
This was the final key ASCII opcode that made it possible to write an executable text program, using noting but those 70+ 7-bit character codes that are blessed in the MIME standard because they never need quoting.
Please don't blame POPA!
This was the final key ASCII opcode that made it possible to write an
executable text program, using noting but those 70+ 7-bit character codes
that are blessed in the MIME standard because they never need quoting.
I think John's Concertina would be the ideal ISA to extend with
a "base64" subset where the encoding is careful to ensure all the bytes
are properly preserved when sent via email (bonus points if it tolerates LF<=>CRLF conversions).
On 9/13/2026 11:55 AM, Paul Clayton wrote:As my also late great High School math teacher did:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
|e-aPower's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
Of course, if humans had standardized on base 16, this would not
have been an issue. While base 60 simplifies division by 2, 3, 4, 5,
6, 10, and 12, base 16 provides reasonable approximations
for a third (5/16) and a fifth (3/16) without requiring such a
large digit. 16 also has the nice binary quality of being the
square of two squared.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.|o-L-|
As the late great Tom Lehrer said decades ago, "Base eight is just like
base ten if you don't have any thumbs."
On Sun, 13 Sep 2026 14:55:25 -0400, Paul Clayton wrote:
While base 60 simplifies division by 2, 3, 4, 5, 6, 10, and 12, base
16 provides reasonable approximations for a third (5/16) and a fifth
(3/16) without requiring such a large digit.
Base 30 gives you terminating fractional representations for all the
same divisors as base 60, and powers and combinations thereof.
You only need one occurrence of a prime factor to be able to deal in a
finite fashion with all combinations involving that factor.
Lawrence D=E2=80=99Oliveiro wrote:
You only need one occurrence of a prime factor to be able to deal in aBase 30 is perfect for a fast prime sieve, since there are exactly 8=20 >possible locations for a prime larger than 30.
finite fashion with all combinations involving that factor.
=20
On 9/7/2026 3:23 PM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/6/2026 5:31 PM, Stephen Fuld wrote:
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to
use long scrolls. :-)
But ISTM that the choice of horizontal direction is independent of
the choice of vertical direction, thus yielding four potential
choices.
Sorry for the self followup, but actually eight choices. For each of
the four corners, proceed in the horizontal direction or the vertical
direction.
How about, for each of the four corners, proceed in a clockwise or
anti-clockwise spiral? That's eight more choices.
Now we are really going down the rabbit hole. There is an issue with "spiral" directions. Let's say you start at the upper left corner,
then proceed to the right. When you hit the right edge, you are going
to go "down". But what is the orientation of the letters, i.e. should
you turn the paper 90 degrees to keep reading left to right, or should
the orientation of the letters change so you can now read top to
bottom without reorientating the paper? Both have problems.
If you say you will change the orientation of the paper, what if the
"paper" is actually a bronze plaque mounted on a monument. Kinda hard
to change its orientation. And when you get to the bottom, do you
stand on your head to read it? :-)
On the other hand, say you change the orientation of the letters.
That works for the right side - you just read it top to bottom. But
now what do you do when you get to the bottom? Do you print the
letters upside down, in which case you have the same problem as the
above solution. You could keep the letters "right side up:, but now
the words are spelled backwards. That's awkward.
So I reject the spiral solutions, and other, even more bizarre
ones. But still, eight possibilities is more than enough. :-)
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Lawrence D=E2=80=99Oliveiro wrote:
You only need one occurrence of a prime factor to be able to deal in aBase 30 is perfect for a fast prime sieve, since there are exactly 8=20
finite fashion with all combinations involving that factor.
=20
possible locations for a prime larger than 30.
30 is also pretty close to a power of 2, so you can store it in binary without too much waste. Continuing along this line with further prime factors:
base ld(base)
2 1
* 3 = 6 2.584962500721156
* 5 = 30 4.906890595608519
* 7 = 210 7.714245517666122
* 11 = 2310 11.17367713630342
* 13 = 30030 14.874116854444512
* 17 = 510510 18.96157969569485
510510 looks like another good base for this kind of thing. If it
should be smaller, another good one is 2*3*5*17=510
(ld(510)=8.994...).
However, like for decimal FP I have my doubts that the benefits of
non-binary bases are enough to justify their costs.
As the late great Tom Lehrer said decades ago, "Base eight is just
like base ten if you don't have any thumbs."
On Mon, 07 Sep 2026 15:23:31 -0700, Tim Rentsch
<tr.17687@z991.linuxsc.com> wrote:
... trying to write on a Mobius strip I'm at a loss as to
where to begin. ;)
You start anywhere and keep going. 8-)
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/7/2026 3:23 PM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/6/2026 5:31 PM, Stephen Fuld wrote:
On 9/6/2026 4:57 PM, MitchAlsup wrote:
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
Perhaps it has to do with the problems it would cause when trying to >>>>> use long scrolls. :-)
But ISTM that the choice of horizontal direction is independent of
the choice of vertical direction, thus yielding four potential
choices.
Sorry for the self followup, but actually eight choices. For each of
the four corners, proceed in the horizontal direction or the vertical
direction.
How about, for each of the four corners, proceed in a clockwise or
anti-clockwise spiral? That's eight more choices.
Now we are really going down the rabbit hole. There is an issue with
"spiral" directions. Let's say you start at the upper left corner,
then proceed to the right. When you hit the right edge, you are going
to go "down". But what is the orientation of the letters, i.e. should
you turn the paper 90 degrees to keep reading left to right, or should
the orientation of the letters change so you can now read top to
bottom without reorientating the paper? Both have problems.
If you say you will change the orientation of the paper, what if the
"paper" is actually a bronze plaque mounted on a monument. Kinda hard
to change its orientation. And when you get to the bottom, do you
stand on your head to read it? :-)
On the other hand, say you change the orientation of the letters.
That works for the right side - you just read it top to bottom. But
now what do you do when you get to the bottom? Do you print the
letters upside down, in which case you have the same problem as the
above solution. You could keep the letters "right side up:, but now
the words are spelled backwards. That's awkward.
So I reject the spiral solutions, and other, even more bizarre
ones. But still, eight possibilities is more than enough. :-)
I'm sorry you took my comments so seriously. They were very much
meant tongue in cheek.
That said, let me address your comments.
(My apologies for posting comments that have nothing to do with
computer architecture.)
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
Paul Clayton <paaronclayton@gmail.com> wrote:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
IIUC some digital systems used "excess 3" code, that is started
from 3 and went up to 12. It had advantage that using 4 bit
biary adder you got binary carry exactly when there was decimal
carry. But after that you had to do correction: if there were
no carry subtract 3, if there was carry add 3.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
I read supposedly serious proposal to switch to base 8. Motivation
was smaller and consequently easier to learn multiplication table.
IIRC this proposal was in 19-th century Sweden. But it seems that
they are still using decimal, so probably it did not pass.
Larry Niven has a few science fiction stories where aliens use base
eight. He may have been influenced (however remotely) by the DEC
machines.
Paul Clayton <paaronclayton@gmail.com> wrote:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
IIUC some digital systems used "excess 3" code, that is started
from 3 and went up to 12. It had advantage that using 4 bit
biary adder you got binary carry exactly when there was decimal
carry. But after that you had to do correction: if there were
no carry subtract 3, if there was carry add 3.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
I read supposedly serious proposal to switch to base 8. Motivation
was smaller and consequently easier to learn multiplication table.
IIRC this proposal was in 19-th century Sweden. But it seems that
they are still using decimal, so probably it did not pass.
On 9/16/26 06:41, Waldek Hebisch wrote:There was an Advent of Code puzzle in 2022 (last day) where you had to
Paul Clayton <paaronclayton@gmail.com> wrote:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
-a Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
IIUC some digital systems used "excess 3" code, that is started
from 3 and went up to 12.-a It had advantage that using 4 bit
biary adder you got binary carry exactly when there was decimal
carry.-a But after that you had to do correction: if there were
no carry subtract 3, if there was carry add 3.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.|o-L-|
I read supposedly serious proposal to switch to base 8.-a Motivation
was smaller and consequently easier to learn multiplication table.
IIRC this proposal was in 19-th century Sweden.-a But it seems that
they are still using decimal, so probably it did not pass.
I'd really have liked a system with digit values symmetrical
around zero, maybe base 10, or even better, base 12. This
makes arithmetic delightfully simple.
There was an article about it in New Scientist of April 22,
1982. I loved it.
On Wed, 16 Sep 2026 05:53:24 -0000 (UTC), Thomas Koenig wrote:
Larry Niven has a few science fiction stories where aliens use base
eight. He may have been influenced (however remotely) by the DEC
machines.
All the DEC machines (prior to the PDP-11) had word sizes which were multiples of 3: 12, 18, 36. Even the 16-bit PDP-11 gave meanings to
3-bit fields in its instruction layout that mapped naturally to octal
digits of the integer value of the instruction word.
Was it IBM that popularized hexadecimal?
On 9/16/2026 12:52 AM, Lawrence DrCOOliveiro wrote:
On Wed, 16 Sep 2026 05:53:24 -0000 (UTC), Thomas Koenig wrote:
Larry Niven has a few science fiction stories where aliens use base
eight. He may have been influenced (however remotely) by the DEC
machines.
All the DEC machines (prior to the PDP-11) had word sizes which were
multiples of 3: 12, 18, 36. Even the 16-bit PDP-11 gave meanings to
3-bit fields in its instruction layout that mapped naturally to octal
digits of the integer value of the instruction word.
Was it IBM that popularized hexadecimal?
Yes. Due to the cost of circuitry and memory, Prior to S/360, almost
all systems were either decimal, or binary with a six bit character set.
Six bit characters fit nicely into two octal "digits" of three bits
each (Binary Coded Decimal, or BCD for short, or similar codes). Hence >systems like those you referred to, and most others had word sizes that
were multiples of six, and octal was the defacto standard.
With the S/360, IBM defined the "byte" as eight bits, and "extended" BCD >into EBCDIC. With that, expressing each byte as two four bit "nibbles" >encouraged each nibble to be expressed as a four bit (i.e. hexadecimal) >value.
On 9/16/26 06:41, Waldek Hebisch wrote:
Paul Clayton <paaronclayton@gmail.com> wrote:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15. (I
suspect there were reasonable historical reasons for both EBDIC
and ASCII layouts.)
IIUC some digital systems used "excess 3" code, that is started
from 3 and went up to 12. It had advantage that using 4 bit
biary adder you got binary carry exactly when there was decimal
carry. But after that you had to do correction: if there were
no carry subtract 3, if there was carry add 3.
If one uses one's thumb as a place holder, finger counting
gives four per hand, so base 8 might have been very natural and
base 16 might not have been too awkward.ry|
I read supposedly serious proposal to switch to base 8. Motivation
was smaller and consequently easier to learn multiplication table.
IIRC this proposal was in 19-th century Sweden. But it seems that
they are still using decimal, so probably it did not pass.
I'd really have liked a system with digit values symmetrical
around zero, maybe base 10, or even better, base 12. This
makes arithmetic delightfully simple.
quadibloc@invalid.com (John Savard) posted:
On Sun, 6 Sep 2026 07:04:02 -0000 (UTC), Lawrence
=?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> wrote:
Endianness has nothing to do with writing direction.
Not directly.
Humans developed 3 writing orders {Left-Right, Right->left, and
Top->bottom}. I wonder why nobody standardized on bottom->top ??
There was an Advent of Code puzzle in 2022 (last day) where you had to
work in symmetrical base-5, using '-' and '=' for -1 and -2, along
with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding, there is
no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
Just like decimal, it is a lot easier to convert ascii to binary than
binary to ascii, the difference between div/mod every digit and a
single mul.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you had
to work in symmetrical base-5, using '-' and '=' for -1 and -2,
along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding,
there is no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
I'm wondering why a seemingly simple problem took as much time as
this. Maybe I'm misunderstanding the problem statement. Is there
one (base negative 5) input number per line? Surely converting such
an input to binary can be done more quickly than 4 microseconds.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you had to
work in symmetrical base-5, using '-' and '=' for -1 and -2, along
with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding, there is
no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
I'm wondering why a seemingly simple problem took as much time as
this. Maybe I'm misunderstanding the problem statement. Is there
one (base negative 5) input number per line? Surely converting such
an input to binary can be done more quickly than 4 microseconds.
Just like decimal, it is a lot easier to convert ascii to binary than
binary to ascii, the difference between div/mod every digit and a
single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
On Fri, 18 Sep 2026 00:38:39 -0700
Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you had
to work in symmetrical base-5, using '-' and '=' for -1 and -2,
along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding,
there is no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
I'm wondering why a seemingly simple problem took as much time as
this. Maybe I'm misunderstanding the problem statement. Is there
one (base negative 5) input number per line? Surely converting such
an input to binary can be done more quickly than 4 microseconds.
May be, numbers were big? May be, bigger than what fits in 64 bits, or
even in 128?
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you had to
work in symmetrical base-5, using '-' and '=' for -1 and -2, along
with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding, there is
no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
I'm wondering why a seemingly simple problem took as much time as
this. Maybe I'm misunderstanding the problem statement. Is there
one (base negative 5) input number per line? Surely converting such
an input to binary can be done more quickly than 4 microseconds.
Oops!
I meant 4-5 ns not microseconds.
3 orders of magnitude does make a small difference...
Just like decimal, it is a lot easier to convert ascii to binary than
binary to ascii, the difference between div/mod every digit and a
single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
In order to be really fast, the compiler has to recognize that both
the div and mod should use reciprocal mul (standard), and if possible,
only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should work for
both positive and negative inputs...
On 9/15/2026 7:49 AM, Tim Rentsch wrote:[...]
I'm sorry you took my comments so seriously. They were very much
meant tongue in cheek.
I did assume your comments weren't serious, and I apologize if my
"down the rabbit hole" didn't make that clear. :-(
That said, let me address your comments.
Big snip. Thank you. I found your comments interesting. My only disagreement is that there are a fair number of instances of at least
English being written top to bottom, usually in signs organized
vertically for space reasons. But there, even if they have multiple
lines, are not "continuous" from one line to the next. I think most
people have no trouble reading them
(My apologies for posting comments that have nothing to do with
computer architecture.)
Yes. This will be my last post on this topic. I hope you have no
hard feelings.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you had to >>>> work in symmetrical base-5, using '-' and '=' for -1 and -2, along
with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, but
they did of course not need to be. With symmetrical encoding, there is >>>> no extra minus sign in front, just a negative first digit.
The fast solutions did each input line/number in about 4-5 us, then
you had to return the (decimal) sum, before converting that back to
the same base-5 system.
I'm wondering why a seemingly simple problem took as much time as
this. Maybe I'm misunderstanding the problem statement. Is there
one (base negative 5) input number per line? Surely converting such
an input to binary can be done more quickly than 4 microseconds.
Oops!
I meant 4-5 ns not microseconds.
3 orders of magnitude does make a small difference...
Just a little... :)
Just like decimal, it is a lot easier to convert ascii to binary than
binary to ascii, the difference between div/mod every digit and a
single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
In order to be really fast, the compiler has to recognize that both
the div and mod should use reciprocal mul (standard), and if possible,
only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should work for
both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other
multiply is doing a multiply by BASE. That is plausible anyway.
Terje Mathisen <terje.mathisen@tmsw.no> writes:-----------------------
In order to be really fast, the compiler has to recognize that both
the div and mod should use reciprocal mul (standard), and if possible,
only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should work for
both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other--- Synchronet 3.22a-Linux NewsLink 1.2
multiply is doing a multiply by BASE. That is plausible anyway.
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Terje Mathisen <terje.mathisen@tmsw.no> writes:-----------------------
In order to be really fast, the compiler has to recognize that both
the div and mod should use reciprocal mul (standard), and if possible,
only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should work for
both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
What is wrong with the standard way::
quotient = n / BASE; // 1 divide
remainder = n % BASE; // no additional instruction
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
-----------------------
In order to be really fast, the compiler has to recognize that both
the div and mod should use reciprocal mul (standard), and if possible,
only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should work for
both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
What is wrong with the standard way::
quotient = n / BASE; // 1 divide
remainder = n % BASE; // no additional instruction
??
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you
had to work in symmetrical base-5, using '-' and '=' for -1 and
-2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive,
but they did of course not need to be. With symmetrical encoding,
there is no extra minus sign in front, just a negative first
digit.
Just like decimal, it is a lot easier to convert ascii to binary
than binary to ascii, the difference between div/mod every digit
and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
In order to be really fast, the compiler has to recognize that
both the div and mod should use reciprocal mul (standard), and
if possible, only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should
work for both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other
multiply is doing a multiply by BASE. That is plausible anyway.
No, it was actually doing two IMULs by the reciprocal, the multipy
by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
cycle.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you
had to work in symmetrical base-5, using '-' and '=' for -1 and
-2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive,
but they did of course not need to be. With symmetrical encoding, >>>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary
than binary to ascii, the difference between div/mod every digit
and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
giving r the value of n % BASE, r is used to index an array. That
doesn't work if the value of r is negative (to be clear, taking
the code as C code). So something is off.
In order to be really fast, the compiler has to recognize that
both the div and mod should use reciprocal mul (standard), and
if possible, only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should
work for both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other
multiply is doing a multiply by BASE. That is plausible anyway.
I realized later this idea is wrong. The two IMULs are a result
of the two numerators being different, in one case being just n,
and in the other case being n-<something>. An IMUL is needed for
each of the two numerators (with one being implied for n%BASE).
No, it was actually doing two IMULs by the reciprocal, the multipy
by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
cycle.
Here is a way of converting that needs only one mul, although it
does have an additional multiply-by-five LEA in the first loop
(the name UL is a typedef for unsigned long):
char *
string_for_balanced_base_5( long x, char *p ){
UL u = x < 0 ? -x : x;
UL v = 0;
UL k = -1;
do k++, v = v*5 + 2; while( v < u );
v += x;
*--p = 0;
do {
*--p = "=-012"[ v%5 ];
v /= 5;
} while( k-- > 0 );
return p;
}
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you
had to work in symmetrical base-5, using '-' and '=' for -1 and
-2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive,
but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
As my original comments stated, in both Perl and Rust n % BASE will
always return a positive remainder when the BASE is positive.
I.e not like C where I believe this used to be implementation-defined.
giving r the value of n % BASE, r is used to index an array. That
doesn't work if the value of r is negative (to be clear, taking
the code as C code). So something is off.
In order to be really fast, the compiler has to recognize that
both the div and mod should use reciprocal mul (standard), and
if possible, only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should
work for both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other
multiply is doing a multiply by BASE. That is plausible anyway.
I realized later this idea is wrong. The two IMULs are a result
of the two numerators being different, in one case being just n,
and in the other case being n-<something>. An IMUL is needed for
each of the two numerators (with one being implied for n%BASE).
Right!
I know that both divisions should return the same result, with the
second intentionally have zero remainder, but that does leave the door
open for an optimized version. Thanks!
No, it was actually doing two IMULs by the reciprocal, the multipy
by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
cycle.
Here is a way of converting that needs only one mul, although it
does have an additional multiply-by-five LEA in the first loop
(the name UL is a typedef for unsigned long):
char *
string_for_balanced_base_5( long x, char *p ){
UL u = x < 0 ? -x : x;
UL v = 0;
UL k = -1;
do k++, v = v*5 + 2; while( v < u );
v += x;
*--p = 0;
do {
*--p = "=-012"[ v%5 ];
v /= 5;
} while( k-- > 0 );
return p;
}
I'll see if this translates, and if it is faster...
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you
had to work in symmetrical base-5, using '-' and '=' for -1 and
-2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive,
but they did of course not need to be. With symmetrical encoding, >>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary
than binary to ascii, the difference between div/mod every digit
and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part
of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
giving r the value of n % BASE, r is used to index an array. That--- Synchronet 3.22a-Linux NewsLink 1.2
doesn't work if the value of r is negative (to be clear, taking
the code as C code). So something is off.
In order to be really fast, the compiler has to recognize that
both the div and mod should use reciprocal mul (standard), and
if possible, only do it once.
...
Checking Godbolt:
Turns out it is actually generating two IMULs, but it should
work for both positive and negative inputs...
I suspect the two multiplys are a consequence of the usual way for
producing a quotient and remainder:
quotient = n / BASE;
remainder = n - quotient*BASE
So one of the multiplys is doing a divide by BASE, and the other
multiply is doing a multiply by BASE. That is plausible anyway.
I realized later this idea is wrong. The two IMULs are a result
of the two numerators being different, in one case being just n,
and in the other case being n-<something>. An IMUL is needed for
each of the two numerators (with one being implied for n%BASE).
No, it was actually doing two IMULs by the reciprocal, the multipy
by 5 is just a regular LEA reg,[reg+reg*4] which only takes a
cycle.
Here is a way of converting that needs only one mul, although it
does have an additional multiply-by-five LEA in the first loop
(the name UL is a typedef for unsigned long):
char *
string_for_balanced_base_5( long x, char *p ){
UL u = x < 0 ? -x : x;
UL v = 0;
UL k = -1;
do k++, v = v*5 + 2; while( v < u );
v += x;
*--p = 0;
do {
*--p = "=-012"[ v%5 ];
v /= 5;
} while( k-- > 0 );
return p;
}
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you
had to work in symmetrical base-5, using '-' and '=' for -1 and
-2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive,
but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0];
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
An obvious case where n should be unsigned not signed.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you >>>>>>>> had to work in symmetrical base-5, using '-' and '=' for -1 and >>>>>>>> -2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, >>>>>>>> but they did of course not need to be. With symmetrical encoding, >>>>>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary >>>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>>> and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part >>>>>>> of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, >>>>>> 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; >>>>>>
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string() >>>>>> }
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
An obvious case where n should be unsigned not signed.
Yes and no. It is possible to solve this problem where the
divisions and remainders are always done on non-negative values
(and so could be unsigned rather than signed). I did in fact
code up a solution that has this property. But such code can be
harder to write and harder to understand. In contrast, it's
fairly easy to solve this problem using signed quantities and
simply deals with the negative remainders appropriately. Which
approach is better? As usual that can depend on other factors,
including performance.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[...]
There was an Advent of Code puzzle in 2022 (last day) where you >>>>>>> had to work in symmetrical base-5, using '-' and '=' for -1 and >>>>>>> -2, along with '0', '1', '2'.
You had to parse 113 such numbers, all happened to be positive, >>>>>>> but they did of course not need to be. With symmetrical encoding, >>>>>>> there is no extra minus sign in front, just a negative first
digit.
[...]
Just like decimal, it is a lot easier to convert ascii to binary >>>>>>> than binary to ascii, the difference between div/mod every digit >>>>>>> and a single mul.
Now I'm curious to see how you implemented the binary-to-ascii part >>>>>> of the problem. Do you still have the code around?
Sure:
const BASE:i64 = 5;
const VALUE:[i8;256] = [
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,-1,0,0,0,1,2,0,0,0,0,0,0,0,0,0,0,-2,0,0, >>>>> 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; >>>>>
const DIGITS:[u8;5] = [b'0',b'1',b'2',b'=',b'-'];
fn tobase(n:i64) -> String
{
let mut n = n;
let mut buf:[u8;32] = [0;32];
let mut p = buf.len();
loop {
let r = n % BASE; // positive remainder for positive $BASE
let d = DIGITS[r as usize];
n = (n-VALUE[d as usize] as i64) / BASE; // Exact division!
p -= 1;
buf[p] = d;
if n == 0 {break;}
}
std::str::from_utf8(&buf[p..buf.len()]).unwrap().to_string()
}
I'm guessing this code is in Rust. Can I ask you to translate it
to C for me? Mostly I can guess at the meaning, but as I am only
a novice with Rust I worry that my guessing would be wrong in some
cases, and wrong functionality would result.
The only non-C part here is the final library call to convert a
slice of bytes into a Rust String, the rest is simply what it seems
like for a C programmer.
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative. But after
An obvious case where n should be unsigned not signed.
Yes and no. It is possible to solve this problem where the
divisions and remainders are always done on non-negative values
(and so could be unsigned rather than signed). I did in fact
code up a solution that has this property. But such code can be
harder to write and harder to understand. In contrast, it's
fairly easy to solve this problem using signed quantities and
simply deals with the negative remainders appropriately. Which
approach is better? As usual that can depend on other factors,
including performance.
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
Yes and no. It is possible to solve this problem where the
divisions and remainders are always done on non-negative values
(and so could be unsigned rather than signed). I did in fact
code up a solution that has this property. But such code can be
harder to write and harder to understand. In contrast, it's
fairly easy to solve this problem using signed quantities and
simply deals with the negative remainders appropriately. Which
approach is better? As usual that can depend on other factors,
including performance.
In my own code, 90%-odd of integer variables are unsigned. So,
perhaps, over the decades, I have developed a way of coding that
works better in unsigned space than in signed space. {{Almost
every field in almost every CPU control register is inherently
unsigned, as are the use of these as memory indexes, method-
call indexes, and switch variables.}}
I got to the point where if a negative value is not absolutely
necessary, the variable is simply unsigned.
Tim Rentsch wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
An obvious case where n should be unsigned not signed.
Yes and no. It is possible to solve this problem where the
divisions and remainders are always done on non-negative values
(and so could be unsigned rather than signed). I did in fact
code up a solution that has this property. But such code can be
harder to write and harder to understand. In contrast, it's
fairly easy to solve this problem using signed quantities and
simply deals with the negative remainders appropriately. Which
approach is better? As usual that can depend on other factors,
including performance.
When using a balanced encoding like I was asked to use here, any
digit can be negative, for both positive and negative inputs:
Only the first digit must have the same sign as the input value
to be converted.
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Tim Rentsch <tr.17687@z991.linuxsc.com> posted:
It looks like this statement cannot be right. In C, if n is
negative and not a multiple of 5, then n%5 is negative.
An obvious case where n should be unsigned not signed.
Yes and no. It is possible to solve this problem where the
divisions and remainders are always done on non-negative values
(and so could be unsigned rather than signed). I did in fact
code up a solution that has this property. But such code can be
harder to write and harder to understand. In contrast, it's
fairly easy to solve this problem using signed quantities and
simply deals with the negative remainders appropriately. Which
approach is better? As usual that can depend on other factors,
including performance.
In my own code, 90%-odd of integer variables are unsigned. So,
perhaps, over the decades, I have developed a way of coding that
works better in unsigned space than in signed space. {{Almost
every field in almost every CPU control register is inherently
unsigned, as are the use of these as memory indexes, method-
call indexes, and switch variables.}}
I got to the point where if a negative value is not absolutely
necessary, the variable is simply unsigned.
In my own code, 90%-odd of integer variables are unsigned.
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
On 2026-09-26 10:48 p.m., Lawrence DrCOOliveiro wrote:
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps? Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
I have a version of TinyBasic that only supports floats. Floats are all
you need Efye
Signed vs unsigned is just a way of expressing commonly used number
sets.
I wonder if it is worth it to define a "prime" number type
indicating the var only holds prime numbers. Types are shortcuts representing the rules used to form numbers. May be useful sometimes to
have a way to define the rules for the numbers?
I have a version of TinyBasic that only supports floats. Floats are all
you need Efye
As proven by APL !! even labels are floating point.
Writing code to support number types seems non-mathematical. But I
suppose it is just a different way of expressing mathematical rules.
Testing for a prime could just be a lookup table if the prime < 16
bits.
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a
difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
If you go back to the early to mid 1960s, in the mainframe era, at least
the IBM S/360 series and the Univac 1100 series only had signed
numbers.
rCLTinyrCYBasic also supports strings and local variables. There are floating-point instructions in hardware so it was easy to modify the
parser from ints to floats. Floats are 96-bit triple precision decimal floats.
On 27/09/2026 12:14, Robert Finch wrote:
rCLTinyrCYBasic also supports strings and local variables. There are
floating-point instructions in hardware so it was easy to modify the
parser from ints to floats. Floats are 96-bit triple precision decimal
floats.
Most "home computers" of the 1980's had BASIC interpreters integral to
the system, with support for strings and numbers that were
integer/floating point hybrids. On the ZX Spectrum these were 3 bytes, >IIRC, where one byte was the exponent and sign, the other two were the >mantissa. One exponent value was used to indicate that the mantissa
word was an integer.
Exceptions include the BBC Micro where you could also have integer
variables that were significantly faster,
David Brown <david.brown@hesbynett.no> writes:
On 27/09/2026 12:14, Robert Finch wrote:
rCLTinyrCYBasic also supports strings and local variables. There are
floating-point instructions in hardware so it was easy to modify the
parser from ints to floats. Floats are 96-bit triple precision decimal
floats.
Most "home computers" of the 1980's had BASIC interpreters integral to
the system, with support for strings and numbers that were
integer/floating point hybrids. On the ZX Spectrum these were 3 bytes,
IIRC, where one byte was the exponent and sign, the other two were the
mantissa. One exponent value was used to indicate that the mantissa
word was an integer.
Exceptions include the BBC Micro where you could also have integer
variables that were significantly faster,
Microsoft Basic (as experienced on the C64 by me) has no integer/FP
hybrids. It has FP variables (no suffix), integer variables (IIRC
suffix: %), and string variables (suffix: $). I have rarely seen code
that used the integer variables, and have not used them much myself,
for no reason I remember. I guess that the performance advantages, if
any, were small, and there was the memory cost of an additional byte
per use in the source code.
Microsoft Basic was pretty widespread on computers designed in the USA
(e.g., Apple, Commodore, Tandy, IBM), but obviously the UK companies
(Acorn, Sinclair) preferred to roll their own.
On Sun, 27 Sep 2026 06:14:24 -0400, Robert Finch wrote:Encode it mod 30, one bit per possible prime location, so 33 kB.
Writing code to support number types seems non-mathematical. But I
suppose it is just a different way of expressing mathematical rules.
Rules are code.
Testing for a prime could just be a lookup table if the prime < 16
bits.
The primality of the first one million positive integers can be easily
and efficiently encoded into (i.e. fast to do lookups on) a block of
data of about 33K bytes in size.
I remember going to an early talk by Stephen Wolfram where he wasThat's stupidly slow!
introducing this whizzy new maths program called |ore4+oMathematica|ore4-Y; he
mentioned that they used a Cray super to generate the table. Nowadays
you can do it with a bit of Python code on your own PC.
(My own program took a bit under 9 minutes to find all 78498 primes.)
On 9/14/26 1:32 PM, Thomas Koenig wrote:
Paul Clayton <paaronclayton@gmail.com> schrieb:
On 9/8/26 3:30 PM, Thomas Koenig wrote:
[snip]
Power's
addg6s is also available for this purpose (but just generates
the zero / sixes for the correction).
It seems BCD arithmetic might have been a little simpler if BCD
(and ASCII) had placed the numbers to be binary 5 to 15.
Would this have made BCD easier, and if so, how?
I was thinking that such would simplify the carry handling
between digits, partially inspired by the previously mentioned
addg6s and the explanation in the Power ISA documentation that a
preceding addition of 0x6666_6666 was used. (By the way, it
should have been 6 to 15, i.e., aligning the decimal digit to
most significant edge of the binary nibble)
My conception (such as it was) was that each digit would
naturally carry into the next (where for the standard low value
encoding, adding anything less than 7 to a 9 digit would not
carry into the next nibble/digit). I failed to recognize that
although a carry-in or carry-generate within a nibble would
always generate a carry-out when appropriate (if I am thinking
correctly now) it would not properly wrap the digit around 6
but rather around 0.
According to MitchAlsup <user5857@newsgrouper.org.invalid>:
I have a version of TinyBasic that only supports floats. Floats are all >>> you need Efye
As proven by APL !! even labels are floating point.
You misspelled JOSS. Or maybe FOCAL.
On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a
difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
If you go back to the early to mid 1960s, in the mainframe era, at least
the IBM S/360 series and the Univac 1100 series only had signed
numbers. I don't know about the other mainframe architectures of the
day (the BNCH).
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> schrieb:
On 9/26/2026 7:48 PM, Lawrence DrCOOliveiro wrote:
On Sat, 26 Sep 2026 07:41:04 -0000 (UTC), Thomas Koenig wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
In my own code, 90%-odd of integer variables are unsigned.
Standard Fortran has ony signed integers. Unfortunately, my attempt
at introducing them into the standard was passed by J3 (the US
standards body which does most of the work) but not taken up by WG5,
the overarching ISO standards body. I went ahead and implemented it
in gfortran anyway, in cooperation with flang.
So, in Fortran code, the existing use of unsigned integers is
extremely low. I use it privately, so it is not zero :-)
Good for you. ;)
IrCOm trying to understand the mindset that believes that signed
integers are all you need, you donrCOt need unsigned ones; is this a
difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
If you go back to the early to mid 1960s, in the mainframe era, at least
the IBM S/360 series and the Univac 1100 series only had signed
numbers.
I don't know about Univac, but the S/360 certainly supported unsigned arithmetic.
They even had two versions of add instructions, which seems
strange for a two's complement machine, but they set flags differently,
of which the /360 had too few.
On 2026-09-26 10:48 p.m., Lawrence D'Oliveiro wrote:[...]
I'm trying to understand the mindset that believes that signed
integers are all you need, you don't need unsigned ones; is this a
difference between mathematicians and computer scientists, perhaps?
Computer scientists are accustomed to thinking in terms of bitfields,
where sign-extension is often a nuisance; mathematicians are not.
I have a version of TinyBasic that only supports floats. Floats are
all you need
Signed vs unsigned is just a way of expressing commonly used number
sets.
I wonder if it is worth it to define a "prime" number type
indicating the var only holds prime numbers. Types are shortcuts representing the rules used to form numbers. [...]
Microsoft Basic (as experienced on the C64 by me) has no integer/FP
hybrids. It has FP variables (no suffix), integer variables (IIRC
suffix: %), and string variables (suffix: $). I have rarely seen
code that used the integer variables, and have not used them much
myself, for no reason I remember.
... the S/360 certainly supported unsigned arithmetic. They even had
two versions of add instructions, which seems strange for a two's
complement machine, but they set flags differently, of which the
/360 had too few.
In my own code, 90%-odd of integer variables are unsigned. So,
perhaps, over the decades, I have developed a way of coding that
works better in unsigned space than in signed space. {{Almost
every field in almost every CPU control register is inherently
unsigned, as are the use of these as memory indexes, method-
call indexes, and switch variables.}}
On Mon, 28 Sep 2026 08:50:56 GMT, Anton Ertl wrote:
Microsoft Basic (as experienced on the C64 by me) has no integer/FP
hybrids. It has FP variables (no suffix), integer variables (IIRC
suffix: %), and string variables (suffix: $). I have rarely seen
code that used the integer variables, and have not used them much
myself, for no reason I remember.
How do you define bitmask operations on floats?
IME, in C both unsigned and signed integers work well (in the sense
that I'm able to write code that works without too many pitfalls),
but things become delicate when you mix the two. So I usually care
more about using the same kind of integers as is used in the
surrounding code.
On Mon, 28 Sep 2026 14:37:34 -0400, Stefan Monnier wrote:
IME, in C both unsigned and signed integers work well (in the sense
that I'm able to write code that works without too many pitfalls),
but things become delicate when you mix the two. So I usually care
more about using the same kind of integers as is used in the
surrounding code.
ThatrCOs usually the case. Nobody can be bothered to remember the
precise details of the integer-promotion rules. ;)
Now, can anybody solve the mystery of why the Java language designers
thought it would be a good idea to leave out unsigned integers?
Now, can anybody solve the mystery of why the Java language designers
thought it would be a good idea to leave out unsigned integers?
Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
Now, can anybody solve the mystery of why the Java language
designers thought it would be a good idea to leave out unsigned
integers?
It's unnecessary ...
With wrapping signed arithmetic, all of the semantics of unsigned
arithmetic is the same, except for right shift and division
operations.
On Wed, 30 Sep 2026 08:52:11 +0000, aph wrote:
Lawrence DrCOOliveiro <ldo@nz.invalid> wrote:
Now, can anybody solve the mystery of why the Java language
designers thought it would be a good idea to leave out unsigned
integers?
It's unnecessary ...
I once had to write a glue program to interface a clientrCOs online shop system to their payment processor. The payment processor provided a
blob of Java code for talking to their server, so I had to write a
Java wrapper around that code which acted as an intermediary between
that and our shop server back-end.
Things worked mostly well ... except every week or two, a payment
would fail to go through.
I finally tracked it down to the encoding/decoding of a length field
in a protocol packet. The shift/mask operations were getting screwed
up by sign extension every time a byte value exceeded 127.
Adding extra code to mask out the extended sign bits fixed the
problem.
But now you see why other programming languages keep the unsigned
integers. It just saves a certain amount of irritation and
occasional outright pain.
Lawrence D?Oliveiro <ldo@nz.invalid> wrote:
Now, can anybody solve the mystery of why the Java language designers
thought it would be a good idea to leave out unsigned integers?
It's unnecessary, and the designer(s) wanted to get away from the zoo of
C's integer types.
With wrapping signed arithmetic, all of the semantics of unsigned
arithmetic is the same, except for right shift and division
operations.
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 02:20:42 |
| Calls: | 1,194 |
| Files: | 1,353 |
| D/L today: |
2 files (1,590K bytes) |
| Messages: | 291,157 |