John Levine <johnl@taugh.com> writes:
It appears that MitchAlsup <user5857@newsgrouper.org.invalid> said:
The PDP-6 and -10 could take an interrupt before
each address calculation so if an interrupt arrived in the middle of
an indirect chain, it just started it over when the interrupt returned.
(...)
One time when I was supposed to be doing something else I wrote a little
program that made a longer and longer interrupt chain until the program
stalled, which told me how often the clock interrupted.
For rCLinterrupt chainrCY here we should read rCLindirect chainrCY, I think?
On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:
This infinite indirect capability made PDP-10 loved by people using
them. It also caused problems when the infinite indirect timer "went
off".
I always wondered 1) was there an upper limit to the number of levels
of indirection, and 2) what that did to the interrupt latency ...
On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:
This infinite indirect capability made PDP-10 loved by people using
them. It also caused problems when the infinite indirect timer "went
off".
I always wondered 1) was there an upper limit to the number of levels
of indirection, and 2) what that did to the interrupt latency ...
I had to research this to check my memory, but on the 1100 series, the
timer value was 100 microseconds. If an instruction took longer than
that, it got an invalid operation interrupt/exception. While this may
seem long by today's standards, remember, this was the 1960s and typical instructions on the 1108 took 750 nanoseconds, and peripherals were a
lot slower. In practice, I never found it to be a problem as
indirection was rare and I never saw more than two levels.
On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:
This infinite indirect capability made PDP-10 loved by people using
them. It also caused problems when the infinite indirect timer "went
off".
I always wondered 1) was there an upper limit to the number of levels
of indirection, and 2) what that did to the interrupt latency ...
I had to research this to check my memory, but on the 1100 series, the
timer value was 100 microseconds. If an instruction took longer than
that, it got an invalid operation interrupt/exception. While this may
seem long by today's standards, remember, this was the 1960s and typical >instructions on the 1108 took 750 nanoseconds, and peripherals were a
lot slower. In practice, I never found it to be a problem as
indirection was rare and I never saw more than two levels.
On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:
This infinite indirect capability made PDP-10 loved by people using
them. It also caused problems when the infinite indirect timer "went
off".
I always wondered 1) was there an upper limit to the number of levels
of indirection, and 2) what that did to the interrupt latency ...
I had to research this to check my memory, but on the 1100 series, the
timer value was 100 microseconds. ...
On Tue, 01 Sep 2026 10:07:14 -0700, Andy Valencia wrote:
Certainly this happened with Stanford University's homebrew DBMS,
SPIRES. My father-in-law told me of the immense work needed to free
up the address bits to permit extended addressing. He said when they
did the original code, nobody imagined that they'd ever need those
upper bits for addressing.
The original Apple MacOS went through this in the 1980s.
The 68000-family instruction set allowed for 32-bit addresses, but the original 68000 processor only looked at the bottom 24 bits. So the Mac
system used those top 8 bits for various memory-management purposes.
Even when they brought out the Mac II, the first with a 68020
processor that *did* pay attention to all 32 bits of the address, they
stuck in a stub MMU on the motherboard that removed those top 8 bits
before actually accessing memory.
Then, by about 1990, they introduced models with rCL32-bit cleanrCY ROMs, that didnrCOt make any such odd uses of those top 8 bits internally, and could address more than 16MiB of RAM.
On 9/1/26 5:17 PM, Lawrence DrCOOliveiro wrote:
On Tue, 01 Sep 2026 10:07:14 -0700, Andy Valencia wrote:
Certainly this happened with Stanford University's homebrew DBMS,
SPIRES. My father-in-law told me of the immense work needed to free
up the address bits to permit extended addressing. He said when they
did the original code, nobody imagined that they'd ever need those
upper bits for addressing.
The original Apple MacOS went through this in the 1980s.
The 68000-family instruction set allowed for 32-bit addresses, but the original 68000 processor only looked at the bottom 24 bits. So the Mac system used those top 8 bits for various memory-management purposes.
Even when they brought out the Mac II, the first with a 68020
processor that *did* pay attention to all 32 bits of the address, they stuck in a stub MMU on the motherboard that removed those top 8 bits
before actually accessing memory.
Then, by about 1990, they introduced models with rCL32-bit cleanrCY ROMs, that didnrCOt make any such odd uses of those top 8 bits internally, and could address more than 16MiB of RAM.
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier implementations, knowing that the extra bits would be desired
for other uses to save memory).
This would have introduced overhead for new OSes (having to
configure the bit), but would allow old software to run
correctly while allowing use of the 32-bit address space.
(Determining when software can use 32-bit mode might be
difficult. User pointers in system calls would also need
to be masked, but I think an OS already needs to prevent
OS-privilege access to arbitrary user-provided addresses.)
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Many
consider this a serious architecture design mistake, citing
history where such has caused compatibility headaches. I _feel_
that system software should be able to handle this extra
complexity without great difficulty, but I have never developed
even a task scheduler much less an OS. Since some OSes support
use of Top Byte Ignore, I am guessing this is not a huge
problem.
Interestingly, Stanford MIPS used the extra bits for an address
space number and had a variable length mask. (-2The size of the
process virtual address space is defined by a bit mask in a
special register.
When masking is enabled for normal operations,
a process identifier from another special register is
substituted for the high order bits of the machine address.
These two special registers are accessible only to processes
running in supervisor state. The masking unit also detects
attempts to access outside the legal segment and raises an
exception to the master pipeline control. Although the virtual
address is a full 32 bits, the package constraint only permits
24 address pins. With word addressing, however, this gives an
address space of 64 megabytes.-+, "Design of a High-Performance
VLSI Processor", John Hennessey et al., 1983)
The masking in Stanford MIPS did not allow software flags but
did allow some process isolation even without a permission
table. Such also allowed the ASID to vary in size without
requiring more address bits if some processes were compact.
If I recall correctly, 32-bit ARM provided a means to define the
size of the user address space such that a variable number of
more significant bits were used as system address bits (similar
to the negative addresses are system addresses of some OSes)
with a separate page table. This did not allow software use of
the extra bits, but provided some flexibility in page tables
(the number of levels for the user page table could be decreased
and the entire system address space page table could be shared
rather than duplicating the root "page" in each process) and
PTEs would not need a global bit to indicate preservation across
processes.
(32-bit PowerPC segments were vaguely similar in
allowing 16 segments with separate virtual address spaces. In
theory, this might support somewhat flexible sharing. HP PA-RISC
provided fewer segments but more flexible/complex use. [I think
encoding the segment bits in the least significant bits would
have been better than using the high bits; it would have made
the change to 64-bit simpler and allowed full-space dynamic
segment addresses at the cost of having to encode segment
numbers in the instruction for byte and half-word aligned
pointers and disallowing alignment traps based on such bits.])
Implicit hardware masking of tags can be useful, though I do
wonder if such could be usefully generalized to better support
multiple data in a single load. (32-bit paired loads in AArch64
do this to some degree, but it might be practical to exploit the
alignment network for loads to parse a load into two pieces at
byte granularity. An unaligned address could indicate the
split point.)
With x86 one can perform a full-register-size load and access
sub-sections (for some registers). Providing denser register
storage use has advantages when memory is slow, but using
subregisters complicates renaming (and forwarding, which is
kind of renaming).
(Like CMOV, out-of-order execution
complicates use of subregisters.)
It might be possible
(practical even) to support a restrained use of subregisters
(less restrained than SIMD that defines a single size and
divides the register into lanes of that size), but it seems
likely that such would not be *worthwhile*.
Even without address masking, part of addresses can be used
to hold type information at the cost of sparser use of the
address space.
(This introduces a possible performance
compatibility concern. Software may seek to minimize translation
overhead from sparsity, but that ties software performance to
specifics of address translation (like which bits are used for
table levels). Even using a tag to indicate a leaf pointer might
be useful with a tracing garbage collector.)
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and
user space), it seems sad that programs would have to explicitly
mask/extract small pointers to use those extra bits.
[I better send this before my mind wanders even farther.]--- Synchronet 3.22a-Linux NewsLink 1.2
I have wondered why Motorola did not add a 24-bit address mode ...
This would have introduced overhead for new OSes (having to
configure the bit), but would allow old software to run correctly
while allowing use of the 32-bit address space.
If I recall correctly, 32-bit ARM provided a means to define the
size of the user address space such that a variable number of more significant bits were used as system address bits (similar to the
negative addresses are system addresses of some OSes) with a
separate page table.
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and user
space), it seems sad that programs would have to explicitly
mask/extract small pointers to use those extra bits.
On 9/1/26 5:17 PM, Lawrence DrCOOliveiro wrote:
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software. Many
consider this a serious architecture design mistake, citing
history where such has caused compatibility headaches. I _feel_
that system software should be able to handle this extra
complexity without great difficulty, but I have never developed
even a task scheduler much less an OS. Since some OSes support
use of Top Byte Ignore, I am guessing this is not a huge
problem.
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier >implementations, knowing that the extra bits would be desired
for other uses to save memory).
But since, as you point out, the chip only had 24 address pins, your
question indeed has a great deal of validity.
On Wed, 2 Sep 2026 18:59:58 -0400, Paul Clayton
<paaronclayton@gmail.com> wrote:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier >>implementations, knowing that the extra bits would be desired
for other uses to save memory).
My immediate reaction would be that they didn't do that for the same
reason they didn't provide a 23-bit address mode and a 25-bit address
mode.
But since, as you point out, the chip only had 24 address pins, your
question indeed has a great deal of validity.
On 9/3/26 10:48 AM, John Savard wrote:
But since, as you point out, the chip only had 24 address pins, your
question indeed has a great deal of validity.
23.
The MC68000 had A1-A23. The upper and lower data strobes selected which >byte(s) to access.
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and
user space), it seems sad that programs would have to explicitly
mask/extract small pointers to use those extra bits.
Paul Clayton <paaronclayton@gmail.com> writes:
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software. Many
consider this a serious architecture design mistake, citing
history where such has caused compatibility headaches. I _feel_
that system software should be able to handle this extra
complexity without great difficulty, but I have never developed
even a task scheduler much less an OS. Since some OSes support
use of Top Byte Ignore, I am guessing this is not a huge
problem.
This capability (TBI) is also leveraged by the PAC
extensions (Pointer Authentication).
David Schultz <david.schultz@earthlink.net> writes:
On 9/3/26 10:48 AM, John Savard wrote:
But since, as you point out, the chip only had 24 address pins, your
question indeed has a great deal of validity.
23.
The MC68000 had A1-A23. The upper and lower data strobes selected which >byte(s) to access.
Doesn't that mean that those strobes are effectively A0?
On 2026-Sep-02 18:59, Paul Clayton wrote:
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and
user space), it seems sad that programs would have to explicitly mask/extract small pointers to use those extra bits.
In those prior "24-bit address space" designs the MMU didn't
look at the upper 8 address bits. That is why they had later
software porting problems.
In the current "48-bit address space" MMU's, all 64 address bits are validated.
The 48-bits is a MMU model specfic VA sub-space translation limit
set by the number of levels in the page table.
Other MMU models can have different numbers of levels.
x64 originally had 4 levels of page table (9-9-9-9-12 = 48 bits) and
around 2017 added level 5 supporting a 57-bit virtual address sub-space within the ISA's 64-bit virtual address space.
David Schultz <david.schultz@earthlink.net> writes:
On 9/3/26 10:48 AM, John Savard wrote:
But since, as you point out, the chip only had 24 address pins, your
question indeed has a great deal of validity.
23.
The MC68000 had A1-A23. The upper and lower data strobes selected which
byte(s) to access.
Doesn't that mean that those strobes are effectively A0?
EricP <ThatWouldBeTelling@thevillage.com> posted:
On 2026-Sep-02 18:59, Paul Clayton wrote:
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and
user space), it seems sad that programs would have to explicitly
mask/extract small pointers to use those extra bits.
In those prior "24-bit address space" designs the MMU didn't
look at the upper 8 address bits. That is why they had later
software porting problems.
In the current "48-bit address space" MMU's, all 64 address bits are validated.
The 48-bits is a MMU model specfic VA sub-space translation limit
set by the number of levels in the page table.
I believe it is set by the used width of PA<63..12> in the PTE in
conjunction to the number of address bits that are routed around
the on-die interconnect. Using PA<63..48> as other than 12b'0
is asking for memory aliasing problem in the coherence protocol.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
EricP <ThatWouldBeTelling@thevillage.com> posted:
On 2026-Sep-02 18:59, Paul Clayton wrote:
With many 64-bit systems limiting user space pointers to 47 bits
(one bit of the 48-bit address space differentiating system and
user space), it seems sad that programs would have to explicitly
mask/extract small pointers to use those extra bits.
In those prior "24-bit address space" designs the MMU didn't
look at the upper 8 address bits. That is why they had later
software porting problems.
In the current "48-bit address space" MMU's, all 64 address bits are validated.
The 48-bits is a MMU model specfic VA sub-space translation limit
set by the number of levels in the page table.
I believe it is set by the used width of PA<63..12> in the PTE in >conjunction to the number of address bits that are routed around
the on-die interconnect. Using PA<63..48> as other than 12b'0
is asking for memory aliasing problem in the coherence protocol.
On the VA side, for future compatibility the topmost defined bit is
extended into the unused bits (to wit, sign extended to 64-bits).
On the PA side, unused bits of the PA will be tied to a logical 0
(there is no need for unused wires in the design).
And also the addressing was kept to 26 bits. Why? So that the
complete CPU state on an interrupt could be stored in 32 bits.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks.
Then see what's left to address.
On 2026-Sep-04 21:08, MitchAlsup wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
My knee jerk reaction is 'none' - those are language issues.
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
That is normally left to the high level language because each
language has it own set rules, and those rules can change.
For example, in Ada a routine argument marked as 'out' is considered
an uninitialized variable that must be written before the routine returns, and must be written before it is read.
Ada85 originally did not allow 'out' arguments to be read after writing
as 'out' meant write-only. But that was a stupid restriction which was just inconvenient so they later changed it to read-after-write.
Do you really want to incorporate that into hardware?
And if a pointer has a bit marking it as read-only then what stops
me clearing that bit? Unless you make "pointer" a first class HW type
with its own set of instructions.
But then you also need the ability to bypass any "pointer" restrictions
to handle things like Anton's software defined tagged integer-pointers.
This is the problem with capabilities systems - they balloon very quickly.
As HW tries to take on ALL the capabilities different languages might
ever require, the generic nature of them drags in inefficiencies and
they wind up as "a jack of all trades but a master of none".
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
And what is in that hidden indexed capability table?
To be flexible enough to handle any data type it would be
a pointer to a privileged routine. And there is the ballooning.
Microcode by any other name would smell as sweet.
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks. Then see what's left to address.
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
First make array index checks simple and cheap.You would probably need more than 32 registers for this...
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks. >>
No, there is no difference in register allocation.
Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.
MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
So 1 extra instruction to check each multiply for overflow.
Or one can incorporate the overflow check into the
MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
and return a single wide result.
Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm. The index register is then used with a scaled-index addressing.
For almost all array indexes the cost is a single reg-reg or reg-imm instruction.
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
I have difficulty judging use-after-free errors.
I don't recall ever having had a use-after-free error ever since
I started writing C in 1992 when I switched to developing on WinNT.
But I'm fairly paranoid so I do things like having my own Assert()
routines that throw my own fatal exceptions on errors and remain
in production code, put validity check markers on all heap
allocated objects and assert they are valid before using,
on free zap the validity markers and NULL the pointer.
Consequently if such programming error did happen, my code should
immediately detect it itself and throw a fatal exception.
If programmers just used languages that checked array indexes then
99.999% of memory access errors would disappear. First make array
index checks simple and cheap. This requires checked arithmetic for
the index expression calculation, plus a set of simple
compare-and-fault instructions various bounds checks.
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
No, there is no difference in register allocation.
Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.
MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
So 1 extra instruction to check each multiply for overflow.
Or one can incorporate the overflow check into the
MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
and return a single wide result.
Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm. The index register is then used with a scaled-index addressing.
For almost all array indexes the cost is a single reg-reg or reg-imm instruction
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
I have difficulty judging use-after-free errors.
I don't recall ever having had a use-after-free error ever since
I started writing C in 1992 when I switched to developing on WinNT.
But I'm fairly paranoid so I do things like having my own Assert()
routines that throw my own fatal exceptions on errors and remain
in production code, put validity check markers on all heap
allocated objects and assert they are valid before using,
on free zap the validity markers and NULL the pointer.
Consequently if such programming error did happen, my code should
immediately detect it itself and throw a fatal exception.
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then
you don't need two types of pointers.
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
There are programming languages which do that, Fortran being
a prime example.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Fortran has an interesting take on pointers: You can associate a
pointer with existing variables, but only if they have the TARGET
attribute, otherwise this must be diagnosed (it's a constraint).
You can also have multi-dimensional pointers, so something like
real, dimension(10,10), target :: a
real, dimension(:,:), pointer :: ap
ap => a(1:10:2,1:10:2)
After this, ap(1,2) is an alias to a(1,3), for example. ap is
then pointing to a sub-array. This will be handled with array
descriptors aka dope vectors.
No address arithmetic required, normal indexing will do.
You can also ALLOCATE pointers, after which they are associated with
an anonymous region. But this is usually not the best way because
Fortran also has ALLOCATABLE variables which are automatically
deallocated when they go out of scope (or are passed as actual
argument to an INTENT(OUT) dummy argument).
But these mechanisms are, by nature, rather heavy-weight and
not as well suited to a lower-level language like C.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation, >>> plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
No, there is no difference in register allocation.
Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.
Assume the following C code:
#define N something
...
double *p = calloc(N, sizeof(double));
...
foo (p+2,i);
void foo(double *p, int i)
{
p[i] = 42.
}
This would require something like, generated from the code
above,
void foo(double *p, int i, double *p_from, double *p_to)
{
if (p + i < p_from)
range_error();
if (p + i > p_to)
range_error();
p[i] = 42.;
}
so each range-checked pointer would require three actual arguments.
Hardware could assist this with an instruction "trap if out of range",
which would combine the if statements above into one. This would be
an instruction which could be predicted to be a no-op with some
confidence.
But register (or argument) pressure would be high, which is why
I suspect that 32 registers might not be enough.
MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
So 1 extra instruction to check each multiply for overflow.
Or one can incorporate the overflow check into the
MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
and return a single wide result.
Or Mitch's carry instruction, followed by a bne0.
Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm.
The index register is then used with a scaled-index addressing.
That would require both limits.
For almost all array indexes the cost is a single reg-reg or reg-imm instruction.
... which should be moved outside the loops.
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
I have difficulty judging use-after-free errors.
It seems to be quire common, see https://cwe.mitre.org/data/definitions/416.html .
I don't recall ever having had a use-after-free error ever since
I started writing C in 1992 when I switched to developing on WinNT.
But I'm fairly paranoid so I do things like having my own Assert()
routines that throw my own fatal exceptions on errors and remain
in production code, put validity check markers on all heap
allocated objects and assert they are valid before using,
on free zap the validity markers and NULL the pointer.
Consequently if such programming error did happen, my code should immediately detect it itself and throw a fatal exception.
That sounds like very good practice.
One reason why I like ALLOCATABLE variables in Fortran is that
a large fraction of that burden is taken off the programmer's
shoulders.
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
There are programming languages which do that, Fortran being
a prime example.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Fortran has an interesting take on pointers: You can associate a
pointer with existing variables, but only if they have the TARGET
attribute, otherwise this must be diagnosed (it's a constraint).
You can also have multi-dimensional pointers, so something like
real, dimension(10,10), target :: a
real, dimension(:,:), pointer :: ap
ap => a(1:10:2,1:10:2)
After this, ap(1,2) is an alias to a(1,3), for example. ap is
then pointing to a sub-array. This will be handled with array
descriptors aka dope vectors.
No address arithmetic required, normal indexing will do.
You can also ALLOCATE pointers, after which they are associated with
an anonymous region. But this is usually not the best way because
Fortran also has ALLOCATABLE variables which are automatically
deallocated when they go out of scope (or are passed as actual
argument to an INTENT(OUT) dummy argument).
But these mechanisms are, by nature, rather heavy-weight and
not as well suited to a lower-level language like C.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation, >>> plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
No, there is no difference in register allocation.
Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.
Assume the following C code:
#define N something
...
double *p = calloc(N, sizeof(double));
...
foo (p+2,i);
void foo(double *p, int i)
{
p[i] = 42.
}
This would require something like, generated from the code
above,
void foo(double *p, int i, double *p_from, double *p_to)
{
if (p + i < p_from)
range_error();
if (p + i > p_to)
range_error();
p[i] = 42.;
}
so each range-checked pointer would require three actual arguments.
Hardware could assist this with an instruction "trap if out of range",
which would combine the if statements above into one. This would be
an instruction which could be predicted to be a no-op with some
confidence.
But register (or argument) pressure would be high, which is why
I suspect that 32 registers might not be enough.
MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
So 1 extra instruction to check each multiply for overflow.
Or one can incorporate the overflow check into the
MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
and return a single wide result.
Or Mitch's carry instruction, followed by a bne0.
Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm.
The index register is then used with a scaled-index addressing.
That would require both limits.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an
array.
Thomas Koenig <tkoenig@netcologne.de> posted:
But these mechanisms are, by nature, rather heavy-weight and
not as well suited to a lower-level language like C.
Perhaps the fault is with C than with Fortran !!!
Assume the following C code:
#define N something
...
double *p = calloc(N, sizeof(double));
...
foo (p+2,i);
void foo(double *p, int i)
{
p[i] = 42.
}
This would require something like, generated from the code
above,
void foo(double *p, int i, double *p_from, double *p_to)
{
if (p + i < p_from)
range_error();
if (p + i > p_to)
range_error();
p[i] = 42.;
}
Whereas::
double p[N] = calloc(N, sizeof(double));
...
foo (&p[2],i);
...
void foo(double p[*], int i)
{
p[i] = 42.
}
does not!!
That would require both limits.
CMP Rd,Rnormalixed_index,array_limit
BCIN Rd,where_ever
Does both the <0 and >=array_limit checks in one inst.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then
you don't need two types of pointers.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array.
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array references to use the existing subscript mechanism, be sufficient?-a Then you don't need two types of pointers.
On 07/09/2026 01:59, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then >> you don't need two types of pointers.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array.
That is up to the compiler, and dependent on types and additional surrounding code. For "typical" usage - "i" and "p" as local variables,
and "i" being either a signed integer type or a size_t (i.e., not a
32-bit unsigned int on a 64-bit machine), and using an optimising
compiler, I would not expect any difference in the code.
But if you can find a more complete example that can be tested on
godbolt, it would be very interesting.
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 01:59, MitchAlsup wrote:
That is up to the compiler, and dependent on types and additional
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address
arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then >>>> you don't need two types of pointers.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array. >>
surrounding code. For "typical" usage - "i" and "p" as local variables,
and "i" being either a signed integer type or a size_t (i.e., not a
32-bit unsigned int on a 64-bit machine), and using an optimising
compiler, I would not expect any difference in the code.
The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
of strength reduction to figure out that only 1 ADD is needed.
But if you can find a more complete example that can be tested on
godbolt, it would be very interesting.
On 07/09/2026 20:25, MitchAlsup wrote:
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 01:59, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address >>>>>> arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array
(so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then >>>> you don't need two types of pointers.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array.
That is up to the compiler, and dependent on types and additional
surrounding code. For "typical" usage - "i" and "p" as local variables, >> and "i" being either a signed integer type or a size_t (i.e., not a
32-bit unsigned int on a 64-bit machine), and using an optimising
compiler, I would not expect any difference in the code.
The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
of strength reduction to figure out that only 1 ADD is needed.
The first one has two adds too - i++, and (p + i) for the array access.
(I'm assuming that there is more going on inside the loop in real code,
such as at least reading or writing from the array. Otherwise the
compiler can see that the whole thing is doing nothing and skip it.)
In both cases, optimisers are likely to generate code approximating :
q = &p[max];
while (p < q) {
p++;
}
(Again, I assume there's a read or write to be included inside the loop.)
Sometimes pointer / array accesses with increment cannot be optimised as well if the index is an unsigned type smaller than size_t, because the compiler can't rule out the possibility of the index wrapping. This is
not an issue with signed integer types, or a big enough unsigned type,
nor is it an issue here when the limits of the index are known. And it
is independent of the ISA.
But if you can find a more complete example that can be tested on
godbolt, it would be very interesting.
Again, if you can give a more complete example, it would be easier to
see what you are getting at here, because I cannot yet see your point.
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 01:59, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:That is up to the compiler, and dependent on types and
On 9/6/2026 9:58 AM, EricP wrote:On modern RISC ISAs:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array
indexes then 99.999% of memory access errors would
disappear.
Retrofitting memory safety onto C is an uphill battle.-a
Address arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards
compatible.
Also agreed.
Have different kinds of pointers: object pointers that
may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular
array (so they can be bounds checked) and do allow pointer
arithmetic.
Would just disallowing arithmetic on pointers, thus forcing
array references to use the existing subscript mechanism, be
sufficient? Then you don't need two types of pointers.
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being
accessed as an array.
additional surrounding code. For "typical" usage - "i" and
"p" as local variables, and "i" being either a signed integer
type or a size_t (i.e., not a 32-bit unsigned int on a 64-bit
machine), and using an optimising compiler, I would not expect
any difference in the code.
The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes
a lot of strength reduction to figure out that only 1 ADD is
needed.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 8/31/2026 9:35 AM, John Dallman wrote:
I've been reading about various early architectures, and realised that
System/360 seems to have been innovative in a way that isn't talked about >>> very much.
The earliest style seems to have been one-word instructions, each
containing an opcode, modifier bits and an address. That was fine while
memories were small, but became increasingly limiting as they grew.
18-bit addressing was common on 36-bit systems. Even the 64-bit Stretch
was limited to 18 address bits.
The 360 introduced word-sized address registers, which meant that address >>> space became a key feature of a computer. It had, theoretically, 32-bit
addresses, although only 24 bits were implemented at first. Address
registers could be loaded, using instructions longer than one word, but
were usually used for base addresses, and not changed very often.
Was this original with the 360, or did someone else invent it first? Were >>> there other styles, now abandoned?
Well, the Univac 1107 (1962) was a 36 bit word oriented system with up
to 64K word addressing. Its 36 bit instructions contained a 16 bit
address field. When the loosely upward compatible 1108 came out in
1964, maximum memory increased to 256K words. Addresses above 64K were
accessed by loading a value into the low order 18 bits of a 36 bit index
register, whose contents were added (by the hardware) to the address in
the instruction.
Several 36-bit machines provided 18-bit displacements with various
kinds of base and index registers.
David Brown <david.brown@hesbynett.no> posted:
Again, if you can give a more complete example, it would be easier to
see what you are getting at here, because I cannot yet see your point.
for( i = 0; i < max; i++ )
a[i] = b[i] + c[i];
versus
for( i = 0; i < max; i++ )
*a++ = *b++ + *c++;
The former has 1 add per iteration, the later has 4.
They added virtual memory similar
to the 67's to S/370 in 1972 and soon made all the major operating
systems use it.
In practice, I never found it to be a problem as
indirection was rare and I never saw more than two levels.
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 20:25, MitchAlsup wrote:
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 01:59, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/6/2026 9:58 AM, EricP wrote:
On 2026-Sep-06 10:56, Thomas Koenig wrote:
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes >>>>>>>>> then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle.-a Address >>>>>>>> arithmetic stands in the way of that.
It might be possible to do but no one is interested.
Agreed.
Making arrays a first class type and distinct from pointers
would be the first step but would not be backwards compatible.
Also agreed.
Have different kinds of pointers: object pointers
that may/may-not be NULL but don't allow pointer arithmetic,
array element pointers that only point inside a particular array >>>>>>> (so they can be bounds checked) and do allow pointer arithmetic.
Would just disallowing arithmetic on pointers, thus forcing array
references to use the existing subscript mechanism, be sufficient? Then >>>>>> you don't need two types of pointers.
On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array.
That is up to the compiler, and dependent on types and additional
surrounding code. For "typical" usage - "i" and "p" as local variables, >>>> and "i" being either a signed integer type or a size_t (i.e., not a
32-bit unsigned int on a 64-bit machine), and using an optimising
compiler, I would not expect any difference in the code.
The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
of strength reduction to figure out that only 1 ADD is needed.
The first one has two adds too - i++, and (p + i) for the array access.
The p+i addition is part of AGEN and thus free for any ISA that has at
least [Rpointer+Rindex] addressing mode (everybody but RISC-V).
(I'm assuming that there is more going on inside the loop in real code,
such as at least reading or writing from the array. Otherwise the
compiler can see that the whole thing is doing nothing and skip it.)
In both cases, optimisers are likely to generate code approximating :
q = &p[max];
while (p < q) {
p++;
}
(Again, I assume there's a read or write to be included inside the loop.)
Sometimes pointer / array accesses with increment cannot be optimised as
well if the index is an unsigned type smaller than size_t, because the
compiler can't rule out the possibility of the index wrapping. This is
not an issue with signed integer types, or a big enough unsigned type,
nor is it an issue here when the limits of the index are known. And it
is independent of the ISA.
But if you can find a more complete example that can be tested on
godbolt, it would be very interesting.
Again, if you can give a more complete example, it would be easier to
see what you are getting at here, because I cannot yet see your point.
for( i = 0; i < max; i++ )
a[i] = b[i] + c[i];
versus
for( i = 0; i < max; i++ )
*a++ = *b++ + *c++;
The former has 1 add per iteration, the later has 4.
John Levine <johnl@taugh.com> writes:
They added virtual memory similar
to the 67's to S/370 in 1972 and soon made all the major operating
systems use it.
Given VM, why was there a need to add virtual memory to the guest OSs?
EricP <ThatWouldBeTelling@thevillage.com> schrieb:
If programmers just used languages that checked array indexes
then 99.999% of memory access errors would disappear.
Retrofitting memory safety onto C is an uphill battle. Address
arithmetic stands in the way of that.
First make array index checks simple and cheap.
This requires checked arithmetic for the index expression calculation,
plus a set of simple compare-and-fault instructions various bounds checks.
You would probably need more than 32 registers for this...
Also, smarten up compilers so they move as much as possible of the
checking outside of loops.
Then see what's left to address.
Use after free will still be a problem, but maybe INVALIDATE
can help there.
ASIDs came along in the late 1990s and early 2000s to better optimize
MMU resource consumption {fewer flushes, support nested paging, ...}.
With their advent, ASIDs were assigned to threads/processes.
Has anyone ever wanted to allow shared ASIDs in such a way that shared
memory or shared files have the ASID associated with the memory/file ?
So that multiple processes accessing the same shared resource co-optim-
ize themselves across cores and caches ??
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
ASIDs came along in the late 1990s and early 2000s to better optimize
MMU resource consumption {fewer flushes, support nested paging, ...}.
With their advent, ASIDs were assigned to threads/processes.
Has anyone ever wanted to allow shared ASIDs in such a way that shared >memory or shared files have the ASID associated with the memory/file ?
So that multiple processes accessing the same shared resource co-optim-
ize themselves across cores and caches ??
Consider that the ASID (and VMID generally speaking) must be present
in every TLB entry, so the size of the ASID (arm supports 8 or 16 bit ASIDS) determines the overall size of the TLB (along with the entry count).
Even with 16-bit ASIDs, large systems with more than 64k processes (not uncommon with 1000-thread processors) will suffer during ASID rollover events.
Associating ASIDs with other resources (such as memory objects) complicates the ASID rollever and assignment functionality without corresponding performance benefits.
Using the G (Global) flag on those architectures which support it often provides similar benefits where applicable.
Additionally, for your suggestion to be viable, every use of the shared object (memory, file pages) must be mapped at the same virtual address;
which is considered a defect (see System V shared libraries for example), even with the larger (48-bit) virtual address space on modern CPUs.
I wonder, did you try to add bound checking to gfortran?
ANd if yes did you measure performance impact?
Given VM, why was there a need to add virtual memory to the guest
OSs?
Using Fortran's array syntax removes the checks completely, so
c = a + b
has no checks left.
It also iserts all checks before the
scalarizer-generated loops.
Unification of a pre-existing free logical variable V with something
(X) is implemented as making V point to X. If X is another logical
variable, and then is unified with something, say Y, a pointer to Y
is stored in the memory location for X. And so on. Accessing V may
need to follow an arbitrarily long chain of pointers until you find
either a free variable, or you find the value that V eventually was
unified with.
Thomas Koenig <tkoenig@netcologne.de> posted:
--------------
Using Fortran's array syntax removes the checks completely, so
c = a + b
has no checks left.
Given that the compiler can see the precise shapes of {a,b,c}
the are no needs for checks. And that is one of the great
benefits of higher-level expressions.
It also iserts all checks before the
scalarizer-generated loops.
When the precise shapes are not (or cannot be) known at compile
time.
On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:That is the way dynamic method invocation works, at least for single-inheritance object-oriented languages:
Unification of a pre-existing free logical variable V with something
(X) is implemented as making V point to X. If X is another logical
variable, and then is unified with something, say Y, a pointer to Y
is stored in the memory location for X. And so on. Accessing V may
need to follow an arbitrarily long chain of pointers until you find
either a free variable, or you find the value that V eventually was
unified with.
There should be a more efficient way: having all bound variables
contain a pointer to shared info about the binding (including
backpointers to all the variables that point here). That should limit
the amount of pointer-chasing necessary.
Lawrence DrCOOliveiro wrote:
On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:
Unification of a pre-existing free logical variable V with
something (X) is implemented as making V point to X. If X is
another logical variable, and then is unified with something, say
Y, a pointer to Y is stored in the memory location for X. And so
on. Accessing V may need to follow an arbitrarily long chain of
pointers until you find either a free variable, or you find the
value that V eventually was unified with.
There should be a more efficient way: having all bound variables
contain a pointer to shared info about the binding (including
backpointers to all the variables that point here). That should
limit the amount of pointer-chasing necessary.
That is the way dynamic method invocation works, at least for single-inheritance object-oriented languages:
On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:
Unification of a pre-existing free logical variable V with something
(X) is implemented as making V point to X. If X is another logical
variable, and then is unified with something, say Y, a pointer to Y
is stored in the memory location for X. And so on. Accessing V may
need to follow an arbitrarily long chain of pointers until you find
either a free variable, or you find the value that V eventually was
unified with.
There should be a more efficient way: having all bound variables
contain a pointer to shared info about the binding (including
backpointers to all the variables that point here). That should limit
the amount of pointer-chasing necessary.
MitchAlsup wrote:
Thomas Koenig <tkoenig@netcologne.de> posted:
--------------
Using Fortran's array syntax removes the checks completely, so
c = a + b
has no checks left.
Given that the compiler can see the precise shapes of {a,b,c}
the are no needs for checks. And that is one of the great
benefits of higher-level expressions.
Absolutely right, this should be a sufficient reason to make all compile-time constant size arrays and 2+ dimensional matrices their own types.
It also iserts all checks before the
scalarizer-generated loops.
When the precise shapes are not (or cannot be) known at compile
time.
As long as the actual size is not miniscule, doing a single set of tests
at startup is so close to free as to not matter.
I.e something like if (a.len() == b.len && b.len == c.len()) could
compile down to
mov rdx,[rbx+8] ;; Dynamic vectors have {ptr, len, size} header
cmp rdx,[rax+8]
jne panic
cmp rdx,[rcx+8]
jne panic
which would be predicted to not panic and therefore run in a cycle or
two, right?
John Levine <johnl@taugh.com> writes:
They added virtual memory similar
to the 67's to S/370 in 1972 and soon made all the major operating
systems use it.
Given VM, why was there a need to add virtual memory to the guest OSs?
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
John Levine <johnl@taugh.com> writes:
They added virtual memory similar
to the 67's to S/370 in 1972 and soon made all the major operating >>systems use it.
Given VM, why was there a need to add virtual memory to the guest OSs?
Orignal proposal for VM 370 included debugging MVS. Once MVS moved
to virtual memory VM 370 needed to emulate this.
People sometimes distinguish paravirtualization, when guest OS is
aware and posibly depends on the host and full virtualization
where guest OS "thinks" that it is running under real hardware.
Once you have paravirtualization, there is nasty question how
guest and host should divide the work and what is the best
interface.
In particular it makes sense to have all virtual
memory management in the host.
But paravitialization needs
cooperation from the guest. So, for general use we have full
vitialization and consequently guest using paging hardware.
David Brown <david.brown@hesbynett.no> posted:[snip]
On 07/09/2026 20:25, MitchAlsup wrote:
David Brown <david.brown@hesbynett.no> posted:
On 07/09/2026 01:59, MitchAlsup wrote:
[snip]On modern RISC ISAs:
for( i = 0; i < max; i++ )
p[i]
is often faster than:
for( i = 0; i < max; i++ )
*p++
Especially when there are more than 1 structure being accessed as an array.
The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
of strength reduction to figure out that only 1 ADD is needed.
The first one has two adds too - i++, and (p + i) for the array access.
The p+i addition is part of AGEN and thus free for any ISA that has at
least [Rpointer+Rindex] addressing mode (everybody but RISC-V).
Paul Clayton <paaronclayton@gmail.com> posted:[snip]
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
Interestingly, Stanford MIPS used the extra bits for an address
space number and had a variable length mask. (-2The size of the
process virtual address space is defined by a bit mask in a
special register.
My 66000 has a 3-bit LVL field in the Root pointer (in all MMU
pointers). LVL in root tells you the VAS size (and PA==VA).
A process==thread can have VAS as small as 23-bits (cat) and
use 1 page as the page table. Saving all those unnecessary
MMU accesses.
LVL in PTPs allows for level skipping (sparse spaces). LVL in
the pointer pointing at a PTE tells you size of the (super)page.
No need for a control register field to specify what the natural
organization of the table structure provides.
A valid design point in 1983, not so much today--unless vastly
extended.
(32-bit PowerPC segments were vaguely similar in
allowing 16 segments with separate virtual address spaces. In
theory, this might support somewhat flexible sharing. HP PA-RISC
provided fewer segments but more flexible/complex use. [I think
encoding the segment bits in the least significant bits would
have been better than using the high bits; it would have made
the change to 64-bit simpler and allowed full-space dynamic
segment addresses at the cost of having to encode segment
numbers in the instruction for byte and half-word aligned
pointers and disallowing alignment traps based on such bits.])
More sins of the past...
Looking forward, loading 128-256 bits per access might be useful
in the not so distant future, if only used for complex types
{real, imag} with 64-bit and 128-bit element sizes.
With x86 one can perform a full-register-size load and access
sub-sections (for some registers). Providing denser register
storage use has advantages when memory is slow, but using
subregisters complicates renaming (and forwarding, which is
kind of renaming).
The understatement of the month award winner.
(Like CMOV, out-of-order execution
complicates use of subregisters.)
For 70 years, a register would contain a single value, then
MMX ruined the game...prior was the model compilers are good
at using.
SIMD is bad for your architecture and for your thinking processes.
The only thing your architecture cannot survive (over time)
is lack of address bits. Do not give them away before your 4th
generation implementations have sold 100M chips.
With ECAM based on PCIe 4.0+, your device address space consumes
up to 40-bits. And then there is the configuration space, and the
interrupt table aperture space, and other system spaces--all vying
for those 47-bits. I suspect 47 will not be sufficient very long.
My 66000's solution is to provide a complete 64-bit VAS that can be translated into a 66-bit universal address space consisting of four
64-bit spaces {DRAM, device, config, ROM}. I don't want any of these
spaces to cause issues while I remain alive. Both cores and devices
use the 66-bit UAS.
Using MSI-X interrupts, allows interrupt tables to perform DPC/softIRQ queueing without normally associated SW overheads. Cores can send cores interrupts using the same mechanisms as devices. Since there is an
unlimited number of interrupt tables (My 66000) and since the are
constructed with a DAG structure, anything that can touch the interrupt aperture can send interrupts to any virtual core. When that virtual
core has control, those interrupts are 'processed'.
antispam@fricas.org (Waldek Hebisch) posted:
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
John Levine <johnl@taugh.com> writes:
They added virtual memory similar
to the 67's to S/370 in 1972 and soon made all the major operating
systems use it.
Given VM, why was there a need to add virtual memory to the guest OSs?
Orignal proposal for VM 370 included debugging MVS. Once MVS moved
to virtual memory VM 370 needed to emulate this.
People sometimes distinguish paravirtualization, when guest OS is
aware and posibly depends on the host and full virtualization
where guest OS "thinks" that it is running under real hardware.
Originally, paravirtualization was to speed up full virtualization.
Once you have paravirtualization, there is nasty question how
guest and host should divide the work and what is the best
interface.
Host OS needs an entry point to Guest OS to ask for resources back
while allowing Guest OS to determine which resources to return.
Guest OS needs an entry point in Host OS to ask for more resources
while allowing Host OS to determine which resources to send.
In particular it makes sense to have all virtual
memory management in the host.
Nested Paging has won. Guest OS virtualizes applications and its own
worker threads, while Host OS virtualized multiple Guest OSs.
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
Some software is highly optimized for a specific implementation
(e.g., cache sizes) such that a microarchitectural change could
require substantial rewriting/retuning. I feel the software
developers should be allowed to make such fragile optimizations
with the understanding that the microarchitecture is not
guaranteed to continue even if the architecture does. (Some
embedded chip vendors do provide product longevity guarantees.)
Originally, paravirtualization was to speed up full virtualization.
On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:
Originally, paravirtualization was to speed up full virtualization.
Still is. ItrCOs common to have drivers for block devices (disks, SSDs),
in particular, in the guest, that are written to work through custom
APIs provided by the host, rather than believe they are talking
directly to actual (virtualized) hardware. You get better performance
that way.
On 9/2/26 9:50 PM, MitchAlsup wrote:
Paul Clayton <paaronclayton@gmail.com> posted:[snip]
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
Some software is highly optimized for a specific implementation
(e.g., cache sizes) such that a microarchitectural change could
require substantial rewriting/retuning. I feel the software
developers should be allowed to make such fragile optimizations
with the understanding that the microarchitecture is not
guaranteed to continue even if the architecture does. (Some
embedded chip vendors do provide product longevity guarantees.)
(This can reduce microarchitectural flexiblity if the software
is considered critical enough and rewriting/retuning is not an
option.)
Software that chooses to use such address bits for tags will
still run on newer designs with full 64-bit virtual addresses
though limited to 56-bit virtual addresses. I think that most
software will never need 64-bit virtual addresses. I also think
that a lot of software would not bother using address tagging.
Address space constrains actual capability of software rather
than "merely" performance, but some software developers (and
some developers of the ARM architecture) have concluded that the
benefit of having tags is worth the cost of a reduced address
space as an option.
I would also note that the canonical address scheme also
introduces a software incompatibility aspect. Software using
a non-canonical address to generate an exception and assuming
the OS will not provide addresses outside the current limit
(as of the 2024 man page I have Linux mmap does not provide a
flag to limit the address range other than MAP_32BIT, so other
than using fixed allocations, an application cannot force the
OS to honor its wishes to limit the virtual address range).
This is not a useful technique given how slow exceptions are,
but I think it is a hole in "compatibility" (and more
significant than space bar heating, https://xkcd.com/1172/ ).
[snip]
Interestingly, Stanford MIPS used the extra bits for an address
space number and had a variable length mask. (-2The size of the
process virtual address space is defined by a bit mask in a
special register.
My 66000 has a 3-bit LVL field in the Root pointer (in all MMU
pointers). LVL in root tells you the VAS size (and PA==VA).
A process==thread can have VAS as small as 23-bits (cat) and
use 1 page as the page table. Saving all those unnecessary
MMU accesses.
LVL in PTPs allows for level skipping (sparse spaces). LVL in
the pointer pointing at a PTE tells you size of the (super)page.
No need for a control register field to specify what the natural organization of the table structure provides.
I do appreciate those features of My 66000 page tables. I do
wonder if using large pages in the page table itself might be
worth supporting (Andy Glew had suggested such).
With page level "merging" (large pages), one could (in theory)
have 5-bit levels that could support multiples of 32 in page
size while not requiring the page table depth to be doubled.
256-KiB pages might be a useful option between 8 MiB and 8 KiB.
Such also supports more flexible tradeoffs of depth versus
internal fragmentation. An application using a large address
space might have little sparsity in the upper address bits and
so benefit from merging top page table levels.
Level merging (really splitting) might also support more
flexible sharing of page tables. With 5-bit basic levels, an
aligned 256-byte section of a 8 KiB page table block could be
shared without sharing the rest of the translations or
permissions. (I have no idea if such would actually be useful
much less worthwhile, but it is possible.)
Level merging seems to be especially useful for nested page
tables in virtualization. The virtual-physical address space
is dense (I think). A virtual machine monitor might give huge
pages to a guest but could also benefit from a flatter upper
portion of the page table.
OS support would be a problem. I do not think any architecture
provides merging of page table levels.
One weird thought I had was to have multiple page table "bases"
loaded for a process. This would be a little like caching
directory entries with prefetching of such on context switches.
The advantage, such as it is, would be in supporting sparse
use with a limited number of roots for faster context switches.
Rather than having to traverse the page table three times to
load three specific node points into the translation cache (and
likely caching other less useful information until it ages out),
the necessary entries would be prefetched and "locked".
Two obvious issues come to mind. First, context switches are
expected to be uncommon, so a modest savings becomes a trivial
savings. Second, this does not scale down or up to different
numbers of nodes.
If the prefetching was optional, then a means would be required
to load missing entries. This seems to require either a full
page table (one node address from which all the other actually
used/valid nodes can be found) rCo which uses extra storage and
look-up layers rCo or something like a hash table (like software
TLBs) that hold the information rCo which also has storage use
overhead (though it might be limited by only storing entries
greater than a default value, e.g., four nodes might be
automatically prefetched and any other nodes would need to
access the hash table) and look-up overhead (number of hash
table probes, at least the parallelism of such allows
tradeoffs of bandwidth versus latency). Since the prefetched
nodes could be broadened to spaces covering multiple previous
nodes, I think hash table misses might be avoidable by
construction at the cost of deeper page tables.
This was just a wild and crazy thought. The unconventionality
alone probably makes such impractical even if it would be
technically possible and perhaps even slightly useful.
[snip]
A valid design point in 1983, not so much today--unless vastly
extended.
I just thought it was a neat little design choice.
[snip My 66000 avoidance of global bit]
How close do you think you are to version 1.0?
You sent me the Principles of Operation from January 2020 and
your posts on comp.arch indicate that a substantial amount has
changed since then. (I think the DOUBLE prefix has been added,
dropped, and reinstated. CARRY was not present in 2020)
I get the impression that a lot of your work recently has been
on system (and implementation) aspects rather than "instruction
set" aspects.
I do not understand how the Virtual Vector Method would be
implemented efficiently. SIMD with its explicit pack and unpack
instructions seems likely to provide similar control. (Most SIMD
designs do not provide special support for short strides or
complete structure unpacking where all the data in a non-strided
stream is used but the data is scattered. GPUs might provide
special support for three and four color (un)packing.) However,
I am not a hardware designer (though I might understand a simple
flowchart ry|).
I do feel that VVM is a nice software interface. It avoids a
lot of SIMD issues and not just instruction diversity explosion.
I feel it does not exploit a few local, non-loop wide execution
cases, does not address blocking, and the loop length limits may
be introduce issues.
Yet I also recognize that being ideal for
all use cases regardless of complexity is both unachievable
(not all tradeoffs are limited to design difficulty or even
implementation area) and impractical (aside from complexity
having a cost, different workloads benefit differently from
"effort" and the value of the benefit is uniform across all
workloads at all scales).
[snip]
(32-bit PowerPC segments were vaguely similar in
allowing 16 segments with separate virtual address spaces. In
theory, this might support somewhat flexible sharing. HP PA-RISC
provided fewer segments but more flexible/complex use. [I think
encoding the segment bits in the least significant bits would
have been better than using the high bits; it would have made
the change to 64-bit simpler and allowed full-space dynamic
segment addresses at the cost of having to encode segment
numbers in the instruction for byte and half-word aligned
pointers and disallowing alignment traps based on such bits.])
More sins of the past...
I am not sure. Segments based on the most significant bits
certainly have issues with respect to increasing the base
address space size, but I am not certain that segments as
address space extensions are necessarily a bad thing. Simple
address spaces are nicer, but the cost of doubling address size
for a few uses that could be handled reasonably with segments
(because of data locality even at that large scale).
Maybe rather than segments (or large flat virtual address
spaces) future huge memory applications might use multiple
address spaces (program/code-connected segmentation) to
support an effectively larger address space.
Given the
desirability of code and data physical locality and the
latency and power issues of a large physical flat memory,
such seems reasonable.
Current warehouse-scale compute
uses program partitioning (microservices) largely to provide
usage scalability and availability, but sharding of databases
does target the memory capacity issue (which is currently a
tighter constraint than virtual address space size).
[snip]
Looking forward, loading 128-256 bits per access might be useful
in the not so distant future, if only used for complex types
{real, imag} with 64-bit and 128-bit element sizes.
Microarchitecturally exploiting spatial locality (SRAM array
width, block size, and page size) to avoid redundant work seems
an obvious method toward power efficiency. Allowing software to
help seems reasonable to *me* (but I am excessively attracted by hardware-software cooperation).
With x86 one can perform a full-register-size load and access
sub-sections (for some registers). Providing denser register
storage use has advantages when memory is slow, but using
subregisters complicates renaming (and forwarding, which is
kind of renaming).
The understatement of the month award winner.
Surely that is a more temporally local understatement.ry| (Maybe
limited to comp.arch?)
A two-size renaming (like IBM zSeries [s/360 descendant]) would
seem to only double the number of registers and the number of
sources. There might even be techniques to simplify the hardware
at the cost of storage utilization and/or extra data movement.
Checkpoint-only values can be more readily copied to a different
namespace (only a single pointer needs to be updated and the
update is not on a critical path). Perhaps an operation might
store two or more values rCo the result to be stored in quick
storage and one or more source operands to be stored in slower
checkpoint storage rCo avoiding an excessive read for the move and
generally keeping the move out of the forwarding network.
I _suspect_ subregisters have potential (certainly for in-order
with "same lane" usage), but developing a reasonable design
would probably require many person-years of research and by
then tradeoffs would have changed and any modest benefit would
likely be of limited application.
(Like CMOV, out-of-order execution
complicates use of subregisters.)
For 70 years, a register would contain a single value, then
MMX ruined the game...prior was the model compilers are good
at using.
x86 had subregisters before MMX. Even some RISCs used register
pairs for double precision floating point, presenting the same renaming/forwarding issues.
[snip]
SIMD is bad for your architecture and for your thinking processes.
I disagree. Vectors are Single Instruction Multiple Data. Fixed
work unit size per instruction does seem architecturally
problematic, but it is not clear that such is much more
problematic for thinking than fixed cache block size.
(Fixed cache block size does affect thinking. In-cache
compression, especially "lossy" compression where part of the
nominal cache block is stored in outer storage, will affect
understanding of capacity. Prefetch/false sharing assumptions
will affect performance.)
[snip]
The only thing your architecture cannot survive (over time)
is lack of address bits. Do not give them away before your 4th
generation implementations have sold 100M chips.
The length of time between address space doubling doubles, and
the delay is longer if cost-per-bit does not halve regularly
and/or capacity demand is worse than linear with price. I think
the speed of solid state storage reduced DRAM demand (cheaper
indirectly addressed storage was good enough for many uses),
lengthening the 64-bit generation. Process scaling issues also
seem to be slowing demand for large memory single systems. The
move toward scale out rather than scale up also reduces the
address space pressure. I also suspect that many of the large
virtual address space use cases are also relatively dense in
mapping, which might buy a few bits of address space compared
with "general" programs.
Address tagging also only applies to software that chooses to
use such. I get the impression that such tag use is rather
niche and even then contained within a subsystem of the program.
Having to rewrite a subsystem may be a huge pain, but
It may be physically possible for a memory technology
breakthrough to make Moore's Law look slow (for a while), but
I would be willing to bet that virtual address space capacity
demand will not accelerate and even that such will continue
to decelerate (though not as much as from the one-time sold
state storage effect or the, hopefully temporary, AI causes
memory price increases).
[snip]
With ECAM based on PCIe 4.0+, your device address space consumes
up to 40-bits. And then there is the configuration space, and the
interrupt table aperture space, and other system spaces--all vying
for those 47-bits. I suspect 47 will not be sufficient very long.
Interesting, but the applications using tagged addresses are not
likely to be interacting so directly with I/O (I think).
My 66000's solution is to provide a complete 64-bit VAS that can be translated into a 66-bit universal address space consisting of four
64-bit spaces {DRAM, device, config, ROM}. I don't want any of these
spaces to cause issues while I remain alive. Both cores and devices
use the 66-bit UAS.
Side note: I have felt some attraction to a virtual Harvard
system, i.e., a separate virtual address space for instructions.
Besides making writing to code more explicit and slightly
increasing the address space (less than a bit as code is rarely
half of the used memory for larger systems), such might present
opportunities to use the instruction pointer to generate data
addresses that are not cluttered with instructions.
I thought of the possibility of using negative offsets from a
masked instruction pointer to provide a cheap pointer usable
with shorter offsets. Separate instruction and data address
spaces would allow
Using MSI-X interrupts, allows interrupt tables to perform DPC/softIRQ queueing without normally associated SW overheads. Cores can send cores interrupts using the same mechanisms as devices. Since there is an unlimited number of interrupt tables (My 66000) and since the are constructed with a DAG structure, anything that can touch the interrupt aperture can send interrupts to any virtual core. When that virtual
core has control, those interrupts are 'processed'.
I like the orientation toward "an agent is an agent". That is
just one more indication of My 66000 being a *design* and not
merely a collection of useful ideas. I am not up to the task of
designing a paper ISA, but I can still appreciate engineering.
On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:
Originally, paravirtualization was to speed up full virtualization.
Still is. ItrCOs common to have drivers for block devices (disks, SSDs),
in particular, in the guest, that are written to work through custom
APIs provided by the host, rather than believe they are talking
directly to actual (virtualized) hardware. You get better performance
that way.
On Wed, 9 Sep 2026 21:45:07 -0000 (UTC), Lawrence DrCOOliveiro wrote:
On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:
Originally, paravirtualization was to speed up full virtualization.
Still is. ItrCOs common to have drivers for block devices (disks,
SSDs), in particular, in the guest, that are written to work
through custom APIs provided by the host, rather than believe they
are talking directly to actual (virtualized) hardware. You get
better performance that way.
Is that prior to PCIe device virtualization or after or from
something else ?
Itanium (not a RISC and probably most appropriate to use past tense
though it may not yet be exclusively a collector's/museum's ISA) ...
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
It's useful to start with an existing experimental project and
look at the current status thereof:
https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
On Wed, 9 Sep 2026 01:37:03 -0000 (UTC), Lawrence DrCOOliveiro wrote:
On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:
Unification of a pre-existing free logical variable V with
something (X) is implemented as making V point to X. If X is
another logical variable, and then is unified with something, say
Y, a pointer to Y is stored in the memory location for X. And so
on. Accessing V may need to follow an arbitrarily long chain of
pointers until you find either a free variable, or you find the
value that V eventually was unified with.
There should be a more efficient way: having all bound variables
contain a pointer to shared info about the binding (including
backpointers to all the variables that point here). That should
limit the amount of pointer-chasing necessary.
That would need back pointers stored everywhere, at least doubling
the necessary memory (plus storing management information, because n
logical variables can point to the same free logical variable), plus
the memory for storing all these changes on the trail stack for
backtracking. And all this memory has to be written.
And in practice the chains of logical variables are usually short,
so you would create all this overhead to address a rare case.
Waldek Hebisch <antispam@fricas.org> schrieb:
I wonder, did you try to add bound checking to gfortran?
It has been there for a long time, and not by me :-)
ANd if yes did you measure performance impact?
I haven't run benchmarks, but analysis can be interesting.
Consider
subroutine foo(a,b,c,n)
integer, intent(in) :: n
real, dimension(n), intent(in) :: a,b
real, dimension(n), intent(out) :: c
integer :: i
do i=1,n
c(i) = a(i) + b(i)
end do
end subroutine foo
Using gfortran, this is translated by the front end into a loop
where every possible condition is checked on every iteration.
Optimization passes convert this into a single check against
INT_MAX, which inhibits vectorization.
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
The length of time between address space doubling doubles, and
the delay is longer if cost-per-bit does not halve regularly
and/or capacity demand is worse than linear with price. I think
the speed of solid state storage reduced DRAM demand (cheaper
indirectly addressed storage was good enough for many uses),
Yes, SSD took some demand out of DRAM size requirements, by
shrinking 10ms rotational latency into 50-|s access delay;
and by raising the data gat from 4GBs into the 30GB/s range.
On the other hand, AI turned around and is consuming all available
DRAM, so its a good thing SSDs are so fast.
On Wed, 09 Sep 2026 23:28:39 GMT, MitchAlsup wrote:
On Wed, 9 Sep 2026 21:45:07 -0000 (UTC), Lawrence DrCOOliveiro wrote:
On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:
Originally, paravirtualization was to speed up full virtualization.
Still is. ItrCOs common to have drivers for block devices (disks,
SSDs), in particular, in the guest, that are written to work
through custom APIs provided by the host, rather than believe they
are talking directly to actual (virtualized) hardware. You get
better performance that way.
Is that prior to PCIe device virtualization or after or from
something else ?
ItrCOs a current thing.
I looked around, and found a good overview here ><https://wiki.xenproject.org/wiki/Understanding_the_Virtualization_Spectrum>:
In article <vdenS.4266$Bn83.1011@fx12.iad>,
Scott Lurndal <slp53@pacbell.net> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
It's useful to start with an existing experimental project and
look at the current status thereof:
https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
CHERI summary: all pointers are now called capabilities and are 129 bits,
so registers which can hold addresses become 129 bits. Capabilities in >memory are also 129 bits, with 128 bits being in "normal" memory, and the >"valid" bit being elsewhere (think of it as hidden in ECC bits). Any
data write to memory or a register clears the "valid" bit, so software cannot >forge a capability. All loads/stores use capabilities as their pointers.
In article <vdenS.4266$Bn83.1011@fx12.iad>,
Scott Lurndal <slp53@pacbell.net> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Looking forward::
a) What kind of distinction should architects make between pointers
and addresses ??
b) are there other ways to make capabilities cheaper without losing
their protection properties ??
As to (a) a pointer could have some bits used to restrict access
rights. So, one could create a read-only pointer and use it only
for reading even when the PTE says it is writeable.
As to (b) a capability might have a 64-bit pointer and a 64-bit
index into a capability table hidden in Guest OS address space.
It's useful to start with an existing experimental project and
look at the current status thereof:
https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
CHERI summary: all pointers are now called capabilities and are 129 bits,
so registers which can hold addresses become 129 bits. Capabilities in memory are also 129 bits, with 128 bits being in "normal" memory, and the "valid" bit being elsewhere (think of it as hidden in ECC bits). Any
data write to memory or a register clears the "valid" bit, so software cannot forge a capability. All loads/stores use capabilities as their pointers.
The capability basically stores a normal pointer in the low 64 bits, and encodes the base address of the capability and the size in the upper 64 bits, with other info. It encodes this by saying larger sized objects have
to have some minimum base and size alignment (so, 16MB+ object must be 4KB aligned, or something like that).
The problem is CHERI is trying to do much more than just provide bounds checking--they also have sealed and unsealed capabilities, and lots
of rules about managing these extra fields. But: they neglect to actually explain what they intend to do with this, so it all made little sense to me.
I THINK they were trying to make capabilities so that the OS could re-use user capabilities (or user code using kernel capabilities) and not be a security hole, but honestly it made my eyes glaze over and I didn't figure
it out.
If you just want to catch bad memory references, then a simplified CHERI capability could work fine. You don't even need the hidden valid bit since forgery is not really an issue for this case.
But CHERI has some holes (just off the top of my head, I'm sure there's more):
- It needs to disallow capabilities in shared memory from creating security
holes. One process mmap()'s some memory, and stores capabilities
pointing to its private memory. Then, another process mmaps the
same shared memory, and now can use those valid capabilities to
access the same addresses in it's address space, which it might
not have access to! This is a tricky problem to solve since you
want shared libraries to work.
- ARM allows a user process to be big-endian. This similarly can break
the capability security model through mmap() and other means.
- There's a whole slew of DMA-related security issues which seem hard to
fully plug. You need to allow paging of capabilities, and this
creates attack surfaces, or just complexity for real use cases
where you don't want to clear the valid bit. All DMA drivers
become part of the attack surface for forging capabilities.
Again, if you don't try to enforce kernel-level trust on these capabilities,
these no longer are issues, and it's why I don't really like CHERI.
A CHERI-lite just providing bounds checking could be useful.
Kent--- Synchronet 3.22a-Linux NewsLink 1.2
kegs@provalid.com (Kent Dickey) posted:
The problem is CHERI is trying to do much more than just provide bounds
checking--they also have sealed and unsealed capabilities, and lots
of rules about managing these extra fields. But: they neglect to actually >> explain what they intend to do with this, so it all made little sense to me.
I tend to say it as "I understand how to make a capability "pointer",
what I don't understand is how to make it C3-secure."
I THINK they were trying to make capabilities so that the OS could re-use
user capabilities (or user code using kernel capabilities) and not be a
security hole, but honestly it made my eyes glaze over and I didn't figure >> it out.
Whereas; most of use simply want solid base-bounds checking. S O L I D
If you just want to catch bad memory references, then a simplified CHERI
capability could work fine. You don't even need the hidden valid bit since >> forgery is not really an issue for this case.
Do you think that with a 63-bit VAS, one could put each root-capability
into its own upper layer paging structure ?? WHere derived-capabilities >simply point inside that root-capability ??
But CHERI has some holes (just off the top of my head, I'm sure there's more):
- It needs to disallow capabilities in shared memory from creating security >> holes. One process mmap()'s some memory, and stores capabilities
pointing to its private memory. Then, another process mmaps the
same shared memory, and now can use those valid capabilities to
access the same addresses in it's address space, which it might
not have access to! This is a tricky problem to solve since you
want shared libraries to work.
Sort-of defeaters the whole purpose, does it not ??
- ARM allows a user process to be big-endian. This similarly can break
the capability security model through mmap() and other means.
- There's a whole slew of DMA-related security issues which seem hard to
fully plug. You need to allow paging of capabilities, and this
creates attack surfaces, or just complexity for real use cases
where you don't want to clear the valid bit. All DMA drivers
become part of the attack surface for forging capabilities.
This simply makes PCIe devices impossible as each "address plus size"
has to be a capability.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_.
2^64 is a, pardon my french,
shitload of virtual memory.
Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required.
Again, one must read the CHERI documentation before commenting on
such things.
A CHERI-lite just providing bounds checking could be useful.
On Wed, 09 Sep 2026 08:43:46 GMT, Anton Ertl wrote:
On Wed, 9 Sep 2026 01:37:03 -0000 (UTC), Lawrence DrCOOliveiro wrote:
On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:
Unification of a pre-existing free logical variable V with
something (X) is implemented as making V point to X. If X is
another logical variable, and then is unified with something, say
Y, a pointer to Y is stored in the memory location for X. And so
on. Accessing V may need to follow an arbitrarily long chain of
pointers until you find either a free variable, or you find the
value that V eventually was unified with.
There should be a more efficient way: having all bound variables
contain a pointer to shared info about the binding (including
backpointers to all the variables that point here). That should
limit the amount of pointer-chasing necessary.
That would need back pointers stored everywhere, at least doubling
the necessary memory (plus storing management information, because n
logical variables can point to the same free logical variable), plus
the memory for storing all these changes on the trail stack for
backtracking. And all this memory has to be written.
Only on actual unification, not on simple lookup.
And in practice the chains of logical variables are usually short,
so you would create all this overhead to address a rare case.
How short is short, though? The question is, what is the average
length of pointer chains being traversed in each scheme.
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
kegs@provalid.com (Kent Dickey) writes:
A CHERI-lite just providing bounds checking could be useful.
I doubt it. The problems I see here (as well as in CHERI) is that we
have nested data structures (not in every programming language, but
certainly in some relevant ones), such as
struct foo {
char x[3];
struct bar {
char u[3]; // making a 2-byte hole in the struct element
int v[3];
} y[4];
long z[5];
} a[3];
Sometimes you want to check that you do not exceed the bounds of a,
sometimes that you do not exceed the bounds of z, sometimes that you
do not exceed the bounds of u. Sometimes you want to treat all of a
as one thing, sometimes, only some part, sometimes your software does
things beyond a single nested struct, such as garbage collection.
I guess there are ways to do these things in CHERI, but I very much
doubt that the only cost CHERI has for them is the doubled memory consumption. And does it buy anything that Rust does not buy us?
Given that the computing world seems to have decided that it does not
even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will
be a widespread success.
- anton--- Synchronet 3.22a-Linux NewsLink 1.2
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
kegs@provalid.com (Kent Dickey) writes:
A CHERI-lite just providing bounds checking could be useful.
I doubt it. The problems I see here (as well as in CHERI) is that we
have nested data structures (not in every programming language, but
certainly in some relevant ones), such as
struct foo {
char x[3];
struct bar {
char u[3]; // making a 2-byte hole in the struct element
int v[3];
} y[4];
long z[5];
} a[3];
Sometimes you want to check that you do not exceed the bounds of a,
sometimes that you do not exceed the bounds of z, sometimes that you
do not exceed the bounds of u. Sometimes you want to treat all of a
as one thing, sometimes, only some part, sometimes your software does
things beyond a single nested struct, such as garbage collection.
And sometimes you even want to prevent accessing data that is not
"in" the struct due to internal alignments--such as the bytes between
u[2] and v[0].
Thomas Koenig <tkoenig@netcologne.de> writes:
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
If this Anton is supposed to be me, I don't think I ever made such a
claim.
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
kegs@provalid.com (Kent Dickey) writes:
A CHERI-lite just providing bounds checking could be useful.
I doubt it. The problems I see here (as well as in CHERI) is that we
have nested data structures (not in every programming language, but
certainly in some relevant ones), such as
struct foo {
char x[3];
struct bar {
char u[3]; // making a 2-byte hole in the struct element
int v[3];
} y[4];
long z[5];
} a[3];
Sometimes you want to check that you do not exceed the bounds of a,
sometimes that you do not exceed the bounds of z, sometimes that you
do not exceed the bounds of u. Sometimes you want to treat all of a
as one thing, sometimes, only some part, sometimes your software does
things beyond a single nested struct, such as garbage collection.
And sometimes you even want to prevent accessing data that is not
"in" the struct due to internal alignments--such as the bytes between
u[2] and v[0].
That would be difficult, since one needs to be able to copy
the structure, which necessarily will need to access those bytes.
Thomas Koenig <tkoenig@netcologne.de> writes:
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
If this Anton is supposed to be me, I don't think I ever made such a
claim.
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
- anton
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
kegs@provalid.com (Kent Dickey) posted:
The problem is CHERI is trying to do much more than just provide bounds
checking--they also have sealed and unsealed capabilities, and lots
of rules about managing these extra fields. But: they neglect to actually >>> explain what they intend to do with this, so it all made little sense to me.
I tend to say it as "I understand how to make a capability "pointer",
what I don't understand is how to make it C3-secure."
The out-of-band 'tag' bit which marks a valid capability makes it
secure (likely at B level or better in orange book terminology).
I THINK they were trying to make capabilities so that the OS could re-use >>> user capabilities (or user code using kernel capabilities) and not be a
security hole, but honestly it made my eyes glaze over and I didn't figure >>> it out.
Whereas; most of use simply want solid base-bounds checking. S O L I D
Which, when it applies to programming environments that include dynamically loaded code (which is pretty much all of them in these times). The capability,
of course, provides additional valuable capabilities over simple bounds checking,
such as the ability to mark a datum as read-only (without marking the entire page containing the datum).
If you just want to catch bad memory references, then a simplified CHERI >>> capability could work fine. You don't even need the hidden valid bit since >>> forgery is not really an issue for this case.
Do you think that with a 63-bit VAS, one could put each root-capability
into its own upper layer paging structure ?? WHere derived-capabilities
simply point inside that root-capability ??
But CHERI has some holes (just off the top of my head, I'm sure there's more):
- It needs to disallow capabilities in shared memory from creating security >>> holes. One process mmap()'s some memory, and stores capabilities
pointing to its private memory. Then, another process mmaps the
same shared memory, and now can use those valid capabilities to
access the same addresses in it's address space, which it might
not have access to! This is a tricky problem to solve since you
want shared libraries to work.
Sort-of defeaters the whole purpose, does it not ??
If that had been an accurate summarization of how CHERI interacts
with mmap (and shared libraries in general), perhaps.
However, if you think that the CHERI team hasn't given significant
though to the issues of inter-thread and inter-process memory sharing;
you might want to re-read the CHERI documentation.
- ARM allows a user process to be big-endian. This similarly can break
the capability security model through mmap() and other means.
I don't see how. A capability is opaque to software and completely independent of the endianness of the current thread. It must be 8-byte aligned (which reduces the number of tag bits to 1/8th).
- There's a whole slew of DMA-related security issues which seem hard to >>> fully plug. You need to allow paging of capabilities, and this
creates attack surfaces, or just complexity for real use cases
where you don't want to clear the valid bit. All DMA drivers
become part of the attack surface for forging capabilities.
This simply makes PCIe devices impossible as each "address plus size"
has to be a capability.
Again, one must read the CHERI documentation before commenting on
such things.
"CHERI (Capability Hardware-Enhanced RISC Instructions) interacts
with Direct Memory Access (DMA) by treating peripheral devices as
capability-unaware entities and enforcing hardware tag-clearing
mechanisms at the memory/interconnect level to protect capability
integrity."
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Given that the computing world seems to have decided that it does not
even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will
be a widespread success.
My 66000 architecture has means maintain robustness in the face of
Spectr|- attack vectors. Stashing microarchitectural state in the
miss buffers until the causing instruction retires--and only then
updating microarchitectural state.
Don't see why CHERI could not do similarly.
Anton Ertl wrote:
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
SOME compilers generate better code than SOME programmers, SOME of the time.
- fixed that for you.
I guess there are ways to do these things in CHERI, but I very much
doubt that the only cost CHERI has for them is the doubled memory consumption. And does it buy anything that Rust does not buy us?
On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
"compilers generate better code than MOST programmers."
The qualification is important.
There is no doubt that a good
programmer can beat the compiler,
but in my experience [HRT systems]
it requires effort that profitably might be spent on something else.
Keep in mind that most programmers are only average and the skill
level of the average programmer now is only slightly above "script
kiddie". Most have no formal CS or CSE schooling, don't know how to
evaluate algorithms, and largely are incapable of writing for
themselves library functions that they routinely use.
Witness the proliferation of languages offering "managed environments" >offering such niceties as automatic storage management, automatic lock >handling (serialized object access), "comprehensions", etc., and large >standard libraries
without which the average programmer largely
would be incapable of producing a working program.
In general the compiler can be aware of more surrounding context and
can generate better initial code.
Even good programmers can have poor intuition about what needs high >optimization.
Thomas Koenig <tkoenig@netcologne.de> writes:
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
If this Anton is supposed to be me, I don't think I ever made such a
claim.
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
Thomas Koenig <tkoenig@netcologne.de> writes:
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
If this Anton is supposed to be me, I don't think I ever made such a
claim.
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
George Neuner <gneuner2@comcast.net> writes:
On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
"compilers generate better code than MOST programmers."
The qualification is important.
The statements I have read did not make such a qualification,
certainly not in capital letters.
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
Thomas Koenig <tkoenig@netcologne.de> writes:
(I think Anton likes to claim that some people think that
compilers are perfect. Not sure who he means, it's certainly
not me :-)
If this Anton is supposed to be me, I don't think I ever made such a
claim.
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
Your statement is ambiguous in several ways.
I assume you mean
"programmers using assembly"
Do you claim that people claim
a) Compilers generate better code than programmers all the time
b) Compilers generate better code than programmers most of the time[...]
b) is true. In the vast majority of cases, it is vastly uneconomical
to revert to assembly programming.
/ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Given that the computing world seems to have decided that it does not
even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will
be a widespread success.
My 66000 architecture has means maintain robustness in the face of >Spectr|a-- attack vectors. Stashing microarchitectural state in the
miss buffers until the causing instruction retires--and only then
updating microarchitectural state.
This is a good approach to implement invisible speculation as far as
cache side channel is concerned. Behnia et al. [behnia+21] describe a
side channel that uses resource contention from speculative
instructions to affect the timing of committing instructions, i.e., it
uses the scheduler (reservation station) behaviour as a side channel.
Behnia et al. also describe how to close this side channel. There may
be other side channels through microarchitectural resources, and they
need to be closed, too, but the cache side channel probably is the
widest and therefore most relevant one by far.
@InProceedings{ behnia+21,
author = {Mohammad Behnia and Prateek Sahu and Riccardo Paccagnella
and Jiyong Yu and Zirui Neil Zhao and Xiang Zou and Thomas
Unterluggauer and Josep Torrellas and Carlos Rozas and Adam
Morrison and Frank Mckeen and Fangfei Liu and Ron Gabor and
Christo- pher W. Fletcher and Abhishek Basak and Alaa
Alameldeen},
title = {Speculative Interference Attacks: Breaking Invisible
Speculation Schemes},
booktitle = {Architectural Support for Programming Languages and
Operating Systems (ASPLOS |o-C-O21)},
year = {2021},
pages = {1046--1060},
url = {https://dl.acm.org/doi/10.1145/3445814.3446708},
optannote = {}
}
Spectre and friends are microarchitectural side channels, you can
implement microarchitectures for any architecture that are vulnerable
to Spectre, and microarchitectures that are not vulnerable. Therefore
your mention of the My 66000 architecture makes no sense.
Don't see why CHERI could not do similarly.
Certainly one can implement a core with CHERI that is not vulnerable
to speculative side channels (either by not implementing speculation
(slow), or by implementing invisible speculation), and if the
financers and the prospective customers are serious about security,
they will insist on that for eventual products.
But my point is that in the mainstream, not even invisible speculation
with its moderate cost in hardware and performance is implemented, so
I don't expect that a high-cost solution like CHERI becomes
mainstream, in particular given that it provides nothing that cannot
be provided more cheaply through programming languages and compilers.
Ok, you might say, what about legacy code in unsafe languages such as
C? Rewriting all of that in Rust (or using a solution like Ivy/Deputy <http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>, which
would be cheaper, but somehow did not catch on) would also be very
costly.
But I expect that much of this code does not work on CHERI in--- Synchronet 3.22a-Linux NewsLink 1.2
the secure configuration; so to run it on CHERI, you would have to
change it anyway, so you could just as well Deputize it, which is
cheaper in hardware and in execution time. Or rewrite it in Rust, if
you prefer that.
- anton
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high >optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
- anton--- Synchronet 3.22a-Linux NewsLink 1.2
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
/ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Given that the computing world seems to have decided that it does not
even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will
be a widespread success.
My 66000 architecture has means maintain robustness in the face of
Spectr|a-- attack vectors. Stashing microarchitectural state in the
miss buffers until the causing instruction retires--and only then
updating microarchitectural state.
This is a good approach to implement invisible speculation as far as
cache side channel is concerned. Behnia et al. [behnia+21] describe a
side channel that uses resource contention from speculative
instructions to affect the timing of committing instructions, i.e., it
uses the scheduler (reservation station) behaviour as a side channel.
My 66000 does not have architectural constraints on the timing of >instructions--each implementation gets that set of choices. And
thanks for the reference.
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:int do_big_calculation() {
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high
optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
When some subroutines in a module might want heavy optimization, while
most of the rest of the subroutines in that module do not, is an indi-
cation that the module is not well organized.
/ERROR "unexpected byte sequence starting at index 665: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
/ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Given that the computing world seems to have decided that it does not >> >> even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will >> >> be a widespread success.
My 66000 architecture has means maintain robustness in the face of
Spectr|a-a|e-- attack vectors. Stashing microarchitectural state in the >> >miss buffers until the causing instruction retires--and only then
updating microarchitectural state.
This is a good approach to implement invisible speculation as far as
cache side channel is concerned. Behnia et al. [behnia+21] describe a
side channel that uses resource contention from speculative
instructions to affect the timing of committing instructions, i.e., it
uses the scheduler (reservation station) behaviour as a side channel.
My 66000 does not have architectural constraints on the timing of >instructions--each implementation gets that set of choices. And
thanks for the reference.
Are there any instructions where the timing will vary based on the
data being operated on?
On 11/09/2026 19:54, MitchAlsup wrote:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high
optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
When some subroutines in a module might want heavy optimization, whileint do_big_calculation() {
most of the rest of the subroutines in that module do not, is an indi- cation that the module is not well organized.
struct Data data;
initialise_data(&data);
for (int i = 0; i < 1'000'000; i++) {
calculate_round(i, &data);
}
int result = finalise(&data);
return result;
}
Are you suggesting that it would be poor organisation to have "initialise_data", "calculate_round" and "finalise" in the same module?
Or are you suggesting that "initialise_data" should be optimised with
the same priority on speed as "calculate_round" ?
scott@slp53.sl.home (Scott Lurndal) posted:
/ERROR "unexpected byte sequence starting at index 665: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
/ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Given that the computing world seems to have decided that it does not >> >> >> even want to pay the very moderate cost of invisible speculation to
protect against Spectre and friends, I very much doubt that CHERI will >> >> >> be a widespread success.
My 66000 architecture has means maintain robustness in the face of
Spectr|a-a|e-- attack vectors. Stashing microarchitectural state in the >> >> >miss buffers until the causing instruction retires--and only then
updating microarchitectural state.
This is a good approach to implement invisible speculation as far as
cache side channel is concerned. Behnia et al. [behnia+21] describe a
side channel that uses resource contention from speculative
instructions to affect the timing of committing instructions, i.e., it
uses the scheduler (reservation station) behaviour as a side channel.
My 66000 does not have architectural constraints on the timing of
instructions--each implementation gets that set of choices. And
thanks for the reference.
Are there any instructions where the timing will vary based on the
data being operated on?
Things like::
TAN, ATAN, ASIN, ACOS where 1/2 the paths do not need a reciprocal and 1/2 do.
One could imagine MUL, DIV, FMUL, and FDIV having early out FUs. MUL and FMUL >are a lot less likely to have early outs than DIV and FDIV.
But I don't think constant time evaluations need those instructions.
On Thu, 10 Sep 2026 05:41:30 -0000 (UTC), Lawrence DrCOOliveiro wrote:
Only on actual unification, not on simple lookup.
Every "simple lookup" is a unification in Prolog.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
And sometimes you even want to prevent accessing data that is notThat would be difficult, since one needs to be able to copy
"in" the struct due to internal alignments--such as the bytes between
u[2] and v[0].
the structure, which necessarily will need to access those bytes.
George Neuner <gneuner2@comcast.net> writes:
Witness the proliferation of languages offering "managed environments" >>offering such niceties as automatic storage management, automatic lock >>handling (serialized object access), "comprehensions", etc., and large >>standard libraries
I don't see any problem with that. These features help to implement >functionality in less programming time (and with less maintenance
time) than without using these features, and that's true for
programmers at any competence level.
without which the average programmer largely
would be incapable of producing a working program.
That's pure elitism.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
In practice, I never found it to be a problem as
indirection was rare and I never saw more than two levels.
Unification of a pre-existing free logical variable V with something
(X) is implemented as making V point to X. If X is another logical
variable, and then is unified with something, say Y, a pointer to Y is
stored in the memory location for X. And so on. Accessing V may need
to follow an arbitrarily long chain of pointers until you find either
a free variable, or you find the value that V eventually was unified
with.
Prolog was implemented in Edinburgh on the DEC-10 (DEC-10 Prolog). I
guess that they used the indirection feature for that: If a variable
points to some other variable or value, its indirection bit is set, if
it is free, it is not.
On modern machines, i.e., without indirection bit, a free variable
points to itself, a variable that is bound to another variable points
to that other variable. While the type tag indicates a variable, one
follows the pointers, until a pointer to itself is found.
On Thu, 10 Sep 2026 17:16:37 GMT, Anton Ertl wrote:
On Thu, 10 Sep 2026 05:41:30 -0000 (UTC), Lawrence DrCOOliveiro wrote:
Only on actual unification, not on simple lookup.
Every "simple lookup" is a unification in Prolog.
At some point, you have to hit actual values.
Once a variable has a
value, you do comparison of values.
On 9/8/2026 1:34 AM, Anton Ertl wrote:...
Prolog was implemented in Edinburgh on the DEC-10 (DEC-10 Prolog). I
guess that they used the indirection feature for that: If a variable
points to some other variable or value, its indirection bit is set, if
it is free, it is not.
I certainly believe you, but I don't know of a Prolog implementation on
the 1100 series, and if there were to be, based on what you say, it
couldn't use the hardware indirection feature.
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
George Neuner <gneuner2@comcast.net> writes:
Witness the proliferation of languages offering "managed environments" >>>offering such niceties as automatic storage management, automatic lock >>>handling (serialized object access), "comprehensions", etc., and large >>>standard libraries
I don't see any problem with that. These features help to implement >>functionality in less programming time (and with less maintenance
time) than without using these features, and that's true for
programmers at any competence level.
without which the average programmer largely
would be incapable of producing a working program.
That's pure elitism.
Really? That's not my conclusion ... it was the result found by a
number of university studies and developer surveys.
Most studies involving GC have shown that programmers working on short >timelines are less likely to produce a correct [or sometimes even just >complete] program using manual memory management vs using GC.
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
George Neuner <gneuner2@comcast.net> writes:
Witness the proliferation of languages offering "managed environments" >>>offering such niceties as automatic storage management, automatic lock >>>handling (serialized object access), "comprehensions", etc., and large >>>standard libraries
I don't see any problem with that. These features help to implement >>functionality in less programming time (and with less maintenance
time) than without using these features, and that's true for
programmers at any competence level.
without which the average programmer largely
would be incapable of producing a working program.
That's pure elitism.
Really? That's not my conclusion ... it was the result found by a
number of university studies and developer surveys.
Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.
David Brown <david.brown@hesbynett.no> posted:
On 11/09/2026 19:54, MitchAlsup wrote:
int do_big_calculation() {
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high
optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
When some subroutines in a module might want heavy optimization, while
most of the rest of the subroutines in that module do not, is an indi-
cation that the module is not well organized.
struct Data data;
initialise_data(&data);
for (int i = 0; i < 1'000'000; i++) {
calculate_round(i, &data);
}
int result = finalise(&data);
return result;
}
Are you suggesting that it would be poor organisation to have
"initialise_data", "calculate_round" and "finalise" in the same module?
Or are you suggesting that "initialise_data" should be optimised with
the same priority on speed as "calculate_round" ?
What I am suggesting is that it might not be appropriate to use -O3
on all 3. do_big_calculation should get -O3 while initialize_data
and finalize might be just as well served by -O2 or even -OS.
Ok, you might say, what about legacy code in unsafe languages such
as C? Rewriting all of that in Rust (or using a solution like
Ivy/Deputy
<http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>,
which would be cheaper, but somehow did not catch on)
But I expect that much of this code does not work on CHERI in
the secure configuration; so to run it on CHERI, you would have to
change it anyway, so you could just as well Deputize it, which is
cheaper in hardware and in execution time. Or rewrite it in Rust,
if you prefer that.
David Brown <david.brown@hesbynett.no> posted:
On 11/09/2026 19:54, MitchAlsup wrote:
int do_big_calculation() {
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high
optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
When some subroutines in a module might want heavy optimization, while
most of the rest of the subroutines in that module do not, is an indi-
cation that the module is not well organized.
struct Data data;
initialise_data(&data);
for (int i = 0; i < 1'000'000; i++) {
calculate_round(i, &data);
}
int result = finalise(&data);
return result;
}
Are you suggesting that it would be poor organisation to have
"initialise_data", "calculate_round" and "finalise" in the same module?
Or are you suggesting that "initialise_data" should be optimised with
the same priority on speed as "calculate_round" ?
What I am suggesting is that it might not be appropriate to use -O3
on all 3. do_big_calculation should get -O3 while initialize_data
and finalize might be just as well served by -O2 or even -OS.
MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:
David Brown <david.brown@hesbynett.no> posted:
On 11/09/2026 19:54, MitchAlsup wrote:
int do_big_calculation() {
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
George Neuner <gneuner2@comcast.net> writes:------------
Even good programmers can have poor intuition about what needs high >>>>>> optimization.
True, and compilers are not any better, even with profile feedback
(rarely used). Compilers address this by trying to optimize
everything, but they do a mediocre and unreliable job on that.
When some subroutines in a module might want heavy optimization, while >>>> most of the rest of the subroutines in that module do not, is an indi- >>>> cation that the module is not well organized.
struct Data data;
initialise_data(&data);
for (int i = 0; i < 1'000'000; i++) {
calculate_round(i, &data);
}
int result = finalise(&data);
return result;
}
Are you suggesting that it would be poor organisation to have
"initialise_data", "calculate_round" and "finalise" in the same module?
Or are you suggesting that "initialise_data" should be optimised with
the same priority on speed as "calculate_round" ?
What I am suggesting is that it might not be appropriate to use -O3
on all 3. do_big_calculation should get -O3 while initialize_data
and finalize might be just as well served by -O2 or even -OS.
I don't think so. Higher optimization tries more things like
function specialization made possible by constant propagation,
inlining and similar. If calulate_round() considers things as
variable that initialize_data() has as constants, or if finalize()
does not use some things in &data which are nonetheless calculated,
then the win can be quite substantial. (One reason why trying out
simple benchmarks on modern compiler is like nailing a pudding to
the wall).
If you really want efficient code, the best way is probably to
put them all into a single translation unit, or use LTO.
In article <2026Sep11.080019@mips.complang.tuwien.ac.at>, >anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
Ok, you might say, what about legacy code in unsafe languages such
as C? Rewriting all of that in Rust (or using a solution like
Ivy/Deputy
<http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>,
which would be cheaper, but somehow did not catch on)
That link seems to have rotted.
When some subroutines in a module might want heavy optimization, while
most of the rest of the subroutines in that module do not, is an indi-
cation that the module is not well organized.
George Neuner <gneuner2@comcast.net> writes:
On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
"compilers generate better code than MOST programmers."
The qualification is important.
The statements I have read did not make such a qualification,
certainly not in capital letters.
There is no doubt that a good
programmer can beat the compiler,
And especially this statement is usually not made. On the contrary,
the perpetrator of such statements seem convinced of compiler
supremacy.
I would not, however, be willing to assert that one could easily
find human programmers who could beat an optimizing compiler...
which had the Itanium as its target.
Based on what was written here about the similar feature on the
DEC-10, interrupts would mean that a program with a too-long chain
would not make any progress, even if it would not be killed. The
Univac 1100 erroring out would have been preferable in this case.
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
I would not, however, be willing to assert that one could easily find
human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
but given the improvements in optimizing compilers, and the advances
in modern computer architectures, I would still be hesitant to be
categorical in asserting the human can always beat the optimizing
compiler.
Particularly given recent advances in AI.
quadibloc@invalid.com (John Savard) writes:
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
I would not, however, be willing to assert that one could easily find
human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
On the contrary: The IA-64 architects sold their architecture to Intel
and HP management with hand-written examples that made good use of the >hardware resources, and the promise that they would write compilers
that could produce just as good code across the board. Those
compilers failed to appear, and that's a big part of why IA-64 failed.
So compilers are obviously not as good as humans on IA-64 code.
On Sat, 12 Sep 2026 18:14:46 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
quadibloc@invalid.com (John Savard) writes:
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
I would not, however, be willing to assert that one could easily find >>>human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
On the contrary: The IA-64 architects sold their architecture to Intel
and HP management with hand-written examples that made good use of the >>hardware resources, and the promise that they would write compilers
that could produce just as good code across the board. Those
compilers failed to appear, and that's a big part of why IA-64 failed.
So compilers are obviously not as good as humans on IA-64 code.
While this suggests that I was wrong about this, I must admit, I don't
think that it proves the case. The compilers could have failed to
appear for any number of reasons.
And later on, Open64 appeared.
However, concerning my statement above, I think the IA-64 architects
used examples for which the architecture (and its in-order
implementation) was particularly well-suited (software-pipelinable
inner loops), and that compilers eventually usually worked ok for such >examples, too. It's just that these compilers don't work so well on >general-purpose code.
They also underestimated how fast out-of-order implementations would advance >which handle data dependent access patterns just fine.
kegs@provalid.com (Kent Dickey) writes:
A CHERI-lite just providing bounds checking could be useful.
I doubt it. The problems I see here (as well as in CHERI) is that we
have nested data structures (not in every programming language, but
certainly in some relevant ones), such as
struct foo {
char x[3];
struct bar {
char u[5];
int v[3];
} y[4];
long z[5];
} a[3];
Sometimes you want to check that you do not exceed the bounds of a,
sometimes that you do not exceed the bounds of z, sometimes that you
do not exceed the bounds of u. Sometimes you want to treat all of a
as one thing, sometimes, only some part, sometimes your software does
things beyond a single nested struct, such as garbage collection.
I guess there are ways to do these things in CHERI, but I very much
doubt that the only cost CHERI has for them is the doubled memory >consumption. And does it buy anything that Rust does not buy us?
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
However, concerning my statement above, I think the IA-64 architects
used examples for which the architecture (and its in-order
implementation) was particularly well-suited (software-pipelinable
inner loops), and that compilers eventually usually worked ok for such >examples, too. It's just that these compilers don't work so well on >general-purpose code.
That's one of the problems that Multiflow also had, works great if you
can predict the access patterns, a lot less great if the access patterns
are data dependent.
They also underestimated how fast out-of-order implementations would advance which handle data dependent access patterns just fine.
John Levine <johnl@taugh.com> writes:
They also underestimated how fast out-of-order implementations would advance >which handle data dependent access patterns just fine.
OoO has a number of benefits over the EPIC (explicitly parallel
instruction computing) approach of IA-64. I have the impression that
this is not well known even among high-profile computer architects.
In particular, in Hennessy and Patterson's computer architecture book,
they 1) explain the advantage of OoO only with supporting more
in-flight cache misses; and 2) only give a very superficial
description of OoO microarchitectures. My impression is that they
lost interest in processor cores after the RISC revolution, and now
prefer to write more about multiprocessor interconnects and such.
As for the advance of OoO, hardware branch prediction (essential for general-purpose code) had advanced beyond the capabilities of compiler
branch prediction (which compiler-based speculation supported by EPIC
was planned to rely on) in 1991 or so, so EPIC was at a disadvantage
there. Intel also knew that they designed the Pentium 4 (released
2000) with a 128-entry reorder buffer. And looking at SPEC CINT 2000, despite their huge caches, IA-64 implementations never outdid
contemporaneous IA-32 or AMD64 implementations:
System Cint CFP
res base res base CPU Tested Published
Intel D850GB 656 640 714 704 Pentium 4 2000MHz Aug-2001 Sep-2001
hp rx4610 --- 379 715 715 Intel Itanium 800MHz Aug-2001 Sep-2001
Dell Precision WS 340 922 893 901 878 Pentium 4 2533MHz May-2002 Jun-2002
hp workstation zx6000 --- 807 1356 1356 Itanium 2 1000Mhz Jul-2002 Jul-2002
IA-64 implementations were good at CFP, at least initially, though.
Intel should have been aware of the Pentium 4's capabilities early on.--- Synchronet 3.22a-Linux NewsLink 1.2
- anton
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
there's more):
kegs@provalid.com (Kent Dickey) posted:
But CHERI has some holes (just off the top of my head, I'm sure
- It needs to disallow capabilities in shared memory from creating security >>> holes. One process mmap()'s some memory, and stores capabilities
pointing to its private memory. Then, another process mmaps the
same shared memory, and now can use those valid capabilities to
access the same addresses in it's address space, which it might
not have access to! This is a tricky problem to solve since you
want shared libraries to work.
Sort-of defeaters the whole purpose, does it not ??
If that had been an accurate summarization of how CHERI interacts
with mmap (and shared libraries in general), perhaps.
However, if you think that the CHERI team hasn't given significant
though to the issues of inter-thread and inter-process memory sharing;
you might want to re-read the CHERI documentation.
- ARM allows a user process to be big-endian. This similarly can break
the capability security model through mmap() and other means.
I don't see how. A capability is opaque to software and completely >independent of the endianness of the current thread. It must be 8-byte >aligned (which reduces the number of tag bits to 1/8th).
- There's a whole slew of DMA-related security issues which seem hard to >>> fully plug. You need to allow paging of capabilities, and this
creates attack surfaces, or just complexity for real use cases
where you don't want to clear the valid bit. All DMA drivers
become part of the attack surface for forging capabilities.
This simply makes PCIe devices impossible as each "address plus size"
has to be a capability.
Again, one must read the CHERI documentation before commenting on
such things.
"CHERI (Capability Hardware-Enhanced RISC Instructions) interacts
with Direct Memory Access (DMA) by treating peripheral devices as
capability-unaware entities and enforcing hardware tag-clearing
mechanisms at the memory/interconnect level to protect capability
integrity."
This is off the top of my head, so it will be "wrong", but you'll get the idea.
If you want to evaulate:
out = a[i].y[j].v[k];
You do (each line is one instruction, everything is registers):--------
Then the evaluation is:
tmp = sizeof(a[0]);
off = i*tmp + offset(a[0].y[0]);
a_i.ptr = a + off, a_i.bound=sizeof(a[0].y);
tmp = sizeof(a[0].y[0]);--------
off = j*tmp + offset(y[0].v[0]);
y_j.ptr = a_i + off, y_j.bound = Sizeof( a[0].y[0].v );
out = LOAD(y_j + k_tmp);
Kent--- Synchronet 3.22a-Linux NewsLink 1.2
In article <eAAoS.242$3vg8.225@fx37.iad>,--------------------
Scott Lurndal <slp53@pacbell.net> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
CHERI creates new system requirements. You could have DMA maintain capabilities, and that has some nice properties, but then means you have
to trust your DMA device.
The CHERI documentation says whether DMA
should be able to read/write capabilities is an open question, but that
the current default is that DMA writes clear capabilities.
So what's the hole with that? To support paging, it means you must have
a mechanism to write in the page data to memory using DMA, then DMA
in the tags elsewhere, and then write in the tags to be valid separately (probably done by a CPU using special instructions).
And this lack of atomicity creates a security issue (which again, can be fixed, but you have to DOCUMENT it clearly to make sure it's handled).
When allowing any DMA to user pages, the lack of atomicity means the OS
must completely unmap the page from the use before doing any tag operations.
Otherwise, the case that happens is there are 2 threads, a page is paged out, one thread touches that page, and starts in the page-in procedure. The
OS is lazy and has let the second thread have access to the page (because this is NOT a security hole without CHERI, it's the process's private data page, if it wants to "corrupt" it, it's no concern of the OS's). DMA writes in the data, and then an OS task waits for DMA to finish, then it sets the tag bits. Meanwhile, the second thread could be aggressively writing
to a pointer on the page, changing the bounds to be the whole memory space. Then the OS process sets the tag to validate the capability, and then
we've let the user forge a capability.
There's a similar race in paging out--if the OS copies the tags first,
then does the DMA reads, and it's mapped in the second user thread still,
it can similarly forge capabilities. If DMA is first and OS copies the
tags second, the user could set up the page with a forged pointer that is not a valid capability, and then write in a dummy capability before the OS
copies the tags. So the data is the forged pointer, and the tag is valid. When paged back in, the page has a valid forged capability.
So, CHERI adds a requirement that any page being used for DMA must be
fully unmapped from the user address space first. Again, some OS'es may
do this by default, but this is a security problem with CHERI and it must
be done properly. The whole idea of CXL is to allow DMA right into
user space, and that's not very compatible with CHERI.
If you think of valid tag bits as a virus that needs to be contained, you'll see there are MANY new requirements for how the OS needs to handle user
data. There needs to be no path where the user creates a malicious capability,
and then an OS operation copies what it thought was data as--- Synchronet 3.22a-Linux NewsLink 1.2
a capability. The OS must not accidentally create capabilities in the
user space. Morello's all-new-instructions actually help solve this,
by making normal integer instructions clear the tag always. But there are many possibilities here, and since I can find no discussion of how they
think they've solved this, I'm doubtful they've covered everything.
Kent
So, those are all security issues. ARM has added a 'data independent timing' >flag that can be used to ensure that a given instruction executes in constant >time, regardless of the data being operated on.
In article <2026Sep10.190203@mips.complang.tuwien.ac.at>,[...]
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
kegs@provalid.com (Kent Dickey) writes:
A CHERI-lite just providing bounds checking could be useful.
I doubt it. The problems I see here (as well as in CHERI) is that we
have nested data structures (not in every programming language, but >>certainly in some relevant ones), such as
struct foo {
char x[3];
struct bar {
char u[5];
int v[3];
} y[4];
long z[5];
} a[3];
Sometimes you want to check that you do not exceed the bounds of a, >>sometimes that you do not exceed the bounds of z, sometimes that you
do not exceed the bounds of u. Sometimes you want to treat all of a
as one thing, sometimes, only some part, sometimes your software does >>things beyond a single nested struct, such as garbage collection.
If you want to evaulate:
out = a[i].y[j].v[k];
You have a pointer to the start of the 'a' array, a, and it allows access to >sizeof(foo)*3 bytes: a.ptr = a, a.bound=3*sizeof(foo)
You do (each line is one instruction, everything is registers):
Then the evaluation is:
tmp = sizeof(foo)
off = i*tmp + offset(a[0].y[0])
a_i.ptr = a + off, a_i.bound=sizeof(bar)*4
[ The above checks that a_i is within the range a allows, then
creates a new pointer with the smaller range ]
tmp = sizeof(bar)
off = j*tmp + offset(y[0].v[0])
y_j.ptr = a_i + off, y_j.bound = 3
[ The above checks that y_j is within the range a_i allows ]
out = LOAD(y_j + k)
[ The above checks that y_j+k is within the y_j range ]
This checks that each lookup is within bounds. The CHERI overhead is the >formation of a_i and y_j, the other address calculations are needed even >without CHERI. So 6 instructions become 7 instructions.
tmp = sizeof(foo)
off = i*tmp + offset(a[0].y[0])
a_i.ptr = a + off, a_i.bound=sizeof(bar)*4
tmp = sizeof(bar)
off = j*tmp + offset(y[0].v[0])
y_j.ptr = a_i + off, y_j.bound = 3
out = LOAD(y_j + k)
[ The above checks that y_j+k is within the y_j range ]
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
George Neuner <gneuner2@comcast.net> writes:
Witness the proliferation of languages offering "managed environments"
offering such niceties as automatic storage management, automatic lock
handling (serialized object access), "comprehensions", etc., and large
standard libraries
I don't see any problem with that. These features help to implement
functionality in less programming time (and with less maintenance
time) than without using these features, and that's true for
programmers at any competence level.
without which the average programmer largely
would be incapable of producing a working program.
That's pure elitism.
Really? That's not my conclusion ... it was the result found by a
number of university studies and developer surveys.
Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
George Neuner <gneuner2@comcast.net> writes:
On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
However, I have seem many cases where people made a claim that
compilers generate better code than programmers.
"compilers generate better code than MOST programmers."
The qualification is important.
The statements I have read did not make such a qualification,
certainly not in capital letters.
There is no doubt that a good
programmer can beat the compiler,
And especially this statement is usually not made. On the contrary,
the perpetrator of such statements seem convinced of compiler
supremacy.
I have no doubt that a good programmer can beat the original FORTRAN
compiler for the IBM 704 computer, despite the fact that its
optimization was very nearly as good as that of most optimizing
compilers until at least the late 1970s.
I would not, however, be willing to assert that one could easily find
human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
I don't believe that there are any other current architectures out
there which are as nightmarish to program in assembler as the Itanium,
John Levine <johnl@taugh.com> posted:
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
However, concerning my statement above, I think the IA-64 architects
used examples for which the architecture (and its in-order
implementation) was particularly well-suited (software-pipelinable
inner loops), and that compilers eventually usually worked ok for such
examples, too. It's just that these compilers don't work so well on
general-purpose code.
That's one of the problems that Multiflow also had, works great if you
can predict the access patterns, a lot less great if the access patterns
are data dependent.
I believe that this will always be the case for memory references.
When striding through memory accessing doublewords 1 cache miss
serves 8 LDDs 7 get hits 1 takes a miss. How does one software
schedule for that pattern ??
There are several startups working on dense (10x or more) SRAM >implementations with 1ns access times as an HBM replacement.
With 1ns latency, who needs cache?
Paul Clayton <paaronclayton@gmail.com> posted:[snip]
On 9/2/26 9:50 PM, MitchAlsup wrote:
Paul Clayton <paaronclayton@gmail.com> posted:
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Yes, but You understand the difference between Architecture and implementation. May designers do not (especially early in their
careers.)
(This can reduce microarchitectural flexiblity if the software
is considered critical enough and rewriting/retuning is not an
option.)
Which is why you should strive to eliminate those things before
SW gets written.
Software that chooses to use such address bits for tags will
still run on newer designs with full 64-bit virtual addresses
though limited to 56-bit virtual addresses. I think that most
software will never need 64-bit virtual addresses. I also think
that a lot of software would not bother using address tagging.
Illustrating the small usage envelope of that mis-feature.
So, send me an e-mail address and I can send you ISA-55 and SFT-55
Paul Clayton <paaronclayton@gmail.com> posted:
Side note: I have felt some attraction to a virtual Harvard
system, i.e., a separate virtual address space for instructions.
Leads to problems for debuggers and for JIT.
Besides making writing to code more explicit and slightly
increasing the address space (less than a bit as code is rarely
half of the used memory for larger systems), such might present
opportunities to use the instruction pointer to generate data
addresses that are not cluttered with instructions.
Mc88110 had separatable Code and Data, nobody used it.
The cache chips each had their own MMU+TLB so one could go full
Harvard is they choose.
I thought of the possibility of using negative offsets from a
masked instruction pointer to provide a cheap pointer usable
with shorter offsets. Separate instruction and data address
spaces would allow
???
Paul Clayton <paaronclayton@gmail.com> posted:[snip]
I do appreciate those features of My 66000 page tables. I do
wonder if using large pages in the page table itself might be
worth supporting (Andy Glew had suggested such).
I expect that either Guest OS will be supplied with large pages
from Host OS, or vice versa. One simplifies Guest OS MMU management;
the other simplifies Host OS management. My philosophy, here, is to
allow SW e-the-large to figure it out for themselves (over time).
With page level "merging" (large pages), one could (in theory)
have 5-bit levels that could support multiples of 32 in page
size while not requiring the page table depth to be doubled.
256-KiB pages might be a useful option between 8 MiB and 8 KiB.
A PTE mapping a large page in My 66000, uses the unused PTE.PA
bits as a limit. So, while the large page still has large page
alignment; it only needs to contain as many pages as required.
So if you are using 23-bit large pages, you can use one PTE
entry (in TLB) to access exactly 173 of those pages. ...
Such also supports more flexible tradeoffs of depth versus
internal fragmentation. An application using a large address
space might have little sparsity in the upper address bits and
so benefit from merging top page table levels.
The LVL field can be used to skip the top bits (to VAS size),
then used to skip intermediate levels (sparse VAS), and then
skip the lower levels (super pages). All in one set of bits.
Level merging (really splitting) might also support more
flexible sharing of page tables. With 5-bit basic levels, an
aligned 256-byte section of a 8 KiB page table block could be
shared without sharing the rest of the translations or
permissions. (I have no idea if such would actually be useful
much less worthwhile, but it is possible.)
Hard on the TLB.
Level merging seems to be especially useful for nested page
tables in virtualization. The virtual-physical address space
is dense (I think). A virtual machine monitor might give huge
pages to a guest but could also benefit from a flatter upper
portion of the page table.
There is currently a debate going on as to whether Guest OS should
page its own pages or just have Host OS do it; Guest OS page fault
handler just never gets control.
OS support would be a problem. I do not think any architecture
provides merging of page table levels.
Not Much HW provides said feature.
[snip]One weird thought I had was to have multiple page table "bases"
loaded for a process. This would be a little like caching
directory entries with prefetching of such on context switches.
Rather than having to traverse the page table three times to
load three specific node points into the translation cache (and
likely caching other less useful information until it ages out),
the necessary entries would be prefetched and "locked".
???
Two obvious issues come to mind. First, context switches are
expected to be uncommon, so a modest savings becomes a trivial
savings. Second, this does not scale down or up to different
numbers of nodes.
???
My 66000 interrupt negotiator goes out and fetches the MSI-X
message from its-core's interrupt table, then pre-fetches the
PSL and registers while core continues running current thread.
The Fetch message and fetch context delays are similar are run
concurrently with current thread.
Once MSI-X message verifies the core should take the interrupt,
the new context is ready to install in core resources, while
pushing our current state, saving latency at each step. While
setting up the context, the first instruction is fetched, and
the ISR is in control.
I studied the Itanium architecture manuals and amused myself by writing
asm kernels that took good advantage of the large register set,
including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working variable
live on the 8-element stack, is significantly harder IMHO.
scott@slp53.sl.home (Scott Lurndal) writes:
So, those are all security issues. ARM has added a 'data independent timing'
flag that can be used to ensure that a given instruction executes in constant
time, regardless of the data being operated on.
Interesting approach. So when you use this flag on a load, it will
take the maximum time that a load can take (e.g., RAM access to RAM
under contention, or access to a remote cache under contention),
wherever it finds the item?
Intel has taken a different approach: It has documented the
instructions that are guaranteed to have data-independent timings
across all Intel microarchitectures; not sure if AMD gives the same guarantee. So if you want to write code that does not have a timing side-channel that reveals something about the data, you use only these instructions.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
John Levine <johnl@taugh.com> posted:
According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
However, concerning my statement above, I think the IA-64 architects
used examples for which the architecture (and its in-order
implementation) was particularly well-suited (software-pipelinable
inner loops), and that compilers eventually usually worked ok for such
examples, too. It's just that these compilers don't work so well on
general-purpose code.
That's one of the problems that Multiflow also had, works great if you
can predict the access patterns, a lot less great if the access patterns >> are data dependent.
I believe that this will always be the case for memory references.
When striding through memory accessing doublewords 1 cache miss
serves 8 LDDs 7 get hits 1 takes a miss. How does one software
schedule for that pattern ??
Don't use cache :-)
There are several startups working on dense (10x or more) SRAM implementations with 1ns access times as an HBM replacement.
With 1ns latency, who needs cache
On 9/13/2026 11:22 PM, Anton Ertl wrote:
scott@slp53.sl.home (Scott Lurndal) writes:
So, those are all security issues. ARM has added a 'data independent timing'
flag that can be used to ensure that a given instruction executes in constant
time, regardless of the data being operated on.
Interesting approach. So when you use this flag on a load, it will
take the maximum time that a load can take (e.g., RAM access to RAM
under contention, or access to a remote cache under contention),
wherever it finds the item?
Sure sounds ugly. :-( On a modern system, how long can this be?
On 9/7/26 2:33 PM, MitchAlsup wrote:
Has anyone ever wanted to allow shared ASIDs in such a way that shared memory or shared files have the ASID associated with the memory/file ?
So that multiple processes accessing the same shared resource co-optim-
ize themselves across cores and caches ??
On 9/13/2026 11:22 PM, Anton Ertl wrote:
Intel has taken a different approach: It has documented the
instructions that are guaranteed to have data-independent timings
across all Intel microarchitectures; not sure if AMD gives the same
guarantee. So if you want to write code that does not have a timing
side-channel that reveals something about the data, you use only these
instructions.
So loads couldn't be on that list. So you couldn't use any load >instructions in such code.
On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
George Neuner <gneuner2@comcast.net> writes:
Witness the proliferation of languages offering "managed environments"
offering such niceties as automatic storage management, automatic lock
handling (serialized object access), "comprehensions", etc., and large
standard libraries
I don't see any problem with that. These features help to implement
functionality in less programming time (and with less maintenance
time) than without using these features, and that's true for
programmers at any competence level.
without which the average programmer largely
would be incapable of producing a working program.
That's pure elitism.
Really? That's not my conclusion ... it was the result found by a
number of university studies and developer surveys.
Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.
Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
I studied the Itanium architecture manuals and amused myself by
writing asm kernels that took good advantage of the large register
set, including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working
variable live on the 8-element stack, is significantly harder
IMHO.
Stacks are hard to program for, at least for me (even large ones
where you don't get the overrun problem).
Some time ago, I wrote some rather complicated Postscript formulas.
I used two approaches: Putting a comment with the current stack
next to each instruction, and writing a quick lex/yacc grammar
to translate infix to postfix. (I also tried my hand at a
two-dimensional rootfinder in Postscript, but that was a bit
too much).
RPN simply does not come naturally to me.
I don't believe that there are any other current architectures out
there which are as nightmarish to program in assembler as the Itanium,
Huh???
I studied the Itanium architecture manuals and amused myself by
writing asm kernels that took good advantage of the large register
set, including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working
variable live on the 8-element stack, is significantly harder IMHO.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
[discussing programming in different assembly languages]
I don't believe that there are any other current architectures out
there which are as nightmarish to program in assembler as the Itanium,
Huh???
I studied the Itanium architecture manuals and amused myself by
writing asm kernels that took good advantage of the large register
set, including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working
variable live on the 8-element stack, is significantly harder IMHO.
If you don't mind my saying so, I think you have demonstrated
a sufficient level of experience and ability to safely drop
the H from IMHO. At least, it seems clear to me that you are
world class in this particular area.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:[...]
Paul Clayton <paaronclayton@gmail.com> posted:
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. [...]
John Savard wrote:
I have no doubt that a good programmer can beat the original FORTRAN
compiler for the IBM 704 computer, despite the fact that its
optimization was very nearly as good as that of most optimizing
compilers until at least the late 1970s.
I would not, however, be willing to assert that one could easily find
human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
I don't believe that there are any other current architectures out
there which are as nightmarish to program in assembler as the Itanium,
Huh???
I studied the Itanium architecture manuals and amused myself by writing
asm kernels that took good advantage of the large register set,
including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working variable
live on the 8-element stack, is significantly harder IMHO.
Brooks quoted a factor of 10 in productivity between individual
programmers in the 1960s, I suspect that gap has widened a lot
since then but haven't looked for literature.
On Mon, 14 Sep 2026 16:31:01 +0200, Terje Mathisen
<terje.mathisen@tmsw.no> wrote:
John Savard wrote:
I have no doubt that a good programmer can beat the original FORTRAN
compiler for the IBM 704 computer, despite the fact that its
optimization was very nearly as good as that of most optimizing
compilers until at least the late 1970s.
I would not, however, be willing to assert that one could easily find
human programmers who could beat an optimizing compiler... which had
the Itanium as its target.
I don't believe that there are any other current architectures out
there which are as nightmarish to program in assembler as the Itanium,
Huh???
I studied the Itanium architecture manuals and amused myself by writing
asm kernels that took good advantage of the large register set,
including the rotating ones.
Not at all a nightmare!
In fact, writing x87 asm by hand, trying to keep every working variable
live on the 8-element stack, is significantly harder IMHO.
I hadn't thought about that. None the less, I still in large part
stand by my earlier comment; ordinary programmers would have
difficulty writing good code for the Itanium, and those that can are
hard to find.
Not that the large register set is the issue. Instead, the fact that instructions come in blocks of three, with some instructions only
being available in certain slots, is what I saw as the biggest
conceptual difficulty for an ordinary human programmer.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
Scott Lurndal <scott@slp53.sl.home> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems. OTOH I see no fundamental
obstacles to having 2^64 bytes of memory. Assuming that
memory cell plus various overheads need 1000 atoms, I get
smaller amount of matter than in a current hard drives and
comparable to current memory modules. Of course, aranging
atoms into useful memory configuration will require significant
technological breaktrough, but saying that this will never
happen looks foolish.
Scott Lurndal <scott@slp53.sl.home> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
2^24 was considerd huge, later 2^32.
Simple observation is that
ie memory is available, then software will use it.
Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64.
But IMO, one can not exlude this,--- Synchronet 3.22a-Linux NewsLink 1.2
especially in biggest systems. OTOH I see no fundamental
obstacles to having 2^64 bytes of memory. Assuming that
memory cell plus various overheads need 1000 atoms, I get
smaller amount of matter than in a current hard drives and
comparable to current memory modules. Of course, aranging
atoms into useful memory configuration will require significant
technological breaktrough, but saying that this will never
happen looks foolish.
<snip>
--
Waldek Hebisch
Scott Lurndal <scott@slp53.sl.home> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems.
OTOH I see no fundamental--- Synchronet 3.22a-Linux NewsLink 1.2
obstacles to having 2^64 bytes of memory. Assuming that
memory cell plus various overheads need 1000 atoms, I get
smaller amount of matter than in a current hard drives and
comparable to current memory modules. Of course, aranging
atoms into useful memory configuration will require significant
technological breaktrough, but saying that this will never
happen looks foolish.
<snip>
antispam@fricas.org (Waldek Hebisch) posted:
Scott Lurndal <scott@slp53.sl.home> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems.
Consider a large server the size of a basketball stadium. Thousands
of racks each with dozens of systems and each motherboard maxed out
with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.
Suppose HyperVisor/OS uses the high order PA bits to route DRAM
requests from this rack-board to any other board in the stadium,
emulating a coherent system with 2^58-bytes of available DRAM by
moving pages from rack-board to rack-board on demand (or better).
So, it seems to me that we already have the capability to build
systems of that scale. Whether we do or not depends on other
factors {like "we tried and it did not perform" or simply $$$}.
antispam@fricas.org (Waldek Hebisch) posted:
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems.
Consider a large server the size of a basketball stadium. Thousands
of racks each with dozens of systems and each motherboard maxed out
with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.
Suppose HyperVisor/OS uses the high order PA bits to route DRAM
requests from this rack-board to any other board in the stadium,
emulating a coherent system with 2^58-bytes of available DRAM by
moving pages from rack-board to rack-board on demand (or better).
Each said rack-board would have its own Device address aperture
and configuration space aperture. The combined Device address
space would exceed ECAM PCIe available address space. So, by
routing device control register writes across the network, one
gets a FREE expansion of ECAM.
So, it seems to me that we already have the capability to build
systems of that scale. Whether we do or not depends on other
factors {like "we tried and it did not perform" or simply $$$}.
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
antispam@fricas.org (Waldek Hebisch) posted:
Scott Lurndal <scott@slp53.sl.home> wrote:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
Paul Clayton <paaronclayton@gmail.com> posted:
On 9/2/26 9:50 PM, MitchAlsup wrote:
[snip]
Paul Clayton <paaronclayton@gmail.com> posted:
I have wondered why Motorola did not add a 24-bit address mode
(or even provide such with a hardwired configuration on earlier
implementations, knowing that the extra bits would be desired
for other uses to save memory).
We realized our earlier mistake and did not want to repeat it.
I think this would be more a case of "extend it" than repeat it,
paying an incompatibility price to remove what was viewed as a
mistake.
[snip]
AArch64 provides a means (Top Byte Ignore) of masking the most
significant octet to allow it to be used by software.
Sins of the present...
I do not understand how allowing software to limit its future
capabilities is an architectural flaw.
When SW has used TBI often enough AND those same applications
need 63-bit (or 64-bit) virtual addresses.
Which will likely be _NEVER_. 2^64 is a, pardon my french,
shitload of virtual memory. Then there is the translation
cost with up to perhaps seven or more levels of page table
walk required. (Note that ARMv9 has support for 128-bit
page table entries [FEAT_D128] which support up to 56-bits
of PA and VA space).
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems.
Consider a large server the size of a basketball stadium. Thousands
of racks each with dozens of systems and each motherboard maxed out
with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.
Suppose HyperVisor/OS uses the high order PA bits to route DRAM
requests from this rack-board to any other board in the stadium,
emulating a coherent system with 2^58-bytes of available DRAM by
moving pages from rack-board to rack-board on demand (or better).
That's basically what we developed at 3Leaf Systems two decades
ago. We designed and fabricated an ASIC that connected to
hypertransport and extended the AMD conherency domain over infiniband
to up to 64 (at the time) independent nodes. We used the top
8 bits of the physical address as the target node address.
Fred Weber was one of our advisors.
Two decades later PCIe CXL is close to being able to support
a similar configuration.
So, it seems to me that we already have the capability to build
systems of that scale. Whether we do or not depends on other
factors {like "we tried and it did not perform" or simply $$$}.
Performance was an issue. At the time, DDR IB had cut-through
routing, so the round trip latency was about 800ns through the
switch for a non-posted request. Our Quickpath version for
Intel used QDR IB and had close to 200ns r/t latency in
simulation for a non-posted request. Unfortunately we got
caught up in the Bush recession - funding ran out and a
pending acquisition canceled.
We had a custom bare-metal hypervisor that could assign resources
from across the entire system to individual VMs. There was no
sharing of physical cores by multiple VMs, which simplified the
hypervisor somewhat. Linux could hot-plug/unplug CPUs and
memory dynamically.
Today with CXL 3.2 and GEN6 PCIe, the switching cost is significantly better.
There were also issues with cache line contention due to
spin locks (LLL had a particular test that hammered a single
lock across all CPUs), in part due to the AMD backstop
of using the global bus lock if it took to long to get exclusive
cache line access to guarantee forward progress.
On 9/16/2026 11:59 AM, MitchAlsup wrote:
antispam@fricas.org (Waldek Hebisch) posted:
snip
2^24 was considerd huge, later 2^32. Simple observation is that
ie memory is available, then software will use it. Regardless
how much memory do you have. So the only question is if enough
memory will be available. Current semiconductor technology
is advancing slower than in the past. And one can doubt if
it ever can get close to 2^64. But IMO, one can not exlude this,
especially in biggest systems.
Consider a large server the size of a basketball stadium. Thousands
of racks each with dozens of systems and each motherboard maxed out
with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.
Suppose HyperVisor/OS uses the high order PA bits to route DRAM
requests from this rack-board to any other board in the stadium,
emulating a coherent system with 2^58-bytes of available DRAM by
moving pages from rack-board to rack-board on demand (or better).
Each said rack-board would have its own Device address aperture
and configuration space aperture. The combined Device address
space would exceed ECAM PCIe available address space. So, by
routing device control register writes across the network, one
gets a FREE expansion of ECAM.
So, it seems to me that we already have the capability to build
systems of that scale. Whether we do or not depends on other
factors {like "we tried and it did not perform" or simply $$$}.
Is this essentially a large NUMA system?
In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul Clayton) wrote:
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
Having had to fit large processes into those segments, on the 32-bit
versions of both those architectures, they were a serious nuisance.
They'd apparently been designed when physical memories were measured in
small numbers of MB, with the idea that using any significant fraction of
a 4GB virtual address space was never going to happen.
John--- Synchronet 3.22a-Linux NewsLink 1.2
In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >Clayton) wrote:
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
Having had to fit large processes into those segments, on the 32-bit
versions of both those architectures, they were a serious nuisance.
They'd apparently been designed when physical memories were measured in
small numbers of MB, with the idea that using any significant fraction of
a 4GB virtual address space was never going to happen.
While I agree with pretty much everything you say in the above post,
note that the previous posts that you included above were talking about >virtual memory size, not physical memory size.
The interesting aspect is that the first 64-bit CPUs were the R4000
(1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it
was delivered); the PowerPC620 only appeared in 1997. One would think
that with these address-space restrictions the HPPA and Power
architects would feel more pressure than the others to go 64-bit
soon.
jgd@cix.co.uk (John Dallman) writes:
In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >Clayton) wrote:
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
Having had to fit large processes into those segments, on the 32-bit >versions of both those architectures, they were a serious nuisance.
They'd apparently been designed when physical memories were measured in >small numbers of MB, with the idea that using any significant fraction of
a 4GB virtual address space was never going to happen.
At least not before 64-bit addresses are available.
The interesting aspect is that the first 64-bit CPUs were the R4000
(1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it was delivered); the PowerPC620 only appeared in 1997. One would think
that with these address-space restrictions the HPPA and Power
architects would feel more pressure than the others to go 64-bit soon.
We had an Alpha with 256MB RAM in 1995, and we were not into big--- Synchronet 3.22a-Linux NewsLink 1.2
machines (this Alpha only had one CPU). So the address-space
limitations of HPPA and PowerPC probably made themselves felt strongly
before the 64-bit variants were available. Ironically, Alpha was
canceled before we reached 4GB RAM in our machines (and
general-purpose MIPS CPUs, too).
- anton
Paul Clayton <paaronclayton@gmail.com> posted:[snip]
I do not understand how the Virtual Vector Method would be
implemented efficiently. SIMD with its explicit pack and unpack
instructions seems likely to provide similar control. (Most SIMD
designs do not provide special support for short strides or
complete structure unpacking where all the data in a non-strided
stream is used but the data is scattered. GPUs might provide
special support for three and four color (un)packing.) However,
I am not a hardware designer (though I might understand a simple
flowchart ry|).
SIMD is simply direct compiler management of multi-lane calculations.
vVM allows for the processor to utilize multi-lane calculations
with flexibility SIMD can never do::
word = word + half*byte;
And if the HW happens to have 256-bits of calculation width, then
8 of those can be performed in 1 pipelined clock, without any
SW visible SIMD registers.
I do feel that VVM is a nice software interface. It avoids a
lot of SIMD issues and not just instruction diversity explosion.
vVM is 2 instructions that provide the utility of thousands (based
on typical SIMD ISA).
vVM fails at the super luminary uses of SIMD (one lone SIMD inst
without any containing loop structure)
I feel it does not exploit a few local, non-loop wide execution
cases, does not address blocking, and the loop length limits may
be introduce issues.
When a LD (or ST) touches a cache line and HW can determine the
access pattern is "dense"; HW reads out cache-width of data and
accesses that buffer multiple times feeding the multi-lane calc-
ulation. So, the wide buffering flip-flops are present (they have
to be for perf) but each implementation gets to decide how many
and how wide, and how many cache ports are "reasonable" for this implementation.
Those buffers are connected to the multi-lane calculation unit
with lots of multiplexers (which perform the pack, unpack,
width-changes, and forwarding). Those multiplexers take an
extra cycle into and out of when considering latency.
Yet I also recognize that being ideal for
all use cases regardless of complexity is both unachievable
(not all tradeoffs are limited to design difficulty or even
implementation area) and impractical (aside from complexity
having a cost, different workloads benefit differently from
"effort" and the value of the benefit is uniform across all
workloads at all scales).
While I agree that it is more complex, is it more complex than
having 1,300 SIMD instructions ??? And when we have the capability
of building 10-wide machines (Apple M5) does that really matter?
Microarchitecturally exploiting spatial locality (SRAM array
width, block size, and page size) to avoid redundant work seems
an obvious method toward power efficiency. Allowing software to
help seems reasonable to *me* (but I am excessively attracted by
hardware-software cooperation).
I don't think I am putting anything in the way of SW doing what it
wants, here.
x86 had subregisters before MMX. Even some RISCs used register
pairs for double precision floating point, presenting the same
renaming/forwarding issues.
There you go again blaming x86 as being something good in computer architecture.
On the other hand, AI turned around and is consuming all available
DRAM, so its a good thing SSDs are so fast.
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
jgd@cix.co.uk (John Dallman) writes:
In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >> >Clayton) wrote:
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
Having had to fit large processes into those segments, on the 32-bit
versions of both those architectures, they were a serious nuisance.
They'd apparently been designed when physical memories were measured in
small numbers of MB, with the idea that using any significant fraction of >> >a 4GB virtual address space was never going to happen.
At least not before 64-bit addresses are available.
The interesting aspect is that the first 64-bit CPUs were the R4000
(1991) and the 21064 (1992) while the first 64-bit HPPA machine was
introduced in November 1995 (and it probably took a while until it was
delivered); the PowerPC620 only appeared in 1997. One would think
that with these address-space restrictions the HPPA and Power
architects would feel more pressure than the others to go 64-bit soon.
CRAY-1 1976
I generally agree that a compiler will not surpass the best
performance of a supreme expert human programmer (for now).
I am guessing that effort (both compiler development and
compilation computer resources) is a significant constraint.
Dedicating three racks of servers to compile a module for a week
seems unlikely to be attractive.
Having half of the world's
programmers working on compiler development also seems unlikely
to be judged worthwhile.
Human beings also seem to have a better information caching
system. For a compiler modifying one variable name for clarity
would (I think) typical force a recompilation, possibly of the
whole program if whole program optimization is used, but a human
would probably recognize the change as non-semantic.
Some of the difficulties in developing a compiler superior to a
human being in quality of generated code are quite substantial,
but it seems that a well-engineered specialized machine should
be able to outperform a human being even in an "intellectual"
task.
Paul Clayton <paaronclayton@gmail.com> writes:
I generally agree that a compiler will not surpass the best
performance of a supreme expert human programmer (for now).
Take a look at Figure 1 of
https://www.complang.tuwien.ac.at/kps2015/proceedings/KPS_2015_submission_29.pdf
There you see the performance from a compiler from 1997 (gcc-2.7.2.3,
and gcc-2.7.0 actually appeared in 1995), from 1999 (egcs-1.1.2), and
from 2015 (gcc-5.2.0, clang-3.5). You also see different compiler optimization levels (-O0, -O3 with various -f... options for defining behaviour that the 2015 compilers treat as undefined, and -O3 with as
little language definition as the 2015 compilers use by default). And
you see a sequence of manual optimizations, from tsp1 to tsp9,
published by Jon Bentley in his 1982 book writing efficient programs;
you can see the source code for these programs by following the links
on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tsp.html>.
There you can see that the difference between the 1997 compiler and
the 2015 compilers at -O3 is usually small. Has there been much
progress since 2015? I doubt it.
You can also see that the manual optimization steps bring vast
improvements, far more than what the compilers have managed in these
18 years.
One interesting case is the step from tsp4 to tsp5
(inlining of a function), which is flat for all the compilers, so all
the compilers (from 1997 to 2015) do it by themselves; it is a
prerequisite to following manual optimizations, so it cannot be left
away in the manual optimization sequence.
In the meantime I know that this program can be vectorized (and we
have discussed manual vectorization at the assembly/intrinsic level
here in 2016).
- anton--- Synchronet 3.22a-Linux NewsLink 1.2
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
jgd@cix.co.uk (John Dallman) writes:
In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >> >Clayton) wrote:
I suspect something like this (but coarser-grained) was the
motivation for PowerPC's segments. Effectively each of 16
address regions (in the 32-bit "effective" address space) had an
ASID. (Virtual segment IDs were 24 bits.)
software ported from HP PA-RISC (which had segments).
Having had to fit large processes into those segments, on the 32-bit
versions of both those architectures, they were a serious nuisance.
They'd apparently been designed when physical memories were measured in >> >small numbers of MB, with the idea that using any significant fraction of >> >a 4GB virtual address space was never going to happen.
At least not before 64-bit addresses are available.
The interesting aspect is that the first 64-bit CPUs were the R4000
(1991) and the 21064 (1992) while the first 64-bit HPPA machine was
introduced in November 1995 (and it probably took a while until it was
delivered); the PowerPC620 only appeared in 1997. One would think
that with these address-space restrictions the HPPA and Power
architects would feel more pressure than the others to go 64-bit soon.
CRAY-1 1976
So one might subclassify 64-bit into:
- 64-bit arithmetic operations
- 64-bit addressing operations.
The B3500 in 1965 did 400-bit arithmetic operations (100 digit),--- Synchronet 3.22a-Linux NewsLink 1.2
applications were limited to 500KB[*] (code + data - slightly less
as a small amount was used by the MCP (OS)). Later machines
increased the physical address space from 6 digits to 10 digits
(500MB).
[*] 1 million digits/nibbles.
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
And
you see a sequence of manual optimizations, from tsp1 to tsp9,
published by Jon Bentley in his 1982 book writing efficient programs;
you can see the source code for these programs by following the links
on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tsp.html>.
There you can see that the difference between the 1997 compiler and
the 2015 compilers at -O3 is usually small. Has there been much
progress since 2015? I doubt it.
You can also see that the manual optimization steps bring vast
improvements, far more than what the compilers have managed in these
18 years.
This is the "find a better algorithm" step of making programs fast.
If one dives into LINPACK and LAPACK one finds FORTRAN code written
in such a way that array-access-order is cache friendly. No compiler
is ever going to do that automagically for general array code where
the program specifies the indexing.
In the meantime I know that this program can be vectorized (and we
have discussed manual vectorization at the assembly/intrinsic level
here in 2016).
I can't seem to find the source code ... I would like to try vectorizing
with My 66000 ISA.
In article <2026Sep17.113215@mips.complang.tuwien.ac.at>, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
The interesting aspect is that the first 64-bit CPUs were the R4000
(1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it
was delivered); the PowerPC620 only appeared in 1997. One would
think that with these address-space restrictions the HPPA and Power architects would feel more pressure than the others to go 64-bit
soon.
Do not underestimate the power of corporate conservatism and sectional interests. From 1995-2005 I regularly backstopped for technical
support staff who were finding that customers felt they were locked
into one particular commercial UNIX and could not change without vast disruptions.
We felt this was weird, because the UNIXes of the era were all pretty similar. It gradually became clear that manufacturers' training
courses on commercial UNIXes emphasised the differences and tried
hard to give the impression that changing to a another supplier would
be difficult and expensive. Human inertia meant that customers'
sysadmins co-operated with that, giving the illusion that they were indispensable while acting against their employers' interests. This
meant that well-established manufacturers like HP and IBM felt less
urgency to move to 64-bit. Sun and SGI were a bit more dynamic, until
they got into financial trouble, and DEC knew they needed Alpha to
have a hope of survival.
Meanwhile, we employed sysadmins who dealt with all of these UNIXes
every week without difficulty. We also had to learn about things like
POWER and PA-RISC segmentation to squeeze large 32-bit applications
onto them. 64-bit was very welcome!
John
HP introduced their first 64-bit PA-RISC CPU only 5 or 6 months behind
Sun.
IBM released the Cobra, their 1st 64-bit POWEER CPU, approximately at
the same time as Sun. But it and its successors (Muskie, Apache,
Northstar) were probably of little interest for you and your employer, >because they were not intended for engineering applications.
The first engineering-oriented 64-bit POWER (POWER3) came, indeed, >significantly later - more than 3 years behind Sun.
ordinary programmers would have
difficulty writing good code for the Itanium, and those that can are
hard to find.
Instead, the fact that
instructions come in blocks of three, with some instructions only
being available in certain slots, is what I saw as the biggest
conceptual difficulty for an ordinary human programmer.
I suspect algorithmic transformations are discouraged in
compilers under the assumption that the software developer used
appropriate algorithms as well as the compute cost to evaluate
options.
Paul Clayton [2026-09-16 13:14:22] wrote:
I suspect algorithmic transformations are discouraged in
compilers under the assumption that the software developer used
appropriate algorithms as well as the compute cost to evaluate
options.
Also because compilers generally can't (or at least shouldn't) take the chance of generating significantly slower code.
The purpose of "language + compiler" is not to figure out magically how
to implement the most efficient code that solves the same problem,
instead it's to allow the programmer to write that most efficient
code conveniently.
Of course, there is a lot of commercial interest in doing the magic
thing so as to save the work of understanding and rewriting the code to improve performance (hence all the work on auto-parallelization), but experience shows that it's an extremely difficult problem.
Luckily, a lot of that work on trying to do the magic thing can also be
used to solve the real problem: what the years of work on compiler-optimization has brought (instead of magically speeding up old
code) is to improve the convenience to write efficient code.
=== Stefan
Paul Clayton [2026-09-16 13:14:22] wrote:
Also because compilers generally can't (or at least shouldn't) take the >chance of generating significantly slower code.
The purpose of "language + compiler" is not to figure out magically how
to implement the most efficient code that solves the same problem,
instead it's to allow the programmer to write that most efficient
code conveniently.
Of course, there is a lot of commercial interest in doing the magic
thing so as to save the work of understanding and rewriting the code to >improve performance (hence all the work on auto-parallelization), but >experience shows that it's an extremely difficult problem.
Luckily, a lot of that work on trying to do the magic thing can also be
used to solve the real problem: what the years of work on >compiler-optimization has brought (instead of magically speeding up old
code) is to improve the convenience to write efficient code.
I suspect algorithmic transformations are discouraged in
compilers under the assumption that the software developer used
appropriate algorithms as well as the compute cost to evaluate
options.
Usage and hardware information might theoretically be
made available, but I get the impression that most profiling
for compiler use is about path frequency.
Paul Clayton <paaronclayton@gmail.com> schrieb:
Two points: A compiler is restricted by the transformations
it can do by the language specification. For example, loop
interchanges with floating point calculation can give different
results.
High-level transformation would transform Anton's beloved
bubblesort benchmark into insertion sort, at least.
But if you're willing to take the risk of AI, you can of course
tell it "Find all O(N^2) algorithms in my code and replace them
with O(N log N) where possible".
Thomas Koenig <tkoenig@netcologne.de> writes:
Paul Clayton <paaronclayton@gmail.com> schrieb:
Two points: A compiler is restricted by the transformations
it can do by the language specification. For example, loop
interchanges with floating point calculation can give different
results.
Interestingly, this is where gcc maintainers do the right thing and
people who want to risk program breakage must use an extra flag
(-ffast-math) to get "optimizations" (such as FP operation
reassociation) that may change the results.
High-level transformation would transform Anton's beloved
bubblesort benchmark into insertion sort, at least.
This bubble-sort benchmark comes from John Hennessey's collection of
integer benchmarks, and Marty Fraeman has translated several of these benchmarks into Forth. This is the reason why I like to use it for
comparing the performance of Forth systems to C compilers.
Recognizing this bubble-sort and using a more efficient sort (e.g.,
insertion sort for small instances, quicksort for larger ones) would
be a real optimization; it would ruin the comparability with the Forth translation (as long as we do not put this kind of effort into Forth compilers), but that's life. But for now gcc autovectorizes it in a
way that produces a significant slowdown; clang avoids this pitfall.
Note that this auto-vectorization of gcc does not just slow down
bubble-sort, but also gforth: when you load two adjacent stack items
in one VM instruction, gcc -O3 by default wants to auto-vectorize
these accesses, and given that one or both of them usually have been
written recently, this would result in a slowdown. What's worse, gcc
seems to think that a lot of values are alive that are actually dead,
and generates dozens or hundreds of copying instructions into the code
of each VM instruction, increasing the slowdown even more. For that
reason, we have disabled tree-slp-vectorization when building Gforth.
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
Thomas Koenig <tkoenig@netcologne.de> writes:
Paul Clayton <paaronclayton@gmail.com> schrieb:
Two points: A compiler is restricted by the transformations
it can do by the language specification. For example, loop
interchanges with floating point calculation can give different
results.
Interestingly, this is where gcc maintainers do the right thing and
people who want to risk program breakage must use an extra flag (-ffast-math) to get "optimizations" (such as FP operation
reassociation) that may change the results.
High-level transformation would transform Anton's beloved
bubblesort benchmark into insertion sort, at least.
This bubble-sort benchmark comes from John Hennessey's collection of integer benchmarks, and Marty Fraeman has translated several of
these benchmarks into Forth. This is the reason why I like to use
it for comparing the performance of Forth systems to C compilers.
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
Recognizing this bubble-sort and using a more efficient sort (e.g., insertion sort for small instances, quicksort for larger ones) would
be a real optimization; it would ruin the comparability with the
Forth translation (as long as we do not put this kind of effort
into Forth compilers), but that's life. But for now gcc
autovectorizes it in a way that produces a significant slowdown;
clang avoids this pitfall.
All such choices are the result of heuristics. Bubble sort has a
very special memory access pattern. A straightforward patch would
very likely pessimize a lot of existing code which profits from auto-vectorization. I have no doubt that the existing heuristics
can be improved.
Note that this auto-vectorization of gcc does not just slow down bubble-sort, but also gforth: when you load two adjacent stack items
in one VM instruction, gcc -O3 by default wants to auto-vectorize
these accesses, and given that one or both of them usually have been written recently, this would result in a slowdown. What's worse,
gcc seems to think that a lot of values are alive that are actually
dead, and generates dozens or hundreds of copying instructions into
the code of each VM instruction, increasing the slowdown even more.
For that reason, we have disabled tree-slp-vectorization when
building Gforth.
There is no stopping you from developing a patch (or having
it developed by a student as part of a thesis - can you act as
supervisor for master's or bachelor's theses?) and then submitting
it to gcc. But if you do so, you should make sure that it is does
not slow down normal programs (as measured by SPEC).
Regarding clang: There are numerous cases where gcc's more
aggressive auto-vectorization generates a lot of profit vs. clang.
These are cases that would very probably suffer when heuristics
are not very carefully chosen.
On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
But one particular pattern of gcc is certainly worse than clang's -
merging narrow stores into wider ones. Nowadays gcc does not even
consider it autovectorization. This pattern rarely leads to major
slowdowns, more typically a little impacts in multiple places. But
sometimes it gets measurable, esp. on Zen3.
I don't believe that disabling this particular generic pessimization
could impact Spec negatively, but I am not aware of simple way (i.e.
normal -f flags, rather than semi-documented flags intended for gcc maintainers) to disable it.
Michael S <already5chosen@yahoo.com> schrieb:
On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
But one particular pattern of gcc is certainly worse than clang's -
merging narrow stores into wider ones. Nowadays gcc does not even
consider it autovectorization. This pattern rarely leads to major slowdowns, more typically a little impacts in multiple places. But sometimes it gets measurable, esp. on Zen3.
I don't believe that disabling this particular generic pessimization
could impact Spec negatively, but I am not aware of simple way (i.e.
normal -f flags, rather than semi-documented flags intended for gcc maintainers) to disable it.
If you have a test case (but please not bubble sort :-) where
-fstore-merging pessimizes things (or -fno-store-merging makes
a significant difference) please post it. I can then write a PR
and hang it off https://gcc.gnu.org/bugzilla/show_bug.cgi?id=94094
where I don't see anything relevant at the moment.
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
Thomas Koenig <tkoenig@netcologne.de> writes:
Paul Clayton <paaronclayton@gmail.com> schrieb:
Two points: A compiler is restricted by the transformations
it can do by the language specification. For example, loop
interchanges with floating point calculation can give different
results.
Interestingly, this is where gcc maintainers do the right thing and
people who want to risk program breakage must use an extra flag
(-ffast-math) to get "optimizations" (such as FP operation
reassociation) that may change the results.
High-level transformation would transform Anton's beloved
bubblesort benchmark into insertion sort, at least.
This bubble-sort benchmark comes from John Hennessey's collection of
integer benchmarks, and Marty Fraeman has translated several of these
benchmarks into Forth. This is the reason why I like to use it for
comparing the performance of Forth systems to C compilers.
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
Recognizing this bubble-sort and using a more efficient sort (e.g.,
insertion sort for small instances, quicksort for larger ones) would
be a real optimization; it would ruin the comparability with the Forth
translation (as long as we do not put this kind of effort into Forth
compilers), but that's life. But for now gcc autovectorizes it in a
way that produces a significant slowdown; clang avoids this pitfall.
All such choices are the result of heuristics. Bubble sort has a
very special memory access pattern. A straightforward patch would
very likely pessimize a lot of existing code which profits from >auto-vectorization. I have no doubt that the existing heuristics
can be improved.
Note that this auto-vectorization of gcc does not just slow down
bubble-sort, but also gforth: when you load two adjacent stack items
in one VM instruction, gcc -O3 by default wants to auto-vectorize
these accesses, and given that one or both of them usually have been
written recently, this would result in a slowdown. What's worse, gcc
seems to think that a lot of values are alive that are actually dead,
and generates dozens or hundreds of copying instructions into the code
of each VM instruction, increasing the slowdown even more. For that
reason, we have disabled tree-slp-vectorization when building Gforth.
There is no stopping you from developing a patch (or having
it developed by a student as part of a thesis - can you act as
supervisor for master's or bachelor's theses?) and then submitting
it to gcc.
But if you do so, you should make sure that it is does
not slow down normal programs (as measured by SPEC).
Regarding clang: There are numerous cases where gcc's more
aggressive auto-vectorization generates a lot of profit vs. clang.
These are cases that would very probably suffer when heuristics
are not very carefully chosen.
But one particular pattern of gcc is certainly worse than clang's -
merging narrow stores into wider ones.
Michael S <already5chosen@yahoo.com> writes:
But one particular pattern of gcc is certainly worse than clang's -
merging narrow stores into wider ones.
I have seen heavy slowdowns from merging narrow loads into wide loads <https://www.complang.tuwien.ac.at/anton/stwlf/>, but not for stores.
Where can I read about this slowdown?
- anton
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
Stefan Monnier <monnier@iro.umontreal.ca> writes:
Paul Clayton [2026-09-16 13:14:22] wrote:One would hope so, but:
Also because compilers generally can't (or at least shouldn't) take the >>chance of generating significantly slower code.
Experience shows that it is an easy problem to get people to put money
into your project by promising them that the compiler will magically
solve it (and thus they can save on expensive programmer time);
earlier successful ways to convert wishful thinking into money are the philosopher's stone, and a contemporary way is AI. In the case of
optimizing compilers it helps that one can present showpieces where it
works as promised. I guess the alchemists had similar tricks (maybe
they sold chemical reactions as first steps to the desired end
result), and we can see in daily news how the AI companies do it.
Luckily, a lot of that work on trying to do the magic thing can also be >>used to solve the real problem: what the years of work on >>compiler-optimization has brought (instead of magically speeding up old >>code) is to improve the convenience to write efficient code.It depends.
But if optimizer writers strove to "improve the convenience to write efficient code" instead of just improving benchmark results, maybe
such paradoxical effects can be avoided.
It seems to me, the most heavy impact on Zen3 is when such merged
SIMD store is followed by GPR load to the same location within a dozen
or so of CPU cycles.
ws=_=>nl>ns line. In this case, on Zen3 merging (-O3) is fasterthan not merging (-O).
Anton Ertl [2026-09-19 06:25:13] wrote:
Stefan Monnier <monnier@iro.umontreal.ca> writes:
Paul Clayton [2026-09-16 13:14:22] wrote:One would hope so, but:
Also because compilers generally can't (or at least shouldn't) take the >>>chance of generating significantly slower code.
Indeed, in practice virtually all compiler "optimizations" are valid
only statistically: in most cases they either have no measurable effect
or they improve some characteristic, but there are almost always corner
cases where they make things worse.
On Sun, 20 Sep 2026 11:52:38 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
Michael S <already5chosen@yahoo.com> schrieb:
On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
But one particular pattern of gcc is certainly worse than clang's -
merging narrow stores into wider ones. Nowadays gcc does not even
consider it autovectorization. This pattern rarely leads to major
slowdowns, more typically a little impacts in multiple places. But
sometimes it gets measurable, esp. on Zen3.
I don't believe that disabling this particular generic pessimization
could impact Spec negatively, but I am not aware of simple way (i.e.
normal -f flags, rather than semi-documented flags intended for gcc
maintainers) to disable it.
If you have a test case (but please not bubble sort :-) where
-fstore-merging pessimizes things (or -fno-store-merging makes
a significant difference) please post it. I can then write a PR
and hang it off https://gcc.gnu.org/bugzilla/show_bug.cgi?id=94094
where I don't see anything relevant at the moment.
IIRC, it already was in at least one of my PRs. May be, more than one.
All such choices are the result of heuristics. Bubble sort has a
very special memory access pattern.
A straightforward patch would
very likely pessimize a lot of existing code which profits from >auto-vectorization.
Michael S <already5chosen@yahoo.com> writes:
It seems to me, the most heavy impact on Zen3 is when such merged
SIMD store is followed by GPR load to the same location within a
dozen or so of CPU cycles.
This is a case I measured in <https://www.complang.tuwien.ac.at/anton/stwlf/>. Look for the wl>ws=_=>nl>ns line. In this case, on Zen3 merging (-O3) is faster
than not merging (-O).
But I see slowdowns from merging (-O3) in other cases where the
partial store-to-load overlap does not occur, e.g., in the "wl>ws=>wl (recurrence), nl>ns (no recurrence)" case (and that's pretty
widespread among microarchitectures), so there is probably another microarchitectural pitfall beyond the partal store-to-load forwarding.
- anton
Anton Ertl [2026-09-19 06:25:13] wrote:[...]
The traditional "-O<N>" flags are a way to state how much you're willing
to get worse performance in exchange for the opportunity to maybe get
better performance.
But if optimizer writers strove to "improve the convenience to write
efficient code" instead of just improving benchmark results, maybe
such paradoxical effects can be avoided.
In general it's hard to completely avoid paradoxical effects.
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
Stefan Monnier <monnier@iro.umontreal.ca> writes:-------------
Anton Ertl [2026-09-19 06:25:13] wrote:
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows. It gives a speedup in some SPEC
program (How much? Don't know, but I am sure it exists); ok, it
breaks this other SPEC program, let's special-case the optimization to
avoid this breakage.
My impression is that the C compiler maintainers and some others have
fallen in love with the idea of optimizing based on assuming that
undefined behaviour does not happen; e.g., I have read the claim that
this is the only thing that gives C an edge over other programming
languages (or somesuch). So I think that while the C89 committee may
not have thought of such things, recent C standardization committees
probably have.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus: <https://blog.regehr.org/archives/1287>.
But fortunately a completely specified C is not necessary, a
willingness to preserve the behaviour of existing working programs
compiled with an earlier version of the same compiler is. Read more
about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.
Linux commits to preserving user-space behaviour (whether standard or
not), everything I have heard from people claiming to speak for gcc
and clang maintainers has been in the opposite direction. So I blame
the gcc and clang maintainers for the undefined-behaviour shenanigans
they perform.
- anton--- Synchronet 3.22a-Linux NewsLink 1.2
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Stefan Monnier <monnier@iro.umontreal.ca> writes:-------------
Anton Ertl [2026-09-19 06:25:13] wrote:
But as
for the problems you mention w.r.t UB, I think it's just the
result of poor semantics, for which I guess we (language
semanticists) are partly to blame: we have developed fairly good
tools to design sane language specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people
were poor at specifying semantics, but because given the differences between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts
unspecified.
Does anyone know if Ada was successful about not leaving parts
unspecified ??
According to Stefan Monnier <monnier@iro.umontreal.ca>:
Indeed, in practice virtually all compiler "optimizations" are validThere are plenty of optimizations that always make things better, e.g., removing dead code, or reusing values in registers. But these days those
only statistically: in most cases they either have no measurable effect
or they improve some characteristic, but there are almost always corner >>cases where they make things worse.
are so obvious we sometimes forget about them.
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows.
My impression is that the C compiler maintainers and some others have
fallen in love with the idea of optimizing based on assuming that
undefined behaviour does not happen; e.g., I have read the claim that
this is the only thing that gives C an edge over other programming
languages (or somesuch).
So I think that while the C89 committee may
not have thought of such things, recent C standardization committees
probably have.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus:
<https://blog.regehr.org/archives/1287>.
Thomas Koenig <tkoenig@netcologne.de> writes:
All such choices are the result of heuristics. Bubble sort has a
very special memory access pattern.
Ok, what would a memory pattern look like that this vectorization was designed for? The big slowdown comes from the fact that in the
reordering case the vectorized version performs a wide store, and in
the next iteration it performs a wide load that partially overlaps the
store. I buy it that the analysis will have trouble seeing that (not
that it's impossible in this case, but it requires effort).
A straightforward patch would
very likely pessimize a lot of existing code which profits from >>auto-vectorization.
How can we test this claim?
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Stefan Monnier <monnier@iro.umontreal.ca> writes:-------------
Anton Ertl [2026-09-19 06:25:13] wrote:
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
Does anyone know if Ada was successful about not leaving parts
unspecified ??
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows. It gives a speedup in some SPEC
program (How much? Don't know, but I am sure it exists); ok, it
breaks this other SPEC program, let's special-case the optimization to
avoid this breakage.
This is a problem with SPEC not the compilers for various machines.
A broad spectrum benchmark should not contain code that relies on
unspecified or undefined behaviors.
My impression is that the C compiler maintainers and some others have
fallen in love with the idea of optimizing based on assuming that
undefined behaviour does not happen; e.g., I have read the claim that
this is the only thing that gives C an edge over other programming
languages (or somesuch). So I think that while the C89 committee may
not have thought of such things, recent C standardization committees
probably have.
How much more optimizations can compilers deliver to the bottom line
(not just benchmarks).
Last week we saw an episode where the compilers
only got 1.x% speedup over years (sub-decade). Is there ever going to
be a time to stop ??
Linux commits to preserving user-space behaviour (whether standard or
not), everything I have heard from people claiming to speak for gcc
and clang maintainers has been in the opposite direction. So I blame
the gcc and clang maintainers for the undefined-behaviour shenanigans
they perform.
Some of the machines have egregious behaviors in tiny little corners
of the architecture(s) and implementation(s). Not being able to run
the same binary with and without integer overflow detection as an example.
And even when they can, they cannot perform Ada ADD Byte and take a signed >overflow trap on 8-bit results easily.
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows.
Define "miscompile".
According to which specification?
In Fortran, for example, signed
integer overflow is just an error, so a compiler that would not
optimize on the assumption of absence of integer overflow would
be doing a poor jobs.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus: >><https://blog.regehr.org/archives/1287>.
People who want that kind of language know where to find it,
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
Stefan Monnier <monnier@iro.umontreal.ca> writes:-------------
Anton Ertl [2026-09-19 06:25:13] wrote:
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
Does anyone know if Ada was successful about not leaving parts
unspecified ??
Stefan Monnier <monnier@iro.umontreal.ca> writes:
Anton Ertl [2026-09-19 06:25:13] wrote:[...]
The traditional "-O<N>" flags are a way to state how much you're willing
to get worse performance in exchange for the opportunity to maybe get
better performance.
The gcc manual states:
With '-O', the compiler tries to reduce code size and execution
time, without performing any optimizations that take a great deal
of compilation time.
[...]
[about -O2] As
compared to '-O', this option increases both compilation time and
the performance of the generated code.
[...]
[about -O3] Optimize yet more.
In earlier gcc versions, it made statements along the lines that -O3
may be a mixed bag, but that's gone.
But if optimizer writers strove to "improve the convenience to write
efficient code" instead of just improving benchmark results, maybe
such paradoxical effects can be avoided.
In general it's hard to completely avoid paradoxical effects.
I referred to a specific paradoxical effect, not completely avoiding
all of them: The effect of programmers who, instead of working on
making the code faster through source-level changes, have to invest
time into avoiding getting it miscompiled.
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows. It gives a speedup in some SPEC
program (How much? Don't know, but I am sure it exists); ok, it
breaks this other SPEC program, let's special-case the optimization to
avoid this breakage.
My impression is that the C compiler maintainers and some others have
fallen in love with the idea of optimizing based on assuming that
undefined behaviour does not happen; e.g., I have read the claim that
this is the only thing that gives C an edge over other programming
languages (or somesuch). So I think that while the C89 committee may
not have thought of such things, recent C standardization committees
probably have.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus: <https://blog.regehr.org/archives/1287>.
But fortunately a completely specified C is not necessary, a
willingness to preserve the behaviour of existing working programs
compiled with an earlier version of the same compiler is. Read more
about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.
Linux commits to preserving user-space behaviour (whether standard or
not), everything I have heard from people claiming to speak for gcc
and clang maintainers has been in the opposite direction. So I blame
the gcc and clang maintainers for the undefined-behaviour shenanigans
they perform.
Paul Clayton <paaronclayton@gmail.com> schrieb:
I suspect algorithmic transformations are discouraged in
compilers under the assumption that the software developer used
appropriate algorithms as well as the compute cost to evaluate
options.
Two points: A compiler is restricted by the transformations
it can do by the language specification. For example, loop
interchanges with floating point calculation can give different
results. Unless directed otherwise, the compiler has to assume
that this is what the user wants and needs.
An exception is something like Fortran's FORALL constuct, where
the loop ordering is not specified and the compiler can chose.
Intrinsic functions like MATMUL also have no restriction on what
they can do internally; they can (and often do) call a highly-
optimized BLAS routine.
Usage and hardware information might theoretically be
made available, but I get the impression that most profiling
for compiler use is about path frequency.
Compilers tend to focus on lower-level stuff. For example,
there is a huge file of simplifying transformations for gcc at https://gcc.gnu.org/git/?p=gcc.git;a=blob_plain;f=gcc/match.pd
This now also contains reverse-engineered bithacks. For example,
the popcnt() method from Hacker's Delight are now recognized and
translated into an internal function, which is then expanded
into either an inlined function (much like the original) or
a machine instruction, if the target has one.
High-level transformation would transform Anton's beloved
bubblesort benchmark into insertion sort, at least.
But if you're willing to take the risk of AI, you can of course
tell it "Find all O(N^2) algorithms in my code and replace them
with O(N log N) where possible". It will happily do something,
and the resulting code might even be correct after a few
iterations.
Thomas Koenig <tkoenig@netcologne.de> writes:
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
For people who don't like bubble sort, or for inclusion in a set
of benchmarks, I suggest the following recently discovered
sorting algorithm (written in C-ish pseudocode):
void
baffle_sort( unsigned n, int *elements ){
for( unsigned i = 0; i < n; i++ ){
for( unsigned j = 0; j < n; j++ ){
if( elements[i] < elements[j] ){
/* exchange elements[i] and elements[j] */
}
}
}
}
(Disclaimer: not original.)
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
How much more optimizations can compilers deliver to the bottom line
(not just benchmarks).
Number of optimizations? There have been a number of cases where I
have noticed that gcc misses a real optimization. The tsp example
also shows that a lot of the optimizations that Jon Bentley performed
in 1982 are not performed by compilers.
Effect of new optimizations or "optimizations" based on assuming that undefined behaviour never happens on performance? They never say.
That's the cool thing. They do not have numbers (they certainly never present any, certainly not for their own compilers), but are convinced
that their "optimizations" do wonders for performance, and their
fanboys are even more convinced.
Matrix300 was eliminated from SPEC because all the computer companies
started to use automatic cache blocking. IIRC this was for the step
from SPEC89 to SPEC92.
Later Sun managed to optimize IIRC the ear
benchmark to perform array-of-structures to structure-of-arrays transformation, achieved a speedup by a factor of IIRC 2, and a
significant increase in the aggregate SPEC score.
Tim Rentsch wrote:
Thomas Koenig <tkoenig@netcologne.de> writes:
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
For people who don't like bubble sort, or for inclusion in a set
of benchmarks, I suggest the following recently discovered
sorting algorithm (written in C-ish pseudocode):
void
baffle_sort( unsigned n, int *elements ){
for( unsigned i = 0; i < n; i++ ){
for( unsigned j = 0; j < n; j++ ){
if( elements[i] < elements[j] ){
/* exchange elements[i] and elements[j] */
}
}
}
}
(Disclaimer: not original.)
Did you watch one of my favorite Youtube channels?
:-)
Are you analyzing this specific function for a computer scienceassignment / code review, or would you like me to help you rewrite
If a student can't rewrite it themselves, they shouldn't be programming. AI>That is a fair and candid perspective. Rewriting or fixing this specificfunction is a foundational exercise in control flow, loops, and basic logic. If someone
It highlights the difference between memorizing syntax and truly understandingcode execution.
Are you grading student submissions that included this specific code, or are youanalyzing a textbook/exam problem designed to trip them up?
I saw the code snippet in comp.archare notorious for regulars posting bizarre, intentionally "baffling" code fragments to
Ah, comp.arch - that explains it perfectly. USENET groups like comp.arch and comp.lang.c
On 21/09/2026 19:17, Anton Ertl wrote:
Stefan Monnier <monnier@iro.umontreal.ca> writes:
Anton Ertl [2026-09-19 06:25:13] wrote:[...]
The traditional "-O<N>" flags are a way to state how much you're willing >>> to get worse performance in exchange for the opportunity to maybe get
better performance.
The gcc manual states:
-a-a-a-a-a With '-O', the compiler tries to reduce code size and execution >> -a-a-a-a-a time, without performing any optimizations that take a great deal >> -a-a-a-a-a of compilation time.
[...]
-a-a-a-a-a [about -O2] As
-a-a-a-a-a compared to '-O', this option increases both compilation time and >> -a-a-a-a-a the performance of the generated code.
[...]
-a-a-a-a-a [about -O3] Optimize yet more.
In earlier gcc versions, it made statements along the lines that -O3
may be a mixed bag, but that's gone.
It must have been removed a /long/ time ago, because it is not in manual pages that I checked (the oldest convenient version is 2.95.3).-a But it
is certainly the case that -O3 is a mixed bag - in particular, more aggressive unrolling and inlining can lead to larger code that can
reduce the effect of caches, branch prediction, and that kind of thing.
In microcontroller work, -Os is often used to put a stronger emphasis on size optimisations.-a I have noticed there are sometimes "blips" - cases where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
faster code than "-O2".-a And there can sometimes be cases where the speed/size tradeoffs are unreasonable - significantly slower code for
very minor size improvements, or vice versa.
It is clearly not the case that all optimisations improve all code, or
that increasing "n" in "-On" always gives faster results.-a Each optimisation flag in a compiler enables one or more transformation
passes that might improve some code but might also have detrimental
effects in some cases.-a Lower "-On" numbers will include the passes with high statistical rates of improvement and low risk of worsening code,
while the passes enabled with higher flags will have worse ratios and
longer compiler times.-a Flags with significant risks of making code a
lot worse are usually not enabled by any "-On" flag, and require manual choice.-a (But no one will claim that gcc, clang, or any other compiler
is perfect here.)
I think one thing that is sometimes forgotten in all this, and could usefully be mentioned on the gcc manual page for optimisation, is that
cpu architecture flags can make a significant difference.-a The step from "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
from "-O2" to "-O3" in performance, with much lower risks of slowdowns.
But if optimizer writers strove to "improve the convenience to write
efficient code" instead of just improving benchmark results, maybe
such paradoxical effects can be avoided.
In general it's hard to completely avoid paradoxical effects.
I referred to a specific paradoxical effect, not completely avoiding
all of them: The effect of programmers who, instead of working on
making the code faster through source-level changes, have to invest
time into avoiding getting it miscompiled.
By "miscompiled", do you mean incorrect object code, or object code that
did not have the performance the programmer expected or hoped for?
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree.-a The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
I have never spoken to the C standards committee or writers, either of current standard versions or the original C89 standard, or writers of pre-standard C specifications.-a So I cannot in any way claim to know
their thoughts or motivations.-a I also have not seen any documentation
that suggests what they might have thought about "optimisation on the assumption that undefined behaviour does not occur" - in either
direction.-a (It has been mentioned that, for example, a two's complement implementation could use wrapping signed arithmetic - but I have never
seen a suggestion that this behaviour should be encouraged or expected
just because a machine uses two's complement.)
I do, however, believe that the folks being the C language design and specifications through all its changes are accomplished and experienced computer scientists.-a They will have been aware of the "garbage in,
garbage out" principle, and that there is no reason to expect any
particular result or effect when you apply a function or operator to something outside its defined semantics.-a You do not ask "what happens
if my signed integer arithmetic overflows?" or "what happens when I
access an array out of bounds?" - rather, it is your responsibility as a
C programmer to make sure that never happens.-a The prime reason C does
not define behaviour here is not that different hardware or compilers
handle things differently, but that there is no sensible definition that could be made.
Remember, C has a perfectly good way to say "this is determined by the hardware or the implementation" - it is "implementation-defined behaviour".-a It has a perfectly good way to say that "this operation
could result in any value" - it is "unspecified behaviour" or
"unspecified value".
When the C standards writers say something is "undefined behaviour",
rather than "implementation defined" or "unspecified", it is my belief
that they did so intentionally and knowingly.-a I have at times been
accused of arrogance, usually quite fairly, but I am not arrogant enough
to suppose that I know when the C standards committee made mistakes here
or intended to write something differently.
So I cannot accept an argument that C's undefined behaviours are either
due to hardware differences, or laziness.-a It simply does not fit with
the level of expertise and effort that have gone into the language and
its standards.
And note that in C23 the macro "unreachable()" was added with the sole semantics being "If a macro invocation unreachable() is reached during execution, the behavior is undefined" and "The program execution shall
not reach such an invocation".
Tim Rentsch wrote:
Thomas Koenig <tkoenig@netcologne.de> writes:
Bubblesort is the worst of non-joke sorting algorithms, see
the quote from Numerical Recipes...
For people who don't like bubble sort, or for inclusion in a set
of benchmarks, I suggest the following recently discovered
sorting algorithm (written in C-ish pseudocode):
void
baffle_sort( unsigned n, int *elements ){
for( unsigned i = 0; i < n; i++ ){
for( unsigned j = 0; j < n; j++ ){
if( elements[i] < elements[j] ){
/* exchange elements[i] and elements[j] */
}
}
}
}
(Disclaimer: not original.)
Did you watch one of my favorite Youtube channels?
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the sole
semantics being "If a macro invocation unreachable() is reached during
execution, the behavior is undefined" and "The program execution shall
not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?
Thomas Koenig <tkoenig@netcologne.de> writes:
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
I don't think that they expected that compiler writers on a
twos-complement machine that does not trap on signed overflow would
say: Hey, signed overflow is undefined behaviour, let's assume it
never happens, and silently miscompile some (not all) programs that
actually perform signed overflows.
Define "miscompile".
If a compiler produces an unintended behaviour for an existing, tested program that works as intended with an earlier version of the compiler.
According to which specification?
If you want that, go for the specification of a conforming program in
the C standard.
Compiling a conforming program compiled with a
different compiler or on a different architecture to different
behaviour is ok with me, compiling it on the same architecture with a
new version of the same compiler into different behaviour is
miscompilation.
In Fortran, for example, signed
integer overflow is just an error, so a compiler that would not
optimize on the assumption of absence of integer overflow would
be doing a poor jobs.
What does "just an error" mean? Does it trap?
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus: >>><https://blog.regehr.org/archives/1287>.
People who want that kind of language know where to find it,
I think these efforts are the result of C compiler maintainers and
their fanboys
In microcontroller work, -Os is often used to put a stronger emphasis on size optimisations. I have noticed there are sometimes "blips" - cases where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
faster code than "-O2". And there can sometimes be cases where the speed/size tradeoffs are unreasonable - significantly slower code for
very minor size improvements, or vice versa.
I think one thing that is sometimes forgotten in all this, and could usefully be mentioned on the gcc manual page for optimisation, is that
cpu architecture flags can make a significant difference. The step from "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
from "-O2" to "-O3" in performance, with much lower risks of slowdowns.
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the sole
semantics being "If a macro invocation unreachable() is reached during
execution, the behavior is undefined" and "The program execution shall
not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?-a Yes, it would make the
compiler implementers do some work to document whatever they decided to
do, but are there any other reasons?
David Brown <david.brown@hesbynett.no> schrieb:
In microcontroller work, -Os is often used to put a stronger emphasis on
size optimisations. I have noticed there are sometimes "blips" - cases
where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
faster code than "-O2". And there can sometimes be cases where the
speed/size tradeoffs are unreasonable - significantly slower code for
very minor size improvements, or vice versa.
Just to muddle the waters a little more: I once submitted a PR
where -O3 led to smaller code than -Os. The reason? -O3 ran a
pass which led to significant simplification and subsequent dead
code elimination.
I think one thing that is sometimes forgotten in all this, and could
usefully be mentioned on the gcc manual page for optimisation, is that
cpu architecture flags can make a significant difference. The step from
"-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
from "-O2" to "-O3" in performance, with much lower risks of slowdowns.
Additionally, -mtune=native can also help.
Ah, comp.arch - that explains it perfectly. USENET groups like comp.arch and comp.lang.care notorious for regulars posting bizarre, intentionally "baffling" code fragments to
spark academic debates, analyze compiler optimization side-effects, or illustrate
architectural edge cases.
AFAIK, it has been argued that
PRINT *,42
END
is a conforming C program because the gfortran command is a
conforming C implementation (the driver also compiles C).
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the sole
semantics being "If a macro invocation unreachable() is reached during
execution, the behavior is undefined" and "The program execution shall
not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as >>opposed to say "implementation defined"?
The idea here is obviously to use this macro only when its invocation
really is unreachable, and in that case it has no effect on the
behaviour.
How should an implementation document what happens when the invocation
of this macro actually is reachable? The effect depends on the
surrounding code and the transformations in the compiler.
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the sole >>>> semantics being "If a macro invocation unreachable() is reached during >>>> execution, the behavior is undefined" and "The program execution shall >>>> not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?
The idea here is obviously to use this macro only when its invocation
really is unreachable, and in that case it has no effect on the
behaviour.
How should an implementation document what happens when the invocation
of this macro actually is reachable? The effect depends on the
surrounding code and the transformations in the compiler.
IME, the annotation is there to squash compiler warnings.
/**
* Restore the state of the current core to the most recent
* rollback state saved. This will result in the rollback handler
* provided to the ::push function to be invoked.
*/
inline void
c_rollback::restore(uint64 rval)
{
_longjmp(&rb_state, rval);
/* NOTREACHED */
}
On 23/09/2026 00:21, Stephen Fuld wrote:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is reached
during execution, the behavior is undefined" and "The program
execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?-a Yes, it would make the
compiler implementers do some work to document whatever they decided
to do, but are there any other reasons?
Good question.
There is a sort of partial ordering (or lattice order) of behaviour specifications.-a "undefined behaviour" is at the bottom - it implements nothing, and anything else can implement it.-a "implementation-defined"
is stronger.-a If the standards say "UB", then an implementation can
define behaviour more concretely.-a If the standards say "IB", then the implementation cannot turn it into "UB".-a So "UB" in the standards
clearly gives the most flexibility to the implementation.
The big disadvantage of making something like this "IB" is not that the compiler implementers need to document their handling, but that they
need to have a specification for it and stick to that - users must be
able to rely on IB being consistent.
Different types of handling for
hitting "unreachable()" can include compiler optimisation on the
assumption that it is never reached, generating a target "trap"
instruction, generating a "breakpoint" when debugging, printing out an
error message and terminating, calling a user-defined function that
sends an angry email to the developer, and tracking all possible code
flows at build time and giving a compiler warning if it can, in fact, be reached.
-a Some users will prefer one method, others will prefer a
different one, and many will change depending on what they are doing.
Crucially, compilers may change what they support over time.-a A compiler cannot document "IB" as "depending on flags, this might be treated as
UB, or give a trap, or perhaps do something else in the future".-a But
it /can/ do exactly that if the standards say it is "UB".
It is also important to note that gcc, clang, icc, and many other
compilers have had something like this for decades - like __builtin_unreachable() - which are implemented exactly as "undefined behaviour" in these tools.
This is, as far as I can tell, the proposal paper that led to "unreachable()" being added to the C standards:
<https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2826.pdf>
"""
We propose the feature unreachable to specify branches in the control
N4eow of a program that will never be reached. The aim is to provide means for the user to express guarantees about the eN4Cective control N4eow that will be executed by a program. Compilers may then apply aggressive optimizations that otherwise would not be possibly or that would rely on
the detection of undeN4Uned behavior for certain input combinations.
"""
The C++ version is :
<https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p0627r6.pdf>
The discussion about the definition is :
"""
What is the best way to define the attribute's effect? The author feels
that the best way is to make the behavior of std::unreachable() be undefined. There are several reasons:
* std::unreachable() causing undefined behavior means that the Standard would not prescribe any particular action, leaving open many possible implementation actions.
* Some compilers already associate being unreachable to undefined
behavior. ClangrCOs documentation states that __builtin_unreachable() rCLhas completely undefined behaviorrCY.
* Optimizing under the assumption that a statement is unreachable, and
thus having unpredictable behavior if the statement is in fact
reachable, falls naturally under "undefined behavior".
* An alternative would be to issue a trap if std::unreachable() is
executed. This could be used in "debug builds", for example. Such a trap falls under "undefined behavior".
* Being undefined behavior implies what happens if a constexpr function calls std::unreachable(): it's not a constant-expression, by (N4713) [expr.const]/2.6
In the authorrCOs opinion, it is not undefined behavior itself that is the problem, but rather unexpected undefined behavior.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the sole
semantics being "If a macro invocation unreachable() is reached during
execution, the behavior is undefined" and "The program execution shall
not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?
The idea here is obviously to use this macro only when its invocation
really is unreachable, and in that case it has no effect on the
behaviour.
How should an implementation document what happens when the invocation
of this macro actually is reachable? The effect depends on the
surrounding code and the transformations in the compiler.
And what would a programmer do with this documentation?
On 9/23/2026 1:38 AM, David Brown wrote:
On 23/09/2026 00:21, Stephen Fuld wrote:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is reached
during execution, the behavior is undefined" and "The program
execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but can
you give a possible rationale for making its behavior "undefined" as
opposed to say "implementation defined"?-a Yes, it would make the
compiler implementers do some work to document whatever they decided
to do, but are there any other reasons?
Good question.
There is a sort of partial ordering (or lattice order) of behaviour
specifications.-a "undefined behaviour" is at the bottom - it
implements nothing, and anything else can implement it.
"implementation-defined" is stronger.-a If the standards say "UB", then
an implementation can define behaviour more concretely.-a If the
standards say "IB", then the implementation cannot turn it into "UB".
So "UB" in the standards clearly gives the most flexibility to the
implementation.
In one sense, yes.-a But in another sense, saying IB doesn't in any way constrain what the code for that implementation does - only that the implementer documents what the code that is written does.
The big disadvantage of making something like this "IB" is not that
the compiler implementers need to document their handling, but that
they need to have a specification for it and stick to that - users
must be able to rely on IB being consistent.
Yes, but it allows that behavior to be anything the implementer chooses, including doing different things for different situations.-a The
implementer has written some code to handle it.-a AFAICT the only
difference is that IB requires the implementer document that code.
Different types of handling for hitting "unreachable()" can include
compiler optimisation on the assumption that it is never reached,
generating a target "trap" instruction, generating a "breakpoint" when
debugging, printing out an error message and terminating, calling a
user-defined function that sends an angry email to the developer, and
tracking all possible code flows at build time and giving a compiler
warning if it can, in fact, be reached.
Absolutely agreed, and perhaps other choices as well. :-)-a And perhaps
even different choices depending upon the specifics of the code around
it.-a ID doesn't preclude any choices here.
-a Some users will prefer one method, others will prefer a different
one, and many will change depending on what they are doing.
Sure.-a And making the implementer document their implementation would
give the user a basis to choose whether to use that capability in a particular situation, or indeed perhaps (though unlikely) choose which compiler implementation to use.
Crucially, compilers may change what they support over time.-a A
compiler cannot document "IB" as "depending on flags, this might be
treated as UB, or give a trap, or perhaps do something else in the
future".-a But it /can/ do exactly that if the standards say it is "UB".
I don't see why IB can't be dependent on flags.-a In fact that might be useful.-a The compiler already does something.-a Just document what it does.-a And as for future choices, nothing in IB prevents future changes,
as long as they are documented.
It is also important to note that gcc, clang, icc, and many other
compilers have had something like this for decades - like
__builtin_unreachable() - which are implemented exactly as "undefined
behaviour" in these tools.
OK.-a But these implementations did *something* with these features.-a I
am not asking them to change that in any way.-a Just document what that something is.
This is, as far as I can tell, the proposal paper that led to
"unreachable()" being added to the C standards:
<https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2826.pdf>
"""
We propose the feature unreachable to specify branches in the control
N4eow of a program that will never be reached. The aim is to provide
means for the user to express guarantees about the eN4Cective control
N4eow that
will be executed by a program. Compilers may then apply aggressive
optimizations that otherwise would not be possibly or that would rely
on the detection of undeN4Uned behavior for certain input combinations.
Sounds reasonable to me.
"""
The C++ version is :
<https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p0627r6.pdf>
The discussion about the definition is :
"""
What is the best way to define the attribute's effect? The author
feels that the best way is to make the behavior of std::unreachable()
be undefined. There are several reasons:
* std::unreachable() causing undefined behavior means that the
Standard would not prescribe any particular action, leaving open many
possible implementation actions.
So does IB.
* Some compilers already associate being unreachable to undefined
behavior. ClangrCOs documentation states that __builtin_unreachable()
rCLhas completely undefined behaviorrCY.
* Optimizing under the assumption that a statement is unreachable, and
thus having unpredictable behavior if the statement is in fact
reachable, falls naturally under "undefined behavior".
* An alternative would be to issue a trap if std::unreachable() is
executed. This could be used in "debug builds", for example. Such a
trap falls under "undefined behavior".
* Being undefined behavior implies what happens if a constexpr
function calls std::unreachable(): it's not a constant-expression, by
(N4713) [expr.const]/2.6
These seem to me to be sort of like "because we've always done it that way".-a While I agree that is true, it doesn't make it right.
In the authorrCOs opinion, it is not undefined behavior itself that is
the problem, but rather unexpected undefined behavior.
This seems curious to me.-a If a behavior is expected, it isn't undefined.
On 23/09/2026 16:54, Stephen Fuld wrote:
On 9/23/2026 1:38 AM, David Brown wrote:
On 23/09/2026 00:21, Stephen Fuld wrote:
On 9/22/2026 5:32 AM, David Brown wrote:
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"?-a Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
Good question.
There is a sort of partial ordering (or lattice order) of behaviour
specifications.-a "undefined behaviour" is at the bottom - it
implements nothing, and anything else can implement it.
"implementation-defined" is stronger.-a If the standards say "UB",
then an implementation can define behaviour more concretely.-a If the
standards say "IB", then the implementation cannot turn it into "UB".
So "UB" in the standards clearly gives the most flexibility to the
implementation.
In one sense, yes.-a But in another sense, saying IB doesn't in any way
constrain what the code for that implementation does - only that the
implementer documents what the code that is written does.
No, you are wrong here.-a And this misunderstanding is key to why it
would be very limited to say "unreachable()" should be implementation- defined.
From the C standards under "Terms, definitions, and symbols" :
"""
* Implementation-defined behaviour
Unspecified behaviour where each implementation documents how the choice
is made.
* Undefined behaviour
Behaviour, upon use of a nonportable or erroneous program construct of erroneous data, for which this document imposes no requirements.
* Unspecified behaviour
Behaviour, that results from the use of an unspecified value, or other behaviour upon which this document provides two or more possibilities
and imposes no further requirements on which is chosen in any instance.
"""
IB means the standard gives certain options, and the implementation documents which choices are made.-a That is completely different from
giving the implementation free reign.-a For example, right-shift of
negative integers gives an implementation-defined value - the
implementation can say it always gives the value 42, if it documents it,
but it is not allowed to say it causes a call to abort() or prints out a run-time error message.
The big disadvantage of making something like this "IB" is not that
the compiler implementers need to document their handling, but that
they need to have a specification for it and stick to that - users
must be able to rely on IB being consistent.
Yes, but it allows that behavior to be anything the implementer
chooses, including doing different things for different situations.
The implementer has written some code to handle it.-a AFAICT the only
difference is that IB requires the implementer document that code.
No.-a See above.
Different types of handling for hitting "unreachable()" can include
compiler optimisation on the assumption that it is never reached,
generating a target "trap" instruction, generating a "breakpoint"
when debugging, printing out an error message and terminating,
calling a user-defined function that sends an angry email to the
developer, and tracking all possible code flows at build time and
giving a compiler warning if it can, in fact, be reached.
Absolutely agreed, and perhaps other choices as well. :-)-a And perhaps
even different choices depending upon the specifics of the code around
it.-a ID doesn't preclude any choices here.
See above.
-a Some users will prefer one method, others will prefer a different
one, and many will change depending on what they are doing.
Sure.-a And making the implementer document their implementation would
give the user a basis to choose whether to use that capability in a
particular situation, or indeed perhaps (though unlikely) choose which
compiler implementation to use.
Crucially, compilers may change what they support over time.-a A
compiler cannot document "IB" as "depending on flags, this might be
treated as UB, or give a trap, or perhaps do something else in the
future".-a But it /can/ do exactly that if the standards say it is "UB".
I don't see why IB can't be dependent on flags.-a In fact that might be
useful.-a The compiler already does something.-a Just document what it
does.-a And as for future choices, nothing in IB prevents future
changes, as long as they are documented.
See above.
I think we mostly agree on what we would like compiler implementations
to do, at least in some cases (I am not sure if you like the "anything
can happen" case of the compiler assuming unreachable() is never hit).
But what you want here is only achievable if the standard says it is UB,
not if the standard says it is IB.
On 9/23/2026 9:24 AM, David Brown wrote:
I think we mostly agree on what we would like compiler implementations
to do, at least in some cases (I am not sure if you like the "anything
can happen" case of the compiler assuming unreachable() is never hit).
But what you want here is only achievable if the standard says it is
UB, not if the standard says it is IB.
It is obvious that you are right; the the existing choices as documented
in the standard and your post above (thanks for that) don't permit what
I want.-a I apologize for not recognizing that up front.
But I maintain
that there should be a category that essentially says "You can do what
you want, but you must document what you did."-a Furthermore, if such a choice existed, most instances of undefined should be "reclassified"
into that new choice.-a I think that would aid programmers understanding.
How much more optimizations can compilers deliver to the bottom line
(not just benchmarks). Last week we saw an episode where the compilers
only got 1.x% speedup over years (sub-decade). Is there ever going to
be a time to stop ??
The "something" that they say is "If control flow reaches the point of
the __builtin_unreachable, the program is undefined." That's the documentation in the gcc manual, along with some examples and use-cases.
David Brown <david.brown@hesbynett.no> schrieb:
The "something" that they say is "If control flow reaches the point of
the __builtin_unreachable, the program is undefined." That's the
documentation in the gcc manual, along with some examples and use-cases.
This has important (well, to compiler developers) use cases, for
generating test cases. Grabbing a random example off bugzilla,
from https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126048 :
int src(int v1_u8) {
if (!((1 <= v1_u8) && (v1_u8 <= 16))) __builtin_unreachable();
int i0_u8 = v1_u8 << v1_u8;
int i1_u8 = i0_u8 / v1_u8;
return i1_u8;
}
The specific information used here is "This cannot happen, so
optimize based on the knowledge that this is the case". It would
be possible to infer the same information from other method, but
this is quite convenient.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard. But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
Thomas Koenig <tkoenig@netcologne.de> writes:
A straightforward patch would
very likely pessimize a lot of existing code which profits from >>>auto-vectorization.
How can we test this claim?
I do not believe that I have to explain the scientific method
to you.
As a first step, you would find an option that enables/disables
what you don't like. -fstore-merging looks like a
candidate, but there may be others. You can look at >https://dl.acm.org/doi/10.1109/ASE56229.2023.00209 or >https://link.springer.com/article/10.1007/s10515-024-00437-w
if you want the full package.
Then try this combination of options on other benchmarks
as well as your pet one. I have my reservations about SPEC,
but as you work at a university, you can get it for a discount.
You can also use freely available benchmarks: Coremark, embench,
Fortran Polyhedron - there are a lot.
If there are benchmark which regress significantly with that
set of options, you have your test case where it hurts.
Like I wrote previously - you could have a bachelor or master
student do this, it could be (part of a) nice thesis.
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are defined in the standard. But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard. But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
That would allow the compiler to fork "/bin/games/chess -play_both_sides";
It is unlikely that the programmer would desire or expect that.
I think you want closer to "Do something reasonable here" for a broad definition of reasonable.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard. But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
That would allow the compiler to fork "/bin/games/chess -play_both_sides";
It is unlikely that the programmer would desire or expect that.
On 21/09/2026 19:17, Anton Ertl wrote:
I referred to a specific paradoxical effect, not completely avoiding
all of them: The effect of programmers who, instead of working on
making the code faster through source-level changes, have to invest
time into avoiding getting it miscompiled.
By "miscompiled", do you mean incorrect object code, or object code that
did not have the performance the programmer expected or hoped for?
But as
for the problems you mention w.r.t UB, I think it's just the result of
poor semantics, for which I guess we (language semanticists) are partly
to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
I have never spoken to the C standards committee or writers, either of >current standard versions or the original C89 standard, or writers of >pre-standard C specifications. So I cannot in any way claim to know
their thoughts or motivations. I also have not seen any documentation
that suggests what they might have thought about "optimisation on the >assumption that undefined behaviour does not occur" - in either
direction.
(It has been mentioned that, for example, a two's complement
implementation could use wrapping signed arithmetic - but I have never
seen a suggestion that this behaviour should be encouraged or expected
just because a machine uses two's complement.)
You do not ask "what happens
if my signed integer arithmetic overflows?" or "what happens when I
access an array out of bounds?" - rather, it is your responsibility as a
C programmer to make sure that never happens.
The prime reason C does
not define behaviour here is not that different hardware or compilers
handle things differently, but that there is no sensible definition that >could be made.
Remember, C has a perfectly good way to say "this is determined by the >hardware or the implementation" - it is "implementation-defined
behaviour". It has a perfectly good way to say that "this operation
could result in any value" - it is "unspecified behaviour" or
"unspecified value".
When the C standards writers say something is "undefined behaviour",
rather than "implementation defined" or "unspecified", it is my belief
that they did so intentionally and knowingly.
And note that in C23 the macro "unreachable()" was added with the sole >semantics being "If a macro invocation unreachable() is reached during >execution, the behavior is undefined" and "The program execution shall
not reach such an invocation". The justification is for better
diagnostics and optimisations.
I don't know whether or not the "founding fathers" of C intended or
expected compilers to optimise on the assumption that UB did not occur.
But I am confident that they considered a program to be broken if
execution reached a point where the behaviour was not defined in an >implementation (something may be UB in the C standards yet defined by an >implementation).
It is, of course, possible that they did not foresee quite how this
would pan out in modern compilers. In particular, they may not have >predicted "time-travel" optimisations. However, authors of more modern
C standards - say, C11 onwards - know about them and have could have >explicitly outlawed them if they were considered to be invalid.
For
heavily used compilers, there is a strong (but not overpowering) push >towards backwards compatibility. This is why they have heavy regression >tests, and pre-release versions are tested with large samples of
important existing code. If this gives unexpected results, these must
be dealt with - was it a bug in the new compiler (in which case the fix
is obvious), or was it a bug in the old source code? Those cases are
more complicated - sometimes the old code must be fixed, but sometimes
the incorrect source code is too common, idomatic or important and the >compiler must, in effect, support additional semantics to retain the old >accidental semantics.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus:
<https://blog.regehr.org/archives/1287>.
There are two key problems with UB that stand in the way of this kind of >initiative. One is that many compilers provide consistent (and
sometimes even documented) behaviour for some things that are UB in the
C standards. Code relying on these behaviours is valid and safe, but >non-portable. But different compilers (or the same compiler but
different targets) could easily have different semantics, making it very >difficult to agree on any one choice. Secondly, many programmers write
code on the assumption that certain UB has certain consistent and
reliable behaviour - even though it is not documented anywhere. It is
code like that which could benefit from a "friendly C" (a poor choice of >name, IMHO, but that's entirely subjective) variant. But it is also
such code that makes a "friendly C" variant hard to define - such code
is hard to identify, and the expected behaviour can be even harder to
find, specify, and consistently describe.
But fortunately a completely specified C is not necessary, a
willingness to preserve the behaviour of existing working programs
compiled with an earlier version of the same compiler is. Read more
about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.
I think it is entirely reasonable to keep old compiler versions around
and use them for old code.
I think it is entirely unreasonable to try to specify that new compilers >should have defined specifications to implement the "semantics" that
older compilers used for particular types of undefined behaviour. These >semantics will, for the most part, be poorly defined and can often be >inconsistent - many programs with UB work by luck, not design.
Linux commits to preserving user-space behaviour (whether standard or
not), everything I have heard from people claiming to speak for gcc
and clang maintainers has been in the opposite direction. So I blame
the gcc and clang maintainers for the undefined-behaviour shenanigans
they perform.
Linux commits to preserving the /defined/ behaviour of user-space APIs.
But if a
particular invalid value of the parameter lets you gain access to
root-owned files, due to a bug in the kernel, you can be very sure that
this user-space behaviour will /not/ be preserved.
On 9/24/2026 12:52 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"?-a Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting.-a A construct having implementation-defined behavior
cannot do "just anything";-a the C standard must specify what
behavior choices are possible.-a Any fixed set of choices might
rule out what an implementation would like to do.-a So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard.-a But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
That would allow the compiler to fork "/bin/games/chess -
play_both_sides";
It is unlikely that the programmer would desire or expect that.
I think you want closer to "Do something reasonable here" for a broad
definition of reasonable.
Ahhh, but that is the problem. Different people have different
definitions of reasonable.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"? Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting. A construct having implementation-defined behavior
cannot do "just anything"; the C standard must specify what
behavior choices are possible. Any fixed set of choices might
rule out what an implementation would like to do. So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard. But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
Let's call the new behavior type "implementation dependent".
Suppose an implementation gives documentation that says "If a program execution would encounter implementation-dependent behavior, then
program execution may be affected in ways that are unpredictable,
unknown, and/or unreliable."
Question: are you okay with that?
If not, how would you write a
requirement in the C standard to limit the meaning of "you can do
whatever you want, but you must document what you do" that gives only
as much freedom as you think should be allowed?
I think we will find that different people have very different ideas
about how much freedom should be allowed in saying what happens. It's
a very hard problem.
David Brown <david.brown@hesbynett.no> writes:
I have never spoken to the C standards committee or writers, either of >>current standard versions or the original C89 standard, or writers of >>pre-standard C specifications. So I cannot in any way claim to know
their thoughts or motivations. I also have not seen any documentation >>that suggests what they might have thought about "optimisation on the >>assumption that undefined behaviour does not occur" - in either
direction.
When C89 was developed, the dominating C compiler in the Unix world
was PCC. GCC came on the scene late in the game (1988), without such >assumptions and produced a significant performance advantage.
David Brown <david.brown@hesbynett.no> writes:
On 21/09/2026 19:17, Anton Ertl wrote:
I referred to a specific paradoxical effect, not completely avoiding
all of them: The effect of programmers who, instead of working on
making the code faster through source-level changes, have to invest
time into avoiding getting it miscompiled.
By "miscompiled", do you mean incorrect object code, or object code that
did not have the performance the programmer expected or hoped for?
Code that behaves differently than intended, e.g., where a bounds check
is "optimized" away.
Performance regressions can be nasty for performance-sensitive code,
and consume time for finding a workaround (if one can be found), but
is a different thing.
But as
for the problems you mention w.r.t UB, I think it's just the result of >>>> poor semantics, for which I guess we (language semanticists) are partly >>>> to blame: we have developed fairly good tools to design sane language
specs in general,
I violently disagree. The C89 standard (just to name one) is a
partial specification not because the original C standards people were
poor at specifying semantics, but because given the differences
between the compilers out there and the hardware out there, the
easiest way to reach a consensus is to leave some parts unspecified.
I have never spoken to the C standards committee or writers, either of
current standard versions or the original C89 standard, or writers of
pre-standard C specifications. So I cannot in any way claim to know
their thoughts or motivations. I also have not seen any documentation
that suggests what they might have thought about "optimisation on the
assumption that undefined behaviour does not occur" - in either
direction.
When C89 was developed, the dominating C compiler in the Unix world
was PCC. GCC came on the scene late in the game (1988), without such assumptions and produced a significant performance advantage. Only
after several years of heavy gcc development assumptions like -fstrict-aliasing were introduced around gcc-2.7 (which was initially disabled by default);
I don't know when the assumption that signed
integer operations do not overflow was introduced, and initially it
could not be disabled (-fwrapv was added around gcc-3.4), so obviously
by that time people had taken over the gcc project who had a different
view about how undefined behaviour should be treated.
Maybe the change can be pinpointed to the version when
-fstrict-aliasing became the default, or maybe the attitude change was
more gradual.
Anyway, given the state of C compilers during C89 development, no, I
don't think they ever thought that undefined behaviour would be used
to justify "optimizing" away bounds checks, such as the second one
below:
char *buf = ...;
char *buf_end = ...;
unsigned int len = ...;
if (buf + len >= buf_end)
return; /* len too large */
if (buf + len < buf)
return; /* overflow, buf+len wrapped around */
/* write to buf[0..len-1] */
[This example is given in <https://people.csail.mit.edu/nickolai/papers/wang-stack.pdf>, which
cites <https://www.kb.cert.org/vuls/id/162289>]
(It has been mentioned that, for example, a two's complement
implementation could use wrapping signed arithmetic - but I have never
seen a suggestion that this behaviour should be encouraged or expected
just because a machine uses two's complement.)
The C99 rationale (which obviously contains text written for C89)
includes the following:
|C code can be non-portable. Although it strove to give
|programmers the opportunity to write truly portable programs, the C89 |Committee did not want to force programmers into writing portably, to |preclude the use of C as a "high-level assembler": the ability to
|write machine-specific code is one of the strengths of C. It is this |principle which largely motivates drawing the distinction between
|strictly conforming program and conforming program (Section 4).
|
|Keep the spirit of C. The C89 Committee kept as a major goal
|to preserve the traditional spirit of C. There are many facets of the |spirit of C, but the essence is a community sentiment of the
|underlying principles upon which the C language is based. Some of the |facets of the spirit of C can be summarized in phrases like:
|
| * Trust the programmer.
| * Don't prevent the programmer from doing what needs to be done.
| * Keep the language small and simple.
| * Provide only one way to do an operation.
| * Make it fast, even if it is not guaranteed to be portable.
|
|The last proverb needs a little explanation. The potential for
|efficient code generation is one of the most important strengths of
|C. To help ensure that no code explosion occurs for what appears to be
|a very simple operation, many operations are defined to be how the
|target machine's hardware does it rather than by a general abstract rule.
So no, it does not say what a compiler should do, but I understand it
as saying that it is in the spirit of C if a programmer relies on 2s-complement wraparound arithmetic for signed integers on an
architecture where C compilers implement signed integer arithmetic in
that way.
That was certainly the case for many architectures in 1989, where no C compilers on MIPS produced the add or addi instructions (which trap on
signed overflow) for signed arithmetic even though it was available.
It was also the case for 68000, IA-32, HPPA, SPARC, Alpha, and others.
And whenever I looked at the code produced for MIPS and Alpha, I have
never seen the trapping addition/subtraction/multiplication
instructions generated. Maybe if you ask for it with -ftrapv, but
IIRC I looked at generated code once and found that the trapping
instructions were ignored even then, instead generating more
long-winded overflow checks.
You do not ask "what happens
if my signed integer arithmetic overflows?" or "what happens when I
access an array out of bounds?" - rather, it is your responsibility as a
C programmer to make sure that never happens.
That's not at all what I read in the rationale.
The prime reason C does
not define behaviour here is not that different hardware or compilers
handle things differently, but that there is no sensible definition that
could be made.
Java had no problem producing a sensible definition for what happens
on signed overflow.
Likewise, for the bounds check above, there is a sensible definition
that can be made, and gcc actually still (or again) uses it if you
tell it to, with -fwrapv-pointer.
Remember, C has a perfectly good way to say "this is determined by the
hardware or the implementation" - it is "implementation-defined
behaviour". It has a perfectly good way to say that "this operation
could result in any value" - it is "unspecified behaviour" or
"unspecified value".
When the C standards writers say something is "undefined behaviour",
rather than "implementation defined" or "unspecified", it is my belief
that they did so intentionally and knowingly.
The question is what the intention was. My impression is a decisive
reason for declaring something undefined has been if straightforwardly generated code could trap on some architecture. This becomes most
apparent when it comes to shifts, where some cases are
implementation-defined and some are undefined, and the undefined cases correspond to cases that trap on some architecture.
I do not care that much about the language-lawyering about the
different names for the incomplete definitions in the C standards, but
if an instruction traps, that does not look like nasal demons to me
(Ariane 501 customers may disagree, but one can blame the
inappropriate handling of the trap, based on the proof that the trap
cannot happen).
Maybe one of the C language lawyers can explain why they never chose
to use "imlementation-defined" or any of the other incompleteness
variants when a trap is possible.
And note that in C23 the macro "unreachable()" was added with the sole
semantics being "If a macro invocation unreachable() is reached during
execution, the behavior is undefined" and "The program execution shall
not reach such an invocation". The justification is for better
diagnostics and optimisations.
How can "undefined behaviour" lead to better diagnostics given the
attitude that "it is your responsibility as a C programmer to make
sure that never happens"? Whoever wrote this justfication apparently
has a different idea of the meaning of undefined behaviour than you
do.
Anyway, undefined behaviour in connection with "unreachable()" or
"restrict" is not a problem as far as I am concerned. A programmer
can just choose to never use "unreachable()" or "restrict", and avoid
any undefined behaviours coming from these language features. And
when he introduces them, my recommendation is to do it sparingly, only
in those places that are relevant to performance.
This is in contrast to taking the assumption that undefined behaviour
does not happen anywhere, which means that programmers have to check everywhere for the >200 explicitly stated (and who knows how many
implied) undefined behaviours in the current C standard.
I don't know whether or not the "founding fathers" of C intended or
expected compilers to optimise on the assumption that UB did not occur.
One can look at the code written by Ritchie, Thompson and the other
people at Bell Labs. Does it contain undefined behaviour? Very
likely.
But I am confident that they considered a program to be broken if
execution reached a point where the behaviour was not defined in an
implementation (something may be UB in the C standards yet defined by an
implementation).
If you mean "documented as defined", I doubt it. I expect that they
did write programs that rely on the actual code generated by the
compiler(s) the used for the program, on the target(s) they used for
the program. And of course the C compilers before C89 did not
document behaviours as defined that would be undefined by the standard
only later.
It is, of course, possible that they did not foresee quite how this
would pan out in modern compilers. In particular, they may not have
predicted "time-travel" optimisations. However, authors of more modern
C standards - say, C11 onwards - know about them and have could have
explicitly outlawed them if they were considered to be invalid.
There is no need to outlaw time travel in C standards, because it was
never allowed. At least that's what I read here some years ago. Time traveling is a C++, not C property.
For
heavily used compilers, there is a strong (but not overpowering) push
towards backwards compatibility. This is why they have heavy regression
tests, and pre-release versions are tested with large samples of
important existing code. If this gives unexpected results, these must
be dealt with - was it a bug in the new compiler (in which case the fix
is obvious), or was it a bug in the old source code? Those cases are
more complicated - sometimes the old code must be fixed, but sometimes
the incorrect source code is too common, idomatic or important and the
compiler must, in effect, support additional semantics to retain the old
accidental semantics.
Yes, my impression is that the pushback after miscompiling "relevant
code" has been strong enough that the regression tests now avoid
miscompiling that code, and a lot of the "irrelevant" code such as
Gforth rides in the slipstream of that.
Of course, relevant code like
the Linux kernel uses a collection of flags (e.g,
-fno-strict-overflow) that define otherwise undefined behaviour, so
anybody who wants to ride in the slipstream of that code should use
these flags as well, and leave the benchmarking settings (i.e.,
without these flags) to the benchmarks.
Still, even if they did not, there are still differences between
hardware and between existing compilers to reconcile (although a lot
of the old hardware variations have died out), and getting consensus
on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
Flatt, and John Regehr tried to create a more completely specified
"friendly C", and did not find consensus:
<https://blog.regehr.org/archives/1287>.
There are two key problems with UB that stand in the way of this kind of
initiative. One is that many compilers provide consistent (and
sometimes even documented) behaviour for some things that are UB in the
C standards. Code relying on these behaviours is valid and safe, but
non-portable. But different compilers (or the same compiler but
different targets) could easily have different semantics, making it very
difficult to agree on any one choice. Secondly, many programmers write
code on the assumption that certain UB has certain consistent and
reliable behaviour - even though it is not documented anywhere. It is
code like that which could benefit from a "friendly C" (a poor choice of
name, IMHO, but that's entirely subjective) variant. But it is also
such code that makes a "friendly C" variant hard to define - such code
is hard to identify, and the expected behaviour can be even harder to
find, specify, and consistently describe.
Such efforts try to achieve more than what is discussed in the part of
the C99 rationale cited above, and more than I argue for in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>; it also
tries to solve portability between architectures and compilers. And
given that such efforts have not come to fruition, one can conclude
that they try to achieve too much.
Alternatively, one can see the various flags provided by gcc, such as -fno-strict-aliasing as achieving at least a part of such efforts (not
the portability between compilers, in general, but at least between architectures). There is at least one case (-fwrapv vs. -ftrapv)
where one can ask gcc for one of several definined behaviours, though.
But fortunately a completely specified C is not necessary, a
willingness to preserve the behaviour of existing working programs
compiled with an earlier version of the same compiler is. Read more
about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.
I think it is entirely reasonable to keep old compiler versions around
and use them for old code.
Unfortunately, the various Linux distributions do not agree with you,
and do not distribute gcc versions back to 1.0. So, for software
distributed as source code, asking for a known-good compiler for
IA-32, such as gcc-2.95, is impractical.
I think it is entirely unreasonable to try to specify that new compilers
should have defined specifications to implement the "semantics" that
older compilers used for particular types of undefined behaviour. These
semantics will, for the most part, be poorly defined and can often be
inconsistent - many programs with UB work by luck, not design.
On the contrary, the most common cases of undefined behaviour become well-defined. E.g., before gcc ever generated code based on the
assumption that signed overflow never happens, it generated addu or equivalent (e.g., addiu or the implied addition in the address
computation of a load) for an addition in the C source code; and optimizations preserved this behaviour, e.g., a+(-b) was compiled to
subu, not sub. Continuing to generate code that behaves like that is well-defined, and it now even has flags in gcc (-fwrapv and
-fwrapv-pointer, which can be combined into -fno-strict-overflow).
Linux commits to preserving user-space behaviour (whether standard or
not), everything I have heard from people claiming to speak for gcc
and clang maintainers has been in the opposite direction. So I blame
the gcc and clang maintainers for the undefined-behaviour shenanigans
they perform.
Linux commits to preserving the /defined/ behaviour of user-space APIs.
If you mean that it takes the same work-to-rule approach that the gcc
people do when it comes to "irrelevant" code, that's not the case. It commits to preserving the actual behaviour. That is well publicised,
e.g. <https://linuxreviews.org/WE_DO_NOT_BREAK_USERSPACE>. Linus
Torvalds wrote:
|If a change results in user programs breaking, it's a bug in the
|kernel. We never EVER blame the user programs.
But if a
particular invalid value of the parameter lets you gain access to
root-owned files, due to a bug in the kernel, you can be very sure that
this user-space behaviour will /not/ be preserved.
I don't know if the maintainers of a user program ever reported a bug
for closing such a security hole and insisted on preserving the
behaviour; I doubt it. If they would, one way to deal with that would
be to have a kernel option (compile-time, startup-time, or run-time)
for opening the hole, which the default being that the hole is closed.
There has been at least one well-known case where the kernel people
went to great lengths to preserve a behaviour for a certain userspace
program where the other behaviour already also had users (and in this
case the kernel-option variant was not satisfactory, so they used
something far uglier).
This has all been well publicised, which makes me wonder why you are spreading the claims above; either you do not know what you are
writing about, or you knowingly spread an untruth.
anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
David Brown <david.brown@hesbynett.no> writes:
I have never spoken to the C standards committee or writers, either of >>>current standard versions or the original C89 standard, or writers of >>>pre-standard C specifications. So I cannot in any way claim to know >>>their thoughts or motivations. I also have not seen any documentation >>>that suggests what they might have thought about "optimisation on the >>>assumption that undefined behaviour does not occur" - in either >>>direction.
When C89 was developed, the dominating C compiler in the Unix world
was PCC. GCC came on the scene late in the game (1988), without such >>assumptions and produced a significant performance advantage.
Well, yes. PCC at that point was 14 years old and showing its age.
Motorola used PCC for the 88100 processor (I had to fix a bug in the
register allocator when compiling the output of cfront, which
extensively uses the comma operator in 1990); I had first used PCC in
1981.
On 9/24/2026 9:29 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/23/2026 7:47 AM, Tim Rentsch wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
On 9/22/2026 5:32 AM, David Brown wrote:
[...]
And note that in C23 the macro "unreachable()" was added with the
sole semantics being "If a macro invocation unreachable() is
reached during execution, the behavior is undefined" and "The
program execution shall not reach such an invocation".
I don't have a problem with the inclusion of "unreachable()", but
can you give a possible rationale for making its behavior
"undefined" as opposed to say "implementation defined"?-a Yes, it
would make the compiler implementers do some work to document
whatever they decided to do, but are there any other reasons?
The short answer is that being implementation-defined is too
limiting.-a A construct having implementation-defined behavior
cannot do "just anything";-a the C standard must specify what
behavior choices are possible.-a Any fixed set of choices might
rule out what an implementation would like to do.-a So the only
way to allow implementations to do whatever they choose is to
have the behavior be undefined.
As I responded to David, you are right, given the way those terms are
defined in the standard.-a But I submit there should be some way to
express essentially "you can do whatever you want, but must document
what you do".
Let's call the new behavior type "implementation dependent".
I am not hung up on the name, so at least for purposes of discussion, OK.
Suppose an implementation gives documentation that says "If a program
execution would encounter implementation-dependent behavior, then
program execution may be affected in ways that are unpredictable,
unknown, and/or unreliable."
Question: are you okay with that?
If that is the implementation's general response, then, while it might
be "legal", then no, I am not happy with it.
If not, how would you write a
requirement in the C standard to limit the meaning of "you can do
whatever you want, but you must document what you do" that gives only
as much freedom as you think should be allowed?
I have minimal experience with standards writing, and none with language writing, so I may be off base here.
I think that there are many situations where the standard, correctly, doesn't specify anything about what the compiler should do.-a These are currently called undefined behavior, and the standard essentially lets
it go at that - no further documentation required.
But in many, but not all, of such cases, the compiler implementer knows
what it is going to do.-a What I am after is that in such cases, the implementer "supplement" the bare words with information that he has
that knowledge and he should communicate it to the user.
I think we will find that different people have very different ideas
about how much freedom should be allowed in saying what happens.-a It's
a very hard problem.
Note that I am not trying to restrict "what happens" in any way.-a What happens is anything that could/would happen under today's definition of undefined behavior.-a I am just trying to ask the implementer, whenever possible (and I realize that it is not always possible) to elaborate on
the bare words "undefined behavior".
Suppose an implementation gives documentation that says "If a program execution would encounter implementation-dependent behavior, then
program execution may be affected in ways that are unpredictable,
unknown, and/or unreliable."
Question: are you okay with that?
Tim Rentsch [2026-09-24 09:29:13] wrote:
Suppose an implementation gives documentation that says "If a program
execution would encounter implementation-dependent behavior, then
program execution may be affected in ways that are unpredictable,
unknown, and/or unreliable."
Question: are you okay with that?
I'm not. It's not clear what it is that I want, but I know that I want
my undefined behaviors to be more constrained in what they can do.
I'll call "SMUB" my desired version of undefined behavior. It would go somewhere along the lines of:
In case of a SMUB operation whose result type is T, then the
operation is allowed to return any bit pattern compatible with the
type T, or it can stop the execution of the program with an error.
In case it stops the execution with an the error, that it is
acceptable if the error doesn't occur exactly at the time the
operation would have taken place (i.e it can occur at any later
time, but also at an earlier time provided we already know that the
SMUB operation would take place anyway).
A SMUB operation that is normally pure is not allowed to mutate
anything. And a SMUB operation that normally mutates one memory
cell of type T is allowed to mutate any *one* memory cell of same
size as type T by putting into it a bit pattern compatible with the
type T.
In the case of `unreachable()`, that means the only thing it would be
allowed to do is to signal an error (at any point of the execution once
we know that `unreachable()` would be executed) or do nothing. IOW, if
it's executed and does not signal an error, the rest of the code may not
be optimized under the assumption that it was not (nor will not
be) executed.
On 25/09/2026 17:16, Stefan Monnier wrote:
Tim Rentsch [2026-09-24 09:29:13] wrote:
Suppose an implementation gives documentation that says "If a program
execution would encounter implementation-dependent behavior, then
program execution may be affected in ways that are unpredictable,
unknown, and/or unreliable."
Question: are you okay with that?
I'm not. It's not clear what it is that I want, but I know that I want
my undefined behaviors to be more constrained in what they can do.
I'll call "SMUB" my desired version of undefined behavior. It would go
somewhere along the lines of:
In case of a SMUB operation whose result type is T, then the
operation is allowed to return any bit pattern compatible with the
type T, or it can stop the execution of the program with an error.
In case it stops the execution with an the error, that it is
acceptable if the error doesn't occur exactly at the time the
operation would have taken place (i.e it can occur at any later
time, but also at an earlier time provided we already know that the
SMUB operation would take place anyway).
A SMUB operation that is normally pure is not allowed to mutate
anything. And a SMUB operation that normally mutates one memory
cell of type T is allowed to mutate any *one* memory cell of same
size as type T by putting into it a bit pattern compatible with the
type T.
That sounds somewhat like "erroneous behaviour" in C++.
<https://cppreference.com/cpp/language/ub>
It also sounds like what C has for conversion of an out-of-range value
to a signed integer type - "either the result is implementation-defined
or an implementation-defined signal is raised."
That seems fine for certain things, as a possible choice by an >implementation - it keeps the diagnostic option for the behaviour. But
it limits optimisation. I have no issues with you wanting that, but I
don't want /my/ compilations to be limited by it.
And in fits what you have today in some cases, such as signed integer >overflow in gcc - you can choose "-fwrapv", or you can choose >"-fsanitize=signed-integer-overflow", or you can choose optimisation on
the assumption that the code is correct. That's the best of all worlds
as I see it, and forcing choices in the standards would limit that.
In the case of `unreachable()`, that means the only thing it would be
allowed to do is to signal an error (at any point of the execution once
we know that `unreachable()` would be executed) or do nothing. IOW, if
it's executed and does not signal an error, the rest of the code may not
be optimized under the assumption that it was not (nor will not
be) executed.
You would only put "unreachable()" in your code at a spot where you do
not expect execution to reach. So if it has got there, your code is
broken - you have no control of what it is doing.
If I write :
// classify must only be called with x between 0 and 100
int classify(int x) {
if (x < 10) return 0;
if (x < 30) return 1;
if (x <= 100) return 2;
unreachable();
On 2026-Sep-22 01:25, Anton Ertl wrote:
Effect of new optimizations or "optimizations" based on assuming that
undefined behaviour never happens on performance? They never say.
That's the cool thing. They do not have numbers (they certainly never
present any, certainly not for their own compilers), but are convinced
that their "optimizations" do wonders for performance, and their
fanboys are even more convinced.
It seems someone measured UB optimization performance impact in LLVM:
Exploiting Undefined Behavior in C C++ Programs for Optimization
A Study on the Performance Impact, 2025 >https://dl.acm.org/doi/abs/10.1145/3729260 >https://dl.acm.org/doi/pdf/10.1145/3729260
"Using LLVM, a compiler known for its extensive use of UB for optimizations, >we demonstrate that, for the benchmarks and UB categories that we evaluated, >the end-to-end performance gains are minimal."
Thomas Koenig <tkoenig@netcologne.de> writes:
Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
Thomas Koenig <tkoenig@netcologne.de> writes:
A straightforward patch would
very likely pessimize a lot of existing code which profits from >>>>auto-vectorization.
How can we test this claim?
I do not believe that I have to explain the scientific method
to you.
As a first step, you would find an option that enables/disables
what you don't like. -fstore-merging looks like a
candidate, but there may be others. You can look at >>https://dl.acm.org/doi/10.1109/ASE56229.2023.00209 or >>https://link.springer.com/article/10.1007/s10515-024-00437-w
if you want the full package.
Then try this combination of options on other benchmarks
as well as your pet one. I have my reservations about SPEC,
but as you work at a university, you can get it for a discount.
You can also use freely available benchmarks: Coremark, embench,
Fortran Polyhedron - there are a lot.
If there are benchmark which regress significantly with that
set of options, you have your test case where it hurts.
Like I wrote previously - you could have a bachelor or master
student do this, it could be (part of a) nice thesis.
That's a lot of words to express that you have no evidence for your
claim.
That seems fine for certain things, as a possible choice by an
implementation - it keeps the diagnostic option for the behaviour. But it limits optimisation.
what is the compiler supposed to do with the "unreachable()" if you are not halting with an error?
There's no way to get sensible semantics here - you either say terminate immediately to limit damage, or accept what UB means. Trying to say you things can go a bit bad, but not too bad, never works.
We can do that already. I don't want "unreachable()" to duplicate
that, I want it to mean something else.
David Brown [2026-09-25 19:00:26] wrote:
That seems fine for certain things, as a possible choice by an
implementation - it keeps the diagnostic option for the behaviour. But it >> limits optimisation.
Limiting optimization is the whole purpose of the exercise!
Remember, we're starting from the problem that UB is used by
compilers in ways which catch programmers off-guard.
My SMUB proposal
is an attempt to define something which I hope is closer to what
programmers expect.
Given that UB has been shown to provide only fairly minor performance
gains, I'm hopeful that "downgrading" UB to SMUB would be good enough in
the vast majority of cases. You could still add a `--SMUB-is-UB` flag
if you're so inclined, of course, just like the `--fast-math`.
what is the compiler supposed to do with the "unreachable()" if you are not >> halting with an error?
At runtime, I think a self-respecting compiler would halt with an error,
yes. That's the only behavior that won't catch programmers by surprise.
There's no way to get sensible semantics here - you either say terminate
immediately to limit damage, or accept what UB means. Trying to say you
things can go a bit bad, but not too bad, never works.
Note that in your example, even if `unreachable()` turns into a nop, the potential damage still corresponds to executing the code which the users wrote, in the order they wrote it, performing the tests they wrote.
So it still matches better the naive semantics most programmers have in
their head than some of the very counter-intuitive runtime behaviors
you may get with UB's current optimizations.
We can do that already. I don't want "unreachable()" to duplicate
that, I want it to mean something else.
I know. You like to take advantage of UB to try and gain the last few percents of optimizations. I put less emphasis on that.
A halfway point might be for the compiler to emit a warning when it
can't show that `unreachable()` is indeed unreachable.
[ I generally like the idea of programmers being able to request
specific optimizations and to be warned by the compiler when those
requests can't be satisfied. So we don't have to go dig in the
assembly output to confirm whether the compiler did the right thing or
not. ]
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
Given that UB has been shown to provide only fairly minor
performance gains,
This I found hard to believe, so I ran a
few benchmarks myself. I used the well-known
Polyhedron suite, which can now be downloaded from https://fortran.uk/fortran-compiler-comparisons/the-polyhedron-solutions-benchmark-suite/
,
This is Fortran, so integer overflow is an error. I ran the
testsuite on gfortran with and without -fwrapv, otherwise letting
the compiler go wild (-Ofast). Because gfortran puts arrays on the
stack with -Ofast, I had to increase the stacksize for one particular
code. I had to disable one test, rnflow, because of a bug.
Here are the results:
Date & Time : 27 Sep 2026 12:13:00
Test Name : gfortran-Ofast
Compile Command : gfortran -march=native -mtune=native -Ofast %n.f90
-o %n Benchmarks : ac aermod air capacita channel2 doduc
fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein
-rnflow test_fpu2 tfft2 Maximum Times : 10000.0 Target Error %
: 0.200 Minimum Repeats : 10
Maximum Repeats : 100
Benchmark Compile Executable Ave Run Number Estim
Name (secs) (bytes) (secs) Repeats Err %
--------- ------- ---------- ------- ------- ------
ac 0.65 38616 4.54 13 0.1876
aermod 24.66 1074024 3.26 15 0.1773
air 2.96 94920 0.89 19 0.1560
capacita 3.97 127104 5.99 18 0.1858
channel2 0.40 26584 40.47 10 0.1494
doduc 4.61 164768 4.43 13 0.1230
fatigue2 1.50 69200 30.37 11 0.1968
gas_dyn2 1.59 74456 38.17 19 0.1610
induct2 4.63 220280 12.26 10 0.1702
linpk 0.53 30328 1.88 14 0.1842
mdbx 1.87 82200 3.02 10 0.1662
mp_prop_desi 0.75 39568 33.30 19 0.1536
nf 0.75 34848 3.03 18 0.1680
protein 2.32 94640 10.60 15 0.1572
-rnflow 0.00 0 -1.00 15 0.1572
test_fpu2 2.45 76648 16.08 16 0.1926
tfft2 0.97 43000 13.08 16 0.1436
Geometric Mean Execution Time = 10.57 seconds
Date & Time : 27 Sep 2026 13:45:57
Test Name : wrapv
Compile Command : gfortran -fwrapv -march=native -mtune=native -Ofast
%n.f90 -o %n Benchmarks : ac aermod air capacita channel2 doduc
fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein
-rnflow test_fpu2 tfft2 Maximum Times : 10000.0 Target Error %
: 0.200 Minimum Repeats : 10
Maximum Repeats : 100
Benchmark Compile Executable Ave Run Number Estim
Name (secs) (bytes) (secs) Repeats Err %
--------- ------- ---------- ------- ------- ------
ac 0.64 38616 4.50 10 0.1097
aermod 27.17 1118976 3.20 10 0.1696
air 3.40 111024 1.33 22 0.1867
capacita 2.98 102528 6.87 16 0.1163
channel2 0.37 26584 42.15 10 0.1187
doduc 4.34 152376 4.58 16 0.1956
fatigue2 1.53 69200 30.29 10 0.1952
gas_dyn2 1.59 74456 38.47 14 0.1602
induct2 4.13 199800 12.35 12 0.1728
linpk 0.50 30328 2.65 12 0.1263
mdbx 2.11 90352 3.25 10 0.0611
mp_prop_desi 0.72 35248 33.36 10 0.0438
nf 0.72 34848 4.19 10 0.1412
protein 2.30 90544 10.94 10 0.1665
-rnflow 0.00 0 -1.00 10 0.1665
test_fpu2 2.36 76648 15.57 10 0.1457
tfft2 0.52 26616 20.11 13 0.1603
Geometric Mean Execution Time = 11.73 seconds
The geometric mean increased by around 11%. Some tests were
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
So at least for that particular codebase, taking away the
freedom of the compiler to optimize based on the assumption
that integers will never overflow, and replacing it with
something defined, leads to a significant slowdown.
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
Given that UB has been shown to provide only fairly minor performance
gains,
This I found hard to believe, so I ran a
few benchmarks myself. I used the well-known
Polyhedron suite, which can now be downloaded from https://fortran.uk/fortran-compiler-comparisons/the-polyhedron-solutions-benchmark-suite/
,
This is Fortran, so integer overflow is an error. I ran the
testsuite on gfortran with and without -fwrapv, otherwise letting
the compiler go wild (-Ofast). Because gfortran puts arrays on the
stack with -Ofast, I had to increase the stacksize for one particular
code. I had to disable one test, rnflow, because of a bug.
Here are the results:
<snip>
Geometric Mean Execution Time = 10.57 seconds
Geometric Mean Execution Time = 11.73 seconds
The geometric mean increased by around 11%. Some tests were
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
So at least for that particular codebase, taking away the
freedom of the compiler to optimize based on the assumption
that integers will never overflow, and replacing it with
something defined, leads to a significant slowdown.
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default 32-bit.
That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
The thing you want is that your loops with "a[i++]" operations are implemented as "*p++" rather than "*(a + i); i = (i + 1) &
0xffff'ffff)". The later is what you end up with if you use an
unsigned 32-bit type or a signed 32-bit type with forced wrapping
semantics.
So you can either use a 64-bit type (signed or unsigned), or use
normal 32-bit "int" as usual, but without forcing wrapping semantics
on it.
Given that UB has been shown to provide only fairly minor performance
gains,
On Sun, 27 Sep 2026 13:42:05 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
Geometric Mean Execution Time = 11.73 seconds
The geometric mean increased by around 11%. Some tests were
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
So at least for that particular codebase, taking away the
freedom of the compiler to optimize based on the assumption
that integers will never overflow, and replacing it with
something defined, leads to a significant slowdown.
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default 32-bit.
That couses real cost under wrapv.
In C there is no such thing as default size of array index.
The
programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
Michael S <already5chosen@yahoo.com> schrieb:
On Sun, 27 Sep 2026 13:42:05 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:
[Polyhedron benchmark, which is Fortran]
Geometric Mean Execution Time = 11.73 seconds
The geometric mean increased by around 11%. Some tests were
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
So at least for that particular codebase, taking away the
freedom of the compiler to optimize based on the assumption
that integers will never overflow, and replacing it with
something defined, leads to a significant slowdown.
Fortran is not the same as C.
C and Fortran share the same rules about integer overflow.
But if you run any program through the NAG compiler, it is
translated to a C program.
My understanding is that in Fortran array indexes are by default
32-bit.
That understanding turns out to be wrong.
In Fortran, an array index is an integer expression, and that
can be of any KIND that the compiler supports. You can write
use iso_fortran_env, only :: int64
integer(int64) :: i
integer :: j
This will declare a 64-bit integer variable i and a default integer
variable j. You can then access an array A either way, for
exmaple
real, dimension(whatever) :: a
a(i) = something
a(j) = something_else
That couses real cost under wrapv.
In C there is no such thing as default size of array index.
Neither is there in Fortran. There is a default integer,
which is typically 32 bit.
The
programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
A Fortran programmer can do exactly the same; he could even import C_PTRDIFF_T from the ISO_C_BINDING module.
In that respect, there is not difference between C and Fortran.
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance
degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
On 9/24/2026 9:29 AM, Tim Rentsch wrote:
[...] how would you write a
requirement in the C standard to limit the meaning of "you can do
whatever you want, but you must document what you do" that gives only
as much freedom as you think should be allowed?
I have minimal experience with standards writing, and none with
language writing, so I may be off base here.
I think that there are many situations where the standard, correctly,
doesn't specify anything about what the compiler should do. These are currently called undefined behavior, and the standard essentially lets
it go at that - no further documentation required.
But in many, but not all, of such cases, the compiler implementer
knows what it is going to do. What I am after is that in such cases,
the implementer "supplement" the bare words with information that he
has that knowledge and he should communicate it to the user.
I think we will find that different people have very different ideas
about how much freedom should be allowed in saying what happens. It's
a very hard problem.
Note that I am not trying to restrict "what happens" in any way. What happens is anything that could/would happen under today's definition
of undefined behavior. I am just trying to ask the implementer,
whenever possible (and I realize that it is not always possible) to
elaborate on the bare words "undefined behavior".
Michael S <already5chosen@yahoo.com> writes:
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance
degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it
targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
I favor a type equivalent to size_t for use as array index variables.
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
Given that UB has been shown to provide only fairly minor performanceThe geometric mean increased by around 11%. Some tests were
gains,
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
Thomas Koenig [2026-09-27 13:42:05] wrote:
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
Given that UB has been shown to provide only fairly minor performanceThe geometric mean increased by around 11%. Some tests were
gains,
within shouting each other. Others (air, linpk, nf, tfft2) show
significant differences.
FWIW, I consider a 10% performance difference to be fairly minor.
Notice also that my SMUB proposal does allow some of the optimizations
that UB allows w.r.t overflow (e.g. it allows the compiler to consider integer addition as associative) so it would not cost as much as
`-fwrapv`.
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
FWIW, I consider a 10% performance difference to be fairly minor.
That was the geometric mean, the worst one was a factor of 1.54.
Notice also that my SMUB proposal does allow some of the optimizations
that UB allows w.r.t overflow (e.g. it allows the compiler to consider
integer addition as associative) so it would not cost as much as
`-fwrapv`.
Do you mean associative as in a + b = b + a (which is trivially
true) or associative as in (a + b) + c = (c + b) + a, even when
c + b overflows?
You may wonder who writes such code, but
it can be created as the result of template initiation,
devirtualization,
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance
degradatioon under wrapv as Fortran does. Or he can define index as
either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
The thing you want is that your loops with "a[i++]" operations are
implemented as "*p++" rather than "*(a + i); i = (i + 1) &
0xffff'ffff)". The later is what you end up with if you use an
unsigned 32-bit type or a signed 32-bit type with forced wrapping
semantics.
So you can either use a 64-bit type (signed or unsigned), or use
normal 32-bit "int" as usual, but without forcing wrapping semantics
on it.
You may wonder who writes such code, but
it can be created as the result of template initiation,
I.e., C programmers don't write such code.
devirtualization,
I.e., C programmers don't write such code.
I think that Fortran programmers don't write such code, either,
Thomas Koenig <tkoenig@netcologne.de> writes:
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
FWIW, I consider a 10% performance difference to be fairly minor.
Concerning -fwrapv: The associative law holds for + and * in modulo arithmetics, so the compiler can perform optimizations that rely on
the associative law when the programmer asks for modulo arithmetics
with -fwrapv. That's not the case if overflow is trapped (-ftrapv, or
its long-winded -fsanitize= cousin).
You may wonder who writes such code, but
it can be created as the result of template initiation,
I.e., C programmers don't write such code.
devirtualization,
I.e., C programmers don't write such code.
I think that Fortran programmers don't write such code, either,
- anton
Tim Rentsch wrote:
Michael S <already5chosen@yahoo.com> writes:
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance
degradatioon under wrapv as Fortran does. Or he can define index as >>>>> either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it
targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
I favor a type equivalent to size_t for use as array index variables.
In Rust, the only acceptable array index is of type "usize", which is effectively the same as u64 on most platforms these days, but does not
need to be so: On a 32-bit target, it would be the same as u32.
So effectively very similar to size_t.
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Michael S <already5chosen@yahoo.com> writes:
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same performance >>>>>> degradatioon under wrapv as Fortran does. Or he can define index as >>>>>> either ptrdiff_t or size_t and suffere no degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit is
enough range for a great many purposes) but can bigger if that's
faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it >>>> targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness >>>> and acceptable geekery.
I favor a type equivalent to size_t for use as array index variables.
In Rust, the only acceptable array index is of type "usize", which is
effectively the same as u64 on most platforms these days, but does not
need to be so: On a 32-bit target, it would be the same as u32.
So effectively very similar to size_t.
Maybe that limitation works okay in Rust. In C sometimes it's
important to allow a signed value for indexing. Probably 99% of
the time size_t is good, but not always.
On 9/9/26 2:41 PM, MitchAlsup wrote:
[snip]
antispam@fricas.org (Waldek Hebisch) posted:
Once you have paravirtualization, there is nasty question how
guest and host should divide the work and what is the best
interface.
Host OS needs an entry point to Guest OS to ask for resources back
while allowing Guest OS to determine which resources to return.
Guest OS needs an entry point in Host OS to ask for more resources
while allowing Host OS to determine which resources to send.
I wonder if it would make sense for a hardware architecture to
define an interface to its own hypervisor (vaguely reminiscent
to Alpha's PALcode?) to make at least a limited form of
paravirtualization natural. (I am thinking of page tables and
some resource management.)
I have a personal dislike of nested page tables and some
Nested Paging has won. Guest OS virtualizes applications and its own
worker threads, while Host OS virtualized multiple Guest OSs.
Nested paging potentially doubles the page table depth, which
seems unattractive to me. (I am aware that caching intermediate
page table nodes and using large pages can avoid much of this
overhead. I almost certainly undervalue virtualization and my
sense of ickiness for nested paging is not deeply informed.)
Michael S <already5chosen@yahoo.com> writes:
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it
targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
I favor a type equivalent to size_t for use as array index variables.
On 26/09/2026 16:25, Stefan Monnier wrote:
Remember, we're starting from the problem that UB is used byIf that is truly the way you think, then you should maybe give up
compilers in ways which catch programmers off-guard.
programming and start a conspiracy-theory pod-cast. (You'll probably make more money!) Seriously - if you start from the suggestion that compiler writers are aiming to catch programmers off-guard, or that they disregard
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
Michael S <already5chosen@yahoo.com> writes:
On Sun, 27 Sep 2026 18:04:05 +0200
David Brown <david.brown@hesbynett.no> wrote:
On 27/09/2026 17:18, Michael S wrote:
Fortran is not the same as C.
My understanding is that in Fortran array indexes are by default
32-bit. That couses real cost under wrapv.
In C there is no such thing as default size of array index. The
programmer can define it as 32-bit and suffer the same
performance degradatioon under wrapv as Fortran does. Or he
can define index as either ptrdiff_t or size_t and suffere no
degradation.
You certainly /can/ do that. But a lot of C code uses "int" for
array indexing. The "ideal" type in many cases would be
"int_fast32_t" to say it should be of a big enough size (32-bit
is enough range for a great many purposes) but can bigger if
that's faster. On most 64-bit targets, "int_fast32_t" will be
64-bit.
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what
OS it targets.
Besides, int_fast32_t is both above my threshold of acceptable
ugliness and acceptable geekery.
I favor a type equivalent to size_t for use as array index
variables.
In Rust, the only acceptable array index is of type "usize", which
is effectively the same as u64 on most platforms these days, but
does not need to be so: On a 32-bit target, it would be the same
as u32.
So effectively very similar to size_t.
Maybe that limitation works okay in Rust. In C sometimes it's
important to allow a signed value for indexing. Probably 99% of
the time size_t is good, but not always.
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
Notice also that my SMUB proposal does allow some of the optimizations
that UB allows w.r.t overflow (e.g. it allows the compiler to consider
integer addition as associative) so it would not cost as much as
`-fwrapv`.
Do you mean associative as in a + b = b + a (which is trivially
true) or associative as in (a + b) + c = (c + b) + a, even when
c + b overflows?
But your proposal, as I understand it, would destroy a lot of
range-based optimization, so that, for example, the second if
statement in
if (a > 5) {
if (a + 1 > 3) {
}
}
would be executed.
You may wonder who writes such code, but
David Brown [2026-09-27 16:40:15] wrote:
On 26/09/2026 16:25, Stefan Monnier wrote:
Remember, we're starting from the problem that UB is used byIf that is truly the way you think, then you should maybe give up
compilers in ways which catch programmers off-guard.
programming and start a conspiracy-theory pod-cast. (You'll probably make >> more money!) Seriously - if you start from the suggestion that compiler
writers are aiming to catch programmers off-guard, or that they disregard
I never claimed it is their aim.
The fact that programmers are surprised by the behavior of the generated
code (e.g. cases described as "nasal demons") just reflects the fact
that their understanding of undefined behavior doesn't match the reality
and my SMUB is an attempt to specify another flavor of undefined
behavior which is hoped to:
- Align more closely with what the average programmers expect when
they hear "undefined behavior".
- Still allow enough flexibility for the implementation that the impact
on performance is usually small.
I'm not blaming anyone. My finger is pointed only at the discrepancy
between what UB is understood to mean by the C spec and what it is
understood to mean by average programmers.
[ And I don't see much hope to fix the programmers' understanding, so
I think the only way to reduce this discrepancy is to change the
C spec. ]
=== Stefan
So, what happens when you edit source code of Polyhedron benchmark to
use of C_PTRDIFF_T for aray indices?
Is there still an impact for -fwrapv ?
Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
Michael S <already5chosen@yahoo.com> writes:
<int_fast_32_discussion>
That's not universal.
For example, on both alive 64-bit Windows targets int_fast32_t is
32-bit.
Godbolt shows the same for clang for MIPS64. I have no idea what OS it
targets.
Besides, int_fast32_t is both above my threshold of acceptable ugliness
and acceptable geekery.
I favor a type equivalent to size_t for use as array index variables.
I -use- size_t for array index variables.
On Mon, 28 Sep 2026 02:59:06 -0700[...]
Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:
Terje Mathisen <terje.mathisen@tmsw.no> writes:
Tim Rentsch wrote:
I favor a type equivalent to size_t for use as array index
variables.
In Rust, the only acceptable array index is of type "usize", which
is effectively the same as u64 on most platforms these days, but
does not need to be so: On a 32-bit target, it would be the same
as u32.
So effectively very similar to size_t.
Maybe that limitation works okay in Rust. In C sometimes it's
important to allow a signed value for indexing. Probably 99% of
the time size_t is good, but not always.
My only Rust program was full of casts of various integer types to
usize. Exactly for the purpose of array indexing.
[snip]
My only Rust program was full of casts of various integer types to
usize. Exactly for the purpose of array indexing.
In article <20260928195833.00001d42@yahoo.com>,
Michael S <already5chosen@yahoo.com> wrote:
[snip]
My only Rust program was full of casts of various integer types to
usize. Exactly for the purpose of array indexing.
Beginning Rust programmers tend to go through a few phases,
particularly if coming from a language like C. "Fighting the
Borrow Checker" is well-known; "traits are awesome; let's use
them everywhere!" is another.
When I started programming in Rust, most of my immediate
colleagues and I had terribly little experience in the language.
We were excited about the its potential benefits, but realized
it would take a few months to build up a sufficient level of
familiarity to use it skillfully. One quipped, "this language
has a near-vertical learning curve."
At one point a coworker (self-deprecatingly) said, "look, I'm
writing crust! C in Rust syntax!" Everyone thought this was a
great play on words, but he was correct: the programs we were
writing were essentially C programs, just expressed in Rust. We
quickly adopted that to describe our still naive code: it wasn't
an insult, just an expression that we were still coming to terms
with a new way of doing things. Some of the hallmarks of that
included using the wrong paradigms (and types!) for lots of our
code.
I'd be curious to know more about your program.
I've written
several hundred kloc of Rust at this point, and while what I
write may be terribly specialized, I rarely find myself indexing
an array; often, I'm using iterators or some other abstraction
for that kind of thing. If that was the one program you have
written in Rust, my _guess_ is that you were using paradigms
appropriate to a different language, which is why you felt you
needed to reach for casts to index an array so frequently.
Again, that's not a comment on your program, so much as
suggesting an impedence mismatch between what you wrote and how
one would generally expect to use the language.
In general, I think a person's experience writing a single
program tells us very little.
- Dan C.
In article <20260928195833.00001d42@yahoo.com>,
Michael S <already5chosen@yahoo.com> wrote:
[snip]
My only Rust program was full of casts of various integer types to
usize. Exactly for the purpose of array indexing.
Beginning Rust programmers tend to go through a few phases,
particularly if coming from a language like C. "Fighting the
Borrow Checker" is well-known; "traits are awesome; let's use
them everywhere!" is another.
When I started programming in Rust, most of my immediate
colleagues and I had terribly little experience in the language.
We were excited about the its potential benefits, but realized
it would take a few months to build up a sufficient level of
familiarity to use it skillfully. One quipped, "this language
has a near-vertical learning curve."
At one point a coworker (self-deprecatingly) said, "look, I'm
writing crust! C in Rust syntax!" Everyone thought this was a
great play on words, but he was correct: the programs we were
writing were essentially C programs, just expressed in Rust. We
quickly adopted that to describe our still naive code: it wasn't
an insult, just an expression that we were still coming to terms
with a new way of doing things. Some of the hallmarks of that
included using the wrong paradigms (and types!) for lots of our
code.
Dan Cross wrote:
In article <20260928195833.00001d42@yahoo.com>,
Michael S <already5chosen@yahoo.com> wrote:
[snip]
My only Rust program was full of casts of various integer types to
usize. Exactly for the purpose of array indexing.
Beginning Rust programmers tend to go through a few phases,
particularly if coming from a language like C. "Fighting the
Borrow Checker" is well-known; "traits are awesome; let's use
them everywhere!" is another.
When I started programming in Rust, most of my immediate
colleagues and I had terribly little experience in the language.
We were excited about the its potential benefits, but realized
it would take a few months to build up a sufficient level of
familiarity to use it skillfully. One quipped, "this language
has a near-vertical learning curve."
At one point a coworker (self-deprecatingly) said, "look, I'm
writing crust! C in Rust syntax!" Everyone thought this was a
great play on words, but he was correct: the programs we were
writing were essentially C programs, just expressed in Rust. We
quickly adopted that to describe our still naive code: it wasn't
an insult, just an expression that we were still coming to terms
with a new way of doing things. Some of the hallmarks of that
included using the wrong paradigms (and types!) for lots of our
code.
I have definitely noted this in my own code, crust is a good name for it.
Quite often, when "fighting the borrow checker", typically because I'm updating a complicated data structure in two locations at the same time, I'll fall back to "write Fortran in Rust", with every access indexed.
This is BTW one of those problems the brand new upcoming borrow checker
is supposed to be far better at figuring out: If I'm updating totally separate parts of the same larger structure, then it should be OK to
mutate both at the same time.
I do realize that there are far more Rusty ways to solve this type of
issue, they just don't come naturally to me yet.
Terje
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 121:08:36 |
| Calls: | 1,194 |
| Files: | 1,352 |
| Messages: | 290,208 |