• Re: PDP-10 Styles of addressing

    From John Levine@johnl@taugh.com to comp.arch on Wed Sep 2 17:32:48 2026
    From Newsgroup: comp.arch

    According to Kragen Javier Sitaker <kragen@canonical.org>:
    John Levine <johnl@taugh.com> writes:

    It appears that MitchAlsup <user5857@newsgrouper.org.invalid> said:
    The PDP-6 and -10 could take an interrupt before
    each address calculation so if an interrupt arrived in the middle of
    an indirect chain, it just started it over when the interrupt returned.
    (...)
    One time when I was supposed to be doing something else I wrote a little
    program that made a longer and longer interrupt chain until the program
    stalled, which told me how often the clock interrupted.

    For rCLinterrupt chainrCY here we should read rCLindirect chainrCY, I think?

    Quite right, indirect chain. I might have chained XCT instructions too,
    same idea.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 2 11:19:44 2026
    From Newsgroup: comp.arch

    On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:

    This infinite indirect capability made PDP-10 loved by people using
    them. It also caused problems when the infinite indirect timer "went
    off".

    I always wondered 1) was there an upper limit to the number of levels
    of indirection, and 2) what that did to the interrupt latency ...

    I had to research this to check my memory, but on the 1100 series, the
    timer value was 100 microseconds. If an instruction took longer than
    that, it got an invalid operation interrupt/exception. While this may
    seem long by today's standards, remember, this was the 1960s and typical instructions on the 1108 took 750 nanoseconds, and peripherals were a
    lot slower. In practice, I never found it to be a problem as
    indirection was rare and I never saw more than two levels.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 2 20:39:22 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> schrieb:
    On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:

    This infinite indirect capability made PDP-10 loved by people using
    them. It also caused problems when the infinite indirect timer "went
    off".

    I always wondered 1) was there an upper limit to the number of levels
    of indirection, and 2) what that did to the interrupt latency ...

    I had to research this to check my memory, but on the 1100 series, the
    timer value was 100 microseconds. If an instruction took longer than
    that, it got an invalid operation interrupt/exception. While this may
    seem long by today's standards, remember, this was the 1960s and typical instructions on the 1108 took 750 nanoseconds, and peripherals were a
    lot slower. In practice, I never found it to be a problem as
    indirection was rare and I never saw more than two levels.

    See the recent assembly hall of shame posting :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 2 21:09:03 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:

    This infinite indirect capability made PDP-10 loved by people using
    them. It also caused problems when the infinite indirect timer "went
    off".

    I always wondered 1) was there an upper limit to the number of levels
    of indirection, and 2) what that did to the interrupt latency ...

    I had to research this to check my memory, but on the 1100 series, the
    timer value was 100 microseconds. If an instruction took longer than
    that, it got an invalid operation interrupt/exception. While this may
    seem long by today's standards, remember, this was the 1960s and typical >instructions on the 1108 took 750 nanoseconds, and peripherals were a
    lot slower. In practice, I never found it to be a problem as
    indirection was rare and I never saw more than two levels.

    The Burroughs B3500 had similar operand recursion (unlimited);
    and also had an internal timeout counter
    that would terminate the instruction if it expired. While
    recursive operands were seldom more than two or three indirections
    deep, the search linked list (SLT) instruction could easily
    run forever if there was a loop in the linked list; the
    internal timeout would kill the job and take a memory dump.

    From the Engineering specification for the V500 Execute Module (XM):

    The timeout counter is a three bit grey code counter that
    is enabled by the 0.2 or 0.4 second output of the Task timer.
    This gives a timeout of 1.6 or 3.2 seconds. The value of
    the timeout is selected by an external strap. The timeout
    counter is reset at the start of each instruction, upon
    entry to the interrupt microcode by the OpLit and interrupt
    signals or by the clear timeout counter internal command.

    https://bitsavers.org/pdf/burroughs/MediumSystems/V500/1993-5204B_V500-Execute-Module.pdf
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Wed Sep 2 22:01:15 2026
    From Newsgroup: comp.arch

    According to Stephen Fuld <sfuld@alumni.cmu.edu.invalid>:
    On 9/1/2026 7:42 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 20:04:53 GMT, MitchAlsup wrote:

    This infinite indirect capability made PDP-10 loved by people using
    them. It also caused problems when the infinite indirect timer "went
    off".

    I always wondered 1) was there an upper limit to the number of levels
    of indirection, and 2) what that did to the interrupt latency ...

    I had to research this to check my memory, but on the 1100 series, the
    timer value was 100 microseconds. ...

    Sounds right. As I said a few messages back, the PDP-10 could take an interrupt
    before each address calculation so the effect on interrupts was insignificant. It did have a BLT block transfer instruction which could be interrupted before each word transfer. The transfer addresses were taken from an AC so it put
    the updated addresses in the AC before the interrupt and left the interrupt PC pointing to the BLT so it would resume after the interrupt. Again, not a big deal for latency.

    I know of a few PDP-10's that were use for process monitoring or
    control so it mattered.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Wed Sep 2 18:59:58 2026
    From Newsgroup: comp.arch

    On 9/1/26 5:17 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 10:07:14 -0700, Andy Valencia wrote:

    Certainly this happened with Stanford University's homebrew DBMS,
    SPIRES. My father-in-law told me of the immense work needed to free
    up the address bits to permit extended addressing. He said when they
    did the original code, nobody imagined that they'd ever need those
    upper bits for addressing.

    The original Apple MacOS went through this in the 1980s.

    The 68000-family instruction set allowed for 32-bit addresses, but the original 68000 processor only looked at the bottom 24 bits. So the Mac
    system used those top 8 bits for various memory-management purposes.

    Even when they brought out the Mac II, the first with a 68020
    processor that *did* pay attention to all 32 bits of the address, they
    stuck in a stub MMU on the motherboard that removed those top 8 bits
    before actually accessing memory.

    Then, by about 1990, they introduced models with rCL32-bit cleanrCY ROMs, that didnrCOt make any such odd uses of those top 8 bits internally, and could address more than 16MiB of RAM.

    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    This would have introduced overhead for new OSes (having to
    configure the bit), but would allow old software to run
    correctly while allowing use of the 32-bit address space.
    (Determining when software can use 32-bit mode might be
    difficult. User pointers in system calls would also need
    to be masked, but I think an OS already needs to prevent
    OS-privilege access to arbitrary user-provided addresses.)

    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software. Many
    consider this a serious architecture design mistake, citing
    history where such has caused compatibility headaches. I _feel_
    that system software should be able to handle this extra
    complexity without great difficulty, but I have never developed
    even a task scheduler much less an OS. Since some OSes support
    use of Top Byte Ignore, I am guessing this is not a huge
    problem.

    Interestingly, Stanford MIPS used the extra bits for an address
    space number and had a variable length mask. (-2The size of the
    process virtual address space is defined by a bit mask in a
    special register. When masking is enabled for normal operations,
    a process identifier from another special register is
    substituted for the high order bits of the machine address.
    These two special registers are accessible only to processes
    running in supervisor state. The masking unit also detects
    attempts to access outside the legal segment and raises an
    exception to the master pipeline control. Although the virtual
    address is a full 32 bits, the package constraint only permits
    24 address pins. With word addressing, however, this gives an
    address space of 64 megabytes.-+, "Design of a High-Performance
    VLSI Processor", John Hennessey et al., 1983)

    The masking in Stanford MIPS did not allow software flags but
    did allow some process isolation even without a permission
    table. Such also allowed the ASID to vary in size without
    requiring more address bits if some processes were compact.

    If I recall correctly, 32-bit ARM provided a means to define the
    size of the user address space such that a variable number of
    more significant bits were used as system address bits (similar
    to the negative addresses are system addresses of some OSes)
    with a separate page table. This did not allow software use of
    the extra bits, but provided some flexibility in page tables
    (the number of levels for the user page table could be decreased
    and the entire system address space page table could be shared
    rather than duplicating the root "page" in each process) and
    PTEs would not need a global bit to indicate preservation across
    processes. (32-bit PowerPC segments were vaguely similar in
    allowing 16 segments with separate virtual address spaces. In
    theory, this might support somewhat flexible sharing. HP PA-RISC
    provided fewer segments but more flexible/complex use. [I think
    encoding the segment bits in the least significant bits would
    have been better than using the high bits; it would have made
    the change to 64-bit simpler and allowed full-space dynamic
    segment addresses at the cost of having to encode segment
    numbers in the instruction for byte and half-word aligned
    pointers and disallowing alignment traps based on such bits.])

    Implicit hardware masking of tags can be useful, though I do
    wonder if such could be usefully generalized to better support
    multiple data in a single load. (32-bit paired loads in AArch64
    do this to some degree, but it might be practical to exploit the
    alignment network for loads to parse a load into two pieces at
    byte granularity. An unaligned address could indicate the
    split point.)

    With x86 one can perform a full-register-size load and access
    sub-sections (for some registers). Providing denser register
    storage use has advantages when memory is slow, but using
    subregisters complicates renaming (and forwarding, which is
    kind of renaming). (Like CMOV, out-of-order execution
    complicates use of subregisters.) It might be possible
    (practical even) to support a restrained use of subregisters
    (less restrained than SIMD that defines a single size and
    divides the register into lanes of that size), but it seems
    likely that such would not be *worthwhile*.

    Even without address masking, part of addresses can be used
    to hold type information at the cost of sparser use of the
    address space. (This introduces a possible performance
    compatibility concern. Software may seek to minimize translation
    overhead from sparsity, but that ties software performance to
    specifics of address translation (like which bits are used for
    table levels). Even using a tag to indicate a leaf pointer might
    be useful with a tracing garbage collector.)

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    [I better send this before my mind wanders even farther.]
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 3 01:50:23 2026
    From Newsgroup: comp.arch


    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/1/26 5:17 PM, Lawrence DrCOOliveiro wrote:
    On Tue, 01 Sep 2026 10:07:14 -0700, Andy Valencia wrote:

    Certainly this happened with Stanford University's homebrew DBMS,
    SPIRES. My father-in-law told me of the immense work needed to free
    up the address bits to permit extended addressing. He said when they
    did the original code, nobody imagined that they'd ever need those
    upper bits for addressing.

    The original Apple MacOS went through this in the 1980s.

    The 68000-family instruction set allowed for 32-bit addresses, but the original 68000 processor only looked at the bottom 24 bits. So the Mac system used those top 8 bits for various memory-management purposes.

    Sins of the past...

    Even when they brought out the Mac II, the first with a 68020
    processor that *did* pay attention to all 32 bits of the address, they stuck in a stub MMU on the motherboard that removed those top 8 bits
    before actually accessing memory.

    Then, by about 1990, they introduced models with rCL32-bit cleanrCY ROMs, that didnrCOt make any such odd uses of those top 8 bits internally, and could address more than 16MiB of RAM.

    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    This would have introduced overhead for new OSes (having to
    configure the bit), but would allow old software to run
    correctly while allowing use of the 32-bit address space.
    (Determining when software can use 32-bit mode might be
    difficult. User pointers in system calls would also need
    to be masked, but I think an OS already needs to prevent
    OS-privilege access to arbitrary user-provided addresses.)

    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    Many
    consider this a serious architecture design mistake, citing
    history where such has caused compatibility headaches. I _feel_
    that system software should be able to handle this extra
    complexity without great difficulty, but I have never developed
    even a task scheduler much less an OS. Since some OSes support
    use of Top Byte Ignore, I am guessing this is not a huge
    problem.

    Interestingly, Stanford MIPS used the extra bits for an address
    space number and had a variable length mask. (-2The size of the
    process virtual address space is defined by a bit mask in a
    special register.

    My 66000 has a 3-bit LVL field in the Root pointer (in all MMU
    pointers). LVL in root tells you the VAS size (and PA==VA).
    A process==thread can have VAS as small as 23-bits (cat) and
    use 1 page as the page table. Saving all those unnecessary
    MMU accesses.

    LVL in PTPs allows for level skipping (sparse spaces). LVL in
    the pointer pointing at a PTE tells you size of the (super)page.
    No need for a control register field to specify what the natural
    organization of the table structure provides.

    When masking is enabled for normal operations,
    a process identifier from another special register is
    substituted for the high order bits of the machine address.

    My 66000 assigns ASID as Address<79..64> and also drags priority
    around the on-die interconnect to allow memory to decide on the
    order of near-identical-timing interfering ATOMIC accesses (higher
    wins).

    These two special registers are accessible only to processes
    running in supervisor state. The masking unit also detects
    attempts to access outside the legal segment and raises an
    exception to the master pipeline control. Although the virtual
    address is a full 32 bits, the package constraint only permits
    24 address pins. With word addressing, however, this gives an
    address space of 64 megabytes.-+, "Design of a High-Performance
    VLSI Processor", John Hennessey et al., 1983)

    The masking in Stanford MIPS did not allow software flags but
    did allow some process isolation even without a permission
    table. Such also allowed the ASID to vary in size without
    requiring more address bits if some processes were compact.

    A valid design point in 1983, not so much today--unless vastly
    extended.

    If I recall correctly, 32-bit ARM provided a means to define the
    size of the user address space such that a variable number of
    more significant bits were used as system address bits (similar
    to the negative addresses are system addresses of some OSes)
    with a separate page table. This did not allow software use of
    the extra bits, but provided some flexibility in page tables
    (the number of levels for the user page table could be decreased
    and the entire system address space page table could be shared
    rather than duplicating the root "page" in each process) and
    PTEs would not need a global bit to indicate preservation across
    processes.

    Since My 66000 Core-State has 4 contexts installed at all times,
    there are 4 {ASIDs, priorities, privileges, and threads} iden-
    tified at all times, allowing greater privilege threads to access
    lesser privileged threads without any "global" bit in PTE/PTP/TLB.
    Basically ASID does this cleaner. HV can reach into Host OS, while
    Host OS can reach into Guest OS. Guest OS can reach into application
    without any of them overusing/overspecifying the concept of 'global'
    and there can be as many as desired 'global' spaces--all specified
    in the MMU tables.

    I, personally, do not see a single 'global' bit being "all that
    useful" when there are 27 Guest OSs running simultaneously, all
    27 thinking they are the only OSs running at that instant in
    time. These Guest OSs might be running under 10 Host OSs which
    might be running under 3 HyperVisors.

    Handing out unique ASIDs is a very easy job for HV.

    (32-bit PowerPC segments were vaguely similar in
    allowing 16 segments with separate virtual address spaces. In
    theory, this might support somewhat flexible sharing. HP PA-RISC
    provided fewer segments but more flexible/complex use. [I think
    encoding the segment bits in the least significant bits would
    have been better than using the high bits; it would have made
    the change to 64-bit simpler and allowed full-space dynamic
    segment addresses at the cost of having to encode segment
    numbers in the instruction for byte and half-word aligned
    pointers and disallowing alignment traps based on such bits.])

    More sins of the past...

    Implicit hardware masking of tags can be useful, though I do
    wonder if such could be usefully generalized to better support
    multiple data in a single load. (32-bit paired loads in AArch64
    do this to some degree, but it might be practical to exploit the
    alignment network for loads to parse a load into two pieces at
    byte granularity. An unaligned address could indicate the
    split point.)

    Looking forward, loading 128-256 bits per access might be useful
    in the not so distant future, if only used for complex types
    {real, imag} with 64-bit and 128-bit element sizes.

    With x86 one can perform a full-register-size load and access
    sub-sections (for some registers). Providing denser register
    storage use has advantages when memory is slow, but using
    subregisters complicates renaming (and forwarding, which is
    kind of renaming).

    The understatement of the month award winner.

    (Like CMOV, out-of-order execution
    complicates use of subregisters.)

    For 70 years, a register would contain a single value, then
    MMX ruined the game...prior was the model compilers are good
    at using.

    It might be possible
    (practical even) to support a restrained use of subregisters
    (less restrained than SIMD that defines a single size and
    divides the register into lanes of that size), but it seems
    likely that such would not be *worthwhile*.

    SIMD is bad for your architecture and for your thinking processes.

    Even without address masking, part of addresses can be used
    to hold type information at the cost of sparser use of the
    address space.

    The only thing your architecture cannot survive (over time)
    is lack of address bits. Do not give them away before your 4th
    generation implementations have sold 100M chips.

    (This introduces a possible performance
    compatibility concern. Software may seek to minimize translation
    overhead from sparsity, but that ties software performance to
    specifics of address translation (like which bits are used for
    table levels). Even using a tag to indicate a leaf pointer might
    be useful with a tracing garbage collector.)

    Simply a symptom of poor MMU table design.

    On the other hand, I have used a strategy where when I need some
    structure that can become garbage, I fork off a process, have it
    create the structure, then distill a result, then have it return
    the result and terminate itself. Almost always faster, always
    safer.

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    With ECAM based on PCIe 4.0+, your device address space consumes
    up to 40-bits. And then there is the configuration space, and the
    interrupt table aperture space, and other system spaces--all vying
    for those 47-bits. I suspect 47 will not be sufficient very long.

    My 66000's solution is to provide a complete 64-bit VAS that can be
    translated into a 66-bit universal address space consisting of four
    64-bit spaces {DRAM, device, config, ROM}. I don't want any of these
    spaces to cause issues while I remain alive. Both cores and devices
    use the 66-bit UAS.

    Using MSI-X interrupts, allows interrupt tables to perform DPC/softIRQ
    queueing without normally associated SW overheads. Cores can send cores interrupts using the same mechanisms as devices. Since there is an
    unlimited number of interrupt tables (My 66000) and since the are
    constructed with a DAG structure, anything that can touch the interrupt aperture can send interrupts to any virtual core. When that virtual
    core has control, those interrupts are 'processed'.

    [I better send this before my mind wanders even farther.]
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Thu Sep 3 02:46:48 2026
    From Newsgroup: comp.arch

    On Wed, 2 Sep 2026 18:59:58 -0400, Paul Clayton wrote:

    I have wondered why Motorola did not add a 24-bit address mode ...

    You mean in later chips? I guess it was considered unnecessary, given
    that external MMU-type hacks like the one Apple stuck in could, and
    did, solve the problem.

    The original 68000 (with a 16-bit bus) only supported 24-bit
    addressing. There was also the 68008 (with an 8-bit bus) that only
    looked at the bottom 20 bits of the address. That was used in the
    Sinclair QL machine. I wonder if that was the most expensive (i.e.
    least cheap) CPU chip that Sir Clive ever put into one of his machines
    ...

    This would have introduced overhead for new OSes (having to
    configure the bit), but would allow old software to run correctly
    while allowing use of the 32-bit address space.

    Apple basically gave early Mac developers an ultimatum: stop assuming
    those top 8 bits could be used for non-address purposes, or else.

    The move to 32-bit cleanliness in the Mac world went quite quickly.
    Compare how it took about a decade for the DOS/Windows world to move
    from 16-bit software APIs to 32-bit.

    If I recall correctly, 32-bit ARM provided a means to define the
    size of the user address space such that a variable number of more significant bits were used as system address bits (similar to the
    negative addresses are system addresses of some OSes) with a
    separate page table.

    This must have been later. The original ARM chip didnrCOt have anything resembling an MMU, as far as I know.

    That original chip was an absolute legend of advanced CPU design on a shoestring. I think this video
    <https://www.youtube.com/watch?v=t59EtDxpYmM> gives, at one point, an
    overview of how the chipset worked. There is a rCLmemory-controllerrCY
    chip (MEMC) and a rCLvideo-controllerrCY chip (VIDC). To keep the pin
    count down, one of these sees the address bits but not the data bits,
    while the other sees the data bits but not the address bits!

    And also the addressing was kept to 26 bits. Why? So that the complete
    CPU state on an interrupt could be stored in 32 bits.

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and user
    space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    I wonder why 64-bit architectures didnrCOt reserve the bottom 3 bits for
    a bit offset within a byte. It would be nice to see some support for arbitrarily bit-aligned pointers to arbitrary-length bitfields ...
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 3 14:40:35 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> writes:
    On 9/1/26 5:17 PM, Lawrence DrCOOliveiro wrote:


    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software. Many
    consider this a serious architecture design mistake, citing
    history where such has caused compatibility headaches. I _feel_
    that system software should be able to handle this extra
    complexity without great difficulty, but I have never developed
    even a task scheduler much less an OS. Since some OSes support
    use of Top Byte Ignore, I am guessing this is not a huge
    problem.

    This capability (TBI) is also leveraged by the PAC
    extensions (Pointer Authentication).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Thu Sep 3 15:48:42 2026
    From Newsgroup: comp.arch

    On Wed, 2 Sep 2026 18:59:58 -0400, Paul Clayton
    <paaronclayton@gmail.com> wrote:

    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier >implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    My immediate reaction would be that they didn't do that for the same
    reason they didn't provide a 23-bit address mode and a 25-bit address
    mode.

    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    I'm just not sure what a 24-bit address mode would entail.
    Instructions had 16 bit displacements in them, so a 24-bit address
    mode wouldn't make them any shorter. Not using the top eight bits in
    an address register doesn't save memory.

    But addresses stored in memory would be shorter as data! Yes, but now
    they would have to be fetched on byte boundaries instead of on word
    boundaries. And a multiplication by three would be required to fetch
    them from packed memory-saving arrays! So you might save a little
    memory, but you would lose time.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Schultz@david.schultz@earthlink.net to comp.arch on Thu Sep 3 12:07:03 2026
    From Newsgroup: comp.arch

    On 9/3/26 10:48 AM, John Savard wrote:
    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    23.

    The MC68000 had A1-A23. The upper and lower data strobes selected which byte(s) to access.
    --
    http://davesrocketworks.com
    David Schultz
    "It's just this little chromium switch here..."
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 3 17:32:25 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Wed, 2 Sep 2026 18:59:58 -0400, Paul Clayton
    <paaronclayton@gmail.com> wrote:

    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier >>implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    My immediate reaction would be that they didn't do that for the same
    reason they didn't provide a 23-bit address mode and a 25-bit address
    mode.

    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    The virtual address space[*] is still 32 bits. The number of pins on the package is a implementation detail, not an architectural limitation.

    [*] on the versions of the 68k with an MMU (68030+ or the MC68851).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 3 17:33:25 2026
    From Newsgroup: comp.arch

    David Schultz <david.schultz@earthlink.net> writes:
    On 9/3/26 10:48 AM, John Savard wrote:
    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    23.

    The MC68000 had A1-A23. The upper and lower data strobes selected which >byte(s) to access.

    Doesn't that mean that those strobes are effectively A0?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Thu Sep 3 14:43:24 2026
    From Newsgroup: comp.arch

    On 2026-Sep-02 18:59, Paul Clayton wrote:

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    In those prior "24-bit address space" designs the MMU didn't
    look at the upper 8 address bits. That is why they had later
    software porting problems.

    In the current "48-bit address space" MMU's, all 64 address bits are validated. The 48-bits is a MMU model specfic VA sub-space translation limit
    set by the number of levels in the page table.
    Other MMU models can have different numbers of levels.
    x64 originally had 4 levels of page table (9-9-9-9-12 = 48 bits) and
    around 2017 added level 5 supporting a 57-bit virtual address sub-space
    within the ISA's 64-bit virtual address space.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Thu Sep 3 20:54:40 2026
    From Newsgroup: comp.arch

    In article <DFfmS.228899$aK2e.41624@fx17.iad>, scott@slp53.sl.home (Scott Lurndal) wrote:

    Paul Clayton <paaronclayton@gmail.com> writes:
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software. Many
    consider this a serious architecture design mistake, citing
    history where such has caused compatibility headaches. I _feel_
    that system software should be able to handle this extra
    complexity without great difficulty, but I have never developed
    even a task scheduler much less an OS. Since some OSes support
    use of Top Byte Ignore, I am guessing this is not a huge
    problem.

    This capability (TBI) is also leveraged by the PAC
    extensions (Pointer Authentication).

    I've implemented PAC in application software without any serious
    difficulty.

    TBI is optional. It may require ARM to move from ARM64 to a hypothetical
    ARM128 sooner than would otherwise have been the case, but PAC helps with today's security problems, which are fairly serious.

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 3 20:23:54 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    David Schultz <david.schultz@earthlink.net> writes:
    On 9/3/26 10:48 AM, John Savard wrote:
    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    23.

    The MC68000 had A1-A23. The upper and lower data strobes selected which >byte(s) to access.

    Doesn't that mean that those strobes are effectively A0?

    In unary, yes; or you can describe them as the decode of size and A<0>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 3 20:32:14 2026
    From Newsgroup: comp.arch


    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-02 18:59, Paul Clayton wrote:

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly mask/extract small pointers to use those extra bits.

    In those prior "24-bit address space" designs the MMU didn't
    look at the upper 8 address bits. That is why they had later
    software porting problems.

    In the current "48-bit address space" MMU's, all 64 address bits are validated.
    The 48-bits is a MMU model specfic VA sub-space translation limit
    set by the number of levels in the page table.

    I believe it is set by the used width of PA<63..12> in the PTE in
    conjunction to the number of address bits that are routed around
    the on-die interconnect. Using PA<63..48> as other than 12b'0
    is asking for memory aliasing problem in the coherence protocol.

    Other MMU models can have different numbers of levels.
    x64 originally had 4 levels of page table (9-9-9-9-12 = 48 bits) and
    around 2017 added level 5 supporting a 57-bit virtual address sub-space within the ISA's 64-bit virtual address space.

    5-levels and nested paging has a cost of 25 accesses when walking a
    complete 5|u5 pair of mapping tables. Thus requiring several kinds of table-walk accelerators (top skippers) and (early outs). But even here,
    this is much better than a SW managed 1-level MMU doing the same kinds
    of translation.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Thu Sep 3 16:44:52 2026
    From Newsgroup: comp.arch

    On 2026-Sep-03 13:33, Scott Lurndal wrote:
    David Schultz <david.schultz@earthlink.net> writes:
    On 9/3/26 10:48 AM, John Savard wrote:
    But since, as you point out, the chip only had 24 address pins, your
    question indeed has a great deal of validity.

    23.

    The MC68000 had A1-A23. The upper and lower data strobes selected which
    byte(s) to access.

    Doesn't that mean that those strobes are effectively A0?

    For a 16 bit bus it only matters a little.
    You can have an 16-bit aligned address and separate enables
    for each byte (the Motorola 68000 way),
    or a byte address and a wire High Byte Enable indicating
    whether the high byte is included/excluded, which you then
    have to decode with extra logic to get separate enables
    (the Intel 8086 way).

    For wider buses, say 32 bits, I'd want a 30 bit,
    4-byte aligned address bus with 4 separate wires indicating
    which bytes are in the read/write operation as you can
    run those directly out onto the bus to the memory or device,
    Also works well if you want to include byte parity on the bus.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 3 21:08:52 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-02 18:59, Paul Clayton wrote:

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    In those prior "24-bit address space" designs the MMU didn't
    look at the upper 8 address bits. That is why they had later
    software porting problems.

    In the current "48-bit address space" MMU's, all 64 address bits are validated.
    The 48-bits is a MMU model specfic VA sub-space translation limit
    set by the number of levels in the page table.

    I believe it is set by the used width of PA<63..12> in the PTE in
    conjunction to the number of address bits that are routed around
    the on-die interconnect. Using PA<63..48> as other than 12b'0
    is asking for memory aliasing problem in the coherence protocol.

    On the VA side, for future compatibility the topmost defined bit is
    extended into the unused bits (to wit, sign extended to 64-bits).

    On the PA side, unused bits of the PA will be tied to a logical 0
    (there is no need for unused wires in the design).

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 4 01:11:38 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-02 18:59, Paul Clayton wrote:

    With many 64-bit systems limiting user space pointers to 47 bits
    (one bit of the 48-bit address space differentiating system and
    user space), it seems sad that programs would have to explicitly
    mask/extract small pointers to use those extra bits.

    In those prior "24-bit address space" designs the MMU didn't
    look at the upper 8 address bits. That is why they had later
    software porting problems.

    In the current "48-bit address space" MMU's, all 64 address bits are validated.
    The 48-bits is a MMU model specfic VA sub-space translation limit
    set by the number of levels in the page table.

    I believe it is set by the used width of PA<63..12> in the PTE in >conjunction to the number of address bits that are routed around
    the on-die interconnect. Using PA<63..48> as other than 12b'0
    is asking for memory aliasing problem in the coherence protocol.

    On the VA side, for future compatibility the topmost defined bit is
    extended into the unused bits (to wit, sign extended to 64-bits).

    AMD calls this feature "canonical".

    On the PA side, unused bits of the PA will be tied to a logical 0
    (there is no need for unused wires in the design).

    Unused PRE.PA bits are checked for canonicality. Then discarded.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 5 01:08:52 2026
    From Newsgroup: comp.arch


    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sat Sep 5 04:26:41 2026
    From Newsgroup: comp.arch

    On Thu, 3 Sep 2026 02:46:48 -0000 (UTC), I wrote:

    And also the addressing was kept to 26 bits. Why? So that the
    complete CPU state on an interrupt could be stored in 32 bits.

    Of course I meant rCLCPU state apart from programmer-visible registersrCY
    ...
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From mas@mas@a4.home to comp.arch on Sun Sep 6 03:56:36 2026
    From Newsgroup: comp.arch

    On 2026-09-05, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    Take a look at:

    https://github.com/pizlonator/fil-c

    It's software but has performance issues -- hardware assistance might
    make a big difference:

    Fil-C is a fanatically compatible memory-safe implementation of C and
    C++. Lots of software compiles and runs with Fil-C with zero or minimal changes. All memory safety errors are caught as Fil-C panics. Fil-C
    achieves this using a combination of concurrent garbage collection
    and invisible capabilities (each pointer in memory has a corresponding capability, not visible to the C address space). Every fundamental C
    operation (as seen in LLVM IR) is checked against the capability. Fil-C
    has no unsafe statement and only limited FFI to unsafe code.

    Fil-C is special because:

    Fil-C achieves full safety with no escape hatches. There is no unsafe
    keyword in Fil-C that could be used to turn off protections. Linking
    to unsafe code is severely restricted.

    Fil-C's capability-based approach achieves a similar level of safety to
    hardware capabilities like CHERI, except that it runs on stock hardware
    (X86_64 or ARM64).

    Fil-C is engineered to prevent memory safety bugs from being used for
    exploitation rather than just simply flagging them often enough to
    find bugs. This makes Fil-C different from AddressSanitizer, HWAsan,
    or MTE, which can all be bypassed by attackers. The key difference that
    makes this possible is that Fil-C is capability based (so each pointer
    knows what range of memory it may access, and how it may access it)
    rather than tag based (where pointer accesses are allowed if they hit
    valid memory).

    From a language user standpoint, Fil-C is just C and C++ with GCC/clang
    extensions. It's more likely than not that your favorite C or C++
    program or library compiles in Fil-C with zero changes. The Fil-C
    compiler is based on clang 20.1.8, so it supports C17 and C++20.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Sun Sep 6 13:51:23 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    It's useful to start with an existing experimental project and
    look at the current status thereof:

    https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Sun Sep 6 10:44:28 2026
    From Newsgroup: comp.arch

    On 2026-Sep-04 21:08, MitchAlsup wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    My knee jerk reaction is 'none' - those are language issues.

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    That is normally left to the high level language because each
    language has it own set rules, and those rules can change.
    For example, in Ada a routine argument marked as 'out' is considered
    an uninitialized variable that must be written before the routine returns,
    and must be written before it is read.
    Ada85 originally did not allow 'out' arguments to be read after writing
    as 'out' meant write-only. But that was a stupid restriction which was just inconvenient so they later changed it to read-after-write.

    Do you really want to incorporate that into hardware?

    And if a pointer has a bit marking it as read-only then what stops
    me clearing that bit? Unless you make "pointer" a first class HW type
    with its own set of instructions.
    But then you also need the ability to bypass any "pointer" restrictions
    to handle things like Anton's software defined tagged integer-pointers.

    This is the problem with capabilities systems - they balloon very quickly.
    As HW tries to take on ALL the capabilities different languages might
    ever require, the generic nature of them drags in inefficiencies and
    they wind up as "a jack of all trades but a master of none".

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    And what is in that hidden indexed capability table?
    To be flexible enough to handle any data type it would be
    a pointer to a privileged routine. And there is the ballooning.
    Microcode by any other name would smell as sweet.

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.
    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.
    Then see what's left to address.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 6 14:56:29 2026
    From Newsgroup: comp.arch

    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 15:56:22 2026
    From Newsgroup: comp.arch


    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-04 21:08, MitchAlsup wrote:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    My knee jerk reaction is 'none' - those are language issues.

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    That is normally left to the high level language because each
    language has it own set rules, and those rules can change.
    For example, in Ada a routine argument marked as 'out' is considered
    an uninitialized variable that must be written before the routine returns, and must be written before it is read.
    Ada85 originally did not allow 'out' arguments to be read after writing
    as 'out' meant write-only. But that was a stupid restriction which was just inconvenient so they later changed it to read-after-write.

    Do you really want to incorporate that into hardware?

    The thought was to "const *pointer" as an argument to a subroutine,
    so, the compiler could pass a pointer such that STs will fault and
    do it from the calling side even when the PTE allows writes.

    And if a pointer has a bit marking it as read-only then what stops
    me clearing that bit? Unless you make "pointer" a first class HW type
    with its own set of instructions.

    Don't want to go "that far".

    But then you also need the ability to bypass any "pointer" restrictions
    to handle things like Anton's software defined tagged integer-pointers.

    This is the problem with capabilities systems - they balloon very quickly.
    As HW tries to take on ALL the capabilities different languages might
    ever require, the generic nature of them drags in inefficiencies and
    they wind up as "a jack of all trades but a master of none".

    So I have seen.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    And what is in that hidden indexed capability table?
    To be flexible enough to handle any data type it would be
    a pointer to a privileged routine. And there is the ballooning.
    Microcode by any other name would smell as sweet.

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.
    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks. Then see what's left to address.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Sun Sep 6 12:58:53 2026
    From Newsgroup: comp.arch

    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.
    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.
    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    No, there is no difference in register allocation.
    Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap
    instead of the unchecked ADD.

    MULS or MULU return a double wide register pair, and the high word then
    checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
    So 1 extra instruction to check each multiply for overflow.
    Or one can incorporate the overflow check into the
    MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
    and return a single wide result.

    Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm. The index register is then used with a scaled-index addressing.

    For almost all array indexes the cost is a single reg-reg or reg-imm instruction.

    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    I have difficulty judging use-after-free errors.
    I don't recall ever having had a use-after-free error ever since
    I started writing C in 1992 when I switched to developing on WinNT.

    But I'm fairly paranoid so I do things like having my own Assert()
    routines that throw my own fatal exceptions on errors and remain
    in production code, put validity check markers on all heap
    allocated objects and assert they are valid before using,
    on free zap the validity markers and NULL the pointer.
    Consequently if such programming error did happen, my code should
    immediately detect it itself and throw a fatal exception.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Niklas Holsti@niklas.holsti@tidorum.invalid to comp.arch on Sun Sep 6 20:21:04 2026
    From Newsgroup: comp.arch

    On 2026-09-06 17:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    As described by mas@a4.home in their earlier post, Fil-C does seem able
    to solve much of the address-arithmetic and use-after-free problems:

    "All memory safety errors are caught as Fil-C panics. Fil-C
    achieves this using a combination of concurrent garbage collection
    and invisible capabilities (each pointer in memory has a corresponding capability, not visible to the C address space)."

    So capabilities are protected from user modification, and "free" is
    probably a no-op because of garbage collection. The remaining question
    is performance, but many applications are not performance-sensitive.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Sun Sep 6 12:04:52 2026
    From Newsgroup: comp.arch

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then
    you don't need two types of pointers.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 6 20:25:46 2026
    From Newsgroup: comp.arch

    EricP <ThatWouldBeTelling@thevillage.com> schrieb:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.
    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    There are programming languages which do that, Fortran being
    a prime example.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Fortran has an interesting take on pointers: You can associate a
    pointer with existing variables, but only if they have the TARGET
    attribute, otherwise this must be diagnosed (it's a constraint).
    You can also have multi-dimensional pointers, so something like

    real, dimension(10,10), target :: a
    real, dimension(:,:), pointer :: ap

    ap => a(1:10:2,1:10:2)

    After this, ap(1,2) is an alias to a(1,3), for example. ap is
    then pointing to a sub-array. This will be handled with array
    descriptors aka dope vectors.

    No address arithmetic required, normal indexing will do.

    You can also ALLOCATE pointers, after which they are associated with
    an anonymous region. But this is usually not the best way because
    Fortran also has ALLOCATABLE variables which are automatically
    deallocated when they go out of scope (or are passed as actual
    argument to an INTENT(OUT) dummy argument).

    But these mechanisms are, by nature, rather heavy-weight and
    not as well suited to a lower-level language like C.


    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks. >>
    You would probably need more than 32 registers for this...

    No, there is no difference in register allocation.
    Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.

    Assume the following C code:

    #define N something

    ...
    double *p = calloc(N, sizeof(double));

    ...

    foo (p+2,i);

    void foo(double *p, int i)
    {
    p[i] = 42.
    }

    This would require something like, generated from the code
    above,

    void foo(double *p, int i, double *p_from, double *p_to)
    {
    if (p + i < p_from)
    range_error();
    if (p + i > p_to)
    range_error();
    p[i] = 42.;
    }

    so each range-checked pointer would require three actual arguments.

    Hardware could assist this with an instruction "trap if out of range",
    which would combine the if statements above into one. This would be
    an instruction which could be predicted to be a no-op with some
    confidence.

    But register (or argument) pressure would be high, which is why
    I suspect that 32 registers might not be enough.

    MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
    So 1 extra instruction to check each multiply for overflow.
    Or one can incorporate the overflow check into the
    MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
    and return a single wide result.

    Or Mitch's carry instruction, followed by a bne0.

    Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm. The index register is then used with a scaled-index addressing.

    That would require both limits.

    For almost all array indexes the cost is a single reg-reg or reg-imm instruction.

    ... which should be moved outside the loops.


    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    I have difficulty judging use-after-free errors.

    It seems to be quire common, see https://cwe.mitre.org/data/definitions/416.html .

    I don't recall ever having had a use-after-free error ever since
    I started writing C in 1992 when I switched to developing on WinNT.

    But I'm fairly paranoid so I do things like having my own Assert()
    routines that throw my own fatal exceptions on errors and remain
    in production code, put validity check markers on all heap
    allocated objects and assert they are valid before using,
    on free zap the validity markers and NULL the pointer.
    Consequently if such programming error did happen, my code should
    immediately detect it itself and throw a fatal exception.

    That sounds like very good practice.

    One reason why I like ALLOCATABLE variables in Fortran is that
    a large fraction of that burden is taken off the programmer's
    shoulders.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sun Sep 6 21:46:15 2026
    From Newsgroup: comp.arch

    On Sun, 6 Sep 2026 10:44:28 -0400, EricP wrote:

    If programmers just used languages that checked array indexes then
    99.999% of memory access errors would disappear. First make array
    index checks simple and cheap. This requires checked arithmetic for
    the index expression calculation, plus a set of simple
    compare-and-fault instructions various bounds checks.

    Pascal introduced the concept of subrange types, where a fixed range
    of integer values (e.g. one just happening to match the bounds of an
    array type in the program) can be used as the type of a variable in
    its own right. Ada carries over this concept.

    I remember some research showing that, with proper use, this helped to
    reduce the total overhead of array bounds checking down to just a few
    percent.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 23:46:51 2026
    From Newsgroup: comp.arch


    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Some of us are.....

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    I proposed a few years back to allow compilers to use one kind
    of checking with p[i] and another when using *(p+i); so long as
    in that scope, there are no p++ calculations.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    THat is nadda gonna fly.

    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    No, there is no difference in register allocation.

    I agree, since 96% of subroutines* use fewer than 32 registers anyway.

    (*) LLVM source and all the source code I can find.

    Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.

    My 66000 already uses the sign bit of pointers and compares it to the
    sign bit of VA and raises a fault when the 63-bit VAS boundary is stepped over. Its not enough, but better than ignoring the problem.

    MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.

    My 66000 does this and takes the exception when IOverflow is enabled.

    So 1 extra instruction to check each multiply for overflow.

    0 more instructions when done right.

    Or one can incorporate the overflow check into the
    MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
    and return a single wide result.

    Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm. The index register is then used with a scaled-index addressing.

    (index - lower_bound) > 0
    is
    ADD Ri,Rindex,-Rlob
    BLE Ri,.....

    For almost all array indexes the cost is a single reg-reg or reg-imm instruction

    Low cost when the lower bound is 0 (or 1).

    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    I have difficulty judging use-after-free errors.

    My programming style does not use free() in any sense that would require garbage collection. So, this problem does not exist in my code.

    I don't recall ever having had a use-after-free error ever since
    I started writing C in 1992 when I switched to developing on WinNT.

    Occasionally, when I need to free() and the afterwards malloc() again,
    I will fork off a task, have it do its think, return its result, and
    the exit(0) saving all the GC code.

    But I'm fairly paranoid so I do things like having my own Assert()
    routines that throw my own fatal exceptions on errors and remain
    in production code, put validity check markers on all heap
    allocated objects and assert they are valid before using,
    on free zap the validity markers and NULL the pointer.

    All good.

    Consequently if such programming error did happen, my code should
    immediately detect it itself and throw a fatal exception.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 6 23:59:58 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then
    you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 00:02:42 2026
    From Newsgroup: comp.arch


    Thomas Koenig <tkoenig@netcologne.de> posted:

    EricP <ThatWouldBeTelling@thevillage.com> schrieb:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.
    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    There are programming languages which do that, Fortran being
    a prime example.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Fortran has an interesting take on pointers: You can associate a
    pointer with existing variables, but only if they have the TARGET
    attribute, otherwise this must be diagnosed (it's a constraint).
    You can also have multi-dimensional pointers, so something like

    real, dimension(10,10), target :: a
    real, dimension(:,:), pointer :: ap

    ap => a(1:10:2,1:10:2)

    After this, ap(1,2) is an alias to a(1,3), for example. ap is
    then pointing to a sub-array. This will be handled with array
    descriptors aka dope vectors.

    No address arithmetic required, normal indexing will do.

    You can also ALLOCATE pointers, after which they are associated with
    an anonymous region. But this is usually not the best way because
    Fortran also has ALLOCATABLE variables which are automatically
    deallocated when they go out of scope (or are passed as actual
    argument to an INTENT(OUT) dummy argument).

    But these mechanisms are, by nature, rather heavy-weight and
    not as well suited to a lower-level language like C.

    Perhaps the fault is with C than with Fortran !!!


    First make array index checks simple and cheap.

    And OBVIOUS !

    This requires checked arithmetic for the index expression calculation, >>> plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    No, there is no difference in register allocation.
    Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.

    Assume the following C code:

    #define N something

    ...
    double *p = calloc(N, sizeof(double));

    ...

    foo (p+2,i);

    void foo(double *p, int i)
    {
    p[i] = 42.
    }

    This would require something like, generated from the code
    above,

    void foo(double *p, int i, double *p_from, double *p_to)
    {
    if (p + i < p_from)
    range_error();
    if (p + i > p_to)
    range_error();
    p[i] = 42.;
    }

    so each range-checked pointer would require three actual arguments.

    Hardware could assist this with an instruction "trap if out of range",
    which would combine the if statements above into one. This would be
    an instruction which could be predicted to be a no-op with some
    confidence.

    But register (or argument) pressure would be high, which is why
    I suspect that 32 registers might not be enough.

    MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
    So 1 extra instruction to check each multiply for overflow.
    Or one can incorporate the overflow check into the
    MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
    and return a single wide result.

    Or Mitch's carry instruction, followed by a bne0.

    Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm.
    The index register is then used with a scaled-index addressing.

    That would require both limits.

    For almost all array indexes the cost is a single reg-reg or reg-imm instruction.

    ... which should be moved outside the loops.


    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    I have difficulty judging use-after-free errors.

    It seems to be quire common, see https://cwe.mitre.org/data/definitions/416.html .

    I don't recall ever having had a use-after-free error ever since
    I started writing C in 1992 when I switched to developing on WinNT.

    But I'm fairly paranoid so I do things like having my own Assert()
    routines that throw my own fatal exceptions on errors and remain
    in production code, put validity check markers on all heap
    allocated objects and assert they are valid before using,
    on free zap the validity markers and NULL the pointer.
    Consequently if such programming error did happen, my code should immediately detect it itself and throw a fatal exception.

    That sounds like very good practice.

    One reason why I like ALLOCATABLE variables in Fortran is that
    a large fraction of that burden is taken off the programmer's
    shoulders.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 00:09:11 2026
    From Newsgroup: comp.arch


    Thomas Koenig <tkoenig@netcologne.de> posted:

    EricP <ThatWouldBeTelling@thevillage.com> schrieb:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.
    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    There are programming languages which do that, Fortran being
    a prime example.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Fortran has an interesting take on pointers: You can associate a
    pointer with existing variables, but only if they have the TARGET
    attribute, otherwise this must be diagnosed (it's a constraint).
    You can also have multi-dimensional pointers, so something like

    real, dimension(10,10), target :: a
    real, dimension(:,:), pointer :: ap

    ap => a(1:10:2,1:10:2)

    After this, ap(1,2) is an alias to a(1,3), for example. ap is
    then pointing to a sub-array. This will be handled with array
    descriptors aka dope vectors.

    No address arithmetic required, normal indexing will do.

    You can also ALLOCATE pointers, after which they are associated with
    an anonymous region. But this is usually not the best way because
    Fortran also has ALLOCATABLE variables which are automatically
    deallocated when they go out of scope (or are passed as actual
    argument to an INTENT(OUT) dummy argument).

    But these mechanisms are, by nature, rather heavy-weight and
    not as well suited to a lower-level language like C.

    Perhaps the fault is with C than with Fortran !!!


    First make array index checks simple and cheap.

    And OBVIOUS !

    This requires checked arithmetic for the index expression calculation, >>> plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    No, there is no difference in register allocation.
    Use an ADDFS Add Fault Signed Overflow or ADDFU Add Fault Unsigned wrap instead of the unchecked ADD.

    Assume the following C code:

    #define N something

    ...
    double *p = calloc(N, sizeof(double));

    ...

    foo (p+2,i);

    void foo(double *p, int i)
    {
    p[i] = 42.
    }

    This would require something like, generated from the code
    above,

    void foo(double *p, int i, double *p_from, double *p_to)
    {
    if (p + i < p_from)
    range_error();
    if (p + i > p_to)
    range_error();
    p[i] = 42.;
    }

    Whereas::

    double p[N] = calloc(N, sizeof(double));
    ...
    foo (&p[2],i);
    ...
    void foo(double p[*], int i)
    {
    p[i] = 42.
    }

    does not!!

    so each range-checked pointer would require three actual arguments.

    Hardware could assist this with an instruction "trap if out of range",
    which would combine the if statements above into one. This would be
    an instruction which could be predicted to be a no-op with some
    confidence.

    But register (or argument) pressure would be high, which is why
    I suspect that 32 registers might not be enough.

    MULS or MULU return a double wide register pair, and the high word then checked if != 0 to detect expression overflow - FLTNZ Fault if reg != 0.
    So 1 extra instruction to check each multiply for overflow.
    Or one can incorporate the overflow check into the
    MULFS Multiply Fault Signed or MULFU Fault Unsigned overflow
    and return a single wide result.

    Or Mitch's carry instruction, followed by a bne0.

    Mitches MUL checks Overflow at either 64-bit boundaries when enabled. So,
    all we need is BNE0

    Then a check the index register is < limit, FLTGE Fault if reg >= reg or imm.
    The index register is then used with a scaled-index addressing.

    That would require both limits.

    CMP Rd,Rnormalixed_index,array_limit
    BCIN Rd,where_ever

    Does both the <0 and >=array_limit checks in one inst.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Mon Sep 7 04:29:46 2026
    From Newsgroup: comp.arch

    On Sun, 06 Sep 2026 23:59:58 GMT, MitchAlsup wrote:

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an
    array.

    Dang. There goes all that old PDP-11 code. ;)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 7 06:17:35 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    Thomas Koenig <tkoenig@netcologne.de> posted:

    But these mechanisms are, by nature, rather heavy-weight and
    not as well suited to a lower-level language like C.

    Perhaps the fault is with C than with Fortran !!!

    I would tend to concur, but of course I am biased in
    favor of Fortran :-)

    Assume the following C code:

    #define N something

    ...
    double *p = calloc(N, sizeof(double));

    ...

    foo (p+2,i);

    void foo(double *p, int i)
    {
    p[i] = 42.
    }

    This would require something like, generated from the code
    above,

    void foo(double *p, int i, double *p_from, double *p_to)
    {
    if (p + i < p_from)
    range_error();
    if (p + i > p_to)
    range_error();
    p[i] = 42.;
    }

    Whereas::

    double p[N] = calloc(N, sizeof(double));

    Either double p[N] or double *p = ...

    ...
    foo (&p[2],i);
    ...
    void foo(double p[*], int i)
    {
    p[i] = 42.
    }

    does not!!

    This would still be legal for

    foo (&p[2],-2)

    so both bounds checks would be needed to avoid a false positive.

    But, looking back at Fortran, there are also the case of
    lower bounds unequal to one, like


    subroutine foo(a,m,n)
    real, dimension(3:n)

    a(m) = 42.

    so a check would require that m be between 3 and n.

    [...]

    That would require both limits.

    CMP Rd,Rnormalixed_index,array_limit
    BCIN Rd,where_ever

    Does both the <0 and >=array_limit checks in one inst.

    Both are needed, unfortunately.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 7 10:40:36 2026
    From Newsgroup: comp.arch

    On 07/09/2026 01:59, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then
    you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.

    That is up to the compiler, and dependent on types and additional
    surrounding code. For "typical" usage - "i" and "p" as local variables,
    and "i" being either a signed integer type or a size_t (i.e., not a
    32-bit unsigned int on a 64-bit machine), and using an optimising
    compiler, I would not expect any difference in the code.

    But if you can find a more complete example that can be tested on
    godbolt, it would be very interesting.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Mon Sep 7 13:28:23 2026
    From Newsgroup: comp.arch

    On 2026-Sep-06 15:04, Stephen Fuld wrote:
    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array references to use the existing subscript mechanism, be sufficient?-a Then you don't need two types of pointers.

    I was thinking of something that programmers could easily
    adopt which would assist them in detecting bugs.
    Forcing them to rewrite their code to add a level of indirection
    through integer indexes I think would pretty much guarantee
    it would not be used.

    The kind of thing I had in mind might be if "in" was a reserved word
    so one could write

    {
    int vec1[32], vec2[64],
    *ptr1a in vec1 = vec1[0],
    *ptr1b in vec1 = vec1[0],
    *ptr2 in vec2 = vec2[0];

    would create three int pointers with distinct data types
    different from *int and each other, such that assigning between
    ptr1a or ptr1b and ptr2 or *int causes a type mismatch warning,
    but assignment between ptr1a and ptr1b is allowed because
    they both reference inside the same object vec1.

    Bounded pointers ptr1a, ptr1b, ptr2 can be manipulated with
    arithmetic expressions whereas object pointers cannot.

    At runtime when a bounded pointer is dereferenced the address is
    bounds checked against its array buffer lower and upper address limits.

    Arrays are first class types and carry their element count
    with them when passed as arguments as an invisible extra argument.
    Zero element count or null arrays are possible.
    A null array does not necessarily have an valid buffer address
    and the element count must be tested first.

    An array slice range is indicated with an ellipsis array[low_exp...upr_exp]
    and zero/null slices of arrays are possible when the slice
    lower bound > upper bound.






    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 18:25:54 2026
    From Newsgroup: comp.arch


    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then >> you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.

    That is up to the compiler, and dependent on types and additional surrounding code. For "typical" usage - "i" and "p" as local variables,
    and "i" being either a signed integer type or a size_t (i.e., not a
    32-bit unsigned int on a 64-bit machine), and using an optimising
    compiler, I would not expect any difference in the code.

    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
    of strength reduction to figure out that only 1 ADD is needed.

    But if you can find a more complete example that can be tested on
    godbolt, it would be very interesting.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 18:33:12 2026
    From Newsgroup: comp.arch


    ASIDs came along in the late 1990s and early 2000s to better optimize
    MMU resource consumption {fewer flushes, support nested paging, ...}.

    With their advent, ASIDs were assigned to threads/processes.

    Has anyone ever wanted to allow shared ASIDs in such a way that shared
    memory or shared files have the ASID associated with the memory/file ?
    So that multiple processes accessing the same shared resource co-optim-
    ize themselves across cores and caches ??
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 7 23:10:19 2026
    From Newsgroup: comp.arch

    On 07/09/2026 20:25, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address
    arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then >>>> you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array. >>
    That is up to the compiler, and dependent on types and additional
    surrounding code. For "typical" usage - "i" and "p" as local variables,
    and "i" being either a signed integer type or a size_t (i.e., not a
    32-bit unsigned int on a 64-bit machine), and using an optimising
    compiler, I would not expect any difference in the code.

    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
    of strength reduction to figure out that only 1 ADD is needed.


    The first one has two adds too - i++, and (p + i) for the array access.
    (I'm assuming that there is more going on inside the loop in real code,
    such as at least reading or writing from the array. Otherwise the
    compiler can see that the whole thing is doing nothing and skip it.)

    In both cases, optimisers are likely to generate code approximating :

    q = &p[max];
    while (p < q) {
    p++;
    }

    (Again, I assume there's a read or write to be included inside the loop.)

    Sometimes pointer / array accesses with increment cannot be optimised as
    well if the index is an unsigned type smaller than size_t, because the compiler can't rule out the possibility of the index wrapping. This is
    not an issue with signed integer types, or a big enough unsigned type,
    nor is it an issue here when the limits of the index are known. And it
    is independent of the ISA.


    But if you can find a more complete example that can be tested on
    godbolt, it would be very interesting.


    Again, if you can give a more complete example, it would be easier to
    see what you are getting at here, because I cannot yet see your point.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 7 21:58:43 2026
    From Newsgroup: comp.arch


    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 20:25, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address >>>>>> arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array
    (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then >>>> you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.

    That is up to the compiler, and dependent on types and additional
    surrounding code. For "typical" usage - "i" and "p" as local variables, >> and "i" being either a signed integer type or a size_t (i.e., not a
    32-bit unsigned int on a 64-bit machine), and using an optimising
    compiler, I would not expect any difference in the code.

    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
    of strength reduction to figure out that only 1 ADD is needed.


    The first one has two adds too - i++, and (p + i) for the array access.

    The p+i addition is part of AGEN and thus free for any ISA that has at
    least [Rpointer+Rindex] addressing mode (everybody but RISC-V).

    (I'm assuming that there is more going on inside the loop in real code,
    such as at least reading or writing from the array. Otherwise the
    compiler can see that the whole thing is doing nothing and skip it.)

    In both cases, optimisers are likely to generate code approximating :

    q = &p[max];
    while (p < q) {
    p++;
    }

    (Again, I assume there's a read or write to be included inside the loop.)

    Sometimes pointer / array accesses with increment cannot be optimised as well if the index is an unsigned type smaller than size_t, because the compiler can't rule out the possibility of the index wrapping. This is
    not an issue with signed integer types, or a big enough unsigned type,
    nor is it an issue here when the limits of the index are known. And it
    is independent of the ISA.


    But if you can find a more complete example that can be tested on
    godbolt, it would be very interesting.


    Again, if you can give a more complete example, it would be easier to
    see what you are getting at here, because I cannot yet see your point.

    for( i = 0; i < max; i++ )
    a[i] = b[i] + c[i];

    versus

    for( i = 0; i < max; i++ )
    *a++ = *b++ + *c++;

    The former has 1 add per iteration, the later has 4.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Josh Vanderhoof@x@y.z to comp.arch on Mon Sep 7 18:22:59 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array
    indexes then 99.999% of memory access errors would
    disappear.

    Retrofitting memory safety onto C is an uphill battle.-a
    Address arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards
    compatible.

    Also agreed.

    Have different kinds of pointers: object pointers that
    may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular
    array (so they can be bounds checked) and do allow pointer
    arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing
    array references to use the existing subscript mechanism, be
    sufficient? Then you don't need two types of pointers.
    On modern RISC ISAs:
    for( i = 0; i < max; i++ )
    p[i]
    is often faster than:
    for( i = 0; i < max; i++ )
    *p++
    Especially when there are more than 1 structure being
    accessed as an array.
    That is up to the compiler, and dependent on types and
    additional surrounding code. For "typical" usage - "i" and
    "p" as local variables, and "i" being either a signed integer
    type or a size_t (i.e., not a 32-bit unsigned int on a 64-bit
    machine), and using an optimising compiler, I would not expect
    any difference in the code.

    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes
    a lot of strength reduction to figure out that only 1 ADD is
    needed.

    It's identical for a good compiler.
    Or even compilers that emit "mov esi, esi" ;)

    x86-64 gcc 16.2 -O3

    void clear_array(int *p, int max) {
    int i; for (i = 0; i < max; i++)
    p[i] = 0;
    } void clear_postinc(int *p, int max) {
    int i; for (i = 0; i < max; i++)
    *p++ = 0;
    }

    "clear_array(int*, int)":
    test esi, esi jle .L1 mov esi, esi lea rdx,
    [0+rsi*4] xor esi, esi jmp "memset"
    .L1:
    ret
    "clear_postinc(int*, int)":
    test esi, esi jle .L4 mov esi, esi lea rdx,
    [0+rsi*4] xor esi, esi jmp "memset"
    .L4:
    ret
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 7 17:15:55 2026
    From Newsgroup: comp.arch

    On 9/1/2026 1:04 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 8/31/2026 9:35 AM, John Dallman wrote:
    I've been reading about various early architectures, and realised that
    System/360 seems to have been innovative in a way that isn't talked about >>> very much.

    The earliest style seems to have been one-word instructions, each
    containing an opcode, modifier bits and an address. That was fine while
    memories were small, but became increasingly limiting as they grew.
    18-bit addressing was common on 36-bit systems. Even the 64-bit Stretch
    was limited to 18 address bits.

    The 360 introduced word-sized address registers, which meant that address >>> space became a key feature of a computer. It had, theoretically, 32-bit
    addresses, although only 24 bits were implemented at first. Address
    registers could be loaded, using instructions longer than one word, but
    were usually used for base addresses, and not changed very often.

    Was this original with the 360, or did someone else invent it first? Were >>> there other styles, now abandoned?

    Well, the Univac 1107 (1962) was a 36 bit word oriented system with up
    to 64K word addressing. Its 36 bit instructions contained a 16 bit
    address field. When the loosely upward compatible 1108 came out in
    1964, maximum memory increased to 256K words. Addresses above 64K were
    accessed by loading a value into the low order 18 bits of a 36 bit index
    register, whose contents were added (by the hardware) to the address in
    the instruction.

    Several 36-bit machines provided 18-bit displacements with various
    kinds of base and index registers.

    Yes. But note that the 1107 had a 16, not 18 bit displacement. This
    was fine as the 1107 had a 64Kword maximum memory. But when the 1108 increased maximum memory to 256Kwords, references to memory above 64K
    required the use of an index register.

    BTW, I never claimed that the 1107 was unique, or even the first, just
    that it predated the S/360. It also predated the PDP6. And it also had indirect addressing via a bit in the instruction word.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 8 05:30:32 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    David Brown <david.brown@hesbynett.no> posted:

    Again, if you can give a more complete example, it would be easier to
    see what you are getting at here, because I cannot yet see your point.

    for( i = 0; i < max; i++ )
    a[i] = b[i] + c[i];

    versus

    for( i = 0; i < max; i++ )
    *a++ = *b++ + *c++;

    The former has 1 add per iteration, the later has 4.

    https://godbolt.org/z/6h84fzYc1 shows that the generated code
    is identical. A lot of work has gone into optimizing strength
    reduction.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 8 08:28:53 2026
    From Newsgroup: comp.arch

    John Levine <johnl@taugh.com> writes:
    They added virtual memory similar
    to the 67's to S/370 in 1972 and soon made all the major operating
    systems use it.

    Given VM, why was there a need to add virtual memory to the guest OSs?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 8 08:34:36 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    In practice, I never found it to be a problem as
    indirection was rare and I never saw more than two levels.

    Unification of a pre-existing free logical variable V with something
    (X) is implemented as making V point to X. If X is another logical
    variable, and then is unified with something, say Y, a pointer to Y is
    stored in the memory location for X. And so on. Accessing V may need
    to follow an arbitrarily long chain of pointers until you find either
    a free variable, or you find the value that V eventually was unified
    with.

    Prolog was implemented in Edinburgh on the DEC-10 (DEC-10 Prolog). I
    guess that they used the indirection feature for that: If a variable
    points to some other variable or value, its indirection bit is set, if
    it is free, it is not.

    On modern machines, i.e., without indirection bit, a free variable
    points to itself, a variable that is bound to another variable points
    to that other variable. While the type tag indicates a variable, one
    follows the pointers, until a pointer to itself is found.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Tue Sep 8 11:20:23 2026
    From Newsgroup: comp.arch

    On 07/09/2026 23:58, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 20:25, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/6/2026 9:58 AM, EricP wrote:
    On 2026-Sep-06 10:56, Thomas Koenig wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes >>>>>>>>> then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle.-a Address >>>>>>>> arithmetic stands in the way of that.

    It might be possible to do but no one is interested.

    Agreed.

    Making arrays a first class type and distinct from pointers
    would be the first step but would not be backwards compatible.

    Also agreed.

    Have different kinds of pointers: object pointers
    that may/may-not be NULL but don't allow pointer arithmetic,
    array element pointers that only point inside a particular array >>>>>>> (so they can be bounds checked) and do allow pointer arithmetic.

    Would just disallowing arithmetic on pointers, thus forcing array
    references to use the existing subscript mechanism, be sufficient? Then >>>>>> you don't need two types of pointers.

    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.

    That is up to the compiler, and dependent on types and additional
    surrounding code. For "typical" usage - "i" and "p" as local variables, >>>> and "i" being either a signed integer type or a size_t (i.e., not a
    32-bit unsigned int on a 64-bit machine), and using an optimising
    compiler, I would not expect any difference in the code.

    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
    of strength reduction to figure out that only 1 ADD is needed.


    The first one has two adds too - i++, and (p + i) for the array access.

    The p+i addition is part of AGEN and thus free for any ISA that has at
    least [Rpointer+Rindex] addressing mode (everybody but RISC-V).

    (I'm assuming that there is more going on inside the loop in real code,
    such as at least reading or writing from the array. Otherwise the
    compiler can see that the whole thing is doing nothing and skip it.)

    In both cases, optimisers are likely to generate code approximating :

    q = &p[max];
    while (p < q) {
    p++;
    }

    (Again, I assume there's a read or write to be included inside the loop.)

    Sometimes pointer / array accesses with increment cannot be optimised as
    well if the index is an unsigned type smaller than size_t, because the
    compiler can't rule out the possibility of the index wrapping. This is
    not an issue with signed integer types, or a big enough unsigned type,
    nor is it an issue here when the limits of the index are known. And it
    is independent of the ISA.


    But if you can find a more complete example that can be tested on
    godbolt, it would be very interesting.


    Again, if you can give a more complete example, it would be easier to
    see what you are getting at here, because I cannot yet see your point.

    for( i = 0; i < max; i++ )
    a[i] = b[i] + c[i];

    versus

    for( i = 0; i < max; i++ )
    *a++ = *b++ + *c++;

    The former has 1 add per iteration, the later has 4.


    enum { max = 100 };

    extern int a[max];
    extern int b[max];
    extern int c[max];

    void foo(void) {
    for (int i = 0; i < max; i++) {
    a[i] = b[i] + c[i];
    }
    }

    void bar(void) {
    int * p = a;
    int * q = b;
    int * r = c;
    for (int i = 0; i < max; i++) {
    *p++ = *q++ + *r++;
    }
    }


    I tried that with a half-dozen gcc targets on godbolt, all with -O2.
    ARM 32-bit ARM 64-bit, x86 64-bit, RISC-V, Power. I also tried clang
    and MSVC with various targets. Sometimes the clang code was too
    unrolled to compare sensibly, but I could find no cases where the code
    for "foo" and "bar" was different in any significant way. There were
    minor changes to the ordering of the instructions or the choice of
    registers, but otherwise the code was identical.

    Compilers are smart enough to know that the two formulations do the same thing. They have been able to do so for decades.

    On x86, both functions use addressing modes that match "a + j", "b + j",
    "c + j", and increment "j" by a stride value until it reaches 400 (so
    "j" is "i * 4"). The stride depends on the SIMD vector width used, and
    any unrolling.

    On RISC cpus, both functions use addressing modes that match "*p++" -
    where the post-increment is included in the instruction. No counter
    variable is used at all, and the loop stops by comparing one of the
    pointers to the end of the array.


    Even if you have a RISC processor that does not have a "post-increment" addressing operation, optimising compilers would still generate
    approximately the same code for each formulation here.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Tue Sep 8 13:31:12 2026
    From Newsgroup: comp.arch

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    John Levine <johnl@taugh.com> writes:
    They added virtual memory similar
    to the 67's to S/370 in 1972 and soon made all the major operating
    systems use it.

    Given VM, why was there a need to add virtual memory to the guest OSs?

    They added virtual memory hardware, page maps and page faults. All the operating systems use that, and can run either on bare hardware or
    under VM/370 or these days in LPARs which are sort of VM in microcode.

    They use virtual memory for all the usual reasons, give each process its
    own address space. Lynn has told us that MVS reserved so much of the
    address space for the operating system that they needed separate address
    spaces to make a reasonable number of programs fit. There was 24 bit
    virtual memory for a while before they added 31 bit mode.

    Early on there was a version of SVS that did a handshake with VM/370 so
    it ran in a single 16M virtual address space, which was huge at the time, and VM
    reflected the page faults back to SVS so it could do process switches rather than wait for each fault to be brought in. But that was several decades ago. --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.arch on Tue Sep 8 15:08:44 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> wrote:
    EricP <ThatWouldBeTelling@thevillage.com> schrieb:

    If programmers just used languages that checked array indexes
    then 99.999% of memory access errors would disappear.

    Retrofitting memory safety onto C is an uphill battle. Address
    arithmetic stands in the way of that.

    I think that any hardware solution compatible with legacy
    C practices will fail. Reasonably modern C allows VMT-s
    and using them one can track sizes. C standard is rather
    unhelpful here, because it says that size info is
    essentially ignored in prototypes, but compilers could
    implement checking as an extra option. So there is at
    least potential way to incrementally improve C code base
    and switch from legacy C to a checked one.

    First make array index checks simple and cheap.
    This requires checked arithmetic for the index expression calculation,
    plus a set of simple compare-and-fault instructions various bounds checks.

    You would probably need more than 32 registers for this...

    Also, smarten up compilers so they move as much as possible of the
    checking outside of loops.

    GCC twenty years ago was good enough. GNU Pascal for several years
    did no bound checking. Slightly more than twenty years ago it
    added bounds checking. Trials on several programs showed that
    typical impact was rather small, that is less than 10% overhead.
    In many cases there were no _new_ checks in object code: normal
    program logic ensured that accesses were in bound and GCC was
    able to infer that bounds checks were redundant. In other
    cases GCC moved accesses outside inner loop, making them quite
    cheap. And that was on i386 with its 8 register and Amd64
    with its 16 registers.

    I wonder, did you try to add bound checking to gfortran?
    ANd if yes did you measure performance impact?

    Then see what's left to address.

    Use after free will still be a problem, but maybe INVALIDATE
    can help there.

    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Tue Sep 8 15:48:28 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    ASIDs came along in the late 1990s and early 2000s to better optimize
    MMU resource consumption {fewer flushes, support nested paging, ...}.

    With their advent, ASIDs were assigned to threads/processes.

    Has anyone ever wanted to allow shared ASIDs in such a way that shared
    memory or shared files have the ASID associated with the memory/file ?
    So that multiple processes accessing the same shared resource co-optim-
    ize themselves across cores and caches ??

    Consider that the ASID (and VMID generally speaking) must be present
    in every TLB entry, so the size of the ASID (arm supports 8 or 16 bit ASIDS) determines the overall size of the TLB (along with the entry count).

    Even with 16-bit ASIDs, large systems with more than 64k processes (not uncommon with 1000-thread processors) will suffer during ASID rollover
    events. Associating ASIDs with other resources (such as memory objects) complicates the ASID rollever and assignment functionality without corresponding performance benefits. Using the G (Global) flag on those architectures which support it often provides similar benefits where applicable.

    Additionally, for your suggestion to be viable, every use of the shared
    object (memory, file pages) must be mapped at the same virtual address;
    which is considered a defect (see System V shared libraries for example),
    even with the larger (48-bit) virtual address space on modern CPUs.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 8 20:05:22 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    ASIDs came along in the late 1990s and early 2000s to better optimize
    MMU resource consumption {fewer flushes, support nested paging, ...}.

    With their advent, ASIDs were assigned to threads/processes.

    Has anyone ever wanted to allow shared ASIDs in such a way that shared >memory or shared files have the ASID associated with the memory/file ?
    So that multiple processes accessing the same shared resource co-optim-
    ize themselves across cores and caches ??

    Consider that the ASID (and VMID generally speaking) must be present
    in every TLB entry, so the size of the ASID (arm supports 8 or 16 bit ASIDS) determines the overall size of the TLB (along with the entry count).

    Known--also consider that TLB maps guest virtual to system physical
    (under nested paging) so the bit pattern in TLB does not correspond
    to the bit pattern in either PTE.

    Even with 16-bit ASIDs, large systems with more than 64k processes (not uncommon with 1000-thread processors) will suffer during ASID rollover events.

    Known issue.

    Associating ASIDs with other resources (such as memory objects) complicates the ASID rollever and assignment functionality without corresponding performance benefits.

    The perceived benefit is that every thread accessing a shared resource
    ends up using the shared resource ASID, lowering the cache/MMU overheads
    by allowing more sharing of the table-walk accelerators storage(s).

    Using the G (Global) flag on those architectures which support it often provides similar benefits where applicable.

    Do you really WANT all shared resources to be under exactly 1 <special>
    ASID {G} ??

    Security would demand that each resource would want only those granted
    access the ability to access--not something G does--unless you restrict
    shared resources to privileged threads.

    Additionally, for your suggestion to be viable, every use of the shared object (memory, file pages) must be mapped at the same virtual address;

    Only if one passes pointers within those spaces instead of passing offsets.

    which is considered a defect (see System V shared libraries for example), even with the larger (48-bit) virtual address space on modern CPUs.

    That is why I brought this up, to use comp.arch's brain trust to spit
    out the "other" important detains. Thanks.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 8 20:11:08 2026
    From Newsgroup: comp.arch

    Waldek Hebisch <antispam@fricas.org> schrieb:

    I wonder, did you try to add bound checking to gfortran?

    It has been there for a long time, and not by me :-)

    ANd if yes did you measure performance impact?


    I haven't run benchmarks, but analysis can be interesting.
    Consider

    subroutine foo(a,b,c,n)
    integer, intent(in) :: n
    real, dimension(n), intent(in) :: a,b
    real, dimension(n), intent(out) :: c
    integer :: i
    do i=1,n
    c(i) = a(i) + b(i)
    end do
    end subroutine foo

    Using gfortran, this is translated by the front end into a loop
    where every possible condition is checked on every iteration.

    Optimization passes convert this into a single check against
    INT_MAX, which inhibits vectorization.

    Interestingly, if both n and i are converted to 64-bit
    integers, the check is removed, and vectorization restored.

    Using Fortran's array syntax removes the checks completely, so

    c = a + b

    has no checks left. It also iserts all checks before the
    scalarizer-generated loops.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 00:19:05 2026
    From Newsgroup: comp.arch

    On Tue, 08 Sep 2026 08:28:53 GMT, Anton Ertl wrote:

    Given VM, why was there a need to add virtual memory to the guest
    OSs?

    Are you talking about virtual machines versus virtual memory? Or
    (wildly guessing) about the idea that if the host OS implements
    paging, the guests donrCOt need to?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 9 00:35:09 2026
    From Newsgroup: comp.arch


    Thomas Koenig <tkoenig@netcologne.de> posted:
    --------------
    Using Fortran's array syntax removes the checks completely, so

    c = a + b

    has no checks left.

    Given that the compiler can see the precise shapes of {a,b,c}
    the are no needs for checks. And that is one of the great
    benefits of higher-level expressions.

    It also iserts all checks before the
    scalarizer-generated loops.

    When the precise shapes are not (or cannot be) known at compile
    time.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 01:37:03 2026
    From Newsgroup: comp.arch

    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with something
    (X) is implemented as making V point to X. If X is another logical
    variable, and then is unified with something, say Y, a pointer to Y
    is stored in the memory location for X. And so on. Accessing V may
    need to follow an arbitrarily long chain of pointers until you find
    either a free variable, or you find the value that V eventually was
    unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should limit
    the amount of pointer-chasing necessary.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 9 08:19:37 2026
    From Newsgroup: comp.arch

    MitchAlsup wrote:

    Thomas Koenig <tkoenig@netcologne.de> posted:
    --------------
    Using Fortran's array syntax removes the checks completely, so

    c = a + b

    has no checks left.

    Given that the compiler can see the precise shapes of {a,b,c}
    the are no needs for checks. And that is one of the great
    benefits of higher-level expressions.

    Absolutely right, this should be a sufficient reason to make all
    compile-time constant size arrays and 2+ dimensional matrices their own
    types.


    It also iserts all checks before the
    scalarizer-generated loops.

    When the precise shapes are not (or cannot be) known at compile
    time.

    As long as the actual size is not miniscule, doing a single set of tests
    at startup is so close to free as to not matter.

    I.e something like if (a.len() == b.len && b.len == c.len()) could
    compile down to

    mov rdx,[rbx+8] ;; Dynamic vectors have {ptr, len, size} header
    cmp rdx,[rax+8]
    jne panic
    cmp rdx,[rcx+8]
    jne panic

    which would be predicted to not panic and therefore run in a cycle or
    two, right?

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 9 08:25:59 2026
    From Newsgroup: comp.arch

    Lawrence DrCOOliveiro wrote:
    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with something
    (X) is implemented as making V point to X. If X is another logical
    variable, and then is unified with something, say Y, a pointer to Y
    is stored in the memory location for X. And so on. Accessing V may
    need to follow an arbitrarily long chain of pointers until you find
    either a free variable, or you find the value that V eventually was
    unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should limit
    the amount of pointer-chasing necessary.
    That is the way dynamic method invocation works, at least for single-inheritance object-oriented languages:
    Each object of the type carries a pointer to a central location
    containing all the method pointers.
    Create a subclassed type means creating a copy of the parent descriptor
    before modifying it with all updated methods.
    I can still remember debugging my first OO program in asm and suddenly
    see how it actually worked with quite small memory overhead.
    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 08:13:57 2026
    From Newsgroup: comp.arch

    On Wed, 9 Sep 2026 08:25:59 +0200, Terje Mathisen wrote:

    Lawrence DrCOOliveiro wrote:

    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with
    something (X) is implemented as making V point to X. If X is
    another logical variable, and then is unified with something, say
    Y, a pointer to Y is stored in the memory location for X. And so
    on. Accessing V may need to follow an arbitrarily long chain of
    pointers until you find either a free variable, or you find the
    value that V eventually was unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should
    limit the amount of pointer-chasing necessary.

    That is the way dynamic method invocation works, at least for single-inheritance object-oriented languages:

    But what happens when you unify two variables that have each
    been unified with a bunch of others?

    In the shared-info scheme, you could handle this by merging one of
    those shared-info objects into the other one: copy all the variable backpointers from the former to the latter, and go through and fix all
    those variables from the first object so they point to the new, merged
    object. Then you can discard the old one.

    This avoids adding extra levels of indirection on variable references,
    at the cost of some more work at each unification.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 9 08:43:46 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with something
    (X) is implemented as making V point to X. If X is another logical
    variable, and then is unified with something, say Y, a pointer to Y
    is stored in the memory location for X. And so on. Accessing V may
    need to follow an arbitrarily long chain of pointers until you find
    either a free variable, or you find the value that V eventually was
    unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should limit
    the amount of pointer-chasing necessary.

    That would need back pointers stored everywhere, at least doubling the necessary memory (plus storing management information, because n
    logical variables can point to the same free logical variable), plus
    the memory for storing all these changes on the trail stack for
    backtracking. And all this memory has to be written. And in practice
    the chains of logical variables are usually short, so you would create
    all this overhead to address a rare case.

    But maybe I am wrong, and you are on to something. In that case you
    could revolutionize the implementation of unification and logic
    programming languages by demonstrating that your implementation
    performs better than the conventional ones.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 9 10:58:55 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:
    MitchAlsup wrote:

    Thomas Koenig <tkoenig@netcologne.de> posted:
    --------------
    Using Fortran's array syntax removes the checks completely, so

    c = a + b

    has no checks left.

    Given that the compiler can see the precise shapes of {a,b,c}
    the are no needs for checks. And that is one of the great
    benefits of higher-level expressions.

    Absolutely right, this should be a sufficient reason to make all compile-time constant size arrays and 2+ dimensional matrices their own types.

    Knowing array sizes at compile-time is a huge win. This allows all
    sorts of optimizations, including choice of vectorization, bounds
    checking, ...

    There is a weakness in Fortran's argument passing. You can
    pass the array bounds together with the array using assumed-shape
    dummy arguments (assumed in the sense that it it is determined by
    the actual argument, not assumed in the sense that the compiler
    thinks of some value), as in

    subroutine sub(a,b,c)
    real, intent(in), dimension(:) :: a, b
    real, intent(out), dimension(:) :: c
    c = a + b
    end subroutine sub

    but there is no way, other than a run-time check, to make sure that
    a, b and c have the same size.

    This is one reason why link-time optimization can be so powerful
    for Fortran - it can be used to see through these cases.



    It also iserts all checks before the
    scalarizer-generated loops.

    When the precise shapes are not (or cannot be) known at compile
    time.

    As long as the actual size is not miniscule, doing a single set of tests
    at startup is so close to free as to not matter.

    I.e something like if (a.len() == b.len && b.len == c.len()) could
    compile down to

    mov rdx,[rbx+8] ;; Dynamic vectors have {ptr, len, size} header
    cmp rdx,[rax+8]
    jne panic
    cmp rdx,[rcx+8]
    jne panic

    which would be predicted to not panic and therefore run in a cycle or
    two, right?

    Right.

    Although efficiency for short vectors / matrices is also important.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.arch on Wed Sep 9 18:21:59 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    John Levine <johnl@taugh.com> writes:
    They added virtual memory similar
    to the 67's to S/370 in 1972 and soon made all the major operating
    systems use it.

    Given VM, why was there a need to add virtual memory to the guest OSs?

    Orignal proposal for VM 370 included debugging MVS. Once MVS moved
    to virtual memory VM 370 needed to emulate this.

    People sometimes distinguish paravirtualization, when guest OS is
    aware and posibly depends on the host and full virtualization
    where guest OS "thinks" that it is running under real hardware.
    Once you have paravirtualization, there is nasty question how
    guest and host should divide the work and what is the best
    interface. In particular it makes sense to have all virtual
    memory management in the host. But paravitialization needs
    cooperation from the guest. So, for general use we have full
    vitialization and consequently guest using paging hardware.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 9 18:41:52 2026
    From Newsgroup: comp.arch


    antispam@fricas.org (Waldek Hebisch) posted:

    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    John Levine <johnl@taugh.com> writes:
    They added virtual memory similar
    to the 67's to S/370 in 1972 and soon made all the major operating >>systems use it.

    Given VM, why was there a need to add virtual memory to the guest OSs?

    Orignal proposal for VM 370 included debugging MVS. Once MVS moved
    to virtual memory VM 370 needed to emulate this.

    People sometimes distinguish paravirtualization, when guest OS is
    aware and posibly depends on the host and full virtualization
    where guest OS "thinks" that it is running under real hardware.

    Originally, paravirtualization was to speed up full virtualization.

    Once you have paravirtualization, there is nasty question how
    guest and host should divide the work and what is the best
    interface.

    Host OS needs an entry point to Guest OS to ask for resources back
    while allowing Guest OS to determine which resources to return.
    Guest OS needs an entry point in Host OS to ask for more resources
    while allowing Host OS to determine which resources to send.

    In particular it makes sense to have all virtual
    memory management in the host.

    Nested Paging has won. Guest OS virtualizes applications and its own
    worker threads, while Host OS virtualized multiple Guest OSs.

    When Host OS virtualized Guest OS, Guest OS ISR* can page fault at
    Host OS level without Guest OS taking a panic. (*) ISR can also be
    an exception handler, or dispatcher--those things written under an
    assumption of no faults or exceptions.

    But paravitialization needs
    cooperation from the guest. So, for general use we have full
    vitialization and consequently guest using paging hardware.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Wed Sep 9 12:23:31 2026
    From Newsgroup: comp.arch

    On 9/7/26 5:58 PM, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 20:25, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 07/09/2026 01:59, MitchAlsup wrote:
    [snip]
    On modern RISC ISAs:

    for( i = 0; i < max; i++ )
    p[i]

    is often faster than:

    for( i = 0; i < max; i++ )
    *p++

    Especially when there are more than 1 structure being accessed as an array.
    [snip]
    The second has 2 ADDs {i++ and p++; i of 1 and p of 4} it takes a lot
    of strength reduction to figure out that only 1 ADD is needed.


    The first one has two adds too - i++, and (p + i) for the array access.

    The p+i addition is part of AGEN and thus free for any ISA that has at
    least [Rpointer+Rindex] addressing mode (everybody but RISC-V).

    Itanium (not a RISC and probably most appropriate to use past
    tense though it may not yet be exclusively a
    collector's/museum's ISA) had even weaker address modes. If I
    recall correctly post-increment by constant was the most
    complex.

    Itanium intended to avoid some load latency with this simple
    addressing (presumably under the assumption that width was
    "free"). Even the small shift amount seems likely to add a
    bit of latency (the shift amount being provided by the
    instruction might slightly reduce this), but two layers of
    muxes (correct?) to shift by zero to three would not add that
    much delay I guess. (My 66000 adds a tiny bit of extra delay
    by supporting a third constant addend.)

    (For loops, it seems to me that one could use a pre-shifted
    index. This would reduce the range of short (e.g., 32-bit)
    indexes (constant range comparisons) and require adding a
    "shifted unit" to the index rather than one (possibly not
    supported with a short increment instruction). Such is also
    limited to the same member size of all arrays using the same
    index and would be more challenging for the compiler. Such an
    index would also have to be recognized as an array base offset
    rather than a true index (when passed as an argument or stored
    for later use).)

    This delay would seem to affect the cycle time of all
    instructions or force such loads to take an additional cycle or
    require folding the delay into all loads (the latency difference
    is presumably too small to be exactly compensated by a change in
    cache size or associativity). With frequency binning (and the
    smallness of the latency difference), such would presumably
    actually just modestly affect the commonness of better bins.

    For the MIPS R2000, providing a third source operand for stores
    would have been problematic. (GPR loads and FPR memory accesses
    could have been supported, though that complicates the
    compiler.) Pushing the read of the stored register value into
    the load delay slot might have worked, but that would have
    limited load delay slot instructions to those with one register
    operand. I think the lack of interlocks meant that a register
    read hazard could not just stall the pipeline. Banking might
    have allowed sufficient filling of the load delay slot (the
    two sources in the load delay slot could not both be in the same
    bank as the stored register), but banking would have complicated
    register allocation even if limited to load delay slots for
    indexed stores.

    [I also feel that implementing the branch delay slot as a
    "two-parcel" complex instruction fused with the branch
    instruction would have been better for exception handling.
    Avoiding delayed branches would have been better in the long
    term, but I am not certain whether alternatives would have been
    acceptable in terms of design time, area, and performance. For
    (inner) loops, a single BTB entry might have sufficed; for
    forward branches predict not taken might have sufficed, but
    misprediction correction might not have played well with the
    lack of interlocks. For unconditional jumps, I do not know
    what could have been done without predecoding or scan ahead or
    large BTBs. (I do not think scanning ahead was an option at the
    time unless one had common 16-bit instructions. "Simple"
    predecoding rCo a bit to indicate an immediate control flow
    instruction and replacement of the immediate offset with an
    inset, perhaps with a carry/borrow bit to allow full offset
    range rather than comparing the two most significant bits of
    the inset with the instruction address to determine an
    "overflow" rCo would introduce aliasing issues unless code
    pages were page colored or virtual tagging was used in the
    instruction cache.)]

    Besides probably not thinking of such at the time (FPR indexed
    memory accesses were added later so not everything was
    understood initially), borrowing operand read resources across
    instructions would have complicated the pipeline. I _think_
    this would not have added significant delay and I *guess* that
    it would not add significant area. A banking requirement might
    also be removed in later implementations (like the removal of
    the load delay slot, such relaxing of requirements allows old
    software to run on new systems though not necessarily all new
    software to run on old systems). (Development delay was also
    important for commercial MIPS, both of the hardware and the
    compiler.)

    [I *feel* that architectural banking, and other software
    cooperation, is underexploited. Of course, I have never written
    a compiler or even designed a paper ISA.]

    Obviously, the tradeoffs are very different in 2026 than in
    1986. The understanding of the tradeoffs among experts has
    probably also increased.

    I get the impression that RISC-V's design choice was urged by
    the desire for teaching a simple implementation (with
    "everything is an extension" 3-input instructions could have
    been pushed to an extension, so a specialized processor would
    not have had to pay for such). I wonder if a better teaching
    tactic would have been to design a complete, sophisticated ISA
    and either ignore the complex/difficult aspects (subsetting)
    or provide exploration of the tradeoffs and means of moderating
    those tradeoffs.

    I guess using an implicit return address register would have
    been an unavoidable complication of even simple subset
    implementations. Sacrificing five bits of target immediate seems
    a high price for such simplicity (especially for a destination
    operand). Of course, the C extension then adds implicit sources
    (for stack accesses) and different register field placements
    (not just size reduction rCo though a 16-bit instruction could
    only have two register fields if the fields were in the same
    positions as in 32-bit instructions) and became standard for the
    "application" profile. (Besides avoiding implicit registers,
    there was also a desire to support different return address
    registers for co-routines and otherwise to avoid return address
    predictor pollution.)

    RISC-V also suffers from its development/release schedule. The
    lack of iterative and thorough development before design choices
    were unalterable lead to inconsistencies and avoidable
    complexity. There also seems to have been a conceptual bias
    related to the extension focus where extensions are viewed as
    fully independent. E.g., the extension for code density was
    "compressed instructions" (duplicating 32-bit instructions,
    which was partially for implementation/conceptual simplicity)
    with no consideration for how 32-bit instructions could be
    provided to improve code density.

    This might also be influenced by the collaborative aspect of
    RISC-V design. Component isolation reduces communication
    overheads. Extreme isolation encourages fragmentation (aspects
    not included in the original design mandate require a separate component/extension), design limitation (an implementation
    detail may benefit a possible feature insufficiently relative to
    cost to justify the feature but be partially needed for features
    in other components), and design conflict (where some costs are
    transferred outside of the component's design concerns).

    [Once more, my mind is wandering. Almost all of the above is
    probably already understood by posters here, though perhaps a
    lurker might learn something rCo hopefully nothing false.]

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Wed Sep 9 16:30:56 2026
    From Newsgroup: comp.arch

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    Some software is highly optimized for a specific implementation
    (e.g., cache sizes) such that a microarchitectural change could
    require substantial rewriting/retuning. I feel the software
    developers should be allowed to make such fragile optimizations
    with the understanding that the microarchitecture is not
    guaranteed to continue even if the architecture does. (Some
    embedded chip vendors do provide product longevity guarantees.)

    (This can reduce microarchitectural flexiblity if the software
    is considered critical enough and rewriting/retuning is not an
    option.)

    Software that chooses to use such address bits for tags will
    still run on newer designs with full 64-bit virtual addresses
    though limited to 56-bit virtual addresses. I think that most
    software will never need 64-bit virtual addresses. I also think
    that a lot of software would not bother using address tagging.

    Address space constrains actual capability of software rather
    than "merely" performance, but some software developers (and
    some developers of the ARM architecture) have concluded that the
    benefit of having tags is worth the cost of a reduced address
    space as an option.

    I would also note that the canonical address scheme also
    introduces a software incompatibility aspect. Software using
    a non-canonical address to generate an exception and assuming
    the OS will not provide addresses outside the current limit
    (as of the 2024 man page I have Linux mmap does not provide a
    flag to limit the address range other than MAP_32BIT, so other
    than using fixed allocations, an application cannot force the
    OS to honor its wishes to limit the virtual address range).

    This is not a useful technique given how slow exceptions are,
    but I think it is a hole in "compatibility" (and more
    significant than space bar heating, https://xkcd.com/1172/ ).

    [snip]
    Interestingly, Stanford MIPS used the extra bits for an address
    space number and had a variable length mask. (-2The size of the
    process virtual address space is defined by a bit mask in a
    special register.

    My 66000 has a 3-bit LVL field in the Root pointer (in all MMU
    pointers). LVL in root tells you the VAS size (and PA==VA).
    A process==thread can have VAS as small as 23-bits (cat) and
    use 1 page as the page table. Saving all those unnecessary
    MMU accesses.

    LVL in PTPs allows for level skipping (sparse spaces). LVL in
    the pointer pointing at a PTE tells you size of the (super)page.
    No need for a control register field to specify what the natural
    organization of the table structure provides.

    I do appreciate those features of My 66000 page tables. I do
    wonder if using large pages in the page table itself might be
    worth supporting (Andy Glew had suggested such).

    With page level "merging" (large pages), one could (in theory)
    have 5-bit levels that could support multiples of 32 in page
    size while not requiring the page table depth to be doubled.
    256-KiB pages might be a useful option between 8 MiB and 8 KiB.

    Such also supports more flexible tradeoffs of depth versus
    internal fragmentation. An application using a large address
    space might have little sparsity in the upper address bits and
    so benefit from merging top page table levels.

    Level merging (really splitting) might also support more
    flexible sharing of page tables. With 5-bit basic levels, an
    aligned 256-byte section of a 8 KiB page table block could be
    shared without sharing the rest of the translations or
    permissions. (I have no idea if such would actually be useful
    much less worthwhile, but it is possible.)

    Level merging seems to be especially useful for nested page
    tables in virtualization. The virtual-physical address space
    is dense (I think). A virtual machine monitor might give huge
    pages to a guest but could also benefit from a flatter upper
    portion of the page table.

    OS support would be a problem. I do not think any architecture
    provides merging of page table levels.

    One weird thought I had was to have multiple page table "bases"
    loaded for a process. This would be a little like caching
    directory entries with prefetching of such on context switches.
    The advantage, such as it is, would be in supporting sparse
    use with a limited number of roots for faster context switches.
    Rather than having to traverse the page table three times to
    load three specific node points into the translation cache (and
    likely caching other less useful information until it ages out),
    the necessary entries would be prefetched and "locked".

    Two obvious issues come to mind. First, context switches are
    expected to be uncommon, so a modest savings becomes a trivial
    savings. Second, this does not scale down or up to different
    numbers of nodes.

    If the prefetching was optional, then a means would be required
    to load missing entries. This seems to require either a full
    page table (one node address from which all the other actually
    used/valid nodes can be found) rCo which uses extra storage and
    look-up layers rCo or something like a hash table (like software
    TLBs) that hold the information rCo which also has storage use
    overhead (though it might be limited by only storing entries
    greater than a default value, e.g., four nodes might be
    automatically prefetched and any other nodes would need to
    access the hash table) and look-up overhead (number of hash
    table probes, at least the parallelism of such allows
    tradeoffs of bandwidth versus latency). Since the prefetched
    nodes could be broadened to spaces covering multiple previous
    nodes, I think hash table misses might be avoidable by
    construction at the cost of deeper page tables.

    This was just a wild and crazy thought. The unconventionality
    alone probably makes such impractical even if it would be
    technically possible and perhaps even slightly useful.

    [snip]
    A valid design point in 1983, not so much today--unless vastly
    extended.

    I just thought it was a neat little design choice.

    [snip My 66000 avoidance of global bit]

    How close do you think you are to version 1.0?

    You sent me the Principles of Operation from January 2020 and
    your posts on comp.arch indicate that a substantial amount has
    changed since then. (I think the DOUBLE prefix has been added,
    dropped, and reinstated. CARRY was not present in 2020)

    I get the impression that a lot of your work recently has been
    on system (and implementation) aspects rather than "instruction
    set" aspects.

    I do not understand how the Virtual Vector Method would be
    implemented efficiently. SIMD with its explicit pack and unpack
    instructions seems likely to provide similar control. (Most SIMD
    designs do not provide special support for short strides or
    complete structure unpacking where all the data in a non-strided
    stream is used but the data is scattered. GPUs might provide
    special support for three and four color (un)packing.) However,
    I am not a hardware designer (though I might understand a simple
    flowchart ry|).

    I do feel that VVM is a nice software interface. It avoids a
    lot of SIMD issues and not just instruction diversity explosion.
    I feel it does not exploit a few local, non-loop wide execution
    cases, does not address blocking, and the loop length limits may
    be introduce issues. Yet I also recognize that being ideal for
    all use cases regardless of complexity is both unachievable
    (not all tradeoffs are limited to design difficulty or even
    implementation area) and impractical (aside from complexity
    having a cost, different workloads benefit differently from
    "effort" and the value of the benefit is uniform across all
    workloads at all scales).

    [snip]
    (32-bit PowerPC segments were vaguely similar in
    allowing 16 segments with separate virtual address spaces. In
    theory, this might support somewhat flexible sharing. HP PA-RISC
    provided fewer segments but more flexible/complex use. [I think
    encoding the segment bits in the least significant bits would
    have been better than using the high bits; it would have made
    the change to 64-bit simpler and allowed full-space dynamic
    segment addresses at the cost of having to encode segment
    numbers in the instruction for byte and half-word aligned
    pointers and disallowing alignment traps based on such bits.])

    More sins of the past...

    I am not sure. Segments based on the most significant bits
    certainly have issues with respect to increasing the base
    address space size, but I am not certain that segments as
    address space extensions are necessarily a bad thing. Simple
    address spaces are nicer, but the cost of doubling address size
    for a few uses that could be handled reasonably with segments
    (because of data locality even at that large scale).

    Maybe rather than segments (or large flat virtual address
    spaces) future huge memory applications might use multiple
    address spaces (program/code-connected segmentation) to
    support an effectively larger address space. Given the
    desirability of code and data physical locality and the
    latency and power issues of a large physical flat memory,
    such seems reasonable. Current warehouse-scale compute
    uses program partitioning (microservices) largely to provide
    usage scalability and availability, but sharding of databases
    does target the memory capacity issue (which is currently a
    tighter constraint than virtual address space size).

    [snip]
    Looking forward, loading 128-256 bits per access might be useful
    in the not so distant future, if only used for complex types
    {real, imag} with 64-bit and 128-bit element sizes.

    Microarchitecturally exploiting spatial locality (SRAM array
    width, block size, and page size) to avoid redundant work seems
    an obvious method toward power efficiency. Allowing software to
    help seems reasonable to *me* (but I am excessively attracted by hardware-software cooperation).

    With x86 one can perform a full-register-size load and access
    sub-sections (for some registers). Providing denser register
    storage use has advantages when memory is slow, but using
    subregisters complicates renaming (and forwarding, which is
    kind of renaming).

    The understatement of the month award winner.

    Surely that is a more temporally local understatement.ry| (Maybe
    limited to comp.arch?)

    A two-size renaming (like IBM zSeries [s/360 descendant]) would
    seem to only double the number of registers and the number of
    sources. There might even be techniques to simplify the hardware
    at the cost of storage utilization and/or extra data movement.

    Checkpoint-only values can be more readily copied to a different
    namespace (only a single pointer needs to be updated and the
    update is not on a critical path). Perhaps an operation might
    store two or more values rCo the result to be stored in quick
    storage and one or more source operands to be stored in slower
    checkpoint storage rCo avoiding an excessive read for the move and
    generally keeping the move out of the forwarding network.

    I _suspect_ subregisters have potential (certainly for in-order
    with "same lane" usage), but developing a reasonable design
    would probably require many person-years of research and by
    then tradeoffs would have changed and any modest benefit would
    likely be of limited application.

    (Like CMOV, out-of-order execution
    complicates use of subregisters.)

    For 70 years, a register would contain a single value, then
    MMX ruined the game...prior was the model compilers are good
    at using.

    x86 had subregisters before MMX. Even some RISCs used register
    pairs for double precision floating point, presenting the same renaming/forwarding issues.

    [snip]

    SIMD is bad for your architecture and for your thinking processes.

    I disagree. Vectors are Single Instruction Multiple Data. Fixed
    work unit size per instruction does seem architecturally
    problematic, but it is not clear that such is much more
    problematic for thinking than fixed cache block size.

    (Fixed cache block size does affect thinking. In-cache
    compression, especially "lossy" compression where part of the
    nominal cache block is stored in outer storage, will affect
    understanding of capacity. Prefetch/false sharing assumptions
    will affect performance.)

    [snip]
    The only thing your architecture cannot survive (over time)
    is lack of address bits. Do not give them away before your 4th
    generation implementations have sold 100M chips.

    The length of time between address space doubling doubles, and
    the delay is longer if cost-per-bit does not halve regularly
    and/or capacity demand is worse than linear with price. I think
    the speed of solid state storage reduced DRAM demand (cheaper
    indirectly addressed storage was good enough for many uses),
    lengthening the 64-bit generation. Process scaling issues also
    seem to be slowing demand for large memory single systems. The
    move toward scale out rather than scale up also reduces the
    address space pressure. I also suspect that many of the large
    virtual address space use cases are also relatively dense in
    mapping, which might buy a few bits of address space compared
    with "general" programs.

    Address tagging also only applies to software that chooses to
    use such. I get the impression that such tag use is rather
    niche and even then contained within a subsystem of the program.
    Having to rewrite a subsystem may be a huge pain, but

    It may be physically possible for a memory technology
    breakthrough to make Moore's Law look slow (for a while), but
    I would be willing to bet that virtual address space capacity
    demand will not accelerate and even that such will continue
    to decelerate (though not as much as from the one-time sold
    state storage effect or the, hopefully temporary, AI causes
    memory price increases).

    [snip]
    With ECAM based on PCIe 4.0+, your device address space consumes
    up to 40-bits. And then there is the configuration space, and the
    interrupt table aperture space, and other system spaces--all vying
    for those 47-bits. I suspect 47 will not be sufficient very long.

    Interesting, but the applications using tagged addresses are not
    likely to be interacting so directly with I/O (I think).

    My 66000's solution is to provide a complete 64-bit VAS that can be translated into a 66-bit universal address space consisting of four
    64-bit spaces {DRAM, device, config, ROM}. I don't want any of these
    spaces to cause issues while I remain alive. Both cores and devices
    use the 66-bit UAS.

    Side note: I have felt some attraction to a virtual Harvard
    system, i.e., a separate virtual address space for instructions.
    Besides making writing to code more explicit and slightly
    increasing the address space (less than a bit as code is rarely
    half of the used memory for larger systems), such might present
    opportunities to use the instruction pointer to generate data
    addresses that are not cluttered with instructions.

    I thought of the possibility of using negative offsets from a
    masked instruction pointer to provide a cheap pointer usable
    with shorter offsets. Separate instruction and data address
    spaces would allow

    Using MSI-X interrupts, allows interrupt tables to perform DPC/softIRQ queueing without normally associated SW overheads. Cores can send cores interrupts using the same mechanisms as devices. Since there is an
    unlimited number of interrupt tables (My 66000) and since the are
    constructed with a DAG structure, anything that can touch the interrupt aperture can send interrupts to any virtual core. When that virtual
    core has control, those interrupts are 'processed'.

    I like the orientation toward "an agent is an agent". That is
    just one more indication of My 66000 being a *design* and not
    merely a collection of useful ideas. I am not up to the task of
    designing a paper ISA, but I can still appreciate engineering.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 9 20:42:35 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    antispam@fricas.org (Waldek Hebisch) posted:

    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    John Levine <johnl@taugh.com> writes:
    They added virtual memory similar
    to the 67's to S/370 in 1972 and soon made all the major operating
    systems use it.

    Given VM, why was there a need to add virtual memory to the guest OSs?

    Orignal proposal for VM 370 included debugging MVS. Once MVS moved
    to virtual memory VM 370 needed to emulate this.

    People sometimes distinguish paravirtualization, when guest OS is
    aware and posibly depends on the host and full virtualization
    where guest OS "thinks" that it is running under real hardware.

    Originally, paravirtualization was to speed up full virtualization.

    That, and unusual tricks[*] to support a pseudo guest address space
    prior to AMD's nested page table.

    [*] Leveraging 32-bit x86 segment descriptors.


    Once you have paravirtualization, there is nasty question how
    guest and host should divide the work and what is the best
    interface.

    Host OS needs an entry point to Guest OS to ask for resources back
    while allowing Guest OS to determine which resources to return.
    Guest OS needs an entry point in Host OS to ask for more resources
    while allowing Host OS to determine which resources to send.

    In particular it makes sense to have all virtual
    memory management in the host.

    Nested Paging has won. Guest OS virtualizes applications and its own
    worker threads, while Host OS virtualized multiple Guest OSs.

    And ARM supports multiple nesting levels so a hypervisor can
    run as a guest of another hypervisor.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 9 20:50:08 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> writes:
    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    Some software is highly optimized for a specific implementation
    (e.g., cache sizes) such that a microarchitectural change could
    require substantial rewriting/retuning. I feel the software
    developers should be allowed to make such fragile optimizations
    with the understanding that the microarchitecture is not
    guaranteed to continue even if the architecture does. (Some
    embedded chip vendors do provide product longevity guarantees.)

    Recalling that most modern chips support a 48/9-bit virtual address space,
    the high eight bits of the virtual address won't be needed for
    addressing for likely several decades. Thus the use of the TBI
    setting for a specific address space is not a significant issue
    for future compatability. Noting that the top byte is also used
    by the Pointer Authentication instructions, it seems an useful
    and portable mechanism even without specific application support
    (other than compiling with PAC enabled).

    Even with the 56/7-bit VAS supported by high end xeon chips,
    the top byte is unused. Now on the physical side, the
    high order 2, 4 or 8 bits are sometimes used as a NUMA node index.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Wed Sep 9 21:45:07 2026
    From Newsgroup: comp.arch

    On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:

    Originally, paravirtualization was to speed up full virtualization.

    Still is. ItrCOs common to have drivers for block devices (disks, SSDs),
    in particular, in the guest, that are written to work through custom
    APIs provided by the host, rather than believe they are talking
    directly to actual (virtualized) hardware. You get better performance
    that way.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 9 21:49:39 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:

    Originally, paravirtualization was to speed up full virtualization.

    Still is. ItrCOs common to have drivers for block devices (disks, SSDs),
    in particular, in the guest, that are written to work through custom
    APIs provided by the host, rather than believe they are talking
    directly to actual (virtualized) hardware. You get better performance
    that way.

    That may have been true over a decade ago. The PCI-Express SRIOV feature
    was developed specifically to allow direct access to the hardware by
    a guest operating system, and also allowed the hardware to be shared
    by multiple guests (or usermode DPDK/ODP applications). Far better
    performance than paravirtualization.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 9 23:25:41 2026
    From Newsgroup: comp.arch


    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Some software is highly optimized for a specific implementation
    (e.g., cache sizes) such that a microarchitectural change could
    require substantial rewriting/retuning. I feel the software
    developers should be allowed to make such fragile optimizations
    with the understanding that the microarchitecture is not
    guaranteed to continue even if the architecture does. (Some
    embedded chip vendors do provide product longevity guarantees.)

    Yes, but You understand the difference between Architecture and
    implementation. May designers do not (especially early in their
    careers.)

    (This can reduce microarchitectural flexiblity if the software
    is considered critical enough and rewriting/retuning is not an
    option.)

    Which is why you should strive to eliminate those things before
    SW gets written.

    Software that chooses to use such address bits for tags will
    still run on newer designs with full 64-bit virtual addresses
    though limited to 56-bit virtual addresses. I think that most
    software will never need 64-bit virtual addresses. I also think
    that a lot of software would not bother using address tagging.

    Illustrating the small usage envelope of that mis-feature.

    Address space constrains actual capability of software rather
    than "merely" performance, but some software developers (and
    some developers of the ARM architecture) have concluded that the
    benefit of having tags is worth the cost of a reduced address
    space as an option.

    My 66000 allows for many sizes of VAS by using the LVL field in
    PTPs. You can say "all he did was move the control bit elsewhere"
    but I consider that it is the MMUs job do manage address space
    stuff. A 66000 single VAS can be {13, 23, 33, 43, 53, 63}-bits.
    and each privilege level gets 2 independent VASs. This DRAM VAS
    does not overlap with {ROM, config, device} address spaces (no
    apertures in PAS).

    I would also note that the canonical address scheme also
    introduces a software incompatibility aspect. Software using
    a non-canonical address to generate an exception and assuming
    the OS will not provide addresses outside the current limit
    (as of the 2024 man page I have Linux mmap does not provide a
    flag to limit the address range other than MAP_32BIT, so other
    than using fixed allocations, an application cannot force the
    OS to honor its wishes to limit the virtual address range).

    In My 66000, the high order bit of VA has to match the high order
    bit of Rbase of we throw an address exception. So, while one has
    2|u63-bit VASs one cannot start in one VAS and access in the other.

    Both VASs can be any combination of {13,23,33,43,53,63}-bit spaces.

    This is not a useful technique given how slow exceptions are,
    but I think it is a hole in "compatibility" (and more
    significant than space bar heating, https://xkcd.com/1172/ ).

    [snip]
    Interestingly, Stanford MIPS used the extra bits for an address
    space number and had a variable length mask. (-2The size of the
    process virtual address space is defined by a bit mask in a
    special register.

    My 66000 has a 3-bit LVL field in the Root pointer (in all MMU
    pointers). LVL in root tells you the VAS size (and PA==VA).
    A process==thread can have VAS as small as 23-bits (cat) and
    use 1 page as the page table. Saving all those unnecessary
    MMU accesses.

    LVL in PTPs allows for level skipping (sparse spaces). LVL in
    the pointer pointing at a PTE tells you size of the (super)page.
    No need for a control register field to specify what the natural organization of the table structure provides.

    I do appreciate those features of My 66000 page tables. I do
    wonder if using large pages in the page table itself might be
    worth supporting (Andy Glew had suggested such).

    I expect that either Guest OS will be supplied with large pages
    from Host OS, or vice versa. One simplifies Guest OS MMU management;
    the other simplifies Host OS management. My philosophy, here, is to
    allow SW e-the-large to figure it out for themselves (over time).

    With page level "merging" (large pages), one could (in theory)
    have 5-bit levels that could support multiples of 32 in page
    size while not requiring the page table depth to be doubled.
    256-KiB pages might be a useful option between 8 MiB and 8 KiB.

    A PTE mapping a large page in My 66000, uses the unused PTE.PA
    bits as a limit. So, while the large page still has large page
    alignment; it only needs to contain as many pages as required.
    So if you are using 23-bit large pages, you can use one PTE
    entry (in TLB) to access exactly 173 of those pages. ...

    Such also supports more flexible tradeoffs of depth versus
    internal fragmentation. An application using a large address
    space might have little sparsity in the upper address bits and
    so benefit from merging top page table levels.

    The LVL field can be used to skip the top bits (to VAS size),
    then used to skip intermediate levels (sparse VAS), and then
    skip the lower levels (super pages). All in one set of bits.

    Level merging (really splitting) might also support more
    flexible sharing of page tables. With 5-bit basic levels, an
    aligned 256-byte section of a 8 KiB page table block could be
    shared without sharing the rest of the translations or
    permissions. (I have no idea if such would actually be useful
    much less worthwhile, but it is possible.)

    Hard on the TLB.

    Level merging seems to be especially useful for nested page
    tables in virtualization. The virtual-physical address space
    is dense (I think). A virtual machine monitor might give huge
    pages to a guest but could also benefit from a flatter upper
    portion of the page table.

    There is currently a debate going on as to whether Guest OS should
    page its own pages or just have Host OS do it; Guest OS page fault
    handler just never gets control.

    OS support would be a problem. I do not think any architecture
    provides merging of page table levels.

    Not Much HW provides said feature.

    One weird thought I had was to have multiple page table "bases"
    loaded for a process. This would be a little like caching
    directory entries with prefetching of such on context switches.

    Basically My 66000 does this. Each core has a CoreStack, and each
    CoreStack contains 4 contexts. Each context contains a pointer to
    a runnable thread (.v=1). Each thread contains its "program Status
    Line" along with its register file.

    There is a 3-bit variable in CoreStack::
    000 = HyperVisor
    001 = Host OS
    010 = Guest OS
    011 = Application
    1xx = idle

    cs indexes CoreStack.context[], and by examining the context.v[] bits
    one can determine which privilege model is in place that instant,
    and which VASs can be accesses through the 2 available translation
    tables.

    The advantage, such as it is, would be in supporting sparse
    use with a limited number of roots for faster context switches.

    Each Root pointer is associated with its own ASID..

    And privilege switches need alter no control registers to receive
    control. The Guest OS context switch involves nothing more than
    writing the application context in CoreStack then SVRing.

    Rather than having to traverse the page table three times to
    load three specific node points into the translation cache (and
    likely caching other less useful information until it ages out),
    the necessary entries would be prefetched and "locked".

    ???

    Two obvious issues come to mind. First, context switches are
    expected to be uncommon, so a modest savings becomes a trivial
    savings. Second, this does not scale down or up to different
    numbers of nodes.

    ???

    If the prefetching was optional, then a means would be required
    to load missing entries. This seems to require either a full
    page table (one node address from which all the other actually
    used/valid nodes can be found) rCo which uses extra storage and
    look-up layers rCo or something like a hash table (like software
    TLBs) that hold the information rCo which also has storage use
    overhead (though it might be limited by only storing entries
    greater than a default value, e.g., four nodes might be
    automatically prefetched and any other nodes would need to
    access the hash table) and look-up overhead (number of hash
    table probes, at least the parallelism of such allows
    tradeoffs of bandwidth versus latency). Since the prefetched
    nodes could be broadened to spaces covering multiple previous
    nodes, I think hash table misses might be avoidable by
    construction at the cost of deeper page tables.

    My 66000 interrupt negotiator goes out and fetches the MSI-X
    message from its-core's interrupt table, then pre-fetches the
    PSL and registers while core continues running current thread.
    The Fetch message and fetch context delays are similar are run
    concurrently with current thread.

    Once MSI-X message verifies the core should take the interrupt,
    the new context is ready to install in core resources, while
    pushing our current state, saving latency at each step. While
    setting up the context, the first instruction is fetched, and
    the ISR is in control.

    This was just a wild and crazy thought. The unconventionality
    alone probably makes such impractical even if it would be
    technically possible and perhaps even slightly useful.

    [snip]
    A valid design point in 1983, not so much today--unless vastly
    extended.

    I just thought it was a neat little design choice.

    [snip My 66000 avoidance of global bit]

    How close do you think you are to version 1.0?

    You sent me the Principles of Operation from January 2020 and
    your posts on comp.arch indicate that a substantial amount has
    changed since then. (I think the DOUBLE prefix has been added,
    dropped, and reinstated. CARRY was not present in 2020)

    Mid 2025 I rearranged the instruction format so that <essentially>
    all instructions have {Size}|u{Sign}. {Size} determines the width
    of the calculation {[byte,half,word,dble];[FP8,FP16,FP32,FP64]}.
    {Sign} is mainly used to determine the non-exceptional range of
    a result, and often used as "just another OpCode bit" for those
    instructions where {Size} has no real meaning {branches,...}

    Read one way, I still have only 63 "instructions"; read another
    we have 350-ish "instructions", and there is still more than
    1/3 of major left available.

    I get the impression that a lot of your work recently has been
    on system (and implementation) aspects rather than "instruction
    set" aspects.

    Yes, I worked out System Binary Interface, control register addressing, Multiprecison arithmetic, virtual Vector Method,Exotic Synchronization
    Method, and To-Memory calculations.

    I do not understand how the Virtual Vector Method would be
    implemented efficiently. SIMD with its explicit pack and unpack
    instructions seems likely to provide similar control. (Most SIMD
    designs do not provide special support for short strides or
    complete structure unpacking where all the data in a non-strided
    stream is used but the data is scattered. GPUs might provide
    special support for three and four color (un)packing.) However,
    I am not a hardware designer (though I might understand a simple
    flowchart ry|).

    SIMD is simply direct compiler management of multi-lane calculations.

    vVM allows for the processor to utilize multi-lane calculations
    with flexibility SIMD can never do::

    word = word + half*byte;

    And if the HW happens to have 256-bits of calculation width, then
    8 of those can be performed in 1 pipelined clock, without any
    SW visible SIMD registers.

    I do feel that VVM is a nice software interface. It avoids a
    lot of SIMD issues and not just instruction diversity explosion.

    vVM is 2 instructions that provide the utility of thousands (based
    on typical SIMD ISA).

    vVM fails at the super luminary uses of SIMD (one lone SIMD inst
    without any containing loop structure)

    I feel it does not exploit a few local, non-loop wide execution
    cases, does not address blocking, and the loop length limits may
    be introduce issues.

    When a LD (or ST) touches a cache line and HW can determine the
    access pattern is "dense"; HW reads out cache-width of data and
    accesses that buffer multiple times feeding the multi-lane calc-
    ulation. So, the wide buffering flip-flops are present (they have
    to be for perf) but each implementation gets to decide how many
    and how wide, and how many cache ports are "reasonable" for this implementation.

    Those buffers are connected to the multi-lane calculation unit
    with lots of multiplexers (which perform the pack, unpack,
    width-changes, and forwarding). Those multiplexers take an
    extra cycle into and out of when considering latency.

    Yet I also recognize that being ideal for
    all use cases regardless of complexity is both unachievable
    (not all tradeoffs are limited to design difficulty or even
    implementation area) and impractical (aside from complexity
    having a cost, different workloads benefit differently from
    "effort" and the value of the benefit is uniform across all
    workloads at all scales).

    While I agree that it is more complex, is it more complex than
    having 1,300 SIMD instructions ??? And when we have the capability
    of building 10-wide machines (Apple M5) does that really matter?

    [snip]
    (32-bit PowerPC segments were vaguely similar in
    allowing 16 segments with separate virtual address spaces. In
    theory, this might support somewhat flexible sharing. HP PA-RISC
    provided fewer segments but more flexible/complex use. [I think
    encoding the segment bits in the least significant bits would
    have been better than using the high bits; it would have made
    the change to 64-bit simpler and allowed full-space dynamic
    segment addresses at the cost of having to encode segment
    numbers in the instruction for byte and half-word aligned
    pointers and disallowing alignment traps based on such bits.])

    More sins of the past...

    I am not sure. Segments based on the most significant bits
    certainly have issues with respect to increasing the base
    address space size, but I am not certain that segments as
    address space extensions are necessarily a bad thing. Simple

    s/Simple/Flat/

    address spaces are nicer, but the cost of doubling address size
    for a few uses that could be handled reasonably with segments
    (because of data locality even at that large scale).

    Segments are a poor man's capability. Go all the way or stop now.

    Maybe rather than segments (or large flat virtual address
    spaces) future huge memory applications might use multiple
    address spaces (program/code-connected segmentation) to
    support an effectively larger address space.

    As long as the smallest segment is at least 1 page, then My
    66000 PTE limit bits go a long way to providing much of what
    general segments provide.

    Given the
    desirability of code and data physical locality and the
    latency and power issues of a large physical flat memory,
    such seems reasonable.

    There is no rational for code and data to be within 2^40
    bits of each other. Code can go wherever it wants, data
    wherever it wants/needs. Over in PA space, code and data
    can be in back-to-back pages (this is virtual memory's job)

    Current warehouse-scale compute
    uses program partitioning (microservices) largely to provide
    usage scalability and availability, but sharding of databases
    does target the memory capacity issue (which is currently a
    tighter constraint than virtual address space size).

    Imagine using the top 16-bits of PA to route accesses over
    2^16 racks in a server farm.

    [snip]
    Looking forward, loading 128-256 bits per access might be useful
    in the not so distant future, if only used for complex types
    {real, imag} with 64-bit and 128-bit element sizes.

    Microarchitecturally exploiting spatial locality (SRAM array
    width, block size, and page size) to avoid redundant work seems
    an obvious method toward power efficiency. Allowing software to
    help seems reasonable to *me* (but I am excessively attracted by hardware-software cooperation).

    I don't think I am putting anything in the way of SW doing what it
    wants, here.

    With x86 one can perform a full-register-size load and access
    sub-sections (for some registers). Providing denser register
    storage use has advantages when memory is slow, but using
    subregisters complicates renaming (and forwarding, which is
    kind of renaming).

    The understatement of the month award winner.

    Surely that is a more temporally local understatement.ry| (Maybe
    limited to comp.arch?)

    A two-size renaming (like IBM zSeries [s/360 descendant]) would
    seem to only double the number of registers and the number of
    sources. There might even be techniques to simplify the hardware
    at the cost of storage utilization and/or extra data movement.

    vVM renames once and then uses that rename across multi-width
    calculations and multiple iterations of the loop. Saving rename
    power. Scalar registers and constants are loaded once and then
    used over all loop iterations. VEC annotates which registers
    are live-out, saving register writes along the way.

    Checkpoint-only values can be more readily copied to a different
    namespace (only a single pointer needs to be updated and the
    update is not on a critical path). Perhaps an operation might
    store two or more values rCo the result to be stored in quick
    storage and one or more source operands to be stored in slower
    checkpoint storage rCo avoiding an excessive read for the move and
    generally keeping the move out of the forwarding network.

    You lost me...

    I _suspect_ subregisters have potential (certainly for in-order
    with "same lane" usage), but developing a reasonable design
    would probably require many person-years of research and by
    then tradeoffs would have changed and any modest benefit would
    likely be of limited application.

    vVM tracks the LD-width and St-width to setup the buffer multiplexers.
    No need for each calculation to annotate that which is already known.

    (Like CMOV, out-of-order execution
    complicates use of subregisters.)

    For 70 years, a register would contain a single value, then
    MMX ruined the game...prior was the model compilers are good
    at using.

    x86 had subregisters before MMX. Even some RISCs used register
    pairs for double precision floating point, presenting the same renaming/forwarding issues.

    There you go again blaming x86 as being something good in computer architecture.

    [snip]

    SIMD is bad for your architecture and for your thinking processes.

    I disagree. Vectors are Single Instruction Multiple Data. Fixed
    work unit size per instruction does seem architecturally
    problematic, but it is not clear that such is much more
    problematic for thinking than fixed cache block size.

    The programmer wrote the sequence as a loop. Thus you should
    vectorize loops and not instructions. vVm does. vVm leverages
    other instruction scheduling mechanics in your implementation
    unlike vector-instructions.

    (Fixed cache block size does affect thinking. In-cache
    compression, especially "lossy" compression where part of the
    nominal cache block is stored in outer storage, will affect
    understanding of capacity. Prefetch/false sharing assumptions
    will affect performance.)

    Fixed block caching has certain hilarious consequences if
    the cache on THIS core has multiple lines that cannot hold
    data {either from locking or from bad cells in the cache}

    [snip]
    The only thing your architecture cannot survive (over time)
    is lack of address bits. Do not give them away before your 4th
    generation implementations have sold 100M chips.

    The length of time between address space doubling doubles, and
    the delay is longer if cost-per-bit does not halve regularly
    and/or capacity demand is worse than linear with price. I think
    the speed of solid state storage reduced DRAM demand (cheaper
    indirectly addressed storage was good enough for many uses),

    Yes, SSD took some demand out of DRAM size requirements, by
    shrinking 10ms rotational latency into 50-|s access delay;
    and by raising the data gat from 4GBs into the 30GB/s range.

    On the other hand, AI turned around and is consuming all available
    DRAM, so its a good thing SSDs are so fast.

    lengthening the 64-bit generation. Process scaling issues also
    seem to be slowing demand for large memory single systems. The
    move toward scale out rather than scale up also reduces the
    address space pressure. I also suspect that many of the large
    virtual address space use cases are also relatively dense in
    mapping, which might buy a few bits of address space compared
    with "general" programs.

    Address tagging also only applies to software that chooses to
    use such. I get the impression that such tag use is rather
    niche and even then contained within a subsystem of the program.

    Way back in 2006, the number of server-class chips AMD produced
    was measured as "an afternoon in the FAB" compared to the general
    PC market. So, we as designers were encouraged to add as much
    server stuff onto the die that would "fit" in an afternoon of
    designer work ...

    Having to rewrite a subsystem may be a huge pain, but

    It may be physically possible for a memory technology
    breakthrough to make Moore's Law look slow (for a while), but
    I would be willing to bet that virtual address space capacity
    demand will not accelerate and even that such will continue
    to decelerate (though not as much as from the one-time sold
    state storage effect or the, hopefully temporary, AI causes
    memory price increases).

    The laptop and desktop PCs are rarely memory limited these days.
    Gamers may be different. Servers ARE different, and Data Base
    is even different than that.

    [snip]
    With ECAM based on PCIe 4.0+, your device address space consumes
    up to 40-bits. And then there is the configuration space, and the
    interrupt table aperture space, and other system spaces--all vying
    for those 47-bits. I suspect 47 will not be sufficient very long.

    Interesting, but the applications using tagged addresses are not
    likely to be interacting so directly with I/O (I think).

    My 66000's solution is to provide a complete 64-bit VAS that can be translated into a 66-bit universal address space consisting of four
    64-bit spaces {DRAM, device, config, ROM}. I don't want any of these
    spaces to cause issues while I remain alive. Both cores and devices
    use the 66-bit UAS.

    Side note: I have felt some attraction to a virtual Harvard
    system, i.e., a separate virtual address space for instructions.

    Leads to problems for debuggers and for JIT.

    Besides making writing to code more explicit and slightly
    increasing the address space (less than a bit as code is rarely
    half of the used memory for larger systems), such might present
    opportunities to use the instruction pointer to generate data
    addresses that are not cluttered with instructions.

    Mc88110 had separatable Code and Data, nobody used it.

    The cache chips each had their own MMU+TLB so one could go full
    Harvard is they choose.

    I thought of the possibility of using negative offsets from a
    masked instruction pointer to provide a cheap pointer usable
    with shorter offsets. Separate instruction and data address
    spaces would allow

    ???

    Using MSI-X interrupts, allows interrupt tables to perform DPC/softIRQ queueing without normally associated SW overheads. Cores can send cores interrupts using the same mechanisms as devices. Since there is an unlimited number of interrupt tables (My 66000) and since the are constructed with a DAG structure, anything that can touch the interrupt aperture can send interrupts to any virtual core. When that virtual
    core has control, those interrupts are 'processed'.

    I like the orientation toward "an agent is an agent". That is
    just one more indication of My 66000 being a *design* and not
    merely a collection of useful ideas. I am not up to the task of
    designing a paper ISA, but I can still appreciate engineering.

    So, send me an e-mail address and I can send you ISA-55 and SFT-55
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 9 23:28:39 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:

    Originally, paravirtualization was to speed up full virtualization.

    Still is. ItrCOs common to have drivers for block devices (disks, SSDs),
    in particular, in the guest, that are written to work through custom
    APIs provided by the host, rather than believe they are talking
    directly to actual (virtualized) hardware. You get better performance
    that way.

    Is that prior to PCIe device virtualization or after or from something else ? --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Thu Sep 10 03:45:35 2026
    From Newsgroup: comp.arch

    On Wed, 09 Sep 2026 23:28:39 GMT, MitchAlsup wrote:

    On Wed, 9 Sep 2026 21:45:07 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:

    Originally, paravirtualization was to speed up full virtualization.

    Still is. ItrCOs common to have drivers for block devices (disks,
    SSDs), in particular, in the guest, that are written to work
    through custom APIs provided by the host, rather than believe they
    are talking directly to actual (virtualized) hardware. You get
    better performance that way.

    Is that prior to PCIe device virtualization or after or from
    something else ?

    ItrCOs a current thing.

    I looked around, and found a good overview here <https://wiki.xenproject.org/wiki/Understanding_the_Virtualization_Spectrum>:

    * In the beginning, before there was proper hardware support for
    virtualization, there was paravirtualization (PV), which required
    a specially-patched kernel.
    * Then there came full virtualization of all the hardware (HVM), which
    allowed using the same kernels that ran on actual physical hardware.
    * And then there was a recognition that virtualizing some low-level,
    performance-critical things -- network interfaces, block devices
    (disks, SSDs) was inefficient. So special drivers were created to
    allow more direct access through the host -- HVM with PV drivers.
    The kernel itself doesnrCOt need to know itrCOs running virtualized,
    just those drivers do.
    * Then, taking this a step further, low-level interrupt control was
    the next performance bottleneck identified in virtualized mode. So a
    special API was introduced to coordinate that between host and guest
    -- hence PVHVM.
    * And then there was this discussion we were having about whether to
    virtualize paging or not. And so there is this new thing called
    rCLPVHrCY (or maybe some other name).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Thu Sep 10 03:47:50 2026
    From Newsgroup: comp.arch

    On Wed, 9 Sep 2026 12:23:31 -0400, Paul Clayton wrote:

    Itanium (not a RISC and probably most appropriate to use past tense
    though it may not yet be exclusively a collector's/museum's ISA) ...

    If the Linux kernel drops support for it, you know it canrCOt have much
    life left in it. ;)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From kegs@kegs@provalid.com (Kent Dickey) to comp.arch on Thu Sep 10 04:11:09 2026
    From Newsgroup: comp.arch

    In article <vdenS.4266$Bn83.1011@fx12.iad>,
    Scott Lurndal <slp53@pacbell.net> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    It's useful to start with an existing experimental project and
    look at the current status thereof:

    https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/

    CHERI summary: all pointers are now called capabilities and are 129 bits,
    so registers which can hold addresses become 129 bits. Capabilities in
    memory are also 129 bits, with 128 bits being in "normal" memory, and the "valid" bit being elsewhere (think of it as hidden in ECC bits). Any
    data write to memory or a register clears the "valid" bit, so software cannot forge a capability. All loads/stores use capabilities as their pointers.

    The capability basically stores a normal pointer in the low 64 bits, and encodes the base address of the capability and the size in the upper 64 bits, with other info. It encodes this by saying larger sized objects have
    to have some minimum base and size alignment (so, 16MB+ object must be 4KB aligned, or something like that).

    The problem is CHERI is trying to do much more than just provide bounds checking--they also have sealed and unsealed capabilities, and lots
    of rules about managing these extra fields. But: they neglect to actually explain what they intend to do with this, so it all made little sense to me.
    I THINK they were trying to make capabilities so that the OS could re-use
    user capabilities (or user code using kernel capabilities) and not be a security hole, but honestly it made my eyes glaze over and I didn't figure
    it out.

    If you just want to catch bad memory references, then a simplified CHERI capability could work fine. You don't even need the hidden valid bit since forgery is not really an issue for this case.

    But CHERI has some holes (just off the top of my head, I'm sure there's more):

    - It needs to disallow capabilities in shared memory from creating security
    holes. One process mmap()'s some memory, and stores capabilities
    pointing to its private memory. Then, another process mmaps the
    same shared memory, and now can use those valid capabilities to
    access the same addresses in it's address space, which it might
    not have access to! This is a tricky problem to solve since you
    want shared libraries to work.

    - ARM allows a user process to be big-endian. This similarly can break
    the capability security model through mmap() and other means.

    - There's a whole slew of DMA-related security issues which seem hard to
    fully plug. You need to allow paging of capabilities, and this
    creates attack surfaces, or just complexity for real use cases
    where you don't want to clear the valid bit. All DMA drivers
    become part of the attack surface for forging capabilities.

    Again, if you don't try to enforce kernel-level trust on these capabilities, these no longer are issues, and it's why I don't really like CHERI.

    A CHERI-lite just providing bounds checking could be useful.

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Thu Sep 10 05:41:30 2026
    From Newsgroup: comp.arch

    On Wed, 09 Sep 2026 08:43:46 GMT, Anton Ertl wrote:

    On Wed, 9 Sep 2026 01:37:03 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with
    something (X) is implemented as making V point to X. If X is
    another logical variable, and then is unified with something, say
    Y, a pointer to Y is stored in the memory location for X. And so
    on. Accessing V may need to follow an arbitrarily long chain of
    pointers until you find either a free variable, or you find the
    value that V eventually was unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should
    limit the amount of pointer-chasing necessary.

    That would need back pointers stored everywhere, at least doubling
    the necessary memory (plus storing management information, because n
    logical variables can point to the same free logical variable), plus
    the memory for storing all these changes on the trail stack for
    backtracking. And all this memory has to be written.

    Only on actual unification, not on simple lookup.

    And in practice the chains of logical variables are usually short,
    so you would create all this overhead to address a rare case.

    How short is short, though? The question is, what is the average
    length of pointer chains being traversed in each scheme.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Thu Sep 10 06:16:11 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> schrieb:
    Waldek Hebisch <antispam@fricas.org> schrieb:

    I wonder, did you try to add bound checking to gfortran?

    It has been there for a long time, and not by me :-)

    ANd if yes did you measure performance impact?


    I haven't run benchmarks, but analysis can be interesting.
    Consider

    subroutine foo(a,b,c,n)
    integer, intent(in) :: n
    real, dimension(n), intent(in) :: a,b
    real, dimension(n), intent(out) :: c
    integer :: i
    do i=1,n
    c(i) = a(i) + b(i)
    end do
    end subroutine foo

    Using gfortran, this is translated by the front end into a loop
    where every possible condition is checked on every iteration.

    Optimization passes convert this into a single check against
    INT_MAX, which inhibits vectorization.

    Now https://gcc.gnu.org/bugzilla/show_bug.cgi?id=127303 .

    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 10 14:53:11 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).

    <big snip>


    The length of time between address space doubling doubles, and
    the delay is longer if cost-per-bit does not halve regularly
    and/or capacity demand is worse than linear with price. I think
    the speed of solid state storage reduced DRAM demand (cheaper
    indirectly addressed storage was good enough for many uses),

    Yes, SSD took some demand out of DRAM size requirements, by
    shrinking 10ms rotational latency into 50-|s access delay;
    and by raising the data gat from 4GBs into the 30GB/s range.

    On the other hand, AI turned around and is consuming all available
    DRAM, so its a good thing SSDs are so fast.

    AI is mostly using HBM, not SSDs.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 10 14:57:08 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Wed, 09 Sep 2026 23:28:39 GMT, MitchAlsup wrote:

    On Wed, 9 Sep 2026 21:45:07 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    On Wed, 09 Sep 2026 18:41:52 GMT, MitchAlsup wrote:

    Originally, paravirtualization was to speed up full virtualization.

    Still is. ItrCOs common to have drivers for block devices (disks,
    SSDs), in particular, in the guest, that are written to work
    through custom APIs provided by the host, rather than believe they
    are talking directly to actual (virtualized) hardware. You get
    better performance that way.

    Is that prior to PCIe device virtualization or after or from
    something else ?

    ItrCOs a current thing.

    Not in my experience. PCIe device virtualization has been preferred
    by customers for several years now, and is used in all but the
    lowest-end (consumer grade) hardware.


    I looked around, and found a good overview here ><https://wiki.xenproject.org/wiki/Understanding_the_Virtualization_Spectrum>:

    XEN[*] was designed more than two decades ago using 32-bit intel
    architecture processors. Times have changed radically since.

    [*] We partnered with them in 2005 to do some hypervisor work.

    Yes, you'll find PV drivers to access low-end slow hardware in
    desktop systems. From the server side, storage and networking
    are using PCIe SRIOV (e.g. NVMe and network accelerators).
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 10 14:59:16 2026
    From Newsgroup: comp.arch

    kegs@provalid.com (Kent Dickey) writes:
    In article <vdenS.4266$Bn83.1011@fx12.iad>,
    Scott Lurndal <slp53@pacbell.net> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    It's useful to start with an existing experimental project and
    look at the current status thereof:

    https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/

    CHERI summary: all pointers are now called capabilities and are 129 bits,
    so registers which can hold addresses become 129 bits. Capabilities in >memory are also 129 bits, with 128 bits being in "normal" memory, and the >"valid" bit being elsewhere (think of it as hidden in ECC bits). Any
    data write to memory or a register clears the "valid" bit, so software cannot >forge a capability. All loads/stores use capabilities as their pointers.


    https://www.arm.com/architecture/cpu/morello

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 10 15:16:19 2026
    From Newsgroup: comp.arch


    kegs@provalid.com (Kent Dickey) posted:

    Thank you Kent !

    In article <vdenS.4266$Bn83.1011@fx12.iad>,
    Scott Lurndal <slp53@pacbell.net> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Looking forward::

    a) What kind of distinction should architects make between pointers
    and addresses ??

    b) are there other ways to make capabilities cheaper without losing
    their protection properties ??

    As to (a) a pointer could have some bits used to restrict access
    rights. So, one could create a read-only pointer and use it only
    for reading even when the PTE says it is writeable.

    As to (b) a capability might have a 64-bit pointer and a 64-bit
    index into a capability table hidden in Guest OS address space.

    It's useful to start with an existing experimental project and
    look at the current status thereof:

    https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/

    CHERI summary: all pointers are now called capabilities and are 129 bits,
    so registers which can hold addresses become 129 bits. Capabilities in memory are also 129 bits, with 128 bits being in "normal" memory, and the "valid" bit being elsewhere (think of it as hidden in ECC bits). Any
    data write to memory or a register clears the "valid" bit, so software cannot forge a capability. All loads/stores use capabilities as their pointers.

    The capability basically stores a normal pointer in the low 64 bits, and encodes the base address of the capability and the size in the upper 64 bits, with other info. It encodes this by saying larger sized objects have
    to have some minimum base and size alignment (so, 16MB+ object must be 4KB aligned, or something like that).

    The problem is CHERI is trying to do much more than just provide bounds checking--they also have sealed and unsealed capabilities, and lots
    of rules about managing these extra fields. But: they neglect to actually explain what they intend to do with this, so it all made little sense to me.

    I tend to say it as "I understand how to make a capability "pointer",
    what I don't understand is how to make it C3-secure."

    I THINK they were trying to make capabilities so that the OS could re-use user capabilities (or user code using kernel capabilities) and not be a security hole, but honestly it made my eyes glaze over and I didn't figure
    it out.

    Whereas; most of use simply want solid base-bounds checking. S O L I D

    If you just want to catch bad memory references, then a simplified CHERI capability could work fine. You don't even need the hidden valid bit since forgery is not really an issue for this case.

    Do you think that with a 63-bit VAS, one could put each root-capability
    into its own upper layer paging structure ?? WHere derived-capabilities
    simply point inside that root-capability ??

    But CHERI has some holes (just off the top of my head, I'm sure there's more):

    - It needs to disallow capabilities in shared memory from creating security
    holes. One process mmap()'s some memory, and stores capabilities
    pointing to its private memory. Then, another process mmaps the
    same shared memory, and now can use those valid capabilities to
    access the same addresses in it's address space, which it might
    not have access to! This is a tricky problem to solve since you
    want shared libraries to work.

    Sort-of defeaters the whole purpose, does it not ??

    - ARM allows a user process to be big-endian. This similarly can break
    the capability security model through mmap() and other means.

    - There's a whole slew of DMA-related security issues which seem hard to
    fully plug. You need to allow paging of capabilities, and this
    creates attack surfaces, or just complexity for real use cases
    where you don't want to clear the valid bit. All DMA drivers
    become part of the attack surface for forging capabilities.

    This simply makes PCIe devices impossible as each "address plus size"
    has to be a capability.

    Again, if you don't try to enforce kernel-level trust on these capabilities,

    s/kernel/hypervisor/

    Or something like nested capabilities; where a guest capability is re-verified/re-translated by host capabilities.

    these no longer are issues, and it's why I don't really like CHERI.

    My eyes glassed over, too.

    A CHERI-lite just providing bounds checking could be useful.

    Indeed.

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 10 16:06:34 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    kegs@provalid.com (Kent Dickey) posted:



    The problem is CHERI is trying to do much more than just provide bounds
    checking--they also have sealed and unsealed capabilities, and lots
    of rules about managing these extra fields. But: they neglect to actually >> explain what they intend to do with this, so it all made little sense to me.

    I tend to say it as "I understand how to make a capability "pointer",
    what I don't understand is how to make it C3-secure."

    The out-of-band 'tag' bit which marks a valid capability makes it
    secure (likely at B level or better in orange book terminology).


    I THINK they were trying to make capabilities so that the OS could re-use
    user capabilities (or user code using kernel capabilities) and not be a
    security hole, but honestly it made my eyes glaze over and I didn't figure >> it out.

    Whereas; most of use simply want solid base-bounds checking. S O L I D

    Which, when it applies to programming environments that include dynamically loaded code (which is pretty much all of them in these times). The capability, of course, provides additional valuable capabilities over simple bounds checking,
    such as the ability to mark a datum as read-only (without marking the entire page containing the datum).


    If you just want to catch bad memory references, then a simplified CHERI
    capability could work fine. You don't even need the hidden valid bit since >> forgery is not really an issue for this case.

    Do you think that with a 63-bit VAS, one could put each root-capability
    into its own upper layer paging structure ?? WHere derived-capabilities >simply point inside that root-capability ??

    But CHERI has some holes (just off the top of my head, I'm sure there's more):

    - It needs to disallow capabilities in shared memory from creating security >> holes. One process mmap()'s some memory, and stores capabilities
    pointing to its private memory. Then, another process mmaps the
    same shared memory, and now can use those valid capabilities to
    access the same addresses in it's address space, which it might
    not have access to! This is a tricky problem to solve since you
    want shared libraries to work.

    Sort-of defeaters the whole purpose, does it not ??

    If that had been an accurate summarization of how CHERI interacts
    with mmap (and shared libraries in general), perhaps.

    However, if you think that the CHERI team hasn't given significant
    though to the issues of inter-thread and inter-process memory sharing;
    you might want to re-read the CHERI documentation.


    - ARM allows a user process to be big-endian. This similarly can break
    the capability security model through mmap() and other means.

    I don't see how. A capability is opaque to software and completely independent of the endianness of the current thread. It must be 8-byte
    aligned (which reduces the number of tag bits to 1/8th).


    - There's a whole slew of DMA-related security issues which seem hard to
    fully plug. You need to allow paging of capabilities, and this
    creates attack surfaces, or just complexity for real use cases
    where you don't want to clear the valid bit. All DMA drivers
    become part of the attack surface for forging capabilities.

    This simply makes PCIe devices impossible as each "address plus size"
    has to be a capability.

    Again, one must read the CHERI documentation before commenting on
    such things.

    "CHERI (Capability Hardware-Enhanced RISC Instructions) interacts
    with Direct Memory Access (DMA) by treating peripheral devices as
    capability-unaware entities and enforcing hardware tag-clearing
    mechanisms at the memory/interconnect level to protect capability
    integrity."


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 10 16:36:41 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_.

    I agree.

    2^64 is a, pardon my french,
    shitload of virtual memory.

    That's not the reason. With the larger machines needing maybe 44 bits
    for physical addresses, and the old pattern of 1 bit every 18 months
    or two years, we would exceed 56/57 bits in 18-26 years and exceed 64
    bits in 30-40 years. But the growth has slowed down substantially,
    and I expect it to slow down even more, so that we will never reach
    physical memories that require 64 bits, and probably not 56 bits,
    either.

    And if we don't have that much physical memory, do we need VAs with
    more than 56 bits? Yes, one can do cool things with address space
    even if physical memories are not that large, but one can also do cool
    things with bits that are ignored.

    Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required.

    With 4KB pages and 8-byte page-table entries, 6 levels of page tables
    are good for 6*9+12=66 bits. With larger pages, you need even fewer
    levels (all assuming 8-byte PTEs):

    Page size levels bits
    4KB 6 66
    8KB 6 73 (63bits with 5 levels)
    16KB 5 69
    32KB 5 75 (63 bits with 4 levels)
    64KB 4 68

    And anyway, make the TLBs large enough, and full table walks will be
    rare.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 10 16:55:56 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:
    Again, one must read the CHERI documentation before commenting on
    such things.

    Nope. If the CHERI people don't manage to make their point in a
    decently short position paper, they will have to live with people
    commenting on CHERI without getting their point (if there actually is
    one). OTOH, Mitch Alsup missing a clearly-made point is not
    unheard-of, either. But they also failed to convince me in the paper
    I read from them.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 10 17:02:03 2026
    From Newsgroup: comp.arch

    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but
    certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[5];
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a,
    sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does
    things beyond a single nested struct, such as garbage collection.

    I guess there are ways to do these things in CHERI, but I very much
    doubt that the only cost CHERI has for them is the doubled memory
    consumption. And does it buy anything that Rust does not buy us?

    Given that the computing world seems to have decided that it does not
    even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will
    be a widespread success.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 10 17:16:37 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Wed, 09 Sep 2026 08:43:46 GMT, Anton Ertl wrote:

    On Wed, 9 Sep 2026 01:37:03 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    On Tue, 08 Sep 2026 08:34:36 GMT, Anton Ertl wrote:

    Unification of a pre-existing free logical variable V with
    something (X) is implemented as making V point to X. If X is
    another logical variable, and then is unified with something, say
    Y, a pointer to Y is stored in the memory location for X. And so
    on. Accessing V may need to follow an arbitrarily long chain of
    pointers until you find either a free variable, or you find the
    value that V eventually was unified with.

    There should be a more efficient way: having all bound variables
    contain a pointer to shared info about the binding (including
    backpointers to all the variables that point here). That should
    limit the amount of pointer-chasing necessary.

    That would need back pointers stored everywhere, at least doubling
    the necessary memory (plus storing management information, because n
    logical variables can point to the same free logical variable), plus
    the memory for storing all these changes on the trail stack for
    backtracking. And all this memory has to be written.

    Only on actual unification, not on simple lookup.

    Every "simple lookup" is a unification in Prolog. But if you mean
    something that a WAM-based Prolog compiler compiles into a load into
    an X register: Yes there is no writing to memory here, but there is
    also no following of the pointer chain (if any).

    And in practice the chains of logical variables are usually short,
    so you would create all this overhead to address a rare case.

    How short is short, though? The question is, what is the average
    length of pointer chains being traversed in each scheme.

    I have been out of Prolog implementation for over three decades, but
    my impression is that the common case is a length of 1, and the
    average may be a length of 2 (with a very low confidence).

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 10 17:22:37 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 10 18:58:36 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but
    certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[3]; // making a 2-byte hole in the struct element
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a,
    sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does
    things beyond a single nested struct, such as garbage collection.

    And sometimes you even want to prevent accessing data that is not
    "in" the struct due to internal alignments--such as the bytes between
    u[2] and v[0].

    I guess there are ways to do these things in CHERI, but I very much
    doubt that the only cost CHERI has for them is the doubled memory consumption. And does it buy anything that Rust does not buy us?

    Given that the computing world seems to have decided that it does not
    even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will
    be a widespread success.

    My 66000 architecture has means maintain robustness in the face of
    Spectr|- attack vectors. Stashing microarchitectural state in the
    miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    Don't see why CHERI could not do similarly.

    Nor do I see other microarchitectures implementing this fairly easy
    "suit of armor" to mitigate Spectr|-!?!?

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 10 19:01:44 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but
    certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[3]; // making a 2-byte hole in the struct element
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a,
    sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does
    things beyond a single nested struct, such as garbage collection.

    And sometimes you even want to prevent accessing data that is not
    "in" the struct due to internal alignments--such as the bytes between
    u[2] and v[0].

    That would be difficult, since one needs to be able to copy
    the structure, which necessarily will need to access those bytes.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Thu Sep 10 22:42:59 2026
    From Newsgroup: comp.arch

    Anton Ertl wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    SOME compilers generate better code than SOME programmers, SOME of the time.

    - fixed that for you.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Thu Sep 10 22:45:40 2026
    From Newsgroup: comp.arch

    Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but
    certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[3]; // making a 2-byte hole in the struct element
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a,
    sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does
    things beyond a single nested struct, such as garbage collection.

    And sometimes you even want to prevent accessing data that is not
    "in" the struct due to internal alignments--such as the bytes between
    u[2] and v[0].

    That would be difficult, since one needs to be able to copy
    the structure, which necessarily will need to access those bytes.


    Not true:

    You _just_ need to do such a deep copy element by element, and suffer
    about an order of magnitude slowdown.

    For better performance, the compiler will figure out if there are any
    holes, if not just do a memcpy().

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From George Neuner@gneuner2@comcast.net to comp.arch on Thu Sep 10 17:02:00 2026
    From Newsgroup: comp.arch

    On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    - anton

    "compilers generate better code than MOST programmers."


    The qualification is important.

    Keep in mind that most programmers are only average and the skill
    level of the average programmer now is only slightly above "script
    kiddie". Most have no formal CS or CSE schooling, don't know how to
    evaluate algorithms, and largely are incapable of writing for
    themselves library functions that they routinely use.

    Witness the proliferation of languages offering "managed environments"
    offering such niceties as automatic storage management, automatic lock
    handling (serialized object access), "comprehensions", etc., and large
    standard libraries - without which the average programmer largely
    would be incapable of producing a working program.

    YMMV.


    In general the compiler can be aware of more surrounding context and
    can generate better initial code. There is no doubt that a good
    programmer can beat the compiler, but in my experience [HRT systems]
    it requires effort that profitably might be spent on something else.
    Even good programmers can have poor intuition about what needs high optimization.

    Again, YMMV.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Robert Finch@robfi680@gmail.com to comp.arch on Thu Sep 10 21:43:23 2026
    From Newsgroup: comp.arch

    On 2026-09-10 12:06 p.m., Scott Lurndal wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    kegs@provalid.com (Kent Dickey) posted:



    The problem is CHERI is trying to do much more than just provide bounds
    checking--they also have sealed and unsealed capabilities, and lots
    of rules about managing these extra fields. But: they neglect to actually >>> explain what they intend to do with this, so it all made little sense to me.

    I tend to say it as "I understand how to make a capability "pointer",
    what I don't understand is how to make it C3-secure."

    The out-of-band 'tag' bit which marks a valid capability makes it
    secure (likely at B level or better in orange book terminology).


    I THINK they were trying to make capabilities so that the OS could re-use >>> user capabilities (or user code using kernel capabilities) and not be a
    security hole, but honestly it made my eyes glaze over and I didn't figure >>> it out.

    Whereas; most of use simply want solid base-bounds checking. S O L I D

    Which, when it applies to programming environments that include dynamically loaded code (which is pretty much all of them in these times). The capability,
    of course, provides additional valuable capabilities over simple bounds checking,
    such as the ability to mark a datum as read-only (without marking the entire page containing the datum).

    I wondered about adding read/write/execute access levels to the
    capability pointer instead of just having single bit for each. But it
    would add even more bits to the capability.

    Suppose one wanted the same capability pointer but with a different read access level to the data. When r/w/x bits are all that is available I
    think one has to map a new pointer with the MMU.>

    If you just want to catch bad memory references, then a simplified CHERI >>> capability could work fine. You don't even need the hidden valid bit since >>> forgery is not really an issue for this case.

    Do you think that with a 63-bit VAS, one could put each root-capability
    into its own upper layer paging structure ?? WHere derived-capabilities
    simply point inside that root-capability ??

    But CHERI has some holes (just off the top of my head, I'm sure there's more):

    - It needs to disallow capabilities in shared memory from creating security >>> holes. One process mmap()'s some memory, and stores capabilities
    pointing to its private memory. Then, another process mmaps the
    same shared memory, and now can use those valid capabilities to
    access the same addresses in it's address space, which it might
    not have access to! This is a tricky problem to solve since you
    want shared libraries to work.

    Sort-of defeaters the whole purpose, does it not ??

    If that had been an accurate summarization of how CHERI interacts
    with mmap (and shared libraries in general), perhaps.

    However, if you think that the CHERI team hasn't given significant
    though to the issues of inter-thread and inter-process memory sharing;
    you might want to re-read the CHERI documentation.


    - ARM allows a user process to be big-endian. This similarly can break
    the capability security model through mmap() and other means.

    I don't see how. A capability is opaque to software and completely independent of the endianness of the current thread. It must be 8-byte aligned (which reduces the number of tag bits to 1/8th).


    - There's a whole slew of DMA-related security issues which seem hard to >>> fully plug. You need to allow paging of capabilities, and this
    creates attack surfaces, or just complexity for real use cases
    where you don't want to clear the valid bit. All DMA drivers
    become part of the attack surface for forging capabilities.

    This simply makes PCIe devices impossible as each "address plus size"
    has to be a capability.

    Again, one must read the CHERI documentation before commenting on
    such things.

    "CHERI (Capability Hardware-Enhanced RISC Instructions) interacts
    with Direct Memory Access (DMA) by treating peripheral devices as
    capability-unaware entities and enforcing hardware tag-clearing
    mechanisms at the memory/interconnect level to protect capability
    integrity."



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 11 06:00:19 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    Given that the computing world seems to have decided that it does not
    even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will
    be a widespread success.

    My 66000 architecture has means maintain robustness in the face of
    Spectr|- attack vectors. Stashing microarchitectural state in the
    miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    This is a good approach to implement invisible speculation as far as
    cache side channel is concerned. Behnia et al. [behnia+21] describe a
    side channel that uses resource contention from speculative
    instructions to affect the timing of committing instructions, i.e., it
    uses the scheduler (reservation station) behaviour as a side channel.

    Behnia et al. also describe how to close this side channel. There may
    be other side channels through microarchitectural resources, and they
    need to be closed, too, but the cache side channel probably is the
    widest and therefore most relevant one by far.

    @InProceedings{ behnia+21,
    author = {Mohammad Behnia and Prateek Sahu and Riccardo Paccagnella
    and Jiyong Yu and Zirui Neil Zhao and Xiang Zou and Thomas
    Unterluggauer and Josep Torrellas and Carlos Rozas and Adam
    Morrison and Frank Mckeen and Fangfei Liu and Ron Gabor and
    Christo- pher W. Fletcher and Abhishek Basak and Alaa
    Alameldeen},
    title = {Speculative Interference Attacks: Breaking Invisible
    Speculation Schemes},
    booktitle = {Architectural Support for Programming Languages and
    Operating Systems (ASPLOS rCO21)},
    year = {2021},
    pages = {1046--1060},
    url = {https://dl.acm.org/doi/10.1145/3445814.3446708},
    optannote = {}
    }

    Spectre and friends are microarchitectural side channels, you can
    implement microarchitectures for any architecture that are vulnerable
    to Spectre, and microarchitectures that are not vulnerable. Therefore
    your mention of the My 66000 architecture makes no sense.

    Don't see why CHERI could not do similarly.

    Certainly one can implement a core with CHERI that is not vulnerable
    to speculative side channels (either by not implementing speculation
    (slow), or by implementing invisible speculation), and if the
    financers and the prospective customers are serious about security,
    they will insist on that for eventual products.

    But my point is that in the mainstream, not even invisible speculation
    with its moderate cost in hardware and performance is implemented, so
    I don't expect that a high-cost solution like CHERI becomes
    mainstream, in particular given that it provides nothing that cannot
    be provided more cheaply through programming languages and compilers.

    Ok, you might say, what about legacy code in unsafe languages such as
    C? Rewriting all of that in Rust (or using a solution like Ivy/Deputy <http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>, which
    would be cheaper, but somehow did not catch on) would also be very
    costly. But I expect that much of this code does not work on CHERI in
    the secure configuration; so to run it on CHERI, you would have to
    change it anyway, so you could just as well Deputize it, which is
    cheaper in hardware and in execution time. Or rewrite it in Rust, if
    you prefer that.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 11 06:32:14 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    Anton Ertl wrote:
    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    SOME compilers generate better code than SOME programmers, SOME of the time.

    That's not what there people write...

    - fixed that for you.

    ... so the fix is wrong.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Fri Sep 11 07:01:06 2026
    From Newsgroup: comp.arch

    On Thu, 10 Sep 2026 17:02:03 GMT, Anton Ertl wrote:

    I guess there are ways to do these things in CHERI, but I very much
    doubt that the only cost CHERI has for them is the doubled memory consumption. And does it buy anything that Rust does not buy us?

    From what I understand of the CHERI project, they have produced a
    complete desktop OS in the form of CheriBSD, complete with a port of
    the KDE Plasma desktop environment. This is a pretty comprehensive
    collection of C++ code, so it sounds like a good testbed for getting
    experience with running a real-world codebase under the CHERI
    constraints.

    Interestingly, the reason they gave for choosing a BSD variant as a
    basis rather than Linux, was that the Linux project was moving too
    fast for them to keep up. DidnrCOt they know about LTS releases?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 11 06:33:21 2026
    From Newsgroup: comp.arch

    George Neuner <gneuner2@comcast.net> writes:
    On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:
    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    "compilers generate better code than MOST programmers."


    The qualification is important.

    The statements I have read did not make such a qualification,
    certainly not in capital letters.

    There is no doubt that a good
    programmer can beat the compiler,

    And especially this statement is usually not made. On the contrary,
    the perpetrator of such statements seem convinced of compiler
    supremacy.

    but in my experience [HRT systems]
    it requires effort that profitably might be spent on something else.

    I don't know what HRT systems are, but sure, if the performance is
    good enough, any time spent on increasing it is wasted. However, if performance is not good enough, spending even relatively moderate time
    on it can yield good results. However, I have heard of cases where
    the lack of performance was rooted so deeply in the system that it had
    to be rewritten (yielding a 10,000 times speedup in one case) or (in
    another case) the project was abandoned, after efforts to save it by
    improving the software, or by throwing more hardware at it had failed.

    [reordered]
    Keep in mind that most programmers are only average and the skill
    level of the average programmer now is only slightly above "script
    kiddie". Most have no formal CS or CSE schooling, don't know how to
    evaluate algorithms, and largely are incapable of writing for
    themselves library functions that they routinely use.

    I don't think that infecting them with the idea of compiler supremacy
    will make them better programmers.

    Michael Jackson's rules in the matter of optimization (and their
    explanation) are more sensible:

    |Rule 1: Don't do it.
    |Rule 2 (for experts only). Don't do it yet.

    (See <https://blog.plover.com/prog/optimization.html> for a longer
    discussion of that.)

    Witness the proliferation of languages offering "managed environments" >offering such niceties as automatic storage management, automatic lock >handling (serialized object access), "comprehensions", etc., and large >standard libraries

    I don't see any problem with that. These features help to implement functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    In general the compiler can be aware of more surrounding context and
    can generate better initial code.

    On the contrary: Programmers know the requirements, while compilers
    only know the source program. Even as far as the source code is
    concerned, programmers also know much more about the complete program,
    and can perform much more far-reaching transformations that compilers
    usually cannot perform (e.g., from a linked list to a
    structure-of-arrays data representation).

    I have no idea how "initial code" is relevant. We usually optimize at
    the source code level, so why should that play a role?

    Even good programmers can have poor intuition about what needs high >optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 11 10:15:39 2026
    From Newsgroup: comp.arch

    On 10/09/2026 19:22, Anton Ertl wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.


    The word "code" can mean many things (especially "source code" and
    "object code"), "programmer" can mean many things, and "better" can mean
    a vast array of different things. So without a lot more context or qualification, the claim does not mean much. As a stand-alone claim, I
    would not say it is wrong - I would say it is "not even wrong".

    And I have always thought that programmers and compilers work together -
    the compiler is a tool used by the programmer. They do different jobs,
    and can't be sensibly compared.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Fri Sep 11 08:54:15 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    Your statement is ambiguous in several ways. I assume you mean
    "programmers using assembly" or "programmers using machine code",
    although the latter has now wildly fallen out of fashion.

    Do you claim that people claim

    a) Compilers generate better code than programmers all the time

    b) Compilers generate better code than programmers most of the time

    a) is patently false. For somebody reasonably proficient at reading
    assembly, it is rather easy to spot inefficiencies in generated code
    (although also easy to be wrong, measurements tops everything).

    b) is true. In the vast majority of cases, it is vastly uneconomical
    to revert to assembly programming. There are exceptions, for
    example hand-tuned BLAS routines.

    In that respect, it is interesting to read about the discussions
    surrounding the very first optimizing compiler, the original
    FORTRAN for the IBM 704. "Abstracting Away The Machine" is
    an excellent reference. (Von Neumann even thought that
    assemblers were a waste of machine time...)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Fri Sep 11 08:55:00 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    George Neuner <gneuner2@comcast.net> writes:
    On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:
    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    "compilers generate better code than MOST programmers."


    The qualification is important.

    The statements I have read did not make such a qualification,
    certainly not in capital letters.

    So which statements did you read, exactly?
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 11 10:54:38 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    (I think Anton likes to claim that some people think that
    compilers are perfect. Not sure who he means, it's certainly
    not me :-)

    If this Anton is supposed to be me, I don't think I ever made such a
    claim.

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    Your statement is ambiguous in several ways.

    Yes, sorry. Given that I have no recent memory of such a statement, I
    tried to paraphrase what these statements transported to me, but
    failed in this paraphrase. Anyway the gist was along the lines of
    "Compilers are very good at optimization, better than humans".

    I assume you mean
    "programmers using assembly"

    Not particularly. It also includes optimizations at the high level.

    Do you claim that people claim

    a) Compilers generate better code than programmers all the time

    They did not qualify it like that, but that's certainly what they
    transported: "You, ordinary (but of course better than average)
    programmer, don't have a chance of outdoing a compiler at
    optimization, don't even think about it."

    Of course, one also sees statements like (from <1788739198-5857@newsgrouper.org>):

    |On modern RISC ISAs:
    |
    | for( i = 0; i < max; i++ )
    | p[i]
    |
    |is often faster than:
    |
    | for( i = 0; i < max; i++ )
    | *p++

    so overall, if one sees a statement about optimization without
    empirical support, there is a good chance that it is nonsense.

    b) Compilers generate better code than programmers most of the time
    [...]
    b) is true. In the vast majority of cases, it is vastly uneconomical
    to revert to assembly programming.

    The more usual approach is to look at the code that the compiler
    produced, and to either change the source code construct in a way that
    helps the compiler overcomes its deficiencies in optimizations, or
    (less frequently) to use some asm statement or routine written in
    assembly language. In the latter case, the assembly language code
    written by the programmer is better than the assembly language code
    generated by the compiler most of the time.

    As for the former case, a recent example is shown in <2026Sep9.153651@mips.complang.tuwien.ac.at>, where one programmer had
    written

    /* close to original */
    Scheme_Object *ADD_tagged(
    Scheme_Object *tagged_a,
    Scheme_Object *tagged_b)
    \{
    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    intptr_t r;
    Scheme_Object *o;
    r = (uintptr_t)a + (uintptr_t)b;
    o = (Scheme_Object *)
    ((((uintptr_t)r)<<1)|1);
    r = ((intptr_t )o) >> 1;
    if (b == (uintptr_t)r - (uintptr_t)a)
    return o ;
    else
    return ADD_slow (a , b) ;
    \}

    which a compiler compiled into

    sarq %rdi
    sarq %rsi
    leaq (%rdi,%rsi), %rax
    leaq 1(%rax,%rax), %rax
    movq %rax, %rdx
    sarq %rdx
    subq %rdi, %rdx
    cmpq %rdx, %rsi
    jne .L6
    ret
    .L6:
    jmp ADD_slow@PLT

    10 instructions in the common case.

    EricP suggested a better source code:

    | intptr_t a = ((intptr_t)tagged_a)>>1;
    | intptr_t b = ((intptr_t)tagged_b)>>1;
    | intptr_t r, tmp;
    | r = a + b;
    | tmp = r << 1;
    | if ((r ^ tmp) >= 0) // check if bit r[63] == r[62]
    | return tmp | 1;
    | else
    | Add_slow (a, b);

    He did not show what a compiler generates for it, but one can imagine
    that it produces better code.

    I have my own optimized source code:

    intptr_t a1=(intptr_t) tagged_a;
    intptr_t b1=(intptr_t) tagged_b;
    intptr_t r ;

    if (!__builtin_add_overflow(a1,(b1-1),&r))
    return (Scheme_Object *)r;

    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    return ADD_slow (a , b );

    On AMD64 this compiles to:

    leaq -1(%rsi), %rax
    addq %rdi, %rax
    jo .L8
    ret
    .L8:
    sarq %rsi
    sarq %rdi
    jmp ADD_slow

    4 instructions in the common case.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 11 17:50:45 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    /ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    Given that the computing world seems to have decided that it does not
    even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will
    be a widespread success.

    My 66000 architecture has means maintain robustness in the face of >Spectr|a-- attack vectors. Stashing microarchitectural state in the
    miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    This is a good approach to implement invisible speculation as far as
    cache side channel is concerned. Behnia et al. [behnia+21] describe a
    side channel that uses resource contention from speculative
    instructions to affect the timing of committing instructions, i.e., it
    uses the scheduler (reservation station) behaviour as a side channel.

    My 66000 does not have architectural constraints on the timing of instructions--each implementation gets that set of choices. And
    thanks for the reference.

    Behnia et al. also describe how to close this side channel. There may
    be other side channels through microarchitectural resources, and they
    need to be closed, too, but the cache side channel probably is the
    widest and therefore most relevant one by far.

    Caches and TLBs (which are (ARE) caches). And this is where My 66000
    has statements in the architectural documents that disallow an imple-
    mentation from modifying cache (+ TLB) state prior to the retirement
    of the causing instruction. Additionally, there is a requirement that
    prevents using a value read from memory (through the caches) being
    used as a pointer or index in a subsequent memory reference until
    the producing instruction is known to complete (access checks applied).

    @InProceedings{ behnia+21,
    author = {Mohammad Behnia and Prateek Sahu and Riccardo Paccagnella
    and Jiyong Yu and Zirui Neil Zhao and Xiang Zou and Thomas
    Unterluggauer and Josep Torrellas and Carlos Rozas and Adam
    Morrison and Frank Mckeen and Fangfei Liu and Ron Gabor and
    Christo- pher W. Fletcher and Abhishek Basak and Alaa
    Alameldeen},
    title = {Speculative Interference Attacks: Breaking Invisible
    Speculation Schemes},
    booktitle = {Architectural Support for Programming Languages and
    Operating Systems (ASPLOS |o-C-O21)},
    year = {2021},
    pages = {1046--1060},
    url = {https://dl.acm.org/doi/10.1145/3445814.3446708},
    optannote = {}
    }

    Spectre and friends are microarchitectural side channels, you can
    implement microarchitectures for any architecture that are vulnerable
    to Spectre, and microarchitectures that are not vulnerable. Therefore
    your mention of the My 66000 architecture makes no sense.

    It is the architectural statements that constrain implementations from
    updating (directly or indirectly) visible -|Arch state until it is known
    that the instruction will retire (even in the face of interrupts).

    Don't see why CHERI could not do similarly.

    Certainly one can implement a core with CHERI that is not vulnerable
    to speculative side channels (either by not implementing speculation
    (slow), or by implementing invisible speculation), and if the
    financers and the prospective customers are serious about security,
    they will insist on that for eventual products.

    Absolutely--however it appears no prospective customer yet cares too.

    But my point is that in the mainstream, not even invisible speculation
    with its moderate cost in hardware and performance is implemented, so
    I don't expect that a high-cost solution like CHERI becomes
    mainstream, in particular given that it provides nothing that cannot
    be provided more cheaply through programming languages and compilers.

    I have no complaint to this paragraph.

    Ok, you might say, what about legacy code in unsafe languages such as
    C? Rewriting all of that in Rust (or using a solution like Ivy/Deputy <http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>, which
    would be cheaper, but somehow did not catch on) would also be very
    costly.

    C says nothing about inter instruction timing. Spectr|- exploits that.
    What My 66000 does at the architectural level is to constrain all -|Arch
    from updating <potentially> visible -|Arch state.

    But I expect that much of this code does not work on CHERI in
    the secure configuration; so to run it on CHERI, you would have to
    change it anyway, so you could just as well Deputize it, which is
    cheaper in hardware and in execution time. Or rewrite it in Rust, if
    you prefer that.

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 11 17:54:52 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high >optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi-
    cation that the module is not well organized.

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Fri Sep 11 18:14:33 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    /ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    Given that the computing world seems to have decided that it does not
    even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will
    be a widespread success.

    My 66000 architecture has means maintain robustness in the face of
    Spectr|a-- attack vectors. Stashing microarchitectural state in the
    miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    This is a good approach to implement invisible speculation as far as
    cache side channel is concerned. Behnia et al. [behnia+21] describe a
    side channel that uses resource contention from speculative
    instructions to affect the timing of committing instructions, i.e., it
    uses the scheduler (reservation station) behaviour as a side channel.

    My 66000 does not have architectural constraints on the timing of >instructions--each implementation gets that set of choices. And
    thanks for the reference.

    Are there any instructions where the timing will vary based on the
    data being operated on?

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 11 23:19:32 2026
    From Newsgroup: comp.arch

    On 11/09/2026 19:54, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high
    optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi-
    cation that the module is not well organized.

    int do_big_calculation() {
    struct Data data;
    initialise_data(&data);
    for (int i = 0; i < 1'000'000; i++) {
    calculate_round(i, &data);
    }
    int result = finalise(&data);
    return result;
    }

    Are you suggesting that it would be poor organisation to have "initialise_data", "calculate_round" and "finalise" in the same module?

    Or are you suggesting that "initialise_data" should be optimised with
    the same priority on speed as "calculate_round" ?


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 11 21:59:24 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    /ERROR "unexpected byte sequence starting at index 665: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    /ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    Given that the computing world seems to have decided that it does not >> >> even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will >> >> be a widespread success.

    My 66000 architecture has means maintain robustness in the face of
    Spectr|a-a|e-- attack vectors. Stashing microarchitectural state in the >> >miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    This is a good approach to implement invisible speculation as far as
    cache side channel is concerned. Behnia et al. [behnia+21] describe a
    side channel that uses resource contention from speculative
    instructions to affect the timing of committing instructions, i.e., it
    uses the scheduler (reservation station) behaviour as a side channel.

    My 66000 does not have architectural constraints on the timing of >instructions--each implementation gets that set of choices. And
    thanks for the reference.

    Are there any instructions where the timing will vary based on the
    data being operated on?

    Things like::

    TAN, ATAN, ASIN, ACOS where 1/2 the paths do not need a reciprocal and 1/2 do.

    One could imagine MUL, DIV, FMUL, and FDIV having early out FUs. MUL and FMUL are a lot less likely to have early outs than DIV and FDIV.

    {Direct write to IP via HR, CALX, CALA, JMPX, JMPA}-instruction taking a
    cache miss versus not taking that miss.

    Invalidation {by-ASID or of-Page} to {cache or TLB}

    {assuming proper privilege:

    Direct write to Root Pointer which results in reloading of TLB state which could be cached (or not).

    Direct write to Interrupt Table control register resulting in a pending interrupt being taken (or not).

    Direct write to {Priority, voltage, clock} control registers
    }

    Otherwise: No

    But I don't think constant time evaluations need those instructions.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 11 22:05:42 2026
    From Newsgroup: comp.arch


    David Brown <david.brown@hesbynett.no> posted:

    On 11/09/2026 19:54, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high
    optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi- cation that the module is not well organized.

    int do_big_calculation() {
    struct Data data;
    initialise_data(&data);
    for (int i = 0; i < 1'000'000; i++) {
    calculate_round(i, &data);
    }
    int result = finalise(&data);
    return result;
    }

    Are you suggesting that it would be poor organisation to have "initialise_data", "calculate_round" and "finalise" in the same module?

    Or are you suggesting that "initialise_data" should be optimised with
    the same priority on speed as "calculate_round" ?

    What I am suggesting is that it might not be appropriate to use -O3
    on all 3. do_big_calculation should get -O3 while initialize_data
    and finalize might be just as well served by -O2 or even -OS.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Fri Sep 11 22:34:09 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    scott@slp53.sl.home (Scott Lurndal) posted:

    /ERROR "unexpected byte sequence starting at index 665: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    /ERROR "unexpected byte sequence starting at index 434: '\xC3'" while decoding/:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    Given that the computing world seems to have decided that it does not >> >> >> even want to pay the very moderate cost of invisible speculation to
    protect against Spectre and friends, I very much doubt that CHERI will >> >> >> be a widespread success.

    My 66000 architecture has means maintain robustness in the face of
    Spectr|a-a|e-- attack vectors. Stashing microarchitectural state in the >> >> >miss buffers until the causing instruction retires--and only then
    updating microarchitectural state.

    This is a good approach to implement invisible speculation as far as
    cache side channel is concerned. Behnia et al. [behnia+21] describe a
    side channel that uses resource contention from speculative
    instructions to affect the timing of committing instructions, i.e., it
    uses the scheduler (reservation station) behaviour as a side channel.

    My 66000 does not have architectural constraints on the timing of
    instructions--each implementation gets that set of choices. And
    thanks for the reference.

    Are there any instructions where the timing will vary based on the
    data being operated on?

    Things like::

    TAN, ATAN, ASIN, ACOS where 1/2 the paths do not need a reciprocal and 1/2 do.

    One could imagine MUL, DIV, FMUL, and FDIV having early out FUs. MUL and FMUL >are a lot less likely to have early outs than DIV and FDIV.

    So, those are all security issues. ARM has added a 'data independent timing' flag that can be used to ensure that a given instruction executes in constant time, regardless of the data being operated on.




    But I don't think constant time evaluations need those instructions.

    Hard to predict. Best to assume someone will use those instructions
    in timing sensitive code.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Sat Sep 12 02:29:02 2026
    From Newsgroup: comp.arch

    On Thu, 10 Sep 2026 17:16:37 GMT, Anton Ertl wrote:

    On Thu, 10 Sep 2026 05:41:30 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    Only on actual unification, not on simple lookup.

    Every "simple lookup" is a unification in Prolog.

    At some point, you have to hit actual values. Once a variable has a
    value, you do comparison of values.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Fri Sep 11 16:49:50 2026
    From Newsgroup: comp.arch

    Scott Lurndal [2026-09-10 19:01:44] wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    And sometimes you even want to prevent accessing data that is not
    "in" the struct due to internal alignments--such as the bytes between
    u[2] and v[0].
    That would be difficult, since one needs to be able to copy
    the structure, which necessarily will need to access those bytes.

    Of course, if the copy enjoys the same property ("you can't read the
    padding"), then it's OK to "blindly" copy all the bytes. EfOe


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From George Neuner@gneuner2@comcast.net to comp.arch on Fri Sep 11 23:25:23 2026
    From Newsgroup: comp.arch

    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    George Neuner <gneuner2@comcast.net> writes:


    Witness the proliferation of languages offering "managed environments" >>offering such niceties as automatic storage management, automatic lock >>handling (serialized object access), "comprehensions", etc., and large >>standard libraries

    I don't see any problem with that. These features help to implement >functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    Really? That's not my conclusion ... it was the result found by a
    number of university studies and developer surveys.


    Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.

    Some of those studies sought to eliminate difference in language as a
    factor by having everyone use C or C++ but with some using a GC
    library, or by using Modula 3 which permits either manual or GC
    operation.

    Similarly, multitasking with shared mutable data programmers were less
    likely to produce a correct [or complete] program using manual access
    locks vs using automatic access locking: e.g., monitors, TM, DBMS,
    etc.


    I don't have URLs handy, but over the years I have seen many of these
    studies documented in ACM journals, and I presume they have been cited
    also in other sources.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Fri Sep 11 23:30:38 2026
    From Newsgroup: comp.arch

    On 9/8/2026 1:34 AM, Anton Ertl wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    In practice, I never found it to be a problem as
    indirection was rare and I never saw more than two levels.

    Unification of a pre-existing free logical variable V with something
    (X) is implemented as making V point to X. If X is another logical
    variable, and then is unified with something, say Y, a pointer to Y is
    stored in the memory location for X. And so on. Accessing V may need
    to follow an arbitrarily long chain of pointers until you find either
    a free variable, or you find the value that V eventually was unified
    with.

    Prolog was implemented in Edinburgh on the DEC-10 (DEC-10 Prolog). I
    guess that they used the indirection feature for that: If a variable
    points to some other variable or value, its indirection bit is set, if
    it is free, it is not.

    On modern machines, i.e., without indirection bit, a free variable
    points to itself, a variable that is bound to another variable points
    to that other variable. While the type tag indicates a variable, one
    follows the pointers, until a pointer to itself is found.

    I certainly believe you, but I don't know of a Prolog implementation on
    the 1100 series, and if there were to be, based on what you say, it
    couldn't use the hardware indirection feature. And, BTW, when the architecture added "Extended Mode" in the early 1990s, which was an incompatible, but interoperable mode with what was then called "Basic
    Mode", i.e. the old stuff, (the hardware indirection feature was taken out.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 12 07:49:42 2026
    From Newsgroup: comp.arch

    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> writes:
    On Thu, 10 Sep 2026 17:16:37 GMT, Anton Ertl wrote:

    On Thu, 10 Sep 2026 05:41:30 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    Only on actual unification, not on simple lookup.

    Every "simple lookup" is a unification in Prolog.

    At some point, you have to hit actual values.

    In a single unification, no: both unified things may be free
    variables. And they can actually still be free variables when the
    resolution ends. And that does not just hold at the top-level of a unification. E.g., when I ask SWI-Prolog

    ?- x(B,C)=x(A,B).

    it answers:

    B = C, C = A.

    i.e., all three variables have been unified into one, and that is
    still free.

    Once a variable has a
    value, you do comparison of values.

    That's when both variables are bound. When the other variable is
    still free, it becomes bound to the value. The way we have done this
    in our implementation (and I think that is the usual way) is to just
    store it at the end point of a chain; that way we only need one write
    and one trail-stack entry that, on backtracking just becomes a free
    variable that is no reference to another free variable. An
    alternative would be to remember the start of a chain during
    unification, and write the value to every one of the unfified
    variables in the chain, but that would require not just an ordinary
    trail-stack entry for each of these writes, but a value-trail entry,
    because you need to remember the thing that you need to restore; it's
    not just a free variable, it's a reference to another variable.

    Value trails did not occur in the literature when I worked on a Prolog implementation in 1990, so Prolog systems at the time did not do that,
    even if it would shorten the chains. Avoiding the memory necessary
    for these value trail entries was one reason for avoiding that
    approach, but the run-time cost of all these writes may be another
    one. One would have to measure how the run-time compares for the two
    options to be sure.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 12 08:23:52 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/8/2026 1:34 AM, Anton Ertl wrote:
    Prolog was implemented in Edinburgh on the DEC-10 (DEC-10 Prolog). I
    guess that they used the indirection feature for that: If a variable
    points to some other variable or value, its indirection bit is set, if
    it is free, it is not.
    ...
    I certainly believe you, but I don't know of a Prolog implementation on
    the 1100 series, and if there were to be, based on what you say, it
    couldn't use the hardware indirection feature.

    Based on what was written here about the similar feature on the
    DEC-10, interrupts would mean that a program with a too-long chain
    would not make any progress, even if it would not be killed. The
    Univac 1100 erroring out would have been preferable in this case.

    So I guess that Prolog programs with too-long chains were rare enough
    that they lived with this problem. The limited address space of a
    DEC-10 (256KW) certainly helped here.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 12 08:38:21 2026
    From Newsgroup: comp.arch

    George Neuner <gneuner2@comcast.net> writes:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    George Neuner <gneuner2@comcast.net> writes:


    Witness the proliferation of languages offering "managed environments" >>>offering such niceties as automatic storage management, automatic lock >>>handling (serialized object access), "comprehensions", etc., and large >>>standard libraries

    I don't see any problem with that. These features help to implement >>functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    Really? That's not my conclusion ... it was the result found by a
    number of university studies and developer surveys.


    Most studies involving GC have shown that programmers working on short >timelines are less likely to produce a correct [or sometimes even just >complete] program using manual memory management vs using GC.

    Which is a different result from being "incapable of producing a
    working program".

    But that's actually not what I meant. Even if you were right, and
    some programmer is incapable of producing a working program without
    GC, and you are capable of doing it, they may be able to do things
    that you cannot do, at least not without preparation. And with
    preparation they may be capable to produce working programs without
    GC, too. But writing "the average programmer largely would be
    incapable of X", where X is something you can do, whether it is
    writing a program without GC, a program in assembly language, a
    program in Haskell, or whatever, is just a way of elevating yourself,
    by elevating X to a skill that's lifts you above the average
    programmer.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 12 09:09:35 2026
    From Newsgroup: comp.arch

    George Neuner <gneuner2@comcast.net> schrieb:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    George Neuner <gneuner2@comcast.net> writes:


    Witness the proliferation of languages offering "managed environments" >>>offering such niceties as automatic storage management, automatic lock >>>handling (serialized object access), "comprehensions", etc., and large >>>standard libraries

    I don't see any problem with that. These features help to implement >>functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    Really? That's not my conclusion ... it was the result found by a
    number of university studies and developer surveys.

    Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.

    Not surprising. Memory management is a complex and error-proine
    additional task, and if you don't have to do it, your programming
    task becomes more difficult.

    But that has been the case since the concept of subroutine libraries
    has been introduced, in the early 1950s, because people did not
    need to reinvent certain wheels such as calculating the sine of
    a floating point number or doing I/O (and let us not forget that
    reinvented wheels are often square).

    There is nothing wrong with using existing tools, in principle.

    There are many things wrong with programming today. Look at
    Microsoft Teams, which is a monster (because they built it upon
    Electron). This squanders memory like there's no tomorrow, and
    in the times of rare and expensive memory, that is a real problem.

    I found Microsoft announching that they want an 8GB Windows 11
    computer to be usable amusing. I am so glad that my new business
    laptop has 32 GB instead of 16 - with 16 it bordered on being
    unusable.

    Commercial code is quite often of very low quality. There is no
    time for good software architecture, we NEED that new feature and
    the product needs to ship NEXT WEEK, just slap it on and never
    mind if it integrated well with the existing code, or has bugs.
    But also, very many programmers these days aren't very good, so
    they need all the frameworks which add overhead and use resources.

    Brooks quoted a factor of 10 in productivity between individual
    programmers in the 1960s, I suspect that gap has widened a lot
    since then but haven't looked for literature.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 12 09:30:42 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    David Brown <david.brown@hesbynett.no> posted:

    On 11/09/2026 19:54, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high
    optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi-
    cation that the module is not well organized.

    int do_big_calculation() {
    struct Data data;
    initialise_data(&data);
    for (int i = 0; i < 1'000'000; i++) {
    calculate_round(i, &data);
    }
    int result = finalise(&data);
    return result;
    }

    Are you suggesting that it would be poor organisation to have
    "initialise_data", "calculate_round" and "finalise" in the same module?

    Or are you suggesting that "initialise_data" should be optimised with
    the same priority on speed as "calculate_round" ?

    What I am suggesting is that it might not be appropriate to use -O3
    on all 3. do_big_calculation should get -O3 while initialize_data
    and finalize might be just as well served by -O2 or even -OS.

    I don't think so. Higher optimization tries more things like
    function specialization made possible by constant propagation,
    inlining and similar. If calulate_round() considers things as
    variable that initialize_data() has as constants, or if finalize()
    does not use some things in &data which are nonetheless calculated,
    then the win can be quite substantial. (One reason why trying out
    simple benchmarks on modern compiler is like nailing a pudding to
    the wall).

    If you really want efficient code, the best way is probably to
    put them all into a single translation unit, or use LTO.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Sat Sep 12 11:32:40 2026
    From Newsgroup: comp.arch

    In article <2026Sep11.080019@mips.complang.tuwien.ac.at>, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    Ok, you might say, what about legacy code in unsafe languages such
    as C? Rewriting all of that in Rust (or using a solution like
    Ivy/Deputy
    <http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>,
    which would be cheaper, but somehow did not catch on)

    That link seems to have rotted.

    But I expect that much of this code does not work on CHERI in
    the secure configuration; so to run it on CHERI, you would have to
    change it anyway, so you could just as well Deputize it, which is
    cheaper in hardware and in execution time. Or rewrite it in Rust,
    if you prefer that.

    CHERI does not seem to have established a niche for itself. A few years
    ago, my employer was offered a Morello board for trials. We concluded
    that while it would be interesting and educational, it wasn't going to be commercially useful unless CHERI caught on. That decision seems to have
    been realistic.

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Sat Sep 12 13:18:38 2026
    From Newsgroup: comp.arch

    On 12/09/2026 00:05, MitchAlsup wrote:

    David Brown <david.brown@hesbynett.no> posted:

    On 11/09/2026 19:54, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high
    optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi-
    cation that the module is not well organized.

    int do_big_calculation() {
    struct Data data;
    initialise_data(&data);
    for (int i = 0; i < 1'000'000; i++) {
    calculate_round(i, &data);
    }
    int result = finalise(&data);
    return result;
    }

    Are you suggesting that it would be poor organisation to have
    "initialise_data", "calculate_round" and "finalise" in the same module?

    Or are you suggesting that "initialise_data" should be optimised with
    the same priority on speed as "calculate_round" ?

    What I am suggesting is that it might not be appropriate to use -O3
    on all 3. do_big_calculation should get -O3 while initialize_data
    and finalize might be just as well served by -O2 or even -OS.


    Fair enough - that makes sense, but I think it is somewhat the opposite
    of what you first wrote.

    Compilers can, to some extent, figure out that some functions are "hot"
    and some are "cold", but they need help for the details. And
    programmers rarely bother with such micro-management unless their
    efforts will lead to significant real-world payoffs. Thus they
    typically just use "-O3" (or whatever) on the whole file, even though it
    is of negligible benefit for the initialisation and finalisation.
    (Equally, it is rarely of any significant cost.)

    For some code, it may be worth using target-specific flags and a
    "resolver" function, so that different function implementations are
    called depending on the details of the target processor (such as the
    SIMD instructions it supports). gcc has support for this kind of thing
    - I have no idea how useful it is in practice. (And it also lets the
    keener assembly programmers make some versions optimised for particular targets, with a fall-back to compiler-generated code on other targets.)



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Sat Sep 12 13:23:18 2026
    From Newsgroup: comp.arch

    On 12/09/2026 11:30, Thomas Koenig wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    David Brown <david.brown@hesbynett.no> posted:

    On 11/09/2026 19:54, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    George Neuner <gneuner2@comcast.net> writes:
    ------------
    Even good programmers can have poor intuition about what needs high >>>>>> optimization.

    True, and compilers are not any better, even with profile feedback
    (rarely used). Compilers address this by trying to optimize
    everything, but they do a mediocre and unreliable job on that.

    When some subroutines in a module might want heavy optimization, while >>>> most of the rest of the subroutines in that module do not, is an indi- >>>> cation that the module is not well organized.

    int do_big_calculation() {
    struct Data data;
    initialise_data(&data);
    for (int i = 0; i < 1'000'000; i++) {
    calculate_round(i, &data);
    }
    int result = finalise(&data);
    return result;
    }

    Are you suggesting that it would be poor organisation to have
    "initialise_data", "calculate_round" and "finalise" in the same module?

    Or are you suggesting that "initialise_data" should be optimised with
    the same priority on speed as "calculate_round" ?

    What I am suggesting is that it might not be appropriate to use -O3
    on all 3. do_big_calculation should get -O3 while initialize_data
    and finalize might be just as well served by -O2 or even -OS.

    I don't think so. Higher optimization tries more things like
    function specialization made possible by constant propagation,
    inlining and similar. If calulate_round() considers things as
    variable that initialize_data() has as constants, or if finalize()
    does not use some things in &data which are nonetheless calculated,
    then the win can be quite substantial. (One reason why trying out
    simple benchmarks on modern compiler is like nailing a pudding to
    the wall).


    I just realised in my post that I had assumed Mitch was talking about
    use of __attribute__((optimize(3)) or #pragma GCC optimize "-O3" to
    control optimisation levels for individual functions. It may be that he
    is not aware of these, or that the tools he usually uses do not have equivalent fine-grained control.

    If you really want efficient code, the best way is probably to
    put them all into a single translation unit, or use LTO.


    Yes.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 12 16:42:33 2026
    From Newsgroup: comp.arch

    jgd@cix.co.uk (John Dallman) writes:
    In article <2026Sep11.080019@mips.complang.tuwien.ac.at>, >anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    Ok, you might say, what about legacy code in unsafe languages such
    as C? Rewriting all of that in Rust (or using a solution like
    Ivy/Deputy
    <http://ivy.cs.berkeley.edu/ivywiki/uploads/deputy-manual.html>,
    which would be cheaper, but somehow did not catch on)

    That link seems to have rotted.

    Works for me. Anyway, it's about annotating existing C such that the
    annotated portions would be memory-safe; the annotations insert checks
    unless the compiler can prove that they hold anyway. In this way,
    existing C code can be turned into memory-safe code in a piecemeal
    way, rather than rewriting it completely, like the Rust conversions
    are doing.

    My guess is that it did not catch on, because the programmers who
    wrote the C code do not (or maybe did not, while Ivy/Deputy was
    somewhat current, and now it's too late) see the need to do all this
    work, because they thought that their programs are memory-safe
    already, and other people prefered rewriting the code to annotating
    the legacy code of other programmers.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 12 17:09:56 2026
    From Newsgroup: comp.arch

    On Fri, 11 Sep 2026 17:54:52 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    When some subroutines in a module might want heavy optimization, while
    most of the rest of the subroutines in that module do not, is an indi-
    cation that the module is not well organized.

    Really? I would have thought it just meant that only some of the
    subroutines are called from within an inner loop.

    Of course, then the obvious optimization of _those_ subroutines is to
    turn them into inline code, since subroutine calls have a lot of
    overhead.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 12 17:15:52 2026
    From Newsgroup: comp.arch

    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:
    George Neuner <gneuner2@comcast.net> writes:
    On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    "compilers generate better code than MOST programmers."

    The qualification is important.

    The statements I have read did not make such a qualification,
    certainly not in capital letters.

    There is no doubt that a good
    programmer can beat the compiler,

    And especially this statement is usually not made. On the contrary,
    the perpetrator of such statements seem convinced of compiler
    supremacy.

    I have no doubt that a good programmer can beat the original FORTRAN
    compiler for the IBM 704 computer, despite the fact that its
    optimization was very nearly as good as that of most optimizing
    compilers until at least the late 1970s.

    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,
    but given the improvements in optimizing compilers, and the advances
    in modern computer architectures, I would still be hesitant to be
    categorical in asserting the human can always beat the optimizing
    compiler. Particularly given recent advances in AI.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Sat Sep 12 19:14:40 2026
    From Newsgroup: comp.arch

    In article <6aa5876f.783937@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:

    I would not, however, be willing to assert that one could easily
    find human programmers who could beat an optimizing compiler...
    which had the Itanium as its target.

    It was claimed at the time that Itanium assembler was quite nice to write,
    once you had wrapped your head around some of its weird idioms. Of course,
    that was a lot easier if you were writing something that had been
    considered during the ISA design.

    Intel in the late 1990s had a period of transplanting supercomputer
    idioms into their ISAs, without fully understanding that those idioms
    were only useful for particular kinds of code.

    The Pentium 4, like the Itanium, wasn't very good at general-purpose code.
    The main gain I got from it was reduced dependency between floating-point operations once FP registers weren't organised as a stack. That allowed
    OoO to work better.

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Sat Sep 12 18:34:20 2026
    From Newsgroup: comp.arch

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    Based on what was written here about the similar feature on the
    DEC-10, interrupts would mean that a program with a too-long chain
    would not make any progress, even if it would not be killed. The
    Univac 1100 erroring out would have been preferable in this case.

    I did a lot of programming on the PDP-10 and I do not ever remember it being
    a problem in practice. A buggy program might hang in an indirect or XCT loop but it was easy enough to interrupt it, start DDT (the debugger) and see what the problem was.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 12 18:14:46 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:
    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    On the contrary: The IA-64 architects sold their architecture to Intel
    and HP management with hand-written examples that made good use of the
    hardware resources, and the promise that they would write compilers
    that could produce just as good code across the board. Those
    compilers failed to appear, and that's a big part of why IA-64 failed.
    So compilers are obviously not as good as humans on IA-64 code.

    but given the improvements in optimizing compilers, and the advances
    in modern computer architectures, I would still be hesitant to be
    categorical in asserting the human can always beat the optimizing
    compiler.

    I have used the traveling salesman program from Jon Bentley's 1982
    book "Writing efficient programs" in my "Efficient programs" course
    since the late 1990s, including most of the optimizations that Jon
    Bentley has made to the source code, always with a relatively recent
    compiler. If the compilers would perform the optimizations that Jon
    Bentley performed in 1982, one would not see any change from his
    source-level optimizations.

    But actually the only one of Jon Bentley's source-level optimizations
    that did not result in a change, even with the latest gcc or clang I
    tried is function inlining. All 6 or 7 other source-level
    optimization steps by Jon Bentley resulted in a change in the
    resulting code and performance, more than 40 years after the book
    appeared.

    Or, consider the Racket 8.6 integer addition example in <2026Sep11.125438@mips.complang.tuwien.ac.at>, where I got the common
    case from 10 instructions down to 4 with a source-level optimization.
    If compilers were as great as you believe, the compiler would manage
    to perform this optimization by itself.

    Particularly given recent advances in AI.

    What about them?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 13 03:28:49 2026
    From Newsgroup: comp.arch

    On Sat, 12 Sep 2026 18:14:46 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    quadibloc@invalid.com (John Savard) writes:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    On the contrary: The IA-64 architects sold their architecture to Intel
    and HP management with hand-written examples that made good use of the >hardware resources, and the promise that they would write compilers
    that could produce just as good code across the board. Those
    compilers failed to appear, and that's a big part of why IA-64 failed.
    So compilers are obviously not as good as humans on IA-64 code.

    While this suggests that I was wrong about this, I must admit, I don't
    think that it proves the case. The compilers could have failed to
    appear for any number of reasons.
    And later on, Open64 appeared.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sun Sep 13 14:27:47 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Sat, 12 Sep 2026 18:14:46 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    quadibloc@invalid.com (John Savard) writes:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    I would not, however, be willing to assert that one could easily find >>>human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    On the contrary: The IA-64 architects sold their architecture to Intel
    and HP management with hand-written examples that made good use of the >>hardware resources, and the promise that they would write compilers
    that could produce just as good code across the board. Those
    compilers failed to appear, and that's a big part of why IA-64 failed.
    So compilers are obviously not as good as humans on IA-64 code.

    While this suggests that I was wrong about this, I must admit, I don't
    think that it proves the case. The compilers could have failed to
    appear for any number of reasons.

    Such as what? It's not as if they did not try.

    And later on, Open64 appeared.

    Later on? Open64 was first released in 2002, and its final release
    was on November 10, 2011 (couldn't they make it 2011-11-11?-). The
    compilers that were intended to make IA-64 competetive have never
    appeared, and Open64 has not changed that.

    However, concerning my statement above, I think the IA-64 architects
    used examples for which the architecture (and its in-order
    implementation) was particularly well-suited (software-pipelinable
    inner loops), and that compilers eventually usually worked ok for such examples, too. It's just that these compilers don't work so well on general-purpose code.

    But in any case, the fact that the architects sold their architecture
    with hand-written, not compiler-generated code shows that humans can
    write such code, and that it is not so easy to make a compiler do so.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Sun Sep 13 14:58:15 2026
    From Newsgroup: comp.arch

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    However, concerning my statement above, I think the IA-64 architects
    used examples for which the architecture (and its in-order
    implementation) was particularly well-suited (software-pipelinable
    inner loops), and that compilers eventually usually worked ok for such >examples, too. It's just that these compilers don't work so well on >general-purpose code.

    That's one of the problems that Multiflow also had, works great if you
    can predict the access patterns, a lot less great if the access patterns
    are data dependent.

    They also underestimated how fast out-of-order implementations would advance which handle data dependent access patterns just fine.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sun Sep 13 16:09:51 2026
    From Newsgroup: comp.arch

    John Levine <johnl@taugh.com> writes:
    They also underestimated how fast out-of-order implementations would advance >which handle data dependent access patterns just fine.

    OoO has a number of benefits over the EPIC (explicitly parallel
    instruction computing) approach of IA-64. I have the impression that
    this is not well known even among high-profile computer architects.

    In particular, in Hennessy and Patterson's computer architecture book,
    they 1) explain the advantage of OoO only with supporting more
    in-flight cache misses; and 2) only give a very superficial
    description of OoO microarchitectures. My impression is that they
    lost interest in processor cores after the RISC revolution, and now
    prefer to write more about multiprocessor interconnects and such.

    As for the advance of OoO, hardware branch prediction (essential for general-purpose code) had advanced beyond the capabilities of compiler
    branch prediction (which compiler-based speculation supported by EPIC
    was planned to rely on) in 1991 or so, so EPIC was at a disadvantage
    there. Intel also knew that they designed the Pentium 4 (released
    2000) with a 128-entry reorder buffer. And looking at SPEC CINT 2000,
    despite their huge caches, IA-64 implementations never outdid
    contemporaneous IA-32 or AMD64 implementations:

    System Cint CFP
    res base res base CPU Tested Published Intel D850GB 656 640 714 704 Pentium 4 2000MHz Aug-2001 Sep-2001 hp rx4610 --- 379 715 715 Intel Itanium 800MHz Aug-2001 Sep-2001 Dell Precision WS 340 922 893 901 878 Pentium 4 2533MHz May-2002 Jun-2002 hp workstation zx6000 --- 807 1356 1356 Itanium 2 1000Mhz Jul-2002 Jul-2002

    IA-64 implementations were good at CFP, at least initially, though.

    Intel should have been aware of the Pentium 4's capabilities early on.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From kegs@kegs@provalid.com (Kent Dickey) to comp.arch on Sun Sep 13 17:21:05 2026
    From Newsgroup: comp.arch

    In article <2026Sep10.190203@mips.complang.tuwien.ac.at>,
    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but
    certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[5];
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a,
    sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does
    things beyond a single nested struct, such as garbage collection.

    I guess there are ways to do these things in CHERI, but I very much
    doubt that the only cost CHERI has for them is the doubled memory >consumption. And does it buy anything that Rust does not buy us?

    This is off the top of my head, so it will be "wrong", but you'll get the
    idea.

    If you want to evaulate:

    out = a[i].y[j].v[k];

    You have a pointer to the start of the 'a' array, a, and it allows access to sizeof(foo)*3 bytes: a.ptr = a, a.bound=3*sizeof(foo)

    You do (each line is one instruction, everything is registers):

    Then the evaluation is:

    tmp = sizeof(foo)
    off = i*tmp + offset(a[0].y[0])

    a_i.ptr = a + off, a_i.bound=sizeof(bar)*4
    [ The above checks that a_i is within the range a allows, then
    creates a new pointer with the smaller range ]

    tmp = sizeof(bar)
    off = j*tmp + offset(y[0].v[0])

    y_j.ptr = a_i + off, y_j.bound = 3
    [ The above checks that y_j is within the range a_i allows ]

    out = LOAD(y_j + k)
    [ The above checks that y_j+k is within the y_j range ]

    This checks that each lookup is within bounds. The CHERI overhead is the formation of a_i and y_j, the other address calculations are needed even without CHERI. So 6 instructions become 7 instructions. What helps CHERI
    is this type of indexing takes several instructions on many architectures already, so adding the pointer creation steps sometimes are just replacing
    an ADD that was needed anyway.

    [Note: the above is more efficient than what Morello does, where the
    setting of the bounds is a separate instruction, but that was a poor choice
    for the implementation: It's pretty easy to implement PTR = PTR + REG, with
    a small immediate encoding a new bounds, and PTR = PTR + REG, size=REG2 where the new size comes from another register. ARM has the encoding space
    available for this, but CHERI is so much more than just bounds checking
    they lost the forest for the trees and added like 160+ instructions].

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 13 17:34:27 2026
    From Newsgroup: comp.arch


    John Levine <johnl@taugh.com> posted:

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    However, concerning my statement above, I think the IA-64 architects
    used examples for which the architecture (and its in-order
    implementation) was particularly well-suited (software-pipelinable
    inner loops), and that compilers eventually usually worked ok for such >examples, too. It's just that these compilers don't work so well on >general-purpose code.

    That's one of the problems that Multiflow also had, works great if you
    can predict the access patterns, a lot less great if the access patterns
    are data dependent.

    I believe that this will always be the case for memory references.
    When striding through memory accessing doublewords 1 cache miss
    serves 8 LDDs 7 get hits 1 takes a miss. How does one software
    schedule for that pattern ??

    On the other hand this is the planned pattern making Reservation
    Stations (dispatch stacks, and scoreboards) near optimal. SW has
    to do nothing to get the most out of the core !!

    They also underestimated how fast out-of-order implementations would advance which handle data dependent access patterns just fine.

    Pentium Pro should have given them a good clue.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 13 17:38:10 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    John Levine <johnl@taugh.com> writes:
    They also underestimated how fast out-of-order implementations would advance >which handle data dependent access patterns just fine.

    OoO has a number of benefits over the EPIC (explicitly parallel
    instruction computing) approach of IA-64. I have the impression that
    this is not well known even among high-profile computer architects.

    In particular, in Hennessy and Patterson's computer architecture book,
    they 1) explain the advantage of OoO only with supporting more
    in-flight cache misses; and 2) only give a very superficial
    description of OoO microarchitectures. My impression is that they
    lost interest in processor cores after the RISC revolution, and now
    prefer to write more about multiprocessor interconnects and such.

    As for the advance of OoO, hardware branch prediction (essential for general-purpose code) had advanced beyond the capabilities of compiler
    branch prediction (which compiler-based speculation supported by EPIC
    was planned to rely on) in 1991 or so, so EPIC was at a disadvantage
    there. Intel also knew that they designed the Pentium 4 (released
    2000) with a 128-entry reorder buffer. And looking at SPEC CINT 2000, despite their huge caches, IA-64 implementations never outdid
    contemporaneous IA-32 or AMD64 implementations:

    System Cint CFP
    res base res base CPU Tested Published
    Intel D850GB 656 640 714 704 Pentium 4 2000MHz Aug-2001 Sep-2001
    hp rx4610 --- 379 715 715 Intel Itanium 800MHz Aug-2001 Sep-2001
    Dell Precision WS 340 922 893 901 878 Pentium 4 2533MHz May-2002 Jun-2002
    hp workstation zx6000 --- 807 1356 1356 Itanium 2 1000Mhz Jul-2002 Jul-2002

    IA-64 implementations were good at CFP, at least initially, though.

    It had 2|u the number of pins for DRAM than P4. That goes a long way in
    SpecFP (way back when on-die cache was constrained...)

    Intel should have been aware of the Pentium 4's capabilities early on.

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From kegs@kegs@provalid.com (Kent Dickey) to comp.arch on Sun Sep 13 17:50:19 2026
    From Newsgroup: comp.arch

    In article <eAAoS.242$3vg8.225@fx37.iad>,
    Scott Lurndal <slp53@pacbell.net> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    kegs@provalid.com (Kent Dickey) posted:
    But CHERI has some holes (just off the top of my head, I'm sure
    there's more):

    - It needs to disallow capabilities in shared memory from creating security >>> holes. One process mmap()'s some memory, and stores capabilities
    pointing to its private memory. Then, another process mmaps the
    same shared memory, and now can use those valid capabilities to
    access the same addresses in it's address space, which it might
    not have access to! This is a tricky problem to solve since you
    want shared libraries to work.

    Sort-of defeaters the whole purpose, does it not ??

    If that had been an accurate summarization of how CHERI interacts
    with mmap (and shared libraries in general), perhaps.

    However, if you think that the CHERI team hasn't given significant
    though to the issues of inter-thread and inter-process memory sharing;
    you might want to re-read the CHERI documentation.

    I've already stated my primary concern with CHERI: that it COULD be secure,
    but it pushes a lot of complexity into other parts of the system, and it
    needs to do that in a clear and open fashion. CHERI is NOT clear and open about how to do these things. It's obtuse and terse, and that's a trick
    used to hide problems.

    Could you explain to us how CHERI completely solves the mmap() aliasing
    issue?


    - ARM allows a user process to be big-endian. This similarly can break
    the capability security model through mmap() and other means.

    I don't see how. A capability is opaque to software and completely >independent of the endianness of the current thread. It must be 8-byte >aligned (which reduces the number of tag bits to 1/8th).

    Morello solved this with one line: Capabilities in memory are always
    little endian. I told them to fix this, and they did fix this one.

    - There's a whole slew of DMA-related security issues which seem hard to >>> fully plug. You need to allow paging of capabilities, and this
    creates attack surfaces, or just complexity for real use cases
    where you don't want to clear the valid bit. All DMA drivers
    become part of the attack surface for forging capabilities.

    This simply makes PCIe devices impossible as each "address plus size"
    has to be a capability.

    Again, one must read the CHERI documentation before commenting on
    such things.

    "CHERI (Capability Hardware-Enhanced RISC Instructions) interacts
    with Direct Memory Access (DMA) by treating peripheral devices as
    capability-unaware entities and enforcing hardware tag-clearing
    mechanisms at the memory/interconnect level to protect capability
    integrity."

    CHERI creates new system requirements. You could have DMA maintain capabilities, and that has some nice properties, but then means you have
    to trust your DMA device. The CHERI documentation says whether DMA
    should be able to read/write capabilities is an open question, but that
    the current default is that DMA writes clear capabilities.

    So what's the hole with that? To support paging, it means you must have
    a mechanism to write in the page data to memory using DMA, then DMA
    in the tags elsewhere, and then write in the tags to be valid separately (probably done by a CPU using special instructions).

    And this lack of atomicity creates a security issue (which again, can be
    fixed, but you have to DOCUMENT it clearly to make sure it's handled).
    When allowing any DMA to user pages, the lack of atomicity means the OS
    must completely unmap the page from the use before doing any tag operations. Otherwise, the case that happens is there are 2 threads, a page is paged out, one thread touches that page, and starts in the page-in procedure. The
    OS is lazy and has let the second thread have access to the page (because
    this is NOT a security hole without CHERI, it's the process's private data page, if it wants to "corrupt" it, it's no concern of the OS's). DMA writes
    in the data, and then an OS task waits for DMA to finish, then it sets the
    tag bits. Meanwhile, the second thread could be aggressively writing
    to a pointer on the page, changing the bounds to be the whole memory space. Then the OS process sets the tag to validate the capability, and then
    we've let the user forge a capability.

    There's a similar race in paging out--if the OS copies the tags first,
    then does the DMA reads, and it's mapped in the second user thread still,
    it can similarly forge capabilities. If DMA is first and OS copies the
    tags second, the user could set up the page with a forged pointer that is not
    a valid capability, and then write in a dummy capability before the OS
    copies the tags. So the data is the forged pointer, and the tag is valid.
    When paged back in, the page has a valid forged capability.

    So, CHERI adds a requirement that any page being used for DMA must be
    fully unmapped from the user address space first. Again, some OS'es may
    do this by default, but this is a security problem with CHERI and it must
    be done properly. The whole idea of CXL is to allow DMA right into
    user space, and that's not very compatible with CHERI.

    If you think of valid tag bits as a virus that needs to be contained, you'll see there are MANY new requirements for how the OS needs to handle user
    data. There needs to be no path where the user creates a malicious
    capability, and then an OS operation copies what it thought was data as
    a capability. The OS must not accidentally create capabilities in the
    user space. Morello's all-new-instructions actually help solve this,
    by making normal integer instructions clear the tag always. But there are
    many possibilities here, and since I can find no discussion of how they
    think they've solved this, I'm doubtful they've covered everything.

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 13 18:00:11 2026
    From Newsgroup: comp.arch


    kegs@provalid.com (Kent Dickey) posted:
    ------------------

    Inserting minor corrections:

    This is off the top of my head, so it will be "wrong", but you'll get the idea.

    If you want to evaulate:

    out = a[i].y[j].v[k];

    -------------------
    Rewriting using variable-types instead of base-types::

    You do (each line is one instruction, everything is registers):

    Then the evaluation is:

    tmp = sizeof(a[0]);
    off = i*tmp + offset(a[0].y[0]);

    a_i.ptr = a + off, a_i.bound=sizeof(a[0].y);
    --------

    tmp = sizeof(a[0].y[0]);
    off = j*tmp + offset(y[0].v[0]);

    y_j.ptr = a_i + off, y_j.bound = Sizeof( a[0].y[0].v );
    --------

    k_tmp = k*sizeof( a[0].y[0].v[0] );
    out = LOAD(y_j + k_tmp);

    The types carry their sizes, no need to use base types.
    AND I am assuming the + in LOAD() is non-scaled.
    --------

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 14 00:42:08 2026
    From Newsgroup: comp.arch


    kegs@provalid.com (Kent Dickey) posted:

    In article <eAAoS.242$3vg8.225@fx37.iad>,
    Scott Lurndal <slp53@pacbell.net> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    --------------------

    CHERI creates new system requirements. You could have DMA maintain capabilities, and that has some nice properties, but then means you have
    to trust your DMA device.

    No, it means devices need to address memory using capabilities--
    good luck with that {since 99.99% of devices now use PCIe}.

    The CHERI documentation says whether DMA
    should be able to read/write capabilities is an open question, but that
    the current default is that DMA writes clear capabilities.

    Why should a device be able to read/write capabilities and cores cannot--
    seems like an attack vector waiting to happen.

    So what's the hole with that? To support paging, it means you must have
    a mechanism to write in the page data to memory using DMA, then DMA
    in the tags elsewhere, and then write in the tags to be valid separately (probably done by a CPU using special instructions).

    Back in 1974, we used 360/67 to perform DMA and got system privilege
    {since the OS got so bloated, TSS put some OS data in page 0 of our
    thread.} I, personally, would not think it is secure to allow DMA to
    manipulate capabilities--we don't allow them to manipulate MMU tables.

    And this lack of atomicity creates a security issue (which again, can be fixed, but you have to DOCUMENT it clearly to make sure it's handled).
    When allowing any DMA to user pages, the lack of atomicity means the OS
    must completely unmap the page from the use before doing any tag operations.

    Why would an application be accessing a page that is currently undergoing
    DMA ?? Seems like a direct race condition. If app is writing a page getting written, what gets to SSD ? if app is reading a page getting written by DMA what gets read ??

    Otherwise, the case that happens is there are 2 threads, a page is paged out, one thread touches that page, and starts in the page-in procedure. The
    OS is lazy and has let the second thread have access to the page (because this is NOT a security hole without CHERI, it's the process's private data page, if it wants to "corrupt" it, it's no concern of the OS's). DMA writes in the data, and then an OS task waits for DMA to finish, then it sets the tag bits. Meanwhile, the second thread could be aggressively writing
    to a pointer on the page, changing the bounds to be the whole memory space. Then the OS process sets the tag to validate the capability, and then
    we've let the user forge a capability.

    Illustrating but one race condition...

    There's a similar race in paging out--if the OS copies the tags first,
    then does the DMA reads, and it's mapped in the second user thread still,
    it can similarly forge capabilities. If DMA is first and OS copies the
    tags second, the user could set up the page with a forged pointer that is not a valid capability, and then write in a dummy capability before the OS
    copies the tags. So the data is the forged pointer, and the tag is valid. When paged back in, the page has a valid forged capability.

    So, CHERI adds a requirement that any page being used for DMA must be
    fully unmapped from the user address space first. Again, some OS'es may
    do this by default, but this is a security problem with CHERI and it must
    be done properly. The whole idea of CXL is to allow DMA right into
    user space, and that's not very compatible with CHERI.

    It is a security problem with or without CHERI.

    If you think of valid tag bits as a virus that needs to be contained, you'll see there are MANY new requirements for how the OS needs to handle user
    data. There needs to be no path where the user creates a malicious capability,

    at a minimum

    and then an OS operation copies what it thought was data as
    a capability. The OS must not accidentally create capabilities in the
    user space. Morello's all-new-instructions actually help solve this,
    by making normal integer instructions clear the tag always. But there are many possibilities here, and since I can find no discussion of how they
    think they've solved this, I'm doubtful they've covered everything.

    Kent
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 14 06:22:02 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:
    So, those are all security issues. ARM has added a 'data independent timing' >flag that can be used to ensure that a given instruction executes in constant >time, regardless of the data being operated on.

    Interesting approach. So when you use this flag on a load, it will
    take the maximum time that a load can take (e.g., RAM access to RAM
    under contention, or access to a remote cache under contention),
    wherever it finds the item?

    Intel has taken a different approach: It has documented the
    instructions that are guaranteed to have data-independent timings
    across all Intel microarchitectures; not sure if AMD gives the same
    guarantee. So if you want to write code that does not have a timing side-channel that reveals something about the data, you use only these instructions.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 14 06:35:34 2026
    From Newsgroup: comp.arch

    kegs@provalid.com (Kent Dickey) writes:
    In article <2026Sep10.190203@mips.complang.tuwien.ac.at>,
    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    kegs@provalid.com (Kent Dickey) writes:
    A CHERI-lite just providing bounds checking could be useful.

    I doubt it. The problems I see here (as well as in CHERI) is that we
    have nested data structures (not in every programming language, but >>certainly in some relevant ones), such as

    struct foo {
    char x[3];
    struct bar {
    char u[5];
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    Sometimes you want to check that you do not exceed the bounds of a, >>sometimes that you do not exceed the bounds of z, sometimes that you
    do not exceed the bounds of u. Sometimes you want to treat all of a
    as one thing, sometimes, only some part, sometimes your software does >>things beyond a single nested struct, such as garbage collection.
    [...]
    If you want to evaulate:

    out = a[i].y[j].v[k];

    You have a pointer to the start of the 'a' array, a, and it allows access to >sizeof(foo)*3 bytes: a.ptr = a, a.bound=3*sizeof(foo)

    You do (each line is one instruction, everything is registers):

    Then the evaluation is:

    tmp = sizeof(foo)
    off = i*tmp + offset(a[0].y[0])

    a_i.ptr = a + off, a_i.bound=sizeof(bar)*4
    [ The above checks that a_i is within the range a allows, then
    creates a new pointer with the smaller range ]

    tmp = sizeof(bar)
    off = j*tmp + offset(y[0].v[0])

    y_j.ptr = a_i + off, y_j.bound = 3
    [ The above checks that y_j is within the range a_i allows ]

    out = LOAD(y_j + k)
    [ The above checks that y_j+k is within the y_j range ]

    This checks that each lookup is within bounds. The CHERI overhead is the >formation of a_i and y_j, the other address calculations are needed even >without CHERI. So 6 instructions become 7 instructions.

    Let's see (annotations in the assembly code by me

    [/tmp:171793] cat xxx.c
    #include <stddef.h>

    struct foo {
    char x[3];
    struct bar {
    char u[5];
    int v[3];
    } y[4];
    long z[5];
    } a[3];

    int bar(size_t i, size_t j, size_t k)
    {
    return a[i].y[j].v[k];
    }
    [/tmp:171798] clang -Wall -O -S xxx.c
    [/tmp:171799] cat xxx.s
    [...]
    shlq $7, %rdi #i1=i*sizeof(foo)
    leaq a(%rip), %rax #tmp = a
    addq %rdi, %rax #a_i = tmp+i1
    leaq (%rsi,%rsi,4), %rcx #tmp2 = j*5
    leaq (%rax,%rcx,4), %rax #tmp3 = a_i + tmp2*4
    movl 12(%rax,%rdx,4), %eax #tmp4 = mem[tmp3+k*4+
    # offsetof(struct foo, y)+
    # offsetof(struct bar,v)
    retq #return tmp4
    [...]

    [gcc produces code with the same number of instructions, but where the
    address computations are more mixed, with even more liberal use of the associative and distributive laws].

    I wanted to map the assembly instructions directly to your code, but
    the structure is sufficiently different that I did not find a way.

    So let's look at your code:

    tmp = sizeof(foo)
    off = i*tmp + offset(a[0].y[0])
    a_i.ptr = a + off, a_i.bound=sizeof(bar)*4
    tmp = sizeof(bar)
    off = j*tmp + offset(y[0].v[0])
    y_j.ptr = a_i + off, y_j.bound = 3
    out = LOAD(y_j + k)
    [ The above checks that y_j+k is within the y_j range ]

    You use an instruction with two source register operands and an
    immediate operand. Ok for AMD64, but other instruction sets avoid
    instruction formats with such requirements.

    For your address computations with bounds, the associative and
    distributive laws don't hold, so in addition to the additional
    hardware required for holding and computing the bounds, you also
    cannot use all the transformations that clang and gcc are using in the
    code above.

    Anyway, that's the easy case for the bounds checking. What if you
    have:

    bla = offsetof(struct foo,y)+sizeof(struct bar)*2+offsetof(struct bar, u);
    blub = offsetof(struct foo,x);
    char *addr = a[2]
    if (cond)
    c = addr[bla+i];
    else
    c = addr[blub+i];

    And arrange enough interference (e.g., indirect calls or returns) that
    using code duplication or somesuch to preserve knowledge becomes
    impossible or at least impractical. How does architectural bounds
    checking deal with that.

    You might ask, how doing it at the language level would deal with
    that. The programmers will have to find a way to explain what they
    want to do in a way that's compatible with the bounds-checking
    programming language. I am not sure how that would be done in Rust;
    at worst one can chicken out and use unsafe Rust.

    For something like Ivy/Deputy, I expect that a solution would be for
    the programmers to turn bla and blub into structs that contain the
    address, and type/bounds information (not sure if Deputy supports that
    also for types, but one can image that an updated Deputy does), and
    then have an annotation that i must be within the bounds for the type
    speficied in bla or blub.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Mon Sep 14 16:24:10 2026
    From Newsgroup: comp.arch

    George Neuner wrote:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    George Neuner <gneuner2@comcast.net> writes:


    Witness the proliferation of languages offering "managed environments"
    offering such niceties as automatic storage management, automatic lock
    handling (serialized object access), "comprehensions", etc., and large
    standard libraries

    I don't see any problem with that. These features help to implement
    functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    Really? That's not my conclusion ... it was the result found by a
    number of university studies and developer surveys.


    Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.

    I have written multi-threaded programs since at least 1992 or so,
    preferably using LOCK XADD as the only synch mechanism, and I still get
    it wrong so often that for my latest effort in this area, I decided to
    go all the way to the file system:

    Every thread works on a completely separate set of input and output
    files, can optionally send progress reports back via UDP packets.

    This model is so very, very simple that I feel confident I'm not messing
    it up. :-)

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Mon Sep 14 16:31:01 2026
    From Newsgroup: comp.arch

    John Savard wrote:
    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:
    George Neuner <gneuner2@comcast.net> writes:
    On Thu, 10 Sep 2026 17:22:37 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    However, I have seem many cases where people made a claim that
    compilers generate better code than programmers.

    "compilers generate better code than MOST programmers."

    The qualification is important.

    The statements I have read did not make such a qualification,
    certainly not in capital letters.

    There is no doubt that a good
    programmer can beat the compiler,

    And especially this statement is usually not made. On the contrary,
    the perpetrator of such statements seem convinced of compiler
    supremacy.

    I have no doubt that a good programmer can beat the original FORTRAN
    compiler for the IBM 704 computer, despite the fact that its
    optimization was very nearly as good as that of most optimizing
    compilers until at least the late 1970s.

    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,

    Huh???

    I studied the Itanium architecture manuals and amused myself by writing
    asm kernels that took good advantage of the large register set,
    including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working variable
    live on the 8-element stack, is significantly harder IMHO.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 14 14:43:28 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    John Levine <johnl@taugh.com> posted:

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    However, concerning my statement above, I think the IA-64 architects
    used examples for which the architecture (and its in-order
    implementation) was particularly well-suited (software-pipelinable
    inner loops), and that compilers eventually usually worked ok for such
    examples, too. It's just that these compilers don't work so well on
    general-purpose code.

    That's one of the problems that Multiflow also had, works great if you
    can predict the access patterns, a lot less great if the access patterns
    are data dependent.

    I believe that this will always be the case for memory references.
    When striding through memory accessing doublewords 1 cache miss
    serves 8 LDDs 7 get hits 1 takes a miss. How does one software
    schedule for that pattern ??

    Don't use cache :-)

    There are several startups working on dense (10x or more) SRAM
    implementations with 1ns access times as an HBM replacement.

    With 1ns latency, who needs cache?

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 14 15:44:57 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:
    There are several startups working on dense (10x or more) SRAM >implementations with 1ns access times as an HBM replacement.

    With 1ns latency, who needs cache?

    Only L1 caches have 1ns access latency, not even per-core L2 caches.
    There is no way a large memory will provide 1ns latency, not to one
    core, and certainly not to multiple cores.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Sun Sep 13 17:12:11 2026
    From Newsgroup: comp.arch



    On 9/9/26 7:25 PM, MitchAlsup wrote:

    I am going to split my responses into multiple posts.

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    I view that as the software developers problem. If the hardware
    interface clearly specifies the tradeoffs, the software
    developers have the responsibility (in my opinion) to make
    reasonable decisions and live with the consequences.

    Humans being human, there will be complaints like "Why did you
    let me shoot myself in the foot?" just as fixing space bar
    heating might ruin someone's workflow.

    [snip]

    Yes, but You understand the difference between Architecture and implementation. May designers do not (especially early in their
    careers.)

    There are many compatibility layers. Socket compatibility
    matters less now for CPUs, but drop-in upgrades used to be a big
    deal. However, Socket 7 forever would not have been a good idea.
    (Of course, software does not care about that compatibility
    layer.)

    (This can reduce microarchitectural flexiblity if the software
    is considered critical enough and rewriting/retuning is not an
    option.)

    Which is why you should strive to eliminate those things before
    SW gets written.

    Which then reduces software flexibility. TANSTAAFL

    Masking the high bits in software would also have compatibility
    issues (besides adding an instruction and using an extra
    register if the tag value it to be retained). (At least the
    convention of "positive" application space addresses avoids the
    need to sign extend the address. Sign extending from an
    arbitrary bit position does not seem to be commonly available as
    a single instruction, though masking upper bits may also be
    expensive in most ISAs.)

    If masking was going to be done, putting the tag in the least
    significant bits would allow a shift to generate the address.
    My 66000 has a 64-bit bit insert instruction (32-bit immediate)
    that could extract an address, allowing alignment bits to be
    used for the tag. (Two 6-bit immediates and three 5-bit register
    specifiers would mean using four major opcode slots for an
    uncommon operation to provide a 32-bit instruction. An explicit
    extract instruction could fit by assuming the zero register as
    the SRC1, but that also seems wasteful of the encoding space.)

    There is also the issue of what hardware optimizes. E.g.,
    abysmally slow floating point denormals might encourage use
    of a flush-to-zero configuration option. Implementing slow
    denormals but not providing a flush-to-zero option seems likely
    to upset users, but not having the extra hardware to provide
    fast or only slightly slow denormals (in the 1980s) might have
    seemed justified.

    Software that uses flush-to-zero might behave noticeably
    differently (besides lacking bit identical results) when run
    on a system with just full-speed denormals.

    Software that chooses to use such address bits for tags will
    still run on newer designs with full 64-bit virtual addresses
    though limited to 56-bit virtual addresses. I think that most
    software will never need 64-bit virtual addresses. I also think
    that a lot of software would not bother using address tagging.

    Illustrating the small usage envelope of that mis-feature.

    I think the feature is likely to have a small amount of software
    that uses it, but a large amount of use. E.g., web browser
    javascript execution might use such a feature. There are not
    many javascript runtimes in use, but they are an important
    category of software.

    I view tagged addresses as something useful to software that
    hardware can provide cheaply.

    For some uses, class-based region allocation might be adequate
    at the cost of some sparsity. With only 48 bits being reliably
    available, using valid addresses for such tags might be
    difficult. This also introduces software fragility where the
    software cannot use all the address bits because some systems
    will not allow such addresses.

    (The use of the tags for detecting many accidental misuses of
    memory may also be significant. I am inclined to believe that
    use after free and buffer overruns would be better handled in
    software, but providing a little protection in a secondary
    layer rCo allowing poorly developed software to run a little
    more securely rCo may be valuable.)

    More posts to follow.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Sun Sep 13 17:26:54 2026
    From Newsgroup: comp.arch

    On 9/9/26 7:25 PM, MitchAlsup wrote:
    [snip]
    So, send me an e-mail address and I can send you ISA-55 and SFT-55

    I may well send you an email within a year asking for this.
    Right now my energy and attention levels seem to be unusually
    low and I want to be able to read with understanding and
    appreciation.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Sun Sep 13 18:27:56 2026
    From Newsgroup: comp.arch

    On 9/9/26 7:25 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:

    [big snip]

    Side note: I have felt some attraction to a virtual Harvard
    system, i.e., a separate virtual address space for instructions.

    Leads to problems for debuggers and for JIT.

    I am not certain why JIT would be that problematic. One would
    need the ability to store into instruction space, but making
    such stores explicit seems desirable (a bit like having pages
    that are thread-local and not needing external coherence).

    Besides making writing to code more explicit and slightly
    increasing the address space (less than a bit as code is rarely
    half of the used memory for larger systems), such might present
    opportunities to use the instruction pointer to generate data
    addresses that are not cluttered with instructions.

    Mc88110 had separatable Code and Data, nobody used it.

    The cache chips each had their own MMU+TLB so one could go full
    Harvard is they choose.

    It was not clear from the little documentation I read if such
    was intended to be an architectural feature (i.e., would the
    software effort to exploit it be wasted in the next generation).

    The Fairchild Clipper had a similar C[A]MMU design.

    I thought both were attractive in providing set associative
    caches with modest pin counts and chip size limits.

    I did notice the possibility of having different page table
    pointers in the different CMMUs, but it was not clear if
    such was intentional or just a side effect of having independent
    interfaces for data and instructions.

    The Clipper CAMMU design also had some hardwired TLB entries,
    mainly to simplify boot-up; TLB entries with fixed virtual
    addresses and even fixed physical addresses (and other metadata)
    would be cheap. It is not clear that providing, e.g., one way of
    few (or zero) tag bits would have been useful; such would use
    less area but would also be of limited use.

    I thought of the possibility of using negative offsets from a
    masked instruction pointer to provide a cheap pointer usable
    with shorter offsets. Separate instruction and data address
    spaces would allow

    ???

    Sorry, I either accidentally deleted the rest of that statement
    or (much more likely) looked at something else and forgot where
    I was.

    My point was that separate address spaces would allow positive
    offsets from a masked instruction pointer without the risk of
    running into code.

    I do think being able to use relatively short offsets from an
    IP-based address could be useful. However, I doubt most
    programs have a large number of static variables.

    I have also wondered about the even odder "feature" of
    potentially exploiting "pseudo-constants" that are hoisted into
    the front-end more like immediates. Although such could
    facilitate lower effective data access latency for some loads,
    I suspect such loads are to rare to justify the effort and
    forwarding updates (even if rare being pseudo-constants) would
    add complexity.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Sun Sep 13 20:29:27 2026
    From Newsgroup: comp.arch

    On 9/9/26 7:25 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I do appreciate those features of My 66000 page tables. I do
    wonder if using large pages in the page table itself might be
    worth supporting (Andy Glew had suggested such).

    I expect that either Guest OS will be supplied with large pages
    from Host OS, or vice versa. One simplifies Guest OS MMU management;
    the other simplifies Host OS management. My philosophy, here, is to
    allow SW e-the-large to figure it out for themselves (over time).

    I was referring to the later "clarified" level merging. For
    dense address space use (but with enough absent pages or enough
    address scattering that using large pages for leaf nodes/real
    pages is impractical) having an 8 MiB (two level) node might
    not cost more storage (possibly even saving a whopping 8 KiB!ry|)
    and provide on level shallower look-up.

    With page level "merging" (large pages), one could (in theory)
    have 5-bit levels that could support multiples of 32 in page
    size while not requiring the page table depth to be doubled.
    256-KiB pages might be a useful option between 8 MiB and 8 KiB.

    A PTE mapping a large page in My 66000, uses the unused PTE.PA
    bits as a limit. So, while the large page still has large page
    alignment; it only needs to contain as many pages as required.
    So if you are using 23-bit large pages, you can use one PTE
    entry (in TLB) to access exactly 173 of those pages. ...

    This is a very nice feature for sparse or partially used pages,
    but it is distinct from using larger pages in the page table
    itself.

    Such also supports more flexible tradeoffs of depth versus
    internal fragmentation. An application using a large address
    space might have little sparsity in the upper address bits and
    so benefit from merging top page table levels.

    The LVL field can be used to skip the top bits (to VAS size),
    then used to skip intermediate levels (sparse VAS), and then
    skip the lower levels (super pages). All in one set of bits.

    Support for sparsity is very nice, but improving performance for
    dense-ish high memory use also seems worth pursuing.

    Level merging (really splitting) might also support more
    flexible sharing of page tables. With 5-bit basic levels, an
    aligned 256-byte section of a 8 KiB page table block could be
    shared without sharing the rest of the translations or
    permissions. (I have no idea if such would actually be useful
    much less worthwhile, but it is possible.)

    Hard on the TLB.

    The TLB proper already has to support multiple page sizes. While
    doubling the number of page sizes supported would add
    difficulties, it would also provide more flexibility in page
    sizes. (My 66000's ability to disallow access to a part of a
    large page does reduce the requirement for finer page
    granularity at least somewhat, though having many huge pages
    that are only 3% used seems sub-optimal. The added used size
    indicator (if that applies to pages proper) presumably would
    also complicate the TLB.)

    Since some ISAs support an intermediate page size (via page
    grouping, e.g.), TLBs would have to deal with that complication
    (if the implementation supports that page size feature).
    Splitting large pages is a straightforward mechanism, though it
    has the complicated tradeoff about how much to prefetch into the
    TLB.

    For page table traversing and intermediate node caching, having
    variably sized levels would add complexity. The traversal
    complexity does not seem likely to be that much worse than for
    skipping levels. The caching complexity seems substantially
    worse and similar to TLB complexity.

    Although one could trivially provide N cache entries for every
    5-bit index-length-change level, this would obviously be
    wasteful given that the utilization would not be balanced.
    With hash-rehash it would be possible for the alternative
    index to be used as a kind of victim cache/L2. The 32-fold
    size difference is also small enough that with a large number of
    sets Andr|- Seznec's method for supporting multiple page sizes in
    TLBs could be used ("Concurrent Support of Multiple Page Sizes
    On a Skewed Associative TLB", 2003); for two sizes one needs one
    indexing bit that is shared by both sizes. Sharing entries does
    cost some bits compared to dedicated size/level entries.

    I _suspect_ the hardware complexity is not insurmountable and
    that the question is whether the effort is justified.
    Intermediate larger page sizes seem attractive.

    Merging two 10-bit levels into one 20-bit level would not seem
    to add much complexity to the node caching. The index and tag
    bits would be the same as for the first of the 10-bit levels and
    the data portion would have unused bits (being 8 MiB aligned
    rather than 8 KiB aligned).

    (Andy Glew had the idea of merging page table levels. Using a
    smaller, e.g., 5-bit, indexing fragment was my own weirdness,
    seeking to provide even more flexibility by exploiting level
    merging.)

    Level merging seems to be especially useful for nested page
    tables in virtualization. The virtual-physical address space
    is dense (I think). A virtual machine monitor might give huge
    pages to a guest but could also benefit from a flatter upper
    portion of the page table.

    There is currently a debate going on as to whether Guest OS should
    page its own pages or just have Host OS do it; Guest OS page fault
    handler just never gets control.

    I think the Guest OS has potentially more knowledge of what
    value the presence of its pages have, but the Host OS has more
    knowledge of what system pressures are present. A particular
    workload might be 10% slower with 25% less memory consumption
    but 50% slower with 40% less memory consumption. Just like
    compute can use priority to "buy" a larger fraction of work,
    being able to communicate to resource managers (OS or hardware)
    might be useful. Cache and memory capacities, energy/power,
    cache/memory bandwidth, and core occupancy are obvious
    resources.

    I have mentioned before that a market-like system of resource
    allocation might be useful. Unlike realworld economics, most
    information asymmetries could be avoided (though communication
    still has overhead the scales and responsiveness seem more
    friendly to market-like regulation).

    I do not have any idea how the complexity of the interaction of
    value of timeliness and resource availability could be handled
    well, but already implemented power management strategies give
    some hint that a budget can be given and the system can adjust
    reasonably well.

    OS support would be a problem. I do not think any architecture
    provides merging of page table levels.

    Not Much HW provides said feature.

    I do not know of any hardware that supports page table level
    merging.

    One weird thought I had was to have multiple page table "bases"
    loaded for a process. This would be a little like caching
    directory entries with prefetching of such on context switches.
    [snip]

    Rather than having to traverse the page table three times to
    load three specific node points into the translation cache (and
    likely caching other less useful information until it ages out),
    the necessary entries would be prefetched and "locked".

    ???

    If a process has a large address space but two intermediate
    level nodes are heavily used and/or densely populated,
    installing these nodes at context switch avoids having to load
    them from the general page table. (This is really just a
    prefetch operation, but it also supports a certain from of
    sparsity. If I understand correctly, with My 66000 if only
    three entries are in the first (base) level of the page table,
    at least that one indirection is initially required even if
    each of those three entries can then skip an arbitrary number
    of levels rCo which I did not have in mind, such is a neat trick.
    With three base pointers, the dense regions would be preloaded.)

    This (probably not worthwhilery|) concept is just prefetching of
    page table nodes and locking them into the node cache while the
    process is active.

    Two obvious issues come to mind. First, context switches are
    expected to be uncommon, so a modest savings becomes a trivial
    savings. Second, this does not scale down or up to different
    numbers of nodes.

    ???

    The benefit of prefetching a few intermediate nodes when a
    process activates shrunk if processes are rarely deactivated
    and reactivated.

    Different processes might want a different number of nodes
    prefetched and locked, but forcing the cache to have N locked
    entries for every thread has issues with utilization. If the
    number of prefetched nodes can vary, how is the storage handled
    (indirection cost to an overflow structure, internal
    fragmentation, or some other cost). With support for a large N,
    the costs increase. Of course, hardware could treat the lock as
    a hint, but that seems problematic for processes using such for
    a responsiveness constraint. Adding criticality metadata could
    avoid such at the cost of further complexity.

    [snip]

    My 66000 interrupt negotiator goes out and fetches the MSI-X
    message from its-core's interrupt table, then pre-fetches the
    PSL and registers while core continues running current thread.
    The Fetch message and fetch context delays are similar are run
    concurrently with current thread.

    Once MSI-X message verifies the core should take the interrupt,
    the new context is ready to install in core resources, while
    pushing our current state, saving latency at each step. While
    setting up the context, the first instruction is fetched, and
    the ISR is in control.

    PSL? Paging Structure L??

    It also seems that the execution aspects of a soon to be paused
    thread might be adjusted. If the interrupted thread is near a
    work type phase boundary (and the interrupt can tolerate some
    delay), stopping immediately (saving some energy by not trying
    to load a new working set into the cache) or continuing to the
    phase change (possibly with a more liberal power budget) may be
    useful. ESM provides a fine-grained phase boundary, but there
    may be other phase boundaries that could be exploited. System
    calls might sometimes be such a boundary.

    I rather suspect that such flexible scheduling/resource use
    would be far too complex to be worthwhile (the time between
    interrupt awareness and interrupt thread readiness seems likely
    to be rather small), but it seems interesting as a theoretical
    possibility.

    [I hope to write some comments on the Virtual Vector Method
    soonish. If I take my computer to the library Monday, I will
    send this and the five other posts then and probably not
    post again before Tuesday.]

    Thank you for sharing your insights. Even when I am not
    convinced I at least get to think about the matter.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 14 16:58:54 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    I studied the Itanium architecture manuals and amused myself by writing
    asm kernels that took good advantage of the large register set,
    including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working variable
    live on the 8-element stack, is significantly harder IMHO.

    Stacks are hard to program for, at least for me (even large ones where
    you don't get the overrun problem).

    Some time ago, I wrote some rather complicated Postscript formulas.
    I used two approaches: Putting a comment with the current stack
    next to each instruction, and writing a quick lex/yacc grammar
    to translate infix to postfix. (I also tried my hand at a
    two-dimensional rootfinder in Postscript, but that was a bit
    too much).

    RPN simply does not come naturally to me.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Mon Sep 14 10:04:47 2026
    From Newsgroup: comp.arch

    On 9/13/2026 11:22 PM, Anton Ertl wrote:
    scott@slp53.sl.home (Scott Lurndal) writes:
    So, those are all security issues. ARM has added a 'data independent timing'
    flag that can be used to ensure that a given instruction executes in constant
    time, regardless of the data being operated on.

    Interesting approach. So when you use this flag on a load, it will
    take the maximum time that a load can take (e.g., RAM access to RAM
    under contention, or access to a remote cache under contention),
    wherever it finds the item?

    Sure sounds ugly. :-( On a modern system, how long can this be?


    Intel has taken a different approach: It has documented the
    instructions that are guaranteed to have data-independent timings
    across all Intel microarchitectures; not sure if AMD gives the same guarantee. So if you want to write code that does not have a timing side-channel that reveals something about the data, you use only these instructions.

    So loads couldn't be on that list. So you couldn't use any load
    instructions in such code. Also sounds ugly, but I am not sure which alternative is worse.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 14 18:01:52 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    John Levine <johnl@taugh.com> posted:

    According to Anton Ertl <anton@mips.complang.tuwien.ac.at>:
    However, concerning my statement above, I think the IA-64 architects
    used examples for which the architecture (and its in-order
    implementation) was particularly well-suited (software-pipelinable
    inner loops), and that compilers eventually usually worked ok for such
    examples, too. It's just that these compilers don't work so well on
    general-purpose code.

    That's one of the problems that Multiflow also had, works great if you
    can predict the access patterns, a lot less great if the access patterns >> are data dependent.

    I believe that this will always be the case for memory references.
    When striding through memory accessing doublewords 1 cache miss
    serves 8 LDDs 7 get hits 1 takes a miss. How does one software
    schedule for that pattern ??

    Don't use cache :-)

    {derisively::} Riigghhtt

    There are several startups working on dense (10x or more) SRAM implementations with 1ns access times as an HBM replacement.

    Pin to pin--that is address/command in makes setup time to pin on
    HBC, and 1 ns later, data makes set up time of requesting chip ??
    {{Oh, BTW, this gives the SRAM only 300ps to be accessed.......}}

    With 1ns latency, who needs cache

    1ns is 5 clocks, which is about what the GBOoO chips at 5GHz are
    running WITH L1 caches.

    The issue with this thinking--is that it will take more than 1s
    to get from a core's AGEN unit to the pins of the die, and more
    than another 1ns for the incoming data to traverse the path in
    the opposite direction. Here, you have already tripled perceived
    latency; and have taken no account of other core's interfering
    with this access.

    I am sensitive to this because of the 68010 loop buffer, and the
    68020 cache--both were just big enough to let the fetch side and
    the data side of the processors to minimize interference to each
    other.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 14 18:08:20 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/13/2026 11:22 PM, Anton Ertl wrote:
    scott@slp53.sl.home (Scott Lurndal) writes:
    So, those are all security issues. ARM has added a 'data independent timing'
    flag that can be used to ensure that a given instruction executes in constant
    time, regardless of the data being operated on.

    Interesting approach. So when you use this flag on a load, it will
    take the maximum time that a load can take (e.g., RAM access to RAM
    under contention, or access to a remote cache under contention),
    wherever it finds the item?

    Sure sounds ugly. :-( On a modern system, how long can this be?

    The ARM DIT flag doesn't apply to standard loads or stores. However,
    software can use non-temporal loads (or explictly warm the cache)
    to eliminate timing variability due to caching on most modern architectures.

    The primary use for DIT is in cryptographic algorithms.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Andy Valencia@vandys@vsta.org to comp.arch on Mon Sep 14 12:52:27 2026
    From Newsgroup: comp.arch

    On 9/7/26 2:33 PM, MitchAlsup wrote:
    Has anyone ever wanted to allow shared ASIDs in such a way that shared memory or shared files have the ASID associated with the memory/file ?
    So that multiple processes accessing the same shared resource co-optim-
    ize themselves across cores and caches ??

    HP's Precision Architecture had this. TLB entries had a Access ID, and the process had a list of Protection ID's. For non-zero AID, you had to have its value in your PID list to access the page.

    Shared memory segments each got their own AID, and then each process
    which successfully attached received it in their PID. The PID list was bounded, and there was a fast path to swap in one which wasn't active
    when needed. Each shared memory object was accessed at the same 64-bit
    address in each process.

    Andy Valencia
    Home page: https://www.vsta.org/andy/
    To contact me: https://www.vsta.org/contact/andy.html
    No AI was used in the composition of this message
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 14 21:35:07 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/13/2026 11:22 PM, Anton Ertl wrote:
    Intel has taken a different approach: It has documented the
    instructions that are guaranteed to have data-independent timings
    across all Intel microarchitectures; not sure if AMD gives the same
    guarantee. So if you want to write code that does not have a timing
    side-channel that reveals something about the data, you use only these
    instructions.

    So loads couldn't be on that list. So you couldn't use any load >instructions in such code.

    You can, but the address must not depend on the data you process.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 06:34:30 2026
    From Newsgroup: comp.arch

    George Neuner <gneuner2@comcast.net> writes:

    On Fri, 11 Sep 2026 06:33:21 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    George Neuner <gneuner2@comcast.net> writes:


    Witness the proliferation of languages offering "managed environments"
    offering such niceties as automatic storage management, automatic lock
    handling (serialized object access), "comprehensions", etc., and large
    standard libraries

    I don't see any problem with that. These features help to implement
    functionality in less programming time (and with less maintenance
    time) than without using these features, and that's true for
    programmers at any competence level.

    without which the average programmer largely
    would be incapable of producing a working program.

    That's pure elitism.

    Really? That's not my conclusion ... it was the result found by a
    number of university studies and developer surveys.


    Most studies involving GC have shown that programmers working on short timelines are less likely to produce a correct [or sometimes even just complete] program using manual memory management vs using GC.

    Yeah. The studies I'm aware of that try to measure such things show
    not having to do manual memory management gives a productivity
    improvement of somewhere between 50 and 100 percent.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 06:45:15 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:

    Terje Mathisen <terje.mathisen@tmsw.no> schrieb:

    I studied the Itanium architecture manuals and amused myself by
    writing asm kernels that took good advantage of the large register
    set, including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working
    variable live on the 8-element stack, is significantly harder
    IMHO.

    Stacks are hard to program for, at least for me (even large ones
    where you don't get the overrun problem).

    Some time ago, I wrote some rather complicated Postscript formulas.
    I used two approaches: Putting a comment with the current stack
    next to each instruction, and writing a quick lex/yacc grammar
    to translate infix to postfix. (I also tried my hand at a
    two-dimensional rootfinder in Postscript, but that was a bit
    too much).

    RPN simply does not come naturally to me.

    Yes, writing code in PostScript is very different from more
    conventional programming. IME it helps to have a solid
    background in functional programming, but even with that
    PostScript needs a new way of thinking, and it takes some time
    to develop the skills to do so. At least it did for me.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 06:53:37 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [discussing programming in different assembly languages]

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,

    Huh???

    I studied the Itanium architecture manuals and amused myself by
    writing asm kernels that took good advantage of the large register
    set, including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working
    variable live on the 8-element stack, is significantly harder IMHO.

    If you don't mind my saying so, I think you have demonstrated
    a sufficient level of experience and ability to safely drop
    the H from IMHO. At least, it seems clear to me that you are
    world class in this particular area.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 17:21:40 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    [discussing programming in different assembly languages]

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,

    Huh???

    I studied the Itanium architecture manuals and amused myself by
    writing asm kernels that took good advantage of the large register
    set, including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working
    variable live on the 8-element stack, is significantly harder IMHO.

    If you don't mind my saying so, I think you have demonstrated
    a sufficient level of experience and ability to safely drop
    the H from IMHO. At least, it seems clear to me that you are
    world class in this particular area.

    Thanks! :-)

    My last job, 5 years with Cognite, where I had to learn several new
    languages, as well as a "cloud native" CI/CD development setup, gave me
    a real scare: I was certainly not "the brightest mind in the room" most
    of the time.

    OTOH, the challenge was interesting and I learned a _lot_, with Rust the
    one thing that's really staying with me.

    Terje
    PS. Cognite really is/was pretty special, Jon Markus Lerheim (previously
    the founder of FAST) managed to hire something like 40:40:20 percent
    Math Olympiad medalists, PhDs and relevant MSc graduates, and get them
    all to move to Oslo. I am not a betting person, but I would have been
    willing to bet this was impossible.
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 08:46:55 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [...]
    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. [...]

    64 bits is a fair amount of virtual address space. That is,
    until it starts being cut up into lots of pieces whose sizes
    are not known a priori, and moreso if (some or all of) the
    pieces are immovable.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Tue Sep 15 16:00:06 2026
    From Newsgroup: comp.arch

    On Mon, 14 Sep 2026 16:31:01 +0200, Terje Mathisen
    <terje.mathisen@tmsw.no> wrote:
    John Savard wrote:

    I have no doubt that a good programmer can beat the original FORTRAN
    compiler for the IBM 704 computer, despite the fact that its
    optimization was very nearly as good as that of most optimizing
    compilers until at least the late 1970s.

    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,

    Huh???

    I studied the Itanium architecture manuals and amused myself by writing
    asm kernels that took good advantage of the large register set,
    including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working variable
    live on the 8-element stack, is significantly harder IMHO.

    I hadn't thought about that. None the less, I still in large part
    stand by my earlier comment; ordinary programmers would have
    difficulty writing good code for the Itanium, and those that can are
    hard to find.
    Not that the large register set is the issue. Instead, the fact that instructions come in blocks of three, with some instructions only
    being available in certain slots, is what I saw as the biggest
    conceptual difficulty for an ordinary human programmer.
    After all, I don't categorize the AMD Am29000 architecture (not to be
    confused with the Am2900 bit-slice) as nightmarish; it also has a
    large register set.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 15 09:17:48 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:

    [...]

    Brooks quoted a factor of 10 in productivity between individual
    programmers in the 1960s, I suspect that gap has widened a lot
    since then but haven't looked for literature.

    Several comments.

    First the study cited very likely included a fair number of
    subjects, almost certainly at least double digits. It seems
    probable that the distribution roughly resembled a bell
    curve, so most of the people being measured fell in a range
    of roughly 2.5 to one.

    Second since the study was done in the 1960s, there is a good
    chance that most of the programming being done was done in
    assembly language. I suspect assembly tends give a (slightly)
    larger range than programming in a higher level language (and
    yes, even in C).

    Third there was very little in the way of educational resources
    for software development in the 1960s. Probably the gap has
    narrowed as better and more widespread as education in computer
    science has proliferated.

    No doubt there are significant differences in the results of
    different software developers. But we are a long way from
    what was being done in the 1960s.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 15 19:40:11 2026
    From Newsgroup: comp.arch

    John Savard wrote:
    On Mon, 14 Sep 2026 16:31:01 +0200, Terje Mathisen
    <terje.mathisen@tmsw.no> wrote:
    John Savard wrote:

    I have no doubt that a good programmer can beat the original FORTRAN
    compiler for the IBM 704 computer, despite the fact that its
    optimization was very nearly as good as that of most optimizing
    compilers until at least the late 1970s.

    I would not, however, be willing to assert that one could easily find
    human programmers who could beat an optimizing compiler... which had
    the Itanium as its target.

    I don't believe that there are any other current architectures out
    there which are as nightmarish to program in assembler as the Itanium,

    Huh???

    I studied the Itanium architecture manuals and amused myself by writing
    asm kernels that took good advantage of the large register set,
    including the rotating ones.

    Not at all a nightmare!

    In fact, writing x87 asm by hand, trying to keep every working variable
    live on the 8-element stack, is significantly harder IMHO.

    I hadn't thought about that. None the less, I still in large part
    stand by my earlier comment; ordinary programmers would have
    difficulty writing good code for the Itanium, and those that can are
    hard to find.
    Not that the large register set is the issue. Instead, the fact that instructions come in blocks of three, with some instructions only
    being available in certain slots, is what I saw as the biggest
    conceptual difficulty for an ordinary human programmer.

    That was not an issue, it reminded me a lot of writing Pentium code
    where the u and v pipes have very different capabilities: You had to
    pair up your instructions very carefully to get close to 2 instructions/cycles.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.arch on Wed Sep 16 17:21:05 2026
    From Newsgroup: comp.arch

    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems. OTOH I see no fundamental
    obstacles to having 2^64 bytes of memory. Assuming that
    memory cell plus various overheads need 1000 atoms, I get
    smaller amount of matter than in a current hard drives and
    comparable to current memory modules. Of course, aranging
    atoms into useful memory configuration will require significant
    technological breaktrough, but saying that this will never
    happen looks foolish.

    <snip>
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 16 10:58:16 2026
    From Newsgroup: comp.arch

    On 9/16/2026 10:21 AM, Waldek Hebisch wrote:
    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems. OTOH I see no fundamental
    obstacles to having 2^64 bytes of memory. Assuming that
    memory cell plus various overheads need 1000 atoms, I get
    smaller amount of matter than in a current hard drives and
    comparable to current memory modules. Of course, aranging
    atoms into useful memory configuration will require significant
    technological breaktrough, but saying that this will never
    happen looks foolish.

    While I agree with pretty much everything you say in the above post,
    note that the previous posts that you included above were talking about virtual memory size, not physical memory size.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 16 18:15:49 2026
    From Newsgroup: comp.arch

    antispam@fricas.org (Waldek Hebisch) writes:
    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32.

    The curve here is exponential, not linear.

    Simple observation is that
    ie memory is available, then software will use it.

    That seems to be the microsoft philosophy, yes.

    However, that applies more to the physical address space
    not the virtual address space.

    Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64.

    It's not just memory that occupies pages in the physical or
    virtual address spaces. A PCI express device (particularly
    CXL) can have very large (64-bit) aperture sizes for the
    set architectural base address registers (BAR) which, in
    the CXL case, may actually be mapped into a process
    virtual address space.

    There are other ways to partition the physical address space
    (e.g. using the top bits as a NUMA node number, or in a distributed
    environment such as 3Leaf Systems had, the top eight bits
    identified another host on the memory bus).

    Even with that, 2^64 in the VA space is likely out of reach for a long, long time.


    But IMO, one can not exlude this,
    especially in biggest systems. OTOH I see no fundamental
    obstacles to having 2^64 bytes of memory. Assuming that
    memory cell plus various overheads need 1000 atoms, I get
    smaller amount of matter than in a current hard drives and
    comparable to current memory modules. Of course, aranging
    atoms into useful memory configuration will require significant
    technological breaktrough, but saying that this will never
    happen looks foolish.

    <snip>
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 16 18:59:00 2026
    From Newsgroup: comp.arch


    antispam@fricas.org (Waldek Hebisch) posted:

    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems.

    Consider a large server the size of a basketball stadium. Thousands
    of racks each with dozens of systems and each motherboard maxed out
    with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.

    Suppose HyperVisor/OS uses the high order PA bits to route DRAM
    requests from this rack-board to any other board in the stadium,
    emulating a coherent system with 2^58-bytes of available DRAM by
    moving pages from rack-board to rack-board on demand (or better).

    Each said rack-board would have its own Device address aperture
    and configuration space aperture. The combined Device address
    space would exceed ECAM PCIe available address space. So, by
    routing device control register writes across the network, one
    gets a FREE expansion of ECAM.

    So, it seems to me that we already have the capability to build
    systems of that scale. Whether we do or not depends on other
    factors {like "we tried and it did not perform" or simply $$$}.

    OTOH I see no fundamental
    obstacles to having 2^64 bytes of memory. Assuming that
    memory cell plus various overheads need 1000 atoms, I get
    smaller amount of matter than in a current hard drives and
    comparable to current memory modules. Of course, aranging
    atoms into useful memory configuration will require significant
    technological breaktrough, but saying that this will never
    happen looks foolish.

    <snip>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 16 19:16:03 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    antispam@fricas.org (Waldek Hebisch) posted:

    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems.

    Consider a large server the size of a basketball stadium. Thousands
    of racks each with dozens of systems and each motherboard maxed out
    with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.

    Suppose HyperVisor/OS uses the high order PA bits to route DRAM
    requests from this rack-board to any other board in the stadium,
    emulating a coherent system with 2^58-bytes of available DRAM by
    moving pages from rack-board to rack-board on demand (or better).

    That's basically what we developed at 3Leaf Systems two decades
    ago. We designed and fabricated an ASIC that connected to
    hypertransport and extended the AMD conherency domain over infiniband
    to up to 64 (at the time) independent nodes. We used the top
    8 bits of the physical address as the target node address.
    Fred Weber was one of our advisors.

    Two decades later PCIe CXL is close to being able to support
    a similar configuration.



    So, it seems to me that we already have the capability to build
    systems of that scale. Whether we do or not depends on other
    factors {like "we tried and it did not perform" or simply $$$}.

    Performance was an issue. At the time, DDR IB had cut-through
    routing, so the round trip latency was about 800ns through the
    switch for a non-posted request. Our Quickpath version for
    Intel used QDR IB and had close to 200ns r/t latency in
    simulation for a non-posted request. Unfortunately we got
    caught up in the Bush recession - funding ran out and a
    pending acquisition canceled.

    We had a custom bare-metal hypervisor that could assign resources
    from across the entire system to individual VMs. There was no
    sharing of physical cores by multiple VMs, which simplified the
    hypervisor somewhat. Linux could hot-plug/unplug CPUs and
    memory dynamically.

    Today with CXL 3.2 and GEN6 PCIe, the switching cost is significantly better.

    There were also issues with cache line contention due to
    spin locks (LLL had a particular test that hammered a single
    lock across all CPUs), in part due to the AMD backstop
    of using the global bus lock if it took to long to get exclusive
    cache line access to guarantee forward progress.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 16 12:24:01 2026
    From Newsgroup: comp.arch

    On 9/16/2026 11:59 AM, MitchAlsup wrote:

    antispam@fricas.org (Waldek Hebisch) posted:

    snip
    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems.

    Consider a large server the size of a basketball stadium. Thousands
    of racks each with dozens of systems and each motherboard maxed out
    with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.

    Suppose HyperVisor/OS uses the high order PA bits to route DRAM
    requests from this rack-board to any other board in the stadium,
    emulating a coherent system with 2^58-bytes of available DRAM by
    moving pages from rack-board to rack-board on demand (or better).

    Each said rack-board would have its own Device address aperture
    and configuration space aperture. The combined Device address
    space would exceed ECAM PCIe available address space. So, by
    routing device control register writes across the network, one
    gets a FREE expansion of ECAM.

    So, it seems to me that we already have the capability to build
    systems of that scale. Whether we do or not depends on other
    factors {like "we tried and it did not perform" or simply $$$}.

    Is this essentially a large NUMA system?
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Wed Sep 16 23:12:40 2026
    From Newsgroup: comp.arch

    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit
    versions of both those architectures, they were a serious nuisance.

    They'd apparently been designed when physical memories were measured in
    small numbers of MB, with the idea that using any significant fraction of
    a 4GB virtual address space was never going to happen.

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 00:19:56 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    Thank you Scott for what follows.

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    antispam@fricas.org (Waldek Hebisch) posted:

    Scott Lurndal <scott@slp53.sl.home> wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    Paul Clayton <paaronclayton@gmail.com> posted:

    On 9/2/26 9:50 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I have wondered why Motorola did not add a 24-bit address mode
    (or even provide such with a hardwired configuration on earlier
    implementations, knowing that the extra bits would be desired
    for other uses to save memory).

    We realized our earlier mistake and did not want to repeat it.

    I think this would be more a case of "extend it" than repeat it,
    paying an incompatibility price to remove what was viewed as a
    mistake.

    [snip]
    AArch64 provides a means (Top Byte Ignore) of masking the most
    significant octet to allow it to be used by software.

    Sins of the present...

    I do not understand how allowing software to limit its future
    capabilities is an architectural flaw.

    When SW has used TBI often enough AND those same applications
    need 63-bit (or 64-bit) virtual addresses.

    Which will likely be _NEVER_. 2^64 is a, pardon my french,
    shitload of virtual memory. Then there is the translation
    cost with up to perhaps seven or more levels of page table
    walk required. (Note that ARMv9 has support for 128-bit
    page table entries [FEAT_D128] which support up to 56-bits
    of PA and VA space).


    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems.

    Consider a large server the size of a basketball stadium. Thousands
    of racks each with dozens of systems and each motherboard maxed out
    with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.

    Suppose HyperVisor/OS uses the high order PA bits to route DRAM
    requests from this rack-board to any other board in the stadium,
    emulating a coherent system with 2^58-bytes of available DRAM by
    moving pages from rack-board to rack-board on demand (or better).

    That's basically what we developed at 3Leaf Systems two decades
    ago. We designed and fabricated an ASIC that connected to
    hypertransport and extended the AMD conherency domain over infiniband
    to up to 64 (at the time) independent nodes. We used the top
    8 bits of the physical address as the target node address.
    Fred Weber was one of our advisors.

    Two decades later PCIe CXL is close to being able to support
    a similar configuration.



    So, it seems to me that we already have the capability to build
    systems of that scale. Whether we do or not depends on other
    factors {like "we tried and it did not perform" or simply $$$}.

    Performance was an issue. At the time, DDR IB had cut-through
    routing, so the round trip latency was about 800ns through the
    switch for a non-posted request. Our Quickpath version for
    Intel used QDR IB and had close to 200ns r/t latency in
    simulation for a non-posted request. Unfortunately we got
    caught up in the Bush recession - funding ran out and a
    pending acquisition canceled.

    We had a custom bare-metal hypervisor that could assign resources
    from across the entire system to individual VMs. There was no
    sharing of physical cores by multiple VMs, which simplified the
    hypervisor somewhat. Linux could hot-plug/unplug CPUs and
    memory dynamically.

    Today with CXL 3.2 and GEN6 PCIe, the switching cost is significantly better.

    There were also issues with cache line contention due to
    spin locks (LLL had a particular test that hammered a single
    lock across all CPUs), in part due to the AMD backstop
    of using the global bus lock if it took to long to get exclusive
    cache line access to guarantee forward progress.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 00:20:35 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/16/2026 11:59 AM, MitchAlsup wrote:

    antispam@fricas.org (Waldek Hebisch) posted:

    snip
    2^24 was considerd huge, later 2^32. Simple observation is that
    ie memory is available, then software will use it. Regardless
    how much memory do you have. So the only question is if enough
    memory will be available. Current semiconductor technology
    is advancing slower than in the past. And one can doubt if
    it ever can get close to 2^64. But IMO, one can not exlude this,
    especially in biggest systems.

    Consider a large server the size of a basketball stadium. Thousands
    of racks each with dozens of systems and each motherboard maxed out
    with DRAM. Suppose that each rack-board contains 2^44-bytes of DRAM.

    Suppose HyperVisor/OS uses the high order PA bits to route DRAM
    requests from this rack-board to any other board in the stadium,
    emulating a coherent system with 2^58-bytes of available DRAM by
    moving pages from rack-board to rack-board on demand (or better).

    Each said rack-board would have its own Device address aperture
    and configuration space aperture. The combined Device address
    space would exceed ECAM PCIe available address space. So, by
    routing device control register writes across the network, one
    gets a FREE expansion of ECAM.

    So, it seems to me that we already have the capability to build
    systems of that scale. Whether we do or not depends on other
    factors {like "we tried and it did not perform" or simply $$$}.

    Is this essentially a large NUMA system?

    Well, it not a UMA system ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 00:26:01 2026
    From Newsgroup: comp.arch


    jgd@cix.co.uk (John Dallman) posted:

    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit
    versions of both those architectures, they were a serious nuisance.

    The always and obvious problem when the position of a segment is not
    "anywhere in PAS" and where the size of the segment cannot be "every
    thing possible in memory".

    Anytime one want segments as small as a word to as large as can fit in
    memory one runs into this problem.

    They'd apparently been designed when physical memories were measured in
    small numbers of MB, with the idea that using any significant fraction of
    a 4GB virtual address space was never going to happen.

    Happened to 6800, 6502, PDP-8, PDP-11, VAX, ~IBM 3080, ... So, you would
    think that history would be showing them not to do that to themselves.


    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 17 09:32:15 2026
    From Newsgroup: comp.arch

    jgd@cix.co.uk (John Dallman) writes:
    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit
    versions of both those architectures, they were a serious nuisance.

    They'd apparently been designed when physical memories were measured in
    small numbers of MB, with the idea that using any significant fraction of
    a 4GB virtual address space was never going to happen.

    At least not before 64-bit addresses are available.

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was
    introduced in November 1995 (and it probably took a while until it was delivered); the PowerPC620 only appeared in 1997. One would think
    that with these address-space restrictions the HPPA and Power
    architects would feel more pressure than the others to go 64-bit soon.

    We had an Alpha with 256MB RAM in 1995, and we were not into big
    machines (this Alpha only had one CPU). So the address-space
    limitations of HPPA and PowerPC probably made themselves felt strongly
    before the 64-bit variants were available. Ironically, Alpha was
    canceled before we reached 4GB RAM in our machines (and
    general-purpose MIPS CPUs, too).

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 17 10:00:51 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    While I agree with pretty much everything you say in the above post,
    note that the previous posts that you included above were talking about >virtual memory size, not physical memory size.

    Historically, once physical address space has approached virtual
    address size, problems made themselves felt that resulted in pressure
    to go for larger virtual addresses. Intel introduced PAE (36-bit
    physical addresses) with the Pentium Pro in 1995, Windows made use of
    that starting with Windows 2000. Linux with 2.3.23 in 1999, but the
    Linux developers had little love for using >1-2GB physical memories on
    32-bit kernels and there are motions to remove support for them <https://lwn.net/Articles/813201/>.

    So, it looks like ideally VAs have at least one bit more tham PAs. HPPA/PowerPC-like segmentation may make for a larger difference if the
    OS uses these segments. AFAIK Linux just ignored the segments; I
    certainly never noticed them on my iBook G4 with 1.25GB RAM running
    Linux.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Thu Sep 17 13:33:40 2026
    From Newsgroup: comp.arch

    In article <2026Sep17.113215@mips.complang.tuwien.ac.at>, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it
    was delivered); the PowerPC620 only appeared in 1997. One would think
    that with these address-space restrictions the HPPA and Power
    architects would feel more pressure than the others to go 64-bit
    soon.

    Do not underestimate the power of corporate conservatism and sectional interests. From 1995-2005 I regularly backstopped for technical support
    staff who were finding that customers felt they were locked into one
    particular commercial UNIX and could not change without vast disruptions.


    We felt this was weird, because the UNIXes of the era were all pretty
    similar. It gradually became clear that manufacturers' training courses
    on commercial UNIXes emphasised the differences and tried hard to give
    the impression that changing to a another supplier would be difficult and expensive. Human inertia meant that customers' sysadmins co-operated with
    that, giving the illusion that they were indispensable while acting
    against their employers' interests. This meant that well-established manufacturers like HP and IBM felt less urgency to move to 64-bit. Sun
    and SGI were a bit more dynamic, until they got into financial trouble,
    and DEC knew they needed Alpha to have a hope of survival.

    Meanwhile, we employed sysadmins who dealt with all of these UNIXes every
    week without difficulty. We also had to learn about things like POWER and PA-RISC segmentation to squeeze large 32-bit applications onto them.
    64-bit was very welcome!

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 16:06:20 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    jgd@cix.co.uk (John Dallman) writes:
    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit >versions of both those architectures, they were a serious nuisance.

    They'd apparently been designed when physical memories were measured in >small numbers of MB, with the idea that using any significant fraction of
    a 4GB virtual address space was never going to happen.

    At least not before 64-bit addresses are available.

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it was delivered); the PowerPC620 only appeared in 1997. One would think
    that with these address-space restrictions the HPPA and Power
    architects would feel more pressure than the others to go 64-bit soon.

    CRAY-1 1976

    We had an Alpha with 256MB RAM in 1995, and we were not into big
    machines (this Alpha only had one CPU). So the address-space
    limitations of HPPA and PowerPC probably made themselves felt strongly
    before the 64-bit variants were available. Ironically, Alpha was
    canceled before we reached 4GB RAM in our machines (and
    general-purpose MIPS CPUs, too).

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Wed Sep 16 18:03:25 2026
    From Newsgroup: comp.arch

    On 9/9/26 7:25 PM, MitchAlsup wrote:

    Paul Clayton <paaronclayton@gmail.com> posted:
    [snip]
    I do not understand how the Virtual Vector Method would be
    implemented efficiently. SIMD with its explicit pack and unpack
    instructions seems likely to provide similar control. (Most SIMD
    designs do not provide special support for short strides or
    complete structure unpacking where all the data in a non-strided
    stream is used but the data is scattered. GPUs might provide
    special support for three and four color (un)packing.) However,
    I am not a hardware designer (though I might understand a simple
    flowchart ry|).

    SIMD is simply direct compiler management of multi-lane calculations.

    Which makes it easy to understand how hardware would implement
    such.

    vVM allows for the processor to utilize multi-lane calculations
    with flexibility SIMD can never do::

    word = word + half*byte;

    Not directly, but most SIMD instruction sets provide pack and
    unpack. Yes, that introduces substantial instruction overhead
    (three unpack instructions for the above code, I think) and
    contributes to the number of SIMD instructions in the
    instruction set (one pack and unpack instruction pair for each
    doubling in operand size), but it makes the implementation
    straightforward. If unpacking is rare (or unpacked values are
    heavily reused from registers), such might not hurt power-
    performance-area.

    On a size-related note, SIMD also has instruction number issues
    with "overflow", e.g., requiring multiply high instructions.

    (The Mill had the nice feature of specifying saturating,
    overflow exception, wrapping, or doubling width.)

    And if the HW happens to have 256-bits of calculation width, then
    8 of those can be performed in 1 pipelined clock, without any
    SW visible SIMD registers.

    In theory, hardware knowing that one of the multiplicands is
    smaller would allow the multiplier to have higher throughput.
    If such was known at hardware design time to be a case worth
    optimizing, I guess that doubling of throughput for such would
    require less than 10% added hardware.

    I do feel that VVM is a nice software interface. It avoids a
    lot of SIMD issues and not just instruction diversity explosion.

    vVM is 2 instructions that provide the utility of thousands (based
    on typical SIMD ISA).

    vVM fails at the super luminary uses of SIMD (one lone SIMD inst
    without any containing loop structure)

    If there is significant overhead in start-up, even two or three
    SIMD instructions might have better performance.

    I also got the impression that there was no way to exploit
    blocking as in matrix multiplication. With SIMD, the extra
    register context can be used as a first level of blocking.
    Wide implementations of vVM would have extra storage that
    (I think) could be used to repeat use (conceptually similar
    to how vVM avoids splat instructions but supports vector-
    scalar operations), but this microarchitectural aspect cannot
    (I think) be exploited by software.

    While not exactly SIMD/vector operations, I also like lane
    crossing operations like count bytes (two-bytes) until zero,
    which can find a zero terminated string length and find a
    matching entry given a previous xor (which could be useful
    with multiple hash table entries in a bucket; with a good
    hash function producing a 64-bit result, there would not be
    a need for splatting). Substantial hardware reuse may be
    possible with find first/last set/unset bit hardware.

    There are some operations that hardware can do much more
    efficiently than software (like popcount), but it is not
    obvious that such instructions are broadly useful or
    otherwise important enough to provide.

    Side note: I wonder if moving the LOOP instruction to the top of
    the loop might be better. Such would seem to require adding an
    early loop body-end instruction (or recognizing a jump to the
    instruction after the end of the loop as such) when there is not
    a single loop end. One possible advantage might be the ability
    to fold loop counter preconditions into that instruction, but
    having the word size of the loop body available at the start
    might be useful. This would also add invisible architectural
    context (loop end marker) to be saved on context switches.

    I think some DSPs started loops with a such a special
    instruction. (It is also kind of like a prepare-to-branch
    instruction.)
    I feel it does not exploit a few local, non-loop wide execution
    cases, does not address blocking, and the loop length limits may
    be introduce issues.

    When a LD (or ST) touches a cache line and HW can determine the
    access pattern is "dense"; HW reads out cache-width of data and
    accesses that buffer multiple times feeding the multi-lane calc-
    ulation. So, the wide buffering flip-flops are present (they have
    to be for perf) but each implementation gets to decide how many
    and how wide, and how many cache ports are "reasonable" for this implementation.

    It might also be nice to optimize cases of short constant
    strides (fewer cache tag accesses). While similar checks could
    be performed for scalar loads, with a vector one can expect
    more reuse of information. Reusing TLB probes is another obvious
    optimization.

    Being able to optimize array of structure accesses where
    multiple structure members are accessed might be worthwhile.
    Complex number and red-blue-green vectors are obvious examples,
    but array of structures might have multiple members accessed
    in a loop and benefit from reducing the number of cache
    accesses.

    Those buffers are connected to the multi-lane calculation unit
    with lots of multiplexers (which perform the pack, unpack,
    width-changes, and forwarding). Those multiplexers take an
    extra cycle into and out of when considering latency.

    Controlling all of that seems complex (to me). On the other
    hand, it obviously provides microarchitectural scalability with
    scalar execution semantics and without greatly increasing the
    number of instructions in the instructions set. Avoiding merely
    temporary context that needs to be saved and restored on context
    switches is also a vVM advantage, but that advantage is related
    to the disadvantage of not optimizing blocking.

    Yet I also recognize that being ideal for
    all use cases regardless of complexity is both unachievable
    (not all tradeoffs are limited to design difficulty or even
    implementation area) and impractical (aside from complexity
    having a cost, different workloads benefit differently from
    "effort" and the value of the benefit is uniform across all
    workloads at all scales).

    While I agree that it is more complex, is it more complex than
    having 1,300 SIMD instructions ??? And when we have the capability
    of building 10-wide machines (Apple M5) does that really matter?

    I do not think one needs to add instructions for different
    execution widths except for load and store. I guess using
    metadata preserved from loads is kind of like what vVM does
    in setting up the operand routing controls. (This is also
    kind of like sign extending short loads to use ordinary
    operate instructions. A Load128 instruction in an architecture
    supporting 512-bit SIMD registers would treat the "higher" 384
    bits as not-present.)

    I (and everyone else) would prefer something that provided
    all the advantages of SIMD and vVM and none of the disadvantages
    of either, but when defining an interface one must make a
    decision (EEP!) about the importance of various uses.

    [snip]
    Microarchitecturally exploiting spatial locality (SRAM array
    width, block size, and page size) to avoid redundant work seems
    an obvious method toward power efficiency. Allowing software to
    help seems reasonable to *me* (but I am excessively attracted by
    hardware-software cooperation).

    I don't think I am putting anything in the way of SW doing what it
    wants, here.

    Software can provide the base semantics to a My 66000
    implementation, but My 66000 seems uninterested in software-
    provided hints or pre-conditions. (Historically such have been
    very fragile.)

    I think AArch64's load register pair instruction requires 128-
    bit alignment, such would allow hardware to know in advance that
    a single 128-bit wide SRAM array access can satisfy this
    operation. (This might not be useful information, but it might
    facilitate scheduling before address generation.)

    The hard part seems to be encoding information that is actually
    useful. Implementation changes can make information useless
    (like aligned SIMD load instructions which became useless once
    unaligned load instructions were equally fast in the aligned
    case).

    There may be preconditions known to software that hardware can
    exploit (a bit like vVM prohibits branches).

    I suspect something like a signature buffer ("Signature Buffer:
    Bridging Performance Gap between Registers and Caches", Lu Peng,
    2004) which stores values accessed relative to a base pointer in
    a special buffer/cache, might benefit from a hint about the
    expected utility of such caching.

    With an intermediate software distribution format, it might be
    more practical to implement microarchitecture-specific (and
    version-specific) optimizations while providing compatibility
    for distributed software.

    Not all information is worth storing/caching. Bloating the
    instruction stream with all the information available to
    software is obviously not a good idea, but I _feel_ that
    software (compiler/developer) will have information that
    could be useful to hardware.

    [snip]
    x86 had subregisters before MMX. Even some RISCs used register
    pairs for double precision floating point, presenting the same
    renaming/forwarding issues.

    There you go again blaming x86 as being something good in computer architecture.

    I am not claiming such is good. I am claiming that MMX was not
    the start of a register holding multiple values.

    (I think for 8086, performing one 16-bit load and operating on
    two 8-bit parts was much faster than using two loads. Having
    the equivalent of AArch64's load pair of registers, which can
    have 32-bit values, would have used a precious register in 8086.
    At the time, such seems likely to have been a reasonable design
    choice. Of course, people were often writing in assembly then,
    so compiler use of this was not that important.)

    Even bit vectors are actually multiple bit values; I would want
    a compiler to pack 32 bools into a single 32-bit "integer"
    rather than use 32 bytes.

    The complementary issue of storing a single value in multiple
    registers is also unpleasant for a compiler, but many RISCs
    thought this was acceptable. While I think My 66000's DOUBLE
    avoids the register allocation constraints of simple register
    pairing, it does present a single "value" as occupying
    multiple registers.


    [snip]
    On the other hand, AI turned around and is consuming all available
    DRAM, so its a good thing SSDs are so fast.

    Yes, this does facilitate extending the length of the refresh
    cycle. The pain of paging memory in-and-out is not as great and
    file caching and prefetch is not as critical. On the other hand,
    such seems likely to increase wear on the flash.

    I read that AI is also increasing the price of flash (and actual
    disk drives since the raw scraped data requires storage as
    well).

    I probably missed commenting on some parts of the long response
    to my long previous posting, but I am going to count this
    sub-project as done.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Thu Sep 17 17:26:34 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    jgd@cix.co.uk (John Dallman) writes:
    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >> >Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit
    versions of both those architectures, they were a serious nuisance.

    They'd apparently been designed when physical memories were measured in
    small numbers of MB, with the idea that using any significant fraction of >> >a 4GB virtual address space was never going to happen.

    At least not before 64-bit addresses are available.

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was
    introduced in November 1995 (and it probably took a while until it was
    delivered); the PowerPC620 only appeared in 1997. One would think
    that with these address-space restrictions the HPPA and Power
    architects would feel more pressure than the others to go 64-bit soon.

    CRAY-1 1976

    So one might subclassify 64-bit into:

    - 64-bit arithmetic operations
    - 64-bit addressing operations.

    The B3500 in 1965 did 400-bit arithmetic operations (100 digit),
    applications were limited to 500KB[*] (code + data - slightly less
    as a small amount was used by the MCP (OS)). Later machines
    increased the physical address space from 6 digits to 10 digits
    (500MB).

    [*] 1 million digits/nibbles.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 17 16:47:07 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> writes:
    I generally agree that a compiler will not surpass the best
    performance of a supreme expert human programmer (for now).

    Take a look at Figure 1 of

    https://www.complang.tuwien.ac.at/kps2015/proceedings/KPS_2015_submission_29.pdf

    There you see the performance from a compiler from 1997 (gcc-2.7.2.3,
    and gcc-2.7.0 actually appeared in 1995), from 1999 (egcs-1.1.2), and
    from 2015 (gcc-5.2.0, clang-3.5). You also see different compiler
    optimization levels (-O0, -O3 with various -f... options for defining
    behaviour that the 2015 compilers treat as undefined, and -O3 with as
    little language definition as the 2015 compilers use by default). And
    you see a sequence of manual optimizations, from tsp1 to tsp9,
    published by Jon Bentley in his 1982 book writing efficient programs;
    you can see the source code for these programs by following the links
    on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tsp.html>.

    There you can see that the difference between the 1997 compiler and
    the 2015 compilers at -O3 is usually small. Has there been much
    progress since 2015? I doubt it.

    You can also see that the manual optimization steps bring vast
    improvements, far more than what the compilers have managed in these
    18 years. One interesting case is the step from tsp4 to tsp5
    (inlining of a function), which is flat for all the compilers, so all
    the compilers (from 1997 to 2015) do it by themselves; it is a
    prerequisite to following manual optimizations, so it cannot be left
    away in the manual optimization sequence.

    In the meantime I know that this program can be vectorized (and we
    have discussed manual vectorization at the assembly/intrinsic level
    here in 2016). I have tried getting recent gcc and clangs to
    vectorize it, and they certainly do not do it in an effective way for tsp1...tsp9. I have played around with transforming the program in a
    way that let's the auto-vectorization work, and have had some
    successes, but something else intervened when I was writing that up,
    so I do not have it in a form for public consumption. In any case,
    there is another factor of 3 or so in this program from manual changes
    that get "auto"-vectorization to work, but the "auto"-vectorization
    does not manage to do it automatically.

    And this is a simple program whose optimizations have been published
    in 1982. Admittedly, one of the steps changes what this program
    outputs, but that's one of the advantages human programmers have over compilers: They know the requirements, and can optimize to them, while
    the compiler only knows the source code, and has to treat it as a specification.

    I am guessing that effort (both compiler development and
    compilation computer resources) is a significant constraint.
    Dedicating three racks of servers to compile a module for a week
    seems unlikely to be attractive.

    The week would be a problem, throwing a lot of hardware at development
    has become pretty common in the last decade (e.g., continuous
    integration tests).

    Having half of the world's
    programmers working on compiler development also seems unlikely
    to be judged worthwhile.

    Certainly not by me. My impression is that gcc and clang development
    has more developers thrown at it than is helpful for programmers. The
    devil finds work for idle hands to do, and in this case, it's
    elaborate "optimizations" based on the assumption that programs do not
    exercise undefined behaviour. There are other parts of development of
    these compilers that is beneficial, such as better error messages.

    Human beings also seem to have a better information caching
    system. For a compiler modifying one variable name for clarity
    would (I think) typical force a recompilation, possibly of the
    whole program if whole program optimization is used, but a human
    would probably recognize the change as non-semantic.

    That case is actually easy to catch and deal with, it's just that it's
    not done. Other things in that vein, like ccache, seem to have fallen
    out of usage, at least I have not read about it for decades.

    Some of the difficulties in developing a compiler superior to a
    human being in quality of generated code are quite substantial,
    but it seems that a well-engineered specialized machine should
    be able to outperform a human being even in an "intellectual"
    task.

    What makes you think so? And what do you mean with "specialized
    machine"? And what does this have to do with compilers?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 19:38:01 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    Paul Clayton <paaronclayton@gmail.com> writes:
    I generally agree that a compiler will not surpass the best
    performance of a supreme expert human programmer (for now).

    Take a look at Figure 1 of

    https://www.complang.tuwien.ac.at/kps2015/proceedings/KPS_2015_submission_29.pdf

    A good read...

    There you see the performance from a compiler from 1997 (gcc-2.7.2.3,
    and gcc-2.7.0 actually appeared in 1995), from 1999 (egcs-1.1.2), and
    from 2015 (gcc-5.2.0, clang-3.5). You also see different compiler optimization levels (-O0, -O3 with various -f... options for defining behaviour that the 2015 compilers treat as undefined, and -O3 with as
    little language definition as the 2015 compilers use by default). And
    you see a sequence of manual optimizations, from tsp1 to tsp9,
    published by Jon Bentley in his 1982 book writing efficient programs;
    you can see the source code for these programs by following the links
    on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tsp.html>.

    There you can see that the difference between the 1997 compiler and
    the 2015 compilers at -O3 is usually small. Has there been much
    progress since 2015? I doubt it.

    You can also see that the manual optimization steps bring vast
    improvements, far more than what the compilers have managed in these
    18 years.

    This is the "find a better algorithm" step of making programs fast.
    If one dives into LINPACK and LAPACK one finds FORTRAN code written
    in such a way that array-access-order is cache friendly. No compiler
    is ever going to do that automagically for general array code where
    the program specifies the indexing. However modern FORTRAN can use array-notation to free the compiler to create cache friendly algo-
    rythms.

    One interesting case is the step from tsp4 to tsp5
    (inlining of a function), which is flat for all the compilers, so all
    the compilers (from 1997 to 2015) do it by themselves; it is a
    prerequisite to following manual optimizations, so it cannot be left
    away in the manual optimization sequence.

    In the meantime I know that this program can be vectorized (and we
    have discussed manual vectorization at the assembly/intrinsic level
    here in 2016).

    I can't seem to find the source code ... I would like to try vectorizing
    with My 66000 ISA.

    -----------
    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 17 19:40:55 2026
    From Newsgroup: comp.arch


    scott@slp53.sl.home (Scott Lurndal) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    jgd@cix.co.uk (John Dallman) writes:
    In article <11898n8$1nc06$1@dont-email.me>, paaronclayton@gmail.com (Paul >> >Clayton) wrote:

    I suspect something like this (but coarser-grained) was the
    motivation for PowerPC's segments. Effectively each of 16
    address regions (in the 32-bit "effective" address space) had an
    ASID. (Virtual segment IDs were 24 bits.)

    software ported from HP PA-RISC (which had segments).

    Having had to fit large processes into those segments, on the 32-bit
    versions of both those architectures, they were a serious nuisance.

    They'd apparently been designed when physical memories were measured in >> >small numbers of MB, with the idea that using any significant fraction of >> >a 4GB virtual address space was never going to happen.

    At least not before 64-bit addresses are available.

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was
    introduced in November 1995 (and it probably took a while until it was
    delivered); the PowerPC620 only appeared in 1997. One would think
    that with these address-space restrictions the HPPA and Power
    architects would feel more pressure than the others to go 64-bit soon.

    CRAY-1 1976

    So one might subclassify 64-bit into:

    - 64-bit arithmetic operations
    - 64-bit addressing operations.

    I even thought about the Cray-1s 24-bit address space when writing the above.

    The B3500 in 1965 did 400-bit arithmetic operations (100 digit),
    applications were limited to 500KB[*] (code + data - slightly less
    as a small amount was used by the MCP (OS)). Later machines
    increased the physical address space from 6 digits to 10 digits
    (500MB).

    [*] 1 million digits/nibbles.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 17 21:25:54 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
    And
    you see a sequence of manual optimizations, from tsp1 to tsp9,
    published by Jon Bentley in his 1982 book writing efficient programs;
    you can see the source code for these programs by following the links
    on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tsp.html>.

    There you can see that the difference between the 1997 compiler and
    the 2015 compilers at -O3 is usually small. Has there been much
    progress since 2015? I doubt it.

    You can also see that the manual optimization steps bring vast
    improvements, far more than what the compilers have managed in these
    18 years.

    This is the "find a better algorithm" step of making programs fast.

    tsp1..tsp9 all use the same algorithm: In the inner loop, look for a
    closest city to the one where the salesman currently is, and go there;
    yes, that's not a particularly good TSP algorithmand produces results
    that are not close to optimal, but it's the algorithm Bentley chose as
    running example for his book.

    If one dives into LINPACK and LAPACK one finds FORTRAN code written
    in such a way that array-access-order is cache friendly. No compiler
    is ever going to do that automagically for general array code where
    the program specifies the indexing.

    Matrix300 was eliminated from SPEC because all the computer companies
    started to use automatic cache blocking. IIRC this was for the step
    from SPEC89 to SPEC92. Later Sun managed to optimize IIRC the ear
    benchmark to perform array-of-structures to structure-of-arrays
    transformation, achieved a speedup by a factor of IIRC 2, and a
    significant increase in the aggregate SPEC score.

    If tsp1.c was a SPEC program, I am sure that the C compiler writers
    would find ways to optimize it to similar performance as tsp9, and
    further vectorize the code for even better performance.

    In the meantime I know that this program can be vectorized (and we
    have discussed manual vectorization at the assembly/intrinsic level
    here in 2016).

    I can't seem to find the source code ... I would like to try vectorizing
    with My 66000 ISA.

    For the 2016 stuff, you can find versions from tspa.c to tspk.c on <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/>. Note that
    these are not always derived from the version that is alhpabetically
    right before it, and some of the early ones are not vectorized at all,
    just attempts to get the compiler to vectorize it, which failed at the
    time. The first actually vectorized example is <https://www.complang.tuwien.ac.at/anton/lvas/effizienz/tspe.c>

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Fri Sep 18 01:08:02 2026
    From Newsgroup: comp.arch

    On Thu, 17 Sep 2026 13:32 +0100 (BST)
    jgd@cix.co.uk (John Dallman) wrote:

    In article <2026Sep17.113215@mips.complang.tuwien.ac.at>, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    The interesting aspect is that the first 64-bit CPUs were the R4000
    (1991) and the 21064 (1992) while the first 64-bit HPPA machine was introduced in November 1995 (and it probably took a while until it
    was delivered); the PowerPC620 only appeared in 1997. One would
    think that with these address-space restrictions the HPPA and Power architects would feel more pressure than the others to go 64-bit
    soon.

    Do not underestimate the power of corporate conservatism and sectional interests. From 1995-2005 I regularly backstopped for technical
    support staff who were finding that customers felt they were locked
    into one particular commercial UNIX and could not change without vast disruptions.


    We felt this was weird, because the UNIXes of the era were all pretty similar. It gradually became clear that manufacturers' training
    courses on commercial UNIXes emphasised the differences and tried
    hard to give the impression that changing to a another supplier would
    be difficult and expensive. Human inertia meant that customers'
    sysadmins co-operated with that, giving the illusion that they were indispensable while acting against their employers' interests. This
    meant that well-established manufacturers like HP and IBM felt less
    urgency to move to 64-bit. Sun and SGI were a bit more dynamic, until
    they got into financial trouble, and DEC knew they needed Alpha to
    have a hope of survival.

    Meanwhile, we employed sysadmins who dealt with all of these UNIXes
    every week without difficulty. We also had to learn about things like
    POWER and PA-RISC segmentation to squeeze large 32-bit applications
    onto them. 64-bit was very welcome!

    John

    HP introduced their first 64-bit PA-RISC CPU only 5 or 6 months behind
    Sun.

    IBM released the Cobra, their 1st 64-bit POWEER CPU, approximately at
    the same time as Sun. But it and its successors (Muskie, Apache,
    Northstar) were probably of little interest for you and your employer,
    because they were not intended for engineering applications.
    The first engineering-oriented 64-bit POWER (POWER3) came, indeed, significantly later - more than 3 years behind Sun.







    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 18 06:10:53 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:
    HP introduced their first 64-bit PA-RISC CPU only 5 or 6 months behind
    Sun.

    According to Wikipedia:
    |The UltraSPARC [...] introduced in mid-1995 [...] is the first
    |microprocessor from Sun to implement the 64-bit SPARC V9 instruction
    |set

    But SPARC did not have the segments of HPPA and Power, and therefore
    had less pressure to extend virtual addresses early on. Yet they were
    earlier.

    Depending on which Wikipedia page you look at, PA-8000 and PA-RISC 2.0
    was introduced on November 2, 1995
    <https://en.wikipedia.org/wiki/PA-8000>, or in January 1996 <https://en.wikipedia.org/wiki/PA-RISC>.

    IBM released the Cobra, their 1st 64-bit POWEER CPU, approximately at
    the same time as Sun. But it and its successors (Muskie, Apache,
    Northstar) were probably of little interest for you and your employer, >because they were not intended for engineering applications.
    The first engineering-oriented 64-bit POWER (POWER3) came, indeed, >significantly later - more than 3 years behind Sun.

    The A10/Cobra was only released in AS/400 familt machines. However,
    the RS64 was also released in RS/6000 family machines, in 1997. IBM
    did not use the PowerPC 620, also released in 1997. Power3/PowerPC
    630 was launched in 1998.

    Maybe the HP/UX and AIX developers implemented ways that worked around
    the 256MB segment size, e.g., several segments for the same virtual
    memory area, and I guess that the Linux developers did so, too.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 18 15:58:03 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    ordinary programmers would have
    difficulty writing good code for the Itanium, and those that can are
    hard to find.

    Of course they are hard to find, given that nobody programs them; that
    does not mean it is more difficult for humans than for the machines.

    Instead, the fact that
    instructions come in blocks of three, with some instructions only
    being available in certain slots, is what I saw as the biggest
    conceptual difficulty for an ordinary human programmer.

    Looks easy to me, for both humans and machines. So the assembler can
    do it; not sure how much the assembler did, but it did at least some
    of the work.

    The more difficult part of the work is to schedule the instructions
    beyond simple loops, in particular using the speculative features
    effectively.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Fri Sep 18 16:38:01 2026
    From Newsgroup: comp.arch

    Paul Clayton [2026-09-16 13:14:22] wrote:
    I suspect algorithmic transformations are discouraged in
    compilers under the assumption that the software developer used
    appropriate algorithms as well as the compute cost to evaluate
    options.

    Also because compilers generally can't (or at least shouldn't) take the
    chance of generating significantly slower code.

    The purpose of "language + compiler" is not to figure out magically how
    to implement the most efficient code that solves the same problem,
    instead it's to allow the programmer to write that most efficient
    code conveniently.

    Of course, there is a lot of commercial interest in doing the magic
    thing so as to save the work of understanding and rewriting the code to
    improve performance (hence all the work on auto-parallelization), but experience shows that it's an extremely difficult problem.

    Luckily, a lot of that work on trying to do the magic thing can also be
    used to solve the real problem: what the years of work on
    compiler-optimization has brought (instead of magically speeding up old
    code) is to improve the convenience to write efficient code.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David LaRue@huey.dll@tampabay.rr.com to comp.arch on Sat Sep 19 06:01:05 2026
    From Newsgroup: comp.arch

    Stefan Monnier <monnier@iro.umontreal.ca> wrote in news:jwva4peqp52.fsf- monnier+comp.arch@gnu.org:

    Paul Clayton [2026-09-16 13:14:22] wrote:
    I suspect algorithmic transformations are discouraged in
    compilers under the assumption that the software developer used
    appropriate algorithms as well as the compute cost to evaluate
    options.

    Also because compilers generally can't (or at least shouldn't) take the chance of generating significantly slower code.

    The purpose of "language + compiler" is not to figure out magically how
    to implement the most efficient code that solves the same problem,
    instead it's to allow the programmer to write that most efficient
    code conveniently.

    Of course, there is a lot of commercial interest in doing the magic
    thing so as to save the work of understanding and rewriting the code to improve performance (hence all the work on auto-parallelization), but experience shows that it's an extremely difficult problem.

    Luckily, a lot of that work on trying to do the magic thing can also be
    used to solve the real problem: what the years of work on compiler-optimization has brought (instead of magically speeding up old
    code) is to improve the convenience to write efficient code.


    === Stefan


    Remember too, that some lower languages and especially older computers actually had code meant to take a certain amount of time. We depended on knowing time constraints. That knowledge isn't always honored by a well meaning compiler writer and can have disasterous results if the optimized
    for time code is put into place and not discovered in time.

    Even modern computers depend on knowing how the code will really run.
    Anyone who has had the joy of figuring out why tinkering with code or
    hardware suddenly had unextpected results can attest to.

    Knowing how to properly interface with hardware is one of the joys of low level programming. Everything mattered, even time, to that developer.

    David -- still enjoying peering inside things and fixing them
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 19 06:25:13 2026
    From Newsgroup: comp.arch

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Paul Clayton [2026-09-16 13:14:22] wrote:
    Also because compilers generally can't (or at least shouldn't) take the >chance of generating significantly slower code.

    One would hope so, but:

    |In the case of the bubble-sort benchmark (c-manual/bubble-sort.c in |bench.zip), on a Rocket Lake (Xeon E-2388G) the auto-vectorization
    |leads to a slowdown by a factor of 5.7 of gcc-14.2 -O3 over gcc-14.2
    |-O3 -fno-tree-vectorize, by a factor of 5.6 over clang-19.1 -O3 (where
    |clang does not vectorize anything in this benchmark), by a factor of
    |5.3 over gcc-14.2 -O, and by a factor of 2.1 over gcc-14.2 -O0.

    Read more about the reasons on
    <http://www.complang.tuwien.ac.at/anton/stwlf/>.

    The purpose of "language + compiler" is not to figure out magically how
    to implement the most efficient code that solves the same problem,
    instead it's to allow the programmer to write that most efficient
    code conveniently.

    I agree.

    Of course, there is a lot of commercial interest in doing the magic
    thing so as to save the work of understanding and rewriting the code to >improve performance (hence all the work on auto-parallelization), but >experience shows that it's an extremely difficult problem.

    Experience shows that it is an easy problem to get people to put money
    into your project by promising them that the compiler will magically
    solve it (and thus they can save on expensive programmer time);
    earlier successful ways to convert wishful thinking into money are the philosopher's stone, and a contemporary way is AI. In the case of
    optimizing compilers it helps that one can present showpieces where it
    works as promised. I guess the alchemists had similar tricks (maybe
    they sold chemical reactions as first steps to the desired end
    result), and we can see in daily news how the AI companies do it.

    Luckily, a lot of that work on trying to do the magic thing can also be
    used to solve the real problem: what the years of work on >compiler-optimization has brought (instead of magically speeding up old
    code) is to improve the convenience to write efficient code.

    It depends. When the optimizer writers, in their mission to achieve
    good benchmark results, put the assumption in their optimizers that
    programs do not exercise undefined behaviour (except the cases that
    occur in the benchmarks, e.g., the SATD function in 464.h264ref) and
    "optimize" programs with that assumption. They tell programmers who
    are hit by this that their previously working programs are buggy and
    have always been broken, and that one must not write such code.

    This does not improve the convenience of writing efficient code, on
    the contrary: In Section 4.3 of <https://www.complang.tuwien.ac.at/kps2015/proceedings/KPS_2015_submission_29.pdf>
    I estimate that Gforth would slow down by a factor >3 compared to
    Gforth 0.7.0 from a set of changes for avoiding undefined behaviour
    where I could make an estimate of the performance effects; we would
    have to introduce other changes as well, which would also affect
    performance negatively.

    And not just for Gforth, also for other programs, it has a chilling
    effect on programmer optimizations if the programmers always have to
    be on the lookout for undefined behaviour instead of concentrating on optimizing the program.

    But if optimizer writers strove to "improve the convenience to write
    efficient code" instead of just improving benchmark results, maybe
    such paradoxical effects can be avoided.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 19 12:06:36 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> schrieb:

    I suspect algorithmic transformations are discouraged in
    compilers under the assumption that the software developer used
    appropriate algorithms as well as the compute cost to evaluate
    options.

    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results. Unless directed otherwise, the compiler has to assume
    that this is what the user wants and needs.

    An exception is something like Fortran's FORALL constuct, where
    the loop ordering is not specified and the compiler can chose.
    Intrinsic functions like MATMUL also have no restriction on what
    they can do internally; they can (and often do) call a highly-
    optimized BLAS routine.

    Usage and hardware information might theoretically be
    made available, but I get the impression that most profiling
    for compiler use is about path frequency.

    Compilers tend to focus on lower-level stuff. For example,
    there is a huge file of simplifying transformations for gcc at https://gcc.gnu.org/git/?p=gcc.git;a=blob_plain;f=gcc/match.pd

    This now also contains reverse-engineered bithacks. For example,
    the popcnt() method from Hacker's Delight are now recognized and
    translated into an internal function, which is then expanded
    into either an inlined function (much like the original) or
    a machine instruction, if the target has one.

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    But if you're willing to take the risk of AI, you can of course
    tell it "Find all O(N^2) algorithms in my code and replace them
    with O(N log N) where possible". It will happily do something,
    and the resulting code might even be correct after a few
    iterations.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 19 15:56:49 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Paul Clayton <paaronclayton@gmail.com> schrieb:
    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results.

    Interestingly, this is where gcc maintainers do the right thing and
    people who want to risk program breakage must use an extra flag
    (-ffast-math) to get "optimizations" (such as FP operation
    reassociation) that may change the results.

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    This bubble-sort benchmark comes from John Hennessey's collection of
    integer benchmarks, and Marty Fraeman has translated several of these benchmarks into Forth. This is the reason why I like to use it for
    comparing the performance of Forth systems to C compilers.

    Recognizing this bubble-sort and using a more efficient sort (e.g.,
    insertion sort for small instances, quicksort for larger ones) would
    be a real optimization; it would ruin the comparability with the Forth translation (as long as we do not put this kind of effort into Forth compilers), but that's life. But for now gcc autovectorizes it in a
    way that produces a significant slowdown; clang avoids this pitfall.

    Note that this auto-vectorization of gcc does not just slow down
    bubble-sort, but also gforth: when you load two adjacent stack items
    in one VM instruction, gcc -O3 by default wants to auto-vectorize
    these accesses, and given that one or both of them usually have been
    written recently, this would result in a slowdown. What's worse, gcc
    seems to think that a lot of values are alive that are actually dead,
    and generates dozens or hundreds of copying instructions into the code
    of each VM instruction, increasing the slowdown even more. For that
    reason, we have disabled tree-slp-vectorization when building Gforth.

    But if you're willing to take the risk of AI, you can of course
    tell it "Find all O(N^2) algorithms in my code and replace them
    with O(N log N) where possible".

    Most O(N^2) algorithms probably cannot be replaced in that way, e.g.,
    the TSP algorithm used in "Writing Efficient Programs" by Jon Bentley.

    OTOH, as a human you can replace this TSP algorithm by one that runs
    faster (O(N log N) or so) and also usually produces a shorter tour,
    and IIRC Jon Bentley mentions this himself, but it will produce a
    different result, so a compiler cannot go there: It has to use the
    source code as specification.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 20 10:46:02 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Paul Clayton <paaronclayton@gmail.com> schrieb:
    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results.

    Interestingly, this is where gcc maintainers do the right thing and
    people who want to risk program breakage must use an extra flag
    (-ffast-math) to get "optimizations" (such as FP operation
    reassociation) that may change the results.

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    This bubble-sort benchmark comes from John Hennessey's collection of
    integer benchmarks, and Marty Fraeman has translated several of these benchmarks into Forth. This is the reason why I like to use it for
    comparing the performance of Forth systems to C compilers.

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    Recognizing this bubble-sort and using a more efficient sort (e.g.,
    insertion sort for small instances, quicksort for larger ones) would
    be a real optimization; it would ruin the comparability with the Forth translation (as long as we do not put this kind of effort into Forth compilers), but that's life. But for now gcc autovectorizes it in a
    way that produces a significant slowdown; clang avoids this pitfall.

    All such choices are the result of heuristics. Bubble sort has a
    very special memory access pattern. A straightforward patch would
    very likely pessimize a lot of existing code which profits from auto-vectorization. I have no doubt that the existing heuristics
    can be improved.

    Note that this auto-vectorization of gcc does not just slow down
    bubble-sort, but also gforth: when you load two adjacent stack items
    in one VM instruction, gcc -O3 by default wants to auto-vectorize
    these accesses, and given that one or both of them usually have been
    written recently, this would result in a slowdown. What's worse, gcc
    seems to think that a lot of values are alive that are actually dead,
    and generates dozens or hundreds of copying instructions into the code
    of each VM instruction, increasing the slowdown even more. For that
    reason, we have disabled tree-slp-vectorization when building Gforth.

    There is no stopping you from developing a patch (or having
    it developed by a student as part of a thesis - can you act as
    supervisor for master's or bachelor's theses?) and then submitting
    it to gcc. But if you do so, you should make sure that it is does
    not slow down normal programs (as measured by SPEC).

    Regarding clang: There are numerous cases where gcc's more
    aggressive auto-vectorization generates a lot of profit vs. clang.
    These are cases that would very probably suffer when heuristics
    are not very carefully chosen.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 20 14:24:21 2026
    From Newsgroup: comp.arch

    On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Paul Clayton <paaronclayton@gmail.com> schrieb:
    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results.

    Interestingly, this is where gcc maintainers do the right thing and
    people who want to risk program breakage must use an extra flag (-ffast-math) to get "optimizations" (such as FP operation
    reassociation) that may change the results.

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    This bubble-sort benchmark comes from John Hennessey's collection of integer benchmarks, and Marty Fraeman has translated several of
    these benchmarks into Forth. This is the reason why I like to use
    it for comparing the performance of Forth systems to C compilers.

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    Recognizing this bubble-sort and using a more efficient sort (e.g., insertion sort for small instances, quicksort for larger ones) would
    be a real optimization; it would ruin the comparability with the
    Forth translation (as long as we do not put this kind of effort
    into Forth compilers), but that's life. But for now gcc
    autovectorizes it in a way that produces a significant slowdown;
    clang avoids this pitfall.

    All such choices are the result of heuristics. Bubble sort has a
    very special memory access pattern. A straightforward patch would
    very likely pessimize a lot of existing code which profits from auto-vectorization. I have no doubt that the existing heuristics
    can be improved.

    Note that this auto-vectorization of gcc does not just slow down bubble-sort, but also gforth: when you load two adjacent stack items
    in one VM instruction, gcc -O3 by default wants to auto-vectorize
    these accesses, and given that one or both of them usually have been written recently, this would result in a slowdown. What's worse,
    gcc seems to think that a lot of values are alive that are actually
    dead, and generates dozens or hundreds of copying instructions into
    the code of each VM instruction, increasing the slowdown even more.
    For that reason, we have disabled tree-slp-vectorization when
    building Gforth.

    There is no stopping you from developing a patch (or having
    it developed by a student as part of a thesis - can you act as
    supervisor for master's or bachelor's theses?) and then submitting
    it to gcc. But if you do so, you should make sure that it is does
    not slow down normal programs (as measured by SPEC).

    Regarding clang: There are numerous cases where gcc's more
    aggressive auto-vectorization generates a lot of profit vs. clang.
    These are cases that would very probably suffer when heuristics
    are not very carefully chosen.


    I am not sure about vectorization in general. Personally, I had seen
    cases where clang autovectorization produced greater slowdowns than
    gcc's.
    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones. Nowadays gcc does not even
    consider it autovectorization. This pattern rarely leads to major
    slowdowns, more typically a little impacts in multiple places. But
    sometimes it gets measurable, esp. on Zen3.
    I don't believe that disabling this particular generic pessimization
    could impact Spec negatively, but I am not aware of simple way (i.e.
    normal -f flags, rather than semi-documented flags intended for gcc maintainers) to disable it.





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 20 11:52:38 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones. Nowadays gcc does not even
    consider it autovectorization. This pattern rarely leads to major
    slowdowns, more typically a little impacts in multiple places. But
    sometimes it gets measurable, esp. on Zen3.
    I don't believe that disabling this particular generic pessimization
    could impact Spec negatively, but I am not aware of simple way (i.e.
    normal -f flags, rather than semi-documented flags intended for gcc maintainers) to disable it.

    If you have a test case (but please not bubble sort :-) where
    -fstore-merging pessimizes things (or -fno-store-merging makes
    a significant difference) please post it. I can then write a PR
    and hang it off https://gcc.gnu.org/bugzilla/show_bug.cgi?id=94094
    where I don't see anything relevant at the moment.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 20 15:09:18 2026
    From Newsgroup: comp.arch

    On Sun, 20 Sep 2026 11:52:38 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones. Nowadays gcc does not even
    consider it autovectorization. This pattern rarely leads to major slowdowns, more typically a little impacts in multiple places. But sometimes it gets measurable, esp. on Zen3.
    I don't believe that disabling this particular generic pessimization
    could impact Spec negatively, but I am not aware of simple way (i.e.
    normal -f flags, rather than semi-documented flags intended for gcc maintainers) to disable it.

    If you have a test case (but please not bubble sort :-) where
    -fstore-merging pessimizes things (or -fno-store-merging makes
    a significant difference) please post it. I can then write a PR
    and hang it off https://gcc.gnu.org/bugzilla/show_bug.cgi?id=94094
    where I don't see anything relevant at the moment.


    IIRC, it already was in at least one of my PRs. May be, more than one.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sun Sep 20 11:49:14 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Paul Clayton <paaronclayton@gmail.com> schrieb:
    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results.

    Interestingly, this is where gcc maintainers do the right thing and
    people who want to risk program breakage must use an extra flag
    (-ffast-math) to get "optimizations" (such as FP operation
    reassociation) that may change the results.

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    This bubble-sort benchmark comes from John Hennessey's collection of
    integer benchmarks, and Marty Fraeman has translated several of these
    benchmarks into Forth. This is the reason why I like to use it for
    comparing the performance of Forth systems to C compilers.

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    So what?

    Recognizing this bubble-sort and using a more efficient sort (e.g.,
    insertion sort for small instances, quicksort for larger ones) would
    be a real optimization; it would ruin the comparability with the Forth
    translation (as long as we do not put this kind of effort into Forth
    compilers), but that's life. But for now gcc autovectorizes it in a
    way that produces a significant slowdown; clang avoids this pitfall.

    All such choices are the result of heuristics. Bubble sort has a
    very special memory access pattern. A straightforward patch would
    very likely pessimize a lot of existing code which profits from >auto-vectorization. I have no doubt that the existing heuristics
    can be improved.

    Oh really? Loading two adjacent memory locations in a way that the
    compiler can see and (if the maintainer is sufficiently clueless)
    merge after storing to one of these locations, with some optimization
    barrier in between (control flow in the case of bubble-sort) is very
    special? How come it also occurs in Gforth?

    Note that this auto-vectorization of gcc does not just slow down
    bubble-sort, but also gforth: when you load two adjacent stack items
    in one VM instruction, gcc -O3 by default wants to auto-vectorize
    these accesses, and given that one or both of them usually have been
    written recently, this would result in a slowdown. What's worse, gcc
    seems to think that a lot of values are alive that are actually dead,
    and generates dozens or hundreds of copying instructions into the code
    of each VM instruction, increasing the slowdown even more. For that
    reason, we have disabled tree-slp-vectorization when building Gforth.

    There is no stopping you from developing a patch (or having
    it developed by a student as part of a thesis - can you act as
    supervisor for master's or bachelor's theses?) and then submitting
    it to gcc.

    Sure, but nobody is paying me or the student to do it, either. So I
    leave it to those people who are paid to do it.

    Also, I don't think that flipping the default for
    tree-slp-vectorization off is a worthwhile thesis topic. Admittedly,
    there may be cases where tree-slp-vectorization would vectorize even
    if the vectorization for loads is turned off, but I expect it to be
    rare.

    But if you do so, you should make sure that it is does
    not slow down normal programs (as measured by SPEC).

    The parenthetical remark is revealing. While SPEC CPU may intend to
    be representative of a wide range of programs, it's still just a bunch
    of programs with a bunch of performance-critical hot spots (typically
    inner loops), and if these happen to avoid a performance pitfall like store-to-wide-load forwarding, it is doubtful that they are
    representative in this respect.

    Regarding clang: There are numerous cases where gcc's more
    aggressive auto-vectorization generates a lot of profit vs. clang.
    These are cases that would very probably suffer when heuristics
    are not very carefully chosen.

    My experiments with auto-vectorization over the last year with both
    gcc and clang contradict your claim that gcc is more aggressive in auto-vectorization in general. E.g., I tested when either compilers
    would stop auto-vectorizing loops due to potential aliases, and I have
    not found a limit for clang, while I have for gcc (I stopped when I
    ran out of registers for the arrays); i.e., clang is more aggressive
    here, at the cost of needing more checking for overlapping arrays
    before entering the loop.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sun Sep 20 12:13:57 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:
    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones.

    I have seen heavy slowdowns from merging narrow loads into wide loads <https://www.complang.tuwien.ac.at/anton/stwlf/>, but not for stores.
    Where can I read about this slowdown?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 20 15:35:43 2026
    From Newsgroup: comp.arch

    On Sun, 20 Sep 2026 12:13:57 GMT
    anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    Michael S <already5chosen@yahoo.com> writes:
    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones.

    I have seen heavy slowdowns from merging narrow loads into wide loads <https://www.complang.tuwien.ac.at/anton/stwlf/>, but not for stores.
    Where can I read about this slowdown?

    - anton

    Those are my measurements, mostly undocumented and dependent on
    situation. I never did a deep research.
    Typically cases of none-SIMD stores merged into SIMD.
    If store location is not used soon thereafter then impact is hard to
    detect. In this cases the only impact is code bloat and over-occupation
    of Ifetch and Decode stages.
    It seems to me, the most heavy impact on Zen3 is when such merged
    SIMD store is followed by GPR load to the same location within a dozen
    or so of CPU cycles.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Sun Sep 20 07:31:08 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    For people who don't like bubble sort, or for inclusion in a set
    of benchmarks, I suggest the following recently discovered
    sorting algorithm (written in C-ish pseudocode):

    void
    baffle_sort( unsigned n, int *elements ){
    for( unsigned i = 0; i < n; i++ ){
    for( unsigned j = 0; j < n; j++ ){
    if( elements[i] < elements[j] ){
    /* exchange elements[i] and elements[j] */
    }
    }
    }
    }

    (Disclaimer: not original.)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Sun Sep 20 02:56:52 2026
    From Newsgroup: comp.arch

    Anton Ertl [2026-09-19 06:25:13] wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Paul Clayton [2026-09-16 13:14:22] wrote:
    Also because compilers generally can't (or at least shouldn't) take the >>chance of generating significantly slower code.
    One would hope so, but:

    Indeed, in practice virtually all compiler "optimizations" are valid
    only statistically: in most cases they either have no measurable effect
    or they improve some characteristic, but there are almost always corner
    cases where they make things worse.

    The traditional "-O<N>" flags are a way to state how much you're willing
    to get worse performance in exchange for the opportunity to maybe get
    better performance.

    Experience shows that it is an easy problem to get people to put money
    into your project by promising them that the compiler will magically
    solve it (and thus they can save on expensive programmer time);
    earlier successful ways to convert wishful thinking into money are the philosopher's stone, and a contemporary way is AI. In the case of
    optimizing compilers it helps that one can present showpieces where it
    works as promised. I guess the alchemists had similar tricks (maybe
    they sold chemical reactions as first steps to the desired end
    result), and we can see in daily news how the AI companies do it.

    +1

    Luckily, a lot of that work on trying to do the magic thing can also be >>used to solve the real problem: what the years of work on >>compiler-optimization has brought (instead of magically speeding up old >>code) is to improve the convenience to write efficient code.
    It depends.

    "Can be used" doesn't mean it's necessarily used that way, indeed.

    But if optimizer writers strove to "improve the convenience to write efficient code" instead of just improving benchmark results, maybe
    such paradoxical effects can be avoided.

    In general it's hard to completely avoid paradoxical effects. But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general, but not to specify anything vaguely related to
    "undefined behavior", so we end up with this bogus notion that "UB
    implies anything you want" (including "Minority Report" absurdities
    where we're allowed to do anything as soon as we know that UB *would*
    happen anyway).


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sun Sep 20 16:43:54 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:
    It seems to me, the most heavy impact on Zen3 is when such merged
    SIMD store is followed by GPR load to the same location within a dozen
    or so of CPU cycles.

    This is a case I measured in
    <https://www.complang.tuwien.ac.at/anton/stwlf/>. Look for the
    ws=_=>nl>ns line. In this case, on Zen3 merging (-O3) is faster
    than not merging (-O).

    But I see slowdowns from merging (-O3) in other cases where the
    partial store-to-load overlap does not occur, e.g., in the "wl>ws=>wl (recurrence), nl>ns (no recurrence)" case (and that's pretty
    widespread among microarchitectures), so there is probably another microarchitectural pitfall beyond the partal store-to-load forwarding.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From John Levine@johnl@taugh.com to comp.arch on Sun Sep 20 20:12:42 2026
    From Newsgroup: comp.arch

    According to Stefan Monnier <monnier@iro.umontreal.ca>:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Paul Clayton [2026-09-16 13:14:22] wrote:
    Also because compilers generally can't (or at least shouldn't) take the >>>chance of generating significantly slower code.
    One would hope so, but:

    Indeed, in practice virtually all compiler "optimizations" are valid
    only statistically: in most cases they either have no measurable effect
    or they improve some characteristic, but there are almost always corner
    cases where they make things worse.

    There are plenty of optimizations that always make things better, e.g., removing dead code, or reusing values in registers. But these days those
    are so obvious we sometimes forget about them.

    I agree that the more sophisticated the optiomization, the more likely it
    is to have perverse cases.
    --
    Regards,
    John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
    Please consider the environment before reading this e-mail. https://jl.ly
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 21 05:27:23 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 20 Sep 2026 11:52:38 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 20 Sep 2026 10:46:02 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    But one particular pattern of gcc is certainly worse than clang's -
    merging narrow stores into wider ones. Nowadays gcc does not even
    consider it autovectorization. This pattern rarely leads to major
    slowdowns, more typically a little impacts in multiple places. But
    sometimes it gets measurable, esp. on Zen3.
    I don't believe that disabling this particular generic pessimization
    could impact Spec negatively, but I am not aware of simple way (i.e.
    normal -f flags, rather than semi-documented flags intended for gcc
    maintainers) to disable it.

    If you have a test case (but please not bubble sort :-) where
    -fstore-merging pessimizes things (or -fno-store-merging makes
    a significant difference) please post it. I can then write a PR
    and hang it off https://gcc.gnu.org/bugzilla/show_bug.cgi?id=94094
    where I don't see anything relevant at the moment.


    IIRC, it already was in at least one of my PRs. May be, more than one.

    The way I read it, it is the opposite - those are about missed
    load/store merging. Here is the list.

    94094: [meta-bug] store-merging and/or bswap load/store-merging missed optimizations [See dependency tree for bug 94094]

    32605: Missing byte swap optimizations [See dependency tree for bug 32605]
    65424: gcc does not recognize byte swaps implemented as loop. [See dependency tree for bug 65424]
    88097: Missing optimization of endian conversion [See dependency tree for bug 88097]
    89810: Suboptimal codegen: integer load/assemble from in-register array of uint8_t [See dependency tree for bug 89810]
    89811: uint32_t load is not recognized if shifts are done in a fixed-size loop [See dependency tree for bug 89811]
    92716: -Os doesn't inline byteswap function even though it's a single instruction [See dependency tree for bug 92716]
    93328: missed optimization opportunity in deserialization code [See dependency tree for bug 93328]
    93896: Store merging could merge 0 constructors into {} [See dependency tree for bug 93896]
    94071: Missed optimization with endian and alignment independent memory access [See dependency tree for bug 94071]
    94086: Missed optimization when converting a bitfield to an integer on x86-64 [See dependency tree for bug 94086]
    94403: Missed optimization bswap [See dependency tree for bug 94403]
    96135: [14/15/16/17 regression] bswap not detected by bswap pass, unexpected results between optimization levels [See dependency tree for bug 96135]
    96167: fails to detect ROL pattern in simple case, but succeeds when operand goes through memcpy [See dependency tree for bug 96167]
    98953: Failure to optimize two reads from adjacent addresses into one due to having an offset (index) [See dependency tree for bug 98953]
    94071: Missed optimization with endian and alignment independent memory access [See dependency tree for bug 94071] (*)
    98982: Optimizing loop variants of fixed-byte-order functions [See dependency tree for bug 98982]
    102495: optimize some consecutive byte load pattern to word load [See dependency tree for bug 102495]
    104344: Suboptimal -Os code for manually unrolled loop [See dependency tree for bug 104344]
    104632: Missed optimization about reading backwards [See dependency tree for bug 104632]
    123541: Load merging missing [See dependency tree for bug 123541]

    So, if anybody is interested in aan actual fix, please post a
    self-contained test case.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 21 15:16:15 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    All such choices are the result of heuristics. Bubble sort has a
    very special memory access pattern.

    Ok, what would a memory pattern look like that this vectorization was
    designed for? The big slowdown comes from the fact that in the
    reordering case the vectorized version performs a wide store, and in
    the next iteration it performs a wide load that partially overlaps the
    store. I buy it that the analysis will have trouble seeing that (not
    that it's impossible in this case, but it requires effort).

    A straightforward patch would
    very likely pessimize a lot of existing code which profits from >auto-vectorization.

    How can we test this claim? What would such code look like that would
    suffer from disabling the combining of loads that tree-slp-vectorize
    performs?

    I decided to test the claim by continuing to use bubble-sort, but
    avoiding the stores and thus the big slowdown: Bubble-sort a
    pre-sorted array. Apart from changing the array size, initializing
    the array with pre-sorted values, changing the indentation and
    condensing the code to one benchmark only (stan.c has a number of
    benchmarks in one file, probably due to Pascal origins), the code is
    unchanged; the original code already performs N iterations of the
    outer loop (rather than doing an early-out on an already-sorted
    array), so one does not need to change that aspect.

    The inner loop is compiled by gcc-14.2 to:

    novect (-O3 -fno-tree-slp-vectorize) vect (-O3)
    70: mov (%rax),%ecx 90: movq (%rax),%xmm0
    mov 0x4(%rax),%esi pshufd $0xe5,%xmm0,%xmm1
    movd %xmm0,%esi
    movd %xmm1,%ecx
    cmp %esi,%ecx cmp %ecx,%esi
    jle 7e jle ae #always taken
    mov %esi,(%rax) pshufd $0xe1,%xmm0,%xmm0 #x
    mov %ecx,0x4(%rax) movq %xmm0,(%rax) #x
    7e: add $0x4,%rax ae: add $0x4,%rax
    cmp %rax,%rdi cmp %rdi,%rax
    jne 70 jne 90

    In the presorted benchmark, the jle branch is always taken, and the
    store instructions marked #x are never executed, so the
    microarchitectural pitfall never occurs.

    The novect code executes 7 instructions per iteration, the vect code
    9, but maybe the vectorization manages to make the code execute faster
    for some reason. Let's see (on a Xeon E-2388G (Rocket Lake):

    cycles instructions
    1,867,053,886 12,602,009,442 novect
    3,603,774,972 16,202,621,392 vect

    So the novect variant is twice as fast as the vect variant even
    without the slowdown from the store-to-partially-overlapping loads.
    If there is a speedup to be had from combining loads with
    tree-slp-vectorize, it hides itself well.

    The performance of the Rocket Lake for the novect variant is
    remarkable. It manages 6.75 IPC on average, which is very close to
    the speed of light for this microarchitecture, with 5 renamer slots,
    with compare-and-branch macro-instructions occupying only one renamer
    slot (so the speed of light for a loop with 5 non-branches and 2
    branches is 7 IPC). Top-down microarchitectural analysis reports:

    TopdownL1 # 2.7 % tma_backend_bound
    # 1.6 % tma_bad_speculation
    # -0.0 % tma_frontend_bound
    # 95.7 % tma_retiring

    The ideal is that 100% of the rename slots are consumed by retiring,
    but the normal case is that many of the slots are lost due to the
    other three reasons, and retiring numbers >50% are already pretty
    good.

    The Rocket Lake manages two taken branches per cycle here, while I had
    thought it was only capable of doing one non-taken and one taken
    branch per cycle.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Mon Sep 21 20:23:16 2026
    From Newsgroup: comp.arch

    On Sun, 20 Sep 2026 16:43:54 GMT
    anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    Michael S <already5chosen@yahoo.com> writes:
    It seems to me, the most heavy impact on Zen3 is when such merged
    SIMD store is followed by GPR load to the same location within a
    dozen or so of CPU cycles.

    This is a case I measured in <https://www.complang.tuwien.ac.at/anton/stwlf/>. Look for the wl>ws=_=>nl>ns line. In this case, on Zen3 merging (-O3) is faster
    than not merging (-O).

    But I see slowdowns from merging (-O3) in other cases where the
    partial store-to-load overlap does not occur, e.g., in the "wl>ws=>wl (recurrence), nl>ns (no recurrence)" case (and that's pretty
    widespread among microarchitectures), so there is probably another microarchitectural pitfall beyond the partal store-to-load forwarding.

    - anton

    The truth is that I don't know exact conditions.
    Probably can dig test cases out of old projects if there is a real
    interest from gcc maintainers. One example where I had seen it was
    emulation of binary128 floating point, either fmul or fadd, I can't
    recollect which of the two.
    Although the whole enterprise with my failed attempt to give away to gcc
    much faster binary128 arithmetic was rather depressing so I'd prefer
    to find example from another field, but right now can not recollect
    anything else.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 21 17:17:19 2026
    From Newsgroup: comp.arch

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    [...]
    The traditional "-O<N>" flags are a way to state how much you're willing
    to get worse performance in exchange for the opportunity to maybe get
    better performance.

    The gcc manual states:

    With '-O', the compiler tries to reduce code size and execution
    time, without performing any optimizations that take a great deal
    of compilation time.
    [...]
    [about -O2] As
    compared to '-O', this option increases both compilation time and
    the performance of the generated code.
    [...]
    [about -O3] Optimize yet more.

    In earlier gcc versions, it made statements along the lines that -O3
    may be a mixed bag, but that's gone.

    But if optimizer writers strove to "improve the convenience to write
    efficient code" instead of just improving benchmark results, maybe
    such paradoxical effects can be avoided.

    In general it's hard to completely avoid paradoxical effects.

    I referred to a specific paradoxical effect, not completely avoiding
    all of them: The effect of programmers who, instead of working on
    making the code faster through source-level changes, have to invest
    time into avoiding getting it miscompiled.

    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows. It gives a speedup in some SPEC
    program (How much? Don't know, but I am sure it exists); ok, it
    breaks this other SPEC program, let's special-case the optimization to
    avoid this breakage.

    My impression is that the C compiler maintainers and some others have
    fallen in love with the idea of optimizing based on assuming that
    undefined behaviour does not happen; e.g., I have read the claim that
    this is the only thing that gives C an edge over other programming
    languages (or somesuch). So I think that while the C89 committee may
    not have thought of such things, recent C standardization committees
    probably have.

    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus: <https://blog.regehr.org/archives/1287>.

    But fortunately a completely specified C is not necessary, a
    willingness to preserve the behaviour of existing working programs
    compiled with an earlier version of the same compiler is. Read more
    about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 21 20:06:23 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    -------------
    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    Does anyone know if Ada was successful about not leaving parts
    unspecified ??

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows. It gives a speedup in some SPEC
    program (How much? Don't know, but I am sure it exists); ok, it
    breaks this other SPEC program, let's special-case the optimization to
    avoid this breakage.

    This is a problem with SPEC not the compilers for various machines.
    A broad spectrum benchmark should not contain code that relies on
    unspecified or undefined behaviors. I remember Whetstone on the S.E.L.
    machines had a call to EXP( float ) with a constant value and expected
    IBM quality result, whereas on the S.E.L. machine the operand constant
    was such that (hex alignment) could NEVER produce the expected result
    to the accuracy the benchmark was expecting. In all other regards S.E.L.
    FP was equivalent to IBM 360 FP arithmetic.

    One would hope that IEEE 754 put an end to it, but BFP8 and FP8
    raises this spectr|- again.

    My impression is that the C compiler maintainers and some others have
    fallen in love with the idea of optimizing based on assuming that
    undefined behaviour does not happen; e.g., I have read the claim that
    this is the only thing that gives C an edge over other programming
    languages (or somesuch). So I think that while the C89 committee may
    not have thought of such things, recent C standardization committees
    probably have.

    How much more optimizations can compilers deliver to the bottom line
    (not just benchmarks). Last week we saw an episode where the compilers
    only got 1.x% speedup over years (sub-decade). Is there ever going to
    be a time to stop ??

    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus: <https://blog.regehr.org/archives/1287>.

    But fortunately a completely specified C is not necessary, a
    willingness to preserve the behaviour of existing working programs
    compiled with an earlier version of the same compiler is. Read more
    about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.

    Some of the machines have egregious behaviors in tiny little corners
    of the architecture(s) and implementation(s). Not being able to run
    the same binary with and without integer overflow detection as an example.
    And even when they can, they cannot perform Ada ADD Byte and take a signed overflow trap on 8-bit results easily.

    So, much of the blame is cast back on architects of those machines.

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Mon Sep 21 23:48:13 2026
    From Newsgroup: comp.arch

    On Mon, 21 Sep 2026 20:06:23 GMT
    MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    -------------
    But as
    for the problems you mention w.r.t UB, I think it's just the
    result of poor semantics, for which I guess we (language
    semanticists) are partly to blame: we have developed fairly good
    tools to design sane language specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people
    were poor at specifying semantics, but because given the differences between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts
    unspecified.

    Does anyone know if Ada was successful about not leaving parts
    unspecified ??


    I think that Ada never had such goal.
    It specified more than contemporaries (at least those that pretended
    to be "system" level languges) and than nearly all successors and
    considered that it is good enough.
    Some corner cases were considered impractical to specify precisely.
    E.g x := a + b + c, where x, a, b and c are 32-bit integers, a =
    0x7fff_ffff, b = 1, c = -2. It may raise exception or silintly produce
    correct answer. Both behaviors are acceptable by Ada spec.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Mon Sep 21 12:48:10 2026
    From Newsgroup: comp.arch

    John Levine [2026-09-20 20:12:42] wrote:
    According to Stefan Monnier <monnier@iro.umontreal.ca>:
    Indeed, in practice virtually all compiler "optimizations" are valid
    only statistically: in most cases they either have no measurable effect
    or they improve some characteristic, but there are almost always corner >>cases where they make things worse.
    There are plenty of optimizations that always make things better, e.g., removing dead code, or reusing values in registers. But these days those
    are so obvious we sometimes forget about them.

    "Always" is a very strong statement. In practice, it's hard to be 100%
    sure that there can't be some weird corner case.
    Dead code removal looks safe enough, but it tends to change code layout,
    which in turn can affect cache conflicts, it can also enable new
    optimizations (which can in turn have detrimental effects), ...

    I'm not saying compilers should not do dead code removal, tho.
    Just that weird performance effects of optimizations are part of life,
    and every optimization comes with a degree of uncertainty over whether
    or not it'll be harmless.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 22 05:19:50 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows.

    Define "miscompile". According to which specification?

    [...]

    My impression is that the C compiler maintainers and some others have
    fallen in love with the idea of optimizing based on assuming that
    undefined behaviour does not happen; e.g., I have read the claim that
    this is the only thing that gives C an edge over other programming
    languages (or somesuch).

    Other languages handle this better. In Fortran, for example, signed
    integer overflow is just an error, so a compiler that would not
    optimize on the assumption of absence of integer overflow would
    be doing a poor jobs.

    So I think that while the C89 committee may
    not have thought of such things, recent C standardization committees
    probably have.

    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus:
    <https://blog.regehr.org/archives/1287>.

    People who want that kind of language know where to find it,
    they can always chose Ada.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 22 06:15:56 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    All such choices are the result of heuristics. Bubble sort has a
    very special memory access pattern.

    Ok, what would a memory pattern look like that this vectorization was designed for? The big slowdown comes from the fact that in the
    reordering case the vectorized version performs a wide store, and in
    the next iteration it performs a wide load that partially overlaps the
    store. I buy it that the analysis will have trouble seeing that (not
    that it's impossible in this case, but it requires effort).

    A straightforward patch would
    very likely pessimize a lot of existing code which profits from >>auto-vectorization.

    How can we test this claim?

    I do not believe that I have to explain the scientific method
    to you.

    As a first step, you would find an option that enables/disables
    what you don't like. -fstore-merging looks like a
    candidate, but there may be others. You can look at https://dl.acm.org/doi/10.1109/ASE56229.2023.00209 or https://link.springer.com/article/10.1007/s10515-024-00437-w
    if you want the full package.

    Then try this combination of options on other benchmarks
    as well as your pet one. I have my reservations about SPEC,
    but as you work at a university, you can get it for a discount.
    You can also use freely available benchmarks: Coremark, embench,
    Fortran Polyhedron - there are a lot.

    If there are benchmark which regress significantly with that
    set of options, you have your test case where it hurts.

    Like I wrote previously - you could have a bachelor or master
    student do this, it could be (part of a) nice thesis.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 22 05:25:54 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    -------------
    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    Does anyone know if Ada was successful about not leaving parts
    unspecified ??

    Does Ada have this goal?

    Java does have this goal, and AFAIK it is far better in this respect
    than C. There are some difficulties where hardware differences
    (handling of denormals on the 80387, thankfully phased out with the
    advent of SSE2) and lack of specification in hardware (weak memory
    models) are concerned.

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows. It gives a speedup in some SPEC
    program (How much? Don't know, but I am sure it exists); ok, it
    breaks this other SPEC program, let's special-case the optimization to
    avoid this breakage.

    This is a problem with SPEC not the compilers for various machines.

    How should it be a problem of SPEC if a compiler wilfully break
    non-SPEC programs, but spares SPEC programs?

    A broad spectrum benchmark should not contain code that relies on
    unspecified or undefined behaviors.

    A benchmark that is supposed to contain representative of real-world application programs written in C must exercise undefined behaviour in
    those C programs. If SPEC rejects benchmark submissions because of
    undefined behaviour, or edits the undefined behaviour out, the result
    is no longer representative. (The focus of SPEC CPU on long-running
    programs that use lots of memory also hurts.)

    My impression is that the C compiler maintainers and some others have
    fallen in love with the idea of optimizing based on assuming that
    undefined behaviour does not happen; e.g., I have read the claim that
    this is the only thing that gives C an edge over other programming
    languages (or somesuch). So I think that while the C89 committee may
    not have thought of such things, recent C standardization committees
    probably have.

    How much more optimizations can compilers deliver to the bottom line
    (not just benchmarks).

    Number of optimizations? There have been a number of cases where I
    have noticed that gcc misses a real optimization. The tsp example
    also shows that a lot of the optimizations that Jon Bentley performed
    in 1982 are not performed by compilers.

    Effect of new optimizations or "optimizations" based on assuming that
    undefined behaviour never happens on performance? They never say.
    That's the cool thing. They do not have numbers (they certainly never
    present any, certainly not for their own compilers), but are convinced
    that their "optimizations" do wonders for performance, and their
    fanboys are even more convinced.

    Why are they convinced? My guess is me that they see that their
    transformation actually is applied, and assume that it gives a
    speedup. The speedup is usually not measurable, because the
    transformation is not applied in hot code, but the assumption is that
    the combination of hundreds of transformations has a beneficial
    effect. Which might be the case if they hit more often than they
    miss, but they defend every questionable transformation (or worse, "optimization") with the argument that it helps, without presenting
    evidence, certainly not in any case where I have seen this claim, and
    I have seen it several times, most recently in this discussion <118odha$31okp$1@dont-email.me>.

    I have tested several of these claims, e.g., just yesterday <2026Sep21.171615@mips.complang.tuwien.ac.at> the claim that
    auto-vectorizing several narrow loads into a wide load is a
    optimization unless the "very special memory access pattern"
    happens. Of course Thomas Koenig did not specify what's so "very
    special" about the memory access pattern, which allows him to claim
    that every case where no speedup is shown, such as the one in <2026Sep21.171615@mips.complang.tuwien.ac.at>, is also "very special".
    OTOH, he has not shown a single instance where this autovectorization
    helps, i.e., a case that he does not consider to be "very special".
    My guess is that the majority of cases are "very special".

    For the "optimization" that broke SATD in 464.h264ref, the gcc people
    defanged it for the release of gcc-4.8 in a way that spared
    464.h264ref. When I talked to a gcc developer about it, he said
    something like "it did not speed up relevant programs". Which makes
    it interesting to know what "relevant programs" are; my guess is that
    it is a corpus of programs consisting of well-known benchmarks plus a collection of programs from paying customers who would not like to see
    their programs to regress in performance; it seems to me that they now
    also have a collection of programs where functionality is tested, but apparently that did not include the SPEC CPU benchmarks before gcc-4.8
    (or the functionality testing was not performed before creating the
    prerelease that broke 464.h264ref).

    Last week we saw an episode where the compilers
    only got 1.x% speedup over years (sub-decade). Is there ever going to
    be a time to stop ??

    When the money stops flowing.

    As mentioned, there are a lot of transformations possible. Turning
    them into reliable automatic optimizations is a lot of work; I have my
    doubts that it is worth it.

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.

    Some of the machines have egregious behaviors in tiny little corners
    of the architecture(s) and implementation(s). Not being able to run
    the same binary with and without integer overflow detection as an example.

    None of the mainstream architectures have hardware support for this
    feature, and I have never missed it.

    And even when they can, they cannot perform Ada ADD Byte and take a signed >overflow trap on 8-bit results easily.

    It's not that hard: Sign-extend the byte, compare the result to the
    original, and jump to the trap if it they are different. Three inline instructions on AMD64 and RISC-V. And checking the result is far
    better than checking every operation, because the result may be in
    range even if some intermediate result is not.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Tue Sep 22 08:29:09 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows.

    Define "miscompile".

    If a compiler produces an unintended behaviour for an existing, tested
    program that works as intended with an earlier version of the compiler.

    @InProceedings{ertl17kps,
    author = {M. Anton Ertl},
    title = {The Intended Meaning of \emph{Undefined Behaviour}
    in {C} Programs},
    crossref = {kps17},
    pages = {20--28},
    url = {https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf},
    url2 = {http://www.kps2017.uni-jena.de/proceedings/kps2017_submission_5.pdf},
    abstract = {All significant C programs contain undefined
    behaviour. There are conflicting positions about how
    to deal with that: One position is that all these
    programs are broken and may be compiled to arbitrary
    code. Another position is that tested and working
    programs should continue to work as intended by the
    programmer with future versions of the same C
    compiler. In that context the advocates of the first
    position often claim that they do not know the
    intended meaning of a program with undefined
    behaviour. This paper explores this topic in greater
    depth. The goal is to preserve the behaviour of
    existing, tested programs. It is achieved by letting
    the compiler define a consistent mapping of C
    operations to machine code; and the compiler then
    has to stick to this behaviour during optimizations
    and in future releases.}
    }

    @Proceedings{kps17,
    title = {19. Kolloquium Programmiersprachen und Grundlagen der Programmierung (KPS'17)},
    booktitle = {19. Kolloquium Programmiersprachen und Grundlagen der Programmierung (KPS'17)},
    year = {2017},
    key = {KPS '17},
    editor = {Wolfram Amme and Thomas Heinze},
    url = {http://www.kps2017.uni-jena.de/kps2017_proceedings.html},
    url-pdf = {http://www.kps2017.uni-jena.de/proceedings/kps2017.pdf}
    }

    According to which specification?

    If you want that, go for the specification of a conforming program in
    the C standard. Compiling a conforming program compiled with a
    different compiler or on a different architecture to different
    behaviour is ok with me, compiling it on the same architecture with a
    new version of the same compiler into different behaviour is
    miscompilation.

    In Fortran, for example, signed
    integer overflow is just an error, so a compiler that would not
    optimize on the assumption of absence of integer overflow would
    be doing a poor jobs.

    What does "just an error" mean? Does it trap?

    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus: >><https://blog.regehr.org/archives/1287>.

    People who want that kind of language know where to find it,

    I think these efforts are the result of C compiler maintainers and
    their fanboys trying to deflect the blame from the real culprits (the
    C compiler maintainers) to the C standardization committee: If the specification is to blame, people try to work on a better
    specification.

    These attempts were not successful with me, howewer. I think that the
    compiler maintainers are to blame. I don't think that the C standard
    is to blame, just as the rules are not to blame for malicious
    compliance. <https://en.wikipedia.org/wiki/Malicious_compliance>
    describes the behaviour of gcc and clang in the last decades:
    "strictly following orders, laws, and rules to the letter, ignoring expectations that go without saying and the spirit of the
    requirement." At least that's the way they behave when some breakage
    actually comes up in public.

    But given that you mention a different language, I also have the
    impression that the unhappyness about this behaviour by the compiler maintainers is among C programmers, not (to my knowledge) among C++ programmers.

    A number of projects that use C use compiler options to make the
    language more defined, among them the Linux, PostgresSQL, and Gforth.
    In Gforth we try, and if present, use the options

    -fno-gcse
    -fcaller-saves
    -fno-defer-pop
    -fno-inline
    -fwrapv
    -fchar-unsigned
    -fno-strict-aliasing
    -fno-cse-follow-jumps
    -fno-reorder-blocks
    -fno-reorder-blocks-and-partition
    -fno-toplevel-reorder
    -fno-trigraphs
    -falign-labels=1
    -falign-loops=1
    -falign-jumps=1
    -fno-delete-null-pointer-checks
    -fcf-protection=none
    -fno-tree-slp-vectorize

    some of which tighten the language definition. I expect that other
    projects have similar lists. With new compiler versions we
    occasionally have to scan the options to see if there is something new
    that may be "optimized" by the new version and where there exists an
    option for tightening the language to prevent that "optimization" from miscompiling the program.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Niklas Holsti@niklas.holsti@tidorum.invalid to comp.arch on Tue Sep 22 14:48:28 2026
    From Newsgroup: comp.arch

    On 2026-09-21 23:06, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    -------------
    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    Does anyone know if Ada was successful about not leaving parts
    unspecified ??

    Not entirely, because there are still some cases where the Ada standard
    says that "erroneous execution" follows, which means essentially
    unpredictable effects, and also cases of "bounded error", which means
    that one of a few, specific and defined things can happen, but it is unpredictable or implementation-dependent which of those things happen.

    Perhaps the most easily triggered erroneous execution is accessing an
    object via a pointer, but that object has been discarded (a dangling reference). Most erroneous-execution cases come from similar "behind the scenes" changes of program state that make some other parts of that
    state no longer valid or correct.

    A typical case of bounded error is when a subprogram call has two formal parameters, but the call supplies the same object as an argument to both
    those parameters, so that the subprogram can access the same object
    through two paths. If the parameter-passing method is left to the
    compiler's choice (by copy or by reference), it is a bounded error to
    assign the object a value via one path, and then read the object via the
    other path, where the possible consequences are that Program_Error is
    raised, or the newly assigned value is read, or some old value of the
    object is read.

    The Ada language manintainers certainly consider it an important goal to minimize the cases of both types of unspecified behaviour. And to
    identify those cases in the Ada reference manual, of course. Also, if
    some behaviour is defined in the reference manual as implementation
    dependent, the standard requires an implementor to document what the implementation does in such cases.

    The SPARK subset/variant of Ada has the stronger goal of either
    preventing such problems by design, or proving that they cannot happen
    in execution.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Tue Sep 22 14:32:47 2026
    From Newsgroup: comp.arch

    On 21/09/2026 19:17, Anton Ertl wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    [...]
    The traditional "-O<N>" flags are a way to state how much you're willing
    to get worse performance in exchange for the opportunity to maybe get
    better performance.

    The gcc manual states:

    With '-O', the compiler tries to reduce code size and execution
    time, without performing any optimizations that take a great deal
    of compilation time.
    [...]
    [about -O2] As
    compared to '-O', this option increases both compilation time and
    the performance of the generated code.
    [...]
    [about -O3] Optimize yet more.

    In earlier gcc versions, it made statements along the lines that -O3
    may be a mixed bag, but that's gone.

    It must have been removed a /long/ time ago, because it is not in manual
    pages that I checked (the oldest convenient version is 2.95.3). But it
    is certainly the case that -O3 is a mixed bag - in particular, more
    aggressive unrolling and inlining can lead to larger code that can
    reduce the effect of caches, branch prediction, and that kind of thing.

    In microcontroller work, -Os is often used to put a stronger emphasis on
    size optimisations. I have noticed there are sometimes "blips" - cases
    where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
    faster code than "-O2". And there can sometimes be cases where the
    speed/size tradeoffs are unreasonable - significantly slower code for
    very minor size improvements, or vice versa.

    It is clearly not the case that all optimisations improve all code, or
    that increasing "n" in "-On" always gives faster results. Each
    optimisation flag in a compiler enables one or more transformation
    passes that might improve some code but might also have detrimental
    effects in some cases. Lower "-On" numbers will include the passes with
    high statistical rates of improvement and low risk of worsening code,
    while the passes enabled with higher flags will have worse ratios and
    longer compiler times. Flags with significant risks of making code a
    lot worse are usually not enabled by any "-On" flag, and require manual choice. (But no one will claim that gcc, clang, or any other compiler
    is perfect here.)

    I think one thing that is sometimes forgotten in all this, and could
    usefully be mentioned on the gcc manual page for optimisation, is that
    cpu architecture flags can make a significant difference. The step from
    "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
    from "-O2" to "-O3" in performance, with much lower risks of slowdowns.


    But if optimizer writers strove to "improve the convenience to write
    efficient code" instead of just improving benchmark results, maybe
    such paradoxical effects can be avoided.

    In general it's hard to completely avoid paradoxical effects.

    I referred to a specific paradoxical effect, not completely avoiding
    all of them: The effect of programmers who, instead of working on
    making the code faster through source-level changes, have to invest
    time into avoiding getting it miscompiled.


    By "miscompiled", do you mean incorrect object code, or object code that
    did not have the performance the programmer expected or hoped for?

    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    I have never spoken to the C standards committee or writers, either of
    current standard versions or the original C89 standard, or writers of pre-standard C specifications. So I cannot in any way claim to know
    their thoughts or motivations. I also have not seen any documentation
    that suggests what they might have thought about "optimisation on the assumption that undefined behaviour does not occur" - in either
    direction. (It has been mentioned that, for example, a two's complement implementation could use wrapping signed arithmetic - but I have never
    seen a suggestion that this behaviour should be encouraged or expected
    just because a machine uses two's complement.)

    I do, however, believe that the folks being the C language design and specifications through all its changes are accomplished and experienced computer scientists. They will have been aware of the "garbage in,
    garbage out" principle, and that there is no reason to expect any
    particular result or effect when you apply a function or operator to
    something outside its defined semantics. You do not ask "what happens
    if my signed integer arithmetic overflows?" or "what happens when I
    access an array out of bounds?" - rather, it is your responsibility as a
    C programmer to make sure that never happens. The prime reason C does
    not define behaviour here is not that different hardware or compilers
    handle things differently, but that there is no sensible definition that
    could be made.

    Remember, C has a perfectly good way to say "this is determined by the hardware or the implementation" - it is "implementation-defined
    behaviour". It has a perfectly good way to say that "this operation
    could result in any value" - it is "unspecified behaviour" or
    "unspecified value".

    When the C standards writers say something is "undefined behaviour",
    rather than "implementation defined" or "unspecified", it is my belief
    that they did so intentionally and knowingly. I have at times been
    accused of arrogance, usually quite fairly, but I am not arrogant enough
    to suppose that I know when the C standards committee made mistakes here
    or intended to write something differently.

    So I cannot accept an argument that C's undefined behaviours are either
    due to hardware differences, or laziness. It simply does not fit with
    the level of expertise and effort that have gone into the language and
    its standards.

    And note that in C23 the macro "unreachable()" was added with the sole semantics being "If a macro invocation unreachable() is reached during execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation". The justification is for better
    diagnostics and optimisations. (Some people are of the opinion that
    later C standards deviate from the philosophy of early C - that may or
    may not be a fair opinion, and opinions will also differ on whether that
    is a good thing or a bad thing.)


    I don't know whether or not the "founding fathers" of C intended or
    expected compilers to optimise on the assumption that UB did not occur.
    But I am confident that they considered a program to be broken if
    execution reached a point where the behaviour was not defined in an implementation (something may be UB in the C standards yet defined by an implementation). I am confident that if they intended that adding two
    ints must always return a valid int even if there is an overflow, they
    would have said its result was implementation-defined or an unspecified
    value. The behaviour is intentionally left wide open, without restriction.

    It is, of course, possible that they did not foresee quite how this
    would pan out in modern compilers. In particular, they may not have
    predicted "time-travel" optimisations. However, authors of more modern
    C standards - say, C11 onwards - know about them and have could have explicitly outlawed them if they were considered to be invalid.


    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows. It gives a speedup in some SPEC
    program (How much? Don't know, but I am sure it exists); ok, it
    breaks this other SPEC program, let's special-case the optimization to
    avoid this breakage.


    Regardless of what was and was not intended by the original designers of
    the C language, compiler developers must take practical considerations
    if their tool is to be useful to people. A compiler that implements multiplication by repeated addition could be fully conforming, but users
    would quickly look for something that gave more efficient results. For heavily used compilers, there is a strong (but not overpowering) push
    towards backwards compatibility. This is why they have heavy regression tests, and pre-release versions are tested with large samples of
    important existing code. If this gives unexpected results, these must
    be dealt with - was it a bug in the new compiler (in which case the fix
    is obvious), or was it a bug in the old source code? Those cases are
    more complicated - sometimes the old code must be fixed, but sometimes
    the incorrect source code is too common, idomatic or important and the compiler must, in effect, support additional semantics to retain the old accidental semantics.

    My impression is that the C compiler maintainers and some others have
    fallen in love with the idea of optimizing based on assuming that
    undefined behaviour does not happen; e.g., I have read the claim that
    this is the only thing that gives C an edge over other programming
    languages (or somesuch). So I think that while the C89 committee may
    not have thought of such things, recent C standardization committees
    probably have.

    I am sure you are right that more recent standards committees consider optimisations here more than older ones. "restrict" was added to C99
    for optimisation purposes (to catch up with Fortran) - it is, in effect, specified by saying that the programmer promises to follow certain
    additional rules in the way they access data or the result is undefined behaviour. It is very much intended that compilers optimise "restrict"
    by assuming that at least a particular type of UB does not occur. In
    C23, there are several features that are entirely for optimisation - the aforementioned "unreachable()" macro, as well as the "reproducible" and "unsequenced" attributes.


    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus: <https://blog.regehr.org/archives/1287>.

    There are two key problems with UB that stand in the way of this kind of initiative. One is that many compilers provide consistent (and
    sometimes even documented) behaviour for some things that are UB in the
    C standards. Code relying on these behaviours is valid and safe, but non-portable. But different compilers (or the same compiler but
    different targets) could easily have different semantics, making it very difficult to agree on any one choice. Secondly, many programmers write
    code on the assumption that certain UB has certain consistent and
    reliable behaviour - even though it is not documented anywhere. It is
    code like that which could benefit from a "friendly C" (a poor choice of
    name, IMHO, but that's entirely subjective) variant. But it is also
    such code that makes a "friendly C" variant hard to define - such code
    is hard to identify, and the expected behaviour can be even harder to
    find, specify, and consistently describe.


    But fortunately a completely specified C is not necessary, a
    willingness to preserve the behaviour of existing working programs
    compiled with an earlier version of the same compiler is. Read more
    about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.


    I think it is entirely reasonable to keep old compiler versions around
    and use them for old code. Even if the code is clinically free of any
    UB, there can be implementation-defined behaviour, differences in code
    size and timing, support for features or targets that are removed in
    later toolchain versions, etc.

    I think it is entirely unreasonable to try to specify that new compilers should have defined specifications to implement the "semantics" that
    older compilers used for particular types of undefined behaviour. These semantics will, for the most part, be poorly defined and can often be inconsistent - many programs with UB work by luck, not design.

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.


    Linux commits to preserving the /defined/ behaviour of user-space APIs.
    If an API is defined to take a valid file descriptor as a parameter,
    then programs that use that API with a valid file descriptor will always
    work the same way. What happens if you give it an invalid parameter is another matter. Maybe your program will receive an error return, or
    maybe it will be given a kill signal. Maybe some extra bits in the
    parameter will be ignored in earlier versions of Linux and later
    versions will use those bits for extra flags. Or maybe the API will
    define exactly what you get for any possible parameter value. But if a particular invalid value of the parameter lets you gain access to
    root-owned files, due to a bug in the kernel, you can be very sure that
    this user-space behaviour will /not/ be preserved.

    It's a very good thing to avoid changing defined and documented
    behaviour unless there is extremely good reason to do so. It's a very
    bad thing to commit to avoid changing undefined and undocumented behaviour.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 22 19:57:03 2026
    From Newsgroup: comp.arch

    Thomas Koenig wrote:
    Paul Clayton <paaronclayton@gmail.com> schrieb:

    I suspect algorithmic transformations are discouraged in
    compilers under the assumption that the software developer used
    appropriate algorithms as well as the compute cost to evaluate
    options.

    Two points: A compiler is restricted by the transformations
    it can do by the language specification. For example, loop
    interchanges with floating point calculation can give different
    results. Unless directed otherwise, the compiler has to assume
    that this is what the user wants and needs.

    An exception is something like Fortran's FORALL constuct, where
    the loop ordering is not specified and the compiler can chose.
    Intrinsic functions like MATMUL also have no restriction on what
    they can do internally; they can (and often do) call a highly-
    optimized BLAS routine.

    Usage and hardware information might theoretically be
    made available, but I get the impression that most profiling
    for compiler use is about path frequency.

    Compilers tend to focus on lower-level stuff. For example,
    there is a huge file of simplifying transformations for gcc at https://gcc.gnu.org/git/?p=gcc.git;a=blob_plain;f=gcc/match.pd

    This now also contains reverse-engineered bithacks. For example,
    the popcnt() method from Hacker's Delight are now recognized and
    translated into an internal function, which is then expanded
    into either an inlined function (much like the original) or
    a machine instruction, if the target has one.

    clang recognizes both the bit per iteration and the n &= n-1 Hacker's
    method, bot turn into POPCNT on x86.

    Much more impressive (to me) was to see that same compiler (working as
    the Rust backend) also recognized at least two different naive
    implementations of a non-square bitmap transpose, converting it into a
    SIMD blob with no need for me to break out the SSE/AVX intrinsics!

    High-level transformation would transform Anton's beloved
    bubblesort benchmark into insertion sort, at least.

    But if you're willing to take the risk of AI, you can of course
    tell it "Find all O(N^2) algorithms in my code and replace them
    with O(N log N) where possible". It will happily do something,
    and the resulting code might even be correct after a few
    iterations.

    "might" is important here...

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Tue Sep 22 20:01:46 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    For people who don't like bubble sort, or for inclusion in a set
    of benchmarks, I suggest the following recently discovered
    sorting algorithm (written in C-ish pseudocode):

    void
    baffle_sort( unsigned n, int *elements ){
    for( unsigned i = 0; i < n; i++ ){
    for( unsigned j = 0; j < n; j++ ){
    if( elements[i] < elements[j] ){
    /* exchange elements[i] and elements[j] */
    }
    }
    }
    }

    (Disclaimer: not original.)

    Did you watch one of my favorite Youtube channels?
    :-)

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Tue Sep 22 14:13:06 2026
    From Newsgroup: comp.arch

    On 2026-Sep-22 01:25, Anton Ertl wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    How much more optimizations can compilers deliver to the bottom line
    (not just benchmarks).

    Number of optimizations? There have been a number of cases where I
    have noticed that gcc misses a real optimization. The tsp example
    also shows that a lot of the optimizations that Jon Bentley performed
    in 1982 are not performed by compilers.

    Effect of new optimizations or "optimizations" based on assuming that undefined behaviour never happens on performance? They never say.
    That's the cool thing. They do not have numbers (they certainly never present any, certainly not for their own compilers), but are convinced
    that their "optimizations" do wonders for performance, and their
    fanboys are even more convinced.

    It seems someone measured UB optimization performance impact in LLVM:

    Exploiting Undefined Behavior in C C++ Programs for Optimization
    A Study on the Performance Impact, 2025 https://dl.acm.org/doi/abs/10.1145/3729260 https://dl.acm.org/doi/pdf/10.1145/3729260

    "Using LLVM, a compiler known for its extensive use of UB for optimizations,
    we demonstrate that, for the benchmarks and UB categories that we evaluated, the end-to-end performance gains are minimal."


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Clayton@paaronclayton@gmail.com to comp.arch on Mon Sep 21 22:10:53 2026
    From Newsgroup: comp.arch

    On 9/17/26 5:25 PM, Anton Ertl wrote:
    [snip]
    Matrix300 was eliminated from SPEC because all the computer companies
    started to use automatic cache blocking. IIRC this was for the step
    from SPEC89 to SPEC92.

    I think I read that cache sizes were also a factor.

    Later Sun managed to optimize IIRC the ear
    benchmark to perform array-of-structures to structure-of-arrays transformation, achieved a speedup by a factor of IIRC 2, and a
    significant increase in the aggregate SPEC score.

    I seem to recall that the (sub)benchmark was art. I got the
    impression (possibly false) that the optimization was fragile
    and strongly targeted to that specific benchmark. This is not
    as extreme as GPU drivers changing optimizations for specific
    game binaries. (Some such optimizations might be valid. If a
    game uses a more general interface rCo for broader compatibility,
    faster development, or easier extension rCo but a more specialized
    (and faster) interface meets all requirements the GPU driver/
    compiler could reasonably make that substitution. If the more
    specialized interface is not published (e.g., to provide greater
    future flexibility), the game developers could not use it even
    if they were willing to sacrifice generality but having the
    compiler/driver perform this optimization may be reasonable.
    Yet using the program name to perform optimizations seems
    sketchy.)

    Side comment: One of the minor benefits proposed for Transmeta's
    Code Morphing Software (converting x86 binaries to internal
    VLIW) was the ability to work around timing limitations ("bugs")
    in hardware. E.g., if a particular code sequence would fail
    timing at peak frequency, binary translation could avoid that
    case rather than force the hardware to be underclocked to
    generally avoid that issue.

    (I do not know how much of this would apply to a design with
    dynamic scheduling. I also suspect that the advantage of
    "fixing" hardware bugs in the binary translator was not that
    great.)

    Trusted binary translation and the fixed addressing of the
    cached translations also allowed the removal of the instruction
    TLB. If the translations had been persistently cached (like
    AS/400 and the Mill proposal) and dynamic paging used, it
    _might_ have been possible for relocation to be supported for
    code and maintained the ability to avoid an instruction TLB.

    (Azul Systems use huge pages for translation to facilitate
    debugging even though, in theory, running JAVA applications
    could be safe; compacting garbage collection might also have
    reduced issues associated with huge pages. Such software-
    hardware codesign does seem to have optimization potential.)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Tue Sep 22 20:52:30 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:
    Tim Rentsch wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    For people who don't like bubble sort, or for inclusion in a set
    of benchmarks, I suggest the following recently discovered
    sorting algorithm (written in C-ish pseudocode):

    void
    baffle_sort( unsigned n, int *elements ){
    for( unsigned i = 0; i < n; i++ ){
    for( unsigned j = 0; j < n; j++ ){
    if( elements[i] < elements[j] ){
    /* exchange elements[i] and elements[j] */
    }
    }
    }
    }

    (Disclaimer: not original.)

    Did you watch one of my favorite Youtube channels?
    :-)

    As the google AI sez:


    Complexity Matrix
    Metric Complexity Description
    Worst-Case Time O(n#) Always runs exactly n + n iterations.
    Best-Case Time O(n#) Unlike optimized Bubble Sort, it has no early exit and runs fully even if pre-sorted.
    Space Complexity O(1) Sorts in-place with a single temporary variable.

    Are you analyzing this specific function for a computer science
    assignment / code review, or would you like me to help you rewrite
    it into a working standard algorithm like Bubble Sort or Quick Sort?
    If a student can't rewrite it themselves, they shouldn't be programming. AI>That is a fair and candid perspective. Rewriting or fixing this specific
    function is a foundational exercise in control flow, loops, and basic logic. If someone
    cannot spot why a double-nested loop from 0 to n breaks standard sorting
    conventions - or how to implement a basic three-line variable swap - they
    will struggle immensely with more complex programming concepts.
    It highlights the difference between memorizing syntax and truly understanding
    code execution.
    Are you grading student submissions that included this specific code, or are you
    analyzing a textbook/exam problem designed to trip them up?
    I saw the code snippet in comp.arch
    Ah, comp.arch - that explains it perfectly. USENET groups like comp.arch and comp.lang.c
    are notorious for regulars posting bizarre, intentionally "baffling" code fragments to
    spark academic debates, analyze compiler optimization side-effects, or illustrate
    architectural edge cases.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Tue Sep 22 15:21:36 2026
    From Newsgroup: comp.arch

    On 9/22/2026 5:32 AM, David Brown wrote:
    On 21/09/2026 19:17, Anton Ertl wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> writes:
    Anton Ertl [2026-09-19 06:25:13] wrote:
    [...]
    The traditional "-O<N>" flags are a way to state how much you're willing >>> to get worse performance in exchange for the opportunity to maybe get
    better performance.

    The gcc manual states:

    -a-a-a-a-a With '-O', the compiler tries to reduce code size and execution >> -a-a-a-a-a time, without performing any optimizations that take a great deal >> -a-a-a-a-a of compilation time.
    [...]
    -a-a-a-a-a [about -O2] As
    -a-a-a-a-a compared to '-O', this option increases both compilation time and >> -a-a-a-a-a the performance of the generated code.
    [...]
    -a-a-a-a-a [about -O3] Optimize yet more.

    In earlier gcc versions, it made statements along the lines that -O3
    may be a mixed bag, but that's gone.

    It must have been removed a /long/ time ago, because it is not in manual pages that I checked (the oldest convenient version is 2.95.3).-a But it
    is certainly the case that -O3 is a mixed bag - in particular, more aggressive unrolling and inlining can lead to larger code that can
    reduce the effect of caches, branch prediction, and that kind of thing.

    In microcontroller work, -Os is often used to put a stronger emphasis on size optimisations.-a I have noticed there are sometimes "blips" - cases where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
    faster code than "-O2".-a And there can sometimes be cases where the speed/size tradeoffs are unreasonable - significantly slower code for
    very minor size improvements, or vice versa.

    It is clearly not the case that all optimisations improve all code, or
    that increasing "n" in "-On" always gives faster results.-a Each optimisation flag in a compiler enables one or more transformation
    passes that might improve some code but might also have detrimental
    effects in some cases.-a Lower "-On" numbers will include the passes with high statistical rates of improvement and low risk of worsening code,
    while the passes enabled with higher flags will have worse ratios and
    longer compiler times.-a Flags with significant risks of making code a
    lot worse are usually not enabled by any "-On" flag, and require manual choice.-a (But no one will claim that gcc, clang, or any other compiler
    is perfect here.)

    I think one thing that is sometimes forgotten in all this, and could usefully be mentioned on the gcc manual page for optimisation, is that
    cpu architecture flags can make a significant difference.-a The step from "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
    from "-O2" to "-O3" in performance, with much lower risks of slowdowns.


    But if optimizer writers strove to "improve the convenience to write
    efficient code" instead of just improving benchmark results, maybe
    such paradoxical effects can be avoided.

    In general it's hard to completely avoid paradoxical effects.

    I referred to a specific paradoxical effect, not completely avoiding
    all of them: The effect of programmers who, instead of working on
    making the code faster through source-level changes, have to invest
    time into avoiding getting it miscompiled.


    By "miscompiled", do you mean incorrect object code, or object code that
    did not have the performance the programmer expected or hoped for?

    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree.-a The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    I have never spoken to the C standards committee or writers, either of current standard versions or the original C89 standard, or writers of pre-standard C specifications.-a So I cannot in any way claim to know
    their thoughts or motivations.-a I also have not seen any documentation
    that suggests what they might have thought about "optimisation on the assumption that undefined behaviour does not occur" - in either
    direction.-a (It has been mentioned that, for example, a two's complement implementation could use wrapping signed arithmetic - but I have never
    seen a suggestion that this behaviour should be encouraged or expected
    just because a machine uses two's complement.)

    I do, however, believe that the folks being the C language design and specifications through all its changes are accomplished and experienced computer scientists.-a They will have been aware of the "garbage in,
    garbage out" principle, and that there is no reason to expect any
    particular result or effect when you apply a function or operator to something outside its defined semantics.-a You do not ask "what happens
    if my signed integer arithmetic overflows?" or "what happens when I
    access an array out of bounds?" - rather, it is your responsibility as a
    C programmer to make sure that never happens.-a The prime reason C does
    not define behaviour here is not that different hardware or compilers
    handle things differently, but that there is no sensible definition that could be made.

    Remember, C has a perfectly good way to say "this is determined by the hardware or the implementation" - it is "implementation-defined behaviour".-a It has a perfectly good way to say that "this operation
    could result in any value" - it is "unspecified behaviour" or
    "unspecified value".

    When the C standards writers say something is "undefined behaviour",
    rather than "implementation defined" or "unspecified", it is my belief
    that they did so intentionally and knowingly.-a I have at times been
    accused of arrogance, usually quite fairly, but I am not arrogant enough
    to suppose that I know when the C standards committee made mistakes here
    or intended to write something differently.

    So I cannot accept an argument that C's undefined behaviours are either
    due to hardware differences, or laziness.-a It simply does not fit with
    the level of expertise and effort that have gone into the language and
    its standards.

    And note that in C23 the macro "unreachable()" was added with the sole semantics being "If a macro invocation unreachable() is reached during execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"? Yes, it would make the
    compiler implementers do some work to document whatever they decided to
    do, but are there any other reasons?
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 22 16:50:12 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Thomas Koenig <tkoenig@netcologne.de> writes:

    Bubblesort is the worst of non-joke sorting algorithms, see
    the quote from Numerical Recipes...

    For people who don't like bubble sort, or for inclusion in a set
    of benchmarks, I suggest the following recently discovered
    sorting algorithm (written in C-ish pseudocode):

    void
    baffle_sort( unsigned n, int *elements ){
    for( unsigned i = 0; i < n; i++ ){
    for( unsigned j = 0; j < n; j++ ){
    if( elements[i] < elements[j] ){
    /* exchange elements[i] and elements[j] */
    }
    }
    }
    }

    (Disclaimer: not original.)

    Did you watch one of my favorite Youtube channels?

    Certainly there is a chance that I did. I found out about
    this sorting algorithm from a YouTube video.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 23 05:23:26 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/22/2026 5:32 AM, David Brown wrote:
    And note that in C23 the macro "unreachable()" was added with the sole
    semantics being "If a macro invocation unreachable() is reached during
    execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?

    The idea here is obviously to use this macro only when its invocation
    really is unreachable, and in that case it has no effect on the
    behaviour.

    How should an implementation document what happens when the invocation
    of this macro actually is reachable? The effect depends on the
    surrounding code and the transformations in the compiler.

    And what would a programmer do with this documentation?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 23 05:40:21 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:

    I don't think that they expected that compiler writers on a
    twos-complement machine that does not trap on signed overflow would
    say: Hey, signed overflow is undefined behaviour, let's assume it
    never happens, and silently miscompile some (not all) programs that
    actually perform signed overflows.

    Define "miscompile".

    If a compiler produces an unintended behaviour for an existing, tested program that works as intended with an earlier version of the compiler.

    Two points: gcc 16 is a different compiler from gcc 15 and all
    previous versions. For gcc 16.2 vs. gcc 16.1, I would agree
    (partially), see below.

    There is a reason why such things as
    https://gcc.gnu.org/gcc-16/porting_to.html are necessary.

    Assume something worked in gcc 14 and was broken in gcc15
    (i.e. did not work according to the language specification).
    Also assume that two programs were tested, one according to
    the correct gcc14 behavior and one according to the incorrect
    gcc15 behavior. Your definition would be impossible to fulfill
    in that case.

    According to which specification?

    If you want that, go for the specification of a conforming program in
    the C standard.

    I disagree. AFAIK, it has been argued that

    PRINT *,42
    END

    is a conforming C program because the gfortran command is a
    conforming C implementation (the driver also compiles C).

    Compiling a conforming program compiled with a
    different compiler or on a different architecture to different
    behaviour is ok with me, compiling it on the same architecture with a
    new version of the same compiler into different behaviour is
    miscompilation.

    gcc 16 is a different compiler from gcc15, but even there
    you would like to make regression fixes between gcc 16.2 and
    gcc 16.1 impossible.


    In Fortran, for example, signed
    integer overflow is just an error, so a compiler that would not
    optimize on the assumption of absence of integer overflow would
    be doing a poor jobs.

    What does "just an error" mean? Does it trap?

    No. It can issue a diagnostic at compile-time, issue a diagnostic
    at run-time with or with aborting the program, continue with
    unpredictable results (including, according to an old quip. starting
    World War III if the right operational hardware has been installed).
    ar anything else you might dream up. (But the "starting
    World War III part is now left to AIs, in the absence of
    any specification).

    The standard makes no such guarantees. The programmer must
    ("shall") not to that, that's all - it is a two-way contract.
    The compiler guarantees that violations of constraints are reported
    (modulo bugs), the programmer guarantees that he does not violate
    "shall" constraints.


    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus: >>><https://blog.regehr.org/archives/1287>.

    People who want that kind of language know where to find it,

    I think these efforts are the result of C compiler maintainers and
    their fanboys

    You're trying to make friends, I know :-) But you might think
    again - if you actually want to achieve something, that attitude
    is not going to help a lot.

    But if you are serious about this, and your opinion reflects
    that of a sufficient number of other people, you have a choice.

    You can fork either gcc or LLVM and yank out all the optimizations
    that you consider undesirable. The cost of that is almost zero.

    If you get enough capable people on board to work on your vision,
    and enough people use it, you will succeed. If your position
    is a very lonely one, as I suspect, you will fail. But
    it is a hypothesis that can be tested.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 23 05:49:22 2026
    From Newsgroup: comp.arch

    David Brown <david.brown@hesbynett.no> schrieb:

    In microcontroller work, -Os is often used to put a stronger emphasis on size optimisations. I have noticed there are sometimes "blips" - cases where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
    faster code than "-O2". And there can sometimes be cases where the speed/size tradeoffs are unreasonable - significantly slower code for
    very minor size improvements, or vice versa.

    Just to muddle the waters a little more: I once submitted a PR
    where -O3 led to smaller code than -Os. The reason? -O3 ran a
    pass which led to significant simplification and subsequent dead
    code elimination.

    I think one thing that is sometimes forgotten in all this, and could usefully be mentioned on the gcc manual page for optimisation, is that
    cpu architecture flags can make a significant difference. The step from "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
    from "-O2" to "-O3" in performance, with much lower risks of slowdowns.

    Additionally, -mtune=native can also help.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 23 10:38:05 2026
    From Newsgroup: comp.arch

    On 23/09/2026 00:21, Stephen Fuld wrote:
    On 9/22/2026 5:32 AM, David Brown wrote:

    And note that in C23 the macro "unreachable()" was added with the sole
    semantics being "If a macro invocation unreachable() is reached during
    execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?-a Yes, it would make the
    compiler implementers do some work to document whatever they decided to
    do, but are there any other reasons?


    Good question.


    There is a sort of partial ordering (or lattice order) of behaviour specifications. "undefined behaviour" is at the bottom - it implements nothing, and anything else can implement it. "implementation-defined"
    is stronger. If the standards say "UB", then an implementation can
    define behaviour more concretely. If the standards say "IB", then the implementation cannot turn it into "UB". So "UB" in the standards
    clearly gives the most flexibility to the implementation.

    The big disadvantage of making something like this "IB" is not that the compiler implementers need to document their handling, but that they
    need to have a specification for it and stick to that - users must be
    able to rely on IB being consistent. Different types of handling for
    hitting "unreachable()" can include compiler optimisation on the
    assumption that it is never reached, generating a target "trap"
    instruction, generating a "breakpoint" when debugging, printing out an
    error message and terminating, calling a user-defined function that
    sends an angry email to the developer, and tracking all possible code
    flows at build time and giving a compiler warning if it can, in fact, be reached. Some users will prefer one method, others will prefer a
    different one, and many will change depending on what they are doing. Crucially, compilers may change what they support over time. A compiler cannot document "IB" as "depending on flags, this might be treated as
    UB, or give a trap, or perhaps do something else in the future". But it
    /can/ do exactly that if the standards say it is "UB".

    It is also important to note that gcc, clang, icc, and many other
    compilers have had something like this for decades - like __builtin_unreachable() - which are implemented exactly as "undefined behaviour" in these tools.


    This is, as far as I can tell, the proposal paper that led to
    "unreachable()" being added to the C standards:

    <https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2826.pdf>

    """
    We propose the feature unreachable to specify branches in the control
    N4eow of a program that will never be reached. The aim is to provide means
    for the user to express guarantees about the eN4Cective control N4eow that
    will be executed by a program. Compilers may then apply aggressive optimizations that otherwise would not be possibly or that would rely on
    the detection of undeN4Uned behavior for certain input combinations.
    """

    The C++ version is :

    <https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p0627r6.pdf>

    The discussion about the definition is :

    """
    What is the best way to define the attribute's effect? The author feels
    that the best way is to make the behavior of std::unreachable() be
    undefined. There are several reasons:

    * std::unreachable() causing undefined behavior means that the Standard
    would not prescribe any particular action, leaving open many possible implementation actions.

    * Some compilers already associate being unreachable to undefined
    behavior. ClangrCOs documentation states that __builtin_unreachable() rCLhas completely undefined behaviorrCY.

    * Optimizing under the assumption that a statement is unreachable, and
    thus having unpredictable behavior if the statement is in fact
    reachable, falls naturally under "undefined behavior".

    * An alternative would be to issue a trap if std::unreachable() is
    executed. This could be used in "debug builds", for example. Such a trap
    falls under "undefined behavior".

    * Being undefined behavior implies what happens if a constexpr function
    calls std::unreachable(): it's not a constant-expression, by (N4713) [expr.const]/2.6

    In the authorrCOs opinion, it is not undefined behavior itself that is the problem, but rather unexpected undefined behavior. C++ will always have undefined behavior; this is unavoidable. Situations in which undefined behavior occurs unexpectedly are problematic for programmers, because
    they often lead to subtle bugs, or unexpected compiler optimizations.
    However, std::unreachable() is expected undefined behavior: a programmer
    using it knows that undefined behavior will occur there and is making
    a calculated risk in using it.

    """

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 23 10:45:13 2026
    From Newsgroup: comp.arch

    On 23/09/2026 07:49, Thomas Koenig wrote:
    David Brown <david.brown@hesbynett.no> schrieb:

    In microcontroller work, -Os is often used to put a stronger emphasis on
    size optimisations. I have noticed there are sometimes "blips" - cases
    where "-O2" leads to smaller code than "-Os", or where "-Os" leads to
    faster code than "-O2". And there can sometimes be cases where the
    speed/size tradeoffs are unreasonable - significantly slower code for
    very minor size improvements, or vice versa.

    Just to muddle the waters a little more: I once submitted a PR
    where -O3 led to smaller code than -Os. The reason? -O3 ran a
    pass which led to significant simplification and subsequent dead
    code elimination.

    Yes. It is not always easy to predict.

    I generally use -O2 in my embedded work. While -Os generally gives
    slightly smaller code for smaller test cases and functions, I find -O2
    often gives smaller total code in practice because you get better inter-procedural optimisations. It can seem counter-intuitive, but even adding a flag like "-fipa-cp-clone" (from -O3) to clone functions can
    end up with smaller object code in the end.


    I think one thing that is sometimes forgotten in all this, and could
    usefully be mentioned on the gcc manual page for optimisation, is that
    cpu architecture flags can make a significant difference. The step from
    "-O2" to "-O2 -march=native" is likely to be a lot bigger than the step
    from "-O2" to "-O3" in performance, with much lower risks of slowdowns.

    Additionally, -mtune=native can also help.


    Sure, although that is already active as part of "-march=native". If
    you want to distribute a binary that can run on a wide range of systems
    but is optimised for newer processors, "-mtune" is the way to go.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 23 13:17:29 2026
    From Newsgroup: comp.arch

    Scott Lurndal <scott@slp53.sl.home> schrieb:

    Ah, comp.arch - that explains it perfectly. USENET groups like comp.arch and comp.lang.c
    are notorious for regulars posting bizarre, intentionally "baffling" code fragments to
    spark academic debates, analyze compiler optimization side-effects, or illustrate
    architectural edge cases.

    ROTFLBTC :-)
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Wed Sep 23 06:45:19 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:

    [...]

    AFAIK, it has been argued that

    PRINT *,42
    END

    is a conforming C program because the gfortran command is a
    conforming C implementation (the driver also compiles C).

    How a compiler is invoked can affect whether it is (part of)
    a conforming C implmentation. For example

    gcc -x c -std=c99 -pedantic ...

    asks gcc to be a conforming C compiler; but if instead the
    invocation is

    gcc -x c++ ...

    then how gcc behaves is not consistent with being a conforming
    C compiler.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Wed Sep 23 14:01:29 2026
    From Newsgroup: comp.arch

    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/22/2026 5:32 AM, David Brown wrote:
    And note that in C23 the macro "unreachable()" was added with the sole
    semantics being "If a macro invocation unreachable() is reached during
    execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as >>opposed to say "implementation defined"?

    The idea here is obviously to use this macro only when its invocation
    really is unreachable, and in that case it has no effect on the
    behaviour.

    How should an implementation document what happens when the invocation
    of this macro actually is reachable? The effect depends on the
    surrounding code and the transformations in the compiler.

    IME, the annotation is there to squash compiler warnings.

    /**
    * Restore the state of the current core to the most recent
    * rollback state saved. This will result in the rollback handler
    * provided to the ::push function to be invoked.
    */
    inline void
    c_rollback::restore(uint64 rval)
    {
    _longjmp(&rb_state, rval);

    /* NOTREACHED */
    }

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Wed Sep 23 07:47:40 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 23 16:52:34 2026
    From Newsgroup: comp.arch

    On 23/09/2026 16:01, Scott Lurndal wrote:
    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/22/2026 5:32 AM, David Brown wrote:
    And note that in C23 the macro "unreachable()" was added with the sole >>>> semantics being "If a macro invocation unreachable() is reached during >>>> execution, the behavior is undefined" and "The program execution shall >>>> not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?

    The idea here is obviously to use this macro only when its invocation
    really is unreachable, and in that case it has no effect on the
    behaviour.

    How should an implementation document what happens when the invocation
    of this macro actually is reachable? The effect depends on the
    surrounding code and the transformations in the compiler.

    IME, the annotation is there to squash compiler warnings.

    /**
    * Restore the state of the current core to the most recent
    * rollback state saved. This will result in the rollback handler
    * provided to the ::push function to be invoked.
    */
    inline void
    c_rollback::restore(uint64 rval)
    {
    _longjmp(&rb_state, rval);

    /* NOTREACHED */
    }


    That is certainly a valid use of "unreachable()", but it would be wrong
    to suggestion it is the main intended or expected use.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 23 07:54:36 2026
    From Newsgroup: comp.arch

    On 9/23/2026 1:38 AM, David Brown wrote:
    On 23/09/2026 00:21, Stephen Fuld wrote:
    On 9/22/2026 5:32 AM, David Brown wrote:

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is reached
    during execution, the behavior is undefined" and "The program
    execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?-a Yes, it would make the
    compiler implementers do some work to document whatever they decided
    to do, but are there any other reasons?


    Good question.


    There is a sort of partial ordering (or lattice order) of behaviour specifications.-a "undefined behaviour" is at the bottom - it implements nothing, and anything else can implement it.-a "implementation-defined"
    is stronger.-a If the standards say "UB", then an implementation can
    define behaviour more concretely.-a If the standards say "IB", then the implementation cannot turn it into "UB".-a So "UB" in the standards
    clearly gives the most flexibility to the implementation.

    In one sense, yes. But in another sense, saying IB doesn't in any way constrain what the code for that implementation does - only that the implementer documents what the code that is written does.


    The big disadvantage of making something like this "IB" is not that the compiler implementers need to document their handling, but that they
    need to have a specification for it and stick to that - users must be
    able to rely on IB being consistent.

    Yes, but it allows that behavior to be anything the implementer chooses, including doing different things for different situations. The
    implementer has written some code to handle it. AFAICT the only
    difference is that IB requires the implementer document that code.

    Different types of handling for
    hitting "unreachable()" can include compiler optimisation on the
    assumption that it is never reached, generating a target "trap"
    instruction, generating a "breakpoint" when debugging, printing out an
    error message and terminating, calling a user-defined function that
    sends an angry email to the developer, and tracking all possible code
    flows at build time and giving a compiler warning if it can, in fact, be reached.

    Absolutely agreed, and perhaps other choices as well. :-) And perhaps
    even different choices depending upon the specifics of the code around
    it. ID doesn't preclude any choices here.

    -a Some users will prefer one method, others will prefer a
    different one, and many will change depending on what they are doing.

    Sure. And making the implementer document their implementation would
    give the user a basis to choose whether to use that capability in a
    particular situation, or indeed perhaps (though unlikely) choose which compiler implementation to use.

    Crucially, compilers may change what they support over time.-a A compiler cannot document "IB" as "depending on flags, this might be treated as
    UB, or give a trap, or perhaps do something else in the future".-a But
    it /can/ do exactly that if the standards say it is "UB".

    I don't see why IB can't be dependent on flags. In fact that might be
    useful. The compiler already does something. Just document what it
    does. And as for future choices, nothing in IB prevents future changes,
    as long as they are documented.

    It is also important to note that gcc, clang, icc, and many other
    compilers have had something like this for decades - like __builtin_unreachable() - which are implemented exactly as "undefined behaviour" in these tools.

    OK. But these implementations did *something* with these features. I
    am not asking them to change that in any way. Just document what that something is.


    This is, as far as I can tell, the proposal paper that led to "unreachable()" being added to the C standards:

    <https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2826.pdf>

    """
    We propose the feature unreachable to specify branches in the control
    N4eow of a program that will never be reached. The aim is to provide means for the user to express guarantees about the eN4Cective control N4eow that will be executed by a program. Compilers may then apply aggressive optimizations that otherwise would not be possibly or that would rely on
    the detection of undeN4Uned behavior for certain input combinations.

    Sounds reasonable to me.


    """

    The C++ version is :

    <https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p0627r6.pdf>

    The discussion about the definition is :

    """
    What is the best way to define the attribute's effect? The author feels
    that the best way is to make the behavior of std::unreachable() be undefined. There are several reasons:

    * std::unreachable() causing undefined behavior means that the Standard would not prescribe any particular action, leaving open many possible implementation actions.

    So does IB.


    * Some compilers already associate being unreachable to undefined
    behavior. ClangrCOs documentation states that __builtin_unreachable() rCLhas completely undefined behaviorrCY.

    * Optimizing under the assumption that a statement is unreachable, and
    thus having unpredictable behavior if the statement is in fact
    reachable, falls naturally under "undefined behavior".

    * An alternative would be to issue a trap if std::unreachable() is
    executed. This could be used in "debug builds", for example. Such a trap falls under "undefined behavior".

    * Being undefined behavior implies what happens if a constexpr function calls std::unreachable(): it's not a constant-expression, by (N4713) [expr.const]/2.6

    These seem to me to be sort of like "because we've always done it that
    way". While I agree that is true, it doesn't make it right.

    In the authorrCOs opinion, it is not undefined behavior itself that is the problem, but rather unexpected undefined behavior.

    This seems curious to me. If a behavior is expected, it isn't undefined.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 23 08:01:26 2026
    From Newsgroup: comp.arch

    On 9/22/2026 10:23 PM, Anton Ertl wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:
    On 9/22/2026 5:32 AM, David Brown wrote:
    And note that in C23 the macro "unreachable()" was added with the sole
    semantics being "If a macro invocation unreachable() is reached during
    execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?

    The idea here is obviously to use this macro only when its invocation
    really is unreachable, and in that case it has no effect on the
    behaviour.

    Yes. Although the programmer may have made an error by coding it when
    it is, in fact, reachable.


    How should an implementation document what happens when the invocation
    of this macro actually is reachable? The effect depends on the
    surrounding code and the transformations in the compiler.

    Fair enough. I am not saying otherwise. Just document it.


    And what would a programmer do with this documentation?

    He might know that a particular error he received was possibly or alternatively was not from the macro.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 23 18:24:16 2026
    From Newsgroup: comp.arch

    On 23/09/2026 16:54, Stephen Fuld wrote:
    On 9/23/2026 1:38 AM, David Brown wrote:
    On 23/09/2026 00:21, Stephen Fuld wrote:
    On 9/22/2026 5:32 AM, David Brown wrote:

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is reached
    during execution, the behavior is undefined" and "The program
    execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but can
    you give a possible rationale for making its behavior "undefined" as
    opposed to say "implementation defined"?-a Yes, it would make the
    compiler implementers do some work to document whatever they decided
    to do, but are there any other reasons?


    Good question.


    There is a sort of partial ordering (or lattice order) of behaviour
    specifications.-a "undefined behaviour" is at the bottom - it
    implements nothing, and anything else can implement it.
    "implementation-defined" is stronger.-a If the standards say "UB", then
    an implementation can define behaviour more concretely.-a If the
    standards say "IB", then the implementation cannot turn it into "UB".
    So "UB" in the standards clearly gives the most flexibility to the
    implementation.

    In one sense, yes.-a But in another sense, saying IB doesn't in any way constrain what the code for that implementation does - only that the implementer documents what the code that is written does.


    No, you are wrong here. And this misunderstanding is key to why it
    would be very limited to say "unreachable()" should be
    implementation-defined.

    From the C standards under "Terms, definitions, and symbols" :

    """
    * Implementation-defined behaviour

    Unspecified behaviour where each implementation documents how the choice
    is made.

    * Undefined behaviour

    Behaviour, upon use of a nonportable or erroneous program construct of erroneous data, for which this document imposes no requirements.

    * Unspecified behaviour

    Behaviour, that results from the use of an unspecified value, or other behaviour upon which this document provides two or more possibilities
    and imposes no further requirements on which is chosen in any instance.
    """

    IB means the standard gives certain options, and the implementation
    documents which choices are made. That is completely different from
    giving the implementation free reign. For example, right-shift of
    negative integers gives an implementation-defined value - the
    implementation can say it always gives the value 42, if it documents it,
    but it is not allowed to say it causes a call to abort() or prints out a run-time error message.


    The big disadvantage of making something like this "IB" is not that
    the compiler implementers need to document their handling, but that
    they need to have a specification for it and stick to that - users
    must be able to rely on IB being consistent.

    Yes, but it allows that behavior to be anything the implementer chooses, including doing different things for different situations.-a The
    implementer has written some code to handle it.-a AFAICT the only
    difference is that IB requires the implementer document that code.

    No. See above.


    Different types of handling for hitting "unreachable()" can include
    compiler optimisation on the assumption that it is never reached,
    generating a target "trap" instruction, generating a "breakpoint" when
    debugging, printing out an error message and terminating, calling a
    user-defined function that sends an angry email to the developer, and
    tracking all possible code flows at build time and giving a compiler
    warning if it can, in fact, be reached.

    Absolutely agreed, and perhaps other choices as well. :-)-a And perhaps
    even different choices depending upon the specifics of the code around
    it.-a ID doesn't preclude any choices here.


    See above.

    -a Some users will prefer one method, others will prefer a different
    one, and many will change depending on what they are doing.

    Sure.-a And making the implementer document their implementation would
    give the user a basis to choose whether to use that capability in a particular situation, or indeed perhaps (though unlikely) choose which compiler implementation to use.

    Crucially, compilers may change what they support over time.-a A
    compiler cannot document "IB" as "depending on flags, this might be
    treated as UB, or give a trap, or perhaps do something else in the
    future".-a But it /can/ do exactly that if the standards say it is "UB".

    I don't see why IB can't be dependent on flags.-a In fact that might be useful.-a The compiler already does something.-a Just document what it does.-a And as for future choices, nothing in IB prevents future changes,
    as long as they are documented.

    See above.

    I think we mostly agree on what we would like compiler implementations
    to do, at least in some cases (I am not sure if you like the "anything
    can happen" case of the compiler assuming unreachable() is never hit).
    But what you want here is only achievable if the standard says it is UB,
    not if the standard says it is IB.


    It is also important to note that gcc, clang, icc, and many other
    compilers have had something like this for decades - like
    __builtin_unreachable() - which are implemented exactly as "undefined
    behaviour" in these tools.

    OK.-a But these implementations did *something* with these features.-a I
    am not asking them to change that in any way.-a Just document what that something is.


    The "something" that they say is "If control flow reaches the point of
    the __builtin_unreachable, the program is undefined." That's the documentation in the gcc manual, along with some examples and use-cases.


    This is, as far as I can tell, the proposal paper that led to
    "unreachable()" being added to the C standards:

    <https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2826.pdf>

    """
    We propose the feature unreachable to specify branches in the control
    N4eow of a program that will never be reached. The aim is to provide
    means for the user to express guarantees about the eN4Cective control
    N4eow that
    will be executed by a program. Compilers may then apply aggressive
    optimizations that otherwise would not be possibly or that would rely
    on the detection of undeN4Uned behavior for certain input combinations.

    Sounds reasonable to me.


    """

    The C++ version is :

    <https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p0627r6.pdf>

    The discussion about the definition is :

    """
    What is the best way to define the attribute's effect? The author
    feels that the best way is to make the behavior of std::unreachable()
    be undefined. There are several reasons:

    * std::unreachable() causing undefined behavior means that the
    Standard would not prescribe any particular action, leaving open many
    possible implementation actions.

    So does IB.


    * Some compilers already associate being unreachable to undefined
    behavior. ClangrCOs documentation states that __builtin_unreachable()
    rCLhas completely undefined behaviorrCY.

    * Optimizing under the assumption that a statement is unreachable, and
    thus having unpredictable behavior if the statement is in fact
    reachable, falls naturally under "undefined behavior".

    * An alternative would be to issue a trap if std::unreachable() is
    executed. This could be used in "debug builds", for example. Such a
    trap falls under "undefined behavior".

    * Being undefined behavior implies what happens if a constexpr
    function calls std::unreachable(): it's not a constant-expression, by
    (N4713) [expr.const]/2.6

    These seem to me to be sort of like "because we've always done it that way".-a While I agree that is true, it doesn't make it right.

    In the authorrCOs opinion, it is not undefined behavior itself that is
    the problem, but rather unexpected undefined behavior.

    This seems curious to me.-a If a behavior is expected, it isn't undefined.


    It's a little odd, yes. It's like saying "unreachable()" is defined to
    be undefined behaviour. I understand what is meant by "expected
    undefined behaviour", but it is perhaps clumsy wording.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Wed Sep 23 11:20:25 2026
    From Newsgroup: comp.arch

    On 9/23/2026 9:24 AM, David Brown wrote:
    On 23/09/2026 16:54, Stephen Fuld wrote:
    On 9/23/2026 1:38 AM, David Brown wrote:
    On 23/09/2026 00:21, Stephen Fuld wrote:
    On 9/22/2026 5:32 AM, David Brown wrote:

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"?-a Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?


    Good question.


    There is a sort of partial ordering (or lattice order) of behaviour
    specifications.-a "undefined behaviour" is at the bottom - it
    implements nothing, and anything else can implement it.
    "implementation-defined" is stronger.-a If the standards say "UB",
    then an implementation can define behaviour more concretely.-a If the
    standards say "IB", then the implementation cannot turn it into "UB".
    So "UB" in the standards clearly gives the most flexibility to the
    implementation.

    In one sense, yes.-a But in another sense, saying IB doesn't in any way
    constrain what the code for that implementation does - only that the
    implementer documents what the code that is written does.


    No, you are wrong here.-a And this misunderstanding is key to why it
    would be very limited to say "unreachable()" should be implementation- defined.

    From the C standards under "Terms, definitions, and symbols" :

    """
    * Implementation-defined behaviour

    Unspecified behaviour where each implementation documents how the choice
    is made.

    * Undefined behaviour

    Behaviour, upon use of a nonportable or erroneous program construct of erroneous data, for which this document imposes no requirements.

    * Unspecified behaviour

    Behaviour, that results from the use of an unspecified value, or other behaviour upon which this document provides two or more possibilities
    and imposes no further requirements on which is chosen in any instance.
    """

    IB means the standard gives certain options, and the implementation documents which choices are made.-a That is completely different from
    giving the implementation free reign.-a For example, right-shift of
    negative integers gives an implementation-defined value - the
    implementation can say it always gives the value 42, if it documents it,
    but it is not allowed to say it causes a call to abort() or prints out a run-time error message.


    The big disadvantage of making something like this "IB" is not that
    the compiler implementers need to document their handling, but that
    they need to have a specification for it and stick to that - users
    must be able to rely on IB being consistent.

    Yes, but it allows that behavior to be anything the implementer
    chooses, including doing different things for different situations.
    The implementer has written some code to handle it.-a AFAICT the only
    difference is that IB requires the implementer document that code.

    No.-a See above.


    Different types of handling for hitting "unreachable()" can include
    compiler optimisation on the assumption that it is never reached,
    generating a target "trap" instruction, generating a "breakpoint"
    when debugging, printing out an error message and terminating,
    calling a user-defined function that sends an angry email to the
    developer, and tracking all possible code flows at build time and
    giving a compiler warning if it can, in fact, be reached.

    Absolutely agreed, and perhaps other choices as well. :-)-a And perhaps
    even different choices depending upon the specifics of the code around
    it.-a ID doesn't preclude any choices here.


    See above.

    -a Some users will prefer one method, others will prefer a different
    one, and many will change depending on what they are doing.

    Sure.-a And making the implementer document their implementation would
    give the user a basis to choose whether to use that capability in a
    particular situation, or indeed perhaps (though unlikely) choose which
    compiler implementation to use.

    Crucially, compilers may change what they support over time.-a A
    compiler cannot document "IB" as "depending on flags, this might be
    treated as UB, or give a trap, or perhaps do something else in the
    future".-a But it /can/ do exactly that if the standards say it is "UB".

    I don't see why IB can't be dependent on flags.-a In fact that might be
    useful.-a The compiler already does something.-a Just document what it
    does.-a And as for future choices, nothing in IB prevents future
    changes, as long as they are documented.

    See above.

    I think we mostly agree on what we would like compiler implementations
    to do, at least in some cases (I am not sure if you like the "anything
    can happen" case of the compiler assuming unreachable() is never hit).
    But what you want here is only achievable if the standard says it is UB,
    not if the standard says it is IB.

    It is obvious that you are right; the the existing choices as documented
    in the standard and your post above (thanks for that) don't permit what
    I want. I apologize for not recognizing that up front. But I maintain
    that there should be a category that essentially says "You can do what
    you want, but you must document what you did." Furthermore, if such a
    choice existed, most instances of undefined should be "reclassified"
    into that new choice. I think that would aid programmers understanding.
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Wed Sep 23 21:24:51 2026
    From Newsgroup: comp.arch

    On 23/09/2026 20:20, Stephen Fuld wrote:
    On 9/23/2026 9:24 AM, David Brown wrote:

    I think we mostly agree on what we would like compiler implementations
    to do, at least in some cases (I am not sure if you like the "anything
    can happen" case of the compiler assuming unreachable() is never hit).
    But what you want here is only achievable if the standard says it is
    UB, not if the standard says it is IB.

    It is obvious that you are right; the the existing choices as documented
    in the standard and your post above (thanks for that) don't permit what
    I want.-a I apologize for not recognizing that up front.

    No problem - you are not the first person to miss that detail about what "implementation-defined behaviour" means in the C standards!

    But I maintain
    that there should be a category that essentially says "You can do what
    you want, but you must document what you did."-a Furthermore, if such a choice existed, most instances of undefined should be "reclassified"
    into that new choice.-a I think that would aid programmers understanding.

    I can appreciate that desire. But I suspect it would be very difficult
    to specify exactly what is meant here - the wording would perhaps end up
    too vague to be appropriate for an ISO language standard. It is also important to remember that everything that is not given an explicit
    definition in the C standards is "undefined behaviour" by omission. UB
    is not just the things the things that are explicitly marked as UB in
    their descriptions:

    """
    If a rCyrCyshallrCOrCO or rCyrCyshall notrCOrCO requirement that appears outside of a
    constraint is violated, the behavior is undeN4Uned. UndeN4Uned behavior is otherwise indicated in this International Standard by the words rCyrCyundeN4Uned behaviorrCOrCO or by the omission of any explicit deN4Unition of behavior. There is no difference in emphasis among these three; they
    all describe rCyrCybehavior that is undeN4UnedrCOrCO.
    """

    It might, however, be possible to add footnotes to particular cases of
    UB with suggestions about how implementations might implement things or
    treat the UB.

    There is a move towards a concept of "erroneous behaviour" - this has
    been introduced in C++26, and will no doubt become part of the C
    standards at some point. This occurs, for example, when you try to use
    the value of an uninitialised local variable. It can, but does not have
    to, lead to a termination of the program and a diagnostic message - but
    AFAIUI it does not allow "time travelling" optimisations. It will not
    be appropriate for all types of UB - only in some specific cases can the compiler limit how badly things can go wrong in the face of UB.


    Annex L of the C standards distinguishes between "bounded undefined
    behaviour" and "critical undefined behaviour", where "bounded UB" does
    not do any stores (or volatile reads) outside of what is intended. (So
    signed integer overflow would be bounded UB, while writing beyond the
    end of an array would be critical UB.) In practice, AFAIK no compiler
    has tried to implement this and it has been fairly universally ignored.
    It is almost certainly too simplistic, as lots of apparently bounded UB
    could quickly lead to critical UB.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 23 19:28:21 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    How much more optimizations can compilers deliver to the bottom line
    (not just benchmarks). Last week we saw an episode where the compilers
    only got 1.x% speedup over years (sub-decade). Is there ever going to
    be a time to stop ??

    As of a hours minutes ago, there were 4547 bugs in GCC with the missed-optimization keyword. There are 229 open bugs blocking the
    "missed optimization in SPEC" bug, PR 26163 (but some of these
    are also ICEs). You will be amused to learn that there are 415
    open bugs regarding vectorization. There are regular regressions
    in SPEC where somebody observes a slowdown of 5% or 15% on some
    architecture or other.

    There is still quite a few optimizations to be done. Which ones have
    which effect on any particular piece of software is of course not
    clear.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Thu Sep 24 11:56:04 2026
    From Newsgroup: comp.arch

    David Brown <david.brown@hesbynett.no> schrieb:

    The "something" that they say is "If control flow reaches the point of
    the __builtin_unreachable, the program is undefined." That's the documentation in the gcc manual, along with some examples and use-cases.

    This has important (well, to compiler developers) use cases, for
    generating test cases. Grabbing a random example off bugzilla,
    from https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126048 :

    int src(int v1_u8) {
    if (!((1 <= v1_u8) && (v1_u8 <= 16))) __builtin_unreachable();
    int i0_u8 = v1_u8 << v1_u8;
    int i1_u8 = i0_u8 / v1_u8;
    return i1_u8;
    }

    The specific information used here is "This cannot happen, so
    optimize based on the knowledge that this is the case". It would
    be possible to infer the same information from other method, but
    this is quite convenient.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Thu Sep 24 16:35:21 2026
    From Newsgroup: comp.arch

    On 24/09/2026 13:56, Thomas Koenig wrote:
    David Brown <david.brown@hesbynett.no> schrieb:

    The "something" that they say is "If control flow reaches the point of
    the __builtin_unreachable, the program is undefined." That's the
    documentation in the gcc manual, along with some examples and use-cases.

    This has important (well, to compiler developers) use cases, for
    generating test cases. Grabbing a random example off bugzilla,
    from https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126048 :

    int src(int v1_u8) {
    if (!((1 <= v1_u8) && (v1_u8 <= 16))) __builtin_unreachable();
    int i0_u8 = v1_u8 << v1_u8;
    int i1_u8 = i0_u8 / v1_u8;
    return i1_u8;
    }

    The specific information used here is "This cannot happen, so
    optimize based on the knowledge that this is the case". It would
    be possible to infer the same information from other method, but
    this is quite convenient.


    This kind of "if (!XX) __builtin_unreachable();" is a common pattern.
    You can also express basically the same thing as [[gnu::assume(XX)]],
    but there can be slight differences in the kind of optimisation or
    static error check effects you get as they are handled in different
    passes of gcc.

    I have used this in my own code, wrapped in a macro so that it is easy
    to swap between checking for fault-finding and assuming for
    optimisation. But you might use :

    if ((i < 1) || (i > 2)) __builtin_unreachable();
    switch (i) {
    case 1 : ...
    case 2 : ...
    }

    as an alternative to having a default case with __builtin_unreachable().
    Either of these could be seen as more symmetric than :

    if (i == 1) {
    ...
    } else {
    ...
    }


    It can also be useful for static error checking. For switches,
    enumeration types and "-Wswitch-enum" can be useful to ensure that if
    you add another identifier to an enumeration and you have a switch using
    the enumeration, you get a warning if you've forgotten to add the new identifier to the switch. But sometimes that can be a bit more than you
    want, and lead to false positives. An alternative setup is :

    [[gnu::error("Should never happen"), noreturn]]
    void should_never_get_here() ;

    const int biggest_i = 2; // Or constexpr for C23

    int foo(int i) {
    if ((i < 1) || (i > biggest_i)) __builtin_unreachable();
    [[gnu::assume((i >= 1) && (i <= biggest_i))]];

    switch (i) {
    case 1 : return 100;
    case 2 : return 200;
    }
    should_never_get_here();
    }


    The "if (XX) __builtin_unreachable();" or "[[gnu::assume(XX)]];" lines
    (only one is needed) put compile-time limits on the value of "i" that
    the function expects. If you later decide that "i" can be as big as 3,
    your switch needs to be changed to match - or else the code will try to
    call "should_never_get_here()". If the compiler can't see that the code
    flow could never reach that "function", and generates a call to it, you
    get a compile-time error. Of course it is still your responsibility to
    make sure "foo" is only called with an appropriate "i", but it adds an
    extra compile-time safety check and documents your assumptions in the code.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Thu Sep 24 07:59:29 2026
    From Newsgroup: comp.arch

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Thu Sep 24 09:29:13 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    Let's call the new behavior type "implementation dependent".

    Suppose an implementation gives documentation that says "If a program
    execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."

    Question: are you okay with that? If not, how would you write a
    requirement in the C standard to limit the meaning of "you can do
    whatever you want, but you must document what you do" that gives only
    as much freedom as you think should be allowed?

    I think we will find that different people have very different ideas
    about how much freedom should be allowed in saying what happens. It's
    a very hard problem.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Thu Sep 24 17:05:31 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    A straightforward patch would
    very likely pessimize a lot of existing code which profits from >>>auto-vectorization.

    How can we test this claim?

    I do not believe that I have to explain the scientific method
    to you.

    As a first step, you would find an option that enables/disables
    what you don't like. -fstore-merging looks like a
    candidate, but there may be others. You can look at >https://dl.acm.org/doi/10.1109/ASE56229.2023.00209 or >https://link.springer.com/article/10.1007/s10515-024-00437-w
    if you want the full package.

    Then try this combination of options on other benchmarks
    as well as your pet one. I have my reservations about SPEC,
    but as you work at a university, you can get it for a discount.
    You can also use freely available benchmarks: Coremark, embench,
    Fortran Polyhedron - there are a lot.

    If there are benchmark which regress significantly with that
    set of options, you have your test case where it hurts.

    Like I wrote previously - you could have a bachelor or master
    student do this, it could be (part of a) nice thesis.

    That's a lot of words to express that you have no evidence for your
    claim.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 19:52:22 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    That would allow the compiler to fork "/bin/games/chess -play_both_sides";
    It is unlikely that the programmer would desire or expect that.

    I think you want closer to "Do something reasonable here" for a broad definition of reasonable.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Thu Sep 24 13:52:31 2026
    From Newsgroup: comp.arch

    On 9/24/2026 12:52 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    That would allow the compiler to fork "/bin/games/chess -play_both_sides";
    It is unlikely that the programmer would desire or expect that.

    I think you want closer to "Do something reasonable here" for a broad definition of reasonable.

    Ahhh, but that is the problem. Different people have different
    definitions of reasonable.

    While forking off a chess game wouldn't pass my (or, I suspect many
    people's) definition of reasonable, if it was documented, as I would
    require, you can't say it was unexpected. And, of course, I suspect
    that if it did that a lot, few people would use that compiler. :-)
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Fri Sep 25 05:43:03 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    That would allow the compiler to fork "/bin/games/chess -play_both_sides";
    It is unlikely that the programmer would desire or expect that.

    Early versions of gcc did a game of hack when encountering a
    pragma. The code looked like

    do_pragma ()
    {
    close (0);
    if (open ("/dev/tty", O_RDONLY, 0666) != 0)
    goto nope;
    close (1);
    if (open ("/dev/tty", O_WRONLY, 0666) != 1)
    goto nope;
    execl ("/usr/games/hack", "#pragma", 0);
    execl ("/usr/games/rogue", "#pragma", 0);
    execl ("/usr/new/emacs", "-f", "hanoi", "9", "-kill", 0);
    execl ("/usr/local/emacs", "-f", "hanoi", "9", "-kill", 0);
    nope:
    fatal ("You are in a maze of twisty compiler features, all different");
    }

    So there is precedent.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Fri Sep 25 05:39:28 2026
    From Newsgroup: comp.arch

    David Brown <david.brown@hesbynett.no> writes:
    On 21/09/2026 19:17, Anton Ertl wrote:
    I referred to a specific paradoxical effect, not completely avoiding
    all of them: The effect of programmers who, instead of working on
    making the code faster through source-level changes, have to invest
    time into avoiding getting it miscompiled.


    By "miscompiled", do you mean incorrect object code, or object code that
    did not have the performance the programmer expected or hoped for?

    Code that behaves differently than intended, e.g., where a bounds check
    is "optimized" away.

    Performance regressions can be nasty for performance-sensitive code,
    and consume time for finding a workaround (if one can be found), but
    is a different thing.

    But as
    for the problems you mention w.r.t UB, I think it's just the result of
    poor semantics, for which I guess we (language semanticists) are partly
    to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    I have never spoken to the C standards committee or writers, either of >current standard versions or the original C89 standard, or writers of >pre-standard C specifications. So I cannot in any way claim to know
    their thoughts or motivations. I also have not seen any documentation
    that suggests what they might have thought about "optimisation on the >assumption that undefined behaviour does not occur" - in either
    direction.

    When C89 was developed, the dominating C compiler in the Unix world
    was PCC. GCC came on the scene late in the game (1988), without such assumptions and produced a significant performance advantage. Only
    after several years of heavy gcc development assumptions like
    -fstrict-aliasing were introduced around gcc-2.7 (which was initially
    disabled by default); I don't know when the assumption that signed
    integer operations do not overflow was introduced, and initially it
    could not be disabled (-fwrapv was added around gcc-3.4), so obviously
    by that time people had taken over the gcc project who had a different
    view about how undefined behaviour should be treated.

    Maybe the change can be pinpointed to the version when
    -fstrict-aliasing became the default, or maybe the attitude change was
    more gradual.

    Anyway, given the state of C compilers during C89 development, no, I
    don't think they ever thought that undefined behaviour would be used
    to justify "optimizing" away bounds checks, such as the second one
    below:

    char *buf = ...;
    char *buf_end = ...;
    unsigned int len = ...;
    if (buf + len >= buf_end)
    return; /* len too large */
    if (buf + len < buf)
    return; /* overflow, buf+len wrapped around */
    /* write to buf[0..len-1] */

    [This example is given in <https://people.csail.mit.edu/nickolai/papers/wang-stack.pdf>, which
    cites <https://www.kb.cert.org/vuls/id/162289>]

    (It has been mentioned that, for example, a two's complement
    implementation could use wrapping signed arithmetic - but I have never
    seen a suggestion that this behaviour should be encouraged or expected
    just because a machine uses two's complement.)

    The C99 rationale (which obviously contains text written for C89)
    includes the following:

    |C code can be non-portable. Although it strove to give
    |programmers the opportunity to write truly portable programs, the C89 |Committee did not want to force programmers into writing portably, to |preclude the use of C as a "high-level assembler": the ability to
    |write machine-specific code is one of the strengths of C. It is this |principle which largely motivates drawing the distinction between
    |strictly conforming program and conforming program (Section 4).
    |
    |Keep the spirit of C. The C89 Committee kept as a major goal
    |to preserve the traditional spirit of C. There are many facets of the
    |spirit of C, but the essence is a community sentiment of the
    |underlying principles upon which the C language is based. Some of the
    |facets of the spirit of C can be summarized in phrases like:
    |
    | * Trust the programmer.
    | * Don't prevent the programmer from doing what needs to be done.
    | * Keep the language small and simple.
    | * Provide only one way to do an operation.
    | * Make it fast, even if it is not guaranteed to be portable.
    |
    |The last proverb needs a little explanation. The potential for
    |efficient code generation is one of the most important strengths of
    |C. To help ensure that no code explosion occurs for what appears to be
    |a very simple operation, many operations are defined to be how the
    |target machine's hardware does it rather than by a general abstract rule.

    So no, it does not say what a compiler should do, but I understand it
    as saying that it is in the spirit of C if a programmer relies on
    2s-complement wraparound arithmetic for signed integers on an
    architecture where C compilers implement signed integer arithmetic in
    that way.

    That was certainly the case for many architectures in 1989, where no C compilers on MIPS produced the add or addi instructions (which trap on
    signed overflow) for signed arithmetic even though it was available.
    It was also the case for 68000, IA-32, HPPA, SPARC, Alpha, and others.
    And whenever I looked at the code produced for MIPS and Alpha, I have
    never seen the trapping addition/subtraction/multiplication
    instructions generated. Maybe if you ask for it with -ftrapv, but
    IIRC I looked at generated code once and found that the trapping
    instructions were ignored even then, instead generating more
    long-winded overflow checks.

    You do not ask "what happens
    if my signed integer arithmetic overflows?" or "what happens when I
    access an array out of bounds?" - rather, it is your responsibility as a
    C programmer to make sure that never happens.

    That's not at all what I read in the rationale.

    The prime reason C does
    not define behaviour here is not that different hardware or compilers
    handle things differently, but that there is no sensible definition that >could be made.

    Java had no problem producing a sensible definition for what happens
    on signed overflow.

    Likewise, for the bounds check above, there is a sensible definition
    that can be made, and gcc actually still (or again) uses it if you
    tell it to, with -fwrapv-pointer.

    Remember, C has a perfectly good way to say "this is determined by the >hardware or the implementation" - it is "implementation-defined
    behaviour". It has a perfectly good way to say that "this operation
    could result in any value" - it is "unspecified behaviour" or
    "unspecified value".

    When the C standards writers say something is "undefined behaviour",
    rather than "implementation defined" or "unspecified", it is my belief
    that they did so intentionally and knowingly.

    The question is what the intention was. My impression is a decisive
    reason for declaring something undefined has been if straightforwardly generated code could trap on some architecture. This becomes most
    apparent when it comes to shifts, where some cases are
    implementation-defined and some are undefined, and the undefined cases correspond to cases that trap on some architecture.

    I do not care that much about the language-lawyering about the
    different names for the incomplete definitions in the C standards, but
    if an instruction traps, that does not look like nasal demons to me
    (Ariane 501 customers may disagree, but one can blame the
    inappropriate handling of the trap, based on the proof that the trap
    cannot happen).

    Maybe one of the C language lawyers can explain why they never chose
    to use "imlementation-defined" or any of the other incompleteness
    variants when a trap is possible.

    And note that in C23 the macro "unreachable()" was added with the sole >semantics being "If a macro invocation unreachable() is reached during >execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation". The justification is for better
    diagnostics and optimisations.

    How can "undefined behaviour" lead to better diagnostics given the
    attitude that "it is your responsibility as a C programmer to make
    sure that never happens"? Whoever wrote this justfication apparently
    has a different idea of the meaning of undefined behaviour than you
    do.

    Anyway, undefined behaviour in connection with "unreachable()" or
    "restrict" is not a problem as far as I am concerned. A programmer
    can just choose to never use "unreachable()" or "restrict", and avoid
    any undefined behaviours coming from these language features. And
    when he introduces them, my recommendation is to do it sparingly, only
    in those places that are relevant to performance.

    This is in contrast to taking the assumption that undefined behaviour
    does not happen anywhere, which means that programmers have to check
    everywhere for the >200 explicitly stated (and who knows how many
    implied) undefined behaviours in the current C standard.

    I don't know whether or not the "founding fathers" of C intended or
    expected compilers to optimise on the assumption that UB did not occur.

    One can look at the code written by Ritchie, Thompson and the other
    people at Bell Labs. Does it contain undefined behaviour? Very
    likely.

    But I am confident that they considered a program to be broken if
    execution reached a point where the behaviour was not defined in an >implementation (something may be UB in the C standards yet defined by an >implementation).

    If you mean "documented as defined", I doubt it. I expect that they
    did write programs that rely on the actual code generated by the
    compiler(s) the used for the program, on the target(s) they used for
    the program. And of course the C compilers before C89 did not
    document behaviours as defined that would be undefined by the standard
    only later.

    It is, of course, possible that they did not foresee quite how this
    would pan out in modern compilers. In particular, they may not have >predicted "time-travel" optimisations. However, authors of more modern
    C standards - say, C11 onwards - know about them and have could have >explicitly outlawed them if they were considered to be invalid.

    There is no need to outlaw time travel in C standards, because it was
    never allowed. At least that's what I read here some years ago. Time traveling is a C++, not C property.

    For
    heavily used compilers, there is a strong (but not overpowering) push >towards backwards compatibility. This is why they have heavy regression >tests, and pre-release versions are tested with large samples of
    important existing code. If this gives unexpected results, these must
    be dealt with - was it a bug in the new compiler (in which case the fix
    is obvious), or was it a bug in the old source code? Those cases are
    more complicated - sometimes the old code must be fixed, but sometimes
    the incorrect source code is too common, idomatic or important and the >compiler must, in effect, support additional semantics to retain the old >accidental semantics.

    Yes, my impression is that the pushback after miscompiling "relevant
    code" has been strong enough that the regression tests now avoid
    miscompiling that code, and a lot of the "irrelevant" code such as
    Gforth rides in the slipstream of that. Of course, relevant code like
    the Linux kernel uses a collection of flags (e.g,
    -fno-strict-overflow) that define otherwise undefined behaviour, so
    anybody who wants to ride in the slipstream of that code should use
    these flags as well, and leave the benchmarking settings (i.e.,
    without these flags) to the benchmarks.

    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus:
    <https://blog.regehr.org/archives/1287>.

    There are two key problems with UB that stand in the way of this kind of >initiative. One is that many compilers provide consistent (and
    sometimes even documented) behaviour for some things that are UB in the
    C standards. Code relying on these behaviours is valid and safe, but >non-portable. But different compilers (or the same compiler but
    different targets) could easily have different semantics, making it very >difficult to agree on any one choice. Secondly, many programmers write
    code on the assumption that certain UB has certain consistent and
    reliable behaviour - even though it is not documented anywhere. It is
    code like that which could benefit from a "friendly C" (a poor choice of >name, IMHO, but that's entirely subjective) variant. But it is also
    such code that makes a "friendly C" variant hard to define - such code
    is hard to identify, and the expected behaviour can be even harder to
    find, specify, and consistently describe.

    Such efforts try to achieve more than what is discussed in the part of
    the C99 rationale cited above, and more than I argue for in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>; it also
    tries to solve portability between architectures and compilers. And
    given that such efforts have not come to fruition, one can conclude
    that they try to achieve too much.

    Alternatively, one can see the various flags provided by gcc, such as -fno-strict-aliasing as achieving at least a part of such efforts (not
    the portability between compilers, in general, but at least between architectures). There is at least one case (-fwrapv vs. -ftrapv)
    where one can ask gcc for one of several definined behaviours, though.

    But fortunately a completely specified C is not necessary, a
    willingness to preserve the behaviour of existing working programs
    compiled with an earlier version of the same compiler is. Read more
    about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.


    I think it is entirely reasonable to keep old compiler versions around
    and use them for old code.

    Unfortunately, the various Linux distributions do not agree with you,
    and do not distribute gcc versions back to 1.0. So, for software
    distributed as source code, asking for a known-good compiler for
    IA-32, such as gcc-2.95, is impractical.

    I think it is entirely unreasonable to try to specify that new compilers >should have defined specifications to implement the "semantics" that
    older compilers used for particular types of undefined behaviour. These >semantics will, for the most part, be poorly defined and can often be >inconsistent - many programs with UB work by luck, not design.

    On the contrary, the most common cases of undefined behaviour become well-defined. E.g., before gcc ever generated code based on the
    assumption that signed overflow never happens, it generated addu or
    equivalent (e.g., addiu or the implied addition in the address
    computation of a load) for an addition in the C source code; and
    optimizations preserved this behaviour, e.g., a+(-b) was compiled to
    subu, not sub. Continuing to generate code that behaves like that is well-defined, and it now even has flags in gcc (-fwrapv and
    -fwrapv-pointer, which can be combined into -fno-strict-overflow).

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.


    Linux commits to preserving the /defined/ behaviour of user-space APIs.

    If you mean that it takes the same work-to-rule approach that the gcc
    people do when it comes to "irrelevant" code, that's not the case. It
    commits to preserving the actual behaviour. That is well publicised,
    e.g. <https://linuxreviews.org/WE_DO_NOT_BREAK_USERSPACE>. Linus
    Torvalds wrote:

    |If a change results in user programs breaking, it's a bug in the
    |kernel. We never EVER blame the user programs.

    But if a
    particular invalid value of the parameter lets you gain access to
    root-owned files, due to a bug in the kernel, you can be very sure that
    this user-space behaviour will /not/ be preserved.

    I don't know if the maintainers of a user program ever reported a bug
    for closing such a security hole and insisted on preserving the
    behaviour; I doubt it. If they would, one way to deal with that would
    be to have a kernel option (compile-time, startup-time, or run-time)
    for opening the hole, which the default being that the hole is closed.

    There has been at least one well-known case where the kernel people
    went to great lengths to preserve a behaviour for a certain userspace
    program where the other behaviour already also had users (and in this
    case the kernel-option variant was not satisfactory, so they used
    something far uglier).

    This has all been well publicised, which makes me wonder why you are
    spreading the claims above; either you do not know what you are
    writing about, or you knowingly spread an untruth.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 25 11:06:49 2026
    From Newsgroup: comp.arch

    On 24/09/2026 22:52, Stephen Fuld wrote:
    On 9/24/2026 12:52 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"?-a Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting.-a A construct having implementation-defined behavior
    cannot do "just anything";-a the C standard must specify what
    behavior choices are possible.-a Any fixed set of choices might
    rule out what an implementation would like to do.-a So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard.-a But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    That would allow the compiler to fork "/bin/games/chess -
    play_both_sides";
    It is unlikely that the programmer would desire or expect that.

    I think you want closer to "Do something reasonable here" for a broad
    definition of reasonable.

    Ahhh, but that is the problem. Different people have different
    definitions of reasonable.


    And there you have it - the crux of the issue, and the key reason why no programming language can have Tim's hypothetical "implementation
    dependent" behaviour. (That would be an extraordinarily confusing
    choice of name, but I am sure Tim is aware of that.) You can't have a specification that says "do something reasonable" - or "the compiler can
    do anything reasonable as long as it is documented".

    So the specification says "this document does not say anything about
    what will happen - it does not need to be defined, consistent, or
    something that any given programmer finds reasonable". Then it is up to compiler writers to decide what /they/ consider reasonable - possibly
    with different variations with different flags. And then the programmer
    can decide what /they/ consider reasonable - choosing a compiler and/or
    flags that suit their own preferences. Or they can write code that does
    not have the such constructs in the first place, and avoid the issue
    entirely.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Fri Sep 25 07:27:24 2026
    From Newsgroup: comp.arch

    On 9/24/2026 9:29 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"? Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting. A construct having implementation-defined behavior
    cannot do "just anything"; the C standard must specify what
    behavior choices are possible. Any fixed set of choices might
    rule out what an implementation would like to do. So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard. But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    Let's call the new behavior type "implementation dependent".

    I am not hung up on the name, so at least for purposes of discussion, OK.


    Suppose an implementation gives documentation that says "If a program execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."

    Question: are you okay with that?

    If that is the implementation's general response, then, while it might
    be "legal", then no, I am not happy with it.


    If not, how would you write a
    requirement in the C standard to limit the meaning of "you can do
    whatever you want, but you must document what you do" that gives only
    as much freedom as you think should be allowed?

    I have minimal experience with standards writing, and none with language writing, so I may be off base here.

    I think that there are many situations where the standard, correctly,
    doesn't specify anything about what the compiler should do. These are currently called undefined behavior, and the standard essentially lets
    it go at that - no further documentation required.

    But in many, but not all, of such cases, the compiler implementer knows
    what it is going to do. What I am after is that in such cases, the implementer "supplement" the bare words with information that he has
    that knowledge and he should communicate it to the user.


    I think we will find that different people have very different ideas
    about how much freedom should be allowed in saying what happens. It's
    a very hard problem.

    Note that I am not trying to restrict "what happens" in any way. What
    happens is anything that could/would happen under today's definition of undefined behavior. I am just trying to ask the implementer, whenever possible (and I realize that it is not always possible) to elaborate on
    the bare words "undefined behavior".
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Fri Sep 25 14:36:50 2026
    From Newsgroup: comp.arch

    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    David Brown <david.brown@hesbynett.no> writes:

    I have never spoken to the C standards committee or writers, either of >>current standard versions or the original C89 standard, or writers of >>pre-standard C specifications. So I cannot in any way claim to know
    their thoughts or motivations. I also have not seen any documentation >>that suggests what they might have thought about "optimisation on the >>assumption that undefined behaviour does not occur" - in either
    direction.

    When C89 was developed, the dominating C compiler in the Unix world
    was PCC. GCC came on the scene late in the game (1988), without such >assumptions and produced a significant performance advantage.

    Well, yes. PCC at that point was 14 years old and showing its age.

    Motorola used PCC for the 88100 processor (I had to fix a bug in the
    register allocator when compiling the output of cfront, which
    extensively uses the comma operator in 1990); I had first used PCC in
    1981.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 25 16:44:34 2026
    From Newsgroup: comp.arch

    On 25/09/2026 07:39, Anton Ertl wrote:
    David Brown <david.brown@hesbynett.no> writes:
    On 21/09/2026 19:17, Anton Ertl wrote:
    I referred to a specific paradoxical effect, not completely avoiding
    all of them: The effect of programmers who, instead of working on
    making the code faster through source-level changes, have to invest
    time into avoiding getting it miscompiled.


    By "miscompiled", do you mean incorrect object code, or object code that
    did not have the performance the programmer expected or hoped for?

    Code that behaves differently than intended, e.g., where a bounds check
    is "optimized" away.


    The trouble with that is determining precisely what was intended.
    Compilers are limited - the programmer has to express their intention in
    the restricted form of the programming language (C or whatever else).
    And the /only/ things that can be relied on to have a clear meaning are
    things that have clear defined behaviour as documented by the language standards or the compiler documentation.

    A program is a contract between the programmer and the compiler - the programmer promises to provide good code with a clearly defined meaning,
    and the compiler promises to give object code that implements that
    defined meaning. If the programmer breaks their side of the contract by writing code that does not have a clear meaning (at compile-time or
    run-time), the compiler /cannot/ fulfil its side of the contract. It is unable to generate "correct" object code from incorrect source code.

    If the programmer writes "a = b + c;" (all signed int), what is the
    intention? Does the programmer mean "generate an ADD assembly
    instruction"? Does he/she mean "I don't care how it happens, but 'a'
    should be the sum of 'b' and 'c'"? Do they mean "give me a two's
    complement wrapped addition"? Do they mean "I promise there will be no overflow" ? Do they mean "I /hope/ there will be no overflow" ? Or
    "While I am debugging, let me know if there is an overflow" ?

    The programmer might have intended lots of subtly different things from
    the same bit of code. But compilers cannot work by mind-reading. The
    only thing that the /code/ intends is "Add these two numbers. I promise
    there will be no overflow". Whatever the programmer imagines, that is
    the intention of the code. The compiler's job is to follow that - and
    if the programmer breaks his promise, it is garbage in, garbage out.


    I find the whole concept of optimising based on "what the programmer intended", rather than what the code they wrote means, somewhat
    hypocritical. The idea that "a = b + c;" should "obviously" wrap on a
    two's complement machine, while "a = b + 10 + c - 10;" should
    "obviously" optimise away the constants is clearly contradictory. It is
    an attitude that is denying responsibility - claiming "the compiler
    broke my code" rather than understanding "my code had mistakes but
    previously gave the right results by luck".

    I am not underestimating the difficulty of writing correct code here, or
    the frustration and costs involved when the behaviour of real-world code changes with different compilers or flags. But it is critical that programmers understand the origin of problems in order to deal with the
    issues instead of passing the blame around. Baring compiler bugs (even
    my favourite compiler has one or two outstanding bugs), what you term "miscompiling" here is actually errors in the programmers' source code.

    The way to improve the situation is not by blaming compiler writers, or inventing some kind of vague notion of what programmers "intended" or
    "do what some old compiler did", or limiting optimisations in compilers.
    Instead, it is by improving compilers, run-time checkers and static
    analysis tools to help identify the errors in the source code, improving teaching of programming to reduce the risk of such errors, improving
    languages to support lower-risk constructs, changing to programming
    languages with different characteristics, etc. Indeed, I find that
    better compilers with more powerful optimisations means fewer bugs in my
    code because I can write the code in clearer ways and still get
    efficient results in the end.


    Performance regressions can be nasty for performance-sensitive code,
    and consume time for finding a workaround (if one can be found), but
    is a different thing.


    Agreed. And correctness trumps speed every time.

    But as
    for the problems you mention w.r.t UB, I think it's just the result of >>>> poor semantics, for which I guess we (language semanticists) are partly >>>> to blame: we have developed fairly good tools to design sane language
    specs in general,

    I violently disagree. The C89 standard (just to name one) is a
    partial specification not because the original C standards people were
    poor at specifying semantics, but because given the differences
    between the compilers out there and the hardware out there, the
    easiest way to reach a consensus is to leave some parts unspecified.

    I have never spoken to the C standards committee or writers, either of
    current standard versions or the original C89 standard, or writers of
    pre-standard C specifications. So I cannot in any way claim to know
    their thoughts or motivations. I also have not seen any documentation
    that suggests what they might have thought about "optimisation on the
    assumption that undefined behaviour does not occur" - in either
    direction.

    When C89 was developed, the dominating C compiler in the Unix world
    was PCC. GCC came on the scene late in the game (1988), without such assumptions and produced a significant performance advantage. Only
    after several years of heavy gcc development assumptions like -fstrict-aliasing were introduced around gcc-2.7 (which was initially disabled by default);

    ("-fstrict-aliasing" is still not enabled by default - you have to
    actively enable at least "-O2".)

    I don't know when the assumption that signed
    integer operations do not overflow was introduced, and initially it
    could not be disabled (-fwrapv was added around gcc-3.4), so obviously
    by that time people had taken over the gcc project who had a different
    view about how undefined behaviour should be treated.

    I completely disagree that this is an "obvious" conclusion. While I
    cannot say anything for sure here, the explanation that comes first to
    my mind is that it was part of the natural progression of gcc towards
    more optimisations. "-fwrapv" would have been added as a result of
    complaints about problems with old code that made invalid assumptions
    that happened to "work" because older gcc did not have as many
    optimisations. I see absolutely no reason to assume that anyone at any
    time thought such optimisations were not valid - they simply had not yet
    been implemented.

    The first time I used a compiler that had "UB optimisation" was a Diab
    Data compiler, perhaps in 1997 or so. It generated code assuming signed integer arithmetic did not overflow. This is not a concept invented by
    some evil minds at GCC that differed from how previous compilers worked
    - the concept of "garbage in, garbage out" in programming was known to Babbage.


    Maybe the change can be pinpointed to the version when
    -fstrict-aliasing became the default, or maybe the attitude change was
    more gradual.

    Anyway, given the state of C compilers during C89 development, no, I
    don't think they ever thought that undefined behaviour would be used
    to justify "optimizing" away bounds checks, such as the second one
    below:

    char *buf = ...;
    char *buf_end = ...;
    unsigned int len = ...;
    if (buf + len >= buf_end)
    return; /* len too large */
    if (buf + len < buf)
    return; /* overflow, buf+len wrapped around */
    /* write to buf[0..len-1] */

    [This example is given in <https://people.csail.mit.edu/nickolai/papers/wang-stack.pdf>, which
    cites <https://www.kb.cert.org/vuls/id/162289>]


    I can happily agree that it did not occur to some /programmers/ that UB
    would lead to things like the removal of that second test. I would be
    very surprised, however, if it had not occurred to the writers of the C standards or at least some earlier compiler writers that compilers can
    make assumptions that code does not have UB, and use that for
    optimisation. The C standards support that viewpoint, and have always
    done so.

    At the very least, it has /always/ been the case that relying on the
    effects of UB is a bad idea and gives, at best, fragile and non-portable
    code. It should be one of the first things anyone learns about
    programming - follow the rules of the language. And it has always been
    known that in C, the compiler trusts the programmer - it is a language targeting maximal efficiency, not maximal user-friendliness or run-time checking.

    Of course people have still written code with UB, and relied on
    particular behaviours from it. Sometimes they have done so
    unintentionally - they may have thought that the addition in the code
    above has wrapping semantics. Sometimes they have done so knowingly,
    and seen that the resulting object code does what they want in an
    efficient manner, while a more correct solution could be less efficient.
    (That can be an appropriate development choice, but is clearly non-portable.) Sometimes it's just a normal bug in the code. And
    sometimes, perhaps most often, it is a result of using someone else's
    code that you thought was correct.

    The "answer" is not to say that something is "miscompiled", blame the
    compiler writers, and try to invent nebulous and vague concepts of what programmers had "intended" or that compilers should "treat code like
    they used to". The aim should be to find and fix the bugs, and to avoid writing such code in the future.


    I want my compiler to eliminate checks that cannot possibly be matched.
    I want my compiler to assume that UB does not occur and use that for optimisation. That is my /intention/ when writing code and using a
    compiler. It means I can add additional checks that have zero cost, but
    that improve the safety of my code because they may be applied later
    (such as when some compile-time configured constant is changed in a
    future code version). Such eliminatable checks may be explicit in the
    code, or may be the result of macros, function inlining, etc.

    But if compiler static analysis could spot the mistake in your example,
    and warn about it without flooding other code with false positives, then
    that would be great. That is not an easy task - especially since "what
    the programmer intended", beyond what they actually wrote as defined by
    the C language, is guesswork at best. gcc used to have a warning you
    could enable for when it eliminated dead code - but as the compiler got
    better at finding such dead code, such warnings became overwhelming and
    the feature was removed.


    (It has been mentioned that, for example, a two's complement
    implementation could use wrapping signed arithmetic - but I have never
    seen a suggestion that this behaviour should be encouraged or expected
    just because a machine uses two's complement.)

    The C99 rationale (which obviously contains text written for C89)
    includes the following:

    |C code can be non-portable. Although it strove to give
    |programmers the opportunity to write truly portable programs, the C89 |Committee did not want to force programmers into writing portably, to |preclude the use of C as a "high-level assembler": the ability to
    |write machine-specific code is one of the strengths of C. It is this |principle which largely motivates drawing the distinction between
    |strictly conforming program and conforming program (Section 4).
    |
    |Keep the spirit of C. The C89 Committee kept as a major goal
    |to preserve the traditional spirit of C. There are many facets of the |spirit of C, but the essence is a community sentiment of the
    |underlying principles upon which the C language is based. Some of the |facets of the spirit of C can be summarized in phrases like:
    |
    | * Trust the programmer.
    | * Don't prevent the programmer from doing what needs to be done.
    | * Keep the language small and simple.
    | * Provide only one way to do an operation.
    | * Make it fast, even if it is not guaranteed to be portable.
    |
    |The last proverb needs a little explanation. The potential for
    |efficient code generation is one of the most important strengths of
    |C. To help ensure that no code explosion occurs for what appears to be
    |a very simple operation, many operations are defined to be how the
    |target machine's hardware does it rather than by a general abstract rule.

    So no, it does not say what a compiler should do, but I understand it
    as saying that it is in the spirit of C if a programmer relies on 2s-complement wraparound arithmetic for signed integers on an
    architecture where C compilers implement signed integer arithmetic in
    that way.

    I do not understand it in that way at all. I /do/ understand the
    standard as saying it is fine for a compiler to have wrapping semantics
    on signed integer overflow, and I understand the rationale as saying
    that nothing should prevent a programming using a given compiler in the
    way that compiler works, even if the code is not portable. The
    rationale here makes it clear that if you have code that benefits
    greatly from wrapped signed integer arithmetic, you can write code that
    only works for "gcc -fwrapv" - the C standards should not limit the
    programmer or the compiler from doing that even though it might fail on
    a different compiler.

    But I cannot see an interpretation of either the standards themselves
    (which are the only really important document) or the rationale
    documents that says programmers are able to rely on particular
    behaviours for UB. In situations where the C standards expect behaviour
    to be defined and consistent but do not want to give a fixed definition themselves, they use "implementation-defined" behaviour. That is
    clearly distinct from "undefined behaviour".


    That was certainly the case for many architectures in 1989, where no C compilers on MIPS produced the add or addi instructions (which trap on
    signed overflow) for signed arithmetic even though it was available.

    Choices made by old compilers - either active and intentional choices,
    or simply choices due to limited time, money, developers, etc,. - are
    never an indication of how things were intended to be done, or a guide
    of how things are supposed to be done in the future.

    It was also the case for 68000, IA-32, HPPA, SPARC, Alpha, and others.
    And whenever I looked at the code produced for MIPS and Alpha, I have
    never seen the trapping addition/subtraction/multiplication
    instructions generated. Maybe if you ask for it with -ftrapv, but
    IIRC I looked at generated code once and found that the trapping
    instructions were ignored even then, instead generating more
    long-winded overflow checks.

    You do not ask "what happens
    if my signed integer arithmetic overflows?" or "what happens when I
    access an array out of bounds?" - rather, it is your responsibility as a
    C programmer to make sure that never happens.

    That's not at all what I read in the rationale.

    The prime reason C does
    not define behaviour here is not that different hardware or compilers
    handle things differently, but that there is no sensible definition that
    could be made.

    Java had no problem producing a sensible definition for what happens
    on signed overflow.

    No, Java has a silly definition of what happens on signed overflow. It
    is simple and consistent, and it is one of many viable choices if you
    decide that you /must/ have definition behaviour - but it is a silly
    choice. There are almost no circumstances where it makes sense to add
    to positive numbers and get a negative result. All Java does is give
    you a guaranteed useless result, while stopping the use of tools like "-fsanitize=signed-integer-overflow" that can help catch bugs in your code.

    C has a well-defined way to support wrapping semantics - if you need
    wrapping, use that. Cast to an appropriate unsigned type as needed.
    (With C23 you also have functions like ckd_add and ckd_sub - it would
    have been nice for them to have been standard long ago, but gcc at least
    has had equivalent built-ins for a very long time.)


    Likewise, for the bounds check above, there is a sensible definition
    that can be made, and gcc actually still (or again) uses it if you
    tell it to, with -fwrapv-pointer.

    Yes. That is another example of picking a specific compiler
    implementation ("gcc -fwrapvpointer" rather than "gcc" or other
    compilers) that is documented to provide a certain behaviour here - C
    does not limit you from doing that. But it also does not encourage it,
    or in any way imply that because one implementation variant supports
    this code, other implementations or variants should also do so.

    And C has a simple and effective way to get the same effect in fully
    defined C :

    if ((uintptr_t) buf + len < (uintptr_t) buf)


    Remember, C has a perfectly good way to say "this is determined by the
    hardware or the implementation" - it is "implementation-defined
    behaviour". It has a perfectly good way to say that "this operation
    could result in any value" - it is "unspecified behaviour" or
    "unspecified value".

    When the C standards writers say something is "undefined behaviour",
    rather than "implementation defined" or "unspecified", it is my belief
    that they did so intentionally and knowingly.

    The question is what the intention was. My impression is a decisive
    reason for declaring something undefined has been if straightforwardly generated code could trap on some architecture. This becomes most
    apparent when it comes to shifts, where some cases are
    implementation-defined and some are undefined, and the undefined cases correspond to cases that trap on some architecture.

    I agree that this is /sometimes/ the reason for marking a particular
    behaviour as UB. But it is wrong to think it is the /only/ reason.


    I do not care that much about the language-lawyering about the
    different names for the incomplete definitions in the C standards, but
    if an instruction traps, that does not look like nasal demons to me
    (Ariane 501 customers may disagree, but one can blame the
    inappropriate handling of the trap, based on the proof that the trap
    cannot happen).

    There are a number of circumstances where the C standards explicitly
    give traps as an option for behaviour. And traps are always an implicit option for UB.

    But is seems questionable to say that you don't care about the details
    of incomplete definitions in the C standards while also holding very
    strong opinions on how you think undefined behaviour should be defined.


    Maybe one of the C language lawyers can explain why they never chose
    to use "imlementation-defined" or any of the other incompleteness
    variants when a trap is possible.

    There is certainly an example of "either the result is
    implementation-defined or an implementation-defined signal is raised" -
    for conversions to signed integer types when the original value cannot
    be represented in the new type. I do not know why this allows for a
    "signal" but not a "trap". (As all signed integer types are now, as of
    C23, two's complement, then this could easily be replaced by a fixed definition as wrapping behaviour to match all real-world implementations.)

    I think it would be entirely reasonable to change some or all of the
    shift behaviours to being implementation-defined or fully defined.
    There are other cases of UB that could also be eliminated.


    And note that in C23 the macro "unreachable()" was added with the sole
    semantics being "If a macro invocation unreachable() is reached during
    execution, the behavior is undefined" and "The program execution shall
    not reach such an invocation". The justification is for better
    diagnostics and optimisations.

    How can "undefined behaviour" lead to better diagnostics given the
    attitude that "it is your responsibility as a C programmer to make
    sure that never happens"? Whoever wrote this justfication apparently
    has a different idea of the meaning of undefined behaviour than you
    do.


    My support for the existence of undefined behaviour in C is first for diagnostic purposes, and secondly for optimisation - there is no
    conflict here.

    "UB" means there are no requirements or restrictions about what happens,
    as far as the C standards are concerned. That means that a compiler -
    if it is appropriate, practical and useful - can give all the options it
    wants for what happens if the UB is hit at run-time. The optimisation
    option is to assume it never happens, and follow the logical
    consequences of that assumption. Diagnostic options include putting in
    trap instructions, generating signals, writing log messages to stderr, terminating the program, causing a breakpoint when using gdb, calling a
    user function, or any other combinations that a compiler might like to support. Not all options are equally possible on all targets, and some
    may have significant performance implementations for certain UB cases.
    Some might be poorly specified (such as "a = 0x7fff'ffff; b = a + 1 - 2"
    which is an overflow if handled na|>vely, but avoids an overflow if the
    "1 - 2" is optimised). Any attempt at giving a fixed definition in the standards - either as fully-define or implementation-defined - would necessarily restrict such diagnostics.

    Anyway, undefined behaviour in connection with "unreachable()" or
    "restrict" is not a problem as far as I am concerned. A programmer
    can just choose to never use "unreachable()" or "restrict", and avoid
    any undefined behaviours coming from these language features. And
    when he introduces them, my recommendation is to do it sparingly, only
    in those places that are relevant to performance.

    My recommendation is to use them exactly as often as it makes sense.
    But as with any other specification of code, it must be clear to the
    user. So put "restrict" in function prototypes when two pointers of the
    same type should not be aliases - that tells the caller something
    important about how to use the function. The optimisation is often just
    a bonus (since most code is not performance-critical). Similarly, "unreachable()" can make code clearer, and avoid spurious returns that
    are never executed but exist to avoid compiler warnings. Unless you are
    being silly (like writing "return; unreachable();") adding these
    constructs is likely to improve code clarity without affecting the risk
    of errors.


    This is in contrast to taking the assumption that undefined behaviour
    does not happen anywhere, which means that programmers have to check everywhere for the >200 explicitly stated (and who knows how many
    implied) undefined behaviours in the current C standard.


    The number of implied undefined behaviours is unlimited - if the C
    standards don't specify the behaviour of something, it is UB exactly as
    though it had been explicitly labelled UB in the standards. In the
    style of "nasal daemons", spilling coffee on your keyboard is undefined behaviour in C.

    Almost all cases of UB listed explicitly in the C standards are clearly
    bad code. If you have an array "xs[100]", then it is clearly wrong to
    try to access "xs[-1]" or "xs[100]". Using a pointer after the end of
    the lifetime of the thing it points at, is clearly wrong. Obviously
    people make mistakes in their code - that's just bugs in the code.

    There are a few situations where something is UB, but programmers may mistakenly think the behaviour is defined. C programmers need to learn
    these cases (or learn that they must look up the details if they need to
    use the relevant constructs). Signed integer overflow, certain shift
    values, and accessing data via pointers to a different type of object (violating the "strict aliasing" rules) are the main candidates. There
    is also the more subtle case of constructing or comparing pointers that
    are not within the same object.

    It is not /that/ difficult to write C code without unintentional UB -
    beyond what would be bugs in the code even if the behaviour had been
    defined in some way. (Intentional UB in knowingly non-portable code has
    been covered earlier.)


    I don't know whether or not the "founding fathers" of C intended or
    expected compilers to optimise on the assumption that UB did not occur.

    One can look at the code written by Ritchie, Thompson and the other
    people at Bell Labs. Does it contain undefined behaviour? Very
    likely.

    But I am confident that they considered a program to be broken if
    execution reached a point where the behaviour was not defined in an
    implementation (something may be UB in the C standards yet defined by an
    implementation).

    If you mean "documented as defined", I doubt it. I expect that they
    did write programs that rely on the actual code generated by the
    compiler(s) the used for the program, on the target(s) they used for
    the program. And of course the C compilers before C89 did not
    document behaviours as defined that would be undefined by the standard
    only later.

    Intentional UB in non-portable code is a different matter.


    It is, of course, possible that they did not foresee quite how this
    would pan out in modern compilers. In particular, they may not have
    predicted "time-travel" optimisations. However, authors of more modern
    C standards - say, C11 onwards - know about them and have could have
    explicitly outlawed them if they were considered to be invalid.

    There is no need to outlaw time travel in C standards, because it was
    never allowed. At least that's what I read here some years ago. Time traveling is a C++, not C property.


    To my knowledge, it is valid in C and C++. C++26 adds the concept of "erroneous behaviour" that allows a lot of possible behaviours, but it
    does not allow time-travel.


    For
    heavily used compilers, there is a strong (but not overpowering) push
    towards backwards compatibility. This is why they have heavy regression
    tests, and pre-release versions are tested with large samples of
    important existing code. If this gives unexpected results, these must
    be dealt with - was it a bug in the new compiler (in which case the fix
    is obvious), or was it a bug in the old source code? Those cases are
    more complicated - sometimes the old code must be fixed, but sometimes
    the incorrect source code is too common, idomatic or important and the
    compiler must, in effect, support additional semantics to retain the old
    accidental semantics.

    Yes, my impression is that the pushback after miscompiling "relevant
    code" has been strong enough that the regression tests now avoid
    miscompiling that code, and a lot of the "irrelevant" code such as
    Gforth rides in the slipstream of that.

    This is not about miscompiling code - it is about old incorrect code
    that relies on undocumented behaviours and unwarranted assumptions, but
    for which it would be too costly (in many ways) to fix. Modern
    compilers expose flaws in old code - the fault lies with the old code,
    not the compilers.

    Of course, relevant code like
    the Linux kernel uses a collection of flags (e.g,
    -fno-strict-overflow) that define otherwise undefined behaviour, so
    anybody who wants to ride in the slipstream of that code should use
    these flags as well, and leave the benchmarking settings (i.e.,
    without these flags) to the benchmarks.

    That is a perfectly valid solution. This is back to the intentional use
    of certain UB in non-portable code. Once particular behaviours are
    specified and documented by the compiler, such as using "-fwrapv" or "-fno-strict-aliasing", you can view the code as being written in an
    extended C that has specified semantics beyond those in the C standards.
    That's fine, it's a well-established way to use C. But it is not
    about stopping the compiler from "miscompiling" - it's about clearly
    defining an extended C dialect.

    And the rest of us, who are writing C code that does not need or use
    these semantics (and are not writing benchmarks) do not use those flags.

    Note also that in some cases when Linus Torvalds has thrown his toys out
    the pram over improvements to gcc, the problematic code was clearly
    wrong and remains wrong despite additional semantic flags in gcc - all
    the flags do is reduce the consequences of the incorrect code. (That's
    the case for the "use the pointer and then test if the pointer is valid"
    bugs. In C, you must look before you leap - not jump off the cliff and
    try to pick up the pieces at the bottom.)


    Still, even if they did not, there are still differences between
    hardware and between existing compilers to reconcile (although a lot
    of the old hardware variations have died out), and getting consensus
    on a completely specified C is unlikely. E.g., Pascal Cuoq, Matthew
    Flatt, and John Regehr tried to create a more completely specified
    "friendly C", and did not find consensus:
    <https://blog.regehr.org/archives/1287>.

    There are two key problems with UB that stand in the way of this kind of
    initiative. One is that many compilers provide consistent (and
    sometimes even documented) behaviour for some things that are UB in the
    C standards. Code relying on these behaviours is valid and safe, but
    non-portable. But different compilers (or the same compiler but
    different targets) could easily have different semantics, making it very
    difficult to agree on any one choice. Secondly, many programmers write
    code on the assumption that certain UB has certain consistent and
    reliable behaviour - even though it is not documented anywhere. It is
    code like that which could benefit from a "friendly C" (a poor choice of
    name, IMHO, but that's entirely subjective) variant. But it is also
    such code that makes a "friendly C" variant hard to define - such code
    is hard to identify, and the expected behaviour can be even harder to
    find, specify, and consistently describe.

    Such efforts try to achieve more than what is discussed in the part of
    the C99 rationale cited above, and more than I argue for in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>; it also
    tries to solve portability between architectures and compilers. And
    given that such efforts have not come to fruition, one can conclude
    that they try to achieve too much.


    That may be the case. Efforts to reduce errors, or reduce the
    consequences of errors, or increase the likelihood of errors being
    spotted early, are always a good thing. But they are often unrealistic.
    Attempts to define "what the programmer intends", beyond what the
    language defines, are doomed to fail. Likewise there is nothing to be
    gained by appeals to what compilers used to do, or what someone thinks
    someone meant when they wrote something decades ago.

    And of course any hope of progress is dependent on the cooperation of
    the C standards committee and the major C compiler developers. When you
    start off by alienating, antagonising and accusing them, anything you
    write - no matter how well considered and well intended - will likely be ignored. Attitude is as important as the technical details here. Linus Torvalds can get away with more because of the importance of Linux -
    issues with Gforth on the latest gcc are not quite as world-shattering.
    (I know I am not always as polite or careful in my wording, but I am
    just posting on Usenet - I am not trying to change the world. I don't
    need to be as diplomatic!)


    Alternatively, one can see the various flags provided by gcc, such as -fno-strict-aliasing as achieving at least a part of such efforts (not
    the portability between compilers, in general, but at least between architectures). There is at least one case (-fwrapv vs. -ftrapv)
    where one can ask gcc for one of several definined behaviours, though.


    Those are specific cases where well-defined semantics are possible, at
    least on the targets supported by gcc. And I'm fine with that - and I
    am confident that the founding fathers of C intended such usage. It
    would be nice if they could be standardised in some way as options in C,
    but I would not want them to become fixed fully-defined behaviour. I
    want the choice - -fwrapv (or -fsanitize-trap=signed-integer-overflow),
    or "-fwrapv", or "normal" optimising overflows.

    The biggest problem with any solution, however, is getting people to use
    them. After all, wrapping for signed integers has always been
    achievable safely and correctly by a few casts between signed and
    unsigned types - yet some people fail to do so correctly. Almost any
    use-case of "-fno-strict-aliasing" can be handled by use of char
    pointers or memcpy/memmove, usually at no extra run-time cost with
    modern tools. (But it is ugly, and can be very slow on weaker compilers.)

    But fortunately a completely specified C is not necessary, a
    willingness to preserve the behaviour of existing working programs
    compiled with an earlier version of the same compiler is. Read more
    about it in <https://www.complang.tuwien.ac.at/papers/ertl17kps.pdf>.


    I think it is entirely reasonable to keep old compiler versions around
    and use them for old code.

    Unfortunately, the various Linux distributions do not agree with you,
    and do not distribute gcc versions back to 1.0. So, for software
    distributed as source code, asking for a known-good compiler for
    IA-32, such as gcc-2.95, is impractical.


    Sure. It is a practical solution for some people, but not for others -
    in particular, it is not practical for distributing programs as source code.

    Sometimes, however, it /is/ practical. I keep all my old toolchains,
    and do not change toolchain for an established project without extremely
    good reason. But my customers are interested in the binaries, not
    source code. (And those that want the source get the toolchain too.)

    I think it is entirely unreasonable to try to specify that new compilers
    should have defined specifications to implement the "semantics" that
    older compilers used for particular types of undefined behaviour. These
    semantics will, for the most part, be poorly defined and can often be
    inconsistent - many programs with UB work by luck, not design.

    On the contrary, the most common cases of undefined behaviour become well-defined. E.g., before gcc ever generated code based on the
    assumption that signed overflow never happens, it generated addu or equivalent (e.g., addiu or the implied addition in the address
    computation of a load) for an addition in the C source code; and optimizations preserved this behaviour, e.g., a+(-b) was compiled to
    subu, not sub. Continuing to generate code that behaves like that is well-defined, and it now even has flags in gcc (-fwrapv and
    -fwrapv-pointer, which can be combined into -fno-strict-overflow).


    Code that assumes signed overflow is wrapping usually works by luck, not design. If the code comes with documentation (or, better, compile-time checks) for particular compilers that are known to wrap then it works by design. If it has a #pragma GCC optimize("-fwrapv") line, it works by
    design (that's the best choice, IMHO). C has never guaranteed the
    behaviour of signed integer overflow - ergo, if the code is general C
    code rather than for a specific implementation that is known to wrap, it
    only works by luck.

    Linux commits to preserving user-space behaviour (whether standard or
    not), everything I have heard from people claiming to speak for gcc
    and clang maintainers has been in the opposite direction. So I blame
    the gcc and clang maintainers for the undefined-behaviour shenanigans
    they perform.


    Linux commits to preserving the /defined/ behaviour of user-space APIs.

    If you mean that it takes the same work-to-rule approach that the gcc
    people do when it comes to "irrelevant" code, that's not the case. It commits to preserving the actual behaviour. That is well publicised,
    e.g. <https://linuxreviews.org/WE_DO_NOT_BREAK_USERSPACE>. Linus
    Torvalds wrote:

    |If a change results in user programs breaking, it's a bug in the
    |kernel. We never EVER blame the user programs.

    But if a
    particular invalid value of the parameter lets you gain access to
    root-owned files, due to a bug in the kernel, you can be very sure that
    this user-space behaviour will /not/ be preserved.

    I don't know if the maintainers of a user program ever reported a bug
    for closing such a security hole and insisted on preserving the
    behaviour; I doubt it. If they would, one way to deal with that would
    be to have a kernel option (compile-time, startup-time, or run-time)
    for opening the hole, which the default being that the hole is closed.

    There has been at least one well-known case where the kernel people
    went to great lengths to preserve a behaviour for a certain userspace
    program where the other behaviour already also had users (and in this
    case the kernel-option variant was not satisfactory, so they used
    something far uglier).

    This has all been well publicised, which makes me wonder why you are spreading the claims above; either you do not know what you are
    writing about, or you knowingly spread an untruth.


    It is strange to hear you suggesting I might be lying here. I certainly
    can't think of the particular case you describe so vaguely here. (It
    might be something I have heard of, and then forgotten.)

    There is no doubt whatsoever that fixing security flaws in the Linux
    kernel /does/ break userspace - it breaks the malware or hacks that
    exist to exploit the flaws. That's the whole point of security fixes.
    It does not matter how many capital letters Torvalds uses when telling
    people not to break userspace. Of course, breaking userspace to stop
    exploits is a good thing!

    I do not know how often the behaviour of Linux API's has changed in how
    it treats things outside of the documented specifications, but it most certainly happens. No one gives guarantees outside the specifications - that's why you have specifications, so that both sides have something
    they agree on. A user-space program used to be able to de-reference a
    pointer with the address value 0 - now it is trapped as a fault. That's
    fine to do, because it was never specified behaviour. Even specified
    and documented behaviour is not sacrosanct - interfaces, ABIs, options,
    etc., get deprecated and removed on a regular basis.

    "We do not break userspace" is not a law - it is an ambition, a
    guideline, and a major consideration for any decisions. That is how it
    should be. Userspace changes are not made unless there is overwhelming
    reason to do so - but /if/ there is overwhelming reason, then the
    changes are made. Security is the typical reason, but also old,
    unmaintained and rarely used code is removed.

    It's the same attitude for C standards version - compatibility with
    existing code is of vital importance, and changes that affect existing
    code are rare, slow and considered. (It took over 30 years - from C89
    to C23 - to finally remove "K&R" style function definitions.)


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Fri Sep 25 14:45:13 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:
    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    David Brown <david.brown@hesbynett.no> writes:

    I have never spoken to the C standards committee or writers, either of >>>current standard versions or the original C89 standard, or writers of >>>pre-standard C specifications. So I cannot in any way claim to know >>>their thoughts or motivations. I also have not seen any documentation >>>that suggests what they might have thought about "optimisation on the >>>assumption that undefined behaviour does not occur" - in either >>>direction.

    When C89 was developed, the dominating C compiler in the Unix world
    was PCC. GCC came on the scene late in the game (1988), without such >>assumptions and produced a significant performance advantage.

    Well, yes. PCC at that point was 14 years old and showing its age.

    Motorola used PCC for the 88100 processor (I had to fix a bug in the
    register allocator when compiling the output of cfront, which
    extensively uses the comma operator in 1990); I had first used PCC in
    1981.

    I'll add that by 1989, AT&T/USL had a much more sophisticated C compiler
    for SVR4 and later Unixware.

    $ ls /usr/src/common/cmd/sgs
    acomp amigo ccsdemos cof2elf cpluspatch dis inc lex libsgs m4 optim strip yacc
    acpp ar cg cplusfe cpp dump intemu libelf lorder mcs sgs.install tsort
    alint as cmd cplusinc demangler fpemu ld libld lprof nm size unix_conv
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 25 17:12:10 2026
    From Newsgroup: comp.arch

    On 25/09/2026 16:27, Stephen Fuld wrote:
    On 9/24/2026 9:29 AM, Tim Rentsch wrote:
    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/23/2026 7:47 AM, Tim Rentsch wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/22/2026 5:32 AM, David Brown wrote:

    [...]

    And note that in C23 the macro "unreachable()" was added with the
    sole semantics being "If a macro invocation unreachable() is
    reached during execution, the behavior is undefined" and "The
    program execution shall not reach such an invocation".

    I don't have a problem with the inclusion of "unreachable()", but
    can you give a possible rationale for making its behavior
    "undefined" as opposed to say "implementation defined"?-a Yes, it
    would make the compiler implementers do some work to document
    whatever they decided to do, but are there any other reasons?

    The short answer is that being implementation-defined is too
    limiting.-a A construct having implementation-defined behavior
    cannot do "just anything";-a the C standard must specify what
    behavior choices are possible.-a Any fixed set of choices might
    rule out what an implementation would like to do.-a So the only
    way to allow implementations to do whatever they choose is to
    have the behavior be undefined.

    As I responded to David, you are right, given the way those terms are
    defined in the standard.-a But I submit there should be some way to
    express essentially "you can do whatever you want, but must document
    what you do".

    Let's call the new behavior type "implementation dependent".

    I am not hung up on the name, so at least for purposes of discussion, OK.


    Suppose an implementation gives documentation that says "If a program
    execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."

    Question: are you okay with that?

    If that is the implementation's general response, then, while it might
    be "legal", then no, I am not happy with it.


    If not, how would you write a
    requirement in the C standard to limit the meaning of "you can do
    whatever you want, but you must document what you do" that gives only
    as much freedom as you think should be allowed?

    I have minimal experience with standards writing, and none with language writing, so I may be off base here.

    I think that there are many situations where the standard, correctly, doesn't specify anything about what the compiler should do.-a These are currently called undefined behavior, and the standard essentially lets
    it go at that - no further documentation required.

    But in many, but not all, of such cases, the compiler implementer knows
    what it is going to do.-a What I am after is that in such cases, the implementer "supplement" the bare words with information that he has
    that knowledge and he should communicate it to the user.


    I think we will find that different people have very different ideas
    about how much freedom should be allowed in saying what happens.-a It's
    a very hard problem.

    Note that I am not trying to restrict "what happens" in any way.-a What happens is anything that could/would happen under today's definition of undefined behavior.-a I am just trying to ask the implementer, whenever possible (and I realize that it is not always possible) to elaborate on
    the bare words "undefined behavior".



    There is nothing - except time, resources and motivation - preventing C implementations documenting what they do in cases where the C standards
    say something is UB. Usually the do that as part of specific flags that affect the behaviour - gcc documents what it does on signed integer
    overflow if you have the "-fwrapv" flag enabled, for example. There are
    a few other situations where gcc has flags that give specific semantics
    to classes of UB.

    But I don't see any point in trying to give a description of "what
    happens in UB" in cases where you are not picking specific semantics.
    If you want wrapping signed overflow semantics, the documentation is
    under "-fwrapv". If you want to halt with a run-time diagnostic, the documentation is there in "-fsantize=signed-integer-overflow". If you
    want "do whatever you like, such as optimise" semantics, there is
    obviously no need for documentation.

    Remember also that once you document something as particular behaviour,
    it is fixed - there needs to be very strong justification for changing
    it later. Documentation /does/ restrict what happens.

    (You could file a bugzilla report asking for a specific page in the
    manual that covers this sort of thing, rather than having it scattered
    around the option pages. That would be handy.)


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Fri Sep 25 11:16:20 2026
    From Newsgroup: comp.arch

    Tim Rentsch [2026-09-24 09:29:13] wrote:
    Suppose an implementation gives documentation that says "If a program execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."
    Question: are you okay with that?

    I'm not. It's not clear what it is that I want, but I know that I want
    my undefined behaviors to be more constrained in what they can do.

    I'll call "SMUB" my desired version of undefined behavior. It would go somewhere along the lines of:

    In case of a SMUB operation whose result type is T, then the
    operation is allowed to return any bit pattern compatible with the
    type T, or it can stop the execution of the program with an error.
    In case it stops the execution with an the error, that it is
    acceptable if the error doesn't occur exactly at the time the
    operation would have taken place (i.e it can occur at any later
    time, but also at an earlier time provided we already know that the
    SMUB operation would take place anyway).

    A SMUB operation that is normally pure is not allowed to mutate
    anything. And a SMUB operation that normally mutates one memory
    cell of type T is allowed to mutate any *one* memory cell of same
    size as type T by putting into it a bit pattern compatible with the
    type T.

    In the case of `unreachable()`, that means the only thing it would be
    allowed to do is to signal an error (at any point of the execution once
    we know that `unreachable()` would be executed) or do nothing. IOW, if
    it's executed and does not signal an error, the rest of the code may not
    be optimized under the assumption that it was not (nor will not
    be) executed.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Fri Sep 25 19:00:26 2026
    From Newsgroup: comp.arch

    On 25/09/2026 17:16, Stefan Monnier wrote:
    Tim Rentsch [2026-09-24 09:29:13] wrote:
    Suppose an implementation gives documentation that says "If a program
    execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."
    Question: are you okay with that?

    I'm not. It's not clear what it is that I want, but I know that I want
    my undefined behaviors to be more constrained in what they can do.

    I'll call "SMUB" my desired version of undefined behavior. It would go somewhere along the lines of:

    In case of a SMUB operation whose result type is T, then the
    operation is allowed to return any bit pattern compatible with the
    type T, or it can stop the execution of the program with an error.
    In case it stops the execution with an the error, that it is
    acceptable if the error doesn't occur exactly at the time the
    operation would have taken place (i.e it can occur at any later
    time, but also at an earlier time provided we already know that the
    SMUB operation would take place anyway).

    A SMUB operation that is normally pure is not allowed to mutate
    anything. And a SMUB operation that normally mutates one memory
    cell of type T is allowed to mutate any *one* memory cell of same
    size as type T by putting into it a bit pattern compatible with the
    type T.

    That sounds somewhat like "erroneous behaviour" in C++.

    <https://cppreference.com/cpp/language/ub>

    It also sounds like what C has for conversion of an out-of-range value
    to a signed integer type - "either the result is implementation-defined
    or an implementation-defined signal is raised."

    That seems fine for certain things, as a possible choice by an
    implementation - it keeps the diagnostic option for the behaviour. But
    it limits optimisation. I have no issues with you wanting that, but I
    don't want /my/ compilations to be limited by it.

    And in fits what you have today in some cases, such as signed integer
    overflow in gcc - you can choose "-fwrapv", or you can choose "-fsanitize=signed-integer-overflow", or you can choose optimisation on
    the assumption that the code is correct. That's the best of all worlds
    as I see it, and forcing choices in the standards would limit that.


    In the case of `unreachable()`, that means the only thing it would be
    allowed to do is to signal an error (at any point of the execution once
    we know that `unreachable()` would be executed) or do nothing. IOW, if
    it's executed and does not signal an error, the rest of the code may not
    be optimized under the assumption that it was not (nor will not
    be) executed.

    You would only put "unreachable()" in your code at a spot where you do
    not expect execution to reach. So if it has got there, your code is
    broken - you have no control of what it is doing.

    If I write :

    // classify must only be called with x between 0 and 100
    int classify(int x) {
    if (x < 10) return 0;
    if (x < 30) return 1;
    if (x <= 100) return 2;
    unreachable();
    }

    and then use it as:

    int vs[3];

    // do_something must only be called with x between 0 and 100
    void do_something(int x, int y) {
    vs[classify(x)] = y;
    }

    what is the compiler supposed to do with the "unreachable()" if you are
    not halting with an error? Is it supposed to guess that it must return
    a value between 0 and 2 inclusive? Should it return an unspecified
    value? There's a good chance that with a "do nothing" optimisation, it
    would return an unchanged "x" - leading to a stomp of some random place
    in memory.

    There's no way to get sensible semantics here - you either say terminate immediately to limit damage, or accept what UB means. Trying to say you things can go a bit bad, but not too bad, never works. You either have
    to be very limiting (like "evaluate to an unspecified value") and add
    extra manual checks in your code for validity, or accept UB. Other
    options are too vague, and there is always scope for the bad behaviour
    to amplify out of any hope of control.


    If I wanted to terminate immediately, I'd write :

    fprintf(stderr, "The programmer has screwed up\n");
    exit(1);

    We can do that already. I don't want "unreachable()" to duplicate that,
    I want it to mean something else.

    By all means make a macro that lets you choose between such actions
    depending on whether you are doing fault-finding or want highly
    optimised code.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Fri Sep 25 18:12:17 2026
    From Newsgroup: comp.arch

    David Brown <david.brown@hesbynett.no> writes:
    On 25/09/2026 17:16, Stefan Monnier wrote:
    Tim Rentsch [2026-09-24 09:29:13] wrote:
    Suppose an implementation gives documentation that says "If a program
    execution would encounter implementation-dependent behavior, then
    program execution may be affected in ways that are unpredictable,
    unknown, and/or unreliable."
    Question: are you okay with that?

    I'm not. It's not clear what it is that I want, but I know that I want
    my undefined behaviors to be more constrained in what they can do.

    I'll call "SMUB" my desired version of undefined behavior. It would go
    somewhere along the lines of:

    In case of a SMUB operation whose result type is T, then the
    operation is allowed to return any bit pattern compatible with the
    type T, or it can stop the execution of the program with an error.
    In case it stops the execution with an the error, that it is
    acceptable if the error doesn't occur exactly at the time the
    operation would have taken place (i.e it can occur at any later
    time, but also at an earlier time provided we already know that the
    SMUB operation would take place anyway).

    A SMUB operation that is normally pure is not allowed to mutate
    anything. And a SMUB operation that normally mutates one memory
    cell of type T is allowed to mutate any *one* memory cell of same
    size as type T by putting into it a bit pattern compatible with the
    type T.

    That sounds somewhat like "erroneous behaviour" in C++.

    <https://cppreference.com/cpp/language/ub>

    It also sounds like what C has for conversion of an out-of-range value
    to a signed integer type - "either the result is implementation-defined
    or an implementation-defined signal is raised."

    That seems fine for certain things, as a possible choice by an >implementation - it keeps the diagnostic option for the behaviour. But
    it limits optimisation. I have no issues with you wanting that, but I
    don't want /my/ compilations to be limited by it.

    And in fits what you have today in some cases, such as signed integer >overflow in gcc - you can choose "-fwrapv", or you can choose >"-fsanitize=signed-integer-overflow", or you can choose optimisation on
    the assumption that the code is correct. That's the best of all worlds
    as I see it, and forcing choices in the standards would limit that.


    In the case of `unreachable()`, that means the only thing it would be
    allowed to do is to signal an error (at any point of the execution once
    we know that `unreachable()` would be executed) or do nothing. IOW, if
    it's executed and does not signal an error, the rest of the code may not
    be optimized under the assumption that it was not (nor will not
    be) executed.

    You would only put "unreachable()" in your code at a spot where you do
    not expect execution to reach. So if it has got there, your code is
    broken - you have no control of what it is doing.

    If I write :

    // classify must only be called with x between 0 and 100
    int classify(int x) {
    if (x < 10) return 0;
    if (x < 30) return 1;
    if (x <= 100) return 2;
    unreachable();

    While I dislike 'assert' immensely (it is rude IMO),
    it seems far better to just assert(x <= 100) upon function entry
    rather than relying on whatever the compiler might do with
    an 'unreachable' declaration.

    C++26 style contracts may find their way into C someday....

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 26 07:44:58 2026
    From Newsgroup: comp.arch

    EricP <ThatWouldBeTelling@thevillage.com> writes:
    On 2026-Sep-22 01:25, Anton Ertl wrote:
    Effect of new optimizations or "optimizations" based on assuming that
    undefined behaviour never happens on performance? They never say.
    That's the cool thing. They do not have numbers (they certainly never
    present any, certainly not for their own compilers), but are convinced
    that their "optimizations" do wonders for performance, and their
    fanboys are even more convinced.

    It seems someone measured UB optimization performance impact in LLVM:

    Exploiting Undefined Behavior in C C++ Programs for Optimization
    A Study on the Performance Impact, 2025 >https://dl.acm.org/doi/abs/10.1145/3729260 >https://dl.acm.org/doi/pdf/10.1145/3729260

    They did it by using 6 pre-existing flags in clang, and adding 12
    additional ones that defined behaviour that C and C++ do not define.

    "Using LLVM, a compiler known for its extensive use of UB for optimizations, >we demonstrate that, for the benchmarks and UB categories that we evaluated, >the end-to-end performance gains are minimal."

    An earlier work with a similar result is:

    @InProceedings{wang+12,
    author = {Xi Wang and Haogang Chen and Alvin Cheung and Zhihao Jia and Nickolai Zeldovich and M. Frans Kaashoek},
    title = {Undefined Behavior: What Happened to My Code?},
    booktitle = {Asia-Pacific Workshop on Systems (APSYS'12)},
    OPTpages = {},
    year = {2012},
    url1 = {http://homes.cs.washington.edu/~akcheung/getFile.php?file=apsys12.pdf},
    url2 = {http://people.csail.mit.edu/nickolai/papers/wang-undef-2012-08-21.pdf},
    OPTannote = {}
    }

    In particular, Wang et al. found that in all of the SPEC 2006 (IIRC)
    CPU benchmarks written in C, there are only two locations where the "optimizations" they looked at made a difference that's beyond the
    noise level. These performance benefits could also be had by two
    small changes in the source code (which does not happen for
    benchmarks, but it happens for production code), which is much cheaper
    than "sanitizing" all of the code (which does not happen for
    benchmarks, either; instead, the compiler is tuned to avoid
    miscompiling the benchmark; but for production code the fans of
    "optimization" want us to do it).

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 26 09:18:57 2026
    From Newsgroup: comp.arch

    On 2026-09-24, Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    A straightforward patch would
    very likely pessimize a lot of existing code which profits from >>>>auto-vectorization.

    How can we test this claim?

    I do not believe that I have to explain the scientific method
    to you.

    As a first step, you would find an option that enables/disables
    what you don't like. -fstore-merging looks like a
    candidate, but there may be others. You can look at >>https://dl.acm.org/doi/10.1109/ASE56229.2023.00209 or >>https://link.springer.com/article/10.1007/s10515-024-00437-w
    if you want the full package.

    Then try this combination of options on other benchmarks
    as well as your pet one. I have my reservations about SPEC,
    but as you work at a university, you can get it for a discount.
    You can also use freely available benchmarks: Coremark, embench,
    Fortran Polyhedron - there are a lot.

    If there are benchmark which regress significantly with that
    set of options, you have your test case where it hurts.

    Like I wrote previously - you could have a bachelor or master
    student do this, it could be (part of a) nice thesis.

    That's a lot of words to express that you have no evidence for your
    claim.

    I was just trying to get you to do something useful. Wasted
    effort, I know, but I keep trying.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Sat Sep 26 10:25:39 2026
    From Newsgroup: comp.arch

    David Brown [2026-09-25 19:00:26] wrote:
    That seems fine for certain things, as a possible choice by an
    implementation - it keeps the diagnostic option for the behaviour. But it limits optimisation.

    Limiting optimization is the whole purpose of the exercise!

    Remember, we're starting from the problem that UB is used by
    compilers in ways which catch programmers off-guard. My SMUB proposal
    is an attempt to define something which I hope is closer to what
    programmers expect.

    Given that UB has been shown to provide only fairly minor performance
    gains, I'm hopeful that "downgrading" UB to SMUB would be good enough in
    the vast majority of cases. You could still add a `--SMUB-is-UB` flag
    if you're so inclined, of course, just like the `--fast-math`.

    what is the compiler supposed to do with the "unreachable()" if you are not halting with an error?

    At runtime, I think a self-respecting compiler would halt with an error,
    yes. That's the only behavior that won't catch programmers by surprise.

    There's no way to get sensible semantics here - you either say terminate immediately to limit damage, or accept what UB means. Trying to say you things can go a bit bad, but not too bad, never works.

    Note that in your example, even if `unreachable()` turns into a nop, the potential damage still corresponds to executing the code which the users
    wrote, in the order they wrote it, performing the tests they wrote.
    So it still matches better the naive semantics most programmers have in
    their head than some of the very counter-intuitive runtime behaviors
    you may get with UB's current optimizations.

    We can do that already. I don't want "unreachable()" to duplicate
    that, I want it to mean something else.

    I know. You like to take advantage of UB to try and gain the last few
    percents of optimizations. I put less emphasis on that.

    A halfway point might be for the compiler to emit a warning when it
    can't show that `unreachable()` is indeed unreachable.

    [ I generally like the idea of programmers being able to request
    specific optimizations and to be warned by the compiler when those
    requests can't be satisfied. So we don't have to go dig in the
    assembly output to confirm whether the compiler did the right thing or
    not. ]


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Sun Sep 27 16:40:15 2026
    From Newsgroup: comp.arch

    On 26/09/2026 16:25, Stefan Monnier wrote:
    David Brown [2026-09-25 19:00:26] wrote:
    That seems fine for certain things, as a possible choice by an
    implementation - it keeps the diagnostic option for the behaviour. But it >> limits optimisation.

    Limiting optimization is the whole purpose of the exercise!


    Limiting optimisations should /never/ be an aim in itself. An
    optimisation is a transformation of the code (source code, internal representations, object code) that results in faster and/or smaller
    object code without affecting the semantics. Baring code whose
    correctness depends on size or speed (which is certainly a valid
    possibility in real world programming) since that is outside the scope
    of definition for C and most other programming languages, an
    optimisation cannot make a correct program incorrect, or change its
    behaviour in any meaningful way - if a compiler pass or flag does make a meaningful change, it is not an optimisation.

    The critical thing to understanding optimisation is to understand the semantics of the code - what the code actually means, according to the specification of the language plus any additional semantics provided by
    the implementation documentation. Where people usually get in trouble
    is when they /think/ their code has defined semantics that it does not.

    An optimisation does not "break" correct code. Code that is correct
    before the optimisation pass will remain correct afterwards - code that
    is incorrect after optimisation was incorrect before the optimisation.

    But while optimisations do not introduce problems in correct code, they certainly can amplify the /consequences/ of incorrect code. Sometimes,
    by luck, the pre-optimisation consequences of a particular piece of
    incorrect code - a particular case of UB at run-time - are negligible or non-existent. Sometimes, by luck, that same UB has serious consequences
    after optimisation.

    And since people write incorrect code by mistake, or perhaps knowingly
    but without appropriate precautions (like pre-processor checks for
    particular compiler versions, or documented specific flags in the build files), this can all lead to unpleasant surprises.


    Remember, we're starting from the problem that UB is used by
    compilers in ways which catch programmers off-guard.

    If that is truly the way you think, then you should maybe give up
    programming and start a conspiracy-theory pod-cast. (You'll probably
    make more money!) Seriously - if you start from the suggestion that
    compiler writers are aiming to catch programmers off-guard, or that they disregard the needs of real programmers so that they can get better
    benchmark results, you will always be fighting with your tools and tool vendors, instead of cooperating. Toolchain writers are fallible humans,
    like the rest of us, and they can make mistakes, make poor decisions,
    have different priorities that don't match ours, and have bosses and
    employers that have different ideas about what is important. But don't imagine that anyone is in the toolchain business (or language
    specification business) with an aim of making things difficult for
    others, or of showing off how great their tools are in some test that
    does not match real usage.

    My SMUB proposal
    is an attempt to define something which I hope is closer to what
    programmers expect.

    There are a million C programmers. There are a million and one
    different things that they expect, or want, or hope for. People who
    think they know what programmers expect are, statistically speaking, wrong.

    Of course you are entirely entitled to your own opinions about what
    /you/ expect. And if you make proposals like this and there enough
    people who think along the same lines, maybe it will lead to changes or
    new features in the language and toolchains. The big challenge is
    getting a workable definition of what this "SMUB" should be - that would
    be a long and hard process.


    Given that UB has been shown to provide only fairly minor performance
    gains, I'm hopeful that "downgrading" UB to SMUB would be good enough in
    the vast majority of cases. You could still add a `--SMUB-is-UB` flag
    if you're so inclined, of course, just like the `--fast-math`.

    First, you already have much of this as a flag in gcc - it's called
    "enabling optimisation". If you blind "literal" code generation, that's
    what you get by default in gcc. You still don't get guarantees - there
    is no black-or-white choices here between "optimised" and "not optimised".

    Secondly, you are not going to get anywhere when you try to talk about "optimisation from UB" - that's too vague to be of any use. You can
    talk about very specific classes of operation that has some UB today,
    but not UB in general. As it stands, your suggestion would mean that
    unless the "--SMUB-is-UB" flag is used, the compiler would need to
    implement this :

    int get(int * p) { return *p; }

    as
    int get(int * p) {
    if (is_valid_pointer_to_an_int_object(p)) {
    return *p;
    } else {
    return some_int;
    // or
    exit_with_error_message();
    }
    }

    Clearly, that is not going to happen. If that's what you want, use a
    managed language (or compile with a good memory sanitizer).

    Thirdly, don't assume that other programmers want any of this. Some
    will, others will not.

    Fourthly, don't rely on an old benchmark paper of some particular
    programs on one or two particular machines as though it "proves"
    anything. It is perhaps an interesting study, and can be a guide to
    where compiler writers should be concentrating their efforts, but it is
    not "proof" that "optimising from UB has only minor gains". Performance
    gains in compiled code is the combination of large numbers of
    transformations whose average improvement is tiny, as well as some
    passes that occasionally make a large difference. If assuming signed
    integer overflow does not happen saves two instructions from a 100
    instruction loop, it makes no practical difference - if it saves two instructions from a four instruction loop, then that code benefits
    greatly. And while code on an x86 might not see a difference from an optimisation since the bottleneck is memory speed and caching patterns,
    the same optimisation on a microcontroller might make a huge difference.


    what is the compiler supposed to do with the "unreachable()" if you are not >> halting with an error?

    At runtime, I think a self-respecting compiler would halt with an error,
    yes. That's the only behavior that won't catch programmers by surprise.

    That would be a huge surprise to /me/ if it happened. Again, you are extrapolating your own ideas to those of other programmers.

    "unreachable()" says, very clearly IMHO (and I'm aware that opinions are subjective) that this code will not be reached. I most certainly do not
    want the compiler to inject code that will stop the program here.
    That's simply not what the word means.

    C already has ways to say "stop the program with a message". It has a
    way to say "unless we have picked no-debugging mode, check that this expression is true, or else stop with an error" - the "assert" macro has
    been standard forever. (And some of us are against the use of that
    macro, as it stands, in their code - others can use it if they like, but
    it is too blunt an instrument for my needs.) gcc has had a "__builtin_unreachable()" feature for decades - any programmer familiar
    with that would likely be very surprised if C23's standardised
    unreachable() macro acted differently.

    There is /no/ choice of definition - or lack of definition - that is not
    going to surprise at least some programmers. There is no choice that
    will make everyone happy. The C standard committee picked the
    definition that was most likely to be useful, and least likely to be surprising, and compilers (like gcc) implement it as it is intended.


    There's no way to get sensible semantics here - you either say terminate
    immediately to limit damage, or accept what UB means. Trying to say you
    things can go a bit bad, but not too bad, never works.

    Note that in your example, even if `unreachable()` turns into a nop, the potential damage still corresponds to executing the code which the users wrote, in the order they wrote it, performing the tests they wrote.

    No, it does not. Turning "unreachable()" into a "nop" would just be an
    extra opcode in the object code, one that should never be executed. If control flow ever reaches the "unreachable()", then the programmer has
    made a critical bug in their code somewhere. /Everything/ is then a
    surprise - and most likely, the problem is long before the unreachable()
    part, as are the consequences of the mistake. "unreachable()" is not an assertion, or a checkpoint in the code - hitting it would mean you have screwed up royally, and nothing can be relied upon. It is no different
    in semantics from a comment saying "// Code flow should never get here"
    - it is just that now you have a way of saying that to the compiler and
    not just to human programmers.

    And any imagined semantics of "unreachable()" just mean that you now
    have a line in your code that does something, but it will never be
    tested and cannot ever be tested. That's a big no-no for many developers.

    Either the line with "unreachable()" can be executed, in which case you
    should not use that macro but should do something else, or it is never executed - in which case it should not have any semantics.

    So it still matches better the naive semantics most programmers have in
    their head than some of the very counter-intuitive runtime behaviors
    you may get with UB's current optimizations.

    Nope.


    We can do that already. I don't want "unreachable()" to duplicate
    that, I want it to mean something else.

    I know. You like to take advantage of UB to try and gain the last few percents of optimizations. I put less emphasis on that.


    I have already explained that I like UB for multiple reasons.
    Optimisation is just one of them.

    A halfway point might be for the compiler to emit a warning when it
    can't show that `unreachable()` is indeed unreachable.

    I can certainly agree that the compiler could give a warning if the unreachable() is definitely reached. And if compiler writers can figure
    out some decent heuristics about "unreachable()" being inevitably
    reached after some useful code path, then a warning would be nice.
    (That can't be given specifications in the standards, because you could
    not define clear semantics here - but it's fine to have optional
    compiler warnings on things that appear to be wrong according to some
    kind of pattern recognition.)

    But it would be virtually useless to have a warning any time the
    compiler cannot prove that the "unreachable()" is indeed unreachable.
    The whole point of its use is when the programmer knows more than the compiler, and is giving the compiler additional information. If the
    compiler knows that point in the code cannot be reached, then it already
    knows it is unreachable, and will already optimise on the assumption
    that the point is not reached!


    [ I generally like the idea of programmers being able to request
    specific optimizations and to be warned by the compiler when those
    requests can't be satisfied. So we don't have to go dig in the
    assembly output to confirm whether the compiler did the right thing or
    not. ]


    I agree on that principle - I'm a big fan of warnings when my code
    appears to be wrong. But this is not a situation where warnings would
    be appropriate.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 27 18:18:05 2026
    From Newsgroup: comp.arch

    On Sun, 27 Sep 2026 13:42:05 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:

    Given that UB has been shown to provide only fairly minor
    performance gains,

    This I found hard to believe, so I ran a
    few benchmarks myself. I used the well-known
    Polyhedron suite, which can now be downloaded from https://fortran.uk/fortran-compiler-comparisons/the-polyhedron-solutions-benchmark-suite/
    ,

    This is Fortran, so integer overflow is an error. I ran the
    testsuite on gfortran with and without -fwrapv, otherwise letting
    the compiler go wild (-Ofast). Because gfortran puts arrays on the
    stack with -Ofast, I had to increase the stacksize for one particular
    code. I had to disable one test, rnflow, because of a bug.

    Here are the results:

    Date & Time : 27 Sep 2026 12:13:00
    Test Name : gfortran-Ofast
    Compile Command : gfortran -march=native -mtune=native -Ofast %n.f90
    -o %n Benchmarks : ac aermod air capacita channel2 doduc
    fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein
    -rnflow test_fpu2 tfft2 Maximum Times : 10000.0 Target Error %
    : 0.200 Minimum Repeats : 10
    Maximum Repeats : 100

    Benchmark Compile Executable Ave Run Number Estim
    Name (secs) (bytes) (secs) Repeats Err %
    --------- ------- ---------- ------- ------- ------
    ac 0.65 38616 4.54 13 0.1876
    aermod 24.66 1074024 3.26 15 0.1773
    air 2.96 94920 0.89 19 0.1560
    capacita 3.97 127104 5.99 18 0.1858
    channel2 0.40 26584 40.47 10 0.1494
    doduc 4.61 164768 4.43 13 0.1230
    fatigue2 1.50 69200 30.37 11 0.1968
    gas_dyn2 1.59 74456 38.17 19 0.1610
    induct2 4.63 220280 12.26 10 0.1702
    linpk 0.53 30328 1.88 14 0.1842
    mdbx 1.87 82200 3.02 10 0.1662
    mp_prop_desi 0.75 39568 33.30 19 0.1536
    nf 0.75 34848 3.03 18 0.1680
    protein 2.32 94640 10.60 15 0.1572
    -rnflow 0.00 0 -1.00 15 0.1572
    test_fpu2 2.45 76648 16.08 16 0.1926
    tfft2 0.97 43000 13.08 16 0.1436

    Geometric Mean Execution Time = 10.57 seconds

    Date & Time : 27 Sep 2026 13:45:57
    Test Name : wrapv
    Compile Command : gfortran -fwrapv -march=native -mtune=native -Ofast
    %n.f90 -o %n Benchmarks : ac aermod air capacita channel2 doduc
    fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein
    -rnflow test_fpu2 tfft2 Maximum Times : 10000.0 Target Error %
    : 0.200 Minimum Repeats : 10
    Maximum Repeats : 100

    Benchmark Compile Executable Ave Run Number Estim
    Name (secs) (bytes) (secs) Repeats Err %
    --------- ------- ---------- ------- ------- ------
    ac 0.64 38616 4.50 10 0.1097
    aermod 27.17 1118976 3.20 10 0.1696
    air 3.40 111024 1.33 22 0.1867
    capacita 2.98 102528 6.87 16 0.1163
    channel2 0.37 26584 42.15 10 0.1187
    doduc 4.34 152376 4.58 16 0.1956
    fatigue2 1.53 69200 30.29 10 0.1952
    gas_dyn2 1.59 74456 38.47 14 0.1602
    induct2 4.13 199800 12.35 12 0.1728
    linpk 0.50 30328 2.65 12 0.1263
    mdbx 2.11 90352 3.25 10 0.0611
    mp_prop_desi 0.72 35248 33.36 10 0.0438
    nf 0.72 34848 4.19 10 0.1412
    protein 2.30 90544 10.94 10 0.1665
    -rnflow 0.00 0 -1.00 10 0.1665
    test_fpu2 2.36 76648 15.57 10 0.1457
    tfft2 0.52 26616 20.11 13 0.1603

    Geometric Mean Execution Time = 11.73 seconds

    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    So at least for that particular codebase, taking away the
    freedom of the compiler to optimize based on the assumption
    that integers will never overflow, and replacing it with
    something defined, leads to a significant slowdown.



    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default 32-bit.
    That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance
    degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.







    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Sun Sep 27 11:46:28 2026
    From Newsgroup: comp.arch

    On 2026-Sep-27 09:42, Thomas Koenig wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:

    Given that UB has been shown to provide only fairly minor performance
    gains,

    This I found hard to believe, so I ran a
    few benchmarks myself. I used the well-known
    Polyhedron suite, which can now be downloaded from https://fortran.uk/fortran-compiler-comparisons/the-polyhedron-solutions-benchmark-suite/
    ,

    This is Fortran, so integer overflow is an error. I ran the
    testsuite on gfortran with and without -fwrapv, otherwise letting
    the compiler go wild (-Ofast). Because gfortran puts arrays on the
    stack with -Ofast, I had to increase the stacksize for one particular
    code. I had to disable one test, rnflow, because of a bug.

    Here are the results:

    <snip>

    Geometric Mean Execution Time = 10.57 seconds

    Geometric Mean Execution Time = 11.73 seconds

    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    So at least for that particular codebase, taking away the
    freedom of the compiler to optimize based on the assumption
    that integers will never overflow, and replacing it with
    something defined, leads to a significant slowdown.


    One of the papers I looked at gave an example that can
    explain some of this:

    "for (int i = 0; i <= N; ++i) { ... }

    Because of the undefined behavior caused by signed overflow
    [8, UB #36], a compiler can often assume that such a loop will iterate
    exactly N + 1 times. If i were instead declared as unsigned int,
    making overflow defined, the compiler would now need to consider
    the possibility that the loop will never terminate
    (e.g., were N to be UINT_MAX)."



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Sun Sep 27 18:04:05 2026
    From Newsgroup: comp.arch

    On 27/09/2026 17:18, Michael S wrote:


    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default 32-bit.
    That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.


    You certainly /can/ do that. But a lot of C code uses "int" for array indexing. The "ideal" type in many cases would be "int_fast32_t" to say
    it should be of a big enough size (32-bit is enough range for a great
    many purposes) but can bigger if that's faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.

    The thing you want is that your loops with "a[i++]" operations are
    implemented as "*p++" rather than "*(a + i); i = (i + 1) &
    0xffff'ffff)". The later is what you end up with if you use an unsigned
    32-bit type or a signed 32-bit type with forced wrapping semantics.

    So you can either use a 64-bit type (signed or unsigned), or use normal
    32-bit "int" as usual, but without forcing wrapping semantics on it.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 27 19:43:01 2026
    From Newsgroup: comp.arch

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:


    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.


    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.


    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it
    targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.


    The thing you want is that your loops with "a[i++]" operations are implemented as "*p++" rather than "*(a + i); i = (i + 1) &
    0xffff'ffff)". The later is what you end up with if you use an
    unsigned 32-bit type or a signed 32-bit type with forced wrapping
    semantics.

    So you can either use a 64-bit type (signed or unsigned), or use
    normal 32-bit "int" as usual, but without forcing wrapping semantics
    on it.




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 27 13:42:05 2026
    From Newsgroup: comp.arch

    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:

    Given that UB has been shown to provide only fairly minor performance
    gains,

    This I found hard to believe, so I ran a
    few benchmarks myself. I used the well-known
    Polyhedron suite, which can now be downloaded from https://fortran.uk/fortran-compiler-comparisons/the-polyhedron-solutions-benchmark-suite/
    ,

    This is Fortran, so integer overflow is an error. I ran the
    testsuite on gfortran with and without -fwrapv, otherwise letting
    the compiler go wild (-Ofast). Because gfortran puts arrays on the
    stack with -Ofast, I had to increase the stacksize for one particular
    code. I had to disable one test, rnflow, because of a bug.

    Here are the results:

    Date & Time : 27 Sep 2026 12:13:00
    Test Name : gfortran-Ofast
    Compile Command : gfortran -march=native -mtune=native -Ofast %n.f90 -o %n Benchmarks : ac aermod air capacita channel2 doduc fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein -rnflow test_fpu2 tfft2
    Maximum Times : 10000.0
    Target Error % : 0.200
    Minimum Repeats : 10
    Maximum Repeats : 100

    Benchmark Compile Executable Ave Run Number Estim
    Name (secs) (bytes) (secs) Repeats Err %
    --------- ------- ---------- ------- ------- ------
    ac 0.65 38616 4.54 13 0.1876
    aermod 24.66 1074024 3.26 15 0.1773
    air 2.96 94920 0.89 19 0.1560
    capacita 3.97 127104 5.99 18 0.1858
    channel2 0.40 26584 40.47 10 0.1494
    doduc 4.61 164768 4.43 13 0.1230
    fatigue2 1.50 69200 30.37 11 0.1968
    gas_dyn2 1.59 74456 38.17 19 0.1610
    induct2 4.63 220280 12.26 10 0.1702
    linpk 0.53 30328 1.88 14 0.1842
    mdbx 1.87 82200 3.02 10 0.1662
    mp_prop_desi 0.75 39568 33.30 19 0.1536
    nf 0.75 34848 3.03 18 0.1680
    protein 2.32 94640 10.60 15 0.1572
    -rnflow 0.00 0 -1.00 15 0.1572
    test_fpu2 2.45 76648 16.08 16 0.1926
    tfft2 0.97 43000 13.08 16 0.1436

    Geometric Mean Execution Time = 10.57 seconds

    Date & Time : 27 Sep 2026 13:45:57
    Test Name : wrapv
    Compile Command : gfortran -fwrapv -march=native -mtune=native -Ofast %n.f90 -o %n
    Benchmarks : ac aermod air capacita channel2 doduc fatigue2 gas_dyn2 induct2 linpk mdbx mp_prop_design nf protein -rnflow test_fpu2 tfft2
    Maximum Times : 10000.0
    Target Error % : 0.200
    Minimum Repeats : 10
    Maximum Repeats : 100

    Benchmark Compile Executable Ave Run Number Estim
    Name (secs) (bytes) (secs) Repeats Err %
    --------- ------- ---------- ------- ------- ------
    ac 0.64 38616 4.50 10 0.1097
    aermod 27.17 1118976 3.20 10 0.1696
    air 3.40 111024 1.33 22 0.1867
    capacita 2.98 102528 6.87 16 0.1163
    channel2 0.37 26584 42.15 10 0.1187
    doduc 4.34 152376 4.58 16 0.1956
    fatigue2 1.53 69200 30.29 10 0.1952
    gas_dyn2 1.59 74456 38.47 14 0.1602
    induct2 4.13 199800 12.35 12 0.1728
    linpk 0.50 30328 2.65 12 0.1263
    mdbx 2.11 90352 3.25 10 0.0611
    mp_prop_desi 0.72 35248 33.36 10 0.0438
    nf 0.72 34848 4.19 10 0.1412
    protein 2.30 90544 10.94 10 0.1665
    -rnflow 0.00 0 -1.00 10 0.1665
    test_fpu2 2.36 76648 15.57 10 0.1457
    tfft2 0.52 26616 20.11 13 0.1603

    Geometric Mean Execution Time = 11.73 seconds

    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    So at least for that particular codebase, taking away the
    freedom of the compiler to optimize based on the assumption
    that integers will never overflow, and replacing it with
    something defined, leads to a significant slowdown.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sun Sep 27 17:39:38 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 27 Sep 2026 13:42:05 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:


    [Polyhedron benchmark, which is Fortran]

    Geometric Mean Execution Time = 11.73 seconds

    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    So at least for that particular codebase, taking away the
    freedom of the compiler to optimize based on the assumption
    that integers will never overflow, and replacing it with
    something defined, leads to a significant slowdown.



    Fortran is not the same as C.

    C and Fortran share the same rules about integer overflow.

    But if you run any program through the NAG compiler, it is
    translated to a C program.

    My understanding is that in Fortran array indexes are by default 32-bit.

    That understanding turns out to be wrong.

    In Fortran, an array index is an integer expression, and that
    can be of any KIND that the compiler supports. You can write

    use iso_fortran_env, only :: int64
    integer(int64) :: i
    integer :: j

    This will declare a 64-bit integer variable i and a default integer
    variable j. You can then access an array A either way, for
    exmaple

    real, dimension(whatever) :: a

    a(i) = something
    a(j) = something_else


    That couses real cost under wrapv.
    In C there is no such thing as default size of array index.

    Neither is there in Fortran. There is a default integer,
    which is typically 32 bit.

    The
    programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.

    A Fortran programmer can do exactly the same; he could even import
    C_PTRDIFF_T from the ISO_C_BINDING module.

    In that respect, there is not difference between C and Fortran.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Sun Sep 27 21:23:34 2026
    From Newsgroup: comp.arch

    On Sun, 27 Sep 2026 17:39:38 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:

    Michael S <already5chosen@yahoo.com> schrieb:
    On Sun, 27 Sep 2026 13:42:05 -0000 (UTC)
    Thomas Koenig <tkoenig@netcologne.de> wrote:


    [Polyhedron benchmark, which is Fortran]

    Geometric Mean Execution Time = 11.73 seconds

    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    So at least for that particular codebase, taking away the
    freedom of the compiler to optimize based on the assumption
    that integers will never overflow, and replacing it with
    something defined, leads to a significant slowdown.



    Fortran is not the same as C.

    C and Fortran share the same rules about integer overflow.

    But if you run any program through the NAG compiler, it is
    translated to a C program.

    My understanding is that in Fortran array indexes are by default
    32-bit.

    That understanding turns out to be wrong.

    In Fortran, an array index is an integer expression, and that
    can be of any KIND that the compiler supports. You can write

    use iso_fortran_env, only :: int64
    integer(int64) :: i
    integer :: j

    This will declare a 64-bit integer variable i and a default integer
    variable j. You can then access an array A either way, for
    exmaple

    real, dimension(whatever) :: a

    a(i) = something
    a(j) = something_else


    That couses real cost under wrapv.
    In C there is no such thing as default size of array index.

    Neither is there in Fortran. There is a default integer,
    which is typically 32 bit.

    The
    programmer can define it as 32-bit and suffer the same performance degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.

    A Fortran programmer can do exactly the same; he could even import C_PTRDIFF_T from the ISO_C_BINDING module.

    In that respect, there is not difference between C and Fortran.



    So, what happens when you edit source code of Polyhedron benchmark to
    use of C_PTRDIFF_T for aray indices?
    Is there still an impact for -fwrapv ?



    of

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Sun Sep 27 10:55:29 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:

    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance
    degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.

    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Sun Sep 27 11:51:18 2026
    From Newsgroup: comp.arch

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> writes:

    On 9/24/2026 9:29 AM, Tim Rentsch wrote:

    [...]

    [...] how would you write a
    requirement in the C standard to limit the meaning of "you can do
    whatever you want, but you must document what you do" that gives only
    as much freedom as you think should be allowed?

    I have minimal experience with standards writing, and none with
    language writing, so I may be off base here.

    I think that there are many situations where the standard, correctly,
    doesn't specify anything about what the compiler should do. These are currently called undefined behavior, and the standard essentially lets
    it go at that - no further documentation required.

    But in many, but not all, of such cases, the compiler implementer
    knows what it is going to do. What I am after is that in such cases,
    the implementer "supplement" the bare words with information that he
    has that knowledge and he should communicate it to the user.


    I think we will find that different people have very different ideas
    about how much freedom should be allowed in saying what happens. It's
    a very hard problem.

    Note that I am not trying to restrict "what happens" in any way. What happens is anything that could/would happen under today's definition
    of undefined behavior. I am just trying to ask the implementer,
    whenever possible (and I realize that it is not always possible) to
    elaborate on the bare words "undefined behavior".

    I want to write a short answer here. I trimmed away anything that
    looked incidental to my reply.

    The problem you want to solve is not a problem for the C standard;
    rather it is a problem with compilers and/or compiler writers.
    Because of that the C standard is not the right place to address
    it.

    The job of the C standard is to define the C language, and not to
    address questions of "how good" C implementations are. This
    property is a conscious and explicit decision made by the authors
    of the C standard. Any matters of "quality of implementation" are
    deliberately not addressed in the standard.

    I understand that you want to address what you see as a shortcoming
    in how compilers behave. Unfortunately the C standard is not the
    right venue to accomplish that, because how individual compilers
    behave is a quality of implementation concern, and not a question
    about what is the definition of the C language.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Sun Sep 27 22:12:54 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Michael S <already5chosen@yahoo.com> writes:

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:

    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance
    degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.

    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it
    targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.


    In Rust, the only acceptable array index is of type "usize", which is effectively the same as u64 on most platforms these days, but does not
    need to be so: On a 32-bit target, it would be the same as u32.

    So effectively very similar to size_t.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Sun Sep 27 14:42:49 2026
    From Newsgroup: comp.arch

    Thomas Koenig [2026-09-27 13:42:05] wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    Given that UB has been shown to provide only fairly minor performance
    gains,
    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    FWIW, I consider a 10% performance difference to be fairly minor.

    Notice also that my SMUB proposal does allow some of the optimizations
    that UB allows w.r.t overflow (e.g. it allows the compiler to consider
    integer addition as associative) so it would not cost as much as
    `-fwrapv`.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 28 06:12:21 2026
    From Newsgroup: comp.arch

    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    Thomas Koenig [2026-09-27 13:42:05] wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    Given that UB has been shown to provide only fairly minor performance
    gains,
    The geometric mean increased by around 11%. Some tests were
    within shouting each other. Others (air, linpk, nf, tfft2) show
    significant differences.

    FWIW, I consider a 10% performance difference to be fairly minor.

    That was the geometric mean, the worst one was a factor of 1.54.

    And people fix speed regressions of an individual test case of
    2-5%, so your opinion differs a lot from that of people who write
    compilers, at least.

    Notice also that my SMUB proposal does allow some of the optimizations
    that UB allows w.r.t overflow (e.g. it allows the compiler to consider integer addition as associative) so it would not cost as much as
    `-fwrapv`.

    Do you mean associative as in a + b = b + a (which is trivially
    true) or associative as in (a + b) + c = (c + b) + a, even when
    c + b overflows?

    But your proposal, as I understand it, would destroy a lot of
    range-based optimization, so that, for example, the second if
    statement in

    if (a > 5) {
    if (a + 1 > 3) {
    }
    }

    would be executed. You may wonder who writes such code, but
    it can be created as the result of template initiation,
    constant propagation, function specialization aka cloning,
    devirtualization, link-time optimization (which allows a lot
    of the previous stuff) or any combination of the above, plus
    other optimization.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 07:57:12 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    FWIW, I consider a 10% performance difference to be fairly minor.

    That was the geometric mean, the worst one was a factor of 1.54.

    It would be interesting to know what kind of code that was, and what
    code was generated with either option.

    Wang et al. only saw a ~10% hit from -fwrapv on one of the SPEC CINT
    benchmarks (and no difference on the others), from a sign extension
    inserted in an inner loop. I find it hard to believe that additional
    sign extensions result in a factor 1.54 slowdown in an inner loop,
    much less across the whole benchmark. Did gfortran auto-vectorize
    some hot loop without -fwrapv, and but not with -fwrapv?

    Notice also that my SMUB proposal does allow some of the optimizations
    that UB allows w.r.t overflow (e.g. it allows the compiler to consider
    integer addition as associative) so it would not cost as much as
    `-fwrapv`.

    Do you mean associative as in a + b = b + a (which is trivially
    true) or associative as in (a + b) + c = (c + b) + a, even when
    c + b overflows?

    a+b = b+a is not the associative law.

    (a + b) + c = (c + b) + a is not the associative law, either.

    (a + b) + c = a + (b + c) is the associative law.

    The first is the commutative law (which also holds for addition), the
    second has no name, but if the commutative and associative laws hold
    for an operation, it holds, too.

    Concerning -fwrapv: The associative law holds for + and * in modulo arithmetics, so the compiler can perform optimizations that rely on
    the associative law when the programmer asks for modulo arithmetics
    with -fwrapv. That's not the case if overflow is trapped (-ftrapv, or
    its long-winded -fsanitize= cousin).

    You may wonder who writes such code, but
    it can be created as the result of template initiation,

    I.e., C programmers don't write such code.

    devirtualization,

    I.e., C programmers don't write such code.

    I think that Fortran programmers don't write such code, either,

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 28 10:54:32 2026
    From Newsgroup: comp.arch

    On 27/09/2026 18:43, Michael S wrote:
    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:


    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance
    degradatioon under wrapv as Fortran does. Or he can define index as
    either ptrdiff_t or size_t and suffere no degradation.


    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.


    That's not universal.

    Sure - it is implementation-dependent, usually defined by a platform ABI.

    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it targets.


    That's interesting - thanks for the correction. I will have to be more nuanced - on /many/ 64-bit targets, int_fast32_t is 64-bit. (I also
    see, from godbolt, that it is 64-bit on 64-bit ARM for Linux but again
    is 32-bit on 64-bit ARM for Windows.)

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    It is not pretty, no. And I think that is the main hinder for the use
    of the "fast" and "least" integer types. I have very rarely used them
    myself - only on code that I know was to be used on 8-bit AVRs and
    32-bit ARMs.



    The thing you want is that your loops with "a[i++]" operations are
    implemented as "*p++" rather than "*(a + i); i = (i + 1) &
    0xffff'ffff)". The later is what you end up with if you use an
    unsigned 32-bit type or a signed 32-bit type with forced wrapping
    semantics.

    So you can either use a 64-bit type (signed or unsigned), or use
    normal 32-bit "int" as usual, but without forcing wrapping semantics
    on it.





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 28 08:56:18 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:

    You may wonder who writes such code, but
    it can be created as the result of template initiation,

    I.e., C programmers don't write such code.

    Compilers share middle and back ends. Do you propose
    to turn off specific optimizations for C only?


    devirtualization,

    I.e., C programmers don't write such code.

    I think that Fortran programmers don't write such code, either,

    You mean Fortran programmers do not use features that have been
    in the language since Fortran 2003? Fortran has had type
    extension, type-bound procedures and CLASS since then.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Mon Sep 28 11:06:59 2026
    From Newsgroup: comp.arch

    On 28/09/2026 09:57, Anton Ertl wrote:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    FWIW, I consider a 10% performance difference to be fairly minor.

    For some code, 10% may be considered a major improvement. Lots of code
    is not performance-critical, but some of it is.



    Concerning -fwrapv: The associative law holds for + and * in modulo arithmetics, so the compiler can perform optimizations that rely on
    the associative law when the programmer asks for modulo arithmetics
    with -fwrapv. That's not the case if overflow is trapped (-ftrapv, or
    its long-winded -fsanitize= cousin).

    Yes, -fwrapv allows a lot more optimisation than -ftrapv.


    You may wonder who writes such code, but
    it can be created as the result of template initiation,

    I.e., C programmers don't write such code.

    devirtualization,

    I.e., C programmers don't write such code.


    C programmers use macros, which can result in such code. (Some C
    programmers use a /lot/ of very convoluted macros.)

    Code like this can also easily turn up as the result of inlining,
    function cloning, constant propagation, and other such transformations:

    int foo(int a) {
    if (a + 1 > 3) {
    ...
    }
    }

    int bar(int a) {
    if (a > 5) {
    return foo(a);
    }
    }

    C++ gives more ways to write code than C, but I don't think it changes
    the principles here.

    (I suspect that another type of optimisation that people disagree about, type-based alias analysis or "-fstrict-aliasing", is more relevant for
    C++ than C, because you typically use more different types in C++.)

    I think that Fortran programmers don't write such code, either,

    - anton

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Mon Sep 28 02:59:06 2026
    From Newsgroup: comp.arch

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Michael S <already5chosen@yahoo.com> writes:

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:

    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance
    degradatioon under wrapv as Fortran does. Or he can define index as >>>>> either ptrdiff_t or size_t and suffere no degradation.

    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it
    targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.

    In Rust, the only acceptable array index is of type "usize", which is effectively the same as u64 on most platforms these days, but does not
    need to be so: On a 32-bit target, it would be the same as u32.

    So effectively very similar to size_t.

    Maybe that limitation works okay in Rust. In C sometimes it's
    important to allow a signed value for indexing. Probably 99% of
    the time size_t is good, but not always.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Mon Sep 28 14:43:51 2026
    From Newsgroup: comp.arch

    Tim Rentsch wrote:
    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Michael S <already5chosen@yahoo.com> writes:

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:

    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same performance >>>>>> degradatioon under wrapv as Fortran does. Or he can define index as >>>>>> either ptrdiff_t or size_t and suffere no degradation.

    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit is
    enough range for a great many purposes) but can bigger if that's
    faster. On most 64-bit targets, "int_fast32_t" will be 64-bit.

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it >>>> targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness >>>> and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.

    In Rust, the only acceptable array index is of type "usize", which is
    effectively the same as u64 on most platforms these days, but does not
    need to be so: On a 32-bit target, it would be the same as u32.

    So effectively very similar to size_t.

    Maybe that limitation works okay in Rust. In C sometimes it's
    important to allow a signed value for indexing. Probably 99% of
    the time size_t is good, but not always.


    The fact that array indices are unsigned is one of the more crucial underpinnings of Rust safety. Even more importantly, all such indices
    are verified against the actual array upper boundary. When the compiler
    can prove to itself that such accesses are always safe, then the coded
    check can be omitted.

    Just like Java, a single test outside of a for loop can replace testing
    in every iteration.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:24:04 2026
    From Newsgroup: comp.arch

    Paul Clayton <paaronclayton@gmail.com> writes:
    On 9/9/26 2:41 PM, MitchAlsup wrote:

    antispam@fricas.org (Waldek Hebisch) posted:
    [snip]
    Once you have paravirtualization, there is nasty question how
    guest and host should divide the work and what is the best
    interface.

    Host OS needs an entry point to Guest OS to ask for resources back
    while allowing Guest OS to determine which resources to return.
    Guest OS needs an entry point in Host OS to ask for more resources
    while allowing Host OS to determine which resources to send.

    I wonder if it would make sense for a hardware architecture to
    define an interface to its own hypervisor (vaguely reminiscent
    to Alpha's PALcode?) to make at least a limited form of
    paravirtualization natural. (I am thinking of page tables and
    some resource management.)

    That is generally done at the API level in modern systems.

    For example, there is a defined specification on ARM systems
    for the interface between the OS and the hardware/firmware
    (e.g. PSCI).



    I have a personal dislike of nested page tables and some

    Care to elaborate? As someone who worked with hypervisors
    both before and after AMD introduced nested page tables,
    I have nothing good to say about the alternatives to
    NPT.

    )

    Nested Paging has won. Guest OS virtualizes applications and its own
    worker threads, while Host OS virtualized multiple Guest OSs.

    Nested paging potentially doubles the page table depth, which
    seems unattractive to me. (I am aware that caching intermediate
    page table nodes and using large pages can avoid much of this
    overhead. I almost certainly undervalue virtualization and my
    sense of ickiness for nested paging is not deeply informed.)

    Indeed, the table walks can be expensive; but as you point
    out, intermediate caching and using large pages
    are de riguour in modern implementations.

    The performance and security of the alternatives (e.g. XEN pre NPT)
    were significantly poorer.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:34:40 2026
    From Newsgroup: comp.arch

    Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
    Michael S <already5chosen@yahoo.com> writes:

    <int_fast_32_discussion>

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it
    targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.

    I -use- size_t for array index variables.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Mon Sep 28 01:58:53 2026
    From Newsgroup: comp.arch

    David Brown [2026-09-27 16:40:15] wrote:
    On 26/09/2026 16:25, Stefan Monnier wrote:
    Remember, we're starting from the problem that UB is used by
    compilers in ways which catch programmers off-guard.
    If that is truly the way you think, then you should maybe give up
    programming and start a conspiracy-theory pod-cast. (You'll probably make more money!) Seriously - if you start from the suggestion that compiler writers are aiming to catch programmers off-guard, or that they disregard

    I never claimed it is their aim.

    The fact that programmers are surprised by the behavior of the generated
    code (e.g. cases described as "nasal demons") just reflects the fact
    that their understanding of undefined behavior doesn't match the reality
    and my SMUB is an attempt to specify another flavor of undefined
    behavior which is hoped to:

    - Align more closely with what the average programmers expect when
    they hear "undefined behavior".
    - Still allow enough flexibility for the implementation that the impact
    on performance is usually small.

    I'm not blaming anyone. My finger is pointed only at the discrepancy
    between what UB is understood to mean by the C spec and what it is
    understood to mean by average programmers.

    [ And I don't see much hope to fix the programmers' understanding, so
    I think the only way to reduce this discrepancy is to change the
    C spec. ]


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Mon Sep 28 19:58:33 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 02:59:06 -0700
    Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:

    Michael S <already5chosen@yahoo.com> writes:

    On Sun, 27 Sep 2026 18:04:05 +0200
    David Brown <david.brown@hesbynett.no> wrote:

    On 27/09/2026 17:18, Michael S wrote:

    Fortran is not the same as C.
    My understanding is that in Fortran array indexes are by default
    32-bit. That couses real cost under wrapv.
    In C there is no such thing as default size of array index. The
    programmer can define it as 32-bit and suffer the same
    performance degradatioon under wrapv as Fortran does. Or he
    can define index as either ptrdiff_t or size_t and suffere no
    degradation.

    You certainly /can/ do that. But a lot of C code uses "int" for
    array indexing. The "ideal" type in many cases would be
    "int_fast32_t" to say it should be of a big enough size (32-bit
    is enough range for a great many purposes) but can bigger if
    that's faster. On most 64-bit targets, "int_fast32_t" will be
    64-bit.

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what
    OS it targets.

    Besides, int_fast32_t is both above my threshold of acceptable
    ugliness and acceptable geekery.

    I favor a type equivalent to size_t for use as array index
    variables.

    In Rust, the only acceptable array index is of type "usize", which
    is effectively the same as u64 on most platforms these days, but
    does not need to be so: On a 32-bit target, it would be the same
    as u32.

    So effectively very similar to size_t.

    Maybe that limitation works okay in Rust. In C sometimes it's
    important to allow a signed value for indexing. Probably 99% of
    the time size_t is good, but not always.


    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Mon Sep 28 14:32:58 2026
    From Newsgroup: comp.arch

    Thomas Koenig [2026-09-28 06:12:21] wrote:
    Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
    Notice also that my SMUB proposal does allow some of the optimizations
    that UB allows w.r.t overflow (e.g. it allows the compiler to consider
    integer addition as associative) so it would not cost as much as
    `-fwrapv`.

    Do you mean associative as in a + b = b + a (which is trivially
    true) or associative as in (a + b) + c = (c + b) + a, even when
    c + b overflows?

    [ Anton replied to this part. ]

    But your proposal, as I understand it, would destroy a lot of
    range-based optimization, so that, for example, the second if
    statement in

    if (a > 5) {
    if (a + 1 > 3) {
    }
    }

    would be executed.

    I think there would be related cases where SMUB would disallow "similar" optimizations (e.g. the `a + b > INTMAX` case), but not for the above
    one: SMUB would say that the compiler is free to make `a+1` return any
    integer it likes when the computation overflows, so the inner test can
    be skipped.

    You may wonder who writes such code, but

    I don't: it's common to see weird/useless code after macro-expansion.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.arch on Tue Sep 29 10:16:58 2026
    From Newsgroup: comp.arch

    On 28/09/2026 07:58, Stefan Monnier wrote:
    David Brown [2026-09-27 16:40:15] wrote:
    On 26/09/2026 16:25, Stefan Monnier wrote:
    Remember, we're starting from the problem that UB is used by
    compilers in ways which catch programmers off-guard.
    If that is truly the way you think, then you should maybe give up
    programming and start a conspiracy-theory pod-cast. (You'll probably make >> more money!) Seriously - if you start from the suggestion that compiler
    writers are aiming to catch programmers off-guard, or that they disregard

    I never claimed it is their aim.

    Okay - I apologise for reading too much into your wording there.


    The fact that programmers are surprised by the behavior of the generated
    code (e.g. cases described as "nasal demons") just reflects the fact
    that their understanding of undefined behavior doesn't match the reality

    True. But the root issue is that some programmers don't understand what
    UB is, or do not know that a particular construct is, or may have, UB. Undefined behaviour itself is not the problem, nor is compiler handling
    of it.

    This is why I am not convinced that some new types of sort-of UB will be helpful. They could only be useful in a few specific cases, and could
    easily add to the confusion.

    The focus, as I see it, should be on helping programmers understand UB,
    and helping them avoid it in their code. The part to be played by
    toolchains here is to be better at spotting potential UB and warning
    about it. (And toolchains are, on the whole, getting better at this.)

    There are also at least some particular cases where the C standards
    could be changed to give at least some definition to undefined
    behaviours, but not in many cases.


    and my SMUB is an attempt to specify another flavor of undefined
    behavior which is hoped to:

    - Align more closely with what the average programmers expect when
    they hear "undefined behavior".

    I don't believe in the "average" programmer, or that there is any
    consensus about what they might expect. And I don't think a language
    like C, or its tools, should be aiming to cater for programmers who want
    a more managed or friendly language. Any solution which reduces the efficiency of generated code should be rejected.

    And any solution which says it is okay for C programmers to be lax
    should be rejected. Today, we tell C programmers "Do not let your
    signed integer arithmetic overflow - it will cause daemons to fly out of
    your nose! This is a rule of C, and you must obey it". It is simple
    and clear, and for the most part, not difficult to achieve. It is an unfortunate truth that not all C programmers know that rule, and that
    even experts occasionally make mistakes, but we live in an imperfect world.

    Changing this to "You should probably not let your signed integer
    arithmetic overflow. It's unlikely to have very bad consequences, but
    it will sometimes give the wrong answer, or perhaps stop your program. However, dereferencing an invalid pointer will still launch nasal
    daemons" is /not/ helpful.

    - Still allow enough flexibility for the implementation that the impact
    on performance is usually small.

    I'm not blaming anyone. My finger is pointed only at the discrepancy
    between what UB is understood to mean by the C spec and what it is
    understood to mean by average programmers.

    [ And I don't see much hope to fix the programmers' understanding, so
    I think the only way to reduce this discrepancy is to change the
    C spec. ]


    === Stefan

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 29 16:53:49 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> schrieb:

    So, what happens when you edit source code of Polyhedron benchmark to
    use of C_PTRDIFF_T for aray indices?

    I (would, and did) get a ton of syntax errors.

    Fortran does not do argument promotion, so (assuming the
    right intrinsic module has been imported

    subroutine foo(i)
    integer(kind=c_ptrdiff_t), intent(in) :: i
    ...

    and somewhere where foo is visible.

    call foo(1)

    is invalid when c_ptrdiff_t is not equal to the default integer
    kind, and compilers catch that in a module. The calling side
    should then read

    call foo(1_c_ptrdiff_t)

    so presumebly an abbreviation would be used.

    Rewriting a Fortran program for larger integers is, unfortunately,
    a major undertaking, best done with automated tools.

    I can use the option of ill repute, -fdefault-integer-8, which
    causes standard violations. Interestingly, that also makes one
    benchmark fail.

    Is there still an impact for -fwrapv ?

    From the numbers I've looked at with -fdefault-integer-8, not
    a big one. But like I wrote above, this is an option of ill
    repute.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 29 12:42:10 2026
    From Newsgroup: comp.arch

    scott@slp53.sl.home (Scott Lurndal) writes:

    Tim Rentsch <tr.17687@z991.linuxsc.com> writes:

    Michael S <already5chosen@yahoo.com> writes:

    <int_fast_32_discussion>

    That's not universal.
    For example, on both alive 64-bit Windows targets int_fast32_t is
    32-bit.
    Godbolt shows the same for clang for MIPS64. I have no idea what OS it
    targets.

    Besides, int_fast32_t is both above my threshold of acceptable ugliness
    and acceptable geekery.

    I favor a type equivalent to size_t for use as array index variables.

    I -use- size_t for array index variables.

    I generally do not use the type name size_t for indicating anything
    other than values meant to hold a size in bytes, such as would be
    produced by the sizeof operator.

    I do use the type size_t for all sorts of other different things,
    but normally with a different name given by using a typedef. An
    example is the type name Index for variables meant to hold array
    index values. I find using such more specific names, rather than
    using size_t everywhere, generally makes code easier to read and
    understand.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Tim Rentsch@tr.17687@z991.linuxsc.com to comp.arch on Tue Sep 29 15:44:56 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:

    On Mon, 28 Sep 2026 02:59:06 -0700
    Tim Rentsch <tr.17687@z991.linuxsc.com> wrote:

    Terje Mathisen <terje.mathisen@tmsw.no> writes:

    Tim Rentsch wrote:
    [...]
    I favor a type equivalent to size_t for use as array index
    variables.

    In Rust, the only acceptable array index is of type "usize", which
    is effectively the same as u64 on most platforms these days, but
    does not need to be so: On a 32-bit target, it would be the same
    as u32.

    So effectively very similar to size_t.

    Maybe that limitation works okay in Rust. In C sometimes it's
    important to allow a signed value for indexing. Probably 99% of
    the time size_t is good, but not always.

    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.

    That tells me a lot. Thank you.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From cross@cross@spitfire.i.gajendra.net (Dan Cross) to comp.arch on Wed Sep 30 13:57:24 2026
    From Newsgroup: comp.arch

    In article <20260928195833.00001d42@yahoo.com>,
    Michael S <already5chosen@yahoo.com> wrote:
    [snip]
    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.

    Beginning Rust programmers tend to go through a few phases,
    particularly if coming from a language like C. "Fighting the
    Borrow Checker" is well-known; "traits are awesome; let's use
    them everywhere!" is another.

    When I started programming in Rust, most of my immediate
    colleagues and I had terribly little experience in the language.
    We were excited about the its potential benefits, but realized
    it would take a few months to build up a sufficient level of
    familiarity to use it skillfully. One quipped, "this language
    has a near-vertical learning curve."

    At one point a coworker (self-deprecatingly) said, "look, I'm
    writing crust! C in Rust syntax!" Everyone thought this was a
    great play on words, but he was correct: the programs we were
    writing were essentially C programs, just expressed in Rust. We
    quickly adopted that to describe our still naive code: it wasn't
    an insult, just an expression that we were still coming to terms
    with a new way of doing things. Some of the hallmarks of that
    included using the wrong paradigms (and types!) for lots of our
    code.

    I'd be curious to know more about your program. I've written
    several hundred kloc of Rust at this point, and while what I
    write may be terribly specialized, I rarely find myself indexing
    an array; often, I'm using iterators or some other abstraction
    for that kind of thing. If that was the one program you have
    written in Rust, my _guess_ is that you were using paradigms
    appropriate to a different language, which is why you felt you
    needed to reach for casts to index an array so frequently.

    Again, that's not a comment on your program, so much as
    suggesting an impedence mismatch between what you wrote and how
    one would generally expect to use the language.

    In general, I think a person's experience writing a single
    program tells us very little.

    - Dan C.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Wed Sep 30 17:50:06 2026
    From Newsgroup: comp.arch

    On Wed, 30 Sep 2026 13:57:24 -0000 (UTC)
    cross@spitfire.i.gajendra.net (Dan Cross) wrote:

    In article <20260928195833.00001d42@yahoo.com>,
    Michael S <already5chosen@yahoo.com> wrote:
    [snip]
    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.

    Beginning Rust programmers tend to go through a few phases,
    particularly if coming from a language like C. "Fighting the
    Borrow Checker" is well-known; "traits are awesome; let's use
    them everywhere!" is another.


    I didn't reach the second stage. Not sure about first.

    When I started programming in Rust, most of my immediate
    colleagues and I had terribly little experience in the language.
    We were excited about the its potential benefits, but realized
    it would take a few months to build up a sufficient level of
    familiarity to use it skillfully. One quipped, "this language
    has a near-vertical learning curve."


    IMHO, he is correct.
    Rust is certainly the hardest to learn out of computer languges that I
    ever tried that I would not call idiotic.

    At one point a coworker (self-deprecatingly) said, "look, I'm
    writing crust! C in Rust syntax!" Everyone thought this was a
    great play on words, but he was correct: the programs we were
    writing were essentially C programs, just expressed in Rust. We
    quickly adopted that to describe our still naive code: it wasn't
    an insult, just an expression that we were still coming to terms
    with a new way of doing things. Some of the hallmarks of that
    included using the wrong paradigms (and types!) for lots of our
    code.

    I'd be curious to know more about your program.

    The program was an implementation scalable flood fill algorithm (i.e.
    one with much better than typical worst case BigO properties, both for
    speed and for space) together with test benches for correctness and
    speed. As you could guess, I didn't really wrote it in Rust. It was a non-literal port from C. The purpose of translation was two-fold:
    improvement of my knowledge of Rust and getting better feel for cost of mandatory bound checking for one particular sort of data structures.
    Before doing Rust port I did port to Golang. Unlike for Rust, where I
    only read The Book and went trough part of tutorial, I had small
    previous practical experience in Go, but it was not very recent. Still,
    the difference in my productivity was huge. In favor of Go, of course.

    I've written
    several hundred kloc of Rust at this point, and while what I
    write may be terribly specialized, I rarely find myself indexing
    an array; often, I'm using iterators or some other abstraction
    for that kind of thing. If that was the one program you have
    written in Rust, my _guess_ is that you were using paradigms
    appropriate to a different language, which is why you felt you
    needed to reach for casts to index an array so frequently.

    Again, that's not a comment on your program, so much as
    suggesting an impedence mismatch between what you wrote and how
    one would generally expect to use the language.

    In general, I think a person's experience writing a single
    program tells us very little.


    I never pretended otherwise. Although I don't see how one can avoid
    cast if one wants to use indices that originate in signed domain. Or,
    as was often a case in this particular program, in domain of narrower
    unsigned integers.

    - Dan C.



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Wed Sep 30 18:03:46 2026
    From Newsgroup: comp.arch

    Dan Cross wrote:
    In article <20260928195833.00001d42@yahoo.com>,
    Michael S <already5chosen@yahoo.com> wrote:
    [snip]
    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.

    Beginning Rust programmers tend to go through a few phases,
    particularly if coming from a language like C. "Fighting the
    Borrow Checker" is well-known; "traits are awesome; let's use
    them everywhere!" is another.

    When I started programming in Rust, most of my immediate
    colleagues and I had terribly little experience in the language.
    We were excited about the its potential benefits, but realized
    it would take a few months to build up a sufficient level of
    familiarity to use it skillfully. One quipped, "this language
    has a near-vertical learning curve."

    At one point a coworker (self-deprecatingly) said, "look, I'm
    writing crust! C in Rust syntax!" Everyone thought this was a
    great play on words, but he was correct: the programs we were
    writing were essentially C programs, just expressed in Rust. We
    quickly adopted that to describe our still naive code: it wasn't
    an insult, just an expression that we were still coming to terms
    with a new way of doing things. Some of the hallmarks of that
    included using the wrong paradigms (and types!) for lots of our
    code.

    I have definitely noted this in my own code, crust is a good name for it.

    Quite often, when "fighting the borrow checker", typically because I'm updating a complicated data structure in two locations at the same time,
    I'll fall back to "write Fortran in Rust", with every access indexed.

    This is BTW one of those problems the brand new upcoming borrow checker
    is supposed to be far better at figuring out: If I'm updating totally
    separate parts of the same larger structure, then it should be OK to
    mutate both at the same time.

    I do realize that there are far more Rusty ways to solve this type of
    issue, they just don't come naturally to me yet.

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 30 17:49:34 2026
    From Newsgroup: comp.arch


    Terje Mathisen <terje.mathisen@tmsw.no> posted:

    Dan Cross wrote:
    In article <20260928195833.00001d42@yahoo.com>,
    Michael S <already5chosen@yahoo.com> wrote:
    [snip]
    My only Rust program was full of casts of various integer types to
    usize. Exactly for the purpose of array indexing.

    Beginning Rust programmers tend to go through a few phases,
    particularly if coming from a language like C. "Fighting the
    Borrow Checker" is well-known; "traits are awesome; let's use
    them everywhere!" is another.

    When I started programming in Rust, most of my immediate
    colleagues and I had terribly little experience in the language.
    We were excited about the its potential benefits, but realized
    it would take a few months to build up a sufficient level of
    familiarity to use it skillfully. One quipped, "this language
    has a near-vertical learning curve."

    At one point a coworker (self-deprecatingly) said, "look, I'm
    writing crust! C in Rust syntax!" Everyone thought this was a
    great play on words, but he was correct: the programs we were
    writing were essentially C programs, just expressed in Rust. We
    quickly adopted that to describe our still naive code: it wasn't
    an insult, just an expression that we were still coming to terms
    with a new way of doing things. Some of the hallmarks of that
    included using the wrong paradigms (and types!) for lots of our
    code.

    I have definitely noted this in my own code, crust is a good name for it.

    Quite often, when "fighting the borrow checker", typically because I'm updating a complicated data structure in two locations at the same time, I'll fall back to "write Fortran in Rust", with every access indexed.

    Wouldn't that be FRustrating ?!?

    This is BTW one of those problems the brand new upcoming borrow checker
    is supposed to be far better at figuring out: If I'm updating totally separate parts of the same larger structure, then it should be OK to
    mutate both at the same time.

    I do realize that there are far more Rusty ways to solve this type of
    issue, they just don't come naturally to me yet.

    Terje

    --- Synchronet 3.22a-Linux NewsLink 1.2