• Microarchitectural effects on performance (was: Optimization ...)

    From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 21 06:55:56 2026
    From Newsgroup: comp.arch

    John Levine <johnl@taugh.com> writes:
    According to Stefan Monnier <monnier@iro.umontreal.ca>:
    Indeed, in practice virtually all compiler "optimizations" are valid
    only statistically: in most cases they either have no measurable effect
    or they improve some characteristic, but there are almost always corner >>cases where they make things worse.

    There are plenty of optimizations that always make things better, e.g., >removing dead code, or reusing values in registers.

    For certain values of "better", such as code size. Unfortunately, for optimizing execution time, there are microarchitectural effects in
    play that can result in paradoxical effects even for code that pretty
    much everybody woul consider better without information to the
    contrary. E.g., on a DecStation I saw a 20% slowdown by removing a
    piece of code that had become unnecessary earlier, in a benchmark that
    actually did not exercise this code. My explanation for that was
    I-cache conflicts.

    Many of these effects (such as I-cache conflicts) are not directly
    connected to the transformation at hand. Some may be fixable by
    further transformations; e.g., if branch targets shortly before a
    cache-line boundary are bad for performance, one can insert padding
    before the branch target.

    OTOH, the store-to-wide-load forwarding issue that I have discussed in
    this thread is directly connected to the auto-vectorization of several
    narrow loads into a wide load. So in this case the transformation has
    direct connection to the microarchitectural pitfall.

    I agree that the more sophisticated the optiomization, the more likely it
    is to have perverse cases.

    In the case of auto-vectorizing loads, I have seen the compiler
    generate code that results in a significantly increased number of
    executed instructions (but reduced executed loads); I wonder how that
    code performs in those cases where the microarchitectural pitfall does
    not strike; maybe I will measure it later. But the more serious issue
    is the microarchitectural pitfall, which does not always strike, so
    the people who enabled this transformation by default may be
    blissfully unaware of it.

    Microarchitectural pitfalls can be unrelated to the sophistication of
    the compiler. E.g., I discuss [ertl24] a number of microarchitectural
    pitfalls that many Forth systems fall into, not because of
    sophistication, but because the Forth systems use techniques that
    worked fine in the microprocessors of the 1970s and 1980s, and
    actually still work fine on threaded-code systems on modern CPUs, but
    fall into these pitfalls in native-code systems. Ok, in a way it
    affects more sophisticated systems more, but not because of the
    sophistication of their optimizations; systems with more sophisticated optimizations see the same issues as less sophisticated ones, that are
    still native-code systems.

    @InProceedings{ertl24,
    author = {M. Anton Ertl},
    title = {How to Implement Words (Efficiently)},
    crossref = {euroforth24},
    pages = {43--52},
    url = {http://www.euroforth.org/ef24/papers/ertl.pdf},
    url-slides = {http://www.euroforth.org/ef24/papers/ertl-slides.pdf},
    video = {https://www.youtube.com/watch?v=bAq4760h5ZQ},
    OPTnote = {not refereed},
    abstract = {The implementation of Forth words has to satisfy the
    following requirements: 1) A word must be
    represented by a single cell (for
    \code{execute}). 2) A word may represent a
    combination of code and data (for, e.g.,
    \code{does>}). In addition, on some hardware,
    keeping executed native code and (written) data
    close together results in slowness and therefore
    should be avoided; moreover, failing to pair up
    calls with returns results in (slow) branch
    mispredictions. The present work describes how
    various Forth systems over the decades have
    satisfied the requirements, and how many systems run
    into performance pitfalls in various situations.
    This paper also discusses how to avoid this
    slowness, including in native-code systems.}
    }
    @Proceedings{euroforth24,
    title = {40th EuroForth Conference},
    booktitle = {40th EuroForth Conference},
    year = {2024},
    key = {EuroForth'24},
    url = {http://www.euroforth.org/ef24/papers/proceedings.pdf}
    }

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 21 17:04:27 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    John Levine <johnl@taugh.com> writes:
    -------------------

    OTOH, the store-to-wide-load forwarding issue that I have discussed in
    this thread is directly connected to the auto-vectorization of several
    narrow loads into a wide load. So in this case the transformation has
    direct connection to the microarchitectural pitfall.

    I think we need to make a distinction between vectorization and
    SIMDization. Vectors are like CRAY (and RISC-v) while SIMD is
    most everyone else.

    In modern machines LDs are well pipelined, often multi-lane; so,
    when a compiler auto-SIMDs a gaggle of LDs into a LDSIMD, one
    then needs a similar means to extract the gaggle from the SIMD.
    This extract is less likely to have the same properties of the
    LD pipeline.

    I agree that the more sophisticated the optiomization, the more likely it >is to have perverse cases.

    In the case of auto-vectorizing loads, I have seen the compiler
    generate code that results in a significantly increased number of
    executed instructions (but reduced executed loads); I wonder how that
    code performs in those cases where the microarchitectural pitfall does
    not strike; maybe I will measure it later. But the more serious issue
    is the microarchitectural pitfall, which does not always strike, so
    the people who enabled this transformation by default may be
    blissfully unaware of it.

    Microarchitectural pitfalls can be unrelated to the sophistication of
    the compiler.

    In 1992 I was working on a machine with a packet-cache organization
    of instructions. The compiler we had was unrolling, loop inducing
    and loop strength reducing generated code. This often-added instructions
    into the loop body and the loop no longer fit in n packets (now n+1) packets--resulting in a slowdown of n/(n+1). The fix was to turn those
    kinds of optimizations back off. {{A packet was a trace with pre-routed instructions}} SPECint gains about 10% on this -|Architecture by turning
    those optimizations off.

    E.g., I discuss [ertl24] a number of microarchitectural
    pitfalls that many Forth systems fall into, not because of
    sophistication, but because the Forth systems use techniques that
    worked fine in the microprocessors of the 1970s and 1980s, and
    actually still work fine on threaded-code systems on modern CPUs, but
    fall into these pitfalls in native-code systems.

    Is this related to the over aggressive branch predictors of modern ?
    Either branches themselves or indirect calls.

    Read the paper--yes a nasty wicket when the container containing
    a return address is used for something other than a return. Safe-
    Stack would prevent that use (RA is neither LD-able nor ST-able.)

    Ok, in a way it
    affects more sophisticated systems more, but not because of the sophistication of their optimizations; systems with more sophisticated optimizations see the same issues as less sophisticated ones, that are
    still native-code systems.

    @InProceedings{ertl24,
    author = {M. Anton Ertl},
    title = {How to Implement Words (Efficiently)},
    crossref = {euroforth24},
    pages = {43--52},
    url = {http://www.euroforth.org/ef24/papers/ertl.pdf},
    url-slides = {http://www.euroforth.org/ef24/papers/ertl-slides.pdf},
    video = {https://www.youtube.com/watch?v=bAq4760h5ZQ},
    OPTnote = {not refereed},
    abstract = {The implementation of Forth words has to satisfy the
    following requirements: 1) A word must be
    represented by a single cell (for
    \code{execute}). 2) A word may represent a
    combination of code and data (for, e.g.,
    \code{does>}). In addition, on some hardware,
    keeping executed native code and (written) data
    close together results in slowness and therefore
    should be avoided; moreover, failing to pair up
    calls with returns results in (slow) branch
    mispredictions. The present work describes how
    various Forth systems over the decades have
    satisfied the requirements, and how many systems run
    into performance pitfalls in various situations.
    This paper also discusses how to avoid this
    slowness, including in native-code systems.}
    }
    @Proceedings{euroforth24,
    title = {40th EuroForth Conference},
    booktitle = {40th EuroForth Conference},
    year = {2024},
    key = {EuroForth'24},
    url = {http://www.euroforth.org/ef24/papers/proceedings.pdf}
    }

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 23 05:39:56 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    John Levine <johnl@taugh.com> writes:
    -------------------

    OTOH, the store-to-wide-load forwarding issue that I have discussed in
    this thread is directly connected to the auto-vectorization of several
    narrow loads into a wide load. So in this case the transformation has
    direct connection to the microarchitectural pitfall.

    I think we need to make a distinction between vectorization and
    SIMDization. Vectors are like CRAY (and RISC-v) while SIMD is
    most everyone else.

    Architecturally, vectors are an instance of SIMD. Is there any other
    instance?

    In a compiler, the conversion of scalar code to code that uses SIMD instructions is called vectorization.

    What you are thinking about is microarchitecture. Does the SIMD
    instruction work on every scalar separately, or does it work on
    several scalars in parallel.

    One interesting aspect is that early SSE2 implementations worked 8
    bytes at a time, so if the scalars were binary64 floats, you would
    classify these implementations as "vectors", while if the scalars were binary32, you would classify them as "SIMD".

    The bubble-sort benchmark sorts 4-byte integers. With
    auto-vectorization, the program loads two adjacent integers with one
    8-byte load; if they are reordered, it stores two integers with an
    8-byte store. The microarchitectural pitfall occurs on a load after
    such a store, and it slows the program down, but the program behaves
    correctly, so the compiler works correctly and the architecture is
    implemented correctly.

    Maybe some future microarchitectures will include a way to avoid this
    pitfall: E.g., it could notice that a given load sees relatively many
    overlaps, and have a mechanism that splits it into sub-loads and merge
    the results, basically undoing the vectorization. Would you classify
    such a microarchitecture as vector or as SIMD?

    One architectural difference between the Cray-1 and modern
    architectures is that the Cray-1 only has 8-byte data, whether shorts,
    ints, or floats (not sure about doubles); and I think that's even the
    case for chars in C. So the issue of dealing with 4-byte or 1-byte
    lanes does not pose itself in microarchitectures for this
    architecture.

    It does pose itself for RISC-V. I have not looked at RISC-V vectors
    and their implementations, but somehow I doubt that all of them will
    perform string loads one byte at a time.

    Reading up about the NEC SX-Aurora <https://en.wikipedia.org/wiki/SX-Aurora_TSUBASA>, the spiritual
    successor of the Cray-1, it has 64 logical vector registers with
    256x64 bits length. It's implementation uses "32-fold parallel SIMD
    units" (obviously using SIMD in your sense, not the architectural
    sense; I would rather call it "32x64-bit-wide functional units"), and
    with three FMA units this results in up to 192 DP FP ops per cycle. I
    expect that it has a similar performance problem if a load partially
    overlaps a store, especially if the distance is not a multiple of the
    FU width.

    What would a 2018-vintag Cray-1 successor look like. Probably much
    like the NEC SX-Aurora, i.e., it would also have this problem.

    Given that we have not seen an SX-Aurora successor in 8 years, despite
    new types of HBM appearing and the demand for SIMD (with small to very
    small components) being huge, it seems like NEC has finished their SX
    line with the Aurora.

    In modern machines LDs are well pipelined, often multi-lane; so,
    when a compiler auto-SIMDs a gaggle of LDs into a LDSIMD, one
    then needs a similar means to extract the gaggle from the SIMD.
    This extract is less likely to have the same properties of the
    LD pipeline.

    Cartainly even the load-only version of bubble-sort performs slower on
    Rocket Lake <2026Sep21.171615@mips.complang.tuwien.ac.at>. OTOH, microarchitectures tend to have fewer load units than ALU units, so
    even if the scalar values are needed, one probably can construct a
    cases where a wide load followed by extraction and processing is
    faster than two narrow loads followed by processing. In general,
    though, it probably is only beneficial if the data then does not need
    to be extracted, but the further processing can also happen in a SIMD
    way.

    E.g., I discuss [ertl24] a number of microarchitectural
    pitfalls that many Forth systems fall into, not because of
    sophistication, but because the Forth systems use techniques that
    worked fine in the microprocessors of the 1970s and 1980s, and
    actually still work fine on threaded-code systems on modern CPUs, but
    fall into these pitfalls in native-code systems.

    Is this related to the over aggressive branch predictors of modern ?
    Either branches themselves or indirect calls.

    One of the pitfalls is related to branch prediction: Using a call
    instruction and then not returning, but popping the return address off
    the stack to get the position of the call instruction as data; this
    results in the hardware return stack getting out-of-sync and results
    in all returns at the current level and further out to be
    mispredicted.

    What do you mean by "over aggressive branch prediction"? Either a
    branch predictor predicts correctly or it mispredicts. In what case
    would it be "over aggressive"?

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 23 18:34:31 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    John Levine <johnl@taugh.com> writes:
    -------------------

    OTOH, the store-to-wide-load forwarding issue that I have discussed in
    this thread is directly connected to the auto-vectorization of several
    narrow loads into a wide load. So in this case the transformation has
    direct connection to the microarchitectural pitfall.

    I think we need to make a distinction between vectorization and >SIMDization. Vectors are like CRAY (and RISC-v) while SIMD is
    most everyone else.

    Architecturally, vectors are an instance of SIMD. Is there any other instance?

    One might consider vVM somewhere between CRAY and SIMD.

    In a compiler, the conversion of scalar code to code that uses SIMD instructions is called vectorization.

    Compilers have to prove various things in order to (your term)
    vectorize. These profs work for CRAY and SIMD. These proofs are
    NOT NEEDED for vVM.

    What you are thinking about is microarchitecture. Does the SIMD
    instruction work on every scalar separately, or does it work on
    several scalars in parallel.

    Later CRAY (and NEC) build multiple lanes of calculations, so that,
    every clock one could perform 2^(small n) calculations of memory
    references. Thus SIMD is the second dimension of executing lots of
    instruction instances per cycle.

    One interesting aspect is that early SSE2 implementations worked 8
    bytes at a time, so if the scalars were binary64 floats, you would
    classify these implementations as "vectors", while if the scalars were binary32, you would classify them as "SIMD".

    The bubble-sort benchmark sorts 4-byte integers. With
    auto-vectorization, the program loads two adjacent integers with one
    8-byte load; if they are reordered, it stores two integers with an
    8-byte store. The microarchitectural pitfall occurs on a load after
    such a store, and it slows the program down, but the program behaves correctly, so the compiler works correctly and the architecture is implemented correctly.

    Maybe some future microarchitectures will include a way to avoid this pitfall: E.g., it could notice that a given load sees relatively many overlaps, and have a mechanism that splits it into sub-loads and merge
    the results, basically undoing the vectorization. Would you classify
    such a microarchitecture as vector or as SIMD?

    deSIMD since you are taking what is in SIMD form and decomposing it
    to perform better on "this implementation".

    One architectural difference between the Cray-1 and modern
    architectures is that the Cray-1 only has 8-byte data,

    Yes to CRAY-1 and CRAY-1S, no to CRAY-XMP/YMP--at least at the
    ports to memory. CRAY-1 had 24-bit base/bounds memory space.

    whether shorts,
    ints, or floats (not sure about doubles); and I think that's even the
    case for chars in C. So the issue of dealing with 4-byte or 1-byte
    lanes does not pose itself in microarchitectures for this
    architecture.

    Lee Higbe told me (1982) that CRAY-1 vectorized symbol table lookup.

    It does pose itself for RISC-V. I have not looked at RISC-V vectors
    and their implementations, but somehow I doubt that all of them will
    perform string loads one byte at a time.

    My 66000 vVM accesses memory at cache-width, and then parcels out
    data to calculations from a high ported set of buffers. So, a GBOoO
    would access 2 cache lines per cycle and then parcel out the data
    at whatever width is useful to the calculations needed done.

    Reading up about the NEC SX-Aurora <https://en.wikipedia.org/wiki/SX-Aurora_TSUBASA>, the spiritual
    successor of the Cray-1, it has 64 logical vector registers with
    256x64 bits length. It's implementation uses "32-fold parallel SIMD
    units" (obviously using SIMD in your sense, not the architectural
    sense; I would rather call it "32x64-bit-wide functional units"), and
    with three FMA units this results in up to 192 DP FP ops per cycle. I
    expect that it has a similar performance problem if a load partially
    overlaps a store, especially if the distance is not a multiple of the
    FU width.

    What would a 2018-vintag Cray-1 successor look like. Probably much
    like the NEC SX-Aurora, i.e., it would also have this problem.

    Vectors died as a concept because one could build way more on-die
    calculation performance than one could feed with the number of
    memory pins one could surround that die with.

    Given that we have not seen an SX-Aurora successor in 8 years, despite
    new types of HBM appearing and the demand for SIMD (with small to very
    small components) being huge, it seems like NEC has finished their SX
    line with the Aurora.

    HBM does not solve the pin problem (above) but it does ameliorate
    a factor of 4|u of it.

    In modern machines LDs are well pipelined, often multi-lane; so,
    when a compiler auto-SIMDs a gaggle of LDs into a LDSIMD, one
    then needs a similar means to extract the gaggle from the SIMD.
    This extract is less likely to have the same properties of the
    LD pipeline.

    Cartainly even the load-only version of bubble-sort performs slower on
    Rocket Lake <2026Sep21.171615@mips.complang.tuwien.ac.at>. OTOH, microarchitectures tend to have fewer load units than ALU units, so
    even if the scalar values are needed, one probably can construct a
    cases where a wide load followed by extraction and processing is
    faster than two narrow loads followed by processing. In general,
    though, it probably is only beneficial if the data then does not need
    to be extracted, but the further processing can also happen in a SIMD
    way.

    E.g., I discuss [ertl24] a number of microarchitectural
    pitfalls that many Forth systems fall into, not because of
    sophistication, but because the Forth systems use techniques that
    worked fine in the microprocessors of the 1970s and 1980s, and
    actually still work fine on threaded-code systems on modern CPUs, but
    fall into these pitfalls in native-code systems.

    Is this related to the over aggressive branch predictors of modern ?
    Either branches themselves or indirect calls.

    One of the pitfalls is related to branch prediction: Using a call
    instruction and then not returning, but popping the return address off
    the stack to get the position of the call instruction as data; this
    results in the hardware return stack getting out-of-sync and results
    in all returns at the current level and further out to be
    mispredicted.

    What do you mean by "over aggressive branch prediction"? Either a
    branch predictor predicts correctly or it mispredicts. In what case
    would it be "over aggressive"?

    There are cases where not predicting is faster than prediction with
    low success rate. The dispatch point of an interpreter/simulator
    is one.

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Wed Sep 23 13:53:55 2026
    From Newsgroup: comp.arch

    I think we need to make a distinction between vectorization and >>SIMDization. Vectors are like CRAY (and RISC-v) while SIMD is
    most everyone else.
    Architecturally, vectors are an instance of SIMD. Is there any other instance?

    My understanding is that "SIMD" can be used to refer to the family of techniques that includes Cray-style vectors, AVX-style small vectors, Connection Machines, and GPU warps.

    But recently, I've heard the term used "SIMD" used to refer specifically
    to AVX-style small vectors, and "SIMT" for GPU warps.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2